Is our data quality good enough for AI?
For most first AI applications in mid-sized companies it is, because those applications never touch your own data. Where data quality genuinely becomes a precondition, which five error types actually matter, and how to test it in a day.
For the first stage it almost always is, because text assistance, summarising and translation work on the material the person feeds in themselves. Data quality only becomes a precondition once an AI application reaches into your file shares, your ERP or your CRM on its own. That boundary is exactly where a tool question turns into a data project.
What you get from this article:
- The data quality question is often asked too early and then blocks a start that never needed it.
- There is a clear boundary: AI with supplied data versus AI with access to your systems.
- Five error types actually matter. Cosmetic flaws in master data are not among them.
- The most expensive error is not the wrong figure, it is the double truth held in two places.
- A meaningful test takes one day and needs no software.
Is our data quality good enough for AI?
For applications where a person supplies the material themselves, yes. For applications that reach into your systems on their own and generate answers from them, data quality decides whether the result is usable. So the answer does not depend on the state of your data but on how the planned application is built.
That distinction is the actual core of the question, and consulting conversations skip it routinely. If an employee pastes a contract text into an approved tool and gets a summary, the quality of your master data is irrelevant to it. If instead an assistant is meant to answer from the ERP how much revenue a customer generated, then whether that customer appears once or three times in the system decides the answer.
What follows from this: a company with messy data can start with AI immediately. It just cannot immediately start with the kind of AI that reaches into that data.
Where exactly does data quality become a precondition?
At four points: assistants with access to document repositories, analyses and forecasts built from business data, automated processes that create or change records, and any application whose output goes outside without a person checking it.
The most common case in mid-sized companies is the first. An assistant meant to answer questions from your own file inventory is attractive, and turns disappointing fast when the inventory holds four versions of the same price list and none of them is marked as current. The tool made no mistake there, it picked one of four truths.
The most expensive case is the fourth. As soon as a result goes outside unchecked, a data error becomes an incident at the customer. For that scope one simple rule applies, and it needs no technology: as long as the data is unverified, a person stays in the chain.
Which five error types actually matter?
Duplicates, missing mandatory entries in fields that are meant to be analysed, inconsistent spellings of names and labels, outdated records with no marker, and the same entry held in two places at different states.
The last is the worst and the least often named. If the customer address differs between ERP and CRM, the organisation does not have a data problem but a responsibility problem. No tool solves that, because the question of which source governs is a decision, not a calculation.
Not on this list: inconsistent formatting, missing fields nobody analyses, and historical records clearly recognisable as closed. These things look untidy and do no harm. Cleaning them first spends the budget in the wrong place.
| Error type | Severity | Effort |
|---|---|---|
| Same entry in two places, different states | high | decision first, then technology |
| Duplicates in customer or article data | high | medium |
| Missing mandatory entries in analysed fields | medium | medium |
| Inconsistent spellings | medium | low, automates well |
| Outdated records with no marker | medium | low |
| Formatting, unused fields, clearly closed history | none | leave alone |
How do we test this without software and without a project?
Take the one question the planned application is meant to answer, ask it manually against five concrete cases, and see how often you have to look in more than one system to answer it. That is the whole test, and it takes a day.
The test works because it narrows the question from data volume down to the concrete purpose. You do not need to know whether your data is good in general. You need to know whether it holds up for this one application. A company can be excellently set up for analysing order lead times and not at all for a customer revenue analysis.
For each of the five cases note three things: how many systems you looked in, how often the entry was contradictory, and how long it took. After that the decision is arithmetic rather than opinion.
Does cleanup have to happen first, or can both run in parallel?
In parallel, and only for the slice the application needs. A cleanup covering the whole inventory is a project of its own with its own budget, and it pushes the benefit back by months.
The effective scope is the reverse: the planned application determines which fields and which records have to be in order. Everything else stays where it is until an application needs it. This approach occasionally gets criticised as sloppy, and it is the only variant that reaches the finish line in a mid-sized company.
There is one exception: deciding which system governs which entry. That decision applies to the whole company and cannot be taken slice by slice. It costs one meeting and is the precondition for everything else.
Who decides which source governs?
Management determines which system governs which data type. The department maintains it there. Without that decision the double truth reappears continuously, no matter how often cleanup happens.
That decision needs no twenty-page policy. A table with three columns is enough: data type, governing system, responsible person. Customer master data, article data, prices, contacts, contracts — to begin with, the five that appear in the planned application will do.
How to organise that responsibility without your own IT department is covered in our article on responsibility for AI without an IT department.
Frequently asked questions (FAQ)
Can we start with AI even though our data is messy? Yes, with the applications that do not reach into your data. Text assistance, summarising, translation and research run independently of the state of your systems.
Do we need a data warehouse before AI makes sense? Not for the first and second stage. A data warehouse becomes sensible when analysis regularly has to span several systems, and it is an investment decision with its own justification.
How many duplicates are too many? It depends on whether they sit in the fields being analysed. Five per cent duplicates in a field the application never reads are meaningless. One per cent in the key field makes the analysis useless.
Won’t the AI do the cleanup itself? It genuinely helps with consolidating inconsistent spellings. It does not help with deciding which of two contradictory entries is correct, because that is a decision question and not a recognition question.
How long until data holds up for an assistant? For a bounded area with a clear filing structure, a few weeks. For a company’s entire file inventory it is not a timeframe but a standing task.
Further reading
- Which processes are suitable for automation? for choosing the processes where data access becomes necessary.
- How do I recognise whether my processes are ready for digitalisation? for the process side of the same question.
- How does an SME use AI in accounting without its own IT department? for an area where data access is needed early.
Want to know whether your data holds up for the application you have in mind? Book a conversation.
Sources: Experience from PASSION4IT Digital Check and AI readiness projects. Practical guidance, not legal advice. As of 25 August 2026.