What does preparing data for AI actually involve?
Preparing data for AI is mostly not a technical task. It is finding where your information actually lives, agreeing which version is correct, and writing down the rules that currently exist only in somebody's judgement. A machine can only work with what has been described. Most small businesses run beautifully on things nobody has ever described.
That is the real project, and the good news is that almost all of it has value whether or not you build anything. A written process survives a resignation. A single agreed price list ends an argument. Those are worth having on their own.
Where does the information actually live?
In more places than the answer you would give in a meeting, which is the first thing worth finding out.
The job records are in the job management system. The pricing is in a spreadsheet, and there is a second spreadsheet somebody uses instead. The supplier terms are in email. The reason you do it differently for one customer is in a text message from 2023. The rules for the awkward cases are in the head of whoever has been there longest.
Draw that map before anything else. Not a formal audit, an hour with a whiteboard and the people who do the work. The map tells you what a machine could reach today and what would first have to be dragged into daylight, and it usually surprises the owner more than anyone.
What has to be cleaned up?
Three things need cleaning up before a machine can use your records: duplicates, inconsistency, and gaps, in roughly that order of annoyance.
- Duplicates. The same customer under three spellings, the same product with two codes. A machine will treat them as different, and every report it produces will be quietly wrong.
- Inconsistency. Dates in three formats, statuses that mean different things to different people, a field everyone uses for something other than its label. Machines are literal, which exposes conventions nobody knew were conventions.
- Gaps. Missing outcomes, unfinished records, jobs that were closed in real life but never in the system. If half your records have no result, a machine cannot learn from results.
You do not need perfection, and chasing it is a good way to never start. You need the specific slice the first machine will touch to be trustworthy. Everything else can wait for its turn.
What about the rules in people's heads?
Getting those written down is the single most common place a build stalls, and it should be treated as the first deliverable rather than as a blocker.
The pattern is always the same. The build works for most cases, then reaches the ones where the answer is "well, it depends, Steve just knows". Steve does just know. Steve has known for eleven years. But nobody has written down the rules Steve applies, so the machine cannot apply them either.
The way through is not a workshop. It is watching Steve do twenty real cases and writing down what he did and why. Two of those twenty will produce a rule nobody in the room knew existed. This is exactly the failure we describe in why AI projects fail, and it is fixable for the price of a couple of afternoons.
Do you need a data warehouse first?
No, and being told you do is a good reason to get a second opinion.
Small and mid sized businesses build useful machines on the systems they already have, connected directly, for a fraction of the cost of a platform project. A warehouse is the right answer when you genuinely have many systems, high volume and a reporting problem, and the wrong answer when someone wants to sell you eighteen months of infrastructure before you have proven anything works.
Start with one process and the data that process touches. If a warehouse turns out to be justified later, you will know because you will have hit the specific wall it solves.
What about privacy and security?
Decide what the machine may hold before you decide what it will do, because that boundary is far harder to move afterwards.
Least privilege is the right default. Give a machine access to the fields the task needs and nothing else. Be deliberate about personal information in particular: putting it into a third party service is a disclosure, and the ordinary obligations apply. The Australian Privacy Principles are the reference, and the ACSC has practical guidance on securing the systems the data sits in. We go further into this in AI data security for Australian businesses.
How much preparation is enough?
Enough preparation for one lane is enough, and preparing everything in advance is how a project quietly turns into a year of housekeeping.
The trap is treating data readiness as a phase that must complete before value can begin. Businesses disappear into that for a year and emerge with tidy data and no machine. Pick the first process, clean what it touches, write down the rules it needs, build it, and let what you learn tell you what to prepare next.
Where to start
Draw the map. Pick one process. Clean its slice. Write down the twenty cases. That is a fortnight of unglamorous work that will make everything after it cheaper, and it is work you own regardless of who builds anything.
If you would rather do that with someone who has done it before, get in touch, or read our method for how we sequence it against a fixed price.