What are AI hallucinations, and why do they happen?
AI hallucinations in business tools are confident, fluent answers that are simply not true. They happen because a language model is built to produce the most plausible next words, not to check whether those words correspond to anything real. Asked something it has no basis for, it does not stop. It produces a plausible answer, in the same tone as a correct one, because tone and truth are unrelated in its machinery.
That is worth stating plainly, because most of the anxiety about this dissolves once you stop treating it as a defect awaiting a patch. It is a property. You design around properties.
Where does the risk actually land?
On the specifics, not the prose. Names, numbers, dates, policies, references and prices are where a machine invents, and where an invention costs you something.
General explanation is usually safe, because the model has seen a great deal of it and there is no single fact to get wrong. The danger is in the particular: a part number that does not exist, a clause from a policy you never wrote, a price from last year, a case reference that sounds right, a delivery date nobody committed to. Those get repeated to a customer, and now they are your position.
The pattern is consistent. Low risk where the answer is general. High risk where the answer is specific and someone will act on it.
Can hallucinations be eliminated?
No, and any supplier telling you otherwise is selling something. They can be contained to the point where the residual risk is smaller than the human error rate you already accept.
That is the honest bar, and it is a reasonable one. Your staff also occasionally quote the wrong price or misremember a policy, and your business has controls for that: checks, approvals, records, correction. A machine gets the same treatment, sized to the same consequence.
How do you contain it?
Four techniques contain almost all of it, and they are listed here in order of how much each one buys you.
- Ground it in your data. Have the machine answer from your documents, your price book, your job records, and cite which one. An answer with a source can be checked in seconds. An answer without one cannot be checked at all.
- Let it say it does not know. Make that a first class outcome the machine is rewarded for, not a failure it tries to avoid. Most invention happens because the design left no honourable exit.
- Narrow the job. A machine asked to do one well defined thing on a defined body of information has far less room to wander than one asked to be generally helpful.
- Gate the consequence. Where an error is expensive, a person approves before it leaves. This is the last line and the reliable one, and it is covered in human in the loop AI.
What about numbers and calculations?
Do not let a language model do arithmetic that matters. Have it call the system that already knows.
Totals, tax, margins, stock levels, hours, dates: these should come from your accounting system, your job management system or a calculation the machine performs with a proper tool, not from a model recalling what a plausible figure looks like. This is a solved problem in a well built machine and an unsolved one in a chat window, which is a large part of the difference between the two.
The same applies to anything with a legal or financial consequence. A machine may retrieve your terms. It should not compose them.
How do you know if it is happening to you?
You find out by sampling the output deliberately, on a schedule, and keeping the record of what you found.
Pull a random handful of the machine's outputs each week and have someone who knows the subject read them properly. Count how many contained something factually wrong. That number is your real accuracy rate, and it is the only one worth quoting. Track it over time, because it moves when your data changes, when a supplier updates a model, or when someone widens the machine's job.
Spot checking is boring and it is the whole control. It is also why we take a baseline and measure, as set out in how to measure ROI on AI.
What should you ask a supplier?
Ask where the answers come from, what happens when the machine does not know, and how you will find out if it is wrong.
A supplier who can show you the source behind an answer, demonstrate the machine declining to guess, and describe the sampling process has built for this. One who answers by talking about how advanced the model is has not. It is the most useful five minutes in any evaluation, and how to choose an AI implementation partner covers the rest of that conversation.
Build for the property, not the hope
A machine that can say "I do not know", cite where it got something, and stop before anything irreversible is a machine you can put in front of customers. One that always has an answer is a liability wearing a confident face.
If you want to see what that looks like on your own data before committing to anything, get in touch or read our method. For the wider expectations on transparency and accountability, the AI Ethics Principles are the Australian reference point, and business.gov.au covers the general obligations that sit behind whatever you publish.