How do you measure ROI on an AI project?
Proper AI ROI measurement is decided before the build, not after it, because the numbers you need are destroyed the moment the machine goes live. Pick two or three measures, write down what they are today, agree what movement would count as success, then read them again after thirty days. That is the whole discipline, and skipping the first step is why most AI spending is never honestly assessed.
Once the machine is running, the old world is gone. You cannot count how many enquiries were missed last quarter, or how long quotes really took, because nothing was counting. What is left is a discussion about whether things seem better, which is won by whoever is most invested rather than whoever is right.
Which numbers are worth measuring?
The ones that connect to money or to time, and that somebody was already responsible for.
- Time. Hours a week spent on the specific task by the specific people. Not "productivity". Hours.
- Speed. Time from trigger to done. Enquiry to first reply. Job finished to invoice sent. Quote requested to quote delivered.
- Volume handled. Items processed without a person touching them, as a share of all items.
- Leakage. The things that used to fall through. Enquiries with no reply. Quotes never followed up. Invoices past terms.
- Quality. The proportion of machine outputs a person changed before approving, which is the single most useful number in the whole set.
Two or three of those. A list of eleven metrics is a list nobody reads.
What should you avoid counting?
Anything you cannot attribute, and anything nobody would have measured before AI was involved.
Revenue is the obvious trap. If revenue rose in the quarter you switched something on, you have learned very little, because a dozen other things also happened. Attribute at the level you can defend: the machine handled a number of enquiries out of hours, of which some converted, and here they are by name. That is a claim you can stand behind in a room.
Avoid vanity counts too. Messages sent, documents processed and tokens consumed all go up on their own and prove nothing. Activity is not outcome.
What does a fair comparison look like?
The same measure, the same definition, the same period length, before and after, with the exceptions declared.
Compare a month to a month, not a month to a quarter. Say what was unusual in either window, because there always is something, a holiday period, a big job, a staff absence. Declaring the noise up front is what makes the signal believable, and a review that admits the awkward bits gets believed on the parts that matter.
If you can, keep a slice unautomated for the first month as a control. It is not always practical, but where it is, it settles the argument permanently.
How long before you judge it?
Thirty days for the operational numbers, a quarter before you judge the commercial ones.
Thirty days is long enough for the process to stabilise and for the early corrections to work through, and short enough that nobody has forgotten what the point was. Judge speed, volume and quality then. Judge anything involving customer behaviour or revenue later, because sales cycles do not respect your review date.
We build to a thirty day verdict for exactly this reason, one page, numbers against target. The process is set out in how to run a 30-day AI pilot.
What if the numbers say it did not work?
Then you have got your money's worth from the measurement, and the right move is to say so plainly and decide.
Three outcomes are all legitimate. It worked, so extend it. It did not, so turn it off and stop paying. Or it worked in a narrower lane than expected, so shrink the scope to that lane and keep it. What is not legitimate is quietly leaving it running because switching it off would be embarrassing. That is how businesses accumulate subscriptions nobody can justify.
A project that ends in a clear no, cheaply and quickly, is a good project. It cost you a month and bought you certainty.
Who should own the numbers?
Someone inside your business, not your supplier, because a supplier grading its own homework is not a measurement, it is marketing.
Name the person before the build starts. Give them the definitions in writing. Have them take the baseline themselves. We will help set it up and we will hand over the working, but the number belongs to you, and the sign off page has your name on it as well as ours.
Take the baseline this week
Whatever you are planning, the cheapest and most valuable thing you can do before it starts costs nothing: write down two or three numbers as they are today, and what movement would make this worth having.
If you want help choosing which numbers actually matter for the process you have in mind, our method shows how we frame it, and get in touch when you want to put your own figures against it.
For the broader expectations around transparency and contestability, the AI Ethics Principles are worth a read, and business.gov.au covers the general planning ground.