Notes

How to Run a 30-Day AI Pilot That Actually Proves Something

How do you run an AI pilot that proves something?

Knowing how to run an AI pilot properly is mostly about deciding, in writing and before anything is built, what result would count as a yes and what result would count as a no. Then you capture where you are today, run it live for thirty days, and read the numbers against the targets on one page. That is the entire discipline, and it is the part most pilots skip.

The alternative is what usually happens. Something gets built, everyone is mildly impressed, nobody can say whether it made money, and the project quietly becomes a subscription nobody wants to be the one to cancel.

Why do most AI pilots fail to prove anything?

They fail because nobody captured a baseline, so there is nothing to compare the result against.

This is the most expensive ordinary mistake in the field. Once the machine is running, the old numbers are gone. You cannot reconstruct how many enquiries you were missing in March, or how long quotes really took before, because nobody was counting. So the pilot ends with an argument about whether things feel better, which the loudest person in the room wins.

The second failure is a success definition made of adjectives. "Improve customer experience" cannot be true or false. "Answer at least ninety percent of after-hours enquiries within two minutes" can be, and everyone knows on day thirty which way it went.

What do you measure, and when?

Measure the two or three numbers that would change if the thing works, and capture them in week one before anything is switched on.

Pick numbers you already have or can start counting immediately. For a lead machine: enquiries received outside business hours, proportion answered, time to first response, jobs booked without an owner touching the phone. For an admin machine: hours spent on the task per week, elapsed time from trigger to completion, error or rework rate. For a document machine: time to find an answer, and how often the answer was right.

Two or three is the right number. A pilot with eleven metrics has none, because there will always be enough movement in eleven numbers to tell whichever story you prefer.

How long should a pilot actually run?

Thirty days live is usually right, and it should start the day it goes live rather than the day the contract is signed.

Shorter than thirty and you are reading noise. One quiet fortnight or one storm week will dominate the result. Longer than about six weeks and the pilot stops being a decision and becomes a habit, which is how a trial silently converts into a permanent unmeasured cost.

Be strict about what "live" means, and agree the definition up front. Live means real work flowing through it with real consequences. A machine that is running in parallel while everyone quietly does the job the old way is not live, it is a rehearsal, and it will produce a flattering result that does not survive contact with actual reliance.

What should a pilot cost, and who pays for the mistakes?

A pilot should have a fixed price agreed in writing before it starts, and the builder should carry the cost of their own testing and their own scoping errors.

That is not generosity, it is alignment. A supplier billing hourly is rewarded when the pilot is complicated. A supplier on a fixed price is rewarded when it works quickly. Only one of those matches your interests.

Our pilots run four weeks to live with the fee credited if you continue to the build, and the shape of the whole engagement is set out on the method page. Ask for the same structure from anyone: fixed price, named numbers, defined go-live, a real verdict.

What does the verdict look like?

One page: the agreed numbers, the actual numbers, the target, and a plain-English sentence saying whether it earned its keep.

Not a deck. Not a workshop. A page you could hand to your accountant. If the numbers beat target, the conversation is about what the next machine should be. If they did not, the honest outcome is that the pilot stands alone, you keep the report and the learnings, and nobody has to talk anyone into a phase two.

That last part is worth insisting on in writing. A pilot that cannot fail was never a test, and a supplier unwilling to write down the failure case is telling you something.

What comes after a pilot that worked?

The next single machine, chosen the same way, with the same discipline applied again.

The temptation after a good result is to expand in every direction at once, and it is the reliable way to turn a success into a mess. Small, then proven, then bigger is the only growth model we trust. It is how every client relationship we have has started, and it works for the same reason each time: the wins compound and the failures stay small enough to absorb.

It also helps to keep reporting past the pilot. A thirty, sixty and ninety day rhythm against the original week-one baseline is what turns a good demo into a defensible investment when someone asks about it a year later.

Where to look next

The crew shows what a single, well-bounded first machine looks like in practice, and what we build groups them by the problem you are trying to solve.

Before you scope one, the Australian Government's AI Ethics Principles are worth reading on oversight and accountability, and business.gov.au is the place to check for adoption support you might be eligible for.

If you would like a pilot structured this way, tell us what the problem is costing you now. The number you arrive with becomes the baseline we are judged against.

Keep reading