What should you ask about AI data security before buying?
AI data security for Australian businesses comes down to five questions, and you should have written answers to all five before anything is connected: where does the data physically go, who can see it, is it used to train anyone's model, how long is it kept, and whose accounts is the whole thing running in. A supplier who cannot answer those in plain language has not designed the system yet.
None of this requires you to be technical. It requires you to insist on specifics, because vagueness here is not a communication problem, it is usually an admission.
Where does your data actually go?
It leaves your building. That is the honest starting point, and any answer that implies otherwise is wrong.
When a machine reads a customer message or a document and produces a response, that content is sent to a model provider, processed, and a result comes back. Which provider, in which region, under which contract, is the thing you need in writing. Many major providers offer Australian or regional processing options, but they are options, which means somebody has to have chosen them.
Ask for the provider name, the region, and the specific data-handling terms that apply to your account. If the answer is "it is all secure", that is not an answer, it is a mood.
Will your data be used to train someone else's model?
Under standard business and enterprise terms with the major providers, no, but it depends entirely on which terms your build runs under, and that is a thing to verify rather than assume.
The distinction that catches people is between a consumer product and a business API account. Staff pasting client information into a free consumer chat tool are operating under completely different terms from a properly configured business integration. In our experience the largest real-world exposure in most companies is not the system anyone bought, it is the tabs already open on people's laptops.
So ask two questions, not one. What are the terms on the build, and what are your staff currently doing without one. The second usually needs a policy more urgently than the first needs a contract.
Who owns the accounts the AI runs in?
You should, on anything of consequence, and this is the single clause that protects you most.
If the model provider account, the integrations and the keys sit inside your supplier's tenancy, you have a dependency dressed up as a service. If that supplier folds, gets acquired or simply becomes unpleasant to deal with, your data and your working machine are both on the wrong side of the fence.
We run larger builds in the client's own provider accounts as a matter of course, and everything we build runs in accounts the client owns. The reasoning is not complicated: leave whenever you like and it should keep working. It is also why usage costs are visible to you directly rather than marked up invisibly.
What does the Privacy Act require of you here?
If the Act applies to your business, feeding personal information into an AI system does not change your obligations, it just increases the number of places they apply.
You still need to collect only what you need, tell people what you are doing with it, keep it secure, and be able to respond if someone asks what you hold. The Australian Privacy Principles are the source, and the OAIC has published guidance specifically on using commercially available AI products that is worth reading before a rollout rather than after an incident.
Two practical consequences. First, your privacy policy probably needs updating, because "we use third-party AI services to process enquiries" is a disclosure most policies written before 2023 do not make. Second, if the machine is recording or transcribing calls, notification requirements apply, and the rules are state-specific. Handle that at build time, not at complaint time.
What are the practical safeguards worth insisting on?
Least privilege, retention limits, logging, and a human gate on anything that leaves the building.
Least privilege means the machine gets access to exactly the systems and records it needs and nothing else. A lead agent does not need your payroll. Retention limits mean conversation and document data is kept for a stated period for a stated reason, then deleted, rather than accumulating forever because nobody specified.
Logging means every action is recorded and searchable, which is both a security control and the thing that lets you audit a bad outcome afterwards. And a human approval gate on outbound actions, for as long as you want it, is the cheapest insurance available against the failure mode where a confident machine does the wrong thing quickly. That is a default in how we work rather than an upgrade.
The Australian Cyber Security Centre publishes practical guidance on securing AI systems that translates well for businesses without a security team.
What about confidential client documents?
They can be used safely, but they need a deliberate design rather than a general-purpose tool and good intentions.
Work over your own documents is the highest-value and highest-sensitivity category, which is why agents on your data is scoped per pilot rather than sold off a price list. The controls that matter are access scoping by role, so a person only gets answers from documents they were already allowed to read, citations on every answer so a claim can be traced back to a source, and a clear boundary on what the system can never surface.
If a supplier proposes indexing your entire shared drive without a permissions model, that is not a shortcut, it is an incident being scheduled.
Where to look next
The house explains the production habits behind this: monitoring, fallbacks and logs are the boring details that decide whether a system is safe a year later.
If you are weighing a build, bring your questions rather than your budget. Ask us the five questions at the top of this page and see how specific the answers are. That is a fair test to run on any supplier, including us.