One of the most common questions we hear from business leaders considering AI is some version of: if we start using these tools, where does our data actually go? And it's often asked a little apologetically, as though it's a naive question.

It isn't. It's the right question, and the fact that it's hard to get a straight answer to is a problem with how the industry communicates, not with the person asking.

This isn't an argument for caution or alarm. It's an attempt at a clear answer, because clarity is what's usually missing — replaced either by vendor reassurance that doesn't really explain anything, or by vague unease that leads businesses to avoid AI entirely for reasons they can't quite articulate.

The question underneath the question

When someone asks where their data goes, they're usually worried about a few specific things, even if they haven't separated them out.

Will our data be used to train someone else's model, so our information ends up benefiting competitors or leaking into outputs elsewhere? Is our data being stored somewhere we can't see, under terms we don't control? And are we creating a compliance or privacy problem — particularly with customer data — that we'll be accountable for later?

These are different concerns with different answers, and lumping them together is part of what makes the topic feel murky. Worth taking them separately.

On training

The concern that your data trains a public model is the most common and the most addressable. For most business-grade AI services — as opposed to free consumer tools — the terms specify that your data is not used to train the underlying models. This is usually a contractual commitment, not just a setting.

The distinction that matters is consumer versus business tier. The free version of a consumer AI tool and its enterprise equivalent often have very different data terms, and the difference is precisely this point. The practical guidance is to read the data-handling terms of the specific tier you're using, and to be cautious about putting anything sensitive into free consumer tools where the terms are less protective.

On storage and control

Where your data is stored, for how long, and who can access it is a question with concrete answers — they're just answers you have to go and find rather than assume.

Reputable business AI services document where data is processed, what's retained, and for how long. For a New Zealand business, the questions worth asking are whether data is processed in a jurisdiction you're comfortable with, whether it's encrypted in transit and at rest, and whether you can delete it on request. These aren't exotic requirements. They're the same questions you'd ask of any cloud service holding your data, and you've almost certainly asked them before about your existing systems.

The reason AI feels different is novelty, not a fundamentally different risk profile. The data governance questions are largely the ones you already know how to ask.

On customer data and compliance

This is where genuine care is warranted, and where the answer is more about your obligations than the tool's behaviour.

If you're putting customer data into an AI system, the Privacy Act obligations that already apply to that data don't disappear because there's AI involved. You remain accountable for how it's collected, used, and disclosed. The practical implications are that you should know what customer data is going into a system, have a lawful basis for it being there, and be able to explain it if asked.

For many SMEs the sensible starting point is to keep customer data out of AI systems initially, prove the value on internal and non-sensitive data first, and bring customer data in deliberately once you understand the governance — rather than as an afterthought once you're already committed.

A practical position

The clear-headed stance on this sits between the two unhelpful extremes. It isn't "AI is a data risk, avoid it," and it isn't "the vendors have it handled, don't worry." It's that data governance for AI is a knowable, manageable set of questions, most of which you already know how to ask about any system handling your information.

Use business-tier tools with terms you've actually read. Know where your data is processed and retained. Keep sensitive customer data out until you've built the governance to handle it properly. Treat the AI system like any other vendor that touches your data, because in the ways that matter, it is one.

The businesses that get this wrong tend to do so in one of two directions — either avoiding AI entirely on the basis of an unexamined fear, or adopting it without asking any of these questions at all. The position worth holding is the one in the middle: clear enough about where your data goes to deploy with confidence, rather than either paralysed or careless.