Data Sovereignty / Source date:

Where Does Your AI Inference Actually Run?

Model endpoints frequently sit outside the region customers assume, undermining residency commitments.

Technician inspects a compute chassis, an illustration of model-serving infrastructure, not a confirmed inference location.

Every organisation that has spent the last two years building a data residency position can tell you where its records are stored. Ask the same organisations where the model that reads those records actually runs, and the answers thin out immediately. That gap is not academic. Inference is processing, processing has a location, and almost nobody has written that location into a contract.

Your storage region was negotiated. Your inference region was inherited from whichever data centre had capacity the week the feature shipped

That is the honest description of how most artificial intelligence features are running in enterprise software today.

Why inference ends up somewhere else

The reasons are mundane rather than sinister. Accelerator hardware is scarce and unevenly distributed, so vendors concentrate model serving in a few locations rather than replicating it into every region they offer for storage. Many software vendors do not run models themselves; they call a model provider, which may call an infrastructure provider, and each hop can cross a boundary. Capacity management routes overflow traffic elsewhere at peak. And feature release cycles move faster than regional deployment does, so a new capability ships from wherever it was built and gets regionalised later, if at all. The result is a product that is entirely compliant about storage and silent about the moment your data is actually read.

Five questions to put in writing

Where is inference performed for each feature, by region. Which entities are involved in the chain — the software vendor, the model provider, the infrastructure operator — and which of them is a subprocessor. What happens to the prompt and the output: retained, logged, cached, or discarded, and for how long. Under what circumstances traffic is routed outside the stated region, and whether you are told. And what happens to embeddings and indexes, which are derived from your content, are frequently stored separately from it, and are routinely omitted from residency commitments. The last one catches the most organisations. A company can hold every document in its chosen region while the vector index built from those documents sits somewhere else entirely.

Ask about each feature's processing footprintQuestions from the draft, not verified model availability, provider terms or a finding that a workload breaches an obligation.
AreaWritten answer to seek
InferenceWhere is this feature's model processing performed?
EntitiesWhich software, model and infrastructure providers are involved?
Prompt and outputWhat is retained, logged or cached, and for how long?
Overflow routingWhen can processing leave the stated region and how is notice given?
Derived contentWhere are embeddings and indexes processed and stored?

Qualitative summary of this article's source text, not a measured outcome or performance estimate.

Residency, sovereignty and jurisdiction are three different claims

Residency is a statement about geography: the data sits in this country. Sovereignty is a statement about control: who operates the infrastructure, who can access it, and under whose law they are compelled. Jurisdiction is a statement about legal reach, which follows the corporate structure of the provider rather than the location of the rack. A model served from a European data centre by a subsidiary of a company incorporated elsewhere satisfies the first, partially addresses the second and does nothing about the third. Vendors answer whichever of the three they can satisfy, so be specific about which one your obligation actually concerns.

What to do about the answers you get

Three practical responses, in rough order of cost. Segment by content class. Most organisations do not need every workload processed locally; they need a defined category — regulated records, client-confidential material, government work — handled under known conditions, and everything else can run wherever the product runs. That distinction makes the problem an order of magnitude cheaper. Contract for notification. If you cannot get a guaranteed region today, get a commitment to tell you before it changes, and get the subprocessor list with change notice. Keep a fallback. For the regulated category, know what you would use if the answer became unacceptable — a regional deployment, a smaller self-hosted model, or a manual process. A dependency with no alternative is a decision you have already made.

Practical Guidance for AI Residency Review

  • Ask for inference location per feature, not per product.
  • Get the full chain of vendor, model provider and infrastructure operator.
  • Confirm where embeddings and indexes live, separately from documents.
  • Pin down retention for prompts, outputs and logs.
  • Distinguish residency, sovereignty and jurisdiction in your requirement.
  • Classify content so only the regulated category needs the expensive answer.
  • Require notice before any routing or subprocessor change.
  • Identify a fallback for anything you cannot afford to lose access to.

The Regional Angle

The first practical issue is that regional data residency commitments are increasingly written into contracts here, and they were drafted before this question existed. Government and semi-government tenders, banking and insurance arrangements, and healthcare contracts across the Gulf routinely require that data remain in-country, and the clause almost always speaks about storage and hosting. When your software vendor enables an assistant feature, the processing may leave the country while every stored record dutifully stays put — and you are the one in breach, not the vendor. Go back through the contracts that carry residency undertakings, work out whether the wording covers processing as well as storage, and if it is ambiguous, assume your client's counsel will read it strictly. Disclose the assistant feature to the client rather than discovering the problem during an audit. The second is the direction of travel in regional infrastructure, which is genuinely favourable and slower than the marketing suggests. National cloud regions, sovereign arrangements and substantial announced investment in local accelerator capacity mean the honest answer to "can we keep inference in-country?" is moving from no towards partly. The gap is that a region offering local storage and local compute frequently does not yet serve the newest models locally, so choosing in-country inference means choosing an older or smaller model. That is a real trade-off and worth making explicitly: for a regulated workload, a competent model processed locally usually beats a superior model processed abroad, and for everything else the reverse holds. Ask your vendor which specific models are served from the local region rather than whether the region exists. The third is about a category of work where the answer cannot be left to the vendor's roadmap. Regional professional services firms, family offices and advisers handling government-related, defence-adjacent or politically sensitive mandates are frequently bound by confidentiality undertakings that do not contemplate any third-party processing at all, regardless of location. For that work the residency conversation is the wrong one; the requirement is that the content never enters a shared service. The workable pattern is a small self-hosted model on infrastructure you control for that narrow category, and mainstream tooling for everything else — which means deciding, in advance, which client matters fall on which side of the line. Make that a matter of file classification handled at engagement, not a judgement made by an associate at eleven at night.

The objection worth taking seriously

The strongest objection is that this is theatre. Data crosses borders constantly in every organisation — email traverses foreign infrastructure, backups replicate, support engineers connect from other countries, and the browser your staff use reports to a service abroad. Singling out model inference for special treatment is an arbitrary line drawn around the newest technology, and it will cost real capability: the organisations that restrict themselves to locally served models will be working with materially weaker tools than competitors who do not care. Both halves are fair, and the inconsistency is genuine. Plenty of firms fretting about inference location have never asked where their support tickets are read. The response is not that inference is special in principle, but that it is different in two practical respects. It is the point at which the full content of a document is read and interpreted rather than merely transmitted or stored, which is how a supervisor and a client will see it whether or not the distinction is technically principled. And it is new enough that your existing contracts do not cover it, which means your current position is not a considered one — it is an accident of vendor architecture. The argument for asking is not that a foreign inference region is unacceptable. It is that you should know the answer, have it in writing, and be able to state it when a regulator or a client asks. For most workloads the correct decision will be to accept it and move on.

Common Questions

Does a vendor's regional hosting cover its assistant features?

Frequently not. Assistant features are often served from a separate footprint under separate terms, so ask specifically rather than relying on the platform's residency page.

Are embeddings personal data?

Treat them as derived from the source content and subject to the same obligations. Assuming otherwise is an aggressive position to be arguing during an investigation.

Is a zero-retention commitment enough?

It addresses storage, not location or jurisdiction. It is valuable and it is not an answer to where the processing happened or who could be compelled to assist.

What should we expect over the next twelve months?

Expect major vendors to publish per-region inference availability as a standard commercial term, because enterprise buyers are beginning to insist. Expect the gap between where the newest models are served and where the regional footprint exists to persist through next year. Expect regulators in this region to start asking about processing location rather than storage location. And expect the organisations that classified their content early to be the only ones able to answer quickly.


AI Residency Review — we trace where each feature actually processes your content, and get it written into the contract.

Continue reading

Talk to OPS

Start with the operating problem.