Data Sovereignty / Source date:

AI Training Data and Cross-Border Transfers

Model training moves personal data internationally in ways most transfer documentation never described.

Illustration of a technician tracing separate cable routes into hardware bays behind a glass partition.

Nine days ago the White House issued its executive order on artificial intelligence. Last week thirty governments signed a declaration at Bletchley Park. The G7 published a code of conduct on the same day as the executive order, and in Brussels the trilogue negotiations on the European text are grinding towards something before the end of the year. None of it answers the question your procurement team is actually stuck on: when someone pastes a client document into an assistant, where does that document go, who can read it, and how long does it stay there?

The debate is about training data. The exposure is in the prompts

Almost all of the public argument concerns what went into the models — scraped text, copyrighted works, whether consent was obtained. That matters, and it is someone else's problem to litigate. Your problem runs the other way. Every day your staff send material outward: contracts, payroll spreadsheets, customer correspondence, draft board papers, code. That is a cross-border transfer of personal data, performed hundreds of times a day, by people who do not experience it as a transfer because it looks like typing into a box.

Four flows, four different answers

The single question "is our data used for training" collapses four distinct flows that need separate answers. Inference input. The prompt itself, plus any attached document or retrieved context. This is the highest-volume flow by far. It usually leaves your region, frequently to the United States, and it does so regardless of where your tenant is hosted unless you have specifically configured otherwise. Training and model improvement. The contractual question everybody asks. Consumer tiers historically default to using your input; enterprise and business tiers generally commit not to. The answer is in the agreement, not on the marketing page, and it differs between the same vendor's products. Abuse monitoring and human review. The flow nobody reads about. Even where training is contractually excluded, providers commonly retain prompts for a period — often around thirty days — for safety and misuse detection, with a small number of authorised staff able to view flagged content. "Not used for training" and "not retained" are entirely different commitments, and only one of them is usually given. Fine-tuning and embeddings. Where your data persists in derived form. A fine-tuned model or a vector store built from your documents lives in a specific region, under specific retention terms, and is frequently the flow with the weakest paperwork because it was set up by an engineer during a proof of concept.

Ask about each data flowThe article's four-flow assessment checklist, not verified processing terms for any provider.
FlowWhat to identifyQuestion to record
Inference inputPrompt, attachment and retrieved contextWhere does processing execute?
Training and improvementContractual use of submitted dataWhich product and tier prohibit reuse?
Abuse monitoringRetained content and human accessWhat is retained, for how long and who can read it?
Fine-tuning and embeddingsDerived model and vector-store dataWhere does it persist and what deletion route exists?

Qualitative summary of this article's source text, not a measured outcome or performance estimate.

Why the standard transfer analysis under-covers this

Your transfer register tells you whether a flow is lawful. It does not tell you where the processing physically happened, how long the content was held, or who was able to read it — and those are the questions that determine whether a client would be comfortable. Two specific traps. First, residency commitments frequently cover storage and exclude inference, so "your data stays in Europe" can be true of the database and false of the model call. Second, model availability varies by region: the newest and best model is often unavailable in the compliant region for months. When that happens, the business chooses the model, and your residency position is decided by a dropdown.

Deletion is where derived data bites

You can delete a document. You can delete an embedding, if you built the store yourself and know which records correspond to it. You cannot remove a document's contribution to a fine-tuned model without retraining it. That is not a vendor failing; it is how the technology works. The consequence is that the decision about what may be used for fine-tuning is effectively irreversible, and should therefore be taken deliberately, in advance, by someone senior — not discovered afterwards during a deletion request.

The artefact to produce

One page per AI tool in use: which model, which processing region, whether inputs train the model, retention period for abuse monitoring, who can access retained content, whether fine-tuning or embedding stores exist and where, the subprocessor list, and the deletion route with its limits. Then a single classification rule that staff can actually remember, expressed in terms of data rather than tools: what may never be pasted anywhere, what may go to approved enterprise tools, what may go anywhere. Rules written per tool go stale monthly. Rules written per data class survive.

The incoming regulation does not replace any of this

The European text under negotiation is largely about model risk tiers, transparency and obligations on providers. It sits alongside the existing data protection regime rather than superseding it, which means the transfer analysis you are avoiding today will still be required after it passes, plus new documentation. Nothing about waiting improves your position.

Practical Guidance for AI Data Transfer Review

  • Separate the four flows rather than asking one question about training.
  • Read the retention and human review terms, not just the training commitment.
  • Confirm whether residency covers inference or only storage.
  • Inventory fine-tuning jobs and vector stores built during proofs of concept.
  • Decide in advance what may be fine-tuned on; it is effectively irreversible.
  • Write the rule by data class, not by product name.
  • Check which models exist in your compliant region before committing to it.
  • Record the subprocessor list and diarise a re-check each quarter.

The Regional Angle

Three things shape this differently here, and the first is contractual rather than statutory. If you serve government, semi-government or large national corporate clients in the Gulf, your obligations about where their data may be processed almost certainly come from your contracts rather than from privacy legislation. National data classification frameworks — Saudi Arabia's in particular, and the various emirate-level government data policies — flow down through procurement into clauses that specify classification levels, in-country processing and restrictions on subcontracting, and they were written before anyone was thinking about assistants. A staff member using an approved corporate AI tool on a document from a classified government engagement may be breaching a contract that predates your AI policy by three years and outranks it commercially. Before writing the policy, read the ten largest client contracts and find out what you have already promised. That constraint is usually tighter than anything the law requires. The second is that your data is unusually valuable to the people asking for it. High-quality Arabic text, and particularly domain-specific Arabic — legal, financial, medical, technical — is scarce relative to English, and everyone building models for this region knows it. Expect requests to use your corpus for model improvement, expect them framed as partnership, innovation collaboration or preferential pricing, and expect them to arrive inside a renewal negotiation rather than as a standalone proposal. That is a legitimate commercial conversation to have, but it should be a priced one with defined scope, defined retention and a defined exit, and it should not be agreed by whoever signs the renewal. Treat a request for your data as a request for an asset, because that is what it is. The third is that a domestic option now genuinely exists, and it resolves one problem while creating another. Regional sovereign AI programmes, national research institutes and state-linked cloud providers offer in-country processing that answers the residency question cleanly for regional data, often with better Arabic performance than the international alternatives. For a locally headquartered group serving local clients, that is a strong answer. For a group with European operations it introduces the mirror image of the problem everyone has been managing for a decade: routing European personal data to a state-linked provider in a third country with no adequacy finding is a harder conversation than routing it to a certified American vendor, not an easier one. Decide the question per data population rather than per vendor, and accept that the right answer may be two providers.

The objection worth taking seriously

The strongest objection is that this is a great deal of legal apparatus for a text box. The overwhelming majority of prompts contain nothing sensitive — rewrite this paragraph, explain this error message, summarise these notes. Enterprise agreements from the major providers are now genuinely reasonable on training and retention, considerably better than the terms attached to plenty of software nobody worried about. And the cost of getting this wrong in the restrictive direction is real: staff who cannot use approved tools use unapproved ones, which is strictly worse for everyone including the lawyers. That is mostly right, and the last point especially. Over-restriction is a worse failure mode than over-permission here, because it relocates the problem somewhere you cannot see. The exposure is not the everyday prompt. It is three specific flows that people do not experience as risky: the attached document, where a single drag-and-drop sends a complete client file rather than a paragraph; the fine-tuning job someone ran in a proof of concept and never decommissioned; and the connectors now being enabled across collaboration suites, which give an assistant reach across the entire file estate without anyone pasting anything at all. Those three deserve real scrutiny. The text box mostly does not. A policy that recognises the difference gets followed, and a policy that treats every prompt as a transfer event gets ignored within a fortnight.

Common Questions

Does an enterprise agreement solve this?

It solves the training question and usually improves retention. It rarely addresses processing region for inference or the human review path, so read those clauses specifically.

Are embeddings personal data?

If they are derived from personal data, assume yes. Treat the vector store as a copy of the source material for retention and deletion purposes.

Can we rely on the adequacy decision for prompts to US providers?

Where the contracting entity is certified for the relevant data categories, yes — that handles lawfulness. It says nothing about retention or who reads flagged content, which are the questions your clients will ask.

What should we expect over the next twelve months?

Expect political agreement on the European text within weeks or a short slip into the new year, followed by obligations phasing in over two years or more — so preparation, not panic. Expect regional processing options to expand quickly as the major providers localise, which will make the residency question easier and the model-availability gap the binding constraint instead. Expect AI-specific addenda to appear in standard data processing agreements through next year, and read the retention clause rather than the headline. And expect the first supervisory authority decision specifically about prompt data to land somewhere in Europe within the year, which will set the tone for everything after it.


AI Data Transfer Review — we map the four flows behind every assistant you use, find the fine-tuning jobs nobody decommissioned, and write the rule by data class so it still works next quarter.

Continue reading

Talk to OPS

Start with the operating problem.