Data Sovereignty / Source date:

EU AI Act: Data Governance Obligations Arrive

Training data documentation and risk classification created new evidence requirements for AI deployments.

Illustration of an analyst reviewing a sample dataset provenance register and human oversight notes.

The European artificial intelligence regulation entered into force today. Nothing is required of anyone today, which is exactly why it is worth reading now rather than in eighteen months. The obligations arrive on a staircase. Prohibited practices and the requirement that staff using these systems be adequately literate in them apply from early February next year. Obligations on providers of general-purpose models apply from August next year. The substantial high-risk regime — the one with conformity assessment, technical documentation and data governance — applies from August 2026, with an extra year for artificial intelligence embedded in regulated products.

The Act does not regulate artificial intelligence. It regulates products that contain it, and it asks you to prove where their data came from

That is the shift most organisations have not absorbed. This is product safety legislation in structure, not data protection legislation, and the evidentiary burden it creates is retrospective.

Four tiers, and most of your systems are in the boring ones

A small set of practices is prohibited outright. A defined list of uses is high-risk, including employment and worker management, access to essential private and public services, creditworthiness assessment, education, biometrics and critical infrastructure. A third tier carries transparency duties — tell people they are interacting with a machine, label synthetic content. Everything else is effectively unregulated. The common misreading is to assume a company running assistants and forecasting tools is heavily caught. Usually it is not. The exposure concentrates in a handful of systems that touch hiring, promotion, credit and access to services, and those systems are frequently bought cheaply by a department that never told anyone.

The data governance obligation is the expensive one

For high-risk systems, training, validation and testing datasets must be relevant, sufficiently representative, and to the best extent possible free of errors and complete in view of the intended purpose. They must be examined for bias that could lead to discrimination. And the design choices, collection processes, origin of the data, and the assumptions behind it must be documented. Read that carefully, because the phrasing is more forgiving than it first appears and more demanding than it sounds. Nobody is required to produce a perfect dataset. What is required is that you can describe, in writing, where the data came from, what population it represents, what it omits, what was done to it, and who decided all of that. Almost no organisation can produce that document for a dataset already in production. The information was never recorded, the people who assembled it have moved on, and the provenance is a folder of extracts. That is why the deadline being two years away is not comfort: provenance cannot be reconstructed after the fact, only recorded going forward.

Record the dataset decisions as they happenArticle-derived record-keeping prompts, not a complete AI Act compliance checklist. Applicable duties depend on system classification, role and current law.
  1. Purpose and population

    State the intended use and the population the dataset represents.

  2. Origin and gaps

    Record where the data came from and what is missing.

  3. Changes and assumptions

    Document filtering, transformations and design choices.

  4. Decision ownership

    Record who decided and retain the rationale.

Qualitative summary of this article's source text, not a measured outcome or performance estimate.

Provider or deployer, and why buying does not transfer everything

Most organisations will be deployers rather than providers. The deployer obligations are lighter but not trivial: use the system according to the provider's instructions, assign human oversight to people with the competence and authority to exercise it, ensure input data is relevant to the intended purpose, keep logs, and inform workers where a high-risk system is used in employment. Two things flip you into the heavier provider role. Putting your own name or trademark on a high-risk system, and making a substantial modification to one, including changing its intended purpose. Both are easier to do accidentally than anyone expects.

What the model layer changes for your contracts

From next August, providers of general-purpose models carry documentation, copyright policy and training-data summary obligations, with additional requirements where a model presents systemic risk. You will experience this as new disclosures appearing in supplier terms, and as an opportunity: the documentation you will eventually need from your vendors becomes something they are obliged to produce. Ask for it at the next renewal rather than in 2026, when everyone will be asking at once.

The twelve-month plan

Inventory every system in your organisation that makes or materially informs a decision about a person. Classify each against the high-risk list honestly rather than optimistically. Start a dataset register for anything that might qualify, recording origin and design choices from now on. Design the human oversight rather than asserting it — a named role, a real ability to override, and a record when it happens. And put documentation obligations into contracts before renewal leverage disappears.

Practical Guidance for AI Act Readiness Assessment

  • Inventory every system that decides or scores a person.
  • Classify against the high-risk list without wishful thinking.
  • Start a dataset register now; provenance cannot be backfilled.
  • Check whether you brand or modify any purchased system.
  • Design human oversight with authority, not just a reviewer's name.
  • Retain logs for systems that will fall in scope.
  • Add documentation duties to supplier contracts at this renewal.
  • Plan the literacy requirement landing in February.

The Regional Angle

The first thing to check is whether you are a provider rather than a deployer, because regional software and services businesses cross that line more often than they realise. A group that white-labels a screening tool inside its own recruitment platform, an implementation partner that packages a scoring model into a sector solution sold under its own brand, or a fintech that embeds a purchased credit model into a product placed on the European market is a provider of a high-risk system, carrying conformity assessment, technical documentation, quality management and post-market monitoring duties. The distinction turns on whose name is on the product and whether the intended purpose was changed, not on who wrote the model. Review every product you place on the European market under your own brand, and every deal where a European client resells your work onward, because the obligation follows the name on the box. The second is a genuine collision between regional record-keeping and European data expectations. Human resources and payroll systems in this region necessarily carry nationality, passport and visa details, sponsor information, marital and dependant status, religion in some configurations, and gender — not through carelessness but because labour, immigration and benefits filings require them. When a dataset for a screening, promotion or workforce planning model is extracted from those systems, every one of those fields is either directly sensitive or a strong proxy for a protected characteristic, and a bias examination will find exactly what you would expect. Separate the fields the regulator requires you to hold from the fields a model is permitted to see, document that separation as a design choice, and keep the record of why each retained field is necessary for the stated purpose. That single piece of documentation will do more work in a future assessment than any amount of model tuning. The third is recruitment, which is where the regional exposure concentrates in sheer volume. Organisations here screen applicant numbers that would be unusual elsewhere, across many nationalities and qualification systems, frequently using inexpensive ranking tools bought by a human resources team without technical review. Employment and worker management sits squarely in the high-risk category, and the deployer duties — human oversight, informing workers, input relevance, logging — apply to the organisation using the tool regardless of who built it. Start with an inventory of what is actually in use, including anything embedded in an applicant tracking system by default, and then establish the one artefact nobody has: a record showing that a competent person reviewed and sometimes overrode the ranking. Oversight that is never exercised is not evidence of oversight.

The objection worth taking seriously

The strongest objection is that this is premature. The substantial obligations are two years out, the harmonised standards that will define what compliance actually looks like are not finished, guidance on classification is incomplete, and European legislation of this scale has a history of arriving softer and later than the text suggests. Spending money now means buying a compliance posture against requirements whose operative detail does not yet exist, and there is a fair chance that much of the work will need redoing once the standards land. That is a reasonable reading of the situation and it should restrain the size of the programme. Nobody should be running a two-year conformity project in August 2024. But it argues for sequencing rather than waiting, because the obligations divide cleanly into two kinds. The ones that depend on standards — conformity assessment procedures, technical file formats, testing methodologies — can genuinely wait, and starting them now is wasted effort. The ones that depend on records cannot, because they describe things that are happening today and will be unrecorded tomorrow. Dataset provenance, design decisions, oversight events and the contents of supplier contracts are all cheap to capture as you go and impossible to reconstruct later. Do the second category now, at a cost of perhaps a few days a quarter, and leave the first until the standards exist. That distinction is the whole of a sensible response this year.

Common Questions

Does this apply to us if we have no European entity?

It can. The regulation reaches providers placing systems on the European market and situations where the output is used in Europe, so a system sold to or used by a European customer can bring you into scope without any establishment there.

Are general business assistants high-risk?

Generally no. Drafting, summarising and internal search sit outside the high-risk list. The classification follows the use, so the same underlying model can be unregulated in one application and high-risk in another.

What happens in February?

The prohibitions take effect, along with the requirement that people dealing with these systems have sufficient understanding of them. The literacy duty is broad, cheap to satisfy with sensible training, and easy to forget about until January.

What should we expect over the next twelve months?

Expect harmonised standards work to dominate the technical conversation and to run late, which will push practical clarity towards the end of the period. Expect national supervisory authorities to be designated unevenly, with some member states well behind. Expect vendor contracts to start carrying artificial intelligence schedules, initially of poor quality. And expect the February prohibitions and literacy deadline to catch organisations that assumed the whole regime was a 2026 problem.


AI Act Readiness Assessment — we classify what you actually run, separate the work that can wait from the records that cannot, and start the register today.

Continue reading

Talk to OPS

Start with the operating problem.