Data Sovereignty / Source date:

Right to Be Forgotten Ruling Reshapes Data Retention

Deletion obligations forced organizations to build data lifecycle capability, not just storage policy.

Illustration of an archival index card being removed while printed source material remains on the shelf.

The right to be forgotten did not arrive as a data retention policy debate. It arrived as a man in Spain who was tired of a sixteen-year-old newspaper notice about his debts being the first thing anyone found when they searched his name. On 13 May 2014, the Grand Chamber of the Court of Justice of the European Union ruled in his favour — and in doing so converted a question most organisations had treated as a storage cost into a question of legal exposure. The practical consequence took years to sink in. Keeping everything forever had been the default posture of nearly every enterprise data architecture built in the preceding decade. Disks were cheap, analytics teams wanted history, and no one got fired for having too much data. After Costeja, "we keep it because deleting it is hard" stopped being a defensible answer.

A repossession notice, sixteen years later

The underlying facts were mundane. In 1998, the Spanish newspaper La Vanguardia published a short notice about a real-estate auction connected to social security debt recovery. One of the names in that notice was Mario Costeja González. The debt was settled. The notice stayed online, indexed, and permanently attached to his name. In 2010 he complained to the Spanish data protection authority, the AEPD, asking both the newspaper and Google to remove it. The AEPD rejected the claim against the newspaper — the publication had been lawful and was arguably still part of the public record — but upheld it against the search engine. Google appealed, the Audiencia Nacional referred the questions to Luxembourg, and the case became Google Spain SL and Google Inc. v AEPD and Mario Costeja González, C-131/12. The distinction the AEPD drew is the one most commentary lost. The court was not asked to erase history. It was asked whether the index that made history instantly retrievable by name was itself performing an act of data processing that carried obligations.

What the court actually decided

Three findings mattered. First, a search engine operator is a controller of the personal data it indexes, not a neutral conduit. Crawling, storing, organising and making data available by name was processing under Directive 95/46/EC, and the operator determined the purposes and means of it. Second, European law applied. Google Inc. ran the index from outside the EU, but Google Spain sold advertising into the Spanish market, and the court held that the two activities were inextricably linked. Offshore infrastructure did not place the processing beyond reach. Third — and this is the operative finding — individuals can request that results be delisted from searches on their name where the information is inadequate, irrelevant, no longer relevant, or excessive in relation to the purposes of the processing, and where the passage of time has weakened the justification for continued prominence. The right is not absolute. It is balanced against public interest in access, with a clear carve-out where the individual plays a role in public life. Note what the ruling did not do. It did not require the newspaper to take the notice down. It did not delete anything from the web. It required a specific intermediary to stop returning a specific link for a specific name query in a specific jurisdiction. "Right to be forgotten" was always a misleading label for what is closer to a right to not be permanently and effortlessly indexed.

Why retention schedules broke

Most enterprises in 2014 had a retention policy. Very few had a retention capability. The policy was a document, usually drafted by legal, specifying how long categories of records should be kept. The capability — the ability to actually locate every copy of a person's data and remove it on request within a defined window — did not exist. The gap had structural causes: Data had been copied, not moved. A customer record lived in the CRM, then in the data warehouse, then in a reporting mart, then in three analyst extracts, then in a vendor's platform, then in a spreadsheet on someone's laptop. Deletion at source touched one of those. Backups were designed for restoration, not selective removal. Pulling one individual out of an immutable snapshot taken eighteen months ago was, for most architectures, simply not a supported operation. Logs were invisible to retention governance. Application logs, access logs, audit trails and message queues were full of identifiers that nobody classified as personal data until a regulator did. And retention periods, where they existed, were set by the most conservative voice in the room. If tax law said seven years and one lawyer said ten to be safe, the system kept it forever, because "forever" required no configuration.

The engineering problem nobody budgeted for

What the ruling exposed — and what GDPR would later formalise into Article 17 — was that deletion is a distributed systems problem wearing a legal costume. To honour a deletion request properly, you need to know every system that holds the data, every downstream system it was replicated to, every processor operating on your behalf, and every derived artefact that could reconstruct the original. That inventory is exactly the thing most organisations have never built. The data map is the expensive part; the delete button is trivial once you know where to point it. The organisations that handled this well over the following decade did one unglamorous thing first: they stopped creating uncontrolled copies. Fewer extracts, fewer shadow datasets, fewer vendor integrations pushing full records where a token or a reference would do. Minimisation at the point of collection and copying turned out to be cheaper than deletion at the point of request.

Follow the request across the copy inventoryArticle-derived review questions, not a legal determination that every copy must be erased. Applicable exceptions and restrictions need individual assessment.
Copy or purposeReview question
Live operational recordsWhat purpose and retention basis still apply?
Exports and processor copiesWhich recipients can carry out the approved action?
Backups and restore pathsHow will restoration avoid reintroducing an approved deletion?
Statutory and evidential recordsWhich obligation or exception requires restricted retention?

Qualitative summary of this article's source text, not a measured outcome or performance estimate.

Practical Guidance for Retention Policy Redesign

  • Build the data map before the policy. A retention schedule that references system names nobody can enumerate is documentation theatre. Start with an honest inventory of where personal data lands, including exports and vendor platforms.
  • Set retention by purpose, not by maximum legal ceiling. The question is not "what is the longest we are allowed to keep this?" but "what specific business or statutory purpose requires it next year?" Different purposes justify different clocks on the same record.
  • Separate statutory records from convenience copies. Payroll, tax and employment records usually carry mandatory minimum retention. Marketing histories, session logs and analytics extracts generally do not. Conflating them means the strictest rule governs everything.
  • Treat backups as a documented exception, not a silent one. Write down that restoration-only backups age out on a defined cycle, that deleted subjects are re-suppressed on restore, and that the window is bounded. Regulators have tolerated this position; they have not tolerated pretending the problem does not exist.
  • Push deletion obligations into processor contracts with actual mechanics. "Processor shall delete upon instruction" is worthless without a named interface, a response window, and evidence of completion.
  • Instrument deletion so you can prove it. Logged request, logged systems touched, logged completion timestamp. If you cannot evidence a deletion, you cannot defend it.
  • Kill uncontrolled extracts at the source. Every CSV a analyst pulls to their desktop becomes a copy outside every control you built. Give them queryable access instead of downloadable files.
  • Rehearse a subject request end to end, twice a year. Run a real one against real systems with a stopwatch. The first rehearsal will find gaps no policy review ever surfaced.

The Regional Dimension

For organisations operating across the GCC, the retention problem has a different shape — and in several respects a harder one. Employment records are the clearest case. Under UAE wage protection arrangements and end-of-service gratuity calculations, and under GOSI contribution histories in Saudi Arabia, employers must retain detailed payroll data for defined statutory periods. Those obligations sit alongside, not instead of, privacy expectations. The region's high workforce turnover — amplified by residency being tied to employment — means the volume of ex-employee personal data accumulating in HR systems grows faster than in most markets, and nobody owns the question of when it leaves. Multi-entity structures compound it. A group with a mainland company, two free-zone entities and a Saudi branch is very likely running separate HR and finance instances with separate retention behaviour, separate backup regimes and the same individual represented four times. Bilingual records make identity resolution worse: the same person appears under an Arabic name, an English transliteration, and two inconsistent spellings across systems, so a deletion request satisfied in one place leaves three copies intact. There are also intermediaries holding data outside the corporate estate. PRO and government-relations agents routinely retain copies of passports, visas, contracts and Emirates ID details to process filings. Those copies are almost never in anyone's data map, and they are exactly the records a subject request would cover. Meanwhile the regulatory floor has risen. Saudi Arabia's personal data protection law and the UAE's federal data protection framework both establish rights that assume deletion is operationally possible, and e-invoicing regimes such as ZATCA's impose their own archival requirements on transaction records. The result is a two-sided constraint: certain data must be kept for a fixed term, and other data must be removable on request. Organisations that never separated the two categories end up unable to satisfy either obligation cleanly.

The objection worth taking seriously

The strongest criticism of the 2014 ruling was never about data hygiene. It was that the court delegated an editorial judgement — what the public is entitled to find — to a private company with no obligation to explain its reasoning, no adversarial process, and a commercial incentive to resolve ambiguity in whichever direction costs least. That criticism has aged well. Delisting decisions are made at volume by staff applying internal criteria, and the person whose interest in access is being weighed — the reader, the journalist, the future employer — is not in the room. Publishers frequently learn a link has been delisted only after the fact. The counterargument is that the alternative was worse: an individual with no remedy at all against permanent, automated, name-keyed retrieval of a settled debt from 1998. Both things are true. A remedy that works imperfectly and a remedy that does not exist are not equivalent, but neither is the current arrangement a clean resolution.

Common Questions

Does the right to be forgotten mean data must be erased everywhere?

No. The 2014 ruling concerned delisting from name-based search results, not deletion from the source publication. Later law, particularly GDPR Article 17, established a broader erasure right against controllers — but that right is still qualified by legal obligations, freedom of expression, public interest, and the establishment or defence of legal claims. Statutory retention obligations generally override an erasure request.

How should backups be handled in a deletion request?

The defensible position is documented suppression plus bounded ageing: remove the record from live systems immediately, record the subject on a suppression list so any restore re-applies the deletion, and demonstrate that backup media itself expires on a defined cycle. What is not defensible is treating backups as a permanent exemption that quietly preserves everything indefinitely.

What retention period should we actually set?

There is no universal answer, which is the point — the period should be derived from a stated purpose and any statutory minimum, per data category, and then enforced automatically. A schedule that lists twelve categories with real justifications beats one that lists two hundred with none, because only the first one will ever be implemented.

Does this affect AI systems trained on personal data?

It is the sharpest open question. A record can be deleted from a database in seconds; removing its influence from a trained model is a genuinely unsolved problem, and embeddings, vector stores and fine-tuning datasets create derived copies that traditional data maps rarely capture. Organisations building on their own historical data should assume that what they feed into training is, in practical terms, difficult to un-feed — which makes minimisation before training far more valuable than remediation after it.


Retention Policy Redesign — most retention policies fail not because the rules are wrong but because nobody ever built the capability to enforce them, and that gap only becomes visible on the day someone asks you to prove a deletion happened.

Continue reading

Talk to OPS

Start with the operating problem.