Data Sovereignty / Source date:

Privacy by Design Moves From Principle to Requirement

Embedding data minimization into system design cut downstream compliance work dramatically.

Conceptual illustration of a data-minimisation prototype form beside physical source, index and backup models.

Privacy by design spent roughly two decades as a set of principles that thoughtful engineers admired and almost nobody was required to follow. The formulation — proactive rather than reactive, privacy as the default setting, privacy embedded into design, full lifecycle protection, visibility and transparency, respect for the user — read well and imposed nothing. What changed is that a version of it acquired a legal citation. Article 25 of the General Data Protection Regulation, which became enforceable in May 2018, obliges controllers to implement data protection by design and by default: appropriate technical and organisational measures at the point of determining the means of processing, and settings that by default limit the amount of data collected, the extent of processing, the retention period and accessibility. The practical difference between a principle and an obligation is enormous. A principle influences the engineers who already agree with it. An obligation changes the definition of done, appears in architecture review, becomes a supplier requirement, and has a documentation burden attached to it. Most of the Gulf frameworks that followed — and the sector rules issued by financial and health regulators here — carry some version of the same expectation.

What it actually requires in engineering terms

Strip away the principles language and the obligation resolves into a small number of design decisions, each of which is cheap at design time and expensive later. Collect less. Every field on a form is a liability with a retention period, an access control requirement, a breach exposure and a subject access implication. The discipline is to ask what each field is for, and to delete the ones where the answer is "it might be useful." This is the single highest-value habit in the whole framework and it costs nothing to adopt. Set defaults conservatively. Sharing off, visibility limited, retention short, optional processing disabled. Defaults are the actual policy, because the overwhelming majority of users never change them — which means a default is a decision the organisation makes on behalf of everyone. Separate identity from behaviour. Pseudonymisation at the storage layer, with the re-identification key held separately and access-controlled, converts many breach scenarios from catastrophic to manageable and makes analytics workloads far less sensitive. Design deletion in from the start. This is the requirement that most often proves impossible to retrofit. Deletion has to propagate through backups, replicas, warehouses, search indexes, caches, logs, exports and downstream systems. A system built without a deletion path has a permanent structural problem, because personal data spreads faster than anyone tracks. Make retention a property of the data. Defined at creation, enforced automatically, not dependent on a quarterly cleanup that never runs. Storage is cheap and that is precisely the problem: nothing forces the question. Log access to personal data. Not for compliance theatre — because it is how you answer the questions that follow an incident, and because it changes behaviour when people know access is recorded. Make the privacy choice legible. If a user cannot understand what they are agreeing to from the interface, the consent is weak in law and worse in practice.

Why it fails in practice

The most common failure is timing. Privacy review happens at the end, when the architecture is fixed and the release date is set, at which point the only available outcome is a list of exceptions with remediation dates that slip. Review has to occur when the means of processing are being determined — which is what the regulation says, and which is a statement about process sequencing rather than about privacy. The second is ownership. Where privacy belongs solely to legal or compliance, it arrives as a constraint imposed from outside and is resisted accordingly. Where an engineering lead owns it, it becomes a design property like performance or reliability. The organisations that do this well have made the shift. The third is the legacy estate. New systems can be designed correctly; the existing ones cannot be redesigned economically. The workable approach is to fix the highest-exposure retrofit — usually retention and deletion — and apply full standards only to new development, accepting a multi-year transition rather than pretending to a uniform position.

Test privacy properties in the systemQualitative engineering checks from the article. These do not replace a legal assessment or override retention obligations.
Design choiceCheck
CollectionExplain each field's specific purpose
DefaultsTest access, optional processing and visibility before user changes
RetentionName purpose-specific retention and applicable exceptions
DeletionTrace downstream copies, indexes and backup handling

Qualitative summary of this article's source text, not a measured outcome or performance estimate.

Practical Guidance for Privacy by Design Review

  • Put the assessment at design stage and make it a gate. A review after the architecture is fixed produces exceptions, not privacy.
  • Challenge every collected field individually. "It might be useful later" is the justification that creates most of the liability in most systems.
  • Test deletion end to end before launch. Through backups, replicas, warehouses, indexes, logs and downstream integrations — if it has not been tested, it does not work.
  • Set defaults to the most protective setting and justify any exception. Defaults are your real policy, whatever the policy document says.
  • Pseudonymise at storage where analytics is the use case. It reduces breach severity and makes secondary use defensible.
  • Attach retention to the data at creation and enforce it automatically. Manual cleanup schedules do not survive contact with operations.
  • Give an engineering owner accountability, with legal as adviser. Privacy owned outside engineering becomes an external constraint and gets treated as one.
  • Apply the full standard to new builds and triage the legacy estate by exposure. A realistic two-speed position beats an aspirational uniform one.

The Regional Dimension

Gulf organisations have a genuine advantage here that is frequently squandered. The advantage is timing. Regional data protection frameworks arrived later than Europe's, which means much of the systems estate being built now is being built after the requirements are known. Designing to a standard is dramatically cheaper than retrofitting to it, and organisations in the middle of ERP replacements, digital government programmes or new platform builds have an opportunity that European peers did not: to get it right the first time. Most do not take it, because the requirement is treated as a legal review at the end rather than an architecture input at the start. The local complications are specific. Employment here is linked to residency, which means HR and payroll systems hold passport data, visa records, medical information, dependant details, sponsorship documents and end-of-service calculations — a concentration of sensitive personal data with unusually high consequences for the individual if it leaks. These systems are often the oldest in the estate, frequently supplemented by spreadsheets held by a PRO or typing centre, and they are almost never in scope for a privacy review focused on customer-facing applications. Multi-entity structures create a second complication. A group with mainland entities, free zone entities, a DIFC or ADGM entity and operations in Saudi Arabia is applying several frameworks simultaneously, with different rules on transfer, notification and data subject rights. Designing each system to the strictest applicable standard is the only approach that remains maintainable, because per-entity variation in the same platform is a governance burden that grows with every change. Third, retention norms in the region lean heavily toward keeping everything. Commercial, tax and labour record-keeping obligations are real and often long, and the cultural default in family groups and government-adjacent entities is preservation. Privacy by design requires the opposite instinct, and the reconciliation is specific rather than general: retain what a named legal obligation requires, for exactly as long as it requires, and delete the rest — which means someone has to map the obligations properly instead of defaulting to indefinite storage. Finally, language. Consent notices, privacy statements and data subject request processes need to work in Arabic and English at minimum, and in practice for a workforce that reads neither comfortably. A privacy notice that is legally complete and practically incomprehensible to the people it describes fails the transparency principle even where it satisfies the lawyer.

The objection worth taking seriously

The serious criticism is that privacy by design, as practised, has become documentation rather than engineering. The assessment templates get completed, the register gets populated, the review meeting happens, the sign-off is recorded — and the system collects the same fields, retains them indefinitely and cannot delete them, because none of those artefacts changed a line of code. Compliance functions can demonstrate a process while the actual privacy properties of the estate remain unchanged, and the growth of privacy tooling has in places accelerated this by making the paperwork easier to produce. There is also a fair point about proportionality. The full apparatus — impact assessments, records of processing, design reviews, documented balancing tests — is a meaningful overhead, and it lands hardest on small organisations with no privacy function. A twenty-person company building a straightforward product ends up choosing between real engineering work and the documentation that proves it did the engineering work, and the documentation is what gets audited. The test worth applying is blunt: can the system delete a person on request, does it collect fields nobody uses, and are its defaults protective? An organisation that can answer those three honestly has more privacy than one with a complete assessment library and no deletion path. Documentation should follow the engineering, and where it substitutes for it, the programme has failed regardless of how good the file looks.

Common Questions

What does "by default" actually oblige?

That the out-of-the-box configuration limits collection, processing, retention and accessibility to what is necessary for the specific purpose — without the user having to do anything. A system that is privacy-protective only after a user changes settings does not satisfy it.

Which requirement is hardest to retrofit?

Deletion. Personal data propagates into backups, warehouses, search indexes, caches, logs and downstream systems, and retrofitting a propagation path across all of them is frequently more expensive than the original build. Design it first.

Does this apply to internal systems and employee data?

Yes, and it is the most commonly overlooked area — particularly in this region, where HR systems hold immigration and medical data with serious consequences for the individual. Internal systems get less scrutiny and often hold more sensitive information than the customer-facing ones.

How do AI systems change the analysis?

They stress every part of it. Training data is a copy that persists inside model weights in a form that cannot be selectively deleted, which is in direct tension with erasure rights. Embeddings and vector indexes are derived personal data living wherever the index lives. Inference logs retained by a provider are a processing location most data maps omit. And the purpose limitation principle sits awkwardly with a technology whose value proposition is finding uses nobody specified in advance. The practical response: decide before building whether personal data enters training at all — the cleanest answer is usually no, with retrieval over permission-aware source data instead, so deletion at the source actually takes effect. Treat prompts and outputs as processing with a retention period, and put model and index location in the data map alongside the database.


Privacy by Design Review — the test is whether the system can delete a person, not whether the assessment was completed; design the deletion path before the first record exists.

Continue reading

Talk to OPS

Start with the operating problem.