Cybersecurity / Source date:

Anthem and OPM Breaches: Identity Data Becomes the Prize

Health and clearance records showed that non-financial identity data carries decades-long exposure.

Illustration of concentrated identity records beside a locked archive cage, not an Anthem or OPM facility.

The Anthem and OPM breaches of 2015 changed what a data breach meant. Until then, the template was financial: cards stolen, fraud attempted, cards reissued, cost absorbed by the payment system. These two incidents involved data that cannot be reissued, and the attackers appeared to want it for reasons that had nothing to do with fraud. Anthem disclosed in early February 2015 that unauthorised access had been discovered on 29 January. The initial figure was around 37.5 million records; within weeks it was revised to roughly 78.8 million people. The exposed data included names, dates of birth, social security numbers, member identification numbers, addresses, employment details and income data. Not payment cards — identity. The Office of Personnel Management disclosures ran through mid-2015 and were, in their own way, worse. An initial announcement in June concerning roughly four million personnel records was followed in July by a second breach affecting 21.5 million individuals. The compromised material included background investigation records covering essentially every federal security clearance investigation conducted since 2000, along with approximately 5.6 million fingerprint records. Background investigation files contain the most sensitive information a person can be asked to disclose: financial difficulties, foreign contacts, medical history, relationships, past conduct. The agency's director resigned.

People affected in the reported breach populationsLater confirmed figures, not figures available on the article's May 2015 source date. OPM personnel and background populations overlap; do not add the bars. Anthem figure is cited in the 2019 indictment.

Million people. Bars start at zero.

Anthem
78.8 Million people
OPM background records
21.5 Million people
OPM personnel records
4.2 Million people

Why these breaches were categorically different

A stolen payment card has a short useful life and a well-developed remediation path. Someone notices anomalous spending, the card is cancelled, a new number is issued, and liability is allocated between parties who have contracts covering exactly this scenario. Social security numbers, dates of birth and background investigation records have none of those properties. They cannot be reissued in any practical sense. They do not expire. And their value to an adversary is not immediate monetary gain — it is the ability to impersonate, to target, to recruit, and to build a picture of who has access to what. That difference broke the standard response. Credit monitoring, the reflexive remedy offered to affected individuals in both cases, is designed to detect financial fraud. It does precisely nothing about a foreign intelligence service holding your security clearance file, and offering it as remediation revealed how poorly the incident response playbook of the time matched the actual harm. The attribution discussion reinforced the point. Both incidents were widely assessed as state-linked espionage rather than criminal monetisation, and in the Anthem case a Chinese national was charged years later. Whether or not any particular attribution is accepted, the operational implication holds: some categories of data are targeted by adversaries who are patient, well-resourced, and not deterred by the absence of a resale market.

What the technical accounts had in common

The two intrusions were different in detail and similar in shape. Both involved credential compromise as the entry mechanism rather than an exotic technical exploit. In the Anthem case, the attackers reportedly used the credentials of a number of employee accounts to move through the environment and reach systems including a data warehouse. In the OPM case, contractor-associated credentials featured prominently in subsequent accounts. Neither breach was primarily a story about a vulnerability that a patch would have closed. Both involved extended dwell time. The attackers were present for months before detection, which meant the compromise was not a smash-and-grab but a sustained presence with time to locate the data that mattered. Both involved data at rest in large, consolidated repositories — exactly the systems that get built for legitimate analytical reasons and then become the single richest target in the estate. And in both cases, detection came late and from a direction other than the primary monitoring. Anthem reportedly noticed when a database administrator saw a query running under their own credentials that they had not initiated — which is a good outcome achieved by an alert human rather than by a control.

The lesson most organisations took, and the one they should have

The lesson widely drawn was about encryption, because Anthem's data at rest was reported not to have been encrypted and the point was easy to make in headlines. It is not a bad lesson, but it is a shallow one. In an intrusion where the attacker holds valid credentials and operates through the application layer, database encryption at rest protects against theft of the physical media and very little else. The credentials decrypt the data by design. The more durable lessons are less quotable. Identity is the control plane. If credential compromise is the entry route, then multi-factor authentication, privileged access management, credential rotation and least privilege matter more than perimeter investment. Both incidents were, at root, authorisation failures. Aggregation creates concentration risk. A data warehouse containing member records for eighty million people is a business asset and a strategic liability, and the decision to build one is rarely reviewed as a security decision. Detection should focus on data access patterns, not just on network events. A query that pulls millions of records is anomalous regardless of whose credentials issued it, and that is an observable signal. And third parties expand the identity perimeter. Contractor access, vendor accounts and business associate relationships are all authentication paths into the environment, governed by contracts that frequently say less about access control than they should.

Practical Guidance for Sensitive Data Protection Review

  • Inventory where irreversible identity data actually lives. National identifiers, dates of birth, background and vetting records, biometric data, health information. Most organisations find copies in test environments, exports and reporting databases that nobody had catalogued.
  • Treat aggregation as a risk decision requiring approval. Every large consolidated repository of personal data should have a named owner, a documented purpose, and a review of whether the full population and full field set are genuinely needed.
  • Put multi-factor authentication on every administrative and remote access path, without exception. Both of these breaches turned on credentials. This remains the highest-return control available.
  • Monitor bulk data access as a first-class alert. Volume, unusual query patterns, access outside normal hours, and administrative accounts reading production data. Set thresholds and route them to someone who will act.
  • Extend access governance to contractors and vendors on the same terms as employees. Separate accounts, least privilege, time-bounded access, and immediate revocation at contract end — verified rather than assumed.
  • Minimise and expire what you hold. Retention schedules enforced in systems rather than written in policy documents. Data that has been deleted cannot be exfiltrated.
  • Plan a response for breaches of irreversible data specifically. Credit monitoring is not a remediation for identity or vetting data. Decide in advance what you would actually offer, and what you would tell people.
  • Rehearse the notification decision with legal, communications and the board. Both of these cases show how the disclosure timeline and the revision of victim counts become the story.

The Regional Angle

For organisations in the Gulf, the relevant question is which datasets in their own environment have the irreversibility property, and the list is longer than the payment-card framing suggests. Emirates ID and national identity numbers, passport and visa records, residency and labour file data, and — for employers — the documentation held during onboarding for expatriate staff all fall into this category. A regional employer of any size holds passport copies, visa records, family details and sometimes educational and police clearance documentation for thousands of people across dozens of nationalities. That file set is closer to a background investigation record than to a customer database, and it is frequently stored with far less protection because it is treated as HR administration rather than as sensitive data. The structural aggravating factor is intermediation. PRO and government-relations agents, typing centres, recruitment agencies, medical testing providers and relocation firms all handle identity documents as part of routine processes. Each is a third party holding irreversible data about your employees, usually under a commercial arrangement that says nothing about security. Mapping that chain is an uncomfortable exercise and an overdue one for most regional organisations. Regulation has moved in the same direction. The UAE's federal data protection framework, Saudi Arabia's personal data protection regime, and the separate regimes operating in DIFC and ADGM all create obligations around sensitive personal data, breach notification and cross-border transfer. Sector regulators in financial services and healthcare add their own requirements. An organisation that has not classified its personal data holdings cannot demonstrate compliance with any of them. The healthcare parallel is direct. Regional insurers and providers hold exactly the data category that made the Anthem breach significant — identity plus health information plus employer linkage — across large expatriate populations whose records are shared between employer, insurer, third-party administrator and provider network. The number of parties holding a copy is the thing to examine.

The honest limitation

The uncomfortable truth in both cases is that a well-resourced, patient adversary with a specific target will frequently succeed, and that the controls being recommended afterwards would have raised the cost of the attack rather than prevented it. Both organisations had security programmes, budgets and staff. Neither was negligent in the way the post-incident commentary implied. What they faced was an attacker willing to spend months inside the environment, using valid credentials, pursuing data of national-scale value. Against that, the realistic objective is not prevention — it is reducing dwell time, limiting how much any single compromised identity can reach, and ensuring that the highest-value aggregation is the hardest thing in the estate to access at volume. There is a corresponding honesty required about remediation. If a dataset of irreversible identity data is lost, there is no fix. The individuals affected carry the consequence permanently. That should change the calculation about whether to collect and retain it in the first place, and it rarely does, because the collection decision is made by a different function than the one that would handle the breach.

Common Questions

Would encryption have prevented the Anthem breach?

Probably not, given the reported use of valid credentials. Encryption at rest defends against theft of storage media and improper access at the infrastructure layer; it does not defend against an attacker authenticating as a legitimate user through the application. Encryption remains worth doing — it just should not be presented as the lesson of this incident.

How do we identify our equivalent of a background investigation file?

Look for datasets where the information cannot be changed if it leaks and where the combination is more sensitive than any single field. Employee onboarding files, vetting and clearance records, health data linked to identity, and biometric templates are the usual candidates. The test is simple: if this were published, could the affected person do anything about it?

What should we tell people if irreversible data is exposed?

The truth, early, with specifics about what was taken and what it enables — and with remediation that matches the harm. Credit monitoring where financial fraud is plausible; explicit warnings about targeted phishing, impersonation and social engineering where the data supports those attacks. Revising victim counts upward in public over several weeks, as happened in both of these cases, does more reputational damage than the initial disclosure.

How does AI change the exposure?

It raises the value of exactly this data to an attacker and multiplies the number of places it sits. Large personal datasets are now also training inputs, embeddings and prompt context, which creates derived copies that are difficult to inventory and effectively impossible to delete from. Any sensitive data review conducted today needs to cover what has been fed into AI systems, which vendors process it, and whether anyone can answer where the derived representations live.


Sensitive Data Protection Review — start by listing the data you hold that cannot be reissued if it is stolen, because that list is the one your incident response plan was probably never written for.

Continue reading

Talk to OPS

Start with the operating problem.