Cybersecurity / Source date:

Colonial Pipeline Ransomware: When Critical Infrastructure Falls

The 2021 Colonial Pipeline ransomware attack shut down fuel supply to the US East Coast and redefined critical infrastructure cybersecurity risk for every sector.

Illustrative pipeline operations scene with a dispatch schedule and flow meter, not the actual Colonial Pipeline incident.

The largest fuel pipeline in the United States restarted this weekend after six days down, and the most important detail of the Colonial Pipeline ransomware attack is the one least discussed in the coverage: nobody attacked the pipeline. The intrusion, attributed by the FBI to the DarkSide ransomware group, reached the corporate network. Reporting on the entry point points to a single legacy virtual private network account that was still live, still valid, and not protected by multi-factor authentication. From there the attackers encrypted business systems and took roughly a hundred gigabytes of data with the usual threat to publish it. The operator then shut the pipeline down itself, as a precaution, on 7 May. Fuel stopped moving to the eastern seaboard because a company could not trust its own billing and scheduling systems, not because anyone touched a valve. That distinction is the entire lesson, and most organisations are about to draw the wrong one from it.

The shutdown was a decision, not an effect

Separation between corporate information technology and operational technology is treated in most industrial businesses as a solved problem. There is a firewall, there is a demilitarised zone, there may be a data diode, and the control network is audited on that basis. Colonial's separation appears to have held in the network sense. It did not hold in the business sense, because the physical operation depends on corporate systems for things that are not optional: knowing how much product moved, whose product it was, what to bill, what to schedule next, and what the customers are entitled to. Running an unmeasured, unbilled pipeline for a week is not a technical impossibility. It is a commercial and legal one. So the honest continuity question for any asset-heavy business is not "is our control network segmented?" It is: if the corporate systems are unavailable or untrusted for seven days, what physically stops, why, and who decides? For most operators — pipelines, ports, utilities, distribution, manufacturing, logistics — nobody has ever written that answer down, and it is discovered at two in the morning by an executive who has to choose between running blind and stopping.

What the entry point tells you

A dormant remote access account with a valid password and no second factor is not an exotic failure. It is the most common finding in every assessment of every organisation of this size, and it has now produced a national fuel emergency. The practical work is unglamorous and can start this week. Enumerate every path into the environment from outside: named user VPN, vendor VPN, jump hosts, remote desktop gateways, management appliances, engineering workstations with dual-homed connections, the OEM's support tunnel, the account created for a project that finished in 2018. Then apply three tests to each. Is multi-factor authentication enforced? Is there a named owner who still needs it? Is authentication centralised, so that when a person leaves, access actually ends? The accounts that fail those tests are not a risk register item. They are the incident, waiting for a date.

On the ransom

Reports indicate a payment of around seventy-five bitcoin, worth roughly four and a half million dollars, was made within hours of the attack, and that the decryption tool provided was slow enough that restoration proceeded largely from the company's own backups anyway. If that holds, the payment bought optionality rather than time — which is the realistic case for paying and also the reason it so often disappoints. The decision itself is the part worth institutionalising. Ransom decisions made during an incident default to yes, because the people in the room are exhausted, the operational pressure is unbearable and the money is small relative to the outage. Decide it in advance instead, in writing: who is authorised to approve payment, what legal and sanctions review must happen first, which advisers and insurers must be on the call, what evidence is required that decryption will work, and what facts would make payment unlawful or unacceptable regardless of cost. A pre-agreed position does not commit you to refusing. It commits you to deciding on the merits rather than on adrenaline.

Practical Guidance for Critical Infrastructure Security Assessment

  • Map the business dependencies of physical operations, system by system, and state explicitly what stops without each one.
  • Write and rehearse the manual fallback, including measurement, scheduling and customer allocation, and establish the tolerance for operating without billing data.
  • Inventory every external access path, including vendor and dormant accounts, and enforce multi-factor authentication on all of them. Treat exceptions as incidents with expiry dates.
  • Practise the shutdown decision, not just the restart. Who can stop the asset, on what evidence, at three in the morning, and who must be informed within the hour.
  • Pre-decide the ransom position with legal, sanctions and insurance input, and record it at board level.
  • Keep backups offline and test restoration against the clock, because speed of restoration, not existence of backups, determines the outage length.
  • Plan the public and customer communication for an operational outage, since panic response can exceed the physical shortfall.
  • Brief the board on the dependency map rather than on the threat landscape. Directors can act on the first and cannot act on the second.
Prepare the operational decisions before an incidentThe article's continuity work sequence, not a reconstruction of Colonial's incident response or a measured risk forecast.
  1. Map the dependencies

    State which physical activities depend on each corporate system and what stops if it is untrusted.

  2. Rehearse the fallback

    Test measurement, scheduling and allocation without ordinary business-system availability.

  3. Review external access

    Inventory vendor and dormant accounts, their owners and authentication controls.

  4. Agree decision rights

    Name who can stop an asset and who must be informed outside business hours.

  5. Set the ransom decision process

    Agree authority and required legal, sanctions and insurance review before an event.

  6. Test recovery and communication

    Restore an offline backup against the clock and rehearse customer and public updates.

Qualitative summary of this article's source text, not a measured outcome or performance estimate.

The Regional Angle

Three features of Gulf critical infrastructure change the calculation, and none of them make the position easier. The first is the shortness of the buffer. In much of this region power and water are produced together, and desalinated water supply is held in storage measured in days rather than weeks. A fuel pipeline outage in a large continental market produces queues and price spikes; the equivalent interruption to a water and power complex produces a public emergency on a far shorter clock, in summer conditions where cooling is not a comfort issue. Any continuity analysis borrowed from North American energy assumes a tolerance that does not exist here. Run the dependency mapping exercise with the real clock: how many hours, not days, of manual operation are actually available before the physical consequence becomes unmanageable. The second is decision rights. A great deal of the region's critical infrastructure is state-owned, or a joint venture between a national entity and an international operator, or operated under a long-term concession. The consequence is that the two decisions this incident turned on — whether to shut the asset down, and whether to pay — are not the operator's to take alone. They involve a ministry, a regulator, a shareholder that is also a government, and in some cases a national cybersecurity authority with its own reporting requirements and its own view on disclosure. That is not a weakness, but it is a latency. The pre-agreed decision tree must therefore name the individual in each external body who can authorise a shutdown outside business hours, with the telephone number, the deputy, and the agreed maximum time to respond. Discovering the escalation path during the event costs hours that the buffer does not contain. The third is contractor-held access to the control estate. In most regional plants, the industrial control systems are maintained by the original equipment manufacturer or a specialist integrator under a long-term service agreement, and that agreement frequently includes standing remote access from outside the country for diagnostics and patching. The access is contractual, it is often documented only in the maintenance agreement rather than in any central inventory, and unilaterally disabling it can breach the contract or affect warranty and availability guarantees. This is the same class of exposure that appears to have started the Colonial incident, with an added commercial lock. The work is to find every one of these tunnels, list them alongside the contract clause that creates them, and renegotiate towards brokered, time-boxed, individually authenticated and recorded access at the next contract touchpoint — rather than leaving permanent foreign access into a national asset because it was easier for the vendor's support desk. One closing observation for regional boards: this market has direct experience of destructive attacks on energy companies, and the institutional memory is real. The gap is not awareness of the threat. It is that the awareness lives in the security function, while the dependency between the corporate systems and the physical asset lives in operations and finance, and the two conversations have never been held in the same room.

The objection worth taking seriously

The strongest objection is that this was not an industrial control systems incident at all, and that treating it as one will send a great deal of money to the wrong place. That objection is correct and the misdiagnosis is already underway: within days the marketing for operational technology security products has reframed itself around a pipeline that was never attacked at the control layer. An organisation that responds to this event by buying specialist industrial monitoring, while leaving dormant VPN accounts without multi-factor authentication, will have spent a budget and changed nothing. The measured response is that the ordinary hygiene came first: identity, remote access, privilege, backups, and knowing which accounts exist. Those controls would have addressed the entry point, and they address most of the incidents that will follow this one. Where the objection stops working is the shutdown. Even with perfect network separation, the operator stopped the asset, because the separation between business systems and physical operations is an engineering fact and an operational fiction at the same time. You can segment the network completely and still be unable to run the plant, ship the cargo or move the product when the corporate estate is untrusted. That is not solved by a security product in either domain. It is solved by mapping the dependency and deciding, in advance and in daylight, what you are prepared to operate without.

Common Questions

Would network segmentation have prevented the shutdown?

Apparently not. The reported segmentation held. The pipeline stopped because the company could not measure, schedule and bill with confidence, which is a business dependency rather than a network one.

Is multi-factor authentication really the headline control?

For externally reachable access, yes. A live account with a valid password and no second factor is the most consistently exploited weakness in this class of incident, and it is among the cheapest to remove.

Should critical infrastructure operators ever pay?

That is a decision for the board, taken before the event, with legal and sanctions advice on the table. What should never happen is an unprepared decision taken under operational duress, which is how most payments occur.

What should we expect over the next twelve months?

Expect mandatory security requirements for pipeline operators within weeks rather than years, and expect them to include incident reporting deadlines and named security officers; the United States has just issued a broad cybersecurity executive order and the political appetite for voluntary frameworks in this sector has evaporated. Expect ransomware to be handled as a national security matter, with pressure on cryptocurrency payment channels and a real possibility of sanctions action against the groups and their financial infrastructure. Expect the criminal ecosystem to respond tactically — the group behind this attack has already claimed it will vet targets to avoid social consequences, and its infrastructure has reportedly gone quiet in the last few days — while the same operators reappear under new names. Expect insurers and boards to demand evidence of multi-factor authentication coverage this quarter rather than a policy statement. And expect at least one more attack on a physical supply chain — food, freight or utilities — before the year is out, because the last fortnight has demonstrated exactly how much leverage the model produces.


Critical Infrastructure Security Assessment — we map the dependencies between your corporate systems and your physical operations, test the access paths that create them, and rehearse the decisions you will otherwise take at three in the morning.

Continue reading

Talk to OPS

Start with the operating problem.