Ask most teams how they model ransomware risk and the answer centers on the moment of infection: phishing training, endpoint detection, email filtering, maybe a tabletop exercise about whether to pay. All of that is worth having. None of it explains why three of the most consequential ransomware incidents on record turned out the way they did, because in each case the infection itself was almost incidental — a stolen password, a vendor's software update, an exploited management tool. What decided the actual outcome was architecture that had been decided months or years earlier, that nobody was thinking about as "ransomware defense" at all.
The short version: Colonial Pipeline shows that a single credential's reach matters more than how it was stolen. Kaseya shows that the tools you exclude from your threat model because they're "just IT admin software" are often exactly the ones with the broadest privileged access. Maersk shows that "we have backups" is not a plan if the backup infrastructure can be destroyed by the same event that took down everything else. None of these are lessons about ransomware specifically — they're lessons about blast radius that ransomware happens to make visible, brutally and all at once.
Misconception 1 — "Ransomware is a malware problem, so the endpoint is where we defend"
On May 7, 2021, Colonial Pipeline shut down the largest fuel pipeline in the United States after discovering a ransomware infection from the DarkSide group. The attackers got in through a single compromised VPN account — a legacy account, not protected by multi-factor authentication, using a password that had likely been reused and had leaked in an unrelated breach. No malware was needed to get in the door; a password was enough. Once inside, the attackers spent roughly two hours exfiltrating about 100 gigabytes of data before deploying ransomware on the corporate IT network, then threatened to publish the stolen data on top of the encryption — the now-standard double-extortion model.
The pipeline itself, the operational technology that actually moves fuel, was never confirmed to be infected. It was shut down anyway, out of precaution, because Colonial could not be confident — in real time, under pressure, during an active incident — that the IT and OT networks were actually separated enough to guarantee it. That uncertainty, not a confirmed OT compromise, is what caused the six-day shutdown, the regional fuel shortage, and the panic buying that followed. Colonial paid a 75-bitcoin ransom worth about $4.4 million; the Department of Justice later recovered roughly 64 of those bitcoins, worth about $2.3 million by the time of recovery.
The mistake wasn't a missing antivirus signature. It was treating "one password" and "the fuel supply of the East Coast" as separated by a boundary nobody had actually tested. A threat model that draws IT and OT as genuinely separate trust boundaries has to ask a harder question than "are they separate" — it has to ask "how would we know, quickly and with confidence, if that separation held during an actual incident," because the six-day shutdown was caused by the absence of a confident answer to that question, not by a confirmed breach of it.
The design question that was skipped: "If this credential were compromised, could we prove — during the incident, not after it — what it couldn't reach?"
Misconception 2 — "Our management tools aren't part of the attack surface, they're just IT admin software"
On July 2, 2021, the REvil ransomware group exploited a chain of zero-day vulnerabilities — including an authentication bypass and an arbitrary command execution flaw, tracked as CVE-2021-30116 — in Kaseya VSA, a remote monitoring and management platform that managed service providers use to administer their clients' networks. REvil pushed ransomware through roughly 60 MSPs to their downstream customers, causing outages at more than 1,000 businesses that had never heard of Kaseya and had no direct relationship with it at all. REvil demanded $70 million for a universal decryptor; Kaseya eventually obtained a decryption key from an unnamed "trusted third party" three weeks later, without ever publicly confirming whether a ransom was paid.
The businesses affected didn't fail to threat model their own applications. They failed to threat model a tool that was, by design, invisible to them — installed and operated entirely by their MSP, with the specific purpose of having broad, privileged, remote reach across every managed endpoint. That's not an oversight so much as a structural blind spot: a tool built for IT operations gets treated as part of the trusted baseline precisely because it's operational rather than customer-facing, when operational reach is exactly the property that makes it valuable to an attacker who compromises it once and inherits access to everything it manages.
This is the same shape of mistake our vendor integration post describes — you can't threat model a vendor, only what the integration reaches — applied to a category of vendor relationship most companies don't even register as an integration at all, because nobody signed a data-sharing agreement for it. An MSP's remote management tool isn't sending you data. It's holding the keys.
The design question that was skipped: "What tool, installed by someone else, has the broadest remote reach into our infrastructure — and is it in our model at all?"
Misconception 3 — "If it comes to it, we'll pay, or restore from backup"
On June 27, 2017, malware disguised as ransomware — now known as NotPetya — spread through a compromised update to M.E.Doc, an accounting application used throughout Ukraine, then propagated internally using EternalBlue, a leaked NSA exploit for a Windows networking flaw, combined with Mimikatz, a tool that extracts credentials directly from memory. At global shipping company Maersk, that combination was enough to destroy essentially every domain controller the company had, worldwide, wiping out the Active Directory infrastructure that the rest of the network depended on to function at all. The ransom note demanded $300 in bitcoin per machine — a figure so low, relative to the damage, that it should have been the first sign this wasn't a real extortion attempt. NotPetya had no functioning decryption mechanism behind the payment. It was a wiper wearing a ransom note.
Maersk recovered only because one domain controller, in a small office in Ghana, happened to be offline during the outbreak due to a power cut, and so was never infected. Restoring from it required physically relaying a hard drive — flown from Ghana through Nigeria, hand to hand, because the Ghana office had neither the bandwidth to transmit a several-hundred-gigabyte backup nor staff with the visas to fly directly to the UK — to Maersk's technology center outside London. The company rebuilt its global infrastructure in about ten days and later put the total cost at $250–300 million; a White House assessment attributed roughly $10 billion in damage to NotPetya globally, across every organization it hit.
Two assumptions failed at once here, and most ransomware plans only account for one of them. The first is "if it comes to it, we'll pay" — which assumes a working decryption path exists, and NotPetya specifically didn't have one. The second is "if paying doesn't work, we'll restore from backup" — which assumes the backup and identity infrastructure survives whatever took down production, and on a flat enough network, it doesn't; the same lateral movement that reaches your file servers reaches the domain controllers your restore process depends on. Maersk's recovery is a story about a lucky power outage, not about a plan that worked as designed.
The design question that was skipped: "Does our backup and identity infrastructure sit inside the same blast radius as everything it exists to restore?"
What the three cases actually share
| Colonial Pipeline (2021) | Kaseya (2021) | Maersk / NotPetya (2017) | |
|---|---|---|---|
| How the attacker got in | Single VPN password, no MFA | Zero-day chain in a remote management platform | Compromised update to third-party accounting software |
| What actually determined the damage | Inability to prove IT/OT separation during the incident | The management tool's own broad, privileged reach across every managed endpoint | Flat internal network let a wiper reach every domain controller worldwide |
| Assumption that failed | "Ransomware is a data problem" — the real cost was an availability decision made under uncertainty | "This isn't part of our attack surface, it's just IT tooling" | "We'll pay or restore" — no working decryptor existed, and backups shared the blast radius |
| Confirmed cost | $4.4M paid, ~$2.3M recovered by DOJ, 6-day shutdown | $70M demanded; resolution terms undisclosed | $250–300M to Maersk; ~$10B globally per White House estimate |
Notice what isn't in that table: the ransomware family, the malware's sophistication, or how patched the victim's endpoints were. In all three cases, the entry technique was almost beside the point — swap DarkSide for any other credential-stealing group, swap the Kaseya zero-day for any other privileged remote-access tool, swap M.E.Doc for any other trusted software update, and the outcome is governed by the same three questions: how far does one compromised thing reach, is a high-privilege tool actually in the model, and does the recovery path survive the event it's supposed to recover from. Those are architecture questions, decided long before any ransom note existed, and a threat model that only shows up after an incident to ask "which strain was this" is asking the least useful question available to it.
What a ransomware-aware model actually asks
- Trace one compromised credential to its actual reach — not its intended reach, its actual reach, including any legacy or unused account nobody has gotten around to disabling.
- Put every high-privilege management tool on the diagram, whether you installed it or a vendor did, and treat its blast radius as a first-class node — precisely the property Kaseya's downstream victims never had visibility into at all.
- Model backup and identity infrastructure as inside the same system, not outside it. If the thing you'd restore from can be destroyed by the same lateral movement that takes down production, it isn't a separate contingency — it's one more thing in the blast radius.
- Ask whether a segmentation claim could survive being tested under pressure. Colonial's IT/OT boundary may or may not have held technically; what mattered was that nobody could demonstrate it held quickly enough to avoid a precautionary shutdown.
- Treat "we'll pay" as a plan with a failure mode, not a backstop. NotPetya is the standing counterexample to the assumption that a ransom note implies a functioning decryption path behind it.
Where this argument overreaches
The same caveats this series applies everywhere else apply here too:
- None of these companies necessarily skipped threat modeling. The article knows the outcomes, not each organization's internal process at the time, and attributing intent in hindsight is a claim the evidence doesn't support. The architecture is a fact regardless of what process, if any, produced it.
- A perfect model doesn't eliminate the credential-theft or zero-day problem. Colonial's stolen password and Kaseya's zero-day chain are entry-vector problems a design-time review is not built to close on its own — patching, credential hygiene, and MFA remain necessary, ordinary controls that a threat model complements rather than replaces.
- Three incidents are illustrations, not a representative sample of the much larger and more mundane population of ransomware incidents most organizations actually face, which are less dramatic and less well documented precisely because they didn't shut down a pipeline or a global shipping company.
Availability requirements around ransomware and business continuity show up explicitly in several compliance regimes — see the industry threat modeling guides for what applies in fintech, healthcare, and critical infrastructure. The free templates include a pre-labelled DFD if you want to trace your own backup and identity infrastructure's actual blast radius this week.