A team that has never threat modeled makes obvious mistakes. They skip it entirely, or they run one workshop under duress and produce a document nobody reads. Those failures are visible and easy to diagnose from the outside, which is why most of the advice aimed at them — how to run your first session, how to get the room to take it seriously — works.
The mistakes in this article are different. They belong to teams that have been doing this for a year or three, who know STRIDE cold, who have a working DFD and a real risk register, and whose practice has quietly stopped improving anything. None of these seven are knowledge gaps. They're what fluency looks like once it curdles into habit — and they're much harder to see from inside the team making them, because the process still looks correct on every dimension anyone thinks to check.
The short version: A beginner's threat model fails because nobody in the room knows what to do. An experienced team's threat model fails because everyone in the room already knows what to do, does it correctly, and stopped asking whether "correctly" still means "usefully." Fluency and rigor are different skills, and only one of them shows up in a review of whether the workshop was run right.
1. The model that never leaves the workshop
The session happens. The diagram gets drawn, the threats get listed, the register gets filled in, everyone leaves satisfied. Then the system changes — a new endpoint, a config flag that quietly became permanent, a "temporary" service that's been in production for a year — and the model doesn't. Nobody updates it because updating it was never anyone's job; running the workshop was.
Experienced teams fall into this precisely because their workshops are good. A well-run session produces a document that looks authoritative, and an authoritative-looking document is exactly what stops getting questioned. Six months later it's cited in an audit as evidence of a security review that describes a system that no longer exists in the ways that matter.
The tell: ask when the diagram was last checked against the actual deployed system, not when it was last opened. If nobody can answer without guessing, the model is documentation of a decision that was made once, not a description of the system today. What should trigger a refresh, and how often, is its own question — worth answering deliberately rather than defaulting to "never" by omission.
2. STRIDE as a mad-libs exercise
Run every element on the diagram through all six STRIDE categories, record something for each, move to the next element. It feels rigorous — six checkboxes per node is a defensible, auditable process — and it's exactly how STRIDE gets taught, so experienced teams do it fluently and fast.
What it produces is a list where every element has entries for categories that don't actually apply to it, sitting at equal weight next to the two or three threats that are genuinely dangerous. A read-only reporting service gets a Repudiation entry because the template has a slot for one, not because anyone believes an attacker cares. The list is technically complete and substantively flat — and a flat list is worse than a short one, because it looks like coverage while actually diluting the signal a reviewer needs to act on.
This is a different failure from having too many threats and no way to rank them — that's a triage problem, downstream of generation. This one happens earlier: judgment about plausibility never gets applied at the point each threat is written down, so triage inherits a list that's already been inflated by rote coverage before anyone starts ranking it.
The tell: every model the team produces is roughly the same length, regardless of whether the system underneath is trivial or genuinely complex. A process that always returns the same amount of output regardless of the input is measuring the process, not the system.
3. Modeling the architecture diagram instead of the system
The DFD reflects the design document. The design document reflects the system as it was proposed. The system in production has a manual runbook step nobody diagrammed, a hotfix that added a new data path under deadline pressure, and a feature flag that's been "temporarily" routing a subset of traffic around the normal authorization check for four months. None of that is in the model, because the model was drawn once, from the plan, and the plan was never the thing that shipped.
This is a subtler version of mistake #1 — not that the model wasn't updated on a schedule, but that it was never actually anchored to the running system in the first place. A team can religiously revisit the model every quarter and still be revising a description of the wrong thing, if what gets revised each time is the diagram's idea of the architecture rather than an inspection of what's actually deployed.
The tell: pick any three elements on the diagram at random and ask an engineer who works on the system day-to-day whether each one still exists exactly as drawn. If the answer is "mostly," the model is modeling the plan.
4. A checklist instead of an adversary
"What could go wrong with this element?" is a passive question, and it's the one most templates actually ask. Experienced teams answer it fluently — for every node, for every flow, for every boundary — and produce a list of things that could theoretically happen. What that framing quietly drops is the active version of the question: if I were trying to get at this specific asset, what would I actually try, and what would I do next if the first thing worked?
The difference shows up in depth, not breadth. A checklist produces a flat list of independent, restated risks. An adversarial frame produces a chain — this weakness gets me here, from here I can reach that, and combining those two individually-low findings gets me the thing that actually matters — which is exactly the structure an attack tree is built to hold and a flat list isn't. Teams that have used STRIDE long enough to run it from memory are the ones most at risk of never rebuilding that adversarial habit, because the checklist version produces output that looks like the same deliverable.
The tell: look at the tree's depth, not its width. A model with forty shallow, single-step threats and no multi-step paths usually means nobody in the room spent any time thinking like the person on the other side of it.
5. Modeling only the happy path
The diagram shows the request flowing through in the normal case: authenticated, authorized, data available, downstream service healthy. What it usually doesn't show is what happens when any of that isn't true — the cache is stale, the primary database is unreachable and a read replica serves the request instead, a circuit breaker trips and a fallback path activates, a queue backs up and a consumer starts processing out of order. Those degraded states are exactly where security assumptions from the happy path quietly stop holding, and exactly where an attacker who understands the system looks first, because degraded-state code paths get far less scrutiny than the primary one.
A fallback that serves stale cached data might serve another tenant's stale cached data. A retry path might replay a request without the idempotency check the primary path enforces. A "fail open" decision made by an engineer trying to preserve availability during an incident might be a standing authorization bypass that nobody threat modeled because it only exists in a state the diagram never depicted.
The tell: ask what the model says about the system's behavior when its single most important dependency is down. If the answer is nothing, the model has an opinion about one operating state and no opinion about all the others — including the one the system is in during an actual incident, which is also usually the state an opportunistic attacker is most likely to be probing.
6. The invite list never changes
The same three or four engineers run every session, because they're the ones who know the process, the tooling, and the system best. This is efficient and it is also how a blind spot becomes permanent: the room's shared knowledge is the model's ceiling, and if nobody in the room has ever operated the system at 3 a.m., worked a support ticket about the one weird edge case, or built the internal tool the rest of the team doesn't think about, that entire category of insight never enters the model, session after session, indefinitely.
This compounds with mistake #3. The people who'd notice that the diagram no longer matches reality are often not the same people who are fluent enough with STRIDE to have been invited to run the workshop efficiently in the first place.
The tell: pull the attendee list from the last five sessions. If it's the same five names, the process has optimized for speed at the cost of the one input — a different pair of eyes — that a review process exists to provide.
7. Identified threats become permanent tenants
A threat gets written down, discussed, maybe assigned a rough severity — and then sits in the register indefinitely, with no owner, no target date, and no scheduled re-review. It isn't dismissed as a false positive and it isn't fixed; it's simply present, the way a browser tab that's been open for three weeks is present. Eventually a register like this becomes indistinguishable from one where every open item was a deliberate accepted risk, except nobody actually decided to accept anything — the outcome just happened by neglect instead of by choice.
This is a different problem from a mitigation being too vague to be testable — that's about the quality of what's written. This is about whether anything written gets tracked to a close, regardless of how well it was worded. A team can write excellent, specific mitigations and still let half of them sit unowned for a year, because writing the mitigation and being accountable for shipping it are two different disciplines, and only the first one happens inside the workshop.
The tell: pick the oldest unresolved item in the register and ask who owns it. "Nobody, currently" is the honest answer more often than teams expect, and it means the register has quietly become a list of things the team has decided not to decide about.
What these seven have in common
None of them is a gap in anyone's understanding of threat modeling. Every team making these mistakes could pass a quiz on STRIDE, explain what a trust boundary is, and run a technically correct workshop end to end. What's missing in each case is a second-order habit — checking that the first-order process is still pointed at something real, still calibrated by judgment rather than rote coverage, still hearing from more than the same five people, still tracking its own output to a close. Those habits don't show up in whether the framework was applied correctly, which is exactly why a team can look rigorous by every visible measure and still be running a hollowed-out version of the practice.
If you only fix one, fix #1 first. A stale model quietly invalidates the other six — there's no point sharpening the STRIDE judgment (#2), inviting new perspectives (#6), or tightening register ownership (#7) on a diagram that no longer describes the system everyone is trying to protect.
Where this list itself goes wrong
Two honest caveats, in the spirit of not overselling a framework any more than we'd want a team to oversell their process:
- Running through this list mechanically is itself an instance of mistake #2. Treating "have we checked all seven?" as a compliance pass — six boxes ticked, no judgment about whether a given item actually applies to how your team works — reproduces the exact failure mode the list is describing. The point isn't the checklist; it's noticing which of these your team's own history makes plausible.
- Not every mature team has all seven, and severity varies by context. A five-person startup team modeling one product doesn't have the "invite list never changes" problem in the same way a two-hundred-engineer org with twelve independent squads does — there's only one team to invite. Read this as a set of questions to ask about your specific practice, not a diagnosis handed down from outside it.
If your organisation is regulated, a stale or checklist-driven model is a specific audit liability — see the industry threat modeling guides for what applies in fintech, healthcare, and critical infrastructure. The free templates are a reasonable place to restart a model that's drifted too far from the running system to trust.