How to Threat Model a Multi-Tenant SaaS Application

Most application threat modeling asks "how does an attacker get in." A multi-tenant SaaS application has a second, equally important question that generic guidance rarely names directly: once someone is legitimately in — as themselves, with a valid account, no exploit required — can they ever end up looking at another customer's data? Authentication answers who the user is. It says nothing about whether the query the app runs next is actually scoped to that user's own tenant. That scoping is the boundary this post is about, and it's a different kind of boundary than the network or trust boundaries most DFDs already draw.

The core question: not "can an attacker get in," but "can Tenant A's authenticated, legitimate request ever touch Tenant B's data or resources." Every threat below is a different way that question ends up answered "yes" by accident.

Isolation models, and what each one costs you

Tenant isolation sits on a spectrum, and where a given system sits on it changes which threats are worth spending time on:

  • Shared everything, logical isolation. One database, one application tier, tenants separated purely by a tenant_id column enforced in application code. Cheapest to run, easiest to scale, and entirely dependent on every single query getting that scoping right — there is no infrastructure backstop if the code forgets.
  • Schema-per-tenant. Same database server, separate schema per tenant. A forgotten WHERE tenant_id = ? can't leak across schemas the way it can across rows, but a connection using the wrong schema still can.
  • Database-per-tenant. Each tenant gets a full database. Meaningfully harder to leak across by accident, at real operational cost — migrations, backups, and connection pooling all multiply by tenant count.
  • Fully siloed infrastructure. Separate compute and storage per tenant, sometimes per-tenant deployments entirely. The strongest isolation available and the most expensive; usually reserved for a handful of regulated or high-value customers layered on top of a shared platform for everyone else.

None of these is universally "correct" — a five-person startup with a shared-everything model and disciplined query scoping can be perfectly reasonable; a healthcare platform sharing a database across covered entities without at least schema-level separation is a harder sell. The threat model's job is to make the trade-off explicit rather than accidental, and to concentrate scrutiny where the chosen model actually depends on code being right every time.

Where the boundary actually breaks

In a shared-everything model — the common case, and the one where this matters most — cross-tenant leakage almost never happens in the main, most-reviewed authenticated API path. It happens in the places that path's discipline doesn't automatically reach:

  • Background jobs and queues. A job enqueued during a request often carries only an entity ID, not the full tenant context — if the worker that dequeues it doesn't re-derive and re-check the tenant, it can act across tenant lines with nothing to stop it.
  • Shared caches. A cache key built from a resource ID alone, without a tenant prefix, can serve Tenant A's cached response to Tenant B's identical-looking request.
  • Shared search indexes. A single Elasticsearch or similar index across all tenants is fast and simple to run — and every query against it needs an explicit tenant filter, since the index itself enforces nothing.
  • Object storage paths. Predictable or sequential file paths/keys in shared object storage let one tenant guess or enumerate another's files, independent of whatever access control the application layer thinks it's enforcing.
  • Internal admin and support tooling. Built for staff who legitimately need to look across tenants, and easy to leave with far broader access than any single support ticket needs — and a frequent target if an attacker compromises a staff account instead of a customer one.

The common thread: each of these sits outside the one code path that usually gets the most security attention. A query-scoping bug in the primary API is the kind of thing a code reviewer is trained to look for. The same bug in a queue worker or a cache-key builder usually isn't, because it doesn't look like the "main" security-relevant code at all.

Noisy neighbors are a security threat too

Multi-tenancy also reopens the third leg of the CIA triad that data-leakage discussion tends to crowd out: availability. If one tenant's workload — a large export, a runaway report, a script hammering an API endpoint within its documented rate limits — can measurably degrade service for every other tenant sharing the same database connection pool, queue, or compute, that's an exploitable denial-of-service path, not just a capacity-planning problem. It deserves the same attack-tree treatment as any other availability threat: what does it take to trigger, and what isolates the blast radius (per-tenant rate limits, resource quotas, separate queues per tier) when it happens.

Modeling this in a DFD and attack tree

The tenant boundary doesn't map cleanly onto a trust boundary drawn once on a DFD, because it isn't a fixed network or infrastructure line — it's a property that has to hold true on every single data access, re-evaluated per request. Draw it anyway, but be explicit about what it represents: a boundary around each data store and process that the current request's tenant context has to be checked against before crossing, not a boundary a request only crosses once.

For the attack tree, put the goal in exactly those terms — "Access Tenant B's data as an authenticated Tenant A user" — and branch by the failure modes above: query missing a tenant filter, background job losing tenant context, cache key collision, search index leak, storage path guessing, over-privileged admin tooling. Each branch is independently testable (can you actually reproduce a request that skips the tenant check here?) and independently scoreable, which is exactly the shape a microservices threat model already uses for the analogous "which service can reach what" question — multi-tenancy is that same discipline applied per-request instead of per-service.

Common mistakes

  • Treating authentication as isolation. Confirming who the user is says nothing about whether the next query is scoped to their tenant. These are two different controls, and only one of them is tenant isolation.
  • Auditing only the primary API path. The isolation bugs that actually ship live in queues, caches, search, and admin tooling — code that rarely gets the same review discipline as request handlers.
  • No default-deny at the data layer. Relying entirely on every developer remembering the tenant filter, with nothing (a query builder default, a database-level policy, a middleware check) failing closed when someone forgets.
  • Ignoring resource contention as a security threat. A noisy-neighbor problem that's exploitable on purpose is a denial-of-service path, not just an SRE concern.

None of this argues for the most expensive isolation model by default — siloed infrastructure for every tenant is rarely the right call for a product still finding its market. It argues for knowing exactly where your chosen model's isolation actually depends on code being correct, and pointing the threat model at those specific places instead of re-auditing the login flow for the tenth time.

This is one of the areas the OWASP Top 10's Broken Access Control category covers most directly -- see that post for how tenant-scoping bugs fit the broader access-control category IDOR belongs to.

Map exactly where your tenant boundary depends on code being right

ThreatTree's attack trees let you score every path to cross-tenant access separately -- query scoping, background jobs, caching, search, admin tooling -- so isolation gaps show up as a risk-register entry, not a surprise.

Get started free

Not ready to sign up? Get new threat-modeling guides by email instead.