A ship's hull isn't one big open hold. It's divided into sealed bulkhead compartments, each watertight, so that if the hull is breached, water floods only that one compartment instead of the entire vessel. The ship stays afloat because the damage is contained.
The bulkhead pattern borrows that idea for software. Instead of letting every request draw from one shared pile of resources, you wall those resources off into separate compartments — so that when one springs a leak, it can't sink the whole application.
The problem
Most services run on a single shared pool of finite resources: a thread pool, a connection pool, a fixed set of instances. As long as everything is healthy, that's efficient — whoever needs capacity grabs it and hands it back a few milliseconds later.
The trouble starts when one dependency turns slow rather than dead. A dead dependency fails fast and gives the thread straight back. A slow one keeps the thread for as long as the call hangs — often 30 seconds or more, until a timeout fires. The arithmetic is brutal: a dependency you call 4 times a second that suddenly takes 30 seconds to answer needs 4 × 30 = 120 threads just to keep up. A pool of 12 is gone in 3 seconds, and from then on there's nothing left to serve any request — even the ones that never touch the broken dependency. One sick component has caused resource exhaustion and dragged down the whole app in a cascading failure.
Below, a storefront calls three dependencies from one pool of 12 threads. Step through what happens when Reviews hangs, and predict the damage before you see it. Then flip the switch to Bulkheads and replay the same hang.
How it works
The fix is to stop sharing one big pool. Partition your resources into separate, isolated pools and give each one a hard cap. A common split is per dependency (the payment API gets its own pool, the reviews service gets another), but you can also carve pools out per tenant (one noisy customer can't starve the rest) or per criticality (checkout traffic never competes with background reporting).
Now a failure is confined to its own compartment. When Reviews hangs, its calls pile up within its own pool only, and the moment that pool is full, further reviews calls are rejected fast instead of taking threads from anyone else. That's what the Bulkheads switch above shows: the same hang that killed checkout now costs you one widget. Every other pool keeps running on its own budget, so the blast radius is limited to the one thing that broke.
A full compartment should fail fast, ideally with a fallback: render the page without reviews, show cached recommendations, send the email later. A bulkhead that makes callers queue up for a free slot just moves the pile-up one step upstream, into the threads that are waiting in line.
Isolation is the whole point. The value of a bulkhead isn't that it makes the broken dependency work again — it doesn't. It's that the rest of your service stays alive and responsive while that one compartment is underwater, turning a total outage into a partial, survivable one.
Your API shares one pool of 50 threads between a flaky geocoding service and a fast database. When the geocoder hangs, every endpoint times out, even ones that only read the database. Which change actually stops that?
Pairing with other patterns
Bulkheads rarely travel alone. They define where the walls are, and other resilience patterns decide what to do inside each compartment.
A circuit breaker complements a bulkhead by detecting that a compartment's dependency is unhealthy and stopping calls to it altogether — so the pool doesn't even fill up before traffic is shed. Throttling works alongside both by capping how much load any one consumer can push, which is effectively how you enforce each pool's budget. Together, they keep one bad actor from taking everyone else down.
The trade-offs
Compartments aren't free. The same walls that contain failures also stop pools from sharing spare capacity, so your overall utilization drops — a pool sitting idle can't lend its threads to a pool that's momentarily slammed, the way one big shared pool could.
There's also more to tune. Every pool needs a size, and getting it wrong cuts both ways: too small and you reject healthy traffic during normal spikes; too large and the pool can exhaust the host before its cap ever kicks in. A good starting point is each dependency's busiest measured moment: one that handles 20 calls a second at 0.3 seconds each keeps about 20 × 0.3 = 6 threads busy, so give it 6 plus a little headroom. Then watch its rejections and adjust.
Below, a flash sale runs into walls sized for an ordinary day. Predict what the walls do, then pick a fix and see how it handles the next hang.
Don't carve up too finely. Every pool you add reserves capacity that can't be shared, so a flood of tiny single-purpose bulkheads quietly wastes a lot of resources. Group dependencies that share a fate or a criticality level, and reserve dedicated pools for the few that genuinely need isolation.
When to use it
Reach for bulkheads when your service depends on several backends and a problem in one could starve the others — especially when those dependencies have very different reliability or latency profiles, or when one is far more critical than the rest. They're also a strong fit for multi-tenant systems, where you want to guarantee that one heavy customer can't degrade everyone else's experience.
If you only have a single dependency and no shared-resource contention to worry about, the extra pools and tuning may not earn their keep. But the moment a slow backend can cascade into a full outage, partitioning your resources is one of the cheapest ways to turn a system-wide failure into a contained, recoverable one.
A reporting service runs every customer's reports on one pool of 40 workers. One customer's runaway report grabs all 40 for an hour, and nobody else's reports run. Which way of slicing the pool fixes this?