What Is a Bulkhead Pattern?

By Oleksandr Andrushchenko — Published on — Modified on
0 Likes
0 Dislikes
What Is a Bulkhead Pattern?
What Is a Bulkhead Pattern?

The Bulkhead Pattern is a reliability pattern that isolates resources so a failure or overload in one part of a system cannot consume all available capacity and bring down unrelated functionality.

The idea comes from ships, where separate watertight compartments prevent damage in one section from sinking the entire vessel. In software, the compartments can be separate thread pools, connection pools, worker groups, queues, service instances, or other limited resources.

Table of Contents

Why the Bulkhead Pattern Exists

Consider an API that communicates with three downstream services:

API
 ├──→ Payment Service
 ├──→ Recommendation Service
 └──→ Notification Service

Suppose all outgoing requests use the same pool of 100 worker threads.

The Recommendation service becomes extremely slow. Requests that normally finish in 100 milliseconds now take 30 seconds.

Incoming requests continue creating work:

Recommendation call 1  → waiting
Recommendation call 2  → waiting
Recommendation call 3  → waiting
...
Recommendation call 100 → waiting

Eventually all 100 threads are waiting for the Recommendation service.

Now a payment request arrives.

The Payment service itself is healthy, but there are no available threads to call it.

Recommendation failure
        ↓
Consumes shared thread pool
        ↓
Payment cannot execute
        ↓
Notification cannot execute
        ↓
Entire API becomes unavailable

This is a cascading failure. One unhealthy dependency consumes a shared resource until unrelated functionality fails too.

The Bulkhead Pattern limits the blast radius.

How the Bulkhead Pattern Works

Instead of sharing one resource pool, capacity is divided into isolated compartments.

API
 │
 ├── Payment Pool
 │      └──→ Payment Service
 │
 ├── Recommendation Pool
 │      └──→ Recommendation Service
 │
 └── Notification Pool
        └──→ Notification Service

For example:

Payment          → 40 threads
Recommendation   → 20 threads
Notification     → 20 threads
Shared reserve   → 20 threads

If all 20 Recommendation threads become blocked, the Recommendation functionality is degraded, but Payment and Notification retain their own capacity.

Recommendation Pool
20 / 20 busy
      ↓
New recommendation calls rejected

Payment Pool
5 / 40 busy
      ↓
Payment continues normally

The system accepts a localized failure instead of allowing that failure to consume every available resource.

The goal is not to prevent the Recommendation service from failing. The goal is to contain the consequences of that failure.

What Can Be Isolated?

A bulkhead is not limited to threads. Any scarce resource shared by unrelated workloads can become an isolation boundary.

Resource Example Isolation
Threads Separate execution pools per dependency
Connections Separate HTTP or database connection pools
Queues Separate queues for independent workloads
Workers Dedicated consumer groups
Service Instances Dedicated instances for critical traffic
CPU / Memory Container resource limits
Rate Limits Separate request budgets per tenant or operation
Database Capacity Separate connection or workload pools

The appropriate boundary depends on which shared resource could allow one workload to interfere with another.

Thread Pool Bulkheads

A common Bulkhead implementation assigns separate thread pools to different dependencies or workloads.

Without isolation:

Shared Pool: 100 threads

Payment
Recommendation
Notification
Search
Analytics
      ↓
All compete for the same 100 threads

With bulkheads:

Payment Pool        → 30
Recommendation Pool → 15
Notification Pool   → 15
Search Pool         → 25
Analytics Pool      → 15

A slow analytics dependency can consume only the 15 threads assigned to that workload.

This is particularly important with blocking I/O because every slow downstream request can hold an execution resource for a long time.

Thread isolation has a cost. Too many pools increase memory consumption, configuration complexity, and potentially context switching.

Isolation should therefore follow meaningful failure boundaries rather than creating a separate thread pool for every individual method.

Connection Pool Bulkheads

Connection pools can create the same type of shared-resource failure.

Suppose an application has 100 database connections shared by API traffic, reporting jobs, and background processing.

Database Pool: 100 connections

API Requests
Reports
Background Jobs
       ↓
Same Pool

A large reporting workload starts 100 expensive queries.

The API now cannot obtain a database connection even though the database may still be capable of processing short transactional queries.

Separate pools can isolate the workloads:

API Pool        → 60 connections
Reporting Pool  → 20 connections
Worker Pool     → 20 connections

If reporting exhausts its 20 connections, reports wait or fail while API capacity remains available.

The same principle applies to HTTP clients.

Using one unlimited or very large connection pool across unrelated downstream dependencies can allow a failing dependency to consume sockets, memory, file descriptors, and request capacity.

Queue and Worker Bulkheads

Asynchronous systems can use separate queues and worker pools as bulkheads.

Consider a system that processes:

  • payment events;
  • emails;
  • image processing;
  • analytics jobs.

Putting everything into one queue creates interference:

Shared Queue

Payment
Email
Image
Image
Image
Analytics
Image
Payment
...

A surge of image-processing jobs can increase payment latency even though the workloads are unrelated.

Separate queues provide isolation:

payments-queue
      ↓
Payment Workers

email-queue
      ↓
Email Workers

image-queue
      ↓
Image Workers

analytics-queue
      ↓
Analytics Workers

Each workload can have independent:

  • worker counts;
  • autoscaling rules;
  • retry policies;
  • dead-letter handling;
  • latency objectives;
  • resource limits.

This is one reason production message architectures frequently separate workloads instead of using one universal queue. Related queue reliability patterns are covered in Message Queue Best Practices for Production Systems.

Service and Instance Bulkheads

Isolation can also happen at the infrastructure level.

Suppose premium and batch workloads use the same service instances:

Premium Requests ─┐
                  ├──→ Shared Service Instances
Batch Processing ─┘

A large batch job can consume CPU, memory, connections, or request concurrency and increase latency for premium requests.

Dedicated capacity creates a stronger bulkhead:

Premium Requests
      ↓
Premium Instance Pool

Batch Processing
      ↓
Batch Instance Pool

This approach costs more because spare capacity cannot always be shared efficiently, but it provides stronger performance and failure isolation.

Similar isolation can be applied across:

  • tenants;
  • regions;
  • availability zones;
  • critical and non-critical APIs;
  • interactive and batch workloads;
  • production and internal workloads.

Bulkhead vs Circuit Breaker

Bulkhead and Circuit Breaker are complementary reliability patterns, but they solve different problems.

Bulkhead Circuit Breaker
Isolates resources Stops calls to a failing dependency
Limits failure blast radius Prevents repeated known-to-fail requests
Controls resource consumption Tracks dependency failure state
Protects unrelated workloads Allows a dependency time to recover

Consider a slow Recommendation service.

A bulkhead might allow at most 20 concurrent Recommendation calls:

Recommendation Bulkhead
       ↓
Max 20 concurrent calls

A Circuit Breaker observes that most of those calls are failing and eventually stops sending requests:

Failures exceed threshold
       ↓
Circuit opens
       ↓
Recommendation calls fail fast

Used together:

Request
   ↓
Circuit Breaker
   ↓
Bulkhead
   ↓
Recommendation Service

The Bulkhead limits how many resources the dependency can consume. The Circuit Breaker reduces unnecessary calls once the dependency is known to be unhealthy.

The Circuit Breaker mechanism is explained separately in What Is a Circuit Breaker?.

Bulkhead vs Load Shedding

Load shedding intentionally rejects work when the system does not have enough capacity to process it safely.

A bulkhead defines where capacity is reserved:

Payment Bulkhead
Max concurrency: 50

Load shedding defines what happens when that capacity is exhausted:

Current concurrency: 50
New request arrives
       ↓
No capacity
       ↓
Reject immediately

Without load shedding, requests may simply accumulate in an unbounded queue.

Capacity exhausted
      ↓
Requests keep arriving
      ↓
Queue grows
      ↓
Memory grows
      ↓
Latency grows
      ↓
Timeouts
      ↓
Retries
      ↓
More traffic

A bounded bulkhead with explicit rejection usually fails more predictably than unlimited waiting.

The relationship among these reliability mechanisms is explored in Circuit Breaker vs Bulkhead vs Load Shedding.

Choosing Bulkhead Boundaries

The most important Bulkhead design decision is deciding what should be isolated.

Useful boundaries often correspond to differences in:

  • business criticality;
  • latency requirements;
  • dependency reliability;
  • resource consumption;
  • traffic patterns;
  • failure behavior.

Consider an e-commerce API with these operations:

Checkout
Product Recommendations
Analytics
Email Notifications

Checkout is revenue-critical and latency-sensitive.

Recommendations are useful but optional.

Analytics can tolerate delays.

Email delivery can run asynchronously.

Putting all four workloads into one shared resource pool makes their reliability unnecessarily dependent on each other.

A better design isolates critical traffic from optional and background workloads.

Bulkhead boundaries should reflect failure domains, not merely code organization.

Capacity and Sizing

Isolation creates an important trade-off: reserved capacity improves reliability but can reduce utilization efficiency.

Suppose 100 threads are divided equally:

Service A → 25
Service B → 25
Service C → 25
Service D → 25

If Service A needs 40 threads while the others use only five each, A still cannot use their idle capacity.

This is the cost of strict isolation.

On the other hand, one shared pool:

Shared Pool → 100

provides better utilization but allows one dependency to consume everything.

Production designs often use a compromise:

Reserved critical capacity
        +
Bounded shared capacity

Bulkhead sizing should consider:

  • expected request rate;
  • request latency;
  • downstream latency during degradation;
  • acceptable queueing time;
  • criticality of the workload;
  • available CPU and memory;
  • database or downstream capacity;
  • failure-mode behavior.

A useful approximation for concurrency is:

Concurrency ≈ Requests per second × Average request duration

For example, a dependency receiving 200 requests per second with an average duration of 100 milliseconds needs approximately:

200 × 0.1 = 20 concurrent requests

But sizing only for normal latency can be dangerous.

If latency increases from 100 milliseconds to two seconds:

200 × 2 = 400 concurrent requests

Without a concurrency limit, a slow dependency can suddenly consume 20 times more application resources.

This is exactly the failure mode a bulkhead is designed to contain.

Production Design Example

Consider an e-commerce API with three downstream dependencies:

Checkout API
    │
    ├──→ Payment Service
    ├──→ Inventory Service
    └──→ Recommendation Service

Payment and Inventory are required to complete checkout. Recommendations are optional.

The application has capacity for approximately 200 concurrent downstream operations.

Instead of allowing every dependency to compete for all 200 slots, capacity is divided:

Payment          → 80
Inventory        → 80
Recommendation   → 20
Reserve          → 20

Each dependency also has a bounded waiting queue.

Payment
  concurrency: 80
  queue: 40

Inventory
  concurrency: 80
  queue: 40

Recommendation
  concurrency: 20
  queue: 10

Normal traffic flows without noticeable difference.

Then the Recommendation service slows from 50 milliseconds to 10 seconds.

Recommendation concurrency quickly reaches:

20 / 20 busy

The queue fills:

10 / 10 waiting

Additional recommendation requests are rejected immediately.

The API can degrade gracefully:

{
  "product_id": "product-52",
  "recommendations": []
}

Meanwhile:

Payment
22 / 80 busy

Inventory
31 / 80 busy

Checkout continues functioning because Recommendation cannot consume Payment or Inventory capacity.

If Recommendation failures continue, its Circuit Breaker opens and requests fail before reaching the bulkhead.

Recommendation Request
        ↓
Circuit Breaker OPEN
        ↓
Fail Fast
        ↓
Return response without recommendations

This is an example of Designing Graceful Degradation Strategies: optional functionality becomes unavailable while critical functionality remains operational.

Useful Bulkhead metrics include:

  • active concurrency per bulkhead;
  • maximum concurrency;
  • queue depth;
  • queue wait time;
  • rejection rate;
  • request latency;
  • timeout rate;
  • resource utilization;
  • circuit-breaker state;
  • downstream error rate.

A particularly important metric is bulkhead saturation:

active concurrency / maximum concurrency

A bulkhead running at 95–100% capacity for an extended period indicates that the workload is approaching its isolation limit even if requests have not started failing yet.

Failure Scenarios

Bulkheads contain failures, but the surrounding behavior still needs explicit design.

A dependency becomes slow. Its bulkhead fills while unrelated pools remain available.

The bulkhead queue fills. New work should usually be rejected or degraded instead of waiting indefinitely.

The bulkhead is too small. Healthy traffic is rejected even though the dependency could handle more work.

The bulkhead is too large. A failing dependency can still consume enough resources to damage the application.

Retries remain inside the same saturated pool. Retried work can increase pressure instead of allowing recovery.

Multiple bulkheads overload the same downstream database. Application-level isolation does not increase the actual capacity of a shared dependency.

One isolated workload consumes all CPU. Thread or connection isolation alone may not help if CPU, memory, or another resource remains globally shared.

Bulkhead design therefore requires understanding which resource actually causes the failure propagation.

When to Use the Bulkhead Pattern

The Bulkhead Pattern is valuable when independent workloads share limited resources and failure in one workload should not disable the others.

Common situations include:

  • microservices calling multiple downstream dependencies;
  • APIs serving critical and optional features;
  • background workers processing unrelated job types;
  • multi-tenant platforms;
  • database workloads with different priorities;
  • systems integrating with slow or unreliable third-party APIs;
  • applications with strict latency objectives;
  • high-traffic systems where resource exhaustion can trigger cascading failures.

A Bulkhead may provide little benefit when workloads are tightly coupled and cannot provide useful functionality independently.

Isolation also has operational cost, so very small systems should not create dozens of resource pools without a concrete failure-isolation requirement.

Common Bulkhead Mistakes

  • Using one shared resource pool for every dependency. One slow dependency can consume the entire application.
  • Creating too many tiny bulkheads. Excessive isolation wastes capacity and increases configuration complexity.
  • Using unbounded queues. The bulkhead limits active work while waiting requests continue consuming memory and increasing latency.
  • Setting concurrency limits without timeouts. Slow calls can occupy every bulkhead slot indefinitely.
  • Retrying aggressively after rejection. Immediate retries create more pressure on an already saturated system.
  • Ignoring downstream capacity. Increasing a bulkhead from 50 to 500 requests does not help if the database supports only 100.
  • Using identical limits for every dependency. Different workloads have different traffic, latency, and criticality.
  • Never measuring saturation. Rejections may appear suddenly even though the bulkhead has been near capacity for hours.
  • Assuming a Circuit Breaker replaces isolation. Resources can be exhausted before the breaker opens.
  • Assuming thread isolation protects every resource. CPU, memory, connections, or downstream capacity may remain shared.
  • Failing critical and optional functionality together. Optional features should often degrade independently.

Frequently Asked Questions

The Bulkhead Pattern is often discussed together with Circuit Breakers, timeouts, retries, and load shedding because production systems typically combine these mechanisms.

Does a Bulkhead Prevent Failures?

No. A bulkhead does not make a dependency more reliable.

It limits how much of the surrounding system can be affected when that dependency becomes slow, overloaded, or unavailable.

Should Bulkhead and Circuit Breaker Be Used Together?

They often work well together. The Bulkhead limits concurrent resource consumption, while the Circuit Breaker stops sending requests when a dependency is consistently failing.

Timeouts are also important because they determine how long a request can occupy a bulkhead slot. Retry behavior must be controlled so failed requests do not immediately create additional pressure.

These mechanisms are part of the broader reliability strategy described in Building Reliable Systems: Core Reliability Patterns Explained.

Does Every Dependency Need a Separate Bulkhead?

No. Bulkheads should represent meaningful failure-isolation boundaries.

Dependencies with similar reliability characteristics, resource usage, criticality, and failure behavior may share capacity. A separate bulkhead becomes valuable when one workload should not be allowed to exhaust resources required by another.

What Happens When a Bulkhead Is Full?

The system needs an explicit overload policy. Common choices include rejecting the request, returning a fallback response, dropping optional work, placing the request into a small bounded queue, or routing it to another capacity pool.

Allowing requests to wait indefinitely usually defeats the purpose of the pattern because the backlog becomes another source of resource exhaustion.

Conclusion

The Bulkhead Pattern protects systems from cascading failures by dividing shared capacity into isolated resource pools. A slow dependency, overloaded tenant, expensive background job, or optional feature can consume only its assigned capacity instead of exhausting the entire application.

Bulkheads can isolate threads, connections, queues, workers, service instances, or infrastructure resources. The right boundary depends on where failure propagation can occur.

The core principle is: do not allow one failing workload to consume resources required by unrelated healthy workloads.

Comments (0)