Architecture & System Design › Reliability & Resilience
Circuit Breaker
Stopping calls to a failing service so it can recover.
Also known as: circuit breaker pattern, circuit breakers, Resilience4j, Polly, fail fast breaker
A circuit breaker wraps calls to a dependency and stops sending requests when that dependency is clearly failing, so callers fail fast instead of piling up on a dead or overloaded service, and the dependency gets room to recover. The name comes from the electrical device that trips to prevent damage.
The three states
failures exceed threshold
CLOSED ───────────────────────────► OPEN
(calls pass through) (calls rejected immediately, no network attempt)
▲ │ after a cool-down period
│ trial calls succeed ▼
└──────────────────────────── HALF-OPEN
(a few trial calls allowed)
trial calls fail → back to OPEN
- Closed: normal operation. The breaker counts failures (errors, timeouts) in a window.
- Open: the failure rate crossed a threshold. Calls fail immediately (or use a fallback) without contacting the dependency, for a cool-down period.
- Half-open: after the cool-down, a few probe requests are allowed. If they succeed, the breaker closes. If they fail, it opens again.
breaker = CircuitBreaker(failure_threshold=0.5, window=20, open_seconds=30)
def get_recommendations(user_id):
try:
return breaker.call(lambda: recs_client.fetch(user_id, timeout=0.3))
except CircuitOpenError:
return popular_items() # fallback: degrade gracefully
Why it helps
- Prevents cascading failures: callers don’t tie up threads and connections waiting on a failing dependency (cascading failure).
- Gives the dependency breathing room: you stop hammering it while it’s down or overloaded, instead of adding retries (retry storms).
- Faster failure for users: an immediate fallback instead of a 30-second hang.
Using it well
- Pair it with timeouts. The breaker counts timeouts as failures, and without a timeout, a hanging call never gets counted (timeouts).
- Provide a meaningful fallback where possible: cached data, a default, a reduced feature, or a clear error (degraded modes).
- Tune the thresholds to the dependency’s normal error rate and traffic. Too sensitive means it trips on noise. Too lax means it doesn’t protect you. Require a minimum number of calls before judging.
- Use one breaker per dependency (or per endpoint), not one global one.
- Combine with bulkheads and retries (retry inside the breaker’s view, and not when it’s open) (bulkhead, retry with backoff).
- Monitor and alert on state changes. An open circuit is a signal that something is wrong.
- Count only the right failures. A
404or validation error isn’t a dependency failure. - Libraries exist for most languages (such as Resilience4j for Java and Polly for .NET), and service meshes and gateways can apply breakers outside your code.
Caveats
- A breaker can hide a problem if nobody’s watching the fallbacks.
- During half-open, a thundering herd of probes can re-trigger the failure, so limit them.
- It protects the caller. The dependency still needs its own protection: load shedding and capacity (load shedding).