Backend Development › Queues & Async Processing · also in Reliability & Resilience
Retry with Exponential Backoff
Retrying with growing delays so you don't hammer a failing service.
Also known as: exponential backoff, retry with exponential backoff, backoff and jitter, retry policy, retries
When a call fails, the right response is often “try again”, but not immediately, and not forever. Retry with exponential backoff waits longer after each failed attempt, giving the failing service time to recover instead of hammering it.
attempt 1 fails → wait 1 s
attempt 2 fails → wait 2 s
attempt 3 fails → wait 4 s
attempt 4 fails → wait 8 s ... capped at some maximum, then give up
import random, time
def call_with_retries(fn, max_attempts=5, base=0.5, cap=30):
for attempt in range(1, max_attempts + 1):
try:
return fn()
except TransientError:
if attempt == max_attempts:
raise
delay = min(cap, base * 2 ** (attempt - 1))
time.sleep(random.uniform(0, delay)) # "full jitter": a random wait up to the delay
Why add jitter
If a thousand clients fail at the same moment and all retry after exactly 1 s, 2 s, 4 s, they all hit the service at the same instant, again and again, a “thundering herd”. Jitter adds randomness, so retries spread out over time (jitter, retry storms).
What to retry, and what not
- Retry transient failures: timeouts, connection resets,
502/503/504,429(honor theRetry-Afterheader) (HTTP 429). - Don’t retry permanent errors:
400,401,403,404, validation failures. They’ll fail the same way every time. - Only retry operations that are safe to repeat (idempotent), or protect them with an idempotency key. A retried “charge card” without one can charge twice.
Rules that prevent trouble
- Always limit the attempts and the total time. Unlimited retries create endless work and stuck requests. Respect an overall deadline, and give up when the caller has already given up.
- Set a timeout on every attempt (timeouts).
- Retry at one level only. If the client, the service and the SDK each retry 3 times, one failure becomes 27 calls (retry amplification).
- Consider a retry budget: a cap on the fraction of traffic that may be retries, so retries can’t double the load on a struggling service.
- Combine with a circuit breaker that stops calling a dependency that’s clearly down.
- Log and measure retries. Frequent retries are a symptom worth fixing, not hiding.
- For messages, retry with delays, then send failures to a dead letter queue.
Many SDKs and HTTP clients include a retry policy. Check its defaults, and don’t stack your own on top blindly.