Architecture & System Design › Reliability & Resilience
Redundancy
Extra components so a single failure isn't fatal.
Also known as: redundancy, redundant systems, n+1 redundancy
Redundancy eliminates single points by duplicating critical components: two power feeds, three database replicas, N+1 app servers — so any one failure leaves service intact. It’s availability’s bluntest instrument: don’t prevent failure, survive it by having spares.
single: [DB] → failure = outage
redundant: [DB] [DB] [DB] → failure = failover (if tested)
Redundancy only works with the machinery around it: failure detection (know it’s dead), failover (shift work automatically), and independence (spares that share fate aren’t spares). Untested, correlated redundancy is expense without safety.
The classic mistakes:
- Correlated spares. Replicas in one rack, AZ or region fail together; “redundant” systems with shared fate fail as one. Spread across genuine failure domains.
- Untested failover. Standbys with stale data, broken config or insufficient capacity fail on takeover. Redundancy untested is redundancy imagined — rehearse regularly.
- N+1 arithmetic without margins. Exactly-one-spare assumes failures arrive singly and repairs are instant; overlapping failures and slow repairs need N+2 thinking for critical paths.
- Redundant but inconsistent. Replicas diverged (lag, split-brain, untested promotion) serve wrong answers after failover. Redundancy includes consistency machinery.
- Ignoring common-mode software. Three replicas running the same buggy version fail identically; redundancy covers hardware and zones, not correlated software faults. Version diversity (or fast rollback) for software risk.
- Active-passive rot. Cold standbys decay (config drift, data staleness, skill atrophy) precisely because they’re never used. Active-active or frequent active testing keeps spares honest.
- Redundancy theatre. Duplicating the easy components while the single load balancer, DNS or human process remains singular. Map all single points, not the convenient ones.
How to build it: duplicate across real domains, automate detection and failover, rehearse constantly, mind correlated faults. Redundancy is the price of surviving failure — pay it where outages cost most.