Architecture & System Design › Reliability & Resilience
Single Point of Failure
One component whose failure takes everything down.
Also known as: single point of failure, SPOF, single points of failure
A single point of failure (SPOF) is any component whose failure takes the system down: one database, one load balancer, one DNS provider, one deploy pipeline, one engineer who knows the passwords. Reliability work starts by enumerating them — because every unexamined architecture hides several, usually where nobody looked (that one cron host, that shared NFS mount, that human).
find: trace every request path → which components appear in ALL paths?
kill: what happens if each dies? (test, don't theorise)
fix: duplicate, degrade, or accept-with-eyes-open
Not every SPOF gets fixed — some are accepted deliberately (cost, physics, scope) — but accepted SPOFs are documented, monitored and mitigated (fast repair, degradation plans), never merely undiscovered.
The classic mistakes:
- Obvious-duplicate blindness. Databases replicated while the single load balancer, single DNS, single region or single human stays singular. Map all paths, including operations and people.
- Shared fate mistaken for redundancy. Two servers, one power feed / switch / AZ — correlated failures kill both. Redundancy spans failure domains or it isn’t.
- “Highly available” components assumed immortal. Managed services fail too (regionally, durably). Multi-AZ still shares the provider; plan for provider-level degradation.
- Undocumented accepted SPOFs. “We’ll fix it later” without a record becomes permanent invisible risk. Log accepted SPOFs with owners and revisit dates.
- Testing components, not paths. Individual redundancy verified while end-to-end paths stay single (that one firewall rule, that shared secret). Test path failure, not box failure.
- People SPOFs ignored. Bus factor one on critical systems is the most common SPOF and the least addressed. Cross-train, document, automate — people redundancy is real redundancy.
- Fixing by duplication alone. Some SPOFs fix better by removal (stateless design) or degradation (serve stale) than by mirroring. Duplicate, degrade, or delete.
The practice: enumerate ruthlessly, fix by domain-spanning redundancy or designed degradation, document what remains. Every outage is a SPOF discovered late — discover them early instead.