Reliability & Resilience
Keeping systems working when parts of them fail.
Backend Engineer
Junior
Write correct code, ship small changes safely, ask good questions.
Core: start here
- TimeoutsNever waiting forever on a network call.
2 more junior concepts
- Health CheckAn endpoint reporting whether the service is alive and ready.
- BackupsCopies of data for restoring, and why untested backups don't count.
Mid-level
Own a feature end to end without hand-holding.
Core: start here
- Rate LimitingLimiting how many requests a client can make.
- Retry with Exponential BackoffRetrying with growing delays so you don't hammer a failing service.
8 more mid-level concepts
- Single Point of FailureOne component whose failure takes everything down.
- RedundancyExtra components so a single failure isn't fatal.
- ReliabilityA system doing what it should, even when things fail.
- AvailabilityThe share of time a system is usable, often measured in "nines".
- Nines (99.9%, 99.99%)How much downtime each level of availability allows.
- Handling Dependency FailuresDeciding what happens when something you call is down.
- Restore TestingActually restoring backups regularly to prove they work.
- Liveness vs Readiness"Am I running?" vs "Can I take traffic?"
Senior
Own a system, its failure modes, and its trade-offs.
Core: start here
- Circuit BreakerStopping calls to a failing service so it can recover.
- Cascading FailureOne failure triggering failures in the systems that depend on it.
- Blast RadiusHow much breaks when something fails.
15 more senior concepts
- FailoverSwitching to a standby when the primary fails.
- Fault ToleranceContinuing to work when components fail.
- RPO and RTOHow much data you can afford to lose, and how long you can be down.
- Disaster RecoveryRestoring service after a catastrophic failure.
- BackpressureA slow consumer signalling a fast producer to slow down.
- Compensating TransactionUndoing a completed step with a new action, since you can't roll it back.
- Deadline PropagationPassing the remaining time budget down a chain of calls.
- JitterRandomizing retry delays so clients don't all retry in sync.
- BulkheadIsolating resources so one failure doesn't sink everything.
- Thread Pool ExhaustionAll workers stuck waiting on something slow, so nothing else gets served.
- Retry StormRetries multiplying the load on an already failing system.
- Degraded ModesServing reduced functionality when dependencies fail.
- Load SheddingRejecting some requests on purpose to protect the rest.
- Active-Active vs Active-PassiveAll nodes serving traffic vs standbys waiting.
- Self-Healing SystemsAutomatically restarting and replacing failed parts.
Staff
Shape how many teams build, across systems.
- Chaos EngineeringInjecting failures on purpose to find weaknesses.
Data Analyst
Mid-level
Own an analysis end to end, from vague question to recommendation.
- Rate LimitingLimiting how many requests a client can make.
Data Engineer
Junior
Build and fix pipelines from clear specs; write correct SQL.
Mid-level
Own pipelines and models end to end, including their quality.
- Rate LimitingLimiting how many requests a client can make.
- Single Point of FailureOne component whose failure takes everything down.
- RedundancyExtra components so a single failure isn't fatal.
- ReliabilityA system doing what it should, even when things fail.
- AvailabilityThe share of time a system is usable, often measured in "nines".
- Nines (99.9%, 99.99%)How much downtime each level of availability allows.
- Retry with Exponential BackoffRetrying with growing delays so you don't hammer a failing service.
- Handling Dependency FailuresDeciding what happens when something you call is down.
- Restore TestingActually restoring backups regularly to prove they work.
- Liveness vs Readiness"Am I running?" vs "Can I take traffic?"
Senior
Design the platform's storage, processing and modeling choices.
- FailoverSwitching to a standby when the primary fails.
- Fault ToleranceContinuing to work when components fail.
- Circuit BreakerStopping calls to a failing service so it can recover.
- RPO and RTOHow much data you can afford to lose, and how long you can be down.
- Disaster RecoveryRestoring service after a catastrophic failure.
- BackpressureA slow consumer signalling a fast producer to slow down.
- Compensating TransactionUndoing a completed step with a new action, since you can't roll it back.
- Cascading FailureOne failure triggering failures in the systems that depend on it.
- Deadline PropagationPassing the remaining time budget down a chain of calls.
- JitterRandomizing retry delays so clients don't all retry in sync.
- BulkheadIsolating resources so one failure doesn't sink everything.
- Thread Pool ExhaustionAll workers stuck waiting on something slow, so nothing else gets served.
- Retry StormRetries multiplying the load on an already failing system.
- Degraded ModesServing reduced functionality when dependencies fail.
- Blast RadiusHow much breaks when something fails.
- Load SheddingRejecting some requests on purpose to protect the rest.
- Active-Active vs Active-PassiveAll nodes serving traffic vs standbys waiting.
- Self-Healing SystemsAutomatically restarting and replacing failed parts.
Staff
Shape how the whole organization produces and uses data.
- Chaos EngineeringInjecting failures on purpose to find weaknesses.
Frontend Engineer
Junior
Build UI that works, ship small changes safely, ask good questions.
Mid-level
Own a feature end to end without hand-holding.
- Rate LimitingLimiting how many requests a client can make.
- Single Point of FailureOne component whose failure takes everything down.
- RedundancyExtra components so a single failure isn't fatal.
- ReliabilityA system doing what it should, even when things fail.
- AvailabilityThe share of time a system is usable, often measured in "nines".
- Nines (99.9%, 99.99%)How much downtime each level of availability allows.
- Retry with Exponential BackoffRetrying with growing delays so you don't hammer a failing service.
- Handling Dependency FailuresDeciding what happens when something you call is down.
Senior
Own an app's architecture, performance, and failure modes.
- Fault ToleranceContinuing to work when components fail.
- Circuit BreakerStopping calls to a failing service so it can recover.
- BackpressureA slow consumer signalling a fast producer to slow down.
- Cascading FailureOne failure triggering failures in the systems that depend on it.
- JitterRandomizing retry delays so clients don't all retry in sync.
- Retry StormRetries multiplying the load on an already failing system.
- Degraded ModesServing reduced functionality when dependencies fail.
- Blast RadiusHow much breaks when something fails.
Tech Entrepreneur
MVP
Ship the smallest thing someone will use or pay for, and set up the company properly.
- BackupsCopies of data for restoring, and why untested backups don't count.
Product-Market Fit
Get users who stay, pay and tell others; raise your first real round.
- AvailabilityThe share of time a system is usable, often measured in "nines".
Growth
Make acquisition and revenue repeatable; hire, raise and build a board.
- Disaster RecoveryRestoring service after a catastrophic failure.