Incidents & SRE
Responding when things break, and engineering so they break less.
Backend Engineer track
Junior
Write correct code, ship small changes safely, ask good questions.
- IncidentAn unplanned event that disrupts or degrades service.
Mid-level
Own a feature end to end without hand-holding.
Core: start here
- On-CallBeing responsible for responding to production issues.
- PostmortemA written review of an incident: what happened and how to prevent it.
10 more mid-level concepts
- Blameless PostmortemFocusing on systems rather than individuals when things go wrong.
- Emergency Change ProcessShipping urgent fixes safely under pressure.
- EscalationKnowing when and how to pull in more help.
- Incident ResponseHow a team detects, coordinates, fixes and communicates during an incident.
- Incident Severity LevelsSEV1 to SEV4: how bad an incident is.
- Mitigate First, Fix LaterStopping the bleeding before hunting for the root cause.
- PagingAlerts that wake someone up.
- RunbookStep-by-step instructions for operating or fixing a system.
- SLAA service level agreement: a promise to customers, with consequences.
- Status PagePublic communication about outages.
Senior
Own a system, its failure modes, and its trade-offs.
Core: start here
- Error BudgetHow much unreliability an SLO allows, used to balance speed and stability.
- SLOA service level objective: the target for an SLI.
7 more senior concepts
- Customer Communication in IncidentsTelling users what's happening, honestly and promptly.
- Incident CommanderThe person coordinating an incident response.
- MTTR / MTTDMean time to recover, and to detect.
- Runbook AutomationTurning manual runbook steps into scripts.
- Site Reliability EngineeringApplying software engineering to operations.
- SLIA service level indicator: a measured aspect of service quality.
- ToilManual, repetitive operational work that should be automated.
Staff
Shape how many teams build, across systems.
- Game DayPracticing incident response with simulated failures.
- Production Readiness ReviewA checklist before a service goes live.
Principal
Set technical direction for the organization.
Nothing here yet.
Data Engineer track
Junior
Build and fix pipelines from clear specs; write correct SQL.
- IncidentAn unplanned event that disrupts or degrades service.
Mid-level
Own pipelines and models end to end, including their quality.
- Blameless PostmortemFocusing on systems rather than individuals when things go wrong.
- Emergency Change ProcessShipping urgent fixes safely under pressure.
- EscalationKnowing when and how to pull in more help.
- Incident ResponseHow a team detects, coordinates, fixes and communicates during an incident.
- Incident Severity LevelsSEV1 to SEV4: how bad an incident is.
- Mitigate First, Fix LaterStopping the bleeding before hunting for the root cause.
- On-CallBeing responsible for responding to production issues.
- PagingAlerts that wake someone up.
- PostmortemA written review of an incident: what happened and how to prevent it.
- RunbookStep-by-step instructions for operating or fixing a system.
- SLAA service level agreement: a promise to customers, with consequences.
- Status PagePublic communication about outages.
Senior
Design the platform's storage, processing and modeling choices.
- Customer Communication in IncidentsTelling users what's happening, honestly and promptly.
- Error BudgetHow much unreliability an SLO allows, used to balance speed and stability.
- Incident CommanderThe person coordinating an incident response.
- MTTR / MTTDMean time to recover, and to detect.
- Runbook AutomationTurning manual runbook steps into scripts.
- Site Reliability EngineeringApplying software engineering to operations.
- SLIA service level indicator: a measured aspect of service quality.
- SLOA service level objective: the target for an SLI.
- ToilManual, repetitive operational work that should be automated.
Staff
Shape how the whole organization produces and uses data.
- Game DayPracticing incident response with simulated failures.
- Production Readiness ReviewA checklist before a service goes live.
Principal
Set data strategy and architecture across the company.
Nothing here yet.
Frontend Engineer track
Junior
Build UI that works, ship small changes safely, ask good questions.
- IncidentAn unplanned event that disrupts or degrades service.
Mid-level
Own a feature end to end without hand-holding.
- Blameless PostmortemFocusing on systems rather than individuals when things go wrong.
- Emergency Change ProcessShipping urgent fixes safely under pressure.
- EscalationKnowing when and how to pull in more help.
- Incident ResponseHow a team detects, coordinates, fixes and communicates during an incident.
- Incident Severity LevelsSEV1 to SEV4: how bad an incident is.
- Mitigate First, Fix LaterStopping the bleeding before hunting for the root cause.
- On-CallBeing responsible for responding to production issues.
- PagingAlerts that wake someone up.
- PostmortemA written review of an incident: what happened and how to prevent it.
- RunbookStep-by-step instructions for operating or fixing a system.
- SLAA service level agreement: a promise to customers, with consequences.
- Status PagePublic communication about outages.
Senior
Own an app's architecture, performance, and failure modes.
Core: start here
- SLOA service level objective: the target for an SLI.
5 more senior concepts
- Customer Communication in IncidentsTelling users what's happening, honestly and promptly.
- Error BudgetHow much unreliability an SLO allows, used to balance speed and stability.
- Incident CommanderThe person coordinating an incident response.
- MTTR / MTTDMean time to recover, and to detect.
- SLIA service level indicator: a measured aspect of service quality.
Staff
Shape how many teams build, across apps.
- Production Readiness ReviewA checklist before a service goes live.
Principal
Set technical direction for the organization.
Nothing here yet.