Working in Production
The everyday operational tasks of running a live system, and how to do them safely.
Backend Engineer track
Junior
Write correct code, ship small changes safely, ask good questions.
Core: start here
- Reading Production LogsFinding the few lines that matter among millions, by request ID, time and level.
- RollbackReturning to the previous version when a deploy goes wrong.
1 more junior concepts
- HotfixAn urgent fix shipped outside the normal release cycle.
Mid-level
Own a feature end to end without hand-holding.
Core: start here
- On-CallBeing responsible for responding to production issues.
10 more mid-level concepts
- Ad-Hoc Queries on ProductionQuerying live data without hurting it: read replicas, timeouts and LIMIT.
- Feature Flag CleanupRemoving flags once rollout is done so they don't pile up as debt.
- Fixing Data in ProductionCorrecting bad rows safely: a reviewed script, a backup, and a record of what changed.
- Launch ChecklistMonitoring, rollback, docs and support readiness before something goes live.
- Maintenance WindowA scheduled, announced time for risky changes.
- On-Call HandoverPassing open issues and context to the next person on call.
- One-Off ScriptsScripts run once against production, and why they deserve review and tests too.
- Production ConsolesInteractive shells against live data, like a Rails or Django console, and their dangers.
- Reproducing Production Issues LocallyRecreating a production bug with realistic, anonymized data.
- RunbookStep-by-step instructions for operating or fixing a system.
Senior
Own a system, its failure modes, and its trade-offs.
- Break-Glass AccessEmergency elevated access that is logged, time-limited and reviewed afterwards.
Staff
Shape how many teams build, across systems.
Nothing here yet.
Principal
Set technical direction for the organization.
Nothing here yet.
Data Engineer track
Junior
Build and fix pipelines from clear specs; write correct SQL.
- HotfixAn urgent fix shipped outside the normal release cycle.
- Reading Production LogsFinding the few lines that matter among millions, by request ID, time and level.
- RollbackReturning to the previous version when a deploy goes wrong.
Mid-level
Own pipelines and models end to end, including their quality.
- Ad-Hoc Queries on ProductionQuerying live data without hurting it: read replicas, timeouts and LIMIT.
- Feature Flag CleanupRemoving flags once rollout is done so they don't pile up as debt.
- Fixing Data in ProductionCorrecting bad rows safely: a reviewed script, a backup, and a record of what changed.
- Launch ChecklistMonitoring, rollback, docs and support readiness before something goes live.
- Maintenance WindowA scheduled, announced time for risky changes.
- On-CallBeing responsible for responding to production issues.
- On-Call HandoverPassing open issues and context to the next person on call.
- One-Off ScriptsScripts run once against production, and why they deserve review and tests too.
- Production ConsolesInteractive shells against live data, like a Rails or Django console, and their dangers.
- Reproducing Production Issues LocallyRecreating a production bug with realistic, anonymized data.
- RunbookStep-by-step instructions for operating or fixing a system.
Senior
Design the platform's storage, processing and modeling choices.
- Break-Glass AccessEmergency elevated access that is logged, time-limited and reviewed afterwards.
Staff
Shape how the whole organization produces and uses data.
Nothing here yet.
Principal
Set data strategy and architecture across the company.
Nothing here yet.
Frontend Engineer track
Junior
Build UI that works, ship small changes safely, ask good questions.
Core: start here
- Reading Production LogsFinding the few lines that matter among millions, by request ID, time and level.
Mid-level
Own a feature end to end without hand-holding.
- Feature Flag CleanupRemoving flags once rollout is done so they don't pile up as debt.
- Launch ChecklistMonitoring, rollback, docs and support readiness before something goes live.
- On-CallBeing responsible for responding to production issues.
- On-Call HandoverPassing open issues and context to the next person on call.
- Reproducing Production Issues LocallyRecreating a production bug with realistic, anonymized data.
- RunbookStep-by-step instructions for operating or fixing a system.
Senior
Own an app's architecture, performance, and failure modes.
Nothing here yet.
Staff
Shape how many teams build, across apps.
Nothing here yet.
Principal
Set technical direction for the organization.
Nothing here yet.