Observability
Understanding what a running system is doing from logs, metrics and traces.
Backend Engineer track
Junior
Write correct code, ship small changes safely, ask good questions.
Core: start here
- Error TrackingTools like Sentry that group and report exceptions.
- Log LevelsDEBUG, INFO, WARN and ERROR, and when to use each.
- LoggingRecording what the application does so you can debug it later.
3 more junior concepts
- DashboardA visual display of key metrics.
- MonitoringWatching known metrics and alerting when they go wrong.
- Uptime MonitoringChecking from the outside that a site responds.
Mid-level
Own a feature end to end without hand-holding.
Core: start here
- AlertingNotifying people when something needs attention.
- Correlation / Request IDAn ID passed through every service to tie one request's logs together.
- MetricsNumeric measurements over time.
- Structured LoggingLogging key-value fields instead of free text.
8 more mid-level concepts
- Alert FatigueSo many alerts that people start ignoring them.
- APMApplication performance monitoring tools.
- Four Golden SignalsLatency, traffic, errors and saturation.
- Log AggregationCollecting logs from every server into one searchable place.
- Logs, Metrics and TracesThe three main kinds of telemetry.
- ObservabilityUnderstanding a system's internal state from its outputs.
- Percentiles (p50, p95, p99)The value below which a given share of measurements fall; the right way to read latency.
- Prometheus and GrafanaA common open-source stack for metrics and dashboards.
Senior
Own a system, its failure modes, and its trade-offs.
Core: start here
- Distributed TracingFollowing a single request across many services.
- SLOA service level objective: the target for an SLI.
- Tail LatencyThe slowest requests (p99), which users notice most.
9 more senior concepts
- Actionable AlertsAlerting only on symptoms that need a human to act.
- CardinalityThe number of unique label combinations, and why high cardinality gets expensive.
- Counter, Gauge, HistogramThe basic metric types.
- OpenTelemetryThe vendor-neutral standard for collecting telemetry.
- RED MethodRate, errors and duration for request-driven services.
- SamplingRecording only a fraction of traces or logs to control cost.
- SLIA service level indicator: a measured aspect of service quality.
- SpanOne timed operation within a trace.
- USE MethodUtilization, saturation and errors for resources.
Staff
Shape how many teams build, across systems.
- Wide EventsRich, high-cardinality events instead of pre-aggregated metrics.
Principal
Set technical direction for the organization.
Nothing here yet.
Data Analyst track
Junior
Write correct SQL, build trusted dashboards, ask good questions.
Core: start here
- MetricsNumeric measurements over time.
1 more junior concepts
- Percentiles (p50, p95, p99)The value below which a given share of measurements fall; the right way to read latency.
Mid-level
Own an analysis end to end, from vague question to recommendation.
Nothing here yet.
Senior
Own experimentation and metrics design; call out bad numbers.
Nothing here yet.
Staff
Shape how the organization measures and decides.
Nothing here yet.
Principal
Set measurement strategy across the company.
Nothing here yet.
Data Engineer track
Junior
Build and fix pipelines from clear specs; write correct SQL.
Core: start here
- LoggingRecording what the application does so you can debug it later.
- Percentiles (p50, p95, p99)The value below which a given share of measurements fall; the right way to read latency.
5 more junior concepts
- DashboardA visual display of key metrics.
- Error TrackingTools like Sentry that group and report exceptions.
- Log LevelsDEBUG, INFO, WARN and ERROR, and when to use each.
- MonitoringWatching known metrics and alerting when they go wrong.
- Uptime MonitoringChecking from the outside that a site responds.
Mid-level
Own pipelines and models end to end, including their quality.
- Alert FatigueSo many alerts that people start ignoring them.
- AlertingNotifying people when something needs attention.
- APMApplication performance monitoring tools.
- Correlation / Request IDAn ID passed through every service to tie one request's logs together.
- Four Golden SignalsLatency, traffic, errors and saturation.
- Log AggregationCollecting logs from every server into one searchable place.
- Logs, Metrics and TracesThe three main kinds of telemetry.
- MetricsNumeric measurements over time.
- ObservabilityUnderstanding a system's internal state from its outputs.
- Prometheus and GrafanaA common open-source stack for metrics and dashboards.
- Structured LoggingLogging key-value fields instead of free text.
Senior
Design the platform's storage, processing and modeling choices.
- Actionable AlertsAlerting only on symptoms that need a human to act.
- CardinalityThe number of unique label combinations, and why high cardinality gets expensive.
- Counter, Gauge, HistogramThe basic metric types.
- Distributed TracingFollowing a single request across many services.
- OpenTelemetryThe vendor-neutral standard for collecting telemetry.
- RED MethodRate, errors and duration for request-driven services.
- SamplingRecording only a fraction of traces or logs to control cost.
- SLIA service level indicator: a measured aspect of service quality.
- SLOA service level objective: the target for an SLI.
- SpanOne timed operation within a trace.
- Tail LatencyThe slowest requests (p99), which users notice most.
- USE MethodUtilization, saturation and errors for resources.
Staff
Shape how the whole organization produces and uses data.
- Wide EventsRich, high-cardinality events instead of pre-aggregated metrics.
Principal
Set data strategy and architecture across the company.
Nothing here yet.
Frontend Engineer track
Junior
Build UI that works, ship small changes safely, ask good questions.
Core: start here
- Error TrackingTools like Sentry that group and report exceptions.
5 more junior concepts
- DashboardA visual display of key metrics.
- Log LevelsDEBUG, INFO, WARN and ERROR, and when to use each.
- LoggingRecording what the application does so you can debug it later.
- MonitoringWatching known metrics and alerting when they go wrong.
- Uptime MonitoringChecking from the outside that a site responds.
Mid-level
Own a feature end to end without hand-holding.
- Alert FatigueSo many alerts that people start ignoring them.
- AlertingNotifying people when something needs attention.
- APMApplication performance monitoring tools.
- Correlation / Request IDAn ID passed through every service to tie one request's logs together.
- Logs, Metrics and TracesThe three main kinds of telemetry.
- MetricsNumeric measurements over time.
- ObservabilityUnderstanding a system's internal state from its outputs.
- Percentiles (p50, p95, p99)The value below which a given share of measurements fall; the right way to read latency.
- Structured LoggingLogging key-value fields instead of free text.
Senior
Own an app's architecture, performance, and failure modes.
Core: start here
- Distributed TracingFollowing a single request across many services.
- SLOA service level objective: the target for an SLI.
5 more senior concepts
- Actionable AlertsAlerting only on symptoms that need a human to act.
- OpenTelemetryThe vendor-neutral standard for collecting telemetry.
- SamplingRecording only a fraction of traces or logs to control cost.
- SLIA service level indicator: a measured aspect of service quality.
- Tail LatencyThe slowest requests (p99), which users notice most.
Staff
Shape how many teams build, across apps.
Nothing here yet.
Principal
Set technical direction for the organization.
Nothing here yet.