Contents

Infrastructure & Operations › Incidents & SRE · also in Observability

SLO

A service level objective: the target for an SLI.

Also known as: service level objective, service level objectives, SLOs, reliability target, SLI SLO SLA

An SLO (service level objective) is a target for how reliable or fast a service should be, measured over a period. It’s built from an SLI and bounded by an error budget.

Three related terms:

TermMeaningExample
SLI (indicator)A measurement of some aspect of service qualityThe fraction of requests that succeed and finish in under 300 ms (SLI)
SLO (objective)The internal target for the SLI99.9% of requests meet that, over a rolling 30 days
SLA (agreement)A contract with customers, with consequences (credits, penalties) if missed (SLA)If availability falls below 99.5%, customers get a refund

SLOs are usually stricter than SLAs, so you notice and react before breaching a contract.

Choosing good SLIs

Measure what users experience, ideally from their side:

  • Availability: the proportion of requests that succeed (valid requests, not counting client errors).
  • Latency: the proportion of requests faster than a threshold (use percentiles, not averages) (percentiles).
  • Quality / correctness: results that are complete or fresh enough (data pipelines: data freshness).
  • Durability or throughput for storage and batch systems.

Express it as a ratio: good events / total events. That makes budgets and burn rates simple (error budget).

SLI = (requests with status < 500 AND latency < 300 ms) / (all valid requests)
SLO = SLI ≥ 99.9% over 30 days

Setting realistic objectives

  • Don’t aim for 100%. Each extra “nine” costs much more (nines of availability), and users can’t perceive past a point.
  • Base it on what users need and what the system does today. Look at historical data. Start with an achievable target, and tighten it deliberately.
  • Fewer, meaningful SLOs per service. A handful covering the critical user journeys beat dozens of metrics.
  • Account for dependencies: your service can’t be more reliable than the things it depends on (unless it degrades gracefully).
  • Agree it with product and business stakeholders, not only engineers.

What SLOs are for

  • Decisions: when the budget is healthy, ship faster. When it’s gone, invest in reliability (error budget).
  • Alerting: page when you’re burning through the budget quickly, not on every blip (alerting).
  • Shared language about “reliable enough” across teams and with customers.
  • Prioritization of reliability work against features.
  • Postmortems and trend analysis.

Pitfalls

  • Measuring the wrong thing (server CPU instead of user-visible success).
  • Targets nobody acts on. An SLO with no consequence is decoration.
  • Too tight and never met, or so loose it never triggers.
  • Gaming, such as excluding failures from the denominator.
  • Forgetting that a slow response can be as bad as an error.

See golden signals for a good starting set of things to measure.