Contents

Infrastructure & Operations › Incidents & SRE

Production Readiness Review

A checklist before a service goes live.

Also known as: production readiness review, PRR, readiness review

A production readiness review (PRR) is the gate a service passes before it takes real traffic. It’s a structured check that the things that make a service operable — not just functional — are in place. The service works in a test; the review asks whether it will stay working under real load, with real users, at 3am.

What a review covers:

  • Observability. Metrics, logs and alerts that show the service’s health, and an SLO that defines “good”. You can’t operate what you can’t see (observability).
  • Failure handling. What happens when a dependency is down, a request times out, or a queue backs up. Graceful degradation, retries with limits, timeouts.
  • Rollback and recovery. A way back (see rollback), and a runbook for known failure modes.
  • Capacity. Realistic load tested (see capacity planning), with limits and scaling understood.
  • Ownership. An on-call rotation and a team that answers when it pages (see on-call).
  • Data and dependencies. Backups, migration safety, and any external services it relies on.
not ready: "it works on staging" + a single instance
ready:     monitored, alerting, SLO'd, rollback-able, on-call owned

The classic mistakes:

  • Checkbox theater. Running through a list without understanding it gives false confidence. The point is a real conversation about how the service fails and how you’ll know.
  • No owner. A review with no accountable team drifts and gets skipped. Someone drives it and signs off.
  • Skipping it for “small” services. Small services cause outages too. Scale the depth to the risk, but don’t drop it — and a service someone depends on isn’t small.
  • A one-time gate. Readiness decays as the system changes. Revisit for major changes (see configuration drift and progressive delivery).

How it fits: a PRR is the deeper, service-level version of a launch checklist, usually required once before a new service carries production traffic. It’s the moment to catch the operational gaps while they’re cheap, rather than discovering them during the first incident.