Contents

Infrastructure & Operations › Observability

Alerting

Notifying people when something needs attention.

Also known as: alerts, monitoring alerts, paging, alert rules, alert thresholds

Alerting means automatically notifying people when something needs attention. Dashboards are for looking, and alerts are for being interrupted. Because an interruption costs sleep and focus, an alert must earn it.

What makes a good alert

  • Actionable: someone must be able to do something about it now. If the answer is “nothing, it’ll fix itself” or “look at it tomorrow”, it isn’t a page (actionable alerts).
  • Symptom-based: alert on what users experience (error rate, latency, failed checkouts, “the site is down”), rather than on every possible cause (CPU at 80%, one pod restarted). Causes are for dashboards and for finding the problem once alerted.
  • Clear: the message says what’s wrong, how bad, and where to start (a link to a dashboard and a runbook).
  • Reliable: rarely wrong. False alarms teach people to ignore alerts (alert fatigue).
# Prometheus-style rule (illustrative)
- alert: HighErrorRate
  expr: sum(rate(http_requests_total{status=~"5.."}[5m])) / sum(rate(http_requests_total[5m])) > 0.02
  for: 10m                     # must persist for 10 minutes: avoids flapping on blips
  labels: { severity: page }
  annotations:
    summary: "More than 2% of requests are failing"
    runbook: "https://wiki.example.com/runbooks/high-error-rate"

Design choices

QuestionGuidance
Page or ticket?Page (wake someone) only when users are being harmed now. Otherwise create a ticket or a chat message to handle in working hours
Threshold or trend?Static thresholds are simple. Alerting on SLO burn rate (how fast you’re using up your error budget) ties alerts to what matters
How long must it persist?Add a for duration to ignore short blips
Missing data?Alert on “no data” or “no heartbeat”: a dead service sends no errors
Who gets it?Route to the owning team’s on-call, with an escalation path

Habits

  • Review alerts regularly. Delete or fix noisy ones, and add alerts after incidents that were detected late.
  • Every page should be investigated, and each recurring one should lead to a fix, or the removal of the alert.
  • Test them: confirm that an alert fires when it should, and reaches someone.
  • Avoid alert storms: group related alerts, and suppress dependent ones when the root cause is known.
  • Watch the watchers: alert when the monitoring itself stops working.
  • Use the four golden signals (latency, traffic, errors, saturation) as a starting set.