Infrastructure & Operations › Observability
Alert Fatigue
So many alerts that people start ignoring them.
Also known as: alert fatigue, alert noise, pager fatigue
Alert fatigue is what happens when a system sends so many alerts — or so many false ones — that people stop taking them seriously. On-call engineers learn that most pages are noise, so a real incident gets the same shrug as the rest, and response gets slower exactly when it matters.
The root cause is usually bad alert design: alerting on every symptom and every threshold rather than on conditions a human needs to act on. If a page doesn’t tell you something is wrong and what to do, it’s training people to ignore the pager.
The classic mistakes:
- Alerting on causes instead of symptoms. A brief CPU spike, or one failed request, isn’t worth waking someone. Alert on user-visible pain — high error rate, slow responses — and let dashboards show the causes.
- Too-sensitive thresholds. A short blip pages immediately, then resolves before anyone looks. Require condition to persist before alerting.
- No owner or runbook. An alert nobody owns, with no runbook, can’t be resolved and shouldn’t page.
- Noisy alerts that auto-resolve. Flapping trains people to wait it out.
Fixing it is as much culture as tooling: review pages regularly, delete or downgrade alerts that don’t lead to action, and track how many pages were noise. Tie alerts to things you actually care about — SLOs and error budgets give a principled line between “page” and “look at the dashboard tomorrow” (see actionable alerts). A healthy on-call rotation (on-call) depends on every page meaning something.