Architecture & System Design › Reliability & Resilience
Degraded Modes
Serving reduced functionality when dependencies fail.
Also known as: degraded modes, graceful degradation modes, partial functionality
Degraded modes are designed reduced-functionality states: recommendations off, images low-res, search simplified, personalization paused — the service works, just less richly. Unlike accidental breakage, degraded modes are planned: triggers defined, behaviour specified, recovery automatic, users informed honestly.
normal → (dependency slow) → degraded (cached/defaults) → (recovered) → normal
Good degradation is tiered (shed in order of user value), automatic (no human flips switches at 3am — though manual overrides exist), and visible internally (dashboards show current mode) while calm externally (no alarming users over handled degradation).
The classic mistakes:
- All-or-nothing design. Full feature or total error, with nothing between, turns every dependency hiccup into user-facing failure. Design the middle states first.
- Degradation by accident. Untested fallback paths (stale caches, default content) failing exactly when needed. Exercise degraded modes regularly (game days, fault injection).
- Silent degradation. Users seeing wrong/stale content without indication lose trust when they discover it. Signal reduced freshness honestly where it matters.
- No recovery path. Entering degraded mode easily but requiring manual restoration leaves systems limping after incidents pass. Automate return-to-normal with verification.
- Priority inversion in shedding. Dropping checkout to save recommendations inverts user value. Order shedding by business impact, decided calmly in advance.
- Flag sprawl. Dozens of manual degradation toggles nobody remembers create unpredictable system states. Few, well-understood modes with clear ownership.
- Testing only normal. Load tests, chaos tests and game days run at full function; degraded paths — the ones incidents actually use — stay unverified. Test the modes users will see.
How to design them: tier features by value, define triggers and behaviours per tier, automate transitions both ways, exercise regularly. Degradation designed is resilience; degradation improvised is outage with extra steps.