Backend Development › Queues & Async Processing
Queue Depth and Consumer Lag
How far behind consumers are; a key health signal.
Also known as: queue depth, consumer lag, backlog
Queue depth is the number of messages waiting in a queue; consumer lag is how far behind a consumer is (for a stream, the gap between the latest produced offset and the consumed one). Both answer the same question: are consumers keeping up with producers?
producers adding 1,000/s, consumers handling 800/s → depth grows 200/s
A stable, small depth means the system is balanced. A growing depth means consumers are falling behind — and because a queue is a buffer, not a fix, the backlog will eventually hit a limit, memory will fill, or latency for queued work will become unacceptable.
The classic mistakes:
- Not monitoring it. Queue depth is the single most important health metric for an async system, yet it’s often unmonitored until an incident. Alarm on sustained growth, not just a large absolute number.
- Ignoring sustained growth. A slowly rising depth is easy to dismiss until it’s a crisis. The slope matters: a steadily growing backlog means the system can’t keep up even if nothing is failing.
- Reacting too late. By the time depth is huge, the backlog is hours of work and the fix is slow. Alert early, on trend.
- Scaling consumers blindly. Adding consumers helps only if the bottleneck is consumer count — if it’s a downstream database or a hot partition, more consumers make it worse. Find the real bottleneck.
- No backpressure. If producers keep pushing while consumers fall behind, the queue grows without bound. The producer should eventually be slowed (see backpressure and load shedding).
- Confusing depth with latency. A low depth with slow processing still means high end-to-end latency; and a big depth means queued work is old. Watch both depth and time-in-queue.
- Ignoring per-partition lag. In a stream, lag can be uneven — one partition far behind while others are fine. Aggregate lag hides a stuck partition.
How to use it: treat queue depth and consumer lag as core dashboards; alert on sustained growth; when it grows, find whether the bottleneck is consumers, a downstream dependency, or a hot partition; and apply backpressure so the producer can’t outrun the system indefinitely. It’s the clearest signal of whether your async pipeline is healthy. See producer-consumer.