Skip to content

System design · 01

Queues don’t fix overload. They can hide it.

Adding a queue in front of a service does not add processing capacity. It can absorb short bursts, but if demand stays above consumer capacity, the backlog grows and latency rises.

A queue is a buffer. It is not infinite capacity.

Infographic: a service handling 1,000 jobs per second, traffic jumping to 2,000, the queue absorbing the spike, the backlog growing, jobs waiting 30 seconds then 2 minutes then 10, and the list of what a queue needs.
The whole idea on one page.

The mistake

A lot of engineers think adding a queue automatically makes a system more scalable. It doesn’t.

A queue does not add processing capacity. It adds a place for work to wait.

The distinction that matters: a transient spike is something a queue can absorb, and consumers catch up afterwards. Sustained demand above consumer capacity is something a queue can only accumulate as debt.

Do the arithmetic

Say your service processes 1,000 jobs per second, and arrivals rise to 2,000 per second and stay there.

You are now falling behind by 1,000 jobs every second. After one minute the backlog is 60,000 jobs. At 1,000 per second of drain capacity, the job arriving now waits a full minute before anything touches it.

After ten minutes of that traffic the backlog is 600,000, and the wait is ten minutes. The queue did not absorb the spike. It recorded it.

This is the whole idea: a queue converts a throughput problem into a latency problem. If arrivals stay above capacity, queueing delay keeps growing until the queue hits a limit, messages expire, or the system starts rejecting or dropping work.

Your system is “up” and still broken

Health checks may pass and error rates may look normal, while consumers are already at their effective throughput limit. The bottleneck may be CPU, or it may be the database, the network, a lock, a rate limit or an external call.

Meanwhile users are reading results that are ten minutes stale. If you are only watching service health and error rates, your dashboard may not show the user-visible delay.

Three ways it gets worse

  • Retries amplify the load. A failing job that is retried three times is four attempts. Retries increase effective load exactly when the system has the least spare capacity.
  • One bad message can block progress in ordered processing. For example, a Kafka consumer that refuses to advance past a failed record, or an SQS FIFO message group. On a standard unordered queue there is no head of line blocking, but a poison message still burns consumer capacity on every redelivery.
  • More consumers can move the failure, not fix it. If the bottleneck is the database behind your consumers, doubling consumers doubles the pressure on it. You have relocated the outage.

What a queue actually needs

  • Backpressure. A way to push the problem back to the producer: reject, throttle, or shed load. Something upstream has to slow down, and it is better that it is a decision than a collapse.
  • Consumer limits. A cap on concurrency so consumers cannot take a struggling downstream dependency down with them.
  • A dead letter queue. Somewhere for a message to go after N failures, so one bad payload cannot cost you capacity forever.
  • Retry policy with backoff. Bounded attempts, exponential backoff, and jitter so retries do not arrive in a synchronised wave.
  • Alerts on queue age, not just depth. Depth alone is misleading: a big queue that is draining fast may be healthy, while a small queue that is barely moving may be a problem. Alert on the age of the oldest unprocessed message, which is the number that matches what a user feels. A large queue can still breach an SLA even while it drains.
  • Idempotency. Assume a message may be delivered more than once. Consumers should make retries safe so duplicate delivery does not create duplicate side effects.
  • A written answer for sustained overload. When demand stays above capacity, something has to give: scale out, shed load, or degrade deliberately. Decide which before it happens, not at 2am.

In an interview

If an interviewer asks you to put a queue in a design, expect the follow up: “what happens when the consumers can’t keep up?” The answer they are listening for is backpressure and a deliberate degradation plan, not “we add more consumers”.

If work keeps entering the system faster than you can process it, the queue is only delaying the failure.

More of these at askgurpreet.com/system-design. Want your own answers scored? Free, five minutes.