Why your queue workers are slower than your web requests
Moving work to a queue makes the request fast. It does not make the work fast — and the difference is where most throughput problems live.
Pushing a job onto a queue is the standard fix for a slow endpoint. The request returns in milliseconds, the user is happy, and the problem looks solved.
It usually isn't. You moved the work, you didn't remove it — and queue workers have a failure mode that web requests don't: they fall behind silently.
The three things that actually bite
1. Worker concurrency doesn't match the work
A worker pool sized for fast jobs will stall the moment a slow job type enters the same queue. Separate queues by expected duration, not by feature area.
2. Retries multiply load
A job that fails at 90% completion and retries from the start turns one unit of work into several. Make jobs idempotent and checkpoint long ones.
3. Nobody is watching lag
Queue depth is the metric that matters, and it is the one nobody alerts on. If depth is growing, your effective response time is growing — the user just can't see it yet.
Queue depth trending upward is an outage in slow motion.
What to do instead
- Split queues by job duration class.
- Make every job idempotent — assume it will run twice.
- Alert on queue depth and oldest-job age, not just error rate.
- Measure end-to-end latency: enqueue time to completion, not request time.
This is a starter post seeded with your site. Edit or delete it from the admin panel.