Originally published at warrenops.io.
A queue with messages and zero consumers is the quietest outage there is. Nothing errors. Producers publish happily. The dashboard shows the broker green. Then the TTL runs out, or the queue hits its limit, and an hour of orders is in a dead-letter queue. Here is how to get paged before that.
What to alert on
"No consumers" alone is not the alert. Plenty of queues have no consumers by design: retry-wait queues, parking-lot DLQs, queues of a batch job that runs at night. The signal is the combination:
consumers == 0- and
messages_ready > 0(something is waiting) - and it has been like that for a few minutes (a deploy restarts consumers; do not page on a 20-second gap)
- and the queue is not on your list of intentionally consumer-less queues.
Two neighbours of this alert are worth setting up at the same time, because they catch the cases "consumers exist but do nothing":
-
Stuck consumers:
consumers > 0,messages_ready > 0, and the ack rate is zero for five minutes. The consumer process is alive, holds the connection, and is deadlocked or stuck on a downstream call. -
Growth:
messages_readygrew by more than N in the last M minutes. Consumers are there but cannot keep up.
Where the numbers come from
Management HTTP API
The management plugin exposes everything you need per queue. Ask for only the columns you want, otherwise the response for a broker with a thousand queues is huge:
curl -s -u monitor:secret \
'http://rabbit:15672/api/queues?columns=vhost,name,consumers,messages_ready,messages_unacknowledged,message_stats.ack_details.rate'
Relevant fields:
| Field | Meaning |
|---|---|
consumers |
Number of consumers currently attached to the queue. |
messages_ready |
Messages waiting to be delivered. This is the "backlog". |
messages_unacknowledged |
Delivered but not yet acked. High and flat means consumers hold messages without finishing them. |
message_stats.ack_details.rate |
Acks per second, averaged over the last sampling interval. Zero with a backlog and consumers present is the stuck-consumer signal. |
consumer_utilisation |
Fraction of time consumers could take new messages. Low with a backlog means consumers are the bottleneck. |
idle_since |
Set when the queue had no activity at all since that time. Useful to ignore truly dead queues. |
The monitoring user needs the monitoring tag, nothing more. Do not use an administrator account for a poller.
Prometheus plugin
Since RabbitMQ 3.8, rabbitmq_prometheus ships with the broker. By default /metrics is aggregated: you get totals per node, not per queue, which is useless for this alert. Per-queue values come from either:
-
/metrics/per-object, every metric for every object. Fine for a few hundred queues, heavy for thousands. -
/metrics/detailed?family=queue_coarse_metrics&family=queue_consumer_count(3.9+), only the families you ask for. Metrics from this endpoint are prefixedrabbitmq_detailed_. - Or set
prometheus.return_per_object_metrics = trueinrabbitmq.confto make/metricsper-object.
rabbitmq-plugins enable rabbitmq_prometheus
curl -s 'http://rabbit:15692/metrics/detailed?family=queue_coarse_metrics&family=queue_consumer_count' | grep orders
Prometheus alert rules
With the detailed endpoint scraped as its own job, three rules cover the three signals above. Adjust the prefix (drop detailed_) if you use per-object metrics.
groups:
- name: rabbitmq-queues
rules:
- alert: RabbitMQQueueNoConsumers
expr: |
rabbitmq_detailed_queue_consumers == 0
and on (vhost, queue) rabbitmq_detailed_queue_messages_ready{queue!~".*(\\.wait|\\.dlq|\\.parking)$"} > 0
for: 5m
labels: {severity: page}
annotations:
summary: "{{ $labels.vhost }}/{{ $labels.queue }} has {{ $value }} messages and no consumer"
- alert: RabbitMQQueueConsumersStuck
expr: |
rabbitmq_detailed_queue_consumers > 0
and on (vhost, queue) rabbitmq_detailed_queue_messages_ready > 0
and on (vhost, queue) rate(rabbitmq_detailed_queue_messages_acked_total[5m]) == 0
for: 5m
labels: {severity: page}
- alert: RabbitMQQueueGrowing
expr: |
delta(rabbitmq_detailed_queue_messages_ready[15m]) > 1000
for: 0m
labels: {severity: warn}
The regex on queue excludes the queues that have no consumers on purpose; more on that below. Because the exclusion sits on the messages_ready side of the and, those queues never produce a result, whatever their consumer count.
Route severity: page to your on-call and warn to a channel. A queue growing at 2 a.m. may be a batch job; a queue with no consumer at 2 a.m. is not.
No Prometheus? Cron and curl
A twenty-line script on a box that can reach the management API does the job for a small setup. This one posts to a Slack incoming webhook when a queue has a backlog and no consumer, and keeps a marker file so it does not repeat itself every minute:
#!/usr/bin/env bash
set -euo pipefail
API=http://rabbit:15672; AUTH=monitor:secret
HOOK=https://hooks.slack.com/services/T000/B000/xxxx
IGNORE='\.(wait|dlq|parking)$'
STATE=/var/tmp/rabbit-noconsumer; mkdir -p "$STATE"
# "vhost/queue<TAB>ready" for every queue that is in trouble right now
ALERTING=$(curl -sf -u "$AUTH" "$API/api/queues?columns=vhost,name,consumers,messages_ready" \
| jq -r --arg ign "$IGNORE" \
'.[] | select(.consumers == 0 and .messages_ready > 0 and (.name | test($ign) | not))
| "\(.vhost)/\(.name)\t\(.messages_ready)"')
# notify once per incident: a marker file per queue
while IFS=$'\t' read -r q ready; do
[ -n "$q" ] || continue
key="$STATE/$(printf '%s' "$q" | md5sum | cut -c1-32)"
[ -e "$key" ] && continue
printf '%s\n' "$q" > "$key"
curl -sf -X POST -H 'content-type: application/json' "$HOOK" \
-d "$(jq -nc --arg t ":rotating_light: $q has $ready messages and no consumer" '{text: $t}')"
done <<< "$ALERTING"
# recovered: drop markers for queues that are no longer alerting
for f in "$STATE"/*; do
[ -e "$f" ] || continue
grep -qF "$(cat "$f")" <<< "$ALERTING" || rm -f "$f"
done
Run it every minute from cron. The "for 5 minutes" part is missing here; add it by requiring the marker to be older than five minutes before posting, or accept the occasional deploy-time notification. Teams users swap the Slack payload for an Adaptive Card or a plain {"text": ...} to an incoming webhook.
Deciding which queues to exclude
Do this once and write it down, otherwise the alert gets muted the first week:
- Retry-wait queues (TTL, dead-letter back to the work queue): never have consumers. Exclude by name pattern.
- Dead-letter and parking-lot queues: no consumers, and messages in them are a different alert (see below). Exclude here, alert separately.
- Batch queues consumed on a schedule: exclude, or alert only outside the schedule.
- Everything else with a backlog and no consumer is an incident.
Name your queues so a regex can tell these apart. orders.process, orders.retry.wait, orders.dlq is a convention that makes every rule on this page a one-liner.
The companion alert: something landed in a DLQ
The no-consumer alert is the early warning. The late warning is messages_ready > 0 on a queue matching \.dlq$, because by then messages have already failed. Give that one a lower threshold than you think: one dead letter in a queue that is normally empty is news, three hundred is a postmortem. If you want fewer false alarms, alert on the rate of arrivals into the DLQ (message_stats.publish_details.rate on the DLQ, or rate(rabbitmq_detailed_queue_messages_published_total[5m])) rather than on the count.
Keep the history
Whatever alerts you build, keep the per-queue time series for at least a week. The question after a no-consumer page is always "since when", and the management UI only shows the last hour at any useful resolution. Prometheus gives you this for free. The cron script does not; if you go that route, at least append the numbers to a file per run.
Where Warren fits
If you do not run Prometheus, Warren's Team edition does this out of the box: alert rules match queues by regex on one or all clusters, with conditions for messages above a threshold, no consumers while messages wait, and growth within a window, each with a sustained-for time. Notifications go to Slack, Teams or a JSON webhook, repeat after a configurable interval and resolve when the queue recovers. Every queue is sampled every 30 seconds and kept for seven days, so "since when" has an answer.
Try Warren in a minute (one compose file, demo broker with real dead letters included) · README on GitHub






