Prometheus Alerting in Depth: Rules, Alertmanager, and Paging
Prometheus Alerting in Depth: Rules, Alertmanager, and Paging
Prometheus evaluates rules; Alertmanager decides who gets paged. Splitting alerting into these two parts is what makes the whole system powerful — this is the complete picture.
Table of Contents
- The flow
- Alerting rules
- Recording rules: precompute the expensive stuff
- Alertmanager: the traffic cop
- The three killer features
- Alert design rules that prevent fatigue
The flow
app /metrics → Prometheus scrapes → alerting rules evaluate
→ alerts fire → Alertmanager → routes/grouped/silenced
→ notifications (Slack, PagerDuty, email, webhook)
Three setup steps: configure Alertmanager, point Prometheus at it (alerting.alertmanagers in prometheus.yml), write rules.
Alerting rules
groups:
- name: app-alerts
rules:
- alert: HighErrorRate
expr: |
sum(rate(http_requests_total{status=~"5.."}[5m]))
/ sum(rate(http_requests_total[5m])) > 0.02
for: 10m
labels:
severity: critical
team: platform
annotations:
summary: "Error rate {{ $value | humanizePercentage }} on {{ $labels.job }}"
runbook: https://wiki.example.com/runbooks/high-error-rate
The pieces: expr is the PromQL test, for: is the "must be true this long" dampener, labels become routing keys, annotations are human text with templating ($labels, $value).
for: is the most important field. Without it, a one-scrape blip pages someone at 3 AM. 5–15 minutes filters noise; match it to your tolerance.
Recording rules: precompute the expensive stuff
- record: job:http_requests:rate5m
expr: sum by (job) (rate(http_requests_total[5m]))
Recording rules evaluate continuously and store results as new series — dashboards and alerts query the fast precomputed metric instead of re-aggregating raw data. Name them level:metric:op so provenance is obvious.
Alertmanager: the traffic cop
route:
receiver: default
group_by: [alertname, namespace]
group_wait: 30s # wait to batch related alerts
group_interval: 5m # repeat batch cadence
repeat_interval: 4h # re-notify if still firing
routes:
- match: { severity: critical }
receiver: pagerduty
- match_re: { team: "(platform|backend)" }
receiver: slack-platform
receivers:
- name: slack-platform
slack_configs:
- channel: "#alerts-platform"
send_resolved: true
- name: pagerduty
pagerduty_configs:
- routing_key: $PAGERDUTY_KEY
The three killer features
- Grouping — 50 pods alerting = one notification, not fifty.
- Silences — mute matchers for a window (maintenance, known breakage). In the UI: choose labels + duration.
- Inhibition — if severity: critical is firing, suppress matching warning alerts. Node down ⇒ don't also page about its pods.
inhibit_rules:
- source_match: { severity: "critical" }
target_match: { severity: "warning" }
equal: [alertname]
Alert design rules that prevent fatigue
- Page on symptoms, not causes. "Users getting errors" pages; "disk at 70%" is a ticket.
- Every alert needs an action. If there's nothing to do, it's a dashboard panel, not an alert.
- Severity tiers. critical = page now, warning = work queue, info = dashboard.
- Test rules with promtool: promtool test rules test.yml runs unit tests against alert expressions.
- Attach runbook links. An alert without a next step trains people to ignore alerts.
Good alerting is a product decision: every page must be actionable, novel, and worth interrupting a human. Build for that bar and the 3 AM calls become rare and real.
Related Articles
Last Updated: October 2026 Author: CloudOpsGuide Team Difficulty: Advanced Estimated Reading Time: 11 minutes