CloudOpsGuide
observability

Prometheus Alerting in Depth: Rules, Alertmanager, and Paging

Advanced
11 minutes
October 2026
CloudOpsGuide Team

Prometheus Alerting in Depth: Rules, Alertmanager, and Paging

Prometheus evaluates rules; Alertmanager decides who gets paged. Splitting alerting into these two parts is what makes the whole system powerful — this is the complete picture.

Table of Contents

The flow

app /metrics → Prometheus scrapes → alerting rules evaluate
    → alerts fire → Alertmanager → routes/grouped/silenced
    → notifications (Slack, PagerDuty, email, webhook)

Three setup steps: configure Alertmanager, point Prometheus at it (alerting.alertmanagers in prometheus.yml), write rules.

Alerting rules

groups:
  - name: app-alerts
    rules:
      - alert: HighErrorRate
        expr: |
          sum(rate(http_requests_total{status=~"5.."}[5m]))
            / sum(rate(http_requests_total[5m])) > 0.02
        for: 10m
        labels:
          severity: critical
          team: platform
        annotations:
          summary: "Error rate {{ $value | humanizePercentage }} on {{ $labels.job }}"
          runbook: https://wiki.example.com/runbooks/high-error-rate

The pieces: expr is the PromQL test, for: is the "must be true this long" dampener, labels become routing keys, annotations are human text with templating ($labels, $value).

for: is the most important field. Without it, a one-scrape blip pages someone at 3 AM. 5–15 minutes filters noise; match it to your tolerance.

Recording rules: precompute the expensive stuff

- record: job:http_requests:rate5m
  expr: sum by (job) (rate(http_requests_total[5m]))

Recording rules evaluate continuously and store results as new series — dashboards and alerts query the fast precomputed metric instead of re-aggregating raw data. Name them level:metric:op so provenance is obvious.

Alertmanager: the traffic cop

route:
  receiver: default
  group_by: [alertname, namespace]
  group_wait: 30s          # wait to batch related alerts
  group_interval: 5m       # repeat batch cadence
  repeat_interval: 4h      # re-notify if still firing
  routes:
    - match: { severity: critical }
      receiver: pagerduty
    - match_re: { team: "(platform|backend)" }
      receiver: slack-platform

receivers:
  - name: slack-platform
    slack_configs:
      - channel: "#alerts-platform"
        send_resolved: true
  - name: pagerduty
    pagerduty_configs:
      - routing_key: $PAGERDUTY_KEY

The three killer features

  • Grouping — 50 pods alerting = one notification, not fifty.
  • Silences — mute matchers for a window (maintenance, known breakage). In the UI: choose labels + duration.
  • Inhibition — if severity: critical is firing, suppress matching warning alerts. Node down ⇒ don't also page about its pods.
inhibit_rules:
  - source_match: { severity: "critical" }
    target_match: { severity: "warning" }
    equal: [alertname]

Alert design rules that prevent fatigue

  • Page on symptoms, not causes. "Users getting errors" pages; "disk at 70%" is a ticket.
  • Every alert needs an action. If there's nothing to do, it's a dashboard panel, not an alert.
  • Severity tiers. critical = page now, warning = work queue, info = dashboard.
  • Test rules with promtool: promtool test rules test.yml runs unit tests against alert expressions.
  • Attach runbook links. An alert without a next step trains people to ignore alerts.

Good alerting is a product decision: every page must be actionable, novel, and worth interrupting a human. Build for that bar and the 3 AM calls become rare and real.

Related Articles


Last Updated: October 2026 Author: CloudOpsGuide Team Difficulty: Advanced Estimated Reading Time: 11 minutes