Production Monitoring with Prometheus and Grafana
Production Monitoring with Prometheus and Grafana
You can't fix what you can't see. Prometheus and Grafana are the de facto open-source stack for metrics: Prometheus collects and stores time-series data, Grafana turns it into dashboards humans actually read.
Table of Contents
- How Prometheus works
- The four metric types
- PromQL you'll use daily
- Exporters: metrics for everything
- Alerting with Alertmanager
- Grafana tips
How Prometheus works
Prometheus uses a pull model: it scrapes HTTP endpoints that expose metrics in a simple text format. Your app (or an exporter on its behalf) serves /metrics; Prometheus fetches it every 15 seconds and stores the samples.
# HELP http_requests_total Total HTTP requests
# TYPE http_requests_total counter
http_requests_total{method="GET",status="200"} 10273
http_requests_total{method="POST",status="500"} 4
The four metric types
- Counter — only goes up (requests served, errors). Rate it with rate().
- Gauge — goes up and down (memory used, queue depth).
- Histogram — samples into buckets (request latency). Powers percentile queries.
- Summary — like a histogram but calculates quantiles client-side.
PromQL you'll use daily
# requests per second, last 5 minutes
rate(http_requests_total[5m])
# error ratio
sum(rate(http_requests_total{status=~"5.."}[5m]))
/ sum(rate(http_requests_total[5m]))
# 95th percentile latency
histogram_quantile(0.95,
rate(http_request_duration_seconds_bucket[5m]))
Exporters: metrics for everything
You don't instrument everything yourself. node_exporter covers host CPU/memory/disk, kube-state-metrics covers Kubernetes objects, and there are exporters for nginx, Postgres, Redis, and hundreds more.
Alerting with Alertmanager
Alerts are PromQL expressions with a duration:
- alert: HighErrorRate
expr: rate(http_requests_total{status=~"5.."}[5m]) > 0.05
for: 10m
annotations:
summary: "Error rate above 5% on {{ $labels.instance }}"
Alertmanager deduplicates, groups, and routes them to Slack, PagerDuty, or email.
Grafana tips
- Build dashboards that answer questions ("is the API healthy?"), not dashboards that graph everything.
- Use the USE method for resources (Utilization, Saturation, Errors) and RED for services (Rate, Errors, Duration).
- Provision dashboards as code — JSON in git, loaded by Grafana's provisioning system.
Start small: node_exporter on one host, one Grafana dashboard, one alert. The stack grows naturally from there.
Related Articles
- SLOs and Error Budgets: Engineering Reliability
- Grafana Loki: Log Aggregation Without the Elasticsearch Tax
Last Updated: October 2026 Author: CloudOpsGuide Team Difficulty: Intermediate Estimated Reading Time: 12 minutes