CloudOpsGuide
observability

Production Monitoring with Prometheus and Grafana

Intermediate
12 minutes
October 2026
CloudOpsGuide Team

Production Monitoring with Prometheus and Grafana

You can't fix what you can't see. Prometheus and Grafana are the de facto open-source stack for metrics: Prometheus collects and stores time-series data, Grafana turns it into dashboards humans actually read.

Table of Contents

How Prometheus works

Prometheus uses a pull model: it scrapes HTTP endpoints that expose metrics in a simple text format. Your app (or an exporter on its behalf) serves /metrics; Prometheus fetches it every 15 seconds and stores the samples.

# HELP http_requests_total Total HTTP requests
# TYPE http_requests_total counter
http_requests_total{method="GET",status="200"} 10273
http_requests_total{method="POST",status="500"} 4

The four metric types

  • Counter — only goes up (requests served, errors). Rate it with rate().
  • Gauge — goes up and down (memory used, queue depth).
  • Histogram — samples into buckets (request latency). Powers percentile queries.
  • Summary — like a histogram but calculates quantiles client-side.

PromQL you'll use daily

# requests per second, last 5 minutes
rate(http_requests_total[5m])

# error ratio
sum(rate(http_requests_total{status=~"5.."}[5m]))
  / sum(rate(http_requests_total[5m]))

# 95th percentile latency
histogram_quantile(0.95,
  rate(http_request_duration_seconds_bucket[5m]))

Exporters: metrics for everything

You don't instrument everything yourself. node_exporter covers host CPU/memory/disk, kube-state-metrics covers Kubernetes objects, and there are exporters for nginx, Postgres, Redis, and hundreds more.

Alerting with Alertmanager

Alerts are PromQL expressions with a duration:

- alert: HighErrorRate
  expr: rate(http_requests_total{status=~"5.."}[5m]) > 0.05
  for: 10m
  annotations:
    summary: "Error rate above 5% on {{ $labels.instance }}"

Alertmanager deduplicates, groups, and routes them to Slack, PagerDuty, or email.

Grafana tips

  • Build dashboards that answer questions ("is the API healthy?"), not dashboards that graph everything.
  • Use the USE method for resources (Utilization, Saturation, Errors) and RED for services (Rate, Errors, Duration).
  • Provision dashboards as code — JSON in git, loaded by Grafana's provisioning system.

Start small: node_exporter on one host, one Grafana dashboard, one alert. The stack grows naturally from there.

Related Articles


Last Updated: October 2026 Author: CloudOpsGuide Team Difficulty: Intermediate Estimated Reading Time: 12 minutes