SLOs and Error Budgets: Engineering Reliability
SLOs and Error Budgets: Engineering Reliability
"Make the site reliable" isn't a goal — it's a wish. SLOs turn reliability into a number you can engineer toward, and error budgets turn it into a decision-making tool for the eternal speed-vs-stability argument.
Table of Contents
The vocabulary
- SLI (indicator) — the measurement. E.g., "proportion of requests served under 300ms without error."
- SLO (objective) — the target. E.g., "99.9% of requests good over a rolling 30 days."
- SLA (agreement) — a contract with consequences. SLOs are internal; SLAs are what you promise customers — usually looser than your SLO.
Error budgets: the useful part
A 99.9% SLO means you're allowed 0.1% bad events — about 43 minutes of downtime per month. That's your error budget, and the policy is simple:
- Budget remaining → ship fast, launch features, run experiments.
- Budget exhausted → freeze risky changes, prioritize reliability work until the budget refills.
This ends the dev-vs-ops argument with data instead of opinion. Neither "always ship" nor "never break anything" wins — the budget decides.
Picking good SLIs
Measure what users experience, not what servers report:
- Availability: good_responses / total_responses
- Latency: responses_faster_than_threshold / total — a proportion, not a raw p99, so the SLI stays bounded 0–1.
- Freshness / correctness for data pipelines and batch jobs.
Alerting on budgets
Page on burn rate, not instantaneous errors. If the budget allows 0.1% errors/month, a fast-burn alert fires when you're consuming it >14x the sustainable rate — catching real incidents while ignoring blips:
# fast burn: ~2% of monthly budget per hour
sum(rate(http_errors[1h]))
/ sum(rate(http_requests[1h])) > 0.02
Common mistakes
- Too many SLOs. Start with 1–3 per user-facing service, on availability and latency.
- Aiming for 100%. Beyond ~99.9%, each nine costs more than the last and buys less user happiness than a new feature.
- SLOs nobody reviews. If missing an SLO has no consequence, it isn't an SLO — it's a metric.
Write down your current SLI numbers before setting targets. SLOs discovered from reality stick; SLOs declared from ambition get ignored.
Related Articles
- Production Monitoring with Prometheus and Grafana
- Grafana Loki: Log Aggregation Without the Elasticsearch Tax
Last Updated: October 2026 Author: CloudOpsGuide Team Difficulty: Advanced Estimated Reading Time: 9 minutes