CloudOpsGuide
observability

SLOs and Error Budgets: Engineering Reliability

Advanced
9 minutes
October 2026
CloudOpsGuide Team

SLOs and Error Budgets: Engineering Reliability

"Make the site reliable" isn't a goal — it's a wish. SLOs turn reliability into a number you can engineer toward, and error budgets turn it into a decision-making tool for the eternal speed-vs-stability argument.

Table of Contents

The vocabulary

  • SLI (indicator) — the measurement. E.g., "proportion of requests served under 300ms without error."
  • SLO (objective) — the target. E.g., "99.9% of requests good over a rolling 30 days."
  • SLA (agreement) — a contract with consequences. SLOs are internal; SLAs are what you promise customers — usually looser than your SLO.

Error budgets: the useful part

A 99.9% SLO means you're allowed 0.1% bad events — about 43 minutes of downtime per month. That's your error budget, and the policy is simple:

  • Budget remaining → ship fast, launch features, run experiments.
  • Budget exhausted → freeze risky changes, prioritize reliability work until the budget refills.

This ends the dev-vs-ops argument with data instead of opinion. Neither "always ship" nor "never break anything" wins — the budget decides.

Picking good SLIs

Measure what users experience, not what servers report:

  • Availability: good_responses / total_responses
  • Latency: responses_faster_than_threshold / total — a proportion, not a raw p99, so the SLI stays bounded 0–1.
  • Freshness / correctness for data pipelines and batch jobs.

Alerting on budgets

Page on burn rate, not instantaneous errors. If the budget allows 0.1% errors/month, a fast-burn alert fires when you're consuming it >14x the sustainable rate — catching real incidents while ignoring blips:

# fast burn: ~2% of monthly budget per hour
sum(rate(http_errors[1h]))
  / sum(rate(http_requests[1h])) > 0.02

Common mistakes

  • Too many SLOs. Start with 1–3 per user-facing service, on availability and latency.
  • Aiming for 100%. Beyond ~99.9%, each nine costs more than the last and buys less user happiness than a new feature.
  • SLOs nobody reviews. If missing an SLO has no consequence, it isn't an SLO — it's a metric.

Write down your current SLI numbers before setting targets. SLOs discovered from reality stick; SLOs declared from ambition get ignored.

Related Articles


Last Updated: October 2026 Author: CloudOpsGuide Team Difficulty: Advanced Estimated Reading Time: 9 minutes