Latency, Load Testing and Capacity

SLOs and error budgets for model services

An SLO is a promise about how reliable your service will be, and an error budget is the small, spendable amount of failure that promise allows before you have to stop and fix things.

On this page 8
  1. The short answer
  2. The analogy you have already lived
  3. Why it exists
  4. How it works
  5. A real example you have seen
  6. The honest part
  7. Remember this
  8. What to learn next

One lesson, three depths. Pick the one that fits you today — you can switch any time.

Beginner — No maths. Plain English.

The short answer

An SLO is a promise about reliability. An error budget is how much failure that promise still allows.

The analogy you have already lived

You have relied on a train or metro service that promises most trains arrive within five minutes of schedule. That promise does not mean every single train is perfectly on time — it allows a small, known amount of lateness. Say a signal failure makes half the trains late in one bad week. The service has now spent a large chunk of its whole month's allowance in a single event.

That allowance is the whole idea behind this lesson, applied to a model service instead of a railway.

Why it exists

SLO stands for service-level objective — a specific, measurable promise about how well a service performs, over some window of time. A common shape: "99.9% of requests succeed, measured over 30 days."

Nobody can promise 100%. Hardware fails, networks glitch, deploys occasionally go wrong. An SLO admits that honestly, and states exactly how much imperfection is acceptable — not zero, but not unlimited either.

The gap between the promise and perfection is the error budget: the number of failures still allowed before the promise is broken. A 99.9% SLO over a million requests allows 1,000 failures. Spend them wisely, and a bad day is absorbed. Spend them all in one incident, and every failure afterward breaks the promise.

How it works

SLO: 99.9% success over 30 days, 1,500,000 requests total

  error budget = 0.1% of 1,500,000 = 1,500 allowed failures

  day  1-17: a few failures each day, mostly on track
  day  18:   a bad deploy -- 900 failures in ONE day
  day 19-30: back to normal

  budget spent: 1,376 of 1,500 (92%)
  budget remaining: 124 (8%)  <- barely any room left for the rest of the month

A real example you have seen

A cloud service's public status page reporting "99.95% uptime this month" is reporting against exactly this kind of promise. The number is not a grade — it is an accounting of a budget, spent or unspent.

The honest part

An error budget that is never spent is not automatically a success. Teams that treat "zero failures" as the only acceptable outcome often become too cautious to ship anything new, since every change carries some risk. A healthy team spends some of its budget on purpose, shipping real improvements, while a completely exhausted budget is the signal to stop and stabilise.

Remember this

  • An SLO is a specific, measurable promise about reliability, over a stated time window.
  • The error budget is how much failure that promise still allows — not zero.
  • A budget spent mostly in one bad incident leaves very little room for the rest of the period.

What to learn next

Developer — Code and libraries.

Setup

No installation needed — this uses only Python's built-in sqlite3 module.

Computing a real error budget from a month of traffic

error_budget.py
import sqlite3

SLO_TARGET = 0.999   # 99.9% of requests must succeed, measured over 30 days
WINDOW_DAYS = 30

# A synthetic 30 days of traffic: mostly healthy, with one bad day where
# a bug shipped (day 18) and a smaller blip from a flaky dependency (day 24).
conn = sqlite3.connect(":memory:")
conn.execute("CREATE TABLE daily_stats (day INTEGER, total INTEGER, failed INTEGER)")

daily = []
for day in range(1, WINDOW_DAYS + 1):
    total = 50000
    failed = 12                # a normal, healthy day
    if day == 18:
        failed = 900            # the bad deploy
    elif day == 24:
        failed = 140             # a rough day, not a full incident
    daily.append((day, total, failed))
conn.executemany("INSERT INTO daily_stats VALUES (?, ?, ?)", daily)
conn.commit()

total_requests = conn.execute("SELECT SUM(total) FROM daily_stats").fetchone()[0]
total_failed = conn.execute("SELECT SUM(failed) FROM daily_stats").fetchone()[0]

allowed_failure_rate = 1 - SLO_TARGET
error_budget = round(total_requests * allowed_failure_rate)
spent = total_failed
remaining = error_budget - spent

print(f"SLO: {SLO_TARGET*100:.1f}% success over {WINDOW_DAYS} days")
print(f"total requests in window : {total_requests:,}")
print(f"error budget (allowed failures) : {error_budget:,}")
print(f"failures actually seen          : {spent:,}")
print(f"budget remaining                : {remaining:,} "
      f"({remaining/error_budget*100:.1f}% of budget left)")
print()

print("days that used more than 5% of the WHOLE 30-day budget in a single day:")
for day, total, failed in daily:
    share = failed / error_budget * 100
    if share > 5:
        print(f"  day {day:2d}: {failed} failures = {share:.1f}% of the entire month's budget, in one day")
Output
SLO: 99.9% success over 30 days
total requests in window : 1,500,000
error budget (allowed failures) : 1,500
failures actually seen          : 1,376
budget remaining                : 124 (8.3% of budget left)

days that used more than 5% of the WHOLE 30-day budget in a single day:
  day 18: 900 failures = 60.0% of the entire month's budget, in one day
  day 24: 140 failures = 9.3% of the entire month's budget, in one day

Real, deterministic arithmetic over synthetic data — the traffic numbers are made up for this demo, but every computed number here follows exactly from them. With only 8.3% of the month's budget left after day 24, this team has almost no room for anything else to go wrong for the rest of the month.

Line-by-line walkthrough

error_budget = round(total_requests * allowed_failure_rate) is the entire idea in one line: the promised failure rate, applied to the actual traffic volume, gives a concrete count — not a percentage — of failures still allowed.

The loop over daily finds days that alone consumed a large share of the whole month's budget — a single bad day can dominate an entire reporting window, which raw daily failure counts alone would not make obvious.

Common mistakes

Setting an SLO with no measurement behind it. Picking "99.99%" because it sounds impressive, without knowing your actual historical failure rate, sets a target that is either meaningless to hit or immediately broken — neither is useful.

Treating the error budget as a target to hit exactly. The budget is a ceiling, not a goal. Spending exactly 100% of it every single month, on purpose, leaves zero margin for a genuine surprise.

Measuring the SLO over too short a window. A single bad hour can look catastrophic in an hourly view and barely register over 30 days. Choose the window to match what you are actually promising, and to whom.

Confusing an SLO with an SLA. An SLO is an internal target a team holds itself to. An SLA — service-level agreement — is an external, often contractual commitment, frequently with financial penalties attached. SLOs are usually set stricter than any corresponding SLA, to leave warning room before the SLA itself is at risk.

Try it yourself

Change SLO_TARGET from 0.999 to 0.9999 and rerun. Watch how much smaller the error budget becomes for the same traffic — and how a single incident like day 18 now blows through the entire month's allowance on its own.

What to learn next

Researcher — Mathematics and papers.

SLIs, SLOs, and error budgets as a formal hierarchy

The Google SRE framework defines three related terms precisely: an SLI (service-level indicator) is a directly measured quantity — request success rate, a latency percentile. An SLO is a target range or threshold for an SLI over a window. An error budget is $1 - \text{SLO}$, expressed as an absolute count against actual traffic, exactly as computed above. This hierarchy is deliberately measurement-first: an SLO with no corresponding SLI is not falsifiable, and is not a real SLO.

Multi-window, multi-burn-rate alerting

Naive alerting on "SLO violated" fires only after the fact, once damage over the whole window is done. The multi-window burn-rate approach (Google SRE Workbook, Chapter 5) instead alerts on the rate at which budget is being consumed, checked over multiple windows simultaneously — for example, a short 5-minute window catching a fast, severe outage, alongside a longer 1-hour window catching a slower, sustained degradation. A burn rate of $14.4\times$ the sustainable rate, sustained for an hour, exhausts a 30-day budget in about two days — the multiplier that most production burn-rate alerting is built around, chosen so it neither pages on noise nor waits too long to page on a real event.

Error budgets as a release-gating mechanism

A common and effective policy: while error budget remains, feature releases proceed on the normal cadence; once budget is exhausted, releases pause and all engineering effort redirects to reliability work until budget recovers. This converts an abstract reliability goal into an automatic, non-negotiable process trigger — removing the recurring human argument about whether "now" is a safe time to ship.

Composing SLOs for ML-specific failure modes

A model service's SLO is often not one number. Latency SLOs (a p99 threshold), availability SLOs (successful HTTP responses), and quality SLOs — outputs within an acceptable accuracy or calibration range, distinct from whether a response was returned at all — need separate SLIs, since a service can return a fast, well-formed, and confidently wrong answer that no availability SLO would ever catch. See monitoring and model drift for the quality side of this measurement.

References

  • Google, Site Reliability Engineering, 2016, Chapter 4, Service Level Objectives — sre.google/sre-book/service-level-objectives
  • Google, The Site Reliability Workbook, 2018, Chapter 5, Alerting on SLOs — the multi-window burn-rate methodology in full detail.

What to learn next