← All insights

Error budgets only work if somebody is allowed to stop

28 August 2025 · Reliability · 8 min read

Almost every organisation we visit has heard of error budgets. Perhaps a third have implemented something they call one. Very few have a policy that changes anyone's behaviour.

The typical failure looks like this: a dashboard tracks the budget, it goes red in the second week of the quarter, and everyone carries on shipping features. The chart is decorative.

Why it fails

An error budget is a trade-off instrument. It only functions if the trade-off is real. That requires three things to be true at once, and most implementations only have the first.

  • The budget is measured correctly. This is the part everybody does.
  • There is a defined action when it is exhausted. Most implementations stop here.
  • Someone has the authority and cover to trigger that action. This is the part that is almost always missing.

Without the third item, the budget becomes a monthly ritual in which an engineer reports bad news and a product manager explains why this quarter is different.

Making the policy real

The policy has to be agreed in advance, in writing, by someone senior enough that invoking it is not a career risk for the person invoking it.

A version that has worked for us:

Reliability policy — Dispatch API

SLO:    99.9% of successful requests, measured over a rolling 28 days.
Budget: 40 minutes of error-time per 28 days.

While budget > 25%:
    Normal feature work. Changes ship through the standard path.

While budget between 25% and 0%:
    Feature work continues. Any change touching the request path needs a
    second reviewer and a documented rollback plan.

When budget is exhausted:
    The service enters reliability mode. No feature changes to the request
    path until the budget recovers above 25%.
    The engineering manager may override, in writing, with a stated reason.
    Overrides are reviewed at the next monthly reliability meeting.

Exceptions:
    Security fixes and data-loss fixes always ship.
    Work that directly improves reliability always ships.

The override clause matters

It looks like a loophole. It is the opposite. An explicit, written, time-stamped override creates accountability. A silent one creates resentment and teaches the team that the policy is theatre.

In practice, overrides are rare and uncomfortable to write. That is the point.

Start smaller than you think

Do not build this for fifty services. Pick the one path whose failure actually reaches customers, define a single objective on it, and run the policy for a full quarter before expanding.

The first quarter is mostly about discovering that your measurement is wrong. That is a worthwhile result on its own.

Next step

Working on something related?

We are happy to talk through a specific problem without any expectation of an engagement.

Get in touch