Almost every organisation we visit has heard of error budgets. Perhaps a third have implemented something they call one. Very few have a policy that changes anyone's behaviour.
The typical failure looks like this: a dashboard tracks the budget, it goes red in the second week of the quarter, and everyone carries on shipping features. The chart is decorative.
An error budget is a trade-off instrument. It only functions if the trade-off is real. That requires three things to be true at once, and most implementations only have the first.
Without the third item, the budget becomes a monthly ritual in which an engineer reports bad news and a product manager explains why this quarter is different.
The policy has to be agreed in advance, in writing, by someone senior enough that invoking it is not a career risk for the person invoking it.
A version that has worked for us:
Reliability policy — Dispatch API
SLO: 99.9% of successful requests, measured over a rolling 28 days.
Budget: 40 minutes of error-time per 28 days.
While budget > 25%:
Normal feature work. Changes ship through the standard path.
While budget between 25% and 0%:
Feature work continues. Any change touching the request path needs a
second reviewer and a documented rollback plan.
When budget is exhausted:
The service enters reliability mode. No feature changes to the request
path until the budget recovers above 25%.
The engineering manager may override, in writing, with a stated reason.
Overrides are reviewed at the next monthly reliability meeting.
Exceptions:
Security fixes and data-loss fixes always ship.
Work that directly improves reliability always ships.
It looks like a loophole. It is the opposite. An explicit, written, time-stamped override creates accountability. A silent one creates resentment and teaches the team that the policy is theatre.
In practice, overrides are rare and uncomfortable to write. That is the point.
Do not build this for fifty services. Pick the one path whose failure actually reaches customers, define a single objective on it, and run the policy for a full quarter before expanding.
The first quarter is mostly about discovering that your measurement is wrong. That is a worthwhile result on its own.
We are happy to talk through a specific problem without any expectation of an engagement.