The most common reason engineers leave a reliability-focused team is not incident load. It is the feeling of being permanently on call even when the rotation says otherwise.
A rotation that looks reasonable on a spreadsheet can still be corrosive, and the damage is usually invisible until someone resigns.
Teams argue about whether a rotation should be one week or two. The more important question is how many times a person is woken up during it.
Our working thresholds:
The correct response to a high page volume is not a larger rotation. It is to stop feature work and fix the noisiest source.
A page that says “CPU above 80%” is not actionable. A page that says “checkout error rate above 1% for five minutes, runbook here” is.
We apply a simple rule: if a responder cannot state what they would do in response to a page, the page should be a dashboard panel or a ticket, not a page. In practice this removes about half of a typical alerting configuration.
Page if:
- Customer-visible error rate breaches SLO burn rate
- Data loss or corruption is possible
- A security boundary is being crossed
- A dependency is failing in a way that requires human action
Ticket (next business day) if:
- A resource is above a warning threshold but not degrading service
- A certificate or credential expires in more than 14 days
- A background job is delayed but making progress
Dashboard only:
- Capacity utilisation
- Queue depth under normal operation
- Anything currently under active investigation
An escalation path that leads to someone who does not respond is worse than no path, because it teaches the responder that they are alone.
Requirements we insist on:
The single most effective change in most rotations is granting explicit permission to apply a mitigation rather than a fix. Restarting a component, shedding load, or disabling a non-essential feature at 03:00 is the correct decision. Root-cause analysis happens the following morning, with sleep.
A responder who is afraid to take the blunt action will spend three hours trying to be elegant, and will get it wrong.
On-call is work. It should be compensated, whether in pay or in time, and the arrangement should be written down.
We also recommend a hard rule on recovery: anyone paged after midnight is not expected to work a normal day afterwards. This is not generosity. Tired people cause the next incident.
Once a month, look at three numbers: total pages, pages per person, and the proportion of pages that turned out to be actionable. If the third number is below 70%, the rotation is generating noise and the fix is in alerting, not in people.
We are happy to talk through a specific problem without any expectation of an engagement.