If you want to reduce on-call burnout, structure your sre pager duty rotation around three principles: keep per-engineer alert volume low enough to allow recovery, distribute load fairly across a team of at least six, and make every page actionable. A rotation that respects sleep, caps consecutive on-call days, and ruthlessly eliminates noisy alerts will outlast any heroics-based model. The rest of this post walks through how I build rotations that hold up under real production pressure.
Start With the Problem: Why On-Call Burnout Happens
Burnout rarely comes from a single bad week. It accumulates. The common root causes I see when auditing a struggling rotation are predictable:
- Too few people in the pool. A three-person rotation means every engineer is on call a third of the time, with no slack for vacation, illness, or departures.
- Alert noise. When half the pages are non-actionable, engineers stop trusting the pager. That erodes response quality and increases stress for the legitimate incidents.
- No recovery time. Back-to-back on-call weeks, or primary-then-secondary handoffs with no gap, prevent the nervous system from resetting.
- Unbounded night work. Pages at 3 a.m. that could have waited until morning are the single biggest driver of pager fatigue.
- Unclear escalation. When the on-call engineer doesn't know who to pull in, they absorb incidents alone that should have been shared.
Every structural decision below targets one or more of these causes directly.
Size the Rotation Before You Schedule It
The math matters more than the calendar tool. Before you touch scheduling software, figure out whether you have enough people to run a humane rotation.
The minimum viable team size
A sustainable primary rotation needs at least six engineers. With six, each person is on primary roughly one week in six, which leaves room for time off and absorbs a departure without collapsing. Below six, you are borrowing against future burnout.
If you run both a primary and secondary tier, you effectively need two pools or a larger single pool so that the same person is not primary and secondary in overlapping cycles.
Follow-the-sun vs. single-timezone
If your team spans time zones, use that to your advantage. A follow-the-sun model routes pages to whoever is awake, which nearly eliminates the worst burnout driver: sleep disruption.
# Example follow-the-sun coverage windows (UTC)
EU team: 06:00 - 14:00
US team: 14:00 - 22:00
APAC team: 22:00 - 06:00
If you are in a single time zone, you cannot eliminate night pages, but you can minimize them through alert design, which I cover below.
Design the Schedule to Protect Recovery
Once the pool is sized, the schedule itself should encode recovery as a first-class constraint.
Rotation length and handoff
I prefer one-week rotations for most teams. Weekly is long enough to maintain context on ongoing issues and short enough that a bad week ends soon. Daily rotations fragment ownership. Multi-week rotations extend the fatigue window.
Build in a handoff ritual at the start and end of each shift. A five-minute written handoff prevents the incoming engineer from being blindsided.
## On-call handoff — 2025-01-13
- Open incidents: INC-4821 (degraded checkout latency, mitigated, RCA pending)
- Flaky alert: `db-replica-lag` fired twice overnight, non-actionable, ticket OPS-992 to tune
- Planned work: payments migration Thu 14:00 UTC, expect elevated error rate
- Nothing else outstanding
Hard limits on consecutive on-call time
Set explicit caps and enforce them in the scheduler:
- No more than one week of primary per three-week window.
- A minimum 24-hour gap between a primary rotation ending and any secondary rotation beginning.
- No primary on-call during the week immediately after a major incident the engineer led. Give them time to write the postmortem without carrying the pager.
Compensate the work honestly
On-call is labor. Whether you compensate with pay, time off in lieu, or reduced sprint commitments, make it explicit. Expecting engineers to carry the pager on top of a full project load is a reliable path to attrition.
Make Every Page Actionable
The fastest way to reduce burnout is to send fewer, better pages. This is where most of the durable improvement lives.
Separate pages from notifications
Not everything that is interesting deserves to wake someone. Split your alert severities clearly:
- P1 / page: customer-impacting, needs human action now. Wakes the on-call.
- P2 / ticket: needs attention within business hours. Creates a ticket, no page.
- P3 / dashboard: informational. Visible but silent.
A simple alerting rule in Prometheus Alertmanager encodes this:
route:
receiver: ticket-queue
routes:
- matchers:
- severity="page"
receiver: pagerduty-primary
continue: false
- matchers:
- severity="ticket"
receiver: ticket-queue
Alert on symptoms, not causes
Page on what the user experiences: elevated error rate, latency past SLO, failed checkouts. Do not page on every CPU spike or disk-usage warning. Cause-based alerts multiply quickly and most of them never correlate with real impact.
# Symptom-based: page when the SLO is actually threatened
- alert: CheckoutErrorBudgetBurn
expr: |
sum(rate(http_requests_total{route="/checkout",code=~"5.."}[5m]))
/
sum(rate(http_requests_total{route="/checkout"}[5m])) > 0.02
for: 10m
labels:
severity: page
annotations:
summary: "Checkout 5xx rate above 2% for 10m"
runbook: "https://runbooks.internal/checkout-5xx"
Attach a runbook to every page
If a page fires without a linked runbook, that is a bug in your alerting. Every alert annotation should link to concrete remediation steps. This turns a 2 a.m. page from an open-ended investigation into a checklist.
Track and prune relentlessly
Review alert volume in every retrospective. For each alert, ask:
- Did it fire this period?
- Was it actionable?
- If not, do we tune the threshold, downgrade it to a ticket, or delete it?
Treating noisy alerts as defects is the single highest-leverage habit for a healthy sre pager duty program. A quiet pager is a feature, not a sign of neglect.
Build the Incident Response Team Around the Rotation
The on-call engineer should never feel like the last line of defense. A functioning incident response team has clear roles so that no one carries a major incident alone.
- Incident commander: coordinates the response, owns communication. Often a separate rotation from the primary responder.
- Primary responder: the on-call engineer who first receives the page.
- Secondary / escalation: backs up the primary and is paged if the primary does not acknowledge within a set window, commonly 5 to 15 minutes.
- Subject-matter experts: pulled in on demand, not on a schedule.
Encode automatic escalation so the primary knows help arrives whether or not they ask:
escalation_policy:
- level: 1
target: primary-oncall
timeout_minutes: 10
- level: 2
target: secondary-oncall
timeout_minutes: 10
- level: 3
target: engineering-manager
If you are standing up this capability from scratch, our managed services and SRE capabilities outline how we approach rotation design, alert hygiene, and incident response as a system rather than a collection of tools.
Measure What Matters
You cannot improve a rotation you do not measure. Track these over time and review them with the team, not at them:
- Pages per on-call shift, split by day and night.
- Percentage of pages that were actionable.
- Mean time to acknowledge, as a proxy for responder fatigue.
- Sleep-disrupting pages per engineer per quarter.
When night pages trend up, treat it as an incident in its own right. The goal is a rotation where the common experience is an uneventful week, not a survived one.
Different sectors carry different load profiles. A payments platform and a batch-oriented analytics product have very different overnight risk, and the rotation should reflect that. We discuss sector-specific reliability patterns across the industries we work with.
A Practical Rollout Checklist
- Confirm the pool is at least six engineers; if not, fix staffing before scheduling.
- Choose weekly rotations with a written handoff ritual.
- Enforce consecutive-shift caps and recovery gaps in the scheduler.
- Split alerts into page, ticket, and dashboard tiers.
- Convert cause-based alerts to symptom-based SLO alerts.
- Attach a runbook link to every page.
- Define escalation roles and automate escalation timeouts.
- Review alert noise and on-call metrics every retrospective.
Do these in order. Most teams want to jump to the scheduling tool, but staffing and alert hygiene determine whether any schedule can actually be humane.
FAQ
How many engineers do I need for a sustainable on-call rotation?
Aim for at least six per primary rotation. That keeps each engineer on primary roughly one week in six and leaves slack for vacation, illness, and attrition. Fewer than six is workable only as a temporary state while you hire, and you should treat it as a known risk, not a steady state.
Should primary and secondary on-call be the same person?
No. The secondary exists to back up the primary during major incidents and to catch missed acknowledgments. Assigning both roles to one person defeats the purpose and increases the load that drives pager fatigue. Use a separate schedule or a large enough pool that the two never overlap for the same engineer.
How do I reduce night pages specifically?
Attack it from two directions. First, move everything that can wait until morning to a ticket tier so it never pages overnight. Second, if your team spans time zones, adopt a follow-the-sun model so pages route to engineers who are awake. What remains should be genuine, customer-impacting incidents with a linked runbook.
Is on-call compensation necessary?
Yes, in some form. On-call is real labor and restricts an engineer's personal time even when the pager stays quiet. Whether you use additional pay, time off in lieu, or reduced sprint commitments, make the compensation explicit and consistent. Unpaid, uncredited on-call is one of the most reliable drivers of attrition.
Production-grade cloud, software, and engineering teams for scaling companies.



