Skip to content
Techsense Developers
TrustLet's Talk
Insights
Managed Services & SRE8 min readOct 2, 2026

What Are the Best Practices for Reducing SRE Pager Fatigue?

If your team dreads SRE pager duty rotations because the pager fires constantly, the fastest way to reduce pager fatigue is to cut the volume of alerts that reach a human. Most on-call pain is not…

If your team dreads SRE pager duty rotations because the pager fires constantly, the fastest way to reduce pager fatigue is to cut the volume of alerts that reach a human. Most on-call pain is not caused by too few engineers. It is caused by alerts that are not actionable, not urgent, or not owned. Fix the signal, and the staffing math gets much easier.

In my experience running on-call rotations, roughly the same pattern repeats across teams: a handful of noisy alerts generate the majority of pages, duplicates stack up during a single incident, and nobody has pruned the alert catalog in months. Below is a practical, ordered approach to reducing alert noise and protecting the people who carry the pager.

Why SRE Pager Duty Burns People Out

Pager fatigue is a measurable operational risk, not a morale footnote. When an engineer is woken three times a night for alerts that resolve themselves, two things happen:

  1. Response quality drops. Tired engineers make slower, riskier decisions during real incidents.
  2. Trust in alerts erodes. Once people learn that most pages are noise, they start ignoring the pager. The one page that mattered gets missed.

The goal of a healthy SRE pager duty program is simple to state and hard to execute: every page should be urgent, actionable, and owned by the person receiving it. If an alert fails any of those three tests, it does not belong on the pager.

The three-question test for any page

Before an alert is allowed to page a human, it must pass all three:

  • Is it urgent? Does it require action within minutes, not hours? If it can wait until morning, it is a ticket, not a page.
  • Is it actionable? Is there a clear, documented thing the responder can do right now?
  • Is it owned? Does the person on call actually have the authority and context to fix it?

Alerts that fail "urgent" become dashboards or daily digests. Alerts that fail "actionable" need a runbook before they are re-enabled. Alerts that fail "owned" are routed to the wrong team.

Alert on Symptoms, Not Causes

The single highest-leverage change most teams can make is to alert on user-facing symptoms rather than internal causes. Google's SRE practice calls this symptom-based alerting, and it dramatically reduces alert count because one symptom can be caused by dozens of root causes.

Consider the difference:

  • Cause-based (noisy): page on high CPU, page on a full disk on one node, page on a single pod restart, page on elevated garbage collection.
  • Symptom-based (quiet): page when the checkout API error rate exceeds your SLO burn rate, or when p99 latency crosses the threshold that users actually feel.

The symptom-based alert fires only when customers are affected. The underlying causes become diagnostic signals you investigate after you are paged, not independent pages in their own right.

Tie alerts to SLOs and error budgets

Service Level Objectives give you a principled way to decide when to page. Instead of static thresholds that fire on every blip, page on error budget burn rate. A fast burn means you are about to violate your SLO soon and someone should wake up. A slow burn can wait.

A simplified Prometheus alert using multi-window burn rate looks like this:

groups:
  - name: slo-burn-rate
    rules:
      - alert: HighErrorBudgetBurn
        expr: |
          (
            job:slo_errors:ratio_rate5m{job="checkout"} > (14.4 * 0.001)
            and
            job:slo_errors:ratio_rate1h{job="checkout"} > (14.4 * 0.001)
          )
        for: 2m
        labels:
          severity: page
        annotations:
          summary: "Checkout burning error budget fast"
          runbook: "https://runbooks.internal/checkout-slo-burn"

The 14.4 factor corresponds to a fast-burn threshold against a 99.9% SLO. The important part is not the exact number. It is that you page on budget consumption that threatens the objective, not on raw infrastructure metrics.

Cut the Noise: A Concrete Reduction Playbook

Here is the sequence I use when a rotation is drowning. Work it in order.

  1. Measure first. Pull 30 to 90 days of paging data. Rank alerts by page count, by pages outside business hours, and by how often each resolved with no action taken.
  2. Kill the self-resolvers. Any alert that regularly fires and then auto-resolves without human action is noise. Either widen the threshold, add a for: duration, or delete it.
  3. Deduplicate. Group related alerts into a single incident so one outage produces one page, not forty. Most incident platforms support alert grouping by service or correlation key.
  4. Add hysteresis and for durations. Require a condition to persist before paging. A CPU spike that lasts 15 seconds is not a page.
  5. Downgrade non-urgent alerts. Move anything that can wait until morning off the pager and into a ticket queue or a business-hours Slack channel.
  6. Require a runbook per page. No runbook, no page. This forces the author to prove the alert is actionable.

Example: suppressing downstream noise during an incident

When a dependency fails, dozens of dependent services can alert at once. Dependency-aware suppression keeps the page count sane:

inhibit_rules:
  - source_matchers:
      - 'severity="page"'
      - 'alertname="DatabasePrimaryDown"'
    target_matchers:
      - 'severity="page"'
    equal: ['cluster']

This tells Alertmanager: if the database primary is down, suppress the cascade of dependent pages for the same cluster. The responder sees the root cause, not the symptom storm.

Design the Rotation, Not Just the Alerts

Even a clean alert catalog will burn people out if the human side of the rotation is poorly designed.

  • Follow the sun where possible. Distribute rotations across time zones so nobody is routinely paged at 3 a.m. If your organization operates across regions, structured on-call handoffs can nearly eliminate overnight pages. Teams building out global coverage often lean on dedicated managed services and SRE capabilities to span time zones without over-hiring.
  • Cap consecutive on-call days. One week is a common ceiling. Rotate before fatigue compounds.
  • Separate primary and secondary. The secondary catches what the primary misses and provides a safety valve for escalation.
  • Budget recovery time. An engineer paged heavily overnight should not be expected to ship features the next morning. Protect that time explicitly.
  • Compensate on-call. Whether through time off or pay, acknowledging the burden reduces resentment and attrition.

Make incidents cheaper to handle

Reducing fatigue is also about reducing the cost of each page:

  • Runbooks linked in every alert. The responder should never start from a blank page.
  • Pre-built dashboards per service. One click to the relevant golden signals.
  • Automated remediation for known issues. If the fix is "restart the pod," automate it and only page if the automation fails.

Review and Iterate Relentlessly

Alert noise creeps back. Build a cadence to keep it out.

  • Weekly on-call review. The outgoing on-call presents every page: Was it actionable? Did it need to be a page? What should change? Track action items to completion.
  • Monthly alert audit. Delete alerts that fired zero actionable pages. Tune thresholds on the noisiest.
  • Track operational load as a metric. Pages per shift and pages outside business hours are leading indicators of burnout. Treat a rising trend as a bug.

Regulated and high-availability environments raise the stakes further. In sectors with strict uptime and audit requirements, such as the ones we cover under our industries work, disciplined on-call hygiene is part of meeting compliance obligations, not just engineering comfort.

A Short Summary of the Playbook

To reduce SRE pager fatigue:

  1. Alert on user-facing symptoms tied to SLOs, not infrastructure causes.
  2. Apply the urgent, actionable, owned test to every page.
  3. Run the reduction playbook: measure, kill self-resolvers, deduplicate, add durations, downgrade, require runbooks.
  4. Design humane rotations with coverage, caps, and recovery time.
  5. Review weekly and audit monthly so noise never creeps back.

The teams that stay healthy treat their alert catalog as a product with an owner, a backlog, and a quality bar. The pager is a scarce, expensive resource. Spend it only where a human's attention truly changes the outcome.

FAQ

How many pages per on-call shift is acceptable?

There is no universal number, but a widely used guideline from SRE practice is that on-call engineers should handle no more than roughly two significant incidents per shift, leaving enough time to investigate each properly. If your shifts routinely exceed that, treat it as a signal to reduce alert noise or add capacity, not as a normal baseline.

What is the difference between an alert that pages and one that files a ticket?

A page demands action within minutes and wakes someone up. A ticket captures work that can wait until business hours. The deciding question is urgency: if the issue can safely wait until morning without harming users, it should never page. Routing non-urgent conditions to a ticket queue is one of the fastest ways to reduce alert fatigue.

Should every alert have a runbook?

Yes, for anything that pages a human. If an alert cannot point to a documented, actionable response, it fails the "actionable" test and should not be on the pager. Requiring a runbook as a precondition for enabling a page also forces the author to prove the alert is worth a human's attention.

How do SLOs help reduce pager fatigue?

SLOs let you page on error budget burn rate rather than static thresholds. You alert only when the user-facing objective is genuinely at risk, which suppresses the flood of infrastructure-level pages that do not correspond to customer impact. This aligns on-call effort with what users actually experience.

How often should we review our alert configuration?

Hold a lightweight review after every on-call rotation, where the outgoing engineer evaluates each page for actionability, and a deeper audit monthly to prune dead alerts and tune noisy ones. Alert noise accumulates over time, so without a recurring cadence, fatigue returns regardless of how clean the catalog starts.