Skip to content
Techsense Developers
TrustLet's Talk
Insights
Managed Services & SRE8 min readOct 3, 2026

How to Achieve 99.9% Uptime: An Incident Mitigation Checklist

Achieving 99.9% uptime comes down to one discipline: shrinking the time between when something breaks and when it is fully recovered. You get there by budgeting your allowable downtime,…

Achieving 99.9% uptime comes down to one discipline: shrinking the time between when something breaks and when it is fully recovered. You get there by budgeting your allowable downtime, instrumenting your systems so failures surface in seconds, and running a repeatable incident mitigation process that removes guesswork under pressure. Nothing about it is magic. It is the sum of small engineering decisions made before the pager goes off.

Let me put the number in perspective first, because "three nines" is easy to say and harder to earn.

What 99.9% Uptime Actually Costs You in Minutes

Service level objectives are easier to reason about when you convert percentages into wall-clock downtime. Here is the budget you are working with:

Availability Downtime per month Downtime per year
99% ("two nines") ~7h 18m ~3d 15h
99.9% ("three nines") ~43m 12s ~8h 46m
99.95% ~21m 36s ~4h 23m
99.99% ("four nines") ~4m 19s ~52m 36s

Three nines gives you roughly 43 minutes of error budget per month. That is not much. A single bad deploy, a slow database failover, or a 20-minute debate about who owns the alert can consume it. The goal of everything below is to protect those 43 minutes.

A quick note on math: these figures assume a 30-day month and a 365-day year. Round numbers like "about 9 hours a year" are close enough for planning. For the derivation and definitions, the Google SRE workbook on implementing SLOs is the reference I point teams to.

Start With an Error Budget, Not a Dashboard

The most common mistake I see is buying monitoring tools before defining what "up" means. Uptime management is a policy problem before it is a tooling problem.

Define your SLO in terms your users feel, not your servers report. "CPU under 80%" is not an SLO. "99.9% of checkout requests complete in under 800ms" is.

A workable SLO definition looks like this:

slo:
  name: checkout-availability
  objective: 0.999          # 99.9% over 28 days
  indicator:
    type: request_based
    good: 'http_status < 500 AND latency_ms < 800'
    total: 'all checkout requests'
  error_budget_minutes: 43  # per 30-day window
  burn_rate_alerts:
    - window: 1h
      threshold: 14.4       # burns a month's budget in ~2 days
    - window: 6h
      threshold: 6

The burn rate approach matters more than static thresholds. Alerting on "any error" produces noise. Alerting on "we are consuming error budget fast enough to blow the month" produces action. The multi-window, multi-burn-rate pattern catches both fast catastrophic failures and slow leaks.

The Incident Mitigation Checklist

Here is the checklist my teams run against. Work through each section. If you cannot check a box, that gap is where your next outage will come from.

1. Detection: find it before your customers do

  • Synthetic probes hit critical user journeys (login, checkout, API auth) every 30 to 60 seconds from multiple regions.
  • Burn-rate alerts are wired to your SLOs, not to raw infrastructure metrics.
  • Health checks distinguish between "process is alive" and "process can serve traffic." A load balancer should drain a node that returns 200 on /healthz but cannot reach its database.
# A readiness check that actually verifies dependencies,
# not just that the process answers.
curl -sf http://localhost:8080/ready || exit 1
# /ready internally checks: DB connection, cache, downstream auth
  • Alerts route to a single on-call rotation per service, with clear escalation after a timeout.

2. Response: remove the human latency

The gap between detection and action is where most error budgets die. Automate the first five minutes.

  • A runbook exists for every alert, linked directly in the alert payload. The engineer who gets paged at 3 a.m. should not have to think, only execute.
  • One person is the Incident Commander (IC). They coordinate; they do not debug. This separation is the single highest-leverage practice in incident response.
  • A dedicated incident channel is created automatically. Everything is logged there for the postmortem.
  • Severity levels are defined in advance so nobody argues about whether to escalate while the site is down.

A minimal severity rubric:

  • SEV1: Customer-facing outage, no workaround. Page everyone.
  • SEV2: Degraded service or major feature down. Page on-call + IC.
  • SEV3: Minor or internal impact. Handle in business hours.

3. Mitigation: stop the bleeding before you diagnose

Recovery is not root-cause analysis. Your first job is to restore service, then investigate. Mitigation levers you should have ready:

  1. Roll back the last deploy. Most incidents correlate with a change. If you can roll back in under two minutes, do that first and ask questions later.
  2. Feature flags to disable the offending code path without a redeploy.
  3. Traffic shifting to a healthy region or a previous version.
  4. Rate limiting or shedding load to protect the core path when a dependency is saturated.
# Rollback via deployment history (Kubernetes example)
kubectl rollout undo deployment/checkout-api
kubectl rollout status deployment/checkout-api --timeout=120s
  • Every deploy is reversible in minutes, not hours.
  • Feature flags gate risky changes so mitigation does not require a build.

4. Recovery and prevention: close the loop

  • A blameless postmortem is written for every SEV1 and SEV2 within 48 hours. Focus on the system, not the individual.
  • Each postmortem produces tracked action items with owners and due dates. Untracked follow-ups are how the same outage happens twice.
  • Chaos and failover drills run on a schedule. If you have never tested your database failover under load, you do not have a failover.

Engineering Decisions That Protect Uptime

The checklist handles incidents. These architectural choices reduce how many incidents reach your users in the first place.

Eliminate single points of failure

  • Run at least two instances of every stateless service across availability zones.
  • Use managed, multi-AZ data stores with automated failover, and actually measure your failover time.
  • Put health-aware load balancing in front of everything so a failing node is removed automatically.

Make deploys boring

Downtime reduction is mostly deploy-risk reduction. Progressive delivery is the mechanism:

  • Canary releases: ship to 5% of traffic, watch the SLO, then expand.
  • Automated rollback triggered by SLO regression during a canary.
  • Immutable artifacts so "it works on my machine" stops being a failure mode.

Set timeouts, retries, and circuit breakers

A slow dependency should fail fast, not drag your whole service down. Configure:

  • Timeouts on every outbound call. No unbounded waits.
  • Bounded retries with exponential backoff and jitter. Infinite retries cause retry storms.
  • Circuit breakers that trip open when a downstream is unhealthy, giving it room to recover.

Where a Managed Services and SRE Partner Fits

Not every team can staff a 24/7 on-call rotation or maintain the automation above while still shipping features. This is where getting help pays for itself. Our managed services and SRE capabilities exist to run this exact loop: SLO definition, detection, mitigation, and the unglamorous follow-through that keeps the error budget intact. We have applied these practices across regulated and high-traffic environments, and you can see sector-specific context on our industries page.

The point is not to outsource responsibility. It is to make sure the checklist above is owned by someone whose full-time job is keeping the lights on, so your engineers can build.

A Realistic Path to 99.9%

If you are starting from unmeasured uptime, do this in order:

  1. Define one SLO for your most critical user journey.
  2. Instrument it and watch the real number for a month before promising it to anyone.
  3. Wire burn-rate alerts and write runbooks for the top three failure modes.
  4. Make rollback fast and gate risky changes behind flags.
  5. Run a game day to prove your mitigation levers work.

Three nines is reachable for most teams within a quarter of focused work. The hard part is not the technology. It is the discipline to protect those 43 minutes every single month.

FAQ

How much downtime does 99.9% uptime allow?

About 43 minutes per month, or roughly 8 hours and 46 minutes per year, assuming a 30-day month and a 365-day year. That is your error budget. Every incident, slow deploy, and failed failover draws against it.

What is the difference between an SLA, an SLO, and an SLI?

An SLI (Service Level Indicator) is the measured number, such as the percentage of requests served under 800ms. An SLO (Objective) is your internal target for that indicator, like 99.9%. An SLA (Agreement) is a contractual promise to customers, usually set more conservatively than your SLO so you have margin.

Should I alert on infrastructure metrics or on SLOs?

Alert on SLOs using burn rate. Raw infrastructure alerts like high CPU produce noise and page people for conditions users never notice. Burn-rate alerts fire when you are consuming error budget fast enough to miss your target, which is the thing that actually matters.

What is the fastest way to reduce downtime without re-architecting?

Make rollback fast and automatic. Most incidents follow a change, so a two-minute rollback and feature flags for risky code paths give you immediate mitigation leverage without touching your architecture.

Do I need a dedicated SRE team to hit 99.9%?

No, but you need the practices an SRE team provides: defined SLOs, on-call coverage, runbooks, and blameless postmortems. Smaller teams often adopt these incrementally or bring in a managed services partner to own the operational loop while internal engineers focus on product.

Production-grade cloud, software, and engineering teams for scaling companies.