The difference between 99.9 uptime and 99.99% uptime is not a rounding error. It is roughly 43 minutes of allowed downtime per month versus 4 minutes. To calculate your SLA, multiply your measurement window by the uptime percentage, subtract from the total, and the remainder is your error budget: the amount of downtime you are permitted before you breach. That single calculation determines your architecture, your on-call rotation, and a large share of your infrastructure bill. This post shows you how to run the numbers, read the downtime table, and decide which service level is actually worth paying for.
What an SLA Actually Measures
Before you can calculate anything, you need to agree on what "up" means. An SLA (Service Level Agreement) is a contractual promise. It sits on top of two other concepts that teams routinely confuse:
- SLI (Service Level Indicator): the raw measurement. For example, the ratio of successful HTTP requests to total requests.
- SLO (Service Level Objective): the internal target you hold yourself to, usually stricter than the SLA.
- SLA: the external commitment, with penalties attached if you miss it.
A common mistake is defining uptime as "the server responds to a ping." That tells you nothing about whether users can log in, check out, or query their data. A credible SLI is request-based or journey-based:
availability = good_events / valid_events
# request-based example
availability = (successful_requests) / (total_requests)
If you measure availability by a health-check endpoint that never touches your database, your reported uptime will look great while customers are failing. Define the SLI against the behavior your users care about, not the behavior that is easy to measure.
The Uptime Downtime Table You Should Memorize
The core of SLA calculation is converting a percentage into allowed downtime. The formula is simple:
allowed_downtime = measurement_window * (1 - uptime_fraction)
# 99.9% monthly, using a 30-day month
window = 30 days = 43,200 minutes
downtime = 43,200 * (1 - 0.999) = 43.2 minutes
Here is the reference table. Keep it somewhere visible.
| Uptime | Per year | Per month (30d) | Per week | Per day |
|---|---|---|---|---|
| 99% ("two nines") | 3d 15h 39m | 7h 18m | 1h 40m | 14m 24s |
| 99.9% ("three nines") | 8h 45m | 43m 49s | 10m 4s | 1m 26s |
| 99.95% | 4h 22m | 21m 54s | 5m 2s | 43s |
| 99.99% ("four nines") | 52m 35s | 4m 23m | 1m 0s | 8.6s |
| 99.999% ("five nines") | 5m 15s | 26s | 6s | 0.86s |
Two things jump out. First, each additional nine cuts your allowed downtime by roughly 90%. Second, at four nines and above, the per-incident budget is smaller than the time it takes a human to read a page, open a laptop, and understand what is broken. That is the single most important fact in this entire discussion.
The Real Cost of Each Nine
The reason "just add another nine" is dangerous is that cost does not scale linearly with availability. It scales with the engineering, redundancy, and operational maturity required to stop relying on humans.
99.9% is a human-in-the-loop target
At 43 minutes of monthly budget, a well-run team can hit 99.9 uptime with:
- A single region with automated restarts and health checks.
- A monitoring and alerting stack that pages a human.
- An on-call engineer who can respond within 15 to 20 minutes.
This is achievable for most web applications without exotic architecture. The dominant cost is a functioning on-call rotation and good observability, not duplicated infrastructure.
99.99% forces automation
Four nines gives you about 4 minutes per month. No human responds, diagnoses, and fixes in 4 minutes. To hit this number you must remove the human from the critical recovery path:
- Multi-AZ or multi-region deployment with automatic failover.
- Load balancer health checks that eject bad instances in seconds.
- Automated rollbacks tied to error-rate and latency signals.
- Redundant dependencies, including databases with replicas and automated promotion.
Your infrastructure cost roughly doubles for the redundant capacity, and you add continuous engineering investment to keep the automation trustworthy.
99.999% is a business decision, not an engineering one
Five nines means 26 seconds per month. At this level, every dependency in your stack must also offer five nines, including DNS, payment processors, and third-party APIs. A single synchronous call to a service with a weaker SLA mathematically caps your own availability. The cost here is systemic: redundant providers, active-active topologies across regions, chaos testing, and a deep SRE practice. Most products do not need this, and paying for it where it is not needed is one of the most common forms of waste we see in production systems. We cover how to right-size this in our managed services and SRE capabilities.
How to Calculate the SLA You Should Actually Promise
Do not pick a number because it looks impressive on a sales page. Work backward from three inputs.
- Measure your current availability for 90 days. You cannot promise what you have never measured. Compute your SLI from real request logs, not uptime monitors.
- Subtract a safety margin for your SLO. Your internal SLO should be stricter than your external SLA. If you promise 99.9%, target 99.95% internally so you have room before you breach a contract.
- Price the penalty. An SLA with no teeth is marketing. A real SLA includes service credits.
A simple service-credit schedule looks like this:
Monthly uptime Service credit
>= 99.9% 0%
99.0% - 99.9% 10% of monthly fee
95.0% - 99.0% 25% of monthly fee
< 95.0% 50% of monthly fee
Account for dependency math
Your availability is bounded by your dependencies when they are in series. If your service depends on three components, each at 99.9%, your theoretical ceiling is:
0.999 * 0.999 * 0.999 = 0.997 (99.7%)
Three healthy dependencies at three nines each give you less than three nines overall. This is why redundancy matters: parallel (redundant) components multiply availability upward, while serial dependencies multiply it down. Model this before you sign anything. Requirements differ sharply across sectors, which is why we align targets to the operational realities of each vertical we work in, detailed on our industries page.
Uptime Management: Making the Number Real
Calculating an SLA is the easy part. Defending it in production is the discipline. A few practices separate teams that hit their targets from teams that merely publish them:
- Track your error budget. If your SLO is 99.9%, your monthly error budget is 43 minutes. Burn it on planned risk, not surprises. When the budget is exhausted, freeze risky deploys until it recovers.
- Alert on budget burn rate, not raw errors. A fast burn (consuming a week of budget in an hour) should page immediately. A slow burn can wait for business hours. Google's SRE workbook describes multi-window, multi-burn-rate alerting as the standard approach (sre.google/workbook).
- Exclude scheduled maintenance explicitly. Define maintenance windows in the contract. Otherwise planned work counts against you.
- Measure from the user's edge. Synthetic checks from inside your own network hide problems your customers see.
Choosing Your Target: A Practical Rule
When clients ask me which number to commit to, I give the same guidance:
- If a few minutes of monthly downtime does not cost real money or trust, 99.9% is honest and affordable.
- If downtime directly stops revenue or violates a regulatory requirement, invest in 99.99% and build the automation to earn it.
- Reserve 99.999% for systems where failure is catastrophic, and expect to rebuild your dependency chain to support it.
The goal is not the most nines. It is the correct number of nines, measured truthfully and defended with automation.
FAQ
What does 99.9% uptime mean in real downtime?
99.9 uptime allows roughly 8 hours 45 minutes of downtime per year, about 43 minutes per month, or around 1 minute 26 seconds per day. It is a realistic target for teams with solid monitoring and a responsive on-call rotation, without requiring multi-region redundancy.
How do I calculate allowed downtime for any SLA percentage?
Multiply your measurement window by one minus the uptime fraction. For a 30-day month (43,200 minutes) at 99.95%, allowed downtime is 43,200 * (1 - 0.9995) = 21.6 minutes. Use the same formula for weekly or annual windows by changing the total minutes.
Why does each additional nine cost so much more?
Each nine cuts allowed downtime by about 90%, which eventually removes humans from the recovery path. Below four nines you can rely on an engineer responding to a page. At four nines and above you must fund redundant infrastructure, automated failover, and a mature SRE practice, so cost rises far faster than the percentage suggests.
Should my internal SLO be the same as my external SLA?
No. Keep your internal SLO stricter than your contractual SLA. If you promise 99.9%, target 99.95% internally. The gap is your safety margin, giving you time to react before you breach the agreement and owe service credits.
How do dependencies affect my achievable uptime?
Dependencies in series multiply down. Three components at 99.9% each yield about 99.7% combined. To raise availability, add redundancy so components run in parallel, and avoid synchronous calls to any dependency with a weaker SLA than your own target.



