Skip to content
Techsense Developers
TrustLet's Talk
Insights
Cloud & Infrastructure7 min readOct 2, 2026

How to Cut Cloud Costs on AWS Without Slowing Down Engineering

To cut cloud costs on AWS without slowing your teams down, focus on removing waste before you touch anything developers depend on: delete idle resources, right-size overprovisioned compute, buy…

To cut cloud costs on AWS without slowing your teams down, focus on removing waste before you touch anything developers depend on: delete idle resources, right-size overprovisioned compute, buy commitment-based discounts for stable baselines, and attach cost signals directly to the tooling engineers already use. The mistake most organizations make is treating cost as a quarterly cleanup project. The durable approach is to make spend visible at the moment decisions are made, so the cheaper choice is also the easier one. Done well, AWS cost optimization reduces your bill and your operational risk at the same time.

This post walks through the specific moves I use with engineering teams, in the order that delivers the most savings for the least disruption.

Start with Visibility, Not Cuts

You cannot reduce cloud spend you cannot see. Before changing a single instance type, get accurate, attributable cost data.

The foundation is tagging. Without consistent tags, your AWS bill is one undifferentiated number and every cost conversation turns into archaeology. Enforce a small, mandatory tag set:

  • team or owner
  • environment (prod, staging, dev)
  • service or application
  • cost-center

Enforce these with tag policies in AWS Organizations and block untagged resource creation where you can. Then enable AWS Cost and Usage Report (CUR) into Athena so you can query spend precisely:

SELECT
  resource_tags_user_team AS team,
  line_item_product_code AS service,
  SUM(line_item_unblended_cost) AS cost
FROM cur_table
WHERE line_item_usage_start_date >= date_add('day', -30, current_date)
GROUP BY 1, 2
ORDER BY cost DESC
LIMIT 25;

That single query tells you which teams and services drive your bill. Everything downstream depends on it. This is the core of a FinOps practice: cost data that engineers, finance, and leadership can all read the same way.

Kill Idle and Orphaned Resources First

The cheapest savings require no engineering trade-offs because nothing is using the resource. This is pure waste removal.

Look for:

  • Unattached EBS volumes left behind after instances terminate
  • Old EBS snapshots nobody tracks
  • Idle load balancers with no healthy targets
  • Unassociated Elastic IPs (AWS charges for these)
  • Stopped instances still holding provisioned storage
  • Dev and staging environments running 24/7 when nobody works nights or weekends

Finding unattached volumes takes one command:

aws ec2 describe-volumes \
  --filters Name=status,Values=available \
  --query 'Volumes[].{ID:VolumeId,Size:Size,Created:CreateTime}' \
  --output table

For non-production environments, scheduled shutdown is one of the highest-return changes available. A dev fleet that runs only during working hours, roughly 50 hours a week instead of 168, cuts that compute bill by around 70 percent. AWS Instance Scheduler or a simple EventBridge rule handles this:

# Lambda triggered by EventBridge cron (stop at 8pm weekdays)
import boto3

def handler(event, context):
    ec2 = boto3.client("ec2")
    instances = ec2.describe_instances(
        Filters=[
            {"Name": "tag:environment", "Values": ["dev", "staging"]},
            {"Name": "instance-state-name", "Values": ["running"]},
        ]
    )
    ids = [
        i["InstanceId"]
        for r in instances["Reservations"]
        for i in r["Instances"]
    ]
    if ids:
        ec2.stop_instances(InstanceIds=ids)

None of this affects production reliability or developer velocity. It removes spend that produced zero value.

Right-Size Compute and Storage

Once waste is gone, address overprovisioning. Teams routinely pick instance sizes defensively and never revisit them.

Use AWS Compute Optimizer to get data-driven recommendations based on actual CPU, memory, and network utilization. Prioritize:

  1. EC2 instances consistently below 40 percent CPU and memory utilization
  2. RDS instances sized for peak loads that rarely occur
  3. Over-provisioned EBS volumes using gp2 that could move to gp3 (gp3 is cheaper and lets you provision IOPS and throughput independently)
  4. Lambda functions with memory settings that do not match their real profile

The gp2-to-gp3 migration is a reliable, low-risk win. gp3 delivers a baseline 3,000 IOPS and 125 MB/s regardless of volume size, and you pay separately only if you need more. For most general-purpose workloads this is a straight price reduction with no performance loss.

Right-sizing does carry a small risk of degrading performance if you cut too aggressively. Protect velocity by making changes in staging first, watching CloudWatch metrics, and moving in one size step at a time rather than jumping two tiers.

Buy Commitments for Your Stable Baseline

After right-sizing, you have a clearer picture of your steady-state usage. That baseline should never run at on-demand rates.

Two mechanisms matter:

  • Savings Plans: commit to a consistent dollar-per-hour spend for one or three years in exchange for lower rates. Compute Savings Plans are the most flexible, applying across EC2, Fargate, and Lambda regardless of instance family or region.
  • Reserved Instances: still useful for RDS, ElastiCache, Redshift, and OpenSearch where Savings Plans do not apply.

The rule I follow: commit only to what you are confident will run for the full term. Cover roughly 70 to 80 percent of your stable baseline and leave the variable top layer on-demand or on Spot. Over-committing to save a few percentage points can lock you into capacity you no longer need, which erases the savings.

For fault-tolerant and interruptible work, Spot Instances offer the deepest discounts. They fit batch processing, CI pipelines, rendering, and stateless workloads behind queues. The trade-off is that AWS can reclaim them with two minutes of notice, so only use Spot where your architecture handles interruption gracefully.

Keep Engineering Fast by Shifting Cost Left

The whole point of this post is that cost control must not become a tax on velocity. The way to achieve that is to put cost information where engineers already work, rather than in a finance report they never open.

Practical steps that preserve speed:

  • Show cost in pull requests. Tools like Infracost estimate the cost delta of an infrastructure-as-code change before it merges, so the trade-off is visible at review time.
  • Set budgets with alerts, not hard blocks. AWS Budgets notifying a team's Slack channel beats a hard cap that pages someone at 2 a.m.
  • Use anomaly detection. AWS Cost Anomaly Detection catches a runaway spend spike within a day, long before the monthly invoice.
  • Make the owning team see its own number. When each team's tagged spend is visible to that team, optimization becomes self-directed instead of top-down.

Here is a minimal Infracost step in a CI pipeline:

- name: Infracost diff
  run: |
    infracost breakdown --path=. --format=json --out-file=base.json
    infracost diff --path=. --compare-to=base.json

When engineers see that a change adds $400 a month before it ships, most of the expensive decisions self-correct without any mandate from above. This is the cultural core of FinOps: cost becomes a non-functional requirement, like latency or security, owned by the people who can actually change it.

Sequence the Work for Maximum Return

Order matters. To cut cloud costs efficiently, work in this sequence:

  1. Visibility: tagging, CUR, dashboards
  2. Waste removal: idle and orphaned resources, non-prod scheduling
  3. Right-sizing: compute, storage, databases
  4. Commitments: Savings Plans, Reserved Instances, Spot
  5. Prevention: cost in CI, budgets, anomaly detection

The first two steps usually deliver meaningful savings in days and require no architectural change. The later steps compound over time.

If you want help designing the guardrails so optimization sticks, our cloud and infrastructure capabilities cover this end to end, and we tailor the approach to the constraints of your sector across the industries we work with.

FAQ

How much can I realistically expect to cut cloud costs on AWS?

It depends on your starting maturity. Teams that have never done structured optimization often find substantial savings from waste removal and right-sizing alone, before any commitments. Rather than chase a headline percentage, measure your own baseline first, then track savings against tagged spend month over month. Concrete numbers come from your CUR, not from benchmarks.

Will cost optimization hurt performance or slow my engineers down?

It should not, if you sequence it correctly. Waste removal and non-production scheduling carry no production risk. Right-sizing is the only step with trade-off potential, so you test in staging and move one size step at a time. Shifting cost signals into pull requests and Slack keeps engineers informed without adding gates that block their work.

What is FinOps, and do I need a dedicated team?

FinOps is the practice of bringing financial accountability to variable cloud spend by making cost a shared, data-driven responsibility across engineering, finance, and leadership. You do not need a large dedicated team to start. A clear tagging standard, a shared cost dashboard, and one owner who drives the cadence are enough to begin. The structure scales with your spend.

Should I use Savings Plans or Reserved Instances?

Use Compute Savings Plans for EC2, Fargate, and Lambda because they are the most flexible and apply across instance families and regions. Use Reserved Instances for services Savings Plans do not cover, such as RDS, ElastiCache, Redshift, and OpenSearch. In both cases, commit only to the stable baseline you are confident will run for the full term.

How do I keep cloud costs from creeping back up?

Make cost a standing part of delivery rather than a cleanup event. Keep tagging enforced, surface cost deltas in pull requests, route budget alerts to the owning team, and enable anomaly detection. When the people who create spend can see it and own it, savings hold instead of eroding.