Back to all articles
Cloud9 min read

We cut a 41,000 dollar AWS bill to 16,400 without touching the product

A line by line account of a six week cloud cost engagement: what we found, what we changed, what we deliberately left alone, and what it cost to do the work.

Written by

Sujata Rai, Head of Platform Engineering

Published

May 27, 2026

A logistics software company came to us with a monthly AWS bill of 41,200 dollars and a board asking why it had tripled in eighteen months while revenue grew forty percent. Their engineering team was six people. Nobody owned the bill, which is the normal situation and the actual root cause.

We ran a six week engagement. The bill for the month after we finished was 16,400 dollars. Nothing in the product changed. No feature was removed, no latency budget was relaxed, no team was told to stop deploying.

Here is the whole thing, including the parts that did not work.

Week one was only measurement

We changed nothing in the first week. That is deliberate and it is the step most teams skip, because the temptation to go and delete something is strong.

Cost Explorer gives you spend by service. That is close to useless for decision making, because it tells you EC2 costs eighteen thousand dollars and not which team, environment or feature caused it. What you need is spend by workload, and to get that you need tags.

Their tagging was about thirty percent complete. We spent four days getting it to ninety four percent, using a mix of Terraform defaults, an AWS Config rule to flag untagged resources and a spreadsheet for the twenty or so orphans nobody recognised. Three of those orphans turned out to be resources from a proof of concept run in 2023 by an engineer who had since left.

The picture at the end of week one:

Category Monthly Share
Production compute 12,900 31%
Non production environments 9,600 23%
Data transfer 6,800 17%
Storage, mostly S3 and EBS snapshots 5,900 14%
Managed databases 4,100 10%
Everything else 1,900 5%

Twenty three percent of the bill was environments nobody used outside working hours. That is not unusual. It is the single most common finding in every cost audit we have run.

Non production was the easy 8,000 dollars

Four environments: development, staging, QA and a demo environment for sales. All four ran the full production topology twenty four hours a day, including three node Kubernetes clusters and multi-availability-zone RDS instances.

What we changed:

Scheduled shutdown. Development, QA and demo now run 07:00 to 21:00 Nepal time on weekdays. That is 70 hours a week instead of 168, a fifty eight percent reduction on those three environments. A Lambda on an EventBridge schedule scales the node groups to zero and stops the RDS instances. Sales got a Slack command to wake the demo environment on request, which takes about four minutes and has been used eleven times in five months.

Single availability zone below production. Multi-AZ RDS in a QA environment protects you against a datacentre failure taking down a system nobody is using. We moved all three to single AZ.

Right sized instances. Staging was running the same m6i.2xlarge nodes as production. Staging peaks at roughly eight percent of production traffic. It runs on m6i.large now, and the only complaint has been slightly slower test suite runs.

Non production went from 9,600 to 1,700 dollars a month. This took nine working days and carried essentially no risk, which is why we always start here. It also buys credibility for the changes that come next.

Data transfer was the interesting one

Data transfer at 6,800 dollars a month for a company of their size was the number that did not make sense. Cross-AZ traffic accounted for most of it.

Their Kubernetes cluster spread pods across three availability zones for resilience, which is correct. But their service mesh was routing requests without any zone preference, so roughly two thirds of every internal call crossed an availability zone boundary at one cent per gigabyte in each direction. A chatty microservice architecture made that a lot of gigabytes.

Enabling topology aware routing so that traffic prefers same zone endpoints and only crosses zones when the local one is unhealthy cut cross-AZ transfer by seventy one percent. That is a configuration change measured in lines, and it saved 3,100 dollars a month.

The second piece was NAT gateway charges. Every pod pulling container images and calling the S3 API was going out through a NAT gateway at 4.5 cents per gigabyte. VPC endpoints for S3 and ECR are free for gateway endpoints and dramatically cheaper for interface endpoints. That removed another 1,400 dollars.

Data transfer went from 6,800 to 1,900 dollars.

Storage was mostly forgotten snapshots

5,900 dollars of storage broke down as 3,400 in EBS snapshots, 1,800 in S3 and 700 in EFS.

The snapshots were the story. An automated daily snapshot policy created in 2022 had no expiry. There were 2,847 snapshots, the oldest from March 2022. Nobody had ever restored from one older than a fortnight. We set a lifecycle policy of thirty daily, twelve monthly and three yearly, then deleted the rest after confirming with their CTO in writing. That is 2,600 dollars a month for a policy change.

S3 was more nuanced. Their application logs went to standard storage and stayed there forever. We applied a lifecycle rule moving objects to Infrequent Access at 30 days and Glacier Instant Retrieval at 90, which cut S3 by about forty percent without changing retrieval behaviour for anything recent.

We left EFS alone. It was 700 dollars, it was correctly provisioned, and there was better work to do.

Compute was where we had to be careful

Production compute at 12,900 dollars is where the money is, and it is also where a bad change causes an outage. We were slower here.

Right sizing based on real data, not on a hunch. Two weeks of CloudWatch and Prometheus data showed the API tier running at eleven percent average CPU with peaks around thirty four. It was provisioned for a load pattern that had changed a year earlier. We reduced node count and moved from m6i to m7g Graviton instances, which are roughly twenty percent cheaper per unit of work for their workload. Their stack is Node and Go, so the ARM migration was a base image change and a CI pipeline update. It took four days including testing.

Savings plans on the stable floor. After right sizing we had a clear picture of baseline usage. A one year compute savings plan covering seventy percent of the new baseline, deliberately not the full amount, cut that portion by about twenty seven percent. We do not recommend covering more than eighty percent of baseline. Commitments you cannot use are just a slower way of overspending.

Spot for the batch tier. Their nightly reconciliation jobs are interruption tolerant and were running on on-demand instances. Moving them to Spot with an on-demand fallback saved sixty eight percent on that workload.

Production compute went from 12,900 to 6,400, and p99 API latency improved slightly, which we did not expect and attribute to the Graviton move.

What we chose not to do

Two things worth naming, because knowing when to stop is most of this discipline.

We did not move them off Kubernetes. A managed container service would probably have been cheaper, perhaps a thousand dollars a month. It would also have been a three month migration with real risk, and their team knows Kubernetes well. The payback period made it a bad trade.

We did not touch their RDS instance class. It was overprovisioned by maybe thirty percent, worth roughly 900 dollars a month. But their query patterns were about to change substantially due to a feature in progress, and right sizing against a load profile that is about to change is guessing. We wrote it up as a recommendation to revisit in the following quarter, which they did.

The final numbers

Category Before After Change
Production compute 12,900 6,400 -50%
Non production 9,600 1,700 -82%
Data transfer 6,800 1,900 -72%
Storage 5,900 2,300 -61%
Managed databases 4,100 3,200 -22%
Everything else 1,900 900 -53%
Total 41,200 16,400 -60%

The engagement cost them about six weeks of a two person team. It paid for itself in the seventh week and has kept paying since, which is the part that matters. Cost work is not a one time cleanup, it is a control loop.

What we left behind so it does not creep back

Cost audits that end with a spreadsheet get undone within a year. The last week of the engagement was entirely about making the improvement durable:

  • A cost anomaly alert into their engineering Slack channel, triggered on a fifteen percent week over week rise in any tagged workload
  • A monthly per team cost report, generated automatically and sent to team leads rather than to finance
  • A required tag policy enforced in CI, so an untagged resource fails the pull request rather than being discovered a year later
  • A short written runbook for the next review, so they can run this themselves

Five months on their bill is 17,100 dollars, against roughly thirty percent more traffic. The drift is normal and visible, which is the entire point.

If you are about to do this yourself

Three things, in order.

Start with measurement and resist changing anything for a week. Every hour spent on tagging pays back several times over, because it turns the bill from one number into a list of decisions.

Do the non production work first. It is the largest easy win in almost every account, it carries no production risk, and it earns you the room to make more careful changes later.

Write down what you deliberately did not do and why. The recommendation we left on their RDS instance was worth more than some of the changes we made, because it was ready when their load profile settled.

AWSFinOpsCloudKubernetes

Working on something similar?

If this article is close to a problem on your desk, we are glad to talk it through. No pitch, just the conversation.

  • A senior engineer reads every brief
  • NDA signed before you share anything sensitive
  • No sales sequence, no automated follow ups