Cost optimization
Finding and removing waste in an existing AWS environment, from idle resources and NAT gateway charges to rightsizing before you commit.
Exam tasks: 3.5 (identify opportunities for cost optimization in an existing solution)
The decision: the bill is higher than it should be. Where is the waste (idle, oversized, wrongly priced or wrongly routed), and in what order do you fix it so the savings stick?
This page is about fixing an existing system. For the tools that show where money goes (tagging, Cost Explorer, CUR, Budgets as reporting) see cost visibility. For how On-Demand, Savings Plans, Reserved Instances and Spot are priced, see pricing models.
The order matters
Rightsize before you commit
Buying Savings Plans or Reserved Instances for an oversized fleet locks in waste for one or three years. When a question lists both rightsizing and a commitment, the right order is rightsizing first, then committing to the new, smaller baseline.
Where the waste usually is
| Waste | How to find it | Fix |
|---|---|---|
| Unattached EBS volumes | Trusted Advisor, Compute Optimizer idle recommendations, describe-volumes with status available | Snapshot if needed, then delete |
| Old snapshots and AMIs | Cost Explorer by usage type, DLM and AWS Backup retention reports | Retention rules, archive tier for rarely used snapshots |
| Idle EC2 and RDS instances | Compute Optimizer idle findings, Trusted Advisor low-utilization checks | Stop, schedule or delete. Stopped RDS restarts after 7 days |
| Idle load balancers and NAT gateways | Trusted Advisor, zero RequestCount or processed bytes | Delete |
| Public IPv4 addresses | Cost Explorer public IPv4 usage, VPC IP Address Manager | Release unused Elastic IPs, move internal traffic to private addresses |
| Oversized instances | Compute Optimizer, Cost Optimization Hub | Downsize or change family |
| gp2 volumes | Compute Optimizer EBS findings | Migrate to gp3 in place |
| S3 growth | S3 Storage Lens: incomplete multipart uploads, noncurrent versions, buckets without lifecycle rules | Lifecycle rules, Intelligent-Tiering, abort incomplete uploads |
| Dev and test running at night | Tags plus utilization by hour | Instance Scheduler or scheduled scaling to zero |
| CloudWatch Logs kept forever | Log groups with no retention setting | Set retention, archive to S3 |
- Cost Optimization Hub brings rightsizing, idle, Graviton and commitment recommendations from across the organization into one list and estimates savings net of your existing discounts.
- Trusted Advisor cost checks are a quick first pass. The full set needs Business Support or higher.
Stopping is not deleting
A stopped EC2 instance still pays for its EBS volumes and any Elastic IP address. "Stop the unused instances" saves compute, but storage keeps billing until you snapshot and delete.
Guardrails on expected usage
| Tool | Does | Good for |
|---|---|---|
| CloudWatch billing alarm | Alarms on the EstimatedCharges metric for the account (us-east-1 only, after turning on billing alerts) | A simple "total bill passed X" alert |
| AWS Budgets | Cost, usage, Savings Plans and RI utilization or coverage budgets, with actual and forecasted alerts | Per team, per tag, per service limits |
| Budget actions | Apply an IAM policy or SCP, or stop specific EC2 or RDS instances, when a budget is exceeded | Sandbox accounts that must not overspend |
| Cost Anomaly Detection | Learns normal spend and alerts on unusual increases by service, account or tag | Catching a runaway resource without setting thresholds |
Exam signal
"Alert when spend is forecast to exceed the monthly amount" means AWS Budgets. A CloudWatch billing alarm only sees estimated charges so far. "Alert on unexpected spikes without setting thresholds" means Cost Anomaly Detection.
Digging into the CUR with Athena
Cost Explorer answers most questions. When you need line-item detail, such as which resource ID or which data transfer type drove a charge, query the Cost and Usage Report:
line_item_usage_type, line_item_resource_id and cost allocation tags. For example, find the top 20 resources by NatGateway-Bytes or DataTransfer-Regional-Bytes.NAT gateway data processing
Every GB that goes through a NAT gateway has a processing charge, on top of the hourly charge and any data transfer. Workloads that pull large amounts from S3, DynamoDB or ECR through NAT pay it on every byte.
- Gateway endpoints for S3 and DynamoDB have no hourly or data charge. Adding one is a route table change.
- Interface endpoints charge per hour per AZ plus per GB, which is cheaper than NAT for high-volume traffic to one service such as ECR, but can cost more for low-volume services spread over many AZs.
- Keep traffic in one AZ where you can. A NAT gateway in another AZ adds cross-AZ data transfer.
- See VPC endpoints for the endpoint details.
Endpoints across Regions
A gateway endpoint only reaches S3 buckets in the same Region. Traffic to buckets in other Regions still goes through the NAT gateway.
Spot for tiers that tolerate interruption
- Good fits: stateless web tiers behind a load balancer, batch and queue workers, CI runners, big data clusters, containers on ECS or EKS.
- Poor fits: single-instance databases, stateful apps without checkpointing, anything that can't handle a two-minute interruption notice.
- Make it safe: an Auto Scaling group with mixed instances (On-Demand base plus Spot), many instance types, the price-capacity-optimized allocation strategy, and capacity rebalancing.
Scenarios
Remove pure waste first, then shrink what's left, then commit to the smaller baseline. Committing first locks in spend for oversized instances. Leaving the volumes misses easy savings, and moving everything to Spot puts workloads that can't be interrupted at risk.
A gateway endpoint sends S3 traffic straight from the VPC with no processing or hourly charge, and it only needs a route table change. A NAT instance adds operational work and still carries the traffic, public subnets add IPv4 charges and exposure, and Transfer Acceleration costs extra and is for long-distance uploads.
Budgets supports forecast alerts, and budget actions can stop instances or apply a restrictive SCP when the budget is exceeded. Billing metrics exist only in us-east-1 and only show charges so far. Cost Anomaly Detection alerts but doesn't act, and weekly Trusted Advisor checks are too slow and don't enforce anything.