Asterrr's Handbook

Cost optimization

Finding and removing waste in an existing AWS environment, from idle resources and NAT gateway charges to rightsizing before you commit.

Exam tasks: 3.5 (identify opportunities for cost optimization in an existing solution)

The decision: the bill is higher than it should be. Where is the waste (idle, oversized, wrongly priced or wrongly routed), and in what order do you fix it so the savings stick?

This page is about fixing an existing system. For the tools that show where money goes (tagging, Cost Explorer, CUR, Budgets as reporting) see cost visibility. For how On-Demand, Savings Plans, Reserved Instances and Spot are priced, see pricing models.

The order matters

Rightsize before you commit

Buying Savings Plans or Reserved Instances for an oversized fleet locks in waste for one or three years. When a question lists both rightsizing and a commitment, the right order is rightsizing first, then committing to the new, smaller baseline.

Where the waste usually is

WasteHow to find itFix
Unattached EBS volumesTrusted Advisor, Compute Optimizer idle recommendations, describe-volumes with status availableSnapshot if needed, then delete
Old snapshots and AMIsCost Explorer by usage type, DLM and AWS Backup retention reportsRetention rules, archive tier for rarely used snapshots
Idle EC2 and RDS instancesCompute Optimizer idle findings, Trusted Advisor low-utilization checksStop, schedule or delete. Stopped RDS restarts after 7 days
Idle load balancers and NAT gatewaysTrusted Advisor, zero RequestCount or processed bytesDelete
Public IPv4 addressesCost Explorer public IPv4 usage, VPC IP Address ManagerRelease unused Elastic IPs, move internal traffic to private addresses
Oversized instancesCompute Optimizer, Cost Optimization HubDownsize or change family
gp2 volumesCompute Optimizer EBS findingsMigrate to gp3 in place
S3 growthS3 Storage Lens: incomplete multipart uploads, noncurrent versions, buckets without lifecycle rulesLifecycle rules, Intelligent-Tiering, abort incomplete uploads
Dev and test running at nightTags plus utilization by hourInstance Scheduler or scheduled scaling to zero
CloudWatch Logs kept foreverLog groups with no retention settingSet retention, archive to S3
  • Cost Optimization Hub brings rightsizing, idle, Graviton and commitment recommendations from across the organization into one list and estimates savings net of your existing discounts.
  • Trusted Advisor cost checks are a quick first pass. The full set needs Business Support or higher.

Stopping is not deleting

A stopped EC2 instance still pays for its EBS volumes and any Elastic IP address. "Stop the unused instances" saves compute, but storage keeps billing until you snapshot and delete.

Guardrails on expected usage

ToolDoesGood for
CloudWatch billing alarmAlarms on the EstimatedCharges metric for the account (us-east-1 only, after turning on billing alerts)A simple "total bill passed X" alert
AWS BudgetsCost, usage, Savings Plans and RI utilization or coverage budgets, with actual and forecasted alertsPer team, per tag, per service limits
Budget actionsApply an IAM policy or SCP, or stop specific EC2 or RDS instances, when a budget is exceededSandbox accounts that must not overspend
Cost Anomaly DetectionLearns normal spend and alerts on unusual increases by service, account or tagCatching a runaway resource without setting thresholds

Exam signal

"Alert when spend is forecast to exceed the monthly amount" means AWS Budgets. A CloudWatch billing alarm only sees estimated charges so far. "Alert on unexpected spikes without setting thresholds" means Cost Anomaly Detection.

Digging into the CUR with Athena

Cost Explorer answers most questions. When you need line-item detail, such as which resource ID or which data transfer type drove a charge, query the Cost and Usage Report:

Create a CUR 2.0 export with Data Exports, delivered to S3 in Parquet with resource IDs included.
Catalog it with a Glue crawler or table so Athena can query it.
Query by line_item_usage_type, line_item_resource_id and cost allocation tags. For example, find the top 20 resources by NatGateway-Bytes or DataTransfer-Regional-Bytes.
Visualize in QuickSight, or use the Cloud Intelligence Dashboards built on the same data.

NAT gateway data processing

Every GB that goes through a NAT gateway has a processing charge, on top of the hourly charge and any data transfer. Workloads that pull large amounts from S3, DynamoDB or ECR through NAT pay it on every byte.

  • Gateway endpoints for S3 and DynamoDB have no hourly or data charge. Adding one is a route table change.
  • Interface endpoints charge per hour per AZ plus per GB, which is cheaper than NAT for high-volume traffic to one service such as ECR, but can cost more for low-volume services spread over many AZs.
  • Keep traffic in one AZ where you can. A NAT gateway in another AZ adds cross-AZ data transfer.
  • See VPC endpoints for the endpoint details.

Endpoints across Regions

A gateway endpoint only reaches S3 buckets in the same Region. Traffic to buckets in other Regions still goes through the NAT gateway.

Spot for tiers that tolerate interruption

  • Good fits: stateless web tiers behind a load balancer, batch and queue workers, CI runners, big data clusters, containers on ECS or EKS.
  • Poor fits: single-instance databases, stateful apps without checkpointing, anything that can't handle a two-minute interruption notice.
  • Make it safe: an Auto Scaling group with mixed instances (On-Demand base plus Spot), many instance types, the price-capacity-optimized allocation strategy, and capacity rebalancing.

Scenarios

Scenario
A company's monthly bill has grown 40% in six months. An investigation shows 300 unattached EBS volumes, most EC2 instances averaging under 10% CPU, and no Savings Plans. The CFO wants the largest lasting savings. In which order should the architect act?
Scenario
A data processing application in private subnets reads about 50 TB a month from S3 buckets in the same Region, and its Cost Explorer shows NAT gateway processing as the largest networking charge. What is the MOST cost-effective change?
Scenario · choose 2
A startup gives each developer a sandbox account. Last month one developer left a large GPU instance running for weeks. Leadership wants an alert when a sandbox is forecast to go over its monthly limit, and wants compute stopped automatically if the limit is actually exceeded. Which TWO actions meet this?

Further reading

On this page