Pricing models
On-Demand, Reserved Instances, Savings Plans, Spot and Capacity Reservations for compute, S3 storage classes and lifecycle, and how to model data transfer costs.
Exam tasks: 1.5 (design a cost optimization strategy for a multi-account environment), 2.6 (meet cost objectives: pricing models, storage tiering, data transfer)
The decision: how predictable and how interruptible is the usage, how much flexibility do you need to keep over one to three years, and which network path does each GB take?
Choosing a compute pricing model
A typical fleet mixes them: Savings Plans for the always-on baseline, On-Demand for the daily peak above it, and Spot for batch and stateless scale-out.
Side by side
| On-Demand | Savings Plans | Reserved Instances | Spot | |
|---|---|---|---|---|
| Commitment | None | $/hour for 1 or 3 years | Instance configuration for 1 or 3 years | None |
| Discount | None | Up to about 72% | Up to about 72% | Up to about 90% |
| Flexibility | Total | Compute SP: any family, size, Region, OS, Fargate and Lambda | Standard: size within family. Convertible: exchange | Any, while capacity lasts |
| Capacity guarantee | No | No | Zonal RIs only | No, can be interrupted |
| Covers | Everything | EC2, Fargate, Lambda, SageMaker, databases (by plan type) | EC2, RDS, ElastiCache, OpenSearch, Redshift and more | EC2, Fargate Spot |
Savings Plans vs Reserved Instances
| Compute Savings Plan | EC2 Instance Savings Plan | Standard RI | Convertible RI | |
|---|---|---|---|---|
| Locked to | Nothing but the $/hour | Instance family and Region | Family, Region, OS, tenancy | Nothing permanent: exchange for equal or greater value |
| Discount | Lower | Higher | Highest (similar to Instance SP) | Lower |
| Applies to Fargate and Lambda | Yes | No | No | No |
| Sell unused | No | No | Yes, on the RI Marketplace | No |
- Regional RIs apply to any AZ and, for Linux with default tenancy, any size in the family. Zonal RIs give up that flexibility in exchange for a capacity reservation in one AZ.
- Discounts from Savings Plans and RIs are shared across the accounts of an organization's consolidated bill by default. The management account can turn sharing off for specific accounts.
- Plan from the Cost Explorer purchase recommendations. See cost visibility.
- For databases, Database Savings Plans cover Aurora, RDS, DynamoDB, ElastiCache, DocumentDB, Neptune, Keyspaces and others across engines and Regions, as an alternative to per-engine reserved nodes.
Exam signal
"Plans to migrate from EC2 to containers or serverless over the next year" means a Compute Savings Plan. "Steady m7i fleet in one Region, maximum discount" means an EC2 Instance Savings Plan or a Standard RI.
Commitments for spiky workloads
Buying RIs or Savings Plans to cover a peak wastes money in the hours below it. Commit to the baseline that runs 24x7, and cover the rest with On-Demand, Spot or scaling down.
Spot
- Spot suits stateless, fault-tolerant and flexible work: batch, CI, rendering, big data task nodes, stateless web tiers behind a load balancer.
- AWS gives a two-minute interruption notice (through instance metadata and EventBridge) and an earlier rebalance recommendation when risk rises. Checkpoint work and drain from the load balancer on either.
- Diversify: allow many instance types and AZs, or use attribute-based instance selection (vCPU and memory ranges), so there are more capacity pools to draw from.
Allocation strategies (EC2 Fleet, Spot Fleet, Auto Scaling groups)
| Strategy | Chooses | Use when |
|---|---|---|
| price-capacity-optimized | Pools with the most capacity, then the lowest price among them | Default recommendation for most workloads |
| capacity-optimized | Pools least likely to be interrupted | Long jobs where an interruption is expensive |
| lowest-price | Cheapest pools | Short jobs that tolerate frequent interruptions |
| diversified | Spread across all pools | EC2 Fleet and Spot Fleet only: large fleets that must not lose everything at once |
An Auto Scaling group with a mixed instances policy combines an On-Demand base (for example, 2 instances) and a percentage of On-Demand above it, with the remainder on Spot.
Spot for the only copy of state
A single Spot instance running a database or a stateful service with no replica fails the "must not lose data" requirement. Spot is for capacity that can disappear.
Capacity and tenancy
- On-Demand Capacity Reservations hold capacity in one AZ for any duration, billed at On-Demand rates whether used or not. Pair them with a regional RI or a Savings Plan to get a discount. Use them for DR failover capacity or known events.
- Capacity Blocks for ML reserve GPU instances for a set future period.
| Dedicated Hosts | Dedicated Instances | |
|---|---|---|
| Isolation | A whole physical server for your account | Hardware not shared with other accounts |
| Visibility | Sockets, cores and host ID | None |
| Placement control | Host affinity: an instance stays on the same host | No |
| Licensing | BYOL per socket or per core (Windows Server, SQL Server, Oracle and others), tracked with License Manager | Doesn't satisfy per-socket or per-core licenses |
| Billing | Per host | Per instance, plus a per-Region fee |
Exam signal
"Existing per-core or per-socket licenses" or "server-bound licenses" means Dedicated Hosts. "Compliance requires single-tenant hardware" with no licensing detail means Dedicated Instances is enough.
S3 storage classes
- S3 Standardms access · no minimums
Active data, no retrieval fee. Express One Zone is the single-AZ, lowest-latency option for hot data.
- Intelligent-Tieringms access · monitoring fee
Moves objects between tiers by access pattern. No retrieval fees. Optional archive tiers.
- Standard-IA / One Zone-IAms access · 30-day minimum
Monthly or rarer access. Retrieval fee per GB. One Zone-IA only for data you can recreate.
- Glacier Instant Retrievalms access · 90-day minimum
Quarterly access, but needed immediately when it is.
- Glacier Flexible Retrievalminutes to hours · 90-day minimum
Backups and archives. Expedited, standard and free bulk retrievals.
- Glacier Deep Archive12–48 hours · 180-day minimum
Compliance archives kept for years. The cheapest storage.
- Intelligent-Tiering is the answer when access patterns are unknown or changing. Known patterns are cheaper with lifecycle rules, such as Standard for 30 days, then Standard-IA, then Glacier after 180 days.
- Lifecycle transitions are billed per object. Moving millions of tiny objects to Glacier can cost more than it saves; aggregate them first.
- Expire old noncurrent versions and incomplete multipart uploads with lifecycle rules too.
- Use S3 Storage Lens and storage class analysis to find candidates.
Data transfer cost modeling
| Path | Cost behavior | How to cut it |
|---|---|---|
| Internet into AWS | Free | — |
| AWS to internet | Per GB, tiered | Serve through CloudFront (origin fetches from AWS are free, and cached content avoids the origin) |
| Same AZ, private IP | Free | Keep chatty tiers in the same AZ where HA allows |
| Cross-AZ | Per GB in each direction | Avoid unnecessary cross-AZ hops; use AZ-aware routing where possible |
| Cross-Region | Per GB out of the source Region | Replicate only what's needed; compress |
| NAT gateway | Per hour plus per GB processed | Gateway endpoints for S3 and DynamoDB (free); interface endpoints for other services |
| Interface endpoints | Per hour per AZ plus per GB, lower than NAT | Share them centrally from a hub VPC |
| Transit Gateway | Per attachment-hour plus per GB processed | Peering for high-volume VPC pairs |
| Direct Connect | Lower per-GB egress than the internet | Use for large, steady hybrid traffic |
| Public IPv4 addresses | Hourly per address | Remove unused Elastic IPs; use IPv6 or private paths |
The NAT gateway surprise
Private instances pulling terabytes from S3 through a NAT gateway pay NAT processing on every GB. An S3 gateway endpoint removes that charge entirely, with a route table change. See VPC endpoints.
Scenarios
A Compute Savings Plan applies across instance families, Regions, Fargate and Lambda, so the discount follows the workload as it migrates. Standard RIs and EC2 Instance Savings Plans are tied to families and Regions and would go unused. Spot can be interrupted, which is wrong for a steady 24x7 production fleet.
Checkpointed batch work suits Spot, and diversifying across types with price-capacity-optimized balances cost and interruption risk. An S3 gateway endpoint removes NAT processing charges for 40 TB a month. RIs don't fit intermittent nightly jobs. Public IPs add IPv4 charges and weaken security. Lowest-price on one instance type concentrates on pools most likely to be interrupted.
Dedicated Hosts expose sockets and cores and keep instances on known hardware, which BYOL per-core licensing requires. Dedicated Instances isolate hardware but don't expose cores. Default tenancy doesn't meet the licensing terms at all. License-included RDS means buying new licenses instead of using the existing ones.
Further reading
End-user computing and contact centers
Choosing between WorkSpaces, WorkSpaces Applications (formerly AppStream 2.0) and WorkSpaces Secure Browser, and building a contact center on Amazon Connect with Lex, Polly, Transcribe and Contact Lens.
Domain 3 · Continuous improvement
25% of the exam. Improving a system that already runs, with the smallest change that fixes the problem.