Asterrr's Handbook

Pricing models

On-Demand, Reserved Instances, Savings Plans, Spot and Capacity Reservations for compute, S3 storage classes and lifecycle, and how to model data transfer costs.

Exam tasks: 1.5 (design a cost optimization strategy for a multi-account environment), 2.6 (meet cost objectives: pricing models, storage tiering, data transfer)

The decision: how predictable and how interruptible is the usage, how much flexibility do you need to keep over one to three years, and which network path does each GB take?

Choosing a compute pricing model

A typical fleet mixes them: Savings Plans for the always-on baseline, On-Demand for the daily peak above it, and Spot for batch and stateless scale-out.

Side by side

On-DemandSavings PlansReserved InstancesSpot
CommitmentNone$/hour for 1 or 3 yearsInstance configuration for 1 or 3 yearsNone
DiscountNoneUp to about 72%Up to about 72%Up to about 90%
FlexibilityTotalCompute SP: any family, size, Region, OS, Fargate and LambdaStandard: size within family. Convertible: exchangeAny, while capacity lasts
Capacity guaranteeNoNoZonal RIs onlyNo, can be interrupted
CoversEverythingEC2, Fargate, Lambda, SageMaker, databases (by plan type)EC2, RDS, ElastiCache, OpenSearch, Redshift and moreEC2, Fargate Spot
2 minutes
Spot interruption notice before reclaiming an instance.
1 or 3 years
Savings Plan and RI terms. Database Savings Plans are 1 year only.
30 days
Minimum billed storage duration in S3 Standard-IA and One Zone-IA.
90 / 180 days
Minimum storage duration in Glacier Instant and Flexible Retrieval / Deep Archive.
128 KB
Minimum billable object size in the IA classes. Intelligent-Tiering doesn't monitor smaller objects.

Savings Plans vs Reserved Instances

Compute Savings PlanEC2 Instance Savings PlanStandard RIConvertible RI
Locked toNothing but the $/hourInstance family and RegionFamily, Region, OS, tenancyNothing permanent: exchange for equal or greater value
DiscountLowerHigherHighest (similar to Instance SP)Lower
Applies to Fargate and LambdaYesNoNoNo
Sell unusedNoNoYes, on the RI MarketplaceNo
  • Regional RIs apply to any AZ and, for Linux with default tenancy, any size in the family. Zonal RIs give up that flexibility in exchange for a capacity reservation in one AZ.
  • Discounts from Savings Plans and RIs are shared across the accounts of an organization's consolidated bill by default. The management account can turn sharing off for specific accounts.
  • Plan from the Cost Explorer purchase recommendations. See cost visibility.
  • For databases, Database Savings Plans cover Aurora, RDS, DynamoDB, ElastiCache, DocumentDB, Neptune, Keyspaces and others across engines and Regions, as an alternative to per-engine reserved nodes.

Exam signal

"Plans to migrate from EC2 to containers or serverless over the next year" means a Compute Savings Plan. "Steady m7i fleet in one Region, maximum discount" means an EC2 Instance Savings Plan or a Standard RI.

Commitments for spiky workloads

Buying RIs or Savings Plans to cover a peak wastes money in the hours below it. Commit to the baseline that runs 24x7, and cover the rest with On-Demand, Spot or scaling down.

Spot

  • Spot suits stateless, fault-tolerant and flexible work: batch, CI, rendering, big data task nodes, stateless web tiers behind a load balancer.
  • AWS gives a two-minute interruption notice (through instance metadata and EventBridge) and an earlier rebalance recommendation when risk rises. Checkpoint work and drain from the load balancer on either.
  • Diversify: allow many instance types and AZs, or use attribute-based instance selection (vCPU and memory ranges), so there are more capacity pools to draw from.

Allocation strategies (EC2 Fleet, Spot Fleet, Auto Scaling groups)

StrategyChoosesUse when
price-capacity-optimizedPools with the most capacity, then the lowest price among themDefault recommendation for most workloads
capacity-optimizedPools least likely to be interruptedLong jobs where an interruption is expensive
lowest-priceCheapest poolsShort jobs that tolerate frequent interruptions
diversifiedSpread across all poolsEC2 Fleet and Spot Fleet only: large fleets that must not lose everything at once

An Auto Scaling group with a mixed instances policy combines an On-Demand base (for example, 2 instances) and a percentage of On-Demand above it, with the remainder on Spot.

Spot for the only copy of state

A single Spot instance running a database or a stateful service with no replica fails the "must not lose data" requirement. Spot is for capacity that can disappear.

Capacity and tenancy

  • On-Demand Capacity Reservations hold capacity in one AZ for any duration, billed at On-Demand rates whether used or not. Pair them with a regional RI or a Savings Plan to get a discount. Use them for DR failover capacity or known events.
  • Capacity Blocks for ML reserve GPU instances for a set future period.
Dedicated HostsDedicated Instances
IsolationA whole physical server for your accountHardware not shared with other accounts
VisibilitySockets, cores and host IDNone
Placement controlHost affinity: an instance stays on the same hostNo
LicensingBYOL per socket or per core (Windows Server, SQL Server, Oracle and others), tracked with License ManagerDoesn't satisfy per-socket or per-core licenses
BillingPer hostPer instance, plus a per-Region fee

Exam signal

"Existing per-core or per-socket licenses" or "server-bound licenses" means Dedicated Hosts. "Compliance requires single-tenant hardware" with no licensing detail means Dedicated Instances is enough.

S3 storage classes

Frequent access · higher storage costRare access · lowest storage cost
  1. S3 Standard
    ms access · no minimums

    Active data, no retrieval fee. Express One Zone is the single-AZ, lowest-latency option for hot data.

  2. Intelligent-Tiering
    ms access · monitoring fee

    Moves objects between tiers by access pattern. No retrieval fees. Optional archive tiers.

  3. Standard-IA / One Zone-IA
    ms access · 30-day minimum

    Monthly or rarer access. Retrieval fee per GB. One Zone-IA only for data you can recreate.

  4. Glacier Instant Retrieval
    ms access · 90-day minimum

    Quarterly access, but needed immediately when it is.

  5. Glacier Flexible Retrieval
    minutes to hours · 90-day minimum

    Backups and archives. Expedited, standard and free bulk retrievals.

  6. Glacier Deep Archive
    12–48 hours · 180-day minimum

    Compliance archives kept for years. The cheapest storage.

  • Intelligent-Tiering is the answer when access patterns are unknown or changing. Known patterns are cheaper with lifecycle rules, such as Standard for 30 days, then Standard-IA, then Glacier after 180 days.
  • Lifecycle transitions are billed per object. Moving millions of tiny objects to Glacier can cost more than it saves; aggregate them first.
  • Expire old noncurrent versions and incomplete multipart uploads with lifecycle rules too.
  • Use S3 Storage Lens and storage class analysis to find candidates.

Data transfer cost modeling

PathCost behaviorHow to cut it
Internet into AWSFree—
AWS to internetPer GB, tieredServe through CloudFront (origin fetches from AWS are free, and cached content avoids the origin)
Same AZ, private IPFreeKeep chatty tiers in the same AZ where HA allows
Cross-AZPer GB in each directionAvoid unnecessary cross-AZ hops; use AZ-aware routing where possible
Cross-RegionPer GB out of the source RegionReplicate only what's needed; compress
NAT gatewayPer hour plus per GB processedGateway endpoints for S3 and DynamoDB (free); interface endpoints for other services
Interface endpointsPer hour per AZ plus per GB, lower than NATShare them centrally from a hub VPC
Transit GatewayPer attachment-hour plus per GB processedPeering for high-volume VPC pairs
Direct ConnectLower per-GB egress than the internetUse for large, steady hybrid traffic
Public IPv4 addressesHourly per addressRemove unused Elastic IPs; use IPv6 or private paths

The NAT gateway surprise

Private instances pulling terabytes from S3 through a NAT gateway pay NAT processing on every GB. An S3 gateway endpoint removes that charge entirely, with a route table change. See VPC endpoints.

Scenarios

Scenario
A company runs a steady fleet of 120 EC2 instances across three instance families in two Regions, 24x7. Over the next 18 months it plans to move about half of this workload to ECS on Fargate and some to Lambda. It wants to reduce compute cost without losing the discount as it migrates. What should it buy?
Scenario · choose 2
A genomics company runs nightly batch jobs on EC2 that take 1 to 3 hours each and can resume from checkpoints. It wants the lowest cost and as few interruptions as possible. Private instances download 40 TB of input data from S3 each month through a NAT gateway. Which TWO actions reduce cost the MOST?
Scenario
A company must move its on-premises Oracle Database workloads to EC2 and bring its existing per-core Oracle licenses. The licensing agreement requires visibility of physical cores. Which option should the architect choose?

Further reading

On this page