Asterrr's Handbook

Performance and rightsizing

Finding the bottleneck in an existing workload and fixing it with the right instance family, storage, networking, placement and scaling policy.

Exam tasks: 3.3 (determine a strategy to improve performance)

The decision: which resource is actually saturated (CPU, memory, disk, network or a downstream dependency), and what's the smallest change that removes that limit?

Find the bottleneck first

  • Metrics to check: CPUUtilization, memory and swap (CloudWatch agent), VolumeQueueLength and EBS throughput, NetworkIn / NetworkOut and ENA allowance-exceeded counters, ALB TargetResponseTime.
  • Compute Optimizer analyzes utilization history for EC2, Auto Scaling groups, EBS, Lambda, ECS on Fargate and RDS, and suggests a better type or size, including Graviton options. Memory-based recommendations need the CloudWatch agent.

Memory-bound apps

A Java, in-memory analytics or caching workload that's slow while CPU stays low, with swapping on the host, needs a memory-optimized (R family, or X and U for very large memory) instance, not more of the same instances.

Instance families

FamilyOptimized forExamples of fit
M (general purpose)Balanced CPU and memoryWeb and app servers, small databases
C (compute optimized)High CPU per GBBatch, encoding, gaming servers, HPC
R, X, z, U (memory optimized)High memory per vCPUIn-memory databases, SAP HANA, large caches, JVM heaps
I, D (storage optimized)Local NVMe or HDD, very high IOPSNoSQL databases, data warehouses, Kafka
P, G, Inf, Trn (accelerated)GPUs and AWS ML chipsTraining, inference, graphics
T (burstable)Low average CPU with burstsDev, small sites. Watch CPU credits
  • Graviton (a g after the generation number, like m7g, c7gn or r8g) usually gives better price performance for Linux workloads that can run on Arm: containers, Java, Python, Node.js, open-source databases, managed services.
  • T instances in unlimited mode that run hot all day cost more than an M instance of the same size. Constant credit exhaustion is a sign to move off T.

Storage

VolumePerformancePick it for
gp33,000 IOPS and 125 MiB/s baseline at any size; up to 80,000 IOPS and 2,000 MiB/s provisionedThe default. Tune IOPS and throughput without growing the volume
io2 Block ExpressUp to 256,000 IOPS, 4,000 MiB/s, sub-millisecond latency, 99.999% durabilityLarge production databases that need consistent high IOPS
st1Throughput-optimized HDDLarge sequential reads: logs, data processing
sc1Lowest-cost HDDCold, rarely read data
Instance storeHighest IOPS, local NVMeTemporary data, caches, replicated stores. Lost when the instance stops
3,000 / 125
gp3 baseline IOPS and MiB/s, regardless of volume size.
256,000
Maximum IOPS for an io2 Block Express volume.
7
Running instances per AZ in a spread placement group, and partitions per AZ in a partition group.
24 hours
Minimum history predictive scaling needs before it can forecast.

The instance caps the volume

Each instance type has its own EBS bandwidth and IOPS limits. A fast io2 volume on a small instance won't reach its provisioned IOPS. Check the instance's EBS-optimized limits before paying for more IOPS.

Exam signal

Moving gp2 volumes to gp3 is a common "improve performance and lower cost" answer. It's an in-place change with Elastic Volumes, with no downtime.

Networking and placement

Placement groupLayoutUse for
ClusterPacked close together in one AZLowest latency, highest throughput between nodes: HPC, tightly coupled jobs
SpreadEach instance on distinct hardwareA few critical instances that must not fail together
PartitionGroups of instances on separate racksLarge distributed systems that are rack-aware: HDFS, Cassandra, Kafka
  • Enhanced networking (ENA) is on by default on current instance types and gives higher bandwidth and packets per second with lower jitter.
  • Elastic Fabric Adapter (EFA) adds OS-bypass networking for MPI and NCCL, for tightly coupled HPC and distributed ML training. Use it with a cluster placement group.

Cluster groups and availability

A cluster placement group is in one AZ, so it trades availability for latency. If the question also asks for resilience against an AZ failure, a single cluster group isn't the answer.

Auto Scaling policies

PolicyHow it worksPick it when
Target trackingKeeps a metric at a target, like 50% CPU or 1,000 requests per targetThe default for most workloads
Step scalingAdds or removes capacity in steps based on how far an alarm is breachedYou need different reactions to small and large spikes
ScheduledChanges capacity at set timesKnown events: business hours, a sale starting at 09:00
PredictiveForecasts load from history and scales ahead of itRegular daily or weekly cycles where reacting is too late
Warm poolsKeeps pre-initialized instances stopped, hibernated or runningInstances take many minutes to boot and warm caches

Combine predictive scaling for the cycle with target tracking for surprises.

Global Accelerator vs CloudFront

CloudFrontGlobal Accelerator
TrafficHTTP and HTTPS, WebSocketAny TCP or UDP
CachingYes, at edge locationsNo
Entry pointDNS name, edge IPs changeTwo static anycast IP addresses
FailoverOrigin failover groupsShifts between Regional endpoints in seconds on health checks
Best forStatic and dynamic web content, APIs, videoGaming, VoIP, IoT, non-HTTP protocols, clients that need fixed IPs

Both carry traffic over the AWS backbone from the nearest edge location. See caching for improving cache hit ratios.

Scenarios

Scenario
A travel analytics company runs a Java service on c5.2xlarge instances. During nightly reports, response times rise sharply. CloudWatch agent metrics show memory utilization at 97% with heavy swapping, while CPU stays around 35%. What should the architect recommend to improve performance at the lowest cost?
Scenario · choose 2
A multiplayer game uses a custom UDP protocol on EC2 in us-east-1 and eu-west-1. Players complain about lag and dropped sessions, and some ISPs require the game's IP addresses to be allowlisted. Which TWO changes should the architect make?
Scenario
An order-processing fleet takes 12 minutes to boot and load reference data into memory. Traffic ramps up sharply every weekday at 08:00, and target tracking can't add capacity fast enough. Which change solves this with the LEAST cost?

Further reading

On this page