Performance and rightsizing
Finding the bottleneck in an existing workload and fixing it with the right instance family, storage, networking, placement and scaling policy.
Exam tasks: 3.3 (determine a strategy to improve performance)
The decision: which resource is actually saturated (CPU, memory, disk, network or a downstream dependency), and what's the smallest change that removes that limit?
Find the bottleneck first
- Metrics to check:
CPUUtilization, memory and swap (CloudWatch agent),VolumeQueueLengthand EBS throughput,NetworkIn/NetworkOutand ENA allowance-exceeded counters, ALBTargetResponseTime. - Compute Optimizer analyzes utilization history for EC2, Auto Scaling groups, EBS, Lambda, ECS on Fargate and RDS, and suggests a better type or size, including Graviton options. Memory-based recommendations need the CloudWatch agent.
Memory-bound apps
A Java, in-memory analytics or caching workload that's slow while CPU stays low, with swapping on the host, needs a memory-optimized (R family, or X and U for very large memory) instance, not more of the same instances.
Instance families
| Family | Optimized for | Examples of fit |
|---|---|---|
| M (general purpose) | Balanced CPU and memory | Web and app servers, small databases |
| C (compute optimized) | High CPU per GB | Batch, encoding, gaming servers, HPC |
| R, X, z, U (memory optimized) | High memory per vCPU | In-memory databases, SAP HANA, large caches, JVM heaps |
| I, D (storage optimized) | Local NVMe or HDD, very high IOPS | NoSQL databases, data warehouses, Kafka |
| P, G, Inf, Trn (accelerated) | GPUs and AWS ML chips | Training, inference, graphics |
| T (burstable) | Low average CPU with bursts | Dev, small sites. Watch CPU credits |
- Graviton (a
gafter the generation number, likem7g,c7gnorr8g) usually gives better price performance for Linux workloads that can run on Arm: containers, Java, Python, Node.js, open-source databases, managed services. - T instances in unlimited mode that run hot all day cost more than an M instance of the same size. Constant credit exhaustion is a sign to move off T.
Storage
| Volume | Performance | Pick it for |
|---|---|---|
| gp3 | 3,000 IOPS and 125 MiB/s baseline at any size; up to 80,000 IOPS and 2,000 MiB/s provisioned | The default. Tune IOPS and throughput without growing the volume |
| io2 Block Express | Up to 256,000 IOPS, 4,000 MiB/s, sub-millisecond latency, 99.999% durability | Large production databases that need consistent high IOPS |
| st1 | Throughput-optimized HDD | Large sequential reads: logs, data processing |
| sc1 | Lowest-cost HDD | Cold, rarely read data |
| Instance store | Highest IOPS, local NVMe | Temporary data, caches, replicated stores. Lost when the instance stops |
The instance caps the volume
Each instance type has its own EBS bandwidth and IOPS limits. A fast io2 volume on a small instance won't reach its provisioned IOPS. Check the instance's EBS-optimized limits before paying for more IOPS.
Exam signal
Moving gp2 volumes to gp3 is a common "improve performance and lower cost" answer. It's an in-place change with Elastic Volumes, with no downtime.
Networking and placement
| Placement group | Layout | Use for |
|---|---|---|
| Cluster | Packed close together in one AZ | Lowest latency, highest throughput between nodes: HPC, tightly coupled jobs |
| Spread | Each instance on distinct hardware | A few critical instances that must not fail together |
| Partition | Groups of instances on separate racks | Large distributed systems that are rack-aware: HDFS, Cassandra, Kafka |
- Enhanced networking (ENA) is on by default on current instance types and gives higher bandwidth and packets per second with lower jitter.
- Elastic Fabric Adapter (EFA) adds OS-bypass networking for MPI and NCCL, for tightly coupled HPC and distributed ML training. Use it with a cluster placement group.
Cluster groups and availability
A cluster placement group is in one AZ, so it trades availability for latency. If the question also asks for resilience against an AZ failure, a single cluster group isn't the answer.
Auto Scaling policies
| Policy | How it works | Pick it when |
|---|---|---|
| Target tracking | Keeps a metric at a target, like 50% CPU or 1,000 requests per target | The default for most workloads |
| Step scaling | Adds or removes capacity in steps based on how far an alarm is breached | You need different reactions to small and large spikes |
| Scheduled | Changes capacity at set times | Known events: business hours, a sale starting at 09:00 |
| Predictive | Forecasts load from history and scales ahead of it | Regular daily or weekly cycles where reacting is too late |
| Warm pools | Keeps pre-initialized instances stopped, hibernated or running | Instances take many minutes to boot and warm caches |
Combine predictive scaling for the cycle with target tracking for surprises.
Global Accelerator vs CloudFront
| CloudFront | Global Accelerator | |
|---|---|---|
| Traffic | HTTP and HTTPS, WebSocket | Any TCP or UDP |
| Caching | Yes, at edge locations | No |
| Entry point | DNS name, edge IPs change | Two static anycast IP addresses |
| Failover | Origin failover groups | Shifts between Regional endpoints in seconds on health checks |
| Best for | Static and dynamic web content, APIs, video | Gaming, VoIP, IoT, non-HTTP protocols, clients that need fixed IPs |
Both carry traffic over the AWS backbone from the nearest edge location. See caching for improving cache hit ratios.
Scenarios
The workload is memory bound, not CPU bound. A memory-optimized R instance gives about as much memory as the next compute-optimized size up for less money, and Graviton improves price performance further. A bigger C instance pays for CPU the service doesn't use, scaling out raises cost without fixing per-process heap pressure, and faster disk only speeds up swapping.
Global Accelerator carries UDP over the AWS backbone from the nearest edge and gives fixed anycast IPs for allowlisting, with fast failover between Regions. CloudFront doesn't proxy custom UDP. Latency-based DNS doesn't give fixed IPs and still sends traffic over the public internet. A cluster placement group is limited to one AZ and doesn't reduce latency to players.
Predictive scaling launches capacity before the known ramp, and a warm pool cuts the 12-minute warm-up for anything extra, while stopped instances cost only for their EBS volumes. Running at peak all day wastes money. Step scaling still reacts after the load arrives, and larger instances boot just as slowly.
Further reading
Backup automation
Replacing ad hoc snapshots in an existing estate with AWS Backup plans, org-wide backup policies, cross-account copies, Vault Lock, restore testing and Data Lifecycle Manager.
Cost optimization
Finding and removing waste in an existing AWS environment, from idle resources and NAT gateway charges to rightsizing before you commit.