Caching
Where to cache (CloudFront, API Gateway, ElastiCache, DAX), how to raise the CloudFront cache hit ratio, and how to make a cache layer fault tolerant.
Exam tasks: 2.5 (design for performance: caching strategy), 3.3 (improve the performance of an existing solution)
The decision: which layer should answer the request so the expensive thing behind it (origin, API, database) does less work, and what happens to users when that cache fails or goes stale?
The layers
| Layer | Caches | Best for | Watch out for |
|---|---|---|---|
| CloudFront | HTTP responses at edge locations | Static assets, cacheable API GETs, video | Cache key choices decide the hit ratio |
| API Gateway cache | Responses per stage and method | Repeated identical REST API calls | REST APIs only, billed per hour by size |
| ElastiCache | Anything your code puts in it | Database query results, sessions, leaderboards | Your code handles misses and invalidation |
| DAX | DynamoDB items and query results | Read-heavy DynamoDB with microsecond needs | Eventually consistent reads only |
| Aurora or RDS read replicas | Not a cache, but offload reads | Read-heavy SQL with complex queries | Replica lag |
CloudFront cache hit ratio
A cache hit happens only when a new request produces the same cache key as a cached object that hasn't expired. Every header, cookie or query string in the key multiplies the number of variants and cuts the hit ratio.
Cache policy vs origin request policy
| Cache policy | Origin request policy | |
|---|---|---|
| Controls | What's in the cache key, and the min, default and max TTL | What CloudFront forwards to the origin |
| Affects hit ratio | Yes, directly | No |
| Use it for | Values that really change the response, such as lang | Values the origin needs but that don't change the cached object, such as a tracking header |
Anything in the cache key is also sent to the origin. The origin request policy adds values without adding them to the key.
Raising the hit ratio
- Include only what varies the response. Allowlist the two query strings that matter instead of all query
strings. Don't put
User-Agent,Authorizationor session cookies in the key for shared content. - Normalize. Turn on Gzip and Brotli in the cache policy so CloudFront normalizes
Accept-Encoding, rather than caching one copy per browser's header string. Use CloudFront Functions to lowercase or sort query strings. - Device-based variants: use the
CloudFront-Is-Mobile-Viewerstyle headers (a few values) instead of the rawUser-Agent(thousands of values). - TTLs. Have the origin send
Cache-Control: max-ageors-maxage. The policy's minimum and maximum TTLs bound what the origin asks for. - Versioned file names (
app.3f9c.js) with long TTLs beat invalidations for deployments. Invalidation is for mistakes, not the release process. - Origin Shield adds one central caching layer in a Regional edge cache in front of the origin. All edge locations fetch through it, so the origin sees one request per object instead of one per Regional edge cache. Use it for global audiences, multi-CDN setups or a fragile origin.
- Monitor with the CloudFront cache statistics reports and the
CacheHitRatemetric (with additional metrics turned on).
Forwarding all headers
"Forward all headers to the origin" through the cache key means almost every request is a miss. If the origin needs headers, send them with an origin request policy and keep the cache key small.
Exam signal
"Origin is overloaded even though CloudFront is in front" points to the cache key (too many headers, cookies or query strings), short TTLs, or missing Origin Shield. Adding more origin servers is rarely the intended answer.
API Gateway caching
- Enabled per stage, with per-method overrides. Only REST APIs have it; HTTP APIs don't.
- Choose which request parameters (headers, query strings, path) form the cache key.
- Clients can send
Cache-Control: max-age=0to bypass the cache. Require authorization for this, or anyone can force misses. - For a public, cacheable API, CloudFront in front of API Gateway gives edge caching too.
ElastiCache
Valkey or Redis OSS vs Memcached
| Valkey or Redis OSS | Memcached | |
|---|---|---|
| Data types | Strings, hashes, lists, sets, sorted sets, streams, geospatial | Strings (simple key/value) |
| Replication and failover | Yes: replicas, Multi-AZ, automatic failover | None. A lost node loses its data |
| Persistence and backup | Snapshots | No |
| Cross-Region | Global Datastore | No |
| Pub/sub, transactions, Lua | Yes | No |
| Threads | Mostly single-threaded command processing, with I/O threading | Multi-threaded per node |
| Scale | Shards (cluster mode) and replicas | Add nodes; client hashes keys across them |
Valkey is the open-source fork of Redis that AWS now leads with, and it's API-compatible with Redis OSS. For the exam, treat them the same. ElastiCache Serverless removes node sizing entirely for either engine.
Exam signal
"Sorted leaderboard", "pub/sub", "session store that must survive a node failure", "replicate to another Region" all mean Valkey or Redis OSS. "Simplest cache, multi-threaded, data loss acceptable" means Memcached.
Fault tolerance
- Multi-AZ with automatic failover: a replica in another AZ is promoted when the primary fails, and the primary endpoint follows it. Needs at least one replica.
- Cluster mode disabled: one shard, up to five replicas. Scale reads with replicas, scale writes only by a bigger node.
- Cluster mode enabled: data is partitioned across shards, so writes and memory scale horizontally, and one shard failing affects only its slice of keys. Resharding can happen online.
- Global Datastore: a primary cluster replicates to read-only clusters in other Regions. Promote a secondary if the primary Region fails. Local reads for global users, plus a cross-Region DR option.
- If the data must never be lost (it's the system of record, not a cache), use MemoryDB instead. It keeps a durable, Multi-AZ transaction log. See purpose-built databases.
Memcached for high availability
Spreading Memcached nodes across AZs limits how much of the cache one AZ outage empties, but there's no replication. An answer that needs a cache to fail over without losing data needs Valkey or Redis OSS with replicas and Multi-AZ.
DAX
- An in-VPC, write-through cache that speaks the DynamoDB API. Swap the client library and keep your code.
- Caches items (
GetItem,BatchGetItem) and query or scan results separately. - Serves eventually consistent reads. Strongly consistent reads pass straight through to DynamoDB.
- Helps read-heavy, repeated-key workloads. It doesn't help write-heavy tables and doesn't raise write capacity.
- Run at least three nodes across AZs for production.
Caching strategies
| Lazy loading (cache-aside) | Write-through | |
|---|---|---|
| How | On a miss, read the DB and write the result to the cache | Every DB write also updates the cache |
| Data in cache | Only what's been requested | Everything written, even if never read |
| Stale data | Possible until the TTL expires | Rare |
| Miss penalty | Three trips: cache, DB, cache write | Low, but every write costs more |
| Node failure | Cache refills gradually | Empty cache until data is written again |
Combine them: write-through for data that must be fresh, plus a TTL on every key so mistakes and orphaned keys expire. Add a little random jitter to TTLs so many keys don't expire at once and stampede the database.
Scenarios
Cookies and irrelevant query strings split each article into thousands of cache variants. Keeping only 'edition' in the key collapses them, and the origin request policy still delivers the logging header without affecting the key. Origin Shield helps a little but doesn't fix a fragmented key. Adding the header to the key makes things worse. Frequent invalidations reduce hits.
Valkey supports replication with automatic failover, so sessions survive a node failure, and sorted sets suit a leaderboard. Global Datastore then gives low-latency reads in the other Regions. More Memcached nodes spread the data but don't replicate it. DAX only works with DynamoDB. Auto Discovery helps clients find nodes but adds no durability.
DAX is API-compatible with DynamoDB, so only the client changes, and it returns hot items in microseconds. A GSI doesn't cache anything. ElastiCache works but needs more code. Capacity mode changes throughput limits, not latency.
Further reading
Decoupling with queues, topics and events
When to use SQS, SNS, EventBridge, Step Functions, Kinesis Data Streams, Amazon MQ or MSK to break a system into independent parts.
Purpose-built databases
Choosing between RDS, Aurora, DynamoDB, DocumentDB, Neptune, Keyspaces, Timestream, MemoryDB, Redshift and OpenSearch, and the DynamoDB and Aurora features the exam tests.