Asterrr's Handbook

Caching

Where to cache (CloudFront, API Gateway, ElastiCache, DAX), how to raise the CloudFront cache hit ratio, and how to make a cache layer fault tolerant.

Exam tasks: 2.5 (design for performance: caching strategy), 3.3 (improve the performance of an existing solution)

The decision: which layer should answer the request so the expensive thing behind it (origin, API, database) does less work, and what happens to users when that cache fails or goes stale?

The layers

LayerCachesBest forWatch out for
CloudFrontHTTP responses at edge locationsStatic assets, cacheable API GETs, videoCache key choices decide the hit ratio
API Gateway cacheResponses per stage and methodRepeated identical REST API callsREST APIs only, billed per hour by size
ElastiCacheAnything your code puts in itDatabase query results, sessions, leaderboardsYour code handles misses and invalidation
DAXDynamoDB items and query resultsRead-heavy DynamoDB with microsecond needsEventually consistent reads only
Aurora or RDS read replicasNot a cache, but offload readsRead-heavy SQL with complex queriesReplica lag
300 s
Default API Gateway cache TTL. The maximum is 3,600 seconds, and 0 turns caching off.
0.5–237 GB
API Gateway cache sizes per stage.
5
Read replicas per ElastiCache (Valkey or Redis OSS) shard.
500
Maximum shards in a cluster-mode-enabled ElastiCache cluster.
Microseconds
DAX read latency, compared with single-digit milliseconds for DynamoDB itself.

CloudFront cache hit ratio

A cache hit happens only when a new request produces the same cache key as a cached object that hasn't expired. Every header, cookie or query string in the key multiplies the number of variants and cuts the hit ratio.

Cache policy vs origin request policy

Cache policyOrigin request policy
ControlsWhat's in the cache key, and the min, default and max TTLWhat CloudFront forwards to the origin
Affects hit ratioYes, directlyNo
Use it forValues that really change the response, such as langValues the origin needs but that don't change the cached object, such as a tracking header

Anything in the cache key is also sent to the origin. The origin request policy adds values without adding them to the key.

Raising the hit ratio

  • Include only what varies the response. Allowlist the two query strings that matter instead of all query strings. Don't put User-Agent, Authorization or session cookies in the key for shared content.
  • Normalize. Turn on Gzip and Brotli in the cache policy so CloudFront normalizes Accept-Encoding, rather than caching one copy per browser's header string. Use CloudFront Functions to lowercase or sort query strings.
  • Device-based variants: use the CloudFront-Is-Mobile-Viewer style headers (a few values) instead of the raw User-Agent (thousands of values).
  • TTLs. Have the origin send Cache-Control: max-age or s-maxage. The policy's minimum and maximum TTLs bound what the origin asks for.
  • Versioned file names (app.3f9c.js) with long TTLs beat invalidations for deployments. Invalidation is for mistakes, not the release process.
  • Origin Shield adds one central caching layer in a Regional edge cache in front of the origin. All edge locations fetch through it, so the origin sees one request per object instead of one per Regional edge cache. Use it for global audiences, multi-CDN setups or a fragile origin.
  • Monitor with the CloudFront cache statistics reports and the CacheHitRate metric (with additional metrics turned on).

Forwarding all headers

"Forward all headers to the origin" through the cache key means almost every request is a miss. If the origin needs headers, send them with an origin request policy and keep the cache key small.

Exam signal

"Origin is overloaded even though CloudFront is in front" points to the cache key (too many headers, cookies or query strings), short TTLs, or missing Origin Shield. Adding more origin servers is rarely the intended answer.

API Gateway caching

  • Enabled per stage, with per-method overrides. Only REST APIs have it; HTTP APIs don't.
  • Choose which request parameters (headers, query strings, path) form the cache key.
  • Clients can send Cache-Control: max-age=0 to bypass the cache. Require authorization for this, or anyone can force misses.
  • For a public, cacheable API, CloudFront in front of API Gateway gives edge caching too.

ElastiCache

Valkey or Redis OSS vs Memcached

Valkey or Redis OSSMemcached
Data typesStrings, hashes, lists, sets, sorted sets, streams, geospatialStrings (simple key/value)
Replication and failoverYes: replicas, Multi-AZ, automatic failoverNone. A lost node loses its data
Persistence and backupSnapshotsNo
Cross-RegionGlobal DatastoreNo
Pub/sub, transactions, LuaYesNo
ThreadsMostly single-threaded command processing, with I/O threadingMulti-threaded per node
ScaleShards (cluster mode) and replicasAdd nodes; client hashes keys across them

Valkey is the open-source fork of Redis that AWS now leads with, and it's API-compatible with Redis OSS. For the exam, treat them the same. ElastiCache Serverless removes node sizing entirely for either engine.

Exam signal

"Sorted leaderboard", "pub/sub", "session store that must survive a node failure", "replicate to another Region" all mean Valkey or Redis OSS. "Simplest cache, multi-threaded, data loss acceptable" means Memcached.

Fault tolerance

  • Multi-AZ with automatic failover: a replica in another AZ is promoted when the primary fails, and the primary endpoint follows it. Needs at least one replica.
  • Cluster mode disabled: one shard, up to five replicas. Scale reads with replicas, scale writes only by a bigger node.
  • Cluster mode enabled: data is partitioned across shards, so writes and memory scale horizontally, and one shard failing affects only its slice of keys. Resharding can happen online.
  • Global Datastore: a primary cluster replicates to read-only clusters in other Regions. Promote a secondary if the primary Region fails. Local reads for global users, plus a cross-Region DR option.
  • If the data must never be lost (it's the system of record, not a cache), use MemoryDB instead. It keeps a durable, Multi-AZ transaction log. See purpose-built databases.

Memcached for high availability

Spreading Memcached nodes across AZs limits how much of the cache one AZ outage empties, but there's no replication. An answer that needs a cache to fail over without losing data needs Valkey or Redis OSS with replicas and Multi-AZ.

DAX

  • An in-VPC, write-through cache that speaks the DynamoDB API. Swap the client library and keep your code.
  • Caches items (GetItem, BatchGetItem) and query or scan results separately.
  • Serves eventually consistent reads. Strongly consistent reads pass straight through to DynamoDB.
  • Helps read-heavy, repeated-key workloads. It doesn't help write-heavy tables and doesn't raise write capacity.
  • Run at least three nodes across AZs for production.

Caching strategies

Lazy loading (cache-aside)Write-through
HowOn a miss, read the DB and write the result to the cacheEvery DB write also updates the cache
Data in cacheOnly what's been requestedEverything written, even if never read
Stale dataPossible until the TTL expiresRare
Miss penaltyThree trips: cache, DB, cache writeLow, but every write costs more
Node failureCache refills graduallyEmpty cache until data is written again

Combine them: write-through for data that must be fresh, plus a TTL on every key so mistakes and orphaned keys expire. Add a little random jitter to TTLs so many keys don't expire at once and stampede the database.

Scenarios

Scenario
A news site serves articles through CloudFront from an ALB origin. The cache hit ratio is 12% and the origin fleet is overloaded during breaking news. The cache policy forwards all query strings and all cookies. Articles vary only by the 'edition' query string. The origin needs the 'x-request-source' header for logging. What should the architect do to raise the hit ratio?
Scenario · choose 2
A gaming company stores player sessions and a global leaderboard in a single-node ElastiCache for Memcached cluster. A node failure logged out every player. Players in Europe and Asia complain of slow leaderboard reads from us-east-1. Which TWO changes address both problems?
Scenario
A product catalog on DynamoDB receives 200,000 reads per second, mostly for the same few thousand popular items, and the team wants microsecond latency with minimal code changes. Writes are light. Which option fits BEST?

Further reading

On this page