Purpose-built databases
Choosing between RDS, Aurora, DynamoDB, DocumentDB, Neptune, Keyspaces, Timestream, MemoryDB, Redshift and OpenSearch, and the DynamoDB and Aurora features the exam tests.
Exam tasks: 2.5 (select database and storage for performance), 4.3 (choose new architectures, including purpose-built databases)
The decision: what shape is the data and the access pattern (relational joins, key lookups, documents, graph traversals, time series, analytics or search), and how much scale, latency and operational effort is acceptable?
Choosing a database
Time-series data (IoT sensors, metrics) with time-window queries goes to Timestream, and a ledger or audit trail usually lands in a regular database with an append-only design (see the QLDB note below).
Side by side
| Service | Model | Signals in a question |
|---|---|---|
| RDS | Relational: MySQL, PostgreSQL, MariaDB, Oracle, SQL Server, Db2 | Commercial engine, lift and shift, license needs |
| Aurora | Relational, MySQL- or PostgreSQL-compatible | High throughput, 15 replicas, fast failover, global reads |
| Aurora DSQL | Distributed serverless PostgreSQL-compatible | Active-active SQL across Regions |
| DynamoDB | Key-value and document | Single-digit ms at any scale, serverless, known access patterns |
| DocumentDB | Document, MongoDB-compatible | Migrate a MongoDB workload with minimal changes |
| Neptune | Graph (Gremlin, openCypher, SPARQL) | Social networks, fraud rings, recommendations, knowledge graphs |
| Keyspaces | Wide-column, Cassandra-compatible (CQL) | Migrate Cassandra without running clusters |
| Timestream | Time series | IoT telemetry, DevOps metrics, time-window queries |
| MemoryDB | Durable in-memory, Valkey- and Redis OSS-compatible | Microsecond reads and durability as the primary DB |
| Redshift | Columnar data warehouse | BI, complex aggregations, petabyte analytics |
| OpenSearch Service | Search and log analytics | Full-text search, log exploration, dashboards |
Legacy: use Aurora PostgreSQL or DynamoDB with an append-only design instead
Amazon QLDB, the managed ledger database, reached end of support in July 2025. Older material still names it as the answer for an immutable, cryptographically verifiable journal.
Legacy: use Timestream for InfluxDB instead
Timestream for LiveAnalytics closed to new customers in June 2025. New time-series designs use Timestream for InfluxDB. Exam questions that just say "Timestream" still mean the time-series service in general.
Read replicas vs Multi-AZ
| Multi-AZ (instance) | Multi-AZ DB cluster | Read replicas | |
|---|---|---|---|
| Purpose | Availability | Availability plus read scaling | Read scaling |
| Replication | Synchronous | Semi-synchronous | Asynchronous |
| Standby readable | No | Yes, two readable standbys | Yes |
| Failover | Automatic, DNS endpoint moves | Automatic, typically faster | Manual promotion |
| Cross-Region | No | No | Yes |
- Aurora is different: storage keeps six copies across three AZs regardless, replicas share that storage, and any replica can be the failover target (tiers set the order). Failover usually takes under a minute.
- A cross-Region read replica is a cheap DR option for RDS. For Aurora, Global Database is the better one.
Read replicas for high availability
A read replica doesn't fail over automatically, so it isn't an HA answer for RDS. And Multi-AZ instance standbys can't serve reads, so they don't fix a read bottleneck. Match the feature to the problem.
Aurora features
- Serverless v2 scales capacity in fine-grained ACUs within seconds, and can scale to zero when idle. Mix serverless and provisioned instances in one cluster, for example a provisioned writer with serverless readers.
- Global Database replicates storage to up to 10 secondary Regions, typically with under a second of lag. Managed switchover for planned moves, failover for a regional outage. Write forwarding lets apps in a secondary Region send writes that the primary executes.
- Cloning uses copy-on-write, so a clone of a multi-TB cluster is ready in minutes and costs only for changed pages. Good for test copies of production.
- Backtrack (Aurora MySQL) rewinds a cluster in place to a point in time, without a restore.
- Custom endpoints send analytics queries to a subset of larger replicas.
Exam signal
"Test environment with a copy of production, created quickly, minimal extra storage" means Aurora cloning. "Unpredictable, spiky or intermittent SQL workload" means Aurora Serverless v2.
DynamoDB in depth
Capacity modes
| On-demand | Provisioned | |
|---|---|---|
| Billing | Per request | Per RCU and WCU per hour |
| Best for | Unknown, spiky or new workloads | Steady, predictable traffic |
| Scaling | Automatic | Auto scaling within limits you set |
| Cost lever | None beyond the mode | Reserved capacity for steady baselines |
You can switch modes on a table periodically, so start on-demand and move to provisioned once traffic is known.
Indexes
| Global secondary index | Local secondary index | |
|---|---|---|
| Keys | Any partition and sort key | Same partition key, different sort key |
| When created | Anytime | Only at table creation |
| Consistency | Eventually consistent | Strong or eventual |
| Capacity | Its own | Shares the table's |
| Size limit | None | 10 GB per partition key value |
| Per table | 20 by default | 5 |
Adding an LSI later
An LSI can't be added to an existing table. If a question needs a new query pattern on a live table, the answer is a GSI (or a new table and a migration).
Other features
- Streams: an ordered, 24-hour log of item changes, consumed by Lambda or the Kinesis Client Library adapter. Use it for triggers, audit and cross-system sync. Kinesis Data Streams for DynamoDB is the alternative for longer retention and more consumers.
- Global tables: multi-Region, multi-active. The default mode is eventually consistent with last writer wins. Multi-Region strong consistency (three Regions, or two plus a witness) gives an RPO of zero at higher write latency.
- TTL deletes expired items in the background for free, usually within a few days of expiry. Deletions appear in the stream, so you can archive them to S3.
- PITR restores to any second in the last 35 days (the window is configurable) into a new table. On-demand backups are kept until you delete them.
- Transactions give all-or-nothing writes across up to 100 items.
- Maximum item size is 400 KB. Store larger objects in S3 and keep a pointer.
The rest, briefly
- DocumentDB: storage architecture like Aurora's, MongoDB API compatibility, global clusters. Elastic clusters shard for very large write scale.
- Neptune: use it when queries walk relationships several hops deep, which is slow and awkward in SQL. Neptune Analytics handles large in-memory graph algorithms.
- Keyspaces: serverless Cassandra. No nodes, compaction or repair to manage.
- MemoryDB: a Multi-AZ transaction log makes writes durable, so it can be the only database. ElastiCache is a cache in front of another database.
- Redshift: RA3 nodes separate compute from managed storage. Redshift Serverless for intermittent use. Concurrency scaling handles bursts. See analytics patterns.
- OpenSearch Service: see analytics patterns.
Scenarios
Multi-AZ gives automatic failover to a synchronous standby, and a read replica takes the reporting load off the primary. A Multi-AZ instance standby can't be read (a Multi-AZ DB cluster would change this, but it isn't an option here). DAX works only with DynamoDB. Snapshots don't help with performance or failover.
Multi-hop relationship traversal is the core graph-database use case, and Neptune runs it with Gremlin or openCypher. Redshift is built for aggregations, not low-latency traversals. DynamoDB GSIs can model one hop but multi-hop queries need many round trips. OpenSearch is for search and log analytics.
A GSI can be added to an existing table and can use a different partition key, which a cross-user query needs. An LSI can only be created with the table and keeps the user ID partition key. Streams record changes, not a queryable index. PITR restores keep the original key schema.
Further reading
Caching
Where to cache (CloudFront, API Gateway, ElastiCache, DAX), how to raise the CloudFront cache hit ratio, and how to make a cache layer fault tolerant.
Compute and containers
Choosing between EC2, Lambda, ECS, EKS, Fargate and Elastic Beanstalk, plus ECS network modes, Lambda concurrency and VPC access, and edge compute.