Disaster recovery
Choosing between backup and restore, pilot light, warm standby and multi-site from an RTO, an RPO and a budget.
Exam tasks: 1.3 (design DR from RTO/RPO, design backup and restore), 2.2 (configure DR, replication, DR testing), 3.4
The decision: the question gives you an RTO, an RPO and usually a budget signal. Pick the cheapest pattern that meets both targets.
RTO and RPO
- RPO (recovery point objective) is about data: how far back the last usable copy can be. It depends on how often, or how continuously, you replicate.
- RTO (recovery time objective) is about time: how long until users are served again. It depends on how much of the environment is already running in the recovery site.
The two are independent. A pilot light has an RPO of seconds, because the database is replicated continuously, but an RTO in the tens of minutes, because compute still has to start.
The four patterns
- Backup and restoreRPO hours · RTO hours
Only backups exist in the recovery Region. Everything is rebuilt from IaC and restored from snapshots.
- Pilot lightRPO seconds–minutes · RTO tens of minutes
Data is replicated live. Compute is defined but switched off, or scaled to zero, until it's needed.
- Warm standbyRPO seconds · RTO minutes
A complete, working copy runs at reduced size. At failover you scale it up and switch traffic.
- Multi-site active/activeRPO ≈ 0 · RTO ≈ 0
Full capacity in two or more Regions, all serving traffic. Failover means removing an unhealthy Region.
| What's running in the recovery Region | Backup and restore | Pilot light | Warm standby | Active/active |
|---|---|---|---|---|
| Data copies | Snapshots and backups | Live replicas | Live replicas | Live, often writable |
| Database | Restored when needed | Replica running | Replica running | Serving reads and writes |
| App servers | None | None, or ASG at 0 | Small, running | Full size |
| Can take traffic right away | No | No | Yes, at low capacity | Yes |
Pilot light vs warm standby
Ask whether the recovery Region can serve a request right now. If it can, even slowly, it's warm standby. If compute has to start first, it's pilot light. Questions often describe the setup without naming the pattern.
AWS services that implement each pattern
| Need | Service | What to know |
|---|---|---|
| Central, policy-based backups | AWS Backup | Backup plans across services and accounts. Copies can go to another Region and another account |
| Backups ransomware can't delete | AWS Backup Vault Lock | In compliance mode, nobody can delete recovery points early, including the root user |
| Server-level DR with an RPO of seconds | Elastic Disaster Recovery (DRS) | Continuous block-level replication into a cheap staging subnet. Full instances launch only at failover or during drills |
| Relational DB across Regions | Aurora Global Database | Storage-level replication with typical lag under a second. A secondary Region can be promoted in about a minute |
| NoSQL, writable in every Region | DynamoDB global tables | Multi-active. Conflicts resolve as last writer wins |
| Object data | S3 Cross-Region Replication | Add Replication Time Control for a 15-minute replication SLA |
| Moving traffic | Route 53 failover routing + health checks | Health checks run in the data plane, so they keep working when a Region's control plane is impaired |
| Controlled, manual failover | Route 53 Application Recovery Controller | Readiness checks and routing controls that flip traffic with highly available data-plane API calls |
Multi-AZ is not DR
Multi-AZ protects against losing a data center. It doesn't protect against losing a Region or corrupting data, because a bad write replicates to the standby instantly. For corruption you need point-in-time backups.
Backups in the same blast radius
Backups in the same account and Region as production don't survive a compromised account. Look for cross-account copies and Vault Lock whenever a question mentions ransomware or a malicious insider.
Failover that relies on the control plane
A plan that needs to create resources, change DNS records or edit configuration during a regional event is fragile. Prefer resources that already exist, with traffic shifted by health checks or ARC routing controls.
Test the plan
The exam guide explicitly includes performing DR testing. Signals worth knowing:
- DRS can launch recovery instances for a drill without stopping replication.
- AWS Fault Injection Service runs controlled experiments such as stopping instances, throttling APIs or cutting AZ connectivity.
- AWS Resilience Hub checks an application's architecture against its RTO and RPO targets and suggests fixes.
Scenarios
An RPO of seconds rules out nightly and hourly copies. A full-size copy meets the targets but is the most expensive option. DRS replicates continuously to small staging instances and launches full-size servers only when needed, which meets both targets for the lowest standby cost.
This is pilot light. Aurora Global Database replicates with sub-second lag, which meets the RPO. An Auto Scaling group at zero in the secondary Region can scale out within the 15-minute RTO, and costs almost nothing until then. Multi-AZ doesn't help with a regional failure, hourly snapshots miss the RPO, and a full standby fleet costs more than necessary.
Further reading
Central security and logging
Collecting audit logs, findings and compliance data from every account and Region into dedicated security accounts, and making the logs tamper-proof.
Multi-account governance
Structuring an AWS organization with OUs, guardrails from SCPs, RCPs and Control Tower, and sharing or deploying resources across accounts with RAM, StackSets and Service Catalog.