Asterrr's Handbook

Disaster recovery

Choosing between backup and restore, pilot light, warm standby and multi-site from an RTO, an RPO and a budget.

Exam tasks: 1.3 (design DR from RTO/RPO, design backup and restore), 2.2 (configure DR, replication, DR testing), 3.4

The decision: the question gives you an RTO, an RPO and usually a budget signal. Pick the cheapest pattern that meets both targets.

RTO and RPO

  • RPO (recovery point objective) is about data: how far back the last usable copy can be. It depends on how often, or how continuously, you replicate.
  • RTO (recovery time objective) is about time: how long until users are served again. It depends on how much of the environment is already running in the recovery site.

The two are independent. A pilot light has an RPO of seconds, because the database is replicated continuously, but an RTO in the tens of minutes, because compute still has to start.

The four patterns

Cheaper · slower recoveryCostlier · faster recovery
  1. Backup and restore
    RPO hours · RTO hours

    Only backups exist in the recovery Region. Everything is rebuilt from IaC and restored from snapshots.

  2. Pilot light
    RPO seconds–minutes · RTO tens of minutes

    Data is replicated live. Compute is defined but switched off, or scaled to zero, until it's needed.

  3. Warm standby
    RPO seconds · RTO minutes

    A complete, working copy runs at reduced size. At failover you scale it up and switch traffic.

  4. Multi-site active/active
    RPO ≈ 0 · RTO ≈ 0

    Full capacity in two or more Regions, all serving traffic. Failover means removing an unhealthy Region.

What's running in the recovery RegionBackup and restorePilot lightWarm standbyActive/active
Data copiesSnapshots and backupsLive replicasLive replicasLive, often writable
DatabaseRestored when neededReplica runningReplica runningServing reads and writes
App serversNoneNone, or ASG at 0Small, runningFull size
Can take traffic right awayNoNoYes, at low capacityYes

Pilot light vs warm standby

Ask whether the recovery Region can serve a request right now. If it can, even slowly, it's warm standby. If compute has to start first, it's pilot light. Questions often describe the setup without naming the pattern.

AWS services that implement each pattern

NeedServiceWhat to know
Central, policy-based backupsAWS BackupBackup plans across services and accounts. Copies can go to another Region and another account
Backups ransomware can't deleteAWS Backup Vault LockIn compliance mode, nobody can delete recovery points early, including the root user
Server-level DR with an RPO of secondsElastic Disaster Recovery (DRS)Continuous block-level replication into a cheap staging subnet. Full instances launch only at failover or during drills
Relational DB across RegionsAurora Global DatabaseStorage-level replication with typical lag under a second. A secondary Region can be promoted in about a minute
NoSQL, writable in every RegionDynamoDB global tablesMulti-active. Conflicts resolve as last writer wins
Object dataS3 Cross-Region ReplicationAdd Replication Time Control for a 15-minute replication SLA
Moving trafficRoute 53 failover routing + health checksHealth checks run in the data plane, so they keep working when a Region's control plane is impaired
Controlled, manual failoverRoute 53 Application Recovery ControllerReadiness checks and routing controls that flip traffic with highly available data-plane API calls

Multi-AZ is not DR

Multi-AZ protects against losing a data center. It doesn't protect against losing a Region or corrupting data, because a bad write replicates to the standby instantly. For corruption you need point-in-time backups.

Backups in the same blast radius

Backups in the same account and Region as production don't survive a compromised account. Look for cross-account copies and Vault Lock whenever a question mentions ransomware or a malicious insider.

Failover that relies on the control plane

A plan that needs to create resources, change DNS records or edit configuration during a regional event is fragile. Prefer resources that already exist, with traffic shifted by health checks or ARC routing controls.

Test the plan

The exam guide explicitly includes performing DR testing. Signals worth knowing:

  • DRS can launch recovery instances for a drill without stopping replication.
  • AWS Fault Injection Service runs controlled experiments such as stopping instances, throttling APIs or cutting AZ connectivity.
  • AWS Resilience Hub checks an application's architecture against its RTO and RPO targets and suggests fixes.

Scenarios

Scenario
A company runs 200 on-premises servers on VMware. It needs a DR site on AWS with an RPO of seconds and an RTO of under 30 minutes, and wants to pay as little as possible while nothing has failed. Which solution fits?
Scenario · choose 2
An application uses Aurora MySQL and stateless EC2 instances in an Auto Scaling group. The business requires an RPO under 1 minute and an RTO under 15 minutes if a Region fails, at the lowest cost. Which TWO actions should the architect take?

Further reading

On this page