Asterrr's Handbook

Backup automation

Replacing ad hoc snapshots in an existing estate with AWS Backup plans, org-wide backup policies, cross-account copies, Vault Lock, restore testing and Data Lifecycle Manager.

Exam tasks: 3.2 (improve reliability and security with automated backups), 1.3 and 2.2 for the DR side

The decision: backups exist today, but as scripts, console clicks or per-service settings. Which service centralizes them, where should copies live so an attacker can't delete them, and how do you prove they restore?

For choosing a DR pattern from RTO and RPO, see disaster recovery. This page covers the backup mechanics.

AWS Backup or Data Lifecycle Manager?

AWS BackupData Lifecycle ManagerNative service backups
CoversEC2, EBS, RDS, Aurora, DynamoDB, EFS, FSx, S3, Storage Gateway, Redshift, DocumentDB, Neptune, VMware and moreEBS snapshots and EBS-backed AMIs onlyOne service each (RDS automated backups, DynamoDB PITR)
TargetingTags or resource ARNs in a backup planTags on volumes or instances, or default policiesPer resource
Org-wideBackup policies in AWS OrganizationsPer account and RegionNo
Cross-Region / cross-account copyCopy actions in the plan ruleCross-Region copy; cross-account copy through event policiesSnapshot copy, service-specific
ImmutabilityVault Lock, logically air-gapped vaultsRecycle Bin rulesNo
AuditBackup Audit Manager frameworks and reportsNoNo

AWS Backup building blocks

  • Backup plan: rules with a schedule, a backup window, lifecycle (move to cold storage, then expire) and copy actions to other vaults.
  • Resource assignment: select by tag, for example every resource tagged backup=gold. New resources with the tag are protected automatically.
  • Backup vault: where recovery points live, encrypted with a KMS key, with a vault access policy.
  • Backup policies (an Organizations policy type) push plans to accounts and OUs from the management or delegated administrator account. Pair them with cross-account monitoring to see jobs everywhere.
  • Continuous backups give point-in-time restore for RDS, Aurora and S3 within the last 35 days.
90 days
Minimum stay in AWS Backup cold storage, and in the EBS snapshot archive tier.
72 hours
Minimum cooling-off period before a compliance-mode Vault Lock becomes permanent.
35 days
Longest retention for RDS automated backups and for continuous backups (point-in-time restore).
24–72 hours
Time to restore an archived EBS snapshot to the standard tier.

Protecting backups from deletion

ControlProtects againstNotes
Vault Lock, governance modeAccidental deletionUsers with the right IAM permissions can still remove the lock. Good for testing retention settings
Vault Lock, compliance modeRansomware, insiders, compromised rootAfter the cooling-off period, nobody, including AWS, can delete recovery points early or remove the lock
Logically air-gapped vaultLoss of the whole source accountLocked by default; share it with a recovery account through AWS RAM to restore from there
Cross-account copiesA compromised workload accountCopy into a separate backup account in the same organization
Recycle BinAccidental deletion of EBS snapshots, EBS-backed AMIs and EBS volumesRetention rules keep deleted items recoverable for a set period; rules can be locked

Encryption keys and cross-account copies

Cross-account copy fails for many resource types if the source is encrypted with an AWS managed key, because the target account can't use it. Encrypt with a customer managed key and allow the backup account to use it. An encrypted RDS snapshot shared with the default key has the same problem.

Exam signal

"Backups must be immutable for 7 years, and even administrators must not be able to delete them" means Vault Lock in compliance mode, usually on a vault in a separate account. S3 Object Lock is the equivalent for data you write to S3 yourself.

Proving backups restore

  • Restore testing in AWS Backup runs restore jobs on a schedule against chosen recovery points, reports the time each restore took, and deletes the test resources afterwards. Add a Lambda function on the restore job event to run your own validation.
  • Backup Audit Manager checks controls such as "resources are in a backup plan", "minimum retention" and "copies exist in another Region", and produces reports for auditors.

Data Lifecycle Manager for EBS

  • Snapshot policies target volumes or instances by tag, with up to four schedules per policy (for example hourly, daily, weekly and monthly), retention by count or age, cross-Region copy, fast snapshot restore, and archiving to the archive tier.
  • AMI policies create and deregister EBS-backed AMIs on a schedule.
  • Cross-account copy event policies copy snapshots shared with this account as soon as they're shared.
  • Default policies protect every volume or instance in the Region that isn't covered, with no tags needed.
  • Snapshots are crash-consistent by default. Pre and post scripts through Systems Manager make them application consistent.

RDS: automated backups vs snapshots

Automated backupsManual snapshots
Created byRDS, daily in the backup window, plus transaction logsYou, AWS Backup or a script
Retention1–35 days (0 turns them off)Until you delete them
Point-in-time restoreYes, typically to within the last 5 minutesNo, restores to the snapshot time
When the DB is deletedDeleted unless you choose to retain themKept
Cross-RegionAutomated backup replicationSnapshot copy

Snapshots for long retention

Automated backups can't be kept longer than 35 days. For monthly backups kept for a year, use AWS Backup (or scheduled manual snapshots), not a longer retention setting.

Scenarios

Scenario
A healthcare company has 80 accounts. Each team takes RDS and EBS snapshots with its own scripts, and auditors found gaps. The company wants daily backups of every resource tagged tier=prod in every account, a copy in a second Region, and protection against a compromised workload account deleting backups. Which approach needs the LEAST ongoing effort?
Scenario · choose 2
A company keeps 400 TB of EBS snapshots for compliance. Snapshots older than 6 months are almost never restored, and when they are, a wait of two days is acceptable. The team wants to lower storage cost and protect against snapshots being deleted by mistake. Which TWO actions should the architect take?
Scenario
After a ransomware drill, a bank's auditors ask for evidence that production backups can actually be restored, and how long each restore takes, every month. What should the architect set up?

Further reading

On this page