Backup automation
Replacing ad hoc snapshots in an existing estate with AWS Backup plans, org-wide backup policies, cross-account copies, Vault Lock, restore testing and Data Lifecycle Manager.
Exam tasks: 3.2 (improve reliability and security with automated backups), 1.3 and 2.2 for the DR side
The decision: backups exist today, but as scripts, console clicks or per-service settings. Which service centralizes them, where should copies live so an attacker can't delete them, and how do you prove they restore?
For choosing a DR pattern from RTO and RPO, see disaster recovery. This page covers the backup mechanics.
AWS Backup or Data Lifecycle Manager?
| AWS Backup | Data Lifecycle Manager | Native service backups | |
|---|---|---|---|
| Covers | EC2, EBS, RDS, Aurora, DynamoDB, EFS, FSx, S3, Storage Gateway, Redshift, DocumentDB, Neptune, VMware and more | EBS snapshots and EBS-backed AMIs only | One service each (RDS automated backups, DynamoDB PITR) |
| Targeting | Tags or resource ARNs in a backup plan | Tags on volumes or instances, or default policies | Per resource |
| Org-wide | Backup policies in AWS Organizations | Per account and Region | No |
| Cross-Region / cross-account copy | Copy actions in the plan rule | Cross-Region copy; cross-account copy through event policies | Snapshot copy, service-specific |
| Immutability | Vault Lock, logically air-gapped vaults | Recycle Bin rules | No |
| Audit | Backup Audit Manager frameworks and reports | No | No |
AWS Backup building blocks
- Backup plan: rules with a schedule, a backup window, lifecycle (move to cold storage, then expire) and copy actions to other vaults.
- Resource assignment: select by tag, for example every resource tagged
backup=gold. New resources with the tag are protected automatically. - Backup vault: where recovery points live, encrypted with a KMS key, with a vault access policy.
- Backup policies (an Organizations policy type) push plans to accounts and OUs from the management or delegated administrator account. Pair them with cross-account monitoring to see jobs everywhere.
- Continuous backups give point-in-time restore for RDS, Aurora and S3 within the last 35 days.
Protecting backups from deletion
| Control | Protects against | Notes |
|---|---|---|
| Vault Lock, governance mode | Accidental deletion | Users with the right IAM permissions can still remove the lock. Good for testing retention settings |
| Vault Lock, compliance mode | Ransomware, insiders, compromised root | After the cooling-off period, nobody, including AWS, can delete recovery points early or remove the lock |
| Logically air-gapped vault | Loss of the whole source account | Locked by default; share it with a recovery account through AWS RAM to restore from there |
| Cross-account copies | A compromised workload account | Copy into a separate backup account in the same organization |
| Recycle Bin | Accidental deletion of EBS snapshots, EBS-backed AMIs and EBS volumes | Retention rules keep deleted items recoverable for a set period; rules can be locked |
Encryption keys and cross-account copies
Cross-account copy fails for many resource types if the source is encrypted with an AWS managed key, because the target account can't use it. Encrypt with a customer managed key and allow the backup account to use it. An encrypted RDS snapshot shared with the default key has the same problem.
Exam signal
"Backups must be immutable for 7 years, and even administrators must not be able to delete them" means Vault Lock in compliance mode, usually on a vault in a separate account. S3 Object Lock is the equivalent for data you write to S3 yourself.
Proving backups restore
- Restore testing in AWS Backup runs restore jobs on a schedule against chosen recovery points, reports the time each restore took, and deletes the test resources afterwards. Add a Lambda function on the restore job event to run your own validation.
- Backup Audit Manager checks controls such as "resources are in a backup plan", "minimum retention" and "copies exist in another Region", and produces reports for auditors.
Data Lifecycle Manager for EBS
- Snapshot policies target volumes or instances by tag, with up to four schedules per policy (for example hourly, daily, weekly and monthly), retention by count or age, cross-Region copy, fast snapshot restore, and archiving to the archive tier.
- AMI policies create and deregister EBS-backed AMIs on a schedule.
- Cross-account copy event policies copy snapshots shared with this account as soon as they're shared.
- Default policies protect every volume or instance in the Region that isn't covered, with no tags needed.
- Snapshots are crash-consistent by default. Pre and post scripts through Systems Manager make them application consistent.
RDS: automated backups vs snapshots
| Automated backups | Manual snapshots | |
|---|---|---|
| Created by | RDS, daily in the backup window, plus transaction logs | You, AWS Backup or a script |
| Retention | 1–35 days (0 turns them off) | Until you delete them |
| Point-in-time restore | Yes, typically to within the last 5 minutes | No, restores to the snapshot time |
| When the DB is deleted | Deleted unless you choose to retain them | Kept |
| Cross-Region | Automated backup replication | Snapshot copy |
Snapshots for long retention
Automated backups can't be kept longer than 35 days. For monthly backups kept for a year, use AWS Backup (or scheduled manual snapshots), not a longer retention setting.
Scenarios
Backup policies push one plan to every account, tag-based assignment picks up new resources, and copies to a locked vault in another account survive a compromised workload account. Scripts are what caused the gaps, DLM doesn't cover RDS or provide central audit, and Multi-AZ and encryption aren't backups.
The archive tier costs much less for rarely restored snapshots, and a restore time of up to 72 hours fits the two-day tolerance. Recycle Bin keeps deleted snapshots recoverable. You can't copy EBS snapshots directly into Glacier storage classes with a script, fast snapshot restore adds cost rather than saving it, and AMIs don't reduce storage.
Restore testing runs real restores on a schedule, records how long they take, and cleans up afterwards. Config and audit controls show that backups exist, not that they restore, and manual restores are the process the bank wants to replace.
Further reading
Customer authentication
Adding sign-in, federation and API authorization to an existing customer-facing app with Cognito user pools, identity pools, ALB and API Gateway authorizers.
Performance and rightsizing
Finding the bottleneck in an existing workload and fixing it with the right instance family, storage, networking, placement and scaling policy.