Deployment strategies
All-at-once, rolling, immutable, blue/green, canary and linear deployments, how CodeDeploy, Elastic Beanstalk and CloudFormation implement them, and how to roll back safely.
Exam tasks: 2.1 (deployment strategy to meet business requirements, IaC, CI/CD), 3.1 (improve operational excellence with automated, repeatable deployments)
The decision: how much downtime, extra capacity and blast radius can the release tolerate, and how fast must you be able to undo it?
Choosing a strategy
Side by side
| Strategy | Downtime | Extra capacity | Mixed versions live | Rollback |
|---|---|---|---|---|
| All at once | Yes | None | No | Redeploy the old version (slow, with downtime) |
| Rolling | No, but capacity drops per batch | None | Yes | Redeploy the old version batch by batch |
| Rolling with additional batch | No, full capacity kept | One batch | Yes | Redeploy the old version |
| Immutable | No | A full new fleet, temporarily | Briefly | Terminate the new fleet. The old one never changed |
| Blue/green | No | A full second environment | No (all or nothing switch) | Point traffic back at blue |
| Canary | No | New version at small scale | Yes, by design | Shift the small slice back |
| Linear | No | Grows with each step | Yes, by design | Shift traffic back at any step |
Exam signal
Keywords map to strategies: "no reduction in capacity" means rolling with additional batch or immutable. "Quickest rollback" means blue/green. "Expose a small percentage of users first" means canary. "Shift 10% every few minutes" means linear.
How the services implement them
CodeDeploy
| Compute platform | In-place | Blue/green | Traffic-shifting configs |
|---|---|---|---|
| EC2 / on-premises | Yes. AllAtOnce, HalfAtATime, OneAtATime or a custom minimum-healthy-hosts value | Yes. New instances (for example a copied Auto Scaling group) registered behind the load balancer, old ones terminated or kept | Not traffic-weighted. The load balancer cuts over per instance |
| ECS | No | Always. A replacement task set in a second target group | ECSAllAtOnce, canary (for example 10% then the rest after 5 minutes), linear (10% every minute or 3 minutes) |
| Lambda | No | Always. The alias weights shift between two versions | LambdaAllAtOnce, canary (10% for 5 to 30 minutes), linear (10% every 1 to 10 minutes) |
- EC2 in-place deployments run lifecycle hooks from
appspec.yml(BeforeInstall,AfterInstall,ApplicationStart,ValidateService). The CodeDeploy agent must be on the instance. - ECS blue/green needs an ALB or NLB with two target groups and a production listener. An optional test listener lets you run checks against green before any user traffic moves.
- Lambda deployments can run
BeforeAllowTrafficandAfterAllowTrafficvalidation functions. - ECS now also has built-in blue/green, canary and linear deployments without CodeDeploy, using ALB, NLB or Service Connect. Expect exam answers to still name CodeDeploy.
Elastic Beanstalk
Beanstalk offers all at once, rolling, rolling with additional batch, immutable and traffic splitting (a canary: a percentage of requests goes to a new set of instances for an evaluation period).
Blue/green in Beanstalk is not a deployment policy. You clone the environment, deploy to the clone, then swap environment URLs (a CNAME swap). Because it's DNS-based, clients with cached records keep reaching blue for a while.
Blue/green and the database
Beanstalk blue/green swaps environments. If the RDS instance was created inside the environment, it belongs to blue and gets deleted with it. Production databases should be created outside Beanstalk and passed in through configuration.
Shifting traffic yourself
| Mechanism | How it shifts | Rollback speed | Watch out for |
|---|---|---|---|
| ALB weighted target groups | One listener rule forwards, for example, 90/10 to two target groups | Seconds. Change the weights | Only within one ALB, so both fleets sit in the same Region |
| Route 53 weighted records | DNS answers split by weight between two endpoints | Minutes, limited by TTL and client caching | Works across Regions and load balancers, but it's not precise |
Exam signal
If a question complains that users kept hitting the old version after a blue/green switch, the cause is DNS caching. Weighted target groups on one ALB avoid it.
Automated rollback
- CodeDeploy can roll back automatically when a deployment fails or when a CloudWatch alarm you attach to the deployment group goes into ALARM (for example 5XX count or p99 latency on the green target group).
- CloudFormation rollback triggers watch up to 5 alarms during a stack create or update, and for a monitoring period of up to 180 minutes afterwards. An alarm rolls the whole stack back.
- Lambda and ECS canaries are the safest place for alarms: only the canary slice sees the bad version.
CloudFormation safety controls
| Control | What it does | Use it when |
|---|---|---|
| Change sets | Preview exactly what an update adds, modifies or replaces before running it | Any production update. Look for Replacement: True |
| Stack policy | JSON policy that denies update actions (for example Update:Replace or Update:Delete) on named resources | Protecting a database from accidental replacement. Can be overridden for one update |
| Termination protection | Blocks stack deletion | Long-lived stacks |
| DeletionPolicy | Retain, Snapshot or Delete when a resource leaves the stack. RetainExceptOnCreate keeps it unless the create failed | Keeping data after a stack is deleted |
| UpdateReplacePolicy | The same choices, applied to the old resource when an update replaces it | Snapshotting a DB before a replacement |
| Drift detection | Compares live resources with the template and reports changes made outside CloudFormation | Finding console hotfixes. It reports, it doesn't fix |
| Custom resources | A Lambda function or SNS topic runs your code on create, update and delete, then signals success or failure | Anything CloudFormation can't model, such as seeding data or calling a third-party API |
Drift detection doesn't remediate
Drift detection only shows the difference. To undo drift you update the stack (or import resources), or you use AWS Config rules with remediation. An answer that says drift detection "automatically reverts changes" is wrong.
Custom resources that never respond
A custom resource Lambda function must send a response to the pre-signed URL CloudFormation gives it, including on delete and on errors. If it doesn't, the stack hangs until the operation times out.
For rolling the same template into many accounts and Regions, use StackSets with service-managed permissions through AWS Organizations. See multi-account governance.
Cross-account pipelines
A common pattern: a tooling account runs CodePipeline, and it deploys into separate test and prod accounts.
What has to be in place:
- A customer managed KMS key for the artifact bucket. The default AWS managed key can't be shared across accounts.
- A key policy and a bucket policy that let the target account's roles decrypt and read artifacts.
- A role in each target account that the pipeline's action assumes (set as the action's
roleArn), plus a CloudFormation service role there with the actual deployment permissions. - A manual approval action before prod when the question mentions a change board or sign-off.
Scenarios
Immutable launches a complete set of new instances in a temporary Auto Scaling group, and the existing instances are never modified, so a failed release is undone by terminating the new ones. Rolling with additional batch keeps capacity, but it updates the existing instances in place, so rolling back means redeploying to them. Plain rolling drops capacity, and all at once causes downtime.
The canary configuration sends 10% for 10 minutes and then shifts the remainder in one step. Linear would take 10 steps and about 100 minutes. A CloudWatch alarm on the deployment group triggers automatic rollback of the alias weights. Route 53 weighting is DNS-level and imprecise for this, and a stack policy prevents changes rather than reverting a bad release.
A change set lists every resource action, including whether a modification needs replacement. Drift detection compares live resources with the current template, not with the proposed one. Termination protection only blocks deleting the whole stack, and DeletionPolicy applies when the resource is removed from the stack, not when it's replaced (that's UpdateReplacePolicy).
Further reading
Domain 2 · New solutions
29% of the exam, the largest domain. Designing a new workload for deployment, continuity, security, reliability, performance and cost.
Configuration and patching
Using Systems Manager to patch, configure, access and inventory EC2 and on-premises servers at scale, and building golden AMIs with EC2 Image Builder.