Asterrr's Handbook

Deployment strategies

All-at-once, rolling, immutable, blue/green, canary and linear deployments, how CodeDeploy, Elastic Beanstalk and CloudFormation implement them, and how to roll back safely.

Exam tasks: 2.1 (deployment strategy to meet business requirements, IaC, CI/CD), 3.1 (improve operational excellence with automated, repeatable deployments)

The decision: how much downtime, extra capacity and blast radius can the release tolerate, and how fast must you be able to undo it?

Choosing a strategy

Side by side

StrategyDowntimeExtra capacityMixed versions liveRollback
All at onceYesNoneNoRedeploy the old version (slow, with downtime)
RollingNo, but capacity drops per batchNoneYesRedeploy the old version batch by batch
Rolling with additional batchNo, full capacity keptOne batchYesRedeploy the old version
ImmutableNoA full new fleet, temporarilyBrieflyTerminate the new fleet. The old one never changed
Blue/greenNoA full second environmentNo (all or nothing switch)Point traffic back at blue
CanaryNoNew version at small scaleYes, by designShift the small slice back
LinearNoGrows with each stepYes, by designShift traffic back at any step

Exam signal

Keywords map to strategies: "no reduction in capacity" means rolling with additional batch or immutable. "Quickest rollback" means blue/green. "Expose a small percentage of users first" means canary. "Shift 10% every few minutes" means linear.

How the services implement them

CodeDeploy

Compute platformIn-placeBlue/greenTraffic-shifting configs
EC2 / on-premisesYes. AllAtOnce, HalfAtATime, OneAtATime or a custom minimum-healthy-hosts valueYes. New instances (for example a copied Auto Scaling group) registered behind the load balancer, old ones terminated or keptNot traffic-weighted. The load balancer cuts over per instance
ECSNoAlways. A replacement task set in a second target groupECSAllAtOnce, canary (for example 10% then the rest after 5 minutes), linear (10% every minute or 3 minutes)
LambdaNoAlways. The alias weights shift between two versionsLambdaAllAtOnce, canary (10% for 5 to 30 minutes), linear (10% every 1 to 10 minutes)
  • EC2 in-place deployments run lifecycle hooks from appspec.yml (BeforeInstall, AfterInstall, ApplicationStart, ValidateService). The CodeDeploy agent must be on the instance.
  • ECS blue/green needs an ALB or NLB with two target groups and a production listener. An optional test listener lets you run checks against green before any user traffic moves.
  • Lambda deployments can run BeforeAllowTraffic and AfterAllowTraffic validation functions.
  • ECS now also has built-in blue/green, canary and linear deployments without CodeDeploy, using ALB, NLB or Service Connect. Expect exam answers to still name CodeDeploy.

Elastic Beanstalk

Beanstalk offers all at once, rolling, rolling with additional batch, immutable and traffic splitting (a canary: a percentage of requests goes to a new set of instances for an evaluation period).

Blue/green in Beanstalk is not a deployment policy. You clone the environment, deploy to the clone, then swap environment URLs (a CNAME swap). Because it's DNS-based, clients with cached records keep reaching blue for a while.

Blue/green and the database

Beanstalk blue/green swaps environments. If the RDS instance was created inside the environment, it belongs to blue and gets deleted with it. Production databases should be created outside Beanstalk and passed in through configuration.

Shifting traffic yourself

MechanismHow it shiftsRollback speedWatch out for
ALB weighted target groupsOne listener rule forwards, for example, 90/10 to two target groupsSeconds. Change the weightsOnly within one ALB, so both fleets sit in the same Region
Route 53 weighted recordsDNS answers split by weight between two endpointsMinutes, limited by TTL and client cachingWorks across Regions and load balancers, but it's not precise

Exam signal

If a question complains that users kept hitting the old version after a blue/green switch, the cause is DNS caching. Weighted target groups on one ALB avoid it.

Automated rollback

  • CodeDeploy can roll back automatically when a deployment fails or when a CloudWatch alarm you attach to the deployment group goes into ALARM (for example 5XX count or p99 latency on the green target group).
  • CloudFormation rollback triggers watch up to 5 alarms during a stack create or update, and for a monitoring period of up to 180 minutes afterwards. An alarm rolls the whole stack back.
  • Lambda and ECS canaries are the safest place for alarms: only the canary slice sees the bad version.
5
CloudWatch alarms per CloudFormation rollback-trigger configuration
180 min
Maximum monitoring period after a stack operation for rollback triggers
2
Target groups an ECS blue/green deployment needs
10%
The step size in CodeDeploy's predefined canary and linear configurations

CloudFormation safety controls

ControlWhat it doesUse it when
Change setsPreview exactly what an update adds, modifies or replaces before running itAny production update. Look for Replacement: True
Stack policyJSON policy that denies update actions (for example Update:Replace or Update:Delete) on named resourcesProtecting a database from accidental replacement. Can be overridden for one update
Termination protectionBlocks stack deletionLong-lived stacks
DeletionPolicyRetain, Snapshot or Delete when a resource leaves the stack. RetainExceptOnCreate keeps it unless the create failedKeeping data after a stack is deleted
UpdateReplacePolicyThe same choices, applied to the old resource when an update replaces itSnapshotting a DB before a replacement
Drift detectionCompares live resources with the template and reports changes made outside CloudFormationFinding console hotfixes. It reports, it doesn't fix
Custom resourcesA Lambda function or SNS topic runs your code on create, update and delete, then signals success or failureAnything CloudFormation can't model, such as seeding data or calling a third-party API

Drift detection doesn't remediate

Drift detection only shows the difference. To undo drift you update the stack (or import resources), or you use AWS Config rules with remediation. An answer that says drift detection "automatically reverts changes" is wrong.

Custom resources that never respond

A custom resource Lambda function must send a response to the pre-signed URL CloudFormation gives it, including on delete and on errors. If it doesn't, the stack hangs until the operation times out.

For rolling the same template into many accounts and Regions, use StackSets with service-managed permissions through AWS Organizations. See multi-account governance.

Cross-account pipelines

A common pattern: a tooling account runs CodePipeline, and it deploys into separate test and prod accounts.

What has to be in place:

  • A customer managed KMS key for the artifact bucket. The default AWS managed key can't be shared across accounts.
  • A key policy and a bucket policy that let the target account's roles decrypt and read artifacts.
  • A role in each target account that the pipeline's action assumes (set as the action's roleArn), plus a CloudFormation service role there with the actual deployment permissions.
  • A manual approval action before prod when the question mentions a change board or sign-off.

Scenarios

Scenario
A company runs a web app on Elastic Beanstalk with 12 instances. Releases must never reduce serving capacity, and if the new version fails health checks, rollback must not touch the instances running the current version. Which deployment policy should the company choose?
Scenario · choose 2
A team deploys a Lambda-based API through CodeDeploy. It wants 10% of invocations to use the new version for 10 minutes, and the release must revert automatically if the error rate rises during that time. Which TWO actions meet these requirements?
Scenario
An architect must update a CloudFormation stack that contains a production RDS instance. The team is worried a property change will replace the database. What is the MOST reliable way to find out before anything changes?

Further reading

On this page