Asterrr's Handbook

Designing and testing a response plan

The incident response lifecycle on AWS, playbooks and runbooks, pre-provisioned access and a forensics account, AWS Security Incident Response, and testing the plan with FIS, Resilience Hub and ARC.

Exam tasks: 2.1.1 (response plans and runbooks), 2.1.2 (preparing services for incidents: access, tooling, blast radius, Shield Advanced), 2.1.3 (testing the plan with FIS and Resilience Hub), 2.1.4 (ARC in recovery)

The decision: what must already exist (people, access, tooling, isolated accounts, written and automated steps) so that responders can contain an incident in minutes without inventing anything, and how do you prove it works before a real attacker tests it for you?

The lifecycle, mapped to AWS

The AWS Security Incident Response Guide follows the NIST 800-61 phases. The exam cares less about the names than about knowing which AWS feature belongs where.

PhaseGoalAWS features you'll see in answers
PrepareNothing has to be invented during the incidentIAM Identity Center IR permission sets, a forensics account, SSM documents, Shield Advanced setup, CloudTrail org trail
DetectKnow quickly, with enough contextGuardDuty, Security Hub, Macie, Inspector, Config, CloudWatch alarms
AnalyzeReal or noise, and how bigDetective, CloudTrail Lake, Athena over logs, Security Lake
ContainStop the spread, keep the evidenceIsolation security group, NACL, deny policies, AWSRevokeOlderSessions, SCPs
Eradicate and recoverRemove the foothold, return to known goodRebuild from a golden AMI, AWS Backup restore, key and secret rotation, ARC traffic shifts
Post-incidentMake it less likely, and faster next timeRoot cause, new Config rules or SCPs, runbook updates, OpsCenter or ticket follow-ups

Playbooks and runbooks

  • A playbook is the human-level plan for one incident type: who decides, what to check, when to escalate, who talks to legal and customers. Write one per likely scenario: exposed access key, crypto-mining instance, public S3 bucket, ransomware on EBS, compromised container.
  • A runbook is the executable part: the exact steps, ideally as a Systems Manager Automation document that a responder, EventBridge rule or Step Functions workflow can run with one set of parameters.
  • Keep both in version control and deploy the runbooks with CloudFormation StackSets so every account has the same version.
  • Jupyter notebooks (for example in SageMaker AI) work well as investigation runbooks: pre-written Athena queries against CloudTrail and VPC Flow Logs, with space for the analyst's notes. The notebook becomes part of the case record.

Where the case lives

ToolUse it forNotes
Systems Manager OpsCenterOpsItems created automatically from EventBridge rules, CloudWatch alarms or Security Hub CSPM findings, with related resources and a Run automation button for SSM runbooksDeduplicates by resource or a dedup string; SNS notification on change
AWS Security Incident ResponseA managed service: AWS triages GuardDuty and Security Hub CSPM threat findings for you, opens cases, and gives you 24/7 access to AWS incident respondersCan run containment on EC2, S3 and IAM if you pre-authorize it; cases publish to EventBridge for Jira, ServiceNow or Slack
Your ITSM (Jira, ServiceNow)The system of record most companies already haveSecurity Hub can create tickets directly; OpsCenter has APIs for sync
Systems Manager Incident ManagerResponse plans, on-call rotations, escalation, chat channels, post-incident analysisSee the Legacy note below

Legacy: use OpsCenter, AWS Security Incident Response, or your own ITSM tool instead

AWS Systems Manager Incident Manager stopped accepting new customers on 7 November 2025 and gets no new features. Existing accounts keep working. Older material may present it as the default answer for "engage the on-call responder and run a runbook". Know what it did (response plans, engagement and escalation plans, runbook automation, post-incident analysis), but for a new design prefer the options above.

Exam signal

"Reduce the time for the security team to engage AWS experts during an active incident" or "have AWS monitor and triage GuardDuty findings around the clock" points to AWS Security Incident Response. It doesn't generate findings itself, so GuardDuty (or another detector) must be on.

Preparing accounts and access

Pre-provisioned responder access

  • Create IR roles or Identity Center permission sets now, not during the incident: a read-only investigator role and a separate, tightly monitored containment role that can change security groups, attach deny policies and snapshot volumes.
  • Deploy them to every account through StackSets or Control Tower, with a trust policy that only the security account can assume.
  • Keep break-glass access that doesn't depend on your IdP (the incident might be the IdP). Store the credentials offline, protect them with MFA, and alert on every use.
  • Protect IR roles with an SCP so workload admins can't delete or edit them.

A forensics account

  • A dedicated account in a security OU, owned by the security team, with no connection to production networks.
  • Receives copies of snapshots (shared and re-encrypted with a customer managed key the forensics account can use), memory captures and logs.
  • Contains an isolated VPC with no internet gateway, VPC endpoints for S3 and SSM, and a hardened forensic workstation AMI with your tools preinstalled.
  • Stores artifacts in an S3 bucket with Object Lock (see Compromised workloads).

Minimizing blast radius before anything happens

  • Separate accounts per workload and environment, so one compromised account doesn't mean everything is compromised (Multi-account strategy).
  • SCPs that deny disabling CloudTrail, GuardDuty, Config and Security Hub, and deny leaving the organization.
  • Short session durations and no long-lived access keys where roles work.
  • Centralized, immutable logs in a log archive account (Logging strategy).

Shield Advanced as preparation

DDoS response depends on setup done in advance:

  • Subscribe and add protected resources (CloudFront, Route 53 hosted zones, Global Accelerator, ALB, CLB, Elastic IPs).
  • Associate Route 53 health checks with the protected resources. Shield uses them for faster, more accurate detection, and proactive engagement requires them.
  • Enable proactive engagement and add emergency contacts so the Shield Response Team (SRT) calls you when a protected resource's health check goes unhealthy during an event.
  • Grant the SRT an IAM role (the AWSShieldDRTAccessPolicy managed policy) and, if needed, access to your WAF log buckets so it can write mitigations for you.
  • Proactive engagement and SRT access need a Business or Enterprise support plan. See Edge protection.

Setting up SRT access during the attack

If the question asks how to get help faster during the next attack, the answer is the preparation: health checks, proactive engagement, and the SRT role, all configured in advance. Opening a support case mid-attack and then granting access is the slow path.

Testing the plan

A plan nobody has run is a guess. Test it at three levels.

TestWhat it provesAWS help
Tabletop exercisePeople know the playbook, decision rights and communication pathsNone needed; AWS offers tabletop facilitation through its security programs
Simulation (game day)Detection, alerting, runbooks and access actually work end to endGenerate GuardDuty sample findings, run the runbook, time it
Fault injectionThe workload survives the containment and recovery actions themselvesAWS Fault Injection Service
Resilience assessmentThe workload can meet its RTO and RPOAWS Resilience Hub

AWS Fault Injection Service (FIS)

  • An experiment template has actions (stop or terminate instances, throttle APIs, inject network latency or packet loss, fail an AZ's network, pause EBS I/O), targets (chosen by ID, tag or filter) and stop conditions (if a named CloudWatch alarm goes into ALARM, FIS ends the run).
  • FIS runs under an IAM role you give it, and it logs to CloudTrail. It can target other accounts with multi-account experiments.
  • Security use: prove that isolating an instance (it disappears from service) or revoking a role's sessions doesn't take the whole application down, and that alarms and runbooks fire as expected.

AWS Resilience Hub

  • You describe an application (from CloudFormation, Terraform, EKS, or tags), attach a resiliency policy with RTO and RPO targets for AZ, Region, infrastructure and application disruptions, and run an assessment.
  • It reports whether the app can meet the policy, recommends changes, and generates recommended alarms, SOPs (SSM runbooks) and FIS experiments you can deploy.
  • It detects drift when later changes break a policy that used to pass.

Exam signal

"Validate that the recovery runbook meets the 1-hour RTO" or "continuously check the app against RTO and RPO targets" means Resilience Hub. "Inject a failure and stop automatically if error rate rises" means FIS with a CloudWatch alarm stop condition.

Recovering traffic with ARC

Amazon Application Recovery Controller (ARC) moves traffic away from a bad place, which is useful when the bad place is a compromised cell, AZ or Region.

  • Zonal shift: temporarily move load balancer traffic out of one AZ (for example while you rebuild instances there). Manual, with an expiry of up to 3 days that you can extend. Zonal autoshift lets AWS do it for AWS impairments.
  • Routing controls: highly available on/off switches for DNS failover between Regions or cells, with safety rules that stop you from turning everything off at once.
  • Region switch: plans that orchestrate a multi-Region failover across accounts.
  • Readiness checks: monitor that the standby has matching quotas and capacity. Not for the failover path itself.

Legacy: use Amazon Application Recovery Controller (ARC) instead

Older material calls this Route 53 Application Recovery Controller. Same service, new name.

NIST 800-61
Lifecycle the AWS IR guide and AWS Security Incident Response follow.
15 minutes
AWS Security Incident Response initial response objective for AWS-supported cases.
3 days
Maximum initial expiry of an ARC zonal shift (extendable).
Nov 2025
Incident Manager closed to new customers.

Scenarios

Scenario
A logistics company, Harbormaster Freight, runs 60 AWS accounts in AWS Organizations. During last quarter's incident, responders lost 90 minutes waiting for workload teams to grant them access and for someone to find a place to analyze EBS snapshots safely. What should the security team do to prevent this delay in future incidents?
Scenario
Thistle Media's playbook for a compromised web tier isolates instances by swapping their security groups. The team wants to confirm, before a real incident, that the application stays available when two instances are pulled out of service, and that the experiment halts automatically if the site's 5xx rate climbs above 2%. What should they use?
Scenario · choose 2
An online ticketing platform protects its CloudFront distribution and ALBs with AWS Shield Advanced. During a recent layer 7 attack, the team only learned of the impact from customer complaints, and the Shield Response Team could not act until the team granted access. Which TWO actions will speed up the response to the next event?

Further reading

On this page