Designing and testing a response plan
The incident response lifecycle on AWS, playbooks and runbooks, pre-provisioned access and a forensics account, AWS Security Incident Response, and testing the plan with FIS, Resilience Hub and ARC.
Exam tasks: 2.1.1 (response plans and runbooks), 2.1.2 (preparing services for incidents: access, tooling, blast radius, Shield Advanced), 2.1.3 (testing the plan with FIS and Resilience Hub), 2.1.4 (ARC in recovery)
The decision: what must already exist (people, access, tooling, isolated accounts, written and automated steps) so that responders can contain an incident in minutes without inventing anything, and how do you prove it works before a real attacker tests it for you?
The lifecycle, mapped to AWS
The AWS Security Incident Response Guide follows the NIST 800-61 phases. The exam cares less about the names than about knowing which AWS feature belongs where.
| Phase | Goal | AWS features you'll see in answers |
|---|---|---|
| Prepare | Nothing has to be invented during the incident | IAM Identity Center IR permission sets, a forensics account, SSM documents, Shield Advanced setup, CloudTrail org trail |
| Detect | Know quickly, with enough context | GuardDuty, Security Hub, Macie, Inspector, Config, CloudWatch alarms |
| Analyze | Real or noise, and how big | Detective, CloudTrail Lake, Athena over logs, Security Lake |
| Contain | Stop the spread, keep the evidence | Isolation security group, NACL, deny policies, AWSRevokeOlderSessions, SCPs |
| Eradicate and recover | Remove the foothold, return to known good | Rebuild from a golden AMI, AWS Backup restore, key and secret rotation, ARC traffic shifts |
| Post-incident | Make it less likely, and faster next time | Root cause, new Config rules or SCPs, runbook updates, OpsCenter or ticket follow-ups |
Playbooks and runbooks
- A playbook is the human-level plan for one incident type: who decides, what to check, when to escalate, who talks to legal and customers. Write one per likely scenario: exposed access key, crypto-mining instance, public S3 bucket, ransomware on EBS, compromised container.
- A runbook is the executable part: the exact steps, ideally as a Systems Manager Automation document that a responder, EventBridge rule or Step Functions workflow can run with one set of parameters.
- Keep both in version control and deploy the runbooks with CloudFormation StackSets so every account has the same version.
- Jupyter notebooks (for example in SageMaker AI) work well as investigation runbooks: pre-written Athena queries against CloudTrail and VPC Flow Logs, with space for the analyst's notes. The notebook becomes part of the case record.
Where the case lives
| Tool | Use it for | Notes |
|---|---|---|
| Systems Manager OpsCenter | OpsItems created automatically from EventBridge rules, CloudWatch alarms or Security Hub CSPM findings, with related resources and a Run automation button for SSM runbooks | Deduplicates by resource or a dedup string; SNS notification on change |
| AWS Security Incident Response | A managed service: AWS triages GuardDuty and Security Hub CSPM threat findings for you, opens cases, and gives you 24/7 access to AWS incident responders | Can run containment on EC2, S3 and IAM if you pre-authorize it; cases publish to EventBridge for Jira, ServiceNow or Slack |
| Your ITSM (Jira, ServiceNow) | The system of record most companies already have | Security Hub can create tickets directly; OpsCenter has APIs for sync |
| Systems Manager Incident Manager | Response plans, on-call rotations, escalation, chat channels, post-incident analysis | See the Legacy note below |
Legacy: use OpsCenter, AWS Security Incident Response, or your own ITSM tool instead
AWS Systems Manager Incident Manager stopped accepting new customers on 7 November 2025 and gets no new features. Existing accounts keep working. Older material may present it as the default answer for "engage the on-call responder and run a runbook". Know what it did (response plans, engagement and escalation plans, runbook automation, post-incident analysis), but for a new design prefer the options above.
Exam signal
"Reduce the time for the security team to engage AWS experts during an active incident" or "have AWS monitor and triage GuardDuty findings around the clock" points to AWS Security Incident Response. It doesn't generate findings itself, so GuardDuty (or another detector) must be on.
Preparing accounts and access
Pre-provisioned responder access
- Create IR roles or Identity Center permission sets now, not during the incident: a read-only investigator role and a separate, tightly monitored containment role that can change security groups, attach deny policies and snapshot volumes.
- Deploy them to every account through StackSets or Control Tower, with a trust policy that only the security account can assume.
- Keep break-glass access that doesn't depend on your IdP (the incident might be the IdP). Store the credentials offline, protect them with MFA, and alert on every use.
- Protect IR roles with an SCP so workload admins can't delete or edit them.
A forensics account
- A dedicated account in a security OU, owned by the security team, with no connection to production networks.
- Receives copies of snapshots (shared and re-encrypted with a customer managed key the forensics account can use), memory captures and logs.
- Contains an isolated VPC with no internet gateway, VPC endpoints for S3 and SSM, and a hardened forensic workstation AMI with your tools preinstalled.
- Stores artifacts in an S3 bucket with Object Lock (see Compromised workloads).
Minimizing blast radius before anything happens
- Separate accounts per workload and environment, so one compromised account doesn't mean everything is compromised (Multi-account strategy).
- SCPs that deny disabling CloudTrail, GuardDuty, Config and Security Hub, and deny leaving the organization.
- Short session durations and no long-lived access keys where roles work.
- Centralized, immutable logs in a log archive account (Logging strategy).
Shield Advanced as preparation
DDoS response depends on setup done in advance:
- Subscribe and add protected resources (CloudFront, Route 53 hosted zones, Global Accelerator, ALB, CLB, Elastic IPs).
- Associate Route 53 health checks with the protected resources. Shield uses them for faster, more accurate detection, and proactive engagement requires them.
- Enable proactive engagement and add emergency contacts so the Shield Response Team (SRT) calls you when a protected resource's health check goes unhealthy during an event.
- Grant the SRT an IAM role (the
AWSShieldDRTAccessPolicymanaged policy) and, if needed, access to your WAF log buckets so it can write mitigations for you. - Proactive engagement and SRT access need a Business or Enterprise support plan. See Edge protection.
Setting up SRT access during the attack
If the question asks how to get help faster during the next attack, the answer is the preparation: health checks, proactive engagement, and the SRT role, all configured in advance. Opening a support case mid-attack and then granting access is the slow path.
Testing the plan
A plan nobody has run is a guess. Test it at three levels.
| Test | What it proves | AWS help |
|---|---|---|
| Tabletop exercise | People know the playbook, decision rights and communication paths | None needed; AWS offers tabletop facilitation through its security programs |
| Simulation (game day) | Detection, alerting, runbooks and access actually work end to end | Generate GuardDuty sample findings, run the runbook, time it |
| Fault injection | The workload survives the containment and recovery actions themselves | AWS Fault Injection Service |
| Resilience assessment | The workload can meet its RTO and RPO | AWS Resilience Hub |
AWS Fault Injection Service (FIS)
- An experiment template has actions (stop or terminate instances, throttle APIs, inject network latency or packet loss, fail an AZ's network, pause EBS I/O), targets (chosen by ID, tag or filter) and stop conditions (if a named CloudWatch alarm goes into ALARM, FIS ends the run).
- FIS runs under an IAM role you give it, and it logs to CloudTrail. It can target other accounts with multi-account experiments.
- Security use: prove that isolating an instance (it disappears from service) or revoking a role's sessions doesn't take the whole application down, and that alarms and runbooks fire as expected.
AWS Resilience Hub
- You describe an application (from CloudFormation, Terraform, EKS, or tags), attach a resiliency policy with RTO and RPO targets for AZ, Region, infrastructure and application disruptions, and run an assessment.
- It reports whether the app can meet the policy, recommends changes, and generates recommended alarms, SOPs (SSM runbooks) and FIS experiments you can deploy.
- It detects drift when later changes break a policy that used to pass.
Exam signal
"Validate that the recovery runbook meets the 1-hour RTO" or "continuously check the app against RTO and RPO targets" means Resilience Hub. "Inject a failure and stop automatically if error rate rises" means FIS with a CloudWatch alarm stop condition.
Recovering traffic with ARC
Amazon Application Recovery Controller (ARC) moves traffic away from a bad place, which is useful when the bad place is a compromised cell, AZ or Region.
- Zonal shift: temporarily move load balancer traffic out of one AZ (for example while you rebuild instances there). Manual, with an expiry of up to 3 days that you can extend. Zonal autoshift lets AWS do it for AWS impairments.
- Routing controls: highly available on/off switches for DNS failover between Regions or cells, with safety rules that stop you from turning everything off at once.
- Region switch: plans that orchestrate a multi-Region failover across accounts.
- Readiness checks: monitor that the standby has matching quotas and capacity. Not for the failover path itself.
Legacy: use Amazon Application Recovery Controller (ARC) instead
Older material calls this Route 53 Application Recovery Controller. Same service, new name.
Scenarios
Pre-provisioned, centrally deployed roles and a ready forensics account remove both delays. Shared admin IAM users are long-lived credentials that break least privilege and auditability. Waiting for workload teams to grant access is exactly the delay you're trying to remove. Detective analyzes logs and findings; it doesn't hold or analyze disk snapshots.
FIS injects the disruption on real resources and uses CloudWatch alarms as stop conditions, which is exactly the safety lever requested. A Resilience Hub assessment is analysis, not a live disruption (though it can recommend FIS experiments). Sample findings test the detection path, not the workload's tolerance. Readiness checks watch standby capacity for failover, not in-Region instance loss.
Health checks give Shield an application-level signal and are required for proactive engagement, where the SRT contacts you. Granting the SRT role in advance lets it deploy WAF mitigations without waiting. An NLB doesn't inspect layer 7, Runtime Monitoring detects host threats rather than DDoS, and the idle timeout has nothing to do with engagement speed.
Further reading
Domain 2 · Incident response
14% of the exam. Being ready before something goes wrong, then containing, investigating and recovering without destroying the evidence.
Automated response
Wiring findings to action with EventBridge, Lambda, Step Functions and Systems Manager Automation, Security Hub automation rules and custom actions, Config remediation, and keeping automation safe with approvals and idempotency.