Automated response
Wiring findings to action with EventBridge, Lambda, Step Functions and Systems Manager Automation, Security Hub automation rules and custom actions, Config remediation, and keeping automation safe with approvals and idempotency.
Exam tasks: 2.1.1 (runbooks), 2.1.4 (automatic remediation with Systems Manager, Step Functions and Lambda), 2.2.4 (containment at speed)
The decision: which responses are safe enough to run with no human involved, which need a person to press the button, and what is the simplest AWS wiring that turns a finding into that action in every account?
The standard pipeline
- Run it from the delegated administrator (security tooling) account. GuardDuty and Security Hub forward member findings there, so one EventBridge rule covers the whole organization.
- The target (Lambda, Step Functions, SSM Automation) acts in member accounts by assuming a remediation role you've deployed everywhere with StackSets.
- Match narrowly in the event pattern:
source,detail-type, then finding type, severity, account or resource tag.
Event names worth recognising
| Source | detail-type | Sent when |
|---|---|---|
| GuardDuty | GuardDuty Finding | New findings in near real time. Repeat occurrences are aggregated and sent every 6 hours by default (the administrator can set 15 minutes or 1 hour) |
| Security Hub CSPM | Security Hub Findings - Imported | Every new or updated finding, automatically |
| Security Hub CSPM | Security Hub Findings - Custom Action | An analyst chooses a custom action on selected findings |
| Security Hub CSPM | Security Hub Insight Results | An analyst sends insight results through a custom action |
| AWS Health | AWS Health Event | For example AWS_RISK_CREDENTIALS_EXPOSED |
| Config | Config Rules Compliance Change | A resource becomes noncompliant |
Waiting for the repeat
A playbook that triggers only on the first GuardDuty event is fine. One that relies on hearing about repeat occurrences quickly needs the notification frequency lowered to 15 minutes. And findings auto-archived by a GuardDuty suppression rule are not sent to EventBridge at all, so they never trigger automation.
Choosing the target
| Target | Use it when | Strengths |
|---|---|---|
| SSM Automation runbook | The action is a known sequence on AWS resources: isolate an instance, snapshot volumes, block public access | AWS-provided runbooks (AWSSupport-ContainEC2Instance, AWS-DisableS3BucketPublicReadWrite and many more), aws:approve step, runs across accounts and Regions, audit trail per step |
| Step Functions | Several steps with branches, waits, retries, parallel work and human approval | Visual history for the case record, .waitForTaskToken callback for approvals, error handling per state |
| Lambda | One quick, custom API action or glue between systems | Simple and fast; keep it small and idempotent |
| SNS, chat, OpsCenter, tickets | A person must decide | No risk of automated damage |
Exam signal
"Least operational overhead" with a standard fix usually means an AWS-managed SSM Automation runbook (directly as
the EventBridge target, or as a Config remediation action), not a custom Lambda function. "Multi-step workflow
with a manual approval step" means Step Functions (or an SSM runbook with aws:approve).
Security Hub: automation rules versus custom actions
| Automation rules | Custom actions | |
|---|---|---|
| Runs | Automatically, as findings arrive | When an analyst selects findings and chooses the action |
| Does | Updates the finding: severity, workflow status (for example SUPPRESSED), notes, user-defined fields. In the new Security Hub, can also create Jira Cloud or ServiceNow tickets | Sends the findings to EventBridge as Security Hub Findings - Custom Action, where your rule decides what runs |
| Where | Administrator account, one set per Region | Create in the aggregation Region if you use cross-Region aggregation |
| Limits | Up to 100 rules per administrator account; rule order decides which wins | Up to 50 custom actions |
| Use for | Enrichment, routing, suppressing known-good patterns, raising severity for crown-jewel accounts | Human-in-the-loop remediation: "Isolate instance", "Send to forensics" |
- Automation rules don't take actions on resources. For that, pair them with EventBridge (the updated finding is a
new
Importedevent). - Automated Security Response on AWS is an AWS Solution that deploys ready-made remediation playbooks (Step Functions plus SSM runbooks) for Security Hub controls, triggered automatically or by custom actions.
Legacy: use Automated Security Response on AWS instead
The same solution was previously published as AWS Security Hub Automated Response and Remediation (SHARR).
Config rules with remediation
- Attach a remediation action (an SSM Automation document) to a Config rule. Set it to manual (someone chooses Remediate) or automatic with a retry count and interval.
- Remediation runs with an automation assume role you pass in the parameters. It needs permission to fix the resource, and only that.
- Good candidates: public S3 buckets, security groups open to
0.0.0.0/0on admin ports, unencrypted volumes or missing log settings. Deploy rules and remediations organization-wide with conformance packs. - Config evaluates on configuration change or on a schedule, so it's minutes, not seconds. For "block this before it happens", use preventive controls (SCPs, organization policies) instead.
AWS Health and Trusted Advisor events
- AWS Health publishes account-specific events to EventBridge, including
AWS_RISK_CREDENTIALS_EXPOSEDwhen AWS finds your access key in public. A rule can run a Lambda function or SSM runbook that deactivates the key and opens a ticket. - Trusted Advisor check results (such as Exposed Access Keys, open security groups, MFA on root) also reach EventBridge. Trusted Advisor events are published in us-east-1, so the rule must live there.
- Organizational views of both exist for the management or delegated administrator account.
Making automation safe
Automation that misfires at 3 a.m. is an incident of its own. Build these in:
- Scope tightly. Auto-contain only for high-confidence, high-severity finding types and for resources you've
agreed can be disrupted. Use a tag such as
ir-auto-contain=falseto opt specific resources out. - Idempotency. EventBridge delivers at least once, and asynchronous Lambda invocations retry on error. Before acting, check the current state (already isolated? key already inactive?) and key the work on the finding ID so a duplicate event does nothing.
- Human approval where the blast radius is large. Deleting data, revoking a role used by a whole fleet or
isolating a production database should wait for a person: Step Functions task token, SSM
aws:approve, or a Security Hub custom action. - Least privilege for the automation role. It's a powerful identity. Allow only the containment actions, only through the expected path, and protect it with SCPs so workload admins can't modify it.
- Failure handling. Dead-letter queues on EventBridge targets and Lambda, alarms on failed executions, and a notification whenever automation acts.
- Log and record. Every automated action should leave a note on the finding or ticket so responders know what was already done.
- Test it. Use GuardDuty sample findings and FIS experiments on non-production copies (see Response plan).
- NotifyMinutes to hoursSNS, chat or ticket only. Safe, slow.
- One-clickMinutesSecurity Hub custom action or OpsCenter "Run automation" starts a vetted runbook.
- Auto with approvalMinutesWorkflow gathers evidence automatically, then pauses for approval before containment.
- Auto-containSecondsEventBridge runs containment directly. Only for narrow, high-confidence cases.
Scenarios
Custom actions send analyst-selected findings to EventBridge, and Step Functions coordinates the several steps. Automation rules only update finding fields and run automatically. An Imported-event rule fires on every finding without an analyst's choice. Config rules evaluate resource configuration, not GuardDuty threat findings.
EventBridge and Lambda can deliver or retry the same event, so the handler must detect work already done. Limiting concurrency to 1 serializes runs but still processes the duplicate. The frequency setting affects repeat occurrences, not duplicate delivery of one event. SNS is also at-least-once, so it doesn't solve anything.
A conformance pack deploys the managed rule and its SSM remediation organization-wide, and Config keeps the compliance history. A polling Lambda function is custom code to maintain and has no compliance record. Denying all ingress changes blocks legitimate work rather than fixing the bad rule. GuardDuty detects threats, and suppression hides findings rather than fixing anything.
Further reading
Designing and testing a response plan
The incident response lifecycle on AWS, playbooks and runbooks, pre-provisioned access and a forensics account, AWS Security Incident Response, and testing the plan with FIS, Resilience Hub and ARC.
Validating findings and scoping impact
Triaging GuardDuty, Security Hub, Macie and Inspector findings, telling true positives from expected activity, handling noise safely, and using Detective and log correlation to scope blast radius and root cause.