Asterrr's Handbook

Automated response

Wiring findings to action with EventBridge, Lambda, Step Functions and Systems Manager Automation, Security Hub automation rules and custom actions, Config remediation, and keeping automation safe with approvals and idempotency.

Exam tasks: 2.1.1 (runbooks), 2.1.4 (automatic remediation with Systems Manager, Step Functions and Lambda), 2.2.4 (containment at speed)

The decision: which responses are safe enough to run with no human involved, which need a person to press the button, and what is the simplest AWS wiring that turns a finding into that action in every account?

The standard pipeline

  • Run it from the delegated administrator (security tooling) account. GuardDuty and Security Hub forward member findings there, so one EventBridge rule covers the whole organization.
  • The target (Lambda, Step Functions, SSM Automation) acts in member accounts by assuming a remediation role you've deployed everywhere with StackSets.
  • Match narrowly in the event pattern: source, detail-type, then finding type, severity, account or resource tag.

Event names worth recognising

Sourcedetail-typeSent when
GuardDutyGuardDuty FindingNew findings in near real time. Repeat occurrences are aggregated and sent every 6 hours by default (the administrator can set 15 minutes or 1 hour)
Security Hub CSPMSecurity Hub Findings - ImportedEvery new or updated finding, automatically
Security Hub CSPMSecurity Hub Findings - Custom ActionAn analyst chooses a custom action on selected findings
Security Hub CSPMSecurity Hub Insight ResultsAn analyst sends insight results through a custom action
AWS HealthAWS Health EventFor example AWS_RISK_CREDENTIALS_EXPOSED
ConfigConfig Rules Compliance ChangeA resource becomes noncompliant

Waiting for the repeat

A playbook that triggers only on the first GuardDuty event is fine. One that relies on hearing about repeat occurrences quickly needs the notification frequency lowered to 15 minutes. And findings auto-archived by a GuardDuty suppression rule are not sent to EventBridge at all, so they never trigger automation.

Choosing the target

TargetUse it whenStrengths
SSM Automation runbookThe action is a known sequence on AWS resources: isolate an instance, snapshot volumes, block public accessAWS-provided runbooks (AWSSupport-ContainEC2Instance, AWS-DisableS3BucketPublicReadWrite and many more), aws:approve step, runs across accounts and Regions, audit trail per step
Step FunctionsSeveral steps with branches, waits, retries, parallel work and human approvalVisual history for the case record, .waitForTaskToken callback for approvals, error handling per state
LambdaOne quick, custom API action or glue between systemsSimple and fast; keep it small and idempotent
SNS, chat, OpsCenter, ticketsA person must decideNo risk of automated damage

Exam signal

"Least operational overhead" with a standard fix usually means an AWS-managed SSM Automation runbook (directly as the EventBridge target, or as a Config remediation action), not a custom Lambda function. "Multi-step workflow with a manual approval step" means Step Functions (or an SSM runbook with aws:approve).

Security Hub: automation rules versus custom actions

Automation rulesCustom actions
RunsAutomatically, as findings arriveWhen an analyst selects findings and chooses the action
DoesUpdates the finding: severity, workflow status (for example SUPPRESSED), notes, user-defined fields. In the new Security Hub, can also create Jira Cloud or ServiceNow ticketsSends the findings to EventBridge as Security Hub Findings - Custom Action, where your rule decides what runs
WhereAdministrator account, one set per RegionCreate in the aggregation Region if you use cross-Region aggregation
LimitsUp to 100 rules per administrator account; rule order decides which winsUp to 50 custom actions
Use forEnrichment, routing, suppressing known-good patterns, raising severity for crown-jewel accountsHuman-in-the-loop remediation: "Isolate instance", "Send to forensics"
  • Automation rules don't take actions on resources. For that, pair them with EventBridge (the updated finding is a new Imported event).
  • Automated Security Response on AWS is an AWS Solution that deploys ready-made remediation playbooks (Step Functions plus SSM runbooks) for Security Hub controls, triggered automatically or by custom actions.

Legacy: use Automated Security Response on AWS instead

The same solution was previously published as AWS Security Hub Automated Response and Remediation (SHARR).

Config rules with remediation

  • Attach a remediation action (an SSM Automation document) to a Config rule. Set it to manual (someone chooses Remediate) or automatic with a retry count and interval.
  • Remediation runs with an automation assume role you pass in the parameters. It needs permission to fix the resource, and only that.
  • Good candidates: public S3 buckets, security groups open to 0.0.0.0/0 on admin ports, unencrypted volumes or missing log settings. Deploy rules and remediations organization-wide with conformance packs.
  • Config evaluates on configuration change or on a schedule, so it's minutes, not seconds. For "block this before it happens", use preventive controls (SCPs, organization policies) instead.

AWS Health and Trusted Advisor events

  • AWS Health publishes account-specific events to EventBridge, including AWS_RISK_CREDENTIALS_EXPOSED when AWS finds your access key in public. A rule can run a Lambda function or SSM runbook that deactivates the key and opens a ticket.
  • Trusted Advisor check results (such as Exposed Access Keys, open security groups, MFA on root) also reach EventBridge. Trusted Advisor events are published in us-east-1, so the rule must live there.
  • Organizational views of both exist for the management or delegated administrator account.

Making automation safe

Automation that misfires at 3 a.m. is an incident of its own. Build these in:

  • Scope tightly. Auto-contain only for high-confidence, high-severity finding types and for resources you've agreed can be disrupted. Use a tag such as ir-auto-contain=false to opt specific resources out.
  • Idempotency. EventBridge delivers at least once, and asynchronous Lambda invocations retry on error. Before acting, check the current state (already isolated? key already inactive?) and key the work on the finding ID so a duplicate event does nothing.
  • Human approval where the blast radius is large. Deleting data, revoking a role used by a whole fleet or isolating a production database should wait for a person: Step Functions task token, SSM aws:approve, or a Security Hub custom action.
  • Least privilege for the automation role. It's a powerful identity. Allow only the containment actions, only through the expected path, and protect it with SCPs so workload admins can't modify it.
  • Failure handling. Dead-letter queues on EventBridge targets and Lambda, alarms on failed executions, and a notification whenever automation acts.
  • Log and record. Every automated action should leave a note on the finding or ticket so responders know what was already done.
  • Test it. Use GuardDuty sample findings and FIS experiments on non-production copies (see Response plan).
Human decides everythingFully automatic
  1. Notify
    Minutes to hours
    SNS, chat or ticket only. Safe, slow.
  2. One-click
    Minutes
    Security Hub custom action or OpsCenter "Run automation" starts a vetted runbook.
  3. Auto with approval
    Minutes
    Workflow gathers evidence automatically, then pauses for approval before containment.
  4. Auto-contain
    Seconds
    EventBridge runs containment directly. Only for narrow, high-confidence cases.
6 hours
Default GuardDuty EventBridge frequency for repeat findings (15 min or 1 h possible).
50
Security Hub CSPM custom actions per account.
100
Security Hub CSPM automation rules per administrator account.
us-east-1
Region where Trusted Advisor events reach EventBridge.

Scenarios

Scenario
Juniper Mobility's SOC wants analysts to be able to select a GuardDuty finding in Security Hub CSPM and, with one click, start a workflow that snapshots the affected instance's volumes, swaps its security groups and opens a ticket. Nothing should run unless an analyst chooses it. What should the security engineer build?
Scenario
A Lambda function triggered by EventBridge revokes sessions and attaches an isolation security group whenever GuardDuty reports a high-severity EC2 finding. Occasionally the same finding is processed twice, and the second run overwrites the saved list of original security groups with the isolation group, so the instance can't be restored after investigation. What is the best fix?
Scenario
A company with 90 accounts wants any security group that allows SSH from 0.0.0.0/0 to be fixed automatically within minutes in every account, with the least custom code and a record of compliance over time. What should the security engineer do?

Further reading

On this page