Asterrr's Handbook

Automated remediation

Detecting drift, risky changes and security findings in an existing environment and fixing them automatically with Config, EventBridge, Systems Manager Automation and Security Hub.

Exam tasks: 3.1 (use automation to remediate issues), 3.2 (automate security responses and compliance)

The decision: something keeps going wrong by hand. What detects it (a Config rule, an API call, a finding or a health event), and what runs the fix (SSM Automation, Lambda or Step Functions)?

Detect, route, fix

Which detector?

DetectorDetectsTypical fix path
AWS Config ruleA resource's configuration is non-compliant: unencrypted volume, public bucket, open port 22Built-in remediation: an SSM Automation runbook, manual or automatic
EventBridge + CloudTrailA specific API call happened: StopLogging, AuthorizeSecurityGroupIngress, PutBucketPolicyRule targets Lambda or SSM Automation within seconds
GuardDutyThreats: crypto-mining, credential exfiltration, malware, unusual API useEventBridge to Step Functions or Lambda to isolate, snapshot and notify
Security Hub CSPMFailed security-standard controls and aggregated findingsAutomation rules to triage; EventBridge or custom actions to fix
AWS HealthAWS-side events: scheduled EC2 retirement, maintenance, service issuesEventBridge to a runbook that stops and starts the instance, or opens a ticket
Trusted AdvisorBest-practice checks: exposed keys, open security groups, idle resourcesEventBridge on check status changes

State vs event

"Ensure resources stay compliant" or "report compliance history" means AWS Config. "Respond immediately when someone does X" means an EventBridge rule on the CloudTrail event. Config evaluates after a configuration change is recorded, so it's slower than an event rule.

AWS Config rules and remediation

  • Managed rules cover common checks, such as s3-bucket-public-read-prohibited, encrypted-volumes, restricted-ssh and iam-user-unused-credentials-check.
  • Custom rules run your logic as a Lambda function or as a Guard policy (policy as code, no function).
  • Triggers: on configuration change, periodically (for example every 24 hours), or both. Proactive mode evaluates a resource before it's deployed, for example from a CloudFormation hook.
  • Remediation attaches an SSM Automation document to the rule. Automatic remediation retries a set number of times. The runbook assumes an IAM role that you pass as a parameter.
  • Conformance packs bundle rules and remediations into one deployable unit, and organization rules or packs deploy them to every account from the management or delegated admin account.
  • An aggregator gives one compliance view across accounts and Regions.
Enable the Config recorder for the resource types you care about (all accounts, all Regions in use).
Deploy the rule, for example restricted-ssh, as an organization rule or in a conformance pack.
Attach remediation, for example AWS-DisablePublicAccessForSecurityGroup, set it to automatic, and give it an Automation role.
Watch the aggregator and route compliance-change events to a team channel.

Config doesn't block

Config detects and fixes after the fact. It never stops the API call. To prevent a change, use an SCP, a permissions boundary or a proactive control. Many answers combine both: prevent what you can and remediate the rest.

EventBridge rules on API calls

A rule matches an event pattern such as source: aws.cloudtrail with detail.eventName: StopLogging, and sends it to a target. Common patterns:

  • Someone disables the trail. Target Lambda that calls StartLogging and notifies security.
  • A security group opens 0.0.0.0/0 on port 3389. Target SSM Automation that revokes the rule.
  • A new S3 bucket is created. Target Lambda that turns on default encryption and Block Public Access.

For many accounts, send events to a central event bus in the security account with cross-account rules, and run the remediation there with a role it can assume in each account.

Security findings

  • GuardDuty findings go to EventBridge in the account and Region where they're generated (and to the delegated admin). Filter by finding type and severity.
  • Security Hub automation rules change findings as they arrive: raise severity for production accounts, set a workflow status, add notes, or suppress known-benign findings. They edit findings, they don't fix resources.
  • Custom actions in Security Hub send a chosen finding to EventBridge, for human-approved remediation.

Naming: Security Hub vs Security Hub CSPM

The original Security Hub service is now called Security Hub CSPM (standards, controls and finding aggregation). The newer AWS Security Hub correlates findings from GuardDuty, Inspector, Macie and CSPM into prioritized exposures. Exam questions that say "Security Hub" and "controls" or "security standards" mean CSPM.

Health and Trusted Advisor

  • AWS Health publishes account-specific events, such as scheduled instance retirement or RDS maintenance. Turn on the organizational view (from the management account or a delegated admin) to see events for every account in one place. EventBridge rules on aws.health can run a runbook that moves the instance before the retirement date.
  • Trusted Advisor checks cost, security, fault tolerance, performance and service limits. The full set of checks and the API need Business, Enterprise On-Ramp or Enterprise Support. Check status changes are sent to EventBridge, and an organizational view reports across accounts.

Exam signal

"Automatically notify the team, across all accounts, about AWS scheduled maintenance" points to the AWS Health organizational view plus EventBridge, not a Lambda that polls each account.

Scenarios

Scenario
A fintech company has 25 accounts in AWS Organizations. Auditors found security groups that allow SSH from 0.0.0.0/0. The company wants any such rule removed automatically in every account, current and future, and a compliance history for audits, with the LEAST custom code. What should the architect do?
Scenario · choose 2
A security team requires that if anyone stops CloudTrail logging in a production account, logging is restarted within one minute and the team is alerted. Which TWO actions meet this requirement?
Scenario
A retailer's SOC receives thousands of Security Hub CSPM findings a week. Findings from sandbox accounts about missing MFA on test users are expected, while any critical finding from production accounts must be escalated. The team wants this handled as findings arrive, without writing code. What should they use?

Further reading

On this page