Asterrr's Handbook

Alerting and dashboards

Turning logs, metrics and findings into alerts with CloudWatch metric filters and alarms, EventBridge rules and AWS User Notifications, plus health checks, AWS Health, Trusted Advisor and Managed Grafana dashboards.

Exam tasks: 1.1 (skills 1.1.1, 1.1.2 and 1.1.4: monitoring requirements, health checks, metrics, alerts and dashboards)

The decision: is the signal a line in a log, a number crossing a threshold, or an event some AWS service emits, and who needs to hear about it, how fast, and through which channel?

Picking the alerting mechanism

MechanismBest forLatencyWatch out
Metric filter + alarmPatterns in CloudTrail or application logs (root use, failed sign-ins, AccessDenied spikes)About a minuteOnly counts events after the filter exists. Needs the logs in CloudWatch Logs
CloudWatch alarmNumeric thresholds, anomaly detection bands, composite conditionsOne evaluation periodDecide how missing data is treated
EventBridge ruleGuardDuty, Security Hub, Macie, Inspector, Health and CloudTrail eventsSecondsRules are Regional, and CloudTrail events need an active trail
AWS User NotificationsHuman-readable notifications from EventBridge events, CloudWatch alarms and AWS Health, org-wideMinutes (aggregation optional)Not a remediation engine

CloudWatch metric filters and alarms

The classic pattern for "alert when someone does X": send the organization trail to a CloudWatch Logs log group, add a metric filter, and alarm on the metric.

{ ($.eventName = "DeleteTrail") || ($.eventName = "StopLogging") || ($.eventName = "PutEventSelectors") }
  • A filter that matches increments a custom metric (for example Security/TrailTampering). An alarm with threshold 1 over 1 period sends to SNS.
  • Typical security filters: root user activity, console sign-in without MFA, IAM policy changes, security group or NACL changes, KMS key disable or scheduled deletion, UnauthorizedOperation and AccessDenied bursts.
  • Metric filters aren't retroactive. For "how often did this happen last month", query the logs with Logs Insights or Athena instead.
  • Alarm states are OK, ALARM and INSUFFICIENT_DATA. For event counters, set missing data to notBreaching, otherwise quiet periods leave the alarm in INSUFFICIENT_DATA.
  • Anomaly detection alarms learn a band from history. Use them for "unusual volume" when there's no fixed threshold, such as GetObject calls per hour.
  • Composite alarms combine alarms with AND/OR to page only when several symptoms agree.
  • CloudWatch Logs anomaly detection flags new or unusual log patterns without writing filters.

Exam signal

"Notify the security team within minutes whenever a specific API is called" with logs already in CloudWatch Logs has two correct shapes: metric filter + alarm + SNS, or an EventBridge rule on the CloudTrail event. Pick EventBridge when you also need the event details in the message or want to trigger Lambda, and the metric filter when the requirement is a count or rate ("more than 5 failures in 10 minutes").

EventBridge rules

{
  "detail": {
    "severity": [{ "numeric": [">=", 7] }],
    "type": [{ "prefix": "CryptoCurrency:" }]
  },
  "detail-type": ["GuardDuty Finding"],
  "source": ["aws.guardduty"]
}
  • Security services publish to the default event bus in the Region where the finding is. GuardDuty, Security Hub CSPM and Macie administrator accounts also receive their members' findings.
  • For one central queue, add a rule in each Region that forwards to a central event bus in the security account. The target bus needs a resource policy allowing the source accounts or the organization (aws:PrincipalOrgID).
  • Use an input transformer to turn raw JSON into a readable message before it reaches SNS or chat.
  • Each target has a retry policy (by default up to 24 hours) and an optional dead-letter queue. Failures show in the rule's FailedInvocations metric.
  • CloudTrail-based events (AWS API Call via CloudTrail) require a trail with logging on. Read-only calls (Get*, List*, Describe*) are only matched by rules in the state ENABLED_WITH_ALL_CLOUDTRAIL_MANAGEMENT_EVENTS.

One rule in us-east-1

A single EventBridge rule in us-east-1 does not see GuardDuty findings from eu-west-1. Findings, rules and default buses are Regional. Either create the rule in every Region, or use Security Hub CSPM cross-Region aggregation and build the rule in the aggregation Region.

AWS User Notifications

  • One place to configure and read notifications from AWS services. AWS managed notifications (Health, billing) arrive by default; user-configured notifications are rules you create on EventBridge events and CloudWatch alarms.
  • Channels: Console Notifications Center, email, Amazon Q Developer in chat applications (Slack, Microsoft Teams, Amazon Chime), and AWS Console Mobile App push.
  • With AWS Organizations trusted access, the management account or a delegated administrator can create notification configurations for the whole organization or chosen OUs. Notifications appear only in the admin account, and they're kept for 90 days.
  • Aggregation (within 5 minutes, within 12 hours, or none) groups related events to cut noise.

Legacy: use Amazon Q Developer in chat applications instead

AWS Chatbot, which posted SNS and EventBridge messages to Slack and Teams, was renamed Amazon Q Developer in chat applications. Configuration and behaviour are the same.

Exam signal

"Operators want console, email and Slack notifications for Health events and alarms across all accounts, with least effort and no custom code" is AWS User Notifications configured at the organization level. "Also isolate the instance" needs EventBridge and Lambda or SSM Automation, because User Notifications only notifies.

Health checks and AWS-side signals

Skill 1.1.2 asks for resource health checks as part of monitoring.

CheckWhat it tells you
Route 53 health checksEndpoint reachable and answering from many locations. Metrics live in us-east-1, so alarms on them go there
ELB target healthWhich targets fail the load balancer's health check (UnHealthyHostCount)
EC2 status checksSystem (AWS host) and instance (OS) reachability. Alarm actions can recover or reboot
CloudWatch Synthetics canariesScripted user journeys and API calls on a schedule, including broken-link and content checks
AWS HealthAWS-side events that affect your resources: scheduled retirement, service issues, and security notices such as AWS_RISK_CREDENTIALS_EXPOSED
Trusted AdvisorBest-practice checks: open security groups, public snapshots and buckets, root MFA, exposed access keys, service quotas

AWS Health

  • Events go to EventBridge as source: aws.health, so you can auto-remediate, for example rotate a key when AWS reports it exposed.
  • Organizational view aggregates events for every account in the management account or a delegated administrator. Events stay for 90 days.
  • The AWS Health API needs a Business, Enterprise On-Ramp or Enterprise support plan. The dashboard and EventBridge events are available to everyone.

Trusted Advisor

  • Every account gets a core set of security checks. The full set, the API and organizational view need Business support or higher.
  • Trusted Advisor also surfaces Security Hub CSPM FSBP control results, so enable both instead of choosing.
  • Check status changes are published to EventBridge in us-east-1.

Legacy: use AWS Health Dashboard instead

The Personal Health Dashboard and the Service Health Dashboard were merged into the AWS Health Dashboard. The account-specific view is the same feature under the new name.

Dashboards

NeedUse
CloudWatch metrics and Logs Insights widgets, one or many accountsCloudWatch dashboards, with cross-account observability linking source accounts to a monitoring account
Security posture score, findings by severity, top resourcesThe Security Hub or Security Hub CSPM console dashboard
Mixed sources (CloudWatch, Athena over Security Lake, OpenSearch, Prometheus), viewers who sign in with the corporate IdPAmazon Managed Grafana
Business-style reporting over Athena or Security LakeAmazon Quick (formerly Amazon QuickSight)

Amazon Managed Grafana

  • A workspace is a managed Grafana server. Users sign in through IAM Identity Center or SAML, never IAM users, and get the Admin, Editor or Viewer role.
  • Data sources include CloudWatch, Athena, OpenSearch Service, Amazon Managed Service for Prometheus, X-Ray and Timestream. In Organizations mode, the workspace can read data sources in member accounts through roles it creates.
  • Connect it to private data sources through VPC configuration, and restrict who can reach the workspace with network access control.
90 days
Retention of AWS Health organizational events and of organization notifications in User Notifications.
us-east-1
Region for Route 53 health check metrics and Trusted Advisor events.
24 hours
Default time EventBridge keeps retrying a failed target before dropping or sending to a DLQ.
0
Past events a new metric filter counts: filters apply only to new log data.

Scenarios

Scenario
A travel booking company uses a GuardDuty delegated administrator in its security account and operates in us-east-1, eu-west-1 and ap-southeast-1. The SOC wants a message in its Slack channel within minutes whenever a High or Critical GuardDuty finding appears in any account, with the least operational effort. What should the security engineer do?
Scenario
An organization trail delivers to a CloudWatch Logs log group in the log archive account. Auditors require an alert when more than 10 AccessDenied errors occur for the same account within 5 minutes, and a monthly report of how many such bursts happened in the previous year. Which approach meets BOTH requirements?
Scenario
A CISO wants a security dashboard for regional managers who have no AWS console access. It must combine CloudWatch metrics from 12 accounts with Athena queries over Security Lake, and managers must sign in with the corporate identity provider. What should the security engineer deploy?

Further reading

On this page