Alerting and dashboards
Turning logs, metrics and findings into alerts with CloudWatch metric filters and alarms, EventBridge rules and AWS User Notifications, plus health checks, AWS Health, Trusted Advisor and Managed Grafana dashboards.
Exam tasks: 1.1 (skills 1.1.1, 1.1.2 and 1.1.4: monitoring requirements, health checks, metrics, alerts and dashboards)
The decision: is the signal a line in a log, a number crossing a threshold, or an event some AWS service emits, and who needs to hear about it, how fast, and through which channel?
Picking the alerting mechanism
| Mechanism | Best for | Latency | Watch out |
|---|---|---|---|
| Metric filter + alarm | Patterns in CloudTrail or application logs (root use, failed sign-ins, AccessDenied spikes) | About a minute | Only counts events after the filter exists. Needs the logs in CloudWatch Logs |
| CloudWatch alarm | Numeric thresholds, anomaly detection bands, composite conditions | One evaluation period | Decide how missing data is treated |
| EventBridge rule | GuardDuty, Security Hub, Macie, Inspector, Health and CloudTrail events | Seconds | Rules are Regional, and CloudTrail events need an active trail |
| AWS User Notifications | Human-readable notifications from EventBridge events, CloudWatch alarms and AWS Health, org-wide | Minutes (aggregation optional) | Not a remediation engine |
CloudWatch metric filters and alarms
The classic pattern for "alert when someone does X": send the organization trail to a CloudWatch Logs log group, add a metric filter, and alarm on the metric.
{ ($.eventName = "DeleteTrail") || ($.eventName = "StopLogging") || ($.eventName = "PutEventSelectors") }- A filter that matches increments a custom metric (for example
Security/TrailTampering). An alarm with threshold 1 over 1 period sends to SNS. - Typical security filters: root user activity, console sign-in without MFA, IAM policy changes, security group
or NACL changes, KMS key disable or scheduled deletion,
UnauthorizedOperationandAccessDeniedbursts. - Metric filters aren't retroactive. For "how often did this happen last month", query the logs with Logs Insights or Athena instead.
- Alarm states are
OK,ALARMandINSUFFICIENT_DATA. For event counters, set missing data to notBreaching, otherwise quiet periods leave the alarm inINSUFFICIENT_DATA. - Anomaly detection alarms learn a band from history. Use them for "unusual volume" when there's no fixed
threshold, such as
GetObjectcalls per hour. - Composite alarms combine alarms with AND/OR to page only when several symptoms agree.
- CloudWatch Logs anomaly detection flags new or unusual log patterns without writing filters.
Exam signal
"Notify the security team within minutes whenever a specific API is called" with logs already in CloudWatch Logs has two correct shapes: metric filter + alarm + SNS, or an EventBridge rule on the CloudTrail event. Pick EventBridge when you also need the event details in the message or want to trigger Lambda, and the metric filter when the requirement is a count or rate ("more than 5 failures in 10 minutes").
EventBridge rules
{
"detail": {
"severity": [{ "numeric": [">=", 7] }],
"type": [{ "prefix": "CryptoCurrency:" }]
},
"detail-type": ["GuardDuty Finding"],
"source": ["aws.guardduty"]
}- Security services publish to the default event bus in the Region where the finding is. GuardDuty, Security Hub CSPM and Macie administrator accounts also receive their members' findings.
- For one central queue, add a rule in each Region that forwards to a central event bus in the security
account. The target bus needs a resource policy allowing the source accounts or the organization
(
aws:PrincipalOrgID). - Use an input transformer to turn raw JSON into a readable message before it reaches SNS or chat.
- Each target has a retry policy (by default up to 24 hours) and an optional dead-letter queue. Failures show
in the rule's
FailedInvocationsmetric. - CloudTrail-based events (
AWS API Call via CloudTrail) require a trail with logging on. Read-only calls (Get*,List*,Describe*) are only matched by rules in the stateENABLED_WITH_ALL_CLOUDTRAIL_MANAGEMENT_EVENTS.
One rule in us-east-1
A single EventBridge rule in us-east-1 does not see GuardDuty findings from eu-west-1. Findings, rules and default buses are Regional. Either create the rule in every Region, or use Security Hub CSPM cross-Region aggregation and build the rule in the aggregation Region.
AWS User Notifications
- One place to configure and read notifications from AWS services. AWS managed notifications (Health, billing) arrive by default; user-configured notifications are rules you create on EventBridge events and CloudWatch alarms.
- Channels: Console Notifications Center, email, Amazon Q Developer in chat applications (Slack, Microsoft Teams, Amazon Chime), and AWS Console Mobile App push.
- With AWS Organizations trusted access, the management account or a delegated administrator can create notification configurations for the whole organization or chosen OUs. Notifications appear only in the admin account, and they're kept for 90 days.
- Aggregation (within 5 minutes, within 12 hours, or none) groups related events to cut noise.
Legacy: use Amazon Q Developer in chat applications instead
AWS Chatbot, which posted SNS and EventBridge messages to Slack and Teams, was renamed Amazon Q Developer in chat applications. Configuration and behaviour are the same.
Exam signal
"Operators want console, email and Slack notifications for Health events and alarms across all accounts, with least effort and no custom code" is AWS User Notifications configured at the organization level. "Also isolate the instance" needs EventBridge and Lambda or SSM Automation, because User Notifications only notifies.
Health checks and AWS-side signals
Skill 1.1.2 asks for resource health checks as part of monitoring.
| Check | What it tells you |
|---|---|
| Route 53 health checks | Endpoint reachable and answering from many locations. Metrics live in us-east-1, so alarms on them go there |
| ELB target health | Which targets fail the load balancer's health check (UnHealthyHostCount) |
| EC2 status checks | System (AWS host) and instance (OS) reachability. Alarm actions can recover or reboot |
| CloudWatch Synthetics canaries | Scripted user journeys and API calls on a schedule, including broken-link and content checks |
| AWS Health | AWS-side events that affect your resources: scheduled retirement, service issues, and security notices such as AWS_RISK_CREDENTIALS_EXPOSED |
| Trusted Advisor | Best-practice checks: open security groups, public snapshots and buckets, root MFA, exposed access keys, service quotas |
AWS Health
- Events go to EventBridge as
source: aws.health, so you can auto-remediate, for example rotate a key when AWS reports it exposed. - Organizational view aggregates events for every account in the management account or a delegated administrator. Events stay for 90 days.
- The AWS Health API needs a Business, Enterprise On-Ramp or Enterprise support plan. The dashboard and EventBridge events are available to everyone.
Trusted Advisor
- Every account gets a core set of security checks. The full set, the API and organizational view need Business support or higher.
- Trusted Advisor also surfaces Security Hub CSPM FSBP control results, so enable both instead of choosing.
- Check status changes are published to EventBridge in us-east-1.
Legacy: use AWS Health Dashboard instead
The Personal Health Dashboard and the Service Health Dashboard were merged into the AWS Health Dashboard. The account-specific view is the same feature under the new name.
Dashboards
| Need | Use |
|---|---|
| CloudWatch metrics and Logs Insights widgets, one or many accounts | CloudWatch dashboards, with cross-account observability linking source accounts to a monitoring account |
| Security posture score, findings by severity, top resources | The Security Hub or Security Hub CSPM console dashboard |
| Mixed sources (CloudWatch, Athena over Security Lake, OpenSearch, Prometheus), viewers who sign in with the corporate IdP | Amazon Managed Grafana |
| Business-style reporting over Athena or Security Lake | Amazon Quick (formerly Amazon QuickSight) |
Amazon Managed Grafana
- A workspace is a managed Grafana server. Users sign in through IAM Identity Center or SAML, never IAM users, and get the Admin, Editor or Viewer role.
- Data sources include CloudWatch, Athena, OpenSearch Service, Amazon Managed Service for Prometheus, X-Ray and Timestream. In Organizations mode, the workspace can read data sources in member accounts through roles it creates.
- Connect it to private data sources through VPC configuration, and restrict who can reach the workspace with network access control.
Scenarios
GuardDuty findings are Regional but the administrator's default bus receives every member's findings for that Region, so one rule per Region in the security account covers the whole organization. A single us-east-1 rule misses the other Regions. Per-member alarms are far more effort and there's no such standard metric to rely on. An insight is a saved view, not an alert.
The metric filter with an account dimension gives the per-account rate alarm. A new filter can't count events from before it existed, so the historical report has to come from querying the logs themselves. EventBridge fires on each event rather than on a rate. GuardDuty findings don't map to a fixed AccessDenied threshold.
Managed Grafana supports IdP sign-in through IAM Identity Center or SAML, reads CloudWatch and Athena as data sources across member accounts, and has a Viewer role for read-only users. A public link exposes security data to anyone. IAM users mean console access and long-term credentials, which the requirement rules out. OpenSearch Dashboards doesn't natively chart CloudWatch metrics.
Further reading
Security findings hub
Which AWS security service finds what, how Security Hub and Security Hub CSPM aggregate findings in ASFF and OCSF, where Detective and Security Lake fit, and how to schedule recurring assessments.
Logging strategy
Choosing log sources for a security question, CloudTrail trails and organization trails, data and network activity events, VPC and transit gateway flow logs, Resolver query logs, service access logs, and a protected log archive account.