Observability
Adding the right CloudWatch metrics, alarms, log pipelines, tracing and synthetic checks to an existing workload, across many accounts.
Exam tasks: 3.1 (determine a strategy to improve operational excellence: monitoring, logging, alerting)
The decision: an existing system has a visibility gap. Which signal is missing (metric, log, trace or user-side check), and what's the smallest change that adds it without re-architecting?
Pick the signal first
Tool by tool
| Need | Tool | What to know |
|---|---|---|
| Built-in service metrics | CloudWatch metrics | Regional. Standard EC2 metrics are 5-minute; detailed monitoring gives 1-minute |
| OS-level metrics (memory, disk used) | CloudWatch agent | EC2 doesn't publish memory or disk-space metrics. The agent does, and also ships log files |
| App metrics from Lambda or containers | Embedded Metric Format (EMF) | Write structured JSON log lines; CloudWatch extracts metrics without API calls |
| Metric from a log pattern | Metric filter | Counts matches going forward only. It doesn't scan logs written before it existed |
| Stream logs to a processor | Subscription filter | Destinations: Lambda, Kinesis Data Streams, Data Firehose, or OpenSearch Service (through a Lambda function) |
| Ad hoc log queries | Logs Insights | Purpose-built query language; queries many log groups at once. Pay per GB scanned |
| Tail logs live | Live Tail | Interactive streaming view while you debug |
| Distributed tracing | X-Ray | Service map, segments, latency breakdown per downstream call |
| APM and SLOs | Application Signals | Auto-instrumented request, latency and error metrics per service, plus SLOs |
| Scripted endpoint checks | Synthetics canaries | Scheduled scripts (Node.js or Python, browser or API) that run on Lambda |
| Real user experience | CloudWatch RUM | JavaScript snippet in the web app reports page loads, errors and sessions |
Alarms that don't page people for nothing
- Static threshold alarms watch one metric or a metric math expression, for example an error rate computed as errors divided by requests.
- Anomaly detection alarms learn a band from history. Use them when "normal" changes by hour or day.
- Composite alarms combine other alarms with
AND,ORandNOT. Page only when high latency and high 5xx are both in alarm, and use action suppression to stay quiet during a known deployment window. - Alarm actions: SNS, Auto Scaling, EC2 actions (stop, reboot, recover), Systems Manager OpsItems or incidents.
Alert fatigue
When a question says the on-call team gets too many alarms from related metrics, the fix is a composite alarm that notifies once. Individual alarms stay, with their own actions turned off.
The missing memory metric
An answer that alarms on the EC2 MemoryUtilization metric without installing anything is wrong. That metric
exists only after the CloudWatch agent publishes it, usually into the CWAgent namespace.
Log pipelines
Subscription filters send matching events, in near real time, to a destination. The destination decides what you can do next.
| Destination | Pick it when |
|---|---|
| Lambda | Light, per-event logic: parse, enrich, alert to Slack or open a ticket |
| Kinesis Data Streams | Several consumers, replay, or cross-account fan-out at high volume |
| Data Firehose | Buffered delivery to S3, OpenSearch, Redshift or third-party tools, with optional transformation. No code to manage |
| OpenSearch Service | Full-text search and dashboards on recent logs |
- Cross-account delivery: create a destination in the central account that fronts a Kinesis stream or Firehose, with a destination policy that allows the source accounts. Source accounts then add subscription filters pointing at it.
- Account-level subscription filters apply one policy to every log group in the account, so new log groups are covered without extra work.
- Log centralization rules in CloudWatch Logs copy log groups from the organization, chosen OUs or accounts, across Regions, into one account. Each event gets its source account and Region as fields.
Exam signal
"Near real time", "analyze as it arrives" or "trigger on a log pattern" point to a subscription filter. "Archive
cheaply and query later" points to Firehose into S3 and Athena. Exporting log groups to S3 with CreateExportTask
is a batch job, not real time.
Seeing across accounts
| Approach | What it gives you |
|---|---|
| CloudWatch cross-account observability | A monitoring account linked to source accounts (through Observability Access Manager sinks and links). Search metrics, logs, traces and Application Signals from the source accounts without copying data. Per Region |
| Cross-account, cross-Region dashboards | Dashboards that show widgets from other accounts and Regions |
| Log centralization rules | Physically copy logs to one account, for retention and a single query point |
| Subscription filters to a central stream | Custom pipelines into S3, OpenSearch or a SIEM |
Link source accounts at the organization or OU level so new accounts join the monitoring account automatically.
Tracing and user-side checks
- X-Ray shows each request's path across API Gateway, Lambda, containers, SQS and databases, and where the latency or errors come from. Instrument with OpenTelemetry (AWS Distro for OpenTelemetry or the CloudWatch agent). The X-Ray SDKs and daemon are in maintenance mode.
- Application Signals builds service dashboards and SLOs from that instrumentation, so you don't hand-build latency and error metrics per service.
- Synthetics canaries catch outages when there's no traffic, for example at 3 a.m. or in a new Region before launch. They can check links, screenshots and API responses, and alarm when they fail.
- RUM tells you what real browsers experience by country, device and page, which server metrics can't show.
Canary or RUM
If the problem only appears for some users (a Region, a browser, a slow network), a canary won't see it. If the site must be checked when nobody is using it, RUM has no data. Match the tool to who generates the signal.
Scenarios
EC2 never publishes guest memory usage, with or without detailed monitoring, so the first option alarms on a metric that doesn't exist. The CloudWatch agent publishes memory metrics that you can alarm on directly. An S3 and Athena job isn't an alarm, and flow logs don't measure memory.
A cross-account destination fronting Firehose, fed by subscription filters in each account, delivers in near real time and writes to both OpenSearch and S3 with no servers. Export tasks are batch and miss the one-minute target, a Fluentd fleet adds servers, and metric filters produce numbers, not log copies.
A composite alarm keeps every underlying alarm and sends one notification based on their combined state. Deleting alarms loses signal, longer evaluation periods delay real alerts, and a custom deduplicator adds code to maintain.
Further reading
Domain 3 · Continuous improvement
25% of the exam. Improving a system that already runs, with the smallest change that fixes the problem.
Automated remediation
Detecting drift, risky changes and security findings in an existing environment and fixing them automatically with Config, EventBridge, Systems Manager Automation and Security Hub.