Asterrr's Handbook

Observability

Adding the right CloudWatch metrics, alarms, log pipelines, tracing and synthetic checks to an existing workload, across many accounts.

Exam tasks: 3.1 (determine a strategy to improve operational excellence: monitoring, logging, alerting)

The decision: an existing system has a visibility gap. Which signal is missing (metric, log, trace or user-side check), and what's the smallest change that adds it without re-architecting?

Pick the signal first

Tool by tool

NeedToolWhat to know
Built-in service metricsCloudWatch metricsRegional. Standard EC2 metrics are 5-minute; detailed monitoring gives 1-minute
OS-level metrics (memory, disk used)CloudWatch agentEC2 doesn't publish memory or disk-space metrics. The agent does, and also ships log files
App metrics from Lambda or containersEmbedded Metric Format (EMF)Write structured JSON log lines; CloudWatch extracts metrics without API calls
Metric from a log patternMetric filterCounts matches going forward only. It doesn't scan logs written before it existed
Stream logs to a processorSubscription filterDestinations: Lambda, Kinesis Data Streams, Data Firehose, or OpenSearch Service (through a Lambda function)
Ad hoc log queriesLogs InsightsPurpose-built query language; queries many log groups at once. Pay per GB scanned
Tail logs liveLive TailInteractive streaming view while you debug
Distributed tracingX-RayService map, segments, latency breakdown per downstream call
APM and SLOsApplication SignalsAuto-instrumented request, latency and error metrics per service, plus SLOs
Scripted endpoint checksSynthetics canariesScheduled scripts (Node.js or Python, browser or API) that run on Lambda
Real user experienceCloudWatch RUMJavaScript snippet in the web app reports page loads, errors and sessions
15 months
How long CloudWatch keeps metric data, rolled up to coarser periods as it ages.
1 second
Finest resolution for high-resolution custom metrics and alarms.
2
Subscription filters per log group. Fan out further through Kinesis or Firehose.
Regional
Metrics, alarms and log groups live in one Region. Cross-Region needs dashboards or centralization.

Alarms that don't page people for nothing

  • Static threshold alarms watch one metric or a metric math expression, for example an error rate computed as errors divided by requests.
  • Anomaly detection alarms learn a band from history. Use them when "normal" changes by hour or day.
  • Composite alarms combine other alarms with AND, OR and NOT. Page only when high latency and high 5xx are both in alarm, and use action suppression to stay quiet during a known deployment window.
  • Alarm actions: SNS, Auto Scaling, EC2 actions (stop, reboot, recover), Systems Manager OpsItems or incidents.

Alert fatigue

When a question says the on-call team gets too many alarms from related metrics, the fix is a composite alarm that notifies once. Individual alarms stay, with their own actions turned off.

The missing memory metric

An answer that alarms on the EC2 MemoryUtilization metric without installing anything is wrong. That metric exists only after the CloudWatch agent publishes it, usually into the CWAgent namespace.

Log pipelines

Subscription filters send matching events, in near real time, to a destination. The destination decides what you can do next.

DestinationPick it when
LambdaLight, per-event logic: parse, enrich, alert to Slack or open a ticket
Kinesis Data StreamsSeveral consumers, replay, or cross-account fan-out at high volume
Data FirehoseBuffered delivery to S3, OpenSearch, Redshift or third-party tools, with optional transformation. No code to manage
OpenSearch ServiceFull-text search and dashboards on recent logs
  • Cross-account delivery: create a destination in the central account that fronts a Kinesis stream or Firehose, with a destination policy that allows the source accounts. Source accounts then add subscription filters pointing at it.
  • Account-level subscription filters apply one policy to every log group in the account, so new log groups are covered without extra work.
  • Log centralization rules in CloudWatch Logs copy log groups from the organization, chosen OUs or accounts, across Regions, into one account. Each event gets its source account and Region as fields.

Exam signal

"Near real time", "analyze as it arrives" or "trigger on a log pattern" point to a subscription filter. "Archive cheaply and query later" points to Firehose into S3 and Athena. Exporting log groups to S3 with CreateExportTask is a batch job, not real time.

Seeing across accounts

ApproachWhat it gives you
CloudWatch cross-account observabilityA monitoring account linked to source accounts (through Observability Access Manager sinks and links). Search metrics, logs, traces and Application Signals from the source accounts without copying data. Per Region
Cross-account, cross-Region dashboardsDashboards that show widgets from other accounts and Regions
Log centralization rulesPhysically copy logs to one account, for retention and a single query point
Subscription filters to a central streamCustom pipelines into S3, OpenSearch or a SIEM

Link source accounts at the organization or OU level so new accounts join the monitoring account automatically.

Tracing and user-side checks

  • X-Ray shows each request's path across API Gateway, Lambda, containers, SQS and databases, and where the latency or errors come from. Instrument with OpenTelemetry (AWS Distro for OpenTelemetry or the CloudWatch agent). The X-Ray SDKs and daemon are in maintenance mode.
  • Application Signals builds service dashboards and SLOs from that instrumentation, so you don't hand-build latency and error metrics per service.
  • Synthetics canaries catch outages when there's no traffic, for example at 3 a.m. or in a new Region before launch. They can check links, screenshots and API responses, and alarm when they fail.
  • RUM tells you what real browsers experience by country, device and page, which server metrics can't show.

Canary or RUM

If the problem only appears for some users (a Region, a browser, a slow network), a canary won't see it. If the site must be checked when nobody is using it, RUM has no data. Match the tool to who generates the signal.

Scenarios

Scenario
A logistics company runs a fleet of EC2 instances in an Auto Scaling group. Instances occasionally become unresponsive, and the team suspects memory exhaustion, but CloudWatch shows no memory data. The team wants an alarm when memory use exceeds 90% on any instance, with the LEAST effort. What should the architect do?
Scenario · choose 2
A media company has 40 accounts. Security wants application logs from all accounts delivered within about a minute to a central OpenSearch Service domain, and also archived in S3 for 7 years, with no servers to manage. Which TWO steps should the architect take?
Scenario
An online booking site has CloudWatch alarms on ALB 5xx errors, target response time, and Aurora CPU. During each incident all three fire and the on-call engineer receives three pages for one problem. The team wants one page per incident without losing any of the signals. What is the simplest change?

Further reading

On this page