Troubleshooting monitoring
Finding why logs, alarms and alerts go missing, from the CloudWatch agent and Lambda or API Gateway logging to trail and bucket policies, KMS key policies on log destinations, and EventBridge rules that never fire.
Exam tasks: 1.3 (skills 1.3.1 and 1.3.2: analyze functionality, permissions and configuration of logging resources, and remediate misconfiguration)
The decision: at which hop does the signal disappear (source, delivery permission, encryption, network, Region), and what is the smallest change that restores it?
Walk the pipeline
Most exam answers are one of four fixes: a missing permission (IAM, bucket policy, resource policy), a key policy that doesn't trust the service, no network path, or the wrong Region.
CloudWatch agent
| Symptom | Likely cause | Fix |
|---|---|---|
Agent running, nothing arrives, AccessDenied in the agent log | Instance profile lacks CloudWatch permissions | Attach CloudWatchAgentServerPolicy to the instance role |
| Agent times out connecting | Private subnet with no route to the endpoint | Add a NAT gateway, or interface endpoints for logs (and monitoring for metrics) |
| Some files ship, one doesn't | Wrong file_path in the config, or the agent user can't read the file | Correct the path or glob; grant read access or change run_as_user |
| Config edits have no effect | Agent not reloaded, or it's reading a different config | Run amazon-cloudwatch-agent-ctl -a fetch-config with the file or SSM parameter, then check status |
| Log lines split or timestamps wrong | Multi-line entries or a timestamp format mismatch | Set multi_line_start_pattern and timestamp_format for that file |
| Fleet partly configured | Only some instances got the config | Store the config in SSM Parameter Store and push it with a State Manager association |
- Start with the agent's own log:
/opt/aws/amazon-cloudwatch-agent/logs/amazon-cloudwatch-agent.logon Linux,C:\ProgramData\Amazon\AmazonCloudWatchAgent\Logson Windows. - The instance also needs the SSM Agent and
AmazonSSMManagedInstanceCoreif you install or configure the CloudWatch agent through Systems Manager.
Legacy: use the unified CloudWatch agent instead
The older CloudWatch Logs agent (awslogs) is deprecated. It shipped logs only, not metrics. Questions that
mention editing awslogs.conf describe the old agent; the current answer is the unified agent with its JSON
config.
Service logs that don't appear
| Service | Usual cause | Fix |
|---|---|---|
| Lambda | Execution role lacks logs:CreateLogStream / logs:PutLogEvents, or the chosen custom log group doesn't exist | Attach AWSLambdaBasicExecutionRole or equivalent scoped to the log group |
| API Gateway (REST) | Execution logging on, but no CloudWatch Logs role ARN in the API Gateway account settings, or stage log level OFF | Set the account-level role (with AmazonAPIGatewayPushToCloudWatchLogs) once per Region, then set the stage log level |
| CloudFront (legacy S3 logs) | Log bucket has Object Ownership bucket owner enforced, so ACLs are disabled | Enable ACLs on the bucket, or move to standard logging v2 (CloudWatch Logs, Firehose or S3) |
| ALB | Bucket policy doesn't allow the ELB log delivery principal, or bucket uses SSE-KMS | Add the Region's delivery principal to the bucket policy, use SSE-S3 |
| VPC Flow Logs to CloudWatch Logs | IAM role not trusted by vpc-flow-logs.amazonaws.com or missing logs:PutLogEvents | Fix the role trust and permissions, then recreate the flow log |
| VPC Flow Logs, Resolver logs, WAF to S3 | Bucket policy missing the delivery.logs.amazonaws.com statements, often after someone replaced the policy | Restore the log delivery statements with aws:SourceAccount and aws:SourceArn conditions |
| S3 server access logs | Target bucket in another Region or account, or it uses SSE-KMS by default | Use a same-Region, same-account bucket with SSE-S3 |
| EKS control plane | Log types not enabled on the cluster | Enable audit and authenticator (and others as needed) in cluster logging |
Health checks (skill 1.3.1) fail for configuration reasons too:
- Route 53 health checks need the target's security group and NACL to allow the Route 53 health checker IP ranges. Blocked checkers look like an outage.
- ALB health checks fail when the health check path returns a redirect or
403. Point them at a lightweight path that returns200without authentication.
CloudTrail and its bucket
Check aws cloudtrail get-trail-status first: IsLogging and LatestDeliveryError usually name the problem.
- Bucket policy: must allow
cloudtrail.amazonaws.comtos3:GetBucketAclon the bucket ands3:PutObjectonAWSLogs/..., conditioned on thebucket-owner-full-controlcanned ACL. An organization trail needs the path with the organization ID. Anaws:SourceArncondition that names the wrong trail ARN or Region blocks delivery. - Organization trails are created in member accounts even if validation fails. Members can see the failure in the trail status, but only the management or delegated admin account can fix it.
- CloudWatch Logs delivery from a trail needs a role CloudTrail can assume with
logs:CreateLogStreamandlogs:PutLogEventson that log group. - SCPs don't restrict service principals such as
cloudtrail.amazonaws.com, so they're rarely why delivery fails. They matter the other way: denycloudtrail:StopLoggingandcloudtrail:DeleteTrailso a stopped trail doesn't happen again.
KMS on log destinations
| Destination | What the key policy needs |
|---|---|
| CloudTrail to S3 with SSE-KMS | Allow cloudtrail.amazonaws.com kms:GenerateDataKey* with the aws:cloudtrail:arn encryption context (and aws:SourceArn). Readers need kms:Decrypt |
| CloudWatch Logs log group | Allow logs.<region>.amazonaws.com Encrypt, Decrypt, ReEncrypt, GenerateDataKey and Describe, conditioned on kms:EncryptionContext:aws:logs:arn for the log group |
| Encrypted SNS topic as an alarm or EventBridge target | Allow cloudwatch.amazonaws.com or events.amazonaws.com kms:GenerateDataKey* and kms:Decrypt. Use a customer managed key, since aws/sns can't be edited |
| Macie discovery results, GuardDuty export | Allow the service principal to use the key, and the key must be in the same Region as the bucket |
- If a customer managed key is disabled or scheduled for deletion, services stop writing and readers
get
KMSInvalidStateExceptionor access errors. Re-enable it or cancel the deletion. - The key and the log group or bucket must be in the same Region. KMS keys don't work cross-Region.
The alarm that fires but nobody hears
CloudWatch alarm history shows ALARM, but no email arrives. If the SNS topic is encrypted with the AWS managed
aws/sns key, CloudWatch can't publish to it, because you can't add the CloudWatch service principal to that
key's policy. Switch the topic to a customer managed key whose policy allows cloudwatch.amazonaws.com. Also
check that the email subscription was confirmed.
Alarms and EventBridge rules that never fire
Alarms
- The metric filter pattern doesn't match the real log format (for example
$.errorCodeversus a text log), or the filter was created after the events. - The alarm watches the wrong namespace, metric name or dimension, so it sees no data.
- Missing data treated as
missingleaves the alarm inINSUFFICIENT_DATA. UsenotBreachingfor event counts. - Metrics that only exist in us-east-1 (Route 53 health checks, billing) need the alarm in us-east-1.
EventBridge
- Wrong Region: GuardDuty, Security Hub CSPM and service events go to the default bus in their own Region. IAM and other global service API calls arrive in us-east-1.
- Pattern mismatch: values are case-sensitive and arrays match any element. Test the pattern against a real
sample event with the sandbox or
TestEventPattern. - No trail:
AWS API Call via CloudTrailevents need an active trail. Read-only calls need the rule stateENABLED_WITH_ALL_CLOUDTRAIL_MANAGEMENT_EVENTS. - Target permission: Lambda needs a resource-based policy allowing
events.amazonaws.com; SNS needs a topic policy; Step Functions, SSM Automation and cross-account buses need an IAM role or bus resource policy. - Look at the rule's
MatchedEventsandFailedInvocationsmetrics: matched but failed means target permissions; not matched means pattern or Region. Add a DLQ to capture what failed.
Exam signal
If MatchedEvents rises but the Lambda never runs, the fix is the Lambda resource-based policy for
events.amazonaws.com, not the Lambda execution role. The execution role controls what the function can do, not
who can invoke it.
Findings missing from the admin account
- The member isn't associated with the delegated admin in that Region, or auto-enable was
NEWand the account existed before. - A suppression rule archived the findings, so they never reached EventBridge or Security Hub.
- Security Hub CSPM shows nothing for other Regions because no aggregation Region is set, or controls show no data because AWS Config isn't recording.
- GuardDuty sends updates of existing findings every 6 hours by default. A changed finding may take that long to update downstream unless you shorten the frequency.
Scenarios
Timeouts mean the agent can't reach the CloudWatch Logs endpoint. An interface endpoint (or a NAT gateway) gives private subnets that path, and its security group must allow HTTPS from the instances. The permissions are already correct, and granting admin breaks least privilege without fixing a network problem. The old agent uses the same endpoint. Log group encryption is performed by CloudWatch Logs, not the instance.
CloudWatch Logs itself performs encryption for a log group, so the key policy must allow the logs.region.amazonaws.com principal, usually with an encryption context condition naming the log group. CloudTrail delivers to encrypted log groups normally. The trail's role only needs Logs permissions. Nothing restricts the delivery log group to the management account.
Matched plus failed invocations means the pattern is fine and the target call is rejected, which for Lambda is the missing resource-based permission for EventBridge. IAM is a global service whose events arrive in us-east-1, so the rule's Region is correct. CreateAccessKey is a write event, so the default rule state already matches it. The execution role governs what the function may do, not whether EventBridge may invoke it.
Further reading
Log analysis
Querying and correlating security logs with CloudWatch Logs Insights, Athena, OpenSearch Service fed by Amazon Data Firehose, and Security Lake, plus parsing and normalizing logs with Lambda and OCSF.
Domain 2 · Incident response
14% of the exam. Being ready before something goes wrong, then containing, investigating and recovering without destroying the evidence.