Asterrr's Handbook

Troubleshooting monitoring

Finding why logs, alarms and alerts go missing, from the CloudWatch agent and Lambda or API Gateway logging to trail and bucket policies, KMS key policies on log destinations, and EventBridge rules that never fire.

Exam tasks: 1.3 (skills 1.3.1 and 1.3.2: analyze functionality, permissions and configuration of logging resources, and remediate misconfiguration)

The decision: at which hop does the signal disappear (source, delivery permission, encryption, network, Region), and what is the smallest change that restores it?

Walk the pipeline

Most exam answers are one of four fixes: a missing permission (IAM, bucket policy, resource policy), a key policy that doesn't trust the service, no network path, or the wrong Region.

CloudWatch agent

SymptomLikely causeFix
Agent running, nothing arrives, AccessDenied in the agent logInstance profile lacks CloudWatch permissionsAttach CloudWatchAgentServerPolicy to the instance role
Agent times out connectingPrivate subnet with no route to the endpointAdd a NAT gateway, or interface endpoints for logs (and monitoring for metrics)
Some files ship, one doesn'tWrong file_path in the config, or the agent user can't read the fileCorrect the path or glob; grant read access or change run_as_user
Config edits have no effectAgent not reloaded, or it's reading a different configRun amazon-cloudwatch-agent-ctl -a fetch-config with the file or SSM parameter, then check status
Log lines split or timestamps wrongMulti-line entries or a timestamp format mismatchSet multi_line_start_pattern and timestamp_format for that file
Fleet partly configuredOnly some instances got the configStore the config in SSM Parameter Store and push it with a State Manager association
  • Start with the agent's own log: /opt/aws/amazon-cloudwatch-agent/logs/amazon-cloudwatch-agent.log on Linux, C:\ProgramData\Amazon\AmazonCloudWatchAgent\Logs on Windows.
  • The instance also needs the SSM Agent and AmazonSSMManagedInstanceCore if you install or configure the CloudWatch agent through Systems Manager.

Legacy: use the unified CloudWatch agent instead

The older CloudWatch Logs agent (awslogs) is deprecated. It shipped logs only, not metrics. Questions that mention editing awslogs.conf describe the old agent; the current answer is the unified agent with its JSON config.

Service logs that don't appear

ServiceUsual causeFix
LambdaExecution role lacks logs:CreateLogStream / logs:PutLogEvents, or the chosen custom log group doesn't existAttach AWSLambdaBasicExecutionRole or equivalent scoped to the log group
API Gateway (REST)Execution logging on, but no CloudWatch Logs role ARN in the API Gateway account settings, or stage log level OFFSet the account-level role (with AmazonAPIGatewayPushToCloudWatchLogs) once per Region, then set the stage log level
CloudFront (legacy S3 logs)Log bucket has Object Ownership bucket owner enforced, so ACLs are disabledEnable ACLs on the bucket, or move to standard logging v2 (CloudWatch Logs, Firehose or S3)
ALBBucket policy doesn't allow the ELB log delivery principal, or bucket uses SSE-KMSAdd the Region's delivery principal to the bucket policy, use SSE-S3
VPC Flow Logs to CloudWatch LogsIAM role not trusted by vpc-flow-logs.amazonaws.com or missing logs:PutLogEventsFix the role trust and permissions, then recreate the flow log
VPC Flow Logs, Resolver logs, WAF to S3Bucket policy missing the delivery.logs.amazonaws.com statements, often after someone replaced the policyRestore the log delivery statements with aws:SourceAccount and aws:SourceArn conditions
S3 server access logsTarget bucket in another Region or account, or it uses SSE-KMS by defaultUse a same-Region, same-account bucket with SSE-S3
EKS control planeLog types not enabled on the clusterEnable audit and authenticator (and others as needed) in cluster logging

Health checks (skill 1.3.1) fail for configuration reasons too:

  • Route 53 health checks need the target's security group and NACL to allow the Route 53 health checker IP ranges. Blocked checkers look like an outage.
  • ALB health checks fail when the health check path returns a redirect or 403. Point them at a lightweight path that returns 200 without authentication.

CloudTrail and its bucket

Check aws cloudtrail get-trail-status first: IsLogging and LatestDeliveryError usually name the problem.

  • Bucket policy: must allow cloudtrail.amazonaws.com to s3:GetBucketAcl on the bucket and s3:PutObject on AWSLogs/..., conditioned on the bucket-owner-full-control canned ACL. An organization trail needs the path with the organization ID. An aws:SourceArn condition that names the wrong trail ARN or Region blocks delivery.
  • Organization trails are created in member accounts even if validation fails. Members can see the failure in the trail status, but only the management or delegated admin account can fix it.
  • CloudWatch Logs delivery from a trail needs a role CloudTrail can assume with logs:CreateLogStream and logs:PutLogEvents on that log group.
  • SCPs don't restrict service principals such as cloudtrail.amazonaws.com, so they're rarely why delivery fails. They matter the other way: deny cloudtrail:StopLogging and cloudtrail:DeleteTrail so a stopped trail doesn't happen again.

KMS on log destinations

DestinationWhat the key policy needs
CloudTrail to S3 with SSE-KMSAllow cloudtrail.amazonaws.com kms:GenerateDataKey* with the aws:cloudtrail:arn encryption context (and aws:SourceArn). Readers need kms:Decrypt
CloudWatch Logs log groupAllow logs.<region>.amazonaws.com Encrypt, Decrypt, ReEncrypt, GenerateDataKey and Describe, conditioned on kms:EncryptionContext:aws:logs:arn for the log group
Encrypted SNS topic as an alarm or EventBridge targetAllow cloudwatch.amazonaws.com or events.amazonaws.com kms:GenerateDataKey* and kms:Decrypt. Use a customer managed key, since aws/sns can't be edited
Macie discovery results, GuardDuty exportAllow the service principal to use the key, and the key must be in the same Region as the bucket
  • If a customer managed key is disabled or scheduled for deletion, services stop writing and readers get KMSInvalidStateException or access errors. Re-enable it or cancel the deletion.
  • The key and the log group or bucket must be in the same Region. KMS keys don't work cross-Region.

The alarm that fires but nobody hears

CloudWatch alarm history shows ALARM, but no email arrives. If the SNS topic is encrypted with the AWS managed aws/sns key, CloudWatch can't publish to it, because you can't add the CloudWatch service principal to that key's policy. Switch the topic to a customer managed key whose policy allows cloudwatch.amazonaws.com. Also check that the email subscription was confirmed.

Alarms and EventBridge rules that never fire

Alarms

  • The metric filter pattern doesn't match the real log format (for example $.errorCode versus a text log), or the filter was created after the events.
  • The alarm watches the wrong namespace, metric name or dimension, so it sees no data.
  • Missing data treated as missing leaves the alarm in INSUFFICIENT_DATA. Use notBreaching for event counts.
  • Metrics that only exist in us-east-1 (Route 53 health checks, billing) need the alarm in us-east-1.

EventBridge

  • Wrong Region: GuardDuty, Security Hub CSPM and service events go to the default bus in their own Region. IAM and other global service API calls arrive in us-east-1.
  • Pattern mismatch: values are case-sensitive and arrays match any element. Test the pattern against a real sample event with the sandbox or TestEventPattern.
  • No trail: AWS API Call via CloudTrail events need an active trail. Read-only calls need the rule state ENABLED_WITH_ALL_CLOUDTRAIL_MANAGEMENT_EVENTS.
  • Target permission: Lambda needs a resource-based policy allowing events.amazonaws.com; SNS needs a topic policy; Step Functions, SSM Automation and cross-account buses need an IAM role or bus resource policy.
  • Look at the rule's MatchedEvents and FailedInvocations metrics: matched but failed means target permissions; not matched means pattern or Region. Add a DLQ to capture what failed.

Exam signal

If MatchedEvents rises but the Lambda never runs, the fix is the Lambda resource-based policy for events.amazonaws.com, not the Lambda execution role. The execution role controls what the function can do, not who can invoke it.

Findings missing from the admin account

  • The member isn't associated with the delegated admin in that Region, or auto-enable was NEW and the account existed before.
  • A suppression rule archived the findings, so they never reached EventBridge or Security Hub.
  • Security Hub CSPM shows nothing for other Regions because no aggregation Region is set, or controls show no data because AWS Config isn't recording.
  • GuardDuty sends updates of existing findings every 6 hours by default. A changed finding may take that long to update downstream unless you shorten the frequency.
CloudWatchAgentServerPolicy
Managed policy the instance role needs for the CloudWatch agent.
logs.region.amazonaws.com
Service principal to allow in a key policy for an encrypted log group.
delivery.logs.amazonaws.com
Principal that delivers flow logs, Resolver logs and WAF logs to S3.
us-east-1
Region where global service API events reach EventBridge.

Scenarios

Scenario
An EC2 fleet in private subnets without internet access runs the unified CloudWatch agent. The instance role has CloudWatchAgentServerPolicy attached and the configuration was pushed from SSM Parameter Store, but no application logs reach CloudWatch Logs. The agent log shows connection timeouts. What should the security engineer do?
Scenario
A security engineer encrypted the organization trail's log group in CloudWatch Logs with a new customer managed KMS key. Since then, the log group receives no events, although the S3 copy of the trail is still delivered. What is the MOST likely cause?
Scenario
An EventBridge rule in us-east-1 should invoke a Lambda function whenever an IAM access key is created in any Region of the account. The rule's MatchedEvents metric increases after each CreateAccessKey call, but the function never runs and FailedInvocations increases too. What should the engineer fix?

Further reading

On this page