Asterrr's Handbook

Log analysis

Querying and correlating security logs with CloudWatch Logs Insights, Athena, OpenSearch Service fed by Amazon Data Firehose, and Security Lake, plus parsing and normalizing logs with Lambda and OCSF.

Exam tasks: 1.2 (skills 1.2.4 and 1.2.5: analyze logs, normalize, parse and correlate them)

The decision: where do the logs already live, how far back and how fast must you search, and do you need ad hoc queries, SQL over years of data, or live search and dashboards?

Picking the query engine

CloudWatch Logs InsightsAthenaOpenSearch ServiceSecurity Lake
Data lives inCloudWatch LogsS3OpenSearch indexesS3, OCSF Parquet
Query styleLogs Insights QL, OpenSearch PPL or SQLStandard SQLFull-text search, PPL, SQL, dashboardsSQL via Athena, or the subscriber's own tool
StrengthZero setup for recent logsCheap over years of dataNear-real-time search, visualisation, alertingOne schema across AWS, SaaS and on-premises sources
Cost driverData scanned per queryData scanned per queryCluster or serverless capacityIngestion and storage, plus query engine
Pick it when"Quickly find the errors in this function's logs""Search 2 years of CloudTrail for this access key""SOC wants live search and dashboards over flow and app logs""Correlate firewall, CloudTrail and DNS logs in one schema"

CloudWatch Logs Insights

fields @timestamp, userIdentity.arn, sourceIPAddress, errorCode
| filter eventSource = "s3.amazonaws.com" and errorCode = "AccessDenied"
| stats count(*) as denials by userIdentity.arn, sourceIPAddress
| sort denials desc
| limit 20
  • Discovers fields automatically for JSON logs and AWS logs such as CloudTrail, VPC flow logs, Route 53 and Lambda.
  • Three languages: Logs Insights QL (above), OpenSearch PPL and OpenSearch SQL, which can join across log groups.
  • Queries time out after 60 minutes and results are kept for 7 days. Save queries, add them to dashboards, and encrypt results with KMS.
  • Field indexes cut the data scanned on large log groups. Pattern analysis groups similar lines, and comparison queries show what changed against an earlier period.
  • In a cross-account observability monitoring account, one query can span log groups in linked source accounts.
  • It can't see events older than the log group, and it's not built for multi-year history. Keep that in S3.

Athena over S3

  • Point Athena at the trail bucket (or flow log, ALB, WAF bucket) with a table definition. Use partition projection on Region, account and date so queries skip irrelevant prefixes instead of relying on crawlers or adding partitions by hand.
  • Always filter on the partition columns. A query without a date filter scans the whole bucket and costs accordingly.
  • Use workgroups to enforce an encrypted results location, per-query data limits and who can run what.
  • Readers need s3:GetObject on the logs, kms:Decrypt on the log key, and write access to the results bucket. In an organization this is usually a read-only investigator role in the log archive account.
SELECT eventtime, eventname, sourceipaddress, awsregion
FROM org_trail
WHERE useridentity.accesskeyid = 'AKIAEXAMPLEKEY123'
  AND day BETWEEN '2025/03/01' AND '2026/08/31'
ORDER BY eventtime;

Exam signal

"Query historical logs in S3 with SQL, pay only per query, no infrastructure" is Athena. "The Athena query is slow and expensive" is fixed by partitioning (projection) and columnar formats such as Parquet, not by a bigger engine.

OpenSearch Service with Amazon Data Firehose

  • A subscription filter streams matching log events from a log group to Firehose, Kinesis Data Streams, Lambda or OpenSearch. Cross-account delivery goes through a CloudWatch Logs destination in the receiving account.
  • Firehose buffers by size and time (up to 900 seconds), can call a Lambda function to parse, enrich or reformat records, and backs up records that fail delivery to S3.
  • CloudWatch Logs delivers to subscribers gzip-compressed and base64-encoded, so the transform must decompress before OpenSearch can index it.
  • Firehose can deliver to a domain inside a VPC. Its IAM role needs es:ESHttp* on the domain, and with fine-grained access control the role must also be mapped to an OpenSearch backend role.
  • OpenSearch Ingestion is the managed pipeline alternative when you need more processing than a Lambda transform. Security Analytics in OpenSearch applies detection rules (Sigma format) to indexed logs.

Real-time with CreateExportTask

Exporting a log group to S3 with an export task is a batch job that can take hours, and only one can run per account at a time. When the question says "near real time", the answer is a subscription filter, not an export task.

Legacy: use Amazon Data Firehose instead

Kinesis Data Firehose was renamed Amazon Data Firehose. The service, its destinations and its APIs are unchanged. Amazon Elasticsearch Service is likewise now Amazon OpenSearch Service.

Normalizing and correlating

  • OCSF gives every source the same field names (actor, src_endpoint, api.operation), so one query covers CloudTrail, flow logs and a partner firewall. Security Lake does this conversion for supported AWS sources.
  • For custom sources, convert to OCSF Parquet yourself, typically in a Lambda function or a Glue job, before writing to the Security Lake custom source location.
  • CloudWatch Logs transformers parse and reshape events at ingestion, including processors that map supported AWS logs to OCSF, so metric filters and queries see structured fields.
  • Lambda is the glue for enrichment: add account names from Organizations, geolocate IPs, tag findings.
  • For visual correlation across sources and accounts, chart the results in Managed Grafana. See Alerting and dashboards.

Querying Security Lake

  • Each source becomes a Glue table in the Security Lake database for its Region, for example the CloudTrail management, VPC flow and Route 53 tables. Query the rollup Region to see all contributing Regions.
  • Query access subscribers get Lake Formation grants (shared through resource links) and query with Athena or another Lake Formation-aware engine in their own account. Nothing is copied.
  • Data access subscribers are notified of new objects (through SQS or an HTTPS endpoint) and read the Parquet files directly, which suits a third-party SIEM that ingests data.
  • Permissions are managed in Lake Formation. A cross-account S3 bucket policy on the lake bucket bypasses the governance model and is the wrong answer.
60 min
Timeout for a CloudWatch Logs Insights query.
7 days
How long Logs Insights query results are available.
900 s
Longest Firehose buffering interval.
1 year
Maximum history Detective keeps in its behaviour graph.

Scenarios

Scenario
During an incident review, a security engineer must list every API call made over the past 18 months with one specific access key, across 70 accounts. The organization trail writes to a bucket in the log archive account, and only the last 30 days are also sent to CloudWatch Logs. What is the MOST cost-effective way to answer the question?
Scenario
A SOC wants full-text search, dashboards and alerting over VPC flow logs and application logs from CloudWatch Logs, available within about a minute of the event. Records that fail indexing must not be lost. Which architecture should the security engineer build?
Scenario
Threat hunters in a separate analytics account need to run SQL against Security Lake data held in the security account. Compliance forbids copying the data out of the security account, and access must be governed per table. What should the security engineer do?

Further reading

On this page