Log analysis
Querying and correlating security logs with CloudWatch Logs Insights, Athena, OpenSearch Service fed by Amazon Data Firehose, and Security Lake, plus parsing and normalizing logs with Lambda and OCSF.
Exam tasks: 1.2 (skills 1.2.4 and 1.2.5: analyze logs, normalize, parse and correlate them)
The decision: where do the logs already live, how far back and how fast must you search, and do you need ad hoc queries, SQL over years of data, or live search and dashboards?
Picking the query engine
| CloudWatch Logs Insights | Athena | OpenSearch Service | Security Lake | |
|---|---|---|---|---|
| Data lives in | CloudWatch Logs | S3 | OpenSearch indexes | S3, OCSF Parquet |
| Query style | Logs Insights QL, OpenSearch PPL or SQL | Standard SQL | Full-text search, PPL, SQL, dashboards | SQL via Athena, or the subscriber's own tool |
| Strength | Zero setup for recent logs | Cheap over years of data | Near-real-time search, visualisation, alerting | One schema across AWS, SaaS and on-premises sources |
| Cost driver | Data scanned per query | Data scanned per query | Cluster or serverless capacity | Ingestion and storage, plus query engine |
| Pick it when | "Quickly find the errors in this function's logs" | "Search 2 years of CloudTrail for this access key" | "SOC wants live search and dashboards over flow and app logs" | "Correlate firewall, CloudTrail and DNS logs in one schema" |
CloudWatch Logs Insights
fields @timestamp, userIdentity.arn, sourceIPAddress, errorCode
| filter eventSource = "s3.amazonaws.com" and errorCode = "AccessDenied"
| stats count(*) as denials by userIdentity.arn, sourceIPAddress
| sort denials desc
| limit 20- Discovers fields automatically for JSON logs and AWS logs such as CloudTrail, VPC flow logs, Route 53 and Lambda.
- Three languages: Logs Insights QL (above), OpenSearch PPL and OpenSearch SQL, which can join across log groups.
- Queries time out after 60 minutes and results are kept for 7 days. Save queries, add them to dashboards, and encrypt results with KMS.
- Field indexes cut the data scanned on large log groups. Pattern analysis groups similar lines, and comparison queries show what changed against an earlier period.
- In a cross-account observability monitoring account, one query can span log groups in linked source accounts.
- It can't see events older than the log group, and it's not built for multi-year history. Keep that in S3.
Athena over S3
- Point Athena at the trail bucket (or flow log, ALB, WAF bucket) with a table definition. Use partition projection on Region, account and date so queries skip irrelevant prefixes instead of relying on crawlers or adding partitions by hand.
- Always filter on the partition columns. A query without a date filter scans the whole bucket and costs accordingly.
- Use workgroups to enforce an encrypted results location, per-query data limits and who can run what.
- Readers need
s3:GetObjecton the logs,kms:Decrypton the log key, and write access to the results bucket. In an organization this is usually a read-only investigator role in the log archive account.
SELECT eventtime, eventname, sourceipaddress, awsregion
FROM org_trail
WHERE useridentity.accesskeyid = 'AKIAEXAMPLEKEY123'
AND day BETWEEN '2025/03/01' AND '2026/08/31'
ORDER BY eventtime;Exam signal
"Query historical logs in S3 with SQL, pay only per query, no infrastructure" is Athena. "The Athena query is slow and expensive" is fixed by partitioning (projection) and columnar formats such as Parquet, not by a bigger engine.
OpenSearch Service with Amazon Data Firehose
- A subscription filter streams matching log events from a log group to Firehose, Kinesis Data Streams, Lambda or OpenSearch. Cross-account delivery goes through a CloudWatch Logs destination in the receiving account.
- Firehose buffers by size and time (up to 900 seconds), can call a Lambda function to parse, enrich or reformat records, and backs up records that fail delivery to S3.
- CloudWatch Logs delivers to subscribers gzip-compressed and base64-encoded, so the transform must decompress before OpenSearch can index it.
- Firehose can deliver to a domain inside a VPC. Its IAM role needs
es:ESHttp*on the domain, and with fine-grained access control the role must also be mapped to an OpenSearch backend role. - OpenSearch Ingestion is the managed pipeline alternative when you need more processing than a Lambda transform. Security Analytics in OpenSearch applies detection rules (Sigma format) to indexed logs.
Real-time with CreateExportTask
Exporting a log group to S3 with an export task is a batch job that can take hours, and only one can run per account at a time. When the question says "near real time", the answer is a subscription filter, not an export task.
Legacy: use Amazon Data Firehose instead
Kinesis Data Firehose was renamed Amazon Data Firehose. The service, its destinations and its APIs are unchanged. Amazon Elasticsearch Service is likewise now Amazon OpenSearch Service.
Normalizing and correlating
- OCSF gives every source the same field names (actor, src_endpoint, api.operation), so one query covers CloudTrail, flow logs and a partner firewall. Security Lake does this conversion for supported AWS sources.
- For custom sources, convert to OCSF Parquet yourself, typically in a Lambda function or a Glue job, before writing to the Security Lake custom source location.
- CloudWatch Logs transformers parse and reshape events at ingestion, including processors that map supported AWS logs to OCSF, so metric filters and queries see structured fields.
- Lambda is the glue for enrichment: add account names from Organizations, geolocate IPs, tag findings.
- For visual correlation across sources and accounts, chart the results in Managed Grafana. See Alerting and dashboards.
Querying Security Lake
- Each source becomes a Glue table in the Security Lake database for its Region, for example the CloudTrail management, VPC flow and Route 53 tables. Query the rollup Region to see all contributing Regions.
- Query access subscribers get Lake Formation grants (shared through resource links) and query with Athena or another Lake Formation-aware engine in their own account. Nothing is copied.
- Data access subscribers are notified of new objects (through SQS or an HTTPS endpoint) and read the Parquet files directly, which suits a third-party SIEM that ingests data.
- Permissions are managed in Lake Formation. A cross-account S3 bucket policy on the lake bucket bypasses the governance model and is the wrong answer.
Scenarios
Athena queries the logs where they already sit in S3, and partition projection limits the scan to the requested dates. Logs Insights only has 30 days here. Loading 18 months into OpenSearch is slow and costly for a one-off question. Detective keeps at most a year of history, so it can't cover 18 months.
Subscription filters stream events as they arrive, Firehose buffers and delivers them to OpenSearch for search, dashboards and alerting, and S3 backup keeps failed records. Export tasks are batch jobs that can take hours. Polling Logs Insights gives no search interface or dashboards and scales badly. A data access subscriber is how a tool reads Security Lake, but Security Lake doesn't ingest arbitrary application logs from CloudWatch Logs.
A query access subscriber receives Lake Formation grants on the tables and queries them in place, which meets both "no copy" and "governed per table". A data access subscriber is designed to pull the objects into its own tool. A raw bucket policy bypasses Lake Formation, and replication copies the data, which compliance forbids.
Further reading
Logging strategy
Choosing log sources for a security question, CloudTrail trails and organization trails, data and network activity events, VPC and transit gateway flow logs, Resolver query logs, service access logs, and a protected log archive account.
Troubleshooting monitoring
Finding why logs, alarms and alerts go missing, from the CloudWatch agent and Lambda or API Gateway logging to trail and bucket policies, KMS key policies on log destinations, and EventBridge rules that never fire.