Asterrr's Handbook

Sensitive data

Finding sensitive data in S3 with Macie, masking it in CloudWatch Logs with data protection policies, SNS message data protection and S3 Object Lambda redaction (and their replacements), and choosing between masking, redaction, tokenization and encryption.

Exam tasks: 5.3.4 (mask sensitive data: CloudWatch Logs data protection policies, SNS message data protection), with discovery supporting 5.2 controls

The decision: where is personal or regulated data hiding, and at which point in its flow (storage, logs, messages, responses) should it be detected, masked, removed or replaced so that only people who need it ever see the real value?

Detect, then protect

Choosing the technique

TechniqueWhat the reader seesReversible?Use when
Masking****-****-****-4417Only for people allowed to unmask (the original is kept)Support staff need to recognise a record, not read it
Redaction[REDACTED] or the field removedNoDownstream systems must never receive the value
Tokenizationtok_9f2c1aYes, through a protected token vaultAnalytics and joins on a value without exposing it; PCI scope reduction
EncryptionCiphertextYes, with the keyStoring and moving the real value between trusted parties
HashingFixed digestNoMatching and de-duplication without storing the value

Amazon Macie

  • Discovers sensitive data in S3 and evaluates bucket security (public access, encryption, sharing). It doesn't scan RDS, DynamoDB or EBS.
  • Automated sensitive data discovery samples objects across all buckets continuously and gives each bucket a sensitivity score, at low cost. Use it for broad visibility.
  • Sensitive data discovery jobs scan chosen buckets fully or on a schedule. Use them for deep, targeted evidence, for example before a migration or an audit.
  • Managed data identifiers detect common types: credentials and private keys, card and bank numbers, passport and national ID numbers, health identifiers.
  • Custom data identifiers: a regex plus optional keywords that must appear near the match, ignore words, and a maximum match distance. Use them for your own formats, such as loyalty numbers like LX- plus eight digits.
  • Allow lists suppress known non-sensitive matches, such as test card numbers or your company address.
Finding typeExampleMeaning
Policy findingsPolicy:IAMUser/S3BlockPublicAccessDisabled, Policy:IAMUser/S3BucketSharedExternallyA bucket's configuration got less secure
Sensitive data findingsSensitiveData:S3Object/Financial, SensitiveData:S3Object/Credentials, SensitiveData:S3Object/CustomIdentifierAn object contains a type of sensitive data
  • Findings go to EventBridge (for automated response) and Security Hub. Macie keeps findings for 90 days.
  • Detailed sensitive data discovery results are written to an S3 bucket you configure, encrypted with a customer managed KMS key. Macie needs permission on that key.
  • In an organization, a delegated administrator account runs Macie for all members.
  • Macie can reveal samples of the sensitive data in a finding for investigators, which requires a KMS key and a separate permission.

Exam signal

"Identify which S3 buckets across all accounts contain PII, with the least effort" is Macie automated sensitive data discovery from a delegated administrator account. "Our employee ID format isn't detected" is a custom data identifier. "False positives on test data" is an allow list.

CloudWatch Logs data protection policies

  • A data protection policy audits and masks sensitive data as it is ingested into a log group. Set one account-wide policy (covers existing and future log groups) and optionally one policy per log group; both apply together.
  • Uses managed data identifiers (credentials, financial, PII, PHI, device identifiers such as IP addresses) and custom data identifiers you define with regex.
  • Masked data stays masked everywhere it leaves the service: the console, Logs Insights, metric filters, subscription filters. Only principals with logs:Unmask can see the original.
  • Audit findings can go to another log group, an S3 bucket or Firehose, and the LogEventsWithFindings metric in AWS/Logs lets you alarm when sensitive data starts appearing.
  • Works on Standard and Infrequent Access log classes.

It only protects what arrives afterwards

A data protection policy doesn't mask events already stored before it was created. If card numbers were logged last month, you still need to delete or re-ingest those events, and you should treat any credentials that were logged as exposed and rotate them.

Masking at the source is still better

Masking in CloudWatch Logs protects readers of the logs. The data still left the application and reached anyone with logs:Unmask. Fixing the logging code so the value is never written remains the stronger control, and the policy is the safety net.

SNS message data protection

Legacy: use a Lambda filter with Amazon Bedrock Guardrails between an inbound and an outbound topic instead

Amazon SNS message data protection stopped accepting new customers on April 30, 2026; existing users can keep it. It let a topic's data protection policy audit, de-identify (mask or redact) or deny messages containing sensitive data, inbound from publishers or outbound per subscriber. AWS now recommends a Lambda function subscribed to an inbound topic that inspects each message with Bedrock Guardrails sensitive information filters (log, block or redact) and republishes to the topic that subscribers use.

Redacting data on the way out

Legacy: use Lambda behind a function URL, API Gateway or CloudFront, or processing in the client instead

S3 Object Lambda has been available only to existing customers (and select partners) since November 7, 2025. It attached a Lambda function to an Object Lambda access point so each GET returned a transformed object, for example with PII detected by Amazon Comprehend and redacted. For new designs, put the same function behind a Lambda function URL, API Gateway or a CloudFront origin, and keep the raw bucket accessible only to that function.

Other building blocks for redaction and detection:

  • Amazon Comprehend DetectPiiEntities finds PII in free text and returns offsets you can redact.
  • AWS Glue sensitive data detection transform finds and masks PII in ETL jobs before data lands in a lake.
  • Amazon Bedrock Guardrails sensitive information filters mask or block PII in prompts and model responses. See Generative AI guardrails.
1 + 1
One account-level and one log group-level data protection policy can apply to a log group.
90 days
How long Macie keeps findings.
April 30, 2026
SNS message data protection closed to new customers.
Nov 7, 2025
S3 Object Lambda closed to new customers.

Scenarios

Scenario
A travel booking company discovers that its payment microservice writes full card numbers into application logs in CloudWatch Logs. Developers will fix the code next quarter. Meanwhile, support engineers must keep reading the logs, but only the fraud team may see full card numbers. What should the security engineer do?
Scenario · choose 2
A bank's security team must find every S3 bucket, across 140 accounts in AWS Organizations, that contains account numbers in the bank's internal format (two letters followed by ten digits, always near the word 'acct'). They want continuous coverage with the least effort. Which TWO actions should they take?
Scenario
An insurer publishes claim events to an SNS topic consumed by several internal teams and a third-party repair network. The repair network must never receive customers' national ID numbers. The insurer's account has never used SNS message data protection. What should the security engineer recommend?

Further reading

On this page