Sensitive data
Finding sensitive data in S3 with Macie, masking it in CloudWatch Logs with data protection policies, SNS message data protection and S3 Object Lambda redaction (and their replacements), and choosing between masking, redaction, tokenization and encryption.
Exam tasks: 5.3.4 (mask sensitive data: CloudWatch Logs data protection policies, SNS message data protection), with discovery supporting 5.2 controls
The decision: where is personal or regulated data hiding, and at which point in its flow (storage, logs, messages, responses) should it be detected, masked, removed or replaced so that only people who need it ever see the real value?
Detect, then protect
Choosing the technique
| Technique | What the reader sees | Reversible? | Use when |
|---|---|---|---|
| Masking | ****-****-****-4417 | Only for people allowed to unmask (the original is kept) | Support staff need to recognise a record, not read it |
| Redaction | [REDACTED] or the field removed | No | Downstream systems must never receive the value |
| Tokenization | tok_9f2c1a | Yes, through a protected token vault | Analytics and joins on a value without exposing it; PCI scope reduction |
| Encryption | Ciphertext | Yes, with the key | Storing and moving the real value between trusted parties |
| Hashing | Fixed digest | No | Matching and de-duplication without storing the value |
Amazon Macie
- Discovers sensitive data in S3 and evaluates bucket security (public access, encryption, sharing). It doesn't scan RDS, DynamoDB or EBS.
- Automated sensitive data discovery samples objects across all buckets continuously and gives each bucket a sensitivity score, at low cost. Use it for broad visibility.
- Sensitive data discovery jobs scan chosen buckets fully or on a schedule. Use them for deep, targeted evidence, for example before a migration or an audit.
- Managed data identifiers detect common types: credentials and private keys, card and bank numbers, passport and national ID numbers, health identifiers.
- Custom data identifiers: a regex plus optional keywords that must appear near the match, ignore
words, and a maximum match distance. Use them for your own formats, such as loyalty numbers like
LX-plus eight digits. - Allow lists suppress known non-sensitive matches, such as test card numbers or your company address.
| Finding type | Example | Meaning |
|---|---|---|
| Policy findings | Policy:IAMUser/S3BlockPublicAccessDisabled, Policy:IAMUser/S3BucketSharedExternally | A bucket's configuration got less secure |
| Sensitive data findings | SensitiveData:S3Object/Financial, SensitiveData:S3Object/Credentials, SensitiveData:S3Object/CustomIdentifier | An object contains a type of sensitive data |
- Findings go to EventBridge (for automated response) and Security Hub. Macie keeps findings for 90 days.
- Detailed sensitive data discovery results are written to an S3 bucket you configure, encrypted with a customer managed KMS key. Macie needs permission on that key.
- In an organization, a delegated administrator account runs Macie for all members.
- Macie can reveal samples of the sensitive data in a finding for investigators, which requires a KMS key and a separate permission.
Exam signal
"Identify which S3 buckets across all accounts contain PII, with the least effort" is Macie automated sensitive data discovery from a delegated administrator account. "Our employee ID format isn't detected" is a custom data identifier. "False positives on test data" is an allow list.
CloudWatch Logs data protection policies
- A data protection policy audits and masks sensitive data as it is ingested into a log group. Set one account-wide policy (covers existing and future log groups) and optionally one policy per log group; both apply together.
- Uses managed data identifiers (credentials, financial, PII, PHI, device identifiers such as IP addresses) and custom data identifiers you define with regex.
- Masked data stays masked everywhere it leaves the service: the console, Logs Insights, metric filters,
subscription filters. Only principals with
logs:Unmaskcan see the original. - Audit findings can go to another log group, an S3 bucket or Firehose, and the
LogEventsWithFindingsmetric inAWS/Logslets you alarm when sensitive data starts appearing. - Works on Standard and Infrequent Access log classes.
It only protects what arrives afterwards
A data protection policy doesn't mask events already stored before it was created. If card numbers were logged last month, you still need to delete or re-ingest those events, and you should treat any credentials that were logged as exposed and rotate them.
Masking at the source is still better
Masking in CloudWatch Logs protects readers of the logs. The data still left the application and reached
anyone with logs:Unmask. Fixing the logging code so the value is never written remains the stronger control,
and the policy is the safety net.
SNS message data protection
Legacy: use a Lambda filter with Amazon Bedrock Guardrails between an inbound and an outbound topic instead
Amazon SNS message data protection stopped accepting new customers on April 30, 2026; existing users can keep it. It let a topic's data protection policy audit, de-identify (mask or redact) or deny messages containing sensitive data, inbound from publishers or outbound per subscriber. AWS now recommends a Lambda function subscribed to an inbound topic that inspects each message with Bedrock Guardrails sensitive information filters (log, block or redact) and republishes to the topic that subscribers use.
Redacting data on the way out
Legacy: use Lambda behind a function URL, API Gateway or CloudFront, or processing in the client instead
S3 Object Lambda has been available only to existing customers (and select partners) since November 7, 2025. It attached a Lambda function to an Object Lambda access point so each GET returned a transformed object, for example with PII detected by Amazon Comprehend and redacted. For new designs, put the same function behind a Lambda function URL, API Gateway or a CloudFront origin, and keep the raw bucket accessible only to that function.
Other building blocks for redaction and detection:
- Amazon Comprehend
DetectPiiEntitiesfinds PII in free text and returns offsets you can redact. - AWS Glue sensitive data detection transform finds and masks PII in ETL jobs before data lands in a lake.
- Amazon Bedrock Guardrails sensitive information filters mask or block PII in prompts and model responses. See Generative AI guardrails.
Scenarios
A data protection policy masks card numbers at ingestion for everyone without logs:Unmask, while the fraud team can still unmask. A subscription filter copies logs elsewhere but leaves the originals unmasked. KMS encryption of the log group is all or nothing: support engineers would lose access to the logs entirely. Macie scans S3, not CloudWatch Logs.
A delegated administrator gives organization-wide coverage, automated discovery keeps it continuous, and a custom data identifier with a keyword detects the bank's format with fewer false positives. Inventory reports list object metadata, not content. GuardDuty detects threats to S3, not the data inside. A custom Lambda scanner is heavy to build and run.
SNS message data protection isn't available to new customers, and AWS's recommended replacement is a Lambda filter with Bedrock Guardrails between two topics. Topic encryption is at rest inside SNS and doesn't stop delivery of plaintext to subscribers. Filter policies match message attributes and would drop whole claims rather than redacting a field.
Further reading
Secrets management
Secrets Manager versus Parameter Store SecureString, rotation with Lambda and managed rotation, single-user and alternating-user strategies, cross-account secrets, caching, and keeping secrets out of code, user data and environment variables.
Domain 6 · Security foundations and governance
14% of the exam. Building the account structure, deployment guardrails and compliance checks that every other security control sits on.