Compromised workloads
Containing and investigating compromised EC2 instances, EKS pods and Lambda functions, capturing memory and disk evidence, storing it immutably with S3 Object Lock, and automating forensics.
Exam tasks: 2.2.1 (capture and store forensic artifacts), 2.2.4 (contain, eradicate, recover), 2.1.4 (Automated Forensics Orchestrator for Amazon EC2)
The decision: how do you cut the workload off from the attacker and from the rest of your environment without losing the evidence that is still in memory and on disk, and where do you keep that evidence so nobody, including an attacker with admin rights, can alter it?
Order of operations for an EC2 instance
Volatile evidence first, destructive actions last.
- Protect the evidence from automation. Turn on termination protection, set shutdown behaviour to stop, and tag
it (for example
Quarantine=true) so your own tooling and teams leave it alone. - Take it out of service. Detach it from its Auto Scaling group (otherwise a failed health check lets the ASG terminate it and your evidence with it) and deregister it from load balancer target groups. The ASG launches a clean replacement, so the service stays up.
- Isolate the network (next section).
- Invalidate its credentials. Revoke active sessions on the instance role. Detaching the instance profile alone doesn't invalidate credentials already handed out.
- Capture memory, while the instance is still running.
- Snapshot all EBS volumes and copy them to the forensics account.
- Recover by replacing, not cleaning. Terminate the original only when the case is closed.
Stop the instance to be safe
Stopping, rebooting or hibernating an instance throws away running processes, network connections and in-memory malware. Isolation already stops the damage. Capture memory before you consider stopping it.
Network isolation that actually works
- Create an isolation security group with no rules, or only inbound from the forensics workstation's IP, and make it the instance's only security group.
- Security groups are stateful. Tracked connections that already exist keep flowing after you change the security group. Only new traffic is blocked.
- Two fixes:
- First attach a temporary security group that allows all traffic in both directions from
0.0.0.0/0. Flows allowed that way are untracked, so they break the moment you swap to the isolation group. This is what theAWSSupport-ContainEC2Instancerunbook automates (with a Restore mode that puts the old groups back). - Or add network ACL deny rules on the subnet. NACLs are stateless and cut existing connections, but they affect every instance in the subnet.
- First attach a temporary security group that allows all traffic in both directions from
- Don't forget other paths out: the instance role (revoke sessions), VPC endpoints and peering, and attached secondary ENIs.
- Reach the instance through Session Manager, not SSH. It needs no inbound rule and logs every command. The isolation security group still needs outbound HTTPS to the SSM VPC endpoints.
Exam signal
"Existing malicious connections are still active after the security group was changed" means the connections are tracked. The answer is a NACL deny, or the untracked-then-isolate security group sequence.
Capturing evidence
| Artifact | Volatile? | How to capture | Where it goes |
|---|---|---|---|
| Memory | Yes | SSM Run Command running LiME or AVML (Linux) or a memory tool (Windows), writing straight to S3 | Forensics bucket |
| EBS volumes | No | CreateSnapshots for all volumes at once (crash-consistent), then share and copy with a CMK the forensics account can use | Forensics account |
| Instance metadata | Changes | DescribeInstances output: ENIs, IPs, security groups, role, AMI, tags, launch time | Case record |
| Logs on the host | Can be wiped | Pull with SSM, or rely on the CloudWatch agent having shipped them already | Forensics bucket |
| AWS-side logs | No | CloudTrail, VPC Flow Logs, DNS query logs, ELB logs from the log archive account | Already central |
- Hash every artifact (SHA-256) when you collect it and record who handled it and when. That's your chain of custody.
- Snapshots encrypted with the default
aws/ebskey can't be shared. Copy them with a customer managed key whose policy allows the forensics account. - In the forensics account, create volumes from the snapshots and attach them to a forensic instance read-only. Never boot the suspect image.
- GuardDuty Malware Protection for EC2 can scan the instance's EBS volumes without an agent, either automatically when GuardDuty raises certain findings or on demand. It works from snapshots, so it doesn't replace your own evidence copies.
Storing evidence immutably
Put forensic artifacts in an S3 bucket in the forensics account with Object Lock.
| Governance mode | Compliance mode | Legal hold | |
|---|---|---|---|
| Can anyone delete before expiry? | Users with s3:BypassGovernanceRetention | No one, not even root | Until someone with s3:PutObjectLegalHold removes it |
| Can retention be shortened? | With the bypass permission | No (only extended) | No expiry date at all |
| Use for | Testing a policy, or evidence where a senior admin may need an override | Evidence you might need in court or for a regulator | An open investigation or litigation of unknown length |
- Object Lock needs versioning. You can enable it on new or existing buckets, and set a default retention so every uploaded artifact is locked automatically.
- Locks apply to object versions. A delete without a version ID just adds a delete marker; the locked version stays.
- Add a bucket policy that denies access except to the IR roles, SSE-KMS with a dedicated key, and CloudTrail data events on the bucket.
Compliance mode for a trial run
Compliance mode can't be undone: the object (and its storage bill) stays until the retention date. For experiments, use governance mode or a short retention. For real evidence of unknown duration, a legal hold is often better than guessing a retention period.
Containers on EKS
- Isolate the pod with a Kubernetes NetworkPolicy that denies all ingress and egress for a label you add to
the pod (for example
quarantine=true). This needs a network policy engine, such as the VPC CNI's network policy support. - Cordon the node so no new pods land on it, and consider moving other workloads off it. If the pod was privileged or GuardDuty flags node credentials, treat the node as compromised too.
- Revoke the pod's AWS credentials: revoke sessions on the IAM role used by IRSA or EKS Pod Identity, and on the node's instance role if pods could reach it.
- Capture before you delete: container filesystem and process list with
kubectl execor the runtime, memory and a snapshot of the node's volumes, and EKS audit logs in CloudWatch Logs. - Then fix the vulnerability, redeploy clean pods, delete the compromised ones, and remove a malicious image from the registry.
Lambda functions
- Stop invocations by setting the function's reserved concurrency to 0. That throttles every invocation without deleting the function or its versions.
- Revoke sessions on the execution role, and rotate secrets the function could read (environment variables, Secrets Manager, Parameter Store).
- Preserve evidence: download the code package and layers (
GetFunctionreturns a link), the configuration and resource policy, and the CloudWatch Logs. - Look for tampering:
UpdateFunctionCode,UpdateFunctionConfiguration, new layers, andAddPermissioncalls that let other accounts invoke it. - Recover by pointing the alias back at a known-good version, or redeploying from the pipeline.
Recovering
- Rebuild from a known-good source (golden AMI, pipeline, IaC). Cleaning a compromised host in place leaves you guessing about persistence.
- Restore data from backups you can trust, ideally from a vault the attacker couldn't touch (see Backup and ransomware protection).
- Fix the root cause before reconnecting: patch the vulnerability, close the exposed port, require IMDSv2.
- Use ARC zonal shift if a whole AZ's fleet has to be rebuilt.
Automated Forensics Orchestrator for Amazon EC2
An AWS Solutions guidance implementation (now covering EC2 and EKS) that you deploy into a forensics account.
- Trigger: a GuardDuty finding, or a Security Hub custom action an analyst chooses on a finding.
- Step Functions workflows isolate the instance, capture memory (LiME or AVML through SSM), snapshot the volumes and copy everything into the forensics account's S3 bucket.
- An investigation workflow then analyzes the copies (for example Volatility on the memory image) and writes a report, with notifications to the SOC.
- Works across accounts and Regions in an organization.
Exam signal
"Automatically capture memory and disk from a flagged instance across 80 accounts and store it for the SOC, with
the least custom code" points to the Automated Forensics Orchestrator. "Just isolate the instance and let me
restore it later" points to the AWSSupport-ContainEC2Instance runbook.
Scenarios
Detaching removes the instance from the ASG's control and the ALB's traffic, and the ASG launches a replacement so the service stays up; termination protection blocks accidental deletion. Stopping discards memory. A long grace period affects every instance and doesn't stop scale-in. Terminating destroys exactly what you need to investigate.
The session is a tracked connection, which security group changes don't break. NACLs are stateless and cut it immediately. Security groups don't support deny rules. The instance profile controls AWS API access, not network sessions. Rebooting would break the session but also destroy memory evidence.
Compliance mode prevents deletion or shortening of retention by anyone, including the root user, and Object Lock requires versioning. Governance mode can be bypassed by anyone who can change the IAM policy back. Termination protection protects an instance, not the data on its volumes. A lifecycle rule schedules deletion but doesn't prevent an earlier one.
Further reading
Compromised credentials
Responding to exposed access keys, stolen role and instance credentials, compromised Identity Center users and root credentials, tracing activity in CloudTrail, and handling AWS abuse reports.
Domain 3 · Infrastructure security
18% of the exam. Putting the right control at each layer, from the edge through the network to the workload, without opening more access than the requirement needs.