Asterrr's Handbook

Compromised workloads

Containing and investigating compromised EC2 instances, EKS pods and Lambda functions, capturing memory and disk evidence, storing it immutably with S3 Object Lock, and automating forensics.

Exam tasks: 2.2.1 (capture and store forensic artifacts), 2.2.4 (contain, eradicate, recover), 2.1.4 (Automated Forensics Orchestrator for Amazon EC2)

The decision: how do you cut the workload off from the attacker and from the rest of your environment without losing the evidence that is still in memory and on disk, and where do you keep that evidence so nobody, including an attacker with admin rights, can alter it?

Order of operations for an EC2 instance

Volatile evidence first, destructive actions last.

  1. Protect the evidence from automation. Turn on termination protection, set shutdown behaviour to stop, and tag it (for example Quarantine=true) so your own tooling and teams leave it alone.
  2. Take it out of service. Detach it from its Auto Scaling group (otherwise a failed health check lets the ASG terminate it and your evidence with it) and deregister it from load balancer target groups. The ASG launches a clean replacement, so the service stays up.
  3. Isolate the network (next section).
  4. Invalidate its credentials. Revoke active sessions on the instance role. Detaching the instance profile alone doesn't invalidate credentials already handed out.
  5. Capture memory, while the instance is still running.
  6. Snapshot all EBS volumes and copy them to the forensics account.
  7. Recover by replacing, not cleaning. Terminate the original only when the case is closed.

Stop the instance to be safe

Stopping, rebooting or hibernating an instance throws away running processes, network connections and in-memory malware. Isolation already stops the damage. Capture memory before you consider stopping it.

Network isolation that actually works

  • Create an isolation security group with no rules, or only inbound from the forensics workstation's IP, and make it the instance's only security group.
  • Security groups are stateful. Tracked connections that already exist keep flowing after you change the security group. Only new traffic is blocked.
  • Two fixes:
    • First attach a temporary security group that allows all traffic in both directions from 0.0.0.0/0. Flows allowed that way are untracked, so they break the moment you swap to the isolation group. This is what the AWSSupport-ContainEC2Instance runbook automates (with a Restore mode that puts the old groups back).
    • Or add network ACL deny rules on the subnet. NACLs are stateless and cut existing connections, but they affect every instance in the subnet.
  • Don't forget other paths out: the instance role (revoke sessions), VPC endpoints and peering, and attached secondary ENIs.
  • Reach the instance through Session Manager, not SSH. It needs no inbound rule and logs every command. The isolation security group still needs outbound HTTPS to the SSM VPC endpoints.

Exam signal

"Existing malicious connections are still active after the security group was changed" means the connections are tracked. The answer is a NACL deny, or the untracked-then-isolate security group sequence.

Capturing evidence

ArtifactVolatile?How to captureWhere it goes
MemoryYesSSM Run Command running LiME or AVML (Linux) or a memory tool (Windows), writing straight to S3Forensics bucket
EBS volumesNoCreateSnapshots for all volumes at once (crash-consistent), then share and copy with a CMK the forensics account can useForensics account
Instance metadataChangesDescribeInstances output: ENIs, IPs, security groups, role, AMI, tags, launch timeCase record
Logs on the hostCan be wipedPull with SSM, or rely on the CloudWatch agent having shipped them alreadyForensics bucket
AWS-side logsNoCloudTrail, VPC Flow Logs, DNS query logs, ELB logs from the log archive accountAlready central
  • Hash every artifact (SHA-256) when you collect it and record who handled it and when. That's your chain of custody.
  • Snapshots encrypted with the default aws/ebs key can't be shared. Copy them with a customer managed key whose policy allows the forensics account.
  • In the forensics account, create volumes from the snapshots and attach them to a forensic instance read-only. Never boot the suspect image.
  • GuardDuty Malware Protection for EC2 can scan the instance's EBS volumes without an agent, either automatically when GuardDuty raises certain findings or on demand. It works from snapshots, so it doesn't replace your own evidence copies.

Storing evidence immutably

Put forensic artifacts in an S3 bucket in the forensics account with Object Lock.

Governance modeCompliance modeLegal hold
Can anyone delete before expiry?Users with s3:BypassGovernanceRetentionNo one, not even rootUntil someone with s3:PutObjectLegalHold removes it
Can retention be shortened?With the bypass permissionNo (only extended)No expiry date at all
Use forTesting a policy, or evidence where a senior admin may need an overrideEvidence you might need in court or for a regulatorAn open investigation or litigation of unknown length
  • Object Lock needs versioning. You can enable it on new or existing buckets, and set a default retention so every uploaded artifact is locked automatically.
  • Locks apply to object versions. A delete without a version ID just adds a delete marker; the locked version stays.
  • Add a bucket policy that denies access except to the IR roles, SSE-KMS with a dedicated key, and CloudTrail data events on the bucket.

Compliance mode for a trial run

Compliance mode can't be undone: the object (and its storage bill) stays until the retention date. For experiments, use governance mode or a short retention. For real evidence of unknown duration, a legal hold is often better than guessing a retention period.

Containers on EKS

  • Isolate the pod with a Kubernetes NetworkPolicy that denies all ingress and egress for a label you add to the pod (for example quarantine=true). This needs a network policy engine, such as the VPC CNI's network policy support.
  • Cordon the node so no new pods land on it, and consider moving other workloads off it. If the pod was privileged or GuardDuty flags node credentials, treat the node as compromised too.
  • Revoke the pod's AWS credentials: revoke sessions on the IAM role used by IRSA or EKS Pod Identity, and on the node's instance role if pods could reach it.
  • Capture before you delete: container filesystem and process list with kubectl exec or the runtime, memory and a snapshot of the node's volumes, and EKS audit logs in CloudWatch Logs.
  • Then fix the vulnerability, redeploy clean pods, delete the compromised ones, and remove a malicious image from the registry.

Lambda functions

  • Stop invocations by setting the function's reserved concurrency to 0. That throttles every invocation without deleting the function or its versions.
  • Revoke sessions on the execution role, and rotate secrets the function could read (environment variables, Secrets Manager, Parameter Store).
  • Preserve evidence: download the code package and layers (GetFunction returns a link), the configuration and resource policy, and the CloudWatch Logs.
  • Look for tampering: UpdateFunctionCode, UpdateFunctionConfiguration, new layers, and AddPermission calls that let other accounts invoke it.
  • Recover by pointing the alias back at a known-good version, or redeploying from the pipeline.

Recovering

  • Rebuild from a known-good source (golden AMI, pipeline, IaC). Cleaning a compromised host in place leaves you guessing about persistence.
  • Restore data from backups you can trust, ideally from a vault the attacker couldn't touch (see Backup and ransomware protection).
  • Fix the root cause before reconnecting: patch the vulnerability, close the exposed port, require IMDSv2.
  • Use ARC zonal shift if a whole AZ's fleet has to be rebuilt.

Automated Forensics Orchestrator for Amazon EC2

An AWS Solutions guidance implementation (now covering EC2 and EKS) that you deploy into a forensics account.

  • Trigger: a GuardDuty finding, or a Security Hub custom action an analyst chooses on a finding.
  • Step Functions workflows isolate the instance, capture memory (LiME or AVML through SSM), snapshot the volumes and copy everything into the forensics account's S3 bucket.
  • An investigation workflow then analyzes the copies (for example Volatility on the memory image) and writes a report, with notifications to the SOC.
  • Works across accounts and Regions in an organization.

Exam signal

"Automatically capture memory and disk from a flagged instance across 80 accounts and store it for the SOC, with the least custom code" points to the Automated Forensics Orchestrator. "Just isolate the instance and let me restore it later" points to the AWSSupport-ContainEC2Instance runbook.

0
Reserved concurrency that stops a Lambda function from running.
Compliance mode
The only Object Lock mode that even root can't shorten or delete.
Tracked
Existing connections survive a security group change.
aws/ebs
Snapshots encrypted with this key can't be shared across accounts.

Scenarios

Scenario
GuardDuty reports that an instance in an Auto Scaling group behind an ALB is communicating with a known command-and-control server. Before the engineer can act, the instance fails its ALB health check and the Auto Scaling group terminates it, destroying the evidence. What should the incident runbook do FIRST in future to prevent this?
Scenario
A security engineer at Cobaltline Energy attaches an isolation security group with no rules to a compromised instance, but VPC Flow Logs show an established outbound session to an attacker's IP continuing. Which action stops the existing session while keeping the instance running for forensics?
Scenario
A healthcare company must keep memory images and disk images from a breach for at least 7 years for regulators, and no one, including administrators of the forensics account, may delete or alter them during that period. Which solution meets this requirement?

Further reading

On this page