Modernization
Choosing target compute, storage and database platforms for existing workloads, and breaking up monoliths with strangler-fig routing and events.
Exam tasks: 4.3 (determine a new architecture for existing workloads), 4.4 (determine opportunities for modernization and enhancements)
The decision: for a workload that already exists, which target platform removes the most operational burden for the effort the business will fund, and how do you change it without a big-bang rewrite?
Picking target platforms
Work through the layers separately. A rehosted app can still get a managed database, and a refactored one can keep EC2 for a component that needs it.
Compute
| Platform | You manage | Good target when | Watch out for |
|---|---|---|---|
| EC2 | OS, patching, scaling config | Rehosted apps, licensed software, custom kernels | Most ops work of all options |
| ECS or EKS on EC2 | Cluster hosts | Containerized apps needing GPUs, specific instances or daemon agents | Host patching and capacity |
| Fargate | Only tasks | Containers without host management | No privileged containers or host-level agents |
| Lambda | Only code | Event handlers, APIs, scheduled jobs, file processing | 15-minute limit, cold starts, per-invocation cost at steady high load |
| Elastic Beanstalk | App code and settings | Replatforming a web app onto managed EC2 with little change | Less control than building it yourself |
Details on container networking and launch types are in compute and containers.
Storage
| Service | Access | Shared by | Typical modernization move |
|---|---|---|---|
| EBS | Block, one AZ | One instance, or a few with Multi-Attach on io2 | Rehosted server disks |
| EFS | NFS, POSIX | Thousands of Linux clients across AZs | Replacing a Linux NFS server, shared content for containers or Lambda |
| FSx for Windows File Server | SMB, Active Directory | Windows clients | Replacing Windows file servers with DFS and AD permissions |
| FSx for NetApp ONTAP | NFS, SMB, iSCSI | Mixed clients | Lifting an on-premises NetApp with SnapMirror, multiprotocol shares |
| FSx for OpenZFS | NFS | Linux clients | ZFS workloads needing snapshots and very low latency |
| FSx for Lustre | Lustre, linked to S3 | HPC clusters | Scratch or persistent storage for HPC and ML training |
| S3 | HTTP API | Anything | Moving files off shared disks into objects, static assets, data lakes |
Match the protocol first
The protocol usually decides the answer. SMB with AD permissions means FSx for Windows. An existing NetApp estate means FSx for ONTAP. Linux NFS shared across AZs means EFS. HPC reading S3 datasets means FSx for Lustre.
Database
- Replatform self-managed engines to RDS or Aurora to hand off patching, backups and failover.
- Refactor access patterns that fit key-value, document, graph or time series stores to a purpose-built database, for example session state to DynamoDB.
- Engine changes need schema conversion; see database migration and purpose-built databases.
Strangler-fig decomposition
Put a routing layer in front of the monolith, then move one capability at a time to a new service. Clients never see the switch, and each step can be rolled back by changing a route.
Front the monolith with an ALB or API Gateway so every request already passes a layer you control.
Carve out a bounded context with few dependencies, such as notifications or catalog search, and build it as a new service with its own data store.
Shift traffic with a path-based ALB rule or an API Gateway route. Use weighted target groups for a gradual shift.
Retire the old code path once the new service carries all traffic. Repeat for the next capability.
| Routing layer | Choose when |
|---|---|
| ALB listener rules | Path or host routing between EC2, containers and Lambda targets, with weighted target groups for canaries |
| API Gateway | You also want throttling, API keys, request validation or authorizers, or a private integration to the monolith through a VPC link |
| CloudFront | The routing should happen at the edge, with different origins for different paths |
Migration Hub Refactor Spaces automated this pattern across accounts: it created the API Gateway, the network links between the monolith's VPC and the new service's account, and the routes. It stopped accepting new customers in November 2025, so treat it as a recognizable exam answer rather than a tool to adopt now.
Sharing the monolith's database
A new service that reads and writes the monolith's tables directly isn't decoupled: schema changes still break both. Give the new service its own store, and sync through events or an API during the transition.
Event-driven refactoring
Replace synchronous calls and polling with events where the caller doesn't need an immediate answer.
| Legacy pattern | Modern replacement |
|---|---|
| Cron job on a server | EventBridge Scheduler invoking Lambda, ECS tasks or Step Functions |
| App polls a table for new work | SQS queue with consumers that scale on queue depth |
| One service calls five others in turn | Publish one event to SNS or EventBridge, and each consumer subscribes |
| Long workflow coded in the app with retries | Step Functions state machine |
| Batch file drop on an FTP server | S3 upload triggers an event and processing through Lambda or Step Functions |
For choosing between queues, topics, buses and workflows, see decoupling.
Serverless signals
"Spiky or unpredictable traffic", "idle most of the day", "reduce operational overhead" and "no servers to patch" point to Lambda, Fargate, DynamoDB on-demand and Aurora Serverless v2. "Steady 24/7 load" and "long-running process" point back to containers or EC2.
Lambda for everything
Moving a job that runs for 40 minutes, or needs a persistent TCP connection, into Lambda fails the 15-minute limit. Use Fargate tasks, AWS Batch or Step Functions to split the work.
Scenarios
This is strangler fig: a routing layer the monolith already sits behind, with a path rule for the extracted module and weights for a gradual shift and quick rollback. A full rewrite is a big-bang change. Peering connects networks but doesn't route users. Having the monolith call the new service still sends checkout traffic through the old code and doesn't allow gradual release.
SMB with AD permissions maps to FSx for Windows. High-throughput processing of S3 datasets maps to FSx for Lustre linked to the bucket. EFS would fit the web tier but not HPC throughput. EBS Multi-Attach stays within one AZ, and File Gateway is for on-premises access, not a cloud-native file server.
Fraud scoring must finish before confirming, but the other steps don't block the customer. One event with a rule per consumer lets each service fail or slow down on its own. Longer timeouts make the delay worse, a bigger instance doesn't remove the coupling, and polling a shared table adds latency and couples everyone to one schema.