Asterrr's Handbook

Modernization

Choosing target compute, storage and database platforms for existing workloads, and breaking up monoliths with strangler-fig routing and events.

Exam tasks: 4.3 (determine a new architecture for existing workloads), 4.4 (determine opportunities for modernization and enhancements)

The decision: for a workload that already exists, which target platform removes the most operational burden for the effort the business will fund, and how do you change it without a big-bang rewrite?

Picking target platforms

Work through the layers separately. A rehosted app can still get a managed database, and a refactored one can keep EC2 for a component that needs it.

Compute

PlatformYou manageGood target whenWatch out for
EC2OS, patching, scaling configRehosted apps, licensed software, custom kernelsMost ops work of all options
ECS or EKS on EC2Cluster hostsContainerized apps needing GPUs, specific instances or daemon agentsHost patching and capacity
FargateOnly tasksContainers without host managementNo privileged containers or host-level agents
LambdaOnly codeEvent handlers, APIs, scheduled jobs, file processing15-minute limit, cold starts, per-invocation cost at steady high load
Elastic BeanstalkApp code and settingsReplatforming a web app onto managed EC2 with little changeLess control than building it yourself

Details on container networking and launch types are in compute and containers.

Storage

ServiceAccessShared byTypical modernization move
EBSBlock, one AZOne instance, or a few with Multi-Attach on io2Rehosted server disks
EFSNFS, POSIXThousands of Linux clients across AZsReplacing a Linux NFS server, shared content for containers or Lambda
FSx for Windows File ServerSMB, Active DirectoryWindows clientsReplacing Windows file servers with DFS and AD permissions
FSx for NetApp ONTAPNFS, SMB, iSCSIMixed clientsLifting an on-premises NetApp with SnapMirror, multiprotocol shares
FSx for OpenZFSNFSLinux clientsZFS workloads needing snapshots and very low latency
FSx for LustreLustre, linked to S3HPC clustersScratch or persistent storage for HPC and ML training
S3HTTP APIAnythingMoving files off shared disks into objects, static assets, data lakes

Match the protocol first

The protocol usually decides the answer. SMB with AD permissions means FSx for Windows. An existing NetApp estate means FSx for ONTAP. Linux NFS shared across AZs means EFS. HPC reading S3 datasets means FSx for Lustre.

Database

  • Replatform self-managed engines to RDS or Aurora to hand off patching, backups and failover.
  • Refactor access patterns that fit key-value, document, graph or time series stores to a purpose-built database, for example session state to DynamoDB.
  • Engine changes need schema conversion; see database migration and purpose-built databases.

Strangler-fig decomposition

Put a routing layer in front of the monolith, then move one capability at a time to a new service. Clients never see the switch, and each step can be rolled back by changing a route.

Front the monolith with an ALB or API Gateway so every request already passes a layer you control.

Carve out a bounded context with few dependencies, such as notifications or catalog search, and build it as a new service with its own data store.

Shift traffic with a path-based ALB rule or an API Gateway route. Use weighted target groups for a gradual shift.

Retire the old code path once the new service carries all traffic. Repeat for the next capability.

Routing layerChoose when
ALB listener rulesPath or host routing between EC2, containers and Lambda targets, with weighted target groups for canaries
API GatewayYou also want throttling, API keys, request validation or authorizers, or a private integration to the monolith through a VPC link
CloudFrontThe routing should happen at the edge, with different origins for different paths

Migration Hub Refactor Spaces automated this pattern across accounts: it created the API Gateway, the network links between the monolith's VPC and the new service's account, and the routes. It stopped accepting new customers in November 2025, so treat it as a recognizable exam answer rather than a tool to adopt now.

Sharing the monolith's database

A new service that reads and writes the monolith's tables directly isn't decoupled: schema changes still break both. Give the new service its own store, and sync through events or an API during the transition.

Event-driven refactoring

Replace synchronous calls and polling with events where the caller doesn't need an immediate answer.

Legacy patternModern replacement
Cron job on a serverEventBridge Scheduler invoking Lambda, ECS tasks or Step Functions
App polls a table for new workSQS queue with consumers that scale on queue depth
One service calls five others in turnPublish one event to SNS or EventBridge, and each consumer subscribes
Long workflow coded in the app with retriesStep Functions state machine
Batch file drop on an FTP serverS3 upload triggers an event and processing through Lambda or Step Functions

For choosing between queues, topics, buses and workflows, see decoupling.

Serverless signals

"Spiky or unpredictable traffic", "idle most of the day", "reduce operational overhead" and "no servers to patch" point to Lambda, Fargate, DynamoDB on-demand and Aurora Serverless v2. "Steady 24/7 load" and "long-running process" point back to containers or EC2.

Lambda for everything

Moving a job that runs for 40 minutes, or needs a persistent TCP connection, into Lambda fails the 15-minute limit. Use Fargate tasks, AWS Batch or Step Functions to split the work.

Scenarios

Scenario
A company rehosted a Java monolith on EC2. It wants to rebuild the checkout module as a separate service, release it gradually, and roll back quickly if errors spike, without changing the URLs clients use. Which approach fits?
Scenario · choose 2
A media company is modernizing three workloads after migration: Windows file shares with NTFS permissions from Active Directory, a Linux rendering farm that reads large datasets from S3 at high throughput, and a web tier on Linux that shares uploaded content across AZs. Which TWO storage choices are correct?
Scenario
A payment service calls fraud scoring, email, loyalty and analytics services in sequence after each payment. A slow loyalty service is delaying payment confirmations. The team wants to decouple the downstream steps with the LEAST custom code. What should it do?

Further reading

On this page