Decoupling with queues, topics and events
When to use SQS, SNS, EventBridge, Step Functions, Kinesis Data Streams, Amazon MQ or MSK to break a system into independent parts.
Exam tasks: 2.4 (design loosely coupled, resilient architectures), 4.4 (decouple components during modernization)
The decision: does one consumer need to work through a backlog, do many consumers need a copy, does an event need routing by content, or does a multi-step process need coordinating? Each answer points to a different service.
Choosing a service
Side by side
| SQS | SNS | EventBridge | Kinesis Data Streams | |
|---|---|---|---|---|
| Model | Queue: pull, one consumer group | Pub/sub: push to subscribers | Event bus: rules route to targets | Ordered log: consumers read shards |
| Message kept after read | No, consumer deletes it | No storage | No (archive is optional) | Yes, until retention expires |
| Replay | No | No | Archive and replay | Yes, from any position |
| Ordering | FIFO queues only | FIFO topics only | Not guaranteed | Per shard (partition key) |
| Filtering | None (consumer decides) | Filter policies per subscription | Rich event patterns per rule | None (consumer decides) |
| Scaling unit | Automatic | Automatic | Automatic | Shards, or on-demand mode |
| Typical use | Buffer work, absorb spikes | Fan out one message to many | SaaS and AWS events, cross-account routing | Clickstreams, telemetry, real-time analytics |
ReceiveMessageWaitTimeSeconds).SQS
Standard vs FIFO
| Standard | FIFO | |
|---|---|---|
| Delivery | At least once: duplicates possible | Exactly-once processing (5-minute deduplication window) |
| Order | Best effort | Strict, per message group ID |
| Throughput | Nearly unlimited | Limited per API action. Batching and high-throughput mode raise it |
| Name | Any | Must end in .fifo |
- FIFO parallelism comes from message groups. One group is processed in order by one consumer at a time, so use one group per customer or order ID, not one group for everything.
- Deduplication uses a
MessageDeduplicationIdor content-based deduplication (a hash of the body). - A standard queue can't be converted to FIFO. You create a new queue and move producers.
Exam signal
"Must process each order exactly once, in the order received" means FIFO. "Millions of messages per second, duplicates acceptable" or no ordering requirement means Standard, and the consumer should be idempotent.
The knobs that show up in questions
- Visibility timeout. After a consumer receives a message, it's hidden for this long. If processing takes
longer, another consumer receives it too. Set it to more than the processing time, or call
ChangeMessageVisibilityto extend it while working. For a Lambda trigger, set it to at least six times the function timeout. - Long polling. A wait time of 1–20 seconds makes
ReceiveMessagewait for messages instead of returning empty immediately. Fewer empty responses means lower cost and fewer API calls. - Dead-letter queue. After
maxReceiveCountfailed receives, the redrive policy moves the message to a DLQ so one poison message doesn't block or loop forever. A FIFO queue's DLQ must also be FIFO. Use DLQ redrive to move messages back once the bug is fixed. - Delay queues and message timers postpone delivery by up to 15 minutes.
- Scaling consumers. An Auto Scaling group can scale on queue depth, ideally as backlog per instance
(
ApproximateNumberOfMessagesVisibledivided by running instances) against a target.
Visibility timeout vs delay
Duplicate processing because jobs run longer than expected is fixed by a longer visibility timeout, not a delay queue. A delay hides a message before its first receive, and does nothing once a consumer has it.
SNS
- Fan-out: publish once to a topic that has several SQS queues subscribed. Each queue gets its own copy, and each consumer works at its own pace. This is the standard "one event, many independent processors" pattern.
- Subscription filter policies match on message attributes or on the message body, so a subscriber only gets what it needs. This removes filtering code from consumers.
- Subscribers: SQS, Lambda, HTTP/S, email, SMS, mobile push and Amazon Data Firehose.
- FIFO topics keep order and deduplicate when delivering to SQS FIFO queues.
- Failed deliveries to a subscriber can go to a DLQ (an SQS queue) set on the subscription.
EventBridge
- Buses. The default bus receives AWS service events. Custom buses hold your application events. Partner buses receive events from SaaS providers such as Zendesk or Datadog.
- Rules match an event pattern (source, detail-type, any field in the detail) and send the event to targets, optionally with an input transformer that reshapes it.
- Cross-account and cross-Region: a rule can target an event bus in another account or Region. The receiving
bus needs a resource policy that allows the sender, often scoped with
aws:PrincipalOrgID. This is how you centralize events from every account into one monitoring account. - Archive and replay keeps events so you can reprocess them after fixing a consumer.
- Schema registry discovers event shapes and generates code bindings.
- Pipes connect one source (SQS, Kinesis, DynamoDB Streams, MSK, Amazon MQ) to one target, with optional filtering and an enrichment step (Lambda, Step Functions, API Gateway or an API destination). Use it instead of writing glue Lambda functions.
- Scheduler runs one-time or recurring invocations of thousands of targets, with time zones and flexible windows. It replaces scheduled rules for new designs, and scales to millions of schedules.
- API destinations call third-party HTTP APIs with managed auth and rate limiting.
Exam signal
"React to an AWS service event", "route events from a SaaS application", or "send events from all accounts to a central account" points to EventBridge. Plain "notify several subscribers" with no routing logic is SNS.
Step Functions
| Standard | Express | |
|---|---|---|
| Maximum duration | 1 year | 5 minutes |
| Execution semantics | Exactly once | At least once (async) or at most once (sync) |
| Rate | Lower start rate | Very high start rate |
| Execution history | Stored and viewable in the console | Sent to CloudWatch Logs |
| Billing | Per state transition | Per execution, duration and memory |
| Good for | Order workflows, human approval, long ETL | High-volume event processing, IoT ingestion, microservice orchestration |
- Service integrations call more than 200 AWS services directly through SDK integrations, without a Lambda function in between.
- Callback with task token (
.waitForTaskToken) pauses until an external system or a human calls back. - Distributed Map runs up to thousands of parallel child executions over items in S3, for large batch jobs.
- Built-in Retry and Catch replace hand-written retry and compensation code (for example, a saga pattern).
Legacy: use AWS Step Functions instead
Amazon Simple Workflow Service (SWF) coordinated tasks through deciders and activity workers that you wrote and ran yourself. It still exists, but new designs and exam answers use Step Functions. The exception is an existing SWF application that needs child workflows or external signals and can't be rewritten yet.
Kinesis Data Streams vs SQS
Pick Kinesis Data Streams when:
- Several applications must read the same records independently (analytics, archive and alerting).
- You need ordering per key plus replay of the last hours or days.
- Consumers need a sliding window over time, such as real-time aggregation with Managed Service for Apache Flink.
Pick SQS when each message is a unit of work that one worker handles and deletes, and you don't want to manage shards.
- Shard capacity: 1 MB/s or 1,000 records/s in, 2 MB/s out, shared by all standard consumers.
- Enhanced fan-out gives each registered consumer its own 2 MB/s per shard with push delivery.
- On-demand capacity mode removes shard planning for unpredictable traffic.
Kinesis for a work queue
Kinesis has no per-message acknowledgement or visibility timeout. A failed record blocks its shard until the consumer skips it or retries succeed. For independent jobs with retries and a DLQ, SQS is the better fit.
Lift and shift of existing brokers
- Amazon MQ runs managed ActiveMQ or RabbitMQ. Choose it when an on-premises app uses JMS, AMQP, STOMP, MQTT or OpenWire and you can't change its code. Use an active/standby broker across two AZs for availability.
- Amazon MSK runs managed Apache Kafka (provisioned or serverless). MSK Connect runs Kafka Connect connectors. Choose it when the existing platform is already Kafka.
- If the app can be rewritten, SQS and SNS scale further with less to operate. The exam usually signals this as "minimize code changes" (Amazon MQ) vs "cloud-native, least operational overhead" (SQS/SNS).
Scenarios
SNS fan-out to per-service SQS queues gives every service its own copy and lets the slow archiver fall behind without affecting the others. A subscription filter policy sends only high-value claims to manual review. A single shared queue delivers each message to only one consumer. Kinesis works but pushes filtering into every consumer. Synchronous calls couple the claims service to the slowest dependency.
With the default 30-second visibility timeout, messages reappear while the first worker is still transcoding, so other workers pick them up. A timeout longer than 8 minutes fixes that. A DLQ removes the malformed file after three attempts. Long polling reduces empty receives but not duplicates. A delay only affects the first delivery. FIFO doesn't help when the duplicate comes from an expired visibility timeout, and it lowers throughput.
EventBridge natively receives GuardDuty findings and CloudTrail-based events, can forward them to another account's bus, and routes by content with rules. A bus policy using aws:PrincipalOrgID covers all 40 accounts at once. SNS lacks content routing across event types. Parsing raw CloudTrail with Kinesis adds a lot of custom code. Step Functions orchestrates workflows and isn't an event router.
Further reading
Load balancers
Choosing between Application, Network and Gateway Load Balancers, and how cross-zone balancing, stickiness, TLS, authentication and static IPs work on each.
Caching
Where to cache (CloudFront, API Gateway, ElastiCache, DAX), how to raise the CloudFront cache hit ratio, and how to make a cache layer fault tolerant.