Asterrr's Handbook

Decoupling with queues, topics and events

When to use SQS, SNS, EventBridge, Step Functions, Kinesis Data Streams, Amazon MQ or MSK to break a system into independent parts.

Exam tasks: 2.4 (design loosely coupled, resilient architectures), 4.4 (decouple components during modernization)

The decision: does one consumer need to work through a backlog, do many consumers need a copy, does an event need routing by content, or does a multi-step process need coordinating? Each answer points to a different service.

Choosing a service

Side by side

SQSSNSEventBridgeKinesis Data Streams
ModelQueue: pull, one consumer groupPub/sub: push to subscribersEvent bus: rules route to targetsOrdered log: consumers read shards
Message kept after readNo, consumer deletes itNo storageNo (archive is optional)Yes, until retention expires
ReplayNoNoArchive and replayYes, from any position
OrderingFIFO queues onlyFIFO topics onlyNot guaranteedPer shard (partition key)
FilteringNone (consumer decides)Filter policies per subscriptionRich event patterns per ruleNone (consumer decides)
Scaling unitAutomaticAutomaticAutomaticShards, or on-demand mode
Typical useBuffer work, absorb spikesFan out one message to manySaaS and AWS events, cross-account routingClickstreams, telemetry, real-time analytics
12 hours
Maximum SQS visibility timeout. The default is 30 seconds.
14 days
Maximum SQS retention. The default is 4 days.
20 seconds
Longest long-polling wait (ReceiveMessageWaitTimeSeconds).
1 MiB
Maximum SQS and EventBridge payload. Larger payloads go to S3 with a pointer in the message.
1 year
Maximum Step Functions Standard execution. Express workflows stop at 5 minutes.
365 days
Maximum Kinesis Data Streams retention. The default is 24 hours.

SQS

Standard vs FIFO

StandardFIFO
DeliveryAt least once: duplicates possibleExactly-once processing (5-minute deduplication window)
OrderBest effortStrict, per message group ID
ThroughputNearly unlimitedLimited per API action. Batching and high-throughput mode raise it
NameAnyMust end in .fifo
  • FIFO parallelism comes from message groups. One group is processed in order by one consumer at a time, so use one group per customer or order ID, not one group for everything.
  • Deduplication uses a MessageDeduplicationId or content-based deduplication (a hash of the body).
  • A standard queue can't be converted to FIFO. You create a new queue and move producers.

Exam signal

"Must process each order exactly once, in the order received" means FIFO. "Millions of messages per second, duplicates acceptable" or no ordering requirement means Standard, and the consumer should be idempotent.

The knobs that show up in questions

  • Visibility timeout. After a consumer receives a message, it's hidden for this long. If processing takes longer, another consumer receives it too. Set it to more than the processing time, or call ChangeMessageVisibility to extend it while working. For a Lambda trigger, set it to at least six times the function timeout.
  • Long polling. A wait time of 1–20 seconds makes ReceiveMessage wait for messages instead of returning empty immediately. Fewer empty responses means lower cost and fewer API calls.
  • Dead-letter queue. After maxReceiveCount failed receives, the redrive policy moves the message to a DLQ so one poison message doesn't block or loop forever. A FIFO queue's DLQ must also be FIFO. Use DLQ redrive to move messages back once the bug is fixed.
  • Delay queues and message timers postpone delivery by up to 15 minutes.
  • Scaling consumers. An Auto Scaling group can scale on queue depth, ideally as backlog per instance (ApproximateNumberOfMessagesVisible divided by running instances) against a target.

Visibility timeout vs delay

Duplicate processing because jobs run longer than expected is fixed by a longer visibility timeout, not a delay queue. A delay hides a message before its first receive, and does nothing once a consumer has it.

SNS

  • Fan-out: publish once to a topic that has several SQS queues subscribed. Each queue gets its own copy, and each consumer works at its own pace. This is the standard "one event, many independent processors" pattern.
  • Subscription filter policies match on message attributes or on the message body, so a subscriber only gets what it needs. This removes filtering code from consumers.
  • Subscribers: SQS, Lambda, HTTP/S, email, SMS, mobile push and Amazon Data Firehose.
  • FIFO topics keep order and deduplicate when delivering to SQS FIFO queues.
  • Failed deliveries to a subscriber can go to a DLQ (an SQS queue) set on the subscription.

EventBridge

  • Buses. The default bus receives AWS service events. Custom buses hold your application events. Partner buses receive events from SaaS providers such as Zendesk or Datadog.
  • Rules match an event pattern (source, detail-type, any field in the detail) and send the event to targets, optionally with an input transformer that reshapes it.
  • Cross-account and cross-Region: a rule can target an event bus in another account or Region. The receiving bus needs a resource policy that allows the sender, often scoped with aws:PrincipalOrgID. This is how you centralize events from every account into one monitoring account.
  • Archive and replay keeps events so you can reprocess them after fixing a consumer.
  • Schema registry discovers event shapes and generates code bindings.
  • Pipes connect one source (SQS, Kinesis, DynamoDB Streams, MSK, Amazon MQ) to one target, with optional filtering and an enrichment step (Lambda, Step Functions, API Gateway or an API destination). Use it instead of writing glue Lambda functions.
  • Scheduler runs one-time or recurring invocations of thousands of targets, with time zones and flexible windows. It replaces scheduled rules for new designs, and scales to millions of schedules.
  • API destinations call third-party HTTP APIs with managed auth and rate limiting.

Exam signal

"React to an AWS service event", "route events from a SaaS application", or "send events from all accounts to a central account" points to EventBridge. Plain "notify several subscribers" with no routing logic is SNS.

Step Functions

StandardExpress
Maximum duration1 year5 minutes
Execution semanticsExactly onceAt least once (async) or at most once (sync)
RateLower start rateVery high start rate
Execution historyStored and viewable in the consoleSent to CloudWatch Logs
BillingPer state transitionPer execution, duration and memory
Good forOrder workflows, human approval, long ETLHigh-volume event processing, IoT ingestion, microservice orchestration
  • Service integrations call more than 200 AWS services directly through SDK integrations, without a Lambda function in between.
  • Callback with task token (.waitForTaskToken) pauses until an external system or a human calls back.
  • Distributed Map runs up to thousands of parallel child executions over items in S3, for large batch jobs.
  • Built-in Retry and Catch replace hand-written retry and compensation code (for example, a saga pattern).

Legacy: use AWS Step Functions instead

Amazon Simple Workflow Service (SWF) coordinated tasks through deciders and activity workers that you wrote and ran yourself. It still exists, but new designs and exam answers use Step Functions. The exception is an existing SWF application that needs child workflows or external signals and can't be rewritten yet.

Kinesis Data Streams vs SQS

Pick Kinesis Data Streams when:

  • Several applications must read the same records independently (analytics, archive and alerting).
  • You need ordering per key plus replay of the last hours or days.
  • Consumers need a sliding window over time, such as real-time aggregation with Managed Service for Apache Flink.

Pick SQS when each message is a unit of work that one worker handles and deletes, and you don't want to manage shards.

  • Shard capacity: 1 MB/s or 1,000 records/s in, 2 MB/s out, shared by all standard consumers.
  • Enhanced fan-out gives each registered consumer its own 2 MB/s per shard with push delivery.
  • On-demand capacity mode removes shard planning for unpredictable traffic.

Kinesis for a work queue

Kinesis has no per-message acknowledgement or visibility timeout. A failed record blocks its shard until the consumer skips it or retries succeed. For independent jobs with retries and a DLQ, SQS is the better fit.

Lift and shift of existing brokers

  • Amazon MQ runs managed ActiveMQ or RabbitMQ. Choose it when an on-premises app uses JMS, AMQP, STOMP, MQTT or OpenWire and you can't change its code. Use an active/standby broker across two AZs for availability.
  • Amazon MSK runs managed Apache Kafka (provisioned or serverless). MSK Connect runs Kafka Connect connectors. Choose it when the existing platform is already Kafka.
  • If the app can be rewritten, SQS and SNS scale further with less to operate. The exam usually signals this as "minimize code changes" (Amazon MQ) vs "cloud-native, least operational overhead" (SQS/SNS).

Scenarios

Scenario
An insurer's claims service publishes a ClaimSubmitted message. A fraud-scoring service, a document-archiving service and a notification service must each process every claim independently, and the archiving service is often slow. High-value claims over a threshold must also go to a manual-review service, which must not receive other claims. Which design meets these requirements with the LEAST custom code?
Scenario · choose 2
A video platform uses an SQS standard queue and EC2 workers to transcode uploads. Transcoding takes up to 8 minutes. The team sees the same video transcoded two or three times, and occasionally one malformed file causes workers to fail repeatedly. Which TWO changes fix these problems?
Scenario
A retailer runs 40 AWS accounts in AWS Organizations. The security team wants every GuardDuty finding and every root-user sign-in event from all accounts delivered to a single account, where different findings trigger different Lambda functions. What should the architect do?

Further reading

On this page