Hands-on: build a lightweight SNS–SQS–Lambda fanout pipeline in a client's AWS account
One topic, three queues, one Lambda. This exercise fits within the Free Tier and teaches the duplicate and lost-message failures you will meet in production on a client site.
In brief
- Standard SQS queues deliver messages at least once, and Lambda can also process duplicates, so your code must be idempotent.
- Enable ReportBatchItemFailures so only failed messages return to the queue, and set maxReceiveCount to move broken messages to a DLQ.
- Lambda and the queue must be in the same Region but can sit in different accounts, a constraint to settle early when working in a client's AWS.
- 1Order system publishesPublishes once to the order-placed SNS topic, without knowing who is listening
- 2SNS fans outEach subscribed queue receives an identical copy of the message
- 3SQS holds messagesinventory-queue and analytics-queue hold messages, 4 days by default
- 4Lambda pulls in batchesInventory branch: event source mapping pulls up to 10 messages; handler must be idempotent
- 5Failures retry or go to DLQInventory branch: only failed messages retry; past maxReceiveCount they go to inventory-dlq
The queue sits between SNS and Lambda so messages aren't lost, and the DLQ holds bad messages instead of letting them loop forever.
Graphic: FDE Times
Amazon SQS standard queues promise only “at least once” delivery.
The client does not want another server to babysit. The SNS, SQS and Lambda trio is a neat answer: serverless bills by execution time, starting when your code runs and stopping when it stops, so an idle pipeline costs next to nothing.
This guide rebuilds the official AWS fanout example (an order is placed and split into an inventory queue and an analytics queue), then adds Lambda and a dead-letter queue. By the end, you will know where this pipeline breaks and how to prevent it from the start.
What do you need before you start?
A sandbox AWS account of your own. Do not practise in the client’s account. The Free Tier includes 1,000,000 SNS publishes and 1,000,000 SQS requests, far more than this exercise needs.
Pick one Region and stick with it throughout. Lambda and the SQS queue must be in the same Region, although they are allowed to live in different accounts. Write the Region name in a notes file now.
Step 1: the SNS topic is the loudspeaker for orders
SNS is a fully managed pub/sub service. Pub/sub is an asynchronous communication model: publishers do not need to know who is listening, and subscribers do not need to know who is speaking. The client’s ordering system only has to publish once.
In the SNS console, create a topic named order-placed. Check: the topic appears in the list, and you have copied its ARN into your notes file.
Step 2: create three queues, not two
Create three Standard queues: inventory-queue, analytics-queue and inventory-dlq. The third is a dead-letter queue (DLQ), a built-in SQS feature. It is easy to skip when everything is running smoothly, so do not skip it.
The default message retention period is 4 days, adjustable from 60 seconds to 1,209,600 seconds (14 days). Practical advice: set the DLQ to 14 days. Broken messages are often discovered after a weekend, and you want them still there when someone looks.
On inventory-queue, set a RedrivePolicy pointing to inventory-dlq with a maxReceiveCount. The snippet below is a sketch to convey the idea, not the syntax of any particular tool:
# Sketch, simplified
inventory-queue:
type: standard
RedrivePolicy:
dead_letter_queue: inventory-dlq
maxReceiveCount: 3
inventory-dlq:
retention_seconds: 1209600
This guide attaches Lambda only to the inventory branch, so only the inventory branch gets a DLQ. analytics-queue has no consumer here yet. When you attach one, create an analytics-dlq with the same configuration; do not leave one branch protected and the other exposed.
Check: in the configuration of inventory-queue, the dead-letter queue field shows inventory-dlq.
Step 3: subscribe both queues to the topic
From each queue’s page, subscribe it to the order-placed topic. When several queues subscribe to the same topic, each receives an identical copy every time a message is published. Inventory and analytics therefore run in parallel, without waiting on each other.
Check: from the SNS console, publish a test message such as {"order_id": "A-001", "sku": "SKU-42", "qty": 2}. Then poll both queues from the SQS console. If only one queue sees the message, the other queue’s subscription did not succeed.
Why not have SNS call Lambda directly? AWS recommends pairing SNS with SQS so messages are not lost while a subscriber is offline. The queue is a buffer: if Lambda fails or is throttled, the message stays there waiting.
Step 4: attach Lambda with an event source mapping
Create a Lambda function inventory-worker in the same Region, then add an SQS trigger pointing to inventory-queue. This trigger is an event source mapping: Lambda polls the queue itself and invokes the function in batches. By default each batch holds at most 10 messages, and you can add a batch window of up to 5 minutes to let batches fill.
In the mapping configuration, enable partial batch response by setting ReportBatchItemFailures in FunctionResponseTypes:
# Sketch, simplified
event_source_mapping:
function: inventory-worker
queue: inventory-queue
batch_size: 10
FunctionResponseTypes: [ReportBatchItemFailures]
Check: the trigger shows as enabled, and the function’s logs record the test message from step 3.
What if one message in a batch fails?
Run the numbers. Lambda pulls a batch of 10 messages, and message 7 fails because the SKU does not exist. A message being processed stays in the queue, merely hidden for the visibility timeout. Without partial batch response, the whole batch is treated as failed and returns after the visibility timeout.
On the next run, the 9 good messages are processed again. If the handler deducts stock whenever it receives a message, order A-001 has 2 items deducted twice. Nothing shows up in the logs; the stock count is simply wrong.
With ReportBatchItemFailures enabled, only message 7 returns. With maxReceiveCount: 3, it is retried a few more times and then moved to inventory-dlq. Without a DLQ, broken messages cycle forever, until unprocessable messages outnumber the valid ones in the queue.
Step 5: write a handler that tolerates duplicates
Partial batch response reduces duplicates but does not eliminate them. An event source mapping processes each event at least once and can still process duplicates. Your code must be idempotent.
The Python sketch below is simplified: already_processed, mark_processed and build_partial_response are functions you write yourself, and the return format for partial batch response should be checked against the Lambda documentation before running it for real.
# Sketch, simplified
def handler(event, context):
failed_ids = []
for record in event["Records"]:
msg_id = record["messageId"]
try:
order = parse_order(record) # your own helper
key = order["order_id"]
if already_processed(key): # look up the idempotency key table
continue
apply_inventory_change(order) # business logic
mark_processed(key)
except Exception:
failed_ids.append(msg_id)
return build_partial_response(failed_ids) # report only the failed messages
The detail worth noticing is that the idempotency key is the business order_id, not the message ID. Imagine the client’s ordering system retries on its own and republishes order A-001 to the topic: the queue receives a new message for the same order. A key based on order_id still recognises the order as processed; a key based on the message ID may let it through.
Check: publish the exact A-001 message twice more. Stock must be deducted only once.
The most common mistakes
The first is creating the Lambda in a different Region from the queue, then spending half an hour wondering why the trigger does not appear. The second is forgetting the DLQ, which usually surfaces only when the queue balloons.
The third is believing Standard queues deliver exactly once. If the business genuinely needs exactly-once processing, FIFO queues are what provide it, and you should raise that option with the client rather than decide alone.
The last mistake lies outside the code: not cleaning up. The official AWS tutorial tells you to delete resources after practising. In a client’s account, a forgotten topic or queue is something their operations team will be asking you about later.
On a client site, what should you ask first?
If the client gives your team a separate account, distinct from the one running the ordering system, ask about the Region at the kickoff meeting. The queue and Lambda can be in different accounts, but not in different Regions.
The second question to ask is: “What breaks if an order is processed twice?” The answer tells you how strict the idempotency needs to be, or whether FIFO needs discussing.
When an FDE job description mentions “event-driven” or “integration”, ask the recruiter whether the actual work includes pipelines like this one; if it does, this exercise prepares you in the right direction.
On your CV, do not write “knows SQS”. Write a measurable result: for example, that you built an SNS–SQS fanout with a DLQ and partial batch response, then proved with a duplicate-publish test that stock was not deducted twice. That shows you understand where pipelines tend to break, not just the names of the services.
The pipeline is light enough to build in an afternoon. What the client will remember longer is message 7: it sat quietly in the DLQ waiting for someone to look, instead of silently deducting stock twice.
Was this article useful?
Thanks for the feedback!
8 sources
- What is Amazon Simple Queue Service?
- Push Notification Service - Amazon Simple Notification Service (SNS) - AWS
- What is Pub/Sub Messaging? - Pub/Sub Messaging Explained - AWS
- Send Fanout Event Notifications
- Using Lambda with Amazon SQS
- Handling errors for an SQS event source in Lambda
- Best practices for implementing partial batch responses
- What Is Serverless Computing? | IBM · 2024-06-10