Poison messages can silently corrupt your entire integration pipeline if you don't have a strategy to quarantine and recover them. This lesson walks through building a complete dead-letter queue handling system in Power Automate with Azure Service Bus — from failure classification and explicit dead-lettering to DLQ monitoring, operator review, and safe bulk reprocessing.

Picture this: your enterprise integration pipeline processes thousands of order messages per day between your ERP system and fulfillment warehouse. One Friday afternoon, a supplier sends a malformed payload — a JSON object with a null where a required numeric field should be. Your flow fails. The message retry count ticks upward. After three attempts, the message either disappears into the void or — if you haven't set this up correctly — your flow locks up in a retry loop that hammers your API rate limits and blocks every valid message behind it. By Monday morning, your operations team is staring at 2,000 unprocessed orders and a very angry fulfillment director.
This is exactly the problem that dead-letter queue (DLQ) handling and poison message recovery exist to solve. A poison message is any message that causes a consumer to repeatedly fail — whether due to malformed data, a missing reference, an expired token, or a downstream service being unavailable. Rather than letting that one bad message corrupt your entire pipeline, you quarantine it, process the healthy messages normally, and then deal with the bad ones systematically. Done right, this pattern is the difference between a resilient enterprise integration and a fragile one.
By the end of this lesson, you'll have a complete working implementation of dead-letter handling in Power Automate backed by Azure Service Bus. You'll understand not just how to detect and isolate failed messages, but how to build a recovery workflow that lets operators inspect, correct, and reprocess them — without manual queue manipulation or developer intervention.
What you'll learn:
You should be comfortable with:
Before we build anything, we need to precisely understand what we're dealing with.
In any message-based integration, a poison message is one that cannot be processed successfully — not due to a transient infrastructure blip, but due to something inherent about the message itself. The distinction matters enormously. A transient failure (network timeout, brief API unavailability) should be retried automatically. A poison message should not — because retrying it indefinitely wastes resources and starves valid messages of processing capacity.
Azure Service Bus has this distinction built in at the infrastructure level. Every queue and topic subscription has two sub-queues:
The DLQ is addressed as <queue-name>/$DeadLetterQueue. It holds messages indefinitely (up to the namespace's message time-to-live) and supports the same peek-lock and complete operations as the main queue. Critically, messages in the DLQ retain their original body, properties, and headers — plus additional system properties explaining why they were dead-lettered, including DeadLetterReason and DeadLetterErrorDescription.
Key insight
The maximum delivery count is set on the Service Bus queue, not in Power Automate. If you set it to 3, Service Bus will attempt to deliver the message up to 3 times before automatically moving it to the DLQ. Your Power Automate flow sees each delivery attempt as a separate trigger invocation. This means your flow's retry configuration and the queue's delivery count work as independent layers of retry logic — understanding the interaction between them is critical.
In Power Automate's Service Bus connector, when you receive a message using peek-lock mode and your flow fails to complete the message (either because the flow run fails, or the lock expires), Service Bus increments the delivery count. Once the max delivery count is hit, the message moves to the DLQ automatically. If your flow successfully completes the message despite processing errors (a common mistake we'll cover later), the message disappears permanently with no record of failure.
This creates three design questions your architecture must answer:
When should a message be dead-lettered explicitly vs. allowed to fail? Explicit dead-lettering (where your flow calls the "Dead-letter message in a queue" action) is cleaner — you can set a reason and description before moving it. Letting it fail automatically means you rely on the delivery count reaching its limit, which takes multiple retry cycles.
How do you surface DLQ contents for operational review? The DLQ isn't visible in the Power Platform UI — you need to build monitoring tooling.
How do you reprocess a corrected message? You need a workflow where an operator can inspect the message payload, correct it if needed, and push it back to the main queue.
We're going to build all three pieces.
The first thing your main processing flow needs to do is distinguish between failures that warrant retry and failures that warrant dead-lettering. This is error classification, and it's the hardest part to get right.
Here's a realistic taxonomy for an order processing pipeline:
| Failure Type | Example | Correct Response |
|---|---|---|
| Transient infrastructure | HTTP 503 from fulfillment API | Retry with backoff |
| Rate limiting | HTTP 429 from ERP connector | Retry with backoff (see Handling Pagination and Throttling) |
| Business validation failure | Missing required field, invalid SKU | Dead-letter immediately |
| Referential integrity failure | Customer ID doesn't exist in CRM | Dead-letter with context |
| Deserialization failure | Malformed JSON, wrong encoding | Dead-letter immediately |
| Downstream permanent failure | Target system record locked by another process | Dead-letter after max retries |
In Power Automate, you implement this classification using scopes with run-after conditions. The pattern looks like this:
Scope: Process Message contains your core business logic. Within it, you have actions that can fail in categorically different ways.
Scope: Handle Transient Failure runs after "Process Message" with run-after set to "has failed" — and inside it, you check whether the error is transient or permanent using the status code from the failed action.
Scope: Dead-Letter Poison Message runs after "Handle Transient Failure" also with run-after set to "has failed" — meaning it only fires if the transient handler itself couldn't recover.
Let's look at the key expression for classifying HTTP errors. Inside your failure handling scope, after capturing the error from the failed action output, you'd use something like this to distinguish retryable from non-retryable HTTP errors:
@or(
equals(outputs('Call_Fulfillment_API')?['statusCode'], 503),
equals(outputs('Call_Fulfillment_API')?['statusCode'], 429),
equals(outputs('Call_Fulfillment_API')?['statusCode'], 502)
)
If this condition is true, you allow the flow to fail naturally (so Service Bus retries it). If it's false — meaning you got a 400, 404, 422, or similar — you take action to dead-letter the message explicitly.
Warning
A common mistake is completing the Service Bus message inside a failed scope. If you call "Complete the message in a queue" inside your error handler regardless of the error type, the message is acknowledged and removed from the queue permanently — it never reaches the DLQ. Only complete a message when processing truly succeeded. In all other cases, either let the flow fail (for retryable errors) or explicitly dead-letter it (for poison messages).
Let's build the main flow concretely. The scenario: an order integration pipeline reads order-intake-queue in a Service Bus namespace, validates the payload, and calls a fulfillment REST API.
Use the Service Bus trigger "When a message is received in a queue (peek-lock)" on order-intake-queue. Peek-lock mode is essential — it gives you control over when the message is completed or abandoned, rather than auto-completing on trigger.
Set the queue's maximum delivery count to 5 in the Azure portal (navigate to your Service Bus namespace, select the queue, then "Properties" — the Max Delivery Count field is there). This gives you five total attempts before automatic DLQ.
Add a "Parse JSON" action on the message content. Define your schema strictly — include required fields:
{
"type": "object",
"required": ["orderId", "customerId", "lineItems", "warehouseCode"],
"properties": {
"orderId": { "type": "string" },
"customerId": { "type": "string" },
"lineItems": {
"type": "array",
"items": {
"type": "object",
"required": ["sku", "quantity"],
"properties": {
"sku": { "type": "string" },
"quantity": { "type": "integer", "minimum": 1 }
}
}
},
"warehouseCode": { "type": "string" }
}
}
Wrap Parse JSON in a scope called "Scope - Validate and Process". Set this scope to run after the previous step (normal success path).
Add a second scope called "Scope - Classify and Handle Failure" with run-after set to "has failed" on Scope - Validate and Process. This is where your classification logic lives.
Inside this scope, add a Condition action:
Use this expression in the condition to check if the Parse JSON action failed (indicating malformed payload):
@equals(result('Scope_-_Validate_and_Process')?[0]?['code'], 'InvalidTemplate')
Or check if a required HTTP call returned a non-retryable error:
@or(
equals(actions('Call_Fulfillment_API')?['statusCode'], 400),
equals(actions('Call_Fulfillment_API')?['statusCode'], 422),
equals(actions('Call_Fulfillment_API')?['statusCode'], 404)
)
If True (poison message path): Call the Service Bus action "Dead-letter the message in a queue". Set these fields:
triggerBody()?['LockToken']ValidationFailure or BusinessRuleViolation@concat('Parse failure on orderId: ', triggerBody()?['MessageId'],
'. Error: ', result('Scope_-_Validate_and_Process')?[0]?['error']?['message'])
If False (transient failure path): Do nothing — let the scope fail, which abandons the lock, allowing Service Bus to retry the message on its next delivery attempt.
After Scope - Validate and Process, add a "Complete the message in a queue" action with run-after set to "is successful" only. Pass the lock token:
triggerBody()?['LockToken']
Tip
When using peek-lock on Service Bus with Power Automate, the default lock duration on the queue is 60 seconds. If your processing flow takes longer than that, the lock expires and Service Bus re-delivers the message — incrementing the delivery count even if your flow is still running successfully. For flows with lengthy processing steps, increase the lock duration in the queue settings, or use lock renewal via the Azure Service Bus REST API in a parallel branch.
Now that poison messages are landing in the DLQ, you need a way to surface them. The DLQ address for order-intake-queue is order-intake-queue/$DeadLetterQueue. Power Automate's Service Bus connector supports reading from this sub-queue directly.
Create a new scheduled cloud flow — run it every 15 minutes. Call it "DLQ Monitor - Order Intake".
Use the Service Bus action "Get messages from a queue (peek-lock)" (not the trigger — this is an action you control explicitly in a scheduled flow). Configure it for the DLQ:
order-intake-queue/$DeadLetterQueueThis returns an array of messages. If the array is empty, you can terminate early with a condition.
For each message in the result, write a record to a Dataverse table (or SharePoint list if Dataverse isn't available). Design your staging table with these columns:
| Column | Type | Purpose |
|---|---|---|
dlq_message_id |
Text | Service Bus MessageId |
dlq_payload |
Multiline Text | Original message body (base64 decoded) |
dlq_reason |
Text | DeadLetterReason from system properties |
dlq_description |
Multiline Text | DeadLetterErrorDescription |
dlq_enqueued_time |
DateTime | When it was originally enqueued |
dlq_delivery_count |
Integer | How many times delivery was attempted |
dlq_status |
Choice | Pending / Under Review / Reprocessed / Discarded |
dlq_corrected_payload |
Multiline Text | Operator-edited payload for reprocessing |
dlq_sequence_number |
Text | Service Bus sequence number (needed for DLQ complete) |
Inside the Apply to Each loop, decode the message body (Service Bus delivers content as base64):
@base64ToString(items('Apply_to_each_DLQ_message')?['ContentData'])
Extract the dead-letter properties from the broker properties:
@items('Apply_to_each_DLQ_message')?['Properties']?['DeadLetterReason']
@items('Apply_to_each_DLQ_message')?['Properties']?['DeadLetterErrorDescription']
After writing the Dataverse record, complete the DLQ message using its lock token. This removes it from the DLQ — you're taking ownership by moving it to your staging table. If you don't complete it, your monitor flow will re-read the same messages every 15 minutes and create duplicates.
Note
Completing a message from the DLQ is a terminal operation — the message is gone from Service Bus. This is why your staging table needs to capture everything you'd ever want to audit. Consider making the payload column large enough (Dataverse's default Multiline Text supports up to 1MB of text) and storing the full broker properties JSON alongside the body.
After the loop, send a notification if any new DLQ messages were processed. A Teams message or email works well here. Include a summary and a deep link to a model-driven app or SharePoint view filtered to dlq_status = Pending.
If you're forwarding telemetry to Azure Application Insights, log a custom event for each dead-lettered message. The monitoring infrastructure patterns from Building a Power Automate Monitoring and Alerting System apply directly here — use an HTTP action against the Application Insights ingestion endpoint with a structured custom event payload.
This is where the pattern gets powerful. You've got failed messages in a staging table with their original payloads, reasons, and enough context for an operator to diagnose and correct them. Now you need a safe, auditable way to push corrected messages back to the main queue.
The reprocessing flow is triggered manually — either by a Power Apps canvas app, a model-driven app command bar button, or a manual trigger HTTP webhook from a Teams adaptive card. The flow receives the Dataverse record ID of the staged DLQ message.
Key insight
Design the reprocessing flow to be idempotent. If an operator accidentally triggers it twice for the same message, you don't want to send duplicate orders to the fulfillment system. Check the dlq_status field at the start of the flow — if it's already Reprocessed, abort immediately. This is the same principle covered in Designing Idempotent Flows: Preventing Duplicate Processing Under At-Least-Once Delivery.
Retrieve the Dataverse row by ID. Check the status — if not Pending or Under Review, terminate with a success response explaining no action was taken.
Update the status to Under Review immediately (optimistic lock pattern — if two operators trigger this simultaneously, the second update will succeed but the flow will check the status again after a brief pause).
The operator has had a chance to edit dlq_corrected_payload. Before sending it to the queue, validate it. Use a child flow that does JSON schema validation — this is a great use of the child flow pattern for shared validation logic you can reuse across multiple integration flows.
If dlq_corrected_payload is empty, fall back to the original dlq_payload. This handles cases where the original message was valid but failed due to a transient system issue, and the operator just wants to requeue it without changes.
Use the Service Bus action "Send message" to order-intake-queue. Set these properties:
application/jsonreprocessedFrom with the original MessageId, and reprocessedBy with the operator's identity (available via the flow's trigger outputs if using a manual trigger or HTTP endpoint){
"reprocessedFrom": "@{variables('OriginalMessageId')}",
"reprocessedBy": "@{triggerOutputs()?['headers']?['x-ms-user-email']}",
"reprocessedAt": "@{utcNow()}"
}
These custom properties flow through to your main processing flow's trigger output, where you can log them for audit purposes.
On success, update the Dataverse record:
dlq_status → Reprocesseddlq_reprocessed_at → utcNow()dlq_reprocessed_by → operator identitydlq_reprocessed_message_id → the new Service Bus MessageId you generatedOn failure, update:
dlq_status → Pending (reset so it can be retried)One subtle poison message scenario that trips up production pipelines is schema drift — where the upstream system changes the message format without coordinating with downstream consumers. Your main flow's Parse JSON action fails because the schema no longer matches, and every new message becomes a poison message.
The safe pattern here is lenient parsing with strict validation. Instead of using Parse JSON with a strict schema (which fails hard on unexpected fields), parse the JSON with a minimal required-fields schema, then run explicit validation logic afterward.
Use the json() function in a Compose action to parse:
@json(base64ToString(triggerBody()?['ContentData']))
Then validate required fields explicitly in a condition:
@and(
not(empty(outputs('Parse_Order_Body')?['orderId'])),
not(empty(outputs('Parse_Order_Body')?['customerId'])),
greater(length(outputs('Parse_Order_Body')?['lineItems']), 0)
)
If validation fails, you know exactly which fields are missing and can write that to the DLQ reason description — making the operator's job of correcting the payload much easier than a generic "InvalidTemplate" error.
Tip
Include a schemaVersion field in your message contract and route to different validation paths based on it. This lets you deploy new schema versions without breaking processing of old messages still in the queue during a deployment window. Combined with the environment variable strategies described in Defining and Enforcing Environment Variable Strategies in Power Automate Solutions, you can toggle schema version routing without redeploying flows.
Here's a scenario distinct from poison messages but handled by similar tooling: your fulfillment API goes down for 4 hours. During that window, all messages fail with HTTP 503 (transient — correctly not dead-lettered), exhaust their 5 delivery attempts, and land in the DLQ anyway. Now you have 800 messages in the DLQ that are actually perfectly valid — they just need to be requeued.
For this case, you need a bulk reprocessing flow rather than operator-by-operator review. Build a separate scheduled or manually triggered flow that:
dlq_reason = 'MaxDeliveryCountExceeded' and dlq_status = 'Pending'Warning
When bulk reprocessing, be mindful of your Service Bus queue's throughput limits and your downstream system's rate limits. Use a Do Until loop with controlled batching rather than an unbounded Apply to Each. The batching patterns in Implementing Batching and Chunking Strategies in Power Automate apply directly here — split your 800 records into batches of 50 with a delay between batches.
Use a variable to track progress and update a summary Dataverse record with counts of succeeded/failed reprocessing attempts. If any individual reprocess fails, keep the record in Pending status and log the error — don't let one failed reprocess abort the entire bulk operation.
Your reprocessing flow's HTTP trigger (if you're using one) is a sensitive endpoint — it can inject arbitrary data into your production queue. Secure it properly.
Use managed identity connections for both the Service Bus connector and Dataverse connector in your reprocessing flow — this eliminates stored credentials entirely. The detailed setup is covered in Integrating Power Automate with Azure Key Vault and Managed Identities.
For the HTTP trigger, configure it with Azure AD authentication by setting the trigger's "Who can trigger the flow" setting to specific users or groups — not "Anyone with the link." Require callers to pass a bearer token in the Authorization header, and use the trigger's built-in caller authentication to validate it.
If you're using a model-driven app command bar button to trigger reprocessing, the Dataverse action call respects the user's security role — add a custom security role that grants access to the DLQ staging table and the Power Automate HTTP connector, and assign it only to your integration operations team.
Build the complete dead-letter handling system for a purchase order integration scenario:
Setup:
po-intake-queue. Set max delivery count to 3 and lock duration to 2 minutes.po_dlq_staging with the columns described in the DLQ monitoring section.Build these three flows:
Flow 1: PO Intake Processor
po-intake-queuepoNumber, vendorId, lineItems[], totalAmounthttps://httpbin.org/post as a stand-in)ValidationFailure and a description listing the missing fieldsFlow 2: DLQ Monitor - PO Intake
po-intake-queue/$DeadLetterQueuepo_dlq_staging, complete the DLQ messageFlow 3: PO DLQ Reprocessor
po-intake-queue with custom properties tracking provenanceTest it: Send a valid PO message and confirm it processes and completes. Then send a message with vendorId missing and confirm it lands in the DLQ, gets picked up by the monitor flow, appears in Dataverse, and can be corrected and reprocessed.
Mistake: Completing the message in an error handler This is the most common mistake. If you call "Complete the message" inside your failure-handling scope, the message is acknowledged and permanently removed — it never hits the DLQ. Only complete messages in your success path. Use the explicit dead-letter action for poison messages and let the flow fail (abandon the lock) for transient failures.
Mistake: Not accounting for lock expiry during long processing
If your flow takes longer than the queue's lock duration (default 60 seconds), Service Bus re-delivers the message while your flow is still running. You end up with duplicate processing AND an incremented delivery count. Fix this by either increasing the queue's lock duration or building a lock renewal loop using the Azure Service Bus REST API (POST /messages/{sequenceNumber}/renewlock).
Mistake: DLQ monitor flow not completing messages
If your monitor flow peeks but doesn't complete DLQ messages after staging them, it will re-read the same messages every run and create duplicate Dataverse records. Always complete DLQ messages you've taken ownership of. Add duplicate detection using dlq_message_id as a unique index on your Dataverse table.
Mistake: Using auto-complete mode on the trigger Auto-complete mode (where the message is completed as soon as the trigger fires, before your flow logic runs) makes the DLQ completely useless — every message succeeds from Service Bus's perspective regardless of what your flow does. Always use peek-lock mode in integration pipelines.
Mistake: Forgetting that DLQ messages retain delivery count When you reprocess a message by sending it back to the main queue, it starts fresh with delivery count 0 — because it's a new message. But if you're reprocessing by reading from the DLQ and not completing (abandoning back), Service Bus may move it to a deeper level of dead-lettering depending on configuration. Always send a new message to the main queue rather than trying to "move" messages back from the DLQ.
Troubleshooting: DLQ messages don't appear Check that your Service Bus queue has dead-lettering enabled — it's on by default but worth confirming. In the Azure portal, navigate to the queue and verify "Dead lettering on message expiration" is enabled. Also check the max delivery count — if it's set to very high (or 0, which means unlimited in some SDK versions), messages may never auto-dead-letter.
Troubleshooting: Flow trigger not firing for DLQ The Service Bus trigger in Power Automate fires on the main queue only — you cannot use the trigger on the DLQ sub-queue path. The DLQ monitoring pattern must be a scheduled flow using the "Get messages" action, not a trigger. This is by design: DLQ messages don't expire or need real-time processing, and a polling approach gives you batching control.
Dead-letter queue handling is what separates hobby integrations from production-grade ones. You've now built a complete system that:
The patterns here extend beyond Service Bus. Similar DLQ concepts apply to Azure Storage Queues, Azure Event Hubs (via capture and reprocessing), and even custom queue implementations backed by Dataverse or SharePoint. The classification logic, staging store design, and reprocessing flow structure all transfer directly.
Where to go next: