When your retry logic becomes the attack — that's when you need exponential backoff and circuit breakers. This lesson teaches you how to implement both patterns in Power Automate using Do Until loops, Dataverse state tracking, and child flows, so your integrations protect downstream systems instead of hammering them during outages.

It's 2:47 AM and your enterprise order processing flow has been hammering a downstream ERP API for the past four hours. Not intentionally — a transient network partition made the first hundred calls fail, and the flow's automatic retry logic kicked in. Then those retries started failing too. Then the backlog of pending orders piled up, and new flow runs launched while old ones were still retrying. By 6 AM, you have three thousand concurrent HTTP requests slamming an API that's already struggling, your ERP vendor has throttled your tenant's access key, and the cascading blast radius has taken down two other integrations that share the same service account. Congratulations, your retry logic just caused a distributed denial-of-service attack against your own ERP system.
This is not a hypothetical. It's a pattern that plays out regularly in enterprise Power Automate deployments when flows are built for the happy path without explicit resilience engineering. The good news is that two patterns — exponential backoff and the circuit breaker — solve exactly this class of problem. Exponential backoff prevents retry storms by spacing out retry attempts with increasing delays. The circuit breaker prevents sustained hammering by detecting persistent failures early, stopping requests entirely for a cooling-off period, and only allowing traffic to resume once the downstream system shows signs of recovery. Together, they transform your flows from a liability during outages into a well-behaved integration partner.
By the end of this lesson, you'll have a fully working implementation of both patterns inside Power Automate using a combination of Do Until loops, Compose actions, and Dataverse state tracking. You'll understand the theory well enough to adapt it to any connector or API you're integrating with.
What you'll learn:
This lesson assumes you're comfortable with:
You don't need prior experience with distributed systems patterns, though the terminology will map cleanly to what you already know.
Power Automate's built-in retry policy (configurable on most connector actions under Settings) is better than nothing. Set it to "Fixed interval" with four retries at 60-second intervals and it'll behave reasonably for brief transient errors. But it has two fundamental weaknesses that make it dangerous at scale.
Problem one: Fixed-interval retries create synchronized thundering herds. Imagine you have 200 concurrent flow runs, all of which hit a rate-limit response (HTTP 429) from the same API at roughly the same time. With a fixed 60-second retry, all 200 runs wait exactly 60 seconds, then simultaneously hammer the API again. If it's still struggling, they all fail again, wait 60 seconds, and retry in unison again. You've turned a brief capacity problem into a sustained synchronized storm.
Problem two: Built-in retry has no cross-run awareness. Each flow run retries independently, with no knowledge of whether other runs are also failing. If an API is completely down for two hours, every single run that launches during that window will go through its full retry cycle, consuming flow run quotas and adding load to a system that's already down. There's no mechanism to say "we've detected persistent failure — stop all traffic until the system recovers."
Exponential backoff solves the first problem by randomizing and increasing delays so that retries naturally spread out over time. The circuit breaker solves the second problem by introducing shared state that all flow runs can read before deciding whether to even attempt a call.
Key insight
These two patterns are complementary, not alternatives. Exponential backoff protects individual call sequences from creating local retry storms. The circuit breaker provides tenant-wide protection by sharing failure state across all concurrent and future runs. You need both.
Exponential backoff works by multiplying the wait time by a base factor after each failed attempt. The standard formula is:
wait_seconds = base_delay * (multiplier ^ attempt_number) + jitter
Where jitter is a random value (typically 0–base_delay seconds) added to prevent synchronized retries. Let's put numbers to it. With a base delay of 2 seconds, a multiplier of 2, and up to 5 attempts:
| Attempt | Wait (no jitter) | Wait with up to 2s jitter |
|---|---|---|
| 1 | 2s | 2–4s |
| 2 | 4s | 4–6s |
| 3 | 8s | 8–10s |
| 4 | 16s | 16–18s |
| 5 | 32s | 32–34s |
After five attempts, you've waited a total of 62–72 seconds with natural spread between concurrent callers. This is far more graceful than five immediate retries or five synchronized retries at fixed intervals.
For API rate limit scenarios where the Retry-After header is present in the 429 response, you should honor that value rather than your calculated delay. We'll cover that in the implementation.
Disable the built-in retry policy on any action you're wrapping in a custom backoff loop. Go to the action's Settings menu, set Retry Policy to None. If you leave it enabled, you'll have two nested retry mechanisms fighting each other.
Here's the complete structure of the backoff loop. Initialize these variables at the top of your flow:
Initialize variable: varAttempt (Integer) = 0
Initialize variable: varMaxAttempts (Integer) = 5
Initialize variable: varBaseDelay (Integer) = 2
Initialize variable: varSuccess (Boolean) = false
Initialize variable: varLastStatusCode (Integer) = 0
Initialize variable: varResponseBody (String) = ""
Now add a Do Until loop with this condition:
Or(
equals(variables('varSuccess'), true),
greaterOrEquals(variables('varAttempt'), variables('varMaxAttempts'))
)
Inside the loop, the sequence is: increment attempt counter, make the API call, evaluate the result, calculate delay if failed, wait.
Step 1 — Increment the counter:
Set variable: varAttempt = add(variables('varAttempt'), 1)
Step 2 — Make the HTTP call:
Add your HTTP action or connector call here. Set Run After to run on Success only — we'll handle failures in subsequent steps. Configure the action with name HTTP_CallERP and disable its retry policy.
Step 3 — Handle the response with a Condition:
Create a condition branching on whether the action succeeded. Power Automate surfaces this via the action's status. Use a Scope action to wrap the HTTP call and use result('Scope_ERPCall')[0]['status'] to check the outcome, or check the HTTP status code directly:
// In the "Yes" (success) branch:
Set variable: varSuccess = true
Set variable: varResponseBody = body('HTTP_CallERP')
// In the "No" (failure) branch:
Set variable: varLastStatusCode = outputs('HTTP_CallERP')['statusCode']
Step 4 — Calculate the backoff delay (inside the failure branch):
Power Automate's expression engine doesn't have a native power/exponent function, so you need to simulate it. For up to 5 attempts, the cleanest approach is a Switch action on varAttempt with hardcoded delay values, which avoids floating-point issues:
Switch on: variables('varAttempt')
Case 1: Set varDelaySeconds = 2
Case 2: Set varDelaySeconds = 4
Case 3: Set varDelaySeconds = 8
Case 4: Set varDelaySeconds = 16
Case 5: Set varDelaySeconds = 32
Default: Set varDelaySeconds = 60
For jitter, add a random integer between 0 and the base delay:
Set variable: varDelaySeconds = add(variables('varDelaySeconds'), rand(0, variables('varBaseDelay')))
Step 5 — Check for Retry-After header:
If the HTTP response is a 429, the API may have told you exactly how long to wait. Don't ignore that:
Condition: equals(variables('varLastStatusCode'), 429)
Yes:
Condition: not(empty(outputs('HTTP_CallERP')?['headers']?['Retry-After']))
Yes: Set varDelaySeconds = int(outputs('HTTP_CallERP')['headers']['Retry-After'])
Step 6 — Wait:
Add a Delay action using the ISO 8601 duration format. Power Automate's Delay action takes a duration like PT30S (30 seconds). Build it dynamically:
concat('PT', string(variables('varDelaySeconds')), 'S')
Warning
The Do Until loop in Power Automate has a default limit of 60 iterations and a default timeout of 60 minutes. For backoff loops, set the iteration limit to your varMaxAttempts value (5 or 6) to prevent runaway loops. Configure this in the loop's Settings panel.
Step 7 — After the loop, check the outcome:
Condition: equals(variables('varSuccess'), false)
Yes: // Handle permanent failure
// Log to Application Insights
// Route to dead-letter queue
// Send alert
This gives you a complete, self-contained exponential backoff implementation that any flow can use. For reusable scenarios, consider extracting this into a child flow as described in Orchestrating Child Flows and Scoped Execution in Power Automate, passing the endpoint URL, headers, and body as inputs and returning the response body and status as outputs.
Tip
For high-throughput flows that run many API calls in parallel, pair this backoff implementation with the concurrency controls in Parallel Branching and Concurrency Control in Power Automate. Backoff within each branch prevents local storms; concurrency limits prevent the parallel branches themselves from overwhelming the target API.
The circuit breaker is a state machine with three states:
The breaker transitions between these states based on rules you configure:
The circuit breaker needs shared state that persists across flow runs. Dataverse is the right choice here — it's transactional, available to all flows in your environment, and already part of your Power Platform subscription.
Create a Dataverse table called Circuit Breaker State with these columns:
| Column Name | Type | Notes |
|---|---|---|
| Name (primary) | Text | Unique identifier for the breaker (e.g., "ERP_OrderAPI") |
| cr_state | Choice | Options: Closed, Open, HalfOpen |
| cr_failurecount | Integer | Count of failures in current window |
| cr_lastfailuretime | DateTime | When the last failure occurred |
| cr_openedtime | DateTime | When the breaker transitioned to Open |
| cr_failurethreshold | Integer | Max failures before opening (e.g., 5) |
| cr_cooldownminutes | Integer | Minutes to stay Open before going HalfOpen (e.g., 5) |
| cr_windowminutes | Integer | Rolling window for failure counting (e.g., 2) |
Note
Store one row per downstream system (or per API endpoint, if you want finer granularity). Your ERP order API gets one row, your SharePoint connector gets another. This lets you trip the ERP breaker without affecting SharePoint flows.
Pre-populate rows for each circuit you want to protect. Set initial state to Closed, failure count to 0, failure threshold to 5, cooldown to 5, and window to 2.
Build a dedicated child flow called CB_Guard_ERP_OrderAPI. Parent flows call this before making any ERP API call. The child flow reads the current circuit state and either returns "proceed" or "rejected" without making any API call itself.
Input parameters:
CircuitName (String) — e.g., "ERP_OrderAPI"
Output parameters:
Allowed (Boolean)
CurrentState (String)
Inside the child flow:
Step 1 — Read the circuit row:
List rows (Dataverse):
Table: Circuit Breaker States
Filter: cr_name eq 'ERP_OrderAPI'
Top count: 1
Extract the row into composed values:
Compose: State = first(outputs('List_CircuitState')?['body/value'])?['cr_state']
Compose: FailureCount = first(outputs('List_CircuitState')?['body/value'])?['cr_failurecount']
Compose: OpenedTime = first(outputs('List_CircuitState')?['body/value'])?['cr_openedtime']
Compose: CooldownMinutes = first(outputs('List_CircuitState')?['body/value'])?['cr_cooldownminutes']
Step 2 — Branch on state:
Switch on: outputs('Compose_State')
Case "Open":
// Check if cool-down has expired
Compose: ElapsedMinutes = div(
sub(
ticks(utcNow()),
ticks(outputs('Compose_OpenedTime'))
),
600000000 // ticks per minute
)
Condition: greaterOrEquals(outputs('Compose_ElapsedMinutes'), outputs('Compose_CooldownMinutes'))
Yes (transition to Half-Open):
Update row (Dataverse): cr_state = HalfOpen
Set output Allowed = true
Set output CurrentState = "HalfOpen"
No (still cooling down):
Set output Allowed = false
Set output CurrentState = "Open"
Case "HalfOpen":
// Only allow one probe — this is a race condition risk in high-concurrency scenarios
// See note below about probe locking
Set output Allowed = true
Set output CurrentState = "HalfOpen"
Default (Closed):
Set output Allowed = true
Set output CurrentState = "Closed"
Warning
In high-concurrency environments, multiple flow runs may simultaneously read "HalfOpen" and all attempt probe requests. This defeats the purpose of Half-Open, which is to allow a single test request. To handle this properly, immediately transition state to a "Probing" sentinel value when the first run reads HalfOpen, so subsequent reads return "Open" (rejected) until the probe resolves. This requires a read-then-update pattern with short Dataverse transactions and is worth implementing for flows with more than 5 concurrent runs.
Build a second child flow called CB_Report_ERP_OrderAPI that flow runs call after each API attempt to report success or failure.
Input parameters:
CircuitName (String)
CallSucceeded (Boolean)
Inside the child flow:
Step 1 — Read current state (same as above)
Step 2 — Branch on call result:
If CallSucceeded = true:
Switch on: CurrentState
Case "HalfOpen":
// Probe succeeded — close the circuit
Update row (Dataverse):
cr_state = Closed
cr_failurecount = 0
Case "Closed":
// Normal success — optionally reset failure count if window expired
Compose: WindowExpired = greaterOrEquals(
div(sub(ticks(utcNow()), ticks(LastFailureTime)), 600000000),
WindowMinutes
)
Condition: WindowExpired
Yes: Update row: cr_failurecount = 0
If CallSucceeded = false:
Switch on: CurrentState
Case "HalfOpen":
// Probe failed — reopen the circuit
Update row (Dataverse):
cr_state = Open
cr_openedtime = utcNow()
Case "Closed":
// Increment failure count
Compose: NewCount = add(FailureCount, 1)
Condition: greaterOrEquals(outputs('Compose_NewCount'), FailureThreshold)
Yes: // Trip the breaker
Update row (Dataverse):
cr_state = Open
cr_failurecount = NewCount
cr_openedtime = utcNow()
cr_lastfailuretime = utcNow()
No: // Just increment
Update row (Dataverse):
cr_failurecount = NewCount
cr_lastfailuretime = utcNow()
Here's how a parent flow that processes ERP orders uses both child flows and the backoff implementation:
Trigger: When a new order arrives (HTTP trigger or Dataverse trigger)
// 1. Check circuit breaker
Run Child Flow: CB_Guard_ERP_OrderAPI
Input: CircuitName = "ERP_OrderAPI"
// 2. Branch on guard result
Condition: equals(outputs('CB_Guard')?['body/Allowed'], true)
YES branch (circuit is Closed or HalfOpen):
// 3. Initialize backoff variables
Initialize varAttempt = 0, varSuccess = false, etc.
// 4. Exponential backoff Do Until loop
Do Until: varSuccess = true OR varAttempt >= 5
Increment varAttempt
HTTP: POST to ERP /orders endpoint
Condition: action succeeded AND statusCode = 201
Yes:
Set varSuccess = true
Set varCallSucceeded = true
No:
Set varCallSucceeded = false
Calculate and apply backoff delay
// 5. Report outcome to circuit breaker
Run Child Flow: CB_Report_ERP_OrderAPI
Input: CircuitName = "ERP_OrderAPI"
Input: CallSucceeded = variables('varCallSucceeded')
// 6. Handle final outcome
Condition: varSuccess = false
Yes: Route to dead-letter queue
NO branch (circuit is Open — system known to be down):
// Skip the API call entirely
// Route order to dead-letter queue immediately
// Optionally: send a single notification (not one per order)
Run Child Flow: DLQ_Route_Order
Input: Order = triggerBody()
Input: Reason = "Circuit open — ERP unavailable"
This structure means that during an outage, orders don't pile up in retry queues consuming flow run quota. They're immediately routed to the dead-letter queue for later reprocessing, and the ERP system gets zero traffic from Power Automate until the circuit half-opens and a probe succeeds.
For dead-letter queue implementation, see Implementing Dead-Letter Queue Handling and Poison Message Recovery in Power Automate, which covers the reprocessing side of this pattern in detail.
A circuit breaker that trips silently in production is only marginally better than having no circuit breaker at all. You need visibility into when breakers trip, how often backoff engages, and what the actual failure rates look like over time.
Every time the circuit breaker reporter updates the Dataverse row with a state change (Closed → Open, HalfOpen → Closed, etc.), emit a telemetry event to Application Insights. Use the HTTP action with the Application Insights Data Collector API:
HTTP POST: https://dc.services.visualstudio.com/v2/track
Headers:
Content-Type: application/json
Body:
{
"name": "Microsoft.ApplicationInsights.Event",
"iKey": "@{variables('AppInsightsKey')}",
"time": "@{utcNow()}",
"data": {
"baseType": "EventData",
"baseData": {
"name": "CircuitBreakerStateChange",
"properties": {
"CircuitName": "ERP_OrderAPI",
"PreviousState": "@{outputs('Compose_PreviousState')}",
"NewState": "@{outputs('Compose_NewState')}",
"FailureCount": "@{outputs('Compose_FailureCount')}",
"FlowRunId": "@{workflow()['run']['name']}",
"Environment": "@{parameters('EnvironmentName')}"
}
}
}
}
For retrieving your Application Insights key securely at runtime, see Integrating Power Automate with Azure Key Vault and Managed Identities.
Inside the backoff loop's failure branch, log each retry attempt:
HTTP POST to Application Insights:
{
"name": "Microsoft.ApplicationInsights.Event",
"data": {
"baseData": {
"name": "BackoffRetryAttempt",
"properties": {
"Circuit": "ERP_OrderAPI",
"AttemptNumber": "@{variables('varAttempt')}",
"StatusCode": "@{variables('varLastStatusCode')}",
"DelaySeconds": "@{variables('varDelaySeconds')}",
"IsRateLimitResponse": "@{equals(variables('varLastStatusCode'), 429)}",
"FlowRunId": "@{workflow()['run']['name')}"
}
}
}
}
With this telemetry in place, you can build Application Insights workbooks or Azure Monitor alerts that fire when:
For a complete monitoring and alerting architecture, see Building a Power Automate Monitoring and Alerting System.
Key insight
Use the workflow()['run']['name'] expression to capture the flow run ID in every telemetry event. This lets you correlate Application Insights events back to specific Power Automate run history entries when you need to investigate. This is a foundational practice discussed in Implementing Correlation IDs and End-to-End Distributed Tracing.
The right threshold values depend on your specific integration characteristics. Here are starting-point guidelines:
Failure threshold (how many failures before opening):
Cooldown period (how long to stay open):
Failure window (rolling window for counting failures):
Maximum backoff attempts:
Don't hardcode threshold values inside the flow. Store them as environment variables or Dataverse table configuration (the cr_failurethreshold, cr_cooldownminutes, and cr_windowminutes columns we created serve this purpose). This lets you tune thresholds without redeploying flows.
For an enterprise-grade approach to externalizing this kind of configuration, see Defining and Enforcing Environment Variable Strategies in Power Automate Solutions.
Tip
Add a "manual override" column to your Circuit Breaker State table — a boolean called cr_forceopen. When set to true, the guard flow always returns Allowed = false regardless of the actual state machine state. This gives operators a manual kill switch they can flip during planned maintenance windows without having to edit or disable flows.
Build a complete working implementation for an imaginary CRM-to-ERP order sync integration. Here's the specification:
Scenario: Your organization's Dynamics 365 Sales (CRM) creates orders that need to be pushed to a legacy ERP API. The ERP API has a rate limit of 100 requests per minute, tends to have planned maintenance windows on Sunday mornings, and occasionally falls over during month-end processing peaks.
Build the following:
Dataverse table setup: Create the Circuit Breaker State table with all columns defined in Part 2. Insert one row for ERP_OrderAPI with: threshold = 5, cooldown = 10 minutes, window = 2 minutes, initial state = Closed.
CB_Guard child flow: Implement the guard flow as described. Add console-style logging (Compose actions with labeled strings) so you can trace the execution path in run history.
CB_Report child flow: Implement the reporter flow with all state transitions.
Backoff loop: Create a standalone flow that accepts an order payload (JSON with orderId, customerId, amount, lineItems), guards via CB_Guard, then attempts to POST to https://httpbin.org/status/429 (which always returns 429 — simulating rate limiting). Verify that your backoff delays correctly increase and the loop exits after max attempts.
Trigger the circuit: Modify the test target URL to https://httpbin.org/status/500. Run the flow ten times in succession with a concurrent limit of 10. Verify that:
Instrument it: Add Application Insights logging to state transitions and backoff retries. Query the custom events in Application Insights to confirm your telemetry is arriving.
If you implement a custom backoff loop but forget to set the action's Retry Policy to None, you'll get both mechanisms firing. The built-in retry fires first (up to 4 times by default), then your Do Until loop retries on top of that. You can end up with 20+ attempts per call instead of 5. Always verify the Settings panel on every action inside a custom retry loop.
The failure count should reset when the window period expires, even without reaching the threshold. If you don't reset it, a system that had 4 failures at 9 AM and then ran perfectly until 5 PM will open the breaker on the very next failure at 5 PM, even though that failure is completely unrelated to the morning's problems. The reset logic in the Reporter flow (checking if the window has expired on a successful call) handles this — don't skip it.
When calculating elapsed time since the circuit opened, use the ticks() function to convert both datetimes to integers, subtract, then divide by the ticks-per-minute value (600,000,000). Don't use dateDifference() or string manipulation on datetime values — those approaches break at timezone boundaries and introduce subtle bugs that are extremely difficult to diagnose.
// Correct approach:
div(
sub(ticks(utcNow()), ticks(variables('varOpenedTime'))),
600000000
)
// Problematic approach:
// Comparing formatted datetime strings or using addMinutes() chains
The circuit breaker's value comes entirely from sharing state across runs. If you initialize the breaker state inside the flow as variables (instead of reading from Dataverse), each run has its own independent breaker that knows nothing about other runs. You'll think you have circuit breaker protection but every run will independently retry against a downed system. The Dataverse table is non-negotiable.
Application Insights HTTP calls can occasionally timeout or return 5xx responses themselves. If you log inside the backoff loop and the logging call fails without error handling, it can exit your Do Until loop prematurely (because the action status drives the flow branching). Wrap all telemetry calls in a Scope with Run After configured to run on both Success and Failure, and set the logging HTTP action's Retry Policy to None.
Warning
Be careful with Dataverse rate limits when using the circuit breaker at very high concurrency. If you have 500 concurrent flow runs all reading the Circuit Breaker State table simultaneously, those reads consume Dataverse API request capacity. For extremely high-volume scenarios, consider caching the circuit state in a short-lived Azure Cache for Redis instance and using an Azure Function as an intermediary — this is overkill for most Power Automate deployments but worth knowing about.
A circuit breaker that's open for 10 minutes might be read by 200 flow runs during that window. Don't send an alert email or Teams message for each read — you'll create an alert storm on top of the outage. Send one alert when the breaker transitions to Open (in CB_Report, not CB_Guard), and send one recovery notification when it transitions back to Closed.
You now have the theoretical foundation and complete implementation details for two of the most important resilience patterns in distributed systems, adapted specifically for Power Automate's execution model.
What you built:
Key principles to remember:
Where to go next:
For scenarios where the downstream system being protected is a high-volume message bus rather than a synchronous API, the circuit breaker pattern integrates naturally with event-driven architectures. See Implementing Event-Driven Automation with Power Automate and Azure Service Bus for how queue-based decoupling provides an additional layer of resilience beyond what circuit breakers offer.
For organizations integrating with APIs that publish OpenAPI specifications, Custom Connectors and HTTP Actions for Production Integration covers how to implement custom connector policies — a complementary approach to rate limit management that works at the connector layer rather than the flow layer.
Finally, if you're deploying this pattern across multiple environments as part of a managed solution, make sure your Dataverse table schema and initial circuit state configuration are packaged as part of your solution and properly promoted through dev/test/production. The ALM considerations for stateful Dataverse configurations are covered in Deploying and Managing Power Automate Solutions Across Environments.