A single gateway machine is a single point of failure waiting to happen. This deep-dive lesson teaches you how to design, deploy, and operate On-Premises Data Gateway clusters for true high availability, from the Azure Service Bus relay architecture through load balancing algorithms, rolling upgrades, and production monitoring patterns.

Picture this: it's 2:47 AM on a Tuesday, and your organization's nightly ETL process—the one that pulls financial data from an on-premises SQL Server, transforms it, and loads it into Dataverse for the morning executive dashboard—silently fails. Not because of a flow logic error or an API throttle, but because the single gateway machine in the data center rebooted after a Windows Update. By the time the ops team notices at 8:15 AM, the CFO is looking at yesterday's numbers and nobody knows why. This is the exact scenario that On-Premises Data Gateway clusters are designed to prevent.
The On-Premises Data Gateway is the bridge between Power Automate's cloud infrastructure and your organization's internal data sources—SQL Server, Oracle, SAP, file shares, legacy line-of-business systems. For flows that never touch on-premises resources, you may never think about gateways at all. But in enterprise environments, the gateway is often the single most critical piece of infrastructure in your automation stack, and most organizations run it wrong. They install one gateway on one machine, give it a service account, and call it a day. That works fine until it doesn't—and in production, "until it doesn't" is a guarantee, not a possibility.
By the end of this lesson, you'll understand the full architecture of gateway clusters, how load balancing actually works under the hood, how to size and deploy a production-grade cluster, and how to monitor it like a professional. You'll be able to design a gateway topology that survives node failures, maintenance windows, and traffic spikes without your flows ever noticing.
What you'll learn:
You should be comfortable with Power Automate cloud flows at a production level. Familiarity with connection references and how flows use connectors will be assumed—if you want a refresher on the credential and connection model, see Securing Power Automate Flows in Production: Managing Credentials, Connection References, and Data Loss Prevention Policies. You should also understand how enterprise flows are deployed across environments, since gateway configuration is tightly coupled to solution and environment strategy—covered in Deploying and Managing Power Automate Solutions Across Environments: ALM Pipelines, Solution-Aware Flows, and Environment Variables for Enterprise-Scale Delivery. Basic Windows Server administration knowledge is assumed.
Before you can design a resilient cluster, you need to understand what the gateway actually does and why a clustered architecture solves the problems it does. Most documentation glosses over this and jumps straight to installation steps. We're not going to do that.
The On-Premises Data Gateway does not open inbound ports on your firewall. This is a common misconception and an important security property. Instead, the gateway establishes outbound connections to Azure Service Bus, and the Power Automate cloud infrastructure communicates with your on-premises resources by sending messages through that relay.
When a flow action calls an on-premises connector—say, "Execute a SQL query" against your internal SQL Server—here's what actually happens:
The entire exchange happens through encrypted outbound HTTPS and WebSocket connections over port 443. Your firewall never needs an inbound rule for gateway traffic.
Key insight
Because the gateway acts as a Service Bus relay consumer, the cloud runtime doesn't "call" the gateway directly—it publishes a message and waits for a response on a correlated reply channel. This relay architecture is what enables multiple gateway nodes to participate in a cluster: they're all listening on the same relay, and whichever node picks up the message handles the request.
Understanding the connection reference model is critical for cluster deployments. A connection in Power Automate's terminology is a stored credential set bound to a connector type—it includes the authentication information (username/password, service account credentials, etc.) and, critically for on-premises connectors, a reference to a specific gateway resource. That gateway resource in the Power Platform is what maps to a cluster.
When you deploy solutions across environments, connection references act as the injection point for environment-specific configuration. Your dev environment might point to a gateway cluster with two nodes in a development VLAN, while production points to a five-node cluster in a hardened network segment. The flow logic doesn't change; only the connection reference target changes. This is why environment strategy and gateway cluster design need to be planned together—they're two halves of the same deployment model.
The term "high availability" gets thrown around loosely. For gateway clusters specifically, HA means two things:
Node failure tolerance: If one gateway node becomes unavailable—due to a crash, a rebooting update, or a network partition—the cluster continues to serve requests using the remaining nodes. Power Automate's cloud runtime detects that the node is no longer responding on the relay and routes subsequent requests to healthy nodes.
Maintenance window support: You can take individual nodes offline for patching, hardware maintenance, or reconfiguration without interrupting running flows. As long as at least one node remains healthy, the cluster stays operational.
What gateway clustering does not provide is automatic horizontal scaling in response to load spikes in the way that, say, Azure Functions consumption plan does. The nodes in your cluster are fixed; adding capacity means adding nodes, which is a deliberate administrative action, not an elastic response. If you need elastic, serverless compute for compute-heavy processing steps, that workload belongs in Azure Functions integrated with Power Automate, not in the gateway layer.
Before you touch an installer, size your nodes correctly. The gateway is a .NET application that acts as a local proxy—it doesn't do heavy compute itself, but it does hold concurrent connections open, buffer data in memory, and serialize/deserialize potentially large result sets. Here are the practical guidelines:
Minimum production node spec:
Network considerations: Each gateway node needs outbound access to:
*.servicebus.windows.net (port 443 and 5671 for AMQP fallback)*.frontend.clouddatahub.netlogin.microsoftonline.com and associated Azure AD endpointsIf your organization uses a proxy server for outbound internet access, the gateway has proxy configuration support—but be aware that some enterprises have proxy configurations that terminate and re-sign TLS connections, which can cause issues with the gateway's certificate validation. Test this early.
Warning
Do not install the gateway on the same machine as your data source. A SQL Server under heavy query load will starve the gateway process of CPU and memory. Treat gateway nodes as dedicated infrastructure, not add-ons to existing servers. This is the single most common misconfiguration we see in enterprise environments.
How many nodes?
For HA alone (survive one node failure), you need at minimum two nodes. But in production you should plan for three, for this reason: if you have two nodes and one goes down for patching, you're running with zero redundancy during the maintenance window. With three nodes, a planned maintenance window still leaves two-node redundancy.
For high-throughput workloads, sizing is empirical. A single gateway node can typically handle 10-20 concurrent requests to a reasonably fast SQL Server before you start seeing queuing delays. If your peak concurrent flow executions number in the hundreds, you need a corresponding number of nodes—or, more likely, you need to rethink whether gateway-mediated queries should be restructured to reduce per-request round trips.
Download the On-Premises Data Gateway installer from the official Microsoft download page. Run it on the first machine that will become your primary cluster node.
During installation, sign in with a work or school account that has Power Platform admin rights (or at minimum Environment Maker rights in the target environment, plus gateway admin permissions). Accept the gateway region—this is important. The gateway region must match the Power Platform environment region. If your Power Platform tenant is in the East US region, your gateway nodes must register to the East US gateway service. Mismatched regions cause connection failures that look inexplicable if you don't know to check this.
After installation, the gateway configuration screen will ask whether you want to register a new gateway or restore/migrate an existing one. Choose Register a new gateway on this computer.
Give the gateway a meaningful name. This name becomes the cluster name. Use a naming convention that encodes environment and purpose, for example:
PROD-OnPremGW-EastUS-01
Set a recovery key. Store this key in your organization's secret management system (not in a shared spreadsheet). The recovery key is required to add member nodes to the cluster and to restore the gateway configuration after hardware replacement. Losing the recovery key doesn't make the cluster permanently unrecoverable, but it makes maintenance significantly harder.
Tip
Store your gateway recovery key in Azure Key Vault alongside your other production secrets. If your team is using managed identities and Key Vault for secrets management for flow credentials, it makes sense to store gateway administration secrets in the same vault and document the retrieval process in your runbooks.
After the primary node registers, verify it appears in the Power Platform Admin Center under Data > Gateways. You should see it listed with a green status indicator and the region correctly identified.
On each additional machine, install the gateway software using the same process—but at the registration step, instead of registering a new gateway, select Add to an existing cluster.
You'll be prompted for:
The installer contacts the Azure gateway service, validates the recovery key, and registers this machine as a member node of the cluster. After registration, refresh the Power Platform Admin Center gateway view—you should see the cluster node count increase and the new member listed.
Repeat this process for each additional node. All nodes can be added in parallel; there's no dependency ordering after the primary is established.
Note
The "primary" node in a gateway cluster is not a hot standby master in the traditional sense. The primary is really just the first node registered and the one through which cluster configuration is administered. For request routing purposes, all nodes (including the primary) are peers. If the primary node goes down, the cluster continues operating normally on remaining members—it just means you'll need to do certain administrative operations from the Power Platform Admin Center rather than from the primary node's local configuration UI.
Once all nodes are registered, open Power Platform Admin Center and navigate to Data > Gateways. Select your cluster. You'll see a cluster-level view with several important settings:
Allow user's cloud datasources to refresh through this cluster: Enable this to allow Power BI and Power Platform resources to route through the cluster. Keep it enabled unless you have a specific reason to restrict it.
Distribute requests across all active nodes in this cluster: This is the load balancing toggle. Enable it. Without this setting, the cluster operates in a primary-preferred mode where the primary node handles all requests unless it fails, defeating most of the throughput benefit of having multiple nodes.
Gateway contact information: Fill in owner contact details. This is surfaced in monitoring and admin views across the tenant.
Under each individual node in the cluster, you can configure:
The per-node concurrency limit defaults to a value based on detected CPU count. For a 4-core node, the default is typically 10 concurrent queries. For 8 cores, around 20. You can override this, but do so conservatively—setting it too high will cause the gateway node to accept more requests than it can process efficiently, leading to latency spikes rather than the increased throughput you're after.
This is where most documentation stops at a hand-wavy "requests are distributed across nodes" and leaves you to figure out the rest. Let's be more precise.
When a Power Automate flow triggers a request that needs to go through a gateway cluster, the Azure relay service needs to select which cluster node will handle it. The algorithm works as follows:
Availability check: The relay service first filters to nodes that are currently connected (have an active outbound relay connection). Nodes that are offline, disabled, or not responding drop out immediately.
Capacity check: From the remaining available nodes, nodes that have reached their maximum concurrent query limit are excluded for that request.
Load-based selection: Among the eligible nodes, the relay service selects the node with the lowest current load—specifically, the one with the most headroom relative to its configured maximum concurrency.
Fallback: If all nodes are at capacity, the request queues briefly. If queuing exceeds the timeout threshold (approximately 10 seconds by default), the request fails with a gateway timeout error.
This is a round-robin-plus-capacity algorithm, not pure round-robin. Two implications fall out of this that matter for your architecture:
Implication 1: Node homogeneity matters. If your cluster has two 8-core nodes and one 4-core node, and you've set concurrency limits proportionally, the 4-core node will run saturated more often than the larger nodes. From the outside, this looks like intermittent latency spikes on about one-third of requests. Make your nodes as uniform as possible.
Implication 2: Slow data sources create head-of-line blocking. If your flows are querying a sluggish Oracle database and each query takes 30 seconds, those long-running queries hold concurrency slots for their entire duration. Ten concurrent 30-second queries can saturate a node entirely, causing fast requests (say, a simple SQL Server lookup) to queue behind them. If you have mixed workloads with very different latency profiles, consider separate gateway clusters for each workload type and bind connections accordingly.
Key insight
The gateway load balancing algorithm is capacity-based, not request-size-based. It doesn't know that one query will return 100,000 rows and another will return 5. It allocates based on concurrency slots, not data volume. This means that data-heavy queries, even though they may complete in the same wall-clock time as light queries, consume bandwidth on the gateway node for longer and can cause memory pressure. Factor expected result set sizes into your node sizing, not just concurrent query counts.
Certain connector types require multiple sequential requests that share state—specifically, connectors that use cursors, streaming queries, or multi-step transactions. The gateway handles this through session affinity: once a stateful session begins on a particular node, subsequent requests in that session are routed to the same node for the duration of the session.
For most Power Automate use cases—single-shot SQL queries, stored procedure calls, file system reads—this doesn't matter. But if you're building flows that use pagination patterns across large datasets through an on-premises connector that maintains cursor state, be aware that the flow's successive "next page" requests will be pinned to the node that handled the first request. If that node becomes unavailable mid-pagination, the paginated session fails and must restart from the beginning.
Design your flows to be stateless and restartable at the gateway layer where possible. Prefer limit/offset pagination over cursor-based pagination for on-premises data sources.
The gateway Windows service runs under a local service account by default. For production environments, replace this with a domain service account. This gives you:
The service account needs:
db_owner when db_datareader plus execute on specific stored procedures will do)Configure the service account through the Windows Services console (services.msc) on each gateway node after installation. Set the service to restart automatically on failure, with a failure action of "restart the service" for the first and second failures, and "run a program" (a notification script) for subsequent failures.
The gateway encrypts credentials stored locally using a machine key derived partly from the recovery key you set during installation. This encryption applies to data source credentials that are stored on the gateway node itself—which is relevant when you're configuring direct connections through the gateway rather than connections that authenticate through the Power Platform cloud connector model.
Warning
If you need to replace a gateway node due to hardware failure without a graceful decommission, you cannot simply install the gateway software on new hardware and expect credentials to transfer automatically. You'll need to use the recovery key to re-register the replacement node and re-enter data source credentials for any locally stored credentials. This is another strong argument for storing all credentials in the Power Platform connection model (where they're encrypted in the cloud) rather than as local gateway data sources where possible.
A gateway cluster you can't observe is a reliability liability. Here's how to build proper observability.
Each gateway node writes detailed logs to %LocalAppData%\Microsoft\On-premises data gateway\GatewayErrors.log for error-level events and GatewayInfo.log for informational events. These logs are verbose and rotate. For production monitoring, you should ship these logs to a central aggregation system.
The most useful log entries to monitor:
Relay connection status: Look for log entries containing "DM.EnterpriseGateway" with keywords "Connected" and "Disconnected." Frequent disconnects and reconnects indicate network instability between the gateway node and Azure Service Bus.
Query execution times: Log entries with "QueryExecutionTime" show how long the gateway spent executing each query. If you see execution times climbing over time for the same query patterns, it may indicate database performance degradation rather than a gateway problem.
Concurrency rejections: Entries containing "MaxConcurrentConnectionCount" indicate requests that were rejected because the node was at capacity. Frequent rejections mean you need more nodes or higher per-node limits.
The Admin Center's gateway view provides cluster-level health metrics updated approximately every 5 minutes:
This dashboard is useful for spot checks but not for automated alerting. You need a monitoring flow or external tooling for that.
This is a pattern worth implementing: a scheduled flow that checks gateway cluster health via the Power Platform Admin API and alerts your operations team when nodes are degraded.
The Power Platform admin APIs expose gateway cluster information at https://api.powerapps.com/providers/Microsoft.PowerApps/gateways. With a service principal that has Power Platform admin rights, you can query this endpoint on a schedule and evaluate node health.
GET https://api.powerapps.com/providers/Microsoft.PowerApps/gateways
?api-version=2016-11-01
Authorization: Bearer {token}
The response includes each gateway cluster with its members and their individual status. A scheduled flow (running every 5 minutes) can parse this response, filter for nodes where status is not Connected, and send Teams notifications or write to Application Insights when degradation is detected.
For a full treatment of building monitoring flows for your Power Automate infrastructure, see Building a Power Automate Monitoring and Alerting System: Detecting Failed Flows, Notifying Owners, and Logging Telemetry to Azure Application Insights—the pattern applies directly to gateway health monitoring with minor adaptation.
Tip
Write your gateway health monitoring as a flow that uses the HTTP connector (not an on-premises connector) to call the admin API. This avoids the circular dependency where your gateway health monitor itself requires the gateway to run. Your monitoring layer should never depend on the thing it's monitoring.
The gateway software can be configured to emit telemetry directly to an Application Insights workspace. Enable this by editing the gateway configuration file at %ProgramFiles%\On-Premises Data Gateway\Microsoft.PowerBI.DataMovement.Pipeline.GatewayCore.dll.config. Add your Application Insights instrumentation key to the relevant configuration section.
With Application Insights telemetry flowing, you can write KQL queries in Azure Monitor to detect gateway performance trends:
customMetrics
| where name == "QueryExecutionTimeMs"
| where timestamp > ago(1h)
| summarize
avg_ms = avg(value),
p95_ms = percentile(value, 95),
p99_ms = percentile(value, 99),
request_count = count()
by bin(timestamp, 5m), tostring(customDimensions.NodeName)
| order by timestamp desc
This query gives you per-node, per-5-minute execution time percentiles—far more useful than average-only metrics for detecting tail latency problems that averages mask.
Keeping gateway nodes updated is non-negotiable. Microsoft releases gateway updates monthly, and security patches can arrive out of cycle. Running outdated gateway software creates compatibility risk as the cloud-side connector infrastructure evolves and, more critically, security risk from unpatched vulnerabilities in the gateway process.
Because a cluster continues serving requests as long as at least one node is healthy, you can upgrade nodes one at a time:
In the Power Platform Admin Center, navigate to the cluster and select the node you're upgrading first. Set its status to Disabled. The node immediately stops receiving new requests; in-flight requests complete. Wait 60–90 seconds for in-flight requests to drain.
On the gateway machine, open the gateway application and check for updates, or run the latest installer over the existing installation. The upgrade preserves your configuration and cluster registration.
After the upgrade completes, the gateway service restarts automatically. Verify the node reconnects to the relay (check the status in Admin Center—it should return to green/connected within about 30 seconds).
Re-enable the node in Admin Center.
Repeat for the next node.
This process gives you zero-downtime upgrades on a cluster with three or more nodes. On a two-node cluster, you'll be running with a single node during each individual node's upgrade—which means you have HA for node failures but not for upgrade-coincident failures. This is another argument for three-node minimum clusters in production.
Warning
Do not upgrade all gateway nodes simultaneously. This is the upgrade equivalent of the single-gateway anti-pattern—you've just scheduled your own outage. One node at a time, always, with verification between each upgrade.
If a gateway node fails to come back online after an upgrade (driver incompatibility, .NET version mismatch, Windows Server issue), you have a few options:
Rollback: The gateway installer doesn't support automatic rollback, but if you've snapshotted the VM or kept the previous installer, you can reinstall the previous version. The gateway's cluster registration persists in Azure, so reinstalling reconnects it to the cluster.
Remove and replace: In Admin Center, remove the failed node from the cluster. Install fresh on a replacement VM and add it as a member node using the recovery key. Re-enable the cluster's auto-recovery mechanism to bring the replacement in.
Hot spare strategy: For mission-critical clusters, maintain a pre-configured spare VM with the gateway installed and cluster-registered but set to disabled in Admin Center. During a node failure, enable the spare in under a minute—no installation, no recovery key entry, just a status toggle.
For organizations with diverse on-premises data sources and flow workloads, a single gateway cluster serving everything creates several problems:
The solution is workload-specific gateway clusters. Define clusters by data source domain:
PROD-GW-FinancialDB → SQL Server cluster (financial systems, ERP)
PROD-GW-HRSystems → SQL Server + Oracle cluster (HR domain)
PROD-GW-Legacy → File system gateway for legacy flat-file integrations
Flows are bound to clusters through their connection references. A flow that reads employee data uses the HR systems cluster; a flow that reads financial ledger data uses the financial DB cluster. Neither workload can saturate the other's cluster.
This architecture does increase administrative overhead—you're managing multiple clusters, multiple service accounts, multiple upgrade schedules. Document your cluster inventory in your governance tool of choice, and use the CoE Toolkit's gateway inventory features to maintain visibility across all clusters.
If your on-premises infrastructure spans multiple physical locations (multiple data centers, colocation facilities, or a hybrid cloud with Azure Stack), you can place gateway nodes from the same cluster across those locations. Because all nodes in a cluster register with the same Azure Service Bus relay, location is transparent to the load balancer.
A practical pattern for organizations with two data centers: place two nodes in each data center, for a four-node cluster. Label nodes by location (achievable with node-level naming conventions in Admin Center). If an entire data center goes offline (power outage, network failure), the two nodes in the surviving location continue serving requests.
Key insight
This multi-site gateway cluster pattern requires that your data source itself is also replicated or accessible from both locations. A gateway cluster that survives a data center failure is useless if the SQL Server it's connecting to is only in the failed data center. Gateway HA and data source HA must be planned as a system, not independently.
High-volume scenarios—where flows are triggered by events and need to make many on-premises database calls—can overwhelm even a well-sized gateway cluster if not designed carefully. The pattern to reach for is decoupling: instead of having each event directly trigger a flow that calls the gateway, batch events through a queue and process them at a controlled rate.
Azure Service Bus queue-based patterns with Power Automate are directly applicable here. The queue acts as a buffer: events arrive at whatever rate the source produces them, but the consuming flows drain the queue at a rate your gateway cluster can sustain. If the queue backs up, it's a self-healing backpressure signal—slow down consumption, nothing drops. This is far more resilient than having hundreds of simultaneous flow instances all hammering the gateway concurrently.
Combined with parallel branching patterns, you can tune the degree of parallelism in your consuming flows to match your gateway cluster's capacity, achieving both throughput and resilience.
In this exercise, you'll deploy a two-node gateway cluster (suitable for a lab or development environment), configure load balancing, and verify that node failover works correctly. You'll need two Windows machines (or VMs) and a Power Platform environment with admin access.
On Machine A (Primary Node):
LAB-GW-Cluster-01.Verify in Admin Center:
Navigate to Power Platform Admin Center → Data → Gateways. Confirm you see LAB-GW-Cluster-01 with one node, status Online.
On Machine B (Member Node):
LAB-GW-Cluster-01 from the cluster list (if it doesn't appear, type the name manually).Verify in Admin Center: The cluster should now show two nodes. Both should show status Online.
Enable load balancing: In Admin Center, select the cluster and enable "Distribute requests across all active nodes in this cluster." Save.
LAB-GW-Cluster-01 gateway cluster from the gateway dropdown.Create a simple instant-trigger test flow that uses a "Get rows (V2)" action against a table in your SQL Server using this connection. Test it manually—verify it returns data.
This confirms that the cluster routes around node failures transparently from the flow's perspective. In a real production scenario with a three-node cluster, this failover happens automatically without any administrative intervention when a node crashes—you'd simply see the Admin Center status change to offline and the remaining nodes absorb the load.
On Machine A (the primary node), navigate to %LocalAppData%\Microsoft\On-premises data gateway\ and open GatewayInfo.log. Search for the flow execution timestamps and trace the request handling—you'll see entries showing the request received, the connection established to SQL Server, the execution time, and the response sent.
Do the same on Machine B and observe that when Machine A is disabled, Machine B's logs show the requests being handled there instead.
Deploying a three-node cluster where all three nodes are VMs on the same physical host, same hypervisor cluster, and same power circuit. If the physical host goes down, all three gateway nodes go down simultaneously. True HA requires that nodes be distributed across independent failure domains—different hypervisor hosts, different racks, different network switches.
Fix: Work with your infrastructure team to define the failure domain for each gateway VM and ensure nodes span at least two independent failure domains.
The default Windows service recovery settings for the gateway service don't restart the service automatically after a crash. So if the gateway process crashes (OOM, unhandled exception), it stays down until someone manually starts it. This converts a recoverable transient failure into an extended outage.
Fix: On each gateway node, open services.msc, find the "On-premises data gateway service," open Properties → Recovery tab. Set: First failure: Restart the service. Second failure: Restart the service. Subsequent failures: Restart the service. Reset fail count after 1 day. This ensures the service self-heals from transient crashes.
Microsoft updates the cloud-side connector infrastructure regularly. The on-premises gateway has a compatibility window—it supports the current version and some number of previous versions. Running a gateway that's more than 6 months out of date risks compatibility failures that manifest as cryptic error messages like "The gateway does not support this feature" or silent connection resets.
Fix: Check the gateway version in Admin Center monthly. Implement the rolling upgrade procedure described above as a recurring operational task, not a reactive emergency measure.
When you add a node to an existing cluster or change cluster settings, existing connections that reference the cluster automatically pick up the new topology—the connection points to the cluster, not a specific node. However, if you create a new cluster (rather than adding nodes to an existing cluster), all existing connections still point to the old cluster. Flows using those connections will not benefit from the new cluster until connections are updated.
Fix: Before creating a new cluster, confirm whether adding nodes to the existing cluster meets your needs. Reserve new cluster creation for genuine workload isolation scenarios, not just "I want more capacity."
You've enabled load balancing, but monitoring shows one node handling 90% of requests.
Cause 1: The other node's concurrency limit is set very low, so it's regularly at capacity and excluded from selection. Check per-node concurrency settings in Admin Center.
Cause 2: One node's relay connection is degraded—it's technically "connected" but experiencing high latency on the Service Bus link. The load balancer may deprioritize it. Check that node's gateway logs for relay connectivity warnings.
Cause 3: A specific flow has established stateful sessions on one node (see the sticky sessions section above) and its long-running requests are pinning to that node.
Power Automate's on-premises connector has a default query timeout of 600 seconds (10 minutes), but in practice, gateway-mediated queries that return very large result sets can fail before that with memory or buffer errors.
Fix: For large data movements, don't use the gateway as a data pipe. Use the gateway to call a stored procedure that writes results to a staging table or Azure Blob Storage, then read from those outputs with a cloud-native connector. The gateway handles the small control-plane request; the heavy data movement happens server-side or through a cloud-native path. This also tends to be significantly faster.
You've now got a complete picture of gateway cluster architecture, from the Azure Service Bus relay model that makes it work, through the load balancing algorithm that distributes requests, to the operational practices that keep a cluster healthy over time.
The key principles to carry forward:
Architecture: Gateway clusters achieve HA through multiple nodes consuming from a shared Service Bus relay. Load balancing is capacity-based, not round-robin—node sizing homogeneity matters. Three nodes is the right minimum for production.
Operations: Rolling upgrades, one node at a time, with verification between steps. Service recovery options set on every node. Recovery keys stored in your secrets management system, not an email thread.
Monitoring: Don't wait for flow failures to discover gateway problems. Build proactive health monitoring using the admin API, ship gateway logs to your SIEM or Application Insights, and set alerts on node offline status.
Design: Avoid making the gateway a bottleneck by designing flows that minimize round trips through it. Use the gateway for control-plane queries; push heavy data movement to server-side processing or direct cloud connections where possible. Decouple high-volume gateway workloads using Service Bus queues to apply natural backpressure.
With a resilient gateway cluster in place, your on-premises integration layer is solid. From here, the natural next area to tackle is end-to-end flow monitoring—combining gateway health signals with flow run telemetry and business-level metrics into a unified observability picture. The Power Automate Monitoring and Alerting System article extends what we've covered here into a full production monitoring architecture.
If you're designing high-volume integration patterns where the gateway is one component of a larger pipeline, explore the event-driven patterns with Azure Service Bus—the queue-based backpressure model pairs directly with gateway cluster sizing and protects your cluster from being overwhelmed by traffic spikes.
Finally, if your organization is preparing to move gateway configuration across environments as part of your ALM pipeline, make sure your solution packaging strategy accounts for the connection references that bind flows to environment-specific gateway clusters—this is covered in depth in the ALM Pipelines and Environment Variables article.
Enterprise Cloud Flows