Learn how to build a production-grade telemetry pipeline that emits structured events from Power Automate flows to Azure Log Analytics — then query them with KQL to power SLA dashboards, failure rate trends, and cross-environment latency reports that actually answer operational questions.

You've promoted your Power Automate solution to production, and within a week your operations team starts asking questions you can't easily answer: "Why did the invoice processing flow slow down Tuesday afternoon? How many orders failed the SAP sync last month? Are we meeting the 99.5% SLA we committed to for the ERP integration?" You open the Power Automate portal, click into the flow run history, and realize you're staring at a paginated list of individual runs with no aggregation, no cross-flow correlation, no time-series trends, and absolutely no way to answer the SLA question without exporting data manually and crunching it in Excel.
This is the gap custom telemetry fills. Built-in flow run history tells you whether a run succeeded or failed. A custom telemetry pipeline tells you why, how fast, at what volume, and whether you're meeting contractual obligations — across every environment, every flow, every integration partner. By emitting structured events from inside your flows to Azure Log Analytics, you get a queryable, time-series-aware data store that powers operational dashboards, SLA reports, and proactive alerting without touching the underlying platform internals.
By the end of this lesson, you'll have built a complete, production-grade telemetry pipeline from scratch. Specifically:
What you'll learn:
You should be comfortable building and running cloud flows in a solution context. Familiarity with HTTP actions in Power Automate is assumed — if you need a refresher on calling external APIs from flows, the lesson on Advanced Power Automate: Custom Connectors and HTTP Actions for Production Integration covers the mechanics in depth. You'll also want a basic understanding of Azure portal navigation and the ability to create a Log Analytics workspace. Prior exposure to JSON and Kusto Query Language (KQL) will help, but we'll explain each query we write.
Before we build anything, let's settle a design decision you'll encounter immediately. Both Azure Application Insights and Azure Log Analytics (part of Azure Monitor) accept custom telemetry. If you've seen the lesson on Building a Power Automate Monitoring and Alerting System, you know Application Insights works well for developer-centric APM scenarios — traces, exceptions, dependencies, and request telemetry tied to a specific application lifecycle.
Log Analytics is the better choice for operational telemetry pipelines that need to span multiple flows, multiple environments, and multiple teams. Here's why:
Queryability. Log Analytics workspaces store everything as tables you query with KQL. You can JOIN your flow telemetry against Azure resource health events, security audit logs, and even custom tables from other systems — all in one query.
Retention and cost control. You set retention policies per table (30 days to 7 years), archive cold data cheaply, and control ingestion costs at the workspace level. Application Insights has a fixed data model that's less flexible for business-domain events.
Cross-environment aggregation. You can point flows from your Dev, Test, UAT, and Production environments at the same workspace (using different custom fields to tag the source), giving you a single pane for comparing behavior across the ALM pipeline. The lesson on Deploying and Managing Power Automate Solutions Across Environments describes how environment promotion works; your telemetry strategy should be designed alongside it.
Note
You can — and for complex enterprises often should — use both. Send flow execution telemetry to Log Analytics and use Application Insights for any Azure Functions or API Management layers that sit alongside your flows. The KQL language is identical across both services, so your query skills transfer.
Resist the urge to start emitting events immediately. The schema you define now determines what questions you can answer in six months, and retrofitting a schema after data is already in Log Analytics requires a painful table migration.
Every telemetry event should share a common envelope regardless of which flow emits it. Think of this as the fields that allow you to aggregate, filter, and correlate across your entire automation portfolio:
{
"EventType": "FlowExecution",
"SchemaVersion": "1.0",
"TenantId": "contoso",
"EnvironmentName": "Production",
"EnvironmentId": "env-00000000-0000-0000-0000-000000000001",
"SolutionName": "InvoiceProcessing",
"FlowName": "Process-InvoiceReceived",
"FlowId": "00000000-0000-0000-0000-000000000002",
"RunId": "08585012345678901234567890",
"CorrelationId": "inv-2024-00847",
"TriggerType": "ServiceBusTrigger",
"Status": "Succeeded",
"StartedAt": "2024-11-15T09:23:14.221Z",
"CompletedAt": "2024-11-15T09:23:19.887Z",
"DurationMs": 5666,
"BusinessEntityType": "Invoice",
"BusinessEntityId": "INV-2024-00847",
"BusinessOutcome": "Routed",
"ErrorCode": null,
"ErrorMessage": null,
"ErrorStep": null,
"RetryAttempt": 0,
"Tags": {
"Vendor": "Acme Corp",
"InvoiceAmount": 14250.00,
"Currency": "USD",
"ApprovalRequired": true
}
}
Let's walk through the fields that require intentional design decisions:
CorrelationId ties this flow run to a business transaction that may span multiple flows, an Azure Function, and a Dataverse write. If you're not already implementing correlation IDs across your integrations, the lesson on Implementing Correlation IDs and End-to-End Distributed Tracing explains the pattern in full — it's worth reading before you finalize this schema.
DurationMs is the field your SLA queries will depend on most. Calculate it as CompletedAt - StartedAt in milliseconds. You'll want this for p50, p95, and p99 latency percentiles.
Tags is a dynamic JSON object for business-domain fields that vary by flow. By keeping them in a nested object, you don't pollute the top-level schema, and KQL's parse_json() function handles them cleanly.
RetryAttempt tracks whether the flow run is a first attempt or a retry, which is critical for accurate SLA calculations. A run that eventually succeeds on retry three is still a first-delivery-failure for SLA purposes.
Key insight
Define Status as a controlled vocabulary, not free text. Use: Succeeded, Failed, PartialSuccess, Skipped, TimedOut. "PartialSuccess" is your escape hatch for scenarios where the flow completed but with degraded output — for example, an invoice processed but the approval notification failed to send. Without this status, you'd be forced into a binary succeeded/failed bucket that hides real operational problems.
For long-running flows, a single start/end event isn't sufficient. You need checkpoint events emitted at key milestones within the flow. These use the same envelope but with EventType set to "FlowCheckpoint" and an additional CheckpointName field:
{
"EventType": "FlowCheckpoint",
"CheckpointName": "SAP-WriteCompleted",
"CheckpointSequence": 3,
"ElapsedMs": 3210,
...all envelope fields...
}
Checkpoint events let you pinpoint exactly where time is being spent inside a complex flow — invaluable when a long-running SAP write or SharePoint large file operation is your performance bottleneck.
In the Azure portal, navigate to Log Analytics workspaces and create a new workspace. Choose the same region as your Power Platform environments to minimize latency and avoid cross-region data transfer costs. For naming, use a convention like law-powerautomate-telemetry-prod — the law- prefix is a common Azure naming convention for Log Analytics workspaces.
Once created, navigate to Agents management (or Legacy agents management depending on your portal version), and note the Workspace ID and regenerate or copy the Primary Key. You'll need both to call the Data Collector API.
Warning
The Workspace ID and Primary Key are sensitive credentials. Do not hardcode them in your flows. Store them in Azure Key Vault and retrieve them at runtime, as described in the lesson on Integrating Power Automate with Azure Key Vault and Managed Identities. Use environment variables to hold the Key Vault secret references so they resolve correctly across Dev, Test, and Production without flow edits.
The Log Analytics Data Collector API accepts HTTP POST requests with a JSON payload. The signature requirement is the part that trips most people up — it's not a bearer token; it's an HMAC-SHA256 signature built from the request content, timestamp, and your workspace key.
The API endpoint is:
POST https://{WorkspaceId}.ods.opinsights.azure.com/api/logs?api-version=2016-04-01
Required headers:
Log-Type: The custom table name (without the _CL suffix Log Analytics appends automatically)x-ms-date: RFC 1123 formatted timestamp (e.g., Fri, 15 Nov 2024 09:23:14 GMT)Authorization: SharedKey {WorkspaceId}:{Signature}Content-Type: application/jsontime-generated-field: (optional) the field in your JSON that holds the event timestamp — use StartedAt so Log Analytics uses your event time rather than ingestion timeThe signature is where the complexity lives, and it's why you'll use an Azure Function rather than a pure Power Automate HTTP action for this.
Computing HMAC-SHA256 in pure Power Automate expressions is not currently possible — the expression language doesn't expose cryptographic primitives. You have two options: use the Log Analytics Data Collector API (which requires HMAC) via an Azure Function, or use the newer Azure Monitor Ingestion API (Logs Ingestion API) with OAuth bearer token authentication.
The Logs Ingestion API is the modern approach and requires no HMAC computation. It uses a Data Collection Endpoint (DCE) and Data Collection Rule (DCR) with OAuth 2.0 client credentials — credentials you can store in Key Vault and retrieve via managed identity. We'll use this approach because it integrates cleanly with Power Automate's existing HTTP action + managed identity pattern.
Note
The Logs Ingestion API requires creating a Data Collection Rule (DCR) that defines your table schema upfront. This extra setup pays dividends: Log Analytics validates your payload against the schema at ingestion time, catching malformed events before they corrupt your table.
Step 1: Create a Data Collection Endpoint
In the Azure portal, search for Data Collection Endpoints and create one in the same region as your workspace. Name it dce-powerautomate-telemetry. Note the Logs Ingestion URI from the endpoint's Overview page — it will look like https://dce-powerautomate-telemetry-xxxx.eastus-1.ingest.monitor.azure.com.
Step 2: Create the custom table
In your Log Analytics workspace, navigate to Tables and create a new custom table. Name it PA_FlowExecution (Log Analytics will append _CL making it PA_FlowExecution_CL). Define the schema columns matching your event envelope: EventType (string), FlowName (string), RunId (string), CorrelationId (string), Status (string), DurationMs (long), BusinessEntityId (string), Tags (dynamic), and so on.
Step 3: Create a Data Collection Rule
Create a DCR linked to your DCE and workspace. The DCR defines the stream name (e.g., Custom-PA_FlowExecution_CL) and maps it to the destination table. During DCR creation you'll specify the transformation (use source to pass through without transformation initially) and the output table.
Step 4: Register an App Registration for authentication
Create an Azure AD App Registration named svc-powerautomate-telemetry. Create a client secret and store it in Key Vault. Grant this service principal the Monitoring Metrics Publisher role on the DCR resource.
Rather than embedding telemetry logic in every individual flow, build a dedicated child flow that acts as the telemetry emission service. Every parent flow calls this child with a telemetry payload, and the child handles the HTTP call, error handling, and retry logic. The lesson on Orchestrating Child Flows and Scoped Execution in Power Automate covers the structural patterns for child flow design — apply them here.
Name this flow Shared-EmitTelemetryEvent. Its input schema should accept the complete event envelope as a JSON object. Here's the flow structure in plain steps:
1. Trigger: Manually trigger a flow (used as child)
telemetryPayload (type: Object)2. Initialize variable: var_AccessToken (string, empty)
3. HTTP action: Get OAuth token
https://login.microsoftonline.com/{TenantId}/oauth2/v2.0/tokengrant_type=client_credentials
&client_id={AppRegistrationClientId}
&client_secret={ClientSecretFromKeyVault}
&scope=https://monitor.azure.com/.default
4. Set variable: var_AccessToken
body('HTTP_GetToken')['access_token']5. HTTP action: Send to Log Analytics Ingestion API
{LogsIngestionURI}/dataCollectionRules/{DCR_ImmutableId}/streams/Custom-PA_FlowExecution_CL?api-version=2023-01-01Authorization: Bearer @{variables('var_AccessToken')}Content-Type: application/json@{json(array(triggerBody()?['telemetryPayload']))}Tip
The Logs Ingestion API expects an array, even for a single event. Wrapping your payload in array() handles this. If you're emitting checkpoint events in bulk (say, 10 checkpoints at flow end), pass them all as one array to reduce API calls and respect Log Analytics ingestion rate limits.
6. Condition: Did the HTTP call succeed?
outputs('HTTP_SendTelemetry')['statusCode'] is between 200 and 2997. If No: Append to a fallback variable or log minimal error
The critical design principle here: if telemetry emission fails, the child flow should not throw an error that propagates to the parent. Set Configure run after on the failure branch to suppress errors. The parent flow should not fail because observability infrastructure is unavailable — that's observability defeating itself.
Here's how to call this child flow from any parent flow. In your parent flow, add a Run a Child Flow action after your main logic completes:
{
"telemetryPayload": {
"EventType": "FlowExecution",
"SchemaVersion": "1.0",
"TenantId": "contoso",
"EnvironmentName": "@{variables('var_EnvironmentName')}",
"SolutionName": "InvoiceProcessing",
"FlowName": "Process-InvoiceReceived",
"FlowId": "@{workflow()['tags']['flowDisplayName']}",
"RunId": "@{workflow()['run']['name']}",
"CorrelationId": "@{variables('var_CorrelationId')}",
"TriggerType": "ServiceBusTrigger",
"Status": "@{variables('var_FlowStatus')}",
"StartedAt": "@{workflow()['run']['startTime']}",
"CompletedAt": "@{utcNow()}",
"DurationMs": "@{sub(ticks(utcNow()), ticks(workflow()['run']['startTime']))}",
"BusinessEntityType": "Invoice",
"BusinessEntityId": "@{variables('var_InvoiceId')}",
"BusinessOutcome": "@{variables('var_BusinessOutcome')}",
"ErrorCode": "@{variables('var_ErrorCode')}",
"ErrorMessage": "@{variables('var_ErrorMessage')}",
"Tags": {
"Vendor": "@{triggerBody()?['VendorName']}",
"InvoiceAmount": "@{triggerBody()?['Amount']}"
}
}
}
Note the DurationMs calculation. ticks() returns 100-nanosecond intervals, so dividing by 10,000 gives milliseconds. The expression above gives ticks difference; adjust to div(sub(...), 10000) for actual milliseconds:
"DurationMs": "@{div(sub(ticks(utcNow()), ticks(workflow()['run']['startTime'])), 10000)}"
The workflow() function is your gateway to run context — flow ID, run ID, start time. For a complete reference on what this function exposes, see the lesson on Understanding Power Automate Run Context and Trigger Metadata.
The biggest mistake practitioners make is treating telemetry as an afterthought — adding it only in happy-path scenarios. Production-grade telemetry requires a try/catch/finally pattern that guarantees the telemetry event fires regardless of outcome.
Here's the pattern using Power Automate's Scope actions:
Scope: "TRY - Main Business Logic"
Scope: "CATCH - Error Handling"
result('TRY_Scope'):@{body('TRY_Main_Business_Logic')[0]['error']['message']}
var_FlowStatus to "Failed"var_ErrorMessage to the parsed error messagevar_ErrorCode to the action name that failedScope: "FINALLY - Emit Telemetry"
Shared-EmitTelemetryEvent child flowThis mirrors try/catch/finally semantics in code. Within the TRY scope, set var_FlowStatus = "Succeeded" and var_BusinessOutcome at the end of successful execution. The CATCH scope overrides these to failure states. The FINALLY scope reads whatever state was set and emits it.
Warning
When using parallel branches for performance (as covered in Implementing Parallel Branching and Concurrency Control in Power Automate), the run-after configuration on your CATCH scope needs to reference the join point after your parallel branches, not the individual parallel actions. Otherwise the CATCH may not trigger when one parallel branch fails silently.
Data in Log Analytics is queryable immediately after ingestion (usually within 2-5 minutes). Your table is PA_FlowExecution_CL. Here are the queries your operations team will use most.
Your SLA might state: "Invoice processing flows must complete within 30 seconds, 99.5% of the time." This query calculates it:
PA_FlowExecution_CL
| where TimeGenerated >= ago(30d)
| where FlowName_s == "Process-InvoiceReceived"
| where EnvironmentName_s == "Production"
| where EventType_s == "FlowExecution"
| summarize
TotalRuns = count(),
SucceededRuns = countif(Status_s == "Succeeded"),
FailedRuns = countif(Status_s == "Failed"),
p50_ms = percentile(DurationMs_d, 50),
p95_ms = percentile(DurationMs_d, 95),
p99_ms = percentile(DurationMs_d, 99),
WithinSLA = countif(DurationMs_d <= 30000 and Status_s == "Succeeded")
by bin(TimeGenerated, 1d)
| extend
SuccessRate = round(100.0 * SucceededRuns / TotalRuns, 2),
SLAComplianceRate = round(100.0 * WithinSLA / TotalRuns, 2)
| project
Date = TimeGenerated,
TotalRuns,
SuccessRate,
SLAComplianceRate,
p50_ms,
p95_ms,
p99_ms
| order by Date asc
Note
Log Analytics appends _s (string), _d (double/numeric), _b (boolean), and _t (datetime) suffixes to custom table columns. If you defined DurationMs as a long in your schema, it becomes DurationMs_d. Design your schema column names to be readable with these suffixes in mind — DurationMs becomes DurationMs_d which reads fine; Duration becomes Duration_d which is less clear.
PA_FlowExecution_CL
| where TimeGenerated >= ago(7d)
| where EventType_s == "FlowExecution"
| where EnvironmentName_s == "Production"
| summarize
TotalRuns = count(),
Failures = countif(Status_s == "Failed")
by FlowName_s, bin(TimeGenerated, 1h)
| extend ErrorRate = round(100.0 * Failures / TotalRuns, 1)
| where ErrorRate > 5
| order by TimeGenerated desc, ErrorRate desc
PA_FlowExecution_CL
| where TimeGenerated >= ago(7d)
| where Status_s == "Failed"
| where EnvironmentName_s == "Production"
| summarize FailureCount = count() by ErrorCode_s, ErrorMessage_s, FlowName_s
| top 20 by FailureCount desc
One of the core reasons to use Log Analytics over built-in flow history is the ability to compare behavior across environments. This query surfaces latency regressions between Test and Production:
PA_FlowExecution_CL
| where TimeGenerated >= ago(24h)
| where FlowName_s == "Process-InvoiceReceived"
| where EnvironmentName_s in ("Test", "Production")
| summarize
p95_ms = percentile(DurationMs_d, 95),
AvgDuration = avg(DurationMs_d),
RunCount = count()
by EnvironmentName_s
| order by EnvironmentName_s asc
If Test p95 is 2,000 ms and Production p95 is 18,000 ms, you have a scaling problem to investigate before the next release.
The BusinessOutcome field pays off here. This query shows how invoices are being routed, which is useful for business operations teams — not just technical staff:
PA_FlowExecution_CL
| where TimeGenerated >= ago(30d)
| where FlowName_s == "Process-InvoiceReceived"
| where Status_s == "Succeeded"
| summarize Count = count() by BusinessOutcome_s
| extend Percentage = round(100.0 * Count / toscalar(
PA_FlowExecution_CL
| where TimeGenerated >= ago(30d)
| where FlowName_s == "Process-InvoiceReceived"
| where Status_s == "Succeeded"
| count
), 1)
| order by Count desc
Once your queries work, pin them to an Azure Monitor Dashboard or build an Azure Workbook for richer, interactive reporting.
Azure Workbooks are the better choice for SLA reporting because they support parameters (date range pickers, environment selectors, flow name dropdowns) that feed into queries dynamically. To create one:
In the Azure portal, navigate to your Log Analytics workspace, select Workbooks from the left menu, and click New. Add a Parameters step first, creating:
TimeRange: Time range picker, default 30 daysEnvironment: Dropdown populated by PA_FlowExecution_CL | distinct EnvironmentName_sFlowName: Dropdown populated by PA_FlowExecution_CL | distinct FlowName_sThen add Query steps for each KQL query, referencing parameters using {Environment} and {TimeRange:start} / {TimeRange:end} syntax:
PA_FlowExecution_CL
| where TimeGenerated between ({TimeRange:start} .. {TimeRange:end})
| where EnvironmentName_s == "{Environment}"
| where EventType_s == "FlowExecution"
...
For the SLA compliance chart, use a Time chart visualization showing SLAComplianceRate over time with a reference line at 99.5. Operations teams can see immediately when SLA dropped below threshold and correlate it with deployment events or external system incidents.
If your flows run at high volume — thousands of runs per hour — naive telemetry emission creates its own API pressure. The Log Analytics Ingestion API has rate limits (approximately 500 MB/minute per DCE), and calling the child flow on every run adds latency.
Several strategies help at scale:
Batch checkpoint events. Rather than making one HTTP call per checkpoint event, accumulate all checkpoint events in an array variable during flow execution and emit them in a single HTTP call at the FINALLY scope. One call for 10 events versus 10 calls.
Sample non-critical telemetry. For very high-frequency flows where every individual run isn't critical to monitor, sample a percentage. Use rand() in an expression:
@{if(greater(rand(1, 100), 95), true, false)}
This emits telemetry for ~5% of runs for volume metrics and 100% of runs for failures (always emit failures). Your dashboard queries use extend Estimated = Count * 20 to extrapolate.
Async emission. If your flow is latency-sensitive and you don't want the telemetry HTTP call in the critical path, push the telemetry payload to a Service Bus queue and process it with a dedicated telemetry flow. This decouples your primary automation from the observability infrastructure entirely. The patterns for this are covered in Implementing Event-Driven Automation with Power Automate and Azure Service Bus.
Tip
Set the Shared-EmitTelemetryEvent child flow's retry policy to None. Retrying a failed telemetry call inside a FINALLY scope means a transient Log Analytics outage adds 30-90 seconds to every single flow run's total duration. The telemetry gap from a 5-minute outage is less harmful than slowing down all your production flows during that window.
Let's build a complete, working telemetry pipeline for a realistic scenario: an order validation flow that syncs to an ERP system.
Scenario: Your organization processes ~500 purchase orders per day through a Power Automate flow triggered by a SharePoint list item creation. The flow validates the order, calls an ERP API, and routes approved orders to a Dataverse queue. The SLA is: 98% of orders must be processed within 45 seconds.
Step 1: Provision infrastructure
law-pa-telemetry-proddce-pa-telemetryPA_FlowExecution with the schema defined earlierStep 2: Build Shared-EmitTelemetryEvent
Step 3: Instrument Process-PurchaseOrder
var_FlowStatus, var_BusinessOutcome, var_ErrorCode, var_ErrorMessage, var_CorrelationIdvar_CorrelationId from the incoming SharePoint item's OrderNumber fieldvar_FlowStatus = "Succeeded" and appropriate var_BusinessOutcome at the endShared-EmitTelemetryEvent with the full payloadStep 4: Run the flow 10 times
Step 5: Validate in Log Analytics
PA_FlowExecution_CL | take 10TimeRange = ago(1h)Step 6: Build the Workbook
If everything is wired up correctly, you should see a live dashboard that refreshes as new orders are processed — no manual exports, no Excel, no guessing.
"My table isn't appearing in Log Analytics after the first event." Custom tables via the Logs Ingestion API can take 5-20 minutes to appear after the first ingestion. Subsequent events land within 2-5 minutes. If after 30 minutes nothing appears, check the HTTP response from the Ingestion API call in your flow run history. A 200 status means the event was accepted; a 403 means your DCR RBAC is wrong; a 400 means your payload doesn't match the DCR stream schema.
"My DurationMs values are showing as 0 or null."
The ticks() function returns a long integer in expressions, but when serialized to JSON it may lose precision or be treated as a string. Cast explicitly: @{int(div(sub(ticks(utcNow()), ticks(workflow()['run']['startTime'])), 10000))}. Also ensure your DCR schema defines this column as long, not string.
"The CATCH scope isn't firing when my TRY scope fails." Check the Configure run after setting on your CATCH scope action. You must explicitly check all four boxes: "Is successful" unchecked, "Has failed" checked, "Has timed out" checked, "Has skipped" checked. The default run-after is "Is successful only," which means CATCH is skipped when TRY fails — the opposite of what you want.
"My telemetry child flow is causing parent flow timeouts." You likely have retry policy enabled on the child flow call. Disable retries on the Run a child flow action. Also check whether your Key Vault secret retrieval is itself slow — if Key Vault is cold-starting, the token fetch can add 3-5 seconds. Cache the token in a flow variable with an expiry check if your flow runs frequently enough.
"KQL queries return _s, _d suffixes inconsistently."
When you define the schema in the DCR, Log Analytics assigns type suffixes based on the type you declared. If a field is sometimes null and sometimes a string, the type inference can produce _s for string and fail to cast numeric values. Always declare explicit types in your DCR schema and always emit the field (use null rather than omitting the key) so the schema is consistent.
"Tags field isn't queryable."
The Tags field needs to be declared as dynamic type in your DCR schema. In KQL, access nested fields with parse_json(): parse_json(Tags_s)['Vendor']. If you declared it as string type accidentally, all the JSON content is a raw string and you must use parse_json() to extract values.
You've built a complete custom telemetry pipeline: a structured event schema designed for operational queries and SLA reporting, a reusable child flow that handles OAuth token acquisition and HTTP emission to Log Analytics, a try/catch/finally instrumentation pattern that guarantees telemetry fires on every run outcome, and a suite of KQL queries that power real operational dashboards.
The payoff is substantial. Instead of scrolling through paginated run history when an executive asks "did we meet SLA last month?", you open a Workbook and show a time-series chart with the exact number. Instead of discovering a performance regression in Production through user complaints, you catch it in your cross-environment latency comparison query before the release even goes live.
The natural extensions of this work are:
ErrorRate > 5% for 15 consecutive minutes, or when p95_ms > 45000 over the last hour. Route these to PagerDuty, Teams, or your ITSM system.CostCenter or BusinessUnit field to your envelope and use telemetry volume to drive chargeback calculations — directly supporting the approach described in Cost Management at Scale: Per-Flow vs Per-User Licensing.Custom telemetry isn't optional at production scale. It's the foundation that makes everything else — SLA commitments, incident response, capacity planning, and governance — actually achievable.