July 2026
23
This is our Preliminary PIR to share what we know so far. After our internal retrospective is completed (generally within 14 days) we will publish a Final PIR with additional details.
What happened?
Between 14:44 UTC and 19:41 UTC on 23 July 2026, customers may have experienced connectivity failures, increased latency, or difficulty accessing Azure and other Microsoft cloud services hosted within the West US region. Impact was limited to network traffic entering or exiting the West US region; traffic that remained entirely within the region was not affected.
Impacted services included, but may have not been limited to:
App Service, Application Gateway, Application Insights, Azure AD B2C, Azure AI Search, Azure AI Speech, Azure API Management, Azure Bastion, Azure Bot Service, Azure Cosmos DB, Azure Data Explorer, Azure Database for PostgreSQL, Azure Databricks, Azure Firewall, Azure Kubernetes Service (AKS), Azure Monitor, Azure Virtual Desktop, Azure VMware Solution, ExpressRoute Circuits, ExpressRoute Gateways, Log Analytics, Microsoft Graph, Microsoft Sentinel, Partner Center, Power BI Embedded, Virtual WAN, VPN Gateway
What do we know so far?
On July 23, 2026, a routine device maintenance required isolating specific network paths. Our maintenance process converts these requests into system-readable requests and verifies that at least one of the two redundant paths remains healthy. Before maintenance begins, these requests undergo safety checks to confirm the work will be impact-less. In this case, a bug in the request conversion system incorrectly marked additional devices as a part of the maintenance event and caused a set of IP routes to be removed from more devices than intended. The routes were removed between our datacenter and wide-area network, impacting traffic entering or exiting the region.
At 14:44 UTC, customers started experiencing impact and the engineering team was engaged immediately. The issue initially presented as large-scale route churn in our Wide-Area Network (WAN). It was later found that route removal was from a datacenter in the West US region. Once confirmed, engineers began roll back at 17:45 UTC. By 18:26 UTC, the network was restored and healthy, and all impacted services had fully recovered by 19:41 UTC.
How did we respond?
- 14:44 UTC on 23 July 2026 – Device maintenance began. Earliest time when customers may have begun to experience impact. Multiple Azure services began to detect and correlate service degradation.
- 14:45 UTC on 23 July 2026 – Networking, service teams, and incident responders begin reviewing traffic anomalies, routing behavior, packet loss signals, and recent changes.
- 15:00 - 16:00 UTC on 23 July 2026 – We validated impact reports and telemetry to determine impact scope and blast radius. Began ruling out components, including addressing the large-scale route churn occurring in our WAN.
- 16:00 - 17:45 UTC on 23 July 2026 – Following continued degradation, we continued narrowing down our investigation. Identifying abnormal network routing behavior due to indications of improper advertising of some routes. Then identified the recent fiber maintenance activity, which we correlated to the identified routing behavior.
- 17:45 UTC on 23 July 2026 – We initiated the rollback of the maintenance change request.
- 18:26 UTC on 23 July 2026 – We completed the rollback and began monitoring recovering telemetry and services.
- 19:41 UTC on 23 July 2026 – All services recovered. Networking confirmed services can failback to West US as needed.
What happens next?
- We will be preforming a full analysis focusing on safety checks, automated maintenance request change process, and more as we progress through our post mitigation internal retrospective.
- The impact times above represent the full incident duration, so are not specific to any individual customer. Actual impact to service availability may vary between customers and resources – for guidance on implementing monitoring to understand granular impact: https://aka.ms/AzPIR/Monitoring
- For mission-critical workloads, customers should consider a multi-region deployment strategy to maintain availability during events that affect a single region: https://learn.microsoft.com/azure/architecture/patterns/geodes and https://learn.microsoft.com/azure/well-architected/design-guides/regions-availability-zones
- More generally, consider evaluating the reliability of your applications using guidance from the Azure Well-Architected Framework and its interactive Well-Architected Review: https://aka.ms/AzPIR/WAF
- Consider ensuring that the right people in your organization will be notified about any future service issues – by configuring Azure Service Health alerts. These can trigger emails, SMS, push notifications, webhooks, and more: https://aka.ms/AzPIR/Alerts
- This is our Preliminary PIR to share what we know so far. After our internal retrospective is completed (generally within 14 days) we will publish a Final PIR with additional details.
How can we make our incident communications more useful?
You can rate this PIR and provide any feedback using our quick 3-question survey: https://aka.ms/AzPIR/LYXT-C1Z