Skip to Main Content

Product:

Region:

Date:

September 2026

29

What happened?

Between 10:03 UTC and 15:58 UTC on 29 September 2026, a platform issue resulted in impact to Azure OpenAI Service, Foundry Agent Service, Foundry Models, and Cognitive Services in the Sweden Central region. Impacted customers experienced intermittent request failures, increased latency, and HTTP 5XX errors when submitting requests to affected models and data-plane APIs hosted in this region.

What do we know so far?

A backend service that is used to retrieve service information and resource metadata when processing requests in the Sweden Central region experienced timeouts when reading from its dependent database and caching layers. Instances for this backend service experienced increased utilization reaching service thresholds, which caused them to become unhealthy and restart repeatedly. An automated health check then triggered additional restarts, further reducing the number of healthy instances available to process requests. As a result, requests that depended on this service were delayed or failed intermittently. We have identified that an internal system process generated an increase in the overall amount of requests for the backend service which reached thresholds when calling the aforementioned database and caching layer.

How did we respond?

  • 10:03 UTC on 29 September 2026 – Initial impact began. 
  • 10:14 UTC on 29 September 2026 – Service monitoring detected elevated failure rates in the backend service that retrieves resource metadata for requests in the region. 
  • 11:29 UTC on 29 September 2026 – Engineers identified timeouts between the backend service and its dependent database layer. 
  • 12:58 UTC on 29 September 2026 – Engineers reviewed ongoing deployments to determine whether any recent change correlated with the issue. 
  • 13:28 UTC on 29 September 2026 – Engineers made attempts to mitigate through scaling adjustments of service instances to limit repeated restarts. 
  • 15:22 UTC on 29 September 2026 – Engineers scaled out the backend service and increased its resource limits to address throttling. In parallel, we investigated infrastructure health for the affected service instances and restarted infrastructure nodes. 
  • 15:58 UTC on 29 September 2026 – Following the scale-out and the disabling of an automated health check, service availability recovered and customers began to experience improvements. Customer impact was mitigated.
  • 18:00 UTC on 29 September 2026 – After a period of monitoring confirmed that service health remained stable, we further investigated the contributing factors. We will be implementing additional improvements to throttling and rate-limiting of internal processes, and making further service reconfigurations to help prevent recurrence.

What happens next?

  • Our team will be completing an internal retrospective to understand the incident in more detail. Once that is completed, generally within 14 days, we will publish a Final Post Incident Review (PIR) to all impacted customers. 
  • To stay informed about future Azure service issues, make sure that you configure and maintain Azure Service Health alerts – these can trigger emails, SMS, push notifications, webhooks, and more:
  • For more information on Post Incident Reviews, refer to
  • The impact times above represent the full incident duration, so are not specific to any individual customer. Actual impact to service availability may vary between customers and resources – for guidance on implementing monitoring to understand granular impact:
  • For broader guidance on preparing for cloud incidents, refer to

July 2026

23

Watch our 'Azure Incident Retrospective' video about this incident:

What happened?

Between 14:44 and 19:41 UTC on 23 July 2026, a subset of customers experienced connectivity failures, increased latency, and/or difficulty accessing Azure services hosted within the West US region. Impact was limited to network traffic entering or exiting the region, traffic that remained entirely within the region was not affected.

Impacted services included: App Service, Application Insights, Azure AD B2C, Azure AI Bot Service, Azure AI Search, Azure API Management, Azure Application Gateway, Azure Bastion, Azure Cosmos DB, Azure Data Explorer, Azure Database for PostgreSQL, Azure Databricks, Azure ExpressRoute, Azure Firewall, Azure Kubernetes Service (AKS), Azure Monitor, Azure Speech in Foundry Tools, Azure Virtual Desktop, Azure Virtual WAN, Azure VMware Solution, Azure VPN Gateway, as well as Microsoft Sentinel and Power BI Embedded. Downstream impact to Microsoft 365 services has been communicated via MO1437424.

What went wrong and why?

Azure datacenters connect through redundant network paths, ensuring that traffic can continue flowing even when one path is unavailable for maintenance or repair. When unplanned maintenance – including ‘break-fix’ hardware repair – is needed on an optical device along one of these paths, the repair process starts by creating system-readable instructions, verifying that the redundant path remains healthy, and performing safety checks to confirm that the work will not affect customer traffic. Because break-fix activities on a single device are generally expected to be impactless, these requests are processed automatically without human engagement or triage.

On 23 July 2026, a break-fix repair was initiated on an optical device to address a network reliability risk. A defect in our blast radius analysis system incorrectly expanded the scope of the repair event to include all optical devices egressing a specific datacenter. The safety validation step, which is designed to confirm that at least one of the two redundant datacenter paths remains available, ran but incorrectly concluded the operation was safe. The checks validated each device individually rather than evaluating the aggregate effect of isolating all devices at once, a scenario that was not accounted for because the system was never designed to process a full datacenter's worth of devices in a single request. As a result, routes were withdrawn from multiple devices simultaneously, disrupting connectivity between the datacenter and the WAN – therefore impacting traffic entering or leaving the West US region.

Once the route withdrawals took effect at 14:44 UTC, physical links and routing adjacencies continued to appear healthy, which initially masked the correlation between the break-fix activity and the connectivity disruption. The impact presented as a WAN routing anomaly, as third-party networks could not reach Azure in the region, rather than as a datacenter connectivity failure. Although all physical work in the region was stopped, our engineers could not correlate to this recent change because the preparation activities in advance of the break-fix did not succeed, so the physical layer and traffic appeared healthy.

Meanwhile, strong signals indicating WAN impact drew much of the investigative focus. Two parallel investigation streams developed: one focused on the WAN routing anomaly, and another examining network device health from within the datacenter. It took additional time to work past these misleading signals and eventually identify the break-fix activity as the source. It was not until 17:19 UTC that our engineers were able to correlate the routing anomaly with the break-fix activity, at which point we had identified the trigger event and began planning a rollback.

Our automated recovery and rollback system detected the device failures, and attempted multiple retries to restore the affected devices. However, because that system depended on the same datacenter connectivity that had been disrupted, its automated rollback attempts were unsuccessful. After our engineers correlated the routing anomaly with the break-fix activity, we identified that the incomplete preparation activities established the trigger conditions. We initiated a manual rollback at 17:45 UTC, then datacenter-to-WAN connectivity was restored when the rollback completed at 18:26 UTC, and all impacted services fully recovered by 19:41 UTC.

How did we respond?

  • 14:44 UTC on 23 July 2026 – Device break-fix began. Earliest potential customer impact time. 
  • 14:45 UTC on 23 July 2026 – Service alerts fired and multiple Azure service teams started investigating the traffic anomalies.
  • 14:54 UTC on 23 July 2026 – Automated recovery system started performing multiple retries to recover the devices that were affected by the over scoping that happened during the break-fix issue.
  • 15:26 UTC on 23 July 2026 – Routing anomaly alert incorrectly triaged to WAN issue and was sent to engineers to investigate further. 
  • 16:58 UTC on 23 July 2026 – After multiple retries, automated recovery system raised an alert to engineers to investigate. Engineers tried to recover the devices manually. 
  • 17:19 UTC on 23 July 2026 – Engineers correlated the routing anomaly with the initial break-fix trigger event and began attempting auto-rollback.
  • 17:45 UTC on 23 July 2026 – Engineers initiated manual rollback of the break-fix.
  • 18:26 UTC on 23 July 2026 – Rollback was successful on all affected devices, engineers continued monitoring recovery of affected services.
  • 19:41 UTC on 23 July 2026 – Customer impact confirmed as mitigated, after all services returned to expected operating levels.

How are we making incidents like this less likely or less impactful?

  • We have fixed the defect in the blast radius analysis system that incorrectly expanded the scope of this ‘break-fix’ unplanned maintenance to include additional devices. Input validation now ensures that both diverse paths for a datacenter cannot undergo break-fix requests at the same time. (Completed)
  • We have scanned all existing break-fix requests in the system to identify similar conditions that could have allowed the same defect to produce similar impact elsewhere. (Completed)
  • We have enhanced safety checks to block any break-fix or maintenance activity that would affect all diverse paths for a datacenter, to de-risk this failure mode. (Completed)
  • We are improving change visibility, break-fix requests will be recorded in the standard Azure change ledger, to reduce the time required to correlate maintenance activity with network issues during future investigations. (Estimated completion: August 2026)
  • We are enhancing our datacenter traffic monitoring to detect abnormalities in inbound and outbound traffic more quickly, and surface these directly to engineers for evaluation – enabling faster customer notification, and mitigation efforts. (Estimated completion: September 2026)
  • We are improving change tooling guardrails so that multiple devices cannot be isolated simultaneously during maintenance. (Estimated completion: September 2026)
  • In the longer term, we are improving our automated recovery system to fail more quickly and escalate to engineers more promptly when retries are unsuccessful, rather than continuing to retry against disrupted connectivity. (Estimated completion: October 2026)
  • Finally, we are improving our automated rollback so it can recover more robustly, even when the underlying connectivity on which it depends is degraded. (Estimated completion: October 2026)

How can customers make incidents like this less impactful?

How can we make our incident communications more useful?

You can rate this PIR and provide any feedback using our quick 3-question survey: