September 2026
29
What happened?
Between 10:03 UTC and 15:58 UTC on 29 September 2026, a platform issue resulted in impact to Azure OpenAI Service, Foundry Agent Service, Foundry Models, and Cognitive Services in the Sweden Central region. Impacted customers experienced intermittent request failures, increased latency, and HTTP 5XX errors when submitting requests to affected models and data-plane APIs hosted in this region.
What do we know so far?
A backend service that is used to retrieve service information and resource metadata when processing requests in the Sweden Central region experienced timeouts when reading from its dependent database and caching layers. Instances for this backend service experienced increased utilization reaching service thresholds, which caused them to become unhealthy and restart repeatedly. An automated health check then triggered additional restarts, further reducing the number of healthy instances available to process requests. As a result, requests that depended on this service were delayed or failed intermittently. We have identified that an internal system process generated an increase in the overall amount of requests for the backend service which reached thresholds when calling the aforementioned database and caching layer.
How did we respond?
- 10:03 UTC on 29 September 2026 – Initial impact began.
- 10:14 UTC on 29 September 2026 – Service monitoring detected elevated failure rates in the backend service that retrieves resource metadata for requests in the region.
- 11:29 UTC on 29 September 2026 – Engineers identified timeouts between the backend service and its dependent database layer.
- 12:58 UTC on 29 September 2026 – Engineers reviewed ongoing deployments to determine whether any recent change correlated with the issue.
- 13:28 UTC on 29 September 2026 – Engineers made attempts to mitigate through scaling adjustments of service instances to limit repeated restarts.
- 15:22 UTC on 29 September 2026 – Engineers scaled out the backend service and increased its resource limits to address throttling. In parallel, we investigated infrastructure health for the affected service instances and restarted infrastructure nodes.
- 15:58 UTC on 29 September 2026 – Following the scale-out and the disabling of an automated health check, service availability recovered and customers began to experience improvements. Customer impact was mitigated.
- 18:00 UTC on 29 September 2026 – After a period of monitoring confirmed that service health remained stable, we further investigated the contributing factors. We will be implementing additional improvements to throttling and rate-limiting of internal processes, and making further service reconfigurations to help prevent recurrence.
What happens next?
- Our team will be completing an internal retrospective to understand the incident in more detail. Once that is completed, generally within 14 days, we will publish a Final Post Incident Review (PIR) to all impacted customers.
- To stay informed about future Azure service issues, make sure that you configure and maintain Azure Service Health alerts – these can trigger emails, SMS, push notifications, webhooks, and more: https://aka.ms/ash-alerts
- For more information on Post Incident Reviews, refer to https://aka.ms/AzurePIRs
- The impact times above represent the full incident duration, so are not specific to any individual customer. Actual impact to service availability may vary between customers and resources – for guidance on implementing monitoring to understand granular impact: https://aka.ms/AzPIR/Monitoring
- For broader guidance on preparing for cloud incidents, refer to https://aka.ms/incidentreadiness