The encyclopedia · Software & IT · Technical decision · 2025
Azure Front Door outage takes down M365, Xbox and Azure Portal worldwide
Incompatible config metadata exposed a latent data-plane bug that crashed Microsoft's CDN for over eight hours.
Microsoft · 2025-10-29
What happened
From 15:41 UTC on 29 October to 00:05 UTC on 30 October 2025, Microsoft's Azure Front Door and Content Delivery Network suffered an outage of roughly eight and a half hours, taking down Microsoft 365, Xbox, Minecraft, Azure Portal, Outlook, Copilot and a swathe of first-party and third-party services worldwide.
Microsoft's post-incident review traced the cause to a specific sequence of customer configuration changes run across two different control-plane build versions. The incompatible metadata they produced exposed a latent bug in the data plane, which crashed during asynchronous processing. The change passed the config-protection safeguards and escaped pre-production validation, then crashed the data plane and took out AFD's internal DNS service too.
Microsoft rolled back to an edited 'Last Known Good' configuration, failed traffic away from AFD, and recovered node by node. It said the service was back above 98% availability and the majority of customers mitigated. Third-party customers including airlines, banks and retailers were caught in the blast.
After mitigation Microsoft blocked AFD customer-configuration changes, then restricted them at the ARM level until 5 November. It added a pre-canary validation stage, extended bake time, removed asynchronous processing from the data plane, and cut data-plane recovery from about four and a half hours to about one.
Why it happened
- Customer config changes across two build versions produced incompatible metadata that a pre-production gap never caught
- A latent data-plane bug crashed during async processing, silently taking out the CDN and its own DNS
- One company's ingress layer was the front door for Microsoft 365, Xbox, Minecraft and thousands of customer sites, so one bug downed all of them
- Recovery measured in hours, not minutes, because the edge infrastructure had no fast rollback path
The lesson
If every service shares one front door, that door needs the most rigorous validation in the company — one bad config change is a company-wide outage.
Sources
- Azure Front Door — Connectivity issues across multiple regions (PIR YKYN-BWZ) — Microsoft Azure status
- Microsoft's Azure outage took down Xbox, Microsoft 365 and more — The Verge
spotted an error? The club wants to know.
More like this
Cloudflare's 1.1.1.1 DNS resolver goes dark worldwide for 62 minutes
Google Cloud's unflagged feature crashed 70+ services for 7 hours
Coupang's 2025 breach: 34M accounts exposed by a former employee's data key
Somewhere, someone solved the problem this company failed at. 2nd Opinion →

Comments · 0
Sign in to join the comments.