Back to the archive

The encyclopedia · Software & IT · Technical decision · 2025

AWS's US-East-1 DNS records fault took down much of the internet for a day

A DNS records problem in one AWS region cascaded to Disney+, Reddit, Lyft, Coinbase, Gov.uk and Amazon itself.

Amazon Web Services · 2025-10-20

What happened

On 20 October 2025, a DNS records problem in AWS's US-East-1 region (northern Virginia) took down more than 70 of its own services and cascaded to customers on the same infrastructure. The fault was first reported at 3:11 a.m. ET and affected Amazon.com, Disney+, Lyft, the McDonald's app, The New York Times, Reddit, Ring, Robinhood, Snapchat, United Airlines, Venmo, Canva, Roblox, Fortnite, Coinbase and Perplexity, among many others.

The root cause was not a database outage but a fault in the DNS records that tell other systems where to find DynamoDB, AWS's database service that underpins many of its other applications. "The data appears to be safe," Notre Dame professor Mike Chapple told CNBC. "Instead, something went wrong with the records that tell other systems where to find their data."

By 6:35 a.m. ET AWS said the DNS issue was "fully mitigated," but EC2 error rates and connectivity problems persisted into the afternoon as customers tried to launch new instances. AWS said around 1:30 p.m. ET it was seeing "early signs" of EC2 recovery. It was not until shortly after 6 p.m. ET that "all AWS services returned to normal operations," roughly fifteen hours after the first reports.

The bill fell on AWS's customers rather than on Amazon. British government sites Gov.uk and HMRC, Lloyds Banking Group and airlines were among those disrupted, and Amazon's own warehouse and delivery employees were told to stand by as internal systems and the Anytime Pay app went offline. AWS controls about a third of the cloud market (Synergy Research), so the failure was a reminder, in Chapple's words, that "when a major cloud provider sneezes, the Internet catches a cold."

Why it happened

  • One shared DNS dependency meant a single records fault in one region became a fault for every service and customer on it
  • AWS runs tens of thousands of businesses on one control plane, so a US-East-1 DNS problem propagated faster than teams could route around it
  • Recovery was serialised: DNS mitigated at 6:35 a.m. did not restore EC2 until the evening, so the blast radius stayed open most of the day
  • Amazon owned the same dependency as its sellers — its own warehouse, Flex and Anytime Pay systems went down with Seller Central
What it costa day offline for most of the major sites everyone usescostly

The lesson

When a few companies underpin the internet, a single DNS records fault is a continent-wide outage — audit shared dependencies as single points of failure.

Sources

spotted an error? The club wants to know.

Comments · 0

    Sign in to join the comments.

    More like this

    Somewhere, someone solved the problem this company failed at. 2nd Opinion →