Back to the archive

The encyclopedia · Software & IT · Technical decision · 2024

AT&T's nationwide outage blocked 92M calls and 25,000+ 911 attempts

A network config change put 125M+ devices offline for 12+ hours, and the FCC found the change shipped with no peer review and no post-install test.

AT&T · 2024-02-22

What happened

At around 2:45 AM Central Time on February 22 2024, AT&T's wireless network went into a nationwide 'protect mode' that stopped most calls, texts and data from routing. At the peak roughly 74,000 incidents were reported on Downdetector, and the outage affected more than 125 million devices across all 50 US states, Washington DC, Puerto Rico and the US Virgin Islands. Service was not fully restored for more than 12 hours.

The cause was a configuration error introduced while AT&T was expanding its network. The Federal Communications Commission's investigation found that the change was deployed without peer review, without post-installation testing, with inadequate lab testing and with insufficient controls over core network changes. The impact went beyond inconvenience: more than 92 million voice calls failed to connect and more than 25,000 attempts to reach 911 emergency services were blocked. San Francisco and other cities warned residents that cellular 911 calls might not go through.

AT&T shares fell 2.41% on the day, and the FCC fined the company $950,000 over the outage. The company said it would take steps to prevent a recurrence, but the FCC's report made the underlying failure clear: a change to the most critical infrastructure in the country went straight to production without the checks that would have caught it.

The outage was notable for how preventable it was. Nothing about the failure was exotic — no cyberattack, no hardware disaster — just a routine network change that skipped the standard review and test gates. The scale of the blast radius came from the fact that a single carrier's misdeployed config could take 125 million people offline at once, including their ability to call for help in an emergency.

Why it happened

  • AT&T deployed a core network configuration change with no peer review and no post-installation test, so the error reached millions of customers untouched
  • Inadequate lab testing and weak controls over core network changes let a preventable mistake become a nationwide failure
  • A single misconfigured change took down more than 125 million devices, showing how little resilience sits between one config error and the whole network
  • The 25,000+ blocked 911 attempts turned a business continuity problem into a public safety one
What it cost125M+ devices offline; 92M calls & 25,000+ 911 blockedcostly

The lesson

The most critical infrastructure demands the most rigorous review: a config change that skips peer review and post-install testing can take 125 million people offline and block emergency calls.

Sources

spotted an error? The club wants to know.

Comments · 0

    Sign in to join the comments.

    More like this

    Somewhere, someone solved the problem this company failed at. 2nd Opinion →