Back to the archive

The encyclopedia · Engineering & Operations · Technical decision · 2023

Vipshop's cooling failure cost ¥100M and a VP his job

A failed cooling system at Vipshop's Nansha data centre knocked out service for 8M+ customers, caused ¥100M+ in losses, and cost the platform VP his job

Vipshop · 唯品会 · 2023-03-29

What happened

On March 29, 2023, Vipshop's e-commerce platform suffered a catastrophic outage when a cooling system failure at its Nansha IDC (Internet Data Centre) caused server room temperatures to rise rapidly, triggering a cascade of equipment shutdowns. The outage lasted over 12 hours, affecting more than 8 million customers and causing the company over ¥100 million in direct losses. Vipshop classified the incident as a P0 (most severe) level failure.

The company's response was swift and public: the head of the infrastructure platform department was fired, and the direct managers of the affected teams were held accountable. The incident also impacted Tencent, which hosted services in the same data centre and internally classified the event as a 'Level 1 accident,' resulting in criticism, demotions, and dismissals of several executives. The failure exposed the risks of single-data-centre architecture and inadequate disaster recovery planning.

Why it happened

  • Vipshop ran its entire platform on a single data centre with no real-time failover — a cooling failure in one room took down the whole company
  • The outage lasted 12 hours, suggesting that even basic disaster recovery procedures were not in place or could not be activated quickly enough
What it cost¥100M+ losses, 8M+ customers affected, VP firedcostly

The lesson

A single data centre is a single point of failure, and a cooling breakdown is not a freak event — it is a predictable risk every platform should survive

Sources

spotted an error? The club wants to know.

Comments · 0

    Sign in to join the comments.

    More like this

    Somewhere, someone solved the problem this company failed at. 2nd Opinion →