CaseLearn logo
CASELEARN
Education that works
Try free
Home › Case studies › Slack
TechnologyEngineering FailureEnterprise SoftwareIntermediate2021

Slack: The 5-Hour Global Outage

How a routine database scaling operation cascaded into a 5-hour outage that took down Slack for millions of users worldwide.

The challenge

In January 2021, Slack experienced unusually high traffic as remote work surged. Engineers scaled up their AWS infrastructure to handle load, adding capacity to their database tier. This routine scaling operation, performed many times before without incident, triggered an unexpected cascading failure across Slack's backend systems.

The strategy

The scaling operation caused a subset of database instances to become overloaded during the transition, which triggered automatic failover mechanisms designed to protect the system. However, these automated protections themselves consumed additional capacity, creating a feedback loop where attempts to fix the overload made the underlying congestion worse rather than better.

Want the full story, outcome and quiz?

Read how Slack executed it, the results, key lessons and test yourself with a quiz. Free on CaseLearn.

Try the full case free →

Key lessons (preview)

More Engineering Failure cases

FacebookThe 6-Hour Outage — How DNS Took Down EverythingKnight Capital$440 Million Lost in 45 Minutes — The Costliest Bug in HistoryCrowdstrike8.5 Million Windows Machines Crashed by One UpdateTherac-25The Radiation Machine That Killed People Due to a Software Bug