Slack: The 5-Hour Global Outage
How a routine database scaling operation cascaded into a 5-hour outage that took down Slack for millions of users worldwide.
The challenge
In January 2021, Slack experienced unusually high traffic as remote work surged. Engineers scaled up their AWS infrastructure to handle load, adding capacity to their database tier. This routine scaling operation, performed many times before without incident, triggered an unexpected cascading failure across Slack's backend systems.
The strategy
The scaling operation caused a subset of database instances to become overloaded during the transition, which triggered automatic failover mechanisms designed to protect the system. However, these automated protections themselves consumed additional capacity, creating a feedback loop where attempts to fix the overload made the underlying congestion worse rather than better.
Want the full story, outcome and quiz?
Read how Slack executed it, the results, key lessons and test yourself with a quiz. Free on CaseLearn.
Try the full case free →Key lessons (preview)
- Automated failover and recovery systems can create feedback loops that worsen an incident if not carefully bounded.
- Routine operations that have succeeded many times before can still trigger unexpected cascading failures at different scale or timing.
- Sometimes the correct incident response is manual, careful intervention rather than trusting automated systems to self-correct.
