CaseLearn logo
CASELEARN
Education that works
Try free
Home › Case studies › Amazon Web Services
TechnologyEngineering FailureCloud ComputingIntermediate2017

Amazon Web Services: The 2017 S3 Outage That Broke the Internet

How a single mistyped command during routine maintenance took down a huge portion of the internet, since so many unrelated websites depended on the same AWS storage region.

The challenge

In February 2017, an AWS engineer was debugging a billing system issue affecting the S3 storage service in their US-EAST-1 region — one of AWS's largest and most heavily used data center regions. The intended fix required removing a small number of servers from one S3 subsystem using an established playbook command.

The strategy

The engineer executing the command made an input error, accidentally removing a much larger number of servers than intended, including servers supporting two other critical S3 subsystems that many other AWS services depended on internally. These subsystems required a full restart, and — critically — had not been fully restarted in years due to the region's rapid growth, meaning the restart process itself took far…

Want the full story, outcome and quiz?

Read how Amazon Web Services executed it, the results, key lessons and test yourself with a quiz. Free on CaseLearn.

Try the full case free →

Key lessons (preview)

More Engineering Failure cases

FacebookThe 6-Hour Outage — How DNS Took Down EverythingKnight Capital$440 Million Lost in 45 Minutes — The Costliest Bug in HistoryCrowdstrike8.5 Million Windows Machines Crashed by One UpdateTherac-25The Radiation Machine That Killed People Due to a Software Bug