Amazon Web Services: The 2017 S3 Outage That Broke the Internet
How a single mistyped command during routine maintenance took down a huge portion of the internet, since so many unrelated websites depended on the same AWS storage region.
The challenge
In February 2017, an AWS engineer was debugging a billing system issue affecting the S3 storage service in their US-EAST-1 region — one of AWS's largest and most heavily used data center regions. The intended fix required removing a small number of servers from one S3 subsystem using an established playbook command.
The strategy
The engineer executing the command made an input error, accidentally removing a much larger number of servers than intended, including servers supporting two other critical S3 subsystems that many other AWS services depended on internally. These subsystems required a full restart, and — critically — had not been fully restarted in years due to the region's rapid growth, meaning the restart process itself took far…
Want the full story, outcome and quiz?
Read how Amazon Web Services executed it, the results, key lessons and test yourself with a quiz. Free on CaseLearn.
Try the full case free →Key lessons (preview)
- A single human input error in a routine maintenance operation can cascade into internet-wide impact when infrastructure is this centralised.
- Systems that haven't been fully restarted in years may take far longer to recover than assumed — untested recovery paths are a hidden risk.
- Building safeguards that limit the blast radius of any single command (like capping how many servers one operation can remove) prevents human error from becoming catastrophic.
