Facebook: The 6-Hour Outage — How DNS Took Down Everything
How a routine maintenance command took down Facebook, Instagram, and WhatsApp for 6 hours — and locked engineers out of their own systems.
The challenge
On October 4, 2021, Facebook's engineering team was auditing their global backbone network capacity. A routine command was sent to assess the availability of global backbone capacity. The command accidentally took down all the connections between Facebook's data centers and the internet.
The strategy
The BGP (Border Gateway Protocol) routes that tell the world's internet how to find Facebook's servers were withdrawn. DNS servers that translate facebook.com to IP addresses went offline. All Facebook services — Facebook, Instagram, WhatsApp, Oculus — disappeared from the internet simultaneously.
Want the full story, outcome and quiz?
Read how Facebook executed it, the results, key lessons and test yourself with a quiz. Free on CaseLearn.
Try the full case free →Key lessons (preview)
- Never let your internal tools depend on the same infrastructure as your external services.
- Every automated command that affects global infrastructure needs multiple confirmation steps.
- Physical access to critical systems must always be possible independent of software.
