Crowdstrike: 8.5 Million Windows Machines Crashed by One Update
How a single faulty configuration file update caused the largest IT outage in history — crashing airlines, hospitals, banks, and broadcasters simultaneously.
The challenge
On July 19, 2024, Crowdstrike pushed a routine content configuration update to their Falcon sensor — security software installed on millions of Windows machines globally. The update contained a logic error in a configuration file. Unlike traditional software, security sensors run at the kernel level with system-level privileges.
The strategy
The configuration file update bypassed traditional software testing because it was classified as a content update, not a code update. Content updates had a faster, less rigorous validation pipeline. The update was pushed globally and simultaneously — the standard procedure for security updates which need to reach all machines quickly to protect against threats.
Want the full story, outcome and quiz?
Read how Crowdstrike executed it, the results, key lessons and test yourself with a quiz. Free on CaseLearn.
Try the full case free →Key lessons (preview)
- Content updates need the same rigorous testing as code updates when they run at kernel level.
- Staged rollouts — releasing to 1% then 10% then 100% — would have caught this before global impact.
- Systems with physical recovery requirements at scale are catastrophic — remote remediation must always be possible.
