Outages
Famous outages, replayed one step at a time.
Each one is rebuilt from the public postmortem as a live diagram. Press play and watch it fail: components go red, data stops moving, and the timeline shows how long it hurt. Then what I would take from it.
CrowdStrike
A content update that crashed 8.5 million Windows machines
A sensor content file with one field fewer than the interpreter expected, read in the kernel.
78 min, then days
AWS
us-east-1 and the retry storm inside AWS
An automated scaling job triggered a connection surge that congested the devices between two internal networks.
~7 hours
The day Facebook withdrew itself from the internet
A maintenance command took down the backbone, and DNS servers pulled their own BGP routes.
~6 hours
Cloudflare
One regular expression, every CPU on the edge
A WAF rule with a backtracking regex pinned CPU at 100% on every edge server.
27 minutes
GitLab
rm -rf on the wrong database, and five backups that weren't
An engineer wiped the primary database instead of the replica, and the backups turned out not to work.
~18 hours
Diagrams are simplified from the postmortems linked on each page. Built with topo.