fin1te

Outages

Famous outages, replayed one step at a time.

Each one is rebuilt from the public postmortem as a live diagram. Press play and watch it fail: components go red, data stops moving, and the timeline shows how long it hurt. Then what I would take from it.

CrowdStrike

A content update that crashed 8.5 million Windows machines

A sensor content file with one field fewer than the interpreter expected, read in the kernel.

78 min, then days

AWS

us-east-1 and the retry storm inside AWS

An automated scaling job triggered a connection surge that congested the devices between two internal networks.

~7 hours

Facebook

The day Facebook withdrew itself from the internet

A maintenance command took down the backbone, and DNS servers pulled their own BGP routes.

~6 hours

Cloudflare

One regular expression, every CPU on the edge

A WAF rule with a backtracking regex pinned CPU at 100% on every edge server.

27 minutes

GitLab

rm -rf on the wrong database, and five backups that weren't

An engineer wiped the primary database instead of the replica, and the backups turned out not to work.

~18 hours

Diagrams are simplified from the postmortems linked on each page. Built with topo.