fin1te

Outages/Facebook··~6 hours

The day Facebook withdrew itself from the internet

A command meant to measure backbone capacity disconnected every Facebook data center. The DNS servers, built to step aside when they lose the data centers, withdrew their BGP routes, and facebook.com vanished from the internet along with the tools engineers needed to fix it.

Fig.Facebook, 4 October 202110 components · 10 linksOpen in topo ↗
6 stepsPress play, or step through with → and ←
Impact 5h 49m
15:3921:28
DNS lookupshealthy?routescommandshould blockphysical accessFacebook networkPeopleFacebook, Instagram, WhatsAppPublic DNS resolversISPs, 1.1.1.1, 8.8.8.8InternetBGP peersBGP speakersannounce DNS prefixesAuthoritative DNSedge locationsMaintenance commandassess backbone capacityAudit toolblocks unsafe commandsBackbonelinks between data centersData centersevery app, internal toolsEngineerssent on site

Times are approximate, in UTC, from the public postmortem. The diagram is simplified.

What happened

During routine maintenance, an engineer ran a command to assess the capacity of the global backbone. It took down every backbone connection instead. An audit tool exists to stop commands like this one, and a bug in it let the command through.

Facebook's authoritative DNS servers run in the edge locations and check that they can reach the data centers. When they could not, they did what they were designed to do: they stopped advertising their BGP routes, so traffic would go to a healthier location. Every location failed the same check at once, so the whole DNS service disappeared from the internet. Resolvers around the world started returning errors and retried hard, and public resolvers saw query volumes many times their normal level.

The same network carried the internal tools, remote access and much of the badge and door system. Engineers had to travel to the data centers and get physical access to the routers, which the security built into those sites slowed down on purpose. Once the backbone was back, traffic was brought back in stages so the sudden load would not trip power and caches across the fleet.

What I'd take from it

  1. A health check can take the whole fleet down. Withdrawing routes when you are unhealthy is right for one location and a disaster for all of them. A rule like this needs a floor: never withdraw the last few.
  2. Keep a way in that does not depend on the thing that is broken. Out-of-band access to routers, on a separate network with separate credentials, turns a six-hour outage into a short one.
  3. Test the guardrail itself. The audit tool was the only thing between a typo and the backbone. Tools that block dangerous commands need their own tests with known-bad commands.

Sources