fin1te

Outages/AWS··~7 hours

us-east-1 and the retry storm inside AWS

An automated scaling activity triggered unexpected behavior in a large number of clients on AWS's internal network. The surge overwhelmed the devices connecting it to the main network, retries made it worse, and the internal DNS, monitoring and control plane services that everything leans on became slow or unreachable.

Fig.AWS us-east-1, 7 December 20218 components · 8 linksOpen in topo ↗
7 stepsPress play, or step through with → and ←
Impact 6h 52m
07:3014:22
launch, scaletriggersconnection surgeMain AWS networkInternal networkCustomersAPIs, consoleControl planeEC2 API, STS, consoleRunning workloadsEC2 instances, S3, DynamoDBAutomated scalingone serviceNetwork devicesbetween the two networksInternal clientsa large fleetInternal DNSservice discoveryMonitoringinternal telemetry

Times are approximate, in PST, from the public postmortem. The diagram is simplified.

What happened

AWS runs most services on its main network, and a set of foundational internal services, like monitoring, internal DNS and parts of the EC2 control plane, on a separate internal network. A fleet of networking devices connects the two.

At 7:30 AM PST an automated activity to scale capacity of one service in the main network set off unexpected behavior in a large number of clients on the internal network. They opened a surge of connections that overwhelmed the devices in between. Latency and errors rose, clients retried, and the retries kept the congestion going. Internal monitoring lost its data, so operators had to work from logs.

Instances that were already running kept running. What failed were the things that need the control plane: EC2 APIs, the console, STS, Connect, and anything that launches or scales. Moving internal DNS off the congested paths at 9:28 helped but did not fix it. Engineers then identified and cut the heaviest sources, congestion eased by 1:34 PM, and the devices were fully recovered at 2:22 PM. Some services took longer to work through their backlogs.

What I'd take from it

  1. Retries need a budget. Exponential backoff with jitter, plus a cap on the share of traffic that can be retries, keeps a slowdown from becoming congestion collapse.
  2. Monitoring must not share fate with what it watches. Operators lost their dashboards exactly when they needed them. Keep a thin, separate path for telemetry.
  3. Separate the data plane from the control plane. Running workloads survived because they did not need the control plane minute to minute. Design so that is true for your own services too.

Sources