us-east-1 and the retry storm inside AWS
An automated scaling activity triggered unexpected behavior in a large number of clients on AWS's internal network. The surge overwhelmed the devices connecting it to the main network, retries made it worse, and the internal DNS, monitoring and control plane services that everything leans on became slow or unreachable.
Times are approximate, in PST, from the public postmortem. The diagram is simplified.
What happened
AWS runs most services on its main network, and a set of foundational internal services, like monitoring, internal DNS and parts of the EC2 control plane, on a separate internal network. A fleet of networking devices connects the two.
At 7:30 AM PST an automated activity to scale capacity of one service in the main network set off unexpected behavior in a large number of clients on the internal network. They opened a surge of connections that overwhelmed the devices in between. Latency and errors rose, clients retried, and the retries kept the congestion going. Internal monitoring lost its data, so operators had to work from logs.
Instances that were already running kept running. What failed were the things that need the control plane: EC2 APIs, the console, STS, Connect, and anything that launches or scales. Moving internal DNS off the congested paths at 9:28 helped but did not fix it. Engineers then identified and cut the heaviest sources, congestion eased by 1:34 PM, and the devices were fully recovered at 2:22 PM. Some services took longer to work through their backlogs.
What I'd take from it
- Retries need a budget. Exponential backoff with jitter, plus a cap on the share of traffic that can be retries, keeps a slowdown from becoming congestion collapse.
- Monitoring must not share fate with what it watches. Operators lost their dashboards exactly when they needed them. Keep a thin, separate path for telemetry.
- Separate the data plane from the control plane. Running workloads survived because they did not need the control plane minute to minute. Design so that is true for your own services too.