fin1te

Outages/GitLab··~18 hours

rm -rf on the wrong database, and five backups that weren't

While fighting replication lag late at night, an engineer deleted the data directory on the primary database instead of the replica. The team then found that none of the backup methods had been working. A snapshot taken six hours earlier for an unrelated reason saved GitLab.com, at the cost of six hours of data.

Fig.GitLab.com, 31 January 20179 components · 8 linksOpen in topo ↗
7 stepsPress play, or step through with → and ←
Impact 18h 47m
17:2018:14
replicationrm -rfslow restoreSpammersheavy loadUsersweb, API, gitGitLab.comRails, API, Gitdb1 · primaryPostgreSQL 9.6db2 · replicafell behindEngineerterminal open on db1pg_dump backupsfailing silentlyDisk snapshotsnever enabledLVM snapshottaken for staging

Times are approximate, in UTC, from the public postmortem. The diagram is simplified.

What happened

On 31 January GitLab.com was under load from spammers, and the extra load made the secondary database, db2, fall behind the primary, db1, until replication broke. To reseed db2, an engineer needed to empty its data directory.

At 23:27 UTC, late in a long evening, the engineer ran rm -rf on the data directory in the terminal for db1, the primary, believing it was db2. It was stopped within seconds, but about 300 GB of the roughly 310 GB of data were already gone.

Then the backups. The nightly pg_dump had been failing silently because it ran an older PostgreSQL version than the server, and the failure emails were rejected by the mail system. Disk snapshots were not enabled for the database servers. What saved them was an LVM snapshot taken around 17:20 for staging. Copying it back over slow disks took most of the next day. GitLab.com came back around 18:14 UTC on 1 February, missing about six hours of data, including roughly 5,000 projects, 5,000 comments and 700 new accounts.

What I'd take from it

  1. A backup you haven't restored is a hope. Restore to a scratch machine on a schedule and alert when it fails or takes longer than it should.
  2. Make the dangerous host look dangerous. Different prompt colors for primaries, and destructive commands that ask for the host name, cost nothing.
  3. Tired people should not be the last line of defence. Late-night database surgery deserves a second person reading along.

Sources