FOLLOW THE FAILURE
How a fault became an incident.
A condensed sequence, not a minute-by-minute reconstruction.
- 01THE TRIGGER
Production data removed
- 02THE FAILURE SPREADS
Recovery paths disappoint
- 03THE CONSEQUENCE
Older snapshot restored
- 04THE AFTERMATHRead the recovery and its limits ↓
Recovery paths disappoint
The recovery outcome matters: usable data, measured age, and elapsed restore time.
Editorial interpretation of the failure pattern.01 / THE INCIDENT
What happened
During work on a database problem, GitLab accidentally removed production database data. Recovery revealed that several backup and replication approaches were not working as expected.
Evidence: GitLab · Incident postmortem ↗02 / THE FAILURE CHAIN
Why it happened
A destructive operation on the wrong database combined with unreliable recovery mechanisms turned a maintenance issue into an extended outage.
When was the last successful restore, and what did it prove?
03 / THE AFTERMATH
Recovery—and its limits
GitLab restored a roughly six-hour-old snapshot and documented the recovery publicly. Database records created in that interval were lost; Git repositories were not the affected data.
04 / TAKE IT WITH YOU
Where could the chain break?
Safeguards to investigate—not a claim that one change would certainly have prevented this incident.
Go to the source.
A condensed editorial account. The lessons are our interpretation, not quotations. Consult the primary source for the full technical record.
GitLab · Incident postmortem ↗