← BACK TO THE ARCHIVE

CASE FILE 004 / DATA LOSS

The backup plan needed a backup plan.

PRIMARY SOURCE LINKED · ABOUT 3 MIN

A deleted production database revealed a much bigger recovery problem.

Incident dateJanuary 31, 2017
database data lost6 hours
CollectionTechnology failures
event_sequence / CASE 004January 31, 2017

FOLLOW THE FAILURE

How a fault became an incident.

A condensed sequence, not a minute-by-minute reconstruction.

  1. 01
    THE TRIGGER

    Production data removed

  2. 02
    THE FAILURE SPREADS

    Recovery paths disappoint

  3. 03
    THE CONSEQUENCE

    Older snapshot restored

  4. 04
Source record: GitLab · Incident postmortem ↗
THE CRITICAL BREAK

Recovery paths disappoint

The recovery outcome matters: usable data, measured age, and elapsed restore time.

Editorial interpretation of the failure pattern.

01 / THE INCIDENT

What happened

During work on a database problem, GitLab accidentally removed production database data. Recovery revealed that several backup and replication approaches were not working as expected.

Evidence: GitLab · Incident postmortem ↗

02 / THE FAILURE CHAIN

Why it happened

A destructive operation on the wrong database combined with unreliable recovery mechanisms turned a maintenance issue into an extended outage.

QUESTION FOR YOUR TEAM

When was the last successful restore, and what did it prove?

03 / THE AFTERMATH

Recovery—and its limits

GitLab restored a roughly six-hour-old snapshot and documented the recovery publicly. Database records created in that interval were lost; Git repositories were not the affected data.

04 / TAKE IT WITH YOU

Where could the chain break?

Safeguards to investigate—not a claim that one change would certainly have prevented this incident.

01
SAFEGUARD 1Regularly restore backups into an isolated environment.
02
SAFEGUARD 2Make production environments unmistakable.
03
SAFEGUARD 3Require safeguards around destructive operations.

Go to the source.

A condensed editorial account. The lessons are our interpretation, not quotations. Consult the primary source for the full technical record.

GitLab · Incident postmortem ↗
Practice your response ↗
Compare this failure with another ↗Rehearse a database recovery →

ONE QUESTION BEFORE YOU GO

What provides useful evidence that a backup works?

TAKE IT TO YOUR TEAM

When was the last successful restore, and what did it prove?

next_step.txtREAD → PRACTICE

TAKE THE LESSON FURTHER

Prove the way back.

What evidence shows your recovery path actually works?

The rehearsal is fictional; the case is historical. The review starts with unanswered questions.

CONTINUE THE CONNECTION

TRAIL 02The way back.Three cases, one recurring question ↗