LOOK ACROSS THE INCIDENTS
Different failures.
Shared lessons.
Put cases side by side. Find the common assumptions—and the safeguards that could break the chain.
Loading comparison…
LOOK ACROSS THE INCIDENTS
Put cases side by side. Find the common assumptions—and the safeguards that could break the chain.
Loading comparison…
SHARED FAILURE PATTERNS
These editorial tags are common to every selected case. Compare the details below before drawing parallels.
On smaller screens, swipe the comparison horizontally.
| Case file | Mars Pathfinder ↗July 1997 | Knight Capital ↗August 1, 2012 |
|---|---|---|
| Failure pattern | Shared dependencies · Hidden assumptions | Unsafe changes · Hidden assumptions |
| What happened | After landing on Mars, Pathfinder experienced computer resets when time-critical work missed its deadline. Resetting protected the spacecraft, but interrupted planned activities. This was a recoverable software incident—not the loss of the mission. | Knight Capital’s order router sent more than four million orders while attempting to fill just 212 customer orders. In the first 45 minutes of trading, the firm accumulated unwanted positions and lost more than $460 million. |
| Why it spread | A low-priority meteorological task held a shared resource inside the operating system’s communication machinery. A higher-priority data-distribution task needed that resource and blocked. Meanwhile, medium-priority work prevented the low-priority task from running long enough to release it. The bus scheduler detected unfinished work and triggered a reset. | An incomplete deployment and the reuse of a flag activated obsolete trading logic. Controls failed to stop the resulting orders. |
| Impact in context | Priority inversion — a scheduling failure, recovered | $460M+ — trading loss |
| Recovery | The team reproduced the problem in the laboratory using tracing facilities deliberately retained in the flight software. Engineers enabled priority inheritance for the relevant class of semaphores, analyzed the wider effects with the operating-system supplier, and tested extensively before changing the spacecraft. Recovery depended on observability and validation as much as on the fix itself. | Knight stopped the problematic trading and unwound positions. The incident exposed gaps in deployment verification and market-access safeguards. |
| Safeguards to discuss | Trace shared resources as well as task priorities. Keep a way to reproduce and inspect rare scheduling failures. Test the wider effects of a recovery change before applying it remotely. | Verify deployment consistency across every node. Remove dead code before reusing its controls. Enforce independent limits on automated actions. |
| Primary source | Glenn Reeves, JPL · Firsthand account, hosted by Cornell ↗ | U.S. SEC · Enforcement release ↗ |
Take one shared pattern into a rehearsal or review the evidence behind your own safeguards.
Review your safeguards →