FOLLOW THE FAILURE
How a fault became an incident.
A condensed sequence, not a minute-by-minute reconstruction.
- 01THE TRIGGER
Low-priority task holds a resource
- 02THE FAILURE SPREADS
Medium-priority work delays its release
- 03THE CONSEQUENCE
Critical task misses its deadline
- 04THE AFTERMATHRead the recovery and its limits ↓
Medium-priority work delays its release
A priority label does not remove dependencies. Priority inheritance can let the resource holder finish so the urgent task can proceed.
Editorial interpretation of the failure pattern.01 / THE INCIDENT
What happened
After landing on Mars, Pathfinder experienced computer resets when time-critical work missed its deadline. Resetting protected the spacecraft, but interrupted planned activities. This was a recoverable software incident—not the loss of the mission.
Evidence: Glenn Reeves, JPL · Firsthand account, hosted by Cornell ↗02 / THE FAILURE CHAIN
Why it happened
A low-priority meteorological task held a shared resource inside the operating system’s communication machinery. A higher-priority data-distribution task needed that resource and blocked. Meanwhile, medium-priority work prevented the low-priority task from running long enough to release it. The bus scheduler detected unfinished work and triggered a reset.
Which low-priority task holds something your critical path needs?
03 / THE AFTERMATH
Recovery—and its limits
The team reproduced the problem in the laboratory using tracing facilities deliberately retained in the flight software. Engineers enabled priority inheritance for the relevant class of semaphores, analyzed the wider effects with the operating-system supplier, and tested extensively before changing the spacecraft. Recovery depended on observability and validation as much as on the fix itself.
CHANGE ONE RULE. FOLLOW THE CONSEQUENCE.
Urgent doesn’t mean unblocked.
Three tasks. One shared lock. Step through the same situation with and without priority inheritance.
Changing mode restarts the model. Steps show causal order, not elapsed time.
Low-priority work acquires the lock.
A shared resource is locked while a low-priority task works. Only its owner can release it.
Critical work has not arrived- CPU
- Low-priority task
- LOCK OWNER
- Low-priority task
How this maps to Pathfinder
In Glenn Reeves’s account, the low-priority ASI/MET task held a semaphore within the operating system’s select mechanism. The higher-priority bc_dist task blocked, while medium-priority work delayed the resource holder. The bc_sched task detected the missed deadline.
This three-task model omits the actual bus cycles, task timings, and operating-system internals. The tested fix changed semaphore options; it was not simply a blanket increase of every task’s priority.
Read the software lead’s account04 / TAKE IT WITH YOU
Where could the chain break?
Safeguards to investigate—not a claim that one change would certainly have prevented this incident.
Go to the source.
A condensed editorial account. The lessons are our interpretation, not quotations. Consult the primary source for the full technical record.
Glenn Reeves, JPL · Firsthand account, hosted by Cornell ↗