← BACK TO THE ARCHIVE

CASE FILE 013 / SOFTWARE FAILURE

The urgent task that had to wait.

PRIMARY SOURCE LINKED · ABOUT 3 MIN

A spacecraft on Mars kept resetting. The route to recovery ran through a shared lock.

Incident dateJuly 1997
a scheduling failure, recoveredPriority inversion
CollectionTechnology failures
event_sequence / CASE 013July 1997

FOLLOW THE FAILURE

How a fault became an incident.

A condensed sequence, not a minute-by-minute reconstruction.

  1. 01
    THE TRIGGER

    Low-priority task holds a resource

  2. 02
    THE FAILURE SPREADS

    Medium-priority work delays its release

  3. 03
    THE CONSEQUENCE

    Critical task misses its deadline

  4. 04
Source record: Glenn Reeves, JPL · Firsthand account, hosted by Cornell ↗
THE CRITICAL BREAK

Medium-priority work delays its release

A priority label does not remove dependencies. Priority inheritance can let the resource holder finish so the urgent task can proceed.

Editorial interpretation of the failure pattern.

01 / THE INCIDENT

What happened

After landing on Mars, Pathfinder experienced computer resets when time-critical work missed its deadline. Resetting protected the spacecraft, but interrupted planned activities. This was a recoverable software incident—not the loss of the mission.

Evidence: Glenn Reeves, JPL · Firsthand account, hosted by Cornell ↗

02 / THE FAILURE CHAIN

Why it happened

A low-priority meteorological task held a shared resource inside the operating system’s communication machinery. A higher-priority data-distribution task needed that resource and blocked. Meanwhile, medium-priority work prevented the low-priority task from running long enough to release it. The bus scheduler detected unfinished work and triggered a reset.

QUESTION FOR YOUR TEAM

Which low-priority task holds something your critical path needs?

03 / THE AFTERMATH

Recovery—and its limits

The team reproduced the problem in the laboratory using tracing facilities deliberately retained in the flight software. Engineers enabled priority inheritance for the relevant class of semaphores, analyzed the wider effects with the operating-system supplier, and tested extensively before changing the spacecraft. Recovery depended on observability and validation as much as on the fix itself.

SCHEDULER_LAB / 013INTERACTIVE MODEL

CHANGE ONE RULE. FOLLOW THE CONSEQUENCE.

Urgent doesn’t mean unblocked.

Three tasks. One shared lock. Step through the same situation with and without priority inheritance.

Changing mode restarts the model. Steps show causal order, not elapsed time.

STEP 1 / 6

Low-priority work acquires the lock.

A shared resource is locked while a low-priority task works. Only its owner can release it.

Critical work has not arrived
High priority / Critical data workNot ready
Medium priority / Other ready workNot ready
Low priority / Resource ownerRunning
CPU
Low-priority task
LOCK OWNER
Low-priority task
How this maps to Pathfinder

In Glenn Reeves’s account, the low-priority ASI/MET task held a semaphore within the operating system’s select mechanism. The higher-priority bc_dist task blocked, while medium-priority work delayed the resource holder. The bc_sched task detected the missed deadline.

This three-task model omits the actual bus cycles, task timings, and operating-system internals. The tested fix changed semaphore options; it was not simply a blanket increase of every task’s priority.

Read the software lead’s account

04 / TAKE IT WITH YOU

Where could the chain break?

Safeguards to investigate—not a claim that one change would certainly have prevented this incident.

01
SAFEGUARD 1Trace shared resources as well as task priorities.
02
SAFEGUARD 2Keep a way to reproduce and inspect rare scheduling failures.
03
SAFEGUARD 3Test the wider effects of a recovery change before applying it remotely.

Go to the source.

A condensed editorial account. The lessons are our interpretation, not quotations. Consult the primary source for the full technical record.

Glenn Reeves, JPL · Firsthand account, hosted by Cornell ↗
Practice your response ↗
Compare this failure with another ↗

ONE QUESTION BEFORE YOU GO

Why can an urgent task wait behind less urgent work?

TAKE IT TO YOUR TEAM

Which low-priority task holds something your critical path needs?

next_step.txtREAD → PRACTICE

TAKE THE LESSON FURTHER

Challenge the fallback.

Which failure could take out both your primary and backup paths?

The rehearsal is fictional; the case is historical. The review starts with unanswered questions.

CONTINUE THE CONNECTION

TRAIL 05The machine followed the rules.Three cases, one recurring question ↗