liveStatus reported "working" for a member that did nothing for 76 minutes — the health snapshot has no progress signal #410

Open
opened 2026-09-10 04:32:22 +02:00 by ltms · 0 comments
Owner

The incident

A member on fleet01 ran for 76 minutes producing nothing. liveStatus: working reported it as
healthy the whole time. It was found because a lead thought to check file mtimes in the member's
worktree by hand, then confirmed nothing was uncommitted and its work was already at 0200c9c.

After stopping it, that host went from load 6.37 to 4.97 and from 376 MB free to 1117 MB. So the
stuck member was not merely idle — it was consuming the resource that another finding this week
showed is the axis a concurrency race turns on (#399: the threshold is about 2x cores). A silent
spinner is not a neutral failure.

Why the report cannot catch this

liveStatus answers "is the backend session alive and not blocked". A member burning CPU in a loop
satisfies that. Nothing in the health snapshot reads anything the member produces, so there is no
value that can distinguish these two states:

  • a member thinking hard about a difficult unit
  • a member in a loop, or waiting on something that will never arrive

Both report working. This is the same shape as #404 and #400, filed this week: a status field
derived from a source that cannot see the thing the field claims to describe.
Here the field is
honest about what it measures; the problem is that no field measures progress at all.

The gap is already recorded as a standing caution — "liveStatus working is not progress, check file
mtimes in its worktree" — which is the tell that it should be a report instead of a habit. A caution
that says "remember to check by hand" is load-bearing documentation for a missing feature, and it
fails exactly when the lead is busy, which is when members are running.

The goal

The health snapshot should carry a signal a lead can read without knowing to go looking. Enough
to tell "working" apart from "alive but producing nothing for N minutes".

Not an automatic kill. A stuck member and a slow one look identical for the first few minutes, and
this fleet runs long units on purpose. Report it; let the lead decide.

Candidate signals, none of them chosen

Deliberately not picking a mechanism — my briefs have caused defects five times by naming one, so
treat all of these as options to evaluate and reject:

  • Newest mtime under the member's worktree. Cheap, already the manual check, and it is what
    actually found this. Weak on a member whose work is genuinely all reading — and reading a large
    codebase for 20 minutes is legitimate here.
  • Turn-scoped output volume — bytes the backend has emitted since the turn started. Closer to
    real progress, but a member can emit plenty of text while looping.
  • CPU time of the member's process tree. Would have caught this specific case loudly. Needs care:
    high CPU is also what a healthy build looks like, so on its own it inverts the signal.
  • Two consecutive quiet ticks, the shape DELEGATION_ORPHANED already uses, so a single slow
    moment does not raise it.

Whatever is chosen, the field must state what it measured and when, not a verdict. "No file
touched in 41 minutes" is actionable. "Possibly stuck" is not, and will be ignored within a week.

Acceptance

  • fleet_list and/or fleet_status report a progress-related value per member, with its units and
    its age visible.
  • A test that a member with no observable progress is reported differently from one making progress.
    A test that only checks the field exists is not enough — see #404, where exactly that test let the
    feature ship able to go permanently dead.
  • The manual caution can be replaced by pointing at the field. If it cannot, the field is not yet
    useful.
  • No automatic teardown in this ticket.

Note on scope

healthCoverage still reports detection-only, which is accurate and should stay accurate. This
ticket adds one detection signal; it does not start remediation.

Found by the fleet01 lead's stop-and-measure, on a host where I had flagged a suspicious member.
Cross-referenced: #399 (why a spinner is not harmless), #404 and #400 (status fields that do not
measure what they claim).

## The incident A member on fleet01 ran for **76 minutes** producing nothing. `liveStatus: working` reported it as healthy the whole time. It was found because a lead thought to check file mtimes in the member's worktree by hand, then confirmed nothing was uncommitted and its work was already at `0200c9c`. After stopping it, that host went from load 6.37 to 4.97 and from 376 MB free to 1117 MB. So the stuck member was not merely idle — it was consuming the resource that another finding this week showed is the axis a concurrency race turns on (#399: the threshold is about 2x cores). A silent spinner is not a neutral failure. ## Why the report cannot catch this `liveStatus` answers "is the backend session alive and not blocked". A member burning CPU in a loop satisfies that. Nothing in the health snapshot reads anything the member **produces**, so there is no value that can distinguish these two states: - a member thinking hard about a difficult unit - a member in a loop, or waiting on something that will never arrive Both report `working`. This is the same shape as #404 and #400, filed this week: **a status field derived from a source that cannot see the thing the field claims to describe.** Here the field is honest about what it measures; the problem is that no field measures progress at all. The gap is already recorded as a standing caution — "liveStatus working is not progress, check file mtimes in its worktree" — which is the tell that it should be a report instead of a habit. A caution that says "remember to check by hand" is load-bearing documentation for a missing feature, and it fails exactly when the lead is busy, which is when members are running. ## The goal The health snapshot should carry a signal a lead can read **without** knowing to go looking. Enough to tell "working" apart from "alive but producing nothing for N minutes". Not an automatic kill. A stuck member and a slow one look identical for the first few minutes, and this fleet runs long units on purpose. Report it; let the lead decide. ## Candidate signals, none of them chosen Deliberately not picking a mechanism — my briefs have caused defects five times by naming one, so treat all of these as options to evaluate and reject: - **Newest mtime under the member's worktree.** Cheap, already the manual check, and it is what actually found this. Weak on a member whose work is genuinely all reading — and reading a large codebase for 20 minutes is legitimate here. - **Turn-scoped output volume** — bytes the backend has emitted since the turn started. Closer to real progress, but a member can emit plenty of text while looping. - **CPU time of the member's process tree.** Would have caught this specific case loudly. Needs care: high CPU is also what a healthy build looks like, so on its own it inverts the signal. - **Two consecutive quiet ticks**, the shape `DELEGATION_ORPHANED` already uses, so a single slow moment does not raise it. Whatever is chosen, the field must state **what it measured and when**, not a verdict. "No file touched in 41 minutes" is actionable. "Possibly stuck" is not, and will be ignored within a week. ## Acceptance - `fleet_list` and/or `fleet_status` report a progress-related value per member, with its units and its age visible. - A test that a member with no observable progress is reported differently from one making progress. A test that only checks the field exists is not enough — see #404, where exactly that test let the feature ship able to go permanently dead. - The manual caution can be replaced by pointing at the field. If it cannot, the field is not yet useful. - No automatic teardown in this ticket. ## Note on scope `healthCoverage` still reports `detection-only`, which is accurate and should stay accurate. This ticket adds one detection signal; it does not start remediation. Found by the fleet01 lead's stop-and-measure, on a host where I had flagged a suspicious member. Cross-referenced: #399 (why a spinner is not harmless), #404 and #400 (status fields that do not measure what they claim).
Sign in to join this conversation.
1 Participants
Notifications
Due Date
No due date set.
Dependencies

No dependencies set.

Reference: fleet/fleetd#410