A sentinel that conflates "measured: no" with "could not measure" — three instances, three subsystems #497

Open
opened 2026-09-12 04:12:52 +02:00 by ltms · 2 comments
Owner

Named by the fleet01 lead, 2026-09-12. Two of the three instances are theirs; one is mine. Filing it separately from #494 on their insistence, because putting it there would misdirect the fix — see the last section.

The shape

One symbol carries two states that need opposite handling, and the reading a caller naturally takes is the confident one.

# Instance Reads as Truth
1 fleet_profiles on fleet01's jar returns no exhaustionDetectionArmed field "nothing to report" "this daemon cannot tell you" (their #479, layer 2)
2 git cat-file -t 49a404d → fatal: Not a valid object name for a commit that exists on the forge "no such object" "not in my clone". Same for an empty git log --all -S: no such commit and outside my fetch window are one result
3 detect_supervisor returns none (PR #495, scripts/redeploy-fleetd.sh:123) "proven unsupervised" also "could not tell" — and none falls through to raw kill + nohup

Instance 1 is the reason I now rank a log line above a status field: the log outlives jar drift, where a missing API field is indistinguishable from a daemon that has nothing to say.

Why this is NOT #494, and filing it there would break the fix

#494's rule is a field that reads as a measurement must come from the measurement. Its fix is to compute the value: elapsed=0.438s budget=20s.

This is the inverse. Here the value is honestly derived — none really is what the probe returned. The problem is that the vocabulary is one symbol short, so no amount of computing it correctly helps. The fix is a third state, not a better number.

The fleet01 lead's warning, and it is a real risk: a worker reads #494's rule, confirms none is genuinely what the probe returned, and closes the ticket.

The fix shape is #415's, not #494's — with one translation caveat

#415's mechanism (be123d0 in this tree, CompletionResolver + Fleetd) made the new decision a required parameter with no defaulted overload, so a future case cannot compile without someone stating what it means. A required parameter is a compile error; a defaulted one is a silent survivor that a green suite vouches for.

Applied here: three states — SUPERVISED / UNSUPERVISED / UNKNOWN — is only half. The half that matters is UNKNOWN must not be able to reach kill by writing nothing. If the enum is added but the call site keeps an else or a default: arm, the unsafe path is still the one you fall into by default, and every test still passes.

The caveat: instance 3 is a bash script, and bash has no exhaustiveness check. A case with no *) arm silently matches nothing and continues. So the compiler cannot carry this one; the substitute is that every switch over the value must carry a *) die … arm.

I measured how that stands today in PR #495 (dcd5052):

CONTROL, ')' occurrences in the file:  149   (non-zero, the grep works)

require_drivable_supervisor  :144   arms: launchd|systemd|none, ambiguous, *) die     <- has one
report state                 :308   arms: launchd, systemd, none                      <- NO *)
stop                         :419   arms: launchd, systemd, none                      <- NO *)
start                        :477   arms: launchd, systemd, none                      <- NO *)

The guard at :305 runs once. The three switches that actually drive the daemon each enumerate exactly three values and silently do nothing for a fourth. So adding a state without also adding the die arms produces: nothing matches in the stop switch, the daemon is never stopped, nothing matches in the start switch — and the script reports no error at any point.

Asked for

  1. detect_supervisor gains an undrivable state, and require_drivable_supervisor refuses it. (Already asked on PR #495.)
  2. Every switch over that value gets a *) die … arm — :308, :419, :477 — so a future fourth state fails loudly at each site instead of skipping the stop or the start.
  3. Instances 1 and 2 are not code changes here; they are why the rule exists. Whoever revisits #479 should read this.

Testing note — two different tests, and only one of them is the useful one

From the fleet01 lead, and it is the same gap as asserting the nudge happens versus asserting the wait waits:

The interesting case is "the probe could not tell", which usually means a timeout, a permission denial, or an unreadable entry — all easy to stub and hard to make real. A test that constructs UNKNOWN directly proves the switch handles it; it does not prove the probe ever returns it. Only the second catches a probe that quietly maps its own failures to none before the state is ever consulted.

Both tests are needed. The second is the one that would have caught instance 3.

Related: #492, PR #495, #479, #415, #494 (the inverse family — a constant pretending to be measured).

Named by the fleet01 lead, 2026-09-12. Two of the three instances are theirs; one is mine. Filing it separately from #494 **on their insistence**, because putting it there would misdirect the fix — see the last section. ## The shape **One symbol carries two states that need opposite handling, and the reading a caller naturally takes is the confident one.** | # | Instance | Reads as | Truth | |---|---|---|---| | 1 | `fleet_profiles` on fleet01's jar returns **no** `exhaustionDetectionArmed` field | "nothing to report" | "this daemon cannot tell you" (their #479, layer 2) | | 2 | `git cat-file -t 49a404d` → `fatal: Not a valid object name` for a commit that exists on the forge | "no such object" | "not in my clone". Same for an empty `git log --all -S`: *no such commit* and *outside my fetch window* are one result | | 3 | `detect_supervisor` returns `none` (PR #495, `scripts/redeploy-fleetd.sh:123`) | "proven unsupervised" | also "could not tell" — and `none` falls through to raw `kill` + `nohup` | Instance 1 is the reason I now rank a log line above a status field: the log outlives jar drift, where a missing API field is indistinguishable from a daemon that has nothing to say. ## Why this is NOT #494, and filing it there would break the fix #494's rule is *a field that reads as a measurement must come from the measurement*. Its fix is to compute the value: `elapsed=0.438s budget=20s`. This is the **inverse**. Here the value is honestly derived — `none` really is what the probe returned. The problem is that the vocabulary is one symbol short, so no amount of computing it correctly helps. **The fix is a third state, not a better number.** The fleet01 lead's warning, and it is a real risk: a worker reads #494's rule, confirms `none` is genuinely what the probe returned, and closes the ticket. ## The fix shape is #415's, not #494's — with one translation caveat #415's mechanism (`be123d0` in this tree, `CompletionResolver` + `Fleetd`) made the new decision a **required parameter with no defaulted overload**, so a future case cannot compile without someone stating what it means. A required parameter is a compile error; a defaulted one is a silent survivor that a green suite vouches for. Applied here: three states — `SUPERVISED` / `UNSUPERVISED` / `UNKNOWN` — is only half. The half that matters is **`UNKNOWN` must not be able to reach `kill` by writing nothing.** If the enum is added but the call site keeps an `else` or a `default:` arm, the unsafe path is still the one you fall into by default, and every test still passes. **The caveat: instance 3 is a bash script, and bash has no exhaustiveness check.** A `case` with no `*)` arm silently matches nothing and continues. So the compiler cannot carry this one; the substitute is that *every* switch over the value must carry a `*) die …` arm. I measured how that stands today in PR #495 (`dcd5052`): ``` CONTROL, ')' occurrences in the file: 149 (non-zero, the grep works) require_drivable_supervisor :144 arms: launchd|systemd|none, ambiguous, *) die <- has one report state :308 arms: launchd, systemd, none <- NO *) stop :419 arms: launchd, systemd, none <- NO *) start :477 arms: launchd, systemd, none <- NO *) ``` The guard at `:305` runs **once**. The three switches that actually drive the daemon each enumerate exactly three values and silently do nothing for a fourth. So adding a state without also adding the `die` arms produces: nothing matches in the stop switch, the daemon is never stopped, nothing matches in the start switch — and the script reports no error at any point. ## Asked for 1. `detect_supervisor` gains an undrivable state, and `require_drivable_supervisor` refuses it. (Already asked on PR #495.) 2. **Every** switch over that value gets a `*) die …` arm — `:308`, `:419`, `:477` — so a future fourth state fails loudly at each site instead of skipping the stop or the start. 3. Instances 1 and 2 are not code changes here; they are why the rule exists. Whoever revisits #479 should read this. ## Testing note — two different tests, and only one of them is the useful one From the fleet01 lead, and it is the same gap as *asserting the nudge happens* versus *asserting the wait waits*: > The interesting case is "the probe could not tell", which usually means a timeout, a permission denial, or an unreadable entry — all easy to stub and hard to make real. **A test that constructs `UNKNOWN` directly proves the switch handles it; it does not prove the probe ever returns it.** Only the second catches a probe that quietly maps its own failures to `none` before the state is ever consulted. Both tests are needed. The second is the one that would have caught instance 3. Related: #492, PR #495, #479, #415, #494 (the inverse family — a constant pretending to be measured).
Author
Owner

Two more instances, both measured today, and one of them is a new sub-shape

Instance 4 — #500, and it is a cleaner example than the ones above

scripts/probe-member-credentials.sh on bash 3.2: mapfile does not exist, so _FIELDS is never set. Every consumer is ${_FIELDS[n]:-default} or a slice, so set -u never fires. The script reaches its own guard at :172 and refuses with:

refusing to run: the policy fetched from $POLICY_URL contains 0 known names

A claim about the policy, caused by the interpreter. Full measurement is on #500.

What makes it a better example than instance 3: here the wrong answer comes out of a deliberate, well-written guard whose comment explicitly says it exists to stop an empty array from printing a table that looks complete. The guard is right about the danger and wrong about the cause. There is no sloppiness to point at.

Instance 5 — a new sub-shape: a value lost in transit, then defaulted into "there is none"

Found in the fix for this very ticket. b17f37a on worker/492-followup-detect-unclear adds the unclear state correctly, and alongside it a global SUPERVISOR_UNCLEAR_DETAIL naming which supervisor and why. The call site is:

SUPERVISOR_KIND="$(detect_supervisor)"      # :404

$( ) is a subshell. Only stdout crosses back. Measured, with a control that discriminates:

plain call, no subshell   -> detail len = 142, correct text     <- CONTROL
K="$(detect_supervisor)"  -> kind = [unclear], detail len = 0

Then require_drivable_supervisor:247 reads ${SUPERVISOR_UNCLEAR_DETAIL:-<generic text>} and prints the generic message. So the third state arrives intact and its reason is silently replaced by "could not tell", with no supervisor named.

The sub-shape: a ${VAR:-default} on a value that should always be present turns "this was lost" into "there was none". That is this ticket's conflation, one level down — and it defeats the fix for this ticket while every test stays green. The :- is the whole mechanism: under set -u a bare $SUPERVISOR_UNCLEAR_DETAIL would have failed loudly at the first run.

Why the test did not catch it — this ticket's own testing note, demonstrated

The test calls require_drivable_supervisor unclear directly, in the same shell, after setting the global by hand. That is exactly the case this ticket warns about:

A test that constructs UNKNOWN directly proves the switch handles it; it does not prove the probe ever returns it.

Here it is stronger than that: constructing the value also skips the handoff, which is where the defect lives. The test cannot fail.

Asked for, in addition to the three asks above

  1. detect_supervisor must return everything its caller needs on stdout — the only channel that survives $( ). State that as a comment above the function; it is a fact about how the function is called, invisible from inside it.
  2. Drop the :- fallback at :247. A missing detail should be a loud failure, not a generic sentence.
  3. Add a test that runs the real call-site shape (K="$(detect_supervisor)" with stubbed probes) and asserts the parent shell sees a non-empty detail naming the right supervisor.

All six are briefed to the worker on worker/492-followup-detect-unclear. Asks 1 and 2 are already implemented in b17f37a: require_drivable_supervisor now has unclear) and *) arms, so the acting guard is closed. The three case "$SUPERVISOR_KIND" switches still have no final arm — re-measured today at :408 (report), :519 (stop), :577 (start), all three with arms launchd) systemd) none) only.

## Two more instances, both measured today, and one of them is a new sub-shape ### Instance 4 — #500, and it is a cleaner example than the ones above `scripts/probe-member-credentials.sh` on bash 3.2: `mapfile` does not exist, so `_FIELDS` is never set. Every consumer is `${_FIELDS[n]:-default}` or a slice, so `set -u` never fires. The script reaches its own guard at `:172` and refuses with: ``` refusing to run: the policy fetched from $POLICY_URL contains 0 known names ``` A claim about the **policy**, caused by the **interpreter**. Full measurement is on #500. What makes it a better example than instance 3: here the wrong answer comes out of a **deliberate, well-written guard** whose comment explicitly says it exists to stop an empty array from printing a table that looks complete. The guard is right about the danger and wrong about the cause. There is no sloppiness to point at. ### Instance 5 — a new sub-shape: a value lost in transit, then defaulted into "there is none" Found in the fix for this very ticket. `b17f37a` on `worker/492-followup-detect-unclear` adds the `unclear` state correctly, and alongside it a global `SUPERVISOR_UNCLEAR_DETAIL` naming which supervisor and why. The call site is: ```bash SUPERVISOR_KIND="$(detect_supervisor)" # :404 ``` `$( )` is a subshell. Only stdout crosses back. Measured, with a control that discriminates: ``` plain call, no subshell -> detail len = 142, correct text <- CONTROL K="$(detect_supervisor)" -> kind = [unclear], detail len = 0 ``` Then `require_drivable_supervisor:247` reads `${SUPERVISOR_UNCLEAR_DETAIL:-<generic text>}` and prints the generic message. So the third state arrives intact and its *reason* is silently replaced by "could not tell", with no supervisor named. **The sub-shape: a `${VAR:-default}` on a value that should always be present turns "this was lost" into "there was none".** That is this ticket's conflation, one level down — and it defeats the fix for this ticket while every test stays green. The `:-` is the whole mechanism: under `set -u` a bare `$SUPERVISOR_UNCLEAR_DETAIL` would have failed loudly at the first run. ### Why the test did not catch it — this ticket's own testing note, demonstrated The test calls `require_drivable_supervisor unclear` directly, in the same shell, after setting the global by hand. That is exactly the case this ticket warns about: > A test that constructs `UNKNOWN` directly proves the switch handles it; it does not prove the probe ever returns it. Here it is stronger than that: constructing the value also skips the **handoff**, which is where the defect lives. The test cannot fail. ### Asked for, in addition to the three asks above 4. `detect_supervisor` must return everything its caller needs on **stdout** — the only channel that survives `$( )`. State that as a comment above the function; it is a fact about how the function is *called*, invisible from inside it. 5. Drop the `:-` fallback at `:247`. A missing detail should be a loud failure, not a generic sentence. 6. Add a test that runs the **real call-site shape** (`K="$(detect_supervisor)"` with stubbed probes) and asserts the parent shell sees a non-empty detail naming the right supervisor. All six are briefed to the worker on `worker/492-followup-detect-unclear`. Asks 1 and 2 are already implemented in `b17f37a`: `require_drivable_supervisor` now has `unclear)` and `*)` arms, so the acting guard is closed. The three `case "$SUPERVISOR_KIND"` switches still have no final arm — re-measured today at `:408` (report), `:519` (stop), `:577` (start), all three with arms `launchd) systemd) none)` only.
Author
Owner

A wider sweep for this family, recovered late

The #492-followup worker dispatched an Explore agent to look for the same shape across the repo and
put the list in its reply. That reply was lost — the worker's turn ended without a fleet_reply
and the pane scrape came back empty. I drained its inbox before tearing the pane down and got the
full text back, so the list is recorded here rather than lost.

Read the confidence marker on each line. Two I checked in the tree myself; four I did not.

Checked myself, at 136312f

  • herdr/PaneLocator.java:129-135 — real, and the highest-stakes member of this family found so
    far. paneOwnsAnyOf catches HerdrException and returns false, so "this pane does not own the
    pid" and "I could not tell" are one answer. Followed out through ConnectionIdentity.resolve to
    CallerResolver:247, that promotes a worker's connection to Principal.primary. Filed
    separately as #505
    , with the chain, why fleetd #317's resolved() gate does not cover it, and
    the acceptance criteria. Work that one from #505, not from this list.

  • mcp/LsofPeerPidLookup.java:53-56 — not a defect; do not change it. The worker listed it
    because both branches return the same -1 sentinel, which is true. But the comment at :45-50
    says so deliberately, and the direction is safe: fleetd #317 made an unresolved caller refuse
    (anonymous), never promote. Recording the negative so nobody re-opens it. LsofProcessCwdLookup.java:43-46
    is the same helper shape one layer over and needs the same reading before anyone touches it.

Not checked by me — the worker's reading, quoted as theirs

Each of these may be real, may be harmless, or may already be deliberate the way LsofPeerPidLookup
turned out to be. Nobody should act on one without reading it first.

  • scripts/rename-checkout.sh:308 — [ -n "$(git -C "$OLD_PATH" status --porcelain 2>/dev/null)" ]
    gates a clean-tree check before a destructive move. If git status itself errors, the empty
    output reads as "clean" and the move proceeds unverified. Note this one also carries the
    2>/dev/null that hides the error it depends on.
  • scripts/rename-checkout.sh:95 — launchd_loaded() treats any non-zero exit as "not loaded",
    which picks the stop path (kill vs. a safe unload -w) at :324. This is the same reasoning
    redeploy-fleetd.sh just gained its unclear state for, in a second script that did not get it.
  • session/GitWorktrees.java:391-401 — remoteUrlsLeakUserInfo, a credential-leak detector,
    catches any RuntimeException from git remote get-url and returns false, i.e. "no leak
    found". A detector that reports "clean" when it failed to look is this family pointed at a
    security check.

What this list is worth

It is a lead, not a finding. The one item I did check split two ways — one real escalation and one
false positive whose code already explained itself. That ratio is the reason none of the four above
is being filed as a defect on somebody's word, mine included.

## A wider sweep for this family, recovered late The #492-followup worker dispatched an Explore agent to look for the same shape across the repo and put the list in its reply. That reply was lost — the worker's turn ended without a `fleet_reply` and the pane scrape came back empty. I drained its inbox before tearing the pane down and got the full text back, so the list is recorded here rather than lost. **Read the confidence marker on each line.** Two I checked in the tree myself; four I did not. ### Checked myself, at `136312f` - **`herdr/PaneLocator.java:129-135`** — real, and the highest-stakes member of this family found so far. `paneOwnsAnyOf` catches `HerdrException` and returns `false`, so "this pane does not own the pid" and "I could not tell" are one answer. Followed out through `ConnectionIdentity.resolve` to `CallerResolver:247`, that promotes a worker's connection to `Principal.primary`. **Filed separately as #505**, with the chain, why fleetd #317's `resolved()` gate does not cover it, and the acceptance criteria. Work that one from #505, not from this list. - **`mcp/LsofPeerPidLookup.java:53-56`** — **not a defect; do not change it.** The worker listed it because both branches return the same `-1` sentinel, which is true. But the comment at `:45-50` says so deliberately, and the direction is safe: fleetd #317 made an unresolved caller refuse (`anonymous`), never promote. Recording the negative so nobody re-opens it. `LsofProcessCwdLookup.java:43-46` is the same helper shape one layer over and needs the same reading before anyone touches it. ### Not checked by me — the worker's reading, quoted as theirs Each of these may be real, may be harmless, or may already be deliberate the way `LsofPeerPidLookup` turned out to be. Nobody should act on one without reading it first. - `scripts/rename-checkout.sh:308` — `[ -n "$(git -C "$OLD_PATH" status --porcelain 2>/dev/null)" ]` gates a clean-tree check before a destructive move. If `git status` itself errors, the empty output reads as "clean" and the move proceeds unverified. Note this one also carries the `2>/dev/null` that hides the error it depends on. - `scripts/rename-checkout.sh:95` — `launchd_loaded()` treats any non-zero exit as "not loaded", which picks the stop path (`kill` vs. a safe `unload -w`) at `:324`. This is the same reasoning `redeploy-fleetd.sh` just gained its `unclear` state for, in a second script that did not get it. - `session/GitWorktrees.java:391-401` — `remoteUrlsLeakUserInfo`, a credential-leak detector, catches any `RuntimeException` from `git remote get-url` and returns `false`, i.e. "no leak found". A detector that reports "clean" when it failed to look is this family pointed at a security check. ### What this list is worth It is a lead, not a finding. The one item I did check split two ways — one real escalation and one false positive whose code already explained itself. That ratio is the reason none of the four above is being filed as a defect on somebody's word, mine included.
Sign in to join this conversation.
1 Participants
Notifications
Due Date
No due date set.
Dependencies

No dependencies set.

Reference: fleet/fleetd#497