Features: loop health in fleet_list//healthz (#562), and TIMED_OUT_UNCONFIRMED (#571)

7-Use-Cases: re-sync the portable CLAUDE.md block after #562 added the
loopHealth row to the intent->tool table.

11-Features: two entries. The loop-health one records the failure
DIRECTION that matters -- a false negative where nothing ever fires is
worse than a muted false positive -- and why the wiring test exists.
The TIMED_OUT_UNCONFIRMED one records the default -> "done" that told
REST callers a delegation had completed, and the order dependence in
using a compile error to find enum readers.
Dai Ha
2026-09-12 20:29:22 +07:00
parent 20d2fc0eb2
commit dca74a658b
2 changed files with 76 additions and 1 deletions
+75
@@ -5534,3 +5534,78 @@ re-litigating:
have this shape. In production the bound is real; in the suite it is inert.
fleetd #480 (PRs #483, #484, #485). Related: #486, #489 (PR #490).
## A timed-out send now says whether delivery was even attempted
Before this, a `fleet_send` that timed out reported one of two things: `TIMED_OUT_WORKING` if the
message was delivered, or `TIMED_OUT_QUEUED` if it was not. There was no third answer, so a case
that is neither got filed under "not delivered".
That case is real. When the injector cancels a delivery it reports `DELIVERED`, `NOT_DELIVERED`, or
**`ATTEMPTED`** — meaning the keystrokes were already going out and nobody can say whether they
landed. The send path collapsed `ATTEMPTED` into `TIMED_OUT_QUEUED`, which **promises the message
never arrived**. A lead reading that promise retries, and the worker gets the same brief twice.
**What it does.** `MessageService.Outcome` gains `TIMED_OUT_UNCONFIRMED`. The `ATTEMPTED`
cancellation now routes to it instead of to `TIMED_OUT_QUEUED`. Every reader handles it:
- **MCP** — `fleet_send` returns `[no reply within Nms — delivery unconfirmed; the message may
already have reached the worker, so a retry risks sending it twice — poll status before
resending]`. It deliberately does **not** carry the "retry or poll status" wording the
queued/working arm uses, because on this route a resend can double-deliver.
- **REST** — `POST /sessions/{id}/messages` answers 202 with `"status":"unconfirmed"`.
- **Metrics** — the `fleet_sends` counter still labels it `timeout`, grouped with the other two
timeout outcomes. That grouping is deliberate and was left alone.
**The knob.** None. It is a behaviour change on an existing path, live as soon as the daemon
restarts.
**Why it exists.** A sentinel that means "no" and a sentinel that means "cannot tell" need opposite
handling from the caller, and this code had only the confident one. Fixing it needs a third state,
not a better guess — the same shape as fleetd #512.
**The gotcha, and it is the interesting one.** `FleetApp.writeReply`'s inner switch carried
`default -> "done"`, so any outcome it did not name told a REST caller **the delegation completed**.
Adding a constant would have been absorbed silently by that default. The fix deletes it and lists
all ten outcomes by name, so the compiler now catches the next missed one. **Order matters if you
repeat this exercise**: a `default` is exactly what suppresses the compile error you are trying to
provoke, so delete the defaults *first*, then add the new constant, or the proof comes back clean
and proves nothing.
**Still open.** The outcome's wire token is its constant name lowercased, not a pinned string — the
sibling enum `ReplyOutcome` does pin its own. That is fleetd #578, which has since grown: the same
enum emits four different tokens across four live surfaces, so the fix is not one accessor.
fleetd #571 (PR #580). Related: #512, #578, #586.
## `fleet_list` and `/healthz` report whether the background loops are alive
Two loops keep the fleet honest: `StatusPoller`, which refreshes member status, and
`SessionReaper`, which retires idle members. Both have had a `LoopWatchdog` for a while. Nothing
outside the daemon could see it.
**What it does.** `fleet_list` gains a `loopHealth` object with keys `statusPoller` and
`sessionReaper`, each `RUNNING`, `STALLED`, or `STOPPED`. `/healthz` carries the same object in its
body. The 200/503 status codes are unchanged — a stalled loop does **not** turn the endpoint red.
**The knob.** None. Always on.
**Why it exists.** A stalled poller does not announce itself. The fleet keeps answering, member
status quietly goes stale, and the first symptom is a lead acting on state that stopped updating
hours ago. Exposing the watchdog turns a silent failure into a visible one.
**The gotcha — the failure direction.** This is monitoring, so ask what happens when the monitor
lies. The dangerous direction here is a false **negative**: report `RUNNING` while the loop is dead,
and the watchdog can never fire. That is worse than a false positive, because a false positive is
noisy and somebody mutes it, whereas a false negative produces no signal for anyone to notice is
missing. The wiring is what protects against it, and the wiring is now pinned:
`FleetdLoopHealthSourceWiringTest` fails if `Fleetd.loopHealthSource` stops asking the real poller
or the real reaper. That test exists because the first version of this feature shipped with five
green tests and the production wiring could still be replaced by a constant with nothing failing —
every one of the five built its own `LoopHealthSource`, which tests the consumer and can never be
evidence about the producer.
**Reading it.** `STOPPED` for `sessionReaper` is normal when no reaper is configured; that is the
null case, not a fault.
fleetd #562 (PRs #579, #584).
+1 -1
@@ -449,7 +449,7 @@ the merge — and merging on a reviewer's word is delegating it by proxy.
| Confirm your own role | `fleet_whoami` |
| See backends available | `fleet_profiles` |
| Start a member | `fleet_spawn{role?, profile?, cwd?, worktree?, ticket?, sessionName?, resumeSessionId?}` → `sessionId` + `paneId` |
| See the fleet | `fleet_list` → `leads` (your peers) + `members` (each carries `agentSessionId` when its backend knows one) · one peer's state: `fleet_status{sessionId}` |
| See the fleet | `fleet_list` → `leads` (your peers) + `members` (each carries `agentSessionId` when its backend knows one) + `loopHealth` (`RUNNING`, `STALLED`, or `STOPPED` for `statusPoller` and `sessionReaper`) · one peer's state: `fleet_status{sessionId}` |
| Delegate (blocking) | `fleet_send{sessionId, content}` |
| Delegate (long task) | `fleet_send{sessionId, content, wait:false}` → ticket → `fleet_poll{ticket}` |
| Answer a member's `fleet_ask` | `fleet_send{turnId, content}` — **not** `sessionId` |