From dca74a658bbda9a9731c5801e83fc456e1379dcf Mon Sep 17 00:00:00 2001 From: Dai Ha Date: Sat, 12 Sep 2026 20:29:22 +0700 Subject: [PATCH] Features: loop health in fleet_list//healthz (#562), and TIMED_OUT_UNCONFIRMED (#571) 7-Use-Cases: re-sync the portable CLAUDE.md block after #562 added the loopHealth row to the intent->tool table. 11-Features: two entries. The loop-health one records the failure DIRECTION that matters -- a false negative where nothing ever fires is worse than a muted false positive -- and why the wiring test exists. The TIMED_OUT_UNCONFIRMED one records the default -> "done" that told REST callers a delegation had completed, and the order dependence in using a compile error to find enum readers. --- 11-Features.md | 75 ++++++++++++++++++++++++++++++++++++++++++++++++++ 7-Use-Cases.md | 2 +- 2 files changed, 76 insertions(+), 1 deletion(-) diff --git a/11-Features.md b/11-Features.md index 9a0fb39..19616eb 100644 --- a/11-Features.md +++ b/11-Features.md @@ -5534,3 +5534,78 @@ re-litigating: have this shape. In production the bound is real; in the suite it is inert. fleetd #480 (PRs #483, #484, #485). Related: #486, #489 (PR #490). + +## A timed-out send now says whether delivery was even attempted + +Before this, a `fleet_send` that timed out reported one of two things: `TIMED_OUT_WORKING` if the +message was delivered, or `TIMED_OUT_QUEUED` if it was not. There was no third answer, so a case +that is neither got filed under "not delivered". + +That case is real. When the injector cancels a delivery it reports `DELIVERED`, `NOT_DELIVERED`, or +**`ATTEMPTED`** — meaning the keystrokes were already going out and nobody can say whether they +landed. The send path collapsed `ATTEMPTED` into `TIMED_OUT_QUEUED`, which **promises the message +never arrived**. A lead reading that promise retries, and the worker gets the same brief twice. + +**What it does.** `MessageService.Outcome` gains `TIMED_OUT_UNCONFIRMED`. The `ATTEMPTED` +cancellation now routes to it instead of to `TIMED_OUT_QUEUED`. Every reader handles it: + +- **MCP** — `fleet_send` returns `[no reply within Nms — delivery unconfirmed; the message may + already have reached the worker, so a retry risks sending it twice — poll status before + resending]`. It deliberately does **not** carry the "retry or poll status" wording the + queued/working arm uses, because on this route a resend can double-deliver. +- **REST** — `POST /sessions/{id}/messages` answers 202 with `"status":"unconfirmed"`. +- **Metrics** — the `fleet_sends` counter still labels it `timeout`, grouped with the other two + timeout outcomes. That grouping is deliberate and was left alone. + +**The knob.** None. It is a behaviour change on an existing path, live as soon as the daemon +restarts. + +**Why it exists.** A sentinel that means "no" and a sentinel that means "cannot tell" need opposite +handling from the caller, and this code had only the confident one. Fixing it needs a third state, +not a better guess — the same shape as fleetd #512. + +**The gotcha, and it is the interesting one.** `FleetApp.writeReply`'s inner switch carried +`default -> "done"`, so any outcome it did not name told a REST caller **the delegation completed**. +Adding a constant would have been absorbed silently by that default. The fix deletes it and lists +all ten outcomes by name, so the compiler now catches the next missed one. **Order matters if you +repeat this exercise**: a `default` is exactly what suppresses the compile error you are trying to +provoke, so delete the defaults *first*, then add the new constant, or the proof comes back clean +and proves nothing. + +**Still open.** The outcome's wire token is its constant name lowercased, not a pinned string — the +sibling enum `ReplyOutcome` does pin its own. That is fleetd #578, which has since grown: the same +enum emits four different tokens across four live surfaces, so the fix is not one accessor. + +fleetd #571 (PR #580). Related: #512, #578, #586. + +## `fleet_list` and `/healthz` report whether the background loops are alive + +Two loops keep the fleet honest: `StatusPoller`, which refreshes member status, and +`SessionReaper`, which retires idle members. Both have had a `LoopWatchdog` for a while. Nothing +outside the daemon could see it. + +**What it does.** `fleet_list` gains a `loopHealth` object with keys `statusPoller` and +`sessionReaper`, each `RUNNING`, `STALLED`, or `STOPPED`. `/healthz` carries the same object in its +body. The 200/503 status codes are unchanged — a stalled loop does **not** turn the endpoint red. + +**The knob.** None. Always on. + +**Why it exists.** A stalled poller does not announce itself. The fleet keeps answering, member +status quietly goes stale, and the first symptom is a lead acting on state that stopped updating +hours ago. Exposing the watchdog turns a silent failure into a visible one. + +**The gotcha — the failure direction.** This is monitoring, so ask what happens when the monitor +lies. The dangerous direction here is a false **negative**: report `RUNNING` while the loop is dead, +and the watchdog can never fire. That is worse than a false positive, because a false positive is +noisy and somebody mutes it, whereas a false negative produces no signal for anyone to notice is +missing. The wiring is what protects against it, and the wiring is now pinned: +`FleetdLoopHealthSourceWiringTest` fails if `Fleetd.loopHealthSource` stops asking the real poller +or the real reaper. That test exists because the first version of this feature shipped with five +green tests and the production wiring could still be replaced by a constant with nothing failing — +every one of the five built its own `LoopHealthSource`, which tests the consumer and can never be +evidence about the producer. + +**Reading it.** `STOPPED` for `sessionReaper` is normal when no reaper is configured; that is the +null case, not a fault. + +fleetd #562 (PRs #579, #584). diff --git a/7-Use-Cases.md b/7-Use-Cases.md index 8ce25c4..f1fbe6f 100644 --- a/7-Use-Cases.md +++ b/7-Use-Cases.md @@ -449,7 +449,7 @@ the merge — and merging on a reviewer's word is delegating it by proxy. | Confirm your own role | `fleet_whoami` | | See backends available | `fleet_profiles` | | Start a member | `fleet_spawn{role?, profile?, cwd?, worktree?, ticket?, sessionName?, resumeSessionId?}` → `sessionId` + `paneId` | -| See the fleet | `fleet_list` → `leads` (your peers) + `members` (each carries `agentSessionId` when its backend knows one) · one peer's state: `fleet_status{sessionId}` | +| See the fleet | `fleet_list` → `leads` (your peers) + `members` (each carries `agentSessionId` when its backend knows one) + `loopHealth` (`RUNNING`, `STALLED`, or `STOPPED` for `statusPoller` and `sessionReaper`) · one peer's state: `fleet_status{sessionId}` | | Delegate (blocking) | `fleet_send{sessionId, content}` | | Delegate (long task) | `fleet_send{sessionId, content, wait:false}` → ticket → `fleet_poll{ticket}` | | Answer a member's `fleet_ask` | `fleet_send{turnId, content}` — **not** `sessionId` |