From ff48e2927e86fd5ed095979658043050bb308971 Mon Sep 17 00:00:00 2001 From: Dai Ha Date: Mon, 31 Aug 2026 22:24:27 +0700 Subject: [PATCH] Features: late-resolved member ids, /members body key, and the honest gap report Also corrects the opencode.db entry: it claimed fleet_list reports agentSessionId, which was only true after the second fix (#209). Reading the database was necessary but not sufficient - SessionManager froze the id at spawn, before opencode writes it. --- 11-Features.md | 89 ++++++++++++++++++++++++++++++++++++++++++++++++++ 1 file changed, 89 insertions(+) diff --git a/11-Features.md b/11-Features.md index 72d7ec8..33eab6c 100644 --- a/11-Features.md +++ b/11-Features.md @@ -69,6 +69,9 @@ six weeks, and the table alone will not carry it. | [Run members on a second herdr daemon](#memberherdrsocket--run-members-on-a-second-herdr-daemon) | `memberHerdrSocket:` | CB-185 | `herdr/HerdrRouter` | | [Let a member under another OS user write its worktree](#worktreegroup--let-a-member-under-another-os-user-write-its-worktree) | `worktreeGroup:` | CB-185 | `session/GitWorktrees` | | [Resume an opencode member's prior session](#opencode-session-ids-come-from-opencodedb) | automatic | CB-206 | `member/OpenCodeSessionDiscovery` | +| [Fill in a member id the backend names late](#late-resolved-member-ids) | automatic | CB-209 | `session/SessionManager` | +| [REST roster rows under `members`](#get-members-reports-its-rows-under-members) | automatic | CB-199 | `rest/FleetApp` | +| [Say "unknown" when the member env is unreadable](#the-credential-gap-detector-admits-when-it-cannot-see) | `memberHerdrSocket:` | CB-185 | `member/HerdrPeerLauncher` | Nearly every knob above lives in one file, on one profile: @@ -2393,6 +2396,13 @@ It is not redundant work to optimise away. `fleet_spawn{resumeSessionId}` can therefore relaunch one onto its prior conversation. Automatic — no knob. +**It took two fixes, not one, and the first one alone did nothing visible.** Reading `opencode.db` +(below) was necessary but not sufficient: `SessionManager` asked the handle for the id **once, at +spawn**, and froze it into an immutable record. opencode has not written its row at that moment, so +the frozen answer was always `null` and the roster never showed one. The handle's own javadoc said +*"the caller re-calls later"* — no caller did, and `SessionManager` did not even keep the handle. +See [Late-resolved member ids](#late-resolved-member-ids) for the second half. + **On.** Nothing to turn on. It needs opencode's own storage root (`~/.local/share/opencode`) to be readable, which it is. @@ -2422,6 +2432,85 @@ undeclared external-binary requirement on a headless host. --- +## Late-resolved member ids + +**What.** A member whose backend cannot name its own session at spawn time gets its +`agentSessionId` filled in later, as soon as the backend writes it. `fleet_list`, `fleet_status` +and the REST roster all report the current answer rather than the one frozen at spawn. Automatic — +no knob. + +**On.** Nothing to turn on. It only does work for a member whose id is still unknown. + +**Why it exists.** `agentSessionId` is the handle `fleet_spawn{resumeSessionId}` needs. It was read +once at spawn and never again, which is fine for claude-code (it knows its session id up front) and +useless for opencode (it does not). The result was a documented, tested feature that could never +produce a value for most of the fleet. This is the second half of the +[`opencode.db`](#opencode-session-ids-come-from-opencodedb) fix. + +**The gotcha: the resolving read is not the cheap one.** Asking an opencode handle for its id opens +a SQLite database — around 841MB here. `roster()` is the roster supplier for the heartbeat loop and +the health monitor, so resolving there would open that database on every tick, for every unresolved +member, **forever** for a member whose row never appears. So there are two reads: + +| Method | Resolves? | Used by | +|---|---|---| +| `roster()` | no | `LeadHeartbeatLoop`, `FleetHealthMonitor`, placement and exhaustion checks, the `/metrics` scrape | +| `rosterResolved()` | yes | `fleet_list`, the REST roster — the surfaces that actually report the id | + +`get(paneId)` (behind `fleet_status`) and the teardown path resolve too; both are caller-driven, not +timers. A test asserts the plain `roster()` never calls the handle again, so the split cannot be +quietly "tidied up" later. Once an id is found it is stored and never looked up again. + +--- + + +## `GET /members` reports its rows under `members` + +**What.** The REST roster's response body uses the key `members`, matching its path. The old key +`workers` is still emitted as a deprecated alias carrying the same rows. + +**On.** Nothing to turn on. + +**Why it exists.** The route was renamed `/workers` → `/members`, and the body key was not renamed +with it. A caller reading `body["members"]` — the obvious guess given the path — got an **empty +list**, which is a perfectly valid answer meaning "this fleet has no members". So a client reported +an idle fleet and nothing anywhere contradicted it. + +**The gotcha: do not just rename the key.** The REST surface is the out-of-band path a lead falls +back to when its MCP mount drops, and at least one client reads `workers` today. Renaming outright +would break that client the same silent way, in the other direction. Both keys are emitted, and a +test asserts they carry identical rows so the alias cannot drift. Drop `workers` once nothing reads +it. + +--- + + +## The credential gap detector admits when it cannot see + +**What.** With `memberHerdrSocket:` set, the `memberCredentials` gap report stops drawing +conclusions about member panes and says the gap is **unknown, not clean**, naming the config key +that made it unknowable. One WARN per launcher, not per spawn. + +**On.** Only in two-herdr mode. With `memberHerdrSocket:` absent — the default — every log line on +this path is byte-identical to before, pinned by a test. + +**Why it exists.** The detector enumerates **fleetd's own** environment, on the assumption that the +member pane's login shell exports the same set. Under `memberHerdrSocket:` the pane belongs to a +different OS user, with a different `$HOME` and a different secret store, so that assumption is +simply false. The old output would then report on the wrong process — and could call a gap clean +while the member user exported something dangerous. There is no channel to read another user's +environment, so the honest answer is the only correct one. + +**The gotcha: this fixes the report, not the control.** Two related defects in the same area are +open under **#213** — the ZDOTDIR scrub decides whether it can run by reading *fleetd's* `$SHELL`, +and writes its generated file into fleetd's own `TMPDIR`, which on macOS is a per-user `0700` +directory another uid cannot even traverse. In two-user mode the scrub can therefore be silently +absent while the daemon believes it ran. Treat `memberCredentials` as unverified in that mode until +#213 lands. + +--- + + ## Backfill status This page was started after the fact, so it is **not yet complete**. Entries above are written from