Features: late-resolved member ids, /members body key, and the honest gap report

Also corrects the opencode.db entry: it claimed fleet_list reports agentSessionId,
which was only true after the second fix (#209). Reading the database was necessary
but not sufficient - SessionManager froze the id at spawn, before opencode writes it.
Dai Ha
2026-08-31 22:24:27 +07:00
parent 34593af1a1
commit ff48e2927e
+89
@@ -69,6 +69,9 @@ six weeks, and the table alone will not carry it.
| [Run members on a second herdr daemon](#memberherdrsocket--run-members-on-a-second-herdr-daemon) | `memberHerdrSocket:` | CB-185 | `herdr/HerdrRouter` |
| [Let a member under another OS user write its worktree](#worktreegroup--let-a-member-under-another-os-user-write-its-worktree) | `worktreeGroup:` | CB-185 | `session/GitWorktrees` |
| [Resume an opencode member's prior session](#opencode-session-ids-come-from-opencodedb) | automatic | CB-206 | `member/OpenCodeSessionDiscovery` |
| [Fill in a member id the backend names late](#late-resolved-member-ids) | automatic | CB-209 | `session/SessionManager` |
| [REST roster rows under `members`](#get-members-reports-its-rows-under-members) | automatic | CB-199 | `rest/FleetApp` |
| [Say "unknown" when the member env is unreadable](#the-credential-gap-detector-admits-when-it-cannot-see) | `memberHerdrSocket:` | CB-185 | `member/HerdrPeerLauncher` |
Nearly every knob above lives in one file, on one profile:
@@ -2393,6 +2396,13 @@ It is not redundant work to optimise away.
`fleet_spawn{resumeSessionId}` can therefore relaunch one onto its prior conversation. Automatic —
no knob.
**It took two fixes, not one, and the first one alone did nothing visible.** Reading `opencode.db`
(below) was necessary but not sufficient: `SessionManager` asked the handle for the id **once, at
spawn**, and froze it into an immutable record. opencode has not written its row at that moment, so
the frozen answer was always `null` and the roster never showed one. The handle's own javadoc said
*"the caller re-calls later"* — no caller did, and `SessionManager` did not even keep the handle.
See [Late-resolved member ids](#late-resolved-member-ids) for the second half.
**On.** Nothing to turn on. It needs opencode's own storage root
(`~/.local/share/opencode`) to be readable, which it is.
@@ -2422,6 +2432,85 @@ undeclared external-binary requirement on a headless host.
---
## Late-resolved member ids
**What.** A member whose backend cannot name its own session at spawn time gets its
`agentSessionId` filled in later, as soon as the backend writes it. `fleet_list`, `fleet_status`
and the REST roster all report the current answer rather than the one frozen at spawn. Automatic —
no knob.
**On.** Nothing to turn on. It only does work for a member whose id is still unknown.
**Why it exists.** `agentSessionId` is the handle `fleet_spawn{resumeSessionId}` needs. It was read
once at spawn and never again, which is fine for claude-code (it knows its session id up front) and
useless for opencode (it does not). The result was a documented, tested feature that could never
produce a value for most of the fleet. This is the second half of the
[`opencode.db`](#opencode-session-ids-come-from-opencodedb) fix.
**The gotcha: the resolving read is not the cheap one.** Asking an opencode handle for its id opens
a SQLite database — around 841MB here. `roster()` is the roster supplier for the heartbeat loop and
the health monitor, so resolving there would open that database on every tick, for every unresolved
member, **forever** for a member whose row never appears. So there are two reads:
| Method | Resolves? | Used by |
|---|---|---|
| `roster()` | no | `LeadHeartbeatLoop`, `FleetHealthMonitor`, placement and exhaustion checks, the `/metrics` scrape |
| `rosterResolved()` | yes | `fleet_list`, the REST roster — the surfaces that actually report the id |
`get(paneId)` (behind `fleet_status`) and the teardown path resolve too; both are caller-driven, not
timers. A test asserts the plain `roster()` never calls the handle again, so the split cannot be
quietly "tidied up" later. Once an id is found it is stored and never looked up again.
---
## `GET /members` reports its rows under `members`
**What.** The REST roster's response body uses the key `members`, matching its path. The old key
`workers` is still emitted as a deprecated alias carrying the same rows.
**On.** Nothing to turn on.
**Why it exists.** The route was renamed `/workers` → `/members`, and the body key was not renamed
with it. A caller reading `body["members"]` — the obvious guess given the path — got an **empty
list**, which is a perfectly valid answer meaning "this fleet has no members". So a client reported
an idle fleet and nothing anywhere contradicted it.
**The gotcha: do not just rename the key.** The REST surface is the out-of-band path a lead falls
back to when its MCP mount drops, and at least one client reads `workers` today. Renaming outright
would break that client the same silent way, in the other direction. Both keys are emitted, and a
test asserts they carry identical rows so the alias cannot drift. Drop `workers` once nothing reads
it.
---
## The credential gap detector admits when it cannot see
**What.** With `memberHerdrSocket:` set, the `memberCredentials` gap report stops drawing
conclusions about member panes and says the gap is **unknown, not clean**, naming the config key
that made it unknowable. One WARN per launcher, not per spawn.
**On.** Only in two-herdr mode. With `memberHerdrSocket:` absent — the default — every log line on
this path is byte-identical to before, pinned by a test.
**Why it exists.** The detector enumerates **fleetd's own** environment, on the assumption that the
member pane's login shell exports the same set. Under `memberHerdrSocket:` the pane belongs to a
different OS user, with a different `$HOME` and a different secret store, so that assumption is
simply false. The old output would then report on the wrong process — and could call a gap clean
while the member user exported something dangerous. There is no channel to read another user's
environment, so the honest answer is the only correct one.
**The gotcha: this fixes the report, not the control.** Two related defects in the same area are
open under **#213** — the ZDOTDIR scrub decides whether it can run by reading *fleetd's* `$SHELL`,
and writes its generated file into fleetd's own `TMPDIR`, which on macOS is a per-user `0700`
directory another uid cannot even traverse. In two-user mode the scrub can therefore be silently
absent while the daemon believes it ran. Treat `memberCredentials` as unverified in that mode until
#213 lands.
---
## Backfill status
This page was started after the fact, so it is **not yet complete**. Entries above are written from