Features: late-resolved member ids, /members body key, and the honest gap report
Also corrects the opencode.db entry: it claimed fleet_list reports agentSessionId, which was only true after the second fix (#209). Reading the database was necessary but not sufficient - SessionManager froze the id at spawn, before opencode writes it.
+89
@@ -69,6 +69,9 @@ six weeks, and the table alone will not carry it.
|
||||
| [Run members on a second herdr daemon](#memberherdrsocket--run-members-on-a-second-herdr-daemon) | `memberHerdrSocket:` | CB-185 | `herdr/HerdrRouter` |
|
||||
| [Let a member under another OS user write its worktree](#worktreegroup--let-a-member-under-another-os-user-write-its-worktree) | `worktreeGroup:` | CB-185 | `session/GitWorktrees` |
|
||||
| [Resume an opencode member's prior session](#opencode-session-ids-come-from-opencodedb) | automatic | CB-206 | `member/OpenCodeSessionDiscovery` |
|
||||
| [Fill in a member id the backend names late](#late-resolved-member-ids) | automatic | CB-209 | `session/SessionManager` |
|
||||
| [REST roster rows under `members`](#get-members-reports-its-rows-under-members) | automatic | CB-199 | `rest/FleetApp` |
|
||||
| [Say "unknown" when the member env is unreadable](#the-credential-gap-detector-admits-when-it-cannot-see) | `memberHerdrSocket:` | CB-185 | `member/HerdrPeerLauncher` |
|
||||
|
||||
Nearly every knob above lives in one file, on one profile:
|
||||
|
||||
@@ -2393,6 +2396,13 @@ It is not redundant work to optimise away.
|
||||
`fleet_spawn{resumeSessionId}` can therefore relaunch one onto its prior conversation. Automatic —
|
||||
no knob.
|
||||
|
||||
**It took two fixes, not one, and the first one alone did nothing visible.** Reading `opencode.db`
|
||||
(below) was necessary but not sufficient: `SessionManager` asked the handle for the id **once, at
|
||||
spawn**, and froze it into an immutable record. opencode has not written its row at that moment, so
|
||||
the frozen answer was always `null` and the roster never showed one. The handle's own javadoc said
|
||||
*"the caller re-calls later"* — no caller did, and `SessionManager` did not even keep the handle.
|
||||
See [Late-resolved member ids](#late-resolved-member-ids) for the second half.
|
||||
|
||||
**On.** Nothing to turn on. It needs opencode's own storage root
|
||||
(`~/.local/share/opencode`) to be readable, which it is.
|
||||
|
||||
@@ -2422,6 +2432,85 @@ undeclared external-binary requirement on a headless host.
|
||||
---
|
||||
|
||||
|
||||
## Late-resolved member ids
|
||||
|
||||
**What.** A member whose backend cannot name its own session at spawn time gets its
|
||||
`agentSessionId` filled in later, as soon as the backend writes it. `fleet_list`, `fleet_status`
|
||||
and the REST roster all report the current answer rather than the one frozen at spawn. Automatic —
|
||||
no knob.
|
||||
|
||||
**On.** Nothing to turn on. It only does work for a member whose id is still unknown.
|
||||
|
||||
**Why it exists.** `agentSessionId` is the handle `fleet_spawn{resumeSessionId}` needs. It was read
|
||||
once at spawn and never again, which is fine for claude-code (it knows its session id up front) and
|
||||
useless for opencode (it does not). The result was a documented, tested feature that could never
|
||||
produce a value for most of the fleet. This is the second half of the
|
||||
[`opencode.db`](#opencode-session-ids-come-from-opencodedb) fix.
|
||||
|
||||
**The gotcha: the resolving read is not the cheap one.** Asking an opencode handle for its id opens
|
||||
a SQLite database — around 841MB here. `roster()` is the roster supplier for the heartbeat loop and
|
||||
the health monitor, so resolving there would open that database on every tick, for every unresolved
|
||||
member, **forever** for a member whose row never appears. So there are two reads:
|
||||
|
||||
| Method | Resolves? | Used by |
|
||||
|---|---|---|
|
||||
| `roster()` | no | `LeadHeartbeatLoop`, `FleetHealthMonitor`, placement and exhaustion checks, the `/metrics` scrape |
|
||||
| `rosterResolved()` | yes | `fleet_list`, the REST roster — the surfaces that actually report the id |
|
||||
|
||||
`get(paneId)` (behind `fleet_status`) and the teardown path resolve too; both are caller-driven, not
|
||||
timers. A test asserts the plain `roster()` never calls the handle again, so the split cannot be
|
||||
quietly "tidied up" later. Once an id is found it is stored and never looked up again.
|
||||
|
||||
---
|
||||
|
||||
|
||||
## `GET /members` reports its rows under `members`
|
||||
|
||||
**What.** The REST roster's response body uses the key `members`, matching its path. The old key
|
||||
`workers` is still emitted as a deprecated alias carrying the same rows.
|
||||
|
||||
**On.** Nothing to turn on.
|
||||
|
||||
**Why it exists.** The route was renamed `/workers` → `/members`, and the body key was not renamed
|
||||
with it. A caller reading `body["members"]` — the obvious guess given the path — got an **empty
|
||||
list**, which is a perfectly valid answer meaning "this fleet has no members". So a client reported
|
||||
an idle fleet and nothing anywhere contradicted it.
|
||||
|
||||
**The gotcha: do not just rename the key.** The REST surface is the out-of-band path a lead falls
|
||||
back to when its MCP mount drops, and at least one client reads `workers` today. Renaming outright
|
||||
would break that client the same silent way, in the other direction. Both keys are emitted, and a
|
||||
test asserts they carry identical rows so the alias cannot drift. Drop `workers` once nothing reads
|
||||
it.
|
||||
|
||||
---
|
||||
|
||||
|
||||
## The credential gap detector admits when it cannot see
|
||||
|
||||
**What.** With `memberHerdrSocket:` set, the `memberCredentials` gap report stops drawing
|
||||
conclusions about member panes and says the gap is **unknown, not clean**, naming the config key
|
||||
that made it unknowable. One WARN per launcher, not per spawn.
|
||||
|
||||
**On.** Only in two-herdr mode. With `memberHerdrSocket:` absent — the default — every log line on
|
||||
this path is byte-identical to before, pinned by a test.
|
||||
|
||||
**Why it exists.** The detector enumerates **fleetd's own** environment, on the assumption that the
|
||||
member pane's login shell exports the same set. Under `memberHerdrSocket:` the pane belongs to a
|
||||
different OS user, with a different `$HOME` and a different secret store, so that assumption is
|
||||
simply false. The old output would then report on the wrong process — and could call a gap clean
|
||||
while the member user exported something dangerous. There is no channel to read another user's
|
||||
environment, so the honest answer is the only correct one.
|
||||
|
||||
**The gotcha: this fixes the report, not the control.** Two related defects in the same area are
|
||||
open under **#213** — the ZDOTDIR scrub decides whether it can run by reading *fleetd's* `$SHELL`,
|
||||
and writes its generated file into fleetd's own `TMPDIR`, which on macOS is a per-user `0700`
|
||||
directory another uid cannot even traverse. In two-user mode the scrub can therefore be silently
|
||||
absent while the daemon believes it ran. Treat `memberCredentials` as unverified in that mode until
|
||||
#213 lands.
|
||||
|
||||
---
|
||||
|
||||
|
||||
## Backfill status
|
||||
|
||||
This page was started after the fact, so it is **not yet complete**. Entries above are written from
|
||||
|
||||
Reference in New Issue
Block a user