Features: opencode member config under memberHerdrSocket, and refusing an unbindable role (#219 #123)

Dai Ha
2026-09-01 15:44:45 +07:00
parent ce35e4887c
commit 8e61adffc1
+65
@@ -78,6 +78,8 @@ six weeks, and the table alone will not carry it.
| [Every claude-code member is resumable](#every-claude-code-member-is-resumable) | automatic | #214 | `member/ClaudeCodeLauncher` |
| [An exhausted backend is quarantined even from a chrome-only pane](#an-exhausted-backend-is-quarantined-even-from-a-chrome-only-pane) | automatic | #211 | `inject/CompletionResolver` |
| [The credential scrub follows the member's own user](#the-credential-scrub-follows-the-members-own-user) | `memberLoginShell:` + `worktreeGroup:` | #213 | `member/HerdrPeerLauncher` |
| [An opencode member's config follows the member's own user](#an-opencode-members-config-follows-the-members-own-user) | `memberHerdrSocket:` + `worktreeGroup:` | #219 | `member/OpenCodeLauncher` |
| [A role fleetd cannot bind is refused](#a-role-fleetd-cannot-bind-is-refused-not-quietly-downgraded) | automatic | #123 | `auth/MemberRegistry` |
Nearly every knob above lives in one file, on one profile:
@@ -2777,6 +2779,69 @@ permissions are deliberately tight in the other direction: the directory is `rwx
`rw-r-----` — group read, **no group write anywhere, no world bits**. A member sources what it needs
and cannot alter fleetd's own scrub.
---
## An opencode member's config follows the member's own user
**What.** When `memberHerdrSocket:` puts member panes under a different OS user, the ephemeral
`opencode.json` no longer goes into fleetd's own temp directory. It goes under `worktreeRoot`, shared
read-only with `worktreeGroup`. Session discovery, which cannot work across users, now says so once
instead of failing quietly.
**On.** `memberHerdrSocket:` plus `worktreeRoot`/`worktreeGroup`. With `memberHerdrSocket:` absent,
both paths are byte-identical to before.
**Why it exists.** The launcher decided two paths from its own process. The config directory came
from `java.io.tmpdir`, which is mode 0700 on macOS, so another uid cannot even traverse it. That file
is the only way a member learns where the bridge MCP is. A member that cannot read it still starts
and still holds a pane, but never becomes deliverable: `fleet_send` waits about 60 seconds on the
readiness gate and then fails, and nothing in that message points at a temp directory.
Session discovery had the same shape with a different result. It read fleetd's own `user.home` to
find `opencode.db`, but under this key that file lives under the **member's** home. It would find
nothing, forever, and `agentSessionId` would stay `null` — which reads as "this backend cannot
resume" rather than "fleetd looked in the wrong home".
**The gotcha: this one refuses to spawn, where the credential scrub above degrades.** The difference
is what each thing is. A weakened credential control still has value, so the scrub falls back to the
sentinel overlay. This config file is not a control — it is the member's only route to the bridge. So
a missing `worktreeRoot` or `worktreeGroup` refuses the spawn and names the missing key. Nothing is
lost by refusing: that member would not have worked either way, and the later failure carries no
clue. Discovery is switched off under this key on purpose, with one WARN — an honest "not available"
beats an answer read from the wrong directory.
---
## A role fleetd cannot bind is refused, not quietly downgraded
**What.** A spawn asking for `role: architect` on a profile that no configured slot carries now fails
at once. The message names the role, the profile, and the profiles `fleet.architects` does carry. The
roster also reports the role a member really holds, never the role that was asked for.
**On.** Automatic.
**Why it exists.** Such a spawn used to succeed. The member started, read the architect charter and
the architect agent definition, and then `fleet_whoami` told it — correctly — that it is a worker,
because no slot could be bound. Meanwhile `GET /members` still said `architect`. Three sources of
truth disagreed about one live member, and the only log line was INFO.
That is the quiet direction of failure, which is the worse one. A member acting **above** its role is
refused by the authorization gate, so the mistake announces itself. A member acting **below** its
role simply does not do the job, and the lead reading the roster has no way to see why. An explicit
operator request became the silent case.
Refusing was chosen over binding anyway for one reason: an architect's identity **is** the slot it
was bound to. With no slot, there is nothing to bind an identity to, so binding would mean inventing
one. The operator's fix is one line of config, and the refusal names it.
**The gotcha: only the config gap is closed, not the race.** If a slot exists but is already bound
when the bind runs, the member has already launched on the architect charter and is then recorded as
`dev`. The roster stays honest and the demotion is logged at WARN, so this no longer hides — but the
charter and the bound role still disagree for that member. A pre-check cannot close that window;
#226 carries the reserve-then-bind fix. The refusal is also scoped to `architect` alone: dev and
reviewer pools are placement candidates, not identity bindings, so an explicit profile outside them
stays a supported override.
### A note for anyone briefing a worker to read this page
**A worker cannot see the current version of this file.** `wiki/` is a submodule, and the parent