Features: correct the config-reload split; record the central secret store

Two fixes to entries that had drifted from the code.

Config reload: the entry claimed a profile's `model` and `tabLabel` were hot. They
are not. HerdrPeerLauncher takes Map.copyOf(profiles) at construction and resolves
every spawn out of that copy, so a reloaded launch setting never reaches a launch.
Only `fleet:`, `placement:` and a profile's weight/maxLoad are read live.

Worktree-hostile config isolation: the reason changed shape. It used to be a crash
(a dangling {file:.secrets/...} reference OpenCode refuses to start on). Credentials
now live in one shell-level store and opencode.json reads them as {env:...}, so in a
worktree that reference resolves instead of failing, and a member would inherit the
primary's admin-scoped token. A loud crash became a quiet privilege leak.

Worker opens its own PR: say where the value comes from. bridged reads gitTokenEnv
from its own process environment, so a daemon started before the export existed
injects an empty token and every worker push fails.
Dai Ha
2026-08-14 20:37:28 +02:00
parent 8b1eb68803
commit 7c50cce52e
+30 -8
@@ -268,10 +268,16 @@ with a valid format-specific stub, then marks a tracked replacement `--skip-work
**On.** Automatic at worktree provisioning; no configuration.
**Why.** The tracked `opencode.json` can reference gitignored `.secrets/` files that a worktree never
contains, and OpenCode refuses to start with that dangling reference. The same isolation rule also
prevents a primary-only MCP configuration or an autoenv authorization prompt from entering a worker's
tool surface.
**Why.** The tracked `opencode.json` mounts the primary's own `gitea` and `context7` servers, with the
primary's credentials, so a worktree that keeps it hands a member access it must never hold. The same
isolation rule also keeps a primary-only MCP configuration or an autoenv authorization prompt out of a
worker's tool surface.
The reason used to be milder, and the change is worth recording. `opencode.json` referenced gitignored
`.secrets/` files that a worktree never contains, and OpenCode refuses to start on that dangling
reference — a crash, but a loud one. Those credentials now live in one shell-level store and the file
reads them as `{env:…}`, so in a worktree the reference resolves instead of failing. A loud crash
became a quiet privilege leak, which makes this isolation load-bearing rather than a workaround.
**Gotcha.** The replacement is deliberately valid, not deleted: `{}` for `opencode.json`, an empty
server map for `.mcp.json`, and an empty `.autoenv`. A deletion could be undone by a later checkout; the
@@ -291,6 +297,13 @@ works.
PR but must not be able to merge — the primary is the gate, and a worker that can merge is not
gated.
**Where the value comes from.** `gitTokenEnv:` names a variable, and `bridged` reads it from **its own
process environment** — so the token must be exported in the shell that launches the daemon, not
stored in the repo. Keep it beside the operator's other credentials in one sourced file; a per-repo
copy is a second copy of the same secret, and the copy you forget is the one that leaks or goes
stale. If the daemon was started before that export existed, it injects an empty token and every
worker push fails: restart it from a login shell.
## Session lifecycle caps
**What.** Reaps idle sessions, caps turns per session, optionally clears a reused Claude Code
@@ -699,8 +712,8 @@ it owns. Changing one pool's `weight` cost the whole fleet's state, so in practi
| Class | Keys | What a reload does |
|---|---|---|
| **Hot** | `fleet:` (pools + `tabLabel`), `placement:`, an existing profile's `weight` / `maxLoad` / `model` / `tabLabel` | takes effect on the next spawn |
| **Deferred** | `lifecycle:`, `leadHeartbeat:`, `guard:`, `worktreeRoot:`, `spawnReadyTimeoutMs` / `spawnReadyPollMs`, **adding or removing** a profile | accepted into the new config, but the startup wiring keeps the old value; logged by name |
| **Hot** | `fleet:` (pools + `tabLabel`), `placement:`, an existing profile's `weight` / `maxLoad` | takes effect on the next spawn |
| **Deferred** | `lifecycle:`, `leadHeartbeat:`, `guard:`, `worktreeRoot:`, `spawnReadyTimeoutMs` / `spawnReadyPollMs`, **adding or removing** a profile, **and an existing profile's launch settings** (`model`, `baseUrl`, `argv`, `env`, `configDir`, `mcpUrl`, `tabLabel`) | accepted into the new config, but the startup wiring keeps the old value; logged by name |
| **Cold** | `bind:`, `herdrSocket:`, `broker:`, `auth:` | **refuses the whole reload** |
A changed cold key refuses everything, not just itself. Applying the hot half and warning about the
@@ -708,8 +721,17 @@ cold half would leave the daemon in a state matching no file on disk — the wor
operator who is reading the file to work out what the daemon is doing. Refusing keeps one invariant:
the live config is always *some* version of the file.
Adding a profile is deferred, not hot, because a new backend needs its own launcher and launchers are
built once at startup. Changing an existing profile's fields is hot, because those are read per spawn.
What decides the class is **who reads the key and when**, not how important it is. `weight` and
`maxLoad` are hot because the placement policy reads them live through a supplier. A profile's
`model` looks like it should behave the same way and does not: `HerdrPeerLauncher` takes
`Map.copyOf(profiles)` at construction and resolves every spawn out of that copy, so a reloaded
`model` never reaches a launch. Adding a profile is deferred for the same underlying reason — a new
backend needs its own launcher, and launchers are built once.
That distinction is worth stating plainly because getting it wrong is invisible. A reload that
reported a changed `model` as applied would log a clean "config reloaded" while every spawn kept
using the old one, and an operator would have no reason to doubt it. So the reload compares each
surviving profile's launch settings and names the profile in the deferred list instead.
A reload that fails to parse, or fails any of the four startup validators, is refused the same way
and the running config stays live. A config file being saved is sometimes read mid-write, and