Features: memberHerdrSocket — members on a second herdr daemon (#185/#186/#187/#188)

Dai Ha
2026-08-29 06:26:06 +07:00
parent 226794a4fb
commit 68e32c6839
+43
@@ -2171,6 +2171,49 @@ under the two names above in `${SHARED_ENV}/tools/secrets.sh`. Also: the skill f
`sed -E 's#://[^@]*@#://<redacted>@#g'`, and the `g` is not optional — without it a line holding two
URIs leaks the second one, which `scripts/redeploy-fleetd.sh --check` prints.
## `memberHerdrSocket` — run members on a second herdr daemon
**What.** A new top-level config key. Set it to a second herdr API socket path and every **member**
is spawned on that daemon, while everything the **lead** does stays on the first one. Absent (the
default), both are the same client and nothing changes. Inside the daemon a `HerdrRouter` owns the
split: it hands each consumer the lead client, the member client, or — for consumers that see both
kinds of traffic — routes per target id.
```yaml
memberHerdrSocket: /run/fleet/members.sock # omit for one daemon (the default)
```
**On.** Add the key and restart. It is read once at startup, so it is a deferred key, not a hot one.
**Why.** herdr forks every pane as **its own** OS user, and there is no user, uid or run-as parameter
anywhere in its socket API — measured against herdr 0.8.0's own docs, not just our client. So the
only way to make a member run as a different user than the operator is a second herdr server started
by that user. That matters because a member currently reads the operator's ssh key straight off the
filesystem: blocking `SSH_AUTH_SOCK` does not stop it, and no environment scrub can, because the
scrub removes variables and the key is a file (#184). A separate OS user is the control; this key is
the fleetd half of it. Design and measurements are in #185.
**Gotcha — the feature is not ready to switch on yet.** Four things must land first, and #185 keeps
the running list. The two that bite hardest: herdr creates its API socket mode `0600` and `umask`
does not change it, so cross-user access still needs a post-start `chmod g+rw` (a race on every
start); and pane ownership lives only in memory, so after a restart every surviving pane is unowned
and `fleet_stop` on it refuses rather than guess which daemon owns it.
**The shape to watch for.** herdr's workspace, tab and pane ids are **per-daemon sequential
counters**. Two daemons really do both hold `w1:p1`, pointing at different panes owned by different
users — measured, not theorised. Four separate defects came from code that assumed one daemon, and
every one of them passed the whole test suite, because with the key unset both clients are the same
object. When reviewing anything on this path, ask each herdr call one question: *which daemon does
this go to, and is that the daemon that owns the thing it is asking about?*
**Observable change even with one daemon.** `GET /healthz` now requires **both** daemons to answer
before it returns 200 (with one configured that is the single ping it always was). `GET /sessions`
merges `workspace.list` across both. The 200 body still reports only the lead daemon's herdr version
and protocol — a known gap, on the #185 list, and it matters because in two-daemon mode it is the
*member* daemon's protocol that decides whether spawns work.
---
---
## Backfill status