diff --git a/13-User-Guide.md b/13-User-Guide.md index be5a160..c36832c 100644 --- a/13-User-Guide.md +++ b/13-User-Guide.md @@ -60,7 +60,7 @@ flowchart LR 1. The lead **never** sets `ANTHROPIC_BASE_URL` or `ANTHROPIC_AUTH_TOKEN`. It stays on the subscription. Only the daemon moves a member off it, at spawn. -2. The bridge is the **only** channel. Text printed in a pane reaches nobody. An answer that is not +2. The **fleet** is the **only** channel. Text printed in a pane reaches nobody. An answer that is not in a `fleet_*` call is discarded silently. --- @@ -112,7 +112,7 @@ keeping one process so launchd's PID tracking still works. Read its header; it e better than this paragraph. > On this host the daemon is **not** under launchd. It runs as a plain `java -jar` started by the -> redeploy script from a login shell. `launchctl list | grep bridg` returns nothing. +> redeploy script from a login shell. `launchctl list | grep fleetd` returns nothing. The daemon logs which required secret names resolved at startup (`Fleetd.reportRequiredSecrets`). Read those lines. But note the gap: it **skips profiles marked @@ -138,6 +138,56 @@ The old `terminal:` key is gone. A daemon at CB-579 or later **refuses to start* still in the config, and also if `tab:` is missing. That refusal is deliberate — a silently demoted lead was worse than a daemon that will not boot. +### 2.5 A second herdr daemon for members (optional) + +§2.1 described one herdr. There is an optional top-level key, `memberHerdrSocket:`, that adds a +second one (`FleetConfig.java:84`, `fleetd/fleetd.example.yaml:117-118`): + +```yaml +memberHerdrSocket: /run/fleet/herdr-members.sock # optional — omit for one daemon +``` + +Set it to a second herdr API socket path and every **member** spawns on that daemon, while +everything the **lead** does stays on the first one. Absent — the default — both names are the +same client, so nothing changes. The daemon builds the two clients from the config on boot, and a +`HerdrRouter` owns the split: each consumer gets the lead client, the member client, or routing by +target id (`Fleetd.java:152-158`, `herdr/HerdrRouter.java:6-24`). + +**This key tells you what the feature will do, not how to turn it on today.** The feature is not +ready to switch on yet. The running blocker list is issue #185, and [11 Features](11-Features) → +`memberHerdrSocket` explains why the split exists in the first place: herdr forks every pane as its +own OS user, and has no user parameter in its socket API, so a member under a different user needs +its own herdr. Two blockers are worth knowing up front: + +- **The socket permissions are a race.** herdr creates its API socket with mode `0600`, and + `umask` does not change that, so cross-user access still needs a `chmod g+rw` after the socket + appears — a race on every start. That is a herdr-side measurement recorded in #185, not + something the fleetd code can show. +- **Pane ids are still not daemon-qualified.** herdr's ids are per-daemon counters, so two daemons + can both hold `w1:p1`, pointing at different panes. `fleet_stop{paneId}` takes that bare id, and + when two daemons really do claim one, the stop is refused rather than closing a pane on an + arbitrary daemon (`CompositePeerLauncher.java:49-60`) — a safety net, not a design that removes + the collision. + +**It is a boot-time-only key.** The daemon reads `memberHerdrSocket:` once, at startup, and builds +the member herdr client from it; nothing re-reads the key after that (`Fleetd.java:153-155`). A +change therefore needs a daemon restart, not a config reload. In the reload machinery it sits in +the set that cannot change under a running daemon, the same set as `herdrSocket:`; with +`configReload:` enabled, a reload that touches it is refused on that account +(`ConfigRef.java:72-74, 190-195`). + +**`/healthz` checks both daemons when the second is configured.** It pings both, and returns 200 +only when both answer; a down member daemon now shows `503 degraded` instead of hiding behind a +healthy lead, which is the old shape of trap 2 (`FleetApp.java:229-255` and the comment at +210-216). The 200 body gains a `member` key with that daemon's version and protocol, and +`protocolMismatch: true` when the two protocol numbers differ (`FleetApp.java:256-264`). The +`herdr` key still carries the **lead** daemon's values, deliberately: `scripts/redeploy-fleetd.sh` +and `scripts/rename-checkout.sh` already read this endpoint, and folding two daemons into one key +would hide a mismatch from whichever reader looks only there (`FleetApp.java:218-227`). With one +daemon the body is byte-identical to before — one ping, no new keys. The point of the change: +it is the member daemon's protocol that decides whether a spawn works, and before, a bad protocol +on the member daemon left `/healthz` green with every spawn failing. + --- ## 3. Configure @@ -286,7 +336,7 @@ loopback only. `After=herdr.service`) and `dev.ltms.fleetd.plist` for macOS launchd. `deploy/lavinmq` holds the optional broker. - **On this host neither is loaded.** The daemon runs as a plain `java -jar` started by - `scripts/redeploy-fleetd.sh` from a login shell. Verified with `launchctl list | grep bridg`, + `scripts/redeploy-fleetd.sh` from a login shell. Verified with `launchctl list | grep fleetd`, which returns nothing. If you expected launchd here, that expectation is the bug. --- @@ -400,9 +450,18 @@ cannot verify identity from that session — ask for `/mcp`. ### Losing a member's work **5. The ticket expired.** -Ticket time-to-live is about 10 minutes. After that `fleet_poll{ticket}` returns -`timed_out_working`. The member is usually fine and its real answer arrives later — in the member's -**inbox**. Drain it with `fleet_poll{target}`, then `fleet_ack{target, msgId}`. +A `wait:false` ticket is kept for about 10 minutes, counted from when the turn **finishes** — not +from when you sent it (`MessageService.pruneTerminalTickets`). So a task that runs for an hour still +hands you its report, as long as you collect within about 10 minutes of it finishing. + +This used to be measured from **creation**, which meant any task longer than the TTL lost its report +the moment it arrived. That was fixed in `#197`; a daemon started before that fix still has the old +behaviour, so check what the running jar is before you trust a long ticket. + +When a ticket has expired, `fleet_poll{ticket}` reports it as timed out. The member is usually fine +and its answer went to the member's **inbox** instead. Drain it with `fleet_poll{target}`, then +`fleet_ack{target, msgId}`. For anything long, still have the member write its report into a file or +its pull request, so a lost ticket is never a lost report. **6. Reading the reply inbox is destructive.** `GET /sessions/{id}/replies` drains on first read. If you `curl` it through `head` or a parser that @@ -466,8 +525,8 @@ before believing it**. | Rendezvous, `fleet_ask`, detached delivery, turn-done fallback | `docs/MCP-Contract.md` **§6 only** — the rest of that page is a pre-build design doc and its tool names never caught up with the code | | Every config key, documented | `fleetd/fleetd.example.yaml` | -**The one thing that is easy to forget.** This repo *is* the bridge, so the charter block in +**The one thing that is easy to forget.** This repo *is* the fleet, so the charter block in `CLAUDE.md` is not documentation about someone else's system — it is the instruction surface this codebase ships. A code change that quietly makes it untrue is an incomplete change. The block in `CLAUDE.md` and the template in [7 Use Cases](7-Use-Cases) must stay byte-identical; there is a -check script in `CLAUDE.md` that proves it. +check script in `CLAUDE.md` that proves it. \ No newline at end of file