#168: revise chapter 13 — the second herdr daemon, and the ticket TTL

- Renamed the product to fleet / fleetd. The invariant on line 63 was right; only
  the name was wrong.
- §2 described one herdr socket. It now covers the optional `memberHerdrSocket:`
  key and says why it exists: herdr forks every pane as its own OS user, so
  running members as a different user needs a second herdr owned by that user.
  It is read once at startup (FleetConfig.java:34-37,84; Fleetd.java:148-158).
- §4 now says what /healthz reports with two daemons: the body gains a `member`
  key, and `protocolMismatch: true` when the protocol numbers differ
  (FleetApp.java:256-264). The `herdr` key keeps the LEAD's values on purpose,
  because redeploy-fleetd.sh and rename-checkout.sh read it. An operator who
  reads the wrong key gets the wrong answer.
- Trap 5 was stale. The ticket TTL is now counted from when the turn FINISHES,
  not from when it was sent (#197). Under the old behaviour any task longer than
  the TTL lost its report on arrival. The trap now states both, and says to check
  the running jar, since a daemon started before that fix still has the old rule.

All twelve traps in §6 are kept. Both diagrams rendered with mmdc.
Dai Ha
2026-08-31 10:38:42 +07:00
parent 426f5a2ff4
commit 7962e69bcf
+67 -8
@@ -60,7 +60,7 @@ flowchart LR
1. The lead **never** sets `ANTHROPIC_BASE_URL` or `ANTHROPIC_AUTH_TOKEN`. It stays on the
subscription. Only the daemon moves a member off it, at spawn.
2. The bridge is the **only** channel. Text printed in a pane reaches nobody. An answer that is not
2. The **fleet** is the **only** channel. Text printed in a pane reaches nobody. An answer that is not
in a `fleet_*` call is discarded silently.
---
@@ -112,7 +112,7 @@ keeping one process so launchd's PID tracking still works. Read its header; it e
better than this paragraph.
> On this host the daemon is **not** under launchd. It runs as a plain `java -jar` started by the
> redeploy script from a login shell. `launchctl list | grep bridg` returns nothing.
> redeploy script from a login shell. `launchctl list | grep fleetd` returns nothing.
The daemon logs which required secret names resolved at startup
(`Fleetd.reportRequiredSecrets`). Read those lines. But note the gap: it **skips profiles marked
@@ -138,6 +138,56 @@ The old `terminal:` key is gone. A daemon at CB-579 or later **refuses to start*
still in the config, and also if `tab:` is missing. That refusal is deliberate — a silently demoted
lead was worse than a daemon that will not boot.
### 2.5 A second herdr daemon for members (optional)
§2.1 described one herdr. There is an optional top-level key, `memberHerdrSocket:`, that adds a
second one (`FleetConfig.java:84`, `fleetd/fleetd.example.yaml:117-118`):
```yaml
memberHerdrSocket: /run/fleet/herdr-members.sock # optional — omit for one daemon
```
Set it to a second herdr API socket path and every **member** spawns on that daemon, while
everything the **lead** does stays on the first one. Absent — the default — both names are the
same client, so nothing changes. The daemon builds the two clients from the config on boot, and a
`HerdrRouter` owns the split: each consumer gets the lead client, the member client, or routing by
target id (`Fleetd.java:152-158`, `herdr/HerdrRouter.java:6-24`).
**This key tells you what the feature will do, not how to turn it on today.** The feature is not
ready to switch on yet. The running blocker list is issue #185, and [11 Features](11-Features) →
`memberHerdrSocket` explains why the split exists in the first place: herdr forks every pane as its
own OS user, and has no user parameter in its socket API, so a member under a different user needs
its own herdr. Two blockers are worth knowing up front:
- **The socket permissions are a race.** herdr creates its API socket with mode `0600`, and
`umask` does not change that, so cross-user access still needs a `chmod g+rw` after the socket
appears — a race on every start. That is a herdr-side measurement recorded in #185, not
something the fleetd code can show.
- **Pane ids are still not daemon-qualified.** herdr's ids are per-daemon counters, so two daemons
can both hold `w1:p1`, pointing at different panes. `fleet_stop{paneId}` takes that bare id, and
when two daemons really do claim one, the stop is refused rather than closing a pane on an
arbitrary daemon (`CompositePeerLauncher.java:49-60`) — a safety net, not a design that removes
the collision.
**It is a boot-time-only key.** The daemon reads `memberHerdrSocket:` once, at startup, and builds
the member herdr client from it; nothing re-reads the key after that (`Fleetd.java:153-155`). A
change therefore needs a daemon restart, not a config reload. In the reload machinery it sits in
the set that cannot change under a running daemon, the same set as `herdrSocket:`; with
`configReload:` enabled, a reload that touches it is refused on that account
(`ConfigRef.java:72-74, 190-195`).
**`/healthz` checks both daemons when the second is configured.** It pings both, and returns 200
only when both answer; a down member daemon now shows `503 degraded` instead of hiding behind a
healthy lead, which is the old shape of trap 2 (`FleetApp.java:229-255` and the comment at
210-216). The 200 body gains a `member` key with that daemon's version and protocol, and
`protocolMismatch: true` when the two protocol numbers differ (`FleetApp.java:256-264`). The
`herdr` key still carries the **lead** daemon's values, deliberately: `scripts/redeploy-fleetd.sh`
and `scripts/rename-checkout.sh` already read this endpoint, and folding two daemons into one key
would hide a mismatch from whichever reader looks only there (`FleetApp.java:218-227`). With one
daemon the body is byte-identical to before — one ping, no new keys. The point of the change:
it is the member daemon's protocol that decides whether a spawn works, and before, a bad protocol
on the member daemon left `/healthz` green with every spawn failing.
---
## 3. Configure
@@ -286,7 +336,7 @@ loopback only.
`After=herdr.service`) and `dev.ltms.fleetd.plist` for macOS launchd. `deploy/lavinmq` holds the
optional broker.
- **On this host neither is loaded.** The daemon runs as a plain `java -jar` started by
`scripts/redeploy-fleetd.sh` from a login shell. Verified with `launchctl list | grep bridg`,
`scripts/redeploy-fleetd.sh` from a login shell. Verified with `launchctl list | grep fleetd`,
which returns nothing. If you expected launchd here, that expectation is the bug.
---
@@ -400,9 +450,18 @@ cannot verify identity from that session — ask for `/mcp`.
### Losing a member's work
**5. The ticket expired.**
Ticket time-to-live is about 10 minutes. After that `fleet_poll{ticket}` returns
`timed_out_working`. The member is usually fine and its real answer arrives later — in the member's
**inbox**. Drain it with `fleet_poll{target}`, then `fleet_ack{target, msgId}`.
A `wait:false` ticket is kept for about 10 minutes, counted from when the turn **finishes** — not
from when you sent it (`MessageService.pruneTerminalTickets`). So a task that runs for an hour still
hands you its report, as long as you collect within about 10 minutes of it finishing.
This used to be measured from **creation**, which meant any task longer than the TTL lost its report
the moment it arrived. That was fixed in `#197`; a daemon started before that fix still has the old
behaviour, so check what the running jar is before you trust a long ticket.
When a ticket has expired, `fleet_poll{ticket}` reports it as timed out. The member is usually fine
and its answer went to the member's **inbox** instead. Drain it with `fleet_poll{target}`, then
`fleet_ack{target, msgId}`. For anything long, still have the member write its report into a file or
its pull request, so a lost ticket is never a lost report.
**6. Reading the reply inbox is destructive.**
`GET /sessions/{id}/replies` drains on first read. If you `curl` it through `head` or a parser that
@@ -466,8 +525,8 @@ before believing it**.
| Rendezvous, `fleet_ask`, detached delivery, turn-done fallback | `docs/MCP-Contract.md` **§6 only** — the rest of that page is a pre-build design doc and its tool names never caught up with the code |
| Every config key, documented | `fleetd/fleetd.example.yaml` |
**The one thing that is easy to forget.** This repo *is* the bridge, so the charter block in
**The one thing that is easy to forget.** This repo *is* the fleet, so the charter block in
`CLAUDE.md` is not documentation about someone else's system — it is the instruction surface this
codebase ships. A code change that quietly makes it untrue is an incomplete change. The block in
`CLAUDE.md` and the template in [7 Use Cases](7-Use-Cases) must stay byte-identical; there is a
check script in `CLAUDE.md` that proves it.
check script in `CLAUDE.md` that proves it.