#168: revise chapter 13 — the second herdr daemon, and the ticket TTL
- Renamed the product to fleet / fleetd. The invariant on line 63 was right; only the name was wrong. - §2 described one herdr socket. It now covers the optional `memberHerdrSocket:` key and says why it exists: herdr forks every pane as its own OS user, so running members as a different user needs a second herdr owned by that user. It is read once at startup (FleetConfig.java:34-37,84; Fleetd.java:148-158). - §4 now says what /healthz reports with two daemons: the body gains a `member` key, and `protocolMismatch: true` when the protocol numbers differ (FleetApp.java:256-264). The `herdr` key keeps the LEAD's values on purpose, because redeploy-fleetd.sh and rename-checkout.sh read it. An operator who reads the wrong key gets the wrong answer. - Trap 5 was stale. The ticket TTL is now counted from when the turn FINISHES, not from when it was sent (#197). Under the old behaviour any task longer than the TTL lost its report on arrival. The trap now states both, and says to check the running jar, since a daemon started before that fix still has the old rule. All twelve traps in §6 are kept. Both diagrams rendered with mmdc.
+67
-8
@@ -60,7 +60,7 @@ flowchart LR
|
||||
|
||||
1. The lead **never** sets `ANTHROPIC_BASE_URL` or `ANTHROPIC_AUTH_TOKEN`. It stays on the
|
||||
subscription. Only the daemon moves a member off it, at spawn.
|
||||
2. The bridge is the **only** channel. Text printed in a pane reaches nobody. An answer that is not
|
||||
2. The **fleet** is the **only** channel. Text printed in a pane reaches nobody. An answer that is not
|
||||
in a `fleet_*` call is discarded silently.
|
||||
|
||||
---
|
||||
@@ -112,7 +112,7 @@ keeping one process so launchd's PID tracking still works. Read its header; it e
|
||||
better than this paragraph.
|
||||
|
||||
> On this host the daemon is **not** under launchd. It runs as a plain `java -jar` started by the
|
||||
> redeploy script from a login shell. `launchctl list | grep bridg` returns nothing.
|
||||
> redeploy script from a login shell. `launchctl list | grep fleetd` returns nothing.
|
||||
|
||||
The daemon logs which required secret names resolved at startup
|
||||
(`Fleetd.reportRequiredSecrets`). Read those lines. But note the gap: it **skips profiles marked
|
||||
@@ -138,6 +138,56 @@ The old `terminal:` key is gone. A daemon at CB-579 or later **refuses to start*
|
||||
still in the config, and also if `tab:` is missing. That refusal is deliberate — a silently demoted
|
||||
lead was worse than a daemon that will not boot.
|
||||
|
||||
### 2.5 A second herdr daemon for members (optional)
|
||||
|
||||
§2.1 described one herdr. There is an optional top-level key, `memberHerdrSocket:`, that adds a
|
||||
second one (`FleetConfig.java:84`, `fleetd/fleetd.example.yaml:117-118`):
|
||||
|
||||
```yaml
|
||||
memberHerdrSocket: /run/fleet/herdr-members.sock # optional — omit for one daemon
|
||||
```
|
||||
|
||||
Set it to a second herdr API socket path and every **member** spawns on that daemon, while
|
||||
everything the **lead** does stays on the first one. Absent — the default — both names are the
|
||||
same client, so nothing changes. The daemon builds the two clients from the config on boot, and a
|
||||
`HerdrRouter` owns the split: each consumer gets the lead client, the member client, or routing by
|
||||
target id (`Fleetd.java:152-158`, `herdr/HerdrRouter.java:6-24`).
|
||||
|
||||
**This key tells you what the feature will do, not how to turn it on today.** The feature is not
|
||||
ready to switch on yet. The running blocker list is issue #185, and [11 Features](11-Features) →
|
||||
`memberHerdrSocket` explains why the split exists in the first place: herdr forks every pane as its
|
||||
own OS user, and has no user parameter in its socket API, so a member under a different user needs
|
||||
its own herdr. Two blockers are worth knowing up front:
|
||||
|
||||
- **The socket permissions are a race.** herdr creates its API socket with mode `0600`, and
|
||||
`umask` does not change that, so cross-user access still needs a `chmod g+rw` after the socket
|
||||
appears — a race on every start. That is a herdr-side measurement recorded in #185, not
|
||||
something the fleetd code can show.
|
||||
- **Pane ids are still not daemon-qualified.** herdr's ids are per-daemon counters, so two daemons
|
||||
can both hold `w1:p1`, pointing at different panes. `fleet_stop{paneId}` takes that bare id, and
|
||||
when two daemons really do claim one, the stop is refused rather than closing a pane on an
|
||||
arbitrary daemon (`CompositePeerLauncher.java:49-60`) — a safety net, not a design that removes
|
||||
the collision.
|
||||
|
||||
**It is a boot-time-only key.** The daemon reads `memberHerdrSocket:` once, at startup, and builds
|
||||
the member herdr client from it; nothing re-reads the key after that (`Fleetd.java:153-155`). A
|
||||
change therefore needs a daemon restart, not a config reload. In the reload machinery it sits in
|
||||
the set that cannot change under a running daemon, the same set as `herdrSocket:`; with
|
||||
`configReload:` enabled, a reload that touches it is refused on that account
|
||||
(`ConfigRef.java:72-74, 190-195`).
|
||||
|
||||
**`/healthz` checks both daemons when the second is configured.** It pings both, and returns 200
|
||||
only when both answer; a down member daemon now shows `503 degraded` instead of hiding behind a
|
||||
healthy lead, which is the old shape of trap 2 (`FleetApp.java:229-255` and the comment at
|
||||
210-216). The 200 body gains a `member` key with that daemon's version and protocol, and
|
||||
`protocolMismatch: true` when the two protocol numbers differ (`FleetApp.java:256-264`). The
|
||||
`herdr` key still carries the **lead** daemon's values, deliberately: `scripts/redeploy-fleetd.sh`
|
||||
and `scripts/rename-checkout.sh` already read this endpoint, and folding two daemons into one key
|
||||
would hide a mismatch from whichever reader looks only there (`FleetApp.java:218-227`). With one
|
||||
daemon the body is byte-identical to before — one ping, no new keys. The point of the change:
|
||||
it is the member daemon's protocol that decides whether a spawn works, and before, a bad protocol
|
||||
on the member daemon left `/healthz` green with every spawn failing.
|
||||
|
||||
---
|
||||
|
||||
## 3. Configure
|
||||
@@ -286,7 +336,7 @@ loopback only.
|
||||
`After=herdr.service`) and `dev.ltms.fleetd.plist` for macOS launchd. `deploy/lavinmq` holds the
|
||||
optional broker.
|
||||
- **On this host neither is loaded.** The daemon runs as a plain `java -jar` started by
|
||||
`scripts/redeploy-fleetd.sh` from a login shell. Verified with `launchctl list | grep bridg`,
|
||||
`scripts/redeploy-fleetd.sh` from a login shell. Verified with `launchctl list | grep fleetd`,
|
||||
which returns nothing. If you expected launchd here, that expectation is the bug.
|
||||
|
||||
---
|
||||
@@ -400,9 +450,18 @@ cannot verify identity from that session — ask for `/mcp`.
|
||||
### Losing a member's work
|
||||
|
||||
**5. The ticket expired.**
|
||||
Ticket time-to-live is about 10 minutes. After that `fleet_poll{ticket}` returns
|
||||
`timed_out_working`. The member is usually fine and its real answer arrives later — in the member's
|
||||
**inbox**. Drain it with `fleet_poll{target}`, then `fleet_ack{target, msgId}`.
|
||||
A `wait:false` ticket is kept for about 10 minutes, counted from when the turn **finishes** — not
|
||||
from when you sent it (`MessageService.pruneTerminalTickets`). So a task that runs for an hour still
|
||||
hands you its report, as long as you collect within about 10 minutes of it finishing.
|
||||
|
||||
This used to be measured from **creation**, which meant any task longer than the TTL lost its report
|
||||
the moment it arrived. That was fixed in `#197`; a daemon started before that fix still has the old
|
||||
behaviour, so check what the running jar is before you trust a long ticket.
|
||||
|
||||
When a ticket has expired, `fleet_poll{ticket}` reports it as timed out. The member is usually fine
|
||||
and its answer went to the member's **inbox** instead. Drain it with `fleet_poll{target}`, then
|
||||
`fleet_ack{target, msgId}`. For anything long, still have the member write its report into a file or
|
||||
its pull request, so a lost ticket is never a lost report.
|
||||
|
||||
**6. Reading the reply inbox is destructive.**
|
||||
`GET /sessions/{id}/replies` drains on first read. If you `curl` it through `head` or a parser that
|
||||
@@ -466,8 +525,8 @@ before believing it**.
|
||||
| Rendezvous, `fleet_ask`, detached delivery, turn-done fallback | `docs/MCP-Contract.md` **§6 only** — the rest of that page is a pre-build design doc and its tool names never caught up with the code |
|
||||
| Every config key, documented | `fleetd/fleetd.example.yaml` |
|
||||
|
||||
**The one thing that is easy to forget.** This repo *is* the bridge, so the charter block in
|
||||
**The one thing that is easy to forget.** This repo *is* the fleet, so the charter block in
|
||||
`CLAUDE.md` is not documentation about someone else's system — it is the instruction surface this
|
||||
codebase ships. A code change that quietly makes it untrue is an incomplete change. The block in
|
||||
`CLAUDE.md` and the template in [7 Use Cases](7-Use-Cases) must stay byte-identical; there is a
|
||||
check script in `CLAUDE.md` that proves it.
|
||||
check script in `CLAUDE.md` that proves it.
|
||||
Reference in New Issue
Block a user