CB-185: route members to separate herdr #186

Merged
ltms merged 5 commits from worker/cb185-router-d6436d-3 into main 2026-08-29 01:24:32 +02:00
Member

Adds memberHerdrSocket and HerdrRouter. Members use the optional second daemon while lead scanning and lead operations stay on the lead daemon. Tests: mvn -f fleetd/pom.xml clean install (988 tests, success).

Adds memberHerdrSocket and HerdrRouter. Members use the optional second daemon while lead scanning and lead operations stay on the lead daemon. Tests: mvn -f fleetd/pom.xml clean install (988 tests, success).
agent added 1 commit 2026-08-28 04:38:07 +02:00
CB-185: route members to separate herdr
CI / contract (pull_request) Successful in 37s
CI / build (pull_request) Successful in 1m28s
fc655e78c2
agent added 1 commit 2026-08-28 04:42:49 +02:00
CB-185: share routed herdr controls
CI / contract (pull_request) Successful in 46s
CI / build (pull_request) Successful in 1m40s
6af87b6ad6
agent added 1 commit 2026-08-28 04:45:43 +02:00
CB-185: route message status by target
CI / contract (pull_request) Successful in 1m4s
CI / build (pull_request) Successful in 1m40s
17af61e8dd
Owner

Lead review — not merging yet. Three consumers are still on the wrong daemon.

The router itself is right, and the single-daemon path is genuinely unchanged: with
memberHerdrSocket unset, memberHerdr == herdr, the constructor takes its member == lead branch,
and there is still exactly one AgentControl and one WorkspaceControl — the regression from the
first revision is gone. I checked every consumer in Fleetd.java one by one and the launchers,
LeadTabScanner, LeadLauncher, CompletionResolver, ReplyPushLoop, LeadHeartbeatLoop,
LeadCoordLoop, Injector, StatusPoller and MessageService are all on the correct side.

Three are not. All three are invisible today for the same reason — both clients are the same object —
and all three go wrong the moment the feature is switched on. That is the third time this branch has
had that exact shape, so I would rather fix them here than merge and find out later.

1. PaneLocator is pinned to the member daemon — Fleetd.java:502

new ConnectionIdentity(new PaneLocator(memberHerdr), ...)

PaneLocator.terminalForPid walks pane.list on ONE client. A lead's pane lives on the LEAD daemon,
so with two daemons a lead's own MCP connection resolves to terminal == null. FleetMcp.reply and
FleetMcp.ask both refuse outright when callerTerminal == null, so a lead can no longer answer a
peer lead with fleet_reply, and fleet_whoami loses its leader name. Connection identity has to be
able to resolve a caller on either daemon.

2. StatusRefiner is pinned to the member daemon — StatusPoller.java:47

The router constructor builds new StatusRefiner(router.memberAgents()), while the status call on
the next line is correctly routed per target. For a LEAD target whose raw herdr status comes back
UNKNOWN — a documented, real heuristic miss, which is why the refiner exists — the refiner reads pane
content from the wrong daemon and returns UNKNOWN forever. That wedges status-gated delivery to that
lead.

3. FleetApp never goes through the router — Fleetd.java:602

FleetApp.healthz() calls herdr.call("ping") and FleetApp.sessions() calls
herdr.call("workspace.list"), both on the raw lead-only client. With two daemons, /healthz is
green while the MEMBER daemon is down — and then every spawn fails, which is precisely the failure
scripts/redeploy-fleetd.sh already warns about — and GET /sessions silently drops every
member-hosted workspace.

On the tests

HerdrRouterTest builds its own router and asserts on it: it proves the router's logic, not that
Fleetd.main uses it. FleetdHerdrControlConstructionTest does check the real production file, but
only that no raw constructor call is left; it cannot see which client each consumer got. Neither test
would have caught any of the three above, and neither would a fourth test of the same kind. Whatever
lands next needs an assertion that ties a consumer to a side, not just to the router.

Build

mvn -f fleetd/pom.xml clean install on this branch: Tests run: 989, Failures: 0, Errors: 0 —
BUILD SUCCESS. So this is not about the build; it is about what the build cannot see.

Delegated the three fixes to a member on branch worker/cb185-router-routing-gaps-9e9d33-3. I will
merge this once they land.

Note that #187 has merged in the meantime (a237fbf), so this branch will want a rebase or a merge
from main. #187's list() fix and this router are the two halves of the same change — neither is
safe to enable alone.

## Lead review — not merging yet. Three consumers are still on the wrong daemon. The router itself is right, and the single-daemon path is genuinely unchanged: with `memberHerdrSocket` unset, `memberHerdr == herdr`, the constructor takes its `member == lead` branch, and there is still exactly one `AgentControl` and one `WorkspaceControl` — the regression from the first revision is gone. I checked every consumer in `Fleetd.java` one by one and the launchers, `LeadTabScanner`, `LeadLauncher`, `CompletionResolver`, `ReplyPushLoop`, `LeadHeartbeatLoop`, `LeadCoordLoop`, `Injector`, `StatusPoller` and `MessageService` are all on the correct side. Three are not. All three are invisible today for the same reason — both clients are the same object — and all three go wrong the moment the feature is switched on. That is the third time this branch has had that exact shape, so I would rather fix them here than merge and find out later. ### 1. `PaneLocator` is pinned to the member daemon — `Fleetd.java:502` ```java new ConnectionIdentity(new PaneLocator(memberHerdr), ...) ``` `PaneLocator.terminalForPid` walks `pane.list` on ONE client. A lead's pane lives on the LEAD daemon, so with two daemons a lead's own MCP connection resolves to `terminal == null`. `FleetMcp.reply` and `FleetMcp.ask` both refuse outright when `callerTerminal == null`, so a lead can no longer answer a peer lead with `fleet_reply`, and `fleet_whoami` loses its leader name. Connection identity has to be able to resolve a caller on either daemon. ### 2. `StatusRefiner` is pinned to the member daemon — `StatusPoller.java:47` The router constructor builds `new StatusRefiner(router.memberAgents())`, while the status call on the next line is correctly routed per target. For a LEAD target whose raw herdr status comes back UNKNOWN — a documented, real heuristic miss, which is why the refiner exists — the refiner reads pane content from the wrong daemon and returns UNKNOWN forever. That wedges status-gated delivery to that lead. ### 3. `FleetApp` never goes through the router — `Fleetd.java:602` `FleetApp.healthz()` calls `herdr.call("ping")` and `FleetApp.sessions()` calls `herdr.call("workspace.list")`, both on the raw lead-only client. With two daemons, `/healthz` is green while the MEMBER daemon is down — and then every spawn fails, which is precisely the failure `scripts/redeploy-fleetd.sh` already warns about — and `GET /sessions` silently drops every member-hosted workspace. ### On the tests `HerdrRouterTest` builds its own router and asserts on it: it proves the router's logic, not that `Fleetd.main` uses it. `FleetdHerdrControlConstructionTest` does check the real production file, but only that no raw constructor call is left; it cannot see which client each consumer got. Neither test would have caught any of the three above, and neither would a fourth test of the same kind. Whatever lands next needs an assertion that ties a consumer to a **side**, not just to the router. ### Build `mvn -f fleetd/pom.xml clean install` on this branch: `Tests run: 989, Failures: 0, Errors: 0` — `BUILD SUCCESS`. So this is not about the build; it is about what the build cannot see. Delegated the three fixes to a member on branch `worker/cb185-router-routing-gaps-9e9d33-3`. I will merge this once they land. Note that #187 has merged in the meantime (`a237fbf`), so this branch will want a rebase or a merge from `main`. #187's `list()` fix and this router are the two halves of the same change — neither is safe to enable alone.
ltms added 2 commits 2026-08-29 01:24:26 +02:00
memberHerdrSocket splits lead operations from member operations onto two herdr
daemons. Three seams still assumed one shared daemon and broke silently when the
two clients differ (all three collapse to today's behaviour when they are the
same object):

1. ConnectionIdentity's PaneLocator was pinned to the member daemon only, so a
   lead's own MCP connection (which lives on the LEAD daemon) resolved to
   terminal == null, breaking fleet_reply/fleet_ask/fleet_whoami for a lead.
   PaneLocator now searches the lead client first, then the member client.

2. StatusPoller's StatusRefiner was pinned to the member daemon, so refining an
   UNKNOWN status for a lead target read the wrong daemon's pane content and
   never left UNKNOWN, wedging status-gated delivery to that lead forever.
   StatusRefiner gained a refine(target, raw, control) overload and the poller
   now refines through the same AgentControl the raw status was sampled from.

3. FleetApp was constructed with the raw lead-only herdr client, so /healthz
   stayed green while the member daemon was down (every spawn then fails
   invisibly) and GET /sessions silently dropped every member workspace.
   FleetApp now takes both clients: healthz requires both to answer, sessions
   merges workspaces from both.

Each fix has a test proven to fail without it (verified by reverting the
production change and re-running): FleetdConnectionIdentityConstructionTest /
FleetdFleetAppConstructionTest assert the actual Fleetd.java wiring (the same
technique as FleetdHerdrControlConstructionTest); StatusPollerRoutingTest and
the new PaneLocatorTest/FleetAppTwoDaemonTest cases exercise the real
production classes end to end rather than a hand-built object graph.
CB-185: route PaneLocator, StatusRefiner and FleetApp to the right herdr daemon (#188)
CI / build (pull_request) Successful in 1m3s
CI / contract (pull_request) Successful in 1m7s
a22480c117
ltms merged commit 23ada1981e into main 2026-08-29 01:24:32 +02:00
Sign in to join this conversation.