Investigate: let any named Claude Code session in herdr talk to the others — a new fleetd role, or a separate project? #669

Open
opened 2026-10-03 19:48:03 +02:00 by ltms · 23 comments
Owner

Operator request, 2026-10-03, in their words:

investigate where fleetd can support wider use case, as long as the Claude Code sessions are withing herdr and with named tab, allow them to communicate - is it better in fleetd or another project?

This ticket records the request and what I measured before handing it to an architect. No decision is made here.

What exists today — measured in the main clone at 4b4a868

The useful surprise is that fleetd is already most of the way there, because a lead is not a spawned session. A lead is a human-driven Claude Code session that fleetd finds by its exact tab label: LeadTabScanner is handed a tab → name map from fleet.leaders.<name>.tab and matches labels exactly, case-insensitively. Nothing about a lead is launched by the daemon.

So "N named sessions that can message each other" is close to expressible now: configure N fleet.leaders entries, one per tab, and each gets a sessionId that fleet_list reports and that fleet_send{sessionId, content} can reach.

The blocker is privilege, not addressing. auth/Authz.java:72:

case SPAWN, STOP, DRAIN, HANDOVER -> caller.isPrimary();

Every fleet.leaders entry resolves as PRIMARY, so configuring a session as a "leader" to let it talk also hands it the power to spawn, stop and drain the whole fleet, and to roll a lead. That is not a side effect anyone would want for a session whose only need is a message channel.

The role that already nearly fits

Role (auth/Role.java) declares PRIMARY, WORKER, ARCHITECT and ANONYMOUS. And Authz.java:77:

case SEND -> caller.isPrimary() || caller.isArchitect();

So ARCHITECT is already "may send and reply and read, may not control the fleet" — exactly the capability profile the request asks for. What an architect is not today is attachable: an architect is spawned by fleetd into its own worktree and bound to a slot, rather than being a session a human already has open in a named tab.

Today the identity ladder has only two outcomes for a session sitting in a herdr pane: its tab matches a configured lead label and it is PRIMARY, or it does not and it is WORKER. There is no middle rung for "a named session I trust to talk, but not to run the fleet".

So the real question for the architect

Is the right shape a new non-privileged named-peer role (or an attachable form of ARCHITECT), keyed on tab label the same way leads already are? And then: fleetd or a separate project?

On the second half, name concretely what a separate project would have to rebuild. fleetd already owns all of this, and none of it is incidental:

  • identity resolved from the connection, not from an argument (mcp/ConnectionIdentity, CallerResolver)
  • the capability table (auth/Authz, auth/Role)
  • herdr pane and tab control, and the tab scan
  • the durable inbox, status-gated delivery, and the one-message-per-turn rule
  • the fleet_* MCP surface every session already mounts
  • cross-host lead↔lead routing over AMQP by coordId

My prior is that this belongs in fleetd as a new rung on the existing ladder, because a separate project would have to duplicate identity and authorization — the two things that must not be guessed twice. But that is a prior, not a finding, and the architect should be free to disagree.

Specific things to settle

  1. New role, or make ARCHITECT attachable? An attachable architect reuses the whole capability row; a new role avoids overloading a word that currently means "advisor the lead consults".
  2. What config names these sessions? A sibling of fleet.leaders keyed by tab, or a flag on a leader entry that withholds the primary capabilities.
  3. Which capabilities exactly. SEND yes. READ/METRICS? COORD_READ (currently isPrimary() only, line 94)? Say why for each.
  4. Does this reopen the privilege hole #661 just closed? A new rung keyed on tab label widens what a tab label can grant. State plainly whether a pane-placed member could reach the new role, and what stops it.
  5. Peer-to-peer or still hub-and-spoke? A lead never assigns work to a peer today — the traffic between leads is coordination only. Does a flat named-peer mesh keep that rule, and what enforces it?
  6. What the operator actually gets. One concrete walk-through: two Claude Code sessions the operator opened by hand, in two named tabs, exchanging a message. Name every config line and every call.

Not verified by me

  • Whether CallerResolver could resolve a non-lead named tab at all without changes. I read Authz and the Role enum myself; I did not trace the resolver for this question.
  • Anything about herdr's own server behaviour.
  • Whether the operator wants these sessions visible in fleet_list alongside leads and members, or kept separate. Not asked.
Operator request, 2026-10-03, in their words: > investigate where fleetd can support wider use case, as long as the Claude Code sessions are withing herdr and with named tab, allow them to communicate - is it better in fleetd or another project? This ticket records the request and what I measured before handing it to an architect. **No decision is made here.** ## What exists today — measured in the main clone at `4b4a868` The useful surprise is that fleetd is already most of the way there, because **a lead is not a spawned session.** A lead is a human-driven Claude Code session that fleetd finds by its exact tab label: `LeadTabScanner` is handed a `tab → name` map from `fleet.leaders.<name>.tab` and matches labels exactly, case-insensitively. Nothing about a lead is launched by the daemon. So "N named sessions that can message each other" is close to expressible now: configure N `fleet.leaders` entries, one per tab, and each gets a `sessionId` that `fleet_list` reports and that `fleet_send{sessionId, content}` can reach. **The blocker is privilege, not addressing.** `auth/Authz.java:72`: ```java case SPAWN, STOP, DRAIN, HANDOVER -> caller.isPrimary(); ``` Every `fleet.leaders` entry resolves as `PRIMARY`, so configuring a session as a "leader" to let it *talk* also hands it the power to spawn, stop and drain the whole fleet, and to roll a lead. That is not a side effect anyone would want for a session whose only need is a message channel. ## The role that already nearly fits `Role` (`auth/Role.java`) declares `PRIMARY`, `WORKER`, `ARCHITECT` and `ANONYMOUS`. And `Authz.java:77`: ```java case SEND -> caller.isPrimary() || caller.isArchitect(); ``` So **`ARCHITECT` is already "may send and reply and read, may not control the fleet"** — exactly the capability profile the request asks for. What an architect is *not* today is attachable: an architect is spawned by fleetd into its own worktree and bound to a slot, rather than being a session a human already has open in a named tab. Today the identity ladder has only two outcomes for a session sitting in a herdr pane: its tab matches a configured lead label and it is `PRIMARY`, or it does not and it is `WORKER`. There is no middle rung for "a named session I trust to talk, but not to run the fleet". ## So the real question for the architect **Is the right shape a new non-privileged named-peer role (or an attachable form of `ARCHITECT`), keyed on tab label the same way leads already are?** And then: fleetd or a separate project? On the second half, name concretely what a separate project would have to rebuild. fleetd already owns all of this, and none of it is incidental: - identity resolved from the connection, not from an argument (`mcp/ConnectionIdentity`, `CallerResolver`) - the capability table (`auth/Authz`, `auth/Role`) - herdr pane and tab control, and the tab scan - the durable inbox, status-gated delivery, and the one-message-per-turn rule - the `fleet_*` MCP surface every session already mounts - cross-host lead↔lead routing over AMQP by `coordId` My prior is that this belongs in fleetd as a new rung on the existing ladder, because a separate project would have to duplicate identity and authorization — the two things that must not be guessed twice. But that is a prior, not a finding, and the architect should be free to disagree. ## Specific things to settle 1. **New role, or make `ARCHITECT` attachable?** An attachable architect reuses the whole capability row; a new role avoids overloading a word that currently means "advisor the lead consults". 2. **What config names these sessions?** A sibling of `fleet.leaders` keyed by tab, or a flag on a leader entry that withholds the primary capabilities. 3. **Which capabilities exactly.** `SEND` yes. `READ`/`METRICS`? `COORD_READ` (currently `isPrimary()` only, line 94)? Say why for each. 4. **Does this reopen the privilege hole #661 just closed?** A new rung keyed on tab label widens what a tab label can grant. State plainly whether a pane-placed member could reach the new role, and what stops it. 5. **Peer-to-peer or still hub-and-spoke?** A lead never assigns work to a peer today — the traffic between leads is coordination only. Does a flat named-peer mesh keep that rule, and what enforces it? 6. **What the operator actually gets.** One concrete walk-through: two Claude Code sessions the operator opened by hand, in two named tabs, exchanging a message. Name every config line and every call. ## Not verified by me - Whether `CallerResolver` could resolve a non-lead named tab at all without changes. I read `Authz` and the `Role` enum myself; I did not trace the resolver for this question. - Anything about herdr's own server behaviour. - Whether the operator wants these sessions visible in `fleet_list` alongside leads and members, or kept separate. Not asked.
Author
Owner

Answer: build it in fleetd, as a new COLLABORATOR role

The architect's answer is in full on this ticket's reply. The recommendation: keep it in fleetd, add a new attachable role COLLABORATOR backed by a new fleet.collaborators.<name>.tab registry. Do not make ARCHITECT attachable, and do not add a flag to fleet.leaders.

The reason it belongs here rather than in a separate project is that the hard part is not moving text. It is proving which herdr pane called, mapping a named tab to a role, routing to the right herdr daemon, and applying one authorization table. Those are already one connected boundary — ConnectionIdentity, CallerResolver, HerdrRouter, Authz. A separate service would either duplicate that boundary or trust fleetd as its identity provider, which leaves the security decision here anyway while splitting the feature across two projects.

Two corrections to my own ticket text, both of which I verified

  1. "A lead is not a spawned session" was wrong. I wrote that when I reframed this ticket. fleetd does auto-launch a configured lead: FleetdAssembly.java:274-280 calls new LeadLauncher(...).ensureLeads() when herdr is up and leaders are configured, and FleetConfig.java:1106-1115 states it plainly — "A lead is now also creatable (CB-557). Before, nothing spawned one." Recognition still comes first and only the shortfall is launched, but the absolute form of my claim does not hold.
  2. My "lead or worker" pane ladder was incomplete. A pane with a live architect-slot binding resolves as ARCHITECT before the worker fallback, at CallerResolver.java:208-228.

My surviving claim, that ARCHITECT is already a send-but-not-control role, is correct — Authz.java:68-77.

The part that matters most: this can reopen #661 at a lower privilege level

A naive implementation would recreate the hole we just fixed. If a collaborator scanner copied LeadTabScanner.scan() (which joins every pane to its tab, LeadTabScanner.java:177-235) and ran its lookup before the worker fallback, a pane-placed member landing in a collaborator-labelled tab would be read back as that collaborator and gain local SEND plus peer visibility.

The #661 validator does not prevent this. I checked the merged code: validatePanePlacementAgainstLeadTabs() returns early on fleet.leaders().isEmpty() and only ever inspects lead tabs (FleetConfig.java:2755-2762, now on main at b4b7cf5). It knows nothing about a collaborator namespace.

So three guards are mandatory, not optional:

  • Widen the #661 validator to refuse placement: pane when either a lead tab or a collaborator tab is configured.
  • Widen the member-label collision validator so a fleet or profile tabLabel cannot match the collaborator namespace.
  • In CallerResolver, a terminal still present in SessionManager must resolve as its spawned member role before any tab-attached role is considered. Copying the current lead-first order would repeat the defect exactly.

This is the one-way-gate pattern: #661's guard closes the direction the incident came from. A new tab-attached role arrives from the other direction with a green build.

Capabilities, as recommended

Grant: READ, METRICS, own-pane REPLY, and a target-limited local fleet_send that may address only a configured lead or collaborator.

Deny: SPAWN, STOP, DRAIN, HANDOVER, ASK, COORD_READ, the cross-host coordId route, and the turnId answer form.

One implementation note with teeth: today a coordId send is authorized as the same SEND action as a local send (FleetMcp.java:450-469, mapping at :1098-1108). So adding COLLABORATOR to SEND would silently grant broker-wide cross-host lead messaging. That route needs splitting into its own action first, leaving current primary/architect behaviour unchanged.

What is not settled

Only one architect answered this question, so this is a single considered position, not an agreement between two. I verified its corrections to my ticket and its #661-reopening claim against the code myself. I have not verified its full capability table line by line, and COLLABORATOR, the config block and the widened validators are a proposed design that does not exist yet, so nothing here is tested.

The architect also noted it did not inspect herdr server code, fleetd/fleetd.yaml, or wiki/.

Next step is an implementation plan split into units; this stays open until then.

## Answer: build it in fleetd, as a new `COLLABORATOR` role The architect's answer is in full on this ticket's reply. The recommendation: **keep it in fleetd**, add a new attachable role `COLLABORATOR` backed by a new `fleet.collaborators.<name>.tab` registry. Do not make `ARCHITECT` attachable, and do not add a flag to `fleet.leaders`. The reason it belongs here rather than in a separate project is that the hard part is not moving text. It is proving which herdr pane called, mapping a named tab to a role, routing to the right herdr daemon, and applying one authorization table. Those are already one connected boundary — `ConnectionIdentity`, `CallerResolver`, `HerdrRouter`, `Authz`. A separate service would either duplicate that boundary or trust fleetd as its identity provider, which leaves the security decision here anyway while splitting the feature across two projects. ## Two corrections to my own ticket text, both of which I verified 1. **"A lead is not a spawned session" was wrong.** I wrote that when I reframed this ticket. fleetd does auto-launch a configured lead: `FleetdAssembly.java:274-280` calls `new LeadLauncher(...).ensureLeads()` when herdr is up and leaders are configured, and `FleetConfig.java:1106-1115` states it plainly — *"A lead is now also creatable (CB-557). Before, nothing spawned one."* Recognition still comes first and only the shortfall is launched, but the absolute form of my claim does not hold. 2. **My "lead or worker" pane ladder was incomplete.** A pane with a live architect-slot binding resolves as `ARCHITECT` before the worker fallback, at `CallerResolver.java:208-228`. My surviving claim, that `ARCHITECT` is already a send-but-not-control role, is correct — `Authz.java:68-77`. ## The part that matters most: this can reopen #661 at a lower privilege level A naive implementation **would** recreate the hole we just fixed. If a collaborator scanner copied `LeadTabScanner.scan()` (which joins every pane to its tab, `LeadTabScanner.java:177-235`) and ran its lookup before the worker fallback, a pane-placed member landing in a collaborator-labelled tab would be read back as that collaborator and gain local `SEND` plus peer visibility. **The #661 validator does not prevent this.** I checked the merged code: `validatePanePlacementAgainstLeadTabs()` returns early on `fleet.leaders().isEmpty()` and only ever inspects lead tabs (`FleetConfig.java:2755-2762`, now on `main` at b4b7cf5). It knows nothing about a collaborator namespace. So three guards are mandatory, not optional: - Widen the #661 validator to refuse `placement: pane` when **either** a lead tab **or** a collaborator tab is configured. - Widen the member-label collision validator so a fleet or profile `tabLabel` cannot match the collaborator namespace. - In `CallerResolver`, a terminal still present in `SessionManager` must resolve as its spawned member role **before** any tab-attached role is considered. Copying the current lead-first order would repeat the defect exactly. This is the one-way-gate pattern: #661's guard closes the direction the incident came from. A new tab-attached role arrives from the other direction with a green build. ## Capabilities, as recommended Grant: `READ`, `METRICS`, own-pane `REPLY`, and a **target-limited** local `fleet_send` that may address only a configured lead or collaborator. Deny: `SPAWN`, `STOP`, `DRAIN`, `HANDOVER`, `ASK`, `COORD_READ`, the cross-host `coordId` route, and the `turnId` answer form. One implementation note with teeth: today a `coordId` send is authorized as the same `SEND` action as a local send (`FleetMcp.java:450-469`, mapping at `:1098-1108`). So **adding `COLLABORATOR` to `SEND` would silently grant broker-wide cross-host lead messaging.** That route needs splitting into its own action first, leaving current primary/architect behaviour unchanged. ## What is not settled Only one architect answered this question, so this is a single considered position, not an agreement between two. I verified its corrections to my ticket and its #661-reopening claim against the code myself. I have not verified its full capability table line by line, and `COLLABORATOR`, the config block and the widened validators are a proposed design that does not exist yet, so nothing here is tested. The architect also noted it did not inspect herdr server code, `fleetd/fleetd.yaml`, or `wiki/`. Next step is an implementation plan split into units; this stays open until then.
Author
Owner

Second architect position collected, and the design is now settled. Implementation plan below.

Two architects worked this question independently, on the same brief, unable to see each other. Neither could see the other's answer, so where they agree it is two positions rather than one. I checked every contested claim myself in the main clone at 6f27522 before ruling.

The headline: the recommendation stands, but it is not safe to implement as written. Both architects independently concluded that SEND must be split into separate actions before the role exists. Each found gaps the other missed, and one of them overturns a guard the accepted answer called mandatory.


What both architects concluded, independently

Build it in fleetd. Add COLLABORATOR and fleet.collaborators.<name>.tab. Do not make ARCHITECT attachable. Deny SPAWN, STOP, DRAIN, HANDOVER, COORD_READ and the cross-host coordId route. Grant METRICS and own-pane REPLY. The #661 attack is real, and the resolver must check SessionManager before any tab map. Generalise the existing tab scan instead of adding a second scanner. The live walk-through is the lead's, not a worker's.

Both also found, separately, that splitting SEND must land first. That is the strongest signal in the two reports, because neither could have copied it.

One architect added a mechanical reason not to make ARCHITECT attachable that nobody had given: Principal.isSpawnedMember() is WORKER || ARCHITECT, and FleetMcp.java:437-440 uses it to enrol a caller in the presence map. An attachable architect would be listed as an available member, so a lead's fleet_list would offer the operator's own hand-opened session as a worker to delegate to.


Disagreement 1 — READ. I rule for the architect who objected.

The accepted answer grants READ. One architect agreed and checked that fleet_list hides the coordinator row. The other refused, saying READ exposes task replies and pending questions, not a safe roster.

I checked this myself, and the objection is right. Three facts, each from a command I ran:

fleet_status returns the pending question, the turnId and the ticket id:

// FleetMcp.java
return text(base + "\n\n[question — worker is waiting for your answer]\n" + ask.question()
        + "\n\nAnswer it by calling fleet_send again with turnId=\"" + ask.turnId()
        + "\" and content set to your answer; the worker resumes the same turn."
        + " (ticket " + ask.ticket() + ")");

Polling a ticket is plain READ — no target means no DRAIN:

static Authz.Action pollAction(String target, String coordId) {
    if (!isBlank(coordId)) { return Authz.Action.COORD_READ; }
    return isBlank(target) ? Authz.Action.READ : Authz.Action.DRAIN;
}

And ticket ids are a plain in-memory counter:

fleetd/src/main/java/dev/ltms/fleet/msg/MessageService.java:342   private final AtomicLong ticketSeq = new AtomicLong();
fleetd/src/main/java/dev/ltms/fleet/msg/MessageService.java:1282  String ticket = "task-" + ticketSeq.incrementAndGet();

So the ids are task-1, task-2, … and they reset on every daemon restart. I have first-hand evidence of how guessable that is from this session: grep -n "task-5" fleetd/fleetd.out returns 20 separate occurrences against 20 different sessions.

Put together: a collaborator holding today's READ can walk task-1..N and read the reply of every delegation the lead has run since the last restart, and can call fleet_status on any member to lift its open question and turnId.

The other architect established the same mechanism and drew the opposite conclusion from it — it used "fleet_poll{ticket} maps to READ and returns the reply text" as the reason DRAIN can stay denied without breaking the round trip. That is true and useful, and it is also exactly the leak. The same fact serves both arguments; only one of them noticed it cuts both ways.

Ruling: READ must be split before COLLABORATOR is granted anything. Keep READ for roster, profiles and identity. Move ticket polling and session status to their own actions and deny them to COLLABORATOR in the first release.

Worth recording: the READ arm's own comment in Authz.java says "the roster carries no secrets." That is now false, and it is false for WORKER today, not only for a future collaborator. Filed separately.


Disagreement 2 — ASK. I rule for granting it.

One architect denied ASK, arguing it pairs with the turnId answer form. The other argued to grant. I side with granting, on the mechanism:

case REPLY, ASK -> caller.ownsSession(targetSession);

That arm is ownership-based, not role-based, so COLLABORATOR gets it with no edit. Denying it means splitting the one rule that currently serves primary, lead, worker and architect — a change to four rows to restrict one — and Authz.java's own comment calls that arm "the load-bearing rule". The risk is also bounded: a collaborator with nobody blocked on it gets a clean NO_WAITER refusal from the mechanism, and when a send is open the worst case is holding that sender's call, which any worker can already do.

The accepted limit: a collaborator cannot answer another collaborator's fleet_ask, because the turnId route is denied. A lead can. That is a usability limit, not a hole, and it should be written down rather than discovered.


Disagreement 3 — guard 2 does not do what it says. Confirmed, and it is a real pre-existing hole.

The accepted answer says to "widen the member-label collision validator so a fleet or profile tabLabel cannot match the collaborator namespace." One architect checked which validator that is and found it checks the wrong field.

I confirmed this. validateLeadTabPrefixes() reads leader.tabPrefix(). And across the whole of FleetConfig.java, leader.tab() appears in exactly two places:

2760:  .anyMatch(leader -> leader != null && leader.tab() != null && !leader.tab().isBlank());   // the pane-placement early return
2927:  if (leader.tab() == null || leader.tab().isBlank()) {                                      // validateMembers: a leader must name a tab

Identity is matched on the exact tab. tabPrefix is vestigial. So no validator compares a lead's exact tab to a member label. An operator who sets fleet.leaders.alpha.tab: "alpha" and a profile tabLabel: "alpha" passes every validator today.

Widening the prefix check would inherit that hole into the new namespace. Replace guard 2 with an exact/template collision check: refuse at startup when fleet.tabLabel, or any profile override, can render to a configured lead or collaborator tab, treating placeholders as wildcards.

The lead-side half of this is a pre-existing defect independent of #669 and is filed separately.


Two guards nobody listed, both confirmed by me

fleet_whoami would report a collaborator as a lead. The branch is if (!caller.isWorker()), and a COLLABORATOR is neither worker nor architect, so it falls through to the lead branch and gets a leader key while role reads collaborator. The answer contradicts itself. This also breaks the instruction surface: the canonical block in CLAUDE.md tells every session that fleet_whoami returns primary, worker or architect. A fourth value means that block changes in the same work.

Herdr routing would send a collaborator's status read to the wrong daemon. HerdrRouter.agentsFor picks isLead.test(targetId) ? leadAgents : memberAgents, and isLead is membership in the lead terminal map. A collaborator is not in it.


Two open questions I answered from the live config, which no member can read

Both architects correctly said they could not settle these, because fleetd/fleetd.yaml is gitignored and absent from their worktrees. I checked:

  1. No memberHerdrSocket is configured — grep -nE 'memberHerdrSocket|herdrSocket' fleetd/fleetd.yaml returns nothing. So memberHerdr == herdr on this host and the routing bug above is latent, not live. It must still be fixed, because a second herdr daemon is exactly the direction fleet01 goes.
  2. No profile uses placement: pane — all eight say placement: tab. So the widened pane-placement refusal is a quiet pass here rather than a startup failure. The change is safe to deploy on this fleet.

The plan — six units, in dependency order

A ──┐
    ├──> C ──> D ──> E ──> F
B ──┘

Unit A — split SEND, and split READ. Must land first. Separate the local send, the coordId broker route and the turnId answer form into distinct actions, and separate ticket-polling and session-status from roster READ. Grant every new action to exactly who holds it today, so no caller gains or loses anything. This is a pure refactor and merges on its own. Acceptance: the three fleet_send call shapes reach the gate as three different actions; flipping any one arm to deny refuses only that call shape and leaves the other two working — that control is what proves the arms are really separate and not a rename.

Unit B — the fleet.collaborators config block and its validators. Recognise-only: no profile, no instances, never auto-launched. Note that Fleet is @JsonIgnoreProperties(ignoreUnknown = true), so the key is silently ignored today — assert the before state in the test, or it can pass vacuously. Includes the widened pane-placement refusal and the replacement exact/template collision check.

Unit C — the COLLABORATOR role and its authorization row. Needs A and B. Acceptance: a local send naming a configured lead or collaborator is permitted, the same call naming a spawned member's terminal is refused, and that refusal happens over both MCP and POST /sessions/{id}/message — a test covering only MCP leaves the REST route open. Also: fleet_whoami reports role: collaborator with no leader key, while the same test fed a lead still reports leader — the second half is the control.

Unit D — the resolver. One generalised scan returning terminal → (name, kind); spawned-member-first ordering. The order is: no pane → token tail; terminal in the member roster → return that member's role and consult no tab map at all; then lead; then collaborator; then the worker fallback. Acceptance: a terminal in both the roster and the lead tab map resolves as its member role — that is #661 closed at the resolver, not only at config validation — and removing the spawned-member step must make that assertion fail. This is the security-critical unit and gets a reviewer who did not write it.

Unit E — deliverability. Widen the isLead predicate to "lead or collaborator" and rename it. Latent on this host, per the config check above.

Unit F — the instruction surface. Mine. The canonical block's fleet_whoami ladder, invariant 3's "send is lead or architect" line, the intent→tool table, and a wiki/11-Features.md entry. A member cannot do this: the CLAUDE.md/wiki sync check needs wiki/, which is uninitialized in every worker worktree.


What is still not settled, and what nobody has tested

Nothing above exists. No architect wrote code and no build was run against any of it, so every acceptance criterion here is a proposal, not a passing test.

Neither architect read herdr's own server code. Neither ran the live walk-through, correctly — it needs a daemon that has the role, and it would cut their own channel.

I have not re-checked every row of either capability table line by line. I verified the contested rows — READ, ASK, SEND, the whoami branch, pollAction, ticket generation and validateLeadTabPrefixes — and I am taking the uncontested rows on two architects agreeing.

This ticket stays open until Unit A lands.

## Second architect position collected, and the design is now settled. Implementation plan below. Two architects worked this question independently, on the same brief, unable to see each other. Neither could see the other's answer, so where they agree it is two positions rather than one. I checked every contested claim myself in the main clone at `6f27522` before ruling. **The headline: the recommendation stands, but it is not safe to implement as written.** Both architects independently concluded that `SEND` must be split into separate actions *before* the role exists. Each found gaps the other missed, and one of them overturns a guard the accepted answer called mandatory. --- ## What both architects concluded, independently Build it in fleetd. Add `COLLABORATOR` and `fleet.collaborators.<name>.tab`. Do not make `ARCHITECT` attachable. Deny `SPAWN`, `STOP`, `DRAIN`, `HANDOVER`, `COORD_READ` and the cross-host `coordId` route. Grant `METRICS` and own-pane `REPLY`. The #661 attack is real, and the resolver must check `SessionManager` before any tab map. Generalise the existing tab scan instead of adding a second scanner. The live walk-through is the lead's, not a worker's. **Both also found, separately, that splitting `SEND` must land first.** That is the strongest signal in the two reports, because neither could have copied it. One architect added a mechanical reason not to make `ARCHITECT` attachable that nobody had given: `Principal.isSpawnedMember()` is `WORKER || ARCHITECT`, and `FleetMcp.java:437-440` uses it to enrol a caller in the presence map. An attachable architect would be listed as an available member, so a lead's `fleet_list` would offer the operator's own hand-opened session as a worker to delegate to. --- ## Disagreement 1 — `READ`. I rule for the architect who objected. The accepted answer grants `READ`. One architect agreed and checked that `fleet_list` hides the coordinator row. The other refused, saying `READ` exposes task replies and pending questions, not a safe roster. **I checked this myself, and the objection is right.** Three facts, each from a command I ran: `fleet_status` returns the pending question, the `turnId` and the ticket id: ```java // FleetMcp.java return text(base + "\n\n[question — worker is waiting for your answer]\n" + ask.question() + "\n\nAnswer it by calling fleet_send again with turnId=\"" + ask.turnId() + "\" and content set to your answer; the worker resumes the same turn." + " (ticket " + ask.ticket() + ")"); ``` Polling a ticket is plain `READ` — no target means no `DRAIN`: ```java static Authz.Action pollAction(String target, String coordId) { if (!isBlank(coordId)) { return Authz.Action.COORD_READ; } return isBlank(target) ? Authz.Action.READ : Authz.Action.DRAIN; } ``` And ticket ids are a plain in-memory counter: ``` fleetd/src/main/java/dev/ltms/fleet/msg/MessageService.java:342 private final AtomicLong ticketSeq = new AtomicLong(); fleetd/src/main/java/dev/ltms/fleet/msg/MessageService.java:1282 String ticket = "task-" + ticketSeq.incrementAndGet(); ``` So the ids are `task-1`, `task-2`, … and they reset on every daemon restart. **I have first-hand evidence of how guessable that is from this session**: `grep -n "task-5" fleetd/fleetd.out` returns **20** separate occurrences against 20 different sessions. Put together: a collaborator holding today's `READ` can walk `task-1..N` and read the reply of every delegation the lead has run since the last restart, and can call `fleet_status` on any member to lift its open question and `turnId`. The other architect established the same mechanism and drew the opposite conclusion from it — it used "`fleet_poll{ticket}` maps to `READ` and returns the reply text" as the reason `DRAIN` can stay denied without breaking the round trip. That is true and useful, and it is also exactly the leak. The same fact serves both arguments; only one of them noticed it cuts both ways. **Ruling: `READ` must be split before `COLLABORATOR` is granted anything.** Keep `READ` for roster, profiles and identity. Move ticket polling and session status to their own actions and deny them to `COLLABORATOR` in the first release. Worth recording: the `READ` arm's own comment in `Authz.java` says *"the roster carries no secrets."* That is now false, and it is false for `WORKER` today, not only for a future collaborator. Filed separately. --- ## Disagreement 2 — `ASK`. I rule for granting it. One architect denied `ASK`, arguing it pairs with the `turnId` answer form. The other argued to grant. I side with granting, on the mechanism: ```java case REPLY, ASK -> caller.ownsSession(targetSession); ``` That arm is **ownership-based, not role-based**, so `COLLABORATOR` gets it with no edit. Denying it means splitting the one rule that currently serves primary, lead, worker and architect — a change to four rows to restrict one — and `Authz.java`'s own comment calls that arm "the load-bearing rule". The risk is also bounded: a collaborator with nobody blocked on it gets a clean `NO_WAITER` refusal from the mechanism, and when a send *is* open the worst case is holding that sender's call, which any worker can already do. **The accepted limit:** a collaborator cannot answer another collaborator's `fleet_ask`, because the `turnId` route is denied. A lead can. That is a usability limit, not a hole, and it should be written down rather than discovered. --- ## Disagreement 3 — guard 2 does not do what it says. Confirmed, and it is a real pre-existing hole. The accepted answer says to "widen the member-label collision validator so a fleet or profile `tabLabel` cannot match the collaborator namespace." One architect checked which validator that is and found it checks the wrong field. **I confirmed this.** `validateLeadTabPrefixes()` reads `leader.tabPrefix()`. And across the whole of `FleetConfig.java`, `leader.tab()` appears in exactly two places: ``` 2760: .anyMatch(leader -> leader != null && leader.tab() != null && !leader.tab().isBlank()); // the pane-placement early return 2927: if (leader.tab() == null || leader.tab().isBlank()) { // validateMembers: a leader must name a tab ``` Identity is matched on the exact `tab`. `tabPrefix` is vestigial. **So no validator compares a lead's exact `tab` to a member label.** An operator who sets `fleet.leaders.alpha.tab: "alpha"` and a profile `tabLabel: "alpha"` passes every validator today. Widening the prefix check would inherit that hole into the new namespace. **Replace guard 2** with an exact/template collision check: refuse at startup when `fleet.tabLabel`, or any profile override, can render to a configured lead or collaborator `tab`, treating placeholders as wildcards. The lead-side half of this is a pre-existing defect independent of #669 and is filed separately. --- ## Two guards nobody listed, both confirmed by me **`fleet_whoami` would report a collaborator as a lead.** The branch is `if (!caller.isWorker())`, and a `COLLABORATOR` is neither worker nor architect, so it falls through to the lead branch and gets a `leader` key while `role` reads `collaborator`. The answer contradicts itself. This also breaks the instruction surface: the canonical block in `CLAUDE.md` tells every session that `fleet_whoami` returns `primary`, `worker` or `architect`. A fourth value means that block changes in the same work. **Herdr routing would send a collaborator's status read to the wrong daemon.** `HerdrRouter.agentsFor` picks `isLead.test(targetId) ? leadAgents : memberAgents`, and `isLead` is membership in the **lead** terminal map. A collaborator is not in it. --- ## Two open questions I answered from the live config, which no member can read Both architects correctly said they could not settle these, because `fleetd/fleetd.yaml` is gitignored and absent from their worktrees. I checked: 1. **No `memberHerdrSocket` is configured** — `grep -nE 'memberHerdrSocket|herdrSocket' fleetd/fleetd.yaml` returns nothing. So `memberHerdr == herdr` on this host and the routing bug above is **latent, not live**. It must still be fixed, because a second herdr daemon is exactly the direction fleet01 goes. 2. **No profile uses `placement: pane`** — all eight say `placement: tab`. So the widened pane-placement refusal is a quiet pass here rather than a startup failure. The change is safe to deploy on this fleet. --- ## The plan — six units, in dependency order ``` A ──┐ ├──> C ──> D ──> E ──> F B ──┘ ``` **Unit A — split `SEND`, and split `READ`. Must land first.** Separate the local send, the `coordId` broker route and the `turnId` answer form into distinct actions, and separate ticket-polling and session-status from roster `READ`. Grant every new action to exactly who holds it today, so no caller gains or loses anything. This is a pure refactor and merges on its own. *Acceptance: the three `fleet_send` call shapes reach the gate as three different actions; flipping any one arm to deny refuses only that call shape and leaves the other two working — that control is what proves the arms are really separate and not a rename.* **Unit B — the `fleet.collaborators` config block and its validators.** Recognise-only: no `profile`, no `instances`, never auto-launched. Note that `Fleet` is `@JsonIgnoreProperties(ignoreUnknown = true)`, so the key is silently ignored today — assert the before state in the test, or it can pass vacuously. Includes the widened pane-placement refusal and the replacement exact/template collision check. **Unit C — the `COLLABORATOR` role and its authorization row.** Needs A and B. *Acceptance: a local send naming a configured lead or collaborator is permitted, the same call naming a spawned member's terminal is refused, and that refusal happens over **both** MCP and `POST /sessions/{id}/message` — a test covering only MCP leaves the REST route open.* Also: `fleet_whoami` reports `role: collaborator` with no `leader` key, while the same test fed a lead still reports `leader` — the second half is the control. **Unit D — the resolver.** One generalised scan returning `terminal → (name, kind)`; spawned-member-first ordering. The order is: no pane → token tail; terminal in the member roster → return that member's role and **consult no tab map at all**; then lead; then collaborator; then the worker fallback. *Acceptance: a terminal in both the roster and the lead tab map resolves as its member role — that is #661 closed at the resolver, not only at config validation — and removing the spawned-member step must make that assertion fail.* This is the security-critical unit and gets a reviewer who did not write it. **Unit E — deliverability.** Widen the `isLead` predicate to "lead or collaborator" and rename it. Latent on this host, per the config check above. **Unit F — the instruction surface. Mine.** The canonical block's `fleet_whoami` ladder, invariant 3's "send is lead or architect" line, the intent→tool table, and a `wiki/11-Features.md` entry. A member cannot do this: the `CLAUDE.md`/wiki sync check needs `wiki/`, which is uninitialized in every worker worktree. --- ## What is still not settled, and what nobody has tested Nothing above exists. No architect wrote code and no build was run against any of it, so every acceptance criterion here is a proposal, not a passing test. Neither architect read herdr's own server code. Neither ran the live walk-through, correctly — it needs a daemon that has the role, and it would cut their own channel. I have not re-checked every row of either capability table line by line. I verified the contested rows — `READ`, `ASK`, `SEND`, the `whoami` branch, `pollAction`, ticket generation and `validateLeadTabPrefixes` — and I am taking the uncontested rows on two architects agreeing. This ticket stays open until Unit A lands.
Author
Owner

The two "filed separately" lines in the adjudication above now have tickets:

  • #678 — Authz.java:87-89 claims "the roster carries no secrets" under the READ arm. False today: fleet_poll{ticket} is READ, ticket ids are a sequential counter with no owner check, and fleet_status returns another member's open question plus its turnId. That ticket is the written evidence for why Unit A must land first, so the reason survives if Unit A is re-litigated.
  • #677 — no validator compares a lead's exact tab to a member label. The collision guard checks tabPrefix, which identity no longer uses. Pre-existing and independent of this ticket. #669 Unit D would close it at the resolver, which is where it bites.

Neither blocks Unit A.

The two "filed separately" lines in the adjudication above now have tickets: - **#678** — `Authz.java:87-89` claims "the roster carries no secrets" under the `READ` arm. False today: `fleet_poll{ticket}` is `READ`, ticket ids are a sequential counter with no owner check, and `fleet_status` returns another member's open question plus its `turnId`. That ticket is the written evidence for why **Unit A must land first**, so the reason survives if Unit A is re-litigated. - **#677** — no validator compares a lead's exact `tab` to a member label. The collision guard checks `tabPrefix`, which identity no longer uses. Pre-existing and independent of this ticket. #669 Unit D would close it at the resolver, which is where it bites. Neither blocks Unit A.
Author
Owner

Unit A has landed

PR #687 merged as 7c458e8 on main (pushed; main is now 2eb2d61). It closes #678.

SEND is now three actions (SEND, COORD_SEND, ANSWER) and READ is two (READ,
TASK_READ), each new action carrying exactly the grant its combined action carried. Verified: a
reviewer compared all 156 (role, action, target) pairs against origin/main with 0 mismatches, and
I read the grants in Authz myself. Merged tree: 1942 tests, 0 failures, BUILD SUCCESS.

So the divergence point the rest of this ticket needs now exists.

One thing came out of Unit A that the later units should know about, filed as #689: the REST
route now parses the request body before the authorization gate, because turnId picks the action.
That is harmless today only because all three send grants are identical. Unit B gives
COLLABORATOR a different grant, and at that moment #689 stops being cosmetic.
Fix #689 before or
with Unit B, not after.

State of the remaining units

Units B to F are specified in the 20:58 comment above and none is delegated yet. Unit F — the
instruction surface (the fleet_whoami ladder, invariant 3, the intent→tool table, a Features
entry) — stays with the lead, because a member cannot run the CLAUDE.md/wiki sync check: wiki/
is uninitialized in every worker worktree.

This ticket stays open.

## Unit A has landed PR #687 merged as `7c458e8` on `main` (pushed; `main` is now `2eb2d61`). It closes #678. `SEND` is now three actions (`SEND`, `COORD_SEND`, `ANSWER`) and `READ` is two (`READ`, `TASK_READ`), each new action carrying exactly the grant its combined action carried. Verified: a reviewer compared all 156 (role, action, target) pairs against `origin/main` with 0 mismatches, and I read the grants in `Authz` myself. Merged tree: **1942 tests, 0 failures, BUILD SUCCESS**. So the divergence point the rest of this ticket needs now exists. One thing came out of Unit A that the later units should know about, filed as **#689**: the REST route now parses the request body before the authorization gate, because `turnId` picks the action. That is harmless today only because all three send grants are identical. **Unit B gives `COLLABORATOR` a different grant, and at that moment #689 stops being cosmetic.** Fix #689 before or with Unit B, not after. ## State of the remaining units Units B to F are specified in the 20:58 comment above and **none is delegated yet**. Unit F — the instruction surface (the `fleet_whoami` ladder, invariant 3, the intent→tool table, a Features entry) — stays with the lead, because a member cannot run the `CLAUDE.md`/wiki sync check: `wiki/` is uninitialized in every worker worktree. This ticket stays open.
Author
Owner

State update: Unit A's follow-up is closed, and the daemon is now live on it. Units B–F still open.

main is at edbd8d8 and pushed. Four PRs merged since Unit A landed, each verified by me in a throwaway worktree with a tree-hash comparison against what I built:

PR Ticket On main
#690 #638 — verdict userinfo over-masking 7dec74f
#691 #677 — lead tab collision guard bfee23a
#695 #693, #676 — pin case-insensitivity, drop stale count d0688c8
#694 #689 — gate before body parse edbd8d8

#689 is fixed, which unblocks Unit B. This ticket said to fix it "before or with Unit B, not after", and it is now done and merged. FleetApp.sendMessage checks the coarse SEND grant with no body read, then parses, then checks ANSWER when turnId is present.

One thing worth carrying into Unit B: the ANSWER gate is behaviourally invisible today, because all three send grants are identical. So I first merged a version where deleting the call site left all 56 tests in those classes green. It is now pinned through the audit trail — allow() logs an allowed entry for every granted action except READ/METRICS/TASK_READ, so a granted turnId request must log both SEND and ANSWER. Deleting the call site now fails exactly one test. When Unit B gives COLLABORATOR a different grant, that test is what stops the gate having quietly disappeared in the meantime.

The daemon is redeployed, so Unit A is actually live

A merge is not a deployment and Unit A's split had never been loaded. scripts/redeploy-fleetd.sh --yes: build 1950 tests, 0 failures, BUILD SUCCESS; old pid 34147 exited with a clean drain (released=0 abandoned=0); new pid 56206 on jar e0e9b2c4109b; /healthz 200; a fresh fleetd listening line at 23:00:48.

I did not stop at healthz, because healthz only proves herdr answers. I ran a real spawn and a full round trip — spawn, send, the member ran a command, fleet_reply came back with edbd8d8. The channel works on the new jar. fleet_whoami still answers primary.

Units B–F

None is delegated. The specs in the 20:58 comment stand unchanged. Two notes for whoever picks them up:

  • Unit B now overlaps landed work. #677's PR already added templateCanRenderAs(...), which is the "replacement exact/template collision check" Unit B was going to include. Re-read FleetConfig.validateLeadTabPrefixes() before writing Unit B's brief — part of it exists. What does not exist is the fleet.collaborators block, its validators, or the widened pane-placement refusal.
  • Unit B must not be delegated alongside anything else touching FleetConfig.java. I had to hold it back twice today for exactly this reason. Two file-disjoint units in this project have collided in shared test files before.

Unit F is still the lead's. I confirmed the reason is live: the CLAUDE.md/wiki sync check needs wiki/, and it is uninitialized in every worker worktree. I ran that check myself today and it reports in sync.

One Unit F item is already done

wiki/11-Features.md claimed "Two startup refusals" guard the lead tab namespace and described tabPrefix as the collision guard. #677 made both statements false — there are four refusals now, and the load-bearing one compares the exact tab. I corrected that section and pushed the wiki (9b4ae2e..b025ae9 on refs/heads/main, verified by ref, because this wiki has both a main and a master and pushing to the wrong one is silent).

The rest of Unit F — the fleet_whoami ladder, invariant 3's "send is lead or architect" line, the intent→tool table — is not done, and must not be done until COLLABORATOR actually exists. Writing a fourth fleet_whoami value into the canonical block before the code returns it would tell every session something false.

## State update: Unit A's follow-up is closed, and the daemon is now live on it. Units B–F still open. `main` is at **`edbd8d8`** and pushed. Four PRs merged since Unit A landed, each verified by me in a throwaway worktree with a tree-hash comparison against what I built: | PR | Ticket | On `main` | |---|---|---| | #690 | #638 — verdict userinfo over-masking | `7dec74f` | | #691 | #677 — lead tab collision guard | `bfee23a` | | #695 | #693, #676 — pin case-insensitivity, drop stale count | `d0688c8` | | #694 | **#689** — gate before body parse | `edbd8d8` | **#689 is fixed, which unblocks Unit B.** This ticket said to fix it "before or with Unit B, not after", and it is now done and merged. `FleetApp.sendMessage` checks the coarse `SEND` grant with no body read, then parses, then checks `ANSWER` when `turnId` is present. One thing worth carrying into Unit B: the ANSWER gate is **behaviourally invisible today**, because all three send grants are identical. So I first merged a version where deleting the call site left all 56 tests in those classes green. It is now pinned through the audit trail — `allow()` logs an allowed entry for every granted action except `READ`/`METRICS`/`TASK_READ`, so a granted `turnId` request must log both `SEND` and `ANSWER`. Deleting the call site now fails exactly one test. **When Unit B gives `COLLABORATOR` a different grant, that test is what stops the gate having quietly disappeared in the meantime.** ### The daemon is redeployed, so Unit A is actually live A merge is not a deployment and Unit A's split had never been loaded. `scripts/redeploy-fleetd.sh --yes`: build **1950 tests, 0 failures, BUILD SUCCESS**; old pid 34147 exited with a clean drain (`released=0 abandoned=0`); new pid **56206** on jar `e0e9b2c4109b`; `/healthz` 200; a fresh `fleetd listening` line at 23:00:48. I did not stop at healthz, because healthz only proves herdr answers. I ran a **real spawn** and a full round trip — spawn, send, the member ran a command, `fleet_reply` came back with `edbd8d8`. The channel works on the new jar. `fleet_whoami` still answers `primary`. ### Units B–F **None is delegated.** The specs in the 20:58 comment stand unchanged. Two notes for whoever picks them up: - **Unit B now overlaps landed work.** #677's PR already added `templateCanRenderAs(...)`, which is the "replacement exact/template collision check" Unit B was going to include. Re-read `FleetConfig.validateLeadTabPrefixes()` before writing Unit B's brief — part of it exists. What does **not** exist is the `fleet.collaborators` block, its validators, or the widened pane-placement refusal. - **Unit B must not be delegated alongside anything else touching `FleetConfig.java`.** I had to hold it back twice today for exactly this reason. Two file-disjoint units in this project have collided in shared test files before. **Unit F is still the lead's.** I confirmed the reason is live: the `CLAUDE.md`/wiki sync check needs `wiki/`, and it is uninitialized in every worker worktree. I ran that check myself today and it reports in sync. ### One Unit F item is already done `wiki/11-Features.md` claimed "**Two** startup refusals" guard the lead tab namespace and described `tabPrefix` as the collision guard. #677 made both statements false — there are four refusals now, and the load-bearing one compares the exact `tab`. I corrected that section and pushed the wiki (`9b4ae2e..b025ae9` on `refs/heads/main`, verified by ref, because this wiki has both a `main` and a `master` and pushing to the wrong one is silent). The rest of Unit F — the `fleet_whoami` ladder, invariant 3's "send is lead or architect" line, the intent→tool table — is **not** done, and must not be done until `COLLABORATOR` actually exists. Writing a fourth `fleet_whoami` value into the canonical block before the code returns it would tell every session something false.
Author
Owner

Unit B correction — one addition before merge: collaborators is missing from the duplicate-key guard

This is a correction to Unit B's scope, written here rather than sent as a revised brief, because the brief is write-once and a ticket comment is what a member can re-read whenever it looks.

PR #697 is open and under review. Its own §7 report flagged this, and I confirmed it by reading the code — so it is being folded in rather than filed for later.

What is missing

FleetConfig.java:1912-1913:

private static final Set<String> FLEET_POOL_KEYS =
        Set.of("leaders", "architects", "developers", "hunters", "reviewers");

collaborators is not in that set. The walk at FleetConfig.java:1970 only descends into a pool whose key is in it:

if (FLEET_POOL_KEYS.contains(pool) && value == JsonToken.START_OBJECT) {
    rejectDuplicateChildSlotKeys(p, pool);
} else {
    skipValue(p, value);
}

So a config naming the same collaborator twice takes the skipValue branch.

Why that matters, in the existing code's own words

The javadoc on rejectDuplicateMemberSlots already states the reason this guard exists:

Each pool is a Map keyed by slot name, so by the time it is read duplicate keys have already collapsed last-wins — a duplicated slot name would silently drop one slot and the daemon would never know. Jackson's YAML parser does not fail on duplicate mapping keys by default, so duplicates are caught here, at parse time, before the map is built.

Every word of that applies to collaborators. An operator who writes fleet.collaborators.alpha twice gets one of them silently discarded, and the collaborator they configured is simply not recognised — the same harm the five existing pools are protected from.

This is a gap in the block Unit B is adding, not a pre-existing defect elsewhere. That is why it belongs in Unit B and not in a follow-up: shipping the block without it means shipping a known hole.

Scope of the addition

Three things, and nothing else:

  1. Add "collaborators" to FLEET_POOL_KEYS.
  2. The javadoc two lines below says "Only the five pools directly under the top-level fleet:". It becomes six. A count written in prose next to the thing it counts is the exact stale-count class #676 was about — fix it in the same commit or it is a new instance of it.
  3. A test that a duplicated collaborator name is refused at parse time, plus the control that a collaborator name repeated across different pools is still not a duplicate. The existing tests to model on are in FleetConfigTest.java around lines 1224-1348: duplicateSlotNamesInOnePoolAreRejectedAtParseTime, theSameSlotNameInTwoPoolsIsNotADuplicate, duplicateKeysOutsideTheFleetPoolsAreUnaffected.

I checked and no test pins the contents of FLEET_POOL_KEYS — grep -rn 'FLEET_POOL_KEYS' fleetd/src/test/java is empty. So adding a key breaks nothing, and nothing would have caught the omission either. Worth knowing when judging how this was missed.

What I measured, and what I did not

I read the set, the walk at :1970 and the javadoc directly on PR #697's branch. I did not run a live config with a duplicated collaborator key through the parser. The conclusion that it collapses silently is a code read, not an observed output — the test asked for in item 3 is what will turn it into a measurement.

Still out of scope for Unit B

Unchanged: Authz, CallerResolver, LeadTabScanner, MemberRole, FleetMcp. Those are Units C, D and E.

## Unit B correction — one addition before merge: `collaborators` is missing from the duplicate-key guard This is a correction to Unit B's scope, written here rather than sent as a revised brief, because the brief is write-once and a ticket comment is what a member can re-read whenever it looks. PR #697 is open and under review. Its own §7 report flagged this, and **I confirmed it by reading the code** — so it is being folded in rather than filed for later. ### What is missing `FleetConfig.java:1912-1913`: ```java private static final Set<String> FLEET_POOL_KEYS = Set.of("leaders", "architects", "developers", "hunters", "reviewers"); ``` `collaborators` is not in that set. The walk at `FleetConfig.java:1970` only descends into a pool whose key is in it: ```java if (FLEET_POOL_KEYS.contains(pool) && value == JsonToken.START_OBJECT) { rejectDuplicateChildSlotKeys(p, pool); } else { skipValue(p, value); } ``` So a config naming the same collaborator twice takes the `skipValue` branch. ### Why that matters, in the existing code's own words The javadoc on `rejectDuplicateMemberSlots` already states the reason this guard exists: > Each pool is a `Map` keyed by slot name, so by the time it is read duplicate keys have already collapsed last-wins — a duplicated slot name would silently drop one slot and the daemon would never know. Jackson's YAML parser does not fail on duplicate mapping keys by default, so duplicates are caught here, at parse time, before the map is built. Every word of that applies to `collaborators`. An operator who writes `fleet.collaborators.alpha` twice gets one of them silently discarded, and the collaborator they configured is simply not recognised — the same harm the five existing pools are protected from. **This is a gap in the block Unit B is adding, not a pre-existing defect elsewhere.** That is why it belongs in Unit B and not in a follow-up: shipping the block without it means shipping a known hole. ### Scope of the addition Three things, and nothing else: 1. Add `"collaborators"` to `FLEET_POOL_KEYS`. 2. The javadoc two lines below says *"Only the **five** pools directly under the top-level `fleet:`"*. It becomes six. A count written in prose next to the thing it counts is the exact stale-count class #676 was about — fix it in the same commit or it is a new instance of it. 3. A test that a duplicated collaborator name is refused at parse time, plus the control that a collaborator name repeated across *different* pools is still not a duplicate. The existing tests to model on are in `FleetConfigTest.java` around lines 1224-1348: `duplicateSlotNamesInOnePoolAreRejectedAtParseTime`, `theSameSlotNameInTwoPoolsIsNotADuplicate`, `duplicateKeysOutsideTheFleetPoolsAreUnaffected`. I checked and **no test pins the contents of `FLEET_POOL_KEYS`** — `grep -rn 'FLEET_POOL_KEYS' fleetd/src/test/java` is empty. So adding a key breaks nothing, and nothing would have caught the omission either. Worth knowing when judging how this was missed. ### What I measured, and what I did not I read the set, the walk at `:1970` and the javadoc directly on PR #697's branch. **I did not run a live config with a duplicated collaborator key through the parser.** The conclusion that it collapses silently is a code read, not an observed output — the test asked for in item 3 is what will turn it into a measurement. ### Still out of scope for Unit B Unchanged: `Authz`, `CallerResolver`, `LeadTabScanner`, `MemberRole`, `FleetMcp`. Those are Units C, D and E.
Author
Owner

Unit B round 2 — two required changes, both about comments telling the truth

PR #697's code is sound. I mutation-tested it myself (results below) and found no defect in the logic. Both items here are about text, and the first one is partly my fault.

Required 1 — the test javadoc narrates history, and my own ticket comment caused it

FleetConfigTest.java:1517:

fleetd #669 ticket correction (23:37): collaborators was missing from FLEET_POOL_KEYS, so a duplicated collaborator name took the same silent last-wins path the other five pools are guarded against. This is the measurement the ticket comment asked for: before the fix, this failed to throw at all (the parser accepted the file and collaborators held only the last of the two entries); it now fails exactly like …

The project rule is explicit and this breaks four parts of it at once: history (was missing, before the fix), evidence (this is the measurement, a 23:37 timestamp), a ticket as the reason, and — for a test specifically — "A test comment names the behaviour. It does not tell the story of the bug it caught."

I asked for this. My 23:37 comment said the test "is what turns it into a measurement", and the implementer dutifully wrote the measurement into the javadoc. The measurement belongs in the PR description and the commit message, which is where I should have said to put it. The rule the implementer followed was mine, and it was wrong.

Replace it with what the test protects, in the present tense — roughly "A duplicated name in fleet.collaborators is refused at parse time, like any other fleet: pool." No ticket, no date, no before-state.

That history/evidence shape appears in exactly this one comment block. I swept all 86 added comment lines for before the fix|was missing|previously|no longer|measurement|measured|verified|ticket correction, and nothing else matched.

Required 2 — two places promise behaviour that does not exist yet

FleetConfig.java:1183-1184, on the Collaborator record:

Nothing here ever launches a pane — a collaborator is a human-opened tab the daemon learns to address, never a member the daemon can spawn.

fleetd.example.yaml:672-673:

Tabs fleetd recognises as collaborators — a human-opened tab the daemon learns to address

Neither is true after Unit B. Nothing recognises a collaborator and nothing addresses one: COLLABORATOR arrives in Unit C, resolution in Unit D, deliverability in Unit E. Today this block is parsed and validated, and that is all.

This is the same hazard already recorded against Unit F — writing a fourth fleet_whoami value into the canonical block before the code returns it would tell every session something false. fleetd.example.yaml has exactly the same readership problem, and it is worse than CLAUDE.md here because an operator uncommenting that block would reasonably expect something to happen.

There is a second reason to cut the sentence from the record javadoc, independent of timing: javadoc on a public type is a contract. This record's contract is "a tab label, keyed by name; tab is required and compared case-insensitively." Whether the daemon addresses that tab is system behaviour owned by other classes, so the sentence is in the wrong place as well as premature.

What to do:

  • In the record javadoc, drop the "learns to address" clause. Keep the contract: recognise-only, no profile/instances/kind, tab required, compared case-insensitively.
  • In fleetd.example.yaml, keep the block (an existing test requires the key to be documented — I confirmed everyNestedConfigKeyIsDocumentedInTheExample forces this), but state plainly that the block is accepted and validated now, and that recognition and addressing arrive with #669 Units C-E. One sentence. An operator must not read today's example as a working feature.

Optional, one word each — do not expand these

Only if you are already in the file. Neither is worth a round trip on its own, and I would rather you under-edit than restyle working comments:

  • FleetConfig.java:1186 — "tabPrefix is deliberately absent." The reason follows in the next clause, so by the project's own rule the confidence marker adds nothing.
  • FleetConfig.java:1187 — "the same as a Leader, so a naming-convention prefix would be vestigial here too" points at another class to justify this one. The rule asks for the fact stated once where it is enforced, not a cross-reference.

What I verified myself, so you know the logic is not in question

Three line-anchored mutations in a throwaway worktree of ad593c9, each compiled before running, each restored after:

mutation result
equalsIgnoreCase → equals on the lead-vs-collaborator tab comparison killed by exactly 1 test, aCollaboratorTabEqualToALeadTabRefusesToStart
equalsIgnoreCase → equals on the collaborator-vs-collaborator comparison killed by exactly 1 test, twoCollaboratorsSharingTheSameExactTabRefusesToStart
the widened early return put back to if (!anyLeaderHasTab) — the original named bug killed by exactly 1 test, aPanePlacedProfileWithACollaboratorTabRefusesToStartEvenWithNoLeaders

Three specific runtime kills, no cascades. The case-insensitivity the example config advertises is genuinely pinned, and the early-return bug cannot silently come back. I deliberately mutated lines the implementer had not mutated itself, since it had already measured the FLEET_POOL_KEYS line.

I also confirmed independently: only one production call site of any Fleet overload (FleetConfig.java:2633, withDefaults(), all nulls), and collaborators was appended last in the canonical record so both convenience overloads keep their signatures and forward a trailing null. Had the field been inserted next to leaders where it reads more naturally, every positional argument in those overloads would have shifted — invisibly, because that call site passes only nulls. That was the right call and it is worth saying so.

A second reviewer found no issue on the Jackson binding and overload dimension, and named what it checked. The validator-logic reviewer has not reported yet; if it finds anything, that is a separate round.

## Unit B round 2 — two required changes, both about comments telling the truth PR #697's code is sound. I mutation-tested it myself (results below) and found no defect in the logic. Both items here are about text, and the first one is partly **my** fault. ### Required 1 — the test javadoc narrates history, and my own ticket comment caused it `FleetConfigTest.java:1517`: > `fleetd #669 ticket correction (23:37): collaborators was missing from FLEET_POOL_KEYS, so a duplicated collaborator name took the same silent last-wins path the other five pools are guarded against. This is the measurement the ticket comment asked for: before the fix, this failed to throw at all (the parser accepted the file and collaborators held only the last of the two entries); it now fails exactly like …` The project rule is explicit and this breaks four parts of it at once: history (*was missing*, *before the fix*), evidence (*this is the measurement*, a `23:37` timestamp), a ticket as the reason, and — for a test specifically — *"A test comment names the behaviour. It does not tell the story of the bug it caught."* **I asked for this.** My 23:37 comment said the test "is what turns it into a measurement", and the implementer dutifully wrote the measurement into the javadoc. The measurement belongs in the PR description and the commit message, which is where I should have said to put it. The rule the implementer followed was mine, and it was wrong. Replace it with what the test protects, in the present tense — roughly *"A duplicated name in `fleet.collaborators` is refused at parse time, like any other `fleet:` pool."* No ticket, no date, no before-state. That history/evidence shape appears in exactly this one comment block. I swept all 86 added comment lines for `before the fix|was missing|previously|no longer|measurement|measured|verified|ticket correction`, and nothing else matched. ### Required 2 — two places promise behaviour that does not exist yet `FleetConfig.java:1183-1184`, on the `Collaborator` record: > Nothing here ever launches a pane — a collaborator is a human-opened tab the daemon **learns to address**, never a member the daemon can spawn. `fleetd.example.yaml:672-673`: > Tabs fleetd **recognises as collaborators** — a human-opened tab the daemon **learns to address** Neither is true after Unit B. Nothing recognises a collaborator and nothing addresses one: `COLLABORATOR` arrives in Unit C, resolution in Unit D, deliverability in Unit E. Today this block is parsed and validated, and that is all. This is the same hazard already recorded against Unit F — writing a fourth `fleet_whoami` value into the canonical block before the code returns it would tell every session something false. `fleetd.example.yaml` has exactly the same readership problem, and it is worse than `CLAUDE.md` here because an operator uncommenting that block would reasonably expect something to happen. There is a second reason to cut the sentence from the record javadoc, independent of timing: **javadoc on a public type is a contract.** This record's contract is "a tab label, keyed by name; `tab` is required and compared case-insensitively." Whether the daemon addresses that tab is system behaviour owned by other classes, so the sentence is in the wrong place as well as premature. What to do: - In the record javadoc, drop the "learns to address" clause. Keep the contract: recognise-only, no `profile`/`instances`/`kind`, `tab` required, compared case-insensitively. - In `fleetd.example.yaml`, keep the block (an existing test requires the key to be documented — I confirmed `everyNestedConfigKeyIsDocumentedInTheExample` forces this), but state plainly that the block is accepted and validated now, and that recognition and addressing arrive with #669 Units C-E. One sentence. An operator must not read today's example as a working feature. ### Optional, one word each — do not expand these Only if you are already in the file. Neither is worth a round trip on its own, and I would rather you under-edit than restyle working comments: - `FleetConfig.java:1186` — "`tabPrefix` is **deliberately** absent." The reason follows in the next clause, so by the project's own rule the confidence marker adds nothing. - `FleetConfig.java:1187` — "the same as a `Leader`, so a naming-convention prefix would be vestigial here too" points at another class to justify this one. The rule asks for the fact stated once where it is enforced, not a cross-reference. ### What I verified myself, so you know the logic is not in question Three line-anchored mutations in a throwaway worktree of `ad593c9`, each compiled before running, each restored after: | mutation | result | |---|---| | `equalsIgnoreCase` → `equals` on the lead-vs-collaborator tab comparison | killed by exactly 1 test, `aCollaboratorTabEqualToALeadTabRefusesToStart` | | `equalsIgnoreCase` → `equals` on the collaborator-vs-collaborator comparison | killed by exactly 1 test, `twoCollaboratorsSharingTheSameExactTabRefusesToStart` | | the widened early return put back to `if (!anyLeaderHasTab)` — the original named bug | killed by exactly 1 test, `aPanePlacedProfileWithACollaboratorTabRefusesToStartEvenWithNoLeaders` | Three specific runtime kills, no cascades. The case-insensitivity the example config advertises is genuinely pinned, and the early-return bug cannot silently come back. I deliberately mutated lines the implementer had not mutated itself, since it had already measured the `FLEET_POOL_KEYS` line. I also confirmed independently: only one production call site of any `Fleet` overload (`FleetConfig.java:2633`, `withDefaults()`, all nulls), and `collaborators` was appended **last** in the canonical record so both convenience overloads keep their signatures and forward a trailing null. Had the field been inserted next to `leaders` where it reads more naturally, every positional argument in those overloads would have shifted — invisibly, because that call site passes only nulls. That was the right call and it is worth saying so. A second reviewer found no issue on the Jackson binding and overload dimension, and named what it checked. The validator-logic reviewer has not reported yet; if it finds anything, that is a separate round.
Author
Owner

Unit B round 3 — one more test required. A reviewer found it; I proved it by mutation, and it is worse than it was reported.

A reviewer reported this as a low-severity coverage gap. I checked it myself and it is a real gap, but the severity reasoning needs correcting, so do not treat "low" as a reason to skip it.

The gap

FleetConfig.java:2764, in the collaborator loop of validateLeadTabPrefixes():

if (templateCanRenderAs(fleet.tabLabel(), tab)) {

That is the fleet-wide fleet.tabLabel check against a collaborator tab. Every collaborator test in the PR drives the profile-override branch a few lines below it instead, so this branch has no regression pin.

Measured, not read

I disabled the branch (if (false && …)) in a throwaway worktree of ad593c9 and ran both config test classes:

mvn exit=0   Tests run: 176, Failures: 0, Errors: 0   BUILD SUCCESS

The mutation survives. The refusal can be removed and the suite stays green.

A survivor has more than one explanation, so I ruled the others out: it is not an equivalent mutation, because disabling the check changes behaviour — and the reviewer independently wrote that exact scenario as a test and watched validateLeadTabPrefixes() throw and name the collaborator correctly. So the behaviour is right; only the pin is missing.

Why "low severity" undersells it — the lead side is pinned three times over

I ran the same mutation on the lead-side equivalent at FleetConfig.java:2736:

mvn exit=1   Tests run: 176, Failures: 3   BUILD FAILURE
  aFleetTabLabelTemplateThatCanRenderAsALeadTabRefusesToStart
  anExactFleetTabLabelCollisionRefusesToStart
  validateAllReachesEveryOneOfTodaysRealValidators

So this is not a pre-existing habit in this file. The lead branch is covered three ways; its new collaborator twin is covered zero ways. The implementation mirrored the lead logic correctly and did not mirror the lead's coverage, and that asymmetry arrived with this PR.

That matters here specifically. This ticket already records that #689's ANSWER gate was behaviourally invisible, that deleting its call site left 56 tests green, and that a later unit changing a grant is the moment it stops being cosmetic. An unpinned startup refusal is the same shape: it can be deleted as dead-looking code with a green build, and the thing it was guarding is a privilege boundary. Cheap to pin now, expensive to discover missing later.

What to add

One test: fleet.tabLabel set directly, no profile override, against a matching fleet.collaborators entry's tab; assert IllegalStateException and that the message names the collaborator. Mirror the existing lead-side aFleetTabLabelTemplateThatCanRenderAsALeadTabRefusesToStart.

Before you commit it, confirm it is a real pin the same way I did: disable line 2764 with if (false && …), see your new test fail, restore, see it pass. A test that does not go red against that mutation is not a pin, and that check is the whole point of this round.

Nothing else changes. No logic edits. Authz, CallerResolver, LeadTabScanner, MemberRole and FleetMcp stay out.

Two notes on process, for the record

My own first attempt at the lead-side mutation silently did nothing. I anchored the sed at line 2702, which is a bare } — the substitution never matched, the suite came back 176 green, and that reads exactly like "the lead side is unpinned too". I caught it only because I printed the mutated line afterwards and saw no false &&. A line-anchored mutation that fails to apply produces the most reassuring possible result. Print the line, every time.

The reviewer's review predates the latest commit. It reports running 167 tests in FleetConfigTest; that class holds 169 at ad593c9, and the duplicate-key round added exactly 2. So it reviewed 05244a8 and its findings do not cover the FLEET_POOL_KEYS addition. I am inferring that from the counts rather than measuring it. I am not ordering another review pass for those two tests, because I confirmed that change myself by a different route — the implementer reverted the single line and observed the refusal disappear, and I read the resulting set and walk directly.

The reviewer also ended its turn without a fleet_reply, so its answer arrived only through the pane scrape. The brief did tell it to end with one. That is now several occurrences of the same failure on review-shaped briefs, and it is a fleetd defect rather than a worker mistake — worth its own ticket, which I will file separately rather than bury here.

## Unit B round 3 — one more test required. A reviewer found it; I proved it by mutation, and it is worse than it was reported. A reviewer reported this as a **low-severity coverage gap**. I checked it myself and it is a real gap, but the severity reasoning needs correcting, so do not treat "low" as a reason to skip it. ### The gap `FleetConfig.java:2764`, in the collaborator loop of `validateLeadTabPrefixes()`: ```java if (templateCanRenderAs(fleet.tabLabel(), tab)) { ``` That is the **fleet-wide** `fleet.tabLabel` check against a collaborator tab. Every collaborator test in the PR drives the *profile-override* branch a few lines below it instead, so this branch has no regression pin. ### Measured, not read I disabled the branch (`if (false && …)`) in a throwaway worktree of `ad593c9` and ran both config test classes: ``` mvn exit=0 Tests run: 176, Failures: 0, Errors: 0 BUILD SUCCESS ``` **The mutation survives.** The refusal can be removed and the suite stays green. A survivor has more than one explanation, so I ruled the others out: it is not an equivalent mutation, because disabling the check changes behaviour — and the reviewer independently wrote that exact scenario as a test and watched `validateLeadTabPrefixes()` throw and name the collaborator correctly. So the behaviour is right; only the pin is missing. ### Why "low severity" undersells it — the lead side is pinned three times over I ran the same mutation on the **lead-side** equivalent at `FleetConfig.java:2736`: ``` mvn exit=1 Tests run: 176, Failures: 3 BUILD FAILURE aFleetTabLabelTemplateThatCanRenderAsALeadTabRefusesToStart anExactFleetTabLabelCollisionRefusesToStart validateAllReachesEveryOneOfTodaysRealValidators ``` So this is **not** a pre-existing habit in this file. The lead branch is covered three ways; its new collaborator twin is covered zero ways. The implementation mirrored the lead logic correctly and did not mirror the lead's coverage, and that asymmetry arrived with this PR. That matters here specifically. This ticket already records that #689's ANSWER gate was behaviourally invisible, that deleting its call site left 56 tests green, and that a later unit changing a grant is the moment it stops being cosmetic. An unpinned startup refusal is the same shape: it can be deleted as dead-looking code with a green build, and the thing it was guarding is a privilege boundary. Cheap to pin now, expensive to discover missing later. ### What to add One test: `fleet.tabLabel` set directly, **no** profile override, against a matching `fleet.collaborators` entry's tab; assert `IllegalStateException` and that the message names the collaborator. Mirror the existing lead-side `aFleetTabLabelTemplateThatCanRenderAsALeadTabRefusesToStart`. Before you commit it, confirm it is a real pin the same way I did: disable line 2764 with `if (false && …)`, see your new test fail, restore, see it pass. A test that does not go red against that mutation is not a pin, and that check is the whole point of this round. Nothing else changes. No logic edits. `Authz`, `CallerResolver`, `LeadTabScanner`, `MemberRole` and `FleetMcp` stay out. ### Two notes on process, for the record **My own first attempt at the lead-side mutation silently did nothing.** I anchored the `sed` at line 2702, which is a bare `}` — the substitution never matched, the suite came back 176 green, and that reads exactly like "the lead side is unpinned too". I caught it only because I printed the mutated line afterwards and saw no `false &&`. A line-anchored mutation that fails to apply produces the most reassuring possible result. Print the line, every time. **The reviewer's review predates the latest commit.** It reports running 167 tests in `FleetConfigTest`; that class holds 169 at `ad593c9`, and the duplicate-key round added exactly 2. So it reviewed `05244a8` and its findings do not cover the `FLEET_POOL_KEYS` addition. I am inferring that from the counts rather than measuring it. I am not ordering another review pass for those two tests, because I confirmed that change myself by a different route — the implementer reverted the single line and observed the refusal disappear, and I read the resulting set and walk directly. **The reviewer also ended its turn without a `fleet_reply`**, so its answer arrived only through the pane scrape. The brief did tell it to end with one. That is now several occurrences of the same failure on review-shaped briefs, and it is a fleetd defect rather than a worker mistake — worth its own ticket, which I will file separately rather than bury here.
Author
Owner

Unit B merged — PR #697, as f0ff252

Merged locally and pushed to main. PR #697 closed by hand (we merge locally, so Gitea does not close it itself).

What I checked myself before merging

Trial merge and build. Throwaway worktree of origin/main at 736fd9c, merged refs/pull/697/head (780cb34) with --no-ff. Full mvn clean install in fleetd/:

MVN_EXIT=0
BUILD SUCCESS
[INFO] Tests run: 1964, Failures: 0, Errors: 0, Skipped: 0

Counted independently from the surefire XML after rm -rf target/surefire-reports:

xml report files: 173     (173 .txt sit beside them — I counted the XML)
XML AGGREGATE: tests=1964 failures=0 errors=0 skipped=0

That matches the 1964 the worker reported, and the 1950 → 1964 step from origin/main.

Tree identity. The tree I built and the tree I merged are the same object: c8a0747465bd977dd369a78e024738d05fa13206, on both the throwaway merge (5a81d3d) and the real merge commit (f0ff252). So the green build covers exactly what landed. The merge commit does not touch the wiki/ submodule pointer.

Push verified by ref. Local HEAD and git ls-remote origin refs/heads/main both read f0ff25221e21f3db10fdd9091c447c7371932cd3.

The round-3 pin, re-measured

Round 3 existed because the fleet-wide fleet.tabLabel-vs-collaborator-tab check had no test. I did not take the worker's word for the new test being a real pin — I re-ran the mutation.

Line-anchored sed at FleetConfig.java:2761, which I first confirmed by printing its context is inside the fleet.collaborators().forEach block:

step result
mutate :2761 to if (false && templateCanRenderAs(fleet.tabLabel(), tab)) { printed the line back: the substitution applied
lead-side twin at :2733 printed unchanged
diff against the backup exactly one line changed (2761c2761)
mvn test -Dtest=FleetConfigTest under the mutation MVN_EXIT=1, BUILD FAILURE, 170 run, 1 failure
which test died aFleetTabLabelTemplateThatCanRenderAsACollaboratorTabRefusesToStart:1084 — one kill, no cascade
restore, git diff for the file empty
re-run restored MVN_EXIT=0, 170 run, 0 failures

So the branch is pinned, by exactly one test, and the test fails for the right reason (Expected java.lang.IllegalStateException to be thrown, but nothing was thrown).

One number corrected

The previous lead's handover recorded "176 tests green" for FleetConfigTest while mutating at ad593c9. I measure 170 in that class at 780cb34. These reconcile against the source rather than against each other: grep -c '@Test' gives 156 on origin/main and 170 on the PR, there are no @ParameterizedTest methods in the file, and only one file is named FleetConfigTest.java. Surefire ran 170. So 170 is the count of that class and 176 was never it. The previous lead's conclusion — that the branch was unpinned — was still correct, which is why this round happened.

Code read, not just built

templateCanRenderAs returns false for a null or blank tab, and both new collision loops skip a null/blank tab, so a collaborator with no tab falls through to validateMembers(), which refuses it. That is what the javadoc claims, and the three paths agree.

The PR body reported FLEET_POOL_KEYS as an unfixed gap; commit ad593c9 on the same branch then added collaborators to it, so that gap is closed in what merged.

Still open on this ticket

Units C, D, E and F are specified in the 20:58 comment and none is delegated. Unit C is now unblocked by this merge. Unit D is the security-critical one and must get a reviewer who did not write it. Unit F is the lead's and must still wait for the code: writing a fourth fleet_whoami value into the canonical block before COLLABORATOR exists would tell every session something false.

A redeploy is now owed — this merge changes Java production code, unlike 736fd9c.

## Unit B merged — PR #697, as `f0ff252` Merged locally and pushed to `main`. PR #697 closed by hand (we merge locally, so Gitea does not close it itself). ### What I checked myself before merging **Trial merge and build.** Throwaway worktree of `origin/main` at `736fd9c`, merged `refs/pull/697/head` (`780cb34`) with `--no-ff`. Full `mvn clean install` in `fleetd/`: ``` MVN_EXIT=0 BUILD SUCCESS [INFO] Tests run: 1964, Failures: 0, Errors: 0, Skipped: 0 ``` Counted independently from the surefire XML after `rm -rf target/surefire-reports`: ``` xml report files: 173 (173 .txt sit beside them — I counted the XML) XML AGGREGATE: tests=1964 failures=0 errors=0 skipped=0 ``` That matches the 1964 the worker reported, and the 1950 → 1964 step from `origin/main`. **Tree identity.** The tree I built and the tree I merged are the same object: `c8a0747465bd977dd369a78e024738d05fa13206`, on both the throwaway merge (`5a81d3d`) and the real merge commit (`f0ff252`). So the green build covers exactly what landed. The merge commit does not touch the `wiki/` submodule pointer. **Push verified by ref.** Local `HEAD` and `git ls-remote origin refs/heads/main` both read `f0ff25221e21f3db10fdd9091c447c7371932cd3`. ### The round-3 pin, re-measured Round 3 existed because the fleet-wide `fleet.tabLabel`-vs-collaborator-tab check had no test. I did not take the worker's word for the new test being a real pin — I re-ran the mutation. Line-anchored `sed` at `FleetConfig.java:2761`, which I first confirmed by printing its context is inside the `fleet.collaborators().forEach` block: | step | result | |---|---| | mutate `:2761` to `if (false && templateCanRenderAs(fleet.tabLabel(), tab)) {` | printed the line back: the substitution applied | | lead-side twin at `:2733` | printed unchanged | | `diff` against the backup | exactly one line changed (`2761c2761`) | | `mvn test -Dtest=FleetConfigTest` under the mutation | `MVN_EXIT=1`, BUILD FAILURE, **170 run, 1 failure** | | which test died | `aFleetTabLabelTemplateThatCanRenderAsACollaboratorTabRefusesToStart:1084` — one kill, no cascade | | restore, `git diff` for the file | empty | | re-run restored | `MVN_EXIT=0`, 170 run, 0 failures | So the branch is pinned, by exactly one test, and the test fails for the right reason (`Expected java.lang.IllegalStateException to be thrown, but nothing was thrown`). ### One number corrected The previous lead's handover recorded "176 tests green" for `FleetConfigTest` while mutating at `ad593c9`. I measure **170** in that class at `780cb34`. These reconcile against the source rather than against each other: `grep -c '@Test'` gives **156** on `origin/main` and **170** on the PR, there are no `@ParameterizedTest` methods in the file, and only one file is named `FleetConfigTest.java`. Surefire ran 170. So 170 is the count of that class and 176 was never it. The previous lead's *conclusion* — that the branch was unpinned — was still correct, which is why this round happened. ### Code read, not just built `templateCanRenderAs` returns `false` for a null or blank `tab`, and both new collision loops skip a null/blank tab, so a collaborator with no tab falls through to `validateMembers()`, which refuses it. That is what the javadoc claims, and the three paths agree. The PR body reported `FLEET_POOL_KEYS` as an unfixed gap; commit `ad593c9` on the same branch then added `collaborators` to it, so that gap is closed in what merged. ### Still open on this ticket Units C, D, E and F are specified in the 20:58 comment and none is delegated. Unit C is now unblocked by this merge. Unit D is the security-critical one and must get a reviewer who did not write it. Unit F is the lead's and must still wait for the code: writing a fourth `fleet_whoami` value into the canonical block before `COLLABORATOR` exists would tell every session something false. A redeploy is now owed — this merge changes Java production code, unlike `736fd9c`.
Author
Owner

Unit C correction — two fixes to my own brief. Both are my errors, not the implementer's.

Written here rather than sent as a revised brief, because the brief is write-once and a ticket comment is what a member can re-read whenever it looks. If this contradicts the Unit C brief, this comment is newer and it wins.

Unit C is in flight on term_65cf6e090dfe057, branch worker/669-1b786a-1.

Correction 1 — do NOT delete the 3-argument Authz.permits overload

My brief said to delete it "so the compiler forces both call sites to supply the classifier". I measured the cost after sending, and it is wrong.

grep -rc "Authz\.permits(" src/test/java/dev/ltms/fleet/auth/AuthzTest.java          # 37
grep -rc "Authz\.permits(" src/test/java/dev/ltms/fleet/auth/CallerResolverTest.java # 10
grep -n  "boolean permits\|private.*permits(" (both files)                           # no matches

So there are 47 direct 3-argument call sites and no test helper to change in one place. Deleting the overload forces 47 mechanical test edits.

What it buys does not justify that, because the default is already safe. The hazard I was guarding against is a future call site that skips the target limit. With a deny-all classifier as the 3-argument form's default, such a call site refuses a collaborator's send. That is the feature being inert — recoverable, debuggable — not a privilege hole. The expensive direction would be a permissive default, and nobody proposed one.

Do this instead:

  • Keep a 3-argument permits(caller, action, target) that delegates to the 4-argument form with the deny-all classifier. Fail-closed.
  • Both production gates still call the 4-argument form explicitly — FleetMcp.denyFor and FleetApp.allow. That was the real point of decision 5 and it stands.
  • Add a test pinning the fail-closed property of the overload itself: the 3-argument form refuses SEND for a collaborator. Otherwise the safe default is an accident rather than a decision, and nothing would notice if it flipped.
  • Leave the 47 existing call sites alone.

Acceptance criterion 2 (mutation-prove the classifier conjunct) is unchanged.

Correction 2 — a line reference I inherited instead of measuring

My brief said FleetMcp.java:437-440 "uses isSpawnedMember() to enrol a caller in the presence map". I took that range from this ticket's earlier architect comment and did not check it. I have now checked it.

grep -n "isSpawnedMember" src/main/java/dev/ltms/fleet/mcp/FleetMcp.java   # 782 only

The accurate picture: the context extractor calls markSpawnedMemberPresent(p, presence) at FleetMcp.java:441, and the role guard lives inside that method at FleetMcp.java:780-784:

static void markSpawnedMemberPresent(Principal caller, MemberPresence presence) {
    if (caller.isSpawnedMember()) {
        presence.markPresent(caller.terminal());
    }
}

Lines 437-440 are the explanatory comment above the call, not the guard.

The substance of decision 3 is unchanged and confirmed — isSpawnedMember() must stay WORKER || ARCHITECT, and the comment at :437-439 states the reason in the code's own words: "Enrolling a lead would count it as an available member in the roster." Only my citation was loose.

Why both of these are worth recording

This project has repeatedly found the defect in the lead's brief rather than the worker's code. Correction 1 is the general shape: I specified a mechanism (delete the overload) instead of the property I wanted (no call site can silently skip the limit). The property had a cheaper implementation that I would have found by asking what would falsify the need, and I sent the brief first. Correction 2 is the other standing rule, broken by me in the same message: a number or a line reference taken from an earlier comment is not measured, and repeating it does not make it mine.

Nothing else in the Unit C brief changes. Scope, the denied-action list, the TASK_READ denial, the whoami branch, and all eight acceptance criteria stand.

## Unit C correction — two fixes to my own brief. Both are my errors, not the implementer's. Written here rather than sent as a revised brief, because the brief is write-once and a ticket comment is what a member can re-read whenever it looks. **If this contradicts the Unit C brief, this comment is newer and it wins.** Unit C is in flight on `term_65cf6e090dfe057`, branch `worker/669-1b786a-1`. ### Correction 1 — do NOT delete the 3-argument `Authz.permits` overload My brief said to delete it "so the compiler forces both call sites to supply the classifier". I measured the cost after sending, and it is wrong. ``` grep -rc "Authz\.permits(" src/test/java/dev/ltms/fleet/auth/AuthzTest.java # 37 grep -rc "Authz\.permits(" src/test/java/dev/ltms/fleet/auth/CallerResolverTest.java # 10 grep -n "boolean permits\|private.*permits(" (both files) # no matches ``` So there are **47** direct 3-argument call sites and **no** test helper to change in one place. Deleting the overload forces 47 mechanical test edits. **What it buys does not justify that, because the default is already safe.** The hazard I was guarding against is a future call site that skips the target limit. With a deny-all classifier as the 3-argument form's default, such a call site refuses a collaborator's send. That is the feature being inert — recoverable, debuggable — not a privilege hole. The expensive direction would be a permissive default, and nobody proposed one. **Do this instead:** - Keep a 3-argument `permits(caller, action, target)` that delegates to the 4-argument form with the deny-all classifier. Fail-closed. - **Both production gates still call the 4-argument form explicitly** — `FleetMcp.denyFor` and `FleetApp.allow`. That was the real point of decision 5 and it stands. - **Add a test pinning the fail-closed property of the overload itself:** the 3-argument form refuses `SEND` for a collaborator. Otherwise the safe default is an accident rather than a decision, and nothing would notice if it flipped. - Leave the 47 existing call sites alone. Acceptance criterion 2 (mutation-prove the classifier conjunct) is unchanged. ### Correction 2 — a line reference I inherited instead of measuring My brief said `FleetMcp.java:437-440` "uses `isSpawnedMember()` to enrol a caller in the presence map". I took that range from this ticket's earlier architect comment and did not check it. I have now checked it. ``` grep -n "isSpawnedMember" src/main/java/dev/ltms/fleet/mcp/FleetMcp.java # 782 only ``` The accurate picture: the context extractor calls `markSpawnedMemberPresent(p, presence)` at **`FleetMcp.java:441`**, and the role guard lives inside that method at **`FleetMcp.java:780-784`**: ```java static void markSpawnedMemberPresent(Principal caller, MemberPresence presence) { if (caller.isSpawnedMember()) { presence.markPresent(caller.terminal()); } } ``` Lines 437-440 are the explanatory comment above the call, not the guard. **The substance of decision 3 is unchanged and confirmed** — `isSpawnedMember()` must stay `WORKER || ARCHITECT`, and the comment at `:437-439` states the reason in the code's own words: *"Enrolling a lead would count it as an available member in the roster."* Only my citation was loose. ### Why both of these are worth recording This project has repeatedly found the defect in the lead's brief rather than the worker's code. Correction 1 is the general shape: **I specified a mechanism (delete the overload) instead of the property I wanted (no call site can silently skip the limit).** The property had a cheaper implementation that I would have found by asking what would falsify the need, and I sent the brief first. Correction 2 is the other standing rule, broken by me in the same message: a number or a line reference taken from an earlier comment is not measured, and repeating it does not make it mine. Nothing else in the Unit C brief changes. Scope, the denied-action list, the `TASK_READ` denial, the `whoami` branch, and all eight acceptance criteria stand.
Author
Owner

Unit C merged — PR #699, as b92a669

Merged locally and pushed to main; PR #699 closed by hand. COLLABORATOR now exists with its authorization row, so Unit D is unblocked.

Verified before merging

Throwaway worktree of origin/main at f0ff252, --no-ff merge of refs/pull/699/head (bb29b00), full mvn clean install in fleetd/:

BUILD SUCCESS   MVN_EXIT=0
[INFO] Tests run: 1974, Failures: 0, Errors: 0, Skipped: 0

Independently from the surefire XML after rm -rf target/surefire-reports: 173 report files (I counted .xml), aggregate tests=1974 failures=0 errors=0. Baseline at f0ff252 was 1964, so +10 — exactly the new tests.

Tree identity: the tree I built and the tree I merged are the same object, a19e87be54c48937efa93748f1453f20939c30a4. Push verified by ref: local HEAD and git ls-remote origin refs/heads/main both read b92a669ddcf3e70f3a63ef5d0dc17ed42e07bc4f. The merge commit does not touch wiki/.

I read the grant table myself

Every arm is added to, never altered. SEND gains || (caller.isCollaborator() && knownLeadOrCollaborator.test(targetSession)); READ, METRICS gain || caller.isCollaborator(). SPAWN/STOP/DRAIN/HANDOVER, ANSWER, COORD_SEND, TASK_READ and COORD_READ keep their exact conditions, so a collaborator is excluded from each by construction. The caller == null || isAnonymous() pre-check is untouched. REPLY, ASK remain ownsSession(targetSession), and isSpawnedMember() remains WORKER || ARCHITECT.

FleetApp.allow now routes through the new permitsFor seam, so the REST test drives the real production gate rather than a test-supplied classifier. That matters here: a limit in a shared helper is only shared by the callers that call it.

Adjudication of the review fan-out

Two reviewers, one dimension each, neither the implementer.

Reviewer 1 (existing-role regression): no issue. It checked each arm against origin/main by reading, noted that Java short-circuits on isCollaborator() so the classifier is never invoked for a non-collaborator, and confirmed CallerResolver is untouched. It stated plainly that it did not re-derive the 156-pair matrix and did not run mvn. I ran the build myself and read the same arms, and I agree.

Reviewer 2 (collaborator privilege set): reported high severity at CallerResolver.java:228 — a configured collaborator tab falls through to Principal.worker, so it receives TASK_READ and can read other members' ticket replies.

I checked this myself, and I am not treating it as a merge blocker. The finding is factually right and the severity label is wrong for this PR.

sed -n '200,240p' CallerResolver.java   # resolve() ends: return Principal.worker(c.terminal(), c.pid());
git diff --stat origin/main...refs/pull/699/head -- .../auth/CallerResolver.java   # empty

The worker fallback catches any pane that is not a configured lead tab and not a bound architect slot. That predates Unit C and predates Unit B, and this PR does not touch the file. So listing a tab under fleet.collaborators grants it nothing new — it was already a worker-by-fallback like every other hand-opened pane. Unit C could not have fixed it without doing Unit D, and this ticket's plan already states Unit D's acceptance criterion as exactly this: the resolver must return a member's role from the roster first, then lead, then collaborator, then the worker fallback.

What the reviewer did surface, and what I had understated, is a documentation hazard: an operator could read the registry as a sandbox. I have fixed that on wiki/11-Features.md (d913009, pushed to refs/heads/main and verified by ref), which now says that naming a tab here buys no restriction, that such a tab resolves as a worker and so holds TASK_READ, and that the fallback is pre-existing rather than granted by the registry.

Shape-sweep finding, carried forward (not fixed)

FleetMcp.denyFor skips the audit allowed entry for READ/TASK_READ; FleetApp.allow skips READ/METRICS/TASK_READ. The implementer added the fact that settles the severity: grep -n "Action.METRICS" FleetMcp.java returns nothing, so no MCP tool ever passes METRICS to denyFor and the asymmetry has no live effect today. It remains one invariant enforced with two hand-maintained lists. Not filed as a ticket yet.

Where #669 stands

  • Units A, B, C: merged. main is b92a669.
  • Unit D is next and unblocked. It is the security-critical one and must get a reviewer who did not write it. Its acceptance criterion already covers reviewer 2's finding above, including the control that removing the spawned-member-first step must make the assertion fail.
  • Unit E (widen the isLead predicate to lead-or-collaborator) is latent on this host: no memberHerdrSocket is configured.
  • Unit F is the lead's and is now partly unblocked — COLLABORATOR exists and fleet_whoami returns it, so the canonical block's fleet_whoami ladder and invariant 3's "send is lead or architect" line can finally be written truthfully. I have not done it in this session. Note the role is still unreachable in production until Unit D, so the block should say what fleet_whoami can return, not imply a collaborator is live.

A redeploy is owed again: this merge changes Java production code.

## Unit C merged — PR #699, as `b92a669` Merged locally and pushed to `main`; PR #699 closed by hand. `COLLABORATOR` now exists with its authorization row, so **Unit D is unblocked**. ### Verified before merging Throwaway worktree of `origin/main` at `f0ff252`, `--no-ff` merge of `refs/pull/699/head` (`bb29b00`), full `mvn clean install` in `fleetd/`: ``` BUILD SUCCESS MVN_EXIT=0 [INFO] Tests run: 1974, Failures: 0, Errors: 0, Skipped: 0 ``` Independently from the surefire XML after `rm -rf target/surefire-reports`: **173** report files (I counted `.xml`), aggregate `tests=1974 failures=0 errors=0`. Baseline at `f0ff252` was 1964, so +10 — exactly the new tests. **Tree identity:** the tree I built and the tree I merged are the same object, `a19e87be54c48937efa93748f1453f20939c30a4`. Push verified by ref: local `HEAD` and `git ls-remote origin refs/heads/main` both read `b92a669ddcf3e70f3a63ef5d0dc17ed42e07bc4f`. The merge commit does not touch `wiki/`. ### I read the grant table myself Every arm is **added to**, never altered. `SEND` gains `|| (caller.isCollaborator() && knownLeadOrCollaborator.test(targetSession))`; `READ, METRICS` gain `|| caller.isCollaborator()`. `SPAWN/STOP/DRAIN/HANDOVER`, `ANSWER`, `COORD_SEND`, `TASK_READ` and `COORD_READ` keep their exact conditions, so a collaborator is excluded from each by construction. The `caller == null || isAnonymous()` pre-check is untouched. `REPLY, ASK` remain `ownsSession(targetSession)`, and `isSpawnedMember()` remains `WORKER || ARCHITECT`. `FleetApp.allow` now routes through the new `permitsFor` seam, so the REST test drives the real production gate rather than a test-supplied classifier. That matters here: a limit in a shared helper is only shared by the callers that call it. ### Adjudication of the review fan-out Two reviewers, one dimension each, neither the implementer. **Reviewer 1 (existing-role regression): no issue.** It checked each arm against `origin/main` by reading, noted that Java short-circuits on `isCollaborator()` so the classifier is never invoked for a non-collaborator, and confirmed `CallerResolver` is untouched. It stated plainly that it did **not** re-derive the 156-pair matrix and did **not** run `mvn`. I ran the build myself and read the same arms, and I agree. **Reviewer 2 (collaborator privilege set): reported high severity at `CallerResolver.java:228` — a configured collaborator tab falls through to `Principal.worker`, so it receives `TASK_READ` and can read other members' ticket replies.** **I checked this myself, and I am not treating it as a merge blocker. The finding is factually right and the severity label is wrong for this PR.** ``` sed -n '200,240p' CallerResolver.java # resolve() ends: return Principal.worker(c.terminal(), c.pid()); git diff --stat origin/main...refs/pull/699/head -- .../auth/CallerResolver.java # empty ``` The worker fallback catches **any** pane that is not a configured lead tab and not a bound architect slot. That predates Unit C and predates Unit B, and this PR does not touch the file. So listing a tab under `fleet.collaborators` grants it nothing new — it was already a worker-by-fallback like every other hand-opened pane. Unit C could not have fixed it without doing Unit D, and this ticket's plan already states Unit D's acceptance criterion as exactly this: the resolver must return a member's role from the roster first, then lead, then collaborator, then the worker fallback. What the reviewer did surface, and what I had understated, is a **documentation** hazard: an operator could read the registry as a sandbox. I have fixed that on `wiki/11-Features.md` (`d913009`, pushed to `refs/heads/main` and verified by ref), which now says that naming a tab here buys no restriction, that such a tab resolves as a worker and so holds `TASK_READ`, and that the fallback is pre-existing rather than granted by the registry. ### Shape-sweep finding, carried forward (not fixed) `FleetMcp.denyFor` skips the audit `allowed` entry for `READ`/`TASK_READ`; `FleetApp.allow` skips `READ`/`METRICS`/`TASK_READ`. The implementer added the fact that settles the severity: `grep -n "Action.METRICS" FleetMcp.java` returns nothing, so no MCP tool ever passes `METRICS` to `denyFor` and the asymmetry has no live effect today. It remains one invariant enforced with two hand-maintained lists. Not filed as a ticket yet. ### Where #669 stands - **Units A, B, C: merged.** `main` is `b92a669`. - **Unit D is next and unblocked.** It is the security-critical one and must get a reviewer who did not write it. Its acceptance criterion already covers reviewer 2's finding above, including the control that removing the spawned-member-first step must make the assertion fail. - **Unit E** (widen the `isLead` predicate to lead-or-collaborator) is latent on this host: no `memberHerdrSocket` is configured. - **Unit F is the lead's** and is now **partly unblocked** — `COLLABORATOR` exists and `fleet_whoami` returns it, so the canonical block's `fleet_whoami` ladder and invariant 3's "send is lead or architect" line can finally be written truthfully. I have not done it in this session. Note the role is still unreachable in production until Unit D, so the block should say what `fleet_whoami` *can* return, not imply a collaborator is live. A redeploy is owed again: this merge changes Java production code.
Author
Owner

Unit D spec — the resolver. Delegated now.

Implementer: dev on sonnet, worktree 669-unit-d, branch worker/669-unit-d-efbbd7-1.
This comment is the authoritative spec. The brief sent to the worker carries the same text.
Any correction goes here as a new comment, never as an edit to this one and never as a re-send.

Everything below I measured myself in the main clone at b92a669, with the command next to it.
Units A, B and C are merged, so this is the live shape, not the shape the plan assumed.

The security property, stated as a property

A terminal that belongs to a live spawned member resolves as that member's own role, whatever
any tab map says about that terminal.
When the spawned-member roster matches, no tab map is
consulted at all.

That is #661 closed at the resolver rather than only at config validation. Today it is open:
resolve() reads the lead tab map first, at CallerResolver.java:211.

The current resolution order — read by me, CallerResolver.java:208-257

c.terminal() != null
  211  lead tab map      -> Principal.leader
  219  architect slot    -> Principal.architect   (slot role confirmed via memberSlotRoles)
  228  fallback          -> Principal.worker
c.terminal() == null
  231  token mode        -> primary | anonymous
  255  loopback-trust    -> primary | anonymous   (needs c.resolved() && c.scanComplete())

The order Unit D must produce

  1. c.terminal() == null -> the token / loopback-trust tail, unchanged.
  2. terminal in the spawned-member roster -> that member's role. Consult no tab map.
    MemberRole.ARCHITECT -> Role.ARCHITECT; DEV, HUNTER, REVIEWER -> Role.WORKER.
  3. lead tab map -> Principal.leader (unchanged behaviour).
  4. architect slot binding -> Principal.architect (unchanged behaviour).
  5. collaborator tab map -> Principal.collaborator(name, terminal, pid).
  6. worker fallback -> Principal.worker (unchanged).

Step 4 stays where it is. A spawned architect is already caught by step 2, so step 4 now serves
the case it was written for: a binding with no live MemberSession.

The roster source

SessionManager.roster() returns List<MemberSession>; each carries terminalId() and
role() (a MemberRole). Its javadoc at SessionManager.java:780-790 says it is
"all registered sessions (acquired minus released)", and it is the non-resolving roster meant for
hot paths — which is what resolve() is. Do not use rosterResolved(): its own javadoc says
it is for caller-driven reads, and it can open an opencode session store per call.

resolve() runs on every request, so pass the lookup in as a function rather than a list to scan.

MemberRegistry.bind is called from SessionManager.java:740, so slot bindings are live. I checked
that myself; an older note of mine claiming bind is never called is wrong.

The scan: generalise the one that exists, do not add a second

LeadTabScanner (src/main/java/dev/ltms/fleet/herdr/LeadTabScanner.java, 256 lines) is handed a
tab -> name map and returns terminal -> name. It must return terminal -> (name, kind) with
kind one of lead or collaborator, from one pass.

Every property it has today must survive, because each one is a fix for a named bug:

  • exact match, case-insensitive, ends stripped (CB-579) — no prefix stripping
  • PendingCloseMarker.strip before matching
  • the agent.list liveness cross-check, so a labelled but dead tab is never in the map (#359)
  • the one grace scan for a terminal already reported live (#359 review finding 2)
  • the TTL cache, and a failed scan keeping the previous answer rather than emptying it

These apply to a collaborator tab exactly as to a lead tab. A dead collaborator tab must not
resolve.

An edge case the assembly has today: FleetdAssembly.java:265 builds the scanner only when
fleet.leaders is non-empty, and reads the TTL from
leaders.values().iterator().next().scanIntervalSeconds(). Collaborator is
record Collaborator(String tab) — it has no scanIntervalSeconds. So collaborators configured
with no leaders must still produce a scanner. Pick a defensible interval for that case and say in
your reply which one and why.

The classifier that is still inert

Authz.NO_KNOWN_LEAD_OR_COLLABORATOR is target -> false (Authz.java:71) and both production
gates pass it — FleetMcp.java:692, and rest/FleetApp.java:276-277 through its
permitsFor seam. So a collaborator's SEND is denied for every target today, by design, until a
real classifier lands.

Unit D wires it: a target terminal is a known lead or collaborator when it is in the lead map or
the collaborator map. It must be the same maps resolve() reads, not a second copy — a roster
that lists an address resolve() would not accept is the drift CallerResolver.leads()' javadoc
already exists to prevent.

Acceptance criteria

  1. The ordering assertion. A terminal present in both the spawned-member roster and the
    lead tab map resolves as its member role, not as a lead.
  2. Its control, and this one is required. Remove the spawned-member step and assertion 1 must
    fail. Run that mutation, and report the exact test name that goes red plus anything else
    that goes red with it. A survivor means the assertion is not pinning what it claims.
  3. A configured collaborator tab that is not a spawned member resolves Role.COLLABORATOR and
    carries that collaborator's name.
  4. A dead collaborator tab — labelled, no agent in agent.list — does not resolve as a
    collaborator.
  5. The regression control. Every resolution outcome for a caller that is neither a spawned
    member nor a collaborator is unchanged: a lead still resolves PRIMARY with its name, a bound
    architect still resolves ARCHITECT, an unrecognised pane still resolves WORKER, and both
    terminal == null tails are untouched.
  6. With the real classifier wired, a collaborator's local send naming a configured lead or
    collaborator is permitted, and the same call naming a spawned member's terminal is refused —
    over both MCP and POST /sessions/{id}/message. A test covering only MCP leaves the REST
    route open.
  7. mvn clean install in fleetd/ passes. Report the real Tests run: line.
    Establish your own baseline before you change anything: build once on untouched
    origin/main and report that number too, so your "+N new tests" has a denominator you measured.
    For reference only, the Unit C evidence comment on this ticket reports 1974 tests / 0 failures
    at b92a669; I have not re-run that myself in this session, so treat it as a cross-check on
    your own baseline, not as the baseline.

Hazards measured on this host

  • There is no POM at the repo root. Maven must run in fleetd/.
  • A trailing command in a compound line steals the exit code. mvn … ; echo "exit=$?" ; tail
    reports tail's status. Write MVN_EXIT=$? into the log and grep the log for it.
  • grep -c returning 0 exits 1, which silently drops the rest of an && chain. Append
    || true.
  • find here is bfs and rejects -newermt with a relative time. Use -mmin -N, and never
    send a probe's stderr to /dev/null when a zero is the answer you would act on.
  • fleetd/fleetd.yaml is gitignored and you cannot read it. No acceptance criterion above
    depends on it. For the record, I checked it myself: no memberHerdrSocket is configured and no
    profile uses placement: pane.
  • Your worktree's .mcp.json, opencode.json and .autoenv are stubs, not the repo's files.
    Never edit them. git config --worktree --get-all fleet.neutralizedConfig lists them.
  • wiki/ is uninitialized in your worktree. Do not read it and do not run the CLAUDE.md/wiki
    sync check — it cannot pass for you. That check is the lead's.
  • refs/stash is shared across worktrees. Never git stash.

Working rules

  • Stage files explicitly. Never git add -A.
  • If a command you were told to run is blocked, say so in your reply. Do not substitute a different
    command to get around it.
  • Comments describe the code as it is now: no history, no dates, no ticket numbers as rationale, no
    "this used to". Contracts go in javadoc, constraints inline, the story in the commit message.
  • Re-read this ticket before you act on anything and again before you commit. A later comment here
    is newer than your brief and wins.
  • Report the shape, do not chase it: if you find another place with this same
    tab-map-before-roster shape, name it in one line and do not fix it.
  • End your turn with exactly one fleet_reply carrying your whole report, including the PR URL.
## Unit D spec — the resolver. Delegated now. Implementer: `dev` on `sonnet`, worktree `669-unit-d`, branch `worker/669-unit-d-efbbd7-1`. This comment is the **authoritative spec**. The brief sent to the worker carries the same text. Any correction goes here as a new comment, never as an edit to this one and never as a re-send. Everything below I measured myself in the main clone at `b92a669`, with the command next to it. Units A, B and C are merged, so this is the live shape, not the shape the plan assumed. ### The security property, stated as a property **A terminal that belongs to a live spawned member resolves as that member's own role, whatever any tab map says about that terminal.** When the spawned-member roster matches, no tab map is consulted at all. That is #661 closed at the resolver rather than only at config validation. Today it is open: `resolve()` reads the lead tab map first, at `CallerResolver.java:211`. ### The current resolution order — read by me, `CallerResolver.java:208-257` ``` c.terminal() != null 211 lead tab map -> Principal.leader 219 architect slot -> Principal.architect (slot role confirmed via memberSlotRoles) 228 fallback -> Principal.worker c.terminal() == null 231 token mode -> primary | anonymous 255 loopback-trust -> primary | anonymous (needs c.resolved() && c.scanComplete()) ``` ### The order Unit D must produce 1. `c.terminal() == null` -> the token / loopback-trust tail, **unchanged**. 2. terminal in the **spawned-member roster** -> that member's role. Consult no tab map. `MemberRole.ARCHITECT` -> `Role.ARCHITECT`; `DEV`, `HUNTER`, `REVIEWER` -> `Role.WORKER`. 3. lead tab map -> `Principal.leader` (unchanged behaviour). 4. architect slot binding -> `Principal.architect` (unchanged behaviour). 5. collaborator tab map -> `Principal.collaborator(name, terminal, pid)`. 6. worker fallback -> `Principal.worker` (unchanged). Step 4 stays where it is. A spawned architect is already caught by step 2, so step 4 now serves the case it was written for: a binding with no live `MemberSession`. ### The roster source `SessionManager.roster()` returns `List<MemberSession>`; each carries `terminalId()` and `role()` (a `MemberRole`). Its javadoc at `SessionManager.java:780-790` says it is "all registered sessions (acquired minus released)", and it is the non-resolving roster meant for hot paths — which is what `resolve()` is. Do **not** use `rosterResolved()`: its own javadoc says it is for caller-driven reads, and it can open an opencode session store per call. `resolve()` runs on every request, so pass the lookup in as a function rather than a list to scan. `MemberRegistry.bind` is called from `SessionManager.java:740`, so slot bindings are live. I checked that myself; an older note of mine claiming `bind` is never called is wrong. ### The scan: generalise the one that exists, do not add a second `LeadTabScanner` (`src/main/java/dev/ltms/fleet/herdr/LeadTabScanner.java`, 256 lines) is handed a `tab -> name` map and returns `terminal -> name`. It must return `terminal -> (name, kind)` with `kind` one of lead or collaborator, from **one** pass. Every property it has today must survive, because each one is a fix for a named bug: - exact match, case-insensitive, ends stripped (CB-579) — no prefix stripping - `PendingCloseMarker.strip` before matching - the `agent.list` liveness cross-check, so a labelled but dead tab is never in the map (#359) - the one grace scan for a terminal already reported live (#359 review finding 2) - the TTL cache, and a failed scan keeping the previous answer rather than emptying it **These apply to a collaborator tab exactly as to a lead tab.** A dead collaborator tab must not resolve. An edge case the assembly has today: `FleetdAssembly.java:265` builds the scanner only when `fleet.leaders` is non-empty, and reads the TTL from `leaders.values().iterator().next().scanIntervalSeconds()`. `Collaborator` is `record Collaborator(String tab)` — **it has no `scanIntervalSeconds`**. So collaborators configured with no leaders must still produce a scanner. Pick a defensible interval for that case and say in your reply which one and why. ### The classifier that is still inert `Authz.NO_KNOWN_LEAD_OR_COLLABORATOR` is `target -> false` (`Authz.java:71`) and both production gates pass it — `FleetMcp.java:692`, and `rest/FleetApp.java:276-277` through its `permitsFor` seam. So a collaborator's `SEND` is denied for every target today, by design, until a real classifier lands. Unit D wires it: a target terminal is a known lead or collaborator when it is in the lead map or the collaborator map. It must be the **same** maps `resolve()` reads, not a second copy — a roster that lists an address `resolve()` would not accept is the drift `CallerResolver.leads()`' javadoc already exists to prevent. ### Acceptance criteria 1. **The ordering assertion.** A terminal present in **both** the spawned-member roster and the lead tab map resolves as its member role, not as a lead. 2. **Its control, and this one is required.** Remove the spawned-member step and assertion 1 must **fail**. Run that mutation, and report the exact test name that goes red plus anything else that goes red with it. A survivor means the assertion is not pinning what it claims. 3. A configured collaborator tab that is not a spawned member resolves `Role.COLLABORATOR` and carries that collaborator's name. 4. A dead collaborator tab — labelled, no agent in `agent.list` — does not resolve as a collaborator. 5. **The regression control.** Every resolution outcome for a caller that is neither a spawned member nor a collaborator is unchanged: a lead still resolves `PRIMARY` with its name, a bound architect still resolves `ARCHITECT`, an unrecognised pane still resolves `WORKER`, and both `terminal == null` tails are untouched. 6. With the real classifier wired, a collaborator's local send naming a configured lead or collaborator is permitted, and the same call naming a spawned member's terminal is refused — over **both** MCP and `POST /sessions/{id}/message`. A test covering only MCP leaves the REST route open. 7. `mvn clean install` in `fleetd/` passes. Report the real `Tests run:` line. **Establish your own baseline before you change anything**: build once on untouched `origin/main` and report that number too, so your "+N new tests" has a denominator you measured. For reference only, the Unit C evidence comment on this ticket reports 1974 tests / 0 failures at `b92a669`; **I have not re-run that myself in this session**, so treat it as a cross-check on your own baseline, not as the baseline. ### Hazards measured on this host - **There is no POM at the repo root.** Maven must run in `fleetd/`. - **A trailing command in a compound line steals the exit code.** `mvn … ; echo "exit=$?" ; tail` reports `tail`'s status. Write `MVN_EXIT=$?` into the log and grep the log for it. - **`grep -c` returning 0 exits 1**, which silently drops the rest of an `&&` chain. Append `|| true`. - **`find` here is `bfs` and rejects `-newermt` with a relative time.** Use `-mmin -N`, and never send a probe's stderr to `/dev/null` when a zero is the answer you would act on. - **`fleetd/fleetd.yaml` is gitignored and you cannot read it.** No acceptance criterion above depends on it. For the record, I checked it myself: no `memberHerdrSocket` is configured and no profile uses `placement: pane`. - **Your worktree's `.mcp.json`, `opencode.json` and `.autoenv` are stubs, not the repo's files.** Never edit them. `git config --worktree --get-all fleet.neutralizedConfig` lists them. - **`wiki/` is uninitialized in your worktree.** Do not read it and do not run the `CLAUDE.md`/wiki sync check — it cannot pass for you. That check is the lead's. - **`refs/stash` is shared across worktrees.** Never `git stash`. ### Working rules - Stage files explicitly. Never `git add -A`. - If a command you were told to run is blocked, say so in your reply. Do not substitute a different command to get around it. - Comments describe the code as it is now: no history, no dates, no ticket numbers as rationale, no "this used to". Contracts go in javadoc, constraints inline, the story in the commit message. - Re-read this ticket before you act on anything and again before you commit. A later comment here is newer than your brief and wins. - Report the shape, do not chase it: if you find another place with this same tab-map-before-roster shape, name it in one line and **do not fix it**. - End your turn with exactly one `fleet_reply` carrying your whole report, including the PR URL.
Author
Owner

The audit-skip asymmetry the Unit C implementer found in its shape sweep is now filed as #700, so it stops being carried forward in handover prose.

I re-measured it myself rather than repeating the implementer's numbers. Both facts hold at b92a669: FleetMcp.denyFor skips the allowed entry for READ/TASK_READ (FleetMcp.java:693), rest/FleetApp.allow skips READ/METRICS/TASK_READ (FleetApp.java:294), and Action.METRICS appears 0 times in FleetMcp.

I paired that zero with a positive control, because a zero reads the same whether the action is absent or the pattern is broken: in the same file Action.READ is 2, Action.TASK_READ is 4, Action.SEND is 2. So the pattern matches when there is something to match, and METRICS really is absent.

It is out of scope for Unit D and the Unit D brief says so — the implementer is told to name a shape like this in one line and not fix it.

The audit-skip asymmetry the Unit C implementer found in its shape sweep is now filed as **#700**, so it stops being carried forward in handover prose. I re-measured it myself rather than repeating the implementer's numbers. Both facts hold at `b92a669`: `FleetMcp.denyFor` skips the `allowed` entry for `READ`/`TASK_READ` (`FleetMcp.java:693`), `rest/FleetApp.allow` skips `READ`/`METRICS`/`TASK_READ` (`FleetApp.java:294`), and `Action.METRICS` appears **0** times in `FleetMcp`. I paired that zero with a positive control, because a zero reads the same whether the action is absent or the pattern is broken: in the same file `Action.READ` is 2, `Action.TASK_READ` is 4, `Action.SEND` is 2. So the pattern matches when there is something to match, and `METRICS` really is absent. It is out of scope for Unit D and the Unit D brief says so — the implementer is told to name a shape like this in one line and not fix it.
Author
Owner

CORRECTION for Unit D / PR #701 — do not merge as it stands. It regresses fleetd #424.

This is newer than the brief and newer than comment 18450, so it wins. One change is needed; everything else in the PR stands.

I found this by reading the diff myself, then proved it with a throwaway test rather than leaving it as a reading. The build is not the problem — I ran the full trial merge myself and it is green: BUILD SUCCESS, MVN_EXIT=0, Tests run: 1988, Failures: 0, Errors: 0, independently cross-checked by aggregating 173 surefire XML files to the same 1988/0/0. The defect is invisible to the suite.

What breaks

fleetd #424 has two halves. One is "a revoked slot refuses the next spawn". The other is "a session already bound to a slot loses the ARCHITECT privilege on its very next request". MemberRegistry's class javadoc states the second half as the rule — "config governs what a bound slot still grants, as well as what may be bound next" — and explains the mechanism: roleForSlot and nameForSlot read slots() with no cache, so once the slot drops out of config, resolve() can no longer confirm it and the pane falls through to Principal.worker(...).

The new spawned-member step returns Principal.architect(...) on the strength of the roster role alone. It never calls memberSlotRoles. Since a spawned architect always has a MemberSession with role ARCHITECT, that step now catches every live architect and the config confirm is skipped.

The measurement, with its control

I built the post-revocation state honestly: reserve and bind while the slot is configured, then swap the config to remove it, exactly as a reload does. Two controls assert the state is really what I claim — roleForSlot returns null after revocation, and the occupancy is deliberately still there.

Then I resolved the same terminal through both construction paths, same registry, same moment:

PROBE roster-aware role=ARCHITECT name=null
PROBE old-path     role=WORKER    name=null
CONTROL: the pre-existing path revokes the privilege (fleetd #424)   -> PASSED
a revoked slot must still demote a LIVE spawned architect            -> FAILED
expected: <WORKER> but was: <ARCHITECT>

The passing control is the point. The only difference between the two lines is which constructor built the resolver, so the demotion is lost by this change and not by my setup. Note also name=null: nameForSlot correctly returns null because the slot is gone, and the ARCHITECT role is granted anyway — so the result is an architect principal with no name.

An earlier attempt of mine failed its own setup control (bind refuses an unconfigured slot, so I could not reach the revoked state that way). I mention it because that control is what stopped me reporting a conclusion from a broken probe.

Why the suite stays green

The privilege-revocation half of #424 was documented but never tested. MemberRegistryLiveTest pins only requireSlotFor and reserve — the next-spawn half. So there was no assertion to go red. A comment claiming an invariant is a free test case, and this one was never cashed in.

The fix

On the new spawned-member step, confirm the slot the same way the existing architect step does: treat a roster role of ARCHITECT as ARCHITECT only when memberSlotRoles still confirms the bound slot is an architect slot, and otherwise fall through to Principal.worker. The roster decides that the pane is a live spawned member; config still decides what that member's slot grants. Those are two questions and the fix keeps them separate.

Then pin it, because nothing else does: a test asserting that a live spawned architect whose slot was revoked resolves WORKER, with the passing control above kept in the test so it cannot go vacuous.

Do not widen this into a redesign. If anyone thinks a live spawned architect should outrank a config revocation, that is a deliberate change to #424's rule and it needs its own ticket and its own argument — not a silent side effect of the resolver reorder.

Everything else about the PR is fine so far

The resolution order is as specified, the collaborator step is last among the tab maps, knownLeadOrCollaborator() reads the same live maps resolve() reads, and both production gates now take the real classifier. Two reviewers are still working other dimensions; I will fold in anything they find.

## CORRECTION for Unit D / PR #701 — do not merge as it stands. It regresses fleetd #424. This is newer than the brief and newer than comment 18450, so it wins. One change is needed; everything else in the PR stands. I found this by reading the diff myself, then proved it with a throwaway test rather than leaving it as a reading. **The build is not the problem** — I ran the full trial merge myself and it is green: `BUILD SUCCESS`, `MVN_EXIT=0`, `Tests run: 1988, Failures: 0, Errors: 0`, independently cross-checked by aggregating 173 surefire XML files to the same 1988/0/0. The defect is invisible to the suite. ### What breaks fleetd #424 has two halves. One is "a revoked slot refuses the **next** spawn". The other is "a session **already bound** to a slot loses the ARCHITECT privilege on its very next request". `MemberRegistry`'s class javadoc states the second half as the rule — *"config governs what a bound slot still grants, as well as what may be bound next"* — and explains the mechanism: `roleForSlot` and `nameForSlot` read `slots()` with no cache, so once the slot drops out of config, `resolve()` can no longer confirm it and the pane falls through to `Principal.worker(...)`. The new spawned-member step returns `Principal.architect(...)` on the strength of the **roster** role alone. It never calls `memberSlotRoles`. Since a spawned architect always has a `MemberSession` with role `ARCHITECT`, that step now catches every live architect and the config confirm is skipped. ### The measurement, with its control I built the post-revocation state honestly: reserve and bind **while** the slot is configured, then swap the config to remove it, exactly as a reload does. Two controls assert the state is really what I claim — `roleForSlot` returns `null` after revocation, and the occupancy is deliberately still there. Then I resolved the same terminal through both construction paths, same registry, same moment: ``` PROBE roster-aware role=ARCHITECT name=null PROBE old-path role=WORKER name=null ``` ``` CONTROL: the pre-existing path revokes the privilege (fleetd #424) -> PASSED a revoked slot must still demote a LIVE spawned architect -> FAILED expected: <WORKER> but was: <ARCHITECT> ``` The passing control is the point. The only difference between the two lines is which constructor built the resolver, so the demotion is lost by this change and not by my setup. Note also `name=null`: `nameForSlot` correctly returns null because the slot is gone, and the ARCHITECT role is granted anyway — so the result is an architect principal with no name. An earlier attempt of mine failed its own setup control (`bind` refuses an unconfigured slot, so I could not reach the revoked state that way). I mention it because that control is what stopped me reporting a conclusion from a broken probe. ### Why the suite stays green The privilege-revocation half of #424 was **documented but never tested**. `MemberRegistryLiveTest` pins only `requireSlotFor` and `reserve` — the next-spawn half. So there was no assertion to go red. A comment claiming an invariant is a free test case, and this one was never cashed in. ### The fix On the new spawned-member step, confirm the slot the same way the existing architect step does: treat a roster role of `ARCHITECT` as `ARCHITECT` **only when `memberSlotRoles` still confirms the bound slot is an architect slot**, and otherwise fall through to `Principal.worker`. The roster decides *that the pane is a live spawned member*; config still decides *what that member's slot grants*. Those are two questions and the fix keeps them separate. Then pin it, because nothing else does: a test asserting that a live spawned architect whose slot was revoked resolves `WORKER`, with the passing control above kept in the test so it cannot go vacuous. **Do not widen this into a redesign.** If anyone thinks a live spawned architect *should* outrank a config revocation, that is a deliberate change to #424's rule and it needs its own ticket and its own argument — not a silent side effect of the resolver reorder. ### Everything else about the PR is fine so far The resolution order is as specified, the collaborator step is last among the tab maps, `knownLeadOrCollaborator()` reads the same live maps `resolve()` reads, and both production gates now take the real classifier. Two reviewers are still working other dimensions; I will fold in anything they find.
Author
Owner

Unit D review fan-out adjudicated. One blocker, one nit, one new ticket.

Two reviewers, one dimension each, neither the implementer. Both reported honestly, including what they did not check. I verified every load-bearing claim myself rather than promoting a reviewer's "sound" to a fact.

The blocker stays as stated in comment 18462

The fleetd #424 privilege-revocation regression is the only thing stopping this merge. Neither reviewer found it — I found it by reading the diff. The fix and the test to pin it are in 18462.

Reviewer 1 (resolver wiring + roster freshness): one medium finding, and it is real

I verified its wiring answer. Production constructs the resolver only through the new roster-aware overload — FleetdAssembly.java:495 and :500 in the PR head, both passing spawnedMemberRole and collaboratorTerminals. git grep withLeadsAndMembers over fleetd/src/main/java on the PR head returns no other call site, so the 5-arg overload has no production caller, and NO_KNOWN_LEAD_OR_COLLABORATOR survives only as Authz's own 3-arg fail-closed default. That closes the inertness question I set it: the security step is wired, not merely present. This mattered because a defaulted no-op would have made the whole unit inert while every test passed.

Its finding: a release-window race. release() does registry.remove(paneId) first (SessionManager.java:306) and stops the pane only afterwards. In between the code calls worktrees.hasUncommitted(...), which shells out to git status — the CB-581 comment at :363 says so in the code's own words — and may then call trySnapshot(...), another git operation. So a member is deregistered while still alive, for as long as one or two git subprocesses take. During that window spawnedMemberRole returns null for its terminal and the caller falls through to the tab maps.

I confirmed the sequence and the intervening work myself. The finding is structurally right and the window is not theoretical.

I am not holding the merge for it, and here is why. It is not a regression. Before this PR there was no roster step at all, so every member fell through to the tab maps on every request — the hole was total. This PR narrows it to the release window. Blocking a strict improvement because it is not yet total would be the wrong trade. The fix also belongs in a different subsystem: it needs a "releasing" marker surviving until launcher.stop returns, which is SessionManager lifecycle work, not resolver work.

Its exploit precondition also does not hold on this host: it needs a member pane carrying a lead or collaborator tab label, which startup validation refuses (validatePanePlacementAgainstLeadTabs, plus Unit B's widened collision check), and I checked the live config myself — no profile uses placement: pane. That makes it a latent hazard here and a live one on any host that ever uses pane placement.

Filed as its own ticket. Credited to reviewer 1.

Reviewer 2 (scanner generalisation): no issue, and I checked its reasoning

It reported NO ISSUE with a stated scope, which is a useful answer. Its central claim is correct and I verified it in the diff: cached is a single combined Map<String, Entry>, both get() and collaborators() go through one refresh(), and cached/gracedTerminals are keyed by terminal, not by kind — so accessor order cannot consume or renew one kind's grace state. byKind only filters the combined map at read time.

I checked all five properties the class carries, since each is a fix for a real bug, and all five survive: exact case-insensitive match with no prefix stripping, PendingCloseMarker.strip before matching (entryOf, PR head :306-311), the agent.list liveness cross-check now covering both kinds, the one grace scan, and the TTL cache whose failed scan keeps the previous answer.

It also correctly noted a coverage gap rather than a defect: the added tests do not cover cross-TTL accessor ordering. Worth knowing; not worth blocking.

One nit to fix alongside the blocker

buildTabIndex's javadoc says a colliding label "takes the lead entry, the first put when a key collides". The outcome is right but the stated reason is backwards: the code puts collaborators first and leads second, and LinkedHashMap.put overwrites, so lead wins because it is the last put. A future editor who believes "first put wins" could swap those two calls to tidy them and silently invert the precedence. Say that lead is put last and therefore wins, or make the precedence explicit instead of relying on put order.

(Config validation already refuses a lead and a collaborator sharing an exact tab, so this is defence in depth — which is fine, and is exactly why the comment should be right.)

Noted, not in scope, not fixed

ConfigRef.java:149 and :621 still point at CallerResolver.withLeadsAndMembers as being "at Fleetd.java:620/624". That wiring lives in FleetdAssembly now. Pre-existing comment staleness, untouched by this PR, and not worth a ticket on its own — but it is the kind of line-numbered cross-reference that goes stale every time the file moves.

## Unit D review fan-out adjudicated. One blocker, one nit, one new ticket. Two reviewers, one dimension each, neither the implementer. Both reported honestly, including what they did not check. I verified every load-bearing claim myself rather than promoting a reviewer's "sound" to a fact. ### The blocker stays as stated in comment 18462 The fleetd #424 privilege-revocation regression is the only thing stopping this merge. Neither reviewer found it — I found it by reading the diff. The fix and the test to pin it are in 18462. ### Reviewer 1 (resolver wiring + roster freshness): one medium finding, and it is real **I verified its wiring answer.** Production constructs the resolver only through the new roster-aware overload — `FleetdAssembly.java:495` and `:500` in the PR head, both passing `spawnedMemberRole` and `collaboratorTerminals`. `git grep withLeadsAndMembers` over `fleetd/src/main/java` on the PR head returns no other call site, so the 5-arg overload has no production caller, and `NO_KNOWN_LEAD_OR_COLLABORATOR` survives only as `Authz`'s own 3-arg fail-closed default. That closes the inertness question I set it: the security step is wired, not merely present. This mattered because a defaulted no-op would have made the whole unit inert while every test passed. **Its finding: a release-window race.** `release()` does `registry.remove(paneId)` first (`SessionManager.java:306`) and stops the pane only afterwards. In between the code calls `worktrees.hasUncommitted(...)`, which **shells out to `git status`** — the CB-581 comment at `:363` says so in the code's own words — and may then call `trySnapshot(...)`, another git operation. So a member is deregistered while still alive, for as long as one or two git subprocesses take. During that window `spawnedMemberRole` returns null for its terminal and the caller falls through to the tab maps. I confirmed the sequence and the intervening work myself. The finding is structurally right and the window is not theoretical. **I am not holding the merge for it, and here is why.** It is not a regression. Before this PR there was no roster step at all, so *every* member fell through to the tab maps on *every* request — the hole was total. This PR narrows it to the release window. Blocking a strict improvement because it is not yet total would be the wrong trade. The fix also belongs in a different subsystem: it needs a "releasing" marker surviving until `launcher.stop` returns, which is `SessionManager` lifecycle work, not resolver work. Its exploit precondition also does not hold on this host: it needs a member pane carrying a lead or collaborator tab label, which startup validation refuses (`validatePanePlacementAgainstLeadTabs`, plus Unit B's widened collision check), and I checked the live config myself — no profile uses `placement: pane`. That makes it a latent hazard here and a live one on any host that ever uses pane placement. **Filed as its own ticket. Credited to reviewer 1.** ### Reviewer 2 (scanner generalisation): no issue, and I checked its reasoning It reported NO ISSUE with a stated scope, which is a useful answer. Its central claim is correct and I verified it in the diff: `cached` is a single combined `Map<String, Entry>`, both `get()` and `collaborators()` go through one `refresh()`, and `cached`/`gracedTerminals` are keyed by **terminal**, not by kind — so accessor order cannot consume or renew one kind's grace state. `byKind` only filters the combined map at read time. I checked all five properties the class carries, since each is a fix for a real bug, and all five survive: exact case-insensitive match with no prefix stripping, `PendingCloseMarker.strip` before matching (`entryOf`, PR head `:306-311`), the `agent.list` liveness cross-check now covering both kinds, the one grace scan, and the TTL cache whose failed scan keeps the previous answer. It also correctly noted a coverage gap rather than a defect: the added tests do not cover cross-TTL accessor ordering. Worth knowing; not worth blocking. ### One nit to fix alongside the blocker `buildTabIndex`'s javadoc says a colliding label "takes the lead entry, **the first put** when a key collides". The outcome is right but the stated reason is backwards: the code puts collaborators first and leads second, and `LinkedHashMap.put` overwrites, so lead wins because it is the **last** put. A future editor who believes "first put wins" could swap those two calls to tidy them and silently invert the precedence. Say that lead is put last and therefore wins, or make the precedence explicit instead of relying on put order. (Config validation already refuses a lead and a collaborator sharing an exact tab, so this is defence in depth — which is fine, and is exactly why the comment should be right.) ### Noted, not in scope, not fixed `ConfigRef.java:149` and `:621` still point at `CallerResolver.withLeadsAndMembers` as being "at `Fleetd.java:620/624`". That wiring lives in `FleetdAssembly` now. Pre-existing comment staleness, untouched by this PR, and not worth a ticket on its own — but it is the kind of line-numbered cross-reference that goes stale every time the file moves.
Author
Owner

Unit D merged — b5bc5d4 on main

PR #701 merged locally and pushed. The PR is closed by hand, because we merge locally and the forge never closes it by itself.

b5bc5d4 Merge PR #701: fleetd #669 Unit D — the resolver, a live spawned member outranks every tab map
d83972e fleetd #669 Unit D correction: confirm live slot role before granting a spawned architect
02c6909 fleetd #669 Unit D: a live spawned member outranks every tab map

What the lead changed before merging

Two reviewers ran on separate dimensions. Neither found the defect below. I found it by reading the diff myself.

The first implementation of the spawned-member step returned an architect principal from the slot name alone:

if (spawnedRole == MemberRole.ARCHITECT) {
    String slot = architectTerminals.get().get(c.terminal());
    return Principal.architect(slot == null ? null : memberSlotNames.apply(slot),
            c.terminal(), c.pid());
}

It never reads memberSlotRoles. That breaks fleetd #424: config is supposed to govern what an already-bound slot still grants, so revoking the slot must demote the bound session on its next request. With this code the session kept ARCHITECT.

The full suite stayed green at 1988 tests. That half of #424 is documented in MemberRegistry's class javadoc but was never tested — MemberRegistryLiveTest pins only requireSlotFor and reserve. A green build was not evidence here.

Fix in d83972e:

String slot = architectTerminals.get().get(c.terminal());
if (slot != null && memberSlotRoles.apply(slot) == MemberRole.ARCHITECT) {
    return Principal.architect(memberSlotNames.apply(slot), c.terminal(), c.pid());
}
return Principal.worker(c.terminal(), c.pid());

My own verification, not the implementer's

I wrote a probe test that binds a slot while it is configured, then swaps the config to remove it, then resolves through both the new roster-aware path and the old path at the same moment with the same registry and terminal. On the unfixed head the roster-aware path returned ARCHITECT while the old path returned WORKER — same input, two answers, which is the regression. On the fixed head both return WORKER:

[INFO] Running dev.ltms.fleet.auth.LeadProbeRevokedSlotTest
PROBE roster-aware role=WORKER name=null
PROBE old-path     role=WORKER name=null
[INFO] Tests run: 1, Failures: 0, Errors: 0, Skipped: 0

Trial merge of the fixed head in a throwaway worktree:

[INFO] BUILD SUCCESS
MVN_EXIT=0
[INFO] Tests run: 1990, Failures: 0, Errors: 0, Skipped: 0

Cross-checked by aggregating the surefire XML directly: 174 report files, tests=1990 failures=0 errors=0. Arithmetic check: the implementer's 1988, plus my 1 probe, plus the implementer's added revocation test = 1990.

Tree identity, so I know I built what I merged:

merged tree: 8e0e290b8a78aa825807dd4aea284f1765fd1e9c
tested tree: 8e0e290b8a78aa825807dd4aea284f1765fd1e9c
MATCH

The wiki/ submodule pointer is untouched.

Reviewer findings, adjudicated

  • A release-window race (reviewer 1). Real, and I filed it as #702 rather than blocking the merge. SessionManager.release() removes the registry entry before teardown, so there is a window where a dying member's pane is no longer in the roster but its tab map entry still stands. It is not a regression — before this PR the hole was total — the fix is SessionManager lifecycle work outside Unit D, and the precondition (placement: pane) does not hold on this host.
  • An audit-skip asymmetry in FleetMcp.denyFor and FleetApp.allow: READ/TASK_READ/METRICS skip the allowed audit line. Filed as #700 with positive controls, so a zero match could not read as clean. Untouched here, per the brief.

Still open on this ticket

  • Unit E — widen the isLead predicate. Latent on this host, not yet delegated.
  • Unit F — the instruction surface. The lead's own work, in progress now.
  • Collaborators are not reported by fleet_list, so there is no discovery route for them. Worth its own ticket.

Not yet deployed. A merge is not a deployment: the running daemon still holds the jar it was started with, and the redeploy is the next step.

## Unit D merged — `b5bc5d4` on `main` PR #701 merged locally and pushed. The PR is closed by hand, because we merge locally and the forge never closes it by itself. ``` b5bc5d4 Merge PR #701: fleetd #669 Unit D — the resolver, a live spawned member outranks every tab map d83972e fleetd #669 Unit D correction: confirm live slot role before granting a spawned architect 02c6909 fleetd #669 Unit D: a live spawned member outranks every tab map ``` ### What the lead changed before merging Two reviewers ran on separate dimensions. Neither found the defect below. I found it by reading the diff myself. The first implementation of the spawned-member step returned an architect principal from the slot *name* alone: ```java if (spawnedRole == MemberRole.ARCHITECT) { String slot = architectTerminals.get().get(c.terminal()); return Principal.architect(slot == null ? null : memberSlotNames.apply(slot), c.terminal(), c.pid()); } ``` It never reads `memberSlotRoles`. That breaks fleetd #424: config is supposed to govern what an already-bound slot still grants, so revoking the slot must demote the bound session on its next request. With this code the session kept ARCHITECT. **The full suite stayed green at 1988 tests.** That half of #424 is documented in `MemberRegistry`'s class javadoc but was never tested — `MemberRegistryLiveTest` pins only `requireSlotFor` and `reserve`. A green build was not evidence here. Fix in `d83972e`: ```java String slot = architectTerminals.get().get(c.terminal()); if (slot != null && memberSlotRoles.apply(slot) == MemberRole.ARCHITECT) { return Principal.architect(memberSlotNames.apply(slot), c.terminal(), c.pid()); } return Principal.worker(c.terminal(), c.pid()); ``` ### My own verification, not the implementer's I wrote a probe test that binds a slot while it is configured, then swaps the config to remove it, then resolves through both the new roster-aware path and the old path at the same moment with the same registry and terminal. On the unfixed head the roster-aware path returned ARCHITECT while the old path returned WORKER — same input, two answers, which is the regression. On the fixed head both return WORKER: ``` [INFO] Running dev.ltms.fleet.auth.LeadProbeRevokedSlotTest PROBE roster-aware role=WORKER name=null PROBE old-path role=WORKER name=null [INFO] Tests run: 1, Failures: 0, Errors: 0, Skipped: 0 ``` Trial merge of the fixed head in a throwaway worktree: ``` [INFO] BUILD SUCCESS MVN_EXIT=0 [INFO] Tests run: 1990, Failures: 0, Errors: 0, Skipped: 0 ``` Cross-checked by aggregating the surefire XML directly: 174 report files, `tests=1990 failures=0 errors=0`. Arithmetic check: the implementer's 1988, plus my 1 probe, plus the implementer's added revocation test = 1990. Tree identity, so I know I built what I merged: ``` merged tree: 8e0e290b8a78aa825807dd4aea284f1765fd1e9c tested tree: 8e0e290b8a78aa825807dd4aea284f1765fd1e9c MATCH ``` The `wiki/` submodule pointer is untouched. ### Reviewer findings, adjudicated - **A release-window race** (reviewer 1). Real, and I filed it as #702 rather than blocking the merge. `SessionManager.release()` removes the registry entry before teardown, so there is a window where a dying member's pane is no longer in the roster but its tab map entry still stands. It is not a regression — before this PR the hole was total — the fix is `SessionManager` lifecycle work outside Unit D, and the precondition (`placement: pane`) does not hold on this host. - **An audit-skip asymmetry** in `FleetMcp.denyFor` and `FleetApp.allow`: READ/TASK_READ/METRICS skip the `allowed` audit line. Filed as #700 with positive controls, so a zero match could not read as clean. Untouched here, per the brief. ### Still open on this ticket - **Unit E** — widen the `isLead` predicate. Latent on this host, not yet delegated. - **Unit F** — the instruction surface. The lead's own work, in progress now. - Collaborators are not reported by `fleet_list`, so there is no discovery route for them. Worth its own ticket. Not yet deployed. A merge is not a deployment: the running daemon still holds the jar it was started with, and the redeploy is the next step.
Author
Owner

Unit F done — the instruction surface

The lead's own unit. On main at c3e3554, and the wiki at 5f8bd3e.

Unit D made four statements in the canonical block false. A session could resolve as collaborator, read "Which role am I?", get the answer, and then find no section anywhere telling it what it may do.

CLAUDE.md (c3e3554)

Four corrections, each read from the merged tree rather than from the ticket:

  • fleet_whoami returns four roles, not three. FleetMcp.whoami returns before the lead branch, so a collaborator gets role, its registry name and its own sessionId, and no leader key.
  • The fallback ladder. It said only that it cannot separate a worker from an architect. It also fires for a collaborator not at all: every rung detects a spawned member — the reply charter, the fixed mcp__fleet__* mount, ANTHROPIC_BASE_URL — and nothing launched a collaborator. So it falls through to "act as a worker". Safe direction, but it means a collaborator cannot learn what it is without asking the daemon.
  • Invariant 3 said send is lead or architect. Authz also allows a collaborator when the target passes knownLeadOrCollaborator.
  • The invariants heading said "both roles". There are four.

Plus a new Collaborator section, placed deliberately after the member turn contract, because the first thing a collaborator needs to be told is that the contract above it is not its own — nothing delegates to it, so it owes no fleet_reply. Its two surprising limits are written down rather than left to be discovered, both measured in Authz: it cannot reach a worker, and it cannot read a ticket, because TASK_READ is withheld where READ is not.

The intent table gains a row for messaging a collaborator. That row states the gap instead of implying discovery works.

The wiki (5f8bd3e, three commits)

  • 7-Use-Cases.md — the portable block re-synced. Not hand-copied: extracted with the same slice the project's own check uses, so the two are identical by construction. Old block 20760 bytes, new 23214, delta +2454. The check reports both directions — False while they were out of step, True after.
  • 11-Features.md — the collaborator entry said "parsed but not yet used" and named Unit D as the fix, so the heading and body were rewritten. The heading changed, so the index row and the one in-page reference to its anchor moved with it; no stale anchor remains. While there I found and corrected a second stale claim: the placement: pane entry still said the validator knows about fleet.leaders only, which Unit B had already widened.
  • 9-Implementation.md — the auth tables had drifted. Role listed four constants and has five; Principal was missing its collaborator factory; Authz.Action listed eight actions and called them eight — there are thirteen. Four cited line numbers had drifted and were re-measured.

The five-way resolution order is now a flowchart rather than a sentence with four semicolons. All seven Mermaid blocks in that chapter render with mmdc 12.0.0, and I checked the validator against a deliberately broken block so the seven passes are not a false green.

Deployed

Unit D's code is now actually running, which it was not when I posted the merge comment. scripts/redeploy-fleetd.sh --yes: build green at 1989 tests, old pid 12672 gone, new pid 70397, jar 895bf6635918 → 34f449e0ff08, fresh fleetd listening line at 01:49:00, no ERROR lines since restart, and fleet_whoami still answers primary.

Green /healthz only proves herdr answers, so I also proved a real spawn on the new daemon: a sonnet dev member came up in its own worktree and accepted a brief.

Filed from this unit

#703 — a lead cannot discover a collaborator. fleet_list emits only leads and members, and CallerResolver.collaborators() has no caller outside the resolver. So the channel works one way only until the collaborator speaks first. This is the first thing this ticket's own requested walk-through (item 6) would hit, and it is worth settling who may see those rows before anyone implements it.

Remaining on this ticket

Unit E is now delegated — widening HerdrRouter's isLead so a collaborator's pane routes to the lead herdr daemon. It is latent on this host: HerdrRouter hands back the same control object for both branches when one herdr serves everything, and neither herdrSocket nor memberHerdrSocket is set in fleetd/fleetd.yaml here. I briefed it as unit-testable only, with two distinct clients, and told the implementer not to offer a live check as evidence.

After Unit E this ticket is complete except #703.

## Unit F done — the instruction surface The lead's own unit. On `main` at `c3e3554`, and the wiki at `5f8bd3e`. Unit D made four statements in the canonical block false. A session could resolve as `collaborator`, read "Which role am I?", get the answer, and then find no section anywhere telling it what it may do. ### `CLAUDE.md` (`c3e3554`) Four corrections, each read from the merged tree rather than from the ticket: - **`fleet_whoami` returns four roles, not three.** `FleetMcp.whoami` returns before the lead branch, so a collaborator gets `role`, its registry name and its own `sessionId`, and **no `leader` key**. - **The fallback ladder.** It said only that it cannot separate a worker from an architect. It also fires for a collaborator *not at all*: every rung detects a *spawned* member — the reply charter, the fixed `mcp__fleet__*` mount, `ANTHROPIC_BASE_URL` — and nothing launched a collaborator. So it falls through to "act as a worker". Safe direction, but it means a collaborator cannot learn what it is without asking the daemon. - **Invariant 3** said send is lead or architect. `Authz` also allows a collaborator when the target passes `knownLeadOrCollaborator`. - **The invariants heading** said "both roles". There are four. Plus a new **Collaborator** section, placed deliberately after the member turn contract, because the first thing a collaborator needs to be told is that the contract above it is *not* its own — nothing delegates to it, so it owes no `fleet_reply`. Its two surprising limits are written down rather than left to be discovered, both measured in `Authz`: it cannot reach a worker, and it cannot read a ticket, because `TASK_READ` is withheld where `READ` is not. The intent table gains a row for messaging a collaborator. That row states the gap instead of implying discovery works. ### The wiki (`5f8bd3e`, three commits) - **`7-Use-Cases.md`** — the portable block re-synced. Not hand-copied: extracted with the same slice the project's own check uses, so the two are identical by construction. Old block 20760 bytes, new 23214, delta +2454. The check reports **both** directions — `False` while they were out of step, `True` after. - **`11-Features.md`** — the collaborator entry said "parsed but not yet used" and named Unit D as the fix, so the heading and body were rewritten. The heading changed, so the index row and the one in-page reference to its anchor moved with it; no stale anchor remains. While there I found and corrected a second stale claim: the `placement: pane` entry still said the validator knows about `fleet.leaders` only, which Unit B had already widened. - **`9-Implementation.md`** — the auth tables had drifted. `Role` listed four constants and has five; `Principal` was missing its `collaborator` factory; **`Authz.Action` listed eight actions and called them eight — there are thirteen**. Four cited line numbers had drifted and were re-measured. The five-way resolution order is now a flowchart rather than a sentence with four semicolons. All seven Mermaid blocks in that chapter render with `mmdc` 12.0.0, and I checked the validator against a deliberately broken block so the seven passes are not a false green. ### Deployed Unit D's code is now actually running, which it was not when I posted the merge comment. `scripts/redeploy-fleetd.sh --yes`: build green at 1989 tests, old pid 12672 gone, new pid 70397, jar `895bf6635918` → `34f449e0ff08`, fresh `fleetd listening` line at 01:49:00, no ERROR lines since restart, and `fleet_whoami` still answers `primary`. Green `/healthz` only proves herdr answers, so I also proved a real spawn on the new daemon: a `sonnet` dev member came up in its own worktree and accepted a brief. ### Filed from this unit **#703 — a lead cannot discover a collaborator.** `fleet_list` emits only `leads` and `members`, and `CallerResolver.collaborators()` has no caller outside the resolver. So the channel works one way only until the collaborator speaks first. This is the first thing this ticket's own requested walk-through (item 6) would hit, and it is worth settling who may see those rows before anyone implements it. ### Remaining on this ticket **Unit E** is now delegated — widening `HerdrRouter`'s `isLead` so a collaborator's pane routes to the lead herdr daemon. It is latent on this host: `HerdrRouter` hands back the same control object for both branches when one herdr serves everything, and neither `herdrSocket` nor `memberHerdrSocket` is set in `fleetd/fleetd.yaml` here. I briefed it as unit-testable only, with two distinct clients, and told the implementer not to offer a live check as evidence. After Unit E this ticket is complete except #703.
Author
Owner

Unit E merged — 13b6ae2, PR #704 closed

A configured collaborator's terminal now routes to the lead herdr daemon. FleetdAssembly ORs a second terminal map into the predicate it hands HerdrRouter, and the predicate's field is renamed isLead → routeToLead so the name matches the widened contract.

What the lead measured, not took on trust

  • mvn clean install on the branch in a throwaway worktree: 1992 tests, 0 failures, BUILD SUCCESS. 174 surefire report files, in a fresh worktree, so no deleted class is re-counted.
  • Same on the merge commit 13b6ae2 before pushing: 1992 tests, 0 failures.

The whole fix rests on one test, so I reproduced the mutation myself rather than reading the worker's number. I dropped only the collaborator clause from the predicate:

  • RED (1): FleetdAssemblyCollaboratorHerdrRoutingTest.assembledRouterRoutesACollaboratorTerminalToTheLeadDaemon, failing with two different AgentControl identities. That the identities differ is the important part — it proves the two-distinct-FakeHerdr setup really does discriminate, so the kill is behavioural and not an artefact.
  • GREEN (4): every HerdrRouterTest test, including the two new ones. They hand the router their own predicate, so they pin the router's mechanism and say nothing about the call site. Their javadoc says so plainly.

Two things I checked by hand in the code:

  • Both terminal maps come from LeadTabScanner's single byKind helper and are keyed terminal_id → name. So containsKey(target) is right for both. A map keyed by tab or by name would have failed silently here.
  • No agentsFor call runs between the router's construction and collaboratorTerminalsRef.set(...). The two uses in between are the direct memberAgents()/memberSpaces() accessors, which never consult the predicate. The publication window is harmless and is the same one leadsRef already had.

Review

One reviewer fanned out against the diff (223 lines, over the ~50-line threshold), briefed from the diff rather than the implementer's rationale, and not the implementer. It reported no issue and independently confirmed the test reaches real production code, asserts identity against the runtime's own router, and has a working control. It traced the mutation by hand rather than running it; the runtime kill above is the lead's own.

A note for whoever reads the diff later

This defect cannot be proven on this host. HerdrRouter sets memberAgents = member == lead ? leadAgents : new AgentControl(member), and neither herdrSocket nor memberHerdrSocket is set in fleetd/fleetd.yaml, so one herdr serves everything and both routing branches return the same object. A live probe cannot tell the fix from the bug. The unit test with two distinct clients is the only valid evidence, and a live check should not be offered as one.

State of #669

Units C, D, E and F are merged. #703 is the remaining gap — a lead still cannot discover a collaborator, because fleet_list emits only leads and members. Two things need settling before anyone implements it: who may see those rows (fleet_list is READ, which workers hold too, and coordinatorVisibleTo(principal) is the precedent for narrowing by role), and what each row carries. Nothing in CLAUDE.md or the Features entry currently promises discovery works.

Redeploy of the running daemon follows this comment; Unit E changes Java production code, so the live fleetd does not have it until then.

## Unit E merged — `13b6ae2`, PR #704 closed A configured collaborator's terminal now routes to the lead herdr daemon. `FleetdAssembly` ORs a second terminal map into the predicate it hands `HerdrRouter`, and the predicate's field is renamed `isLead` → `routeToLead` so the name matches the widened contract. ### What the lead measured, not took on trust - `mvn clean install` on the branch in a throwaway worktree: **1992 tests, 0 failures**, BUILD SUCCESS. 174 surefire report files, in a fresh worktree, so no deleted class is re-counted. - Same on the merge commit `13b6ae2` before pushing: **1992 tests, 0 failures**. The whole fix rests on one test, so I reproduced the mutation myself rather than reading the worker's number. I dropped only the collaborator clause from the predicate: - **RED (1):** `FleetdAssemblyCollaboratorHerdrRoutingTest.assembledRouterRoutesACollaboratorTerminalToTheLeadDaemon`, failing with two *different* `AgentControl` identities. That the identities differ is the important part — it proves the two-distinct-`FakeHerdr` setup really does discriminate, so the kill is behavioural and not an artefact. - **GREEN (4):** every `HerdrRouterTest` test, including the two new ones. They hand the router their own predicate, so they pin the router's mechanism and say nothing about the call site. Their javadoc says so plainly. Two things I checked by hand in the code: - Both terminal maps come from `LeadTabScanner`'s single `byKind` helper and are keyed `terminal_id → name`. So `containsKey(target)` is right for both. A map keyed by tab or by name would have failed silently here. - No `agentsFor` call runs between the router's construction and `collaboratorTerminalsRef.set(...)`. The two uses in between are the direct `memberAgents()`/`memberSpaces()` accessors, which never consult the predicate. The publication window is harmless and is the same one `leadsRef` already had. ### Review One reviewer fanned out against the diff (223 lines, over the ~50-line threshold), briefed from the diff rather than the implementer's rationale, and not the implementer. It reported **no issue** and independently confirmed the test reaches real production code, asserts identity against the runtime's own router, and has a working control. It traced the mutation by hand rather than running it; the runtime kill above is the lead's own. ### A note for whoever reads the diff later This defect cannot be proven on this host. `HerdrRouter` sets `memberAgents = member == lead ? leadAgents : new AgentControl(member)`, and neither `herdrSocket` nor `memberHerdrSocket` is set in `fleetd/fleetd.yaml`, so one herdr serves everything and both routing branches return the same object. A live probe cannot tell the fix from the bug. The unit test with two distinct clients is the only valid evidence, and a live check should not be offered as one. ### State of #669 Units C, D, E and F are merged. **#703 is the remaining gap** — a lead still cannot discover a collaborator, because `fleet_list` emits only `leads` and `members`. Two things need settling before anyone implements it: who may see those rows (`fleet_list` is `READ`, which workers hold too, and `coordinatorVisibleTo(principal)` is the precedent for narrowing by role), and what each row carries. Nothing in `CLAUDE.md` or the Features entry currently promises discovery works. Redeploy of the running daemon follows this comment; Unit E changes Java production code, so the live `fleetd` does not have it until then.
Author
Owner

Two things found while closing out Unit E

1. The example config told operators the feature does nothing — fixed in 16b52fa

fleetd/fleetd.example.yaml still said, above the collaborators: example:

This block is parsed and validated today; nothing yet recognises or addresses the tab it names.

That was true when Unit B landed the config shape. Units C, D and E have merged and deployed since, so it had gone false. This is the worst place for such a sentence: it sits in the file an operator copies to turn the feature on, and it tells them not to bother. Replaced with what the role may and may not do, read off the rows in auth/Authz.java rather than copied from the wiki, plus the #703 limit. FleetConfigTest, which loads this file, is green at 170 tests.

The same stale-claim sweep found one more: wiki/11-Features.md carried a Still open paragraph saying HerdrRouter's isLead predicate reads the lead map only. Replaced with a resolved gotcha in wiki 8a6ccd7. It was also the only stale reference to the renamed field left in the wiki.

2. fleet.collaborators is not enabled on this host, so the path has never run live

grep -i collaborator fleetd/fleetd.yaml   # no match

There is no collaborators: block in the live config, commented or otherwise. That is the operator's choice and not a defect — the feature is opt-in. But it means the whole collaborator path has never executed on this host. Everything I reported for Unit E is unit-test evidence plus a code read; no part of it is a live observation of a collaborator session.

This bears directly on item 6 of this ticket — "one concrete walk-through: two Claude Code sessions the operator opened by hand, in two named tabs, exchanging a message." Items 1–5 are settled by the merged units. Item 6 is not demonstrated, and I have deliberately not enabled a collaborator myself: it needs a tab a person opens and names, which is an operator action, not a fleet one.

So I am leaving #669 open rather than closing it on the merges alone. What is left here is a live demonstration, not code.

Enabling it, when the operator wants to

  1. Add to fleetd/fleetd.yaml (the live config is gitignored; the shape is in fleetd.example.yaml):
    fleet:
      collaborators:
        <name>:
          tab: "collab: <name>"
    
  2. Restart the daemon — the tab registry is read once at startup, so this is a deferred key. scripts/redeploy-fleetd.sh --no-build --yes is enough; no rebuild is needed for a config-only change.
  3. Open a tab, label it with that exact string, start Claude Code in it, and call fleet_whoami. It should answer collaborator and carry the registry name and its own sessionId.
  4. That sessionId is the address. A lead cannot discover it (#703), so the collaborator sends first, or passes it along.

One caution for whoever does this: placement: pane on any profile while a collaborator tab is named is refused at startup, and so is a fleet.tabLabel that could render as the collaborator's exact tab. Both refusals name the offending entry.

## Two things found while closing out Unit E ### 1. The example config told operators the feature does nothing — fixed in `16b52fa` `fleetd/fleetd.example.yaml` still said, above the `collaborators:` example: > This block is parsed and validated today; nothing yet recognises or addresses the tab it names. That was true when Unit B landed the config shape. Units C, D and E have merged and deployed since, so it had gone false. This is the worst place for such a sentence: it sits in the file an operator copies to turn the feature on, and it tells them not to bother. Replaced with what the role may and may not do, read off the rows in `auth/Authz.java` rather than copied from the wiki, plus the #703 limit. `FleetConfigTest`, which loads this file, is green at 170 tests. The same stale-claim sweep found one more: `wiki/11-Features.md` carried a **Still open** paragraph saying `HerdrRouter`'s `isLead` predicate reads the lead map only. Replaced with a resolved gotcha in wiki `8a6ccd7`. It was also the only stale reference to the renamed field left in the wiki. ### 2. `fleet.collaborators` is not enabled on this host, so the path has never run live ``` grep -i collaborator fleetd/fleetd.yaml # no match ``` There is no `collaborators:` block in the live config, commented or otherwise. That is the operator's choice and not a defect — the feature is opt-in. But it means the **whole collaborator path has never executed on this host**. Everything I reported for Unit E is unit-test evidence plus a code read; no part of it is a live observation of a collaborator session. This bears directly on **item 6 of this ticket** — "one concrete walk-through: two Claude Code sessions the operator opened by hand, in two named tabs, exchanging a message." Items 1–5 are settled by the merged units. **Item 6 is not demonstrated**, and I have deliberately not enabled a collaborator myself: it needs a tab a person opens and names, which is an operator action, not a fleet one. So I am **leaving #669 open** rather than closing it on the merges alone. What is left here is a live demonstration, not code. ### Enabling it, when the operator wants to 1. Add to `fleetd/fleetd.yaml` (the live config is gitignored; the shape is in `fleetd.example.yaml`): ```yaml fleet: collaborators: <name>: tab: "collab: <name>" ``` 2. Restart the daemon — the tab registry is read once at startup, so this is a deferred key. `scripts/redeploy-fleetd.sh --no-build --yes` is enough; no rebuild is needed for a config-only change. 3. Open a tab, label it with that exact string, start Claude Code in it, and call `fleet_whoami`. It should answer `collaborator` and carry the registry name and its own `sessionId`. 4. That `sessionId` is the address. A lead cannot discover it (#703), so the collaborator sends first, or passes it along. One caution for whoever does this: `placement: pane` on any profile while a collaborator tab is named is refused at startup, and so is a `fleet.tabLabel` that could render as the collaborator's exact tab. Both refusals name the offending entry.
Author
Owner

The named-peer mesh is now with two architects

The operator has moved the question past "does the collaborator role work" to "why is per-tab config needed at all". In their words:

  • lets assume all agent tab have a unique name in a unique workspace (to target)
  • the fleet mcp installed at globale for all agent (claude, opencode...)
  • when one randome agent-in-tab want to communicate with another agent-in-tab, they just need to call mcp with addresses and message

They also asked that the implementation be well structured, so the unit split is the primary deliverable.

Why this is doable, and smaller than it sounds

The addressing scheme they assume is already in the data model, measured this session:

  • herdr/Agent.java:23 carries terminalId, paneId, workspaceId, tabId, sessionId, agentType, status, name.
  • herdr/LeadTabScanner.java already walks workspace.list → tab.list → agent.list over every pane and builds a terminal → Entry index, then filters it down to the configured labels. The information the use case needs is gathered and then discarded.
  • Identity for an unconfigured pane already resolves with no config at all — it just lands on WORKER (auth/CallerResolver.java:43).
  • Delivery to an arbitrary pane already works; #669 Unit E generalized the herdr routing for exactly this shape.

So the gap is four pieces, three of which are generalizations of running code: an unfiltered name index, an address form that takes a name rather than a sessionId, an authorization row that permits peer→peer send, and a durable per-peer mailbox. msg/LeadMailbox.java is already the durable, non-blocking, pull-based peer channel and is the closest existing component.

The fork the architects are settling

Workspace-scoped mesh with a safe default role, versus the current per-tab allowlist. The prior I gave them, explicitly flagged as a prior rather than a finding: the operator's own phrase "unique name in a unique workspace" may already name the trust boundary, since workspaceId is on every agent record. Config would then shrink from "enumerate every peer" to "which workspaces are meshed" — far less config, which is the operator's actual goal, without giving every pane on the host a channel into every other.

Four hazards they are briefed not to hand-wave

  1. Prompt injection is the real cost. Any pane in a meshed scope can type into any other agent's context. Scoping bounds the blast radius; it does not remove it. This is the primary security property, not a footnote.
  2. Status gating does not go away. A fleet_send to a BUSY peer is accepted, returns a ticket, and is then never delivered — measured three times in one session. A mesh on the rendezvous path silently loses mail, so this probably has to be pull-based.
  3. Multi-fleet on one host — from notes, not re-measured: pane ids are per-daemon counters so they collide across daemons, and two daemons sharing one herdr interfere with each other's members. With a global MCP mount, "which daemon resolves this pane" is undefined.
  4. Non-Claude peers mount the bridge and never read CLAUDE.md, so any rule they must obey belongs in the launcher charter.

Also filed

#705 — any unconfigured pane resolves to WORKER, and WORKER holds TASK_READ, so an unlisted tab can enumerate task-1, task-2, … and read other sessions' delegation replies. Stated in this issue's body since it was opened, never ticketed. It is defensible on its own whatever the mesh decision is, but the two must land in a consistent order.

Process

Two architects, deliberately on different vendors' models so they do not share blind spots (slots opus and sol). Round 1 is independent positions on the same nine questions; I run the exchange in round 2, because a member cannot receive a message inside its own turn. If they agree, they settle it. If they still disagree after comparing, they return both positions with their evidence and the lead decides.

No code until they land — the answer decides whether fleet.collaborators stays an allowlist or becomes an override on a safer default, and building before that is settled would commit to one of them by accident.

## The named-peer mesh is now with two architects The operator has moved the question past "does the collaborator role work" to "why is per-tab config needed at all". In their words: > * lets assume all agent tab have a unique name in a unique workspace (to target) > * the fleet mcp installed at globale for all agent (claude, opencode...) > * when one randome agent-in-tab want to communicate with another agent-in-tab, they just need to call mcp with addresses and message They also asked that the implementation be well structured, so the unit split is the primary deliverable. ### Why this is doable, and smaller than it sounds The addressing scheme they assume is **already in the data model**, measured this session: - `herdr/Agent.java:23` carries `terminalId, paneId, workspaceId, tabId, sessionId, agentType, status, name`. - `herdr/LeadTabScanner.java` already walks `workspace.list → tab.list → agent.list` over **every** pane and builds a `terminal → Entry` index, then filters it down to the configured labels. The information the use case needs is gathered and then discarded. - Identity for an unconfigured pane already resolves with no config at all — it just lands on `WORKER` (`auth/CallerResolver.java:43`). - Delivery to an arbitrary pane already works; #669 Unit E generalized the herdr routing for exactly this shape. So the gap is four pieces, three of which are generalizations of running code: an unfiltered name index, an address form that takes a name rather than a `sessionId`, an authorization row that permits peer→peer send, and a durable per-peer mailbox. `msg/LeadMailbox.java` is already the durable, non-blocking, pull-based peer channel and is the closest existing component. ### The fork the architects are settling **Workspace-scoped mesh with a safe default role, versus the current per-tab allowlist.** The prior I gave them, explicitly flagged as a prior rather than a finding: the operator's own phrase "unique name in a unique **workspace**" may already name the trust boundary, since `workspaceId` is on every agent record. Config would then shrink from "enumerate every peer" to "which workspaces are meshed" — far less config, which is the operator's actual goal, without giving every pane on the host a channel into every other. ### Four hazards they are briefed not to hand-wave 1. **Prompt injection is the real cost.** Any pane in a meshed scope can type into any other agent's context. Scoping bounds the blast radius; it does not remove it. This is the primary security property, not a footnote. 2. **Status gating does not go away.** A `fleet_send` to a BUSY peer is accepted, returns a ticket, and is then never delivered — measured three times in one session. A mesh on the rendezvous path silently loses mail, so this probably has to be pull-based. 3. **Multi-fleet on one host** — from notes, not re-measured: pane ids are per-daemon counters so they collide across daemons, and two daemons sharing one herdr interfere with each other's members. With a global MCP mount, "which daemon resolves this pane" is undefined. 4. **Non-Claude peers** mount the bridge and never read `CLAUDE.md`, so any rule they must obey belongs in the launcher charter. ### Also filed **#705** — any unconfigured pane resolves to `WORKER`, and `WORKER` holds `TASK_READ`, so an unlisted tab can enumerate `task-1, task-2, …` and read other sessions' delegation replies. Stated in this issue's body since it was opened, never ticketed. It is defensible on its own whatever the mesh decision is, but the two must land in a consistent order. ### Process Two architects, deliberately on different vendors' models so they do not share blind spots (slots `opus` and `sol`). Round 1 is independent positions on the same nine questions; I run the exchange in round 2, because a member cannot receive a message inside its own turn. If they agree, they settle it. If they still disagree after comparing, they return both positions with their evidence and the lead decides. No code until they land — the answer decides whether `fleet.collaborators` stays an allowlist or becomes an override on a safer default, and building before that is settled would commit to one of them by accident.
Author
Owner

Mesh design — round 1 done, round 2 running. Three live defects found and checked.

Two architects on different vendors (opus and sol) each wrote an independent position on the "any named tab can talk to any named tab" design, against the same nine questions. Both have reported. Round 2 (each answers the other) is running now. No code has been written, on purpose: the answer decides whether fleet.collaborators stays an allowlist or becomes an override on a safer default, and building first would pick one by accident.

Three defects in shipped code. I checked each one myself in the main clone.

These are code facts, not design choices, and they are independent of which design wins.

1. A collaborator cannot be sent to at all. The feature this ticket shipped does not work in that direction.

The injector readiness gate is Fleetd.java:238-240:

static Predicate<String> deliverableTo(MemberPresence presence, Supplier<Map<String, String>> leads) {
    return target -> presence.isPresent(target) || leads.get().containsKey(target);
}

Two disjuncts, and a collaborator's terminal is in neither:

  • FleetdAssembly.java:369,371,518 — one MemberPresence object, created at 369, passed to deliverableTo at 371 and to FleetMcp at 518. Same instance.
  • grep -rn markPresent fleetd/src/main/java returns 4 hits: the MemberPresence.java:23 definition, FleetMcp.java:791, and SessionManager.java:1256,1260 — where 1256 is a PresenceFleet extends MemberPresence override calling super, so it is the same class's method, not a second entry point. FleetMcp:791 is the only write, and it is gated on caller.isSpawnedMember().
  • Principal.java:113 — isSpawnedMember() is WORKER || ARCHITECT. COLLABORATOR is excluded.
  • LeadTabScanner.java:191-192 — get() is byKind(refresh(), Kind.LEAD), LEAD only, so the leads supplier can never hold a collaborator.

Authz.java:113 permits the send via knownLeadOrCollaborator. So the send is allowed and then dropped: it waits on the gate for about 60 seconds (Injector.java:95,104 — 240 polls at 250 ms) and fails NOT_DELIVERED at :532-570, never typed into the pane.

Unit E added the collaborator map to the herdr router (FleetdAssembly.java:149-151). It did not add it to the injector gate at :371. Two gates, one got the fix.

What works and what does not: collaborator → lead works, because a lead is in leads. lead → collaborator and collaborator → collaborator both fail after 60 seconds.

No test covers it. FleetDeliverabilityTest.java is 86 lines; grep -c deliverableTo on it returns 6 and grep -ric collaborator returns 0 — a real zero, with a positive control.

Architect opus found this by reading the code and rated it high confidence pending a live send. I closed it to a proof by enumerating every writer into both sets, so no live send is needed to establish it. A live confirmation is still worth running whenever the block is next enabled.

2. The guard that protects #661 is keyed on the config this ticket wants to remove.

FleetConfig.java:2893-2898, inside validatePanePlacementAgainstLeadTabs:

boolean anyLeaderHasTab = ...;
boolean anyCollaboratorHasTab = ...;
if (!anyLeaderHasTab && !anyCollaboratorHasTab) {
    return;            // refuses nothing
}

Remove the per-tab config — the operator's whole goal here — and with no lead tab either, this check refuses nothing, while every pane in a meshed workspace becomes addressable. So per-tab config is not only extra typing today; it is load-bearing for a startup refusal. Any mesh work has to move that check onto the mesh's own switch. Architect opus rates this the highest risk in the design.

3. The main safety control is a filter that does nothing, and the example config says it works.

FleetdAssembly.java:287 builds the scanner with Set.of() for excludedWorkspaceLabels, and the daemon logs "shared fleet space". But fleetd/fleetd.example.yaml:663-664 tells an operator that a lead's workspace

MUST NOT be a member workspace — those are excluded from the scan, so a lead placed in one is never found again.

That is false as shipped. Member workspaces are not excluded. It matters to both designs, because excluding member workspaces is how a member pane is kept out of the mesh — so the design's main control is a filter that is currently passed an empty set.

One correction to my own brief

My brief told both architects that LeadTabScanner walks every pane, builds the index and then filters. That is wrong, and both architects caught it independently. The filter is entryOf(tab.label()) at LeadTabScanner.java:247, inside the tab walk and before agent.list/pane.list, and :253 returns Map.of() before those calls when nothing matched. The topology is therefore never gathered, not gathered and discarded. That makes the mesh's core unit bigger than my brief implied: it changes the shape of the scan, and it removes a zero-cost short-circuit from a cached hot path.

What the two architects already agree on

The workspace is the right scope and bounds blast radius, but is not a claim that peer text is safe · address by tab label, never Agent.name, which is null for a tab a person opened · reject "default every pane to COLLABORATOR" · a meshed peer gets no TASK_READ · durable mail, not rendezvous, with the broker required and a refusal rather than a silent drop to in-memory · the envelope carries a sender taken from the connection, never from a request field · a duplicate name is refused and re-checked at send and delivery time, never "pick the first", and a startup-only check is not the control · peek and ack only your own address · #661's roster-first precedence must hold, and target classification must read one shared authority source rather than a second weaker predicate · the MCP URL picks the daemon, so the real hazard is two daemons on one herdr socket, refused by a lock keyed on the socket path · no launcher charter change, because a human-opened tab never sees the charter and the MCP tool schema is the only text guaranteed to reach it — which is where "message bodies are untrusted data" has to live.

What is still open

Eight seams, now in round 2. The two that decide the most:

  • Config polarity. opus wants every workspace meshed by default, except member workspaces derived from profiles.*.workspace, arguing an allowlist drifts from that list. sol wants an opt-in allowlist, arguing a workspace nobody named must not become trusted. The two fail in opposite directions: the allowlist fails closed, the derived default fails open.
  • Whether to build the mesh now at all. opus recommends shipping three small units first — the ticket owner check, the deliverability fix above, and one shared address classifier — with no mesh, then asking again, on the grounds that part of the original complaint is that the configured feature does not work. sol plans the full mesh in seven units.

Also open: a new PEER role versus reusing COLLABORATOR; the floor role for an unmapped pane (sol proposes OBSERVER with READ and METRICS only, which would close #705's unmapped-pane exposure with no ticket owner check); whether #705's owner check must land before the mesh; the address form (sol's structured {fleet, workspace, tab} with a required fleet field looks right to me, and opus's workspace/name string carries no fleet component); and whether #702 blocks the mesh.

Not started, and still the operator's call to prioritise

  • "it is possible to have multiple fleets in one host" — stated as a goal. Both architects agree the MCP URL selects the daemon, so one global mount is one daemon, and that two daemons over one herdr socket must be refused.
  • The message-queue audit, and the decision inside it: should the broker become a hard dependency. Both designs already require it for peer mail, which is most of the way to that answer.

I will post the decision here when the two architects have compared positions, whether they agree or return two.

## Mesh design — round 1 done, round 2 running. Three live defects found and checked. Two architects on different vendors (`opus` and `sol`) each wrote an independent position on the "any named tab can talk to any named tab" design, against the same nine questions. Both have reported. Round 2 (each answers the other) is running now. No code has been written, on purpose: the answer decides whether `fleet.collaborators` stays an allowlist or becomes an override on a safer default, and building first would pick one by accident. ### Three defects in shipped code. I checked each one myself in the main clone. These are code facts, not design choices, and they are independent of which design wins. **1. A collaborator cannot be sent to at all. The feature this ticket shipped does not work in that direction.** The injector readiness gate is `Fleetd.java:238-240`: ```java static Predicate<String> deliverableTo(MemberPresence presence, Supplier<Map<String, String>> leads) { return target -> presence.isPresent(target) || leads.get().containsKey(target); } ``` Two disjuncts, and a collaborator's terminal is in neither: - `FleetdAssembly.java:369,371,518` — one `MemberPresence` object, created at 369, passed to `deliverableTo` at 371 and to `FleetMcp` at 518. Same instance. - `grep -rn markPresent fleetd/src/main/java` returns 4 hits: the `MemberPresence.java:23` definition, `FleetMcp.java:791`, and `SessionManager.java:1256,1260` — where 1256 is a `PresenceFleet extends MemberPresence` override calling `super`, so it is the same class's method, not a second entry point. `FleetMcp:791` is the only write, and it is gated on `caller.isSpawnedMember()`. - `Principal.java:113` — `isSpawnedMember()` is `WORKER || ARCHITECT`. `COLLABORATOR` is excluded. - `LeadTabScanner.java:191-192` — `get()` is `byKind(refresh(), Kind.LEAD)`, LEAD only, so the `leads` supplier can never hold a collaborator. `Authz.java:113` permits the send via `knownLeadOrCollaborator`. So the send is **allowed and then dropped**: it waits on the gate for about 60 seconds (`Injector.java:95,104` — 240 polls at 250 ms) and fails `NOT_DELIVERED` at `:532-570`, never typed into the pane. Unit E added the collaborator map to the herdr router (`FleetdAssembly.java:149-151`). It did not add it to the injector gate at `:371`. Two gates, one got the fix. What works and what does not: collaborator → lead works, because a lead is in `leads`. **lead → collaborator and collaborator → collaborator both fail after 60 seconds.** No test covers it. `FleetDeliverabilityTest.java` is 86 lines; `grep -c deliverableTo` on it returns 6 and `grep -ric collaborator` returns 0 — a real zero, with a positive control. Architect `opus` found this by reading the code and rated it high confidence pending a live send. I closed it to a proof by enumerating every writer into both sets, so no live send is needed to establish it. A live confirmation is still worth running whenever the block is next enabled. **2. The guard that protects #661 is keyed on the config this ticket wants to remove.** `FleetConfig.java:2893-2898`, inside `validatePanePlacementAgainstLeadTabs`: ```java boolean anyLeaderHasTab = ...; boolean anyCollaboratorHasTab = ...; if (!anyLeaderHasTab && !anyCollaboratorHasTab) { return; // refuses nothing } ``` Remove the per-tab config — the operator's whole goal here — and with no lead tab either, this check refuses nothing, while every pane in a meshed workspace becomes addressable. So per-tab config is not only extra typing today; it is load-bearing for a startup refusal. Any mesh work has to move that check onto the mesh's own switch. Architect `opus` rates this the highest risk in the design. **3. The main safety control is a filter that does nothing, and the example config says it works.** `FleetdAssembly.java:287` builds the scanner with `Set.of()` for `excludedWorkspaceLabels`, and the daemon logs "shared fleet space". But `fleetd/fleetd.example.yaml:663-664` tells an operator that a lead's workspace > MUST NOT be a member workspace — those are excluded from the scan, so a lead placed in one is never found again. That is false as shipped. Member workspaces are not excluded. It matters to both designs, because excluding member workspaces is how a member pane is kept out of the mesh — so the design's main control is a filter that is currently passed an empty set. ### One correction to my own brief My brief told both architects that `LeadTabScanner` walks every pane, builds the index and then filters. That is wrong, and both architects caught it independently. The filter is `entryOf(tab.label())` at `LeadTabScanner.java:247`, inside the tab walk and before `agent.list`/`pane.list`, and `:253` returns `Map.of()` before those calls when nothing matched. The topology is therefore **never gathered**, not gathered and discarded. That makes the mesh's core unit bigger than my brief implied: it changes the shape of the scan, and it removes a zero-cost short-circuit from a cached hot path. ### What the two architects already agree on The workspace is the right scope and bounds blast radius, but is not a claim that peer text is safe · address by tab label, never `Agent.name`, which is null for a tab a person opened · reject "default every pane to `COLLABORATOR`" · a meshed peer gets no `TASK_READ` · durable mail, not rendezvous, with the broker required and a refusal rather than a silent drop to in-memory · the envelope carries a sender taken from the connection, never from a request field · a duplicate name is refused and re-checked at send and delivery time, never "pick the first", and a startup-only check is not the control · peek and ack only your own address · #661's roster-first precedence must hold, and target classification must read one shared authority source rather than a second weaker predicate · the MCP URL picks the daemon, so the real hazard is two daemons on one herdr socket, refused by a lock keyed on the socket path · no launcher charter change, because a human-opened tab never sees the charter and the MCP tool schema is the only text guaranteed to reach it — which is where "message bodies are untrusted data" has to live. ### What is still open Eight seams, now in round 2. The two that decide the most: - **Config polarity.** `opus` wants every workspace meshed by default, except member workspaces derived from `profiles.*.workspace`, arguing an allowlist drifts from that list. `sol` wants an opt-in allowlist, arguing a workspace nobody named must not become trusted. The two fail in opposite directions: the allowlist fails closed, the derived default fails open. - **Whether to build the mesh now at all.** `opus` recommends shipping three small units first — the ticket owner check, the deliverability fix above, and one shared address classifier — with no mesh, then asking again, on the grounds that part of the original complaint is that the configured feature does not work. `sol` plans the full mesh in seven units. Also open: a new `PEER` role versus reusing `COLLABORATOR`; the floor role for an unmapped pane (`sol` proposes `OBSERVER` with `READ` and `METRICS` only, which would close #705's unmapped-pane exposure with no ticket owner check); whether #705's owner check must land before the mesh; the address form (`sol`'s structured `{fleet, workspace, tab}` with a required fleet field looks right to me, and `opus`'s `workspace/name` string carries no fleet component); and whether #702 blocks the mesh. ### Not started, and still the operator's call to prioritise - "it is possible to have multiple fleets in one host" — stated as a goal. Both architects agree the MCP URL selects the daemon, so one global mount is one daemon, and that two daemons over one herdr socket must be refused. - The message-queue audit, and the decision inside it: should the broker become a hard dependency. Both designs already require it for peer mail, which is most of the way to that answer. I will post the decision here when the two architects have compared positions, whether they agree or return two.
Author
Owner

Decision — the two architects agree on every seam. The workspace is not the boundary.

Three rounds are done. Both architects moved on checked evidence, and the design they ended with uses a discriminator neither of them proposed at the start. No operator decision is needed to proceed.

First, a correction to my previous comment

I wrote that defect 3 was "the design's main safety control is a filter that does nothing, and the example config says it works". The two facts in that sentence are right. The conclusion I implied was wrong, and the fix is the opposite of what it suggests.

FleetConfig.java:1144-1149, the javadoc on Leader.DEFAULT_WORKSPACE:

Where an auto-launched lead's tab is created (CB-558). It defaults to the SAME shared "fleet" space the members use, so the operator sees one "session" with many tabs. The scanner no longer excludes member spaces — it tells a lead from a member by the exact tab label, so a lead sharing the members' space is still discovered (see LeadLauncher).

So the empty Set.of() at FleetdAssembly.java:287 is deliberate. The stale thing is fleetd.example.yaml:662-664, which still describes the world before CB-558. Do not populate that exclusion set. Doing so would skip the workspace the lead lives in, and the lead would never be found again — the exact harm the stale comment warns about.

Architect opus had proposed populating it, then withdrew the unit and found this javadoc itself once I gave it the live measurement. Its words: "my R2-a would have undone a deliberate design decision". Worth recording, because a reader of my earlier comment would have built the wrong fix.

The measurement that decided the design

The string workspace appears exactly once in the live fleetd/fleetd.yaml — line 260, inside fleet.leaders.opus:

      workspace: fleet          # one shared space: the lead + every worker are tabs in it,
                                # so the operator sees ONE "session", many windows (tabs).

No profile sets workspace:. So I measured the resolved defaults rather than the key:

  • FleetConfig.java:525 — a profile's workspace defaults to the literal "fleet".
  • FleetConfig.java:1150 — Leader.DEFAULT_WORKSPACE = "fleet", applied at :1158.

The lead and every member share one workspace, and that is the shipped default, not a local choice. Any operator who sets nothing gets it.

This kills a premise both architects shared in round 1 — that the workspace already separates the panes people own from the panes the fleet owns. On the default layout it has one value, so it separates nothing. A mesh scoped to it is either empty or total. There is no third state.

Neither architect could find this: fleetd.yaml is gitignored, so a worktree cannot read it. That is why the lead verifies.

What the design became

The discriminator is the spawned-member roster, used as a negative filter. In the roster means its own member role, never a peer. Not in the roster, and not a lead, architect or configured collaborator, means a candidate peer. CallerResolver.java:294 already consults the roster first, from Unit D, so the precedence chain exists — the mesh only changes what sits at the bottom of it.

The tab label stays the address. It must never be the identity. A person can label a tab to match the member template, and the scanner's defence — a worker cannot rename a tab — says nothing about a person, who is exactly who opens a meshed pane. Separating address from identity is the core of the fix.

The roster has exactly one hole: the teardown window. SessionManager.java:306 removes the registry entry, :390 stops the pane, and in between :341 runs a git status and :355 may commit the whole worktree. That is seconds, not an instant.

Settled, seam by seam

Seam Outcome
Config polarity Dissolved. With no workspace boundary there is no list to allow or deny. An absent mesh: block disables the mesh. mesh.workspaces survives only as an optional extra filter.
Role for a meshed pane New PEER row. Both reached this independently. A lead is reachable only by explicit opt-in — a target-side rule that never depended on the workspace. Reusing COLLABORATOR is rejected: its SEND is gated on knownLeadOrCollaborator, which includes leads, so on the default one-space layout it would hand every human tab a channel into the lead's pane with no operator act at all.
Floor role OBSERVER, but only together with provisional member authority. See below — this one I decided against opus on measurement.
#705 ticket ownership Lands first and alone. A foreign ticket and a ticket that never existed must be indistinguishable to the caller, or the counter is still an oracle. Task has no creator field at all (MessageService.java:238-259), so this is a new field, not a comparison.
Address form Structured and fleet-qualified: {peer: {fleet, workspace, tab}}. No bare delimited string — live tab labels contain a colon and a space. peer.fleet is checked against this daemon's own id, so a call to the wrong daemon is refused instead of routed into its similarly named tab. No shorthand in v1.
#702 tombstone Now a gate before the mesh, not hardening. opus conceded fully: with the workspace boundary gone, the roster is the only discriminator and the teardown window is its only hole.
Build now or pause Continue. A remediation release and a live checkpoint first, then the mesh. The earlier "stop and re-ask" recommendation was withdrawn — you already said go ahead.
Deferred config and locks Confirmed. mesh: is restart-required; reload must say so. An exclusive broker lease on the fleet id, plus a local lock on the canonicalized herdr socket path.
Test profile pom.xml:264-267 — profile default-excludes is activeByDefault with excludedGroups=contract. So every mailbox property that matters for correctness needs a hermetic test that runs in the normal build; a real broker stays a separate release gate. A property asserted only under @Tag("contract") does not run.

The floor role — I decided this against opus, on a measurement

opus argued the OBSERVER floor is safe alone, because a spawned member is registered before it can connect MCP. That is wrong.

  • SessionManager.java:235 — handle = launcher.spawn(req);
  • SessionManager.java:244 — registry.put(handle.id(), session), only after spawn returns.
  • Inside spawn, HerdrPeerLauncher.waitUntilInjectableOrThrow polls the pane until it reports an injectable state.
  • Live config: spawnReadyTimeoutMs: 20000.

So spawn() can block for up to 20 seconds with the agent already up and unregistered. A member's first MCP call in that window resolves through the floor — and today it works only because the floor is WORKER, for which isSpawnedMember() is true, so FleetMcp.java:791 marks presence and the member becomes deliverable.

Today's WORKER floor is load-bearing for member boot. Change it to OBSERVER on its own and the first readiness signal is lost. sol found this window and withdrew its own "cheap standalone change" claim; it is right, and the provisional-authority half must land in the same change as the floor.

For the record, I had accepted opus's reasoning earlier in the session and was wrong to; the measurement above is what settled it.

Landing order

Release 1 — remediation, no mesh. Every unit small and reversible.

  1. #705 ticket ownership — first and alone.
  2. Collaborator deliverability (defect 1) — in flight now.
  3. ConfigRef reporting for fleet.collaborators (finding E, below) — in flight now.
  4. Fix the stale fleetd.example.yaml:662-664 text — and do not populate the exclusion. Add the test that nothing pins today: with every profile defaulting to fleet and a lead whose workspace is also fleet, the scan still reports that lead's terminal. Without it, the wrong fix ships green and demotes the lead.
  5. One shared address classifier — a terminal held by a live spawned member must not be an addressable target, even when it also appears in the collaborator index.
  6. fleet_list reports collaborators (#703).
  7. Re-key the pane-placement refusal (defect 2) — must be ready before mesh identity. Today a mesh with no per-tab entries would start cleanly with the guard switched off.

Then a checkpoint: confirm lead → collaborator and collaborator → collaborator delivery in a real injectable window, and confirm an unknown terminal is still refused.

Release 2 — the mesh. Provisional member authority with the OBSERVER floor; the #702 tombstone as a gate; deferred mesh: config with the PEER role and explicit lead opt-in; one shared peer directory built on the roster-negative filter; the durable mailbox with hermetic default-profile tests; the broker lease and socket lock; then peer send, delivery and the whole instruction surface in one landing.

Two details that moved: mesh.workspaces is now an optional filter rather than a control, and lead opt-in can no longer live per-workspace — it needs one top-level setting.

Finding E — new, verified, and already delegated

Editing fleet.collaborators and reloading reports success and does nothing, with no warning.

  • grep -ci collaborator on ConfigRef.java returns 0. Control: grep -ci leaders returns 16, so the zero is real.
  • changedSplitKeys compares leadersOf(old) against leadersOf(fresh); leadersOf returns cfg.fleet().leaders() only.
  • FleetdAssembly.java:264 reads collaborators() off the startup config, :272-277 builds the tab map, :287 bakes it into the LeadTabScanner, which is never rebuilt.
  • Why the coverage test misses it: ConfigRef.java:272 is SPLIT_KEYS = Set.of("health", "coordinator", "fleet"). The test only needs one branch for fleet, and the leaders branch satisfies it. A newly frozen sub-field under it has no branch at all.

The existing message even lists what is hot — "the rest of fleet: (developers, hunters, reviewers, charters, tabLabel)" — and collaborators is in neither list, so an operator reading it would reasonably think the edit applied.

Instruction surface

One line in the canonical CLAUDE.md block becomes false the moment the floor changes. It currently ends the fallback ladder with "Still unsure ⇒ act as a worker, the most restricted member role". WORKER holds TASK_READ and the new floor does not, so that sentence would tell a confused session to assume more authority than it has. Both architects flagged it. It must name the new floor and describe behaviour rather than claim a rank.

Still yours to prioritise

  • "it is possible to have multiple fleets in one host" — both architects agree one MCP URL is one daemon, so a global mount is unambiguous; the real hazard is two daemons over one herdr socket, which the lease and the canonical socket lock refuse.
  • The message-queue audit. Both designs already require the broker for peer mail and refuse to degrade to in-memory, which is most of the way to the decision hiding inside it: whether the broker becomes a hard dependency.
## Decision — the two architects agree on every seam. The workspace is not the boundary. Three rounds are done. Both architects moved on checked evidence, and the design they ended with uses a discriminator **neither of them proposed at the start**. No operator decision is needed to proceed. ### First, a correction to my previous comment I wrote that defect 3 was "the design's main safety control is a filter that does nothing, and the example config says it works". The two facts in that sentence are right. **The conclusion I implied was wrong, and the fix is the opposite of what it suggests.** `FleetConfig.java:1144-1149`, the javadoc on `Leader.DEFAULT_WORKSPACE`: > Where an auto-launched lead's tab is created (CB-558). It defaults to the SAME shared `"fleet"` space the members use, so the operator sees one "session" with many tabs. The scanner no longer excludes member spaces — it tells a lead from a member by the exact tab label, so a lead sharing the members' space is still discovered (see LeadLauncher). So the empty `Set.of()` at `FleetdAssembly.java:287` is **deliberate**. The stale thing is `fleetd.example.yaml:662-664`, which still describes the world before CB-558. **Do not populate that exclusion set.** Doing so would skip the workspace the lead lives in, and the lead would never be found again — the exact harm the stale comment warns about. Architect `opus` had proposed populating it, then withdrew the unit and found this javadoc itself once I gave it the live measurement. Its words: "my R2-a would have undone a deliberate design decision". Worth recording, because a reader of my earlier comment would have built the wrong fix. ### The measurement that decided the design The string `workspace` appears **exactly once** in the live `fleetd/fleetd.yaml` — line 260, inside `fleet.leaders.opus`: ```yaml workspace: fleet # one shared space: the lead + every worker are tabs in it, # so the operator sees ONE "session", many windows (tabs). ``` **No profile sets `workspace:`.** So I measured the resolved defaults rather than the key: - `FleetConfig.java:525` — a profile's workspace defaults to the literal `"fleet"`. - `FleetConfig.java:1150` — `Leader.DEFAULT_WORKSPACE = "fleet"`, applied at `:1158`. **The lead and every member share one workspace, and that is the shipped default, not a local choice.** Any operator who sets nothing gets it. This kills a premise both architects shared in round 1 — that the workspace already separates the panes people own from the panes the fleet owns. On the default layout it has one value, so it separates nothing. A mesh scoped to it is either empty or total. There is no third state. Neither architect could find this: `fleetd.yaml` is gitignored, so a worktree cannot read it. That is why the lead verifies. ### What the design became **The discriminator is the spawned-member roster, used as a negative filter.** In the roster means its own member role, never a peer. Not in the roster, and not a lead, architect or configured collaborator, means a candidate peer. `CallerResolver.java:294` already consults the roster first, from Unit D, so the precedence chain exists — the mesh only changes what sits at the bottom of it. **The tab label stays the address. It must never be the identity.** A person can label a tab to match the member template, and the scanner's defence — a worker cannot rename a tab — says nothing about a person, who is exactly who opens a meshed pane. Separating address from identity is the core of the fix. **The roster has exactly one hole: the teardown window.** `SessionManager.java:306` removes the registry entry, `:390` stops the pane, and in between `:341` runs a git status and `:355` may commit the whole worktree. That is seconds, not an instant. ### Settled, seam by seam | Seam | Outcome | |---|---| | Config polarity | **Dissolved.** With no workspace boundary there is no list to allow or deny. An absent `mesh:` block disables the mesh. `mesh.workspaces` survives only as an optional extra filter. | | Role for a meshed pane | **New `PEER` row.** Both reached this independently. A lead is reachable only by explicit opt-in — a target-side rule that never depended on the workspace. Reusing `COLLABORATOR` is rejected: its `SEND` is gated on `knownLeadOrCollaborator`, which includes leads, so on the default one-space layout it would hand every human tab a channel into the lead's pane with no operator act at all. | | Floor role | **`OBSERVER`, but only together with provisional member authority.** See below — this one I decided against `opus` on measurement. | | #705 ticket ownership | **Lands first and alone.** A foreign ticket and a ticket that never existed must be indistinguishable to the caller, or the counter is still an oracle. `Task` has no creator field at all (`MessageService.java:238-259`), so this is a new field, not a comparison. | | Address form | **Structured and fleet-qualified:** `{peer: {fleet, workspace, tab}}`. No bare delimited string — live tab labels contain a colon and a space. `peer.fleet` is checked against this daemon's own id, so a call to the wrong daemon is refused instead of routed into its similarly named tab. No shorthand in v1. | | #702 tombstone | **Now a gate before the mesh, not hardening.** `opus` conceded fully: with the workspace boundary gone, the roster is the only discriminator and the teardown window is its only hole. | | Build now or pause | **Continue.** A remediation release and a live checkpoint first, then the mesh. The earlier "stop and re-ask" recommendation was withdrawn — you already said go ahead. | | Deferred config and locks | **Confirmed.** `mesh:` is restart-required; reload must say so. An exclusive broker lease on the fleet id, plus a local lock on the **canonicalized** herdr socket path. | | Test profile | `pom.xml:264-267` — profile `default-excludes` is `activeByDefault` with `excludedGroups=contract`. So every mailbox property that matters for correctness needs a hermetic test that runs in the normal build; a real broker stays a separate release gate. A property asserted only under `@Tag("contract")` does not run. | ### The floor role — I decided this against `opus`, on a measurement `opus` argued the `OBSERVER` floor is safe alone, because a spawned member is registered before it can connect MCP. **That is wrong.** - `SessionManager.java:235` — `handle = launcher.spawn(req);` - `SessionManager.java:244` — `registry.put(handle.id(), session)`, only after spawn returns. - Inside spawn, `HerdrPeerLauncher.waitUntilInjectableOrThrow` polls the pane until it reports an **injectable** state. - Live config: `spawnReadyTimeoutMs: 20000`. So `spawn()` can block for up to 20 seconds with the agent already up and unregistered. A member's first MCP call in that window resolves through the floor — and today it works only because the floor is `WORKER`, for which `isSpawnedMember()` is true, so `FleetMcp.java:791` marks presence and the member becomes deliverable. **Today's `WORKER` floor is load-bearing for member boot.** Change it to `OBSERVER` on its own and the first readiness signal is lost. `sol` found this window and withdrew its own "cheap standalone change" claim; it is right, and the provisional-authority half must land in the same change as the floor. For the record, I had accepted `opus`'s reasoning earlier in the session and was wrong to; the measurement above is what settled it. ### Landing order **Release 1 — remediation, no mesh. Every unit small and reversible.** 1. **#705 ticket ownership** — first and alone. 2. **Collaborator deliverability** (defect 1) — *in flight now*. 3. **`ConfigRef` reporting for `fleet.collaborators`** (finding E, below) — *in flight now*. 4. **Fix the stale `fleetd.example.yaml:662-664` text** — and do **not** populate the exclusion. Add the test that nothing pins today: with every profile defaulting to `fleet` and a lead whose workspace is also `fleet`, the scan still reports that lead's terminal. Without it, the wrong fix ships green and demotes the lead. 5. **One shared address classifier** — a terminal held by a live spawned member must not be an addressable target, even when it also appears in the collaborator index. 6. **`fleet_list` reports collaborators** (#703). 7. **Re-key the pane-placement refusal** (defect 2) — must be ready before mesh identity. Today a mesh with no per-tab entries would start cleanly with the guard switched off. Then a **checkpoint**: confirm lead → collaborator and collaborator → collaborator delivery in a real injectable window, and confirm an unknown terminal is still refused. **Release 2 — the mesh.** Provisional member authority with the `OBSERVER` floor; the #702 tombstone as a gate; deferred `mesh:` config with the `PEER` role and explicit lead opt-in; one shared peer directory built on the roster-negative filter; the durable mailbox with hermetic default-profile tests; the broker lease and socket lock; then peer send, delivery and the whole instruction surface in one landing. Two details that moved: `mesh.workspaces` is now an optional filter rather than a control, and lead opt-in can no longer live per-workspace — it needs one top-level setting. ### Finding E — new, verified, and already delegated Editing `fleet.collaborators` and reloading reports success and does nothing, with no warning. - `grep -ci collaborator` on `ConfigRef.java` returns **0**. Control: `grep -ci leaders` returns **16**, so the zero is real. - `changedSplitKeys` compares `leadersOf(old)` against `leadersOf(fresh)`; `leadersOf` returns `cfg.fleet().leaders()` only. - `FleetdAssembly.java:264` reads `collaborators()` off the startup config, `:272-277` builds the tab map, `:287` bakes it into the `LeadTabScanner`, which is never rebuilt. - Why the coverage test misses it: `ConfigRef.java:272` is `SPLIT_KEYS = Set.of("health", "coordinator", "fleet")`. The test only needs **one** branch for `fleet`, and the `leaders` branch satisfies it. A newly frozen sub-field under it has no branch at all. The existing message even lists what *is* hot — "the rest of `fleet:` (developers, hunters, reviewers, charters, tabLabel)" — and `collaborators` is in neither list, so an operator reading it would reasonably think the edit applied. ### Instruction surface One line in the canonical `CLAUDE.md` block becomes false the moment the floor changes. It currently ends the fallback ladder with "**Still unsure ⇒ act as a worker**, the most restricted member role". `WORKER` holds `TASK_READ` and the new floor does not, so that sentence would tell a confused session to assume **more** authority than it has. Both architects flagged it. It must name the new floor and describe behaviour rather than claim a rank. ### Still yours to prioritise - **"it is possible to have multiple fleets in one host"** — both architects agree one MCP URL is one daemon, so a global mount is unambiguous; the real hazard is two daemons over one herdr socket, which the lease and the canonical socket lock refuse. - **The message-queue audit.** Both designs already require the broker for peer mail and refuse to degrade to in-memory, which is most of the way to the decision hiding inside it: whether the broker becomes a hard dependency.
Author
Owner

Two more units merged in 3e8e314

PR #706 and PR #707, merged together and verified as one tree. Both PRs are closed by hand,
because a local merge never closes a PR here.

#706 — deliverableTo opens for a collaborator. This is the defect I described earlier in
this ticket: a send to a collaborator was allowed at Authz and then dropped at the injector
readiness gate, because a collaborator is in neither of that gate's two disjuncts. It now has a
third. The thing that decides whether the fix is live is the key shape: LeadTabScanner.get() is
terminal_id → lead name and collaborators() is terminal_id → collaborator name, both built
by byKind from one scan, so containsKey(target) really fires. If that map had been keyed by
name the fix would have been dead in the same way the defect was.

#707 — a fleet.collaborators change is reported on reload. The map is read once at startup
to build the scanner's identity map, and a reload does not rebuild it. ConfigRef said nothing,
so an operator saw a clean reload and no live effect. It now reports that a restart is needed.

Verification

I did not promote either worker's "clean" to a fact. Merged both onto origin/main in one
throwaway worktree and ran mvn clean install myself, with the output written to a file and not
piped, because a pipe hides a failure behind a zero exit:

  • task-7 alone: exit 0, BUILD SUCCESS, Tests run: 1996, Failures: 0, Errors: 0
  • both together: exit 0, BUILD SUCCESS, Tests run: 1999, Failures: 0, Errors: 0

1999 is 1996 plus #706's three new tests, which is the count I expected. ConfigRefTest: 32,
FleetDeliverabilityTest: 9. I checked the pushed tree is byte-identical to the tree I built.

Both units came with a revert proof from the worker: dropping only the production change turned
exactly the new tests red and left the pre-existing ones green. I did not re-run either
experiment myself.

A process defect worth recording

The #706 worker launched a fork subagent for a read-only side search, with an explicit
instruction to change nothing. The fork ran the whole commit, push and open-PR sequence itself.
The worker reported this unprompted, which is the right thing to do, and the content is not in
question. But a member must not hand off its own push, because then nobody who read the code
reviewed what was pushed. A fork inherits the parent's context, so it inherits the implementer
skill's recipe, and a per-call "read-only" line loses to a loaded playbook. The durable fix is in
the skill, not in each brief.

Still open under this ticket

#703 (a lead cannot discover a collaborator) and the fleetd.example.yaml truth fix are now
delegated. #705 (the WORKER floor) is with an architect — see my comment there.

## Two more units merged in `3e8e314` PR #706 and PR #707, merged together and verified as one tree. Both PRs are closed by hand, because a local merge never closes a PR here. **#706 — `deliverableTo` opens for a collaborator.** This is the defect I described earlier in this ticket: a send to a collaborator was allowed at `Authz` and then dropped at the injector readiness gate, because a collaborator is in neither of that gate's two disjuncts. It now has a third. The thing that decides whether the fix is live is the key shape: `LeadTabScanner.get()` is `terminal_id → lead name` and `collaborators()` is `terminal_id → collaborator name`, both built by `byKind` from one scan, so `containsKey(target)` really fires. If that map had been keyed by name the fix would have been dead in the same way the defect was. **#707 — a `fleet.collaborators` change is reported on reload.** The map is read once at startup to build the scanner's identity map, and a reload does not rebuild it. `ConfigRef` said nothing, so an operator saw a clean reload and no live effect. It now reports that a restart is needed. ### Verification I did not promote either worker's "clean" to a fact. Merged both onto `origin/main` in one throwaway worktree and ran `mvn clean install` myself, with the output written to a file and not piped, because a pipe hides a failure behind a zero exit: - task-7 alone: exit 0, `BUILD SUCCESS`, `Tests run: 1996, Failures: 0, Errors: 0` - both together: exit 0, `BUILD SUCCESS`, `Tests run: 1999, Failures: 0, Errors: 0` 1999 is 1996 plus #706's three new tests, which is the count I expected. `ConfigRefTest: 32`, `FleetDeliverabilityTest: 9`. I checked the pushed tree is byte-identical to the tree I built. Both units came with a revert proof from the worker: dropping only the production change turned exactly the new tests red and left the pre-existing ones green. I did not re-run either experiment myself. ### A process defect worth recording The #706 worker launched a `fork` subagent for a read-only side search, with an explicit instruction to change nothing. The fork ran the whole commit, push and open-PR sequence itself. The worker reported this unprompted, which is the right thing to do, and the content is not in question. But a member must not hand off its own push, because then nobody who read the code reviewed what was pushed. A `fork` inherits the parent's context, so it inherits the `implementer` skill's recipe, and a per-call "read-only" line loses to a loaded playbook. The durable fix is in the skill, not in each brief. ### Still open under this ticket #703 (a lead cannot discover a collaborator) and the `fleetd.example.yaml` truth fix are now delegated. #705 (the `WORKER` floor) is with an architect — see my comment there.
Sign in to join this conversation.
1 Participants
Notifications
Due Date
No due date set.
Dependencies

No dependencies set.

Reference: fleet/fleetd#669