Investigate: let any named Claude Code session in herdr talk to the others — a new fleetd role, or a separate project? #669
Open
opened 2026-10-03 19:48:03 +02:00 by ltms
·
23 comments
No Branch/Tag Specified
main
worker/702-4f5c7f-2
worker/715-5c43fc-1
worker/721-70f9ea-5
worker/718-99362b-2
worker/task-15-af0d10-12
worker/task-16-50a702-13
worker/task-12-4d0479-9
worker/task-13-823ce2-10
worker/705-ticket-owner-af9928-8
worker/703-list-collaborators-9c06c2-7
worker/669-example-truth-0b303d-6
worker/669-collab-deliverability-9ba859-3
worker/669-collab-reload-report-2a21bd-4
worker/669-7e80a6-1
worker/669-unit-d-efbbd7-1
worker/669-1b786a-1
worker/669-1d1d9f-1
worker/692-4afb9d-2
worker/689-02fced-13
worker/693-cf23fa-14
worker/677-fix-lead-collision-f69073-12
worker/638-fix-overmask-dbb1bf-11
worker/675-5b7478-4
worker/669-unit-a-70cc8f-3
worker/677-8cdaaf-5
worker/638-a7b391-1
worker/683-4536d6-2
worker/651-a75bbe-8
worker/680-20607d-7
worker/664-c12e95-3
worker/668-08534d-4
worker/672-0f2469-2
worker/670-7d1022-1
worker/661-ac7c28-2
worker/664-37fb9b-3
worker/663-remove-3arg-read-3f6783-1
worker/659-remove-dead-backcompat-ba5e6f-1
worker/637-revision-60a488-23
worker/656-redact-regression-tests-892903-19
worker/637-context-gauge-threshold-466eb5-16
worker/639-redact-line-numbers-de4ac4-17
worker/641-set-reformat-guard-6f96a4-18
worker/642-herdr-guard-scope-5de0e4-15
worker/650-javadoc-scope-95f3b3-14
worker/612-01e9f7-13
worker/612-a-r4-quarantine-outage-7ab0e8-5
worker/612-a-r9-r11-capacity-coverage-peers-cfcc79-7
worker/612-a-r10-loophealth-ccc872-8
worker/612-a-r12-turnregistrar-9e3bb7-9
worker/612-a-r5-leadconfigdir-9e70cf-6
lead/config-edit-redact-anchor-wording
worker/config-edit-seam-ca8dc1-1
worker/612-r67-630-lifecycle-290b8d-3
worker/629-625-ports-seams-da7d5d-4
worker/612-r12-exhaustion-f37cd7-1
worker/612-r38-amqp-24b083-2
worker/fleetd-612-unita-87807e-1
worker/612-b3-mcpwirings-da2b58-3
worker/612-b2-cb185-176d3a-2
worker/612-b1-completion-457459-1
worker/612-agaps-73a926-2
worker/608-sleeps-3a64ff-3
worker/621-b4520b-1
worker/618-b83894-2
worker/fleetd-615-e05481-5
worker/lead-autocompact-5f1ab2-3
worker/fleetd-613-f85deb-3
worker/fleetd-608-flaky-nudge-test-d0c2d1-3
worker/lead-context-gauge-ad404f-1
worker/gauge-wiring-9158c1-4
worker/redeploy-slowstart-ead0e5-5
worker/charter-bytes-13668c-6
worker/rollover-outcome-291483-2
worker/589-f64303-2
worker/593-1a8025-5
worker/589-fcd2aa-1
worker/568-9fdaa2-3
worker/571-attempted-outcome-5739f7-2
worker/581-completionresolver-cas-sites-0542b7-6
worker/562-loop-health-wiring-test-99611c-5
worker/562-surface-loop-health-7df5cc-4
worker/575-waiter-cleanup-sites-62ad80-1
worker/572-answer-lock-release-46a9ae-5
worker/567-probe-channel-leak-a38fc5-6
worker/551-record-before-send-7cbf56-1
worker/561-listener-fanout-survives-a-throw-61d538-2
worker/555-redeploy-main-flow-seam-65c2f5-2
worker/556-injector-owns-registration-e027a5-1
worker/552-post-restart-mktemp-abort-bc2672-4
worker/553-onstatus-completion-leak-0da881-2
worker/550-shasum-linux-196132-1
worker/538-loop-dies-on-error-4a5eeb-6
worker/426-health-coverage-ef1fd4-4
worker/504-failed-reported-clean-3cfd66-3
worker/537-capturedlog-close-e4c437-2
worker/459-broken-link-targets-cadc17-5
worker/535-appender-leak-fe74c1-1
worker/512-part2-shutdown-detection-434701-9
worker/529-logger-level-sweep-2a5533-8
worker/528-drain-gate-call-site-5de83d-7
charter/forge-mcp-vs-token
worker/521-swap-guard-unpinned-28e931-5
worker/519-probe-test-harness-d25ab8-4
worker/525-logger-level-leak-1b4eb0-6
worker/518-fleetmcp-resolver-wiring-8ef96c-1
worker/512-drain-complete-line-7edd71-3
worker/517-abort-branch-and-jar-id-41b641-2
worker/500-9e52c9-3
worker/509-4912f4-2
worker/511-9a4b23-1
worker/493-479f45-2
worker/505-03f8b2-1
worker/492-followup-detect-unclear
worker/501-a31fa0-7
worker/498-451d1c-5
worker/494-1015ce-2
worker/492-209647-1
worker/489-001902-2
worker/480-relative-handover-path-906323-1
worker/480-b-handover-skill-45bf1f-5
worker/474-followup-source-pin-f54a55-17
worker/474-charter-check-on-reload-f54a55-17
worker/466-quarantine-repeatcount-report
worker/393-opencode-skill-seeding-71854b-13
worker/469-canonical-tool-names-2a472a-16
worker/466-quarantine-escalation-5ae9c1-15
worker/446-hot-exhausted-pattern-0af580-6
worker/464-charter-tool-name-guard-a85635-12
worker/463-listfleet-default-fails-open-f1c76c-11
worker/458-invariant-5-by-purpose-862f9a-10
worker/439-coordinator-row-gate-bc032a-8
worker/449-herdr-protocol-576015-4
worker/450-abstract-spawn-599e1c-5
worker/437-ack-refuses-177d91-1
worker/444-placement-window-feb56a-2
worker/440-helddurable-derived-d462d7-13
worker/425-rework-placement-resolve-c58ba1-9
worker/421-lead-peek-held-msgs-cdbad2-10
worker/435-fixed-policy-cap-fe11de-12
worker/422-gate-state-observability-9e79d6-11
worker/431-memberregistry-live-readers-cdbad2-10
worker/424-architect-slot-hot-038b41-7
worker/422-model-gate-spawn-c29f48-6
worker/425-default-profile-live-f55534-8
worker/415-coverage-wording-2cbf9c-5
worker/416-3ad1da-1
worker/418-588283-3
worker/deterministic-stamp-race-409-3cb7b6-10
worker/armed-reads-live-config-404-ed931f-9
worker/reply-peer-refusal-391-5a34bd-7
worker/models-allowlist-aa9e9b-3
worker/ttl-stamp-race-399-f1122f-8
worker/scrub-receipt-400-316b3e-5
worker/exhaustion-detection-395-105105-6
worker/scrub-abort-394-316b3e-5
fix/scrub-uid-abort
worker/task-scrub-517574-2
worker/t386-clock-bd5b78-4
worker/t384-scrub-813790-5
worker/t381-cc-748314-2
worker/t373-336973-2
worker/t365-3920c5-3
worker/t358-6e989b-1
worker/t355-8b321c-1
worker/fleetd-369-hermetic-git-tests-e8b19a-3
worker/fleetd-368-stale-lead-binding-f5682e-2
worker/fleetd-360-deploy-units-0d3793-1
worker/359-dead-lead-tabs-f1253b-4
worker/362-worktree-skills-c03e51-3
worker/361-coord-visibility-655144-1
362-plugin-visibility-and-drift
worker/errscan-bed2ca-2
worker/amqp-log-identity-bed2ca-2
worker/withdefaults-guard-561704
worker/sleepguard-82076d-1
worker/fd334-9ee1b6-5
worker/fd348-f1ab27-4
worker/fd335-a71c35-1
worker/fd342-174a17-2
worker/fd345-490d0f-3
worker/fleetd-337-5ec7d4-21
worker/fleetd-341-af5a6b-24
worker/fleetd-339-5ca0a2-23
worker/fleetd-338-83a4a1-22
worker/fleetd-333-281f46-18
worker/fleetd-329-11bdbb-16
worker/fleetd-330-2770fb-17
worker/fix-326-50506e-15
worker/fix-324-3e9bbf-14
worker/fix-323-b8287d-13
worker/fix-316b-bd0860-11
worker/fix-318-76ca36-9
worker/fix-317-486aec-8
worker/fix-315-ce47c5-6
worker/fix-307-275890-6
worker/fix-308-b4f664-7
worker/fix-309-ec3939-8
worker/fix-310-7a3974-9
worker/fix-302-52ad0e-9
worker/fix-298-ce1acb-8
worker/fix-297-66bd11-7
worker/fix-296-104622-6
worker/fix-293-bare-closetab-eb22b5-3
worker/fix-280-gone-ask-lapse-bca98e-2
worker/fix-290-reapidle-guard-coverage-9b0dd1-1
worker/fix-285-trust-seed-8f3565-10
worker/fix-284-backend-error-seat-85912c-11
worker/fix-282-chained-ask-e6d0bb-8
worker/fix-283-teardown-leaks-f40dfa-9
worker/fix-281-pin-handler-actions-4921ac-7
worker/audit-rendezvous-lifecycle-d072ae-2
worker/audit-health-placement-1a2476-6
worker/audit-teardown-exits-e207a5-3
worker/audit-launcher-asymmetry-27e370-4
worker/audit-rest-authz-6ca53c-5
worker/investigate-275-abandon-asking-fdef52-8
worker/fix-274-worktree-leak-b0095d-7
worker/fix-273-exhausted-pattern-9665b5-6
worker/fleetd-267-model-check-bd8068-1
worker/fleetd-131-archunit-18b834-7
worker/fleetd-266-sshagent-rename-a014ff-6
worker/fleetd-184-uid-claim-8e1f31-4
worker/fleetd-184-warn-b381ee-10
worker/fleetd-184-docs-be1d12-9
worker/fleetd-257-9bf010-7
worker/fleetd-103-23a113-6
worker/fleetd-247-342356-5
worker/fleetd-116-04dea8-4
worker/fleetd-252-a830e0-3
worker/fleetd-111-7e8673-9
worker/fleetd-155c-f8ef4b-8
worker/fleetd-176-b928ca-3
worker/fleetd-249-7a7878-2
worker/cb248-composition-root-b-9acdf7-15
worker/cb148-envrc-default-fa6c82-12
worker/cb201-unit5-wiring-6c12e6-8
worker/cb241-fallback-echo-1175e9-11
worker/cb149-trust-dialog-2392a5-9
worker/cb134-148-overlay-visible-c9b986-10
worker/cb234-session-id-keyed-04e1fc-1
worker/cb201-unit3-nudge-abdf5c-6
worker/cb201-unit2-policy-c1102c-5
worker/cb201-unit4-outcome-a13bfa-7
worker/cb201-unit1-classifier-91b9b1-4
worker/cb201-227-refine-831980-3
worker/cb175-model-readback-0f085f-1
worker/cb222-charter-tmpdir-17f013-1
worker/cb226-architect-slot-race-cd3aa8-3
worker/cb224-worktree-root-group-024523-2
worker/cb-123-role-demotion-c600f7-2
worker/cb-219-opencode-roots-1f677e-1
worker/cb214-claude-session-id-b9eab4-4
worker/cb213-zdotdir-wrong-process-dd6de4-3
worker/cb211-exhaustion-classification-9546e0-2
worker/cb137-ambiguous-task-4df3d8-4
worker/cb209-agentsessionid-4dfdb6-2
worker/cb185-hostenvnames-2692b5-3
worker/cb206-opencode-sqlite-128718-2
worker/cb185-worktree-group-fc0c99-1
worker/cb-137-ask-ticket-e7760c-2
worker/cb-172-broker-uri-d36ae4-4
worker/cb-175-model-readback-76ead6-3
worker/cb-161-pane-ancestry-293510-1
worker/cb-164-rebase-885863-8
worker/cb-164-empty-scrape-false-success-1a80af-3
fix/cb-197-ticket-ttl-from-completion
worker/cb-189-remote-url-coverage-4692f3-1
worker/cb-185-blockers-027756-4
worker/cb-192-gap-log-11b631-2
worker/cb-633-fix-5f4396-3
worker/cb185-router-d6436d-3
worker/cb185-router-routing-gaps-9e9d33-3
worker/cb185-paneids-992586-2
worker/cb-633-allow-list-union-ed374b-1
worker/cb-157-credential-in-remote-url-496e44-2
worker/cb-641-health-herdr-evidence-8f1f54-6
worker/cb-640-health-msg-evidence-99c9cd-1
worker/cb-642-fleets-status-skill-bbbc40-5
cb-634-ide-mcp
worker/lead-comms-wiring-c014b9-7
worker/lead-mailbox-c19577-6
worker/autocompact-window-82bc2f-5
worker/cb-634-probe-18056f-4
worker/cb635-broker-urienv
worker/cb-632-config-retry-8e0efa-7
lead/cb-622e-claude-md
lead/cb-622-followup
worker/cb-622a-165dff-1
lead/cb-622d-opencode-mount
worker/cb-622b-717c67-2
worker/cb-622c-ab7759-3
worker/cb-617b2-20ca4b-3
worker/cb-617a-5c2f4a-1
worker/cb596-4e49ef-3
worker/cb586-10500c-1
worker/cb-606-b9343a-25
worker/cb604-1445f8-24
worker/cb582-477374-21
worker/cb584-8c2281-22
worker/cb600-e6b9a9-20
worker/cb602-ce257f-19
worker/cb601-b42837-18
worker/cb598-6c7ba7-17
worker/cb599-740fe4-16
worker/cb597-282224-15
worker/cb590fix-185e9a-10
worker/cb528-recovery-race
worker/cb594-96bead-8
worker/cb590-916766-2
worker/cb527-997d99-3
worker/cb592-env-leak-3cbf9c-1
worker/cb588-async-ticket-nudge-3218f7-5
worker/cb578b-9dcb13-6
worker/cb581-d24826-5
worker/m2-u5-ef8c42-15
worker/cb578a-516499-2
worker/cb576-01a04b-17
worker/cb579-lead-tab-acba06-20
worker/cb580-terminal-health-ed6058-21
worker/cb577-f36fdc-18
worker/cb573b-3db06f-16
worker/cb568c-f36fdc-18
worker/cb568-drop-cause-c3ac1c
worker/cb575-cancelled-notification-c3ac1c
worker/m4-sol-a2cbec-3
worker/cb574-async-ask-c3ac1c
worker/cb573-health-model-8ca857-14
worker/cb572-unknown-target-7f2e35-13
worker/u4-700706-9
worker/u3-b9fcb6-6
worker/u2-ef5b68-4
worker/u1-469dce-1-clean
worker/u1-469dce-1
worker/cb-564-health-events-70cf7e-2
worker/cb-565-recycle-drops-role-98e58f-3
worker/cb-563-missing-reply-df2866-1
worker/cb-562-readiness-gate-silent-6c23c9-3
worker/cb-560-architect-presence-da8155-1
worker/cb-561-architect-silent-off-a71cab-2
worker/cb-548-bind-architect-slot-fe1b8c-1
worker/parity-overlay-settings-5fb711-1
secrets-central-store
cb-559-hot-key-correction
cb-557-fleet-role-pools
worker/cb-553-maxload-explicit-spawn-305ee3-6
worker/cb-551-idle-lead-heartbeat-f1633c-1
worker/cb-544-drain-preserves-worktree-925fad-3
worker/cb-552-docs-sync-1cb9cf-4
worker/cb-548-rendezvous-guard-rebased
worker/cb-548-rendezvous-guard-116b53-10
worker/cb-548-authz-v2-586df6-8
worker/cb-548-authz-264363-5
salvage/cb-528b-codex-home
salvage/cb-528a-codex-launcher
CB-518-primary-flow
feature/peer-launcher-spi
cb-103-injector
v1.1.0
v1.0.0
Labels
Clear labels
blocked
needs-live-proof
ready-to-delegate
silent-default
Cannot start until something else lands. The body says what.
Merged and green, but never shown working on the running daemon. Not the same as done.
Scope, files and acceptance criteria are written. A worker can be briefed from the body alone.
A feature that compiles, passes tests, and ships turned off. Nine recurrences and counting.
No Label
Milestone
No items
No Milestone
Projects
Clear projects
No project
Notifications
Due Date
No due date set.
Dependencies
No dependencies set.
Reference: fleet/fleetd#669
Reference in New Issue
Block a user
Blocking a user prevents them from interacting with repositories, such as opening or commenting on pull requests or issues. Learn more about blocking a user.
Delete Branch "%!s()"
Deleting a branch is permanent. Although the deleted branch may continue to exist for a short time before it actually gets removed, it CANNOT be undone in most cases. Continue?
Operator request, 2026-10-03, in their words:
This ticket records the request and what I measured before handing it to an architect. No decision is made here.
What exists today — measured in the main clone at
4b4a868The useful surprise is that fleetd is already most of the way there, because a lead is not a spawned session. A lead is a human-driven Claude Code session that fleetd finds by its exact tab label:
LeadTabScanneris handed atab → namemap fromfleet.leaders.<name>.taband matches labels exactly, case-insensitively. Nothing about a lead is launched by the daemon.So "N named sessions that can message each other" is close to expressible now: configure N
fleet.leadersentries, one per tab, and each gets asessionIdthatfleet_listreports and thatfleet_send{sessionId, content}can reach.The blocker is privilege, not addressing.
auth/Authz.java:72:Every
fleet.leadersentry resolves asPRIMARY, so configuring a session as a "leader" to let it talk also hands it the power to spawn, stop and drain the whole fleet, and to roll a lead. That is not a side effect anyone would want for a session whose only need is a message channel.The role that already nearly fits
Role(auth/Role.java) declaresPRIMARY,WORKER,ARCHITECTandANONYMOUS. AndAuthz.java:77:So
ARCHITECTis already "may send and reply and read, may not control the fleet" — exactly the capability profile the request asks for. What an architect is not today is attachable: an architect is spawned by fleetd into its own worktree and bound to a slot, rather than being a session a human already has open in a named tab.Today the identity ladder has only two outcomes for a session sitting in a herdr pane: its tab matches a configured lead label and it is
PRIMARY, or it does not and it isWORKER. There is no middle rung for "a named session I trust to talk, but not to run the fleet".So the real question for the architect
Is the right shape a new non-privileged named-peer role (or an attachable form of
ARCHITECT), keyed on tab label the same way leads already are? And then: fleetd or a separate project?On the second half, name concretely what a separate project would have to rebuild. fleetd already owns all of this, and none of it is incidental:
mcp/ConnectionIdentity,CallerResolver)auth/Authz,auth/Role)fleet_*MCP surface every session already mountscoordIdMy prior is that this belongs in fleetd as a new rung on the existing ladder, because a separate project would have to duplicate identity and authorization — the two things that must not be guessed twice. But that is a prior, not a finding, and the architect should be free to disagree.
Specific things to settle
ARCHITECTattachable? An attachable architect reuses the whole capability row; a new role avoids overloading a word that currently means "advisor the lead consults".fleet.leaderskeyed by tab, or a flag on a leader entry that withholds the primary capabilities.SENDyes.READ/METRICS?COORD_READ(currentlyisPrimary()only, line 94)? Say why for each.Not verified by me
CallerResolvercould resolve a non-lead named tab at all without changes. I readAuthzand theRoleenum myself; I did not trace the resolver for this question.fleet_listalongside leads and members, or kept separate. Not asked.Answer: build it in fleetd, as a new
COLLABORATORroleThe architect's answer is in full on this ticket's reply. The recommendation: keep it in fleetd, add a new attachable role
COLLABORATORbacked by a newfleet.collaborators.<name>.tabregistry. Do not makeARCHITECTattachable, and do not add a flag tofleet.leaders.The reason it belongs here rather than in a separate project is that the hard part is not moving text. It is proving which herdr pane called, mapping a named tab to a role, routing to the right herdr daemon, and applying one authorization table. Those are already one connected boundary —
ConnectionIdentity,CallerResolver,HerdrRouter,Authz. A separate service would either duplicate that boundary or trust fleetd as its identity provider, which leaves the security decision here anyway while splitting the feature across two projects.Two corrections to my own ticket text, both of which I verified
FleetdAssembly.java:274-280callsnew LeadLauncher(...).ensureLeads()when herdr is up and leaders are configured, andFleetConfig.java:1106-1115states it plainly — "A lead is now also creatable (CB-557). Before, nothing spawned one." Recognition still comes first and only the shortfall is launched, but the absolute form of my claim does not hold.ARCHITECTbefore the worker fallback, atCallerResolver.java:208-228.My surviving claim, that
ARCHITECTis already a send-but-not-control role, is correct —Authz.java:68-77.The part that matters most: this can reopen #661 at a lower privilege level
A naive implementation would recreate the hole we just fixed. If a collaborator scanner copied
LeadTabScanner.scan()(which joins every pane to its tab,LeadTabScanner.java:177-235) and ran its lookup before the worker fallback, a pane-placed member landing in a collaborator-labelled tab would be read back as that collaborator and gain localSENDplus peer visibility.The #661 validator does not prevent this. I checked the merged code:
validatePanePlacementAgainstLeadTabs()returns early onfleet.leaders().isEmpty()and only ever inspects lead tabs (FleetConfig.java:2755-2762, now onmainatb4b7cf5). It knows nothing about a collaborator namespace.So three guards are mandatory, not optional:
placement: panewhen either a lead tab or a collaborator tab is configured.tabLabelcannot match the collaborator namespace.CallerResolver, a terminal still present inSessionManagermust resolve as its spawned member role before any tab-attached role is considered. Copying the current lead-first order would repeat the defect exactly.This is the one-way-gate pattern: #661's guard closes the direction the incident came from. A new tab-attached role arrives from the other direction with a green build.
Capabilities, as recommended
Grant:
READ,METRICS, own-paneREPLY, and a target-limited localfleet_sendthat may address only a configured lead or collaborator.Deny:
SPAWN,STOP,DRAIN,HANDOVER,ASK,COORD_READ, the cross-hostcoordIdroute, and theturnIdanswer form.One implementation note with teeth: today a
coordIdsend is authorized as the sameSENDaction as a local send (FleetMcp.java:450-469, mapping at:1098-1108). So addingCOLLABORATORtoSENDwould silently grant broker-wide cross-host lead messaging. That route needs splitting into its own action first, leaving current primary/architect behaviour unchanged.What is not settled
Only one architect answered this question, so this is a single considered position, not an agreement between two. I verified its corrections to my ticket and its #661-reopening claim against the code myself. I have not verified its full capability table line by line, and
COLLABORATOR, the config block and the widened validators are a proposed design that does not exist yet, so nothing here is tested.The architect also noted it did not inspect herdr server code,
fleetd/fleetd.yaml, orwiki/.Next step is an implementation plan split into units; this stays open until then.
Second architect position collected, and the design is now settled. Implementation plan below.
Two architects worked this question independently, on the same brief, unable to see each other. Neither could see the other's answer, so where they agree it is two positions rather than one. I checked every contested claim myself in the main clone at
6f27522before ruling.The headline: the recommendation stands, but it is not safe to implement as written. Both architects independently concluded that
SENDmust be split into separate actions before the role exists. Each found gaps the other missed, and one of them overturns a guard the accepted answer called mandatory.What both architects concluded, independently
Build it in fleetd. Add
COLLABORATORandfleet.collaborators.<name>.tab. Do not makeARCHITECTattachable. DenySPAWN,STOP,DRAIN,HANDOVER,COORD_READand the cross-hostcoordIdroute. GrantMETRICSand own-paneREPLY. The #661 attack is real, and the resolver must checkSessionManagerbefore any tab map. Generalise the existing tab scan instead of adding a second scanner. The live walk-through is the lead's, not a worker's.Both also found, separately, that splitting
SENDmust land first. That is the strongest signal in the two reports, because neither could have copied it.One architect added a mechanical reason not to make
ARCHITECTattachable that nobody had given:Principal.isSpawnedMember()isWORKER || ARCHITECT, andFleetMcp.java:437-440uses it to enrol a caller in the presence map. An attachable architect would be listed as an available member, so a lead'sfleet_listwould offer the operator's own hand-opened session as a worker to delegate to.Disagreement 1 —
READ. I rule for the architect who objected.The accepted answer grants
READ. One architect agreed and checked thatfleet_listhides the coordinator row. The other refused, sayingREADexposes task replies and pending questions, not a safe roster.I checked this myself, and the objection is right. Three facts, each from a command I ran:
fleet_statusreturns the pending question, theturnIdand the ticket id:Polling a ticket is plain
READ— no target means noDRAIN:And ticket ids are a plain in-memory counter:
So the ids are
task-1,task-2, … and they reset on every daemon restart. I have first-hand evidence of how guessable that is from this session:grep -n "task-5" fleetd/fleetd.outreturns 20 separate occurrences against 20 different sessions.Put together: a collaborator holding today's
READcan walktask-1..Nand read the reply of every delegation the lead has run since the last restart, and can callfleet_statuson any member to lift its open question andturnId.The other architect established the same mechanism and drew the opposite conclusion from it — it used "
fleet_poll{ticket}maps toREADand returns the reply text" as the reasonDRAINcan stay denied without breaking the round trip. That is true and useful, and it is also exactly the leak. The same fact serves both arguments; only one of them noticed it cuts both ways.Ruling:
READmust be split beforeCOLLABORATORis granted anything. KeepREADfor roster, profiles and identity. Move ticket polling and session status to their own actions and deny them toCOLLABORATORin the first release.Worth recording: the
READarm's own comment inAuthz.javasays "the roster carries no secrets." That is now false, and it is false forWORKERtoday, not only for a future collaborator. Filed separately.Disagreement 2 —
ASK. I rule for granting it.One architect denied
ASK, arguing it pairs with theturnIdanswer form. The other argued to grant. I side with granting, on the mechanism:That arm is ownership-based, not role-based, so
COLLABORATORgets it with no edit. Denying it means splitting the one rule that currently serves primary, lead, worker and architect — a change to four rows to restrict one — andAuthz.java's own comment calls that arm "the load-bearing rule". The risk is also bounded: a collaborator with nobody blocked on it gets a cleanNO_WAITERrefusal from the mechanism, and when a send is open the worst case is holding that sender's call, which any worker can already do.The accepted limit: a collaborator cannot answer another collaborator's
fleet_ask, because theturnIdroute is denied. A lead can. That is a usability limit, not a hole, and it should be written down rather than discovered.Disagreement 3 — guard 2 does not do what it says. Confirmed, and it is a real pre-existing hole.
The accepted answer says to "widen the member-label collision validator so a fleet or profile
tabLabelcannot match the collaborator namespace." One architect checked which validator that is and found it checks the wrong field.I confirmed this.
validateLeadTabPrefixes()readsleader.tabPrefix(). And across the whole ofFleetConfig.java,leader.tab()appears in exactly two places:Identity is matched on the exact
tab.tabPrefixis vestigial. So no validator compares a lead's exacttabto a member label. An operator who setsfleet.leaders.alpha.tab: "alpha"and a profiletabLabel: "alpha"passes every validator today.Widening the prefix check would inherit that hole into the new namespace. Replace guard 2 with an exact/template collision check: refuse at startup when
fleet.tabLabel, or any profile override, can render to a configured lead or collaboratortab, treating placeholders as wildcards.The lead-side half of this is a pre-existing defect independent of #669 and is filed separately.
Two guards nobody listed, both confirmed by me
fleet_whoamiwould report a collaborator as a lead. The branch isif (!caller.isWorker()), and aCOLLABORATORis neither worker nor architect, so it falls through to the lead branch and gets aleaderkey whilerolereadscollaborator. The answer contradicts itself. This also breaks the instruction surface: the canonical block inCLAUDE.mdtells every session thatfleet_whoamireturnsprimary,workerorarchitect. A fourth value means that block changes in the same work.Herdr routing would send a collaborator's status read to the wrong daemon.
HerdrRouter.agentsForpicksisLead.test(targetId) ? leadAgents : memberAgents, andisLeadis membership in the lead terminal map. A collaborator is not in it.Two open questions I answered from the live config, which no member can read
Both architects correctly said they could not settle these, because
fleetd/fleetd.yamlis gitignored and absent from their worktrees. I checked:memberHerdrSocketis configured —grep -nE 'memberHerdrSocket|herdrSocket' fleetd/fleetd.yamlreturns nothing. SomemberHerdr == herdron this host and the routing bug above is latent, not live. It must still be fixed, because a second herdr daemon is exactly the direction fleet01 goes.placement: pane— all eight sayplacement: tab. So the widened pane-placement refusal is a quiet pass here rather than a startup failure. The change is safe to deploy on this fleet.The plan — six units, in dependency order
Unit A — split
SEND, and splitREAD. Must land first. Separate the local send, thecoordIdbroker route and theturnIdanswer form into distinct actions, and separate ticket-polling and session-status from rosterREAD. Grant every new action to exactly who holds it today, so no caller gains or loses anything. This is a pure refactor and merges on its own. Acceptance: the threefleet_sendcall shapes reach the gate as three different actions; flipping any one arm to deny refuses only that call shape and leaves the other two working — that control is what proves the arms are really separate and not a rename.Unit B — the
fleet.collaboratorsconfig block and its validators. Recognise-only: noprofile, noinstances, never auto-launched. Note thatFleetis@JsonIgnoreProperties(ignoreUnknown = true), so the key is silently ignored today — assert the before state in the test, or it can pass vacuously. Includes the widened pane-placement refusal and the replacement exact/template collision check.Unit C — the
COLLABORATORrole and its authorization row. Needs A and B. Acceptance: a local send naming a configured lead or collaborator is permitted, the same call naming a spawned member's terminal is refused, and that refusal happens over both MCP andPOST /sessions/{id}/message— a test covering only MCP leaves the REST route open. Also:fleet_whoamireportsrole: collaboratorwith noleaderkey, while the same test fed a lead still reportsleader— the second half is the control.Unit D — the resolver. One generalised scan returning
terminal → (name, kind); spawned-member-first ordering. The order is: no pane → token tail; terminal in the member roster → return that member's role and consult no tab map at all; then lead; then collaborator; then the worker fallback. Acceptance: a terminal in both the roster and the lead tab map resolves as its member role — that is #661 closed at the resolver, not only at config validation — and removing the spawned-member step must make that assertion fail. This is the security-critical unit and gets a reviewer who did not write it.Unit E — deliverability. Widen the
isLeadpredicate to "lead or collaborator" and rename it. Latent on this host, per the config check above.Unit F — the instruction surface. Mine. The canonical block's
fleet_whoamiladder, invariant 3's "send is lead or architect" line, the intent→tool table, and awiki/11-Features.mdentry. A member cannot do this: theCLAUDE.md/wiki sync check needswiki/, which is uninitialized in every worker worktree.What is still not settled, and what nobody has tested
Nothing above exists. No architect wrote code and no build was run against any of it, so every acceptance criterion here is a proposal, not a passing test.
Neither architect read herdr's own server code. Neither ran the live walk-through, correctly — it needs a daemon that has the role, and it would cut their own channel.
I have not re-checked every row of either capability table line by line. I verified the contested rows —
READ,ASK,SEND, thewhoamibranch,pollAction, ticket generation andvalidateLeadTabPrefixes— and I am taking the uncontested rows on two architects agreeing.This ticket stays open until Unit A lands.
The two "filed separately" lines in the adjudication above now have tickets:
Authz.java:87-89claims "the roster carries no secrets" under theREADarm. False today:fleet_poll{ticket}isREAD, ticket ids are a sequential counter with no owner check, andfleet_statusreturns another member's open question plus itsturnId. That ticket is the written evidence for why Unit A must land first, so the reason survives if Unit A is re-litigated.tabto a member label. The collision guard checkstabPrefix, which identity no longer uses. Pre-existing and independent of this ticket. #669 Unit D would close it at the resolver, which is where it bites.Neither blocks Unit A.
Unit A has landed
PR #687 merged as
7c458e8onmain(pushed;mainis now2eb2d61). It closes #678.SENDis now three actions (SEND,COORD_SEND,ANSWER) andREADis two (READ,TASK_READ), each new action carrying exactly the grant its combined action carried. Verified: areviewer compared all 156 (role, action, target) pairs against
origin/mainwith 0 mismatches, andI read the grants in
Authzmyself. Merged tree: 1942 tests, 0 failures, BUILD SUCCESS.So the divergence point the rest of this ticket needs now exists.
One thing came out of Unit A that the later units should know about, filed as #689: the REST
route now parses the request body before the authorization gate, because
turnIdpicks the action.That is harmless today only because all three send grants are identical. Unit B gives
COLLABORATORa different grant, and at that moment #689 stops being cosmetic. Fix #689 before orwith Unit B, not after.
State of the remaining units
Units B to F are specified in the 20:58 comment above and none is delegated yet. Unit F — the
instruction surface (the
fleet_whoamiladder, invariant 3, the intent→tool table, a Featuresentry) — stays with the lead, because a member cannot run the
CLAUDE.md/wiki sync check:wiki/is uninitialized in every worker worktree.
This ticket stays open.
State update: Unit A's follow-up is closed, and the daemon is now live on it. Units B–F still open.
mainis atedbd8d8and pushed. Four PRs merged since Unit A landed, each verified by me in a throwaway worktree with a tree-hash comparison against what I built:main7dec74fbfee23ad0688c8edbd8d8#689 is fixed, which unblocks Unit B. This ticket said to fix it "before or with Unit B, not after", and it is now done and merged.
FleetApp.sendMessagechecks the coarseSENDgrant with no body read, then parses, then checksANSWERwhenturnIdis present.One thing worth carrying into Unit B: the ANSWER gate is behaviourally invisible today, because all three send grants are identical. So I first merged a version where deleting the call site left all 56 tests in those classes green. It is now pinned through the audit trail —
allow()logs an allowed entry for every granted action exceptREAD/METRICS/TASK_READ, so a grantedturnIdrequest must log bothSENDandANSWER. Deleting the call site now fails exactly one test. When Unit B givesCOLLABORATORa different grant, that test is what stops the gate having quietly disappeared in the meantime.The daemon is redeployed, so Unit A is actually live
A merge is not a deployment and Unit A's split had never been loaded.
scripts/redeploy-fleetd.sh --yes: build 1950 tests, 0 failures, BUILD SUCCESS; old pid 34147 exited with a clean drain (released=0 abandoned=0); new pid 56206 on jare0e9b2c4109b;/healthz200; a freshfleetd listeningline at 23:00:48.I did not stop at healthz, because healthz only proves herdr answers. I ran a real spawn and a full round trip — spawn, send, the member ran a command,
fleet_replycame back withedbd8d8. The channel works on the new jar.fleet_whoamistill answersprimary.Units B–F
None is delegated. The specs in the 20:58 comment stand unchanged. Two notes for whoever picks them up:
templateCanRenderAs(...), which is the "replacement exact/template collision check" Unit B was going to include. Re-readFleetConfig.validateLeadTabPrefixes()before writing Unit B's brief — part of it exists. What does not exist is thefleet.collaboratorsblock, its validators, or the widened pane-placement refusal.FleetConfig.java. I had to hold it back twice today for exactly this reason. Two file-disjoint units in this project have collided in shared test files before.Unit F is still the lead's. I confirmed the reason is live: the
CLAUDE.md/wiki sync check needswiki/, and it is uninitialized in every worker worktree. I ran that check myself today and it reports in sync.One Unit F item is already done
wiki/11-Features.mdclaimed "Two startup refusals" guard the lead tab namespace and describedtabPrefixas the collision guard. #677 made both statements false — there are four refusals now, and the load-bearing one compares the exacttab. I corrected that section and pushed the wiki (9b4ae2e..b025ae9onrefs/heads/main, verified by ref, because this wiki has both amainand amasterand pushing to the wrong one is silent).The rest of Unit F — the
fleet_whoamiladder, invariant 3's "send is lead or architect" line, the intent→tool table — is not done, and must not be done untilCOLLABORATORactually exists. Writing a fourthfleet_whoamivalue into the canonical block before the code returns it would tell every session something false.Unit B correction — one addition before merge:
collaboratorsis missing from the duplicate-key guardThis is a correction to Unit B's scope, written here rather than sent as a revised brief, because the brief is write-once and a ticket comment is what a member can re-read whenever it looks.
PR #697 is open and under review. Its own §7 report flagged this, and I confirmed it by reading the code — so it is being folded in rather than filed for later.
What is missing
FleetConfig.java:1912-1913:collaboratorsis not in that set. The walk atFleetConfig.java:1970only descends into a pool whose key is in it:So a config naming the same collaborator twice takes the
skipValuebranch.Why that matters, in the existing code's own words
The javadoc on
rejectDuplicateMemberSlotsalready states the reason this guard exists:Every word of that applies to
collaborators. An operator who writesfleet.collaborators.alphatwice gets one of them silently discarded, and the collaborator they configured is simply not recognised — the same harm the five existing pools are protected from.This is a gap in the block Unit B is adding, not a pre-existing defect elsewhere. That is why it belongs in Unit B and not in a follow-up: shipping the block without it means shipping a known hole.
Scope of the addition
Three things, and nothing else:
"collaborators"toFLEET_POOL_KEYS.fleet:". It becomes six. A count written in prose next to the thing it counts is the exact stale-count class #676 was about — fix it in the same commit or it is a new instance of it.FleetConfigTest.javaaround lines 1224-1348:duplicateSlotNamesInOnePoolAreRejectedAtParseTime,theSameSlotNameInTwoPoolsIsNotADuplicate,duplicateKeysOutsideTheFleetPoolsAreUnaffected.I checked and no test pins the contents of
FLEET_POOL_KEYS—grep -rn 'FLEET_POOL_KEYS' fleetd/src/test/javais empty. So adding a key breaks nothing, and nothing would have caught the omission either. Worth knowing when judging how this was missed.What I measured, and what I did not
I read the set, the walk at
:1970and the javadoc directly on PR #697's branch. I did not run a live config with a duplicated collaborator key through the parser. The conclusion that it collapses silently is a code read, not an observed output — the test asked for in item 3 is what will turn it into a measurement.Still out of scope for Unit B
Unchanged:
Authz,CallerResolver,LeadTabScanner,MemberRole,FleetMcp. Those are Units C, D and E.Unit B round 2 — two required changes, both about comments telling the truth
PR #697's code is sound. I mutation-tested it myself (results below) and found no defect in the logic. Both items here are about text, and the first one is partly my fault.
Required 1 — the test javadoc narrates history, and my own ticket comment caused it
FleetConfigTest.java:1517:The project rule is explicit and this breaks four parts of it at once: history (was missing, before the fix), evidence (this is the measurement, a
23:37timestamp), a ticket as the reason, and — for a test specifically — "A test comment names the behaviour. It does not tell the story of the bug it caught."I asked for this. My 23:37 comment said the test "is what turns it into a measurement", and the implementer dutifully wrote the measurement into the javadoc. The measurement belongs in the PR description and the commit message, which is where I should have said to put it. The rule the implementer followed was mine, and it was wrong.
Replace it with what the test protects, in the present tense — roughly "A duplicated name in
fleet.collaboratorsis refused at parse time, like any otherfleet:pool." No ticket, no date, no before-state.That history/evidence shape appears in exactly this one comment block. I swept all 86 added comment lines for
before the fix|was missing|previously|no longer|measurement|measured|verified|ticket correction, and nothing else matched.Required 2 — two places promise behaviour that does not exist yet
FleetConfig.java:1183-1184, on theCollaboratorrecord:fleetd.example.yaml:672-673:Neither is true after Unit B. Nothing recognises a collaborator and nothing addresses one:
COLLABORATORarrives in Unit C, resolution in Unit D, deliverability in Unit E. Today this block is parsed and validated, and that is all.This is the same hazard already recorded against Unit F — writing a fourth
fleet_whoamivalue into the canonical block before the code returns it would tell every session something false.fleetd.example.yamlhas exactly the same readership problem, and it is worse thanCLAUDE.mdhere because an operator uncommenting that block would reasonably expect something to happen.There is a second reason to cut the sentence from the record javadoc, independent of timing: javadoc on a public type is a contract. This record's contract is "a tab label, keyed by name;
tabis required and compared case-insensitively." Whether the daemon addresses that tab is system behaviour owned by other classes, so the sentence is in the wrong place as well as premature.What to do:
profile/instances/kind,tabrequired, compared case-insensitively.fleetd.example.yaml, keep the block (an existing test requires the key to be documented — I confirmedeveryNestedConfigKeyIsDocumentedInTheExampleforces this), but state plainly that the block is accepted and validated now, and that recognition and addressing arrive with #669 Units C-E. One sentence. An operator must not read today's example as a working feature.Optional, one word each — do not expand these
Only if you are already in the file. Neither is worth a round trip on its own, and I would rather you under-edit than restyle working comments:
FleetConfig.java:1186— "tabPrefixis deliberately absent." The reason follows in the next clause, so by the project's own rule the confidence marker adds nothing.FleetConfig.java:1187— "the same as aLeader, so a naming-convention prefix would be vestigial here too" points at another class to justify this one. The rule asks for the fact stated once where it is enforced, not a cross-reference.What I verified myself, so you know the logic is not in question
Three line-anchored mutations in a throwaway worktree of
ad593c9, each compiled before running, each restored after:equalsIgnoreCase→equalson the lead-vs-collaborator tab comparisonaCollaboratorTabEqualToALeadTabRefusesToStartequalsIgnoreCase→equalson the collaborator-vs-collaborator comparisontwoCollaboratorsSharingTheSameExactTabRefusesToStartif (!anyLeaderHasTab)— the original named bugaPanePlacedProfileWithACollaboratorTabRefusesToStartEvenWithNoLeadersThree specific runtime kills, no cascades. The case-insensitivity the example config advertises is genuinely pinned, and the early-return bug cannot silently come back. I deliberately mutated lines the implementer had not mutated itself, since it had already measured the
FLEET_POOL_KEYSline.I also confirmed independently: only one production call site of any
Fleetoverload (FleetConfig.java:2633,withDefaults(), all nulls), andcollaboratorswas appended last in the canonical record so both convenience overloads keep their signatures and forward a trailing null. Had the field been inserted next toleaderswhere it reads more naturally, every positional argument in those overloads would have shifted — invisibly, because that call site passes only nulls. That was the right call and it is worth saying so.A second reviewer found no issue on the Jackson binding and overload dimension, and named what it checked. The validator-logic reviewer has not reported yet; if it finds anything, that is a separate round.
Unit B round 3 — one more test required. A reviewer found it; I proved it by mutation, and it is worse than it was reported.
A reviewer reported this as a low-severity coverage gap. I checked it myself and it is a real gap, but the severity reasoning needs correcting, so do not treat "low" as a reason to skip it.
The gap
FleetConfig.java:2764, in the collaborator loop ofvalidateLeadTabPrefixes():That is the fleet-wide
fleet.tabLabelcheck against a collaborator tab. Every collaborator test in the PR drives the profile-override branch a few lines below it instead, so this branch has no regression pin.Measured, not read
I disabled the branch (
if (false && …)) in a throwaway worktree ofad593c9and ran both config test classes:The mutation survives. The refusal can be removed and the suite stays green.
A survivor has more than one explanation, so I ruled the others out: it is not an equivalent mutation, because disabling the check changes behaviour — and the reviewer independently wrote that exact scenario as a test and watched
validateLeadTabPrefixes()throw and name the collaborator correctly. So the behaviour is right; only the pin is missing.Why "low severity" undersells it — the lead side is pinned three times over
I ran the same mutation on the lead-side equivalent at
FleetConfig.java:2736:So this is not a pre-existing habit in this file. The lead branch is covered three ways; its new collaborator twin is covered zero ways. The implementation mirrored the lead logic correctly and did not mirror the lead's coverage, and that asymmetry arrived with this PR.
That matters here specifically. This ticket already records that #689's ANSWER gate was behaviourally invisible, that deleting its call site left 56 tests green, and that a later unit changing a grant is the moment it stops being cosmetic. An unpinned startup refusal is the same shape: it can be deleted as dead-looking code with a green build, and the thing it was guarding is a privilege boundary. Cheap to pin now, expensive to discover missing later.
What to add
One test:
fleet.tabLabelset directly, no profile override, against a matchingfleet.collaboratorsentry's tab; assertIllegalStateExceptionand that the message names the collaborator. Mirror the existing lead-sideaFleetTabLabelTemplateThatCanRenderAsALeadTabRefusesToStart.Before you commit it, confirm it is a real pin the same way I did: disable line 2764 with
if (false && …), see your new test fail, restore, see it pass. A test that does not go red against that mutation is not a pin, and that check is the whole point of this round.Nothing else changes. No logic edits.
Authz,CallerResolver,LeadTabScanner,MemberRoleandFleetMcpstay out.Two notes on process, for the record
My own first attempt at the lead-side mutation silently did nothing. I anchored the
sedat line 2702, which is a bare}— the substitution never matched, the suite came back 176 green, and that reads exactly like "the lead side is unpinned too". I caught it only because I printed the mutated line afterwards and saw nofalse &&. A line-anchored mutation that fails to apply produces the most reassuring possible result. Print the line, every time.The reviewer's review predates the latest commit. It reports running 167 tests in
FleetConfigTest; that class holds 169 atad593c9, and the duplicate-key round added exactly 2. So it reviewed05244a8and its findings do not cover theFLEET_POOL_KEYSaddition. I am inferring that from the counts rather than measuring it. I am not ordering another review pass for those two tests, because I confirmed that change myself by a different route — the implementer reverted the single line and observed the refusal disappear, and I read the resulting set and walk directly.The reviewer also ended its turn without a
fleet_reply, so its answer arrived only through the pane scrape. The brief did tell it to end with one. That is now several occurrences of the same failure on review-shaped briefs, and it is a fleetd defect rather than a worker mistake — worth its own ticket, which I will file separately rather than bury here.Unit B merged — PR #697, as
f0ff252Merged locally and pushed to
main. PR #697 closed by hand (we merge locally, so Gitea does not close it itself).What I checked myself before merging
Trial merge and build. Throwaway worktree of
origin/mainat736fd9c, mergedrefs/pull/697/head(780cb34) with--no-ff. Fullmvn clean installinfleetd/:Counted independently from the surefire XML after
rm -rf target/surefire-reports:That matches the 1964 the worker reported, and the 1950 → 1964 step from
origin/main.Tree identity. The tree I built and the tree I merged are the same object:
c8a0747465bd977dd369a78e024738d05fa13206, on both the throwaway merge (5a81d3d) and the real merge commit (f0ff252). So the green build covers exactly what landed. The merge commit does not touch thewiki/submodule pointer.Push verified by ref. Local
HEADandgit ls-remote origin refs/heads/mainboth readf0ff25221e21f3db10fdd9091c447c7371932cd3.The round-3 pin, re-measured
Round 3 existed because the fleet-wide
fleet.tabLabel-vs-collaborator-tab check had no test. I did not take the worker's word for the new test being a real pin — I re-ran the mutation.Line-anchored
sedatFleetConfig.java:2761, which I first confirmed by printing its context is inside thefleet.collaborators().forEachblock::2761toif (false && templateCanRenderAs(fleet.tabLabel(), tab)) {:2733diffagainst the backup2761c2761)mvn test -Dtest=FleetConfigTestunder the mutationMVN_EXIT=1, BUILD FAILURE, 170 run, 1 failureaFleetTabLabelTemplateThatCanRenderAsACollaboratorTabRefusesToStart:1084— one kill, no cascadegit difffor the fileMVN_EXIT=0, 170 run, 0 failuresSo the branch is pinned, by exactly one test, and the test fails for the right reason (
Expected java.lang.IllegalStateException to be thrown, but nothing was thrown).One number corrected
The previous lead's handover recorded "176 tests green" for
FleetConfigTestwhile mutating atad593c9. I measure 170 in that class at780cb34. These reconcile against the source rather than against each other:grep -c '@Test'gives 156 onorigin/mainand 170 on the PR, there are no@ParameterizedTestmethods in the file, and only one file is namedFleetConfigTest.java. Surefire ran 170. So 170 is the count of that class and 176 was never it. The previous lead's conclusion — that the branch was unpinned — was still correct, which is why this round happened.Code read, not just built
templateCanRenderAsreturnsfalsefor a null or blanktab, and both new collision loops skip a null/blank tab, so a collaborator with no tab falls through tovalidateMembers(), which refuses it. That is what the javadoc claims, and the three paths agree.The PR body reported
FLEET_POOL_KEYSas an unfixed gap; commitad593c9on the same branch then addedcollaboratorsto it, so that gap is closed in what merged.Still open on this ticket
Units C, D, E and F are specified in the 20:58 comment and none is delegated. Unit C is now unblocked by this merge. Unit D is the security-critical one and must get a reviewer who did not write it. Unit F is the lead's and must still wait for the code: writing a fourth
fleet_whoamivalue into the canonical block beforeCOLLABORATORexists would tell every session something false.A redeploy is now owed — this merge changes Java production code, unlike
736fd9c.Unit C correction — two fixes to my own brief. Both are my errors, not the implementer's.
Written here rather than sent as a revised brief, because the brief is write-once and a ticket comment is what a member can re-read whenever it looks. If this contradicts the Unit C brief, this comment is newer and it wins.
Unit C is in flight on
term_65cf6e090dfe057, branchworker/669-1b786a-1.Correction 1 — do NOT delete the 3-argument
Authz.permitsoverloadMy brief said to delete it "so the compiler forces both call sites to supply the classifier". I measured the cost after sending, and it is wrong.
So there are 47 direct 3-argument call sites and no test helper to change in one place. Deleting the overload forces 47 mechanical test edits.
What it buys does not justify that, because the default is already safe. The hazard I was guarding against is a future call site that skips the target limit. With a deny-all classifier as the 3-argument form's default, such a call site refuses a collaborator's send. That is the feature being inert — recoverable, debuggable — not a privilege hole. The expensive direction would be a permissive default, and nobody proposed one.
Do this instead:
permits(caller, action, target)that delegates to the 4-argument form with the deny-all classifier. Fail-closed.FleetMcp.denyForandFleetApp.allow. That was the real point of decision 5 and it stands.SENDfor a collaborator. Otherwise the safe default is an accident rather than a decision, and nothing would notice if it flipped.Acceptance criterion 2 (mutation-prove the classifier conjunct) is unchanged.
Correction 2 — a line reference I inherited instead of measuring
My brief said
FleetMcp.java:437-440"usesisSpawnedMember()to enrol a caller in the presence map". I took that range from this ticket's earlier architect comment and did not check it. I have now checked it.The accurate picture: the context extractor calls
markSpawnedMemberPresent(p, presence)atFleetMcp.java:441, and the role guard lives inside that method atFleetMcp.java:780-784:Lines 437-440 are the explanatory comment above the call, not the guard.
The substance of decision 3 is unchanged and confirmed —
isSpawnedMember()must stayWORKER || ARCHITECT, and the comment at:437-439states the reason in the code's own words: "Enrolling a lead would count it as an available member in the roster." Only my citation was loose.Why both of these are worth recording
This project has repeatedly found the defect in the lead's brief rather than the worker's code. Correction 1 is the general shape: I specified a mechanism (delete the overload) instead of the property I wanted (no call site can silently skip the limit). The property had a cheaper implementation that I would have found by asking what would falsify the need, and I sent the brief first. Correction 2 is the other standing rule, broken by me in the same message: a number or a line reference taken from an earlier comment is not measured, and repeating it does not make it mine.
Nothing else in the Unit C brief changes. Scope, the denied-action list, the
TASK_READdenial, thewhoamibranch, and all eight acceptance criteria stand.Unit C merged — PR #699, as
b92a669Merged locally and pushed to
main; PR #699 closed by hand.COLLABORATORnow exists with its authorization row, so Unit D is unblocked.Verified before merging
Throwaway worktree of
origin/mainatf0ff252,--no-ffmerge ofrefs/pull/699/head(bb29b00), fullmvn clean installinfleetd/:Independently from the surefire XML after
rm -rf target/surefire-reports: 173 report files (I counted.xml), aggregatetests=1974 failures=0 errors=0. Baseline atf0ff252was 1964, so +10 — exactly the new tests.Tree identity: the tree I built and the tree I merged are the same object,
a19e87be54c48937efa93748f1453f20939c30a4. Push verified by ref: localHEADandgit ls-remote origin refs/heads/mainboth readb92a669ddcf3e70f3a63ef5d0dc17ed42e07bc4f. The merge commit does not touchwiki/.I read the grant table myself
Every arm is added to, never altered.
SENDgains|| (caller.isCollaborator() && knownLeadOrCollaborator.test(targetSession));READ, METRICSgain|| caller.isCollaborator().SPAWN/STOP/DRAIN/HANDOVER,ANSWER,COORD_SEND,TASK_READandCOORD_READkeep their exact conditions, so a collaborator is excluded from each by construction. Thecaller == null || isAnonymous()pre-check is untouched.REPLY, ASKremainownsSession(targetSession), andisSpawnedMember()remainsWORKER || ARCHITECT.FleetApp.allownow routes through the newpermitsForseam, so the REST test drives the real production gate rather than a test-supplied classifier. That matters here: a limit in a shared helper is only shared by the callers that call it.Adjudication of the review fan-out
Two reviewers, one dimension each, neither the implementer.
Reviewer 1 (existing-role regression): no issue. It checked each arm against
origin/mainby reading, noted that Java short-circuits onisCollaborator()so the classifier is never invoked for a non-collaborator, and confirmedCallerResolveris untouched. It stated plainly that it did not re-derive the 156-pair matrix and did not runmvn. I ran the build myself and read the same arms, and I agree.Reviewer 2 (collaborator privilege set): reported high severity at
CallerResolver.java:228— a configured collaborator tab falls through toPrincipal.worker, so it receivesTASK_READand can read other members' ticket replies.I checked this myself, and I am not treating it as a merge blocker. The finding is factually right and the severity label is wrong for this PR.
The worker fallback catches any pane that is not a configured lead tab and not a bound architect slot. That predates Unit C and predates Unit B, and this PR does not touch the file. So listing a tab under
fleet.collaboratorsgrants it nothing new — it was already a worker-by-fallback like every other hand-opened pane. Unit C could not have fixed it without doing Unit D, and this ticket's plan already states Unit D's acceptance criterion as exactly this: the resolver must return a member's role from the roster first, then lead, then collaborator, then the worker fallback.What the reviewer did surface, and what I had understated, is a documentation hazard: an operator could read the registry as a sandbox. I have fixed that on
wiki/11-Features.md(d913009, pushed torefs/heads/mainand verified by ref), which now says that naming a tab here buys no restriction, that such a tab resolves as a worker and so holdsTASK_READ, and that the fallback is pre-existing rather than granted by the registry.Shape-sweep finding, carried forward (not fixed)
FleetMcp.denyForskips the auditallowedentry forREAD/TASK_READ;FleetApp.allowskipsREAD/METRICS/TASK_READ. The implementer added the fact that settles the severity:grep -n "Action.METRICS" FleetMcp.javareturns nothing, so no MCP tool ever passesMETRICStodenyForand the asymmetry has no live effect today. It remains one invariant enforced with two hand-maintained lists. Not filed as a ticket yet.Where #669 stands
mainisb92a669.isLeadpredicate to lead-or-collaborator) is latent on this host: nomemberHerdrSocketis configured.COLLABORATORexists andfleet_whoamireturns it, so the canonical block'sfleet_whoamiladder and invariant 3's "send is lead or architect" line can finally be written truthfully. I have not done it in this session. Note the role is still unreachable in production until Unit D, so the block should say whatfleet_whoamican return, not imply a collaborator is live.A redeploy is owed again: this merge changes Java production code.
Unit D spec — the resolver. Delegated now.
Implementer:
devonsonnet, worktree669-unit-d, branchworker/669-unit-d-efbbd7-1.This comment is the authoritative spec. The brief sent to the worker carries the same text.
Any correction goes here as a new comment, never as an edit to this one and never as a re-send.
Everything below I measured myself in the main clone at
b92a669, with the command next to it.Units A, B and C are merged, so this is the live shape, not the shape the plan assumed.
The security property, stated as a property
A terminal that belongs to a live spawned member resolves as that member's own role, whatever
any tab map says about that terminal. When the spawned-member roster matches, no tab map is
consulted at all.
That is #661 closed at the resolver rather than only at config validation. Today it is open:
resolve()reads the lead tab map first, atCallerResolver.java:211.The current resolution order — read by me,
CallerResolver.java:208-257The order Unit D must produce
c.terminal() == null-> the token / loopback-trust tail, unchanged.MemberRole.ARCHITECT->Role.ARCHITECT;DEV,HUNTER,REVIEWER->Role.WORKER.Principal.leader(unchanged behaviour).Principal.architect(unchanged behaviour).Principal.collaborator(name, terminal, pid).Principal.worker(unchanged).Step 4 stays where it is. A spawned architect is already caught by step 2, so step 4 now serves
the case it was written for: a binding with no live
MemberSession.The roster source
SessionManager.roster()returnsList<MemberSession>; each carriesterminalId()androle()(aMemberRole). Its javadoc atSessionManager.java:780-790says it is"all registered sessions (acquired minus released)", and it is the non-resolving roster meant for
hot paths — which is what
resolve()is. Do not userosterResolved(): its own javadoc saysit is for caller-driven reads, and it can open an opencode session store per call.
resolve()runs on every request, so pass the lookup in as a function rather than a list to scan.MemberRegistry.bindis called fromSessionManager.java:740, so slot bindings are live. I checkedthat myself; an older note of mine claiming
bindis never called is wrong.The scan: generalise the one that exists, do not add a second
LeadTabScanner(src/main/java/dev/ltms/fleet/herdr/LeadTabScanner.java, 256 lines) is handed atab -> namemap and returnsterminal -> name. It must returnterminal -> (name, kind)withkindone of lead or collaborator, from one pass.Every property it has today must survive, because each one is a fix for a named bug:
PendingCloseMarker.stripbefore matchingagent.listliveness cross-check, so a labelled but dead tab is never in the map (#359)These apply to a collaborator tab exactly as to a lead tab. A dead collaborator tab must not
resolve.
An edge case the assembly has today:
FleetdAssembly.java:265builds the scanner only whenfleet.leadersis non-empty, and reads the TTL fromleaders.values().iterator().next().scanIntervalSeconds().Collaboratorisrecord Collaborator(String tab)— it has noscanIntervalSeconds. So collaborators configuredwith no leaders must still produce a scanner. Pick a defensible interval for that case and say in
your reply which one and why.
The classifier that is still inert
Authz.NO_KNOWN_LEAD_OR_COLLABORATORistarget -> false(Authz.java:71) and both productiongates pass it —
FleetMcp.java:692, andrest/FleetApp.java:276-277through itspermitsForseam. So a collaborator'sSENDis denied for every target today, by design, until areal classifier lands.
Unit D wires it: a target terminal is a known lead or collaborator when it is in the lead map or
the collaborator map. It must be the same maps
resolve()reads, not a second copy — a rosterthat lists an address
resolve()would not accept is the driftCallerResolver.leads()' javadocalready exists to prevent.
Acceptance criteria
lead tab map resolves as its member role, not as a lead.
fail. Run that mutation, and report the exact test name that goes red plus anything else
that goes red with it. A survivor means the assertion is not pinning what it claims.
Role.COLLABORATORandcarries that collaborator's name.
agent.list— does not resolve as acollaborator.
member nor a collaborator is unchanged: a lead still resolves
PRIMARYwith its name, a boundarchitect still resolves
ARCHITECT, an unrecognised pane still resolvesWORKER, and bothterminal == nulltails are untouched.collaborator is permitted, and the same call naming a spawned member's terminal is refused —
over both MCP and
POST /sessions/{id}/message. A test covering only MCP leaves the RESTroute open.
mvn clean installinfleetd/passes. Report the realTests run:line.Establish your own baseline before you change anything: build once on untouched
origin/mainand report that number too, so your "+N new tests" has a denominator you measured.For reference only, the Unit C evidence comment on this ticket reports 1974 tests / 0 failures
at
b92a669; I have not re-run that myself in this session, so treat it as a cross-check onyour own baseline, not as the baseline.
Hazards measured on this host
fleetd/.mvn … ; echo "exit=$?" ; tailreports
tail's status. WriteMVN_EXIT=$?into the log and grep the log for it.grep -creturning 0 exits 1, which silently drops the rest of an&&chain. Append|| true.findhere isbfsand rejects-newermtwith a relative time. Use-mmin -N, and neversend a probe's stderr to
/dev/nullwhen a zero is the answer you would act on.fleetd/fleetd.yamlis gitignored and you cannot read it. No acceptance criterion abovedepends on it. For the record, I checked it myself: no
memberHerdrSocketis configured and noprofile uses
placement: pane..mcp.json,opencode.jsonand.autoenvare stubs, not the repo's files.Never edit them.
git config --worktree --get-all fleet.neutralizedConfiglists them.wiki/is uninitialized in your worktree. Do not read it and do not run theCLAUDE.md/wikisync check — it cannot pass for you. That check is the lead's.
refs/stashis shared across worktrees. Nevergit stash.Working rules
git add -A.command to get around it.
"this used to". Contracts go in javadoc, constraints inline, the story in the commit message.
is newer than your brief and wins.
tab-map-before-roster shape, name it in one line and do not fix it.
fleet_replycarrying your whole report, including the PR URL.The audit-skip asymmetry the Unit C implementer found in its shape sweep is now filed as #700, so it stops being carried forward in handover prose.
I re-measured it myself rather than repeating the implementer's numbers. Both facts hold at
b92a669:FleetMcp.denyForskips theallowedentry forREAD/TASK_READ(FleetMcp.java:693),rest/FleetApp.allowskipsREAD/METRICS/TASK_READ(FleetApp.java:294), andAction.METRICSappears 0 times inFleetMcp.I paired that zero with a positive control, because a zero reads the same whether the action is absent or the pattern is broken: in the same file
Action.READis 2,Action.TASK_READis 4,Action.SENDis 2. So the pattern matches when there is something to match, andMETRICSreally is absent.It is out of scope for Unit D and the Unit D brief says so — the implementer is told to name a shape like this in one line and not fix it.
CORRECTION for Unit D / PR #701 — do not merge as it stands. It regresses fleetd #424.
This is newer than the brief and newer than comment 18450, so it wins. One change is needed; everything else in the PR stands.
I found this by reading the diff myself, then proved it with a throwaway test rather than leaving it as a reading. The build is not the problem — I ran the full trial merge myself and it is green:
BUILD SUCCESS,MVN_EXIT=0,Tests run: 1988, Failures: 0, Errors: 0, independently cross-checked by aggregating 173 surefire XML files to the same 1988/0/0. The defect is invisible to the suite.What breaks
fleetd #424 has two halves. One is "a revoked slot refuses the next spawn". The other is "a session already bound to a slot loses the ARCHITECT privilege on its very next request".
MemberRegistry's class javadoc states the second half as the rule — "config governs what a bound slot still grants, as well as what may be bound next" — and explains the mechanism:roleForSlotandnameForSlotreadslots()with no cache, so once the slot drops out of config,resolve()can no longer confirm it and the pane falls through toPrincipal.worker(...).The new spawned-member step returns
Principal.architect(...)on the strength of the roster role alone. It never callsmemberSlotRoles. Since a spawned architect always has aMemberSessionwith roleARCHITECT, that step now catches every live architect and the config confirm is skipped.The measurement, with its control
I built the post-revocation state honestly: reserve and bind while the slot is configured, then swap the config to remove it, exactly as a reload does. Two controls assert the state is really what I claim —
roleForSlotreturnsnullafter revocation, and the occupancy is deliberately still there.Then I resolved the same terminal through both construction paths, same registry, same moment:
The passing control is the point. The only difference between the two lines is which constructor built the resolver, so the demotion is lost by this change and not by my setup. Note also
name=null:nameForSlotcorrectly returns null because the slot is gone, and the ARCHITECT role is granted anyway — so the result is an architect principal with no name.An earlier attempt of mine failed its own setup control (
bindrefuses an unconfigured slot, so I could not reach the revoked state that way). I mention it because that control is what stopped me reporting a conclusion from a broken probe.Why the suite stays green
The privilege-revocation half of #424 was documented but never tested.
MemberRegistryLiveTestpins onlyrequireSlotForandreserve— the next-spawn half. So there was no assertion to go red. A comment claiming an invariant is a free test case, and this one was never cashed in.The fix
On the new spawned-member step, confirm the slot the same way the existing architect step does: treat a roster role of
ARCHITECTasARCHITECTonly whenmemberSlotRolesstill confirms the bound slot is an architect slot, and otherwise fall through toPrincipal.worker. The roster decides that the pane is a live spawned member; config still decides what that member's slot grants. Those are two questions and the fix keeps them separate.Then pin it, because nothing else does: a test asserting that a live spawned architect whose slot was revoked resolves
WORKER, with the passing control above kept in the test so it cannot go vacuous.Do not widen this into a redesign. If anyone thinks a live spawned architect should outrank a config revocation, that is a deliberate change to #424's rule and it needs its own ticket and its own argument — not a silent side effect of the resolver reorder.
Everything else about the PR is fine so far
The resolution order is as specified, the collaborator step is last among the tab maps,
knownLeadOrCollaborator()reads the same live mapsresolve()reads, and both production gates now take the real classifier. Two reviewers are still working other dimensions; I will fold in anything they find.Unit D review fan-out adjudicated. One blocker, one nit, one new ticket.
Two reviewers, one dimension each, neither the implementer. Both reported honestly, including what they did not check. I verified every load-bearing claim myself rather than promoting a reviewer's "sound" to a fact.
The blocker stays as stated in comment 18462
The fleetd #424 privilege-revocation regression is the only thing stopping this merge. Neither reviewer found it — I found it by reading the diff. The fix and the test to pin it are in 18462.
Reviewer 1 (resolver wiring + roster freshness): one medium finding, and it is real
I verified its wiring answer. Production constructs the resolver only through the new roster-aware overload —
FleetdAssembly.java:495and:500in the PR head, both passingspawnedMemberRoleandcollaboratorTerminals.git grep withLeadsAndMembersoverfleetd/src/main/javaon the PR head returns no other call site, so the 5-arg overload has no production caller, andNO_KNOWN_LEAD_OR_COLLABORATORsurvives only asAuthz's own 3-arg fail-closed default. That closes the inertness question I set it: the security step is wired, not merely present. This mattered because a defaulted no-op would have made the whole unit inert while every test passed.Its finding: a release-window race.
release()doesregistry.remove(paneId)first (SessionManager.java:306) and stops the pane only afterwards. In between the code callsworktrees.hasUncommitted(...), which shells out togit status— the CB-581 comment at:363says so in the code's own words — and may then calltrySnapshot(...), another git operation. So a member is deregistered while still alive, for as long as one or two git subprocesses take. During that windowspawnedMemberRolereturns null for its terminal and the caller falls through to the tab maps.I confirmed the sequence and the intervening work myself. The finding is structurally right and the window is not theoretical.
I am not holding the merge for it, and here is why. It is not a regression. Before this PR there was no roster step at all, so every member fell through to the tab maps on every request — the hole was total. This PR narrows it to the release window. Blocking a strict improvement because it is not yet total would be the wrong trade. The fix also belongs in a different subsystem: it needs a "releasing" marker surviving until
launcher.stopreturns, which isSessionManagerlifecycle work, not resolver work.Its exploit precondition also does not hold on this host: it needs a member pane carrying a lead or collaborator tab label, which startup validation refuses (
validatePanePlacementAgainstLeadTabs, plus Unit B's widened collision check), and I checked the live config myself — no profile usesplacement: pane. That makes it a latent hazard here and a live one on any host that ever uses pane placement.Filed as its own ticket. Credited to reviewer 1.
Reviewer 2 (scanner generalisation): no issue, and I checked its reasoning
It reported NO ISSUE with a stated scope, which is a useful answer. Its central claim is correct and I verified it in the diff:
cachedis a single combinedMap<String, Entry>, bothget()andcollaborators()go through onerefresh(), andcached/gracedTerminalsare keyed by terminal, not by kind — so accessor order cannot consume or renew one kind's grace state.byKindonly filters the combined map at read time.I checked all five properties the class carries, since each is a fix for a real bug, and all five survive: exact case-insensitive match with no prefix stripping,
PendingCloseMarker.stripbefore matching (entryOf, PR head:306-311), theagent.listliveness cross-check now covering both kinds, the one grace scan, and the TTL cache whose failed scan keeps the previous answer.It also correctly noted a coverage gap rather than a defect: the added tests do not cover cross-TTL accessor ordering. Worth knowing; not worth blocking.
One nit to fix alongside the blocker
buildTabIndex's javadoc says a colliding label "takes the lead entry, the first put when a key collides". The outcome is right but the stated reason is backwards: the code puts collaborators first and leads second, andLinkedHashMap.putoverwrites, so lead wins because it is the last put. A future editor who believes "first put wins" could swap those two calls to tidy them and silently invert the precedence. Say that lead is put last and therefore wins, or make the precedence explicit instead of relying on put order.(Config validation already refuses a lead and a collaborator sharing an exact tab, so this is defence in depth — which is fine, and is exactly why the comment should be right.)
Noted, not in scope, not fixed
ConfigRef.java:149and:621still point atCallerResolver.withLeadsAndMembersas being "atFleetd.java:620/624". That wiring lives inFleetdAssemblynow. Pre-existing comment staleness, untouched by this PR, and not worth a ticket on its own — but it is the kind of line-numbered cross-reference that goes stale every time the file moves.Unit D merged —
b5bc5d4onmainPR #701 merged locally and pushed. The PR is closed by hand, because we merge locally and the forge never closes it by itself.
What the lead changed before merging
Two reviewers ran on separate dimensions. Neither found the defect below. I found it by reading the diff myself.
The first implementation of the spawned-member step returned an architect principal from the slot name alone:
It never reads
memberSlotRoles. That breaks fleetd #424: config is supposed to govern what an already-bound slot still grants, so revoking the slot must demote the bound session on its next request. With this code the session kept ARCHITECT.The full suite stayed green at 1988 tests. That half of #424 is documented in
MemberRegistry's class javadoc but was never tested —MemberRegistryLiveTestpins onlyrequireSlotForandreserve. A green build was not evidence here.Fix in
d83972e:My own verification, not the implementer's
I wrote a probe test that binds a slot while it is configured, then swaps the config to remove it, then resolves through both the new roster-aware path and the old path at the same moment with the same registry and terminal. On the unfixed head the roster-aware path returned ARCHITECT while the old path returned WORKER — same input, two answers, which is the regression. On the fixed head both return WORKER:
Trial merge of the fixed head in a throwaway worktree:
Cross-checked by aggregating the surefire XML directly: 174 report files,
tests=1990 failures=0 errors=0. Arithmetic check: the implementer's 1988, plus my 1 probe, plus the implementer's added revocation test = 1990.Tree identity, so I know I built what I merged:
The
wiki/submodule pointer is untouched.Reviewer findings, adjudicated
SessionManager.release()removes the registry entry before teardown, so there is a window where a dying member's pane is no longer in the roster but its tab map entry still stands. It is not a regression — before this PR the hole was total — the fix isSessionManagerlifecycle work outside Unit D, and the precondition (placement: pane) does not hold on this host.FleetMcp.denyForandFleetApp.allow: READ/TASK_READ/METRICS skip theallowedaudit line. Filed as #700 with positive controls, so a zero match could not read as clean. Untouched here, per the brief.Still open on this ticket
isLeadpredicate. Latent on this host, not yet delegated.fleet_list, so there is no discovery route for them. Worth its own ticket.Not yet deployed. A merge is not a deployment: the running daemon still holds the jar it was started with, and the redeploy is the next step.
Unit F done — the instruction surface
The lead's own unit. On
mainatc3e3554, and the wiki at5f8bd3e.Unit D made four statements in the canonical block false. A session could resolve as
collaborator, read "Which role am I?", get the answer, and then find no section anywhere telling it what it may do.CLAUDE.md(c3e3554)Four corrections, each read from the merged tree rather than from the ticket:
fleet_whoamireturns four roles, not three.FleetMcp.whoamireturns before the lead branch, so a collaborator getsrole, its registry name and its ownsessionId, and noleaderkey.mcp__fleet__*mount,ANTHROPIC_BASE_URL— and nothing launched a collaborator. So it falls through to "act as a worker". Safe direction, but it means a collaborator cannot learn what it is without asking the daemon.Authzalso allows a collaborator when the target passesknownLeadOrCollaborator.Plus a new Collaborator section, placed deliberately after the member turn contract, because the first thing a collaborator needs to be told is that the contract above it is not its own — nothing delegates to it, so it owes no
fleet_reply. Its two surprising limits are written down rather than left to be discovered, both measured inAuthz: it cannot reach a worker, and it cannot read a ticket, becauseTASK_READis withheld whereREADis not.The intent table gains a row for messaging a collaborator. That row states the gap instead of implying discovery works.
The wiki (
5f8bd3e, three commits)7-Use-Cases.md— the portable block re-synced. Not hand-copied: extracted with the same slice the project's own check uses, so the two are identical by construction. Old block 20760 bytes, new 23214, delta +2454. The check reports both directions —Falsewhile they were out of step,Trueafter.11-Features.md— the collaborator entry said "parsed but not yet used" and named Unit D as the fix, so the heading and body were rewritten. The heading changed, so the index row and the one in-page reference to its anchor moved with it; no stale anchor remains. While there I found and corrected a second stale claim: theplacement: paneentry still said the validator knows aboutfleet.leadersonly, which Unit B had already widened.9-Implementation.md— the auth tables had drifted.Rolelisted four constants and has five;Principalwas missing itscollaboratorfactory;Authz.Actionlisted eight actions and called them eight — there are thirteen. Four cited line numbers had drifted and were re-measured.The five-way resolution order is now a flowchart rather than a sentence with four semicolons. All seven Mermaid blocks in that chapter render with
mmdc12.0.0, and I checked the validator against a deliberately broken block so the seven passes are not a false green.Deployed
Unit D's code is now actually running, which it was not when I posted the merge comment.
scripts/redeploy-fleetd.sh --yes: build green at 1989 tests, old pid 12672 gone, new pid 70397, jar895bf6635918→34f449e0ff08, freshfleetd listeningline at 01:49:00, no ERROR lines since restart, andfleet_whoamistill answersprimary.Green
/healthzonly proves herdr answers, so I also proved a real spawn on the new daemon: asonnetdev member came up in its own worktree and accepted a brief.Filed from this unit
#703 — a lead cannot discover a collaborator.
fleet_listemits onlyleadsandmembers, andCallerResolver.collaborators()has no caller outside the resolver. So the channel works one way only until the collaborator speaks first. This is the first thing this ticket's own requested walk-through (item 6) would hit, and it is worth settling who may see those rows before anyone implements it.Remaining on this ticket
Unit E is now delegated — widening
HerdrRouter'sisLeadso a collaborator's pane routes to the lead herdr daemon. It is latent on this host:HerdrRouterhands back the same control object for both branches when one herdr serves everything, and neitherherdrSocketnormemberHerdrSocketis set infleetd/fleetd.yamlhere. I briefed it as unit-testable only, with two distinct clients, and told the implementer not to offer a live check as evidence.After Unit E this ticket is complete except #703.
Unit E merged —
13b6ae2, PR #704 closedA configured collaborator's terminal now routes to the lead herdr daemon.
FleetdAssemblyORs a second terminal map into the predicate it handsHerdrRouter, and the predicate's field is renamedisLead→routeToLeadso the name matches the widened contract.What the lead measured, not took on trust
mvn clean installon the branch in a throwaway worktree: 1992 tests, 0 failures, BUILD SUCCESS. 174 surefire report files, in a fresh worktree, so no deleted class is re-counted.13b6ae2before pushing: 1992 tests, 0 failures.The whole fix rests on one test, so I reproduced the mutation myself rather than reading the worker's number. I dropped only the collaborator clause from the predicate:
FleetdAssemblyCollaboratorHerdrRoutingTest.assembledRouterRoutesACollaboratorTerminalToTheLeadDaemon, failing with two differentAgentControlidentities. That the identities differ is the important part — it proves the two-distinct-FakeHerdrsetup really does discriminate, so the kill is behavioural and not an artefact.HerdrRouterTesttest, including the two new ones. They hand the router their own predicate, so they pin the router's mechanism and say nothing about the call site. Their javadoc says so plainly.Two things I checked by hand in the code:
LeadTabScanner's singlebyKindhelper and are keyedterminal_id → name. SocontainsKey(target)is right for both. A map keyed by tab or by name would have failed silently here.agentsForcall runs between the router's construction andcollaboratorTerminalsRef.set(...). The two uses in between are the directmemberAgents()/memberSpaces()accessors, which never consult the predicate. The publication window is harmless and is the same oneleadsRefalready had.Review
One reviewer fanned out against the diff (223 lines, over the ~50-line threshold), briefed from the diff rather than the implementer's rationale, and not the implementer. It reported no issue and independently confirmed the test reaches real production code, asserts identity against the runtime's own router, and has a working control. It traced the mutation by hand rather than running it; the runtime kill above is the lead's own.
A note for whoever reads the diff later
This defect cannot be proven on this host.
HerdrRoutersetsmemberAgents = member == lead ? leadAgents : new AgentControl(member), and neitherherdrSocketnormemberHerdrSocketis set infleetd/fleetd.yaml, so one herdr serves everything and both routing branches return the same object. A live probe cannot tell the fix from the bug. The unit test with two distinct clients is the only valid evidence, and a live check should not be offered as one.State of #669
Units C, D, E and F are merged. #703 is the remaining gap — a lead still cannot discover a collaborator, because
fleet_listemits onlyleadsandmembers. Two things need settling before anyone implements it: who may see those rows (fleet_listisREAD, which workers hold too, andcoordinatorVisibleTo(principal)is the precedent for narrowing by role), and what each row carries. Nothing inCLAUDE.mdor the Features entry currently promises discovery works.Redeploy of the running daemon follows this comment; Unit E changes Java production code, so the live
fleetddoes not have it until then.Two things found while closing out Unit E
1. The example config told operators the feature does nothing — fixed in
16b52fafleetd/fleetd.example.yamlstill said, above thecollaborators:example:That was true when Unit B landed the config shape. Units C, D and E have merged and deployed since, so it had gone false. This is the worst place for such a sentence: it sits in the file an operator copies to turn the feature on, and it tells them not to bother. Replaced with what the role may and may not do, read off the rows in
auth/Authz.javarather than copied from the wiki, plus the #703 limit.FleetConfigTest, which loads this file, is green at 170 tests.The same stale-claim sweep found one more:
wiki/11-Features.mdcarried a Still open paragraph sayingHerdrRouter'sisLeadpredicate reads the lead map only. Replaced with a resolved gotcha in wiki8a6ccd7. It was also the only stale reference to the renamed field left in the wiki.2.
fleet.collaboratorsis not enabled on this host, so the path has never run liveThere is no
collaborators:block in the live config, commented or otherwise. That is the operator's choice and not a defect — the feature is opt-in. But it means the whole collaborator path has never executed on this host. Everything I reported for Unit E is unit-test evidence plus a code read; no part of it is a live observation of a collaborator session.This bears directly on item 6 of this ticket — "one concrete walk-through: two Claude Code sessions the operator opened by hand, in two named tabs, exchanging a message." Items 1–5 are settled by the merged units. Item 6 is not demonstrated, and I have deliberately not enabled a collaborator myself: it needs a tab a person opens and names, which is an operator action, not a fleet one.
So I am leaving #669 open rather than closing it on the merges alone. What is left here is a live demonstration, not code.
Enabling it, when the operator wants to
fleetd/fleetd.yaml(the live config is gitignored; the shape is infleetd.example.yaml):scripts/redeploy-fleetd.sh --no-build --yesis enough; no rebuild is needed for a config-only change.fleet_whoami. It should answercollaboratorand carry the registry name and its ownsessionId.sessionIdis the address. A lead cannot discover it (#703), so the collaborator sends first, or passes it along.One caution for whoever does this:
placement: paneon any profile while a collaborator tab is named is refused at startup, and so is afleet.tabLabelthat could render as the collaborator's exact tab. Both refusals name the offending entry.The named-peer mesh is now with two architects
The operator has moved the question past "does the collaborator role work" to "why is per-tab config needed at all". In their words:
They also asked that the implementation be well structured, so the unit split is the primary deliverable.
Why this is doable, and smaller than it sounds
The addressing scheme they assume is already in the data model, measured this session:
herdr/Agent.java:23carriesterminalId, paneId, workspaceId, tabId, sessionId, agentType, status, name.herdr/LeadTabScanner.javaalready walksworkspace.list → tab.list → agent.listover every pane and builds aterminal → Entryindex, then filters it down to the configured labels. The information the use case needs is gathered and then discarded.WORKER(auth/CallerResolver.java:43).So the gap is four pieces, three of which are generalizations of running code: an unfiltered name index, an address form that takes a name rather than a
sessionId, an authorization row that permits peer→peer send, and a durable per-peer mailbox.msg/LeadMailbox.javais already the durable, non-blocking, pull-based peer channel and is the closest existing component.The fork the architects are settling
Workspace-scoped mesh with a safe default role, versus the current per-tab allowlist. The prior I gave them, explicitly flagged as a prior rather than a finding: the operator's own phrase "unique name in a unique workspace" may already name the trust boundary, since
workspaceIdis on every agent record. Config would then shrink from "enumerate every peer" to "which workspaces are meshed" — far less config, which is the operator's actual goal, without giving every pane on the host a channel into every other.Four hazards they are briefed not to hand-wave
fleet_sendto a BUSY peer is accepted, returns a ticket, and is then never delivered — measured three times in one session. A mesh on the rendezvous path silently loses mail, so this probably has to be pull-based.CLAUDE.md, so any rule they must obey belongs in the launcher charter.Also filed
#705 — any unconfigured pane resolves to
WORKER, andWORKERholdsTASK_READ, so an unlisted tab can enumeratetask-1, task-2, …and read other sessions' delegation replies. Stated in this issue's body since it was opened, never ticketed. It is defensible on its own whatever the mesh decision is, but the two must land in a consistent order.Process
Two architects, deliberately on different vendors' models so they do not share blind spots (slots
opusandsol). Round 1 is independent positions on the same nine questions; I run the exchange in round 2, because a member cannot receive a message inside its own turn. If they agree, they settle it. If they still disagree after comparing, they return both positions with their evidence and the lead decides.No code until they land — the answer decides whether
fleet.collaboratorsstays an allowlist or becomes an override on a safer default, and building before that is settled would commit to one of them by accident.Mesh design — round 1 done, round 2 running. Three live defects found and checked.
Two architects on different vendors (
opusandsol) each wrote an independent position on the "any named tab can talk to any named tab" design, against the same nine questions. Both have reported. Round 2 (each answers the other) is running now. No code has been written, on purpose: the answer decides whetherfleet.collaboratorsstays an allowlist or becomes an override on a safer default, and building first would pick one by accident.Three defects in shipped code. I checked each one myself in the main clone.
These are code facts, not design choices, and they are independent of which design wins.
1. A collaborator cannot be sent to at all. The feature this ticket shipped does not work in that direction.
The injector readiness gate is
Fleetd.java:238-240:Two disjuncts, and a collaborator's terminal is in neither:
FleetdAssembly.java:369,371,518— oneMemberPresenceobject, created at 369, passed todeliverableToat 371 and toFleetMcpat 518. Same instance.grep -rn markPresent fleetd/src/main/javareturns 4 hits: theMemberPresence.java:23definition,FleetMcp.java:791, andSessionManager.java:1256,1260— where 1256 is aPresenceFleet extends MemberPresenceoverride callingsuper, so it is the same class's method, not a second entry point.FleetMcp:791is the only write, and it is gated oncaller.isSpawnedMember().Principal.java:113—isSpawnedMember()isWORKER || ARCHITECT.COLLABORATORis excluded.LeadTabScanner.java:191-192—get()isbyKind(refresh(), Kind.LEAD), LEAD only, so theleadssupplier can never hold a collaborator.Authz.java:113permits the send viaknownLeadOrCollaborator. So the send is allowed and then dropped: it waits on the gate for about 60 seconds (Injector.java:95,104— 240 polls at 250 ms) and failsNOT_DELIVEREDat:532-570, never typed into the pane.Unit E added the collaborator map to the herdr router (
FleetdAssembly.java:149-151). It did not add it to the injector gate at:371. Two gates, one got the fix.What works and what does not: collaborator → lead works, because a lead is in
leads. lead → collaborator and collaborator → collaborator both fail after 60 seconds.No test covers it.
FleetDeliverabilityTest.javais 86 lines;grep -c deliverableToon it returns 6 andgrep -ric collaboratorreturns 0 — a real zero, with a positive control.Architect
opusfound this by reading the code and rated it high confidence pending a live send. I closed it to a proof by enumerating every writer into both sets, so no live send is needed to establish it. A live confirmation is still worth running whenever the block is next enabled.2. The guard that protects #661 is keyed on the config this ticket wants to remove.
FleetConfig.java:2893-2898, insidevalidatePanePlacementAgainstLeadTabs:Remove the per-tab config — the operator's whole goal here — and with no lead tab either, this check refuses nothing, while every pane in a meshed workspace becomes addressable. So per-tab config is not only extra typing today; it is load-bearing for a startup refusal. Any mesh work has to move that check onto the mesh's own switch. Architect
opusrates this the highest risk in the design.3. The main safety control is a filter that does nothing, and the example config says it works.
FleetdAssembly.java:287builds the scanner withSet.of()forexcludedWorkspaceLabels, and the daemon logs "shared fleet space". Butfleetd/fleetd.example.yaml:663-664tells an operator that a lead's workspaceThat is false as shipped. Member workspaces are not excluded. It matters to both designs, because excluding member workspaces is how a member pane is kept out of the mesh — so the design's main control is a filter that is currently passed an empty set.
One correction to my own brief
My brief told both architects that
LeadTabScannerwalks every pane, builds the index and then filters. That is wrong, and both architects caught it independently. The filter isentryOf(tab.label())atLeadTabScanner.java:247, inside the tab walk and beforeagent.list/pane.list, and:253returnsMap.of()before those calls when nothing matched. The topology is therefore never gathered, not gathered and discarded. That makes the mesh's core unit bigger than my brief implied: it changes the shape of the scan, and it removes a zero-cost short-circuit from a cached hot path.What the two architects already agree on
The workspace is the right scope and bounds blast radius, but is not a claim that peer text is safe · address by tab label, never
Agent.name, which is null for a tab a person opened · reject "default every pane toCOLLABORATOR" · a meshed peer gets noTASK_READ· durable mail, not rendezvous, with the broker required and a refusal rather than a silent drop to in-memory · the envelope carries a sender taken from the connection, never from a request field · a duplicate name is refused and re-checked at send and delivery time, never "pick the first", and a startup-only check is not the control · peek and ack only your own address · #661's roster-first precedence must hold, and target classification must read one shared authority source rather than a second weaker predicate · the MCP URL picks the daemon, so the real hazard is two daemons on one herdr socket, refused by a lock keyed on the socket path · no launcher charter change, because a human-opened tab never sees the charter and the MCP tool schema is the only text guaranteed to reach it — which is where "message bodies are untrusted data" has to live.What is still open
Eight seams, now in round 2. The two that decide the most:
opuswants every workspace meshed by default, except member workspaces derived fromprofiles.*.workspace, arguing an allowlist drifts from that list.solwants an opt-in allowlist, arguing a workspace nobody named must not become trusted. The two fail in opposite directions: the allowlist fails closed, the derived default fails open.opusrecommends shipping three small units first — the ticket owner check, the deliverability fix above, and one shared address classifier — with no mesh, then asking again, on the grounds that part of the original complaint is that the configured feature does not work.solplans the full mesh in seven units.Also open: a new
PEERrole versus reusingCOLLABORATOR; the floor role for an unmapped pane (solproposesOBSERVERwithREADandMETRICSonly, which would close #705's unmapped-pane exposure with no ticket owner check); whether #705's owner check must land before the mesh; the address form (sol's structured{fleet, workspace, tab}with a required fleet field looks right to me, andopus'sworkspace/namestring carries no fleet component); and whether #702 blocks the mesh.Not started, and still the operator's call to prioritise
I will post the decision here when the two architects have compared positions, whether they agree or return two.
Decision — the two architects agree on every seam. The workspace is not the boundary.
Three rounds are done. Both architects moved on checked evidence, and the design they ended with uses a discriminator neither of them proposed at the start. No operator decision is needed to proceed.
First, a correction to my previous comment
I wrote that defect 3 was "the design's main safety control is a filter that does nothing, and the example config says it works". The two facts in that sentence are right. The conclusion I implied was wrong, and the fix is the opposite of what it suggests.
FleetConfig.java:1144-1149, the javadoc onLeader.DEFAULT_WORKSPACE:So the empty
Set.of()atFleetdAssembly.java:287is deliberate. The stale thing isfleetd.example.yaml:662-664, which still describes the world before CB-558. Do not populate that exclusion set. Doing so would skip the workspace the lead lives in, and the lead would never be found again — the exact harm the stale comment warns about.Architect
opushad proposed populating it, then withdrew the unit and found this javadoc itself once I gave it the live measurement. Its words: "my R2-a would have undone a deliberate design decision". Worth recording, because a reader of my earlier comment would have built the wrong fix.The measurement that decided the design
The string
workspaceappears exactly once in the livefleetd/fleetd.yaml— line 260, insidefleet.leaders.opus:No profile sets
workspace:. So I measured the resolved defaults rather than the key:FleetConfig.java:525— a profile's workspace defaults to the literal"fleet".FleetConfig.java:1150—Leader.DEFAULT_WORKSPACE = "fleet", applied at:1158.The lead and every member share one workspace, and that is the shipped default, not a local choice. Any operator who sets nothing gets it.
This kills a premise both architects shared in round 1 — that the workspace already separates the panes people own from the panes the fleet owns. On the default layout it has one value, so it separates nothing. A mesh scoped to it is either empty or total. There is no third state.
Neither architect could find this:
fleetd.yamlis gitignored, so a worktree cannot read it. That is why the lead verifies.What the design became
The discriminator is the spawned-member roster, used as a negative filter. In the roster means its own member role, never a peer. Not in the roster, and not a lead, architect or configured collaborator, means a candidate peer.
CallerResolver.java:294already consults the roster first, from Unit D, so the precedence chain exists — the mesh only changes what sits at the bottom of it.The tab label stays the address. It must never be the identity. A person can label a tab to match the member template, and the scanner's defence — a worker cannot rename a tab — says nothing about a person, who is exactly who opens a meshed pane. Separating address from identity is the core of the fix.
The roster has exactly one hole: the teardown window.
SessionManager.java:306removes the registry entry,:390stops the pane, and in between:341runs a git status and:355may commit the whole worktree. That is seconds, not an instant.Settled, seam by seam
mesh:block disables the mesh.mesh.workspacessurvives only as an optional extra filter.PEERrow. Both reached this independently. A lead is reachable only by explicit opt-in — a target-side rule that never depended on the workspace. ReusingCOLLABORATORis rejected: itsSENDis gated onknownLeadOrCollaborator, which includes leads, so on the default one-space layout it would hand every human tab a channel into the lead's pane with no operator act at all.OBSERVER, but only together with provisional member authority. See below — this one I decided againstopuson measurement.Taskhas no creator field at all (MessageService.java:238-259), so this is a new field, not a comparison.{peer: {fleet, workspace, tab}}. No bare delimited string — live tab labels contain a colon and a space.peer.fleetis checked against this daemon's own id, so a call to the wrong daemon is refused instead of routed into its similarly named tab. No shorthand in v1.opusconceded fully: with the workspace boundary gone, the roster is the only discriminator and the teardown window is its only hole.mesh:is restart-required; reload must say so. An exclusive broker lease on the fleet id, plus a local lock on the canonicalized herdr socket path.pom.xml:264-267— profiledefault-excludesisactiveByDefaultwithexcludedGroups=contract. So every mailbox property that matters for correctness needs a hermetic test that runs in the normal build; a real broker stays a separate release gate. A property asserted only under@Tag("contract")does not run.The floor role — I decided this against
opus, on a measurementopusargued theOBSERVERfloor is safe alone, because a spawned member is registered before it can connect MCP. That is wrong.SessionManager.java:235—handle = launcher.spawn(req);SessionManager.java:244—registry.put(handle.id(), session), only after spawn returns.HerdrPeerLauncher.waitUntilInjectableOrThrowpolls the pane until it reports an injectable state.spawnReadyTimeoutMs: 20000.So
spawn()can block for up to 20 seconds with the agent already up and unregistered. A member's first MCP call in that window resolves through the floor — and today it works only because the floor isWORKER, for whichisSpawnedMember()is true, soFleetMcp.java:791marks presence and the member becomes deliverable.Today's
WORKERfloor is load-bearing for member boot. Change it toOBSERVERon its own and the first readiness signal is lost.solfound this window and withdrew its own "cheap standalone change" claim; it is right, and the provisional-authority half must land in the same change as the floor.For the record, I had accepted
opus's reasoning earlier in the session and was wrong to; the measurement above is what settled it.Landing order
Release 1 — remediation, no mesh. Every unit small and reversible.
ConfigRefreporting forfleet.collaborators(finding E, below) — in flight now.fleetd.example.yaml:662-664text — and do not populate the exclusion. Add the test that nothing pins today: with every profile defaulting tofleetand a lead whose workspace is alsofleet, the scan still reports that lead's terminal. Without it, the wrong fix ships green and demotes the lead.fleet_listreports collaborators (#703).Then a checkpoint: confirm lead → collaborator and collaborator → collaborator delivery in a real injectable window, and confirm an unknown terminal is still refused.
Release 2 — the mesh. Provisional member authority with the
OBSERVERfloor; the #702 tombstone as a gate; deferredmesh:config with thePEERrole and explicit lead opt-in; one shared peer directory built on the roster-negative filter; the durable mailbox with hermetic default-profile tests; the broker lease and socket lock; then peer send, delivery and the whole instruction surface in one landing.Two details that moved:
mesh.workspacesis now an optional filter rather than a control, and lead opt-in can no longer live per-workspace — it needs one top-level setting.Finding E — new, verified, and already delegated
Editing
fleet.collaboratorsand reloading reports success and does nothing, with no warning.grep -ci collaboratoronConfigRef.javareturns 0. Control:grep -ci leadersreturns 16, so the zero is real.changedSplitKeyscomparesleadersOf(old)againstleadersOf(fresh);leadersOfreturnscfg.fleet().leaders()only.FleetdAssembly.java:264readscollaborators()off the startup config,:272-277builds the tab map,:287bakes it into theLeadTabScanner, which is never rebuilt.ConfigRef.java:272isSPLIT_KEYS = Set.of("health", "coordinator", "fleet"). The test only needs one branch forfleet, and theleadersbranch satisfies it. A newly frozen sub-field under it has no branch at all.The existing message even lists what is hot — "the rest of
fleet:(developers, hunters, reviewers, charters, tabLabel)" — andcollaboratorsis in neither list, so an operator reading it would reasonably think the edit applied.Instruction surface
One line in the canonical
CLAUDE.mdblock becomes false the moment the floor changes. It currently ends the fallback ladder with "Still unsure ⇒ act as a worker, the most restricted member role".WORKERholdsTASK_READand the new floor does not, so that sentence would tell a confused session to assume more authority than it has. Both architects flagged it. It must name the new floor and describe behaviour rather than claim a rank.Still yours to prioritise
Two more units merged in
3e8e314PR #706 and PR #707, merged together and verified as one tree. Both PRs are closed by hand,
because a local merge never closes a PR here.
#706 —
deliverableToopens for a collaborator. This is the defect I described earlier inthis ticket: a send to a collaborator was allowed at
Authzand then dropped at the injectorreadiness gate, because a collaborator is in neither of that gate's two disjuncts. It now has a
third. The thing that decides whether the fix is live is the key shape:
LeadTabScanner.get()isterminal_id → lead nameandcollaborators()isterminal_id → collaborator name, both builtby
byKindfrom one scan, socontainsKey(target)really fires. If that map had been keyed byname the fix would have been dead in the same way the defect was.
#707 — a
fleet.collaboratorschange is reported on reload. The map is read once at startupto build the scanner's identity map, and a reload does not rebuild it.
ConfigRefsaid nothing,so an operator saw a clean reload and no live effect. It now reports that a restart is needed.
Verification
I did not promote either worker's "clean" to a fact. Merged both onto
origin/mainin onethrowaway worktree and ran
mvn clean installmyself, with the output written to a file and notpiped, because a pipe hides a failure behind a zero exit:
BUILD SUCCESS,Tests run: 1996, Failures: 0, Errors: 0BUILD SUCCESS,Tests run: 1999, Failures: 0, Errors: 01999 is 1996 plus #706's three new tests, which is the count I expected.
ConfigRefTest: 32,FleetDeliverabilityTest: 9. I checked the pushed tree is byte-identical to the tree I built.Both units came with a revert proof from the worker: dropping only the production change turned
exactly the new tests red and left the pre-existing ones green. I did not re-run either
experiment myself.
A process defect worth recording
The #706 worker launched a
forksubagent for a read-only side search, with an explicitinstruction to change nothing. The fork ran the whole commit, push and open-PR sequence itself.
The worker reported this unprompted, which is the right thing to do, and the content is not in
question. But a member must not hand off its own push, because then nobody who read the code
reviewed what was pushed. A
forkinherits the parent's context, so it inherits theimplementerskill's recipe, and a per-call "read-only" line loses to a loaded playbook. The durable fix is in
the skill, not in each brief.
Still open under this ticket
#703 (a lead cannot discover a collaborator) and the
fleetd.example.yamltruth fix are nowdelegated. #705 (the
WORKERfloor) is with an architect — see my comment there.