Any unconfigured herdr pane resolves to WORKER, and WORKER holds TASK_READ — so an unlisted tab can walk every ticket and read other sessions' delegation replies #705
Open
opened 2026-10-04 05:35:35 +02:00 by ltms
·
7 comments
No Branch/Tag Specified
main
worker/702-4f5c7f-2
worker/715-5c43fc-1
worker/721-70f9ea-5
worker/718-99362b-2
worker/task-15-af0d10-12
worker/task-16-50a702-13
worker/task-12-4d0479-9
worker/task-13-823ce2-10
worker/705-ticket-owner-af9928-8
worker/703-list-collaborators-9c06c2-7
worker/669-example-truth-0b303d-6
worker/669-collab-deliverability-9ba859-3
worker/669-collab-reload-report-2a21bd-4
worker/669-7e80a6-1
worker/669-unit-d-efbbd7-1
worker/669-1b786a-1
worker/669-1d1d9f-1
worker/692-4afb9d-2
worker/689-02fced-13
worker/693-cf23fa-14
worker/677-fix-lead-collision-f69073-12
worker/638-fix-overmask-dbb1bf-11
worker/675-5b7478-4
worker/669-unit-a-70cc8f-3
worker/677-8cdaaf-5
worker/638-a7b391-1
worker/683-4536d6-2
worker/651-a75bbe-8
worker/680-20607d-7
worker/664-c12e95-3
worker/668-08534d-4
worker/672-0f2469-2
worker/670-7d1022-1
worker/661-ac7c28-2
worker/664-37fb9b-3
worker/663-remove-3arg-read-3f6783-1
worker/659-remove-dead-backcompat-ba5e6f-1
worker/637-revision-60a488-23
worker/656-redact-regression-tests-892903-19
worker/637-context-gauge-threshold-466eb5-16
worker/639-redact-line-numbers-de4ac4-17
worker/641-set-reformat-guard-6f96a4-18
worker/642-herdr-guard-scope-5de0e4-15
worker/650-javadoc-scope-95f3b3-14
worker/612-01e9f7-13
worker/612-a-r4-quarantine-outage-7ab0e8-5
worker/612-a-r9-r11-capacity-coverage-peers-cfcc79-7
worker/612-a-r10-loophealth-ccc872-8
worker/612-a-r12-turnregistrar-9e3bb7-9
worker/612-a-r5-leadconfigdir-9e70cf-6
lead/config-edit-redact-anchor-wording
worker/config-edit-seam-ca8dc1-1
worker/612-r67-630-lifecycle-290b8d-3
worker/629-625-ports-seams-da7d5d-4
worker/612-r12-exhaustion-f37cd7-1
worker/612-r38-amqp-24b083-2
worker/fleetd-612-unita-87807e-1
worker/612-b3-mcpwirings-da2b58-3
worker/612-b2-cb185-176d3a-2
worker/612-b1-completion-457459-1
worker/612-agaps-73a926-2
worker/608-sleeps-3a64ff-3
worker/621-b4520b-1
worker/618-b83894-2
worker/fleetd-615-e05481-5
worker/lead-autocompact-5f1ab2-3
worker/fleetd-613-f85deb-3
worker/fleetd-608-flaky-nudge-test-d0c2d1-3
worker/lead-context-gauge-ad404f-1
worker/gauge-wiring-9158c1-4
worker/redeploy-slowstart-ead0e5-5
worker/charter-bytes-13668c-6
worker/rollover-outcome-291483-2
worker/589-f64303-2
worker/593-1a8025-5
worker/589-fcd2aa-1
worker/568-9fdaa2-3
worker/571-attempted-outcome-5739f7-2
worker/581-completionresolver-cas-sites-0542b7-6
worker/562-loop-health-wiring-test-99611c-5
worker/562-surface-loop-health-7df5cc-4
worker/575-waiter-cleanup-sites-62ad80-1
worker/572-answer-lock-release-46a9ae-5
worker/567-probe-channel-leak-a38fc5-6
worker/551-record-before-send-7cbf56-1
worker/561-listener-fanout-survives-a-throw-61d538-2
worker/555-redeploy-main-flow-seam-65c2f5-2
worker/556-injector-owns-registration-e027a5-1
worker/552-post-restart-mktemp-abort-bc2672-4
worker/553-onstatus-completion-leak-0da881-2
worker/550-shasum-linux-196132-1
worker/538-loop-dies-on-error-4a5eeb-6
worker/426-health-coverage-ef1fd4-4
worker/504-failed-reported-clean-3cfd66-3
worker/537-capturedlog-close-e4c437-2
worker/459-broken-link-targets-cadc17-5
worker/535-appender-leak-fe74c1-1
worker/512-part2-shutdown-detection-434701-9
worker/529-logger-level-sweep-2a5533-8
worker/528-drain-gate-call-site-5de83d-7
charter/forge-mcp-vs-token
worker/521-swap-guard-unpinned-28e931-5
worker/519-probe-test-harness-d25ab8-4
worker/525-logger-level-leak-1b4eb0-6
worker/518-fleetmcp-resolver-wiring-8ef96c-1
worker/512-drain-complete-line-7edd71-3
worker/517-abort-branch-and-jar-id-41b641-2
worker/500-9e52c9-3
worker/509-4912f4-2
worker/511-9a4b23-1
worker/493-479f45-2
worker/505-03f8b2-1
worker/492-followup-detect-unclear
worker/501-a31fa0-7
worker/498-451d1c-5
worker/494-1015ce-2
worker/492-209647-1
worker/489-001902-2
worker/480-relative-handover-path-906323-1
worker/480-b-handover-skill-45bf1f-5
worker/474-followup-source-pin-f54a55-17
worker/474-charter-check-on-reload-f54a55-17
worker/466-quarantine-repeatcount-report
worker/393-opencode-skill-seeding-71854b-13
worker/469-canonical-tool-names-2a472a-16
worker/466-quarantine-escalation-5ae9c1-15
worker/446-hot-exhausted-pattern-0af580-6
worker/464-charter-tool-name-guard-a85635-12
worker/463-listfleet-default-fails-open-f1c76c-11
worker/458-invariant-5-by-purpose-862f9a-10
worker/439-coordinator-row-gate-bc032a-8
worker/449-herdr-protocol-576015-4
worker/450-abstract-spawn-599e1c-5
worker/437-ack-refuses-177d91-1
worker/444-placement-window-feb56a-2
worker/440-helddurable-derived-d462d7-13
worker/425-rework-placement-resolve-c58ba1-9
worker/421-lead-peek-held-msgs-cdbad2-10
worker/435-fixed-policy-cap-fe11de-12
worker/422-gate-state-observability-9e79d6-11
worker/431-memberregistry-live-readers-cdbad2-10
worker/424-architect-slot-hot-038b41-7
worker/422-model-gate-spawn-c29f48-6
worker/425-default-profile-live-f55534-8
worker/415-coverage-wording-2cbf9c-5
worker/416-3ad1da-1
worker/418-588283-3
worker/deterministic-stamp-race-409-3cb7b6-10
worker/armed-reads-live-config-404-ed931f-9
worker/reply-peer-refusal-391-5a34bd-7
worker/models-allowlist-aa9e9b-3
worker/ttl-stamp-race-399-f1122f-8
worker/scrub-receipt-400-316b3e-5
worker/exhaustion-detection-395-105105-6
worker/scrub-abort-394-316b3e-5
fix/scrub-uid-abort
worker/task-scrub-517574-2
worker/t386-clock-bd5b78-4
worker/t384-scrub-813790-5
worker/t381-cc-748314-2
worker/t373-336973-2
worker/t365-3920c5-3
worker/t358-6e989b-1
worker/t355-8b321c-1
worker/fleetd-369-hermetic-git-tests-e8b19a-3
worker/fleetd-368-stale-lead-binding-f5682e-2
worker/fleetd-360-deploy-units-0d3793-1
worker/359-dead-lead-tabs-f1253b-4
worker/362-worktree-skills-c03e51-3
worker/361-coord-visibility-655144-1
362-plugin-visibility-and-drift
worker/errscan-bed2ca-2
worker/amqp-log-identity-bed2ca-2
worker/withdefaults-guard-561704
worker/sleepguard-82076d-1
worker/fd334-9ee1b6-5
worker/fd348-f1ab27-4
worker/fd335-a71c35-1
worker/fd342-174a17-2
worker/fd345-490d0f-3
worker/fleetd-337-5ec7d4-21
worker/fleetd-341-af5a6b-24
worker/fleetd-339-5ca0a2-23
worker/fleetd-338-83a4a1-22
worker/fleetd-333-281f46-18
worker/fleetd-329-11bdbb-16
worker/fleetd-330-2770fb-17
worker/fix-326-50506e-15
worker/fix-324-3e9bbf-14
worker/fix-323-b8287d-13
worker/fix-316b-bd0860-11
worker/fix-318-76ca36-9
worker/fix-317-486aec-8
worker/fix-315-ce47c5-6
worker/fix-307-275890-6
worker/fix-308-b4f664-7
worker/fix-309-ec3939-8
worker/fix-310-7a3974-9
worker/fix-302-52ad0e-9
worker/fix-298-ce1acb-8
worker/fix-297-66bd11-7
worker/fix-296-104622-6
worker/fix-293-bare-closetab-eb22b5-3
worker/fix-280-gone-ask-lapse-bca98e-2
worker/fix-290-reapidle-guard-coverage-9b0dd1-1
worker/fix-285-trust-seed-8f3565-10
worker/fix-284-backend-error-seat-85912c-11
worker/fix-282-chained-ask-e6d0bb-8
worker/fix-283-teardown-leaks-f40dfa-9
worker/fix-281-pin-handler-actions-4921ac-7
worker/audit-rendezvous-lifecycle-d072ae-2
worker/audit-health-placement-1a2476-6
worker/audit-teardown-exits-e207a5-3
worker/audit-launcher-asymmetry-27e370-4
worker/audit-rest-authz-6ca53c-5
worker/investigate-275-abandon-asking-fdef52-8
worker/fix-274-worktree-leak-b0095d-7
worker/fix-273-exhausted-pattern-9665b5-6
worker/fleetd-267-model-check-bd8068-1
worker/fleetd-131-archunit-18b834-7
worker/fleetd-266-sshagent-rename-a014ff-6
worker/fleetd-184-uid-claim-8e1f31-4
worker/fleetd-184-warn-b381ee-10
worker/fleetd-184-docs-be1d12-9
worker/fleetd-257-9bf010-7
worker/fleetd-103-23a113-6
worker/fleetd-247-342356-5
worker/fleetd-116-04dea8-4
worker/fleetd-252-a830e0-3
worker/fleetd-111-7e8673-9
worker/fleetd-155c-f8ef4b-8
worker/fleetd-176-b928ca-3
worker/fleetd-249-7a7878-2
worker/cb248-composition-root-b-9acdf7-15
worker/cb148-envrc-default-fa6c82-12
worker/cb201-unit5-wiring-6c12e6-8
worker/cb241-fallback-echo-1175e9-11
worker/cb149-trust-dialog-2392a5-9
worker/cb134-148-overlay-visible-c9b986-10
worker/cb234-session-id-keyed-04e1fc-1
worker/cb201-unit3-nudge-abdf5c-6
worker/cb201-unit2-policy-c1102c-5
worker/cb201-unit4-outcome-a13bfa-7
worker/cb201-unit1-classifier-91b9b1-4
worker/cb201-227-refine-831980-3
worker/cb175-model-readback-0f085f-1
worker/cb222-charter-tmpdir-17f013-1
worker/cb226-architect-slot-race-cd3aa8-3
worker/cb224-worktree-root-group-024523-2
worker/cb-123-role-demotion-c600f7-2
worker/cb-219-opencode-roots-1f677e-1
worker/cb214-claude-session-id-b9eab4-4
worker/cb213-zdotdir-wrong-process-dd6de4-3
worker/cb211-exhaustion-classification-9546e0-2
worker/cb137-ambiguous-task-4df3d8-4
worker/cb209-agentsessionid-4dfdb6-2
worker/cb185-hostenvnames-2692b5-3
worker/cb206-opencode-sqlite-128718-2
worker/cb185-worktree-group-fc0c99-1
worker/cb-137-ask-ticket-e7760c-2
worker/cb-172-broker-uri-d36ae4-4
worker/cb-175-model-readback-76ead6-3
worker/cb-161-pane-ancestry-293510-1
worker/cb-164-rebase-885863-8
worker/cb-164-empty-scrape-false-success-1a80af-3
fix/cb-197-ticket-ttl-from-completion
worker/cb-189-remote-url-coverage-4692f3-1
worker/cb-185-blockers-027756-4
worker/cb-192-gap-log-11b631-2
worker/cb-633-fix-5f4396-3
worker/cb185-router-d6436d-3
worker/cb185-router-routing-gaps-9e9d33-3
worker/cb185-paneids-992586-2
worker/cb-633-allow-list-union-ed374b-1
worker/cb-157-credential-in-remote-url-496e44-2
worker/cb-641-health-herdr-evidence-8f1f54-6
worker/cb-640-health-msg-evidence-99c9cd-1
worker/cb-642-fleets-status-skill-bbbc40-5
cb-634-ide-mcp
worker/lead-comms-wiring-c014b9-7
worker/lead-mailbox-c19577-6
worker/autocompact-window-82bc2f-5
worker/cb-634-probe-18056f-4
worker/cb635-broker-urienv
worker/cb-632-config-retry-8e0efa-7
lead/cb-622e-claude-md
lead/cb-622-followup
worker/cb-622a-165dff-1
lead/cb-622d-opencode-mount
worker/cb-622b-717c67-2
worker/cb-622c-ab7759-3
worker/cb-617b2-20ca4b-3
worker/cb-617a-5c2f4a-1
worker/cb596-4e49ef-3
worker/cb586-10500c-1
worker/cb-606-b9343a-25
worker/cb604-1445f8-24
worker/cb582-477374-21
worker/cb584-8c2281-22
worker/cb600-e6b9a9-20
worker/cb602-ce257f-19
worker/cb601-b42837-18
worker/cb598-6c7ba7-17
worker/cb599-740fe4-16
worker/cb597-282224-15
worker/cb590fix-185e9a-10
worker/cb528-recovery-race
worker/cb594-96bead-8
worker/cb590-916766-2
worker/cb527-997d99-3
worker/cb592-env-leak-3cbf9c-1
worker/cb588-async-ticket-nudge-3218f7-5
worker/cb578b-9dcb13-6
worker/cb581-d24826-5
worker/m2-u5-ef8c42-15
worker/cb578a-516499-2
worker/cb576-01a04b-17
worker/cb579-lead-tab-acba06-20
worker/cb580-terminal-health-ed6058-21
worker/cb577-f36fdc-18
worker/cb573b-3db06f-16
worker/cb568c-f36fdc-18
worker/cb568-drop-cause-c3ac1c
worker/cb575-cancelled-notification-c3ac1c
worker/m4-sol-a2cbec-3
worker/cb574-async-ask-c3ac1c
worker/cb573-health-model-8ca857-14
worker/cb572-unknown-target-7f2e35-13
worker/u4-700706-9
worker/u3-b9fcb6-6
worker/u2-ef5b68-4
worker/u1-469dce-1-clean
worker/u1-469dce-1
worker/cb-564-health-events-70cf7e-2
worker/cb-565-recycle-drops-role-98e58f-3
worker/cb-563-missing-reply-df2866-1
worker/cb-562-readiness-gate-silent-6c23c9-3
worker/cb-560-architect-presence-da8155-1
worker/cb-561-architect-silent-off-a71cab-2
worker/cb-548-bind-architect-slot-fe1b8c-1
worker/parity-overlay-settings-5fb711-1
secrets-central-store
cb-559-hot-key-correction
cb-557-fleet-role-pools
worker/cb-553-maxload-explicit-spawn-305ee3-6
worker/cb-551-idle-lead-heartbeat-f1633c-1
worker/cb-544-drain-preserves-worktree-925fad-3
worker/cb-552-docs-sync-1cb9cf-4
worker/cb-548-rendezvous-guard-rebased
worker/cb-548-rendezvous-guard-116b53-10
worker/cb-548-authz-v2-586df6-8
worker/cb-548-authz-264363-5
salvage/cb-528b-codex-home
salvage/cb-528a-codex-launcher
CB-518-primary-flow
feature/peer-launcher-spi
cb-103-injector
v1.1.0
v1.0.0
Labels
Clear labels
blocked
needs-live-proof
ready-to-delegate
silent-default
Cannot start until something else lands. The body says what.
Merged and green, but never shown working on the running daemon. Not the same as done.
Scope, files and acceptance criteria are written. A worker can be briefed from the body alone.
A feature that compiles, passes tests, and ships turned off. Nine recurrences and counting.
No Label
Milestone
No items
No Milestone
Projects
Clear projects
No project
Notifications
Due Date
No due date set.
Dependencies
No dependencies set.
Reference: fleet/fleetd#705
Reference in New Issue
Block a user
Blocking a user prevents them from interacting with repositories, such as opening or commenting on pull requests or issues. Learn more about blocking a user.
Delete Branch "%!s()"
Deleting a branch is permanent. Although the deleted branch may continue to exist for a short time before it actually gets removed, it CANNOT be undone in most cases. Continue?
Split out of #669, where it is stated in the body but has never had its own ticket.
The hole
CallerResolver's resolution ladder (auth/CallerResolver.java:43):So every pane that is not a live spawned member, not a configured lead, not a bound architect slot and not a configured collaborator falls to
WORKER. That includes a tab a person opened for something unrelated.And
WORKERholdsTASK_READ(auth/Authz.java:145):The code comment immediately above it already states the consequence — it is the stated reason a collaborator is excluded from the same row:
That reasoning applies with equal force to the worker fallback, which is reachable without any configuration at all. The narrower role was given the stricter rule; the catch-all default was not.
#669's body says the same thing in passing:
Why this is not the same as #702
#702 is a narrow timing window during a member's release, and only with
placement: pane. This is the steady state for any pane the operator never configured, with no window and no special placement.What makes it exploitable rather than theoretical
Ticket ids are a plain sequential counter with no owner check. So the holder does not need to guess or discover anything —
task-1,task-2, … enumerates the fleet's delegation traffic. A delegation reply is exactly the content most likely to carry work product, file contents and findings.Not yet measured
I have not run a live probe of this. I read the resolver ladder and the
Authzrow; I did not open an unconfigured tab and poll a ticket from it. Anyone picking this up should prove it end to end first, and should expect to need a positive control — a probe that returns nothing proves nothing until you have seen the same call succeed from a session you know holdsTASK_READ.The fix is a design choice, not a one-liner
Three shapes, and they are not equivalent:
READ/METRICSonly, noTASK_READ, noSEND. Strictly narrower than bothWORKERandCOLLABORATOR, and it does not widen anything.COLLABORATORholdTASK_READsafely for its own tickets. Larger change.COLLABORATOR. RemovesTASK_READbut addsSEND, so any unconfigured pane could inject text into a lead's pane. A trade, not a strict improvement — noted here so it is not mistaken for the obvious answer.Option 1 is the smallest safe step and does not foreclose option 2. This interacts with the named-peer mesh question being designed under #669, so the two should land in a consistent order — but option 1 is defensible on its own, whatever the mesh decision turns out to be.
A measurement that may make option 1 unsafe — with an architect now
The ticket prefers option 1, a new bottom rung below
WORKERholding onlyREAD/METRICS, andcalls it "the smallest safe step [that] does not foreclose option 2". I agree with the diagnosis.
I am not yet sure option 1 is safe, for a reason the ticket does not cover.
The
WORKERfloor looks load-bearing for a spawned member's boot.SessionManager.java:235callslauncher.spawn(req). The registry entry is written only at:244, after spawn returns.HerdrPeerLauncher.waitUntilInjectableOrThrow(aroundHerdrPeerLauncher.java:1034) polls herdr until the pane is injectable.fleetd.yamlon this host setsspawnReadyTimeoutMs: 20000.seconds.
CallerResolverasks the roster first (#669 Unit D), so during that window the memberfalls through to the floor.
FleetMcp.java:789-792markSpawnedMemberPresentis the only writer intoMemberPresence. Iproved that by enumerating writers rather than readers:
grep -rn markPresentreturned 4 hitsand one was an override calling
super.Principal.isSpawnedMember()is true, andPrincipal.java:113defines thatas
WORKER || ARCHITECT.If a booting member's first MCP call resolved as the new bottom rung, presence would not be marked
for that call. If nothing later marks it, the lead's first
fleet_sendto that member would sit onthe injector readiness gate for about 60 seconds and fail without a keystroke reaching the pane.
That is the same silent, slow shape as the collaborator defect just fixed in #706 — allowed at one
gate, dropped at another.
What is not yet established. I have not shown that a member actually makes an MCP call inside
that window. It is possible every first call lands after registration, in which case the worry is
theoretical and option 1 is clean. I did not measure it, and I am not going to assume it in either
direction — assuming it this session is how I got two things wrong already on this same code.
Now with an architect (ticket
task-8), asked to settle five things: whether the window isreally reachable; if it is, which of four fixes is right, including whether registering before the
injectable wait is even possible given that the paneId comes back from
spawn(); whether theticket's option 2 (give tickets an owner check) dodges the problem entirely by fixing the root
cause instead of the role table; the landing order so no intermediate state is either exploitable
or broken; and a unit list. I asked it to form its own position and to say plainly where it
disagrees with the reading above.
The probe is still not run. The ticket's own "Not yet measured" section still stands: nobody
has opened an unconfigured tab and polled a ticket from it. That needs a positive control, because
a probe returning nothing proves nothing until the same call is seen to succeed from a session
known to hold
TASK_READ. I have not done it either, and I am not treating the hole as confirmedend to end.
Decision: option 2 first, and alone. Option 1 is optional hardening, not the fix.
The architect disagreed with me and it is right. I checked its correction in the code myself
before accepting it, because I had already been wrong twice this session on this same path.
My step 4 was wrong
I said a spawned member "can be up and serving MCP, but absent from the registry, for up to 20
seconds". That read
spawnReadyTimeoutMs: 20000as a duration the member spends serving MCP. Itis a timeout on reaching
idle, andidlecomes before the MCP connect, not after.MemberPresence's own javadoc says it, and I read it: herdr'sagent_status"reportsidlefor aworker whose Claude is still booting", and presence is "populated from the MCP transport … its
initializeis the first such contact". SowaitUntilInjectableOrThrowreturns while Claude isstill booting,
registry.putfollows a few in-memory statements later, and the member's first MCPcall arrives after that. The order is the opposite of what my worry assumed. The whole reason
MemberPresenceexists is thatidleis an unreliable readiness signal — if MCP connected firstthere would be no boot window to guard at all.
The architect also found a measurement rather than only a reading: a member can only reach
READYthrough
transitionByTerminal, which scans the registry, andonDeliveredrefusesBUSYunlessthe state is already
READYorDONE. I verified that gate atSessionManager.java:911. Allthree live members reported
state: "busy", so each had a post-registration MCP contact. Under anOBSERVERfloor that same contact resolvesWORKERfrom the roster and marks presence exactly asit does today.
It was careful about what this does not prove: it did not prove the microsecond race impossible,
only that it does not fire in practice and that no live member depended on an in-window contact.
That is the right distinction and it is enough to act on.
But my worry was right about a different case, and that is the valuable finding
A member pane that outlives a daemon restart really is absent from the registry, indefinitely.
FleetMcp.java:1303documents it inwhoami's own javadoc: "A worker the registry has no recordof — one that outlived a daemon restart". Today that pane hits the floor, resolves
WORKER, andmarks presence, so a new daemon can deliver to it again. Under option 1 it would resolve
OBSERVER, presence would never be marked, andFleetd.deliverableTowould refuse it forever.I verified the floor returns
Principal.worker(...)and read that javadoc.That is the same shape as the collaborator defect merged in #706 an hour ago: allowed at one gate,
dropped at another, silently. So the failure mode I feared is real — it just lives on the restart
path, not the spawn path. It is a precondition on option 1 and not on option 2.
Why option 2 wins, and this is the part that decided it
Option 1 does not fix the hole. It shrinks the set of principals that can exploit it. After option
1, every legitimately spawned member still holds
TASK_READand can still walktask-1,task-2, … and read other members' delegation replies. A spawned member is a language modelrunning an arbitrary brief, which is not a principal I would trust with every other delegation's
reply body. Option 2 closes it for everyone, and it has no contact with the boot path at all.
Option 1 also does not cover all of
TASK_READ:FleetMcp.java:1124mapsfleet_statusto it aswell. Option 2 does not close that either, but a live status is a much smaller harm than a reply
body.
Order
Option 2 alone leaves no exploitable intermediate state. Option 1 without the presence change is
the worst kind of intermediate state — possibly broken rather than exploitable, and nothing in the
suite would catch it — so that state must not exist even for one commit. If option 1 ever ships, it
ships with the presence condition widened in the same commit, splitting the floor's two jobs: it
currently both grants authorization and supplies the liveness signal that marks presence, and
option 1 narrows only the first while silently taking the second with it.
Delegated
Unit A (the owner check) is going out now. Its acceptance criteria include the two the architect
named that I would not have thought of: an unnamed primary has a
nullterminal and must stillread any ticket, and the refusal branch itself needs a control, so that the "another terminal is
refused" test cannot pass merely because everything is refused.
Two things are explicitly not checked and must not be assumed: whether a lead keeps its
terminal id across
fleet_handover, which nobody has readLeadRolloverto confirm; and therestart-with-live-panes path, which is reasoned from code only. And after this, lead A can no
longer poll lead B's ticket. I judge that correct and intended, but it is a behaviour change rather
than a side effect, so it is recorded here.
The end-to-end probe from an unconfigured tab is still not run, by anyone. The ticket's "Not
yet measured" section stands.
Unit A merged in
25d53e6— and this ticket stays openPR #712 closes the MCP door.
MessageService.Tasknow records the terminal of the caller whosefleet_send{wait:false}created the ticket, andpoll(ticket, callerTerminal)refuses a differentterminal with a
FAILEDview carrying no reply text.Verified by me in a throwaway worktree, output to a file and not piped: exit 0,
BUILD SUCCESS,Tests run: 2008, Failures: 0. The pushed tree is byte-identical to the tree I built.Two things are still missing, and I am not calling this fixed.
1. The REST door is still open
FleetApp.java:898:That is the no-check overload. The route is gated on
TASK_READ(
FleetApp.java:81), andAuthz.java:145grantsTASK_READto a worker. A worker runs on thesame host as the daemon, so
curl http://127.0.0.1:8765/tasks/task-1reaches it. The exactcaller this ticket is about can still walk every ticket, through REST.
I checked both lines myself. The implementer reported this unprompted, which was the right call.
2. The handler wiring is not pinned — measured, not suspected
The three new tests prove the refusal at the seam,
MessageService.poll(ticket, callerTerminal).Nothing proves the MCP handler passes the caller's real terminal into it.
I replaced
callerTerminal(exchange)withnullatFleetMcp.java:527using a line-anchoredsed, confirmedmvn compilestayed green so the mutation was live rather than a compile error,and ran the full suite:
The mutation survives. The whole fix can be switched off by one argument and no test notices.
I restored the file and confirmed it byte-identical before merging.
The repo has the idiom already:
FleetMcpAuthzTestscrapes the handler block and asserts theargument, with a control assertion so it cannot pass by failing to find the block.
theFleetListHandlerActuallyConsultsCollaboratorsVisibleTois the model.Why I merged anyway
It is a strict improvement and it breaks nothing. Blocking one because it is not yet total would be
the wrong trade. But a half-closed door must not be recorded as a closed one, so this ticket stays
open until both items above are done. They are one follow-up unit.
One thing worth keeping from the review
ownsTicketlets any caller with anullterminal read every ticket. That is safe onlybecause
Authz.java:145deniesANONYMOUStheTASK_READaction, so an anonymous caller neverreaches
poll. I enumerated it:CALLER_TERMINALhas exactly one writer, on every call, and everynon-primary
PrincipalinCallerResolveris built with a real terminal — the only terminal-lessconstructions are
Principal.primary(...)andPrincipal.anonymous(). Sonulldoes not conflate"the primary" with "I could not tell".
Nothing pins that coupling. Granting
ANONYMOUSany read action would silently open the ticketgate, in a different file from the one that looks like it holds the rule.
Answers to the two questions I required
fleet_handover. The roll sends/clearand thebootstrap text to
p.leadTerminal()— the same terminal recorded atopen(), never a new pane,and
confirm()requiresp.leadTerminal().equals(callerTerminal)first. So a lead's pre-handovertickets stay readable afterwards. Good: the fix does not break the rollover path.
only the unnamed primary can still read it. This does not hurt the common case, because
abandon()already fails an open task when its target is torn down. The narrow exposure is anamed lead whose pane is replaced outside
fleet_handover. The implementer could not settlewhether that is reachable in production config, because
fleetd.yamlis gitignored and absentfrom a worker's worktree. Neither have I.
Related, filed separately
#715 — the same "an identifier used as an authorization token" shape on
fleet_send{turnId}, whichis a write path. Also reported by this implementer.
Option 1 (the
OBSERVERfloor) remains undelegated and still must ship with the presence split inthe same commit.
Fixed on both entry paths. Closing.
a9a37af.fleet_poll{ticket}threads the caller's terminal.9a64d42.GET /tasks/{ticket}threads it too, and thewait:falsesend path now records a creator terminal.The REST half was two problems, not one
Closing the read door alone would have broken a working path.
FleetApp.sendMessage'swait:falsebranch called the
sendAsyncoverload that records no creator, so a REST-created ticket carriedcreatorTerminal == null. With the new check,ownsTicketevaluates"term_x".equals(null)for aterminal-bearing caller, which is false — so the ticket's own creator would have been refused, and
driving the fleet over REST is a documented fallback for when the MCP mount drops. Both halves had to
ship together, and they did.
The mutation that mattered
Before this work, replacing
callerTerminal(exchange)withnullatFleetMcp.java:527compiledgreen and the full suite still passed — the fix could be switched off and nothing noticed. I re-ran
that exact mutation on the merged tree myself, confirmed it was live with
mvn -o compilefirst, andit now kills
theFleetPollHandlerActuallyThreadsCallerTerminalIntoPoll. Two further mutations on theREST lines kill one and three tests. Every file restored byte-identical.
Merged build:
Tests run: 2018, Failures: 0, BUILD SUCCESS, against amainbaseline of 2008 that Imeasured myself.
The coupling this fix rests on, which nothing pins
ownsTicketnever refuses a caller with no terminal, because that is the unnamed primary and it mustkeep reading every ticket. That is safe only because
Authz.java:145deniesANONYMOUStheTASK_READaction, so an anonymous caller never reaches the lookup. GrantingANONYMOUSanyticket-read action would reopen this silently. It is written into the new
wiki/11-Features.mdentryas a gotcha, since it is the kind of thing a later change breaks without noticing.
Follow-up
#718 —
MessageService.poll(String)still exists and defaults the terminal tonull, which skipsthe check. It has no production caller now, but 44 test callers, so it stays and gets pinned by a
scrape test instead. #718 also covers a second instance of the same shape at
FleetMcp.java:975.Correction: I closed this a moment ago and have reopened it. Option 2 is done; option 1 is not.
My previous comment is accurate about what shipped, but closing the ticket on it was wrong. This
ticket offers three fix shapes, and PRs #712 and #716 implement option 2 — give tickets an owner
check. Option 1, the narrowed bottom rung (the
OBSERVERfloor), has not been done and is stillan open decision.
What is actually closed
The reported exploit. An unconfigured pane still resolves to
WORKERand still holdsTASK_READ, soit can still call
fleet_poll{ticket}— but it now reads only tickets its own terminal created, andit creates none. So walking
task-1,task-2, … returns nothing.What is still open, and why it is worth keeping open
Option 1 is defence that does not depend on the owner check staying correct. Right now the only thing
standing between an unconfigured tab and other sessions' delegation replies is one comparison in
ownsTicket, and #718 exists precisely because that comparison can be bypassed by reaching for aconvenience overload that compiles and passes. A narrowed floor would mean an unconfigured pane never
holds
TASK_READat all, so the overload hazard and the owner check would both have to fail beforeanything leaked.
It also still carries the constraint recorded against it: the floor must ship together with the
presence split in the same commit. Otherwise a pane that outlived a daemon restart becomes
permanently undeliverable. That coupling is the reason this is a design decision and not a one-line
role-table edit, and it is why I am not delegating it yet.
Also still true
The "not yet measured" section above still stands in full. Nobody has opened an unconfigured tab and
polled a ticket from it, before or after this fix. What is proven is the owner check itself, by unit
test and by three mutations I ran on the merged tree. The resolver-ladder claim that an unconfigured
pane becomes a
WORKERis still read from the source, not observed.Decision on option 1, after consulting two architects — and the inherited constraint was wrong
I put the question to two architect members on different profiles (
solandopus), with the samebrief and the same evidence, and deliberately withheld my own suspicion so their positions would be
independent. I then measured the crux myself. Recording the decision here per the charter.
The two positions
They returned opposite headline verdicts and then described the same mechanism:
solopusAuthzreason"REPLY/ASKsurvive underOBSERVER?fleet_sendto the surviving panefleet_sendto the surviving paneisSpawnedMember(); the injector gates on presenceisObserver()to the presence gateOBSERVER?Both independently reached the
ownsSessionreasoning I had suspected and withheld. Bothindependently found a launch-before-registration window that neither I nor the previous lead had
named. That is genuine corroboration: two different models, two separate readings.
Verdict: the inherited constraint is false as stated, and must not go into a commit
The handover said an
OBSERVERfloor must ship with a "presence split" or a restart-surviving panebecomes "permanently undeliverable". Three things are wrong with that, all of which I checked
myself:
AgentStatus.injectable()— thepane's terminal status — not on the target's authz role.
case REPLY, ASK -> caller.ownsSession(targetSession)names norole, and
ownsSessionisterminal != null && terminal.equals(sessionId)(
Principal.java:133-135). An observer carries its pane's terminal, so it still ends its turn.tasks,asyncTasksByWaiter,asyncTasksByTurn,strandedRepliesandticketSeqare all plain in-memory maps, and theroster is a plain
ConcurrentHashMapwith no persistence. After a restart the lead has noaddress for that pane from any tool.
But the constraint pointed at something real, and credit to both architects for finding it. The
loss is the lead's ability to push a new turn, and it is caused by one line:
An observer is neither a worker nor an architect, so it would never mark presence, and the injector
would never consider it ready.
Where I ruled against
sol, by measurementsoladditionally required registration reconciliation for the launch race, and would reject arole-only PR without it. I checked that path and it is pre-existing and unchanged by this work:
So a pre-registration MCP contact marks presence and silently loses the
SPAWNING -> READYtransition — today, for a
WORKER, exactly as it would for anOBSERVER. Once the observer isadded to the presence gate, that path is byte-for-byte what it is now. It does have a real
consequence (
FleetMcp.java:2204usesstate() == READYfor seat accounting), so it is worthfixing — but it is not a prerequisite for this change, and bundling it would hide a pre-existing
defect inside a role addition. Filed separately.
With that removed, the two positions agree on the required behaviour.
sol's own wording — "asuccessful bridge MCP contact from an observer terminal must mark that terminal safe for injection"
— is
opus's one predicate.Decision
OBSERVERships, but not next. It is defence in depth, not a fix for a live hole: option 2 ismerged and closed the reported exploit. Its real justification is narrower and still worth it — the
floor currently makes a false identity claim, telling the daemon that anything in a pane is a
spawned member.
Ship #721 first.
opusfound, and I verified on both the MCP and REST surfaces, thatfleet_statushas no owner check and returns another session's pending question text, itsturnIdand itsticketto anyTASK_READholder. That is a live content leak, it is cheaper thanthis change, and it is independent of this decision. It also voids the mitigation I had claimed on
#715. That ticket now leads.
Required in the same commit as the floor (the union both architects agree on):
Role.OBSERVER,Principal.observer(terminal, pid),isObserver(), and adescribe()case —describe()has nodefault, so a missing case is a compile error. Keep that.CallerResolver.resolvechanges. The roster rung mustkeep returning
WORKER/ARCHITECT— a careless diff breaks that half, and it is the half thatmatters.
Authz: addisObserver()tocase READ, METRICSand to nothing else.REPLY/ASKneedno edit.
FleetMcp.java:834and rename the method, since it isno longer only spawned members.
isObserver()branch infleet_whoamibeforeif (!caller.isWorker())(
FleetMcp.java:1402). Today an observer would reach the lead branch by elimination. Theoutput happens to be correct only because
m.put("leader", ...)is guarded on a non-null name —correct by luck is not correct.
Fleetd.java:229-234paragraph claiming presence "doubles as the memberroster's availability signal".
opusenumerated everyisPresentreader inmainand foundnone that makes it true. A comment asserting a false invariant is what turned this into a "hard
constraint" in the first place.
CLAUDE.md+ the byte-identical wiki template + awiki/11-Features.mdentry. This is a roletable and
ConnectionIdentitychange, so the project's own "prompt is part of the product" tablemakes all three mandatory. The fallback ladder has no row for
observer.OBSERVERrow for every action inAuthzTest; that an observer mayREPLY/ASKfor its own pane and not another's; and a
CallerResolverTestcase asserting both that anunknown pane resolves
OBSERVERand that a registered member still resolvesWORKER/ARCHITECT.May follow later: narrowing
fleet_profilesfor an observer; a persisted roster.Cost warning, which I am recording rather than discounting
opusnoted that every role added here has needed follow-up:COLLABORATORtook #669, #703, #710and then
d2db8c7as a review fix on the PR that merged yesterday. Four corrections for one role.Treat the diff estimate as a lower bound.
Not measured
Neither architect ran a live probe, and nor did I. Nobody has opened an unconfigured tab and
attempted these calls.
solran a 173-test subset green and changed nothing;opusran no build.Neither read
AuthzTestin full, so the number of assertions a new enum constant turns red isunknown. And nobody has verified that a Claude Code member re-establishes its MCP connection after
a daemon restart —
opusflagged that its whole restart argument assumes it does. If thatassumption is false, window 2 is moot under either floor, and that is one probe worth running before
anyone writes code for item 4.
Refining item 6 of the decision — half that paragraph is true, so do not delete it wholesale
I attributed the claim about
Fleetd.java:229-234to the architect rather than measuring it, anditem 6 said "correct or delete the paragraph". I have now checked it myself, and the instruction was
too blunt: the paragraph makes two claims and only one is false.
The paragraph:
Claim 1 — the stated reason — is false. Every reader of
MemberPresence.isPresentinfleetd/src/main/java, after filtering out unrelatedOptional.isPresent()hits:Three call sites, all one predicate: deliverability. Nothing reads presence as "an available member
in the roster." Seat and capacity accounting come from elsewhere —
reclaimable()keys onsession.state() == READY(FleetMcp.java:2204), and the roster itself comes fromsessions.roster(). So "a lead or collaborator counted there would show up as an available member"has no reader behind it. It may have been true when written.
Claim 2 — the consequence — is true and load-bearing. Without the
leadsandcollaboratorsdisjuncts a lead or collaborator really is undeliverable, because the gate is exactly those three
disjuncts. That sentence must stay.
Revised item 6
Correct only the causal clause — the "since that map doubles as…" reason. Keep the sentence about
the second and third disjuncts. And when you remove the false reason, say what actually pins the
exclusion now, or the next reader will reinstate it: per this project's comment rules, name the
constraint and what breaks if someone changes the line, not the history.
This also matters beyond a comment. That false reason is why the handover recorded a "hard
constraint" in the first place: a reader looking at it concludes that adding anyone to presence has
roster side effects, so the one-predicate fix looks unsafe and a "presence split" looks mandatory. A
comment asserting an invariant nothing enforces cost us a wrong constraint, two architect turns and a
round of my own verification. It is a free test case — the claim was checkable in one
grep.