After #726 ships, a lead handover will orphan every in-flight delegation: the fresh lead gets a new terminal id and ownsTicket refuses it #737
Closed
opened 2026-10-04 19:36:17 +02:00 by ltms
·
14 comments
No Branch/Tag Specified
main
worker/737-9c61d3-4
worker/737-a263f3-3
worker/737-20d1d9-1
worker/737-b038f7-2
worker/726-unit2-75cb13-4
worker/737-owner-key-ff061f-10
worker/736-presence-forget-f35144-9
worker/705-observer-14c258-6
worker/722-024c34-5
worker/726-ea34a0-2
worker/726-10cbf0-1
worker/729-5961c6-3
worker/727-ee14ed-3
worker/719-bdd95e-4
worker/702-4f5c7f-2
worker/715-5c43fc-1
worker/721-70f9ea-5
worker/718-99362b-2
worker/task-15-af0d10-12
worker/task-16-50a702-13
worker/task-12-4d0479-9
worker/task-13-823ce2-10
worker/705-ticket-owner-af9928-8
worker/703-list-collaborators-9c06c2-7
worker/669-example-truth-0b303d-6
worker/669-collab-deliverability-9ba859-3
worker/669-collab-reload-report-2a21bd-4
worker/669-7e80a6-1
worker/669-unit-d-efbbd7-1
worker/669-1b786a-1
worker/669-1d1d9f-1
worker/692-4afb9d-2
worker/689-02fced-13
worker/693-cf23fa-14
worker/677-fix-lead-collision-f69073-12
worker/638-fix-overmask-dbb1bf-11
worker/675-5b7478-4
worker/669-unit-a-70cc8f-3
worker/677-8cdaaf-5
worker/638-a7b391-1
worker/683-4536d6-2
worker/651-a75bbe-8
worker/680-20607d-7
worker/664-c12e95-3
worker/668-08534d-4
worker/672-0f2469-2
worker/670-7d1022-1
worker/661-ac7c28-2
worker/664-37fb9b-3
worker/663-remove-3arg-read-3f6783-1
worker/659-remove-dead-backcompat-ba5e6f-1
worker/637-revision-60a488-23
worker/656-redact-regression-tests-892903-19
worker/637-context-gauge-threshold-466eb5-16
worker/639-redact-line-numbers-de4ac4-17
worker/641-set-reformat-guard-6f96a4-18
worker/642-herdr-guard-scope-5de0e4-15
worker/650-javadoc-scope-95f3b3-14
worker/612-01e9f7-13
worker/612-a-r4-quarantine-outage-7ab0e8-5
worker/612-a-r9-r11-capacity-coverage-peers-cfcc79-7
worker/612-a-r10-loophealth-ccc872-8
worker/612-a-r12-turnregistrar-9e3bb7-9
worker/612-a-r5-leadconfigdir-9e70cf-6
lead/config-edit-redact-anchor-wording
worker/config-edit-seam-ca8dc1-1
worker/612-r67-630-lifecycle-290b8d-3
worker/629-625-ports-seams-da7d5d-4
worker/612-r12-exhaustion-f37cd7-1
worker/612-r38-amqp-24b083-2
worker/fleetd-612-unita-87807e-1
worker/612-b3-mcpwirings-da2b58-3
worker/612-b2-cb185-176d3a-2
worker/612-b1-completion-457459-1
worker/612-agaps-73a926-2
worker/608-sleeps-3a64ff-3
worker/621-b4520b-1
worker/618-b83894-2
worker/fleetd-615-e05481-5
worker/lead-autocompact-5f1ab2-3
worker/fleetd-613-f85deb-3
worker/fleetd-608-flaky-nudge-test-d0c2d1-3
worker/lead-context-gauge-ad404f-1
worker/gauge-wiring-9158c1-4
worker/redeploy-slowstart-ead0e5-5
worker/charter-bytes-13668c-6
worker/rollover-outcome-291483-2
worker/589-f64303-2
worker/593-1a8025-5
worker/589-fcd2aa-1
worker/568-9fdaa2-3
worker/571-attempted-outcome-5739f7-2
worker/581-completionresolver-cas-sites-0542b7-6
worker/562-loop-health-wiring-test-99611c-5
worker/562-surface-loop-health-7df5cc-4
worker/575-waiter-cleanup-sites-62ad80-1
worker/572-answer-lock-release-46a9ae-5
worker/567-probe-channel-leak-a38fc5-6
worker/551-record-before-send-7cbf56-1
worker/561-listener-fanout-survives-a-throw-61d538-2
worker/555-redeploy-main-flow-seam-65c2f5-2
worker/556-injector-owns-registration-e027a5-1
worker/552-post-restart-mktemp-abort-bc2672-4
worker/553-onstatus-completion-leak-0da881-2
worker/550-shasum-linux-196132-1
worker/538-loop-dies-on-error-4a5eeb-6
worker/426-health-coverage-ef1fd4-4
worker/504-failed-reported-clean-3cfd66-3
worker/537-capturedlog-close-e4c437-2
worker/459-broken-link-targets-cadc17-5
worker/535-appender-leak-fe74c1-1
worker/512-part2-shutdown-detection-434701-9
worker/529-logger-level-sweep-2a5533-8
worker/528-drain-gate-call-site-5de83d-7
charter/forge-mcp-vs-token
worker/521-swap-guard-unpinned-28e931-5
worker/519-probe-test-harness-d25ab8-4
worker/525-logger-level-leak-1b4eb0-6
worker/518-fleetmcp-resolver-wiring-8ef96c-1
worker/512-drain-complete-line-7edd71-3
worker/517-abort-branch-and-jar-id-41b641-2
worker/500-9e52c9-3
worker/509-4912f4-2
worker/511-9a4b23-1
worker/493-479f45-2
worker/505-03f8b2-1
worker/492-followup-detect-unclear
worker/501-a31fa0-7
worker/498-451d1c-5
worker/494-1015ce-2
worker/492-209647-1
worker/489-001902-2
worker/480-relative-handover-path-906323-1
worker/480-b-handover-skill-45bf1f-5
worker/474-followup-source-pin-f54a55-17
worker/474-charter-check-on-reload-f54a55-17
worker/466-quarantine-repeatcount-report
worker/393-opencode-skill-seeding-71854b-13
worker/469-canonical-tool-names-2a472a-16
worker/466-quarantine-escalation-5ae9c1-15
worker/446-hot-exhausted-pattern-0af580-6
worker/464-charter-tool-name-guard-a85635-12
worker/463-listfleet-default-fails-open-f1c76c-11
worker/458-invariant-5-by-purpose-862f9a-10
worker/439-coordinator-row-gate-bc032a-8
worker/449-herdr-protocol-576015-4
worker/450-abstract-spawn-599e1c-5
worker/437-ack-refuses-177d91-1
worker/444-placement-window-feb56a-2
worker/440-helddurable-derived-d462d7-13
worker/425-rework-placement-resolve-c58ba1-9
worker/421-lead-peek-held-msgs-cdbad2-10
worker/435-fixed-policy-cap-fe11de-12
worker/422-gate-state-observability-9e79d6-11
worker/431-memberregistry-live-readers-cdbad2-10
worker/424-architect-slot-hot-038b41-7
worker/422-model-gate-spawn-c29f48-6
worker/425-default-profile-live-f55534-8
worker/415-coverage-wording-2cbf9c-5
worker/416-3ad1da-1
worker/418-588283-3
worker/deterministic-stamp-race-409-3cb7b6-10
worker/armed-reads-live-config-404-ed931f-9
worker/reply-peer-refusal-391-5a34bd-7
worker/models-allowlist-aa9e9b-3
worker/ttl-stamp-race-399-f1122f-8
worker/scrub-receipt-400-316b3e-5
worker/exhaustion-detection-395-105105-6
worker/scrub-abort-394-316b3e-5
fix/scrub-uid-abort
worker/task-scrub-517574-2
worker/t386-clock-bd5b78-4
worker/t384-scrub-813790-5
worker/t381-cc-748314-2
worker/t373-336973-2
worker/t365-3920c5-3
worker/t358-6e989b-1
worker/t355-8b321c-1
worker/fleetd-369-hermetic-git-tests-e8b19a-3
worker/fleetd-368-stale-lead-binding-f5682e-2
worker/fleetd-360-deploy-units-0d3793-1
worker/359-dead-lead-tabs-f1253b-4
worker/362-worktree-skills-c03e51-3
worker/361-coord-visibility-655144-1
362-plugin-visibility-and-drift
worker/errscan-bed2ca-2
worker/amqp-log-identity-bed2ca-2
worker/withdefaults-guard-561704
worker/sleepguard-82076d-1
worker/fd334-9ee1b6-5
worker/fd348-f1ab27-4
worker/fd335-a71c35-1
worker/fd342-174a17-2
worker/fd345-490d0f-3
worker/fleetd-337-5ec7d4-21
worker/fleetd-341-af5a6b-24
worker/fleetd-339-5ca0a2-23
worker/fleetd-338-83a4a1-22
worker/fleetd-333-281f46-18
worker/fleetd-329-11bdbb-16
worker/fleetd-330-2770fb-17
worker/fix-326-50506e-15
worker/fix-324-3e9bbf-14
worker/fix-323-b8287d-13
worker/fix-316b-bd0860-11
worker/fix-318-76ca36-9
worker/fix-317-486aec-8
worker/fix-315-ce47c5-6
worker/fix-307-275890-6
worker/fix-308-b4f664-7
worker/fix-309-ec3939-8
worker/fix-310-7a3974-9
worker/fix-302-52ad0e-9
worker/fix-298-ce1acb-8
worker/fix-297-66bd11-7
worker/fix-296-104622-6
worker/fix-293-bare-closetab-eb22b5-3
worker/fix-280-gone-ask-lapse-bca98e-2
worker/fix-290-reapidle-guard-coverage-9b0dd1-1
worker/fix-285-trust-seed-8f3565-10
worker/fix-284-backend-error-seat-85912c-11
worker/fix-282-chained-ask-e6d0bb-8
worker/fix-283-teardown-leaks-f40dfa-9
worker/fix-281-pin-handler-actions-4921ac-7
worker/audit-rendezvous-lifecycle-d072ae-2
worker/audit-health-placement-1a2476-6
worker/audit-teardown-exits-e207a5-3
worker/audit-launcher-asymmetry-27e370-4
worker/audit-rest-authz-6ca53c-5
worker/investigate-275-abandon-asking-fdef52-8
worker/fix-274-worktree-leak-b0095d-7
worker/fix-273-exhausted-pattern-9665b5-6
worker/fleetd-267-model-check-bd8068-1
worker/fleetd-131-archunit-18b834-7
worker/fleetd-266-sshagent-rename-a014ff-6
worker/fleetd-184-uid-claim-8e1f31-4
worker/fleetd-184-warn-b381ee-10
worker/fleetd-184-docs-be1d12-9
worker/fleetd-257-9bf010-7
worker/fleetd-103-23a113-6
worker/fleetd-247-342356-5
worker/fleetd-116-04dea8-4
worker/fleetd-252-a830e0-3
worker/fleetd-111-7e8673-9
worker/fleetd-155c-f8ef4b-8
worker/fleetd-176-b928ca-3
worker/fleetd-249-7a7878-2
worker/cb248-composition-root-b-9acdf7-15
worker/cb148-envrc-default-fa6c82-12
worker/cb201-unit5-wiring-6c12e6-8
worker/cb241-fallback-echo-1175e9-11
worker/cb149-trust-dialog-2392a5-9
worker/cb134-148-overlay-visible-c9b986-10
worker/cb234-session-id-keyed-04e1fc-1
worker/cb201-unit3-nudge-abdf5c-6
worker/cb201-unit2-policy-c1102c-5
worker/cb201-unit4-outcome-a13bfa-7
worker/cb201-unit1-classifier-91b9b1-4
worker/cb201-227-refine-831980-3
worker/cb175-model-readback-0f085f-1
worker/cb222-charter-tmpdir-17f013-1
worker/cb226-architect-slot-race-cd3aa8-3
worker/cb224-worktree-root-group-024523-2
worker/cb-123-role-demotion-c600f7-2
worker/cb-219-opencode-roots-1f677e-1
worker/cb214-claude-session-id-b9eab4-4
worker/cb213-zdotdir-wrong-process-dd6de4-3
worker/cb211-exhaustion-classification-9546e0-2
worker/cb137-ambiguous-task-4df3d8-4
worker/cb209-agentsessionid-4dfdb6-2
worker/cb185-hostenvnames-2692b5-3
worker/cb206-opencode-sqlite-128718-2
worker/cb185-worktree-group-fc0c99-1
worker/cb-137-ask-ticket-e7760c-2
worker/cb-172-broker-uri-d36ae4-4
worker/cb-175-model-readback-76ead6-3
worker/cb-161-pane-ancestry-293510-1
worker/cb-164-rebase-885863-8
worker/cb-164-empty-scrape-false-success-1a80af-3
fix/cb-197-ticket-ttl-from-completion
worker/cb-189-remote-url-coverage-4692f3-1
worker/cb-185-blockers-027756-4
worker/cb-192-gap-log-11b631-2
worker/cb-633-fix-5f4396-3
worker/cb185-router-d6436d-3
worker/cb185-router-routing-gaps-9e9d33-3
worker/cb185-paneids-992586-2
worker/cb-633-allow-list-union-ed374b-1
worker/cb-157-credential-in-remote-url-496e44-2
worker/cb-641-health-herdr-evidence-8f1f54-6
worker/cb-640-health-msg-evidence-99c9cd-1
worker/cb-642-fleets-status-skill-bbbc40-5
cb-634-ide-mcp
worker/lead-comms-wiring-c014b9-7
worker/lead-mailbox-c19577-6
worker/autocompact-window-82bc2f-5
worker/cb-634-probe-18056f-4
worker/cb635-broker-urienv
worker/cb-632-config-retry-8e0efa-7
lead/cb-622e-claude-md
lead/cb-622-followup
worker/cb-622a-165dff-1
lead/cb-622d-opencode-mount
worker/cb-622b-717c67-2
worker/cb-622c-ab7759-3
worker/cb-617b2-20ca4b-3
worker/cb-617a-5c2f4a-1
worker/cb596-4e49ef-3
worker/cb586-10500c-1
worker/cb-606-b9343a-25
worker/cb604-1445f8-24
worker/cb582-477374-21
worker/cb584-8c2281-22
worker/cb600-e6b9a9-20
worker/cb602-ce257f-19
worker/cb601-b42837-18
worker/cb598-6c7ba7-17
worker/cb599-740fe4-16
worker/cb597-282224-15
worker/cb590fix-185e9a-10
worker/cb528-recovery-race
worker/cb594-96bead-8
worker/cb590-916766-2
worker/cb527-997d99-3
worker/cb592-env-leak-3cbf9c-1
worker/cb588-async-ticket-nudge-3218f7-5
worker/cb578b-9dcb13-6
worker/cb581-d24826-5
worker/m2-u5-ef8c42-15
worker/cb578a-516499-2
worker/cb576-01a04b-17
worker/cb579-lead-tab-acba06-20
worker/cb580-terminal-health-ed6058-21
worker/cb577-f36fdc-18
worker/cb573b-3db06f-16
worker/cb568c-f36fdc-18
worker/cb568-drop-cause-c3ac1c
worker/cb575-cancelled-notification-c3ac1c
worker/m4-sol-a2cbec-3
worker/cb574-async-ask-c3ac1c
worker/cb573-health-model-8ca857-14
worker/cb572-unknown-target-7f2e35-13
worker/u4-700706-9
worker/u3-b9fcb6-6
worker/u2-ef5b68-4
worker/u1-469dce-1-clean
worker/u1-469dce-1
worker/cb-564-health-events-70cf7e-2
worker/cb-565-recycle-drops-role-98e58f-3
worker/cb-563-missing-reply-df2866-1
worker/cb-562-readiness-gate-silent-6c23c9-3
worker/cb-560-architect-presence-da8155-1
worker/cb-561-architect-silent-off-a71cab-2
worker/cb-548-bind-architect-slot-fe1b8c-1
worker/parity-overlay-settings-5fb711-1
secrets-central-store
cb-559-hot-key-correction
cb-557-fleet-role-pools
worker/cb-553-maxload-explicit-spawn-305ee3-6
worker/cb-551-idle-lead-heartbeat-f1633c-1
worker/cb-544-drain-preserves-worktree-925fad-3
worker/cb-552-docs-sync-1cb9cf-4
worker/cb-548-rendezvous-guard-rebased
worker/cb-548-rendezvous-guard-116b53-10
worker/cb-548-authz-v2-586df6-8
worker/cb-548-authz-264363-5
salvage/cb-528b-codex-home
salvage/cb-528a-codex-launcher
CB-518-primary-flow
feature/peer-launcher-spi
cb-103-injector
v1.1.0
v1.0.0
Labels
Clear labels
blocked
needs-live-proof
ready-to-delegate
silent-default
Cannot start until something else lands. The body says what.
Merged and green, but never shown working on the running daemon. Not the same as done.
Scope, files and acceptance criteria are written. A worker can be briefed from the body alone.
A feature that compiles, passes tests, and ships turned off. Nine recurrences and counting.
No Label
Milestone
No items
No Milestone
Projects
Clear projects
No project
Notifications
Due Date
No due date set.
Dependencies
No dependencies set.
Reference: fleet/fleetd#737
Reference in New Issue
Block a user
Blocking a user prevents them from interacting with repositories, such as opening or commenting on pull requests or issues. Learn more about blocking a user.
Delete Branch "%!s()"
Deleting a branch is permanent. Although the deleted branch may continue to exist for a short time before it actually gets removed, it CANNOT be undone in most cases. Continue?
Found while deciding whether to roll this lead with two delegations running. This is a regression #726 unit 2 will introduce, not a defect in today's code — so it must be settled before unit 2 is deployed.
The two facts, read in the code
1. A named lead's ticket reads are checked against its terminal id.
A named lead carries a real terminal, not null:
So the
callerTerminal == nullescape inownsTicketis for the unnamed primary (Principal.primary(pid)), never for a configured lead.2. #726 unit 2 gives the fresh lead a new terminal id.
That is the whole point of the ticket: end the old
claudeprocess, close its pane, create a new tab and start a new agent. A herdr terminal id belongs to the pane, andLeadLauncher.launch()creates a fresh tab every time. The new lead resolves through the sameCallerResolver.java:308rung — same lead name, differentc.terminal().The consequence
ownsTicketcompares the thing that changes and ignores the thing that does not. After a roll, the fresh lead callingfleet_poll{ticket}on any ticket the previous lead created gets:A lead rolls precisely when its context is full, which is usually mid-work with workers running. So the failure fires exactly when it costs most: every in-flight delegation's report becomes uncollectable, and the worker's whole turn is lost. The workers themselves are fine — they keep working and
fleet_replynormally. Nobody can read the answer.This falsifies a conclusion already recorded as safe
#705's comment 18617 answered a required question with:
That was true when written and true today, because the roll is a
/clearinto the same pane. Unit 2 is the change that makes it false. This is the "a fix can create reachability" shape from the other direction: a correctness fix elsewhere turns a previously-safe statement into a defect, and the statement reads as settled so nobody re-checks it.REST is closed too, and one escape hatch exists
GET /tasks/{ticket}threads the caller terminal as well (PR #716), so dropping to REST from the lead's own pane hits the same refusal — its pid maps to its pane, so it resolves as the same named lead.The hatch: a session that is not in a lead-labelled tab resolves as the unnamed primary, whose terminal is null, so
ownsTicketreturns true for every ticket. An ordinary terminal on this host can therefore collect an orphaned ticket. That is a real workaround and worth knowing, but it is a manual rescue by an operator, not a fix — and it depends on an authorization property (ANONYMOUSis deniedTASK_READ, so null means "the primary", not "could not tell") that #705's comment already flagged as pinned by nothing.Fix shapes, not equivalent
LeadRollovertellsMessageServiceto reassigncreatorTerminalfrom the old terminal to the new one, as part of the roll. Smallest change that keeps the ownership model intact. The reassignment is an authorization-relevant write, so it must be reachable only from the roll and must name the old terminal it is replacing, not accept an arbitrary pair.My own reading is that 1 or 3 must ship with or before unit 2 reaching the live daemon, and 2 is the better long-term answer. I am not deciding it here — the next lead should settle it, consulting architects if it is not obvious, and record the decision on this ticket.
Scope note for #726 unit 2
The unit 2 worker has not been asked to fix this, and should not widen its scope. Its brief is write-once and the correction is on #726's ticket instead.
Not measured
I have not rolled a lead with unit 2's code, because unit 2 is not merged. Everything above is read from the four code sites quoted, at
main=428a12a. Nobody has observed the refusal, and the statement it falsifies was itself read from code rather than measured.Delegated to two architects, and the question is wider than the issue body says
Two architects are working this now, independently, with an identical brief —
sol(term_65d07475d124a8e) andopus(term_65d0747d6e8fe8f). Neither has seen the other's answer.I widened the question, because the issue body under-scopes it. I enumerated exactly one channel — tickets — and then wrote a conclusion about handover in general. That is a mistake with a name here: enumerate one channel, conclude about all of them. The real question is not "do tickets break" but:
So the brief asks for an enumeration of every read and write keyed on a lead's terminal, each ruled in or out with a
file:line, covering at leastTask.creatorTerminal, the per-session reply inbox, rendezvous entries andfleet_askturnIds,PrimaryRegistry,LeadRollover's own post-rollp.leadTerminal()uses, the coordinator's held lead-to-lead mail, and any terminal match inAuthzorInjector. Tickets may turn out to be the only one. That would be a finding; assuming it is not.The brief also asks each architect to check my three code reads rather than take them, and to mark every claim as read-in-the-code, inferred, or not-checked. I told them explicitly that finding my reading wrong is a useful answer.
On the three fix shapes, each must answer five things that are where a fix of this kind goes wrong: whether it keeps lead A out of lead B's tickets (#705 option 2 closed that deliberately, and reopening it would be worse than this bug), whether a reassignment stays reachable only from the roll, who owns the state when a roll dies half way, whether the unnamed-primary
nullescape has to change, and whetherTaskcan even see a lead name today.I am not asking them for a verdict. On #726 two architects agreed on a conclusion while both had the grace-scan mechanism backwards, and their agreement made that harder to catch rather than easier. So the brief asks for mechanisms, and I will read those rather than count votes. If they disagree after their independent pass, I will exchange their positions and ask each to compare; if they still disagree, the decision is mine and I will record it here with the reason.
Nothing is decided yet. This comment records who is working it and what they were actually asked.
Architect A (
sol) — eight things break, not one. My issue body was badly under-scoped.I have not re-verified these line references myself yet, and the second architect is still working. This comment records architect A's report so it is durable. Treat every row as architect A's reading until a second pair of eyes confirms it. Its own claim-marking (read in the code / inferred / not checked) is preserved below.
It checked at
428a12a.Breaks when a lead's terminal id changes
Task.creatorTerminalwritten atMessageService.java:285-296and:1331-1334; refused at:1412-1420viaownsTicket:1463-1465FleetMcp.java:1225-1240) and REST (FleetApp.java:901-907) pass the new terminal into the same gatefleet_statusownsTicketgate inpendingAskatMessageService.java:1793-1797turnIdblockfleet_askanswerRendezvous.openAskcopiesownerOf(workerSession)atRendezvous.java:195-203;MessageService.answerrequires an exact terminal match at:1175-1182turnIdsurvives, its stored owner does not — the answer is refused asNOT_TURN_OWNERPrimaryRegistry.leadByTargetat:30-36,79-84;nudgeTargetForat:103-105ReplyPushLoopstores the lead terminal inReplyEntry/PendingTicket/PendingQuestionat:96-114,178-206; tick sends at:394-405,753-776WAIT_BUSY, so it retries the dead address forever.resolveLiveLeadis not called by that tickLeadHeartbeatLoop.tickreads onlyPrimaryRegistry.primaryTerminal()at:309-327, sends at:364-397PrimaryRegistry.java:58-60) never recovers by itselfrollingByTerminalkeyed by the old terminal atLeadRollover.java:313-322, claimed:505-513, released inrunRollover'sfinally:572-589Injector.targetskeyed by target terminal at:124, enqueue:342-358fleet_listand retrySurvives
Rendezvous.java:96-103,:308-310), which does not change. The reply still completes the ticket; only reading it then fails atownsTicketInMemoryReplyInbox.java:9-24,AmqpReplyInbox.java:32-40,85-86), andDRAINis primary-wide (Authz.java:103-108)LeadMailboxownslead.<selfCoordId>.inbox(:197-208);LeadCoordLoopre-reads the live terminal→name map each tick (:194-220)LeadRolloverpending requestconfirmruns before the old pane dies;status(token)is token-keyed (:666-686)AuthzMessageServiceandRendezvous, not the role tableCorrections to what I wrote in this issue
selfCoordId, not "the host". My wording was loose.null-escape claim needs a qualifier. "The escape is for the unnamed primary only" is true at an authorized production entry point, butownsTicketitself cannot tell why its argument is null — safety depends onAuthzrefusing anonymous callers first. #705 comment 18617 already flagged that coupling as unpinned.main, because unit 2 is not merged. Architect A accepted it as inferred from #726's design and comments, not as a fact read from the branch. That is the right way to hold it, and I should have written it that way.Architect A's proposed fix — option 2, broadened
Not just a
creatorLeadNamefield. One typed, immutable owner identity, built only from the connection-resolvedPrincipaland never from a request argument:NamedLead(name)for a configured leadSession(terminal)for an architect or other terminal-bound callerUnnamedPrimaryfor the off-pane primaryThen resolve
name → current terminallate, at each send, rather than storing an address. Applied toTask,Rendezvous.Owner,PrimaryRegistry.leadByTarget,ReplyPushLoop's schedules, the heartbeat, andLeadRollover's single-flight key.Two deliberate restrictions in its design, both of which I think are the right instincts:
transferTicket(old, new)primitive. The owner is immutable from creation; only the address is resolved late. That answers my "does this become a general owner-change primitive" question with "there is nothing to reach".Role.PRIMARYwith a non-null configured name getsNamedLead.On lead A versus lead B:
NamedLead("A")matches onlyNamedLead("A"), and the name comes fromCallerResolver.java:302-308, not the request — so #705's boundary holds.On old tickets: no migration needed. Tasks live only in an in-memory map (
MessageService.java:322-328) and ticket ids carry a per-boot nonce (:357-365), and loading new code requires a restart, which builds a newMessageService. So no ticket from the old class can exist inside the new one.Sequencing
Architect A: may merge before, must not deploy before. Its reason is that a warning is not enough, because a worker can finish or raise a question during the roll, after any pre-roll drain check. That matches what I wrote, with a better reason than I gave.
What architect A did not check
No unit 2 code (not merged in its worktree), no live rollover, no OpenCode restart behaviour, no
fleetd.yaml, nowiki/, and no Maven run — it changed no code. Its final broadrgsweep did not run (rgis not installed in the worktree); the table comes from the editor Grep and Read tools, which did run.Still to do
Architect B (
opus) is still working and has not seen this. I will read its mechanism rather than count votes, post it, and then adjudicate. One item needs carrying to #726 regardless: the rollover single-flight claim unit 3 just merged is keyed on the lead terminal, which is exactly what unit 2 changes.Decision: option 2, role-prefixed. Both architects chose it. And one claim in my issue body is wrong — correcting it first.
Architect B (
opus) has reported. I re-read every line it corrected, in the tree now atmain=11998cd, and I am quoting what I read rather than relaying it.First: my escape-hatch paragraph is wrong. Do not use it.
I wrote in this issue:
That is false, and it would make an operator's rescue fail silently. Any herdr pane gets a terminal, and an unmatched pane does not fall through to the unnamed primary — it hits the pane floor:
A non-null terminal, so
ownsTicketrefuses it. The unnamed primary is reached only when no pane resolves at all — a plain terminal application that is not a herdr pane. An operator on this host is usually already inside a herdr pane, which is exactly where my stated rescue does not work.And it is now closed harder than that. #705 option 1 merged a few minutes ago (
11998cd), narrowing that floor fromWORKERtoOBSERVER, andOBSERVERis not in theTASK_READgate at all:So a herdr pane that is not a lead, member, architect or collaborator is now refused at the role gate, before any owner comparison. Architect B flagged that the hatch line should not be relied on long-term; it is already gone.
Where the two architects agree, and where they do not
Both independently chose option 2, both rejected a transfer primitive, and both said the same thing about sequencing: free to merge, coupled to deploy. I am not treating that agreement as a check — on #726 two architects agreed while both had a mechanism backwards — so I read their mechanisms, and the mechanisms differ in useful ways.
They disagree on exactly one row: unit 3's single-flight claim
runRollover'sfinally, and nothing re-reads it later.They are both right, about different properties, and neither said which property it was answering. I read the code: the claim is taken with
putIfAbsent(p.leadTerminal(), token)and released with the two-argumentremove(p.leadTerminal(), p.token())— the same key both times. So B is right that the release is correct and nothing leaks. A is right that it stops being a per-lead lock: once the terminal changes mid-roll, a secondconfirm()from the fresh lead lands on a different key and is not excluded.So this is not a contradiction but an under-specification: the claim does what it was built to do (release reliably) and no longer does a thing nobody had asked it to do yet (exclude a second roll of the same lead). That is a real, narrow gap, and it belongs in the fix — B's design already closes it by keying single-flight on the lead identity rather than the pane.
This is why the brief asked for mechanisms. Two bare verdicts here would have been a coin toss.
B found a third break neither I nor A ranked correctly
fleet_send{turnId}— answering a worker'sfleet_ask— breaks through a different gate:B's severity argument is right and I had it wrong. A refused poll loses a report. A refused answer loses control of live work: the worker's ask window is about 55 seconds, nothing extends it, so it resumes unanswered and takes whatever default it had. That can produce wrong code, not a missing message.
And note what the javadoc already says:
ownsTicketandOwner.permitstreatnulldifferently on purpose. Any fix touches both and must not quietly align them. The ask path's fail-closed default is deliberate and documented at the line.B closed the escape I was half relying on, with a mechanism I can confirm from this session
B's claim: the ungated
fleet_poll{target}inbox drain cannot rescue an in-flight delegation, becausereply()gives the text to the open waiter and never publishes to the inbox:and the forward waiter stays open for
ASYNC_TIMEOUT_MS = 30 * 60 * 1_000L(:56). So for 30 minutes there is no route: the ticket is refused and the inbox is empty.I can corroborate this from my own session today, and the two readings differ in the axis that matters:
fleet_poll{target}fleet_poll{ticket}timed_out_working(past 30 min)[]That is exactly B's mechanism, observed twice, varying ticket-expiry rather than repeating one case. It also means the stranding is real but time-bounded, as B said.
The decision
Option 2, with architect B's restriction, which is not optional.
Give
TaskandRendezvous.Ownera stable, role-prefixed owner key derived only from the connection-resolvedPrincipal:Role.PRIMARYwith a non-null name →"leader:" + name"worker:" + terminal,"architect:" + slot,"collaborator:" + namenull, so today's escape is untouchedThree things I am adopting because the reasoning behind them is sound, not because an architect said them:
opuscollide with the lead namedopus. And key only a configured lead on its name — an architect must keep terminal ownership, or a newly spawned architect in the same slot inherits its predecessor's tickets and open turns. That is the one way option 2 could end up worse than the bug.Principal.describe(). It returns the constant"primary"for the unnamed primary, which would destroy the null escape, and its javadoc scopes it to audit lines. Add a separateownerKey().The argument that actually decided it is B's failed-roll case, and it is the one I would have got wrong on my own. Option 1 is a one-shot write that can only run after a new terminal exists. If the bounded relaunch exhausts its retries there is no new terminal, the roll is terminally
FAILEDwith no retry, and the state stays owned by a dead terminal — unreadable by every terminal-bearing caller, permanently. Option 1 delivers nothing in exactly the case I asked about. Option 2 has no event to miss: ownership isleader:opuswhether a pane exists or not, so the next session to hold that lead's tab label can read it.Put plainly: after a failed roll, option 1 leaves the state owned by nobody; option 2 leaves it owned by the next session in that lead's tab.
B also noted there is no migration surface at all —
tasksis an in-memory map and ticket ids carry a per-boot nonce, so new code can only arrive via a restart that empties it. The change that looks larger on paper has zero compatibility cost.Option 3 survives as a warning, not a refusal
fleet_handover{action:"open"}should report how many tickets and open ask-turns the caller owns. Useful whichever fix lands. It must not refuse: a lead rolls when its context is full, which is the worst possible moment to be told it may not, and "drain first" is not achievable in bounded time when a worker is mid-brief against a 30-minute clock.Sequencing — and I am not accepting "deploy with the loss documented"
Unit 2 may merge before this fix. The fix must be in the jar before unit 2 is.
B's reason for refusing the documented-caveat option is better than mine, so I am taking it: the daemon actively pushes the lead toward the unsafe action.
LeadHeartbeatLoopcarries a context-high nudge, and the livebootstrapTextis already written for a handover. So the system tells the lead to hand over precisely when its context is full — and a context-full lead is the least reliable reader of a documented caveat. A mitigation that lives in the judgement of the agent with the least context left is not a mitigation.That is also the shape this whole ticket came from: a safe answer recorded on #705 expired when another ticket shipped. Deploying unit 2 with a second written caveat would recreate the same failure one level up.
What is NOT settled, and what nobody has measured
fleet_handover— an operator closing the tab and opening a new one with the same label — also inherits the tickets. B argues this is acceptable because whoever can set a lead's tab label already getsRole.PRIMARYfromCallerResolver.java:308, which already grantsfleet_spawnandfleet_stop, so inheriting tickets is strictly less than the label already grants. I find that convincing and I have not tested it. The untested case B named is a mislabelled tab with no agent behind it. Whoever implements this should probe it before relying on the argument.LeadLaunchercreating a fresh tab. If unit 2 somehow reuses the pane, this whole ticket is moot.creatorTerminal's rename will turn a wide set of tests red —MessageService.poll(String)alone is reported to have 44 test callers. That figure is from #705 comment 18676 and I have not counted it myself.ReplyPushLoopandLeadHeartbeatLoopdepend on herdr answeringagent_not_foundfor a closed pane's terminal. That is not verified forpane.close— the reported evidence is for/exit, a different teardown.rgsweep did not run (rgis not installed in a worktree); its table came from editor Grep and Read, which did.Implementation units, for whoever picks this up
Roughly A's split, with B's restriction folded in. Not yet briefed, and deliberately not started tonight:
Principal.ownerKey(), role-prefixed, lead-name only forPRIMARYwith a non-null name. Leavedescribe()alone. Cover named lead, architect, collaborator, unnamed primary and anonymous.Task,Rendezvous.Owner,poll,pendingAskandanswer. Acceptance: lead A's old terminal creates, A's new terminal polls and answers, lead B is refused both. Keep the two null rules distinct and say in one place why.PrimaryRegistry,ReplyPushLoopand the heartbeat keep a stable lead identity and resolve the terminal at send time. Acceptance: a reminder pending before the terminal changes reaches the new terminal.openreports owned tickets and open asks.The one item to carry to #726 now: unit 3's single-flight claim is keyed on the lead terminal, which is what unit 2 changes.
Units 1+2 delegated as one. And one addition to the design that neither architect proposed.
Delegated to
term_65d0789fa36a791(profilesol, branchworker/737-owner-key-ff061f-10), covering the identity unit and the messaging unit together. I combined them because unit 1 alone is about thirty lines and is meaningless without unit 2.Blast radius, measured rather than quoted
Architect B relayed a figure of 44 test callers from #705 comment 18676. I did not reproduce that number and I am not repeating it. What I measured on
mainjust now:creatorTerminalrefsMessageServiceTestonlyOwner.of/Owner.permits/Rendezvous.OwnersendAsync(call sitesFleetMcp5,MessageService10,FleetApp1)MessageServiceTest.javaSo the production change is genuinely small and the churn is large but mechanical and concentrated in a single test file. That is what made one unit the right size rather than two.
The addition:
ANONYMOUSgets its own key, and that pins the open #705 couplingBoth architects said to leave the
nullescape alone, and both separately flagged that nothing pins the coupling that makes it safe — namely thatAuthzdeniesANONYMOUStheTASK_READaction, so an anonymous caller never reachesownsTicket. #705 comments 18617 and 18676 both record it as unpinned.I think that is a
nulldoing two opposite jobs, and the fix costs one table row. Todaynullmeans both:One symbol, two states needing opposite handling. So
ownerKey()returnsnullonly for the unnamed primary, and"anonymous"forANONYMOUS— a value that matches no stored owner, because no anonymous caller can ever create a ticket.The gate then fails closed on its own. It no longer depends on
Authzrefusing anonymous first, which means the thing #705 recorded as unpinned stops needing a pin: it is no longer load-bearing. The brief requires a test that proves the refusal without relying onAuthz, and a mutation (anonymous returnsnullagain) that must kill it.This is a small widening of the unit's scope beyond what the architects proposed. I am recording it here rather than only in the brief, because it changes a safety property and the next reader should see the reasoning, not just the code.
What the brief withholds, deliberately
The worker is told not to touch
PrimaryRegistry,ReplyPushLoop,LeadHeartbeatLoop,LeadRollover's single-flight key, orAuthz's role table — those are the routing and rollover units.LeadRolloverespecially, because #726 unit 2 is editing that file right now.It is also told there is no transfer primitive, with the instruction to stop and re-read the decision if it finds itself adding a setter.
Required acceptance, so it cannot be faked green
The headline test must use a terminal that actually differs across the roll. A fake returning the same terminal would pass while proving nothing, and that is the single most likely way this unit comes back green and wrong. Six mutations are required, each with a named test that must die, including "a lead is keyed on its terminal again" and "an architect is keyed on its slot name".
Remaining units, not yet briefed
PrimaryRegistry,ReplyPushLoop, the heartbeat: hold a stable lead identity, resolve the terminal at send time.LeadRollover.fleet_handover{open}reports owned tickets and open asks.Deployment gate, restated
The running daemon is on a jar that predates all of today's merges. I am not redeploying until this fix and #726 unit 2 are both in the tree, for the reason in the decision comment: the daemon's own context-high nudge pushes a lead toward a handover exactly when its context is full, so a documented caveat is not a mitigation.
Unit 3 (routing) is read and specified, but NOT delegated yet. It would collide with #726 unit 2.
I read the routing path rather than briefing it from the architects' summary, and it is both narrower and sharper than "these classes hold stale lead addresses". One finding changes the unit, and one changes the schedule.
The schedule first: do not delegate this until #726 unit 2 merges
The fix needs the live-lead map injected into the registry or the loops, and that wiring lives in
Fleetd.javaandFleetdAssembly.java. #726 unit 2 is editing both of those files right now — it rewiresFleetd.leadRolloverand hoists theLeadLauncheratFleetdAssembly.java:304. Two workers in those files is a guaranteed conflict for no gain, so unit 3 waits. Nothing else in the #737 list is disjoint from it either, so unit 3 is next in line after unit 2 lands.What
leadByTargetdoes on a roll: it already self-healsReplyPushLoop.resolveLiveLead(:443-452) already probes the recorded lead before trusting it, and forgets the binding when herdr affirmatively answersagent_not_found(fleetd #368). After a roll the old pane is genuinely closed, so a worker's stale binding is dropped on the first tick and resolution retries. That half is fine and must not be touched.isLivenarrowing death to the one affirmative code is also correct and is exactly the care #359 asked for — leave it alone.The actual defect: the fallback is returned without being probed
The second
nudgeTargetForfalls through toPrimaryRegistry's singleterminalslot. That slot holds the dead lead's terminal too, and it is returned with no liveness check at all. The method is namedresolveLiveLeadand its second return is not a live lead.So the sequence on a roll is: drop the stale per-worker binding correctly, then hand back a different stale address unchecked.
Why the single slot does not recover on its own
recordPrimarySingleton(FleetMcp.java:728) is reached from exactly two call sites,FleetMcp.java:467and:546— the send and spawn handlers. Measured withgrep -n 'recordPrimarySingleton\|registry.record(' mcp/FleetMcp.java. So the slot is refreshed only when a lead sends or spawns.A freshly bootstrapped lead reads its handover file first. Until it calls
fleet_sendorfleet_spawn, the slot still names the pane that was closed. During that window:LeadHeartbeatLoop(:310,:315,:370) readsprimaryTerminal()for the context-pressure notice and aims it at the same dead paneNeither loses data — the durable inbox still holds the reply and
fleet_pollstill works. What is lost is every push notification, which is the whole point of the feature. The new lead has to already suspect there is something to collect.The shape of the fix
The daemon already knows which terminals are live leads:
LeadTabScanner, the sameSupplier<Map<String, String>>of terminal to name thatFleetd.leadRolloveris handed. The single slot is a value learned from call traffic while an authoritative source sits right next to it. So:resolveLiveLeadreturns must have passedisLive, or the method is misnamed.LeadHeartbeatLoopask for the current terminal rather than the learned one.Scope limits for whoever takes this
ownerKeyorMessageService. Unit 1+2 is building a stable identity for comparison — "may this caller read this ticket". This unit needs resolution — "which pane do I nudge now". Only a lead can be rolled, so only the lead case needs the indirection, and reusing the comparison key here would drag an architect and collaborator case into a question that does not have one. Two deliberately separate notions; say so in the javadoc so the next reader does not merge them.isLive. AnyRuntimeExceptionother thanagent_not_foundmust keep meaning "still live".One thing I have not checked
I have not run a roll with a worker mid-delegation and watched where the nudge went. Everything above is read from the code, not observed. The window is real on the code's own terms, but I am not reporting a measured incident.
Units 1+2: the member died on a usage limit — the work is safe, and here is where it is (2026-10-04)
Read this before you pick units 1+2 up. The work is nearly done and it is not committed. Do not start it again from nothing.
What happened
The member on profile
solstopped mid-task. Its pane went quiet and the bridge recordedstate: done, which is a turn boundary and not a finish. The real cause came from the ticket poll:fleet_profilesthen showed why it takes two profiles with it:solandterrashare the credentialopenai-shared, so one exhaustion quarantined both for about 29 minutes. I stopped the pane, because an exhausted credential cannot do any more work and a live pane could write into the tree while I was building it.Where the work is
Branch
worker/737-owner-key-ff061f-10, in the worktree/Users/dai.ha/LTMS/.bridged-worktrees/ac9bb2-10.fleetd/src/test/java/dev/ltms/fleet/auth/PrincipalTest.java(31 lines, untracked).auth/Principal.java,mcp/FleetMcp.java,msg/MessageService.java,msg/Rendezvous.java,rest/FleetApp.java.e3050ef.origin/mainis nowaabecce, so the branch is four commits behind.The resumable member id, if anyone wants to continue that conversation once the quarantine lifts:
ses_ef7ec1b66ffesitLBWRn0Kk5Er. It needs profilesolagain — a resume cannot cross backends.The core of it is already right
Principal.ownerKey()is implemented, and it handles theANONYMOUScase this ticket asked for:The unnamed primary returning
nullis deliberate and matches the design on this ticket: it keeps the message layer's primary-wide ticket rule for a caller that carries no name.I have not reviewed the other ten files yet, and I am not calling this correct. I am recording that it exists and looks coherent.
What still has to happen
origin/maininto the branch first. The branch is four commits behind. Those four commits touchsession/SessionManager.java,.claude/skills/handover/SKILL.mdand the wiki pointer — none of the five main files above — so no conflict is expected. Expected is not measured: do the merge, then build.One hygiene note for whoever reuses that worktree
Delete
fleetd/target/surefire-reportsbefore trusting a test count. A reused worktree keeps XML from earlier runs, including for test classes that no longer exist, so the count reads high. I cleared it before my own build.An aside that cost us this turn
exhaustionDetectionArmedistruefor onlysolandterra:So a usage limit on any other profile can never be classified or quarantined, however many times it happens. This loss was detected because it landed on an armed profile. The same failure on
sonnetwould surface as an opaque dead member, and a lead would likely respawn straight back into it. That is a separate concern from this ticket and I am not fixing it here — noting it so it is written down somewhere.Lead review of units 1+2's main-code diff (2026-10-04)
I read all five main files myself. I said earlier I had not, so here is the result. No blocker. One finding worth carrying, and two things done right that are easy to get wrong.
Base for this review: the branch fast-forwarded to
aabecce,mvn clean installgreen — 2094 tests, 0 failures, 0 errors, 0 skipped, across 177 classes. That ismain's 2089 plus the 5 in the newPrincipalTest, and the class count ismain's 176 plus one. It reconciles from both ends.Done right: the three null states stay three
This is the part I expected to be wrong, and it is not.
Three states, not two, and they are kept apart:
OwnerisnullOwnerisUNNAMED_PRIMARY(keynull)nullatownsTicketSo
permitsandownsTickettreatnullin opposite directions on purpose, and the diff says so in the javadoc. That is correct, and it is the distinction a single sentinel usually loses.ANONYMOUS -> "anonymous"is also the right call. A non-null key means an anonymous caller can never fall into the null-means-unnamed-primary rule, even thoughAuthzalready refuses it earlier. Two gates, same answer.Done right: the primary's key is its NAME
This is the whole unit. The key survives a handover because the name does and the terminal does not.
COLLABORATORuses its name for the same reason.WORKER,ARCHITECTandOBSERVERuse the terminal, which is right — they have no configured name, and none of them outlives its pane.The finding: the type-level guard is on 1 of 6 entry points
sendAsynctakes aPrincipaland derives the key inside, and its javadoc gives the reason:That reason is good. It is just not applied anywhere else:
Five of six take a bare
String, so a future call site writingcaller.terminal()instead ofcaller.ownerKey()compiles. The javadoc on the guarded one reads as if the hazard is handled, which makes the other five easier to get wrong, not harder. Same shape as a grant applied at one gate of two.Why this is not a blocker: every confusion direction fails closed. I worked through each one:
poll(ticket, caller.terminal())— a terminal compared against a storedleader:opusnever matches, so the caller is refused its own ticket. Loud.answer(turnId, ..., caller.terminal())— same,permitsreturns false and the turn is not resolved. Loud.None of them widens access. That is the safe direction and it is why I am not holding the PR for it.
What I suggest, and am not asking the current member to do (its brief says change no behaviour): take the
Principaloverload pattern tosend,answer,pollandpendingAskas a follow-up, or drop the claim fromsendAsync's javadoc so it does not describe a protection the API mostly lacks. One fact, one place — right now the sentence is true of one method and reads as true of the layer.Correctly left alone
recordPrimarySingletonandprimaryRegistry.recordDelegationstill take the terminal:That is right for this unit. The registry's job is nudge routing to a live pane, and a pane is addressed by terminal. Making it owner-keyed is unit 3's problem, not a miss here.
Still outstanding before I merge
The mutation evidence. It has not been produced, and a green 2094 proves the code runs, not that the tests would notice it breaking. A member is on that now.
Review of PR #741: one finding, adjudicated — not a blocker, but it names a real gap
A reviewer looked at the ownership semantics of PR #741. It returned one finding,
high:ownsTickettreats anullcaller key as a wildcard, so a caller with no key reads every ticket.I checked the claim myself. Here is what is true, what is not, and what I am doing about it.
The finding is real, but it is NOT introduced by this PR
The reviewer quoted
MessageService.java:1463and the namecallerTerminal. Those aremain's line andmain's name, so it reviewedmain, not the diff. The behaviour it describes is the same on both sides:So the PR changes what the key is (terminal → owner key) and keeps the
nullrule exactly as it was. The PR's own javadoc states the rule on purpose: "Anullcaller key is the unnamed primary and may read every ticket."So this is not a regression and does not block the merge. The fix the reviewer asks for — "replace terminal-based ownership with a stable owner key derived from role+identity" — is what this PR already did. Only the
nulltreatment is left.One part of the finding is wrong
The reviewer wrote that an "unauthenticated caller" can poll tickets. It cannot.
Authz.permitsrefuses beforepollis ever reached:Both surfaces gate first:
FleetMcp.java:1246sits behind theTASK_READcheck, andFleetApp.java:903callsallow(...)before:907. An anonymous caller also does not even get anullkey —Principal.ownerKey()mapsANONYMOUSto the string"anonymous".What the gap actually is, measured
In production, a
nullowner key has exactly one source. I checked every.poll(in main sources on the branch:The single-argument
poll(String ticket)overload has no production caller. It is a test seam. So the only live route to anullkey isPrincipal.ownerKey()returningnull, and that happens for one case only:That is the unnamed primary: a caller the daemon resolved as primary with no pane name — token auth, or loopback trust. Every other role returns a non-null key.
So the gap is: an unnamed primary reads every lead's ticket, including the full reply text. The live-turn path refuses the same caller —
Rendezvous.Owner.permitsneeds an exact match and never treatsnullas a wildcard. Async tickets are therefore less protected than live turns, which is the asymmetry the reviewer spotted. That part stands.The shape of it: one
null, two meaningsThe
nullkey carries two different states that need opposite handling:nullpoll(String)ownerKey()Today both get "no check". This is the same conflated-sentinel shape as #497: one symbol, two states, one handling. The fix is a third state — keep the internal bypass explicit, and stop the unnamed primary sharing its value.
How this relates to #705
#705 offers three fix shapes for ticket walking. Its option 2 is "give tickets an owner check", and that is what this ticket is. This finding is the hole left inside option 2: the owner check now exists, and one caller still skips it. #705 can close on option 1 or 2 without this being fixed, so I am keeping the two separate rather than merging the tickets.
Decision
nullrule untouched and documented.nullwildcard becomes #737 unit 6, specified below with the rest of the units.Severity, stated plainly: lower than the reviewer's
high. An unnamed primary already holdsSPAWN,STOP,SENDandDRAINover every session on this host, so it can reach most of the same content by other routes. It is still worth closing, because the live-turn path already refuses it and the two paths should agree.I checked the code in this comment myself, on
origin/mainataabecceand onorigin/worker/737-owner-key-ff061f-10. I did not run a live probe from an unnamed primary.Unit 6 — stop the unnamed primary sharing the internal bypass
Comes from the PR #741 review above. Do this after units 1+2 are merged, because it edits the same method.
What changes
MessageService.ownsTicketmust stop treating anullowner key as "read everything". Today onenullmeans two things: an internal caller that has no identity to check, and a real authenticated caller that has no name. Give them two different values.Shape to use — a named constant for the internal bypass, so the caller says what it means:
poll(String ticket)overload passes an explicit "no check" marker, notnull. It has no production caller (measured above), so this is a test seam saying so in code.nullfromPrincipal.ownerKey()is then just another key. An unnamed primary matches only a ticket it created itself, the same rule every other role already follows.Do not simply delete the
nullbranch. A ticket created by an unnamed primary recordscreatorOwner == null, so the comparison must still treat twonulls as a match, or a token-auth primary loses its own tickets. That is the bug this whole ticket exists to prevent, in a new place.Apply the same change to
pendingAsk(:1793on the branch), which callsownsTicketwith the same key.Why
Rendezvous.Owner.permitsalready refuses anullcaller for a named owner, so the live-turn path and the ticket path disagree about the same caller. One of them is wrong, and the stricter one is right.Tests
forbidden:reason and no reply text.pendingAsk.Mutation evidence, not a green build
nullcreators compare unequal ⇒ the "its own ticket" test dies.Show each one red, then revert and show it green. A mutation that kills nothing means the test is vacuous.
Out of scope
Authz'sTASK_READrow is not part of this. Whether an unconfigured pane should holdTASK_READat all is #705, and the two fixes are independent.Unit 6 addendum — two existing tests pin the behaviour you are changing
Units 1+2 are merged as
cd1f04c. While checking the merged tree I found that thenullwildcard is pinned by two tests. Whoever takes unit 6 will see them go red and must invert them, not weaken or delete them:Both currently assert the wildcard is correct. After unit 6 they must assert the opposite: an unnamed primary reads only a ticket it created. Rename each one to say what it now protects.
Two more that must keep passing unchanged, because they are the reason the
nullbranch cannot simply be deleted:This is the trap in this unit. A worker that changes
ownsTicketand then finds two red tests has an easy wrong move available: edit the two tests until they are green again. That turns a security fix into a test weakening, and the build stays green either way. Invert them deliberately, and say in the reply what each one asserted before and after.Unit 4 addendum — the rollover single-flight key
Unit 6 is merged (
6ab3a81+ lead follow-up7467ffa, PR #744 closed). Unit 4's one-line spec needs expanding before it can be delegated, because the obvious implementation introduces a worse bug than the one it fixes. Writing it here so the worker can re-read it.The defect, stated precisely
rollingByTerminal(LeadRollover.java:372) is the single-flight lock for a lead roll. It is keyed onp.leadTerminal()— the pane address. A roll replaces the pane, so the fresh lead has a different terminal. The lock therefore stops being a per-lead lock: once the fresh lead is up, a secondopen()+confirm()from it lands on a different map key and is approved while the first roll's continuation is still running.The two architects disagreed on exactly this row, and both were right about different properties:
:569and the releases at:601/:647all usep.leadTerminal(), so the release is correct and nothing leaks;The fix is the second half only. Do not "fix" the release — it is not broken.
The hazard that makes this harder than it looks
leadNameForTerminalis wired inFleetd.java:962as:It reads the live lead roster. The roll kills the old pane, so by the time
runRollover'sfinallyruns,leadNameForTerminal.apply(p.leadTerminal())returns null.So the key must not be computed independently at the claim site and the release sites. If it is, the release computes a different key from the claim, the claim is never removed, and that lead can never be rolled again —
confirm()answersROLL_ALREADY_RUNNINGforever. Resolve the key once and carry it to both release sites.What must change, and what must not
Change: key the single-flight claim on the lead's configured name.
Keep, because these are genuinely about the pane:
:552NOT_YOUR_ROLLOVERmust keep comparingp.leadTerminal()tocallerTerminal. That gate means "only the exact pane that opened this request may confirm it". Do not touch it.runRollover's use ofp.leadTerminal()from:653onwards as the pane to tear down stays as it is.Fallback: if the lead name is null or blank (the terminal is not in the live roster), fall back to keying on the terminal. That is at least as strict as today, must not throw, and must not skip the claim.
Rename the field so it stops saying
Terminalonce it is keyed by name.Acceptance criteria
confirm()for the same lead, while the first roll's continuation is in flight, is refusedROLL_ALREADY_RUNNINGeven when the second request was opened from a different terminal for that lead. New behaviour, needs a new test.leadNameForTerminalto return null for the old terminal at release time, then show a laterconfirm()for that lead is approved, not refused. Without this test the regression above is invisible.continuationRunner-rejected release at:601still works; its existing test stays green.NOT_YOUR_ROLLOVERbehaviour unchanged, existing test green.leadNameForTerminalinstead of using the carried key → criterion 2's test must fail;p.leadTerminal()→ criterion 1's test must fail.mvn clean installrun fromfleetd/— there is no root pom, and running it one directory up fails withno POM in this directory. Unpiped, full output read,rm -rf target/surefire-reportsfirst.Implementation shape is the worker's call. Two that work: a token → resolved-key map cleaned up on the same paths that release the claim; or carry the resolved key on
PendingRollover, resolved inopen()at:462while the lead is certainly still live.PendingRolloveris consumed in production only atFleetMcp.java:1514, so adding a component is contained — but checkhandoverOpendoes not put it into the tool response.Out of scope
MessageService,Principal,FleetMcp's ownership javadoc,ReplyPushLoop,LeadHeartbeatLoop. Unit 5 also editsLeadRollover, so it is deliberately not running at the same time.Unit 5 addendum —
fleet_handover{open}reports outstanding tickets and open asksUnits 3 and 6 are merged (
d438a74+803c91e,6ab3a81+7467ffa; PRs #744 and #745 closed). Unit 4 is running. This is unit 5's full spec.First, a correction to comment 2's quoted code
Comment 2 quotes
Rendezvous.Owner.permitsas:That is not the deployed code.
Rendezvous.java:78-92onmaintoday is:It is keyed on the owner key, not the terminal. This matters directly for unit 5, because I started writing this spec around the hypothesis that a roll makes a pending ask unanswerable — the fresh lead has a new terminal, so a terminal-keyed gate would refuse it. That hypothesis is false, and I only caught it by reading the file instead of trusting the quote. Comment 2's severity argument about a refused answer still stands for a genuinely different caller; it does not apply to a rolled lead.
Why unit 5 is needed — stated correctly
A lead roll restarts the lead's process, not
fleetd. The daemon keeps itstasksmap and its rendezvous registrations across the roll. Both are gated on the caller's owner key:MessageService.ownsTicketandRendezvous.Owner.permitseach comparePrincipal.ownerKey(), which for a named lead isleader:<name>and does not change when the pane changes.So the fresh session keeps the authority to poll those tickets and answer those asks. What it loses is the knowledge — the ticket ids and turnIds lived in the outgoing session's context and nowhere else.
fleet_handover{open}is the one call made while the outgoing session still holds both, so it is the right place to list them for the handover file.Do not implement this as a warning that a roll "drops" tickets or asks. It does not.
Scope
MessageService.java, thehandover/handoverOpenpath inFleetMcp.java, and their tests.Not in scope:
LeadRollover.java. Unit 4 is editing that file concurrently. If the change seems to belong there, stop and ask rather than editing it.Required change
MessageService: a public method listing the delegations a given owner key created that are not yet collected, plus the open asks it may answer. Follow the existingtasks.values()iteration pattern (:509,:648,:802,:1808). Return a record — formatting belongs in the MCP layer. Filter with the same ownership ruleownsTicketalready uses; do not introduce a second rule that can drift from it. Do not special-caseINTERNAL_NO_OWNER_CHECKhere.FleetMcp: thread the caller's owner key intohandover(...)/handoverOpen(...). The established pattern isprincipal(exchange).ownerKey()— see:525and:536. Do not build a terminal-to-owner lookup; the principal already carries it.handoverOpen's JSON gains the outstanding tickets (ticket id, phase, target) and the open asks (ticket id, turnId, worker session).token,handoverPathandrequestedAtMillisstay present and unchanged — thehandoverskill reads them.handoverOpenby theisBlank(callerTerminal)check. Leave that refusal exactly as it is.Acceptance criteria
open()from a lead that created two async tickets reports both, with their phases.open()from a lead whose worker is paused infleet_askreports that ask'sturnId.token,handoverPath,requestedAtMillisstill present and unchanged; existinghandoverOpentests green.mvn clean installfromfleetd/— there is no root pom. Unpiped, whole output read,rm -rf target/surefire-reportsfirst.Unit 5 correction — a DONE, uncollected ticket is the case that matters most
This supersedes the filter choice in the "Unit 5 addendum" comment. Everything else in that comment stands.
Unit 5 implemented "not yet collected" as
!task.future.isDone(), so a DONE or FAILED ticket is excluded. The worker flagged this itself and gave its reasoning:That reasoning conflates authority with knowledge, which is the exact distinction the unit 5 addendum drew. Authority survives the roll, because the owner key is
leader:<name>. Knowledge does not — the ticket id lived only in the outgoing session's context. So "the lead can poll it anytime" is true and beside the point: after the roll it does not know what to poll.And the excluded case is the perishable one.
pruneTerminalTickets's own javadoc says so (msg/MessageService.java:1511-1533):The prune removes any task where
future.isDone() && completedNanos < cutoff. So a finished, uncollected ticket is destroyed on a timer, and its report is unrecoverable. A PENDING ticket, by contrast, is not going anywhere — the worker is still running.The filter is backwards for the purpose. The tickets most worth writing into a handover file are the finished ones nobody has read yet.
Required change
Report every ticket the caller owns that is still present in
tasks, with its real phase —PENDING,ASKING,DONE,FAILED. Drop the!task.future.isDone()condition. Keep usingownsTicketas the only ownership rule.A ticket already collected still sits in
tasksuntil the TTL prunes it, so it will appear too. That is acceptable and better than the alternative:MessageServicehas nocollectedflag to filter on — collection is recorded inReplyPushLoopviaticketCollected(...), and reaching into the push loop from here to hide a row is not worth a new coupling. The phase already tells the reader what it is looking at.Extra acceptance criterion
outstandingTicketswith a terminal phase. Add a mutation for it: restore the!task.future.isDone()filter and show this new test goes RED, then revert to GREEN.Criteria 1-6 from the addendum are unchanged, and the existing tests for them must stay green.
Shipped and deployed — closing
The decision this ticket asked the next lead to settle: fix shape 2 (key ownership on the stable identity), plus the reporting half of fix shape 3. Shape 1 (reassign
creatorTerminalduring the roll) was not taken, because shape 2 removes the reason for it.Six units merged to
main:efd9cdbf4176ae,803c91e8a1d73b70a735b,d41aff4fleet_handover{open}reportsoutstandingTicketsandopenAsks6ab3a81,7467ffaA named lead now owns a ticket through
Principal.ownerKey(), which isleader:<name>. That value does not change when the pane changes, so the refusal this ticket predicted cannot fire.mainis at7f0c4a8with 2118 tests and 0 failures.Docs updated with the code:
wiki/9-Implementation.md(the ownership rules and the line refs),wiki/11-Features.md(the rollover cycle and a third gotcha), and.claude/skills/handover/SKILL.mdstep 1, which now tells the outgoing lead to copyoutstandingTicketsandopenAsksinto the handover file.It is deployed, and I checked the jar itself
scripts/redeploy-fleetd.sh --yesran today. The daemon now running is pid 18094, startedMon Oct 5 06:11:16 2026, and it holdsfleetd/run/fleetd.jaropen, hash93215c19c8e6. I read the shipped code out of that exact file rather than trustinggit log:#726 unit 2 — the real process restart, which is the change that would have introduced this regression — is also in that jar (
4ffe49f). So the regression and its fix went live together, which is what this ticket asked for.What is NOT measured
Nobody has rolled a lead on the restart path. So the end-to-end claim — a fresh lead, new terminal, same name, collects a ticket the previous lead created — is still read from code and covered only by unit tests. I have not seen it happen.
I tried a non-destructive probe of the deployed behaviour:
fleet_handover{action:"open"}followed straight bycancel, which creates a token and changes nothing else. My own permission classifier refused it as modifying a shared resource, and I did not route around it. The real proof needs a deliberate roll of a live lead, which is the operator's call, not mine.That gap is already tracked in #595, which records that the roll has never been observed working end to end. I am not filing a second ticket for it. The matching gotcha in
wiki/11-Features.mdstays as it is, because it is still true.Follow-ups found while doing this, filed separately
validate*method is wired by its name, so a rename silently removes a boot check.FleetApp.java:885,LeadRollover.java:96, and one wiki bullet). All the same shape: the definition was updated, the prose pointing at it was not. Ahuntersweep for{@code X#method}references that no longer resolve would probably find more.Closing.