Lead session rollover: the lead asks for a handover, fleetd verifies it, clears the pane and boots the next lead #480
Open
opened 2026-09-11 01:09:38 +02:00 by ltms
·
10 comments
No Branch/Tag Specified
main
worker/fleetd-612-unita-87807e-1
worker/612-b3-mcpwirings-da2b58-3
worker/612-b2-cb185-176d3a-2
worker/612-b1-completion-457459-1
worker/612-agaps-73a926-2
worker/608-sleeps-3a64ff-3
worker/621-b4520b-1
worker/618-b83894-2
worker/fleetd-615-e05481-5
worker/lead-autocompact-5f1ab2-3
worker/fleetd-613-f85deb-3
worker/fleetd-608-flaky-nudge-test-d0c2d1-3
worker/lead-context-gauge-ad404f-1
worker/gauge-wiring-9158c1-4
worker/redeploy-slowstart-ead0e5-5
worker/charter-bytes-13668c-6
worker/rollover-outcome-291483-2
worker/589-f64303-2
worker/593-1a8025-5
worker/589-fcd2aa-1
worker/568-9fdaa2-3
worker/571-attempted-outcome-5739f7-2
worker/581-completionresolver-cas-sites-0542b7-6
worker/562-loop-health-wiring-test-99611c-5
worker/562-surface-loop-health-7df5cc-4
worker/575-waiter-cleanup-sites-62ad80-1
worker/572-answer-lock-release-46a9ae-5
worker/567-probe-channel-leak-a38fc5-6
worker/551-record-before-send-7cbf56-1
worker/561-listener-fanout-survives-a-throw-61d538-2
worker/555-redeploy-main-flow-seam-65c2f5-2
worker/556-injector-owns-registration-e027a5-1
worker/552-post-restart-mktemp-abort-bc2672-4
worker/553-onstatus-completion-leak-0da881-2
worker/550-shasum-linux-196132-1
worker/538-loop-dies-on-error-4a5eeb-6
worker/426-health-coverage-ef1fd4-4
worker/504-failed-reported-clean-3cfd66-3
worker/537-capturedlog-close-e4c437-2
worker/459-broken-link-targets-cadc17-5
worker/535-appender-leak-fe74c1-1
worker/512-part2-shutdown-detection-434701-9
worker/529-logger-level-sweep-2a5533-8
worker/528-drain-gate-call-site-5de83d-7
charter/forge-mcp-vs-token
worker/521-swap-guard-unpinned-28e931-5
worker/519-probe-test-harness-d25ab8-4
worker/525-logger-level-leak-1b4eb0-6
worker/518-fleetmcp-resolver-wiring-8ef96c-1
worker/512-drain-complete-line-7edd71-3
worker/517-abort-branch-and-jar-id-41b641-2
worker/500-9e52c9-3
worker/509-4912f4-2
worker/511-9a4b23-1
worker/493-479f45-2
worker/505-03f8b2-1
worker/492-followup-detect-unclear
worker/501-a31fa0-7
worker/498-451d1c-5
worker/494-1015ce-2
worker/492-209647-1
worker/489-001902-2
worker/480-relative-handover-path-906323-1
worker/480-b-handover-skill-45bf1f-5
worker/474-followup-source-pin-f54a55-17
worker/474-charter-check-on-reload-f54a55-17
worker/466-quarantine-repeatcount-report
worker/393-opencode-skill-seeding-71854b-13
worker/469-canonical-tool-names-2a472a-16
worker/466-quarantine-escalation-5ae9c1-15
worker/446-hot-exhausted-pattern-0af580-6
worker/464-charter-tool-name-guard-a85635-12
worker/463-listfleet-default-fails-open-f1c76c-11
worker/458-invariant-5-by-purpose-862f9a-10
worker/439-coordinator-row-gate-bc032a-8
worker/449-herdr-protocol-576015-4
worker/450-abstract-spawn-599e1c-5
worker/437-ack-refuses-177d91-1
worker/444-placement-window-feb56a-2
worker/440-helddurable-derived-d462d7-13
worker/425-rework-placement-resolve-c58ba1-9
worker/421-lead-peek-held-msgs-cdbad2-10
worker/435-fixed-policy-cap-fe11de-12
worker/422-gate-state-observability-9e79d6-11
worker/431-memberregistry-live-readers-cdbad2-10
worker/424-architect-slot-hot-038b41-7
worker/422-model-gate-spawn-c29f48-6
worker/425-default-profile-live-f55534-8
worker/415-coverage-wording-2cbf9c-5
worker/416-3ad1da-1
worker/418-588283-3
worker/deterministic-stamp-race-409-3cb7b6-10
worker/armed-reads-live-config-404-ed931f-9
worker/reply-peer-refusal-391-5a34bd-7
worker/models-allowlist-aa9e9b-3
worker/ttl-stamp-race-399-f1122f-8
worker/scrub-receipt-400-316b3e-5
worker/exhaustion-detection-395-105105-6
worker/scrub-abort-394-316b3e-5
fix/scrub-uid-abort
worker/task-scrub-517574-2
worker/t386-clock-bd5b78-4
worker/t384-scrub-813790-5
worker/t381-cc-748314-2
worker/t373-336973-2
worker/t365-3920c5-3
worker/t358-6e989b-1
worker/t355-8b321c-1
worker/fleetd-369-hermetic-git-tests-e8b19a-3
worker/fleetd-368-stale-lead-binding-f5682e-2
worker/fleetd-360-deploy-units-0d3793-1
worker/359-dead-lead-tabs-f1253b-4
worker/362-worktree-skills-c03e51-3
worker/361-coord-visibility-655144-1
362-plugin-visibility-and-drift
worker/errscan-bed2ca-2
worker/amqp-log-identity-bed2ca-2
worker/withdefaults-guard-561704
worker/sleepguard-82076d-1
worker/fd334-9ee1b6-5
worker/fd348-f1ab27-4
worker/fd335-a71c35-1
worker/fd342-174a17-2
worker/fd345-490d0f-3
worker/fleetd-337-5ec7d4-21
worker/fleetd-341-af5a6b-24
worker/fleetd-339-5ca0a2-23
worker/fleetd-338-83a4a1-22
worker/fleetd-333-281f46-18
worker/fleetd-329-11bdbb-16
worker/fleetd-330-2770fb-17
worker/fix-326-50506e-15
worker/fix-324-3e9bbf-14
worker/fix-323-b8287d-13
worker/fix-316b-bd0860-11
worker/fix-318-76ca36-9
worker/fix-317-486aec-8
worker/fix-315-ce47c5-6
worker/fix-307-275890-6
worker/fix-308-b4f664-7
worker/fix-309-ec3939-8
worker/fix-310-7a3974-9
worker/fix-302-52ad0e-9
worker/fix-298-ce1acb-8
worker/fix-297-66bd11-7
worker/fix-296-104622-6
worker/fix-293-bare-closetab-eb22b5-3
worker/fix-280-gone-ask-lapse-bca98e-2
worker/fix-290-reapidle-guard-coverage-9b0dd1-1
worker/fix-285-trust-seed-8f3565-10
worker/fix-284-backend-error-seat-85912c-11
worker/fix-282-chained-ask-e6d0bb-8
worker/fix-283-teardown-leaks-f40dfa-9
worker/fix-281-pin-handler-actions-4921ac-7
worker/audit-rendezvous-lifecycle-d072ae-2
worker/audit-health-placement-1a2476-6
worker/audit-teardown-exits-e207a5-3
worker/audit-launcher-asymmetry-27e370-4
worker/audit-rest-authz-6ca53c-5
worker/investigate-275-abandon-asking-fdef52-8
worker/fix-274-worktree-leak-b0095d-7
worker/fix-273-exhausted-pattern-9665b5-6
worker/fleetd-267-model-check-bd8068-1
worker/fleetd-131-archunit-18b834-7
worker/fleetd-266-sshagent-rename-a014ff-6
worker/fleetd-184-uid-claim-8e1f31-4
worker/fleetd-184-warn-b381ee-10
worker/fleetd-184-docs-be1d12-9
worker/fleetd-257-9bf010-7
worker/fleetd-103-23a113-6
worker/fleetd-247-342356-5
worker/fleetd-116-04dea8-4
worker/fleetd-252-a830e0-3
worker/fleetd-111-7e8673-9
worker/fleetd-155c-f8ef4b-8
worker/fleetd-176-b928ca-3
worker/fleetd-249-7a7878-2
worker/cb248-composition-root-b-9acdf7-15
worker/cb148-envrc-default-fa6c82-12
worker/cb201-unit5-wiring-6c12e6-8
worker/cb241-fallback-echo-1175e9-11
worker/cb149-trust-dialog-2392a5-9
worker/cb134-148-overlay-visible-c9b986-10
worker/cb234-session-id-keyed-04e1fc-1
worker/cb201-unit3-nudge-abdf5c-6
worker/cb201-unit2-policy-c1102c-5
worker/cb201-unit4-outcome-a13bfa-7
worker/cb201-unit1-classifier-91b9b1-4
worker/cb201-227-refine-831980-3
worker/cb175-model-readback-0f085f-1
worker/cb222-charter-tmpdir-17f013-1
worker/cb226-architect-slot-race-cd3aa8-3
worker/cb224-worktree-root-group-024523-2
worker/cb-123-role-demotion-c600f7-2
worker/cb-219-opencode-roots-1f677e-1
worker/cb214-claude-session-id-b9eab4-4
worker/cb213-zdotdir-wrong-process-dd6de4-3
worker/cb211-exhaustion-classification-9546e0-2
worker/cb137-ambiguous-task-4df3d8-4
worker/cb209-agentsessionid-4dfdb6-2
worker/cb185-hostenvnames-2692b5-3
worker/cb206-opencode-sqlite-128718-2
worker/cb185-worktree-group-fc0c99-1
worker/cb-137-ask-ticket-e7760c-2
worker/cb-172-broker-uri-d36ae4-4
worker/cb-175-model-readback-76ead6-3
worker/cb-161-pane-ancestry-293510-1
worker/cb-164-rebase-885863-8
worker/cb-164-empty-scrape-false-success-1a80af-3
fix/cb-197-ticket-ttl-from-completion
worker/cb-189-remote-url-coverage-4692f3-1
worker/cb-185-blockers-027756-4
worker/cb-192-gap-log-11b631-2
worker/cb-633-fix-5f4396-3
worker/cb185-router-d6436d-3
worker/cb185-router-routing-gaps-9e9d33-3
worker/cb185-paneids-992586-2
worker/cb-633-allow-list-union-ed374b-1
worker/cb-157-credential-in-remote-url-496e44-2
worker/cb-641-health-herdr-evidence-8f1f54-6
worker/cb-640-health-msg-evidence-99c9cd-1
worker/cb-642-fleets-status-skill-bbbc40-5
cb-634-ide-mcp
worker/lead-comms-wiring-c014b9-7
worker/lead-mailbox-c19577-6
worker/autocompact-window-82bc2f-5
worker/cb-634-probe-18056f-4
worker/cb635-broker-urienv
worker/cb-632-config-retry-8e0efa-7
lead/cb-622e-claude-md
lead/cb-622-followup
worker/cb-622a-165dff-1
lead/cb-622d-opencode-mount
worker/cb-622b-717c67-2
worker/cb-622c-ab7759-3
worker/cb-617b2-20ca4b-3
worker/cb-617a-5c2f4a-1
worker/cb596-4e49ef-3
worker/cb586-10500c-1
worker/cb-606-b9343a-25
worker/cb604-1445f8-24
worker/cb582-477374-21
worker/cb584-8c2281-22
worker/cb600-e6b9a9-20
worker/cb602-ce257f-19
worker/cb601-b42837-18
worker/cb598-6c7ba7-17
worker/cb599-740fe4-16
worker/cb597-282224-15
worker/cb590fix-185e9a-10
worker/cb528-recovery-race
worker/cb594-96bead-8
worker/cb590-916766-2
worker/cb527-997d99-3
worker/cb592-env-leak-3cbf9c-1
worker/cb588-async-ticket-nudge-3218f7-5
worker/cb578b-9dcb13-6
worker/cb581-d24826-5
worker/m2-u5-ef8c42-15
worker/cb578a-516499-2
worker/cb576-01a04b-17
worker/cb579-lead-tab-acba06-20
worker/cb580-terminal-health-ed6058-21
worker/cb577-f36fdc-18
worker/cb573b-3db06f-16
worker/cb568c-f36fdc-18
worker/cb568-drop-cause-c3ac1c
worker/cb575-cancelled-notification-c3ac1c
worker/m4-sol-a2cbec-3
worker/cb574-async-ask-c3ac1c
worker/cb573-health-model-8ca857-14
worker/cb572-unknown-target-7f2e35-13
worker/u4-700706-9
worker/u3-b9fcb6-6
worker/u2-ef5b68-4
worker/u1-469dce-1-clean
worker/u1-469dce-1
worker/cb-564-health-events-70cf7e-2
worker/cb-565-recycle-drops-role-98e58f-3
worker/cb-563-missing-reply-df2866-1
worker/cb-562-readiness-gate-silent-6c23c9-3
worker/cb-560-architect-presence-da8155-1
worker/cb-561-architect-silent-off-a71cab-2
worker/cb-548-bind-architect-slot-fe1b8c-1
worker/parity-overlay-settings-5fb711-1
secrets-central-store
cb-559-hot-key-correction
cb-557-fleet-role-pools
worker/cb-553-maxload-explicit-spawn-305ee3-6
worker/cb-551-idle-lead-heartbeat-f1633c-1
worker/cb-544-drain-preserves-worktree-925fad-3
worker/cb-552-docs-sync-1cb9cf-4
worker/cb-548-rendezvous-guard-rebased
worker/cb-548-rendezvous-guard-116b53-10
worker/cb-548-authz-v2-586df6-8
worker/cb-548-authz-264363-5
salvage/cb-528b-codex-home
salvage/cb-528a-codex-launcher
CB-518-primary-flow
feature/peer-launcher-spi
cb-103-injector
v1.1.0
v1.0.0
Labels
Clear labels
blocked
needs-live-proof
ready-to-delegate
silent-default
Cannot start until something else lands. The body says what.
Merged and green, but never shown working on the running daemon. Not the same as done.
Scope, files and acceptance criteria are written. A worker can be briefed from the body alone.
A feature that compiles, passes tests, and ships turned off. Nine recurrences and counting.
No Label
Milestone
No items
No Milestone
Projects
Clear projects
No project
Notifications
Due Date
No due date set.
Dependencies
No dependencies set.
Reference: fleet/fleetd#480
Reference in New Issue
Block a user
Blocking a user prevents them from interacting with repositories, such as opening or commenting on pull requests or issues. Learn more about blocking a user.
Delete Branch "%!s()"
Deleting a branch is permanent. Although the deleted branch may continue to exist for a short time before it actually gets removed, it CANNOT be undone in most cases. Continue?
Why
A lead session fills up its context and then has to be replaced by hand. Today that means the
operator notices, asks the lead to write a handover file, starts a new session, and pastes a
pointer to the file. The fleet cannot run unattended across that boundary.
This ticket makes the cycle a fleetd feature: the lead says when it is ready, fleetd checks the
handover file is really there, clears the pane, and starts the next lead on that file.
What the operator chose (2026-09-11)
Three decisions, taken by the operator after the measurements below:
that can see its own context filling.
/clearin the same pane. Not a stop-and-relaunch.Measured facts that constrain the design
Every line below was measured on 2026-09-11 against
7b97aae.context numbers are two operator-configured ceilings:
autoCompactWindow(
config/FleetConfig.java:487, a spawn flag) andLifecycle.contextCap(
config/FleetConfig.java:868), which counts turns, not tokens. fleetd reads no Claude Codetranscript:
jsonlhas zero matches in the main sources (control: 4 files matchCLAUDE_CONFIG_DIR, so the search works). This is why the trigger has to be the lead.SessionManagerat all.lead/LeadLauncher.javanever calls it,so the idle reaper (
session/SessionManager.java:1002), the context cap (:965) and theshutdown drain (
:1068) can none of them act on a lead. There is no per-lead lifecycle objectto hang this on; it needs its own.
herdr/LeadTabScanner.java:247matches the tab label, andauth/CallerResolver.java:211grants
Role.PRIMARYfrom that map.terminal_idis expected to change on restart. A/cleardoes not even change the terminal, so identity is safe on this route.fleet_stopcannot target a lead (mcp/FleetMcp.java:461→SessionManager.stop, whichreads a registry a lead is never in). So there is no accidental-teardown path today.
System.nanoTime, which stops while the Mac sleeps. Onlyhealth/FleetHealthMonitor.java:47got a realtime clock, under #386. Any deadline in thisfeature must use a realtime clock, or it will freeze across host sleep.
Two live findings from a probe — these are the load-bearing ones
Probed on a throwaway sonnet member, using bridge tools only.
/clearsent over herdragent.promptreally does run as a slash command. It is not aninert paste. Proof is in the transcripts, not in a reply: the member's session ends with its
last answer at 06:05:04, and a new session is born at 06:05:04 whose line 5 is literally
<command-name>/clear</command-name>. Somember/ClaudeCodeLauncher.java:966works, and theshipped
clearAfterTurnfeature is not silently dead./clearrouted throughInjectorwedges that pane for good. After the clear, everylater message to that member vanished. fleetd read
state: busywhile herdr readliveStatus: done, and two tickets sat at[pending — worker done]forever./clearproduces no turn boundary, so the Injector's turn never completes and later sends queue behind
it. The comment at
member/ClaudeCodeLauncher.java:971— "This deliberately bypassesInjector: /clear is housekeeping, not a delegated turn" — is load-bearing, not stylistic.
The flow
The two
agents.sendcalls follow the two shipped patterns:clearContext(
member/ClaudeCodeLauncher.java:972) and the lead nudge (msg/LeadHeartbeatLoop.java:252).Units
Unit A — config block plus the rollover executor
leadRollover:, followingleadHeartbeat:exactly. With noblock, nothing is ever constructed, so an upgrade cannot silently acquire this behaviour.
Keys:
handoverPath,requireOperatorConfirm(defaulttrue),maxDocAgeSeconds,bootstrapText,clearSettleSeconds.LeadRolloverclass with a pure decision function plus a thin executor, the wayLeadHeartbeatLoopis built. It must:after the request — checked on a realtime clock, never
System.nanoTime;/clearwithagents.send(...)directly, never throughInjector;KNOWN_TOP_LEVEL_KEYS, classify it inConfigRef, document it infleetd.example.yaml, and update the running tally inConfigRef's class doc.Unit B — the
handoverskill.claude/skills/handover/SKILL.md: the lead-side playbook for writing a handover a stranger canact on. The doc's quality decides whether the next lead succeeds, so this unit is not
decoration. It must require: every number carries the command that produced it; open decisions
and who owns them; live hazards; what is explicitly not owed. Model it on the existing
handover file, which is a good example of the shape.
Unit C — the
fleet_handoverMCP tool (after A merges)Primary-only in
auth/Authz. Two phases on one tool: an opening call that returns the path andthe token, and a confirming call that carries
ready,operatorConfirmedand the token. A thirdform cancels. Because it adds a
fleet_*tool,CLAUDE.md's own rule applies: the primary'sintent-to-tool table must gain a row, and the canonical block must stay byte-identical with the
wiki template.
Acceptance criteria
leadRollover:block, no rollover object is constructed and no behaviour changes.can trigger the wipe. There must be a test that pins this.
three checks failed.
/clearand the bootstrap prompt both go throughagents.senddirectly. A test must fail ifeither is routed through
Injector.System.nanoTimeis used for thefile-freshness check.
freshly cleared pane picks up the next
agents.sendprompt. This is the one step stillunproven; my probe could not test it because the Injector was already wedged.
Hazards
the file checks; the operator flag is an assertion by the lead. Say it that way in the docs
rather than implying fleetd verified it.
What the test suite pins about
/clear, and what it does notChecked while Unit A was being built, against
7b97aae.member/ClaudeCodeLauncherTest.java:786(clearContextUsesTheClaudeCommandThroughTheOwningHandle)asserts that
agent.promptwas called withtext: "/clear"against aFakeHerdr. So it pinsthe call. It cannot pin the effect, because no real Claude Code session is involved.
Control: 128 test files in the tree, and the grep finds every
clearContext/clearAfterTurnsite, so the search discriminates.
session/SessionManagerTest.java:831(
clearAfterTurnResetsContextWithoutDoubleCountingTheTurn) is the same shape — it countslauncher.clearContextCalls(), again a fake.So until 2026-09-11 nothing in this repo had evidence that
/clearsent overagent.promptdoes anything at all. The shipped
clearAfterTurnfeature rested on an untested assumption,and it defaults to
false, so nobody would have noticed had it been wrong. The live proberecorded in the ticket body is the first evidence of the effect.
This does not close acceptance criterion 6. Two different things were unproven, and only one
is now settled:
/clearoveragent.promptexecutes as a slash commandagents.sendpromptThe second one is what criterion 6 asks for, and it is the step the whole feature ends on.
One reassuring detail found while checking, which is an argument but not a proof. In
clearAfterTurnthe clear happens insidecompleteTurn(
session/SessionManager.java:969) — that is, after the turn boundary was already seen. So theInjectoris not wedged there, which is the opposite of what my probe did, where the/clearwas the turn. The rollover feature is on the safe side of that line too: the lead is never
tracked by the
Injectorat all. But this reasoning is not a substitute for the probe, andcriterion 6 stays open.
Two corrections to the ticket body, one of them a defect in my own brief
1. The roll order in the ticket body is wrong — this is my mistake, not a worker's
The body says the roll is: send
/clear→ wait until injectable → send the bootstrap text.That first step can land in the middle of the lead's own turn.
confirm()is called by thelead, from inside a turn. At the moment it returns, the lead's pane is
WORKING, not idle. Everyother injection path in this daemon refuses to touch a
WORKINGpane on purpose —msg/LeadHeartbeatLoop.javacalls it constraint 2, andReplyPushLoop.decidechecksstatus.injectable()before it sends. My ordering skips that gate for the one send that destroyscontext.
Correct order — three waits, not two:
/clearwithagents.send(...), directly;agents.send(...), directly.So
confirm()must record the decision and return, and let a later tick do the roll. It mustnot roll inline. That also keeps
confirm()fast, instead of blocking the lead's MCP call for aslong as the settle window.
clearSettleSecondstherefore bounds step 3. Step 1 needs its own bound — call itturnSettleSeconds, same default of 20. If step 1 times out, refuse and roll nothing: a lead thatnever goes idle is a lead that is still working.
2. A freshly cleared session does reload
CLAUDE.mdand the memory indexMeasured on this host, 2026-09-11, from the lead's own transcript. The current lead session was
itself created by a
/clear(its line 5 is<command-name>/clear</command-name>), and line 27of that same session carries both the
CLAUDE.mdcanonical block and theMEMORY.mdindex.This closes a risk that was open: will the new lead know it is a lead? Yes.
CLAUDE.mdisreloaded, and
CLAUDE.mdalready tells every session to settle its role withfleet_whoamibefore acting.
Consequence for the design:
bootstrapTextcan be short. It does not need to re-explain therole, the bridge, or the charter — all of that arrives on its own. It needs to do one thing: name
the handover file and say to read it first. A long bootstrap prompt would be duplicated
instruction surface, and this repo already has a rule against that.
Two more design points, found by re-reading the flow
1. The roll must target the CALLING lead, not "the primary" — this is invariant 3
msg/LeadHeartbeatLoop.java:245resolves its target withprimaryRegistry.primaryTerminal().That is right for a background loop, which has no caller to speak of.
It is wrong here.
confirm()has a caller, andCLAUDE.md's invariant 3 is explicit: identitycomes from the connection, never an argument. So
LeadRollovermust roll the terminal thecall arrived on, resolved the way
auth/CallerResolver.java:208already resolves it — not aregistry lookup that answers "who is the primary".
Why it matters in practice: this daemon can hold more than one labelled lead tab. fleetd #359
exists because stale labels used to pile up, and
lead/LeadLauncher.java:132still carries thetwo-reading cleanup for exactly that. A lookup-based target means one lead can call
confirm()and a different lead's pane gets cleared. That is an unrecoverable loss of someone else's
context, caused by a lookup that looked harmless.
There is no matching risk in the other direction: a caller-derived target can only ever clear the
pane that asked to be cleared.
2.
open()should hand back the fleet state the handover needs to mentionThe handover file has to record in-flight work, or the new lead will not know it exists. The
handoverskill (merged in #481) requires a "live hazards" section for this, but the outgoinglead has to gather the facts by hand.
fleetd already has them.
msg/LeadHeartbeatLoop.java'sFleetStaterecord carries exactly theright four: undrained worker replies, sessions in
DONEawaiting teardown, live workers, andwhich terminals hold replies.
open()should return that snapshot alongside the path and thetoken.
This is cheap — the record exists and is already computed for the heartbeat — and it removes the
most likely gap in a handover file: a member still mid-turn that nobody wrote down.
Not a refusal.
open()should report the state, not block on it. Whether to roll withworkers still live is the lead's judgement, and sometimes the right answer is yes, because the
replies stay in the durable inbox for the next lead to drain. Refusing would make the feature
unusable exactly when a long session most needs it.
Lead status, 2026-09-11 — Unit A merged, two units in flight, one criterion I cannot prove
Unit A merged (PR #483, merge commit
4bab23e)leadRollover:config block +dev.ltms.fleet.lead.LeadRollover. Nothing calls it yet.CI run 1722 green on head
a942622. Full suite on mergedmain:Tests run: 1651, Failures: 0, Errors: 0, Skipped: 0,BUILD SUCCESS(run withoutclean, to protect the running daemon's jar— see #413).
Two corrections were applied before the merge, and the PR body still describes the pre-correction
shape:
confirm()cannot roll inline. It is called BY the lead FROM the lead's own turn, so thepane is
WORKINGand cannot report settled untilconfirm()returns. The original code sent/clearand then timed out waiting — destroying the context and starting no fresh session,while returning a refusal that claimed nothing had happened. Now
confirm()validates, hands aone-shot continuation to
continuationRunner, and returns. The continuation waits for thecalling turn to settle first (new knob
turnSettleSeconds); if that wait times out it sendsno
/clearat all.PrimaryRegistry#primaryTerminal()is asingle slot, and this daemon can hold more than one labelled lead tab, so lead X's
confirm()could clear lead Y's pane.
open()records the caller's terminal;confirm()refuses withNOT_YOUR_ROLLOVERon a mismatch.Mutation battery on the two safety-critical branches, in a scratch worktree at
a942622. Controlsmeasured against the original first; each mutant proven applied with two unrelated proofs using
different search strings.
if (!turnSettled)→if (false):turnThatNeverSettlesSendsNoClearAtAllFAILS(
LeadRolloverTest:247, expected 0 sends, was 1).if (false):aDifferentLeadTerminalCannotConfirmAnotherLeadsRolloverFAILS (
LeadRolloverTest:266).WARNlines logged, so the branches reallyare exercised.
New defect found after the merge — Unit E in flight
waitUntilInjectableusesAgentStatus#injectable(), which isIDLE || BLOCKED || DONE. That isthe right rule for
Injector("may I deliver a message"), but the wrong rule for this gate ("hasthe turn actually ended"), because
BLOCKEDis a live turn that is paused — a pane sitting onan approval prompt.
Path in: the lead calls
confirm(); the turn carries on and hits anything needing approval; herdrreports
blocked; within 250ms the continuation reads that as settled;/clearis typed into anopen prompt. The lead's live context is destroyed mid-turn — exactly what correction 1 exists to
prevent. Window is
turnSettleSeconds, default 20s.The mutation battery could not have found this: it proved the guard fires, not that the guard's
predicate was tight enough.
LeadRolloverTestdrives the fake status with only"idle"(
:198) and"working"(:229); neither value discriminates this defect, only"blocked"does.Unit E makes both waits require a real turn boundary (
IDLEorDONE), local toLeadRollover.AgentStatus#injectable()itself is left alone — it is correct forInjectorandLeadHeartbeatLoop.Unit C in flight — the
fleet_handoverMCP toolopen/confirm/cancel, primary-only inAuthz, caller terminal resolved from the connectionwith no terminal parameter in the input schema. Registered unconditionally and answering a
clean
NOT_CONFIGUREDrefusal whenleadRollover:is absent, so the registered tool surface doesnot vary with config — otherwise it collides with #474's charter tool-surface gate and
McpContractDocTest.Acceptance criterion 6 — I cannot prove this, and I am not going to pretend otherwise
"A freshly cleared pane picks up the next
agents.sendprompt" is still unproven. The liveprobe recorded earlier in this ticket proved
/clearoveragent.promptexecutes as a real slashcommand, but it could not get past that, because routing
/clearthroughInjectorwedged thetarget permanently.
There is no route left that would prove it:
Injectoragents.send. The registered routes are/agents,/healthz,/mcp,/member-credentials,/members,/members/{paneId},/metrics,/profiles,/sessions,/sessions/{id}/{ask,message,replies,reply,status},/tasks/{ticket}.clearAfterTurnis the one shipped feature that does this, but turning it on needs afleetd.yamledit, which the classifier denied in an earlier session. Not routing around it.four named contract tests.
LeadRolloveroperates on the lead's own pane, resolved from the connection, and the tool isprimary-only — so it cannot be exercised on a throwaway member at all. The only honest proof is to
use the feature for its real purpose, once, on a live lead that has already written its handover
file. That needs the operator's say-so, which the design already requires
(
requireOperatorConfirm, default true).Until that happens, the residual risk is bounded and worth stating: if the bootstrap prompt
does not land, the lead's context is gone and no fresh session starts. The handover file is
written before
confirm()is even accepted, so the loss is recoverable by hand — the operatorstarts a session and points it at the file. That is the same manual path the
handoverskilldescribes today.
Also verified
The lead's own pane status is readable through the daemon —
GET /sessions/term_…/status→ HTTP 200,{"status":"working","ready":true}, read while thislead was mid-turn. So the premise
waitUntilInjectablerests on holds, and theworkingreadingduring a live turn is direct evidence for correction 1.
Not done
leadRollover:— waiting until the tool lands so it describes acapability an operator can actually use.
CLAUDE.mdintent-to-tool row forfleet_handover, kept byte-identical withwiki/7-Use-Cases.md.handoverskill still saysfleet_handoverdoes not exist. True today; must change whenUnit C merges.
Shipped and deployed — 2026-09-11
All three units merged, docs updated, daemon redeployed, tool proven reachable on the live daemon.
leadRollover:config +LeadRolloverexecutorBLOCKEDis not a settled panefleet_handoverMCP toolmainat008a457. Full suite on merged main:Tests run: 1662, Failures: 0, Errors: 0, Skipped: 0.The gap in #485 that the first pass shipped
Worth recording because a green suite vouched for it. The unit first landed with a defaulted
15-argument
FleetMcpconstructor delegating to the new 16-argument one withleadRollover = null.I mutated the wiring instead of reasoning about it — deleted just the
leadRolloverargument fromFleetd.main'sFleetMcpcall:…while the live daemon would have answered
NOT_CONFIGUREDto everyfleet_handovercall for ever.FleetMcpHandoverTestcould not see it (it builds its ownFleetMcp), andFleetdLeadRolloverWiringTestcould not either (it pins thatLeadRolloveris constructed, notthat it is passed on).
The defect was wider than I named. I pointed at one overload; the worker found 11/12/13/14/15-arg
constructors forming a single defaulting chain into the 16-arg one, each silently supplying another
feature's "off" value —
leadChannel,outage,leadSeats,peers, thenleadRollover. All fivedeleted.
FleetMcpnow has exactly one public constructor, so all of those features arecompile-enforced at their call sites, not just this one. Re-running the identical mutation now gives
constructor FleetMcp cannot be applied to given types … argument lists differ in length.Deployed
Live probe of the tool itself, on the running daemon:
Clean structured refusal, no exception — and the live wire schema carries exactly four parameters
(
action,reason,token,operatorConfirmed) with no terminal, session or leadTerminal ofany kind, confirmed against the registered schema rather than the source.
Docs
CLAUDE.md+wiki/7-Use-Cases.md: onefleet_handoverintent-to-tool row, byte-identical(sync check prints
True). The row carries the ordering trap, not just the call.wiki/11-Features.md: full entry — knob, defaults, the three design decisions and why, sixgotchas. Pushed and verified by ref (
1eadfbc), sinceHEAD:masteris a silent no-op here..claude/skills/handover/SKILL.md: it said "do not call afleet_handovertool: it does notexist", which was true this morning and false by this afternoon. Rewritten, plus a new section 11
with the three-step order and what surprises a caller.
What the operator still has to do
The feature is deployed but inert until
leadRollover:is added tofleetd.yaml, which I amnot permitted to edit. Minimum block:
Everything else defaults:
requireOperatorConfirm: true,maxDocAgeSeconds: 3600,turnSettleSeconds: 20,clearSettleSeconds: 20, and abootstrapTextnaminghandoverPath.Adding the block needs a daemon restart — the fields are hot, but the executor's construction is
gated on the block being present in the startup snapshot.
Still open
previous comment for every route I checked. The only honest proof is running the feature once, for
real, on a live lead that has already written its handover file. That is the operator's call.
surfaces as a CI hang rather than a named red test. Not a production bug.
leadRollover:is now configured and LIVE on the MacThe operator authorised the
fleetd.yamledit, with one requirement: use a relative path to a file in the workspace, not an absolute one. That turned out to need code, because a relative path could not have worked at all.The defect a relative path exposed
Measured today:
/Users/dai.ha/LTMS/claude-bridge/fleetd(lsof -a -p 1907 -d cwd)/Users/dai.ha/LTMS/claude-bridge(fleet.leaders.opus.cwd)They differ by one level, because this repo nests
fleetd/inside its own root.LeadRollover.java:218stored the raw configured string and:321calledPath.of(...)on it. Three readers consume that one string, in two different directories: the daemon stats the file, the lead is handed it in theopenresponse and writes there, and the fresh session is handed it insidebootstrapText. A relative path would have failed withHANDOVER_NOT_FOUNDevery time.Fixed in #487 (merged as
3bf3968, main at7f9137f)open()now resolves the path to an absolute path exactly once, against the calling lead's configuredcwd, falling back toSystem.getProperty("user.dir")— the same fallbackLeadLauncher.java:311already uses.PendingRollover.handoverPath()is absolute from then on, so all three readers agree.bootstrapTextis no longer defaulted in the config record (it would bake in the raw relative path);bootstrapTextFor(resolvedPath)builds it at send time.Two defects found during verification, both already fixed
1. The production wiring was uncovered. I mutated the lookup lambda inside
Fleetd.leadRollover(...)to always returnnull— which forces every relative path back onto the daemon's own directory, i.e. this exact bug:Proven applied with two greps using different search strings (original gone: 0; mutant present: 1). Result:
Tests run: 1669, Failures: 0, Errors: 0, Skipped: 0/BUILD SUCCESS. The whole suite vouched for dead wiring.This is worth naming precisely, because it is a new wrinkle on the #415 antidote. Making
leadWorkspacea required constructor parameter did work — it made every call site a compile error, andFleetdLeadRolloverWiringTest's source-text pin catchesleadsbeing dropped from the argument list. But required only proves an argument is passed. It says nothing about whether the argument is correct, and here the argument is a lambda built inside the factory, so its body was untested surface that the antidote does not reach.Fixed by
FleetdLeadRolloverWorkspaceLookupTest— a behavioural test that calls the package-privateFleetd.leadRollover(...)directly with aConfigReffrom a tempfleetd.yaml. I re-ran the identical mutation against the merged tree myself:Tests run: 1673, Failures: 2, both failures printing/Users/dai.ha/LTMS/claude-bridge/fleetd/handover.mdas the wrong answer. One of its three cases pins that the lookup is read live: the map is empty when the factory runs and only gains the terminal afterwards, which is the real startup order, since leads are found by a tab scan after boot.FleetdLeadRolloverWiringTest's javadoc claimed "no behavioural test can catch this wiring dropping out." True before the factory took a lookup, false after. Corrected — a wrong claim inside a test is how the next person decides not to write the missing test.2. A relative
fleet.leaders.<name>.cwdbroke the absolute-path guarantee. Found by a reviewer (terra, deliberately not the sonnet that wrote the diff).base.resolve(path).normalize()stays relative if the base is relative, silently reopening the same ambiguity — and falsifying the "always absolute" promise the newPendingRolloverjavadoc makes. Now.toAbsolutePath()on both branches. I mutated it back out:Tests run: 1673, Failures: 1, killed.Filed #488 for the underlying config gap: nothing validates
Leader.cwd()at load, so a relativecwdis silently accepted.Deployed and proven live
Build
Tests run: 1673, Failures: 0, Errors: 0, Skipped: 0. Redeploy: jar2520c3c051c7→0faa26d1bb10, pid 1907 → 62179,/healthz200, freshfleetd listeningline, no ERROR lines since restart,fleet_whoami→primary.Config now in
fleetd.yaml:Live probe on the running daemon —
fleet_handover{action:"open"}:and the daemon's own log line:
That is the relative value resolving against the lead's workspace, not the daemon's — through the real wiring, on the real daemon. Token cancelled afterwards; nothing was rolled.
.handover/is gitignored (the file is a snapshot of live state). Docs updated:wiki/11-Features.md(pushed, verified by ref285b0a6) and thehandoverskill, which now says the returned path is always absolute and must not be re-resolved by the lead. TheCLAUDE.mdcharter row needed no change; sync check still printsin sync: True.Still open
Acceptance criterion 6 is still not met — nothing has yet shown end-to-end that a freshly cleared pane picks up the bootstrap prompt. It is now reachable for the first time, since the feature is configured, but the only honest proof is running a real rollover on a live lead, and that is the operator's call. If the bootstrap does not land, the lead's context is gone and no fresh session starts; the recovery is the manual path, which is exactly why the handover file must exist before
confirmis accepted.#486 (unbounded settle poll hangs CI instead of failing) remains open and unassigned.
Acceptance criterion 6 is now answered, negatively. Measured on the Mac, 2026-09-12, first real rollover of the shipped feature with a Claude Code lead.
The roll was confirmed at
07:03:30.903and loggedrolledat07:03:43.998. The bootstrap prompt did not land. The pane received one submitted line,/clearFresh lead session. ..., and Claude Code answeredUnknown command: /clearFresh.The failure is safe. No
/clearever executed, so no context was lost, no session died, and the handover file stayed intact at.handover/HANDOVER.md. The roll simply did nothing.Cause, evidence and fix are in #489. In short: the second settle wait is a no-op (
/clearstarts no turn, so the pane never leavesIDLEand the wait returns on its first poll - 438 ms of a 20-second budget), and under that the submit keystroke raced the paste, whichAgentControl.submit's own javadoc already records as CB-113.wiki/11-Features.mdand thehandoverskill have been corrected: both said "not yet proven end-to-end", which is now too kind. They say "measured failing" and point at #489.One observation, on one host, with one backend. It is a data point, not a law - but it is a reproducible defect in the code path, not a flake.
The fleet01 test: reported here on that lead's behalf
The operator asked me to have fleet01 test the roll and to fix whatever came out. I asked the fleet01 lead over the coordination channel. They cannot run it, and that is the right answer. They have no forge credential, so they cannot post to this issue themselves. Everything in §1–§4 below is their measurement on their host. I have not re-run any of it. §5 is mine.
1. fleet01 cannot run the roll at all — the code is not there
Their report:
So fleet01 is not a second instrument for the roll today. It is a host that would need a 91-commit update and new config written before a roll could even start. Note the last line: even after an update,
bootstrapTexthas nothing to attach to there, so thefleet_whoami-first bootstrap would not exist.2. They are not refusing — it is their operator's call
They have 12 unpushed commits on
worker/git-native-removals-4633ad-2at38fd7f3in the kb repo. Their operator reserved the rebuild and wants a human present for recovery. A roll that fails leaves a wiped lead and no fresh session, so doing it with unpushed work and nobody watching is the wrong trade. I agree with them. The fleet01 test stays blocked on their operator, not on us.3. They corrected one of my claims — I was wrong
I wrote in #491 that fleet01's layout would expose #488 (
fleet.leaders.<name>.cwdis never validated). It does not. Theircwdatfleetd.yaml:229is already absolute (/home/ltms/LTMS/kb), so the bad input is simply not present. #488 still needs a host with a relativecwd, or a test that supplies one. Corrected on #491.The part that survives is stronger than I filed it. For #487, their two directories (
/home/ltms/LTMS/kband/home/ltms/LTMS/fleetd/fleetd) share only/home/ltms/LTMS— not nested at any depth. On this Mac they are one level apart, the weakest separation that still counts.4. They corrected a second claim, and it produced a new ticket
I had advised "build before you stop anything". They pointed out this contradicts #413, with evidence from their host on 2026-09-10: jar rewritten at 02:10:41, pid 1610855 died at 02:11:17 with
NoClassDefFoundError: dev/ltms/fleet/msg/Rendezvous$Resolution, and systemd recordedstatus=143— which reads as clean.5. What I measured myself, and what I filed
I checked the same mechanism on this Mac before accepting it:
Different OS, different supervisor, different class, different finder — so these are two data points, not one. Filed as #493:
scripts/redeploy-fleetd.shbuilds into the livetarget/fleetd.jarbefore stopping the daemon, so the shutdown drain can die on a failure-path class the running JVM never loaded. The exit code does not show it and nothing currently looks.6. Where that leaves this issue
Criterion 6 — "a fresh session boots from the handover file" — is still not met on any host. PR #490 fixed the
/clearFreshconcatenation and is merged and deployed here, but it has not yet rolled a real lead session end to end. fleet01 cannot be that proof today.fleet01's operator has answered: deferred, not scheduled. Do not hold this issue for that host.
Reported for the fleet01 lead, who still has no forge credential. Their words, condensed: nothing changes on their host now, and the decision gets made when their operator is present rather than at a booked time. Booking a window would either park their operator waiting for it or make the lead roll unattended, and an unattended roll turns our test into their outage. Their host state is unchanged:
4887731, jarb92db8f4built 2026-09-10 02:10:41, daemon pid 1696852, 91 commits behind, fetch window ends at #470.So criterion 6 stays open at one host, one backend. If another lead comes online, they are a better second instrument than fleet01 precisely because they are available. This issue should not wait for fleet01.
Two facts from that host survive the deferral and are worth keeping:
cwdand daemoncwdshare nothing below/home/ltms/LTMS),cwdalready absolute). #488 needs a different host, or a test that supplies the bad input.Their read-only review of #489/#490 produced two results
1. The question: does a test exist that would have caught the ORIGINAL defect, or does the new test pin the fix rather than the bug?
Answered by mutation, not by reading, since I wrote both the fix and the test. Baseline
LeadRolloverTest= 29 tests, 0 failures.waitUntilAtTurnBoundary)clearPickupIsNudgedBeforeBootstrapTextWhenPaneStaysIdle:498agents.submit(target); return true;— satisfies "a nudge happened, ordered between/clearandbootstrapText" and nothing elseThe second mutation is the important one. Three of those failures assert the negative — that
bootstrapTextis not sent when the wait should not release. That is "the wait actually waits", expressed as a consequence rather than as a duration. So the suite pins the bug, not only the fix.Proof each mutant applied used two different search strings (mutant present, original gone) plus a control. After each restore: 29 / 0, source
shasumbyte-identical to the pristine copy.2. Their second ask found a real defect my mutations could not. They wanted "a bound and a named failure message, not a happy path". Checking what an operator actually sees,
LeadRollover:400printscfg.clearSettleSeconds()— the configured budget, not the measured wait. Had that path fired during the live incident it would have loggedwithin 20swhile the truth was 0.438s: confidently wrong, not silent. And:569releases the roll atinfo, followed byrolled, so the case most likely to produce/clearFresh-shaped garbage reports success to anyone grepping forWARN.Filed as #494 and delegated. #490 made the wait real; it did not make the wait observable.
Correction to my mutation numbers above
I reported mutation 2 as "4 failures + 1 error". The error carries nothing and should not have been in the total. The fleet01 lead caught it from the shape of the result before I checked the stack.
submitThatThrowsDoesNotAbortTheRoll:629errored, and the exception escaped from my stub, not from an assertion catching my stub:My stub was
agents.submit(target); return true;with notry/catch; the real method wraps that call. So the mutant broke the test rather than the test catching the mutant.Corrected: mutation 2 = 4 failures. The conclusion is unchanged and rests on those four, three of which assert the negative —
bootstrapTextis not sent when the wait should not release.General rule worth keeping: when a mutation produces errors as well as failures, count only the failures. An error can be collateral damage from the mutant, and it looks like signal.