A handover rolls the lead with /clear, so the process survives and a Claude Code CLI update is never picked up #726
Open
opened 2026-10-04 17:28:10 +02:00 by ltms
·
9 comments
No Branch/Tag Specified
main
worker/726-unit2-75cb13-4
worker/737-owner-key-ff061f-10
worker/736-presence-forget-f35144-9
worker/705-observer-14c258-6
worker/722-024c34-5
worker/726-ea34a0-2
worker/726-10cbf0-1
worker/729-5961c6-3
worker/727-ee14ed-3
worker/719-bdd95e-4
worker/702-4f5c7f-2
worker/715-5c43fc-1
worker/721-70f9ea-5
worker/718-99362b-2
worker/task-15-af0d10-12
worker/task-16-50a702-13
worker/task-12-4d0479-9
worker/task-13-823ce2-10
worker/705-ticket-owner-af9928-8
worker/703-list-collaborators-9c06c2-7
worker/669-example-truth-0b303d-6
worker/669-collab-deliverability-9ba859-3
worker/669-collab-reload-report-2a21bd-4
worker/669-7e80a6-1
worker/669-unit-d-efbbd7-1
worker/669-1b786a-1
worker/669-1d1d9f-1
worker/692-4afb9d-2
worker/689-02fced-13
worker/693-cf23fa-14
worker/677-fix-lead-collision-f69073-12
worker/638-fix-overmask-dbb1bf-11
worker/675-5b7478-4
worker/669-unit-a-70cc8f-3
worker/677-8cdaaf-5
worker/638-a7b391-1
worker/683-4536d6-2
worker/651-a75bbe-8
worker/680-20607d-7
worker/664-c12e95-3
worker/668-08534d-4
worker/672-0f2469-2
worker/670-7d1022-1
worker/661-ac7c28-2
worker/664-37fb9b-3
worker/663-remove-3arg-read-3f6783-1
worker/659-remove-dead-backcompat-ba5e6f-1
worker/637-revision-60a488-23
worker/656-redact-regression-tests-892903-19
worker/637-context-gauge-threshold-466eb5-16
worker/639-redact-line-numbers-de4ac4-17
worker/641-set-reformat-guard-6f96a4-18
worker/642-herdr-guard-scope-5de0e4-15
worker/650-javadoc-scope-95f3b3-14
worker/612-01e9f7-13
worker/612-a-r4-quarantine-outage-7ab0e8-5
worker/612-a-r9-r11-capacity-coverage-peers-cfcc79-7
worker/612-a-r10-loophealth-ccc872-8
worker/612-a-r12-turnregistrar-9e3bb7-9
worker/612-a-r5-leadconfigdir-9e70cf-6
lead/config-edit-redact-anchor-wording
worker/config-edit-seam-ca8dc1-1
worker/612-r67-630-lifecycle-290b8d-3
worker/629-625-ports-seams-da7d5d-4
worker/612-r12-exhaustion-f37cd7-1
worker/612-r38-amqp-24b083-2
worker/fleetd-612-unita-87807e-1
worker/612-b3-mcpwirings-da2b58-3
worker/612-b2-cb185-176d3a-2
worker/612-b1-completion-457459-1
worker/612-agaps-73a926-2
worker/608-sleeps-3a64ff-3
worker/621-b4520b-1
worker/618-b83894-2
worker/fleetd-615-e05481-5
worker/lead-autocompact-5f1ab2-3
worker/fleetd-613-f85deb-3
worker/fleetd-608-flaky-nudge-test-d0c2d1-3
worker/lead-context-gauge-ad404f-1
worker/gauge-wiring-9158c1-4
worker/redeploy-slowstart-ead0e5-5
worker/charter-bytes-13668c-6
worker/rollover-outcome-291483-2
worker/589-f64303-2
worker/593-1a8025-5
worker/589-fcd2aa-1
worker/568-9fdaa2-3
worker/571-attempted-outcome-5739f7-2
worker/581-completionresolver-cas-sites-0542b7-6
worker/562-loop-health-wiring-test-99611c-5
worker/562-surface-loop-health-7df5cc-4
worker/575-waiter-cleanup-sites-62ad80-1
worker/572-answer-lock-release-46a9ae-5
worker/567-probe-channel-leak-a38fc5-6
worker/551-record-before-send-7cbf56-1
worker/561-listener-fanout-survives-a-throw-61d538-2
worker/555-redeploy-main-flow-seam-65c2f5-2
worker/556-injector-owns-registration-e027a5-1
worker/552-post-restart-mktemp-abort-bc2672-4
worker/553-onstatus-completion-leak-0da881-2
worker/550-shasum-linux-196132-1
worker/538-loop-dies-on-error-4a5eeb-6
worker/426-health-coverage-ef1fd4-4
worker/504-failed-reported-clean-3cfd66-3
worker/537-capturedlog-close-e4c437-2
worker/459-broken-link-targets-cadc17-5
worker/535-appender-leak-fe74c1-1
worker/512-part2-shutdown-detection-434701-9
worker/529-logger-level-sweep-2a5533-8
worker/528-drain-gate-call-site-5de83d-7
charter/forge-mcp-vs-token
worker/521-swap-guard-unpinned-28e931-5
worker/519-probe-test-harness-d25ab8-4
worker/525-logger-level-leak-1b4eb0-6
worker/518-fleetmcp-resolver-wiring-8ef96c-1
worker/512-drain-complete-line-7edd71-3
worker/517-abort-branch-and-jar-id-41b641-2
worker/500-9e52c9-3
worker/509-4912f4-2
worker/511-9a4b23-1
worker/493-479f45-2
worker/505-03f8b2-1
worker/492-followup-detect-unclear
worker/501-a31fa0-7
worker/498-451d1c-5
worker/494-1015ce-2
worker/492-209647-1
worker/489-001902-2
worker/480-relative-handover-path-906323-1
worker/480-b-handover-skill-45bf1f-5
worker/474-followup-source-pin-f54a55-17
worker/474-charter-check-on-reload-f54a55-17
worker/466-quarantine-repeatcount-report
worker/393-opencode-skill-seeding-71854b-13
worker/469-canonical-tool-names-2a472a-16
worker/466-quarantine-escalation-5ae9c1-15
worker/446-hot-exhausted-pattern-0af580-6
worker/464-charter-tool-name-guard-a85635-12
worker/463-listfleet-default-fails-open-f1c76c-11
worker/458-invariant-5-by-purpose-862f9a-10
worker/439-coordinator-row-gate-bc032a-8
worker/449-herdr-protocol-576015-4
worker/450-abstract-spawn-599e1c-5
worker/437-ack-refuses-177d91-1
worker/444-placement-window-feb56a-2
worker/440-helddurable-derived-d462d7-13
worker/425-rework-placement-resolve-c58ba1-9
worker/421-lead-peek-held-msgs-cdbad2-10
worker/435-fixed-policy-cap-fe11de-12
worker/422-gate-state-observability-9e79d6-11
worker/431-memberregistry-live-readers-cdbad2-10
worker/424-architect-slot-hot-038b41-7
worker/422-model-gate-spawn-c29f48-6
worker/425-default-profile-live-f55534-8
worker/415-coverage-wording-2cbf9c-5
worker/416-3ad1da-1
worker/418-588283-3
worker/deterministic-stamp-race-409-3cb7b6-10
worker/armed-reads-live-config-404-ed931f-9
worker/reply-peer-refusal-391-5a34bd-7
worker/models-allowlist-aa9e9b-3
worker/ttl-stamp-race-399-f1122f-8
worker/scrub-receipt-400-316b3e-5
worker/exhaustion-detection-395-105105-6
worker/scrub-abort-394-316b3e-5
fix/scrub-uid-abort
worker/task-scrub-517574-2
worker/t386-clock-bd5b78-4
worker/t384-scrub-813790-5
worker/t381-cc-748314-2
worker/t373-336973-2
worker/t365-3920c5-3
worker/t358-6e989b-1
worker/t355-8b321c-1
worker/fleetd-369-hermetic-git-tests-e8b19a-3
worker/fleetd-368-stale-lead-binding-f5682e-2
worker/fleetd-360-deploy-units-0d3793-1
worker/359-dead-lead-tabs-f1253b-4
worker/362-worktree-skills-c03e51-3
worker/361-coord-visibility-655144-1
362-plugin-visibility-and-drift
worker/errscan-bed2ca-2
worker/amqp-log-identity-bed2ca-2
worker/withdefaults-guard-561704
worker/sleepguard-82076d-1
worker/fd334-9ee1b6-5
worker/fd348-f1ab27-4
worker/fd335-a71c35-1
worker/fd342-174a17-2
worker/fd345-490d0f-3
worker/fleetd-337-5ec7d4-21
worker/fleetd-341-af5a6b-24
worker/fleetd-339-5ca0a2-23
worker/fleetd-338-83a4a1-22
worker/fleetd-333-281f46-18
worker/fleetd-329-11bdbb-16
worker/fleetd-330-2770fb-17
worker/fix-326-50506e-15
worker/fix-324-3e9bbf-14
worker/fix-323-b8287d-13
worker/fix-316b-bd0860-11
worker/fix-318-76ca36-9
worker/fix-317-486aec-8
worker/fix-315-ce47c5-6
worker/fix-307-275890-6
worker/fix-308-b4f664-7
worker/fix-309-ec3939-8
worker/fix-310-7a3974-9
worker/fix-302-52ad0e-9
worker/fix-298-ce1acb-8
worker/fix-297-66bd11-7
worker/fix-296-104622-6
worker/fix-293-bare-closetab-eb22b5-3
worker/fix-280-gone-ask-lapse-bca98e-2
worker/fix-290-reapidle-guard-coverage-9b0dd1-1
worker/fix-285-trust-seed-8f3565-10
worker/fix-284-backend-error-seat-85912c-11
worker/fix-282-chained-ask-e6d0bb-8
worker/fix-283-teardown-leaks-f40dfa-9
worker/fix-281-pin-handler-actions-4921ac-7
worker/audit-rendezvous-lifecycle-d072ae-2
worker/audit-health-placement-1a2476-6
worker/audit-teardown-exits-e207a5-3
worker/audit-launcher-asymmetry-27e370-4
worker/audit-rest-authz-6ca53c-5
worker/investigate-275-abandon-asking-fdef52-8
worker/fix-274-worktree-leak-b0095d-7
worker/fix-273-exhausted-pattern-9665b5-6
worker/fleetd-267-model-check-bd8068-1
worker/fleetd-131-archunit-18b834-7
worker/fleetd-266-sshagent-rename-a014ff-6
worker/fleetd-184-uid-claim-8e1f31-4
worker/fleetd-184-warn-b381ee-10
worker/fleetd-184-docs-be1d12-9
worker/fleetd-257-9bf010-7
worker/fleetd-103-23a113-6
worker/fleetd-247-342356-5
worker/fleetd-116-04dea8-4
worker/fleetd-252-a830e0-3
worker/fleetd-111-7e8673-9
worker/fleetd-155c-f8ef4b-8
worker/fleetd-176-b928ca-3
worker/fleetd-249-7a7878-2
worker/cb248-composition-root-b-9acdf7-15
worker/cb148-envrc-default-fa6c82-12
worker/cb201-unit5-wiring-6c12e6-8
worker/cb241-fallback-echo-1175e9-11
worker/cb149-trust-dialog-2392a5-9
worker/cb134-148-overlay-visible-c9b986-10
worker/cb234-session-id-keyed-04e1fc-1
worker/cb201-unit3-nudge-abdf5c-6
worker/cb201-unit2-policy-c1102c-5
worker/cb201-unit4-outcome-a13bfa-7
worker/cb201-unit1-classifier-91b9b1-4
worker/cb201-227-refine-831980-3
worker/cb175-model-readback-0f085f-1
worker/cb222-charter-tmpdir-17f013-1
worker/cb226-architect-slot-race-cd3aa8-3
worker/cb224-worktree-root-group-024523-2
worker/cb-123-role-demotion-c600f7-2
worker/cb-219-opencode-roots-1f677e-1
worker/cb214-claude-session-id-b9eab4-4
worker/cb213-zdotdir-wrong-process-dd6de4-3
worker/cb211-exhaustion-classification-9546e0-2
worker/cb137-ambiguous-task-4df3d8-4
worker/cb209-agentsessionid-4dfdb6-2
worker/cb185-hostenvnames-2692b5-3
worker/cb206-opencode-sqlite-128718-2
worker/cb185-worktree-group-fc0c99-1
worker/cb-137-ask-ticket-e7760c-2
worker/cb-172-broker-uri-d36ae4-4
worker/cb-175-model-readback-76ead6-3
worker/cb-161-pane-ancestry-293510-1
worker/cb-164-rebase-885863-8
worker/cb-164-empty-scrape-false-success-1a80af-3
fix/cb-197-ticket-ttl-from-completion
worker/cb-189-remote-url-coverage-4692f3-1
worker/cb-185-blockers-027756-4
worker/cb-192-gap-log-11b631-2
worker/cb-633-fix-5f4396-3
worker/cb185-router-d6436d-3
worker/cb185-router-routing-gaps-9e9d33-3
worker/cb185-paneids-992586-2
worker/cb-633-allow-list-union-ed374b-1
worker/cb-157-credential-in-remote-url-496e44-2
worker/cb-641-health-herdr-evidence-8f1f54-6
worker/cb-640-health-msg-evidence-99c9cd-1
worker/cb-642-fleets-status-skill-bbbc40-5
cb-634-ide-mcp
worker/lead-comms-wiring-c014b9-7
worker/lead-mailbox-c19577-6
worker/autocompact-window-82bc2f-5
worker/cb-634-probe-18056f-4
worker/cb635-broker-urienv
worker/cb-632-config-retry-8e0efa-7
lead/cb-622e-claude-md
lead/cb-622-followup
worker/cb-622a-165dff-1
lead/cb-622d-opencode-mount
worker/cb-622b-717c67-2
worker/cb-622c-ab7759-3
worker/cb-617b2-20ca4b-3
worker/cb-617a-5c2f4a-1
worker/cb596-4e49ef-3
worker/cb586-10500c-1
worker/cb-606-b9343a-25
worker/cb604-1445f8-24
worker/cb582-477374-21
worker/cb584-8c2281-22
worker/cb600-e6b9a9-20
worker/cb602-ce257f-19
worker/cb601-b42837-18
worker/cb598-6c7ba7-17
worker/cb599-740fe4-16
worker/cb597-282224-15
worker/cb590fix-185e9a-10
worker/cb528-recovery-race
worker/cb594-96bead-8
worker/cb590-916766-2
worker/cb527-997d99-3
worker/cb592-env-leak-3cbf9c-1
worker/cb588-async-ticket-nudge-3218f7-5
worker/cb578b-9dcb13-6
worker/cb581-d24826-5
worker/m2-u5-ef8c42-15
worker/cb578a-516499-2
worker/cb576-01a04b-17
worker/cb579-lead-tab-acba06-20
worker/cb580-terminal-health-ed6058-21
worker/cb577-f36fdc-18
worker/cb573b-3db06f-16
worker/cb568c-f36fdc-18
worker/cb568-drop-cause-c3ac1c
worker/cb575-cancelled-notification-c3ac1c
worker/m4-sol-a2cbec-3
worker/cb574-async-ask-c3ac1c
worker/cb573-health-model-8ca857-14
worker/cb572-unknown-target-7f2e35-13
worker/u4-700706-9
worker/u3-b9fcb6-6
worker/u2-ef5b68-4
worker/u1-469dce-1-clean
worker/u1-469dce-1
worker/cb-564-health-events-70cf7e-2
worker/cb-565-recycle-drops-role-98e58f-3
worker/cb-563-missing-reply-df2866-1
worker/cb-562-readiness-gate-silent-6c23c9-3
worker/cb-560-architect-presence-da8155-1
worker/cb-561-architect-silent-off-a71cab-2
worker/cb-548-bind-architect-slot-fe1b8c-1
worker/parity-overlay-settings-5fb711-1
secrets-central-store
cb-559-hot-key-correction
cb-557-fleet-role-pools
worker/cb-553-maxload-explicit-spawn-305ee3-6
worker/cb-551-idle-lead-heartbeat-f1633c-1
worker/cb-544-drain-preserves-worktree-925fad-3
worker/cb-552-docs-sync-1cb9cf-4
worker/cb-548-rendezvous-guard-rebased
worker/cb-548-rendezvous-guard-116b53-10
worker/cb-548-authz-v2-586df6-8
worker/cb-548-authz-264363-5
salvage/cb-528b-codex-home
salvage/cb-528a-codex-launcher
CB-518-primary-flow
feature/peer-launcher-spi
cb-103-injector
v1.1.0
v1.0.0
Labels
Clear labels
blocked
needs-live-proof
ready-to-delegate
silent-default
Cannot start until something else lands. The body says what.
Merged and green, but never shown working on the running daemon. Not the same as done.
Scope, files and acceptance criteria are written. A worker can be briefed from the body alone.
A feature that compiles, passes tests, and ships turned off. Nine recurrences and counting.
No Label
Milestone
No items
No Milestone
Projects
Clear projects
No project
Notifications
Due Date
No due date set.
Dependencies
No dependencies set.
Reference: fleet/fleetd#726
Reference in New Issue
Block a user
Blocking a user prevents them from interacting with repositories, such as opening or commenting on pull requests or issues. Learn more about blocking a user.
Delete Branch "%!s()"
Deleting a branch is permanent. Although the deleted branch may continue to exist for a short time before it actually gets removed, it CANNOT be undone in most cases. Continue?
The problem
fleet_handoverreplaces a lead's context by typing/clearinto its pane. TheclaudeOSprocess never dies. So a newer Claude Code CLI installed on disk is never loaded, no matter how
many times the lead is rolled.
The operator asked for this on 2026-10-04: "I want to see a true new restart of lead session, not
just /new - doing so ensure Claude Code update pick up".
Measured, 2026-10-04 17:24 CEST on the Mac
The running lead started three days before the CLI on disk was written. A roll would not change
that. I did not read the running process's own version number — I am inferring it is older from the
start time against the binary mtime, which is enough to show the gap but is not a version reading.
Where it is
fleetd/src/main/java/dev/ltms/fleet/lead/LeadRollover.java:571Then
waitForClearPickupAndSettle(cfg.clearSettleSeconds()), thenagents.send(lead, cfg.bootstrapTextFor(p.handoverPath())). Outcomes areTURN_NEVER_SETTLED,CLEAR_NEVER_SETTLED,ROLLED. Nothing in that path ends a process.Why this looks tractable
dev/ltms/fleet/lead/LeadLauncher.javaalready launches a lead from config, so the daemon does notneed to learn a new launch command:
ensureLeads()at:105agents.start("lead-" + name, herdrKind(profile), argv, tab.rootPaneId())at:326leadArgv(profile)at:371builds argv from the profile and adds--mcp-config,--modelandthe auto-compact pin. It deliberately omits
--append-system-prompt, so a lead never gets theworker reply charter.
countLeads), andneeds the same dead reading twice before closing one (fleetd #359).
AgentControloffersstart(name, kind, args, paneId)andclose(paneId)alongsidesend/submit/status.But
LeadLauncherpasses no initial prompt, sobootstrapTextwould still have to be typedafter the fresh process reaches its prompt.
Open design questions
/exit, Ctrl-D, orAgentControl.close(paneId). The pane mustend up with the right tab label, because a lead's identity is only its tab label
(
fleet.leaders.opus.tab), and a fresh session cannot rename its own tab — invariant 5 reservesthat for the bridge.
bootstrapText. Note that herdr types a pane's launch commandand the pty drops everything past byte 1024 with no error, so a longer argv is not free.
LeadRolloverdrives the relaunch directly, or ends the process and letsensureLeads()notice. The second option interacts with the two-dead-readings rule above./clearloses context; endinga process also loses anything unsaved. The handover file is written and freshness-checked before
any of this runs, which is already the real gate.
Two architects are forming independent positions on these; their answers will be added as comments.
Not the same as #595
#595 asked whether to roll a lead on a timer to catch up on instruction changes, and answered
no. This ticket is about a CLI binary update, which a roll cannot deliver at all — so it is a
capability gap, not a scheduling question. #595's recommendation ("notify, do not restart") does not
apply here: no notice makes a running process load a different binary.
#595 also carries a claim that is now stale. It says the roll "has never been observed working
end-to-end". Four rolls were measured working on 2026-09-22. I have not re-run that measurement
today, so I am reporting the earlier result, not a fresh one. #595 should be corrected separately.
Related: #480 (the handover design), #489 (the measured
/clear+ bootstrap concatenation failure),#594 (
LeadRolloverandLeadHeartbeatLoopdisagree onBLOCKED), #595.The version gap is now a direct reading, not an inference
The ticket body said I inferred the running CLI was older from its start time. I have now read it
instead, with
lsofon the live lead process:So the running lead executes 2.1.285. A new process gets 2.1.289 (
claude --version).Installed versions on disk, newest first:
The lead is four releases behind: 285 → 286, 287, 288, 289. Note that 2.1.286's mtime is
Oct 1 16:02, the same minute
pid=96771started (16:02:27) — the newer build landed as this sessionwas starting, so the session has been one behind since its first second.
Re-measure with:
/exitends the process cleanly — measured on a throwaway memberRather than reason about Q1 from the code, I spawned a disposable member and sent it
/exitthroughfleet_send. Two independent signals say the process really ends:The daemon noticed in under a second, through paths that already exist:
Three further facts from the same probe:
fleet_stop{paneId}on it afterwards returnedstopped.localwent back tofree: 2infleet_listwhile the member stillappeared with
state: "failed",liveStatus: "unknown".failedwith theagent_not_foundcause,not as a hang. So the failure is reported rather than silent.
This was a member, not a lead. I did not test
/exiton a lead, and a lead differs in two waysthat matter: it is not in
SessionManager, and its tab label is its identity. So this shows themechanism works and is observable; it does not show the lead path is safe.
A caution about building on the existing wait
From the same log grep as #595:
19 of 20 rolls took the
PICKUP_GRACE_POLLSgrace-release branch rather than observing a realWORKING -> IDLE/DONEpickup. Any restart built onwaitForClearPickupAndSettleinherits that. Ifthe design reuses that wait, it should say why the grace branch is acceptable when the thing being
waited for is a process death rather than a slash command.
Decision — two architects, independent positions, then a comparison round
Both architects answered the brief independently, then each read the other's case. They converged,
and the convergence reversed the position I had provisionally reported. I verified every load-bearing
fact myself; the three that decided it are quoted below.
The design
End the old session with
AgentControl.close(paneId)on a captured pane id, into a stagedreplacement tab. Not a typed
/exit.Why — the three facts that decided it
1. An in-place restart buys no identity continuity.
terminal_idchanges either way.I had reported the opposite.
FleetConfig.java:1112-1115says so in the repo's own words:LeadTabScanner.java:271-278does taketerminal_idfrompane.list, andCallerResolver.java:302-308does key on it — both true, and both were my own check. But
pane.listreports the terminal ofwhatever session is attached now. Reusing the pane does not reuse the terminal. The
worker-resolution window is therefore a cost of any restart, not of one design.
2.
agent_not_foundcannot prove death, so a typed/exithas no reliable success signal.resolveTargetreturns the target verbatim whenagent.listcomes back short:herdr then answers
agent_not_foundfor a live agent. And the retry inagentCallskips exactlythat case:
In the false-empty case
resolved == target, so it throws immediately with no re-resolve. The onepath that manufactures a false death gets zero retries. This is not hypothetical —
LeadLauncher.java:60-69recordsagent.listreporting 0 live for a tabpsproved was running.3.
pane.closedoes not touch that path at all.No
agent.list, no resolution, no false empty. It is a command against a captured pane id rather thana reading of an unreliable registry. Only
send,submit,readandgetgo throughagentCall(
:117,:127,:136,:142).My own
/exitprobe stands as reported — the process really died and the pane survived — but itmeasured that
/exitworks, not that the daemon can reliably know it worked. Those are differentclaims and only the second one matters here.
Conditions carried into the design
Agentbefore the turn-settle wait, and retry the capture.agents.getgoesthrough
agentCall(:142), so the capture itself can false-fail. Carry the pane and tab idsforward; never re-resolve from the terminal later.
agents.status. Corroborate withpane.list, which the scanneralready uses and which never consults
agent.list.HerdrPeerLauncher.java:947-995. Otherwise every roll leaves a dead labelled tab, and thetwo-reading cleanup needs two daemon restarts to clear it —
ensureLeads()has exactly oneproduction call site,
FleetdAssembly.java:304, at boot, with no scheduler (my own check).bootstrapTextby observing lead identity, not by invalidating a cache. Pollleads.get().containsKey(newTerminal), bounded.get()rescans once its TTL expires, so thisforces the refresh as a side effect and proves the property that matters. The scanner's cache
fields are private with no invalidate method and
LeadRolloverholds no handle on it. A guessedsleep is not enough:
scanIntervalSecondsis 10 live, and a fresh CLI boot takes longer, so theterminal will be dropped after its one grace scan (
LeadTabScanner.java:289) and must beobserved coming back.
BLOCKEDon that wait.injectable()acceptsBLOCKED(AgentStatus.java:40-42), so agate built on it would type the handover instruction into a boot trust prompt.
LeadRollover.java:697-710already argues this at length — keep the rule.confirm()keyed onp.leadTerminal().pendingis keyed by token only(
LeadRollover.java:471-503); twoopen()calls give two UUIDs, both pass the ownership check, andeach reaches
continuationRunner.accept(...)at:502. Today that costs two/clears and issurvivable; with a restart it is a race to kill and rebuild the same lead twice.
Leader.kind.FleetConfig.java:1103states
kindandmodelare "descriptive only";LeadLauncher.herdrKindbranches onprofile.isOpenCode()(:357). Verified: insrc/main/javathe only lead-path.kind()read isLeadTabScanner.java:206, on a different type.waitUntilAtTurnBoundaryfor any death or readiness wait. It converts a failedstatus read into "not settled yet" (
LeadRollover.java:715-721). NorwaitForClearPickupAndSettle— 19 of its 20 recorded runs took thePICKUP_GRACE_POLLSgrace-release branch, which releases on a budget rather than on evidence.
Rollout — no boolean, on one condition
Both architects ended up here, and the condition is theirs, not mine:
The bounded relaunch retry must land in the same change. The objection to default-on was never
about losing model context — it was that nothing retries:
ensureLeads()is boot-only, andrunRollover's catch turns any exception into terminalFAILEDwith no retry(
LeadRollover.java:531-543). So the roll owns its own relaunch until it succeeds or exhausts Nattempts, retrying
agent_pane_busyandagent_name_takenthe way members already do(
HerdrPeerLauncher.java:71,:78,:770-783,:836-849). Not a scheduler — a one-shot boundedretry inside the continuation that already exists.
With that in the same commit, there is no flag and no second code path. If the retry is deferred to
a later ticket, the boolean comes back, because shipping default-on with no recovery ships the
zero-leads case to every install at once. I am recording that as a condition on this ticket, not a
preference.
One correction for whoever implements the config work
A "reject a config that sets both the old and new key" check cannot be written where it looks like
it belongs.
LeadRollover's compact constructor already turns null into 20(
FleetConfig.java:1468-1473) andvalidateLeadRollover()runs on the constructed record(
:2930-2939), so after construction "absent" and "set to 20" are indistinguishable. The check mustread the raw YAML, or the fields must stay nullable and default at read time.
Dropped
The
/exit+ same-pane-restart contract probe is dropped — nothing depends on it now. The openquestion of whether herdr frees the agent name after an exit no longer blocks this ticket, though
the fixed
"lead-" + nameis still a gap (filed separately).Not checked by anyone, stated plainly
OpenCode's exit behaviour;
/exiton a real lead; whether a fresh lead meets a trust prompt at boot;and the live
fleetd.yamlvalues from a member's worktree — both architects reportedwiki/empty andfleetd/fleetd.yamlabsent, exactly as the addendum predicts. No build or test was run in anyarchitect turn. The canonical-block sync check is unsatisfiable for a member and stays mine.
Implementation split — three units, and two facts that made unit 2 smaller
I read the launch path before briefing it. Two things the design discussion did not have, both read
in the code at
809b7d9, not inferred:1. The "staged replacement tab" already exists.
LeadLauncher.launch()never reuses a pane.LeadLauncher.java:347-392creates a fresh tab on every launch and only labels it once the starthas succeeded:
So unit 2 does not need to write launch code, stage a tab, or decide a label ordering. It needs a
seam onto this method and the discipline to close the old tab before the new one is labelled —
otherwise two tabs briefly carry the same label and
countLeadsreads two live leads.It also means the bounded-retry condition is already half-met:
ResilientAgentLaunch(#727, merged809b7d9) retriesagent_name_takenandagent_pane_busyinside oneagents.start. What it doesnot retry is
ensureWorkspace,createTab, a tab returning no seed pane, or a herdr blip. Thatis what unit 1's outer loop covers.
2. The identity gate the design asks for is already wired. No new lookup is needed.
Fleetd.leadRollover(Fleetd.java:953-968) already receivesliveLeadTerminals, aSupplier<Map<String, String>>of terminal → lead name, and it isLeadTabScanneritself — live,TTL-rescanning, read on every call. The design's "poll
leads.get().containsKey(newTerminal),bounded" is a use of a parameter that is already there. Today only
leadWorkspaceconsumes it, andit throws the lead name away after resolving a cwd — unit 2 needs that name to call the launcher,
so the lambda has to stop discarding it.
The units
LeadLauncher:public Agent relaunch(String name)+ bounded attempt loop;launch()returnsAgentnotbooleanLeadRollover.confirm(): single-flight claim keyed onp.leadTerminal(), released inrunRollover'sfinallyLeadRollover: end the old process, close pane then tab, relaunch, observe lead identity, then sendbootstrapText1 and 3 are file-disjoint and run in parallel. Unit 2 goes on the merged base, because it rewrites
runRolloverUnguardedand rewiresFleetd.leadRollover, and both would collide with unit 3 in thesame file.
Unit 3 is split out rather than folded in because it is a live defect on its own terms: two
open()calls on one lead terminal mint two tokens, both pass
confirm()'s ownership check, and both reachcontinuationRunner.accept(...)atLeadRollover.java:502. Today that is two/clears. After unit2 it is a race to kill and rebuild the same lead twice.
A correction to one of the design conditions
The decision comment says the fresh terminal "will be dropped after its one grace scan
(
LeadTabScanner.java:289) and must be observed coming back". I read that line, and the gracemechanism runs the other way. The conclusion it supports is right; the reason is not, and the wrong
reason predicts flapping that cannot happen.
cached(:111) is the previous scan's result. So:cached.containsKey(newTerminal)is false for aterminal that has never been reported live, which the comment at
:286-288states as the intent.It is simply absent until a scan actually sees it, and then it stays. It never flaps.
cached, so the first scan afterwe kill it still reports it live, and it is dropped only on the scan after that.
Two consequences for unit 2, and the second is the one that bites:
get()(
:213-217) returnscachedunchanged until the TTL expires, so a guessed sleep can read a mapthat predates the launch entirely.
liveLeadTerminalsto confirm the old process died. It lags by one full scan bydesign, so a dead lead keeps reading as live for up to ~2x
scanIntervalSeconds— 10s live here,so about 20s.
pane.listis the right corroboration, which is what the decision says, though fora different reason than it gives.
There is therefore a window in which the dead old terminal and the live new terminal both resolve to
the same lead name. Nothing reachable counts leads in that window —
countLeadshas one productioncaller,
ensureLeads(), and that is boot-only (FleetdAssembly.java:304) — but unit 2 should notcreate a second counter that would.
Config, for unit 2
FleetConfig.LeadRollover(FleetConfig.java:1465-1474) holdshandoverPath,requireOperatorConfirm,maxDocAgeSeconds,turnSettleSeconds,clearSettleSeconds,bootstrapText. Unit 2 needs one more: a budget for observing the fresh lead become a recognisedlead.
clearSettleSecondsdefaults to 20 and cannot be reused — the livescanIntervalSecondsis10, so 20s buys only two scans, and a CLI boot plus a scan does not reliably fit in that.
No boolean, per the decision above — unit 1 carries the retry that was the condition for that.
Measured: a fresh boot in a lead's config dir does NOT meet a trust prompt
Both architects left this unchecked, and it decides whether unit 2 needs a readiness condition beyond
"the terminal is a recognised lead" —
injectable()acceptsBLOCKED, so a boot dialog would betyped into rather than waited for.
I probed it. A member on the
sonnetprofile withworktree:false, which puts a freshclaudeprocess in a fresh herdr tab with the same two things that decide trust as a lead gets:
It reached its input prompt unattended, took a delegated message as a prompt, ran two shell commands
and answered with a structured
fleet_reply. Nothing blocked, and no human touched the pane. So onthis host, in this config dir, at this cwd, a fresh boot is usable without an operator.
claude --versionreporting 2.1.289 also confirms the version gap from the other direction: a newprocess gets 289 while the running lead executes 285.
What this does not prove. It was a member launch, so its argv is not a lead's — a member carries
--append-system-promptand a lead deliberately does not (LeadLauncher.java:406). Trust isdecided by cwd and config dir, and both match exactly, so I am confident about the trust prompt
specifically. I have not booted a lead through
LeadLauncher.launch()today, and nothing here testsOpenCode, whose exit and boot behaviour is still unchecked by anyone.
The weaker, older evidence stands behind it:
ensureLeads()has launched a lead on this host exactlyonce (
grep -coverfleetd/fleetd.outreturns 1),lead 'opus' launched: profile=opus tab=w4:t2 pane=w4:p2 terminal=term_65910edceb7f267. That line's logger isd.l.b.lead.LeadLauncher, so itpredates the CB-634 rename and is months old. I report it as "this path has worked once", not as a
current reading.
Measured:
pane.closeends theclaudeprocess, andlocatePaneis the right death checkTwo more facts, both needed by unit 2 and neither established before.
1. Closing the pane kills the process
This is the fact the whole ticket rests on. The decision chose
AgentControl.close(paneId)over atyped
/exitbecause the signal is reliable — but nobody checked that closing a pane actually endsthe process, and if it left an orphan
clauderunning, a fresh launch would not pick up a new binaryeither.
Counted with
pgrep -x claude | wc -l(pids only — neverpgrep -florps -f, which print argv,and argv carries
NAME=value):fleet_stopis the production teardown:HerdrPeerLauncher.stop()→agents.close(paneId)→pane.close. So the pane close ends the process. 12 → 13 → 12.One correction to my own method: the second and third counts in my first command ran back-to-back
with no delay, so I re-ran the count in its own call afterwards. The "12 again" above is that
re-reading, not a same-command duplicate.
2. The death check already exists, and the member teardown already has the shape to copy
WorkspaceControl.locatePane(paneId)(WorkspaceControl.java:143-155) callspane.getandtab.listand neveragent.list:So it cannot hit the
agentCallfalse-empty path the decision warns about, and it returnsnullonce the pane is gone. That is the corroboration instrument, and the
nullis the proof.HerdrPeerLauncher.stop()(:883-935) already does the whole sequence correctly, and unit 2 shouldcopy it rather than invent one. Three details in it that the design discussion did not mention:
locatePaneis called BEFORE the close, because afterwards there is no pane to resolve a tabfrom. The tab id must be captured first.
loc.tabPaneCount() == 1. A lead's tab holds exactly one pane(
FleetConfig.java:1112-1115), so this is 1 in practice, but the guard is what stops the codeclosing a human's shared tab. Keep it.
agents.closetreats an already-gone paneas success (
isAlreadyGone(e)) and propagates anything else, because that is the step whose failuremeans the teardown did not happen.
spaces.closeTabnever propagates — by then the pane is alreadyclosed, so a failing tab close is cosmetic tidying and must not mask later cleanup (fleetd #293).
3. Config: the key can be replaced, not added
FleetConfig.LeadRollovercarries@JsonIgnoreProperties(ignoreUnknown = true)(:1464). Soremoving a component does not break an operator yaml that still sets it — the key is ignored.
The live
fleetd.yamlon this host sets onlyhandoverPath,turnSettleSeconds: 300,requireOperatorConfirm: falseandbootstrapText. It does not setclearSettleSeconds, so thatkey is on its default 20 and nothing here depends on its value. I checked this myself because a
worker cannot:
fleetd/fleetd.yamlis gitignored and absent from every worktree.So unit 2 replaces
clearSettleSecondswith the new readiness budget in the same position. The recordkeeps six components, so the ~7 positional constructor calls in
LeadRolloverTestneed their valuechanged and not their shape.
/cleargoes away entirely, and with itwaitForClearPickupAndSettleandRollState.CLEAR_NEVER_SETTLED.One hazard to carry into the change: a different host's yaml that does set
clearSettleSecondswould have it silently ignored after this.
validateLeadRollover()should warn when the retired keyis present, rather than let an operator's setting disappear without a word. I cannot read fleet01's
fleetd.yamlfrom here, so I do not know whether any host sets it.Already live and in the lead's favour
The configured
bootstrapTexton this host already opens with an identity check — "First runfleet_whoami; it must answer primary. If it answers worker, your pane's tab label no longermatches…". That was written for the
/clearpath, and it happens to be exactly the right firstinstruction for a session that woke up in a brand-new tab.
LeadLauncher.launch()labels the new tabwith
lead.tabLabel()after the start succeeds, so the check should pass — and if the labelling everregresses, the fresh lead stops and says so instead of orchestrating as a worker.
Units 1 and 3 are merged. Unit 2 is dispatched.
LeadLauncher.relaunch(String)+ bounded attempt loop7b3beaaLeadRollover.confirm()single-flight per lead terminala332dfdbootstrapTexta332dfdmainis nowa332dfd. I built the merge result of each unit myself in a throwaway worktree and confirmed the merge tree matched the tree I pushed, becausemainmoved under both branches while they worked (#729, then unit 1).2054 base + 3 (#729) + 7 (unit 1) + 6 (unit 3) = 2070. The arithmetic agrees from both ends.
One check I have not run on either unit.
ide_diagnosticsanswersproject_not_found— thefleetdproject is not open in IntelliJ at the moment — so neither merge has had the IDE inspections run over it. The Maven build is all I am reporting.Line numbers in the unit 2 brief, re-measured at
a332dfdUnit 3 added 51 lines to
LeadRollover.java, so three references in my earlier design comments have moved. The brief carries the corrected numbers; correcting them here too, so a later reader of this ticket does not follow a stale one:injectable()-excludes-BLOCKEDargumentLeadRollover.java:697-710:742-752waitUntilAtTurnBoundaryswallowing a failed status read as "not settled":715-721:764-771HerdrPeerLauncher.stop(), the pane-then-tab teardown to copy:883-935:884-937Unchanged and re-checked:
LeadTabScanner.java:289(the grace scan),FleetConfig.java:1464(@JsonIgnorePropertieson theLeadRolloverrecord),FleetConfig.java:1465(the record itself),FleetdAssembly.java:304(the discardedLeadLauncher).Where unit 3 touches unit 2's work
The brief tells the worker, and it belongs on the ticket too:
rollingByTerminalis released inrunRollover'sfinally, not inrunRolloverUnguarded. Unit 2 rewritesrunRolloverUnguarded's body, so every new early exit it adds still returns through thatfinallyand still releases the claim.confirm()is not unit 2's to touch.Still unproven, and only a live roll can settle it
Everything below is untested by any unit's test suite, because no test boots a real CLI:
LeadLauncher.launch()reaches its prompt unattended. I measured a freshclaudeboot at the lead's cwd with the lead'sCLAUDE_CONFIG_DIRdoing exactly that, but through a member launch, so not with a lead's argv.I will verify the live roll myself after unit 2 merges and the daemon is redeployed.
Correction to the brief — unit 2 ships a regression, and it is NOT yours to fix
Read this before you commit. It does not change your scope. It records a consequence of unit 2 that nobody caught during the design, so that you do not fix it by accident and so the next lead does not deploy unit 2 without it.
Filed in full as #737. The short version:
A named lead's ticket reads are checked against its terminal id:
and a configured lead carries a real, non-null terminal (
CallerResolver.java:308→Principal.leader(lead, c.terminal(), c.pid())). Unit 2's whole point is a new pane, so the fresh lead gets a different terminal id. After unit 2, a rolled lead can no longerfleet_pollany ticket its predecessor created — it getsforbidden: this ticket was created by a different session. A lead rolls when its context is full, which is usually with workers running, so every in-flight delegation's report becomes uncollectable.#705's comment 18617 recorded the opposite as safe. That was correct at the time: today's roll types
/clearinto the same pane, so the terminal survives. Unit 2 is the change that makes it false.What this means for you
Nothing to implement. Do not touch
MessageService,ownsTicket,Task, or ticket ownership. Do not add a reassignment call to the roll. Your brief stands exactly as written, and widening it would collide with whatever fix lands for #737.One thing to add to your reply. Your brief already asks you to say which parts of the new path your tests cannot cover. Add this one to that list: your tests use fakes, so the new terminal id is whatever the fake returns, and no test here exercises a real ticket read across a roll. Say so plainly rather than implying the sequence is proven end to end.
For whoever merges unit 2
A fix for #737 must land with or before unit 2 reaches the live daemon — a merge is not a deployment, so there is room between the two. #737 lists three fix shapes and does not pick one. Settle it there, not here.
Not measured
No roll has run under unit 2's code, so nobody has seen the refusal. The four code sites in #737 are read at
main=428a12a.Second correction — unit 3's single-flight claim stops being a per-lead lock under unit 2. Still NOT yours to fix.
#737 is now decided (option 2, role-prefixed owner key — see its comment thread). One finding from it lands in the file you are editing, so you need to know it exists even though you must not act on it.
What the two architects found
Unit 3 added this, and you were told to leave it alone — that still holds:
The release is correct and nothing leaks. Claim and release use the same key, so your rewrite of
runRolloverUnguardedcannot strand it — every early exit you add returns from inside that method, and thefinallyinrunRolloverstill runs. Nothing changes for you here.What does change: it is no longer a lock on the lead. Once your sequence replaces the pane, the fresh lead has a different terminal, so a second
confirm()from it lands on a different map key and is not excluded by the claim the predecessor's continuation still holds. The window is small but real, because your bootstrap delivery happens inside that continuation.The two architects reported this as "breaks" and "survives" respectively. Both were right about different properties, and I checked the code myself: release correct, exclusion no longer per-lead.
What you do about it
Nothing. Do not re-key
rollingByTerminal, do not touchconfirm(), and do not widen your diff. Keying single-flight on the lead identity is unit 4 of #737's plan, and doing it here would collide with that work.One line to add to your reply: say explicitly that your single-flight tests use a fake whose terminal does not change across the roll, so they do not cover a second
confirm()from the new terminal. That is a true statement about your tests and I want it on the record rather than discovered later.And the earlier correction still stands
The ticket-ownership break from my previous comment is confirmed and is wider than I first wrote —
fleet_poll{ticket}is one of three gates that break, along withfleet_status's pending-ask block andfleet_send{turnId}. The last one is the worst: a refused answer means a worker'sfleet_asktimes out after ~55s and it resumes unanswered. None of it is yours to fix. Your scope is unchanged.Sequencing, so you know where your work lands
Your unit may merge to
mainwhen it is ready and verified. It must not reach the running daemon until #737's fix is also in the jar. I will not redeploy on your merge alone. Nothing for you to do — just so you are not surprised that a merged unit is not deployed.Correction for unit 2 — my brief missed the prompt surface. Read this before you commit.
Your config work is right and I am not asking you to change it.
relaunchReadySecondsat 45, with the javadoc naming the 10s scan interval and saying why 20 was rejected, is exactly what I asked for. The raw-YAML retired-key warning is also right, and your javadoc gives the correct reason for it:@JsonIgnoreProperties(ignoreUnknown = true)means Jackson has already dropped the key by the time anyLeadRolloverinstance exists, so a raw check is the only place it can be reported.My brief had a gap. It told you to rewrite the class javadoc and said nothing about the tool description an operator and a lead actually read. That is my defect, not yours. This repo has a mandatory rule that the prompt is part of the product: a change to a
fleet_*tool's semantics must update the text that describes it. Your change alters whatfleet_handoverdoes — from typing/clearinto a pane to killing a process and launching a new one — so that text is now wrong, and it is the only thing a lead reads before deciding to call it.Add to your scope — the operator-facing text
1.
FleetMcp.handoverTool(), the tool description (around:2630-2648in my copy of your tree). Two things in it are now false:statusoutcome list: "the calling turn never settled within turnSettleSeconds so no/clearwas ever sent, or/clearitself never settled so bootstrapText was never sent". Those are the old two failures. You are adding three new ones. The list must name the real outcomes, orstatusdescribes states that can no longer happen and omits every state that can.Keep "it does NOT itself clear the pane, the roll runs once this call's own turn ends" as a statement about ordering, because that is still true and it is load-bearing — just stop calling it "clear".
2.
FleetMcpcomments at:894and:1495both say the continuation sends/clear. Both are now wrong.:1495is the unnamed-primary refusal and its reasoning still holds — an unnamed primary has no pane for the bootstrap to land in — so fix the wording, not the logic.3.
FleetConfigjavadoc, five places.:1424-1429(theturnSettleSecondsparagraph),:1450(its@param),:1462(bootstrapText's@param),:1484(thebootstrapText()accessor).:1427is the worst of them — "a separate wait fromclearSettleSecondsbelow" — it points at a key you deleted.The
turnSettleSecondsparagraph needs real thought rather than a find-and-replace. Its argument is still completely valid and is the most important sentence in the block: "a lead that never goes idle is still doing real work, and clearing it would destroy live context." After your change the stake is higher, not lower — the old failure wiped context, the new one ends a process. Keep the argument and sharpen it.One comment to delete rather than reword
LeadRolloverTest.java:33currently reads:That is history in a code comment, and this project bans it: no ticket numbers, no "replaced X with Y", no "this used to". A comment must read as if the code had always been this way. Git and the commit message keep the history. Name the behaviour the tests protect, not the change you made. The same rule applies to every comment you touch — I will check for it.
Not yours — I am doing these myself
Do not edit these, so we do not collide:
.claude/skills/handover/SKILL.md— four stale/clearmentions. It is a lead-side skill, and most of its text records live rolls on this host that you have no way to check.wiki/11-Features.md— your worktree haswiki/uninitialised, so you cannot edit it and must not try.Acceptance, unchanged except for one addition
Everything in the brief still stands, including the five mutations. Add one check: after your edits,
grep -rn '/clear' fleetd/src/main/java fleetd/src/test/javamust return only member-side hits. For reference, these are the legitimate ones that must stay —/clearis also how a member's context is reset, which is a different feature and nothing to do with rollover:ClaudeCodeLauncher(:948-949),Injector(:597), and the testsInjectorTest,CompletionResolverTest,ClaudeCodeLauncherTest,CompositePeerLauncherTest,SessionManagerTest. Do not touch any of them.In your
fleet_reply, list the operator-facing sentences you rewrote and paste the finalgrepoutput.Second correction, and this one is a design defect in my brief. It changes your step 7 and 9.
I measured the live log while reviewing your config default. Two findings, and the second one matters more than the first.
The numbers, measured just now on
fleetd/fleetd.out19 of those 20 carry an
elapsedMs. Sorted:Median 16507 ms, max 48261 ms, two over 45000 ms.
Read that carefully before you treat it as a verdict on your 45-second default. That
elapsedMscovers the whole old roll, and it is dominated by waiting for the calling lead's own turn to end — a budget of 300 seconds. So it is not the same quantity asrelaunchReadySeconds, and I am not telling you 45 is wrong. What it does show is that this sequence already ran past 45 seconds twice out of 19, and the old sequence contained no process start at all. Yours adds a CLI boot. Keep 45 if you still think it fits, but say in your reply that you saw these numbers and why you kept it.The design defect: your step 7 gate is doing two different jobs
My brief told you to wait on
liveLeadTerminals.get().containsKey(newTerminal)and then sendbootstrapText. That folds together two conditions that are not the same thing, and only one of them is a safety gate:IDLEorDONEidlewhile Claude is still booting; typing into that window loses the keystrokes and can wedge delivery. This is the whole reasonMemberPresenceexists — read its class javadoc.Readiness is the real gate on sending. Recognition is bookkeeping. Treat them separately:
BLOCKEDexactly as the brief said. Never send into an unready pane.relaunchReadySeconds.bootstrapTextanyway and record the timeout as its own outcome.Why I reversed that, because it contradicts my brief
My brief said a recognition timeout must never send
bootstrapText, and yourrelaunchReadySecondsjavadoc now repeats it. I was wrong, and I was wrong for a specific reason worth stating: I carried over a safety argument from the old code without checking that it still applies.In the
/cleardesign, withholding the bootstrap protected a lead that was still alive. Every/clear-era failure was safe — context intact, nothing lost. In your design the old process is already dead by the time step 7 runs. So withholding the bootstrap protects nothing. It strands a live, freshly booted pane that nobody has told to read the handover file. The operator sees an empty session, and the file sits on disk with nothing pointing at it.There is also a plain timing reason. Recognition comes from a scan with a TTL, so a timeout here can simply mean the scan has not ticked yet, while the session is perfectly fine. Refusing to bootstrap because of a scan interval is the wrong trade.
Keep "never send on failure" for the two states where nothing is alive to send to: the turn never settled (nothing was killed — this one stays exactly as it is, and it is still the most important branch in the unit) and the relaunch failed every attempt.
What this changes in your deliverables
relaunchReadySeconds's javadoc must stop saying a timeout never sendsbootstrapText. Say what it really bounds, and that the bootstrap still goes out when the pane is ready.RollStatenow means "alive and bootstrapped, but not recognised as a lead". Name it so that reads clearly, and log it atwarn— it needs an operator's attention even though the roll mostly worked.bootstrapTextwas sent" with: recognition times out while the pane is ready ⇒bootstrapTextIS sent, and the not-recognised outcome is recorded. Add a separate test that a pane which never becomes ready gets nobootstrapText, because that is the case where the old assertion was protecting something real.The
BLOCKEDmutation in the brief still applies and now belongs to the readiness gate.Still not yours
.claude/skills/handover/SKILL.mdandwiki/11-Features.mdare mine. I have already put a dated note in the skill saying the old/cleartext stays accurate until the new jar is deployed, and how a lead can tell which behaviour is live.If any of this disagrees with the brief, this comment is newer and it wins. Re-read both of my comments before you commit.