A handover rolls the lead with /clear, so the process survives and a Claude Code CLI update is never picked up #726

Open
opened 2026-10-04 17:28:10 +02:00 by ltms · 9 comments
Owner

The problem

fleet_handover replaces a lead's context by typing /clear into its pane. The claude OS
process never dies. So a newer Claude Code CLI installed on disk is never loaded, no matter how
many times the lead is rolled.

The operator asked for this on 2026-10-04: "I want to see a true new restart of lead session, not
just /new - doing so ensure Claude Code update pick up"
.

Measured, 2026-10-04 17:24 CEST on the Mac

lead process   pid=96771  comm=claude   started Thu Oct  1 16:02:27 2026
claude CLI on disk        2.1.289       binary mtime Oct 4 01:15

The running lead started three days before the CLI on disk was written. A roll would not change
that. I did not read the running process's own version number — I am inferring it is older from the
start time against the binary mtime, which is enough to show the gap but is not a version reading.

Where it is

fleetd/src/main/java/dev/ltms/fleet/lead/LeadRollover.java:571

agents.send(lead, "/clear");

Then waitForClearPickupAndSettle(cfg.clearSettleSeconds()), then
agents.send(lead, cfg.bootstrapTextFor(p.handoverPath())). Outcomes are TURN_NEVER_SETTLED,
CLEAR_NEVER_SETTLED, ROLLED. Nothing in that path ends a process.

Why this looks tractable

dev/ltms/fleet/lead/LeadLauncher.java already launches a lead from config, so the daemon does not
need to learn a new launch command:

  • ensureLeads() at :105
  • agents.start("lead-" + name, herdrKind(profile), argv, tab.rootPaneId()) at :326
  • leadArgv(profile) at :371 builds argv from the profile and adds --mcp-config, --model and
    the auto-compact pin. It deliberately omits --append-system-prompt, so a lead never gets the
    worker reply charter.
  • it already treats a labelled tab with no running agent as not a live lead (countLeads), and
    needs the same dead reading twice before closing one (fleetd #359).

AgentControl offers start(name, kind, args, paneId) and close(paneId) alongside
send/submit/status.

But LeadLauncher passes no initial prompt, so bootstrapText would still have to be typed
after the fresh process reaches its prompt.

Open design questions

  1. How the old process ends: typed /exit, Ctrl-D, or AgentControl.close(paneId). The pane must
    end up with the right tab label, because a lead's identity is only its tab label
    (fleet.leaders.opus.tab), and a fresh session cannot rename its own tab — invariant 5 reserves
    that for the bridge.
  2. How the fresh lead receives bootstrapText. Note that herdr types a pane's launch command
    and the pty drops everything past byte 1024 with no error, so a longer argv is not free.
  3. Whether LeadRollover drives the relaunch directly, or ends the process and lets
    ensureLeads() notice. The second option interacts with the two-dead-readings rule above.
  4. What prevents two live leads at once, and what prevents zero leads permanently.
  5. Whether this becomes the default or sits behind a new config key. A /clear loses context; ending
    a process also loses anything unsaved. The handover file is written and freshness-checked before
    any of this runs, which is already the real gate.

Two architects are forming independent positions on these; their answers will be added as comments.

Not the same as #595

#595 asked whether to roll a lead on a timer to catch up on instruction changes, and answered
no. This ticket is about a CLI binary update, which a roll cannot deliver at all — so it is a
capability gap, not a scheduling question. #595's recommendation ("notify, do not restart") does not
apply here: no notice makes a running process load a different binary.

#595 also carries a claim that is now stale. It says the roll "has never been observed working
end-to-end". Four rolls were measured working on 2026-09-22. I have not re-run that measurement
today, so I am reporting the earlier result, not a fresh one. #595 should be corrected separately.

Related: #480 (the handover design), #489 (the measured /clear + bootstrap concatenation failure),
#594 (LeadRollover and LeadHeartbeatLoop disagree on BLOCKED), #595.

## The problem `fleet_handover` replaces a lead's context by **typing `/clear` into its pane**. The `claude` OS process never dies. So a newer Claude Code CLI installed on disk is never loaded, no matter how many times the lead is rolled. The operator asked for this on 2026-10-04: *"I want to see a true new restart of lead session, not just /new - doing so ensure Claude Code update pick up"*. ## Measured, 2026-10-04 17:24 CEST on the Mac ``` lead process pid=96771 comm=claude started Thu Oct 1 16:02:27 2026 claude CLI on disk 2.1.289 binary mtime Oct 4 01:15 ``` The running lead started three days before the CLI on disk was written. A roll would not change that. I did not read the running process's own version number — I am inferring it is older from the start time against the binary mtime, which is enough to show the gap but is not a version reading. ## Where it is `fleetd/src/main/java/dev/ltms/fleet/lead/LeadRollover.java:571` ```java agents.send(lead, "/clear"); ``` Then `waitForClearPickupAndSettle(cfg.clearSettleSeconds())`, then `agents.send(lead, cfg.bootstrapTextFor(p.handoverPath()))`. Outcomes are `TURN_NEVER_SETTLED`, `CLEAR_NEVER_SETTLED`, `ROLLED`. Nothing in that path ends a process. ## Why this looks tractable `dev/ltms/fleet/lead/LeadLauncher.java` already launches a lead from config, so the daemon does not need to learn a new launch command: - `ensureLeads()` at `:105` - `agents.start("lead-" + name, herdrKind(profile), argv, tab.rootPaneId())` at `:326` - `leadArgv(profile)` at `:371` builds argv from the profile and adds `--mcp-config`, `--model` and the auto-compact pin. It deliberately omits `--append-system-prompt`, so a lead never gets the worker reply charter. - it already treats a labelled tab with **no running agent** as not a live lead (`countLeads`), and needs the same dead reading twice before closing one (fleetd #359). `AgentControl` offers `start(name, kind, args, paneId)` and `close(paneId)` alongside `send`/`submit`/`status`. **But `LeadLauncher` passes no initial prompt**, so `bootstrapText` would still have to be typed after the fresh process reaches its prompt. ## Open design questions 1. How the old process ends: typed `/exit`, Ctrl-D, or `AgentControl.close(paneId)`. The pane must end up with the right tab label, because a lead's identity is *only* its tab label (`fleet.leaders.opus.tab`), and a fresh session cannot rename its own tab — invariant 5 reserves that for the bridge. 2. How the fresh lead receives `bootstrapText`. Note that herdr **types** a pane's launch command and the pty drops everything past byte 1024 with no error, so a longer argv is not free. 3. Whether `LeadRollover` drives the relaunch directly, or ends the process and lets `ensureLeads()` notice. The second option interacts with the two-dead-readings rule above. 4. What prevents two live leads at once, and what prevents zero leads permanently. 5. Whether this becomes the default or sits behind a new config key. A `/clear` loses context; ending a process also loses anything unsaved. The handover file is written and freshness-checked before any of this runs, which is already the real gate. Two architects are forming independent positions on these; their answers will be added as comments. ## Not the same as #595 #595 asked whether to roll a lead **on a timer** to catch up on *instruction* changes, and answered no. This ticket is about a **CLI binary** update, which a roll cannot deliver at all — so it is a capability gap, not a scheduling question. #595's recommendation ("notify, do not restart") does not apply here: no notice makes a running process load a different binary. **#595 also carries a claim that is now stale.** It says the roll "has never been observed working end-to-end". Four rolls were measured working on 2026-09-22. I have not re-run that measurement today, so I am reporting the earlier result, not a fresh one. #595 should be corrected separately. Related: #480 (the handover design), #489 (the measured `/clear` + bootstrap concatenation failure), #594 (`LeadRollover` and `LeadHeartbeatLoop` disagree on `BLOCKED`), #595.
Author
Owner

The version gap is now a direct reading, not an inference

The ticket body said I inferred the running CLI was older from its start time. I have now read it
instead, with lsof on the live lead process:

lsof -p 96771 | awk '$4=="txt"'
  /Users/dai.ha/.local/share/claude/versions/2.1.285

So the running lead executes 2.1.285. A new process gets 2.1.289 (claude --version).

Installed versions on disk, newest first:

  2.1.289      Oct 4 01:15
  2.1.288      Oct 2 22:25
  2.1.287      Oct 1 20:14
  2.1.286      Oct 1 16:02
  2.1.285      Sep 29 22:07
  2.1.283      Sep 26 01:53

The lead is four releases behind: 285 → 286, 287, 288, 289. Note that 2.1.286's mtime is
Oct 1 16:02, the same minute pid=96771 started (16:02:27) — the newer build landed as this session
was starting, so the session has been one behind since its first second.

Re-measure with:

lsof -p <lead pid> | awk '$4=="txt" {print $NF}' | grep versions
claude --version

/exit ends the process cleanly — measured on a throwaway member

Rather than reason about Q1 from the code, I spawned a disposable member and sent it /exit through
fleet_send. Two independent signals say the process really ends:

claude processes:  12 before  ->  11 after
herdr:             agent target term_65d056eba1ee282 not found

The daemon noticed in under a second, through paths that already exist:

17:29:53.819  DEBUG StatusPoller      target term_… gone; dropping its queue
17:29:53.820  WARN  Injector          term_… is gone, dropping its queue: 0 message(s) failed
17:29:53.822  WARN  SessionManager    member … can no longer be delegated to: its turn never resolved
17:29:53.823  WARN  CompletionResolver failing send via turn-stall fallback

Three further facts from the same probe:

  1. The pane survived the agent's death — fleet_stop{paneId} on it afterwards returned stopped.
  2. The seat was released: local went back to free: 2 in fleet_list while the member still
    appeared with state: "failed", liveStatus: "unknown".
  3. An async ticket aimed at the dying session resolves as failed with the agent_not_found cause,
    not as a hang. So the failure is reported rather than silent.

This was a member, not a lead. I did not test /exit on a lead, and a lead differs in two ways
that matter: it is not in SessionManager, and its tab label is its identity. So this shows the
mechanism works and is observable; it does not show the lead path is safe.

A caution about building on the existing wait

From the same log grep as #595:

lead-rollover: rolled token=     20
never observed as WORKING        19

19 of 20 rolls took the PICKUP_GRACE_POLLS grace-release branch rather than observing a real
WORKING -> IDLE/DONE pickup. Any restart built on waitForClearPickupAndSettle inherits that. If
the design reuses that wait, it should say why the grace branch is acceptable when the thing being
waited for is a process death rather than a slash command.

## The version gap is now a direct reading, not an inference The ticket body said I inferred the running CLI was older from its start time. I have now read it instead, with `lsof` on the live lead process: ``` lsof -p 96771 | awk '$4=="txt"' /Users/dai.ha/.local/share/claude/versions/2.1.285 ``` So the running lead executes **2.1.285**. A new process gets **2.1.289** (`claude --version`). Installed versions on disk, newest first: ``` 2.1.289 Oct 4 01:15 2.1.288 Oct 2 22:25 2.1.287 Oct 1 20:14 2.1.286 Oct 1 16:02 2.1.285 Sep 29 22:07 2.1.283 Sep 26 01:53 ``` **The lead is four releases behind: 285 → 286, 287, 288, 289.** Note that 2.1.286's mtime is Oct 1 16:02, the same minute `pid=96771` started (16:02:27) — the newer build landed as this session was starting, so the session has been one behind since its first second. Re-measure with: ```bash lsof -p <lead pid> | awk '$4=="txt" {print $NF}' | grep versions claude --version ``` ## `/exit` ends the process cleanly — measured on a throwaway member Rather than reason about Q1 from the code, I spawned a disposable member and sent it `/exit` through `fleet_send`. Two independent signals say the process really ends: ``` claude processes: 12 before -> 11 after herdr: agent target term_65d056eba1ee282 not found ``` The daemon noticed in under a second, through paths that already exist: ``` 17:29:53.819 DEBUG StatusPoller target term_… gone; dropping its queue 17:29:53.820 WARN Injector term_… is gone, dropping its queue: 0 message(s) failed 17:29:53.822 WARN SessionManager member … can no longer be delegated to: its turn never resolved 17:29:53.823 WARN CompletionResolver failing send via turn-stall fallback ``` Three further facts from the same probe: 1. The pane survived the agent's death — `fleet_stop{paneId}` on it afterwards returned `stopped`. 2. The seat was released: `local` went back to `free: 2` in `fleet_list` while the member still appeared with `state: "failed"`, `liveStatus: "unknown"`. 3. An async ticket aimed at the dying session resolves as `failed` with the `agent_not_found` cause, not as a hang. So the failure is reported rather than silent. This was a member, not a lead. I did **not** test `/exit` on a lead, and a lead differs in two ways that matter: it is not in `SessionManager`, and its tab label is its identity. So this shows the mechanism works and is observable; it does not show the lead path is safe. ## A caution about building on the existing wait From the same log grep as #595: ``` lead-rollover: rolled token= 20 never observed as WORKING 19 ``` 19 of 20 rolls took the `PICKUP_GRACE_POLLS` grace-release branch rather than observing a real `WORKING -> IDLE/DONE` pickup. Any restart built on `waitForClearPickupAndSettle` inherits that. If the design reuses that wait, it should say why the grace branch is acceptable when the thing being waited for is a process death rather than a slash command.
Author
Owner

Decision — two architects, independent positions, then a comparison round

Both architects answered the brief independently, then each read the other's case. They converged,
and the convergence reversed the position I had provisionally reported. I verified every load-bearing
fact myself; the three that decided it are quoted below.

The design

End the old session with AgentControl.close(paneId) on a captured pane id, into a staged
replacement tab.
Not a typed /exit.

Why — the three facts that decided it

1. An in-place restart buys no identity continuity. terminal_id changes either way.

I had reported the opposite. FleetConfig.java:1112-1115 says so in the repo's own words:

A herdr terminal_id changes every time the lead's session restarts, so pinning one cost a config
edit and a daemon restart per restart. A tab is stable: a human opens it once, it holds exactly one
pane, and its label survives restarts of the agent inside it — so identity is now the tab label
alone.

LeadTabScanner.java:271-278 does take terminal_id from pane.list, and CallerResolver.java:302-308
does key on it — both true, and both were my own check. But pane.list reports the terminal of
whatever session is attached now. Reusing the pane does not reuse the terminal. The
worker-resolution window is therefore a cost of any restart, not of one design.

2. agent_not_found cannot prove death, so a typed /exit has no reliable success signal.

resolveTarget returns the target verbatim when agent.list comes back short:

return target; // unknown terminal — let herdr report it against the original target
                                              // AgentControl.java:85

herdr then answers agent_not_found for a live agent. And the retry in agentCall skips exactly
that case:

if (!"agent_not_found".equals(e.code()) || resolved.equals(target)) throw e;   // AgentControl.java:52

In the false-empty case resolved == target, so it throws immediately with no re-resolve. The one
path that manufactures a false death gets zero retries. This is not hypothetical —
LeadLauncher.java:60-69 records agent.list reporting 0 live for a tab ps proved was running.

3. pane.close does not touch that path at all.

public void close(String paneId) { herdr.call("pane.close", Map.of("pane_id", paneId)); }
                                              // AgentControl.java:161-163

No agent.list, no resolution, no false empty. It is a command against a captured pane id rather than
a reading of an unreliable registry. Only send, submit, read and get go through agentCall
(:117, :127, :136, :142).

My own /exit probe stands as reported — the process really died and the pane survived — but it
measured that /exit works, not that the daemon can reliably know it worked. Those are different
claims and only the second one matters here.

Conditions carried into the design

  • Capture the Agent before the turn-settle wait, and retry the capture. agents.get goes
    through agentCall (:142), so the capture itself can false-fail. Carry the pane and tab ids
    forward; never re-resolve from the terminal later.
  • Do not confirm death with agents.status. Corroborate with pane.list, which the scanner
    already uses and which never consults agent.list.
  • Close the old tab as well as the pane, pane-then-tab, following member teardown at
    HerdrPeerLauncher.java:947-995. Otherwise every roll leaves a dead labelled tab, and the
    two-reading cleanup needs two daemon restarts to clear it — ensureLeads() has exactly one
    production call site, FleetdAssembly.java:304, at boot, with no scheduler (my own check).
  • Gate bootstrapText by observing lead identity, not by invalidating a cache. Poll
    leads.get().containsKey(newTerminal), bounded. get() rescans once its TTL expires, so this
    forces the refresh as a side effect and proves the property that matters. The scanner's cache
    fields are private with no invalidate method and LeadRollover holds no handle on it. A guessed
    sleep is not enough: scanIntervalSeconds is 10 live, and a fresh CLI boot takes longer, so the
    terminal will be dropped after its one grace scan (LeadTabScanner.java:289) and must be
    observed coming back.
  • Refuse BLOCKED on that wait. injectable() accepts BLOCKED (AgentStatus.java:40-42), so a
    gate built on it would type the handover instruction into a boot trust prompt.
    LeadRollover.java:697-710 already argues this at length — keep the rule.
  • Single-flight confirm() keyed on p.leadTerminal(). pending is keyed by token only
    (LeadRollover.java:471-503); two open() calls give two UUIDs, both pass the ownership check, and
    each reaches continuationRunner.accept(...) at :502. Today that costs two /clears and is
    survivable; with a restart it is a race to kill and rebuild the same lead twice.
  • The exit/launch backend comes from the profile, never from Leader.kind. FleetConfig.java:1103
    states kind and model are "descriptive only"; LeadLauncher.herdrKind branches on
    profile.isOpenCode() (:357). Verified: in src/main/java the only lead-path .kind() read is
    LeadTabScanner.java:206, on a different type.
  • Never reuse waitUntilAtTurnBoundary for any death or readiness wait. It converts a failed
    status read into "not settled yet" (LeadRollover.java:715-721). Nor
    waitForClearPickupAndSettle — 19 of its 20 recorded runs took the PICKUP_GRACE_POLLS
    grace-release branch, which releases on a budget rather than on evidence.

Rollout — no boolean, on one condition

Both architects ended up here, and the condition is theirs, not mine:

The bounded relaunch retry must land in the same change. The objection to default-on was never
about losing model context — it was that nothing retries: ensureLeads() is boot-only, and
runRollover's catch turns any exception into terminal FAILED with no retry
(LeadRollover.java:531-543). So the roll owns its own relaunch until it succeeds or exhausts N
attempts, retrying agent_pane_busy and agent_name_taken the way members already do
(HerdrPeerLauncher.java:71, :78, :770-783, :836-849). Not a scheduler — a one-shot bounded
retry inside the continuation that already exists.

With that in the same commit, there is no flag and no second code path. If the retry is deferred to
a later ticket, the boolean comes back
, because shipping default-on with no recovery ships the
zero-leads case to every install at once. I am recording that as a condition on this ticket, not a
preference.

One correction for whoever implements the config work

A "reject a config that sets both the old and new key" check cannot be written where it looks like
it belongs. LeadRollover's compact constructor already turns null into 20
(FleetConfig.java:1468-1473) and validateLeadRollover() runs on the constructed record
(:2930-2939), so after construction "absent" and "set to 20" are indistinguishable. The check must
read the raw YAML, or the fields must stay nullable and default at read time.

Dropped

The /exit + same-pane-restart contract probe is dropped — nothing depends on it now. The open
question of whether herdr frees the agent name after an exit no longer blocks this ticket, though
the fixed "lead-" + name is still a gap (filed separately).

Not checked by anyone, stated plainly

OpenCode's exit behaviour; /exit on a real lead; whether a fresh lead meets a trust prompt at boot;
and the live fleetd.yaml values from a member's worktree — both architects reported wiki/ empty and
fleetd/fleetd.yaml absent, exactly as the addendum predicts. No build or test was run in any
architect turn. The canonical-block sync check is unsatisfiable for a member and stays mine.

## Decision — two architects, independent positions, then a comparison round Both architects answered the brief independently, then each read the other's case. **They converged**, and the convergence reversed the position I had provisionally reported. I verified every load-bearing fact myself; the three that decided it are quoted below. ### The design **End the old session with `AgentControl.close(paneId)` on a captured pane id, into a staged replacement tab.** Not a typed `/exit`. ### Why — the three facts that decided it **1. An in-place restart buys no identity continuity. `terminal_id` changes either way.** I had reported the opposite. `FleetConfig.java:1112-1115` says so in the repo's own words: > A herdr `terminal_id` changes every time the lead's session restarts, so pinning one cost a config > edit and a daemon restart per restart. A tab is stable: a human opens it once, it holds exactly one > pane, and its label survives restarts of the agent inside it — **so identity is now the tab label > alone.** `LeadTabScanner.java:271-278` does take `terminal_id` from `pane.list`, and `CallerResolver.java:302-308` does key on it — both true, and both were my own check. But `pane.list` reports the terminal of whatever session is attached *now*. Reusing the pane does not reuse the terminal. The worker-resolution window is therefore a cost of **any** restart, not of one design. **2. `agent_not_found` cannot prove death, so a typed `/exit` has no reliable success signal.** `resolveTarget` returns the target verbatim when `agent.list` comes back short: ```java return target; // unknown terminal — let herdr report it against the original target // AgentControl.java:85 ``` herdr then answers `agent_not_found` for a **live** agent. And the retry in `agentCall` skips exactly that case: ```java if (!"agent_not_found".equals(e.code()) || resolved.equals(target)) throw e; // AgentControl.java:52 ``` In the false-empty case `resolved == target`, so it throws immediately with **no** re-resolve. The one path that manufactures a false death gets zero retries. This is not hypothetical — `LeadLauncher.java:60-69` records `agent.list` reporting 0 live for a tab `ps` proved was running. **3. `pane.close` does not touch that path at all.** ```java public void close(String paneId) { herdr.call("pane.close", Map.of("pane_id", paneId)); } // AgentControl.java:161-163 ``` No `agent.list`, no resolution, no false empty. It is a command against a captured pane id rather than a reading of an unreliable registry. Only `send`, `submit`, `read` and `get` go through `agentCall` (`:117`, `:127`, `:136`, `:142`). My own `/exit` probe stands as reported — the process really died and the pane survived — but it measured that `/exit` *works*, not that the daemon can reliably *know* it worked. Those are different claims and only the second one matters here. ### Conditions carried into the design - **Capture the `Agent` before the turn-settle wait, and retry the capture.** `agents.get` goes through `agentCall` (`:142`), so the capture itself can false-fail. Carry the pane and tab ids forward; never re-resolve from the terminal later. - **Do not confirm death with `agents.status`.** Corroborate with `pane.list`, which the scanner already uses and which never consults `agent.list`. - **Close the old tab as well as the pane**, pane-then-tab, following member teardown at `HerdrPeerLauncher.java:947-995`. Otherwise every roll leaves a dead labelled tab, and the two-reading cleanup needs **two** daemon restarts to clear it — `ensureLeads()` has exactly one production call site, `FleetdAssembly.java:304`, at boot, with no scheduler (my own check). - **Gate `bootstrapText` by observing lead identity, not by invalidating a cache.** Poll `leads.get().containsKey(newTerminal)`, bounded. `get()` rescans once its TTL expires, so this forces the refresh as a side effect *and* proves the property that matters. The scanner's cache fields are private with no invalidate method and `LeadRollover` holds no handle on it. A guessed sleep is not enough: `scanIntervalSeconds` is 10 live, and a fresh CLI boot takes longer, so the terminal **will** be dropped after its one grace scan (`LeadTabScanner.java:289`) and must be observed coming back. - **Refuse `BLOCKED` on that wait.** `injectable()` accepts `BLOCKED` (`AgentStatus.java:40-42`), so a gate built on it would type the handover instruction into a boot trust prompt. `LeadRollover.java:697-710` already argues this at length — keep the rule. - **Single-flight `confirm()` keyed on `p.leadTerminal()`.** `pending` is keyed by token only (`LeadRollover.java:471-503`); two `open()` calls give two UUIDs, both pass the ownership check, and each reaches `continuationRunner.accept(...)` at `:502`. Today that costs two `/clear`s and is survivable; with a restart it is a race to kill and rebuild the same lead twice. - **The exit/launch backend comes from the profile, never from `Leader.kind`.** `FleetConfig.java:1103` states `kind` and `model` are "descriptive only"; `LeadLauncher.herdrKind` branches on `profile.isOpenCode()` (`:357`). Verified: in `src/main/java` the only lead-path `.kind()` read is `LeadTabScanner.java:206`, on a different type. - **Never reuse `waitUntilAtTurnBoundary` for any death or readiness wait.** It converts a failed status read into "not settled yet" (`LeadRollover.java:715-721`). Nor `waitForClearPickupAndSettle` — 19 of its 20 recorded runs took the `PICKUP_GRACE_POLLS` grace-release branch, which releases on a budget rather than on evidence. ### Rollout — no boolean, on one condition Both architects ended up here, and the condition is theirs, not mine: **The bounded relaunch retry must land in the same change.** The objection to default-on was never about losing model context — it was that nothing retries: `ensureLeads()` is boot-only, and `runRollover`'s catch turns any exception into terminal `FAILED` with no retry (`LeadRollover.java:531-543`). So the roll owns its own relaunch until it succeeds or exhausts N attempts, retrying `agent_pane_busy` and `agent_name_taken` the way members already do (`HerdrPeerLauncher.java:71`, `:78`, `:770-783`, `:836-849`). Not a scheduler — a one-shot bounded retry inside the continuation that already exists. With that in the same commit, there is no flag and no second code path. **If the retry is deferred to a later ticket, the boolean comes back**, because shipping default-on with no recovery ships the zero-leads case to every install at once. I am recording that as a condition on this ticket, not a preference. ### One correction for whoever implements the config work A "reject a config that sets both the old and new key" check **cannot** be written where it looks like it belongs. `LeadRollover`'s compact constructor already turns null into 20 (`FleetConfig.java:1468-1473`) and `validateLeadRollover()` runs on the constructed record (`:2930-2939`), so after construction "absent" and "set to 20" are indistinguishable. The check must read the raw YAML, or the fields must stay nullable and default at read time. ### Dropped The `/exit` + same-pane-restart contract probe is **dropped** — nothing depends on it now. The open question of whether herdr frees the agent *name* after an exit no longer blocks this ticket, though the fixed `"lead-" + name` is still a gap (filed separately). ### Not checked by anyone, stated plainly OpenCode's exit behaviour; `/exit` on a real lead; whether a fresh lead meets a trust prompt at boot; and the live `fleetd.yaml` values from a member's worktree — both architects reported `wiki/` empty and `fleetd/fleetd.yaml` absent, exactly as the addendum predicts. No build or test was run in any architect turn. The canonical-block sync check is unsatisfiable for a member and stays mine.
Author
Owner

Implementation split — three units, and two facts that made unit 2 smaller

I read the launch path before briefing it. Two things the design discussion did not have, both read
in the code at 809b7d9, not inferred:

1. The "staged replacement tab" already exists. LeadLauncher.launch() never reuses a pane.

LeadLauncher.java:347-392 creates a fresh tab on every launch and only labels it once the start
has succeeded:

Workspace ws = spaces.ensureWorkspace(lead.workspace());
tab = spaces.createTab(ws.workspaceId(), cwd, leadEnv(profile));
...
Agent started = ResilientAgentLaunch.startUniquelyNamed(agents, herdrKind(profile), args,
        tab.rootPaneId(), ...);
// Label AFTER the start succeeds.
spaces.renameTab(tab.tab().tabId(), label);

So unit 2 does not need to write launch code, stage a tab, or decide a label ordering. It needs a
seam onto this method and the discipline to close the old tab before the new one is labelled —
otherwise two tabs briefly carry the same label and countLeads reads two live leads.

It also means the bounded-retry condition is already half-met: ResilientAgentLaunch (#727, merged
809b7d9) retries agent_name_taken and agent_pane_busy inside one agents.start. What it does
not retry is ensureWorkspace, createTab, a tab returning no seed pane, or a herdr blip. That
is what unit 1's outer loop covers.

2. The identity gate the design asks for is already wired. No new lookup is needed.

Fleetd.leadRollover (Fleetd.java:953-968) already receives liveLeadTerminals, a
Supplier<Map<String, String>> of terminal → lead name, and it is LeadTabScanner itself — live,
TTL-rescanning, read on every call. The design's "poll leads.get().containsKey(newTerminal),
bounded" is a use of a parameter that is already there. Today only leadWorkspace consumes it, and
it throws the lead name away after resolving a cwd — unit 2 needs that name to call the launcher,
so the lambda has to stop discarding it.

The units

# Scope Depends on State
1 LeadLauncher: public Agent relaunch(String name) + bounded attempt loop; launch() returns Agent not boolean — delegated
3 LeadRollover.confirm(): single-flight claim keyed on p.leadTerminal(), released in runRollover's finally — delegated
2 LeadRollover: end the old process, close pane then tab, relaunch, observe lead identity, then send bootstrapText 1 and 3 merged not written

1 and 3 are file-disjoint and run in parallel. Unit 2 goes on the merged base, because it rewrites
runRolloverUnguarded and rewires Fleetd.leadRollover, and both would collide with unit 3 in the
same file.

Unit 3 is split out rather than folded in because it is a live defect on its own terms: two open()
calls on one lead terminal mint two tokens, both pass confirm()'s ownership check, and both reach
continuationRunner.accept(...) at LeadRollover.java:502. Today that is two /clears. After unit
2 it is a race to kill and rebuild the same lead twice.

A correction to one of the design conditions

The decision comment says the fresh terminal "will be dropped after its one grace scan
(LeadTabScanner.java:289) and must be observed coming back". I read that line, and the grace
mechanism runs the other way. The conclusion it supports is right; the reason is not, and the wrong
reason predicts flapping that cannot happen.

// LeadTabScanner.java:289
if (cached.containsKey(terminal) && !gracedTerminals.contains(terminal)) {
    byTerminal.put(terminal, entry);
    stillGraced.add(terminal);
}

cached (:111) is the previous scan's result. So:

  • A fresh terminal gets no grace at all. cached.containsKey(newTerminal) is false for a
    terminal that has never been reported live, which the comment at :286-288 states as the intent.
    It is simply absent until a scan actually sees it, and then it stays. It never flaps.
  • The OLD terminal is the one that gets the grace — it is in cached, so the first scan after
    we kill it still reports it live, and it is dropped only on the scan after that.

Two consequences for unit 2, and the second is the one that bites:

  1. Still poll until the new terminal appears. The reason is the TTL, not a grace scan: get()
    (:213-217) returns cached unchanged until the TTL expires, so a guessed sleep can read a map
    that predates the launch entirely.
  2. Do not use liveLeadTerminals to confirm the old process died. It lags by one full scan by
    design, so a dead lead keeps reading as live for up to ~2x scanIntervalSeconds — 10s live here,
    so about 20s. pane.list is the right corroboration, which is what the decision says, though for
    a different reason than it gives.

There is therefore a window in which the dead old terminal and the live new terminal both resolve to
the same lead name. Nothing reachable counts leads in that window — countLeads has one production
caller, ensureLeads(), and that is boot-only (FleetdAssembly.java:304) — but unit 2 should not
create a second counter that would.

Config, for unit 2

FleetConfig.LeadRollover (FleetConfig.java:1465-1474) holds handoverPath,
requireOperatorConfirm, maxDocAgeSeconds, turnSettleSeconds, clearSettleSeconds,
bootstrapText. Unit 2 needs one more: a budget for observing the fresh lead become a recognised
lead. clearSettleSeconds defaults to 20 and cannot be reused — the live scanIntervalSeconds is
10, so 20s buys only two scans, and a CLI boot plus a scan does not reliably fit in that.

No boolean, per the decision above — unit 1 carries the retry that was the condition for that.

Measured: a fresh boot in a lead's config dir does NOT meet a trust prompt

Both architects left this unchecked, and it decides whether unit 2 needs a readiness condition beyond
"the terminal is a recognised lead" — injectable() accepts BLOCKED, so a boot dialog would be
typed into rather than waited for.

I probed it. A member on the sonnet profile with worktree:false, which puts a fresh claude
process in a fresh herdr tab with the same two things that decide trust as a lead gets:

$PWD                 /Users/dai.ha/LTMS/claude-bridge          <- the lead's cwd
$CLAUDE_CONFIG_DIR   /Users/dai.ha/.ccs/instances/ltms         <- opus and sonnet both point here
claude --version     2.1.289 (Claude Code)

It reached its input prompt unattended, took a delegated message as a prompt, ran two shell commands
and answered with a structured fleet_reply. Nothing blocked, and no human touched the pane. So on
this host, in this config dir, at this cwd, a fresh boot is usable without an operator.

claude --version reporting 2.1.289 also confirms the version gap from the other direction: a new
process gets 289 while the running lead executes 285.

What this does not prove. It was a member launch, so its argv is not a lead's — a member carries
--append-system-prompt and a lead deliberately does not (LeadLauncher.java:406). Trust is
decided by cwd and config dir, and both match exactly, so I am confident about the trust prompt
specifically. I have not booted a lead through LeadLauncher.launch() today, and nothing here tests
OpenCode, whose exit and boot behaviour is still unchecked by anyone.

The weaker, older evidence stands behind it: ensureLeads() has launched a lead on this host exactly
once (grep -c over fleetd/fleetd.out returns 1), lead 'opus' launched: profile=opus tab=w4:t2 pane=w4:p2 terminal=term_65910edceb7f267. That line's logger is d.l.b.lead.LeadLauncher, so it
predates the CB-634 rename and is months old. I report it as "this path has worked once", not as a
current reading.

## Implementation split — three units, and two facts that made unit 2 smaller I read the launch path before briefing it. Two things the design discussion did not have, both read in the code at `809b7d9`, not inferred: **1. The "staged replacement tab" already exists. `LeadLauncher.launch()` never reuses a pane.** `LeadLauncher.java:347-392` creates a *fresh* tab on every launch and only labels it once the start has succeeded: ```java Workspace ws = spaces.ensureWorkspace(lead.workspace()); tab = spaces.createTab(ws.workspaceId(), cwd, leadEnv(profile)); ... Agent started = ResilientAgentLaunch.startUniquelyNamed(agents, herdrKind(profile), args, tab.rootPaneId(), ...); // Label AFTER the start succeeds. spaces.renameTab(tab.tab().tabId(), label); ``` So unit 2 does not need to write launch code, stage a tab, or decide a label ordering. It needs a seam onto this method and the discipline to close the old tab *before* the new one is labelled — otherwise two tabs briefly carry the same label and `countLeads` reads two live leads. It also means the bounded-retry condition is already half-met: `ResilientAgentLaunch` (#727, merged `809b7d9`) retries `agent_name_taken` and `agent_pane_busy` inside one `agents.start`. What it does *not* retry is `ensureWorkspace`, `createTab`, a tab returning no seed pane, or a herdr blip. That is what unit 1's outer loop covers. **2. The identity gate the design asks for is already wired. No new lookup is needed.** `Fleetd.leadRollover` (`Fleetd.java:953-968`) already receives `liveLeadTerminals`, a `Supplier<Map<String, String>>` of terminal → lead name, and it is `LeadTabScanner` itself — live, TTL-rescanning, read on every call. The design's "poll `leads.get().containsKey(newTerminal)`, bounded" is a use of a parameter that is already there. Today only `leadWorkspace` consumes it, and it throws the lead *name* away after resolving a cwd — unit 2 needs that name to call the launcher, so the lambda has to stop discarding it. ### The units | # | Scope | Depends on | State | |---|---|---|---| | 1 | `LeadLauncher`: `public Agent relaunch(String name)` + bounded attempt loop; `launch()` returns `Agent` not `boolean` | — | delegated | | 3 | `LeadRollover.confirm()`: single-flight claim keyed on `p.leadTerminal()`, released in `runRollover`'s `finally` | — | delegated | | 2 | `LeadRollover`: end the old process, close pane then tab, relaunch, observe lead identity, then send `bootstrapText` | 1 and 3 merged | not written | 1 and 3 are file-disjoint and run in parallel. Unit 2 goes on the merged base, because it rewrites `runRolloverUnguarded` and rewires `Fleetd.leadRollover`, and both would collide with unit 3 in the same file. Unit 3 is split out rather than folded in because it is a live defect on its own terms: two `open()` calls on one lead terminal mint two tokens, both pass `confirm()`'s ownership check, and both reach `continuationRunner.accept(...)` at `LeadRollover.java:502`. Today that is two `/clear`s. After unit 2 it is a race to kill and rebuild the same lead twice. ### A correction to one of the design conditions The decision comment says the fresh terminal "**will** be dropped after its one grace scan (`LeadTabScanner.java:289`) and must be observed coming back". I read that line, and the grace mechanism runs the other way. The conclusion it supports is right; the reason is not, and the wrong reason predicts flapping that cannot happen. ```java // LeadTabScanner.java:289 if (cached.containsKey(terminal) && !gracedTerminals.contains(terminal)) { byTerminal.put(terminal, entry); stillGraced.add(terminal); } ``` `cached` (`:111`) is the *previous* scan's result. So: - **A fresh terminal gets no grace at all.** `cached.containsKey(newTerminal)` is false for a terminal that has never been reported live, which the comment at `:286-288` states as the intent. It is simply absent until a scan actually sees it, and then it stays. It never flaps. - **The OLD terminal is the one that gets the grace** — it *is* in `cached`, so the first scan after we kill it still reports it live, and it is dropped only on the scan after that. Two consequences for unit 2, and the second is the one that bites: 1. Still poll until the new terminal appears. The reason is the TTL, not a grace scan: `get()` (`:213-217`) returns `cached` unchanged until the TTL expires, so a guessed sleep can read a map that predates the launch entirely. 2. **Do not use `liveLeadTerminals` to confirm the old process died.** It lags by one full scan by design, so a dead lead keeps reading as live for up to ~2x `scanIntervalSeconds` — 10s live here, so about 20s. `pane.list` is the right corroboration, which is what the decision says, though for a different reason than it gives. There is therefore a window in which the dead old terminal and the live new terminal both resolve to the same lead name. Nothing reachable counts leads in that window — `countLeads` has one production caller, `ensureLeads()`, and that is boot-only (`FleetdAssembly.java:304`) — but unit 2 should not create a second counter that would. ### Config, for unit 2 `FleetConfig.LeadRollover` (`FleetConfig.java:1465-1474`) holds `handoverPath`, `requireOperatorConfirm`, `maxDocAgeSeconds`, `turnSettleSeconds`, `clearSettleSeconds`, `bootstrapText`. Unit 2 needs one more: a budget for observing the fresh lead become a recognised lead. `clearSettleSeconds` defaults to 20 and cannot be reused — the live `scanIntervalSeconds` is 10, so 20s buys only two scans, and a CLI boot plus a scan does not reliably fit in that. No boolean, per the decision above — unit 1 carries the retry that was the condition for that. ### Measured: a fresh boot in a lead's config dir does NOT meet a trust prompt Both architects left this unchecked, and it decides whether unit 2 needs a readiness condition beyond "the terminal is a recognised lead" — `injectable()` accepts `BLOCKED`, so a boot dialog would be typed into rather than waited for. I probed it. A member on the `sonnet` profile with `worktree:false`, which puts a fresh `claude` process in a fresh herdr tab with the same two things that decide trust as a lead gets: ``` $PWD /Users/dai.ha/LTMS/claude-bridge <- the lead's cwd $CLAUDE_CONFIG_DIR /Users/dai.ha/.ccs/instances/ltms <- opus and sonnet both point here claude --version 2.1.289 (Claude Code) ``` It reached its input prompt unattended, took a delegated message as a prompt, ran two shell commands and answered with a structured `fleet_reply`. Nothing blocked, and no human touched the pane. So on this host, in this config dir, at this cwd, a fresh boot is usable without an operator. `claude --version` reporting 2.1.289 also confirms the version gap from the other direction: a new process gets 289 while the running lead executes 285. **What this does not prove.** It was a member launch, so its argv is not a lead's — a member carries `--append-system-prompt` and a lead deliberately does not (`LeadLauncher.java:406`). Trust is decided by cwd and config dir, and both match exactly, so I am confident about the trust prompt specifically. I have not booted a lead through `LeadLauncher.launch()` today, and nothing here tests OpenCode, whose exit and boot behaviour is still unchecked by anyone. The weaker, older evidence stands behind it: `ensureLeads()` has launched a lead on this host exactly once (`grep -c` over `fleetd/fleetd.out` returns 1), `lead 'opus' launched: profile=opus tab=w4:t2 pane=w4:p2 terminal=term_65910edceb7f267`. That line's logger is `d.l.b.lead.LeadLauncher`, so it predates the CB-634 rename and is months old. I report it as "this path has worked once", not as a current reading.
Author
Owner

Measured: pane.close ends the claude process, and locatePane is the right death check

Two more facts, both needed by unit 2 and neither established before.

1. Closing the pane kills the process

This is the fact the whole ticket rests on. The decision chose AgentControl.close(paneId) over a
typed /exit because the signal is reliable — but nobody checked that closing a pane actually ends
the process, and if it left an orphan claude running, a fresh launch would not pick up a new binary
either.

Counted with pgrep -x claude | wc -l (pids only — never pgrep -fl or ps -f, which print argv,
and argv carries NAME=value):

12      before
13      after fleet_spawn{profile:"local", worktree:false}, status idle
12      after fleet_stop{paneId}
12      again, in a separate call

fleet_stop is the production teardown: HerdrPeerLauncher.stop() → agents.close(paneId) →
pane.close. So the pane close ends the process. 12 → 13 → 12.

One correction to my own method: the second and third counts in my first command ran back-to-back
with no delay, so I re-ran the count in its own call afterwards. The "12 again" above is that
re-reading, not a same-command duplicate.

2. The death check already exists, and the member teardown already has the shape to copy

WorkspaceControl.locatePane(paneId) (WorkspaceControl.java:143-155) calls pane.get and
tab.list and never agent.list:

try {
    pane = herdr.call("pane.get", Map.of("pane_id", paneId)).path("pane");
} catch (HerdrException e) {
    log.debug("pane.get({}) — pane already gone: {}", paneId, e.getMessage());
    return null;
}

So it cannot hit the agentCall false-empty path the decision warns about, and it returns null
once the pane is gone. That is the corroboration instrument, and the null is the proof.

HerdrPeerLauncher.stop() (:883-935) already does the whole sequence correctly, and unit 2 should
copy it rather than invent one. Three details in it that the design discussion did not mention:

  • locatePane is called BEFORE the close, because afterwards there is no pane to resolve a tab
    from. The tab id must be captured first.
  • The tab is closed only when loc.tabPaneCount() == 1. A lead's tab holds exactly one pane
    (FleetConfig.java:1112-1115), so this is 1 in practice, but the guard is what stops the code
    closing a human's shared tab. Keep it.
  • The two failures are handled differently on purpose. agents.close treats an already-gone pane
    as success (isAlreadyGone(e)) and propagates anything else, because that is the step whose failure
    means the teardown did not happen. spaces.closeTab never propagates — by then the pane is already
    closed, so a failing tab close is cosmetic tidying and must not mask later cleanup (fleetd #293).

3. Config: the key can be replaced, not added

FleetConfig.LeadRollover carries @JsonIgnoreProperties(ignoreUnknown = true) (:1464). So
removing a component does not break an operator yaml that still sets it — the key is ignored.

The live fleetd.yaml on this host sets only handoverPath, turnSettleSeconds: 300,
requireOperatorConfirm: false and bootstrapText. It does not set clearSettleSeconds, so that
key is on its default 20 and nothing here depends on its value. I checked this myself because a
worker cannot: fleetd/fleetd.yaml is gitignored and absent from every worktree.

So unit 2 replaces clearSettleSeconds with the new readiness budget in the same position. The record
keeps six components, so the ~7 positional constructor calls in LeadRolloverTest need their value
changed and not their shape. /clear goes away entirely, and with it
waitForClearPickupAndSettle and RollState.CLEAR_NEVER_SETTLED.

One hazard to carry into the change: a different host's yaml that does set clearSettleSeconds
would have it silently ignored after this. validateLeadRollover() should warn when the retired key
is present, rather than let an operator's setting disappear without a word. I cannot read fleet01's
fleetd.yaml from here, so I do not know whether any host sets it.

Already live and in the lead's favour

The configured bootstrapText on this host already opens with an identity check — "First run
fleet_whoami; it must answer primary. If it answers worker, your pane's tab label no longer
matches…"
. That was written for the /clear path, and it happens to be exactly the right first
instruction for a session that woke up in a brand-new tab. LeadLauncher.launch() labels the new tab
with lead.tabLabel() after the start succeeds, so the check should pass — and if the labelling ever
regresses, the fresh lead stops and says so instead of orchestrating as a worker.

## Measured: `pane.close` ends the `claude` process, and `locatePane` is the right death check Two more facts, both needed by unit 2 and neither established before. ### 1. Closing the pane kills the process This is the fact the whole ticket rests on. The decision chose `AgentControl.close(paneId)` over a typed `/exit` because the *signal* is reliable — but nobody checked that closing a pane actually ends the process, and if it left an orphan `claude` running, a fresh launch would not pick up a new binary either. Counted with `pgrep -x claude | wc -l` (pids only — never `pgrep -fl` or `ps -f`, which print argv, and argv carries `NAME=value`): ``` 12 before 13 after fleet_spawn{profile:"local", worktree:false}, status idle 12 after fleet_stop{paneId} 12 again, in a separate call ``` `fleet_stop` is the production teardown: `HerdrPeerLauncher.stop()` → `agents.close(paneId)` → `pane.close`. So the pane close ends the process. 12 → 13 → 12. One correction to my own method: the second and third counts in my first command ran back-to-back with no delay, so I re-ran the count in its own call afterwards. The "12 again" above is that re-reading, not a same-command duplicate. ### 2. The death check already exists, and the member teardown already has the shape to copy `WorkspaceControl.locatePane(paneId)` (`WorkspaceControl.java:143-155`) calls `pane.get` and `tab.list` and **never** `agent.list`: ```java try { pane = herdr.call("pane.get", Map.of("pane_id", paneId)).path("pane"); } catch (HerdrException e) { log.debug("pane.get({}) — pane already gone: {}", paneId, e.getMessage()); return null; } ``` So it cannot hit the `agentCall` false-empty path the decision warns about, and it returns `null` once the pane is gone. That is the corroboration instrument, and the `null` is the proof. `HerdrPeerLauncher.stop()` (`:883-935`) already does the whole sequence correctly, and unit 2 should copy it rather than invent one. Three details in it that the design discussion did not mention: - **`locatePane` is called BEFORE the close**, because afterwards there is no pane to resolve a tab from. The tab id must be captured first. - **The tab is closed only when `loc.tabPaneCount() == 1`.** A lead's tab holds exactly one pane (`FleetConfig.java:1112-1115`), so this is 1 in practice, but the guard is what stops the code closing a human's shared tab. Keep it. - **The two failures are handled differently on purpose.** `agents.close` treats an already-gone pane as success (`isAlreadyGone(e)`) and propagates anything else, because that is the step whose failure means the teardown did not happen. `spaces.closeTab` never propagates — by then the pane is already closed, so a failing tab close is cosmetic tidying and must not mask later cleanup (fleetd #293). ### 3. Config: the key can be replaced, not added `FleetConfig.LeadRollover` carries `@JsonIgnoreProperties(ignoreUnknown = true)` (`:1464`). So removing a component does **not** break an operator yaml that still sets it — the key is ignored. The live `fleetd.yaml` on this host sets only `handoverPath`, `turnSettleSeconds: 300`, `requireOperatorConfirm: false` and `bootstrapText`. It does **not** set `clearSettleSeconds`, so that key is on its default 20 and nothing here depends on its value. I checked this myself because a worker cannot: `fleetd/fleetd.yaml` is gitignored and absent from every worktree. So unit 2 replaces `clearSettleSeconds` with the new readiness budget in the same position. The record keeps six components, so the ~7 positional constructor calls in `LeadRolloverTest` need their value changed and not their shape. `/clear` goes away entirely, and with it `waitForClearPickupAndSettle` and `RollState.CLEAR_NEVER_SETTLED`. One hazard to carry into the change: a *different* host's yaml that does set `clearSettleSeconds` would have it silently ignored after this. `validateLeadRollover()` should warn when the retired key is present, rather than let an operator's setting disappear without a word. I cannot read fleet01's `fleetd.yaml` from here, so I do not know whether any host sets it. ### Already live and in the lead's favour The configured `bootstrapText` on this host already opens with an identity check — *"First run `fleet_whoami`; it must answer primary. If it answers worker, your pane's tab label no longer matches…"*. That was written for the `/clear` path, and it happens to be exactly the right first instruction for a session that woke up in a brand-new tab. `LeadLauncher.launch()` labels the new tab with `lead.tabLabel()` after the start succeeds, so the check should pass — and if the labelling ever regresses, the fresh lead stops and says so instead of orchestrating as a worker.
Author
Owner

Units 1 and 3 are merged. Unit 2 is dispatched.

# Scope PR Merge commit
1 LeadLauncher.relaunch(String) + bounded attempt loop #731 7b3beaa
3 LeadRollover.confirm() single-flight per lead terminal #733 a332dfd
2 end the process, close pane then tab, relaunch, observe identity, send bootstrapText — dispatched on a332dfd

main is now a332dfd. I built the merge result of each unit myself in a throwaway worktree and confirmed the merge tree matched the tree I pushed, because main moved under both branches while they worked (#729, then unit 1).

after unit 1:  tests 2064  failures 0  LeadLauncherTest 37
after unit 3:  tests 2070  failures 0  LeadRolloverTest 51  LeadLauncherTest 37

2054 base + 3 (#729) + 7 (unit 1) + 6 (unit 3) = 2070. The arithmetic agrees from both ends.

One check I have not run on either unit. ide_diagnostics answers project_not_found — the fleetd project is not open in IntelliJ at the moment — so neither merge has had the IDE inspections run over it. The Maven build is all I am reporting.

Line numbers in the unit 2 brief, re-measured at a332dfd

Unit 3 added 51 lines to LeadRollover.java, so three references in my earlier design comments have moved. The brief carries the corrected numbers; correcting them here too, so a later reader of this ticket does not follow a stale one:

What Was Now
the injectable()-excludes-BLOCKED argument LeadRollover.java:697-710 :742-752
waitUntilAtTurnBoundary swallowing a failed status read as "not settled" :715-721 :764-771
HerdrPeerLauncher.stop(), the pane-then-tab teardown to copy :883-935 :884-937

Unchanged and re-checked: LeadTabScanner.java:289 (the grace scan), FleetConfig.java:1464 (@JsonIgnoreProperties on the LeadRollover record), FleetConfig.java:1465 (the record itself), FleetdAssembly.java:304 (the discarded LeadLauncher).

Where unit 3 touches unit 2's work

The brief tells the worker, and it belongs on the ticket too: rollingByTerminal is released in runRollover's finally, not in runRolloverUnguarded. Unit 2 rewrites runRolloverUnguarded's body, so every new early exit it adds still returns through that finally and still releases the claim. confirm() is not unit 2's to touch.

Still unproven, and only a live roll can settle it

Everything below is untested by any unit's test suite, because no test boots a real CLI:

  • that a lead relaunched through LeadLauncher.launch() reaches its prompt unattended. I measured a fresh claude boot at the lead's cwd with the lead's CLAUDE_CONFIG_DIR doing exactly that, but through a member launch, so not with a lead's argv.
  • that the version actually moves. 285 → 289 is the check, and the fresh lead has to report it, because the roll destroys the session that asked for it.
  • anything at all about OpenCode's exit and boot behaviour. Nobody has checked it.

I will verify the live roll myself after unit 2 merges and the daemon is redeployed.

## Units 1 and 3 are merged. Unit 2 is dispatched. | # | Scope | PR | Merge commit | |---|---|---|---| | 1 | `LeadLauncher.relaunch(String)` + bounded attempt loop | #731 | `7b3beaa` | | 3 | `LeadRollover.confirm()` single-flight per lead terminal | #733 | `a332dfd` | | 2 | end the process, close pane then tab, relaunch, observe identity, send `bootstrapText` | — | dispatched on `a332dfd` | `main` is now `a332dfd`. I built the *merge result* of each unit myself in a throwaway worktree and confirmed the merge tree matched the tree I pushed, because `main` moved under both branches while they worked (#729, then unit 1). ``` after unit 1: tests 2064 failures 0 LeadLauncherTest 37 after unit 3: tests 2070 failures 0 LeadRolloverTest 51 LeadLauncherTest 37 ``` 2054 base + 3 (#729) + 7 (unit 1) + 6 (unit 3) = 2070. The arithmetic agrees from both ends. **One check I have not run on either unit.** `ide_diagnostics` answers `project_not_found` — the `fleetd` project is not open in IntelliJ at the moment — so neither merge has had the IDE inspections run over it. The Maven build is all I am reporting. ### Line numbers in the unit 2 brief, re-measured at `a332dfd` Unit 3 added 51 lines to `LeadRollover.java`, so three references in my earlier design comments have moved. The brief carries the corrected numbers; correcting them here too, so a later reader of this ticket does not follow a stale one: | What | Was | Now | |---|---|---| | the `injectable()`-excludes-`BLOCKED` argument | `LeadRollover.java:697-710` | `:742-752` | | `waitUntilAtTurnBoundary` swallowing a failed status read as "not settled" | `:715-721` | `:764-771` | | `HerdrPeerLauncher.stop()`, the pane-then-tab teardown to copy | `:883-935` | `:884-937` | Unchanged and re-checked: `LeadTabScanner.java:289` (the grace scan), `FleetConfig.java:1464` (`@JsonIgnoreProperties` on the `LeadRollover` record), `FleetConfig.java:1465` (the record itself), `FleetdAssembly.java:304` (the discarded `LeadLauncher`). ### Where unit 3 touches unit 2's work The brief tells the worker, and it belongs on the ticket too: `rollingByTerminal` is released in **`runRollover`'s `finally`**, not in `runRolloverUnguarded`. Unit 2 rewrites `runRolloverUnguarded`'s body, so every new early exit it adds still returns through that `finally` and still releases the claim. `confirm()` is not unit 2's to touch. ### Still unproven, and only a live roll can settle it Everything below is untested by any unit's test suite, because no test boots a real CLI: - that a lead relaunched through `LeadLauncher.launch()` reaches its prompt unattended. I measured a fresh `claude` boot at the lead's cwd with the lead's `CLAUDE_CONFIG_DIR` doing exactly that, but through a *member* launch, so not with a lead's argv. - that the version actually moves. 285 → 289 is the check, and the fresh lead has to report it, because the roll destroys the session that asked for it. - anything at all about OpenCode's exit and boot behaviour. Nobody has checked it. I will verify the live roll myself after unit 2 merges and the daemon is redeployed.
Author
Owner

Correction to the brief — unit 2 ships a regression, and it is NOT yours to fix

Read this before you commit. It does not change your scope. It records a consequence of unit 2 that nobody caught during the design, so that you do not fix it by accident and so the next lead does not deploy unit 2 without it.

Filed in full as #737. The short version:

A named lead's ticket reads are checked against its terminal id:

// msg/MessageService.java:1463-1465
private static boolean ownsTicket(Task task, String callerTerminal) {
    return callerTerminal == null || callerTerminal.equals(task.creatorTerminal);
}

and a configured lead carries a real, non-null terminal (CallerResolver.java:308 → Principal.leader(lead, c.terminal(), c.pid())). Unit 2's whole point is a new pane, so the fresh lead gets a different terminal id. After unit 2, a rolled lead can no longer fleet_poll any ticket its predecessor created — it gets forbidden: this ticket was created by a different session. A lead rolls when its context is full, which is usually with workers running, so every in-flight delegation's report becomes uncollectable.

#705's comment 18617 recorded the opposite as safe. That was correct at the time: today's roll types /clear into the same pane, so the terminal survives. Unit 2 is the change that makes it false.

What this means for you

Nothing to implement. Do not touch MessageService, ownsTicket, Task, or ticket ownership. Do not add a reassignment call to the roll. Your brief stands exactly as written, and widening it would collide with whatever fix lands for #737.

One thing to add to your reply. Your brief already asks you to say which parts of the new path your tests cannot cover. Add this one to that list: your tests use fakes, so the new terminal id is whatever the fake returns, and no test here exercises a real ticket read across a roll. Say so plainly rather than implying the sequence is proven end to end.

For whoever merges unit 2

A fix for #737 must land with or before unit 2 reaches the live daemon — a merge is not a deployment, so there is room between the two. #737 lists three fix shapes and does not pick one. Settle it there, not here.

Not measured

No roll has run under unit 2's code, so nobody has seen the refusal. The four code sites in #737 are read at main = 428a12a.

## Correction to the brief — unit 2 ships a regression, and it is NOT yours to fix **Read this before you commit.** It does not change your scope. It records a consequence of unit 2 that nobody caught during the design, so that you do not fix it by accident and so the next lead does not deploy unit 2 without it. Filed in full as **#737**. The short version: A named lead's ticket reads are checked against its **terminal id**: ```java // msg/MessageService.java:1463-1465 private static boolean ownsTicket(Task task, String callerTerminal) { return callerTerminal == null || callerTerminal.equals(task.creatorTerminal); } ``` and a configured lead carries a real, non-null terminal (`CallerResolver.java:308` → `Principal.leader(lead, c.terminal(), c.pid())`). Unit 2's whole point is a **new pane**, so the fresh lead gets a **different** terminal id. After unit 2, a rolled lead can no longer `fleet_poll` any ticket its predecessor created — it gets `forbidden: this ticket was created by a different session`. A lead rolls when its context is full, which is usually with workers running, so every in-flight delegation's report becomes uncollectable. #705's comment 18617 recorded the opposite as safe. That was correct at the time: today's roll types `/clear` into the same pane, so the terminal survives. Unit 2 is the change that makes it false. ### What this means for you **Nothing to implement.** Do not touch `MessageService`, `ownsTicket`, `Task`, or ticket ownership. Do not add a reassignment call to the roll. Your brief stands exactly as written, and widening it would collide with whatever fix lands for #737. **One thing to add to your reply.** Your brief already asks you to say which parts of the new path your tests cannot cover. Add this one to that list: your tests use fakes, so the new terminal id is whatever the fake returns, and no test here exercises a real ticket read across a roll. Say so plainly rather than implying the sequence is proven end to end. ### For whoever merges unit 2 A fix for #737 must land **with or before** unit 2 reaches the live daemon — a merge is not a deployment, so there is room between the two. #737 lists three fix shapes and does not pick one. Settle it there, not here. ### Not measured No roll has run under unit 2's code, so nobody has seen the refusal. The four code sites in #737 are read at `main` = `428a12a`.
Author
Owner

Second correction — unit 3's single-flight claim stops being a per-lead lock under unit 2. Still NOT yours to fix.

#737 is now decided (option 2, role-prefixed owner key — see its comment thread). One finding from it lands in the file you are editing, so you need to know it exists even though you must not act on it.

What the two architects found

Unit 3 added this, and you were told to leave it alone — that still holds:

rollingByTerminal.putIfAbsent(p.leadTerminal(), token)   // claim, in confirm()
rollingByTerminal.remove(p.leadTerminal(), p.token())    // release, in runRollover's finally

The release is correct and nothing leaks. Claim and release use the same key, so your rewrite of runRolloverUnguarded cannot strand it — every early exit you add returns from inside that method, and the finally in runRollover still runs. Nothing changes for you here.

What does change: it is no longer a lock on the lead. Once your sequence replaces the pane, the fresh lead has a different terminal, so a second confirm() from it lands on a different map key and is not excluded by the claim the predecessor's continuation still holds. The window is small but real, because your bootstrap delivery happens inside that continuation.

The two architects reported this as "breaks" and "survives" respectively. Both were right about different properties, and I checked the code myself: release correct, exclusion no longer per-lead.

What you do about it

Nothing. Do not re-key rollingByTerminal, do not touch confirm(), and do not widen your diff. Keying single-flight on the lead identity is unit 4 of #737's plan, and doing it here would collide with that work.

One line to add to your reply: say explicitly that your single-flight tests use a fake whose terminal does not change across the roll, so they do not cover a second confirm() from the new terminal. That is a true statement about your tests and I want it on the record rather than discovered later.

And the earlier correction still stands

The ticket-ownership break from my previous comment is confirmed and is wider than I first wrote — fleet_poll{ticket} is one of three gates that break, along with fleet_status's pending-ask block and fleet_send{turnId}. The last one is the worst: a refused answer means a worker's fleet_ask times out after ~55s and it resumes unanswered. None of it is yours to fix. Your scope is unchanged.

Sequencing, so you know where your work lands

Your unit may merge to main when it is ready and verified. It must not reach the running daemon until #737's fix is also in the jar. I will not redeploy on your merge alone. Nothing for you to do — just so you are not surprised that a merged unit is not deployed.

## Second correction — unit 3's single-flight claim stops being a per-lead lock under unit 2. Still NOT yours to fix. #737 is now decided (option 2, role-prefixed owner key — see its comment thread). One finding from it lands in the file you are editing, so you need to know it exists even though you must not act on it. ### What the two architects found Unit 3 added this, and you were told to leave it alone — that still holds: ```java rollingByTerminal.putIfAbsent(p.leadTerminal(), token) // claim, in confirm() rollingByTerminal.remove(p.leadTerminal(), p.token()) // release, in runRollover's finally ``` **The release is correct and nothing leaks.** Claim and release use the same key, so your rewrite of `runRolloverUnguarded` cannot strand it — every early exit you add returns from inside that method, and the `finally` in `runRollover` still runs. Nothing changes for you here. **What does change: it is no longer a lock on the *lead*.** Once your sequence replaces the pane, the fresh lead has a different terminal, so a second `confirm()` from it lands on a different map key and is not excluded by the claim the predecessor's continuation still holds. The window is small but real, because your bootstrap delivery happens inside that continuation. The two architects reported this as "breaks" and "survives" respectively. Both were right about different properties, and I checked the code myself: release correct, exclusion no longer per-lead. ### What you do about it **Nothing.** Do not re-key `rollingByTerminal`, do not touch `confirm()`, and do not widen your diff. Keying single-flight on the lead identity is unit 4 of #737's plan, and doing it here would collide with that work. **One line to add to your reply:** say explicitly that your single-flight tests use a fake whose terminal does not change across the roll, so they do not cover a second `confirm()` from the *new* terminal. That is a true statement about your tests and I want it on the record rather than discovered later. ### And the earlier correction still stands The ticket-ownership break from my previous comment is confirmed and is wider than I first wrote — `fleet_poll{ticket}` is one of **three** gates that break, along with `fleet_status`'s pending-ask block and `fleet_send{turnId}`. The last one is the worst: a refused answer means a worker's `fleet_ask` times out after ~55s and it resumes **unanswered**. None of it is yours to fix. Your scope is unchanged. ### Sequencing, so you know where your work lands Your unit **may merge** to `main` when it is ready and verified. It must **not reach the running daemon** until #737's fix is also in the jar. I will not redeploy on your merge alone. Nothing for you to do — just so you are not surprised that a merged unit is not deployed.
Author
Owner

Correction for unit 2 — my brief missed the prompt surface. Read this before you commit.

Your config work is right and I am not asking you to change it. relaunchReadySeconds at 45, with the javadoc naming the 10s scan interval and saying why 20 was rejected, is exactly what I asked for. The raw-YAML retired-key warning is also right, and your javadoc gives the correct reason for it: @JsonIgnoreProperties(ignoreUnknown = true) means Jackson has already dropped the key by the time any LeadRollover instance exists, so a raw check is the only place it can be reported.

My brief had a gap. It told you to rewrite the class javadoc and said nothing about the tool description an operator and a lead actually read. That is my defect, not yours. This repo has a mandatory rule that the prompt is part of the product: a change to a fleet_* tool's semantics must update the text that describes it. Your change alters what fleet_handover does — from typing /clear into a pane to killing a process and launching a new one — so that text is now wrong, and it is the only thing a lead reads before deciding to call it.

Add to your scope — the operator-facing text

1. FleetMcp.handoverTool(), the tool description (around :2630-2648 in my copy of your tree). Two things in it are now false:

  • "have fleetd clear your pane and bootstrap a fresh lead session against it" — it now ends the process and launches a new one. A lead reading "clear" will not expect its CLI to be restarted. Say what it does.
  • the status outcome list: "the calling turn never settled within turnSettleSeconds so no /clear was ever sent, or /clear itself never settled so bootstrapText was never sent". Those are the old two failures. You are adding three new ones. The list must name the real outcomes, or status describes states that can no longer happen and omits every state that can.

Keep "it does NOT itself clear the pane, the roll runs once this call's own turn ends" as a statement about ordering, because that is still true and it is load-bearing — just stop calling it "clear".

2. FleetMcp comments at :894 and :1495 both say the continuation sends /clear. Both are now wrong. :1495 is the unnamed-primary refusal and its reasoning still holds — an unnamed primary has no pane for the bootstrap to land in — so fix the wording, not the logic.

3. FleetConfig javadoc, five places. :1424-1429 (the turnSettleSeconds paragraph), :1450 (its @param), :1462 (bootstrapText's @param), :1484 (the bootstrapText() accessor). :1427 is the worst of them — "a separate wait from clearSettleSeconds below" — it points at a key you deleted.

The turnSettleSeconds paragraph needs real thought rather than a find-and-replace. Its argument is still completely valid and is the most important sentence in the block: "a lead that never goes idle is still doing real work, and clearing it would destroy live context." After your change the stake is higher, not lower — the old failure wiped context, the new one ends a process. Keep the argument and sharpen it.

One comment to delete rather than reword

LeadRolloverTest.java:33 currently reads:

plus fleetd #726 unit 2's replacement of the /clear-based continuation with a real process restart

That is history in a code comment, and this project bans it: no ticket numbers, no "replaced X with Y", no "this used to". A comment must read as if the code had always been this way. Git and the commit message keep the history. Name the behaviour the tests protect, not the change you made. The same rule applies to every comment you touch — I will check for it.

Not yours — I am doing these myself

Do not edit these, so we do not collide:

  • .claude/skills/handover/SKILL.md — four stale /clear mentions. It is a lead-side skill, and most of its text records live rolls on this host that you have no way to check.
  • wiki/11-Features.md — your worktree has wiki/ uninitialised, so you cannot edit it and must not try.

Acceptance, unchanged except for one addition

Everything in the brief still stands, including the five mutations. Add one check: after your edits, grep -rn '/clear' fleetd/src/main/java fleetd/src/test/java must return only member-side hits. For reference, these are the legitimate ones that must stay — /clear is also how a member's context is reset, which is a different feature and nothing to do with rollover:

ClaudeCodeLauncher (:948-949), Injector (:597), and the tests InjectorTest, CompletionResolverTest, ClaudeCodeLauncherTest, CompositePeerLauncherTest, SessionManagerTest. Do not touch any of them.

In your fleet_reply, list the operator-facing sentences you rewrote and paste the final grep output.

## Correction for unit 2 — my brief missed the prompt surface. Read this before you commit. Your config work is right and I am not asking you to change it. `relaunchReadySeconds` at 45, with the javadoc naming the 10s scan interval and saying why 20 was rejected, is exactly what I asked for. The raw-YAML retired-key warning is also right, and your javadoc gives the correct reason for it: `@JsonIgnoreProperties(ignoreUnknown = true)` means Jackson has already dropped the key by the time any `LeadRollover` instance exists, so a raw check is the only place it can be reported. **My brief had a gap. It told you to rewrite the class javadoc and said nothing about the tool description an operator and a lead actually read.** That is my defect, not yours. This repo has a mandatory rule that the prompt is part of the product: a change to a `fleet_*` tool's semantics must update the text that describes it. Your change alters what `fleet_handover` *does* — from typing `/clear` into a pane to killing a process and launching a new one — so that text is now wrong, and it is the only thing a lead reads before deciding to call it. ### Add to your scope — the operator-facing text **1. `FleetMcp.handoverTool()`, the tool description (around `:2630-2648` in my copy of your tree).** Two things in it are now false: - *"have fleetd **clear your pane** and bootstrap a fresh lead session against it"* — it now ends the process and launches a new one. A lead reading "clear" will not expect its CLI to be restarted. Say what it does. - the `status` outcome list: *"the calling turn never settled within turnSettleSeconds so no `/clear` was ever sent, or `/clear` itself never settled so bootstrapText was never sent"*. Those are the old two failures. You are adding three new ones. The list must name the real outcomes, or `status` describes states that can no longer happen and omits every state that can. Keep *"it does NOT itself clear the pane, the roll runs once this call's own turn ends"* as a statement about **ordering**, because that is still true and it is load-bearing — just stop calling it "clear". **2. `FleetMcp` comments at `:894` and `:1495`** both say the continuation sends `/clear`. Both are now wrong. `:1495` is the unnamed-primary refusal and its reasoning still holds — an unnamed primary has no pane for the *bootstrap* to land in — so fix the wording, not the logic. **3. `FleetConfig` javadoc, five places.** `:1424-1429` (the `turnSettleSeconds` paragraph), `:1450` (its `@param`), `:1462` (`bootstrapText`'s `@param`), `:1484` (the `bootstrapText()` accessor). `:1427` is the worst of them — *"a separate wait from `clearSettleSeconds` below"* — it points at a key you deleted. The `turnSettleSeconds` paragraph needs real thought rather than a find-and-replace. Its argument is still completely valid and is the most important sentence in the block: *"a lead that never goes idle is still doing real work, and clearing it would destroy live context."* **After your change the stake is higher, not lower** — the old failure wiped context, the new one ends a process. Keep the argument and sharpen it. ### One comment to delete rather than reword `LeadRolloverTest.java:33` currently reads: > plus fleetd #726 unit 2's replacement of the `/clear`-based continuation with a real process restart That is history in a code comment, and this project bans it: no ticket numbers, no "replaced X with Y", no "this used to". A comment must read as if the code had always been this way. Git and the commit message keep the history. Name the behaviour the tests protect, not the change you made. The same rule applies to every comment you touch — I will check for it. ### Not yours — I am doing these myself Do **not** edit these, so we do not collide: - `.claude/skills/handover/SKILL.md` — four stale `/clear` mentions. It is a lead-side skill, and most of its text records live rolls on this host that you have no way to check. - `wiki/11-Features.md` — your worktree has `wiki/` uninitialised, so you cannot edit it and must not try. ### Acceptance, unchanged except for one addition Everything in the brief still stands, including the five mutations. Add one check: **after your edits, `grep -rn '/clear' fleetd/src/main/java fleetd/src/test/java` must return only member-side hits.** For reference, these are the legitimate ones that must stay — `/clear` is also how a *member's* context is reset, which is a different feature and nothing to do with rollover: `ClaudeCodeLauncher` (`:948-949`), `Injector` (`:597`), and the tests `InjectorTest`, `CompletionResolverTest`, `ClaudeCodeLauncherTest`, `CompositePeerLauncherTest`, `SessionManagerTest`. Do not touch any of them. In your `fleet_reply`, list the operator-facing sentences you rewrote and paste the final `grep` output.
Author
Owner

Second correction, and this one is a design defect in my brief. It changes your step 7 and 9.

I measured the live log while reviewing your config default. Two findings, and the second one matters more than the first.

The numbers, measured just now on fleetd/fleetd.out

grep -c 'lead-rollover: rolled' fleetd/fleetd.out   -> 20
grep -c 'lead-rollover:'        fleetd/fleetd.out   -> 86    (control, larger as it must be)
grep -c 'Unknown command'       fleetd/fleetd.out   -> 0

19 of those 20 carry an elapsedMs. Sorted:

5179 6284 9168 11191 11917 12519 13895 14429 16314 16507
17929 17960 18468 18692 19373 19775 21122 45685 48261

Median 16507 ms, max 48261 ms, two over 45000 ms.

Read that carefully before you treat it as a verdict on your 45-second default. That elapsedMs covers the whole old roll, and it is dominated by waiting for the calling lead's own turn to end — a budget of 300 seconds. So it is not the same quantity as relaunchReadySeconds, and I am not telling you 45 is wrong. What it does show is that this sequence already ran past 45 seconds twice out of 19, and the old sequence contained no process start at all. Yours adds a CLI boot. Keep 45 if you still think it fits, but say in your reply that you saw these numbers and why you kept it.

The design defect: your step 7 gate is doing two different jobs

My brief told you to wait on liveLeadTerminals.get().containsKey(newTerminal) and then send bootstrapText. That folds together two conditions that are not the same thing, and only one of them is a safety gate:

Condition What it means What goes wrong if you skip it
readiness — the pane reports IDLE or DONE the CLI has booted and can accept input the text is silently lost. herdr reports idle while Claude is still booting; typing into that window loses the keystrokes and can wedge delivery. This is the whole reason MemberPresence exists — read its class javadoc.
recognition — the terminal is in the live-lead map the tab label took, so the daemon will resolve it as a lead the session works but resolves as something else

Readiness is the real gate on sending. Recognition is bookkeeping. Treat them separately:

  1. Wait for readiness first, and refuse BLOCKED exactly as the brief said. Never send into an unready pane.
  2. Wait for recognition second, bounded by relaunchReadySeconds.
  3. If readiness passed but recognition timed out, send bootstrapText anyway and record the timeout as its own outcome.

Why I reversed that, because it contradicts my brief

My brief said a recognition timeout must never send bootstrapText, and your relaunchReadySeconds javadoc now repeats it. I was wrong, and I was wrong for a specific reason worth stating: I carried over a safety argument from the old code without checking that it still applies.

In the /clear design, withholding the bootstrap protected a lead that was still alive. Every /clear-era failure was safe — context intact, nothing lost. In your design the old process is already dead by the time step 7 runs. So withholding the bootstrap protects nothing. It strands a live, freshly booted pane that nobody has told to read the handover file. The operator sees an empty session, and the file sits on disk with nothing pointing at it.

There is also a plain timing reason. Recognition comes from a scan with a TTL, so a timeout here can simply mean the scan has not ticked yet, while the session is perfectly fine. Refusing to bootstrap because of a scan interval is the wrong trade.

Keep "never send on failure" for the two states where nothing is alive to send to: the turn never settled (nothing was killed — this one stays exactly as it is, and it is still the most important branch in the unit) and the relaunch failed every attempt.

What this changes in your deliverables

  • relaunchReadySeconds's javadoc must stop saying a timeout never sends bootstrapText. Say what it really bounds, and that the bootstrap still goes out when the pane is ready.
  • The not-recognised RollState now means "alive and bootstrapped, but not recognised as a lead". Name it so that reads clearly, and log it at warn — it needs an operator's attention even though the roll mostly worked.
  • Replace the brief's test "the new terminal never appears in the live-lead map ⇒ no bootstrapText was sent" with: recognition times out while the pane is ready ⇒ bootstrapText IS sent, and the not-recognised outcome is recorded. Add a separate test that a pane which never becomes ready gets no bootstrapText, because that is the case where the old assertion was protecting something real.
  • Add a mutation: the readiness gate is dropped and only recognition is waited on ⇒ the "never ready" test must die. That is the gate that now carries the safety, so it is the one that must be pinned.

The BLOCKED mutation in the brief still applies and now belongs to the readiness gate.

Still not yours

.claude/skills/handover/SKILL.md and wiki/11-Features.md are mine. I have already put a dated note in the skill saying the old /clear text stays accurate until the new jar is deployed, and how a lead can tell which behaviour is live.

If any of this disagrees with the brief, this comment is newer and it wins. Re-read both of my comments before you commit.

## Second correction, and this one is a design defect in my brief. It changes your step 7 and 9. I measured the live log while reviewing your config default. Two findings, and the second one matters more than the first. ### The numbers, measured just now on `fleetd/fleetd.out` ``` grep -c 'lead-rollover: rolled' fleetd/fleetd.out -> 20 grep -c 'lead-rollover:' fleetd/fleetd.out -> 86 (control, larger as it must be) grep -c 'Unknown command' fleetd/fleetd.out -> 0 ``` 19 of those 20 carry an `elapsedMs`. Sorted: ``` 5179 6284 9168 11191 11917 12519 13895 14429 16314 16507 17929 17960 18468 18692 19373 19775 21122 45685 48261 ``` Median 16507 ms, max 48261 ms, **two over 45000 ms**. Read that carefully before you treat it as a verdict on your 45-second default. That `elapsedMs` covers the **whole old roll**, and it is dominated by waiting for the calling lead's own turn to end — a budget of 300 seconds. So it is not the same quantity as `relaunchReadySeconds`, and I am not telling you 45 is wrong. What it does show is that this sequence already ran past 45 seconds twice out of 19, **and the old sequence contained no process start at all.** Yours adds a CLI boot. Keep 45 if you still think it fits, but say in your reply that you saw these numbers and why you kept it. ### The design defect: your step 7 gate is doing two different jobs My brief told you to wait on `liveLeadTerminals.get().containsKey(newTerminal)` and then send `bootstrapText`. That folds together two conditions that are not the same thing, and only one of them is a safety gate: | Condition | What it means | What goes wrong if you skip it | |---|---|---| | **readiness** — the pane reports `IDLE` or `DONE` | the CLI has booted and can accept input | **the text is silently lost.** herdr reports `idle` while Claude is still booting; typing into that window loses the keystrokes and can wedge delivery. This is the whole reason `MemberPresence` exists — read its class javadoc. | | **recognition** — the terminal is in the live-lead map | the tab label took, so the daemon will resolve it as a lead | the session works but resolves as something else | **Readiness is the real gate on sending. Recognition is bookkeeping.** Treat them separately: 1. Wait for readiness first, and refuse `BLOCKED` exactly as the brief said. Never send into an unready pane. 2. Wait for recognition second, bounded by `relaunchReadySeconds`. 3. **If readiness passed but recognition timed out, send `bootstrapText` anyway** and record the timeout as its own outcome. ### Why I reversed that, because it contradicts my brief My brief said a recognition timeout must never send `bootstrapText`, and your `relaunchReadySeconds` javadoc now repeats it. **I was wrong, and I was wrong for a specific reason worth stating: I carried over a safety argument from the old code without checking that it still applies.** In the `/clear` design, withholding the bootstrap protected a lead that was still alive. Every `/clear`-era failure was safe — context intact, nothing lost. **In your design the old process is already dead by the time step 7 runs.** So withholding the bootstrap protects nothing. It strands a live, freshly booted pane that nobody has told to read the handover file. The operator sees an empty session, and the file sits on disk with nothing pointing at it. There is also a plain timing reason. Recognition comes from a scan with a TTL, so a timeout here can simply mean the scan has not ticked yet, while the session is perfectly fine. Refusing to bootstrap because of a scan interval is the wrong trade. Keep "never send on failure" for the two states where **nothing is alive to send to**: the turn never settled (nothing was killed — this one stays exactly as it is, and it is still the most important branch in the unit) and the relaunch failed every attempt. ### What this changes in your deliverables - `relaunchReadySeconds`'s javadoc must stop saying a timeout never sends `bootstrapText`. Say what it really bounds, and that the bootstrap still goes out when the pane is ready. - The not-recognised `RollState` now means "alive and bootstrapped, but not recognised as a lead". Name it so that reads clearly, and log it at `warn` — it needs an operator's attention even though the roll mostly worked. - Replace the brief's test *"the new terminal never appears in the live-lead map ⇒ no `bootstrapText` was sent"* with: **recognition times out while the pane is ready ⇒ `bootstrapText` IS sent, and the not-recognised outcome is recorded.** Add a separate test that a pane which never becomes ready gets **no** `bootstrapText`, because that is the case where the old assertion was protecting something real. - Add a mutation: **the readiness gate is dropped and only recognition is waited on ⇒ the "never ready" test must die.** That is the gate that now carries the safety, so it is the one that must be pinned. The `BLOCKED` mutation in the brief still applies and now belongs to the readiness gate. ### Still not yours `.claude/skills/handover/SKILL.md` and `wiki/11-Features.md` are mine. I have already put a dated note in the skill saying the old `/clear` text stays accurate until the new jar is deployed, and how a lead can tell which behaviour is live. If any of this disagrees with the brief, **this comment is newer and it wins.** Re-read both of my comments before you commit.
Sign in to join this conversation.
1 Participants
Notifications
Due Date
No due date set.
Dependencies

No dependencies set.

Reference: fleet/fleetd#726