Lead session rollover: the lead asks for a handover, fleetd verifies it, clears the pane and boots the next lead #480

Open
opened 2026-09-11 01:09:38 +02:00 by ltms · 10 comments
Owner

Why

A lead session fills up its context and then has to be replaced by hand. Today that means the
operator notices, asks the lead to write a handover file, starts a new session, and pastes a
pointer to the file. The fleet cannot run unattended across that boundary.

This ticket makes the cycle a fleetd feature: the lead says when it is ready, fleetd checks the
handover file is really there, clears the pane, and starts the next lead on that file.

What the operator chose (2026-09-11)

Three decisions, taken by the operator after the measurements below:

  1. The lead asks for it. Not a timer, not an operator command. The lead is the only party
    that can see its own context filling.
  2. /clear in the same pane. Not a stop-and-relaunch.
  3. The operator confirms before the wipe, on top of the lead's own confirmation.

Measured facts that constrain the design

Every line below was measured on 2026-09-11 against 7b97aae.

  • fleetd cannot see how full a context is. There is no token measurement anywhere. The only
    context numbers are two operator-configured ceilings: autoCompactWindow
    (config/FleetConfig.java:487, a spawn flag) and Lifecycle.contextCap
    (config/FleetConfig.java:868), which counts turns, not tokens. fleetd reads no Claude Code
    transcript: jsonl has zero matches in the main sources (control: 4 files match
    CLAUDE_CONFIG_DIR, so the search works). This is why the trigger has to be the lead.
  • A lead has no record in SessionManager at all. lead/LeadLauncher.java never calls it,
    so the idle reaper (session/SessionManager.java:1002), the context cap (:965) and the
    shutdown drain (:1068) can none of them act on a lead. There is no per-lead lifecycle object
    to hang this on; it needs its own.
  • Lead identity survives a restart, because it is pinned to the tab, not the session.
    herdr/LeadTabScanner.java:247 matches the tab label, and auth/CallerResolver.java:211
    grants Role.PRIMARY from that map. terminal_id is expected to change on restart. A
    /clear does not even change the terminal, so identity is safe on this route.
  • fleet_stop cannot target a lead (mcp/FleetMcp.java:461 → SessionManager.stop, which
    reads a registry a lead is never in). So there is no accidental-teardown path today.
  • Every session timer still uses System.nanoTime, which stops while the Mac sleeps. Only
    health/FleetHealthMonitor.java:47 got a realtime clock, under #386. Any deadline in this
    feature must use a realtime clock
    , or it will freeze across host sleep.

Two live findings from a probe — these are the load-bearing ones

Probed on a throwaway sonnet member, using bridge tools only.

  • /clear sent over herdr agent.prompt really does run as a slash command. It is not an
    inert paste. Proof is in the transcripts, not in a reply: the member's session ends with its
    last answer at 06:05:04, and a new session is born at 06:05:04 whose line 5 is literally
    <command-name>/clear</command-name>. So member/ClaudeCodeLauncher.java:966 works, and the
    shipped clearAfterTurn feature is not silently dead.
  • A /clear routed through Injector wedges that pane for good. After the clear, every
    later message to that member vanished. fleetd read state: busy while herdr read
    liveStatus: done, and two tickets sat at [pending — worker done] forever. /clear
    produces no turn boundary, so the Injector's turn never completes and later sends queue behind
    it. The comment at member/ClaudeCodeLauncher.java:971 — "This deliberately bypasses
    Injector: /clear is housekeeping, not a delegated turn"
    — is load-bearing, not stylistic.

The flow

sequenceDiagram
    participant L as Lead session
    participant F as fleetd
    participant O as Operator
    L->>F: "fleet_handover{}" (I am filling up)
    F-->>L: path to write, what to include, a token
    L->>L: writes the handover file
    L->>O: asks for the go-ahead
    O-->>L: yes
    L->>F: "fleet_handover{token, ready, operatorConfirmed}"
    F->>F: file exists? non-empty? newer than the request?
    F->>L: "agents.send" -> "/clear" (direct, NOT via Injector)
    F->>F: wait until the pane is injectable again
    F->>L: "agents.send" -> "read <path> and carry on"

The two agents.send calls follow the two shipped patterns: clearContext
(member/ClaudeCodeLauncher.java:972) and the lead nudge (msg/LeadHeartbeatLoop.java:252).

Units

Unit A — config block plus the rollover executor

  • New opt-in top-level config block leadRollover:, following leadHeartbeat: exactly. With no
    block, nothing is ever constructed, so an upgrade cannot silently acquire this behaviour.
    Keys: handoverPath, requireOperatorConfirm (default true), maxDocAgeSeconds,
    bootstrapText, clearSettleSeconds.
  • A LeadRollover class with a pure decision function plus a thin executor, the way
    LeadHeartbeatLoop is built. It must:
    • refuse to roll unless the handover file exists, is non-empty, and its modified time is
      after the request — checked on a realtime clock, never System.nanoTime;
    • send /clear with agents.send(...) directly, never through Injector;
    • wait for the pane to report injectable again before the bootstrap prompt;
    • never roll on a timeout. Only an explicit confirmation rolls.
  • Register the key in KNOWN_TOP_LEVEL_KEYS, classify it in ConfigRef, document it in
    fleetd.example.yaml, and update the running tally in ConfigRef's class doc.

Unit B — the handover skill

.claude/skills/handover/SKILL.md: the lead-side playbook for writing a handover a stranger can
act on. The doc's quality decides whether the next lead succeeds, so this unit is not
decoration. It must require: every number carries the command that produced it; open decisions
and who owns them; live hazards; what is explicitly not owed. Model it on the existing
handover file, which is a good example of the shape.

Unit C — the fleet_handover MCP tool (after A merges)

Primary-only in auth/Authz. Two phases on one tool: an opening call that returns the path and
the token, and a confirming call that carries ready, operatorConfirmed and the token. A third
form cancels. Because it adds a fleet_* tool, CLAUDE.md's own rule applies: the primary's
intent-to-tool table must gain a row, and the canonical block must stay byte-identical with the
wiki template.

Acceptance criteria

  1. With no leadRollover: block, no rollover object is constructed and no behaviour changes.
  2. A roll never happens without an explicit confirmation. No timeout, and no missed heartbeat,
    can trigger the wipe. There must be a test that pins this.
  3. A missing, empty or stale handover file refuses the roll, and the refusal names which of the
    three checks failed.
  4. /clear and the bootstrap prompt both go through agents.send directly. A test must fail if
    either is routed through Injector.
  5. Every deadline uses a realtime clock. A test must fail if System.nanoTime is used for the
    file-freshness check.
  6. An end-to-end probe on a throwaway member — never on the real lead — showing that a
    freshly cleared pane picks up the next agents.send prompt. This is the one step still
    unproven; my probe could not test it because the Injector was already wedged.

Hazards

  • Do not test this on the live lead. A failed roll destroys the operator's session.
  • The operator's confirmation reaches fleetd as a flag the lead sets. fleetd's hard gates are
    the file checks; the operator flag is an assertion by the lead. Say it that way in the docs
    rather than implying fleetd verified it.
## Why A lead session fills up its context and then has to be replaced by hand. Today that means the operator notices, asks the lead to write a handover file, starts a new session, and pastes a pointer to the file. The fleet cannot run unattended across that boundary. This ticket makes the cycle a fleetd feature: the lead says when it is ready, fleetd checks the handover file is really there, clears the pane, and starts the next lead on that file. ## What the operator chose (2026-09-11) Three decisions, taken by the operator after the measurements below: 1. **The lead asks for it.** Not a timer, not an operator command. The lead is the only party that can see its own context filling. 2. **`/clear` in the same pane.** Not a stop-and-relaunch. 3. **The operator confirms before the wipe**, on top of the lead's own confirmation. ## Measured facts that constrain the design Every line below was measured on 2026-09-11 against `7b97aae`. - **fleetd cannot see how full a context is.** There is no token measurement anywhere. The only context numbers are two operator-configured ceilings: `autoCompactWindow` (`config/FleetConfig.java:487`, a spawn flag) and `Lifecycle.contextCap` (`config/FleetConfig.java:868`), which counts *turns*, not tokens. fleetd reads no Claude Code transcript: `jsonl` has zero matches in the main sources (control: 4 files match `CLAUDE_CONFIG_DIR`, so the search works). **This is why the trigger has to be the lead.** - **A lead has no record in `SessionManager` at all.** `lead/LeadLauncher.java` never calls it, so the idle reaper (`session/SessionManager.java:1002`), the context cap (`:965`) and the shutdown drain (`:1068`) can none of them act on a lead. There is no per-lead lifecycle object to hang this on; it needs its own. - **Lead identity survives a restart, because it is pinned to the tab, not the session.** `herdr/LeadTabScanner.java:247` matches the tab label, and `auth/CallerResolver.java:211` grants `Role.PRIMARY` from that map. `terminal_id` is *expected* to change on restart. A `/clear` does not even change the terminal, so identity is safe on this route. - **`fleet_stop` cannot target a lead** (`mcp/FleetMcp.java:461` → `SessionManager.stop`, which reads a registry a lead is never in). So there is no accidental-teardown path today. - **Every session timer still uses `System.nanoTime`**, which stops while the Mac sleeps. Only `health/FleetHealthMonitor.java:47` got a realtime clock, under #386. **Any deadline in this feature must use a realtime clock**, or it will freeze across host sleep. ## Two live findings from a probe — these are the load-bearing ones Probed on a throwaway sonnet member, using bridge tools only. - **`/clear` sent over herdr `agent.prompt` really does run as a slash command.** It is not an inert paste. Proof is in the transcripts, not in a reply: the member's session ends with its last answer at 06:05:04, and a new session is born at 06:05:04 whose line 5 is literally `<command-name>/clear</command-name>`. So `member/ClaudeCodeLauncher.java:966` works, and the shipped `clearAfterTurn` feature is not silently dead. - **A `/clear` routed through `Injector` wedges that pane for good.** After the clear, every later message to that member vanished. fleetd read `state: busy` while herdr read `liveStatus: done`, and two tickets sat at `[pending — worker done]` forever. `/clear` produces no turn boundary, so the Injector's turn never completes and later sends queue behind it. The comment at `member/ClaudeCodeLauncher.java:971` — *"This deliberately bypasses Injector: /clear is housekeeping, not a delegated turn"* — is load-bearing, not stylistic. ## The flow ```mermaid sequenceDiagram participant L as Lead session participant F as fleetd participant O as Operator L->>F: "fleet_handover{}" (I am filling up) F-->>L: path to write, what to include, a token L->>L: writes the handover file L->>O: asks for the go-ahead O-->>L: yes L->>F: "fleet_handover{token, ready, operatorConfirmed}" F->>F: file exists? non-empty? newer than the request? F->>L: "agents.send" -> "/clear" (direct, NOT via Injector) F->>F: wait until the pane is injectable again F->>L: "agents.send" -> "read <path> and carry on" ``` The two `agents.send` calls follow the two shipped patterns: `clearContext` (`member/ClaudeCodeLauncher.java:972`) and the lead nudge (`msg/LeadHeartbeatLoop.java:252`). ## Units ### Unit A — config block plus the rollover executor - New opt-in top-level config block `leadRollover:`, following `leadHeartbeat:` exactly. With no block, nothing is ever constructed, so an upgrade cannot silently acquire this behaviour. Keys: `handoverPath`, `requireOperatorConfirm` (default `true`), `maxDocAgeSeconds`, `bootstrapText`, `clearSettleSeconds`. - A `LeadRollover` class with a pure decision function plus a thin executor, the way `LeadHeartbeatLoop` is built. It must: - refuse to roll unless the handover file exists, is non-empty, and its modified time is **after** the request — checked on a **realtime clock**, never `System.nanoTime`; - send `/clear` with `agents.send(...)` **directly**, never through `Injector`; - wait for the pane to report injectable again before the bootstrap prompt; - never roll on a timeout. Only an explicit confirmation rolls. - Register the key in `KNOWN_TOP_LEVEL_KEYS`, classify it in `ConfigRef`, document it in `fleetd.example.yaml`, and update the running tally in `ConfigRef`'s class doc. ### Unit B — the `handover` skill `.claude/skills/handover/SKILL.md`: the lead-side playbook for writing a handover a stranger can act on. The doc's quality decides whether the next lead succeeds, so this unit is not decoration. It must require: every number carries the command that produced it; open decisions and who owns them; live hazards; what is explicitly *not* owed. Model it on the existing handover file, which is a good example of the shape. ### Unit C — the `fleet_handover` MCP tool (after A merges) Primary-only in `auth/Authz`. Two phases on one tool: an opening call that returns the path and the token, and a confirming call that carries `ready`, `operatorConfirmed` and the token. A third form cancels. Because it adds a `fleet_*` tool, `CLAUDE.md`'s own rule applies: the primary's intent-to-tool table must gain a row, and the canonical block must stay byte-identical with the wiki template. ## Acceptance criteria 1. With no `leadRollover:` block, no rollover object is constructed and no behaviour changes. 2. A roll never happens without an explicit confirmation. No timeout, and no missed heartbeat, can trigger the wipe. There must be a test that pins this. 3. A missing, empty or stale handover file refuses the roll, and the refusal names which of the three checks failed. 4. `/clear` and the bootstrap prompt both go through `agents.send` directly. A test must fail if either is routed through `Injector`. 5. Every deadline uses a realtime clock. A test must fail if `System.nanoTime` is used for the file-freshness check. 6. **An end-to-end probe on a throwaway member** — never on the real lead — showing that a freshly cleared pane picks up the next `agents.send` prompt. This is the one step still unproven; my probe could not test it because the Injector was already wedged. ## Hazards - Do not test this on the live lead. A failed roll destroys the operator's session. - The operator's confirmation reaches fleetd as a flag the lead sets. fleetd's hard gates are the file checks; the operator flag is an assertion by the lead. Say it that way in the docs rather than implying fleetd verified it.
Author
Owner

What the test suite pins about /clear, and what it does not

Checked while Unit A was being built, against 7b97aae.

member/ClaudeCodeLauncherTest.java:786 (clearContextUsesTheClaudeCommandThroughTheOwningHandle)
asserts that agent.prompt was called with text: "/clear" against a FakeHerdr. So it pins
the call. It cannot pin the effect, because no real Claude Code session is involved.

Control: 128 test files in the tree, and the grep finds every clearContext / clearAfterTurn
site, so the search discriminates. session/SessionManagerTest.java:831
(clearAfterTurnResetsContextWithoutDoubleCountingTheTurn) is the same shape — it counts
launcher.clearContextCalls(), again a fake.

So until 2026-09-11 nothing in this repo had evidence that /clear sent over agent.prompt
does anything at all.
The shipped clearAfterTurn feature rested on an untested assumption,
and it defaults to false, so nobody would have noticed had it been wrong. The live probe
recorded in the ticket body is the first evidence of the effect.

This does not close acceptance criterion 6. Two different things were unproven, and only one
is now settled:

Claim Status
/clear over agent.prompt executes as a slash command proven live, 2026-09-11
A freshly cleared pane picks up the next agents.send prompt still unproven

The second one is what criterion 6 asks for, and it is the step the whole feature ends on.

One reassuring detail found while checking, which is an argument but not a proof. In
clearAfterTurn the clear happens inside completeTurn
(session/SessionManager.java:969) — that is, after the turn boundary was already seen. So the
Injector is not wedged there, which is the opposite of what my probe did, where the /clear
was the turn. The rollover feature is on the safe side of that line too: the lead is never
tracked by the Injector at all. But this reasoning is not a substitute for the probe, and
criterion 6 stays open.

## What the test suite pins about `/clear`, and what it does not Checked while Unit A was being built, against `7b97aae`. `member/ClaudeCodeLauncherTest.java:786` (`clearContextUsesTheClaudeCommandThroughTheOwningHandle`) asserts that `agent.prompt` was called with `text: "/clear"` against a **`FakeHerdr`**. So it pins the *call*. It cannot pin the *effect*, because no real Claude Code session is involved. Control: 128 test files in the tree, and the grep finds every `clearContext` / `clearAfterTurn` site, so the search discriminates. `session/SessionManagerTest.java:831` (`clearAfterTurnResetsContextWithoutDoubleCountingTheTurn`) is the same shape — it counts `launcher.clearContextCalls()`, again a fake. **So until 2026-09-11 nothing in this repo had evidence that `/clear` sent over `agent.prompt` does anything at all.** The shipped `clearAfterTurn` feature rested on an untested assumption, and it defaults to `false`, so nobody would have noticed had it been wrong. The live probe recorded in the ticket body is the first evidence of the effect. **This does not close acceptance criterion 6.** Two different things were unproven, and only one is now settled: | Claim | Status | |---|---| | `/clear` over `agent.prompt` executes as a slash command | **proven live**, 2026-09-11 | | A *freshly cleared* pane picks up the next `agents.send` prompt | **still unproven** | The second one is what criterion 6 asks for, and it is the step the whole feature ends on. One reassuring detail found while checking, which is an argument but not a proof. In `clearAfterTurn` the clear happens inside `completeTurn` (`session/SessionManager.java:969`) — that is, *after* the turn boundary was already seen. So the `Injector` is not wedged there, which is the opposite of what my probe did, where the `/clear` **was** the turn. The rollover feature is on the safe side of that line too: the lead is never tracked by the `Injector` at all. But this reasoning is not a substitute for the probe, and criterion 6 stays open.
Author
Owner

Two corrections to the ticket body, one of them a defect in my own brief

1. The roll order in the ticket body is wrong — this is my mistake, not a worker's

The body says the roll is: send /clear → wait until injectable → send the bootstrap text.

That first step can land in the middle of the lead's own turn. confirm() is called by the
lead
, from inside a turn. At the moment it returns, the lead's pane is WORKING, not idle. Every
other injection path in this daemon refuses to touch a WORKING pane on purpose —
msg/LeadHeartbeatLoop.java calls it constraint 2, and ReplyPushLoop.decide checks
status.injectable() before it sends. My ordering skips that gate for the one send that destroys
context.

Correct order — three waits, not two:

  1. wait until the lead's pane reports injectable (the lead has finished the turn it asked from);
  2. send /clear with agents.send(...), directly;
  3. wait until the pane reports injectable again;
  4. send the bootstrap text with agents.send(...), directly.

So confirm() must record the decision and return, and let a later tick do the roll. It must
not roll inline. That also keeps confirm() fast, instead of blocking the lead's MCP call for as
long as the settle window.

clearSettleSeconds therefore bounds step 3. Step 1 needs its own bound — call it
turnSettleSeconds, same default of 20. If step 1 times out, refuse and roll nothing: a lead that
never goes idle is a lead that is still working.

2. A freshly cleared session does reload CLAUDE.md and the memory index

Measured on this host, 2026-09-11, from the lead's own transcript. The current lead session was
itself created by a /clear (its line 5 is <command-name>/clear</command-name>), and line 27
of that same session carries both the CLAUDE.md canonical block and the MEMORY.md index.

This closes a risk that was open: will the new lead know it is a lead? Yes. CLAUDE.md is
reloaded, and CLAUDE.md already tells every session to settle its role with fleet_whoami
before acting.

Consequence for the design: bootstrapText can be short. It does not need to re-explain the
role, the bridge, or the charter — all of that arrives on its own. It needs to do one thing: name
the handover file and say to read it first. A long bootstrap prompt would be duplicated
instruction surface, and this repo already has a rule against that.

## Two corrections to the ticket body, one of them a defect in my own brief ### 1. The roll order in the ticket body is wrong — this is my mistake, not a worker's The body says the roll is: send `/clear` → wait until injectable → send the bootstrap text. **That first step can land in the middle of the lead's own turn.** `confirm()` is called *by the lead*, from inside a turn. At the moment it returns, the lead's pane is `WORKING`, not idle. Every other injection path in this daemon refuses to touch a `WORKING` pane on purpose — `msg/LeadHeartbeatLoop.java` calls it constraint 2, and `ReplyPushLoop.decide` checks `status.injectable()` before it sends. My ordering skips that gate for the one send that destroys context. **Correct order — three waits, not two:** 1. wait until the lead's pane reports **injectable** (the lead has finished the turn it asked from); 2. send `/clear` with `agents.send(...)`, directly; 3. wait until the pane reports injectable **again**; 4. send the bootstrap text with `agents.send(...)`, directly. So `confirm()` must **record the decision and return**, and let a later tick do the roll. It must not roll inline. That also keeps `confirm()` fast, instead of blocking the lead's MCP call for as long as the settle window. `clearSettleSeconds` therefore bounds step 3. Step 1 needs its own bound — call it `turnSettleSeconds`, same default of 20. If step 1 times out, refuse and roll nothing: a lead that never goes idle is a lead that is still working. ### 2. A freshly cleared session does reload `CLAUDE.md` and the memory index Measured on this host, 2026-09-11, from the lead's own transcript. The current lead session was itself created by a `/clear` (its line 5 is `<command-name>/clear</command-name>`), and **line 27 of that same session carries both the `CLAUDE.md` canonical block and the `MEMORY.md` index.** This closes a risk that was open: *will the new lead know it is a lead?* Yes. `CLAUDE.md` is reloaded, and `CLAUDE.md` already tells every session to settle its role with `fleet_whoami` before acting. **Consequence for the design:** `bootstrapText` can be short. It does not need to re-explain the role, the bridge, or the charter — all of that arrives on its own. It needs to do one thing: name the handover file and say to read it first. A long bootstrap prompt would be duplicated instruction surface, and this repo already has a rule against that.
Author
Owner

Two more design points, found by re-reading the flow

1. The roll must target the CALLING lead, not "the primary" — this is invariant 3

msg/LeadHeartbeatLoop.java:245 resolves its target with primaryRegistry.primaryTerminal().
That is right for a background loop, which has no caller to speak of.

It is wrong here. confirm() has a caller, and CLAUDE.md's invariant 3 is explicit: identity
comes from the connection, never an argument.
So LeadRollover must roll the terminal the
call arrived on
, resolved the way auth/CallerResolver.java:208 already resolves it — not a
registry lookup that answers "who is the primary".

Why it matters in practice: this daemon can hold more than one labelled lead tab. fleetd #359
exists because stale labels used to pile up, and lead/LeadLauncher.java:132 still carries the
two-reading cleanup for exactly that. A lookup-based target means one lead can call confirm()
and a different lead's pane gets cleared. That is an unrecoverable loss of someone else's
context, caused by a lookup that looked harmless.

There is no matching risk in the other direction: a caller-derived target can only ever clear the
pane that asked to be cleared.

2. open() should hand back the fleet state the handover needs to mention

The handover file has to record in-flight work, or the new lead will not know it exists. The
handover skill (merged in #481) requires a "live hazards" section for this, but the outgoing
lead has to gather the facts by hand.

fleetd already has them. msg/LeadHeartbeatLoop.java's FleetState record carries exactly the
right four: undrained worker replies, sessions in DONE awaiting teardown, live workers, and
which terminals hold replies. open() should return that snapshot alongside the path and the
token.

This is cheap — the record exists and is already computed for the heartbeat — and it removes the
most likely gap in a handover file: a member still mid-turn that nobody wrote down.

Not a refusal. open() should report the state, not block on it. Whether to roll with
workers still live is the lead's judgement, and sometimes the right answer is yes, because the
replies stay in the durable inbox for the next lead to drain. Refusing would make the feature
unusable exactly when a long session most needs it.

## Two more design points, found by re-reading the flow ### 1. The roll must target the CALLING lead, not "the primary" — this is invariant 3 `msg/LeadHeartbeatLoop.java:245` resolves its target with `primaryRegistry.primaryTerminal()`. That is right **for a background loop**, which has no caller to speak of. It is wrong here. `confirm()` has a caller, and `CLAUDE.md`'s invariant 3 is explicit: *identity comes from the connection, never an argument.* So `LeadRollover` must roll **the terminal the call arrived on**, resolved the way `auth/CallerResolver.java:208` already resolves it — not a registry lookup that answers "who is the primary". Why it matters in practice: this daemon can hold more than one labelled lead tab. fleetd #359 exists because stale labels used to pile up, and `lead/LeadLauncher.java:132` still carries the two-reading cleanup for exactly that. A lookup-based target means one lead can call `confirm()` and **a different lead's pane gets cleared**. That is an unrecoverable loss of someone else's context, caused by a lookup that looked harmless. There is no matching risk in the other direction: a caller-derived target can only ever clear the pane that asked to be cleared. ### 2. `open()` should hand back the fleet state the handover needs to mention The handover file has to record in-flight work, or the new lead will not know it exists. The `handover` skill (merged in #481) requires a "live hazards" section for this, but the outgoing lead has to gather the facts by hand. fleetd already has them. `msg/LeadHeartbeatLoop.java`'s `FleetState` record carries exactly the right four: undrained worker replies, sessions in `DONE` awaiting teardown, live workers, and which terminals hold replies. `open()` should return that snapshot alongside the path and the token. This is cheap — the record exists and is already computed for the heartbeat — and it removes the most likely gap in a handover file: a member still mid-turn that nobody wrote down. **Not a refusal.** `open()` should report the state, not block on it. Whether to roll with workers still live is the lead's judgement, and sometimes the right answer is yes, because the replies stay in the durable inbox for the next lead to drain. Refusing would make the feature unusable exactly when a long session most needs it.
Author
Owner

Lead status, 2026-09-11 — Unit A merged, two units in flight, one criterion I cannot prove

Unit A merged (PR #483, merge commit 4bab23e)

leadRollover: config block + dev.ltms.fleet.lead.LeadRollover. Nothing calls it yet.
CI run 1722 green on head a942622. Full suite on merged main: Tests run: 1651, Failures: 0, Errors: 0, Skipped: 0, BUILD SUCCESS (run without clean, to protect the running daemon's jar
— see #413).

Two corrections were applied before the merge, and the PR body still describes the pre-correction
shape:

  1. confirm() cannot roll inline. It is called BY the lead FROM the lead's own turn, so the
    pane is WORKING and cannot report settled until confirm() returns. The original code sent
    /clear and then timed out waiting — destroying the context and starting no fresh session,
    while returning a refusal that claimed nothing had happened. Now confirm() validates, hands a
    one-shot continuation to continuationRunner, and returns. The continuation waits for the
    calling turn to settle first (new knob turnSettleSeconds); if that wait times out it sends
    no /clear at all.
  2. Identity comes from the caller, not a lookup. PrimaryRegistry#primaryTerminal() is a
    single slot, and this daemon can hold more than one labelled lead tab, so lead X's confirm()
    could clear lead Y's pane. open() records the caller's terminal; confirm() refuses with
    NOT_YOUR_ROLLOVER on a mismatch.

Mutation battery on the two safety-critical branches, in a scratch worktree at a942622. Controls
measured against the original first; each mutant proven applied with two unrelated proofs using
different search strings.

  • if (!turnSettled) → if (false): turnThatNeverSettlesSendsNoClearAtAll FAILS
    (LeadRolloverTest:247, expected 0 sends, was 1).
  • ownership check → if (false): aDifferentLeadTerminalCannotConfirmAnotherLeadsRollover
    FAILS (LeadRolloverTest:266).
  • Unmutated harness-proof run: exit 0, with both guard WARN lines logged, so the branches really
    are exercised.

New defect found after the merge — Unit E in flight

waitUntilInjectable uses AgentStatus#injectable(), which is IDLE || BLOCKED || DONE. That is
the right rule for Injector ("may I deliver a message"), but the wrong rule for this gate ("has
the turn actually ended"), because BLOCKED is a live turn that is paused — a pane sitting on
an approval prompt.

Path in: the lead calls confirm(); the turn carries on and hits anything needing approval; herdr
reports blocked; within 250ms the continuation reads that as settled; /clear is typed into an
open prompt. The lead's live context is destroyed mid-turn — exactly what correction 1 exists to
prevent. Window is turnSettleSeconds, default 20s.

The mutation battery could not have found this: it proved the guard fires, not that the guard's
predicate was tight enough. LeadRolloverTest drives the fake status with only "idle"
(:198) and "working" (:229); neither value discriminates this defect, only "blocked" does.

Unit E makes both waits require a real turn boundary (IDLE or DONE), local to LeadRollover.
AgentStatus#injectable() itself is left alone — it is correct for Injector and
LeadHeartbeatLoop.

Unit C in flight — the fleet_handover MCP tool

open/confirm/cancel, primary-only in Authz, caller terminal resolved from the connection
with no terminal parameter in the input schema. Registered unconditionally and answering a
clean NOT_CONFIGURED refusal when leadRollover: is absent, so the registered tool surface does
not vary with config — otherwise it collides with #474's charter tool-surface gate and
McpContractDocTest.

Acceptance criterion 6 — I cannot prove this, and I am not going to pretend otherwise

"A freshly cleared pane picks up the next agents.send prompt" is still unproven. The live
probe recorded earlier in this ticket proved /clear over agent.prompt executes as a real slash
command, but it could not get past that, because routing /clear through Injector wedged the
target permanently.

There is no route left that would prove it:

  • No shipped MCP tool or REST route does a non-Injector agents.send. The registered routes are
    /agents, /healthz, /mcp, /member-credentials, /members, /members/{paneId},
    /metrics, /profiles, /sessions, /sessions/{id}/{ask,message,replies,reply,status},
    /tasks/{ticket}.
  • clearAfterTurn is the one shipped feature that does this, but turning it on needs a
    fleetd.yaml edit, which the classifier denied in an earlier session. Not routing around it.
  • Driving herdr directly is banned by charter invariant 5, and the repo's carve-out covers only
    four named contract tests.

LeadRollover operates on the lead's own pane, resolved from the connection, and the tool is
primary-only — so it cannot be exercised on a throwaway member at all. The only honest proof is to
use the feature for its real purpose, once, on a live lead that has already written its handover
file. That needs the operator's say-so, which the design already requires
(requireOperatorConfirm, default true).

Until that happens, the residual risk is bounded and worth stating: if the bootstrap prompt
does not land, the lead's context is gone and no fresh session starts. The handover file is
written before confirm() is even accepted, so the loss is recoverable by hand — the operator
starts a session and points it at the file. That is the same manual path the handover skill
describes today.

Also verified

The lead's own pane status is readable through the daemon —
GET /sessions/term_…/status → HTTP 200, {"status":"working","ready":true}, read while this
lead was mid-turn. So the premise waitUntilInjectable rests on holds, and the working reading
during a live turn is direct evidence for correction 1.

Not done

  • Wiki Features entry for leadRollover: — waiting until the tool lands so it describes a
    capability an operator can actually use.
  • CLAUDE.md intent-to-tool row for fleet_handover, kept byte-identical with
    wiki/7-Use-Cases.md.
  • The handover skill still says fleet_handover does not exist. True today; must change when
    Unit C merges.
  • Redeploy. A merge is not a deployment, and there is nothing worth deploying until the tool lands.
## Lead status, 2026-09-11 — Unit A merged, two units in flight, one criterion I cannot prove ### Unit A merged (PR #483, merge commit `4bab23e`) `leadRollover:` config block + `dev.ltms.fleet.lead.LeadRollover`. Nothing calls it yet. CI run 1722 green on head `a942622`. Full suite on merged `main`: `Tests run: 1651, Failures: 0, Errors: 0, Skipped: 0`, `BUILD SUCCESS` (run without `clean`, to protect the running daemon's jar — see #413). Two corrections were applied before the merge, and the PR body still describes the pre-correction shape: 1. **`confirm()` cannot roll inline.** It is called BY the lead FROM the lead's own turn, so the pane is `WORKING` and cannot report settled until `confirm()` returns. The original code sent `/clear` and then timed out waiting — destroying the context and starting no fresh session, while returning a refusal that claimed nothing had happened. Now `confirm()` validates, hands a one-shot continuation to `continuationRunner`, and returns. The continuation waits for the calling turn to settle first (new knob `turnSettleSeconds`); if that wait times out it sends **no** `/clear` at all. 2. **Identity comes from the caller, not a lookup.** `PrimaryRegistry#primaryTerminal()` is a single slot, and this daemon can hold more than one labelled lead tab, so lead X's `confirm()` could clear lead Y's pane. `open()` records the caller's terminal; `confirm()` refuses with `NOT_YOUR_ROLLOVER` on a mismatch. Mutation battery on the two safety-critical branches, in a scratch worktree at `a942622`. Controls measured against the original first; each mutant proven applied with two unrelated proofs using different search strings. - `if (!turnSettled)` → `if (false)`: `turnThatNeverSettlesSendsNoClearAtAll` **FAILS** (`LeadRolloverTest:247`, expected 0 sends, was 1). - ownership check → `if (false)`: `aDifferentLeadTerminalCannotConfirmAnotherLeadsRollover` **FAILS** (`LeadRolloverTest:266`). - Unmutated harness-proof run: exit 0, with both guard `WARN` lines logged, so the branches really are exercised. ### New defect found after the merge — Unit E in flight `waitUntilInjectable` uses `AgentStatus#injectable()`, which is `IDLE || BLOCKED || DONE`. That is the right rule for `Injector` ("may I deliver a message"), but the wrong rule for this gate ("has the turn actually ended"), because **`BLOCKED` is a live turn that is paused** — a pane sitting on an approval prompt. Path in: the lead calls `confirm()`; the turn carries on and hits anything needing approval; herdr reports `blocked`; within 250ms the continuation reads that as settled; `/clear` is typed into an open prompt. The lead's live context is destroyed mid-turn — exactly what correction 1 exists to prevent. Window is `turnSettleSeconds`, default 20s. The mutation battery could not have found this: it proved the guard *fires*, not that the guard's predicate was tight enough. `LeadRolloverTest` drives the fake status with only `"idle"` (`:198`) and `"working"` (`:229`); neither value discriminates this defect, only `"blocked"` does. Unit E makes both waits require a real turn boundary (`IDLE` or `DONE`), local to `LeadRollover`. `AgentStatus#injectable()` itself is left alone — it is correct for `Injector` and `LeadHeartbeatLoop`. ### Unit C in flight — the `fleet_handover` MCP tool `open`/`confirm`/`cancel`, primary-only in `Authz`, caller terminal resolved from the connection with **no** terminal parameter in the input schema. Registered **unconditionally** and answering a clean `NOT_CONFIGURED` refusal when `leadRollover:` is absent, so the registered tool surface does not vary with config — otherwise it collides with #474's charter tool-surface gate and `McpContractDocTest`. ### Acceptance criterion 6 — I cannot prove this, and I am not going to pretend otherwise "A freshly cleared pane picks up the next `agents.send` prompt" is still **unproven**. The live probe recorded earlier in this ticket proved `/clear` over `agent.prompt` executes as a real slash command, but it could not get past that, because routing `/clear` through `Injector` wedged the target permanently. There is no route left that would prove it: - No shipped MCP tool or REST route does a non-`Injector` `agents.send`. The registered routes are `/agents`, `/healthz`, `/mcp`, `/member-credentials`, `/members`, `/members/{paneId}`, `/metrics`, `/profiles`, `/sessions`, `/sessions/{id}/{ask,message,replies,reply,status}`, `/tasks/{ticket}`. - `clearAfterTurn` is the one shipped feature that does this, but turning it on needs a `fleetd.yaml` edit, which the classifier denied in an earlier session. Not routing around it. - Driving herdr directly is banned by charter invariant 5, and the repo's carve-out covers only four named contract tests. `LeadRollover` operates on the **lead's own** pane, resolved from the connection, and the tool is primary-only — so it cannot be exercised on a throwaway member at all. The only honest proof is to use the feature for its real purpose, once, on a live lead that has already written its handover file. That needs the operator's say-so, which the design already requires (`requireOperatorConfirm`, default true). **Until that happens, the residual risk is bounded and worth stating:** if the bootstrap prompt does not land, the lead's context is gone and no fresh session starts. The handover file is written before `confirm()` is even accepted, so the loss is recoverable by hand — the operator starts a session and points it at the file. That is the same manual path the `handover` skill describes today. ### Also verified The lead's own pane status is readable through the daemon — `GET /sessions/term_…/status` → HTTP 200, `{"status":"working","ready":true}`, read while this lead was mid-turn. So the premise `waitUntilInjectable` rests on holds, and the `working` reading during a live turn is direct evidence for correction 1. ### Not done - Wiki Features entry for `leadRollover:` — waiting until the tool lands so it describes a capability an operator can actually use. - `CLAUDE.md` intent-to-tool row for `fleet_handover`, kept byte-identical with `wiki/7-Use-Cases.md`. - The `handover` skill still says `fleet_handover` does not exist. True today; must change when Unit C merges. - Redeploy. A merge is not a deployment, and there is nothing worth deploying until the tool lands.
Author
Owner

Shipped and deployed — 2026-09-11

All three units merged, docs updated, daemon redeployed, tool proven reachable on the live daemon.

PR What Verified by
#483 leadRollover: config + LeadRollover executor 2 mutations, both caught
#484 BLOCKED is not a settled pane mutation on the half the worker did not touch
#485 fleet_handover MCP tool mutation on the wiring — found a real gap, sent back, re-verified

main at 008a457. Full suite on merged main: Tests run: 1662, Failures: 0, Errors: 0, Skipped: 0.

The gap in #485 that the first pass shipped

Worth recording because a green suite vouched for it. The unit first landed with a defaulted
15-argument FleetMcp constructor delegating to the new 16-argument one with leadRollover = null.
I mutated the wiring instead of reasoning about it — deleted just the leadRollover argument from
Fleetd.main's FleetMcp call:

compile:  0 errors
mvn test: BUILD SUCCESS 1, BUILD FAILURE 0
          Tests run: 1659, Failures: 0, Errors: 0, Skipped: 0

…while the live daemon would have answered NOT_CONFIGURED to every fleet_handover call for ever.
FleetMcpHandoverTest could not see it (it builds its own FleetMcp), and
FleetdLeadRolloverWiringTest could not either (it pins that LeadRollover is constructed, not
that it is passed on).

The defect was wider than I named. I pointed at one overload; the worker found 11/12/13/14/15-arg
constructors forming a single defaulting chain into the 16-arg one, each silently supplying another
feature's "off" value — leadChannel, outage, leadSeats, peers, then leadRollover. All five
deleted. FleetMcp now has exactly one public constructor, so all of those features are
compile-enforced at their call sites, not just this one. Re-running the identical mutation now gives
constructor FleetMcp cannot be applied to given types … argument lists differ in length.

Deployed

scripts/redeploy-fleetd.sh --yes
  build   Tests run: 1662, Failures: 0 — BUILD SUCCESS
  jar     82ebc1cc7047 -> e031c0f10170
  stop    pid 52482 exited (launchctl unload, not kill)
  start   pid 1907
  verify  /healthz 200, fresh "fleetd listening" line, no ERROR lines since restart
fleet_whoami -> primary

Live probe of the tool itself, on the running daemon:

fleet_handover{action: "open", reason: "post-deploy probe"}
-> {"accepted":false,"reason":"NOT_CONFIGURED","detail":"leadRollover: is not configured"}

Clean structured refusal, no exception — and the live wire schema carries exactly four parameters
(action, reason, token, operatorConfirmed) with no terminal, session or leadTerminal of
any kind
, confirmed against the registered schema rather than the source.

Docs

  • CLAUDE.md + wiki/7-Use-Cases.md: one fleet_handover intent-to-tool row, byte-identical
    (sync check prints True). The row carries the ordering trap, not just the call.
  • wiki/11-Features.md: full entry — knob, defaults, the three design decisions and why, six
    gotchas. Pushed and verified by ref (1eadfbc), since HEAD:master is a silent no-op here.
  • .claude/skills/handover/SKILL.md: it said "do not call a fleet_handover tool: it does not
    exist", which was true this morning and false by this afternoon. Rewritten, plus a new section 11
    with the three-step order and what surprises a caller.

What the operator still has to do

The feature is deployed but inert until leadRollover: is added to fleetd.yaml, which I am
not permitted to edit. Minimum block:

leadRollover:
  handoverPath: /Users/dai.ha/LTMS/.bridged-handover/HANDOVER.md

Everything else defaults: requireOperatorConfirm: true, maxDocAgeSeconds: 3600,
turnSettleSeconds: 20, clearSettleSeconds: 20, and a bootstrapText naming handoverPath.
Adding the block needs a daemon restart — the fields are hot, but the executor's construction is
gated on the block being present in the startup snapshot.

Still open

  • Acceptance criterion 6 remains unproven and is unprovable with the shipped surfaces — see my
    previous comment for every route I checked. The only honest proof is running the feature once, for
    real, on a live lead that has already written its handover file. That is the operator's call.
  • #486 — the settle poll is bounded only by an injected clock, so a regression in this class
    surfaces as a CI hang rather than a named red test. Not a production bug.
## Shipped and deployed — 2026-09-11 All three units merged, docs updated, daemon redeployed, tool proven reachable on the live daemon. | PR | What | Verified by | |---|---|---| | #483 | `leadRollover:` config + `LeadRollover` executor | 2 mutations, both caught | | #484 | `BLOCKED` is not a settled pane | mutation on the half the worker did not touch | | #485 | `fleet_handover` MCP tool | mutation on the wiring — found a real gap, sent back, re-verified | `main` at `008a457`. Full suite on merged main: `Tests run: 1662, Failures: 0, Errors: 0, Skipped: 0`. ### The gap in #485 that the first pass shipped Worth recording because a green suite vouched for it. The unit first landed with a defaulted 15-argument `FleetMcp` constructor delegating to the new 16-argument one with `leadRollover = null`. I mutated the wiring instead of reasoning about it — deleted just the `leadRollover` argument from `Fleetd.main`'s `FleetMcp` call: ``` compile: 0 errors mvn test: BUILD SUCCESS 1, BUILD FAILURE 0 Tests run: 1659, Failures: 0, Errors: 0, Skipped: 0 ``` …while the live daemon would have answered `NOT_CONFIGURED` to every `fleet_handover` call for ever. `FleetMcpHandoverTest` could not see it (it builds its own `FleetMcp`), and `FleetdLeadRolloverWiringTest` could not either (it pins that `LeadRollover` is *constructed*, not that it is *passed on*). **The defect was wider than I named.** I pointed at one overload; the worker found 11/12/13/14/15-arg constructors forming a single defaulting chain into the 16-arg one, each silently supplying another feature's "off" value — `leadChannel`, `outage`, `leadSeats`, `peers`, then `leadRollover`. All five deleted. `FleetMcp` now has exactly one public constructor, so all of those features are compile-enforced at their call sites, not just this one. Re-running the identical mutation now gives `constructor FleetMcp cannot be applied to given types … argument lists differ in length`. ### Deployed ``` scripts/redeploy-fleetd.sh --yes build Tests run: 1662, Failures: 0 — BUILD SUCCESS jar 82ebc1cc7047 -> e031c0f10170 stop pid 52482 exited (launchctl unload, not kill) start pid 1907 verify /healthz 200, fresh "fleetd listening" line, no ERROR lines since restart fleet_whoami -> primary ``` Live probe of the tool itself, on the running daemon: ``` fleet_handover{action: "open", reason: "post-deploy probe"} -> {"accepted":false,"reason":"NOT_CONFIGURED","detail":"leadRollover: is not configured"} ``` Clean structured refusal, no exception — and the live wire schema carries exactly four parameters (`action`, `reason`, `token`, `operatorConfirmed`) with **no terminal, session or leadTerminal of any kind**, confirmed against the registered schema rather than the source. ### Docs - `CLAUDE.md` + `wiki/7-Use-Cases.md`: one `fleet_handover` intent-to-tool row, byte-identical (sync check prints `True`). The row carries the ordering trap, not just the call. - `wiki/11-Features.md`: full entry — knob, defaults, the three design decisions and why, six gotchas. Pushed and verified by ref (`1eadfbc`), since `HEAD:master` is a silent no-op here. - `.claude/skills/handover/SKILL.md`: it said "do not call a `fleet_handover` tool: it does not exist", which was true this morning and false by this afternoon. Rewritten, plus a new section 11 with the three-step order and what surprises a caller. ### What the operator still has to do The feature is deployed but **inert until `leadRollover:` is added to `fleetd.yaml`**, which I am not permitted to edit. Minimum block: ```yaml leadRollover: handoverPath: /Users/dai.ha/LTMS/.bridged-handover/HANDOVER.md ``` Everything else defaults: `requireOperatorConfirm: true`, `maxDocAgeSeconds: 3600`, `turnSettleSeconds: 20`, `clearSettleSeconds: 20`, and a `bootstrapText` naming `handoverPath`. **Adding the block needs a daemon restart** — the fields are hot, but the executor's construction is gated on the block being present in the startup snapshot. ### Still open - **Acceptance criterion 6 remains unproven** and is unprovable with the shipped surfaces — see my previous comment for every route I checked. The only honest proof is running the feature once, for real, on a live lead that has already written its handover file. That is the operator's call. - **#486** — the settle poll is bounded only by an injected clock, so a regression in this class surfaces as a CI hang rather than a named red test. Not a production bug.
Author
Owner

leadRollover: is now configured and LIVE on the Mac

The operator authorised the fleetd.yaml edit, with one requirement: use a relative path to a file in the workspace, not an absolute one. That turned out to need code, because a relative path could not have worked at all.

The defect a relative path exposed

Measured today:

  • daemon working directory: /Users/dai.ha/LTMS/claude-bridge/fleetd (lsof -a -p 1907 -d cwd)
  • lead pane working directory: /Users/dai.ha/LTMS/claude-bridge (fleet.leaders.opus.cwd)

They differ by one level, because this repo nests fleetd/ inside its own root. LeadRollover.java:218 stored the raw configured string and :321 called Path.of(...) on it. Three readers consume that one string, in two different directories: the daemon stats the file, the lead is handed it in the open response and writes there, and the fresh session is handed it inside bootstrapText. A relative path would have failed with HANDOVER_NOT_FOUND every time.

Fixed in #487 (merged as 3bf3968, main at 7f9137f)

open() now resolves the path to an absolute path exactly once, against the calling lead's configured cwd, falling back to System.getProperty("user.dir") — the same fallback LeadLauncher.java:311 already uses. PendingRollover.handoverPath() is absolute from then on, so all three readers agree. bootstrapText is no longer defaulted in the config record (it would bake in the raw relative path); bootstrapTextFor(resolvedPath) builds it at send time.

Two defects found during verification, both already fixed

1. The production wiring was uncovered. I mutated the lookup lambda inside Fleetd.leadRollover(...) to always return null — which forces every relative path back onto the daemon's own directory, i.e. this exact bug:

String leadName = null; // MUTANT_M3_WIRING_ALWAYS_NULL

Proven applied with two greps using different search strings (original gone: 0; mutant present: 1). Result: Tests run: 1669, Failures: 0, Errors: 0, Skipped: 0 / BUILD SUCCESS. The whole suite vouched for dead wiring.

This is worth naming precisely, because it is a new wrinkle on the #415 antidote. Making leadWorkspace a required constructor parameter did work — it made every call site a compile error, and FleetdLeadRolloverWiringTest's source-text pin catches leads being dropped from the argument list. But required only proves an argument is passed. It says nothing about whether the argument is correct, and here the argument is a lambda built inside the factory, so its body was untested surface that the antidote does not reach.

Fixed by FleetdLeadRolloverWorkspaceLookupTest — a behavioural test that calls the package-private Fleetd.leadRollover(...) directly with a ConfigRef from a temp fleetd.yaml. I re-ran the identical mutation against the merged tree myself: Tests run: 1673, Failures: 2, both failures printing /Users/dai.ha/LTMS/claude-bridge/fleetd/handover.md as the wrong answer. One of its three cases pins that the lookup is read live: the map is empty when the factory runs and only gains the terminal afterwards, which is the real startup order, since leads are found by a tab scan after boot.

FleetdLeadRolloverWiringTest's javadoc claimed "no behavioural test can catch this wiring dropping out." True before the factory took a lookup, false after. Corrected — a wrong claim inside a test is how the next person decides not to write the missing test.

2. A relative fleet.leaders.<name>.cwd broke the absolute-path guarantee. Found by a reviewer (terra, deliberately not the sonnet that wrote the diff). base.resolve(path).normalize() stays relative if the base is relative, silently reopening the same ambiguity — and falsifying the "always absolute" promise the new PendingRollover javadoc makes. Now .toAbsolutePath() on both branches. I mutated it back out: Tests run: 1673, Failures: 1, killed.

Filed #488 for the underlying config gap: nothing validates Leader.cwd() at load, so a relative cwd is silently accepted.

Deployed and proven live

Build Tests run: 1673, Failures: 0, Errors: 0, Skipped: 0. Redeploy: jar 2520c3c051c7 → 0faa26d1bb10, pid 1907 → 62179, /healthz 200, fresh fleetd listening line, no ERROR lines since restart, fleet_whoami → primary.

Config now in fleetd.yaml:

leadRollover:
  handoverPath: .handover/HANDOVER.md

Live probe on the running daemon — fleet_handover{action:"open"}:

{"token":"56280346-...","handoverPath":"/Users/dai.ha/LTMS/claude-bridge/.handover/HANDOVER.md","requestedAtMillis":1789166971737}

and the daemon's own log line:

lead-rollover: open token=56280346-... lead=term_65a4c36d11af27
  configuredHandoverPath=.handover/HANDOVER.md
  resolvedHandoverPath=/Users/dai.ha/LTMS/claude-bridge/.handover/HANDOVER.md

That is the relative value resolving against the lead's workspace, not the daemon's — through the real wiring, on the real daemon. Token cancelled afterwards; nothing was rolled.

.handover/ is gitignored (the file is a snapshot of live state). Docs updated: wiki/11-Features.md (pushed, verified by ref 285b0a6) and the handover skill, which now says the returned path is always absolute and must not be re-resolved by the lead. The CLAUDE.md charter row needed no change; sync check still prints in sync: True.

Still open

Acceptance criterion 6 is still not met — nothing has yet shown end-to-end that a freshly cleared pane picks up the bootstrap prompt. It is now reachable for the first time, since the feature is configured, but the only honest proof is running a real rollover on a live lead, and that is the operator's call. If the bootstrap does not land, the lead's context is gone and no fresh session starts; the recovery is the manual path, which is exactly why the handover file must exist before confirm is accepted.

#486 (unbounded settle poll hangs CI instead of failing) remains open and unassigned.

## `leadRollover:` is now configured and LIVE on the Mac The operator authorised the `fleetd.yaml` edit, with one requirement: **use a relative path to a file in the workspace**, not an absolute one. That turned out to need code, because a relative path could not have worked at all. ### The defect a relative path exposed Measured today: - daemon working directory: `/Users/dai.ha/LTMS/claude-bridge/fleetd` (`lsof -a -p 1907 -d cwd`) - lead pane working directory: `/Users/dai.ha/LTMS/claude-bridge` (`fleet.leaders.opus.cwd`) They differ by one level, because this repo nests `fleetd/` inside its own root. `LeadRollover.java:218` stored the raw configured string and `:321` called `Path.of(...)` on it. **Three** readers consume that one string, in two different directories: the daemon stats the file, the lead is handed it in the `open` response and writes there, and the fresh session is handed it inside `bootstrapText`. A relative path would have failed with `HANDOVER_NOT_FOUND` every time. ### Fixed in #487 (merged as `3bf3968`, main at `7f9137f`) `open()` now resolves the path to an absolute path exactly once, against the **calling lead's** configured `cwd`, falling back to `System.getProperty("user.dir")` — the same fallback `LeadLauncher.java:311` already uses. `PendingRollover.handoverPath()` is absolute from then on, so all three readers agree. `bootstrapText` is no longer defaulted in the config record (it would bake in the raw relative path); `bootstrapTextFor(resolvedPath)` builds it at send time. ### Two defects found during verification, both already fixed **1. The production wiring was uncovered.** I mutated the lookup lambda inside `Fleetd.leadRollover(...)` to always return `null` — which forces every relative path back onto the daemon's own directory, i.e. this exact bug: ```java String leadName = null; // MUTANT_M3_WIRING_ALWAYS_NULL ``` Proven applied with two greps using different search strings (original gone: 0; mutant present: 1). Result: **`Tests run: 1669, Failures: 0, Errors: 0, Skipped: 0` / `BUILD SUCCESS`.** The whole suite vouched for dead wiring. This is worth naming precisely, because it is a **new wrinkle on the #415 antidote**. Making `leadWorkspace` a required constructor parameter did work — it made every call site a compile error, and `FleetdLeadRolloverWiringTest`'s source-text pin catches `leads` being dropped from the argument list. But *required* only proves an argument is **passed**. It says nothing about whether the argument is **correct**, and here the argument is a lambda built inside the factory, so its body was untested surface that the antidote does not reach. Fixed by `FleetdLeadRolloverWorkspaceLookupTest` — a behavioural test that calls the package-private `Fleetd.leadRollover(...)` directly with a `ConfigRef` from a temp `fleetd.yaml`. I re-ran the identical mutation against the merged tree myself: **`Tests run: 1673, Failures: 2`**, both failures printing `/Users/dai.ha/LTMS/claude-bridge/fleetd/handover.md` as the wrong answer. One of its three cases pins that the lookup is read *live*: the map is empty when the factory runs and only gains the terminal afterwards, which is the real startup order, since leads are found by a tab scan after boot. `FleetdLeadRolloverWiringTest`'s javadoc claimed "no behavioural test can catch this wiring dropping out." True before the factory took a lookup, false after. Corrected — a wrong claim inside a test is how the next person decides not to write the missing test. **2. A relative `fleet.leaders.<name>.cwd` broke the absolute-path guarantee.** Found by a reviewer (terra, deliberately not the sonnet that wrote the diff). `base.resolve(path).normalize()` stays relative if the base is relative, silently reopening the same ambiguity — and falsifying the "always absolute" promise the new `PendingRollover` javadoc makes. Now `.toAbsolutePath()` on both branches. I mutated it back out: **`Tests run: 1673, Failures: 1`**, killed. Filed #488 for the underlying config gap: nothing validates `Leader.cwd()` at load, so a relative `cwd` is silently accepted. ### Deployed and proven live Build `Tests run: 1673, Failures: 0, Errors: 0, Skipped: 0`. Redeploy: jar `2520c3c051c7` → `0faa26d1bb10`, pid 1907 → 62179, `/healthz` 200, fresh `fleetd listening` line, no ERROR lines since restart, `fleet_whoami` → `primary`. Config now in `fleetd.yaml`: ```yaml leadRollover: handoverPath: .handover/HANDOVER.md ``` Live probe on the running daemon — `fleet_handover{action:"open"}`: ```json {"token":"56280346-...","handoverPath":"/Users/dai.ha/LTMS/claude-bridge/.handover/HANDOVER.md","requestedAtMillis":1789166971737} ``` and the daemon's own log line: ``` lead-rollover: open token=56280346-... lead=term_65a4c36d11af27 configuredHandoverPath=.handover/HANDOVER.md resolvedHandoverPath=/Users/dai.ha/LTMS/claude-bridge/.handover/HANDOVER.md ``` That is the relative value resolving against the **lead's** workspace, not the daemon's — through the real wiring, on the real daemon. Token cancelled afterwards; nothing was rolled. `.handover/` is gitignored (the file is a snapshot of live state). Docs updated: `wiki/11-Features.md` (pushed, verified by ref `285b0a6`) and the `handover` skill, which now says the returned path is always absolute and must not be re-resolved by the lead. The `CLAUDE.md` charter row needed no change; sync check still prints `in sync: True`. ### Still open **Acceptance criterion 6 is still not met** — nothing has yet shown end-to-end that a freshly cleared pane picks up the bootstrap prompt. It is now *reachable* for the first time, since the feature is configured, but the only honest proof is running a real rollover on a live lead, and that is the operator's call. If the bootstrap does not land, the lead's context is gone and no fresh session starts; the recovery is the manual path, which is exactly why the handover file must exist before `confirm` is accepted. #486 (unbounded settle poll hangs CI instead of failing) remains open and unassigned.
Author
Owner

Acceptance criterion 6 is now answered, negatively. Measured on the Mac, 2026-09-12, first real rollover of the shipped feature with a Claude Code lead.

The roll was confirmed at 07:03:30.903 and logged rolled at 07:03:43.998. The bootstrap prompt did not land. The pane received one submitted line, /clearFresh lead session. ..., and Claude Code answered Unknown command: /clearFresh.

The failure is safe. No /clear ever executed, so no context was lost, no session died, and the handover file stayed intact at .handover/HANDOVER.md. The roll simply did nothing.

Cause, evidence and fix are in #489. In short: the second settle wait is a no-op (/clear starts no turn, so the pane never leaves IDLE and the wait returns on its first poll - 438 ms of a 20-second budget), and under that the submit keystroke raced the paste, which AgentControl.submit's own javadoc already records as CB-113.

wiki/11-Features.md and the handover skill have been corrected: both said "not yet proven end-to-end", which is now too kind. They say "measured failing" and point at #489.

One observation, on one host, with one backend. It is a data point, not a law - but it is a reproducible defect in the code path, not a flake.

**Acceptance criterion 6 is now answered, negatively.** Measured on the Mac, 2026-09-12, first real rollover of the shipped feature with a Claude Code lead. The roll was confirmed at `07:03:30.903` and logged `rolled` at `07:03:43.998`. The bootstrap prompt did **not** land. The pane received one submitted line, `/clearFresh lead session. ...`, and Claude Code answered `Unknown command: /clearFresh`. The failure is safe. No `/clear` ever executed, so no context was lost, no session died, and the handover file stayed intact at `.handover/HANDOVER.md`. The roll simply did nothing. Cause, evidence and fix are in #489. In short: the second settle wait is a no-op (`/clear` starts no turn, so the pane never leaves `IDLE` and the wait returns on its first poll - 438 ms of a 20-second budget), and under that the submit keystroke raced the paste, which `AgentControl.submit`'s own javadoc already records as CB-113. `wiki/11-Features.md` and the `handover` skill have been corrected: both said "not yet proven end-to-end", which is now too kind. They say "measured failing" and point at #489. One observation, on one host, with one backend. It is a data point, not a law - but it is a reproducible defect in the code path, not a flake.
Author
Owner

The fleet01 test: reported here on that lead's behalf

The operator asked me to have fleet01 test the roll and to fix whatever came out. I asked the fleet01 lead over the coordination channel. They cannot run it, and that is the right answer. They have no forge credential, so they cannot post to this issue themselves. Everything in §1–§4 below is their measurement on their host. I have not re-run any of it. §5 is mine.

1. fleet01 cannot run the roll at all — the code is not there

Their report:

origin/main on fleet01 stops at #470          (91 commits behind)
fleetd/src/main/java/dev/ltms/fleet/lead/LeadRollover.java   -> does not exist
'leadRollover:' in their fleetd.yaml                         -> 0 matches
'bootstrapText' anywhere in their fleetd.yaml                -> 0 matches

So fleet01 is not a second instrument for the roll today. It is a host that would need a 91-commit update and new config written before a roll could even start. Note the last line: even after an update, bootstrapText has nothing to attach to there, so the fleet_whoami-first bootstrap would not exist.

2. They are not refusing — it is their operator's call

They have 12 unpushed commits on worker/git-native-removals-4633ad-2 at 38fd7f3 in the kb repo. Their operator reserved the rebuild and wants a human present for recovery. A roll that fails leaves a wiped lead and no fresh session, so doing it with unpushed work and nobody watching is the wrong trade. I agree with them. The fleet01 test stays blocked on their operator, not on us.

3. They corrected one of my claims — I was wrong

I wrote in #491 that fleet01's layout would expose #488 (fleet.leaders.<name>.cwd is never validated). It does not. Their cwd at fleetd.yaml:229 is already absolute (/home/ltms/LTMS/kb), so the bad input is simply not present. #488 still needs a host with a relative cwd, or a test that supplies one. Corrected on #491.

The part that survives is stronger than I filed it. For #487, their two directories (/home/ltms/LTMS/kb and /home/ltms/LTMS/fleetd/fleetd) share only /home/ltms/LTMS — not nested at any depth. On this Mac they are one level apart, the weakest separation that still counts.

4. They corrected a second claim, and it produced a new ticket

I had advised "build before you stop anything". They pointed out this contradicts #413, with evidence from their host on 2026-09-10: jar rewritten at 02:10:41, pid 1610855 died at 02:11:17 with NoClassDefFoundError: dev/ltms/fleet/msg/Rendezvous$Resolution, and systemd recorded status=143 — which reads as clean.

5. What I measured myself, and what I filed

I checked the same mechanism on this Mac before accepting it:

grep -c NoClassDefFoundError fleetd/fleetd.out   -> 1
grep -c 'fleetd listening' fleetd/fleetd.out     -> 58     (control, must be non-zero)
line 49902, in the boot that starts at line 49890:
  Exception in thread "Thread-0" java.lang.NoClassDefFoundError: reactor/core/Exceptions
today's redeploy window (line 68406 -> EOF)      -> 0 matches

Different OS, different supervisor, different class, different finder — so these are two data points, not one. Filed as #493: scripts/redeploy-fleetd.sh builds into the live target/fleetd.jar before stopping the daemon, so the shutdown drain can die on a failure-path class the running JVM never loaded. The exit code does not show it and nothing currently looks.

6. Where that leaves this issue

Criterion 6 — "a fresh session boots from the handover file" — is still not met on any host. PR #490 fixed the /clearFresh concatenation and is merged and deployed here, but it has not yet rolled a real lead session end to end. fleet01 cannot be that proof today.

## The fleet01 test: reported here on that lead's behalf The operator asked me to have fleet01 test the roll and to fix whatever came out. I asked the fleet01 lead over the coordination channel. **They cannot run it, and that is the right answer.** They have no forge credential, so they cannot post to this issue themselves. Everything in §1–§4 below is **their** measurement on **their** host. I have not re-run any of it. §5 is mine. ### 1. fleet01 cannot run the roll at all — the code is not there Their report: ``` origin/main on fleet01 stops at #470 (91 commits behind) fleetd/src/main/java/dev/ltms/fleet/lead/LeadRollover.java -> does not exist 'leadRollover:' in their fleetd.yaml -> 0 matches 'bootstrapText' anywhere in their fleetd.yaml -> 0 matches ``` So fleet01 is not a second instrument for the roll today. It is a host that would need a 91-commit update **and** new config written before a roll could even start. Note the last line: even after an update, `bootstrapText` has nothing to attach to there, so the `fleet_whoami`-first bootstrap would not exist. ### 2. They are not refusing — it is their operator's call They have 12 unpushed commits on `worker/git-native-removals-4633ad-2` at `38fd7f3` in the kb repo. Their operator reserved the rebuild and wants a human present for recovery. A roll that fails leaves a wiped lead and no fresh session, so doing it with unpushed work and nobody watching is the wrong trade. I agree with them. **The fleet01 test stays blocked on their operator, not on us.** ### 3. They corrected one of my claims — I was wrong I wrote in #491 that fleet01's layout would expose #488 (`fleet.leaders.<name>.cwd` is never validated). It does not. Their `cwd` at `fleetd.yaml:229` is already absolute (`/home/ltms/LTMS/kb`), so the bad input is simply not present. #488 still needs a host with a relative `cwd`, or a test that supplies one. Corrected on #491. The part that survives is **stronger** than I filed it. For #487, their two directories (`/home/ltms/LTMS/kb` and `/home/ltms/LTMS/fleetd/fleetd`) share only `/home/ltms/LTMS` — not nested at any depth. On this Mac they are one level apart, the weakest separation that still counts. ### 4. They corrected a second claim, and it produced a new ticket I had advised "build before you stop anything". They pointed out this contradicts #413, with evidence from their host on 2026-09-10: jar rewritten at 02:10:41, pid 1610855 died at 02:11:17 with `NoClassDefFoundError: dev/ltms/fleet/msg/Rendezvous$Resolution`, and systemd recorded `status=143` — which reads as clean. ### 5. What I measured myself, and what I filed I checked the same mechanism on this Mac before accepting it: ``` grep -c NoClassDefFoundError fleetd/fleetd.out -> 1 grep -c 'fleetd listening' fleetd/fleetd.out -> 58 (control, must be non-zero) line 49902, in the boot that starts at line 49890: Exception in thread "Thread-0" java.lang.NoClassDefFoundError: reactor/core/Exceptions today's redeploy window (line 68406 -> EOF) -> 0 matches ``` Different OS, different supervisor, different class, different finder — so these are two data points, not one. Filed as **#493**: `scripts/redeploy-fleetd.sh` builds into the live `target/fleetd.jar` before stopping the daemon, so the shutdown drain can die on a failure-path class the running JVM never loaded. The exit code does not show it and nothing currently looks. ### 6. Where that leaves this issue Criterion 6 — "a fresh session boots from the handover file" — is still **not met on any host**. PR #490 fixed the `/clearFresh` concatenation and is merged and deployed here, but it has not yet rolled a real lead session end to end. fleet01 cannot be that proof today.
Author
Owner

fleet01's operator has answered: deferred, not scheduled. Do not hold this issue for that host.

Reported for the fleet01 lead, who still has no forge credential. Their words, condensed: nothing changes on their host now, and the decision gets made when their operator is present rather than at a booked time. Booking a window would either park their operator waiting for it or make the lead roll unattended, and an unattended roll turns our test into their outage. Their host state is unchanged: 4887731, jar b92db8f4 built 2026-09-10 02:10:41, daemon pid 1696852, 91 commits behind, fetch window ends at #470.

So criterion 6 stays open at one host, one backend. If another lead comes online, they are a better second instrument than fleet01 precisely because they are available. This issue should not wait for fleet01.

Two facts from that host survive the deferral and are worth keeping:

  • it is a clean instrument for #487 when it does move (lead cwd and daemon cwd share nothing below /home/ltms/LTMS),
  • it is not an instrument for #488 at all (cwd already absolute). #488 needs a different host, or a test that supplies the bad input.

Their read-only review of #489/#490 produced two results

1. The question: does a test exist that would have caught the ORIGINAL defect, or does the new test pin the fix rather than the bug?

Answered by mutation, not by reading, since I wrote both the fix and the test. Baseline LeadRolloverTest = 29 tests, 0 failures.

Mutation Result
Put the original defect back (call site → waitUntilAtTurnBoundary) 1 failure — clearPickupIsNudgedBeforeBootstrapTextWhenPaneStaysIdle:498
Reduce the wait to agents.submit(target); return true; — satisfies "a nudge happened, ordered between /clear and bootstrapText" and nothing else 4 failures + 1 error

The second mutation is the important one. Three of those failures assert the negative — that bootstrapText is not sent when the wait should not release. That is "the wait actually waits", expressed as a consequence rather than as a duration. So the suite pins the bug, not only the fix.

Proof each mutant applied used two different search strings (mutant present, original gone) plus a control. After each restore: 29 / 0, source shasum byte-identical to the pristine copy.

2. Their second ask found a real defect my mutations could not. They wanted "a bound and a named failure message, not a happy path". Checking what an operator actually sees, LeadRollover:400 prints cfg.clearSettleSeconds() — the configured budget, not the measured wait. Had that path fired during the live incident it would have logged within 20s while the truth was 0.438s: confidently wrong, not silent. And :569 releases the roll at info, followed by rolled, so the case most likely to produce /clearFresh-shaped garbage reports success to anyone grepping for WARN.

Filed as #494 and delegated. #490 made the wait real; it did not make the wait observable.

## fleet01's operator has answered: deferred, not scheduled. Do not hold this issue for that host. Reported for the fleet01 lead, who still has no forge credential. Their words, condensed: nothing changes on their host now, and the decision gets made when their operator is present rather than at a booked time. Booking a window would either park their operator waiting for it or make the lead roll unattended, and an unattended roll turns our test into their outage. Their host state is unchanged: `4887731`, jar `b92db8f4` built 2026-09-10 02:10:41, daemon pid 1696852, 91 commits behind, fetch window ends at #470. **So criterion 6 stays open at one host, one backend.** If another lead comes online, they are a better second instrument than fleet01 precisely because they are available. This issue should not wait for fleet01. Two facts from that host survive the deferral and are worth keeping: - it **is** a clean instrument for #487 when it does move (lead `cwd` and daemon `cwd` share nothing below `/home/ltms/LTMS`), - it is **not** an instrument for #488 at all (`cwd` already absolute). #488 needs a different host, or a test that supplies the bad input. ## Their read-only review of #489/#490 produced two results **1. The question: does a test exist that would have caught the ORIGINAL defect, or does the new test pin the fix rather than the bug?** Answered by mutation, not by reading, since I wrote both the fix and the test. Baseline `LeadRolloverTest` = 29 tests, 0 failures. | Mutation | Result | |---|---| | Put the original defect back (call site → `waitUntilAtTurnBoundary`) | **1 failure** — `clearPickupIsNudgedBeforeBootstrapTextWhenPaneStaysIdle:498` | | Reduce the wait to `agents.submit(target); return true;` — satisfies "a nudge happened, ordered between `/clear` and `bootstrapText`" and nothing else | **4 failures + 1 error** | The second mutation is the important one. Three of those failures assert the *negative* — that `bootstrapText` is **not** sent when the wait should not release. That is "the wait actually waits", expressed as a consequence rather than as a duration. So the suite pins the bug, not only the fix. Proof each mutant applied used two different search strings (mutant present, original gone) plus a control. After each restore: 29 / 0, source `shasum` byte-identical to the pristine copy. **2. Their second ask found a real defect my mutations could not.** They wanted "a bound and a named failure message, not a happy path". Checking what an operator actually sees, `LeadRollover:400` prints `cfg.clearSettleSeconds()` — the **configured** budget, not the measured wait. Had that path fired during the live incident it would have logged `within 20s` while the truth was 0.438s: confidently wrong, not silent. And `:569` releases the roll at `info`, followed by `rolled`, so the case most likely to produce `/clearFresh`-shaped garbage reports success to anyone grepping for `WARN`. Filed as **#494** and delegated. #490 made the wait real; it did not make the wait observable.
Author
Owner

Correction to my mutation numbers above

I reported mutation 2 as "4 failures + 1 error". The error carries nothing and should not have been in the total. The fleet01 lead caught it from the shape of the result before I checked the stack.

submitThatThrowsDoesNotAbortTheRoll:629 errored, and the exception escaped from my stub, not from an assertion catching my stub:

java.lang.RuntimeException: simulated herdr transport failure on submit
  at LeadRolloverTest$5.call(LeadRolloverTest.java:614)
  at herdr.AgentControl.submit(AgentControl.java:127)
  at lead.LeadRollover.waitForClearPickupAndSettle(LeadRollover.java:550)   <- my stub's bare submit()

My stub was agents.submit(target); return true; with no try/catch; the real method wraps that call. So the mutant broke the test rather than the test catching the mutant.

Corrected: mutation 2 = 4 failures. The conclusion is unchanged and rests on those four, three of which assert the negative — bootstrapText is not sent when the wait should not release.

General rule worth keeping: when a mutation produces errors as well as failures, count only the failures. An error can be collateral damage from the mutant, and it looks like signal.

## Correction to my mutation numbers above I reported mutation 2 as **"4 failures + 1 error"**. The error carries nothing and should not have been in the total. The fleet01 lead caught it from the shape of the result before I checked the stack. `submitThatThrowsDoesNotAbortTheRoll:629` errored, and the exception escaped **from my stub**, not from an assertion catching my stub: ``` java.lang.RuntimeException: simulated herdr transport failure on submit at LeadRolloverTest$5.call(LeadRolloverTest.java:614) at herdr.AgentControl.submit(AgentControl.java:127) at lead.LeadRollover.waitForClearPickupAndSettle(LeadRollover.java:550) <- my stub's bare submit() ``` My stub was `agents.submit(target); return true;` with no `try`/`catch`; the real method wraps that call. So the mutant broke the test rather than the test catching the mutant. **Corrected: mutation 2 = 4 failures.** The conclusion is unchanged and rests on those four, three of which assert the negative — `bootstrapText` is not sent when the wait should not release. General rule worth keeping: **when a mutation produces errors as well as failures, count only the failures.** An error can be collateral damage from the mutant, and it looks like signal.
Sign in to join this conversation.
1 Participants
Notifications
Due Date
No due date set.
Dependencies

No dependencies set.

Reference: fleet/fleetd#480