A lead rollover reports rolled while the bootstrap paste is lost: the send skips the MemberPresence gate that exists for exactly this race #809

Closed
opened 2026-10-07 06:50:35 +02:00 by ltms · 3 comments
Owner

Measured on this host (mac) on 2026-10-07. This is #796 happening again, in the other direction. #796's fix made the send stop failing; it did not make the send arrive. The roll logged rolled, the successor got nothing, and the operator had to type the resume by hand.

MemberPresence's own class javadoc already names the mechanism. LeadRollover is the one delivery path in the daemon that does not read it.

What the log says

fleetd/fleetd.out lines 130629-130632. The file carries no dates; these times are CEST on 2026-10-07.

06:26:02.931 LeadLauncher  - lead 'opus' launched: profile=opus tab=wN:t2 pane=wN:p2 terminal=term_65d38832d8c734d
06:26:06.524 LeadTabScanner- lead/collaborator panes: {term_65d38832d8c734d=Entry[name=opus, kind=LEAD]}
06:26:06.527 LeadRollover  - lead-rollover: rolled token=e170714f-… oldLead=term_65d2a09b91f8f3d newTerminal=term_65d38832d8c734d elapsedMs=15634
06:26:07.110 McpAsyncServer- Client initialize request … Info: Implementation[name=claude-code, version=2.1.292]

Read the order. rolled is 3ms after the tab scanner line, so sendBootstrapWithRetry returned sent=true on its first attempt, at about 06:26:06.526. The fresh Claude Code connected its MCP client 583ms later. The daemon typed the bootstrap into a pane whose agent had not booted yet, herdr accepted the keystrokes, and the text went nowhere.

There is no bootstrapText was never sent line and no BOOTSTRAP_NEVER_SENT outcome. The retry window (relaunchReadySeconds, default 45s, not set in fleetd.yaml) was never used, because nothing refused.

What the session transcripts say

The successor's own transcript is ~/.ccs/instances/ltms/projects/-Users-dai-ha-LTMS-claude-bridge/611c0301-….jsonl. Its first entry is the operator's /clear at 2026-10-07T04:42:01.817Z — 06:42:01 CEST, 16 minutes after the launch. Nothing was recorded in between, because nothing arrived. The first real message is the operator's own, at 04:42:52Z.

Counting every transcript in that directory that holds the bootstrap as a received user message:

  2026-09-22 3
  2026-10-01 3
  2026-10-02 1
  2026-10-03 9
  2026-10-04 3
  total 19   last 2026-10-04T05:20:49.031Z (UTC)

That last one is 07:20:49.031 CEST, and its roll logged rolled at 07:20:48.707 — 324ms earlier. So a delivered bootstrap does show up in the transcript, within a third of a second. Nothing is recorded for 2026-10-06 or 2026-10-07. Two rolls, two losses: #796's (loud) and this one (silent).

Re-measure with:

cd ~/.ccs/instances/ltms/projects/-Users-dai-ha-LTMS-claude-bridge
python3 -I -c "
import json,glob,collections
c=collections.Counter(); last=''
for f in glob.glob('*.jsonl'):
    for l in open(f,errors='replace'):
        if 'Fresh lead session' not in l: continue
        try: d=json.loads(l)
        except: continue
        if d.get('type')!='user': continue
        m=d.get('message',{}).get('content')
        if isinstance(m,str) and m.lstrip().startswith('Fresh lead session'):
            ts=d['timestamp']; c[ts[:10]]+=1; last=max(last,ts)
for k in sorted(c): print(k,c[k])
print('total',sum(c.values()),'last',last)
"

Cause, read in the code

MemberPresence.java:6-17 describes this exact failure, and says the daemon already has the right signal:

Tracks which peers are available — their Claude has booted and connected its MCP client to the bridge. For a spawned member this is the reliable readiness signal, unlike herdr's agent_status, which reports idle while its Claude is still booting. Delivering into that boot window pastes into a not-yet-ready TUI (the text is lost) and wedges that member's delivery state, so the Injector holds a spawned member's first delivery until it is present here.

So for a spawned member, Injector holds the first delivery until MCP presence. For a fresh lead, LeadRollover does not:

  • LeadRollover.waitUntilPaneReady (LeadRollover.java:898) polls agents.status(newTerminal) and accepts IDLE or DONE. That is herdr's agent_status — the signal MemberPresence's javadoc says you must not trust for this. Its own javadoc claims it stops bootstrapText being "typed into a pane that has not actually finished booting". It cannot keep that promise.
  • LeadRollover.sendBootstrapWithRetry (LeadRollover.java:781) calls agents.send — raw herdr keystrokes. It is the only message in the daemon that does not go through Injector.enqueue, so it gets no status gate, no PaneInbox collection route for a mod-served pane, and no delivery future. "herdr accepted the keystrokes" is the only thing the roll can check, and it is not evidence of receipt.

Two separate defects, then:

  1. The readiness gate reads the wrong signal. agent_status is idle during the boot window.
  2. Nothing confirms the bootstrap landed. A send into the boot window is indistinguishable from a delivered one, so the roll reports success.

Why this is the worst outcome a roll can have

A refused roll keeps your context. This one threw the context away and put nothing in its place. The successor has no handover file, no ticket, no token to call fleet_handover{action:"status"} with, and no reason to suspect anything is wrong. Here it sat idle for 16 minutes until a person noticed.

Suggested fix

  • Make waitUntilPaneReady require MCP presence as well as a turn boundary. MemberPresence.isPresent(newTerminal) is the signal. Its existing timeout state RELAUNCH_NEVER_READY already covers the failure, so this adds no new outcome, and it makes that method's javadoc true.
  • Confirm the send landed. After a successful agents.send, wait a bounded time for the fresh pane to start a turn (agent_status WORKING). If it never does, record a new terminal outcome and log a WARN, so a lost paste is visible instead of being reported as rolled. Report only — do not re-send, because a second paste into a session that did receive the first would be worse.
  • FleetConfig.relaunchReadySeconds's javadoc was corrected under #796 to say it bounds three waits. A new wait makes it four. Correct it in the same pass, and update the outcome table in .claude/skills/handover/SKILL.md.
  • Longer term, worth its own ticket: route bootstrapText through Injector.enqueue instead of raw agents.send, and wait on the Delivery future. That would give the bootstrap the same status gate, inbox-collection route and delivery proof every other message already has, instead of re-implementing a weaker version of each.

Not measured

  • I did not reproduce this. It is one occurrence. Rolling the only live lead to test it would destroy this session's context, which is the damage the ticket is about.
  • I did not confirm by observation that the keystrokes were lost in the TUI boot window. What I measured is the ordering (send 583ms before the MCP connect) and the absence of the text in every transcript. The mechanism is MemberPresence's own documented one, not my own reading of herdr.
  • I did not check whether the MCP initialize request is the earliest moment the TUI accepts input. It may be later still, in which case presence is necessary but may not be sufficient — which is the second half of the fix.
  • 21 lead-rollover: rolled lines are in the current fleetd.out against 19 recorded deliveries, but that file spans more than this project's rolls, so I am not treating the difference as two extra losses.
Measured on this host (mac) on 2026-10-07. This is #796 happening again, in the other direction. #796's fix made the send stop failing; it did not make the send arrive. The roll logged `rolled`, the successor got nothing, and the operator had to type the resume by hand. `MemberPresence`'s own class javadoc already names the mechanism. `LeadRollover` is the one delivery path in the daemon that does not read it. ## What the log says `fleetd/fleetd.out` lines 130629-130632. The file carries no dates; these times are CEST on 2026-10-07. ``` 06:26:02.931 LeadLauncher - lead 'opus' launched: profile=opus tab=wN:t2 pane=wN:p2 terminal=term_65d38832d8c734d 06:26:06.524 LeadTabScanner- lead/collaborator panes: {term_65d38832d8c734d=Entry[name=opus, kind=LEAD]} 06:26:06.527 LeadRollover - lead-rollover: rolled token=e170714f-… oldLead=term_65d2a09b91f8f3d newTerminal=term_65d38832d8c734d elapsedMs=15634 06:26:07.110 McpAsyncServer- Client initialize request … Info: Implementation[name=claude-code, version=2.1.292] ``` Read the order. `rolled` is 3ms after the tab scanner line, so `sendBootstrapWithRetry` returned `sent=true` on its first attempt, at about `06:26:06.526`. The fresh Claude Code connected its MCP client **583ms later**. The daemon typed the bootstrap into a pane whose agent had not booted yet, herdr accepted the keystrokes, and the text went nowhere. There is no `bootstrapText was never sent` line and no `BOOTSTRAP_NEVER_SENT` outcome. The retry window (`relaunchReadySeconds`, default 45s, not set in `fleetd.yaml`) was never used, because nothing refused. ## What the session transcripts say The successor's own transcript is `~/.ccs/instances/ltms/projects/-Users-dai-ha-LTMS-claude-bridge/611c0301-….jsonl`. Its first entry is the operator's `/clear` at `2026-10-07T04:42:01.817Z` — 06:42:01 CEST, **16 minutes after the launch**. Nothing was recorded in between, because nothing arrived. The first real message is the operator's own, at 04:42:52Z. Counting every transcript in that directory that holds the bootstrap as a received user message: ``` 2026-09-22 3 2026-10-01 3 2026-10-02 1 2026-10-03 9 2026-10-04 3 total 19 last 2026-10-04T05:20:49.031Z (UTC) ``` That last one is 07:20:49.031 CEST, and its roll logged `rolled` at 07:20:48.707 — 324ms earlier. So a delivered bootstrap does show up in the transcript, within a third of a second. **Nothing is recorded for 2026-10-06 or 2026-10-07.** Two rolls, two losses: #796's (loud) and this one (silent). Re-measure with: ```bash cd ~/.ccs/instances/ltms/projects/-Users-dai-ha-LTMS-claude-bridge python3 -I -c " import json,glob,collections c=collections.Counter(); last='' for f in glob.glob('*.jsonl'): for l in open(f,errors='replace'): if 'Fresh lead session' not in l: continue try: d=json.loads(l) except: continue if d.get('type')!='user': continue m=d.get('message',{}).get('content') if isinstance(m,str) and m.lstrip().startswith('Fresh lead session'): ts=d['timestamp']; c[ts[:10]]+=1; last=max(last,ts) for k in sorted(c): print(k,c[k]) print('total',sum(c.values()),'last',last) " ``` ## Cause, read in the code `MemberPresence.java:6-17` describes this exact failure, and says the daemon already has the right signal: > Tracks which peers are *available* — their Claude has booted and connected its MCP client to the bridge. For a spawned member this is the reliable readiness signal, **unlike herdr's `agent_status`, which reports `idle` while its Claude is still booting. Delivering into that boot window pastes into a not-yet-ready TUI (the text is lost)** and wedges that member's delivery state, so the `Injector` holds a spawned member's first delivery until it is present here. So for a spawned member, `Injector` holds the first delivery until MCP presence. For a fresh lead, `LeadRollover` does not: - `LeadRollover.waitUntilPaneReady` (`LeadRollover.java:898`) polls `agents.status(newTerminal)` and accepts `IDLE` or `DONE`. That is herdr's `agent_status` — the signal `MemberPresence`'s javadoc says you must not trust for this. Its own javadoc claims it stops `bootstrapText` being "typed into a pane that has not actually finished booting". It cannot keep that promise. - `LeadRollover.sendBootstrapWithRetry` (`LeadRollover.java:781`) calls `agents.send` — raw herdr keystrokes. It is the only message in the daemon that does not go through `Injector.enqueue`, so it gets no status gate, no `PaneInbox` collection route for a mod-served pane, and **no delivery future**. "herdr accepted the keystrokes" is the only thing the roll can check, and it is not evidence of receipt. Two separate defects, then: 1. **The readiness gate reads the wrong signal.** `agent_status` is `idle` during the boot window. 2. **Nothing confirms the bootstrap landed.** A send into the boot window is indistinguishable from a delivered one, so the roll reports success. ## Why this is the worst outcome a roll can have A refused roll keeps your context. This one threw the context away and put nothing in its place. The successor has no handover file, no ticket, no token to call `fleet_handover{action:"status"}` with, and no reason to suspect anything is wrong. Here it sat idle for 16 minutes until a person noticed. ## Suggested fix - **Make `waitUntilPaneReady` require MCP presence as well as a turn boundary.** `MemberPresence.isPresent(newTerminal)` is the signal. Its existing timeout state `RELAUNCH_NEVER_READY` already covers the failure, so this adds no new outcome, and it makes that method's javadoc true. - **Confirm the send landed.** After a successful `agents.send`, wait a bounded time for the fresh pane to start a turn (`agent_status` `WORKING`). If it never does, record a new terminal outcome and log a WARN, so a lost paste is visible instead of being reported as `rolled`. Report only — do not re-send, because a second paste into a session that did receive the first would be worse. - `FleetConfig.relaunchReadySeconds`'s javadoc was corrected under #796 to say it bounds **three** waits. A new wait makes it four. Correct it in the same pass, and update the outcome table in `.claude/skills/handover/SKILL.md`. - Longer term, worth its own ticket: route `bootstrapText` through `Injector.enqueue` instead of raw `agents.send`, and wait on the `Delivery` future. That would give the bootstrap the same status gate, inbox-collection route and delivery proof every other message already has, instead of re-implementing a weaker version of each. ## Not measured - I did not reproduce this. It is one occurrence. Rolling the only live lead to test it would destroy this session's context, which is the damage the ticket is about. - I did not confirm by observation that the keystrokes were lost in the TUI boot window. What I measured is the ordering (send 583ms before the MCP connect) and the absence of the text in every transcript. The mechanism is `MemberPresence`'s own documented one, not my own reading of herdr. - I did not check whether the MCP `initialize` request is the earliest moment the TUI accepts input. It may be later still, in which case presence is necessary but may not be sufficient — which is the second half of the fix. - 21 `lead-rollover: rolled` lines are in the current `fleetd.out` against 19 recorded deliveries, but that file spans more than this project's rolls, so I am not treating the difference as two extra losses.
Author
Owner

Correction from the lead for the worker on worker/809-ed41e4-1. This comment is newer than your brief and it wins.

Your mvn clean install is not slow, it is spinning. Do not wait for it. I took a thread dump of your forked surefire JVM (pid 83235) at 07:17:24:

"main" #3 [6147] ... cpu=423691.44ms elapsed=441.80s ... runnable
	at dev.ltms.fleet.lead.LeadRolloverTest.statusHistoryDoesNotGrowPastItsCap(LeadRolloverTest.java:1079)

7 minutes of CPU on one test, runnable the whole time, and no new surefire report since 07:10:16.

The cause, and it is your change, not that test

Your new post-send wait polls for AgentStatus.WORKING and bounds it with relaunchReadySeconds (45s). No existing full-roll test ever drives the fresh terminal to WORKING — FakeHerdr leaves it at the turn boundary — so every roll in those tests now pays the whole bound. statusHistoryDoesNotGrowPastItsCap drives OUTCOME_HISTORY_CAP + 50 = 250 rolls with nowMillis = () -> clock.addAndGet(1) and pollSleeper = () -> { }, so that is 250 × 45,000 = about 11 million iterations of a no-sleep loop, each one calling into FakeHerdr. It terminates eventually; it is useless as a test run.

The presence gate is not the problem. newRolloverForAFullRoll already defaults mcpPresent to _ -> true, and that part of your change looks right.

What to do

  1. Give the post-send wait its own, much smaller bound. Reusing relaunchReadySeconds was my suggestion in the brief and it was wrong: 45s is sized for a pane to boot, not for an agent that has already been handed a prompt to start acting on it. Pick a bound of a few seconds, name it in the code, and say in your reply what you chose and why. That also keeps the FleetConfig javadoc at three waits instead of four, so skip that part of the brief — but still add the new RollState row to .claude/skills/handover/SKILL.md.
  2. Do not let an existing test pay a full timeout. Whatever bound you pick, the default test path must satisfy the new wait immediately rather than time out. Make FakeHerdr able to report a WORKING status for the fresh terminal and have the full-roll helper use it, so the existing tests stay fast and keep asserting ROLLED.
  3. Re-check the signal itself. agent_status is polled every 250ms in production. A successor that starts its turn and finishes it inside one poll interval would never be observed as WORKING, and your wait would then report a lost bootstrap that in fact landed — a false alarm in the one place an operator must be able to trust. Prefer a signal that cannot be missed between polls: any departure from the turn boundary, or the session's turn counter if one is reachable from here. Choose one, say which, and say what you rejected.

Unchanged

Everything else in the brief stands: report only, no re-send; do not grow FleetdAssembly.assembleAndStart; mvn clean install unpiped with the real Tests run: line; your own PR; no merge and no redeploy.

Kill the running build before you start — it is burning a core for nothing.

**Correction from the lead for the worker on `worker/809-ed41e4-1`. This comment is newer than your brief and it wins.** Your `mvn clean install` is not slow, it is spinning. Do not wait for it. I took a thread dump of your forked surefire JVM (pid 83235) at 07:17:24: ``` "main" #3 [6147] ... cpu=423691.44ms elapsed=441.80s ... runnable at dev.ltms.fleet.lead.LeadRolloverTest.statusHistoryDoesNotGrowPastItsCap(LeadRolloverTest.java:1079) ``` 7 minutes of CPU on one test, `runnable` the whole time, and no new surefire report since 07:10:16. ## The cause, and it is your change, not that test Your new post-send wait polls for `AgentStatus.WORKING` and bounds it with `relaunchReadySeconds` (45s). No existing full-roll test ever drives the fresh terminal to `WORKING` — `FakeHerdr` leaves it at the turn boundary — so **every** roll in those tests now pays the whole bound. `statusHistoryDoesNotGrowPastItsCap` drives `OUTCOME_HISTORY_CAP + 50` = 250 rolls with `nowMillis = () -> clock.addAndGet(1)` and `pollSleeper = () -> { }`, so that is 250 × 45,000 = about 11 million iterations of a no-sleep loop, each one calling into `FakeHerdr`. It terminates eventually; it is useless as a test run. The presence gate is not the problem. `newRolloverForAFullRoll` already defaults `mcpPresent` to `_ -> true`, and that part of your change looks right. ## What to do 1. **Give the post-send wait its own, much smaller bound.** Reusing `relaunchReadySeconds` was my suggestion in the brief and it was wrong: 45s is sized for a pane to boot, not for an agent that has already been handed a prompt to start acting on it. Pick a bound of a few seconds, name it in the code, and say in your reply what you chose and why. That also keeps the `FleetConfig` javadoc at three waits instead of four, so **skip that part of the brief** — but still add the new `RollState` row to `.claude/skills/handover/SKILL.md`. 2. **Do not let an existing test pay a full timeout.** Whatever bound you pick, the default test path must satisfy the new wait immediately rather than time out. Make `FakeHerdr` able to report a `WORKING` status for the fresh terminal and have the full-roll helper use it, so the existing tests stay fast and keep asserting `ROLLED`. 3. **Re-check the signal itself.** `agent_status` is polled every 250ms in production. A successor that starts its turn and finishes it inside one poll interval would never be observed as `WORKING`, and your wait would then report a lost bootstrap that in fact landed — a false alarm in the one place an operator must be able to trust. Prefer a signal that cannot be missed between polls: any departure from the turn boundary, or the session's turn counter if one is reachable from here. Choose one, say which, and say what you rejected. ## Unchanged Everything else in the brief stands: report only, no re-send; do not grow `FleetdAssembly.assembleAndStart`; `mvn clean install` unpiped with the real `Tests run:` line; your own PR; no merge and no redeploy. Kill the running build before you start — it is burning a core for nothing.
Author
Owner

Lead note for the worker — one thing I did to your worktree, so a confusing build failure has an explanation.

The spinning build was never killed, so at 07:22 two mvn runs were live in /Users/dai.ha/LTMS/.bridged-worktrees/447965-1 at once: the old one (pid 83008, 12m13s elapsed) and your new one (pid 9918, started 07:21:06). They share fleetd/target. I killed 83008 and its surefire children; your 9918 is untouched.

If your current build fails in a way that makes no sense — a missing class, a half-written jar, a surefire report for a test you did not run — that is why. Re-run mvn -o clean install once and treat the second result as the real one. Also note target/surefire-reports can still hold XML from the killed run; a stale report counts tests that never ran in your build, so read the Tests run: total from the build output, not from the report files.

Your ticket task-8dd115-1 has also gone terminal (30-minute async timeout) while you were building, so the id no longer accepts a reply. Still end your turn with exactly one fleet_reply — it lands in your pane's inbox and I drain it with fleet_poll{target}. Nothing is lost.

No other change to your tree, and nothing else in the previous comment is altered.

**Lead note for the worker — one thing I did to your worktree, so a confusing build failure has an explanation.** The spinning build was never killed, so at 07:22 two `mvn` runs were live in `/Users/dai.ha/LTMS/.bridged-worktrees/447965-1` at once: the old one (pid 83008, 12m13s elapsed) and your new one (pid 9918, started 07:21:06). They share `fleetd/target`. I killed 83008 and its surefire children; your 9918 is untouched. If your current build fails in a way that makes no sense — a missing class, a half-written jar, a surefire report for a test you did not run — that is why. Re-run `mvn -o clean install` once and treat the second result as the real one. Also note `target/surefire-reports` can still hold XML from the killed run; a stale report counts tests that never ran in your build, so read the `Tests run:` total from the build output, not from the report files. Your ticket `task-8dd115-1` has also gone terminal (30-minute async timeout) while you were building, so the id no longer accepts a reply. **Still end your turn with exactly one `fleet_reply`** — it lands in your pane's inbox and I drain it with `fleet_poll{target}`. Nothing is lost. No other change to your tree, and nothing else in the previous comment is altered.
Author
Owner

Fixed on main in a144dbe.

What changed

waitUntilPaneReady now requires two signals, not one: a real turn boundary (IDLE or DONE) and mcpPresent — MemberPresence::isPresent, wired at FleetdAssembly.java:460. That is the same signal Injector already uses to hold a spawned member's first delivery. LeadRollover was the only delivery path that did not read it.

RELAUNCH_NEVER_READY now names which of the two was missing: last status=… (a real turn boundary is IDLE or DONE), bridge MCP connected=…. The old message named only the turn boundary, so the presence half would have been invisible.

After the send, a new waitUntilTurnStarted polls for a live turn (WORKING or BLOCKED), bounded by BOOTSTRAP_CONFIRM_TIMEOUT_SECONDS = 5 — not relaunchReadySeconds. That budget sizes a pane booting a CLI; this one sizes an agent reacting to text it already holds, which takes one status poll. A miss reports the new BOOTSTRAP_NOT_CONFIRMED and re-sends nothing: the status is polled every 250ms, so a turn shorter than one interval reads as unconfirmed, and a second paste into a session that did receive the first is worse than one unconfirmed roll. The detail text says so, and tells the reader to check the successor.

Verified

  • mvn -o clean install in a throwaway worktree, unpiped: BUILD SUCCESS, Tests run: 2249, Failures: 0, Errors: 0, Skipped: 0, 0 ^[ERROR] lines. Merge was a fast-forward and HEAD^{tree} matched the tree I built.
  • Both gates mutation-checked, one at a time:
    • if (atTurnBoundary && mcpSeen) → if (atTurnBoundary) fails freshTerminalAtTurnBoundaryButNotMcpPresentNeverBecomesReadyNeverSendsBootstrapText (expected: <0> but was: <1> prompt calls).
    • short-circuiting the confirmation fails bootstrapTextSentButNeverConfirmedEndsInBootstrapNotConfirmedWithNoResend (expected: <BOOTSTRAP_NOT_CONFIRMED> but was: <ROLLED>).
  • The roll history, re-measured 2026-10-07: grep -c "lead-rollover: rolled" fleetd/fleetd.out = 21, of which exactly one carries the restart path's oldLead=… newTerminal=… shape. The other 20 log lead=… and ran under the old /clear behaviour. So the restart path had been exercised once, and that once lost the bootstrap.

Not fixed here, and worth knowing

The daemon running while I write this is still the 06:21 jar, so this fix is not live until a redeploy. That is next.

FakeHerdr gained turnStartsOnPrompt(): after each agent.prompt it reports "working" for the next two agent.get reads. Two, because one AgentControl.status() costs two raw reads — resolveTarget issues its own agent.get first. It is bounded rather than sticky so a test can roll more than once; the earlier sticky version made every multi-roll test spin forever against a non-advancing fake clock.

FleetdLeadRolloverAssemblyTest now calls markPresent("term_new_1"), because no real Claude boots in that test and the gate would otherwise stop the roll at readiness before it reaches the behaviour the test is about.

Docs: .claude/skills/handover/SKILL.md gains the BOOTSTRAP_NOT_CONFIRMED row and corrects the timeout section — four waits now, under three budgets, not "four times 45". wiki/11-Features.md replaces its "nobody has rolled under the restart path yet" gotcha, which this ticket falsified.

Fixed on `main` in `a144dbe`. ## What changed `waitUntilPaneReady` now requires **two** signals, not one: a real turn boundary (`IDLE` or `DONE`) **and** `mcpPresent` — `MemberPresence::isPresent`, wired at `FleetdAssembly.java:460`. That is the same signal `Injector` already uses to hold a spawned member's first delivery. `LeadRollover` was the only delivery path that did not read it. `RELAUNCH_NEVER_READY` now names which of the two was missing: `last status=… (a real turn boundary is IDLE or DONE), bridge MCP connected=…`. The old message named only the turn boundary, so the presence half would have been invisible. After the send, a new `waitUntilTurnStarted` polls for a live turn (`WORKING` or `BLOCKED`), bounded by `BOOTSTRAP_CONFIRM_TIMEOUT_SECONDS = 5` — **not** `relaunchReadySeconds`. That budget sizes a pane booting a CLI; this one sizes an agent reacting to text it already holds, which takes one status poll. A miss reports the new `BOOTSTRAP_NOT_CONFIRMED` and **re-sends nothing**: the status is polled every 250ms, so a turn shorter than one interval reads as unconfirmed, and a second paste into a session that did receive the first is worse than one unconfirmed roll. The detail text says so, and tells the reader to check the successor. ## Verified - `mvn -o clean install` in a throwaway worktree, unpiped: `BUILD SUCCESS`, `Tests run: 2249, Failures: 0, Errors: 0, Skipped: 0`, 0 `^[ERROR]` lines. Merge was a fast-forward and `HEAD^{tree}` matched the tree I built. - Both gates mutation-checked, one at a time: - `if (atTurnBoundary && mcpSeen)` → `if (atTurnBoundary)` fails `freshTerminalAtTurnBoundaryButNotMcpPresentNeverBecomesReadyNeverSendsBootstrapText` (`expected: <0> but was: <1>` prompt calls). - short-circuiting the confirmation fails `bootstrapTextSentButNeverConfirmedEndsInBootstrapNotConfirmedWithNoResend` (`expected: <BOOTSTRAP_NOT_CONFIRMED> but was: <ROLLED>`). - The roll history, re-measured 2026-10-07: `grep -c "lead-rollover: rolled" fleetd/fleetd.out` = 21, of which exactly **one** carries the restart path's `oldLead=… newTerminal=…` shape. The other 20 log `lead=…` and ran under the old `/clear` behaviour. So the restart path had been exercised once, and that once lost the bootstrap. ## Not fixed here, and worth knowing The daemon running while I write this is still the 06:21 jar, so this fix is not live until a redeploy. That is next. `FakeHerdr` gained `turnStartsOnPrompt()`: after each `agent.prompt` it reports `"working"` for the next **two** `agent.get` reads. Two, because one `AgentControl.status()` costs two raw reads — `resolveTarget` issues its own `agent.get` first. It is bounded rather than sticky so a test can roll more than once; the earlier sticky version made every multi-roll test spin forever against a non-advancing fake clock. `FleetdLeadRolloverAssemblyTest` now calls `markPresent("term_new_1")`, because no real Claude boots in that test and the gate would otherwise stop the roll at readiness before it reaches the behaviour the test is about. Docs: `.claude/skills/handover/SKILL.md` gains the `BOOTSTRAP_NOT_CONFIRMED` row and corrects the timeout section — four waits now, under three budgets, not "four times 45". `wiki/11-Features.md` replaces its "nobody has rolled under the restart path yet" gotcha, which this ticket falsified.
ltms closed this issue 2026-10-07 08:49:34 +02:00
Sign in to join this conversation.
1 Participants
Notifications
Due Date
No due date set.
Dependencies

No dependencies set.

Reference: fleet/fleetd#809