A lead rollover can replace the lead with a session that never learns it is a successor: the bootstrap send races the agent becoming named, and it is not retried #796

Open
opened 2026-10-06 14:04:15 +02:00 by ltms · 3 comments
Owner

Measured on this host (mac) on 2026-10-06. The roll ended the old lead, started a fresh one, and then failed to deliver bootstrapText. The new session came up with no instructions and no knowledge of the handover file. The operator found it and asked what it was.

This is the worst outcome a roll can have. A refused roll keeps your context. This one threw the context away and delivered nothing in its place.

What the log says

13:09:14.676 LeadRollover - lead-rollover: open token=0a78dc4c-… lead=term_65d106559b02e1
13:09:46.893 LeadRollover - lead-rollover: confirmed token=0a78dc4c-… — roll scheduled once the calling turn ends
13:09:53.406 LeadLauncher - lead 'opus' launched: profile=opus tab=w2:t3G pane=w2:p3Z terminal=term_65d2a09b91f8f3d label='lead'
13:09:55.935 LeadTabScanner - lead/collaborator panes: {term_65d2a09b91f8f3d=Entry[name=opus, kind=LEAD]}
13:09:56.178 WARN LeadRollover - lead-rollover: continuation for token=0a78dc4c-… threw
             dev.ltms.fleet.herdr.HerdrException: herdr error [agent_not_ready]: agent w2:p3Z is not
             an active named agent — the roll is dead; no further step in this continuation will run
	at dev.ltms.fleet.herdr.AgentControl.send(AgentControl.java:117)
	at dev.ltms.fleet.lead.LeadRollover.runRolloverUnguarded(LeadRollover.java:734)
13:09:56.889 McpAsyncServer - Client initialize request … Info: Implementation[name=claude-code, version=2.1.291]

The send failed 2.772s after the launch. The new Claude Code connected the MCP 0.711s after the send failed.

Cause, read in the code

LeadRollover.java:734:

IdentityResult identityResult = waitUntilRecognisedAsLead(newAgent.terminalId(),
        cfg.relaunchReadySeconds());
agents.send(newAgent.terminalId(), cfg.bootstrapTextFor(p.handoverPath()));

Two problems, and they are separate.

1. The readiness gates do not read what agent.send requires. waitUntilPaneReady waits for a turn boundary and waitUntilRecognisedAsLead waits for the tab scanner to report kind=LEAD. Both passed here — the scanner line is 0.243s before the throw. herdr's agent.send then refused because the pane is not yet "an active named agent". So a roll can pass every gate it has and still be unable to send. This is the shape of #669: a capability checked at one gate and refused at another.

2. The send is a bare call with no retry and no catch. One transient agent_not_ready kills the whole continuation. Nothing retries, and no RollStatus is recorded for this path, so fleet_handover{action:"status"} cannot report it either — the only record is this WARN line. The outcome table in the handover skill has no row for it.

Why it matters more than it looks

A fresh lead's only inheritance is the handover file, and the only thing that tells it the file exists is bootstrapText. When that send is dropped, the successor has:

  • no knowledge of the file,
  • no knowledge of any outstanding ticket (this roll had one, task-7da785-4),
  • no reason to suspect anything is wrong.

In this case the successor sat idle while a delegated worker finished, timed out, and lost its reply. The work survived only because the worker had pushed a branch.

Suggested fix

  • Retry agents.send on agent_not_ready for a bounded window, rather than once.
  • Make the pre-send gate read the same readiness that agent.send requires, or drop the separate gate and let the retry be the gate.
  • Wrap the send so a failure records a RollStatus (a new outcome, for example BOOTSTRAP_NEVER_SENT) instead of only a WARN. The caller is already gone, so the status table and the log are the only readers left.
  • Consider a fallback that does not depend on a paste: if the bootstrap cannot be delivered, the successor has no other route to the file.

Not measured

  • I did not reproduce this. It is one occurrence, and the margin was under one second, so I cannot say how often it happens. 20 earlier rolls are recorded as rolled, all under the older /clear behaviour.
  • I did not check whether herdr's "active named agent" state has an observable field the gate could read instead.
  • elapsedMs is not in the log for this roll, because the continuation threw before the success line.

Side observation, not part of this defect

While the successor was idle, ReplyPushLoop held its nudge 11 times in a row from 13:39:28 to 13:41:59, because the operator had 29 characters of unsubmitted text in the prompt box. That is the box gate working as designed, but it means the one signal that would have told the successor about the finished ticket was also blocked. Reported for context only.

Measured on this host (mac) on 2026-10-06. The roll ended the old lead, started a fresh one, and then failed to deliver `bootstrapText`. The new session came up with no instructions and no knowledge of the handover file. The operator found it and asked what it was. This is the worst outcome a roll can have. A refused roll keeps your context. This one threw the context away and delivered nothing in its place. ## What the log says ``` 13:09:14.676 LeadRollover - lead-rollover: open token=0a78dc4c-… lead=term_65d106559b02e1 13:09:46.893 LeadRollover - lead-rollover: confirmed token=0a78dc4c-… — roll scheduled once the calling turn ends 13:09:53.406 LeadLauncher - lead 'opus' launched: profile=opus tab=w2:t3G pane=w2:p3Z terminal=term_65d2a09b91f8f3d label='lead' 13:09:55.935 LeadTabScanner - lead/collaborator panes: {term_65d2a09b91f8f3d=Entry[name=opus, kind=LEAD]} 13:09:56.178 WARN LeadRollover - lead-rollover: continuation for token=0a78dc4c-… threw dev.ltms.fleet.herdr.HerdrException: herdr error [agent_not_ready]: agent w2:p3Z is not an active named agent — the roll is dead; no further step in this continuation will run at dev.ltms.fleet.herdr.AgentControl.send(AgentControl.java:117) at dev.ltms.fleet.lead.LeadRollover.runRolloverUnguarded(LeadRollover.java:734) 13:09:56.889 McpAsyncServer - Client initialize request … Info: Implementation[name=claude-code, version=2.1.291] ``` The send failed 2.772s after the launch. The new Claude Code connected the MCP 0.711s **after** the send failed. ## Cause, read in the code `LeadRollover.java:734`: ```java IdentityResult identityResult = waitUntilRecognisedAsLead(newAgent.terminalId(), cfg.relaunchReadySeconds()); agents.send(newAgent.terminalId(), cfg.bootstrapTextFor(p.handoverPath())); ``` Two problems, and they are separate. **1. The readiness gates do not read what `agent.send` requires.** `waitUntilPaneReady` waits for a turn boundary and `waitUntilRecognisedAsLead` waits for the tab scanner to report `kind=LEAD`. Both passed here — the scanner line is 0.243s before the throw. herdr's `agent.send` then refused because the pane is not yet "an active named agent". So a roll can pass every gate it has and still be unable to send. This is the shape of #669: a capability checked at one gate and refused at another. **2. The send is a bare call with no retry and no catch.** One transient `agent_not_ready` kills the whole continuation. Nothing retries, and no `RollStatus` is recorded for this path, so `fleet_handover{action:"status"}` cannot report it either — the only record is this `WARN` line. The outcome table in the handover skill has no row for it. ## Why it matters more than it looks A fresh lead's **only** inheritance is the handover file, and the only thing that tells it the file exists is `bootstrapText`. When that send is dropped, the successor has: - no knowledge of the file, - no knowledge of any outstanding ticket (this roll had one, `task-7da785-4`), - no reason to suspect anything is wrong. In this case the successor sat idle while a delegated worker finished, timed out, and lost its reply. The work survived only because the worker had pushed a branch. ## Suggested fix - Retry `agents.send` on `agent_not_ready` for a bounded window, rather than once. - Make the pre-send gate read the same readiness that `agent.send` requires, or drop the separate gate and let the retry be the gate. - Wrap the send so a failure records a `RollStatus` (a new outcome, for example `BOOTSTRAP_NEVER_SENT`) instead of only a `WARN`. The caller is already gone, so the status table and the log are the only readers left. - Consider a fallback that does not depend on a paste: if the bootstrap cannot be delivered, the successor has no other route to the file. ## Not measured - I did not reproduce this. It is one occurrence, and the margin was under one second, so I cannot say how often it happens. 20 earlier rolls are recorded as `rolled`, all under the older `/clear` behaviour. - I did not check whether `herdr`'s "active named agent" state has an observable field the gate could read instead. - `elapsedMs` is not in the log for this roll, because the continuation threw before the success line. ## Side observation, not part of this defect While the successor was idle, `ReplyPushLoop` held its nudge 11 times in a row from 13:39:28 to 13:41:59, because the operator had 29 characters of unsubmitted text in the prompt box. That is the box gate working as designed, but it means the one signal that would have told the successor about the finished ticket was also blocked. Reported for context only.
Author
Owner

Correction to my own issue text. One claim above is wrong.

I wrote: "no RollStatus is recorded for this path, so fleet_handover{action:"status"} cannot report it either — the only record is this WARN line."

That is false. I had not read the catch when I wrote it. LeadRollover.runRollover (LeadRollover.java:647-655) catches RuntimeException and does record an outcome:

outcomes.put(p.token(), new RollStatus(RollState.FAILED,
        "the roll's continuation threw " + e.toString() + " — the roll is dead and will "
                + "not retry itself; check the daemon log for the stack trace, then open() "
                + "a fresh rollover request"));

So fleet_handover{action:"status", token} would have answered FAILED with the herdr message in it. The status surface is not the gap.

What this does not change. The successor still gets nothing. It has no token, so it cannot ask for that status, and nothing tells it to look. The FAILED status is readable only by a caller that holds the token — and after a successful pane kill, that caller no longer exists. So the recorded outcome helps an operator who goes looking; it does not help the session that woke up blind.

What this changes in the fix list. Drop the third bullet ("wrap the send so a failure records a RollStatus") — already done. The remaining two stand:

  • retry agents.send on agent_not_ready for a bounded window;
  • make the pre-send gate read the same readiness agent.send requires, or let the retry be the gate.

And one bullet gets more important, because it is now the only thing that reaches the successor: a route to the handover that does not depend on one paste landing at the right moment.

A more specific outcome name than FAILED would still help whoever reads the log later, since FAILED covers every thrown path. That is a nice-to-have, not the defect.

**Correction to my own issue text. One claim above is wrong.** I wrote: "no `RollStatus` is recorded for this path, so `fleet_handover{action:"status"}` cannot report it either — the only record is this `WARN` line." That is false. I had not read the catch when I wrote it. `LeadRollover.runRollover` (`LeadRollover.java:647-655`) catches `RuntimeException` and does record an outcome: ```java outcomes.put(p.token(), new RollStatus(RollState.FAILED, "the roll's continuation threw " + e.toString() + " — the roll is dead and will " + "not retry itself; check the daemon log for the stack trace, then open() " + "a fresh rollover request")); ``` So `fleet_handover{action:"status", token}` would have answered `FAILED` with the herdr message in it. The status surface is not the gap. **What this does not change.** The successor still gets nothing. It has no token, so it cannot ask for that status, and nothing tells it to look. The `FAILED` status is readable only by a caller that holds the token — and after a successful pane kill, that caller no longer exists. So the recorded outcome helps an operator who goes looking; it does not help the session that woke up blind. **What this changes in the fix list.** Drop the third bullet ("wrap the send so a failure records a `RollStatus`") — already done. The remaining two stand: - retry `agents.send` on `agent_not_ready` for a bounded window; - make the pre-send gate read the same readiness `agent.send` requires, or let the retry be the gate. And one bullet gets more important, because it is now the only thing that reaches the successor: a route to the handover that does not depend on one paste landing at the right moment. A more specific outcome name than `FAILED` would still help whoever reads the log later, since `FAILED` covers every thrown path. That is a nice-to-have, not the defect.
Author
Owner

Lead review of PR #798. One required change, and one concern I checked and dismissed.

Required: relaunchReadySeconds now bounds THREE waits, and its javadoc still says two

FleetConfig.java:1469 reads:

@param relaunchReadySeconds default 45 — bound on EACH of two separate waits that run after
               the old lead's pane has been torn down and a fresh one launched: first,
               for the fresh pane itself to reach a real turn boundary ...
               second, for the fresh terminal to show up as a recognised lead ...

The PR adds a third consumer of the same value — sendBootstrapWithRetry(..., cfg.relaunchReadySeconds()) — and does not touch that javadoc. So the documented worst case moves from about 2 × 45s to about 3 × 45s, and the parameter's own contract now understates it by 45 seconds.

Reusing the existing key instead of adding a new one was the right call and is what I asked for. The cost of that choice is that this javadoc is the only place a reader learns what the key bounds, so it has to name all three. Please update it to say three waits and describe the new one in the same style as the other two. Nothing else in the diff needs to change for this.

That javadoc is a config contract for an operator, so it belongs in main source rather than in the commit message.

Checked and NOT a finding: the retry loop making zero attempts

sendBootstrapWithRetry is a while (now < deadline) loop, so a deadline already in the past would return sent=false having never called agents.send once — strictly worse than the single bare call it replaces. I went looking for a way to reach that.

It is not reachable. FleetConfig.java:1501 clamps the value:

relaunchReadySeconds = (relaunchReadySeconds == null || relaunchReadySeconds <= 0) ? 45 : ...

So the key can never be zero or negative, and in production nowMillis is the real clock, which cannot already be 45 seconds past a deadline computed one statement earlier. A controlled clock in a test could produce it, but a test that sets up that state is asserting about something that cannot happen. A do/while would make the single attempt structural rather than depend on that clamp, which I would mildly prefer, but I am not asking for it: it changes nothing observable and the clamp is the real guarantee.

What I verified myself

  • I read the whole main-source diff. 49 insertions, under the fan-out threshold, so no reviewer was spawned.
  • Retrying only on e.code().equals("agent_not_ready") and rethrowing every other HerdrException is correct, and reading a structured code rather than matching the message text is the right instinct — a message can be rewrapped, a code cannot.
  • The early return on a failed bootstrap now pre-empts the RELAUNCH_NOT_RECOGNISED branch, so a roll that both fails to bootstrap and is never recognised reports BOOTSTRAP_NEVER_SENT. That is the more actionable of the two. Deliberate and fine, but worth stating because it changes which outcome an operator sees in the combined case.
  • Dropping the word "five" from the RollStatus javadoc rather than changing it to "six" is the right fix for a count that would rot again.

Mine to do after merge

.claude/skills/handover/SKILL.md needs a BOOTSTRAP_NEVER_SENT row in its outcome table, and its sentence "the worst case there is about twice that number" becomes three times. The worker correctly flagged this as outside its unit; I own that file.

Lead review of PR #798. One required change, and one concern I checked and dismissed. ## Required: `relaunchReadySeconds` now bounds THREE waits, and its javadoc still says two `FleetConfig.java:1469` reads: ``` @param relaunchReadySeconds default 45 — bound on EACH of two separate waits that run after the old lead's pane has been torn down and a fresh one launched: first, for the fresh pane itself to reach a real turn boundary ... second, for the fresh terminal to show up as a recognised lead ... ``` The PR adds a third consumer of the same value — `sendBootstrapWithRetry(..., cfg.relaunchReadySeconds())` — and does not touch that javadoc. So the documented worst case moves from about 2 × 45s to about 3 × 45s, and the parameter's own contract now understates it by 45 seconds. Reusing the existing key instead of adding a new one was the right call and is what I asked for. The cost of that choice is that this javadoc is the only place a reader learns what the key bounds, so it has to name all three. Please update it to say three waits and describe the new one in the same style as the other two. Nothing else in the diff needs to change for this. That javadoc is a config contract for an operator, so it belongs in main source rather than in the commit message. ## Checked and NOT a finding: the retry loop making zero attempts `sendBootstrapWithRetry` is a `while (now < deadline)` loop, so a deadline already in the past would return `sent=false` having never called `agents.send` once — strictly worse than the single bare call it replaces. I went looking for a way to reach that. It is not reachable. `FleetConfig.java:1501` clamps the value: ```java relaunchReadySeconds = (relaunchReadySeconds == null || relaunchReadySeconds <= 0) ? 45 : ... ``` So the key can never be zero or negative, and in production `nowMillis` is the real clock, which cannot already be 45 seconds past a deadline computed one statement earlier. A controlled clock in a test could produce it, but a test that sets up that state is asserting about something that cannot happen. A `do/while` would make the single attempt structural rather than depend on that clamp, which I would mildly prefer, but I am not asking for it: it changes nothing observable and the clamp is the real guarantee. ## What I verified myself - I read the whole main-source diff. 49 insertions, under the fan-out threshold, so no reviewer was spawned. - Retrying only on `e.code().equals("agent_not_ready")` and rethrowing every other `HerdrException` is correct, and reading a structured code rather than matching the message text is the right instinct — a message can be rewrapped, a code cannot. - The early return on a failed bootstrap now pre-empts the `RELAUNCH_NOT_RECOGNISED` branch, so a roll that both fails to bootstrap and is never recognised reports `BOOTSTRAP_NEVER_SENT`. That is the more actionable of the two. Deliberate and fine, but worth stating because it changes which outcome an operator sees in the combined case. - Dropping the word "five" from the `RollStatus` javadoc rather than changing it to "six" is the right fix for a count that would rot again. ## Mine to do after merge `.claude/skills/handover/SKILL.md` needs a `BOOTSTRAP_NEVER_SENT` row in its outcome table, and its sentence "the worst case there is about twice that number" becomes three times. The worker correctly flagged this as outside its unit; I own that file.
Author
Owner

Fixed and merged locally as b696c31 (PR #798 closed, Gitea does not see a local merge). a3eeace follows with the doc change.

What shipped:

  • LeadRollover.sendBootstrapWithRetry retries only the herdr agent_not_ready refusal, bounded by relaunchReadySeconds. Every other HerdrException still propagates to runRollover's catch and becomes FAILED.
  • A refusal that lasts the whole bound now ends the roll with a new RollState.BOOTSTRAP_NEVER_SENT, so fleet_handover{action:"status", token} can report it. Before, the only record was one WARN line.
  • FleetConfig's javadoc for relaunchReadySeconds said it bounds two waits. It is three now, and the old text also said only a timeout on the first wait withholds bootstrapText. Both corrected.
  • .claude/skills/handover/SKILL.md gains the BOOTSTRAP_NEVER_SENT row, and its "about twice that number" worst case is now three times.

Verified at 027d413 in a throwaway worktree: 2210 tests, 0 failures, 0 errors, and the install wrote ~/.m2/repository/dev/ltms/fleetd metadata at 19:22, so the build completed.

Correcting my own earlier comment on this issue: I wrote that no RollStatus is recorded when the continuation throws. That was wrong and I had not read the code. runRollover's catch already records FAILED. The real gap was narrower — a successful continuation that simply lost the send — and that is what this fix closes.

Still not fixed, and worth a separate issue if it recurs: a successor cannot distinguish a failed roll from a cold start, because nothing in its own environment says it is a successor. The working rule for now is that a fresh lead in the fleet space reads .handover/HANDOVER.md and greps fleetd.out for lead-rollover: before assuming it is new. The retry makes that rarer; it does not make the successor self-aware.

Fixed and merged locally as `b696c31` (PR #798 closed, Gitea does not see a local merge). `a3eeace` follows with the doc change. What shipped: - `LeadRollover.sendBootstrapWithRetry` retries **only** the herdr `agent_not_ready` refusal, bounded by `relaunchReadySeconds`. Every other `HerdrException` still propagates to `runRollover`'s catch and becomes `FAILED`. - A refusal that lasts the whole bound now ends the roll with a new `RollState.BOOTSTRAP_NEVER_SENT`, so `fleet_handover{action:"status", token}` can report it. Before, the only record was one WARN line. - `FleetConfig`'s javadoc for `relaunchReadySeconds` said it bounds **two** waits. It is three now, and the old text also said only a timeout on the first wait withholds `bootstrapText`. Both corrected. - `.claude/skills/handover/SKILL.md` gains the `BOOTSTRAP_NEVER_SENT` row, and its "about twice that number" worst case is now three times. Verified at `027d413` in a throwaway worktree: **2210 tests, 0 failures, 0 errors**, and the install wrote `~/.m2/repository/dev/ltms/fleetd` metadata at 19:22, so the build completed. Correcting my own earlier comment on this issue: I wrote that no `RollStatus` is recorded when the continuation throws. That was wrong and I had not read the code. `runRollover`'s catch already records `FAILED`. The real gap was narrower — a *successful* continuation that simply lost the send — and that is what this fix closes. Still not fixed, and worth a separate issue if it recurs: a successor cannot distinguish a failed roll from a cold start, because nothing in its own environment says it is a successor. The working rule for now is that a fresh lead in the `fleet` space reads `.handover/HANDOVER.md` and greps `fleetd.out` for `lead-rollover:` before assuming it is new. The retry makes that rarer; it does not make the successor self-aware.
Sign in to join this conversation.
1 Participants
Notifications
Due Date
No due date set.
Dependencies

No dependencies set.

Reference: fleet/fleetd#796