fleetd #726 unit 2: replace /clear-based lead rollover with a real process restart
CI / shell-tests (pull_request) Failing after 6s
CI / contract (pull_request) Successful in 59s
CI / build (pull_request) Failing after 1m56s

LeadRollover's deferred continuation now ends the old lead's pane, relaunches
a fresh one, and bootstraps it, instead of sending /clear into the same
process. The relaunch step runs two separate bounded waits instead of one
combined check: a readiness wait (the fresh pane reaches a real turn
boundary) is the safety gate and withholds bootstrapText on timeout
(RELAUNCH_NEVER_READY); a recognition wait (the fresh terminal shows up in
the live-lead map) is bookkeeping only, so a timeout there still lets
bootstrapText go out (RELAUNCH_NOT_RECOGNISED). clearSettleSeconds is
retired in favor of relaunchReadySeconds (default 45), which bounds both
waits. Updates FleetConfig/FleetMcp operator-facing text to match.
This commit is contained in:
Dai Ha
2026-10-04 20:44:42 +02:00
parent a332dfdb2c
commit 4ffe49f3bb
13 changed files with 1040 additions and 946 deletions
+22 -17
View File
@@ -99,17 +99,17 @@ bind:
# contextHighNudge: false
# Lead rollover (fleetd #480): replace a lead session that has decided it is ready to be replaced,
# without an operator doing it by hand. A lead writes a handover file, then asks fleetd to clear its
# own pane and bootstrap a fresh session against that file.
# without an operator doing it by hand. A lead writes a handover file, then asks fleetd to end its
# own pane, launch a fresh one, and bootstrap that fresh session against the handover file.
#
# Opt-in on purpose — it clears the lead's own pane on request, so upgrading the daemon must never
# acquire that ability for you. Absent block = feature off, and nothing is constructed at all. Even
# once present, nothing but an explicit confirm() call — one that passes every check — can ever
# cause a /clear: there is no recurring timer, heartbeat or scheduler anywhere in this feature that
# fires one on its own initiative. confirm() itself is called FROM the calling lead's own turn, so
# it cannot clear the pane inline (that pane is still WORKING); instead it schedules a one-shot
# Opt-in on purpose — it tears down the lead's own pane on request, so upgrading the daemon must
# never acquire that ability for you. Absent block = feature off, and nothing is constructed at all.
# Even once present, nothing but an explicit confirm() call — one that passes every check — can ever
# tear a pane down: there is no recurring timer, heartbeat or scheduler anywhere in this feature that
# fires one on its own initiative. confirm() itself is called FROM the calling lead's own turn, so it
# cannot act on the pane inline (that pane is still WORKING); instead it schedules a one-shot
# continuation that waits for the SAME confirm() call's turn to end, then does the actual work. See
# dev.ltms.fleet.lead.LeadRollover's class javadoc for the exact order (fleetd #480 correction).
# dev.ltms.fleet.lead.LeadRollover's class javadoc for the exact order.
#
# handoverPath: REQUIRED when this block is present — where the handover file a fresh lead session
# reads must live. No default (an operator-specific path); a present block with no
@@ -118,25 +118,30 @@ bind:
# working directory when that lead has none configured) — never against whatever
# directory the daemon process happens to have been started in. An absolute path is
# used unchanged. Prefer an absolute path if the daemon and the lead's pane might not
# share a working directory (fleetd #480 follow-up).
# share a working directory.
# requireOperatorConfirm: true # default true — confirm() refuses unless the caller also passes
# # operatorConfirmed: true
# maxDocAgeSeconds: 3600 # default 3600 — refuse a handover file older than this
# turnSettleSeconds: 20 # default 20 — how long the deferred roll waits for the CALLING
# # lead's own turn to end (its pane to report injectable again)
# # before sending /clear at all. If this elapses, /clear is NEVER
# # sent — a lead that never goes idle is still doing real work.
# clearSettleSeconds: 20 # default 20 — how long to wait for the pane to become injectable
# # again AFTER /clear before giving up (never sends bootstrapText
# # if this elapses). A separate, second wait from turnSettleSeconds.
# # before tearing the old pane down at all. If this elapses, nothing
# # is torn down — a lead that never goes idle is still doing real
# # work.
# relaunchReadySeconds: 45 # default 45 — bounds two later waits, after the old pane is gone
# # and a fresh one has been launched: first, for the fresh pane to
# # reach a real turn boundary (never sends bootstrapText if THIS one
# # elapses); second, for the new terminal to be recognised as this
# # lead (bootstrapText is sent either way once the first wait
# # passes). A separate, later pair of waits from turnSettleSeconds.
# bootstrapText: "..." # default names the RESOLVED (absolute) handoverPath — sent to
# # the lead once its pane settles after /clear
# # the freshly relaunched lead's pane once it reaches a real turn
# # boundary
# leadRollover:
# handoverPath: /path/to/handover.md
# requireOperatorConfirm: true
# maxDocAgeSeconds: 3600
# turnSettleSeconds: 20
# clearSettleSeconds: 20
# relaunchReadySeconds: 45
# bootstrapText: "Fresh lead session: read the handover file and carry on."
# Fleet health detection is dormant unless enabled (CB-573). It reads one whole-fleet agent list
@@ -942,21 +942,27 @@ public final class Fleetd {
* cfg.leadHeartbeat()}
* @param leadAgents the {@link AgentControl} instance that reaches the LEAD's pane (not
* {@code memberAgents}), normally {@code router.leadAgents()}
* @param leadSpaces the {@link WorkspaceControl} instance that reaches the LEAD's
* workspace, normally {@code router.leadSpaces()} — used to tear down
* a rolled lead's old pane and confirm it is gone
* @param launcher starts the fresh lead a roll relaunches once the old one is gone
* @param config the live {@link ConfigRef}, captured only inside the returned
* supplier and the workspace lookup — never dereferenced here
* supplier and the two lookups below — never dereferenced here
* @param liveLeadTerminals terminal id → lead NAME for every CURRENTLY recognised lead, normally
* the same {@code leads} supplier {@code main} already builds for
* {@code HerdrRouter}/{@link #leadSeatLookup} — never a value snapshot
* @return a constructed {@link LeadRollover}, or {@code null} when {@code leadRollover:} is
* absent from the startup config
*/
static LeadRollover leadRollover(FleetConfig cfg, AgentControl leadAgents, ConfigRef config,
static LeadRollover leadRollover(FleetConfig cfg, AgentControl leadAgents,
WorkspaceControl leadSpaces, LeadLauncher launcher, ConfigRef config,
Supplier<Map<String, String>> liveLeadTerminals) {
if (cfg.leadRollover() == null) {
return null;
}
Function<String, String> leadNameForTerminal = terminal -> liveLeadTerminals.get().get(terminal);
Function<String, String> leadWorkspace = terminal -> {
String leadName = liveLeadTerminals.get().get(terminal);
String leadName = leadNameForTerminal.apply(terminal);
if (leadName == null) {
return null;
}
@@ -964,7 +970,8 @@ public final class Fleetd {
FleetConfig.Leader leader = fleet == null ? null : fleet.leaders().get(leadName);
return leader == null ? null : leader.cwd();
};
return new LeadRollover(leadAgents, () -> config.get().leadRollover(), leadWorkspace);
return new LeadRollover(leadAgents, leadSpaces, launcher, () -> config.get().leadRollover(),
leadWorkspace, leadNameForTerminal, liveLeadTerminals);
}
/**
@@ -297,11 +297,15 @@ final class FleetdAssembly {
leadsRef.set(leads);
collaboratorTerminalsRef.set(collaboratorTerminals);
// Constructed unconditionally — it is cheap and side-effect free — so a LeadRollover built
// below can relaunch a lead even on a boot where herdr was down for the ensureLeads() call.
LeadLauncher leadLauncher = new LeadLauncher(router.leadAgents(), router.leadSpaces(), cfg);
// CB-558: start any declared lead that is not already running. After the scanner is built,
// and only when herdr answered — the launcher's whole safety property is that it can count
// live leads first, and must never guess and risk a second orchestrator.
if (herdrUp && !leaders.isEmpty()) {
int launched = new LeadLauncher(router.leadAgents(), router.leadSpaces(), cfg).ensureLeads();
int launched = leadLauncher.ensureLeads();
if (launched > 0) {
log.info("lead auto-launch: {} lead(s) started", launched);
}
@@ -440,7 +444,8 @@ final class FleetdAssembly {
heartbeatScheduler.shutdownNow();
}
// fleetd #480: lead rollover. Opt-in; absent `leadRollover:` this is never constructed.
LeadRollover leadRollover = Fleetd.leadRollover(cfg, router.leadAgents(), config, leads);
LeadRollover leadRollover = Fleetd.leadRollover(cfg, router.leadAgents(), router.leadSpaces(),
leadLauncher, config, leads);
MessageService messages = new MessageService(router, injector, rendezvous, replyInbox,
pushLoop, metrics);
@@ -67,7 +67,7 @@ import java.util.function.Supplier;
* {@code models:} above: {@code dev.ltms.fleet.lead.LeadRollover} holds a
* {@code Supplier<FleetConfig.LeadRollover>} (the same {@code () -> config.get().x()} shape)
* and reads {@code handoverPath}/{@code requireOperatorConfirm}/{@code maxDocAgeSeconds}/
* {@code turnSettleSeconds}/{@code clearSettleSeconds}/{@code bootstrapText} fresh on every
* {@code turnSettleSeconds}/{@code relaunchReadySeconds}/{@code bootstrapText} fresh on every
* {@code open()}/{@code confirm()} call (and on the deferred post-{@code confirm()}
* continuation fleetd #480's correction added — see {@code LeadRollover}'s class doc) rather
* than capturing them into fields at construction — unlike its closest
@@ -1416,17 +1416,16 @@ public record FleetConfig(
* before anything exists to call — the same fact already true of adding a brand-new
* {@code profiles:} entry.
*
* <p><strong>{@code turnSettleSeconds} (fleetd #480 correction):</strong> {@code confirm()} is
* called FROM the calling lead's own turn, so its pane is still {@code WORKING} the instant
* {@code confirm()} validates every gate and schedules the roll. {@code
* dev.ltms.fleet.lead.LeadRollover}'s deferred continuation waits up to this many seconds for
* that SAME pane to report {@code IDLE} or {@code DONE} — i.e. for the calling turn to actually
* end — before it sends {@code /clear} at all. {@code BLOCKED} does not count: that is a live
* turn merely paused, not one that has finished. If that wait times out, no {@code /clear} is
* ever sent: a lead that never goes idle is still doing real work, and clearing it would
* destroy live context. This is a separate wait from {@code clearSettleSeconds} below, which
* bounds the SECOND wait, for the pane to reach {@code IDLE} or {@code DONE} again AFTER
* {@code /clear} has already gone out.
* <p><strong>{@code turnSettleSeconds}:</strong> {@code confirm()} is called FROM the calling
* lead's own turn, so its pane is still {@code WORKING} the instant {@code confirm()} validates
* every gate and schedules the roll. {@code dev.ltms.fleet.lead.LeadRollover}'s deferred
* continuation waits up to this many seconds for that SAME pane to report {@code IDLE} or
* {@code DONE} — i.e. for the calling turn to actually end — before it ends the old pane's
* process at all. {@code BLOCKED} does not count: that is a live turn merely paused, not one
* that has finished. If that wait times out, the old pane is never touched: a lead that never
* goes idle is still doing real work, and the roll ends that pane's whole process — there is no
* way back from this once it runs, so this wait is the only thing standing between "still
* working" and "gone".
*
* @param handoverPath required when this block is present — where the handover file a fresh
* lead session reads must live. There is no sane non-null default for an
@@ -1447,36 +1446,49 @@ public record FleetConfig(
* attempt can never be mistaken for a fresh one.
* @param turnSettleSeconds default 300 — bound on how long the deferred roll waits for the
* CALLING lead's own turn to end (its pane to report {@code IDLE} or
* {@code DONE}) before sending {@code /clear} at all. See the paragraph
* above.
* @param clearSettleSeconds default 20 — bound on how long to wait for the lead's pane to
* report {@code IDLE} or {@code DONE} again after {@code /clear} before
* giving up. A roll that times out here never sends {@code bootstrapText}.
* {@code DONE}) before ending that pane's process at all. See the
* paragraph above.
* @param relaunchReadySeconds default 45 — bound on EACH of two separate waits that run after
* the old lead's pane has been torn down and a fresh one launched: first,
* for the fresh pane itself to reach a real turn boundary ({@code IDLE} or
* {@code DONE}, never merely {@code BLOCKED}) — the safety gate, since
* typing into a pane that has not finished booting loses the keystrokes;
* second, for the fresh terminal to show up as a recognised lead, which is
* bookkeeping rather than a safety gate, so a timeout on this second wait
* does not withhold {@code bootstrapText} — it is sent once the pane is
* ready regardless. Recognition comes from the same periodically-refreshed
* scan {@code LeadTabScanner} already keeps ({@code scanIntervalSeconds},
* 10s live), so a budget has to clear more than one scan interval to leave
* any real margin for the CLI's own boot time; 20 was rejected for exactly
* that reason — at a 10s scan interval it only buys two scans. 45 buys
* roughly four. Only a timeout on the FIRST wait (the pane never becomes
* ready) withholds {@code bootstrapText}.
* @param bootstrapText default a sentence naming the RESOLVED handover path — sent to the
* lead's pane once it settles after {@code /clear}, telling the fresh
* session where to read the handover and carry on. Left {@code null} here
* when the operator configures none: the default sentence cannot be built
* at construction time because it must name the path AFTER {@code
* dev.ltms.fleet.lead.LeadRollover#open} has resolved a relative {@code
* handoverPath} against the calling lead's workspace, which this record has
* no way to know — see {@link #bootstrapTextFor(String)}.
* fresh lead's pane once it reaches a real turn boundary after relaunch,
* telling the fresh session where to read the handover and carry on. Left
* {@code null} here when the operator configures none: the default sentence
* cannot be built at construction time because it must name the path AFTER
* {@code dev.ltms.fleet.lead.LeadRollover#open} has resolved a relative
* {@code handoverPath} against the calling lead's workspace, which this
* record has no way to know — see {@link #bootstrapTextFor(String)}.
*/
@JsonIgnoreProperties(ignoreUnknown = true)
public record LeadRollover(String handoverPath, Boolean requireOperatorConfirm,
Integer maxDocAgeSeconds, Integer turnSettleSeconds,
Integer clearSettleSeconds, String bootstrapText) {
Integer relaunchReadySeconds, String bootstrapText) {
public LeadRollover {
requireOperatorConfirm = requireOperatorConfirm == null || requireOperatorConfirm;
maxDocAgeSeconds = (maxDocAgeSeconds == null || maxDocAgeSeconds <= 0) ? 3600 : maxDocAgeSeconds;
turnSettleSeconds = (turnSettleSeconds == null || turnSettleSeconds <= 0) ? 300 : turnSettleSeconds;
clearSettleSeconds = (clearSettleSeconds == null || clearSettleSeconds <= 0) ? 20 : clearSettleSeconds;
relaunchReadySeconds = (relaunchReadySeconds == null || relaunchReadySeconds <= 0)
? 45 : relaunchReadySeconds;
bootstrapText = (bootstrapText == null || bootstrapText.isBlank()) ? null : bootstrapText;
}
/**
* The text actually sent to the lead's pane once it settles after {@code /clear}: the
* operator's configured {@link #bootstrapText} when one is set, otherwise the default
* sentence built from {@code resolvedHandoverPath}.
* The text actually sent to the fresh lead's pane once it reaches a real turn boundary
* after relaunch: the operator's configured {@link #bootstrapText} when one is set,
* otherwise the default sentence built from {@code resolvedHandoverPath}.
*
* @param resolvedHandoverPath the ABSOLUTE path {@code dev.ltms.fleet.lead.LeadRollover
* #open} already resolved — never the raw configured {@link
@@ -1919,6 +1931,7 @@ public record FleetConfig(
rejectNegativeMaxLoad(yaml);
rejectAutoCompactWindowOutOfRange(yaml);
warnConflictingAutoCompactWindows(yaml);
warnRetiredClearSettleSecondsKey(yaml);
rejectMalformedProfilePatterns(yaml);
rejectUnknownKind(yaml);
rejectUnknownAuthMode(yaml);
@@ -2342,6 +2355,33 @@ public record FleetConfig(
names, String.join(", ", detail));
}
/**
* Warn when a {@code leadRollover:} block still sets the retired {@code clearSettleSeconds}
* key. {@link LeadRollover} carries {@code @JsonIgnoreProperties(ignoreUnknown = true)} and no
* longer declares that component, so Jackson drops it with no signal of its own — this raw-YAML
* check is the only place an operator's now-inert setting is reported at all; by the time a
* {@link LeadRollover} instance exists to run a validator against, the key is already gone.
*
* @param yaml the raw config text
*/
static void warnRetiredClearSettleSecondsKey(String yaml) {
Map<?, ?> raw;
try {
raw = YAML.readValue(yaml, Map.class);
} catch (IOException | IllegalArgumentException e) {
return;
}
if (raw == null || !(raw.get("leadRollover") instanceof Map<?, ?> leadRollover)) {
return;
}
if (leadRollover.containsKey("clearSettleSeconds")) {
log.warn("leadRollover.clearSettleSeconds is retired and no longer read. Set "
+ "leadRollover.relaunchReadySeconds instead: it bounds how long to wait, after "
+ "a lead is relaunched, for its pane to become ready and then for it to be "
+ "recognised as a lead. Remove clearSettleSeconds from fleetd.yaml.");
}
}
/**
* Reject a profile whose {@code errorPattern} (fleetd #201 Unit 5) or {@code exhaustedPattern}
* (CB-578 stage A) is not a valid Java regex, naming the profile, the key, and the parser's own
@@ -1,8 +1,11 @@
package dev.ltms.fleet.lead;
import dev.ltms.fleet.config.FleetConfig;
import dev.ltms.fleet.herdr.Agent;
import dev.ltms.fleet.herdr.AgentControl;
import dev.ltms.fleet.herdr.AgentStatus;
import dev.ltms.fleet.herdr.HerdrException;
import dev.ltms.fleet.herdr.WorkspaceControl;
import org.slf4j.Logger;
import org.slf4j.LoggerFactory;
@@ -24,60 +27,69 @@ import java.util.function.Supplier;
/**
* fleetd #480: replace a lead session that has decided it is ready to be rolled over, without an
* operator doing it by hand. A lead writes a handover file, calls {@link #open}, and then — once
* every gate ({@link #confirm}'s own checks) has passed — a deferred, single-shot continuation
* clears the lead's own pane and bootstraps a fresh session against that file.
* every gate ({@link #confirm}'s own checks) has passed — a deferred, single-shot continuation ends
* the lead's own pane, launches a fresh one, and bootstraps that fresh session against the file.
*
* <p>This is the executor behind the {@code fleet_handover} MCP tool ({@code
* dev.ltms.fleet.mcp.FleetMcp#handover}), which drives {@link #open}, {@link #confirm}, {@link
* #cancel}, and {@link #status} from a tool call — wired in fleetd #480 Unit C. <strong>An earlier
* version of this paragraph said nothing called this class at all; that stopped being true once
* that unit landed, and this correction exists so the javadoc does not go on claiming it.</strong>
* #cancel}, and {@link #status} from a tool call.
*
* <p><strong>{@code confirm()} cannot roll inline — a fleetd #480 correction.</strong> The first
* version of this class called {@code agents.send(lead, "/clear")} directly from inside {@code
* confirm()}, then polled for the pane to become injectable again. That is wrong, because {@code
* confirm()} is called BY the lead, FROM the lead's own turn: the lead's pane is {@code WORKING}
* for the whole duration of that call and cannot possibly report injectable until {@code confirm()}
* itself returns. The poll always timed out — but only after the {@code /clear} had already been
* sent and queued in the pane, where it fired the instant the turn ended anyway. The result was the
* worst outcome this feature can produce: a silently destroyed lead context with no fresh session
* ever started, and a refusal return value that claimed nothing had happened.
*
* <p>The fix: {@link #confirm} validates every gate, then does no I/O against the lead's own pane
* at all — it only records that the request is approved and hands a one-shot continuation to
* {@code continuationRunner} before returning. That continuation is what actually touches the pane,
* once the calling turn has ended, in this order:
* <p><strong>{@code confirm()} cannot roll inline.</strong> {@code confirm()} is called BY the
* lead, FROM the lead's own turn: the lead's pane is {@code WORKING} for the whole duration of that
* call and cannot possibly report a real turn boundary until {@code confirm()} itself returns. So
* {@link #confirm} validates every gate, then does no I/O against the lead's own pane at all — it
* only records that the request is approved and hands a one-shot continuation to {@code
* continuationRunner} before returning. That continuation is what actually touches the pane, once
* the calling turn has ended, in this order:
* <ol>
* <li>wait for the lead's own pane to report a real turn boundary — {@code IDLE} or {@code
* DONE}, never merely {@code BLOCKED} — i.e. wait for the very {@code confirm()} call that
* approved this roll to finish its turn — bounded by {@code turnSettleSeconds}. <strong>If
* this never happens, nothing else in this list runs: no {@code /clear} is ever sent.</strong>
* A lead that never goes idle is a lead still doing real work, and clearing it would throw
* away live context — exactly the failure this correction exists to prevent.</li>
* <li>{@code agents.send(lead, "/clear")}</li>
* <li>wait for {@code /clear} to be picked up and settle, bounded by {@code clearSettleSeconds}
* (fleetd #489: no longer a plain re-check of the same boundary — {@code /clear} starts no
* turn of its own, so this instead nudges the submit keystroke while no pickup has been seen,
* then waits for a real {@code WORKING} → {@code IDLE}/{@code DONE} boundary once one has;
* see {@link #waitForClearPickupAndSettle})</li>
* <li>{@code agents.send(lead, cfg.bootstrapTextFor(p.handoverPath()))}</li>
* this never happens, nothing else in this list runs: the old pane is never touched.</strong>
* A lead that never goes idle is a lead still doing real work, and tearing it down would throw
* away live context.</li>
* <li>capture the old pane id (and, through it, the old tab) from {@link AgentControl#get}, with
* a bounded retry — the terminal-to-pane lookup it goes through can itself report a genuinely
* live agent as not found (see {@code AgentControl#agentCall}'s own re-resolve-once
* behaviour), and one false negative here must not abort an otherwise-healthy roll. Neither id
* is ever re-resolved from the terminal again after this — once the pane below is closed there
* is nothing left to resolve it from.</li>
* <li>resolve the lead's configured name from its terminal, for the relaunch step below.</li>
* <li>end the old session: close the pane (an already-gone pane counts as success; any other
* failure propagates), then close its tab only when the pane was that tab's sole occupant —
* the same pane-then-tab teardown {@code HerdrPeerLauncher#stop} uses for a member.</li>
* <li>confirm the old pane is actually gone by polling {@link
* dev.ltms.fleet.herdr.WorkspaceControl#locatePane} for a {@code null} result — never {@link
* AgentControl#status}, and never the live-lead terminal map, each of which answers a
* different question. <strong>If the old pane is never confirmed gone, no relaunch is
* attempted</strong> — see {@link RollState#OLD_PANE_NEVER_DIED}.</li>
* <li>launch a fresh lead with {@code LeadLauncher#relaunch}. <strong>If every attempt fails,
* {@code bootstrapText} is never sent</strong> — see {@link RollState#RELAUNCH_FAILED}.</li>
* <li>wait for the fresh pane to reach a real turn boundary ({@code IDLE} or {@code DONE},
* never merely {@code BLOCKED}), bounded by {@code relaunchReadySeconds}. This is the
* safety gate: typing into a pane that has not actually finished booting loses the
* keystrokes. <strong>If the pane never becomes ready, {@code bootstrapText} is never
* sent</strong> — see {@link RollState#RELAUNCH_NEVER_READY}.</li>
* <li>wait for the fresh terminal to be recognised as a live lead — present in the live-lead
* terminal map — bounded by {@code relaunchReadySeconds}. This is bookkeeping, not a
* safety gate: {@code bootstrapText} is sent either way once the pane is ready, whether or
* not this wait itself times out — see {@link RollState#RELAUNCH_NOT_RECOGNISED}.</li>
* <li>{@code agents.send(newTerminal, cfg.bootstrapTextFor(p.handoverPath()))} — sent to the
* FRESH terminal, never the one that was just torn down.</li>
* </ol>
* A {@link #confirm} that returns {@link RollDecision#approved()} therefore means <em>"every gate
* passed and the roll is scheduled"</em>, never <em>"the pane has been cleared"</em> — the pane may
* still be mid-turn, possibly for a long time, when the caller gets that answer back.
* passed and the roll is scheduled"</em>, never <em>"the lead has already been replaced"</em> — the
* old pane may still be mid-turn, possibly for a long time, when the caller gets that answer back.
*
* <p><strong>The safety invariant survives this change, restated precisely.</strong> The ticket
* that first defined this class required "no timer, no scheduler, no background thread" so that
* nothing but an explicit {@link #confirm} call could ever cause a {@code /clear}. That invariant
* is about INITIATIVE, not about synchronicity, and this correction keeps it: {@code
* continuationRunner} launches a single-shot task that exists only because one specific,
* <p><strong>The safety invariant.</strong> "No timer, no scheduler, no background thread" means
* that nothing but an explicit {@link #confirm} call can ever tear a lead's pane down.
* {@code continuationRunner} launches a single-shot task that exists only because one specific,
* already-approved {@link #confirm} call created it — it is not recurring, it is not started at
* construction time or on any schedule, and no two invocations of it ever share state. A recurring
* heartbeat or timer that could decide on its own initiative to roll a pane is still, and will
* always be, absent from this class. <strong>Nothing but an explicit {@link #confirm} call that
* passes every gate can ever cause a {@code /clear} — that call may simply finish its own work
* slightly later than the method return, as a continuation of the same approved request, rather
* than entirely inside the method body.</strong>
* heartbeat or timer that could decide on its own initiative to roll a pane is absent from this
* class. <strong>Nothing but an explicit {@link #confirm} call that passes every gate can ever tear
* a pane down — that call may simply finish its own work slightly later than the method return, as
* a continuation of the same approved request, rather than entirely inside the method body.</strong>
*
* <p><strong>Identity is resolved by the caller, never looked up here — a second fleetd #480
* correction.</strong> The first version resolved the pane to clear via {@code
@@ -115,19 +127,17 @@ public final class LeadRollover {
private static final Logger log = LoggerFactory.getLogger(LeadRollover.class);
/** Poll interval while waiting for the lead's pane to settle after {@code /clear}. */
static final long SETTLE_POLL_MS = 250;
/** Poll interval shared by every bounded wait in this class. */
static final long POLL_INTERVAL_MS = 250;
/**
* How many consecutive not-yet-picked-up polls {@link #waitForClearPickupAndSettle} allows
* before releasing rather than wedging the roll — the same constant and the same
* release-not-wedge choice {@link dev.ltms.fleet.inject.Injector} already makes for its own
* post-turn {@code /clear} housekeeping (fleetd #306). <strong>This bounds the number of
* consecutive polls, not the number of nudges:</strong> the first {@code PICKUP_GRACE_POLLS - 1}
* of those polls each send a nudge, and the {@code PICKUP_GRACE_POLLS}th releases instead of
* nudging again — so 8 polls produce 7 nudges, not 8.
* How long {@link #waitUntilPaneGone} polls {@link WorkspaceControl#locatePane} before giving
* up on ever seeing the old pane disappear. Not configurable: once {@link #endOldSession} has
* closed the pane (and, usually, its tab), herdr dropping the pane from its own bookkeeping is
* expected to show up within one or two polls, not on an operator-tunable timescale the way a
* CLI boot is.
*/
static final int PICKUP_GRACE_POLLS = 8;
static final int PANE_DEATH_TIMEOUT_SECONDS = 10;
/**
* One request opened by {@link #open}, pending its {@link #confirm} (or {@link #cancel}).
@@ -175,10 +185,10 @@ public final class LeadRollover {
/**
* The outcome of a {@link #confirm} call. {@link #approved()} means every gate passed and the
* roll has been handed to a one-shot continuation — <strong>not</strong> that the pane has been
* cleared; the continuation may still be waiting for the calling turn to end when this returns.
* Whether the deferred roll itself later goes on to clear the pane, refuse for never going
* idle, or refuse for never re-settling after {@code /clear} is logged only (see this class's
* roll has been handed to a one-shot continuation — <strong>not</strong> that the lead has
* already been replaced; the continuation may still be waiting for the calling turn to end when
* this returns. Whether the deferred roll itself later goes on to tear the old pane down and
* relaunch the lead, or refuses at any of its own steps, is logged only (see this class's
* javadoc) — there is deliberately no synchronous caller left by that point to hand a result to.
*/
public record RollDecision(boolean accepted, RefusalReason reason, String detail) {
@@ -214,7 +224,7 @@ public final class LeadRollover {
/**
* What is known about one token, right now — the answer {@link #status} gives. Distinguishes
* three terminal outcomes an approved roll can finish with, one in-flight outcome for a roll
* five terminal outcomes an approved roll can finish with, one in-flight outcome for a roll
* that has been approved but has not finished yet, and two answers for a token that names no
* active work at all: still pending confirmation, or nothing known about this token at all.
*/
@@ -237,41 +247,64 @@ public final class LeadRollover {
* #status} could wrongly answer {@link #UNKNOWN} ("nothing was ever requested") for a roll
* that is, in fact, actively running. This is not sticky: the deferred continuation
* overwrites this same entry with a terminal state ({@link #ROLLED}, {@link
* #TURN_NEVER_SETTLED}, {@link #CLEAR_NEVER_SETTLED}, or {@link #FAILED}) once it finishes
* — including by throwing, which fleetd #615's catch in {@link #runRollover} now turns into
* {@link #FAILED} instead of leaving this entry stuck forever.
* #TURN_NEVER_SETTLED}, {@link #OLD_PANE_NEVER_DIED}, {@link #RELAUNCH_FAILED}, {@link
* #RELAUNCH_NOT_RECOGNISED}, or {@link #FAILED}) once it finishes — including by throwing,
* which {@link #runRollover}'s catch turns into {@link #FAILED} instead of leaving this
* entry stuck forever.
*/
IN_PROGRESS,
/**
* {@link #confirm} was approved and the deferred continuation completed the entire roll:
* the calling lead's turn settled, {@code /clear} was sent and settled, and {@code
* bootstrapText} was sent.
* {@link #confirm} was approved and the deferred continuation completed the entire roll: the
* calling lead's turn settled, the old pane was torn down and confirmed gone, a fresh lead
* was launched and recognised, and {@code bootstrapText} was sent to it.
*/
ROLLED,
/**
* {@link #confirm} was approved, but the calling lead's own turn never reached a boundary
* (IDLE or DONE) within {@code turnSettleSeconds} — no {@code /clear} was ever sent, at
* all. This is the branch the fleetd #480 correction exists to make safe, and the one this
* status exists to make VISIBLE: before this, a lead that hit this case had no way to find
* out, and would carry on believing it was about to be replaced. See this class's javadoc.
* (IDLE or DONE) within {@code turnSettleSeconds} — the old pane was never touched at all.
* This is the state that makes a lead's own stuck turn VISIBLE: without it, a lead that hit
* this case would have no way to find out, and would carry on believing it was about to be
* replaced. See this class's javadoc.
*/
TURN_NEVER_SETTLED,
/**
* {@link #confirm} was approved and {@code /clear} was sent, but the pane never re-settled
* within {@code clearSettleSeconds} — {@code bootstrapText} was never sent.
* {@link #confirm} was approved and the calling lead's turn settled, the old pane was closed
* (and its tab, if it was the sole occupant), but {@link
* dev.ltms.fleet.herdr.WorkspaceControl#locatePane} kept reporting it as still present for
* the whole pane-death timeout. No relaunch was ever attempted, and {@code bootstrapText}
* was never sent.
*/
CLEAR_NEVER_SETTLED,
OLD_PANE_NEVER_DIED,
/**
* fleetd #615: the deferred continuation threw a {@link RuntimeException} — most likely a
* {@link dev.ltms.fleet.herdr.HerdrException} out of one of the two unwrapped {@code
* agents.send} calls in {@link #runRollover} — and the continuation thread died with it.
* Before this state existed, that throw left {@link #outcomes} holding {@link #IN_PROGRESS}
* forever, because the production {@code continuationRunner} is a bare virtual thread with
* no uncaught-exception handler and nothing downstream of the throw ever ran to write a
* terminal outcome. {@code detail} names the exception, so a reader has something to act on
* — the same diagnostic style as {@link #TURN_NEVER_SETTLED} and {@link
* #CLEAR_NEVER_SETTLED}. The roll is dead at this point and does not retry itself; a stuck
* lead must {@link #open} a fresh request.
* The old pane was confirmed gone, but {@code LeadLauncher#relaunch} returned {@code null}
* — every launch attempt failed. {@code bootstrapText} was never sent, and no fresh terminal
* exists for this roll to have recognised.
*/
RELAUNCH_FAILED,
/**
* A fresh lead was launched, but its pane never reached a real turn boundary ({@code IDLE}
* or {@code DONE}, never merely {@code BLOCKED}) within {@code relaunchReadySeconds} — the
* CLI never finished booting, or it stayed paused on a startup prompt. {@code bootstrapText}
* was never sent: typing into a pane that is not actually ready to accept input loses the
* keystrokes.
*/
RELAUNCH_NEVER_READY,
/**
* A fresh lead was launched and its pane reached a real turn boundary, so {@code
* bootstrapText} WAS sent to it, but the terminal was never recognised as a live lead —
* present in the live-lead terminal map — within {@code relaunchReadySeconds}. The session
* itself is alive and bootstrapped; only the daemon's own bookkeeping has not caught up, and
* an operator should check why the tab was not recognised.
*/
RELAUNCH_NOT_RECOGNISED,
/**
* The deferred continuation threw a {@link RuntimeException} and the continuation thread
* died with it. Without this state, that throw would leave {@link #outcomes} holding {@link
* #IN_PROGRESS} forever, because the production {@code continuationRunner} is a bare virtual
* thread with no uncaught-exception handler and nothing downstream of the throw ever runs to
* write a terminal outcome. {@code detail} names the exception, so a reader has something to
* act on. The roll is dead at this point and does not retry itself; a stuck lead must
* {@link #open} a fresh request.
*/
FAILED,
/**
@@ -292,6 +325,10 @@ public final class LeadRollover {
public record RollStatus(RollState state, String detail) {}
private final AgentControl agents;
/** Workspace/tab/pane control — used to tear down the old pane and confirm it is gone. */
private final WorkspaceControl spaces;
/** Starts the fresh lead that replaces the one this roll tears down. */
private final LeadLauncher launcher;
private final Supplier<FleetConfig.LeadRollover> configSupplier;
/**
* Terminal id → that lead's configured workspace directory (their {@code
@@ -301,8 +338,21 @@ public final class LeadRollover {
* daemon-cwd bug this parameter exists to fix.
*/
private final Function<String, String> leadWorkspace;
/**
* Terminal id → that lead's configured name under {@code fleet.leaders}, or {@code null} when
* the terminal names no currently-recognised lead. The deferred continuation calls this, on the
* OLD terminal, before tearing it down, so it knows which lead to pass to {@link
* LeadLauncher#relaunch}.
*/
private final Function<String, String> leadNameForTerminal;
/**
* The daemon's current terminal id → lead name map, read fresh on every poll. The deferred
* continuation polls this for the FRESH terminal {@link LeadLauncher#relaunch} returns, to
* learn when that terminal has been recognised as a live lead — see this class's javadoc.
*/
private final Supplier<Map<String, String>> liveLeadTerminals;
private final LongSupplier nowMillis;
private final Runnable settleSleeper;
private final Runnable pollSleeper;
/**
* Launches the post-{@code confirm()} continuation. Production uses a single unstarted virtual
* thread per confirmed request — see this class's javadoc for why that is a single-shot task,
@@ -338,29 +388,40 @@ public final class LeadRollover {
}
});
/** Production constructor — wall clock, real sleep between settle polls, a real virtual thread. */
public LeadRollover(AgentControl agents, Supplier<FleetConfig.LeadRollover> configSupplier,
Function<String, String> leadWorkspace) {
this(agents, configSupplier, leadWorkspace, System::currentTimeMillis,
() -> sleepUninterruptibly(SETTLE_POLL_MS),
/** Production constructor — wall clock, real sleep between polls, a real virtual thread. */
public LeadRollover(AgentControl agents, WorkspaceControl spaces, LeadLauncher launcher,
Supplier<FleetConfig.LeadRollover> configSupplier,
Function<String, String> leadWorkspace,
Function<String, String> leadNameForTerminal,
Supplier<Map<String, String>> liveLeadTerminals) {
this(agents, spaces, launcher, configSupplier, leadWorkspace, leadNameForTerminal,
liveLeadTerminals, System::currentTimeMillis,
() -> sleepUninterruptibly(POLL_INTERVAL_MS),
r -> Thread.ofVirtual().name("lead-rollover-continuation-").start(r));
}
/**
* Full constructor — an injectable wall-clock supplier, settle-poll sleeper, and continuation
* runner, for tests. {@code nowMillis} MUST be a wall-clock source (e.g. {@code
* Full constructor — an injectable wall-clock supplier, poll sleeper, and continuation runner,
* for tests. {@code nowMillis} MUST be a wall-clock source (e.g. {@code
* System.currentTimeMillis()}), never {@code System.nanoTime()}: the freshness check compares
* against a file's modified time, which only a wall clock is comparable to, and {@code
* nanoTime} freezes while the host sleeps (fleetd #386).
* nanoTime} freezes while the host sleeps.
*/
LeadRollover(AgentControl agents, Supplier<FleetConfig.LeadRollover> configSupplier,
Function<String, String> leadWorkspace, LongSupplier nowMillis,
Runnable settleSleeper, Consumer<Runnable> continuationRunner) {
LeadRollover(AgentControl agents, WorkspaceControl spaces, LeadLauncher launcher,
Supplier<FleetConfig.LeadRollover> configSupplier,
Function<String, String> leadWorkspace,
Function<String, String> leadNameForTerminal,
Supplier<Map<String, String>> liveLeadTerminals,
LongSupplier nowMillis, Runnable pollSleeper, Consumer<Runnable> continuationRunner) {
this.agents = agents;
this.spaces = spaces;
this.launcher = launcher;
this.configSupplier = configSupplier;
this.leadWorkspace = leadWorkspace;
this.leadNameForTerminal = leadNameForTerminal;
this.liveLeadTerminals = liveLeadTerminals;
this.nowMillis = nowMillis;
this.settleSleeper = settleSleeper;
this.pollSleeper = pollSleeper;
this.continuationRunner = continuationRunner;
}
@@ -519,8 +580,9 @@ public final class LeadRollover {
// remove-then-put ordering would leave in which the token is in neither map.
outcomes.put(token, new RollStatus(RollState.IN_PROGRESS,
"confirm() approved this roll and handed it to the deferred continuation; it has "
+ "not finished yet — still waiting for the calling turn to settle, for "
+ "/clear to be sent and settle, or for bootstrapText to be sent"));
+ "not finished yet — still waiting for the calling turn to settle, for the "
+ "old pane to be torn down and confirmed gone, for the fresh lead to be "
+ "recognised, or for bootstrapText to be sent"));
pending.remove(token);
log.info("lead-rollover: confirmed token={} lead={} — roll scheduled once the calling turn ends",
token, callerTerminal);
@@ -547,22 +609,19 @@ public final class LeadRollover {
/**
* The single-shot continuation {@link #confirm} hands to {@code continuationRunner}. Runs
* entirely after {@link #confirm} has returned to its caller — see this class's javadoc for the
* four-step order. There is no result to return to by this point, so every outcome is logged
* only.
* full order. There is no result to return to by this point, so every outcome is logged only.
*
* <p><strong>fleetd #615 — the whole body is wrapped in one {@code try}.</strong> The two {@code
* agents.send} calls below are not wrapped individually: {@code send} → {@code agentCall} →
* {@code herdr.call} can throw an unchecked {@link dev.ltms.fleet.herdr.HerdrException} (see
* {@code AgentControl.java}), and the production {@code continuationRunner} is a bare virtual
* thread with no uncaught-exception handler (see this class's public constructor). Before this
* fix, either throw killed the continuation thread silently, leaving the {@link
* RollState#IN_PROGRESS} entry {@link #confirm} wrote at hand-off stuck forever — {@link
* #status} had no way to tell a dead roll from one still genuinely running. The {@code catch}
* below is scoped to the method body rather than to each {@code send} call individually, so it
* also covers anything else added to this continuation later, not just today's two call sites —
* the same reasoning that put the write-a-terminal-outcome step at each of this method's other
* exits (see the {@link RollState#TURN_NEVER_SETTLED} and {@link RollState#CLEAR_NEVER_SETTLED}
* branches below) rather than inside the helpers that detect them.</p>
* <p><strong>The whole body is wrapped in one {@code try}.</strong> Several calls below —
* {@code agents.get}, {@code agents.close}, {@code agents.send} — can throw an unchecked {@link
* dev.ltms.fleet.herdr.HerdrException} (see {@code AgentControl.java}), and the production
* {@code continuationRunner} is a bare virtual thread with no uncaught-exception handler (see
* this class's public constructor). An uncaught throw would kill the continuation thread
* silently, leaving the {@link RollState#IN_PROGRESS} entry {@link #confirm} wrote at hand-off
* stuck forever — {@link #status} would have no way to tell a dead roll from one still
* genuinely running. The {@code catch} below is scoped to the method body rather than to each
* call individually, so it also covers every call in this continuation, not a fixed list of
* call sites — the same reasoning that put the write-a-terminal-outcome step at each of this
* method's other exits rather than inside the helpers that detect them.</p>
*
* <p>Only {@link RuntimeException} is caught, matching the local convention {@link
* #waitUntilAtTurnBoundary} already set around its own {@code agents.status} call — not the
@@ -593,53 +652,237 @@ public final class LeadRollover {
private void runRolloverUnguarded(PendingRollover p, FleetConfig.LeadRollover cfg) {
String lead = p.leadTerminal();
long rollStartMillis = nowMillis.getAsLong();
TurnSettleResult turnResult = waitUntilAtTurnBoundary(lead, cfg.turnSettleSeconds());
if (!turnResult.settled()) {
// fleetd #494 follow-up: this line had the SAME defect as the /clear-timeout line below
// — cfg.turnSettleSeconds() is the CONFIGURED budget, not how long this wait actually
// ran. Print the measured elapsed time alongside it, labelled, exactly like the /clear
// path already does.
log.warn("lead-rollover: pane {} did not reach a turn boundary (IDLE or DONE) after "
+ "confirm() — refusing to send /clear at all; the calling lead's own "
+ "turn is still live and clearing it now would destroy live context "
+ "confirm() — the old pane is never touched; the calling lead's own "
+ "turn is still live and tearing it down now would destroy live context "
+ "(token={}, configured={}s elapsed={}ms)",
lead, p.token(), cfg.turnSettleSeconds(), turnResult.elapsedMillis());
outcomes.put(p.token(), new RollStatus(RollState.TURN_NEVER_SETTLED,
"the calling lead's own turn never reached a boundary (IDLE or DONE) within "
+ "turnSettleSeconds=" + cfg.turnSettleSeconds() + "s (measured elapsed="
+ turnResult.elapsedMillis() + "ms) — no /clear was ever sent. If this "
+ "keeps happening, raise turnSettleSeconds in fleetd.yaml"));
+ turnResult.elapsedMillis() + "ms) — the old pane was never touched. If "
+ "this keeps happening, raise turnSettleSeconds in fleetd.yaml"));
return;
}
// This deliberately bypasses Injector, exactly like ClaudeCodeLauncher#clearContext:
// /clear is housekeeping, not a delegated turn, and routing it through Injector wedges the
// pane forever (see this class's javadoc).
agents.send(lead, "/clear");
ClearSettleResult clearResult = waitForClearPickupAndSettle(lead, cfg.clearSettleSeconds());
if (!clearResult.settled()) {
// fleetd #494: cfg.clearSettleSeconds() is the CONFIGURED budget, not how long the wait
// actually ran — an operator reading only that number wrongly believes it is a measured
// duration. Print the measured elapsed time and nudge count alongside it, each labelled,
// so the two can be compared at a glance.
log.warn("lead-rollover: pane {} did not reach a turn boundary (IDLE or DONE) after "
+ "/clear — NOT sending bootstrapText (token={}, configured={}s "
+ "elapsed={}ms nudges={})",
lead, p.token(), cfg.clearSettleSeconds(), clearResult.elapsedMillis(),
clearResult.nudges());
outcomes.put(p.token(), new RollStatus(RollState.CLEAR_NEVER_SETTLED,
"/clear was sent, but the pane never re-settled within clearSettleSeconds="
+ cfg.clearSettleSeconds() + "s (measured elapsed=" + clearResult.elapsedMillis()
+ "ms, nudges=" + clearResult.nudges() + ") — bootstrapText was never sent"));
// Captured once, here, and never re-resolved from `lead` again below: once the pane is
// closed there is nothing left for a terminal lookup to find.
Agent oldAgent = captureAgentWithRetry(lead);
String oldPaneId = oldAgent.paneId();
String leadName = leadNameForTerminal.apply(lead);
endOldSession(oldPaneId);
DeathResult deathResult = waitUntilPaneGone(oldPaneId);
if (!deathResult.gone()) {
log.warn("lead-rollover: old pane {} for lead {} was never confirmed gone after being "
+ "closed — not attempting a relaunch (token={}, timeout={}s "
+ "elapsed={}ms)",
oldPaneId, lead, p.token(), PANE_DEATH_TIMEOUT_SECONDS, deathResult.elapsedMillis());
outcomes.put(p.token(), new RollStatus(RollState.OLD_PANE_NEVER_DIED,
"the old pane was closed, but locatePane kept reporting it as still present "
+ "after a pane-death timeout=" + PANE_DEATH_TIMEOUT_SECONDS
+ "s (measured elapsed=" + deathResult.elapsedMillis() + "ms) — no "
+ "relaunch was attempted"));
return;
}
agents.send(lead, cfg.bootstrapTextFor(p.handoverPath()));
Agent newAgent = launcher.relaunch(leadName);
if (newAgent == null) {
log.warn("lead-rollover: relaunch of lead '{}' (old terminal {}) failed every attempt "
+ "— bootstrapText was never sent (token={})", leadName, lead, p.token());
outcomes.put(p.token(), new RollStatus(RollState.RELAUNCH_FAILED,
"lead '" + leadName + "' could not be relaunched — every attempt failed; "
+ "bootstrapText was never sent"));
return;
}
ReadinessResult readinessResult = waitUntilPaneReady(newAgent.terminalId(),
cfg.relaunchReadySeconds());
if (!readinessResult.ready()) {
log.warn("lead-rollover: fresh pane for lead '{}' (terminal {}) never reached a real "
+ "turn boundary — bootstrapText was never sent (token={}, configured={}s "
+ "elapsed={}ms)",
leadName, newAgent.terminalId(), p.token(), cfg.relaunchReadySeconds(),
readinessResult.elapsedMillis());
outcomes.put(p.token(), new RollStatus(RollState.RELAUNCH_NEVER_READY,
"fresh terminal " + newAgent.terminalId() + " never reached a real turn "
+ "boundary (IDLE or DONE) within relaunchReadySeconds="
+ cfg.relaunchReadySeconds() + "s (measured elapsed="
+ readinessResult.elapsedMillis() + "ms) — bootstrapText was never "
+ "sent"));
return;
}
IdentityResult identityResult = waitUntilRecognisedAsLead(newAgent.terminalId(),
cfg.relaunchReadySeconds());
agents.send(newAgent.terminalId(), cfg.bootstrapTextFor(p.handoverPath()));
if (!identityResult.ready()) {
log.warn("lead-rollover: fresh terminal {} for lead '{}' is alive and bootstrapped, but "
+ "was never recognised as a live lead — an operator should check why "
+ "the tab was not recognised (token={}, configured={}s elapsed={}ms)",
newAgent.terminalId(), leadName, p.token(), cfg.relaunchReadySeconds(),
identityResult.elapsedMillis());
outcomes.put(p.token(), new RollStatus(RollState.RELAUNCH_NOT_RECOGNISED,
"bootstrapText was sent to fresh terminal " + newAgent.terminalId() + ", but "
+ "it was never recognised as a live lead within relaunchReadySeconds="
+ cfg.relaunchReadySeconds() + "s (measured elapsed="
+ identityResult.elapsedMillis() + "ms) — check why the tab was not "
+ "recognised"));
return;
}
long rollElapsedMillis = nowMillis.getAsLong() - rollStartMillis;
log.info("lead-rollover: rolled token={} lead={} elapsedMs={}", p.token(), lead, rollElapsedMillis);
log.info("lead-rollover: rolled token={} oldLead={} newTerminal={} elapsedMs={}",
p.token(), lead, newAgent.terminalId(), rollElapsedMillis);
outcomes.put(p.token(), new RollStatus(RollState.ROLLED,
"rolled successfully in " + rollElapsedMillis + "ms"));
"rolled successfully in " + rollElapsedMillis + "ms; new terminal="
+ newAgent.terminalId()));
}
/** Attempts {@link #captureAgentWithRetry} makes before letting the failure propagate. */
static final int CAPTURE_RETRIES = 3;
/**
* {@link AgentControl#get} for {@code lead}, retried up to {@link #CAPTURE_RETRIES} times. The
* terminal-to-pane lookup it goes through can report a genuinely live agent as not found (see
* {@code AgentControl#agentCall}'s own re-resolve-once behaviour), and one such false negative
* must not abort an otherwise-healthy roll. The result is captured once by the caller and never
* looked up again — see this class's javadoc.
*
* @throws RuntimeException the last failure, if every attempt fails — {@link #runRollover}'s
* catch turns that into {@link RollState#FAILED}
*/
private Agent captureAgentWithRetry(String lead) {
RuntimeException last = null;
for (int attempt = 1; attempt <= CAPTURE_RETRIES; attempt++) {
try {
return agents.get(lead);
} catch (RuntimeException e) {
last = e;
log.debug("lead-rollover: agents.get({}) failed on attempt {}/{}: {}",
lead, attempt, CAPTURE_RETRIES, e.toString());
if (attempt < CAPTURE_RETRIES) {
pollSleeper.run();
}
}
}
throw last;
}
/**
* End the old lead's session: close its pane, then close its tab only when the pane was that
* tab's sole occupant — the same pane-then-tab teardown {@code HerdrPeerLauncher#stop} uses for
* a member. An already-gone pane counts as success; any other {@code agents.close} failure
* propagates, so a genuinely failed teardown is never reported as done. A failing
* {@code spaces.closeTab} never propagates — by the time it runs the pane is already closed, so
* it is cosmetic tidying, not a real teardown failure.
*/
private void endOldSession(String paneId) {
WorkspaceControl.PaneLocation loc = spaces.locatePane(paneId);
try {
agents.close(paneId);
} catch (HerdrException e) {
if (!isAlreadyGone(e)) {
throw e;
}
log.debug("lead-rollover: pane.close({}) ignored — already gone: {}", paneId, e.getMessage());
}
if (loc != null && loc.tabPaneCount() == 1) {
try {
spaces.closeTab(loc.tabId());
} catch (RuntimeException e) {
log.warn("lead-rollover: tab.close({}) failed — the pane is already torn down, so "
+ "continuing; the tab may need manual cleanup: {}", loc.tabId(), e.getMessage());
}
} else if (loc != null) {
log.debug("lead-rollover: not closing tab {} — it holds {} panes (not a dedicated lead "
+ "tab)", loc.tabId(), loc.tabPaneCount());
}
}
/** True when a herdr error means the target is already gone (safe to treat as done). */
private static boolean isAlreadyGone(HerdrException e) {
return e.code() != null && e.code().endsWith("_not_found");
}
/**
* Poll {@link WorkspaceControl#locatePane} for {@code paneId} until it reports {@code null}
* (the pane is gone) or {@link #PANE_DEATH_TIMEOUT_SECONDS} elapses. Deliberately never calls
* {@link AgentControl#status} and never reads the live-lead terminal map — both answer a
* different question (whether an AGENT is live, not whether this PANE still exists) and
* {@code locatePane} alone catches a {@link HerdrException} from the underlying {@code
* pane.get} and turns it into {@code null} — see this class's javadoc.
*/
private DeathResult waitUntilPaneGone(String paneId) {
long startMillis = nowMillis.getAsLong();
long deadline = startMillis + TimeUnit.SECONDS.toMillis(PANE_DEATH_TIMEOUT_SECONDS);
while (nowMillis.getAsLong() < deadline) {
if (spaces.locatePane(paneId) == null) {
return new DeathResult(true, nowMillis.getAsLong() - startMillis);
}
pollSleeper.run();
}
return new DeathResult(false, nowMillis.getAsLong() - startMillis);
}
/** The measured outcome of {@link #waitUntilPaneGone}. */
private record DeathResult(boolean gone, long elapsedMillis) {}
/**
* Poll until {@code newTerminal}'s own pane reaches a real turn boundary ({@link
* AgentStatus#IDLE} or {@link AgentStatus#DONE}, never merely {@link AgentStatus#BLOCKED}) —
* the same exclusion {@link #waitUntilAtTurnBoundary} applies to the calling lead's own turn,
* applied here to the fresh one, so {@code bootstrapText} is never typed into a pane that has
* not actually finished booting — or {@code readySeconds} elapses. A failed status read
* degrades to "not yet ready" and is retried on the next poll.
*/
private ReadinessResult waitUntilPaneReady(String newTerminal, int readySeconds) {
long startMillis = nowMillis.getAsLong();
long deadline = startMillis + TimeUnit.SECONDS.toMillis(readySeconds);
while (nowMillis.getAsLong() < deadline) {
AgentStatus status;
try {
status = agents.status(newTerminal);
} catch (RuntimeException e) {
log.debug("lead-rollover: status check failed while waiting for {} to be ready: {}",
newTerminal, e.toString());
status = null;
}
if (status == AgentStatus.IDLE || status == AgentStatus.DONE) {
return new ReadinessResult(true, nowMillis.getAsLong() - startMillis);
}
pollSleeper.run();
}
return new ReadinessResult(false, nowMillis.getAsLong() - startMillis);
}
/** The measured outcome of {@link #waitUntilPaneReady}. */
private record ReadinessResult(boolean ready, long elapsedMillis) {}
/**
* Poll until {@code newTerminal} is present in {@link #liveLeadTerminals} or {@code
* readySeconds} elapses. This is bookkeeping, not a safety gate: the pane's own readiness (see
* {@link #waitUntilPaneReady}) is what decides whether {@code bootstrapText} is safe to send —
* a timeout here only means the daemon's own lead-discovery scan has not caught up yet.
*/
private IdentityResult waitUntilRecognisedAsLead(String newTerminal, int readySeconds) {
long startMillis = nowMillis.getAsLong();
long deadline = startMillis + TimeUnit.SECONDS.toMillis(readySeconds);
while (nowMillis.getAsLong() < deadline) {
if (liveLeadTerminals.get().containsKey(newTerminal)) {
return new IdentityResult(true, nowMillis.getAsLong() - startMillis);
}
pollSleeper.run();
}
return new IdentityResult(false, nowMillis.getAsLong() - startMillis);
}
/** The measured outcome of {@link #waitUntilRecognisedAsLead}. */
private record IdentityResult(boolean ready, long elapsedMillis) {}
/** Drop a pending request without rolling. @return whether a pending request existed for {@code token} */
public boolean cancel(String token) {
return pending.remove(token) != null;
@@ -729,15 +972,12 @@ public final class LeadRollover {
/**
* Poll {@link AgentControl#status} until {@code target} reports a real turn boundary — {@link
* AgentStatus#IDLE} or {@link AgentStatus#DONE} — bounded by {@code settleSeconds}. Used once by
* {@link #runRollover}, to wait for the CALLING turn's own pane to settle before {@code /clear}
* is ever sent at all — the {@code turnSettleSeconds} gate that makes this correction safe. The
* SECOND wait, after {@code /clear}, is {@link #waitForClearPickupAndSettle} instead (fleetd
* #489) — a plain boundary check is not enough there, because {@code /clear} starts no turn of
* its own, so this method would (wrongly) report "settled" on its very first poll whether or not
* {@code /clear} was actually picked up. A failed status read degrades to "not yet settled" and
* is retried on the next poll, the same posture {@code LeadHeartbeatLoop} and {@code
* HerdrPeerLauncher}'s readiness gate already take toward an unreadable status.
* AgentStatus#IDLE} or {@link AgentStatus#DONE} — bounded by {@code settleSeconds}. Used by
* {@link #runRollover} to wait for the CALLING turn's own pane to settle before the old pane is
* touched at all — the {@code turnSettleSeconds} gate that makes tearing it down safe. A failed
* status read degrades to "not yet settled" and is retried on the next poll, the same posture
* {@code LeadHeartbeatLoop} and {@code HerdrPeerLauncher}'s readiness gate already take toward
* an unreadable status.
*
* <p><strong>Deliberately not {@link AgentStatus#injectable()}.</strong> {@code injectable()}
* answers the {@code Injector}'s question — "may I deliver a message without stepping on a live
@@ -745,16 +985,16 @@ public final class LeadRollover {
* an approval prompt is safe to queue a message behind. This class asks a stricter question —
* "has the turn actually ended" — and {@code BLOCKED} answers no: it is a live turn that is
* merely paused, not one that has finished. Reusing {@code injectable()} here would let this
* wait fire {@code /clear} while the lead's own {@code confirm()}-calling turn is still live and
* paused on a prompt — exactly the live-context-destroying failure the {@code turnSettleSeconds}
* gate exists to prevent. Do not "simplify" this back to {@code injectable()}. ({@link
* #waitForClearPickupAndSettle} keeps the same exclusion of {@code BLOCKED}, for the same
* reason, on the second wait.)
* wait tear the old pane down while the lead's own {@code confirm()}-calling turn is still live
* and paused on a prompt — exactly the live-context-destroying failure {@code turnSettleSeconds}
* exists to prevent. Do not "simplify" this back to {@code injectable()}. ({@link
* #waitUntilPaneReady} applies the same exclusion of {@code BLOCKED} to the fresh lead's own
* turn.)
*
* @return a {@link TurnSettleResult} whose {@code settled()} is {@code true} once a real
* boundary was observed, {@code false} if {@code settleSeconds} elapses first.
* {@code elapsedMillis()} is a MEASURED value from the injected {@link #nowMillis}
* clock, never the configured {@code settleSeconds} budget (fleetd #494 follow-up).
* clock, never the configured {@code settleSeconds} budget.
*/
private TurnSettleResult waitUntilAtTurnBoundary(String target, int settleSeconds) {
long startMillis = nowMillis.getAsLong();
@@ -771,142 +1011,11 @@ public final class LeadRollover {
if (status == AgentStatus.IDLE || status == AgentStatus.DONE) {
return new TurnSettleResult(true, nowMillis.getAsLong() - startMillis);
}
settleSleeper.run();
pollSleeper.run();
}
return new TurnSettleResult(false, nowMillis.getAsLong() - startMillis);
}
/**
* The measured outcome of {@link #waitUntilAtTurnBoundary} — fleetd #494 follow-up. The sibling
* of {@link ClearSettleResult} for the FIRST wait, which never nudges, so it carries no nudge
* count.
*/
/** The measured outcome of {@link #waitUntilAtTurnBoundary}. */
private record TurnSettleResult(boolean settled, long elapsedMillis) {}
/**
* The SECOND wait in {@link #runRollover} — after {@code /clear} has been sent, waits for it to
* settle, bounded by {@code settleSeconds}. <strong>fleetd #489 — the paste-race fix.</strong>
* {@code /clear} does not start a real turn of its own, so a pane with no submit race simply
* stays {@link AgentStatus#IDLE} the whole time: {@link #waitUntilAtTurnBoundary} would (wrongly)
* call that "settled" on its very first poll, whether or not the {@code /clear} Enter actually
* landed. That was Fault 1, measured live on 2026-09-12 — the second gate was a no-op, so a
* {@code bootstrapText} send followed immediately, racing Fault 2: {@link AgentControl#submit}'s
* own javadoc already records that the submit accompanying a delivery "can race the paste —
* especially right as the worker's TUI becomes interactive — leaving the text unsubmitted"
* (CB-113). Because {@code runRollover} deliberately bypasses {@code Injector} for {@code
* /clear} (see this class's javadoc), it inherited none of {@code Injector}'s nudging — so the
* lost {@code /clear} Enter sat in the input box and {@code bootstrapText} was typed right after
* it, landing as one concatenated line.
*
* <p>This method copies the pickup-nudge pattern {@link dev.ltms.fleet.inject.Injector} already
* ships for exactly this, on its own post-turn {@code /clear} housekeeping (fleetd #306; see
* {@code Injector.java:288-340} and {@code Injector.java:437-442}):
* <ul>
* <li>an {@link AgentStatus#WORKING} sample means {@code /clear} was picked up as a real
* turn;</li>
* <li>until that happens, each poll that still reports {@link AgentStatus#IDLE} or {@link
* AgentStatus#DONE} re-sends the submit keystroke ({@link AgentControl#submit}) to nudge
* the raced Enter — for the first {@code PICKUP_GRACE_POLLS - 1} of {@link
* #PICKUP_GRACE_POLLS} consecutive such polls (i.e. {@code PICKUP_GRACE_POLLS - 1}
* nudges: 7, not 8, given {@code PICKUP_GRACE_POLLS = 8}). A second Enter on an empty
* Claude Code prompt is a no-op, so repeating it is safe;</li>
* <li>the {@code PICKUP_GRACE_POLLS}th consecutive such poll, with {@code WORKING} still never
* observed, releases rather than wedges the roll instead of nudging again — the same
* choice {@code Injector} makes — and returns {@code settled() == true} anyway, logged at
* {@code warn} with the measured elapsed time (fleetd #494) so an operator can see which
* path ran and how long it actually took;</li>
* <li>once {@code WORKING} has been observed, nudging stops and this instead waits for a real
* {@code working → IDLE/DONE} completion boundary before returning {@code true}.</li>
* </ul>
*
* <p><strong>{@link AgentStatus#BLOCKED} is deliberately excluded from both the nudge and the
* boundary check</strong> — the same reasoning as {@link #waitUntilAtTurnBoundary}'s own
* javadoc: a paused live turn is not a settled one, and re-sending Enter into an open approval
* prompt could wrongly answer it. A {@code BLOCKED} sample (or an unreadable/{@link
* AgentStatus#UNKNOWN} one) simply keeps this polling, with no nudge and no release, until either
* a real boundary is reached or {@code settleSeconds} runs out.
*
* <p>{@link AgentControl#submit} can itself throw; a {@link RuntimeException} from it is
* swallowed and logged at {@code debug}, exactly like {@code Injector.java:437-442} — a failed
* nudge must not abort the roll.
*
* @return a {@link ClearSettleResult} whose {@code settled()} is {@code true} once {@code
* /clear} has settled, or once the nudge budget was exhausted with no pickup ever
* observed (released rather than wedged); {@code false} if {@code settleSeconds} elapses
* first — the caller must NOT send {@code bootstrapText} in that case, exactly as before
* this fix. {@code elapsedMillis()} and {@code nudges()} are MEASURED values (from the
* injected {@link #nowMillis} clock and an actual nudge count), never the configured
* {@code settleSeconds} budget (fleetd #494).
*/
private ClearSettleResult waitForClearPickupAndSettle(String target, int settleSeconds) {
long startMillis = nowMillis.getAsLong();
long deadline = startMillis + TimeUnit.SECONDS.toMillis(settleSeconds);
boolean pickedUp = false; // a WORKING sample has been observed since /clear was sent
int idlePollsAwaitingPickup = 0;
int nudges = 0;
while (nowMillis.getAsLong() < deadline) {
AgentStatus status;
try {
status = agents.status(target);
} catch (RuntimeException e) {
log.debug("lead-rollover: status check failed while waiting for {} to settle after "
+ "/clear: {}", target, e.toString());
status = null;
}
if (status == AgentStatus.WORKING) {
pickedUp = true;
} else if (status == AgentStatus.IDLE || status == AgentStatus.DONE) {
if (pickedUp) {
// a real WORKING -> IDLE/DONE completion boundary
return new ClearSettleResult(true, nowMillis.getAsLong() - startMillis, nudges);
}
if (++idlePollsAwaitingPickup >= PICKUP_GRACE_POLLS) {
long elapsedMillis = nowMillis.getAsLong() - startMillis;
// fleetd #494: this release trades a possibly-unsubmitted /clear for progress
// instead of wedging the roll — that trade is deliberate and stays. But it is
// also exactly the case that reported false success in the real incident (the
// whole roll "succeeded" after 438ms of a 20s budget), so raise it to WARN and
// print the MEASURED elapsed time next to the target pane, not just the count.
//
// fleetd #494 follow-up (2nd pass): BOTH numbers in this line must come from
// the loop's own counters, never from the PICKUP_GRACE_POLLS constant.
// `idlePollsAwaitingPickup` and `nudges` each have exactly one write site in
// this loop, on the same branch, so on this branch they cannot differ from
// PICKUP_GRACE_POLLS / PICKUP_GRACE_POLLS - 1 today — no test can prove the
// difference on this line, and printing the counters does not change that.
// What it does buy: one source of truth instead of two, so a later change to
// the loop (an early return, a second increment site, a different exit
// condition) cannot leave this message reporting a number the loop no longer
// produces. The place where `nudges` genuinely varies with the run — and is
// covered by a test that can tell it apart from a constant — is the
// /clear-timeout warn in runRollover, which prints clearResult.nudges().
log.warn("lead-rollover: /clear on {} was never observed as WORKING after {} "
+ "consecutive IDLE/DONE polls ({} of those were nudged) — "
+ "releasing rather than wedging the roll (elapsed={}ms)",
target, idlePollsAwaitingPickup, nudges, elapsedMillis);
return new ClearSettleResult(true, elapsedMillis, nudges);
}
try {
agents.submit(target); // nudge a raced Enter (CB-113) so /clear actually submits
} catch (RuntimeException e) {
log.debug("lead-rollover: resubmit to {} failed (will retry next poll): {}",
target, e.getMessage());
} finally {
nudges++; // an attempted nudge, whether or not the submit call itself threw
}
}
// AgentStatus.BLOCKED or UNKNOWN (or an unreadable status, above): neither a pickup
// signal nor a boundary — keep polling without nudging or releasing.
settleSleeper.run();
}
return new ClearSettleResult(false, nowMillis.getAsLong() - startMillis, nudges);
}
/**
* The measured outcome of {@link #waitForClearPickupAndSettle} — fleetd #494. Carries the
* MEASURED elapsed time (from the injected {@link #nowMillis} clock) and nudge count alongside
* the settle/timeout decision, so callers can log them instead of the configured budget, which
* is not how long the wait actually ran.
*/
private record ClearSettleResult(boolean settled, long elapsedMillis, int nudges) {}
}
@@ -891,8 +891,8 @@ public final class FleetMcp {
* fleetd #612 B3 — as {@link #quarantineSource()}, {@code public} for the same cross-package
* reason, for the real {@link LeadRollover} (or {@code null}) this daemon was assembled with.
* {@code FleetdLeadRolloverAssemblyTest} drives {@code open}/{@code confirm} on this exact
* instance and waits for the real continuation to send {@code /clear} and {@code bootstrapText}
* through the real {@code router.leadAgents()}.
* instance and waits for the real continuation to end the old pane, relaunch a fresh one, and
* send {@code bootstrapText} through the real {@code router.leadAgents()}.
*/
public LeadRollover leadRollover() {
return leadRollover;
@@ -1492,7 +1492,7 @@ public final class FleetMcp {
}
if (isBlank(callerTerminal)) {
// An unnamed primary (token/loopback path, no resolved pane) has nowhere for the
// eventual /clear + bootstrap to land — LeadRollover#open would throw
// eventual relaunch + bootstrap to land — LeadRollover#open would throw
// IllegalArgumentException for the same reason; refuse cleanly here instead.
return error("fleet_handover requires a named lead pane (a resolved connection terminal) "
+ "to open a rollover request against — an unnamed primary has none");
@@ -2630,17 +2630,20 @@ public final class FleetMcp {
private static McpSchema.Tool handoverTool() {
return tool(FleetTool.HANDOVER.wireName(),
"Replace your OWN lead session once its context is full: write a handover file, "
+ "then use this to have fleetd clear your pane and bootstrap a fresh lead "
+ "session against it. Four actions: 'open' (requests a token and the "
+ "handoverPath you must write the handover file to before confirming), "
+ "'confirm' (validates every gate and — only if every one passes — schedules "
+ "the roll; it does NOT itself clear the pane, the roll runs once this call's "
+ "own turn ends), 'cancel' (drops a pending request without rolling), and "
+ "'status' (read-only: what happened to a token after 'confirm' — still "
+ "running (approved but not finished yet), the roll completed, the calling "
+ "turn never settled within turnSettleSeconds so no /clear was ever sent, or "
+ "/clear itself never settled so bootstrapText was never sent; never "
+ "schedules, cancels or retries anything). Primary-only. "
+ "then use this to have fleetd end your pane's process and relaunch a fresh "
+ "lead session bootstrapped against it. Four actions: 'open' (requests a "
+ "token and the handoverPath you must write the handover file to before "
+ "confirming), 'confirm' (validates every gate and — only if every one "
+ "passes — schedules the roll; it does NOT itself end your pane, the roll "
+ "runs once this call's own turn ends), 'cancel' (drops a pending request "
+ "without rolling), and 'status' (read-only: what happened to a token after "
+ "'confirm' — still running (approved but not finished yet), the roll "
+ "completed, the calling turn never settled within turnSettleSeconds so "
+ "nothing was touched, the old pane never confirmed dead so no relaunch was "
+ "attempted, the relaunch itself failed, the fresh pane never became ready "
+ "so bootstrapText was never sent, or the fresh pane became ready and was "
+ "bootstrapped but was never recognised as a live lead; never schedules, "
+ "cancels or retries anything). Primary-only. "
+ "There is deliberately no terminal/session/leadTerminal parameter: the pane "
+ "to roll is always resolved from YOUR OWN connection, never a value you "
+ "pass, so you can only ever roll yourself — never another lead. Requires "
@@ -6,6 +6,8 @@ import dev.ltms.fleet.guard.SubscriptionGuard;
import dev.ltms.fleet.herdr.AgentControl;
import dev.ltms.fleet.herdr.FakeHerdr;
import dev.ltms.fleet.herdr.HerdrClient;
import dev.ltms.fleet.herdr.WorkspaceControl;
import dev.ltms.fleet.lead.LeadLauncher;
import dev.ltms.fleet.lead.LeadRollover;
import dev.ltms.fleet.msg.ReplyInbox;
import io.javalin.Javalin;
@@ -34,7 +36,7 @@ import static org.junit.jupiter.api.Assertions.fail;
* fleetd #612 Unit A) with three methods: {@code unrelatedAnchorStillPresent} (a scaffold anchor,
* not an independent claim — needs no replacement of its own), {@code
* mainStillCallsTheLeadRolloverFactory} (the call-site pin replaced by {@link
* #assembledLeadRolloverRunsTheRealClearAndBootstrapSequence}), and {@code
* #assembledLeadRolloverEndsTheOldPaneThroughTheRealHerdrRouter}), and {@code
* factoryGatesOnConfigPresence} (the absent-config claim replaced by {@link
* #absentLeadRolloverConfigMeansNoRolloverIsBuilt} — a claim this ticket found was NOT actually
* covered behaviourally anywhere else: {@code LeadRolloverTest}'s only related assertion is
@@ -50,8 +52,8 @@ import static org.junit.jupiter.api.Assertions.fail;
* invisible to this test, even though the two are genuinely different daemons in production. This
* version configures two distinct sockets and two distinct {@link FakeHerdr} instances (the same
* pattern {@code FleetdAssemblyConnectionIdentityTest}, fleetd #612 B2, already uses to separate
* lead from member) and asserts the roll's {@code /clear}/bootstrap sends land on the LEAD fake
* and never on the MEMBER one.
* lead from member) and asserts the roll's {@code pane.close} call lands on the LEAD fake and
* never on the MEMBER one.
*/
class FleetdLeadRolloverAssemblyTest {
@@ -177,11 +179,11 @@ class FleetdLeadRolloverAssemblyTest {
return FleetConfig.load(f);
}
@SuppressWarnings("unchecked")
@Test
@DisplayName("[BEHAVIOURAL] the real assembled LeadRollover runs the full open/confirm/continuation "
+ "sequence — /clear, then bootstrapText — through the real herdr router")
void assembledLeadRolloverRunsTheRealClearAndBootstrapSequence(@TempDir Path dir) throws Exception {
@DisplayName("[BEHAVIOURAL] the real assembled LeadRollover runs the open/confirm/continuation "
+ "sequence through the real herdr router — ending the old pane, then giving up once it "
+ "never reports gone")
void assembledLeadRolloverEndsTheOldPaneThroughTheRealHerdrRouter(@TempDir Path dir) throws Exception {
Path leadCwd = dir.resolve("lead-workspace");
Files.createDirectories(leadCwd);
FleetConfig cfg = writeConfig(dir, leadCwd);
@@ -223,40 +225,42 @@ class FleetdLeadRolloverAssemblyTest {
// constructor), so this polls the real FleetMcp.leadRollover() instance's status(token)
// until the real continuation finishes.
LeadRollover.RollStatus status = pollUntilTerminal(rollover, pending.token());
assertEquals(LeadRollover.RollState.ROLLED, status.state(),
"the full happy path must complete: FakeHerdr's default agent status is 'idle', so "
+ "the turn-boundary wait settles immediately and the post-/clear wait "
+ "releases via its pickup-grace path — detail: " + status.detail());
// Prove the real herdr router actually sent BOTH messages, in order, to the real LEAD
// pane — this is the one thing a source-text pin on the call site could never show.
// FakeHerdr's pane.get is a fixed canned response that never reports a pane as gone, so the
// real router's death poll runs out its whole budget and the roll stops here — proving the
// real teardown call landed on the real LEAD pane without ever reaching a relaunch or a send.
assertEquals(LeadRollover.RollState.OLD_PANE_NEVER_DIED, status.state(),
"the old pane never reports gone against this fake, so the roll must stop with "
+ "OLD_PANE_NEVER_DIED rather than ever relaunching or sending anything — "
+ "detail: " + status.detail());
// Prove the real herdr router actually closed the real LEAD pane — this is the one thing a
// source-text pin on the call site could never show.
boolean closedOldPane = lead.calls.stream()
.anyMatch(c -> c.method().equals("pane.close")
&& c.params() instanceof Map<?, ?> m && "w2:p7".equals(m.get("pane_id")));
assertTrue(closedOldPane, "endOldSession must close the real old pane (w2:p7) through the "
+ "real LEAD herdr client, got calls: " + lead.calls);
// No agent.prompt is ever sent on this path: the roll stops at the pane-death wait, strictly
// before the relaunch and the final send step.
List<FakeHerdr.Call> prompts = lead.calls.stream()
.filter(c -> c.method().equals("agent.prompt"))
.toList();
assertTrue(prompts.size() >= 2, "expected at least a /clear send and a bootstrapText send "
+ "on the LEAD daemon, got " + prompts.size() + " agent.prompt calls: " + prompts);
assertEquals("/clear", ((Map<String, Object>) prompts.get(0).params()).get("text"),
"the first send must be the literal /clear housekeeping command");
Object secondText = ((Map<String, Object>) prompts.get(1).params()).get("text");
assertTrue(secondText instanceof String && ((String) secondText).contains(expectedHandoverPath),
"the second send must be the default bootstrapText naming the resolved handover "
+ "path, got: " + secondText);
assertTrue(prompts.isEmpty(), "a roll that stops at OLD_PANE_NEVER_DIED must never reach the "
+ "send step, got agent.prompt call(s) on the LEAD daemon: " + prompts);
// fleetd #612 B3 correction: prove the roll never touches the MEMBER daemon. A mutation
// swapping router.leadAgents() for router.memberAgents() at the real call site would move
// both sends above onto `member` instead, which this assertion catches — the thing the
// single-fake version of this test could never see, because both wrapped the same client.
List<FakeHerdr.Call> memberPrompts = member.calls.stream()
.filter(c -> c.method().equals("agent.prompt"))
.toList();
assertTrue(memberPrompts.isEmpty(), "the roll must be wired to the LEAD daemon only — got "
+ memberPrompts.size() + " agent.prompt call(s) on the MEMBER daemon instead: "
+ memberPrompts);
// the pane.close call above onto `member` instead, which this assertion catches — the thing
// the single-fake version of this test could never see, because both wrapped the same client.
assertTrue(member.calls.isEmpty(), "the roll must be wired to the LEAD daemon only — got "
+ member.calls.size() + " call(s) on the MEMBER daemon instead: " + member.calls);
}
private static LeadRollover.RollStatus pollUntilTerminal(LeadRollover rollover, String token)
throws InterruptedException {
long deadline = System.nanoTime() + java.util.concurrent.TimeUnit.SECONDS.toNanos(10);
long deadline = System.nanoTime() + java.util.concurrent.TimeUnit.SECONDS.toNanos(15);
while (System.nanoTime() < deadline) {
LeadRollover.RollStatus status = rollover.status(token);
if (status.state() != LeadRollover.RollState.PENDING
@@ -283,8 +287,10 @@ class FleetdLeadRolloverAssemblyTest {
""");
ConfigRef config = new ConfigRef(yaml, FleetConfig.load(yaml));
AgentControl agents = new AgentControl(new FakeHerdr());
WorkspaceControl spaces = new WorkspaceControl(new FakeHerdr());
LeadLauncher launcher = new LeadLauncher(agents, spaces, config.get());
LeadRollover rollover = Fleetd.leadRollover(config.get(), agents, config, Map::of);
LeadRollover rollover = Fleetd.leadRollover(config.get(), agents, spaces, launcher, config, Map::of);
assertNull(rollover, "leadRollover: is absent from this config, so the factory's opt-in "
+ "gate (`if (cfg.leadRollover() == null) return null;`) must fire and no "
@@ -4,6 +4,8 @@ import dev.ltms.fleet.config.ConfigRef;
import dev.ltms.fleet.config.FleetConfig;
import dev.ltms.fleet.herdr.AgentControl;
import dev.ltms.fleet.herdr.FakeHerdr;
import dev.ltms.fleet.herdr.WorkspaceControl;
import dev.ltms.fleet.lead.LeadLauncher;
import dev.ltms.fleet.lead.LeadRollover;
import org.junit.jupiter.api.DisplayName;
import org.junit.jupiter.api.Test;
@@ -54,6 +56,11 @@ class FleetdLeadRolloverWorkspaceLookupTest {
return new AgentControl(new FakeHerdr());
}
/** None of this class's tests reach the deferred continuation, so a plain fake is enough. */
private static LeadLauncher fakeLauncher(FleetConfig cfg) {
return new LeadLauncher(fakeAgents(), new WorkspaceControl(new FakeHerdr()), cfg);
}
@Test
@DisplayName("[BEHAVIOURAL] Fleetd.leadRollover(...) resolves a relative handoverPath against "
+ "the CALLING lead's configured cwd, not the daemon's own working directory")
@@ -74,7 +81,8 @@ class FleetdLeadRolloverWorkspaceLookupTest {
""".formatted(leadCwd.toString()));
ConfigRef config = new ConfigRef(yaml, FleetConfig.load(yaml));
LeadRollover rollover = Fleetd.leadRollover(config.get(), fakeAgents(), config,
LeadRollover rollover = Fleetd.leadRollover(config.get(), fakeAgents(),
new WorkspaceControl(new FakeHerdr()), fakeLauncher(config.get()), config,
() -> Map.of("term_opus", "opus"));
assertNotNull(rollover, "leadRollover: is present in the loaded config, so the factory "
+ "must construct an object");
@@ -105,7 +113,8 @@ class FleetdLeadRolloverWorkspaceLookupTest {
// No lead has been discovered yet — exactly the real shape of a lead the live tab scan
// has not yet scanned, or one with no fleet.leaders entry at all.
LeadRollover rollover = Fleetd.leadRollover(config.get(), fakeAgents(), config, Map::of);
LeadRollover rollover = Fleetd.leadRollover(config.get(), fakeAgents(),
new WorkspaceControl(new FakeHerdr()), fakeLauncher(config.get()), config, Map::of);
assertNotNull(rollover);
LeadRollover.PendingRollover pending = rollover.open("term_unknown", "test");
@@ -144,7 +153,8 @@ class FleetdLeadRolloverWorkspaceLookupTest {
// below — exactly the natural mistake to make, since leads are discovered by a live tab
// scan that runs AFTER this factory is constructed at startup.
Map<String, String> liveLeadTerminals = new HashMap<>();
LeadRollover rollover = Fleetd.leadRollover(config.get(), fakeAgents(), config,
LeadRollover rollover = Fleetd.leadRollover(config.get(), fakeAgents(),
new WorkspaceControl(new FakeHerdr()), fakeLauncher(config.get()), config,
() -> liveLeadTerminals);
assertNotNull(rollover);
@@ -103,7 +103,7 @@ class FleetConfigWithDefaultsPreservesEveryComponentTest {
// comment there), same as broker/primary/leadHeartbeat/... above — a real, non-null value
// here proves it, rather than leaving it null and proving nothing.
v.put("leadRollover", new FleetConfig.LeadRollover(
"/handover/guard.md", true, 3600, 20, 20, "read the handover file"));
"/handover/guard.md", true, 3600, 20, 45, "read the handover file"));
assertNamesMatchComponents(v);
return v;
}
@@ -7,6 +7,7 @@ import java.util.ArrayList;
import java.util.LinkedHashMap;
import java.util.List;
import java.util.Map;
import java.util.Set;
import java.util.concurrent.ConcurrentHashMap;
import java.util.concurrent.CopyOnWriteArrayList;
@@ -59,6 +60,10 @@ public final class FakeHerdr implements HerdrClient {
private Runnable onAgentStart; // fires the instant agent.start is called — see onAgentStart(Runnable)
private volatile int agentGetOkCalls = Integer.MAX_VALUE; // how many agent.get calls succeed first
private volatile String agentGetFailCode = null; // error code every agent.get call after that reports
/** pane ids that {@link #paneGoneAfterClose} has opted into reporting gone — see that method. */
private final Set<String> paneGoneAfterCloseIds = ConcurrentHashMap.newKeySet();
/** pane ids a {@code pane.close} call has actually reached, for {@link #paneGoneAfterCloseIds}. */
private final Set<String> closedPaneIds = ConcurrentHashMap.newKeySet();
public FakeHerdr healthy(boolean h) {
this.healthy = h;
@@ -220,6 +225,19 @@ public final class FakeHerdr implements HerdrClient {
return this;
}
/**
* Make {@code pane.get(paneId)} report the pane gone (a {@code pane_not_found} {@link
* HerdrException}, exactly as {@link WorkspaceControl#locatePane} expects to see once a pane
* has really disappeared) once a {@code pane.close} call for that same {@code paneId} has
* actually reached this fake. Every other pane, and this pane before its own close, keeps
* reporting the default canned {@code pane.get} response — opt-in, by pane id, so no existing
* test's {@code pane.get} behaviour changes.
*/
public FakeHerdr paneGoneAfterClose(String paneId) {
paneGoneAfterCloseIds.add(paneId);
return this;
}
/**
* Run {@code hook} synchronously the instant an {@code agent.start} call reaches this fake —
* i.e. the instant the peer PROCESS would start against a real herdr daemon. A test uses this
@@ -422,9 +440,18 @@ public final class FakeHerdr implements HerdrClient {
}
yield mapper.readTree("{\"type\":\"ok\"}");
}
case "pane.get" -> mapper.readTree("""
case "pane.get" -> {
Object paneIdParam = params instanceof Map<?, ?> m ? m.get("pane_id") : null;
String paneIdKey = paneIdParam == null ? null : String.valueOf(paneIdParam);
if (paneIdKey != null && paneGoneAfterCloseIds.contains(paneIdKey)
&& closedPaneIds.contains(paneIdKey)) {
throw new HerdrException("herdr error [pane_not_found]: pane.get failed",
"pane_not_found", null);
}
yield mapper.readTree("""
{"type":"pane_info","pane":{"pane_id":"w9:pW","workspace_id":"w9",
"tab_id":"w9:t2","agent_status":"idle"}}""");
}
case "pane.list" -> noPanes
? mapper.readTree("{\"type\":\"pane_list\",\"panes\":[]}")
: mapper.readTree("""
@@ -458,6 +485,9 @@ public final class FakeHerdr implements HerdrClient {
throw new HerdrException("herdr error [" + code + "]: pane.close failed",
code, null);
}
if (paneIdParam != null) {
closedPaneIds.add(String.valueOf(paneIdParam));
}
yield mapper.readTree("{\"type\":\"ok\"}");
}
default -> throw new HerdrException("fake has no canned response for " + method);
File diff suppressed because it is too large Load Diff
@@ -11,6 +11,7 @@ import dev.ltms.fleet.herdr.FakeHerdr;
import dev.ltms.fleet.herdr.PaneLocator;
import dev.ltms.fleet.herdr.WorkspaceControl;
import dev.ltms.fleet.inject.Injector;
import dev.ltms.fleet.lead.LeadLauncher;
import dev.ltms.fleet.lead.LeadRollover;
import dev.ltms.fleet.member.ClaudeCodeLauncher;
import dev.ltms.fleet.msg.InMemoryReplyInbox;
@@ -24,6 +25,8 @@ import org.junit.jupiter.api.DisplayName;
import org.junit.jupiter.api.Test;
import org.junit.jupiter.api.io.TempDir;
import java.io.IOException;
import java.io.UncheckedIOException;
import java.nio.file.Files;
import java.nio.file.Path;
import java.util.List;
@@ -36,14 +39,14 @@ import static org.junit.jupiter.api.Assertions.*;
* fleetd #480 Unit C — the {@code fleet_handover} MCP tool, the surface that finally calls
* {@link LeadRollover#open}/{@link LeadRollover#confirm}/{@link LeadRollover#cancel}.
*
* <p>Uses {@link LeadRollover}'s PUBLIC constructor (real wall clock, real 250ms settle poll, a
* real virtual-thread continuation runner) rather than its package-private test constructor —
* this test lives in {@code dev.ltms.fleet.mcp}, not {@code dev.ltms.fleet.lead}, and does not
* need to control the post-{@code confirm()} continuation's timing: it only asserts the
* SYNCHRONOUS return value of {@code open}/{@code confirm}/{@code cancel}, which is exactly what
* {@code FleetMcp.handover} forwards to the client. {@code turnSettleSeconds}/{@code
* clearSettleSeconds} are kept at 1s so a confirmed request's background continuation (which this
* class does not wait on or assert against) gives up quickly rather than polling for 20s on a
* <p>Uses {@link LeadRollover}'s PUBLIC constructor (real wall clock, real 250ms poll, a real
* virtual-thread continuation runner) rather than its package-private test constructor — this
* test lives in {@code dev.ltms.fleet.mcp}, not {@code dev.ltms.fleet.lead}, and does not need to
* control the post-{@code confirm()} continuation's timing: it only asserts the SYNCHRONOUS return
* value of {@code open}/{@code confirm}/{@code cancel}, which is exactly what {@code
* FleetMcp.handover} forwards to the client. {@code turnSettleSeconds}/{@code
* relaunchReadySeconds} are kept at 1s so a confirmed request's background continuation (which
* this class does not wait on or assert against) gives up quickly rather than polling for 20s on a
* daemon virtual thread.
*/
class FleetMcpHandoverTest {
@@ -67,10 +70,26 @@ class FleetMcpHandoverTest {
return new FleetConfig.LeadRollover(handoverPath, false, 3600, 1, 1, "read the handover file");
}
private static FleetConfig minimalFleetConfig() {
try {
Path yaml = Files.createTempFile("fleet-mcp-handover-test", ".yaml");
Files.writeString(yaml, "bind:\n port: 8080\n");
return FleetConfig.load(yaml);
} catch (IOException e) {
throw new UncheckedIOException(e);
}
}
private LeadRollover newRollover(String handoverPath) {
// Every handoverPath this test class uses comes from tmp.resolve(...), which is already
// absolute, so the workspace lookup is never actually consulted — a no-op lookup is enough.
return new LeadRollover(agents, () -> cfg(handoverPath), _ -> null);
// None of this class's tests reach the recognition-wait or the relaunch call, so the
// launcher's own functional correctness is irrelevant here — any constructed instance,
// backed by the same fake herdr, is enough.
WorkspaceControl spaces = new WorkspaceControl(herdr);
LeadLauncher launcher = new LeadLauncher(agents, spaces, minimalFleetConfig());
return new LeadRollover(agents, spaces, launcher, () -> cfg(handoverPath),
_ -> null, _ -> null, Map::of);
}
/** A fully wired FleetMcp on fakes (mirrors FleetMcpAuthzTest's helper), plus a leadRollover. */