fleetd #726 unit 2: replace /clear-based lead rollover with a real process restart
LeadRollover's deferred continuation now ends the old lead's pane, relaunches a fresh one, and bootstraps it, instead of sending /clear into the same process. The relaunch step runs two separate bounded waits instead of one combined check: a readiness wait (the fresh pane reaches a real turn boundary) is the safety gate and withholds bootstrapText on timeout (RELAUNCH_NEVER_READY); a recognition wait (the fresh terminal shows up in the live-lead map) is bookkeeping only, so a timeout there still lets bootstrapText go out (RELAUNCH_NOT_RECOGNISED). clearSettleSeconds is retired in favor of relaunchReadySeconds (default 45), which bounds both waits. Updates FleetConfig/FleetMcp operator-facing text to match.
This commit is contained in:
+22
-17
@@ -99,17 +99,17 @@ bind:
|
||||
# contextHighNudge: false
|
||||
|
||||
# Lead rollover (fleetd #480): replace a lead session that has decided it is ready to be replaced,
|
||||
# without an operator doing it by hand. A lead writes a handover file, then asks fleetd to clear its
|
||||
# own pane and bootstrap a fresh session against that file.
|
||||
# without an operator doing it by hand. A lead writes a handover file, then asks fleetd to end its
|
||||
# own pane, launch a fresh one, and bootstrap that fresh session against the handover file.
|
||||
#
|
||||
# Opt-in on purpose — it clears the lead's own pane on request, so upgrading the daemon must never
|
||||
# acquire that ability for you. Absent block = feature off, and nothing is constructed at all. Even
|
||||
# once present, nothing but an explicit confirm() call — one that passes every check — can ever
|
||||
# cause a /clear: there is no recurring timer, heartbeat or scheduler anywhere in this feature that
|
||||
# fires one on its own initiative. confirm() itself is called FROM the calling lead's own turn, so
|
||||
# it cannot clear the pane inline (that pane is still WORKING); instead it schedules a one-shot
|
||||
# Opt-in on purpose — it tears down the lead's own pane on request, so upgrading the daemon must
|
||||
# never acquire that ability for you. Absent block = feature off, and nothing is constructed at all.
|
||||
# Even once present, nothing but an explicit confirm() call — one that passes every check — can ever
|
||||
# tear a pane down: there is no recurring timer, heartbeat or scheduler anywhere in this feature that
|
||||
# fires one on its own initiative. confirm() itself is called FROM the calling lead's own turn, so it
|
||||
# cannot act on the pane inline (that pane is still WORKING); instead it schedules a one-shot
|
||||
# continuation that waits for the SAME confirm() call's turn to end, then does the actual work. See
|
||||
# dev.ltms.fleet.lead.LeadRollover's class javadoc for the exact order (fleetd #480 correction).
|
||||
# dev.ltms.fleet.lead.LeadRollover's class javadoc for the exact order.
|
||||
#
|
||||
# handoverPath: REQUIRED when this block is present — where the handover file a fresh lead session
|
||||
# reads must live. No default (an operator-specific path); a present block with no
|
||||
@@ -118,25 +118,30 @@ bind:
|
||||
# working directory when that lead has none configured) — never against whatever
|
||||
# directory the daemon process happens to have been started in. An absolute path is
|
||||
# used unchanged. Prefer an absolute path if the daemon and the lead's pane might not
|
||||
# share a working directory (fleetd #480 follow-up).
|
||||
# share a working directory.
|
||||
# requireOperatorConfirm: true # default true — confirm() refuses unless the caller also passes
|
||||
# # operatorConfirmed: true
|
||||
# maxDocAgeSeconds: 3600 # default 3600 — refuse a handover file older than this
|
||||
# turnSettleSeconds: 20 # default 20 — how long the deferred roll waits for the CALLING
|
||||
# # lead's own turn to end (its pane to report injectable again)
|
||||
# # before sending /clear at all. If this elapses, /clear is NEVER
|
||||
# # sent — a lead that never goes idle is still doing real work.
|
||||
# clearSettleSeconds: 20 # default 20 — how long to wait for the pane to become injectable
|
||||
# # again AFTER /clear before giving up (never sends bootstrapText
|
||||
# # if this elapses). A separate, second wait from turnSettleSeconds.
|
||||
# # before tearing the old pane down at all. If this elapses, nothing
|
||||
# # is torn down — a lead that never goes idle is still doing real
|
||||
# # work.
|
||||
# relaunchReadySeconds: 45 # default 45 — bounds two later waits, after the old pane is gone
|
||||
# # and a fresh one has been launched: first, for the fresh pane to
|
||||
# # reach a real turn boundary (never sends bootstrapText if THIS one
|
||||
# # elapses); second, for the new terminal to be recognised as this
|
||||
# # lead (bootstrapText is sent either way once the first wait
|
||||
# # passes). A separate, later pair of waits from turnSettleSeconds.
|
||||
# bootstrapText: "..." # default names the RESOLVED (absolute) handoverPath — sent to
|
||||
# # the lead once its pane settles after /clear
|
||||
# # the freshly relaunched lead's pane once it reaches a real turn
|
||||
# # boundary
|
||||
# leadRollover:
|
||||
# handoverPath: /path/to/handover.md
|
||||
# requireOperatorConfirm: true
|
||||
# maxDocAgeSeconds: 3600
|
||||
# turnSettleSeconds: 20
|
||||
# clearSettleSeconds: 20
|
||||
# relaunchReadySeconds: 45
|
||||
# bootstrapText: "Fresh lead session: read the handover file and carry on."
|
||||
|
||||
# Fleet health detection is dormant unless enabled (CB-573). It reads one whole-fleet agent list
|
||||
|
||||
@@ -942,21 +942,27 @@ public final class Fleetd {
|
||||
* cfg.leadHeartbeat()}
|
||||
* @param leadAgents the {@link AgentControl} instance that reaches the LEAD's pane (not
|
||||
* {@code memberAgents}), normally {@code router.leadAgents()}
|
||||
* @param leadSpaces the {@link WorkspaceControl} instance that reaches the LEAD's
|
||||
* workspace, normally {@code router.leadSpaces()} — used to tear down
|
||||
* a rolled lead's old pane and confirm it is gone
|
||||
* @param launcher starts the fresh lead a roll relaunches once the old one is gone
|
||||
* @param config the live {@link ConfigRef}, captured only inside the returned
|
||||
* supplier and the workspace lookup — never dereferenced here
|
||||
* supplier and the two lookups below — never dereferenced here
|
||||
* @param liveLeadTerminals terminal id → lead NAME for every CURRENTLY recognised lead, normally
|
||||
* the same {@code leads} supplier {@code main} already builds for
|
||||
* {@code HerdrRouter}/{@link #leadSeatLookup} — never a value snapshot
|
||||
* @return a constructed {@link LeadRollover}, or {@code null} when {@code leadRollover:} is
|
||||
* absent from the startup config
|
||||
*/
|
||||
static LeadRollover leadRollover(FleetConfig cfg, AgentControl leadAgents, ConfigRef config,
|
||||
static LeadRollover leadRollover(FleetConfig cfg, AgentControl leadAgents,
|
||||
WorkspaceControl leadSpaces, LeadLauncher launcher, ConfigRef config,
|
||||
Supplier<Map<String, String>> liveLeadTerminals) {
|
||||
if (cfg.leadRollover() == null) {
|
||||
return null;
|
||||
}
|
||||
Function<String, String> leadNameForTerminal = terminal -> liveLeadTerminals.get().get(terminal);
|
||||
Function<String, String> leadWorkspace = terminal -> {
|
||||
String leadName = liveLeadTerminals.get().get(terminal);
|
||||
String leadName = leadNameForTerminal.apply(terminal);
|
||||
if (leadName == null) {
|
||||
return null;
|
||||
}
|
||||
@@ -964,7 +970,8 @@ public final class Fleetd {
|
||||
FleetConfig.Leader leader = fleet == null ? null : fleet.leaders().get(leadName);
|
||||
return leader == null ? null : leader.cwd();
|
||||
};
|
||||
return new LeadRollover(leadAgents, () -> config.get().leadRollover(), leadWorkspace);
|
||||
return new LeadRollover(leadAgents, leadSpaces, launcher, () -> config.get().leadRollover(),
|
||||
leadWorkspace, leadNameForTerminal, liveLeadTerminals);
|
||||
}
|
||||
|
||||
/**
|
||||
|
||||
@@ -297,11 +297,15 @@ final class FleetdAssembly {
|
||||
leadsRef.set(leads);
|
||||
collaboratorTerminalsRef.set(collaboratorTerminals);
|
||||
|
||||
// Constructed unconditionally — it is cheap and side-effect free — so a LeadRollover built
|
||||
// below can relaunch a lead even on a boot where herdr was down for the ensureLeads() call.
|
||||
LeadLauncher leadLauncher = new LeadLauncher(router.leadAgents(), router.leadSpaces(), cfg);
|
||||
|
||||
// CB-558: start any declared lead that is not already running. After the scanner is built,
|
||||
// and only when herdr answered — the launcher's whole safety property is that it can count
|
||||
// live leads first, and must never guess and risk a second orchestrator.
|
||||
if (herdrUp && !leaders.isEmpty()) {
|
||||
int launched = new LeadLauncher(router.leadAgents(), router.leadSpaces(), cfg).ensureLeads();
|
||||
int launched = leadLauncher.ensureLeads();
|
||||
if (launched > 0) {
|
||||
log.info("lead auto-launch: {} lead(s) started", launched);
|
||||
}
|
||||
@@ -440,7 +444,8 @@ final class FleetdAssembly {
|
||||
heartbeatScheduler.shutdownNow();
|
||||
}
|
||||
// fleetd #480: lead rollover. Opt-in; absent `leadRollover:` this is never constructed.
|
||||
LeadRollover leadRollover = Fleetd.leadRollover(cfg, router.leadAgents(), config, leads);
|
||||
LeadRollover leadRollover = Fleetd.leadRollover(cfg, router.leadAgents(), router.leadSpaces(),
|
||||
leadLauncher, config, leads);
|
||||
MessageService messages = new MessageService(router, injector, rendezvous, replyInbox,
|
||||
pushLoop, metrics);
|
||||
|
||||
|
||||
@@ -67,7 +67,7 @@ import java.util.function.Supplier;
|
||||
* {@code models:} above: {@code dev.ltms.fleet.lead.LeadRollover} holds a
|
||||
* {@code Supplier<FleetConfig.LeadRollover>} (the same {@code () -> config.get().x()} shape)
|
||||
* and reads {@code handoverPath}/{@code requireOperatorConfirm}/{@code maxDocAgeSeconds}/
|
||||
* {@code turnSettleSeconds}/{@code clearSettleSeconds}/{@code bootstrapText} fresh on every
|
||||
* {@code turnSettleSeconds}/{@code relaunchReadySeconds}/{@code bootstrapText} fresh on every
|
||||
* {@code open()}/{@code confirm()} call (and on the deferred post-{@code confirm()}
|
||||
* continuation fleetd #480's correction added — see {@code LeadRollover}'s class doc) rather
|
||||
* than capturing them into fields at construction — unlike its closest
|
||||
|
||||
@@ -1416,17 +1416,16 @@ public record FleetConfig(
|
||||
* before anything exists to call — the same fact already true of adding a brand-new
|
||||
* {@code profiles:} entry.
|
||||
*
|
||||
* <p><strong>{@code turnSettleSeconds} (fleetd #480 correction):</strong> {@code confirm()} is
|
||||
* called FROM the calling lead's own turn, so its pane is still {@code WORKING} the instant
|
||||
* {@code confirm()} validates every gate and schedules the roll. {@code
|
||||
* dev.ltms.fleet.lead.LeadRollover}'s deferred continuation waits up to this many seconds for
|
||||
* that SAME pane to report {@code IDLE} or {@code DONE} — i.e. for the calling turn to actually
|
||||
* end — before it sends {@code /clear} at all. {@code BLOCKED} does not count: that is a live
|
||||
* turn merely paused, not one that has finished. If that wait times out, no {@code /clear} is
|
||||
* ever sent: a lead that never goes idle is still doing real work, and clearing it would
|
||||
* destroy live context. This is a separate wait from {@code clearSettleSeconds} below, which
|
||||
* bounds the SECOND wait, for the pane to reach {@code IDLE} or {@code DONE} again AFTER
|
||||
* {@code /clear} has already gone out.
|
||||
* <p><strong>{@code turnSettleSeconds}:</strong> {@code confirm()} is called FROM the calling
|
||||
* lead's own turn, so its pane is still {@code WORKING} the instant {@code confirm()} validates
|
||||
* every gate and schedules the roll. {@code dev.ltms.fleet.lead.LeadRollover}'s deferred
|
||||
* continuation waits up to this many seconds for that SAME pane to report {@code IDLE} or
|
||||
* {@code DONE} — i.e. for the calling turn to actually end — before it ends the old pane's
|
||||
* process at all. {@code BLOCKED} does not count: that is a live turn merely paused, not one
|
||||
* that has finished. If that wait times out, the old pane is never touched: a lead that never
|
||||
* goes idle is still doing real work, and the roll ends that pane's whole process — there is no
|
||||
* way back from this once it runs, so this wait is the only thing standing between "still
|
||||
* working" and "gone".
|
||||
*
|
||||
* @param handoverPath required when this block is present — where the handover file a fresh
|
||||
* lead session reads must live. There is no sane non-null default for an
|
||||
@@ -1447,36 +1446,49 @@ public record FleetConfig(
|
||||
* attempt can never be mistaken for a fresh one.
|
||||
* @param turnSettleSeconds default 300 — bound on how long the deferred roll waits for the
|
||||
* CALLING lead's own turn to end (its pane to report {@code IDLE} or
|
||||
* {@code DONE}) before sending {@code /clear} at all. See the paragraph
|
||||
* above.
|
||||
* @param clearSettleSeconds default 20 — bound on how long to wait for the lead's pane to
|
||||
* report {@code IDLE} or {@code DONE} again after {@code /clear} before
|
||||
* giving up. A roll that times out here never sends {@code bootstrapText}.
|
||||
* {@code DONE}) before ending that pane's process at all. See the
|
||||
* paragraph above.
|
||||
* @param relaunchReadySeconds default 45 — bound on EACH of two separate waits that run after
|
||||
* the old lead's pane has been torn down and a fresh one launched: first,
|
||||
* for the fresh pane itself to reach a real turn boundary ({@code IDLE} or
|
||||
* {@code DONE}, never merely {@code BLOCKED}) — the safety gate, since
|
||||
* typing into a pane that has not finished booting loses the keystrokes;
|
||||
* second, for the fresh terminal to show up as a recognised lead, which is
|
||||
* bookkeeping rather than a safety gate, so a timeout on this second wait
|
||||
* does not withhold {@code bootstrapText} — it is sent once the pane is
|
||||
* ready regardless. Recognition comes from the same periodically-refreshed
|
||||
* scan {@code LeadTabScanner} already keeps ({@code scanIntervalSeconds},
|
||||
* 10s live), so a budget has to clear more than one scan interval to leave
|
||||
* any real margin for the CLI's own boot time; 20 was rejected for exactly
|
||||
* that reason — at a 10s scan interval it only buys two scans. 45 buys
|
||||
* roughly four. Only a timeout on the FIRST wait (the pane never becomes
|
||||
* ready) withholds {@code bootstrapText}.
|
||||
* @param bootstrapText default a sentence naming the RESOLVED handover path — sent to the
|
||||
* lead's pane once it settles after {@code /clear}, telling the fresh
|
||||
* session where to read the handover and carry on. Left {@code null} here
|
||||
* when the operator configures none: the default sentence cannot be built
|
||||
* at construction time because it must name the path AFTER {@code
|
||||
* dev.ltms.fleet.lead.LeadRollover#open} has resolved a relative {@code
|
||||
* handoverPath} against the calling lead's workspace, which this record has
|
||||
* no way to know — see {@link #bootstrapTextFor(String)}.
|
||||
* fresh lead's pane once it reaches a real turn boundary after relaunch,
|
||||
* telling the fresh session where to read the handover and carry on. Left
|
||||
* {@code null} here when the operator configures none: the default sentence
|
||||
* cannot be built at construction time because it must name the path AFTER
|
||||
* {@code dev.ltms.fleet.lead.LeadRollover#open} has resolved a relative
|
||||
* {@code handoverPath} against the calling lead's workspace, which this
|
||||
* record has no way to know — see {@link #bootstrapTextFor(String)}.
|
||||
*/
|
||||
@JsonIgnoreProperties(ignoreUnknown = true)
|
||||
public record LeadRollover(String handoverPath, Boolean requireOperatorConfirm,
|
||||
Integer maxDocAgeSeconds, Integer turnSettleSeconds,
|
||||
Integer clearSettleSeconds, String bootstrapText) {
|
||||
Integer relaunchReadySeconds, String bootstrapText) {
|
||||
public LeadRollover {
|
||||
requireOperatorConfirm = requireOperatorConfirm == null || requireOperatorConfirm;
|
||||
maxDocAgeSeconds = (maxDocAgeSeconds == null || maxDocAgeSeconds <= 0) ? 3600 : maxDocAgeSeconds;
|
||||
turnSettleSeconds = (turnSettleSeconds == null || turnSettleSeconds <= 0) ? 300 : turnSettleSeconds;
|
||||
clearSettleSeconds = (clearSettleSeconds == null || clearSettleSeconds <= 0) ? 20 : clearSettleSeconds;
|
||||
relaunchReadySeconds = (relaunchReadySeconds == null || relaunchReadySeconds <= 0)
|
||||
? 45 : relaunchReadySeconds;
|
||||
bootstrapText = (bootstrapText == null || bootstrapText.isBlank()) ? null : bootstrapText;
|
||||
}
|
||||
|
||||
/**
|
||||
* The text actually sent to the lead's pane once it settles after {@code /clear}: the
|
||||
* operator's configured {@link #bootstrapText} when one is set, otherwise the default
|
||||
* sentence built from {@code resolvedHandoverPath}.
|
||||
* The text actually sent to the fresh lead's pane once it reaches a real turn boundary
|
||||
* after relaunch: the operator's configured {@link #bootstrapText} when one is set,
|
||||
* otherwise the default sentence built from {@code resolvedHandoverPath}.
|
||||
*
|
||||
* @param resolvedHandoverPath the ABSOLUTE path {@code dev.ltms.fleet.lead.LeadRollover
|
||||
* #open} already resolved — never the raw configured {@link
|
||||
@@ -1919,6 +1931,7 @@ public record FleetConfig(
|
||||
rejectNegativeMaxLoad(yaml);
|
||||
rejectAutoCompactWindowOutOfRange(yaml);
|
||||
warnConflictingAutoCompactWindows(yaml);
|
||||
warnRetiredClearSettleSecondsKey(yaml);
|
||||
rejectMalformedProfilePatterns(yaml);
|
||||
rejectUnknownKind(yaml);
|
||||
rejectUnknownAuthMode(yaml);
|
||||
@@ -2342,6 +2355,33 @@ public record FleetConfig(
|
||||
names, String.join(", ", detail));
|
||||
}
|
||||
|
||||
/**
|
||||
* Warn when a {@code leadRollover:} block still sets the retired {@code clearSettleSeconds}
|
||||
* key. {@link LeadRollover} carries {@code @JsonIgnoreProperties(ignoreUnknown = true)} and no
|
||||
* longer declares that component, so Jackson drops it with no signal of its own — this raw-YAML
|
||||
* check is the only place an operator's now-inert setting is reported at all; by the time a
|
||||
* {@link LeadRollover} instance exists to run a validator against, the key is already gone.
|
||||
*
|
||||
* @param yaml the raw config text
|
||||
*/
|
||||
static void warnRetiredClearSettleSecondsKey(String yaml) {
|
||||
Map<?, ?> raw;
|
||||
try {
|
||||
raw = YAML.readValue(yaml, Map.class);
|
||||
} catch (IOException | IllegalArgumentException e) {
|
||||
return;
|
||||
}
|
||||
if (raw == null || !(raw.get("leadRollover") instanceof Map<?, ?> leadRollover)) {
|
||||
return;
|
||||
}
|
||||
if (leadRollover.containsKey("clearSettleSeconds")) {
|
||||
log.warn("leadRollover.clearSettleSeconds is retired and no longer read. Set "
|
||||
+ "leadRollover.relaunchReadySeconds instead: it bounds how long to wait, after "
|
||||
+ "a lead is relaunched, for its pane to become ready and then for it to be "
|
||||
+ "recognised as a lead. Remove clearSettleSeconds from fleetd.yaml.");
|
||||
}
|
||||
}
|
||||
|
||||
/**
|
||||
* Reject a profile whose {@code errorPattern} (fleetd #201 Unit 5) or {@code exhaustedPattern}
|
||||
* (CB-578 stage A) is not a valid Java regex, naming the profile, the key, and the parser's own
|
||||
|
||||
@@ -1,8 +1,11 @@
|
||||
package dev.ltms.fleet.lead;
|
||||
|
||||
import dev.ltms.fleet.config.FleetConfig;
|
||||
import dev.ltms.fleet.herdr.Agent;
|
||||
import dev.ltms.fleet.herdr.AgentControl;
|
||||
import dev.ltms.fleet.herdr.AgentStatus;
|
||||
import dev.ltms.fleet.herdr.HerdrException;
|
||||
import dev.ltms.fleet.herdr.WorkspaceControl;
|
||||
import org.slf4j.Logger;
|
||||
import org.slf4j.LoggerFactory;
|
||||
|
||||
@@ -24,60 +27,69 @@ import java.util.function.Supplier;
|
||||
/**
|
||||
* fleetd #480: replace a lead session that has decided it is ready to be rolled over, without an
|
||||
* operator doing it by hand. A lead writes a handover file, calls {@link #open}, and then — once
|
||||
* every gate ({@link #confirm}'s own checks) has passed — a deferred, single-shot continuation
|
||||
* clears the lead's own pane and bootstraps a fresh session against that file.
|
||||
* every gate ({@link #confirm}'s own checks) has passed — a deferred, single-shot continuation ends
|
||||
* the lead's own pane, launches a fresh one, and bootstraps that fresh session against the file.
|
||||
*
|
||||
* <p>This is the executor behind the {@code fleet_handover} MCP tool ({@code
|
||||
* dev.ltms.fleet.mcp.FleetMcp#handover}), which drives {@link #open}, {@link #confirm}, {@link
|
||||
* #cancel}, and {@link #status} from a tool call — wired in fleetd #480 Unit C. <strong>An earlier
|
||||
* version of this paragraph said nothing called this class at all; that stopped being true once
|
||||
* that unit landed, and this correction exists so the javadoc does not go on claiming it.</strong>
|
||||
* #cancel}, and {@link #status} from a tool call.
|
||||
*
|
||||
* <p><strong>{@code confirm()} cannot roll inline — a fleetd #480 correction.</strong> The first
|
||||
* version of this class called {@code agents.send(lead, "/clear")} directly from inside {@code
|
||||
* confirm()}, then polled for the pane to become injectable again. That is wrong, because {@code
|
||||
* confirm()} is called BY the lead, FROM the lead's own turn: the lead's pane is {@code WORKING}
|
||||
* for the whole duration of that call and cannot possibly report injectable until {@code confirm()}
|
||||
* itself returns. The poll always timed out — but only after the {@code /clear} had already been
|
||||
* sent and queued in the pane, where it fired the instant the turn ended anyway. The result was the
|
||||
* worst outcome this feature can produce: a silently destroyed lead context with no fresh session
|
||||
* ever started, and a refusal return value that claimed nothing had happened.
|
||||
*
|
||||
* <p>The fix: {@link #confirm} validates every gate, then does no I/O against the lead's own pane
|
||||
* at all — it only records that the request is approved and hands a one-shot continuation to
|
||||
* {@code continuationRunner} before returning. That continuation is what actually touches the pane,
|
||||
* once the calling turn has ended, in this order:
|
||||
* <p><strong>{@code confirm()} cannot roll inline.</strong> {@code confirm()} is called BY the
|
||||
* lead, FROM the lead's own turn: the lead's pane is {@code WORKING} for the whole duration of that
|
||||
* call and cannot possibly report a real turn boundary until {@code confirm()} itself returns. So
|
||||
* {@link #confirm} validates every gate, then does no I/O against the lead's own pane at all — it
|
||||
* only records that the request is approved and hands a one-shot continuation to {@code
|
||||
* continuationRunner} before returning. That continuation is what actually touches the pane, once
|
||||
* the calling turn has ended, in this order:
|
||||
* <ol>
|
||||
* <li>wait for the lead's own pane to report a real turn boundary — {@code IDLE} or {@code
|
||||
* DONE}, never merely {@code BLOCKED} — i.e. wait for the very {@code confirm()} call that
|
||||
* approved this roll to finish its turn — bounded by {@code turnSettleSeconds}. <strong>If
|
||||
* this never happens, nothing else in this list runs: no {@code /clear} is ever sent.</strong>
|
||||
* A lead that never goes idle is a lead still doing real work, and clearing it would throw
|
||||
* away live context — exactly the failure this correction exists to prevent.</li>
|
||||
* <li>{@code agents.send(lead, "/clear")}</li>
|
||||
* <li>wait for {@code /clear} to be picked up and settle, bounded by {@code clearSettleSeconds}
|
||||
* (fleetd #489: no longer a plain re-check of the same boundary — {@code /clear} starts no
|
||||
* turn of its own, so this instead nudges the submit keystroke while no pickup has been seen,
|
||||
* then waits for a real {@code WORKING} → {@code IDLE}/{@code DONE} boundary once one has;
|
||||
* see {@link #waitForClearPickupAndSettle})</li>
|
||||
* <li>{@code agents.send(lead, cfg.bootstrapTextFor(p.handoverPath()))}</li>
|
||||
* this never happens, nothing else in this list runs: the old pane is never touched.</strong>
|
||||
* A lead that never goes idle is a lead still doing real work, and tearing it down would throw
|
||||
* away live context.</li>
|
||||
* <li>capture the old pane id (and, through it, the old tab) from {@link AgentControl#get}, with
|
||||
* a bounded retry — the terminal-to-pane lookup it goes through can itself report a genuinely
|
||||
* live agent as not found (see {@code AgentControl#agentCall}'s own re-resolve-once
|
||||
* behaviour), and one false negative here must not abort an otherwise-healthy roll. Neither id
|
||||
* is ever re-resolved from the terminal again after this — once the pane below is closed there
|
||||
* is nothing left to resolve it from.</li>
|
||||
* <li>resolve the lead's configured name from its terminal, for the relaunch step below.</li>
|
||||
* <li>end the old session: close the pane (an already-gone pane counts as success; any other
|
||||
* failure propagates), then close its tab only when the pane was that tab's sole occupant —
|
||||
* the same pane-then-tab teardown {@code HerdrPeerLauncher#stop} uses for a member.</li>
|
||||
* <li>confirm the old pane is actually gone by polling {@link
|
||||
* dev.ltms.fleet.herdr.WorkspaceControl#locatePane} for a {@code null} result — never {@link
|
||||
* AgentControl#status}, and never the live-lead terminal map, each of which answers a
|
||||
* different question. <strong>If the old pane is never confirmed gone, no relaunch is
|
||||
* attempted</strong> — see {@link RollState#OLD_PANE_NEVER_DIED}.</li>
|
||||
* <li>launch a fresh lead with {@code LeadLauncher#relaunch}. <strong>If every attempt fails,
|
||||
* {@code bootstrapText} is never sent</strong> — see {@link RollState#RELAUNCH_FAILED}.</li>
|
||||
* <li>wait for the fresh pane to reach a real turn boundary ({@code IDLE} or {@code DONE},
|
||||
* never merely {@code BLOCKED}), bounded by {@code relaunchReadySeconds}. This is the
|
||||
* safety gate: typing into a pane that has not actually finished booting loses the
|
||||
* keystrokes. <strong>If the pane never becomes ready, {@code bootstrapText} is never
|
||||
* sent</strong> — see {@link RollState#RELAUNCH_NEVER_READY}.</li>
|
||||
* <li>wait for the fresh terminal to be recognised as a live lead — present in the live-lead
|
||||
* terminal map — bounded by {@code relaunchReadySeconds}. This is bookkeeping, not a
|
||||
* safety gate: {@code bootstrapText} is sent either way once the pane is ready, whether or
|
||||
* not this wait itself times out — see {@link RollState#RELAUNCH_NOT_RECOGNISED}.</li>
|
||||
* <li>{@code agents.send(newTerminal, cfg.bootstrapTextFor(p.handoverPath()))} — sent to the
|
||||
* FRESH terminal, never the one that was just torn down.</li>
|
||||
* </ol>
|
||||
* A {@link #confirm} that returns {@link RollDecision#approved()} therefore means <em>"every gate
|
||||
* passed and the roll is scheduled"</em>, never <em>"the pane has been cleared"</em> — the pane may
|
||||
* still be mid-turn, possibly for a long time, when the caller gets that answer back.
|
||||
* passed and the roll is scheduled"</em>, never <em>"the lead has already been replaced"</em> — the
|
||||
* old pane may still be mid-turn, possibly for a long time, when the caller gets that answer back.
|
||||
*
|
||||
* <p><strong>The safety invariant survives this change, restated precisely.</strong> The ticket
|
||||
* that first defined this class required "no timer, no scheduler, no background thread" so that
|
||||
* nothing but an explicit {@link #confirm} call could ever cause a {@code /clear}. That invariant
|
||||
* is about INITIATIVE, not about synchronicity, and this correction keeps it: {@code
|
||||
* continuationRunner} launches a single-shot task that exists only because one specific,
|
||||
* <p><strong>The safety invariant.</strong> "No timer, no scheduler, no background thread" means
|
||||
* that nothing but an explicit {@link #confirm} call can ever tear a lead's pane down.
|
||||
* {@code continuationRunner} launches a single-shot task that exists only because one specific,
|
||||
* already-approved {@link #confirm} call created it — it is not recurring, it is not started at
|
||||
* construction time or on any schedule, and no two invocations of it ever share state. A recurring
|
||||
* heartbeat or timer that could decide on its own initiative to roll a pane is still, and will
|
||||
* always be, absent from this class. <strong>Nothing but an explicit {@link #confirm} call that
|
||||
* passes every gate can ever cause a {@code /clear} — that call may simply finish its own work
|
||||
* slightly later than the method return, as a continuation of the same approved request, rather
|
||||
* than entirely inside the method body.</strong>
|
||||
* heartbeat or timer that could decide on its own initiative to roll a pane is absent from this
|
||||
* class. <strong>Nothing but an explicit {@link #confirm} call that passes every gate can ever tear
|
||||
* a pane down — that call may simply finish its own work slightly later than the method return, as
|
||||
* a continuation of the same approved request, rather than entirely inside the method body.</strong>
|
||||
*
|
||||
* <p><strong>Identity is resolved by the caller, never looked up here — a second fleetd #480
|
||||
* correction.</strong> The first version resolved the pane to clear via {@code
|
||||
@@ -115,19 +127,17 @@ public final class LeadRollover {
|
||||
|
||||
private static final Logger log = LoggerFactory.getLogger(LeadRollover.class);
|
||||
|
||||
/** Poll interval while waiting for the lead's pane to settle after {@code /clear}. */
|
||||
static final long SETTLE_POLL_MS = 250;
|
||||
/** Poll interval shared by every bounded wait in this class. */
|
||||
static final long POLL_INTERVAL_MS = 250;
|
||||
|
||||
/**
|
||||
* How many consecutive not-yet-picked-up polls {@link #waitForClearPickupAndSettle} allows
|
||||
* before releasing rather than wedging the roll — the same constant and the same
|
||||
* release-not-wedge choice {@link dev.ltms.fleet.inject.Injector} already makes for its own
|
||||
* post-turn {@code /clear} housekeeping (fleetd #306). <strong>This bounds the number of
|
||||
* consecutive polls, not the number of nudges:</strong> the first {@code PICKUP_GRACE_POLLS - 1}
|
||||
* of those polls each send a nudge, and the {@code PICKUP_GRACE_POLLS}th releases instead of
|
||||
* nudging again — so 8 polls produce 7 nudges, not 8.
|
||||
* How long {@link #waitUntilPaneGone} polls {@link WorkspaceControl#locatePane} before giving
|
||||
* up on ever seeing the old pane disappear. Not configurable: once {@link #endOldSession} has
|
||||
* closed the pane (and, usually, its tab), herdr dropping the pane from its own bookkeeping is
|
||||
* expected to show up within one or two polls, not on an operator-tunable timescale the way a
|
||||
* CLI boot is.
|
||||
*/
|
||||
static final int PICKUP_GRACE_POLLS = 8;
|
||||
static final int PANE_DEATH_TIMEOUT_SECONDS = 10;
|
||||
|
||||
/**
|
||||
* One request opened by {@link #open}, pending its {@link #confirm} (or {@link #cancel}).
|
||||
@@ -175,10 +185,10 @@ public final class LeadRollover {
|
||||
|
||||
/**
|
||||
* The outcome of a {@link #confirm} call. {@link #approved()} means every gate passed and the
|
||||
* roll has been handed to a one-shot continuation — <strong>not</strong> that the pane has been
|
||||
* cleared; the continuation may still be waiting for the calling turn to end when this returns.
|
||||
* Whether the deferred roll itself later goes on to clear the pane, refuse for never going
|
||||
* idle, or refuse for never re-settling after {@code /clear} is logged only (see this class's
|
||||
* roll has been handed to a one-shot continuation — <strong>not</strong> that the lead has
|
||||
* already been replaced; the continuation may still be waiting for the calling turn to end when
|
||||
* this returns. Whether the deferred roll itself later goes on to tear the old pane down and
|
||||
* relaunch the lead, or refuses at any of its own steps, is logged only (see this class's
|
||||
* javadoc) — there is deliberately no synchronous caller left by that point to hand a result to.
|
||||
*/
|
||||
public record RollDecision(boolean accepted, RefusalReason reason, String detail) {
|
||||
@@ -214,7 +224,7 @@ public final class LeadRollover {
|
||||
|
||||
/**
|
||||
* What is known about one token, right now — the answer {@link #status} gives. Distinguishes
|
||||
* three terminal outcomes an approved roll can finish with, one in-flight outcome for a roll
|
||||
* five terminal outcomes an approved roll can finish with, one in-flight outcome for a roll
|
||||
* that has been approved but has not finished yet, and two answers for a token that names no
|
||||
* active work at all: still pending confirmation, or nothing known about this token at all.
|
||||
*/
|
||||
@@ -237,41 +247,64 @@ public final class LeadRollover {
|
||||
* #status} could wrongly answer {@link #UNKNOWN} ("nothing was ever requested") for a roll
|
||||
* that is, in fact, actively running. This is not sticky: the deferred continuation
|
||||
* overwrites this same entry with a terminal state ({@link #ROLLED}, {@link
|
||||
* #TURN_NEVER_SETTLED}, {@link #CLEAR_NEVER_SETTLED}, or {@link #FAILED}) once it finishes
|
||||
* — including by throwing, which fleetd #615's catch in {@link #runRollover} now turns into
|
||||
* {@link #FAILED} instead of leaving this entry stuck forever.
|
||||
* #TURN_NEVER_SETTLED}, {@link #OLD_PANE_NEVER_DIED}, {@link #RELAUNCH_FAILED}, {@link
|
||||
* #RELAUNCH_NOT_RECOGNISED}, or {@link #FAILED}) once it finishes — including by throwing,
|
||||
* which {@link #runRollover}'s catch turns into {@link #FAILED} instead of leaving this
|
||||
* entry stuck forever.
|
||||
*/
|
||||
IN_PROGRESS,
|
||||
/**
|
||||
* {@link #confirm} was approved and the deferred continuation completed the entire roll:
|
||||
* the calling lead's turn settled, {@code /clear} was sent and settled, and {@code
|
||||
* bootstrapText} was sent.
|
||||
* {@link #confirm} was approved and the deferred continuation completed the entire roll: the
|
||||
* calling lead's turn settled, the old pane was torn down and confirmed gone, a fresh lead
|
||||
* was launched and recognised, and {@code bootstrapText} was sent to it.
|
||||
*/
|
||||
ROLLED,
|
||||
/**
|
||||
* {@link #confirm} was approved, but the calling lead's own turn never reached a boundary
|
||||
* (IDLE or DONE) within {@code turnSettleSeconds} — no {@code /clear} was ever sent, at
|
||||
* all. This is the branch the fleetd #480 correction exists to make safe, and the one this
|
||||
* status exists to make VISIBLE: before this, a lead that hit this case had no way to find
|
||||
* out, and would carry on believing it was about to be replaced. See this class's javadoc.
|
||||
* (IDLE or DONE) within {@code turnSettleSeconds} — the old pane was never touched at all.
|
||||
* This is the state that makes a lead's own stuck turn VISIBLE: without it, a lead that hit
|
||||
* this case would have no way to find out, and would carry on believing it was about to be
|
||||
* replaced. See this class's javadoc.
|
||||
*/
|
||||
TURN_NEVER_SETTLED,
|
||||
/**
|
||||
* {@link #confirm} was approved and {@code /clear} was sent, but the pane never re-settled
|
||||
* within {@code clearSettleSeconds} — {@code bootstrapText} was never sent.
|
||||
* {@link #confirm} was approved and the calling lead's turn settled, the old pane was closed
|
||||
* (and its tab, if it was the sole occupant), but {@link
|
||||
* dev.ltms.fleet.herdr.WorkspaceControl#locatePane} kept reporting it as still present for
|
||||
* the whole pane-death timeout. No relaunch was ever attempted, and {@code bootstrapText}
|
||||
* was never sent.
|
||||
*/
|
||||
CLEAR_NEVER_SETTLED,
|
||||
OLD_PANE_NEVER_DIED,
|
||||
/**
|
||||
* fleetd #615: the deferred continuation threw a {@link RuntimeException} — most likely a
|
||||
* {@link dev.ltms.fleet.herdr.HerdrException} out of one of the two unwrapped {@code
|
||||
* agents.send} calls in {@link #runRollover} — and the continuation thread died with it.
|
||||
* Before this state existed, that throw left {@link #outcomes} holding {@link #IN_PROGRESS}
|
||||
* forever, because the production {@code continuationRunner} is a bare virtual thread with
|
||||
* no uncaught-exception handler and nothing downstream of the throw ever ran to write a
|
||||
* terminal outcome. {@code detail} names the exception, so a reader has something to act on
|
||||
* — the same diagnostic style as {@link #TURN_NEVER_SETTLED} and {@link
|
||||
* #CLEAR_NEVER_SETTLED}. The roll is dead at this point and does not retry itself; a stuck
|
||||
* lead must {@link #open} a fresh request.
|
||||
* The old pane was confirmed gone, but {@code LeadLauncher#relaunch} returned {@code null}
|
||||
* — every launch attempt failed. {@code bootstrapText} was never sent, and no fresh terminal
|
||||
* exists for this roll to have recognised.
|
||||
*/
|
||||
RELAUNCH_FAILED,
|
||||
/**
|
||||
* A fresh lead was launched, but its pane never reached a real turn boundary ({@code IDLE}
|
||||
* or {@code DONE}, never merely {@code BLOCKED}) within {@code relaunchReadySeconds} — the
|
||||
* CLI never finished booting, or it stayed paused on a startup prompt. {@code bootstrapText}
|
||||
* was never sent: typing into a pane that is not actually ready to accept input loses the
|
||||
* keystrokes.
|
||||
*/
|
||||
RELAUNCH_NEVER_READY,
|
||||
/**
|
||||
* A fresh lead was launched and its pane reached a real turn boundary, so {@code
|
||||
* bootstrapText} WAS sent to it, but the terminal was never recognised as a live lead —
|
||||
* present in the live-lead terminal map — within {@code relaunchReadySeconds}. The session
|
||||
* itself is alive and bootstrapped; only the daemon's own bookkeeping has not caught up, and
|
||||
* an operator should check why the tab was not recognised.
|
||||
*/
|
||||
RELAUNCH_NOT_RECOGNISED,
|
||||
/**
|
||||
* The deferred continuation threw a {@link RuntimeException} and the continuation thread
|
||||
* died with it. Without this state, that throw would leave {@link #outcomes} holding {@link
|
||||
* #IN_PROGRESS} forever, because the production {@code continuationRunner} is a bare virtual
|
||||
* thread with no uncaught-exception handler and nothing downstream of the throw ever runs to
|
||||
* write a terminal outcome. {@code detail} names the exception, so a reader has something to
|
||||
* act on. The roll is dead at this point and does not retry itself; a stuck lead must
|
||||
* {@link #open} a fresh request.
|
||||
*/
|
||||
FAILED,
|
||||
/**
|
||||
@@ -292,6 +325,10 @@ public final class LeadRollover {
|
||||
public record RollStatus(RollState state, String detail) {}
|
||||
|
||||
private final AgentControl agents;
|
||||
/** Workspace/tab/pane control — used to tear down the old pane and confirm it is gone. */
|
||||
private final WorkspaceControl spaces;
|
||||
/** Starts the fresh lead that replaces the one this roll tears down. */
|
||||
private final LeadLauncher launcher;
|
||||
private final Supplier<FleetConfig.LeadRollover> configSupplier;
|
||||
/**
|
||||
* Terminal id → that lead's configured workspace directory (their {@code
|
||||
@@ -301,8 +338,21 @@ public final class LeadRollover {
|
||||
* daemon-cwd bug this parameter exists to fix.
|
||||
*/
|
||||
private final Function<String, String> leadWorkspace;
|
||||
/**
|
||||
* Terminal id → that lead's configured name under {@code fleet.leaders}, or {@code null} when
|
||||
* the terminal names no currently-recognised lead. The deferred continuation calls this, on the
|
||||
* OLD terminal, before tearing it down, so it knows which lead to pass to {@link
|
||||
* LeadLauncher#relaunch}.
|
||||
*/
|
||||
private final Function<String, String> leadNameForTerminal;
|
||||
/**
|
||||
* The daemon's current terminal id → lead name map, read fresh on every poll. The deferred
|
||||
* continuation polls this for the FRESH terminal {@link LeadLauncher#relaunch} returns, to
|
||||
* learn when that terminal has been recognised as a live lead — see this class's javadoc.
|
||||
*/
|
||||
private final Supplier<Map<String, String>> liveLeadTerminals;
|
||||
private final LongSupplier nowMillis;
|
||||
private final Runnable settleSleeper;
|
||||
private final Runnable pollSleeper;
|
||||
/**
|
||||
* Launches the post-{@code confirm()} continuation. Production uses a single unstarted virtual
|
||||
* thread per confirmed request — see this class's javadoc for why that is a single-shot task,
|
||||
@@ -338,29 +388,40 @@ public final class LeadRollover {
|
||||
}
|
||||
});
|
||||
|
||||
/** Production constructor — wall clock, real sleep between settle polls, a real virtual thread. */
|
||||
public LeadRollover(AgentControl agents, Supplier<FleetConfig.LeadRollover> configSupplier,
|
||||
Function<String, String> leadWorkspace) {
|
||||
this(agents, configSupplier, leadWorkspace, System::currentTimeMillis,
|
||||
() -> sleepUninterruptibly(SETTLE_POLL_MS),
|
||||
/** Production constructor — wall clock, real sleep between polls, a real virtual thread. */
|
||||
public LeadRollover(AgentControl agents, WorkspaceControl spaces, LeadLauncher launcher,
|
||||
Supplier<FleetConfig.LeadRollover> configSupplier,
|
||||
Function<String, String> leadWorkspace,
|
||||
Function<String, String> leadNameForTerminal,
|
||||
Supplier<Map<String, String>> liveLeadTerminals) {
|
||||
this(agents, spaces, launcher, configSupplier, leadWorkspace, leadNameForTerminal,
|
||||
liveLeadTerminals, System::currentTimeMillis,
|
||||
() -> sleepUninterruptibly(POLL_INTERVAL_MS),
|
||||
r -> Thread.ofVirtual().name("lead-rollover-continuation-").start(r));
|
||||
}
|
||||
|
||||
/**
|
||||
* Full constructor — an injectable wall-clock supplier, settle-poll sleeper, and continuation
|
||||
* runner, for tests. {@code nowMillis} MUST be a wall-clock source (e.g. {@code
|
||||
* Full constructor — an injectable wall-clock supplier, poll sleeper, and continuation runner,
|
||||
* for tests. {@code nowMillis} MUST be a wall-clock source (e.g. {@code
|
||||
* System.currentTimeMillis()}), never {@code System.nanoTime()}: the freshness check compares
|
||||
* against a file's modified time, which only a wall clock is comparable to, and {@code
|
||||
* nanoTime} freezes while the host sleeps (fleetd #386).
|
||||
* nanoTime} freezes while the host sleeps.
|
||||
*/
|
||||
LeadRollover(AgentControl agents, Supplier<FleetConfig.LeadRollover> configSupplier,
|
||||
Function<String, String> leadWorkspace, LongSupplier nowMillis,
|
||||
Runnable settleSleeper, Consumer<Runnable> continuationRunner) {
|
||||
LeadRollover(AgentControl agents, WorkspaceControl spaces, LeadLauncher launcher,
|
||||
Supplier<FleetConfig.LeadRollover> configSupplier,
|
||||
Function<String, String> leadWorkspace,
|
||||
Function<String, String> leadNameForTerminal,
|
||||
Supplier<Map<String, String>> liveLeadTerminals,
|
||||
LongSupplier nowMillis, Runnable pollSleeper, Consumer<Runnable> continuationRunner) {
|
||||
this.agents = agents;
|
||||
this.spaces = spaces;
|
||||
this.launcher = launcher;
|
||||
this.configSupplier = configSupplier;
|
||||
this.leadWorkspace = leadWorkspace;
|
||||
this.leadNameForTerminal = leadNameForTerminal;
|
||||
this.liveLeadTerminals = liveLeadTerminals;
|
||||
this.nowMillis = nowMillis;
|
||||
this.settleSleeper = settleSleeper;
|
||||
this.pollSleeper = pollSleeper;
|
||||
this.continuationRunner = continuationRunner;
|
||||
}
|
||||
|
||||
@@ -519,8 +580,9 @@ public final class LeadRollover {
|
||||
// remove-then-put ordering would leave in which the token is in neither map.
|
||||
outcomes.put(token, new RollStatus(RollState.IN_PROGRESS,
|
||||
"confirm() approved this roll and handed it to the deferred continuation; it has "
|
||||
+ "not finished yet — still waiting for the calling turn to settle, for "
|
||||
+ "/clear to be sent and settle, or for bootstrapText to be sent"));
|
||||
+ "not finished yet — still waiting for the calling turn to settle, for the "
|
||||
+ "old pane to be torn down and confirmed gone, for the fresh lead to be "
|
||||
+ "recognised, or for bootstrapText to be sent"));
|
||||
pending.remove(token);
|
||||
log.info("lead-rollover: confirmed token={} lead={} — roll scheduled once the calling turn ends",
|
||||
token, callerTerminal);
|
||||
@@ -547,22 +609,19 @@ public final class LeadRollover {
|
||||
/**
|
||||
* The single-shot continuation {@link #confirm} hands to {@code continuationRunner}. Runs
|
||||
* entirely after {@link #confirm} has returned to its caller — see this class's javadoc for the
|
||||
* four-step order. There is no result to return to by this point, so every outcome is logged
|
||||
* only.
|
||||
* full order. There is no result to return to by this point, so every outcome is logged only.
|
||||
*
|
||||
* <p><strong>fleetd #615 — the whole body is wrapped in one {@code try}.</strong> The two {@code
|
||||
* agents.send} calls below are not wrapped individually: {@code send} → {@code agentCall} →
|
||||
* {@code herdr.call} can throw an unchecked {@link dev.ltms.fleet.herdr.HerdrException} (see
|
||||
* {@code AgentControl.java}), and the production {@code continuationRunner} is a bare virtual
|
||||
* thread with no uncaught-exception handler (see this class's public constructor). Before this
|
||||
* fix, either throw killed the continuation thread silently, leaving the {@link
|
||||
* RollState#IN_PROGRESS} entry {@link #confirm} wrote at hand-off stuck forever — {@link
|
||||
* #status} had no way to tell a dead roll from one still genuinely running. The {@code catch}
|
||||
* below is scoped to the method body rather than to each {@code send} call individually, so it
|
||||
* also covers anything else added to this continuation later, not just today's two call sites —
|
||||
* the same reasoning that put the write-a-terminal-outcome step at each of this method's other
|
||||
* exits (see the {@link RollState#TURN_NEVER_SETTLED} and {@link RollState#CLEAR_NEVER_SETTLED}
|
||||
* branches below) rather than inside the helpers that detect them.</p>
|
||||
* <p><strong>The whole body is wrapped in one {@code try}.</strong> Several calls below —
|
||||
* {@code agents.get}, {@code agents.close}, {@code agents.send} — can throw an unchecked {@link
|
||||
* dev.ltms.fleet.herdr.HerdrException} (see {@code AgentControl.java}), and the production
|
||||
* {@code continuationRunner} is a bare virtual thread with no uncaught-exception handler (see
|
||||
* this class's public constructor). An uncaught throw would kill the continuation thread
|
||||
* silently, leaving the {@link RollState#IN_PROGRESS} entry {@link #confirm} wrote at hand-off
|
||||
* stuck forever — {@link #status} would have no way to tell a dead roll from one still
|
||||
* genuinely running. The {@code catch} below is scoped to the method body rather than to each
|
||||
* call individually, so it also covers every call in this continuation, not a fixed list of
|
||||
* call sites — the same reasoning that put the write-a-terminal-outcome step at each of this
|
||||
* method's other exits rather than inside the helpers that detect them.</p>
|
||||
*
|
||||
* <p>Only {@link RuntimeException} is caught, matching the local convention {@link
|
||||
* #waitUntilAtTurnBoundary} already set around its own {@code agents.status} call — not the
|
||||
@@ -593,53 +652,237 @@ public final class LeadRollover {
|
||||
private void runRolloverUnguarded(PendingRollover p, FleetConfig.LeadRollover cfg) {
|
||||
String lead = p.leadTerminal();
|
||||
long rollStartMillis = nowMillis.getAsLong();
|
||||
|
||||
TurnSettleResult turnResult = waitUntilAtTurnBoundary(lead, cfg.turnSettleSeconds());
|
||||
if (!turnResult.settled()) {
|
||||
// fleetd #494 follow-up: this line had the SAME defect as the /clear-timeout line below
|
||||
// — cfg.turnSettleSeconds() is the CONFIGURED budget, not how long this wait actually
|
||||
// ran. Print the measured elapsed time alongside it, labelled, exactly like the /clear
|
||||
// path already does.
|
||||
log.warn("lead-rollover: pane {} did not reach a turn boundary (IDLE or DONE) after "
|
||||
+ "confirm() — refusing to send /clear at all; the calling lead's own "
|
||||
+ "turn is still live and clearing it now would destroy live context "
|
||||
+ "confirm() — the old pane is never touched; the calling lead's own "
|
||||
+ "turn is still live and tearing it down now would destroy live context "
|
||||
+ "(token={}, configured={}s elapsed={}ms)",
|
||||
lead, p.token(), cfg.turnSettleSeconds(), turnResult.elapsedMillis());
|
||||
outcomes.put(p.token(), new RollStatus(RollState.TURN_NEVER_SETTLED,
|
||||
"the calling lead's own turn never reached a boundary (IDLE or DONE) within "
|
||||
+ "turnSettleSeconds=" + cfg.turnSettleSeconds() + "s (measured elapsed="
|
||||
+ turnResult.elapsedMillis() + "ms) — no /clear was ever sent. If this "
|
||||
+ "keeps happening, raise turnSettleSeconds in fleetd.yaml"));
|
||||
+ turnResult.elapsedMillis() + "ms) — the old pane was never touched. If "
|
||||
+ "this keeps happening, raise turnSettleSeconds in fleetd.yaml"));
|
||||
return;
|
||||
}
|
||||
|
||||
// This deliberately bypasses Injector, exactly like ClaudeCodeLauncher#clearContext:
|
||||
// /clear is housekeeping, not a delegated turn, and routing it through Injector wedges the
|
||||
// pane forever (see this class's javadoc).
|
||||
agents.send(lead, "/clear");
|
||||
ClearSettleResult clearResult = waitForClearPickupAndSettle(lead, cfg.clearSettleSeconds());
|
||||
if (!clearResult.settled()) {
|
||||
// fleetd #494: cfg.clearSettleSeconds() is the CONFIGURED budget, not how long the wait
|
||||
// actually ran — an operator reading only that number wrongly believes it is a measured
|
||||
// duration. Print the measured elapsed time and nudge count alongside it, each labelled,
|
||||
// so the two can be compared at a glance.
|
||||
log.warn("lead-rollover: pane {} did not reach a turn boundary (IDLE or DONE) after "
|
||||
+ "/clear — NOT sending bootstrapText (token={}, configured={}s "
|
||||
+ "elapsed={}ms nudges={})",
|
||||
lead, p.token(), cfg.clearSettleSeconds(), clearResult.elapsedMillis(),
|
||||
clearResult.nudges());
|
||||
outcomes.put(p.token(), new RollStatus(RollState.CLEAR_NEVER_SETTLED,
|
||||
"/clear was sent, but the pane never re-settled within clearSettleSeconds="
|
||||
+ cfg.clearSettleSeconds() + "s (measured elapsed=" + clearResult.elapsedMillis()
|
||||
+ "ms, nudges=" + clearResult.nudges() + ") — bootstrapText was never sent"));
|
||||
// Captured once, here, and never re-resolved from `lead` again below: once the pane is
|
||||
// closed there is nothing left for a terminal lookup to find.
|
||||
Agent oldAgent = captureAgentWithRetry(lead);
|
||||
String oldPaneId = oldAgent.paneId();
|
||||
String leadName = leadNameForTerminal.apply(lead);
|
||||
|
||||
endOldSession(oldPaneId);
|
||||
DeathResult deathResult = waitUntilPaneGone(oldPaneId);
|
||||
if (!deathResult.gone()) {
|
||||
log.warn("lead-rollover: old pane {} for lead {} was never confirmed gone after being "
|
||||
+ "closed — not attempting a relaunch (token={}, timeout={}s "
|
||||
+ "elapsed={}ms)",
|
||||
oldPaneId, lead, p.token(), PANE_DEATH_TIMEOUT_SECONDS, deathResult.elapsedMillis());
|
||||
outcomes.put(p.token(), new RollStatus(RollState.OLD_PANE_NEVER_DIED,
|
||||
"the old pane was closed, but locatePane kept reporting it as still present "
|
||||
+ "after a pane-death timeout=" + PANE_DEATH_TIMEOUT_SECONDS
|
||||
+ "s (measured elapsed=" + deathResult.elapsedMillis() + "ms) — no "
|
||||
+ "relaunch was attempted"));
|
||||
return;
|
||||
}
|
||||
agents.send(lead, cfg.bootstrapTextFor(p.handoverPath()));
|
||||
|
||||
Agent newAgent = launcher.relaunch(leadName);
|
||||
if (newAgent == null) {
|
||||
log.warn("lead-rollover: relaunch of lead '{}' (old terminal {}) failed every attempt "
|
||||
+ "— bootstrapText was never sent (token={})", leadName, lead, p.token());
|
||||
outcomes.put(p.token(), new RollStatus(RollState.RELAUNCH_FAILED,
|
||||
"lead '" + leadName + "' could not be relaunched — every attempt failed; "
|
||||
+ "bootstrapText was never sent"));
|
||||
return;
|
||||
}
|
||||
|
||||
ReadinessResult readinessResult = waitUntilPaneReady(newAgent.terminalId(),
|
||||
cfg.relaunchReadySeconds());
|
||||
if (!readinessResult.ready()) {
|
||||
log.warn("lead-rollover: fresh pane for lead '{}' (terminal {}) never reached a real "
|
||||
+ "turn boundary — bootstrapText was never sent (token={}, configured={}s "
|
||||
+ "elapsed={}ms)",
|
||||
leadName, newAgent.terminalId(), p.token(), cfg.relaunchReadySeconds(),
|
||||
readinessResult.elapsedMillis());
|
||||
outcomes.put(p.token(), new RollStatus(RollState.RELAUNCH_NEVER_READY,
|
||||
"fresh terminal " + newAgent.terminalId() + " never reached a real turn "
|
||||
+ "boundary (IDLE or DONE) within relaunchReadySeconds="
|
||||
+ cfg.relaunchReadySeconds() + "s (measured elapsed="
|
||||
+ readinessResult.elapsedMillis() + "ms) — bootstrapText was never "
|
||||
+ "sent"));
|
||||
return;
|
||||
}
|
||||
|
||||
IdentityResult identityResult = waitUntilRecognisedAsLead(newAgent.terminalId(),
|
||||
cfg.relaunchReadySeconds());
|
||||
agents.send(newAgent.terminalId(), cfg.bootstrapTextFor(p.handoverPath()));
|
||||
if (!identityResult.ready()) {
|
||||
log.warn("lead-rollover: fresh terminal {} for lead '{}' is alive and bootstrapped, but "
|
||||
+ "was never recognised as a live lead — an operator should check why "
|
||||
+ "the tab was not recognised (token={}, configured={}s elapsed={}ms)",
|
||||
newAgent.terminalId(), leadName, p.token(), cfg.relaunchReadySeconds(),
|
||||
identityResult.elapsedMillis());
|
||||
outcomes.put(p.token(), new RollStatus(RollState.RELAUNCH_NOT_RECOGNISED,
|
||||
"bootstrapText was sent to fresh terminal " + newAgent.terminalId() + ", but "
|
||||
+ "it was never recognised as a live lead within relaunchReadySeconds="
|
||||
+ cfg.relaunchReadySeconds() + "s (measured elapsed="
|
||||
+ identityResult.elapsedMillis() + "ms) — check why the tab was not "
|
||||
+ "recognised"));
|
||||
return;
|
||||
}
|
||||
|
||||
long rollElapsedMillis = nowMillis.getAsLong() - rollStartMillis;
|
||||
log.info("lead-rollover: rolled token={} lead={} elapsedMs={}", p.token(), lead, rollElapsedMillis);
|
||||
log.info("lead-rollover: rolled token={} oldLead={} newTerminal={} elapsedMs={}",
|
||||
p.token(), lead, newAgent.terminalId(), rollElapsedMillis);
|
||||
outcomes.put(p.token(), new RollStatus(RollState.ROLLED,
|
||||
"rolled successfully in " + rollElapsedMillis + "ms"));
|
||||
"rolled successfully in " + rollElapsedMillis + "ms; new terminal="
|
||||
+ newAgent.terminalId()));
|
||||
}
|
||||
|
||||
/** Attempts {@link #captureAgentWithRetry} makes before letting the failure propagate. */
|
||||
static final int CAPTURE_RETRIES = 3;
|
||||
|
||||
/**
|
||||
* {@link AgentControl#get} for {@code lead}, retried up to {@link #CAPTURE_RETRIES} times. The
|
||||
* terminal-to-pane lookup it goes through can report a genuinely live agent as not found (see
|
||||
* {@code AgentControl#agentCall}'s own re-resolve-once behaviour), and one such false negative
|
||||
* must not abort an otherwise-healthy roll. The result is captured once by the caller and never
|
||||
* looked up again — see this class's javadoc.
|
||||
*
|
||||
* @throws RuntimeException the last failure, if every attempt fails — {@link #runRollover}'s
|
||||
* catch turns that into {@link RollState#FAILED}
|
||||
*/
|
||||
private Agent captureAgentWithRetry(String lead) {
|
||||
RuntimeException last = null;
|
||||
for (int attempt = 1; attempt <= CAPTURE_RETRIES; attempt++) {
|
||||
try {
|
||||
return agents.get(lead);
|
||||
} catch (RuntimeException e) {
|
||||
last = e;
|
||||
log.debug("lead-rollover: agents.get({}) failed on attempt {}/{}: {}",
|
||||
lead, attempt, CAPTURE_RETRIES, e.toString());
|
||||
if (attempt < CAPTURE_RETRIES) {
|
||||
pollSleeper.run();
|
||||
}
|
||||
}
|
||||
}
|
||||
throw last;
|
||||
}
|
||||
|
||||
/**
|
||||
* End the old lead's session: close its pane, then close its tab only when the pane was that
|
||||
* tab's sole occupant — the same pane-then-tab teardown {@code HerdrPeerLauncher#stop} uses for
|
||||
* a member. An already-gone pane counts as success; any other {@code agents.close} failure
|
||||
* propagates, so a genuinely failed teardown is never reported as done. A failing
|
||||
* {@code spaces.closeTab} never propagates — by the time it runs the pane is already closed, so
|
||||
* it is cosmetic tidying, not a real teardown failure.
|
||||
*/
|
||||
private void endOldSession(String paneId) {
|
||||
WorkspaceControl.PaneLocation loc = spaces.locatePane(paneId);
|
||||
try {
|
||||
agents.close(paneId);
|
||||
} catch (HerdrException e) {
|
||||
if (!isAlreadyGone(e)) {
|
||||
throw e;
|
||||
}
|
||||
log.debug("lead-rollover: pane.close({}) ignored — already gone: {}", paneId, e.getMessage());
|
||||
}
|
||||
if (loc != null && loc.tabPaneCount() == 1) {
|
||||
try {
|
||||
spaces.closeTab(loc.tabId());
|
||||
} catch (RuntimeException e) {
|
||||
log.warn("lead-rollover: tab.close({}) failed — the pane is already torn down, so "
|
||||
+ "continuing; the tab may need manual cleanup: {}", loc.tabId(), e.getMessage());
|
||||
}
|
||||
} else if (loc != null) {
|
||||
log.debug("lead-rollover: not closing tab {} — it holds {} panes (not a dedicated lead "
|
||||
+ "tab)", loc.tabId(), loc.tabPaneCount());
|
||||
}
|
||||
}
|
||||
|
||||
/** True when a herdr error means the target is already gone (safe to treat as done). */
|
||||
private static boolean isAlreadyGone(HerdrException e) {
|
||||
return e.code() != null && e.code().endsWith("_not_found");
|
||||
}
|
||||
|
||||
/**
|
||||
* Poll {@link WorkspaceControl#locatePane} for {@code paneId} until it reports {@code null}
|
||||
* (the pane is gone) or {@link #PANE_DEATH_TIMEOUT_SECONDS} elapses. Deliberately never calls
|
||||
* {@link AgentControl#status} and never reads the live-lead terminal map — both answer a
|
||||
* different question (whether an AGENT is live, not whether this PANE still exists) and
|
||||
* {@code locatePane} alone catches a {@link HerdrException} from the underlying {@code
|
||||
* pane.get} and turns it into {@code null} — see this class's javadoc.
|
||||
*/
|
||||
private DeathResult waitUntilPaneGone(String paneId) {
|
||||
long startMillis = nowMillis.getAsLong();
|
||||
long deadline = startMillis + TimeUnit.SECONDS.toMillis(PANE_DEATH_TIMEOUT_SECONDS);
|
||||
while (nowMillis.getAsLong() < deadline) {
|
||||
if (spaces.locatePane(paneId) == null) {
|
||||
return new DeathResult(true, nowMillis.getAsLong() - startMillis);
|
||||
}
|
||||
pollSleeper.run();
|
||||
}
|
||||
return new DeathResult(false, nowMillis.getAsLong() - startMillis);
|
||||
}
|
||||
|
||||
/** The measured outcome of {@link #waitUntilPaneGone}. */
|
||||
private record DeathResult(boolean gone, long elapsedMillis) {}
|
||||
|
||||
/**
|
||||
* Poll until {@code newTerminal}'s own pane reaches a real turn boundary ({@link
|
||||
* AgentStatus#IDLE} or {@link AgentStatus#DONE}, never merely {@link AgentStatus#BLOCKED}) —
|
||||
* the same exclusion {@link #waitUntilAtTurnBoundary} applies to the calling lead's own turn,
|
||||
* applied here to the fresh one, so {@code bootstrapText} is never typed into a pane that has
|
||||
* not actually finished booting — or {@code readySeconds} elapses. A failed status read
|
||||
* degrades to "not yet ready" and is retried on the next poll.
|
||||
*/
|
||||
private ReadinessResult waitUntilPaneReady(String newTerminal, int readySeconds) {
|
||||
long startMillis = nowMillis.getAsLong();
|
||||
long deadline = startMillis + TimeUnit.SECONDS.toMillis(readySeconds);
|
||||
while (nowMillis.getAsLong() < deadline) {
|
||||
AgentStatus status;
|
||||
try {
|
||||
status = agents.status(newTerminal);
|
||||
} catch (RuntimeException e) {
|
||||
log.debug("lead-rollover: status check failed while waiting for {} to be ready: {}",
|
||||
newTerminal, e.toString());
|
||||
status = null;
|
||||
}
|
||||
if (status == AgentStatus.IDLE || status == AgentStatus.DONE) {
|
||||
return new ReadinessResult(true, nowMillis.getAsLong() - startMillis);
|
||||
}
|
||||
pollSleeper.run();
|
||||
}
|
||||
return new ReadinessResult(false, nowMillis.getAsLong() - startMillis);
|
||||
}
|
||||
|
||||
/** The measured outcome of {@link #waitUntilPaneReady}. */
|
||||
private record ReadinessResult(boolean ready, long elapsedMillis) {}
|
||||
|
||||
/**
|
||||
* Poll until {@code newTerminal} is present in {@link #liveLeadTerminals} or {@code
|
||||
* readySeconds} elapses. This is bookkeeping, not a safety gate: the pane's own readiness (see
|
||||
* {@link #waitUntilPaneReady}) is what decides whether {@code bootstrapText} is safe to send —
|
||||
* a timeout here only means the daemon's own lead-discovery scan has not caught up yet.
|
||||
*/
|
||||
private IdentityResult waitUntilRecognisedAsLead(String newTerminal, int readySeconds) {
|
||||
long startMillis = nowMillis.getAsLong();
|
||||
long deadline = startMillis + TimeUnit.SECONDS.toMillis(readySeconds);
|
||||
while (nowMillis.getAsLong() < deadline) {
|
||||
if (liveLeadTerminals.get().containsKey(newTerminal)) {
|
||||
return new IdentityResult(true, nowMillis.getAsLong() - startMillis);
|
||||
}
|
||||
pollSleeper.run();
|
||||
}
|
||||
return new IdentityResult(false, nowMillis.getAsLong() - startMillis);
|
||||
}
|
||||
|
||||
/** The measured outcome of {@link #waitUntilRecognisedAsLead}. */
|
||||
private record IdentityResult(boolean ready, long elapsedMillis) {}
|
||||
|
||||
/** Drop a pending request without rolling. @return whether a pending request existed for {@code token} */
|
||||
public boolean cancel(String token) {
|
||||
return pending.remove(token) != null;
|
||||
@@ -729,15 +972,12 @@ public final class LeadRollover {
|
||||
|
||||
/**
|
||||
* Poll {@link AgentControl#status} until {@code target} reports a real turn boundary — {@link
|
||||
* AgentStatus#IDLE} or {@link AgentStatus#DONE} — bounded by {@code settleSeconds}. Used once by
|
||||
* {@link #runRollover}, to wait for the CALLING turn's own pane to settle before {@code /clear}
|
||||
* is ever sent at all — the {@code turnSettleSeconds} gate that makes this correction safe. The
|
||||
* SECOND wait, after {@code /clear}, is {@link #waitForClearPickupAndSettle} instead (fleetd
|
||||
* #489) — a plain boundary check is not enough there, because {@code /clear} starts no turn of
|
||||
* its own, so this method would (wrongly) report "settled" on its very first poll whether or not
|
||||
* {@code /clear} was actually picked up. A failed status read degrades to "not yet settled" and
|
||||
* is retried on the next poll, the same posture {@code LeadHeartbeatLoop} and {@code
|
||||
* HerdrPeerLauncher}'s readiness gate already take toward an unreadable status.
|
||||
* AgentStatus#IDLE} or {@link AgentStatus#DONE} — bounded by {@code settleSeconds}. Used by
|
||||
* {@link #runRollover} to wait for the CALLING turn's own pane to settle before the old pane is
|
||||
* touched at all — the {@code turnSettleSeconds} gate that makes tearing it down safe. A failed
|
||||
* status read degrades to "not yet settled" and is retried on the next poll, the same posture
|
||||
* {@code LeadHeartbeatLoop} and {@code HerdrPeerLauncher}'s readiness gate already take toward
|
||||
* an unreadable status.
|
||||
*
|
||||
* <p><strong>Deliberately not {@link AgentStatus#injectable()}.</strong> {@code injectable()}
|
||||
* answers the {@code Injector}'s question — "may I deliver a message without stepping on a live
|
||||
@@ -745,16 +985,16 @@ public final class LeadRollover {
|
||||
* an approval prompt is safe to queue a message behind. This class asks a stricter question —
|
||||
* "has the turn actually ended" — and {@code BLOCKED} answers no: it is a live turn that is
|
||||
* merely paused, not one that has finished. Reusing {@code injectable()} here would let this
|
||||
* wait fire {@code /clear} while the lead's own {@code confirm()}-calling turn is still live and
|
||||
* paused on a prompt — exactly the live-context-destroying failure the {@code turnSettleSeconds}
|
||||
* gate exists to prevent. Do not "simplify" this back to {@code injectable()}. ({@link
|
||||
* #waitForClearPickupAndSettle} keeps the same exclusion of {@code BLOCKED}, for the same
|
||||
* reason, on the second wait.)
|
||||
* wait tear the old pane down while the lead's own {@code confirm()}-calling turn is still live
|
||||
* and paused on a prompt — exactly the live-context-destroying failure {@code turnSettleSeconds}
|
||||
* exists to prevent. Do not "simplify" this back to {@code injectable()}. ({@link
|
||||
* #waitUntilPaneReady} applies the same exclusion of {@code BLOCKED} to the fresh lead's own
|
||||
* turn.)
|
||||
*
|
||||
* @return a {@link TurnSettleResult} whose {@code settled()} is {@code true} once a real
|
||||
* boundary was observed, {@code false} if {@code settleSeconds} elapses first.
|
||||
* {@code elapsedMillis()} is a MEASURED value from the injected {@link #nowMillis}
|
||||
* clock, never the configured {@code settleSeconds} budget (fleetd #494 follow-up).
|
||||
* clock, never the configured {@code settleSeconds} budget.
|
||||
*/
|
||||
private TurnSettleResult waitUntilAtTurnBoundary(String target, int settleSeconds) {
|
||||
long startMillis = nowMillis.getAsLong();
|
||||
@@ -771,142 +1011,11 @@ public final class LeadRollover {
|
||||
if (status == AgentStatus.IDLE || status == AgentStatus.DONE) {
|
||||
return new TurnSettleResult(true, nowMillis.getAsLong() - startMillis);
|
||||
}
|
||||
settleSleeper.run();
|
||||
pollSleeper.run();
|
||||
}
|
||||
return new TurnSettleResult(false, nowMillis.getAsLong() - startMillis);
|
||||
}
|
||||
|
||||
/**
|
||||
* The measured outcome of {@link #waitUntilAtTurnBoundary} — fleetd #494 follow-up. The sibling
|
||||
* of {@link ClearSettleResult} for the FIRST wait, which never nudges, so it carries no nudge
|
||||
* count.
|
||||
*/
|
||||
/** The measured outcome of {@link #waitUntilAtTurnBoundary}. */
|
||||
private record TurnSettleResult(boolean settled, long elapsedMillis) {}
|
||||
|
||||
/**
|
||||
* The SECOND wait in {@link #runRollover} — after {@code /clear} has been sent, waits for it to
|
||||
* settle, bounded by {@code settleSeconds}. <strong>fleetd #489 — the paste-race fix.</strong>
|
||||
* {@code /clear} does not start a real turn of its own, so a pane with no submit race simply
|
||||
* stays {@link AgentStatus#IDLE} the whole time: {@link #waitUntilAtTurnBoundary} would (wrongly)
|
||||
* call that "settled" on its very first poll, whether or not the {@code /clear} Enter actually
|
||||
* landed. That was Fault 1, measured live on 2026-09-12 — the second gate was a no-op, so a
|
||||
* {@code bootstrapText} send followed immediately, racing Fault 2: {@link AgentControl#submit}'s
|
||||
* own javadoc already records that the submit accompanying a delivery "can race the paste —
|
||||
* especially right as the worker's TUI becomes interactive — leaving the text unsubmitted"
|
||||
* (CB-113). Because {@code runRollover} deliberately bypasses {@code Injector} for {@code
|
||||
* /clear} (see this class's javadoc), it inherited none of {@code Injector}'s nudging — so the
|
||||
* lost {@code /clear} Enter sat in the input box and {@code bootstrapText} was typed right after
|
||||
* it, landing as one concatenated line.
|
||||
*
|
||||
* <p>This method copies the pickup-nudge pattern {@link dev.ltms.fleet.inject.Injector} already
|
||||
* ships for exactly this, on its own post-turn {@code /clear} housekeeping (fleetd #306; see
|
||||
* {@code Injector.java:288-340} and {@code Injector.java:437-442}):
|
||||
* <ul>
|
||||
* <li>an {@link AgentStatus#WORKING} sample means {@code /clear} was picked up as a real
|
||||
* turn;</li>
|
||||
* <li>until that happens, each poll that still reports {@link AgentStatus#IDLE} or {@link
|
||||
* AgentStatus#DONE} re-sends the submit keystroke ({@link AgentControl#submit}) to nudge
|
||||
* the raced Enter — for the first {@code PICKUP_GRACE_POLLS - 1} of {@link
|
||||
* #PICKUP_GRACE_POLLS} consecutive such polls (i.e. {@code PICKUP_GRACE_POLLS - 1}
|
||||
* nudges: 7, not 8, given {@code PICKUP_GRACE_POLLS = 8}). A second Enter on an empty
|
||||
* Claude Code prompt is a no-op, so repeating it is safe;</li>
|
||||
* <li>the {@code PICKUP_GRACE_POLLS}th consecutive such poll, with {@code WORKING} still never
|
||||
* observed, releases rather than wedges the roll instead of nudging again — the same
|
||||
* choice {@code Injector} makes — and returns {@code settled() == true} anyway, logged at
|
||||
* {@code warn} with the measured elapsed time (fleetd #494) so an operator can see which
|
||||
* path ran and how long it actually took;</li>
|
||||
* <li>once {@code WORKING} has been observed, nudging stops and this instead waits for a real
|
||||
* {@code working → IDLE/DONE} completion boundary before returning {@code true}.</li>
|
||||
* </ul>
|
||||
*
|
||||
* <p><strong>{@link AgentStatus#BLOCKED} is deliberately excluded from both the nudge and the
|
||||
* boundary check</strong> — the same reasoning as {@link #waitUntilAtTurnBoundary}'s own
|
||||
* javadoc: a paused live turn is not a settled one, and re-sending Enter into an open approval
|
||||
* prompt could wrongly answer it. A {@code BLOCKED} sample (or an unreadable/{@link
|
||||
* AgentStatus#UNKNOWN} one) simply keeps this polling, with no nudge and no release, until either
|
||||
* a real boundary is reached or {@code settleSeconds} runs out.
|
||||
*
|
||||
* <p>{@link AgentControl#submit} can itself throw; a {@link RuntimeException} from it is
|
||||
* swallowed and logged at {@code debug}, exactly like {@code Injector.java:437-442} — a failed
|
||||
* nudge must not abort the roll.
|
||||
*
|
||||
* @return a {@link ClearSettleResult} whose {@code settled()} is {@code true} once {@code
|
||||
* /clear} has settled, or once the nudge budget was exhausted with no pickup ever
|
||||
* observed (released rather than wedged); {@code false} if {@code settleSeconds} elapses
|
||||
* first — the caller must NOT send {@code bootstrapText} in that case, exactly as before
|
||||
* this fix. {@code elapsedMillis()} and {@code nudges()} are MEASURED values (from the
|
||||
* injected {@link #nowMillis} clock and an actual nudge count), never the configured
|
||||
* {@code settleSeconds} budget (fleetd #494).
|
||||
*/
|
||||
private ClearSettleResult waitForClearPickupAndSettle(String target, int settleSeconds) {
|
||||
long startMillis = nowMillis.getAsLong();
|
||||
long deadline = startMillis + TimeUnit.SECONDS.toMillis(settleSeconds);
|
||||
boolean pickedUp = false; // a WORKING sample has been observed since /clear was sent
|
||||
int idlePollsAwaitingPickup = 0;
|
||||
int nudges = 0;
|
||||
while (nowMillis.getAsLong() < deadline) {
|
||||
AgentStatus status;
|
||||
try {
|
||||
status = agents.status(target);
|
||||
} catch (RuntimeException e) {
|
||||
log.debug("lead-rollover: status check failed while waiting for {} to settle after "
|
||||
+ "/clear: {}", target, e.toString());
|
||||
status = null;
|
||||
}
|
||||
if (status == AgentStatus.WORKING) {
|
||||
pickedUp = true;
|
||||
} else if (status == AgentStatus.IDLE || status == AgentStatus.DONE) {
|
||||
if (pickedUp) {
|
||||
// a real WORKING -> IDLE/DONE completion boundary
|
||||
return new ClearSettleResult(true, nowMillis.getAsLong() - startMillis, nudges);
|
||||
}
|
||||
if (++idlePollsAwaitingPickup >= PICKUP_GRACE_POLLS) {
|
||||
long elapsedMillis = nowMillis.getAsLong() - startMillis;
|
||||
// fleetd #494: this release trades a possibly-unsubmitted /clear for progress
|
||||
// instead of wedging the roll — that trade is deliberate and stays. But it is
|
||||
// also exactly the case that reported false success in the real incident (the
|
||||
// whole roll "succeeded" after 438ms of a 20s budget), so raise it to WARN and
|
||||
// print the MEASURED elapsed time next to the target pane, not just the count.
|
||||
//
|
||||
// fleetd #494 follow-up (2nd pass): BOTH numbers in this line must come from
|
||||
// the loop's own counters, never from the PICKUP_GRACE_POLLS constant.
|
||||
// `idlePollsAwaitingPickup` and `nudges` each have exactly one write site in
|
||||
// this loop, on the same branch, so on this branch they cannot differ from
|
||||
// PICKUP_GRACE_POLLS / PICKUP_GRACE_POLLS - 1 today — no test can prove the
|
||||
// difference on this line, and printing the counters does not change that.
|
||||
// What it does buy: one source of truth instead of two, so a later change to
|
||||
// the loop (an early return, a second increment site, a different exit
|
||||
// condition) cannot leave this message reporting a number the loop no longer
|
||||
// produces. The place where `nudges` genuinely varies with the run — and is
|
||||
// covered by a test that can tell it apart from a constant — is the
|
||||
// /clear-timeout warn in runRollover, which prints clearResult.nudges().
|
||||
log.warn("lead-rollover: /clear on {} was never observed as WORKING after {} "
|
||||
+ "consecutive IDLE/DONE polls ({} of those were nudged) — "
|
||||
+ "releasing rather than wedging the roll (elapsed={}ms)",
|
||||
target, idlePollsAwaitingPickup, nudges, elapsedMillis);
|
||||
return new ClearSettleResult(true, elapsedMillis, nudges);
|
||||
}
|
||||
try {
|
||||
agents.submit(target); // nudge a raced Enter (CB-113) so /clear actually submits
|
||||
} catch (RuntimeException e) {
|
||||
log.debug("lead-rollover: resubmit to {} failed (will retry next poll): {}",
|
||||
target, e.getMessage());
|
||||
} finally {
|
||||
nudges++; // an attempted nudge, whether or not the submit call itself threw
|
||||
}
|
||||
}
|
||||
// AgentStatus.BLOCKED or UNKNOWN (or an unreadable status, above): neither a pickup
|
||||
// signal nor a boundary — keep polling without nudging or releasing.
|
||||
settleSleeper.run();
|
||||
}
|
||||
return new ClearSettleResult(false, nowMillis.getAsLong() - startMillis, nudges);
|
||||
}
|
||||
|
||||
/**
|
||||
* The measured outcome of {@link #waitForClearPickupAndSettle} — fleetd #494. Carries the
|
||||
* MEASURED elapsed time (from the injected {@link #nowMillis} clock) and nudge count alongside
|
||||
* the settle/timeout decision, so callers can log them instead of the configured budget, which
|
||||
* is not how long the wait actually ran.
|
||||
*/
|
||||
private record ClearSettleResult(boolean settled, long elapsedMillis, int nudges) {}
|
||||
}
|
||||
|
||||
@@ -891,8 +891,8 @@ public final class FleetMcp {
|
||||
* fleetd #612 B3 — as {@link #quarantineSource()}, {@code public} for the same cross-package
|
||||
* reason, for the real {@link LeadRollover} (or {@code null}) this daemon was assembled with.
|
||||
* {@code FleetdLeadRolloverAssemblyTest} drives {@code open}/{@code confirm} on this exact
|
||||
* instance and waits for the real continuation to send {@code /clear} and {@code bootstrapText}
|
||||
* through the real {@code router.leadAgents()}.
|
||||
* instance and waits for the real continuation to end the old pane, relaunch a fresh one, and
|
||||
* send {@code bootstrapText} through the real {@code router.leadAgents()}.
|
||||
*/
|
||||
public LeadRollover leadRollover() {
|
||||
return leadRollover;
|
||||
@@ -1492,7 +1492,7 @@ public final class FleetMcp {
|
||||
}
|
||||
if (isBlank(callerTerminal)) {
|
||||
// An unnamed primary (token/loopback path, no resolved pane) has nowhere for the
|
||||
// eventual /clear + bootstrap to land — LeadRollover#open would throw
|
||||
// eventual relaunch + bootstrap to land — LeadRollover#open would throw
|
||||
// IllegalArgumentException for the same reason; refuse cleanly here instead.
|
||||
return error("fleet_handover requires a named lead pane (a resolved connection terminal) "
|
||||
+ "to open a rollover request against — an unnamed primary has none");
|
||||
@@ -2630,17 +2630,20 @@ public final class FleetMcp {
|
||||
private static McpSchema.Tool handoverTool() {
|
||||
return tool(FleetTool.HANDOVER.wireName(),
|
||||
"Replace your OWN lead session once its context is full: write a handover file, "
|
||||
+ "then use this to have fleetd clear your pane and bootstrap a fresh lead "
|
||||
+ "session against it. Four actions: 'open' (requests a token and the "
|
||||
+ "handoverPath you must write the handover file to before confirming), "
|
||||
+ "'confirm' (validates every gate and — only if every one passes — schedules "
|
||||
+ "the roll; it does NOT itself clear the pane, the roll runs once this call's "
|
||||
+ "own turn ends), 'cancel' (drops a pending request without rolling), and "
|
||||
+ "'status' (read-only: what happened to a token after 'confirm' — still "
|
||||
+ "running (approved but not finished yet), the roll completed, the calling "
|
||||
+ "turn never settled within turnSettleSeconds so no /clear was ever sent, or "
|
||||
+ "/clear itself never settled so bootstrapText was never sent; never "
|
||||
+ "schedules, cancels or retries anything). Primary-only. "
|
||||
+ "then use this to have fleetd end your pane's process and relaunch a fresh "
|
||||
+ "lead session bootstrapped against it. Four actions: 'open' (requests a "
|
||||
+ "token and the handoverPath you must write the handover file to before "
|
||||
+ "confirming), 'confirm' (validates every gate and — only if every one "
|
||||
+ "passes — schedules the roll; it does NOT itself end your pane, the roll "
|
||||
+ "runs once this call's own turn ends), 'cancel' (drops a pending request "
|
||||
+ "without rolling), and 'status' (read-only: what happened to a token after "
|
||||
+ "'confirm' — still running (approved but not finished yet), the roll "
|
||||
+ "completed, the calling turn never settled within turnSettleSeconds so "
|
||||
+ "nothing was touched, the old pane never confirmed dead so no relaunch was "
|
||||
+ "attempted, the relaunch itself failed, the fresh pane never became ready "
|
||||
+ "so bootstrapText was never sent, or the fresh pane became ready and was "
|
||||
+ "bootstrapped but was never recognised as a live lead; never schedules, "
|
||||
+ "cancels or retries anything). Primary-only. "
|
||||
+ "There is deliberately no terminal/session/leadTerminal parameter: the pane "
|
||||
+ "to roll is always resolved from YOUR OWN connection, never a value you "
|
||||
+ "pass, so you can only ever roll yourself — never another lead. Requires "
|
||||
|
||||
@@ -6,6 +6,8 @@ import dev.ltms.fleet.guard.SubscriptionGuard;
|
||||
import dev.ltms.fleet.herdr.AgentControl;
|
||||
import dev.ltms.fleet.herdr.FakeHerdr;
|
||||
import dev.ltms.fleet.herdr.HerdrClient;
|
||||
import dev.ltms.fleet.herdr.WorkspaceControl;
|
||||
import dev.ltms.fleet.lead.LeadLauncher;
|
||||
import dev.ltms.fleet.lead.LeadRollover;
|
||||
import dev.ltms.fleet.msg.ReplyInbox;
|
||||
import io.javalin.Javalin;
|
||||
@@ -34,7 +36,7 @@ import static org.junit.jupiter.api.Assertions.fail;
|
||||
* fleetd #612 Unit A) with three methods: {@code unrelatedAnchorStillPresent} (a scaffold anchor,
|
||||
* not an independent claim — needs no replacement of its own), {@code
|
||||
* mainStillCallsTheLeadRolloverFactory} (the call-site pin replaced by {@link
|
||||
* #assembledLeadRolloverRunsTheRealClearAndBootstrapSequence}), and {@code
|
||||
* #assembledLeadRolloverEndsTheOldPaneThroughTheRealHerdrRouter}), and {@code
|
||||
* factoryGatesOnConfigPresence} (the absent-config claim replaced by {@link
|
||||
* #absentLeadRolloverConfigMeansNoRolloverIsBuilt} — a claim this ticket found was NOT actually
|
||||
* covered behaviourally anywhere else: {@code LeadRolloverTest}'s only related assertion is
|
||||
@@ -50,8 +52,8 @@ import static org.junit.jupiter.api.Assertions.fail;
|
||||
* invisible to this test, even though the two are genuinely different daemons in production. This
|
||||
* version configures two distinct sockets and two distinct {@link FakeHerdr} instances (the same
|
||||
* pattern {@code FleetdAssemblyConnectionIdentityTest}, fleetd #612 B2, already uses to separate
|
||||
* lead from member) and asserts the roll's {@code /clear}/bootstrap sends land on the LEAD fake
|
||||
* and never on the MEMBER one.
|
||||
* lead from member) and asserts the roll's {@code pane.close} call lands on the LEAD fake and
|
||||
* never on the MEMBER one.
|
||||
*/
|
||||
class FleetdLeadRolloverAssemblyTest {
|
||||
|
||||
@@ -177,11 +179,11 @@ class FleetdLeadRolloverAssemblyTest {
|
||||
return FleetConfig.load(f);
|
||||
}
|
||||
|
||||
@SuppressWarnings("unchecked")
|
||||
@Test
|
||||
@DisplayName("[BEHAVIOURAL] the real assembled LeadRollover runs the full open/confirm/continuation "
|
||||
+ "sequence — /clear, then bootstrapText — through the real herdr router")
|
||||
void assembledLeadRolloverRunsTheRealClearAndBootstrapSequence(@TempDir Path dir) throws Exception {
|
||||
@DisplayName("[BEHAVIOURAL] the real assembled LeadRollover runs the open/confirm/continuation "
|
||||
+ "sequence through the real herdr router — ending the old pane, then giving up once it "
|
||||
+ "never reports gone")
|
||||
void assembledLeadRolloverEndsTheOldPaneThroughTheRealHerdrRouter(@TempDir Path dir) throws Exception {
|
||||
Path leadCwd = dir.resolve("lead-workspace");
|
||||
Files.createDirectories(leadCwd);
|
||||
FleetConfig cfg = writeConfig(dir, leadCwd);
|
||||
@@ -223,40 +225,42 @@ class FleetdLeadRolloverAssemblyTest {
|
||||
// constructor), so this polls the real FleetMcp.leadRollover() instance's status(token)
|
||||
// until the real continuation finishes.
|
||||
LeadRollover.RollStatus status = pollUntilTerminal(rollover, pending.token());
|
||||
assertEquals(LeadRollover.RollState.ROLLED, status.state(),
|
||||
"the full happy path must complete: FakeHerdr's default agent status is 'idle', so "
|
||||
+ "the turn-boundary wait settles immediately and the post-/clear wait "
|
||||
+ "releases via its pickup-grace path — detail: " + status.detail());
|
||||
|
||||
// Prove the real herdr router actually sent BOTH messages, in order, to the real LEAD
|
||||
// pane — this is the one thing a source-text pin on the call site could never show.
|
||||
// FakeHerdr's pane.get is a fixed canned response that never reports a pane as gone, so the
|
||||
// real router's death poll runs out its whole budget and the roll stops here — proving the
|
||||
// real teardown call landed on the real LEAD pane without ever reaching a relaunch or a send.
|
||||
assertEquals(LeadRollover.RollState.OLD_PANE_NEVER_DIED, status.state(),
|
||||
"the old pane never reports gone against this fake, so the roll must stop with "
|
||||
+ "OLD_PANE_NEVER_DIED rather than ever relaunching or sending anything — "
|
||||
+ "detail: " + status.detail());
|
||||
|
||||
// Prove the real herdr router actually closed the real LEAD pane — this is the one thing a
|
||||
// source-text pin on the call site could never show.
|
||||
boolean closedOldPane = lead.calls.stream()
|
||||
.anyMatch(c -> c.method().equals("pane.close")
|
||||
&& c.params() instanceof Map<?, ?> m && "w2:p7".equals(m.get("pane_id")));
|
||||
assertTrue(closedOldPane, "endOldSession must close the real old pane (w2:p7) through the "
|
||||
+ "real LEAD herdr client, got calls: " + lead.calls);
|
||||
|
||||
// No agent.prompt is ever sent on this path: the roll stops at the pane-death wait, strictly
|
||||
// before the relaunch and the final send step.
|
||||
List<FakeHerdr.Call> prompts = lead.calls.stream()
|
||||
.filter(c -> c.method().equals("agent.prompt"))
|
||||
.toList();
|
||||
assertTrue(prompts.size() >= 2, "expected at least a /clear send and a bootstrapText send "
|
||||
+ "on the LEAD daemon, got " + prompts.size() + " agent.prompt calls: " + prompts);
|
||||
assertEquals("/clear", ((Map<String, Object>) prompts.get(0).params()).get("text"),
|
||||
"the first send must be the literal /clear housekeeping command");
|
||||
Object secondText = ((Map<String, Object>) prompts.get(1).params()).get("text");
|
||||
assertTrue(secondText instanceof String && ((String) secondText).contains(expectedHandoverPath),
|
||||
"the second send must be the default bootstrapText naming the resolved handover "
|
||||
+ "path, got: " + secondText);
|
||||
assertTrue(prompts.isEmpty(), "a roll that stops at OLD_PANE_NEVER_DIED must never reach the "
|
||||
+ "send step, got agent.prompt call(s) on the LEAD daemon: " + prompts);
|
||||
|
||||
// fleetd #612 B3 correction: prove the roll never touches the MEMBER daemon. A mutation
|
||||
// swapping router.leadAgents() for router.memberAgents() at the real call site would move
|
||||
// both sends above onto `member` instead, which this assertion catches — the thing the
|
||||
// single-fake version of this test could never see, because both wrapped the same client.
|
||||
List<FakeHerdr.Call> memberPrompts = member.calls.stream()
|
||||
.filter(c -> c.method().equals("agent.prompt"))
|
||||
.toList();
|
||||
assertTrue(memberPrompts.isEmpty(), "the roll must be wired to the LEAD daemon only — got "
|
||||
+ memberPrompts.size() + " agent.prompt call(s) on the MEMBER daemon instead: "
|
||||
+ memberPrompts);
|
||||
// the pane.close call above onto `member` instead, which this assertion catches — the thing
|
||||
// the single-fake version of this test could never see, because both wrapped the same client.
|
||||
assertTrue(member.calls.isEmpty(), "the roll must be wired to the LEAD daemon only — got "
|
||||
+ member.calls.size() + " call(s) on the MEMBER daemon instead: " + member.calls);
|
||||
}
|
||||
|
||||
private static LeadRollover.RollStatus pollUntilTerminal(LeadRollover rollover, String token)
|
||||
throws InterruptedException {
|
||||
long deadline = System.nanoTime() + java.util.concurrent.TimeUnit.SECONDS.toNanos(10);
|
||||
long deadline = System.nanoTime() + java.util.concurrent.TimeUnit.SECONDS.toNanos(15);
|
||||
while (System.nanoTime() < deadline) {
|
||||
LeadRollover.RollStatus status = rollover.status(token);
|
||||
if (status.state() != LeadRollover.RollState.PENDING
|
||||
@@ -283,8 +287,10 @@ class FleetdLeadRolloverAssemblyTest {
|
||||
""");
|
||||
ConfigRef config = new ConfigRef(yaml, FleetConfig.load(yaml));
|
||||
AgentControl agents = new AgentControl(new FakeHerdr());
|
||||
WorkspaceControl spaces = new WorkspaceControl(new FakeHerdr());
|
||||
LeadLauncher launcher = new LeadLauncher(agents, spaces, config.get());
|
||||
|
||||
LeadRollover rollover = Fleetd.leadRollover(config.get(), agents, config, Map::of);
|
||||
LeadRollover rollover = Fleetd.leadRollover(config.get(), agents, spaces, launcher, config, Map::of);
|
||||
|
||||
assertNull(rollover, "leadRollover: is absent from this config, so the factory's opt-in "
|
||||
+ "gate (`if (cfg.leadRollover() == null) return null;`) must fire and no "
|
||||
|
||||
@@ -4,6 +4,8 @@ import dev.ltms.fleet.config.ConfigRef;
|
||||
import dev.ltms.fleet.config.FleetConfig;
|
||||
import dev.ltms.fleet.herdr.AgentControl;
|
||||
import dev.ltms.fleet.herdr.FakeHerdr;
|
||||
import dev.ltms.fleet.herdr.WorkspaceControl;
|
||||
import dev.ltms.fleet.lead.LeadLauncher;
|
||||
import dev.ltms.fleet.lead.LeadRollover;
|
||||
import org.junit.jupiter.api.DisplayName;
|
||||
import org.junit.jupiter.api.Test;
|
||||
@@ -54,6 +56,11 @@ class FleetdLeadRolloverWorkspaceLookupTest {
|
||||
return new AgentControl(new FakeHerdr());
|
||||
}
|
||||
|
||||
/** None of this class's tests reach the deferred continuation, so a plain fake is enough. */
|
||||
private static LeadLauncher fakeLauncher(FleetConfig cfg) {
|
||||
return new LeadLauncher(fakeAgents(), new WorkspaceControl(new FakeHerdr()), cfg);
|
||||
}
|
||||
|
||||
@Test
|
||||
@DisplayName("[BEHAVIOURAL] Fleetd.leadRollover(...) resolves a relative handoverPath against "
|
||||
+ "the CALLING lead's configured cwd, not the daemon's own working directory")
|
||||
@@ -74,7 +81,8 @@ class FleetdLeadRolloverWorkspaceLookupTest {
|
||||
""".formatted(leadCwd.toString()));
|
||||
ConfigRef config = new ConfigRef(yaml, FleetConfig.load(yaml));
|
||||
|
||||
LeadRollover rollover = Fleetd.leadRollover(config.get(), fakeAgents(), config,
|
||||
LeadRollover rollover = Fleetd.leadRollover(config.get(), fakeAgents(),
|
||||
new WorkspaceControl(new FakeHerdr()), fakeLauncher(config.get()), config,
|
||||
() -> Map.of("term_opus", "opus"));
|
||||
assertNotNull(rollover, "leadRollover: is present in the loaded config, so the factory "
|
||||
+ "must construct an object");
|
||||
@@ -105,7 +113,8 @@ class FleetdLeadRolloverWorkspaceLookupTest {
|
||||
|
||||
// No lead has been discovered yet — exactly the real shape of a lead the live tab scan
|
||||
// has not yet scanned, or one with no fleet.leaders entry at all.
|
||||
LeadRollover rollover = Fleetd.leadRollover(config.get(), fakeAgents(), config, Map::of);
|
||||
LeadRollover rollover = Fleetd.leadRollover(config.get(), fakeAgents(),
|
||||
new WorkspaceControl(new FakeHerdr()), fakeLauncher(config.get()), config, Map::of);
|
||||
assertNotNull(rollover);
|
||||
|
||||
LeadRollover.PendingRollover pending = rollover.open("term_unknown", "test");
|
||||
@@ -144,7 +153,8 @@ class FleetdLeadRolloverWorkspaceLookupTest {
|
||||
// below — exactly the natural mistake to make, since leads are discovered by a live tab
|
||||
// scan that runs AFTER this factory is constructed at startup.
|
||||
Map<String, String> liveLeadTerminals = new HashMap<>();
|
||||
LeadRollover rollover = Fleetd.leadRollover(config.get(), fakeAgents(), config,
|
||||
LeadRollover rollover = Fleetd.leadRollover(config.get(), fakeAgents(),
|
||||
new WorkspaceControl(new FakeHerdr()), fakeLauncher(config.get()), config,
|
||||
() -> liveLeadTerminals);
|
||||
assertNotNull(rollover);
|
||||
|
||||
|
||||
+1
-1
@@ -103,7 +103,7 @@ class FleetConfigWithDefaultsPreservesEveryComponentTest {
|
||||
// comment there), same as broker/primary/leadHeartbeat/... above — a real, non-null value
|
||||
// here proves it, rather than leaving it null and proving nothing.
|
||||
v.put("leadRollover", new FleetConfig.LeadRollover(
|
||||
"/handover/guard.md", true, 3600, 20, 20, "read the handover file"));
|
||||
"/handover/guard.md", true, 3600, 20, 45, "read the handover file"));
|
||||
assertNamesMatchComponents(v);
|
||||
return v;
|
||||
}
|
||||
|
||||
@@ -7,6 +7,7 @@ import java.util.ArrayList;
|
||||
import java.util.LinkedHashMap;
|
||||
import java.util.List;
|
||||
import java.util.Map;
|
||||
import java.util.Set;
|
||||
import java.util.concurrent.ConcurrentHashMap;
|
||||
import java.util.concurrent.CopyOnWriteArrayList;
|
||||
|
||||
@@ -59,6 +60,10 @@ public final class FakeHerdr implements HerdrClient {
|
||||
private Runnable onAgentStart; // fires the instant agent.start is called — see onAgentStart(Runnable)
|
||||
private volatile int agentGetOkCalls = Integer.MAX_VALUE; // how many agent.get calls succeed first
|
||||
private volatile String agentGetFailCode = null; // error code every agent.get call after that reports
|
||||
/** pane ids that {@link #paneGoneAfterClose} has opted into reporting gone — see that method. */
|
||||
private final Set<String> paneGoneAfterCloseIds = ConcurrentHashMap.newKeySet();
|
||||
/** pane ids a {@code pane.close} call has actually reached, for {@link #paneGoneAfterCloseIds}. */
|
||||
private final Set<String> closedPaneIds = ConcurrentHashMap.newKeySet();
|
||||
|
||||
public FakeHerdr healthy(boolean h) {
|
||||
this.healthy = h;
|
||||
@@ -220,6 +225,19 @@ public final class FakeHerdr implements HerdrClient {
|
||||
return this;
|
||||
}
|
||||
|
||||
/**
|
||||
* Make {@code pane.get(paneId)} report the pane gone (a {@code pane_not_found} {@link
|
||||
* HerdrException}, exactly as {@link WorkspaceControl#locatePane} expects to see once a pane
|
||||
* has really disappeared) once a {@code pane.close} call for that same {@code paneId} has
|
||||
* actually reached this fake. Every other pane, and this pane before its own close, keeps
|
||||
* reporting the default canned {@code pane.get} response — opt-in, by pane id, so no existing
|
||||
* test's {@code pane.get} behaviour changes.
|
||||
*/
|
||||
public FakeHerdr paneGoneAfterClose(String paneId) {
|
||||
paneGoneAfterCloseIds.add(paneId);
|
||||
return this;
|
||||
}
|
||||
|
||||
/**
|
||||
* Run {@code hook} synchronously the instant an {@code agent.start} call reaches this fake —
|
||||
* i.e. the instant the peer PROCESS would start against a real herdr daemon. A test uses this
|
||||
@@ -422,9 +440,18 @@ public final class FakeHerdr implements HerdrClient {
|
||||
}
|
||||
yield mapper.readTree("{\"type\":\"ok\"}");
|
||||
}
|
||||
case "pane.get" -> mapper.readTree("""
|
||||
case "pane.get" -> {
|
||||
Object paneIdParam = params instanceof Map<?, ?> m ? m.get("pane_id") : null;
|
||||
String paneIdKey = paneIdParam == null ? null : String.valueOf(paneIdParam);
|
||||
if (paneIdKey != null && paneGoneAfterCloseIds.contains(paneIdKey)
|
||||
&& closedPaneIds.contains(paneIdKey)) {
|
||||
throw new HerdrException("herdr error [pane_not_found]: pane.get failed",
|
||||
"pane_not_found", null);
|
||||
}
|
||||
yield mapper.readTree("""
|
||||
{"type":"pane_info","pane":{"pane_id":"w9:pW","workspace_id":"w9",
|
||||
"tab_id":"w9:t2","agent_status":"idle"}}""");
|
||||
}
|
||||
case "pane.list" -> noPanes
|
||||
? mapper.readTree("{\"type\":\"pane_list\",\"panes\":[]}")
|
||||
: mapper.readTree("""
|
||||
@@ -458,6 +485,9 @@ public final class FakeHerdr implements HerdrClient {
|
||||
throw new HerdrException("herdr error [" + code + "]: pane.close failed",
|
||||
code, null);
|
||||
}
|
||||
if (paneIdParam != null) {
|
||||
closedPaneIds.add(String.valueOf(paneIdParam));
|
||||
}
|
||||
yield mapper.readTree("{\"type\":\"ok\"}");
|
||||
}
|
||||
default -> throw new HerdrException("fake has no canned response for " + method);
|
||||
|
||||
File diff suppressed because it is too large
Load Diff
@@ -11,6 +11,7 @@ import dev.ltms.fleet.herdr.FakeHerdr;
|
||||
import dev.ltms.fleet.herdr.PaneLocator;
|
||||
import dev.ltms.fleet.herdr.WorkspaceControl;
|
||||
import dev.ltms.fleet.inject.Injector;
|
||||
import dev.ltms.fleet.lead.LeadLauncher;
|
||||
import dev.ltms.fleet.lead.LeadRollover;
|
||||
import dev.ltms.fleet.member.ClaudeCodeLauncher;
|
||||
import dev.ltms.fleet.msg.InMemoryReplyInbox;
|
||||
@@ -24,6 +25,8 @@ import org.junit.jupiter.api.DisplayName;
|
||||
import org.junit.jupiter.api.Test;
|
||||
import org.junit.jupiter.api.io.TempDir;
|
||||
|
||||
import java.io.IOException;
|
||||
import java.io.UncheckedIOException;
|
||||
import java.nio.file.Files;
|
||||
import java.nio.file.Path;
|
||||
import java.util.List;
|
||||
@@ -36,14 +39,14 @@ import static org.junit.jupiter.api.Assertions.*;
|
||||
* fleetd #480 Unit C — the {@code fleet_handover} MCP tool, the surface that finally calls
|
||||
* {@link LeadRollover#open}/{@link LeadRollover#confirm}/{@link LeadRollover#cancel}.
|
||||
*
|
||||
* <p>Uses {@link LeadRollover}'s PUBLIC constructor (real wall clock, real 250ms settle poll, a
|
||||
* real virtual-thread continuation runner) rather than its package-private test constructor —
|
||||
* this test lives in {@code dev.ltms.fleet.mcp}, not {@code dev.ltms.fleet.lead}, and does not
|
||||
* need to control the post-{@code confirm()} continuation's timing: it only asserts the
|
||||
* SYNCHRONOUS return value of {@code open}/{@code confirm}/{@code cancel}, which is exactly what
|
||||
* {@code FleetMcp.handover} forwards to the client. {@code turnSettleSeconds}/{@code
|
||||
* clearSettleSeconds} are kept at 1s so a confirmed request's background continuation (which this
|
||||
* class does not wait on or assert against) gives up quickly rather than polling for 20s on a
|
||||
* <p>Uses {@link LeadRollover}'s PUBLIC constructor (real wall clock, real 250ms poll, a real
|
||||
* virtual-thread continuation runner) rather than its package-private test constructor — this
|
||||
* test lives in {@code dev.ltms.fleet.mcp}, not {@code dev.ltms.fleet.lead}, and does not need to
|
||||
* control the post-{@code confirm()} continuation's timing: it only asserts the SYNCHRONOUS return
|
||||
* value of {@code open}/{@code confirm}/{@code cancel}, which is exactly what {@code
|
||||
* FleetMcp.handover} forwards to the client. {@code turnSettleSeconds}/{@code
|
||||
* relaunchReadySeconds} are kept at 1s so a confirmed request's background continuation (which
|
||||
* this class does not wait on or assert against) gives up quickly rather than polling for 20s on a
|
||||
* daemon virtual thread.
|
||||
*/
|
||||
class FleetMcpHandoverTest {
|
||||
@@ -67,10 +70,26 @@ class FleetMcpHandoverTest {
|
||||
return new FleetConfig.LeadRollover(handoverPath, false, 3600, 1, 1, "read the handover file");
|
||||
}
|
||||
|
||||
private static FleetConfig minimalFleetConfig() {
|
||||
try {
|
||||
Path yaml = Files.createTempFile("fleet-mcp-handover-test", ".yaml");
|
||||
Files.writeString(yaml, "bind:\n port: 8080\n");
|
||||
return FleetConfig.load(yaml);
|
||||
} catch (IOException e) {
|
||||
throw new UncheckedIOException(e);
|
||||
}
|
||||
}
|
||||
|
||||
private LeadRollover newRollover(String handoverPath) {
|
||||
// Every handoverPath this test class uses comes from tmp.resolve(...), which is already
|
||||
// absolute, so the workspace lookup is never actually consulted — a no-op lookup is enough.
|
||||
return new LeadRollover(agents, () -> cfg(handoverPath), _ -> null);
|
||||
// None of this class's tests reach the recognition-wait or the relaunch call, so the
|
||||
// launcher's own functional correctness is irrelevant here — any constructed instance,
|
||||
// backed by the same fake herdr, is enough.
|
||||
WorkspaceControl spaces = new WorkspaceControl(herdr);
|
||||
LeadLauncher launcher = new LeadLauncher(agents, spaces, minimalFleetConfig());
|
||||
return new LeadRollover(agents, spaces, launcher, () -> cfg(handoverPath),
|
||||
_ -> null, _ -> null, Map::of);
|
||||
}
|
||||
|
||||
/** A fully wired FleetMcp on fakes (mirrors FleetMcpAuthzTest's helper), plus a leadRollover. */
|
||||
|
||||
Reference in New Issue
Block a user