fleetd #480 Unit A: lead rollover core (config block + executor) #483
Reference in New Issue
Block a user
Delete Branch "worker/480-a-rollover-core-6fc2ad-4"
Deleting a branch is permanent. Although the deleted branch may continue to exist for a short time before it actually gets removed, it CANNOT be undone in most cases. Continue?
fleetd #480 Unit A — lead rollover core
What this builds
A lead session fills up its context and must be replaced. This unit is the config block and the
executor only — not the MCP tool (a later unit will call
LeadRolloverfrom an MCP tool).leadRollover:config block (opt-in top-level key inFleetConfig):handoverPath(required when the block is present),
requireOperatorConfirm(defaulttrue),maxDocAgeSeconds(default3600),clearSettleSeconds(default20),bootstrapText(default names
handoverPath). Registered inConfigRef's hot/deferred/split/coldclassification, documented in
fleetd.example.yaml, validated by a newvalidateLeadRollover()(auto-wired into
FleetConfig.validateAll()'s reflective sweep).dev.ltms.fleet.lead.LeadRollover— the executor:open(String reason)→ records a token, the resolved handover path, and the wall-clocktimestamp of the request.
confirm(String token, boolean operatorConfirmed)→ verifies (operator confirmation, thenthe handover file: exists, non-empty, modified after
open()and not older thanmaxDocAgeSeconds), then rolls:agents.send(lead, "/clear")→ poll until the pane reportsinjectable again (bounded by
clearSettleSeconds) →agents.send(lead, bootstrapText). Ifthe pane never settles,
bootstrapTextis never sent and the refusal says so.cancel(String token)drops a pending request.confirm()call can ever roll a pane — no timer, no heartbeat, nobackground thread anywhere in this class.
/clearandbootstrapTextgo throughagents.send(...)directly, never throughInjector— a live probe (recorded in the ticket) proved a/clearrouted throughInjectorwedges that pane forever, since
/clearproduces no turn boundary.LongSupplier, neverSystem.nanoTime()(which freezes across a Mac sleep — fleetd #386).
Wired into
Fleetd.javaexactly likeLeadHeartbeatLoop: constructed only whencfg.leadRollover() != nullat startup. Nothing calls it yet — expected, and out of thisunit's scope.
Config classification: HOT, and why
Classified
leadRollover:as HOT inConfigRef(joinsplacement,memberCredentials,memberLoginShell,modelsin the hot-excluded set — 5 total now, was 4). Reasoning: unlikeLeadHeartbeatLoop, which bakes its config numbers intofinalfields at construction time (whatmakes
leadHeartbeat:DEFERRED),LeadRolloverholds aSupplier<FleetConfig.LeadRollover>(
() -> config.get().leadRollover(), the same shapefleet/placement/modelsalready use) andreads every field fresh on each
open()/confirm()call.One caveat, documented in both
ConfigRef.javaandFleetConfig.java: theLeadRolloverobject's construction is still gated on
cfg.leadRollover() != nullread from the startupconfig snapshot in
Fleetd.java(per the ticket's instruction to mirrorLeadHeartbeatLoop'sconstruction gating), so adding the block where it was absent at boot still needs a restart before
anything exists to call. This is a structural/existence fact, not a stale-value fact — analogous to
how adding a brand-new
profiles:entry needs a restart even though an existing profile's fieldsare genuinely hot.
Tests (all new/updated tests pinned against the unmutated file first)
LeadRolloverTest(14 cases) covers all 6 hard requirements from the ticket:leadRollover:block ⇒ no object constructed (proven viaFleetd.leadRollover(...)'sgate, cross-checked by the wiring test below).
nothingButAnExplicitConfirmCanEverRollAPane—open()alone never callsagents.send.missingHandoverFileRefuses(HANDOVER_MISSING),emptyHandoverFileRefuses(HANDOVER_EMPTY),staleHandoverFileRefuses(HANDOVER_STALE).operatorConfirmRequiredAndNotGivenRefuses(OPERATOR_NOT_CONFIRMED).freshnessCheckUsesTheInjectedWallClockNotNanoTime— provesopen()/confirm()use theinjected
LongSupplier, notSystem.nanoTime().cancel(),open()throwing when unconfigured, aCLEAR_DID_NOT_SETTLErefusal path,and a full successful roll (
/clear→ settle →bootstrapText, in order, token consumed).FleetdLeadRolloverWiringTest— source-text pin onFleetd.main's construction call(
LeadRollover leadRollover = leadRollover(cfg, primaryRegistry, router.leadAgents(), config);),guarded by an unrelated anchor assertion (
public final class Fleetd), mirroringFleetdCompletionResolverWiringTest's pattern — the same shape that already pins 5 log-onlyreporters in
Fleetd.mainagainst silent wiring drift.Updated 4 existing tests to account for the new 26th
FleetConfigrecord component:ConfigRefTopLevelCoverageTest,ConfigRefTopLevelReportingCoverageTest,FleetConfigValidateAllTest,FleetConfigWithDefaultsPreservesEveryComponentTest.Build result
mvn clean installfrom thefleetdmodule, full unpiped output, grepped explicitly for bothBUILD SUCCESS/BUILD FAILURE:(First full-build attempt caught 2 pre-existing coverage tests that needed updating for the new
record component — both fixed and re-verified green before this PR.)
Not done / out of scope
open/confirm/cancel— explicitly a later unit.wiki/is uninitialized in this worker worktree(
git submodule statusshows a leading-), unsatisfiable for a worker; leaving to the lead.Correction round (fleetd #480)
Two defects found after the initial review, both fixed on this branch (same PR, new commit
a942622):1.
confirm()deadlocked and destroyed context. It sent/clearbefore polling for the paneto become injectable — but
confirm()is called FROM the calling lead's own turn, so that pane isstill
WORKINGand can never report injectable inside that same call. The poll always timed out,but only after
/clearhad already fired and queued in the pane, destroying the lead's contextwith no fresh session ever started.
Fix:
confirm()now only validates every gate, then hands a one-shot continuation (a real virtualthread in production,
Runnable::runin tests) to do the actual work and returnsRollDecision.approved()immediately — "scheduled", not "rolled". The continuation waits for theSAME pane to report injectable again (new
turnSettleSecondsconfig key, default 20) beforesending
/clearat all; if that wait times out,/clearis never sent. The "only confirm() canroll, no timer/scheduler" invariant still holds — it was always about initiative, not synchronicity.
2.
confirm()usedPrimaryRegistry.primaryTerminal()— a single-slot lookup wrong for afeature with a caller. On a daemon with more than one labelled lead tab, lead X's
confirm()couldclear lead Y's pane. Fix:
open()/confirm()now take the caller's terminal id as a parameter(to be resolved by the MCP layer from the connection in the later tool-wiring unit), and
confirm()refuses with a new
NOT_YOUR_ROLLOVERreason unless it matches the terminalopen()recorded.Build:
mvn clean install—BUILD SUCCESS,Tests run: 1651, Failures: 0, Errors: 0, Skipped: 0.New tests:
turnThatNeverSettlesSendsNoClearAtAll(the branch that matters most) andaDifferentLeadTerminalCannotConfirmAnotherLeadsRollover, plus a settle-after-clear timeout testand an
open()input-validation test.LeadRolloverTest: 11 → 14 tests.