Compare commits
9 Commits
| Author | SHA1 | Date | |
|---|---|---|---|
| dcd505286f | |||
| 60fa958da2 | |||
| 9950361bc9 | |||
| 9f74b6619a | |||
| f687046450 | |||
| c6058652be | |||
| 2a95da2ff4 | |||
| 7f9137fcb6 | |||
| 3bf3968bc7 |
@@ -135,6 +135,20 @@ refused.**
|
||||
|
||||
1. **`fleet_handover{action: "open", reason: "<why now>"}`.** It returns a `token` and the
|
||||
`handoverPath` you must write to. Nothing has happened to your pane yet.
|
||||
|
||||
**Write to exactly that path, and do not resolve it yourself.** It is always absolute, even when
|
||||
the operator configured a relative `handoverPath`: fleetd resolves a relative one against your
|
||||
own workspace before it hands it to you. The daemon and your pane can run in different
|
||||
directories, so a path you resolve yourself can point at a different file from the one the daemon
|
||||
will check.
|
||||
|
||||
**Check that the path is ignored by git before you write to it (#491).** A relative
|
||||
`handoverPath` resolves inside YOUR workspace, which is usually a repository — and usually not
|
||||
the `fleetd` one, so an ignore rule added to `fleetd` does not protect it. Run
|
||||
`grep -n handover <your workspace>/.gitignore`. No output means the file you are about to write
|
||||
will show up as untracked content in that repo. The file is a snapshot of live state and must
|
||||
never be committed, so tell the operator rather than committing it or silently editing their
|
||||
`.gitignore`.
|
||||
2. **Write the handover file at that path**, following sections 1–10 above.
|
||||
3. **Ask the operator, then `fleet_handover{action: "confirm", token, operatorConfirmed: true}`.**
|
||||
|
||||
@@ -158,10 +172,14 @@ fails.
|
||||
session.
|
||||
- **The roll can still refuse after `confirm` returns**, and by then there is no caller to tell.
|
||||
Those outcomes are logged only, as `lead-rollover:` lines in the daemon log.
|
||||
- **If the bootstrap prompt never lands, your context is gone and no fresh session starts.** This
|
||||
has not yet been proven end-to-end (see fleetd #480). The recovery is the manual path: the file
|
||||
is already written, so the operator starts a session and points it at the file. That is why you
|
||||
write the file before you confirm, and never the other way round.
|
||||
- **The bootstrap prompt has never yet landed, and the fix is unproven (fleetd #489).** The first
|
||||
real rollover, on 2026-09-12, joined `/clear` and the bootstrap text into one line and Claude Code
|
||||
refused it as `Unknown command: /clearFresh`. The pane was never cleared and no context was lost,
|
||||
so the failure was safe — the roll simply did nothing. PR #490 fixed the cause and is deployed,
|
||||
but no roll has bootstrapped a fresh session end to end yet. **Assume it may still fail, and tell
|
||||
the operator so before you confirm.** The recovery is the same either way: the file is already
|
||||
written, so the operator starts a session and points it at the file. That is why you write the
|
||||
file before you confirm, and never the other way round.
|
||||
|
||||
## Writing style
|
||||
|
||||
|
||||
@@ -19,3 +19,9 @@
|
||||
fleetd.out
|
||||
fleetd/fleetd.out
|
||||
logs/
|
||||
|
||||
# fleetd #480: the lead rollover handover file. `leadRollover.handoverPath` points here, and the
|
||||
# outgoing lead rewrites it on every rollover. It is a snapshot of one moment's live state —
|
||||
# unpushed branches, running builds, open questions — so it is stale the moment it is written and
|
||||
# has no business in git history.
|
||||
.handover/
|
||||
|
||||
@@ -44,15 +44,18 @@ import java.util.function.Supplier;
|
||||
* {@code continuationRunner} before returning. That continuation is what actually touches the pane,
|
||||
* once the calling turn has ended, in this order:
|
||||
* <ol>
|
||||
* <li>wait for the lead's own pane to report an injectable state — i.e. wait for the very
|
||||
* {@code confirm()} call that approved this roll to finish its turn — bounded by
|
||||
* {@code turnSettleSeconds}. <strong>If this never happens, nothing else in this list runs:
|
||||
* no {@code /clear} is ever sent.</strong> A lead that never goes idle is a lead still doing
|
||||
* real work, and clearing it would throw away live context — exactly the failure this
|
||||
* correction exists to prevent.</li>
|
||||
* <li>wait for the lead's own pane to report a real turn boundary — {@code IDLE} or {@code
|
||||
* DONE}, never merely {@code BLOCKED} — i.e. wait for the very {@code confirm()} call that
|
||||
* approved this roll to finish its turn — bounded by {@code turnSettleSeconds}. <strong>If
|
||||
* this never happens, nothing else in this list runs: no {@code /clear} is ever sent.</strong>
|
||||
* A lead that never goes idle is a lead still doing real work, and clearing it would throw
|
||||
* away live context — exactly the failure this correction exists to prevent.</li>
|
||||
* <li>{@code agents.send(lead, "/clear")}</li>
|
||||
* <li>wait again for the pane to report injectable, bounded by {@code clearSettleSeconds} (this
|
||||
* is the original, pre-correction wait — still here, just no longer the only one)</li>
|
||||
* <li>wait for {@code /clear} to be picked up and settle, bounded by {@code clearSettleSeconds}
|
||||
* (fleetd #489: no longer a plain re-check of the same boundary — {@code /clear} starts no
|
||||
* turn of its own, so this instead nudges the submit keystroke while no pickup has been seen,
|
||||
* then waits for a real {@code WORKING} → {@code IDLE}/{@code DONE} boundary once one has;
|
||||
* see {@link #waitForClearPickupAndSettle})</li>
|
||||
* <li>{@code agents.send(lead, cfg.bootstrapTextFor(p.handoverPath()))}</li>
|
||||
* </ol>
|
||||
* A {@link #confirm} that returns {@link RollDecision#approved()} therefore means <em>"every gate
|
||||
@@ -111,6 +114,17 @@ public final class LeadRollover {
|
||||
/** Poll interval while waiting for the lead's pane to settle after {@code /clear}. */
|
||||
static final long SETTLE_POLL_MS = 250;
|
||||
|
||||
/**
|
||||
* How many consecutive not-yet-picked-up polls {@link #waitForClearPickupAndSettle} allows
|
||||
* before releasing rather than wedging the roll — the same constant and the same
|
||||
* release-not-wedge choice {@link dev.ltms.fleet.inject.Injector} already makes for its own
|
||||
* post-turn {@code /clear} housekeeping (fleetd #306). <strong>This bounds the number of
|
||||
* consecutive polls, not the number of nudges:</strong> the first {@code PICKUP_GRACE_POLLS - 1}
|
||||
* of those polls each send a nudge, and the {@code PICKUP_GRACE_POLLS}th releases instead of
|
||||
* nudging again — so 8 polls produce 7 nudges, not 8.
|
||||
*/
|
||||
static final int PICKUP_GRACE_POLLS = 8;
|
||||
|
||||
/**
|
||||
* One request opened by {@link #open}, pending its {@link #confirm} (or {@link #cancel}).
|
||||
*
|
||||
@@ -381,7 +395,7 @@ public final class LeadRollover {
|
||||
// /clear is housekeeping, not a delegated turn, and routing it through Injector wedges the
|
||||
// pane forever (see this class's javadoc).
|
||||
agents.send(lead, "/clear");
|
||||
boolean clearSettled = waitUntilAtTurnBoundary(lead, cfg.clearSettleSeconds());
|
||||
boolean clearSettled = waitForClearPickupAndSettle(lead, cfg.clearSettleSeconds());
|
||||
if (!clearSettled) {
|
||||
log.warn("lead-rollover: pane {} did not reach a turn boundary (IDLE or DONE) within {}s "
|
||||
+ "after /clear — NOT sending bootstrapText (token={})",
|
||||
@@ -440,11 +454,14 @@ public final class LeadRollover {
|
||||
|
||||
/**
|
||||
* Poll {@link AgentControl#status} until {@code target} reports a real turn boundary — {@link
|
||||
* AgentStatus#IDLE} or {@link AgentStatus#DONE} — bounded by {@code settleSeconds}. Used twice
|
||||
* by {@link #runRollover}: once to wait for the CALLING turn's own pane to settle (the {@code
|
||||
* turnSettleSeconds} gate that makes this correction safe), and once to wait for the pane to
|
||||
* re-settle after {@code /clear}. A failed status read degrades to "not yet settled" and is
|
||||
* retried on the next poll, the same posture {@code LeadHeartbeatLoop} and {@code
|
||||
* AgentStatus#IDLE} or {@link AgentStatus#DONE} — bounded by {@code settleSeconds}. Used once by
|
||||
* {@link #runRollover}, to wait for the CALLING turn's own pane to settle before {@code /clear}
|
||||
* is ever sent at all — the {@code turnSettleSeconds} gate that makes this correction safe. The
|
||||
* SECOND wait, after {@code /clear}, is {@link #waitForClearPickupAndSettle} instead (fleetd
|
||||
* #489) — a plain boundary check is not enough there, because {@code /clear} starts no turn of
|
||||
* its own, so this method would (wrongly) report "settled" on its very first poll whether or not
|
||||
* {@code /clear} was actually picked up. A failed status read degrades to "not yet settled" and
|
||||
* is retried on the next poll, the same posture {@code LeadHeartbeatLoop} and {@code
|
||||
* HerdrPeerLauncher}'s readiness gate already take toward an unreadable status.
|
||||
*
|
||||
* <p><strong>Deliberately not {@link AgentStatus#injectable()}.</strong> {@code injectable()}
|
||||
@@ -452,12 +469,12 @@ public final class LeadRollover {
|
||||
* turn" — and it accepts {@link AgentStatus#BLOCKED} for that purpose, because a pane paused on
|
||||
* an approval prompt is safe to queue a message behind. This class asks a stricter question —
|
||||
* "has the turn actually ended" — and {@code BLOCKED} answers no: it is a live turn that is
|
||||
* merely paused, not one that has finished. Reusing {@code injectable()} here would let both
|
||||
* waits fire into an open approval prompt mid-turn (the first wait would send {@code /clear}
|
||||
* while the lead's own {@code confirm()}-calling turn is still live and paused on a prompt; the
|
||||
* second would send {@code bootstrapText} the same way after {@code /clear}) — exactly the
|
||||
* live-context-destroying failure the {@code turnSettleSeconds} gate exists to prevent. Do not
|
||||
* "simplify" this back to {@code injectable()}.
|
||||
* merely paused, not one that has finished. Reusing {@code injectable()} here would let this
|
||||
* wait fire {@code /clear} while the lead's own {@code confirm()}-calling turn is still live and
|
||||
* paused on a prompt — exactly the live-context-destroying failure the {@code turnSettleSeconds}
|
||||
* gate exists to prevent. Do not "simplify" this back to {@code injectable()}. ({@link
|
||||
* #waitForClearPickupAndSettle} keeps the same exclusion of {@code BLOCKED}, for the same
|
||||
* reason, on the second wait.)
|
||||
*/
|
||||
private boolean waitUntilAtTurnBoundary(String target, int settleSeconds) {
|
||||
long deadline = nowMillis.getAsLong() + TimeUnit.SECONDS.toMillis(settleSeconds);
|
||||
@@ -477,4 +494,95 @@ public final class LeadRollover {
|
||||
}
|
||||
return false;
|
||||
}
|
||||
|
||||
/**
|
||||
* The SECOND wait in {@link #runRollover} — after {@code /clear} has been sent, waits for it to
|
||||
* settle, bounded by {@code settleSeconds}. <strong>fleetd #489 — the paste-race fix.</strong>
|
||||
* {@code /clear} does not start a real turn of its own, so a pane with no submit race simply
|
||||
* stays {@link AgentStatus#IDLE} the whole time: {@link #waitUntilAtTurnBoundary} would (wrongly)
|
||||
* call that "settled" on its very first poll, whether or not the {@code /clear} Enter actually
|
||||
* landed. That was Fault 1, measured live on 2026-09-12 — the second gate was a no-op, so a
|
||||
* {@code bootstrapText} send followed immediately, racing Fault 2: {@link AgentControl#submit}'s
|
||||
* own javadoc already records that the submit accompanying a delivery "can race the paste —
|
||||
* especially right as the worker's TUI becomes interactive — leaving the text unsubmitted"
|
||||
* (CB-113). Because {@code runRollover} deliberately bypasses {@code Injector} for {@code
|
||||
* /clear} (see this class's javadoc), it inherited none of {@code Injector}'s nudging — so the
|
||||
* lost {@code /clear} Enter sat in the input box and {@code bootstrapText} was typed right after
|
||||
* it, landing as one concatenated line.
|
||||
*
|
||||
* <p>This method copies the pickup-nudge pattern {@link dev.ltms.fleet.inject.Injector} already
|
||||
* ships for exactly this, on its own post-turn {@code /clear} housekeeping (fleetd #306; see
|
||||
* {@code Injector.java:288-340} and {@code Injector.java:437-442}):
|
||||
* <ul>
|
||||
* <li>an {@link AgentStatus#WORKING} sample means {@code /clear} was picked up as a real
|
||||
* turn;</li>
|
||||
* <li>until that happens, each poll that still reports {@link AgentStatus#IDLE} or {@link
|
||||
* AgentStatus#DONE} re-sends the submit keystroke ({@link AgentControl#submit}) to nudge
|
||||
* the raced Enter — for the first {@code PICKUP_GRACE_POLLS - 1} of {@link
|
||||
* #PICKUP_GRACE_POLLS} consecutive such polls (i.e. {@code PICKUP_GRACE_POLLS - 1}
|
||||
* nudges: 7, not 8, given {@code PICKUP_GRACE_POLLS = 8}). A second Enter on an empty
|
||||
* Claude Code prompt is a no-op, so repeating it is safe;</li>
|
||||
* <li>the {@code PICKUP_GRACE_POLLS}th consecutive such poll, with {@code WORKING} still never
|
||||
* observed, releases rather than wedges the roll instead of nudging again — the same
|
||||
* choice {@code Injector} makes — and returns {@code true} anyway, logged at {@code info}
|
||||
* so an operator can see which path ran;</li>
|
||||
* <li>once {@code WORKING} has been observed, nudging stops and this instead waits for a real
|
||||
* {@code working → IDLE/DONE} completion boundary before returning {@code true}.</li>
|
||||
* </ul>
|
||||
*
|
||||
* <p><strong>{@link AgentStatus#BLOCKED} is deliberately excluded from both the nudge and the
|
||||
* boundary check</strong> — the same reasoning as {@link #waitUntilAtTurnBoundary}'s own
|
||||
* javadoc: a paused live turn is not a settled one, and re-sending Enter into an open approval
|
||||
* prompt could wrongly answer it. A {@code BLOCKED} sample (or an unreadable/{@link
|
||||
* AgentStatus#UNKNOWN} one) simply keeps this polling, with no nudge and no release, until either
|
||||
* a real boundary is reached or {@code settleSeconds} runs out.
|
||||
*
|
||||
* <p>{@link AgentControl#submit} can itself throw; a {@link RuntimeException} from it is
|
||||
* swallowed and logged at {@code debug}, exactly like {@code Injector.java:437-442} — a failed
|
||||
* nudge must not abort the roll.
|
||||
*
|
||||
* @return {@code true} once {@code /clear} has settled, or once the nudge budget was exhausted
|
||||
* with no pickup ever observed (released rather than wedged); {@code false} if {@code
|
||||
* settleSeconds} elapses first — the caller must NOT send {@code bootstrapText} in that
|
||||
* case, exactly as before this fix
|
||||
*/
|
||||
private boolean waitForClearPickupAndSettle(String target, int settleSeconds) {
|
||||
long deadline = nowMillis.getAsLong() + TimeUnit.SECONDS.toMillis(settleSeconds);
|
||||
boolean pickedUp = false; // a WORKING sample has been observed since /clear was sent
|
||||
int idlePollsAwaitingPickup = 0;
|
||||
while (nowMillis.getAsLong() < deadline) {
|
||||
AgentStatus status;
|
||||
try {
|
||||
status = agents.status(target);
|
||||
} catch (RuntimeException e) {
|
||||
log.debug("lead-rollover: status check failed while waiting for {} to settle after "
|
||||
+ "/clear: {}", target, e.toString());
|
||||
status = null;
|
||||
}
|
||||
if (status == AgentStatus.WORKING) {
|
||||
pickedUp = true;
|
||||
} else if (status == AgentStatus.IDLE || status == AgentStatus.DONE) {
|
||||
if (pickedUp) {
|
||||
return true; // a real WORKING -> IDLE/DONE completion boundary
|
||||
}
|
||||
if (++idlePollsAwaitingPickup >= PICKUP_GRACE_POLLS) {
|
||||
log.info("lead-rollover: /clear on {} was never observed as WORKING after {} "
|
||||
+ "consecutive IDLE/DONE polls ({} of those were nudged) — "
|
||||
+ "releasing rather than wedging the roll",
|
||||
target, PICKUP_GRACE_POLLS, PICKUP_GRACE_POLLS - 1);
|
||||
return true;
|
||||
}
|
||||
try {
|
||||
agents.submit(target); // nudge a raced Enter (CB-113) so /clear actually submits
|
||||
} catch (RuntimeException e) {
|
||||
log.debug("lead-rollover: resubmit to {} failed (will retry next poll): {}",
|
||||
target, e.getMessage());
|
||||
}
|
||||
}
|
||||
// AgentStatus.BLOCKED or UNKNOWN (or an unreadable status, above): neither a pickup
|
||||
// signal nor a boundary — keep polling without nudging or releasing.
|
||||
settleSleeper.run();
|
||||
}
|
||||
return false;
|
||||
}
|
||||
}
|
||||
|
||||
@@ -458,6 +458,184 @@ class LeadRolloverTest {
|
||||
assertEquals(0, bootstrapSends, "bootstrapText must never be sent when /clear did not settle");
|
||||
}
|
||||
|
||||
// ---- fleetd #489: the second wait nudges the /clear pickup instead of being a no-op --------
|
||||
|
||||
/** Every {@code agent.send_keys} call {@code herdr} recorded — the submit-keystroke nudge. */
|
||||
private static long sendKeysCallCount(FakeHerdr herdr) {
|
||||
return herdr.calls.stream().filter(c -> "agent.send_keys".equals(c.method())).count();
|
||||
}
|
||||
|
||||
@Test
|
||||
@DisplayName("[fleetd #489] a pane that stays IDLE the whole time (the paste-race case, "
|
||||
+ "measured live 2026-09-12) is nudged between /clear and bootstrapText, never lets "
|
||||
+ "them concatenate into one line")
|
||||
void clearPickupIsNudgedBeforeBootstrapTextWhenPaneStaysIdle() throws IOException {
|
||||
FakeHerdr herdr = new FakeHerdr(); // default agentStatus is "idle" throughout — no WORKING sample ever
|
||||
Path handover = writeHandover("handover contents");
|
||||
AtomicLong clock = new AtomicLong(1_000);
|
||||
LeadRollover rollover = newRollover(herdr, cfg(handover.toString()), fixedClock(clock));
|
||||
|
||||
LeadRollover.PendingRollover pending = rollover.open(LEAD, "context is full");
|
||||
LeadRollover.RollDecision decision = rollover.confirm(LEAD, pending.token(), true);
|
||||
|
||||
assertTrue(decision.accepted(), "expected approval; got: " + decision.reason() + " / " + decision.detail());
|
||||
|
||||
int clearIdx = -1;
|
||||
int bootstrapIdx = -1;
|
||||
int firstNudgeIdx = -1;
|
||||
for (int i = 0; i < herdr.calls.size(); i++) {
|
||||
FakeHerdr.Call c = herdr.calls.get(i);
|
||||
if ("agent.prompt".equals(c.method()) && String.valueOf(c.params()).contains("/clear") && clearIdx < 0) {
|
||||
clearIdx = i;
|
||||
} else if ("agent.prompt".equals(c.method()) && String.valueOf(c.params()).contains("read the handover file")) {
|
||||
bootstrapIdx = i;
|
||||
} else if ("agent.send_keys".equals(c.method()) && firstNudgeIdx < 0) {
|
||||
firstNudgeIdx = i;
|
||||
}
|
||||
}
|
||||
assertTrue(clearIdx >= 0, "/clear must have been sent");
|
||||
assertTrue(bootstrapIdx >= 0, "bootstrapText must have been sent");
|
||||
assertTrue(firstNudgeIdx >= 0, "at least one agent.send_keys nudge must go out — the pane "
|
||||
+ "never reported WORKING, so the /clear Enter may have raced the paste, and only a "
|
||||
+ "re-sent Enter proves the clear rather than concatenating bootstrapText onto "
|
||||
+ "whatever sits unsubmitted in the input box");
|
||||
assertTrue(firstNudgeIdx > clearIdx, "the nudge must happen AFTER /clear was sent, got call "
|
||||
+ "order: " + herdr.calls);
|
||||
assertTrue(firstNudgeIdx < bootstrapIdx, "the nudge must happen BEFORE bootstrapText is "
|
||||
+ "sent — never concatenated onto the same input line, got call order: " + herdr.calls);
|
||||
}
|
||||
|
||||
@Test
|
||||
@DisplayName("[fleetd #489] once a WORKING sample confirms /clear was picked up, nudging stops "
|
||||
+ "and bootstrapText is still sent after the pane returns to IDLE")
|
||||
void pickupSeenStopsNudgingAndBootstrapTextIsSent() throws IOException {
|
||||
FakeHerdr fake = new FakeHerdr();
|
||||
// Scripts the SECOND wait only: idle (default) until /clear is sent, then the first status
|
||||
// poll after /clear reports WORKING (a confirmed pickup), and every poll after that reports
|
||||
// IDLE (the completion boundary). The first wait (turnSettleSeconds) never sees this
|
||||
// sequence — it passes on its own first poll, before /clear is ever sent, on the default
|
||||
// "idle" status.
|
||||
AtomicLong postClearGetCalls = new AtomicLong(0);
|
||||
HerdrClient scriptsPickupThenIdle = new HerdrClient() {
|
||||
private volatile boolean clearSent = false;
|
||||
|
||||
@Override
|
||||
public JsonNode call(String method, Object params) throws HerdrException {
|
||||
if ("agent.prompt".equals(method) && String.valueOf(params).contains("/clear")) {
|
||||
clearSent = true;
|
||||
}
|
||||
if (clearSent && "agent.get".equals(method)) {
|
||||
long n = postClearGetCalls.incrementAndGet();
|
||||
fake.agentStatus(n == 1 ? "working" : "idle");
|
||||
}
|
||||
return fake.call(method, params);
|
||||
}
|
||||
|
||||
@Override
|
||||
public void close() {
|
||||
fake.close();
|
||||
}
|
||||
};
|
||||
Path handover = writeHandover("handover contents");
|
||||
AtomicLong clock = new AtomicLong(1_000);
|
||||
LeadRollover rollover = newRollover(scriptsPickupThenIdle, cfg(handover.toString()), fixedClock(clock));
|
||||
|
||||
LeadRollover.PendingRollover pending = rollover.open(LEAD, "context is full");
|
||||
LeadRollover.RollDecision decision = rollover.confirm(LEAD, pending.token(), true);
|
||||
|
||||
assertTrue(decision.accepted(), "expected approval; got: " + decision.reason() + " / " + decision.detail());
|
||||
assertEquals(2, promptCallCount(fake), "a confirmed WORKING pickup followed by IDLE must "
|
||||
+ "still complete the full roll — /clear then bootstrapText");
|
||||
assertEquals(0, sendKeysCallCount(fake), "once WORKING was observed, nudging must stop "
|
||||
+ "immediately — no agent.send_keys call should ever have been needed or sent");
|
||||
assertTrue(postClearGetCalls.get() >= 2, "the pane's status must have been polled AGAIN "
|
||||
+ "after the WORKING sample, before bootstrapText was sent — this is what proves "
|
||||
+ "the method actually waited for the WORKING -> IDLE completion boundary instead "
|
||||
+ "of returning as soon as pickup was seen (or worse, without polling at all, as a "
|
||||
+ "stub that just returns true would); got " + postClearGetCalls.get()
|
||||
+ " agent.get call(s) after /clear");
|
||||
}
|
||||
|
||||
@Test
|
||||
@DisplayName("[fleetd #489] a pane that never reaches a turn boundary after /clear (stuck at "
|
||||
+ "UNKNOWN, never WORKING either) lets clearSettleSeconds expire — bootstrapText is "
|
||||
+ "never sent")
|
||||
void clearPickupNeverSettlesWhenStatusNeverReachesABoundary() throws IOException {
|
||||
FakeHerdr fake = new FakeHerdr();
|
||||
HerdrClient stuckUnknownAfterClear = new HerdrClient() {
|
||||
@Override
|
||||
public JsonNode call(String method, Object params) throws HerdrException {
|
||||
if ("agent.prompt".equals(method) && String.valueOf(params).contains("/clear")) {
|
||||
// "wedged" maps to AgentStatus.UNKNOWN (see AgentStatus#fromWire) — neither a
|
||||
// pickup signal (WORKING) nor a boundary (IDLE/DONE), and distinct from the
|
||||
// already-covered BLOCKED case below.
|
||||
fake.agentStatus("wedged");
|
||||
}
|
||||
return fake.call(method, params);
|
||||
}
|
||||
|
||||
@Override
|
||||
public void close() {
|
||||
fake.close();
|
||||
}
|
||||
};
|
||||
Path handover = writeHandover("handover contents");
|
||||
FleetConfig.LeadRollover config =
|
||||
new FleetConfig.LeadRollover(handover.toString(), true, 3600, 20, 1 /*clearSettleSeconds*/, "boot text");
|
||||
AtomicLong clock = new AtomicLong(1_000);
|
||||
LeadRollover rollover = newRollover(stuckUnknownAfterClear, config, () -> clock.addAndGet(500));
|
||||
|
||||
LeadRollover.PendingRollover pending = rollover.open(LEAD, "context is full");
|
||||
LeadRollover.RollDecision decision = rollover.confirm(LEAD, pending.token(), true);
|
||||
|
||||
assertTrue(decision.accepted(), "every synchronous gate passes; the refusal is logged only, "
|
||||
+ "deep inside the deferred continuation");
|
||||
assertEquals(1, promptCallCount(fake), "exactly one agent.prompt call — the /clear — and "
|
||||
+ "nothing else");
|
||||
long bootstrapSends = fake.calls.stream()
|
||||
.filter(c -> "agent.prompt".equals(c.method()))
|
||||
.filter(c -> String.valueOf(c.params()).contains("boot text"))
|
||||
.count();
|
||||
assertEquals(0, bootstrapSends, "bootstrapText must never be sent when the pane never "
|
||||
+ "reaches a turn boundary after /clear, whether WORKING was ever observed or not");
|
||||
assertEquals(0, sendKeysCallCount(fake), "an UNKNOWN status is neither a pickup signal nor "
|
||||
+ "a boundary — it must never be nudged");
|
||||
}
|
||||
|
||||
@Test
|
||||
@DisplayName("[fleetd #489] a submit() nudge that throws does not abort the roll — /clear and "
|
||||
+ "bootstrapText are both still sent")
|
||||
void submitThatThrowsDoesNotAbortTheRoll() throws IOException {
|
||||
FakeHerdr fake = new FakeHerdr(); // default idle throughout — nudging will be attempted
|
||||
HerdrClient throwsOnSubmit = new HerdrClient() {
|
||||
@Override
|
||||
public JsonNode call(String method, Object params) throws HerdrException {
|
||||
if ("agent.send_keys".equals(method)) {
|
||||
throw new RuntimeException("simulated herdr transport failure on submit");
|
||||
}
|
||||
return fake.call(method, params);
|
||||
}
|
||||
|
||||
@Override
|
||||
public void close() {
|
||||
fake.close();
|
||||
}
|
||||
};
|
||||
Path handover = writeHandover("handover contents");
|
||||
AtomicLong clock = new AtomicLong(1_000);
|
||||
LeadRollover rollover = newRollover(throwsOnSubmit, cfg(handover.toString()), fixedClock(clock));
|
||||
|
||||
LeadRollover.PendingRollover pending = rollover.open(LEAD, "context is full");
|
||||
LeadRollover.RollDecision decision = rollover.confirm(LEAD, pending.token(), true);
|
||||
|
||||
assertTrue(decision.accepted(), "expected approval; got: " + decision.reason() + " / " + decision.detail());
|
||||
var prompts = fake.calls.stream().filter(c -> "agent.prompt".equals(c.method())).toList();
|
||||
assertEquals(2, prompts.size(), "a throwing submit() must be swallowed, not abort the roll "
|
||||
+ "— /clear and bootstrapText must both still be sent");
|
||||
assertTrue(prompts.get(0).params().toString().contains("/clear"));
|
||||
assertTrue(prompts.get(1).params().toString().contains("read the handover file"));
|
||||
}
|
||||
|
||||
@Test
|
||||
@DisplayName("a full successful roll sends /clear then bootstrapText, in order, and consumes the token")
|
||||
void successfulRollSendsClearThenBootstrapTextAndConsumesTheToken() throws IOException {
|
||||
|
||||
+217
-54
@@ -5,7 +5,7 @@
|
||||
# A merge is not a deployment: the running daemon holds the jar it was started with, so code merged
|
||||
# to main does nothing until this runs. See CLAUDE.md -> "Redeploying the daemon".
|
||||
#
|
||||
# This script exists to turn six remembered traps into one auditable command:
|
||||
# This script exists to turn eight remembered traps into one auditable command:
|
||||
#
|
||||
# 1. A piped `mvn` hides BUILD FAILURE behind a zero exit, so the build here is never piped.
|
||||
# 2. The daemon must start from a LOGIN shell, or the tokens it hands to members are empty:
|
||||
@@ -27,6 +27,17 @@
|
||||
# restart of the OLD jar. So this script detects whether the agent is loaded and, only then,
|
||||
# swaps `kill` + manual `nohup` for `launchctl unload`/`load` — the one supervisor in control
|
||||
# at any moment is whichever one you asked to act, never both.
|
||||
# 7. fleetd #492 — a systemd --user unit is a THIRD possible supervisor (seen on a second host):
|
||||
# Restart=on-failure treats this JVM's SIGTERM exit code (143, per CB-594 above) as a failure
|
||||
# too, so a bare `kill` there would race systemd's own restart of the OLD jar exactly like
|
||||
# launchd would. This script now tells launchd, systemd, and "genuinely unsupervised" apart as
|
||||
# three different answers, drives whichever one it finds through its own control plane
|
||||
# (`launchctl` / `systemctl --user`), and REFUSES outright — never falls back to `kill` — when
|
||||
# it finds a supervision signal it cannot map to exactly one of the two it knows how to drive.
|
||||
# A wrong guess here is how two daemons end up running against one herdr session.
|
||||
# 8. fleetd #492 — a post-restart check counts running fleetd processes and fails the whole run if
|
||||
# more than one is alive. That is the one thing none of the checks above (healthz 200, jar id,
|
||||
# the fresh "listening" line) can see: every one of them is satisfied by EITHER daemon.
|
||||
#
|
||||
# Usage:
|
||||
# scripts/redeploy-fleetd.sh # build, confirm, restart, verify
|
||||
@@ -55,13 +66,18 @@ HEALTH_WAIT=60 # seconds to wait for /healthz to answer after start
|
||||
LAUNCHD_LABEL='dev.ltms.fleetd'
|
||||
LAUNCHD_PLIST="$HOME/Library/LaunchAgents/$LAUNCHD_LABEL.plist"
|
||||
|
||||
# fleetd #492: the systemd --user unit this script must not fight with either (see trap 7 above).
|
||||
# Measured on the second host: `systemctl --user cat fleetd` names the unit "fleetd" (not
|
||||
# "dev.ltms.fleetd" — systemd user units here are not namespaced the way the launchd label is).
|
||||
SYSTEMD_UNIT='fleetd'
|
||||
|
||||
DO_BUILD=1; ASSUME_YES=0; CHECK_ONLY=0
|
||||
for arg in "$@"; do
|
||||
case "$arg" in
|
||||
--yes|-y) ASSUME_YES=1 ;;
|
||||
--no-build) DO_BUILD=0 ;;
|
||||
--check) CHECK_ONLY=1 ;;
|
||||
-h|--help) sed -n '3,37p' "${BASH_SOURCE[0]}"; exit 0 ;;
|
||||
-h|--help) sed -n '3,48p' "${BASH_SOURCE[0]}"; exit 0 ;;
|
||||
*) echo "unknown option: $arg (try --help)" >&2; exit 2 ;;
|
||||
esac
|
||||
done
|
||||
@@ -79,6 +95,91 @@ running_pid() { pgrep -f "$PATTERN" || true; }
|
||||
launchd_installed() { [ -f "$LAUNCHD_PLIST" ]; }
|
||||
launchd_loaded() { launchctl list "$LAUNCHD_LABEL" >/dev/null 2>&1; }
|
||||
|
||||
# fleetd #492: same two questions for systemd --user. Kept as separate, overridable functions
|
||||
# (never an inline `systemctl` call at each use site) so a test on a box with no systemd at all
|
||||
# (this repo is developed on macOS) can substitute each one independently — the same seam
|
||||
# launchd_installed/launchd_loaded above already use.
|
||||
#
|
||||
# "installed": a unit FILE by this name exists, regardless of its current state — the systemd
|
||||
# analogue of the plist file existing on disk. `list-unit-files` reads unit definitions without
|
||||
# depending on runtime state, so this stays read-only and safe under --check.
|
||||
systemd_installed() {
|
||||
command -v systemctl >/dev/null 2>&1 \
|
||||
&& systemctl --user list-unit-files "$SYSTEMD_UNIT.service" --no-legend 2>/dev/null | grep -q .
|
||||
}
|
||||
# "loaded": systemd currently supervises this unit as an active job — the systemd analogue of
|
||||
# `launchctl list <label>` succeeding. Measured on the second host: `systemctl --user is-active
|
||||
# fleetd` -> "active".
|
||||
systemd_loaded() {
|
||||
command -v systemctl >/dev/null 2>&1 && systemctl --user is-active "$SYSTEMD_UNIT" >/dev/null 2>&1
|
||||
}
|
||||
|
||||
# fleetd #492: three real answers, not two — launchd, systemd, or genuinely unsupervised — plus a
|
||||
# fourth, "ambiguous", for the one case this script cannot tell apart: both signals firing at once.
|
||||
# That is exactly "I cannot tell who supervises this process", and guessing wrong here is how two
|
||||
# daemons end up running against one herdr session (see trap 7 in the header). Pure and
|
||||
# side-effect-free: reads the two probes above and decides — never mutates anything, so it is safe
|
||||
# under --check and testable by overriding launchd_loaded/systemd_loaded after sourcing.
|
||||
detect_supervisor() {
|
||||
local ld=0 sd=0
|
||||
launchd_loaded && ld=1
|
||||
systemd_loaded && sd=1
|
||||
if [ "$ld" = 1 ] && [ "$sd" = 1 ]; then
|
||||
echo "ambiguous"
|
||||
elif [ "$ld" = 1 ]; then
|
||||
echo "launchd"
|
||||
elif [ "$sd" = 1 ]; then
|
||||
echo "systemd"
|
||||
else
|
||||
echo "none"
|
||||
fi
|
||||
}
|
||||
|
||||
# fleetd #492: turns anything detect_supervisor returns that is NOT exactly one of the two
|
||||
# supervisors this script knows how to drive into a die() — never a fall-through to the `kill`
|
||||
# path. Kept as its own function so a test can call it directly (in a subshell, since it die()s)
|
||||
# without running the whole report-state flow or needing a real launchd/systemd.
|
||||
require_drivable_supervisor() {
|
||||
local kind="$1"
|
||||
case "$kind" in
|
||||
launchd|systemd|none) ;;
|
||||
ambiguous)
|
||||
die "both launchd ($LAUNCHD_LABEL) and systemd --user ($SYSTEMD_UNIT) report themselves as
|
||||
loaded for this daemon at the same time. This script cannot tell which one actually
|
||||
supervises the running process, and driving either alone risks the OTHER reviving the
|
||||
OLD jar out from under it — the exact failure this ticket (fleetd #492) exists to
|
||||
prevent. Stop one of the two supervisors by hand, confirm only one remains loaded, then
|
||||
rerun." ;;
|
||||
*)
|
||||
die "detect_supervisor returned an unrecognized value '$kind' — refusing to guess which
|
||||
supervisor, if any, controls this daemon." ;;
|
||||
esac
|
||||
}
|
||||
|
||||
# fleetd #492: the exact symptom a racing supervisor produces — count how many fleetd processes are
|
||||
# alive right now. Takes the pid list as a parameter (rather than calling running_pid() itself) so a
|
||||
# test can pass a canned two-line string without a real second process running. Pure except for the
|
||||
# die() in assert_single_daemon below.
|
||||
count_daemon_pids() {
|
||||
local pids="$1"
|
||||
if [ -z "$pids" ]; then
|
||||
echo 0
|
||||
else
|
||||
printf '%s\n' "$pids" | grep -c .
|
||||
fi
|
||||
}
|
||||
assert_single_daemon() {
|
||||
local pids="$1" count
|
||||
count="$(count_daemon_pids "$pids")"
|
||||
if [ "$count" -gt 1 ]; then
|
||||
die "more than one fleetd process is running after this restart (pids: $(printf '%s' "$pids" | tr '\n' ' ')).
|
||||
This is the exact failure a racing supervisor produces: the OLD jar was revived by its
|
||||
supervisor while this script started a NEW copy. Two daemons on one herdr session kill
|
||||
each other's members. Investigate with 'pgrep -f \"$PATTERN\"' and stop the wrong one by
|
||||
hand — do not assume either pid is the one you want."
|
||||
fi
|
||||
}
|
||||
|
||||
# CB-600: the script computes its own log path from where it sits on disk (REPO, above); the
|
||||
# plist hard-codes an absolute StandardOutPath. Nothing forced the two to agree — if this script
|
||||
# were ever run from a checkout other than the one the loaded plist names, launchd would start and
|
||||
@@ -182,25 +283,45 @@ fi
|
||||
ok "jar on disk: $(jar_id) ($([ -f "$JAR" ] && date -r "$JAR" '+%Y-%m-%d %H:%M:%S' || echo 'none'))"
|
||||
ok "HEAD: $(git -C "$REPO" log --oneline -1)"
|
||||
|
||||
# CB-594: supervision state. Installed and loaded are different facts — a copied-but-never-loaded
|
||||
# plist supervises nothing, and a loaded label with no file backing it (rare, but possible after an
|
||||
# edited/moved plist) is still what launchd will act on.
|
||||
# CB-594 / fleetd #492: supervision state. Installed and loaded are different facts — a
|
||||
# copied-but-never-loaded plist (or an unloaded systemd unit) supervises nothing, and a loaded
|
||||
# label/unit with no file backing it is still what its supervisor will act on.
|
||||
if launchd_installed; then
|
||||
ok "launchd agent installed: $LAUNCHD_PLIST"
|
||||
else
|
||||
warn "launchd agent NOT installed (no supervision — a crash will not restart the daemon)."
|
||||
warn "launchd agent NOT installed."
|
||||
fi
|
||||
SUPERVISED=0
|
||||
if launchd_loaded; then
|
||||
SUPERVISED=1
|
||||
ok "launchd agent loaded ($LAUNCHD_LABEL) — launchd supervises this daemon"
|
||||
# CB-600: fail loudly here, before ANY other check runs, if this script and the loaded plist
|
||||
# would read different log files — every check after this point is worthless otherwise.
|
||||
check_log_path_matches_plist "$OUT" "$LAUNCHD_PLIST"
|
||||
if systemd_installed; then
|
||||
ok "systemd --user unit installed: $SYSTEMD_UNIT"
|
||||
else
|
||||
warn "launchd agent not loaded — this script is the only thing that will restart the daemon."
|
||||
warn "systemd --user unit NOT installed ($SYSTEMD_UNIT)."
|
||||
fi
|
||||
|
||||
# fleetd #492: decide which of the two (if either) actually supervises this daemon, and refuse
|
||||
# outright — before touching anything — if that cannot be told apart (see require_drivable_
|
||||
# supervisor above). --check reaches this same line, so a host with an undrivable supervisor is
|
||||
# reported as a failure even in --check, without ever reaching the build/stop/start steps.
|
||||
SUPERVISOR_KIND="$(detect_supervisor)"
|
||||
require_drivable_supervisor "$SUPERVISOR_KIND"
|
||||
ok "supervisor detected: $SUPERVISOR_KIND"
|
||||
SUPERVISED=0
|
||||
case "$SUPERVISOR_KIND" in
|
||||
launchd)
|
||||
SUPERVISED=1
|
||||
ok "launchd agent loaded ($LAUNCHD_LABEL) — launchd supervises this daemon"
|
||||
# CB-600: fail loudly here, before ANY other check runs, if this script and the loaded plist
|
||||
# would read different log files — every check after this point is worthless otherwise.
|
||||
check_log_path_matches_plist "$OUT" "$LAUNCHD_PLIST"
|
||||
;;
|
||||
systemd)
|
||||
SUPERVISED=1
|
||||
ok "systemd --user unit active ($SYSTEMD_UNIT) — systemd supervises this daemon"
|
||||
;;
|
||||
none)
|
||||
warn "no supervisor loaded — this script is the only thing that will restart the daemon."
|
||||
;;
|
||||
esac
|
||||
|
||||
# The trap with no log line. Checked in a LOGIN shell, because that is how the daemon is started
|
||||
# below. Never prints the value — only whether it resolved.
|
||||
if zsh -lc '[ -n "${WORKER_GITEA_TOKEN:-}" ]' 2>/dev/null; then
|
||||
@@ -281,25 +402,39 @@ fi
|
||||
|
||||
# ------------------------------------------------------------------ stop
|
||||
#
|
||||
# CB-594: when SUPERVISED, launchd owns the stop — never a raw `kill` here. A bare SIGTERM makes
|
||||
# this JVM exit 143 even with its shutdown hook running to completion (verified separately: a
|
||||
# throwaway Java process with an equivalent shutdown hook, sent SIGTERM from a login shell that
|
||||
# could `wait` on it directly, reported exit code 143 every time — never 0). launchd's
|
||||
# KeepAlive.SuccessfulExit=false treats any nonzero exit as a crash and restarts the OLD jar,
|
||||
# which would race this script's own restart of the NEW one. `launchctl unload` avoids that race
|
||||
# by deregistering the job first, so no KeepAlive is left armed when the process actually stops.
|
||||
# CB-594 / fleetd #492: when SUPERVISED, the supervisor owns the stop — never a raw `kill` here. A
|
||||
# bare SIGTERM makes this JVM exit 143 even with its shutdown hook running to completion (verified
|
||||
# separately: a throwaway Java process with an equivalent shutdown hook, sent SIGTERM from a login
|
||||
# shell that could `wait` on it directly, reported exit code 143 every time — never 0). launchd's
|
||||
# KeepAlive.SuccessfulExit=false and systemd's Restart=on-failure both treat any nonzero exit as a
|
||||
# crash and restart the OLD jar, which would race this script's own restart of the NEW one.
|
||||
# `launchctl unload` avoids that race by deregistering the job first, so no KeepAlive is left
|
||||
# armed when the process actually stops. `systemctl --user stop` needs no such dance: unlike
|
||||
# KeepAlive, systemd's Restart= does not fire on a deliberate stop, only on an unexpected exit of
|
||||
# an active unit.
|
||||
|
||||
if [ -n "$OLD_PID" ]; then
|
||||
say "stop"
|
||||
RESTART_MARK="$(wc -l < "$OUT" 2>/dev/null || echo 0)" # verify a FRESH line appears later
|
||||
if [ "$SUPERVISED" = 1 ]; then
|
||||
echo " supervision is ON: using 'launchctl unload' (not kill) so launchd's own KeepAlive"
|
||||
echo " cannot restart the OLD jar out from under this script — see the CB-594 comment above."
|
||||
launchctl unload -w "$LAUNCHD_PLIST" \
|
||||
|| die "launchctl unload failed — the daemon may still be under supervision; investigate before retrying"
|
||||
else
|
||||
kill "$OLD_PID"
|
||||
fi
|
||||
case "$SUPERVISOR_KIND" in
|
||||
launchd)
|
||||
echo " supervision is ON (launchd): using 'launchctl unload' (not kill) so launchd's own"
|
||||
echo " KeepAlive cannot restart the OLD jar out from under this script — see the CB-594"
|
||||
echo " comment above."
|
||||
launchctl unload -w "$LAUNCHD_PLIST" \
|
||||
|| die "launchctl unload failed — the daemon may still be under supervision; investigate before retrying"
|
||||
;;
|
||||
systemd)
|
||||
echo " supervision is ON (systemd --user): using 'systemctl --user stop' (not kill) so"
|
||||
echo " systemd's own Restart=on-failure cannot restart the OLD jar out from under this"
|
||||
echo " script — see the fleetd #492 comment above."
|
||||
systemctl --user stop "$SYSTEMD_UNIT" \
|
||||
|| die "'systemctl --user stop $SYSTEMD_UNIT' failed — the daemon may still be under supervision; investigate before retrying"
|
||||
;;
|
||||
none)
|
||||
kill "$OLD_PID"
|
||||
;;
|
||||
esac
|
||||
for _ in $(seq "$STOP_WAIT"); do
|
||||
[ -z "$(running_pid)" ] && break
|
||||
sleep 1
|
||||
@@ -310,13 +445,20 @@ if [ -n "$OLD_PID" ]; then
|
||||
leave worktrees and panes behind. Investigate, then kill -9 by hand if you accept that."
|
||||
fi
|
||||
ok "pid $OLD_PID exited"
|
||||
elif [ "$SUPERVISED" = 1 ]; then
|
||||
elif [ "$SUPERVISOR_KIND" = "launchd" ]; then
|
||||
# Loaded but not currently running (e.g. throttled after a crash loop). Unload it anyway so the
|
||||
# start step below does a clean load, never a load stacked on an already-loaded label.
|
||||
say "stop"
|
||||
RESTART_MARK="$(wc -l < "$OUT" 2>/dev/null || echo 0)"
|
||||
launchctl unload -w "$LAUNCHD_PLIST" 2>/dev/null || true
|
||||
ok "launchd agent unloaded (was already not running)"
|
||||
elif [ "$SUPERVISOR_KIND" = "systemd" ]; then
|
||||
# Same case for systemd: the unit is known/active-capable but not currently running. `stop` on an
|
||||
# already-stopped unit is a harmless no-op — kept for symmetry with the launchd branch above.
|
||||
say "stop"
|
||||
RESTART_MARK="$(wc -l < "$OUT" 2>/dev/null || echo 0)"
|
||||
systemctl --user stop "$SYSTEMD_UNIT" 2>/dev/null || true
|
||||
ok "systemd --user unit stopped (was already not running)"
|
||||
else
|
||||
RESTART_MARK="$(wc -l < "$OUT" 2>/dev/null || echo 0)"
|
||||
fi
|
||||
@@ -324,34 +466,48 @@ fi
|
||||
# ------------------------------------------------------------------ start
|
||||
# Unsupervised: login shell (zsh -l) is what puts the secrets on the daemon's environment, and cwd
|
||||
# must be fleetd/ because the daemon resolves fleetd.yaml, logs/ and target/ relative to it.
|
||||
# Supervised: launchd does both — deploy/dev.ltms.fleetd.plist points ProgramArguments at
|
||||
# Supervised (launchd): launchd does both — deploy/dev.ltms.fleetd.plist points ProgramArguments at
|
||||
# scripts/fleetd-launchd-wrapper.sh (CB-594), which is what execs the login shell in launchd's
|
||||
# place, and WorkingDirectory in the plist already pins fleetd/.
|
||||
# Supervised (systemd --user): the unit does both too — measured on the second host, ExecStart is
|
||||
# `/bin/zsh -lc "exec java -jar target/fleetd.jar fleetd.yaml"` (a login shell, same reason as
|
||||
# above) and WorkingDirectory is already pinned to fleetd/.
|
||||
|
||||
say "start"
|
||||
if [ "$SUPERVISED" = 1 ]; then
|
||||
echo " supervision is ON: using 'launchctl load' so launchd starts and keeps supervising this"
|
||||
echo " process, instead of a manual nohup that launchd would know nothing about."
|
||||
# CB-600: 'launchctl unload -w' above already persisted Disabled=true for this label. A load -w
|
||||
# that succeeds clears it; a load -w that FAILS leaves the agent both stopped and disabled — worse
|
||||
# than before this script ran, because a later reboot or login will not bring it back either. One
|
||||
# retry covers a transient race (e.g. launchd not yet fully done deregistering); if it still fails,
|
||||
# die with the exact recovery command rather than a bare "failed".
|
||||
if ! launchctl load -w "$LAUNCHD_PLIST" 2>/dev/null; then
|
||||
warn "launchctl load failed on the first attempt — retrying once after a short pause"
|
||||
sleep 2
|
||||
launchctl load -w "$LAUNCHD_PLIST" || die "launchctl load failed twice.
|
||||
The agent is now STOPPED and DISABLED — it will NOT come back on its own, not even after a
|
||||
reboot or login, because 'launchctl unload -w' above persisted Disabled=true and load -w
|
||||
never got the chance to clear it. Recover with:
|
||||
launchctl load -w \"$LAUNCHD_PLIST\"
|
||||
If that still fails, check 'launchctl list $LAUNCHD_LABEL', validate the plist with
|
||||
'plutil -lint \"$LAUNCHD_PLIST\"', and check $OUT before assuming a retry will succeed."
|
||||
fi
|
||||
else
|
||||
# Absolute jar path so `ps` names which checkout is running.
|
||||
( cd "$MODULE" && zsh -lc "nohup java -jar '$JAR' >> fleetd.out 2>&1 &" )
|
||||
fi
|
||||
case "$SUPERVISOR_KIND" in
|
||||
launchd)
|
||||
echo " supervision is ON (launchd): using 'launchctl load' so launchd starts and keeps"
|
||||
echo " supervising this process, instead of a manual nohup that launchd would know nothing"
|
||||
echo " about."
|
||||
# CB-600: 'launchctl unload -w' above already persisted Disabled=true for this label. A load -w
|
||||
# that succeeds clears it; a load -w that FAILS leaves the agent both stopped and disabled — worse
|
||||
# than before this script ran, because a later reboot or login will not bring it back either. One
|
||||
# retry covers a transient race (e.g. launchd not yet fully done deregistering); if it still fails,
|
||||
# die with the exact recovery command rather than a bare "failed".
|
||||
if ! launchctl load -w "$LAUNCHD_PLIST" 2>/dev/null; then
|
||||
warn "launchctl load failed on the first attempt — retrying once after a short pause"
|
||||
sleep 2
|
||||
launchctl load -w "$LAUNCHD_PLIST" || die "launchctl load failed twice.
|
||||
The agent is now STOPPED and DISABLED — it will NOT come back on its own, not even after a
|
||||
reboot or login, because 'launchctl unload -w' above persisted Disabled=true and load -w
|
||||
never got the chance to clear it. Recover with:
|
||||
launchctl load -w \"$LAUNCHD_PLIST\"
|
||||
If that still fails, check 'launchctl list $LAUNCHD_LABEL', validate the plist with
|
||||
'plutil -lint \"$LAUNCHD_PLIST\"', and check $OUT before assuming a retry will succeed."
|
||||
fi
|
||||
;;
|
||||
systemd)
|
||||
echo " supervision is ON (systemd --user): using 'systemctl --user start' so systemd starts"
|
||||
echo " and keeps supervising this process, instead of a manual nohup it would know nothing"
|
||||
echo " about."
|
||||
systemctl --user start "$SYSTEMD_UNIT" || die "'systemctl --user start $SYSTEMD_UNIT' failed.
|
||||
Check 'systemctl --user status $SYSTEMD_UNIT' and $OUT before assuming a retry will succeed."
|
||||
;;
|
||||
none)
|
||||
# Absolute jar path so `ps` names which checkout is running.
|
||||
( cd "$MODULE" && zsh -lc "nohup java -jar '$JAR' >> fleetd.out 2>&1 &" )
|
||||
;;
|
||||
esac
|
||||
|
||||
for _ in $(seq 10); do
|
||||
NEW_PID="$(running_pid)"
|
||||
@@ -411,6 +567,13 @@ FRESH_LOG="$(mktemp -t fleetd-fresh-log)"
|
||||
trap 'rm -f "$FRESH_LOG"' EXIT
|
||||
tail -n "+$((RESTART_MARK + 1))" "$OUT" > "$FRESH_LOG" 2>/dev/null || true
|
||||
classify_amqp_connection_errors "$FRESH_LOG"
|
||||
|
||||
# fleetd #492: checked here, after healthz and the fresh-log check have both had time to run, so a
|
||||
# supervisor that revives the OLD jar a few seconds late is caught too. Every check above (healthz
|
||||
# 200, jar id, the fresh 'listening' line) is satisfied by EITHER daemon if two are alive — this is
|
||||
# the only one that can tell.
|
||||
assert_single_daemon "$(running_pid)"
|
||||
|
||||
say "result"
|
||||
ok "pid $NEW_PID, jar $(jar_id)"
|
||||
if [ "$REDEPLOY_ERROR_COUNT" -eq 0 ]; then
|
||||
|
||||
@@ -25,6 +25,78 @@ classify_fixture() {
|
||||
classify_amqp_connection_errors "$TMP/$name"
|
||||
}
|
||||
|
||||
|
||||
# fleetd #492 — supervisor detection. Detect_supervisor() reads launchd_loaded/systemd_loaded, so
|
||||
# each test overrides BOTH pairs (installed + loaded) explicitly, rather than relying on either
|
||||
# being naturally absent: this machine may itself be running a real fleetd under launchd right now
|
||||
# (see CLAUDE.md/MEMORY.md — launchd supervision has been live here since 2026-08-26), so leaving
|
||||
# launchd_loaded unmocked in a "systemd only" test would silently read this host's own live state
|
||||
# instead of the fixture.
|
||||
test_detect_supervisor_launchd_only() {
|
||||
launchd_installed() { return 0; }
|
||||
launchd_loaded() { return 0; }
|
||||
systemd_installed() { return 1; }
|
||||
systemd_loaded() { return 1; }
|
||||
assert_equals "launchd" "$(detect_supervisor)" "launchd-only detection"
|
||||
}
|
||||
|
||||
test_detect_supervisor_systemd_only() {
|
||||
launchd_installed() { return 1; }
|
||||
launchd_loaded() { return 1; }
|
||||
systemd_installed() { return 0; }
|
||||
systemd_loaded() { return 0; }
|
||||
assert_equals "systemd" "$(detect_supervisor)" "systemd-only detection"
|
||||
}
|
||||
|
||||
test_detect_supervisor_none() {
|
||||
launchd_installed() { return 1; }
|
||||
launchd_loaded() { return 1; }
|
||||
systemd_installed() { return 1; }
|
||||
systemd_loaded() { return 1; }
|
||||
assert_equals "none" "$(detect_supervisor)" "unsupervised detection"
|
||||
}
|
||||
|
||||
# The heart of the ticket: a supervisor this script cannot drive must refuse, never fall through to
|
||||
# `kill`. require_drivable_supervisor die()s, so it is invoked inside a command substitution — that
|
||||
# forks a subshell, so its exit() only ends the subshell and this test script keeps running under
|
||||
# `set -e`.
|
||||
test_require_drivable_supervisor_refuses_ambiguous() {
|
||||
local output rc=0
|
||||
output="$(require_drivable_supervisor "ambiguous" 2>&1)" || rc=$?
|
||||
[ "$rc" -ne 0 ] || fail "require_drivable_supervisor accepted an ambiguous (undrivable) supervisor"
|
||||
printf '%s' "$output" | grep -qF "$LAUNCHD_LABEL" \
|
||||
|| fail "refusal message does not name the launchd label it found"
|
||||
printf '%s' "$output" | grep -qF "$SYSTEMD_UNIT" \
|
||||
|| fail "refusal message does not name the systemd unit it found"
|
||||
}
|
||||
|
||||
test_require_drivable_supervisor_accepts_known_kinds() {
|
||||
require_drivable_supervisor "launchd" || fail "refused a drivable launchd supervisor"
|
||||
require_drivable_supervisor "systemd" || fail "refused a drivable systemd supervisor"
|
||||
require_drivable_supervisor "none" || fail "refused the unsupervised case"
|
||||
}
|
||||
|
||||
# fleetd #492 — the one-daemon check. Two live pids is the exact symptom a racing supervisor
|
||||
# produces, and none of the other post-restart checks (healthz, jar id, the fresh log line) can see
|
||||
# it because either daemon alone satisfies them.
|
||||
test_count_daemon_pids() {
|
||||
assert_equals 0 "$(count_daemon_pids "")" "count of an empty pid list"
|
||||
assert_equals 1 "$(count_daemon_pids "4242")" "count of a single pid"
|
||||
assert_equals 2 "$(count_daemon_pids "$(printf '4242\n4343\n')")" "count of two pids"
|
||||
}
|
||||
|
||||
test_assert_single_daemon_accepts_one_pid() {
|
||||
assert_single_daemon "4242" || fail "assert_single_daemon rejected a single running pid"
|
||||
}
|
||||
|
||||
test_assert_single_daemon_rejects_two_pids() {
|
||||
local output rc=0
|
||||
output="$(assert_single_daemon "$(printf '4242\n4343\n')" 2>&1)" || rc=$?
|
||||
[ "$rc" -ne 0 ] || fail "assert_single_daemon accepted two simultaneously running pids"
|
||||
printf '%s' "$output" | grep -qF '4242' || fail "refusal message does not list the pids it found"
|
||||
printf '%s' "$output" | grep -qF '4343' || fail "refusal message does not list the pids it found"
|
||||
}
|
||||
|
||||
test_no_errors() {
|
||||
cat > "$TMP/no-errors.log" <<'LOG'
|
||||
2026-09-05 12:00:00 INFO fleetd listening
|
||||
@@ -224,6 +296,14 @@ test_unattributable_quiet_mutation_is_caught() {
|
||||
printf 'Unattributable mutation: FAIL: cross-unattributable recovered: expected 0, got 2\n'
|
||||
}
|
||||
|
||||
test_detect_supervisor_launchd_only
|
||||
test_detect_supervisor_systemd_only
|
||||
test_detect_supervisor_none
|
||||
test_require_drivable_supervisor_refuses_ambiguous
|
||||
test_require_drivable_supervisor_accepts_known_kinds
|
||||
test_count_daemon_pids
|
||||
test_assert_single_daemon_accepts_one_pid
|
||||
test_assert_single_daemon_rejects_two_pids
|
||||
test_no_errors
|
||||
test_recovery_patterns_match_source
|
||||
test_attributed_recovered_connection_error
|
||||
|
||||
Reference in New Issue
Block a user