Compare commits

...

7 Commits

Author SHA1 Message Date
Dai Ha a3eeace447 handover skill: BOOTSTRAP_NEVER_SENT, and relaunchReadySeconds bounds three waits
CI / shell-tests (push) Failing after 7s
CI / contract (push) Successful in 1m2s
CI / build (push) Failing after 1m55s
Co-Authored-By: Claude <noreply@anthropic.com>
2026-10-06 19:28:21 +02:00
Dai Ha b696c31756 fleetd #796: retry bootstrap delivery to a fresh lead pane
CI / shell-tests (push) Failing after 6s
CI / contract (push) Successful in 1m5s
CI / build (push) Failing after 2m17s
A roll killed the old pane, launched the new one, then lost bootstrapText to a
herdr agent_not_ready refusal. The successor woke with no handover and no way to
tell a failed roll from a cold start.

sendBootstrapWithRetry retries only that refusal, bounded by
relaunchReadySeconds. A persistent refusal now ends the roll with the new
BOOTSTRAP_NEVER_SENT outcome, so fleet_handover{action:"status"} can report it.

Verified at 027d413 in a throwaway worktree: 2210 tests, 0 failures, 0 errors,
clean install.

Co-Authored-By: Claude <noreply@anthropic.com>
2026-10-06 19:27:36 +02:00
Dai Ha 96ebae28a6 reviewer skill: send fleet_reply before writing the reasoning out
Five reviewer turns were lost in one session. Each wrote a long, correct
analysis to its own pane and ended the turn with no fleet_reply, so the
bridge scraped the pane and the lead received a clipped fragment.

The briefs are half the cause: each carried a five-item checklist of things
to hunt alongside a ~90-word capped output format, which reads as two
contradictory output contracts. The skill now says the checklist is where to
look, not the shape of the answer, and that the reply goes out as soon as
the answer is known.

Measured: tickets task-7da785-29 and -30 resolved at 18:28 via the
turn-completion fallback, 1233 and 1403 chars scraped.
2026-10-06 19:01:38 +02:00
Dai Ha b41aa663f7 Bridge block: invariant 6 — an operator outranks the confirm rule
Invariant 6 as first written told every receiver to confirm, with no
mention of who may override it. A peer cannot impose a rule on another
operator's session.

Measured: the trinotes pane (observer, work config dir) answered the
communication test and then said its operator's standing rule is not to
answer fleet messages, that a peer cannot change that rule, and that it
would ask its operator before following mine. That is the correct reading
and the rule now says so.

Also records the symmetric error: a sender must not read silence as
agreement or as a dead session.

wiki/7-Use-Cases.md synced byte-identical.
2026-10-06 18:48:12 +02:00
Dai Ha 027d413ce9 fleetd #796: document bootstrap retry readiness bound
CI / shell-tests (pull_request) Failing after 7s
CI / contract (pull_request) Successful in 51s
CI / build (pull_request) Failing after 1m59s
2026-10-06 18:26:01 +02:00
Dai Ha f544906621 Bridge block: a receiver confirms what it receives
Adds invariant 6 to the canonical block. Nothing in it said a received
message must be answered, so a sender could not tell a handled message
from one that never arrived.

Measured today: the anki pane (observer) sent to vms, which reads
deliverable:false because it has not contacted the daemon since the 08:05
restart. The send was accepted, held 83s, and failed with the body 'nst' --
three characters scraped off the vms screen. The sender read that as an
inconclusive result and asked the lead what happened. fleetd #757 covers
the accept-time refusal; this rule covers the half the protocol owns.

Operator's words: 'when a message sent, at least the receiver should
confirm unless explicit told to not reply'.

wiki/7-Use-Cases.md template synced byte-identical (wiki 3ad606f).
2026-10-06 18:18:10 +02:00
Dai Ha c88f01ecb8 fleetd #796: retry bootstrap delivery after agent readiness refusal
CI / shell-tests (pull_request) Failing after 8s
CI / contract (pull_request) Successful in 54s
CI / build (pull_request) Failing after 2m4s
2026-10-06 18:16:47 +02:00
6 changed files with 138 additions and 8 deletions
+4 -2
View File
@@ -189,8 +189,9 @@ If the first number has grown past 20, somebody has rolled under the restart pat
should be replaced with what they measured.
**Three separate timeouts bound a roll.** `leadRollover.relaunchReadySeconds` (default 45,
`FleetConfig.java:1483`) bounds **each** of two waits that run after the relaunch, so the worst case
there is about twice that number, not 45 seconds in total. A third bound gives your old pane 10
`FleetConfig.java:1483`) bounds **each** of three waits that run after the relaunch, so the worst
case there is about three times that number, not 45 seconds in total. The third wait retries
`bootstrapText` while herdr answers `agent_not_ready`. A separate bound gives your old pane 10
seconds to die (`LeadRollover.PANE_DEATH_TIMEOUT_SECONDS`).
`fleet_handover{action: "status", token}` answers with one of these:
@@ -204,6 +205,7 @@ seconds to die (`LeadRollover.PANE_DEATH_TIMEOUT_SECONDS`).
| `RELAUNCH_FAILED` | launching the fresh pane failed |
| `RELAUNCH_NEVER_READY` | the fresh pane never became ready within `relaunchReadySeconds` |
| `RELAUNCH_NOT_RECOGNISED` | the fresh terminal never resolved as a lead |
| `BOOTSTRAP_NEVER_SENT` | the fresh pane was ready, but herdr refused `bootstrapText` with `agent_not_ready` for the whole bound, so the successor never learned where the handover file is |
| `FAILED` | the roll threw; `runRollover`'s catch records this rather than leaving it stuck |
Only `TURN_NEVER_SETTLED` guarantees your context is intact. The other failures can leave you
+9
View File
@@ -37,6 +37,15 @@ confident-but-wrong finding. Anything you could settle by reading more code is y
## 4. The finding — what goes in `fleet_reply`
**Call `fleet_reply` as soon as you know your answer, before you write the reasoning out.** The
four lines below are the whole deliverable, and they are short on purpose. Analysis you type into
your terminal reaches nobody: when a turn ends with no `fleet_reply`, the bridge scrapes the pane
and the lead receives a clipped fragment instead of a finding. A long, correct analysis and no
`fleet_reply` is a failed turn, and it is the most common way this role fails.
If the delegation also handed you a list of things to check, that list is where to *look*. It is
not the shape of the answer. Work the list, then still send these four lines.
Report the **single most important** real issue in the scope, in these four lines, under
~90 words:
+12
View File
@@ -95,6 +95,18 @@ and the sender silently receives nothing. Fail toward the recoverable error.
5. **Never move a fleet session, pane or peer except through the bridge.** The bridge owns policy;
the multiplexer owns PTYs. Any route that changes fleet state without the bridge's checks
bypasses every rule above — the `herdr` CLI and its socket are the usual example.
6. **Confirm what you receive.** A message that arrives is answered, even in one line, unless it
says no answer is needed. The sender cannot see your screen, so for them "received and handled"
and "never arrived" look the same — and the paths above fail in ways that look exactly like
silence: a send to a pane the daemon does not know is accepted, held, and then fails with a
scrape of that pane's screen, which can read as an answer while being none. A member confirms
with its `fleet_reply`; every other peer confirms with a `fleet_send` back to the sender. If you
cannot do the thing asked, say that — a refusal is a confirmation. **Never read a failed
ticket's body as a reply.** Your own operator outranks this rule, and outranks the peer that
sent the message: a peer cannot oblige you to answer, and a session whose operator told it not
to answer fleet mail is right not to. Where you can, say that much and nothing more. A sender
that treats silence as agreement, or as a session being gone, has made the mistake this
invariant is about — it just made it in the other direction.
### Primary (lead) — run this on every task, in order
@@ -1466,7 +1466,7 @@ public record FleetConfig(
* CALLING lead's own turn to end (its pane to report {@code IDLE} or
* {@code DONE}) before ending that pane's process at all. See the
* paragraph above.
* @param relaunchReadySeconds default 45 — bound on EACH of two separate waits that run after
* @param relaunchReadySeconds default 45 — bound on EACH of three separate waits that run after
* the old lead's pane has been torn down and a fresh one launched: first,
* for the fresh pane itself to reach a real turn boundary ({@code IDLE} or
* {@code DONE}, never merely {@code BLOCKED}) — the safety gate, since
@@ -1479,8 +1479,9 @@ public record FleetConfig(
* 10s live), so a budget has to clear more than one scan interval to leave
* any real margin for the CLI's own boot time; 20 was rejected for exactly
* that reason — at a 10s scan interval it only buys two scans. 45 buys
* roughly four. Only a timeout on the FIRST wait (the pane never becomes
* ready) withholds {@code bootstrapText}.
* roughly four; third, to retry sending {@code bootstrapText} while herdr
* reports {@code agent_not_ready}. A timeout on the first wait or the third
* one withholds {@code bootstrapText}.
* @param bootstrapText default a sentence naming the RESOLVED handover path — sent to the
* fresh lead's pane once it reaches a real turn boundary after relaunch,
* telling the fresh session where to read the handover and carry on. Left
@@ -232,7 +232,7 @@ public final class LeadRollover {
/**
* What is known about one token, right now — the answer {@link #status} gives. Distinguishes
* five terminal outcomes an approved roll can finish with, one in-flight outcome for a roll
* terminal outcomes an approved roll can finish with, one in-flight outcome for a roll
* that has been approved but has not finished yet, and two answers for a token that names no
* active work at all: still pending confirmation, or nothing known about this token at all.
*/
@@ -256,7 +256,8 @@ public final class LeadRollover {
* that is, in fact, actively running. This is not sticky: the deferred continuation
* overwrites this same entry with a terminal state ({@link #ROLLED}, {@link
* #TURN_NEVER_SETTLED}, {@link #OLD_PANE_NEVER_DIED}, {@link #RELAUNCH_FAILED}, {@link
* #RELAUNCH_NEVER_READY}, {@link #RELAUNCH_NOT_RECOGNISED}, or {@link #FAILED}) once it
* #RELAUNCH_NEVER_READY}, {@link #RELAUNCH_NOT_RECOGNISED}, {@link #BOOTSTRAP_NEVER_SENT},
* or {@link #FAILED}) once it
* finishes — including by throwing, which {@link #runRollover}'s catch turns into {@link
* #FAILED} instead of leaving this entry stuck forever.
*/
@@ -305,6 +306,12 @@ public final class LeadRollover {
* an operator should check why the tab was not recognised.
*/
RELAUNCH_NOT_RECOGNISED,
/**
* A fresh lead was ready, but herdr kept reporting {@code agent_not_ready} while this class
* retried {@code bootstrapText} for {@code relaunchReadySeconds}. The fresh session did not
* receive its handover instruction.
*/
BOOTSTRAP_NEVER_SENT,
/**
* The deferred continuation threw a {@link RuntimeException} and the continuation thread
* died with it. Without this state, that throw would leave {@link #outcomes} holding {@link
@@ -731,7 +738,19 @@ public final class LeadRollover {
IdentityResult identityResult = waitUntilRecognisedAsLead(newAgent.terminalId(),
cfg.relaunchReadySeconds());
agents.send(newAgent.terminalId(), cfg.bootstrapTextFor(p.handoverPath()));
BootstrapResult bootstrapResult = sendBootstrapWithRetry(newAgent.terminalId(),
cfg.bootstrapTextFor(p.handoverPath()), cfg.relaunchReadySeconds());
if (!bootstrapResult.sent()) {
log.warn("lead-rollover: bootstrapText was never sent to fresh terminal {} for lead '{}' "
+ "after agent_not_ready persisted for {}ms (token={}, configured={}s)",
newAgent.terminalId(), leadName, bootstrapResult.elapsedMillis(), p.token(),
cfg.relaunchReadySeconds());
outcomes.put(p.token(), new RollStatus(RollState.BOOTSTRAP_NEVER_SENT,
"fresh terminal " + newAgent.terminalId() + " kept rejecting bootstrapText with "
+ "agent_not_ready for relaunchReadySeconds=" + cfg.relaunchReadySeconds()
+ "s (measured elapsed=" + bootstrapResult.elapsedMillis() + "ms)"));
return;
}
if (!identityResult.ready()) {
log.warn("lead-rollover: fresh terminal {} for lead '{}' is alive and bootstrapped, but "
+ "was never recognised as a live lead — an operator should check why "
@@ -755,6 +774,30 @@ public final class LeadRollover {
+ newAgent.terminalId()));
}
/**
* Sends {@code bootstrapText}, retrying only the transient herdr {@code agent_not_ready} refusal
* until {@code readySeconds} elapses. All other failures propagate to {@link #runRollover}.
*/
private BootstrapResult sendBootstrapWithRetry(String terminal, String bootstrapText, int readySeconds) {
long startMillis = nowMillis.getAsLong();
long deadline = startMillis + TimeUnit.SECONDS.toMillis(readySeconds);
while (nowMillis.getAsLong() < deadline) {
try {
agents.send(terminal, bootstrapText);
return new BootstrapResult(true, nowMillis.getAsLong() - startMillis);
} catch (HerdrException e) {
if (!"agent_not_ready".equals(e.code())) {
throw e;
}
pollSleeper.run();
}
}
return new BootstrapResult(false, nowMillis.getAsLong() - startMillis);
}
/** The result of {@link #sendBootstrapWithRetry}. */
private record BootstrapResult(boolean sent, long elapsedMillis) {}
/** Attempts {@link #captureAgentWithRetry} makes before letting the failure propagate. */
static final int CAPTURE_RETRIES = 3;
@@ -2,10 +2,12 @@ package dev.ltms.fleet.lead;
import ch.qos.logback.classic.Level;
import ch.qos.logback.classic.spi.ILoggingEvent;
import com.fasterxml.jackson.databind.JsonNode;
import dev.ltms.fleet.config.FleetConfig;
import dev.ltms.fleet.herdr.AgentControl;
import dev.ltms.fleet.herdr.FakeHerdr;
import dev.ltms.fleet.herdr.HerdrClient;
import dev.ltms.fleet.herdr.HerdrException;
import dev.ltms.fleet.herdr.WorkspaceControl;
import dev.ltms.fleet.testing.CapturedLog;
import org.junit.jupiter.api.DisplayName;
@@ -20,6 +22,7 @@ import java.util.HashMap;
import java.util.LinkedHashMap;
import java.util.List;
import java.util.Map;
import java.util.concurrent.atomic.AtomicInteger;
import java.util.concurrent.atomic.AtomicLong;
import java.util.function.Function;
import java.util.function.LongSupplier;
@@ -1323,6 +1326,66 @@ class LeadRolloverTest {
+ "so an operator reading status() has something to act on: " + status.detail());
}
@Test
@DisplayName("[BOOTSTRAP 1] one agent_not_ready bootstrap refusal is retried and sends the "
+ "handover instruction exactly once")
void transientAgentNotReadyRetriesBootstrapAndRolls() throws IOException {
FakeHerdr fake = herdrReadyForAFullRoll();
AtomicInteger refusedSends = new AtomicInteger();
HerdrClient transientRefusal = new HerdrClient() {
@Override
public JsonNode call(String method, Object params) {
if ("agent.prompt".equals(method) && refusedSends.getAndIncrement() == 0) {
throw new HerdrException("herdr error [agent_not_ready]: agent.prompt failed",
"agent_not_ready", null);
}
return fake.call(method, params);
}
@Override
public void close() {
}
};
Path handover = writeHandover("handover contents");
AtomicLong clock = new AtomicLong(1_000);
AgentControl agents = new AgentControl(transientRefusal);
WorkspaceControl spaces = new WorkspaceControl(transientRefusal);
LeadRollover rollover = new LeadRollover(agents, spaces,
new LeadLauncher(agents, spaces, fleetConfigWithRelaunchableLead()),
() -> cfg(handover.toString()), _ -> null, _ -> LEAD_NAME,
() -> Map.of(NEW_TERMINAL, LEAD_NAME), fixedClock(clock), () -> { }, Runnable::run);
LeadRollover.PendingRollover pending = rollover.open(LEAD, "context is full");
assertTrue(rollover.confirm(LEAD, pending.token(), true).accepted());
assertEquals(2, refusedSends.get(), "one agent_not_ready refusal must be followed by one retry");
assertEquals(1, promptCallCount(fake), "only the successful retry reaches herdr delivery");
assertEquals(LeadRollover.RollState.ROLLED, rollover.status(pending.token()).state());
}
@Test
@DisplayName("[BOOTSTRAP 2] persistent agent_not_ready records BOOTSTRAP_NEVER_SENT and "
+ "releases the single-flight claim")
void persistentAgentNotReadyRecordsBootstrapNeverSentAndReleasesClaim() throws IOException {
FakeHerdr fake = herdrReadyForAFullRoll();
fake.agentSendFailsWith("agent_not_ready");
Path handover = writeHandover("handover contents");
AtomicLong clock = new AtomicLong(1_000);
LeadRollover rollover = newRolloverForAFullRoll(fake, cfg(handover.toString()),
() -> clock.addAndGet(1_000), Map.of(NEW_TERMINAL, LEAD_NAME));
LeadRollover.PendingRollover pending = rollover.open(LEAD, "context is full");
assertTrue(rollover.confirm(LEAD, pending.token(), true).accepted());
LeadRollover.RollStatus status = rollover.status(pending.token());
assertEquals(LeadRollover.RollState.BOOTSTRAP_NEVER_SENT, status.state());
assertNotEquals(LeadRollover.RollState.ROLLED, status.state());
assertNotEquals(LeadRollover.RollState.FAILED, status.state());
LeadRollover.PendingRollover retry = rollover.open(LEAD, "retry after bootstrap timeout");
assertTrue(rollover.confirm(LEAD, retry.token(), true).accepted(),
"the terminal outcome must release the single-flight claim");
}
// ---- fleetd #726 unit 3: confirm() single-flights one roll at a time per lead ---------------
@Test