turnSettleSeconds=20 is shorter than the goodbye turn the handover skill tells the lead to write, so a correct roll refuses itself — measured 2026-10-02 #651

Closed
opened 2026-10-03 14:37:56 +02:00 by ltms · 3 comments
Owner

What happened

A lead ran the full fleet_handover procedure correctly and the roll did not happen. The two
log lines, from fleetd/fleetd.out:

05:09:24.486 INFO  lead-rollover: confirmed token=851bda9f-874f-42ea-a748-60d643210059
                   lead=term_65c806b784dc28 — roll scheduled once the calling turn ends
05:09:44.880 WARN  lead-rollover: pane term_65c806b784dc28 did not reach a turn boundary
                   (IDLE or DONE) after confirm() — refusing to send /clear at all; the calling
                   lead's own turn is still live and clearing it now would destroy live context
                   (token=851bda9f-874f-42ea-a748-60d643210059, configured=20s elapsed=20394ms)

rolled stayed at 8 across the attempt. Control: grep -c "lead-rollover:" fleetd/fleetd.out went
32 → 35, so three lines were written and the pattern is sound.

The refusal itself is the correct branch — clearing a live turn would destroy context. The
defect is that a lead following the documented procedure lands in that branch.

The conflict

turnSettleSeconds defaults to 20 (config/FleetConfig.java:1406, :1428) and is not set
in this host's fleetd.yaml:

$ grep -c "turnSettleSeconds" fleetd/fleetd.yaml   # 0
$ grep -c "clearSettleSeconds" fleetd/fleetd.yaml  # 0
$ grep -c "requireOperatorConfirm" fleetd/fleetd.yaml  # 1  <- positive control

So the deferred roll waits 20 seconds for the calling pane to reach IDLE or DONE, then gives up and
writes TURN_NEVER_SETTLED (lead/LeadRollover.java:550, :560).

Meanwhile the handover skill tells the lead, about confirm returning accepted:

accepted does not mean your pane has been cleared. It means every gate passed and the roll
is scheduled to run once your current turn ends. Say your goodbye in the same turn — you will
not get another one.

A lead that obeys that instruction writes its closing message inside the same turn. Here that turn
ran past 20 seconds (elapsed=20394ms) and the roll refused. The instruction and the timeout
contradict each other.
The longer and more useful the goodbye, the more likely the roll fails.

I have not measured the distribution of lead closing-turn lengths — this is one observation. What I
can say is that 20s is below what one ordinary closing message took, and nothing in the skill warns
the lead to keep it short.

Why it is easy to miss

confirm already returned accepted, so the caller believes it succeeded. The skill documents
that a late refusal is logged only, with no caller left to tell. That is honest, but it means this
failure is invisible unless somebody greps the daemon log afterwards. The lead carries on with a
full context believing it has been replaced.

Options (not a decision — recording them)

  1. Raise turnSettleSeconds in fleetd.yaml. Cheapest. Needs a number with a reason behind it,
    not a guess, and the key is read at boot so it needs a redeploy.
  2. Do not bound this wait the way the others are bounded. The /clear wait
    (clearSettleSeconds) has a real reason to time out: the pane may never pick the command up. The
    calling turn will end on its own, so waiting longer costs nothing but a parked virtual thread.
    A much larger bound, or one derived from the turn rather than a constant, may be the right shape.
  3. Tell the lead the truth in the skill — that the goodbye has a deadline — and keep the
    timeout. This is the option that needs no code, and the one that makes the feature worse to use.

Related: #486 (LeadRollover's settle poll has no bound of its own), #489 and #480 (the rollover
feature), #491.

## What happened A lead ran the full `fleet_handover` procedure correctly and the roll **did not happen**. The two log lines, from `fleetd/fleetd.out`: ``` 05:09:24.486 INFO lead-rollover: confirmed token=851bda9f-874f-42ea-a748-60d643210059 lead=term_65c806b784dc28 — roll scheduled once the calling turn ends 05:09:44.880 WARN lead-rollover: pane term_65c806b784dc28 did not reach a turn boundary (IDLE or DONE) after confirm() — refusing to send /clear at all; the calling lead's own turn is still live and clearing it now would destroy live context (token=851bda9f-874f-42ea-a748-60d643210059, configured=20s elapsed=20394ms) ``` `rolled` stayed at 8 across the attempt. Control: `grep -c "lead-rollover:" fleetd/fleetd.out` went 32 → 35, so three lines were written and the pattern is sound. The refusal itself is the **correct** branch — clearing a live turn would destroy context. The defect is that a lead following the documented procedure lands in that branch. ## The conflict `turnSettleSeconds` defaults to **20** (`config/FleetConfig.java:1406`, `:1428`) and is **not set** in this host's `fleetd.yaml`: ``` $ grep -c "turnSettleSeconds" fleetd/fleetd.yaml # 0 $ grep -c "clearSettleSeconds" fleetd/fleetd.yaml # 0 $ grep -c "requireOperatorConfirm" fleetd/fleetd.yaml # 1 <- positive control ``` So the deferred roll waits 20 seconds for the calling pane to reach IDLE or DONE, then gives up and writes `TURN_NEVER_SETTLED` (`lead/LeadRollover.java:550`, `:560`). Meanwhile the `handover` skill tells the lead, about `confirm` returning `accepted`: > **`accepted` does not mean your pane has been cleared.** It means every gate passed and the roll > is scheduled to run once your current turn ends. **Say your goodbye in the same turn — you will > not get another one.** A lead that obeys that instruction writes its closing message inside the same turn. Here that turn ran past 20 seconds (`elapsed=20394ms`) and the roll refused. **The instruction and the timeout contradict each other.** The longer and more useful the goodbye, the more likely the roll fails. I have not measured the distribution of lead closing-turn lengths — this is one observation. What I can say is that 20s is below what one ordinary closing message took, and nothing in the skill warns the lead to keep it short. ## Why it is easy to miss `confirm` already returned `accepted`, so the caller believes it succeeded. The skill documents that a late refusal is logged only, with no caller left to tell. That is honest, but it means this failure is invisible unless somebody greps the daemon log afterwards. The lead carries on with a full context believing it has been replaced. ## Options (not a decision — recording them) 1. **Raise `turnSettleSeconds`** in `fleetd.yaml`. Cheapest. Needs a number with a reason behind it, not a guess, and the key is read at boot so it needs a redeploy. 2. **Do not bound this wait the way the others are bounded.** The `/clear` wait (`clearSettleSeconds`) has a real reason to time out: the pane may never pick the command up. The *calling turn* will end on its own, so waiting longer costs nothing but a parked virtual thread. A much larger bound, or one derived from the turn rather than a constant, may be the right shape. 3. **Tell the lead the truth in the skill** — that the goodbye has a deadline — and keep the timeout. This is the option that needs no code, and the one that makes the feature worse to use. Related: #486 (LeadRollover's settle poll has no bound of its own), #489 and #480 (the rollover feature), #491.
Author
Owner

Decision (lead). I read the code myself rather than ruling off the options list. Two of the ticket's premises turned out to be wrong, and both make this cheaper than it looked.

Premise 1 that was wrong: this needs a redeploy

leadRollover is hot, not boot-read. ConfigRef.java:188-196 names it explicitly:

placement, memberCredentials, memberLoginShell, models and leadRollover are hot and correctly absent — all five are read live off config.get() … for leadRollover, the next open()/confirm() call

And LeadRollover.confirm() does read it live:

public RollDecision confirm(String callerTerminal, String token, boolean operatorConfirmed) {
    FleetConfig.LeadRollover cfg = configSupplier.get();

So option 1 costs a config reload, not a redeploy. That was the main thing making it look expensive.

(One detail: cfg is captured at confirm() and passed into the deferred continuation, so a reload mid-roll does not change a roll already scheduled. The next one picks it up. That is correct behaviour, not a bug.)

Premise 2 that was wrong: the failure is invisible

The ticket says the refusal "is invisible unless somebody greps the daemon log". It is not. fleet_handover{action:"status", token} already exists — FleetMcp.java:1359 dispatches it to LeadRollover.status(String token) at :619, which reads the same outcomes map the refusal writes TURN_NEVER_SETTLED into at :560, carrying the measured elapsed.

The real gap is that nobody is told to look. And there is a clean signal that costs nothing: if the roll worked, the lead has been cleared and is not there to wonder. So a lead that is still alive after its goodbye turn already knows the roll did not happen. That is a reliable self-check, and no code is needed for it.

The ruling

Not option 3. Telling the lead to keep its goodbye short makes the feature worse at the one moment it matters, and the deadline would still be invisible while the lead writes.

1. Raise turnSettleSeconds to 300, in both the live config and the code default.

The number is not a guess, and it is not derived from the single 20394ms observation — one data point cannot set a bound. It comes from the only other constant in this codebase that answers "how long may a lead legitimately be mid-turn?", leadHeartbeat.idleAfterSeconds, default 300, whose own javadoc gives the reason:

a lead that just finished a turn sits momentarily idle … 5 minutes absorbs normal pauses without stalling

The same judgement applies here, and the two should not disagree by a factor of fifteen. turnSettleSeconds=20 is far below the only existing estimate of a lead's natural rhythm.

2. Keep it bounded. I considered the ticket's option 2 (do not bound this wait at all) and rejected it. The argument for it is good — the first wait has sent nothing, so waiting costs only a parked thread, unlike clearSettleSeconds where /clear is already out. But a lead whose turn genuinely never ends would park a continuation and hold a pending token forever, and that failure is harder to see than a logged refusal. 300s is far above any real goodbye and still terminates.

3. Fix a documentation defect I found while reading this. The config javadoc describes both waits as waiting for the pane "to report an injectable state again" (FleetConfig.java:1403, :1426). The code does not do that:

if (status == AgentStatus.IDLE || status == AgentStatus.DONE) {

injectable() also accepts BLOCKED, which is a paused live turn — exactly the state where sending /clear would destroy context. The code is right and the doc is wrong, and the doc is the dangerous half: it invites a future reader to "simplify" the check to injectable() and reintroduce the bug the strict check prevents.

4. Add the survival check to the handover skill. After the goodbye turn, a lead that is still running must treat that as evidence the roll refused: read fleet_handover{action:"status", token}, and if it says TURN_NEVER_SETTLED, open a fresh request and retry. This replaces the "keep it short" warning with something that costs the lead nothing and cannot be forgotten at the wrong moment.

Acceptance criteria

  1. With turnSettleSeconds unset, the resolved value is 300. With it set to a positive number, that number is used. With it set to 0 or negative, it falls back to 300 — the existing <= 0 guard at FleetConfig.java:1447 already does this; assert it still holds.
  2. A roll whose calling turn reaches IDLE or DONE within the budget still proceeds to /clear. A roll whose pane stays WORKING past the budget still writes TURN_NEVER_SETTLED and sends no /clear. Both directions — the refusal branch is correct and must not be removed.
  3. A pane reporting BLOCKED is not treated as settled. Assert this directly; it is the case the javadoc currently mis-describes.
  4. No javadoc in FleetConfig or LeadRollover still describes either wait as waiting for an "injectable" state when the code requires IDLE or DONE.

Note on the live config

I will set turnSettleSeconds: 300 in this host's fleetd.yaml myself — it is gitignored and lead-only, so a worker cannot see or change it. Whoever takes the code half should not expect to find it in the repo. Flagging it because a config-dependent change that workers cannot see has shipped green and inert here before.

**Decision (lead).** I read the code myself rather than ruling off the options list. Two of the ticket's premises turned out to be wrong, and both make this cheaper than it looked. ## Premise 1 that was wrong: this needs a redeploy `leadRollover` is **hot**, not boot-read. `ConfigRef.java:188-196` names it explicitly: > `placement`, `memberCredentials`, `memberLoginShell`, `models` and `leadRollover` are **hot** and correctly absent — all five are read live off `config.get()` … for `leadRollover`, the next `open()`/`confirm()` call And `LeadRollover.confirm()` does read it live: ```java public RollDecision confirm(String callerTerminal, String token, boolean operatorConfirmed) { FleetConfig.LeadRollover cfg = configSupplier.get(); ``` So option 1 costs a config reload, not a redeploy. That was the main thing making it look expensive. (One detail: `cfg` is captured at `confirm()` and passed into the deferred continuation, so a reload mid-roll does not change a roll already scheduled. The next one picks it up. That is correct behaviour, not a bug.) ## Premise 2 that was wrong: the failure is invisible The ticket says the refusal "is invisible unless somebody greps the daemon log". It is not. `fleet_handover{action:"status", token}` already exists — `FleetMcp.java:1359` dispatches it to `LeadRollover.status(String token)` at `:619`, which reads the same `outcomes` map the refusal writes `TURN_NEVER_SETTLED` into at `:560`, carrying the measured `elapsed`. The real gap is that **nobody is told to look**. And there is a clean signal that costs nothing: if the roll worked, the lead has been cleared and is not there to wonder. **So a lead that is still alive after its goodbye turn already knows the roll did not happen.** That is a reliable self-check, and no code is needed for it. ## The ruling Not option 3. Telling the lead to keep its goodbye short makes the feature worse at the one moment it matters, and the deadline would still be invisible while the lead writes. **1. Raise `turnSettleSeconds` to 300, in both the live config and the code default.** The number is not a guess, and it is not derived from the single 20394ms observation — one data point cannot set a bound. It comes from the only other constant in this codebase that answers "how long may a lead legitimately be mid-turn?", `leadHeartbeat.idleAfterSeconds`, default 300, whose own javadoc gives the reason: > a lead that just finished a turn sits momentarily idle … 5 minutes absorbs normal pauses without stalling The same judgement applies here, and the two should not disagree by a factor of fifteen. `turnSettleSeconds=20` is far below the only existing estimate of a lead's natural rhythm. **2. Keep it bounded.** I considered the ticket's option 2 (do not bound this wait at all) and rejected it. The argument for it is good — the first wait has sent nothing, so waiting costs only a parked thread, unlike `clearSettleSeconds` where `/clear` is already out. But a lead whose turn genuinely never ends would park a continuation and hold a pending token forever, and that failure is harder to see than a logged refusal. 300s is far above any real goodbye and still terminates. **3. Fix a documentation defect I found while reading this.** The config javadoc describes both waits as waiting for the pane "to report an injectable state again" (`FleetConfig.java:1403`, `:1426`). The code does **not** do that: ```java if (status == AgentStatus.IDLE || status == AgentStatus.DONE) { ``` `injectable()` also accepts `BLOCKED`, which is a paused live turn — exactly the state where sending `/clear` would destroy context. The code is right and the doc is wrong, and the doc is the dangerous half: it invites a future reader to "simplify" the check to `injectable()` and reintroduce the bug the strict check prevents. **4. Add the survival check to the `handover` skill.** After the goodbye turn, a lead that is still running must treat that as evidence the roll refused: read `fleet_handover{action:"status", token}`, and if it says `TURN_NEVER_SETTLED`, open a fresh request and retry. This replaces the "keep it short" warning with something that costs the lead nothing and cannot be forgotten at the wrong moment. ## Acceptance criteria 1. With `turnSettleSeconds` unset, the resolved value is 300. With it set to a positive number, that number is used. With it set to 0 or negative, it falls back to 300 — the existing `<= 0` guard at `FleetConfig.java:1447` already does this; assert it still holds. 2. A roll whose calling turn reaches IDLE or DONE within the budget still proceeds to `/clear`. A roll whose pane stays `WORKING` past the budget still writes `TURN_NEVER_SETTLED` and sends no `/clear`. Both directions — the refusal branch is correct and must not be removed. 3. A pane reporting `BLOCKED` is **not** treated as settled. Assert this directly; it is the case the javadoc currently mis-describes. 4. No javadoc in `FleetConfig` or `LeadRollover` still describes either wait as waiting for an "injectable" state when the code requires IDLE or DONE. ## Note on the live config I will set `turnSettleSeconds: 300` in this host's `fleetd.yaml` myself — it is gitignored and lead-only, so a worker cannot see or change it. Whoever takes the code half should not expect to find it in the repo. Flagging it because a config-dependent change that workers cannot see has shipped green and inert here before.
Author
Owner

Live config half done, and it confirms the hot-reload claim by measurement rather than by reading the javadoc.

Set in this host's fleetd.yaml (gitignored, lead-only) at 21:27:56 CEST. The daemon picked it up on its own:

$ stat -f '%Sm' -t '%Y-%m-%d %H:%M:%S' fleetd/fleetd.yaml
2026-10-03 21:27:56

$ grep 'config reloaded' fleetd/fleetd.out | tail -1
21:28:00.593 INFO  [config-watcher] d.l.fleet.config.ConfigRef - config reloaded

Four seconds apart, and it is the last line in the log. No restart was needed and none happened — the daemon is still pid 42543, up since 20:05:53.

The important detail is what that line does not say. When a reload touches a key the running daemon cannot apply, ConfigRef names it. An earlier line in this same log shows the other shape:

19:04:49.930 INFO  [config-watcher] … config reloaded; these changes need a restart to take effect:
             profiles.local/local-direct/opus/sonnet launch settings (model, baseUrl, argv, env, …)
             — the launcher holds a startup snapshot

My reload carries no such clause, so turnSettleSeconds: 300 is live now. That is the empirical version of the ConfigRef.java:188-196 claim in my decision above, and it settles the ticket's "needs a redeploy" premise: it does not.

I verified the edit before letting the watcher see it: diff against a backup showed exactly one seven-line addition and nothing else, and a YAML parse returned turnSettleSeconds=300 as an Integer. I also checked the parser rejects a deliberately broken copy, so "it parsed" is not a broken check reporting success.

The config comment records the date, the measurement and how to re-measure (roll a lead, then read fleet_handover{action:"status", token}; TURN_NEVER_SETTLED means 300 is still too low), because a config file is a notebook. The code half must carry none of that — the reasoning belongs in its commit message.

Remaining for this ticket: the code default (delegated, task-8) and change 4, the handover skill's survival check. The skill is primary-side, so that one is mine.

**Live config half done, and it confirms the hot-reload claim by measurement rather than by reading the javadoc.** Set in this host's `fleetd.yaml` (gitignored, lead-only) at 21:27:56 CEST. The daemon picked it up on its own: ``` $ stat -f '%Sm' -t '%Y-%m-%d %H:%M:%S' fleetd/fleetd.yaml 2026-10-03 21:27:56 $ grep 'config reloaded' fleetd/fleetd.out | tail -1 21:28:00.593 INFO [config-watcher] d.l.fleet.config.ConfigRef - config reloaded ``` Four seconds apart, and it is the last line in the log. **No restart was needed and none happened** — the daemon is still pid 42543, up since 20:05:53. The important detail is what that line does **not** say. When a reload touches a key the running daemon cannot apply, `ConfigRef` names it. An earlier line in this same log shows the other shape: ``` 19:04:49.930 INFO [config-watcher] … config reloaded; these changes need a restart to take effect: profiles.local/local-direct/opus/sonnet launch settings (model, baseUrl, argv, env, …) — the launcher holds a startup snapshot ``` My reload carries no such clause, so `turnSettleSeconds: 300` is live now. That is the empirical version of the `ConfigRef.java:188-196` claim in my decision above, and it settles the ticket's "needs a redeploy" premise: it does not. I verified the edit before letting the watcher see it: `diff` against a backup showed exactly one seven-line addition and nothing else, and a YAML parse returned `turnSettleSeconds=300` as an Integer. I also checked the parser rejects a deliberately broken copy, so "it parsed" is not a broken check reporting success. The config comment records the date, the measurement and how to re-measure (roll a lead, then read `fleet_handover{action:"status", token}`; `TURN_NEVER_SETTLED` means 300 is still too low), because a config file is a notebook. The code half must carry none of that — the reasoning belongs in its commit message. **Remaining for this ticket:** the code default (delegated, task-8) and change 4, the `handover` skill's survival check. The skill is primary-side, so that one is mine.
Author
Owner

Closed — all four changes are in

Merged locally as 209e123 (PR #682, closed by hand as usual). git ls-remote origin main → 209e1231eafe1fdad20a71e32a96de6ff9f9c3ff, matching local HEAD.

# Change Where
1 live turnSettleSeconds: 300 fleetd/fleetd.yaml (lead, hot-reloaded)
2 code default 20 → 300 FleetConfig.java:1450
3 settle javadoc corrected FleetConfig.java:1402, :1427, :1430
4 survival check in the skill .claude/skills/handover/SKILL.md (31b3c24)

Two of this ticket's premises were wrong, and the record should say so

"the key is read at boot so it needs a redeploy" — it does not. ConfigRef.java:188-196 says all five leadRollover keys are read live off config.get(), taken at the next open()/confirm(). Measured on the live daemon: edit written 21:27:56, config reloaded logged 21:28:00.593, no "needs a restart" clause in the line, and the daemon never restarted (still pid 42543). So change 1 was live four seconds after it was saved.

"this failure is invisible unless somebody greps the daemon log" — also not so. fleet_handover{action: "status", token} returns TURN_NEVER_SETTLED to the lead itself. The real gap was that nothing told the lead to look, which is why change 4 exists: a roll that works clears you, so surviving your own goodbye is itself the signal that it refused.

Why 300 and not something derived from the turn

Option 2 (unbounded, or derived from the turn) was rejected. The bound is a safety property, not a performance one: LeadRollover cannot distinguish "turn still running" from "pane wedged", so removing the bound means a wedged pane parks a roll forever with no outcome recorded. 300 is not fitted to the single elapsed=20394ms observation — one observation cannot set a timeout. It is taken from leadHeartbeat.idleAfterSeconds, which is already 300 in this codebase with the written reason "5 minutes absorbs normal pauses without stalling". The same reasoning applies to a closing turn, so the two now agree instead of contradicting each other.

clearSettleSeconds stays at 20 on purpose. That wait has a genuine reason to expire — the pane may never pick /clear up — which is the distinction the ticket's option 2 drew, and it holds.

A documentation defect found on the way

The javadoc said both waits are for the pane to "report an injectable state again". AgentStatus.java:37 defines injectable() as IDLE || BLOCKED || DONE, but LeadRollover.waitUntilAtTurnBoundary (:724) requires IDLE || DONE. BLOCKED is a live turn merely paused — the exact state where /clear would destroy context. The wording named the wrong predicate, so a reader checking the code against the docs would have concluded the code was buggy. Fixed in change 3.

Lead verification (not the worker's run)

Three-dot diff against current origin/main: 2 files, +49/−9, nothing outside the two declared paths, no wiki/, no .mcp.json, no fleetd.yaml. Trial merge in a throwaway worktree: clean, and scripts/redeploy-fleetd.sh's JAR="$MODULE/run/fleetd.jar" from the three newer main commits survived it.

Full build in that worktree: Tests run: 1932, Failures: 0, Errors: 0, Skipped: 0, 172 surefire files after rm -rf target/surefire-reports. Positive control — all three new test names appear in TEST-…FleetConfigTest.xml, so they really ran.

Three mutations, run by me, line-anchored:

  • 300 → 20 — the two default tests went red, turnSettleSecondsUsesAnExplicitPositiveValue stayed green. The selectivity is the point: the tests pin 300, not merely "a default".
  • turnSettleSeconds = 300; unconditionally (my own complement, which the worker did not run) — turnSettleSecondsUsesAnExplicitPositiveValue went red, so that test is not vacuous and an explicit value really is honoured.
  • :724 also accept BLOCKED — LeadRolloverTest.blockedTheWholeTurnSettleWindowSendsNoClearAtAll failed, 45 run / 1 failed, no compilation error in the log. This one first looked like a bare exit 1 with no named test, because my own grep anchored on ( and that test takes no parameters. A zero match is not a finding; the log said otherwise.

git diff empty after all three reverts.

Left open on purpose

LeadRollover.java's class javadoc (~:32-44) narrates the history of an earlier #480 fix — "An earlier version of this class…", "That is wrong, because…". That is history in a code comment, against the project rule. It is pre-existing, it describes code that no longer exists rather than mischaracterising current behaviour, and it is outside this ticket's four changes. The worker flagged it instead of fixing it, which was the right call. Not filed as its own ticket; it is a comment-hygiene sweep, not a defect.

Closing.

## Closed — all four changes are in Merged locally as `209e123` (PR #682, closed by hand as usual). `git ls-remote origin main` → `209e1231eafe1fdad20a71e32a96de6ff9f9c3ff`, matching local HEAD. | # | Change | Where | |---|---|---| | 1 | live `turnSettleSeconds: 300` | `fleetd/fleetd.yaml` (lead, hot-reloaded) | | 2 | code default 20 → 300 | `FleetConfig.java:1450` | | 3 | settle javadoc corrected | `FleetConfig.java:1402`, `:1427`, `:1430` | | 4 | survival check in the skill | `.claude/skills/handover/SKILL.md` (`31b3c24`) | ## Two of this ticket's premises were wrong, and the record should say so **"the key is read at boot so it needs a redeploy"** — it does not. `ConfigRef.java:188-196` says all five `leadRollover` keys are read live off `config.get()`, taken at the next `open()`/`confirm()`. Measured on the live daemon: edit written 21:27:56, `config reloaded` logged 21:28:00.593, no "needs a restart" clause in the line, and the daemon never restarted (still pid 42543). So change 1 was live four seconds after it was saved. **"this failure is invisible unless somebody greps the daemon log"** — also not so. `fleet_handover{action: "status", token}` returns `TURN_NEVER_SETTLED` to the lead itself. The real gap was that nothing *told* the lead to look, which is why change 4 exists: a roll that works clears you, so surviving your own goodbye is itself the signal that it refused. ## Why 300 and not something derived from the turn Option 2 (unbounded, or derived from the turn) was rejected. The bound is a safety property, not a performance one: `LeadRollover` cannot distinguish "turn still running" from "pane wedged", so removing the bound means a wedged pane parks a roll forever with no outcome recorded. 300 is not fitted to the single `elapsed=20394ms` observation — one observation cannot set a timeout. It is taken from `leadHeartbeat.idleAfterSeconds`, which is already 300 in this codebase with the written reason "5 minutes absorbs normal pauses without stalling". The same reasoning applies to a closing turn, so the two now agree instead of contradicting each other. `clearSettleSeconds` stays at 20 on purpose. That wait has a genuine reason to expire — the pane may never pick `/clear` up — which is the distinction the ticket's option 2 drew, and it holds. ## A documentation defect found on the way The javadoc said both waits are for the pane to "report an injectable state again". `AgentStatus.java:37` defines `injectable()` as `IDLE || BLOCKED || DONE`, but `LeadRollover.waitUntilAtTurnBoundary` (`:724`) requires `IDLE || DONE`. `BLOCKED` is a live turn merely paused — the exact state where `/clear` would destroy context. The wording named the wrong predicate, so a reader checking the code against the docs would have concluded the code was buggy. Fixed in change 3. ## Lead verification (not the worker's run) Three-dot diff against current `origin/main`: 2 files, +49/−9, nothing outside the two declared paths, no `wiki/`, no `.mcp.json`, no `fleetd.yaml`. Trial merge in a throwaway worktree: clean, and `scripts/redeploy-fleetd.sh`'s `JAR="$MODULE/run/fleetd.jar"` from the three newer `main` commits survived it. Full build in that worktree: `Tests run: 1932, Failures: 0, Errors: 0, Skipped: 0`, 172 surefire files after `rm -rf target/surefire-reports`. Positive control — all three new test names appear in `TEST-…FleetConfigTest.xml`, so they really ran. Three mutations, run by me, line-anchored: - **300 → 20** — the two default tests went red, `turnSettleSecondsUsesAnExplicitPositiveValue` stayed green. The selectivity is the point: the tests pin *300*, not merely "a default". - **`turnSettleSeconds = 300;` unconditionally** (my own complement, which the worker did not run) — `turnSettleSecondsUsesAnExplicitPositiveValue` went red, so that test is not vacuous and an explicit value really is honoured. - **`:724` also accept `BLOCKED`** — `LeadRolloverTest.blockedTheWholeTurnSettleWindowSendsNoClearAtAll` failed, 45 run / 1 failed, no compilation error in the log. This one first *looked* like a bare exit 1 with no named test, because my own grep anchored on `(` and that test takes no parameters. A zero match is not a finding; the log said otherwise. `git diff` empty after all three reverts. ## Left open on purpose `LeadRollover.java`'s class javadoc (~`:32-44`) narrates the history of an earlier #480 fix — "An earlier version of this class…", "That is wrong, because…". That is history in a code comment, against the project rule. It is pre-existing, it describes code that no longer exists rather than mischaracterising current behaviour, and it is outside this ticket's four changes. The worker flagged it instead of fixing it, which was the right call. Not filed as its own ticket; it is a comment-hygiene sweep, not a defect. Closing.
ltms closed this issue 2026-10-03 21:44:10 +02:00
Sign in to join this conversation.
1 Participants
Notifications
Due Date
No due date set.
Dependencies

No dependencies set.

Reference: fleet/fleetd#651