A worker that wedges in unknown during post-turn housekeeping is polled forever and can never be delivered to again #306

Closed
opened 2026-09-04 08:06:16 +02:00 by ltms · 1 comment
Owner

Found by a delegated hunter. Latent, not live — see the reachability section; I checked the running config rather than taking the report's word for it.

The one-way gate

Four latches in Injector gate delivery. All four appear in the same condition:

if (!t.awaitingCompletion && !t.postTurnPending
        && !t.awaitingPostTurnPickup && !t.postTurnObserved) {

and in the reclaim check that decides whether a target stays in activeTargets().

The UNKNOWN branch has an escape for exactly one of them:

if (t.awaitingCompletion && ++t.unknownSinceTurn >= TURN_STALL_GRACE_POLLS) {
    ...
    turnFailed = true;
}

CB-109 added that because a worker stuck in a state herdr cannot classify never produces a working → idle boundary. That reasoning applies unchanged to the post-turn phase, but the escape was never extended to it — a gate that closes the direction its incident came from and no other.

  • awaitingPostTurnPickup — set when housekeeping is dispatched, released only on an injectable sample. A worker that goes to UNKNOWN and stays there never produces one.
  • postTurnObserved — set when the reset is seen picked up (WORKING), cleared only on a later injectable sample. Same wedge, one step further along.
  • postTurnPending — needs no escape. It is cleared unconditionally on the line after the listener call that sets it.

Note the short-circuit: ++t.unknownSinceTurn sits inside t.awaitingCompletion && …, so during the post-turn phase the counter does not even increment.

Direction of harm

Silent. The target stays in activeTargets() and is polled forever; the delivery gate blocks every later message to it; and no onTurnFailed fires, so SessionManager.onFailed never runs and the session sits at DONE, looking healthy in fleet_status and fleet_list. Later sends time out as TIMED_OUT_QUEUED with nothing pointing at the cause.

It is bounded, and I want that on the record rather than overstated: Injector.drop clears all four latches, so releasing the session ends the wedge, and the idle reaper releases a DONE session at idleTtlSeconds (1800 in the live config). So this is a wedge until the session is torn down, not a permanent daemon leak.

Reachability — checked, and it is off today

The path needs hasPostTurnAction to return true, which needs clearAfterTurn:

public boolean hasPostTurnAction(String target) {
    if (!clearAfterTurn) return false;

and that comes from cfg.lifecycle().clearAfterTurn(). The live fleetd.yaml lifecycle: block sets idleTtlSeconds, contextCap and drainTimeoutSeconds — and no clearAfterTurn. grep -c clearAfterTurn fleetd.yaml → 0. So the post-turn phase never runs on this fleet, and the wedge is not reachable in the current configuration.

It becomes reachable the moment anyone sets lifecycle.clearAfterTurn: true, which is a supported, documented knob.

One correction to the report. It argued the trigger is live config, citing "memory notes CB-636/637 autocompact is live". That conflates two different settings: autoCompactWindow is a per-profile context-window number, clearAfterTurn is the lifecycle flag this path depends on. The finding is right; that supporting claim is not.

Scope of the fix — deliberately wider than the report

The report named awaitingPostTurnPickup. Fixing only that would rebuild the same one-way gate one notch further along, because postTurnObserved wedges identically. Both are released.

The escape does not set turnFailed. The delegated turn already completed and its waiter already resolved; what is outstanding is the /clear. Reporting a turn failure would drive SessionManager.onFailed on a session that genuinely finished — a worse lie than the wedge.

A separate counter (unknownSincePostTurn) rather than reusing unknownSinceTurn. The two states look mutually exclusive to me, but the argument depends on the delivery gate staying exactly as it is, and a shared counter would fail silently if that ever stopped holding.

Found by a delegated hunter. **Latent, not live** — see the reachability section; I checked the running config rather than taking the report's word for it. ## The one-way gate Four latches in `Injector` gate delivery. All four appear in the same condition: ```java if (!t.awaitingCompletion && !t.postTurnPending && !t.awaitingPostTurnPickup && !t.postTurnObserved) { ``` and in the reclaim check that decides whether a target stays in `activeTargets()`. The `UNKNOWN` branch has an escape for exactly one of them: ```java if (t.awaitingCompletion && ++t.unknownSinceTurn >= TURN_STALL_GRACE_POLLS) { ... turnFailed = true; } ``` CB-109 added that because a worker stuck in a state herdr cannot classify never produces a `working → idle` boundary. That reasoning applies unchanged to the post-turn phase, but the escape was never extended to it — a gate that closes the direction its incident came from and no other. - `awaitingPostTurnPickup` — set when housekeeping is dispatched, released only on an **injectable** sample. A worker that goes to `UNKNOWN` and stays there never produces one. - `postTurnObserved` — set when the reset is seen picked up (`WORKING`), cleared only on a later injectable sample. Same wedge, one step further along. - `postTurnPending` — needs no escape. It is cleared unconditionally on the line after the listener call that sets it. Note the short-circuit: `++t.unknownSinceTurn` sits inside `t.awaitingCompletion && …`, so during the post-turn phase the counter does not even increment. ## Direction of harm Silent. The target stays in `activeTargets()` and is polled forever; the delivery gate blocks every later message to it; and no `onTurnFailed` fires, so `SessionManager.onFailed` never runs and the session sits at `DONE`, looking healthy in `fleet_status` and `fleet_list`. Later sends time out as `TIMED_OUT_QUEUED` with nothing pointing at the cause. It is bounded, and I want that on the record rather than overstated: `Injector.drop` clears all four latches, so releasing the session ends the wedge, and the idle reaper releases a `DONE` session at `idleTtlSeconds` (1800 in the live config). So this is a wedge until the session is torn down, not a permanent daemon leak. ## Reachability — checked, and it is off today The path needs `hasPostTurnAction` to return true, which needs `clearAfterTurn`: ```java public boolean hasPostTurnAction(String target) { if (!clearAfterTurn) return false; ``` and that comes from `cfg.lifecycle().clearAfterTurn()`. The live `fleetd.yaml` `lifecycle:` block sets `idleTtlSeconds`, `contextCap` and `drainTimeoutSeconds` — and no `clearAfterTurn`. `grep -c clearAfterTurn fleetd.yaml` → `0`. So the post-turn phase never runs on this fleet, and **the wedge is not reachable in the current configuration.** It becomes reachable the moment anyone sets `lifecycle.clearAfterTurn: true`, which is a supported, documented knob. **One correction to the report.** It argued the trigger is live config, citing "memory notes CB-636/637 autocompact is live". That conflates two different settings: `autoCompactWindow` is a per-profile context-window number, `clearAfterTurn` is the lifecycle flag this path depends on. The finding is right; that supporting claim is not. ## Scope of the fix — deliberately wider than the report The report named `awaitingPostTurnPickup`. Fixing only that would rebuild the same one-way gate one notch further along, because `postTurnObserved` wedges identically. Both are released. The escape does **not** set `turnFailed`. The delegated turn already completed and its waiter already resolved; what is outstanding is the `/clear`. Reporting a turn failure would drive `SessionManager.onFailed` on a session that genuinely finished — a worse lie than the wedge. A separate counter (`unknownSincePostTurn`) rather than reusing `unknownSinceTurn`. The two states look mutually exclusive to me, but the argument depends on the delivery gate staying exactly as it is, and a shared counter would fail silently if that ever stopped holding.
ltms closed this issue 2026-09-04 08:06:34 +02:00
Author
Owner

Fixed in 76672ff, on main.

Mutation proof

Reverted only Injector.java, kept the three new tests:

[ERROR] Tests run: 30, Failures: 2 -- in dev.ltms.fleet.inject.InjectorTest
AssertionFailedError: a target wedged awaiting post-turn pickup must be reclaimed, not polled forever
    ==> expected: <true> but was: <false>
AssertionFailedError: a wedge after reset pickup must also be reclaimed
    ==> expected: <true> but was: <false>

Both latches fail without the fix, which is the point of covering the sibling rather than only the reported one.

The third test goes the other way: a brief unknown glitch (10 samples, well under the 120-sample grace) must not release the latch, or a queued delegation would overtake housekeeping that is still running. It asserts the target stays active and only first has been sent.

Restored, then mvn clean install: Tests run: 1316, Failures: 0, Errors: 0, Skipped: 0 — BUILD SUCCESS. main was at 1313.

A measurement mistake of my own, worth recording

My first run of these tests looked green. It was not — the test class did not compile, because PostTurnListener did not implement TurnListener.onTurnComplete. I had written mvn ... | grep ... ; echo "rc=$?", and $? after a pipe is grep's exit status, not Maven's. Grep found no failures because there was no test output at all, and reported success.

The mutation run is what exposed it: it failed to compile too, and the compiler error named the real problem. If I had trusted the first "pass" I would have shipped a fix whose tests never ran.

Reading the build's own Tests run: and BUILD lines is the check. An exit code downstream of a pipe is not.

Fixed in `76672ff`, on `main`. ## Mutation proof Reverted only `Injector.java`, kept the three new tests: ``` [ERROR] Tests run: 30, Failures: 2 -- in dev.ltms.fleet.inject.InjectorTest AssertionFailedError: a target wedged awaiting post-turn pickup must be reclaimed, not polled forever ==> expected: <true> but was: <false> AssertionFailedError: a wedge after reset pickup must also be reclaimed ==> expected: <true> but was: <false> ``` Both latches fail without the fix, which is the point of covering the sibling rather than only the reported one. The third test goes the other way: a brief unknown glitch (10 samples, well under the 120-sample grace) must **not** release the latch, or a queued delegation would overtake housekeeping that is still running. It asserts the target stays active and only `first` has been sent. Restored, then `mvn clean install`: `Tests run: 1316, Failures: 0, Errors: 0, Skipped: 0` — `BUILD SUCCESS`. `main` was at 1313. ## A measurement mistake of my own, worth recording My first run of these tests looked green. It was not — the test class did not compile, because `PostTurnListener` did not implement `TurnListener.onTurnComplete`. I had written `mvn ... | grep ... ; echo "rc=$?"`, and `$?` after a pipe is **grep's** exit status, not Maven's. Grep found no failures because there was no test output at all, and reported success. The mutation run is what exposed it: it failed to compile too, and the compiler error named the real problem. If I had trusted the first "pass" I would have shipped a fix whose tests never ran. Reading the build's own `Tests run:` and `BUILD` lines is the check. An exit code downstream of a pipe is not.
Sign in to join this conversation.
1 Participants
Notifications
Due Date
No due date set.
Dependencies

No dependencies set.

Reference: fleet/fleetd#306