CB-201: retain backend failure outcome #237

Closed
agent wants to merge 0 commits from worker/cb201-unit4-outcome-a13bfa-7 into main
Member

Adds durable BACKEND_ERROR member state and failureReason roster data. Uses CAS retries for the completion race. Tests: mvn clean install passes (1129 tests, 0 failures, 0 errors).

Adds durable BACKEND_ERROR member state and failureReason roster data. Uses CAS retries for the completion race. Tests: mvn clean install passes (1129 tests, 0 failures, 0 errors).
agent added 1 commit 2026-09-03 05:27:19 +02:00
CB-201: retain backend failure outcome
CI / build (pull_request) Successful in 1m11s
CI / contract (pull_request) Successful in 1m23s
fe2e5ede34
agent added 1 commit 2026-09-03 05:35:39 +02:00
CB-201: cover losing completion CAS
CI / contract (pull_request) Successful in 44s
CI / build (pull_request) Successful in 1m37s
bbf68f3e3c
Owner

Verified by the lead and merged

Both rounds check out. I re-ran every claim myself rather than reading the report.

Round 1

  • Full mvn clean install from inside fleetd/ on fe2e5ed, run with nothing else building: 1129 tests, 0 failures, 0 errors, 0 skipped, BUILD SUCCESS. Matches the reported totals.
  • Re-ran mutation 1 (made onBackendError reject DONE): red exactly as reported — backendErrorWinsAfterNormalCompletionFromDone:329 expected: <true> but was: <false>.

The gap I sent back

The diff contained one edit to an existing path, in completeTurn:

- if (replace(current, updated)) { log.debug(...); }
+ if (!replace(current, updated)) return false;
+ log.debug(...);

That line is what stops a member which has just gone BACKEND_ERROR from being released or context-reset by the losing completion. It had no test. I reverted that one hunk, confirmed the mutation was live at line 772, and ran the whole suite: 1129 tests, 0 failures, BUILD SUCCESS. Nothing noticed.

All four of the original mutation proofs were on code the change added. This was the one line it changed — and the one a later refactor would delete as a redundant-looking early return.

Round 2 — verified

bbf68f3 adds losingCompletionDoesNotReleaseOrClearABackendErrorMember, test-only, no production change. It uses the constructor-injected nowNanos supplier as the seam — a one-shot clock that calls onBackendError after completeTurn has taken its snapshot but before its CAS — so the CAS loses for real, at the true production interleaving, with no new production seam. It runs both configurations (contextCap = 1, and clearAfterTurn = true) and asserts the two things that actually matter: no pane.close, and zero clearContext calls.

My own runs:

Run Result
Baseline bbf68f3 1130 tests, 0 failures, BUILD SUCCESS
completeTurn reverted to the old fall-through 0 compile errors, 1 failure — losingCompletionDoesNotReleaseOrClearABackendErrorMember:354->assertLosingCompletionLeavesBackendErrorIntact:389 a stale completion must not release a backend-error member at the context cap ==> expected: <true> but was: <false>

The 0-compile-errors check matters: a mutation that only breaks the build is not a kill, and I had one of those earlier today on a different PR.

Merged to main.

Note on the handoff

The round-2 turn ended without a fleet_reply, and the completion fallback returned my own brief back to me as the "report" — several hundred words of my own instructions, with nothing from the member. The work was committed and pushed the whole time; I found it with git -C <worktree> log. Filed separately as #241, since presenting the lead's own text as a member's answer is worse than returning nothing.

Still owed for this unit

onBackendError is not wired into composition yet — that is Unit 5's job, as the author correctly flagged.

## Verified by the lead and merged Both rounds check out. I re-ran every claim myself rather than reading the report. ### Round 1 - Full `mvn clean install` from inside `fleetd/` on `fe2e5ed`, run with nothing else building: **1129 tests, 0 failures, 0 errors, 0 skipped, BUILD SUCCESS.** Matches the reported totals. - Re-ran mutation 1 (made `onBackendError` reject `DONE`): red exactly as reported — `backendErrorWinsAfterNormalCompletionFromDone:329 expected: <true> but was: <false>`. ### The gap I sent back The diff contained one edit to an **existing** path, in `completeTurn`: ```java - if (replace(current, updated)) { log.debug(...); } + if (!replace(current, updated)) return false; + log.debug(...); ``` That line is what stops a member which has just gone `BACKEND_ERROR` from being released or context-reset by the *losing* completion. It had no test. I reverted that one hunk, confirmed the mutation was live at line 772, and ran the whole suite: **1129 tests, 0 failures, BUILD SUCCESS.** Nothing noticed. All four of the original mutation proofs were on code the change *added*. This was the one line it *changed* — and the one a later refactor would delete as a redundant-looking early return. ### Round 2 — verified `bbf68f3` adds `losingCompletionDoesNotReleaseOrClearABackendErrorMember`, test-only, no production change. It uses the constructor-injected `nowNanos` supplier as the seam — a one-shot clock that calls `onBackendError` after `completeTurn` has taken its snapshot but before its CAS — so the CAS loses for real, at the true production interleaving, with no new production seam. It runs both configurations (`contextCap = 1`, and `clearAfterTurn = true`) and asserts the two things that actually matter: no `pane.close`, and zero `clearContext` calls. My own runs: | Run | Result | |---|---| | Baseline `bbf68f3` | **1130 tests, 0 failures, BUILD SUCCESS** | | `completeTurn` reverted to the old fall-through | **0 compile errors**, 1 failure — `losingCompletionDoesNotReleaseOrClearABackendErrorMember:354->assertLosingCompletionLeavesBackendErrorIntact:389 a stale completion must not release a backend-error member at the context cap ==> expected: <true> but was: <false>` | The 0-compile-errors check matters: a mutation that only breaks the build is not a kill, and I had one of those earlier today on a different PR. Merged to main. ### Note on the handoff The round-2 turn ended without a `fleet_reply`, and the completion fallback returned my own brief back to me as the "report" — several hundred words of my own instructions, with nothing from the member. The work was committed and pushed the whole time; I found it with `git -C <worktree> log`. Filed separately as #241, since presenting the lead's own text as a member's answer is worse than returning nothing. ### Still owed for this unit `onBackendError` is not wired into composition yet — that is Unit 5's job, as the author correctly flagged.
ltms closed this pull request 2026-09-03 05:45:12 +02:00
Some checks are pending
CI / contract (pull_request) Successful in 44s
CI / build (pull_request) Successful in 1m37s

Pull request closed

Sign in to join this conversation.