fleetd #284: free backend-error capacity #287

Closed
agent wants to merge 0 commits from worker/fix-284-backend-error-seat-85912c-11 into main
Member

Report

PR: #287
Branch: worker/fix-284-backend-error-seat-85912c-11
Root: /Users/dai.ha/LTMS/.bridged-worktrees/04928f-11
Files: fleetd/src/main/java/dev/ltms/fleet/Fleetd.java, fleetd/src/main/java/dev/ltms/fleet/mcp/FleetMcp.java, fleetd/src/test/java/dev/ltms/fleet/FleetdBackendErrorSinkTest.java, and fleetd/src/test/java/dev/ltms/fleet/mcp/FleetMcpTest.java.

Decision

I chose option 2. BACKEND_ERROR stays in the roster so fleet_list keeps its failure details. It no longer uses profile capacity. This avoids a teardown path that could remove a pane or worktree after a backend failure.

Change

  • Fleetd now excludes BACKEND_ERROR from the live counter used by the real CompositePeerLauncher maxLoad gate.
  • fleet_list reports BACKEND_ERROR sessions as reclaimable.
  • The gate-level test acquires a maxLoad-limited member, calls onBackendError, then acquires another member on the same profile. The second spawn succeeds.

Regression proof

I removed the BACKEND_ERROR exclusion temporarily and ran the focused test. It failed with:

worker profile 'terra' is at maxLoad: 1 live >= 1 cap; refusing spawn — no fallback to another profile

I restored the exclusion before the full build.

Checks

  • cd fleetd && mvn clean install
  • Tests run: 1285, Failures: 0, Errors: 0, Skipped: 0
  • BUILD SUCCESS

Caveat

FAILED is also set by SessionManager.onFailed, is non-deliverable, and is not reaped. It remains outside this issue scope.

## Report PR: https://git.ltms.dev/fleet/fleetd/pulls/287 Branch: `worker/fix-284-backend-error-seat-85912c-11` Root: `/Users/dai.ha/LTMS/.bridged-worktrees/04928f-11` Files: `fleetd/src/main/java/dev/ltms/fleet/Fleetd.java`, `fleetd/src/main/java/dev/ltms/fleet/mcp/FleetMcp.java`, `fleetd/src/test/java/dev/ltms/fleet/FleetdBackendErrorSinkTest.java`, and `fleetd/src/test/java/dev/ltms/fleet/mcp/FleetMcpTest.java`. ### Decision I chose option 2. `BACKEND_ERROR` stays in the roster so `fleet_list` keeps its failure details. It no longer uses profile capacity. This avoids a teardown path that could remove a pane or worktree after a backend failure. ### Change - `Fleetd` now excludes `BACKEND_ERROR` from the live counter used by the real `CompositePeerLauncher` maxLoad gate. - `fleet_list` reports `BACKEND_ERROR` sessions as reclaimable. - The gate-level test acquires a maxLoad-limited member, calls `onBackendError`, then acquires another member on the same profile. The second spawn succeeds. ### Regression proof I removed the `BACKEND_ERROR` exclusion temporarily and ran the focused test. It failed with: `worker profile 'terra' is at maxLoad: 1 live >= 1 cap; refusing spawn — no fallback to another profile` I restored the exclusion before the full build. ### Checks - `cd fleetd && mvn clean install` - `Tests run: 1285, Failures: 0, Errors: 0, Skipped: 0` - `BUILD SUCCESS` ### Caveat `FAILED` is also set by `SessionManager.onFailed`, is non-deliverable, and is not reaped. It remains outside this issue scope.
agent added 1 commit 2026-09-04 05:47:47 +02:00
fleetd #284: free backend-error capacity
CI / contract (pull_request) Successful in 46s
CI / build (pull_request) Successful in 2m42s
4dd12083ab
agent added 1 commit 2026-09-04 05:55:24 +02:00
fleetd #284: free failed-session capacity
CI / contract (pull_request) Successful in 1m11s
CI / build (pull_request) Successful in 1m19s
11cbfa79b4
Owner

Merged locally. Posting this here because the worker had no gh and could not update its own PR body — the round-two commit 11cbfa7 extends the fix from BACKEND_ERROR to FAILED as well, and the body still describes only the first round.

I also changed one thing the worker did on my own instruction, because the instruction was wrong. Reasoning is on #284.

Merged locally. Posting this here because the worker had no `gh` and could not update its own PR body — the round-two commit `11cbfa7` extends the fix from `BACKEND_ERROR` to `FAILED` as well, and the body still describes only the first round. I also changed one thing the worker did on my own instruction, because the instruction was wrong. Reasoning is on #284.
ltms closed this pull request 2026-09-04 06:14:11 +02:00
Some checks are pending
CI / contract (pull_request) Successful in 1m11s
CI / build (pull_request) Successful in 1m19s

Pull request closed

Sign in to join this conversation.