Files
fleetd/AUDIT.md
T
2026-09-04 10:33:09 +07:00

1.3 KiB

Health, placement, and cooldown audit

  1. fleetd/src/main/java/dev/ltms/fleet/session/SessionManager.java:741
  2. issue: After a classified backend error, the session stays BACKEND_ERROR; it cannot receive another delivery or be reaped, but it remains in the roster and counts against maxLoad, so after the short cool-off expires the fleet can refuse a healthy replacement because a dead member still holds the only seat.
  3. fix: On the backend-error path, release the unusable session (and abandon any remaining work), or make BACKEND_ERROR sessions reclaimable and exclude them from the placement live count.
  4. severity: high

Sequence checked: a classified backend error calls Fleetd.backendErrorSink, which calls SessionManager.onBackendError and sets BACKEND_ERROR. onDelivered accepts only READY and DONE; reapIdle also handles only those states. Fleetd counts every roster entry for placement. Once BackendOutagePolicy ends its short cool-off, a profile with maxLoad: 1 still has live == 1 and rejects a new spawn. This conclusion does not depend on the unread live fleetd.yaml; it needs any profile with a finite maxLoad.

Confidence: high. The state transitions and placement count are direct code paths. I did not run builds or tests, as requested.