1.3 KiB
1.3 KiB
Health, placement, and cooldown audit
fleetd/src/main/java/dev/ltms/fleet/session/SessionManager.java:741- issue: After a classified backend error, the session stays
BACKEND_ERROR; it cannot receive another delivery or be reaped, but it remains in the roster and counts againstmaxLoad, so after the short cool-off expires the fleet can refuse a healthy replacement because a dead member still holds the only seat. - fix: On the backend-error path, release the unusable session (and abandon any remaining work), or make
BACKEND_ERRORsessions reclaimable and exclude them from the placement live count. - severity: high
Sequence checked: a classified backend error calls Fleetd.backendErrorSink, which calls SessionManager.onBackendError and sets BACKEND_ERROR. onDelivered accepts only READY and DONE; reapIdle also handles only those states. Fleetd counts every roster entry for placement. Once BackendOutagePolicy ends its short cool-off, a profile with maxLoad: 1 still has live == 1 and rejects a new spawn. This conclusion does not depend on the unread live fleetd.yaml; it needs any profile with a finite maxLoad.
Confidence: high. The state transitions and placement count are direct code paths. I did not run builds or tests, as requested.