audit health placement cooldowns
This commit is contained in:
@@ -0,0 +1,10 @@
|
||||
# Health, placement, and cooldown audit
|
||||
|
||||
1. `fleetd/src/main/java/dev/ltms/fleet/session/SessionManager.java:741`
|
||||
2. issue: After a classified backend error, the session stays `BACKEND_ERROR`; it cannot receive another delivery or be reaped, but it remains in the roster and counts against `maxLoad`, so after the short cool-off expires the fleet can refuse a healthy replacement because a dead member still holds the only seat.
|
||||
3. fix: On the backend-error path, release the unusable session (and abandon any remaining work), or make `BACKEND_ERROR` sessions reclaimable and exclude them from the placement live count.
|
||||
4. severity: high
|
||||
|
||||
Sequence checked: a classified backend error calls `Fleetd.backendErrorSink`, which calls `SessionManager.onBackendError` and sets `BACKEND_ERROR`. `onDelivered` accepts only `READY` and `DONE`; `reapIdle` also handles only those states. `Fleetd` counts every roster entry for placement. Once `BackendOutagePolicy` ends its short cool-off, a profile with `maxLoad: 1` still has `live == 1` and rejects a new spawn. This conclusion does not depend on the unread live `fleetd.yaml`; it needs any profile with a finite `maxLoad`.
|
||||
|
||||
Confidence: high. The state transitions and placement count are direct code paths. I did not run builds or tests, as requested.
|
||||
Reference in New Issue
Block a user