Compare commits

...

1 Commits

Author SHA1 Message Date
Dai Ha 7269351542 audit health placement cooldowns 2026-09-04 10:33:09 +07:00
+10
View File
@@ -0,0 +1,10 @@
# Health, placement, and cooldown audit
1. `fleetd/src/main/java/dev/ltms/fleet/session/SessionManager.java:741`
2. issue: After a classified backend error, the session stays `BACKEND_ERROR`; it cannot receive another delivery or be reaped, but it remains in the roster and counts against `maxLoad`, so after the short cool-off expires the fleet can refuse a healthy replacement because a dead member still holds the only seat.
3. fix: On the backend-error path, release the unusable session (and abandon any remaining work), or make `BACKEND_ERROR` sessions reclaimable and exclude them from the placement live count.
4. severity: high
Sequence checked: a classified backend error calls `Fleetd.backendErrorSink`, which calls `SessionManager.onBackendError` and sets `BACKEND_ERROR`. `onDelivered` accepts only `READY` and `DONE`; `reapIdle` also handles only those states. `Fleetd` counts every roster entry for placement. Once `BackendOutagePolicy` ends its short cool-off, a profile with `maxLoad: 1` still has `live == 1` and rejects a new spawn. This conclusion does not depend on the unread live `fleetd.yaml`; it needs any profile with a finite `maxLoad`.
Confidence: high. The state transitions and placement count are direct code paths. I did not run builds or tests, as requested.