diff --git a/AUDIT.md b/AUDIT.md new file mode 100644 index 0000000..dafdf1e --- /dev/null +++ b/AUDIT.md @@ -0,0 +1,10 @@ +# Health, placement, and cooldown audit + +1. `fleetd/src/main/java/dev/ltms/fleet/session/SessionManager.java:741` +2. issue: After a classified backend error, the session stays `BACKEND_ERROR`; it cannot receive another delivery or be reaped, but it remains in the roster and counts against `maxLoad`, so after the short cool-off expires the fleet can refuse a healthy replacement because a dead member still holds the only seat. +3. fix: On the backend-error path, release the unusable session (and abandon any remaining work), or make `BACKEND_ERROR` sessions reclaimable and exclude them from the placement live count. +4. severity: high + +Sequence checked: a classified backend error calls `Fleetd.backendErrorSink`, which calls `SessionManager.onBackendError` and sets `BACKEND_ERROR`. `onDelivered` accepts only `READY` and `DONE`; `reapIdle` also handles only those states. `Fleetd` counts every roster entry for placement. Once `BackendOutagePolicy` ends its short cool-off, a profile with `maxLoad: 1` still has `live == 1` and rejects a new spawn. This conclusion does not depend on the unread live `fleetd.yaml`; it needs any profile with a finite `maxLoad`. + +Confidence: high. The state transitions and placement count are direct code paths. I did not run builds or tests, as requested.