diff --git a/docs/M4-Fleet-Health.md b/docs/M4-Fleet-Health.md index c605172..0ea6043 100644 --- a/docs/M4-Fleet-Health.md +++ b/docs/M4-Fleet-Health.md @@ -1,7 +1,9 @@ # M4 - Fleet health, recovery, routing, and capacity **Status:** Design accepted on 2026-08-15. CB-573 part 1 has shipped the classification model and -the `bridge_list` capacity view; the remaining M4 units are not yet shipped. +the `bridge_list` capacity view; the remaining M4 units are not yet shipped. See +[Unit 2 — what has landed so far](#unit-2---what-has-landed-so-far) before planning Unit 2 work: +some of its criteria were met by separate CB tickets, and one of them contradicts the unit text. **Scope:** Fleet evidence, safe mechanical repair, lead routing, capacity reporting, and optional human notification. **Grounded in:** `health/FleetHealth`, `health/PaneBudget`, `inject/StatusPoller`, @@ -799,6 +801,39 @@ Acceptance criteria: 21. Tests cover both real traces, all repair refusals, clipping, explicit-reply and next-turn races, restart without capture, concurrent send and release, and preserved discovery after restart. +#### Unit 2 - what has landed so far + +Checked against `main` at `e09cac6` on 2026-08-15. Unit 2 was written as one block, but parts of it +have since been built by separate CB tickets. Read this before planning the rest, or that work gets +done twice. + +The check was a symbol survey of `bridged/src/main/java` plus the merge history. It tells you whether +the machinery exists at all. It is **not** a line-by-line audit of whether each criterion is fully +met, and I did not run one. + +| Criterion | Marker searched for | Found in main source | Reading | +|---|---|---|---| +| 1 | `TurnToken` | 8 files | **Done** — unit 2a, merged as `fec284e`. Criterion 1 was corrected first; see the note under it. | +| 2-5, 9 | `REPAIRED` | 0 files | Not started. The whole guarded-repair path is absent. | +| 6, 7, 10 | `reconcileLostBoundary` | 0 files | Not started. | +| 12 | CB-568 failure operation | via CB-580 | **Partial.** CB-580 (`0af902e`) routes `GONE` and `NEVER_READY` into the one idempotent target-wide failure. I did not check that release and abnormal stop go through the same call. | +| 14 | `DELEGATION_ORPHANED` | 3 files | **Partial.** The health state exists. The teardown-invariant check that creates it, and the retry rule, do not. | +| 15 | `SPAWN_ROLLBACK` | 0 files | **Contradicted — see below.** | +| 16 | — | — | Partial at best. CB-576 made release preserve a dirty worktree; whether explicit stop is state-aware is not checked. | +| 17, 18 | `preservedWorktrees` | 0 files | Not started. No manifest, and no lead-only `bridge_list` field. | +| 19 | `WORK_PRODUCT_AT_RISK` | 0 files | Not started. | + +**Criterion 15 no longer matches the code, and the code is right.** It says "normal `COMPLETED` +remove worktrees". Since CB-576 (`500bfa2`) that is false on purpose: a `COMPLETED` release now +preserves the worktree when it still holds uncommitted work, because deleting it destroys work +nobody can get back. CB-576 was filed after exactly that loss. CB-581 goes further — if the +dirty-check itself fails, the worktree is preserved rather than removed, since "we could not tell" +must not be treated as "it is clean". + +So criterion 15 should be rewritten as: `SPAWN_ROLLBACK` and a `COMPLETED` release with a **clean** +worktree remove it; abnormal causes, shutdown, a dirty worktree, and a failed dirty-check all +preserve it. `SPAWN_ROLLBACK` itself does not exist yet. + ### Unit 3 - Typed inbox and member routing Scope: semantic record, AMQP migration, both adapters, member routing, polling, and member health in