# CB-201 and CB-227 refinement Date: 2026-09-03 ## Decision #201 and #227 are one delivery program, but they are not one implementation unit. #201 has a real seam: `CompletionResolver` can publish a typed backend-error event only after its waiter resolution wins. #227 can consume that event without knowing any pane text. The classifier must land before the final #227 wiring. However, the policy engine, roster state, and lead nudge can be built in parallel with the classifier. I propose five units. Units 1 to 4 own separate files and can run in parallel. Unit 5 owns all composition files and lands after them. It also depends on the #234 defect 2 fix named in the task. ```mermaid flowchart LR U1["Unit 1: typed backend-error classification"] --> U5["Unit 5: wire policy, spawn gate, and fleet views"] U2["Unit 2: credential outage policy"] --> U5 U3["Unit 3: lead outage nudge"] --> U5 U4["Unit 4: durable member outcome"] --> U5 D234["#234 defect 2: fail-loud target resolution"] --> U5 ``` *Figure 1. Four file-disjoint foundations feed one composition unit.* This split keeps `Fleetd.java` under one owner. It also keeps every other production file under one unit in this plan. ## Evidence checked in the current branch I read both issue pages in full. Each page reports zero comments. | Evidence | What the code says now | |---|---| | `inject/CompletionResolver.java:229-237` | A turn below two seconds fails before normal scrape classification. A matching fast backend error is therefore only a generic failure today. | | `inject/CompletionResolver.java:260-276` and `:332-367` | #211 already added raw-screen classification when `lastAssistantBlock` is empty. The “dead-code question” in #201 is stale on this branch. | | `inject/CompletionResolver.java:288-317` | Exhaustion wins before the hard-coded `API Error:` match. A backend error then goes through generic `fail(...)`. | | `inject/CompletionResolver.java:311-316` | The code admits that the pattern is a heuristic. A member report which quotes an API error may match it. | | `inject/CompletionResolver.java:449-467` | Startup coverage exists only for `exhaustedPattern`. | | `inject/ExhaustedPatternLookup.java:13-25` | The current lookup and explicit `none()` value are a good shape for the new classifier seam. | | `Fleetd.java:196-207` | One `BackendQuarantine` is shared by placement and the exhaustion sink. Its cooldown comes from `quarantineCooldownSeconds`. | | `Fleetd.java:322-363` | Pattern compilation, target-to-profile lookup, and the live `ExhaustionSink` are composed in `Fleetd.main`. The sink on this branch still ends in `.ifPresent(...)`. This plan assumes #234 replaces that silent path. | | `placement/BackendQuarantine.java:60-87` | A repeated exhaustion restarts one long quarantine. The store is credential-keyed and uses an injected monotonic clock. | | `member/CompositePeerLauncher.java:260-317` | Explicit and policy-selected spawns have separate gates. Both paths must learn about outage cool-off. | | `member/CompositePeerLauncher.java:347-379` | Exhaustion refusal already checks a credential for explicit spawns and filters policy candidates. Its error text says “exhausted”. | | `placement/PlacementContext.java:10-22` and `PlacementPolicyUtil.java:14-83` | Automatic placement has only one transient exclusion set named `quarantined`. Reusing it would make outage errors say “backend exhausted”. | | `mcp/FleetMcp.java:913-1025` | `fleet_list` sets `free: 0` and adds `credentialId` plus `quarantinedForSeconds` when quarantine is active. | | `session/MemberSession.java:51-59` | The roster has `DONE` and generic `FAILED`, but no backend-error state or stored reason. | | `session/SessionManager.java:695-773` | A normal boundary moves `BUSY` to `DONE`. A failure moves any non-released session to `FAILED`. The async completion resolver can race the `DONE` update. | | `session/SessionManager.java:648-687` | `rosterView` reports the session state, but it reports no terminal reason. | | `msg/MessageService.java:922-940` | CB-588 already nudges for every terminal async ticket, including failures. Current code would report failed tickets, but it would not report one correlated outage. | | `msg/ReplyPushLoop.java:20-48` | Replies, terminal tickets, and questions share one per-lead schedule. This prevents two push sources from injecting competing turns. | | `msg/ReplyPushLoop.java:305-395` | Each push entry point resolves the owning lead through `PrimaryRegistry`. Missing ownership is logged and the durable or pending item remains the backstop. | | `msg/ReplyPushLoop.java:496-547` | One tick builds one combined nudge. Pending items have separate reminder counts. | | `health/FleetHealthMonitor.java:91-143` | Health is a slow periodic observer of members and message-layer facts. It does not receive completion classifications. | | `health/FleetHealthMonitor.java:206-208` | `healthCoverage` means health enabled plus webhook configured. It does not describe lead-pane alerts. | | `Fleetd.java:465-486` | Health stays `detection-only` without the webhook notification setting. | I also read the related unit tests for `CompletionResolver`, `ReplyPushLoop`, `BackendQuarantine`, `CompositePeerLauncher`, `PlacementPolicyUtil`, `SessionManager`, `MessageService`, and `FleetMcp`. I did not inspect the in-progress #234 branch. I only used the two measured facts in the task. No peer architect was named, so I did not exchange a design with one. ## Required behaviour The policy should use these first values: - Threshold: **2** classified backend errors. - Window: **60 seconds**, measured from the first error to the second. - Cool-off: **60 seconds**, starting when the threshold is reached. - Correlation key: `credentialId`, never profile name and never error text. - Incident rule: one active incident per credential. Errors during its cool-off do not extend it and do not create more lead notices. - Rearm rule: after cool-off ends, two fresh errors are needed for another incident. Two errors are the smallest threshold which protects the honest one-turn failure. A 60-second window fits the measured two-member outage. A 60-second cool-off blocks immediate repeat spawns without turning a short backend fault into the default 1,800-second exhaustion quarantine. A single classified error still fails its send and marks its member `backend_error`. It does not cool a credential and does not send an outage notice. This is what “a single error changes nothing” must mean at the credential level. It cannot mean that the failed member still looks successful. ```mermaid sequenceDiagram participant R1 as Resolver for member A participant R2 as Resolver for member B participant P as Outage policy participant S as Spawn gate participant N as Lead push loop participant L as Lead pane R1->>P: backend error for credential C Note over P: Count 1, no cool-off R2->>P: backend error for credential C within 60s P->>P: Start one 60s incident P->>S: Credential C is cooling off P->>N: Queue one incident notice N->>L: Inject when lead is idle, blocked, or done L->>S: Request another spawn on credential C S-->>L: Refuse and report remaining cool-off ``` *Figure 2. The second independent classification creates the fleet-level event.* Against the 2026-09-01 case, the second failed member would start cool-off. `fleet_list` would show zero free capacity and both members as `backend_error`. The push loop would inject one outage notice even if the lead had not polled either ticket yet. The design reports the outage. It does not recover uncommitted work from the members. ## Unit 1 — Typed backend-error classification ### Scope Replace the direct hard-coded check inside `CompletionResolver` with a lookup and a sink. Keep the public send result as a failed send. The typed internal event is the seam #227 consumes. The lookup returns the pattern for a target. The sink receives the target, matched line, and full failure reason. It fires only after `Rendezvous.resolveFailure(...)` wins for that exact captured waiter. This copies the race rule already used by `ExhaustionSink`. The classifier must run in all three current paths: 1. a normal non-empty assistant block; 2. the #211 raw scrape fallback; 3. a turn inside `MIN_TURN_NANOS`, before it becomes a generic too-fast failure. In every path, the order stays: stale-baseline guard, exhaustion, backend error, then generic failure or completion. A fast turn still fails when no configured pattern matches. Keep `(?i)\bAPI Error\s*:` as a compatibility pattern for profiles without `errorPattern` until the operator config is updated. Do not call this full coverage. Startup reporting in Unit 5 must name profiles using this weaker legacy default. ### Files owned - Add `fleetd/src/main/java/dev/ltms/fleet/inject/BackendErrorPatternLookup.java`. - Add `fleetd/src/main/java/dev/ltms/fleet/inject/BackendErrorSink.java`. - Change `fleetd/src/main/java/dev/ltms/fleet/inject/CompletionResolver.java`. - Change `fleetd/src/test/java/dev/ltms/fleet/inject/CompletionResolverTest.java`. No other unit may edit these files. ### Acceptance criteria 1. A target-specific error pattern matches a normal assistant block and resolves the send as failed. 2. The same match calls `BackendErrorSink` exactly once after the waiter resolution wins. 3. A late classification which loses to `fleet_reply` does not call the sink. 4. An exhausted line that also matches the generic error pattern stays `BACKEND_EXHAUSTED`. It calls only `ExhaustionSink`. 5. A raw pane with leading Terminal User Interface (TUI) chrome and no assistant marker still uses the #211 fallback and calls the backend-error sink. 6. A matching error inside the two-second floor is typed and sent to the sink. A non-matching fast turn stays a generic failure. 7. An unchanged delivery baseline which contains old backend-error text is suppressed. It never increments outage evidence. 8. A non-match keeps the existing completion result and text. 9. Constructors used by current callers keep compiling. They use the legacy default lookup and an explicit inert sink until Unit 5 supplies the production objects. 10. Unit tests pass. The developer runs the focused test first, then `mvn clean install` from `fleetd/`. ### Dependencies None. Unit 1 can run with Units 2, 3, and 4. Unit 5 depends on its new lookup, sink, and constructor. ### What to report back - The exact classifier order in all three paths. - The focused test command and result. - The test which proves a losing waiter race does not publish an event. - The test which proves a fast matching failure is typed. - The final `mvn clean install` result. - Any constructor kept only for transition and where Unit 5 replaces it. ## Unit 2 — Credential outage policy ### Scope Build a small credential-keyed state machine. It accepts already-classified backend-error events. It does not read pane text, profiles, sessions, or lead state. Use an injected monotonic clock. A call records `credentialId`, target, and reason. It returns a new incident only on the threshold crossing. The incident contains a stable event id, credential id, the distinct affected targets, evidence count, window, and remaining cool-off. This class owns both correlation and short cool-off. Keeping them together makes threshold crossing and the cool-off deadline one atomic state change. ### Files owned - Add `fleetd/src/main/java/dev/ltms/fleet/placement/BackendOutagePolicy.java`. - Add `fleetd/src/test/java/dev/ltms/fleet/placement/BackendOutagePolicyTest.java`. No other unit may edit these files. ### Acceptance criteria 1. One error creates no incident and no cool-off. 2. Two errors for one credential within 60 seconds create exactly one incident and a 60-second cool-off. 3. Two errors more than 60 seconds apart do not create an incident. 4. The exact 60-second boundary has a pinned result. Use inclusive `<= 60s` so scheduler delay does not discard evidence at the boundary. 5. Different credentials never share evidence. 6. Different profiles which supply the same credential id do share evidence. The policy itself only sees the credential id. 7. More errors during active cool-off do not extend its deadline and do not return another incident. 8. After expiry, old evidence is cleared. Two fresh errors are needed to create the next incident. 9. Remaining seconds round up, matching `BackendQuarantine` reporting. 10. Concurrent second and third errors cannot return two incidents. 11. The focused tests and `mvn clean install` pass. ### Dependencies None. Unit 2 can run with Units 1, 3, and 4. Unit 5 depends on the policy API. ### What to report back - The state transition table and locking method. - The exact threshold, window, cool-off, and boundary rule. - The test which proves one incident under concurrent calls. - The focused test and `mvn clean install` results. ## Unit 3 — Lead outage nudge ### Scope Add backend incidents as a fourth pending source in `ReplyPushLoop`. Do not create another scheduler or call `AgentControl.send` from `Fleetd`. The existing combined per-lead schedule is the control which prevents competing injected turns. The entry point takes an incident id, affected worker targets, credential id, affected profile names, and remaining cool-off. It resolves distinct owning leads through `PrimaryRegistry`. Each `(incidentId, lead)` item is one-shot. It waits while the lead is not injectable. After one successful `agents.send`, remove it. A send exception keeps it pending for a bounded retry. It never uses the repeated reminder behaviour of an uncollected ticket. Also add a fail-loud entry point for a classified target that Unit 5 cannot map to a credential. It uses `PrimaryRegistry.nudgeTargetFor(target)` and says that correlation could not run. If no lead is known, log at `WARN`, not `DEBUG`. ### Files owned - Change `fleetd/src/main/java/dev/ltms/fleet/msg/ReplyPushLoop.java`. - Change `fleetd/src/test/java/dev/ltms/fleet/msg/ReplyPushLoopTest.java`. No other unit may edit these files. ### Acceptance criteria 1. One incident affecting two workers owned by one lead causes one successful pane injection. 2. Two affected workers owned by two leads cause one successful injection per affected lead. This is one notice per event per lead, not one notice per member. 3. Repeating the same incident id is idempotent. 4. A busy or unknown lead is not injected. The item stays pending until the lead becomes injectable or its attempt cap is reached. 5. After one successful injection, later ticks do not mention that incident again. 6. A failed `agents.send` is retried within the existing bound. A successful retry still gives only one successful send. 7. A pending ticket and an outage incident for one lead appear in one combined nudge, not two competing turns. 8. The text names the credential, profiles, affected workers, and remaining cool-off. It tells the lead to run `fleet_list`. 9. An unmapped target produces a direct warning notice when a lead is known. If no lead is known, the code logs a `WARN` naming the target and reason. 10. `stop()` clears incident state as it clears other push state. 11. Existing reply, ticket, and question tests stay green. The focused tests and `mvn clean install` pass. ### Dependencies None. The API uses plain values, not the Unit 2 incident class. This lets Unit 3 run in parallel. Unit 5 adapts the Unit 2 incident into this entry point. ### What to report back - The exact one-shot and retry rules. - The test showing one combined nudge with a failed ticket. - The test showing one successful send for two affected workers. - The focused test and `mvn clean install` results. ## Unit 4 — Durable member backend-error outcome ### Scope Make a classified backend failure remain visible after its ticket is collected or expires. Add `BACKEND_ERROR` to `MemberSession.State`. Add a nullable failure detail to `MemberSession` and render it as `failureReason` in `SessionManager.rosterView`. Add `SessionManager.onBackendError(target, reason)`. The transition must handle both completion orderings: - `BUSY -> BACKEND_ERROR` when classification wins before the normal completion state update; - `DONE -> BACKEND_ERROR` when the async resolver runs after `SessionManager.onTurnComplete`. It must use a compare-and-set retry or another atomic update. `RELEASED` must never return to the roster. A backend-error member is terminal and cannot accept another delivery. ### Files owned - Change `fleetd/src/main/java/dev/ltms/fleet/session/MemberSession.java`. - Change `fleetd/src/main/java/dev/ltms/fleet/session/SessionManager.java`. - Change `fleetd/src/test/java/dev/ltms/fleet/session/SessionManagerTest.java`. No other unit may edit these files. ### Acceptance criteria 1. `onBackendError` moves a `BUSY` member to `BACKEND_ERROR` and stores the reason. 2. It also moves `DONE` to `BACKEND_ERROR`, covering the resolver race. 3. A later normal `onTurnComplete` cannot change `BACKEND_ERROR` back to `DONE`. 4. A released or unknown member is not recreated. The unknown case logs at `WARN` and returns an explicit false result to its caller. 5. `onDelivered` refuses a `BACKEND_ERROR` member, just as it refuses generic `FAILED`. 6. `rosterView` reports `state: backend_error` and `failureReason` after the send ticket is gone. 7. Ordinary members do not gain a blank or invented `failureReason` field. 8. Existing constructors keep source compatibility for tests and adapters. 9. The focused tests and `mvn clean install` pass. ### Dependencies None. Unit 4 can run with Units 1, 2, and 3. Unit 5 calls the new session method from the production sink. ### What to report back - The two race orderings and the tests for both. - The exact roster JSON shape. - The unknown-target result and log level. - The focused test and `mvn clean install` results. ## Unit 5 — Production wiring, spawn gate, and fleet views ### Scope Compose Units 1 to 4 in production. This is the only unit which edits `Fleetd.java`. Add per-profile `errorPattern` config beside `exhaustedPattern`. Compile both once at startup. A configured pattern wins over the legacy default. Report configured profiles and legacy-default profiles separately at startup. A bad regex must stop startup with the profile and key in the message. Wire one production `BackendErrorSink` with this order: 1. mark the member `backend_error` with its reason; 2. resolve the profile and its current `effectiveCredentialId()` through the fail-loud #234 seam; 3. record the error in `BackendOutagePolicy`; 4. on a new incident, submit one event to `ReplyPushLoop`. If target metadata cannot be resolved, do not end in `Optional.ifPresent`. Log an error and call the Unit 3 unmapped-target notice. The failed send still reaches its ticket through CB-588. Teach both spawn paths about a separate cool-off source. Exhaustion quarantine has priority when both states are active. Automatic placement needs a distinct `coolingOff` set so its refusal does not say “exhausted”. Extend the MCP (Model Context Protocol) views: - A cooling profile has `free: 0`, `credentialId`, and `coolingOffForSeconds` in `fleet_list`. - It does not have `quarantinedForSeconds` unless exhaustion quarantine is also active. - `fleet_profiles` has a separate `coolingOff` map, not an entry in `quarantined`. - A direct spawn refusal says the credential is cooling off after repeated backend errors and gives the remaining seconds. Do not change `FleetHealthMonitor.coverage`. It still describes the periodic health webhook path. Lead-pane outage delivery is a separate capability. ### Files owned - `fleetd/src/main/java/dev/ltms/fleet/Fleetd.java` - `fleetd/src/main/java/dev/ltms/fleet/config/FleetConfig.java` - `fleetd/src/main/java/dev/ltms/fleet/config/ConfigRef.java` - `fleetd/src/main/java/dev/ltms/fleet/member/CompositePeerLauncher.java` - `fleetd/src/main/java/dev/ltms/fleet/placement/PlacementContext.java` - `fleetd/src/main/java/dev/ltms/fleet/placement/PlacementPolicyUtil.java` - `fleetd/src/main/java/dev/ltms/fleet/mcp/FleetMcp.java` - `fleetd/src/test/java/dev/ltms/fleet/config/FleetConfigTest.java` - `fleetd/src/test/java/dev/ltms/fleet/config/ConfigRefTest.java` - `fleetd/src/test/java/dev/ltms/fleet/member/CompositePeerLauncherTest.java` - `fleetd/src/test/java/dev/ltms/fleet/placement/PlacementPolicyTest.java` - `fleetd/src/test/java/dev/ltms/fleet/mcp/FleetMcpTest.java` - Add `fleetd/src/test/java/dev/ltms/fleet/BackendOutageFlowTest.java`. - `fleetd/fleetd.example.yaml` - `CLAUDE.md` No earlier unit edits these files. The lead, not a worker, must update `wiki/7-Use-Cases.md`, `wiki/9-Implementation.md`, and `wiki/11-Features.md`. Project rules forbid workers from committing `wiki/`. The portable block in `CLAUDE.md` and `wiki/7-Use-Cases.md` must remain byte-identical. ### Acceptance criteria 1. `errorPattern` binds per profile. Blank uses the legacy default and is reported as degraded coverage. A config reload which changes it is reported as deferred because patterns are compiled at startup. 2. A malformed `errorPattern` stops startup and names `profiles..errorPattern`. 3. The real production sink never silently drops an unknown target. A test captures its error log and the fallback notice call. 4. One real backend-error classification through `CompletionResolver` marks only its member. It does not cool the credential and does not send an outage notice. 5. Two real classifications for one credential within 60 seconds start one incident. 6. The integration test then calls the real explicit-profile spawn gate. It is refused before any adapter spawn call, with “cooling off” and remaining seconds in the message. 7. The same test calls an automatic placement path. A cooling candidate is skipped. If every candidate is cooling, the error names cool-off rather than exhaustion. 8. Profiles sharing the credential are all blocked. A profile on another credential stays usable. 9. `fleet_list` from the same fixture shows both members as `backend_error`, preserves each failure reason, and reports `free: 0`, the credential, and `coolingOffForSeconds`. 10. `fleet_profiles` reports cool-off separately from quarantine. 11. The real `ReplyPushLoop` receives one incident and makes one successful lead-pane send. Existing failed-ticket notice content may share that same combined send. 12. Exhaustion still wins when a line matches both patterns. A simultaneous exhaustion quarantine also wins in spawn errors and fleet views. 13. After the 60-second cool-off, spawn is allowed again. A new incident needs two fresh errors. 14. `healthCoverage` has the same value before and after this change for the same health config. 15. `fleetd.example.yaml` explains `errorPattern`, the legacy fallback, 2/60/60 policy, and the difference between cool-off and exhaustion quarantine. 16. `CLAUDE.md` tells leads how `fleet_profiles` and `fleet_list` report cool-off. The lead later applies the matching wiki updates and runs the documented byte-sync check. 17. The developer records the new end-to-end test failing before implementation, then passing. The focused suites and `mvn clean install` pass. ### Dependencies Unit 5 starts only after Units 1 to 4 are merged or rebased into its branch. It also starts after the #234 defect 2 fix lands, because both areas touch the same target-resolution control path. ### What to report back - The exact commits used for Units 1 to 4 and #234. - The startup coverage line with one configured and one legacy-default profile. - The red test output before the implementation and its green result after. - The explicit and automatic spawn refusal text. - Sample `fleet_list` and `fleet_profiles` JSON for cool-off and exhaustion. - The number and text of lead-pane sends in the real-path test. - The focused test commands and final `mvn clean install` result. - The exact `CLAUDE.md` change and the wiki edits the lead must apply. ## File ownership summary | Area | Unit | Shared edit risk | |---|---:|---| | Completion classification | 1 | Only Unit 1 edits `CompletionResolver` and its test. | | Correlation and cool-off state | 2 | New files only. | | Lead push scheduling | 3 | Only Unit 3 edits `ReplyPushLoop` and its test. | | Member terminal state | 4 | Only Unit 4 edits `MemberSession`, `SessionManager`, and their test. | | Main composition, config, placement, MCP views, shipped prompt | 5 | Only Unit 5 edits `Fleetd`, `FleetConfig`, `CompositePeerLauncher`, placement context, `FleetMcp`, and `CLAUDE.md`. | | Wiki propagation | Lead after Unit 5 | Workers do not commit the wiki submodule. | ## What I would not build 1. **Do not reuse `BackendQuarantine` for outages.** Its repeat call restarts a long credential quarantine. Its fields and errors say “exhausted”. That is wrong for a short outage. 2. **Do not merge the exhaustion and generic error patterns.** Exhaustion must win because it has a different policy and duration. 3. **Do not group by error string.** One outage can produce different text. The shared operational limit is the credential. 4. **Do not mark a profile unusable until config changes.** The current classifier cannot safely tell a permanent malformed request from a transient service fault. A permanent state would need a stronger error taxonomy first. 5. **Do not quarantine on the first generic backend error.** That would turn one bad request or one false pattern match into a fleet-wide capacity loss. 6. **Do not add this to `FleetHealthMonitor`.** The monitor samples slow member health. The exact backend event already exists at completion resolution, and moving it to polling would lose type and time. 7. **Do not add another direct lead injector.** `ReplyPushLoop` already owns status gating, per-lead coalescing, retry bounds, and heartbeat stand-down. 8. **Do not change `healthCoverage` to `full`.** That field still means a webhook notification sink exists for periodic health. A backend outage nudge does not make every health event visible. 9. **Do not persist incident history across daemon restart in this work.** Existing exhaustion quarantine is also in memory. A 60-second state does not justify a new durable store. 10. **Do not build work recovery.** The PR-body survival story proves why checkpoint-first work is useful, but these tickets are about detection, capacity, and signalling. 11. **Do not remove the legacy `API Error:` fallback in the first release.** Doing so would turn an unedited config back into a false successful completion. Report it as degraded coverage instead. 12. **Do not reorder or add the old `visibleTurn` fallback from #201.** #211 already implemented the narrow raw-scrape fallback at `CompletionResolver.classifyRawScrapeFallback`. ## Riskiest assumption and cheapest experiment The riskiest assumption is that a configured error regex means “the backend failed this turn”. The current code and test already show the counterexample: a worker may quote `API Error:` while writing a valid report. Two such false matches on one credential would now remove capacity for 60 seconds. The cheapest experiment is a replay corpus before Unit 5 ships: 1. Save the full pane text from the measured 2026-09-01 outage. 2. Produce one safe failure per backend with a disposable invalid endpoint or request. 3. Save one valid member report which quotes each error line. 4. Replay all samples through the real `CompletionResolver` test fixture. 5. Require outage samples to match and quoted-report samples not to match after assistant-block extraction and baseline checks. This costs no outage deployment and no real sleep. If quoted reports still match, narrow the profile patterns before enabling correlation. Do not raise the threshold to hide a bad classifier. ## Sequencing with three developers First wave: 1. Developer A: Unit 1, typed classification. 2. Developer B: Unit 2, credential outage policy. 3. Developer C: Unit 3, lead outage nudge. As soon as one slot is free, start Unit 4. It is file-disjoint from every first-wave unit. Merge and review Units 1 to 4 independently. Start Unit 5 only after all four foundations and #234 are available. Unit 5 is the only high-conflict integration branch, so no other active unit should touch its file list. ## Checks performed for this refinement - Read issue #201 and issue #227 through their Gitea pages. Both showed zero comments. - Read the source and tests named in the evidence section. - Ran `git status --short --branch`; the branch was clean before this document was added. - Ran `git log --oneline -12` to identify the branch base. - I did not run Maven because this change adds only a design document. - Rendered both Mermaid blocks with `npx @mermaid-js/mermaid-cli`; both commands succeeded.