diff --git a/docs/CB-201-227-Refinement.md b/docs/CB-201-227-Refinement.md new file mode 100644 index 0000000..5610bcc --- /dev/null +++ b/docs/CB-201-227-Refinement.md @@ -0,0 +1,531 @@ +# CB-201 and CB-227 refinement + +Date: 2026-09-03 + +## Decision + +#201 and #227 are one delivery program, but they are not one implementation unit. + +#201 has a real seam: `CompletionResolver` can publish a typed backend-error event only after its +waiter resolution wins. #227 can consume that event without knowing any pane text. The classifier +must land before the final #227 wiring. However, the policy engine, roster state, and lead nudge can +be built in parallel with the classifier. + +I propose five units. Units 1 to 4 own separate files and can run in parallel. Unit 5 owns all +composition files and lands after them. It also depends on the #234 defect 2 fix named in the task. + +```mermaid +flowchart LR + U1["Unit 1: typed backend-error classification"] --> U5["Unit 5: wire policy, spawn gate, and fleet views"] + U2["Unit 2: credential outage policy"] --> U5 + U3["Unit 3: lead outage nudge"] --> U5 + U4["Unit 4: durable member outcome"] --> U5 + D234["#234 defect 2: fail-loud target resolution"] --> U5 +``` + +*Figure 1. Four file-disjoint foundations feed one composition unit.* + +This split keeps `Fleetd.java` under one owner. It also keeps every other production file under one +unit in this plan. + +## Evidence checked in the current branch + +I read both issue pages in full. Each page reports zero comments. + +| Evidence | What the code says now | +|---|---| +| `inject/CompletionResolver.java:229-237` | A turn below two seconds fails before normal scrape classification. A matching fast backend error is therefore only a generic failure today. | +| `inject/CompletionResolver.java:260-276` and `:332-367` | #211 already added raw-screen classification when `lastAssistantBlock` is empty. The “dead-code question” in #201 is stale on this branch. | +| `inject/CompletionResolver.java:288-317` | Exhaustion wins before the hard-coded `API Error:` match. A backend error then goes through generic `fail(...)`. | +| `inject/CompletionResolver.java:311-316` | The code admits that the pattern is a heuristic. A member report which quotes an API error may match it. | +| `inject/CompletionResolver.java:449-467` | Startup coverage exists only for `exhaustedPattern`. | +| `inject/ExhaustedPatternLookup.java:13-25` | The current lookup and explicit `none()` value are a good shape for the new classifier seam. | +| `Fleetd.java:196-207` | One `BackendQuarantine` is shared by placement and the exhaustion sink. Its cooldown comes from `quarantineCooldownSeconds`. | +| `Fleetd.java:322-363` | Pattern compilation, target-to-profile lookup, and the live `ExhaustionSink` are composed in `Fleetd.main`. The sink on this branch still ends in `.ifPresent(...)`. This plan assumes #234 replaces that silent path. | +| `placement/BackendQuarantine.java:60-87` | A repeated exhaustion restarts one long quarantine. The store is credential-keyed and uses an injected monotonic clock. | +| `member/CompositePeerLauncher.java:260-317` | Explicit and policy-selected spawns have separate gates. Both paths must learn about outage cool-off. | +| `member/CompositePeerLauncher.java:347-379` | Exhaustion refusal already checks a credential for explicit spawns and filters policy candidates. Its error text says “exhausted”. | +| `placement/PlacementContext.java:10-22` and `PlacementPolicyUtil.java:14-83` | Automatic placement has only one transient exclusion set named `quarantined`. Reusing it would make outage errors say “backend exhausted”. | +| `mcp/FleetMcp.java:913-1025` | `fleet_list` sets `free: 0` and adds `credentialId` plus `quarantinedForSeconds` when quarantine is active. | +| `session/MemberSession.java:51-59` | The roster has `DONE` and generic `FAILED`, but no backend-error state or stored reason. | +| `session/SessionManager.java:695-773` | A normal boundary moves `BUSY` to `DONE`. A failure moves any non-released session to `FAILED`. The async completion resolver can race the `DONE` update. | +| `session/SessionManager.java:648-687` | `rosterView` reports the session state, but it reports no terminal reason. | +| `msg/MessageService.java:922-940` | CB-588 already nudges for every terminal async ticket, including failures. Current code would report failed tickets, but it would not report one correlated outage. | +| `msg/ReplyPushLoop.java:20-48` | Replies, terminal tickets, and questions share one per-lead schedule. This prevents two push sources from injecting competing turns. | +| `msg/ReplyPushLoop.java:305-395` | Each push entry point resolves the owning lead through `PrimaryRegistry`. Missing ownership is logged and the durable or pending item remains the backstop. | +| `msg/ReplyPushLoop.java:496-547` | One tick builds one combined nudge. Pending items have separate reminder counts. | +| `health/FleetHealthMonitor.java:91-143` | Health is a slow periodic observer of members and message-layer facts. It does not receive completion classifications. | +| `health/FleetHealthMonitor.java:206-208` | `healthCoverage` means health enabled plus webhook configured. It does not describe lead-pane alerts. | +| `Fleetd.java:465-486` | Health stays `detection-only` without the webhook notification setting. | + +I also read the related unit tests for `CompletionResolver`, `ReplyPushLoop`, `BackendQuarantine`, +`CompositePeerLauncher`, `PlacementPolicyUtil`, `SessionManager`, `MessageService`, and `FleetMcp`. + +I did not inspect the in-progress #234 branch. I only used the two measured facts in the task. No +peer architect was named, so I did not exchange a design with one. + +## Required behaviour + +The policy should use these first values: + +- Threshold: **2** classified backend errors. +- Window: **60 seconds**, measured from the first error to the second. +- Cool-off: **60 seconds**, starting when the threshold is reached. +- Correlation key: `credentialId`, never profile name and never error text. +- Incident rule: one active incident per credential. Errors during its cool-off do not extend it and + do not create more lead notices. +- Rearm rule: after cool-off ends, two fresh errors are needed for another incident. + +Two errors are the smallest threshold which protects the honest one-turn failure. A 60-second window +fits the measured two-member outage. A 60-second cool-off blocks immediate repeat spawns without +turning a short backend fault into the default 1,800-second exhaustion quarantine. + +A single classified error still fails its send and marks its member `backend_error`. It does not +cool a credential and does not send an outage notice. This is what “a single error changes nothing” +must mean at the credential level. It cannot mean that the failed member still looks successful. + +```mermaid +sequenceDiagram + participant R1 as Resolver for member A + participant R2 as Resolver for member B + participant P as Outage policy + participant S as Spawn gate + participant N as Lead push loop + participant L as Lead pane + + R1->>P: backend error for credential C + Note over P: Count 1, no cool-off + R2->>P: backend error for credential C within 60s + P->>P: Start one 60s incident + P->>S: Credential C is cooling off + P->>N: Queue one incident notice + N->>L: Inject when lead is idle, blocked, or done + L->>S: Request another spawn on credential C + S-->>L: Refuse and report remaining cool-off +``` + +*Figure 2. The second independent classification creates the fleet-level event.* + +Against the 2026-09-01 case, the second failed member would start cool-off. `fleet_list` would show +zero free capacity and both members as `backend_error`. The push loop would inject one outage notice +even if the lead had not polled either ticket yet. The design reports the outage. It does not recover +uncommitted work from the members. + +## Unit 1 — Typed backend-error classification + +### Scope + +Replace the direct hard-coded check inside `CompletionResolver` with a lookup and a sink. Keep the +public send result as a failed send. The typed internal event is the seam #227 consumes. + +The lookup returns the pattern for a target. The sink receives the target, matched line, and full +failure reason. It fires only after `Rendezvous.resolveFailure(...)` wins for that exact captured +waiter. This copies the race rule already used by `ExhaustionSink`. + +The classifier must run in all three current paths: + +1. a normal non-empty assistant block; +2. the #211 raw scrape fallback; +3. a turn inside `MIN_TURN_NANOS`, before it becomes a generic too-fast failure. + +In every path, the order stays: stale-baseline guard, exhaustion, backend error, then generic +failure or completion. A fast turn still fails when no configured pattern matches. + +Keep `(?i)\bAPI Error\s*:` as a compatibility pattern for profiles without `errorPattern` until the +operator config is updated. Do not call this full coverage. Startup reporting in Unit 5 must name +profiles using this weaker legacy default. + +### Files owned + +- Add `fleetd/src/main/java/dev/ltms/fleet/inject/BackendErrorPatternLookup.java`. +- Add `fleetd/src/main/java/dev/ltms/fleet/inject/BackendErrorSink.java`. +- Change `fleetd/src/main/java/dev/ltms/fleet/inject/CompletionResolver.java`. +- Change `fleetd/src/test/java/dev/ltms/fleet/inject/CompletionResolverTest.java`. + +No other unit may edit these files. + +### Acceptance criteria + +1. A target-specific error pattern matches a normal assistant block and resolves the send as failed. +2. The same match calls `BackendErrorSink` exactly once after the waiter resolution wins. +3. A late classification which loses to `fleet_reply` does not call the sink. +4. An exhausted line that also matches the generic error pattern stays `BACKEND_EXHAUSTED`. It calls + only `ExhaustionSink`. +5. A raw pane with leading Terminal User Interface (TUI) chrome and no assistant marker still uses + the #211 fallback and calls the backend-error sink. +6. A matching error inside the two-second floor is typed and sent to the sink. A non-matching fast + turn stays a generic failure. +7. An unchanged delivery baseline which contains old backend-error text is suppressed. It never + increments outage evidence. +8. A non-match keeps the existing completion result and text. +9. Constructors used by current callers keep compiling. They use the legacy default lookup and an + explicit inert sink until Unit 5 supplies the production objects. +10. Unit tests pass. The developer runs the focused test first, then `mvn clean install` from + `fleetd/`. + +### Dependencies + +None. Unit 1 can run with Units 2, 3, and 4. + +Unit 5 depends on its new lookup, sink, and constructor. + +### What to report back + +- The exact classifier order in all three paths. +- The focused test command and result. +- The test which proves a losing waiter race does not publish an event. +- The test which proves a fast matching failure is typed. +- The final `mvn clean install` result. +- Any constructor kept only for transition and where Unit 5 replaces it. + +## Unit 2 — Credential outage policy + +### Scope + +Build a small credential-keyed state machine. It accepts already-classified backend-error events. +It does not read pane text, profiles, sessions, or lead state. + +Use an injected monotonic clock. A call records `credentialId`, target, and reason. It returns a new +incident only on the threshold crossing. The incident contains a stable event id, credential id, +the distinct affected targets, evidence count, window, and remaining cool-off. + +This class owns both correlation and short cool-off. Keeping them together makes threshold crossing +and the cool-off deadline one atomic state change. + +### Files owned + +- Add `fleetd/src/main/java/dev/ltms/fleet/placement/BackendOutagePolicy.java`. +- Add `fleetd/src/test/java/dev/ltms/fleet/placement/BackendOutagePolicyTest.java`. + +No other unit may edit these files. + +### Acceptance criteria + +1. One error creates no incident and no cool-off. +2. Two errors for one credential within 60 seconds create exactly one incident and a 60-second + cool-off. +3. Two errors more than 60 seconds apart do not create an incident. +4. The exact 60-second boundary has a pinned result. Use inclusive `<= 60s` so scheduler delay does + not discard evidence at the boundary. +5. Different credentials never share evidence. +6. Different profiles which supply the same credential id do share evidence. The policy itself only + sees the credential id. +7. More errors during active cool-off do not extend its deadline and do not return another incident. +8. After expiry, old evidence is cleared. Two fresh errors are needed to create the next incident. +9. Remaining seconds round up, matching `BackendQuarantine` reporting. +10. Concurrent second and third errors cannot return two incidents. +11. The focused tests and `mvn clean install` pass. + +### Dependencies + +None. Unit 2 can run with Units 1, 3, and 4. + +Unit 5 depends on the policy API. + +### What to report back + +- The state transition table and locking method. +- The exact threshold, window, cool-off, and boundary rule. +- The test which proves one incident under concurrent calls. +- The focused test and `mvn clean install` results. + +## Unit 3 — Lead outage nudge + +### Scope + +Add backend incidents as a fourth pending source in `ReplyPushLoop`. Do not create another scheduler +or call `AgentControl.send` from `Fleetd`. The existing combined per-lead schedule is the control +which prevents competing injected turns. + +The entry point takes an incident id, affected worker targets, credential id, affected profile +names, and remaining cool-off. It resolves distinct owning leads through `PrimaryRegistry`. + +Each `(incidentId, lead)` item is one-shot. It waits while the lead is not injectable. After one +successful `agents.send`, remove it. A send exception keeps it pending for a bounded retry. It never +uses the repeated reminder behaviour of an uncollected ticket. + +Also add a fail-loud entry point for a classified target that Unit 5 cannot map to a credential. It +uses `PrimaryRegistry.nudgeTargetFor(target)` and says that correlation could not run. If no lead is +known, log at `WARN`, not `DEBUG`. + +### Files owned + +- Change `fleetd/src/main/java/dev/ltms/fleet/msg/ReplyPushLoop.java`. +- Change `fleetd/src/test/java/dev/ltms/fleet/msg/ReplyPushLoopTest.java`. + +No other unit may edit these files. + +### Acceptance criteria + +1. One incident affecting two workers owned by one lead causes one successful pane injection. +2. Two affected workers owned by two leads cause one successful injection per affected lead. This + is one notice per event per lead, not one notice per member. +3. Repeating the same incident id is idempotent. +4. A busy or unknown lead is not injected. The item stays pending until the lead becomes injectable + or its attempt cap is reached. +5. After one successful injection, later ticks do not mention that incident again. +6. A failed `agents.send` is retried within the existing bound. A successful retry still gives only + one successful send. +7. A pending ticket and an outage incident for one lead appear in one combined nudge, not two + competing turns. +8. The text names the credential, profiles, affected workers, and remaining cool-off. It tells the + lead to run `fleet_list`. +9. An unmapped target produces a direct warning notice when a lead is known. If no lead is known, + the code logs a `WARN` naming the target and reason. +10. `stop()` clears incident state as it clears other push state. +11. Existing reply, ticket, and question tests stay green. The focused tests and + `mvn clean install` pass. + +### Dependencies + +None. The API uses plain values, not the Unit 2 incident class. This lets Unit 3 run in parallel. + +Unit 5 adapts the Unit 2 incident into this entry point. + +### What to report back + +- The exact one-shot and retry rules. +- The test showing one combined nudge with a failed ticket. +- The test showing one successful send for two affected workers. +- The focused test and `mvn clean install` results. + +## Unit 4 — Durable member backend-error outcome + +### Scope + +Make a classified backend failure remain visible after its ticket is collected or expires. + +Add `BACKEND_ERROR` to `MemberSession.State`. Add a nullable failure detail to `MemberSession` and +render it as `failureReason` in `SessionManager.rosterView`. Add +`SessionManager.onBackendError(target, reason)`. + +The transition must handle both completion orderings: + +- `BUSY -> BACKEND_ERROR` when classification wins before the normal completion state update; +- `DONE -> BACKEND_ERROR` when the async resolver runs after `SessionManager.onTurnComplete`. + +It must use a compare-and-set retry or another atomic update. `RELEASED` must never return to the +roster. A backend-error member is terminal and cannot accept another delivery. + +### Files owned + +- Change `fleetd/src/main/java/dev/ltms/fleet/session/MemberSession.java`. +- Change `fleetd/src/main/java/dev/ltms/fleet/session/SessionManager.java`. +- Change `fleetd/src/test/java/dev/ltms/fleet/session/SessionManagerTest.java`. + +No other unit may edit these files. + +### Acceptance criteria + +1. `onBackendError` moves a `BUSY` member to `BACKEND_ERROR` and stores the reason. +2. It also moves `DONE` to `BACKEND_ERROR`, covering the resolver race. +3. A later normal `onTurnComplete` cannot change `BACKEND_ERROR` back to `DONE`. +4. A released or unknown member is not recreated. The unknown case logs at `WARN` and returns an + explicit false result to its caller. +5. `onDelivered` refuses a `BACKEND_ERROR` member, just as it refuses generic `FAILED`. +6. `rosterView` reports `state: backend_error` and `failureReason` after the send ticket is gone. +7. Ordinary members do not gain a blank or invented `failureReason` field. +8. Existing constructors keep source compatibility for tests and adapters. +9. The focused tests and `mvn clean install` pass. + +### Dependencies + +None. Unit 4 can run with Units 1, 2, and 3. + +Unit 5 calls the new session method from the production sink. + +### What to report back + +- The two race orderings and the tests for both. +- The exact roster JSON shape. +- The unknown-target result and log level. +- The focused test and `mvn clean install` results. + +## Unit 5 — Production wiring, spawn gate, and fleet views + +### Scope + +Compose Units 1 to 4 in production. This is the only unit which edits `Fleetd.java`. + +Add per-profile `errorPattern` config beside `exhaustedPattern`. Compile both once at startup. A +configured pattern wins over the legacy default. Report configured profiles and legacy-default +profiles separately at startup. A bad regex must stop startup with the profile and key in the +message. + +Wire one production `BackendErrorSink` with this order: + +1. mark the member `backend_error` with its reason; +2. resolve the profile and its current `effectiveCredentialId()` through the fail-loud #234 seam; +3. record the error in `BackendOutagePolicy`; +4. on a new incident, submit one event to `ReplyPushLoop`. + +If target metadata cannot be resolved, do not end in `Optional.ifPresent`. Log an error and call the +Unit 3 unmapped-target notice. The failed send still reaches its ticket through CB-588. + +Teach both spawn paths about a separate cool-off source. Exhaustion quarantine has priority when +both states are active. Automatic placement needs a distinct `coolingOff` set so its refusal does +not say “exhausted”. + +Extend the MCP (Model Context Protocol) views: + +- A cooling profile has `free: 0`, `credentialId`, and `coolingOffForSeconds` in `fleet_list`. +- It does not have `quarantinedForSeconds` unless exhaustion quarantine is also active. +- `fleet_profiles` has a separate `coolingOff` map, not an entry in `quarantined`. +- A direct spawn refusal says the credential is cooling off after repeated backend errors and gives + the remaining seconds. + +Do not change `FleetHealthMonitor.coverage`. It still describes the periodic health webhook path. +Lead-pane outage delivery is a separate capability. + +### Files owned + +- `fleetd/src/main/java/dev/ltms/fleet/Fleetd.java` +- `fleetd/src/main/java/dev/ltms/fleet/config/FleetConfig.java` +- `fleetd/src/main/java/dev/ltms/fleet/config/ConfigRef.java` +- `fleetd/src/main/java/dev/ltms/fleet/member/CompositePeerLauncher.java` +- `fleetd/src/main/java/dev/ltms/fleet/placement/PlacementContext.java` +- `fleetd/src/main/java/dev/ltms/fleet/placement/PlacementPolicyUtil.java` +- `fleetd/src/main/java/dev/ltms/fleet/mcp/FleetMcp.java` +- `fleetd/src/test/java/dev/ltms/fleet/config/FleetConfigTest.java` +- `fleetd/src/test/java/dev/ltms/fleet/config/ConfigRefTest.java` +- `fleetd/src/test/java/dev/ltms/fleet/member/CompositePeerLauncherTest.java` +- `fleetd/src/test/java/dev/ltms/fleet/placement/PlacementPolicyTest.java` +- `fleetd/src/test/java/dev/ltms/fleet/mcp/FleetMcpTest.java` +- Add `fleetd/src/test/java/dev/ltms/fleet/BackendOutageFlowTest.java`. +- `fleetd/fleetd.example.yaml` +- `CLAUDE.md` + +No earlier unit edits these files. + +The lead, not a worker, must update `wiki/7-Use-Cases.md`, `wiki/9-Implementation.md`, and +`wiki/11-Features.md`. Project rules forbid workers from committing `wiki/`. The portable block in +`CLAUDE.md` and `wiki/7-Use-Cases.md` must remain byte-identical. + +### Acceptance criteria + +1. `errorPattern` binds per profile. Blank uses the legacy default and is reported as degraded + coverage. A config reload which changes it is reported as deferred because patterns are compiled + at startup. +2. A malformed `errorPattern` stops startup and names `profiles..errorPattern`. +3. The real production sink never silently drops an unknown target. A test captures its error log + and the fallback notice call. +4. One real backend-error classification through `CompletionResolver` marks only its member. It does + not cool the credential and does not send an outage notice. +5. Two real classifications for one credential within 60 seconds start one incident. +6. The integration test then calls the real explicit-profile spawn gate. It is refused before any + adapter spawn call, with “cooling off” and remaining seconds in the message. +7. The same test calls an automatic placement path. A cooling candidate is skipped. If every + candidate is cooling, the error names cool-off rather than exhaustion. +8. Profiles sharing the credential are all blocked. A profile on another credential stays usable. +9. `fleet_list` from the same fixture shows both members as `backend_error`, preserves each failure + reason, and reports `free: 0`, the credential, and `coolingOffForSeconds`. +10. `fleet_profiles` reports cool-off separately from quarantine. +11. The real `ReplyPushLoop` receives one incident and makes one successful lead-pane send. Existing + failed-ticket notice content may share that same combined send. +12. Exhaustion still wins when a line matches both patterns. A simultaneous exhaustion quarantine + also wins in spawn errors and fleet views. +13. After the 60-second cool-off, spawn is allowed again. A new incident needs two fresh errors. +14. `healthCoverage` has the same value before and after this change for the same health config. +15. `fleetd.example.yaml` explains `errorPattern`, the legacy fallback, 2/60/60 policy, and the + difference between cool-off and exhaustion quarantine. +16. `CLAUDE.md` tells leads how `fleet_profiles` and `fleet_list` report cool-off. The lead later + applies the matching wiki updates and runs the documented byte-sync check. +17. The developer records the new end-to-end test failing before implementation, then passing. The + focused suites and `mvn clean install` pass. + +### Dependencies + +Unit 5 starts only after Units 1 to 4 are merged or rebased into its branch. It also starts after the +#234 defect 2 fix lands, because both areas touch the same target-resolution control path. + +### What to report back + +- The exact commits used for Units 1 to 4 and #234. +- The startup coverage line with one configured and one legacy-default profile. +- The red test output before the implementation and its green result after. +- The explicit and automatic spawn refusal text. +- Sample `fleet_list` and `fleet_profiles` JSON for cool-off and exhaustion. +- The number and text of lead-pane sends in the real-path test. +- The focused test commands and final `mvn clean install` result. +- The exact `CLAUDE.md` change and the wiki edits the lead must apply. + +## File ownership summary + +| Area | Unit | Shared edit risk | +|---|---:|---| +| Completion classification | 1 | Only Unit 1 edits `CompletionResolver` and its test. | +| Correlation and cool-off state | 2 | New files only. | +| Lead push scheduling | 3 | Only Unit 3 edits `ReplyPushLoop` and its test. | +| Member terminal state | 4 | Only Unit 4 edits `MemberSession`, `SessionManager`, and their test. | +| Main composition, config, placement, MCP views, shipped prompt | 5 | Only Unit 5 edits `Fleetd`, `FleetConfig`, `CompositePeerLauncher`, placement context, `FleetMcp`, and `CLAUDE.md`. | +| Wiki propagation | Lead after Unit 5 | Workers do not commit the wiki submodule. | + +## What I would not build + +1. **Do not reuse `BackendQuarantine` for outages.** Its repeat call restarts a long credential + quarantine. Its fields and errors say “exhausted”. That is wrong for a short outage. +2. **Do not merge the exhaustion and generic error patterns.** Exhaustion must win because it has a + different policy and duration. +3. **Do not group by error string.** One outage can produce different text. The shared operational + limit is the credential. +4. **Do not mark a profile unusable until config changes.** The current classifier cannot safely + tell a permanent malformed request from a transient service fault. A permanent state would need + a stronger error taxonomy first. +5. **Do not quarantine on the first generic backend error.** That would turn one bad request or one + false pattern match into a fleet-wide capacity loss. +6. **Do not add this to `FleetHealthMonitor`.** The monitor samples slow member health. The exact + backend event already exists at completion resolution, and moving it to polling would lose type + and time. +7. **Do not add another direct lead injector.** `ReplyPushLoop` already owns status gating, + per-lead coalescing, retry bounds, and heartbeat stand-down. +8. **Do not change `healthCoverage` to `full`.** That field still means a webhook notification sink + exists for periodic health. A backend outage nudge does not make every health event visible. +9. **Do not persist incident history across daemon restart in this work.** Existing exhaustion + quarantine is also in memory. A 60-second state does not justify a new durable store. +10. **Do not build work recovery.** The PR-body survival story proves why checkpoint-first work is + useful, but these tickets are about detection, capacity, and signalling. +11. **Do not remove the legacy `API Error:` fallback in the first release.** Doing so would turn an + unedited config back into a false successful completion. Report it as degraded coverage instead. +12. **Do not reorder or add the old `visibleTurn` fallback from #201.** #211 already implemented the + narrow raw-scrape fallback at `CompletionResolver.classifyRawScrapeFallback`. + +## Riskiest assumption and cheapest experiment + +The riskiest assumption is that a configured error regex means “the backend failed this turn”. The +current code and test already show the counterexample: a worker may quote `API Error:` while writing +a valid report. Two such false matches on one credential would now remove capacity for 60 seconds. + +The cheapest experiment is a replay corpus before Unit 5 ships: + +1. Save the full pane text from the measured 2026-09-01 outage. +2. Produce one safe failure per backend with a disposable invalid endpoint or request. +3. Save one valid member report which quotes each error line. +4. Replay all samples through the real `CompletionResolver` test fixture. +5. Require outage samples to match and quoted-report samples not to match after assistant-block + extraction and baseline checks. + +This costs no outage deployment and no real sleep. If quoted reports still match, narrow the profile +patterns before enabling correlation. Do not raise the threshold to hide a bad classifier. + +## Sequencing with three developers + +First wave: + +1. Developer A: Unit 1, typed classification. +2. Developer B: Unit 2, credential outage policy. +3. Developer C: Unit 3, lead outage nudge. + +As soon as one slot is free, start Unit 4. It is file-disjoint from every first-wave unit. Merge and +review Units 1 to 4 independently. + +Start Unit 5 only after all four foundations and #234 are available. Unit 5 is the only high-conflict +integration branch, so no other active unit should touch its file list. + +## Checks performed for this refinement + +- Read issue #201 and issue #227 through their Gitea pages. Both showed zero comments. +- Read the source and tests named in the evidence section. +- Ran `git status --short --branch`; the branch was clean before this document was added. +- Ran `git log --oneline -12` to identify the branch base. +- I did not run Maven because this change adds only a design document. +- Rendered both Mermaid blocks with `npx @mermaid-js/mermaid-cli`; both commands succeeded.