532 lines
28 KiB
Markdown
532 lines
28 KiB
Markdown
# CB-201 and CB-227 refinement
|
|
|
|
Date: 2026-09-03
|
|
|
|
## Decision
|
|
|
|
#201 and #227 are one delivery program, but they are not one implementation unit.
|
|
|
|
#201 has a real seam: `CompletionResolver` can publish a typed backend-error event only after its
|
|
waiter resolution wins. #227 can consume that event without knowing any pane text. The classifier
|
|
must land before the final #227 wiring. However, the policy engine, roster state, and lead nudge can
|
|
be built in parallel with the classifier.
|
|
|
|
I propose five units. Units 1 to 4 own separate files and can run in parallel. Unit 5 owns all
|
|
composition files and lands after them. It also depends on the #234 defect 2 fix named in the task.
|
|
|
|
```mermaid
|
|
flowchart LR
|
|
U1["Unit 1: typed backend-error classification"] --> U5["Unit 5: wire policy, spawn gate, and fleet views"]
|
|
U2["Unit 2: credential outage policy"] --> U5
|
|
U3["Unit 3: lead outage nudge"] --> U5
|
|
U4["Unit 4: durable member outcome"] --> U5
|
|
D234["#234 defect 2: fail-loud target resolution"] --> U5
|
|
```
|
|
|
|
*Figure 1. Four file-disjoint foundations feed one composition unit.*
|
|
|
|
This split keeps `Fleetd.java` under one owner. It also keeps every other production file under one
|
|
unit in this plan.
|
|
|
|
## Evidence checked in the current branch
|
|
|
|
I read both issue pages in full. Each page reports zero comments.
|
|
|
|
| Evidence | What the code says now |
|
|
|---|---|
|
|
| `inject/CompletionResolver.java:229-237` | A turn below two seconds fails before normal scrape classification. A matching fast backend error is therefore only a generic failure today. |
|
|
| `inject/CompletionResolver.java:260-276` and `:332-367` | #211 already added raw-screen classification when `lastAssistantBlock` is empty. The “dead-code question” in #201 is stale on this branch. |
|
|
| `inject/CompletionResolver.java:288-317` | Exhaustion wins before the hard-coded `API Error:` match. A backend error then goes through generic `fail(...)`. |
|
|
| `inject/CompletionResolver.java:311-316` | The code admits that the pattern is a heuristic. A member report which quotes an API error may match it. |
|
|
| `inject/CompletionResolver.java:449-467` | Startup coverage exists only for `exhaustedPattern`. |
|
|
| `inject/ExhaustedPatternLookup.java:13-25` | The current lookup and explicit `none()` value are a good shape for the new classifier seam. |
|
|
| `Fleetd.java:196-207` | One `BackendQuarantine` is shared by placement and the exhaustion sink. Its cooldown comes from `quarantineCooldownSeconds`. |
|
|
| `Fleetd.java:322-363` | Pattern compilation, target-to-profile lookup, and the live `ExhaustionSink` are composed in `Fleetd.main`. The sink on this branch still ends in `.ifPresent(...)`. This plan assumes #234 replaces that silent path. |
|
|
| `placement/BackendQuarantine.java:60-87` | A repeated exhaustion restarts one long quarantine. The store is credential-keyed and uses an injected monotonic clock. |
|
|
| `member/CompositePeerLauncher.java:260-317` | Explicit and policy-selected spawns have separate gates. Both paths must learn about outage cool-off. |
|
|
| `member/CompositePeerLauncher.java:347-379` | Exhaustion refusal already checks a credential for explicit spawns and filters policy candidates. Its error text says “exhausted”. |
|
|
| `placement/PlacementContext.java:10-22` and `PlacementPolicyUtil.java:14-83` | Automatic placement has only one transient exclusion set named `quarantined`. Reusing it would make outage errors say “backend exhausted”. |
|
|
| `mcp/FleetMcp.java:913-1025` | `fleet_list` sets `free: 0` and adds `credentialId` plus `quarantinedForSeconds` when quarantine is active. |
|
|
| `session/MemberSession.java:51-59` | The roster has `DONE` and generic `FAILED`, but no backend-error state or stored reason. |
|
|
| `session/SessionManager.java:695-773` | A normal boundary moves `BUSY` to `DONE`. A failure moves any non-released session to `FAILED`. The async completion resolver can race the `DONE` update. |
|
|
| `session/SessionManager.java:648-687` | `rosterView` reports the session state, but it reports no terminal reason. |
|
|
| `msg/MessageService.java:922-940` | CB-588 already nudges for every terminal async ticket, including failures. Current code would report failed tickets, but it would not report one correlated outage. |
|
|
| `msg/ReplyPushLoop.java:20-48` | Replies, terminal tickets, and questions share one per-lead schedule. This prevents two push sources from injecting competing turns. |
|
|
| `msg/ReplyPushLoop.java:305-395` | Each push entry point resolves the owning lead through `PrimaryRegistry`. Missing ownership is logged and the durable or pending item remains the backstop. |
|
|
| `msg/ReplyPushLoop.java:496-547` | One tick builds one combined nudge. Pending items have separate reminder counts. |
|
|
| `health/FleetHealthMonitor.java:91-143` | Health is a slow periodic observer of members and message-layer facts. It does not receive completion classifications. |
|
|
| `health/FleetHealthMonitor.java:206-208` | `healthCoverage` means health enabled plus webhook configured. It does not describe lead-pane alerts. |
|
|
| `Fleetd.java:465-486` | Health stays `detection-only` without the webhook notification setting. |
|
|
|
|
I also read the related unit tests for `CompletionResolver`, `ReplyPushLoop`, `BackendQuarantine`,
|
|
`CompositePeerLauncher`, `PlacementPolicyUtil`, `SessionManager`, `MessageService`, and `FleetMcp`.
|
|
|
|
I did not inspect the in-progress #234 branch. I only used the two measured facts in the task. No
|
|
peer architect was named, so I did not exchange a design with one.
|
|
|
|
## Required behaviour
|
|
|
|
The policy should use these first values:
|
|
|
|
- Threshold: **2** classified backend errors.
|
|
- Window: **60 seconds**, measured from the first error to the second.
|
|
- Cool-off: **60 seconds**, starting when the threshold is reached.
|
|
- Correlation key: `credentialId`, never profile name and never error text.
|
|
- Incident rule: one active incident per credential. Errors during its cool-off do not extend it and
|
|
do not create more lead notices.
|
|
- Rearm rule: after cool-off ends, two fresh errors are needed for another incident.
|
|
|
|
Two errors are the smallest threshold which protects the honest one-turn failure. A 60-second window
|
|
fits the measured two-member outage. A 60-second cool-off blocks immediate repeat spawns without
|
|
turning a short backend fault into the default 1,800-second exhaustion quarantine.
|
|
|
|
A single classified error still fails its send and marks its member `backend_error`. It does not
|
|
cool a credential and does not send an outage notice. This is what “a single error changes nothing”
|
|
must mean at the credential level. It cannot mean that the failed member still looks successful.
|
|
|
|
```mermaid
|
|
sequenceDiagram
|
|
participant R1 as Resolver for member A
|
|
participant R2 as Resolver for member B
|
|
participant P as Outage policy
|
|
participant S as Spawn gate
|
|
participant N as Lead push loop
|
|
participant L as Lead pane
|
|
|
|
R1->>P: backend error for credential C
|
|
Note over P: Count 1, no cool-off
|
|
R2->>P: backend error for credential C within 60s
|
|
P->>P: Start one 60s incident
|
|
P->>S: Credential C is cooling off
|
|
P->>N: Queue one incident notice
|
|
N->>L: Inject when lead is idle, blocked, or done
|
|
L->>S: Request another spawn on credential C
|
|
S-->>L: Refuse and report remaining cool-off
|
|
```
|
|
|
|
*Figure 2. The second independent classification creates the fleet-level event.*
|
|
|
|
Against the 2026-09-01 case, the second failed member would start cool-off. `fleet_list` would show
|
|
zero free capacity and both members as `backend_error`. The push loop would inject one outage notice
|
|
even if the lead had not polled either ticket yet. The design reports the outage. It does not recover
|
|
uncommitted work from the members.
|
|
|
|
## Unit 1 — Typed backend-error classification
|
|
|
|
### Scope
|
|
|
|
Replace the direct hard-coded check inside `CompletionResolver` with a lookup and a sink. Keep the
|
|
public send result as a failed send. The typed internal event is the seam #227 consumes.
|
|
|
|
The lookup returns the pattern for a target. The sink receives the target, matched line, and full
|
|
failure reason. It fires only after `Rendezvous.resolveFailure(...)` wins for that exact captured
|
|
waiter. This copies the race rule already used by `ExhaustionSink`.
|
|
|
|
The classifier must run in all three current paths:
|
|
|
|
1. a normal non-empty assistant block;
|
|
2. the #211 raw scrape fallback;
|
|
3. a turn inside `MIN_TURN_NANOS`, before it becomes a generic too-fast failure.
|
|
|
|
In every path, the order stays: stale-baseline guard, exhaustion, backend error, then generic
|
|
failure or completion. A fast turn still fails when no configured pattern matches.
|
|
|
|
Keep `(?i)\bAPI Error\s*:` as a compatibility pattern for profiles without `errorPattern` until the
|
|
operator config is updated. Do not call this full coverage. Startup reporting in Unit 5 must name
|
|
profiles using this weaker legacy default.
|
|
|
|
### Files owned
|
|
|
|
- Add `fleetd/src/main/java/dev/ltms/fleet/inject/BackendErrorPatternLookup.java`.
|
|
- Add `fleetd/src/main/java/dev/ltms/fleet/inject/BackendErrorSink.java`.
|
|
- Change `fleetd/src/main/java/dev/ltms/fleet/inject/CompletionResolver.java`.
|
|
- Change `fleetd/src/test/java/dev/ltms/fleet/inject/CompletionResolverTest.java`.
|
|
|
|
No other unit may edit these files.
|
|
|
|
### Acceptance criteria
|
|
|
|
1. A target-specific error pattern matches a normal assistant block and resolves the send as failed.
|
|
2. The same match calls `BackendErrorSink` exactly once after the waiter resolution wins.
|
|
3. A late classification which loses to `fleet_reply` does not call the sink.
|
|
4. An exhausted line that also matches the generic error pattern stays `BACKEND_EXHAUSTED`. It calls
|
|
only `ExhaustionSink`.
|
|
5. A raw pane with leading Terminal User Interface (TUI) chrome and no assistant marker still uses
|
|
the #211 fallback and calls the backend-error sink.
|
|
6. A matching error inside the two-second floor is typed and sent to the sink. A non-matching fast
|
|
turn stays a generic failure.
|
|
7. An unchanged delivery baseline which contains old backend-error text is suppressed. It never
|
|
increments outage evidence.
|
|
8. A non-match keeps the existing completion result and text.
|
|
9. Constructors used by current callers keep compiling. They use the legacy default lookup and an
|
|
explicit inert sink until Unit 5 supplies the production objects.
|
|
10. Unit tests pass. The developer runs the focused test first, then `mvn clean install` from
|
|
`fleetd/`.
|
|
|
|
### Dependencies
|
|
|
|
None. Unit 1 can run with Units 2, 3, and 4.
|
|
|
|
Unit 5 depends on its new lookup, sink, and constructor.
|
|
|
|
### What to report back
|
|
|
|
- The exact classifier order in all three paths.
|
|
- The focused test command and result.
|
|
- The test which proves a losing waiter race does not publish an event.
|
|
- The test which proves a fast matching failure is typed.
|
|
- The final `mvn clean install` result.
|
|
- Any constructor kept only for transition and where Unit 5 replaces it.
|
|
|
|
## Unit 2 — Credential outage policy
|
|
|
|
### Scope
|
|
|
|
Build a small credential-keyed state machine. It accepts already-classified backend-error events.
|
|
It does not read pane text, profiles, sessions, or lead state.
|
|
|
|
Use an injected monotonic clock. A call records `credentialId`, target, and reason. It returns a new
|
|
incident only on the threshold crossing. The incident contains a stable event id, credential id,
|
|
the distinct affected targets, evidence count, window, and remaining cool-off.
|
|
|
|
This class owns both correlation and short cool-off. Keeping them together makes threshold crossing
|
|
and the cool-off deadline one atomic state change.
|
|
|
|
### Files owned
|
|
|
|
- Add `fleetd/src/main/java/dev/ltms/fleet/placement/BackendOutagePolicy.java`.
|
|
- Add `fleetd/src/test/java/dev/ltms/fleet/placement/BackendOutagePolicyTest.java`.
|
|
|
|
No other unit may edit these files.
|
|
|
|
### Acceptance criteria
|
|
|
|
1. One error creates no incident and no cool-off.
|
|
2. Two errors for one credential within 60 seconds create exactly one incident and a 60-second
|
|
cool-off.
|
|
3. Two errors more than 60 seconds apart do not create an incident.
|
|
4. The exact 60-second boundary has a pinned result. Use inclusive `<= 60s` so scheduler delay does
|
|
not discard evidence at the boundary.
|
|
5. Different credentials never share evidence.
|
|
6. Different profiles which supply the same credential id do share evidence. The policy itself only
|
|
sees the credential id.
|
|
7. More errors during active cool-off do not extend its deadline and do not return another incident.
|
|
8. After expiry, old evidence is cleared. Two fresh errors are needed to create the next incident.
|
|
9. Remaining seconds round up, matching `BackendQuarantine` reporting.
|
|
10. Concurrent second and third errors cannot return two incidents.
|
|
11. The focused tests and `mvn clean install` pass.
|
|
|
|
### Dependencies
|
|
|
|
None. Unit 2 can run with Units 1, 3, and 4.
|
|
|
|
Unit 5 depends on the policy API.
|
|
|
|
### What to report back
|
|
|
|
- The state transition table and locking method.
|
|
- The exact threshold, window, cool-off, and boundary rule.
|
|
- The test which proves one incident under concurrent calls.
|
|
- The focused test and `mvn clean install` results.
|
|
|
|
## Unit 3 — Lead outage nudge
|
|
|
|
### Scope
|
|
|
|
Add backend incidents as a fourth pending source in `ReplyPushLoop`. Do not create another scheduler
|
|
or call `AgentControl.send` from `Fleetd`. The existing combined per-lead schedule is the control
|
|
which prevents competing injected turns.
|
|
|
|
The entry point takes an incident id, affected worker targets, credential id, affected profile
|
|
names, and remaining cool-off. It resolves distinct owning leads through `PrimaryRegistry`.
|
|
|
|
Each `(incidentId, lead)` item is one-shot. It waits while the lead is not injectable. After one
|
|
successful `agents.send`, remove it. A send exception keeps it pending for a bounded retry. It never
|
|
uses the repeated reminder behaviour of an uncollected ticket.
|
|
|
|
Also add a fail-loud entry point for a classified target that Unit 5 cannot map to a credential. It
|
|
uses `PrimaryRegistry.nudgeTargetFor(target)` and says that correlation could not run. If no lead is
|
|
known, log at `WARN`, not `DEBUG`.
|
|
|
|
### Files owned
|
|
|
|
- Change `fleetd/src/main/java/dev/ltms/fleet/msg/ReplyPushLoop.java`.
|
|
- Change `fleetd/src/test/java/dev/ltms/fleet/msg/ReplyPushLoopTest.java`.
|
|
|
|
No other unit may edit these files.
|
|
|
|
### Acceptance criteria
|
|
|
|
1. One incident affecting two workers owned by one lead causes one successful pane injection.
|
|
2. Two affected workers owned by two leads cause one successful injection per affected lead. This
|
|
is one notice per event per lead, not one notice per member.
|
|
3. Repeating the same incident id is idempotent.
|
|
4. A busy or unknown lead is not injected. The item stays pending until the lead becomes injectable
|
|
or its attempt cap is reached.
|
|
5. After one successful injection, later ticks do not mention that incident again.
|
|
6. A failed `agents.send` is retried within the existing bound. A successful retry still gives only
|
|
one successful send.
|
|
7. A pending ticket and an outage incident for one lead appear in one combined nudge, not two
|
|
competing turns.
|
|
8. The text names the credential, profiles, affected workers, and remaining cool-off. It tells the
|
|
lead to run `fleet_list`.
|
|
9. An unmapped target produces a direct warning notice when a lead is known. If no lead is known,
|
|
the code logs a `WARN` naming the target and reason.
|
|
10. `stop()` clears incident state as it clears other push state.
|
|
11. Existing reply, ticket, and question tests stay green. The focused tests and
|
|
`mvn clean install` pass.
|
|
|
|
### Dependencies
|
|
|
|
None. The API uses plain values, not the Unit 2 incident class. This lets Unit 3 run in parallel.
|
|
|
|
Unit 5 adapts the Unit 2 incident into this entry point.
|
|
|
|
### What to report back
|
|
|
|
- The exact one-shot and retry rules.
|
|
- The test showing one combined nudge with a failed ticket.
|
|
- The test showing one successful send for two affected workers.
|
|
- The focused test and `mvn clean install` results.
|
|
|
|
## Unit 4 — Durable member backend-error outcome
|
|
|
|
### Scope
|
|
|
|
Make a classified backend failure remain visible after its ticket is collected or expires.
|
|
|
|
Add `BACKEND_ERROR` to `MemberSession.State`. Add a nullable failure detail to `MemberSession` and
|
|
render it as `failureReason` in `SessionManager.rosterView`. Add
|
|
`SessionManager.onBackendError(target, reason)`.
|
|
|
|
The transition must handle both completion orderings:
|
|
|
|
- `BUSY -> BACKEND_ERROR` when classification wins before the normal completion state update;
|
|
- `DONE -> BACKEND_ERROR` when the async resolver runs after `SessionManager.onTurnComplete`.
|
|
|
|
It must use a compare-and-set retry or another atomic update. `RELEASED` must never return to the
|
|
roster. A backend-error member is terminal and cannot accept another delivery.
|
|
|
|
### Files owned
|
|
|
|
- Change `fleetd/src/main/java/dev/ltms/fleet/session/MemberSession.java`.
|
|
- Change `fleetd/src/main/java/dev/ltms/fleet/session/SessionManager.java`.
|
|
- Change `fleetd/src/test/java/dev/ltms/fleet/session/SessionManagerTest.java`.
|
|
|
|
No other unit may edit these files.
|
|
|
|
### Acceptance criteria
|
|
|
|
1. `onBackendError` moves a `BUSY` member to `BACKEND_ERROR` and stores the reason.
|
|
2. It also moves `DONE` to `BACKEND_ERROR`, covering the resolver race.
|
|
3. A later normal `onTurnComplete` cannot change `BACKEND_ERROR` back to `DONE`.
|
|
4. A released or unknown member is not recreated. The unknown case logs at `WARN` and returns an
|
|
explicit false result to its caller.
|
|
5. `onDelivered` refuses a `BACKEND_ERROR` member, just as it refuses generic `FAILED`.
|
|
6. `rosterView` reports `state: backend_error` and `failureReason` after the send ticket is gone.
|
|
7. Ordinary members do not gain a blank or invented `failureReason` field.
|
|
8. Existing constructors keep source compatibility for tests and adapters.
|
|
9. The focused tests and `mvn clean install` pass.
|
|
|
|
### Dependencies
|
|
|
|
None. Unit 4 can run with Units 1, 2, and 3.
|
|
|
|
Unit 5 calls the new session method from the production sink.
|
|
|
|
### What to report back
|
|
|
|
- The two race orderings and the tests for both.
|
|
- The exact roster JSON shape.
|
|
- The unknown-target result and log level.
|
|
- The focused test and `mvn clean install` results.
|
|
|
|
## Unit 5 — Production wiring, spawn gate, and fleet views
|
|
|
|
### Scope
|
|
|
|
Compose Units 1 to 4 in production. This is the only unit which edits `Fleetd.java`.
|
|
|
|
Add per-profile `errorPattern` config beside `exhaustedPattern`. Compile both once at startup. A
|
|
configured pattern wins over the legacy default. Report configured profiles and legacy-default
|
|
profiles separately at startup. A bad regex must stop startup with the profile and key in the
|
|
message.
|
|
|
|
Wire one production `BackendErrorSink` with this order:
|
|
|
|
1. mark the member `backend_error` with its reason;
|
|
2. resolve the profile and its current `effectiveCredentialId()` through the fail-loud #234 seam;
|
|
3. record the error in `BackendOutagePolicy`;
|
|
4. on a new incident, submit one event to `ReplyPushLoop`.
|
|
|
|
If target metadata cannot be resolved, do not end in `Optional.ifPresent`. Log an error and call the
|
|
Unit 3 unmapped-target notice. The failed send still reaches its ticket through CB-588.
|
|
|
|
Teach both spawn paths about a separate cool-off source. Exhaustion quarantine has priority when
|
|
both states are active. Automatic placement needs a distinct `coolingOff` set so its refusal does
|
|
not say “exhausted”.
|
|
|
|
Extend the MCP (Model Context Protocol) views:
|
|
|
|
- A cooling profile has `free: 0`, `credentialId`, and `coolingOffForSeconds` in `fleet_list`.
|
|
- It does not have `quarantinedForSeconds` unless exhaustion quarantine is also active.
|
|
- `fleet_profiles` has a separate `coolingOff` map, not an entry in `quarantined`.
|
|
- A direct spawn refusal says the credential is cooling off after repeated backend errors and gives
|
|
the remaining seconds.
|
|
|
|
Do not change `FleetHealthMonitor.coverage`. It still describes the periodic health webhook path.
|
|
Lead-pane outage delivery is a separate capability.
|
|
|
|
### Files owned
|
|
|
|
- `fleetd/src/main/java/dev/ltms/fleet/Fleetd.java`
|
|
- `fleetd/src/main/java/dev/ltms/fleet/config/FleetConfig.java`
|
|
- `fleetd/src/main/java/dev/ltms/fleet/config/ConfigRef.java`
|
|
- `fleetd/src/main/java/dev/ltms/fleet/member/CompositePeerLauncher.java`
|
|
- `fleetd/src/main/java/dev/ltms/fleet/placement/PlacementContext.java`
|
|
- `fleetd/src/main/java/dev/ltms/fleet/placement/PlacementPolicyUtil.java`
|
|
- `fleetd/src/main/java/dev/ltms/fleet/mcp/FleetMcp.java`
|
|
- `fleetd/src/test/java/dev/ltms/fleet/config/FleetConfigTest.java`
|
|
- `fleetd/src/test/java/dev/ltms/fleet/config/ConfigRefTest.java`
|
|
- `fleetd/src/test/java/dev/ltms/fleet/member/CompositePeerLauncherTest.java`
|
|
- `fleetd/src/test/java/dev/ltms/fleet/placement/PlacementPolicyTest.java`
|
|
- `fleetd/src/test/java/dev/ltms/fleet/mcp/FleetMcpTest.java`
|
|
- Add `fleetd/src/test/java/dev/ltms/fleet/BackendOutageFlowTest.java`.
|
|
- `fleetd/fleetd.example.yaml`
|
|
- `CLAUDE.md`
|
|
|
|
No earlier unit edits these files.
|
|
|
|
The lead, not a worker, must update `wiki/7-Use-Cases.md`, `wiki/9-Implementation.md`, and
|
|
`wiki/11-Features.md`. Project rules forbid workers from committing `wiki/`. The portable block in
|
|
`CLAUDE.md` and `wiki/7-Use-Cases.md` must remain byte-identical.
|
|
|
|
### Acceptance criteria
|
|
|
|
1. `errorPattern` binds per profile. Blank uses the legacy default and is reported as degraded
|
|
coverage. A config reload which changes it is reported as deferred because patterns are compiled
|
|
at startup.
|
|
2. A malformed `errorPattern` stops startup and names `profiles.<name>.errorPattern`.
|
|
3. The real production sink never silently drops an unknown target. A test captures its error log
|
|
and the fallback notice call.
|
|
4. One real backend-error classification through `CompletionResolver` marks only its member. It does
|
|
not cool the credential and does not send an outage notice.
|
|
5. Two real classifications for one credential within 60 seconds start one incident.
|
|
6. The integration test then calls the real explicit-profile spawn gate. It is refused before any
|
|
adapter spawn call, with “cooling off” and remaining seconds in the message.
|
|
7. The same test calls an automatic placement path. A cooling candidate is skipped. If every
|
|
candidate is cooling, the error names cool-off rather than exhaustion.
|
|
8. Profiles sharing the credential are all blocked. A profile on another credential stays usable.
|
|
9. `fleet_list` from the same fixture shows both members as `backend_error`, preserves each failure
|
|
reason, and reports `free: 0`, the credential, and `coolingOffForSeconds`.
|
|
10. `fleet_profiles` reports cool-off separately from quarantine.
|
|
11. The real `ReplyPushLoop` receives one incident and makes one successful lead-pane send. Existing
|
|
failed-ticket notice content may share that same combined send.
|
|
12. Exhaustion still wins when a line matches both patterns. A simultaneous exhaustion quarantine
|
|
also wins in spawn errors and fleet views.
|
|
13. After the 60-second cool-off, spawn is allowed again. A new incident needs two fresh errors.
|
|
14. `healthCoverage` has the same value before and after this change for the same health config.
|
|
15. `fleetd.example.yaml` explains `errorPattern`, the legacy fallback, 2/60/60 policy, and the
|
|
difference between cool-off and exhaustion quarantine.
|
|
16. `CLAUDE.md` tells leads how `fleet_profiles` and `fleet_list` report cool-off. The lead later
|
|
applies the matching wiki updates and runs the documented byte-sync check.
|
|
17. The developer records the new end-to-end test failing before implementation, then passing. The
|
|
focused suites and `mvn clean install` pass.
|
|
|
|
### Dependencies
|
|
|
|
Unit 5 starts only after Units 1 to 4 are merged or rebased into its branch. It also starts after the
|
|
#234 defect 2 fix lands, because both areas touch the same target-resolution control path.
|
|
|
|
### What to report back
|
|
|
|
- The exact commits used for Units 1 to 4 and #234.
|
|
- The startup coverage line with one configured and one legacy-default profile.
|
|
- The red test output before the implementation and its green result after.
|
|
- The explicit and automatic spawn refusal text.
|
|
- Sample `fleet_list` and `fleet_profiles` JSON for cool-off and exhaustion.
|
|
- The number and text of lead-pane sends in the real-path test.
|
|
- The focused test commands and final `mvn clean install` result.
|
|
- The exact `CLAUDE.md` change and the wiki edits the lead must apply.
|
|
|
|
## File ownership summary
|
|
|
|
| Area | Unit | Shared edit risk |
|
|
|---|---:|---|
|
|
| Completion classification | 1 | Only Unit 1 edits `CompletionResolver` and its test. |
|
|
| Correlation and cool-off state | 2 | New files only. |
|
|
| Lead push scheduling | 3 | Only Unit 3 edits `ReplyPushLoop` and its test. |
|
|
| Member terminal state | 4 | Only Unit 4 edits `MemberSession`, `SessionManager`, and their test. |
|
|
| Main composition, config, placement, MCP views, shipped prompt | 5 | Only Unit 5 edits `Fleetd`, `FleetConfig`, `CompositePeerLauncher`, placement context, `FleetMcp`, and `CLAUDE.md`. |
|
|
| Wiki propagation | Lead after Unit 5 | Workers do not commit the wiki submodule. |
|
|
|
|
## What I would not build
|
|
|
|
1. **Do not reuse `BackendQuarantine` for outages.** Its repeat call restarts a long credential
|
|
quarantine. Its fields and errors say “exhausted”. That is wrong for a short outage.
|
|
2. **Do not merge the exhaustion and generic error patterns.** Exhaustion must win because it has a
|
|
different policy and duration.
|
|
3. **Do not group by error string.** One outage can produce different text. The shared operational
|
|
limit is the credential.
|
|
4. **Do not mark a profile unusable until config changes.** The current classifier cannot safely
|
|
tell a permanent malformed request from a transient service fault. A permanent state would need
|
|
a stronger error taxonomy first.
|
|
5. **Do not quarantine on the first generic backend error.** That would turn one bad request or one
|
|
false pattern match into a fleet-wide capacity loss.
|
|
6. **Do not add this to `FleetHealthMonitor`.** The monitor samples slow member health. The exact
|
|
backend event already exists at completion resolution, and moving it to polling would lose type
|
|
and time.
|
|
7. **Do not add another direct lead injector.** `ReplyPushLoop` already owns status gating,
|
|
per-lead coalescing, retry bounds, and heartbeat stand-down.
|
|
8. **Do not change `healthCoverage` to `full`.** That field still means a webhook notification sink
|
|
exists for periodic health. A backend outage nudge does not make every health event visible.
|
|
9. **Do not persist incident history across daemon restart in this work.** Existing exhaustion
|
|
quarantine is also in memory. A 60-second state does not justify a new durable store.
|
|
10. **Do not build work recovery.** The PR-body survival story proves why checkpoint-first work is
|
|
useful, but these tickets are about detection, capacity, and signalling.
|
|
11. **Do not remove the legacy `API Error:` fallback in the first release.** Doing so would turn an
|
|
unedited config back into a false successful completion. Report it as degraded coverage instead.
|
|
12. **Do not reorder or add the old `visibleTurn` fallback from #201.** #211 already implemented the
|
|
narrow raw-scrape fallback at `CompletionResolver.classifyRawScrapeFallback`.
|
|
|
|
## Riskiest assumption and cheapest experiment
|
|
|
|
The riskiest assumption is that a configured error regex means “the backend failed this turn”. The
|
|
current code and test already show the counterexample: a worker may quote `API Error:` while writing
|
|
a valid report. Two such false matches on one credential would now remove capacity for 60 seconds.
|
|
|
|
The cheapest experiment is a replay corpus before Unit 5 ships:
|
|
|
|
1. Save the full pane text from the measured 2026-09-01 outage.
|
|
2. Produce one safe failure per backend with a disposable invalid endpoint or request.
|
|
3. Save one valid member report which quotes each error line.
|
|
4. Replay all samples through the real `CompletionResolver` test fixture.
|
|
5. Require outage samples to match and quoted-report samples not to match after assistant-block
|
|
extraction and baseline checks.
|
|
|
|
This costs no outage deployment and no real sleep. If quoted reports still match, narrow the profile
|
|
patterns before enabling correlation. Do not raise the threshold to hide a bad classifier.
|
|
|
|
## Sequencing with three developers
|
|
|
|
First wave:
|
|
|
|
1. Developer A: Unit 1, typed classification.
|
|
2. Developer B: Unit 2, credential outage policy.
|
|
3. Developer C: Unit 3, lead outage nudge.
|
|
|
|
As soon as one slot is free, start Unit 4. It is file-disjoint from every first-wave unit. Merge and
|
|
review Units 1 to 4 independently.
|
|
|
|
Start Unit 5 only after all four foundations and #234 are available. Unit 5 is the only high-conflict
|
|
integration branch, so no other active unit should touch its file list.
|
|
|
|
## Checks performed for this refinement
|
|
|
|
- Read issue #201 and issue #227 through their Gitea pages. Both showed zero comments.
|
|
- Read the source and tests named in the evidence section.
|
|
- Ran `git status --short --branch`; the branch was clean before this document was added.
|
|
- Ran `git log --oneline -12` to identify the branch base.
|
|
- I did not run Maven because this change adds only a design document.
|
|
- Rendered both Mermaid blocks with `npx @mermaid-js/mermaid-cli`; both commands succeeded.
|