28 KiB
CB-201 and CB-227 refinement
Date: 2026-09-03
Decision
#201 and #227 are one delivery program, but they are not one implementation unit.
#201 has a real seam: CompletionResolver can publish a typed backend-error event only after its
waiter resolution wins. #227 can consume that event without knowing any pane text. The classifier
must land before the final #227 wiring. However, the policy engine, roster state, and lead nudge can
be built in parallel with the classifier.
I propose five units. Units 1 to 4 own separate files and can run in parallel. Unit 5 owns all composition files and lands after them. It also depends on the #234 defect 2 fix named in the task.
flowchart LR
U1["Unit 1: typed backend-error classification"] --> U5["Unit 5: wire policy, spawn gate, and fleet views"]
U2["Unit 2: credential outage policy"] --> U5
U3["Unit 3: lead outage nudge"] --> U5
U4["Unit 4: durable member outcome"] --> U5
D234["#234 defect 2: fail-loud target resolution"] --> U5
Figure 1. Four file-disjoint foundations feed one composition unit.
This split keeps Fleetd.java under one owner. It also keeps every other production file under one
unit in this plan.
Evidence checked in the current branch
I read both issue pages in full. Each page reports zero comments.
| Evidence | What the code says now |
|---|---|
inject/CompletionResolver.java:229-237 |
A turn below two seconds fails before normal scrape classification. A matching fast backend error is therefore only a generic failure today. |
inject/CompletionResolver.java:260-276 and :332-367 |
#211 already added raw-screen classification when lastAssistantBlock is empty. The “dead-code question” in #201 is stale on this branch. |
inject/CompletionResolver.java:288-317 |
Exhaustion wins before the hard-coded API Error: match. A backend error then goes through generic fail(...). |
inject/CompletionResolver.java:311-316 |
The code admits that the pattern is a heuristic. A member report which quotes an API error may match it. |
inject/CompletionResolver.java:449-467 |
Startup coverage exists only for exhaustedPattern. |
inject/ExhaustedPatternLookup.java:13-25 |
The current lookup and explicit none() value are a good shape for the new classifier seam. |
Fleetd.java:196-207 |
One BackendQuarantine is shared by placement and the exhaustion sink. Its cooldown comes from quarantineCooldownSeconds. |
Fleetd.java:322-363 |
Pattern compilation, target-to-profile lookup, and the live ExhaustionSink are composed in Fleetd.main. The sink on this branch still ends in .ifPresent(...). This plan assumes #234 replaces that silent path. |
placement/BackendQuarantine.java:60-87 |
A repeated exhaustion restarts one long quarantine. The store is credential-keyed and uses an injected monotonic clock. |
member/CompositePeerLauncher.java:260-317 |
Explicit and policy-selected spawns have separate gates. Both paths must learn about outage cool-off. |
member/CompositePeerLauncher.java:347-379 |
Exhaustion refusal already checks a credential for explicit spawns and filters policy candidates. Its error text says “exhausted”. |
placement/PlacementContext.java:10-22 and PlacementPolicyUtil.java:14-83 |
Automatic placement has only one transient exclusion set named quarantined. Reusing it would make outage errors say “backend exhausted”. |
mcp/FleetMcp.java:913-1025 |
fleet_list sets free: 0 and adds credentialId plus quarantinedForSeconds when quarantine is active. |
session/MemberSession.java:51-59 |
The roster has DONE and generic FAILED, but no backend-error state or stored reason. |
session/SessionManager.java:695-773 |
A normal boundary moves BUSY to DONE. A failure moves any non-released session to FAILED. The async completion resolver can race the DONE update. |
session/SessionManager.java:648-687 |
rosterView reports the session state, but it reports no terminal reason. |
msg/MessageService.java:922-940 |
CB-588 already nudges for every terminal async ticket, including failures. Current code would report failed tickets, but it would not report one correlated outage. |
msg/ReplyPushLoop.java:20-48 |
Replies, terminal tickets, and questions share one per-lead schedule. This prevents two push sources from injecting competing turns. |
msg/ReplyPushLoop.java:305-395 |
Each push entry point resolves the owning lead through PrimaryRegistry. Missing ownership is logged and the durable or pending item remains the backstop. |
msg/ReplyPushLoop.java:496-547 |
One tick builds one combined nudge. Pending items have separate reminder counts. |
health/FleetHealthMonitor.java:91-143 |
Health is a slow periodic observer of members and message-layer facts. It does not receive completion classifications. |
health/FleetHealthMonitor.java:206-208 |
healthCoverage means health enabled plus webhook configured. It does not describe lead-pane alerts. |
Fleetd.java:465-486 |
Health stays detection-only without the webhook notification setting. |
I also read the related unit tests for CompletionResolver, ReplyPushLoop, BackendQuarantine,
CompositePeerLauncher, PlacementPolicyUtil, SessionManager, MessageService, and FleetMcp.
I did not inspect the in-progress #234 branch. I only used the two measured facts in the task. No peer architect was named, so I did not exchange a design with one.
Required behaviour
The policy should use these first values:
- Threshold: 2 classified backend errors.
- Window: 60 seconds, measured from the first error to the second.
- Cool-off: 60 seconds, starting when the threshold is reached.
- Correlation key:
credentialId, never profile name and never error text. - Incident rule: one active incident per credential. Errors during its cool-off do not extend it and do not create more lead notices.
- Rearm rule: after cool-off ends, two fresh errors are needed for another incident.
Two errors are the smallest threshold which protects the honest one-turn failure. A 60-second window fits the measured two-member outage. A 60-second cool-off blocks immediate repeat spawns without turning a short backend fault into the default 1,800-second exhaustion quarantine.
A single classified error still fails its send and marks its member backend_error. It does not
cool a credential and does not send an outage notice. This is what “a single error changes nothing”
must mean at the credential level. It cannot mean that the failed member still looks successful.
sequenceDiagram
participant R1 as Resolver for member A
participant R2 as Resolver for member B
participant P as Outage policy
participant S as Spawn gate
participant N as Lead push loop
participant L as Lead pane
R1->>P: backend error for credential C
Note over P: Count 1, no cool-off
R2->>P: backend error for credential C within 60s
P->>P: Start one 60s incident
P->>S: Credential C is cooling off
P->>N: Queue one incident notice
N->>L: Inject when lead is idle, blocked, or done
L->>S: Request another spawn on credential C
S-->>L: Refuse and report remaining cool-off
Figure 2. The second independent classification creates the fleet-level event.
Against the 2026-09-01 case, the second failed member would start cool-off. fleet_list would show
zero free capacity and both members as backend_error. The push loop would inject one outage notice
even if the lead had not polled either ticket yet. The design reports the outage. It does not recover
uncommitted work from the members.
Unit 1 — Typed backend-error classification
Scope
Replace the direct hard-coded check inside CompletionResolver with a lookup and a sink. Keep the
public send result as a failed send. The typed internal event is the seam #227 consumes.
The lookup returns the pattern for a target. The sink receives the target, matched line, and full
failure reason. It fires only after Rendezvous.resolveFailure(...) wins for that exact captured
waiter. This copies the race rule already used by ExhaustionSink.
The classifier must run in all three current paths:
- a normal non-empty assistant block;
- the #211 raw scrape fallback;
- a turn inside
MIN_TURN_NANOS, before it becomes a generic too-fast failure.
In every path, the order stays: stale-baseline guard, exhaustion, backend error, then generic failure or completion. A fast turn still fails when no configured pattern matches.
Keep (?i)\bAPI Error\s*: as a compatibility pattern for profiles without errorPattern until the
operator config is updated. Do not call this full coverage. Startup reporting in Unit 5 must name
profiles using this weaker legacy default.
Files owned
- Add
fleetd/src/main/java/dev/ltms/fleet/inject/BackendErrorPatternLookup.java. - Add
fleetd/src/main/java/dev/ltms/fleet/inject/BackendErrorSink.java. - Change
fleetd/src/main/java/dev/ltms/fleet/inject/CompletionResolver.java. - Change
fleetd/src/test/java/dev/ltms/fleet/inject/CompletionResolverTest.java.
No other unit may edit these files.
Acceptance criteria
- A target-specific error pattern matches a normal assistant block and resolves the send as failed.
- The same match calls
BackendErrorSinkexactly once after the waiter resolution wins. - A late classification which loses to
fleet_replydoes not call the sink. - An exhausted line that also matches the generic error pattern stays
BACKEND_EXHAUSTED. It calls onlyExhaustionSink. - A raw pane with leading Terminal User Interface (TUI) chrome and no assistant marker still uses the #211 fallback and calls the backend-error sink.
- A matching error inside the two-second floor is typed and sent to the sink. A non-matching fast turn stays a generic failure.
- An unchanged delivery baseline which contains old backend-error text is suppressed. It never increments outage evidence.
- A non-match keeps the existing completion result and text.
- Constructors used by current callers keep compiling. They use the legacy default lookup and an explicit inert sink until Unit 5 supplies the production objects.
- Unit tests pass. The developer runs the focused test first, then
mvn clean installfromfleetd/.
Dependencies
None. Unit 1 can run with Units 2, 3, and 4.
Unit 5 depends on its new lookup, sink, and constructor.
What to report back
- The exact classifier order in all three paths.
- The focused test command and result.
- The test which proves a losing waiter race does not publish an event.
- The test which proves a fast matching failure is typed.
- The final
mvn clean installresult. - Any constructor kept only for transition and where Unit 5 replaces it.
Unit 2 — Credential outage policy
Scope
Build a small credential-keyed state machine. It accepts already-classified backend-error events. It does not read pane text, profiles, sessions, or lead state.
Use an injected monotonic clock. A call records credentialId, target, and reason. It returns a new
incident only on the threshold crossing. The incident contains a stable event id, credential id,
the distinct affected targets, evidence count, window, and remaining cool-off.
This class owns both correlation and short cool-off. Keeping them together makes threshold crossing and the cool-off deadline one atomic state change.
Files owned
- Add
fleetd/src/main/java/dev/ltms/fleet/placement/BackendOutagePolicy.java. - Add
fleetd/src/test/java/dev/ltms/fleet/placement/BackendOutagePolicyTest.java.
No other unit may edit these files.
Acceptance criteria
- One error creates no incident and no cool-off.
- Two errors for one credential within 60 seconds create exactly one incident and a 60-second cool-off.
- Two errors more than 60 seconds apart do not create an incident.
- The exact 60-second boundary has a pinned result. Use inclusive
<= 60sso scheduler delay does not discard evidence at the boundary. - Different credentials never share evidence.
- Different profiles which supply the same credential id do share evidence. The policy itself only sees the credential id.
- More errors during active cool-off do not extend its deadline and do not return another incident.
- After expiry, old evidence is cleared. Two fresh errors are needed to create the next incident.
- Remaining seconds round up, matching
BackendQuarantinereporting. - Concurrent second and third errors cannot return two incidents.
- The focused tests and
mvn clean installpass.
Dependencies
None. Unit 2 can run with Units 1, 3, and 4.
Unit 5 depends on the policy API.
What to report back
- The state transition table and locking method.
- The exact threshold, window, cool-off, and boundary rule.
- The test which proves one incident under concurrent calls.
- The focused test and
mvn clean installresults.
Unit 3 — Lead outage nudge
Scope
Add backend incidents as a fourth pending source in ReplyPushLoop. Do not create another scheduler
or call AgentControl.send from Fleetd. The existing combined per-lead schedule is the control
which prevents competing injected turns.
The entry point takes an incident id, affected worker targets, credential id, affected profile
names, and remaining cool-off. It resolves distinct owning leads through PrimaryRegistry.
Each (incidentId, lead) item is one-shot. It waits while the lead is not injectable. After one
successful agents.send, remove it. A send exception keeps it pending for a bounded retry. It never
uses the repeated reminder behaviour of an uncollected ticket.
Also add a fail-loud entry point for a classified target that Unit 5 cannot map to a credential. It
uses PrimaryRegistry.nudgeTargetFor(target) and says that correlation could not run. If no lead is
known, log at WARN, not DEBUG.
Files owned
- Change
fleetd/src/main/java/dev/ltms/fleet/msg/ReplyPushLoop.java. - Change
fleetd/src/test/java/dev/ltms/fleet/msg/ReplyPushLoopTest.java.
No other unit may edit these files.
Acceptance criteria
- One incident affecting two workers owned by one lead causes one successful pane injection.
- Two affected workers owned by two leads cause one successful injection per affected lead. This is one notice per event per lead, not one notice per member.
- Repeating the same incident id is idempotent.
- A busy or unknown lead is not injected. The item stays pending until the lead becomes injectable or its attempt cap is reached.
- After one successful injection, later ticks do not mention that incident again.
- A failed
agents.sendis retried within the existing bound. A successful retry still gives only one successful send. - A pending ticket and an outage incident for one lead appear in one combined nudge, not two competing turns.
- The text names the credential, profiles, affected workers, and remaining cool-off. It tells the
lead to run
fleet_list. - An unmapped target produces a direct warning notice when a lead is known. If no lead is known,
the code logs a
WARNnaming the target and reason. stop()clears incident state as it clears other push state.- Existing reply, ticket, and question tests stay green. The focused tests and
mvn clean installpass.
Dependencies
None. The API uses plain values, not the Unit 2 incident class. This lets Unit 3 run in parallel.
Unit 5 adapts the Unit 2 incident into this entry point.
What to report back
- The exact one-shot and retry rules.
- The test showing one combined nudge with a failed ticket.
- The test showing one successful send for two affected workers.
- The focused test and
mvn clean installresults.
Unit 4 — Durable member backend-error outcome
Scope
Make a classified backend failure remain visible after its ticket is collected or expires.
Add BACKEND_ERROR to MemberSession.State. Add a nullable failure detail to MemberSession and
render it as failureReason in SessionManager.rosterView. Add
SessionManager.onBackendError(target, reason).
The transition must handle both completion orderings:
BUSY -> BACKEND_ERRORwhen classification wins before the normal completion state update;DONE -> BACKEND_ERRORwhen the async resolver runs afterSessionManager.onTurnComplete.
It must use a compare-and-set retry or another atomic update. RELEASED must never return to the
roster. A backend-error member is terminal and cannot accept another delivery.
Files owned
- Change
fleetd/src/main/java/dev/ltms/fleet/session/MemberSession.java. - Change
fleetd/src/main/java/dev/ltms/fleet/session/SessionManager.java. - Change
fleetd/src/test/java/dev/ltms/fleet/session/SessionManagerTest.java.
No other unit may edit these files.
Acceptance criteria
onBackendErrormoves aBUSYmember toBACKEND_ERRORand stores the reason.- It also moves
DONEtoBACKEND_ERROR, covering the resolver race. - A later normal
onTurnCompletecannot changeBACKEND_ERRORback toDONE. - A released or unknown member is not recreated. The unknown case logs at
WARNand returns an explicit false result to its caller. onDeliveredrefuses aBACKEND_ERRORmember, just as it refuses genericFAILED.rosterViewreportsstate: backend_errorandfailureReasonafter the send ticket is gone.- Ordinary members do not gain a blank or invented
failureReasonfield. - Existing constructors keep source compatibility for tests and adapters.
- The focused tests and
mvn clean installpass.
Dependencies
None. Unit 4 can run with Units 1, 2, and 3.
Unit 5 calls the new session method from the production sink.
What to report back
- The two race orderings and the tests for both.
- The exact roster JSON shape.
- The unknown-target result and log level.
- The focused test and
mvn clean installresults.
Unit 5 — Production wiring, spawn gate, and fleet views
Scope
Compose Units 1 to 4 in production. This is the only unit which edits Fleetd.java.
Add per-profile errorPattern config beside exhaustedPattern. Compile both once at startup. A
configured pattern wins over the legacy default. Report configured profiles and legacy-default
profiles separately at startup. A bad regex must stop startup with the profile and key in the
message.
Wire one production BackendErrorSink with this order:
- mark the member
backend_errorwith its reason; - resolve the profile and its current
effectiveCredentialId()through the fail-loud #234 seam; - record the error in
BackendOutagePolicy; - on a new incident, submit one event to
ReplyPushLoop.
If target metadata cannot be resolved, do not end in Optional.ifPresent. Log an error and call the
Unit 3 unmapped-target notice. The failed send still reaches its ticket through CB-588.
Teach both spawn paths about a separate cool-off source. Exhaustion quarantine has priority when
both states are active. Automatic placement needs a distinct coolingOff set so its refusal does
not say “exhausted”.
Extend the MCP (Model Context Protocol) views:
- A cooling profile has
free: 0,credentialId, andcoolingOffForSecondsinfleet_list. - It does not have
quarantinedForSecondsunless exhaustion quarantine is also active. fleet_profileshas a separatecoolingOffmap, not an entry inquarantined.- A direct spawn refusal says the credential is cooling off after repeated backend errors and gives the remaining seconds.
Do not change FleetHealthMonitor.coverage. It still describes the periodic health webhook path.
Lead-pane outage delivery is a separate capability.
Files owned
fleetd/src/main/java/dev/ltms/fleet/Fleetd.javafleetd/src/main/java/dev/ltms/fleet/config/FleetConfig.javafleetd/src/main/java/dev/ltms/fleet/config/ConfigRef.javafleetd/src/main/java/dev/ltms/fleet/member/CompositePeerLauncher.javafleetd/src/main/java/dev/ltms/fleet/placement/PlacementContext.javafleetd/src/main/java/dev/ltms/fleet/placement/PlacementPolicyUtil.javafleetd/src/main/java/dev/ltms/fleet/mcp/FleetMcp.javafleetd/src/test/java/dev/ltms/fleet/config/FleetConfigTest.javafleetd/src/test/java/dev/ltms/fleet/config/ConfigRefTest.javafleetd/src/test/java/dev/ltms/fleet/member/CompositePeerLauncherTest.javafleetd/src/test/java/dev/ltms/fleet/placement/PlacementPolicyTest.javafleetd/src/test/java/dev/ltms/fleet/mcp/FleetMcpTest.java- Add
fleetd/src/test/java/dev/ltms/fleet/BackendOutageFlowTest.java. fleetd/fleetd.example.yamlCLAUDE.md
No earlier unit edits these files.
The lead, not a worker, must update wiki/7-Use-Cases.md, wiki/9-Implementation.md, and
wiki/11-Features.md. Project rules forbid workers from committing wiki/. The portable block in
CLAUDE.md and wiki/7-Use-Cases.md must remain byte-identical.
Acceptance criteria
errorPatternbinds per profile. Blank uses the legacy default and is reported as degraded coverage. A config reload which changes it is reported as deferred because patterns are compiled at startup.- A malformed
errorPatternstops startup and namesprofiles.<name>.errorPattern. - The real production sink never silently drops an unknown target. A test captures its error log and the fallback notice call.
- One real backend-error classification through
CompletionResolvermarks only its member. It does not cool the credential and does not send an outage notice. - Two real classifications for one credential within 60 seconds start one incident.
- The integration test then calls the real explicit-profile spawn gate. It is refused before any adapter spawn call, with “cooling off” and remaining seconds in the message.
- The same test calls an automatic placement path. A cooling candidate is skipped. If every candidate is cooling, the error names cool-off rather than exhaustion.
- Profiles sharing the credential are all blocked. A profile on another credential stays usable.
fleet_listfrom the same fixture shows both members asbackend_error, preserves each failure reason, and reportsfree: 0, the credential, andcoolingOffForSeconds.fleet_profilesreports cool-off separately from quarantine.- The real
ReplyPushLoopreceives one incident and makes one successful lead-pane send. Existing failed-ticket notice content may share that same combined send. - Exhaustion still wins when a line matches both patterns. A simultaneous exhaustion quarantine also wins in spawn errors and fleet views.
- After the 60-second cool-off, spawn is allowed again. A new incident needs two fresh errors.
healthCoveragehas the same value before and after this change for the same health config.fleetd.example.yamlexplainserrorPattern, the legacy fallback, 2/60/60 policy, and the difference between cool-off and exhaustion quarantine.CLAUDE.mdtells leads howfleet_profilesandfleet_listreport cool-off. The lead later applies the matching wiki updates and runs the documented byte-sync check.- The developer records the new end-to-end test failing before implementation, then passing. The
focused suites and
mvn clean installpass.
Dependencies
Unit 5 starts only after Units 1 to 4 are merged or rebased into its branch. It also starts after the #234 defect 2 fix lands, because both areas touch the same target-resolution control path.
What to report back
- The exact commits used for Units 1 to 4 and #234.
- The startup coverage line with one configured and one legacy-default profile.
- The red test output before the implementation and its green result after.
- The explicit and automatic spawn refusal text.
- Sample
fleet_listandfleet_profilesJSON for cool-off and exhaustion. - The number and text of lead-pane sends in the real-path test.
- The focused test commands and final
mvn clean installresult. - The exact
CLAUDE.mdchange and the wiki edits the lead must apply.
File ownership summary
| Area | Unit | Shared edit risk |
|---|---|---|
| Completion classification | 1 | Only Unit 1 edits CompletionResolver and its test. |
| Correlation and cool-off state | 2 | New files only. |
| Lead push scheduling | 3 | Only Unit 3 edits ReplyPushLoop and its test. |
| Member terminal state | 4 | Only Unit 4 edits MemberSession, SessionManager, and their test. |
| Main composition, config, placement, MCP views, shipped prompt | 5 | Only Unit 5 edits Fleetd, FleetConfig, CompositePeerLauncher, placement context, FleetMcp, and CLAUDE.md. |
| Wiki propagation | Lead after Unit 5 | Workers do not commit the wiki submodule. |
What I would not build
- Do not reuse
BackendQuarantinefor outages. Its repeat call restarts a long credential quarantine. Its fields and errors say “exhausted”. That is wrong for a short outage. - Do not merge the exhaustion and generic error patterns. Exhaustion must win because it has a different policy and duration.
- Do not group by error string. One outage can produce different text. The shared operational limit is the credential.
- Do not mark a profile unusable until config changes. The current classifier cannot safely tell a permanent malformed request from a transient service fault. A permanent state would need a stronger error taxonomy first.
- Do not quarantine on the first generic backend error. That would turn one bad request or one false pattern match into a fleet-wide capacity loss.
- Do not add this to
FleetHealthMonitor. The monitor samples slow member health. The exact backend event already exists at completion resolution, and moving it to polling would lose type and time. - Do not add another direct lead injector.
ReplyPushLoopalready owns status gating, per-lead coalescing, retry bounds, and heartbeat stand-down. - Do not change
healthCoveragetofull. That field still means a webhook notification sink exists for periodic health. A backend outage nudge does not make every health event visible. - Do not persist incident history across daemon restart in this work. Existing exhaustion quarantine is also in memory. A 60-second state does not justify a new durable store.
- Do not build work recovery. The PR-body survival story proves why checkpoint-first work is useful, but these tickets are about detection, capacity, and signalling.
- Do not remove the legacy
API Error:fallback in the first release. Doing so would turn an unedited config back into a false successful completion. Report it as degraded coverage instead. - Do not reorder or add the old
visibleTurnfallback from #201. #211 already implemented the narrow raw-scrape fallback atCompletionResolver.classifyRawScrapeFallback.
Riskiest assumption and cheapest experiment
The riskiest assumption is that a configured error regex means “the backend failed this turn”. The
current code and test already show the counterexample: a worker may quote API Error: while writing
a valid report. Two such false matches on one credential would now remove capacity for 60 seconds.
The cheapest experiment is a replay corpus before Unit 5 ships:
- Save the full pane text from the measured 2026-09-01 outage.
- Produce one safe failure per backend with a disposable invalid endpoint or request.
- Save one valid member report which quotes each error line.
- Replay all samples through the real
CompletionResolvertest fixture. - Require outage samples to match and quoted-report samples not to match after assistant-block extraction and baseline checks.
This costs no outage deployment and no real sleep. If quoted reports still match, narrow the profile patterns before enabling correlation. Do not raise the threshold to hide a bad classifier.
Sequencing with three developers
First wave:
- Developer A: Unit 1, typed classification.
- Developer B: Unit 2, credential outage policy.
- Developer C: Unit 3, lead outage nudge.
As soon as one slot is free, start Unit 4. It is file-disjoint from every first-wave unit. Merge and review Units 1 to 4 independently.
Start Unit 5 only after all four foundations and #234 are available. Unit 5 is the only high-conflict integration branch, so no other active unit should touch its file list.
Checks performed for this refinement
- Read issue #201 and issue #227 through their Gitea pages. Both showed zero comments.
- Read the source and tests named in the evidence section.
- Ran
git status --short --branch; the branch was clean before this document was added. - Ran
git log --oneline -12to identify the branch base. - I did not run Maven because this change adds only a design document.
- Rendered both Mermaid blocks with
npx @mermaid-js/mermaid-cli; both commands succeeded.