Files
fleetd/docs/CB-201-227-Refinement.md
Dai Ha 7662e2d0c8
CI / contract (push) Successful in 53s
CI / build (push) Successful in 1m49s
fleetd #201/#227: refine backend outage work
2026-09-03 10:21:44 +07:00

28 KiB

CB-201 and CB-227 refinement

Date: 2026-09-03

Decision

#201 and #227 are one delivery program, but they are not one implementation unit.

#201 has a real seam: CompletionResolver can publish a typed backend-error event only after its waiter resolution wins. #227 can consume that event without knowing any pane text. The classifier must land before the final #227 wiring. However, the policy engine, roster state, and lead nudge can be built in parallel with the classifier.

I propose five units. Units 1 to 4 own separate files and can run in parallel. Unit 5 owns all composition files and lands after them. It also depends on the #234 defect 2 fix named in the task.

flowchart LR
    U1["Unit 1: typed backend-error classification"] --> U5["Unit 5: wire policy, spawn gate, and fleet views"]
    U2["Unit 2: credential outage policy"] --> U5
    U3["Unit 3: lead outage nudge"] --> U5
    U4["Unit 4: durable member outcome"] --> U5
    D234["#234 defect 2: fail-loud target resolution"] --> U5

Figure 1. Four file-disjoint foundations feed one composition unit.

This split keeps Fleetd.java under one owner. It also keeps every other production file under one unit in this plan.

Evidence checked in the current branch

I read both issue pages in full. Each page reports zero comments.

Evidence What the code says now
inject/CompletionResolver.java:229-237 A turn below two seconds fails before normal scrape classification. A matching fast backend error is therefore only a generic failure today.
inject/CompletionResolver.java:260-276 and :332-367 #211 already added raw-screen classification when lastAssistantBlock is empty. The “dead-code question” in #201 is stale on this branch.
inject/CompletionResolver.java:288-317 Exhaustion wins before the hard-coded API Error: match. A backend error then goes through generic fail(...).
inject/CompletionResolver.java:311-316 The code admits that the pattern is a heuristic. A member report which quotes an API error may match it.
inject/CompletionResolver.java:449-467 Startup coverage exists only for exhaustedPattern.
inject/ExhaustedPatternLookup.java:13-25 The current lookup and explicit none() value are a good shape for the new classifier seam.
Fleetd.java:196-207 One BackendQuarantine is shared by placement and the exhaustion sink. Its cooldown comes from quarantineCooldownSeconds.
Fleetd.java:322-363 Pattern compilation, target-to-profile lookup, and the live ExhaustionSink are composed in Fleetd.main. The sink on this branch still ends in .ifPresent(...). This plan assumes #234 replaces that silent path.
placement/BackendQuarantine.java:60-87 A repeated exhaustion restarts one long quarantine. The store is credential-keyed and uses an injected monotonic clock.
member/CompositePeerLauncher.java:260-317 Explicit and policy-selected spawns have separate gates. Both paths must learn about outage cool-off.
member/CompositePeerLauncher.java:347-379 Exhaustion refusal already checks a credential for explicit spawns and filters policy candidates. Its error text says “exhausted”.
placement/PlacementContext.java:10-22 and PlacementPolicyUtil.java:14-83 Automatic placement has only one transient exclusion set named quarantined. Reusing it would make outage errors say “backend exhausted”.
mcp/FleetMcp.java:913-1025 fleet_list sets free: 0 and adds credentialId plus quarantinedForSeconds when quarantine is active.
session/MemberSession.java:51-59 The roster has DONE and generic FAILED, but no backend-error state or stored reason.
session/SessionManager.java:695-773 A normal boundary moves BUSY to DONE. A failure moves any non-released session to FAILED. The async completion resolver can race the DONE update.
session/SessionManager.java:648-687 rosterView reports the session state, but it reports no terminal reason.
msg/MessageService.java:922-940 CB-588 already nudges for every terminal async ticket, including failures. Current code would report failed tickets, but it would not report one correlated outage.
msg/ReplyPushLoop.java:20-48 Replies, terminal tickets, and questions share one per-lead schedule. This prevents two push sources from injecting competing turns.
msg/ReplyPushLoop.java:305-395 Each push entry point resolves the owning lead through PrimaryRegistry. Missing ownership is logged and the durable or pending item remains the backstop.
msg/ReplyPushLoop.java:496-547 One tick builds one combined nudge. Pending items have separate reminder counts.
health/FleetHealthMonitor.java:91-143 Health is a slow periodic observer of members and message-layer facts. It does not receive completion classifications.
health/FleetHealthMonitor.java:206-208 healthCoverage means health enabled plus webhook configured. It does not describe lead-pane alerts.
Fleetd.java:465-486 Health stays detection-only without the webhook notification setting.

I also read the related unit tests for CompletionResolver, ReplyPushLoop, BackendQuarantine, CompositePeerLauncher, PlacementPolicyUtil, SessionManager, MessageService, and FleetMcp.

I did not inspect the in-progress #234 branch. I only used the two measured facts in the task. No peer architect was named, so I did not exchange a design with one.

Required behaviour

The policy should use these first values:

  • Threshold: 2 classified backend errors.
  • Window: 60 seconds, measured from the first error to the second.
  • Cool-off: 60 seconds, starting when the threshold is reached.
  • Correlation key: credentialId, never profile name and never error text.
  • Incident rule: one active incident per credential. Errors during its cool-off do not extend it and do not create more lead notices.
  • Rearm rule: after cool-off ends, two fresh errors are needed for another incident.

Two errors are the smallest threshold which protects the honest one-turn failure. A 60-second window fits the measured two-member outage. A 60-second cool-off blocks immediate repeat spawns without turning a short backend fault into the default 1,800-second exhaustion quarantine.

A single classified error still fails its send and marks its member backend_error. It does not cool a credential and does not send an outage notice. This is what “a single error changes nothing” must mean at the credential level. It cannot mean that the failed member still looks successful.

sequenceDiagram
    participant R1 as Resolver for member A
    participant R2 as Resolver for member B
    participant P as Outage policy
    participant S as Spawn gate
    participant N as Lead push loop
    participant L as Lead pane

    R1->>P: backend error for credential C
    Note over P: Count 1, no cool-off
    R2->>P: backend error for credential C within 60s
    P->>P: Start one 60s incident
    P->>S: Credential C is cooling off
    P->>N: Queue one incident notice
    N->>L: Inject when lead is idle, blocked, or done
    L->>S: Request another spawn on credential C
    S-->>L: Refuse and report remaining cool-off

Figure 2. The second independent classification creates the fleet-level event.

Against the 2026-09-01 case, the second failed member would start cool-off. fleet_list would show zero free capacity and both members as backend_error. The push loop would inject one outage notice even if the lead had not polled either ticket yet. The design reports the outage. It does not recover uncommitted work from the members.

Unit 1 — Typed backend-error classification

Scope

Replace the direct hard-coded check inside CompletionResolver with a lookup and a sink. Keep the public send result as a failed send. The typed internal event is the seam #227 consumes.

The lookup returns the pattern for a target. The sink receives the target, matched line, and full failure reason. It fires only after Rendezvous.resolveFailure(...) wins for that exact captured waiter. This copies the race rule already used by ExhaustionSink.

The classifier must run in all three current paths:

  1. a normal non-empty assistant block;
  2. the #211 raw scrape fallback;
  3. a turn inside MIN_TURN_NANOS, before it becomes a generic too-fast failure.

In every path, the order stays: stale-baseline guard, exhaustion, backend error, then generic failure or completion. A fast turn still fails when no configured pattern matches.

Keep (?i)\bAPI Error\s*: as a compatibility pattern for profiles without errorPattern until the operator config is updated. Do not call this full coverage. Startup reporting in Unit 5 must name profiles using this weaker legacy default.

Files owned

  • Add fleetd/src/main/java/dev/ltms/fleet/inject/BackendErrorPatternLookup.java.
  • Add fleetd/src/main/java/dev/ltms/fleet/inject/BackendErrorSink.java.
  • Change fleetd/src/main/java/dev/ltms/fleet/inject/CompletionResolver.java.
  • Change fleetd/src/test/java/dev/ltms/fleet/inject/CompletionResolverTest.java.

No other unit may edit these files.

Acceptance criteria

  1. A target-specific error pattern matches a normal assistant block and resolves the send as failed.
  2. The same match calls BackendErrorSink exactly once after the waiter resolution wins.
  3. A late classification which loses to fleet_reply does not call the sink.
  4. An exhausted line that also matches the generic error pattern stays BACKEND_EXHAUSTED. It calls only ExhaustionSink.
  5. A raw pane with leading Terminal User Interface (TUI) chrome and no assistant marker still uses the #211 fallback and calls the backend-error sink.
  6. A matching error inside the two-second floor is typed and sent to the sink. A non-matching fast turn stays a generic failure.
  7. An unchanged delivery baseline which contains old backend-error text is suppressed. It never increments outage evidence.
  8. A non-match keeps the existing completion result and text.
  9. Constructors used by current callers keep compiling. They use the legacy default lookup and an explicit inert sink until Unit 5 supplies the production objects.
  10. Unit tests pass. The developer runs the focused test first, then mvn clean install from fleetd/.

Dependencies

None. Unit 1 can run with Units 2, 3, and 4.

Unit 5 depends on its new lookup, sink, and constructor.

What to report back

  • The exact classifier order in all three paths.
  • The focused test command and result.
  • The test which proves a losing waiter race does not publish an event.
  • The test which proves a fast matching failure is typed.
  • The final mvn clean install result.
  • Any constructor kept only for transition and where Unit 5 replaces it.

Unit 2 — Credential outage policy

Scope

Build a small credential-keyed state machine. It accepts already-classified backend-error events. It does not read pane text, profiles, sessions, or lead state.

Use an injected monotonic clock. A call records credentialId, target, and reason. It returns a new incident only on the threshold crossing. The incident contains a stable event id, credential id, the distinct affected targets, evidence count, window, and remaining cool-off.

This class owns both correlation and short cool-off. Keeping them together makes threshold crossing and the cool-off deadline one atomic state change.

Files owned

  • Add fleetd/src/main/java/dev/ltms/fleet/placement/BackendOutagePolicy.java.
  • Add fleetd/src/test/java/dev/ltms/fleet/placement/BackendOutagePolicyTest.java.

No other unit may edit these files.

Acceptance criteria

  1. One error creates no incident and no cool-off.
  2. Two errors for one credential within 60 seconds create exactly one incident and a 60-second cool-off.
  3. Two errors more than 60 seconds apart do not create an incident.
  4. The exact 60-second boundary has a pinned result. Use inclusive <= 60s so scheduler delay does not discard evidence at the boundary.
  5. Different credentials never share evidence.
  6. Different profiles which supply the same credential id do share evidence. The policy itself only sees the credential id.
  7. More errors during active cool-off do not extend its deadline and do not return another incident.
  8. After expiry, old evidence is cleared. Two fresh errors are needed to create the next incident.
  9. Remaining seconds round up, matching BackendQuarantine reporting.
  10. Concurrent second and third errors cannot return two incidents.
  11. The focused tests and mvn clean install pass.

Dependencies

None. Unit 2 can run with Units 1, 3, and 4.

Unit 5 depends on the policy API.

What to report back

  • The state transition table and locking method.
  • The exact threshold, window, cool-off, and boundary rule.
  • The test which proves one incident under concurrent calls.
  • The focused test and mvn clean install results.

Unit 3 — Lead outage nudge

Scope

Add backend incidents as a fourth pending source in ReplyPushLoop. Do not create another scheduler or call AgentControl.send from Fleetd. The existing combined per-lead schedule is the control which prevents competing injected turns.

The entry point takes an incident id, affected worker targets, credential id, affected profile names, and remaining cool-off. It resolves distinct owning leads through PrimaryRegistry.

Each (incidentId, lead) item is one-shot. It waits while the lead is not injectable. After one successful agents.send, remove it. A send exception keeps it pending for a bounded retry. It never uses the repeated reminder behaviour of an uncollected ticket.

Also add a fail-loud entry point for a classified target that Unit 5 cannot map to a credential. It uses PrimaryRegistry.nudgeTargetFor(target) and says that correlation could not run. If no lead is known, log at WARN, not DEBUG.

Files owned

  • Change fleetd/src/main/java/dev/ltms/fleet/msg/ReplyPushLoop.java.
  • Change fleetd/src/test/java/dev/ltms/fleet/msg/ReplyPushLoopTest.java.

No other unit may edit these files.

Acceptance criteria

  1. One incident affecting two workers owned by one lead causes one successful pane injection.
  2. Two affected workers owned by two leads cause one successful injection per affected lead. This is one notice per event per lead, not one notice per member.
  3. Repeating the same incident id is idempotent.
  4. A busy or unknown lead is not injected. The item stays pending until the lead becomes injectable or its attempt cap is reached.
  5. After one successful injection, later ticks do not mention that incident again.
  6. A failed agents.send is retried within the existing bound. A successful retry still gives only one successful send.
  7. A pending ticket and an outage incident for one lead appear in one combined nudge, not two competing turns.
  8. The text names the credential, profiles, affected workers, and remaining cool-off. It tells the lead to run fleet_list.
  9. An unmapped target produces a direct warning notice when a lead is known. If no lead is known, the code logs a WARN naming the target and reason.
  10. stop() clears incident state as it clears other push state.
  11. Existing reply, ticket, and question tests stay green. The focused tests and mvn clean install pass.

Dependencies

None. The API uses plain values, not the Unit 2 incident class. This lets Unit 3 run in parallel.

Unit 5 adapts the Unit 2 incident into this entry point.

What to report back

  • The exact one-shot and retry rules.
  • The test showing one combined nudge with a failed ticket.
  • The test showing one successful send for two affected workers.
  • The focused test and mvn clean install results.

Unit 4 — Durable member backend-error outcome

Scope

Make a classified backend failure remain visible after its ticket is collected or expires.

Add BACKEND_ERROR to MemberSession.State. Add a nullable failure detail to MemberSession and render it as failureReason in SessionManager.rosterView. Add SessionManager.onBackendError(target, reason).

The transition must handle both completion orderings:

  • BUSY -> BACKEND_ERROR when classification wins before the normal completion state update;
  • DONE -> BACKEND_ERROR when the async resolver runs after SessionManager.onTurnComplete.

It must use a compare-and-set retry or another atomic update. RELEASED must never return to the roster. A backend-error member is terminal and cannot accept another delivery.

Files owned

  • Change fleetd/src/main/java/dev/ltms/fleet/session/MemberSession.java.
  • Change fleetd/src/main/java/dev/ltms/fleet/session/SessionManager.java.
  • Change fleetd/src/test/java/dev/ltms/fleet/session/SessionManagerTest.java.

No other unit may edit these files.

Acceptance criteria

  1. onBackendError moves a BUSY member to BACKEND_ERROR and stores the reason.
  2. It also moves DONE to BACKEND_ERROR, covering the resolver race.
  3. A later normal onTurnComplete cannot change BACKEND_ERROR back to DONE.
  4. A released or unknown member is not recreated. The unknown case logs at WARN and returns an explicit false result to its caller.
  5. onDelivered refuses a BACKEND_ERROR member, just as it refuses generic FAILED.
  6. rosterView reports state: backend_error and failureReason after the send ticket is gone.
  7. Ordinary members do not gain a blank or invented failureReason field.
  8. Existing constructors keep source compatibility for tests and adapters.
  9. The focused tests and mvn clean install pass.

Dependencies

None. Unit 4 can run with Units 1, 2, and 3.

Unit 5 calls the new session method from the production sink.

What to report back

  • The two race orderings and the tests for both.
  • The exact roster JSON shape.
  • The unknown-target result and log level.
  • The focused test and mvn clean install results.

Unit 5 — Production wiring, spawn gate, and fleet views

Scope

Compose Units 1 to 4 in production. This is the only unit which edits Fleetd.java.

Add per-profile errorPattern config beside exhaustedPattern. Compile both once at startup. A configured pattern wins over the legacy default. Report configured profiles and legacy-default profiles separately at startup. A bad regex must stop startup with the profile and key in the message.

Wire one production BackendErrorSink with this order:

  1. mark the member backend_error with its reason;
  2. resolve the profile and its current effectiveCredentialId() through the fail-loud #234 seam;
  3. record the error in BackendOutagePolicy;
  4. on a new incident, submit one event to ReplyPushLoop.

If target metadata cannot be resolved, do not end in Optional.ifPresent. Log an error and call the Unit 3 unmapped-target notice. The failed send still reaches its ticket through CB-588.

Teach both spawn paths about a separate cool-off source. Exhaustion quarantine has priority when both states are active. Automatic placement needs a distinct coolingOff set so its refusal does not say “exhausted”.

Extend the MCP (Model Context Protocol) views:

  • A cooling profile has free: 0, credentialId, and coolingOffForSeconds in fleet_list.
  • It does not have quarantinedForSeconds unless exhaustion quarantine is also active.
  • fleet_profiles has a separate coolingOff map, not an entry in quarantined.
  • A direct spawn refusal says the credential is cooling off after repeated backend errors and gives the remaining seconds.

Do not change FleetHealthMonitor.coverage. It still describes the periodic health webhook path. Lead-pane outage delivery is a separate capability.

Files owned

  • fleetd/src/main/java/dev/ltms/fleet/Fleetd.java
  • fleetd/src/main/java/dev/ltms/fleet/config/FleetConfig.java
  • fleetd/src/main/java/dev/ltms/fleet/config/ConfigRef.java
  • fleetd/src/main/java/dev/ltms/fleet/member/CompositePeerLauncher.java
  • fleetd/src/main/java/dev/ltms/fleet/placement/PlacementContext.java
  • fleetd/src/main/java/dev/ltms/fleet/placement/PlacementPolicyUtil.java
  • fleetd/src/main/java/dev/ltms/fleet/mcp/FleetMcp.java
  • fleetd/src/test/java/dev/ltms/fleet/config/FleetConfigTest.java
  • fleetd/src/test/java/dev/ltms/fleet/config/ConfigRefTest.java
  • fleetd/src/test/java/dev/ltms/fleet/member/CompositePeerLauncherTest.java
  • fleetd/src/test/java/dev/ltms/fleet/placement/PlacementPolicyTest.java
  • fleetd/src/test/java/dev/ltms/fleet/mcp/FleetMcpTest.java
  • Add fleetd/src/test/java/dev/ltms/fleet/BackendOutageFlowTest.java.
  • fleetd/fleetd.example.yaml
  • CLAUDE.md

No earlier unit edits these files.

The lead, not a worker, must update wiki/7-Use-Cases.md, wiki/9-Implementation.md, and wiki/11-Features.md. Project rules forbid workers from committing wiki/. The portable block in CLAUDE.md and wiki/7-Use-Cases.md must remain byte-identical.

Acceptance criteria

  1. errorPattern binds per profile. Blank uses the legacy default and is reported as degraded coverage. A config reload which changes it is reported as deferred because patterns are compiled at startup.
  2. A malformed errorPattern stops startup and names profiles.<name>.errorPattern.
  3. The real production sink never silently drops an unknown target. A test captures its error log and the fallback notice call.
  4. One real backend-error classification through CompletionResolver marks only its member. It does not cool the credential and does not send an outage notice.
  5. Two real classifications for one credential within 60 seconds start one incident.
  6. The integration test then calls the real explicit-profile spawn gate. It is refused before any adapter spawn call, with “cooling off” and remaining seconds in the message.
  7. The same test calls an automatic placement path. A cooling candidate is skipped. If every candidate is cooling, the error names cool-off rather than exhaustion.
  8. Profiles sharing the credential are all blocked. A profile on another credential stays usable.
  9. fleet_list from the same fixture shows both members as backend_error, preserves each failure reason, and reports free: 0, the credential, and coolingOffForSeconds.
  10. fleet_profiles reports cool-off separately from quarantine.
  11. The real ReplyPushLoop receives one incident and makes one successful lead-pane send. Existing failed-ticket notice content may share that same combined send.
  12. Exhaustion still wins when a line matches both patterns. A simultaneous exhaustion quarantine also wins in spawn errors and fleet views.
  13. After the 60-second cool-off, spawn is allowed again. A new incident needs two fresh errors.
  14. healthCoverage has the same value before and after this change for the same health config.
  15. fleetd.example.yaml explains errorPattern, the legacy fallback, 2/60/60 policy, and the difference between cool-off and exhaustion quarantine.
  16. CLAUDE.md tells leads how fleet_profiles and fleet_list report cool-off. The lead later applies the matching wiki updates and runs the documented byte-sync check.
  17. The developer records the new end-to-end test failing before implementation, then passing. The focused suites and mvn clean install pass.

Dependencies

Unit 5 starts only after Units 1 to 4 are merged or rebased into its branch. It also starts after the #234 defect 2 fix lands, because both areas touch the same target-resolution control path.

What to report back

  • The exact commits used for Units 1 to 4 and #234.
  • The startup coverage line with one configured and one legacy-default profile.
  • The red test output before the implementation and its green result after.
  • The explicit and automatic spawn refusal text.
  • Sample fleet_list and fleet_profiles JSON for cool-off and exhaustion.
  • The number and text of lead-pane sends in the real-path test.
  • The focused test commands and final mvn clean install result.
  • The exact CLAUDE.md change and the wiki edits the lead must apply.

File ownership summary

Area Unit Shared edit risk
Completion classification 1 Only Unit 1 edits CompletionResolver and its test.
Correlation and cool-off state 2 New files only.
Lead push scheduling 3 Only Unit 3 edits ReplyPushLoop and its test.
Member terminal state 4 Only Unit 4 edits MemberSession, SessionManager, and their test.
Main composition, config, placement, MCP views, shipped prompt 5 Only Unit 5 edits Fleetd, FleetConfig, CompositePeerLauncher, placement context, FleetMcp, and CLAUDE.md.
Wiki propagation Lead after Unit 5 Workers do not commit the wiki submodule.

What I would not build

  1. Do not reuse BackendQuarantine for outages. Its repeat call restarts a long credential quarantine. Its fields and errors say “exhausted”. That is wrong for a short outage.
  2. Do not merge the exhaustion and generic error patterns. Exhaustion must win because it has a different policy and duration.
  3. Do not group by error string. One outage can produce different text. The shared operational limit is the credential.
  4. Do not mark a profile unusable until config changes. The current classifier cannot safely tell a permanent malformed request from a transient service fault. A permanent state would need a stronger error taxonomy first.
  5. Do not quarantine on the first generic backend error. That would turn one bad request or one false pattern match into a fleet-wide capacity loss.
  6. Do not add this to FleetHealthMonitor. The monitor samples slow member health. The exact backend event already exists at completion resolution, and moving it to polling would lose type and time.
  7. Do not add another direct lead injector. ReplyPushLoop already owns status gating, per-lead coalescing, retry bounds, and heartbeat stand-down.
  8. Do not change healthCoverage to full. That field still means a webhook notification sink exists for periodic health. A backend outage nudge does not make every health event visible.
  9. Do not persist incident history across daemon restart in this work. Existing exhaustion quarantine is also in memory. A 60-second state does not justify a new durable store.
  10. Do not build work recovery. The PR-body survival story proves why checkpoint-first work is useful, but these tickets are about detection, capacity, and signalling.
  11. Do not remove the legacy API Error: fallback in the first release. Doing so would turn an unedited config back into a false successful completion. Report it as degraded coverage instead.
  12. Do not reorder or add the old visibleTurn fallback from #201. #211 already implemented the narrow raw-scrape fallback at CompletionResolver.classifyRawScrapeFallback.

Riskiest assumption and cheapest experiment

The riskiest assumption is that a configured error regex means “the backend failed this turn”. The current code and test already show the counterexample: a worker may quote API Error: while writing a valid report. Two such false matches on one credential would now remove capacity for 60 seconds.

The cheapest experiment is a replay corpus before Unit 5 ships:

  1. Save the full pane text from the measured 2026-09-01 outage.
  2. Produce one safe failure per backend with a disposable invalid endpoint or request.
  3. Save one valid member report which quotes each error line.
  4. Replay all samples through the real CompletionResolver test fixture.
  5. Require outage samples to match and quoted-report samples not to match after assistant-block extraction and baseline checks.

This costs no outage deployment and no real sleep. If quoted reports still match, narrow the profile patterns before enabling correlation. Do not raise the threshold to hide a bad classifier.

Sequencing with three developers

First wave:

  1. Developer A: Unit 1, typed classification.
  2. Developer B: Unit 2, credential outage policy.
  3. Developer C: Unit 3, lead outage nudge.

As soon as one slot is free, start Unit 4. It is file-disjoint from every first-wave unit. Merge and review Units 1 to 4 independently.

Start Unit 5 only after all four foundations and #234 are available. Unit 5 is the only high-conflict integration branch, so no other active unit should touch its file list.

Checks performed for this refinement

  • Read issue #201 and issue #227 through their Gitea pages. Both showed zero comments.
  • Read the source and tests named in the evidence section.
  • Ran git status --short --branch; the branch was clean before this document was added.
  • Ran git log --oneline -12 to identify the branch base.
  • I did not run Maven because this change adds only a design document.
  • Rendered both Mermaid blocks with npx @mermaid-js/mermaid-cli; both commands succeeded.