CB-201 unit 2: credential-keyed backend outage policy #239

Closed
agent wants to merge 0 commits from worker/cb201-unit2-policy-c1102c-5 into main
Member

Part of fleetd #201 / #227 (unit 2 of 5, parallel with units 1/3/4/5).

Adds BackendOutagePolicy (placement package): a credential-keyed state machine that decides when a backend is out from classified backend-error events, independent of panes, profiles, sessions, launchers, or leads.

  • Threshold: 2 classified errors on the same credentialId (never the profile name, never error text)
  • Window: 60s, inclusive both ends (exactly 60s apart still counts)
  • Cool-off: 60s starting at the moment the threshold crosses
  • One active incident per credential; errors during cool-off are ignored (no deadline extension, no repeat incident)
  • Cool-off expiry clears evidence; two fresh errors are required to rearm
  • Correlation + cool-off live in one class; threshold-crossing and setting the deadline are one atomic ConcurrentHashMap.compute() update per credential, so a concurrent second and third event can never both mint an incident
  • Deliberately NOT BackendQuarantine (that is the 1800s exhaustion store with different, misleading semantics)

Tests: BackendOutagePolicyTest (10 tests), including a real CountDownLatch-based concurrency test asserting exactly one incident across 8 threads racing the same credential, looped 25x internally and re-run 3x separately with no flake.

Mutation-proofed manually (not committed) for the window boundary, cool-off non-extension, evidence-clear-on-expiry, and the lock — each showed a real red assertion failure before being reverted; see the worker's fleet_reply for the exact output.

Build: mvn clean install — Tests run: 1133, Failures: 0, Errors: 0, Skipped: 0, BUILD SUCCESS.

Part of fleetd #201 / #227 (unit 2 of 5, parallel with units 1/3/4/5). Adds `BackendOutagePolicy` (placement package): a credential-keyed state machine that decides when a backend is out from classified backend-error events, independent of panes, profiles, sessions, launchers, or leads. - Threshold: 2 classified errors on the same `credentialId` (never the profile name, never error text) - Window: 60s, inclusive both ends (exactly 60s apart still counts) - Cool-off: 60s starting at the moment the threshold crosses - One active incident per credential; errors during cool-off are ignored (no deadline extension, no repeat incident) - Cool-off expiry clears evidence; two fresh errors are required to rearm - Correlation + cool-off live in one class; threshold-crossing and setting the deadline are one atomic `ConcurrentHashMap.compute()` update per credential, so a concurrent second and third event can never both mint an incident - Deliberately NOT `BackendQuarantine` (that is the 1800s exhaustion store with different, misleading semantics) Tests: `BackendOutagePolicyTest` (10 tests), including a real `CountDownLatch`-based concurrency test asserting exactly one incident across 8 threads racing the same credential, looped 25x internally and re-run 3x separately with no flake. Mutation-proofed manually (not committed) for the window boundary, cool-off non-extension, evidence-clear-on-expiry, and the lock — each showed a real red assertion failure before being reverted; see the worker's fleet_reply for the exact output. Build: `mvn clean install` — Tests run: 1133, Failures: 0, Errors: 0, Skipped: 0, BUILD SUCCESS.
agent added 1 commit 2026-09-03 05:31:02 +02:00
CB-201 unit 2: credential-keyed backend outage policy
CI / build (pull_request) Successful in 1m11s
CI / contract (pull_request) Successful in 1m24s
826e0aeb2a
Add BackendOutagePolicy: two classified backend errors on the same
credentialId within a 60s window mint one Incident and start a 60s
cool-off for that credential, one atomic ConcurrentHashMap.compute()
per credentialId so a concurrent second and third event can never
both cross the threshold. Errors during cool-off are ignored outright
(no extension, no incident); once cool-off elapses the next error
clears old evidence, requiring two fresh errors to rearm. This is a
new class, deliberately not BackendQuarantine (wrong store, wrong
1800s duration, misleading "exhausted" semantics for a 60s transient
fault). Knows nothing about panes, profiles, sessions, launchers, or
leads — takes events in, returns incidents out.
agent added 1 commit 2026-09-03 05:42:44 +02:00
CB-201 unit 2 review fix: threshold counts distinct targets, not raw events
CI / contract (pull_request) Successful in 58s
CI / build (pull_request) Successful in 1m46s
cf54aed451
Two errors from the same target inside the window must never trip the
outage threshold on their own (a valid member report can legitimately
quote an "API Error:" line twice) — only two DIFFERENT targets on the
same credential do. Change evidenceCount() to targets.size() instead
of reasons.size(); reasons() still keeps every event, including
same-target repeats, so it can be longer than evidenceCount(). A real
outage still hits every target on the credential, so this loses no
true-positive coverage while cutting a real false-positive path.
Owner

Merged to main in 959c835.

Verified by the lead: baseline 1135 tests, 0 failures; mutating evidenceCount() back to reasons.size() gives 2 reds with 0 compile errors.

One lead change before the merge, in cf54aed: the threshold now counts distinct targets, not raw events. The reasoning, so it is not re-litigated later:

  • The classifier is a heuristic on pane text. One member that quotes API Error: twice in a turn must not be able to remove a healthy credential's capacity on its own.
  • A real credential outage hits every member on that credential, so distinct targets is the signal that actually separates an outage from noise.
  • reasons still keeps every event, so nothing is lost for the notice text — only the threshold test changed.

I also did not reuse BackendQuarantine. That one is keyed on a spawn-time model check and holds a different lifecycle; folding both into one class would have coupled two unrelated triggers to one cool-off clock.

Merged to main in `959c835`. Verified by the lead: baseline **1135 tests, 0 failures**; mutating `evidenceCount()` back to `reasons.size()` gives **2 reds with 0 compile errors**. One lead change before the merge, in `cf54aed`: the threshold now counts **distinct targets**, not raw events. The reasoning, so it is not re-litigated later: - The classifier is a heuristic on pane text. One member that quotes `API Error:` twice in a turn must not be able to remove a healthy credential's capacity on its own. - A real credential outage hits every member on that credential, so distinct targets is the signal that actually separates an outage from noise. - `reasons` still keeps every event, so nothing is lost for the notice text — only the threshold test changed. I also did not reuse `BackendQuarantine`. That one is keyed on a spawn-time model check and holds a different lifecycle; folding both into one class would have coupled two unrelated triggers to one cool-off clock.
ltms closed this pull request 2026-09-03 05:57:55 +02:00
Some checks are pending
CI / contract (pull_request) Successful in 58s
CI / build (pull_request) Successful in 1m46s

Pull request closed

Sign in to join this conversation.