A backend outage is invisible at fleet level: repeated backend errors are never correlated, never cool off a credential, and never reach the lead #227

Closed
opened 2026-09-01 11:10:25 +02:00 by ltms · 2 comments
Owner

Found live on 2026-09-01. A sonnet outage killed both running members mid-turn. fleetd handled each one correctly on its own and never noticed it was one event.

What actually happened

Both members ended with the same pane text:

API Error: No response from API

fleet_poll reported it honestly, per ticket:

[failed — member term_65a66c69faee03d ended on a backend error: API Error: No response from API]

Then nothing else happened. fleet_list showed both members as state: done, liveStatus: idle — the same row a member that finished cleanly gets. Capacity still reported sonnet free: 1. I found out by polling a ticket, and had to read a worktree by hand to learn whether any work had survived.

What already exists (do not rebuild it)

Three paths, and this one is the third:

Pane output What fleetd does
matches the profile's configured exhaustedPattern: BACKEND_EXHAUSTED → ExhaustionSink → BackendQuarantine → fleet_list shows free: 0 with credentialId and quarantinedForSeconds
matches the hardcoded /API Error:/ in CompletionResolver.java:83 fails that one send with the pane tail. No sink, no quarantine, no correlation
member reaches a terminal state FleetHealthMonitor detects it; notifications.mode is unset here, so coverage is detection-only

The second row is deliberate and documented at CompletionResolver.java:83: "Kept deliberately narrow — a growing list of ad-hoc error strings rots as backends change their wording; broader backend-error surfacing is out of scope here (fleetd#164 point 3)."

#201 already owns that classification work. This ticket is not that. Assume #201 lands and gives a trustworthy "this turn failed because of the backend" signal; this ticket is about what fleetd does with two of them.

The three gaps

  1. Nothing correlates. Two members, one credential, the same failure, inside a minute. fleetd saw two independent send failures. A per-turn view structurally cannot see an outage — the evidence only exists across members.
  2. Nothing stops the next spawn. A third sonnet spawn would have been accepted and would have died the same way. Capacity said free: 1, because capacity counts panes.
  3. The failure is invisible in the roster. The reason lives on the ticket and nowhere else. Once the ticket is collected or expires, a member that died on a backend error is indistinguishable from one that finished. See #197 — the ticket is not durable storage.

The design trap

An outage is not exhaustion, and reusing the exhaustion path would be a bug.

Exhaustion means the credential is spent, so a long quarantine is right. An outage is transient, often seconds. Quarantining sonnet for an hour because of a 30-second blip would remove the fleet's main capacity and turn a small outage into a large one. The two need different half-lives, and probably different names in fleet_list so an operator can tell "spent" from "flaky right now".

Whatever threshold is chosen must also survive the honest case: one backend error is often just one bad turn, and must not cool anything off.

The signal has to reach the lead

Health.Notifications.configured() returns true only for mode: webhook. So even at full coverage the health monitor posts to an external URL. That is right for an ops system and useless for the lead session, which reads its own pane.

The channel that works is the one CB-588 already uses: the async-ticket nudge injected into the lead's pane. An outage signal that does not use it will be missed exactly the way this one was.

Acceptance

  • Repeated backend-error classifications on one credentialId inside a short window are recognised as one event, not N.
  • That event applies a short cool-off, clearly distinct from exhaustion quarantine in both duration and in what fleet_list reports. A spawn onto a cooling-off credential is refused with a message saying it is cooling off and for how long — never silently accepted.
  • A single backend error changes nothing. State the threshold and the window, and say why those numbers.
  • The lead is told through the CB-588 pane-nudge path, not only a webhook. One message per event, not one per affected member.
  • A member that died on a backend error is distinguishable in fleet_list from one that finished cleanly, after its ticket is gone.
  • Tests drive the real path: two real classifications through the real resolver, then assert the spawn refusal and the roster. Asserting on the quarantine store alone walks around the gate — see #113.
  • Watch each test fail before keeping it.

Related

#201 (the classification mechanism this depends on) · #164 (points 3 and 4, where the narrowness was chosen) · #211 (the exhaustion path that already works, and the one not to copy wholesale) · #197 (why the ticket is not durable storage) · CB-588 (the nudge channel) · #176 (capacity counts panes, which is why free: 1 was wrong twice over).

Found live on 2026-09-01. A sonnet outage killed both running members mid-turn. fleetd handled each one correctly on its own and never noticed it was one event. ## What actually happened Both members ended with the same pane text: ``` API Error: No response from API ``` `fleet_poll` reported it honestly, per ticket: ``` [failed — member term_65a66c69faee03d ended on a backend error: API Error: No response from API] ``` Then nothing else happened. `fleet_list` showed both members as `state: done, liveStatus: idle` — the same row a member that finished cleanly gets. Capacity still reported `sonnet free: 1`. I found out by polling a ticket, and had to read a worktree by hand to learn whether any work had survived. ## What already exists (do not rebuild it) Three paths, and this one is the third: | Pane output | What fleetd does | |---|---| | matches the profile's configured `exhaustedPattern:` | `BACKEND_EXHAUSTED` → `ExhaustionSink` → `BackendQuarantine` → `fleet_list` shows `free: 0` with `credentialId` and `quarantinedForSeconds` | | matches the hardcoded `/API Error:/` in `CompletionResolver.java:83` | fails **that one send** with the pane tail. No sink, no quarantine, no correlation | | member reaches a terminal state | `FleetHealthMonitor` detects it; `notifications.mode` is unset here, so coverage is `detection-only` | The second row is deliberate and documented at `CompletionResolver.java:83`: *"Kept deliberately narrow — a growing list of ad-hoc error strings rots as backends change their wording; broader backend-error surfacing is out of scope here (fleetd#164 point 3)."* **#201 already owns that classification work.** This ticket is not that. Assume #201 lands and gives a trustworthy "this turn failed because of the backend" signal; this ticket is about what fleetd does with two of them. ## The three gaps 1. **Nothing correlates.** Two members, one credential, the same failure, inside a minute. fleetd saw two independent send failures. A per-turn view structurally cannot see an outage — the evidence only exists across members. 2. **Nothing stops the next spawn.** A third sonnet spawn would have been accepted and would have died the same way. Capacity said `free: 1`, because capacity counts panes. 3. **The failure is invisible in the roster.** The reason lives on the ticket and nowhere else. Once the ticket is collected or expires, a member that died on a backend error is indistinguishable from one that finished. See #197 — the ticket is not durable storage. ## The design trap **An outage is not exhaustion, and reusing the exhaustion path would be a bug.** Exhaustion means the credential is spent, so a long quarantine is right. An outage is transient, often seconds. Quarantining `sonnet` for an hour because of a 30-second blip would remove the fleet's main capacity and turn a small outage into a large one. The two need different half-lives, and probably different names in `fleet_list` so an operator can tell "spent" from "flaky right now". Whatever threshold is chosen must also survive the honest case: **one** backend error is often just one bad turn, and must not cool anything off. ## The signal has to reach the lead `Health.Notifications.configured()` returns true only for `mode: webhook`. So even at `full` coverage the health monitor posts to an external URL. That is right for an ops system and useless for the lead session, which reads its own pane. The channel that works is the one CB-588 already uses: the async-ticket nudge injected into the lead's pane. An outage signal that does not use it will be missed exactly the way this one was. ## Acceptance - Repeated backend-error classifications on one `credentialId` inside a short window are recognised as one event, not N. - That event applies a **short** cool-off, clearly distinct from exhaustion quarantine in both duration and in what `fleet_list` reports. A spawn onto a cooling-off credential is refused with a message saying it is cooling off and for how long — never silently accepted. - A single backend error changes nothing. State the threshold and the window, and say why those numbers. - The lead is told through the CB-588 pane-nudge path, not only a webhook. One message per event, not one per affected member. - A member that died on a backend error is distinguishable in `fleet_list` from one that finished cleanly, after its ticket is gone. - Tests drive the real path: two real classifications through the real resolver, then assert the spawn refusal and the roster. Asserting on the quarantine store alone walks around the gate — see #113. - Watch each test fail before keeping it. ## Related #201 (the classification mechanism this depends on) · #164 (points 3 and 4, where the narrowness was chosen) · #211 (the exhaustion path that already works, and the one not to copy wholesale) · #197 (why the ticket is not durable storage) · CB-588 (the nudge channel) · #176 (capacity counts panes, which is why `free: 1` was wrong twice over).
Author
Owner

Refined together with #201 — see #201 for the full plan

An architect member refined #201 and this issue as one delivery program in five units. The design is on main: docs/CB-201-227-Refinement.md (7662e2d). The detailed write-up is in #201's comment; this note records what belongs to this issue.

Unit 1: typed classification (#201) ----+
Unit 2: credential outage policy -------+
Unit 3: lead outage nudge --------------+--> Unit 5: production wiring and views
Unit 4: durable member outcome ---------+
#234 defect 2 fix ----------------------+

#201's Unit 1 produces a typed backend-error event from CompletionResolver. This issue consumes that event and never reads pane text. That is the seam that lets the two be built in parallel.

This issue's units

Unit 2 — credential outage policy. A new credential-keyed state machine, BackendOutagePolicy, on an injected monotonic clock. Threshold 2, window 60 seconds (inclusive), cool-off 60 seconds, correlated on credentialId and never on profile name or error text. One active incident per credential; errors during cool-off neither extend it nor raise a second notice; expiry clears evidence, so two fresh errors are needed to rearm.

Deliberately not BackendQuarantine: that restarts a 1800-second cooldown per exhaustion (:60-87). A short transient fault is not a spent credential, and reusing that store would make every refusal message say "backend exhausted".

Unit 3 — lead outage nudge. Backend incidents become a fourth source inside ReplyPushLoop (:20-48, :305-395), not a new scheduler. CB-588 already nudges the lead for every terminal async ticket (MessageService.java:922-940), but it cannot say that several failed tickets are one credential outage — that is what this adds. One notice per incident per affected lead (not per member), one-shot, waiting while the lead pane is not injectable, and combined into the same nudge as any pending failed ticket. A classified target that cannot be mapped to a credential produces a WARN, not a silent DEBUG.

Unit 4 — durable member outcome. BACKEND_ERROR joins MemberSession.State with a nullable reason, surfaced as failureReason in rosterView. Today the roster cannot preserve the cause once the ticket is gone (MemberSession.java:51-59, SessionManager.java:648-687), so the lead cannot tell "worker finished" from "backend died". The async resolver races SessionManager.onTurnComplete (:715-755), so the update must handle both BUSY -> BACKEND_ERROR and DONE -> BACKEND_ERROR atomically, and BACKEND_ERROR is terminal.

Unit 5 — wiring and views. Per-profile errorPattern beside exhaustedPattern, compiled once at startup, invalid regex stopping the boot by name. Both spawn paths learn a separate cool-off exclusion from quarantine, with exhaustion winning when both apply. fleet_list shows free: 0, the credential and coolingOffForSeconds; fleet_profiles gets its own coolingOff map; a direct spawn refusal says "cooling off after repeated backend errors" with the remaining seconds. FleetHealthMonitor.healthCoverage must not change to full because of any of this.

Against the measured 2026-09-01 outage

The second failed member crosses the threshold and starts cool-off. fleet_list then shows free: 0 for that credential and both members as backend_error with their reasons. ReplyPushLoop injects one outage notice even if neither ticket was ever polled. That tells the lead the fleet is down. It does not recover the lost work — recovery is explicitly out of scope for these tickets.

Units 1–4 are now delegated in parallel. Unit 5 waits on all four plus #234.

## Refined together with #201 — see #201 for the full plan An architect member refined #201 and this issue as **one delivery program in five units**. The design is on main: `docs/CB-201-227-Refinement.md` (`7662e2d`). The detailed write-up is in [#201's comment](https://git.ltms.dev/fleet/fleetd/issues/201#issuecomment-14258); this note records what belongs to **this** issue. ``` Unit 1: typed classification (#201) ----+ Unit 2: credential outage policy -------+ Unit 3: lead outage nudge --------------+--> Unit 5: production wiring and views Unit 4: durable member outcome ---------+ #234 defect 2 fix ----------------------+ ``` #201's Unit 1 produces a typed backend-error event from `CompletionResolver`. **This issue consumes that event and never reads pane text.** That is the seam that lets the two be built in parallel. ### This issue's units **Unit 2 — credential outage policy.** A new credential-keyed state machine, `BackendOutagePolicy`, on an injected monotonic clock. Threshold **2**, window **60 seconds** (inclusive), cool-off **60 seconds**, correlated on **`credentialId`** and never on profile name or error text. One active incident per credential; errors during cool-off neither extend it nor raise a second notice; expiry clears evidence, so two fresh errors are needed to rearm. Deliberately **not** `BackendQuarantine`: that restarts a 1800-second cooldown per exhaustion (`:60-87`). A short transient fault is not a spent credential, and reusing that store would make every refusal message say "backend exhausted". **Unit 3 — lead outage nudge.** Backend incidents become a **fourth source inside `ReplyPushLoop`** (`:20-48`, `:305-395`), not a new scheduler. CB-588 already nudges the lead for every terminal async ticket (`MessageService.java:922-940`), but it cannot say that several failed tickets are *one* credential outage — that is what this adds. One notice per incident per affected lead (not per member), one-shot, waiting while the lead pane is not injectable, and combined into the same nudge as any pending failed ticket. A classified target that cannot be mapped to a credential produces a **`WARN`**, not a silent `DEBUG`. **Unit 4 — durable member outcome.** `BACKEND_ERROR` joins `MemberSession.State` with a nullable reason, surfaced as `failureReason` in `rosterView`. Today the roster cannot preserve the cause once the ticket is gone (`MemberSession.java:51-59`, `SessionManager.java:648-687`), so the lead cannot tell "worker finished" from "backend died". The async resolver races `SessionManager.onTurnComplete` (`:715-755`), so the update must handle **both** `BUSY -> BACKEND_ERROR` and `DONE -> BACKEND_ERROR` atomically, and `BACKEND_ERROR` is terminal. **Unit 5 — wiring and views.** Per-profile `errorPattern` beside `exhaustedPattern`, compiled once at startup, invalid regex stopping the boot by name. Both spawn paths learn a **separate** cool-off exclusion from quarantine, with exhaustion winning when both apply. `fleet_list` shows `free: 0`, the credential and `coolingOffForSeconds`; `fleet_profiles` gets its own `coolingOff` map; a direct spawn refusal says "cooling off after repeated backend errors" with the remaining seconds. `FleetHealthMonitor.healthCoverage` must **not** change to `full` because of any of this. ### Against the measured 2026-09-01 outage The second failed member crosses the threshold and starts cool-off. `fleet_list` then shows `free: 0` for that credential and both members as `backend_error` with their reasons. `ReplyPushLoop` injects one outage notice **even if neither ticket was ever polled**. That tells the lead the fleet is down. It does not recover the lost work — recovery is explicitly out of scope for these tickets. Units 1–4 are now delegated in parallel. Unit 5 waits on all four plus #234.
Author
Owner

Done and merged, together with #201 — the two were refined into one five-unit series because a fleet-level correlation needs a per-target classification to correlate. See #201 for the unit-by-unit table and the merge notes.

What this ticket asked for specifically, and where it now lives:

  • Repeated backend errors are correlated — BackendOutagePolicy (959c835) keys on the credential, not the profile, because that is the thing that actually runs out. Two distinct targets failing inside 60 seconds is an incident; one target failing repeatedly is not, since that is a member problem, not a backend problem.
  • The fleet is told — the lead gets exactly one nudge when an incident opens (e5eb353), naming the credential and the affected targets. One, not one per error.
  • Placement stops feeding it — the credential cools off for 60 seconds. Automatic placement skips it; an explicit fleet_spawn naming it is refused before the adapter is ever called, with wording that says "cooling off", never "exhausted" (ac47498).
  • It is visible — fleet_list and fleet_profiles report coolingOffForSeconds alongside the separate CB-578 quarantinedForSeconds. Both can be true at once for the same profile, and they are reported as independent fields on purpose. Either one alone already sets that profile's free to 0.

Build on main: 1215 tests, 0 failures, 0 compile errors.

Not claimed: the startup coverage log line was verified by reading the code, not captured from a running daemon. The worker said so rather than overclaiming it, and I am repeating it here so nobody later reads it as measured. I will confirm it at the next redeploy.

Done and merged, together with #201 — the two were refined into one five-unit series because a fleet-level correlation needs a per-target classification to correlate. See #201 for the unit-by-unit table and the merge notes. What this ticket asked for specifically, and where it now lives: - **Repeated backend errors are correlated** — `BackendOutagePolicy` (`959c835`) keys on the *credential*, not the profile, because that is the thing that actually runs out. Two **distinct** targets failing inside 60 seconds is an incident; one target failing repeatedly is not, since that is a member problem, not a backend problem. - **The fleet is told** — the lead gets exactly one nudge when an incident opens (`e5eb353`), naming the credential and the affected targets. One, not one per error. - **Placement stops feeding it** — the credential cools off for 60 seconds. Automatic placement skips it; an explicit `fleet_spawn` naming it is refused before the adapter is ever called, with wording that says "cooling off", never "exhausted" (`ac47498`). - **It is visible** — `fleet_list` and `fleet_profiles` report `coolingOffForSeconds` alongside the separate CB-578 `quarantinedForSeconds`. Both can be true at once for the same profile, and they are reported as independent fields on purpose. Either one alone already sets that profile's `free` to `0`. Build on `main`: 1215 tests, 0 failures, 0 compile errors. **Not claimed:** the startup coverage log line was verified by reading the code, not captured from a running daemon. The worker said so rather than overclaiming it, and I am repeating it here so nobody later reads it as measured. I will confirm it at the next redeploy.
ltms closed this issue 2026-09-03 06:58:16 +02:00
Sign in to join this conversation.
1 Participants
Notifications
Due Date
No due date set.
Dependencies

No dependencies set.

Reference: fleet/fleetd#227