CB-578 stage B: quarantine the exhausted credential, not the profile #60

Closed
agent wants to merge 0 commits from worker/cb578b-9dcb13-6 into main
Member

Closes #50.

When CompletionResolver classifies a turn BACKEND_EXHAUSTED (stage A, already merged), this stage acts on it instead of letting a fresh spawn walk straight back onto the same exhausted account.

What changed

  • New BackendQuarantine (placement pkg): keyed by credential id, injected monotonic clock, .none() inert stand-in — no defaulted dependency.
  • New Profile.credentialId (HOT) + effectiveCredentialId(): profiles sharing a credentialId quarantine together (e.g. two models on one OpenAI account); an unset profile quarantines alone, unchanged from today.
  • New top-level quarantineCooldownSeconds (DEFERRED, default 1800s).
  • New ExhaustionSink (inject pkg, .none() stand-in): CompletionResolver calls it only when Rendezvous.resolveExhausted actually wins the race, so a late/duplicate classification never double-quarantines.
  • CompositePeerLauncher: explicit-profile spawns refuse on a quarantined credential (message names the profile, the credential, and ~seconds remaining); placement policies (fixed/round-robin/weighted) exclude quarantined candidates via PlacementContext.quarantined().
  • bridge_profiles (BridgeMcp.QuarantineSource) reports each quarantined profile's credential and remaining seconds.
  • Found and fixed along the way: ConfigRef.sameLaunchSettings() didn't compare exhaustedPattern, so a reload changing only that key was silently reported "applied" even though it's actually deferred (compiled once into a startup pattern map). Fixed the comparison and corrected the docs (yaml + ConfigRef javadoc) to mark it DEFERRED — this was stage A's own undocumented mistake, now closed.

Design note not settled by the brief: I asked the lead (bridge_ask) about the credentialId field name/semantics before building it, since it's a config surface; the ask timed out with no answer (55s, no primary response), so I proceeded on my own judgment as instructed and flagged it here. Semantics: nullable string, defaults to the profile's own name via effectiveCredentialId() when unset, so existing configs behave identically.

Constraints honored: no new defaulted dependency (every new collaborator has a required-param constructor + explicit .none()); cooldown is config, documented default, never hardcoded; no vendor wording in Java; clocks injected (LongSupplier, matches FleetHealthMonitor's pattern); did NOT touch session/SessionManager.java (CB-581 is changing it concurrently) — used only its existing public roster()/read methods via a lambda in Bridged.java, mirroring stage A's own pattern.

Tests: mvn -f bridged/pom.xml clean install, unpiped — Tests run: 732, Failures: 0, Errors: 0, Skipped: 0, BUILD SUCCESS, exit 0. New: BackendQuarantineTest (8), quarantine-specific CompositePeerLauncherTest cases (explicit refusal naming profile+credential+remaining, two profiles sharing a credential both quarantined by one event, placement skip, clock-driven expiry, quarantine-inert-by-default regression), PlacementPolicyTest quarantine cases for fixed/round-robin/weighted, CompletionResolverTest cases asserting the sink fires only on a winning classification, and a ConfigRefTest regression for the exhaustedPattern deferred-detection fix.

Config keys added (see bridged.example.yaml):

  • profiles.<name>.credentialId — HOT
  • quarantineCooldownSeconds (top-level) — DEFERRED, default 1800
Closes #50. When CompletionResolver classifies a turn BACKEND_EXHAUSTED (stage A, already merged), this stage acts on it instead of letting a fresh spawn walk straight back onto the same exhausted account. **What changed** - New `BackendQuarantine` (placement pkg): keyed by *credential id*, injected monotonic clock, `.none()` inert stand-in — no defaulted dependency. - New `Profile.credentialId` (HOT) + `effectiveCredentialId()`: profiles sharing a `credentialId` quarantine together (e.g. two models on one OpenAI account); an unset profile quarantines alone, unchanged from today. - New top-level `quarantineCooldownSeconds` (DEFERRED, default 1800s). - New `ExhaustionSink` (inject pkg, `.none()` stand-in): `CompletionResolver` calls it only when `Rendezvous.resolveExhausted` actually wins the race, so a late/duplicate classification never double-quarantines. - `CompositePeerLauncher`: explicit-profile spawns refuse on a quarantined credential (message names the profile, the credential, and ~seconds remaining); placement policies (fixed/round-robin/weighted) exclude quarantined candidates via `PlacementContext.quarantined()`. - `bridge_profiles` (`BridgeMcp.QuarantineSource`) reports each quarantined profile's credential and remaining seconds. - Found and fixed along the way: `ConfigRef.sameLaunchSettings()` didn't compare `exhaustedPattern`, so a reload changing only that key was silently reported "applied" even though it's actually deferred (compiled once into a startup pattern map). Fixed the comparison and corrected the docs (yaml + ConfigRef javadoc) to mark it DEFERRED — this was stage A's own undocumented mistake, now closed. **Design note not settled by the brief**: I asked the lead (`bridge_ask`) about the `credentialId` field name/semantics before building it, since it's a config surface; the ask timed out with no answer (55s, no primary response), so I proceeded on my own judgment as instructed and flagged it here. Semantics: nullable string, defaults to the profile's own name via `effectiveCredentialId()` when unset, so existing configs behave identically. **Constraints honored**: no new defaulted dependency (every new collaborator has a required-param constructor + explicit `.none()`); cooldown is config, documented default, never hardcoded; no vendor wording in Java; clocks injected (`LongSupplier`, matches `FleetHealthMonitor`'s pattern); did NOT touch `session/SessionManager.java` (CB-581 is changing it concurrently) — used only its existing public `roster()`/read methods via a lambda in `Bridged.java`, mirroring stage A's own pattern. **Tests**: `mvn -f bridged/pom.xml clean install`, unpiped — `Tests run: 732, Failures: 0, Errors: 0, Skipped: 0`, `BUILD SUCCESS`, exit 0. New: `BackendQuarantineTest` (8), quarantine-specific `CompositePeerLauncherTest` cases (explicit refusal naming profile+credential+remaining, two profiles sharing a credential both quarantined by one event, placement skip, clock-driven expiry, quarantine-inert-by-default regression), `PlacementPolicyTest` quarantine cases for fixed/round-robin/weighted, `CompletionResolverTest` cases asserting the sink fires only on a winning classification, and a `ConfigRefTest` regression for the exhaustedPattern deferred-detection fix. **Config keys added** (see `bridged.example.yaml`): - `profiles.<name>.credentialId` — HOT - `quarantineCooldownSeconds` (top-level) — DEFERRED, default 1800
agent added 1 commit 2026-08-15 10:30:21 +02:00
CB-578 stage B: quarantine the exhausted credential, not the profile
CI / build (pull_request) Successful in 55s
CI / contract (pull_request) Successful in 1m21s
e501d39988
A BACKEND_EXHAUSTED classification (stage A) now puts that profile's
credential into a BackendQuarantine for a configurable cooldown. A spawn
onto a quarantined profile is refused naming the credential and roughly
when it lifts; weighted/round-robin/fixed placement skip a quarantined
candidate; the quarantine lifts itself on the injected clock; and it is
visible on bridge_profiles.

Keyed by credential, not by profile name, via the new Profile.credentialId
(profiles sharing one credential quarantine together — e.g. two models on
one account) and effectiveCredentialId() (unset ⇒ quarantines alone,
today's behaviour unchanged). Fixed a related gap along the way: a reload
changing exhaustedPattern was silently reported "applied" even though it's
deferred — sameLaunchSettings() now catches it too.
ltms closed this pull request 2026-08-15 10:34:19 +02:00
Some checks are pending
CI / build (pull_request) Successful in 55s
CI / contract (pull_request) Successful in 1m21s

Pull request closed

Sign in to join this conversation.