CB-578: a member killed by a provider usage limit is reported as a normal reply, and its work is lost #50

Closed
opened 2026-08-15 06:59:06 +02:00 by ltms · 3 comments
Owner

What happens now

A member runs on a metered backend. When that backend refuses the request — a subscription usage
limit — the agent process stays alive and its pane stays alive, but the turn ends with no
bridge_reply.

The bridge then does the worst available thing. The completion fallback scrapes the pane, clips it
to the 4000 character cap, and hands the lead a "reply" that is really an error message. The lead
cannot tell it from a real answer.

Observed live twice on 2026-08-15, on profile terra:

WARN [completion-term_6590e620fc60f5f] CompletionResolver - completion scrape for
term_6590e620fc60f5f clipped from 5195 chars to the 4000 char cap; member did not call
bridge_reply, so the pane tail is partial

The second of these was a probe whose correct answer was one short line. The pane held
The usage limit has been reached.

Three separate failures follow from one refusal:

  1. The lead is told the wrong thing. The ticket looks answered.
  2. The lead keeps dispatching into a dead account. Both members on terra died the same way.
  3. The member's work is lost. Its worktree still held an uncommitted build. It survived only
    because a human noticed and snapshotted it by hand.

The blast radius is the account, not the member

Checked on this host: sol and terra are both kind: opencode on openai/gpt-5.6-*, and
opencode auth list shows one OpenAI oauth credential. So one usage limit takes out the whole
opencode half of the fleet at once. Quarantining the member that happened to die first is useless;
the unit of exhaustion is the credential.

Why the fix must be reactive

There is no way to ask how much quota is left. Anthropic and OpenAI both return rate-limit headers
(anthropic-ratelimit-tokens-remaining, x-ratelimit-remaining-tokens) and both offer an admin
usage report, but all of that is for API keys. Every profile in this fleet runs on a
subscription seat, and no vendor exposes subscription quota to a script. Checked on the host:
neither codex --help nor claude --help has a usage, quota or limit subcommand.
opencode stats reports its own local token tally, not remaining vendor quota.

So the bridge cannot predict the refusal. It can only recognise it and react well.

Not checked yet: whether the gateway terra and sol reach exposes usage of its own. If it does,
that is a proactive signal and deserves its own ticket.

Scope

A — classify the refusal

Add a terminal health state BACKEND_EXHAUSTED, kept separate from GONE. GONE means the
pane died. Here the pane is healthy and the account is refusing; the two need different handling and
different operator advice.

Evidence: the turn ended with no bridge_reply and the scrape matches the profile's exhausted
pattern.

The pattern belongs in profile config, never in code. Each backend words its refusal
differently, and hardcoding one vendor's sentence makes the check wrong for every other profile.

The ticket outcome becomes BACKEND_EXHAUSTED carrying the matched line as its reason. This is
CB-568's rule — report the real cause, never a generic one — applied one level up.

B — quarantine the credential

Mark the credential, not the member and not only the profile. Profiles that share an account
share its limit, so quarantining terra alone still lets a spawn land on sol and die at once.

Profiles need a way to say which account they draw on — a credential: key, defaulting to the
profile name so today's behaviour is unchanged for profiles that do not set it.

While a credential is quarantined:

  • bridge_spawn on any profile using it refuses at once, naming the reason and the retry time.
  • The capacity view reports zero free slots for every one of those profiles.

The quarantine must expire on its own from a configured cooldown. A quarantine an operator has
to clear by hand is a quarantine that stays on forever. If the refusal text carries a reset time,
prefer it over the cooldown.

C — do not lose the work

State the limit honestly first: the dead member's context is gone. Nothing can continue that turn.
"Resume" means a fresh member, same brief, pointed at the same worktree — not a continuation.

That only works if the work reached disk. So: when a turn ends abnormally and the worktree is dirty,
commit it to a wip/<branch> ref before anything can release it.

This is the same disease as CB-576 (context-cap release deletes uncommitted worker output),
which is in flight now. Reuse that mechanism; do not grow a second one.

The failed ticket should then carry the worktree path, the branch, and the snapshot ref, so the lead
can re-dispatch onto the same tree instead of starting from the base commit.

codex resume --last exists and may give opencode-backed members a real context resume rather than
a re-brief. Untested. Out of scope here, worth its own ticket.

Acceptance criteria

  1. A turn that ends with no bridge_reply whose scrape matches the profile's exhausted pattern is
    classified BACKEND_EXHAUSTED, not reported as a completed reply.
  2. The exhausted pattern comes from profile config. No vendor wording appears in Java source.
  3. The ticket's outcome and reason name the real cause and carry the matched line.
  4. Quarantine applies to every profile sharing the credential, not only the one that died.
  5. credential: defaults to the profile name, so profiles that do not set it behave as today.
  6. bridge_spawn on a quarantined profile fails immediately with the reason and the retry time.
  7. The capacity view reports zero free slots for every quarantined profile.
  8. The quarantine expires without operator action.
  9. A dirty worktree is snapshotted before any release path can remove it.
  10. The failed ticket carries worktree, branch and snapshot ref.
  11. A profile with no configured pattern keeps today's behaviour exactly. Log the coverage at
    startup the way FleetHealthMonitor.coverage does, so the operator can see whether it is on.

Note on the silent-default trap — read before writing code

This repo has produced the same bug five times: a new dependency gets a default so existing wiring
compiles, and the capability ships turned off (CB-561, CB-572, CB-573 twice, TestTurnTokens).

Make new dependencies required. No convenience overload. For tests, pass an explicit inert value
that omits the fact rather than inventing a zero — TestTurnTokens.inert and
BridgeMcp.CapacitySource.none() are the patterns to copy.

An overload that defaults the new collaborator is an automatic rejection at review, because it also
means no test covers the new path.

## What happens now A member runs on a metered backend. When that backend refuses the request — a subscription usage limit — the agent process stays alive and its pane stays alive, but the turn ends with no `bridge_reply`. The bridge then does the worst available thing. The completion fallback scrapes the pane, clips it to the 4000 character cap, and hands the lead a "reply" that is really an error message. The lead cannot tell it from a real answer. Observed live twice on 2026-08-15, on profile `terra`: ``` WARN [completion-term_6590e620fc60f5f] CompletionResolver - completion scrape for term_6590e620fc60f5f clipped from 5195 chars to the 4000 char cap; member did not call bridge_reply, so the pane tail is partial ``` The second of these was a probe whose correct answer was one short line. The pane held `The usage limit has been reached`. Three separate failures follow from one refusal: 1. **The lead is told the wrong thing.** The ticket looks answered. 2. **The lead keeps dispatching into a dead account.** Both members on `terra` died the same way. 3. **The member's work is lost.** Its worktree still held an uncommitted build. It survived only because a human noticed and snapshotted it by hand. ## The blast radius is the account, not the member Checked on this host: `sol` and `terra` are both `kind: opencode` on `openai/gpt-5.6-*`, and `opencode auth list` shows **one** OpenAI oauth credential. So one usage limit takes out the whole opencode half of the fleet at once. Quarantining the member that happened to die first is useless; the unit of exhaustion is the credential. ## Why the fix must be reactive There is no way to ask how much quota is left. Anthropic and OpenAI both return rate-limit headers (`anthropic-ratelimit-tokens-remaining`, `x-ratelimit-remaining-tokens`) and both offer an admin usage report, but all of that is for **API keys**. Every profile in this fleet runs on a **subscription seat**, and no vendor exposes subscription quota to a script. Checked on the host: neither `codex --help` nor `claude --help` has a `usage`, `quota` or `limit` subcommand. `opencode stats` reports its own local token tally, not remaining vendor quota. So the bridge cannot predict the refusal. It can only recognise it and react well. Not checked yet: whether the gateway `terra` and `sol` reach exposes usage of its own. If it does, that is a proactive signal and deserves its own ticket. ## Scope ### A — classify the refusal Add a terminal health state `BACKEND_EXHAUSTED`, kept **separate from `GONE`**. `GONE` means the pane died. Here the pane is healthy and the account is refusing; the two need different handling and different operator advice. Evidence: the turn ended with no `bridge_reply` **and** the scrape matches the profile's exhausted pattern. The pattern belongs in **profile config**, never in code. Each backend words its refusal differently, and hardcoding one vendor's sentence makes the check wrong for every other profile. The ticket outcome becomes `BACKEND_EXHAUSTED` carrying the matched line as its reason. This is CB-568's rule — report the real cause, never a generic one — applied one level up. ### B — quarantine the credential Mark the **credential**, not the member and not only the profile. Profiles that share an account share its limit, so quarantining `terra` alone still lets a spawn land on `sol` and die at once. Profiles need a way to say which account they draw on — a `credential:` key, defaulting to the profile name so today's behaviour is unchanged for profiles that do not set it. While a credential is quarantined: - `bridge_spawn` on any profile using it refuses at once, naming the reason and the retry time. - The capacity view reports zero free slots for every one of those profiles. The quarantine **must expire on its own** from a configured cooldown. A quarantine an operator has to clear by hand is a quarantine that stays on forever. If the refusal text carries a reset time, prefer it over the cooldown. ### C — do not lose the work State the limit honestly first: the dead member's context is gone. Nothing can continue that turn. "Resume" means a fresh member, same brief, pointed at the same worktree — not a continuation. That only works if the work reached disk. So: when a turn ends abnormally and the worktree is dirty, commit it to a `wip/<branch>` ref **before anything can release it**. This is the same disease as **CB-576** (context-cap release deletes uncommitted worker output), which is in flight now. Reuse that mechanism; do not grow a second one. The failed ticket should then carry the worktree path, the branch, and the snapshot ref, so the lead can re-dispatch onto the same tree instead of starting from the base commit. `codex resume --last` exists and may give opencode-backed members a real context resume rather than a re-brief. Untested. Out of scope here, worth its own ticket. ## Acceptance criteria 1. A turn that ends with no `bridge_reply` whose scrape matches the profile's exhausted pattern is classified `BACKEND_EXHAUSTED`, not reported as a completed reply. 2. The exhausted pattern comes from profile config. No vendor wording appears in Java source. 3. The ticket's outcome and reason name the real cause and carry the matched line. 4. Quarantine applies to every profile sharing the credential, not only the one that died. 5. `credential:` defaults to the profile name, so profiles that do not set it behave as today. 6. `bridge_spawn` on a quarantined profile fails immediately with the reason and the retry time. 7. The capacity view reports zero free slots for every quarantined profile. 8. The quarantine expires without operator action. 9. A dirty worktree is snapshotted before any release path can remove it. 10. The failed ticket carries worktree, branch and snapshot ref. 11. A profile with no configured pattern keeps today's behaviour exactly. Log the coverage at startup the way `FleetHealthMonitor.coverage` does, so the operator can see whether it is on. ## Note on the silent-default trap — read before writing code This repo has produced the same bug five times: a new dependency gets a default so existing wiring compiles, and the capability ships turned off (CB-561, CB-572, CB-573 twice, `TestTurnTokens`). Make new dependencies **required**. No convenience overload. For tests, pass an explicit inert value that **omits** the fact rather than inventing a zero — `TestTurnTokens.inert` and `BridgeMcp.CapacitySource.none()` are the patterns to copy. An overload that defaults the new collaborator is an automatic rejection at review, because it also means no test covers the new path.
Author
Owner

Stages A and B are both merged to main.

  • Stage A — 2c2196a. Classify a usage-limit refusal instead of handing the scrape back as a real answer. Lead-verified at 704 tests, exit 0.
  • Stage B — 72d3481. Quarantine the exhausted credential so the fleet stops walking back onto it. Lead-verified at 740 tests, exit 0.

Stage B keys the quarantine on the credential rather than the profile name, because sol and terra are two models on one OpenAI account. Quarantining only the profile that reported the refusal would leave its sibling live and the next spawn would hit the same dead account. A profile that sets no credentialId quarantines alone under its own name, so an existing config behaves exactly as before.

Also fixed along the way, found by the implementer outside the brief: ConfigRef.sameLaunchSettings() never compared exhaustedPattern, so a reload changing only that key reported "config reloaded" while the value is actually deferred — the exact failure that file's own doc calls the worst outcome a reload can produce. Now compared, with a regression test.

Changed at merge review: BackendQuarantine.none() held a clock frozen at 0, so a quarantine() call on it recorded a deadline that could never pass — the credential would be locked out for the life of the daemon, and two production CompositePeerLauncher constructors default to none(). It is now a real no-op with a test.

Leaving this issue open for stage C (snapshot the work a member loses when its backend refuses), which is the remaining half of the original report — the title says the work is lost, and nothing yet saves it.

Stages A and B are both merged to `main`. - **Stage A** — `2c2196a`. Classify a usage-limit refusal instead of handing the scrape back as a real answer. Lead-verified at 704 tests, exit 0. - **Stage B** — `72d3481`. Quarantine the exhausted **credential** so the fleet stops walking back onto it. Lead-verified at 740 tests, exit 0. Stage B keys the quarantine on the credential rather than the profile name, because `sol` and `terra` are two models on one OpenAI account. Quarantining only the profile that reported the refusal would leave its sibling live and the next spawn would hit the same dead account. A profile that sets no `credentialId` quarantines alone under its own name, so an existing config behaves exactly as before. Also fixed along the way, found by the implementer outside the brief: `ConfigRef.sameLaunchSettings()` never compared `exhaustedPattern`, so a reload changing only that key reported "config reloaded" while the value is actually deferred — the exact failure that file's own doc calls the worst outcome a reload can produce. Now compared, with a regression test. Changed at merge review: `BackendQuarantine.none()` held a clock frozen at 0, so a `quarantine()` call on it recorded a deadline that could never pass — the credential would be locked out for the life of the daemon, and two production `CompositePeerLauncher` constructors default to `none()`. It is now a real no-op with a test. **Leaving this issue open for stage C** (snapshot the work a member loses when its backend refuses), which is the remaining half of the original report — the title says the work is lost, and nothing yet saves it.
Author
Owner

All three stages are merged to main. Closing.

Stage What it does Merge
A Classify a usage-limit refusal as a backend failure, not a reply 2c2196a
B Quarantine the credential, not the profile, so siblings on one account go out together 72d3481
C Snapshot a dirty worktree to refs/wip/<branch> before it can be lost a3842c8
criterion 7 (CB-583) bridge_list capacity reports free: 0 for a quarantined profile 758d62a

My own build on the merged tree: 764 tests, 0 failures, BUILD SUCCESS, exit 0 (mvn -f bridged/pom.xml clean install, unpiped).

Stage C verified by hand, not only by its tests

I ran the mechanism on a throwaway repo to check the two claims the feature rests on:

  • After git worktree remove --force, refs/wip/<branch> still resolves and git show refs/wip/<branch>:<file> returns the work-in-progress content. The ref lives in the shared ref store, so removing the worktree does not touch it.
  • A gitignored file in the worktree is not in the snapshot's tree, while a non-ignored untracked file is. add -A respects .gitignore, so a secret does not reach the commit.
  • refs/wip/* does not appear in git branch, so a later git branch -d sweep cannot destroy it.

Two things this change makes load-bearing

  1. The parityOverlay gitignore rule is now a security rule. The overlay copies the primary's environment files into every worktree. Before stage C a non-gitignored overlay path merely sat on disk; now it would be committed into a durable git object that outlives the worktree. .gitignore is what stops that. Documented on BridgedConfig.parityOverlay in 3a10f6a.
  2. refs/wip/* refs pin objects and are never cleaned up. Every snapshot holds a full tree, and nothing prunes them, so git gc can never reclaim that space. Not a problem yet — worth a follow-up before the fleet has run for months.

Still not live

exhaustedPattern is deliberately unset on every profile, so stage A's classification — and therefore stages B and C's quarantine path — stay dormant. The reason is in bridged.yaml: the pattern is matched against the pane scrape of any turn that ended without bridge_reply, and workers in this repo write about usage limits constantly. A broad pattern would let a worker quarantine its own credential by describing this feature. A real vendor refusal has to be captured from a live pane first.

Stage C's snapshot does not depend on that pattern — it fires on any dirty worktree at release, so it is live now.

All three stages are merged to `main`. Closing. | Stage | What it does | Merge | |---|---|---| | A | Classify a usage-limit refusal as a backend failure, not a reply | `2c2196a` | | B | Quarantine the **credential**, not the profile, so siblings on one account go out together | `72d3481` | | C | Snapshot a dirty worktree to `refs/wip/<branch>` before it can be lost | `a3842c8` | | criterion 7 (CB-583) | `bridge_list` capacity reports `free: 0` for a quarantined profile | `758d62a` | My own build on the merged tree: **764 tests, 0 failures, BUILD SUCCESS, exit 0** (`mvn -f bridged/pom.xml clean install`, unpiped). ### Stage C verified by hand, not only by its tests I ran the mechanism on a throwaway repo to check the two claims the feature rests on: - After `git worktree remove --force`, `refs/wip/<branch>` **still resolves** and `git show refs/wip/<branch>:<file>` returns the work-in-progress content. The ref lives in the shared ref store, so removing the worktree does not touch it. - A **gitignored** file in the worktree is **not** in the snapshot's tree, while a non-ignored untracked file **is**. `add -A` respects `.gitignore`, so a secret does not reach the commit. - `refs/wip/*` does not appear in `git branch`, so a later `git branch -d` sweep cannot destroy it. ### Two things this change makes load-bearing 1. **The `parityOverlay` gitignore rule is now a security rule.** The overlay copies the primary's environment files into every worktree. Before stage C a non-gitignored overlay path merely sat on disk; now it would be committed into a durable git object that outlives the worktree. `.gitignore` is what stops that. Documented on `BridgedConfig.parityOverlay` in `3a10f6a`. 2. **`refs/wip/*` refs pin objects and are never cleaned up.** Every snapshot holds a full tree, and nothing prunes them, so `git gc` can never reclaim that space. Not a problem yet — worth a follow-up before the fleet has run for months. ### Still not live `exhaustedPattern` is deliberately unset on every profile, so stage A's classification — and therefore stages B and C's quarantine path — stay **dormant**. The reason is in `bridged.yaml`: the pattern is matched against the pane scrape of any turn that ended without `bridge_reply`, and workers in this repo write about usage limits constantly. A broad pattern would let a worker quarantine its own credential by describing this feature. A real vendor refusal has to be captured from a live pane first. Stage C's snapshot does **not** depend on that pattern — it fires on any dirty worktree at release, so it is live now.
ltms closed this issue 2026-08-15 13:10:13 +02:00
Author
Owner

Correction to my closing comment above, after dogfooding stage C on the live daemon.

The daemon was restarted onto the stage C jar and the feature works end to end: a probe worktree was preserved and snapshotted to refs/wip/<branch>, and both the modified and the untracked file were recoverable.

But I found a defect the tests do not catch, filed as issue #68 (CB-587).

I wrote above that "a secret does not reach the commit". That is correct for gitignored files — add -A respects .gitignore and I confirmed it. It is not the whole picture. --skip-worktree tracked files are a separate category, and the snapshot ignores that flag entirely, because it stages into a fresh temporary index and index flags do not survive that.

Measured on the live daemon: a worktree whose git status --porcelain reported 2 changed paths produced a snapshot whose diff contains 15, including .mcp.json — the file CLAUDE.md names as never-commit — plus opencode.json and the wiki submodule pointer.

The leak in this instance was benign (the worktree's .mcp.json is a 23-byte stub, not the primary's 231-byte real one), but that is luck rather than design. The bigger practical harm is that a recovered snapshot misreports what the worker actually changed, which undercuts the reason stage C exists.

Stage C stays merged and live — it is a clear improvement over losing the work — and #68 tracks making the snapshot faithful.

**Correction to my closing comment above, after dogfooding stage C on the live daemon.** The daemon was restarted onto the stage C jar and the feature works end to end: a probe worktree was preserved and snapshotted to `refs/wip/<branch>`, and both the modified and the untracked file were recoverable. But I found a defect the tests do not catch, filed as issue #68 (CB-587). I wrote above that "a secret does not reach the commit". That is correct **for gitignored files** — `add -A` respects `.gitignore` and I confirmed it. It is **not** the whole picture. `--skip-worktree` tracked files are a separate category, and the snapshot ignores that flag entirely, because it stages into a fresh temporary index and index flags do not survive that. Measured on the live daemon: a worktree whose `git status --porcelain` reported **2** changed paths produced a snapshot whose diff contains **15**, including `.mcp.json` — the file `CLAUDE.md` names as never-commit — plus `opencode.json` and the `wiki` submodule pointer. The leak in this instance was benign (the worktree's `.mcp.json` is a 23-byte stub, not the primary's 231-byte real one), but that is luck rather than design. The bigger practical harm is that a recovered snapshot misreports what the worker actually changed, which undercuts the reason stage C exists. Stage C stays merged and live — it is a clear improvement over losing the work — and #68 tracks making the snapshot faithful.
Sign in to join this conversation.
1 Participants
Notifications
Due Date
No due date set.
Dependencies

No dependencies set.

Reference: fleet/fleetd#50