capacity counts panes, not subscription seats, so a subscription profile reports a free slot that cannot be filled #176

Closed
opened 2026-08-28 00:52:34 +02:00 by ltms · 7 comments
Owner

What happened

Measured on the Mac fleet, 2026-08-28. The sonnet profile is kind: claude-code with subscription: true and maxLoad: 3.

With two sonnet members up and the fleet otherwise idle, fleet_list reported:

{"profile":"sonnet","maxLoad":3,"live":2,"free":1,"reclaimable":2}

A third fleet_spawn{profile: sonnet} failed:

spawn timed out — worker pane never reached injectable state:
worker pane w6:pK did not reach injectable state within 20000ms

Retried once the fleet was idle. It failed again, on a new pane. Both times the daemon log shows the tab was already gone before the readiness gate tried to close it:

05:48:42.795 WARN  peer pane=w6:pK did not become injectable within 20000ms — closing
05:48:42.928 DEBUG tab.close(w6:tK) ignored — already gone: herdr error [tab_not_found]

tab_not_found means the claude process exited, rather than being slow to start. A hang would leave the tab in place.

The likely cause, and the part that is certain

Likely: the lead is itself a claude session on the same subscription, so it holds a seat. Two workers plus the lead reaches the subscription's concurrency limit, and the fourth claude exits at once.

I have not confirmed that limit exists or what its value is — no subscription usage or concurrency API is available to check against, and reading the pane to see the process's own error message would mean driving herdr directly. So treat the specific cause as unconfirmed.

What is certain, and is the actual defect:

  1. Two identical spawns failed against a profile the daemon reported as having a free slot.
  2. free: 1 was wrong, and the lead was told to expect capacity that did not exist.
  3. The failure surfaced as a 20-second readiness timeout — a message about the pane being slow — when the process had in fact already died. That points an investigation at the readiness gate instead of at capacity.

Why it matters

free is the number a lead plans against. CLAUDE.md tells the lead to spawn every delegated unit before sending any, so a wrong free is discovered part-way through a fan-out, after some units are already briefed. That is the worst moment to find out.

It also costs 20 seconds per attempt to learn nothing useful, and it provisions and then removes a git worktree each time.

Suggested direction

  1. Count the seat the lead occupies. A profile with subscription: true shares its budget with the lead when the lead runs on the same subscription. Either subtract the lead's seat from maxLoad, or make the relationship explicit in config so the operator states it once.
  2. Report an immediate exit as an exit, not as a timeout. tab_not_found at close time is direct evidence the process died. Say so, rather than reporting a readiness timeout — the two have completely different fixes, and the current message names the wrong one.
  3. Do not silently retry into the same wall. After a spawn fails this way for a profile, that profile should be marked so the next spawn either fails fast or picks a different profile.

Workaround in place

sonnet.maxLoad is left at 3 on purpose, with a comment recording the measurement. Lowering it to 2 would hide the counting bug rather than fix it, and 2 is only correct while the lead runs on this same subscription.

Related

  • #175 — the other case of fleetd asking a backend for something and never checking it got it.
## What happened Measured on the Mac fleet, 2026-08-28. The `sonnet` profile is `kind: claude-code` with `subscription: true` and `maxLoad: 3`. With **two** sonnet members up and the fleet otherwise idle, `fleet_list` reported: ```json {"profile":"sonnet","maxLoad":3,"live":2,"free":1,"reclaimable":2} ``` A third `fleet_spawn{profile: sonnet}` failed: ``` spawn timed out — worker pane never reached injectable state: worker pane w6:pK did not reach injectable state within 20000ms ``` Retried once the fleet was idle. It failed again, on a new pane. Both times the daemon log shows the tab was **already gone** before the readiness gate tried to close it: ``` 05:48:42.795 WARN peer pane=w6:pK did not become injectable within 20000ms — closing 05:48:42.928 DEBUG tab.close(w6:tK) ignored — already gone: herdr error [tab_not_found] ``` `tab_not_found` means the `claude` process **exited**, rather than being slow to start. A hang would leave the tab in place. ## The likely cause, and the part that is certain Likely: the lead is itself a `claude` session on the same subscription, so it holds a seat. Two workers plus the lead reaches the subscription's concurrency limit, and the fourth `claude` exits at once. I have **not** confirmed that limit exists or what its value is — no subscription usage or concurrency API is available to check against, and reading the pane to see the process's own error message would mean driving herdr directly. So treat the specific cause as unconfirmed. What **is** certain, and is the actual defect: 1. Two identical spawns failed against a profile the daemon reported as having a free slot. 2. `free: 1` was wrong, and the lead was told to expect capacity that did not exist. 3. The failure surfaced as a 20-second readiness timeout — a message about the pane being slow — when the process had in fact already died. That points an investigation at the readiness gate instead of at capacity. ## Why it matters `free` is the number a lead plans against. CLAUDE.md tells the lead to spawn every delegated unit before sending any, so a wrong `free` is discovered part-way through a fan-out, after some units are already briefed. That is the worst moment to find out. It also costs 20 seconds per attempt to learn nothing useful, and it provisions and then removes a git worktree each time. ## Suggested direction 1. **Count the seat the lead occupies.** A profile with `subscription: true` shares its budget with the lead when the lead runs on the same subscription. Either subtract the lead's seat from `maxLoad`, or make the relationship explicit in config so the operator states it once. 2. **Report an immediate exit as an exit, not as a timeout.** `tab_not_found` at close time is direct evidence the process died. Say so, rather than reporting a readiness timeout — the two have completely different fixes, and the current message names the wrong one. 3. **Do not silently retry into the same wall.** After a spawn fails this way for a profile, that profile should be marked so the next spawn either fails fast or picks a different profile. ## Workaround in place `sonnet.maxLoad` is left at `3` on purpose, with a comment recording the measurement. Lowering it to `2` would hide the counting bug rather than fix it, and `2` is only correct while the lead runs on this same subscription. ## Related - #175 — the other case of fleetd asking a backend for something and never checking it got it.
Author
Owner

Fresh data point, 2026-08-28, and it is worse than the original report.

fleet_list showed a completely idle fleet — zero members, sonnet reporting maxLoad: 3, live: 0, free: 3. A single fleet_spawn{profile: sonnet} still failed:

spawn timed out — worker pane never reached injectable state:
worker pane w6:pY did not reach injectable state within 20000ms

Same signature as before. terra spawned fine in the same minute, twice, so herdr and the spawn path are healthy — it is the Claude subscription backend refusing to start.

So this is now two distinct problems wearing one symptom:

  1. The capacity counter counts panes, not subscription seats — the original report. The lead holds a seat, so free: 3 was never true; 2 was the real ceiling.
  2. Right now the backend will not start a claude-code member at all, even for the first one. free: 3 on an idle fleet is simply wrong, and nothing in fleet_list hints at it.

Both point the same way: the daemon reports capacity it has never verified. It treats "the process was asked to start" as success and never reads back whether a seat existed. A spawn-time probe, or marking the profile quarantined after a readiness timeout the way credential quarantine already works, would stop free: N from being a claim the daemon cannot support.

Worked around by moving both #185 units onto terra.

Fresh data point, 2026-08-28, and it is worse than the original report. `fleet_list` showed a **completely idle fleet** — zero members, `sonnet` reporting `maxLoad: 3, live: 0, free: 3`. A single `fleet_spawn{profile: sonnet}` still failed: ``` spawn timed out — worker pane never reached injectable state: worker pane w6:pY did not reach injectable state within 20000ms ``` Same signature as before. `terra` spawned fine in the same minute, twice, so herdr and the spawn path are healthy — it is the Claude subscription backend refusing to start. So this is now two distinct problems wearing one symptom: 1. **The capacity counter counts panes, not subscription seats** — the original report. The lead holds a seat, so `free: 3` was never true; 2 was the real ceiling. 2. **Right now the backend will not start a `claude-code` member at all**, even for the first one. `free: 3` on an idle fleet is simply wrong, and nothing in `fleet_list` hints at it. Both point the same way: the daemon reports capacity it has never verified. It treats "the process was asked to start" as success and never reads back whether a seat existed. A spawn-time probe, or marking the profile quarantined after a readiness timeout the way credential quarantine already works, would stop `free: N` from being a claim the daemon cannot support. Worked around by moving both #185 units onto `terra`.
Author
Owner

2026-08-29 — the cause is daemon-lifetime state, not the claude backend. A restart fixes it.

Earlier notes on this issue said "the Claude subscription backend is failing". That was wrong. The
backend is healthy. What fails is a fleetd process that has been running for a long time.

What I ruled out, each by measurement

Every one of these was tested on this host with claude 2.1.251, outside a herdr pane:

Suspect Test Result
auth / login claude -p "OK" --model claude-sonnet-5 OK, rc=0
usage limit same OK
the launcher's flags -p plus --mcp-config + --append-system-prompt + --agent dev + --autocompact 250000 + --session-id OK
the memberCredentials scrub reproduced the CB-633 allow-list in a zsh (105 of 125 variables blanked), then started claude TUI renders
interactive mode / pty real pty via pty.fork(), 40x120 winsize, full member flag set TUI renders in <4s
scrub and pty together both of the above in one run TUI renders
a subscription seat limit three concurrent interactive claude sessions all three render

So nothing in the argv, the environment, the terminal, or the account explains it.

What the log actually shows

Sonnet spawns broke mid-session on 2026-08-29. The daemon had been up since 2026-08-28 06:36.

06:16:42  spawning claude profile=sonnet space=w6 tab=w6:tR
06:16:45  peer pane=w6:pR reached injectable state        <- SUCCESS, 0.5s
06:25:06  spawning claude profile=sonnet space=w6 tab=w6:tS
06:25:26  peer pane=w6:pS did not become injectable within 20000ms — closing
06:26:01  ... tab=w6:tT  -> same
06:27:04  ... tab=w6:tV  -> same
06:27:51  spawning opencode profile=terra space=w6 tab=w6:tW
06:27:53  peer pane=w6:pW reached injectable state        <- opencode still fine

opencode members kept spawning in the same workspace, at the same time, throughout. Only
claude-code members stopped.

The restart

I raised spawnReadyTimeoutMs to 120000 so a stuck pane would stay on screen, and restarted the
daemon (scripts/redeploy-fleetd.sh --yes, pid 70731, jar 4604c4958832). The new daemon created a
fresh workspace w7. Then:

05:29:32  spawning claude profile=sonnet space=w7 tab=w7:t2
05:29:33  pane w7:p2 not at its shell prompt yet, retrying agent.start
05:29:33  claude started pane=w7:p2
05:29:35  peer pane=w7:p2 reached injectable state        <- 2.1s
05:29:38  Client initialize request ... claude-code 2.1.251

And end to end, over REST:

POST /sessions/term_65a22fc8f4ea531/message
-> {"reply":"11c3ff6 OK","replySource":"reply"}

Spawn, MCP mount, send and fleet_reply all work. The raised timeout was never needed — the member
was ready in 2.1 seconds.

What is still open

The root cause. A restart clears it, so it is state held by the running daemon or by its herdr
workspace, and it accumulates over a daemon's lifetime. I have not proved which. One lead worth
following: the failing spawns were all in the long-lived workspace w6, whose herdr tab ids had
wrapped past the single-character range (tR, tS, ... tZ, t0, t1, ... t11). fleetd's
AgentControl caches a terminal→pane mapping (paneByTerminal), invalidated only on
agent_not_found. A reused pane id would send status polls to the wrong pane, which looks exactly
like "never becomes injectable". That does not yet explain why opencode members were unaffected, so
treat it as a hypothesis, not a finding.

For now

  • Workaround: restart the daemon when claude-code spawns start timing out. opencode members
    staying healthy is not evidence the daemon is fine.
  • spawnReadyTimeoutMs: 120000 is live and marked as a debug setting in fleetd.yaml. Put it back
    to the 20000 default once this is understood — at this value a genuinely bad spawn blocks its
    caller for two minutes.
  • Related and still true: the lead holds a subscription seat, so the real sonnet fan-out is 2, not
    the configured maxLoad: 3.
## 2026-08-29 — the cause is daemon-lifetime state, not the claude backend. A restart fixes it. Earlier notes on this issue said "the Claude subscription backend is failing". That was wrong. The backend is healthy. What fails is a `fleetd` process that has been running for a long time. ### What I ruled out, each by measurement Every one of these was tested on this host with claude 2.1.251, outside a herdr pane: | Suspect | Test | Result | |---|---|---| | auth / login | `claude -p "OK" --model claude-sonnet-5` | OK, rc=0 | | usage limit | same | OK | | the launcher's flags | `-p` plus `--mcp-config` + `--append-system-prompt` + `--agent dev` + `--autocompact 250000` + `--session-id` | OK | | the `memberCredentials` scrub | reproduced the CB-633 allow-list in a zsh (105 of 125 variables blanked), then started claude | TUI renders | | interactive mode / pty | real pty via `pty.fork()`, 40x120 winsize, full member flag set | TUI renders in <4s | | scrub **and** pty together | both of the above in one run | TUI renders | | a subscription seat limit | three concurrent interactive claude sessions | all three render | So nothing in the argv, the environment, the terminal, or the account explains it. ### What the log actually shows Sonnet spawns broke *mid-session* on 2026-08-29. The daemon had been up since 2026-08-28 06:36. ``` 06:16:42 spawning claude profile=sonnet space=w6 tab=w6:tR 06:16:45 peer pane=w6:pR reached injectable state <- SUCCESS, 0.5s 06:25:06 spawning claude profile=sonnet space=w6 tab=w6:tS 06:25:26 peer pane=w6:pS did not become injectable within 20000ms — closing 06:26:01 ... tab=w6:tT -> same 06:27:04 ... tab=w6:tV -> same 06:27:51 spawning opencode profile=terra space=w6 tab=w6:tW 06:27:53 peer pane=w6:pW reached injectable state <- opencode still fine ``` opencode members kept spawning in the same workspace, at the same time, throughout. Only `claude-code` members stopped. ### The restart I raised `spawnReadyTimeoutMs` to 120000 so a stuck pane would stay on screen, and restarted the daemon (`scripts/redeploy-fleetd.sh --yes`, pid 70731, jar `4604c4958832`). The new daemon created a fresh workspace `w7`. Then: ``` 05:29:32 spawning claude profile=sonnet space=w7 tab=w7:t2 05:29:33 pane w7:p2 not at its shell prompt yet, retrying agent.start 05:29:33 claude started pane=w7:p2 05:29:35 peer pane=w7:p2 reached injectable state <- 2.1s 05:29:38 Client initialize request ... claude-code 2.1.251 ``` And end to end, over REST: ``` POST /sessions/term_65a22fc8f4ea531/message -> {"reply":"11c3ff6 OK","replySource":"reply"} ``` Spawn, MCP mount, send and `fleet_reply` all work. The raised timeout was never needed — the member was ready in 2.1 seconds. ### What is still open The **root cause**. A restart clears it, so it is state held by the running daemon or by its herdr workspace, and it accumulates over a daemon's lifetime. I have not proved which. One lead worth following: the failing spawns were all in the long-lived workspace `w6`, whose herdr tab ids had wrapped past the single-character range (`tR`, `tS`, ... `tZ`, `t0`, `t1`, ... `t11`). fleetd's `AgentControl` caches a terminal→pane mapping (`paneByTerminal`), invalidated only on `agent_not_found`. A reused pane id would send status polls to the wrong pane, which looks exactly like "never becomes injectable". That does not yet explain why opencode members were unaffected, so treat it as a hypothesis, not a finding. ### For now - **Workaround: restart the daemon** when claude-code spawns start timing out. opencode members staying healthy is not evidence the daemon is fine. - `spawnReadyTimeoutMs: 120000` is live and marked as a debug setting in `fleetd.yaml`. Put it back to the 20000 default once this is understood — at this value a genuinely bad spawn blocks its caller for two minutes. - Related and still true: the lead holds a subscription seat, so the real sonnet fan-out is 2, not the configured `maxLoad: 3`.
Author
Owner

A structural gap that matches the symptom exactly: the spawn gate cannot resolve UNKNOWN

Read against main @ 23ada19. This is not the root cause of the mid-session break, but it is a
real defect, it explains the claude-code/opencode asymmetry, and fixing it would make this class of
failure far less likely whatever the trigger turns out to be.

The two status paths are not the same, and only one of them can think

The spawn readiness gate — HerdrPeerLauncher.waitUntilInjectableOrThrow, line ~842:

long deadline = nowMillis.getAsLong() + spawnReadyTimeoutMs;
while (nowMillis.getAsLong() < deadline) {
    if (agents.status(paneId).injectable()) { ... return; }
    sleeper.run();
}

AgentControl.status(target) is get(target).status() — the raw value herdr reports, and nothing
more. AgentStatus.injectable() is IDLE || BLOCKED || DONE, so UNKNOWN is not injectable.

The status poller — StatusPoller, line 78:

AgentStatus status = refiner.refine(target, control.status(target), control);

StatusRefiner exists precisely to resolve UNKNOWN: when herdr cannot classify a pane, it re-reads
the pane tail and classifies it. Its own javadoc (line ~82) says "Classify a Claude Code TUI pane
tail"
— it is written for this backend.

StatusRefiner.refine has exactly one caller in the whole codebase, and it is StatusPoller.
I grepped for it. The spawn gate never refines.

So the two paths behave differently on the same pane:

sees UNKNOWN outcome
spawn gate cannot resolve it spins the full timeout, closes the pane, throws PeerUnreachableException
status poller re-reads the pane and classifies it resolves to a real status

That is exactly the reported symptom: "a claude-code member that never reaches an injectable state",
closed at the timeout — while everything else about the daemon keeps working.

This also corrects something I wrote earlier in this issue

I noted that opencode members kept spawning fine and warned it was not evidence the daemon was
healthy. That was right, and here is the mechanism. I first suspected opencode skipped the gate
entirely — OpenCodeLauncher does have a production constructor that disables it
(spawnReadyTimeoutMs == 0, line ~75). That is not the one in use. Fleetd.java:185 passes
cfg.spawnReadyTimeoutMs() to OpenCodeLauncher, so both backends run the gate. I checked this
rather than assuming it.

The asymmetry is not the gate — it is what herdr can classify. StatusRefiner is Claude-Code-TUI
specific, which says that this TUI is the one herdr struggles to classify. opencode presumably gets
a definite status straight from herdr and so never needs the refinement that the gate cannot do.

What I did not establish

Why it broke mid-session on 2026-08-29 and why a restart fixed it. This gap is permanent — it was
there before 06:16:45 when spawns worked, and after 06:25:06 when they stopped. So something else
changed what herdr reports for a Claude Code pane. That is still open, and I have nothing new on it.
The only accumulating state I can find on our side is AgentControl.paneByTerminal (a
ConcurrentHashMap invalidated only on agent_not_found), and it does not explain this: the
gate polls by paneId, not by terminal, and a fresh member has a fresh terminal id, so no stale
entry can apply to it. I am ruling that hypothesis out rather than leaving it standing.

Proposed fix, independent of the root cause

Give the spawn gate the same refinement the poller has. A gate that closes a healthy pane because
herdr shrugged is a worse failure than one that waits and looks. Concretely: have
waitUntilInjectableOrThrow refine the raw status the way StatusPoller does before testing
injectable().

Two things to be careful about, and a reason this is not a five-minute change:

  1. StatusRefiner reads pane content, so the gate would go from one cheap herdr call per poll to a
    content read per poll. Refine only when the raw status is UNKNOWN, not on every tick.
  2. StatusRefiner is Claude-Code-specific. Wiring it into the shared HerdrPeerLauncher gate applies
    it to opencode panes too, and a classifier reading the wrong TUI can produce a confident wrong
    answer, which is worse than UNKNOWN. It needs to be per-adapter, or it needs to be safe on a TUI
    it does not recognise.

Housekeeping

spawnReadyTimeoutMs is back to the 20000 default as of today. It was raised to 120000 on
2026-08-29 to keep a stuck pane on screen; there has been no recurrence, and two redeploys have
cleared the condition anyway, so the long timeout was only costing a two-minute block on every
genuinely bad spawn. Raise it again the moment this comes back.

## A structural gap that matches the symptom exactly: the spawn gate cannot resolve UNKNOWN Read against `main` @ `23ada19`. This is not the root cause of the mid-session break, but it is a real defect, it explains the claude-code/opencode asymmetry, and fixing it would make this class of failure far less likely whatever the trigger turns out to be. ### The two status paths are not the same, and only one of them can think **The spawn readiness gate** — `HerdrPeerLauncher.waitUntilInjectableOrThrow`, line ~842: ```java long deadline = nowMillis.getAsLong() + spawnReadyTimeoutMs; while (nowMillis.getAsLong() < deadline) { if (agents.status(paneId).injectable()) { ... return; } sleeper.run(); } ``` `AgentControl.status(target)` is `get(target).status()` — the raw value herdr reports, and nothing more. `AgentStatus.injectable()` is `IDLE || BLOCKED || DONE`, so **UNKNOWN is not injectable**. **The status poller** — `StatusPoller`, line 78: ```java AgentStatus status = refiner.refine(target, control.status(target), control); ``` `StatusRefiner` exists precisely to resolve UNKNOWN: when herdr cannot classify a pane, it re-reads the pane tail and classifies it. Its own javadoc (line ~82) says *"Classify a Claude Code TUI pane tail"* — it is written for this backend. **`StatusRefiner.refine` has exactly one caller in the whole codebase, and it is `StatusPoller`.** I grepped for it. The spawn gate never refines. So the two paths behave differently on the same pane: | | sees UNKNOWN | outcome | |---|---|---| | spawn gate | cannot resolve it | spins the full timeout, closes the pane, throws `PeerUnreachableException` | | status poller | re-reads the pane and classifies it | resolves to a real status | That is exactly the reported symptom: *"a claude-code member that never reaches an injectable state"*, closed at the timeout — while everything else about the daemon keeps working. ### This also corrects something I wrote earlier in this issue I noted that opencode members kept spawning fine and warned it was not evidence the daemon was healthy. That was right, and here is the mechanism. I first suspected opencode skipped the gate entirely — `OpenCodeLauncher` does have a production constructor that disables it (`spawnReadyTimeoutMs == 0`, line ~75). **That is not the one in use.** `Fleetd.java:185` passes `cfg.spawnReadyTimeoutMs()` to `OpenCodeLauncher`, so both backends run the gate. I checked this rather than assuming it. The asymmetry is not the gate — it is what herdr can classify. `StatusRefiner` is Claude-Code-TUI specific, which says that this TUI is the one herdr struggles to classify. opencode presumably gets a definite status straight from herdr and so never needs the refinement that the gate cannot do. ### What I did not establish **Why it broke mid-session on 2026-08-29 and why a restart fixed it.** This gap is permanent — it was there before 06:16:45 when spawns worked, and after 06:25:06 when they stopped. So something else changed what herdr reports for a Claude Code pane. That is still open, and I have nothing new on it. The only accumulating state I can find on our side is `AgentControl.paneByTerminal` (a `ConcurrentHashMap` invalidated only on `agent_not_found`), and it does **not** explain this: the gate polls by **paneId**, not by terminal, and a fresh member has a fresh terminal id, so no stale entry can apply to it. I am ruling that hypothesis out rather than leaving it standing. ### Proposed fix, independent of the root cause Give the spawn gate the same refinement the poller has. A gate that closes a healthy pane because herdr shrugged is a worse failure than one that waits and looks. Concretely: have `waitUntilInjectableOrThrow` refine the raw status the way `StatusPoller` does before testing `injectable()`. Two things to be careful about, and a reason this is not a five-minute change: 1. `StatusRefiner` reads pane content, so the gate would go from one cheap herdr call per poll to a content read per poll. Refine only when the raw status is UNKNOWN, not on every tick. 2. `StatusRefiner` is Claude-Code-specific. Wiring it into the shared `HerdrPeerLauncher` gate applies it to opencode panes too, and a classifier reading the wrong TUI can produce a confident wrong answer, which is worse than UNKNOWN. It needs to be per-adapter, or it needs to be safe on a TUI it does not recognise. ### Housekeeping `spawnReadyTimeoutMs` is back to the `20000` default as of today. It was raised to `120000` on 2026-08-29 to keep a stuck pane on screen; there has been no recurrence, and two redeploys have cleared the condition anyway, so the long timeout was only costing a two-minute block on every genuinely bad spawn. Raise it again the moment this comes back.
Author
Owner

Hit the same shape again today, on a different profile and for a different underlying reason. Recording it because it shows the defect is wider than the subscription-seat case this ticket was filed for.

What happened

2026-09-03, ~08:57. fleet_list reported terra as maxLoad: 2, live: 0, free: 2. I spawned one member onto it. The member never started. Its pane showed one line:

The usage limit has been reached

The shared OpenAI credential was spent. free: 2 was wrong, exactly as free: 1 was wrong in the original report — a different cause, the same lie, and the same cost: a briefed unit that silently produced nothing.

It was worse than a wasted spawn, because of a second thing. The turn-done fallback then handed me back my own brief as the member's "report" (#241). So the surface said: a member ran, and here is a long, on-topic report. Nothing in it came from the member, and the actual cause — one line about the usage limit — was buried in the middle of my own echoed text. #241 is fixed on main but was not yet deployed when this happened.

Two things this changes

1. free is wrong for at least three separate reasons now, not one:

  • the lead holds a seat on the same subscription (the original report),
  • the credential is exhausted,
  • the credential is cooling off after repeated backend errors (#201/#227, now merged).

The last one is handled: a cooling credential forces free to 0. The first two are not. Whatever fix this ticket gets should treat free as "spawns that would actually succeed", not "panes not currently occupied" — otherwise the next cause gets its own ticket too.

2. The exhaustion machinery existed and was simply not switched on. Neither sol nor terra declared an exhaustedPattern, so nothing could ever classify that pane line as exhaustion, and the credential both profiles share was never quarantined. fleet_list went on advertising free slots on both.

That is not a code defect — it is the gap where a knob nobody set looks identical to a feature that does not work. Workers cannot see fleetd.yaml, so this kind of gap does not show up in any PR.

I have added the measured literal to both profiles:

    credentialId: openai-shared
    exhaustedPattern: "The usage limit has been reached"

with a comment saying it was measured from that pane on this date, and that if the wording changes this stops matching silently. exhaustedPattern is a deferred key, so it takes effect at the next daemon restart, not before.

Note for whoever takes this ticket

Do not fix it by lowering maxLoad. The original report already says why, and this second cause makes it clearer: no static number is correct, because the reasons free is wrong are dynamic and there is more than one of them.

Hit the same shape again today, on a different profile and for a different underlying reason. Recording it because it shows the defect is wider than the subscription-seat case this ticket was filed for. ## What happened 2026-09-03, ~08:57. `fleet_list` reported `terra` as `maxLoad: 2, live: 0, free: 2`. I spawned one member onto it. The member never started. Its pane showed one line: ``` The usage limit has been reached ``` The shared OpenAI credential was spent. `free: 2` was wrong, exactly as `free: 1` was wrong in the original report — a different cause, the same lie, and the same cost: a briefed unit that silently produced nothing. It was worse than a wasted spawn, because of a second thing. The turn-done fallback then handed me back **my own brief** as the member's "report" (#241). So the surface said: a member ran, and here is a long, on-topic report. Nothing in it came from the member, and the actual cause — one line about the usage limit — was buried in the middle of my own echoed text. #241 is fixed on `main` but was not yet deployed when this happened. ## Two things this changes **1. `free` is wrong for at least three separate reasons now**, not one: - the lead holds a seat on the same subscription (the original report), - the credential is exhausted, - the credential is cooling off after repeated backend errors (#201/#227, now merged). The last one is handled: a cooling credential forces `free` to `0`. The first two are not. Whatever fix this ticket gets should treat `free` as "spawns that would actually succeed", not "panes not currently occupied" — otherwise the next cause gets its own ticket too. **2. The exhaustion machinery existed and was simply not switched on.** Neither `sol` nor `terra` declared an `exhaustedPattern`, so nothing could ever classify that pane line as exhaustion, and the credential both profiles share was never quarantined. `fleet_list` went on advertising free slots on both. That is not a code defect — it is the gap where a knob nobody set looks identical to a feature that does not work. Workers cannot see `fleetd.yaml`, so this kind of gap does not show up in any PR. I have added the measured literal to both profiles: ```yaml credentialId: openai-shared exhaustedPattern: "The usage limit has been reached" ``` with a comment saying it was measured from that pane on this date, and that if the wording changes this stops matching **silently**. `exhaustedPattern` is a deferred key, so it takes effect at the next daemon restart, not before. ## Note for whoever takes this ticket Do not fix it by lowering `maxLoad`. The original report already says why, and this second cause makes it clearer: no static number is correct, because the reasons `free` is wrong are dynamic and there is more than one of them.
Author
Owner

Live evidence, 2026-09-03 ~09:35 — cause 2 is already handled. Cause 1 is not, and I watched it lie again.

I hit the same exhaustion an hour after the last comment, and this time the machinery was switched on. Recording the measurement because it answers one of the open questions on this ticket directly.

Quarantine does already force free to 0

fleet_list during the outage:

{"profile":"sol",  "maxLoad":1,"live":0,"free":0,"reclaimable":0,
 "credentialId":"openai-shared","quarantinedForSeconds":1729}
{"profile":"terra","maxLoad":2,"live":2,"free":0,"reclaimable":2,
 "credentialId":"openai-shared","quarantinedForSeconds":1729}

Both profiles on the shared credential went to free: 0 on their own. So cause 2 in my earlier comment needs no work — the exhaustedPattern I added this morning classified the pane line, the credential quarantined, and capacity followed. Whoever takes this ticket should scope to cause 1 only.

This is also the first end-to-end proof of that path. It had never fired before today, because the key was never set.

The report is honest now too

fleet_poll returned:

[failed — backend exhausted (usage limit): ... The usage limit has been reached ...]

Not my own brief echoed back. #241's echo suppression and the exhaustion classifier both did their job on their first real occurrence. That is the difference between this outage and the 08:57 one described above: same failure, and now it says so.

For completeness — I checked both worktrees rather than trusting the message: zero commits, zero dirty files, nothing modified. The members genuinely never started.

Cause 1 lied again, in the same minute

While sol and terra sat at free: 0, sonnet reported:

{"profile":"sonnet","maxLoad":3,"live":2,"free":1,"reclaimable":0}

free: 1 was wrong. Two sonnet members were live and the lead holds a seat on the same subscription, so the fleet was already at the ceiling of 3 and no third member could start. This is exactly the original report, still live, and it is now the only remaining cause.

That sharpens the fix: free needs to account for the lead's own seat on a subscription: true profile. The other two causes are both handled by the credential layer, which is a completely different mechanism — a credential is quarantined or cooling; a seat is simply occupied. Do not try to route the seat problem through the quarantine machinery.

Practical note for the fleet

The three routes that actually work are down to one when this happens. sol/terra share openai-shared; local, local-direct, opus and xf all sit at weight: 0; and sonnet is bounded by the subscription seat. When the OpenAI credential is spent, gx is the only profile with real free capacity. That is worth knowing before planning a fan-out, and it is not visible from fleet_list today — another reason free should mean "spawns that would succeed".

## Live evidence, 2026-09-03 ~09:35 — cause 2 is already handled. Cause 1 is not, and I watched it lie again. I hit the same exhaustion an hour after the last comment, and this time the machinery was switched on. Recording the measurement because it answers one of the open questions on this ticket directly. ### Quarantine **does** already force `free` to 0 `fleet_list` during the outage: ```json {"profile":"sol", "maxLoad":1,"live":0,"free":0,"reclaimable":0, "credentialId":"openai-shared","quarantinedForSeconds":1729} {"profile":"terra","maxLoad":2,"live":2,"free":0,"reclaimable":2, "credentialId":"openai-shared","quarantinedForSeconds":1729} ``` Both profiles on the shared credential went to `free: 0` on their own. So **cause 2 in my earlier comment needs no work** — the `exhaustedPattern` I added this morning classified the pane line, the credential quarantined, and capacity followed. Whoever takes this ticket should scope to cause 1 only. This is also the first end-to-end proof of that path. It had never fired before today, because the key was never set. ### The report is honest now too `fleet_poll` returned: ``` [failed — backend exhausted (usage limit): ... The usage limit has been reached ...] ``` Not my own brief echoed back. #241's echo suppression and the exhaustion classifier both did their job on their first real occurrence. That is the difference between this outage and the 08:57 one described above: same failure, and now it says so. For completeness — I checked both worktrees rather than trusting the message: zero commits, zero dirty files, nothing modified. The members genuinely never started. ### Cause 1 lied again, in the same minute While `sol` and `terra` sat at `free: 0`, `sonnet` reported: ```json {"profile":"sonnet","maxLoad":3,"live":2,"free":1,"reclaimable":0} ``` `free: 1` was **wrong**. Two sonnet members were live and the lead holds a seat on the same subscription, so the fleet was already at the ceiling of 3 and no third member could start. This is exactly the original report, still live, and it is now the *only* remaining cause. That sharpens the fix: `free` needs to account for the lead's own seat on a `subscription: true` profile. The other two causes are both handled by the credential layer, which is a completely different mechanism — a credential is quarantined or cooling; a seat is simply occupied. Do not try to route the seat problem through the quarantine machinery. ### Practical note for the fleet The three routes that actually work are down to one when this happens. `sol`/`terra` share `openai-shared`; `local`, `local-direct`, `opus` and `xf` all sit at `weight: 0`; and `sonnet` is bounded by the subscription seat. When the OpenAI credential is spent, `gx` is the only profile with real free capacity. That is worth knowing before planning a fan-out, and it is not visible from `fleet_list` today — another reason `free` should mean "spawns that would succeed".
Author
Owner

A third mechanism, measured 20 minutes after the last comment — and this one the spawn gate cannot catch

gx reported maxLoad: 2, live: 0, free: 2. I spawned two members onto it. Both spawns succeeded — panes created, MCP mounted, fleet_status returning working.

Twenty minutes later:

3a0786-6: dirty=0 commits=0 files written in last 5 min=0
9d4c58-7: dirty=0 commits=0 files written in last 5 min=0

Nothing. In the same fleet_list:

{"profile":"gx","maxLoad":2,"live":2,"free":0,"reclaimable":0,
 "credentialId":"gx","quarantinedForSeconds":1325}

The gx credential had thrown repeated backend errors and quarantined after the spawns went through. So both members were alive, healthy by every signal the daemon exposes, and talking to a backend that never answered. I stopped them and lost two units of work.

Why this one is different, and why it matters for the fix

The two causes already on this ticket are both spawn-time: the process exits immediately (subscription seat) or never starts (spent credential). The spawn gate can see those, and #0f08b93 improved how it reports them.

This one is post-spawn. The gate did its job correctly — the pane really did reach an injectable state. The backend died afterwards. No amount of work on waitUntilInjectableOrThrow would have caught it, and nothing in fleet_list distinguishes these two members from two that are genuinely thinking hard.

So free is now wrong for four distinct reasons, and they split into two families:

cause when visible today?
1 lead holds a subscription seat spawn no — the open work on this ticket
2 credential exhausted spawn yes, quarantine zeroes free
3 credential cooling after errors spawn yes, cool-off zeroes free
4 credential dies while members are live after spawn no

Cause 4 does not make free wrong — free was already 0 because both panes were occupied. It makes live: 2 wrong, which is worse: the lead believes two units are in flight. reclaimable: 0 said they were not even reclaimable.

Concrete suggestion, on top of the existing scope

When a credential is quarantined or cooling, the members already running on that credential should be marked, not just the profile's future capacity. They are not working; they cannot work; and the operator is the only one who can tell, by hand, by checking file mtimes in a worktree. A liveStatus that keeps saying working for a member whose credential is known-dead is a claim the daemon cannot support — the same defect this ticket names, one layer in.

That is arguably a separate ticket. I am leaving it here rather than splitting it because it is the same sentence: the daemon reports state it has never verified. Whoever takes cause 1 should read this before deciding where the fix belongs.

Practical note

That leaves sonnet as the only usable profile right now — openai-shared (sol, terra) and gx are all quarantined, and local, local-direct, opus, xf sit at weight: 0. sonnet reports free: 1 and the true figure is 0, because of cause 1. So at this moment fleet_list shows 8 profiles, advertises free capacity on 5 of them, and the real number of members I can start is zero.

## A third mechanism, measured 20 minutes after the last comment — and this one the spawn gate cannot catch `gx` reported `maxLoad: 2, live: 0, free: 2`. I spawned two members onto it. **Both spawns succeeded** — panes created, MCP mounted, `fleet_status` returning `working`. Twenty minutes later: ``` 3a0786-6: dirty=0 commits=0 files written in last 5 min=0 9d4c58-7: dirty=0 commits=0 files written in last 5 min=0 ``` Nothing. In the same `fleet_list`: ```json {"profile":"gx","maxLoad":2,"live":2,"free":0,"reclaimable":0, "credentialId":"gx","quarantinedForSeconds":1325} ``` The `gx` credential had thrown repeated backend errors and quarantined **after** the spawns went through. So both members were alive, healthy by every signal the daemon exposes, and talking to a backend that never answered. I stopped them and lost two units of work. ### Why this one is different, and why it matters for the fix The two causes already on this ticket are both **spawn-time**: the process exits immediately (subscription seat) or never starts (spent credential). The spawn gate can see those, and #0f08b93 improved how it reports them. This one is **post-spawn**. The gate did its job correctly — the pane really did reach an injectable state. The backend died afterwards. No amount of work on `waitUntilInjectableOrThrow` would have caught it, and nothing in `fleet_list` distinguishes these two members from two that are genuinely thinking hard. So `free` is now wrong for four distinct reasons, and they split into two families: | | cause | when | visible today? | |---|---|---|---| | 1 | lead holds a subscription seat | spawn | no — the open work on this ticket | | 2 | credential exhausted | spawn | yes, quarantine zeroes `free` | | 3 | credential cooling after errors | spawn | yes, cool-off zeroes `free` | | 4 | **credential dies while members are live** | **after spawn** | **no** | Cause 4 does not make `free` wrong — `free` was already 0 because both panes were occupied. It makes `live: 2` wrong, which is worse: the lead believes two units are in flight. `reclaimable: 0` said they were not even reclaimable. ### Concrete suggestion, on top of the existing scope When a credential is quarantined or cooling, the members **already running** on that credential should be marked, not just the profile's future capacity. They are not working; they cannot work; and the operator is the only one who can tell, by hand, by checking file mtimes in a worktree. A `liveStatus` that keeps saying `working` for a member whose credential is known-dead is a claim the daemon cannot support — the same defect this ticket names, one layer in. That is arguably a separate ticket. I am leaving it here rather than splitting it because it is the same sentence: **the daemon reports state it has never verified.** Whoever takes cause 1 should read this before deciding where the fix belongs. ### Practical note That leaves `sonnet` as the only usable profile right now — `openai-shared` (sol, terra) and `gx` are all quarantined, and `local`, `local-direct`, `opus`, `xf` sit at `weight: 0`. `sonnet` reports `free: 1` and the true figure is 0, because of cause 1. So at this moment `fleet_list` shows 8 profiles, advertises free capacity on 5 of them, and the real number of members I can start is **zero**.
Author
Owner

Stage 2 merged to main as 0e8bfb7. 1250 tests, 0 failures, 0 compile errors on the merged tree.

Why stage 1 needed a stage 2

Stage 1 was correct code that did nothing on this host, and every test passed. effectiveCredentialId() fell back to the profile's own name when credentialId was unset. The lead runs on opus, members on sonnet; both are subscription: true with no credentialId, so the matcher compared "opus" against "sonnet", never matched, and charged 0 seats.

The tell was in stage 1's own fixture: every test put the lead on the same profile name as the target. The live config is the one shape that could not work.

Stage 2 returns a "<subscription>" sentinel when credentialId is unset and subscription is true. An explicit credentialId still wins, so two genuinely separate Claude logins on one host stay apart.

Verified on the live shape, by me, not by reasoning

opus.effectiveCredentialId()   = <subscription>
sonnet.effectiveCredentialId() = <subscription>
seats charged to sonnet = 1      (was 0 before this change)

Only opus and sonnet join the sentinel group here. local, local-direct, gx, xf, sol, terra are unaffected — checked against the live fleetd.yaml, which no worker can see.

The consequence nobody asked about — please read this

The subtraction is reporting-only. LeadSeatSource is wired into fleet_list's row and nowhere else. The spawn gate, CompositePeerLauncher, uses maxLoad and the live count and never consults the lead seat.

So on this host, with sonnet.maxLoad: 3 and 2 members live:

  • fleet_list now reports free: 0
  • a third fleet_spawn{profile: "sonnet"} still succeeds

The two numbers now disagree by one, and free understates what the daemon will actually do. That is the opposite of the overstatement this ticket was filed about.

It also sits against a tested finding recorded in the live config: on 2026-08-29 three concurrent interactive claude sessions all started fine and claude -p with the full member flag set returned rc=0, so no backend seat limit was ever measured here. The config comment ends "Left at 3 because 2 was never shown to be the ceiling."

I have not changed maxLoad, and I have written the above into the live config next to that comment, including: do not raise maxLoad to 4 to win the slot back — the slot was never taken, and you would be granting a fourth real member.

Open question for whoever owns this: should free mean "slots the spawn gate will grant" or "sessions this account can carry"? Today it means the second and the gate means the first. Both are defensible; having them differ silently is not. Worth a follow-up ticket rather than a quiet change here.

Second bug this also fixes

CompositePeerLauncher.credentialIdFor feeds enforceNotQuarantined and enforceNotCoolingOff. Before this, quarantining opus did not block fleet_spawn{profile:"sonnet"} even though they are one login. Now it does. Confirmed by reading the call sites, not taken on the worker's word — that is correct behaviour, since one subscription hitting a usage limit really does take out every profile on it.

All five logical callers of effectiveCredentialId() were checked; every one wants "this account", none wants "this exact profile".

Merge notes

The branch was five commits behind main. It auto-merged with zero conflicts — which proves nothing, so the merged tree was built before it landed. free is clamped with Math.max(0, …), so the subtraction cannot report a negative.

Stage 2 merged to `main` as `0e8bfb7`. 1250 tests, 0 failures, 0 compile errors on the merged tree. ## Why stage 1 needed a stage 2 Stage 1 was correct code that did nothing on this host, and every test passed. `effectiveCredentialId()` fell back to the profile's own **name** when `credentialId` was unset. The lead runs on `opus`, members on `sonnet`; both are `subscription: true` with no `credentialId`, so the matcher compared `"opus"` against `"sonnet"`, never matched, and charged **0 seats**. The tell was in stage 1's own fixture: every test put the lead on the **same profile name** as the target. The live config is the one shape that could not work. Stage 2 returns a `"<subscription>"` sentinel when `credentialId` is unset and `subscription` is true. An explicit `credentialId` still wins, so two genuinely separate Claude logins on one host stay apart. ## Verified on the live shape, by me, not by reasoning ``` opus.effectiveCredentialId() = <subscription> sonnet.effectiveCredentialId() = <subscription> seats charged to sonnet = 1 (was 0 before this change) ``` Only `opus` and `sonnet` join the sentinel group here. `local`, `local-direct`, `gx`, `xf`, `sol`, `terra` are unaffected — checked against the live `fleetd.yaml`, which no worker can see. ## The consequence nobody asked about — please read this **The subtraction is reporting-only.** `LeadSeatSource` is wired into `fleet_list`'s row and nowhere else. The spawn gate, `CompositePeerLauncher`, uses `maxLoad` and the live count and never consults the lead seat. So on this host, with `sonnet.maxLoad: 3` and 2 members live: - `fleet_list` now reports `free: 0` - a third `fleet_spawn{profile: "sonnet"}` still **succeeds** The two numbers now disagree by one, and `free` **understates** what the daemon will actually do. That is the opposite of the overstatement this ticket was filed about. It also sits against a tested finding recorded in the live config: on 2026-08-29 three concurrent interactive `claude` sessions all started fine and `claude -p` with the full member flag set returned rc=0, so no backend seat limit was ever measured here. The config comment ends "Left at 3 because 2 was never shown to be the ceiling." I have not changed `maxLoad`, and I have written the above into the live config next to that comment, including: do **not** raise `maxLoad` to 4 to win the slot back — the slot was never taken, and you would be granting a fourth real member. **Open question for whoever owns this:** should `free` mean "slots the spawn gate will grant" or "sessions this account can carry"? Today it means the second and the gate means the first. Both are defensible; having them differ silently is not. Worth a follow-up ticket rather than a quiet change here. ## Second bug this also fixes `CompositePeerLauncher.credentialIdFor` feeds `enforceNotQuarantined` and `enforceNotCoolingOff`. Before this, quarantining `opus` did **not** block `fleet_spawn{profile:"sonnet"}` even though they are one login. Now it does. Confirmed by reading the call sites, not taken on the worker's word — that is correct behaviour, since one subscription hitting a usage limit really does take out every profile on it. All five logical callers of `effectiveCredentialId()` were checked; every one wants "this account", none wants "this exact profile". ## Merge notes The branch was five commits behind `main`. It auto-merged with zero conflicts — which proves nothing, so the merged tree was built before it landed. `free` is clamped with `Math.max(0, …)`, so the subtraction cannot report a negative.
ltms closed this issue 2026-09-03 08:26:20 +02:00
Sign in to join this conversation.
1 Participants
Notifications
Due Date
No due date set.
Dependencies

No dependencies set.

Reference: fleet/fleetd#176