capacity counts panes, not subscription seats, so a subscription profile reports a free slot that cannot be filled #176
Closed
opened 2026-08-28 00:52:34 +02:00 by ltms
·
7 comments
No Branch/Tag Specified
main
worker/fleetd-612-unita-87807e-1
worker/612-b3-mcpwirings-da2b58-3
worker/612-b2-cb185-176d3a-2
worker/612-b1-completion-457459-1
worker/612-agaps-73a926-2
worker/608-sleeps-3a64ff-3
worker/621-b4520b-1
worker/618-b83894-2
worker/fleetd-615-e05481-5
worker/lead-autocompact-5f1ab2-3
worker/fleetd-613-f85deb-3
worker/fleetd-608-flaky-nudge-test-d0c2d1-3
worker/lead-context-gauge-ad404f-1
worker/gauge-wiring-9158c1-4
worker/redeploy-slowstart-ead0e5-5
worker/charter-bytes-13668c-6
worker/rollover-outcome-291483-2
worker/589-f64303-2
worker/593-1a8025-5
worker/589-fcd2aa-1
worker/568-9fdaa2-3
worker/571-attempted-outcome-5739f7-2
worker/581-completionresolver-cas-sites-0542b7-6
worker/562-loop-health-wiring-test-99611c-5
worker/562-surface-loop-health-7df5cc-4
worker/575-waiter-cleanup-sites-62ad80-1
worker/572-answer-lock-release-46a9ae-5
worker/567-probe-channel-leak-a38fc5-6
worker/551-record-before-send-7cbf56-1
worker/561-listener-fanout-survives-a-throw-61d538-2
worker/555-redeploy-main-flow-seam-65c2f5-2
worker/556-injector-owns-registration-e027a5-1
worker/552-post-restart-mktemp-abort-bc2672-4
worker/553-onstatus-completion-leak-0da881-2
worker/550-shasum-linux-196132-1
worker/538-loop-dies-on-error-4a5eeb-6
worker/426-health-coverage-ef1fd4-4
worker/504-failed-reported-clean-3cfd66-3
worker/537-capturedlog-close-e4c437-2
worker/459-broken-link-targets-cadc17-5
worker/535-appender-leak-fe74c1-1
worker/512-part2-shutdown-detection-434701-9
worker/529-logger-level-sweep-2a5533-8
worker/528-drain-gate-call-site-5de83d-7
charter/forge-mcp-vs-token
worker/521-swap-guard-unpinned-28e931-5
worker/519-probe-test-harness-d25ab8-4
worker/525-logger-level-leak-1b4eb0-6
worker/518-fleetmcp-resolver-wiring-8ef96c-1
worker/512-drain-complete-line-7edd71-3
worker/517-abort-branch-and-jar-id-41b641-2
worker/500-9e52c9-3
worker/509-4912f4-2
worker/511-9a4b23-1
worker/493-479f45-2
worker/505-03f8b2-1
worker/492-followup-detect-unclear
worker/501-a31fa0-7
worker/498-451d1c-5
worker/494-1015ce-2
worker/492-209647-1
worker/489-001902-2
worker/480-relative-handover-path-906323-1
worker/480-b-handover-skill-45bf1f-5
worker/474-followup-source-pin-f54a55-17
worker/474-charter-check-on-reload-f54a55-17
worker/466-quarantine-repeatcount-report
worker/393-opencode-skill-seeding-71854b-13
worker/469-canonical-tool-names-2a472a-16
worker/466-quarantine-escalation-5ae9c1-15
worker/446-hot-exhausted-pattern-0af580-6
worker/464-charter-tool-name-guard-a85635-12
worker/463-listfleet-default-fails-open-f1c76c-11
worker/458-invariant-5-by-purpose-862f9a-10
worker/439-coordinator-row-gate-bc032a-8
worker/449-herdr-protocol-576015-4
worker/450-abstract-spawn-599e1c-5
worker/437-ack-refuses-177d91-1
worker/444-placement-window-feb56a-2
worker/440-helddurable-derived-d462d7-13
worker/425-rework-placement-resolve-c58ba1-9
worker/421-lead-peek-held-msgs-cdbad2-10
worker/435-fixed-policy-cap-fe11de-12
worker/422-gate-state-observability-9e79d6-11
worker/431-memberregistry-live-readers-cdbad2-10
worker/424-architect-slot-hot-038b41-7
worker/422-model-gate-spawn-c29f48-6
worker/425-default-profile-live-f55534-8
worker/415-coverage-wording-2cbf9c-5
worker/416-3ad1da-1
worker/418-588283-3
worker/deterministic-stamp-race-409-3cb7b6-10
worker/armed-reads-live-config-404-ed931f-9
worker/reply-peer-refusal-391-5a34bd-7
worker/models-allowlist-aa9e9b-3
worker/ttl-stamp-race-399-f1122f-8
worker/scrub-receipt-400-316b3e-5
worker/exhaustion-detection-395-105105-6
worker/scrub-abort-394-316b3e-5
fix/scrub-uid-abort
worker/task-scrub-517574-2
worker/t386-clock-bd5b78-4
worker/t384-scrub-813790-5
worker/t381-cc-748314-2
worker/t373-336973-2
worker/t365-3920c5-3
worker/t358-6e989b-1
worker/t355-8b321c-1
worker/fleetd-369-hermetic-git-tests-e8b19a-3
worker/fleetd-368-stale-lead-binding-f5682e-2
worker/fleetd-360-deploy-units-0d3793-1
worker/359-dead-lead-tabs-f1253b-4
worker/362-worktree-skills-c03e51-3
worker/361-coord-visibility-655144-1
362-plugin-visibility-and-drift
worker/errscan-bed2ca-2
worker/amqp-log-identity-bed2ca-2
worker/withdefaults-guard-561704
worker/sleepguard-82076d-1
worker/fd334-9ee1b6-5
worker/fd348-f1ab27-4
worker/fd335-a71c35-1
worker/fd342-174a17-2
worker/fd345-490d0f-3
worker/fleetd-337-5ec7d4-21
worker/fleetd-341-af5a6b-24
worker/fleetd-339-5ca0a2-23
worker/fleetd-338-83a4a1-22
worker/fleetd-333-281f46-18
worker/fleetd-329-11bdbb-16
worker/fleetd-330-2770fb-17
worker/fix-326-50506e-15
worker/fix-324-3e9bbf-14
worker/fix-323-b8287d-13
worker/fix-316b-bd0860-11
worker/fix-318-76ca36-9
worker/fix-317-486aec-8
worker/fix-315-ce47c5-6
worker/fix-307-275890-6
worker/fix-308-b4f664-7
worker/fix-309-ec3939-8
worker/fix-310-7a3974-9
worker/fix-302-52ad0e-9
worker/fix-298-ce1acb-8
worker/fix-297-66bd11-7
worker/fix-296-104622-6
worker/fix-293-bare-closetab-eb22b5-3
worker/fix-280-gone-ask-lapse-bca98e-2
worker/fix-290-reapidle-guard-coverage-9b0dd1-1
worker/fix-285-trust-seed-8f3565-10
worker/fix-284-backend-error-seat-85912c-11
worker/fix-282-chained-ask-e6d0bb-8
worker/fix-283-teardown-leaks-f40dfa-9
worker/fix-281-pin-handler-actions-4921ac-7
worker/audit-rendezvous-lifecycle-d072ae-2
worker/audit-health-placement-1a2476-6
worker/audit-teardown-exits-e207a5-3
worker/audit-launcher-asymmetry-27e370-4
worker/audit-rest-authz-6ca53c-5
worker/investigate-275-abandon-asking-fdef52-8
worker/fix-274-worktree-leak-b0095d-7
worker/fix-273-exhausted-pattern-9665b5-6
worker/fleetd-267-model-check-bd8068-1
worker/fleetd-131-archunit-18b834-7
worker/fleetd-266-sshagent-rename-a014ff-6
worker/fleetd-184-uid-claim-8e1f31-4
worker/fleetd-184-warn-b381ee-10
worker/fleetd-184-docs-be1d12-9
worker/fleetd-257-9bf010-7
worker/fleetd-103-23a113-6
worker/fleetd-247-342356-5
worker/fleetd-116-04dea8-4
worker/fleetd-252-a830e0-3
worker/fleetd-111-7e8673-9
worker/fleetd-155c-f8ef4b-8
worker/fleetd-176-b928ca-3
worker/fleetd-249-7a7878-2
worker/cb248-composition-root-b-9acdf7-15
worker/cb148-envrc-default-fa6c82-12
worker/cb201-unit5-wiring-6c12e6-8
worker/cb241-fallback-echo-1175e9-11
worker/cb149-trust-dialog-2392a5-9
worker/cb134-148-overlay-visible-c9b986-10
worker/cb234-session-id-keyed-04e1fc-1
worker/cb201-unit3-nudge-abdf5c-6
worker/cb201-unit2-policy-c1102c-5
worker/cb201-unit4-outcome-a13bfa-7
worker/cb201-unit1-classifier-91b9b1-4
worker/cb201-227-refine-831980-3
worker/cb175-model-readback-0f085f-1
worker/cb222-charter-tmpdir-17f013-1
worker/cb226-architect-slot-race-cd3aa8-3
worker/cb224-worktree-root-group-024523-2
worker/cb-123-role-demotion-c600f7-2
worker/cb-219-opencode-roots-1f677e-1
worker/cb214-claude-session-id-b9eab4-4
worker/cb213-zdotdir-wrong-process-dd6de4-3
worker/cb211-exhaustion-classification-9546e0-2
worker/cb137-ambiguous-task-4df3d8-4
worker/cb209-agentsessionid-4dfdb6-2
worker/cb185-hostenvnames-2692b5-3
worker/cb206-opencode-sqlite-128718-2
worker/cb185-worktree-group-fc0c99-1
worker/cb-137-ask-ticket-e7760c-2
worker/cb-172-broker-uri-d36ae4-4
worker/cb-175-model-readback-76ead6-3
worker/cb-161-pane-ancestry-293510-1
worker/cb-164-rebase-885863-8
worker/cb-164-empty-scrape-false-success-1a80af-3
fix/cb-197-ticket-ttl-from-completion
worker/cb-189-remote-url-coverage-4692f3-1
worker/cb-185-blockers-027756-4
worker/cb-192-gap-log-11b631-2
worker/cb-633-fix-5f4396-3
worker/cb185-router-d6436d-3
worker/cb185-router-routing-gaps-9e9d33-3
worker/cb185-paneids-992586-2
worker/cb-633-allow-list-union-ed374b-1
worker/cb-157-credential-in-remote-url-496e44-2
worker/cb-641-health-herdr-evidence-8f1f54-6
worker/cb-640-health-msg-evidence-99c9cd-1
worker/cb-642-fleets-status-skill-bbbc40-5
cb-634-ide-mcp
worker/lead-comms-wiring-c014b9-7
worker/lead-mailbox-c19577-6
worker/autocompact-window-82bc2f-5
worker/cb-634-probe-18056f-4
worker/cb635-broker-urienv
worker/cb-632-config-retry-8e0efa-7
lead/cb-622e-claude-md
lead/cb-622-followup
worker/cb-622a-165dff-1
lead/cb-622d-opencode-mount
worker/cb-622b-717c67-2
worker/cb-622c-ab7759-3
worker/cb-617b2-20ca4b-3
worker/cb-617a-5c2f4a-1
worker/cb596-4e49ef-3
worker/cb586-10500c-1
worker/cb-606-b9343a-25
worker/cb604-1445f8-24
worker/cb582-477374-21
worker/cb584-8c2281-22
worker/cb600-e6b9a9-20
worker/cb602-ce257f-19
worker/cb601-b42837-18
worker/cb598-6c7ba7-17
worker/cb599-740fe4-16
worker/cb597-282224-15
worker/cb590fix-185e9a-10
worker/cb528-recovery-race
worker/cb594-96bead-8
worker/cb590-916766-2
worker/cb527-997d99-3
worker/cb592-env-leak-3cbf9c-1
worker/cb588-async-ticket-nudge-3218f7-5
worker/cb578b-9dcb13-6
worker/cb581-d24826-5
worker/m2-u5-ef8c42-15
worker/cb578a-516499-2
worker/cb576-01a04b-17
worker/cb579-lead-tab-acba06-20
worker/cb580-terminal-health-ed6058-21
worker/cb577-f36fdc-18
worker/cb573b-3db06f-16
worker/cb568c-f36fdc-18
worker/cb568-drop-cause-c3ac1c
worker/cb575-cancelled-notification-c3ac1c
worker/m4-sol-a2cbec-3
worker/cb574-async-ask-c3ac1c
worker/cb573-health-model-8ca857-14
worker/cb572-unknown-target-7f2e35-13
worker/u4-700706-9
worker/u3-b9fcb6-6
worker/u2-ef5b68-4
worker/u1-469dce-1-clean
worker/u1-469dce-1
worker/cb-564-health-events-70cf7e-2
worker/cb-565-recycle-drops-role-98e58f-3
worker/cb-563-missing-reply-df2866-1
worker/cb-562-readiness-gate-silent-6c23c9-3
worker/cb-560-architect-presence-da8155-1
worker/cb-561-architect-silent-off-a71cab-2
worker/cb-548-bind-architect-slot-fe1b8c-1
worker/parity-overlay-settings-5fb711-1
secrets-central-store
cb-559-hot-key-correction
cb-557-fleet-role-pools
worker/cb-553-maxload-explicit-spawn-305ee3-6
worker/cb-551-idle-lead-heartbeat-f1633c-1
worker/cb-544-drain-preserves-worktree-925fad-3
worker/cb-552-docs-sync-1cb9cf-4
worker/cb-548-rendezvous-guard-rebased
worker/cb-548-rendezvous-guard-116b53-10
worker/cb-548-authz-v2-586df6-8
worker/cb-548-authz-264363-5
salvage/cb-528b-codex-home
salvage/cb-528a-codex-launcher
CB-518-primary-flow
feature/peer-launcher-spi
cb-103-injector
v1.1.0
v1.0.0
Labels
Clear labels
blocked
needs-live-proof
ready-to-delegate
silent-default
Cannot start until something else lands. The body says what.
Merged and green, but never shown working on the running daemon. Not the same as done.
Scope, files and acceptance criteria are written. A worker can be briefed from the body alone.
A feature that compiles, passes tests, and ships turned off. Nine recurrences and counting.
No Label
Milestone
No items
No Milestone
Projects
Clear projects
No project
Notifications
Due Date
No due date set.
Dependencies
No dependencies set.
Reference: fleet/fleetd#176
Reference in New Issue
Block a user
Blocking a user prevents them from interacting with repositories, such as opening or commenting on pull requests or issues. Learn more about blocking a user.
Delete Branch "%!s()"
Deleting a branch is permanent. Although the deleted branch may continue to exist for a short time before it actually gets removed, it CANNOT be undone in most cases. Continue?
What happened
Measured on the Mac fleet, 2026-08-28. The
sonnetprofile iskind: claude-codewithsubscription: trueandmaxLoad: 3.With two sonnet members up and the fleet otherwise idle,
fleet_listreported:A third
fleet_spawn{profile: sonnet}failed:Retried once the fleet was idle. It failed again, on a new pane. Both times the daemon log shows the tab was already gone before the readiness gate tried to close it:
tab_not_foundmeans theclaudeprocess exited, rather than being slow to start. A hang would leave the tab in place.The likely cause, and the part that is certain
Likely: the lead is itself a
claudesession on the same subscription, so it holds a seat. Two workers plus the lead reaches the subscription's concurrency limit, and the fourthclaudeexits at once.I have not confirmed that limit exists or what its value is — no subscription usage or concurrency API is available to check against, and reading the pane to see the process's own error message would mean driving herdr directly. So treat the specific cause as unconfirmed.
What is certain, and is the actual defect:
free: 1was wrong, and the lead was told to expect capacity that did not exist.Why it matters
freeis the number a lead plans against. CLAUDE.md tells the lead to spawn every delegated unit before sending any, so a wrongfreeis discovered part-way through a fan-out, after some units are already briefed. That is the worst moment to find out.It also costs 20 seconds per attempt to learn nothing useful, and it provisions and then removes a git worktree each time.
Suggested direction
subscription: trueshares its budget with the lead when the lead runs on the same subscription. Either subtract the lead's seat frommaxLoad, or make the relationship explicit in config so the operator states it once.tab_not_foundat close time is direct evidence the process died. Say so, rather than reporting a readiness timeout — the two have completely different fixes, and the current message names the wrong one.Workaround in place
sonnet.maxLoadis left at3on purpose, with a comment recording the measurement. Lowering it to2would hide the counting bug rather than fix it, and2is only correct while the lead runs on this same subscription.Related
Fresh data point, 2026-08-28, and it is worse than the original report.
fleet_listshowed a completely idle fleet — zero members,sonnetreportingmaxLoad: 3, live: 0, free: 3. A singlefleet_spawn{profile: sonnet}still failed:Same signature as before.
terraspawned fine in the same minute, twice, so herdr and the spawn path are healthy — it is the Claude subscription backend refusing to start.So this is now two distinct problems wearing one symptom:
free: 3was never true; 2 was the real ceiling.claude-codemember at all, even for the first one.free: 3on an idle fleet is simply wrong, and nothing infleet_listhints at it.Both point the same way: the daemon reports capacity it has never verified. It treats "the process was asked to start" as success and never reads back whether a seat existed. A spawn-time probe, or marking the profile quarantined after a readiness timeout the way credential quarantine already works, would stop
free: Nfrom being a claim the daemon cannot support.Worked around by moving both #185 units onto
terra.2026-08-29 — the cause is daemon-lifetime state, not the claude backend. A restart fixes it.
Earlier notes on this issue said "the Claude subscription backend is failing". That was wrong. The
backend is healthy. What fails is a
fleetdprocess that has been running for a long time.What I ruled out, each by measurement
Every one of these was tested on this host with claude 2.1.251, outside a herdr pane:
claude -p "OK" --model claude-sonnet-5-pplus--mcp-config+--append-system-prompt+--agent dev+--autocompact 250000+--session-idmemberCredentialsscrubpty.fork(), 40x120 winsize, full member flag setSo nothing in the argv, the environment, the terminal, or the account explains it.
What the log actually shows
Sonnet spawns broke mid-session on 2026-08-29. The daemon had been up since 2026-08-28 06:36.
opencode members kept spawning in the same workspace, at the same time, throughout. Only
claude-codemembers stopped.The restart
I raised
spawnReadyTimeoutMsto 120000 so a stuck pane would stay on screen, and restarted thedaemon (
scripts/redeploy-fleetd.sh --yes, pid 70731, jar4604c4958832). The new daemon created afresh workspace
w7. Then:And end to end, over REST:
Spawn, MCP mount, send and
fleet_replyall work. The raised timeout was never needed — the memberwas ready in 2.1 seconds.
What is still open
The root cause. A restart clears it, so it is state held by the running daemon or by its herdr
workspace, and it accumulates over a daemon's lifetime. I have not proved which. One lead worth
following: the failing spawns were all in the long-lived workspace
w6, whose herdr tab ids hadwrapped past the single-character range (
tR,tS, ...tZ,t0,t1, ...t11). fleetd'sAgentControlcaches a terminal→pane mapping (paneByTerminal), invalidated only onagent_not_found. A reused pane id would send status polls to the wrong pane, which looks exactlylike "never becomes injectable". That does not yet explain why opencode members were unaffected, so
treat it as a hypothesis, not a finding.
For now
staying healthy is not evidence the daemon is fine.
spawnReadyTimeoutMs: 120000is live and marked as a debug setting infleetd.yaml. Put it backto the 20000 default once this is understood — at this value a genuinely bad spawn blocks its
caller for two minutes.
the configured
maxLoad: 3.A structural gap that matches the symptom exactly: the spawn gate cannot resolve UNKNOWN
Read against
main@23ada19. This is not the root cause of the mid-session break, but it is areal defect, it explains the claude-code/opencode asymmetry, and fixing it would make this class of
failure far less likely whatever the trigger turns out to be.
The two status paths are not the same, and only one of them can think
The spawn readiness gate —
HerdrPeerLauncher.waitUntilInjectableOrThrow, line ~842:AgentControl.status(target)isget(target).status()— the raw value herdr reports, and nothingmore.
AgentStatus.injectable()isIDLE || BLOCKED || DONE, so UNKNOWN is not injectable.The status poller —
StatusPoller, line 78:StatusRefinerexists precisely to resolve UNKNOWN: when herdr cannot classify a pane, it re-readsthe pane tail and classifies it. Its own javadoc (line ~82) says "Classify a Claude Code TUI pane
tail" — it is written for this backend.
StatusRefiner.refinehas exactly one caller in the whole codebase, and it isStatusPoller.I grepped for it. The spawn gate never refines.
So the two paths behave differently on the same pane:
PeerUnreachableExceptionThat is exactly the reported symptom: "a claude-code member that never reaches an injectable state",
closed at the timeout — while everything else about the daemon keeps working.
This also corrects something I wrote earlier in this issue
I noted that opencode members kept spawning fine and warned it was not evidence the daemon was
healthy. That was right, and here is the mechanism. I first suspected opencode skipped the gate
entirely —
OpenCodeLauncherdoes have a production constructor that disables it(
spawnReadyTimeoutMs == 0, line ~75). That is not the one in use.Fleetd.java:185passescfg.spawnReadyTimeoutMs()toOpenCodeLauncher, so both backends run the gate. I checked thisrather than assuming it.
The asymmetry is not the gate — it is what herdr can classify.
StatusRefineris Claude-Code-TUIspecific, which says that this TUI is the one herdr struggles to classify. opencode presumably gets
a definite status straight from herdr and so never needs the refinement that the gate cannot do.
What I did not establish
Why it broke mid-session on 2026-08-29 and why a restart fixed it. This gap is permanent — it was
there before 06:16:45 when spawns worked, and after 06:25:06 when they stopped. So something else
changed what herdr reports for a Claude Code pane. That is still open, and I have nothing new on it.
The only accumulating state I can find on our side is
AgentControl.paneByTerminal(aConcurrentHashMapinvalidated only onagent_not_found), and it does not explain this: thegate polls by paneId, not by terminal, and a fresh member has a fresh terminal id, so no stale
entry can apply to it. I am ruling that hypothesis out rather than leaving it standing.
Proposed fix, independent of the root cause
Give the spawn gate the same refinement the poller has. A gate that closes a healthy pane because
herdr shrugged is a worse failure than one that waits and looks. Concretely: have
waitUntilInjectableOrThrowrefine the raw status the wayStatusPollerdoes before testinginjectable().Two things to be careful about, and a reason this is not a five-minute change:
StatusRefinerreads pane content, so the gate would go from one cheap herdr call per poll to acontent read per poll. Refine only when the raw status is UNKNOWN, not on every tick.
StatusRefineris Claude-Code-specific. Wiring it into the sharedHerdrPeerLaunchergate appliesit to opencode panes too, and a classifier reading the wrong TUI can produce a confident wrong
answer, which is worse than UNKNOWN. It needs to be per-adapter, or it needs to be safe on a TUI
it does not recognise.
Housekeeping
spawnReadyTimeoutMsis back to the20000default as of today. It was raised to120000on2026-08-29 to keep a stuck pane on screen; there has been no recurrence, and two redeploys have
cleared the condition anyway, so the long timeout was only costing a two-minute block on every
genuinely bad spawn. Raise it again the moment this comes back.
Hit the same shape again today, on a different profile and for a different underlying reason. Recording it because it shows the defect is wider than the subscription-seat case this ticket was filed for.
What happened
2026-09-03, ~08:57.
fleet_listreportedterraasmaxLoad: 2, live: 0, free: 2. I spawned one member onto it. The member never started. Its pane showed one line:The shared OpenAI credential was spent.
free: 2was wrong, exactly asfree: 1was wrong in the original report — a different cause, the same lie, and the same cost: a briefed unit that silently produced nothing.It was worse than a wasted spawn, because of a second thing. The turn-done fallback then handed me back my own brief as the member's "report" (#241). So the surface said: a member ran, and here is a long, on-topic report. Nothing in it came from the member, and the actual cause — one line about the usage limit — was buried in the middle of my own echoed text. #241 is fixed on
mainbut was not yet deployed when this happened.Two things this changes
1.
freeis wrong for at least three separate reasons now, not one:The last one is handled: a cooling credential forces
freeto0. The first two are not. Whatever fix this ticket gets should treatfreeas "spawns that would actually succeed", not "panes not currently occupied" — otherwise the next cause gets its own ticket too.2. The exhaustion machinery existed and was simply not switched on. Neither
solnorterradeclared anexhaustedPattern, so nothing could ever classify that pane line as exhaustion, and the credential both profiles share was never quarantined.fleet_listwent on advertising free slots on both.That is not a code defect — it is the gap where a knob nobody set looks identical to a feature that does not work. Workers cannot see
fleetd.yaml, so this kind of gap does not show up in any PR.I have added the measured literal to both profiles:
with a comment saying it was measured from that pane on this date, and that if the wording changes this stops matching silently.
exhaustedPatternis a deferred key, so it takes effect at the next daemon restart, not before.Note for whoever takes this ticket
Do not fix it by lowering
maxLoad. The original report already says why, and this second cause makes it clearer: no static number is correct, because the reasonsfreeis wrong are dynamic and there is more than one of them.Live evidence, 2026-09-03 ~09:35 — cause 2 is already handled. Cause 1 is not, and I watched it lie again.
I hit the same exhaustion an hour after the last comment, and this time the machinery was switched on. Recording the measurement because it answers one of the open questions on this ticket directly.
Quarantine does already force
freeto 0fleet_listduring the outage:Both profiles on the shared credential went to
free: 0on their own. So cause 2 in my earlier comment needs no work — theexhaustedPatternI added this morning classified the pane line, the credential quarantined, and capacity followed. Whoever takes this ticket should scope to cause 1 only.This is also the first end-to-end proof of that path. It had never fired before today, because the key was never set.
The report is honest now too
fleet_pollreturned:Not my own brief echoed back. #241's echo suppression and the exhaustion classifier both did their job on their first real occurrence. That is the difference between this outage and the 08:57 one described above: same failure, and now it says so.
For completeness — I checked both worktrees rather than trusting the message: zero commits, zero dirty files, nothing modified. The members genuinely never started.
Cause 1 lied again, in the same minute
While
solandterrasat atfree: 0,sonnetreported:free: 1was wrong. Two sonnet members were live and the lead holds a seat on the same subscription, so the fleet was already at the ceiling of 3 and no third member could start. This is exactly the original report, still live, and it is now the only remaining cause.That sharpens the fix:
freeneeds to account for the lead's own seat on asubscription: trueprofile. The other two causes are both handled by the credential layer, which is a completely different mechanism — a credential is quarantined or cooling; a seat is simply occupied. Do not try to route the seat problem through the quarantine machinery.Practical note for the fleet
The three routes that actually work are down to one when this happens.
sol/terrashareopenai-shared;local,local-direct,opusandxfall sit atweight: 0; andsonnetis bounded by the subscription seat. When the OpenAI credential is spent,gxis the only profile with real free capacity. That is worth knowing before planning a fan-out, and it is not visible fromfleet_listtoday — another reasonfreeshould mean "spawns that would succeed".A third mechanism, measured 20 minutes after the last comment — and this one the spawn gate cannot catch
gxreportedmaxLoad: 2, live: 0, free: 2. I spawned two members onto it. Both spawns succeeded — panes created, MCP mounted,fleet_statusreturningworking.Twenty minutes later:
Nothing. In the same
fleet_list:The
gxcredential had thrown repeated backend errors and quarantined after the spawns went through. So both members were alive, healthy by every signal the daemon exposes, and talking to a backend that never answered. I stopped them and lost two units of work.Why this one is different, and why it matters for the fix
The two causes already on this ticket are both spawn-time: the process exits immediately (subscription seat) or never starts (spent credential). The spawn gate can see those, and #0f08b93 improved how it reports them.
This one is post-spawn. The gate did its job correctly — the pane really did reach an injectable state. The backend died afterwards. No amount of work on
waitUntilInjectableOrThrowwould have caught it, and nothing infleet_listdistinguishes these two members from two that are genuinely thinking hard.So
freeis now wrong for four distinct reasons, and they split into two families:freefreeCause 4 does not make
freewrong —freewas already 0 because both panes were occupied. It makeslive: 2wrong, which is worse: the lead believes two units are in flight.reclaimable: 0said they were not even reclaimable.Concrete suggestion, on top of the existing scope
When a credential is quarantined or cooling, the members already running on that credential should be marked, not just the profile's future capacity. They are not working; they cannot work; and the operator is the only one who can tell, by hand, by checking file mtimes in a worktree. A
liveStatusthat keeps sayingworkingfor a member whose credential is known-dead is a claim the daemon cannot support — the same defect this ticket names, one layer in.That is arguably a separate ticket. I am leaving it here rather than splitting it because it is the same sentence: the daemon reports state it has never verified. Whoever takes cause 1 should read this before deciding where the fix belongs.
Practical note
That leaves
sonnetas the only usable profile right now —openai-shared(sol, terra) andgxare all quarantined, andlocal,local-direct,opus,xfsit atweight: 0.sonnetreportsfree: 1and the true figure is 0, because of cause 1. So at this momentfleet_listshows 8 profiles, advertises free capacity on 5 of them, and the real number of members I can start is zero.agent referenced this issue2026-09-03 07:56:00 +02:00
Stage 2 merged to
mainas0e8bfb7. 1250 tests, 0 failures, 0 compile errors on the merged tree.Why stage 1 needed a stage 2
Stage 1 was correct code that did nothing on this host, and every test passed.
effectiveCredentialId()fell back to the profile's own name whencredentialIdwas unset. The lead runs onopus, members onsonnet; both aresubscription: truewith nocredentialId, so the matcher compared"opus"against"sonnet", never matched, and charged 0 seats.The tell was in stage 1's own fixture: every test put the lead on the same profile name as the target. The live config is the one shape that could not work.
Stage 2 returns a
"<subscription>"sentinel whencredentialIdis unset andsubscriptionis true. An explicitcredentialIdstill wins, so two genuinely separate Claude logins on one host stay apart.Verified on the live shape, by me, not by reasoning
Only
opusandsonnetjoin the sentinel group here.local,local-direct,gx,xf,sol,terraare unaffected — checked against the livefleetd.yaml, which no worker can see.The consequence nobody asked about — please read this
The subtraction is reporting-only.
LeadSeatSourceis wired intofleet_list's row and nowhere else. The spawn gate,CompositePeerLauncher, usesmaxLoadand the live count and never consults the lead seat.So on this host, with
sonnet.maxLoad: 3and 2 members live:fleet_listnow reportsfree: 0fleet_spawn{profile: "sonnet"}still succeedsThe two numbers now disagree by one, and
freeunderstates what the daemon will actually do. That is the opposite of the overstatement this ticket was filed about.It also sits against a tested finding recorded in the live config: on 2026-08-29 three concurrent interactive
claudesessions all started fine andclaude -pwith the full member flag set returned rc=0, so no backend seat limit was ever measured here. The config comment ends "Left at 3 because 2 was never shown to be the ceiling."I have not changed
maxLoad, and I have written the above into the live config next to that comment, including: do not raisemaxLoadto 4 to win the slot back — the slot was never taken, and you would be granting a fourth real member.Open question for whoever owns this: should
freemean "slots the spawn gate will grant" or "sessions this account can carry"? Today it means the second and the gate means the first. Both are defensible; having them differ silently is not. Worth a follow-up ticket rather than a quiet change here.Second bug this also fixes
CompositePeerLauncher.credentialIdForfeedsenforceNotQuarantinedandenforceNotCoolingOff. Before this, quarantiningopusdid not blockfleet_spawn{profile:"sonnet"}even though they are one login. Now it does. Confirmed by reading the call sites, not taken on the worker's word — that is correct behaviour, since one subscription hitting a usage limit really does take out every profile on it.All five logical callers of
effectiveCredentialId()were checked; every one wants "this account", none wants "this exact profile".Merge notes
The branch was five commits behind
main. It auto-merged with zero conflicts — which proves nothing, so the merged tree was built before it landed.freeis clamped withMath.max(0, …), so the subtraction cannot report a negative.