CB-630: every subscription claude-code profile fails to spawn — the pane exits and the error names a symptom #140

Closed
opened 2026-08-23 05:48:09 +02:00 by ltms · 1 comment
Owner

Found 2026-08-23 while doing the CB-623 (#127) worker-PR proof.

What happens

fleet_spawn{profile:"sonnet", role:"dev", worktree:true} fails with:

spawn timed out — worker pane never reached injectable state:
worker pane w3:p5K did not reach injectable state within 20000ms

The daemon log shows the launch succeeding and then the pane disappearing on its own:

05:33:47.178 INFO  HerdrPeerLauncher - claude started pane=w3:p5K tab=w3:t5D terminal=term_659ae89760df654
05:33:47.178 INFO  HerdrPeerLauncher - spawned role=dev profile=sonnet charterSource=none charterBytes=799
05:34:07.211 WARN  HerdrPeerLauncher - peer pane=w3:p5K did not become injectable within 20000ms — closing
05:34:07.242 DEBUG WorkspaceControl - tab.close(w3:t5D) ignored — already gone: herdr error [tab_not_found]: tab w3:t5D not found

tab_not_found on close is the tell: the claude process had already exited by itself.

Scope of the outage

sonnet and opus are the two subscription: true kind: claude-code profiles. Both are affected. xf (kind: opencode) spawned normally in the same minute, on the same daemon, into the same repo. So the fleet currently has no working Claude-model member — only the opencode backends work.

What has already been ruled out (do not re-test these)

Each was checked by running the real binary with the real arguments, which is the CB-618 lesson:

  1. Auth / usage limit. CLAUDE_CONFIG_DIR=/Users/dai.ha/.ccs/instances/ltms claude -p "..." --model claude-sonnet-5 returns ok, exit 0.
  2. The full launcher argv. The same binary with --mcp-config '{"mcpServers":{"bridge":{"type":"http","url":"http://127.0.0.1:8765/mcp"}}}' --append-system-prompt <charter> --agent dev --session-id <uuid> --model claude-sonnet-5, run in a clean clone, also returns ok, exit 0. So no flag is being rejected and --agent dev resolves.
  3. The CB-618 agent-file cause. All three of .claude/agents/{architect,dev,reviewer}.md have name: in their frontmatter.
  4. Worktree provisioning. The log shows the worktree added and the parity overlay applied cleanly before the launch.

That leaves the pane/environment path — what the login shell does to the process, or what the launcher's env does to it — not the command line.

The reporting defect underneath it

This is the second time (after CB-618) that a spawn failure has been reported as did not reach injectable state within 20000ms, which names a symptom 20 seconds downstream of the cause. The launcher never captures what the spawned process printed before exiting. Diagnosing this costs a manual binary run every time.

Fix that as part of this ticket: on a spawn that never becomes injectable, capture the pane's contents before closing the tab, and log it. The pane is still there at the point the timeout fires — the close is what destroys the evidence.

Acceptance criteria

  • fleet_spawn{profile:"sonnet"} produces a member that reaches injectable state and answers a fleet_send.
  • Same for opus.
  • A spawn that fails to become injectable logs what the spawned process printed, not only the timeout.
  • Prove the last point by forcing a failure (for example an argv the binary rejects) and showing the captured output in the log.
Found 2026-08-23 while doing the CB-623 (#127) worker-PR proof. ## What happens `fleet_spawn{profile:"sonnet", role:"dev", worktree:true}` fails with: ``` spawn timed out — worker pane never reached injectable state: worker pane w3:p5K did not reach injectable state within 20000ms ``` The daemon log shows the launch succeeding and then the pane disappearing on its own: ``` 05:33:47.178 INFO HerdrPeerLauncher - claude started pane=w3:p5K tab=w3:t5D terminal=term_659ae89760df654 05:33:47.178 INFO HerdrPeerLauncher - spawned role=dev profile=sonnet charterSource=none charterBytes=799 05:34:07.211 WARN HerdrPeerLauncher - peer pane=w3:p5K did not become injectable within 20000ms — closing 05:34:07.242 DEBUG WorkspaceControl - tab.close(w3:t5D) ignored — already gone: herdr error [tab_not_found]: tab w3:t5D not found ``` `tab_not_found` on close is the tell: the `claude` process had already exited by itself. ## Scope of the outage `sonnet` and `opus` are the two `subscription: true` `kind: claude-code` profiles. Both are affected. `xf` (`kind: opencode`) spawned normally in the same minute, on the same daemon, into the same repo. **So the fleet currently has no working Claude-model member — only the opencode backends work.** ## What has already been ruled out (do not re-test these) Each was checked by running the real binary with the real arguments, which is the CB-618 lesson: 1. **Auth / usage limit.** `CLAUDE_CONFIG_DIR=/Users/dai.ha/.ccs/instances/ltms claude -p "..." --model claude-sonnet-5` returns `ok`, exit 0. 2. **The full launcher argv.** The same binary with `--mcp-config '{"mcpServers":{"bridge":{"type":"http","url":"http://127.0.0.1:8765/mcp"}}}' --append-system-prompt <charter> --agent dev --session-id <uuid> --model claude-sonnet-5`, run in a clean clone, also returns `ok`, exit 0. So no flag is being rejected and `--agent dev` resolves. 3. **The CB-618 agent-file cause.** All three of `.claude/agents/{architect,dev,reviewer}.md` have `name:` in their frontmatter. 4. **Worktree provisioning.** The log shows the worktree added and the parity overlay applied cleanly before the launch. That leaves the pane/environment path — what the login shell does to the process, or what the launcher's env does to it — not the command line. ## The reporting defect underneath it This is the second time (after CB-618) that a spawn failure has been reported as `did not reach injectable state within 20000ms`, which names a symptom 20 seconds downstream of the cause. **The launcher never captures what the spawned process printed before exiting.** Diagnosing this costs a manual binary run every time. Fix that as part of this ticket: on a spawn that never becomes injectable, capture the pane's contents *before* closing the tab, and log it. The pane is still there at the point the timeout fires — the close is what destroys the evidence. ## Acceptance criteria - `fleet_spawn{profile:"sonnet"}` produces a member that reaches injectable state and answers a `fleet_send`. - Same for `opus`. - A spawn that fails to become injectable logs what the spawned process printed, not only the timeout. - Prove the last point by forcing a failure (for example an argv the binary rejects) and showing the captured output in the log.
ltms added this to the 2.0 — one operation centre, many hosts milestone 2026-08-23 05:48:09 +02:00
Author
Owner

This is #220, found and fixed on 2026-09-01. Same signature, right down to the detail you called the tell: tab_not_found on close, opencode unaffected in the same minute, and the same argv running fine by hand.

The cause. herdr does not exec the launch command — it types it into the pane, and a pty line buffer holds 1024 bytes (BSD/macOS MAX_CANON). Past that the tail is dropped, silently. claude then exits on the mangled argument it was handed, the pane closes itself, and the readiness gate reports a timeout twenty seconds later. The charter alone was ~800 of those bytes, so a claude-code command sat a few bytes under the cliff and any added flag went over.

That is why your step 2 ruled out the argv correctly and still missed it: the command that fails is not the command fleetd builds — it is the first 1024 bytes of it. Running it by hand proves the binary, not the delivery.

Acceptance criteria, checked today against the fix on main (c3fa113):

  1. fleet_spawn{profile:"sonnet"} — reaches injectable, answers a fleet_send, and fleet_list shows its agentSessionId. ✅
  2. fleet_spawn{profile:"opus"} — reaches injectable and answered a fleet_send. ✅
  3. A spawn that never becomes injectable now logs the pane tail and the last herdr status, read before stop() closes the pane — exactly the change this ticket asked for. ✅
  4. Proof by forced failure: this is how #220 was found. With the over-long command in place the log read:
    peer pane=w8:p1C did not become injectable within 20000ms (last status UNKNOWN) — closing. Pane tail:
    ... --model claude-sonnet-5 --autocompact 25
    
    One line, versus the several hours the same diagnosis took without it. ✅

One thing I did not establish: why this cleared on its own between 23 August and 31 August, when claude-code members were spawning normally again. Some change shortened the command by a few bytes and moved it back under the cap. I did not reconstruct that arithmetic, and it does not matter now — the guard makes the margin explicit instead of accidental — but it is worth knowing the fleet ran a week within a handful of bytes of this.

Closing as fixed by #220.

This is #220, found and fixed on 2026-09-01. Same signature, right down to the detail you called the tell: `tab_not_found` on close, opencode unaffected in the same minute, and the same argv running fine by hand. **The cause.** herdr does not `exec` the launch command — it **types** it into the pane, and a pty line buffer holds 1024 bytes (BSD/macOS `MAX_CANON`). Past that the tail is dropped, silently. claude then exits on the mangled argument it was handed, the pane closes itself, and the readiness gate reports a timeout twenty seconds later. The charter alone was ~800 of those bytes, so a claude-code command sat a few bytes under the cliff and any added flag went over. That is why your step 2 ruled out the argv correctly and still missed it: **the command that fails is not the command fleetd builds — it is the first 1024 bytes of it.** Running it by hand proves the binary, not the delivery. **Acceptance criteria, checked today against the fix on main (`c3fa113`):** 1. `fleet_spawn{profile:"sonnet"}` — reaches injectable, answers a `fleet_send`, and `fleet_list` shows its `agentSessionId`. ✅ 2. `fleet_spawn{profile:"opus"}` — reaches injectable and answered a `fleet_send`. ✅ 3. A spawn that never becomes injectable now logs the pane tail and the last herdr status, read **before** `stop()` closes the pane — exactly the change this ticket asked for. ✅ 4. Proof by forced failure: this is how #220 was found. With the over-long command in place the log read: ``` peer pane=w8:p1C did not become injectable within 20000ms (last status UNKNOWN) — closing. Pane tail: ... --model claude-sonnet-5 --autocompact 25 ``` One line, versus the several hours the same diagnosis took without it. ✅ **One thing I did not establish:** why this cleared on its own between 23 August and 31 August, when claude-code members were spawning normally again. Some change shortened the command by a few bytes and moved it back under the cap. I did not reconstruct that arithmetic, and it does not matter now — the guard makes the margin explicit instead of accidental — but it is worth knowing the fleet ran a week within a handful of bytes of this. Closing as fixed by #220.
ltms closed this issue 2026-09-01 09:18:46 +02:00
Sign in to join this conversation.
1 Participants
Notifications
Due Date
No due date set.
Dependencies

No dependencies set.

Reference: fleet/fleetd#140