Features: pane line limit, spawn-failure pane tail, claude-code resume, raw-scrape exhaustion, member-user scrub (#220 #214 #211 #213)

Dai Ha
2026-09-01 14:15:01 +07:00
parent dc565af709
commit ce35e4887c
+141
@@ -73,6 +73,11 @@ six weeks, and the table alone will not carry it.
| [REST roster rows under `members`](#get-members-reports-its-rows-under-members) | automatic | CB-199 | `rest/FleetApp` |
| [Say "unknown" when the member env is unreadable](#the-credential-gap-detector-admits-when-it-cannot-see) | `memberHerdrSocket:` | CB-185 | `member/HerdrPeerLauncher` |
| [One reply settles one ticket](#one-worker-reply-settles-exactly-one-ticket) | automatic | CB-137 | `msg/MessageService` |
| [A launch command that cannot fit is refused](#a-launch-command-that-cannot-fit-is-refused) | automatic | #220 | `member/HerdrPeerLauncher` |
| [A failed spawn shows you the pane](#a-failed-spawn-shows-you-the-pane) | automatic | #220 | `member/HerdrPeerLauncher` |
| [Every claude-code member is resumable](#every-claude-code-member-is-resumable) | automatic | #214 | `member/ClaudeCodeLauncher` |
| [An exhausted backend is quarantined even from a chrome-only pane](#an-exhausted-backend-is-quarantined-even-from-a-chrome-only-pane) | automatic | #211 | `inject/CompletionResolver` |
| [The credential scrub follows the member's own user](#the-credential-scrub-follows-the-members-own-user) | `memberLoginShell:` + `worktreeGroup:` | #213 | `member/HerdrPeerLauncher` |
Nearly every knob above lives in one file, on one profile:
@@ -2636,6 +2641,142 @@ herdr protocol 19 (herdr 0.8.0), while `HerdrClient`'s still says protocol 14 (0
compares the protocol number herdr reports against what fleetd actually needs — which is precisely
how `/healthz` once went green while every spawn failed.
---
## A launch command that cannot fit is refused
**What.** Before starting a member, fleetd measures the command it is about to hand herdr. If it
cannot fit the pane's line, the spawn is refused straight away with the byte count and the longest
argument named. Automatic — no knob.
**On.** Nothing to turn on.
**Why it exists.** herdr does not `exec` a member's launch command. It **types** it into the pane,
and a pty line buffer holds 1024 bytes (BSD/macOS `MAX_CANON`). Everything past that byte is dropped,
and no layer says a word: herdr answers "agent started", the backend exits on the mangled argument it
was handed, the pane closes, and the only symptom is the readiness gate timing out twenty seconds
later with no reason at all.
That is not a theory. It took the whole claude-code half of the fleet down. The reply charter used to
ride inline on `--append-system-prompt`, so a sonnet launch command was already 978 bytes. Adding
`--session-id <uuid>` for #214 — fifty bytes — made it 1028, and the four bytes cut off the end
turned `--autocompact 250000` into `--autocompact 25`. Claude Code refuses that value, so every
claude-code spawn died. The cut is at byte 1024 exactly; that was measured on a live pane, not
assumed.
Two things changed. The charter now always travels as `--append-system-prompt-file`, which takes
~800 bytes of prose off the command line for good. And this guard catches whatever grows next.
**The gotcha: the estimate is deliberately too big.** fleetd cannot see how herdr quotes each
argument, so every argument is charged its own bytes plus a separator and a quote pair. An
over-estimate costs you a clear error at a length that was already unsafe. An under-estimate would
let the silent truncation back in, and that is the failure this exists to prevent.
**It already caught a second one.** A profile with `ideMcpUrl:` set assembles 1084 bytes — over the
limit before this ticket was written. The IDE mount was one config key away from the same silent
death. It cannot ship that way now.
---
## A failed spawn shows you the pane
**What.** When the spawn-readiness gate gives up on a member, the WARN carries the last status herdr
reported **and the tail of the pane**, read before the pane is closed. Automatic — no knob.
**On.** Nothing to turn on.
**Why it exists.** The gate used to log only that it had timed out, which is true of every possible
cause: a backend that never launched, a binary that rejected an argument and exited, a trust prompt
waiting for a keypress, a login shell that hung. The pane holds the one copy of that answer, and the
next line of code destroyed it.
Finding #220 without this took a live process sampler, a hand-built pty and a byte count. With it,
the answer was one log line: `... --model claude-sonnet-5 --autocompact 25`.
**The gotcha: read it before `stop()`.** The order matters and is easy to get backwards. `stop(paneId)`
closes the pane, and a pane read after that returns nothing — the log would be just as empty as
before, only slower.
---
## Every claude-code member is resumable
**What.** Every claude-code spawn gets a session id, so `fleet_list` always reports an
`agentSessionId` you can hand back to `fleet_spawn{resumeSessionId}`. Automatic — no knob.
**On.** Nothing to turn on.
**Why it exists.** A session id used to be minted only when the caller passed `sessionName` or
`resumeSessionId`, and `fleet_spawn` treats both as optional. So an ordinary spawn minted nothing and
`agentSessionId` stayed empty for that member's whole life. Unlike opencode, nothing resolves it
afterwards: a claude-code session id is fixed at launch.
That made resume opt-in **at spawn**, while the moment you want it is *after* a member has done work
worth keeping. By then it was too late, and an empty field in the roster read as "this backend does
not support resume", which was not what was happening.
**The gotcha: the binary takes a UUID and nothing else.** `claude --session-id` rejects any other
value at argument parsing — *"Invalid session ID. Must be a valid UUID."* — so the random UUID is
required, not incidental. The one documented flag interaction is with `-r`, which claims the same id;
a resume spawn passes `-r` and never mints.
---
## An exhausted backend is quarantined even from a chrome-only pane
**What.** When a member's pane yields no usable assistant block, fleetd still classifies
`BACKEND_EXHAUSTED` and `BACKEND_ERROR` from the raw screen — so an exhausted credential is
quarantined instead of being recorded as a member that did nothing. Automatic — no knob.
**On.** Nothing to turn on.
**Why it exists.** The completion resolver checked for an empty scrape **before** it looked for an
exhaustion line, and returned. `lastAssistantBlock` finds nothing whenever the pane carries no `⏺`
marker and its first visible line is TUI chrome — a box border, a warning, the input box — because
the boundary scan breaks on that first line.
The exhaustion branch is the only caller of the quarantine sink. So a backend that refused the turn
for a usage limit was filed as "produced nothing", the credential was never quarantined, and fleetd
kept handing work to a backend that could not run it. Each attempt looked like another member
silently doing nothing. That is the shape of a usage limit taking out several members with the roster
giving no reason.
**The gotcha: it is a fallback, not a reordering.** The empty-scrape failure was not moved or
weakened, and a pane that yields a usable block never reaches this path. Matching the raw screen
widens what the patterns can hit, including scrollback that is not this turn's output — and a false
positive here quarantines a **working** credential. That is why the raw match runs only where the
alternative was a generic failure with no information in it at all.
---
## The credential scrub follows the member's own user
**What.** When `memberHerdrSocket:` puts member panes under a different OS user, the allow-list
credential scrub uses that user's facts: `memberLoginShell:` decides whether the ZDOTDIR scrub can
run, and the generated directory goes under `worktreeRoot`, shared read-only with `worktreeGroup`.
**On.** `memberLoginShell:` and `worktreeGroup:` — both required once `memberHerdrSocket:` is set.
With `memberHerdrSocket:` absent, nothing changes: fleetd's own `$SHELL` still decides and the
directory still goes to `java.io.tmpdir`, byte-identical to before.
**Why it exists.** The scrub read fleetd's own `$SHELL` to decide whether a member's login shell
honours `ZDOTDIR`, and wrote the generated directory into fleetd's own `java.io.tmpdir`. Both
describe the wrong process once panes run as another user.
The dangerous combination was fleetd on zsh and the member user on anything else. fleetd took the
zsh branch, generated a ZDOTDIR, and set it on the pane; the member's shell ignored `ZDOTDIR`
entirely, so no scrub ran — and because the zsh branch deliberately skips the sentinel overlay, the
fallback never happened either. **No protection at all, on the path fleetd believed was the protected
one.** Even with the member on zsh it still failed: `java.io.tmpdir` is mode 0700 on macOS, so
another uid cannot even traverse it.
**The gotcha: it degrades, it never refuses to spawn.** If the member shell is not zsh, or
`worktreeRoot`/`worktreeGroup` is missing, fleetd warns loudly and falls back to the sentinel
overlay. A weakened credential control must not become an outage for an opt-in feature. The
permissions are deliberately tight in the other direction: the directory is `rwxr-x---` and each file
`rw-r-----` — group read, **no group write anywhere, no world bits**. A member sources what it needs
and cannot alter fleetd's own scrub.
### A note for anyone briefing a worker to read this page
**A worker cannot see the current version of this file.** `wiki/` is a submodule, and the parent