From ce35e4887c60d7d365acf1ce88abc7f93b3c31ec Mon Sep 17 00:00:00 2001 From: Dai Ha Date: Tue, 1 Sep 2026 14:15:01 +0700 Subject: [PATCH] Features: pane line limit, spawn-failure pane tail, claude-code resume, raw-scrape exhaustion, member-user scrub (#220 #214 #211 #213) --- 11-Features.md | 141 +++++++++++++++++++++++++++++++++++++++++++++++++ 1 file changed, 141 insertions(+) diff --git a/11-Features.md b/11-Features.md index ba6abf1..a692b02 100644 --- a/11-Features.md +++ b/11-Features.md @@ -73,6 +73,11 @@ six weeks, and the table alone will not carry it. | [REST roster rows under `members`](#get-members-reports-its-rows-under-members) | automatic | CB-199 | `rest/FleetApp` | | [Say "unknown" when the member env is unreadable](#the-credential-gap-detector-admits-when-it-cannot-see) | `memberHerdrSocket:` | CB-185 | `member/HerdrPeerLauncher` | | [One reply settles one ticket](#one-worker-reply-settles-exactly-one-ticket) | automatic | CB-137 | `msg/MessageService` | +| [A launch command that cannot fit is refused](#a-launch-command-that-cannot-fit-is-refused) | automatic | #220 | `member/HerdrPeerLauncher` | +| [A failed spawn shows you the pane](#a-failed-spawn-shows-you-the-pane) | automatic | #220 | `member/HerdrPeerLauncher` | +| [Every claude-code member is resumable](#every-claude-code-member-is-resumable) | automatic | #214 | `member/ClaudeCodeLauncher` | +| [An exhausted backend is quarantined even from a chrome-only pane](#an-exhausted-backend-is-quarantined-even-from-a-chrome-only-pane) | automatic | #211 | `inject/CompletionResolver` | +| [The credential scrub follows the member's own user](#the-credential-scrub-follows-the-members-own-user) | `memberLoginShell:` + `worktreeGroup:` | #213 | `member/HerdrPeerLauncher` | Nearly every knob above lives in one file, on one profile: @@ -2636,6 +2641,142 @@ herdr protocol 19 (herdr 0.8.0), while `HerdrClient`'s still says protocol 14 (0 compares the protocol number herdr reports against what fleetd actually needs — which is precisely how `/healthz` once went green while every spawn failed. +--- + +## A launch command that cannot fit is refused + +**What.** Before starting a member, fleetd measures the command it is about to hand herdr. If it +cannot fit the pane's line, the spawn is refused straight away with the byte count and the longest +argument named. Automatic — no knob. + +**On.** Nothing to turn on. + +**Why it exists.** herdr does not `exec` a member's launch command. It **types** it into the pane, +and a pty line buffer holds 1024 bytes (BSD/macOS `MAX_CANON`). Everything past that byte is dropped, +and no layer says a word: herdr answers "agent started", the backend exits on the mangled argument it +was handed, the pane closes, and the only symptom is the readiness gate timing out twenty seconds +later with no reason at all. + +That is not a theory. It took the whole claude-code half of the fleet down. The reply charter used to +ride inline on `--append-system-prompt`, so a sonnet launch command was already 978 bytes. Adding +`--session-id ` for #214 — fifty bytes — made it 1028, and the four bytes cut off the end +turned `--autocompact 250000` into `--autocompact 25`. Claude Code refuses that value, so every +claude-code spawn died. The cut is at byte 1024 exactly; that was measured on a live pane, not +assumed. + +Two things changed. The charter now always travels as `--append-system-prompt-file`, which takes +~800 bytes of prose off the command line for good. And this guard catches whatever grows next. + +**The gotcha: the estimate is deliberately too big.** fleetd cannot see how herdr quotes each +argument, so every argument is charged its own bytes plus a separator and a quote pair. An +over-estimate costs you a clear error at a length that was already unsafe. An under-estimate would +let the silent truncation back in, and that is the failure this exists to prevent. + +**It already caught a second one.** A profile with `ideMcpUrl:` set assembles 1084 bytes — over the +limit before this ticket was written. The IDE mount was one config key away from the same silent +death. It cannot ship that way now. + +--- + +## A failed spawn shows you the pane + +**What.** When the spawn-readiness gate gives up on a member, the WARN carries the last status herdr +reported **and the tail of the pane**, read before the pane is closed. Automatic — no knob. + +**On.** Nothing to turn on. + +**Why it exists.** The gate used to log only that it had timed out, which is true of every possible +cause: a backend that never launched, a binary that rejected an argument and exited, a trust prompt +waiting for a keypress, a login shell that hung. The pane holds the one copy of that answer, and the +next line of code destroyed it. + +Finding #220 without this took a live process sampler, a hand-built pty and a byte count. With it, +the answer was one log line: `... --model claude-sonnet-5 --autocompact 25`. + +**The gotcha: read it before `stop()`.** The order matters and is easy to get backwards. `stop(paneId)` +closes the pane, and a pane read after that returns nothing — the log would be just as empty as +before, only slower. + +--- + +## Every claude-code member is resumable + +**What.** Every claude-code spawn gets a session id, so `fleet_list` always reports an +`agentSessionId` you can hand back to `fleet_spawn{resumeSessionId}`. Automatic — no knob. + +**On.** Nothing to turn on. + +**Why it exists.** A session id used to be minted only when the caller passed `sessionName` or +`resumeSessionId`, and `fleet_spawn` treats both as optional. So an ordinary spawn minted nothing and +`agentSessionId` stayed empty for that member's whole life. Unlike opencode, nothing resolves it +afterwards: a claude-code session id is fixed at launch. + +That made resume opt-in **at spawn**, while the moment you want it is *after* a member has done work +worth keeping. By then it was too late, and an empty field in the roster read as "this backend does +not support resume", which was not what was happening. + +**The gotcha: the binary takes a UUID and nothing else.** `claude --session-id` rejects any other +value at argument parsing — *"Invalid session ID. Must be a valid UUID."* — so the random UUID is +required, not incidental. The one documented flag interaction is with `-r`, which claims the same id; +a resume spawn passes `-r` and never mints. + +--- + +## An exhausted backend is quarantined even from a chrome-only pane + +**What.** When a member's pane yields no usable assistant block, fleetd still classifies +`BACKEND_EXHAUSTED` and `BACKEND_ERROR` from the raw screen — so an exhausted credential is +quarantined instead of being recorded as a member that did nothing. Automatic — no knob. + +**On.** Nothing to turn on. + +**Why it exists.** The completion resolver checked for an empty scrape **before** it looked for an +exhaustion line, and returned. `lastAssistantBlock` finds nothing whenever the pane carries no `⏺` +marker and its first visible line is TUI chrome — a box border, a warning, the input box — because +the boundary scan breaks on that first line. + +The exhaustion branch is the only caller of the quarantine sink. So a backend that refused the turn +for a usage limit was filed as "produced nothing", the credential was never quarantined, and fleetd +kept handing work to a backend that could not run it. Each attempt looked like another member +silently doing nothing. That is the shape of a usage limit taking out several members with the roster +giving no reason. + +**The gotcha: it is a fallback, not a reordering.** The empty-scrape failure was not moved or +weakened, and a pane that yields a usable block never reaches this path. Matching the raw screen +widens what the patterns can hit, including scrollback that is not this turn's output — and a false +positive here quarantines a **working** credential. That is why the raw match runs only where the +alternative was a generic failure with no information in it at all. + +--- + +## The credential scrub follows the member's own user + +**What.** When `memberHerdrSocket:` puts member panes under a different OS user, the allow-list +credential scrub uses that user's facts: `memberLoginShell:` decides whether the ZDOTDIR scrub can +run, and the generated directory goes under `worktreeRoot`, shared read-only with `worktreeGroup`. + +**On.** `memberLoginShell:` and `worktreeGroup:` — both required once `memberHerdrSocket:` is set. +With `memberHerdrSocket:` absent, nothing changes: fleetd's own `$SHELL` still decides and the +directory still goes to `java.io.tmpdir`, byte-identical to before. + +**Why it exists.** The scrub read fleetd's own `$SHELL` to decide whether a member's login shell +honours `ZDOTDIR`, and wrote the generated directory into fleetd's own `java.io.tmpdir`. Both +describe the wrong process once panes run as another user. + +The dangerous combination was fleetd on zsh and the member user on anything else. fleetd took the +zsh branch, generated a ZDOTDIR, and set it on the pane; the member's shell ignored `ZDOTDIR` +entirely, so no scrub ran — and because the zsh branch deliberately skips the sentinel overlay, the +fallback never happened either. **No protection at all, on the path fleetd believed was the protected +one.** Even with the member on zsh it still failed: `java.io.tmpdir` is mode 0700 on macOS, so +another uid cannot even traverse it. + +**The gotcha: it degrades, it never refuses to spawn.** If the member shell is not zsh, or +`worktreeRoot`/`worktreeGroup` is missing, fleetd warns loudly and falls back to the sentinel +overlay. A weakened credential control must not become an outage for an opt-in feature. The +permissions are deliberately tight in the other direction: the directory is `rwxr-x---` and each file +`rw-r-----` — group read, **no group write anywhere, no world bits**. A member sources what it needs +and cannot alter fleetd's own scrub. + ### A note for anyone briefing a worker to read this page **A worker cannot see the current version of this file.** `wiki/` is a submodule, and the parent