From 573eed33fb1580982b2892d4a4730b84ba0fbac1 Mon Sep 17 00:00:00 2001 From: Dai Ha Date: Thu, 3 Sep 2026 16:39:04 +0700 Subject: [PATCH] Features: correct the free/leadSeats entry, and record the trust-seed CAS Two entries described behaviour that changed today. #257: the 'free counts the lead's seat' entry described a subtraction that has been removed. free now means what the spawn gate grants. leadSeats is still reported, as a fact beside it. Renamed the entry to say so, and kept the history of why the subtraction was tried and dropped. #247: the workspace-trust entry said the lost-update race against the operator's own Claude Code was 'tracked separately'. It is now fixed with a compare-and-swap plus a bounded retry that writes nothing rather than overwrite. Recorded that, the two new WARNs, and the window that remains. --- 11-Features.md | 65 +++++++++++++++++++++++++++++++++++++------------- 1 file changed, 49 insertions(+), 16 deletions(-) diff --git a/11-Features.md b/11-Features.md index 7366c8c..f935188 100644 --- a/11-Features.md +++ b/11-Features.md @@ -2997,9 +2997,29 @@ why the write is gated on the directory actually being a fleetd-provisioned work the path merely being set — a default-resolved cwd is the operator's own checkout. The write is additive and atomic, and a process-wide lock serialises concurrent spawns. -That lock cannot reach the operator's **own** running Claude Code, which writes the same file. A save -landing between fleetd's read and its write is still lost. The file is never torn, but a change can -vanish. Tracked separately; do not read the atomic write as covering that case. +**The second gotcha: the lock cannot reach the operator's own Claude Code (fleetd #247).** That +in-process lock only serialises spawns inside one JVM. The operator's own running Claude Code writes +the same file, and on this host it is literally the same file: `opus` and `sonnet` both set +`configDir` to the lead's own config directory. So this was never a rare collision — it was a lost +update on every ordinary spawn of those profiles. + +The write is now a **compare-and-swap with a bounded retry**. fleetd reads the file's bytes, builds +its update, then re-reads the bytes immediately before the atomic move and compares. If they changed, +another writer got in, so it throws its work away and rebuilds from the fresh bytes — up to five +times. + +**On exhausting the retries it writes nothing** and logs a WARN naming the cwd and the file. That is +deliberate. An unseeded member shows the trust dialog and fails to reach an injectable state: visible, +logged, and recoverable by retrying the spawn. Writing a stale copy over the operator's live config is +silent and not recoverable. The code fails toward the recoverable outcome. + +A second WARN now fires every time the seed is about to target the default `~/.claude.json`, which +happens only when a profile sets no `configDir`. That is the one path reaching the operator's home +file, and it is the default, so a profile that simply forgot the setting used to get no signal at all. + +**What this still does not fix.** The window is narrowed, not closed. A write landing between the +final re-read and the move itself is still lost. There is no operating-system compare-and-swap on a +plain file, only this cooperative narrowing. Do not read the CAS as making the race gone. ## A worktree tells you which of its config files are stubs @@ -3093,13 +3113,18 @@ describes the wrong account and only the configured `memberLoginShell:` is read. `memberHerdrSocket` and leave `memberLoginShell` unset, the shell reads as ``, counts as non-zsh, and **every spawn is refused**. The refusal message says so, but it is cheaper to know first. -## `free` counts the seat the lead itself holds +## `leadSeats` reports the seat the lead itself holds -**What.** For a `subscription: true` profile, `fleet_list`'s `free` now subtracts the seat the lead's -own session occupies on that same account, and the row carries a `leadSeats` field when the count is -positive. Profiles are grouped by **account**, not by profile name: a profile with no explicit -`credentialId` that is `subscription: true` joins a shared `` group. An explicit -`credentialId` still wins, so two genuinely separate Claude logins on one host stay apart. +**What.** For a `subscription: true` profile, `fleet_list`'s row carries a `leadSeats` field when the +lead's own session occupies a seat on that same account. Profiles are grouped by **account**, not by +profile name: a profile with no explicit `credentialId` that is `subscription: true` joins a shared +`` group. An explicit `credentialId` still wins, so two genuinely separate Claude +logins on one host stay apart. + +**`free` is not reduced by `leadSeats`.** `free` means one thing only: how many slots a fresh +`fleet_spawn` on that profile will actually be granted right now, which is `max(0, maxLoad - live)` — +the same check the spawn gate runs. `leadSeats` is a fact reported beside it, for a caller to use +however it likes. **On.** Automatic for `subscription: true` profiles, once `fleet.leaders..profile` is set. @@ -3111,13 +3136,21 @@ The same grouping fixes a second, wider bug: quarantining one subscription profi spawns on the other. One subscription hitting a usage limit really does take out every profile running on it. -**The gotcha: `free` and the spawn gate now disagree by one, deliberately.** The subtraction is -**reporting only**. The placement gate uses `maxLoad` and the live count and never consults the lead -seat, so with `maxLoad: 3` and two members up, `fleet_list` says `free: 0` while a third spawn still -succeeds. That is the opposite of the overstatement this was filed to fix, and which of the two -numbers is "right" is an open question — see the follow-up ticket before changing `maxLoad` to -compensate. Raising `maxLoad` to win the slot back grants a real extra member; the slot was never -taken. +**The gotcha, and how it was resolved (fleetd #257).** The first cut of this **did** subtract +`leadSeats` from `free`. That shipped, and was corrected the same day. The placement gate never +consulted the lead seat, so with `maxLoad: 3` and two members up, `fleet_list` said `free: 0` while a +third spawn still succeeded. A lead that believed the number gave up a slot the daemon would have +granted — the opposite of the overstatement this was filed to fix. + +The subtraction was removed rather than pushed into the spawn gate, for two reasons. Making the gate +subtract it would change what `maxLoad: 3` means in every existing config file, on every host. And no +backend seat ceiling shared with the lead has ever been measured: a test on 2026-08-29 ran three +concurrent interactive `claude` sessions with no trouble. Subtracting the seat therefore described an +accounting policy, not a constraint, and dressing a policy as a capacity fact is what made the two +numbers disagree. + +Do not raise `maxLoad` to "win a slot back". Nothing takes one; you would be granting a real extra +member. The first attempt at this shipped **inert** and every test passed: the matcher grouped by `effectiveCredentialId()`, which fell back to the profile's own *name*, so a lead on `opus` and