Features: correct the free/leadSeats entry, and record the trust-seed CAS

Two entries described behaviour that changed today.

#257: the 'free counts the lead's seat' entry described a subtraction that
has been removed. free now means what the spawn gate grants. leadSeats is
still reported, as a fact beside it. Renamed the entry to say so, and kept
the history of why the subtraction was tried and dropped.

#247: the workspace-trust entry said the lost-update race against the
operator's own Claude Code was 'tracked separately'. It is now fixed with a
compare-and-swap plus a bounded retry that writes nothing rather than
overwrite. Recorded that, the two new WARNs, and the window that remains.
Dai Ha
2026-09-03 16:39:04 +07:00
parent 3a57e56677
commit 573eed33fb
+49 -16
@@ -2997,9 +2997,29 @@ why the write is gated on the directory actually being a fleetd-provisioned work
the path merely being set — a default-resolved cwd is the operator's own checkout. The write is
additive and atomic, and a process-wide lock serialises concurrent spawns.
That lock cannot reach the operator's **own** running Claude Code, which writes the same file. A save
landing between fleetd's read and its write is still lost. The file is never torn, but a change can
vanish. Tracked separately; do not read the atomic write as covering that case.
**The second gotcha: the lock cannot reach the operator's own Claude Code (fleetd #247).** That
in-process lock only serialises spawns inside one JVM. The operator's own running Claude Code writes
the same file, and on this host it is literally the same file: `opus` and `sonnet` both set
`configDir` to the lead's own config directory. So this was never a rare collision — it was a lost
update on every ordinary spawn of those profiles.
The write is now a **compare-and-swap with a bounded retry**. fleetd reads the file's bytes, builds
its update, then re-reads the bytes immediately before the atomic move and compares. If they changed,
another writer got in, so it throws its work away and rebuilds from the fresh bytes — up to five
times.
**On exhausting the retries it writes nothing** and logs a WARN naming the cwd and the file. That is
deliberate. An unseeded member shows the trust dialog and fails to reach an injectable state: visible,
logged, and recoverable by retrying the spawn. Writing a stale copy over the operator's live config is
silent and not recoverable. The code fails toward the recoverable outcome.
A second WARN now fires every time the seed is about to target the default `~/.claude.json`, which
happens only when a profile sets no `configDir`. That is the one path reaching the operator's home
file, and it is the default, so a profile that simply forgot the setting used to get no signal at all.
**What this still does not fix.** The window is narrowed, not closed. A write landing between the
final re-read and the move itself is still lost. There is no operating-system compare-and-swap on a
plain file, only this cooperative narrowing. Do not read the CAS as making the race gone.
## A worktree tells you which of its config files are stubs
@@ -3093,13 +3113,18 @@ describes the wrong account and only the configured `memberLoginShell:` is read.
`memberHerdrSocket` and leave `memberLoginShell` unset, the shell reads as `<unset>`, counts as
non-zsh, and **every spawn is refused**. The refusal message says so, but it is cheaper to know first.
## `free` counts the seat the lead itself holds
## `leadSeats` reports the seat the lead itself holds
**What.** For a `subscription: true` profile, `fleet_list`'s `free` now subtracts the seat the lead's
own session occupies on that same account, and the row carries a `leadSeats` field when the count is
positive. Profiles are grouped by **account**, not by profile name: a profile with no explicit
`credentialId` that is `subscription: true` joins a shared `<subscription>` group. An explicit
`credentialId` still wins, so two genuinely separate Claude logins on one host stay apart.
**What.** For a `subscription: true` profile, `fleet_list`'s row carries a `leadSeats` field when the
lead's own session occupies a seat on that same account. Profiles are grouped by **account**, not by
profile name: a profile with no explicit `credentialId` that is `subscription: true` joins a shared
`<subscription>` group. An explicit `credentialId` still wins, so two genuinely separate Claude
logins on one host stay apart.
**`free` is not reduced by `leadSeats`.** `free` means one thing only: how many slots a fresh
`fleet_spawn` on that profile will actually be granted right now, which is `max(0, maxLoad - live)` —
the same check the spawn gate runs. `leadSeats` is a fact reported beside it, for a caller to use
however it likes.
**On.** Automatic for `subscription: true` profiles, once `fleet.leaders.<name>.profile` is set.
@@ -3111,13 +3136,21 @@ The same grouping fixes a second, wider bug: quarantining one subscription profi
spawns on the other. One subscription hitting a usage limit really does take out every profile
running on it.
**The gotcha: `free` and the spawn gate now disagree by one, deliberately.** The subtraction is
**reporting only**. The placement gate uses `maxLoad` and the live count and never consults the lead
seat, so with `maxLoad: 3` and two members up, `fleet_list` says `free: 0` while a third spawn still
succeeds. That is the opposite of the overstatement this was filed to fix, and which of the two
numbers is "right" is an open question — see the follow-up ticket before changing `maxLoad` to
compensate. Raising `maxLoad` to win the slot back grants a real extra member; the slot was never
taken.
**The gotcha, and how it was resolved (fleetd #257).** The first cut of this **did** subtract
`leadSeats` from `free`. That shipped, and was corrected the same day. The placement gate never
consulted the lead seat, so with `maxLoad: 3` and two members up, `fleet_list` said `free: 0` while a
third spawn still succeeded. A lead that believed the number gave up a slot the daemon would have
granted — the opposite of the overstatement this was filed to fix.
The subtraction was removed rather than pushed into the spawn gate, for two reasons. Making the gate
subtract it would change what `maxLoad: 3` means in every existing config file, on every host. And no
backend seat ceiling shared with the lead has ever been measured: a test on 2026-08-29 ran three
concurrent interactive `claude` sessions with no trouble. Subtracting the seat therefore described an
accounting policy, not a constraint, and dressing a policy as a capacity fact is what made the two
numbers disagree.
Do not raise `maxLoad` to "win a slot back". Nothing takes one; you would be granting a real extra
member.
The first attempt at this shipped **inert** and every test passed: the matcher grouped by
`effectiveCredentialId()`, which fell back to the profile's own *name*, so a lead on `opus` and