fleetd #176: subtract the lead's own subscription seat from free #254

Closed
agent wants to merge 0 commits from worker/fleetd-176-b928ca-3 into main
Member

fleetd #176 — correction: the stage-1 fix was inert on the live host

Stage 1 (commit c796eac) added Fleetd.leadSeatLookup, matching lead and target
profiles by effectiveCredentialId(). The logic was correct but the default
effectiveCredentialId() was not: it fell back to a profile's own name, so two
differently-named subscription: true profiles on the same Claude login never
matched. The lead pointed this out with a live measurement (lead on opus,
members on sonnet, both subscription: true, neither setting credentialId)
and ran leadSeatLookup against that exact shape, getting seats charged to sonnet = 0 — proving the bug rather than reasoning about it.

The sentinel chosen, and why

FleetConfig.Profile.effectiveCredentialId() now returns a shared constant,
SUBSCRIPTION_CREDENTIAL_ID = "<subscription>", when credentialId is unset
and subscription is true — instead of falling back to the profile's own
name. An explicit credentialId still wins over the sentinel in every case, so
an operator running two genuinely separate Claude logins on one host can still
set different credentialId values and keep them apart.

The reasoning: a subscription: true profile has no credential of its own —
it authenticates as the operator's own Claude login, and a host has exactly one
of those. Falling back to the profile's own name (the way an ordinary
off-subscription profile does) kept two subscription profiles on the same login
apart from each other, which is the opposite of what "one account" means, and
is exactly what made stage 1 inert.

The live-shape test going red on the old code

New test FleetdLeadSeatLookupTest.leadOnADifferentSubscriptionProfileSameAccountStillCountsAsASeat
puts the lead on opus and the target on sonnet — different profile names,
same account, neither setting credentialId. Every pre-existing test in that
class put the lead on the same profile name as the target, so this exact
shape had zero coverage before this PR.

Mutation test — reverted effectiveCredentialId()'s subscription branch back
to return profile; (the old behaviour) and re-ran:

[ERROR] Tests run: 9, Failures: 1, Errors: 0, Skipped: 0, Time elapsed: 0.127 s <<< FAILURE! -- in dev.ltms.fleet.FleetdLeadSeatLookupTest
[ERROR] dev.ltms.fleet.FleetdLeadSeatLookupTest.leadOnADifferentSubscriptionProfileSameAccountStillCountsAsASeat -- Time elapsed: 0.004 s <<< FAILURE!
org.opentest4j.AssertionFailedError: opus and sonnet are both subscription:true with no explicit credentialId, so they share one Claude login and the lead's live seat on opus must be charged against sonnet too — this is the live host's actual shape (fleetd #176 stage 2) ==> expected: <1> but was: <0>
[ERROR] Tests run: 9, Failures: 1, Errors: 0, Skipped: 0

0 compile errors, exactly the one intended test went red, every other test in
the class stayed green. Reverted the mutation and confirmed the source file was
restored byte-identical via diff -q (exit 0) before re-running the full
build.

The same mutation was also run against the new FleetMcpTest
quarantiningOneSubscriptionProfileZeroesFreeOnTheOtherSharingTheSameAccount
test — same result: 1 failure, 0 errors, restored and re-verified green.

Every other effectiveCredentialId() caller — reviewed

I grepped every call site and checked whether each one wants "this exact
profile" (which the sentinel would break) or "this account" (which it fixes).
All five want "this account":

  1. Fleetd.exhaustionSink — quarantine.quarantine(profile.effectiveCredentialId())
    on a BACKEND_EXHAUSTED classification. Wants "this account": a usage-limit
    hit really should quarantine every profile sharing that login.
  2. Fleetd.main's QuarantineSource/OutageSource construction —
    profile -> configured.effectiveCredentialId(), used purely to look up and
    report quarantine/cool-off state per profile in fleet_list/fleet_profiles.
    Consistent with (1) by construction.
  3. Fleetd.backendErrorSink — feeds BackendOutagePolicy.record(credentialId, ...)
    and builds the affectedProfiles list for the lead notification by filtering
    on credentialId.equals(p.effectiveCredentialId()). Wants "this account":
    repeated backend errors on one subscription profile should cool off every
    profile sharing that login, and the notification should name all of them.
  4. Fleetd.leadSeatLookup (this ticket) — wants "this account" by
    definition; that is the whole point of the correction.
  5. CompositePeerLauncher.credentialIdFor, used by enforceNotQuarantined,
    enforceNotCoolingOff, and quarantinedProfiles — the actual spawn-time
    placement gate, separate from FleetMcp's reporting-only capacity view. This
    is a real, previously-unnoticed behavioural change worth flagging explicitly:
    before this PR, quarantining opus did not block fleet_spawn{profile: "sonnet"} even though both share the same Claude login; after this PR, it
    does. That is a second bug this same fix corrects, not a side effect — a
    spawn onto a quarantined account should be refused regardless of which
    profile name is used to reach it.

No caller wants per-profile identity over per-account identity.

fleetd.example.yaml

Rewrote the fleetd #176 notes under subscription: true (GOTCHA 3 + new "THE
SENTINEL" / "WHEN TO OVERRIDE" blocks) and under leaders:'s profile: field.
The old text framed the risk as "set profile: or seats read 0"; that's still
true but incomplete. The new text explains: subscription profiles with no
explicit credentialId share one implicit account-wide id automatically (so
opus and sonnet link with zero extra config), that this same id drives
quarantine/cool-off grouping, and that an explicit, differing credentialId on
each profile is how an operator with two real separate Claude logins keeps them
apart.

Build

cd fleetd && mvn clean install, plain and unpiped:

[INFO] Tests run: 1243, Failures: 0, Errors: 0, Skipped: 0
[INFO] BUILD SUCCESS

(1241 tests before this PR's 2 new tests: the live-shape lead-seat test and the
quarantine-sharing capacity test.)

Caveat for review

I could not spawn members or restart the daemon to verify this against the
live fleet directly — that verification is the lead's, as instructed. Everything
above is grepped/read/tested inside this worktree only.

# fleetd #176 — correction: the stage-1 fix was inert on the live host Stage 1 (commit c796eac) added `Fleetd.leadSeatLookup`, matching lead and target profiles by `effectiveCredentialId()`. The logic was correct but the *default* `effectiveCredentialId()` was not: it fell back to a profile's own name, so two differently-named `subscription: true` profiles on the same Claude login never matched. The lead pointed this out with a live measurement (lead on `opus`, members on `sonnet`, both `subscription: true`, neither setting `credentialId`) and ran `leadSeatLookup` against that exact shape, getting `seats charged to sonnet = 0` — proving the bug rather than reasoning about it. ## The sentinel chosen, and why `FleetConfig.Profile.effectiveCredentialId()` now returns a shared constant, `SUBSCRIPTION_CREDENTIAL_ID = "<subscription>"`, when `credentialId` is unset **and** `subscription` is `true` — instead of falling back to the profile's own name. An explicit `credentialId` still wins over the sentinel in every case, so an operator running two genuinely separate Claude logins on one host can still set different `credentialId` values and keep them apart. The reasoning: a `subscription: true` profile has no credential of its own — it authenticates as the operator's own Claude login, and a host has exactly one of those. Falling back to the profile's own name (the way an ordinary off-subscription profile does) kept two subscription profiles on the same login apart from each other, which is the opposite of what "one account" means, and is exactly what made stage 1 inert. ## The live-shape test going red on the old code New test `FleetdLeadSeatLookupTest.leadOnADifferentSubscriptionProfileSameAccountStillCountsAsASeat` puts the lead on `opus` and the target on `sonnet` — different profile names, same account, neither setting `credentialId`. Every pre-existing test in that class put the lead on the *same* profile name as the target, so this exact shape had zero coverage before this PR. Mutation test — reverted `effectiveCredentialId()`'s subscription branch back to `return profile;` (the old behaviour) and re-ran: ``` [ERROR] Tests run: 9, Failures: 1, Errors: 0, Skipped: 0, Time elapsed: 0.127 s <<< FAILURE! -- in dev.ltms.fleet.FleetdLeadSeatLookupTest [ERROR] dev.ltms.fleet.FleetdLeadSeatLookupTest.leadOnADifferentSubscriptionProfileSameAccountStillCountsAsASeat -- Time elapsed: 0.004 s <<< FAILURE! org.opentest4j.AssertionFailedError: opus and sonnet are both subscription:true with no explicit credentialId, so they share one Claude login and the lead's live seat on opus must be charged against sonnet too — this is the live host's actual shape (fleetd #176 stage 2) ==> expected: <1> but was: <0> [ERROR] Tests run: 9, Failures: 1, Errors: 0, Skipped: 0 ``` 0 compile errors, exactly the one intended test went red, every other test in the class stayed green. Reverted the mutation and confirmed the source file was restored byte-identical via `diff -q` (exit 0) before re-running the full build. The same mutation was also run against the new `FleetMcpTest` `quarantiningOneSubscriptionProfileZeroesFreeOnTheOtherSharingTheSameAccount` test — same result: 1 failure, 0 errors, restored and re-verified green. ## Every other `effectiveCredentialId()` caller — reviewed I grepped every call site and checked whether each one wants "this exact profile" (which the sentinel would break) or "this account" (which it fixes). All five want "this account": 1. **`Fleetd.exhaustionSink`** — `quarantine.quarantine(profile.effectiveCredentialId())` on a `BACKEND_EXHAUSTED` classification. Wants "this account": a usage-limit hit really should quarantine every profile sharing that login. 2. **`Fleetd.main`'s `QuarantineSource`/`OutageSource` construction** — `profile -> configured.effectiveCredentialId()`, used purely to look up and report quarantine/cool-off state per profile in `fleet_list`/`fleet_profiles`. Consistent with (1) by construction. 3. **`Fleetd.backendErrorSink`** — feeds `BackendOutagePolicy.record(credentialId, ...)` and builds the `affectedProfiles` list for the lead notification by filtering on `credentialId.equals(p.effectiveCredentialId())`. Wants "this account": repeated backend errors on one subscription profile should cool off every profile sharing that login, and the notification should name all of them. 4. **`Fleetd.leadSeatLookup`** (this ticket) — wants "this account" by definition; that is the whole point of the correction. 5. **`CompositePeerLauncher.credentialIdFor`**, used by `enforceNotQuarantined`, `enforceNotCoolingOff`, and `quarantinedProfiles` — the actual **spawn-time** placement gate, separate from `FleetMcp`'s reporting-only capacity view. This is a real, previously-unnoticed behavioural change worth flagging explicitly: before this PR, quarantining `opus` did **not** block `fleet_spawn{profile: "sonnet"}` even though both share the same Claude login; after this PR, it does. That is a second bug this same fix corrects, not a side effect — a spawn onto a quarantined account should be refused regardless of which profile name is used to reach it. No caller wants per-profile identity over per-account identity. ## `fleetd.example.yaml` Rewrote the fleetd #176 notes under `subscription: true` (GOTCHA 3 + new "THE SENTINEL" / "WHEN TO OVERRIDE" blocks) and under `leaders:`'s `profile:` field. The old text framed the risk as "set `profile:` or seats read 0"; that's still true but incomplete. The new text explains: subscription profiles with no explicit `credentialId` share one implicit account-wide id automatically (so `opus` and `sonnet` link with zero extra config), that this same id drives quarantine/cool-off grouping, and that an explicit, differing `credentialId` on each profile is how an operator with two real separate Claude logins keeps them apart. ## Build `cd fleetd && mvn clean install`, plain and unpiped: ``` [INFO] Tests run: 1243, Failures: 0, Errors: 0, Skipped: 0 [INFO] BUILD SUCCESS ``` (1241 tests before this PR's 2 new tests: the live-shape lead-seat test and the quarantine-sharing capacity test.) ## Caveat for review I could not spawn members or restart the daemon to verify this against the live fleet directly — that verification is the lead's, as instructed. Everything above is grepped/read/tested inside this worktree only.
agent added 1 commit 2026-09-03 07:56:00 +02:00
fleetd #176: subtract the lead's own subscription seat from free
CI / contract (pull_request) Successful in 1m11s
CI / build (pull_request) Successful in 1m16s
c796eac09c
maxLoad counted panes, never subscription seats: a subscription:true
profile's lead is itself a live claude session on that same account,
so free overstated capacity by the lead's own seat (measured free:1
with a real ceiling of 0, and free:3 on an idle fleet with a real
ceiling of 2).

Add FleetMcp.LeadSeatSource (same shape as QuarantineSource/
OutageSource) and Fleetd.leadSeatLookup, which derives the seat count
from fleet.leaders.<name>.profile matched against the target profile
by effectiveCredentialId() - no hardcoded "-1", and no new config key:
profile: already exists for this exact "which account does this lead
share" question. maxLoad itself is left untouched; only free (and a
new, additive-only leadSeats field) changes.

Exhaustion quarantine (cause 2 in the ticket) already forced free to 0
via the same BackendQuarantine capacityView already reads - confirmed
by reading the exhaustionSink wiring, no code change needed there.
agent added 1 commit 2026-09-03 08:10:19 +02:00
fleetd #176 stage 2: make effectiveCredentialId() subscription-aware
CI / contract (pull_request) Successful in 44s
CI / build (pull_request) Successful in 1m44s
c50f5b2d61
Stage 1's lead-seat matcher (leadSeatLookup) was correct but inert on
the live host: the lead runs on profile 'opus', members on 'sonnet',
both subscription:true with no explicit credentialId. Because
effectiveCredentialId() fell back to the profile's own name, opus and
sonnet never matched even though they share one Claude login, so the
matcher charged zero seats.

FleetConfig.Profile.effectiveCredentialId() now falls back to a shared
sentinel (SUBSCRIPTION_CREDENTIAL_ID = "<subscription>") instead of the
profile name when subscription:true and credentialId is unset. An
explicit credentialId still wins, so two separate Claude logins on one
host can still be kept apart.

This is also BackendQuarantine's and BackendOutagePolicy's grouping
key and CompositePeerLauncher's spawn-time enforcement key, so the fix
also links quarantine/cool-off across subscription profiles sharing an
account -- intentional: one usage limit really does take out every
profile on that login, mirroring credentialId: openai-shared already
doing this for off-subscription profiles. Every caller was reviewed;
none wants "this exact profile" over "this account".

Tests added:
- FleetdLeadSeatLookupTest: the live shape itself (lead on a
  DIFFERENT subscription profile than the target, same account,
  neither sets credentialId) -- the case stage 1's suite never covered
- FleetMcpTest: quarantining one subscription profile's shared
  account zeroes free on another sharing it, via the same
  effectiveCredentialId()-driven wiring Fleetd.main uses

Mutation-tested: reverting the subscription branch to the old
fall-back-to-profile-name behavior sends both new tests RED with 0
compile errors; reverting the mutation restores byte-identical
(diff -q) source and green tests.

fleetd.example.yaml's fleetd #176 notes are rewritten for the sentinel
semantics and when to override it with an explicit credentialId.
Owner

Merged into main as 0e8bfb7 (merge commit, after rebuilding the merged tree — the branch was five commits behind).

Closing manually: the merge went in from the worktree rather than through the Gitea merge button, so this PR did not auto-close.

Adjudication and the live verification are recorded on #176. Short version: the fix works on this host — seats charged to sonnet = 1, where stage 1 gave 0 — confirmed again on the live daemon after redeploy (fleet_list moved from sonnet free: 3 to free: 2, leadSeats: 1).

Follow-up filed as #257: the subtraction is reporting-only, so free and the spawn gate now disagree by one.

Merged into `main` as `0e8bfb7` (merge commit, after rebuilding the merged tree — the branch was five commits behind). Closing manually: the merge went in from the worktree rather than through the Gitea merge button, so this PR did not auto-close. Adjudication and the live verification are recorded on #176. Short version: the fix works on this host — `seats charged to sonnet = 1`, where stage 1 gave 0 — confirmed again on the live daemon after redeploy (`fleet_list` moved from `sonnet free: 3` to `free: 2, leadSeats: 1`). Follow-up filed as #257: the subtraction is reporting-only, so `free` and the spawn gate now disagree by one.
ltms closed this pull request 2026-09-03 08:39:59 +02:00
Some checks are pending
CI / contract (pull_request) Successful in 44s
CI / build (pull_request) Successful in 1m44s

Pull request closed

Sign in to join this conversation.