Compare commits

..

39 Commits

Author SHA1 Message Date
Dai Ha d4f93a7b13 fleetd #449: fix stale herdr protocol 14 javadocs/assertion, diagnose and fix the timing-raced AgentControlContractTest, select contract tests by tag in CI
CI / contract (pull_request) Successful in 46s
CI / build (pull_request) Successful in 2m9s
- HerdrClient.java, HerdrCodec.java, HerdrContractTest.java: the herdr port to
  protocol 19 (CB-521) left the client javadoc and the contract test's own
  assertion still saying protocol 14 / herdr 0.7.0. Updated to 19 / 0.8.0 and
  renamed pingReturnsProtocol14 -> pingReturnsProtocol19. Verified the
  assertion is real by temporarily changing the expected value to 20 (fails),
  then restoring 19 (passes).

- AgentControlContractTest.java: tabCreateInjectsEnvIntoTheSeedShell was
  failing, not skipping, on a host with a live herdr socket. Diagnosed with a
  temporary instrumented run (not committed) that polled the pane every
  200ms before and after sending input: the seed shell reliably takes ~2.5s
  to reach its prompt (measured 3x), while the test's fixed 1000ms sleep
  raced that startup. Input typed too early was swallowed by the shell's own
  startup, leaving the typed line followed by the "Restored session" banner
  and no command output — indistinguishable at a glance from the env map
  never reaching the shell. Once the shell was actually ready, the injected
  env value showed up in ~200ms, ruling out an env-seam defect. Replaced both
  fixed sleeps with bounded polling on the actual conditions (pane text
  settling, then the expected output appearing). Ran the fixed test 3x
  standalone, all green.

- .gitea/workflows/ci.yml: the "Contract tests" step ran exactly one class by
  name (-Dtest=AmqpReplyInboxContractTest), silently excluding every other
  @Tag("contract") test from CI including the herdr ones above -- which is
  how the stale protocol 14 assertion went unnoticed. Changed to
  -Dgroups=contract, which selects the whole tagged group and picks up
  future contract tests automatically.
2026-09-10 17:14:09 +07:00
ltms 822327eed5 Merge #447: pin the place()-to-spawn() window PlacementDecision closes (fleetd #444)
CI / contract (push) Successful in 1m33s
CI / build (push) Successful in 1m36s
Verified by the lead on the exact tree that lands (head e4c703a, base 82fae94 is
an ancestor, so this is the tree I measured):

  FULL BUILD  Tests run: 1575, Failures: 0, Errors: 0, Skipped: 0  BUILD SUCCESS
              compile errors: 0
  CONTROL     CompositePeerLauncherTest  Tests run: 76, Failures: 0  -> GREEN

  M1  the 2-arg spawn re-enters the 1-arg spawn (the inherited default this
      ticket forbids for a multi-profile launcher)
      -> KILLED  Errors: 1
      CompositePeerLauncherTest
        .spawnHonorsAPlacementDecisionEvenAfterItsProfileIsQuarantinedInTheWindowAfterPlace

  M2  drop the stamping: route to the decided profile but do not carry it
      (SpawnRequest routedReq = req)
      -> KILLED  Failures: 1
      same test method

Both mutations proved applied by printing the mutated method, and the tree was
restored clean after each (git status --porcelain empty).

Round 1 of this PR had a fixture weakness I found by mutation: StubLauncher's own
fallback default was "sol", the same profile place() decides, so an unstamped
request landed on spawnCount("sol") by coincidence and M2 survived. e4c703a gives
the adapter "b" as its fallback instead. One fixture now kills both mutations.

src/main is javadoc-only in this PR: 12 added lines, 0 added code lines, measured
by filtering the main-side diff.
2026-09-10 12:00:59 +02:00
Dai Ha e4c703a51a fleetd #444: separate the adapter's fallback default from the decided profile
CI / contract (pull_request) Successful in 46s
CI / build (pull_request) Successful in 1m56s
Review found the fixture's StubLauncher fell back to 'sol' too — the
same profile place() decides — so an UNSTAMPED request could land on
spawnCount('sol') by coincidence, and the assertion's claim that the
request 'actually carried sol' was unproven. Dropping the stamping
(SpawnRequest routedReq = req) while keeping the routing survived the
test unchanged.

Fix: give the adapter 'b' as its own fallback default instead, so an
unstamped request counts against 'b', not 'sol'. Verified both
mutations against the single test in isolation:
  - drop-stamping (routedReq = req): RED, expected <sol> but was <b>
  - re-entering (return spawn(req.withProfile(decision.profile())))
    i.e. M1 from the first round: still RED, PlacementException
    naming the now-quarantined 'sol'
Restored both; full suite green at 1575 tests.

No changes to src/main — PeerLauncher's javadoc from the first round
is unchanged.
2026-09-10 16:50:39 +07:00
Dai Ha 3f036b2a62 fleetd #444: pin the place()-to-spawn() window PlacementDecision closes
CI / contract (pull_request) Successful in 1m4s
CI / build (pull_request) Successful in 1m53s
Add a test that resolves place(role) while nothing is quarantined,
then quarantines the resolved profile's credential BEFORE spawning
against the held PlacementDecision. CompositePeerLauncher.spawn(req,
decision) must still honor the decision and land on the quarantined
profile, since it never re-runs the explicit-profile enforce* checks.

Verified the test kills the regression: with the override's body
replaced by the re-entering spawn(req.withProfile(...)) form, this
exact test goes RED with a PlacementException naming the now-
quarantined profile; restored, the full suite is green (1575 tests).

Also documents on PeerLauncher's default spawn(req, decision) that a
launcher routing across more than one profile MUST override it,
naming the four enforce* checks the default's re-entry re-applies.
2026-09-10 16:41:52 +07:00
ltms 82fae94c55 Merge #445: pin every startup report call in Fleetd.main (fleetd #442)
CI / contract (push) Successful in 48s
CI / build (push) Successful in 1m59s
Test written by a worker whose backend died before it could report; evidence
re-run by the lead against the merged tree.

Verified: merge of current main clean (0 conflicts); full build 1574 tests,
0 failures, BUILD SUCCESS, 0 compile errors; control green; deleting each of
reportGitHostShape, reportMemberTrustModel, reportMemberCredentialsGap and
reportExhaustedPatternGap from main() is KILLED by
mainReportsEveryStartupGapBeforeValidationAborts.
2026-09-10 11:30:31 +02:00
Dai Ha b1f34c2e6b fleetd #442: drop the unused java.util.List import
CI / contract (pull_request) Successful in 47s
CI / build (pull_request) Successful in 2m8s
The new test never names List — only ListAppender, which has its own
import. An unused import is an IDE warning, and this repo treats
warnings as gates. No behaviour change: FleetdStartupReportTest still
runs 1 test, 0 failures, BUILD SUCCESS, 0 compile errors.
2026-09-10 16:29:56 +07:00
ltms 3f807d9f1b Merge #443: derive coordinator.heldDurable from queue durability + ack mode (fleetd #440)
CI / contract (push) Successful in 51s
CI / build (push) Successful in 1m35s
Found by the fleet01 lead reviewing #438 after I had merged it. Verified
independently before merging.

The implementation choice is the load-bearing part: LeadMailbox.own() now
assigns queueDeclare's durable flag and basicConsume's autoAck flag to named
locals, passes those SAME locals into the two real AMQP calls (:203/:204), and
derives heldDurable from them (:207). So the reported fact cannot drift from a
duplicate constant - a mutation to either call's argument moves the behaviour
and the report together. LeadChannel.heldDurable() is abstract, so a future
implementer gets a compile error rather than a silent default.

My own battery, merged tree, control green, tree restored clean:
- M1 revert to the literal true -> KILLED by
  FleetMcpTest.listReportsHeldDurableFalseWhenTheChannelSaysMailIsNotDurable
- M3 heldCount forced to 0 (a half the worker did not touch) -> KILLED by
  FleetMcpTest.listReportsAnHonestHeldCountAndDurabilityNotJustPendingZero
- M2 break the derivation itself -> SURVIVED under plain clean install, exactly
  as the worker reported. LeadMailbox needs a real broker, so the only test that
  reaches the real queueDeclare/basicConsume is @Tag("contract"), excluded from
  the default build. The worker ran that arm with -Pcontract and got 3 reds
  including its own new test. Pre-existing structural limit of this class, not
  introduced here, and the worker flagged it rather than hiding it.

Full build, my own run: Tests run: 1573, Failures: 0, Errors: 0, Skipped: 0 -
BUILD SUCCESS, 0 compile errors.

Not blocking, noted for a possible follow-up: one boolean over two independent
facts cannot say WHICH fact was lost. The merged field is still strictly better
than the literal it replaces, because it can now go false at all.
2026-09-10 11:21:41 +02:00
Dai Ha e70263062c fleetd #442: pin startup report calls
CI / contract (pull_request) Successful in 1m15s
CI / build (pull_request) Successful in 1m36s
2026-09-10 14:41:10 +07:00
ltms 1a1e586b62 Merge #433: carry the PlacementDecision instead of re-resolving the profile (fleetd #425)
CI / contract (push) Successful in 45s
CI / build (push) Successful in 1m47s
Round 4 pins the fix. Verified independently before merging.

My own battery in the worker's tree (control green, tree restored clean):
- M1 revert SessionManager:677 to launcher.spawn(spawnReq) -> KILLED by
  SessionManagerTest.acquireWithWorktreeSpawnsOnTheSameProfileItProvisionedThe
  WorktreeForUnderARotatingPolicy. This is the ticket's own deliverable and it
  survived 186 tests in round 3.
- M3 drop the withProfile stamping in the 2-arg spawn -> KILLED by 3 tests.
- M2 make the 2-arg spawn re-enter the refusing branch -> SURVIVED, but it is a
  near-equivalent mutant, not a gap in this work. Post-#435 all four enforce*
  conditions are ones the routing branch already filtered on, so the two paths
  differ only if placement state moves between place() and spawn(). Filed
  separately.

Full build, my own run this turn, whole log redirected and grepped:
Tests run: 1572, Failures: 0, Errors: 0, Skipped: 0 - BUILD SUCCESS, 0 compile
errors. Branch already contains current main.

Read the src/main diff. The new test uses roundRobin() (stateful) and asserts
agreement between the overlay profile and the spawned profile, rather than a
hardcoded name, which is the right shape - select() is stateful, so two calls
disagree by design.
2026-09-10 09:35:02 +02:00
Dai Ha c16d118f09 fleetd #440: derive coordinator.heldDurable from queue durability + ack mode
CI / contract (pull_request) Successful in 47s
CI / build (pull_request) Successful in 1m36s
FleetMcp.coordinatorView wrote heldDurable as a literal true, so a change
that broke either the durable queue declare or the manual-ack consume in
LeadMailbox would leave the field, and the full suite, green.

- LeadChannel gets a new heldDurable() method: the conclusion of a durable
  queue declare AND a manual-ack consumer, derived by the implementation
  from what it actually did, never asserted.
- LeadMailbox.own() captures the exact booleans it passes to
  queueDeclare/basicConsume and stores their conjunction; heldDurable()
  returns it.
- FleetMcp.coordinatorView now reads channel.heldDurable() instead of a
  literal; updated the javadoc to say where the fact comes from.
- FakeLeadChannel gets a heldDurable field (default true) + withHeldDurable
  setter so FleetMcpTest can prove the field goes false.
- FleetMcpTest: new test asserts heldDurable:false when the channel says so.
- LeadMailboxTest (contract, real broker): new test asserts heldDurable()
  true against a real LeadMailbox. Verified by hand that flipping own()'s
  autoAck local to true turns this test (and two pre-existing redelivery
  tests) red, and restoring it turns them green again.
2026-09-10 14:31:44 +07:00
Dai Ha 4b10d02207 fleetd #425 rework round 4: mutation-pinning test for the dropped PlacementDecision
CI / contract (pull_request) Successful in 1m10s
CI / build (pull_request) Successful in 2m3s
SessionManager.acquireWithWorktree's unqualified branch must carry the
PlacementDecision it already resolved via launcher.place() into
launcher.spawn(spawnReq, decision) rather than re-deriving it through a
blank-profile launcher.spawn(spawnReq). Every existing test in this file uses
PlacementPolicies.fixed(), which answers select() the same way on every call,
so dropping the decision (handle = launcher.spawn(spawnReq);) was invisible:
186 tests stayed green under that mutation.

acquireWithWorktreeSpawnsOnTheSameProfileItProvisionedTheWorktreeForUnderARotatingPolicy
uses PlacementPolicies.roundRobin() instead — deterministic AND stateful, so
two select() calls on the same policy instance disagree (index 0 then index 1
across a two-profile pool). It asserts AGREEMENT between the profile the
worktree's parity overlay was provisioned for and the profile the member
actually spawned on, never a hardcoded expected profile name.

Verified as a real mutation, not a no-op: applying the exact mutation
(handle = launcher.spawn(spawnReq);) turns it red — expected [b.mcp.json]
but was [a.mcp.json] — and reverting turns it green again. Full build:
1572 tests, 0 failures, 0 errors, BUILD SUCCESS.
2026-09-10 14:24:14 +07:00
Dai Ha c5fbfdbf4a Merge main into #425 rework branch (brings #438 held-peer-mail read) 2026-09-10 14:11:37 +07:00
ltms 12cff28abb Merge #438: let a lead read its own held peer mail, primary-only (fleetd #421)
CI / contract (push) Successful in 1m15s
CI / build (push) Successful in 2m6s
fleet_poll{coordId} peeks this daemon's own held lead-to-lead mail and
returns full bodies without acking. New Authz.Action COORD_READ, primary
only — not the architect, which holds READ today. pollAction is now
argument-derived over both target and coordId.

heldView and HELD_PREVIEW_MAX_CHARS are untouched: fleet_list stays a
cheap always-safe scan, and the full read is a separately authorized call.

Verified by me, not taken from the report: merged tree builds 1562 green
(main 1555 + 7 new methods), 0 compile errors, no merge conflicts.
Mutation battery on lines the worker did NOT mutate, control 111 green:
  - COORD_READ widened to the architect  KILLED (AuthzTest + FleetMcpAuthzTest)
  - self-coord-id guard removed          KILLED (FleetMcpTest)
  - full body swapped for the preview    KILLED (FleetMcpTest)

The third mutation targets the ticket's own deliverable, and it is pinned.
Authz.permits has no default, so a new action is a compile error rather
than a silently unhandled case.

Follow-up filed as #439: fleet_list's coordinator row is still READ-gated,
so a worker sees peer coord-ids and 80-char previews of lead-to-lead
bodies. Pre-existing; the implementer flagged it and left it alone.
2026-09-10 09:08:17 +02:00
Dai Ha 77a6a7142e Merge main into #421 branch (brings #434 model-gate observability and #436 fixed-placement cap)
CI / contract (pull_request) Successful in 58s
CI / build (pull_request) Successful in 1m48s
2026-09-10 14:06:34 +07:00
Dai Ha 9f3671b801 fleetd #425 rework round 3: rewrite prose after #435 made fixed honour maxLoad
CI / contract (pull_request) Successful in 48s
CI / build (pull_request) Successful in 1m56s
fleetd #435 (merged to main) made FixedPlacementPolicy evaluate maxLoad during automatic
selection, the same way weighted/round-robin already did. Six comments across
CompositePeerLauncher.java, PeerLauncher.java, PlacementDecision.java, and SessionManager.java
justified round 2's place()/spawn(req, decision) mechanism by saying fixed "deliberately never
evaluates maxLoad" — that claim is now false, and needed restating, not just deleting.

The honest case after #435: the two-path shape (routing branch falls through an excluded
candidate; explicit-profile branch refuses on it) is still real and still deliberate — an
operator who names a profile should get a refusal, not a silent substitution. What round 1 got
wrong, and what round 2 still needs to prevent, is turning a fall-through into a refusal by
accident: resolving a name via place() and then feeding it back to spawn(SpawnRequest) as an
explicit profile. Before #435 that accident was reachable through maxLoad specifically, because
fixed never evaluated it; #435 closed that specific gap, so a PlacementDecision can no longer be
at-cap in the first place. What survives as the justification for spawn(req, decision): it never
re-evaluates a condition place() already decided, and it closes the window between that decision
and the spawn in which the underlying state could otherwise move — not a failure #435 already
prevents.

Re-measured the sibling paths this round exists to keep in agreement (one profile at maxLoad: 1,
liveCount pinned at 1, PlacementPolicies.fixed(), unqualified spawn): both the with-worktree and
without-worktree paths now throw the identical PlacementException — "worker profile 'a' is at
maxLoad (1 live >= 1 cap), and no available candidate remains" — closed upstream by #435, at
place()/select(), before either path ever reaches a spawn call. The observable asymmetry this PR
was filed to fix is gone; what remains is the structural argument above.

No behavior change: place()/PlacementDecision/spawn(req, decision) are untouched, and
FixedPlacementPolicy/PlacementPolicyUtil are taken wholesale from main's merge.
2026-09-10 14:06:10 +07:00
Dai Ha 84034b34d1 Merge remote-tracking branch 'origin/main' into worker/425-rework-placement-resolve-c58ba1-9 2026-09-10 13:56:54 +07:00
Dai Ha 1e9b2c9b7e fleetd #421: let a lead peek its own held peer mail, primary-only
CI / contract (pull_request) Successful in 1m27s
CI / build (pull_request) Successful in 1m36s
fleet_list truncated held lead-to-lead messages to an 80-char preview with
no way to read the full body, and fleet_poll{target} drained the wrong
inbox (a worker's reply queue, not the coordinator mailbox) -- it silently
returned []. fleet_ack would have destroyed the message unread.

Add a non-destructive read: fleet_poll{coordId} peeks (never acks) this
daemon's own held mail via LeadChannel.peek(). The coordId must equal the
caller's own selfCoordId -- passing a peer's id is refused with a reason,
instead of repeating the original silent-[] confusion.

This is authorization-sensitive: mapping it to the existing READ action
would let any worker read every peer lead's mail in full. READ's openness
rests on "the roster carries no secrets" (Authz.java), which does not hold
for lead-to-lead coordination bodies. Added Authz.Action.COORD_READ,
primary-only (not even the architect, which holds READ today), and made
pollAction's signature depend on both target and coordId so every call
site states explicitly what it passes.

Also fixes fleet_list's "pending: 0" trap: mailbox.pending only counts
broker-ready messages, so a healthy held mailbox reads as empty. Added
heldCount/heldDurable beside held[] so the durability fact isn't implied
only by reading the code.

Mutation-tested: pollAction's COORD_READ->READ mapping, the peek->ack
substitution, the 80-char preview cap widened to 81, and the Authz case
widened to include caller.isWorker() -- each breaks exactly its matching
test and nothing else. The first attempt at the preview-cap test used a
homogeneous "x"*200 body, which a widened cap slipped through unnoticed
(contains() found a shifted match); replaced with a sentinel character at
index 80 to actually pin the boundary.

Updates CLAUDE.md's intent->tool table for fleet_poll's new coordId
semantics, per this repo's own "prompt is part of the product" rule.
wiki/ is a submodule and not committable from a worker's worktree --
wiki-bound content is in the PR body instead.
2026-09-10 13:54:50 +07:00
ltms 5d422f85fa Merge #436: make fixed placement honour maxLoad (fleetd #435)
CI / contract (push) Successful in 48s
CI / build (push) Successful in 1m37s
fixed was the one automatic policy that ignored maxLoad, and it is the
default for an absent placement: key. It now gates on the same shared
atCap predicate weighted and round-robin use, and falls through to the
next candidate rather than refusing.

Verified here: 1555 tests green (main was 1548, +7 new methods), 0
compile errors. Mutation battery, control 113 green:
  - atCap boundary >= -> >        KILLED (15 tests, across all 3 policies)
  - drop cap check in fixed walk  KILLED (3 tests)
  - weightExcluded -> false       KILLED (2 pre-existing tests)

Neither live host is affected today: both run placement: weighted (Mac
fleetd.yaml:210, fleet01 fleetd.yaml:172), and weighted already skipped
at-cap candidates via PlacementPolicyUtil.available. This aligns fixed
with the other policies and with maxLoad's documented contract.
2026-09-10 08:54:18 +02:00
Dai Ha ed2027b202 fleetd #435: make FixedPlacementPolicy honor maxLoad
CI / contract (pull_request) Successful in 45s
CI / build (pull_request) Successful in 1m54s
FixedPlacementPolicy (the default placement policy) never consulted
maxLoad, so an at-cap default was chosen anyway on every unqualified
spawn -- the cap was advisory, not enforced, for the one policy every
config uses by default. weighted/round-robin already gated on it via
PlacementPolicyUtil.available().

Extract the "at cap" predicate into PlacementPolicyUtil.atCap(ctx, c)
so all three policies share one definition, and consult it at both of
FixedPlacementPolicy's filter sites (the default fast path and the
candidate walk), mirroring the existing weightExcluded pattern. An
at-cap default now falls through to the next candidate instead of
refusing the spawn -- only when every candidate is unusable does the
policy still throw, naming the cap in the message. Update the class
javadoc (five exceptions -> six) and the reason-priority comments to
match CompositePeerLauncher's explicit-spawn order (quarantine,
cooling off, max load, model-off).
2026-09-10 13:46:06 +07:00
Dai Ha 6b0a99b2b7 fleetd #425 rework round 2: stop routedProfileFor's caller re-entering the throwing branch
CI / contract (pull_request) Successful in 1m24s
CI / build (pull_request) Successful in 1m34s
Round 1 closed quarantine/cool-off/model-off routing for acquireWithWorktree by resolving the
profile through routedProfileFor(role) and handing that name back to launcher.spawn(SpawnRequest)
as an EXPLICIT profile. That re-resolution has a cost the lead measured directly: naming a
profile explicitly makes CompositePeerLauncher.spawn take its THROWING branch (enforceMaxLoad
included), while the routing branch a blank spawn takes never calls enforceMaxLoad at all, and
FixedPlacementPolicy (the default) deliberately never evaluates maxLoad during automatic
selection. So an at-cap pool-first profile that placement itself would have picked for a plain
unqualified spawn could die at enforceMaxLoad one call later, purely because the worktree path's
route to the spawn passed through an explicit profile name — a new failure a worktree-less
unqualified spawn never hits.

This closes the two-path shape instead of moving it: PeerLauncher gains place(role), returning an
opaque PlacementDecision, and spawn(req, decision), which honors that decision through the SAME
routing branch a blank spawn uses — no enforce* check is newly applied. SessionManager.
acquireWithWorktree now keeps the PlacementDecision from place() and hands it to
spawn(req, decision) for an unqualified request, instead of re-resolving through an explicit
profile name. An explicitly-named profile is unaffected: it still goes through spawn(req) and its
throwing branch, exactly as before.

Also corrects the acquireWithWorktree comment's false claim that round 1 "loses nothing else" —
maxLoad was lost too, as a new hard failure, not a retry. The comment now names it explicitly.

Kept the four round-1 tests (still pass — routedProfileFor now just delegates to place()). Added
one class asserting the invariant itself: an unqualified spawn on a maxLoad-capped profile must
land the same outcome with and without a worktree, asserting on the pair rather than a hardcoded
direction, so it stays correct however fleetd #435 (not this ticket) resolves whether maxLoad
should gate an unqualified spawn at all.
2026-09-10 13:42:06 +07:00
ltms 6b7caba248 Merge #434: make the model gate's own state observable
CI / contract (push) Successful in 47s
CI / build (push) Successful in 1m49s
Verified in the worker's tree at 7fd914d: 1544 tests green (base 1535 + 9),
0 compile errors, unpiped mvn clean install. merge-tree against f8b0d42 reports
no conflicts and the two sides share no files.

Read all four production diffs. The design is right: one ModelGateState record
carrying both "armed" and "off" from a single models0() read, so the startup
log line, fleet_profiles' modelGateArmed and the spawn gate cannot disagree —
the fleetd #404 lesson applied properly. disabledModels() now delegates to it
rather than being a second independent read.

Mutated the two subtlest lines, with a control in the same script and the
changed line echoed back with its number:

- modelGateState(): sentinel identity check replaced by the naive
  "m.offIds().isEmpty()" -> 2 failures. FleetProfilesModelGateStateTest
  .modelsBlockWithNothingOffReportsGateArmedAndZeroOff:86 and
  CompositePeerLauncherTest.modelGateStateIsHotReloadedThroughARealConfigRef
  :1519, both "expected: <true> but was: <false>". So the one state this
  ticket exists to expose — a block present with nothing off — is pinned.
- notConfigured() returning a non-empty off set, breaking the invariant the
  record's javadoc states but does not enforce -> 3 failures, including
  disabledModelsIsEmptyWithNoModelsConfigured:1581. So the invariant is
  observable even though the constructor does not check it.

Control run unmutated: 1544 green.

Follow-up on me, not a merge blocker: fleet_profiles gains an operator-visible
field, so this needs a wiki/11-Features.md entry. Workers cannot commit the
wiki submodule, so I am adding it.
2026-09-10 08:32:54 +02:00
Dai Ha f8b0d42a5c fleetd #431 follow-up: profileForSlot's javadoc named a caller that does not exist
CI / contract (push) Successful in 1m28s
CI / build (push) Failing after 1m53s
The javadoc said "what the spawn lifecycle reads". Nothing in src/main calls
profileForSlot at all, in either the ".profileForSlot(" or the
"::profileForSlot" form. My own #431 ticket text repeated that sentence as a
fact and ranked the three accessors by it, and the #432 worker copied it into
the test file's comment and one assertion message. Corrected in all three
places; the ticket correction is posted on #431.

What the spawn lifecycle actually reads for an architect's profile is the
SlotReservation that reserve() returns — SessionManager.java:225,
"reservation == null ? profile : reservation.profile()".

The corrected ranking, measured rather than read off the javadoc:
- nameForSlot is wired, at CallerResolver.java:137 (method reference, which is
  why a ".nameForSlot(" grep missed it)
- isSlot is reached through bind, called at MemberRegistry.java:235 and :376
- profileForSlot has no caller at all

Prose only. No behaviour change.
2026-09-10 13:25:19 +07:00
ltms d1e7d71eee Merge #432: pin profileForSlot, isSlot and nameForSlot against a live reload
CI / contract (push) Successful in 1m29s
CI / build (push) Successful in 1m34s
Verified in my own tree at a196d34: 1539 tests green, 0 compile errors, 0 files
changed under src/main (test-only, as reported).

Mutated two halves the worker's own proof did not cover, with a control in the
same script and the changed line echoed back each time:

- roleForSlot returning ARCHITECT for any configured slot (flatten is
  role-blind) -> 2 failures, incl. CallerResolverTest
  .aBoundNonArchitectSlotStillResolvesAsAWorker:454. So the role-blind flatten
  cannot grant ARCHITECT through a non-architect pool; that was already pinned.
- nameForSlot parsing the key suffix instead of reading config -> 1 failure,
  MemberRegistryLiveTest.nameForSlotReflectsANameChangedByReload:255. The new
  test has teeth beyond the freeze the worker ran.

Control run unmutated: green.

The ticket's severity ranking was wrong and I corrected it on #431. profileForSlot
has no caller in src/main in either the "." or "::" form, so its javadoc ("what
the spawn lifecycle reads") names a caller that does not exist; nameForSlot is
wired at CallerResolver.java:137; isSlot is reached through bind, at :235 and
:376. A follow-up commit fixes the test prose that repeated my claim.
2026-09-10 08:24:26 +02:00
Dai Ha 7fd914df1a fleetd #422 follow-up: make the model gate's own state observable
CI / contract (pull_request) Successful in 45s
CI / build (pull_request) Successful in 1m53s
PeerLauncher.disabledModels() reported an empty set both when no
models: block exists and when a block exists with nothing off, so
fleet_profiles/GET /profiles and the startup log could not tell an
inert gate from an armed one reporting zero. Add
PeerLauncher.ModelGateState (configured + off), a modelGateState()
default method disabledModels() now delegates to, and a
CompositePeerLauncher override that reads models0() once and
distinguishes the null-supplier case (no models: block) from a real,
config-supplied block via identity against the NO_MODELS_CONFIGURED
sentinel — reusing the exact accessor the spawn gate itself reads, per
the fleetd #404 lesson.

Wires the state into a new startup log line (Fleetd.modelGateCoverageLine)
and a new modelGateArmed field in FleetMcp.profilesView, reported
unconditionally alongside the existing modelsOff set.
2026-09-10 13:18:52 +07:00
Dai Ha b066eb1903 fleetd #425 rework: resolve acquireWithWorktree through real placement, not a blind pool-first read
CI / contract (pull_request) Successful in 1m14s
CI / build (pull_request) Successful in 2m13s
e1d7dde (PR #430) kept two good fixes and one regressed one. Kept: (1)
CompositePeerLauncher.defaultProfile() delegating to
defaultProfileFor(MemberRole.DEV) so fleet_profiles' "default" tracks a live
reload, and (2) PeerLauncher.defaultProfileFor(MemberRole). Redone:
acquireWithWorktree's profile pre-resolution.

The regression: acquireWithWorktree pre-resolved via
launcher.defaultProfileFor(memberRole), which just returns the role pool's
FIRST entry, blind to quarantine/cool-off/model-off. That name was then
passed to launcher.spawn as an EXPLICIT profile, which takes
CompositePeerLauncher.spawn's THROWING branch (enforceNotQuarantined /
enforceMaxLoad / enforceModelEnabled) instead of the ROUTING branch a blank
profile gets. So a quarantined or model-off pool-first profile turned a
routine unqualified spawn into a hard PlacementException -- undermining
fleetd #429's "the fleet keeps working when a model is turned off"
guarantee for every worktree spawn.

Fix: add PeerLauncher.routedProfileFor(MemberRole), the profile an
unqualified spawn of that role would actually be routed to right now --
same candidate list, same quarantined/coolingOff/modelOff filtering, same
PlacementPolicy spawn() itself consults. CompositePeerLauncher implements it
by extracting spawn()'s context-building into a shared private
placementContextFor(role, unreachable), so spawn() and routedProfileFor()
can never disagree about which conditions apply to which candidate.
acquireWithWorktree now calls routedProfileFor once and reuses that name for
repoRoot, parityOverlay, and the spawn -- the fleetd #425 defect (the three
disagreeing) stays fixed, now on the routed path instead of the blind one.

An explicit profile named by the caller is untouched -- it still hits the
throwing branch, which is correct for an operator override.

Trade-off carried over from e1d7dde, now precisely scoped: an unqualified
worktree spawn still loses CompositePeerLauncher's cross-candidate retry on
a live PeerUnreachableException (a transport failure at spawn time, which
placement cannot see in advance) -- but NOT the quarantine/cool-off/
model-off routing, which routedProfileFor already resolved before spawn
ever runs. Accepted: a worktree provisioned for the wrong backend is worse
than a spawn that fails cleanly and can be retried by the caller.

Tests: CompositePeerLauncherTest gains
routedProfileForSkipsAQuarantinedPoolFirstProfileUnderFixedPolicy and
...ModelOff..., both under PlacementPolicies.fixed() (the default policy,
not weighted() -- the previous round's tests all used weighted() and never
exercised FixedPlacementPolicy's own inline filter, which is exactly what
regressed). SessionManagerTest gains
acquireWithWorktreeRoutesAroundAQuarantinedPoolFirstProfile, proving
repoRoot/parityOverlay/spawn agree on the ROUTED profile, not just the
pool-reordered-by-reload one the existing #425 tests already covered.
2026-09-10 13:10:30 +07:00
Dai Ha a196d34455 fleetd #431: pin profileForSlot, isSlot and nameForSlot against a live reload
CI / contract (pull_request) Successful in 49s
CI / build (pull_request) Successful in 2m7s
#424 made MemberRegistry.slots() re-read fleet: on every call, but only
roleForSlot was tested against a real reload. profileForSlot, isSlot and
nameForSlot all have the same live-read line and none was pinned — proved by
freezing each to a construction-time snapshot and watching the full suite
stay green.

Adds 4 tests to MemberRegistryLiveTest, each driving a real ConfigRef.reload()
against a @TempDir config file (never two frozen registries compared in
memory, which would test the constructor instead of the reload):
- profileForSlotReflectsAProfileChangedByReload
- isSlotStopsReportingASlotRemovedByReload / isSlotStartsReportingASlotAddedByReload
- nameForSlotReflectsANameChangedByReload

No production change. profileForSlot has no call site anywhere in src/main
yet, so there is no spawn-lifecycle seam to drive the test through beyond the
accessor itself.
2026-09-10 13:06:56 +07:00
Dai Ha 051d320ea0 fleetd #425: fleet_profiles' default and worktree provisioning must read live placement
fleet_profiles' "default" was CompositePeerLauncher.defaultProfile, a value
frozen at construction from cfg.effectiveDefaultProfile(). An unqualified
fleet_spawn instead resolves the dev pool live via defaultProfileFor(DEV) on
every call, so reordering fleet.developers and reloading changed where a
spawn landed without ever changing what fleet_profiles reported.

- CompositePeerLauncher.defaultProfile() now delegates to
  defaultProfileFor(MemberRole.DEV) -- the same live, reload-aware pool read
  placement already uses -- falling back to the frozen field only when no
  profiles are configured at all.
- PeerLauncher gains a default defaultProfileFor(MemberRole) method so a
  generic PeerLauncher reference can ask for a role's live default; the
  default implementation delegates to defaultProfile() for launchers with no
  pool concept of their own.
- SessionManager.acquireWithWorktree resolved a profile via
  launcher.defaultProfile() (DEV-only) to provision repoRoot/parityOverlay,
  then spawned with the original (possibly blank) profile, which re-resolves
  independently through placement -- for any non-DEV role, or across a config
  reload between the two reads, the two resolutions could disagree and
  provision a worktree for a profile the member never runs on. Fixed by
  resolving once, through defaultProfileFor(the caller's actual role), and
  reusing that same resolved name for repoRoot, parityOverlay, and the spawn
  itself. Trade-off: this path now spawns with an explicit profile rather
  than a blank one, so it loses CompositePeerLauncher's cross-candidate retry
  on PeerUnreachableException -- accepted because a worktree provisioned for
  the wrong backend is worse than a spawn that fails cleanly and can be
  retried.

Tests: CompositePeerLauncherTest (live dev-pool reorder + empty-pool
fallback), FleetProfilesLiveDefaultTest (drives FleetMcp.profilesView
directly), SessionManagerTest (worktree overlay follows a reorder, and a
non-DEV role's worktree spawn uses that role's pool, not DEV's).
2026-09-10 12:58:20 +07:00
Dai Ha 766772763f Merge #428: revoke the ARCHITECT privilege on reload, not just future spawns
CI / contract (push) Successful in 1m27s
CI / build (push) Successful in 1m32s
fleetd #424. MemberRegistry.slots() now re-reads fleet: through a supplier,
so removing an architect slot demotes the bound pane on its very next request.
The boundEntries cache and entryFor() fallback from the first round are gone:
the slot OCCUPANCY (terminalToSlot) survives a reload, the ARCHITECT role does
not. That split was my own ticket wording's fault -- I asked for a test that a
bound architect "survives the rebuild", which conflated the binding with the
privilege.

Conflict resolved by hand in ConfigRef.java: #422 (models:) and #424
(architects) both rewrote the same Hot bullet. Kept both.

Also corrected two claims #424's own second commit left stale -- 7f672f0
reversed the behaviour but never touched ConfigRef, whose whole job is to tell
the operator what a reload does:
  - the Hot bullet said MemberRegistry's rule "keeps a live session's identity
    even after its slot is removed from config"
  - the reload-report comment said "only a NEW bind is refused"
Both now say what the code does: removal revokes ARCHITECT on the next
request, and only the slot occupancy survives.

Verified by the lead: 1535 tests, 0 failures, 0 compile errors, BUILD SUCCESS
on the merged tree.

Mutation of three halves the worker's own proof did not cover -- profileForSlot,
nameForSlot and isSlot each pointed at a frozen snapshot taken at construction
(live readers 5 -> 4, each mutation naming its method and line). All three
PASSED at 1535. The ticket's own fix is well pinned; these three sibling live
reads are not. Follow-up filed.
2026-09-10 12:54:35 +07:00
ltms eab8185d7b Merge #429: enforce the models.allow on/off gate at spawn, hot — including under fixed placement
CI / contract (push) Successful in 1m29s
CI / build (push) Successful in 1m33s
fleetd #422 + follow-up. Verified by the lead: 1524 tests green, 0 compile errors.

Mutation proof of three halves the worker's own proof did not cover:
- modelOffProfiles() -> Set.of() (the feed into ctx.modelOff): kills 3, incl. fixedPlacementSkipsAnOffModelProfileToo
- enforceModelEnabled() -> no-op (the explicit-spawn gate): kills 3, incl. the real-ConfigRef hot-reload test
- disabledModels() -> Set.of() (the reporting accessor): kills 1, alone
2026-09-10 07:44:35 +02:00
Dai Ha 7f672f0fb8 fleetd #424: revoke the ARCHITECT privilege on reload, not just future spawns
CI / contract (pull_request) Successful in 43s
CI / build (pull_request) Successful in 2m3s
Correction to the #424 fix in PR #428: the ticket asked to revoke a removed
architect slot, but the previous change (boundEntries) kept BOTH the binding
and the ARCHITECT privilege alive for an already-bound session after its slot
left config. That left the ticket's actual headline defect half-open.

The corrected rule: config governs both what may be bound next AND what a
bound slot still grants. Removing a slot now demotes its bound session to
worker on the very next request (roleForSlot/nameForSlot read slots() with no
cache, so CallerResolver.resolve falls through to Principal.worker(...)). The
terminalToSlot binding itself is untouched by a reload, on purpose: dropping
it would double-book the slot key and break unbind's compare-safe contract.

- Delete boundEntries and entryFor; profileForSlot/roleForSlot/nameForSlot/
  isSlot all read slots() directly, live, with no cache.
- Rewrite the class doc's binding rule for the corrected semantic.
- Replace the old "survives removal" test with anArchitectAlreadyBoundToASlotIsDemotedByReload,
  asserted through a real CallerResolver.resolve (not the roleForSlot seam),
  plus two tests for what must NOT change: the binding still occupies the
  slot after removal (a second terminal cannot claim it, even once the slot
  returns to config), and unbind still succeeds for the original terminal.

Verified snapshot()/CallerResolver.members() need no change: snapshot() only
ever reported raw terminalToSlot occupancy, and CallerResolver.members() has
no production caller.
2026-09-10 12:29:35 +07:00
Dai Ha e2801b9bbc fleetd #422 follow-up: gate FixedPlacementPolicy on modelOff too
CI / contract (pull_request) Successful in 44s
CI / build (pull_request) Successful in 1m53s
FixedPlacementPolicy is the DEFAULT placement policy (PlacementPolicies.fromName
returns it for an absent/blank name) and it built its own inline candidate
filter instead of calling PlacementPolicyUtil.available(). That filter checked
quarantined/coolingOff/unreachable/excluded() but never modelOff(), so an
unqualified fleet_spawn on any fleet without an explicit placement: policy
could still land on a profile whose model the operator turned off.

- Add ctx.modelOff() to both filter sites: the default-profile fast path and
  the fallback walk over ctx.candidates().
- Add a modelOff refusal reason to the default-profile reasons list, worded as
  an operator decision ("turned off in models.allow"), matching
  enforceModelEnabled. Quarantine and cooling off still take priority when a
  profile is also model-off, matching CompositePeerLauncher's explicit-spawn
  check order.
- Update the two stale "excluded from automatic selection" messages to name
  model-off, consistent with PlacementPolicyUtil.emptyException.
- Update the class javadoc: four exceptions -> five, with a new bullet for
  model-off (fleetd #422).

Tests: PlacementPolicyTest gains fixedSkipsModelOffDefault (fast-path),
fixedFallbackWalkSkipsModelOffCandidate (fallback walk),
fixedThrowsWhenDefaultAndEveryCandidateModelOff (all-off refusal wording), and
fixedReportsQuarantineNotModelOffWhenBothApply (priority). CompositePeerLauncherTest
gains fixedPlacementSkipsAnOffModelProfileToo, an integration-level mirror of
the existing placementSkipsAnOffModelProfileAndRoutesToAnotherOne but under
PlacementPolicies.fixed(). The two existing weighted()-based tests are
untouched.
2026-09-10 12:25:31 +07:00
Dai Ha ea02c7b248 fleetd #422: enforce the model allow-list on/off state at spawn, hot
CI / contract (pull_request) Successful in 1m28s
CI / build (pull_request) Successful in 1m30s
Ships the two halves left out of the earlier allow-list ticket in one PR,
since apart they are inert: a gate with no flag always allows, and a flag
nothing reads does nothing.

- FleetConfig.Models.ModelEntry gains `enabled` (default on; absent/true =
  on, false = off). Turning a model off never removes it from `allow:` —
  validateModels() checks membership only, so an off model stays valid
  config and a still-configured profile naming it does not refuse reload.
  Models.offIds() is the one live accessor both the gate and the status
  report read.
- CompositePeerLauncher.enforceModelEnabled is a FOURTH, independent
  spawn-refusal reason (operator intent) — never layered onto
  BackendQuarantine/BackendOutagePolicy, which are backend-reported outage.
  Wired into the explicit-profile branch. modelOffProfiles() feeds the same
  off-model exclusion into PlacementContext for unqualified spawns via
  PlacementPolicyUtil (a new modelOff set, counted into its own bucket in
  emptyException so "all off" is named as the cause, not generic).
  Both read models0(), a live Supplier<FleetConfig.Models>, so a reload
  reaches the very next spawn — no restart.
- PeerLauncher.disabledModels() (default empty) lets fleet_profiles/
  GET /profiles report off models by reading the exact same accessor the
  gate reads (the fleetd #404 lesson: a status field must read the source
  the behaviour reads).
- ConfigRef: `models` reclassified from deferred to hot-excluded — nothing
  about it is baked into a startup-built object anymore; membership is
  re-validated in full on every reload via validateAll(), and the on/off
  half is read live everywhere. Tally: 5 cold, 13 deferred, 3 split, 4
  hot-excluded (25 total). ConfigRefTopLevelCoverageTest and
  ConfigRefTopLevelReportingCoverageTest updated with no new exclusion
  added just to force green.

Tests: FleetConfigTest (old-style fixture stays on; an off model is still
valid config; one model disables every profile naming it),
CompositePeerLauncherTest (explicit refusal wording distinct from
quarantine/cool-off; unqualified spawn skips an off candidate and names
model-off when every candidate is off; a real ConfigRef.reload() proves
the hot path; disabledModels() matches the gate).
2026-09-10 12:08:18 +07:00
Dai Ha ce74e164c6 fleetd #424: make architect-slot identity checks read fleet.architects live
CI / contract (pull_request) Successful in 1m15s
CI / build (pull_request) Successful in 1m32s
MemberRegistry used to flatten fleet.architects into an unmodifiable map at
construction, so removing (revoking) an architect slot from config never
took effect: reserve()/requireSlotFor() kept granting spawns against the
frozen snapshot forever, while ConfigRef told the operator "already
applied" for the wrong consumer.

- MemberRegistry gains a live constructor (MemberRegistry.live(Supplier))
  that re-flattens fleet.architects/developers/reviewers on every
  slots()/slotsFor() call, so reserve() and requireSlotFor() (which both
  read through slotsFor) govern the NEXT spawn with no restart. The frozen
  single-arg constructor is kept for tests and fixed/code-built configs.
- Binding rule: config governs what may be bound next; it never
  retroactively unbinds a live session. A slot removed from config while a
  terminal is bound to it keeps that binding. To keep the bound terminal's
  IDENTITY too (CallerResolver.resolve reads roleForSlot/nameForSlot on
  every request), every successful bind now caches the slot's Entry into a
  new boundEntries map; roleForSlot/nameForSlot/profileForSlot/isSlot fall
  back to it when the slot is no longer live, and unbind clears it in the
  same critical section it clears the binding.
- Fleetd.java now wires MemberRegistry.live(() -> config.get().fleet())
  instead of the frozen constructor.
- ConfigRef: corrected the fleet.leaders split-key message and the class
  doc's Hot bullet — architects is now hot for two independent consumers
  (CompositePeerLauncher for placement, MemberRegistry for identity), not
  only the one the message used to name. An architects-only edit still
  reports nothing beyond "config reloaded", which is now honest since the
  key really is fully hot for both consumers.

Added MemberRegistryLiveTest: real ConfigRef.reload() against a @TempDir
file, both directions (slot removed / slot added) for requireSlotFor and
reserve tested separately, plus a bound-architect-survives-removal test
that checks the binding AND the identity (roleForSlot/nameForSlot).

Mutation-tested: reverting requireSlotFor to a frozen snapshot fails
requireSlotForRefusesAProfileWhoseSlotWasRemovedByReload and its mirror;
reverting reserve the same way fails the two reserve tests; dropping the
boundEntries fallback fails the survives-removal test's roleForSlot
assertion. All three restored before commit.

mvn clean install: Tests run: 1510, Failures: 0, Errors: 0, Skipped: 0 —
BUILD SUCCESS.
2026-09-10 12:05:18 +07:00
ltms e60f892efd Merge pull request 'fleetd #415: split coverage() feature-state wording by pattern fallback semantics' (#423) from worker/415-coverage-wording-2cbf9c-5 into main
CI / contract (push) Successful in 1m13s
CI / build (push) Successful in 1m37s
2026-09-10 06:57:33 +02:00
Dai Ha ce05886831 fleetd #415: pin which UnsetMeaning Fleetd pairs with which pattern key
CI / contract (pull_request) Successful in 50s
CI / build (pull_request) Successful in 1m54s
Review found a gap: the earlier tests all called CompletionResolver.coverage()
directly, supplying the UnsetMeaning themselves — proving the enum's wording,
never that Fleetd's two call sites pair the right meaning with the right key.
Swapping the two UnsetMeaning arguments at those call sites (recreating #415's
defect with exhaustedPattern and errorPattern exchanged) compiled with 0 errors
and left all 1506 tests green.

Extract the two coverage-line call sites out of main() into package-private
static factories (Fleetd.exhaustedPatternCoverageLine /
errorPatternCoverageLine), the same pattern already used for capacitySource
and worktreeBranchLookup. Add FleetdPatternCoverageLineTest, which calls both
factories directly and asserts the actual wording each produces for the same
empty-coverage input, including that the two differ.

Also recorded the swap-mutation measurement (0 errors, 1506 green) in
UnsetMeaning's javadoc so a future reader does not delete the new test as
redundant with CompletionResolverTest.
2026-09-10 11:54:02 +07:00
Dai Ha be123d0ac7 fleetd #415: split coverage() feature-state wording by pattern-key fallback semantics
CI / contract (pull_request) Successful in 46s
CI / build (pull_request) Successful in 2m4s
CompletionResolver.coverage() measured pattern coverage (how many profiles set
a key) but its 'off' wording read as feature state. That is false for
errorPattern: an unset errorPattern still runs the classification against the
built-in BACKEND_ERROR pattern (CompletionResolver.java:84), so the empty case
is not off.

Add CompletionResolver.UnsetMeaning (OFF / BUILT_IN_DEFAULT), a required
parameter every coverage() call must supply — no defaulted overload, so a
future third pattern key cannot compile without stating what unset means for
it. Fleetd.java now passes UnsetMeaning.OFF for exhaustedPattern (no fallback
exists) and UnsetMeaning.BUILT_IN_DEFAULT for errorPattern.

Tests: updated the three existing empty/full/partial cases to pass the new
parameter, corrected the one test that pinned the old (wrong) errorPattern
wording, and added a test that asserts the same empty-coverage input produces
different wording for the two keys.
2026-09-10 11:40:32 +07:00
ltms 2d09c8b027 Merge pull request 'fleetd #418: barrier the throw-path push-loop test on state decide() reads' (#419) from worker/418-588283-3 into main
CI / contract (push) Successful in 49s
CI / build (push) Successful in 1m39s
2026-09-10 06:29:23 +02:00
ltms ed54f0224e Merge pull request 'fleetd #416: fleet_list must enumerate the STARTUP profile set' (#420) from worker/416-3ad1da-1 into main
CI / contract (push) Successful in 48s
CI / build (push) Successful in 1m34s
2026-09-10 06:27:36 +02:00
Dai Ha 8d5bc3ee89 fleetd #416: fleet_list must enumerate the STARTUP profile set
CI / contract (pull_request) Successful in 1m29s
CI / build (pull_request) Successful in 1m52s
`profiles` is a DEFERRED key: HerdrPeerLauncher takes Map.copyOf(profiles) once at
construction, so a profile added only to the hot-reloaded map can never be spawned.
CapacitySource was built with `() -> config.get().profiles().keySet()` — the live map —
so fleet_list reported a hot-added profile as free while fleet_spawn on that same
profile failed with "unknown worker profile". fleet_list's contract for `free` says it
runs "the same check the spawn gate itself runs"; it did not.

The set now comes from the startup snapshot, the same shape as the coordinator.peers
wiring three lines below, whose comment already stated the rule. maxLoad stays live on
purpose — ConfigRef documents it as hot, like credentialId and weight — so a maxLoad
edit still takes effect without a restart.

Found by the fleet01 lead, who proved it with a live ghost profile rather than an
argument. Implemented by a sonnet member; its reply was lost to an empty scrape, so the
work was salvaged uncommitted from its worktree and both proof steps were run by the
lead instead:

  full build                          Tests run: 1505, Failures: 0, 0 compile errors
  mutation A: live keySet restored    liveOnlyProfileIsNotListed FAILS (alone)
  mutation B: permanently empty set   startupProfileIsListed FAILS (alone)

Mutation B is the point of the second direction: per #404, a test that only ever checks
the absent case cannot tell a correct lookup from one that returns nothing at all.
2026-09-10 11:25:56 +07:00
40 changed files with 3744 additions and 200 deletions
+10 -4
View File
@@ -87,12 +87,18 @@ jobs:
apt-get update && apt-get install -y --no-install-recommends maven
mvn -version
# The `contract` profile clears the default-excludes group, so the @Tag("contract") AMQP test
# runs against the RabbitMQ service container (AMQP_URI). Pinned to the one contract test to
# avoid re-running the unit suite already covered by the `build` job.
# The `contract` profile clears the default-excludes group, so `-Dgroups=contract` runs every
# @Tag("contract") test and nothing from the unit suite the `build` job already covered — a
# tag selects the whole group, so a test added to it later runs here automatically. A prior
# version of this step pinned `-Dtest=AmqpReplyInboxContractTest` by class name instead: that
# silently excluded every other contract test (including the herdr ones) from CI, and nobody
# noticed until the herdr protocol drifted out from under a test that never ran here
# (fleetd #449). If this runner has no herdr socket, the herdr-backed tests in the group
# skip on their own `assumeTrue` and only the broker-backed ones actually run — check the
# step output rather than assuming which.
- name: Contract tests
working-directory: fleetd
run: mvn -B -Pcontract test -Dtest=AmqpReplyInboxContractTest
run: mvn -B -Pcontract test -Dgroups=contract
- name: Failing test output
if: failure()
+1
View File
@@ -138,6 +138,7 @@ the merge — and merging on a reviewer's word is delegating it by proxy.
| Message a **peer lead** on this host | `fleet_send{sessionId: <their terminal>, content}` — `fleet_list` → `leads` reports it. Coordination only, **never** a task |
| Message a **peer lead** on another daemon or host | `fleet_send{coordId: <their coord-id>, content}` — needs a `coordinator:` block; your own coord-id is in `fleet_list`. Coordination only, **never** a task |
| Answer a peer lead that messaged you | `fleet_send{coordId}` — or `{sessionId}` if they are on this host. **Not** `fleet_reply`: it has no peer route and the publish is refused |
| Read your own held lead-to-lead mail (no ack) | `fleet_poll{coordId: <your own coord-id, from fleet_list's coordinator.selfId>}` — primary-only; never acks, so `fleet_list`'s `held[]` still shows it after. `fleet_list`'s `held[]` gives only a truncated preview — this is the only way to read the full body |
| Collect a held reply | `fleet_poll{target}` · then `fleet_ack{target, msgId}` |
| Tear down a member | `fleet_stop{paneId}` |
+104 -10
View File
@@ -69,6 +69,7 @@ import java.util.List;
import java.util.Map;
import java.util.Optional;
import java.util.Set;
import java.util.TreeSet;
import java.util.concurrent.Executors;
import java.util.concurrent.ScheduledExecutorService;
import java.util.concurrent.TimeUnit;
@@ -231,6 +232,13 @@ public final class Fleetd {
profileName -> liveCountRef.get().apply(profileName),
quarantine,
outagePolicy);
// fleetd #422 follow-up: say which of the three model-gate states the daemon booted into —
// no models: block at all, a block armed with nothing off, or a block with N off — the same
// way exhaustedPatternCoverageLine/errorPatternCoverageLine report CB-578 stage A/fleetd
// #201 Unit 5 coverage just below. Read from workers.modelGateState() (never a separate
// config.get().models() here) so this line and fleet_profiles' modelGateArmed can never
// disagree about what CompositePeerLauncher's spawn gate actually enforces.
log.info("model gate (fleetd #422): {}", modelGateCoverageLine(workers.modelGateState()));
// CB-504: under supervision (launchd/systemd) fleetd can start before herdr's socket
// exists. The client itself is lazy — it connects per call — but the orphan reap below is
// the first thing that actually talks to herdr, so without this wait a boot-order race
@@ -352,7 +360,11 @@ public final class Fleetd {
// pane resolves to an architect until the later spawn lifecycle binds one. The registry is
// what CallerResolver resolves against and what that lifecycle will read profiles from;
// nothing here spawns a slot.
MemberRegistry members = new MemberRegistry(cfg.fleet());
// fleetd #424: MemberRegistry.live re-reads fleet.architects through `config` on every
// reserve/requireSlotFor call, so a reload that removes or adds an architect slot governs
// the next spawn with no restart — the frozen `new MemberRegistry(cfg.fleet())` this used
// to be let a "revoked" slot keep granting new architect spawns forever.
MemberRegistry members = MemberRegistry.live(() -> config.get().fleet());
sessions.setMemberLifecycle(members);
if (!members.slots().isEmpty()) {
log.info("member slots: {} configured {} — none bound yet (a slot is idle until the "
@@ -380,8 +392,7 @@ public final class Fleetd {
.map(session -> exhaustedPatternsByProfile.get(session.profile()))
.orElse(null);
log.info("backend-exhausted classification (CB-578 stage A): {}",
CompletionResolver.coverage("exhaustedPattern", cfg.profiles().keySet(),
exhaustedPatternsByProfile.keySet()));
exhaustedPatternCoverageLine(cfg.profiles().keySet(), exhaustedPatternsByProfile.keySet()));
// fleetd #201 Unit 5: classify a completion-fallback scrape that matches a profile's
// configured backend-error refusal (a credential outage, a provider 5xx) as a backend error
// rather than handing it back as a real answer. Compiled once at startup, keyed by profile
@@ -401,8 +412,7 @@ public final class Fleetd {
BackendErrorPatternLookup backendErrorPatterns = backendErrorPatternLookup(sessions::roster,
errorPatternsByProfile);
log.info("backend-error classification (fleetd #201 Unit 5): {}",
CompletionResolver.coverage("errorPattern", cfg.profiles().keySet(),
errorPatternsByProfile.keySet()));
errorPatternCoverageLine(cfg.profiles().keySet(), errorPatternsByProfile.keySet()));
// CB-578 stage B: on a classification that actually wins, quarantine the exhausted profile's
// CREDENTIAL — not the profile name — so a profile sharing that credential (e.g. two models
// on one OpenAI account) is refused too, not just the one that happened to report it. Reads
@@ -670,11 +680,8 @@ public final class Fleetd {
}, outagePolicy);
FleetMcp mcp = new FleetMcp(messages, workers, sessions, identity, presence,
primaryRegistry, callers, metrics, new FleetMcp.CapacitySource(profile -> liveCountRef.get().apply(profile),
profile -> {
var configured = config.get().profiles().get(profile);
return configured == null ? null : configured.maxLoad();
}, () -> config.get().profiles().keySet(), System::nanoTime),
primaryRegistry, callers, metrics,
capacitySource(config, cfg, profile -> liveCountRef.get().apply(profile)),
new FleetMcp.HealthCoverageSource(() -> {
var health = config.get().health();
return FleetHealthMonitor.coverage(health != null && health.isEnabled(),
@@ -810,6 +817,93 @@ public final class Fleetd {
}, quarantine, profile -> startupExhaustedPatterns.containsKey(profile));
}
/**
* fleetd #415 (review follow-up): package-private factory for the CB-578 stage A {@code
* exhaustedPattern} startup coverage line, paired explicitly with {@link
* CompletionResolver.UnsetMeaning#OFF} — {@code exhaustedPattern} has no fallback, so a
* profile with none configured really does have the classification off.
*
* <p>Extracted out of {@code main} for the same reason {@link #capacitySource} and {@link
* #worktreeBranchLookup} were: {@code coverage()}'s own tests ({@code CompletionResolverTest})
* prove it words {@code OFF} and {@link CompletionResolver.UnsetMeaning#BUILT_IN_DEFAULT}
* correctly when a test supplies the meaning itself — they cannot prove {@code main} pairs the
* right meaning with the right key, which is the actual fleetd #415 defect. <b>Measured:</b>
* swapping the {@code UnsetMeaning} arguments between this method and {@link
* #errorPatternCoverageLine} — recreating #415's defect with the two keys exchanged — compiled
* with 0 errors and left all 1506 existing tests green before {@code
* FleetdPatternCoverageLineTest} was added to catch exactly that swap.
*/
static String exhaustedPatternCoverageLine(Set<String> allProfiles, Set<String> configuredProfiles) {
return CompletionResolver.coverage("exhaustedPattern", CompletionResolver.UnsetMeaning.OFF,
allProfiles, configuredProfiles);
}
/**
* fleetd #415 (review follow-up): the {@code errorPattern} counterpart of {@link
* #exhaustedPatternCoverageLine}, paired explicitly with {@link
* CompletionResolver.UnsetMeaning#BUILT_IN_DEFAULT} — an unset {@code errorPattern} still runs
* backend-error classification against {@code CompletionResolver}'s built-in {@code
* BACKEND_ERROR} pattern, so the empty case is not "off". See {@link
* #exhaustedPatternCoverageLine}'s javadoc for the measured swap mutation this pairing guards
* against.
*/
static String errorPatternCoverageLine(Set<String> allProfiles, Set<String> configuredProfiles) {
return CompletionResolver.coverage("errorPattern", CompletionResolver.UnsetMeaning.BUILT_IN_DEFAULT,
allProfiles, configuredProfiles);
}
/**
* fleetd #422 follow-up: package-private factory for the startup line reporting which of the
* three central {@code models.allow:} gate states the daemon booted into. Extracted the same
* way {@link #exhaustedPatternCoverageLine}/{@link #errorPatternCoverageLine} are, so a
* dedicated test can call it directly rather than parsing log output, and so {@code main}'s
* only source for this line is {@link PeerLauncher#modelGateState()} — never a second,
* independently-derived read of {@code cfg.models()} that could disagree with what {@code
* CompositePeerLauncher}'s spawn gate actually enforces (the fleetd #404 lesson).
*
* <p>Unlike the two pattern-key lines above, there is no {@code UnsetMeaning} choice to make
* here: {@link PeerLauncher.ModelGateState#configured()} already states, unambiguously, whether
* an empty {@link PeerLauncher.ModelGateState#off()} means "no {@code models:} block to gate
* with" or "a block armed and currently reporting zero off" — the exact two states a bare
* {@code disabledModels()} read could not tell apart before this ticket.
*/
static String modelGateCoverageLine(PeerLauncher.ModelGateState state) {
if (!state.configured()) {
return "not configured (no models: block — nothing is gated, and nothing can be)";
}
Set<String> off = state.off();
return off.isEmpty()
? "armed (models: block present; 0 models currently turned off)"
: "armed (" + off.size() + " model(s) turned off: " + new TreeSet<>(off) + ")";
}
/**
* fleetd #416: production source for {@code fleet_list}'s per-profile capacity facts.
*
* <p>The profile <em>set</em> ({@code configuredProfiles}) must come from {@code cfg} — the
* startup snapshot — not the live {@code config.get()}. {@code profiles} as a whole is a
* {@code DEFERRED} key ({@link ConfigRef#DEFERRED_KEYS}): {@code HerdrPeerLauncher} takes
* {@code Map.copyOf(profiles)} once at construction and a profile only added to the
* hot-reloaded map can never actually be spawned, so enumerating it live made {@code fleet_list}
* report a profile as available when {@code fleet_spawn} on that same profile fails with
* {@code unknown worker profile}. {@code fleet_list}'s own contract for {@code free} is "the
* same check the spawn gate itself runs" — the set the spawn gate can see is the startup one,
* so this must enumerate that one too, the same shape as {@code coordinator.peers} above.
*
* <p>{@code maxLoad} stays live on purpose: it is read off {@code config.get()} exactly like
* {@code credentialId} ({@link ConfigRef} documents both as hot), so an existing profile's
* {@code maxLoad} edit must still change what {@code fleet_list} reports without a restart.
*/
static FleetMcp.CapacitySource capacitySource(ConfigRef config, FleetConfig cfg,
Function<String, Integer> liveCount) {
return new FleetMcp.CapacitySource(liveCount,
profile -> {
var configured = config.get().profiles().get(profile);
return configured == null ? null : configured.maxLoad();
},
cfg.profiles()::keySet, System::nanoTime);
}
/**
* fleetd #248: package-private factory for the member worktree/branch lookup {@link
* CompletionResolver} uses to name a fallback report's worktree and branch (fleetd#241).
@@ -30,6 +30,16 @@ public final class Authz {
DRAIN,
/** Read-only observation: status, roster, profiles, task polling. */
READ,
/**
* Read (never ack) this daemon's own held lead-to-lead coordination mail (fleetd #421).
*
* <p>Deliberately <strong>not</strong> folded into {@link #READ}. {@code READ}'s grant
* rests on "the roster carries no secrets" (see its case below) — a lead-to-lead body is
* not the roster; it is where leads discuss host shapes, credentials and unmerged work.
* Mapping this to {@code READ} would let any worker read every peer lead's mail in full
* and would silently falsify that comment for every other {@code READ} caller.
*/
COORD_READ,
/** Scrape the metrics endpoint. */
METRICS
}
@@ -68,6 +78,11 @@ public final class Authz {
// Observation is open to every authenticated role: a worker legitimately polls its own
// status, and the roster carries no secrets.
case READ, METRICS -> caller.isPrimary() || caller.isWorker() || caller.isArchitect();
// fleetd #421: reading held lead-to-lead mail is the primary's alone. An architect
// holds READ today (CB-548), so "not primary" must mean not-architect here too — this
// is coordination between leads, not observation of the roster.
case COORD_READ -> caller.isPrimary();
};
}
@@ -11,6 +11,7 @@ import java.util.LinkedHashMap;
import java.util.List;
import java.util.Map;
import java.util.Objects;
import java.util.function.Supplier;
/**
* The architect-slot registry (CB-548): every gateway-local architect name and the strong-model
@@ -19,16 +20,43 @@ import java.util.Objects;
*
* <p>Two halves, split by who owns each:
* <ul>
* <li><b>slots</b> — configured once, keyed by the gateway-local unique name; each carries the
* {@code profile} reference the spawn lifecycle reads when it stands the slot up. A read-only
* snapshot taken at construction.</li>
* <li><b>terminal bindings</b> — owned by this registry and initially <em>empty</em>. Config
* declares no architect terminal, so at startup every slot is idle and nothing resolves to an
* architect; a session only becomes one when the spawn lifecycle {@linkplain #bind(String,
* String) binds} its terminal to a slot. {@link CallerResolver} reads this through
* {@link #snapshot()} to turn a pane into an {@link Role#ARCHITECT}.</li>
* <li><b>slots</b> — read from {@code fleet.architects}/{@code developers}/{@code reviewers}
* (see {@link #slots()}), each carrying the {@code profile} reference the spawn lifecycle
* reads when it stands the slot up. <strong>Live, since fleetd #424</strong>: {@link #live}
* re-reads {@code fleet:} on every call, through a supplier the same shape as
* {@code CompositePeerLauncher}'s (see {@code ConfigRef}'s class doc) — so a config reload
* that removes or adds an architect slot governs the <em>next</em> spawn with no restart.
* Only {@link #MemberRegistry(FleetConfig.Fleet)} freezes the pool at construction, and that
* constructor exists for tests and for the (rare) case of wiring a fixed, code-built config.</li>
* <li><b>terminal bindings</b> — owned by this registry, initially <em>empty</em>, and
* <strong>never</strong> touched by a reload. Config declares no architect terminal, so at
* startup every slot is idle and nothing resolves to an architect; a session only becomes one
* when the spawn lifecycle {@linkplain #bind(String, String) binds} its terminal to a slot.
* {@link CallerResolver} reads this through {@link #snapshot()} to turn a pane into an
* {@link Role#ARCHITECT}.</li>
* </ul>
*
* <p><strong>The binding rule (fleetd #424): config governs what a bound slot still grants, as
* well as what may be bound next.</strong> Removing a slot from config revokes it — that is the
* ticket's entire point ("Revoking an architect slot does not revoke it"). Revoking it means an
* architect already bound to that slot loses the ARCHITECT privilege on its very next request:
* {@link #roleForSlot} and {@link #nameForSlot} read {@link #slots()} directly, with no cache, so
* the moment a slot drops out of config, {@link CallerResolver#resolve} (which calls both on every
* request from a bound pane, {@code CallerResolver.java:220}) can no longer confirm the pane's slot
* is an architect slot, and the pane falls through to {@code Principal.worker(...)}. What does
* <em>not</em> change is the {@code terminalToSlot} <em>occupancy</em> — the binding created by
* {@link #bind} is untouched by a reload, on purpose: unbinding it here would double-book the slot
* key (a second terminal could then bind to the "freed" key while the first is still the terminal
* the operator actually meant to demote) and would silently break {@link #unbind}'s compare-safe
* contract, which needs the original {@code terminal → slot} pair intact to remove it cleanly. So
* the demoted session keeps occupying its slot — {@link #slotForTerminal} and {@link #snapshot()}
* still name it — it just no longer resolves as an architect through that occupancy, and a fresh
* spawn still cannot bind to the same key while it is occupied ({@link #reserve}/
* {@link #requireSlotFor} refuse it anyway, since it is gone from {@link #slots()}). The demoted
* session's own turn is unaffected: {@code fleet_reply}'s authorization
* ({@code Authz.Action.REPLY}) is {@code caller.ownsSession(targetSession)} — identity by terminal,
* not by role — so a demoted architect can still end its own turn normally.
*
* <p>Spawning/lifecycle is deliberately a separate unit: this class only owns the bindings and
* exposes the map the resolver resolves against plus the profile lookup lifecycle will call.
* Nothing here creates or manages an architect session.
@@ -55,14 +83,42 @@ public final class MemberRegistry implements MemberLifecycle {
}
}
private final Map<String, Entry> slots;
private final Supplier<FleetConfig.Fleet> fleet;
/** Live {@code terminal_id → qualified slot key}; guarded by {@code terminalToSlot}. */
private final Map<String, String> terminalToSlot = new HashMap<>();
/** Slot keys held between reservation and the terminal binding. Guarded by terminalToSlot. */
private final java.util.Set<String> reservedSlots = new java.util.HashSet<>();
/** Flatten every role pool in {@code fleet} into one registry. Leaders are not members. */
/**
* Freeze the pool at construction — for tests, and for the rare case of wiring a fixed,
* code-built config. Production wiring should prefer {@link #live}, which re-reads
* {@code fleet:} on every call.
*/
public MemberRegistry(FleetConfig.Fleet fleet) {
this(() -> fleet);
}
private MemberRegistry(Supplier<FleetConfig.Fleet> fleet) {
this.fleet = fleet;
}
/**
* Live variant (fleetd #424): {@code fleet} is read fresh on every {@link #slots()} call — pass
* {@code () -> config.get().fleet()}, the same supplier shape {@code CompositePeerLauncher}
* already uses for placement — so a reload that adds or removes an architect slot governs the
* next spawn's {@link #reserve}/{@link #requireSlotFor} check with no restart. A separate,
* private constructor rather than a same-arity public overload of
* {@link #MemberRegistry(FleetConfig.Fleet)}: a {@code FleetConfig.Fleet} and a
* {@code Supplier<FleetConfig.Fleet>} overload are ambiguous for a literal {@code null} — the
* same reason {@code CallerResolver.withLeads} is a static factory rather than a fourth
* constructor overload.
*/
public static MemberRegistry live(Supplier<FleetConfig.Fleet> fleet) {
return new MemberRegistry(Objects.requireNonNull(fleet, "fleet"));
}
/** Flatten every role pool in {@code fleet} into one map. Leaders are not members. */
private static Map<String, Entry> flatten(FleetConfig.Fleet fleet) {
Map<String, Entry> flat = new LinkedHashMap<>();
if (fleet != null) {
for (MemberRole role : MemberRole.values()) {
@@ -74,18 +130,22 @@ public final class MemberRegistry implements MemberLifecycle {
});
}
}
this.slots = Collections.unmodifiableMap(flat);
return Collections.unmodifiableMap(flat);
}
/** The configured slots, keyed by qualified {@link Entry#key()}. Unmodifiable snapshot. */
/**
* The configured slots, keyed by qualified {@link Entry#key()}. Unmodifiable snapshot of
* {@code fleet:} <em>as of this call</em> — see the class doc for which constructor makes that
* live versus frozen.
*/
public Map<String, Entry> slots() {
return slots;
return flatten(fleet.get());
}
/** The slots belonging to {@code role}, in definition order. */
/** The slots belonging to {@code role}, in definition order, as of this call. */
public Map<String, Entry> slotsFor(MemberRole role) {
Map<String, Entry> out = new LinkedHashMap<>();
slots.forEach((key, e) -> {
slots().forEach((key, e) -> {
if (e.role() == role) {
out.put(key, e);
}
@@ -117,31 +177,49 @@ public final class MemberRegistry implements MemberLifecycle {
}
/**
* The strong-model profile a slot runs under — what the spawn lifecycle reads.
* The strong-model profile a slot runs under, as of this call.
*
* <p>Nothing in {@code src/main} calls this (fleetd #431 — grepped both the {@code
* .profileForSlot(} and the {@code ::profileForSlot} form). This javadoc used to say "what the
* spawn lifecycle reads", and that seam does not exist: the spawn lifecycle takes its profile
* from the {@link MemberLifecycle.SlotReservation} that {@code reserve} returns, never from
* here. Kept and pinned rather than deleted because it is the natural accessor for that seam
* if one is added; live for the same reason as {@link #roleForSlot}, so a reload cannot leave
* it answering for the old config.
*
* @return the slot's configured {@code profile}, or {@code null} if the slot is unknown or
* declares none
*/
public String profileForSlot(String slotName) {
Entry e = slots.get(slotName);
Entry e = slots().get(slotName);
return (e == null || e.profile() == null) ? null : e.profile();
}
/** The role a qualified slot key belongs to, or {@code null} when the key is unknown. */
/**
* The role a qualified slot key belongs to, or {@code null} when the key is not currently
* configured. Deliberately live, with no cache (fleetd #424, see the class doc's binding rule):
* removing a slot from config must make {@link CallerResolver#resolve} stop granting the
* ARCHITECT role for it on the very next request from a terminal that was bound to it, which is
* the ticket's whole point — revoking a slot must actually revoke it, not just refuse the next
* spawn.
*/
public MemberRole roleForSlot(String slotName) {
Entry e = slots.get(slotName);
Entry e = slots().get(slotName);
return e == null ? null : e.role();
}
/** The unqualified configured name for a slot, or {@code null} if it is unknown. */
/**
* The unqualified configured name for a slot, or {@code null} if it is not currently configured.
* Live for the same reason as {@link #roleForSlot} — see the class doc's binding rule.
*/
public String nameForSlot(String slotName) {
Entry e = slots.get(slotName);
Entry e = slots().get(slotName);
return e == null ? null : e.name();
}
/** True when {@code slotName} is a configured architect slot. */
public boolean isSlot(String slotName) {
return slots.containsKey(slotName);
return slots().containsKey(slotName);
}
/**
@@ -32,9 +32,26 @@ import java.util.function.Supplier;
* not the fact that they are config. Most of {@code fleet:} — every role pool
* ({@code architects}/{@code developers}/{@code reviewers}), {@code charters}, and
* {@code tabLabel} — is read the same live way, through the same supplier
* ({@code () -> config.get().fleet()}). <strong>But {@code fleet:} as a whole is NOT in this
* class</strong>: {@code fleet.leaders} inside the same key is frozen, which is exactly what
* makes {@code fleet:} split rather than hot — see below.</li>
* ({@code () -> config.get().fleet()}). {@code architects} in particular is hot for
* <strong>two independent consumers</strong> (fleetd #424): {@code CompositePeerLauncher}
* reads it live for placement (which profile an unqualified architect spawn may land on), and
* {@code MemberRegistry} separately reads it live, through its own instance of the same
* supplier shape, for identity — both which slot a spawn may bind to <em>and</em> what a slot
* already bound still grants. Removing an architect slot from config therefore revokes the
* {@link dev.ltms.fleet.auth.Role#ARCHITECT} role on the bound pane's very next request; only
* the slot <em>occupancy</em> survives, so the demoted session still holds its slot key until
* it unbinds. See {@code MemberRegistry}'s class doc for that binding rule.
* <strong>But {@code fleet:} as a whole is NOT in this class</strong>: {@code fleet.leaders}
* inside the same key is frozen, which is exactly what makes {@code fleet:} split rather than
* hot — see below. {@code models:} (fleetd #422) joined this class whole: {@link
* FleetConfig#validateModels()} re-runs fully against the fresh config on every {@link
* #reload()} (via {@link FleetConfig#validateAll()}), refusing a bad edit outright rather than
* caching a stale copy anywhere, and the on/off half added by fleetd #422 is read live both by
* {@code CompositePeerLauncher}'s spawn gate ({@code enforceModelEnabled} and its candidate
* filter) and by {@code fleet_profiles}/{@code GET /profiles} (via
* {@code PeerLauncher.disabledModels()}). Nothing about {@code models:} is baked into an
* object built at startup, so — unlike the deferred keys below — there is no frozen half left
* to report; it moved here from deferred rather than joining split.</li>
* <li><strong>Deferred</strong> — accepted into the new snapshot, but the wiring built at startup
* keeps the old value until a restart: {@code lifecycle:}, {@code leadHeartbeat:},
* {@code idleSleepGuard:} ({@code Fleetd.java} reads it once, at startup, to decide whether
@@ -43,12 +60,6 @@ import java.util.function.Supplier;
* running daemon keeps whatever this was at startup regardless of a later edit),
* {@code spawnReadyTimeoutMs} / {@code spawnReadyPollMs}, {@code quarantineCooldownSeconds}
* (CB-578 stage B — baked once into the {@code BackendQuarantine} built at startup),
* {@code models:} (fleetd ticket "central allow-list of usable models" — {@link
* FleetConfig#validateModels()} re-runs against the fresh config in {@link #reload()}
* (via {@link FleetConfig#validateAll()}), so a
* models.allow: edit that would refuse to boot still refuses the reload; a change that
* passes has nothing built at startup to rebuild, so it is reported deferred rather than
* silently accepted with no report at all),
* {@code guard:}, {@code worktreeRoot:}, {@code worktreeGroup:} and {@code memberSkills:}
* (all three of the latter baked once into the {@code GitWorktrees} built at
* {@code Fleetd.java:251} and never rebuilt — fleetd #323 instance 2 found
@@ -141,16 +152,18 @@ import java.util.function.Supplier;
* </ul>
*
* <p><strong>The denominator, measured on 2026-09-04 (fleetd #330; recounted for fleetd #333);
* recounted again for fleetd #362, again after {@code idleSleepGuard:} was added, and again after
* {@code models:} was added.</strong>
* {@code FleetConfig} has 25 top-level record components: 5 cold, 14 deferred, 3 split, 3
* hot-excluded. Three of them are named nowhere in this file, and the reason is the same for all
* three: {@code placement}, {@code memberCredentials} and {@code memberLoginShell} are
* <strong>hot</strong> and correctly absent — all three are read live off {@code config.get()}
* recounted again for fleetd #362, again after {@code idleSleepGuard:} was added, again after
* {@code models:} was added as deferred, and again for fleetd #422, which moved {@code models:}
* from deferred to hot-excluded once its on/off half was read live everywhere.</strong>
* {@code FleetConfig} has 25 top-level record components: 5 cold, 13 deferred, 3 split, 4
* hot-excluded. Four of them are named nowhere in this file, and the reason is the same for all
* four: {@code placement}, {@code memberCredentials}, {@code memberLoginShell} and {@code models}
* are <strong>hot</strong> and correctly absent — all four are read live off {@code config.get()}
* (placement through the {@code CompositePeerLauncher} supplier the Hot bullet names;
* {@code memberCredentials}/{@code memberLoginShell} at spawn time, {@code Fleetd.java:198, 205, 729}
* and {@code HerdrPeerLauncher#configuredMemberLoginShell}), so a reload takes effect on the next
* spawn with no entry needed here.
* and {@code HerdrPeerLauncher#configuredMemberLoginShell}; {@code models} the same way, through the
* Hot bullet's {@code models:} paragraph), so a reload takes effect on the next spawn (or, for
* {@code models}, the next reported status) with no entry needed here.
* {@code health} and {@code coordinator} used to be a third kind — <strong>undecided</strong>, not
* hot — until fleetd #330 added the <strong>split</strong> class above and gave them a home. A
* reload touching either used to report a bare "config reloaded", which under-claimed; now it names
@@ -226,7 +239,7 @@ public final class ConfigRef implements Supplier<FleetConfig> {
static final Set<String> DEFERRED_KEYS = Set.of(
"guard", "worktreeRoot", "worktreeGroup", "memberSkills", "primary", "configReload",
"leadHeartbeat", "lifecycle", "spawnReadyTimeoutMs", "spawnReadyPollMs",
"quarantineCooldownSeconds", "profiles", "idleSleepGuard", "models");
"quarantineCooldownSeconds", "profiles", "idleSleepGuard");
private final Path path;
private final AtomicReference<FleetConfig> current;
@@ -446,15 +459,6 @@ public final class ConfigRef implements Supplier<FleetConfig> {
if (!Objects.equals(old.idleSleepGuard(), fresh.idleSleepGuard())) {
changed.add("idleSleepGuard");
}
// fleetd ticket "central allow-list of usable models": validateModels() runs again in
// reload() above (via validateAll()), so a bad edit is already refused as cold-adjacent
// (the whole reload is refused via the catch block, never partially applied). A GOOD edit
// to the allow-list
// itself has nothing built at startup to rebuild — it only ever mattered to the validation
// call that already ran — so report it deferred rather than silently swallowing the change.
if (!Objects.equals(old.models(), fresh.models())) {
changed.add("models");
}
if (!Objects.equals(old.spawnReadyTimeoutMs(), fresh.spawnReadyTimeoutMs())
|| !Objects.equals(old.spawnReadyPollMs(), fresh.spawnReadyPollMs())) {
changed.add("spawnReady*");
@@ -529,26 +533,35 @@ public final class ConfigRef implements Supplier<FleetConfig> {
+ "opened once and needs a restart; the broker URI env-var name kept out of a "
+ "member's environment is read live on every spawn and already applied");
}
// fleetd #333: unlike health/coordinator above, most of `fleet:` (architects, developers,
// reviewers, charters, tabLabel) is genuinely hot — ConfigRefTest.aHotChangeIsAppliedAndRead-
// fleetd #333: unlike health/coordinator above, most of `fleet:` (developers, reviewers,
// charters, tabLabel) is genuinely hot — ConfigRefTest.aHotChangeIsAppliedAndRead-
// ThroughGet and aCharterChangeIsHotAndReachesTheLiveConfig prove it reaches the live config
// with no restart note. Only fleet.leaders is frozen (Fleetd.java:281 reads
// cfg.fleet().leaders() off the startup snapshot to build both the LeadTabScanner's
// tab-label-to-name map, wired into CallerResolver.withLeadsAndMembers at Fleetd.java:620/624,
// and — when herdr answered — LeadLauncher(...).ensureLeads() at Fleetd.java:315, which
// auto-launches each lead up to its `instances` count; neither is rebuilt on reload). So this
// compares fleet.leaders alone, not the whole Fleet record: comparing the whole record would
// report "split" for a tabLabel-only or charters-only change that is actually fully hot,
// which is the over-claim mirror of the under-claim bug this class exists to prevent.
// with no restart note. `architects` is hot too, and — since fleetd #424 — hot for BOTH of
// its consumers, not just the one this comment used to name: CompositePeerLauncher reads it
// live for PLACEMENT through the () -> config.get().fleet() supplier named in the class doc's
// Hot bullet, and MemberRegistry separately reads it live for IDENTITY (which slot a spawn
// may bind to, AND what a slot already bound still grants) through its own instance of that
// same supplier shape — see MemberRegistry.live and its class doc for the binding rule:
// removing a slot revokes ARCHITECT on the bound pane's very next request, and only the slot
// OCCUPANCY survives, so the demoted session keeps its slot key until it unbinds. Only
// fleet.leaders is frozen (Fleetd.java:281 reads cfg.fleet().leaders() off the startup
// snapshot to build both the LeadTabScanner's tab-label-to-name map, wired into
// CallerResolver.withLeadsAndMembers at Fleetd.java:620/624, and — when herdr answered —
// LeadLauncher(...).ensureLeads() at Fleetd.java:315, which auto-launches each lead up to its
// `instances` count; neither is rebuilt on reload). So this compares fleet.leaders alone, not
// the whole Fleet record: comparing the whole record would report "split" for a tabLabel-only
// or architects-only change that is actually fully hot, which is the over-claim mirror of the
// under-claim bug this class exists to prevent.
if (!Objects.equals(leadersOf(old), leadersOf(fresh))) {
changed.add("fleet: fleet.leaders (each lead's tab, workspace, cwd, profile and "
+ "instances count) is read once at startup to build the LeadTabScanner's "
+ "identity map and to auto-launch leads, and neither is rebuilt on reload, so a "
+ "lead added, removed, or given a new tab: label needs a restart — until then it "
+ "stays unrecognised, and a caller from its new tab resolves as a worker, not a "
+ "lead; the rest of fleet: (architects, developers, reviewers, charters, "
+ "tabLabel) is read live through the supplier on CompositePeerLauncher and "
+ "already applied");
+ "lead; the rest of fleet: (developers, reviewers, charters, tabLabel) is read "
+ "live through the supplier on CompositePeerLauncher, and architects is read "
+ "live through that same supplier for placement AND through a separate supplier "
+ "on MemberRegistry for spawn-time identity — both already applied");
}
// Kept in step with SPLIT_KEYS the same way changedColdKeys is kept in step with COLD_KEYS —
// every message here must be traceable to one of the split keys the class doc documents.
@@ -133,8 +133,10 @@ import java.util.regex.PatternSyntaxException;
* nothing and every existing config keeps working exactly as it does today.
* When non-empty, a profile whose {@code model:} is not one of {@link
* Models#ids()} fails config load, naming both the model and the profile —
* see {@link #validateModels()}. This block only decides what may be
* CONFIGURED; nothing here enforces it at spawn time. See {@link Models}.
* see {@link #validateModels()}. This block decides what may be CONFIGURED;
* fleetd #422 added the separate on/off question — whether a configured model
* may be spawned onto RIGHT NOW ({@link Models.ModelEntry#enabled}) — enforced
* live at spawn by {@code CompositePeerLauncher}, not here. See {@link Models}.
*/
@JsonIgnoreProperties(ignoreUnknown = true)
public record FleetConfig(
@@ -1362,9 +1364,14 @@ public record FleetConfig(
* the set of permitted models — only editing {@code models.allow:} itself can. This is the
* invariant the ticket asked for: the two blocks are validated in one direction only.
*
* <p><b>Out of scope here, deliberately:</b> nothing in this block is read at spawn time —
* enforcing it against a live spawn, an on/off runtime switch, and any interaction with {@code
* BackendQuarantine} are separate units. This block is config-load validation only.
* <p><b>Spawn-time enforcement (fleetd #422, units 2+3) lives outside this record</b> —
* {@code CompositePeerLauncher.enforceModelEnabled} and its candidate-set filter read this
* block LIVE (through the same kind of supplier {@code weight}/{@code maxLoad} already use),
* so the on/off state below is hot: no restart needed. This block itself still only decides
* what may be CONFIGURED (membership in {@link #allow}); {@link ModelEntry#enabled} decides
* whether a member of that list is currently spawnable. The two questions are deliberately
* separate — see {@link ModelEntry}'s javadoc for why turning a model off must never mean
* removing it from {@link #allow}.
*
* @param allow the permitted models, each its own {@link ModelEntry} rather than a bare
* string — see that record's javadoc for why. {@code null}/empty ⇒ the block is
@@ -1378,24 +1385,56 @@ public record FleetConfig(
}
/**
* One permitted model, named as a record rather than a bare string on purpose: a later unit
* needs to hang an on/off state and a load-limit state off each entry, and a bare {@code
* List<String>} cannot grow those fields without changing the YAML shape underneath every
* operator who already wrote one. {@link #model()} is intentionally a single flat,
* opaque-string namespace — a bare Claude id ({@code claude-sonnet-5}) and an opencode
* provider-prefixed id ({@code openai/gpt-5.6-terra}) both fit it unchanged, because
* {@link FleetConfig#validateModels()} only ever compares a profile's {@code model:} value
* against this string for exact equality; it never parses a provider prefix or branches on
* a profile's {@code kind:}.
* One permitted model, named as a record rather than a bare string on purpose: this ticket
* (fleetd #422) is the "later unit" the original comment here predicted — it hangs an on/off
* state ({@link #enabled}) off each entry, and a bare {@code List<String>} could not have
* grown that field without changing the YAML shape underneath every operator who already
* wrote one. {@link #model()} is intentionally a single flat, opaque-string namespace — a
* bare Claude id ({@code claude-sonnet-5}) and an opencode provider-prefixed id
* ({@code openai/gpt-5.6-terra}) both fit it unchanged, because {@link
* FleetConfig#validateModels()} only ever compares a profile's {@code model:} value against
* this string for exact equality; it never parses a provider prefix or branches on a
* profile's {@code kind:}.
*
* @param model the model id exactly as a {@code profiles:} entry's {@code model:} field
* would name it
* @param model the model id exactly as a {@code profiles:} entry's {@code model:} field
* would name it
* @param enabled {@code false} turns spawning onto this model off; {@code null} (the field
* omitted — every config written before fleetd #422 is this shape) or
* {@code true} leaves it on. Turning a model off must NEVER remove it from
* {@link Models#allow} — {@link FleetConfig#validateModels()} checks
* <em>membership</em> only, never the on/off state, so an off entry stays a
* valid thing for a {@code profiles:} entry to name; only
* {@code CompositePeerLauncher}'s spawn-time gate reads {@link #enabled}.
* Collapsing the two — turning a model off by deleting its {@code allow:}
* entry — would make {@link FleetConfig#validateModels()} refuse the whole
* config reload the moment a still-configured profile names it, which is
* exactly the restart-to-flip-a-switch problem this field exists to avoid.
* <p>A model id can be named by more than one {@code profiles:} entry (e.g.
* {@code deepseek-v4-flash} backs both {@code local} and {@code
* local-direct} in the live config) — turning it off disables every profile
* that names it, on purpose: the model is what a subscription's rate limit
* actually constrains, not any one profile alias for it.
*/
@JsonIgnoreProperties(ignoreUnknown = true)
public record ModelEntry(String model) {
public record ModelEntry(String model, Boolean enabled) {
public ModelEntry {
model = (model == null || model.isBlank()) ? null : model.trim();
}
/**
* Back-compat form before {@link #enabled} was added (fleetd #422) — the model is
* unconditionally on, exactly as every {@code ModelEntry} behaved before this field
* existed. Keeps pre-#422 call sites (and any YAML that omits {@code enabled:})
* compiling and behaving identically.
*/
public ModelEntry(String model) {
this(model, null);
}
/** {@code true} unless {@link #enabled} is explicitly {@code false} — absent means on. */
public boolean isEnabled() {
return !Boolean.FALSE.equals(enabled);
}
}
/** {@link #allow}'s model ids, as a set for membership checks. Blank/null entries are dropped. */
@@ -1408,6 +1447,23 @@ public record FleetConfig(
}
return Collections.unmodifiableSet(ids);
}
/**
* Model ids currently turned off (fleetd #422: {@link ModelEntry#isEnabled()} {@code
* false}). Read live by {@code CompositePeerLauncher}'s spawn gate and by {@code
* fleet_profiles}/{@code GET /profiles} — both must read this same accessor off the same
* live config so the two surfaces cannot disagree about which model is off (the fleetd
* #404 lesson: a status field must read the source the behaviour reads).
*/
public Set<String> offIds() {
Set<String> off = new java.util.LinkedHashSet<>();
for (ModelEntry e : allow) {
if (e != null && e.model() != null && !e.isEnabled()) {
off.add(e.model());
}
}
return Collections.unmodifiableSet(off);
}
}
/**
@@ -3,7 +3,7 @@ package dev.ltms.fleet.herdr;
import com.fasterxml.jackson.databind.JsonNode;
/**
* Client face onto the herdr daemon (protocol 14, herdr 0.7.0).
* Client face onto the herdr daemon (protocol 19, herdr 0.8.0).
*
* <p>This is the ONLY thing in {@code fleetd} that speaks to herdr. Every method
* maps to a herdr JSON-RPC call over its Unix domain socket. Requests are
@@ -8,7 +8,7 @@ import com.fasterxml.jackson.databind.node.ObjectNode;
import java.nio.charset.StandardCharsets;
/**
* Wire codec for herdr's newline-delimited JSON-RPC (protocol 14).
* Wire codec for herdr's newline-delimited JSON-RPC (protocol 19).
*
* <p>Split out from the socket so the framing rules — the ones that actually bit us
* during the spike (id MUST be a string; response carries {@code result} or
@@ -698,17 +698,57 @@ public final class CompletionResolver implements TurnListener {
}
/**
* Coverage summary for the CB-578 stage A exhausted-pattern classification, logged at startup
* What an unset pattern key means for the classification it configures (fleetd#415).
* {@code coverage()} cannot infer this from the key's name — the two keys it currently
* describes disagree on it, and a string comparison on the name would just move the same bug
* to a new spot — so every caller must state it explicitly.
*
* <p><strong>This alone does not prove a caller passes the right one for its key.</strong> A
* test that calls {@code coverage()} directly and supplies the meaning itself only proves this
* enum is worded correctly, never that {@code Fleetd}'s two call sites pair each key with its
* true meaning — that pairing is #415's actual defect. Measured on review: swapping the two
* {@code UnsetMeaning} arguments at those call sites (giving {@code exhaustedPattern} the
* built-in-default wording and {@code errorPattern} the off wording — #415's exact defect with
* the keys exchanged) compiled with 0 errors and left all 1506 existing tests green. See
* {@code dev.ltms.fleet.Fleetd#exhaustedPatternCoverageLine}/{@code #errorPatternCoverageLine}
* and {@code FleetdPatternCoverageLineTest}, which exists specifically to catch that swap.
*/
public enum UnsetMeaning {
/** No fallback exists: a profile with no configured pattern truly has this classification off. */
OFF,
/** A built-in pattern applies when unset: the classification still runs for that profile. */
BUILT_IN_DEFAULT
}
/**
* Coverage summary for a fleetd#201/CB-578-style pattern-key classification, logged at startup
* the way {@link dev.ltms.fleet.health.FleetHealthMonitor#coverage} is — so an operator can
* see whether the classification is on, and for which profiles, without reading every
* profile's config by hand.
*
* <p>fleetd#415: this method measures <em>pattern coverage</em> — how many profiles set the
* key — which is not the same thing as <em>feature state</em> for a key with a fallback. For
* {@code errorPattern}, an empty {@code configuredProfiles} still runs the classification
* against {@code CompletionResolver}'s built-in compatibility pattern ({@link #BACKEND_ERROR}
* at line ~84); for {@code exhaustedPattern} there is no fallback, so empty really does mean
* off. {@code unsetMeaning} is the single, required source of that fact — see
* {@link dev.ltms.fleet.config.FleetConfig#rejectMalformedProfilePatterns} lines ~2029-2032 for
* where it is documented for config authors. It is a required parameter, not a defaulted
* overload: a third pattern key added later must supply one to compile at all, rather than
* silently inheriting whichever wording this method happened to default to.
*
* @param allProfiles every configured profile name
* @param configuredProfiles the subset of {@code allProfiles} that carry an exhausted pattern
* @param configuredProfiles the subset of {@code allProfiles} that carry the pattern
*/
public static String coverage(String patternKey, Set<String> allProfiles, Set<String> configuredProfiles) {
public static String coverage(String patternKey, UnsetMeaning unsetMeaning, Set<String> allProfiles,
Set<String> configuredProfiles) {
if (configuredProfiles.isEmpty()) {
return "off (no profile has an " + patternKey + " configured; profiles: " + sorted(allProfiles) + ")";
return switch (unsetMeaning) {
case OFF -> "off (no profile has an " + patternKey + " configured; profiles: "
+ sorted(allProfiles) + ")";
case BUILT_IN_DEFAULT -> "built-in default for all profiles (no profile customises "
+ patternKey + "; profiles: " + sorted(allProfiles) + ")";
};
}
Set<String> unconfigured = new TreeSet<>(allProfiles);
unconfigured.removeAll(configuredProfiles);
@@ -37,6 +37,7 @@ import io.modelcontextprotocol.spec.McpSchema;
import com.fasterxml.jackson.databind.ObjectMapper;
import jakarta.servlet.http.HttpServlet;
import java.util.ArrayList;
import java.util.LinkedHashMap;
import java.util.List;
import java.util.Map;
@@ -385,10 +386,11 @@ public final class FleetMcp {
(exchange, req) -> {
Map<String, Object> a = req.arguments();
String target = str(a, "target");
String coordId = str(a, "coordId");
// The action depends on the ARGUMENTS, not on the tool name -- see pollAction.
McpSchema.CallToolResult denied = deny(exchange, toolAction("fleet_poll", a), target);
if (denied != null) return denied;
return poll(messages, str(a, "ticket"), target);
return poll(messages, leadChannel, str(a, "ticket"), target, coordId);
};
// CB-307 Increment 3: per-msgId ack (not needed in v1 but supported by the inbox).
// Acking removes a reply from the inbox, so it is a drain, not a read.
@@ -810,29 +812,42 @@ public final class FleetMcp {
/**
* Which authorization action a {@code fleet_poll} call needs, decided by its arguments
* (fleetd #272).
* (fleetd #272, widened by fleetd #421).
*
* <p>{@code fleet_poll} is <strong>two operations behind one tool name</strong>. With {@code
* ticket} it observes an async delegation and changes nothing, which is a {@link
* <p>{@code fleet_poll} is now <strong>three operations behind one tool name</strong>. With
* {@code ticket} it observes an async delegation and changes nothing, which is a {@link
* Authz.Action#READ}. With {@code target} it calls {@link MessageService#drainReplies} on that
* session -- the replies are removed from the inbox and a second call returns nothing -- so it
* is a {@link Authz.Action#DRAIN}, the same gate {@code fleet_ack} already uses for removing a
* single message, and the same one the REST path uses at {@code FleetApp.drainReplies}.
* single message, and the same one the REST path uses at {@code FleetApp.drainReplies}. With
* {@code coordId} it reads (never acks) this daemon's own held lead-to-lead mail, which is a
* {@link Authz.Action#COORD_READ} -- <strong>not</strong> {@code READ}, even though nothing is
* consumed: {@code READ}'s grant is open to every authenticated role on the premise that the
* roster carries no secrets, and a lead-to-lead body is not the roster. Mapping a non-destructive
* peer-mail read to {@code READ} would let any worker read every peer lead's mail in full.
*
* <p>Until this method existed the handler passed a constant {@code READ} for both branches.
* {@code READ} is open to every authenticated role, so any worker could read a peer's id out of
* {@code fleet_list} and destroy the replies that peer had queued for the primary. The gate
* failed open, and it did so because the required action is a function of the arguments while
* the handler chose it before looking at them.
* <p>Before this method existed (fleetd #272) the handler passed a constant {@code READ} for
* both of the original branches. {@code READ} is open to every authenticated role, so any
* worker could read a peer's id out of {@code fleet_list} and destroy the replies that peer had
* queued for the primary. The gate failed open, and it did so because the required action is a
* function of the arguments while the handler chose it before looking at them.
*
* <p>The choice lives in this method, and not inline in the handler, so that a test can assert
* the mapping the handler actually uses. {@code FleetMcpAuthzTest} already checked every
* {@link Authz.Action} against every {@link Role} and passed throughout -- it tested the policy
* table, which was correct, while the defect was in which action the caller handed it.
*
* @param target the {@code target} argument of the call, or {@code null}/blank when absent
* <p>Checked first, and exclusively of {@code target}: a call naming {@code coordId} is reading
* a different inbox entirely (this daemon's own lead channel, never a worker's), so it takes
* priority over whatever {@code target} might also say.
*
* @param target the {@code target} argument of the call, or {@code null}/blank when absent
* @param coordId the {@code coordId} argument of the call, or {@code null}/blank when absent
*/
static Authz.Action pollAction(String target) {
static Authz.Action pollAction(String target, String coordId) {
if (!isBlank(coordId)) {
return Authz.Action.COORD_READ;
}
return isBlank(target) ? Authz.Action.READ : Authz.Action.DRAIN;
}
@@ -847,7 +862,7 @@ public final class FleetMcp {
case "fleet_reply" -> Authz.Action.REPLY;
case "fleet_ask" -> Authz.Action.ASK;
case "fleet_status", "fleet_list", "fleet_profiles", "fleet_whoami" -> Authz.Action.READ;
case "fleet_poll" -> pollAction(str(arguments, "target"));
case "fleet_poll" -> pollAction(str(arguments, "target"), str(arguments, "coordId"));
case "fleet_ack" -> Authz.Action.DRAIN;
case "fleet_spawn" -> Authz.Action.SPAWN;
case "fleet_stop" -> Authz.Action.STOP;
@@ -857,6 +872,21 @@ public final class FleetMcp {
/** {@code fleet_poll}: check an async delegation by ticket, or drain a worker's inbox by target. */
static McpSchema.CallToolResult poll(MessageService messages, String ticket, String target) {
return poll(messages, null, ticket, target, null);
}
/**
* As above, plus (fleetd #421) a held-peer-mail read when {@code coordId} is present: returns
* this daemon's own held lead-to-lead messages, in full, without acking them. Checked first and
* exclusively of {@code ticket}/{@code target} — see {@link #pollAction}'s javadoc for why this
* is a different inbox (this daemon's own {@link LeadChannel}) that authorizes differently
* ({@link Authz.Action#COORD_READ}, primary-only) from either of the original two branches.
*/
static McpSchema.CallToolResult poll(MessageService messages, LeadChannel leadChannel, String ticket,
String target, String coordId) {
if (!isBlank(coordId)) {
return pollHeldPeerMail(leadChannel, coordId);
}
if (!isBlank(target)) {
var replies = messages.drainReplies(target);
if (replies.isEmpty()) {
@@ -884,6 +914,46 @@ public final class FleetMcp {
};
}
/**
* fleetd #421: a lead's own held lead-to-lead mail, read without consuming it.
*
* <p>{@link LeadChannel#peek} is non-destructive, so calling this twice returns the same
* bodies, and {@code fleet_list}'s {@code coordinator.held[]} is unaffected — this adds a
* read, it never acks. Never let this short-circuit the delivery contract: {@code
* LeadCoordLoop}'s javadoc explains why a message must stay unacked until actually delivered,
* and that is unchanged here.
*
* <p>{@code coordId} must be THIS daemon's own coord-id ({@code fleet_list}'s
* {@code coordinator.selfId}) — there is no route here to read a PEER's outbound mail, only
* your own inbound mail. Requiring the caller to echo its own id catches the exact confusion
* that opened fleetd #421: the original failed attempt passed a PEER's id ("fleet01") as
* {@code fleet_poll}'s {@code target}, expecting to read that peer's messages, and got the same
* silent {@code []} as a genuinely empty worker inbox. This refuses the same mistake here with a
* reason, instead of a second silent wrong answer.
*/
static McpSchema.CallToolResult pollHeldPeerMail(LeadChannel leadChannel, String coordId) {
if (leadChannel == null) {
return error("lead coordination is not configured (no coordinator: block) — there is "
+ "no held peer mail to read.");
}
String selfId = leadChannel.selfCoordId();
if (!coordId.equals(selfId)) {
return error("coordId \"" + coordId + "\" is not this daemon's own coord-id (\"" + selfId
+ "\"). fleet_poll reads only YOUR OWN held mail — pass your own coordId "
+ "(fleet_list's coordinator.selfId), not a peer's.");
}
return text(json(leadChannel.peek().stream().map(FleetMcp::heldMailView).toList()));
}
/** The full body of one held lead-to-lead message — never truncated, unlike {@link #heldView}. */
private static Map<String, Object> heldMailView(LeadMessage m) {
Map<String, Object> row = new LinkedHashMap<>();
row.put("msgId", m.msgId());
row.put("from", m.from());
row.put("content", m.content());
return row;
}
/**
* {@code fleet_reply}: the worker returns its structured answer, resolving the awaiting send
* or — when no send is open — queueing the reply in the inbox for later drain (CB-307).
@@ -1179,6 +1249,23 @@ public final class FleetMcp {
if (!coolingOff.isEmpty()) {
result.put("coolingOff", coolingOff);
}
// fleetd #422: read the exact same accessor CompositePeerLauncher's spawn gate reads
// (PeerLauncher.modelGateState(), which for the composite is models0() read live) — never a
// separately-derived answer, so this status can never overstate or understate what the gate
// actually enforces (the fleetd #404 lesson).
//
// fleetd #422 follow-up: "armed" and "off" come from the ONE modelGateState() call below,
// never two independent reads of the gate — a reload landing between two separate reads
// could otherwise make them disagree. modelGateArmed is reported unconditionally (never
// omitted like quarantined/coolingOff above) precisely so a lead can tell "no models: block
// at all" (false) apart from "a models: block with nothing currently off" (true, with
// modelsOff simply absent below) — the two states PeerLauncher.disabledModels() alone
// cannot distinguish, both reporting an empty set.
PeerLauncher.ModelGateState modelGate = workers.modelGateState();
result.put("modelGateArmed", modelGate.configured());
if (!modelGate.off().isEmpty()) {
result.put("modelsOff", new ArrayList<>(modelGate.off()));
}
return result;
}
@@ -1305,6 +1392,16 @@ public final class FleetMcp {
* counts, it can never make {@code fleet_list} itself slow or fail. {@code held} comes from
* {@link LeadChannel#peek}, a pure in-memory read with no broker round trip, so it is never
* subject to that bound.
*
* <p><strong>fleetd #421: {@code heldCount}/{@code heldDurable} fix the "pending: 0" trap.</strong>
* {@code mailbox.pending} counts only broker-<em>ready</em> messages; a held message is already
* an unacked delivery sitting with this consumer, so the normal, healthy state of a blocked lead
* is {@code "pending": 0} next to a non-empty {@code held[]} — which invites the false reading
* "these are only in memory, a restart will lose them". {@code heldCount} is the honest second
* number beside {@code pending} ({@code held.size()}, not left for the reader to count the
* array). {@code heldDurable} comes straight from {@link LeadChannel#heldDurable}, which the
* channel implementation derives from what it actually did when it declared and consumed its own
* queue (fleetd #440) — this method never asserts the fact itself.
*/
private static Map<String, Object> coordinatorView(CoordinationSource coordination) {
LeadChannel channel = coordination.leadChannel();
@@ -1312,11 +1409,14 @@ public final class FleetMcp {
return null;
}
String selfId = channel.selfCoordId();
List<LeadMessage> held = channel.peek();
Map<String, Object> row = new LinkedHashMap<>();
row.put("selfId", selfId);
row.put("configured", true);
row.put("mailbox", mailboxView(probe(channel, selfId)));
row.put("held", channel.peek().stream().map(FleetMcp::heldView).toList());
row.put("heldCount", held.size());
row.put("heldDurable", channel.heldDurable());
row.put("held", held.stream().map(FleetMcp::heldView).toList());
row.put("peers", coordination.peers().stream().map(p -> peerView(channel, p)).toList());
return row;
}
@@ -1645,10 +1745,18 @@ public final class FleetMcp {
"Check an async delegation (a fleet_send with wait:false) by its ticket: "
+ "pending, done (with the worker's reply), or failed. When target (a worker "
+ "session id) is present instead of ticket, drain that worker's inbox of "
+ "replies delivered when no send was open.",
+ "replies delivered when no send was open. When coordId is present instead, "
+ "read (never consume) your own held lead-to-lead mail — primary-only.",
objectSchema(Map.of(
"ticket", stringProp("The ticket returned by fleet_send wait:false"),
"target", stringProp("Worker session id to drain pending replies from (optional)")),
"target", stringProp("Worker session id to drain pending replies from (optional)"),
"coordId", stringProp("Your own coord-id (fleet_list's coordinator.selfId) — "
+ "reads every message currently held[] for you in full, without "
+ "acking. Read twice, get the same bodies both times; fleet_list's "
+ "held[] still reports them afterward. Primary-only, and this can "
+ "read only YOUR OWN mailbox — there is no route to a peer's outbound "
+ "mail, so passing a peer's coordId here is refused rather than "
+ "silently returning the wrong thing (or nothing).")),
List.of()));
}
@@ -14,6 +14,7 @@ import dev.ltms.fleet.placement.BackendOutagePolicy;
import dev.ltms.fleet.placement.BackendQuarantine;
import dev.ltms.fleet.placement.PlacementCandidate;
import dev.ltms.fleet.placement.PlacementContext;
import dev.ltms.fleet.placement.PlacementDecision;
import dev.ltms.fleet.placement.PlacementException;
import dev.ltms.fleet.placement.PlacementPolicies;
import dev.ltms.fleet.placement.PlacementPolicy;
@@ -90,6 +91,19 @@ public final class CompositePeerLauncher implements PeerLauncher {
private final Supplier<Map<String, FleetConfig.Profile>> profileConfigs;
private final Supplier<PlacementPolicy> placementPolicy;
/**
* fleetd #422: the central model allow-list's on/off state, read live per spawn — same reason
* {@link #profileConfigs} is a supplier rather than a captured map (see the class doc above and
* {@link #enforceModelEnabled}). A caller with no {@code models:} block to read from (the
* simpler, map-based constructors used throughout this class's own tests) wires this to a
* constant {@code null}, which {@link #models0} treats as "nothing configured, gate never
* fires" — the pre-#422 behaviour.
*/
private final Supplier<FleetConfig.Models> models;
/** The value {@link #models0} normalizes a {@code null} supplier result to. */
private static final FleetConfig.Models NO_MODELS_CONFIGURED = new FleetConfig.Models(List.of());
/** CB-578 stage B: credential cooldown, checked before an explicit spawn and filtered into placement. */
private final BackendQuarantine quarantine;
@@ -205,12 +219,44 @@ public final class CompositePeerLauncher implements PeerLauncher {
FleetConfig.Fleet fleet,
BackendQuarantine quarantine,
BackendOutagePolicy outagePolicy) {
// No models: block to read from a plain profiles Map — fleetd #422's gate is wired to a
// constant null (see enforceModelEnabled/models0), the pre-#422 behaviour for every caller
// of this overload. Use the 9-arg overload below to test the gate against a LIVE supplier.
this(delegates, defaultProfile, profileConfigs, placementPolicy, liveCount, fleet, quarantine,
outagePolicy, constant(null));
}
/**
* As above, plus a LIVE model-gate source (fleetd #422). The 8-arg overload above wires
* {@code models} to a constant {@code null} because it has only a static {@code Map<String,
* Profile>}, never a full {@code FleetConfig}, to read one from; this overload exists so a test
* can prove {@link #enforceModelEnabled} and its candidate filter re-read {@link
* FleetConfig.Models} on every call rather than a value captured once at construction — the
* exact distinction fleetd #422 exists to get right (see the class doc's CB-559 note on {@link
* #profileConfigs}, which this follows). Production wiring uses the
* {@code Supplier<FleetConfig>} constructor below instead, which already threads a live
* {@code config.get().models()} through.
*
* @param models required — pass a supplier returning {@code null} for a caller that has no
* {@code models:} block to gate against, never a defaulting overload (the same
* "explicit opt-out, never a silent default" rule {@code quarantine}/
* {@code outagePolicy} already follow).
*/
public CompositePeerLauncher(List<HerdrPeerLauncher> delegates,
String defaultProfile,
Map<String, FleetConfig.Profile> profileConfigs,
PlacementPolicy placementPolicy,
Function<String, Integer> liveCount,
FleetConfig.Fleet fleet,
BackendQuarantine quarantine,
BackendOutagePolicy outagePolicy,
Supplier<FleetConfig.Models> models) {
// LinkedHashMap, not Map.copyOf: candidates() promises definition order and the weighted
// policy breaks exact-weight ties on it, so a salted iteration order would make placement
// differ from one JVM run to the next.
this(delegates, defaultProfile,
constant(Collections.unmodifiableMap(new LinkedHashMap<>(profileConfigs))),
constant(placementPolicy), liveCount, constant(fleet), quarantine, outagePolicy);
constant(placementPolicy), liveCount, constant(fleet), quarantine, outagePolicy, models);
}
/**
@@ -250,7 +296,10 @@ public final class CompositePeerLauncher implements PeerLauncher {
liveCount,
() -> config.get().fleet(),
quarantine,
outagePolicy);
outagePolicy,
// fleetd #422: read live, same as profiles/placement/fleet above — a models.allow
// edit (on/off or otherwise) is visible to the very next spawn, no restart needed.
() -> config.get().models());
}
/** The all-suppliers form every other constructor funnels into. */
@@ -261,7 +310,8 @@ public final class CompositePeerLauncher implements PeerLauncher {
Function<String, Integer> liveCount,
Supplier<FleetConfig.Fleet> fleet,
BackendQuarantine quarantine,
BackendOutagePolicy outagePolicy) {
BackendOutagePolicy outagePolicy,
Supplier<FleetConfig.Models> models) {
this.fleet = fleet;
this.quarantine = Objects.requireNonNull(quarantine, "quarantine");
this.outagePolicy = Objects.requireNonNull(outagePolicy, "outagePolicy");
@@ -273,6 +323,7 @@ public final class CompositePeerLauncher implements PeerLauncher {
this.profileConfigs = profileConfigs;
this.placementPolicy = placementPolicy;
this.liveCount = liveCount;
this.models = Objects.requireNonNull(models, "models");
Map<String, HerdrPeerLauncher> index = new LinkedHashMap<>();
for (HerdrPeerLauncher d : this.delegates) {
for (String profile : d.profiles()) {
@@ -303,6 +354,16 @@ public final class CompositePeerLauncher implements PeerLauncher {
return m == null ? Map.of() : m;
}
/**
* The currently-configured {@code models:} block, never null (fleetd #422). Read fresh on
* every call, the same reason {@link #profiles0} is — a config reload's on/off edit must reach
* the very next spawn.
*/
private FleetConfig.Models models0() {
FleetConfig.Models m = models.get();
return m == null ? NO_MODELS_CONFIGURED : m;
}
/** The adapter owning {@code profileName} (null/blank → the default). Throws on an unknown profile. */
private HerdrPeerLauncher route(String profileName) {
String resolved = (profileName == null || profileName.isBlank()) ? defaultProfile : profileName;
@@ -332,6 +393,7 @@ public final class CompositePeerLauncher implements PeerLauncher {
enforceNotQuarantined(requestedProfile);
enforceNotCoolingOff(requestedProfile);
enforceMaxLoad(requestedProfile);
enforceModelEnabled(requestedProfile);
PeerHandle handle = d.spawn(req);
spawnedBy.put(handle.id(), d);
return handle;
@@ -341,18 +403,10 @@ public final class CompositePeerLauncher implements PeerLauncher {
// the whole profile list. An EXPLICIT profile (above) is left alone on purpose — it is the
// operator overriding, and refusing it would break `fleet_spawn{profile:"opus"}`, which
// carries no role and so would be judged against the dev pool it was never meant for.
List<PlacementCandidate> candidates = candidates(req.role());
String roleDefault = defaultProfileFor(req.role());
Set<String> unreachable = new HashSet<>();
// CB-578 stage B: computed once up front — a quarantine's expiry cannot pass within one spawn
// call, so re-deriving it per retry would only cost work, never change the answer.
Set<String> quarantined = quarantinedProfiles(candidates);
// fleetd #201 Unit 5: a distinct set from quarantined — see PlacementContext.coolingOff.
Set<String> coolingOff = coolingOffProfiles(candidates);
PlacementContext ctx = new PlacementContext(roleDefault, candidates, liveCount, unreachable,
quarantined, coolingOff);
PlacementContext ctx = placementContextFor(req.role(), unreachable);
int maxAttempts = candidates.isEmpty() ? 1 : candidates.size();
int maxAttempts = ctx.candidates().isEmpty() ? 1 : ctx.candidates().size();
for (int attempt = 0; attempt < maxAttempts; attempt++) {
// Deliberately uncaught: when no candidate is left (all at cap, or all unreachable) the
// policy already throws a clear message. Catching it to rethrow a generic
@@ -379,8 +433,7 @@ public final class CompositePeerLauncher implements PeerLauncher {
chosen.profile(), e.getMessage());
unreachable.add(chosen.profile());
// Update the context for the next selection so the policy excludes this profile.
ctx = new PlacementContext(roleDefault, candidates, liveCount, unreachable,
quarantined, coolingOff);
ctx = placementContextFor(req.role(), unreachable);
}
}
@@ -492,6 +545,49 @@ public final class CompositePeerLauncher implements PeerLauncher {
}
}
/**
* Refuse an explicit-profile spawn whose {@code model:} the operator has turned off in the
* central {@code models.allow:} list (fleetd #422). Deliberately worded apart from {@link
* #enforceNotQuarantined} and {@link #enforceNotCoolingOff}: those two report a BACKEND-reported
* outage (exhaustion, repeated errors); this one reports an OPERATOR decision, so the message
* says "turned off" and names the model, never "quarantined" or "cooling off". A FOURTH,
* independent reason to refuse a spawn — never layered onto {@code BackendQuarantine} or {@code
* BackendOutagePolicy}, which would misattribute an operator's own choice to the backend.
*
* <p>Reads {@link #models0()} fresh on every call — the same liveness {@link #profiles0()}
* already has — so flipping {@code enabled: false} and reloading takes effect on the very next
* spawn, no restart (criterion 4). A profile naming no model, or a model absent from {@code
* models.allow:} entirely (nothing to gate against), is never refused here.
*
* @throws PlacementException naming the model and the profile, distinct from quarantine/cool-off
*/
private void enforceModelEnabled(String profile) {
FleetConfig.Profile cfg = profiles0().get(profile);
String model = (cfg == null) ? null : cfg.model();
if (model == null) {
return;
}
if (models0().offIds().contains(model)) {
throw new PlacementException("worker profile '" + profile + "' names model '" + model
+ "', which the operator has turned off in models.allow — refusing spawn");
}
}
/** The subset of {@code candidates} whose {@code model:} is currently turned off (fleetd #422). */
private Set<String> modelOffProfiles(List<PlacementCandidate> candidates) {
Set<String> off = models0().offIds();
if (off.isEmpty()) {
return Set.of();
}
return candidates.stream()
.map(PlacementCandidate::profile)
.filter(p -> {
FleetConfig.Profile cfg = profiles0().get(p);
return cfg != null && cfg.model() != null && off.contains(cfg.model());
})
.collect(Collectors.toSet());
}
/**
* The profile names {@code role} may be placed on, in definition order.
*
@@ -508,8 +604,23 @@ public final class CompositePeerLauncher implements PeerLauncher {
return known.isEmpty() ? List.copyOf(configured.keySet()) : known;
}
/** The profile an unqualified spawn for {@code role} falls back to under {@code fixed} placement. */
private String defaultProfileFor(MemberRole role) {
/**
* {@inheritDoc}
*
* <p>Live: reads {@link #poolFor}, which reads {@link #profileConfigs} and {@link #fleet} fresh
* on every call, so a config reload is visible without a restart (fleetd #425) — unlike {@link
* #defaultProfile}, the field captured once at construction, which this falls back to only when
* {@link #poolFor} has nothing to offer at all (no profiles configured for this composite).
*
* <p>Exact only under the {@code fixed} placement policy — the one that reads this value
* ({@code FixedPlacementPolicy}, package-private, hence not linked) as its first, preferred
* candidate. {@code weighted}/{@code round-robin} placement can choose a different candidate
* from {@code role}'s pool even on the very first spawn; this method does not simulate that
* choice, matching what the {@code defaultProfile:}-derived reporting this replaces has always
* done.
*/
@Override
public String defaultProfileFor(MemberRole role) {
List<String> pool = poolFor(role);
return pool.isEmpty() ? defaultProfile : pool.getFirst();
}
@@ -526,6 +637,142 @@ public final class CompositePeerLauncher implements PeerLauncher {
return out;
}
/**
* Build the {@link PlacementContext} an unqualified spawn of {@code role} would be judged
* against right now — the single source both {@link #spawn} and {@link #place} read, so the two
* can never disagree about which conditions (quarantine, cool-off, model-off) apply to which
* candidate (fleetd #425 rework: round 1 duplicated this into a second, blind resolver —
* {@link #defaultProfileFor} — which is why it regressed; round 2 found that even a single
* shared resolver is not enough on its own if the CALLER re-resolves through an explicit
* profile afterwards — see {@link PlacementDecision}).
*
* @param unreachable the caller's mutable unreachable set; {@link #spawn} grows this across
* retries and rebuilds the context from it, {@link #place} passes a fresh
* empty one since it never retries
*/
private PlacementContext placementContextFor(MemberRole role, Set<String> unreachable) {
List<PlacementCandidate> candidates = candidates(role);
String roleDefault = defaultProfileFor(role);
// CB-578 stage B: computed once up front — a quarantine's expiry cannot pass within one spawn
// call, so re-deriving it per retry would only cost work, never change the answer.
Set<String> quarantined = quarantinedProfiles(candidates);
// fleetd #201 Unit 5: a distinct set from quarantined — see PlacementContext.coolingOff.
Set<String> coolingOff = coolingOffProfiles(candidates);
// fleetd #422: read live per spawn, same as quarantined/coolingOff above — a config reload
// that flips a model's enabled state is visible to the very next unqualified spawn.
Set<String> modelOff = modelOffProfiles(candidates);
return new PlacementContext(roleDefault, candidates, liveCount, unreachable,
quarantined, coolingOff, modelOff);
}
/**
* {@inheritDoc}
*
* <p>fleetd #425 rework, round 2: runs the exact same selection {@link #spawn} uses for a
* blank-profile request — {@link #placementContextFor} plus one {@link PlacementPolicy#select}
* — rather than {@link #defaultProfileFor}'s blind "pool's first entry", so a quarantined,
* cooling-off, or model-off pool-first candidate is routed around here exactly as it would be
* by a real spawn. Unlike {@link #spawn}, this never retries on {@link
* PeerUnreachableException}: there is no spawn attempt to fail, so "unreachable" never grows
* past the empty set it starts with, and a single {@link PlacementPolicy#select} call already
* reflects the live quarantine/cool-off/model-off state.
*
* <p>Deliberately does <em>not</em> apply {@link #enforceMaxLoad} (or any of the other three
* {@code enforce*} checks): those belong to {@link #spawn}'s EXPLICIT-profile branch, the
* operator-override path, and this method answers a different question — "where would an
* UNQUALIFIED spawn land". That is not the same as {@code select} ignoring these conditions —
* every condition {@code select} filters on (quarantine, cooling off, {@code maxLoad} under
* every placement policy including the default {@code fixed}, since fleetd #435, model-off,
* unreachable, weight-0) is already reflected in the {@link PlacementDecision} this method
* returns, because {@code select} walked past every excluded candidate to find it. What this
* method's caller must not do is take that resolved name and hand it back to {@link
* #spawn(SpawnRequest)} as an explicit profile: the explicit-profile branch treats the same
* exclusion conditions as a reason to REFUSE, where {@code select} had already treated them as
* a reason to fall through — round 1 of this fix did exactly that, turning a fall-through this
* method had already resolved around into a refusal one call later. Round 2 fixes that at the
* caller: {@link #spawn(SpawnRequest, PlacementDecision)} carries this exact decision to the
* spawn without re-resolving or re-checking it, through the same routing path {@code select}
* itself was consulted from.
*
* @throws PlacementException if no candidate in {@code role}'s pool is currently placeable
* (mirrors what an actual unqualified spawn would throw)
*/
@Override
public PlacementDecision place(MemberRole role) {
PlacementContext ctx = placementContextFor(role, new HashSet<>());
return new PlacementDecision(placementPolicy.get().select(ctx).profile());
}
/**
* {@inheritDoc}
*
* <p>Delegates to {@link #place}, so the two can never disagree about the answer for the same
* {@code role} at the same instant — kept as a convenience for a caller that only wants the
* resolved name (a status report, a log line), never for a caller that will act on it by
* spawning: that caller must hold the {@link PlacementDecision} itself and pass it to {@link
* #spawn(SpawnRequest, PlacementDecision)} — see {@link PlacementDecision}'s javadoc for why
* resolving here and spawning separately, with the name fed back in as an explicit profile,
* regressed fleetd #425 twice.
*/
@Override
public String routedProfileFor(MemberRole role) {
return place(role).profile();
}
/**
* {@inheritDoc}
*
* <p>Routes {@code decision.profile()} directly to its owning delegate — the identical
* {@code d.spawn(routedReq)} call {@link #spawn(SpawnRequest)}'s blank-profile branch makes for
* its first pick — WITHOUT re-running {@link #enforceNotQuarantined}, {@link
* #enforceNotCoolingOff}, {@link #enforceMaxLoad}, or {@link #enforceModelEnabled}: those are
* the EXPLICIT-profile branch's checks, and {@code decision} did not come from an operator
* naming a profile — it came from {@link #place}, which already applied whichever of these
* conditions {@link PlacementPolicy#select} actually filters on (fleetd #425 rework, round 2).
*
* <p>The two branches disagree on purpose about what an excluded profile means, and that
* disagreement is not what this method removes. The blank-profile routing branch (and
* {@link #place}) treats a quarantined/cooling-off/at-cap/model-off/unreachable/weight-0 profile
* as a reason to fall through to the next candidate; the EXPLICIT-profile branch treats naming
* that same profile as a reason to refuse outright — someone who names a profile should get a
* refusal, not a silent substitution onto a different backend. That is still correct after
* fleetd #435. What round 1 got wrong, and what this method exists to stop happening again, is
* turning a fall-through into a refusal by accident: resolving a name via {@link #place} and
* then handing that same name back to {@link #spawn(SpawnRequest)} as an explicit profile takes
* the refusing branch on a decision the routing branch had already approved by falling through
* past everything else.
*
* <p>Before fleetd #435, this exact accident was reachable through {@code maxLoad} specifically:
* {@code FixedPlacementPolicy} — the default policy — did not evaluate {@code maxLoad} at all
* for automatic selection, so {@link #place} could approve an at-cap profile that {@link
* #enforceMaxLoad} would then refuse one call later. fleetd #435 closed that: {@code
* FixedPlacementPolicy} now walks past an at-cap candidate exactly like {@code weighted}/
* {@code round-robin} already did, so {@link #place} can no longer return one, and this specific
* failure — an approved placement dying at {@code enforceMaxLoad} — cannot happen any more.
* What this method still buys, now that {@code maxLoad} can no longer cause it: it never
* re-evaluates a condition {@link #place} already decided, and it closes the window between
* that decision and the spawn in which the underlying state (another spawn landing on the same
* profile, a config reload) could otherwise move and make a stale explicit re-check wrong.
*
* <p>Deliberately does not retry on {@link PeerUnreachableException} across candidates the way
* {@link #spawn(SpawnRequest)}'s blank-profile branch does: retrying here would silently
* re-place the caller onto a different profile than the one {@code decision} named, behind the
* back of a caller that may already have provisioned something (a worktree's {@code repoRoot},
* parity overlay) specifically for that name. A caller that wants the composite's own failover
* should call {@link #spawn(SpawnRequest)} with a blank profile directly, not resolve through
* {@link #place} first. Losing that retry on a resolve-then-spawn path is an accepted, unrelated
* cost — see {@code SessionManager.acquireWithWorktree}'s own comment on it — never widened by
* this round to include {@code maxLoad}, which is what round 1 actually lost.
*/
@Override
public PeerHandle spawn(SpawnRequest req, PlacementDecision decision) {
HerdrPeerLauncher d = route(decision.profile());
SpawnRequest routedReq = req.withProfile(decision.profile());
PeerHandle handle = d.spawn(routedReq);
spawnedBy.put(handle.id(), d);
return handle;
}
@Override
public String effectiveCwd(SpawnRequest req) {
return route(req.profileName()).effectiveCwd(req);
@@ -536,6 +783,32 @@ public final class CompositePeerLauncher implements PeerLauncher {
return route(profileName).parityOverlay(profileName);
}
/**
* fleetd #422 follow-up: the single live read that answers both "is the models.allow: gate
* armed" and "which models are off", off the exact same accessor ({@link #models0()}) {@link
* #enforceModelEnabled} and {@link #modelOffProfiles} read — so {@code fleet_profiles}/{@code
* GET /profiles} (via {@link PeerLauncher#disabledModels()}, which now delegates here) can
* never report a different answer than the gate enforces (the fleetd #404 lesson), and the
* startup log line built from this can never disagree with either.
*
* <p>{@link #models0()} itself normalizes a {@code null} {@link #models} read to the shared
* {@link #NO_MODELS_CONFIGURED} sentinel — deliberately the one object no config-supplied
* {@code Models} instance can ever be identical to, since it is private to this class — so
* comparing by reference here recovers exactly the fact {@code models0()}'s normalization
* would otherwise erase: whether the live source was {@code null} (no {@code models:} block,
* armed = false) or a real, config-supplied block (armed = true, even one whose {@code allow:}
* is itself empty or absent — {@link FleetConfig.Models}'s "absent or empty allow: is off"
* wording governs config-load validation, a distinct question from whether this gate is armed
* for reporting).
*/
@Override
public PeerLauncher.ModelGateState modelGateState() {
FleetConfig.Models m = models0();
return m == NO_MODELS_CONFIGURED
? PeerLauncher.ModelGateState.notConfigured()
: PeerLauncher.ModelGateState.armed(m.offIds());
}
@Override
public void stop(String id) {
HerdrPeerLauncher d = spawnedBy.get(id);
@@ -647,9 +920,21 @@ public final class CompositePeerLauncher implements PeerLauncher {
return byProfile.keySet();
}
/**
* {@inheritDoc}
*
* <p>fleetd #425: reports the <em>live</em> {@code dev} pool's first entry — the same value
* {@link #defaultProfileFor} computes for {@link MemberRole#DEV} — not the {@link
* #defaultProfile} field captured at construction. An unqualified {@code fleet_spawn} defaults
* to {@code MemberRole#DEV} (see {@link dev.ltms.fleet.peer.SpawnRequest}), so "the dev pool's
* live first entry" is exactly the profile such a spawn actually lands on right now — the
* question {@code fleet_profiles}' {@code "default"} field exists to answer. The frozen field is
* a role-agnostic fallback used only when {@link #poolFor} has nothing to report at all (no
* profiles configured), which {@link #defaultProfileFor} already handles.
*/
@Override
public String defaultProfile() {
return defaultProfile;
return defaultProfileFor(MemberRole.DEV);
}
/**
@@ -42,6 +42,18 @@ public interface LeadChannel {
/** This daemon's own lead coordination id — the mailbox it owns, and the {@code from} it sends as. */
String selfCoordId();
/**
* Whether a message sitting in {@link #peek}'s held set (fetched but not yet {@link #ack}ed) is
* still safe if this daemon crashes or restarts right now — the conclusion of two independent
* facts about how this channel owns its own queue: the queue was declared <em>durable</em>, and
* the consumer that filled {@code held} uses <em>manual ack</em>, so an unacked delivery is still
* owned by the broker rather than only in this process's memory. Both must hold for {@code true};
* an implementation must derive this from what it actually did when it declared and consumed its
* queue, never return a literal — fleetd #440 found {@code FleetMcp}'s {@code heldDurable} field
* doing exactly that, unable to ever report {@code false} even after the fact stopped being true.
*/
boolean heldDurable();
/**
* A non-destructive look at {@code coordId}'s mailbox — does it exist, how many messages are
* waiting on it, and how many consumers are attached — without owning, consuming, or otherwise
@@ -86,6 +86,12 @@ public final class LeadMailbox implements LeadChannel, AutoCloseable {
private final Object channelLock = new Object();
/** msgId → held delivery, for this mailbox's own queue only (there is exactly one). */
private final LinkedHashMap<String, Held> held = new LinkedHashMap<>();
/**
* fleetd #440: the answer to {@link #heldDurable()}, set once by {@link #own()} from the exact
* booleans it passed to {@code queueDeclare}/{@code basicConsume} — never a separate literal that
* could drift from what those calls actually did.
*/
private boolean heldDurable;
/** Successful broker acks on this connection, retained only to make a repeated caller ack quiet. */
private final LinkedHashMap<String, Boolean> recentlyAcked = new LinkedHashMap<>();
/** Bounds {@link #recentlyAcked}: it is only an idempotency aid, never delivery state. */
@@ -191,13 +197,22 @@ public final class LeadMailbox implements LeadChannel, AutoCloseable {
/** Declare + consume this daemon's own {@code lead.<selfCoordId>.inbox}. Called once, at construction. */
private void own() throws IOException {
String queue = queueName(selfCoordId);
boolean durableQueue = true; // durable, non-exclusive, keep on idle
boolean autoAck = false; // manual ack
synchronized (channelLock) {
channel.queueDeclare(queue, true, false, false, null); // durable, non-exclusive, keep on idle
channel.basicConsume(queue, false, deliverCallback(), _ -> { }); // autoAck=false: manual ack
channel.queueDeclare(queue, durableQueue, false, false, null);
channel.basicConsume(queue, autoAck, deliverCallback(), _ -> { });
}
// fleetd #440: held mail is durable only while both hold — a durable queue AND manual ack.
this.heldDurable = durableQueue && !autoAck;
log.debug("lead mailbox owns queue {} for coord-id {}", queue, selfCoordId);
}
@Override
public boolean heldDurable() {
return heldDurable;
}
/**
* Publish {@code msg} to {@code toCoordId}'s mailbox and block until the broker's publisher
* confirm for it lands. Does <em>not</em> imply owning or consuming {@code toCoordId}'s queue.
@@ -1,5 +1,7 @@
package dev.ltms.fleet.peer;
import dev.ltms.fleet.placement.PlacementDecision;
import java.nio.file.Path;
import java.util.List;
import java.util.Set;
@@ -143,9 +145,135 @@ public interface PeerLauncher {
/**
* The profile a no-argument {@link #spawn(SpawnRequest)} uses, or {@code null} if none is configured.
*
* <p>fleetd #425: for an implementation with role pools (a no-argument spawn is read as {@link
* MemberRole#DEV}, see {@link SpawnRequest}), this must be the profile a live spawn of that role
* would actually be placed on right now, not a value captured once at startup — a caller such as
* {@code fleet_profiles} relies on this to report a live, not frozen, fact.
*/
String defaultProfile();
/**
* The profile an unqualified spawn of {@code role} would resolve to right now — the role-aware,
* live counterpart of {@link #defaultProfile()} (fleetd #425).
*
* <p>A caller that must provision something profile-specific (working directory, parity overlay
* files) <em>before</em> the actual spawn — {@code SessionManager.acquireWithWorktree} is the one
* that exists today — needs the exact profile that spawn will use, for the caller's real role,
* not a role-agnostic guess. Calling {@link #defaultProfile()} for that purpose reads {@code
* MemberRole#DEV}'s answer regardless of the caller's actual role, which is wrong for any other
* role and can provision for a profile the spawn never lands on.
*
* <p>Default implementation returns {@link #defaultProfile()}, ignoring {@code role} — the right
* answer for a launcher with no role-pool concept of its own (e.g. a single {@code
* HerdrPeerLauncher} adapter, which is never reached this way in production: {@code
* CompositePeerLauncher} always fronts it and resolves roles itself).
*/
default String defaultProfileFor(MemberRole role) {
return defaultProfile();
}
/**
* The profile an <em>unqualified</em> spawn of {@code role} would actually be routed to right
* now — the same candidate list, the same {@code quarantined}/{@code coolingOff}/{@code
* modelOff} filtering, and the same {@code PlacementPolicy} that {@link #spawn} itself
* consults for a blank-profile request (fleetd #425 rework).
*
* <p>This is <em>not</em> {@link #defaultProfileFor}: that method answers "what is first in
* {@code role}'s pool", blind to quarantine, cool-off, and the model on/off gate — the right
* answer for a role-agnostic, best-effort report ({@code fleet_profiles}' {@code "default"}
* field), but the wrong one for a caller that needs the profile a spawn will actually land on.
* A quarantined or model-off pool-first profile makes {@link #defaultProfileFor} return a name
* an unqualified spawn will never be routed to.
*
* <p>Just the resolved name, not the full {@link PlacementDecision} — a caller that only wants
* to know the answer (a status report, a log line) can call this; a caller that will later
* <em>act</em> on the answer by spawning — provisioning a worktree for a specific profile
* before the peer exists is the one that matters — must call {@link #place} and carry the
* {@link PlacementDecision} itself through to {@link #spawn(SpawnRequest, PlacementDecision)}
* instead of calling this method and feeding the string back in as an explicit profile. Doing
* that re-enters {@link #spawn(SpawnRequest)}'s explicit-profile branch, which disagrees with
* the routing branch on purpose about what an excluded profile means: the routing branch (and
* {@link #place}) falls through a quarantined/cooling-off/at-cap/model-off/unreachable/weight-0
* profile to the next candidate, while the explicit branch refuses outright — correct for an
* operator who named that profile on purpose, wrong for a name that only ever came from placement
* itself. That accidental refusal is exactly the regression fleetd #425 rework round 2 fixes:
* the default implementation below delegates to {@link #place}, so the two can never drift apart,
* but a caller that resolves through this method alone and spawns separately can still recreate
* the round-1 defect for itself. (Before fleetd #435, this accident was also reachable through
* {@code maxLoad} specifically, because {@code FixedPlacementPolicy} — the default policy — did
* not evaluate it at all for automatic selection; #435 closed that gap, so a placement decision
* can no longer be at cap in the first place. The refusal-vs-fall-through disagreement above is
* the part that was never about {@code maxLoad} and is still real.)
*
* @throws RuntimeException (implementation-specific, typically a placement exception) if no
* candidate in {@code role}'s pool is currently placeable
*/
default String routedProfileFor(MemberRole role) {
return place(role).profile();
}
/**
* Resolve, <em>without spawning</em>, the {@link PlacementDecision} an unqualified spawn of
* {@code role} would make right now — the same candidate list, the same {@code
* quarantined}/{@code coolingOff}/{@code modelOff} filtering, and the same {@code
* PlacementPolicy} {@link #spawn(SpawnRequest)}'s blank-profile branch itself consults (fleetd
* #425 rework).
*
* <p>Pair this with {@link #spawn(SpawnRequest, PlacementDecision)}, never with {@link
* #spawn(SpawnRequest)} fed the decision's profile as an explicit name — see {@link
* PlacementDecision}'s own javadoc for why that second form regressed.
*
* <p>Default implementation wraps {@link #defaultProfile()}, ignoring {@code role} and every
* placement condition — the right answer for a launcher with no pool or placement-policy
* concept of its own, matching {@link #defaultProfileFor}'s own default.
*
* @throws RuntimeException (implementation-specific, typically a placement exception) if no
* candidate in {@code role}'s pool is currently placeable
*/
default PlacementDecision place(MemberRole role) {
return new PlacementDecision(defaultProfile());
}
/**
* Spawn against an already-resolved {@link PlacementDecision} from {@link #place}, honoring it
* completely: none of the conditions {@link #place} already applied — quarantine, cooling off,
* {@code maxLoad} (evaluated by every placement policy including the default {@code fixed},
* since fleetd #435), model-off — are re-evaluated here; {@code decision} already reflects them.
* This is not skipping a check {@code place} left undone; it is not repeating one {@code place}
* already did, and not re-opening the window between that decision and this spawn in which the
* underlying state could otherwise move. This is what lets a resolve-then-spawn caller
* ({@code SessionManager.acquireWithWorktree}, which must know the profile before it can
* provision a worktree for it) and a plain blank-profile {@link #spawn(SpawnRequest)} caller
* land on the exact same outcome for the exact same placement state (fleetd #425 rework,
* round 2).
*
* <p>{@code req}'s own {@link SpawnRequest#profileName()} is ignored in favor of {@code
* decision.profile()} — the caller is expected to have built {@code req} with a blank or
* matching profile; passing a request that names a <em>different</em>, explicit profile than
* the decision it is paired with is a caller bug this method does not attempt to detect.
*
* <p>Default implementation for a launcher with no placement concept of its own: delegates to
* {@link #spawn(SpawnRequest)} with the decision's profile named explicitly — its only spawn
* contract, since there is no separate routing path to honor. This default is correct ONLY for
* a launcher that spawns a single profile of its own (e.g. {@code HerdrPeerLauncher}), where
* the explicit-profile branch it re-enters and the routing branch {@link #place} would have
* used are the same thing. <strong>A launcher that routes across more than one profile — the
* way {@code CompositePeerLauncher} routes across every configured adapter — MUST override
* this method instead of inheriting this default.</strong> Re-entering {@link
* #spawn(SpawnRequest)} re-applies that single-argument method's explicit-profile checks
* ({@code enforceNotQuarantined}, {@code enforceNotCoolingOff}, {@code enforceMaxLoad}, {@code
* enforceModelEnabled} in {@code CompositePeerLauncher}), which can refuse the very profile
* {@link #place} just chose, if the underlying placement state moved in the window between the
* {@link #place} call and this one — the exact window this method and {@link PlacementDecision}
* exist to close (fleetd #444).
*
* @throws IllegalArgumentException if the decision names an unknown profile
*/
default PeerHandle spawn(SpawnRequest req, PlacementDecision decision) {
return spawn(req.withProfile(decision.profile()));
}
/**
* Resolve the effective working directory for a spawn {@code req} without actually spawning.
* Resolution order: requestedCwd → profile cwd → callerCwd → daemon cwd.
@@ -192,4 +320,66 @@ public interface PeerLauncher {
* @return {@code true} when a reset was sent and its status transition must settle before reuse
*/
boolean clearContext(String id);
/**
* Model ids the operator has currently turned off in the central {@code models.allow:} list
* (fleetd #422) — empty for a launcher with nothing to gate against. {@code fleet_profiles}/
* {@code GET /profiles} (via {@code FleetMcp.profilesView}) call this to report which models
* are off, and MUST read this exact accessor rather than deriving their own answer: the fleetd
* #404 lesson is that a status field reading a different source than the behaviour it describes
* can drift from what the gate ({@code CompositePeerLauncher.enforceModelEnabled} and its
* candidate filter) actually enforces. A default of {@code Set.of()} keeps every other {@link
* PeerLauncher} implementer (the herdr adapters, and the two test-fake implementers) unchanged.
*
* <p>fleetd #422 follow-up: this alone cannot tell "no {@code models:} block at all" from "a
* {@code models:} block where nothing is currently off" — both report an empty set here. Delegates
* to {@link #modelGateState()} so the two facts always come from the one read {@link
* #modelGateState()}'s implementer makes; do not override this method separately from that one.
*/
default Set<String> disabledModels() {
return modelGateState().off();
}
/**
* Whether the central {@code models.allow:} gate (fleetd #422) is armed at all, together with
* which model ids are currently off — fleetd #422 follow-up. {@link #disabledModels()} alone
* cannot distinguish two states that both report an empty set: a host with no {@code models:}
* block (nothing is gated, and nothing can be) and a host WITH a {@code models:} block where
* nothing is currently turned off (the gate is armed and reporting zero). This method exists so
* a caller — the startup log, {@code fleet_profiles}/{@code GET /profiles} — can tell the two
* apart, the same reason {@code CompletionResolver.UnsetMeaning} exists: an accessor that can
* legitimately report "empty" must never let a caller guess why.
*
* <p>Default {@link ModelGateState#notConfigured()} — every launcher without a {@code models:}
* block to read from (the herdr adapters, and the two test-fake implementers), matching {@link
* #disabledModels()}'s own default of an empty set.
*/
default ModelGateState modelGateState() {
return ModelGateState.notConfigured();
}
/**
* fleetd #422 follow-up: the result of {@link #modelGateState()} — see that method's javadoc
* for why "armed" and "off" must be reported together from one read rather than as two
* separately-derived facts that a reload landing between them could make disagree.
*
* @param configured {@code true} when a {@code models:} block exists at all (armed), regardless
* of whether anything in it is currently turned off; {@code false} when there
* is no block to gate against
* @param off the model ids currently turned off; always empty when {@code configured} is
* {@code false}
*/
record ModelGateState(boolean configured, Set<String> off) {
public ModelGateState {
off = Set.copyOf(off);
}
public static ModelGateState notConfigured() {
return new ModelGateState(false, Set.of());
}
public static ModelGateState armed(Set<String> off) {
return new ModelGateState(true, off);
}
}
}
@@ -5,14 +5,12 @@ import java.util.List;
/**
* Backward-compatible placement: an unqualified spawn always resolves to the configured default
* profile, exactly as {@code CompositePeerLauncher} did before CB-518. This ignores caps
* ({@code maxLoad}) so that a pre-existing config behaves identically after upgrade — capacity
* gating for automatic placement is deliberately out of scope for {@code fixed}, exactly as it
* always has been. Reachability is a narrower exception (fleetd #315, below): a profile is never
* checked for reachability up front, only skipped once it has already failed in <em>this same</em>
* spawn call's retry loop — see the unreachable case below.
* profile, exactly as {@code CompositePeerLauncher} did before CB-518. Reachability is a narrower
* exception (fleetd #315, below): a profile is never checked for reachability up front, only
* skipped once it has already failed in <em>this same</em> spawn call's retry loop — see the
* unreachable case below.
*
* <p>Four exceptions walk past the default instead of returning it unconditionally:
* <p>Six exceptions walk past the default instead of returning it unconditionally:
* <ul>
* <li>Quarantine (CB-578 stage B): a quarantined default is a credential that just refused on
* a usage limit, not a transient capacity or reachability concern.
@@ -20,6 +18,24 @@ import java.util.List;
* ({@code BackendOutagePolicy}) — a separate, shorter-lived source from quarantine. When a
* profile is both quarantined and cooling off, only the quarantine reason is reported
* (exhaustion takes priority), matching {@code CompositePeerLauncher}'s explicit-spawn order.
* <li>At cap (fleetd #435): a profile whose live count has reached its {@code maxLoad}
* ({@link PlacementPolicyUtil#atCap}) — a documented, unconditional capacity limit (see
* {@code FleetConfig.Profile#maxLoad}), so {@code fixed} must gate on it exactly as {@code
* weighted}/{@code round-robin} already do via {@link PlacementPolicyUtil#available}. Before
* this fix {@code fixed} built its own {@link PlacementCandidate} for the default with {@code
* maxLoad} forced to {@code null}, so a capped default was chosen anyway on every unqualified
* spawn — the cap was advisory, not enforced, for the one placement policy every config uses
* by default. Reported only when quarantine and cooling off are both absent, matching {@code
* CompositePeerLauncher}'s explicit-spawn check order (quarantine, then cooling off, then max
* load, then model-off).
* <li>Model off (fleetd #422): a profile whose {@code model:} the operator has turned off in
* {@code models.allow:} — an operator decision, never a backend-reported outage, so it is a
* fifth, independent source from quarantine, cooling off, and at-cap (never merged with any
* of them), exactly as {@code CompositePeerLauncher.enforceModelEnabled} and {@link
* PlacementPolicyUtil#available} treat it. When a profile is model-off <em>and</em> quarantined,
* cooling off, or at cap, only the higher-priority reason is reported, matching {@code
* CompositePeerLauncher}'s explicit-spawn check order (quarantine, then cooling off, then max
* load, then model-off).
* <li>Unreachable (fleetd #315): {@code CompositePeerLauncher.spawn} retries a failed candidate
* on the next one and rebuilds the {@link PlacementContext} so {@code ctx.unreachable()}
* names every profile that already failed with {@code PeerUnreachableException} in this same
@@ -32,9 +48,10 @@ import java.util.List;
* {@code weighted}/{@code round-robin} skip it — an explicit {@code fleet_spawn} naming
* the profile is unaffected, only this automatic fallback walk.
* </ul>
* A fleet where nothing is ever quarantined, cooling off, unreachable, or weight-0 never exercises
* any of these paths, so today's behaviour is unchanged — in particular, the very first selection
* of a spawn call always sees an empty {@code unreachable} set, so the first choice is untouched.
* A fleet where nothing is ever quarantined, cooling off, at cap, model-off, unreachable, or
* weight-0 never exercises any of these paths, so today's behaviour is unchanged — in particular,
* the very first selection of a spawn call always sees an empty {@code unreachable} set, so the
* first choice is untouched.
*/
final class FixedPlacementPolicy implements PlacementPolicy {
@@ -42,12 +59,14 @@ final class FixedPlacementPolicy implements PlacementPolicy {
public PlacementCandidate select(PlacementContext ctx) {
String d = ctx.defaultProfile();
if (d != null && !d.isBlank() && !ctx.quarantined().contains(d) && !ctx.coolingOff().contains(d)
&& !ctx.unreachable().contains(d) && !weightExcluded(ctx, d)) {
&& !ctx.modelOff().contains(d) && !ctx.unreachable().contains(d) && !weightExcluded(ctx, d)
&& !capExcluded(ctx, d)) {
return new PlacementCandidate(d, null, 1.0f, null);
}
for (PlacementCandidate c : ctx.candidates()) {
if (!ctx.quarantined().contains(c.profile()) && !ctx.coolingOff().contains(c.profile())
&& !ctx.unreachable().contains(c.profile()) && !c.excluded()) {
&& !ctx.modelOff().contains(c.profile()) && !ctx.unreachable().contains(c.profile())
&& !c.excluded() && !PlacementPolicyUtil.atCap(ctx, c)) {
return new PlacementCandidate(c.profile(), null, c.weight(), c.maxLoad());
}
}
@@ -56,9 +75,18 @@ final class FixedPlacementPolicy implements PlacementPolicy {
// Exhaustion quarantine takes priority: reported only when quarantine is absent, so the
// message never claims "cooling off" for a profile that is really backend-exhausted.
boolean dCoolingOff = !dQuarantined && ctx.coolingOff().contains(d);
// fleetd #435: at-cap sits between cooling off and model-off, matching
// CompositePeerLauncher's explicit-spawn check order (quarantine, cooling off, max load,
// then model-off) — reported only when quarantine/cooling-off are both absent.
boolean dAtCap = !dQuarantined && !dCoolingOff && capExcluded(ctx, d);
// fleetd #422: model-off is a fifth, independent source (an operator decision) — but
// quarantine/cooling-off/at-cap still take priority when more than one applies, matching
// CompositePeerLauncher's explicit-spawn check order (quarantine, cooling off, max load,
// then model-off).
boolean dModelOff = !dQuarantined && !dCoolingOff && !dAtCap && ctx.modelOff().contains(d);
boolean dUnreachable = ctx.unreachable().contains(d);
boolean dWeightExcluded = weightExcluded(ctx, d);
if (dQuarantined || dCoolingOff || dUnreachable || dWeightExcluded) {
if (dQuarantined || dCoolingOff || dAtCap || dModelOff || dUnreachable || dWeightExcluded) {
List<String> reasons = new ArrayList<>();
if (dQuarantined) {
reasons.add("is quarantined (backend exhausted)");
@@ -66,6 +94,14 @@ final class FixedPlacementPolicy implements PlacementPolicy {
if (dCoolingOff) {
reasons.add("is cooling off after repeated backend errors");
}
if (dAtCap) {
PlacementCandidate c = candidateFor(ctx, d);
int live = ctx.liveCount().apply(d);
reasons.add("is at maxLoad (" + live + " live >= " + c.maxLoad() + " cap)");
}
if (dModelOff) {
reasons.add("names a model the operator has turned off in models.allow");
}
if (dUnreachable) {
reasons.add("is unreachable");
}
@@ -78,18 +114,38 @@ final class FixedPlacementPolicy implements PlacementPolicy {
}
if (!ctx.candidates().isEmpty()) {
throw new PlacementException("all worker profiles are excluded from automatic "
+ "selection (quarantined, cooling off, unreachable, or weight-0)");
+ "selection (quarantined, cooling off, at cap, model-off, unreachable, or weight-0)");
}
throw new PlacementException("no worker profiles configured");
}
/** Whether {@code profile} carries {@code weight <= 0} (CB-554) among {@code ctx}'s candidates. */
private static boolean weightExcluded(PlacementContext ctx, String profile) {
PlacementCandidate c = candidateFor(ctx, profile);
return c != null && c.excluded();
}
/**
* Whether {@code profile} has reached its {@code maxLoad} cap (fleetd #435), using the shared
* {@link PlacementPolicyUtil#atCap} definition — the same one {@code weighted}/{@code
* round-robin} already consult via {@link PlacementPolicyUtil#available}. Looked up by name,
* the same way {@link #weightExcluded} is: the default fast path above builds its own {@link
* PlacementCandidate} with {@code maxLoad} forced to {@code null} (it carries no cap of its
* own), so the candidate actually configured for {@code profile} has to be found in {@code
* ctx.candidates()} first.
*/
private static boolean capExcluded(PlacementContext ctx, String profile) {
PlacementCandidate c = candidateFor(ctx, profile);
return c != null && PlacementPolicyUtil.atCap(ctx, c);
}
/** The configured candidate named {@code profile} in {@code ctx}, or {@code null} if none. */
private static PlacementCandidate candidateFor(PlacementContext ctx, String profile) {
for (PlacementCandidate c : ctx.candidates()) {
if (c.profile().equals(profile)) {
return c.excluded();
return c;
}
}
return false;
return null;
}
}
@@ -22,11 +22,29 @@ import java.util.function.Function;
* the two apart so its refusal message says "cooling off", not "exhausted",
* when only this one is active. A profile can be in both sets at once; when it
* is, exhaustion quarantine is reported (it takes priority).
* @param modelOff profiles whose {@code model:} is currently turned off in {@code
* models.allow:} (fleetd #422) — an operator decision, not a backend-reported
* outage, so a SEPARATE, independent source from both {@code quarantined} and
* {@code coolingOff}. A profile can be in this set together with either (or
* both) of the others; {@link PlacementPolicyUtil} counts it into its own
* bucket rather than merging it into theirs, the same reason
* {@code coolingOff} is kept apart from {@code quarantined}.
*/
public record PlacementContext(String defaultProfile,
List<PlacementCandidate> candidates,
Function<String, Integer> liveCount,
Set<String> unreachable,
Set<String> quarantined,
Set<String> coolingOff) {
Set<String> coolingOff,
Set<String> modelOff) {
/**
* Back-compat form before the fleetd #422 model on/off gate was added — no candidate's model
* is off. Keeps pre-#422 call sites (tests included) compiling and behaving identically.
*/
public PlacementContext(String defaultProfile, List<PlacementCandidate> candidates,
Function<String, Integer> liveCount, Set<String> unreachable,
Set<String> quarantined, Set<String> coolingOff) {
this(defaultProfile, candidates, liveCount, unreachable, quarantined, coolingOff, Set.of());
}
}
@@ -0,0 +1,49 @@
package dev.ltms.fleet.placement;
import dev.ltms.fleet.peer.MemberRole;
import dev.ltms.fleet.peer.PeerLauncher;
import dev.ltms.fleet.peer.SpawnRequest;
/**
* An already-completed placement choice — the outcome of one {@link PeerLauncher#place} call,
* carried forward so a later {@link PeerLauncher#spawn(SpawnRequest, PlacementDecision)} can honor
* it directly instead of re-resolving the profile a second time (fleetd #425 rework, round 2).
*
* <p>The problem this exists to close: a caller that must know the profile <em>before</em> it can
* spawn — {@code SessionManager.acquireWithWorktree} provisions a worktree's {@code repoRoot} and
* parity overlay for a specific profile before the peer process exists — used to resolve that name
* with {@code PeerLauncher.routedProfileFor(role)} and then hand the SAME string back to {@link
* PeerLauncher#spawn(SpawnRequest)} as an EXPLICIT profile. That re-resolution is not free: naming
* a profile explicitly makes {@code CompositePeerLauncher.spawn} take its THROWING branch
* ({@code enforceNotQuarantined}/{@code enforceNotCoolingOff}/{@code enforceMaxLoad}/{@code
* enforceModelEnabled}), while an unqualified spawn's ROUTING branch never runs those checks at
* all — it instead FALLS THROUGH to the next candidate on exactly the same conditions the throwing
* branch refuses on. That disagreement is deliberate: an operator who names a profile should get a
* refusal, not a silent substitution. The bug is turning the fall-through into a refusal by
* accident — resolving a name through the routing side and then re-entering the refusing side with
* it, for a decision the routing side had already approved by walking past everything else.
* Before fleetd #435, this accident was also reachable through {@code maxLoad} specifically: the
* default {@code fixed} placement policy did not evaluate {@code maxLoad} at all for automatic
* selection, so a profile placement itself just approved could still die at {@code enforceMaxLoad}
* one call later, purely because the caller's route to the spawn passed through an explicit
* profile name instead of the routing branch — a failure a worktree-less unqualified spawn would
* never hit. fleetd #435 closed that specific gap ({@code fixed} now evaluates {@code maxLoad}
* exactly like every other placement policy), so a {@link PlacementDecision} can no longer be
* at-cap in the first place — but the refusal-vs-fall-through disagreement above was never about
* {@code maxLoad}, and resolving a name and re-entering the refusing branch with it is still wrong
* for every OTHER condition placement filters on.
*
* <p>{@link PeerLauncher#spawn(SpawnRequest, PlacementDecision)} closes that by spawning through
* the identical code path the routing branch itself uses, keyed off the SAME decision {@link
* PeerLauncher#place} returned — no re-checking of any condition placement already evaluated. A
* resolve-then-spawn caller and a blank-profile {@link PeerLauncher#spawn(SpawnRequest)} caller can
* then never disagree about which conditions apply to the same placement state, and neither one
* re-opens the window between the placement decision and the spawn in which the underlying state
* could otherwise move.
*
* @param profile the profile this decision resolved to (may be {@code null} only when no profile is
* configured at all — the same corner case {@link PeerLauncher#defaultProfile()}
* already tolerates)
*/
public record PlacementDecision(String profile) {
}
@@ -11,28 +11,40 @@ final class PlacementPolicyUtil {
private PlacementPolicyUtil() {
}
/**
* True when {@code c} has reached its {@code maxLoad} cap: {@code liveCount(c.profile()) >=
* c.maxLoad()}. A {@code null} maxLoad means unlimited, so it is never at cap.
*
* <p>Extracted as the single shared definition of "at cap" (fleetd #435): before this fix it
* was computed inline in both {@link #available} and {@link #emptyException}, and {@code
* FixedPlacementPolicy} — not a caller of either — quietly kept its own {@code select} free of
* any cap check at all, so a capped default profile was chosen anyway under the default
* placement policy. Every automatic policy must call this, not re-derive it.
*/
static boolean atCap(PlacementContext ctx, PlacementCandidate c) {
Integer cap = c.maxLoad();
return cap != null && ctx.liveCount().apply(c.profile()) >= cap;
}
/**
* Candidates that are not weight-excluded (CB-554: explicit {@code weight <= 0}, checked
* first because it is a static config choice rather than transient state), not
* known-unreachable, not quarantined (CB-578 stage B), not cooling off after repeated backend
* errors (fleetd #201 Unit 5 — a separate, shorter-lived source from quarantine), and have not
* reached their maxLoad. A {@code null} maxLoad means unlimited.
* errors (fleetd #201 Unit 5 — a separate, shorter-lived source from quarantine), not naming a
* model the operator has turned off (fleetd #422 — a third, independent source: an operator
* decision, never a backend-reported outage), and have not reached their maxLoad (see {@link
* #atCap}). A {@code null} maxLoad means unlimited.
*/
static List<PlacementCandidate> available(PlacementContext ctx) {
List<PlacementCandidate> out = new ArrayList<>();
for (PlacementCandidate c : ctx.candidates()) {
if (c.excluded() || ctx.unreachable().contains(c.profile())
|| ctx.quarantined().contains(c.profile())
|| ctx.coolingOff().contains(c.profile())) {
|| ctx.coolingOff().contains(c.profile())
|| ctx.modelOff().contains(c.profile())
|| atCap(ctx, c)) {
continue;
}
Integer cap = c.maxLoad();
if (cap != null) {
int live = ctx.liveCount().apply(c.profile());
if (live >= cap) {
continue;
}
}
out.add(c);
}
return out;
@@ -40,12 +52,13 @@ final class PlacementPolicyUtil {
/**
* Build a clear exception describing why every candidate was dropped: all weight-0, all
* quarantined, all cooling off, all at capacity, all unreachable, or a mix. Each candidate is
* counted into exactly one bucket (weight-excluded first, then quarantined, then cooling off)
* so a candidate excluded for more than one reason is never double-counted — a candidate that is
* both quarantined (CB-578 stage B, backend exhausted) and cooling off (fleetd #201 Unit 5,
* repeated backend errors) counts only as quarantined, matching {@code CompositePeerLauncher}'s
* explicit-spawn ordering: exhaustion quarantine takes priority when both are active.
* quarantined, all cooling off, all model-off, all at capacity, all unreachable, or a mix. Each
* candidate is counted into exactly one bucket (weight-excluded first, then quarantined, then
* cooling off, then model-off) so a candidate excluded for more than one reason is never
* double-counted — a candidate that is both quarantined (CB-578 stage B, backend exhausted) and
* cooling off (fleetd #201 Unit 5, repeated backend errors) counts only as quarantined, matching
* {@code CompositePeerLauncher}'s explicit-spawn ordering: exhaustion quarantine takes priority
* when more than one applies.
*/
static PlacementException emptyException(PlacementContext ctx) {
int weightExcluded = 0;
@@ -53,17 +66,19 @@ final class PlacementPolicyUtil {
int unreachable = 0;
int quarantined = 0;
int coolingOff = 0;
int modelOff = 0;
for (PlacementCandidate c : ctx.candidates()) {
Integer cap = c.maxLoad();
if (c.excluded()) {
weightExcluded++;
} else if (ctx.quarantined().contains(c.profile())) {
quarantined++;
} else if (ctx.coolingOff().contains(c.profile())) {
coolingOff++;
} else if (ctx.modelOff().contains(c.profile())) {
modelOff++;
} else if (ctx.unreachable().contains(c.profile())) {
unreachable++;
} else if (cap != null && ctx.liveCount().apply(c.profile()) >= cap) {
} else if (atCap(ctx, c)) {
atCap++;
}
}
@@ -83,6 +98,10 @@ final class PlacementPolicyUtil {
return new PlacementException(
"all worker profiles are cooling off after repeated backend errors");
}
if (modelOff == total) {
return new PlacementException(
"all worker profiles name a model the operator has turned off");
}
if (atCap == total) {
return new PlacementException("all worker profiles are at maxLoad");
}
@@ -92,8 +111,9 @@ final class PlacementPolicyUtil {
return new PlacementException("no worker profile available: " + atCap + " at maxLoad, "
+ unreachable + " unreachable, " + quarantined + " quarantined, "
+ coolingOff + " cooling off, "
+ modelOff + " model-off, "
+ weightExcluded + " weight-0, "
+ (total - atCap - unreachable - quarantined - coolingOff - weightExcluded)
+ (total - atCap - unreachable - quarantined - coolingOff - modelOff - weightExcluded)
+ " remaining");
}
}
@@ -10,6 +10,7 @@ import dev.ltms.fleet.peer.MemberRole;
import dev.ltms.fleet.peer.PeerHandle;
import dev.ltms.fleet.peer.PeerLauncher;
import dev.ltms.fleet.peer.SpawnRequest;
import dev.ltms.fleet.placement.PlacementDecision;
import org.slf4j.Logger;
import org.slf4j.LoggerFactory;
@@ -584,8 +585,65 @@ public final class SessionManager implements TurnListener {
String ownerTerminal, WorktreeRequest wt,
String sessionName, String resumeSessionId,
MemberLifecycle.SlotReservation reservation) {
String preResolvedProfile = (profile == null || profile.isBlank())
? launcher.defaultProfile() : profile;
// fleetd #425 rework (round 2): resolved through launcher.place(memberRole) — the same
// candidate list, quarantine/cool-off/model-off filtering, and PlacementPolicy an unqualified
// spawn of this role is actually judged against right now — never launcher.defaultProfile()
// (MemberRole.DEV only, wrong for any other role) and never launcher.defaultProfileFor()
// (the role's pool FIRST entry, blind to quarantine/cool-off/model-off: a first-round fix
// used exactly this and regressed fleetd #429's "the fleet keeps working when a model is
// turned off" guarantee — a quarantined or model-off pool-first profile made this throw
// instead of routing around it, which an unqualified spawn is supposed to do). This same
// resolved decision is reused below for repoRoot, parityOverlay, AND the spawn itself so the
// worktree is always provisioned for the profile the member actually runs on — the two could
// disagree before fleetd #425: this name picked repoRoot/overlay, but the spawn below passed
// the ORIGINAL (blank) profile through to placement, which re-resolves live and can pick a
// different profile if the pool changed between the two reads, or a genuinely different one
// under weighted/round-robin placement.
//
// Round 1 of this rework fed the resolved name back into launcher.spawn(SpawnRequest) as an
// EXPLICIT profile. That was a mistake this round corrects, and the mistake is not that the
// two branches apply different checks — they are SUPPOSED to disagree: the routing branch a
// blank spawn takes treats a quarantined/cooling-off/at-cap/model-off/unreachable/weight-0
// profile as a reason to fall through to the next candidate, while CompositePeerLauncher's
// THROWING branch (enforceNotQuarantined/enforceNotCoolingOff/enforceMaxLoad/
// enforceModelEnabled) treats naming that same profile explicitly as a reason to refuse
// outright. That is correct: an operator who names a profile should get a refusal, not a
// silent substitution onto a different backend. The mistake was turning a fall-through into
// a refusal by accident — resolving a name via the routing side and then re-entering the
// refusing side with it, for a placement the routing side had already approved by walking
// past everything else.
//
// Before fleetd #435, this accident was reachable through maxLoad specifically: the default
// `fixed` placement policy did not evaluate maxLoad at all for automatic selection, so an
// at-cap pool-first profile that placement itself would have picked for a plain unqualified
// spawn could die at enforceMaxLoad one call later, purely because this method's route to
// the spawn passed through an explicit profile name — a failure a worktree-less unqualified
// spawn never hit. fleetd #435 closed that gap (`fixed` now evaluates maxLoad exactly like
// every other placement policy), so that specific failure can no longer happen — a
// PlacementDecision this method resolves can no longer be at-cap in the first place. What
// this round's fix still buys, now that maxLoad can no longer cause the accident: it keeps
// the PlacementDecision from place() and hands it to launcher.spawn(SpawnRequest,
// PlacementDecision) for an unqualified request, which spawns through the SAME routing
// branch a blank spawn uses — no enforce* check is newly applied, and the window between the
// placement decision and the spawn (in which the pool, a config reload, or another spawn
// landing on the same profile could otherwise move the state) never reopens. An
// explicitly-named profile still goes through launcher.spawn(SpawnRequest) and its throwing
// branch, unchanged — that caller asked for one profile by name and still gets everything
// enforceNotQuarantined/enforceNotCoolingOff/enforceMaxLoad/enforceModelEnabled decide about
// it, refusal included.
//
// The one cost that remains, unchanged from round 1: an unqualified worktree-provisioned
// spawn does not get CompositePeerLauncher's cross-candidate retry on a live
// PeerUnreachableException raised by the backend itself at spawn time (a transport-level
// failure placement cannot see in advance) — spawn(req, decision) commits to the one profile
// place() already chose, the same way an explicit-profile spawn commits to its one name. That
// trade is deliberate: a worktree provisioned for the wrong backend (the #425 hazard) is worse
// than a spawn that fails cleanly and can be retried by the caller. Nothing else is lost:
// maxLoad, quarantine, cool-off and model-off all behave identically whether or not a
// worktree was requested — that agreement is the invariant this rework exists to hold.
boolean unqualifiedProfile = profile == null || profile.isBlank();
PlacementDecision decision = unqualifiedProfile ? launcher.place(memberRole) : new PlacementDecision(profile);
String preResolvedProfile = decision.profile();
// CB-507: resolve through the launcher's CB-112 chain (requested → profile cwd → caller →
// daemon cwd → "."), never the raw args. A plain REST spawn supplies neither a requested
// nor a caller cwd, so taking the first non-blank of those two yielded null and put
@@ -608,7 +666,15 @@ public final class SessionManager implements TurnListener {
// copies more files into the worktree after add() returns, so sharing the group any earlier
// leaves those overlay files operator-owned and read-only for a different-uid member.
worktrees.shareWithGroup(repoRoot, path);
handle = launcher.spawn(new SpawnRequest(profile, path, callerCwd, sessionName, resumeSessionId, memberRole));
// fleetd #425: preResolvedProfile, not the original (possibly blank) profile — see the
// comment above where it is resolved. The overlay/repoRoot above and the spawn here must
// name the same profile. An unqualified request stays unqualified here and is honored via
// the PlacementDecision already captured above (spawn(req, decision) — the routing branch,
// no enforce* re-check); an explicitly-named profile still goes through the single-arg
// spawn(req) and its throwing branch, exactly as before this rework.
SpawnRequest spawnReq = new SpawnRequest(unqualifiedProfile ? null : preResolvedProfile,
path, callerCwd, sessionName, resumeSessionId, memberRole);
handle = unqualifiedProfile ? launcher.spawn(spawnReq, decision) : launcher.spawn(spawnReq);
} catch (RuntimeException e) {
log.warn("spawn failed for profile={} role={} branch={} path={}: {}",
preResolvedProfile, memberRole, branch, path, e.getMessage());
@@ -636,7 +702,7 @@ public final class SessionManager implements TurnListener {
}
throw e;
}
String resolvedProfile = resolveProfile(handle, profile);
String resolvedProfile = resolveProfile(handle, preResolvedProfile);
String cwd = launcher.effectiveCwd(new SpawnRequest(resolvedProfile, path, callerCwd));
long now = nowNanos.getAsLong();
// CB-619: see the no-worktree path above — bind before recording, and store the returned
@@ -0,0 +1,155 @@
package dev.ltms.fleet;
import dev.ltms.fleet.config.ConfigRef;
import dev.ltms.fleet.config.FleetConfig;
import dev.ltms.fleet.mcp.FleetMcp;
import org.junit.jupiter.api.DisplayName;
import org.junit.jupiter.api.Test;
import org.junit.jupiter.api.io.TempDir;
import java.nio.file.Files;
import java.nio.file.Path;
import static org.junit.jupiter.api.Assertions.assertFalse;
import static org.junit.jupiter.api.Assertions.assertTrue;
import static org.junit.jupiter.api.Assertions.assertEquals;
/**
* fleetd #416: {@code fleet_list}'s {@code CapacitySource.configuredProfiles} must enumerate the
* <em>startup</em> profile set, not the live, hot-reloaded one.
*
* <p>{@code profiles} as a whole is a {@code DEFERRED} key ({@link ConfigRef#DEFERRED_KEYS}):
* {@code HerdrPeerLauncher} takes {@code Map.copyOf(profiles)} once at construction, so a profile
* only added to the hot-reloaded map can never actually be spawned. Before this fix, {@code
* Fleetd.main} built {@code CapacitySource} with {@code () -> config.get().profiles().keySet()} —
* the live map — so {@code fleet_list} would report a freshly hot-reloaded profile as available
* ({@code free > 0}) while {@code fleet_spawn} on that same profile failed with
* {@code unknown worker profile}. Measured on another host: adding a throwaway profile and letting
* it hot-reload gave {@code fleet_list} -> {@code free: 3} and {@code fleet_spawn} ->
* {@code error: unknown worker profile}.
*
* <p>This test needs a reload, the same reason {@link FleetdExhaustionDetectionArmedWiringTest}
* does: at startup the two snapshots agree, so a test of only a newly started daemon would not
* detect a live {@code config.get()} lookup for the set.
*
* <p><b>maxLoad must stay hot.</b> It is read off {@code config.get()} exactly like
* {@code credentialId} ({@link ConfigRef} documents both as hot, "read live off the config
* supplier ... exactly like weight/maxLoad"), so a reload that only changes an existing profile's
* {@code maxLoad} — no add/remove — must still change what {@code fleet_list} reports without a
* restart. A fix that freezes the whole {@code CapacitySource} against {@code cfg} (rather than
* only its {@code configuredProfiles} set) would trade this bug for its mirror image and is pinned
* wrong by {@link #reloadedMaxLoadStillChangesWhatFleetListReports}.
*/
class FleetdCapacitySourceWiringTest {
private static final String STARTUP = """
bind:
host: 127.0.0.1
port: 8765
herdrSocket: ~/.config/herdr/herdr.sock
profiles:
terra:
baseUrl: http://gx00.gw:8000
model: terra
maxLoad: 3
guard:
offSubscriptionHosts:
- gx00.gw
""";
private static final String WITH_NEW_PROFILE = """
bind:
host: 127.0.0.1
port: 8765
herdrSocket: ~/.config/herdr/herdr.sock
profiles:
terra:
baseUrl: http://gx00.gw:8000
model: terra
maxLoad: 3
ghost404:
baseUrl: http://gx00.gw:8001
model: ghost404
maxLoad: 3
guard:
offSubscriptionHosts:
- gx00.gw
""";
private static final String WITH_CHANGED_MAX_LOAD = """
bind:
host: 127.0.0.1
port: 8765
herdrSocket: ~/.config/herdr/herdr.sock
profiles:
terra:
baseUrl: http://gx00.gw:8000
model: terra
maxLoad: 9
guard:
offSubscriptionHosts:
- gx00.gw
""";
@Test
@DisplayName("a profile present only in the live (hot-reloaded) config is NOT listed")
void liveOnlyProfileIsNotListed(@TempDir Path dir) throws Exception {
Path file = dir.resolve("fleetd.yaml");
Files.writeString(file, STARTUP);
FleetConfig cfg = FleetConfig.load(file);
ConfigRef config = new ConfigRef(file, cfg);
Files.writeString(file, WITH_NEW_PROFILE);
assertTrue(config.reload().applied());
// The live snapshot now has the new profile — proves the reload really happened and this
// test is not accidentally passing because nothing changed.
assertTrue(config.get().profiles().containsKey("ghost404"));
FleetMcp.CapacitySource source = Fleetd.capacitySource(config, cfg, _ -> 0);
assertFalse(source.configuredProfiles().get().contains("ghost404"),
"a profile added only to the hot-reloaded config must not be listed by fleet_list — "
+ "HerdrPeerLauncher never learns about it until a restart, so fleet_spawn on it "
+ "would fail with 'unknown worker profile' while fleet_list claimed it free");
}
@Test
@DisplayName("a profile present in the startup set IS listed")
void startupProfileIsListed(@TempDir Path dir) throws Exception {
Path file = dir.resolve("fleetd.yaml");
Files.writeString(file, STARTUP);
FleetConfig cfg = FleetConfig.load(file);
ConfigRef config = new ConfigRef(file, cfg);
FleetMcp.CapacitySource source = Fleetd.capacitySource(config, cfg, _ -> 0);
// fleetd #416, both-directions requirement: a test that only ever passes an empty/absent
// startup set (the case above) cannot tell a correct lookup from one that is permanently
// empty (e.g. a mutation replacing the supplier with Set::of). This is the direction that
// fails if the fix regresses to reporting nothing at all.
assertTrue(source.configuredProfiles().get().contains("terra"),
"a profile present in the startup snapshot must still be listed by fleet_list");
}
@Test
@DisplayName("a hot maxLoad edit still changes what fleet_list reports")
void reloadedMaxLoadStillChangesWhatFleetListReports(@TempDir Path dir) throws Exception {
Path file = dir.resolve("fleetd.yaml");
Files.writeString(file, STARTUP);
FleetConfig cfg = FleetConfig.load(file);
ConfigRef config = new ConfigRef(file, cfg);
FleetMcp.CapacitySource source = Fleetd.capacitySource(config, cfg, _ -> 0);
assertEquals(3, source.maxLoad().apply("terra"),
"sanity: maxLoad reads 3 from the startup config before any reload");
Files.writeString(file, WITH_CHANGED_MAX_LOAD);
assertTrue(config.reload().applied());
assertEquals(9, source.maxLoad().apply("terra"),
"maxLoad must stay hot — the SAME CapacitySource instance must reflect a reloaded "
+ "maxLoad without a restart, exactly like credentialId. Freezing the whole "
+ "CapacitySource against the startup snapshot (rather than only its "
+ "configuredProfiles set) would trade fleetd #416 for its mirror image.");
}
}
@@ -0,0 +1,62 @@
package dev.ltms.fleet;
import dev.ltms.fleet.peer.PeerLauncher;
import org.junit.jupiter.api.DisplayName;
import org.junit.jupiter.api.Test;
import java.util.Set;
import static org.junit.jupiter.api.Assertions.assertEquals;
import static org.junit.jupiter.api.Assertions.assertNotEquals;
/**
* fleetd #422 follow-up: {@code Fleetd.modelGateCoverageLine} is the startup-log counterpart of
* {@code exhaustedPatternCoverageLine}/{@code errorPatternCoverageLine} — see {@code
* FleetdPatternCoverageLineTest} for the identical shape this follows — except here there is no
* {@code UnsetMeaning} choice for a caller to get backwards: {@link
* PeerLauncher.ModelGateState#configured()} already states, unambiguously, whether an empty
* {@link PeerLauncher.ModelGateState#off()} means "no {@code models:} block to gate with at all"
* or "a block armed and currently reporting zero off". This class proves {@code
* modelGateCoverageLine} words those two states — plus the third, N off — distinctly, so a
* mutation that made it ignore {@code configured()} either way is caught here.
*/
class FleetdModelGateCoverageLineTest {
@Test
@DisplayName("no models: block reports not configured, distinct from armed-with-zero")
void noModelsBlockReportsNotConfigured() {
String line = Fleetd.modelGateCoverageLine(PeerLauncher.ModelGateState.notConfigured());
assertEquals("not configured (no models: block — nothing is gated, and nothing can be)", line);
}
@Test
@DisplayName("a models: block armed with nothing off reports armed, distinct from not configured")
void armedWithNothingOffReportsArmed() {
String line = Fleetd.modelGateCoverageLine(PeerLauncher.ModelGateState.armed(Set.of()));
assertEquals("armed (models: block present; 0 models currently turned off)", line);
}
@Test
@DisplayName("a models: block with N off names the off models")
void armedWithModelsOffNamesThem() {
String line = Fleetd.modelGateCoverageLine(
PeerLauncher.ModelGateState.armed(Set.of("deepseek-v4-flash", "claude-opus-9000")));
assertEquals("armed (2 model(s) turned off: [claude-opus-9000, deepseek-v4-flash])", line);
}
@Test
@DisplayName("the three states produce pairwise-distinct wording for the same empty-looking input")
void theThreeStatesProduceDistinctWording() {
String notConfigured = Fleetd.modelGateCoverageLine(PeerLauncher.ModelGateState.notConfigured());
String armedZero = Fleetd.modelGateCoverageLine(PeerLauncher.ModelGateState.armed(Set.of()));
String armedOne = Fleetd.modelGateCoverageLine(PeerLauncher.ModelGateState.armed(Set.of("x")));
// Pinned individually above; restated here so this test alone still catches a regression
// even if one of the three tests above were ever deleted — the exact FleetdPatternCoverageLineTest
// pattern, adapted from "two keys" to "three states of one gate".
assertNotEquals(notConfigured, armedZero,
"collapsing 'no models: block' into 'armed, zero off' is fleetd #422 follow-up's exact defect");
assertNotEquals(armedZero, armedOne);
assertNotEquals(notConfigured, armedOne);
}
}
@@ -0,0 +1,67 @@
package dev.ltms.fleet;
import org.junit.jupiter.api.DisplayName;
import org.junit.jupiter.api.Test;
import java.util.Set;
import static org.junit.jupiter.api.Assertions.assertEquals;
import static org.junit.jupiter.api.Assertions.assertNotEquals;
/**
* fleetd #415 (review follow-up): {@code CompletionResolverTest} proves {@code coverage()} words
* {@code UnsetMeaning.OFF} and {@code UnsetMeaning.BUILT_IN_DEFAULT} correctly — but every one of
* those tests supplies the meaning itself. That proves the enum's wording, never that {@code
* Fleetd} pairs the right meaning with the right pattern key. That pairing is #415's actual
* defect: {@code coverage()} had no way to know what unset meant for its key, so the fix moved
* the fact to the caller — and nothing yet proved the caller states it correctly.
*
* <p><b>Measured or it didn't happen:</b> swapping the two {@code UnsetMeaning} arguments at
* {@code Fleetd}'s two coverage call sites — giving {@code exhaustedPattern} the built-in-default
* wording and {@code errorPattern} the off wording, #415's exact defect with the keys exchanged —
* compiled with 0 errors and left all 1506 existing tests green. This class exists to turn that
* swap red.
*
* <p>It calls {@link Fleetd#exhaustedPatternCoverageLine} and {@link Fleetd#errorPatternCoverageLine}
* directly rather than reading {@code Fleetd.java} as source text (the shape {@code
* FleetdCompletionResolverWiringTest} uses for a different wiring gap): those two methods are the
* extracted call sites {@code main} actually invokes, following the same {@code static} factory +
* dedicated-test pattern as {@link Fleetd#capacitySource} and {@link Fleetd#worktreeBranchLookup}.
*/
class FleetdPatternCoverageLineTest {
private static final Set<String> PROFILES = Set.of("terra", "gx10");
@Test
@DisplayName("exhaustedPatternCoverageLine says off when no profile configures exhaustedPattern")
void exhaustedPatternCoverageLineSaysOffWhenNoProfileConfiguresIt() {
assertEquals("off (no profile has an exhaustedPattern configured; profiles: [gx10, terra])",
Fleetd.exhaustedPatternCoverageLine(PROFILES, Set.of()));
}
@Test
@DisplayName("errorPatternCoverageLine says built-in default when no profile configures errorPattern")
void errorPatternCoverageLineSaysBuiltInDefaultWhenNoProfileConfiguresIt() {
assertEquals("built-in default for all profiles (no profile customises errorPattern; "
+ "profiles: [gx10, terra])",
Fleetd.errorPatternCoverageLine(PROFILES, Set.of()));
}
@Test
@DisplayName("the two keys produce different wording for the identical empty-coverage input")
void theTwoKeysProduceDifferentWordingForTheSameEmptyInput() {
String exhaustedLine = Fleetd.exhaustedPatternCoverageLine(PROFILES, Set.of());
String errorLine = Fleetd.errorPatternCoverageLine(PROFILES, Set.of());
// Pinned individually above; restated here so this test alone still catches a swap even
// if one of the two tests above were ever deleted.
assertEquals("off (no profile has an exhaustedPattern configured; profiles: [gx10, terra])",
exhaustedLine);
assertEquals("built-in default for all profiles (no profile customises errorPattern; "
+ "profiles: [gx10, terra])", errorLine);
assertNotEquals(exhaustedLine, errorLine,
"swapping which UnsetMeaning pairs with which pattern key at Fleetd's call sites "
+ "must be caught here — that pairing, not coverage()'s own wording in isolation, "
+ "is fleetd #415's actual defect");
}
}
@@ -0,0 +1,79 @@
package dev.ltms.fleet;
import ch.qos.logback.classic.Logger;
import ch.qos.logback.classic.Level;
import ch.qos.logback.classic.spi.ILoggingEvent;
import ch.qos.logback.core.read.ListAppender;
import dev.ltms.fleet.config.FleetConfig;
import org.junit.jupiter.api.Test;
import org.junit.jupiter.api.io.TempDir;
import org.slf4j.LoggerFactory;
import java.nio.file.Files;
import java.nio.file.Path;
import static org.junit.jupiter.api.Assertions.assertThrows;
import static org.junit.jupiter.api.Assertions.assertTrue;
/**
* Proves {@link Fleetd#main(String[])} calls every startup report before validation aborts startup.
* The invalid non-loopback bind makes {@link FleetConfig#validateAll()} throw before {@code main}
* can open the herdr socket or bind a port. The fixture also triggers every report, so removing any
* one call from {@code main} leaves its expected log line absent.
*/
class FleetdStartupReportTest {
private static Level originalLevel;
private static ListAppender<ILoggingEvent> attach() {
Logger logger = (Logger) LoggerFactory.getLogger(Fleetd.class);
originalLevel = logger.getLevel();
logger.setLevel(Level.INFO);
ListAppender<ILoggingEvent> appender = new ListAppender<>();
appender.start();
logger.addAppender(appender);
return appender;
}
private static void detach(ListAppender<ILoggingEvent> appender) {
Logger logger = (Logger) LoggerFactory.getLogger(Fleetd.class);
logger.detachAppender(appender);
logger.setLevel(originalLevel);
}
private static boolean contains(ListAppender<ILoggingEvent> appender, String fragment) {
return appender.list.stream()
.map(ILoggingEvent::getFormattedMessage)
.anyMatch(message -> message.contains(fragment));
}
@Test
void mainReportsEveryStartupGapBeforeValidationAborts(@TempDir Path dir) throws Exception {
Path config = dir.resolve("fleetd.yaml");
Files.writeString(config, """
bind:
host: 0.0.0.0
port: 8765
profiles:
worker:
baseUrl: https://llm.ltms.dev/v1
gitTokenEnv: GITEA_TOKEN
""");
ListAppender<ILoggingEvent> appender = attach();
try {
assertThrows(IllegalStateException.class, () -> Fleetd.main(new String[]{config.toString()}));
} finally {
detach(appender);
}
assertTrue(contains(appender, "startup git host GITEA_HOST:"),
"Fleetd.main must report the git host shape");
assertTrue(contains(appender, "member trust model: members run as the same OS user"),
"Fleetd.main must report the member trust model");
assertTrue(contains(appender, "memberCredentials: absent or empty"),
"Fleetd.main must report an absent memberCredentials policy");
assertTrue(contains(appender, "exhaustedPattern: profile(s) [worker] have no exhaustedPattern configured"),
"Fleetd.main must report profiles without exhaustedPattern");
}
}
@@ -113,6 +113,20 @@ class AuthzTest {
assertTrue(Authz.permits(WORKER_A, METRICS, null));
}
/**
* fleetd #421 — unlike READ (the case above), COORD_READ (a lead's own held peer mail) is the
* primary's alone. A worker or an architect reading it would disclose lead-to-lead
* coordination bodies, not the secret-free roster READ is open about.
*/
@Test
void coordReadIsThePrimarysAloneNotAWidenedRead() {
assertTrue(Authz.permits(PRIMARY, COORD_READ, null));
assertFalse(Authz.permits(WORKER_A, COORD_READ, null),
"a worker must not read held lead-to-lead mail");
assertFalse(Authz.permits(ARCH_DESIGN, COORD_READ, null),
"an architect holds READ today, but not-primary must mean not-architect here too");
}
@Test
void unauthenticatedIsDistinguishedFromMerelyForbidden() {
// Drives the 401-vs-403 split: a missing credential is fixable by the caller, a wrong role
@@ -0,0 +1,360 @@
package dev.ltms.fleet.auth;
import dev.ltms.fleet.config.ConfigRef;
import dev.ltms.fleet.config.FleetConfig;
import dev.ltms.fleet.herdr.FakeHerdr;
import dev.ltms.fleet.herdr.PaneLocator;
import dev.ltms.fleet.mcp.ConnectionIdentity;
import dev.ltms.fleet.peer.MemberRole;
import org.junit.jupiter.api.Test;
import org.junit.jupiter.api.io.TempDir;
import java.nio.file.Files;
import java.nio.file.Path;
import java.util.Map;
import static org.junit.jupiter.api.Assertions.*;
/**
* fleetd #424 — revoking (or granting) an architect slot must take effect on the next spawn with
* no restart. A session already bound to a slot keeps its <em>binding</em> (the {@code
* terminalToSlot} occupancy) even after that slot drops out of config, but NOT the ARCHITECT
* <em>privilege</em> the slot used to grant — that is revoked on the bound session's very next
* request. See {@link MemberRegistry}'s class doc for the exact rule: "config governs what a
* bound slot still grants, as well as what may be bound next."
*
* <p>Every test here drives a REAL {@link ConfigRef#reload()} against a {@code @TempDir} file and
* asserts {@link ConfigRef.Outcome#applied()}, rather than comparing two frozen
* {@code MemberRegistry} instances in memory — the defect this ticket fixes is specifically that
* {@link MemberRegistry} used to ignore a live reload, so a test that never reloads cannot tell the
* fixed registry from the broken one. {@code requireSlotFor} and {@code reserve} are pinned in
* separate tests, in both directions (removed and added), so a registry that simply refuses (or
* simply allows) everything cannot pass by accident — see {@link MemberRegistryTest} for the
* registry's other invariants (bind/unbind cardinality, thread-safety), which are unaffected by
* this ticket and still exercised against the frozen constructor.
*/
class MemberRegistryLiveTest {
private static String yaml(String fleetBlock) {
return """
bind:
host: 127.0.0.1
port: 8765
herdrSocket: ~/.config/herdr/herdr.sock
profiles:
sonnet:
baseUrl: http://gx00.gw:8000
model: sonnet
opus:
baseUrl: http://gx00.gw:8001
model: opus
guard:
offSubscriptionHosts:
- gx00.gw
""" + fleetBlock;
}
private static final String WITH_SONNET_SLOT = """
fleet:
architects:
designer:
profile: sonnet
""";
/** No architect pool at all — developers is unrelated dead data for this registry (#424 out of scope). */
private static final String WITHOUT_ARCHITECT_SLOTS = """
fleet:
developers:
dev1:
profile: sonnet
""";
/** Same slot name ({@code designer}) as {@link #WITH_SONNET_SLOT}, repointed to a different profile. */
private static final String WITH_OPUS_SLOT = """
fleet:
architects:
designer:
profile: opus
""";
/** Same profile ({@code sonnet}) as {@link #WITH_SONNET_SLOT}, but the pool key is renamed. */
private static final String WITH_RENAMED_SLOT = """
fleet:
architects:
architect-lead:
profile: sonnet
""";
private static ConfigRef refFor(Path f) {
return new ConfigRef(f, FleetConfig.load(f));
}
// ── requireSlotFor is live (criteria 1, 2, 4) ──────────────────────────────────────────────
@Test
void requireSlotForRefusesAProfileWhoseSlotWasRemovedByReload(@TempDir Path dir) throws Exception {
Path f = dir.resolve("fleetd.yaml");
Files.writeString(f, yaml(WITH_SONNET_SLOT));
ConfigRef ref = refFor(f);
MemberRegistry registry = MemberRegistry.live(() -> ref.get().fleet());
assertDoesNotThrow(() -> registry.requireSlotFor(MemberRole.ARCHITECT, "sonnet"),
"the slot is configured before the reload");
Files.writeString(f, yaml(WITHOUT_ARCHITECT_SLOTS));
ConfigRef.Outcome out = ref.reload();
assertTrue(out.applied(), "the reload must actually take effect: " + out.summary());
assertThrows(IllegalArgumentException.class,
() -> registry.requireSlotFor(MemberRole.ARCHITECT, "sonnet"),
"revoking the slot must refuse the NEXT spawn that names it");
}
@Test
void requireSlotForAllowsAProfileWhoseSlotWasAddedByReload(@TempDir Path dir) throws Exception {
Path f = dir.resolve("fleetd.yaml");
Files.writeString(f, yaml(WITHOUT_ARCHITECT_SLOTS));
ConfigRef ref = refFor(f);
MemberRegistry registry = MemberRegistry.live(() -> ref.get().fleet());
assertThrows(IllegalArgumentException.class,
() -> registry.requireSlotFor(MemberRole.ARCHITECT, "sonnet"),
"no architect slot is configured yet");
Files.writeString(f, yaml(WITH_SONNET_SLOT));
ConfigRef.Outcome out = ref.reload();
assertTrue(out.applied(), "the reload must actually take effect: " + out.summary());
assertDoesNotThrow(() -> registry.requireSlotFor(MemberRole.ARCHITECT, "sonnet"),
"a slot added by reload must be usable with no restart");
}
// ── reserve is live too — tested separately from requireSlotFor (criteria 1, 2, 4) ────────
@Test
void reserveRefusesAProfileWhoseSlotWasRemovedByReload(@TempDir Path dir) throws Exception {
Path f = dir.resolve("fleetd.yaml");
Files.writeString(f, yaml(WITH_SONNET_SLOT));
ConfigRef ref = refFor(f);
MemberRegistry registry = MemberRegistry.live(() -> ref.get().fleet());
MemberLifecycle.SlotReservation before = registry.reserve(MemberRole.ARCHITECT, "sonnet");
assertEquals("architect:designer", before.slot());
registry.release(before); // free it back up so the reload-side reserve below starts clean
Files.writeString(f, yaml(WITHOUT_ARCHITECT_SLOTS));
ConfigRef.Outcome out = ref.reload();
assertTrue(out.applied(), "the reload must actually take effect: " + out.summary());
assertThrows(IllegalArgumentException.class,
() -> registry.reserve(MemberRole.ARCHITECT, "sonnet"),
"revoking the slot must refuse the NEXT reservation for it");
}
@Test
void reserveAllowsAProfileWhoseSlotWasAddedByReload(@TempDir Path dir) throws Exception {
Path f = dir.resolve("fleetd.yaml");
Files.writeString(f, yaml(WITHOUT_ARCHITECT_SLOTS));
ConfigRef ref = refFor(f);
MemberRegistry registry = MemberRegistry.live(() -> ref.get().fleet());
assertThrows(IllegalArgumentException.class,
() -> registry.reserve(MemberRole.ARCHITECT, "sonnet"),
"no architect slot is configured yet");
Files.writeString(f, yaml(WITH_SONNET_SLOT));
ConfigRef.Outcome out = ref.reload();
assertTrue(out.applied(), "the reload must actually take effect: " + out.summary());
MemberLifecycle.SlotReservation after = registry.reserve(MemberRole.ARCHITECT, "sonnet");
assertEquals("architect:designer", after.slot(),
"a slot added by reload must be reservable with no restart");
}
// ── profileForSlot, isSlot and nameForSlot are live too (fleetd #431) ─────────────────────────
// #424 pinned roleForSlot against a live reload but left these three untested — proved by
// mutating each to read a snapshot flattened once at construction: the full suite stayed green
// for all three.
//
// The three differ in how much production behaviour depends on them, and the ticket first got
// this ranking wrong. nameForSlot is wired: CallerResolver passes members::nameForSlot, next to
// members::roleForSlot. isSlot is reached through bind, which calls it to refuse an unknown
// slot. profileForSlot has NO caller in src/main at all — grepped both the ".profileForSlot("
// and the "::profileForSlot" form — so there is no seam to drive its test through beyond the
// accessor itself, and its own javadoc ("what the spawn lifecycle reads") describes a caller
// that does not exist. These tests pin the accessors as they are; whether profileForSlot should
// be wired or deleted is a separate question.
@Test
void profileForSlotReflectsAProfileChangedByReload(@TempDir Path dir) throws Exception {
Path f = dir.resolve("fleetd.yaml");
Files.writeString(f, yaml(WITH_SONNET_SLOT));
ConfigRef ref = refFor(f);
MemberRegistry registry = MemberRegistry.live(() -> ref.get().fleet());
assertEquals("sonnet", registry.profileForSlot("architect:designer"),
"the slot's profile before the reload");
Files.writeString(f, yaml(WITH_OPUS_SLOT));
ConfigRef.Outcome out = ref.reload();
assertTrue(out.applied(), "the reload must actually take effect: " + out.summary());
assertEquals("opus", registry.profileForSlot("architect:designer"),
"repointing the slot to a different profile must take effect with no restart");
}
@Test
void isSlotStopsReportingASlotRemovedByReload(@TempDir Path dir) throws Exception {
Path f = dir.resolve("fleetd.yaml");
Files.writeString(f, yaml(WITH_SONNET_SLOT));
ConfigRef ref = refFor(f);
MemberRegistry registry = MemberRegistry.live(() -> ref.get().fleet());
assertTrue(registry.isSlot("architect:designer"), "the slot is configured before the reload");
Files.writeString(f, yaml(WITHOUT_ARCHITECT_SLOTS));
ConfigRef.Outcome out = ref.reload();
assertTrue(out.applied(), "the reload must actually take effect: " + out.summary());
assertFalse(registry.isSlot("architect:designer"),
"removing the slot from config must make isSlot say so on the very next call, "
+ "with no restart");
}
@Test
void isSlotStartsReportingASlotAddedByReload(@TempDir Path dir) throws Exception {
Path f = dir.resolve("fleetd.yaml");
Files.writeString(f, yaml(WITHOUT_ARCHITECT_SLOTS));
ConfigRef ref = refFor(f);
MemberRegistry registry = MemberRegistry.live(() -> ref.get().fleet());
assertFalse(registry.isSlot("architect:designer"), "no architect slot is configured yet");
Files.writeString(f, yaml(WITH_SONNET_SLOT));
ConfigRef.Outcome out = ref.reload();
assertTrue(out.applied(), "the reload must actually take effect: " + out.summary());
assertTrue(registry.isSlot("architect:designer"),
"a slot added by reload must be visible to isSlot with no restart");
}
@Test
void nameForSlotReflectsANameChangedByReload(@TempDir Path dir) throws Exception {
Path f = dir.resolve("fleetd.yaml");
Files.writeString(f, yaml(WITH_SONNET_SLOT));
ConfigRef ref = refFor(f);
MemberRegistry registry = MemberRegistry.live(() -> ref.get().fleet());
assertEquals("designer", registry.nameForSlot("architect:designer"),
"the configured name before the reload");
Files.writeString(f, yaml(WITH_RENAMED_SLOT));
ConfigRef.Outcome out = ref.reload();
assertTrue(out.applied(), "the reload must actually take effect: " + out.summary());
assertNull(registry.nameForSlot("architect:designer"),
"the old key no longer names a configured slot — it was renamed away by the reload");
assertEquals("architect-lead", registry.nameForSlot("architect:architect-lead"),
"the new name must be visible under its new qualified key with no restart");
}
// ── a bound architect is demoted, but the binding itself is not touched (fleetd #424) ───────
// The lead's corrected ruling: the PRIVILEGE a slot grants is revoked on the bound session's
// very next request, but the terminalToSlot BINDING itself is untouched by a reload — dropping
// it would double-book the slot key and break unbind's compare-safe contract. See the class
// doc's binding rule.
/** A caller identity resolving the one canned pane (terminal {@code term_a}) in {@link FakeHerdr}. */
private static ConnectionIdentity boundPaneIdentity() {
return new ConnectionIdentity(new PaneLocator(new FakeHerdr()), _ -> FakeHerdr.WORKER_PID);
}
@Test
void anArchitectAlreadyBoundToASlotIsDemotedByReload(@TempDir Path dir) throws Exception {
Path f = dir.resolve("fleetd.yaml");
Files.writeString(f, yaml(WITH_SONNET_SLOT));
ConfigRef ref = refFor(f);
MemberRegistry registry = MemberRegistry.live(() -> ref.get().fleet());
MemberLifecycle.SlotReservation reservation = registry.reserve(MemberRole.ARCHITECT, "sonnet");
assertTrue(registry.bind(reservation, "term_a"));
assertEquals("architect:designer", registry.slotForTerminal("term_a"));
// Drive the real caller path, not the roleForSlot seam directly: CallerResolver.resolve is
// what a live request actually goes through (CallerResolver.java:220), and a resolver that
// ignored roleForSlot entirely would still pass a test that only checked the seam.
CallerResolver resolver = CallerResolver.withLeadsAndMembers(
boundPaneIdentity(), false, null, Map::of, registry);
Principal before = resolver.resolve("127.0.0.1", 42, null);
assertEquals(Role.ARCHITECT, before.role(), "sanity check: the harness binds term_a as an architect");
assertEquals("designer", before.name());
Files.writeString(f, yaml(WITHOUT_ARCHITECT_SLOTS));
ConfigRef.Outcome out = ref.reload();
assertTrue(out.applied(), "the reload must actually take effect: " + out.summary());
Principal after = resolver.resolve("127.0.0.1", 42, null);
assertEquals(Role.WORKER, after.role(),
"removing the slot from config must demote the bound session to worker on its "
+ "NEXT request — this is the ticket's whole point");
assertEquals("term_a", after.terminal(), "same pane, same terminal — only the role changed");
}
@Test
void theOriginalBindingStillOccupiesTheRemovedSlotSoASecondTerminalCannotClaimIt(@TempDir Path dir)
throws Exception {
Path f = dir.resolve("fleetd.yaml");
Files.writeString(f, yaml(WITH_SONNET_SLOT));
ConfigRef ref = refFor(f);
MemberRegistry registry = MemberRegistry.live(() -> ref.get().fleet());
MemberLifecycle.SlotReservation reservation = registry.reserve(MemberRole.ARCHITECT, "sonnet");
assertTrue(registry.bind(reservation, "term_a"));
Files.writeString(f, yaml(WITHOUT_ARCHITECT_SLOTS));
ConfigRef.Outcome removed = ref.reload();
assertTrue(removed.applied(), "the reload must actually take effect: " + removed.summary());
// The binding survives the removal untouched.
assertEquals("architect:designer", registry.slotForTerminal("term_a"),
"a live binding must never be retroactively unbound by a config edit");
assertEquals(Map.of("term_a", "architect:designer"), registry.snapshot());
// Bring the slot back into config. If the binding had been silently dropped by the removal
// (rather than merely losing the privilege it grants), a second terminal could now claim
// the "freed" key — the exact double-booking the class doc's binding rule rules out.
Files.writeString(f, yaml(WITH_SONNET_SLOT));
ConfigRef.Outcome restored = ref.reload();
assertTrue(restored.applied(), "the reload must actually take effect: " + restored.summary());
assertFalse(registry.bind("architect:designer", "term_b"),
"the slot is still occupied by term_a — a second terminal must not bind to it");
assertThrows(IllegalArgumentException.class,
() -> registry.reserve(MemberRole.ARCHITECT, "sonnet"),
"the slot is still occupied by term_a — a fresh reservation must not find it free");
assertEquals("architect:designer", registry.slotForTerminal("term_a"),
"the original binding is unchanged throughout");
}
@Test
void unbindStillSucceedsForTheOriginalTerminalAfterItsSlotIsRemoved(@TempDir Path dir) throws Exception {
Path f = dir.resolve("fleetd.yaml");
Files.writeString(f, yaml(WITH_SONNET_SLOT));
ConfigRef ref = refFor(f);
MemberRegistry registry = MemberRegistry.live(() -> ref.get().fleet());
MemberLifecycle.SlotReservation reservation = registry.reserve(MemberRole.ARCHITECT, "sonnet");
assertTrue(registry.bind(reservation, "term_a"));
Files.writeString(f, yaml(WITHOUT_ARCHITECT_SLOTS));
ConfigRef.Outcome out = ref.reload();
assertTrue(out.applied(), "the reload must actually take effect: " + out.summary());
assertTrue(registry.unbind("architect:designer", "term_a"),
"unbind must still work for a slot that config has since removed, or a session "
+ "that outlives its slot's removal could never release it");
assertNull(registry.slotForTerminal("term_a"));
assertEquals(Map.of(), registry.snapshot());
}
}
@@ -79,6 +79,14 @@ class ConfigRefTopLevelCoverageTest {
* <li>{@code memberCredentials} — read live at {@code Fleetd.java:198, 205, 729}.</li>
* <li>{@code memberLoginShell} — read live at
* {@code HerdrPeerLauncher#configuredMemberLoginShell}.</li>
* <li>{@code models} (fleetd #422) — membership is re-validated in full against the fresh
* config on every {@code ConfigRef.reload()} (via {@code FleetConfig#validateAll()}), so
* a bad edit is refused, never cached stale; the on/off half is read live by {@code
* CompositePeerLauncher.enforceModelEnabled} and its candidate filter, and by {@code
* fleet_profiles}/{@code GET /profiles} through {@code PeerLauncher.disabledModels()}.
* Unlike {@code fleet} below, nothing about {@code models} is baked into a startup-built
* object anywhere — there is no frozen half, so it belongs here whole rather than in
* {@code SPLIT_KEYS}.</li>
* </ul>
*
* <p>{@code fleet} used to sit here too, on the strength of most of it (role pools, charters,
@@ -93,7 +101,7 @@ class ConfigRefTopLevelCoverageTest {
* while this test stayed green throughout.</p>
*/
private static final Set<String> HOT_EXCLUDED_TOP_LEVEL_KEYS =
Set.of("placement", "memberCredentials", "memberLoginShell");
Set.of("placement", "memberCredentials", "memberLoginShell", "models");
@Test
void everyTopLevelComponentIsAccountedForInExactlyOneClass() {
@@ -120,7 +128,7 @@ class ConfigRefTopLevelCoverageTest {
// The escape hatch is pinned. Growing it requires editing this line — a visible, deliberate
// diff, not a quiet one. See the field javadoc above for what "belongs here" actually means.
assertEquals(Set.of("placement", "memberCredentials", "memberLoginShell"), hot,
assertEquals(Set.of("placement", "memberCredentials", "memberLoginShell", "models"), hot,
"HOT_EXCLUDED_TOP_LEVEL_KEYS changed. A component belongs here ONLY if it is read "
+ "live off the config supplier, never because adding it makes this test "
+ "pass. If you are adding one to silence this test, that is fleetd #323 "
@@ -2901,4 +2901,119 @@ class FleetConfigTest {
void modelsIsAKnownTopLevelKey() {
assertTrue(FleetConfig.KNOWN_TOP_LEVEL_KEYS.contains("models"));
}
// --- fleetd #422: model on/off (spawn-time gate + hot reload) -----------------------------
/**
* Criterion 3: an {@code allow:} entry written before {@code enabled:} existed — no such key in
* the YAML at all — must behave exactly as it always did: on, and absent from {@link
* FleetConfig.Models#offIds()}. This is the old-style fixture the ticket asks for, proven
* through a real YAML load rather than only through the {@code ModelEntry(String)} back-compat
* constructor.
*/
@Test
void anOldStyleAllowEntryWithNoEnabledFieldStaysOn(@TempDir Path dir) throws Exception {
Path f = dir.resolve("fleetd.yaml");
Files.writeString(f, """
bind:
port: 8080
profiles:
sonnet:
baseUrl: http://gx10.gw:8000
model: claude-sonnet-5
models:
allow:
- model: claude-sonnet-5
""");
FleetConfig cfg = FleetConfig.load(f);
assertDoesNotThrow(cfg::validateModels);
assertTrue(cfg.models().ids().contains("claude-sonnet-5"), "membership is unaffected");
assertTrue(cfg.models().offIds().isEmpty(), "no enabled: field ⇒ nothing is off");
}
/**
* Criterion 7 (the whole point of the ticket, mirroring criterion 3's shape): a profile naming a
* model that is turned off ({@code enabled: false}) must still be VALID config —
* {@code validateModels()} checks membership only, never the on/off state, so it must not throw.
* Collapsing "off" into "removed from allow:" would make this test fail, which is exactly the
* design trap the ticket calls out.
*/
@Test
void anOffModelIsStillValidConfigForValidateModels(@TempDir Path dir) throws Exception {
Path f = dir.resolve("fleetd.yaml");
Files.writeString(f, """
bind:
port: 8080
profiles:
local:
baseUrl: http://local.gw:8000
model: deepseek-v4-flash
models:
allow:
- model: deepseek-v4-flash
enabled: false
""");
FleetConfig cfg = FleetConfig.load(f);
assertDoesNotThrow(cfg::validateModels,
"an off model must stay a valid allow-list member — only the spawn gate reads enabled");
assertTrue(cfg.models().ids().contains("deepseek-v4-flash"));
assertTrue(cfg.models().offIds().contains("deepseek-v4-flash"));
}
/**
* Criterion 8: one {@code enabled: false} entry names a model, not a profile — every profile
* naming that model is off, on purpose (the live example: {@code deepseek-v4-flash} backs both
* {@code local} and {@code local-direct}). Proven here at the {@code Models}/config-load level;
* {@code CompositePeerLauncherTest} proves the spawn-time consequence for both profiles.
*/
@Test
void turningOffOneModelIsIndependentOfHowManyProfilesNameIt(@TempDir Path dir) throws Exception {
Path f = dir.resolve("fleetd.yaml");
Files.writeString(f, """
bind:
port: 8080
profiles:
local:
baseUrl: http://local.gw:8000
model: deepseek-v4-flash
local-direct:
baseUrl: http://local.gw:8001
model: deepseek-v4-flash
models:
allow:
- model: deepseek-v4-flash
enabled: false
""");
FleetConfig cfg = FleetConfig.load(f);
assertDoesNotThrow(cfg::validateModels);
// One allow-list entry, one off id — the fan-out to both profiles happens at the reader
// (CompositePeerLauncher), not by duplicating the entry per profile.
assertEquals(Set.of("deepseek-v4-flash"), cfg.models().offIds());
assertEquals("deepseek-v4-flash", cfg.profiles().get("local").model());
assertEquals("deepseek-v4-flash", cfg.profiles().get("local-direct").model());
}
/** An {@code enabled: true} entry (explicit, not just absent) also stays on — not just null. */
@Test
void explicitlyEnabledTrueStaysOn(@TempDir Path dir) throws Exception {
Path f = dir.resolve("fleetd.yaml");
Files.writeString(f, """
bind:
port: 8080
profiles:
sonnet:
baseUrl: http://gx10.gw:8000
model: claude-sonnet-5
models:
allow:
- model: claude-sonnet-5
enabled: true
""");
FleetConfig cfg = FleetConfig.load(f);
assertTrue(cfg.models().offIds().isEmpty());
}
}
@@ -17,15 +17,70 @@ import static org.junit.jupiter.api.Assumptions.assumeTrue;
* SHELL directly (never {@code claude}, so no subscription/token involvement) and always tears
* the throwaway space down.
*
* <p>The seed shell's own startup (restoring its session, printing its banner) is asynchronous
* and its length is not a fleetd contract — measured here at ~2.5s on one host (fleetd #449). A
* fixed sleep before typing raced that startup: input typed before the shell reached its prompt
* was swallowed by the shell's own startup, and the pane showed the typed line followed by the
* startup banner with no command output at all — indistinguishable, at a glance, from the env
* map never reaching the shell. So this polls for a real signal (the pane's visible text
* settling, then the expected output appearing) instead of guessing a sleep length.
*
* <p>Tagged {@code contract}; run with {@code mvn test -Pcontract}.
*/
@Tag("contract")
class AgentControlContractTest {
private static final long POLL_INTERVAL_MS = 150;
/** Bound for the seed shell to settle: observed ~2.5s three times running; this leaves headroom. */
private static final long SHELL_READY_TIMEOUT_MS = 8_000;
/** Bound for the typed command's output to appear once the shell is ready: observed ~0.2s. */
private static final long OUTPUT_TIMEOUT_MS = 5_000;
private boolean noSocket() {
return !Files.exists(UnixSocketHerdrClient.defaultSocketPath());
}
private static String readPane(UnixSocketHerdrClient herdr, String paneId) {
return herdr.call("pane.read", Map.of("pane_id", paneId, "source", "visible"))
.path("read").path("text").asText("");
}
/**
* Poll {@code pane.read} until two consecutive reads come back identical — the shell's own
* startup output (restore banner, prompt) has stopped changing — or {@code timeoutMs} elapses.
* Never asserts by itself; the caller's own assertion is what actually verifies the outcome,
* this only avoids sending input into a shell still mid-startup.
*/
private static String waitUntilSettled(UnixSocketHerdrClient herdr, String paneId, long timeoutMs)
throws InterruptedException {
long deadline = System.currentTimeMillis() + timeoutMs;
String previous = null;
while (System.currentTimeMillis() < deadline) {
Thread.sleep(POLL_INTERVAL_MS);
String current = readPane(herdr, paneId);
if (current.equals(previous) && !current.isBlank()) {
return current;
}
previous = current;
}
return previous == null ? "" : previous;
}
/** Poll {@code pane.read} until {@code needle} appears or {@code timeoutMs} elapses. */
private static String waitForText(UnixSocketHerdrClient herdr, String paneId, String needle, long timeoutMs)
throws InterruptedException {
long deadline = System.currentTimeMillis() + timeoutMs;
String last = "";
while (System.currentTimeMillis() < deadline) {
last = readPane(herdr, paneId);
if (last.contains(needle)) {
return last;
}
Thread.sleep(POLL_INTERVAL_MS);
}
return last;
}
@Test
void tabCreateInjectsEnvIntoTheSeedShell() throws Exception {
assumeTrue(!noSocket(), "no herdr socket — skipping");
@@ -36,15 +91,13 @@ class AgentControlContractTest {
Map.of("ANTHROPIC_BASE_URL", "http://gx00.gw:8000"));
try {
assertNotNull(tab.rootPaneId(), "tab.create must return the seed pane");
Thread.sleep(1000); // let the seed shell reach its prompt
waitUntilSettled(herdr, tab.rootPaneId(), SHELL_READY_TIMEOUT_MS);
herdr.call("pane.send_input", Map.of(
"pane_id", tab.rootPaneId(),
"text", "printf 'PROBE_BASE=[%s]\\n' \"$ANTHROPIC_BASE_URL\"",
"keys", List.of("enter")));
Thread.sleep(800);
String visible = herdr.call("pane.read",
Map.of("pane_id", tab.rootPaneId(), "source", "visible"))
.path("read").path("text").asText("");
String visible = waitForText(herdr, tab.rootPaneId(),
"PROBE_BASE=[http://gx00.gw:8000]", OUTPUT_TIMEOUT_MS);
assertTrue(visible.contains("PROBE_BASE=[http://gx00.gw:8000]"),
"env map must reach the seed shell; saw: " + visible);
} finally {
@@ -14,7 +14,7 @@ import static org.junit.jupiter.api.Assumptions.assumeTrue;
* Contract test against a REAL running herdr. Tagged {@code contract} so it is
* excluded from {@code mvn test}; run it with {@code mvn test -Pcontract}. It fails
* loudly if herdr drifts from the protocol {@code fleetd} was built against
* (0.7.0, protocol 14) — catching breakage that unit tests with canned frames cannot.
* (0.8.0, protocol 19) — catching breakage that unit tests with canned frames cannot.
*/
@Tag("contract")
class HerdrContractTest {
@@ -24,13 +24,13 @@ class HerdrContractTest {
}
@Test
void pingReturnsProtocol14() {
void pingReturnsProtocol19() {
assumeTrue(Files.exists(socket()), "no herdr socket at " + socket() + " — skipping");
try (UnixSocketHerdrClient herdr = UnixSocketHerdrClient.connect()) {
JsonNode pong = herdr.call("ping");
assertEquals("pong", pong.get("type").asText());
assertEquals(14, pong.get("protocol").asInt(),
"fleetd is built against herdr protocol 14");
assertEquals(19, pong.get("protocol").asInt(),
"fleetd is built against herdr protocol 19");
assertFalse(pong.get("version").asText().isBlank());
}
}
@@ -20,6 +20,7 @@ import java.util.regex.Pattern;
import static org.junit.jupiter.api.Assertions.assertEquals;
import static org.junit.jupiter.api.Assertions.assertFalse;
import static org.junit.jupiter.api.Assertions.assertNotEquals;
import static org.junit.jupiter.api.Assertions.assertTrue;
/** Unit behaviour of the CB-106 completion resolver in isolation from the injector. */
@@ -884,19 +885,22 @@ class CompletionResolverTest {
@Test
void coverageIsOffWhenNoProfileHasAPatternConfigured() {
assertEquals("off (no profile has an exhaustedPattern configured; profiles: [terra])",
CompletionResolver.coverage("exhaustedPattern", Set.of("terra"), Set.of()));
CompletionResolver.coverage("exhaustedPattern", CompletionResolver.UnsetMeaning.OFF,
Set.of("terra"), Set.of()));
}
@Test
void coverageIsFullWhenEveryProfileHasAPatternConfigured() {
assertEquals("full (all profiles configured: [gx10, terra])",
CompletionResolver.coverage("exhaustedPattern", Set.of("terra", "gx10"), Set.of("terra", "gx10")));
CompletionResolver.coverage("exhaustedPattern", CompletionResolver.UnsetMeaning.OFF,
Set.of("terra", "gx10"), Set.of("terra", "gx10")));
}
@Test
void coverageIsPartialAndNamesWhichProfilesAreConfigured() {
assertEquals("partial (configured: [terra]; not configured: [gx10])",
CompletionResolver.coverage("exhaustedPattern", Set.of("terra", "gx10"), Set.of("terra")));
CompletionResolver.coverage("exhaustedPattern", CompletionResolver.UnsetMeaning.OFF,
Set.of("terra", "gx10"), Set.of("terra")));
}
/**
@@ -910,11 +914,42 @@ class CompletionResolverTest {
*
* <p>Every earlier test here passed the exhaustion case only, so none of them could see it. This
* one pins that the message names the key the caller actually meant.
*
* <p>fleetd#415: the expected wording changed here too. {@code errorPattern} has a built-in
* fallback ({@link CompletionResolver#BACKEND_ERROR}), so an empty {@code configuredProfiles}
* for it is not "off" — see {@link #coverageDistinguishesOffFromBuiltInDefaultForTheSameEmptyInput}
* for the test built specifically to pin that distinction.
*/
@Test
void coverageNamesTheConfigKeyItsCallerMeansRatherThanAlwaysSayingExhaustedPattern() {
assertEquals("off (no profile has an errorPattern configured; profiles: [gx10, terra])",
CompletionResolver.coverage("errorPattern", Set.of("terra", "gx10"), Set.of()));
assertEquals("built-in default for all profiles (no profile customises errorPattern; "
+ "profiles: [gx10, terra])",
CompletionResolver.coverage("errorPattern", CompletionResolver.UnsetMeaning.BUILT_IN_DEFAULT,
Set.of("terra", "gx10"), Set.of()));
}
/**
* fleetd#415: {@code coverage()} measures pattern coverage (how many profiles set the key), but
* for {@code errorPattern} the empty case is not the feature-off state — a profile with no
* configured {@code errorPattern} still runs the classification against
* {@link CompletionResolver#BACKEND_ERROR}. For {@code exhaustedPattern} there is no fallback,
* so empty really is off. Same shape of input (empty {@code configuredProfiles}, one profile),
* different {@link CompletionResolver.UnsetMeaning} — the wording must differ, or this method is
* back to conflating pattern coverage with feature state for the one key where they disagree.
*/
@Test
void coverageDistinguishesOffFromBuiltInDefaultForTheSameEmptyInput() {
String exhaustedLine = CompletionResolver.coverage("exhaustedPattern",
CompletionResolver.UnsetMeaning.OFF, Set.of("gx10", "terra"), Set.of());
String errorLine = CompletionResolver.coverage("errorPattern",
CompletionResolver.UnsetMeaning.BUILT_IN_DEFAULT, Set.of("gx10", "terra"), Set.of());
assertEquals("off (no profile has an exhaustedPattern configured; profiles: [gx10, terra])",
exhaustedLine);
assertEquals("built-in default for all profiles (no profile customises errorPattern; "
+ "profiles: [gx10, terra])", errorLine);
assertNotEquals(exhaustedLine, errorLine,
"the same empty-coverage input must not read as the same feature state for both keys");
}
// --- fleetd#201 Unit 1: target-keyed backend-error pattern + typed sink ----------------------
@@ -201,14 +201,34 @@ class FleetMcpAuthzTest {
*/
@Test
void pollingByTargetIsADrainAndPollingByTicketIsARead() {
assertEquals(Authz.Action.DRAIN, FleetMcp.pollAction("term_b"),
assertEquals(Authz.Action.DRAIN, FleetMcp.pollAction("term_b", null),
"poll by target removes the replies — that is a drain, not an observation");
assertEquals(Authz.Action.READ, FleetMcp.pollAction(null),
assertEquals(Authz.Action.READ, FleetMcp.pollAction(null, null),
"poll by ticket changes nothing");
assertEquals(Authz.Action.READ, FleetMcp.pollAction(" "),
assertEquals(Authz.Action.READ, FleetMcp.pollAction(" ", null),
"a blank target is an absent target");
}
/**
* fleetd #421: a coordId branch is a THIRD operation behind fleet_poll's one name, and it must
* map to {@link Authz.Action#COORD_READ} — never {@link Authz.Action#READ}, even though this
* branch also consumes nothing. READ's grant is open to every authenticated role on the premise
* that the roster carries no secrets; a lead-to-lead body is not the roster, so folding this
* branch into READ would let any worker read every peer lead's mail in full. coordId also takes
* priority over target when both happen to be present — it addresses a different inbox entirely.
*/
@Test
void pollingByCoordIdIsACoordReadNeverAPlainRead() {
assertEquals(Authz.Action.COORD_READ, FleetMcp.pollAction(null, "mac-opus"),
"reading held peer mail must not be mapped to the everyone-readable READ action");
assertEquals(Authz.Action.COORD_READ, FleetMcp.pollAction(" ", "mac-opus"),
"a blank target must not fall through to READ/DRAIN when coordId is present");
assertEquals(Authz.Action.READ, FleetMcp.pollAction(null, " "),
"a blank coordId is an absent coordId, same as target/ticket");
assertEquals(Authz.Action.COORD_READ, FleetMcp.pollAction("term_b", "mac-opus"),
"coordId takes priority over target — this is a different inbox, not a drain");
}
@Test
void everyRegisteredToolHasItsHandlerActionPinned() {
Set<String> registered = toolsTheServerRegisters();
@@ -230,6 +250,8 @@ class FleetMcpAuthzTest {
assertEquals(Authz.Action.READ, FleetMcp.toolAction("fleet_whoami", Map.of()));
assertEquals(Authz.Action.READ, FleetMcp.toolAction("fleet_poll", Map.of("ticket", "task")));
assertEquals(Authz.Action.DRAIN, FleetMcp.toolAction("fleet_poll", Map.of("target", "term_b")));
assertEquals(Authz.Action.COORD_READ,
FleetMcp.toolAction("fleet_poll", Map.of("coordId", "mac-opus")));
}
private static Set<String> toolsTheServerRegisters() {
@@ -249,11 +271,11 @@ class FleetMcpAuthzTest {
void aWorkerMayNotDrainAnotherSessionsInboxByPolling() {
FleetMcp m = mcp(true);
assertNotNull(m.denyFor(WORKER_A, FleetMcp.pollAction("term_b"), "term_b"),
assertNotNull(m.denyFor(WORKER_A, FleetMcp.pollAction("term_b", null), "term_b"),
"a worker draining a peer's inbox would destroy replies queued for the primary");
assertNotNull(m.denyFor(ARCH_DESIGN, FleetMcp.pollAction("term_b"), "term_b"),
assertNotNull(m.denyFor(ARCH_DESIGN, FleetMcp.pollAction("term_b", null), "term_b"),
"an architect has no lifecycle rights either — same gate as fleet_ack");
assertNull(m.denyFor(PRIMARY, FleetMcp.pollAction("term_b"), "term_b"),
assertNull(m.denyFor(PRIMARY, FleetMcp.pollAction("term_b", null), "term_b"),
"collecting a held reply is the primary's job");
}
@@ -265,11 +287,33 @@ class FleetMcpAuthzTest {
void pollingAnOwnTicketStaysOpenToWorkersAndArchitects() {
FleetMcp m = mcp(true);
assertNull(m.denyFor(WORKER_A, FleetMcp.pollAction(null), null));
assertNull(m.denyFor(ARCH_DESIGN, FleetMcp.pollAction(null), null),
assertNull(m.denyFor(WORKER_A, FleetMcp.pollAction(null, null), null));
assertNull(m.denyFor(ARCH_DESIGN, FleetMcp.pollAction(null, null), null),
"an architect delegates with wait:false, so it must be able to poll its ticket");
}
/**
* fleetd #421 acceptance: the whole ticket, pinned at the policy-table layer. A worker must be
* refused the coordId branch, the primary must be allowed, and (CB-548) an architect — which
* holds READ today — must be refused too, because "not primary" means not-architect here.
*/
@Test
void aWorkerAndAnArchitectMayNotReadHeldPeerMailOnlyThePrimaryMay() {
FleetMcp m = mcp(true);
Authz.Action coordRead = FleetMcp.pollAction(null, "mac-opus");
McpSchema.CallToolResult workerDenied = m.denyFor(WORKER_A, coordRead, null);
assertNotNull(workerDenied, "a worker must not read held lead-to-lead mail");
assertTrue(workerDenied.isError());
McpSchema.CallToolResult archDenied = m.denyFor(ARCH_DESIGN, coordRead, null);
assertNotNull(archDenied, "an architect holds READ today, but not-primary must mean "
+ "not-architect here too");
assertTrue(archDenied.isError());
assertNull(m.denyFor(PRIMARY, coordRead, null), "reading its own held mail is the primary's job");
}
// --- identity reconstruction from the transport context ------------------------------------
@Test
@@ -680,7 +680,11 @@ class FleetMcpTest {
void listReportsHeldMessagesWithATruncatedPreviewNeverTheFullBody() {
FakeHerdr h = new FakeHerdr();
SessionManager sessions = new SessionManager(workerService(h, "http://gx00.gw:8000", Set.of("gx00.gw")));
String longContent = "x".repeat(200);
// A homogeneous "x".repeat(200) body would not pin the exact 80-char boundary: a preview
// widened by one (81 chars) still CONTAINS "x".repeat(80) + "…" as a substring one position
// later, because every character is 'x'. Put a sentinel ("Y") exactly at index 80 — the
// first character a widened cap would leak — so any preview past 80 chars is caught.
String longContent = "x".repeat(80) + "Y" + "z".repeat(119);
FakeLeadChannel channel = new FakeLeadChannel("mac-opus")
.hold(new LeadMessage("m1", "fleet01-lead", "mac-opus", longContent));
@@ -694,9 +698,113 @@ class FleetMcpTest {
assertTrue(out.contains("\"msgId\":\"m1\""), out);
assertTrue(out.contains("\"from\":\"fleet01-lead\""), out);
assertFalse(out.contains(longContent), "fleet_list must never dump a held message's full body: " + out);
assertFalse(out.contains("Y"),
"the sentinel at index 80 must never appear — a preview past 80 chars leaked it: " + out);
assertTrue(out.contains("x".repeat(80) + "…"), "expected an 80-char preview with an ellipsis: " + out);
}
/**
* fleetd #421: {@code mailbox.pending} counts only broker-ready messages, so a blocked lead's
* normal, healthy state is {@code "pending": 0} next to a non-empty {@code held[]} — which
* invites the false reading that held mail is in-memory-only. {@code heldCount}/{@code
* heldDurable} put an honest second number and fact beside {@code pending} instead of leaving it
* as the only one.
*/
@Test
void listReportsAnHonestHeldCountAndDurabilityNotJustPendingZero() {
FakeHerdr h = new FakeHerdr();
SessionManager sessions = new SessionManager(workerService(h, "http://gx00.gw:8000", Set.of("gx00.gw")));
FakeLeadChannel channel = new FakeLeadChannel("mac-opus")
.withMailbox("mac-opus", LeadChannel.MailboxState.exists("mac-opus", 0, 1))
.hold(new LeadMessage("m1", "fleet01-lead", "mac-opus", "one"))
.hold(new LeadMessage("m2", "fleet01-lead", "mac-opus", "two"))
.hold(new LeadMessage("m3", "fleet01-lead", "mac-opus", "three"));
McpSchema.CallToolResult res = FleetMcp.listFleet(
workerService(h, "http://gx00.gw:8000", Set.of("gx00.gw")), sessions, null,
FleetMcp.CapacitySource.none(), new FleetMcp.HealthCoverageSource(() -> "off"),
FleetMcp.QuarantineSource.none(), Map.of(), "",
new FleetMcp.CoordinationSource(channel, List.of()));
String out = textOf(res);
assertTrue(out.contains("\"pending\":0"), out);
assertTrue(out.contains("\"heldCount\":3"),
"the honest count beside pending: 0 — three messages really are held: " + out);
assertTrue(out.contains("\"heldDurable\":true"),
"must state the durability fact, not leave pending as the only number next to held[]: " + out);
}
/**
* fleetd #440: {@code heldDurable} must be a derived fact, not a literal — so it can report
* {@code false} when the channel behind it says held mail is not durable (a non-durable queue,
* or a consumer running with {@code autoAck=true}). A test that only ever asserts {@code true}
* repeats the defect this ticket fixes.
*/
@Test
void listReportsHeldDurableFalseWhenTheChannelSaysMailIsNotDurable() {
FakeHerdr h = new FakeHerdr();
SessionManager sessions = new SessionManager(workerService(h, "http://gx00.gw:8000", Set.of("gx00.gw")));
FakeLeadChannel channel = new FakeLeadChannel("mac-opus")
.withMailbox("mac-opus", LeadChannel.MailboxState.exists("mac-opus", 0, 1))
.withHeldDurable(false)
.hold(new LeadMessage("m1", "fleet01-lead", "mac-opus", "one"));
McpSchema.CallToolResult res = FleetMcp.listFleet(
workerService(h, "http://gx00.gw:8000", Set.of("gx00.gw")), sessions, null,
FleetMcp.CapacitySource.none(), new FleetMcp.HealthCoverageSource(() -> "off"),
FleetMcp.QuarantineSource.none(), Map.of(), "",
new FleetMcp.CoordinationSource(channel, List.of()));
String out = textOf(res);
assertTrue(out.contains("\"heldDurable\":false"),
"heldDurable must follow the channel, not a hardcoded true: " + out);
}
// ── fleetd #421: a lead reads (never consumes) its own held peer mail ──────────────────────
@Test
void pollWithCoordIdReturnsTheFullBodyWithoutAckingAndLeavesItHeld() {
FakeLeadChannel channel = new FakeLeadChannel("mac-opus")
.hold(new LeadMessage("m1", "fleet01-lead", "mac-opus", "x".repeat(200)));
McpSchema.CallToolResult first = FleetMcp.poll(messages, channel, null, null, "mac-opus");
assertNotEquals(Boolean.TRUE, first.isError(), textOf(first));
String out1 = textOf(first);
assertTrue(out1.contains("\"msgId\":\"m1\""), out1);
assertTrue(out1.contains("\"from\":\"fleet01-lead\""), out1);
assertTrue(out1.contains("\"content\":\"" + "x".repeat(200) + "\""),
"the coordId route must return the FULL body, unlike fleet_list's preview: " + out1);
// Read again: identical bodies, and nothing was acked — peek() still holds it.
McpSchema.CallToolResult second = FleetMcp.poll(messages, channel, null, null, "mac-opus");
assertEquals(out1, textOf(second), "peek is non-destructive — reading twice must return the same bodies");
assertTrue(channel.acked().isEmpty(), "a read must never ack — that is the whole point of the ticket");
assertEquals(1, channel.peek().size(), "the message must still be held after being read");
}
@Test
void pollWithCoordIdRefusesAPeersCoordIdInsteadOfReturningTheWrongMailOrNothing() {
FakeLeadChannel channel = new FakeLeadChannel("mac-opus")
.hold(new LeadMessage("m1", "fleet01-lead", "mac-opus", "secret coordination body"));
// The ORIGINAL fleetd #421 confusion: passing a PEER's id where a self-address was meant.
McpSchema.CallToolResult res = FleetMcp.poll(messages, channel, null, null, "fleet01-lead");
assertEquals(Boolean.TRUE, res.isError());
String out = textOf(res);
assertFalse(out.contains("secret coordination body"),
"a refused read must never leak the body it refused to return: " + out);
assertTrue(out.contains("mac-opus"), "the error must name this daemon's own coordId: " + out);
}
@Test
void pollWithCoordIdErrorsHonestlyWhenLeadCoordinationIsNotConfigured() {
McpSchema.CallToolResult res = FleetMcp.poll(messages, null, null, null, "mac-opus");
assertEquals(Boolean.TRUE, res.isError());
assertTrue(textOf(res).toLowerCase().contains("coordinat"), textOf(res));
}
@Test
void listReportsEachDeclaredPeersLiveReachability() {
FakeHerdr h = new FakeHerdr();
@@ -752,6 +860,9 @@ class FleetMcpTest {
@Override
public String selfCoordId() { return "mac-opus"; }
@Override
public boolean heldDurable() { return true; }
@Override
public MailboxState inspect(String coordId) {
started.countDown();
@@ -0,0 +1,114 @@
package dev.ltms.fleet.mcp;
import dev.ltms.fleet.config.ConfigRef;
import dev.ltms.fleet.config.FleetConfig;
import dev.ltms.fleet.guard.SubscriptionGuard;
import dev.ltms.fleet.herdr.AgentControl;
import dev.ltms.fleet.herdr.FakeHerdr;
import dev.ltms.fleet.herdr.WorkspaceControl;
import dev.ltms.fleet.member.ClaudeCodeLauncher;
import dev.ltms.fleet.member.CompositePeerLauncher;
import dev.ltms.fleet.peer.MemberRole;
import dev.ltms.fleet.peer.PeerLauncher;
import dev.ltms.fleet.peer.SpawnRequest;
import dev.ltms.fleet.placement.BackendQuarantine;
import org.junit.jupiter.api.Test;
import org.junit.jupiter.api.io.TempDir;
import java.nio.file.Files;
import java.nio.file.Path;
import java.util.List;
import java.util.Map;
import java.util.Set;
import static org.junit.jupiter.api.Assertions.assertEquals;
import static org.junit.jupiter.api.Assertions.assertTrue;
/**
* fleetd #425: {@code fleet_profiles}' {@code "default"} field was captured once at boot
* ({@code cfg.effectiveDefaultProfile()}, frozen into {@code CompositePeerLauncher.defaultProfile}
* at construction) while an unqualified spawn resolves the same underlying key
* ({@code fleet.developers}' first entry) live, on every call. Reordering {@code fleet.developers}
* and reloading changed where a spawn landed without ever changing what {@code fleet_profiles}
* reported — a lead following {@code CLAUDE.md}'s "check {@code fleet_profiles} once per session"
* instruction was told a stale answer.
*
* <p>This test drives the exact caller {@code fleet_profiles} uses —
* {@link FleetMcp#profilesView(PeerLauncher, FleetMcp.QuarantineSource, FleetMcp.OutageSource)} —
* against a real, reloadable {@link ConfigRef}, so it fails if the reporting path is ever recoupled
* to a frozen value instead of {@link CompositePeerLauncher#defaultProfile()}'s live answer.
*/
class FleetProfilesLiveDefaultTest {
/** A minimal fleetd.yaml whose dev pool is {@code profilesInOrder}, in that definition order. */
private static String yamlWithDevPool(String... profilesInOrder) {
StringBuilder devPool = new StringBuilder();
for (int i = 0; i < profilesInOrder.length; i++) {
devPool.append(" slot").append(i).append(":\n profile: ")
.append(profilesInOrder[i]).append('\n');
}
return """
bind:
host: 127.0.0.1
port: 8765
herdrSocket: ~/.config/herdr/herdr.sock
profiles:
opus:
baseUrl: http://gx00.gw:8000
model: opus-coder
sonnet:
baseUrl: http://gx00.gw:8000
model: sonnet-coder
guard:
offSubscriptionHosts:
- gx00.gw
fleet:
developers:
""" + devPool;
}
@Test
void fleetProfilesDefaultTracksALiveDevPoolReorderAfterReload(@TempDir Path dir) throws Exception {
Path f = dir.resolve("fleetd.yaml");
Files.writeString(f, yamlWithDevPool("opus", "sonnet"));
ConfigRef ref = new ConfigRef(f, FleetConfig.load(f));
Map<String, FleetConfig.Profile> profiles = Map.of(
"opus", new FleetConfig.Profile("opus", "http://gx00.gw:8000", "opus-coder", null,
"FLEETD_WORKER_TOKEN", List.of("claude"), "tab", "fleetd-workers",
"worker: {profile} #{n}", null, null, null),
"sonnet", new FleetConfig.Profile("sonnet", "http://gx00.gw:8000", "sonnet-coder", null,
"FLEETD_WORKER_TOKEN", List.of("claude"), "tab", "fleetd-workers",
"worker: {profile} #{n}", null, null, null));
FakeHerdr herdr = new FakeHerdr();
ClaudeCodeLauncher adapter = new ClaudeCodeLauncher(new AgentControl(herdr), new WorkspaceControl(herdr),
new SubscriptionGuard(Set.of("gx00.gw")), profiles, "opus", _ -> null);
PeerLauncher workers = new CompositePeerLauncher(
List.of(adapter), "opus", ref, _ -> 0, BackendQuarantine.none());
assertReportedDefaultMatchesAnUnqualifiedSpawn(workers, "opus");
Files.writeString(f, yamlWithDevPool("sonnet", "opus"));
ConfigRef.Outcome out = ref.reload();
assertTrue(out.applied(), () -> "reload should apply cleanly: " + out.error());
assertReportedDefaultMatchesAnUnqualifiedSpawn(workers, "sonnet");
}
/**
* Asserts BOTH that {@code fleet_profiles}' {@code "default"} equals {@code expected}, AND that
* it equals what a real unqualified {@code MemberRole#DEV} spawn actually gets placed on right
* now — the two facts fleetd #425 found disagreeing.
*/
private static void assertReportedDefaultMatchesAnUnqualifiedSpawn(PeerLauncher workers, String expected) {
Map<String, Object> view = FleetMcp.profilesView(
workers, FleetMcp.QuarantineSource.none(), FleetMcp.OutageSource.none());
assertEquals(expected, view.get("default"),
"fleet_profiles' \"default\" must be the live dev-pool answer, not a boot-time snapshot");
String placed = workers.spawn(
new SpawnRequest(null, null, null, null, null, MemberRole.DEV)).profile();
assertEquals(expected, placed,
"sanity: the profile an unqualified dev spawn actually lands on");
}
}
@@ -0,0 +1,109 @@
package dev.ltms.fleet.mcp;
import dev.ltms.fleet.config.FleetConfig;
import dev.ltms.fleet.guard.SubscriptionGuard;
import dev.ltms.fleet.herdr.AgentControl;
import dev.ltms.fleet.herdr.FakeHerdr;
import dev.ltms.fleet.herdr.WorkspaceControl;
import dev.ltms.fleet.member.ClaudeCodeLauncher;
import dev.ltms.fleet.member.CompositePeerLauncher;
import dev.ltms.fleet.member.HerdrPeerLauncher;
import dev.ltms.fleet.peer.PeerLauncher;
import dev.ltms.fleet.placement.BackendOutagePolicy;
import dev.ltms.fleet.placement.BackendQuarantine;
import dev.ltms.fleet.placement.PlacementPolicies;
import org.junit.jupiter.api.Test;
import java.util.List;
import java.util.Map;
import java.util.Set;
import static org.junit.jupiter.api.Assertions.assertEquals;
import static org.junit.jupiter.api.Assertions.assertFalse;
/**
* fleetd #422 follow-up: {@code fleet_profiles}/{@code GET /profiles} — the reporting surface a
* lead actually reads — must let it tell apart the three states {@link
* dev.ltms.fleet.peer.PeerLauncher#disabledModels()} alone collapses into one empty set: no
* {@code models:} block at all, a block armed with nothing currently off, and a block with N
* models off. See {@link dev.ltms.fleet.peer.PeerLauncher.ModelGateState}'s javadoc for why a
* bare {@code disabledModels()} read cannot make this distinction, and {@link
* FleetMcp#profilesView} for where {@code modelGateArmed} is added alongside the existing {@code
* modelsOff} key.
*
* <p>Every assertion here goes through {@link FleetMcp#profilesView}, never {@code
* PeerLauncher.modelGateState()} directly — {@code CompositePeerLauncherTest} already proves the
* accessor itself; this class proves the surface a lead reads (fleet_profiles / GET /profiles)
* renders what that accessor reports.
*/
class FleetProfilesModelGateStateTest {
private static FleetConfig.Profile profile(String name, String model) {
return new FleetConfig.Profile(name, "http://gx00.gw:8000", model, null, "FLEETD_WORKER_TOKEN",
null, "tab", "fleetd-workers", "worker: {profile} #{n}", null, null, null);
}
private static FleetMcp.QuarantineSource noQuarantine() {
return new FleetMcp.QuarantineSource(_ -> null, BackendQuarantine.none(), _ -> false);
}
private static HerdrPeerLauncher claudeAdapter(FakeHerdr h, Map<String, FleetConfig.Profile> profiles) {
return new ClaudeCodeLauncher(new AgentControl(h), new WorkspaceControl(h),
new SubscriptionGuard(Set.of("gx00.gw")), profiles, "local", _ -> "tok");
}
/**
* State 1: no {@code models:} block at all — a plain {@code ClaudeCodeLauncher} (no {@code
* models:} supplier exists for it to read) has nothing to gate against, matching the fleet01
* host measured for this ticket: {@code grep -c '^models:' fleetd.yaml} returns 0 there.
*/
@Test
void noModelsBlockReportsGateNotArmed() {
FakeHerdr h = new FakeHerdr();
Map<String, FleetConfig.Profile> profiles = Map.of("local", profile("local", "deepseek-v4-flash"));
PeerLauncher workers = claudeAdapter(h, profiles);
Map<String, Object> view = FleetMcp.profilesView(workers, noQuarantine(), FleetMcp.OutageSource.none());
assertEquals(Boolean.FALSE, view.get("modelGateArmed"),
"no models: block to read from — nothing is gated, and nothing can be");
assertFalse(view.containsKey("modelsOff"), "nothing configured, so no off set to report either");
}
/** State 2: a {@code models:} block is present, but nothing in it is currently turned off. */
@Test
void modelsBlockWithNothingOffReportsGateArmedAndZeroOff() {
FakeHerdr h = new FakeHerdr();
Map<String, FleetConfig.Profile> profiles = Map.of("local", profile("local", "deepseek-v4-flash"));
FleetConfig.Models models = new FleetConfig.Models(
List.of(new FleetConfig.Models.ModelEntry("deepseek-v4-flash", true)));
PeerLauncher workers = new CompositePeerLauncher(List.of(claudeAdapter(h, profiles)), "local", profiles,
PlacementPolicies.fixed(), _ -> 0, null, BackendQuarantine.none(),
new BackendOutagePolicy(System::nanoTime), () -> models);
Map<String, Object> view = FleetMcp.profilesView(workers, noQuarantine(), FleetMcp.OutageSource.none());
assertEquals(Boolean.TRUE, view.get("modelGateArmed"),
"a models: block is present, so the gate is armed even though nothing is off yet");
assertFalse(view.containsKey("modelsOff"),
"nothing is off, so the key stays absent — an empty list here would be indistinguishable "
+ "from today's modelsOff omission, exactly the ambiguity modelGateArmed exists to remove");
}
/** State 3: a {@code models:} block is present with one model currently turned off. */
@Test
void modelsBlockWithModelsOffReportsGateArmedAndTheOffSet() {
FakeHerdr h = new FakeHerdr();
Map<String, FleetConfig.Profile> profiles = Map.of("local", profile("local", "deepseek-v4-flash"));
FleetConfig.Models models = new FleetConfig.Models(
List.of(new FleetConfig.Models.ModelEntry("deepseek-v4-flash", false)));
PeerLauncher workers = new CompositePeerLauncher(List.of(claudeAdapter(h, profiles)), "local", profiles,
PlacementPolicies.fixed(), _ -> 0, null, BackendQuarantine.none(),
new BackendOutagePolicy(System::nanoTime), () -> models);
Map<String, Object> view = FleetMcp.profilesView(workers, noQuarantine(), FleetMcp.OutageSource.none());
assertEquals(Boolean.TRUE, view.get("modelGateArmed"));
assertEquals(List.of("deepseek-v4-flash"), view.get("modelsOff"));
}
}
@@ -3,6 +3,7 @@ package dev.ltms.fleet.member;
import ch.qos.logback.classic.Logger;
import ch.qos.logback.classic.spi.ILoggingEvent;
import ch.qos.logback.core.read.ListAppender;
import dev.ltms.fleet.config.ConfigRef;
import dev.ltms.fleet.config.FleetConfig;
import dev.ltms.fleet.guard.SubscriptionGuard;
import dev.ltms.fleet.herdr.Agent;
@@ -19,11 +20,15 @@ import dev.ltms.fleet.peer.PeerUnreachableException;
import dev.ltms.fleet.peer.SpawnRequest;
import dev.ltms.fleet.placement.BackendOutagePolicy;
import dev.ltms.fleet.placement.BackendQuarantine;
import dev.ltms.fleet.placement.PlacementDecision;
import dev.ltms.fleet.placement.PlacementException;
import dev.ltms.fleet.placement.PlacementPolicies;
import org.junit.jupiter.api.Test;
import org.junit.jupiter.api.io.TempDir;
import org.slf4j.LoggerFactory;
import java.nio.file.Files;
import java.nio.file.Path;
import java.util.EnumSet;
import java.util.HashMap;
import java.util.LinkedHashMap;
@@ -141,6 +146,13 @@ class CompositePeerLauncherTest {
null, null, credentialId, null);
}
/** Like {@link #stubWorker(String)}, but with an explicit {@code model:} for fleetd #422 tests. */
private static FleetConfig.Profile stubWorkerModel(String profile, String model) {
return new FleetConfig.Profile(profile, "http://gx00.gw:8000", model,
null, "FLEETD_WORKER_TOKEN", List.of("claude"), "tab", "fleetd-workers",
"w #{n}", null, null, null, null, null, null, null, null, null);
}
/**
* An <em>order-preserving</em> profile map. Never {@code Map.of} here: its iteration order is
* salted per JVM run, and the weighted policy breaks an exact-weight tie on candidate order —
@@ -525,6 +537,31 @@ class CompositePeerLauncherTest {
assertEquals("claude", h.profile(), "the returned handle carries the resolved default profile");
}
/**
* fleetd #435: {@code FixedPlacementPolicy} — the default placement policy every config uses
* unless {@code placement:} is set — never consulted {@code maxLoad}, so an unqualified spawn
* (a blank profile, the normal delegation path) landed on a capped default anyway. Measured on
* 7667727: a single dev profile at {@code maxLoad: 1} with 1 live, under {@code fixed()},
* returned "SPAWNED on profile=a". This test goes through {@code CompositePeerLauncher.spawn}
* with a blank profile, not the policy in isolation, so it proves the caller actually reaches
* the fixed default's new cap check rather than only the {@code select} method.
*/
@Test
void fixedPolicyGatesDefaultProfileAtMaxLoadOnUnqualifiedSpawn() {
FakeHerdr herdr = new FakeHerdr();
Map<String, FleetConfig.Profile> profiles = ordered(
"a", stubWorker("a", 1.0f, 1),
"b", stubWorker("b", 1.0f, null));
StubLauncher adapter = new StubLauncher("claude", herdr, profiles, "a", Set.of());
CompositePeerLauncher composite = new CompositePeerLauncher(
List.of(adapter), "a", profiles, PlacementPolicies.fixed(), name -> "a".equals(name) ? 1 : 0);
PeerHandle h = composite.spawn(new SpawnRequest(null, null, null));
assertEquals("b", h.profile(),
"the default profile a is at maxLoad, so fixed placement must fall through to b");
assertEquals(0, adapter.spawnCount("a"), "a is never spawned — it is already at cap");
}
@Test
void weightedPolicyGatesProfileAtMaxLoad() {
FakeHerdr herdr = new FakeHerdr();
@@ -908,6 +945,89 @@ class CompositePeerLauncherTest {
assertTrue(e.getMessage().contains("maxLoad"), e.getMessage());
}
// ── fleetd #425: defaultProfile()/defaultProfileFor() must track a live reload ─────────────
/** A minimal fleetd.yaml whose dev pool is {@code profilesInOrder}, in that definition order. */
private static String yamlWithDevPool(String... profilesInOrder) {
StringBuilder devPool = new StringBuilder();
for (int i = 0; i < profilesInOrder.length; i++) {
devPool.append(" slot").append(i).append(":\n profile: ")
.append(profilesInOrder[i]).append('\n');
}
return """
bind:
host: 127.0.0.1
port: 8765
herdrSocket: ~/.config/herdr/herdr.sock
profiles:
opus:
baseUrl: http://gx00.gw:8000
model: opus-coder
sonnet:
baseUrl: http://gx00.gw:8000
model: sonnet-coder
guard:
offSubscriptionHosts:
- gx00.gw
fleet:
developers:
""" + devPool;
}
/**
* Criterion 1 (fleetd #425): reorder {@code fleet.developers}, reload, and assert the reported
* default ({@link CompositePeerLauncher#defaultProfile()} — what {@code fleet_profiles}' {@code
* "default"} is built from, see {@code FleetMcp.profilesView}) matches what an unqualified
* {@code MemberRole#DEV} spawn is actually placed on, both before and after the reorder. Asserts
* {@code applied()} so the test proves the reload actually took, not that nothing changed.
*/
@Test
void defaultProfileTracksALiveDevPoolReorderAfterReload(@TempDir Path dir) throws Exception {
Path f = dir.resolve("fleetd.yaml");
Files.writeString(f, yamlWithDevPool("opus", "sonnet"));
ConfigRef ref = new ConfigRef(f, FleetConfig.load(f));
FakeHerdr herdr = new FakeHerdr();
StubLauncher adapter = new StubLauncher("claude", herdr, threeProfiles(), "opus", Set.of());
CompositePeerLauncher composite = new CompositePeerLauncher(
List.of(adapter), "opus", ref, _ -> 0, BackendQuarantine.none());
assertEquals("opus", composite.defaultProfile(),
"reported default starts at the dev pool's first entry");
assertEquals("opus", composite.spawn(
new SpawnRequest(null, null, null, null, null, MemberRole.DEV)).profile(),
"an unqualified dev spawn must land on the same profile that was just reported");
Files.writeString(f, yamlWithDevPool("sonnet", "opus"));
ConfigRef.Outcome out = ref.reload();
assertTrue(out.applied(), () -> "reload should apply cleanly: " + out.error());
assertEquals("sonnet", composite.defaultProfile(),
"the reported default must follow the reorder with no daemon restart");
assertEquals("sonnet", composite.spawn(
new SpawnRequest(null, null, null, null, null, MemberRole.DEV)).profile(),
"and it must still be exactly what an unqualified spawn actually gets");
}
/**
* Criterion 2 — the mirror, and the load-bearing half (fleetd #425): with NOTHING configured (no
* profiles at all, hence an empty pool for every role), the frozen {@code defaultProfile} field
* is still what gets reported. A fix that always returns {@code poolFor(role).getFirst()} with no
* empty-pool fallback throws or returns the wrong thing here even though criterion 1 above still
* passes — this is the test that catches it.
*/
@Test
void defaultProfileFallsBackToTheFrozenFieldWhenNothingIsConfiguredAtAll() {
FakeHerdr herdr = new FakeHerdr();
StubLauncher adapter = new StubLauncher("claude", herdr, Map.of(), "opus", Set.of());
CompositePeerLauncher composite = new CompositePeerLauncher(
List.of(adapter), "opus", Map.of(), PlacementPolicies.fixed(), _ -> 0);
assertEquals("opus", composite.defaultProfile(),
"with no profiles configured at all, the frozen field is the only answer available");
assertEquals("opus", composite.defaultProfileFor(MemberRole.DEV));
}
// ── CB-578 stage B: a BACKEND_EXHAUSTED classification quarantines the credential ──────────
@Test
@@ -974,6 +1094,98 @@ class CompositePeerLauncherTest {
assertEquals(0, adapter.spawnCount("sol"));
}
/**
* fleetd #425 rework, acceptance 1: {@link CompositePeerLauncher#routedProfileFor} must apply
* the SAME quarantine filtering {@link CompositePeerLauncher#spawn} does, under {@code fixed()}
* — the DEFAULT placement policy, deliberately not {@code weighted()} (which the regressed
* round's own tests all used, and which never exercises {@code FixedPlacementPolicy}'s own
* inline filter). This is the exact defect: the previous round's {@code defaultProfileFor}
* blindly returns the pool's first entry ("sol", quarantined here) with no awareness of
* quarantine at all, which is what turned a routine unqualified spawn into a hard throw once
* {@code acquireWithWorktree} pre-resolved through it.
*/
@Test
void routedProfileForSkipsAQuarantinedPoolFirstProfileUnderFixedPolicy() {
FakeHerdr herdr = new FakeHerdr();
Map<String, FleetConfig.Profile> profiles = ordered(
"sol", stubWorker("sol", "shared-openai"),
"b", stubWorker("b"));
StubLauncher adapter = new StubLauncher("claude", herdr, profiles, "sol", Set.of());
BackendQuarantine quarantine = new BackendQuarantine(() -> 0L, TimeUnit.MINUTES.toNanos(30));
quarantine.quarantine("shared-openai");
CompositePeerLauncher composite = new CompositePeerLauncher(List.of(adapter), "sol", profiles,
PlacementPolicies.fixed(), _ -> 0, null, quarantine);
assertEquals("b", composite.routedProfileFor(MemberRole.DEV),
"sol (the pool's first entry) is quarantined, so the routed answer must be b");
assertEquals("sol", composite.defaultProfileFor(MemberRole.DEV),
"sanity: defaultProfileFor stays blind to quarantine — that's the gap routedProfileFor closes");
assertEquals(0, adapter.spawnCount("sol"), "routedProfileFor never spawns anything");
assertEquals(0, adapter.spawnCount("b"), "routedProfileFor never spawns anything");
}
/**
* fleetd #444: {@link PlacementDecision} exists to close the window between {@link
* CompositePeerLauncher#place} and {@link CompositePeerLauncher#spawn(SpawnRequest,
* PlacementDecision)} — the placement state must be free to move in that window without the
* held decision being re-checked against the new state. Every quarantine test above resolves
* and spawns in one call, so none of them ever open that window; this test is the one that
* does: "sol" is placed FIRST, while nothing is quarantined yet, and only THEN is its
* credential quarantined, before the held decision is spawned.
*
* <p>This is the test that tells the real override apart from the alternative body the ticket
* measured: routing {@code decision.profile()} straight to its adapter (the real override)
* never re-runs {@code enforceNotQuarantined}, so the spawn against the held decision still
* succeeds on sol. Re-entering {@code spawn(req.withProfile(decision.profile()))} instead
* lands in the explicit-profile branch, which refuses a now-quarantined sol outright — before
* this test existed, replacing the real override's body with that re-entering call left the
* whole suite green.
*/
@Test
void spawnHonorsAPlacementDecisionEvenAfterItsProfileIsQuarantinedInTheWindowAfterPlace() {
FakeHerdr herdr = new FakeHerdr();
Map<String, FleetConfig.Profile> profiles = ordered(
"sol", stubWorker("sol", "shared-openai"),
"b", stubWorker("b"));
// The adapter's OWN fallback default is "b", deliberately different from the profile place()
// decides ("sol") — see the note below on why this must not be "sol" too.
StubLauncher adapter = new StubLauncher("claude", herdr, profiles, "b", Set.of());
BackendQuarantine quarantine = new BackendQuarantine(() -> 0L, TimeUnit.MINUTES.toNanos(30));
CompositePeerLauncher composite = new CompositePeerLauncher(List.of(adapter), "sol", profiles,
PlacementPolicies.fixed(), _ -> 0, null, quarantine);
// 1. Resolve BEFORE anything is quarantined — sol (definition order first, fixed policy) wins.
// composite's own defaultProfile ("sol", the constructor arg above) never enters this: the
// pool poolFor(DEV) resolves to is never empty here, so place() only ever reads that field as
// a fallback for an empty pool, which this test does not exercise.
PlacementDecision decision = composite.place(MemberRole.DEV);
assertEquals("sol", decision.profile(), "sanity: nothing is quarantined yet, so sol is placed");
// 2. Move the placement state IN THE WINDOW between place() and spawn() — sol's credential
// is now quarantined. A fresh place()/spawn(req) pair would fall through to b instead; the
// held decision must not be re-evaluated against this new state at all.
quarantine.quarantine("shared-openai");
// 3. Spawn against the HELD decision, not a fresh resolve.
SpawnRequest req = new SpawnRequest(null, null, null, null, null, MemberRole.DEV);
PeerHandle handle = composite.spawn(req, decision);
assertEquals("sol", handle.profile(),
"the decision from place() is honored even though sol is now quarantined");
// A fixture whose adapter falls back to "sol" too would let an UNSTAMPED request (one
// routed but never given req.withProfile("sol")) land on spawnCount("sol") == 1 by
// COINCIDENCE, since StubLauncher.spawn falls back to its own defaultProfile whenever
// req.profileName() is blank. Giving the adapter "b" as its fallback instead means only an
// actually-stamped request can produce this count — an unstamped one would count against
// "b" and this assertion would fail.
assertEquals(1, adapter.spawnCount("sol"),
"the request that reached the delegate actually carried sol as its profile "
+ "(the adapter's own fallback default is 'b', so this can't happen by accident)");
assertEquals(0, adapter.spawnCount("b"),
"b must never be touched — neither as the decision's profile nor as an unstamped "
+ "request's accidental fallback");
}
@Test
void aQuarantineLiftsOnTheInjectedClockAndTheProfileBecomesSpawnableAgain() {
FakeHerdr herdr = new FakeHerdr();
@@ -1189,4 +1401,415 @@ class CompositePeerLauncherTest {
assertDoesNotThrow(() -> composite.spawn(new SpawnRequest("claude", null, null)));
assertDoesNotThrow(() -> composite.spawn(new SpawnRequest("gemini", null, null)));
}
// ── fleetd #422: the model on/off gate — a FOURTH, independent reason to refuse a spawn ──────
// ── (an operator decision, never a backend-reported outage) — never merged with quarantine ────
// ── or cool-off above ─────────────────────────────────────────────────────────────────────────
private static final BackendOutagePolicy NO_OUTAGE = new BackendOutagePolicy(() -> 0L);
/** Criterion 1: an explicit spawn onto a profile whose model is off is refused. */
@Test
void explicitSpawnOntoAnOffModelProfileIsRefused() {
FakeHerdr herdr = new FakeHerdr();
Map<String, FleetConfig.Profile> profiles = ordered(
"local", stubWorkerModel("local", "deepseek-v4-flash"),
"sonnet", stubWorkerModel("sonnet", "claude-sonnet-5"));
StubLauncher adapter = new StubLauncher("claude", herdr, profiles, "local", Set.of());
FleetConfig.Models models = new FleetConfig.Models(List.of(
new FleetConfig.Models.ModelEntry("deepseek-v4-flash", false)));
CompositePeerLauncher composite = new CompositePeerLauncher(List.of(adapter), "local", profiles,
PlacementPolicies.fixed(), _ -> 0, null, BackendQuarantine.none(), NO_OUTAGE,
() -> models);
PlacementException e = assertThrows(PlacementException.class,
() -> composite.spawn(new SpawnRequest("local", null, null)));
assertTrue(e.getMessage().contains("local"), "message names the profile: " + e.getMessage());
assertTrue(e.getMessage().contains("deepseek-v4-flash"),
"message names the model: " + e.getMessage());
assertTrue(e.getMessage().contains("operator") && e.getMessage().contains("turned off"),
"wording says the OPERATOR turned it off: " + e.getMessage());
// Distinct from quarantine/cool-off wording (criterion 1's explicit requirement).
assertFalse(e.getMessage().contains("quarantined"), "must not read like quarantine: " + e.getMessage());
assertFalse(e.getMessage().contains("cooling off"), "must not read like cool-off: " + e.getMessage());
assertEquals(0, adapter.spawnCount("local"), "the off-model profile is never delegated to");
}
/** A profile whose model is NOT off spawns normally even while another model is off. */
@Test
void explicitSpawnOntoAnEnabledModelProfileSucceedsWhileAnotherModelIsOff() {
FakeHerdr herdr = new FakeHerdr();
Map<String, FleetConfig.Profile> profiles = ordered(
"local", stubWorkerModel("local", "deepseek-v4-flash"),
"sonnet", stubWorkerModel("sonnet", "claude-sonnet-5"));
StubLauncher adapter = new StubLauncher("claude", herdr, profiles, "local", Set.of());
FleetConfig.Models models = new FleetConfig.Models(List.of(
new FleetConfig.Models.ModelEntry("deepseek-v4-flash", false)));
CompositePeerLauncher composite = new CompositePeerLauncher(List.of(adapter), "local", profiles,
PlacementPolicies.fixed(), _ -> 0, null, BackendQuarantine.none(), NO_OUTAGE,
() -> models);
assertDoesNotThrow(() -> composite.spawn(new SpawnRequest("sonnet", null, null)));
assertEquals(1, adapter.spawnCount("sonnet"));
}
/**
* Criterion 8: one {@code enabled: false} entry disables EVERY profile naming that model — the
* live example, {@code deepseek-v4-flash} backing both {@code local} and {@code local-direct}.
*/
@Test
void turningOffOneModelRefusesEveryProfileThatNamesIt() {
FakeHerdr herdr = new FakeHerdr();
Map<String, FleetConfig.Profile> profiles = ordered(
"local", stubWorkerModel("local", "deepseek-v4-flash"),
"local-direct", stubWorkerModel("local-direct", "deepseek-v4-flash"));
StubLauncher adapter = new StubLauncher("claude", herdr, profiles, "local", Set.of());
FleetConfig.Models models = new FleetConfig.Models(List.of(
new FleetConfig.Models.ModelEntry("deepseek-v4-flash", false)));
CompositePeerLauncher composite = new CompositePeerLauncher(List.of(adapter), "local", profiles,
PlacementPolicies.fixed(), _ -> 0, null, BackendQuarantine.none(), NO_OUTAGE,
() -> models);
assertThrows(PlacementException.class, () -> composite.spawn(new SpawnRequest("local", null, null)));
assertThrows(PlacementException.class,
() -> composite.spawn(new SpawnRequest("local-direct", null, null)),
"local-direct shares local's model, so it must be refused too");
assertEquals(0, adapter.spawnCount("local"));
assertEquals(0, adapter.spawnCount("local-direct"));
}
/** Criterion 3: an old-style entry with no {@code enabled} field never refuses a spawn. */
@Test
void anEntryWithNoEnabledFieldNeverRefusesASpawn() {
FakeHerdr herdr = new FakeHerdr();
Map<String, FleetConfig.Profile> profiles = Map.of("local", stubWorkerModel("local", "deepseek-v4-flash"));
StubLauncher adapter = new StubLauncher("claude", herdr, profiles, "local", Set.of());
// The back-compat single-arg ModelEntry constructor — no enabled field in the shape at all.
FleetConfig.Models models = new FleetConfig.Models(
List.of(new FleetConfig.Models.ModelEntry("deepseek-v4-flash")));
CompositePeerLauncher composite = new CompositePeerLauncher(List.of(adapter), "local", profiles,
PlacementPolicies.fixed(), _ -> 0, null, BackendQuarantine.none(), NO_OUTAGE,
() -> models);
assertDoesNotThrow(() -> composite.spawn(new SpawnRequest("local", null, null)));
assertEquals(1, adapter.spawnCount("local"));
}
/** A model absent from {@code models.allow:} entirely (nothing to gate against) is never refused. */
@Test
void aModelNotConfiguredInTheAllowListIsNeverGated() {
FakeHerdr herdr = new FakeHerdr();
Map<String, FleetConfig.Profile> profiles = Map.of("local", stubWorkerModel("local", "unlisted-model"));
StubLauncher adapter = new StubLauncher("claude", herdr, profiles, "local", Set.of());
FleetConfig.Models models = new FleetConfig.Models(List.of(
new FleetConfig.Models.ModelEntry("deepseek-v4-flash", false)));
CompositePeerLauncher composite = new CompositePeerLauncher(List.of(adapter), "local", profiles,
PlacementPolicies.fixed(), _ -> 0, null, BackendQuarantine.none(), NO_OUTAGE,
() -> models);
assertDoesNotThrow(() -> composite.spawn(new SpawnRequest("local", null, null)));
}
/** Criterion 2: an unqualified spawn skips an off-model candidate and lands on another one. */
@Test
void placementSkipsAnOffModelProfileAndRoutesToAnotherOne() {
FakeHerdr herdr = new FakeHerdr();
Map<String, FleetConfig.Profile> profiles = ordered(
"local", stubWorkerModel("local", "deepseek-v4-flash"),
"sonnet", stubWorkerModel("sonnet", "claude-sonnet-5"));
StubLauncher adapter = new StubLauncher("claude", herdr, profiles, "local", Set.of());
FleetConfig.Models models = new FleetConfig.Models(List.of(
new FleetConfig.Models.ModelEntry("deepseek-v4-flash", false)));
CompositePeerLauncher composite = new CompositePeerLauncher(List.of(adapter), "local", profiles,
PlacementPolicies.weighted(), _ -> 0, null, BackendQuarantine.none(), NO_OUTAGE,
() -> models);
PeerHandle h = composite.spawn(new SpawnRequest(null, null, null));
assertEquals("sonnet", h.profile(), "local's model is off, so an unqualified spawn must land on sonnet");
assertEquals(0, adapter.spawnCount("local"));
}
/**
* Criterion 2's second half: when EVERY candidate's model is off, the placement exception must
* name that as the cause — not a generic "no candidates" message.
*/
@Test
void automaticPlacementNamesModelOffWhenEveryCandidateIsOffModel() {
FakeHerdr herdr = new FakeHerdr();
Map<String, FleetConfig.Profile> profiles = ordered(
"local", stubWorkerModel("local", "deepseek-v4-flash"),
"local-direct", stubWorkerModel("local-direct", "deepseek-v4-flash"));
StubLauncher adapter = new StubLauncher("claude", herdr, profiles, "local", Set.of());
FleetConfig.Models models = new FleetConfig.Models(List.of(
new FleetConfig.Models.ModelEntry("deepseek-v4-flash", false)));
CompositePeerLauncher composite = new CompositePeerLauncher(List.of(adapter), "local", profiles,
PlacementPolicies.weighted(), _ -> 0, null, BackendQuarantine.none(), NO_OUTAGE,
() -> models);
PlacementException e = assertThrows(PlacementException.class,
() -> composite.spawn(new SpawnRequest(null, null, null)),
"both candidates share the off model — nothing is available");
assertTrue(e.getMessage().contains("model") && e.getMessage().contains("turned off"),
"message names model-off as the cause, not a generic no-candidates message: "
+ e.getMessage());
}
/**
* fleetd #422 follow-up: {@code fixed} is the DEFAULT placement policy ({@code
* PlacementPolicies.fromName} returns it for an absent/blank name), and it built its own inline
* candidate filter instead of calling {@code PlacementPolicyUtil.available()} — so it never
* checked {@code modelOff()}. Mirrors {@code placementSkipsAnOffModelProfileAndRoutesToAnotherOne}
* above with only the policy swapped, to prove the gate now fires on the path most fleets use.
*/
@Test
void fixedPlacementSkipsAnOffModelProfileToo() {
FakeHerdr herdr = new FakeHerdr();
Map<String, FleetConfig.Profile> profiles = ordered(
"local", stubWorkerModel("local", "deepseek-v4-flash"),
"sonnet", stubWorkerModel("sonnet", "claude-sonnet-5"));
StubLauncher adapter = new StubLauncher("claude", herdr, profiles, "local", Set.of());
FleetConfig.Models models = new FleetConfig.Models(List.of(
new FleetConfig.Models.ModelEntry("deepseek-v4-flash", false)));
CompositePeerLauncher composite = new CompositePeerLauncher(List.of(adapter), "local", profiles,
PlacementPolicies.fixed(), _ -> 0, null, BackendQuarantine.none(), NO_OUTAGE,
() -> models);
PeerHandle h = composite.spawn(new SpawnRequest(null, null, null));
assertEquals("sonnet", h.profile(), "local's model is off, so an unqualified spawn must land on sonnet");
assertEquals(0, adapter.spawnCount("local"));
}
/**
* fleetd #425 rework, acceptance 2: same shape as {@link #fixedPlacementSkipsAnOffModelProfileToo}
* above, but through {@link CompositePeerLauncher#routedProfileFor} rather than an actual
* {@link CompositePeerLauncher#spawn} — the exact call {@code SessionManager.acquireWithWorktree}
* makes to pre-resolve a profile for provisioning. This is the fleetd #429 case named in the
* ticket: an operator turns a model off, and an unqualified worktree spawn must still route
* around it instead of throwing "names model, which the operator has turned off" — the throw
* {@link CompositePeerLauncher#enforceModelEnabled} raises only on the EXPLICIT-profile branch,
* which is exactly the branch the regressed round accidentally routed every worktree spawn onto.
*/
@Test
void routedProfileForSkipsAModelOffPoolFirstProfileUnderFixedPolicy() {
FakeHerdr herdr = new FakeHerdr();
Map<String, FleetConfig.Profile> profiles = ordered(
"local", stubWorkerModel("local", "deepseek-v4-flash"),
"sonnet", stubWorkerModel("sonnet", "claude-sonnet-5"));
StubLauncher adapter = new StubLauncher("claude", herdr, profiles, "local", Set.of());
FleetConfig.Models models = new FleetConfig.Models(List.of(
new FleetConfig.Models.ModelEntry("deepseek-v4-flash", false)));
CompositePeerLauncher composite = new CompositePeerLauncher(List.of(adapter), "local", profiles,
PlacementPolicies.fixed(), _ -> 0, null, BackendQuarantine.none(), NO_OUTAGE,
() -> models);
assertEquals("sonnet", composite.routedProfileFor(MemberRole.DEV),
"local (the pool's first entry) names an off model, so the routed answer must be sonnet");
assertEquals("local", composite.defaultProfileFor(MemberRole.DEV),
"sanity: defaultProfileFor stays blind to model-off — that's the gap routedProfileFor closes");
assertEquals(0, adapter.spawnCount("local"), "routedProfileFor never spawns anything");
assertEquals(0, adapter.spawnCount("sonnet"), "routedProfileFor never spawns anything");
}
/**
* Criterion 4: turning a model off/on is HOT — no restart — proven through a REAL
* {@code ConfigRef.reload()}, not a hand-rolled supplier swap. Also proves {@code models} is
* correctly reclassified: the reload's {@link ConfigRef.Outcome#applied()} is {@code true} and
* {@code "models"} never appears in {@link ConfigRef.Outcome#deferred()}.
*/
@Test
void modelOnOffIsHotReloadedThroughARealConfigRef(@TempDir Path dir) throws Exception {
Path yaml = dir.resolve("fleetd.yaml");
Files.writeString(yaml, """
bind:
port: 8080
profiles:
local:
baseUrl: http://local.gw:8000
model: deepseek-v4-flash
models:
allow:
- model: deepseek-v4-flash
""");
FleetConfig initial = FleetConfig.load(yaml);
ConfigRef configRef = new ConfigRef(yaml, initial);
FakeHerdr herdr = new FakeHerdr();
StubLauncher adapter = new StubLauncher("claude", herdr,
Map.of("local", stubWorker("local")), "local", Set.of());
CompositePeerLauncher composite = new CompositePeerLauncher(
List.of(adapter), "local", configRef, _ -> 0, BackendQuarantine.none(), NO_OUTAGE);
assertDoesNotThrow(() -> composite.spawn(new SpawnRequest("local", null, null)),
"the model starts enabled");
assertEquals(1, adapter.spawnCount("local"));
Files.writeString(yaml, """
bind:
port: 8080
profiles:
local:
baseUrl: http://local.gw:8000
model: deepseek-v4-flash
models:
allow:
- model: deepseek-v4-flash
enabled: false
""");
ConfigRef.Outcome outcome = configRef.reload();
assertTrue(outcome.applied(), "a models.allow on/off edit must apply live, never be refused");
assertFalse(outcome.deferred().contains("models"),
"models is hot-excluded now — it must never be reported as a deferred key");
PlacementException e = assertThrows(PlacementException.class,
() -> composite.spawn(new SpawnRequest("local", null, null)),
"the very next spawn must see the reload, with no restart");
assertTrue(e.getMessage().contains("deepseek-v4-flash"));
assertEquals(1, adapter.spawnCount("local"), "still just the one successful spawn from before");
// And back on, still hot, still no restart.
Files.writeString(yaml, """
bind:
port: 8080
profiles:
local:
baseUrl: http://local.gw:8000
model: deepseek-v4-flash
models:
allow:
- model: deepseek-v4-flash
enabled: true
""");
ConfigRef.Outcome reEnabled = configRef.reload();
assertTrue(reEnabled.applied());
assertDoesNotThrow(() -> composite.spawn(new SpawnRequest("local", null, null)));
assertEquals(2, adapter.spawnCount("local"));
}
/**
* fleetd #422 follow-up, acceptance criterion 2: {@link CompositePeerLauncher#modelGateState()}
* is LIVE — no restart — proven through a REAL {@link ConfigRef#reload()}, exactly like {@link
* #modelOnOffIsHotReloadedThroughARealConfigRef} above proves for the on/off gate itself. This
* single reload sequence walks through all three states the ticket asks for: no {@code models:}
* block, a block armed with nothing off, and a block with one model off — so a reload that flips
* between any of the three is proven live, not just the on/off edit within an already-armed block.
*/
@Test
void modelGateStateIsHotReloadedThroughARealConfigRef(@TempDir Path dir) throws Exception {
Path yaml = dir.resolve("fleetd.yaml");
Files.writeString(yaml, """
bind:
port: 8080
profiles:
local:
baseUrl: http://local.gw:8000
model: deepseek-v4-flash
""");
FleetConfig initial = FleetConfig.load(yaml);
ConfigRef configRef = new ConfigRef(yaml, initial);
FakeHerdr herdr = new FakeHerdr();
StubLauncher adapter = new StubLauncher("claude", herdr,
Map.of("local", stubWorker("local")), "local", Set.of());
CompositePeerLauncher composite = new CompositePeerLauncher(
List.of(adapter), "local", configRef, _ -> 0, BackendQuarantine.none(), NO_OUTAGE);
PeerLauncher.ModelGateState notConfigured = composite.modelGateState();
assertFalse(notConfigured.configured(), "no models: block in the config at all");
assertEquals(Set.of(), notConfigured.off());
Files.writeString(yaml, """
bind:
port: 8080
profiles:
local:
baseUrl: http://local.gw:8000
model: deepseek-v4-flash
models:
allow:
- model: deepseek-v4-flash
enabled: false
""");
assertTrue(configRef.reload().applied(), "adding a models: block must apply live, no restart");
PeerLauncher.ModelGateState armedWithOneOff = composite.modelGateState();
assertTrue(armedWithOneOff.configured(), "a models: block now exists — the gate is armed");
assertEquals(Set.of("deepseek-v4-flash"), armedWithOneOff.off());
Files.writeString(yaml, """
bind:
port: 8080
profiles:
local:
baseUrl: http://local.gw:8000
model: deepseek-v4-flash
models:
allow:
- model: deepseek-v4-flash
enabled: true
""");
assertTrue(configRef.reload().applied(), "flipping the entry back on must apply live too");
PeerLauncher.ModelGateState armedWithZeroOff = composite.modelGateState();
assertTrue(armedWithZeroOff.configured(),
"the block is still present — armed and reporting zero, not the same as no block at all");
assertEquals(Set.of(), armedWithZeroOff.off());
}
/**
* fleetd #422 follow-up, acceptance criterion 3: the invariant is that an absent {@code
* models:} block stays permitted and must never be fatal. Proved, not assumed — a config
* without one loads, validates, reports the gate as not configured, AND still spawns normally
* (no {@link PlacementException} from a gate that has nothing to check against), using the same
* production-shaped {@code Supplier<FleetConfig>} wiring {@code Fleetd.main} actually uses.
*/
@Test
void noModelsBlockConfigStillLoadsAndSpawnsNormally(@TempDir Path dir) throws Exception {
Path yaml = dir.resolve("fleetd.yaml");
Files.writeString(yaml, """
bind:
port: 8080
profiles:
local:
baseUrl: http://local.gw:8000
model: deepseek-v4-flash
""");
FleetConfig cfg = FleetConfig.load(yaml);
assertDoesNotThrow(cfg::validateAll, "a config with no models: block must load and validate cleanly");
ConfigRef configRef = new ConfigRef(yaml, cfg);
FakeHerdr herdr = new FakeHerdr();
StubLauncher adapter = new StubLauncher("claude", herdr,
Map.of("local", stubWorker("local")), "local", Set.of());
CompositePeerLauncher composite = new CompositePeerLauncher(
List.of(adapter), "local", configRef, _ -> 0, BackendQuarantine.none(), NO_OUTAGE);
assertFalse(composite.modelGateState().configured());
assertDoesNotThrow(() -> composite.spawn(new SpawnRequest("local", null, null)),
"no models: block means nothing to gate against — the spawn must go through");
assertEquals(1, adapter.spawnCount("local"));
}
/** {@code fleet_profiles}/{@code GET /profiles} must read the exact same live source the gate reads. */
@Test
void disabledModelsReportsWhatTheGateActuallyEnforces() {
FakeHerdr herdr = new FakeHerdr();
Map<String, FleetConfig.Profile> profiles = ordered(
"local", stubWorkerModel("local", "deepseek-v4-flash"),
"sonnet", stubWorkerModel("sonnet", "claude-sonnet-5"));
StubLauncher adapter = new StubLauncher("claude", herdr, profiles, "local", Set.of());
FleetConfig.Models models = new FleetConfig.Models(List.of(
new FleetConfig.Models.ModelEntry("deepseek-v4-flash", false),
new FleetConfig.Models.ModelEntry("claude-sonnet-5", true)));
CompositePeerLauncher composite = new CompositePeerLauncher(List.of(adapter), "local", profiles,
PlacementPolicies.fixed(), _ -> 0, null, BackendQuarantine.none(), NO_OUTAGE,
() -> models);
assertEquals(Set.of("deepseek-v4-flash"), composite.disabledModels());
}
/** A caller with no {@code models:} block (the map-based constructors) reports nothing off. */
@Test
void disabledModelsIsEmptyWithNoModelsConfigured() {
FakeHerdr herdr = new FakeHerdr();
PeerLauncher composite = composite(herdr);
assertEquals(Set.of(), composite.disabledModels());
}
}
@@ -28,11 +28,19 @@ public final class FakeLeadChannel implements LeadChannel {
private volatile IllegalStateException publishFailure;
/** Canned {@link #inspect} results by coord-id — absent for any coord-id not configured here. */
private final Map<String, MailboxState> mailboxes = new ConcurrentHashMap<>();
/** fleetd #440: matches {@link LeadMailbox}'s real default (durable queue + manual ack) unless overridden. */
private volatile boolean heldDurable = true;
public FakeLeadChannel(String selfCoordId) {
this.selfCoordId = selfCoordId;
}
/** Make {@link #heldDurable()} report {@code durable} — the fleetd #440 seam for the false case. */
public FakeLeadChannel withHeldDurable(boolean durable) {
this.heldDurable = durable;
return this;
}
/** Make {@link #inspect(String)} return {@code state} for {@code coordId} instead of "absent". */
public FakeLeadChannel withMailbox(String coordId, MailboxState state) {
mailboxes.put(coordId, state);
@@ -80,6 +88,11 @@ public final class FakeLeadChannel implements LeadChannel {
return mailboxes.getOrDefault(coordId, MailboxState.absent(coordId));
}
@Override
public boolean heldDurable() {
return heldDurable;
}
public List<LeadMessage> published() {
return List.copyOf(published);
}
@@ -207,6 +207,21 @@ class LeadMailboxTest {
}
}
/**
* fleetd #440: {@code heldDurable()} must be derived from what {@link LeadMailbox#own} actually
* did against the real broker — a durable queue declare plus a manual-ack consumer — not a
* hardcoded literal. This is the mutation-sensitive test: flip {@code own()}'s {@code autoAck}
* local to {@code true} (or its {@code durableQueue} local to {@code false}) and this must fail.
*/
@Test
void heldDurableReportsTrueBecauseTheQueueIsDurableAndTheConsumeIsManualAck() throws Exception {
String self = coordId("lead-held-durable");
try (LeadMailbox mailbox = LeadMailbox.open(uri(), self)) {
assertTrue(mailbox.heldDurable(),
"own() declares a durable queue and consumes with autoAck=false, so held mail is durable");
}
}
@Test
void inspectReportsAMissingMailboxAsAbsentRatherThanThrowing() throws Exception {
String nobody = coordId("lead-inspect-nobody");
@@ -90,6 +90,63 @@ class PlacementPolicyTest {
assertTrue(e.getMessage().contains("quarantined"), e.getMessage());
}
// --- fleetd #422: FixedPlacementPolicy must consult modelOff too, at BOTH filter sites ------
/** The default-profile fast path (:44) must skip a default whose model is off. */
@Test
void fixedSkipsModelOffDefault() {
PlacementPolicy policy = PlacementPolicies.fixed();
PlacementContext ctx = new PlacementContext("b",
List.of(PlacementCandidate.profile("a"), PlacementCandidate.profile("b")),
noSessions(), Set.of(), Set.of(), Set.of(), Set.of("b"));
assertEquals("a", policy.select(ctx).profile(),
"the default 'b' names an off model, so fixed falls through to the first available candidate");
}
/**
* The fallback walk (:48-53) must skip an off-model candidate too — exercised independently of
* the default-profile fast path by using no default at all, so this is the only filter that runs.
*/
@Test
void fixedFallbackWalkSkipsModelOffCandidate() {
PlacementPolicy policy = PlacementPolicies.fixed();
PlacementContext ctx = new PlacementContext(null,
List.of(PlacementCandidate.profile("a"), PlacementCandidate.profile("b")),
noSessions(), Set.of(), Set.of(), Set.of(), Set.of("a"));
assertEquals("b", policy.select(ctx).profile(),
"candidate 'a' names an off model, so the fallback walk skips it and picks 'b'");
}
@Test
void fixedThrowsWhenDefaultAndEveryCandidateModelOff() {
PlacementPolicy policy = PlacementPolicies.fixed();
PlacementContext ctx = new PlacementContext("b",
List.of(PlacementCandidate.profile("a"), PlacementCandidate.profile("b")),
noSessions(), Set.of(), Set.of(), Set.of(), Set.of("a", "b"));
PlacementException e = assertThrows(PlacementException.class, () -> policy.select(ctx));
assertTrue(e.getMessage().contains("turned off"),
"message names model-off as the cause: " + e.getMessage());
assertFalse(e.getMessage().contains("quarantined"), "must not read like quarantine: " + e.getMessage());
assertFalse(e.getMessage().contains("cooling off"), "must not read like cool-off: " + e.getMessage());
assertFalse(e.getMessage().contains("weight 0"), "must not read like weight-0: " + e.getMessage());
}
/**
* Quarantine still wins when a profile is both quarantined and model-off (mirrors {@code
* fixedThrowsWhenDefaultAndEveryCandidateQuarantined}'s priority over cooling off).
*/
@Test
void fixedReportsQuarantineNotModelOffWhenBothApply() {
PlacementPolicy policy = PlacementPolicies.fixed();
PlacementContext ctx = new PlacementContext("b",
List.of(PlacementCandidate.profile("a"), PlacementCandidate.profile("b")),
noSessions(), Set.of(), Set.of("a", "b"), Set.of(), Set.of("a", "b"));
PlacementException e = assertThrows(PlacementException.class, () -> policy.select(ctx));
assertTrue(e.getMessage().contains("quarantined"), e.getMessage());
assertFalse(e.getMessage().contains("turned off"),
"quarantine takes priority over model-off in the message: " + e.getMessage());
}
@Test
void roundRobinCyclesThroughAvailableProfiles() {
PlacementPolicy policy = PlacementPolicies.roundRobin();
@@ -404,6 +461,93 @@ class PlacementPolicyTest {
assertTrue(e.getMessage().contains("weight 0"), e.getMessage());
}
// --- fleetd #435: FixedPlacementPolicy must consult maxLoad too, at BOTH filter sites -------
/**
* The default-profile fast path must skip a capped default. Measured on 7667727 before this
* fix: a single dev profile at {@code maxLoad: 1} with 1 live, under {@code fixed()}, still
* returned "SPAWNED on profile=a" — the cap was advisory for every unqualified spawn.
*/
@Test
void fixedSkipsCappedDefault() {
PlacementPolicy policy = PlacementPolicies.fixed();
PlacementContext ctx = new PlacementContext("b",
List.of(PlacementCandidate.profile("a", 1.0f, null),
PlacementCandidate.profile("b", 1.0f, 1)),
name -> "b".equals(name) ? 1 : 0, Set.of(), Set.of(), Set.of());
assertEquals("a", policy.select(ctx).profile(),
"the default 'b' is at its maxLoad cap, so fixed falls through to the free candidate 'a'");
}
/**
* The fallback walk must skip a capped candidate too — exercised independently of the
* default-profile fast path by using no default at all, so this is the only filter that runs.
*/
@Test
void fixedFallbackWalkSkipsCappedCandidate() {
PlacementPolicy policy = PlacementPolicies.fixed();
PlacementContext ctx = new PlacementContext(null,
List.of(PlacementCandidate.profile("a", 1.0f, 1),
PlacementCandidate.profile("b", 1.0f, null)),
name -> "a".equals(name) ? 1 : 0, Set.of(), Set.of(), Set.of());
assertEquals("b", policy.select(ctx).profile(),
"candidate 'a' is at its maxLoad cap, so the fallback walk skips it and picks 'b'");
}
@Test
void fixedThrowsWhenDefaultAndEveryCandidateAtCap() {
PlacementPolicy policy = PlacementPolicies.fixed();
PlacementContext ctx = new PlacementContext("b",
List.of(PlacementCandidate.profile("a", 1.0f, 1),
PlacementCandidate.profile("b", 1.0f, 1)),
_ -> 1, Set.of(), Set.of(), Set.of());
PlacementException e = assertThrows(PlacementException.class, () -> policy.select(ctx));
assertTrue(e.getMessage().contains("maxLoad"), "message names the cap: " + e.getMessage());
assertTrue(e.getMessage().contains("1 cap"), "message names the cap value: " + e.getMessage());
assertFalse(e.getMessage().contains("quarantined"), "must not read like quarantine: " + e.getMessage());
assertFalse(e.getMessage().contains("turned off"), "must not read like model-off: " + e.getMessage());
}
/** The mirror: an uncapped default is still chosen, so the new term cannot exclude everything. */
@Test
void fixedStillReturnsUncappedDefault() {
PlacementPolicy policy = PlacementPolicies.fixed();
PlacementContext ctx = new PlacementContext("b",
List.of(PlacementCandidate.profile("a", 1.0f, null),
PlacementCandidate.profile("b", 1.0f, null)),
noSessions(), Set.of(), Set.of(), Set.of());
assertEquals("b", policy.select(ctx).profile(), "an uncapped default is returned unconditionally");
}
/** CB-585: an explicit {@code maxLoad: 0} on the default caps it at zero live members. */
@Test
void fixedSkipsMaxLoadZeroDefaultEvenWithZeroLiveWorkers() {
PlacementPolicy policy = PlacementPolicies.fixed();
PlacementContext ctx = new PlacementContext("b",
List.of(PlacementCandidate.profile("a", 1.0f, null),
PlacementCandidate.profile("b", 1.0f, 0)),
_ -> 0, Set.of(), Set.of(), Set.of());
assertEquals("a", policy.select(ctx).profile(),
"the default 'b' has maxLoad 0, so it is already at its cap with nobody live");
}
/**
* Quarantine still wins when a profile is both quarantined and at cap (mirrors {@code
* fixedReportsQuarantineNotModelOffWhenBothApply}'s priority over model-off).
*/
@Test
void fixedReportsQuarantineNotAtCapWhenBothApply() {
PlacementPolicy policy = PlacementPolicies.fixed();
PlacementContext ctx = new PlacementContext("b",
List.of(PlacementCandidate.profile("a", 1.0f, 1),
PlacementCandidate.profile("b", 1.0f, 1)),
_ -> 1, Set.of(), Set.of("a", "b"), Set.of());
PlacementException e = assertThrows(PlacementException.class, () -> policy.select(ctx));
assertTrue(e.getMessage().contains("quarantined"), e.getMessage());
assertFalse(e.getMessage().contains("maxLoad"),
"quarantine takes priority over at-cap in the message: " + e.getMessage());
}
@Test
void unknownPolicyNameThrows() {
assertThrows(IllegalArgumentException.class, () -> PlacementPolicies.fromName("random"));
@@ -6,12 +6,14 @@ import ch.qos.logback.classic.spi.ILoggingEvent;
import ch.qos.logback.core.read.ListAppender;
import dev.ltms.fleet.auth.MemberRegistry;
import dev.ltms.fleet.auth.MemberLifecycle;
import dev.ltms.fleet.config.ConfigRef;
import dev.ltms.fleet.config.FleetConfig;
import dev.ltms.fleet.guard.SubscriptionGuard;
import dev.ltms.fleet.herdr.AgentControl;
import dev.ltms.fleet.herdr.FakeHerdr;
import dev.ltms.fleet.herdr.WorkspaceControl;
import dev.ltms.fleet.member.ClaudeCodeLauncher;
import dev.ltms.fleet.member.CompositePeerLauncher;
import dev.ltms.fleet.msg.TestTurnTokens;
import dev.ltms.fleet.peer.Capability;
import dev.ltms.fleet.peer.CharterReceipt;
@@ -20,9 +22,15 @@ import dev.ltms.fleet.peer.PeerHandle;
import dev.ltms.fleet.peer.PeerLauncher;
import dev.ltms.fleet.peer.PeerUnreachableException;
import dev.ltms.fleet.peer.SpawnRequest;
import dev.ltms.fleet.placement.BackendQuarantine;
import dev.ltms.fleet.placement.PlacementPolicies;
import org.junit.jupiter.api.Test;
import org.junit.jupiter.api.io.TempDir;
import org.slf4j.LoggerFactory;
import java.nio.file.Files;
import java.nio.file.Path;
import java.util.LinkedHashMap;
import java.util.List;
import java.util.Map;
import java.util.Set;
@@ -33,6 +41,7 @@ import java.util.concurrent.Future;
import java.util.concurrent.TimeUnit;
import java.util.concurrent.atomic.AtomicInteger;
import java.util.function.LongSupplier;
import java.util.function.Supplier;
import static org.junit.jupiter.api.Assertions.*;
@@ -1991,4 +2000,296 @@ class SessionManagerTest {
sessions.rosterResolved();
assertEquals(2, handle.callCount(), "once resolved, the id must not be looked up again");
}
// ── fleetd #425 criterion 3: acquireWithWorktree must provision for the profile it actually
// spawns, never a name resolved before a live pool change is accounted for ────────────────────
/** Two profiles with distinct {@code cwd}/{@code parityOverlay}, and a dev pool of {@code first,second}. */
private static String worktreeReorderYaml(String first, String second) {
return """
bind:
host: 127.0.0.1
port: 8765
herdrSocket: ~/.config/herdr/herdr.sock
profiles:
a:
baseUrl: http://gx00.gw:8000
model: coder-a
b:
baseUrl: http://gx00.gw:8000
model: coder-b
guard:
offSubscriptionHosts:
- gx00.gw
fleet:
developers:
slot0:
profile: %s
slot1:
profile: %s
""".formatted(first, second);
}
@Test
void acquireWithWorktreeProvisionsTheOverlayForTheProfileActuallySpawned(
@TempDir Path dir) throws Exception {
Path f = dir.resolve("fleetd.yaml");
Files.writeString(f, worktreeReorderYaml("a", "b"));
ConfigRef ref = new ConfigRef(f, FleetConfig.load(f));
Map<String, FleetConfig.Profile> profiles = Map.of(
"a", new FleetConfig.Profile("a", "http://gx00.gw:8000", "coder-a", null,
"FLEETD_WORKER_TOKEN", List.of("claude"), "tab", "fleetd-workers",
"worker: {profile} #{n}", null, "/repo/a", List.of("a.mcp.json")),
"b", new FleetConfig.Profile("b", "http://gx00.gw:8000", "coder-b", null,
"FLEETD_WORKER_TOKEN", List.of("claude"), "tab", "fleetd-workers",
"worker: {profile} #{n}", null, "/repo/b", List.of("b.mcp.json")));
FakeHerdr herdr = new FakeHerdr();
ClaudeCodeLauncher adapter = new ClaudeCodeLauncher(new AgentControl(herdr), new WorkspaceControl(herdr),
new SubscriptionGuard(Set.of("gx00.gw")), profiles, "a", _ -> null);
PeerLauncher launcher = new CompositePeerLauncher(
List.of(adapter), "a", ref, _ -> 0, BackendQuarantine.none());
// The pool changes AFTER the composite/launcher is built, and BEFORE the unqualified
// worktree spawn — exactly the fleetd #425 scenario: the live pool's first entry is "b" by
// the time acquireWithWorktree runs, even though nothing here was rebuilt.
Files.writeString(f, worktreeReorderYaml("b", "a"));
ConfigRef.Outcome out = ref.reload();
assertTrue(out.applied(), () -> "reload should apply cleanly: " + out.error());
FakeWorktrees worktrees = new FakeWorktrees();
SessionManager sessions = new SessionManager(launcher, worktrees, () -> 0L);
MemberSession s = sessions.acquire(null, null, "/caller",
null, new WorktreeRequest("fleetd-425", null));
assertEquals("b", s.profile(),
"the live dev pool now starts at b, so the unqualified spawn must land there");
FakeWorktrees.OverlayCall overlay = worktrees.lastOverlay();
assertNotNull(overlay, "overlayParity must have been called");
assertEquals(List.of("b.mcp.json"), overlay.requested(),
"the worktree must be provisioned with profile b's overlay — the one actually "
+ "spawned — never a's, the pool's stale first entry");
}
/**
* The deterministic, mutation-pinning half of criterion 3: {@code launcher.defaultProfile()}
* only ever answers for {@link MemberRole#DEV} (see {@link CompositePeerLauncher#defaultProfile()}),
* so resolving a worktree spawn's profile through it — instead of through {@link
* PeerLauncher#defaultProfileFor(MemberRole)}, resolved against the CALLER's actual role — picks
* the wrong pool's answer for any role other than DEV. No reload or race is needed to see it: an
* ARCHITECT pool and a DEV pool that simply disagree, held constant, are enough.
*/
@Test
void acquireWithWorktreeForANonDevRoleUsesThatRolesPoolNotTheDevPool() {
Map<String, FleetConfig.Profile> profiles = Map.of(
"a", new FleetConfig.Profile("a", "http://gx00.gw:8000", "coder-a", null,
"FLEETD_WORKER_TOKEN", List.of("claude"), "tab", "fleetd-workers",
"worker: {profile} #{n}", null, "/repo/a", List.of("a.mcp.json")),
"b", new FleetConfig.Profile("b", "http://gx00.gw:8000", "coder-b", null,
"FLEETD_WORKER_TOKEN", List.of("claude"), "tab", "fleetd-workers",
"worker: {profile} #{n}", null, "/repo/b", List.of("b.mcp.json")));
FakeHerdr herdr = new FakeHerdr();
ClaudeCodeLauncher adapter = new ClaudeCodeLauncher(new AgentControl(herdr), new WorkspaceControl(herdr),
new SubscriptionGuard(Set.of("gx00.gw")), profiles, "a", _ -> null);
// developers -> a (first/only entry); architects -> b (first/only entry). The two pools
// disagree on purpose, so a role-blind resolution (DEV's answer, "a") is visibly wrong for
// an ARCHITECT spawn, which must land on "b".
FleetConfig.Fleet fleet = new FleetConfig.Fleet(Map.of(),
Map.of("s0", new FleetConfig.Slot("b")),
Map.of("s0", new FleetConfig.Slot("a")),
Map.of(), null);
PeerLauncher launcher = new CompositePeerLauncher(List.of(adapter), "a", profiles,
PlacementPolicies.fixed(), _ -> 0, fleet);
FakeWorktrees worktrees = new FakeWorktrees();
SessionManager sessions = new SessionManager(launcher, worktrees, () -> 0L);
MemberSession s = sessions.acquire(null, MemberRole.ARCHITECT, null, "/caller",
null, new WorktreeRequest("fleetd-425b", null));
assertEquals("b", s.profile(),
"an unqualified ARCHITECT worktree spawn must land on the architect pool's profile");
FakeWorktrees.OverlayCall overlay = worktrees.lastOverlay();
assertNotNull(overlay, "overlayParity must have been called");
assertEquals(List.of("b.mcp.json"), overlay.requested(),
"the worktree must be provisioned with profile b's overlay — the ARCHITECT pool's "
+ "answer, the one actually spawned — never a's, the DEV pool's answer that "
+ "launcher.defaultProfile() alone would have given");
}
/**
* fleetd #425 rework, acceptance 3: repoRoot, parityOverlay, AND the actual spawn must all name
* the SAME routed profile, proven on the ROUTED path — a quarantine skips the pool's first entry
* — not just the "pool reordered by a live reload" path the two tests above already cover.
*
* <p>This is the exact regression the rework fixes: the first round resolved
* {@code acquireWithWorktree}'s profile through {@code launcher.defaultProfileFor(memberRole)},
* which is blind to quarantine and just returns the pool's first entry ("a" here, quarantined).
* That name went on to provision repoRoot/overlay for "a", and then the spawn itself — now an
* EXPLICIT-profile spawn naming "a" — hit {@code CompositePeerLauncher.enforceNotQuarantined}
* and threw, where the pre-fix code (a blank-profile spawn) would have routed around "a" onto
* "b" without any trouble. {@code launcher.routedProfileFor(memberRole)} closes that gap by
* running the SAME quarantine-aware selection {@code spawn} itself uses, so all three — repoRoot,
* overlay, and the spawn — land on "b" together.
*/
@Test
void acquireWithWorktreeRoutesAroundAQuarantinedPoolFirstProfile() {
Map<String, FleetConfig.Profile> profiles = new LinkedHashMap<>();
profiles.put("a", new FleetConfig.Profile("a", "http://gx00.gw:8000", "coder-a", null,
"FLEETD_WORKER_TOKEN", List.of("claude"), "tab", "fleetd-workers",
"worker: {profile} #{n}", null, "/repo/a", List.of("a.mcp.json"),
null, null, null, null, null, null, null, null, "shared-cred", null));
profiles.put("b", new FleetConfig.Profile("b", "http://gx00.gw:8000", "coder-b", null,
"FLEETD_WORKER_TOKEN", List.of("claude"), "tab", "fleetd-workers",
"worker: {profile} #{n}", null, "/repo/b", List.of("b.mcp.json")));
FakeHerdr herdr = new FakeHerdr();
ClaudeCodeLauncher adapter = new ClaudeCodeLauncher(new AgentControl(herdr), new WorkspaceControl(herdr),
new SubscriptionGuard(Set.of("gx00.gw")), profiles, "a", _ -> null);
BackendQuarantine quarantine = new BackendQuarantine(() -> 0L, TimeUnit.MINUTES.toNanos(30));
quarantine.quarantine("shared-cred");
PeerLauncher launcher = new CompositePeerLauncher(List.of(adapter), "a", profiles,
PlacementPolicies.fixed(), _ -> 0, null, quarantine);
FakeWorktrees worktrees = new FakeWorktrees();
SessionManager sessions = new SessionManager(launcher, worktrees, () -> 0L);
MemberSession s = sessions.acquire(null, null, "/caller",
null, new WorktreeRequest("fleetd-425c", null));
assertEquals("b", s.profile(),
"a is quarantined, so the unqualified worktree spawn must route to b");
FakeWorktrees.OverlayCall overlay = worktrees.lastOverlay();
assertNotNull(overlay, "overlayParity must have been called");
assertEquals(List.of("b.mcp.json"), overlay.requested(),
"parityOverlay must be provisioned for b — the profile actually spawned, never a's, "
+ "the quarantined pool-first entry");
FakeWorktrees.RepoRootCall repoRootCall = worktrees.repoRootCalls().getLast();
assertTrue(repoRootCall.cwd().contains("/repo/b"),
"repoRoot must be resolved through b's effectiveCwd, not a's: " + repoRootCall.cwd());
}
/**
* fleetd #425 rework, round 2: this is the exact probe that found round 1's maxLoad
* regression. One dev profile ("a") is configured with {@code maxLoad: 1} and a liveCount
* pinned at 1 — permanently at cap — under {@code PlacementPolicies.fixed()}, the default
* policy, which deliberately never evaluates {@code maxLoad} during automatic selection (see
* {@code CompositePeerLauncher}'s own javadoc on {@code place}/{@code FixedPlacementPolicy}).
*
* <p>Round 1 resolved {@code acquireWithWorktree}'s profile through
* {@code launcher.routedProfileFor(memberRole)} and then fed that name back into
* {@code launcher.spawn(SpawnRequest)} as an EXPLICIT profile. Naming a profile explicitly
* takes {@code CompositePeerLauncher.spawn}'s THROWING branch, which calls
* {@code enforceMaxLoad} — so the worktree path died with a {@code PlacementException} while
* the exact same unqualified request, with no worktree, still spawned cleanly through the
* routing branch that never checks {@code maxLoad} at all. One intent, two different answers,
* depending only on whether a worktree was asked for — the #425 shape, moved to a different
* filter instead of closed.
*
* <p>This test does not hardcode which of the two outcomes is correct — whether an unqualified
* spawn SHOULD respect {@code maxLoad} is fleetd #435, a separate ticket. It only asserts that
* the WITH-worktree and WITHOUT-worktree paths agree: both spawn on the same profile, or both
* fail with the same exception type and message. That way this test stays correct however
* #435 is eventually resolved, and only breaks if the two paths disagree again.
*/
@Test
void unqualifiedAcquireAgreesWithAndWithoutAWorktreeWhenTheOnlyProfileIsAtMaxLoad() {
Map<String, FleetConfig.Profile> profiles = new LinkedHashMap<>();
profiles.put("a", new FleetConfig.Profile("a", "http://gx00.gw:8000", "coder-a", null,
"FLEETD_WORKER_TOKEN", List.of("claude"), "tab", "fleetd-workers",
"worker: {profile} #{n}", null, "/repo/a", List.of("a.mcp.json"),
null, null, null, null, 1.0f, 1));
FakeHerdr herdr = new FakeHerdr();
ClaudeCodeLauncher adapter = new ClaudeCodeLauncher(new AgentControl(herdr), new WorkspaceControl(herdr),
new SubscriptionGuard(Set.of("gx00.gw")), profiles, "a", _ -> null);
// liveCount pinned at 1 for "a", exactly matching maxLoad — "a" is permanently at cap,
// regardless of how many times either branch below actually spawns.
PeerLauncher launcher = new CompositePeerLauncher(List.of(adapter), "a", profiles,
PlacementPolicies.fixed(), name -> "a".equals(name) ? 1 : 0);
Object without = attemptAcquire(() ->
new SessionManager(launcher, new FakeWorktrees(), () -> 0L)
.acquire(null, null, "/caller", null));
Object with = attemptAcquire(() ->
new SessionManager(launcher, new FakeWorktrees(), () -> 0L)
.acquire(null, null, "/caller", null, new WorktreeRequest("fleetd-425-maxload", null)));
assertEquals(without, with, "an unqualified spawn on a profile at maxLoad must agree "
+ "whether or not a worktree was requested — no condition may become newly fatal "
+ "on the worktree path alone (fleetd #425 rework, round 2)");
}
/**
* fleetd #425 rework, round 4: the exact regression a mutation test found that 186 green tests
* missed — {@code acquireWithWorktree} dropping the {@link
* dev.ltms.fleet.placement.PlacementDecision} it already resolved via {@code launcher.place},
* and letting the unqualified spawn re-run placement a second time (a blank-profile {@code
* launcher.spawn(spawnReq)}) instead of carrying that decision forward via {@code
* launcher.spawn(spawnReq, decision)}. Every earlier test in this file uses {@code
* PlacementPolicies.fixed()}, which returns the same answer on every {@code select()} call, so
* dropping the decision is invisible under it — two {@code select()} calls simply agree by
* accident. {@code PlacementPolicies.roundRobin()} is deterministic AND stateful: its {@code
* select()} advances an internal index on every call, so two consecutive calls for the SAME
* spawn (one from {@code place()} to provision the worktree, a second from a dropped-decision
* blank-profile {@code spawn(spawnReq)}) land on DIFFERENT profiles from a two-profile pool —
* index 0 ("a"), then index 1 ("b").
*
* <p>This test does not hardcode which profile wins — asserting one specific name would pass
* for the wrong reason the moment the rotation order changes (round-4 brief invariant 3). It
* asserts AGREEMENT instead: whichever profile the worktree's parity overlay was provisioned
* for must be the SAME profile the member actually spawned on. Each profile's overlay list is
* named after the profile itself ({@code "a.mcp.json"}/{@code "b.mcp.json"}), so comparing the
* recorded overlay against {@code s.profile() + ".mcp.json"} checks agreement without ever
* naming an expected winner.
*/
@Test
void acquireWithWorktreeSpawnsOnTheSameProfileItProvisionedTheWorktreeForUnderARotatingPolicy() {
Map<String, FleetConfig.Profile> profiles = new LinkedHashMap<>();
profiles.put("a", new FleetConfig.Profile("a", "http://gx00.gw:8000", "coder-a", null,
"FLEETD_WORKER_TOKEN", List.of("claude"), "tab", "fleetd-workers",
"worker: {profile} #{n}", null, "/repo/a", List.of("a.mcp.json")));
profiles.put("b", new FleetConfig.Profile("b", "http://gx00.gw:8000", "coder-b", null,
"FLEETD_WORKER_TOKEN", List.of("claude"), "tab", "fleetd-workers",
"worker: {profile} #{n}", null, "/repo/b", List.of("b.mcp.json")));
FakeHerdr herdr = new FakeHerdr();
ClaudeCodeLauncher adapter = new ClaudeCodeLauncher(new AgentControl(herdr), new WorkspaceControl(herdr),
new SubscriptionGuard(Set.of("gx00.gw")), profiles, "a", _ -> null);
// roundRobin is deterministic AND stateful: the first select() call picks index 0 ("a"),
// and the SAME policy instance's second select() call (reached only if the
// PlacementDecision is dropped) picks index 1 ("b") — the two-call disagreement this test
// needs to make a dropped decision observable, rather than merely probable.
PeerLauncher launcher = new CompositePeerLauncher(List.of(adapter), "a", profiles,
PlacementPolicies.roundRobin(), _ -> 0);
FakeWorktrees worktrees = new FakeWorktrees();
SessionManager sessions = new SessionManager(launcher, worktrees, () -> 0L);
MemberSession s = sessions.acquire(null, null, "/caller",
null, new WorktreeRequest("fleetd-425-round4", null));
FakeWorktrees.OverlayCall overlay = worktrees.lastOverlay();
assertNotNull(overlay, "overlayParity must have been called");
assertEquals(List.of(s.profile() + ".mcp.json"), overlay.requested(),
"the worktree must be provisioned for the SAME profile the member actually spawned "
+ "on — under a rotating policy, dropping the PlacementDecision makes the "
+ "second, spawn-time select() call disagree with the first, place()-time "
+ "call, so the member ends up on a profile whose worktree (repoRoot/parity "
+ "overlay) was built for a DIFFERENT profile (fleetd #425 rework, round 4)");
}
/**
* Reduce one {@code acquire(...)} attempt to a value comparable across the with-worktree and
* without-worktree paths: the spawned profile name on success, or the thrown exception's class
* and message on failure. Comparing THIS — instead of asserting "both spawn" or "both throw" as
* a hardcoded direction — is what keeps {@link
* #unqualifiedAcquireAgreesWithAndWithoutAWorktreeWhenTheOnlyProfileIsAtMaxLoad} valid whichever
* way fleetd #435 eventually resolves whether an unqualified spawn should respect maxLoad.
*/
private static Object attemptAcquire(Supplier<MemberSession> call) {
try {
return "spawned:" + call.get().profile();
} catch (RuntimeException e) {
return "threw:" + e.getClass().getName() + ":" + e.getMessage();
}
}
}