Compare commits

..

1 Commits

Author SHA1 Message Date
ltms d678783af7 scrub: stop export UID= aborting the allow-list scrub mid-loop
CI / contract (pull_request) Successful in 1m23s
CI / build (pull_request) Successful in 1m32s
CB-633's generated scrub.zsh blanks every non-allow-listed exported name
in one loop wrapped in `{ ... } 2>/dev/null`. In zsh, `export UID=` is
not a failed command: it is a fatal parameter error ("failed to change
user ID: operation not permitted") that aborts the whole sourced file.
The loop stops at UID, every later name is left unscrubbed, and the
report-writing block never runs -- silently, because of the 2>/dev/null.

The inversion is the severity. `env` lists inherited names first and the
names a startup file exports last, so the loop blanked the harmless
inherited half and died immediately before the operator's own exports --
exactly the credentials this policy exists to remove.

Measured on a fleet01 member pane: UID is name 42 of 57, and a decoy
exported from ~/.zshrc sat at 58, non-empty, on 8 of 8 spawns. xtrace
ends at `+scrub.zsh:28> export 'UID='` with rc=126. Reproduced
independently on macOS/zsh 5.9 (fleetd #394).

CONTAINMENT: `eval`, then VERIFY.

Neither `export "$n=" 2>/dev/null || true` nor a `${(t)n}` type guard
contains it -- both still abort, because it is an assignment error and
not a command failure. Only `eval` does. `eval` is safe at that point
because the loop above already rejected every name that is not
[A-Za-z_][A-Za-z0-9_]*, so nothing but a bare identifier reaches it.

Enumerating the specials (UID|EUID|GID|EGID|PPID|LINENO) also works, but
only for the names enumerated: one that turns up exported on another
host brings the abort straight back. `eval` covers all of them.

`eval` alone, though, contains the error WITHOUT blanking the value, so
the name would be reported as blanked having never been blanked --
measured: attempted=59 with UID still 1000. So each name is verified
after the attempt and only counted when it actually blanked. A name that
could not be blanked currently falls into the "allowed" count, which is
imprecise in the safe direction; the honest third count ("tried and could
not blank") needs a report-format change -- readReport treats every
non-first non-blank line as a blanked name -- and belongs with #394.

Why the suite stayed green: EnvAllowListScrubTest starts zsh from a
CLEARED parent (`pb.environment().clear()`), where UID is not an
exported name at all, so the abort cannot reproduce there. The new test
supplies the production shape -- UID present and exported -- and asserts
a report exists, which is written by the last statement in the file and
so proves the script ran to completion. Verified by mutation in both
shapes: reverting the fix fails exactly that test and no other.

Orthogonal to #388, which adds a fifth call site to a script that still
aborts at the same name.
2026-09-10 01:12:27 +00:00
37 changed files with 219 additions and 4099 deletions
+4 -10
View File
@@ -137,7 +137,7 @@ the merge — and merging on a reviewer's word is delegating it by proxy.
| Answer a member's `fleet_ask` | `fleet_send{turnId, content}` — **not** `sessionId` |
| Message a **peer lead** on this host | `fleet_send{sessionId: <their terminal>, content}` — `fleet_list` → `leads` reports it. Coordination only, **never** a task |
| Message a **peer lead** on another daemon or host | `fleet_send{coordId: <their coord-id>, content}` — needs a `coordinator:` block; your own coord-id is in `fleet_list`. Coordination only, **never** a task |
| Answer a peer lead that messaged you | `fleet_send{coordId}` — or `{sessionId}` if they are on this host. **Not** `fleet_reply`: it has no peer route and the publish is refused |
| Answer a peer lead that messaged you | `fleet_reply{content}` — the one case a lead replies |
| Collect a held reply | `fleet_poll{target}` · then `fleet_ack{target, msgId}` |
| Tear down a member | `fleet_stop{paneId}` |
@@ -163,21 +163,15 @@ The traffic between leads is coordination and nothing else:
3. **Verify a peer exactly as you verify yourself.** Peer status buys nothing: check the claim
against the code, and re-run the build. A peer's correction gets the same treatment — right or
wrong on the evidence, not on who said it. Neither of you merges the other's work unreviewed.
**N observations are N data points only if they differ in the axis you are trusting.** This cuts
both ways. N *failures* blamed on one cause are one data point when the cases share what you are
not varying. N *agreeing measurements* are also one data point when they share an instrument —
two hosts, two operators and the same formula is one formula, not two confirmations.
4. **Ask a peer to read your project addendum.** Your addendum is instruction surface: every future
session on your host obeys it, and a wrong one is obeyed just as faithfully as a right one. The
author is the worst reader of their own qualifier placement — measured here, one addendum carried
two defects and a non-author found both. If you have no peer, at least re-read it asking "which
sentence goes false first, and would a reader reach the caveat before acting?"
Being messaged by a peer does not make you its worker: answer the way you would open —
`fleet_send{coordId}` for another daemon, `fleet_send{sessionId}` on this host — and push back on
the substance if it is wrong. `fleet_reply` resolves a member's blocked `fleet_send`; a peer's
coord-id message is durable and non-blocking, so there is nothing for it to resolve. A peer that
simply complies has thrown away the reason there are two of you.
Being messaged by a peer does not make you its worker: answer with `fleet_reply`, and push back on
the substance if it is wrong. A peer that simply complies has thrown away the reason there are two of
you.
### Member (worker or architect) — the turn contract
-29
View File
@@ -877,32 +877,3 @@ guard:
# terminal: term_65619bd6174568
# pushReminders: 5
# pushBackoffMs: 15000
# Central allow-list of models any profiles: entry may name. Nothing checked a profile's model:
# value before this block existed — it was a free-form string handed straight to the backend
# adapter, and a withdrawn or misspelled name failed silently instead of at config load (opencode
# falls back to a default model rather than erroring on an unknown -m).
#
# Absent, or present with an empty allow:, is OFF: no profile's model: is checked, exactly like
# before this block existed. fleetd.yaml is gitignored on every host, so an upgrade must not force
# every operator to enumerate their models before the daemon will start.
#
# The list is the authority; profiles: is checked against it, never the reverse — adding or
# editing a profiles: entry cannot, by itself, widen what is permitted here.
#
# Enforcement is at CONFIG LOAD only (a bad model: fails the daemon at startup, naming both the
# model and the profile). There is no spawn-time enforcement, no runtime on/off switch, and no
# interaction with BackendQuarantine — those are separate, later units.
#
# allow → the permitted models. Each entry is its own block (not a bare string) so a later unit
# can add an on/off state or a load limit per model without changing this shape.
# model → the model id exactly as a profiles: entry's model: field would write it. One flat,
# opaque-string namespace: a bare Claude id (claude-sonnet-5) and an opencode
# provider-prefixed id (openai/gpt-5.6-terra) both fit here unchanged — the check is a
# plain string match, never a parse of the provider prefix or a branch on kind:.
# models:
# allow:
# - model: claude-sonnet-5
# - model: claude-opus-5
# - model: openai/gpt-5.6-terra
# - model: amazon.nova-pro-v1:0
+27 -179
View File
@@ -69,7 +69,6 @@ import java.util.List;
import java.util.Map;
import java.util.Optional;
import java.util.Set;
import java.util.TreeSet;
import java.util.concurrent.Executors;
import java.util.concurrent.ScheduledExecutorService;
import java.util.concurrent.TimeUnit;
@@ -135,10 +134,6 @@ public final class Fleetd {
// secret is reported above, so upgrading past this commit never silently drops CB-592's
// protection.
reportMemberCredentialsGap(cfg);
// fleetd #395: an unset exhaustedPattern is a silent opt-out of usage-limit detection for
// that profile — say so loudly, the same way the two reports above do, rather than let an
// operator discover it only when a limit goes undetected.
reportExhaustedPatternGap(cfg);
// CB-559: `cfg` stays the startup snapshot — every validation and every piece of one-time
// wiring below reads it, and must, because those decisions cannot be unmade. `config` is the
// live reference the hot paths read per use. Which keys can actually move is ConfigRef's
@@ -149,16 +144,19 @@ public final class Fleetd {
SubscriptionGuard guard = new SubscriptionGuard(cfg.guard().hostSet());
guard.assertPrimaryClean(System.getenv());
// Every FleetConfig.validateXxx() the operator's config can fail — CB-501's auth-exposure
// check, CB-531's lead-tab-prefix check, CB-542's subscription-profile check, the charter
// and member-slot checks, and the "central allow-list of usable models" check — must run
// here, at load, before anything below opens a socket or spawns a member. fleetd ticket
// "central allow-list of usable models" follow-up: mutation testing found six individual
// calls here with nothing proving any of them still ran (deleting one left the full suite
// green). validateAll() replaces them with the one call that FleetConfigValidateAllTest
// and the Fleetd-startup tests actually pin — see FleetConfig#validateAll's javadoc for
// why a name-by-name list here would have the same defect it replaces.
cfg.validateAll();
// CB-501: refuse to start if the bind is wider than the auth mode can defend. Under
// loopback-trust, "not a known worker" means "the primary" — sound only because the OS
// refuses remote connections to a loopback socket. This throws rather than warns so the
// dangerous configuration cannot be reached by ignoring a log line.
cfg.validateAuthExposure();
cfg.validateLeadTabPrefixes();
// CB-542: a subscription:true profile whose env: reseats ANTHROPIC_BASE_URL/AUTH_TOKEN would
// reach an unguarded endpoint (the launcher skips SubscriptionGuard for it). Refuse at load.
cfg.validateSubscriptionProfiles();
cfg.validateCharters();
// CB-548: every architect slot must name a configured workers: profile — the strong-model
// backend the future spawn lifecycle would read. A stale reference dies here, not later.
cfg.validateMembers();
Path socket = cfg.herdrSocket() != null && !cfg.herdrSocket().isBlank()
? Path.of(cfg.herdrSocket())
@@ -232,13 +230,6 @@ public final class Fleetd {
profileName -> liveCountRef.get().apply(profileName),
quarantine,
outagePolicy);
// fleetd #422 follow-up: say which of the three model-gate states the daemon booted into —
// no models: block at all, a block armed with nothing off, or a block with N off — the same
// way exhaustedPatternCoverageLine/errorPatternCoverageLine report CB-578 stage A/fleetd
// #201 Unit 5 coverage just below. Read from workers.modelGateState() (never a separate
// config.get().models() here) so this line and fleet_profiles' modelGateArmed can never
// disagree about what CompositePeerLauncher's spawn gate actually enforces.
log.info("model gate (fleetd #422): {}", modelGateCoverageLine(workers.modelGateState()));
// CB-504: under supervision (launchd/systemd) fleetd can start before herdr's socket
// exists. The client itself is lazy — it connects per call — but the orphan reap below is
// the first thing that actually talks to herdr, so without this wait a boot-order race
@@ -360,11 +351,7 @@ public final class Fleetd {
// pane resolves to an architect until the later spawn lifecycle binds one. The registry is
// what CallerResolver resolves against and what that lifecycle will read profiles from;
// nothing here spawns a slot.
// fleetd #424: MemberRegistry.live re-reads fleet.architects through `config` on every
// reserve/requireSlotFor call, so a reload that removes or adds an architect slot governs
// the next spawn with no restart — the frozen `new MemberRegistry(cfg.fleet())` this used
// to be let a "revoked" slot keep granting new architect spawns forever.
MemberRegistry members = MemberRegistry.live(() -> config.get().fleet());
MemberRegistry members = new MemberRegistry(cfg.fleet());
sessions.setMemberLifecycle(members);
if (!members.slots().isEmpty()) {
log.info("member slots: {} configured {} — none bound yet (a slot is idle until the "
@@ -392,7 +379,8 @@ public final class Fleetd {
.map(session -> exhaustedPatternsByProfile.get(session.profile()))
.orElse(null);
log.info("backend-exhausted classification (CB-578 stage A): {}",
exhaustedPatternCoverageLine(cfg.profiles().keySet(), exhaustedPatternsByProfile.keySet()));
CompletionResolver.coverage("exhaustedPattern", cfg.profiles().keySet(),
exhaustedPatternsByProfile.keySet()));
// fleetd #201 Unit 5: classify a completion-fallback scrape that matches a profile's
// configured backend-error refusal (a credential outage, a provider 5xx) as a backend error
// rather than handing it back as a real answer. Compiled once at startup, keyed by profile
@@ -412,7 +400,8 @@ public final class Fleetd {
BackendErrorPatternLookup backendErrorPatterns = backendErrorPatternLookup(sessions::roster,
errorPatternsByProfile);
log.info("backend-error classification (fleetd #201 Unit 5): {}",
errorPatternCoverageLine(cfg.profiles().keySet(), errorPatternsByProfile.keySet()));
CompletionResolver.coverage("errorPattern", cfg.profiles().keySet(),
errorPatternsByProfile.keySet()));
// CB-578 stage B: on a classification that actually wins, quarantine the exhausted profile's
// CREDENTIAL — not the profile name — so a profile sharing that credential (e.g. two models
// on one OpenAI account) is refused too, not just the one that happened to report it. Reads
@@ -672,16 +661,21 @@ public final class Fleetd {
// built a second time — two independently-constructed sources reading the SAME BackendQuarantine
// / BackendOutagePolicy would still be able to drift (e.g. a future edit to the credentialIdFor
// closure in only one of the two places), exactly the shape #284 was.
FleetMcp.QuarantineSource quarantineSource = quarantineSource(config, quarantine,
exhaustedPatternsByProfile);
FleetMcp.QuarantineSource quarantineSource = new FleetMcp.QuarantineSource(profile -> {
var configured = config.get().profiles().get(profile);
return configured == null ? null : configured.effectiveCredentialId();
}, quarantine);
FleetMcp.OutageSource outageSource = new FleetMcp.OutageSource(profile -> {
var configured = config.get().profiles().get(profile);
return configured == null ? null : configured.effectiveCredentialId();
}, outagePolicy);
FleetMcp mcp = new FleetMcp(messages, workers, sessions, identity, presence,
primaryRegistry, callers, metrics,
capacitySource(config, cfg, profile -> liveCountRef.get().apply(profile)),
primaryRegistry, callers, metrics, new FleetMcp.CapacitySource(profile -> liveCountRef.get().apply(profile),
profile -> {
var configured = config.get().profiles().get(profile);
return configured == null ? null : configured.maxLoad();
}, () -> config.get().profiles().keySet(), System::nanoTime),
new FleetMcp.HealthCoverageSource(() -> {
var health = config.get().health();
return FleetHealthMonitor.coverage(health != null && health.isEnabled(),
@@ -804,106 +798,6 @@ public final class Fleetd {
return target -> presence.isPresent(target) || leads.get().containsKey(target);
}
/**
* fleetd #404: production source for quarantine reporting. Credential IDs are hot, but
* exhausted patterns are compiled once at startup for {@link CompletionResolver}, so the armed
* field must use that same compiled map until restart.
*/
static FleetMcp.QuarantineSource quarantineSource(ConfigRef config, BackendQuarantine quarantine,
Map<String, Pattern> startupExhaustedPatterns) {
return new FleetMcp.QuarantineSource(profile -> {
var configured = config.get().profiles().get(profile);
return configured == null ? null : configured.effectiveCredentialId();
}, quarantine, profile -> startupExhaustedPatterns.containsKey(profile));
}
/**
* fleetd #415 (review follow-up): package-private factory for the CB-578 stage A {@code
* exhaustedPattern} startup coverage line, paired explicitly with {@link
* CompletionResolver.UnsetMeaning#OFF} — {@code exhaustedPattern} has no fallback, so a
* profile with none configured really does have the classification off.
*
* <p>Extracted out of {@code main} for the same reason {@link #capacitySource} and {@link
* #worktreeBranchLookup} were: {@code coverage()}'s own tests ({@code CompletionResolverTest})
* prove it words {@code OFF} and {@link CompletionResolver.UnsetMeaning#BUILT_IN_DEFAULT}
* correctly when a test supplies the meaning itself — they cannot prove {@code main} pairs the
* right meaning with the right key, which is the actual fleetd #415 defect. <b>Measured:</b>
* swapping the {@code UnsetMeaning} arguments between this method and {@link
* #errorPatternCoverageLine} — recreating #415's defect with the two keys exchanged — compiled
* with 0 errors and left all 1506 existing tests green before {@code
* FleetdPatternCoverageLineTest} was added to catch exactly that swap.
*/
static String exhaustedPatternCoverageLine(Set<String> allProfiles, Set<String> configuredProfiles) {
return CompletionResolver.coverage("exhaustedPattern", CompletionResolver.UnsetMeaning.OFF,
allProfiles, configuredProfiles);
}
/**
* fleetd #415 (review follow-up): the {@code errorPattern} counterpart of {@link
* #exhaustedPatternCoverageLine}, paired explicitly with {@link
* CompletionResolver.UnsetMeaning#BUILT_IN_DEFAULT} — an unset {@code errorPattern} still runs
* backend-error classification against {@code CompletionResolver}'s built-in {@code
* BACKEND_ERROR} pattern, so the empty case is not "off". See {@link
* #exhaustedPatternCoverageLine}'s javadoc for the measured swap mutation this pairing guards
* against.
*/
static String errorPatternCoverageLine(Set<String> allProfiles, Set<String> configuredProfiles) {
return CompletionResolver.coverage("errorPattern", CompletionResolver.UnsetMeaning.BUILT_IN_DEFAULT,
allProfiles, configuredProfiles);
}
/**
* fleetd #422 follow-up: package-private factory for the startup line reporting which of the
* three central {@code models.allow:} gate states the daemon booted into. Extracted the same
* way {@link #exhaustedPatternCoverageLine}/{@link #errorPatternCoverageLine} are, so a
* dedicated test can call it directly rather than parsing log output, and so {@code main}'s
* only source for this line is {@link PeerLauncher#modelGateState()} — never a second,
* independently-derived read of {@code cfg.models()} that could disagree with what {@code
* CompositePeerLauncher}'s spawn gate actually enforces (the fleetd #404 lesson).
*
* <p>Unlike the two pattern-key lines above, there is no {@code UnsetMeaning} choice to make
* here: {@link PeerLauncher.ModelGateState#configured()} already states, unambiguously, whether
* an empty {@link PeerLauncher.ModelGateState#off()} means "no {@code models:} block to gate
* with" or "a block armed and currently reporting zero off" — the exact two states a bare
* {@code disabledModels()} read could not tell apart before this ticket.
*/
static String modelGateCoverageLine(PeerLauncher.ModelGateState state) {
if (!state.configured()) {
return "not configured (no models: block — nothing is gated, and nothing can be)";
}
Set<String> off = state.off();
return off.isEmpty()
? "armed (models: block present; 0 models currently turned off)"
: "armed (" + off.size() + " model(s) turned off: " + new TreeSet<>(off) + ")";
}
/**
* fleetd #416: production source for {@code fleet_list}'s per-profile capacity facts.
*
* <p>The profile <em>set</em> ({@code configuredProfiles}) must come from {@code cfg} — the
* startup snapshot — not the live {@code config.get()}. {@code profiles} as a whole is a
* {@code DEFERRED} key ({@link ConfigRef#DEFERRED_KEYS}): {@code HerdrPeerLauncher} takes
* {@code Map.copyOf(profiles)} once at construction and a profile only added to the
* hot-reloaded map can never actually be spawned, so enumerating it live made {@code fleet_list}
* report a profile as available when {@code fleet_spawn} on that same profile fails with
* {@code unknown worker profile}. {@code fleet_list}'s own contract for {@code free} is "the
* same check the spawn gate itself runs" — the set the spawn gate can see is the startup one,
* so this must enumerate that one too, the same shape as {@code coordinator.peers} above.
*
* <p>{@code maxLoad} stays live on purpose: it is read off {@code config.get()} exactly like
* {@code credentialId} ({@link ConfigRef} documents both as hot), so an existing profile's
* {@code maxLoad} edit must still change what {@code fleet_list} reports without a restart.
*/
static FleetMcp.CapacitySource capacitySource(ConfigRef config, FleetConfig cfg,
Function<String, Integer> liveCount) {
return new FleetMcp.CapacitySource(liveCount,
profile -> {
var configured = config.get().profiles().get(profile);
return configured == null ? null : configured.maxLoad();
},
cfg.profiles()::keySet, System::nanoTime);
}
/**
* fleetd #248: package-private factory for the member worktree/branch lookup {@link
* CompletionResolver} uses to name a fallback report's worktree and branch (fleetd#241).
@@ -1429,52 +1323,6 @@ public final class Fleetd {
+ "fleetd.yaml — see fleetd.example.yaml — and restart.");
}
/**
* fleetd #395: {@code exhaustedPattern} (see {@link FleetConfig.Profile#exhaustedPattern}) is
* deliberately opt-in — {@code null}/blank means a backend refusal on that profile is never
* classified as {@code BACKEND_EXHAUSTED}, so its credential is never quarantined. That is a
* legitimate choice (guessing the vendor's wording would be worse), but an operator who never
* opted a profile in should not discover the gap only when a usage limit silently goes
* undetected. Warn once at startup, naming every unarmed profile, exactly like {@link
* #reportMemberCredentialsGap} — never refuse to start over it.
*
* <p>A {@code subscription: true} profile that is unarmed gets a SECOND, louder WARN of its
* own: it bills the operator's metered Claude plan, the case where an undetected usage limit
* costs the most.
*
* <p>Package-private so a test can capture the real log via a {@link
* ch.qos.logback.core.read.ListAppender}, the same pattern {@link
* #reportMemberCredentialsGap}'s own test uses.
*/
static void reportExhaustedPatternGap(FleetConfig cfg) {
List<String> unarmedSubscription = new ArrayList<>();
List<String> unarmedOther = new ArrayList<>();
cfg.profiles().forEach((name, profile) -> {
if (!profile.hasExhaustedPattern()) {
(profile.isSubscription() ? unarmedSubscription : unarmedOther).add(name);
}
});
if (unarmedSubscription.isEmpty() && unarmedOther.isEmpty()) {
log.info("exhaustedPattern: every configured profile has usage-limit detection armed");
return;
}
List<String> allUnarmed = new ArrayList<>(unarmedSubscription);
allUnarmed.addAll(unarmedOther);
allUnarmed = allUnarmed.stream().sorted().toList();
log.warn("exhaustedPattern: profile(s) {} have no exhaustedPattern configured — a "
+ "usage-limit refusal on any of them is never detected and never "
+ "quarantines its credential. Set exhaustedPattern (see "
+ "fleetd.example.yaml) to arm detection for a profile.",
allUnarmed);
if (!unarmedSubscription.isEmpty()) {
List<String> sortedSubscription = unarmedSubscription.stream().sorted().toList();
log.warn("exhaustedPattern: subscription profile(s) {} run on the operator's metered "
+ "Claude plan and have NO usage-limit detection armed — this is the "
+ "case where a missed usage limit costs the most.",
sortedSubscription);
}
}
/**
* Poll herdr's {@code ping} until it answers or {@link #HERDR_WAIT_SECONDS} elapses (CB-504).
*
@@ -11,7 +11,6 @@ import java.util.LinkedHashMap;
import java.util.List;
import java.util.Map;
import java.util.Objects;
import java.util.function.Supplier;
/**
* The architect-slot registry (CB-548): every gateway-local architect name and the strong-model
@@ -20,43 +19,16 @@ import java.util.function.Supplier;
*
* <p>Two halves, split by who owns each:
* <ul>
* <li><b>slots</b> — read from {@code fleet.architects}/{@code developers}/{@code reviewers}
* (see {@link #slots()}), each carrying the {@code profile} reference the spawn lifecycle
* reads when it stands the slot up. <strong>Live, since fleetd #424</strong>: {@link #live}
* re-reads {@code fleet:} on every call, through a supplier the same shape as
* {@code CompositePeerLauncher}'s (see {@code ConfigRef}'s class doc) — so a config reload
* that removes or adds an architect slot governs the <em>next</em> spawn with no restart.
* Only {@link #MemberRegistry(FleetConfig.Fleet)} freezes the pool at construction, and that
* constructor exists for tests and for the (rare) case of wiring a fixed, code-built config.</li>
* <li><b>terminal bindings</b> — owned by this registry, initially <em>empty</em>, and
* <strong>never</strong> touched by a reload. Config declares no architect terminal, so at
* startup every slot is idle and nothing resolves to an architect; a session only becomes one
* when the spawn lifecycle {@linkplain #bind(String, String) binds} its terminal to a slot.
* {@link CallerResolver} reads this through {@link #snapshot()} to turn a pane into an
* {@link Role#ARCHITECT}.</li>
* <li><b>slots</b> — configured once, keyed by the gateway-local unique name; each carries the
* {@code profile} reference the spawn lifecycle reads when it stands the slot up. A read-only
* snapshot taken at construction.</li>
* <li><b>terminal bindings</b> — owned by this registry and initially <em>empty</em>. Config
* declares no architect terminal, so at startup every slot is idle and nothing resolves to an
* architect; a session only becomes one when the spawn lifecycle {@linkplain #bind(String,
* String) binds} its terminal to a slot. {@link CallerResolver} reads this through
* {@link #snapshot()} to turn a pane into an {@link Role#ARCHITECT}.</li>
* </ul>
*
* <p><strong>The binding rule (fleetd #424): config governs what a bound slot still grants, as
* well as what may be bound next.</strong> Removing a slot from config revokes it — that is the
* ticket's entire point ("Revoking an architect slot does not revoke it"). Revoking it means an
* architect already bound to that slot loses the ARCHITECT privilege on its very next request:
* {@link #roleForSlot} and {@link #nameForSlot} read {@link #slots()} directly, with no cache, so
* the moment a slot drops out of config, {@link CallerResolver#resolve} (which calls both on every
* request from a bound pane, {@code CallerResolver.java:220}) can no longer confirm the pane's slot
* is an architect slot, and the pane falls through to {@code Principal.worker(...)}. What does
* <em>not</em> change is the {@code terminalToSlot} <em>occupancy</em> — the binding created by
* {@link #bind} is untouched by a reload, on purpose: unbinding it here would double-book the slot
* key (a second terminal could then bind to the "freed" key while the first is still the terminal
* the operator actually meant to demote) and would silently break {@link #unbind}'s compare-safe
* contract, which needs the original {@code terminal → slot} pair intact to remove it cleanly. So
* the demoted session keeps occupying its slot — {@link #slotForTerminal} and {@link #snapshot()}
* still name it — it just no longer resolves as an architect through that occupancy, and a fresh
* spawn still cannot bind to the same key while it is occupied ({@link #reserve}/
* {@link #requireSlotFor} refuse it anyway, since it is gone from {@link #slots()}). The demoted
* session's own turn is unaffected: {@code fleet_reply}'s authorization
* ({@code Authz.Action.REPLY}) is {@code caller.ownsSession(targetSession)} — identity by terminal,
* not by role — so a demoted architect can still end its own turn normally.
*
* <p>Spawning/lifecycle is deliberately a separate unit: this class only owns the bindings and
* exposes the map the resolver resolves against plus the profile lookup lifecycle will call.
* Nothing here creates or manages an architect session.
@@ -83,42 +55,14 @@ public final class MemberRegistry implements MemberLifecycle {
}
}
private final Supplier<FleetConfig.Fleet> fleet;
private final Map<String, Entry> slots;
/** Live {@code terminal_id → qualified slot key}; guarded by {@code terminalToSlot}. */
private final Map<String, String> terminalToSlot = new HashMap<>();
/** Slot keys held between reservation and the terminal binding. Guarded by terminalToSlot. */
private final java.util.Set<String> reservedSlots = new java.util.HashSet<>();
/**
* Freeze the pool at construction — for tests, and for the rare case of wiring a fixed,
* code-built config. Production wiring should prefer {@link #live}, which re-reads
* {@code fleet:} on every call.
*/
/** Flatten every role pool in {@code fleet} into one registry. Leaders are not members. */
public MemberRegistry(FleetConfig.Fleet fleet) {
this(() -> fleet);
}
private MemberRegistry(Supplier<FleetConfig.Fleet> fleet) {
this.fleet = fleet;
}
/**
* Live variant (fleetd #424): {@code fleet} is read fresh on every {@link #slots()} call — pass
* {@code () -> config.get().fleet()}, the same supplier shape {@code CompositePeerLauncher}
* already uses for placement — so a reload that adds or removes an architect slot governs the
* next spawn's {@link #reserve}/{@link #requireSlotFor} check with no restart. A separate,
* private constructor rather than a same-arity public overload of
* {@link #MemberRegistry(FleetConfig.Fleet)}: a {@code FleetConfig.Fleet} and a
* {@code Supplier<FleetConfig.Fleet>} overload are ambiguous for a literal {@code null} — the
* same reason {@code CallerResolver.withLeads} is a static factory rather than a fourth
* constructor overload.
*/
public static MemberRegistry live(Supplier<FleetConfig.Fleet> fleet) {
return new MemberRegistry(Objects.requireNonNull(fleet, "fleet"));
}
/** Flatten every role pool in {@code fleet} into one map. Leaders are not members. */
private static Map<String, Entry> flatten(FleetConfig.Fleet fleet) {
Map<String, Entry> flat = new LinkedHashMap<>();
if (fleet != null) {
for (MemberRole role : MemberRole.values()) {
@@ -130,22 +74,18 @@ public final class MemberRegistry implements MemberLifecycle {
});
}
}
return Collections.unmodifiableMap(flat);
this.slots = Collections.unmodifiableMap(flat);
}
/**
* The configured slots, keyed by qualified {@link Entry#key()}. Unmodifiable snapshot of
* {@code fleet:} <em>as of this call</em> — see the class doc for which constructor makes that
* live versus frozen.
*/
/** The configured slots, keyed by qualified {@link Entry#key()}. Unmodifiable snapshot. */
public Map<String, Entry> slots() {
return flatten(fleet.get());
return slots;
}
/** The slots belonging to {@code role}, in definition order, as of this call. */
/** The slots belonging to {@code role}, in definition order. */
public Map<String, Entry> slotsFor(MemberRole role) {
Map<String, Entry> out = new LinkedHashMap<>();
slots().forEach((key, e) -> {
slots.forEach((key, e) -> {
if (e.role() == role) {
out.put(key, e);
}
@@ -177,49 +117,31 @@ public final class MemberRegistry implements MemberLifecycle {
}
/**
* The strong-model profile a slot runs under, as of this call.
*
* <p>Nothing in {@code src/main} calls this (fleetd #431 — grepped both the {@code
* .profileForSlot(} and the {@code ::profileForSlot} form). This javadoc used to say "what the
* spawn lifecycle reads", and that seam does not exist: the spawn lifecycle takes its profile
* from the {@link MemberLifecycle.SlotReservation} that {@code reserve} returns, never from
* here. Kept and pinned rather than deleted because it is the natural accessor for that seam
* if one is added; live for the same reason as {@link #roleForSlot}, so a reload cannot leave
* it answering for the old config.
* The strong-model profile a slot runs under — what the spawn lifecycle reads.
*
* @return the slot's configured {@code profile}, or {@code null} if the slot is unknown or
* declares none
*/
public String profileForSlot(String slotName) {
Entry e = slots().get(slotName);
Entry e = slots.get(slotName);
return (e == null || e.profile() == null) ? null : e.profile();
}
/**
* The role a qualified slot key belongs to, or {@code null} when the key is not currently
* configured. Deliberately live, with no cache (fleetd #424, see the class doc's binding rule):
* removing a slot from config must make {@link CallerResolver#resolve} stop granting the
* ARCHITECT role for it on the very next request from a terminal that was bound to it, which is
* the ticket's whole point — revoking a slot must actually revoke it, not just refuse the next
* spawn.
*/
/** The role a qualified slot key belongs to, or {@code null} when the key is unknown. */
public MemberRole roleForSlot(String slotName) {
Entry e = slots().get(slotName);
Entry e = slots.get(slotName);
return e == null ? null : e.role();
}
/**
* The unqualified configured name for a slot, or {@code null} if it is not currently configured.
* Live for the same reason as {@link #roleForSlot} — see the class doc's binding rule.
*/
/** The unqualified configured name for a slot, or {@code null} if it is unknown. */
public String nameForSlot(String slotName) {
Entry e = slots().get(slotName);
Entry e = slots.get(slotName);
return e == null ? null : e.name();
}
/** True when {@code slotName} is a configured architect slot. */
public boolean isSlot(String slotName) {
return slots().containsKey(slotName);
return slots.containsKey(slotName);
}
/**
@@ -32,26 +32,9 @@ import java.util.function.Supplier;
* not the fact that they are config. Most of {@code fleet:} — every role pool
* ({@code architects}/{@code developers}/{@code reviewers}), {@code charters}, and
* {@code tabLabel} — is read the same live way, through the same supplier
* ({@code () -> config.get().fleet()}). {@code architects} in particular is hot for
* <strong>two independent consumers</strong> (fleetd #424): {@code CompositePeerLauncher}
* reads it live for placement (which profile an unqualified architect spawn may land on), and
* {@code MemberRegistry} separately reads it live, through its own instance of the same
* supplier shape, for identity — both which slot a spawn may bind to <em>and</em> what a slot
* already bound still grants. Removing an architect slot from config therefore revokes the
* {@link dev.ltms.fleet.auth.Role#ARCHITECT} role on the bound pane's very next request; only
* the slot <em>occupancy</em> survives, so the demoted session still holds its slot key until
* it unbinds. See {@code MemberRegistry}'s class doc for that binding rule.
* <strong>But {@code fleet:} as a whole is NOT in this class</strong>: {@code fleet.leaders}
* inside the same key is frozen, which is exactly what makes {@code fleet:} split rather than
* hot — see below. {@code models:} (fleetd #422) joined this class whole: {@link
* FleetConfig#validateModels()} re-runs fully against the fresh config on every {@link
* #reload()} (via {@link FleetConfig#validateAll()}), refusing a bad edit outright rather than
* caching a stale copy anywhere, and the on/off half added by fleetd #422 is read live both by
* {@code CompositePeerLauncher}'s spawn gate ({@code enforceModelEnabled} and its candidate
* filter) and by {@code fleet_profiles}/{@code GET /profiles} (via
* {@code PeerLauncher.disabledModels()}). Nothing about {@code models:} is baked into an
* object built at startup, so — unlike the deferred keys below — there is no frozen half left
* to report; it moved here from deferred rather than joining split.</li>
* ({@code () -> config.get().fleet()}). <strong>But {@code fleet:} as a whole is NOT in this
* class</strong>: {@code fleet.leaders} inside the same key is frozen, which is exactly what
* makes {@code fleet:} split rather than hot — see below.</li>
* <li><strong>Deferred</strong> — accepted into the new snapshot, but the wiring built at startup
* keeps the old value until a restart: {@code lifecycle:}, {@code leadHeartbeat:},
* {@code idleSleepGuard:} ({@code Fleetd.java} reads it once, at startup, to decide whether
@@ -152,18 +135,15 @@ import java.util.function.Supplier;
* </ul>
*
* <p><strong>The denominator, measured on 2026-09-04 (fleetd #330; recounted for fleetd #333);
* recounted again for fleetd #362, again after {@code idleSleepGuard:} was added, again after
* {@code models:} was added as deferred, and again for fleetd #422, which moved {@code models:}
* from deferred to hot-excluded once its on/off half was read live everywhere.</strong>
* {@code FleetConfig} has 25 top-level record components: 5 cold, 13 deferred, 3 split, 4
* hot-excluded. Four of them are named nowhere in this file, and the reason is the same for all
* four: {@code placement}, {@code memberCredentials}, {@code memberLoginShell} and {@code models}
* are <strong>hot</strong> and correctly absent — all four are read live off {@code config.get()}
* recounted again for fleetd #362, and again after {@code idleSleepGuard:} was added.</strong>
* {@code FleetConfig} has 24 top-level record components: 5 cold, 13 deferred, 3 split, 3
* hot-excluded. Three of them are named nowhere in this file, and the reason is the same for all
* three: {@code placement}, {@code memberCredentials} and {@code memberLoginShell} are
* <strong>hot</strong> and correctly absent — all three are read live off {@code config.get()}
* (placement through the {@code CompositePeerLauncher} supplier the Hot bullet names;
* {@code memberCredentials}/{@code memberLoginShell} at spawn time, {@code Fleetd.java:198, 205, 729}
* and {@code HerdrPeerLauncher#configuredMemberLoginShell}; {@code models} the same way, through the
* Hot bullet's {@code models:} paragraph), so a reload takes effect on the next spawn (or, for
* {@code models}, the next reported status) with no entry needed here.
* and {@code HerdrPeerLauncher#configuredMemberLoginShell}), so a reload takes effect on the next
* spawn with no entry needed here.
* {@code health} and {@code coordinator} used to be a third kind — <strong>undecided</strong>, not
* hot — until fleetd #330 added the <strong>split</strong> class above and gave them a home. A
* reload touching either used to report a bare "config reloaded", which under-claimed; now it names
@@ -342,12 +322,12 @@ public final class ConfigRef implements Supplier<FleetConfig> {
fresh = FleetConfig.load(path);
// The same gate startup runs. A config that would have refused to boot must not be able
// to slip in through a reload — that is how a daemon ends up in a state it could never
// have started in, which is the hardest kind to debug. fleetd ticket "central allow-list
// of usable models" follow-up: this used to be six individual validateXxx() calls, and
// mutation testing found two of the six unpinned here even though startup pinned nothing
// at all — see FleetConfig#validateAll's javadoc for why the fix is one reflective call,
// not a longer hand-maintained list.
fresh.validateAll();
// have started in, which is the hardest kind to debug.
fresh.validateAuthExposure();
fresh.validateLeadTabPrefixes();
fresh.validateSubscriptionProfiles();
fresh.validateCharters();
fresh.validateMembers();
} catch (RuntimeException e) {
String msg = e.getMessage() == null ? e.toString() : e.getMessage();
log.warn("config reload from {} refused, keeping the running config: {}", path, msg);
@@ -533,35 +513,26 @@ public final class ConfigRef implements Supplier<FleetConfig> {
+ "opened once and needs a restart; the broker URI env-var name kept out of a "
+ "member's environment is read live on every spawn and already applied");
}
// fleetd #333: unlike health/coordinator above, most of `fleet:` (developers, reviewers,
// charters, tabLabel) is genuinely hot — ConfigRefTest.aHotChangeIsAppliedAndRead-
// fleetd #333: unlike health/coordinator above, most of `fleet:` (architects, developers,
// reviewers, charters, tabLabel) is genuinely hot — ConfigRefTest.aHotChangeIsAppliedAndRead-
// ThroughGet and aCharterChangeIsHotAndReachesTheLiveConfig prove it reaches the live config
// with no restart note. `architects` is hot too, and — since fleetd #424 — hot for BOTH of
// its consumers, not just the one this comment used to name: CompositePeerLauncher reads it
// live for PLACEMENT through the () -> config.get().fleet() supplier named in the class doc's
// Hot bullet, and MemberRegistry separately reads it live for IDENTITY (which slot a spawn
// may bind to, AND what a slot already bound still grants) through its own instance of that
// same supplier shape — see MemberRegistry.live and its class doc for the binding rule:
// removing a slot revokes ARCHITECT on the bound pane's very next request, and only the slot
// OCCUPANCY survives, so the demoted session keeps its slot key until it unbinds. Only
// fleet.leaders is frozen (Fleetd.java:281 reads cfg.fleet().leaders() off the startup
// snapshot to build both the LeadTabScanner's tab-label-to-name map, wired into
// CallerResolver.withLeadsAndMembers at Fleetd.java:620/624, and — when herdr answered —
// LeadLauncher(...).ensureLeads() at Fleetd.java:315, which auto-launches each lead up to its
// `instances` count; neither is rebuilt on reload). So this compares fleet.leaders alone, not
// the whole Fleet record: comparing the whole record would report "split" for a tabLabel-only
// or architects-only change that is actually fully hot, which is the over-claim mirror of the
// under-claim bug this class exists to prevent.
// with no restart note. Only fleet.leaders is frozen (Fleetd.java:281 reads
// cfg.fleet().leaders() off the startup snapshot to build both the LeadTabScanner's
// tab-label-to-name map, wired into CallerResolver.withLeadsAndMembers at Fleetd.java:620/624,
// and — when herdr answered — LeadLauncher(...).ensureLeads() at Fleetd.java:315, which
// auto-launches each lead up to its `instances` count; neither is rebuilt on reload). So this
// compares fleet.leaders alone, not the whole Fleet record: comparing the whole record would
// report "split" for a tabLabel-only or charters-only change that is actually fully hot,
// which is the over-claim mirror of the under-claim bug this class exists to prevent.
if (!Objects.equals(leadersOf(old), leadersOf(fresh))) {
changed.add("fleet: fleet.leaders (each lead's tab, workspace, cwd, profile and "
+ "instances count) is read once at startup to build the LeadTabScanner's "
+ "identity map and to auto-launch leads, and neither is rebuilt on reload, so a "
+ "lead added, removed, or given a new tab: label needs a restart — until then it "
+ "stays unrecognised, and a caller from its new tab resolves as a worker, not a "
+ "lead; the rest of fleet: (developers, reviewers, charters, tabLabel) is read "
+ "live through the supplier on CompositePeerLauncher, and architects is read "
+ "live through that same supplier for placement AND through a separate supplier "
+ "on MemberRegistry for spawn-time identity — both already applied");
+ "lead; the rest of fleet: (architects, developers, reviewers, charters, "
+ "tabLabel) is read live through the supplier on CompositePeerLauncher and "
+ "already applied");
}
// Kept in step with SPLIT_KEYS the same way changedColdKeys is kept in step with COLD_KEYS —
// every message here must be traceable to one of the split keys the class doc documents.
@@ -16,15 +16,11 @@ import org.slf4j.LoggerFactory;
import java.io.IOException;
import java.io.UncheckedIOException;
import java.lang.reflect.InvocationTargetException;
import java.lang.reflect.Method;
import java.lang.reflect.Modifier;
import java.nio.file.Files;
import java.nio.file.Path;
import java.util.ArrayList;
import java.util.Arrays;
import java.util.Collections;
import java.util.Comparator;
import java.util.HashSet;
import java.util.LinkedHashMap;
import java.util.List;
@@ -128,15 +124,6 @@ import java.util.regex.PatternSyntaxException;
* under a member's long turn. {@code null} (the block omitted) behaves the
* same as an explicit {@code enabled: true}; set {@code enabled: false} to
* turn it off. See {@link dev.ltms.fleet.power.IdleSleepGuard}.
* @param models central allow-list of models any {@code profiles:} entry may name. {@code
* null} or an empty {@code allow:} ⇒ off: {@link #validateModels()} checks
* nothing and every existing config keeps working exactly as it does today.
* When non-empty, a profile whose {@code model:} is not one of {@link
* Models#ids()} fails config load, naming both the model and the profile —
* see {@link #validateModels()}. This block decides what may be CONFIGURED;
* fleetd #422 added the separate on/off question — whether a configured model
* may be spawned onto RIGHT NOW ({@link Models.ModelEntry#enabled}) — enforced
* live at spawn by {@code CompositePeerLauncher}, not here. See {@link Models}.
*/
@JsonIgnoreProperties(ignoreUnknown = true)
public record FleetConfig(
@@ -163,22 +150,7 @@ public record FleetConfig(
String worktreeGroup,
String memberLoginShell,
String memberSkills,
IdleSleepGuard idleSleepGuard,
Models models) {
/** Back-compat form before the {@code models:} block was added. */
public FleetConfig(Bind bind, String herdrSocket, String memberHerdrSocket, Map<String, Profile> profiles,
Guard guard, String worktreeRoot, Lifecycle lifecycle, Integer spawnReadyTimeoutMs,
Integer spawnReadyPollMs, Broker broker, Primary primary, Fleet fleet,
LeadHeartbeat leadHeartbeat, Health health, String placement, Auth auth,
ConfigReload configReload, Integer quarantineCooldownSeconds,
MemberCredentials memberCredentials, Coordinator coordinator, String worktreeGroup,
String memberLoginShell, String memberSkills, IdleSleepGuard idleSleepGuard) {
this(bind, herdrSocket, memberHerdrSocket, profiles, guard, worktreeRoot, lifecycle, spawnReadyTimeoutMs,
spawnReadyPollMs, broker, primary, fleet, leadHeartbeat, health, placement, auth,
configReload, quarantineCooldownSeconds, memberCredentials, coordinator, worktreeGroup,
memberLoginShell, memberSkills, idleSleepGuard, null);
}
IdleSleepGuard idleSleepGuard) {
/** Back-compat form before the {@code idleSleepGuard:} block was added. */
public FleetConfig(Bind bind, String herdrSocket, String memberHerdrSocket, Map<String, Profile> profiles,
@@ -1346,126 +1318,6 @@ public record FleetConfig(
}
}
/**
* Central allow-list of models any {@code profiles:} entry may name (fleetd ticket: "a central
* allow-list of usable models"). Nothing before this block checked a profile's {@code model:}
* against anything — it was a free-form string handed straight to the backend adapter, and a
* withdrawn or misspelled name failed silently (opencode falls back to a default model rather
* than erroring on an unknown {@code -m}) rather than at config load, where a mistake is cheap.
*
* <p><b>Absent or empty {@code allow:} is "off"</b>, on purpose: this is a large deployment
* with gitignored {@code fleetd.yaml} on more than one host, and a change that forced every
* operator to enumerate their models before the daemon would start would break every one of
* them on upgrade. See {@link FleetConfig#validateModels()}, which is where the allow-list is
* actually enforced, at config load.
*
* <p><b>The list is the authority; a {@code profiles:} entry is checked against it, never the
* other way around.</b> Adding or editing a {@code profiles:} entry cannot, by itself, widen
* the set of permitted models — only editing {@code models.allow:} itself can. This is the
* invariant the ticket asked for: the two blocks are validated in one direction only.
*
* <p><b>Spawn-time enforcement (fleetd #422, units 2+3) lives outside this record</b> —
* {@code CompositePeerLauncher.enforceModelEnabled} and its candidate-set filter read this
* block LIVE (through the same kind of supplier {@code weight}/{@code maxLoad} already use),
* so the on/off state below is hot: no restart needed. This block itself still only decides
* what may be CONFIGURED (membership in {@link #allow}); {@link ModelEntry#enabled} decides
* whether a member of that list is currently spawnable. The two questions are deliberately
* separate — see {@link ModelEntry}'s javadoc for why turning a model off must never mean
* removing it from {@link #allow}.
*
* @param allow the permitted models, each its own {@link ModelEntry} rather than a bare
* string — see that record's javadoc for why. {@code null}/empty ⇒ the block is
* treated as absent: {@link FleetConfig#validateModels()} checks nothing.
*/
@JsonIgnoreProperties(ignoreUnknown = true)
public record Models(List<ModelEntry> allow) {
public Models {
allow = (allow == null) ? List.of() : List.copyOf(allow);
}
/**
* One permitted model, named as a record rather than a bare string on purpose: this ticket
* (fleetd #422) is the "later unit" the original comment here predicted — it hangs an on/off
* state ({@link #enabled}) off each entry, and a bare {@code List<String>} could not have
* grown that field without changing the YAML shape underneath every operator who already
* wrote one. {@link #model()} is intentionally a single flat, opaque-string namespace — a
* bare Claude id ({@code claude-sonnet-5}) and an opencode provider-prefixed id
* ({@code openai/gpt-5.6-terra}) both fit it unchanged, because {@link
* FleetConfig#validateModels()} only ever compares a profile's {@code model:} value against
* this string for exact equality; it never parses a provider prefix or branches on a
* profile's {@code kind:}.
*
* @param model the model id exactly as a {@code profiles:} entry's {@code model:} field
* would name it
* @param enabled {@code false} turns spawning onto this model off; {@code null} (the field
* omitted — every config written before fleetd #422 is this shape) or
* {@code true} leaves it on. Turning a model off must NEVER remove it from
* {@link Models#allow} — {@link FleetConfig#validateModels()} checks
* <em>membership</em> only, never the on/off state, so an off entry stays a
* valid thing for a {@code profiles:} entry to name; only
* {@code CompositePeerLauncher}'s spawn-time gate reads {@link #enabled}.
* Collapsing the two — turning a model off by deleting its {@code allow:}
* entry — would make {@link FleetConfig#validateModels()} refuse the whole
* config reload the moment a still-configured profile names it, which is
* exactly the restart-to-flip-a-switch problem this field exists to avoid.
* <p>A model id can be named by more than one {@code profiles:} entry (e.g.
* {@code deepseek-v4-flash} backs both {@code local} and {@code
* local-direct} in the live config) — turning it off disables every profile
* that names it, on purpose: the model is what a subscription's rate limit
* actually constrains, not any one profile alias for it.
*/
@JsonIgnoreProperties(ignoreUnknown = true)
public record ModelEntry(String model, Boolean enabled) {
public ModelEntry {
model = (model == null || model.isBlank()) ? null : model.trim();
}
/**
* Back-compat form before {@link #enabled} was added (fleetd #422) — the model is
* unconditionally on, exactly as every {@code ModelEntry} behaved before this field
* existed. Keeps pre-#422 call sites (and any YAML that omits {@code enabled:})
* compiling and behaving identically.
*/
public ModelEntry(String model) {
this(model, null);
}
/** {@code true} unless {@link #enabled} is explicitly {@code false} — absent means on. */
public boolean isEnabled() {
return !Boolean.FALSE.equals(enabled);
}
}
/** {@link #allow}'s model ids, as a set for membership checks. Blank/null entries are dropped. */
public Set<String> ids() {
Set<String> ids = new java.util.LinkedHashSet<>();
for (ModelEntry e : allow) {
if (e != null && e.model() != null) {
ids.add(e.model());
}
}
return Collections.unmodifiableSet(ids);
}
/**
* Model ids currently turned off (fleetd #422: {@link ModelEntry#isEnabled()} {@code
* false}). Read live by {@code CompositePeerLauncher}'s spawn gate and by {@code
* fleet_profiles}/{@code GET /profiles} — both must read this same accessor off the same
* live config so the two surfaces cannot disagree about which model is off (the fleetd
* #404 lesson: a status field must read the source the behaviour reads).
*/
public Set<String> offIds() {
Set<String> off = new java.util.LinkedHashSet<>();
for (ModelEntry e : allow) {
if (e != null && e.model() != null && !e.isEnabled()) {
off.add(e.model());
}
}
return Collections.unmodifiableSet(off);
}
}
/**
* The terminal → lead-name map seeded from the legacy singular {@code primary:} pin (CB-530).
*
@@ -1717,7 +1569,7 @@ public record FleetConfig(
"lifecycle", "spawnReadyTimeoutMs", "spawnReadyPollMs", "broker", "primary", "fleet",
"leadHeartbeat", "health", "placement", "auth", "configReload", "quarantineCooldownSeconds",
"memberCredentials", "coordinator", "worktreeGroup", "memberLoginShell", "memberSkills",
"idleSleepGuard", "models");
"idleSleepGuard");
/** Load and validate config from {@code path}. */
public static FleetConfig load(Path path) {
@@ -2402,15 +2254,10 @@ public record FleetConfig(
// enabled: true — see its javadoc), so defaulting the block here would change nothing a
// reader observes and would only obscure that "block omitted" and "block present and
// enabled" are deliberately the same outcome.
// models is left as-is, like broker/primary/coordinator above: null/empty is "off", and an
// absent block must validate nothing (see Models's javadoc) — defaulting it here to an
// empty Models would be a no-op for validateModels() either way, since an empty allow-list
// already means "check nothing", so there is nothing to gain and one more null check to
// avoid by leaving it exactly as configured.
return new FleetConfig(b, herdrSocket, memberHerdrSocket, profiles, g, worktreeRoot, l, timeout, pollMs,
broker, primary, f, leadHeartbeat, health, placementOrDefault, a, configReload,
quarantineCooldown, mc, coordinator, worktreeGroup, memberLoginShell, memberSkills,
idleSleepGuard, models);
idleSleepGuard);
}
/**
@@ -2630,118 +2477,6 @@ public record FleetConfig(
}
}
/**
* Reject a {@code profiles:} entry whose {@code model:} is not on the configured {@link
* #models} allow-list.
*
* <p>Absent or empty {@code models.allow:} validates nothing — see {@link Models}'s javadoc:
* an existing config with no such block must keep working exactly as it does today. Once the
* operator declares at least one entry, every profile's {@code model:} (when set — a profile
* may legitimately leave it {@code null}, e.g. a {@code subscription: true} profile relying on
* the account's own default) must equal one of {@link Models#ids()} exactly. The comparison is
* a flat string match: a bare Claude id and an opencode {@code provider/model} id are both
* just opaque strings here, so nothing here needs to know which {@code kind:} a profile runs.
*
* <p>The check runs one direction only, by construction: it reads {@link #profiles} and
* {@link #models}, and only ever adds to {@code bad} when a profile's model is missing from
* the allow-list. Nothing here can be satisfied by widening a {@code profiles:} entry — only
* editing {@code models.allow:} itself changes what passes. That is the invariant the ticket
* asked for: the list is the authority, profiles are checked against it.
*
* @throws IllegalStateException when any profile names a model outside the configured
* allow-list, naming both the model and the profile that wanted it
*/
public void validateModels() {
if (models == null || models.allow().isEmpty()) {
return;
}
Set<String> allowed = models.ids();
List<String> bad = new ArrayList<>();
profiles.forEach((name, p) -> {
String model = p.model();
if (model != null && !model.isBlank() && !allowed.contains(model)) {
bad.add("profile '" + name + "' names model '" + model + "', which is not in "
+ "models.allow: (have: " + allowed + ").");
}
});
if (!bad.isEmpty()) {
throw new IllegalStateException("refusing to start: " + String.join(" ", bad));
}
}
/**
* Runs every validator this class declares — found by reflection, not by name.
*
* <p>fleetd ticket "central allow-list of usable models", follow-up: mutation testing found
* that although each of the six validators above was well pinned on its own, nothing proved
* either real caller ({@code Fleetd.main} and {@link ConfigRef#reload()}) still
* invoked it — deleting a call site left the full suite green. The fix is not a seventh test
* per caller; a hand-maintained list of six names here would have the exact same defect its
* own javadoc would warn against: the seventh validator someone adds next month has no reason
* to be added to it. So this method does not name any validator. It sweeps {@link
* #getClass()}'s own public, no-argument, {@code void} methods whose name starts with {@code
* "validate"} (excluding itself) and invokes every one it finds, via {@link
* #invokeAllValidators}. A new {@code validateXxx()} method is therefore wired into both
* callers the moment it is written — there is no second step to forget, and so no state in
* which it silently never runs.
*
* <p>{@code Fleetd.main} and {@link ConfigRef#reload()} each call this one method instead of
* the six individually — see the comments at those two call sites for why
* each must run it.
*
* <p>Methods run in a fixed (alphabetical) order, so a config with more than one violation
* always names the same one first, on every run.
*
* @throws IllegalStateException (or whatever unchecked exception a validator itself throws),
* propagated unchanged from the first validator, in that order,
* that finds a problem
*/
public void validateAll() {
invokeAllValidators(this);
}
/**
* The reflective sweep behind {@link #validateAll()}, kept as its own method — taking any
* {@code target}, not just {@code this} — so a test can prove the MECHANISM is generic (it
* would sweep a seventh {@code validateXxx()} method added to any class, not just something
* special-cased to today's six on {@link FleetConfig}) without needing to add a real, unwanted
* seventh validator to this class just to exercise that claim. See {@code
* FleetConfigValidateAllTest} for that proof.
*
* @param target an object whose public, no-argument, {@code void} methods named {@code
* validateXxx} (any name starting with {@code "validate"}, excluding {@code
* validateAll} itself) should all run, in alphabetical-by-name order
*/
static void invokeAllValidators(Object target) {
List<Method> methods = new ArrayList<>();
for (Method m : target.getClass().getMethods()) {
if (Modifier.isPublic(m.getModifiers())
&& m.getParameterCount() == 0
&& m.getReturnType() == void.class
&& m.getName().startsWith("validate")
&& !m.getName().equals("validateAll")) {
methods.add(m);
}
}
methods.sort(Comparator.comparing(Method::getName));
for (Method m : methods) {
try {
m.invoke(target);
} catch (InvocationTargetException e) {
Throwable cause = e.getCause();
if (cause instanceof RuntimeException re) {
throw re;
}
if (cause instanceof Error err) {
throw err;
}
throw new IllegalStateException("validator " + m.getName() + " failed", cause);
} catch (IllegalAccessException e) {
throw new IllegalStateException("cannot invoke validator " + m.getName(), e);
}
}
}
/** True for the loopback addresses and the unspecified-but-local forms we treat as same-host. */
private static boolean isLoopbackBind(String host) {
if (host == null || host.isBlank()) {
@@ -698,57 +698,17 @@ public final class CompletionResolver implements TurnListener {
}
/**
* What an unset pattern key means for the classification it configures (fleetd#415).
* {@code coverage()} cannot infer this from the key's name — the two keys it currently
* describes disagree on it, and a string comparison on the name would just move the same bug
* to a new spot — so every caller must state it explicitly.
*
* <p><strong>This alone does not prove a caller passes the right one for its key.</strong> A
* test that calls {@code coverage()} directly and supplies the meaning itself only proves this
* enum is worded correctly, never that {@code Fleetd}'s two call sites pair each key with its
* true meaning — that pairing is #415's actual defect. Measured on review: swapping the two
* {@code UnsetMeaning} arguments at those call sites (giving {@code exhaustedPattern} the
* built-in-default wording and {@code errorPattern} the off wording — #415's exact defect with
* the keys exchanged) compiled with 0 errors and left all 1506 existing tests green. See
* {@code dev.ltms.fleet.Fleetd#exhaustedPatternCoverageLine}/{@code #errorPatternCoverageLine}
* and {@code FleetdPatternCoverageLineTest}, which exists specifically to catch that swap.
*/
public enum UnsetMeaning {
/** No fallback exists: a profile with no configured pattern truly has this classification off. */
OFF,
/** A built-in pattern applies when unset: the classification still runs for that profile. */
BUILT_IN_DEFAULT
}
/**
* Coverage summary for a fleetd#201/CB-578-style pattern-key classification, logged at startup
* Coverage summary for the CB-578 stage A exhausted-pattern classification, logged at startup
* the way {@link dev.ltms.fleet.health.FleetHealthMonitor#coverage} is — so an operator can
* see whether the classification is on, and for which profiles, without reading every
* profile's config by hand.
*
* <p>fleetd#415: this method measures <em>pattern coverage</em> — how many profiles set the
* key — which is not the same thing as <em>feature state</em> for a key with a fallback. For
* {@code errorPattern}, an empty {@code configuredProfiles} still runs the classification
* against {@code CompletionResolver}'s built-in compatibility pattern ({@link #BACKEND_ERROR}
* at line ~84); for {@code exhaustedPattern} there is no fallback, so empty really does mean
* off. {@code unsetMeaning} is the single, required source of that fact — see
* {@link dev.ltms.fleet.config.FleetConfig#rejectMalformedProfilePatterns} lines ~2029-2032 for
* where it is documented for config authors. It is a required parameter, not a defaulted
* overload: a third pattern key added later must supply one to compile at all, rather than
* silently inheriting whichever wording this method happened to default to.
*
* @param allProfiles every configured profile name
* @param configuredProfiles the subset of {@code allProfiles} that carry the pattern
* @param configuredProfiles the subset of {@code allProfiles} that carry an exhausted pattern
*/
public static String coverage(String patternKey, UnsetMeaning unsetMeaning, Set<String> allProfiles,
Set<String> configuredProfiles) {
public static String coverage(String patternKey, Set<String> allProfiles, Set<String> configuredProfiles) {
if (configuredProfiles.isEmpty()) {
return switch (unsetMeaning) {
case OFF -> "off (no profile has an " + patternKey + " configured; profiles: "
+ sorted(allProfiles) + ")";
case BUILT_IN_DEFAULT -> "built-in default for all profiles (no profile customises "
+ patternKey + "; profiles: " + sorted(allProfiles) + ")";
};
return "off (no profile has an " + patternKey + " configured; profiles: " + sorted(allProfiles) + ")";
}
Set<String> unconfigured = new TreeSet<>(allProfiles);
unconfigured.removeAll(configuredProfiles);
@@ -37,7 +37,6 @@ import io.modelcontextprotocol.spec.McpSchema;
import com.fasterxml.jackson.databind.ObjectMapper;
import jakarta.servlet.http.HttpServlet;
import java.util.ArrayList;
import java.util.LinkedHashMap;
import java.util.List;
import java.util.Map;
@@ -120,28 +119,9 @@ public final class FleetMcp {
/**
* CB-578 stage B quarantine facts used by {@code fleet_profiles}: a profile → credential id
* lookup, plus the shared {@link BackendQuarantine} to read remaining cooldowns off.
*
* @param exhaustedPatternArmed fleetd #395: profile → whether that profile's {@code
* exhaustedPattern} is configured (see {@code
* FleetConfig.Profile#hasExhaustedPattern}), i.e. whether a backend refusal
* on it can EVER be classified {@code BACKEND_EXHAUSTED} and quarantine its
* credential. Bundled here, not a separate Source, because it answers the
* exact question {@code fleet_profiles}'s quarantine facts already answer
* for a QUARANTINED profile — "can this profile's usage limit ever be
* caught?" — just for every profile, not only one currently caught.
*/
public record QuarantineSource(Function<String, String> credentialIdFor, BackendQuarantine quarantine,
Function<String, Boolean> exhaustedPatternArmed) {
/**
* Backward-compatible 2-arg form, before fleetd #395 added {@code exhaustedPatternArmed} —
* reports every profile unarmed. Keeps every pre-existing call site (production and test)
* compiling and behaving identically for the quarantine facts they actually asked for.
*/
public QuarantineSource(Function<String, String> credentialIdFor, BackendQuarantine quarantine) {
this(credentialIdFor, quarantine, _ -> false);
}
/** Inert source — no profile is ever reported quarantined or armed. Explicit stand-in, not a default. */
public record QuarantineSource(Function<String, String> credentialIdFor, BackendQuarantine quarantine) {
/** Inert source — no profile is ever reported quarantined. Explicit stand-in, not a default. */
public static QuarantineSource none() { return new QuarantineSource(_ -> null, BackendQuarantine.none()); }
}
@@ -366,7 +346,7 @@ public final class FleetMcp {
String self = callerTerminal(exchange);
McpSchema.CallToolResult denied = deny(exchange, toolAction("fleet_reply", req.arguments()), self);
if (denied != null) return denied;
return reply(messages, self, principal(exchange).role(), str(req.arguments(), "content"));
return reply(messages, self, str(req.arguments(), "content"));
};
// fleet_ask (CB-205): a worker's mid-turn question — identity from the CONNECTION.
BiFunction<McpSyncServerExchange, McpSchema.CallToolRequest, McpSchema.CallToolResult> askHandler =
@@ -888,26 +868,19 @@ public final class FleetMcp {
/**
* {@code fleet_reply}: the worker returns its structured answer, resolving the awaiting send
* or — when no send is open — queueing the reply in the inbox for later drain (CB-307).
* {@code callerTerminal} and {@code callerRole} are resolved from the connection (never an
* argument). A {@code null} terminal means the caller is not a known worker. A PRIMARY with a
* terminal is a lead and must use {@code fleet_send}, because reply has no peer-lead route.
* {@code callerTerminal} is resolved from the connection (never an argument); a {@code null}
* means the caller is not a known worker (e.g. the primary called it by mistake).
*
* <p>fleetd #365: the result text names which of those actually happened
* ({@link MessageService.ReplyOutcome#description()}) instead of the single word "delivered"
* for both — a queued reply is a real success, but it is not the same fact as one that resolved
* a live waiter, and the caller could not previously tell them apart.
*/
static McpSchema.CallToolResult reply(MessageService messages, String callerTerminal, Role callerRole, String content) {
static McpSchema.CallToolResult reply(MessageService messages, String callerTerminal, String content) {
if (callerTerminal == null) {
return error("fleet_reply is for workers only — could not identify the calling worker "
+ "from the connection");
}
if (callerRole == Role.PRIMARY) {
return error("fleet_reply has no route to a peer lead. Use fleet_send{coordId: ...} for a peer on another "
+ "daemon or fleet_send{sessionId: ...} for a peer on this host. fleet_reply resolves a member's "
+ "blocked fleet_send, and a peer's coord-id message is durable and non-blocking, so there is "
+ "nothing for it to resolve.");
}
// fleetd #302: isBlank, not == null, to match fleet_send's own guard above. MessageService
// .reply now REJECTS blank content, and this handler is a bare BiFunction with no try/catch
// around it — so a whitespace-only fleet_reply would leave here as an uncaught
@@ -1136,14 +1109,6 @@ public final class FleetMcp {
* not the other, and the two then disagree about a live outage. That is exactly what fleetd
* #284 was, where one rule computed in two places was widened in only one and a single response
* contradicted itself. Shared inputs do not make duplicated computation safe.
*
* <p>fleetd #395: also reports {@code exhaustionDetectionArmed}, one boolean per configured
* profile — {@code true} when that profile's {@code exhaustedPattern} is set, {@code false}
* when it is not, so an operator can tell "this profile is healthy" from "nothing can ever
* quarantine this profile" without reading {@code fleetd.yaml}. Unlike {@code quarantined}/
* {@code coolingOff}, this map always names every profile: an unarmed profile never enters a
* transient state to be absent from, so silence here would read as "healthy" rather than "not
* being watched at all".
*/
public static Map<String, Object> profilesView(PeerLauncher workers, QuarantineSource quarantine, OutageSource outage) {
Map<String, Object> result = new LinkedHashMap<>();
@@ -1151,9 +1116,7 @@ public final class FleetMcp {
result.put("default", workers.defaultProfile() == null ? "" : workers.defaultProfile());
Map<String, Object> quarantined = new LinkedHashMap<>();
Map<String, Object> coolingOff = new LinkedHashMap<>();
Map<String, Object> exhaustionDetectionArmed = new LinkedHashMap<>();
for (String profile : workers.profiles()) {
exhaustionDetectionArmed.put(profile, quarantine.exhaustedPatternArmed().apply(profile));
String credentialId = quarantine.credentialIdFor().apply(profile);
if (credentialId != null) {
quarantine.quarantine().remainingSeconds(credentialId).ifPresent(remaining -> {
@@ -1173,30 +1136,12 @@ public final class FleetMcp {
});
}
}
result.put("exhaustionDetectionArmed", exhaustionDetectionArmed);
if (!quarantined.isEmpty()) {
result.put("quarantined", quarantined);
}
if (!coolingOff.isEmpty()) {
result.put("coolingOff", coolingOff);
}
// fleetd #422: read the exact same accessor CompositePeerLauncher's spawn gate reads
// (PeerLauncher.modelGateState(), which for the composite is models0() read live) — never a
// separately-derived answer, so this status can never overstate or understate what the gate
// actually enforces (the fleetd #404 lesson).
//
// fleetd #422 follow-up: "armed" and "off" come from the ONE modelGateState() call below,
// never two independent reads of the gate — a reload landing between two separate reads
// could otherwise make them disagree. modelGateArmed is reported unconditionally (never
// omitted like quarantined/coolingOff above) precisely so a lead can tell "no models: block
// at all" (false) apart from "a models: block with nothing currently off" (true, with
// modelsOff simply absent below) — the two states PeerLauncher.disabledModels() alone
// cannot distinguish, both reporting an empty set.
PeerLauncher.ModelGateState modelGate = workers.modelGateState();
result.put("modelGateArmed", modelGate.configured());
if (!modelGate.off().isEmpty()) {
result.put("modelsOff", new ArrayList<>(modelGate.off()));
}
return result;
}
@@ -1718,11 +1663,7 @@ public final class FleetMcp {
"List the configured worker profiles (backends) and which one fleet_spawn uses by "
+ "default. A 'quarantined' map is present when a backend-exhausted refusal put "
+ "a profile's credential on cooldown — fleet_spawn onto it is refused until "
+ "quarantinedForSeconds elapses; a profile sharing that credential is listed too. "
+ "'exhaustionDetectionArmed' reports, per profile, whether a usage-limit refusal "
+ "on it can EVER be classified and quarantined (its exhaustedPattern is "
+ "configured) — false means that profile's credential can never be quarantined "
+ "by this mechanism, however many usage-limit refusals it sees.",
+ "quarantinedForSeconds elapses; a profile sharing that credential is listed too.",
objectSchema(Map.of(), List.of()));
}
@@ -90,19 +90,6 @@ public final class CompositePeerLauncher implements PeerLauncher {
private final Supplier<Map<String, FleetConfig.Profile>> profileConfigs;
private final Supplier<PlacementPolicy> placementPolicy;
/**
* fleetd #422: the central model allow-list's on/off state, read live per spawn — same reason
* {@link #profileConfigs} is a supplier rather than a captured map (see the class doc above and
* {@link #enforceModelEnabled}). A caller with no {@code models:} block to read from (the
* simpler, map-based constructors used throughout this class's own tests) wires this to a
* constant {@code null}, which {@link #models0} treats as "nothing configured, gate never
* fires" — the pre-#422 behaviour.
*/
private final Supplier<FleetConfig.Models> models;
/** The value {@link #models0} normalizes a {@code null} supplier result to. */
private static final FleetConfig.Models NO_MODELS_CONFIGURED = new FleetConfig.Models(List.of());
/** CB-578 stage B: credential cooldown, checked before an explicit spawn and filtered into placement. */
private final BackendQuarantine quarantine;
@@ -218,44 +205,12 @@ public final class CompositePeerLauncher implements PeerLauncher {
FleetConfig.Fleet fleet,
BackendQuarantine quarantine,
BackendOutagePolicy outagePolicy) {
// No models: block to read from a plain profiles Map — fleetd #422's gate is wired to a
// constant null (see enforceModelEnabled/models0), the pre-#422 behaviour for every caller
// of this overload. Use the 9-arg overload below to test the gate against a LIVE supplier.
this(delegates, defaultProfile, profileConfigs, placementPolicy, liveCount, fleet, quarantine,
outagePolicy, constant(null));
}
/**
* As above, plus a LIVE model-gate source (fleetd #422). The 8-arg overload above wires
* {@code models} to a constant {@code null} because it has only a static {@code Map<String,
* Profile>}, never a full {@code FleetConfig}, to read one from; this overload exists so a test
* can prove {@link #enforceModelEnabled} and its candidate filter re-read {@link
* FleetConfig.Models} on every call rather than a value captured once at construction — the
* exact distinction fleetd #422 exists to get right (see the class doc's CB-559 note on {@link
* #profileConfigs}, which this follows). Production wiring uses the
* {@code Supplier<FleetConfig>} constructor below instead, which already threads a live
* {@code config.get().models()} through.
*
* @param models required — pass a supplier returning {@code null} for a caller that has no
* {@code models:} block to gate against, never a defaulting overload (the same
* "explicit opt-out, never a silent default" rule {@code quarantine}/
* {@code outagePolicy} already follow).
*/
public CompositePeerLauncher(List<HerdrPeerLauncher> delegates,
String defaultProfile,
Map<String, FleetConfig.Profile> profileConfigs,
PlacementPolicy placementPolicy,
Function<String, Integer> liveCount,
FleetConfig.Fleet fleet,
BackendQuarantine quarantine,
BackendOutagePolicy outagePolicy,
Supplier<FleetConfig.Models> models) {
// LinkedHashMap, not Map.copyOf: candidates() promises definition order and the weighted
// policy breaks exact-weight ties on it, so a salted iteration order would make placement
// differ from one JVM run to the next.
this(delegates, defaultProfile,
constant(Collections.unmodifiableMap(new LinkedHashMap<>(profileConfigs))),
constant(placementPolicy), liveCount, constant(fleet), quarantine, outagePolicy, models);
constant(placementPolicy), liveCount, constant(fleet), quarantine, outagePolicy);
}
/**
@@ -295,10 +250,7 @@ public final class CompositePeerLauncher implements PeerLauncher {
liveCount,
() -> config.get().fleet(),
quarantine,
outagePolicy,
// fleetd #422: read live, same as profiles/placement/fleet above — a models.allow
// edit (on/off or otherwise) is visible to the very next spawn, no restart needed.
() -> config.get().models());
outagePolicy);
}
/** The all-suppliers form every other constructor funnels into. */
@@ -309,8 +261,7 @@ public final class CompositePeerLauncher implements PeerLauncher {
Function<String, Integer> liveCount,
Supplier<FleetConfig.Fleet> fleet,
BackendQuarantine quarantine,
BackendOutagePolicy outagePolicy,
Supplier<FleetConfig.Models> models) {
BackendOutagePolicy outagePolicy) {
this.fleet = fleet;
this.quarantine = Objects.requireNonNull(quarantine, "quarantine");
this.outagePolicy = Objects.requireNonNull(outagePolicy, "outagePolicy");
@@ -322,7 +273,6 @@ public final class CompositePeerLauncher implements PeerLauncher {
this.profileConfigs = profileConfigs;
this.placementPolicy = placementPolicy;
this.liveCount = liveCount;
this.models = Objects.requireNonNull(models, "models");
Map<String, HerdrPeerLauncher> index = new LinkedHashMap<>();
for (HerdrPeerLauncher d : this.delegates) {
for (String profile : d.profiles()) {
@@ -353,16 +303,6 @@ public final class CompositePeerLauncher implements PeerLauncher {
return m == null ? Map.of() : m;
}
/**
* The currently-configured {@code models:} block, never null (fleetd #422). Read fresh on
* every call, the same reason {@link #profiles0} is — a config reload's on/off edit must reach
* the very next spawn.
*/
private FleetConfig.Models models0() {
FleetConfig.Models m = models.get();
return m == null ? NO_MODELS_CONFIGURED : m;
}
/** The adapter owning {@code profileName} (null/blank → the default). Throws on an unknown profile. */
private HerdrPeerLauncher route(String profileName) {
String resolved = (profileName == null || profileName.isBlank()) ? defaultProfile : profileName;
@@ -392,7 +332,6 @@ public final class CompositePeerLauncher implements PeerLauncher {
enforceNotQuarantined(requestedProfile);
enforceNotCoolingOff(requestedProfile);
enforceMaxLoad(requestedProfile);
enforceModelEnabled(requestedProfile);
PeerHandle handle = d.spawn(req);
spawnedBy.put(handle.id(), d);
return handle;
@@ -410,11 +349,8 @@ public final class CompositePeerLauncher implements PeerLauncher {
Set<String> quarantined = quarantinedProfiles(candidates);
// fleetd #201 Unit 5: a distinct set from quarantined — see PlacementContext.coolingOff.
Set<String> coolingOff = coolingOffProfiles(candidates);
// fleetd #422: read live per spawn, same as quarantined/coolingOff above — a config reload
// that flips a model's enabled state is visible to the very next unqualified spawn.
Set<String> modelOff = modelOffProfiles(candidates);
PlacementContext ctx = new PlacementContext(roleDefault, candidates, liveCount, unreachable,
quarantined, coolingOff, modelOff);
quarantined, coolingOff);
int maxAttempts = candidates.isEmpty() ? 1 : candidates.size();
for (int attempt = 0; attempt < maxAttempts; attempt++) {
@@ -444,7 +380,7 @@ public final class CompositePeerLauncher implements PeerLauncher {
unreachable.add(chosen.profile());
// Update the context for the next selection so the policy excludes this profile.
ctx = new PlacementContext(roleDefault, candidates, liveCount, unreachable,
quarantined, coolingOff, modelOff);
quarantined, coolingOff);
}
}
@@ -556,49 +492,6 @@ public final class CompositePeerLauncher implements PeerLauncher {
}
}
/**
* Refuse an explicit-profile spawn whose {@code model:} the operator has turned off in the
* central {@code models.allow:} list (fleetd #422). Deliberately worded apart from {@link
* #enforceNotQuarantined} and {@link #enforceNotCoolingOff}: those two report a BACKEND-reported
* outage (exhaustion, repeated errors); this one reports an OPERATOR decision, so the message
* says "turned off" and names the model, never "quarantined" or "cooling off". A FOURTH,
* independent reason to refuse a spawn — never layered onto {@code BackendQuarantine} or {@code
* BackendOutagePolicy}, which would misattribute an operator's own choice to the backend.
*
* <p>Reads {@link #models0()} fresh on every call — the same liveness {@link #profiles0()}
* already has — so flipping {@code enabled: false} and reloading takes effect on the very next
* spawn, no restart (criterion 4). A profile naming no model, or a model absent from {@code
* models.allow:} entirely (nothing to gate against), is never refused here.
*
* @throws PlacementException naming the model and the profile, distinct from quarantine/cool-off
*/
private void enforceModelEnabled(String profile) {
FleetConfig.Profile cfg = profiles0().get(profile);
String model = (cfg == null) ? null : cfg.model();
if (model == null) {
return;
}
if (models0().offIds().contains(model)) {
throw new PlacementException("worker profile '" + profile + "' names model '" + model
+ "', which the operator has turned off in models.allow — refusing spawn");
}
}
/** The subset of {@code candidates} whose {@code model:} is currently turned off (fleetd #422). */
private Set<String> modelOffProfiles(List<PlacementCandidate> candidates) {
Set<String> off = models0().offIds();
if (off.isEmpty()) {
return Set.of();
}
return candidates.stream()
.map(PlacementCandidate::profile)
.filter(p -> {
FleetConfig.Profile cfg = profiles0().get(p);
return cfg != null && cfg.model() != null && off.contains(cfg.model());
})
.collect(Collectors.toSet());
}
/**
* The profile names {@code role} may be placed on, in definition order.
*
@@ -643,32 +536,6 @@ public final class CompositePeerLauncher implements PeerLauncher {
return route(profileName).parityOverlay(profileName);
}
/**
* fleetd #422 follow-up: the single live read that answers both "is the models.allow: gate
* armed" and "which models are off", off the exact same accessor ({@link #models0()}) {@link
* #enforceModelEnabled} and {@link #modelOffProfiles} read — so {@code fleet_profiles}/{@code
* GET /profiles} (via {@link PeerLauncher#disabledModels()}, which now delegates here) can
* never report a different answer than the gate enforces (the fleetd #404 lesson), and the
* startup log line built from this can never disagree with either.
*
* <p>{@link #models0()} itself normalizes a {@code null} {@link #models} read to the shared
* {@link #NO_MODELS_CONFIGURED} sentinel — deliberately the one object no config-supplied
* {@code Models} instance can ever be identical to, since it is private to this class — so
* comparing by reference here recovers exactly the fact {@code models0()}'s normalization
* would otherwise erase: whether the live source was {@code null} (no {@code models:} block,
* armed = false) or a real, config-supplied block (armed = true, even one whose {@code allow:}
* is itself empty or absent — {@link FleetConfig.Models}'s "absent or empty allow: is off"
* wording governs config-load validation, a distinct question from whether this gate is armed
* for reporting).
*/
@Override
public PeerLauncher.ModelGateState modelGateState() {
FleetConfig.Models m = models0();
return m == NO_MODELS_CONFIGURED
? PeerLauncher.ModelGateState.notConfigured()
: PeerLauncher.ModelGateState.armed(m.offIds());
}
@Override
public void stop(String id) {
HerdrPeerLauncher d = spawnedBy.get(id);
@@ -75,39 +75,9 @@ import java.util.stream.Stream;
* never to the control.
*
* <p>The scrub also writes {@code scrub-report.txt} into its own directory: one {@code allowed N of
* M failed F} line (N = exports left untouched, M = exports present when the scrub ran, F = names
* the scrub attempted to blank but could not), then the NAMES — blanked ones bare, unblankable ones
* {@code !}-prefixed — never values. The launcher reads this back at teardown and logs it, because a
* blocked count next to an unknown denominator is not a finding.
*
* <p><b>fleetd #394:</b> plain {@code export "$n="} is a FATAL error for a zsh read-only or special
* parameter (for example {@code UID}) — it aborts the whole sourced file, so every name still to
* come is never blanked and the report above is never written at all. The blanking loop instead
* routes each attempt through {@code eval}, which contains that error to the single iteration: the
* loop always finishes, and a name that could not be blanked is counted as {@code failed} and
* listed {@code !}-prefixed rather than silently disappearing. This is deliberately not a skip-list
* of known-bad names — every enumerated name is still attempted, so a name nobody has thought of
* yet still gets tried and, if it fails, still gets counted.
*
* <p>The blanking loop also re-asserts, on its own, the same {@code [A-Za-z_][A-Za-z0-9_]*} shape
* check the enumeration loop already applied. Before {@code eval} was introduced a non-conforming
* name reaching {@code export "$n="} was harmless either way — the quoting made it inert. With
* {@code eval}, the name is spliced into a string and interpreted as shell syntax, so the enumeration
* loop's check is no longer sufficient on its own to keep that call site safe — it is a guard on a
* different loop, and the two must not silently drift apart. Re-checking right before the
* {@code eval} keeps that call site safe by its own reading, independent of whatever the enumeration
* loop does or stops doing in a later change.
*
* <p><b>fleetd #400:</b> {@code eval}'s exit status is not proof that the blank actually happened.
* zsh coerces a bare {@code NAME=} assignment on an integer special parameter (measured on macOS zsh
* 5.9: {@code SECONDS}, {@code RANDOM}, {@code SHLVL}, {@code HISTSIZE}, {@code COLUMNS},
* {@code LINES}, {@code USERNAME}) to a number instead of failing — {@code eval} returns success,
* the value is untouched, and a status-based classification reports it as blanked when it was not.
* The fix classifies on the observed effect instead: after the attempt, the name's value is read
* back with the {@code (P)} indirection flag and the decision is made from whether that is now
* empty. This one check covers all three shapes a name can take at this point — a genuine blank, a
* fatal read-only error {@code eval} merely contained, and this silent no-op — and the exit status
* plays no part in the decision at all.
* M} line (N = exports left untouched, M = exports present when the scrub ran), then the blanked
* NAMES — never values. The launcher reads this back at teardown and logs it, because a blocked
* count next to an unknown denominator is not a finding.
*/
public final class EnvAllowListScrub {
@@ -162,16 +132,10 @@ public final class EnvAllowListScrub {
}
/**
* A parsed {@code scrub-report.txt}: how many exported variables existed when the scrub ran
* ({@code total}), how many were left untouched ({@code allowed}), how many the scrub attempted
* to blank but could not ({@code failed} — fleetd #394: a zsh read-only/special parameter such
* as {@code UID} fatally errors on plain {@code export NAME=}, so those attempts go through
* {@code eval} instead so the loop keeps going and the failure is counted rather than left
* invisible), and the NAMES in each of the latter two categories. {@code allowed +
* blanked.size() + unblankable.size() == total}, and {@code unblankable.size() == failed}.
* Values never appear.
* A parsed {@code scrub-report.txt}: how many exported variables existed when the scrub ran,
* how many were left untouched (allowed), and the NAMES that were blanked. Values never appear.
*/
record ScrubReport(int allowed, int total, int failed, List<String> blanked, List<String> unblankable) {
record ScrubReport(int allowed, int total, List<String> blanked) {
}
/**
@@ -370,57 +334,38 @@ public final class EnvAllowListScrub {
_cb633_blank+=("$_cb633_n")
done
# fleetd #394: plain `export "$n="` is FATAL for a zsh read-only/special parameter
# (e.g. UID) and aborts this whole sourced file — every name still to come is never
# blanked, and the report below is never written, silently. `eval` contains that
# error to the single iteration instead: it still fails for that one name, but the
# loop continues and we can tell allowed / blanked / unblankable apart afterwards.
# This is not a skip-list of known-bad names (that would miss the next one nobody
# thought of) — every name in _cb633_blank is still attempted, unconditionally.
# Every name reaching this loop already passed the identical identifier check in the
# enumeration loop above — but that guard is 20 lines away in a different loop, and
# this line is about to splice the name into a string handed to `eval`. Before this
# fix the name only ever reached `export` quoted ("$n="), which is inert on a
# non-identifier string either way; `eval` makes THIS line the only thing standing
# between such a string and code execution in the member's pane, so it re-asserts the
# same check on its own rather than trusting a guard it does not own. Under normal
# operation this can never fire (the enumeration guard already filtered everything
# reaching _cb633_blank), so a name caught here is counted as unblankable rather than
# silently dropped — it is real evidence that the upstream guard was bypassed.
# `export UID=` is not a failed command: it is a FATAL zsh parameter error
# ("failed to change user ID") that aborts this whole sourced file mid-loop,
# leaving every later name unscrubbed and the report below unwritten — silently,
# because of the 2>/dev/null. Neither `|| true` nor a `${(t)n}` type guard
# contains it; only `eval` does. `eval` is safe here precisely because the loop
# above already rejected every name that is not [A-Za-z_][A-Za-z0-9_]*, so
# nothing but a bare identifier can reach it.
#
# fleetd #400: the attempt's own exit status is NOT proof of its effect. zsh coerces
# a bare `NAME=` assignment on an integer special parameter (SECONDS, RANDOM, SHLVL,
# HISTSIZE, COLUMNS, LINES, USERNAME on this host) to a number instead of failing —
# `eval` returns 0, the value is untouched, and the old exit-status check reported it
# as blanked when it was not. Classify on the observed effect instead: attempt the
# export, then read the name's value back with the `(P)` indirection flag and decide
# from whether it is now empty. One check then covers all three shapes a name can
# take here — a genuine blank, a fatal read-only error `eval` merely contained, and
# this silent no-op — without the exit status entering the decision at all.
typeset -a _cb633_ok _cb633_unblankable
_cb633_ok=()
_cb633_unblankable=()
# Enumerating the special names instead (UID|EUID|GID|EGID|PPID|LINENO) also
# works, but only for the ones enumerated: a special that turns up exported on
# some other host brings the abort straight back. `eval` contains all of them.
#
# Then VERIFY. A contained failure is still a failure, so a name that did not
# actually blank must not be reported as blanked. It currently falls into the
# "allowed" count, which is imprecise in the safe direction; the honest third
# count ("tried and could not blank") needs a report-format change and belongs
# with fleetd #394, not here.
typeset -a _cb633_done
_cb633_done=()
for _cb633_n in "${_cb633_blank[@]}"; do
if [[ ! "$_cb633_n" =~ ^[A-Za-z_][A-Za-z0-9_]*$ ]]; then
_cb633_unblankable+=("$_cb633_n")
continue
fi
eval "export ${_cb633_n}=" 2>/dev/null
if [[ -z "${(P)_cb633_n}" ]]; then
_cb633_ok+=("$_cb633_n")
else
_cb633_unblankable+=("$_cb633_n")
fi
[[ -z "${(P)_cb633_n}" ]] && _cb633_done+=("$_cb633_n")
done
_cb633_blank=("${_cb633_done[@]}")
integer _cb633_kept=$(( _cb633_total - ${#_cb633_blank} ))
{
print -r -- "allowed $_cb633_kept of $_cb633_total failed ${#_cb633_unblankable}"
for _cb633_n in "${_cb633_ok[@]}"; do print -r -- "$_cb633_n"; done
for _cb633_n in "${_cb633_unblankable[@]}"; do print -r -- "!$_cb633_n"; done
print -r -- "allowed $_cb633_kept of $_cb633_total"
for _cb633_n in "${_cb633_blank[@]}"; do print -r -- "$_cb633_n"; done
} > "$ZDOTDIR/%s" 2>/dev/null
unset _cb633_allowed _cb633_names _cb633_blank _cb633_ok _cb633_unblankable _cb633_n _cb633_total _cb633_kept
unset _cb633_done _cb633_allowed _cb633_names _cb633_blank _cb633_n _cb633_total _cb633_kept
""".formatted(names, MemberEnvAllowList.zshCasePattern(), REPORT_FILE);
}
@@ -434,11 +379,6 @@ public final class EnvAllowListScrub {
* Read and parse {@link #REPORT_FILE} out of a generated ZDOTDIR directory. Returns {@code null}
* when absent or unreadable (the pane may have been torn down before its login shell ever got to
* the scrub) — callers treat that as "no measurement available", never as success.
*
* <p>First line is {@code "allowed <N> of <M> failed <F>"} (fleetd #394 added the trailing
* {@code failed <F>} — a count of names the scrub attempted to blank but could not, e.g. a zsh
* read-only/special parameter). Every following non-blank line is a name: a bare name was
* blanked, a {@code !}-prefixed name was attempted and failed. Values never appear on either.
*/
static ScrubReport readReport(Path zdotdir) {
Path report = zdotdir.resolve(REPORT_FILE);
@@ -451,24 +391,17 @@ public final class EnvAllowListScrub {
return null;
}
String[] parts = lines.getFirst().substring("allowed ".length()).trim().split("\\s+");
if (parts.length != 5 || !"of".equals(parts[1]) || !"failed".equals(parts[3])) {
if (parts.length != 3 || !"of".equals(parts[1])) {
return null;
}
List<String> blanked = new ArrayList<>();
List<String> unblankable = new ArrayList<>();
for (int i = 1; i < lines.size(); i++) {
String line = lines.get(i);
if (line.isBlank()) {
continue;
}
if (line.startsWith("!")) {
unblankable.add(line.substring(1));
} else {
blanked.add(line);
if (!lines.get(i).isBlank()) {
blanked.add(lines.get(i));
}
}
return new ScrubReport(Integer.parseInt(parts[0]), Integer.parseInt(parts[2]),
Integer.parseInt(parts[4]), List.copyOf(blanked), List.copyOf(unblankable));
List.copyOf(blanked));
} catch (IOException | NumberFormatException e) {
return null;
}
@@ -1643,19 +1643,11 @@ public abstract class HerdrPeerLauncher implements PeerLauncher {
log.warn("memberCredentials allow-list: pane {} left no scrub report in {} — the "
+ "environment scrub cannot be confirmed to have run. Either the pane ended "
+ "before its shell finished starting, or its shell never read our generated "
+ "startup files. Either way, we cannot tell from here whether the scrub ran, "
+ "so we do not know what that member's environment contained.",
+ "startup files, in which case that member saw the full host environment.",
paneId, dir);
} else {
log.info("memberCredentials allow-list: pane {} allowed {} of {} environment variables",
paneId, report.allowed(), report.total());
if (report.failed() > 0) {
log.warn("memberCredentials allow-list: pane {} could not blank {} environment "
+ "variable(s) — {} (likely a zsh read-only/special parameter) — those "
+ "names were left in the member's environment. Confirm none of them is a "
+ "credential.",
paneId, report.failed(), report.unblankable());
}
List<String> shaped = report.blanked().stream()
.filter(name -> CREDENTIAL_SHAPED_NAME.matcher(name).matches())
.toList();
@@ -1317,27 +1317,6 @@ public final class MessageService {
return new TaskView(ticket, Phase.FAILED, null, null, detail, null);
}
/**
* Test seam only — carries no production behaviour, and nothing in this class calls it;
* {@link #pruneTerminalTickets} still reads {@link Task#completedNanos} directly.
*
* <p>Reports whether {@code ticket}'s completion hook (the {@code whenComplete} registered in
* {@link Task}'s constructor) has actually run yet. Exists because {@link #poll} can report
* {@link Phase#DONE} for a ticket before that hook fires: {@code CompletableFuture.complete()}
* publishes its result and only afterwards runs dependent actions such as {@code whenComplete}
* (fleetd #399), so a caller that observes the future done via {@link #poll} is not thereby
* guaranteed to also observe {@link Task#completedNanos} stamped. A test that must order both
* events — e.g. before advancing an injected clock past the TTL, to avoid stamping the
* *advanced* time and masking a real eviction bug — waits on this instead of on
* {@link Phase#DONE}.
*
* @return {@code false} for an unknown ticket or one whose completion hook has not run yet
*/
boolean isCompletionStampedForTest(String ticket) {
Task task = tasks.get(ticket);
return task != null && task.completedNanos != null;
}
/** Best-effort live worker status for a pending poll; never throws (a lookup error is just noise). */
private String liveStatus(String target) {
try {
@@ -230,26 +230,6 @@ public final class ReplyPushLoop {
.collect(Collectors.toUnmodifiableSet());
}
/**
* Test seam only (fleetd #418): carries no production behaviour, and nothing in this class
* calls it. Exposes {@link #pendingQuestionTurnIdsFor} — the exact state {@link #decide} reads
* to decide whether a question keeps a lead's schedule alive.
*
* <p>{@code MessageService.ask()} does three things in order before a question is fully open to
* this loop: it flips the ticket's {@code poll()} phase to {@code Phase.ASKING}, then resolves
* the reverse-rendezvous waiter, then calls {@link #onQuestionOpened}, which is what actually
* populates {@link #pendingQuestions}. A test that barriers on {@code Phase.ASKING} observes only
* the first of those three steps — under load the asker thread can be descheduled between steps
* one and three, so the barrier releases before this method's underlying map is populated, and
* {@link #decide} correctly reports nothing pending yet. A test that must order itself after the
* state {@link #decide} actually reads waits on this instead of on the phase.
*
* @return an unmodifiable snapshot; empty for a lead with no open questions
*/
Set<String> pendingQuestionTurnIdsForTest(String lead) {
return pendingQuestionTurnIdsFor(lead);
}
private List<PendingIncident> pendingIncidentsFor(String lead) {
return pendingIncidents.values().stream().filter(i -> lead.equals(i.key().lead())).toList();
}
@@ -192,66 +192,4 @@ public interface PeerLauncher {
* @return {@code true} when a reset was sent and its status transition must settle before reuse
*/
boolean clearContext(String id);
/**
* Model ids the operator has currently turned off in the central {@code models.allow:} list
* (fleetd #422) — empty for a launcher with nothing to gate against. {@code fleet_profiles}/
* {@code GET /profiles} (via {@code FleetMcp.profilesView}) call this to report which models
* are off, and MUST read this exact accessor rather than deriving their own answer: the fleetd
* #404 lesson is that a status field reading a different source than the behaviour it describes
* can drift from what the gate ({@code CompositePeerLauncher.enforceModelEnabled} and its
* candidate filter) actually enforces. A default of {@code Set.of()} keeps every other {@link
* PeerLauncher} implementer (the herdr adapters, and the two test-fake implementers) unchanged.
*
* <p>fleetd #422 follow-up: this alone cannot tell "no {@code models:} block at all" from "a
* {@code models:} block where nothing is currently off" — both report an empty set here. Delegates
* to {@link #modelGateState()} so the two facts always come from the one read {@link
* #modelGateState()}'s implementer makes; do not override this method separately from that one.
*/
default Set<String> disabledModels() {
return modelGateState().off();
}
/**
* Whether the central {@code models.allow:} gate (fleetd #422) is armed at all, together with
* which model ids are currently off — fleetd #422 follow-up. {@link #disabledModels()} alone
* cannot distinguish two states that both report an empty set: a host with no {@code models:}
* block (nothing is gated, and nothing can be) and a host WITH a {@code models:} block where
* nothing is currently turned off (the gate is armed and reporting zero). This method exists so
* a caller — the startup log, {@code fleet_profiles}/{@code GET /profiles} — can tell the two
* apart, the same reason {@code CompletionResolver.UnsetMeaning} exists: an accessor that can
* legitimately report "empty" must never let a caller guess why.
*
* <p>Default {@link ModelGateState#notConfigured()} — every launcher without a {@code models:}
* block to read from (the herdr adapters, and the two test-fake implementers), matching {@link
* #disabledModels()}'s own default of an empty set.
*/
default ModelGateState modelGateState() {
return ModelGateState.notConfigured();
}
/**
* fleetd #422 follow-up: the result of {@link #modelGateState()} — see that method's javadoc
* for why "armed" and "off" must be reported together from one read rather than as two
* separately-derived facts that a reload landing between them could make disagree.
*
* @param configured {@code true} when a {@code models:} block exists at all (armed), regardless
* of whether anything in it is currently turned off; {@code false} when there
* is no block to gate against
* @param off the model ids currently turned off; always empty when {@code configured} is
* {@code false}
*/
record ModelGateState(boolean configured, Set<String> off) {
public ModelGateState {
off = Set.copyOf(off);
}
public static ModelGateState notConfigured() {
return new ModelGateState(false, Set.of());
}
public static ModelGateState armed(Set<String> off) {
return new ModelGateState(true, off);
}
}
}
@@ -5,12 +5,14 @@ import java.util.List;
/**
* Backward-compatible placement: an unqualified spawn always resolves to the configured default
* profile, exactly as {@code CompositePeerLauncher} did before CB-518. Reachability is a narrower
* exception (fleetd #315, below): a profile is never checked for reachability up front, only
* skipped once it has already failed in <em>this same</em> spawn call's retry loop — see the
* unreachable case below.
* profile, exactly as {@code CompositePeerLauncher} did before CB-518. This ignores caps
* ({@code maxLoad}) so that a pre-existing config behaves identically after upgrade — capacity
* gating for automatic placement is deliberately out of scope for {@code fixed}, exactly as it
* always has been. Reachability is a narrower exception (fleetd #315, below): a profile is never
* checked for reachability up front, only skipped once it has already failed in <em>this same</em>
* spawn call's retry loop — see the unreachable case below.
*
* <p>Six exceptions walk past the default instead of returning it unconditionally:
* <p>Four exceptions walk past the default instead of returning it unconditionally:
* <ul>
* <li>Quarantine (CB-578 stage B): a quarantined default is a credential that just refused on
* a usage limit, not a transient capacity or reachability concern.
@@ -18,24 +20,6 @@ import java.util.List;
* ({@code BackendOutagePolicy}) — a separate, shorter-lived source from quarantine. When a
* profile is both quarantined and cooling off, only the quarantine reason is reported
* (exhaustion takes priority), matching {@code CompositePeerLauncher}'s explicit-spawn order.
* <li>At cap (fleetd #435): a profile whose live count has reached its {@code maxLoad}
* ({@link PlacementPolicyUtil#atCap}) — a documented, unconditional capacity limit (see
* {@code FleetConfig.Profile#maxLoad}), so {@code fixed} must gate on it exactly as {@code
* weighted}/{@code round-robin} already do via {@link PlacementPolicyUtil#available}. Before
* this fix {@code fixed} built its own {@link PlacementCandidate} for the default with {@code
* maxLoad} forced to {@code null}, so a capped default was chosen anyway on every unqualified
* spawn — the cap was advisory, not enforced, for the one placement policy every config uses
* by default. Reported only when quarantine and cooling off are both absent, matching {@code
* CompositePeerLauncher}'s explicit-spawn check order (quarantine, then cooling off, then max
* load, then model-off).
* <li>Model off (fleetd #422): a profile whose {@code model:} the operator has turned off in
* {@code models.allow:} — an operator decision, never a backend-reported outage, so it is a
* fifth, independent source from quarantine, cooling off, and at-cap (never merged with any
* of them), exactly as {@code CompositePeerLauncher.enforceModelEnabled} and {@link
* PlacementPolicyUtil#available} treat it. When a profile is model-off <em>and</em> quarantined,
* cooling off, or at cap, only the higher-priority reason is reported, matching {@code
* CompositePeerLauncher}'s explicit-spawn check order (quarantine, then cooling off, then max
* load, then model-off).
* <li>Unreachable (fleetd #315): {@code CompositePeerLauncher.spawn} retries a failed candidate
* on the next one and rebuilds the {@link PlacementContext} so {@code ctx.unreachable()}
* names every profile that already failed with {@code PeerUnreachableException} in this same
@@ -48,10 +32,9 @@ import java.util.List;
* {@code weighted}/{@code round-robin} skip it — an explicit {@code fleet_spawn} naming
* the profile is unaffected, only this automatic fallback walk.
* </ul>
* A fleet where nothing is ever quarantined, cooling off, at cap, model-off, unreachable, or
* weight-0 never exercises any of these paths, so today's behaviour is unchanged — in particular,
* the very first selection of a spawn call always sees an empty {@code unreachable} set, so the
* first choice is untouched.
* A fleet where nothing is ever quarantined, cooling off, unreachable, or weight-0 never exercises
* any of these paths, so today's behaviour is unchanged — in particular, the very first selection
* of a spawn call always sees an empty {@code unreachable} set, so the first choice is untouched.
*/
final class FixedPlacementPolicy implements PlacementPolicy {
@@ -59,14 +42,12 @@ final class FixedPlacementPolicy implements PlacementPolicy {
public PlacementCandidate select(PlacementContext ctx) {
String d = ctx.defaultProfile();
if (d != null && !d.isBlank() && !ctx.quarantined().contains(d) && !ctx.coolingOff().contains(d)
&& !ctx.modelOff().contains(d) && !ctx.unreachable().contains(d) && !weightExcluded(ctx, d)
&& !capExcluded(ctx, d)) {
&& !ctx.unreachable().contains(d) && !weightExcluded(ctx, d)) {
return new PlacementCandidate(d, null, 1.0f, null);
}
for (PlacementCandidate c : ctx.candidates()) {
if (!ctx.quarantined().contains(c.profile()) && !ctx.coolingOff().contains(c.profile())
&& !ctx.modelOff().contains(c.profile()) && !ctx.unreachable().contains(c.profile())
&& !c.excluded() && !PlacementPolicyUtil.atCap(ctx, c)) {
&& !ctx.unreachable().contains(c.profile()) && !c.excluded()) {
return new PlacementCandidate(c.profile(), null, c.weight(), c.maxLoad());
}
}
@@ -75,18 +56,9 @@ final class FixedPlacementPolicy implements PlacementPolicy {
// Exhaustion quarantine takes priority: reported only when quarantine is absent, so the
// message never claims "cooling off" for a profile that is really backend-exhausted.
boolean dCoolingOff = !dQuarantined && ctx.coolingOff().contains(d);
// fleetd #435: at-cap sits between cooling off and model-off, matching
// CompositePeerLauncher's explicit-spawn check order (quarantine, cooling off, max load,
// then model-off) — reported only when quarantine/cooling-off are both absent.
boolean dAtCap = !dQuarantined && !dCoolingOff && capExcluded(ctx, d);
// fleetd #422: model-off is a fifth, independent source (an operator decision) — but
// quarantine/cooling-off/at-cap still take priority when more than one applies, matching
// CompositePeerLauncher's explicit-spawn check order (quarantine, cooling off, max load,
// then model-off).
boolean dModelOff = !dQuarantined && !dCoolingOff && !dAtCap && ctx.modelOff().contains(d);
boolean dUnreachable = ctx.unreachable().contains(d);
boolean dWeightExcluded = weightExcluded(ctx, d);
if (dQuarantined || dCoolingOff || dAtCap || dModelOff || dUnreachable || dWeightExcluded) {
if (dQuarantined || dCoolingOff || dUnreachable || dWeightExcluded) {
List<String> reasons = new ArrayList<>();
if (dQuarantined) {
reasons.add("is quarantined (backend exhausted)");
@@ -94,14 +66,6 @@ final class FixedPlacementPolicy implements PlacementPolicy {
if (dCoolingOff) {
reasons.add("is cooling off after repeated backend errors");
}
if (dAtCap) {
PlacementCandidate c = candidateFor(ctx, d);
int live = ctx.liveCount().apply(d);
reasons.add("is at maxLoad (" + live + " live >= " + c.maxLoad() + " cap)");
}
if (dModelOff) {
reasons.add("names a model the operator has turned off in models.allow");
}
if (dUnreachable) {
reasons.add("is unreachable");
}
@@ -114,38 +78,18 @@ final class FixedPlacementPolicy implements PlacementPolicy {
}
if (!ctx.candidates().isEmpty()) {
throw new PlacementException("all worker profiles are excluded from automatic "
+ "selection (quarantined, cooling off, at cap, model-off, unreachable, or weight-0)");
+ "selection (quarantined, cooling off, unreachable, or weight-0)");
}
throw new PlacementException("no worker profiles configured");
}
/** Whether {@code profile} carries {@code weight <= 0} (CB-554) among {@code ctx}'s candidates. */
private static boolean weightExcluded(PlacementContext ctx, String profile) {
PlacementCandidate c = candidateFor(ctx, profile);
return c != null && c.excluded();
}
/**
* Whether {@code profile} has reached its {@code maxLoad} cap (fleetd #435), using the shared
* {@link PlacementPolicyUtil#atCap} definition — the same one {@code weighted}/{@code
* round-robin} already consult via {@link PlacementPolicyUtil#available}. Looked up by name,
* the same way {@link #weightExcluded} is: the default fast path above builds its own {@link
* PlacementCandidate} with {@code maxLoad} forced to {@code null} (it carries no cap of its
* own), so the candidate actually configured for {@code profile} has to be found in {@code
* ctx.candidates()} first.
*/
private static boolean capExcluded(PlacementContext ctx, String profile) {
PlacementCandidate c = candidateFor(ctx, profile);
return c != null && PlacementPolicyUtil.atCap(ctx, c);
}
/** The configured candidate named {@code profile} in {@code ctx}, or {@code null} if none. */
private static PlacementCandidate candidateFor(PlacementContext ctx, String profile) {
for (PlacementCandidate c : ctx.candidates()) {
if (c.profile().equals(profile)) {
return c;
return c.excluded();
}
}
return null;
return false;
}
}
@@ -22,29 +22,11 @@ import java.util.function.Function;
* the two apart so its refusal message says "cooling off", not "exhausted",
* when only this one is active. A profile can be in both sets at once; when it
* is, exhaustion quarantine is reported (it takes priority).
* @param modelOff profiles whose {@code model:} is currently turned off in {@code
* models.allow:} (fleetd #422) — an operator decision, not a backend-reported
* outage, so a SEPARATE, independent source from both {@code quarantined} and
* {@code coolingOff}. A profile can be in this set together with either (or
* both) of the others; {@link PlacementPolicyUtil} counts it into its own
* bucket rather than merging it into theirs, the same reason
* {@code coolingOff} is kept apart from {@code quarantined}.
*/
public record PlacementContext(String defaultProfile,
List<PlacementCandidate> candidates,
Function<String, Integer> liveCount,
Set<String> unreachable,
Set<String> quarantined,
Set<String> coolingOff,
Set<String> modelOff) {
/**
* Back-compat form before the fleetd #422 model on/off gate was added — no candidate's model
* is off. Keeps pre-#422 call sites (tests included) compiling and behaving identically.
*/
public PlacementContext(String defaultProfile, List<PlacementCandidate> candidates,
Function<String, Integer> liveCount, Set<String> unreachable,
Set<String> quarantined, Set<String> coolingOff) {
this(defaultProfile, candidates, liveCount, unreachable, quarantined, coolingOff, Set.of());
}
Set<String> coolingOff) {
}
@@ -11,40 +11,28 @@ final class PlacementPolicyUtil {
private PlacementPolicyUtil() {
}
/**
* True when {@code c} has reached its {@code maxLoad} cap: {@code liveCount(c.profile()) >=
* c.maxLoad()}. A {@code null} maxLoad means unlimited, so it is never at cap.
*
* <p>Extracted as the single shared definition of "at cap" (fleetd #435): before this fix it
* was computed inline in both {@link #available} and {@link #emptyException}, and {@code
* FixedPlacementPolicy} — not a caller of either — quietly kept its own {@code select} free of
* any cap check at all, so a capped default profile was chosen anyway under the default
* placement policy. Every automatic policy must call this, not re-derive it.
*/
static boolean atCap(PlacementContext ctx, PlacementCandidate c) {
Integer cap = c.maxLoad();
return cap != null && ctx.liveCount().apply(c.profile()) >= cap;
}
/**
* Candidates that are not weight-excluded (CB-554: explicit {@code weight <= 0}, checked
* first because it is a static config choice rather than transient state), not
* known-unreachable, not quarantined (CB-578 stage B), not cooling off after repeated backend
* errors (fleetd #201 Unit 5 — a separate, shorter-lived source from quarantine), not naming a
* model the operator has turned off (fleetd #422 — a third, independent source: an operator
* decision, never a backend-reported outage), and have not reached their maxLoad (see {@link
* #atCap}). A {@code null} maxLoad means unlimited.
* errors (fleetd #201 Unit 5 — a separate, shorter-lived source from quarantine), and have not
* reached their maxLoad. A {@code null} maxLoad means unlimited.
*/
static List<PlacementCandidate> available(PlacementContext ctx) {
List<PlacementCandidate> out = new ArrayList<>();
for (PlacementCandidate c : ctx.candidates()) {
if (c.excluded() || ctx.unreachable().contains(c.profile())
|| ctx.quarantined().contains(c.profile())
|| ctx.coolingOff().contains(c.profile())
|| ctx.modelOff().contains(c.profile())
|| atCap(ctx, c)) {
|| ctx.coolingOff().contains(c.profile())) {
continue;
}
Integer cap = c.maxLoad();
if (cap != null) {
int live = ctx.liveCount().apply(c.profile());
if (live >= cap) {
continue;
}
}
out.add(c);
}
return out;
@@ -52,13 +40,12 @@ final class PlacementPolicyUtil {
/**
* Build a clear exception describing why every candidate was dropped: all weight-0, all
* quarantined, all cooling off, all model-off, all at capacity, all unreachable, or a mix. Each
* candidate is counted into exactly one bucket (weight-excluded first, then quarantined, then
* cooling off, then model-off) so a candidate excluded for more than one reason is never
* double-counted — a candidate that is both quarantined (CB-578 stage B, backend exhausted) and
* cooling off (fleetd #201 Unit 5, repeated backend errors) counts only as quarantined, matching
* {@code CompositePeerLauncher}'s explicit-spawn ordering: exhaustion quarantine takes priority
* when more than one applies.
* quarantined, all cooling off, all at capacity, all unreachable, or a mix. Each candidate is
* counted into exactly one bucket (weight-excluded first, then quarantined, then cooling off)
* so a candidate excluded for more than one reason is never double-counted — a candidate that is
* both quarantined (CB-578 stage B, backend exhausted) and cooling off (fleetd #201 Unit 5,
* repeated backend errors) counts only as quarantined, matching {@code CompositePeerLauncher}'s
* explicit-spawn ordering: exhaustion quarantine takes priority when both are active.
*/
static PlacementException emptyException(PlacementContext ctx) {
int weightExcluded = 0;
@@ -66,19 +53,17 @@ final class PlacementPolicyUtil {
int unreachable = 0;
int quarantined = 0;
int coolingOff = 0;
int modelOff = 0;
for (PlacementCandidate c : ctx.candidates()) {
Integer cap = c.maxLoad();
if (c.excluded()) {
weightExcluded++;
} else if (ctx.quarantined().contains(c.profile())) {
quarantined++;
} else if (ctx.coolingOff().contains(c.profile())) {
coolingOff++;
} else if (ctx.modelOff().contains(c.profile())) {
modelOff++;
} else if (ctx.unreachable().contains(c.profile())) {
unreachable++;
} else if (atCap(ctx, c)) {
} else if (cap != null && ctx.liveCount().apply(c.profile()) >= cap) {
atCap++;
}
}
@@ -98,10 +83,6 @@ final class PlacementPolicyUtil {
return new PlacementException(
"all worker profiles are cooling off after repeated backend errors");
}
if (modelOff == total) {
return new PlacementException(
"all worker profiles name a model the operator has turned off");
}
if (atCap == total) {
return new PlacementException("all worker profiles are at maxLoad");
}
@@ -111,9 +92,8 @@ final class PlacementPolicyUtil {
return new PlacementException("no worker profile available: " + atCap + " at maxLoad, "
+ unreachable + " unreachable, " + quarantined + " quarantined, "
+ coolingOff + " cooling off, "
+ modelOff + " model-off, "
+ weightExcluded + " weight-0, "
+ (total - atCap - unreachable - quarantined - coolingOff - modelOff - weightExcluded)
+ (total - atCap - unreachable - quarantined - coolingOff - weightExcluded)
+ " remaining");
}
}
@@ -1,150 +0,0 @@
package dev.ltms.fleet;
import ch.qos.logback.classic.Level;
import ch.qos.logback.classic.Logger;
import ch.qos.logback.classic.spi.ILoggingEvent;
import ch.qos.logback.core.read.ListAppender;
import dev.ltms.fleet.config.FleetConfig;
import org.junit.jupiter.api.Test;
import org.junit.jupiter.api.io.TempDir;
import org.slf4j.LoggerFactory;
import java.nio.file.Files;
import java.nio.file.Path;
import java.util.List;
import static org.junit.jupiter.api.Assertions.assertFalse;
import static org.junit.jupiter.api.Assertions.assertTrue;
/**
* fleetd #395: {@code exhaustedPattern} (see {@link FleetConfig.Profile#exhaustedPattern}) is
* deliberately opt-in — an unset one leaves usage-limit detection silently OFF for that profile,
* and nothing quarantines its credential. {@link Fleetd#reportExhaustedPatternGap} must say so at
* startup, naming every unarmed profile, and must never fire when every profile is armed. Mirrors
* {@link MemberCredentialsGapReportTest}'s pattern, capturing the real log via a
* {@link ListAppender}.
*
* <p>The 8-profile shape in {@link #theLiveEightProfileShapeWarnsExactlyTheSixUnarmedProfiles} is
* the live {@code fleetd.yaml} shape measured 2026-09-10 (fleetd #395's own ticket): 6 of 8
* profiles unarmed, 2 of those 6 ({@code opus}, {@code sonnet}) running on the operator's Claude
* subscription. {@code fleetd.yaml} itself is gitignored and unavailable to this test, so the
* shape is reproduced as a throwaway config in a {@code @TempDir} rather than read off disk.
*/
class ExhaustedPatternGapReportTest {
private static FleetConfig load(Path dir, String yaml) throws Exception {
Path f = dir.resolve("fleetd.yaml");
Files.writeString(f, yaml);
return FleetConfig.load(f);
}
private static ListAppender<ILoggingEvent> attach() {
Logger logger = (Logger) LoggerFactory.getLogger(Fleetd.class);
ListAppender<ILoggingEvent> appender = new ListAppender<>();
appender.start();
logger.addAppender(appender);
return appender;
}
private static void detach(ListAppender<ILoggingEvent> appender) {
((Logger) LoggerFactory.getLogger(Fleetd.class)).detachAppender(appender);
}
/** Every profile name mentioned by a WARN-level log line, across every WARN this call produced. */
private static List<String> warnMessages(ListAppender<ILoggingEvent> appender) {
return appender.list.stream()
.filter(e -> e.getLevel() == Level.WARN)
.map(ILoggingEvent::getFormattedMessage)
.toList();
}
@Test
void theLiveEightProfileShapeWarnsExactlyTheSixUnarmedProfiles(@TempDir Path dir) throws Exception {
// Reproduces the live shape measured 2026-09-10: 8 profiles, 2 armed (sol, terra), 6
// unarmed (local, local-direct, gx, opus, sonnet, xf) — 2 of the unarmed 6 (opus, sonnet)
// are subscription: true.
FleetConfig cfg = load(dir, """
profiles:
local:
baseUrl: http://gx00.gw:8000
local-direct:
baseUrl: http://gx01.gw:8000
gx:
kind: opencode
baseUrl: https://llm.ltms.dev/v1
opus:
subscription: true
model: claude-opus-5
sonnet:
subscription: true
model: claude-sonnet-5
sol:
baseUrl: https://llm.ltms.dev/v1
exhaustedPattern: "The usage limit has been reached"
terra:
baseUrl: https://llm.ltms.dev/v1
exhaustedPattern: "The usage limit has been reached"
xf:
baseUrl: https://llm.ltms.dev/v1
""");
ListAppender<ILoggingEvent> appender = attach();
try {
Fleetd.reportExhaustedPatternGap(cfg);
} finally {
detach(appender);
}
List<String> warns = warnMessages(appender);
assertFalse(warns.isEmpty(), "6 of 8 profiles are unarmed — at least one WARN must fire");
// Exactly one WARN aggregates every unarmed profile, naming all 6 and none of the 2 armed.
String aggregate = warns.stream()
.filter(m -> m.contains("no exhaustedPattern configured"))
.findFirst()
.orElseThrow(() -> new AssertionError("expected an aggregate unarmed-profiles WARN: " + warns));
for (String unarmed : List.of("local", "local-direct", "gx", "opus", "sonnet", "xf")) {
assertTrue(aggregate.contains(unarmed), "aggregate WARN must name '" + unarmed + "': " + aggregate);
}
for (String armed : List.of("sol", "terra")) {
assertFalse(aggregate.contains(armed), "aggregate WARN must NOT name armed profile '" + armed + "': " + aggregate);
}
// A second, louder WARN calls out the subscription profiles specifically.
String subscriptionWarn = warns.stream()
.filter(m -> m.contains("metered Claude plan"))
.findFirst()
.orElseThrow(() -> new AssertionError("expected a subscription-specific WARN: " + warns));
assertTrue(subscriptionWarn.contains("opus"), subscriptionWarn);
assertTrue(subscriptionWarn.contains("sonnet"), subscriptionWarn);
assertFalse(subscriptionWarn.contains("local-direct"),
"the subscription WARN must not name a non-subscription profile: " + subscriptionWarn);
}
@Test
void everyProfileArmedProducesNoWarningAtAll(@TempDir Path dir) throws Exception {
FleetConfig cfg = load(dir, """
profiles:
sol:
baseUrl: https://llm.ltms.dev/v1
exhaustedPattern: "The usage limit has been reached"
terra:
baseUrl: https://llm.ltms.dev/v1
exhaustedPattern: "The usage limit has been reached"
opus:
subscription: true
model: claude-opus-5
exhaustedPattern: "5-hour limit reached"
""");
ListAppender<ILoggingEvent> appender = attach();
try {
Fleetd.reportExhaustedPatternGap(cfg);
} finally {
detach(appender);
}
assertTrue(warnMessages(appender).isEmpty(),
"every profile is armed — a checker that warns anyway always fires: " + warnMessages(appender));
}
}
@@ -1,155 +0,0 @@
package dev.ltms.fleet;
import dev.ltms.fleet.config.ConfigRef;
import dev.ltms.fleet.config.FleetConfig;
import dev.ltms.fleet.mcp.FleetMcp;
import org.junit.jupiter.api.DisplayName;
import org.junit.jupiter.api.Test;
import org.junit.jupiter.api.io.TempDir;
import java.nio.file.Files;
import java.nio.file.Path;
import static org.junit.jupiter.api.Assertions.assertFalse;
import static org.junit.jupiter.api.Assertions.assertTrue;
import static org.junit.jupiter.api.Assertions.assertEquals;
/**
* fleetd #416: {@code fleet_list}'s {@code CapacitySource.configuredProfiles} must enumerate the
* <em>startup</em> profile set, not the live, hot-reloaded one.
*
* <p>{@code profiles} as a whole is a {@code DEFERRED} key ({@link ConfigRef#DEFERRED_KEYS}):
* {@code HerdrPeerLauncher} takes {@code Map.copyOf(profiles)} once at construction, so a profile
* only added to the hot-reloaded map can never actually be spawned. Before this fix, {@code
* Fleetd.main} built {@code CapacitySource} with {@code () -> config.get().profiles().keySet()} —
* the live map — so {@code fleet_list} would report a freshly hot-reloaded profile as available
* ({@code free > 0}) while {@code fleet_spawn} on that same profile failed with
* {@code unknown worker profile}. Measured on another host: adding a throwaway profile and letting
* it hot-reload gave {@code fleet_list} -> {@code free: 3} and {@code fleet_spawn} ->
* {@code error: unknown worker profile}.
*
* <p>This test needs a reload, the same reason {@link FleetdExhaustionDetectionArmedWiringTest}
* does: at startup the two snapshots agree, so a test of only a newly started daemon would not
* detect a live {@code config.get()} lookup for the set.
*
* <p><b>maxLoad must stay hot.</b> It is read off {@code config.get()} exactly like
* {@code credentialId} ({@link ConfigRef} documents both as hot, "read live off the config
* supplier ... exactly like weight/maxLoad"), so a reload that only changes an existing profile's
* {@code maxLoad} — no add/remove — must still change what {@code fleet_list} reports without a
* restart. A fix that freezes the whole {@code CapacitySource} against {@code cfg} (rather than
* only its {@code configuredProfiles} set) would trade this bug for its mirror image and is pinned
* wrong by {@link #reloadedMaxLoadStillChangesWhatFleetListReports}.
*/
class FleetdCapacitySourceWiringTest {
private static final String STARTUP = """
bind:
host: 127.0.0.1
port: 8765
herdrSocket: ~/.config/herdr/herdr.sock
profiles:
terra:
baseUrl: http://gx00.gw:8000
model: terra
maxLoad: 3
guard:
offSubscriptionHosts:
- gx00.gw
""";
private static final String WITH_NEW_PROFILE = """
bind:
host: 127.0.0.1
port: 8765
herdrSocket: ~/.config/herdr/herdr.sock
profiles:
terra:
baseUrl: http://gx00.gw:8000
model: terra
maxLoad: 3
ghost404:
baseUrl: http://gx00.gw:8001
model: ghost404
maxLoad: 3
guard:
offSubscriptionHosts:
- gx00.gw
""";
private static final String WITH_CHANGED_MAX_LOAD = """
bind:
host: 127.0.0.1
port: 8765
herdrSocket: ~/.config/herdr/herdr.sock
profiles:
terra:
baseUrl: http://gx00.gw:8000
model: terra
maxLoad: 9
guard:
offSubscriptionHosts:
- gx00.gw
""";
@Test
@DisplayName("a profile present only in the live (hot-reloaded) config is NOT listed")
void liveOnlyProfileIsNotListed(@TempDir Path dir) throws Exception {
Path file = dir.resolve("fleetd.yaml");
Files.writeString(file, STARTUP);
FleetConfig cfg = FleetConfig.load(file);
ConfigRef config = new ConfigRef(file, cfg);
Files.writeString(file, WITH_NEW_PROFILE);
assertTrue(config.reload().applied());
// The live snapshot now has the new profile — proves the reload really happened and this
// test is not accidentally passing because nothing changed.
assertTrue(config.get().profiles().containsKey("ghost404"));
FleetMcp.CapacitySource source = Fleetd.capacitySource(config, cfg, _ -> 0);
assertFalse(source.configuredProfiles().get().contains("ghost404"),
"a profile added only to the hot-reloaded config must not be listed by fleet_list — "
+ "HerdrPeerLauncher never learns about it until a restart, so fleet_spawn on it "
+ "would fail with 'unknown worker profile' while fleet_list claimed it free");
}
@Test
@DisplayName("a profile present in the startup set IS listed")
void startupProfileIsListed(@TempDir Path dir) throws Exception {
Path file = dir.resolve("fleetd.yaml");
Files.writeString(file, STARTUP);
FleetConfig cfg = FleetConfig.load(file);
ConfigRef config = new ConfigRef(file, cfg);
FleetMcp.CapacitySource source = Fleetd.capacitySource(config, cfg, _ -> 0);
// fleetd #416, both-directions requirement: a test that only ever passes an empty/absent
// startup set (the case above) cannot tell a correct lookup from one that is permanently
// empty (e.g. a mutation replacing the supplier with Set::of). This is the direction that
// fails if the fix regresses to reporting nothing at all.
assertTrue(source.configuredProfiles().get().contains("terra"),
"a profile present in the startup snapshot must still be listed by fleet_list");
}
@Test
@DisplayName("a hot maxLoad edit still changes what fleet_list reports")
void reloadedMaxLoadStillChangesWhatFleetListReports(@TempDir Path dir) throws Exception {
Path file = dir.resolve("fleetd.yaml");
Files.writeString(file, STARTUP);
FleetConfig cfg = FleetConfig.load(file);
ConfigRef config = new ConfigRef(file, cfg);
FleetMcp.CapacitySource source = Fleetd.capacitySource(config, cfg, _ -> 0);
assertEquals(3, source.maxLoad().apply("terra"),
"sanity: maxLoad reads 3 from the startup config before any reload");
Files.writeString(file, WITH_CHANGED_MAX_LOAD);
assertTrue(config.reload().applied());
assertEquals(9, source.maxLoad().apply("terra"),
"maxLoad must stay hot — the SAME CapacitySource instance must reflect a reloaded "
+ "maxLoad without a restart, exactly like credentialId. Freezing the whole "
+ "CapacitySource against the startup snapshot (rather than only its "
+ "configuredProfiles set) would trade fleetd #416 for its mirror image.");
}
}
@@ -1,94 +0,0 @@
package dev.ltms.fleet;
import dev.ltms.fleet.config.ConfigRef;
import dev.ltms.fleet.config.FleetConfig;
import dev.ltms.fleet.mcp.FleetMcp;
import dev.ltms.fleet.placement.BackendQuarantine;
import org.junit.jupiter.api.DisplayName;
import org.junit.jupiter.api.Test;
import org.junit.jupiter.api.io.TempDir;
import java.nio.file.Files;
import java.nio.file.Path;
import java.util.Map;
import java.util.regex.Pattern;
import static org.junit.jupiter.api.Assertions.assertFalse;
import static org.junit.jupiter.api.Assertions.assertTrue;
/**
* fleetd #404: {@code exhaustionDetectionArmed} must describe the startup pattern map, not the
* reloaded config snapshot.
*
* <p>This test needs a reload. At startup the two snapshots agree, so a test of only a newly
* started daemon would not detect a live {@code config.get()} lookup in the report field.
*/
class FleetdExhaustionDetectionArmedWiringTest {
private static final String NO_PATTERN = """
bind:
host: 127.0.0.1
port: 8765
herdrSocket: ~/.config/herdr/herdr.sock
profiles:
terra:
baseUrl: http://gx00.gw:8000
model: terra
guard:
offSubscriptionHosts:
- gx00.gw
""";
private static final String WITH_PATTERN = """
bind:
host: 127.0.0.1
port: 8765
herdrSocket: ~/.config/herdr/herdr.sock
profiles:
terra:
baseUrl: http://gx00.gw:8000
model: terra
exhaustedPattern: "usage limit"
guard:
offSubscriptionHosts:
- gx00.gw
""";
@Test
@DisplayName("reloading an exhaustedPattern does not arm the startup detection source")
void reloadedPatternDoesNotChangeTheArmedFieldUntilRestart(@TempDir Path dir) throws Exception {
Path file = dir.resolve("fleetd.yaml");
Files.writeString(file, NO_PATTERN);
ConfigRef config = new ConfigRef(file, FleetConfig.load(file));
Files.writeString(file, WITH_PATTERN);
assertTrue(config.reload().applied());
assertTrue(config.get().profiles().get("terra").hasExhaustedPattern());
FleetMcp.QuarantineSource source = Fleetd.quarantineSource(config, BackendQuarantine.none(),
Map.of());
assertFalse(source.exhaustedPatternArmed().apply("terra"),
"exhaustionDetectionArmed must use the startup pattern map, not config.get()");
}
@Test
@DisplayName("a profile in the startup pattern map is reported as armed")
void aProfileInTheStartupMapIsArmed(@TempDir Path dir) throws Exception {
// fleetd #404, second direction. The test above only ever passes an EMPTY startup map, so
// it cannot tell a correct lookup from one that is permanently off. Measured: replacing the
// armed lambda with `profile -> false` left the whole suite green at 1475 tests. That
// mutation would make #395's visibility feature dead — an operator fixing a detection gap
// would be told the gap is still open after fixing it, forever. Both directions are needed:
// this test is the only thing that fails when the field stops reporting armed at all.
Path file = dir.resolve("fleetd.yaml");
Files.writeString(file, WITH_PATTERN);
ConfigRef config = new ConfigRef(file, FleetConfig.load(file));
FleetMcp.QuarantineSource source = Fleetd.quarantineSource(config, BackendQuarantine.none(),
Map.of("terra", Pattern.compile("usage limit")));
assertTrue(source.exhaustedPatternArmed().apply("terra"),
"a profile whose pattern was compiled at startup must report armed");
assertFalse(source.exhaustedPatternArmed().apply("sonnet"),
"a profile absent from the startup map must not report armed");
}
}
@@ -1,62 +0,0 @@
package dev.ltms.fleet;
import dev.ltms.fleet.peer.PeerLauncher;
import org.junit.jupiter.api.DisplayName;
import org.junit.jupiter.api.Test;
import java.util.Set;
import static org.junit.jupiter.api.Assertions.assertEquals;
import static org.junit.jupiter.api.Assertions.assertNotEquals;
/**
* fleetd #422 follow-up: {@code Fleetd.modelGateCoverageLine} is the startup-log counterpart of
* {@code exhaustedPatternCoverageLine}/{@code errorPatternCoverageLine} — see {@code
* FleetdPatternCoverageLineTest} for the identical shape this follows — except here there is no
* {@code UnsetMeaning} choice for a caller to get backwards: {@link
* PeerLauncher.ModelGateState#configured()} already states, unambiguously, whether an empty
* {@link PeerLauncher.ModelGateState#off()} means "no {@code models:} block to gate with at all"
* or "a block armed and currently reporting zero off". This class proves {@code
* modelGateCoverageLine} words those two states — plus the third, N off — distinctly, so a
* mutation that made it ignore {@code configured()} either way is caught here.
*/
class FleetdModelGateCoverageLineTest {
@Test
@DisplayName("no models: block reports not configured, distinct from armed-with-zero")
void noModelsBlockReportsNotConfigured() {
String line = Fleetd.modelGateCoverageLine(PeerLauncher.ModelGateState.notConfigured());
assertEquals("not configured (no models: block — nothing is gated, and nothing can be)", line);
}
@Test
@DisplayName("a models: block armed with nothing off reports armed, distinct from not configured")
void armedWithNothingOffReportsArmed() {
String line = Fleetd.modelGateCoverageLine(PeerLauncher.ModelGateState.armed(Set.of()));
assertEquals("armed (models: block present; 0 models currently turned off)", line);
}
@Test
@DisplayName("a models: block with N off names the off models")
void armedWithModelsOffNamesThem() {
String line = Fleetd.modelGateCoverageLine(
PeerLauncher.ModelGateState.armed(Set.of("deepseek-v4-flash", "claude-opus-9000")));
assertEquals("armed (2 model(s) turned off: [claude-opus-9000, deepseek-v4-flash])", line);
}
@Test
@DisplayName("the three states produce pairwise-distinct wording for the same empty-looking input")
void theThreeStatesProduceDistinctWording() {
String notConfigured = Fleetd.modelGateCoverageLine(PeerLauncher.ModelGateState.notConfigured());
String armedZero = Fleetd.modelGateCoverageLine(PeerLauncher.ModelGateState.armed(Set.of()));
String armedOne = Fleetd.modelGateCoverageLine(PeerLauncher.ModelGateState.armed(Set.of("x")));
// Pinned individually above; restated here so this test alone still catches a regression
// even if one of the three tests above were ever deleted — the exact FleetdPatternCoverageLineTest
// pattern, adapted from "two keys" to "three states of one gate".
assertNotEquals(notConfigured, armedZero,
"collapsing 'no models: block' into 'armed, zero off' is fleetd #422 follow-up's exact defect");
assertNotEquals(armedZero, armedOne);
assertNotEquals(notConfigured, armedOne);
}
}
@@ -1,67 +0,0 @@
package dev.ltms.fleet;
import org.junit.jupiter.api.DisplayName;
import org.junit.jupiter.api.Test;
import java.util.Set;
import static org.junit.jupiter.api.Assertions.assertEquals;
import static org.junit.jupiter.api.Assertions.assertNotEquals;
/**
* fleetd #415 (review follow-up): {@code CompletionResolverTest} proves {@code coverage()} words
* {@code UnsetMeaning.OFF} and {@code UnsetMeaning.BUILT_IN_DEFAULT} correctly — but every one of
* those tests supplies the meaning itself. That proves the enum's wording, never that {@code
* Fleetd} pairs the right meaning with the right pattern key. That pairing is #415's actual
* defect: {@code coverage()} had no way to know what unset meant for its key, so the fix moved
* the fact to the caller — and nothing yet proved the caller states it correctly.
*
* <p><b>Measured or it didn't happen:</b> swapping the two {@code UnsetMeaning} arguments at
* {@code Fleetd}'s two coverage call sites — giving {@code exhaustedPattern} the built-in-default
* wording and {@code errorPattern} the off wording, #415's exact defect with the keys exchanged —
* compiled with 0 errors and left all 1506 existing tests green. This class exists to turn that
* swap red.
*
* <p>It calls {@link Fleetd#exhaustedPatternCoverageLine} and {@link Fleetd#errorPatternCoverageLine}
* directly rather than reading {@code Fleetd.java} as source text (the shape {@code
* FleetdCompletionResolverWiringTest} uses for a different wiring gap): those two methods are the
* extracted call sites {@code main} actually invokes, following the same {@code static} factory +
* dedicated-test pattern as {@link Fleetd#capacitySource} and {@link Fleetd#worktreeBranchLookup}.
*/
class FleetdPatternCoverageLineTest {
private static final Set<String> PROFILES = Set.of("terra", "gx10");
@Test
@DisplayName("exhaustedPatternCoverageLine says off when no profile configures exhaustedPattern")
void exhaustedPatternCoverageLineSaysOffWhenNoProfileConfiguresIt() {
assertEquals("off (no profile has an exhaustedPattern configured; profiles: [gx10, terra])",
Fleetd.exhaustedPatternCoverageLine(PROFILES, Set.of()));
}
@Test
@DisplayName("errorPatternCoverageLine says built-in default when no profile configures errorPattern")
void errorPatternCoverageLineSaysBuiltInDefaultWhenNoProfileConfiguresIt() {
assertEquals("built-in default for all profiles (no profile customises errorPattern; "
+ "profiles: [gx10, terra])",
Fleetd.errorPatternCoverageLine(PROFILES, Set.of()));
}
@Test
@DisplayName("the two keys produce different wording for the identical empty-coverage input")
void theTwoKeysProduceDifferentWordingForTheSameEmptyInput() {
String exhaustedLine = Fleetd.exhaustedPatternCoverageLine(PROFILES, Set.of());
String errorLine = Fleetd.errorPatternCoverageLine(PROFILES, Set.of());
// Pinned individually above; restated here so this test alone still catches a swap even
// if one of the two tests above were ever deleted.
assertEquals("off (no profile has an exhaustedPattern configured; profiles: [gx10, terra])",
exhaustedLine);
assertEquals("built-in default for all profiles (no profile customises errorPattern; "
+ "profiles: [gx10, terra])", errorLine);
assertNotEquals(exhaustedLine, errorLine,
"swapping which UnsetMeaning pairs with which pattern key at Fleetd's call sites "
+ "must be caught here — that pairing, not coverage()'s own wording in isolation, "
+ "is fleetd #415's actual defect");
}
}
@@ -1,132 +0,0 @@
package dev.ltms.fleet;
import org.junit.jupiter.api.Test;
import org.junit.jupiter.api.io.TempDir;
import java.nio.file.Files;
import java.nio.file.Path;
import static org.junit.jupiter.api.Assertions.assertThrows;
import static org.junit.jupiter.api.Assertions.assertTrue;
/**
* Proves the one thing no test proved before this ticket's follow-up: that {@code Fleetd.main}
* ITSELF — not a copy of its logic, not the validator called directly — still refuses to start on
* a bad config. Mutation testing found that deleting {@code cfg.validateAll();} (née six
* individual {@code cfg.validateXxx();} calls) from {@code Fleetd.main} left the full 1478-test
* suite green; every existing test called a validator directly and none exercised {@code
* Fleetd.main} as the caller. See {@code FleetConfigValidateAllTest} for why the fix collapses
* those six calls into one reflective {@link FleetConfig#validateAll()}, and {@code
* ConfigRefTest} for the equivalent proof on the {@link
* dev.ltms.fleet.config.ConfigRef#reload()} path.
*
* <p>This is deliberately the real {@code static void main(String[] args)} — package-private, so
* only a test in this package can call it, which is exactly what makes this proof strong: it is
* not a helper extracted for testability, it is the literal method {@code java -jar fleetd.jar}
* invokes. Every fixture below is otherwise-valid and fails exactly one validator, and — because
* {@link FleetConfig#validateAll()} runs immediately after {@code SubscriptionGuard.
* assertPrimaryClean}, before {@code main} opens the herdr socket, binds Javalin, or touches
* anything else with a real side effect — calling {@code Fleetd.main} with one of these configs is
* safe: it is guaranteed to throw before reaching any of that, precisely because the config is
* deliberately invalid.
*/
class FleetdStartupValidationTest {
private static void assertMainRefuses(Path dir, String fileName, String yaml,
String mustContain) throws Exception {
Path f = dir.resolve(fileName);
Files.writeString(f, yaml);
IllegalStateException e = assertThrows(IllegalStateException.class,
() -> Fleetd.main(new String[]{f.toString()}),
fileName + ": Fleetd.main must refuse this config before doing anything else");
assertTrue(e.getMessage().contains(mustContain),
fileName + ": expected message to contain \"" + mustContain + "\" but was: "
+ e.getMessage());
}
@Test
void mainRefusesANonLoopbackBindWithoutTokenMode(@TempDir Path dir) throws Exception {
assertMainRefuses(dir, "auth-exposure.yaml", """
bind:
host: 0.0.0.0
port: 8765
""", "auth.mode: token");
}
@Test
void mainRefusesALeadTabPrefixCollision(@TempDir Path dir) throws Exception {
assertMainRefuses(dir, "lead-tab-prefixes.yaml", """
bind:
host: 127.0.0.1
port: 8765
fleet:
tabLabel: "lead: {role} {profile}"
leaders:
opus:
tab: "lead: opus"
""", "fleet.tabLabel");
}
@Test
void mainRefusesASubscriptionProfileThatReseatsAnthropicBaseUrl(@TempDir Path dir)
throws Exception {
assertMainRefuses(dir, "subscription-profiles.yaml", """
bind:
host: 127.0.0.1
port: 8765
profiles:
sonnet:
subscription: true
argv: ["ccs", "sonnet"]
env:
ANTHROPIC_BASE_URL: http://anything-not-on-the-allowlist
""", "ANTHROPIC_BASE_URL");
}
@Test
void mainRefusesAnUnknownCharterKey(@TempDir Path dir) throws Exception {
assertMainRefuses(dir, "charters.yaml", """
bind:
host: 127.0.0.1
port: 8765
fleet:
charters:
architetc: text
""", "architetc");
}
@Test
void mainRefusesAnArchitectSlotNamingAnUnconfiguredProfile(@TempDir Path dir) throws Exception {
assertMainRefuses(dir, "members.yaml", """
bind:
host: 127.0.0.1
port: 8765
profiles:
gx10:
baseUrl: http://gx10.gw:8000
fleet:
architects:
lead-designer:
profile: sonnet
""", "lead-designer");
}
@Test
void mainRefusesAProfileNamingAModelOutsideTheAllowList(@TempDir Path dir) throws Exception {
assertMainRefuses(dir, "models.yaml", """
bind:
host: 127.0.0.1
port: 8765
profiles:
sonnet:
baseUrl: http://gx10.gw:8000
model: claude-sonnet-5
rogue:
baseUrl: http://gx11.gw:8000
model: claude-opus-9000
models:
allow:
- model: claude-sonnet-5
""", "rogue");
}
}
@@ -1,360 +0,0 @@
package dev.ltms.fleet.auth;
import dev.ltms.fleet.config.ConfigRef;
import dev.ltms.fleet.config.FleetConfig;
import dev.ltms.fleet.herdr.FakeHerdr;
import dev.ltms.fleet.herdr.PaneLocator;
import dev.ltms.fleet.mcp.ConnectionIdentity;
import dev.ltms.fleet.peer.MemberRole;
import org.junit.jupiter.api.Test;
import org.junit.jupiter.api.io.TempDir;
import java.nio.file.Files;
import java.nio.file.Path;
import java.util.Map;
import static org.junit.jupiter.api.Assertions.*;
/**
* fleetd #424 — revoking (or granting) an architect slot must take effect on the next spawn with
* no restart. A session already bound to a slot keeps its <em>binding</em> (the {@code
* terminalToSlot} occupancy) even after that slot drops out of config, but NOT the ARCHITECT
* <em>privilege</em> the slot used to grant — that is revoked on the bound session's very next
* request. See {@link MemberRegistry}'s class doc for the exact rule: "config governs what a
* bound slot still grants, as well as what may be bound next."
*
* <p>Every test here drives a REAL {@link ConfigRef#reload()} against a {@code @TempDir} file and
* asserts {@link ConfigRef.Outcome#applied()}, rather than comparing two frozen
* {@code MemberRegistry} instances in memory — the defect this ticket fixes is specifically that
* {@link MemberRegistry} used to ignore a live reload, so a test that never reloads cannot tell the
* fixed registry from the broken one. {@code requireSlotFor} and {@code reserve} are pinned in
* separate tests, in both directions (removed and added), so a registry that simply refuses (or
* simply allows) everything cannot pass by accident — see {@link MemberRegistryTest} for the
* registry's other invariants (bind/unbind cardinality, thread-safety), which are unaffected by
* this ticket and still exercised against the frozen constructor.
*/
class MemberRegistryLiveTest {
private static String yaml(String fleetBlock) {
return """
bind:
host: 127.0.0.1
port: 8765
herdrSocket: ~/.config/herdr/herdr.sock
profiles:
sonnet:
baseUrl: http://gx00.gw:8000
model: sonnet
opus:
baseUrl: http://gx00.gw:8001
model: opus
guard:
offSubscriptionHosts:
- gx00.gw
""" + fleetBlock;
}
private static final String WITH_SONNET_SLOT = """
fleet:
architects:
designer:
profile: sonnet
""";
/** No architect pool at all — developers is unrelated dead data for this registry (#424 out of scope). */
private static final String WITHOUT_ARCHITECT_SLOTS = """
fleet:
developers:
dev1:
profile: sonnet
""";
/** Same slot name ({@code designer}) as {@link #WITH_SONNET_SLOT}, repointed to a different profile. */
private static final String WITH_OPUS_SLOT = """
fleet:
architects:
designer:
profile: opus
""";
/** Same profile ({@code sonnet}) as {@link #WITH_SONNET_SLOT}, but the pool key is renamed. */
private static final String WITH_RENAMED_SLOT = """
fleet:
architects:
architect-lead:
profile: sonnet
""";
private static ConfigRef refFor(Path f) {
return new ConfigRef(f, FleetConfig.load(f));
}
// ── requireSlotFor is live (criteria 1, 2, 4) ──────────────────────────────────────────────
@Test
void requireSlotForRefusesAProfileWhoseSlotWasRemovedByReload(@TempDir Path dir) throws Exception {
Path f = dir.resolve("fleetd.yaml");
Files.writeString(f, yaml(WITH_SONNET_SLOT));
ConfigRef ref = refFor(f);
MemberRegistry registry = MemberRegistry.live(() -> ref.get().fleet());
assertDoesNotThrow(() -> registry.requireSlotFor(MemberRole.ARCHITECT, "sonnet"),
"the slot is configured before the reload");
Files.writeString(f, yaml(WITHOUT_ARCHITECT_SLOTS));
ConfigRef.Outcome out = ref.reload();
assertTrue(out.applied(), "the reload must actually take effect: " + out.summary());
assertThrows(IllegalArgumentException.class,
() -> registry.requireSlotFor(MemberRole.ARCHITECT, "sonnet"),
"revoking the slot must refuse the NEXT spawn that names it");
}
@Test
void requireSlotForAllowsAProfileWhoseSlotWasAddedByReload(@TempDir Path dir) throws Exception {
Path f = dir.resolve("fleetd.yaml");
Files.writeString(f, yaml(WITHOUT_ARCHITECT_SLOTS));
ConfigRef ref = refFor(f);
MemberRegistry registry = MemberRegistry.live(() -> ref.get().fleet());
assertThrows(IllegalArgumentException.class,
() -> registry.requireSlotFor(MemberRole.ARCHITECT, "sonnet"),
"no architect slot is configured yet");
Files.writeString(f, yaml(WITH_SONNET_SLOT));
ConfigRef.Outcome out = ref.reload();
assertTrue(out.applied(), "the reload must actually take effect: " + out.summary());
assertDoesNotThrow(() -> registry.requireSlotFor(MemberRole.ARCHITECT, "sonnet"),
"a slot added by reload must be usable with no restart");
}
// ── reserve is live too — tested separately from requireSlotFor (criteria 1, 2, 4) ────────
@Test
void reserveRefusesAProfileWhoseSlotWasRemovedByReload(@TempDir Path dir) throws Exception {
Path f = dir.resolve("fleetd.yaml");
Files.writeString(f, yaml(WITH_SONNET_SLOT));
ConfigRef ref = refFor(f);
MemberRegistry registry = MemberRegistry.live(() -> ref.get().fleet());
MemberLifecycle.SlotReservation before = registry.reserve(MemberRole.ARCHITECT, "sonnet");
assertEquals("architect:designer", before.slot());
registry.release(before); // free it back up so the reload-side reserve below starts clean
Files.writeString(f, yaml(WITHOUT_ARCHITECT_SLOTS));
ConfigRef.Outcome out = ref.reload();
assertTrue(out.applied(), "the reload must actually take effect: " + out.summary());
assertThrows(IllegalArgumentException.class,
() -> registry.reserve(MemberRole.ARCHITECT, "sonnet"),
"revoking the slot must refuse the NEXT reservation for it");
}
@Test
void reserveAllowsAProfileWhoseSlotWasAddedByReload(@TempDir Path dir) throws Exception {
Path f = dir.resolve("fleetd.yaml");
Files.writeString(f, yaml(WITHOUT_ARCHITECT_SLOTS));
ConfigRef ref = refFor(f);
MemberRegistry registry = MemberRegistry.live(() -> ref.get().fleet());
assertThrows(IllegalArgumentException.class,
() -> registry.reserve(MemberRole.ARCHITECT, "sonnet"),
"no architect slot is configured yet");
Files.writeString(f, yaml(WITH_SONNET_SLOT));
ConfigRef.Outcome out = ref.reload();
assertTrue(out.applied(), "the reload must actually take effect: " + out.summary());
MemberLifecycle.SlotReservation after = registry.reserve(MemberRole.ARCHITECT, "sonnet");
assertEquals("architect:designer", after.slot(),
"a slot added by reload must be reservable with no restart");
}
// ── profileForSlot, isSlot and nameForSlot are live too (fleetd #431) ─────────────────────────
// #424 pinned roleForSlot against a live reload but left these three untested — proved by
// mutating each to read a snapshot flattened once at construction: the full suite stayed green
// for all three.
//
// The three differ in how much production behaviour depends on them, and the ticket first got
// this ranking wrong. nameForSlot is wired: CallerResolver passes members::nameForSlot, next to
// members::roleForSlot. isSlot is reached through bind, which calls it to refuse an unknown
// slot. profileForSlot has NO caller in src/main at all — grepped both the ".profileForSlot("
// and the "::profileForSlot" form — so there is no seam to drive its test through beyond the
// accessor itself, and its own javadoc ("what the spawn lifecycle reads") describes a caller
// that does not exist. These tests pin the accessors as they are; whether profileForSlot should
// be wired or deleted is a separate question.
@Test
void profileForSlotReflectsAProfileChangedByReload(@TempDir Path dir) throws Exception {
Path f = dir.resolve("fleetd.yaml");
Files.writeString(f, yaml(WITH_SONNET_SLOT));
ConfigRef ref = refFor(f);
MemberRegistry registry = MemberRegistry.live(() -> ref.get().fleet());
assertEquals("sonnet", registry.profileForSlot("architect:designer"),
"the slot's profile before the reload");
Files.writeString(f, yaml(WITH_OPUS_SLOT));
ConfigRef.Outcome out = ref.reload();
assertTrue(out.applied(), "the reload must actually take effect: " + out.summary());
assertEquals("opus", registry.profileForSlot("architect:designer"),
"repointing the slot to a different profile must take effect with no restart");
}
@Test
void isSlotStopsReportingASlotRemovedByReload(@TempDir Path dir) throws Exception {
Path f = dir.resolve("fleetd.yaml");
Files.writeString(f, yaml(WITH_SONNET_SLOT));
ConfigRef ref = refFor(f);
MemberRegistry registry = MemberRegistry.live(() -> ref.get().fleet());
assertTrue(registry.isSlot("architect:designer"), "the slot is configured before the reload");
Files.writeString(f, yaml(WITHOUT_ARCHITECT_SLOTS));
ConfigRef.Outcome out = ref.reload();
assertTrue(out.applied(), "the reload must actually take effect: " + out.summary());
assertFalse(registry.isSlot("architect:designer"),
"removing the slot from config must make isSlot say so on the very next call, "
+ "with no restart");
}
@Test
void isSlotStartsReportingASlotAddedByReload(@TempDir Path dir) throws Exception {
Path f = dir.resolve("fleetd.yaml");
Files.writeString(f, yaml(WITHOUT_ARCHITECT_SLOTS));
ConfigRef ref = refFor(f);
MemberRegistry registry = MemberRegistry.live(() -> ref.get().fleet());
assertFalse(registry.isSlot("architect:designer"), "no architect slot is configured yet");
Files.writeString(f, yaml(WITH_SONNET_SLOT));
ConfigRef.Outcome out = ref.reload();
assertTrue(out.applied(), "the reload must actually take effect: " + out.summary());
assertTrue(registry.isSlot("architect:designer"),
"a slot added by reload must be visible to isSlot with no restart");
}
@Test
void nameForSlotReflectsANameChangedByReload(@TempDir Path dir) throws Exception {
Path f = dir.resolve("fleetd.yaml");
Files.writeString(f, yaml(WITH_SONNET_SLOT));
ConfigRef ref = refFor(f);
MemberRegistry registry = MemberRegistry.live(() -> ref.get().fleet());
assertEquals("designer", registry.nameForSlot("architect:designer"),
"the configured name before the reload");
Files.writeString(f, yaml(WITH_RENAMED_SLOT));
ConfigRef.Outcome out = ref.reload();
assertTrue(out.applied(), "the reload must actually take effect: " + out.summary());
assertNull(registry.nameForSlot("architect:designer"),
"the old key no longer names a configured slot — it was renamed away by the reload");
assertEquals("architect-lead", registry.nameForSlot("architect:architect-lead"),
"the new name must be visible under its new qualified key with no restart");
}
// ── a bound architect is demoted, but the binding itself is not touched (fleetd #424) ───────
// The lead's corrected ruling: the PRIVILEGE a slot grants is revoked on the bound session's
// very next request, but the terminalToSlot BINDING itself is untouched by a reload — dropping
// it would double-book the slot key and break unbind's compare-safe contract. See the class
// doc's binding rule.
/** A caller identity resolving the one canned pane (terminal {@code term_a}) in {@link FakeHerdr}. */
private static ConnectionIdentity boundPaneIdentity() {
return new ConnectionIdentity(new PaneLocator(new FakeHerdr()), _ -> FakeHerdr.WORKER_PID);
}
@Test
void anArchitectAlreadyBoundToASlotIsDemotedByReload(@TempDir Path dir) throws Exception {
Path f = dir.resolve("fleetd.yaml");
Files.writeString(f, yaml(WITH_SONNET_SLOT));
ConfigRef ref = refFor(f);
MemberRegistry registry = MemberRegistry.live(() -> ref.get().fleet());
MemberLifecycle.SlotReservation reservation = registry.reserve(MemberRole.ARCHITECT, "sonnet");
assertTrue(registry.bind(reservation, "term_a"));
assertEquals("architect:designer", registry.slotForTerminal("term_a"));
// Drive the real caller path, not the roleForSlot seam directly: CallerResolver.resolve is
// what a live request actually goes through (CallerResolver.java:220), and a resolver that
// ignored roleForSlot entirely would still pass a test that only checked the seam.
CallerResolver resolver = CallerResolver.withLeadsAndMembers(
boundPaneIdentity(), false, null, Map::of, registry);
Principal before = resolver.resolve("127.0.0.1", 42, null);
assertEquals(Role.ARCHITECT, before.role(), "sanity check: the harness binds term_a as an architect");
assertEquals("designer", before.name());
Files.writeString(f, yaml(WITHOUT_ARCHITECT_SLOTS));
ConfigRef.Outcome out = ref.reload();
assertTrue(out.applied(), "the reload must actually take effect: " + out.summary());
Principal after = resolver.resolve("127.0.0.1", 42, null);
assertEquals(Role.WORKER, after.role(),
"removing the slot from config must demote the bound session to worker on its "
+ "NEXT request — this is the ticket's whole point");
assertEquals("term_a", after.terminal(), "same pane, same terminal — only the role changed");
}
@Test
void theOriginalBindingStillOccupiesTheRemovedSlotSoASecondTerminalCannotClaimIt(@TempDir Path dir)
throws Exception {
Path f = dir.resolve("fleetd.yaml");
Files.writeString(f, yaml(WITH_SONNET_SLOT));
ConfigRef ref = refFor(f);
MemberRegistry registry = MemberRegistry.live(() -> ref.get().fleet());
MemberLifecycle.SlotReservation reservation = registry.reserve(MemberRole.ARCHITECT, "sonnet");
assertTrue(registry.bind(reservation, "term_a"));
Files.writeString(f, yaml(WITHOUT_ARCHITECT_SLOTS));
ConfigRef.Outcome removed = ref.reload();
assertTrue(removed.applied(), "the reload must actually take effect: " + removed.summary());
// The binding survives the removal untouched.
assertEquals("architect:designer", registry.slotForTerminal("term_a"),
"a live binding must never be retroactively unbound by a config edit");
assertEquals(Map.of("term_a", "architect:designer"), registry.snapshot());
// Bring the slot back into config. If the binding had been silently dropped by the removal
// (rather than merely losing the privilege it grants), a second terminal could now claim
// the "freed" key — the exact double-booking the class doc's binding rule rules out.
Files.writeString(f, yaml(WITH_SONNET_SLOT));
ConfigRef.Outcome restored = ref.reload();
assertTrue(restored.applied(), "the reload must actually take effect: " + restored.summary());
assertFalse(registry.bind("architect:designer", "term_b"),
"the slot is still occupied by term_a — a second terminal must not bind to it");
assertThrows(IllegalArgumentException.class,
() -> registry.reserve(MemberRole.ARCHITECT, "sonnet"),
"the slot is still occupied by term_a — a fresh reservation must not find it free");
assertEquals("architect:designer", registry.slotForTerminal("term_a"),
"the original binding is unchanged throughout");
}
@Test
void unbindStillSucceedsForTheOriginalTerminalAfterItsSlotIsRemoved(@TempDir Path dir) throws Exception {
Path f = dir.resolve("fleetd.yaml");
Files.writeString(f, yaml(WITH_SONNET_SLOT));
ConfigRef ref = refFor(f);
MemberRegistry registry = MemberRegistry.live(() -> ref.get().fleet());
MemberLifecycle.SlotReservation reservation = registry.reserve(MemberRole.ARCHITECT, "sonnet");
assertTrue(registry.bind(reservation, "term_a"));
Files.writeString(f, yaml(WITHOUT_ARCHITECT_SLOTS));
ConfigRef.Outcome out = ref.reload();
assertTrue(out.applied(), "the reload must actually take effect: " + out.summary());
assertTrue(registry.unbind("architect:designer", "term_a"),
"unbind must still work for a slot that config has since removed, or a session "
+ "that outlives its slot's removal could never release it");
assertNull(registry.slotForTerminal("term_a"));
assertEquals(Map.of(), registry.snapshot());
}
}
@@ -79,14 +79,6 @@ class ConfigRefTopLevelCoverageTest {
* <li>{@code memberCredentials} — read live at {@code Fleetd.java:198, 205, 729}.</li>
* <li>{@code memberLoginShell} — read live at
* {@code HerdrPeerLauncher#configuredMemberLoginShell}.</li>
* <li>{@code models} (fleetd #422) — membership is re-validated in full against the fresh
* config on every {@code ConfigRef.reload()} (via {@code FleetConfig#validateAll()}), so
* a bad edit is refused, never cached stale; the on/off half is read live by {@code
* CompositePeerLauncher.enforceModelEnabled} and its candidate filter, and by {@code
* fleet_profiles}/{@code GET /profiles} through {@code PeerLauncher.disabledModels()}.
* Unlike {@code fleet} below, nothing about {@code models} is baked into a startup-built
* object anywhere — there is no frozen half, so it belongs here whole rather than in
* {@code SPLIT_KEYS}.</li>
* </ul>
*
* <p>{@code fleet} used to sit here too, on the strength of most of it (role pools, charters,
@@ -101,7 +93,7 @@ class ConfigRefTopLevelCoverageTest {
* while this test stayed green throughout.</p>
*/
private static final Set<String> HOT_EXCLUDED_TOP_LEVEL_KEYS =
Set.of("placement", "memberCredentials", "memberLoginShell", "models");
Set.of("placement", "memberCredentials", "memberLoginShell");
@Test
void everyTopLevelComponentIsAccountedForInExactlyOneClass() {
@@ -128,7 +120,7 @@ class ConfigRefTopLevelCoverageTest {
// The escape hatch is pinned. Growing it requires editing this line — a visible, deliberate
// diff, not a quiet one. See the field javadoc above for what "belongs here" actually means.
assertEquals(Set.of("placement", "memberCredentials", "memberLoginShell", "models"), hot,
assertEquals(Set.of("placement", "memberCredentials", "memberLoginShell"), hot,
"HOT_EXCLUDED_TOP_LEVEL_KEYS changed. A component belongs here ONLY if it is read "
+ "live off the config supplier, never because adding it makes this test "
+ "pass. If you are adding one to silence this test, that is fleetd #323 "
@@ -109,8 +109,6 @@ class ConfigRefTopLevelReportingCoverageTest {
v.put("memberLoginShell", null);
v.put("memberSkills", "/skills/a");
v.put("idleSleepGuard", new FleetConfig.IdleSleepGuard(true));
v.put("models", new FleetConfig.Models(
List.of(new FleetConfig.Models.ModelEntry("model-a"))));
assertNamesMatchComponents(v);
return v;
}
@@ -153,8 +151,6 @@ class ConfigRefTopLevelReportingCoverageTest {
v.put("memberLoginShell", null);
v.put("memberSkills", "/skills/b");
v.put("idleSleepGuard", new FleetConfig.IdleSleepGuard(false));
v.put("models", new FleetConfig.Models(
List.of(new FleetConfig.Models.ModelEntry("model-b"))));
assertNamesMatchComponents(v);
return v;
}
@@ -2683,337 +2683,4 @@ class FleetConfigTest {
assertTrue(cfg.idleSleepGuard().isEnabled(),
"unlike ConfigReload/Health, this block defaults to ON even when present but empty");
}
// --- models: central allow-list of usable models -----------------------------------------
/**
* An absent {@code models:} block is today's behaviour exactly: no profile's {@code model:} is
* checked against anything, whatever it says. This is the "existing config keeps working"
* invariant — an operator on a gitignored {@code fleetd.yaml} that predates this feature must
* not be broken by upgrading the daemon.
*/
@Test
void absentModelsBlockValidatesNothing(@TempDir Path dir) throws Exception {
Path f = dir.resolve("fleetd.yaml");
Files.writeString(f, """
bind:
port: 8080
profiles:
gx10:
baseUrl: http://gx10.gw:8000
model: totally-unheard-of-model-xyz
""");
FleetConfig cfg = FleetConfig.load(f);
assertDoesNotThrow(cfg::validateModels);
}
/**
* {@code models:} present but with an empty (or absent) {@code allow:} must behave exactly like
* the block being absent — an operator adding the block for the first time with nothing in it
* yet must not be surprised by every profile suddenly refusing to start.
*/
@Test
void modelsBlockPresentButEmptyValidatesNothing(@TempDir Path dir) throws Exception {
Path f = dir.resolve("fleetd.yaml");
Files.writeString(f, """
bind:
port: 8080
profiles:
gx10:
baseUrl: http://gx10.gw:8000
model: whatever-the-operator-typed
models:
allow: []
""");
FleetConfig cfg = FleetConfig.load(f);
assertDoesNotThrow(cfg::validateModels);
}
/**
* The core of the ticket: a profile naming a model outside the configured allow-list fails
* config load, naming both the model and the profile that wanted it.
*/
@Test
void aProfileNamingAModelOutsideTheAllowListRefusesToStart(@TempDir Path dir) throws Exception {
Path f = dir.resolve("fleetd.yaml");
Files.writeString(f, """
bind:
port: 8080
profiles:
sonnet:
baseUrl: http://gx10.gw:8000
model: claude-sonnet-5
rogue:
baseUrl: http://gx11.gw:8000
model: claude-opus-9000
models:
allow:
- model: claude-sonnet-5
- model: claude-haiku-5
""");
FleetConfig cfg = FleetConfig.load(f);
IllegalStateException e = assertThrows(IllegalStateException.class, cfg::validateModels);
assertTrue(e.getMessage().contains("rogue"), "the refusal names the offending profile");
assertTrue(e.getMessage().contains("claude-opus-9000"), "the refusal names the offending model");
assertFalse(e.getMessage().contains("'sonnet'"),
"the profile whose model IS allowed must not be reported");
}
/** A profile whose {@code model:} is on the allow-list loads fine. */
@Test
void aProfileNamingAModelOnTheAllowListLoadsFine(@TempDir Path dir) throws Exception {
Path f = dir.resolve("fleetd.yaml");
Files.writeString(f, """
bind:
port: 8080
profiles:
sonnet:
baseUrl: http://gx10.gw:8000
model: claude-sonnet-5
models:
allow:
- model: claude-sonnet-5
""");
FleetConfig cfg = FleetConfig.load(f);
assertDoesNotThrow(cfg::validateModels);
}
/**
* A profile that never sets {@code model:} (a {@code subscription: true} profile relying on the
* account's own default is the live example) must not be refused just because an allow-list is
* active — there is nothing to check it against.
*/
@Test
void aProfileWithNoModelPassesEvenWithAnActiveAllowList(@TempDir Path dir) throws Exception {
Path f = dir.resolve("fleetd.yaml");
Files.writeString(f, """
bind:
port: 8080
profiles:
opus:
subscription: true
models:
allow:
- model: claude-sonnet-5
""");
FleetConfig cfg = FleetConfig.load(f);
assertDoesNotThrow(cfg::validateModels);
}
/**
* Reproduces the live config shape this ticket measured: a mix of {@code amazon-bedrock},
* {@code opencode} and {@code openai}-backed profiles, some naming a bare Claude id and some a
* provider-prefixed opencode id, in ONE allow-list. Both forms are just opaque strings compared
* for exact equality — proves the shape decision holds against the shape actually seen live,
* not just against a synthetic single-provider example.
*/
@Test
void aBareClaudeIdAndAnOpencodeProviderPrefixedIdBothFitOneAllowList(@TempDir Path dir) throws Exception {
Path f = dir.resolve("fleetd.yaml");
Files.writeString(f, """
bind:
port: 8080
profiles:
sonnet:
baseUrl: http://gx10.gw:8000
model: claude-sonnet-5
terra:
kind: opencode
model: openai/gpt-5.6-terra
nova:
kind: opencode
model: amazon-bedrock/amazon.nova-pro-v1:0
models:
allow:
- model: claude-sonnet-5
- model: openai/gpt-5.6-terra
- model: amazon-bedrock/amazon.nova-pro-v1:0
""");
FleetConfig cfg = FleetConfig.load(f);
assertDoesNotThrow(cfg::validateModels);
}
/** The same live shape, but one opencode profile's model is missing from the allow-list. */
@Test
void anUnlistedOpencodeProviderPrefixedModelRefusesToStart(@TempDir Path dir) throws Exception {
Path f = dir.resolve("fleetd.yaml");
Files.writeString(f, """
bind:
port: 8080
profiles:
sonnet:
baseUrl: http://gx10.gw:8000
model: claude-sonnet-5
terra:
kind: opencode
model: openai/gpt-5.6-terra-withdrawn
models:
allow:
- model: claude-sonnet-5
- model: openai/gpt-5.6-terra
""");
FleetConfig cfg = FleetConfig.load(f);
IllegalStateException e = assertThrows(IllegalStateException.class, cfg::validateModels);
assertTrue(e.getMessage().contains("terra"), "the refusal names the offending profile");
assertTrue(e.getMessage().contains("openai/gpt-5.6-terra-withdrawn"),
"the refusal names the offending model, with its provider prefix intact");
}
/**
* Editing a {@code profiles:} entry alone must never be able to widen what is permitted — only
* editing {@code models.allow:} itself can. This is the invariant the ticket states explicitly;
* this test pins it by giving a profile a plausible-looking model that was never added to the
* allow-list and confirming it is still refused.
*/
@Test
void addingAProfileCannotWidenTheAllowListByItself(@TempDir Path dir) throws Exception {
Path f = dir.resolve("fleetd.yaml");
Files.writeString(f, """
bind:
port: 8080
profiles:
sonnet:
baseUrl: http://gx10.gw:8000
model: claude-sonnet-5
brand-new:
baseUrl: http://gx12.gw:8000
model: claude-sonnet-6-preview
models:
allow:
- model: claude-sonnet-5
""");
FleetConfig cfg = FleetConfig.load(f);
IllegalStateException e = assertThrows(IllegalStateException.class, cfg::validateModels);
assertTrue(e.getMessage().contains("brand-new"));
assertTrue(e.getMessage().contains("claude-sonnet-6-preview"));
}
/** {@code models} is a brand-new top-level key and must be recognized, not WARN-ed as unknown. */
@Test
void modelsIsAKnownTopLevelKey() {
assertTrue(FleetConfig.KNOWN_TOP_LEVEL_KEYS.contains("models"));
}
// --- fleetd #422: model on/off (spawn-time gate + hot reload) -----------------------------
/**
* Criterion 3: an {@code allow:} entry written before {@code enabled:} existed — no such key in
* the YAML at all — must behave exactly as it always did: on, and absent from {@link
* FleetConfig.Models#offIds()}. This is the old-style fixture the ticket asks for, proven
* through a real YAML load rather than only through the {@code ModelEntry(String)} back-compat
* constructor.
*/
@Test
void anOldStyleAllowEntryWithNoEnabledFieldStaysOn(@TempDir Path dir) throws Exception {
Path f = dir.resolve("fleetd.yaml");
Files.writeString(f, """
bind:
port: 8080
profiles:
sonnet:
baseUrl: http://gx10.gw:8000
model: claude-sonnet-5
models:
allow:
- model: claude-sonnet-5
""");
FleetConfig cfg = FleetConfig.load(f);
assertDoesNotThrow(cfg::validateModels);
assertTrue(cfg.models().ids().contains("claude-sonnet-5"), "membership is unaffected");
assertTrue(cfg.models().offIds().isEmpty(), "no enabled: field ⇒ nothing is off");
}
/**
* Criterion 7 (the whole point of the ticket, mirroring criterion 3's shape): a profile naming a
* model that is turned off ({@code enabled: false}) must still be VALID config —
* {@code validateModels()} checks membership only, never the on/off state, so it must not throw.
* Collapsing "off" into "removed from allow:" would make this test fail, which is exactly the
* design trap the ticket calls out.
*/
@Test
void anOffModelIsStillValidConfigForValidateModels(@TempDir Path dir) throws Exception {
Path f = dir.resolve("fleetd.yaml");
Files.writeString(f, """
bind:
port: 8080
profiles:
local:
baseUrl: http://local.gw:8000
model: deepseek-v4-flash
models:
allow:
- model: deepseek-v4-flash
enabled: false
""");
FleetConfig cfg = FleetConfig.load(f);
assertDoesNotThrow(cfg::validateModels,
"an off model must stay a valid allow-list member — only the spawn gate reads enabled");
assertTrue(cfg.models().ids().contains("deepseek-v4-flash"));
assertTrue(cfg.models().offIds().contains("deepseek-v4-flash"));
}
/**
* Criterion 8: one {@code enabled: false} entry names a model, not a profile — every profile
* naming that model is off, on purpose (the live example: {@code deepseek-v4-flash} backs both
* {@code local} and {@code local-direct}). Proven here at the {@code Models}/config-load level;
* {@code CompositePeerLauncherTest} proves the spawn-time consequence for both profiles.
*/
@Test
void turningOffOneModelIsIndependentOfHowManyProfilesNameIt(@TempDir Path dir) throws Exception {
Path f = dir.resolve("fleetd.yaml");
Files.writeString(f, """
bind:
port: 8080
profiles:
local:
baseUrl: http://local.gw:8000
model: deepseek-v4-flash
local-direct:
baseUrl: http://local.gw:8001
model: deepseek-v4-flash
models:
allow:
- model: deepseek-v4-flash
enabled: false
""");
FleetConfig cfg = FleetConfig.load(f);
assertDoesNotThrow(cfg::validateModels);
// One allow-list entry, one off id — the fan-out to both profiles happens at the reader
// (CompositePeerLauncher), not by duplicating the entry per profile.
assertEquals(Set.of("deepseek-v4-flash"), cfg.models().offIds());
assertEquals("deepseek-v4-flash", cfg.profiles().get("local").model());
assertEquals("deepseek-v4-flash", cfg.profiles().get("local-direct").model());
}
/** An {@code enabled: true} entry (explicit, not just absent) also stays on — not just null. */
@Test
void explicitlyEnabledTrueStaysOn(@TempDir Path dir) throws Exception {
Path f = dir.resolve("fleetd.yaml");
Files.writeString(f, """
bind:
port: 8080
profiles:
sonnet:
baseUrl: http://gx10.gw:8000
model: claude-sonnet-5
models:
allow:
- model: claude-sonnet-5
enabled: true
""");
FleetConfig cfg = FleetConfig.load(f);
assertTrue(cfg.models().offIds().isEmpty());
}
}
@@ -1,358 +0,0 @@
package dev.ltms.fleet.config;
import org.junit.jupiter.api.Test;
import org.junit.jupiter.api.io.TempDir;
import java.lang.reflect.Method;
import java.nio.file.Files;
import java.nio.file.Path;
import java.util.ArrayList;
import java.util.List;
import java.util.Set;
import java.util.TreeSet;
import static org.junit.jupiter.api.Assertions.assertDoesNotThrow;
import static org.junit.jupiter.api.Assertions.assertEquals;
import static org.junit.jupiter.api.Assertions.assertThrows;
import static org.junit.jupiter.api.Assertions.assertTrue;
/**
* The gap this class exists to close: mutation testing on the fleetd ticket "central allow-list
* of usable models" found that although {@link FleetConfig#validateModels()}'s own logic was well
* pinned, nothing proved either real caller ({@code Fleetd.main} and {@link ConfigRef#reload()})
* still invoked it — deleting the call site left the full suite green (1478/0/0/0). A follow-up
* measurement (same technique — remove one call site, run the suite, not read the code) found the
* SAME gap for all five of {@link FleetConfig}'s other validators at startup, and for four of the
* six inside {@link ConfigRef#reload()}. This is a class of gap, not one line's mistake: every one
* of those thirteen tests called the validator itself directly, never the real caller that was
* supposed to.
*
* <p>The fix replaces the six individual {@code cfg.validateXxx()} calls at each of the two real
* call sites with one {@link FleetConfig#validateAll()}, which reaches every validator by
* reflection rather than by a hand-maintained list of names. A hand-maintained list of six names
* would have exactly the defect it replaces: the seventh validator someone adds next month has no
* reason to be added to it, and nothing would say so. This class proves TWO separate claims, and
* keeps them separate on purpose:
*
* <ol>
* <li>{@link #theSweepMechanismIsGenericNotHardcodedToFleetConfigsSixNames()} and its neighbours
* prove the reflective sweep itself ({@link FleetConfig#invokeAllValidators}) is a general
* mechanism — it runs whatever public, no-arg, void {@code validateXxx()} methods a class
* happens to declare today, including a class with more of them than {@link FleetConfig}
* has right now. This is the proof that a future, real seventh validator on {@link
* FleetConfig} would be swept automatically, without needing to add a real (unwanted)
* seventh validator just to exercise the claim.</li>
* <li>{@link #validateAllReachesEveryOneOfTodaysSixValidators()} proves {@link
* FleetConfig#validateAll()} itself is wired to that same generic mechanism and genuinely
* reaches each of today's six real validators — reusing the exact minimal failing
* configurations {@code FleetConfigTest} already established for each one directly, so a
* single call to {@code validateAll()} is shown to reproduce every one of those six
* failures.</li>
* </ol>
*
* <p>Together with the direct-{@code Fleetd.main}-invocation tests in {@code
* FleetdStartupValidationTest} (which prove the real startup call site still calls {@code
* validateAll()}) and the {@code ConfigRefTest} reload tests (which prove the same for {@link
* ConfigRef#reload()}), removing {@code cfg.validateAll();} from either real call site now fails
* a test in this module.
*
* <p><b>What is NOT pinned, measured rather than assumed.</b> Reverting {@link
* FleetConfig#validateAll()} to a hardcoded list of today's six method calls leaves the whole
* suite green (measured at review: 1491 tests, 0 failures). Nothing ties {@code validateAll()} to
* the generic sweep — claim 1 proves {@link FleetConfig#invokeAllValidators} is generic, and claim
* 2 proves {@code validateAll()} reaches today's six, and a hardcoded list satisfies both. So the
* reflective sweep is a convenience, not the guarantee. The guarantee is {@link
* #fleetConfigDeclaresExactlyTheseSixValidatorsToday()}: it fails the moment a seventh validator
* is declared, which forces whoever adds it to look at this file.
*/
class FleetConfigValidateAllTest {
// ── Claim 1: the reflective sweep is a general mechanism, not six names in disguise ──────────
/**
* A throwaway fixture class, unrelated to {@link FleetConfig} in every way except shape: three
* public, no-arg, void methods named {@code validateXxx}. Proves the sweep works on ANY class
* with this shape, not on something special-cased to {@link FleetConfig}.
*/
static class ThreeValidators {
final List<String> ran = new ArrayList<>();
public void validateAlpha() {
ran.add("validateAlpha");
}
public void validateBeta() {
ran.add("validateBeta");
}
public void validateGamma() {
ran.add("validateGamma");
}
}
@Test
void theSweepMechanismIsGenericNotHardcodedToFleetConfigsSixNames() {
ThreeValidators target = new ThreeValidators();
FleetConfig.invokeAllValidators(target);
assertEquals(List.of("validateAlpha", "validateBeta", "validateGamma"), target.ran,
"every validateXxx() method on this unrelated class must run, in alphabetical "
+ "order — the sweep reads the class's own shape, not a name FleetConfig "
+ "happens to know about");
}
/**
* The core of the "self-maintaining" requirement: the exact same class shape as {@link
* ThreeValidators}, plus one more method — standing in for "a developer adds a validator next
* month". Nothing about the sweep changes to pick it up; the new method is invoked purely
* because it exists and matches the shape. This is what makes adding a seventh real validator
* to {@link FleetConfig} safe without touching {@link FleetConfig#validateAll()} or either
* call site — there is no "wire it in" step left to forget.
*/
static class FourValidators {
final List<String> ran = new ArrayList<>();
public void validateAlpha() {
ran.add("validateAlpha");
}
public void validateBeta() {
ran.add("validateBeta");
}
public void validateGamma() {
ran.add("validateGamma");
}
public void validateDelta() {
ran.add("validateDelta");
}
}
@Test
void addingAFourthValidatorMethodGetsSweptWithNoOtherChange() {
FourValidators target = new FourValidators();
FleetConfig.invokeAllValidators(target);
assertEquals(List.of("validateAlpha", "validateBeta", "validateDelta", "validateGamma"),
sorted(target.ran),
"the fourth method must be reached automatically — proving a class can grow the "
+ "set of things it validates with no change to the sweep itself");
}
private static List<String> sorted(List<String> in) {
List<String> copy = new ArrayList<>(in);
copy.sort(String::compareTo);
return copy;
}
@Test
void aFailingValidatorStopsTheSweepAndPropagatesUnchanged() {
class OneFails {
@SuppressWarnings("unused")
public void validateOk() {
// passes
}
public void validateBoom() {
throw new IllegalStateException("refusing to start: boom");
}
}
IllegalStateException e = assertThrows(IllegalStateException.class,
() -> FleetConfig.invokeAllValidators(new OneFails()));
assertEquals("refusing to start: boom", e.getMessage(),
"the real exception must propagate unchanged, not be wrapped or swallowed");
}
/**
* Every rule the sweep's filter applies, proven independently: only public, no-arg, void
* methods whose name starts with {@code "validate"} run, {@code validateAll} itself is
* excluded (so a class that declares one of its own — as {@link FleetConfig} does — cannot
* recurse into itself), and a same-shaped-but-wrongly-named or wrongly-shaped method never
* runs. A reader who "simplifies" the filter in {@link FleetConfig#invokeAllValidators} in a
* way that widens or narrows it breaks one of these.
*/
static class FilterEdgeCases {
final List<String> ran = new ArrayList<>();
public void validateReal() {
ran.add("validateReal");
}
/** Wrong name — must not run. */
public void checkSomething() {
ran.add("checkSomething");
}
/** Wrong shape — takes an argument. */
public void validateWithArg(String ignored) {
ran.add("validateWithArg");
}
/** Wrong shape — returns something. */
public boolean validateReturnsBoolean() {
ran.add("validateReturnsBoolean");
return true;
}
/** Excluded by name on purpose, so the sweep cannot call itself. */
public void validateAll() {
ran.add("validateAll");
}
}
@Test
void onlyPublicNoArgVoidValidateNamedMethodsRun() {
FilterEdgeCases target = new FilterEdgeCases();
FleetConfig.invokeAllValidators(target);
assertEquals(List.of("validateReal"), target.ran,
"checkSomething (wrong name), validateWithArg (wrong shape), "
+ "validateReturnsBoolean (wrong shape), and validateAll (excluded by "
+ "name) must all be skipped");
}
// ── Claim 2: FleetConfig.validateAll() is wired to that mechanism and reaches all six today ──
/**
* Reflectively enumerates {@link FleetConfig}'s own public, no-arg, void {@code validateXxx()}
* methods (excluding {@code validateAll} itself) — the exact same filter {@link
* FleetConfig#invokeAllValidators} applies. This is not the mechanism proof (that is claim 1,
* above, on an unrelated class) — it is a visible denominator: today there are six, named
* here, so a reader adding a seventh sees this assertion name the new count rather than a
* silent pass at the old one.
*/
@Test
void fleetConfigDeclaresExactlyTheseSixValidatorsToday() {
Set<String> names = new TreeSet<>();
for (Method m : FleetConfig.class.getMethods()) {
if (java.lang.reflect.Modifier.isPublic(m.getModifiers())
&& m.getParameterCount() == 0
&& m.getReturnType() == void.class
&& m.getName().startsWith("validate")
&& !m.getName().equals("validateAll")) {
names.add(m.getName());
}
}
assertEquals(new TreeSet<>(Set.of("validateAuthExposure", "validateLeadTabPrefixes",
"validateSubscriptionProfiles", "validateCharters", "validateMembers",
"validateModels")), names,
"FleetConfig's public validate*() methods changed. Do TWO things, in this "
+ "order. First confirm validateAll() still delegates to "
+ "invokeAllValidators(this) — a hardcoded list there passes every other "
+ "test in this class, so this assertion is the only place that will ever "
+ "make you check. Only then update the expected set to match.");
}
/** A minimal, otherwise-valid file — same shape FleetConfigTest and ConfigRefTest use. */
private static String minimalValidYaml() {
return """
bind:
host: 127.0.0.1
port: 8765
profiles:
sonnet:
baseUrl: http://gx00.gw:8000
model: sonnet
""";
}
@Test
void aFullyValidConfigPassesValidateAll(@TempDir Path dir) throws Exception {
Path f = dir.resolve("fleetd.yaml");
Files.writeString(f, minimalValidYaml());
assertDoesNotThrow(() -> FleetConfig.load(f).validateAll());
}
/**
* The heart of claim 2: for each of today's six real validators, a minimal file that fails
* ONLY that one — the exact fixtures {@code FleetConfigTest} uses to test each validator
* directly — must also fail through {@link FleetConfig#validateAll()}. If a future edit to
* {@code validateAll()} silently dropped one validator from the sweep (e.g. a typo'd name
* filter), exactly one of these six would start passing when it must not.
*/
@Test
void validateAllReachesEveryOneOfTodaysSixValidators(@TempDir Path dir) throws Exception {
// validateAuthExposure: a non-loopback bind without token mode.
assertValidateAllRefuses(dir, "auth-exposure.yaml", """
bind:
host: 0.0.0.0
port: 8765
""", "auth.mode: token");
// validateLeadTabPrefixes: a fleet-wide tabLabel that starts with a lead's own tabPrefix.
assertValidateAllRefuses(dir, "lead-tab-prefixes.yaml", """
bind:
host: 127.0.0.1
port: 8765
fleet:
tabLabel: "lead: {role} {profile}"
leaders:
opus:
tab: "lead: opus"
""", "fleet.tabLabel");
// validateSubscriptionProfiles: subscription: true with env: reseating ANTHROPIC_BASE_URL.
assertValidateAllRefuses(dir, "subscription-profiles.yaml", """
bind:
host: 127.0.0.1
port: 8765
profiles:
sonnet:
subscription: true
argv: ["ccs", "sonnet"]
env:
ANTHROPIC_BASE_URL: http://anything-not-on-the-allowlist
""", "ANTHROPIC_BASE_URL");
// validateCharters: a charter key that is not a role wire name.
assertValidateAllRefuses(dir, "charters.yaml", """
bind:
host: 127.0.0.1
port: 8765
fleet:
charters:
architetc: text
""", "architetc");
// validateMembers: an architect slot naming an unconfigured profile.
assertValidateAllRefuses(dir, "members.yaml", """
bind:
host: 127.0.0.1
port: 8765
profiles:
gx10:
baseUrl: http://gx10.gw:8000
fleet:
architects:
lead-designer:
profile: sonnet
""", "lead-designer");
// validateModels: a profile naming a model outside the configured allow-list.
assertValidateAllRefuses(dir, "models.yaml", """
bind:
host: 127.0.0.1
port: 8765
profiles:
sonnet:
baseUrl: http://gx10.gw:8000
model: claude-sonnet-5
rogue:
baseUrl: http://gx11.gw:8000
model: claude-opus-9000
models:
allow:
- model: claude-sonnet-5
""", "rogue");
}
private static void assertValidateAllRefuses(Path dir, String fileName, String yaml,
String mustContain) throws Exception {
Path f = dir.resolve(fileName);
Files.writeString(f, yaml);
FleetConfig cfg = FleetConfig.load(f);
IllegalStateException e = assertThrows(IllegalStateException.class, cfg::validateAll,
fileName + ": validateAll() must refuse this config");
assertTrue(e.getMessage().contains(mustContain),
fileName + ": expected message to contain \"" + mustContain + "\" but was: "
+ e.getMessage());
}
}
@@ -97,8 +97,6 @@ class FleetConfigWithDefaultsPreservesEveryComponentTest {
v.put("memberLoginShell", "/bin/zsh");
v.put("memberSkills", "/skills/guard");
v.put("idleSleepGuard", new FleetConfig.IdleSleepGuard(true));
v.put("models", new FleetConfig.Models(
List.of(new FleetConfig.Models.ModelEntry("model-guard"))));
assertNamesMatchComponents(v);
return v;
}
@@ -20,7 +20,6 @@ import java.util.regex.Pattern;
import static org.junit.jupiter.api.Assertions.assertEquals;
import static org.junit.jupiter.api.Assertions.assertFalse;
import static org.junit.jupiter.api.Assertions.assertNotEquals;
import static org.junit.jupiter.api.Assertions.assertTrue;
/** Unit behaviour of the CB-106 completion resolver in isolation from the injector. */
@@ -885,22 +884,19 @@ class CompletionResolverTest {
@Test
void coverageIsOffWhenNoProfileHasAPatternConfigured() {
assertEquals("off (no profile has an exhaustedPattern configured; profiles: [terra])",
CompletionResolver.coverage("exhaustedPattern", CompletionResolver.UnsetMeaning.OFF,
Set.of("terra"), Set.of()));
CompletionResolver.coverage("exhaustedPattern", Set.of("terra"), Set.of()));
}
@Test
void coverageIsFullWhenEveryProfileHasAPatternConfigured() {
assertEquals("full (all profiles configured: [gx10, terra])",
CompletionResolver.coverage("exhaustedPattern", CompletionResolver.UnsetMeaning.OFF,
Set.of("terra", "gx10"), Set.of("terra", "gx10")));
CompletionResolver.coverage("exhaustedPattern", Set.of("terra", "gx10"), Set.of("terra", "gx10")));
}
@Test
void coverageIsPartialAndNamesWhichProfilesAreConfigured() {
assertEquals("partial (configured: [terra]; not configured: [gx10])",
CompletionResolver.coverage("exhaustedPattern", CompletionResolver.UnsetMeaning.OFF,
Set.of("terra", "gx10"), Set.of("terra")));
CompletionResolver.coverage("exhaustedPattern", Set.of("terra", "gx10"), Set.of("terra")));
}
/**
@@ -914,42 +910,11 @@ class CompletionResolverTest {
*
* <p>Every earlier test here passed the exhaustion case only, so none of them could see it. This
* one pins that the message names the key the caller actually meant.
*
* <p>fleetd#415: the expected wording changed here too. {@code errorPattern} has a built-in
* fallback ({@link CompletionResolver#BACKEND_ERROR}), so an empty {@code configuredProfiles}
* for it is not "off" — see {@link #coverageDistinguishesOffFromBuiltInDefaultForTheSameEmptyInput}
* for the test built specifically to pin that distinction.
*/
@Test
void coverageNamesTheConfigKeyItsCallerMeansRatherThanAlwaysSayingExhaustedPattern() {
assertEquals("built-in default for all profiles (no profile customises errorPattern; "
+ "profiles: [gx10, terra])",
CompletionResolver.coverage("errorPattern", CompletionResolver.UnsetMeaning.BUILT_IN_DEFAULT,
Set.of("terra", "gx10"), Set.of()));
}
/**
* fleetd#415: {@code coverage()} measures pattern coverage (how many profiles set the key), but
* for {@code errorPattern} the empty case is not the feature-off state — a profile with no
* configured {@code errorPattern} still runs the classification against
* {@link CompletionResolver#BACKEND_ERROR}. For {@code exhaustedPattern} there is no fallback,
* so empty really is off. Same shape of input (empty {@code configuredProfiles}, one profile),
* different {@link CompletionResolver.UnsetMeaning} — the wording must differ, or this method is
* back to conflating pattern coverage with feature state for the one key where they disagree.
*/
@Test
void coverageDistinguishesOffFromBuiltInDefaultForTheSameEmptyInput() {
String exhaustedLine = CompletionResolver.coverage("exhaustedPattern",
CompletionResolver.UnsetMeaning.OFF, Set.of("gx10", "terra"), Set.of());
String errorLine = CompletionResolver.coverage("errorPattern",
CompletionResolver.UnsetMeaning.BUILT_IN_DEFAULT, Set.of("gx10", "terra"), Set.of());
assertEquals("off (no profile has an exhaustedPattern configured; profiles: [gx10, terra])",
exhaustedLine);
assertEquals("built-in default for all profiles (no profile customises errorPattern; "
+ "profiles: [gx10, terra])", errorLine);
assertNotEquals(exhaustedLine, errorLine,
"the same empty-coverage input must not read as the same feature state for both keys");
assertEquals("off (no profile has an errorPattern configured; profiles: [gx10, terra])",
CompletionResolver.coverage("errorPattern", Set.of("terra", "gx10"), Set.of()));
}
// --- fleetd#201 Unit 1: target-keyed backend-error pattern + typed sink ----------------------
@@ -3,7 +3,6 @@ package dev.ltms.fleet.mcp;
import dev.ltms.fleet.auth.CallerResolver;
import dev.ltms.fleet.auth.MemberRegistry;
import dev.ltms.fleet.auth.Principal;
import dev.ltms.fleet.auth.Role;
import dev.ltms.fleet.config.FleetConfig;
import dev.ltms.fleet.guard.SubscriptionGuard;
import dev.ltms.fleet.herdr.AgentControl;
@@ -29,7 +28,6 @@ import dev.ltms.fleet.placement.BackendQuarantine;
import dev.ltms.fleet.placement.PlacementPolicies;
import io.modelcontextprotocol.spec.McpSchema;
import dev.ltms.fleet.msg.InMemoryReplyInbox;
import dev.ltms.fleet.msg.ReplyInbox;
import org.junit.jupiter.api.BeforeEach;
import org.junit.jupiter.api.Test;
@@ -79,7 +77,7 @@ class FleetMcpTest {
Thread.sleep(5);
}
assertTrue(rendezvous.isWaiting(target), "send should be accepted for " + target);
FleetMcp.reply(messages, target, Role.WORKER, "received");
FleetMcp.reply(messages, target, "received");
assertEquals("received", textOf(send.get(6, TimeUnit.SECONDS)));
}
@@ -112,7 +110,7 @@ class FleetMcpTest {
// fleetd #365: a resolved live send must read distinctly from a merely-queued reply —
// see replyWithNoPendingSendIsQueuedNotError below for the other case.
McpSchema.CallToolResult reply = FleetMcp.reply(messages, "term_a", Role.WORKER, "LGTM");
McpSchema.CallToolResult reply = FleetMcp.reply(messages, "term_a", "LGTM");
assertEquals(MessageService.ReplyOutcome.RESOLVED_SEND.description(), textOf(reply));
McpSchema.CallToolResult res = send.get(6, TimeUnit.SECONDS);
@@ -138,7 +136,7 @@ class FleetMcpTest {
}
assertTrue(rendezvous.isWaiting("term_a"), "send should have opened its waiter");
McpSchema.CallToolResult reply = FleetMcp.reply(messages, "term_a", Role.WORKER, "async LGTM");
McpSchema.CallToolResult reply = FleetMcp.reply(messages, "term_a", "async LGTM");
assertEquals(MessageService.ReplyOutcome.RESOLVED_SEND.description(), textOf(reply));
// Poll until the async send completes and reports the reply.
@@ -185,7 +183,7 @@ class FleetMcpTest {
Thread.sleep(5);
}
assertTrue(rendezvous.isWaiting("term_a"));
FleetMcp.reply(messages, "term_a", Role.WORKER, "done");
FleetMcp.reply(messages, "term_a", "done");
assertEquals("done", textOf(answer.get(6, TimeUnit.SECONDS)));
McpSchema.CallToolResult done = FleetMcp.poll(messages, ticket, null);
@@ -217,7 +215,7 @@ class FleetMcpTest {
// fleet_poll{ticket} stuck PENDING forever and later force-failed with a false "session
// released before it replied" reason. This used to land in the inbox instead (see the old
// assertion this replaced: messages.drainReplies("term_a").getFirst()...) — that was the bug.
FleetMcp.reply(messages, "term_a", Role.WORKER, "finished after timeout");
FleetMcp.reply(messages, "term_a", "finished after timeout");
assertEquals("finished after timeout", textOf(FleetMcp.poll(messages, ticket, null)));
assertTrue(messages.drainReplies("term_a").isEmpty(),
"the reply completed its own ticket directly and never touched the inbox");
@@ -271,7 +269,7 @@ class FleetMcpTest {
}
assertEquals(MessageService.Phase.FAILED, second.phase());
FleetMcp.reply(messages, "term_a", Role.WORKER, "late reply");
FleetMcp.reply(messages, "term_a", "late reply");
assertEquals("late reply", messages.drainReplies("term_a").getFirst().content());
CompletableFuture<McpSchema.CallToolResult> answer = CompletableFuture.supplyAsync(
@@ -280,7 +278,7 @@ class FleetMcpTest {
while (!rendezvous.isWaiting("term_a") && System.currentTimeMillis() < deadline) {
Thread.sleep(5);
}
FleetMcp.reply(messages, "term_a", Role.WORKER, "done");
FleetMcp.reply(messages, "term_a", "done");
assertEquals("done", textOf(answer.get(6, TimeUnit.SECONDS)));
}
@@ -333,7 +331,7 @@ class FleetMcpTest {
void replyWithNoPendingSendIsQueuedNotError() {
// CB-307: a reply with no open send is now queued in the inbox, not an error.
// fleetd #365: it must also no longer claim "delivered" — nothing was waiting for it.
McpSchema.CallToolResult res = FleetMcp.reply(messages, "term_a", Role.WORKER, "orphan");
McpSchema.CallToolResult res = FleetMcp.reply(messages, "term_a", "orphan");
assertNotEquals(Boolean.TRUE, res.isError(), "a queued reply is not an error");
assertEquals(MessageService.ReplyOutcome.QUEUED.description(), textOf(res));
@@ -343,40 +341,6 @@ class FleetMcpTest {
assertEquals("orphan", drained.getFirst().content());
}
@Test
void replyFromLeadIsRefusedBeforeItCanPublishToTheWorkerInbox() {
ReplyInbox inboxThatRejectsPublishes = new ReplyInbox() {
@Override public void own(String target) { }
@Override public void release(String target) { }
@Override public void publish(String target, String msgId, String content) {
fail("a lead fleet_reply must not publish to the worker inbox");
}
@Override public List<InboxMessage> peek(String target) { return List.of(); }
@Override public void ack(String target, String msgId) { }
};
MessageService leadMessages = new MessageService(agents, new Injector(agents), new Rendezvous(),
inboxThatRejectsPublishes);
McpSchema.CallToolResult res = assertDoesNotThrow(
() -> FleetMcp.reply(leadMessages, "term_lead", Role.PRIMARY, "peer reply"));
assertTrue(res.isError());
assertEquals("fleet_reply has no route to a peer lead. Use fleet_send{coordId: ...} for a peer on another "
+ "daemon or fleet_send{sessionId: ...} for a peer on this host. fleet_reply resolves a member's "
+ "blocked fleet_send, and a peer's coord-id message is durable and non-blocking, so there is "
+ "nothing for it to resolve.",
textOf(res));
}
@Test
void replyFromUnidentifiedCallerKeepsItsOwnError() {
McpSchema.CallToolResult res = FleetMcp.reply(messages, null, Role.PRIMARY, "reply");
assertTrue(res.isError());
assertEquals("fleet_reply is for workers only — could not identify the calling worker from the connection",
textOf(res));
}
@Test
void replyWithBlankContentIsACleanToolErrorNotAnUncaughtException() {
// fleetd #302: MessageService.reply now REJECTS blank content by throwing. fleet_reply's
@@ -387,7 +351,7 @@ class FleetMcpTest {
// has always used isBlank for exactly this reason.
for (String blank : new String[] {null, "", " ", "\n\t"}) {
McpSchema.CallToolResult res = assertDoesNotThrow(
() -> FleetMcp.reply(messages, "term_a", Role.WORKER, blank),
() -> FleetMcp.reply(messages, "term_a", blank),
"blank content must be refused as a tool error, never thrown out of the handler");
assertEquals(Boolean.TRUE, res.isError(), "blank content is an error result");
assertTrue(textOf(res).contains("content is required"),
@@ -400,7 +364,7 @@ class FleetMcpTest {
@Test
void bridgePollWithTargetDrainsReplies() {
// A reply with no open send queues it in the inbox.
FleetMcp.reply(messages, "term_a", Role.WORKER, "queued-msg");
FleetMcp.reply(messages, "term_a", "queued-msg");
// fleet_poll with target drains the inbox.
McpSchema.CallToolResult res = FleetMcp.poll(messages, null, "term_a");
@@ -451,7 +415,7 @@ class FleetMcpTest {
Thread.sleep(5);
}
assertTrue(rendezvous.isWaiting("term_a"), "the answer should have reopened a waiter");
McpSchema.CallToolResult reply = FleetMcp.reply(messages, "term_a", Role.WORKER, "done");
McpSchema.CallToolResult reply = FleetMcp.reply(messages, "term_a", "done");
assertEquals(MessageService.ReplyOutcome.RESOLVED_SEND.description(), textOf(reply));
assertEquals("done", textOf(answer.get(6, TimeUnit.SECONDS)));
}
@@ -1290,13 +1254,13 @@ class FleetMcpTest {
@Test
void bridgeAckRemovesSpecificReply() {
// Queue a reply and capture its msgId.
FleetMcp.reply(messages, "term_a", Role.WORKER, "orphan");
FleetMcp.reply(messages, "term_a", "orphan");
var before = messages.drainReplies("term_a");
assertEquals(1, before.size(), "one reply in the inbox");
String msgId = before.getFirst().msgId();
// Publish the same reply again and ack it via fleet_ack surface.
FleetMcp.reply(messages, "term_a", Role.WORKER, "orphan-again");
FleetMcp.reply(messages, "term_a", "orphan-again");
var peeked = messages.drainReplies("term_a");
assertEquals(1, peeked.size(), "one fresh reply in the inbox");
@@ -1,62 +0,0 @@
package dev.ltms.fleet.mcp;
import dev.ltms.fleet.config.FleetConfig;
import dev.ltms.fleet.guard.SubscriptionGuard;
import dev.ltms.fleet.herdr.AgentControl;
import dev.ltms.fleet.herdr.FakeHerdr;
import dev.ltms.fleet.herdr.WorkspaceControl;
import dev.ltms.fleet.member.ClaudeCodeLauncher;
import dev.ltms.fleet.peer.PeerLauncher;
import io.modelcontextprotocol.spec.McpSchema;
import org.junit.jupiter.api.Test;
import java.util.LinkedHashMap;
import java.util.Map;
import java.util.Set;
import static org.junit.jupiter.api.Assertions.assertNotEquals;
import static org.junit.jupiter.api.Assertions.assertTrue;
/**
* fleetd #395: {@code fleet_profiles} must let an operator tell "this profile's usage-limit
* detection is armed" from "nothing can ever quarantine this profile" — see {@link
* FleetMcp.QuarantineSource#exhaustedPatternArmed()} and {@link
* FleetMcp#profilesView(PeerLauncher, FleetMcp.QuarantineSource, FleetMcp.OutageSource)}.
*/
class FleetProfilesArmedFieldTest {
private static PeerLauncher twoProfileLauncher(FakeHerdr h) {
FleetConfig.Profile armed = new FleetConfig.Profile(
"armed-profile", "http://gx00.gw:8000", "coder", null, "FLEETD_WORKER_TOKEN", null,
"tab", "fleetd-workers", "worker: {profile} #{n}", null, null, null);
FleetConfig.Profile unarmed = new FleetConfig.Profile(
"unarmed-profile", "http://gx01.gw:8000", "coder", null, "FLEETD_WORKER_TOKEN", null,
"tab", "fleetd-workers", "worker: {profile} #{n}", null, null, null);
Map<String, FleetConfig.Profile> profiles = new LinkedHashMap<>();
profiles.put(armed.profile(), armed);
profiles.put(unarmed.profile(), unarmed);
return new ClaudeCodeLauncher(new AgentControl(h), new WorkspaceControl(h),
new SubscriptionGuard(Set.of("gx00.gw", "gx01.gw")), profiles, armed.profile(), _ -> "tok");
}
private static String textOf(McpSchema.CallToolResult r) {
return ((McpSchema.TextContent) r.content().getFirst()).text();
}
@Test
void armedProfileReportsArmedAndUnarmedReportsUnarmed() {
FakeHerdr h = new FakeHerdr();
PeerLauncher workers = twoProfileLauncher(h);
FleetMcp.QuarantineSource source = new FleetMcp.QuarantineSource(
_ -> null, dev.ltms.fleet.placement.BackendQuarantine.none(),
profile -> "armed-profile".equals(profile));
McpSchema.CallToolResult res = FleetMcp.profiles(workers, source);
assertNotEquals(Boolean.TRUE, res.isError());
String out = textOf(res);
assertTrue(out.contains("\"exhaustionDetectionArmed\""), out);
assertTrue(out.contains("\"armed-profile\":true"), out);
assertTrue(out.contains("\"unarmed-profile\":false"), out);
}
}
@@ -1,109 +0,0 @@
package dev.ltms.fleet.mcp;
import dev.ltms.fleet.config.FleetConfig;
import dev.ltms.fleet.guard.SubscriptionGuard;
import dev.ltms.fleet.herdr.AgentControl;
import dev.ltms.fleet.herdr.FakeHerdr;
import dev.ltms.fleet.herdr.WorkspaceControl;
import dev.ltms.fleet.member.ClaudeCodeLauncher;
import dev.ltms.fleet.member.CompositePeerLauncher;
import dev.ltms.fleet.member.HerdrPeerLauncher;
import dev.ltms.fleet.peer.PeerLauncher;
import dev.ltms.fleet.placement.BackendOutagePolicy;
import dev.ltms.fleet.placement.BackendQuarantine;
import dev.ltms.fleet.placement.PlacementPolicies;
import org.junit.jupiter.api.Test;
import java.util.List;
import java.util.Map;
import java.util.Set;
import static org.junit.jupiter.api.Assertions.assertEquals;
import static org.junit.jupiter.api.Assertions.assertFalse;
/**
* fleetd #422 follow-up: {@code fleet_profiles}/{@code GET /profiles} — the reporting surface a
* lead actually reads — must let it tell apart the three states {@link
* dev.ltms.fleet.peer.PeerLauncher#disabledModels()} alone collapses into one empty set: no
* {@code models:} block at all, a block armed with nothing currently off, and a block with N
* models off. See {@link dev.ltms.fleet.peer.PeerLauncher.ModelGateState}'s javadoc for why a
* bare {@code disabledModels()} read cannot make this distinction, and {@link
* FleetMcp#profilesView} for where {@code modelGateArmed} is added alongside the existing {@code
* modelsOff} key.
*
* <p>Every assertion here goes through {@link FleetMcp#profilesView}, never {@code
* PeerLauncher.modelGateState()} directly — {@code CompositePeerLauncherTest} already proves the
* accessor itself; this class proves the surface a lead reads (fleet_profiles / GET /profiles)
* renders what that accessor reports.
*/
class FleetProfilesModelGateStateTest {
private static FleetConfig.Profile profile(String name, String model) {
return new FleetConfig.Profile(name, "http://gx00.gw:8000", model, null, "FLEETD_WORKER_TOKEN",
null, "tab", "fleetd-workers", "worker: {profile} #{n}", null, null, null);
}
private static FleetMcp.QuarantineSource noQuarantine() {
return new FleetMcp.QuarantineSource(_ -> null, BackendQuarantine.none(), _ -> false);
}
private static HerdrPeerLauncher claudeAdapter(FakeHerdr h, Map<String, FleetConfig.Profile> profiles) {
return new ClaudeCodeLauncher(new AgentControl(h), new WorkspaceControl(h),
new SubscriptionGuard(Set.of("gx00.gw")), profiles, "local", _ -> "tok");
}
/**
* State 1: no {@code models:} block at all — a plain {@code ClaudeCodeLauncher} (no {@code
* models:} supplier exists for it to read) has nothing to gate against, matching the fleet01
* host measured for this ticket: {@code grep -c '^models:' fleetd.yaml} returns 0 there.
*/
@Test
void noModelsBlockReportsGateNotArmed() {
FakeHerdr h = new FakeHerdr();
Map<String, FleetConfig.Profile> profiles = Map.of("local", profile("local", "deepseek-v4-flash"));
PeerLauncher workers = claudeAdapter(h, profiles);
Map<String, Object> view = FleetMcp.profilesView(workers, noQuarantine(), FleetMcp.OutageSource.none());
assertEquals(Boolean.FALSE, view.get("modelGateArmed"),
"no models: block to read from — nothing is gated, and nothing can be");
assertFalse(view.containsKey("modelsOff"), "nothing configured, so no off set to report either");
}
/** State 2: a {@code models:} block is present, but nothing in it is currently turned off. */
@Test
void modelsBlockWithNothingOffReportsGateArmedAndZeroOff() {
FakeHerdr h = new FakeHerdr();
Map<String, FleetConfig.Profile> profiles = Map.of("local", profile("local", "deepseek-v4-flash"));
FleetConfig.Models models = new FleetConfig.Models(
List.of(new FleetConfig.Models.ModelEntry("deepseek-v4-flash", true)));
PeerLauncher workers = new CompositePeerLauncher(List.of(claudeAdapter(h, profiles)), "local", profiles,
PlacementPolicies.fixed(), _ -> 0, null, BackendQuarantine.none(),
new BackendOutagePolicy(System::nanoTime), () -> models);
Map<String, Object> view = FleetMcp.profilesView(workers, noQuarantine(), FleetMcp.OutageSource.none());
assertEquals(Boolean.TRUE, view.get("modelGateArmed"),
"a models: block is present, so the gate is armed even though nothing is off yet");
assertFalse(view.containsKey("modelsOff"),
"nothing is off, so the key stays absent — an empty list here would be indistinguishable "
+ "from today's modelsOff omission, exactly the ambiguity modelGateArmed exists to remove");
}
/** State 3: a {@code models:} block is present with one model currently turned off. */
@Test
void modelsBlockWithModelsOffReportsGateArmedAndTheOffSet() {
FakeHerdr h = new FakeHerdr();
Map<String, FleetConfig.Profile> profiles = Map.of("local", profile("local", "deepseek-v4-flash"));
FleetConfig.Models models = new FleetConfig.Models(
List.of(new FleetConfig.Models.ModelEntry("deepseek-v4-flash", false)));
PeerLauncher workers = new CompositePeerLauncher(List.of(claudeAdapter(h, profiles)), "local", profiles,
PlacementPolicies.fixed(), _ -> 0, null, BackendQuarantine.none(),
new BackendOutagePolicy(System::nanoTime), () -> models);
Map<String, Object> view = FleetMcp.profilesView(workers, noQuarantine(), FleetMcp.OutageSource.none());
assertEquals(Boolean.TRUE, view.get("modelGateArmed"));
assertEquals(List.of("deepseek-v4-flash"), view.get("modelsOff"));
}
}
@@ -3,7 +3,6 @@ package dev.ltms.fleet.member;
import ch.qos.logback.classic.Logger;
import ch.qos.logback.classic.spi.ILoggingEvent;
import ch.qos.logback.core.read.ListAppender;
import dev.ltms.fleet.config.ConfigRef;
import dev.ltms.fleet.config.FleetConfig;
import dev.ltms.fleet.guard.SubscriptionGuard;
import dev.ltms.fleet.herdr.Agent;
@@ -23,11 +22,8 @@ import dev.ltms.fleet.placement.BackendQuarantine;
import dev.ltms.fleet.placement.PlacementException;
import dev.ltms.fleet.placement.PlacementPolicies;
import org.junit.jupiter.api.Test;
import org.junit.jupiter.api.io.TempDir;
import org.slf4j.LoggerFactory;
import java.nio.file.Files;
import java.nio.file.Path;
import java.util.EnumSet;
import java.util.HashMap;
import java.util.LinkedHashMap;
@@ -145,13 +141,6 @@ class CompositePeerLauncherTest {
null, null, credentialId, null);
}
/** Like {@link #stubWorker(String)}, but with an explicit {@code model:} for fleetd #422 tests. */
private static FleetConfig.Profile stubWorkerModel(String profile, String model) {
return new FleetConfig.Profile(profile, "http://gx00.gw:8000", model,
null, "FLEETD_WORKER_TOKEN", List.of("claude"), "tab", "fleetd-workers",
"w #{n}", null, null, null, null, null, null, null, null, null);
}
/**
* An <em>order-preserving</em> profile map. Never {@code Map.of} here: its iteration order is
* salted per JVM run, and the weighted policy breaks an exact-weight tie on candidate order —
@@ -536,31 +525,6 @@ class CompositePeerLauncherTest {
assertEquals("claude", h.profile(), "the returned handle carries the resolved default profile");
}
/**
* fleetd #435: {@code FixedPlacementPolicy} — the default placement policy every config uses
* unless {@code placement:} is set — never consulted {@code maxLoad}, so an unqualified spawn
* (a blank profile, the normal delegation path) landed on a capped default anyway. Measured on
* 7667727: a single dev profile at {@code maxLoad: 1} with 1 live, under {@code fixed()},
* returned "SPAWNED on profile=a". This test goes through {@code CompositePeerLauncher.spawn}
* with a blank profile, not the policy in isolation, so it proves the caller actually reaches
* the fixed default's new cap check rather than only the {@code select} method.
*/
@Test
void fixedPolicyGatesDefaultProfileAtMaxLoadOnUnqualifiedSpawn() {
FakeHerdr herdr = new FakeHerdr();
Map<String, FleetConfig.Profile> profiles = ordered(
"a", stubWorker("a", 1.0f, 1),
"b", stubWorker("b", 1.0f, null));
StubLauncher adapter = new StubLauncher("claude", herdr, profiles, "a", Set.of());
CompositePeerLauncher composite = new CompositePeerLauncher(
List.of(adapter), "a", profiles, PlacementPolicies.fixed(), name -> "a".equals(name) ? 1 : 0);
PeerHandle h = composite.spawn(new SpawnRequest(null, null, null));
assertEquals("b", h.profile(),
"the default profile a is at maxLoad, so fixed placement must fall through to b");
assertEquals(0, adapter.spawnCount("a"), "a is never spawned — it is already at cap");
}
@Test
void weightedPolicyGatesProfileAtMaxLoad() {
FakeHerdr herdr = new FakeHerdr();
@@ -1225,384 +1189,4 @@ class CompositePeerLauncherTest {
assertDoesNotThrow(() -> composite.spawn(new SpawnRequest("claude", null, null)));
assertDoesNotThrow(() -> composite.spawn(new SpawnRequest("gemini", null, null)));
}
// ── fleetd #422: the model on/off gate — a FOURTH, independent reason to refuse a spawn ──────
// ── (an operator decision, never a backend-reported outage) — never merged with quarantine ────
// ── or cool-off above ─────────────────────────────────────────────────────────────────────────
private static final BackendOutagePolicy NO_OUTAGE = new BackendOutagePolicy(() -> 0L);
/** Criterion 1: an explicit spawn onto a profile whose model is off is refused. */
@Test
void explicitSpawnOntoAnOffModelProfileIsRefused() {
FakeHerdr herdr = new FakeHerdr();
Map<String, FleetConfig.Profile> profiles = ordered(
"local", stubWorkerModel("local", "deepseek-v4-flash"),
"sonnet", stubWorkerModel("sonnet", "claude-sonnet-5"));
StubLauncher adapter = new StubLauncher("claude", herdr, profiles, "local", Set.of());
FleetConfig.Models models = new FleetConfig.Models(List.of(
new FleetConfig.Models.ModelEntry("deepseek-v4-flash", false)));
CompositePeerLauncher composite = new CompositePeerLauncher(List.of(adapter), "local", profiles,
PlacementPolicies.fixed(), _ -> 0, null, BackendQuarantine.none(), NO_OUTAGE,
() -> models);
PlacementException e = assertThrows(PlacementException.class,
() -> composite.spawn(new SpawnRequest("local", null, null)));
assertTrue(e.getMessage().contains("local"), "message names the profile: " + e.getMessage());
assertTrue(e.getMessage().contains("deepseek-v4-flash"),
"message names the model: " + e.getMessage());
assertTrue(e.getMessage().contains("operator") && e.getMessage().contains("turned off"),
"wording says the OPERATOR turned it off: " + e.getMessage());
// Distinct from quarantine/cool-off wording (criterion 1's explicit requirement).
assertFalse(e.getMessage().contains("quarantined"), "must not read like quarantine: " + e.getMessage());
assertFalse(e.getMessage().contains("cooling off"), "must not read like cool-off: " + e.getMessage());
assertEquals(0, adapter.spawnCount("local"), "the off-model profile is never delegated to");
}
/** A profile whose model is NOT off spawns normally even while another model is off. */
@Test
void explicitSpawnOntoAnEnabledModelProfileSucceedsWhileAnotherModelIsOff() {
FakeHerdr herdr = new FakeHerdr();
Map<String, FleetConfig.Profile> profiles = ordered(
"local", stubWorkerModel("local", "deepseek-v4-flash"),
"sonnet", stubWorkerModel("sonnet", "claude-sonnet-5"));
StubLauncher adapter = new StubLauncher("claude", herdr, profiles, "local", Set.of());
FleetConfig.Models models = new FleetConfig.Models(List.of(
new FleetConfig.Models.ModelEntry("deepseek-v4-flash", false)));
CompositePeerLauncher composite = new CompositePeerLauncher(List.of(adapter), "local", profiles,
PlacementPolicies.fixed(), _ -> 0, null, BackendQuarantine.none(), NO_OUTAGE,
() -> models);
assertDoesNotThrow(() -> composite.spawn(new SpawnRequest("sonnet", null, null)));
assertEquals(1, adapter.spawnCount("sonnet"));
}
/**
* Criterion 8: one {@code enabled: false} entry disables EVERY profile naming that model — the
* live example, {@code deepseek-v4-flash} backing both {@code local} and {@code local-direct}.
*/
@Test
void turningOffOneModelRefusesEveryProfileThatNamesIt() {
FakeHerdr herdr = new FakeHerdr();
Map<String, FleetConfig.Profile> profiles = ordered(
"local", stubWorkerModel("local", "deepseek-v4-flash"),
"local-direct", stubWorkerModel("local-direct", "deepseek-v4-flash"));
StubLauncher adapter = new StubLauncher("claude", herdr, profiles, "local", Set.of());
FleetConfig.Models models = new FleetConfig.Models(List.of(
new FleetConfig.Models.ModelEntry("deepseek-v4-flash", false)));
CompositePeerLauncher composite = new CompositePeerLauncher(List.of(adapter), "local", profiles,
PlacementPolicies.fixed(), _ -> 0, null, BackendQuarantine.none(), NO_OUTAGE,
() -> models);
assertThrows(PlacementException.class, () -> composite.spawn(new SpawnRequest("local", null, null)));
assertThrows(PlacementException.class,
() -> composite.spawn(new SpawnRequest("local-direct", null, null)),
"local-direct shares local's model, so it must be refused too");
assertEquals(0, adapter.spawnCount("local"));
assertEquals(0, adapter.spawnCount("local-direct"));
}
/** Criterion 3: an old-style entry with no {@code enabled} field never refuses a spawn. */
@Test
void anEntryWithNoEnabledFieldNeverRefusesASpawn() {
FakeHerdr herdr = new FakeHerdr();
Map<String, FleetConfig.Profile> profiles = Map.of("local", stubWorkerModel("local", "deepseek-v4-flash"));
StubLauncher adapter = new StubLauncher("claude", herdr, profiles, "local", Set.of());
// The back-compat single-arg ModelEntry constructor — no enabled field in the shape at all.
FleetConfig.Models models = new FleetConfig.Models(
List.of(new FleetConfig.Models.ModelEntry("deepseek-v4-flash")));
CompositePeerLauncher composite = new CompositePeerLauncher(List.of(adapter), "local", profiles,
PlacementPolicies.fixed(), _ -> 0, null, BackendQuarantine.none(), NO_OUTAGE,
() -> models);
assertDoesNotThrow(() -> composite.spawn(new SpawnRequest("local", null, null)));
assertEquals(1, adapter.spawnCount("local"));
}
/** A model absent from {@code models.allow:} entirely (nothing to gate against) is never refused. */
@Test
void aModelNotConfiguredInTheAllowListIsNeverGated() {
FakeHerdr herdr = new FakeHerdr();
Map<String, FleetConfig.Profile> profiles = Map.of("local", stubWorkerModel("local", "unlisted-model"));
StubLauncher adapter = new StubLauncher("claude", herdr, profiles, "local", Set.of());
FleetConfig.Models models = new FleetConfig.Models(List.of(
new FleetConfig.Models.ModelEntry("deepseek-v4-flash", false)));
CompositePeerLauncher composite = new CompositePeerLauncher(List.of(adapter), "local", profiles,
PlacementPolicies.fixed(), _ -> 0, null, BackendQuarantine.none(), NO_OUTAGE,
() -> models);
assertDoesNotThrow(() -> composite.spawn(new SpawnRequest("local", null, null)));
}
/** Criterion 2: an unqualified spawn skips an off-model candidate and lands on another one. */
@Test
void placementSkipsAnOffModelProfileAndRoutesToAnotherOne() {
FakeHerdr herdr = new FakeHerdr();
Map<String, FleetConfig.Profile> profiles = ordered(
"local", stubWorkerModel("local", "deepseek-v4-flash"),
"sonnet", stubWorkerModel("sonnet", "claude-sonnet-5"));
StubLauncher adapter = new StubLauncher("claude", herdr, profiles, "local", Set.of());
FleetConfig.Models models = new FleetConfig.Models(List.of(
new FleetConfig.Models.ModelEntry("deepseek-v4-flash", false)));
CompositePeerLauncher composite = new CompositePeerLauncher(List.of(adapter), "local", profiles,
PlacementPolicies.weighted(), _ -> 0, null, BackendQuarantine.none(), NO_OUTAGE,
() -> models);
PeerHandle h = composite.spawn(new SpawnRequest(null, null, null));
assertEquals("sonnet", h.profile(), "local's model is off, so an unqualified spawn must land on sonnet");
assertEquals(0, adapter.spawnCount("local"));
}
/**
* Criterion 2's second half: when EVERY candidate's model is off, the placement exception must
* name that as the cause — not a generic "no candidates" message.
*/
@Test
void automaticPlacementNamesModelOffWhenEveryCandidateIsOffModel() {
FakeHerdr herdr = new FakeHerdr();
Map<String, FleetConfig.Profile> profiles = ordered(
"local", stubWorkerModel("local", "deepseek-v4-flash"),
"local-direct", stubWorkerModel("local-direct", "deepseek-v4-flash"));
StubLauncher adapter = new StubLauncher("claude", herdr, profiles, "local", Set.of());
FleetConfig.Models models = new FleetConfig.Models(List.of(
new FleetConfig.Models.ModelEntry("deepseek-v4-flash", false)));
CompositePeerLauncher composite = new CompositePeerLauncher(List.of(adapter), "local", profiles,
PlacementPolicies.weighted(), _ -> 0, null, BackendQuarantine.none(), NO_OUTAGE,
() -> models);
PlacementException e = assertThrows(PlacementException.class,
() -> composite.spawn(new SpawnRequest(null, null, null)),
"both candidates share the off model — nothing is available");
assertTrue(e.getMessage().contains("model") && e.getMessage().contains("turned off"),
"message names model-off as the cause, not a generic no-candidates message: "
+ e.getMessage());
}
/**
* fleetd #422 follow-up: {@code fixed} is the DEFAULT placement policy ({@code
* PlacementPolicies.fromName} returns it for an absent/blank name), and it built its own inline
* candidate filter instead of calling {@code PlacementPolicyUtil.available()} — so it never
* checked {@code modelOff()}. Mirrors {@code placementSkipsAnOffModelProfileAndRoutesToAnotherOne}
* above with only the policy swapped, to prove the gate now fires on the path most fleets use.
*/
@Test
void fixedPlacementSkipsAnOffModelProfileToo() {
FakeHerdr herdr = new FakeHerdr();
Map<String, FleetConfig.Profile> profiles = ordered(
"local", stubWorkerModel("local", "deepseek-v4-flash"),
"sonnet", stubWorkerModel("sonnet", "claude-sonnet-5"));
StubLauncher adapter = new StubLauncher("claude", herdr, profiles, "local", Set.of());
FleetConfig.Models models = new FleetConfig.Models(List.of(
new FleetConfig.Models.ModelEntry("deepseek-v4-flash", false)));
CompositePeerLauncher composite = new CompositePeerLauncher(List.of(adapter), "local", profiles,
PlacementPolicies.fixed(), _ -> 0, null, BackendQuarantine.none(), NO_OUTAGE,
() -> models);
PeerHandle h = composite.spawn(new SpawnRequest(null, null, null));
assertEquals("sonnet", h.profile(), "local's model is off, so an unqualified spawn must land on sonnet");
assertEquals(0, adapter.spawnCount("local"));
}
/**
* Criterion 4: turning a model off/on is HOT — no restart — proven through a REAL
* {@code ConfigRef.reload()}, not a hand-rolled supplier swap. Also proves {@code models} is
* correctly reclassified: the reload's {@link ConfigRef.Outcome#applied()} is {@code true} and
* {@code "models"} never appears in {@link ConfigRef.Outcome#deferred()}.
*/
@Test
void modelOnOffIsHotReloadedThroughARealConfigRef(@TempDir Path dir) throws Exception {
Path yaml = dir.resolve("fleetd.yaml");
Files.writeString(yaml, """
bind:
port: 8080
profiles:
local:
baseUrl: http://local.gw:8000
model: deepseek-v4-flash
models:
allow:
- model: deepseek-v4-flash
""");
FleetConfig initial = FleetConfig.load(yaml);
ConfigRef configRef = new ConfigRef(yaml, initial);
FakeHerdr herdr = new FakeHerdr();
StubLauncher adapter = new StubLauncher("claude", herdr,
Map.of("local", stubWorker("local")), "local", Set.of());
CompositePeerLauncher composite = new CompositePeerLauncher(
List.of(adapter), "local", configRef, _ -> 0, BackendQuarantine.none(), NO_OUTAGE);
assertDoesNotThrow(() -> composite.spawn(new SpawnRequest("local", null, null)),
"the model starts enabled");
assertEquals(1, adapter.spawnCount("local"));
Files.writeString(yaml, """
bind:
port: 8080
profiles:
local:
baseUrl: http://local.gw:8000
model: deepseek-v4-flash
models:
allow:
- model: deepseek-v4-flash
enabled: false
""");
ConfigRef.Outcome outcome = configRef.reload();
assertTrue(outcome.applied(), "a models.allow on/off edit must apply live, never be refused");
assertFalse(outcome.deferred().contains("models"),
"models is hot-excluded now — it must never be reported as a deferred key");
PlacementException e = assertThrows(PlacementException.class,
() -> composite.spawn(new SpawnRequest("local", null, null)),
"the very next spawn must see the reload, with no restart");
assertTrue(e.getMessage().contains("deepseek-v4-flash"));
assertEquals(1, adapter.spawnCount("local"), "still just the one successful spawn from before");
// And back on, still hot, still no restart.
Files.writeString(yaml, """
bind:
port: 8080
profiles:
local:
baseUrl: http://local.gw:8000
model: deepseek-v4-flash
models:
allow:
- model: deepseek-v4-flash
enabled: true
""");
ConfigRef.Outcome reEnabled = configRef.reload();
assertTrue(reEnabled.applied());
assertDoesNotThrow(() -> composite.spawn(new SpawnRequest("local", null, null)));
assertEquals(2, adapter.spawnCount("local"));
}
/**
* fleetd #422 follow-up, acceptance criterion 2: {@link CompositePeerLauncher#modelGateState()}
* is LIVE — no restart — proven through a REAL {@link ConfigRef#reload()}, exactly like {@link
* #modelOnOffIsHotReloadedThroughARealConfigRef} above proves for the on/off gate itself. This
* single reload sequence walks through all three states the ticket asks for: no {@code models:}
* block, a block armed with nothing off, and a block with one model off — so a reload that flips
* between any of the three is proven live, not just the on/off edit within an already-armed block.
*/
@Test
void modelGateStateIsHotReloadedThroughARealConfigRef(@TempDir Path dir) throws Exception {
Path yaml = dir.resolve("fleetd.yaml");
Files.writeString(yaml, """
bind:
port: 8080
profiles:
local:
baseUrl: http://local.gw:8000
model: deepseek-v4-flash
""");
FleetConfig initial = FleetConfig.load(yaml);
ConfigRef configRef = new ConfigRef(yaml, initial);
FakeHerdr herdr = new FakeHerdr();
StubLauncher adapter = new StubLauncher("claude", herdr,
Map.of("local", stubWorker("local")), "local", Set.of());
CompositePeerLauncher composite = new CompositePeerLauncher(
List.of(adapter), "local", configRef, _ -> 0, BackendQuarantine.none(), NO_OUTAGE);
PeerLauncher.ModelGateState notConfigured = composite.modelGateState();
assertFalse(notConfigured.configured(), "no models: block in the config at all");
assertEquals(Set.of(), notConfigured.off());
Files.writeString(yaml, """
bind:
port: 8080
profiles:
local:
baseUrl: http://local.gw:8000
model: deepseek-v4-flash
models:
allow:
- model: deepseek-v4-flash
enabled: false
""");
assertTrue(configRef.reload().applied(), "adding a models: block must apply live, no restart");
PeerLauncher.ModelGateState armedWithOneOff = composite.modelGateState();
assertTrue(armedWithOneOff.configured(), "a models: block now exists — the gate is armed");
assertEquals(Set.of("deepseek-v4-flash"), armedWithOneOff.off());
Files.writeString(yaml, """
bind:
port: 8080
profiles:
local:
baseUrl: http://local.gw:8000
model: deepseek-v4-flash
models:
allow:
- model: deepseek-v4-flash
enabled: true
""");
assertTrue(configRef.reload().applied(), "flipping the entry back on must apply live too");
PeerLauncher.ModelGateState armedWithZeroOff = composite.modelGateState();
assertTrue(armedWithZeroOff.configured(),
"the block is still present — armed and reporting zero, not the same as no block at all");
assertEquals(Set.of(), armedWithZeroOff.off());
}
/**
* fleetd #422 follow-up, acceptance criterion 3: the invariant is that an absent {@code
* models:} block stays permitted and must never be fatal. Proved, not assumed — a config
* without one loads, validates, reports the gate as not configured, AND still spawns normally
* (no {@link PlacementException} from a gate that has nothing to check against), using the same
* production-shaped {@code Supplier<FleetConfig>} wiring {@code Fleetd.main} actually uses.
*/
@Test
void noModelsBlockConfigStillLoadsAndSpawnsNormally(@TempDir Path dir) throws Exception {
Path yaml = dir.resolve("fleetd.yaml");
Files.writeString(yaml, """
bind:
port: 8080
profiles:
local:
baseUrl: http://local.gw:8000
model: deepseek-v4-flash
""");
FleetConfig cfg = FleetConfig.load(yaml);
assertDoesNotThrow(cfg::validateAll, "a config with no models: block must load and validate cleanly");
ConfigRef configRef = new ConfigRef(yaml, cfg);
FakeHerdr herdr = new FakeHerdr();
StubLauncher adapter = new StubLauncher("claude", herdr,
Map.of("local", stubWorker("local")), "local", Set.of());
CompositePeerLauncher composite = new CompositePeerLauncher(
List.of(adapter), "local", configRef, _ -> 0, BackendQuarantine.none(), NO_OUTAGE);
assertFalse(composite.modelGateState().configured());
assertDoesNotThrow(() -> composite.spawn(new SpawnRequest("local", null, null)),
"no models: block means nothing to gate against — the spawn must go through");
assertEquals(1, adapter.spawnCount("local"));
}
/** {@code fleet_profiles}/{@code GET /profiles} must read the exact same live source the gate reads. */
@Test
void disabledModelsReportsWhatTheGateActuallyEnforces() {
FakeHerdr herdr = new FakeHerdr();
Map<String, FleetConfig.Profile> profiles = ordered(
"local", stubWorkerModel("local", "deepseek-v4-flash"),
"sonnet", stubWorkerModel("sonnet", "claude-sonnet-5"));
StubLauncher adapter = new StubLauncher("claude", herdr, profiles, "local", Set.of());
FleetConfig.Models models = new FleetConfig.Models(List.of(
new FleetConfig.Models.ModelEntry("deepseek-v4-flash", false),
new FleetConfig.Models.ModelEntry("claude-sonnet-5", true)));
CompositePeerLauncher composite = new CompositePeerLauncher(List.of(adapter), "local", profiles,
PlacementPolicies.fixed(), _ -> 0, null, BackendQuarantine.none(), NO_OUTAGE,
() -> models);
assertEquals(Set.of("deepseek-v4-flash"), composite.disabledModels());
}
/** A caller with no {@code models:} block (the map-based constructors) reports nothing off. */
@Test
void disabledModelsIsEmptyWithNoModelsConfigured() {
FakeHerdr herdr = new FakeHerdr();
PeerLauncher composite = composite(herdr);
assertEquals(Set.of(), composite.disabledModels());
}
}
@@ -120,202 +120,40 @@ class EnvAllowListScrubTest {
}
/**
* fleetd #394: the actual defect. Plain {@code export "$n="} is FATAL for a zsh read-only or
* special parameter (e.g. {@code UID}) and aborts the whole sourced file — every name still to
* come is never blanked, and the {@code scrub-report.txt} below is never written at all,
* silently ({@code 2>/dev/null} swallows the error). This plants an unblankable, exported,
* read-only variable in the MIDDLE of the names the scrub attempts to blank, with two more
* names after it, and asserts that both of those later names are STILL blanked and the report
* is STILL written with the failure counted — a test that only checked names BEFORE the failure
* point would pass today and prove nothing.
* A pane inherits {@code UID}; a cleared test parent does not. The scrub must survive it.
*
* <p>The planted name is a made-up one ({@code FLEETD_TEST_UNBLANKABLE}), not {@code UID} or
* any other name a skip-list might already know about — invariant 1 is that the loop survives
* ANY unblankable name, not a known one, so the test must not lean on one either.
* <p>Every other test here starts zsh from a CLEARED environment, so {@code UID} is never an
* exported name and never reaches the blanking loop. In a real member pane it is exported and
* it IS reached — and {@code export UID=} is a fatal zsh parameter error that aborts the whole
* sourced file, leaving every later name unscrubbed and writing no report at all. The abort is
* silent: the loop is wrapped in {@code 2>/dev/null}.
*
* <p>Exercises the real artefact: {@link EnvAllowListScrub#scrubScript} is run verbatim under a
* real {@code /bin/zsh}, not just asserted on as a Java string. The four planted names are
* exported one at a time via {@code typeset -x}/{@code typeset -rx} immediately before the
* script runs, in a fixed order — zsh's {@code export}/{@code typeset -x} appends to the
* process's environment table in call order (verified empirically: a freshly-exported name
* always sorts after every inherited one and after every earlier freshly-exported name in
* {@code command env}'s own output), which is what makes the "middle" position deterministic
* here, unlike relying on the OS's own inherited-environment order.
* <p>The assertion is deliberately "a report exists" rather than "the canary is blanked". The
* report is written by the last statement in the file, so its presence proves the script ran
* to completion; the canary alone would depend on where it happens to sit in {@code env} order.
* Both are checked, but only the first one fails deterministically without the fix.
*/
@Test
void unblankableNameInTheMiddleDoesNotAbortNamesAfterIt(@TempDir Path tmp) throws Exception {
void scrubSurvivesAnInheritedUidTheWayARealPaneHasIt(@TempDir Path tmp) throws Exception {
assumeTrue(Files.isExecutable(ZSH), "/bin/zsh not present — nothing to prove here");
Set<String> allowed = MemberEnvAllowList.derive(List.of());
Path zdotdir = EnvAllowListScrub.generate(tmp, allowed);
// Only ZDOTDIR is allowed — it must survive the scrub itself, since the report is written
// to "$ZDOTDIR/..." AFTER the blanking loop runs; if ZDOTDIR were blanked as a side effect,
// the report write would silently go to the wrong place instead of testing anything.
String script = EnvAllowListScrub.scrubScript(Set.of("ZDOTDIR"));
String setup = """
typeset -x FLEETD_TEST_BEFORE=1
typeset -rx FLEETD_TEST_UNBLANKABLE=1
typeset -x FLEETD_TEST_AFTER_A=1
typeset -x FLEETD_TEST_AFTER_B=1
""";
// The production shape: UID present and exported, as every pane shell inherits it.
Map<String, String> paneLikeParent = Map.of(
"HOME", System.getProperty("user.home"),
"PATH", "/usr/bin:/bin",
"SHELL", "/bin/zsh",
"UID", "1000",
"CB633_CANARY", "must-not-survive-the-scrub");
Set<String> survivors = exportedNamesFromCleanParent(paneLikeParent, zdotdir);
ProcessBuilder pb = new ProcessBuilder("/bin/zsh");
pb.environment().clear();
pb.environment().put("PATH", "/usr/bin:/bin");
pb.environment().put("ZDOTDIR", tmp.toAbsolutePath().toString());
pb.redirectError(ProcessBuilder.Redirect.DISCARD);
Process zsh = pb.start();
zsh.getOutputStream().write((setup + script).getBytes(StandardCharsets.UTF_8));
zsh.getOutputStream().flush();
zsh.getOutputStream().close();
assertTrue(zsh.waitFor(60, java.util.concurrent.TimeUnit.SECONDS),
"the scrub script did not exit within 60s");
assertEquals(0, zsh.exitValue(),
"the scrub script itself must never abort — an unblankable name must not kill the "
+ "sourced file");
EnvAllowListScrub.ScrubReport report = EnvAllowListScrub.readReport(tmp);
assertNotNull(report, "the report must still be written even though one name could not be "
+ "blanked — a report that silently never appears is the #394 bug");
assertTrue(report.blanked().contains("FLEETD_TEST_BEFORE"),
"sanity: the name before the unblankable one must be blanked");
assertTrue(report.blanked().contains("FLEETD_TEST_AFTER_A"),
"the FIRST name AFTER the unblankable one must still be blanked — before the fix, "
+ "the whole loop aborted at the unblankable name and every later name was "
+ "silently left untouched");
assertTrue(report.blanked().contains("FLEETD_TEST_AFTER_B"),
"the SECOND name after the unblankable one must also still be blanked");
assertTrue(report.unblankable().contains("FLEETD_TEST_UNBLANKABLE"),
"the unblankable name is reported by name, not silently dropped");
// fleetd #400 note: zsh itself auto-exports SHLVL on every shell start (measured: it appears
// in `command env` even from a fully cleared parent), and a bare assignment to it is coerced
// rather than failing — exactly the shape #400 fixes. Before that fix, eval's exit status
// alone silently misclassified SHLVL as blanked, so this test's old "exactly one" assertion
// passed by accident: it never actually proved SHLVL was absent from the candidates, only
// that the old bug hid it. Now that classification reads the value back, SHLVL and
// FLEETD_TEST_UNBLANKABLE both correctly land in unblankable() — real, unplanted evidence
// the #400 fix works, not just the synthetic case in the dedicated #400 test above.
assertTrue(report.unblankable().contains("SHLVL"),
"fleetd #400: zsh's own auto-exported SHLVL must also be reported unblankable, not "
+ "silently miscounted as blanked");
assertEquals(report.unblankable().size(), report.failed(),
"the failed count must equal the number of names actually reported unblankable");
}
/**
* fleetd #400: {@code eval}'s exit status is not proof that a name was actually blanked. zsh
* coerces a bare {@code NAME=} assignment on an integer special parameter to a number instead of
* failing, so {@code eval} reports success while the value stays non-empty — a status-based
* classification calls that "blanked" when it was not. This drives all three shapes a name can
* take through the real {@code scrubScript} in ONE run: a normal, genuinely blankable name; a
* fatal one ({@code LINENO} — deliberately not {@code UID}, so this test does not depend on the
* harness's uid); and the silent-no-op one the ticket is about ({@code SECONDS}, rc 0 but
* unchanged). Under the pre-#400 exit-status check, {@code SECONDS} would land in
* {@code blanked()} — that is the exact false receipt this fix removes.
*
* <p>Criterion 3: a cleared {@code ProcessBuilder} parent does not, by itself, give the child
* zsh any of these names — {@code SECONDS}/{@code LINENO} are zsh's own built-in parameters and
* only become CANDIDATES the enumeration loop can see (i.e. show up in {@code command env}) when
* they arrive via the process's own environment table, not merely by existing as zsh parameters
* inside the shell. So each is put into {@code pb.environment()} explicitly, after
* {@code clear()} — confirmed empirically first (a throwaway probe piping
* {@code env -i PATH=... SECONDS=999 LINENO=999 FLEETD_TEST_NORMAL=1 zsh -c 'command env | cut
* -d= -f1'}) that all three names really appear in {@code command env}'s output under exactly
* this construction, not relying on whatever the test-runner's own ambient environment happens
* to contain.
*/
@Test
void classifiesByObservedValueNotExitStatusAcrossAllThreeShapes(@TempDir Path tmp) throws Exception {
assumeTrue(Files.isExecutable(ZSH), "/bin/zsh not present — nothing to prove here");
String script = EnvAllowListScrub.scrubScript(Set.of("ZDOTDIR"));
ProcessBuilder pb = new ProcessBuilder("/bin/zsh");
pb.environment().clear();
pb.environment().put("PATH", "/usr/bin:/bin");
pb.environment().put("ZDOTDIR", tmp.toAbsolutePath().toString());
// Explicitly placed in the child's environment table — see the javadoc above on why a
// cleared parent alone does not put these on the enumeration loop's candidate list.
pb.environment().put("SECONDS", "999"); // rc 0, value coerced/unchanged — the #400 bug
pb.environment().put("LINENO", "999"); // fatal on assignment, eval rc != 0, contained
pb.environment().put("FLEETD_TEST_NORMAL", "1"); // genuinely blankable, the control case
pb.redirectError(ProcessBuilder.Redirect.DISCARD);
Process zsh = pb.start();
zsh.getOutputStream().write(script.getBytes(StandardCharsets.UTF_8));
zsh.getOutputStream().flush();
zsh.getOutputStream().close();
assertTrue(zsh.waitFor(60, java.util.concurrent.TimeUnit.SECONDS),
"the scrub script did not exit within 60s");
assertEquals(0, zsh.exitValue(),
"the scrub script must still reach its end with both a fatal name and a silent "
+ "no-op name among the candidates");
EnvAllowListScrub.ScrubReport report = EnvAllowListScrub.readReport(tmp);
assertNotNull(report, "the report must still be written");
assertTrue(report.blanked().contains("FLEETD_TEST_NORMAL"),
"the control case: an ordinary name is genuinely blankable and must be reported so");
assertTrue(report.unblankable().contains("LINENO"),
"a name fatal to assign to must be reported unblankable — sanity check that "
+ "containment still works under the new classification");
assertTrue(report.unblankable().contains("SECONDS"),
"the #400 defect: eval returns rc 0 for SECONDS (zsh coerces the assignment instead "
+ "of failing) but the value is left non-empty — classifying on the observed "
+ "value catches this; classifying on eval's exit status would have called "
+ "this \"blanked\" and produced a false receipt");
assertFalse(report.blanked().contains("SECONDS"),
"SECONDS must never appear as blanked — it was never actually emptied");
}
/**
* fleetd #394 follow-up: the blanking loop's {@code eval "export ${n}="} splices {@code n} into
* a string that zsh then interprets as shell syntax. That is only safe because every name
* reaching {@code _cb633_blank} already passed an identifier check in the ENUMERATION loop
* (20 lines away, in a different loop) — so the fix re-asserts the identical check immediately
* before the {@code eval} call, rather than trusting that distant guard to keep holding.
*
* <p>This test plants a value with an embedded newline, exploiting the exact "junk from
* multi-line values" gap the enumeration loop's own comment already documents: {@code command
* env}'s text output is read line-by-line, so a value's second line becomes a spurious extra
* "name" that was never a real exported variable. The fragment used here ({@code
* junk.fragment}) is merely non-conforming (it contains a dot) — never command-shaped; this
* test must never demonstrate command execution and plants no command-shaped payload.
*
* <p>Exercises the real artefact end-to-end: {@link EnvAllowListScrub#scrubScript} runs
* verbatim under a real {@code /bin/zsh}, exactly as {@code generate()} would produce it — this
* is not a synthetic call into just the blanking loop.
*/
@Test
void nonIdentifierJunkFromAMultilineValueIsSkippedNotBlankedOrUnblankable(@TempDir Path tmp)
throws Exception {
assumeTrue(Files.isExecutable(ZSH), "/bin/zsh not present — nothing to prove here");
String script = EnvAllowListScrub.scrubScript(Set.of("ZDOTDIR"));
ProcessBuilder pb = new ProcessBuilder("/bin/zsh");
pb.environment().clear();
pb.environment().put("PATH", "/usr/bin:/bin");
pb.environment().put("ZDOTDIR", tmp.toAbsolutePath().toString());
// Embedded newline: `command env`'s own text output splits this into two lines, and the
// second ("junk.fragment") has no "=" at all, so `cut -d= -f1` returns it unchanged as a
// spurious candidate "name" — it was never an actual exported variable by that name.
pb.environment().put("FLEETD_TEST_MULTILINE", "keep\njunk.fragment");
pb.redirectError(ProcessBuilder.Redirect.DISCARD);
Process zsh = pb.start();
zsh.getOutputStream().write(script.getBytes(StandardCharsets.UTF_8));
zsh.getOutputStream().flush();
zsh.getOutputStream().close();
assertTrue(zsh.waitFor(60, java.util.concurrent.TimeUnit.SECONDS),
"the scrub script did not exit within 60s");
assertEquals(0, zsh.exitValue(), "the scrub script must reach its end");
EnvAllowListScrub.ScrubReport report = EnvAllowListScrub.readReport(tmp);
assertNotNull(report, "the report must still be written");
assertTrue(report.blanked().contains("FLEETD_TEST_MULTILINE"),
"sanity: the real, identifier-shaped variable must still be blanked normally");
assertFalse(report.blanked().contains("junk.fragment"),
"a non-identifier fragment is not a real variable and must never be blanked");
assertFalse(report.unblankable().contains("junk.fragment"),
"a non-identifier fragment must never even become a candidate the blanking loop "
+ "attempts — it must be filtered before either guard has to catch it, so "
+ "it is neither blanked nor counted as a failed attempt");
EnvAllowListScrub.ScrubReport report = EnvAllowListScrub.readReport(zdotdir);
assertNotNull(report,
"an inherited UID must not abort the scrub — no report means the file died mid-loop "
+ "and every name after UID in `env` order was left unscrubbed");
assertFalse(survivors.contains("CB633_CANARY"),
"a non-allow-listed name must still be blanked when UID is in the environment");
}
/** A group-shared ZDOTDIR still lets the member truncate and write its pre-created receipt. */
@@ -1802,11 +1802,7 @@ class MessageServiceTest {
Thread asker = new Thread(() -> assertThrows(IllegalStateException.class,
() -> service.ask(T, "which config file?", 30_000)));
asker.start();
// Phase.ASKING (markAsyncQuestion) is only the FIRST of ask()'s three steps; the assertion
// below depends on the THIRD (pushLoop.onQuestionOpened). Barrier on the push loop's own
// pending-question state instead of the phase — see ReplyPushLoop#pendingQuestionTurnIdsForTest.
MessageService.TaskView asking = awaitTicketPhaseOn(service, ticket, MessageService.Phase.ASKING);
awaitQuestionPendingOn(pushLoop, LEAD, asking.turnId());
awaitTicketPhaseOn(service, ticket, MessageService.Phase.ASKING);
assertEquals(ReplyPushLoop.Action.INJECT, pushLoop.decide(LEAD, 0, 0, 0),
"the open question should be the one thing keeping this lead's schedule alive");
@@ -1921,12 +1917,6 @@ class MessageServiceTest {
// Only now does it reply. Under the old clock this reply was born already expired.
assertTrue(rendezvous.resolve(T, "the long report"));
awaitTicketPhaseOn(wiring.service(), slow, MessageService.Phase.DONE);
// #399: same barrier fix as the sibling eviction test, for consistency — this test's
// own assertion happens to survive a late stamp today (the clock is not advanced any
// further between here and the sweep below, so cutoff cannot move past whichever
// value completedNanos ends up stamped with), but DONE is still the wrong thing to
// order on before a sweep that matters for the TTL.
awaitCompletionStamped(wiring.service(), slow);
// A second delegation runs pruneTerminalTickets before it returns.
wiring.service().sendAsync(T, "an unrelated second task");
@@ -1956,11 +1946,6 @@ class MessageServiceTest {
injectDelivery();
assertTrue(rendezvous.resolve(T, "quick result"));
awaitTicketPhaseOn(wiring.service(), done, MessageService.Phase.DONE);
// #399: DONE can be observed before the completion hook stamps completedNanos. Wait
// for the real stamp before advancing the clock, or the hook can run late and stamp
// the ADVANCED time — making cutoff = advanced - TTL unreachable and hiding the very
// eviction this test exists to pin.
awaitCompletionStamped(wiring.service(), done);
// Nobody collected it, and the TTL has now passed since it FINISHED.
clock.addAndGet(MessageService.TICKET_TTL_NANOS + TimeUnit.SECONDS.toNanos(1));
@@ -1971,80 +1956,6 @@ class MessageServiceTest {
}
}
/**
* fleetd #409: pins the #399 ordering invariant deterministically, on the first run, without
* relying on host load.
*
* <p>fleetd #399 was itself only reproducible probabilistically: the real race window between
* {@code CompletableFuture.complete()} making {@link MessageService.Phase#DONE} observable and
* the constructor's {@code whenComplete} hook actually stamping {@code completedNanos} (see
* {@link MessageService#isCompletionStampedForTest}) is normally a handful of instructions wide,
* and needed heavy background load on the host to show up in a run at all. This test does not
* try to hit that narrow window by chance — it widens it on purpose: {@link #STAMP_DELAY_MILLIS}
* is injected into the clock itself, so the completion hook's one read of {@code nowNanos} for
* this ticket sleeps before returning, and everything the test does in the meantime (advance the
* clock, run a sweep, assert) happens for certain inside that window, on any host.
*
* <p>Removing the {@link #awaitCompletionStamped} call below reproduces the pre-#399 ordering:
* the test then advances the clock and runs its sweep while the hook is still asleep, so the
* hook wakes up and stamps {@code completedNanos} with the clock's ALREADY-ADVANCED value
* instead of the real completion time. {@code cutoff = advanced - TICKET_TTL_NANOS} can then
* never exceed that stamp (the gap between them is fixed at exactly {@code TICKET_TTL_NANOS}),
* so the sweep that already ran never evicts the ticket and the very next assertion — expecting
* eviction — fails immediately. That is a deterministic, first-run RED failure, not a flaky one
* and not a false pass: I ran the test with the barrier call removed and confirmed
* {@code assertNull} fails because {@code poll} still returns the ticket's DONE view, which is
* exactly this masked-eviction mechanism and not some unrelated defect in the test's own wiring.
*/
@Test
void aTicketOrderedOnDoneInsteadOfTheCompletionStampSurvivesAnEvictionItMustNotSurvive() throws Exception {
java.util.concurrent.atomic.AtomicLong clock = new java.util.concurrent.atomic.AtomicLong(1_000_000_000L);
// Armed for exactly one call: MessageService's only three nowNanos() call sites are the Task
// constructor's createdNanos, this constructor's whenComplete hook's completedNanos, and
// pruneTerminalTickets' cutoff — none of which run between arming this (right before
// triggering the reply below) and the completion hook firing, so the delayed call is
// unambiguously that hook's stamp for `ticket`, never a createdNanos or cutoff read.
java.util.concurrent.atomic.AtomicBoolean delayArmed = new java.util.concurrent.atomic.AtomicBoolean(false);
java.util.function.LongSupplier gatedClock = () -> {
if (delayArmed.compareAndSet(true, false)) {
try {
Thread.sleep(STAMP_DELAY_MILLIS);
} catch (InterruptedException e) {
Thread.currentThread().interrupt();
}
}
return clock.get();
};
MessageService service = new MessageService(agents, injector, rendezvous, inbox, null, null, gatedClock);
try {
String ticket = service.sendAsync(T, "a quick task");
awaitWaiting();
injectDelivery();
delayArmed.set(true);
assertTrue(rendezvous.resolve(T, "quick result"));
awaitTicketPhaseOn(service, ticket, MessageService.Phase.DONE);
// The barrier under test (fleetd #409, same fix as fleetd #399): remove this one call to
// reproduce the pre-#399 ordering — see the class-level note above for what happens then.
awaitCompletionStamped(service, ticket);
// A real clock would need the whole TTL to pass; the injected one does it instantly, and
// by now completedNanos already holds the REAL (small, unadvanced) completion time.
clock.addAndGet(MessageService.TICKET_TTL_NANOS + TimeUnit.SECONDS.toNanos(1));
service.sendAsync(T, "an unrelated second task"); // runs pruneTerminalTickets before returning
assertNull(service.poll(ticket),
"a ticket whose real completion time is long past the advanced cutoff must be "
+ "evicted, whatever the completion hook's clock read was delayed by");
} finally {
service.close();
}
}
/** Bounded delay the injected clock sleeps for in {@link #aTicketOrderedOnDoneInsteadOfTheCompletionStampSurvivesAnEvictionItMustNotSurvive}. */
private static final long STAMP_DELAY_MILLIS = 300;
private MessageService.TaskView awaitTicketPhaseOn(MessageService svc, String ticket,
MessageService.Phase phase) throws Exception {
long deadline = System.currentTimeMillis() + 3000;
@@ -2079,45 +1990,6 @@ class MessageServiceTest {
return view;
}
/**
* fleetd #399: waits until {@code ticket}'s completion hook has actually stamped
* {@code completedNanos}, not just until {@link MessageService#poll} reports
* {@link MessageService.Phase#DONE} for it. {@code poll} can observe {@code DONE} the instant
* the task's future resolves, before the {@code whenComplete} hook that stamps the completion
* time has run — {@code CompletableFuture.complete()} publishes its result and only then runs
* dependents. A test that is about to advance an injected clock past the TTL must order itself
* after the stamp, not after {@code DONE}: winning the race the other way stamps the
* *advanced* clock value and can hide a real eviction bug behind a false pass.
*/
private void awaitCompletionStamped(MessageService svc, String ticket) throws Exception {
long deadline = System.currentTimeMillis() + 3000;
while (!svc.isCompletionStampedForTest(ticket)) {
assertTrue(System.currentTimeMillis() < deadline,
"completedNanos for " + ticket + " was never stamped");
Thread.sleep(5);
}
}
/**
* fleetd #418: waits until {@code turnId} actually appears in {@code pushLoop}'s own open-question
* state for {@code lead} — {@link ReplyPushLoop#pendingQuestionTurnIdsForTest} — rather than until
* {@link MessageService#poll} reports {@link MessageService.Phase#ASKING}. {@code Phase.ASKING} is
* set by {@code markAsyncQuestion}, the FIRST of three steps {@code MessageService.ask()} performs;
* {@code pushLoop.onQuestionOpened} (the one that actually publishes to {@code pendingQuestions})
* is the THIRD. Under load the asker thread can be descheduled between those two steps, so a
* barrier on the phase alone can release before the push loop has anything pending — a test that
* then asserts on {@link ReplyPushLoop#decide} is asserting on state that has not been published
* yet, not on the throwing path it is named for.
*/
private void awaitQuestionPendingOn(ReplyPushLoop pushLoop, String lead, String turnId) throws Exception {
long deadline = System.currentTimeMillis() + 3000;
while (!pushLoop.pendingQuestionTurnIdsForTest(lead).contains(turnId)) {
assertTrue(System.currentTimeMillis() < deadline,
"turnId " + turnId + " never appeared in the push loop's pending questions for " + lead);
Thread.sleep(5);
}
}
// --- CB-640: fleet health evidence accessors --------------------------------------------
@Test
@@ -90,63 +90,6 @@ class PlacementPolicyTest {
assertTrue(e.getMessage().contains("quarantined"), e.getMessage());
}
// --- fleetd #422: FixedPlacementPolicy must consult modelOff too, at BOTH filter sites ------
/** The default-profile fast path (:44) must skip a default whose model is off. */
@Test
void fixedSkipsModelOffDefault() {
PlacementPolicy policy = PlacementPolicies.fixed();
PlacementContext ctx = new PlacementContext("b",
List.of(PlacementCandidate.profile("a"), PlacementCandidate.profile("b")),
noSessions(), Set.of(), Set.of(), Set.of(), Set.of("b"));
assertEquals("a", policy.select(ctx).profile(),
"the default 'b' names an off model, so fixed falls through to the first available candidate");
}
/**
* The fallback walk (:48-53) must skip an off-model candidate too — exercised independently of
* the default-profile fast path by using no default at all, so this is the only filter that runs.
*/
@Test
void fixedFallbackWalkSkipsModelOffCandidate() {
PlacementPolicy policy = PlacementPolicies.fixed();
PlacementContext ctx = new PlacementContext(null,
List.of(PlacementCandidate.profile("a"), PlacementCandidate.profile("b")),
noSessions(), Set.of(), Set.of(), Set.of(), Set.of("a"));
assertEquals("b", policy.select(ctx).profile(),
"candidate 'a' names an off model, so the fallback walk skips it and picks 'b'");
}
@Test
void fixedThrowsWhenDefaultAndEveryCandidateModelOff() {
PlacementPolicy policy = PlacementPolicies.fixed();
PlacementContext ctx = new PlacementContext("b",
List.of(PlacementCandidate.profile("a"), PlacementCandidate.profile("b")),
noSessions(), Set.of(), Set.of(), Set.of(), Set.of("a", "b"));
PlacementException e = assertThrows(PlacementException.class, () -> policy.select(ctx));
assertTrue(e.getMessage().contains("turned off"),
"message names model-off as the cause: " + e.getMessage());
assertFalse(e.getMessage().contains("quarantined"), "must not read like quarantine: " + e.getMessage());
assertFalse(e.getMessage().contains("cooling off"), "must not read like cool-off: " + e.getMessage());
assertFalse(e.getMessage().contains("weight 0"), "must not read like weight-0: " + e.getMessage());
}
/**
* Quarantine still wins when a profile is both quarantined and model-off (mirrors {@code
* fixedThrowsWhenDefaultAndEveryCandidateQuarantined}'s priority over cooling off).
*/
@Test
void fixedReportsQuarantineNotModelOffWhenBothApply() {
PlacementPolicy policy = PlacementPolicies.fixed();
PlacementContext ctx = new PlacementContext("b",
List.of(PlacementCandidate.profile("a"), PlacementCandidate.profile("b")),
noSessions(), Set.of(), Set.of("a", "b"), Set.of(), Set.of("a", "b"));
PlacementException e = assertThrows(PlacementException.class, () -> policy.select(ctx));
assertTrue(e.getMessage().contains("quarantined"), e.getMessage());
assertFalse(e.getMessage().contains("turned off"),
"quarantine takes priority over model-off in the message: " + e.getMessage());
}
@Test
void roundRobinCyclesThroughAvailableProfiles() {
PlacementPolicy policy = PlacementPolicies.roundRobin();
@@ -461,93 +404,6 @@ class PlacementPolicyTest {
assertTrue(e.getMessage().contains("weight 0"), e.getMessage());
}
// --- fleetd #435: FixedPlacementPolicy must consult maxLoad too, at BOTH filter sites -------
/**
* The default-profile fast path must skip a capped default. Measured on 7667727 before this
* fix: a single dev profile at {@code maxLoad: 1} with 1 live, under {@code fixed()}, still
* returned "SPAWNED on profile=a" — the cap was advisory for every unqualified spawn.
*/
@Test
void fixedSkipsCappedDefault() {
PlacementPolicy policy = PlacementPolicies.fixed();
PlacementContext ctx = new PlacementContext("b",
List.of(PlacementCandidate.profile("a", 1.0f, null),
PlacementCandidate.profile("b", 1.0f, 1)),
name -> "b".equals(name) ? 1 : 0, Set.of(), Set.of(), Set.of());
assertEquals("a", policy.select(ctx).profile(),
"the default 'b' is at its maxLoad cap, so fixed falls through to the free candidate 'a'");
}
/**
* The fallback walk must skip a capped candidate too — exercised independently of the
* default-profile fast path by using no default at all, so this is the only filter that runs.
*/
@Test
void fixedFallbackWalkSkipsCappedCandidate() {
PlacementPolicy policy = PlacementPolicies.fixed();
PlacementContext ctx = new PlacementContext(null,
List.of(PlacementCandidate.profile("a", 1.0f, 1),
PlacementCandidate.profile("b", 1.0f, null)),
name -> "a".equals(name) ? 1 : 0, Set.of(), Set.of(), Set.of());
assertEquals("b", policy.select(ctx).profile(),
"candidate 'a' is at its maxLoad cap, so the fallback walk skips it and picks 'b'");
}
@Test
void fixedThrowsWhenDefaultAndEveryCandidateAtCap() {
PlacementPolicy policy = PlacementPolicies.fixed();
PlacementContext ctx = new PlacementContext("b",
List.of(PlacementCandidate.profile("a", 1.0f, 1),
PlacementCandidate.profile("b", 1.0f, 1)),
_ -> 1, Set.of(), Set.of(), Set.of());
PlacementException e = assertThrows(PlacementException.class, () -> policy.select(ctx));
assertTrue(e.getMessage().contains("maxLoad"), "message names the cap: " + e.getMessage());
assertTrue(e.getMessage().contains("1 cap"), "message names the cap value: " + e.getMessage());
assertFalse(e.getMessage().contains("quarantined"), "must not read like quarantine: " + e.getMessage());
assertFalse(e.getMessage().contains("turned off"), "must not read like model-off: " + e.getMessage());
}
/** The mirror: an uncapped default is still chosen, so the new term cannot exclude everything. */
@Test
void fixedStillReturnsUncappedDefault() {
PlacementPolicy policy = PlacementPolicies.fixed();
PlacementContext ctx = new PlacementContext("b",
List.of(PlacementCandidate.profile("a", 1.0f, null),
PlacementCandidate.profile("b", 1.0f, null)),
noSessions(), Set.of(), Set.of(), Set.of());
assertEquals("b", policy.select(ctx).profile(), "an uncapped default is returned unconditionally");
}
/** CB-585: an explicit {@code maxLoad: 0} on the default caps it at zero live members. */
@Test
void fixedSkipsMaxLoadZeroDefaultEvenWithZeroLiveWorkers() {
PlacementPolicy policy = PlacementPolicies.fixed();
PlacementContext ctx = new PlacementContext("b",
List.of(PlacementCandidate.profile("a", 1.0f, null),
PlacementCandidate.profile("b", 1.0f, 0)),
_ -> 0, Set.of(), Set.of(), Set.of());
assertEquals("a", policy.select(ctx).profile(),
"the default 'b' has maxLoad 0, so it is already at its cap with nobody live");
}
/**
* Quarantine still wins when a profile is both quarantined and at cap (mirrors {@code
* fixedReportsQuarantineNotModelOffWhenBothApply}'s priority over model-off).
*/
@Test
void fixedReportsQuarantineNotAtCapWhenBothApply() {
PlacementPolicy policy = PlacementPolicies.fixed();
PlacementContext ctx = new PlacementContext("b",
List.of(PlacementCandidate.profile("a", 1.0f, 1),
PlacementCandidate.profile("b", 1.0f, 1)),
_ -> 1, Set.of(), Set.of("a", "b"), Set.of());
PlacementException e = assertThrows(PlacementException.class, () -> policy.select(ctx));
assertTrue(e.getMessage().contains("quarantined"), e.getMessage());
assertFalse(e.getMessage().contains("maxLoad"),
"quarantine takes priority over at-cap in the message: " + e.getMessage());
}
@Test
void unknownPolicyNameThrows() {
assertThrows(IllegalArgumentException.class, () -> PlacementPolicies.fromName("random"));