Compare commits
94 Commits
| Author | SHA1 | Date | |
|---|---|---|---|
| 9e4e423ad6 | |||
| d4a2cd720c | |||
| 1fb6176783 | |||
| 49df79203c | |||
| 235644c0f0 | |||
| 7772b41993 | |||
| eccd0548ce | |||
| c4e23eebad | |||
| 7df503dfc2 | |||
| f5e02fedd6 | |||
| 29e7a06c49 | |||
| 92a96fcbd8 | |||
| 13482872bb | |||
| c1ca6273fc | |||
| 5f1b260c81 | |||
| 30d6872779 | |||
| 9d1306d442 | |||
| e54e3d87ea | |||
| bf027f10b9 | |||
| 3c5873dfe2 | |||
| e29227d5f4 | |||
| 45aca9eb3e | |||
| 2af13ab1ff | |||
| 0788d84be8 | |||
| 20c1094cbf | |||
| 9011c59b9f | |||
| c11ad71ed0 | |||
| bdcf285265 | |||
| d4f93a7b13 | |||
| cfebc575ea | |||
| 822327eed5 | |||
| 5289eb509f | |||
| e4c703a51a | |||
| 703a05db41 | |||
| 3f036b2a62 | |||
| 82fae94c55 | |||
| b1f34c2e6b | |||
| 3f807d9f1b | |||
| e70263062c | |||
| 1a1e586b62 | |||
| c16d118f09 | |||
| 4b10d02207 | |||
| c5fbfdbf4a | |||
| 12cff28abb | |||
| 77a6a7142e | |||
| 9f3671b801 | |||
| 84034b34d1 | |||
| 1e9b2c9b7e | |||
| 5d422f85fa | |||
| ed2027b202 | |||
| 6b0a99b2b7 | |||
| 6b7caba248 | |||
| f8b0d42a5c | |||
| d1e7d71eee | |||
| 7fd914df1a | |||
| b066eb1903 | |||
| a196d34455 | |||
| 051d320ea0 | |||
| 766772763f | |||
| eab8185d7b | |||
| 7f672f0fb8 | |||
| e2801b9bbc | |||
| ea02c7b248 | |||
| ce74e164c6 | |||
| e60f892efd | |||
| ce05886831 | |||
| be123d0ac7 | |||
| 2d09c8b027 | |||
| ed54f0224e | |||
| 8d5bc3ee89 | |||
| ab0cc71aa4 | |||
| 4cd9046353 | |||
| 4e98a74047 | |||
| a502ba53e0 | |||
| 446cc11d1d | |||
| ddd81fe174 | |||
| ae74cc081f | |||
| cf8da1d5fa | |||
| b8aedeafcb | |||
| 3d61af6f6f | |||
| ed99c209ac | |||
| af4c88d54b | |||
| 44c735f6f5 | |||
| b540a1744b | |||
| 0f51d53098 | |||
| c670792ffe | |||
| 3982ace544 | |||
| 7180b1aad0 | |||
| c26f695402 | |||
| 48877315ca | |||
| 5c08054533 | |||
| b6db9c31f5 | |||
| e7b33fe3a0 | |||
| e3e403e5c8 |
+10
-4
@@ -87,12 +87,18 @@ jobs:
|
||||
apt-get update && apt-get install -y --no-install-recommends maven
|
||||
mvn -version
|
||||
|
||||
# The `contract` profile clears the default-excludes group, so the @Tag("contract") AMQP test
|
||||
# runs against the RabbitMQ service container (AMQP_URI). Pinned to the one contract test to
|
||||
# avoid re-running the unit suite already covered by the `build` job.
|
||||
# The `contract` profile clears the default-excludes group, so `-Dgroups=contract` runs every
|
||||
# @Tag("contract") test and nothing from the unit suite the `build` job already covered — a
|
||||
# tag selects the whole group, so a test added to it later runs here automatically. A prior
|
||||
# version of this step pinned `-Dtest=AmqpReplyInboxContractTest` by class name instead: that
|
||||
# silently excluded every other contract test (including the herdr ones) from CI, and nobody
|
||||
# noticed until the herdr protocol drifted out from under a test that never ran here
|
||||
# (fleetd #449). If this runner has no herdr socket, the herdr-backed tests in the group
|
||||
# skip on their own `assumeTrue` and only the broker-backed ones actually run — check the
|
||||
# step output rather than assuming which.
|
||||
- name: Contract tests
|
||||
working-directory: fleetd
|
||||
run: mvn -B -Pcontract test -Dtest=AmqpReplyInboxContractTest
|
||||
run: mvn -B -Pcontract test -Dgroups=contract
|
||||
|
||||
- name: Failing test output
|
||||
if: failure()
|
||||
|
||||
@@ -58,8 +58,9 @@ and the sender silently receives nothing. Fail toward the recoverable error.
|
||||
re-send because a call looks slow — the bridge delivers when the peer is `idle`, `blocked` or
|
||||
`done`. A spawned member must **also** have mounted the bridge MCP: until it has, it is not
|
||||
deliverable, and a send waits on that gate for ~60s and then fails without ever reaching its pane.
|
||||
5. **Never drive the terminal multiplexer directly** (no `herdr` CLI, no socket). The bridge owns
|
||||
policy; the multiplexer owns PTYs. Going around the bridge bypasses every rule above.
|
||||
5. **Never move a fleet session, pane or peer except through the bridge.** The bridge owns policy;
|
||||
the multiplexer owns PTYs. Any route that changes fleet state without the bridge's checks
|
||||
bypasses every rule above — the `herdr` CLI and its socket are the usual example.
|
||||
|
||||
### Primary (lead) — run this on every task, in order
|
||||
|
||||
@@ -137,7 +138,8 @@ the merge — and merging on a reviewer's word is delegating it by proxy.
|
||||
| Answer a member's `fleet_ask` | `fleet_send{turnId, content}` — **not** `sessionId` |
|
||||
| Message a **peer lead** on this host | `fleet_send{sessionId: <their terminal>, content}` — `fleet_list` → `leads` reports it. Coordination only, **never** a task |
|
||||
| Message a **peer lead** on another daemon or host | `fleet_send{coordId: <their coord-id>, content}` — needs a `coordinator:` block; your own coord-id is in `fleet_list`. Coordination only, **never** a task |
|
||||
| Answer a peer lead that messaged you | `fleet_reply{content}` — the one case a lead replies |
|
||||
| Answer a peer lead that messaged you | `fleet_send{coordId}` — or `{sessionId}` if they are on this host. **Not** `fleet_reply`: it has no peer route and the publish is refused |
|
||||
| Read your own held lead-to-lead mail (no ack) | `fleet_poll{coordId: <your own coord-id, from fleet_list's coordinator.selfId>}` — primary-only; never acks, so `fleet_list`'s `held[]` still shows it after. `fleet_list`'s `held[]` gives only a truncated preview — this is the only way to read the full body |
|
||||
| Collect a held reply | `fleet_poll{target}` · then `fleet_ack{target, msgId}` |
|
||||
| Tear down a member | `fleet_stop{paneId}` |
|
||||
|
||||
@@ -163,15 +165,21 @@ The traffic between leads is coordination and nothing else:
|
||||
3. **Verify a peer exactly as you verify yourself.** Peer status buys nothing: check the claim
|
||||
against the code, and re-run the build. A peer's correction gets the same treatment — right or
|
||||
wrong on the evidence, not on who said it. Neither of you merges the other's work unreviewed.
|
||||
**N observations are N data points only if they differ in the axis you are trusting.** This cuts
|
||||
both ways. N *failures* blamed on one cause are one data point when the cases share what you are
|
||||
not varying. N *agreeing measurements* are also one data point when they share an instrument —
|
||||
two hosts, two operators and the same formula is one formula, not two confirmations.
|
||||
4. **Ask a peer to read your project addendum.** Your addendum is instruction surface: every future
|
||||
session on your host obeys it, and a wrong one is obeyed just as faithfully as a right one. The
|
||||
author is the worst reader of their own qualifier placement — measured here, one addendum carried
|
||||
two defects and a non-author found both. If you have no peer, at least re-read it asking "which
|
||||
sentence goes false first, and would a reader reach the caveat before acting?"
|
||||
|
||||
Being messaged by a peer does not make you its worker: answer with `fleet_reply`, and push back on
|
||||
the substance if it is wrong. A peer that simply complies has thrown away the reason there are two of
|
||||
you.
|
||||
Being messaged by a peer does not make you its worker: answer the way you would open —
|
||||
`fleet_send{coordId}` for another daemon, `fleet_send{sessionId}` on this host — and push back on
|
||||
the substance if it is wrong. `fleet_reply` resolves a member's blocked `fleet_send`; a peer's
|
||||
coord-id message is durable and non-blocking, so there is nothing for it to resolve. A peer that
|
||||
simply complies has thrown away the reason there are two of you.
|
||||
|
||||
### Member (worker or architect) — the turn contract
|
||||
|
||||
@@ -216,6 +224,19 @@ must obey belongs in the charter, not here.
|
||||
- **This repo is the bridge.** The daemon is `fleetd`, its MCP mount is `http://127.0.0.1:8765/mcp`,
|
||||
and the code behind the rules above is `mcp/FleetMcp` (tools), `auth/Authz` (the role table),
|
||||
`mcp/ConnectionIdentity` (connection→role), and `worker/*Launcher` (`REPLY_CHARTER`).
|
||||
- **Herdr socket tests (measured 2026-09-10).** In this repo, herdr is a subject under test. A
|
||||
worker assigned to herdr code, and the lead, may let a test open the herdr socket directly in a
|
||||
throwaway workspace that the test tears down. This only covers
|
||||
`fleetd/src/test/java/dev/ltms/fleet/herdr/AgentControlContractTest.java`,
|
||||
`fleetd/src/test/java/dev/ltms/fleet/herdr/HerdrContractTest.java`,
|
||||
`fleetd/src/test/java/dev/ltms/fleet/herdr/PaneLocatorContractTest.java`, and
|
||||
`fleetd/src/test/java/dev/ltms/fleet/herdr/WorkspacePlacementContractTest.java`. It is not a
|
||||
general licence. Using the herdr CLI or socket to move a real fleet session, pane, or peer stays
|
||||
banned. That is the control plane that invariant 5 protects. Re-measure with
|
||||
`grep -rl 'UnixSocketHerdrClient.connect()' fleetd/src/test/java --include='*.java'`. A non-empty
|
||||
result means tests still open the socket and this note still applies. An empty result means nobody
|
||||
does this any more; delete this section. Canonical invariant 5 restatement is tracked in #458 and
|
||||
is not part of this change.
|
||||
- **`fleet_profiles`/`fleet_list` report two separate outage states, and they are not the same
|
||||
thing.** *Quarantined* (CB-578) means the backend told us it is out of capacity — a long,
|
||||
1800s-default cooldown. *Cooling off* (fleetd #201/#227) means a profile's credential threw two
|
||||
@@ -372,6 +393,16 @@ print("in sync:", w[i:w.index("\n```\n", i) + 1] == block)
|
||||
PY
|
||||
```
|
||||
|
||||
**Only the lead can run that check (measured 2026-09-10).** A member's provisioned worktree has
|
||||
`wiki/` uninitialized, so the script dies with `FileNotFoundError: wiki/7-Use-Cases.md`. Measured
|
||||
in three worker worktrees: `git submodule status` printed a leading `-` and `wiki/` held 0
|
||||
entries; the primary's own clone printed a leading `+` and the file was there. So never make this
|
||||
check a member's acceptance criterion — it is unsatisfiable for them, and a brief that asks for it
|
||||
is asking a worker to invent a pass. A member told to check it must say it could not run it, and
|
||||
must never report it as passed. The lead runs it in the main clone before merging. Re-measure with
|
||||
`git submodule status` in a member's worktree: a leading `-` means this still applies; once it
|
||||
prints a commit with no `-`, delete this paragraph.
|
||||
|
||||
## IDE MCP tools & validation workflow (enforced)
|
||||
|
||||
> **Primary only.** Workers have no IDE MCP mount — if you are a worker, skip this section and
|
||||
|
||||
+61
-17
@@ -207,8 +207,12 @@ herdrSocket: ~/.config/herdr/herdr.sock
|
||||
# omit and this profile's completion fallback behaves exactly as before.
|
||||
# Every backend words its refusal differently, so this is config, never a
|
||||
# vendor string baked into fleetd itself.
|
||||
# DEFERRED: compiled once into a startup pattern map — editing it needs a
|
||||
# daemon restart, same as this profile's model/baseUrl/argv.
|
||||
# HOT (fleetd #446): read live, cached by profile name, at every
|
||||
# completion-fallback check AND by fleet_profiles' exhaustionDetectionArmed —
|
||||
# editing it and reloading arms or disarms usage-limit detection for this
|
||||
# profile with no daemon restart. (Before fleetd #446 this was DEFERRED,
|
||||
# compiled once into a startup pattern map like model/baseUrl/argv still are —
|
||||
# see errorPattern below, which is still deferred that way on purpose.)
|
||||
# credentialId → CB-578 stage B: the credential this profile quarantines WITH when a
|
||||
# BACKEND_EXHAUSTED classification fires. Two profiles that set the SAME
|
||||
# credentialId share one quarantine — the case this exists for is two models
|
||||
@@ -227,8 +231,9 @@ herdrSocket: ~/.config/herdr/herdr.sock
|
||||
# happens, just without a profile-specific match; every backend words its
|
||||
# failure differently, so a hardcoded sentence would only ever match one
|
||||
# of them.
|
||||
# DEFERRED: compiled once into a startup pattern map, same as exhaustedPattern
|
||||
# — editing it needs a daemon restart.
|
||||
# DEFERRED: compiled once into a startup pattern map — editing it needs a
|
||||
# daemon restart. Unlike exhaustedPattern above (made hot by fleetd #446),
|
||||
# errorPattern was scoped out of that ticket on purpose and stays deferred.
|
||||
# # errorPattern: "503 Service Unavailable" # opt-in: classify a backend outage
|
||||
#
|
||||
# What happens once a match fires (BackendOutagePolicy, credentialId-keyed,
|
||||
@@ -468,9 +473,10 @@ placement: weighted
|
||||
# when the reload happens — not about how important the key is:
|
||||
# HOT → takes effect on the next spawn: the whole `fleet:` block (every role pool,
|
||||
# `charters`, and `tabLabel`), `placement:`, and an existing profile's weight / maxLoad
|
||||
# / credentialId. Those are hot because the placement policy (and, for credentialId,
|
||||
# the CB-578 stage B quarantine check) reads them through a supplier — being config is
|
||||
# not by itself enough to make a key hot.
|
||||
# / credentialId / exhaustedPattern. Those are hot because the placement policy (and,
|
||||
# for credentialId, the CB-578 stage B quarantine check; for exhaustedPattern, fleetd
|
||||
# #446's LiveExhaustedPatterns) reads them through a supplier — being config is not by
|
||||
# itself enough to make a key hot.
|
||||
# EXCEPT `fleet.leaders`: Fleetd.main reads it once at startup to build the lead tab
|
||||
# scanner and launcher, and neither is rebuilt on reload. A changed/added/removed
|
||||
# `fleet.leaders` entry is silently accepted — the reload reports "config reloaded"
|
||||
@@ -482,10 +488,10 @@ placement: weighted
|
||||
# stage B — baked once into the quarantine tracker built at startup), ADDING or
|
||||
# REMOVING a profile (a new backend needs its own launcher, and launchers are built
|
||||
# once), AND an existing profile's launch settings — model, baseUrl, argv, env,
|
||||
# configDir, mcpUrl, tabLabel, exhaustedPattern, errorPattern (fleetd #201 / #227 —
|
||||
# compiled once into a startup pattern map the same way exhaustedPattern is). The
|
||||
# launcher takes a copy of `profiles:` at startup and resolves every spawn out of
|
||||
# that copy, so those never
|
||||
# configDir, mcpUrl, tabLabel, errorPattern (fleetd #201 / #227 — compiled once into
|
||||
# a startup pattern map; exhaustedPattern used to be compiled the same way until
|
||||
# fleetd #446 made it hot — see above). The launcher takes a copy of `profiles:` at
|
||||
# startup and resolves every spawn out of that copy, so those never
|
||||
# reach a launch until you restart. The reload logs them by name rather than
|
||||
# pretending they applied.
|
||||
# COLD → cannot change at all: `bind:`, `herdrSocket:`, `broker:` and `auth:`. The socket is
|
||||
@@ -787,12 +793,21 @@ guard:
|
||||
# .claude/skills/, so a member spawned against ANY repo — not only one that already ships its own
|
||||
# copy — can load a bridge skill (e.g. implementer). Unset (the default): no worktree is touched
|
||||
# beyond today's behaviour. A skill folder the target repo already carries under
|
||||
# .claude/skills/<name> is never overwritten — the repo's own copy always wins. Claude Code
|
||||
# members only; an opencode member reads a different path (.opencode/agent) this key does not
|
||||
# touch. Best-effort like worktreeGroup above: a missing/unreadable directory here is logged and
|
||||
# skipped, never a failed spawn. Every non-hidden subdirectory of this directory is copied
|
||||
# wholesale, with no per-file allowlist — don't park scratch files or drafts alongside the real
|
||||
# skill folders, they will be copied into every provisioned worktree too.
|
||||
# .claude/skills/<name> is never overwritten — the repo's own copy always wins. Best-effort like
|
||||
# worktreeGroup above: a missing/unreadable directory here is logged and skipped, never a failed
|
||||
# spawn. Every non-hidden subdirectory of this directory is copied wholesale, with no per-file
|
||||
# allowlist — don't park scratch files or drafts alongside the real skill folders, they will be
|
||||
# copied into every provisioned worktree too.
|
||||
#
|
||||
# fleetd #393: which member KINDS actually consume this once it is copied. kind: claude-code —
|
||||
# the Claude Code CLI discovers .claude/skills/ on its own; nothing else is needed. kind: opencode
|
||||
# — opencode has no such discovery, so OpenCodeLauncher reads whatever landed under
|
||||
# .claude/skills/ and appends each seeded skill's SKILL.md to the generated instructions[] file
|
||||
# (opencode's only channel for static guidance text; unlike Claude Code's Skill tool, the content
|
||||
# is always part of the system prompt, not loaded on demand). Both kinds are covered as of #393 —
|
||||
# earlier builds copied the files for every kind but only claude-code could read them, and the
|
||||
# seeding log said "N of M" regardless. Check the per-spawn launcher log (not just the seeding
|
||||
# log) to see what a given member actually got.
|
||||
# memberSkills: /path/to/fleetd/checkout/.claude/skills
|
||||
|
||||
# Session lifecycle limits (CB-303). All knobs are opt-in; omit or set to null to keep
|
||||
@@ -877,3 +892,32 @@ guard:
|
||||
# terminal: term_65619bd6174568
|
||||
# pushReminders: 5
|
||||
# pushBackoffMs: 15000
|
||||
|
||||
# Central allow-list of models any profiles: entry may name. Nothing checked a profile's model:
|
||||
# value before this block existed — it was a free-form string handed straight to the backend
|
||||
# adapter, and a withdrawn or misspelled name failed silently instead of at config load (opencode
|
||||
# falls back to a default model rather than erroring on an unknown -m).
|
||||
#
|
||||
# Absent, or present with an empty allow:, is OFF: no profile's model: is checked, exactly like
|
||||
# before this block existed. fleetd.yaml is gitignored on every host, so an upgrade must not force
|
||||
# every operator to enumerate their models before the daemon will start.
|
||||
#
|
||||
# The list is the authority; profiles: is checked against it, never the reverse — adding or
|
||||
# editing a profiles: entry cannot, by itself, widen what is permitted here.
|
||||
#
|
||||
# Enforcement is at CONFIG LOAD only (a bad model: fails the daemon at startup, naming both the
|
||||
# model and the profile). There is no spawn-time enforcement, no runtime on/off switch, and no
|
||||
# interaction with BackendQuarantine — those are separate, later units.
|
||||
#
|
||||
# allow → the permitted models. Each entry is its own block (not a bare string) so a later unit
|
||||
# can add an on/off state or a load limit per model without changing this shape.
|
||||
# model → the model id exactly as a profiles: entry's model: field would write it. One flat,
|
||||
# opaque-string namespace: a bare Claude id (claude-sonnet-5) and an opencode
|
||||
# provider-prefixed id (openai/gpt-5.6-terra) both fit here unchanged — the check is a
|
||||
# plain string match, never a parse of the provider prefix or a branch on kind:.
|
||||
# models:
|
||||
# allow:
|
||||
# - model: claude-sonnet-5
|
||||
# - model: claude-opus-5
|
||||
# - model: openai/gpt-5.6-terra
|
||||
# - model: amazon.nova-pro-v1:0
|
||||
|
||||
@@ -18,6 +18,7 @@ import dev.ltms.fleet.inject.BackendErrorSink;
|
||||
import dev.ltms.fleet.inject.CompletionResolver;
|
||||
import dev.ltms.fleet.inject.ExhaustedPatternLookup;
|
||||
import dev.ltms.fleet.inject.ExhaustionSink;
|
||||
import dev.ltms.fleet.inject.LiveExhaustedPatterns;
|
||||
import dev.ltms.fleet.inject.Injector;
|
||||
import dev.ltms.fleet.inject.StatusPoller;
|
||||
import dev.ltms.fleet.inject.TurnListener;
|
||||
@@ -65,10 +66,13 @@ import java.nio.file.Files;
|
||||
import java.nio.file.Path;
|
||||
import java.util.ArrayList;
|
||||
import java.util.LinkedHashMap;
|
||||
import java.util.LinkedHashSet;
|
||||
import java.util.List;
|
||||
import java.util.Map;
|
||||
import java.util.Optional;
|
||||
import java.util.Set;
|
||||
import java.util.TreeSet;
|
||||
import java.util.concurrent.ConcurrentHashMap;
|
||||
import java.util.concurrent.Executors;
|
||||
import java.util.concurrent.ScheduledExecutorService;
|
||||
import java.util.concurrent.TimeUnit;
|
||||
@@ -77,6 +81,7 @@ import java.util.function.Function;
|
||||
import java.util.function.Predicate;
|
||||
import java.util.function.Supplier;
|
||||
import java.util.regex.Pattern;
|
||||
import java.util.stream.Collectors;
|
||||
|
||||
/**
|
||||
* {@code fleetd} entry point. Wires the real herdr socket client to the REST app and
|
||||
@@ -134,6 +139,10 @@ public final class Fleetd {
|
||||
// secret is reported above, so upgrading past this commit never silently drops CB-592's
|
||||
// protection.
|
||||
reportMemberCredentialsGap(cfg);
|
||||
// fleetd #395: an unset exhaustedPattern is a silent opt-out of usage-limit detection for
|
||||
// that profile — say so loudly, the same way the two reports above do, rather than let an
|
||||
// operator discover it only when a limit goes undetected.
|
||||
reportExhaustedPatternGap(cfg);
|
||||
// CB-559: `cfg` stays the startup snapshot — every validation and every piece of one-time
|
||||
// wiring below reads it, and must, because those decisions cannot be unmade. `config` is the
|
||||
// live reference the hot paths read per use. Which keys can actually move is ConfigRef's
|
||||
@@ -144,19 +153,16 @@ public final class Fleetd {
|
||||
SubscriptionGuard guard = new SubscriptionGuard(cfg.guard().hostSet());
|
||||
guard.assertPrimaryClean(System.getenv());
|
||||
|
||||
// CB-501: refuse to start if the bind is wider than the auth mode can defend. Under
|
||||
// loopback-trust, "not a known worker" means "the primary" — sound only because the OS
|
||||
// refuses remote connections to a loopback socket. This throws rather than warns so the
|
||||
// dangerous configuration cannot be reached by ignoring a log line.
|
||||
cfg.validateAuthExposure();
|
||||
cfg.validateLeadTabPrefixes();
|
||||
// CB-542: a subscription:true profile whose env: reseats ANTHROPIC_BASE_URL/AUTH_TOKEN would
|
||||
// reach an unguarded endpoint (the launcher skips SubscriptionGuard for it). Refuse at load.
|
||||
cfg.validateSubscriptionProfiles();
|
||||
cfg.validateCharters();
|
||||
// CB-548: every architect slot must name a configured workers: profile — the strong-model
|
||||
// backend the future spawn lifecycle would read. A stale reference dies here, not later.
|
||||
cfg.validateMembers();
|
||||
// Every FleetConfig.validateXxx() the operator's config can fail — CB-501's auth-exposure
|
||||
// check, CB-531's lead-tab-prefix check, CB-542's subscription-profile check, the charter
|
||||
// and member-slot checks, and the "central allow-list of usable models" check — must run
|
||||
// here, at load, before anything below opens a socket or spawns a member. fleetd ticket
|
||||
// "central allow-list of usable models" follow-up: mutation testing found six individual
|
||||
// calls here with nothing proving any of them still ran (deleting one left the full suite
|
||||
// green). validateAll() replaces them with the one call that FleetConfigValidateAllTest
|
||||
// and the Fleetd-startup tests actually pin — see FleetConfig#validateAll's javadoc for
|
||||
// why a name-by-name list here would have the same defect it replaces.
|
||||
cfg.validateAll();
|
||||
|
||||
Path socket = cfg.herdrSocket() != null && !cfg.herdrSocket().isBlank()
|
||||
? Path.of(cfg.herdrSocket())
|
||||
@@ -230,6 +236,13 @@ public final class Fleetd {
|
||||
profileName -> liveCountRef.get().apply(profileName),
|
||||
quarantine,
|
||||
outagePolicy);
|
||||
// fleetd #422 follow-up: say which of the three model-gate states the daemon booted into —
|
||||
// no models: block at all, a block armed with nothing off, or a block with N off — the same
|
||||
// way exhaustedPatternCoverageLine/errorPatternCoverageLine report CB-578 stage A/fleetd
|
||||
// #201 Unit 5 coverage just below. Read from workers.modelGateState() (never a separate
|
||||
// config.get().models() here) so this line and fleet_profiles' modelGateArmed can never
|
||||
// disagree about what CompositePeerLauncher's spawn gate actually enforces.
|
||||
log.info("model gate (fleetd #422): {}", modelGateCoverageLine(workers.modelGateState()));
|
||||
// CB-504: under supervision (launchd/systemd) fleetd can start before herdr's socket
|
||||
// exists. The client itself is lazy — it connects per call — but the orphan reap below is
|
||||
// the first thing that actually talks to herdr, so without this wait a boot-order race
|
||||
@@ -351,7 +364,11 @@ public final class Fleetd {
|
||||
// pane resolves to an architect until the later spawn lifecycle binds one. The registry is
|
||||
// what CallerResolver resolves against and what that lifecycle will read profiles from;
|
||||
// nothing here spawns a slot.
|
||||
MemberRegistry members = new MemberRegistry(cfg.fleet());
|
||||
// fleetd #424: MemberRegistry.live re-reads fleet.architects through `config` on every
|
||||
// reserve/requireSlotFor call, so a reload that removes or adds an architect slot governs
|
||||
// the next spawn with no restart — the frozen `new MemberRegistry(cfg.fleet())` this used
|
||||
// to be let a "revoked" slot keep granting new architect spawns forever.
|
||||
MemberRegistry members = MemberRegistry.live(() -> config.get().fleet());
|
||||
sessions.setMemberLifecycle(members);
|
||||
if (!members.slots().isEmpty()) {
|
||||
log.info("member slots: {} configured {} — none bound yet (a slot is idle until the "
|
||||
@@ -365,29 +382,34 @@ public final class Fleetd {
|
||||
Rendezvous rendezvous = new Rendezvous();
|
||||
// CB-578 stage A: classify a completion-fallback scrape that matches a profile's configured
|
||||
// usage-limit refusal as BACKEND_EXHAUSTED rather than handing it back as a real answer.
|
||||
// Compiled once at startup, keyed by profile name; a profile with no exhaustedPattern is
|
||||
// simply absent here, so its workers keep today's completion-fallback behaviour unchanged.
|
||||
Map<String, Pattern> exhaustedPatternsByProfile = new LinkedHashMap<>();
|
||||
cfg.profiles().forEach((name, profile) -> {
|
||||
if (profile.hasExhaustedPattern()) {
|
||||
exhaustedPatternsByProfile.put(name, Pattern.compile(profile.exhaustedPattern()));
|
||||
}
|
||||
});
|
||||
// fleetd #446: read live off `config` per lookup, cached by profile name — see
|
||||
// LiveExhaustedPatterns's class doc for why this replaces the old compiled-once-at-startup
|
||||
// map. A profile with no exhaustedPattern simply returns null here, so its workers keep
|
||||
// today's completion-fallback behaviour unchanged.
|
||||
LiveExhaustedPatterns liveExhaustedPatterns = new LiveExhaustedPatterns(() -> config.get().profiles());
|
||||
ExhaustedPatternLookup exhaustedPatterns = target -> sessions.roster().stream()
|
||||
.filter(session -> target.equals(session.terminalId()))
|
||||
.findFirst()
|
||||
.map(session -> exhaustedPatternsByProfile.get(session.profile()))
|
||||
.map(session -> liveExhaustedPatterns.patternFor(session.profile()))
|
||||
.orElse(null);
|
||||
// The startup coverage line still reports the boot-time snapshot only — it is printed once,
|
||||
// here, and a reload no longer needs to change what it said; exhaustionDetectionArmed (via
|
||||
// liveExhaustedPatterns.armed, wired into quarantineSource below) is what stays live.
|
||||
Set<String> exhaustedConfiguredAtStartup = cfg.profiles().entrySet().stream()
|
||||
.filter(e -> e.getValue().hasExhaustedPattern())
|
||||
.map(Map.Entry::getKey)
|
||||
.collect(Collectors.toCollection(LinkedHashSet::new));
|
||||
log.info("backend-exhausted classification (CB-578 stage A): {}",
|
||||
CompletionResolver.coverage("exhaustedPattern", cfg.profiles().keySet(),
|
||||
exhaustedPatternsByProfile.keySet()));
|
||||
exhaustedPatternCoverageLine(cfg.profiles().keySet(), exhaustedConfiguredAtStartup));
|
||||
// fleetd #201 Unit 5: classify a completion-fallback scrape that matches a profile's
|
||||
// configured backend-error refusal (a credential outage, a provider 5xx) as a backend error
|
||||
// rather than handing it back as a real answer. Compiled once at startup, keyed by profile
|
||||
// name, mirroring exhaustedPatternsByProfile above — a profile with no configured
|
||||
// errorPattern is simply absent here, so CompletionResolver falls back to its built-in
|
||||
// narrow {@code (?i)\bAPI Error\s*:} compatibility pattern for that profile's targets
|
||||
// (BackendErrorPatternLookup#legacy's contract — see backendErrorPatterns below).
|
||||
// rather than handing it back as a real answer. Still compiled once at startup, keyed by
|
||||
// profile name — unlike exhaustedPattern above (fleetd #446), errorPattern was left DEFERRED
|
||||
// on purpose: the ticket that made exhaustedPattern hot scoped errorPattern/cooling-off out
|
||||
// explicitly. A profile with no configured errorPattern is simply absent here, so
|
||||
// CompletionResolver falls back to its built-in narrow {@code (?i)\bAPI Error\s*:}
|
||||
// compatibility pattern for that profile's targets (BackendErrorPatternLookup#legacy's
|
||||
// contract — see backendErrorPatterns below).
|
||||
Map<String, Pattern> errorPatternsByProfile = new LinkedHashMap<>();
|
||||
cfg.profiles().forEach((name, profile) -> {
|
||||
if (profile.hasErrorPattern()) {
|
||||
@@ -400,52 +422,25 @@ public final class Fleetd {
|
||||
BackendErrorPatternLookup backendErrorPatterns = backendErrorPatternLookup(sessions::roster,
|
||||
errorPatternsByProfile);
|
||||
log.info("backend-error classification (fleetd #201 Unit 5): {}",
|
||||
CompletionResolver.coverage("errorPattern", cfg.profiles().keySet(),
|
||||
errorPatternsByProfile.keySet()));
|
||||
errorPatternCoverageLine(cfg.profiles().keySet(), errorPatternsByProfile.keySet()));
|
||||
// CB-578 stage B: on a classification that actually wins, quarantine the exhausted profile's
|
||||
// CREDENTIAL — not the profile name — so a profile sharing that credential (e.g. two models
|
||||
// on one OpenAI account) is refused too, not just the one that happened to report it. Reads
|
||||
// the profile config live off `config`, so a credentialId edit is hot: no restart needed.
|
||||
//
|
||||
// fleetd #175: this sink is now also the quarantine target for OpenCodeLauncher's
|
||||
// model-mismatch check — it needs nothing profile-specific from the caller beyond `target`
|
||||
// (a herdr terminal id) and `reason`, so reusing it here is exactly "the existing
|
||||
// ExhaustionSink path", not a new mechanism.
|
||||
//
|
||||
// fleetd #234: that check fires from SessionAwareHandle.agentSessionId(), which runs during
|
||||
// SessionManager.acquire() BEFORE this session is registered in sessions.roster() — so the
|
||||
// roster-only lookup below used to find nothing, .ifPresent silently no-op'd, and the
|
||||
// ERROR the check had just logged ("quarantining this profile's credential") was a lie:
|
||||
// nothing was quarantined, and nothing said so. Two changes: (1) OpenCodeLauncher now
|
||||
// passes its OWN profile name via ExhaustionSink's 3-arg overload — it already has the
|
||||
// FleetConfig.Profile in hand and does not need the roster at all — used here as a
|
||||
// fallback whenever the roster lookup misses; (2) if a profile still cannot be resolved
|
||||
// (neither the roster nor the hint names a configured one), this logs loudly at ERROR
|
||||
// instead of silently doing nothing — a control that cannot act must say so.
|
||||
// fleetd #234, round 4: the 3-arg overload is now ExhaustionSink's single abstract method,
|
||||
// so this is safely a lambda — there is no separate 2-arg overload left for it to bind to
|
||||
// instead and silently drop profileHint (that was rounds 1-3's whole hazard).
|
||||
ExhaustionSink exhaustionSink = (target, reason, profileHint) -> {
|
||||
String profileName = sessions.roster().stream()
|
||||
.filter(session -> target.equals(session.terminalId()))
|
||||
.findFirst()
|
||||
.map(MemberSession::profile)
|
||||
.orElse(profileHint);
|
||||
FleetConfig.Profile profile = profileName == null ? null : config.get().profiles().get(profileName);
|
||||
if (profile == null) {
|
||||
log.error("quarantine requested for target '{}' ({}) but no profile could be "
|
||||
+ "resolved — the target is not (yet) in the roster, and {} — "
|
||||
+ "credential NOT quarantined (fleetd #234)",
|
||||
target, reason,
|
||||
profileHint == null ? "no profile hint was given"
|
||||
: "the hinted profile '" + profileHint + "' is not configured");
|
||||
return;
|
||||
}
|
||||
String credentialId = profile.effectiveCredentialId();
|
||||
quarantine.quarantine(credentialId);
|
||||
log.warn("credential '{}' quarantined for {}s (profile '{}'): {}", credentialId,
|
||||
cfg.quarantineCooldownSeconds(), profile.profile(), reason);
|
||||
};
|
||||
// fleetd #446 criterion 3: the backend text that triggered the most recent quarantine of
|
||||
// each credential, so fleet_profiles can report WHY a limit was hit, not only that it was.
|
||||
// Keyed by credential id — the same key BackendQuarantine's own remainingSeconds uses —
|
||||
// and written at the one call site that actually quarantines (inside exhaustionSink below),
|
||||
// so a reason can never be reported for a quarantine that never happened. Bounded the same
|
||||
// way BackendQuarantine's own internal map is documented to be: by the number of distinct
|
||||
// credentials ever exhausted, not by how often the config is edited.
|
||||
Map<String, String> quarantineReasonByCredential = new ConcurrentHashMap<>();
|
||||
// fleetd #446 follow-up (round 3): extracted into exhaustionSink(...) below — see that
|
||||
// method's javadoc for the full fleetd #175/#234/#446 history this used to carry inline —
|
||||
// so a dedicated test can drive the exact ExhaustionSink main() builds, not a hand-rebuilt
|
||||
// copy of its shape.
|
||||
ExhaustionSink exhaustionSink = exhaustionSink(sessions, config, quarantine,
|
||||
quarantineReasonByCredential, cfg);
|
||||
// fleetd #175: point the forwarding sink handed to OpenCodeLauncher above at the real one,
|
||||
// now that `sessions` exists to resolve target -> session -> profile.
|
||||
exhaustionSinkRef.set(exhaustionSink);
|
||||
@@ -661,21 +656,16 @@ public final class Fleetd {
|
||||
// built a second time — two independently-constructed sources reading the SAME BackendQuarantine
|
||||
// / BackendOutagePolicy would still be able to drift (e.g. a future edit to the credentialIdFor
|
||||
// closure in only one of the two places), exactly the shape #284 was.
|
||||
FleetMcp.QuarantineSource quarantineSource = new FleetMcp.QuarantineSource(profile -> {
|
||||
var configured = config.get().profiles().get(profile);
|
||||
return configured == null ? null : configured.effectiveCredentialId();
|
||||
}, quarantine);
|
||||
FleetMcp.QuarantineSource quarantineSource = quarantineSource(config, quarantine,
|
||||
liveExhaustedPatterns, quarantineReasonByCredential);
|
||||
FleetMcp.OutageSource outageSource = new FleetMcp.OutageSource(profile -> {
|
||||
var configured = config.get().profiles().get(profile);
|
||||
return configured == null ? null : configured.effectiveCredentialId();
|
||||
}, outagePolicy);
|
||||
|
||||
FleetMcp mcp = new FleetMcp(messages, workers, sessions, identity, presence,
|
||||
primaryRegistry, callers, metrics, new FleetMcp.CapacitySource(profile -> liveCountRef.get().apply(profile),
|
||||
profile -> {
|
||||
var configured = config.get().profiles().get(profile);
|
||||
return configured == null ? null : configured.maxLoad();
|
||||
}, () -> config.get().profiles().keySet(), System::nanoTime),
|
||||
primaryRegistry, callers, metrics,
|
||||
capacitySource(config, cfg, profile -> liveCountRef.get().apply(profile)),
|
||||
new FleetMcp.HealthCoverageSource(() -> {
|
||||
var health = config.get().health();
|
||||
return FleetHealthMonitor.coverage(health != null && health.isEnabled(),
|
||||
@@ -798,6 +788,223 @@ public final class Fleetd {
|
||||
return target -> presence.isPresent(target) || leads.get().containsKey(target);
|
||||
}
|
||||
|
||||
/**
|
||||
* fleetd #446 follow-up (round 3): the quarantine {@link ExhaustionSink} main() actually wires
|
||||
* — extracted out of {@code main} for the same reason {@link #quarantineSource} below and
|
||||
* {@link #usageLimitFixWarning}/{@link #usageLimitFixWarningNoModel} were. Round 2 pinned the
|
||||
* WARNING text by having {@code FleetdUsageLimitFixWarningTest} call those two methods
|
||||
* directly, but a mutation battery proved that left two gaps: nothing proved this sink's
|
||||
* {@code log.warn} call actually uses either method (replacing the whole ternary with a
|
||||
* literal string), and nothing proved it picks the right one for a profile WITH a configured
|
||||
* {@code model:} versus one WITHOUT (swapping the ternary's two branches) — both mutations
|
||||
* left round 2's full 1592-test suite green. {@code FleetdExhaustionSinkWarningTest} drives
|
||||
* THIS factory's return value directly and asserts on the real text a {@code ListAppender}
|
||||
* attached to this class's own logger captures — the only way to prove the caller, not just
|
||||
* the two callees in isolation.
|
||||
*
|
||||
* <p>Behaviourally unchanged from the inline lambda this replaces:
|
||||
* <ul>
|
||||
* <li>fleetd #175: also the quarantine target for {@code OpenCodeLauncher}'s model-mismatch
|
||||
* check — it needs nothing profile-specific beyond {@code target}/{@code reason}, so
|
||||
* reusing this sink is the existing {@code ExhaustionSink} path, not a new mechanism;</li>
|
||||
* <li>fleetd #234 (round 4): resolves {@code target} to a profile via the live session
|
||||
* roster first, falling back to {@code profileHint} — {@code
|
||||
* SessionAwareHandle.agentSessionId()} fires this sink during {@code
|
||||
* SessionManager.acquire()}, <em>before</em> that session is registered, so the
|
||||
* roster-only lookup alone used to silently miss it. A profile resolved neither way
|
||||
* logs loudly at ERROR instead of doing nothing;</li>
|
||||
* <li>fleetd #446 criterion 3: writes {@code quarantineReasonByCredential} at the one call
|
||||
* site that actually quarantines, keyed by credential id — see the field's declaration
|
||||
* in {@code main} for the bound on its size.</li>
|
||||
* </ul>
|
||||
*
|
||||
* <p>{@code cfg} is the startup {@link FleetConfig} snapshot, read only for {@code
|
||||
* quarantineCooldownSeconds()} in the log text — a Cold key (see {@code ConfigRef}'s class
|
||||
* doc), so reading it off the snapshot rather than {@code config.get()} makes no observable
|
||||
* difference and matches what the inline version already did.
|
||||
*/
|
||||
static ExhaustionSink exhaustionSink(SessionManager sessions, ConfigRef config, BackendQuarantine quarantine,
|
||||
Map<String, String> quarantineReasonByCredential, FleetConfig cfg) {
|
||||
return (target, reason, profileHint) -> {
|
||||
String profileName = sessions.roster().stream()
|
||||
.filter(session -> target.equals(session.terminalId()))
|
||||
.findFirst()
|
||||
.map(MemberSession::profile)
|
||||
.orElse(profileHint);
|
||||
FleetConfig.Profile profile = profileName == null ? null : config.get().profiles().get(profileName);
|
||||
if (profile == null) {
|
||||
log.error("quarantine requested for target '{}' ({}) but no profile could be "
|
||||
+ "resolved — the target is not (yet) in the roster, and {} — "
|
||||
+ "credential NOT quarantined (fleetd #234)",
|
||||
target, reason,
|
||||
profileHint == null ? "no profile hint was given"
|
||||
: "the hinted profile '" + profileHint + "' is not configured");
|
||||
return;
|
||||
}
|
||||
String credentialId = profile.effectiveCredentialId();
|
||||
quarantine.quarantine(credentialId);
|
||||
quarantineReasonByCredential.put(credentialId, reason);
|
||||
log.warn("credential '{}' quarantined for {}s (profile '{}'): {}", credentialId,
|
||||
cfg.quarantineCooldownSeconds(), profile.profile(), reason);
|
||||
// fleetd #446 criterion 2: name the fix, not just the fact — an operator reading this
|
||||
// should not have to work out which of several configured models to touch.
|
||||
String model = profile.model();
|
||||
log.warn(model != null && !model.isBlank()
|
||||
? usageLimitFixWarning(profile.profile(), model)
|
||||
: usageLimitFixWarningNoModel(profile.profile(), cfg.quarantineCooldownSeconds()));
|
||||
};
|
||||
}
|
||||
|
||||
/**
|
||||
* fleetd #404, superseded by fleetd #446: production source for quarantine reporting.
|
||||
* Credential IDs were already hot; {@code exhaustedPattern} used to be compiled once at
|
||||
* startup for {@link CompletionResolver}, so the armed field had to read that same frozen map
|
||||
* until restart — the exact asymmetry #446 exists to close. Both {@code exhaustedPatternArmed}
|
||||
* here and {@link CompletionResolver}'s own classification now read the ONE live {@link
|
||||
* LiveExhaustedPatterns} instance ({@code exhaustedPatterns.armed}/{@code .patternFor}), so a
|
||||
* reload that arms or disarms a profile's detection is visible to both at once — never two
|
||||
* independently-updated copies that could disagree, the fleetd #404 rule this method used to
|
||||
* violate on purpose and now upholds.
|
||||
*
|
||||
* <p>{@code modelFor} and {@code reasonFor} back fleetd #446 criterion 3 — {@code
|
||||
* fleet_profiles}'s quarantined row naming which model a quarantined profile runs
|
||||
* ({@code modelFor}, read live off {@code config} the same way {@code credentialIdFor} already
|
||||
* is) and the backend text that triggered the most recent quarantine of that credential
|
||||
* ({@code reasonFor}, backed by {@code reasonByCredential} — see its call site in {@code main}
|
||||
* for where that map is written, at the one place a credential is actually quarantined).
|
||||
*/
|
||||
static FleetMcp.QuarantineSource quarantineSource(ConfigRef config, BackendQuarantine quarantine,
|
||||
LiveExhaustedPatterns exhaustedPatterns, Map<String, String> reasonByCredential) {
|
||||
return new FleetMcp.QuarantineSource(profile -> {
|
||||
var configured = config.get().profiles().get(profile);
|
||||
return configured == null ? null : configured.effectiveCredentialId();
|
||||
}, quarantine, exhaustedPatterns::armed, profile -> {
|
||||
var configured = config.get().profiles().get(profile);
|
||||
return configured == null ? null : configured.model();
|
||||
}, reasonByCredential::get);
|
||||
}
|
||||
|
||||
/**
|
||||
* fleetd #415 (review follow-up): package-private factory for the CB-578 stage A {@code
|
||||
* exhaustedPattern} startup coverage line, paired explicitly with {@link
|
||||
* CompletionResolver.UnsetMeaning#OFF} — {@code exhaustedPattern} has no fallback, so a
|
||||
* profile with none configured really does have the classification off.
|
||||
*
|
||||
* <p>Extracted out of {@code main} for the same reason {@link #capacitySource} and {@link
|
||||
* #worktreeBranchLookup} were: {@code coverage()}'s own tests ({@code CompletionResolverTest})
|
||||
* prove it words {@code OFF} and {@link CompletionResolver.UnsetMeaning#BUILT_IN_DEFAULT}
|
||||
* correctly when a test supplies the meaning itself — they cannot prove {@code main} pairs the
|
||||
* right meaning with the right key, which is the actual fleetd #415 defect. <b>Measured:</b>
|
||||
* swapping the {@code UnsetMeaning} arguments between this method and {@link
|
||||
* #errorPatternCoverageLine} — recreating #415's defect with the two keys exchanged — compiled
|
||||
* with 0 errors and left all 1506 existing tests green before {@code
|
||||
* FleetdPatternCoverageLineTest} was added to catch exactly that swap.
|
||||
*/
|
||||
/**
|
||||
* fleetd #446 follow-up: the criterion-2 WARNING text — "name the fix, not just the fact" —
|
||||
* for a profile whose {@code model:} is configured. Extracted out of the {@code
|
||||
* exhaustionSink} lambda the same way {@link #exhaustedPatternCoverageLine} was extracted out
|
||||
* of {@code main}: the inline version compiled, ran, and read correctly, but nothing pinned
|
||||
* its text, so a later edit could silently stop naming the fix and every test would stay
|
||||
* green (measured: renaming the leading {@code "usage-limit fix:"} tag left all 1586
|
||||
* pre-follow-up tests passing). {@link FleetdUsageLimitFixWarningTest} asserts on the return
|
||||
* value of this method directly, which is exactly what the caller logs — {@code log.warn} is
|
||||
* called with this method's result as a single, already-formatted argument, so the string a
|
||||
* test sees here is byte-identical to what {@code fleetd.out} receives.
|
||||
*
|
||||
* <p>Deliberately asserts on three substrings, not the whole sentence (the profile name, the
|
||||
* model name, and the literal {@code enabled: false} / {@code models.allow} action) — see that
|
||||
* test's class doc for why a whole-sentence pin is the wrong granularity here.
|
||||
*/
|
||||
static String usageLimitFixWarning(String profile, String model) {
|
||||
return "usage-limit fix: profile '" + profile + "' runs model '" + model + "' — set "
|
||||
+ "`enabled: false` on that model's entry under models.allow in fleetd.yaml to "
|
||||
+ "stop new spawns landing on it (models: is hot, no restart needed); remove the "
|
||||
+ "line again once the subscription window resets";
|
||||
}
|
||||
|
||||
/**
|
||||
* fleetd #446 follow-up: the {@link #usageLimitFixWarning} counterpart for a profile with no
|
||||
* {@code model:} configured — {@code models.allow} gates by model name, so there is nothing
|
||||
* for an operator to flip, and this says so instead of naming a fix that does not exist.
|
||||
*/
|
||||
static String usageLimitFixWarningNoModel(String profile, Integer quarantineCooldownSeconds) {
|
||||
return "usage-limit fix: profile '" + profile + "' has no model: configured, so "
|
||||
+ "models.allow cannot gate it by name — the " + quarantineCooldownSeconds
|
||||
+ "s quarantine above is the only thing keeping new spawns off it for now";
|
||||
}
|
||||
|
||||
static String exhaustedPatternCoverageLine(Set<String> allProfiles, Set<String> configuredProfiles) {
|
||||
return CompletionResolver.coverage("exhaustedPattern", CompletionResolver.UnsetMeaning.OFF,
|
||||
allProfiles, configuredProfiles);
|
||||
}
|
||||
|
||||
/**
|
||||
* fleetd #415 (review follow-up): the {@code errorPattern} counterpart of {@link
|
||||
* #exhaustedPatternCoverageLine}, paired explicitly with {@link
|
||||
* CompletionResolver.UnsetMeaning#BUILT_IN_DEFAULT} — an unset {@code errorPattern} still runs
|
||||
* backend-error classification against {@code CompletionResolver}'s built-in {@code
|
||||
* BACKEND_ERROR} pattern, so the empty case is not "off". See {@link
|
||||
* #exhaustedPatternCoverageLine}'s javadoc for the measured swap mutation this pairing guards
|
||||
* against.
|
||||
*/
|
||||
static String errorPatternCoverageLine(Set<String> allProfiles, Set<String> configuredProfiles) {
|
||||
return CompletionResolver.coverage("errorPattern", CompletionResolver.UnsetMeaning.BUILT_IN_DEFAULT,
|
||||
allProfiles, configuredProfiles);
|
||||
}
|
||||
|
||||
/**
|
||||
* fleetd #422 follow-up: package-private factory for the startup line reporting which of the
|
||||
* three central {@code models.allow:} gate states the daemon booted into. Extracted the same
|
||||
* way {@link #exhaustedPatternCoverageLine}/{@link #errorPatternCoverageLine} are, so a
|
||||
* dedicated test can call it directly rather than parsing log output, and so {@code main}'s
|
||||
* only source for this line is {@link PeerLauncher#modelGateState()} — never a second,
|
||||
* independently-derived read of {@code cfg.models()} that could disagree with what {@code
|
||||
* CompositePeerLauncher}'s spawn gate actually enforces (the fleetd #404 lesson).
|
||||
*
|
||||
* <p>Unlike the two pattern-key lines above, there is no {@code UnsetMeaning} choice to make
|
||||
* here: {@link PeerLauncher.ModelGateState#configured()} already states, unambiguously, whether
|
||||
* an empty {@link PeerLauncher.ModelGateState#off()} means "no {@code models:} block to gate
|
||||
* with" or "a block armed and currently reporting zero off" — the exact two states a bare
|
||||
* {@code disabledModels()} read could not tell apart before this ticket.
|
||||
*/
|
||||
static String modelGateCoverageLine(PeerLauncher.ModelGateState state) {
|
||||
if (!state.configured()) {
|
||||
return "not configured (no models: block — nothing is gated, and nothing can be)";
|
||||
}
|
||||
Set<String> off = state.off();
|
||||
return off.isEmpty()
|
||||
? "armed (models: block present; 0 models currently turned off)"
|
||||
: "armed (" + off.size() + " model(s) turned off: " + new TreeSet<>(off) + ")";
|
||||
}
|
||||
|
||||
/**
|
||||
* fleetd #416: production source for {@code fleet_list}'s per-profile capacity facts.
|
||||
*
|
||||
* <p>The profile <em>set</em> ({@code configuredProfiles}) must come from {@code cfg} — the
|
||||
* startup snapshot — not the live {@code config.get()}. {@code profiles} as a whole is a
|
||||
* {@code DEFERRED} key ({@link ConfigRef#DEFERRED_KEYS}): {@code HerdrPeerLauncher} takes
|
||||
* {@code Map.copyOf(profiles)} once at construction and a profile only added to the
|
||||
* hot-reloaded map can never actually be spawned, so enumerating it live made {@code fleet_list}
|
||||
* report a profile as available when {@code fleet_spawn} on that same profile fails with
|
||||
* {@code unknown worker profile}. {@code fleet_list}'s own contract for {@code free} is "the
|
||||
* same check the spawn gate itself runs" — the set the spawn gate can see is the startup one,
|
||||
* so this must enumerate that one too, the same shape as {@code coordinator.peers} above.
|
||||
*
|
||||
* <p>{@code maxLoad} stays live on purpose: it is read off {@code config.get()} exactly like
|
||||
* {@code credentialId} ({@link ConfigRef} documents both as hot), so an existing profile's
|
||||
* {@code maxLoad} edit must still change what {@code fleet_list} reports without a restart.
|
||||
*/
|
||||
static FleetMcp.CapacitySource capacitySource(ConfigRef config, FleetConfig cfg,
|
||||
Function<String, Integer> liveCount) {
|
||||
return new FleetMcp.CapacitySource(liveCount,
|
||||
profile -> {
|
||||
var configured = config.get().profiles().get(profile);
|
||||
return configured == null ? null : configured.maxLoad();
|
||||
},
|
||||
cfg.profiles()::keySet, System::nanoTime);
|
||||
}
|
||||
|
||||
/**
|
||||
* fleetd #248: package-private factory for the member worktree/branch lookup {@link
|
||||
* CompletionResolver} uses to name a fallback report's worktree and branch (fleetd#241).
|
||||
@@ -1323,6 +1530,52 @@ public final class Fleetd {
|
||||
+ "fleetd.yaml — see fleetd.example.yaml — and restart.");
|
||||
}
|
||||
|
||||
/**
|
||||
* fleetd #395: {@code exhaustedPattern} (see {@link FleetConfig.Profile#exhaustedPattern}) is
|
||||
* deliberately opt-in — {@code null}/blank means a backend refusal on that profile is never
|
||||
* classified as {@code BACKEND_EXHAUSTED}, so its credential is never quarantined. That is a
|
||||
* legitimate choice (guessing the vendor's wording would be worse), but an operator who never
|
||||
* opted a profile in should not discover the gap only when a usage limit silently goes
|
||||
* undetected. Warn once at startup, naming every unarmed profile, exactly like {@link
|
||||
* #reportMemberCredentialsGap} — never refuse to start over it.
|
||||
*
|
||||
* <p>A {@code subscription: true} profile that is unarmed gets a SECOND, louder WARN of its
|
||||
* own: it bills the operator's metered Claude plan, the case where an undetected usage limit
|
||||
* costs the most.
|
||||
*
|
||||
* <p>Package-private so a test can capture the real log via a {@link
|
||||
* ch.qos.logback.core.read.ListAppender}, the same pattern {@link
|
||||
* #reportMemberCredentialsGap}'s own test uses.
|
||||
*/
|
||||
static void reportExhaustedPatternGap(FleetConfig cfg) {
|
||||
List<String> unarmedSubscription = new ArrayList<>();
|
||||
List<String> unarmedOther = new ArrayList<>();
|
||||
cfg.profiles().forEach((name, profile) -> {
|
||||
if (!profile.hasExhaustedPattern()) {
|
||||
(profile.isSubscription() ? unarmedSubscription : unarmedOther).add(name);
|
||||
}
|
||||
});
|
||||
if (unarmedSubscription.isEmpty() && unarmedOther.isEmpty()) {
|
||||
log.info("exhaustedPattern: every configured profile has usage-limit detection armed");
|
||||
return;
|
||||
}
|
||||
List<String> allUnarmed = new ArrayList<>(unarmedSubscription);
|
||||
allUnarmed.addAll(unarmedOther);
|
||||
allUnarmed = allUnarmed.stream().sorted().toList();
|
||||
log.warn("exhaustedPattern: profile(s) {} have no exhaustedPattern configured — a "
|
||||
+ "usage-limit refusal on any of them is never detected and never "
|
||||
+ "quarantines its credential. Set exhaustedPattern (see "
|
||||
+ "fleetd.example.yaml) to arm detection for a profile.",
|
||||
allUnarmed);
|
||||
if (!unarmedSubscription.isEmpty()) {
|
||||
List<String> sortedSubscription = unarmedSubscription.stream().sorted().toList();
|
||||
log.warn("exhaustedPattern: subscription profile(s) {} run on the operator's metered "
|
||||
+ "Claude plan and have NO usage-limit detection armed — this is the "
|
||||
+ "case where a missed usage limit costs the most.",
|
||||
sortedSubscription);
|
||||
}
|
||||
}
|
||||
|
||||
/**
|
||||
* Poll herdr's {@code ping} until it answers or {@link #HERDR_WAIT_SECONDS} elapses (CB-504).
|
||||
*
|
||||
|
||||
@@ -30,6 +30,16 @@ public final class Authz {
|
||||
DRAIN,
|
||||
/** Read-only observation: status, roster, profiles, task polling. */
|
||||
READ,
|
||||
/**
|
||||
* Read (never ack) this daemon's own held lead-to-lead coordination mail (fleetd #421).
|
||||
*
|
||||
* <p>Deliberately <strong>not</strong> folded into {@link #READ}. {@code READ}'s grant
|
||||
* rests on "the roster carries no secrets" (see its case below) — a lead-to-lead body is
|
||||
* not the roster; it is where leads discuss host shapes, credentials and unmerged work.
|
||||
* Mapping this to {@code READ} would let any worker read every peer lead's mail in full
|
||||
* and would silently falsify that comment for every other {@code READ} caller.
|
||||
*/
|
||||
COORD_READ,
|
||||
/** Scrape the metrics endpoint. */
|
||||
METRICS
|
||||
}
|
||||
@@ -68,6 +78,11 @@ public final class Authz {
|
||||
// Observation is open to every authenticated role: a worker legitimately polls its own
|
||||
// status, and the roster carries no secrets.
|
||||
case READ, METRICS -> caller.isPrimary() || caller.isWorker() || caller.isArchitect();
|
||||
|
||||
// fleetd #421: reading held lead-to-lead mail is the primary's alone. An architect
|
||||
// holds READ today (CB-548), so "not primary" must mean not-architect here too — this
|
||||
// is coordination between leads, not observation of the roster.
|
||||
case COORD_READ -> caller.isPrimary();
|
||||
};
|
||||
}
|
||||
|
||||
|
||||
@@ -11,6 +11,7 @@ import java.util.LinkedHashMap;
|
||||
import java.util.List;
|
||||
import java.util.Map;
|
||||
import java.util.Objects;
|
||||
import java.util.function.Supplier;
|
||||
|
||||
/**
|
||||
* The architect-slot registry (CB-548): every gateway-local architect name and the strong-model
|
||||
@@ -19,16 +20,43 @@ import java.util.Objects;
|
||||
*
|
||||
* <p>Two halves, split by who owns each:
|
||||
* <ul>
|
||||
* <li><b>slots</b> — configured once, keyed by the gateway-local unique name; each carries the
|
||||
* {@code profile} reference the spawn lifecycle reads when it stands the slot up. A read-only
|
||||
* snapshot taken at construction.</li>
|
||||
* <li><b>terminal bindings</b> — owned by this registry and initially <em>empty</em>. Config
|
||||
* declares no architect terminal, so at startup every slot is idle and nothing resolves to an
|
||||
* architect; a session only becomes one when the spawn lifecycle {@linkplain #bind(String,
|
||||
* String) binds} its terminal to a slot. {@link CallerResolver} reads this through
|
||||
* {@link #snapshot()} to turn a pane into an {@link Role#ARCHITECT}.</li>
|
||||
* <li><b>slots</b> — read from {@code fleet.architects}/{@code developers}/{@code reviewers}
|
||||
* (see {@link #slots()}), each carrying the {@code profile} reference the spawn lifecycle
|
||||
* reads when it stands the slot up. <strong>Live, since fleetd #424</strong>: {@link #live}
|
||||
* re-reads {@code fleet:} on every call, through a supplier the same shape as
|
||||
* {@code CompositePeerLauncher}'s (see {@code ConfigRef}'s class doc) — so a config reload
|
||||
* that removes or adds an architect slot governs the <em>next</em> spawn with no restart.
|
||||
* Only {@link #MemberRegistry(FleetConfig.Fleet)} freezes the pool at construction, and that
|
||||
* constructor exists for tests and for the (rare) case of wiring a fixed, code-built config.</li>
|
||||
* <li><b>terminal bindings</b> — owned by this registry, initially <em>empty</em>, and
|
||||
* <strong>never</strong> touched by a reload. Config declares no architect terminal, so at
|
||||
* startup every slot is idle and nothing resolves to an architect; a session only becomes one
|
||||
* when the spawn lifecycle {@linkplain #bind(String, String) binds} its terminal to a slot.
|
||||
* {@link CallerResolver} reads this through {@link #snapshot()} to turn a pane into an
|
||||
* {@link Role#ARCHITECT}.</li>
|
||||
* </ul>
|
||||
*
|
||||
* <p><strong>The binding rule (fleetd #424): config governs what a bound slot still grants, as
|
||||
* well as what may be bound next.</strong> Removing a slot from config revokes it — that is the
|
||||
* ticket's entire point ("Revoking an architect slot does not revoke it"). Revoking it means an
|
||||
* architect already bound to that slot loses the ARCHITECT privilege on its very next request:
|
||||
* {@link #roleForSlot} and {@link #nameForSlot} read {@link #slots()} directly, with no cache, so
|
||||
* the moment a slot drops out of config, {@link CallerResolver#resolve} (which calls both on every
|
||||
* request from a bound pane, {@code CallerResolver.java:220}) can no longer confirm the pane's slot
|
||||
* is an architect slot, and the pane falls through to {@code Principal.worker(...)}. What does
|
||||
* <em>not</em> change is the {@code terminalToSlot} <em>occupancy</em> — the binding created by
|
||||
* {@link #bind} is untouched by a reload, on purpose: unbinding it here would double-book the slot
|
||||
* key (a second terminal could then bind to the "freed" key while the first is still the terminal
|
||||
* the operator actually meant to demote) and would silently break {@link #unbind}'s compare-safe
|
||||
* contract, which needs the original {@code terminal → slot} pair intact to remove it cleanly. So
|
||||
* the demoted session keeps occupying its slot — {@link #slotForTerminal} and {@link #snapshot()}
|
||||
* still name it — it just no longer resolves as an architect through that occupancy, and a fresh
|
||||
* spawn still cannot bind to the same key while it is occupied ({@link #reserve}/
|
||||
* {@link #requireSlotFor} refuse it anyway, since it is gone from {@link #slots()}). The demoted
|
||||
* session's own turn is unaffected: {@code fleet_reply}'s authorization
|
||||
* ({@code Authz.Action.REPLY}) is {@code caller.ownsSession(targetSession)} — identity by terminal,
|
||||
* not by role — so a demoted architect can still end its own turn normally.
|
||||
*
|
||||
* <p>Spawning/lifecycle is deliberately a separate unit: this class only owns the bindings and
|
||||
* exposes the map the resolver resolves against plus the profile lookup lifecycle will call.
|
||||
* Nothing here creates or manages an architect session.
|
||||
@@ -55,14 +83,42 @@ public final class MemberRegistry implements MemberLifecycle {
|
||||
}
|
||||
}
|
||||
|
||||
private final Map<String, Entry> slots;
|
||||
private final Supplier<FleetConfig.Fleet> fleet;
|
||||
/** Live {@code terminal_id → qualified slot key}; guarded by {@code terminalToSlot}. */
|
||||
private final Map<String, String> terminalToSlot = new HashMap<>();
|
||||
/** Slot keys held between reservation and the terminal binding. Guarded by terminalToSlot. */
|
||||
private final java.util.Set<String> reservedSlots = new java.util.HashSet<>();
|
||||
|
||||
/** Flatten every role pool in {@code fleet} into one registry. Leaders are not members. */
|
||||
/**
|
||||
* Freeze the pool at construction — for tests, and for the rare case of wiring a fixed,
|
||||
* code-built config. Production wiring should prefer {@link #live}, which re-reads
|
||||
* {@code fleet:} on every call.
|
||||
*/
|
||||
public MemberRegistry(FleetConfig.Fleet fleet) {
|
||||
this(() -> fleet);
|
||||
}
|
||||
|
||||
private MemberRegistry(Supplier<FleetConfig.Fleet> fleet) {
|
||||
this.fleet = fleet;
|
||||
}
|
||||
|
||||
/**
|
||||
* Live variant (fleetd #424): {@code fleet} is read fresh on every {@link #slots()} call — pass
|
||||
* {@code () -> config.get().fleet()}, the same supplier shape {@code CompositePeerLauncher}
|
||||
* already uses for placement — so a reload that adds or removes an architect slot governs the
|
||||
* next spawn's {@link #reserve}/{@link #requireSlotFor} check with no restart. A separate,
|
||||
* private constructor rather than a same-arity public overload of
|
||||
* {@link #MemberRegistry(FleetConfig.Fleet)}: a {@code FleetConfig.Fleet} and a
|
||||
* {@code Supplier<FleetConfig.Fleet>} overload are ambiguous for a literal {@code null} — the
|
||||
* same reason {@code CallerResolver.withLeads} is a static factory rather than a fourth
|
||||
* constructor overload.
|
||||
*/
|
||||
public static MemberRegistry live(Supplier<FleetConfig.Fleet> fleet) {
|
||||
return new MemberRegistry(Objects.requireNonNull(fleet, "fleet"));
|
||||
}
|
||||
|
||||
/** Flatten every role pool in {@code fleet} into one map. Leaders are not members. */
|
||||
private static Map<String, Entry> flatten(FleetConfig.Fleet fleet) {
|
||||
Map<String, Entry> flat = new LinkedHashMap<>();
|
||||
if (fleet != null) {
|
||||
for (MemberRole role : MemberRole.values()) {
|
||||
@@ -74,18 +130,22 @@ public final class MemberRegistry implements MemberLifecycle {
|
||||
});
|
||||
}
|
||||
}
|
||||
this.slots = Collections.unmodifiableMap(flat);
|
||||
return Collections.unmodifiableMap(flat);
|
||||
}
|
||||
|
||||
/** The configured slots, keyed by qualified {@link Entry#key()}. Unmodifiable snapshot. */
|
||||
/**
|
||||
* The configured slots, keyed by qualified {@link Entry#key()}. Unmodifiable snapshot of
|
||||
* {@code fleet:} <em>as of this call</em> — see the class doc for which constructor makes that
|
||||
* live versus frozen.
|
||||
*/
|
||||
public Map<String, Entry> slots() {
|
||||
return slots;
|
||||
return flatten(fleet.get());
|
||||
}
|
||||
|
||||
/** The slots belonging to {@code role}, in definition order. */
|
||||
/** The slots belonging to {@code role}, in definition order, as of this call. */
|
||||
public Map<String, Entry> slotsFor(MemberRole role) {
|
||||
Map<String, Entry> out = new LinkedHashMap<>();
|
||||
slots.forEach((key, e) -> {
|
||||
slots().forEach((key, e) -> {
|
||||
if (e.role() == role) {
|
||||
out.put(key, e);
|
||||
}
|
||||
@@ -117,31 +177,49 @@ public final class MemberRegistry implements MemberLifecycle {
|
||||
}
|
||||
|
||||
/**
|
||||
* The strong-model profile a slot runs under — what the spawn lifecycle reads.
|
||||
* The strong-model profile a slot runs under, as of this call.
|
||||
*
|
||||
* <p>Nothing in {@code src/main} calls this (fleetd #431 — grepped both the {@code
|
||||
* .profileForSlot(} and the {@code ::profileForSlot} form). This javadoc used to say "what the
|
||||
* spawn lifecycle reads", and that seam does not exist: the spawn lifecycle takes its profile
|
||||
* from the {@link MemberLifecycle.SlotReservation} that {@code reserve} returns, never from
|
||||
* here. Kept and pinned rather than deleted because it is the natural accessor for that seam
|
||||
* if one is added; live for the same reason as {@link #roleForSlot}, so a reload cannot leave
|
||||
* it answering for the old config.
|
||||
*
|
||||
* @return the slot's configured {@code profile}, or {@code null} if the slot is unknown or
|
||||
* declares none
|
||||
*/
|
||||
public String profileForSlot(String slotName) {
|
||||
Entry e = slots.get(slotName);
|
||||
Entry e = slots().get(slotName);
|
||||
return (e == null || e.profile() == null) ? null : e.profile();
|
||||
}
|
||||
|
||||
/** The role a qualified slot key belongs to, or {@code null} when the key is unknown. */
|
||||
/**
|
||||
* The role a qualified slot key belongs to, or {@code null} when the key is not currently
|
||||
* configured. Deliberately live, with no cache (fleetd #424, see the class doc's binding rule):
|
||||
* removing a slot from config must make {@link CallerResolver#resolve} stop granting the
|
||||
* ARCHITECT role for it on the very next request from a terminal that was bound to it, which is
|
||||
* the ticket's whole point — revoking a slot must actually revoke it, not just refuse the next
|
||||
* spawn.
|
||||
*/
|
||||
public MemberRole roleForSlot(String slotName) {
|
||||
Entry e = slots.get(slotName);
|
||||
Entry e = slots().get(slotName);
|
||||
return e == null ? null : e.role();
|
||||
}
|
||||
|
||||
/** The unqualified configured name for a slot, or {@code null} if it is unknown. */
|
||||
/**
|
||||
* The unqualified configured name for a slot, or {@code null} if it is not currently configured.
|
||||
* Live for the same reason as {@link #roleForSlot} — see the class doc's binding rule.
|
||||
*/
|
||||
public String nameForSlot(String slotName) {
|
||||
Entry e = slots.get(slotName);
|
||||
Entry e = slots().get(slotName);
|
||||
return e == null ? null : e.name();
|
||||
}
|
||||
|
||||
/** True when {@code slotName} is a configured architect slot. */
|
||||
public boolean isSlot(String slotName) {
|
||||
return slots.containsKey(slotName);
|
||||
return slots().containsKey(slotName);
|
||||
}
|
||||
|
||||
/**
|
||||
|
||||
@@ -32,9 +32,36 @@ import java.util.function.Supplier;
|
||||
* not the fact that they are config. Most of {@code fleet:} — every role pool
|
||||
* ({@code architects}/{@code developers}/{@code reviewers}), {@code charters}, and
|
||||
* {@code tabLabel} — is read the same live way, through the same supplier
|
||||
* ({@code () -> config.get().fleet()}). <strong>But {@code fleet:} as a whole is NOT in this
|
||||
* class</strong>: {@code fleet.leaders} inside the same key is frozen, which is exactly what
|
||||
* makes {@code fleet:} split rather than hot — see below.</li>
|
||||
* ({@code () -> config.get().fleet()}). {@code architects} in particular is hot for
|
||||
* <strong>two independent consumers</strong> (fleetd #424): {@code CompositePeerLauncher}
|
||||
* reads it live for placement (which profile an unqualified architect spawn may land on), and
|
||||
* {@code MemberRegistry} separately reads it live, through its own instance of the same
|
||||
* supplier shape, for identity — both which slot a spawn may bind to <em>and</em> what a slot
|
||||
* already bound still grants. Removing an architect slot from config therefore revokes the
|
||||
* {@link dev.ltms.fleet.auth.Role#ARCHITECT} role on the bound pane's very next request; only
|
||||
* the slot <em>occupancy</em> survives, so the demoted session still holds its slot key until
|
||||
* it unbinds. See {@code MemberRegistry}'s class doc for that binding rule.
|
||||
* <strong>But {@code fleet:} as a whole is NOT in this class</strong>: {@code fleet.leaders}
|
||||
* inside the same key is frozen, which is exactly what makes {@code fleet:} split rather than
|
||||
* hot — see below. {@code models:} (fleetd #422) joined this class whole: {@link
|
||||
* FleetConfig#validateModels()} re-runs fully against the fresh config on every {@link
|
||||
* #reload()} (via {@link FleetConfig#validateAll()}), refusing a bad edit outright rather than
|
||||
* caching a stale copy anywhere, and the on/off half added by fleetd #422 is read live both by
|
||||
* {@code CompositePeerLauncher}'s spawn gate ({@code enforceModelEnabled} and its candidate
|
||||
* filter) and by {@code fleet_profiles}/{@code GET /profiles} (via
|
||||
* {@code PeerLauncher.disabledModels()}). Nothing about {@code models:} is baked into an
|
||||
* object built at startup, so — unlike the deferred keys below — there is no frozen half left
|
||||
* to report; it moved here from deferred rather than joining split. An existing profile's
|
||||
* {@code exhaustedPattern} (fleetd #446) joined this class the same way: it used to be
|
||||
* compiled once into {@code Fleetd.main}'s startup pattern map (see the Deferred bullet's old
|
||||
* wording, and {@code LiveExhaustedPatterns}'s class doc for the history), and is now read
|
||||
* live, cached by profile name, by both {@code CompletionResolver}'s classification (via
|
||||
* {@code LiveExhaustedPatterns.patternFor}) and {@code fleet_profiles}'s {@code
|
||||
* exhaustionDetectionArmed} (via {@code LiveExhaustedPatterns.armed}) — the one live object
|
||||
* both read, so a reload that arms or disarms a profile's usage-limit detection takes effect
|
||||
* on the next check with no restart. {@code errorPattern}, {@code exhaustedPattern}'s sibling
|
||||
* key for backend-error (not usage-limit) classification, was deliberately left OUT of this
|
||||
* fleetd #446 change and stays deferred below — the ticket scoped it out explicitly.</li>
|
||||
* <li><strong>Deferred</strong> — accepted into the new snapshot, but the wiring built at startup
|
||||
* keeps the old value until a restart: {@code lifecycle:}, {@code leadHeartbeat:},
|
||||
* {@code idleSleepGuard:} ({@code Fleetd.java} reads it once, at startup, to decide whether
|
||||
@@ -64,15 +91,17 @@ import java.util.function.Supplier;
|
||||
* (if any) simply keeps its old settings), adding or removing a profile (a new backend needs its own launcher,
|
||||
* which is constructed once), <em>and an existing profile's launch settings</em> —
|
||||
* {@code model}, {@code baseUrl}, {@code argv}, {@code env}, {@code mcpUrl},
|
||||
* {@code exhaustedPattern} (CB-578 stage A — compiled once into {@code Fleetd.main}'s
|
||||
* pattern map at startup), {@code errorPattern} (fleetd #201 Unit 5 — compiled once into
|
||||
* {@code Fleetd.main}'s backend-error pattern map at startup, the same way),
|
||||
* {@code errorPattern} (fleetd #201 Unit 5 — compiled once into
|
||||
* {@code Fleetd.main}'s backend-error pattern map at startup; deliberately NOT made hot
|
||||
* alongside {@code exhaustedPattern} by fleetd #446 — that ticket scoped {@code errorPattern}
|
||||
* and cooling-off out explicitly),
|
||||
* {@code ideProjectDir} / {@code ideOpenCommand} / {@code autoCompactWindow} (fleetd #323
|
||||
* instance 1 — all three are read at spawn off the same frozen profile map and were missing
|
||||
* from {@link #sameLaunchSettings}), and the rest of {@link #sameLaunchSettings}.
|
||||
* {@code credentialId} (CB-578 stage B) is NOT on
|
||||
* this list — it is read live off the config supplier at every quarantine check and
|
||||
* exhaustion event, exactly like {@code weight} / {@code maxLoad}, so it is hot instead.
|
||||
* {@code credentialId} (CB-578 stage B) and {@code exhaustedPattern} (fleetd #446) are NOT on
|
||||
* this list — both are read live off the config supplier at every quarantine check and
|
||||
* exhaustion event, exactly like {@code weight} / {@code maxLoad}, so both are hot instead
|
||||
* (see the Hot bullet above for {@code exhaustedPattern}'s history).
|
||||
* {@code HerdrPeerLauncher} takes {@code Map.copyOf(profiles)} at construction and resolves
|
||||
* each spawn out of that copy, so those never reach a launch until the daemon restarts. A
|
||||
* reload logs these rather than pretending they applied.</li>
|
||||
@@ -135,15 +164,18 @@ import java.util.function.Supplier;
|
||||
* </ul>
|
||||
*
|
||||
* <p><strong>The denominator, measured on 2026-09-04 (fleetd #330; recounted for fleetd #333);
|
||||
* recounted again for fleetd #362, and again after {@code idleSleepGuard:} was added.</strong>
|
||||
* {@code FleetConfig} has 24 top-level record components: 5 cold, 13 deferred, 3 split, 3
|
||||
* hot-excluded. Three of them are named nowhere in this file, and the reason is the same for all
|
||||
* three: {@code placement}, {@code memberCredentials} and {@code memberLoginShell} are
|
||||
* <strong>hot</strong> and correctly absent — all three are read live off {@code config.get()}
|
||||
* recounted again for fleetd #362, again after {@code idleSleepGuard:} was added, again after
|
||||
* {@code models:} was added as deferred, and again for fleetd #422, which moved {@code models:}
|
||||
* from deferred to hot-excluded once its on/off half was read live everywhere.</strong>
|
||||
* {@code FleetConfig} has 25 top-level record components: 5 cold, 13 deferred, 3 split, 4
|
||||
* hot-excluded. Four of them are named nowhere in this file, and the reason is the same for all
|
||||
* four: {@code placement}, {@code memberCredentials}, {@code memberLoginShell} and {@code models}
|
||||
* are <strong>hot</strong> and correctly absent — all four are read live off {@code config.get()}
|
||||
* (placement through the {@code CompositePeerLauncher} supplier the Hot bullet names;
|
||||
* {@code memberCredentials}/{@code memberLoginShell} at spawn time, {@code Fleetd.java:198, 205, 729}
|
||||
* and {@code HerdrPeerLauncher#configuredMemberLoginShell}), so a reload takes effect on the next
|
||||
* spawn with no entry needed here.
|
||||
* and {@code HerdrPeerLauncher#configuredMemberLoginShell}; {@code models} the same way, through the
|
||||
* Hot bullet's {@code models:} paragraph), so a reload takes effect on the next spawn (or, for
|
||||
* {@code models}, the next reported status) with no entry needed here.
|
||||
* {@code health} and {@code coordinator} used to be a third kind — <strong>undecided</strong>, not
|
||||
* hot — until fleetd #330 added the <strong>split</strong> class above and gave them a home. A
|
||||
* reload touching either used to report a bare "config reloaded", which under-claimed; now it names
|
||||
@@ -322,12 +354,12 @@ public final class ConfigRef implements Supplier<FleetConfig> {
|
||||
fresh = FleetConfig.load(path);
|
||||
// The same gate startup runs. A config that would have refused to boot must not be able
|
||||
// to slip in through a reload — that is how a daemon ends up in a state it could never
|
||||
// have started in, which is the hardest kind to debug.
|
||||
fresh.validateAuthExposure();
|
||||
fresh.validateLeadTabPrefixes();
|
||||
fresh.validateSubscriptionProfiles();
|
||||
fresh.validateCharters();
|
||||
fresh.validateMembers();
|
||||
// have started in, which is the hardest kind to debug. fleetd ticket "central allow-list
|
||||
// of usable models" follow-up: this used to be six individual validateXxx() calls, and
|
||||
// mutation testing found two of the six unpinned here even though startup pinned nothing
|
||||
// at all — see FleetConfig#validateAll's javadoc for why the fix is one reflective call,
|
||||
// not a longer hand-maintained list.
|
||||
fresh.validateAll();
|
||||
} catch (RuntimeException e) {
|
||||
String msg = e.getMessage() == null ? e.toString() : e.getMessage();
|
||||
log.warn("config reload from {} refused, keeping the running config: {}", path, msg);
|
||||
@@ -513,26 +545,35 @@ public final class ConfigRef implements Supplier<FleetConfig> {
|
||||
+ "opened once and needs a restart; the broker URI env-var name kept out of a "
|
||||
+ "member's environment is read live on every spawn and already applied");
|
||||
}
|
||||
// fleetd #333: unlike health/coordinator above, most of `fleet:` (architects, developers,
|
||||
// reviewers, charters, tabLabel) is genuinely hot — ConfigRefTest.aHotChangeIsAppliedAndRead-
|
||||
// fleetd #333: unlike health/coordinator above, most of `fleet:` (developers, reviewers,
|
||||
// charters, tabLabel) is genuinely hot — ConfigRefTest.aHotChangeIsAppliedAndRead-
|
||||
// ThroughGet and aCharterChangeIsHotAndReachesTheLiveConfig prove it reaches the live config
|
||||
// with no restart note. Only fleet.leaders is frozen (Fleetd.java:281 reads
|
||||
// cfg.fleet().leaders() off the startup snapshot to build both the LeadTabScanner's
|
||||
// tab-label-to-name map, wired into CallerResolver.withLeadsAndMembers at Fleetd.java:620/624,
|
||||
// and — when herdr answered — LeadLauncher(...).ensureLeads() at Fleetd.java:315, which
|
||||
// auto-launches each lead up to its `instances` count; neither is rebuilt on reload). So this
|
||||
// compares fleet.leaders alone, not the whole Fleet record: comparing the whole record would
|
||||
// report "split" for a tabLabel-only or charters-only change that is actually fully hot,
|
||||
// which is the over-claim mirror of the under-claim bug this class exists to prevent.
|
||||
// with no restart note. `architects` is hot too, and — since fleetd #424 — hot for BOTH of
|
||||
// its consumers, not just the one this comment used to name: CompositePeerLauncher reads it
|
||||
// live for PLACEMENT through the () -> config.get().fleet() supplier named in the class doc's
|
||||
// Hot bullet, and MemberRegistry separately reads it live for IDENTITY (which slot a spawn
|
||||
// may bind to, AND what a slot already bound still grants) through its own instance of that
|
||||
// same supplier shape — see MemberRegistry.live and its class doc for the binding rule:
|
||||
// removing a slot revokes ARCHITECT on the bound pane's very next request, and only the slot
|
||||
// OCCUPANCY survives, so the demoted session keeps its slot key until it unbinds. Only
|
||||
// fleet.leaders is frozen (Fleetd.java:281 reads cfg.fleet().leaders() off the startup
|
||||
// snapshot to build both the LeadTabScanner's tab-label-to-name map, wired into
|
||||
// CallerResolver.withLeadsAndMembers at Fleetd.java:620/624, and — when herdr answered —
|
||||
// LeadLauncher(...).ensureLeads() at Fleetd.java:315, which auto-launches each lead up to its
|
||||
// `instances` count; neither is rebuilt on reload). So this compares fleet.leaders alone, not
|
||||
// the whole Fleet record: comparing the whole record would report "split" for a tabLabel-only
|
||||
// or architects-only change that is actually fully hot, which is the over-claim mirror of the
|
||||
// under-claim bug this class exists to prevent.
|
||||
if (!Objects.equals(leadersOf(old), leadersOf(fresh))) {
|
||||
changed.add("fleet: fleet.leaders (each lead's tab, workspace, cwd, profile and "
|
||||
+ "instances count) is read once at startup to build the LeadTabScanner's "
|
||||
+ "identity map and to auto-launch leads, and neither is rebuilt on reload, so a "
|
||||
+ "lead added, removed, or given a new tab: label needs a restart — until then it "
|
||||
+ "stays unrecognised, and a caller from its new tab resolves as a worker, not a "
|
||||
+ "lead; the rest of fleet: (architects, developers, reviewers, charters, "
|
||||
+ "tabLabel) is read live through the supplier on CompositePeerLauncher and "
|
||||
+ "already applied");
|
||||
+ "lead; the rest of fleet: (developers, reviewers, charters, tabLabel) is read "
|
||||
+ "live through the supplier on CompositePeerLauncher, and architects is read "
|
||||
+ "live through that same supplier for placement AND through a separate supplier "
|
||||
+ "on MemberRegistry for spawn-time identity — both already applied");
|
||||
}
|
||||
// Kept in step with SPLIT_KEYS the same way changedColdKeys is kept in step with COLD_KEYS —
|
||||
// every message here must be traceable to one of the split keys the class doc documents.
|
||||
@@ -554,12 +595,16 @@ public final class ConfigRef implements Supplier<FleetConfig> {
|
||||
* {@link #sameLaunchSettings} because they are read <em>live</em>, not baked in at spawn — see
|
||||
* the class doc's <em>Hot</em> bullet. {@code weight} and {@code maxLoad} are read live by the
|
||||
* placement policy on every spawn; {@code credentialId} is read live by
|
||||
* {@code CompositePeerLauncher} and the CB-578 stage B exhaustion sink. Nothing else is
|
||||
* {@code CompositePeerLauncher} and the CB-578 stage B exhaustion sink; {@code exhaustedPattern}
|
||||
* (fleetd #446) is read live, cached by profile name, by {@code LiveExhaustedPatterns} — see
|
||||
* that class's doc and the class doc's <em>Hot</em> bullet for the history (it used to be
|
||||
* compared here, deferred, like its sibling {@code errorPattern} still is). Nothing else is
|
||||
* excluded — see {@code sameLaunchSettingsComparesEveryProfileComponentOrExcludesIt} in
|
||||
* {@code ConfigRefProfileCoverageTest}, which enumerates every {@code Profile} record component
|
||||
* by reflection and fails the build if one is neither compared below nor named here.
|
||||
*/
|
||||
static final Set<String> LAUNCH_SETTINGS_EXCLUDED = Set.of("weight", "maxLoad", "credentialId");
|
||||
static final Set<String> LAUNCH_SETTINGS_EXCLUDED =
|
||||
Set.of("weight", "maxLoad", "credentialId", "exhaustedPattern");
|
||||
|
||||
/**
|
||||
* Whether two versions of a profile would launch a peer identically.
|
||||
@@ -596,13 +641,15 @@ public final class ConfigRef implements Supplier<FleetConfig> {
|
||||
&& Objects.equals(a.kind(), b.kind())
|
||||
&& Objects.equals(a.env(), b.env())
|
||||
&& Objects.equals(a.subscription(), b.subscription())
|
||||
// CB-578 stage B: exhaustedPattern is compiled once into Fleetd.main's pattern map
|
||||
// at startup (see ExhaustedPatternLookup wiring) — a reload never re-reads it, so a
|
||||
// changed pattern must be reported as deferred, exactly like model/baseUrl/argv.
|
||||
&& Objects.equals(a.exhaustedPattern(), b.exhaustedPattern())
|
||||
// fleetd #446: exhaustedPattern moved to LAUNCH_SETTINGS_EXCLUDED — it is now read
|
||||
// live, cached by profile name, through LiveExhaustedPatterns (see that class's doc
|
||||
// and ConfigRef's class doc Hot bullet), so it must NOT be compared here any more: a
|
||||
// reload that changes only exhaustedPattern must report "config reloaded", not
|
||||
// "these changes need a restart".
|
||||
// fleetd #201 Unit 5: errorPattern is compiled once into Fleetd.main's backend-error
|
||||
// pattern map at startup (see BackendErrorPatternLookup wiring), the same way
|
||||
// exhaustedPattern is — a reload never re-reads it either.
|
||||
// pattern map at startup (see BackendErrorPatternLookup wiring) — unlike its sibling
|
||||
// exhaustedPattern above (fleetd #446), a reload still never re-reads it; fleetd #446
|
||||
// scoped errorPattern out on purpose (see the class doc's Hot bullet).
|
||||
&& Objects.equals(a.errorPattern(), b.errorPattern())
|
||||
// fleetd #323 instance 1: ideProjectDir and ideOpenCommand are read at spawn off the
|
||||
// same frozen profile map as ideMcpUrl above (ClaudeCodeLauncher.java:267/269,
|
||||
|
||||
@@ -16,11 +16,15 @@ import org.slf4j.LoggerFactory;
|
||||
|
||||
import java.io.IOException;
|
||||
import java.io.UncheckedIOException;
|
||||
import java.lang.reflect.InvocationTargetException;
|
||||
import java.lang.reflect.Method;
|
||||
import java.lang.reflect.Modifier;
|
||||
import java.nio.file.Files;
|
||||
import java.nio.file.Path;
|
||||
import java.util.ArrayList;
|
||||
import java.util.Arrays;
|
||||
import java.util.Collections;
|
||||
import java.util.Comparator;
|
||||
import java.util.HashSet;
|
||||
import java.util.LinkedHashMap;
|
||||
import java.util.List;
|
||||
@@ -124,6 +128,15 @@ import java.util.regex.PatternSyntaxException;
|
||||
* under a member's long turn. {@code null} (the block omitted) behaves the
|
||||
* same as an explicit {@code enabled: true}; set {@code enabled: false} to
|
||||
* turn it off. See {@link dev.ltms.fleet.power.IdleSleepGuard}.
|
||||
* @param models central allow-list of models any {@code profiles:} entry may name. {@code
|
||||
* null} or an empty {@code allow:} ⇒ off: {@link #validateModels()} checks
|
||||
* nothing and every existing config keeps working exactly as it does today.
|
||||
* When non-empty, a profile whose {@code model:} is not one of {@link
|
||||
* Models#ids()} fails config load, naming both the model and the profile —
|
||||
* see {@link #validateModels()}. This block decides what may be CONFIGURED;
|
||||
* fleetd #422 added the separate on/off question — whether a configured model
|
||||
* may be spawned onto RIGHT NOW ({@link Models.ModelEntry#enabled}) — enforced
|
||||
* live at spawn by {@code CompositePeerLauncher}, not here. See {@link Models}.
|
||||
*/
|
||||
@JsonIgnoreProperties(ignoreUnknown = true)
|
||||
public record FleetConfig(
|
||||
@@ -150,7 +163,22 @@ public record FleetConfig(
|
||||
String worktreeGroup,
|
||||
String memberLoginShell,
|
||||
String memberSkills,
|
||||
IdleSleepGuard idleSleepGuard) {
|
||||
IdleSleepGuard idleSleepGuard,
|
||||
Models models) {
|
||||
|
||||
/** Back-compat form before the {@code models:} block was added. */
|
||||
public FleetConfig(Bind bind, String herdrSocket, String memberHerdrSocket, Map<String, Profile> profiles,
|
||||
Guard guard, String worktreeRoot, Lifecycle lifecycle, Integer spawnReadyTimeoutMs,
|
||||
Integer spawnReadyPollMs, Broker broker, Primary primary, Fleet fleet,
|
||||
LeadHeartbeat leadHeartbeat, Health health, String placement, Auth auth,
|
||||
ConfigReload configReload, Integer quarantineCooldownSeconds,
|
||||
MemberCredentials memberCredentials, Coordinator coordinator, String worktreeGroup,
|
||||
String memberLoginShell, String memberSkills, IdleSleepGuard idleSleepGuard) {
|
||||
this(bind, herdrSocket, memberHerdrSocket, profiles, guard, worktreeRoot, lifecycle, spawnReadyTimeoutMs,
|
||||
spawnReadyPollMs, broker, primary, fleet, leadHeartbeat, health, placement, auth,
|
||||
configReload, quarantineCooldownSeconds, memberCredentials, coordinator, worktreeGroup,
|
||||
memberLoginShell, memberSkills, idleSleepGuard, null);
|
||||
}
|
||||
|
||||
/** Back-compat form before the {@code idleSleepGuard:} block was added. */
|
||||
public FleetConfig(Bind bind, String herdrSocket, String memberHerdrSocket, Map<String, Profile> profiles,
|
||||
@@ -396,6 +424,12 @@ public record FleetConfig(
|
||||
* {@code null}/blank ⇒ the classification never fires for this profile and
|
||||
* today's completion-fallback behaviour is unchanged. Every backend words
|
||||
* its refusal differently, so this is config, never a vendor string in code.
|
||||
* Read live off the current config through {@code LiveExhaustedPatterns},
|
||||
* cached by profile name (fleetd #446), so it is HOT: an operator can arm
|
||||
* or disarm this profile's usage-limit detection by editing this key and
|
||||
* reloading, with no restart. Before fleetd #446 this was compiled once at
|
||||
* daemon startup, the same deferred shape its sibling {@code errorPattern}
|
||||
* (below) still has.
|
||||
* @param credentialId the shared account this profile authenticates as (CB-578 stage B). Two
|
||||
* or more profiles setting the <em>same</em> non-blank value are quarantined
|
||||
* together by one {@code exhaustedPattern} classification on any one of
|
||||
@@ -1318,6 +1352,126 @@ public record FleetConfig(
|
||||
}
|
||||
}
|
||||
|
||||
/**
|
||||
* Central allow-list of models any {@code profiles:} entry may name (fleetd ticket: "a central
|
||||
* allow-list of usable models"). Nothing before this block checked a profile's {@code model:}
|
||||
* against anything — it was a free-form string handed straight to the backend adapter, and a
|
||||
* withdrawn or misspelled name failed silently (opencode falls back to a default model rather
|
||||
* than erroring on an unknown {@code -m}) rather than at config load, where a mistake is cheap.
|
||||
*
|
||||
* <p><b>Absent or empty {@code allow:} is "off"</b>, on purpose: this is a large deployment
|
||||
* with gitignored {@code fleetd.yaml} on more than one host, and a change that forced every
|
||||
* operator to enumerate their models before the daemon would start would break every one of
|
||||
* them on upgrade. See {@link FleetConfig#validateModels()}, which is where the allow-list is
|
||||
* actually enforced, at config load.
|
||||
*
|
||||
* <p><b>The list is the authority; a {@code profiles:} entry is checked against it, never the
|
||||
* other way around.</b> Adding or editing a {@code profiles:} entry cannot, by itself, widen
|
||||
* the set of permitted models — only editing {@code models.allow:} itself can. This is the
|
||||
* invariant the ticket asked for: the two blocks are validated in one direction only.
|
||||
*
|
||||
* <p><b>Spawn-time enforcement (fleetd #422, units 2+3) lives outside this record</b> —
|
||||
* {@code CompositePeerLauncher.enforceModelEnabled} and its candidate-set filter read this
|
||||
* block LIVE (through the same kind of supplier {@code weight}/{@code maxLoad} already use),
|
||||
* so the on/off state below is hot: no restart needed. This block itself still only decides
|
||||
* what may be CONFIGURED (membership in {@link #allow}); {@link ModelEntry#enabled} decides
|
||||
* whether a member of that list is currently spawnable. The two questions are deliberately
|
||||
* separate — see {@link ModelEntry}'s javadoc for why turning a model off must never mean
|
||||
* removing it from {@link #allow}.
|
||||
*
|
||||
* @param allow the permitted models, each its own {@link ModelEntry} rather than a bare
|
||||
* string — see that record's javadoc for why. {@code null}/empty ⇒ the block is
|
||||
* treated as absent: {@link FleetConfig#validateModels()} checks nothing.
|
||||
*/
|
||||
@JsonIgnoreProperties(ignoreUnknown = true)
|
||||
public record Models(List<ModelEntry> allow) {
|
||||
|
||||
public Models {
|
||||
allow = (allow == null) ? List.of() : List.copyOf(allow);
|
||||
}
|
||||
|
||||
/**
|
||||
* One permitted model, named as a record rather than a bare string on purpose: this ticket
|
||||
* (fleetd #422) is the "later unit" the original comment here predicted — it hangs an on/off
|
||||
* state ({@link #enabled}) off each entry, and a bare {@code List<String>} could not have
|
||||
* grown that field without changing the YAML shape underneath every operator who already
|
||||
* wrote one. {@link #model()} is intentionally a single flat, opaque-string namespace — a
|
||||
* bare Claude id ({@code claude-sonnet-5}) and an opencode provider-prefixed id
|
||||
* ({@code openai/gpt-5.6-terra}) both fit it unchanged, because {@link
|
||||
* FleetConfig#validateModels()} only ever compares a profile's {@code model:} value against
|
||||
* this string for exact equality; it never parses a provider prefix or branches on a
|
||||
* profile's {@code kind:}.
|
||||
*
|
||||
* @param model the model id exactly as a {@code profiles:} entry's {@code model:} field
|
||||
* would name it
|
||||
* @param enabled {@code false} turns spawning onto this model off; {@code null} (the field
|
||||
* omitted — every config written before fleetd #422 is this shape) or
|
||||
* {@code true} leaves it on. Turning a model off must NEVER remove it from
|
||||
* {@link Models#allow} — {@link FleetConfig#validateModels()} checks
|
||||
* <em>membership</em> only, never the on/off state, so an off entry stays a
|
||||
* valid thing for a {@code profiles:} entry to name; only
|
||||
* {@code CompositePeerLauncher}'s spawn-time gate reads {@link #enabled}.
|
||||
* Collapsing the two — turning a model off by deleting its {@code allow:}
|
||||
* entry — would make {@link FleetConfig#validateModels()} refuse the whole
|
||||
* config reload the moment a still-configured profile names it, which is
|
||||
* exactly the restart-to-flip-a-switch problem this field exists to avoid.
|
||||
* <p>A model id can be named by more than one {@code profiles:} entry (e.g.
|
||||
* {@code deepseek-v4-flash} backs both {@code local} and {@code
|
||||
* local-direct} in the live config) — turning it off disables every profile
|
||||
* that names it, on purpose: the model is what a subscription's rate limit
|
||||
* actually constrains, not any one profile alias for it.
|
||||
*/
|
||||
@JsonIgnoreProperties(ignoreUnknown = true)
|
||||
public record ModelEntry(String model, Boolean enabled) {
|
||||
public ModelEntry {
|
||||
model = (model == null || model.isBlank()) ? null : model.trim();
|
||||
}
|
||||
|
||||
/**
|
||||
* Back-compat form before {@link #enabled} was added (fleetd #422) — the model is
|
||||
* unconditionally on, exactly as every {@code ModelEntry} behaved before this field
|
||||
* existed. Keeps pre-#422 call sites (and any YAML that omits {@code enabled:})
|
||||
* compiling and behaving identically.
|
||||
*/
|
||||
public ModelEntry(String model) {
|
||||
this(model, null);
|
||||
}
|
||||
|
||||
/** {@code true} unless {@link #enabled} is explicitly {@code false} — absent means on. */
|
||||
public boolean isEnabled() {
|
||||
return !Boolean.FALSE.equals(enabled);
|
||||
}
|
||||
}
|
||||
|
||||
/** {@link #allow}'s model ids, as a set for membership checks. Blank/null entries are dropped. */
|
||||
public Set<String> ids() {
|
||||
Set<String> ids = new java.util.LinkedHashSet<>();
|
||||
for (ModelEntry e : allow) {
|
||||
if (e != null && e.model() != null) {
|
||||
ids.add(e.model());
|
||||
}
|
||||
}
|
||||
return Collections.unmodifiableSet(ids);
|
||||
}
|
||||
|
||||
/**
|
||||
* Model ids currently turned off (fleetd #422: {@link ModelEntry#isEnabled()} {@code
|
||||
* false}). Read live by {@code CompositePeerLauncher}'s spawn gate and by {@code
|
||||
* fleet_profiles}/{@code GET /profiles} — both must read this same accessor off the same
|
||||
* live config so the two surfaces cannot disagree about which model is off (the fleetd
|
||||
* #404 lesson: a status field must read the source the behaviour reads).
|
||||
*/
|
||||
public Set<String> offIds() {
|
||||
Set<String> off = new java.util.LinkedHashSet<>();
|
||||
for (ModelEntry e : allow) {
|
||||
if (e != null && e.model() != null && !e.isEnabled()) {
|
||||
off.add(e.model());
|
||||
}
|
||||
}
|
||||
return Collections.unmodifiableSet(off);
|
||||
}
|
||||
}
|
||||
|
||||
/**
|
||||
* The terminal → lead-name map seeded from the legacy singular {@code primary:} pin (CB-530).
|
||||
*
|
||||
@@ -1569,7 +1723,7 @@ public record FleetConfig(
|
||||
"lifecycle", "spawnReadyTimeoutMs", "spawnReadyPollMs", "broker", "primary", "fleet",
|
||||
"leadHeartbeat", "health", "placement", "auth", "configReload", "quarantineCooldownSeconds",
|
||||
"memberCredentials", "coordinator", "worktreeGroup", "memberLoginShell", "memberSkills",
|
||||
"idleSleepGuard");
|
||||
"idleSleepGuard", "models");
|
||||
|
||||
/** Load and validate config from {@code path}. */
|
||||
public static FleetConfig load(Path path) {
|
||||
@@ -2254,10 +2408,15 @@ public record FleetConfig(
|
||||
// enabled: true — see its javadoc), so defaulting the block here would change nothing a
|
||||
// reader observes and would only obscure that "block omitted" and "block present and
|
||||
// enabled" are deliberately the same outcome.
|
||||
// models is left as-is, like broker/primary/coordinator above: null/empty is "off", and an
|
||||
// absent block must validate nothing (see Models's javadoc) — defaulting it here to an
|
||||
// empty Models would be a no-op for validateModels() either way, since an empty allow-list
|
||||
// already means "check nothing", so there is nothing to gain and one more null check to
|
||||
// avoid by leaving it exactly as configured.
|
||||
return new FleetConfig(b, herdrSocket, memberHerdrSocket, profiles, g, worktreeRoot, l, timeout, pollMs,
|
||||
broker, primary, f, leadHeartbeat, health, placementOrDefault, a, configReload,
|
||||
quarantineCooldown, mc, coordinator, worktreeGroup, memberLoginShell, memberSkills,
|
||||
idleSleepGuard);
|
||||
idleSleepGuard, models);
|
||||
}
|
||||
|
||||
/**
|
||||
@@ -2477,6 +2636,118 @@ public record FleetConfig(
|
||||
}
|
||||
}
|
||||
|
||||
/**
|
||||
* Reject a {@code profiles:} entry whose {@code model:} is not on the configured {@link
|
||||
* #models} allow-list.
|
||||
*
|
||||
* <p>Absent or empty {@code models.allow:} validates nothing — see {@link Models}'s javadoc:
|
||||
* an existing config with no such block must keep working exactly as it does today. Once the
|
||||
* operator declares at least one entry, every profile's {@code model:} (when set — a profile
|
||||
* may legitimately leave it {@code null}, e.g. a {@code subscription: true} profile relying on
|
||||
* the account's own default) must equal one of {@link Models#ids()} exactly. The comparison is
|
||||
* a flat string match: a bare Claude id and an opencode {@code provider/model} id are both
|
||||
* just opaque strings here, so nothing here needs to know which {@code kind:} a profile runs.
|
||||
*
|
||||
* <p>The check runs one direction only, by construction: it reads {@link #profiles} and
|
||||
* {@link #models}, and only ever adds to {@code bad} when a profile's model is missing from
|
||||
* the allow-list. Nothing here can be satisfied by widening a {@code profiles:} entry — only
|
||||
* editing {@code models.allow:} itself changes what passes. That is the invariant the ticket
|
||||
* asked for: the list is the authority, profiles are checked against it.
|
||||
*
|
||||
* @throws IllegalStateException when any profile names a model outside the configured
|
||||
* allow-list, naming both the model and the profile that wanted it
|
||||
*/
|
||||
public void validateModels() {
|
||||
if (models == null || models.allow().isEmpty()) {
|
||||
return;
|
||||
}
|
||||
Set<String> allowed = models.ids();
|
||||
List<String> bad = new ArrayList<>();
|
||||
profiles.forEach((name, p) -> {
|
||||
String model = p.model();
|
||||
if (model != null && !model.isBlank() && !allowed.contains(model)) {
|
||||
bad.add("profile '" + name + "' names model '" + model + "', which is not in "
|
||||
+ "models.allow: (have: " + allowed + ").");
|
||||
}
|
||||
});
|
||||
if (!bad.isEmpty()) {
|
||||
throw new IllegalStateException("refusing to start: " + String.join(" ", bad));
|
||||
}
|
||||
}
|
||||
|
||||
/**
|
||||
* Runs every validator this class declares — found by reflection, not by name.
|
||||
*
|
||||
* <p>fleetd ticket "central allow-list of usable models", follow-up: mutation testing found
|
||||
* that although each of the six validators above was well pinned on its own, nothing proved
|
||||
* either real caller ({@code Fleetd.main} and {@link ConfigRef#reload()}) still
|
||||
* invoked it — deleting a call site left the full suite green. The fix is not a seventh test
|
||||
* per caller; a hand-maintained list of six names here would have the exact same defect its
|
||||
* own javadoc would warn against: the seventh validator someone adds next month has no reason
|
||||
* to be added to it. So this method does not name any validator. It sweeps {@link
|
||||
* #getClass()}'s own public, no-argument, {@code void} methods whose name starts with {@code
|
||||
* "validate"} (excluding itself) and invokes every one it finds, via {@link
|
||||
* #invokeAllValidators}. A new {@code validateXxx()} method is therefore wired into both
|
||||
* callers the moment it is written — there is no second step to forget, and so no state in
|
||||
* which it silently never runs.
|
||||
*
|
||||
* <p>{@code Fleetd.main} and {@link ConfigRef#reload()} each call this one method instead of
|
||||
* the six individually — see the comments at those two call sites for why
|
||||
* each must run it.
|
||||
*
|
||||
* <p>Methods run in a fixed (alphabetical) order, so a config with more than one violation
|
||||
* always names the same one first, on every run.
|
||||
*
|
||||
* @throws IllegalStateException (or whatever unchecked exception a validator itself throws),
|
||||
* propagated unchanged from the first validator, in that order,
|
||||
* that finds a problem
|
||||
*/
|
||||
public void validateAll() {
|
||||
invokeAllValidators(this);
|
||||
}
|
||||
|
||||
/**
|
||||
* The reflective sweep behind {@link #validateAll()}, kept as its own method — taking any
|
||||
* {@code target}, not just {@code this} — so a test can prove the MECHANISM is generic (it
|
||||
* would sweep a seventh {@code validateXxx()} method added to any class, not just something
|
||||
* special-cased to today's six on {@link FleetConfig}) without needing to add a real, unwanted
|
||||
* seventh validator to this class just to exercise that claim. See {@code
|
||||
* FleetConfigValidateAllTest} for that proof.
|
||||
*
|
||||
* @param target an object whose public, no-argument, {@code void} methods named {@code
|
||||
* validateXxx} (any name starting with {@code "validate"}, excluding {@code
|
||||
* validateAll} itself) should all run, in alphabetical-by-name order
|
||||
*/
|
||||
static void invokeAllValidators(Object target) {
|
||||
List<Method> methods = new ArrayList<>();
|
||||
for (Method m : target.getClass().getMethods()) {
|
||||
if (Modifier.isPublic(m.getModifiers())
|
||||
&& m.getParameterCount() == 0
|
||||
&& m.getReturnType() == void.class
|
||||
&& m.getName().startsWith("validate")
|
||||
&& !m.getName().equals("validateAll")) {
|
||||
methods.add(m);
|
||||
}
|
||||
}
|
||||
methods.sort(Comparator.comparing(Method::getName));
|
||||
for (Method m : methods) {
|
||||
try {
|
||||
m.invoke(target);
|
||||
} catch (InvocationTargetException e) {
|
||||
Throwable cause = e.getCause();
|
||||
if (cause instanceof RuntimeException re) {
|
||||
throw re;
|
||||
}
|
||||
if (cause instanceof Error err) {
|
||||
throw err;
|
||||
}
|
||||
throw new IllegalStateException("validator " + m.getName() + " failed", cause);
|
||||
} catch (IllegalAccessException e) {
|
||||
throw new IllegalStateException("cannot invoke validator " + m.getName(), e);
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
/** True for the loopback addresses and the unspecified-but-local forms we treat as same-host. */
|
||||
private static boolean isLoopbackBind(String host) {
|
||||
if (host == null || host.isBlank()) {
|
||||
|
||||
@@ -3,7 +3,7 @@ package dev.ltms.fleet.herdr;
|
||||
import com.fasterxml.jackson.databind.JsonNode;
|
||||
|
||||
/**
|
||||
* Client face onto the herdr daemon (protocol 14, herdr 0.7.0).
|
||||
* Client face onto the herdr daemon (protocol 19, herdr 0.8.0).
|
||||
*
|
||||
* <p>This is the ONLY thing in {@code fleetd} that speaks to herdr. Every method
|
||||
* maps to a herdr JSON-RPC call over its Unix domain socket. Requests are
|
||||
|
||||
@@ -8,7 +8,7 @@ import com.fasterxml.jackson.databind.node.ObjectNode;
|
||||
import java.nio.charset.StandardCharsets;
|
||||
|
||||
/**
|
||||
* Wire codec for herdr's newline-delimited JSON-RPC (protocol 14).
|
||||
* Wire codec for herdr's newline-delimited JSON-RPC (protocol 19).
|
||||
*
|
||||
* <p>Split out from the socket so the framing rules — the ones that actually bit us
|
||||
* during the spike (id MUST be a string; response carries {@code result} or
|
||||
|
||||
@@ -698,17 +698,57 @@ public final class CompletionResolver implements TurnListener {
|
||||
}
|
||||
|
||||
/**
|
||||
* Coverage summary for the CB-578 stage A exhausted-pattern classification, logged at startup
|
||||
* What an unset pattern key means for the classification it configures (fleetd#415).
|
||||
* {@code coverage()} cannot infer this from the key's name — the two keys it currently
|
||||
* describes disagree on it, and a string comparison on the name would just move the same bug
|
||||
* to a new spot — so every caller must state it explicitly.
|
||||
*
|
||||
* <p><strong>This alone does not prove a caller passes the right one for its key.</strong> A
|
||||
* test that calls {@code coverage()} directly and supplies the meaning itself only proves this
|
||||
* enum is worded correctly, never that {@code Fleetd}'s two call sites pair each key with its
|
||||
* true meaning — that pairing is #415's actual defect. Measured on review: swapping the two
|
||||
* {@code UnsetMeaning} arguments at those call sites (giving {@code exhaustedPattern} the
|
||||
* built-in-default wording and {@code errorPattern} the off wording — #415's exact defect with
|
||||
* the keys exchanged) compiled with 0 errors and left all 1506 existing tests green. See
|
||||
* {@code dev.ltms.fleet.Fleetd#exhaustedPatternCoverageLine}/{@code #errorPatternCoverageLine}
|
||||
* and {@code FleetdPatternCoverageLineTest}, which exists specifically to catch that swap.
|
||||
*/
|
||||
public enum UnsetMeaning {
|
||||
/** No fallback exists: a profile with no configured pattern truly has this classification off. */
|
||||
OFF,
|
||||
/** A built-in pattern applies when unset: the classification still runs for that profile. */
|
||||
BUILT_IN_DEFAULT
|
||||
}
|
||||
|
||||
/**
|
||||
* Coverage summary for a fleetd#201/CB-578-style pattern-key classification, logged at startup
|
||||
* the way {@link dev.ltms.fleet.health.FleetHealthMonitor#coverage} is — so an operator can
|
||||
* see whether the classification is on, and for which profiles, without reading every
|
||||
* profile's config by hand.
|
||||
*
|
||||
* <p>fleetd#415: this method measures <em>pattern coverage</em> — how many profiles set the
|
||||
* key — which is not the same thing as <em>feature state</em> for a key with a fallback. For
|
||||
* {@code errorPattern}, an empty {@code configuredProfiles} still runs the classification
|
||||
* against {@code CompletionResolver}'s built-in compatibility pattern ({@link #BACKEND_ERROR}
|
||||
* at line ~84); for {@code exhaustedPattern} there is no fallback, so empty really does mean
|
||||
* off. {@code unsetMeaning} is the single, required source of that fact — see
|
||||
* {@link dev.ltms.fleet.config.FleetConfig#rejectMalformedProfilePatterns} lines ~2029-2032 for
|
||||
* where it is documented for config authors. It is a required parameter, not a defaulted
|
||||
* overload: a third pattern key added later must supply one to compile at all, rather than
|
||||
* silently inheriting whichever wording this method happened to default to.
|
||||
*
|
||||
* @param allProfiles every configured profile name
|
||||
* @param configuredProfiles the subset of {@code allProfiles} that carry an exhausted pattern
|
||||
* @param configuredProfiles the subset of {@code allProfiles} that carry the pattern
|
||||
*/
|
||||
public static String coverage(String patternKey, Set<String> allProfiles, Set<String> configuredProfiles) {
|
||||
public static String coverage(String patternKey, UnsetMeaning unsetMeaning, Set<String> allProfiles,
|
||||
Set<String> configuredProfiles) {
|
||||
if (configuredProfiles.isEmpty()) {
|
||||
return "off (no profile has an " + patternKey + " configured; profiles: " + sorted(allProfiles) + ")";
|
||||
return switch (unsetMeaning) {
|
||||
case OFF -> "off (no profile has an " + patternKey + " configured; profiles: "
|
||||
+ sorted(allProfiles) + ")";
|
||||
case BUILT_IN_DEFAULT -> "built-in default for all profiles (no profile customises "
|
||||
+ patternKey + "; profiles: " + sorted(allProfiles) + ")";
|
||||
};
|
||||
}
|
||||
Set<String> unconfigured = new TreeSet<>(allProfiles);
|
||||
unconfigured.removeAll(configuredProfiles);
|
||||
|
||||
@@ -0,0 +1,91 @@
|
||||
package dev.ltms.fleet.inject;
|
||||
|
||||
import dev.ltms.fleet.config.FleetConfig;
|
||||
|
||||
import java.util.Map;
|
||||
import java.util.Objects;
|
||||
import java.util.concurrent.ConcurrentHashMap;
|
||||
import java.util.function.Supplier;
|
||||
import java.util.regex.Pattern;
|
||||
|
||||
/**
|
||||
* fleetd #446: the single live source for "is {@code exhaustedPattern} configured for this
|
||||
* profile, and what does it compile to". Read fresh off the config supplier on every call — the
|
||||
* same reason {@code CompositePeerLauncher#models0} is a live supplier read rather than a value
|
||||
* captured at construction (see that class's doc) — so an operator can arm or disarm usage-limit
|
||||
* detection for a profile by editing {@code exhaustedPattern} and reloading, with no restart.
|
||||
*
|
||||
* <p>Before this class, {@code exhaustedPattern} was compiled once into a {@code Map<String,
|
||||
* Pattern>} built inside {@code Fleetd.main} at startup ({@code ConfigRef}'s class doc used to
|
||||
* list it under <em>Deferred</em>, CB-578 stage A) — the asymmetry fleetd #446 exists to close:
|
||||
* an operator could turn a model off at runtime ({@code models.allow}'s {@code enabled: false},
|
||||
* hot since fleetd #422) but could not arm the detector that would tell them to, without a
|
||||
* restart. That was backwards for a feature whose whole point is to react while the fleet runs.
|
||||
*
|
||||
* <p>{@link #patternFor} is what {@link CompletionResolver} enforces on (via the {@link
|
||||
* ExhaustedPatternLookup} production wiring in {@code Fleetd.main}); {@link #armed} is what
|
||||
* {@code fleet_profiles}/{@code GET /profiles} report as {@code exhaustionDetectionArmed}. Both
|
||||
* read this ONE object, so the report can never disagree with the behaviour — the same rule
|
||||
* {@code CompositePeerLauncher#modelGateState()}'s javadoc states for the model gate's own
|
||||
* armed/off pair (fleetd #404): "armed" and "which models are off" must come from one read of the
|
||||
* same accessor the gate enforces on.
|
||||
*
|
||||
* <h2>Cache eviction</h2>
|
||||
* Cached by profile NAME, not by pattern text. A {@code Map<String, Pattern>} keyed by the
|
||||
* pattern STRING would grow by one entry per distinct regex ever typed for any profile across the
|
||||
* daemon's uptime — unbounded in practice, because tuning a regex to match a backend's exact
|
||||
* wording is exactly the kind of edit an operator makes several times while getting it right, and
|
||||
* every edit-and-reload cycle would leave the previous attempt's compiled {@link Pattern} behind
|
||||
* forever. Keyed by profile name instead, this cache holds at most one entry per profile name that
|
||||
* has ever been looked up — and that key space is bounded by the (small, human-authored) set of
|
||||
* configured profiles, which changes only on a restart: adding or removing a profile is itself a
|
||||
* deferred key (a new backend needs its own launcher, built once — see {@code ConfigRef}'s class
|
||||
* doc), so profile names do not churn the way pattern text does. Re-editing an EXISTING profile's
|
||||
* {@code exhaustedPattern} — the case this class exists to make hot — simply overwrites that
|
||||
* profile's one cache entry; it never adds a new one.
|
||||
*/
|
||||
public final class LiveExhaustedPatterns {
|
||||
|
||||
/** A compiled pattern paired with the source string it was compiled from, for change detection. */
|
||||
private record Cached(String source, Pattern compiled) {
|
||||
}
|
||||
|
||||
private final Supplier<Map<String, FleetConfig.Profile>> profiles;
|
||||
private final ConcurrentHashMap<String, Cached> cache = new ConcurrentHashMap<>();
|
||||
|
||||
public LiveExhaustedPatterns(Supplier<Map<String, FleetConfig.Profile>> profiles) {
|
||||
this.profiles = Objects.requireNonNull(profiles, "profiles");
|
||||
}
|
||||
|
||||
/**
|
||||
* The compiled {@code exhaustedPattern} currently configured for {@code profileName}, or
|
||||
* {@code null} when that profile is unknown or has none configured. Recompiles only when the
|
||||
* live pattern text differs from what is cached for this profile name; {@code
|
||||
* FleetConfig#rejectMalformedProfilePatterns} already refuses a config (at load and at reload)
|
||||
* whose {@code exhaustedPattern} does not compile, so this is not expected to throw in
|
||||
* production — it is not defended against here for that reason, the same trust
|
||||
* {@code CompositePeerLauncher#models0} places in config validation having already run.
|
||||
*/
|
||||
public Pattern patternFor(String profileName) {
|
||||
if (profileName == null) {
|
||||
return null;
|
||||
}
|
||||
FleetConfig.Profile profile = profiles.get().get(profileName);
|
||||
if (profile == null || !profile.hasExhaustedPattern()) {
|
||||
return null;
|
||||
}
|
||||
String source = profile.exhaustedPattern();
|
||||
Cached cached = cache.get(profileName);
|
||||
if (cached != null && cached.source().equals(source)) {
|
||||
return cached.compiled();
|
||||
}
|
||||
Cached fresh = new Cached(source, Pattern.compile(source));
|
||||
cache.put(profileName, fresh);
|
||||
return fresh.compiled();
|
||||
}
|
||||
|
||||
/** Whether {@code profileName} currently has a usage-limit pattern configured (live). */
|
||||
public boolean armed(String profileName) {
|
||||
return patternFor(profileName) != null;
|
||||
}
|
||||
}
|
||||
@@ -37,6 +37,7 @@ import io.modelcontextprotocol.spec.McpSchema;
|
||||
import com.fasterxml.jackson.databind.ObjectMapper;
|
||||
import jakarta.servlet.http.HttpServlet;
|
||||
|
||||
import java.util.ArrayList;
|
||||
import java.util.LinkedHashMap;
|
||||
import java.util.List;
|
||||
import java.util.Map;
|
||||
@@ -119,9 +120,57 @@ public final class FleetMcp {
|
||||
/**
|
||||
* CB-578 stage B quarantine facts used by {@code fleet_profiles}: a profile → credential id
|
||||
* lookup, plus the shared {@link BackendQuarantine} to read remaining cooldowns off.
|
||||
*
|
||||
* @param exhaustedPatternArmed fleetd #395, made live by fleetd #446: profile → whether that
|
||||
* profile's {@code exhaustedPattern} is configured right now (see {@code
|
||||
* FleetConfig.Profile#hasExhaustedPattern}), i.e. whether a backend refusal
|
||||
* on it can EVER be classified {@code BACKEND_EXHAUSTED} and quarantine its
|
||||
* credential. Bundled here, not a separate Source, because it answers the
|
||||
* exact question {@code fleet_profiles}'s quarantine facts already answer
|
||||
* for a QUARANTINED profile — "can this profile's usage limit ever be
|
||||
* caught?" — just for every profile, not only one currently caught. In
|
||||
* production this reads {@link dev.ltms.fleet.inject.LiveExhaustedPatterns#armed} —
|
||||
* the SAME live accessor {@link dev.ltms.fleet.inject.CompletionResolver}
|
||||
* enforces on — so a reload that arms or disarms detection is reported
|
||||
* correctly on the very next call, no restart (see that class's doc).
|
||||
* @param modelFor fleetd #446 criterion 3: profile → the {@code model:} it runs, or
|
||||
* {@code null} when the profile names none. Read live off the current
|
||||
* config, the same way {@code credentialIdFor} already is. Used to name,
|
||||
* in a QUARANTINED profile's row, which of the operator's configured models
|
||||
* the fix (\"set {@code enabled: false} on it under {@code models.allow}\")
|
||||
* actually applies to.
|
||||
* @param reasonFor fleetd #446 criterion 3: credential id → the backend text that triggered
|
||||
* its most recent quarantine, or {@code null} when none is known (e.g. an
|
||||
* inert source, or a quarantine recorded before this field existed). Never
|
||||
* consulted on its own to decide whether a credential is quarantined —
|
||||
* callers gate on {@link BackendQuarantine#remainingSeconds} first, exactly
|
||||
* like {@code credentialId} itself, so a stale reason left behind after a
|
||||
* quarantine expires is never surfaced.
|
||||
*/
|
||||
public record QuarantineSource(Function<String, String> credentialIdFor, BackendQuarantine quarantine) {
|
||||
/** Inert source — no profile is ever reported quarantined. Explicit stand-in, not a default. */
|
||||
public record QuarantineSource(Function<String, String> credentialIdFor, BackendQuarantine quarantine,
|
||||
Function<String, Boolean> exhaustedPatternArmed,
|
||||
Function<String, String> modelFor,
|
||||
Function<String, String> reasonFor) {
|
||||
/**
|
||||
* Backward-compatible 3-arg form, before fleetd #446 added {@code modelFor}/{@code
|
||||
* reasonFor} — neither is ever reported. Keeps every pre-existing call site (production and
|
||||
* test) compiling and behaving identically for the quarantine facts they actually asked for.
|
||||
*/
|
||||
public QuarantineSource(Function<String, String> credentialIdFor, BackendQuarantine quarantine,
|
||||
Function<String, Boolean> exhaustedPatternArmed) {
|
||||
this(credentialIdFor, quarantine, exhaustedPatternArmed, _ -> null, _ -> null);
|
||||
}
|
||||
|
||||
/**
|
||||
* Backward-compatible 2-arg form, before fleetd #395 added {@code exhaustedPatternArmed} —
|
||||
* reports every profile unarmed. Keeps every pre-existing call site (production and test)
|
||||
* compiling and behaving identically for the quarantine facts they actually asked for.
|
||||
*/
|
||||
public QuarantineSource(Function<String, String> credentialIdFor, BackendQuarantine quarantine) {
|
||||
this(credentialIdFor, quarantine, _ -> false);
|
||||
}
|
||||
|
||||
/** Inert source — no profile is ever reported quarantined or armed. Explicit stand-in, not a default. */
|
||||
public static QuarantineSource none() { return new QuarantineSource(_ -> null, BackendQuarantine.none()); }
|
||||
}
|
||||
|
||||
@@ -346,7 +395,7 @@ public final class FleetMcp {
|
||||
String self = callerTerminal(exchange);
|
||||
McpSchema.CallToolResult denied = deny(exchange, toolAction("fleet_reply", req.arguments()), self);
|
||||
if (denied != null) return denied;
|
||||
return reply(messages, self, str(req.arguments(), "content"));
|
||||
return reply(messages, self, principal(exchange).role(), str(req.arguments(), "content"));
|
||||
};
|
||||
// fleet_ask (CB-205): a worker's mid-turn question — identity from the CONNECTION.
|
||||
BiFunction<McpSyncServerExchange, McpSchema.CallToolRequest, McpSchema.CallToolResult> askHandler =
|
||||
@@ -366,10 +415,11 @@ public final class FleetMcp {
|
||||
(exchange, req) -> {
|
||||
Map<String, Object> a = req.arguments();
|
||||
String target = str(a, "target");
|
||||
String coordId = str(a, "coordId");
|
||||
// The action depends on the ARGUMENTS, not on the tool name -- see pollAction.
|
||||
McpSchema.CallToolResult denied = deny(exchange, toolAction("fleet_poll", a), target);
|
||||
if (denied != null) return denied;
|
||||
return poll(messages, str(a, "ticket"), target);
|
||||
return poll(messages, leadChannel, str(a, "ticket"), target, coordId);
|
||||
};
|
||||
// CB-307 Increment 3: per-msgId ack (not needed in v1 but supported by the inbox).
|
||||
// Acking removes a reply from the inbox, so it is a drain, not a read.
|
||||
@@ -405,7 +455,8 @@ public final class FleetMcp {
|
||||
return listFleet(workers, sessions, messages, capacity, healthCoverage, quarantine, outage,
|
||||
leadSeats, callers == null ? Map.of() : callers.leads(),
|
||||
callerTerminal(exchange),
|
||||
new CoordinationSource(leadChannel, peers));
|
||||
new CoordinationSource(leadChannel, peers),
|
||||
coordinatorVisibleTo(principal(exchange)));
|
||||
};
|
||||
BiFunction<McpSyncServerExchange, McpSchema.CallToolRequest, McpSchema.CallToolResult> stopHandler =
|
||||
(exchange, req) -> {
|
||||
@@ -558,6 +609,20 @@ public final class FleetMcp {
|
||||
}
|
||||
}
|
||||
|
||||
/**
|
||||
* fleetd #439: only the primary may read {@code fleet_list}'s {@code coordinator} row —
|
||||
* lead-to-lead coordination state (coord-ids, mailbox facts, held-message previews), never the
|
||||
* roster. Split out of the {@code fleet_list} handler, same reason as {@link #denyFor} and
|
||||
* {@link #recordPrimarySingleton}: the decision must be unit-testable without fabricating an
|
||||
* SDK {@code McpSyncServerExchange}, and the handler must call this named predicate rather than
|
||||
* inlining the check, so a future edit cannot silently pass a literal instead of asking who
|
||||
* called ({@code FleetMcpAuthzTest.theFleetListHandlerActuallyConsultsCoordinatorVisibleTo}
|
||||
* reads the source and asserts the handler calls this method by name, not a literal).
|
||||
*/
|
||||
static boolean coordinatorVisibleTo(Principal caller) {
|
||||
return caller.isPrimary();
|
||||
}
|
||||
|
||||
/** The worker identity resolved from this call's connection, or {@code null} if the primary. */
|
||||
private static String callerTerminal(McpSyncServerExchange exchange) {
|
||||
Object v = exchange.transportContext().get(CALLER_TERMINAL);
|
||||
@@ -791,29 +856,42 @@ public final class FleetMcp {
|
||||
|
||||
/**
|
||||
* Which authorization action a {@code fleet_poll} call needs, decided by its arguments
|
||||
* (fleetd #272).
|
||||
* (fleetd #272, widened by fleetd #421).
|
||||
*
|
||||
* <p>{@code fleet_poll} is <strong>two operations behind one tool name</strong>. With {@code
|
||||
* ticket} it observes an async delegation and changes nothing, which is a {@link
|
||||
* <p>{@code fleet_poll} is now <strong>three operations behind one tool name</strong>. With
|
||||
* {@code ticket} it observes an async delegation and changes nothing, which is a {@link
|
||||
* Authz.Action#READ}. With {@code target} it calls {@link MessageService#drainReplies} on that
|
||||
* session -- the replies are removed from the inbox and a second call returns nothing -- so it
|
||||
* is a {@link Authz.Action#DRAIN}, the same gate {@code fleet_ack} already uses for removing a
|
||||
* single message, and the same one the REST path uses at {@code FleetApp.drainReplies}.
|
||||
* single message, and the same one the REST path uses at {@code FleetApp.drainReplies}. With
|
||||
* {@code coordId} it reads (never acks) this daemon's own held lead-to-lead mail, which is a
|
||||
* {@link Authz.Action#COORD_READ} -- <strong>not</strong> {@code READ}, even though nothing is
|
||||
* consumed: {@code READ}'s grant is open to every authenticated role on the premise that the
|
||||
* roster carries no secrets, and a lead-to-lead body is not the roster. Mapping a non-destructive
|
||||
* peer-mail read to {@code READ} would let any worker read every peer lead's mail in full.
|
||||
*
|
||||
* <p>Until this method existed the handler passed a constant {@code READ} for both branches.
|
||||
* {@code READ} is open to every authenticated role, so any worker could read a peer's id out of
|
||||
* {@code fleet_list} and destroy the replies that peer had queued for the primary. The gate
|
||||
* failed open, and it did so because the required action is a function of the arguments while
|
||||
* the handler chose it before looking at them.
|
||||
* <p>Before this method existed (fleetd #272) the handler passed a constant {@code READ} for
|
||||
* both of the original branches. {@code READ} is open to every authenticated role, so any
|
||||
* worker could read a peer's id out of {@code fleet_list} and destroy the replies that peer had
|
||||
* queued for the primary. The gate failed open, and it did so because the required action is a
|
||||
* function of the arguments while the handler chose it before looking at them.
|
||||
*
|
||||
* <p>The choice lives in this method, and not inline in the handler, so that a test can assert
|
||||
* the mapping the handler actually uses. {@code FleetMcpAuthzTest} already checked every
|
||||
* {@link Authz.Action} against every {@link Role} and passed throughout -- it tested the policy
|
||||
* table, which was correct, while the defect was in which action the caller handed it.
|
||||
*
|
||||
* @param target the {@code target} argument of the call, or {@code null}/blank when absent
|
||||
* <p>Checked first, and exclusively of {@code target}: a call naming {@code coordId} is reading
|
||||
* a different inbox entirely (this daemon's own lead channel, never a worker's), so it takes
|
||||
* priority over whatever {@code target} might also say.
|
||||
*
|
||||
* @param target the {@code target} argument of the call, or {@code null}/blank when absent
|
||||
* @param coordId the {@code coordId} argument of the call, or {@code null}/blank when absent
|
||||
*/
|
||||
static Authz.Action pollAction(String target) {
|
||||
static Authz.Action pollAction(String target, String coordId) {
|
||||
if (!isBlank(coordId)) {
|
||||
return Authz.Action.COORD_READ;
|
||||
}
|
||||
return isBlank(target) ? Authz.Action.READ : Authz.Action.DRAIN;
|
||||
}
|
||||
|
||||
@@ -828,7 +906,7 @@ public final class FleetMcp {
|
||||
case "fleet_reply" -> Authz.Action.REPLY;
|
||||
case "fleet_ask" -> Authz.Action.ASK;
|
||||
case "fleet_status", "fleet_list", "fleet_profiles", "fleet_whoami" -> Authz.Action.READ;
|
||||
case "fleet_poll" -> pollAction(str(arguments, "target"));
|
||||
case "fleet_poll" -> pollAction(str(arguments, "target"), str(arguments, "coordId"));
|
||||
case "fleet_ack" -> Authz.Action.DRAIN;
|
||||
case "fleet_spawn" -> Authz.Action.SPAWN;
|
||||
case "fleet_stop" -> Authz.Action.STOP;
|
||||
@@ -838,6 +916,21 @@ public final class FleetMcp {
|
||||
|
||||
/** {@code fleet_poll}: check an async delegation by ticket, or drain a worker's inbox by target. */
|
||||
static McpSchema.CallToolResult poll(MessageService messages, String ticket, String target) {
|
||||
return poll(messages, null, ticket, target, null);
|
||||
}
|
||||
|
||||
/**
|
||||
* As above, plus (fleetd #421) a held-peer-mail read when {@code coordId} is present: returns
|
||||
* this daemon's own held lead-to-lead messages, in full, without acking them. Checked first and
|
||||
* exclusively of {@code ticket}/{@code target} — see {@link #pollAction}'s javadoc for why this
|
||||
* is a different inbox (this daemon's own {@link LeadChannel}) that authorizes differently
|
||||
* ({@link Authz.Action#COORD_READ}, primary-only) from either of the original two branches.
|
||||
*/
|
||||
static McpSchema.CallToolResult poll(MessageService messages, LeadChannel leadChannel, String ticket,
|
||||
String target, String coordId) {
|
||||
if (!isBlank(coordId)) {
|
||||
return pollHeldPeerMail(leadChannel, coordId);
|
||||
}
|
||||
if (!isBlank(target)) {
|
||||
var replies = messages.drainReplies(target);
|
||||
if (replies.isEmpty()) {
|
||||
@@ -865,22 +958,69 @@ public final class FleetMcp {
|
||||
};
|
||||
}
|
||||
|
||||
/**
|
||||
* fleetd #421: a lead's own held lead-to-lead mail, read without consuming it.
|
||||
*
|
||||
* <p>{@link LeadChannel#peek} is non-destructive, so calling this twice returns the same
|
||||
* bodies, and {@code fleet_list}'s {@code coordinator.held[]} is unaffected — this adds a
|
||||
* read, it never acks. Never let this short-circuit the delivery contract: {@code
|
||||
* LeadCoordLoop}'s javadoc explains why a message must stay unacked until actually delivered,
|
||||
* and that is unchanged here.
|
||||
*
|
||||
* <p>{@code coordId} must be THIS daemon's own coord-id ({@code fleet_list}'s
|
||||
* {@code coordinator.selfId}) — there is no route here to read a PEER's outbound mail, only
|
||||
* your own inbound mail. Requiring the caller to echo its own id catches the exact confusion
|
||||
* that opened fleetd #421: the original failed attempt passed a PEER's id ("fleet01") as
|
||||
* {@code fleet_poll}'s {@code target}, expecting to read that peer's messages, and got the same
|
||||
* silent {@code []} as a genuinely empty worker inbox. This refuses the same mistake here with a
|
||||
* reason, instead of a second silent wrong answer.
|
||||
*/
|
||||
static McpSchema.CallToolResult pollHeldPeerMail(LeadChannel leadChannel, String coordId) {
|
||||
if (leadChannel == null) {
|
||||
return error("lead coordination is not configured (no coordinator: block) — there is "
|
||||
+ "no held peer mail to read.");
|
||||
}
|
||||
String selfId = leadChannel.selfCoordId();
|
||||
if (!coordId.equals(selfId)) {
|
||||
return error("coordId \"" + coordId + "\" is not this daemon's own coord-id (\"" + selfId
|
||||
+ "\"). fleet_poll reads only YOUR OWN held mail — pass your own coordId "
|
||||
+ "(fleet_list's coordinator.selfId), not a peer's.");
|
||||
}
|
||||
return text(json(leadChannel.peek().stream().map(FleetMcp::heldMailView).toList()));
|
||||
}
|
||||
|
||||
/** The full body of one held lead-to-lead message — never truncated, unlike {@link #heldView}. */
|
||||
private static Map<String, Object> heldMailView(LeadMessage m) {
|
||||
Map<String, Object> row = new LinkedHashMap<>();
|
||||
row.put("msgId", m.msgId());
|
||||
row.put("from", m.from());
|
||||
row.put("content", m.content());
|
||||
return row;
|
||||
}
|
||||
|
||||
/**
|
||||
* {@code fleet_reply}: the worker returns its structured answer, resolving the awaiting send
|
||||
* or — when no send is open — queueing the reply in the inbox for later drain (CB-307).
|
||||
* {@code callerTerminal} is resolved from the connection (never an argument); a {@code null}
|
||||
* means the caller is not a known worker (e.g. the primary called it by mistake).
|
||||
* {@code callerTerminal} and {@code callerRole} are resolved from the connection (never an
|
||||
* argument). A {@code null} terminal means the caller is not a known worker. A PRIMARY with a
|
||||
* terminal is a lead and must use {@code fleet_send}, because reply has no peer-lead route.
|
||||
*
|
||||
* <p>fleetd #365: the result text names which of those actually happened
|
||||
* ({@link MessageService.ReplyOutcome#description()}) instead of the single word "delivered"
|
||||
* for both — a queued reply is a real success, but it is not the same fact as one that resolved
|
||||
* a live waiter, and the caller could not previously tell them apart.
|
||||
*/
|
||||
static McpSchema.CallToolResult reply(MessageService messages, String callerTerminal, String content) {
|
||||
static McpSchema.CallToolResult reply(MessageService messages, String callerTerminal, Role callerRole, String content) {
|
||||
if (callerTerminal == null) {
|
||||
return error("fleet_reply is for workers only — could not identify the calling worker "
|
||||
+ "from the connection");
|
||||
}
|
||||
if (callerRole == Role.PRIMARY) {
|
||||
return error("fleet_reply has no route to a peer lead. Use fleet_send{coordId: ...} for a peer on another "
|
||||
+ "daemon or fleet_send{sessionId: ...} for a peer on this host. fleet_reply resolves a member's "
|
||||
+ "blocked fleet_send, and a peer's coord-id message is durable and non-blocking, so there is "
|
||||
+ "nothing for it to resolve.");
|
||||
}
|
||||
// fleetd #302: isBlank, not == null, to match fleet_send's own guard above. MessageService
|
||||
// .reply now REJECTS blank content, and this handler is a bare BiFunction with no try/catch
|
||||
// around it — so a whitespace-only fleet_reply would leave here as an uncaught
|
||||
@@ -898,7 +1038,11 @@ public final class FleetMcp {
|
||||
if (isBlank(target) || isBlank(msgId)) {
|
||||
return error("target and msgId are required");
|
||||
}
|
||||
messages.ackReply(target, msgId);
|
||||
if (!messages.ackReply(target, msgId)) {
|
||||
return error(msgId + " is not in " + target + "'s reply inbox (wrong id, wrong target, "
|
||||
+ "or already acked). Held lead-to-lead (peer) mail cannot be acked this way — "
|
||||
+ "read it with fleet_poll{coordId}.");
|
||||
}
|
||||
return text("acknowledged " + msgId);
|
||||
}
|
||||
|
||||
@@ -1109,6 +1253,22 @@ public final class FleetMcp {
|
||||
* not the other, and the two then disagree about a live outage. That is exactly what fleetd
|
||||
* #284 was, where one rule computed in two places was widened in only one and a single response
|
||||
* contradicted itself. Shared inputs do not make duplicated computation safe.
|
||||
*
|
||||
* <p>fleetd #395, made live by fleetd #446: also reports {@code exhaustionDetectionArmed}, one
|
||||
* boolean per configured profile — {@code true} when that profile's {@code exhaustedPattern} is
|
||||
* set RIGHT NOW, {@code false} when it is not, so an operator can tell "this profile is
|
||||
* healthy" from "nothing can ever quarantine this profile" without reading {@code fleetd.yaml}.
|
||||
* Unlike {@code quarantined}/{@code coolingOff}, this map always names every profile: an
|
||||
* unarmed profile never enters a transient state to be absent from, so silence here would read
|
||||
* as "healthy" rather than "not being watched at all". Hot since fleetd #446: a reload that
|
||||
* arms or disarms a profile's {@code exhaustedPattern} changes this map's answer on the very
|
||||
* next call, no restart — see {@link dev.ltms.fleet.inject.LiveExhaustedPatterns}'s class doc.
|
||||
*
|
||||
* <p>fleetd #446 criterion 3: each {@code quarantined} row also names {@code model} (the
|
||||
* profile's configured {@code model:}, omitted when the profile names none) and {@code reason}
|
||||
* (the backend text that triggered the most recent quarantine of that credential, omitted when
|
||||
* none is known) — so a lead can see WHICH model to turn off and WHY, without reading the
|
||||
* daemon log.
|
||||
*/
|
||||
public static Map<String, Object> profilesView(PeerLauncher workers, QuarantineSource quarantine, OutageSource outage) {
|
||||
Map<String, Object> result = new LinkedHashMap<>();
|
||||
@@ -1116,13 +1276,23 @@ public final class FleetMcp {
|
||||
result.put("default", workers.defaultProfile() == null ? "" : workers.defaultProfile());
|
||||
Map<String, Object> quarantined = new LinkedHashMap<>();
|
||||
Map<String, Object> coolingOff = new LinkedHashMap<>();
|
||||
Map<String, Object> exhaustionDetectionArmed = new LinkedHashMap<>();
|
||||
for (String profile : workers.profiles()) {
|
||||
exhaustionDetectionArmed.put(profile, quarantine.exhaustedPatternArmed().apply(profile));
|
||||
String credentialId = quarantine.credentialIdFor().apply(profile);
|
||||
if (credentialId != null) {
|
||||
quarantine.quarantine().remainingSeconds(credentialId).ifPresent(remaining -> {
|
||||
Map<String, Object> row = new LinkedHashMap<>();
|
||||
row.put("credentialId", credentialId);
|
||||
row.put("quarantinedForSeconds", remaining);
|
||||
String model = quarantine.modelFor().apply(profile);
|
||||
if (model != null && !model.isBlank()) {
|
||||
row.put("model", model);
|
||||
}
|
||||
String reason = quarantine.reasonFor().apply(credentialId);
|
||||
if (reason != null && !reason.isBlank()) {
|
||||
row.put("reason", reason);
|
||||
}
|
||||
quarantined.put(profile, row);
|
||||
});
|
||||
}
|
||||
@@ -1136,12 +1306,30 @@ public final class FleetMcp {
|
||||
});
|
||||
}
|
||||
}
|
||||
result.put("exhaustionDetectionArmed", exhaustionDetectionArmed);
|
||||
if (!quarantined.isEmpty()) {
|
||||
result.put("quarantined", quarantined);
|
||||
}
|
||||
if (!coolingOff.isEmpty()) {
|
||||
result.put("coolingOff", coolingOff);
|
||||
}
|
||||
// fleetd #422: read the exact same accessor CompositePeerLauncher's spawn gate reads
|
||||
// (PeerLauncher.modelGateState(), which for the composite is models0() read live) — never a
|
||||
// separately-derived answer, so this status can never overstate or understate what the gate
|
||||
// actually enforces (the fleetd #404 lesson).
|
||||
//
|
||||
// fleetd #422 follow-up: "armed" and "off" come from the ONE modelGateState() call below,
|
||||
// never two independent reads of the gate — a reload landing between two separate reads
|
||||
// could otherwise make them disagree. modelGateArmed is reported unconditionally (never
|
||||
// omitted like quarantined/coolingOff above) precisely so a lead can tell "no models: block
|
||||
// at all" (false) apart from "a models: block with nothing currently off" (true, with
|
||||
// modelsOff simply absent below) — the two states PeerLauncher.disabledModels() alone
|
||||
// cannot distinguish, both reporting an empty set.
|
||||
PeerLauncher.ModelGateState modelGate = workers.modelGateState();
|
||||
result.put("modelGateArmed", modelGate.configured());
|
||||
if (!modelGate.off().isEmpty()) {
|
||||
result.put("modelsOff", new ArrayList<>(modelGate.off()));
|
||||
}
|
||||
return result;
|
||||
}
|
||||
|
||||
@@ -1216,12 +1404,48 @@ public final class FleetMcp {
|
||||
LeadSeatSource.none(), leads, selfTerm, coordination);
|
||||
}
|
||||
|
||||
/** As above, plus fleetd #176 lead-seat facts (see {@link LeadSeatSource}). */
|
||||
/**
|
||||
* As above, plus fleetd #176 lead-seat facts (see {@link LeadSeatSource}).
|
||||
*
|
||||
* <p>Assumes the caller is <strong>not</strong> the primary (fleetd #463) — every wrapper
|
||||
* overload above delegates here without carrying a caller identity, which is exactly right for
|
||||
* them: they exist for call sites (and unit tests) that have no {@link Principal} to hand over,
|
||||
* and a missing identity should fail closed rather than fail open onto lead-to-lead state. A
|
||||
* test that wants the {@code coordinator} row must call the canonical overload below with an
|
||||
* explicit {@code true}. The one call site that has a real caller ({@code fleet_list}'s MCP
|
||||
* handler) uses {@link #listFleet(PeerLauncher, SessionManager, MessageService, CapacitySource,
|
||||
* HealthCoverageSource, QuarantineSource, OutageSource, LeadSeatSource, Map, String,
|
||||
* CoordinationSource, boolean)} instead, so it can pass the true answer.
|
||||
*/
|
||||
static McpSchema.CallToolResult listFleet(PeerLauncher workers, SessionManager sessions, MessageService messages,
|
||||
CapacitySource capacity, HealthCoverageSource healthCoverage,
|
||||
QuarantineSource quarantine, OutageSource outage,
|
||||
LeadSeatSource leadSeats, Map<String, String> leads, String selfTerm,
|
||||
CoordinationSource coordination) {
|
||||
return listFleet(workers, sessions, messages, capacity, healthCoverage, quarantine, outage,
|
||||
leadSeats, leads, selfTerm, coordination, false);
|
||||
}
|
||||
|
||||
/**
|
||||
* As above, gated by the caller's role (fleetd #439). The {@code coordinator} row is
|
||||
* lead-to-lead coordination state — coordination between orchestrators, not roster
|
||||
* observation — so it is assembled and included only when {@code callerIsPrimary} is
|
||||
* {@code true}. A worker or an architect gets a result with the {@code coordinator} key
|
||||
* <strong>absent</strong>, never an empty or redacted one, and never pays the cost of
|
||||
* {@link #coordinatorView} probing peer mailboxes for a row it will not receive.
|
||||
*
|
||||
* @param callerIsPrimary whether the {@code fleet_list} caller is the primary; only the MCP
|
||||
* handler computes this from the real connection (see
|
||||
* {@code Principal#isPrimary()}) — every other overload passes
|
||||
* {@code false} (fleetd #463: a forgotten argument fails closed, not
|
||||
* open), so a test that wants the {@code coordinator} row must pass
|
||||
* an explicit {@code true}
|
||||
*/
|
||||
static McpSchema.CallToolResult listFleet(PeerLauncher workers, SessionManager sessions, MessageService messages,
|
||||
CapacitySource capacity, HealthCoverageSource healthCoverage,
|
||||
QuarantineSource quarantine, OutageSource outage,
|
||||
LeadSeatSource leadSeats, Map<String, String> leads, String selfTerm,
|
||||
CoordinationSource coordination, boolean callerIsPrimary) {
|
||||
try {
|
||||
Map<String, Agent> live = workers.list().stream()
|
||||
.map(Agent.class::cast)
|
||||
@@ -1243,9 +1467,14 @@ public final class FleetMcp {
|
||||
Map<String, Object> result = new LinkedHashMap<>();
|
||||
result.put("leads", leadRows); result.put("members", out);
|
||||
result.put("healthCoverage", healthCoverage.value().get());
|
||||
Map<String, Object> coordinatorRow = coordinatorView(coordination);
|
||||
if (coordinatorRow != null) {
|
||||
result.put("coordinator", coordinatorRow);
|
||||
// fleetd #439: coordinator/coordinatorView is lead-to-lead coordination state and must
|
||||
// never reach a worker or an architect -- gate BEFORE assembling it, not after, so the
|
||||
// key is absent rather than present-and-empty.
|
||||
if (callerIsPrimary) {
|
||||
Map<String, Object> coordinatorRow = coordinatorView(coordination);
|
||||
if (coordinatorRow != null) {
|
||||
result.put("coordinator", coordinatorRow);
|
||||
}
|
||||
}
|
||||
if (capacity.available()) result.put("capacity", profiles.stream()
|
||||
.map(profile -> capacityView(profile, capacity.liveCount(), capacity.maxLoad(), roster, messages,
|
||||
@@ -1268,6 +1497,16 @@ public final class FleetMcp {
|
||||
* counts, it can never make {@code fleet_list} itself slow or fail. {@code held} comes from
|
||||
* {@link LeadChannel#peek}, a pure in-memory read with no broker round trip, so it is never
|
||||
* subject to that bound.
|
||||
*
|
||||
* <p><strong>fleetd #421: {@code heldCount}/{@code heldDurable} fix the "pending: 0" trap.</strong>
|
||||
* {@code mailbox.pending} counts only broker-<em>ready</em> messages; a held message is already
|
||||
* an unacked delivery sitting with this consumer, so the normal, healthy state of a blocked lead
|
||||
* is {@code "pending": 0} next to a non-empty {@code held[]} — which invites the false reading
|
||||
* "these are only in memory, a restart will lose them". {@code heldCount} is the honest second
|
||||
* number beside {@code pending} ({@code held.size()}, not left for the reader to count the
|
||||
* array). {@code heldDurable} comes straight from {@link LeadChannel#heldDurable}, which the
|
||||
* channel implementation derives from what it actually did when it declared and consumed its own
|
||||
* queue (fleetd #440) — this method never asserts the fact itself.
|
||||
*/
|
||||
private static Map<String, Object> coordinatorView(CoordinationSource coordination) {
|
||||
LeadChannel channel = coordination.leadChannel();
|
||||
@@ -1275,11 +1514,14 @@ public final class FleetMcp {
|
||||
return null;
|
||||
}
|
||||
String selfId = channel.selfCoordId();
|
||||
List<LeadMessage> held = channel.peek();
|
||||
Map<String, Object> row = new LinkedHashMap<>();
|
||||
row.put("selfId", selfId);
|
||||
row.put("configured", true);
|
||||
row.put("mailbox", mailboxView(probe(channel, selfId)));
|
||||
row.put("held", channel.peek().stream().map(FleetMcp::heldView).toList());
|
||||
row.put("heldCount", held.size());
|
||||
row.put("heldDurable", channel.heldDurable());
|
||||
row.put("held", held.stream().map(FleetMcp::heldView).toList());
|
||||
row.put("peers", coordination.peers().stream().map(p -> peerView(channel, p)).toList());
|
||||
return row;
|
||||
}
|
||||
@@ -1608,10 +1850,18 @@ public final class FleetMcp {
|
||||
"Check an async delegation (a fleet_send with wait:false) by its ticket: "
|
||||
+ "pending, done (with the worker's reply), or failed. When target (a worker "
|
||||
+ "session id) is present instead of ticket, drain that worker's inbox of "
|
||||
+ "replies delivered when no send was open.",
|
||||
+ "replies delivered when no send was open. When coordId is present instead, "
|
||||
+ "read (never consume) your own held lead-to-lead mail — primary-only.",
|
||||
objectSchema(Map.of(
|
||||
"ticket", stringProp("The ticket returned by fleet_send wait:false"),
|
||||
"target", stringProp("Worker session id to drain pending replies from (optional)")),
|
||||
"target", stringProp("Worker session id to drain pending replies from (optional)"),
|
||||
"coordId", stringProp("Your own coord-id (fleet_list's coordinator.selfId) — "
|
||||
+ "reads every message currently held[] for you in full, without "
|
||||
+ "acking. Read twice, get the same bodies both times; fleet_list's "
|
||||
+ "held[] still reports them afterward. Primary-only, and this can "
|
||||
+ "read only YOUR OWN mailbox — there is no route to a peer's outbound "
|
||||
+ "mail, so passing a peer's coordId here is refused rather than "
|
||||
+ "silently returning the wrong thing (or nothing).")),
|
||||
List.of()));
|
||||
}
|
||||
|
||||
@@ -1621,7 +1871,10 @@ public final class FleetMcp {
|
||||
+ "has processed a reply and wants to confirm it, leaving other pending replies "
|
||||
+ "in the inbox for later drain.",
|
||||
objectSchema(Map.of(
|
||||
"target", stringProp("Worker session id whose inbox to ack from"),
|
||||
"target", stringProp("Worker session id whose inbox to ack from. Must name a "
|
||||
+ "reply actually queued for it — an id in no inbox, or a coord-id "
|
||||
+ "(peer held mail, read with fleet_poll{coordId} instead), errors "
|
||||
+ "rather than reporting a false success"),
|
||||
"msgId", stringProp("The message id to acknowledge")),
|
||||
List.of("target", "msgId")));
|
||||
}
|
||||
@@ -1663,7 +1916,16 @@ public final class FleetMcp {
|
||||
"List the configured worker profiles (backends) and which one fleet_spawn uses by "
|
||||
+ "default. A 'quarantined' map is present when a backend-exhausted refusal put "
|
||||
+ "a profile's credential on cooldown — fleet_spawn onto it is refused until "
|
||||
+ "quarantinedForSeconds elapses; a profile sharing that credential is listed too.",
|
||||
+ "quarantinedForSeconds elapses; a profile sharing that credential is listed too, "
|
||||
+ "each row naming 'model' (the model that profile runs, when configured) and "
|
||||
+ "'reason' (the backend text that triggered the quarantine, when known) — the fix "
|
||||
+ "is usually `enabled: false` on that model under models.allow. "
|
||||
+ "'exhaustionDetectionArmed' reports, per profile, whether a usage-limit refusal "
|
||||
+ "on it can EVER be classified and quarantined (its exhaustedPattern is "
|
||||
+ "configured) — false means that profile's credential can never be quarantined "
|
||||
+ "by this mechanism, however many usage-limit refusals it sees. Both "
|
||||
+ "exhaustedPattern and models.allow's on/off state are hot: editing fleetd.yaml "
|
||||
+ "and reloading arms/disarms detection or flips a model off with no restart.",
|
||||
objectSchema(Map.of(), List.of()));
|
||||
}
|
||||
|
||||
|
||||
@@ -14,6 +14,7 @@ import dev.ltms.fleet.placement.BackendOutagePolicy;
|
||||
import dev.ltms.fleet.placement.BackendQuarantine;
|
||||
import dev.ltms.fleet.placement.PlacementCandidate;
|
||||
import dev.ltms.fleet.placement.PlacementContext;
|
||||
import dev.ltms.fleet.placement.PlacementDecision;
|
||||
import dev.ltms.fleet.placement.PlacementException;
|
||||
import dev.ltms.fleet.placement.PlacementPolicies;
|
||||
import dev.ltms.fleet.placement.PlacementPolicy;
|
||||
@@ -90,6 +91,19 @@ public final class CompositePeerLauncher implements PeerLauncher {
|
||||
private final Supplier<Map<String, FleetConfig.Profile>> profileConfigs;
|
||||
private final Supplier<PlacementPolicy> placementPolicy;
|
||||
|
||||
/**
|
||||
* fleetd #422: the central model allow-list's on/off state, read live per spawn — same reason
|
||||
* {@link #profileConfigs} is a supplier rather than a captured map (see the class doc above and
|
||||
* {@link #enforceModelEnabled}). A caller with no {@code models:} block to read from (the
|
||||
* simpler, map-based constructors used throughout this class's own tests) wires this to a
|
||||
* constant {@code null}, which {@link #models0} treats as "nothing configured, gate never
|
||||
* fires" — the pre-#422 behaviour.
|
||||
*/
|
||||
private final Supplier<FleetConfig.Models> models;
|
||||
|
||||
/** The value {@link #models0} normalizes a {@code null} supplier result to. */
|
||||
private static final FleetConfig.Models NO_MODELS_CONFIGURED = new FleetConfig.Models(List.of());
|
||||
|
||||
/** CB-578 stage B: credential cooldown, checked before an explicit spawn and filtered into placement. */
|
||||
private final BackendQuarantine quarantine;
|
||||
|
||||
@@ -205,12 +219,44 @@ public final class CompositePeerLauncher implements PeerLauncher {
|
||||
FleetConfig.Fleet fleet,
|
||||
BackendQuarantine quarantine,
|
||||
BackendOutagePolicy outagePolicy) {
|
||||
// No models: block to read from a plain profiles Map — fleetd #422's gate is wired to a
|
||||
// constant null (see enforceModelEnabled/models0), the pre-#422 behaviour for every caller
|
||||
// of this overload. Use the 9-arg overload below to test the gate against a LIVE supplier.
|
||||
this(delegates, defaultProfile, profileConfigs, placementPolicy, liveCount, fleet, quarantine,
|
||||
outagePolicy, constant(null));
|
||||
}
|
||||
|
||||
/**
|
||||
* As above, plus a LIVE model-gate source (fleetd #422). The 8-arg overload above wires
|
||||
* {@code models} to a constant {@code null} because it has only a static {@code Map<String,
|
||||
* Profile>}, never a full {@code FleetConfig}, to read one from; this overload exists so a test
|
||||
* can prove {@link #enforceModelEnabled} and its candidate filter re-read {@link
|
||||
* FleetConfig.Models} on every call rather than a value captured once at construction — the
|
||||
* exact distinction fleetd #422 exists to get right (see the class doc's CB-559 note on {@link
|
||||
* #profileConfigs}, which this follows). Production wiring uses the
|
||||
* {@code Supplier<FleetConfig>} constructor below instead, which already threads a live
|
||||
* {@code config.get().models()} through.
|
||||
*
|
||||
* @param models required — pass a supplier returning {@code null} for a caller that has no
|
||||
* {@code models:} block to gate against, never a defaulting overload (the same
|
||||
* "explicit opt-out, never a silent default" rule {@code quarantine}/
|
||||
* {@code outagePolicy} already follow).
|
||||
*/
|
||||
public CompositePeerLauncher(List<HerdrPeerLauncher> delegates,
|
||||
String defaultProfile,
|
||||
Map<String, FleetConfig.Profile> profileConfigs,
|
||||
PlacementPolicy placementPolicy,
|
||||
Function<String, Integer> liveCount,
|
||||
FleetConfig.Fleet fleet,
|
||||
BackendQuarantine quarantine,
|
||||
BackendOutagePolicy outagePolicy,
|
||||
Supplier<FleetConfig.Models> models) {
|
||||
// LinkedHashMap, not Map.copyOf: candidates() promises definition order and the weighted
|
||||
// policy breaks exact-weight ties on it, so a salted iteration order would make placement
|
||||
// differ from one JVM run to the next.
|
||||
this(delegates, defaultProfile,
|
||||
constant(Collections.unmodifiableMap(new LinkedHashMap<>(profileConfigs))),
|
||||
constant(placementPolicy), liveCount, constant(fleet), quarantine, outagePolicy);
|
||||
constant(placementPolicy), liveCount, constant(fleet), quarantine, outagePolicy, models);
|
||||
}
|
||||
|
||||
/**
|
||||
@@ -250,7 +296,10 @@ public final class CompositePeerLauncher implements PeerLauncher {
|
||||
liveCount,
|
||||
() -> config.get().fleet(),
|
||||
quarantine,
|
||||
outagePolicy);
|
||||
outagePolicy,
|
||||
// fleetd #422: read live, same as profiles/placement/fleet above — a models.allow
|
||||
// edit (on/off or otherwise) is visible to the very next spawn, no restart needed.
|
||||
() -> config.get().models());
|
||||
}
|
||||
|
||||
/** The all-suppliers form every other constructor funnels into. */
|
||||
@@ -261,7 +310,8 @@ public final class CompositePeerLauncher implements PeerLauncher {
|
||||
Function<String, Integer> liveCount,
|
||||
Supplier<FleetConfig.Fleet> fleet,
|
||||
BackendQuarantine quarantine,
|
||||
BackendOutagePolicy outagePolicy) {
|
||||
BackendOutagePolicy outagePolicy,
|
||||
Supplier<FleetConfig.Models> models) {
|
||||
this.fleet = fleet;
|
||||
this.quarantine = Objects.requireNonNull(quarantine, "quarantine");
|
||||
this.outagePolicy = Objects.requireNonNull(outagePolicy, "outagePolicy");
|
||||
@@ -273,6 +323,7 @@ public final class CompositePeerLauncher implements PeerLauncher {
|
||||
this.profileConfigs = profileConfigs;
|
||||
this.placementPolicy = placementPolicy;
|
||||
this.liveCount = liveCount;
|
||||
this.models = Objects.requireNonNull(models, "models");
|
||||
Map<String, HerdrPeerLauncher> index = new LinkedHashMap<>();
|
||||
for (HerdrPeerLauncher d : this.delegates) {
|
||||
for (String profile : d.profiles()) {
|
||||
@@ -303,6 +354,16 @@ public final class CompositePeerLauncher implements PeerLauncher {
|
||||
return m == null ? Map.of() : m;
|
||||
}
|
||||
|
||||
/**
|
||||
* The currently-configured {@code models:} block, never null (fleetd #422). Read fresh on
|
||||
* every call, the same reason {@link #profiles0} is — a config reload's on/off edit must reach
|
||||
* the very next spawn.
|
||||
*/
|
||||
private FleetConfig.Models models0() {
|
||||
FleetConfig.Models m = models.get();
|
||||
return m == null ? NO_MODELS_CONFIGURED : m;
|
||||
}
|
||||
|
||||
/** The adapter owning {@code profileName} (null/blank → the default). Throws on an unknown profile. */
|
||||
private HerdrPeerLauncher route(String profileName) {
|
||||
String resolved = (profileName == null || profileName.isBlank()) ? defaultProfile : profileName;
|
||||
@@ -332,6 +393,7 @@ public final class CompositePeerLauncher implements PeerLauncher {
|
||||
enforceNotQuarantined(requestedProfile);
|
||||
enforceNotCoolingOff(requestedProfile);
|
||||
enforceMaxLoad(requestedProfile);
|
||||
enforceModelEnabled(requestedProfile);
|
||||
PeerHandle handle = d.spawn(req);
|
||||
spawnedBy.put(handle.id(), d);
|
||||
return handle;
|
||||
@@ -341,18 +403,10 @@ public final class CompositePeerLauncher implements PeerLauncher {
|
||||
// the whole profile list. An EXPLICIT profile (above) is left alone on purpose — it is the
|
||||
// operator overriding, and refusing it would break `fleet_spawn{profile:"opus"}`, which
|
||||
// carries no role and so would be judged against the dev pool it was never meant for.
|
||||
List<PlacementCandidate> candidates = candidates(req.role());
|
||||
String roleDefault = defaultProfileFor(req.role());
|
||||
Set<String> unreachable = new HashSet<>();
|
||||
// CB-578 stage B: computed once up front — a quarantine's expiry cannot pass within one spawn
|
||||
// call, so re-deriving it per retry would only cost work, never change the answer.
|
||||
Set<String> quarantined = quarantinedProfiles(candidates);
|
||||
// fleetd #201 Unit 5: a distinct set from quarantined — see PlacementContext.coolingOff.
|
||||
Set<String> coolingOff = coolingOffProfiles(candidates);
|
||||
PlacementContext ctx = new PlacementContext(roleDefault, candidates, liveCount, unreachable,
|
||||
quarantined, coolingOff);
|
||||
PlacementContext ctx = placementContextFor(req.role(), unreachable);
|
||||
|
||||
int maxAttempts = candidates.isEmpty() ? 1 : candidates.size();
|
||||
int maxAttempts = ctx.candidates().isEmpty() ? 1 : ctx.candidates().size();
|
||||
for (int attempt = 0; attempt < maxAttempts; attempt++) {
|
||||
// Deliberately uncaught: when no candidate is left (all at cap, or all unreachable) the
|
||||
// policy already throws a clear message. Catching it to rethrow a generic
|
||||
@@ -379,8 +433,7 @@ public final class CompositePeerLauncher implements PeerLauncher {
|
||||
chosen.profile(), e.getMessage());
|
||||
unreachable.add(chosen.profile());
|
||||
// Update the context for the next selection so the policy excludes this profile.
|
||||
ctx = new PlacementContext(roleDefault, candidates, liveCount, unreachable,
|
||||
quarantined, coolingOff);
|
||||
ctx = placementContextFor(req.role(), unreachable);
|
||||
}
|
||||
}
|
||||
|
||||
@@ -492,6 +545,49 @@ public final class CompositePeerLauncher implements PeerLauncher {
|
||||
}
|
||||
}
|
||||
|
||||
/**
|
||||
* Refuse an explicit-profile spawn whose {@code model:} the operator has turned off in the
|
||||
* central {@code models.allow:} list (fleetd #422). Deliberately worded apart from {@link
|
||||
* #enforceNotQuarantined} and {@link #enforceNotCoolingOff}: those two report a BACKEND-reported
|
||||
* outage (exhaustion, repeated errors); this one reports an OPERATOR decision, so the message
|
||||
* says "turned off" and names the model, never "quarantined" or "cooling off". A FOURTH,
|
||||
* independent reason to refuse a spawn — never layered onto {@code BackendQuarantine} or {@code
|
||||
* BackendOutagePolicy}, which would misattribute an operator's own choice to the backend.
|
||||
*
|
||||
* <p>Reads {@link #models0()} fresh on every call — the same liveness {@link #profiles0()}
|
||||
* already has — so flipping {@code enabled: false} and reloading takes effect on the very next
|
||||
* spawn, no restart (criterion 4). A profile naming no model, or a model absent from {@code
|
||||
* models.allow:} entirely (nothing to gate against), is never refused here.
|
||||
*
|
||||
* @throws PlacementException naming the model and the profile, distinct from quarantine/cool-off
|
||||
*/
|
||||
private void enforceModelEnabled(String profile) {
|
||||
FleetConfig.Profile cfg = profiles0().get(profile);
|
||||
String model = (cfg == null) ? null : cfg.model();
|
||||
if (model == null) {
|
||||
return;
|
||||
}
|
||||
if (models0().offIds().contains(model)) {
|
||||
throw new PlacementException("worker profile '" + profile + "' names model '" + model
|
||||
+ "', which the operator has turned off in models.allow — refusing spawn");
|
||||
}
|
||||
}
|
||||
|
||||
/** The subset of {@code candidates} whose {@code model:} is currently turned off (fleetd #422). */
|
||||
private Set<String> modelOffProfiles(List<PlacementCandidate> candidates) {
|
||||
Set<String> off = models0().offIds();
|
||||
if (off.isEmpty()) {
|
||||
return Set.of();
|
||||
}
|
||||
return candidates.stream()
|
||||
.map(PlacementCandidate::profile)
|
||||
.filter(p -> {
|
||||
FleetConfig.Profile cfg = profiles0().get(p);
|
||||
return cfg != null && cfg.model() != null && off.contains(cfg.model());
|
||||
})
|
||||
.collect(Collectors.toSet());
|
||||
}
|
||||
|
||||
/**
|
||||
* The profile names {@code role} may be placed on, in definition order.
|
||||
*
|
||||
@@ -508,8 +604,23 @@ public final class CompositePeerLauncher implements PeerLauncher {
|
||||
return known.isEmpty() ? List.copyOf(configured.keySet()) : known;
|
||||
}
|
||||
|
||||
/** The profile an unqualified spawn for {@code role} falls back to under {@code fixed} placement. */
|
||||
private String defaultProfileFor(MemberRole role) {
|
||||
/**
|
||||
* {@inheritDoc}
|
||||
*
|
||||
* <p>Live: reads {@link #poolFor}, which reads {@link #profileConfigs} and {@link #fleet} fresh
|
||||
* on every call, so a config reload is visible without a restart (fleetd #425) — unlike {@link
|
||||
* #defaultProfile}, the field captured once at construction, which this falls back to only when
|
||||
* {@link #poolFor} has nothing to offer at all (no profiles configured for this composite).
|
||||
*
|
||||
* <p>Exact only under the {@code fixed} placement policy — the one that reads this value
|
||||
* ({@code FixedPlacementPolicy}, package-private, hence not linked) as its first, preferred
|
||||
* candidate. {@code weighted}/{@code round-robin} placement can choose a different candidate
|
||||
* from {@code role}'s pool even on the very first spawn; this method does not simulate that
|
||||
* choice, matching what the {@code defaultProfile:}-derived reporting this replaces has always
|
||||
* done.
|
||||
*/
|
||||
@Override
|
||||
public String defaultProfileFor(MemberRole role) {
|
||||
List<String> pool = poolFor(role);
|
||||
return pool.isEmpty() ? defaultProfile : pool.getFirst();
|
||||
}
|
||||
@@ -526,6 +637,142 @@ public final class CompositePeerLauncher implements PeerLauncher {
|
||||
return out;
|
||||
}
|
||||
|
||||
/**
|
||||
* Build the {@link PlacementContext} an unqualified spawn of {@code role} would be judged
|
||||
* against right now — the single source both {@link #spawn} and {@link #place} read, so the two
|
||||
* can never disagree about which conditions (quarantine, cool-off, model-off) apply to which
|
||||
* candidate (fleetd #425 rework: round 1 duplicated this into a second, blind resolver —
|
||||
* {@link #defaultProfileFor} — which is why it regressed; round 2 found that even a single
|
||||
* shared resolver is not enough on its own if the CALLER re-resolves through an explicit
|
||||
* profile afterwards — see {@link PlacementDecision}).
|
||||
*
|
||||
* @param unreachable the caller's mutable unreachable set; {@link #spawn} grows this across
|
||||
* retries and rebuilds the context from it, {@link #place} passes a fresh
|
||||
* empty one since it never retries
|
||||
*/
|
||||
private PlacementContext placementContextFor(MemberRole role, Set<String> unreachable) {
|
||||
List<PlacementCandidate> candidates = candidates(role);
|
||||
String roleDefault = defaultProfileFor(role);
|
||||
// CB-578 stage B: computed once up front — a quarantine's expiry cannot pass within one spawn
|
||||
// call, so re-deriving it per retry would only cost work, never change the answer.
|
||||
Set<String> quarantined = quarantinedProfiles(candidates);
|
||||
// fleetd #201 Unit 5: a distinct set from quarantined — see PlacementContext.coolingOff.
|
||||
Set<String> coolingOff = coolingOffProfiles(candidates);
|
||||
// fleetd #422: read live per spawn, same as quarantined/coolingOff above — a config reload
|
||||
// that flips a model's enabled state is visible to the very next unqualified spawn.
|
||||
Set<String> modelOff = modelOffProfiles(candidates);
|
||||
return new PlacementContext(roleDefault, candidates, liveCount, unreachable,
|
||||
quarantined, coolingOff, modelOff);
|
||||
}
|
||||
|
||||
/**
|
||||
* {@inheritDoc}
|
||||
*
|
||||
* <p>fleetd #425 rework, round 2: runs the exact same selection {@link #spawn} uses for a
|
||||
* blank-profile request — {@link #placementContextFor} plus one {@link PlacementPolicy#select}
|
||||
* — rather than {@link #defaultProfileFor}'s blind "pool's first entry", so a quarantined,
|
||||
* cooling-off, or model-off pool-first candidate is routed around here exactly as it would be
|
||||
* by a real spawn. Unlike {@link #spawn}, this never retries on {@link
|
||||
* PeerUnreachableException}: there is no spawn attempt to fail, so "unreachable" never grows
|
||||
* past the empty set it starts with, and a single {@link PlacementPolicy#select} call already
|
||||
* reflects the live quarantine/cool-off/model-off state.
|
||||
*
|
||||
* <p>Deliberately does <em>not</em> apply {@link #enforceMaxLoad} (or any of the other three
|
||||
* {@code enforce*} checks): those belong to {@link #spawn}'s EXPLICIT-profile branch, the
|
||||
* operator-override path, and this method answers a different question — "where would an
|
||||
* UNQUALIFIED spawn land". That is not the same as {@code select} ignoring these conditions —
|
||||
* every condition {@code select} filters on (quarantine, cooling off, {@code maxLoad} under
|
||||
* every placement policy including the default {@code fixed}, since fleetd #435, model-off,
|
||||
* unreachable, weight-0) is already reflected in the {@link PlacementDecision} this method
|
||||
* returns, because {@code select} walked past every excluded candidate to find it. What this
|
||||
* method's caller must not do is take that resolved name and hand it back to {@link
|
||||
* #spawn(SpawnRequest)} as an explicit profile: the explicit-profile branch treats the same
|
||||
* exclusion conditions as a reason to REFUSE, where {@code select} had already treated them as
|
||||
* a reason to fall through — round 1 of this fix did exactly that, turning a fall-through this
|
||||
* method had already resolved around into a refusal one call later. Round 2 fixes that at the
|
||||
* caller: {@link #spawn(SpawnRequest, PlacementDecision)} carries this exact decision to the
|
||||
* spawn without re-resolving or re-checking it, through the same routing path {@code select}
|
||||
* itself was consulted from.
|
||||
*
|
||||
* @throws PlacementException if no candidate in {@code role}'s pool is currently placeable
|
||||
* (mirrors what an actual unqualified spawn would throw)
|
||||
*/
|
||||
@Override
|
||||
public PlacementDecision place(MemberRole role) {
|
||||
PlacementContext ctx = placementContextFor(role, new HashSet<>());
|
||||
return new PlacementDecision(placementPolicy.get().select(ctx).profile());
|
||||
}
|
||||
|
||||
/**
|
||||
* {@inheritDoc}
|
||||
*
|
||||
* <p>Delegates to {@link #place}, so the two can never disagree about the answer for the same
|
||||
* {@code role} at the same instant — kept as a convenience for a caller that only wants the
|
||||
* resolved name (a status report, a log line), never for a caller that will act on it by
|
||||
* spawning: that caller must hold the {@link PlacementDecision} itself and pass it to {@link
|
||||
* #spawn(SpawnRequest, PlacementDecision)} — see {@link PlacementDecision}'s javadoc for why
|
||||
* resolving here and spawning separately, with the name fed back in as an explicit profile,
|
||||
* regressed fleetd #425 twice.
|
||||
*/
|
||||
@Override
|
||||
public String routedProfileFor(MemberRole role) {
|
||||
return place(role).profile();
|
||||
}
|
||||
|
||||
/**
|
||||
* {@inheritDoc}
|
||||
*
|
||||
* <p>Routes {@code decision.profile()} directly to its owning delegate — the identical
|
||||
* {@code d.spawn(routedReq)} call {@link #spawn(SpawnRequest)}'s blank-profile branch makes for
|
||||
* its first pick — WITHOUT re-running {@link #enforceNotQuarantined}, {@link
|
||||
* #enforceNotCoolingOff}, {@link #enforceMaxLoad}, or {@link #enforceModelEnabled}: those are
|
||||
* the EXPLICIT-profile branch's checks, and {@code decision} did not come from an operator
|
||||
* naming a profile — it came from {@link #place}, which already applied whichever of these
|
||||
* conditions {@link PlacementPolicy#select} actually filters on (fleetd #425 rework, round 2).
|
||||
*
|
||||
* <p>The two branches disagree on purpose about what an excluded profile means, and that
|
||||
* disagreement is not what this method removes. The blank-profile routing branch (and
|
||||
* {@link #place}) treats a quarantined/cooling-off/at-cap/model-off/unreachable/weight-0 profile
|
||||
* as a reason to fall through to the next candidate; the EXPLICIT-profile branch treats naming
|
||||
* that same profile as a reason to refuse outright — someone who names a profile should get a
|
||||
* refusal, not a silent substitution onto a different backend. That is still correct after
|
||||
* fleetd #435. What round 1 got wrong, and what this method exists to stop happening again, is
|
||||
* turning a fall-through into a refusal by accident: resolving a name via {@link #place} and
|
||||
* then handing that same name back to {@link #spawn(SpawnRequest)} as an explicit profile takes
|
||||
* the refusing branch on a decision the routing branch had already approved by falling through
|
||||
* past everything else.
|
||||
*
|
||||
* <p>Before fleetd #435, this exact accident was reachable through {@code maxLoad} specifically:
|
||||
* {@code FixedPlacementPolicy} — the default policy — did not evaluate {@code maxLoad} at all
|
||||
* for automatic selection, so {@link #place} could approve an at-cap profile that {@link
|
||||
* #enforceMaxLoad} would then refuse one call later. fleetd #435 closed that: {@code
|
||||
* FixedPlacementPolicy} now walks past an at-cap candidate exactly like {@code weighted}/
|
||||
* {@code round-robin} already did, so {@link #place} can no longer return one, and this specific
|
||||
* failure — an approved placement dying at {@code enforceMaxLoad} — cannot happen any more.
|
||||
* What this method still buys, now that {@code maxLoad} can no longer cause it: it never
|
||||
* re-evaluates a condition {@link #place} already decided, and it closes the window between
|
||||
* that decision and the spawn in which the underlying state (another spawn landing on the same
|
||||
* profile, a config reload) could otherwise move and make a stale explicit re-check wrong.
|
||||
*
|
||||
* <p>Deliberately does not retry on {@link PeerUnreachableException} across candidates the way
|
||||
* {@link #spawn(SpawnRequest)}'s blank-profile branch does: retrying here would silently
|
||||
* re-place the caller onto a different profile than the one {@code decision} named, behind the
|
||||
* back of a caller that may already have provisioned something (a worktree's {@code repoRoot},
|
||||
* parity overlay) specifically for that name. A caller that wants the composite's own failover
|
||||
* should call {@link #spawn(SpawnRequest)} with a blank profile directly, not resolve through
|
||||
* {@link #place} first. Losing that retry on a resolve-then-spawn path is an accepted, unrelated
|
||||
* cost — see {@code SessionManager.acquireWithWorktree}'s own comment on it — never widened by
|
||||
* this round to include {@code maxLoad}, which is what round 1 actually lost.
|
||||
*/
|
||||
@Override
|
||||
public PeerHandle spawn(SpawnRequest req, PlacementDecision decision) {
|
||||
HerdrPeerLauncher d = route(decision.profile());
|
||||
SpawnRequest routedReq = req.withProfile(decision.profile());
|
||||
PeerHandle handle = d.spawn(routedReq);
|
||||
spawnedBy.put(handle.id(), d);
|
||||
return handle;
|
||||
}
|
||||
|
||||
@Override
|
||||
public String effectiveCwd(SpawnRequest req) {
|
||||
return route(req.profileName()).effectiveCwd(req);
|
||||
@@ -536,6 +783,32 @@ public final class CompositePeerLauncher implements PeerLauncher {
|
||||
return route(profileName).parityOverlay(profileName);
|
||||
}
|
||||
|
||||
/**
|
||||
* fleetd #422 follow-up: the single live read that answers both "is the models.allow: gate
|
||||
* armed" and "which models are off", off the exact same accessor ({@link #models0()}) {@link
|
||||
* #enforceModelEnabled} and {@link #modelOffProfiles} read — so {@code fleet_profiles}/{@code
|
||||
* GET /profiles} (via {@link PeerLauncher#disabledModels()}, which now delegates here) can
|
||||
* never report a different answer than the gate enforces (the fleetd #404 lesson), and the
|
||||
* startup log line built from this can never disagree with either.
|
||||
*
|
||||
* <p>{@link #models0()} itself normalizes a {@code null} {@link #models} read to the shared
|
||||
* {@link #NO_MODELS_CONFIGURED} sentinel — deliberately the one object no config-supplied
|
||||
* {@code Models} instance can ever be identical to, since it is private to this class — so
|
||||
* comparing by reference here recovers exactly the fact {@code models0()}'s normalization
|
||||
* would otherwise erase: whether the live source was {@code null} (no {@code models:} block,
|
||||
* armed = false) or a real, config-supplied block (armed = true, even one whose {@code allow:}
|
||||
* is itself empty or absent — {@link FleetConfig.Models}'s "absent or empty allow: is off"
|
||||
* wording governs config-load validation, a distinct question from whether this gate is armed
|
||||
* for reporting).
|
||||
*/
|
||||
@Override
|
||||
public PeerLauncher.ModelGateState modelGateState() {
|
||||
FleetConfig.Models m = models0();
|
||||
return m == NO_MODELS_CONFIGURED
|
||||
? PeerLauncher.ModelGateState.notConfigured()
|
||||
: PeerLauncher.ModelGateState.armed(m.offIds());
|
||||
}
|
||||
|
||||
@Override
|
||||
public void stop(String id) {
|
||||
HerdrPeerLauncher d = spawnedBy.get(id);
|
||||
@@ -647,9 +920,21 @@ public final class CompositePeerLauncher implements PeerLauncher {
|
||||
return byProfile.keySet();
|
||||
}
|
||||
|
||||
/**
|
||||
* {@inheritDoc}
|
||||
*
|
||||
* <p>fleetd #425: reports the <em>live</em> {@code dev} pool's first entry — the same value
|
||||
* {@link #defaultProfileFor} computes for {@link MemberRole#DEV} — not the {@link
|
||||
* #defaultProfile} field captured at construction. An unqualified {@code fleet_spawn} defaults
|
||||
* to {@code MemberRole#DEV} (see {@link dev.ltms.fleet.peer.SpawnRequest}), so "the dev pool's
|
||||
* live first entry" is exactly the profile such a spawn actually lands on right now — the
|
||||
* question {@code fleet_profiles}' {@code "default"} field exists to answer. The frozen field is
|
||||
* a role-agnostic fallback used only when {@link #poolFor} has nothing to report at all (no
|
||||
* profiles configured), which {@link #defaultProfileFor} already handles.
|
||||
*/
|
||||
@Override
|
||||
public String defaultProfile() {
|
||||
return defaultProfile;
|
||||
return defaultProfileFor(MemberRole.DEV);
|
||||
}
|
||||
|
||||
/**
|
||||
|
||||
@@ -75,9 +75,39 @@ import java.util.stream.Stream;
|
||||
* never to the control.
|
||||
*
|
||||
* <p>The scrub also writes {@code scrub-report.txt} into its own directory: one {@code allowed N of
|
||||
* M} line (N = exports left untouched, M = exports present when the scrub ran), then the blanked
|
||||
* NAMES — never values. The launcher reads this back at teardown and logs it, because a blocked
|
||||
* count next to an unknown denominator is not a finding.
|
||||
* M failed F} line (N = exports left untouched, M = exports present when the scrub ran, F = names
|
||||
* the scrub attempted to blank but could not), then the NAMES — blanked ones bare, unblankable ones
|
||||
* {@code !}-prefixed — never values. The launcher reads this back at teardown and logs it, because a
|
||||
* blocked count next to an unknown denominator is not a finding.
|
||||
*
|
||||
* <p><b>fleetd #394:</b> plain {@code export "$n="} is a FATAL error for a zsh read-only or special
|
||||
* parameter (for example {@code UID}) — it aborts the whole sourced file, so every name still to
|
||||
* come is never blanked and the report above is never written at all. The blanking loop instead
|
||||
* routes each attempt through {@code eval}, which contains that error to the single iteration: the
|
||||
* loop always finishes, and a name that could not be blanked is counted as {@code failed} and
|
||||
* listed {@code !}-prefixed rather than silently disappearing. This is deliberately not a skip-list
|
||||
* of known-bad names — every enumerated name is still attempted, so a name nobody has thought of
|
||||
* yet still gets tried and, if it fails, still gets counted.
|
||||
*
|
||||
* <p>The blanking loop also re-asserts, on its own, the same {@code [A-Za-z_][A-Za-z0-9_]*} shape
|
||||
* check the enumeration loop already applied. Before {@code eval} was introduced a non-conforming
|
||||
* name reaching {@code export "$n="} was harmless either way — the quoting made it inert. With
|
||||
* {@code eval}, the name is spliced into a string and interpreted as shell syntax, so the enumeration
|
||||
* loop's check is no longer sufficient on its own to keep that call site safe — it is a guard on a
|
||||
* different loop, and the two must not silently drift apart. Re-checking right before the
|
||||
* {@code eval} keeps that call site safe by its own reading, independent of whatever the enumeration
|
||||
* loop does or stops doing in a later change.
|
||||
*
|
||||
* <p><b>fleetd #400:</b> {@code eval}'s exit status is not proof that the blank actually happened.
|
||||
* zsh coerces a bare {@code NAME=} assignment on an integer special parameter (measured on macOS zsh
|
||||
* 5.9: {@code SECONDS}, {@code RANDOM}, {@code SHLVL}, {@code HISTSIZE}, {@code COLUMNS},
|
||||
* {@code LINES}, {@code USERNAME}) to a number instead of failing — {@code eval} returns success,
|
||||
* the value is untouched, and a status-based classification reports it as blanked when it was not.
|
||||
* The fix classifies on the observed effect instead: after the attempt, the name's value is read
|
||||
* back with the {@code (P)} indirection flag and the decision is made from whether that is now
|
||||
* empty. This one check covers all three shapes a name can take at this point — a genuine blank, a
|
||||
* fatal read-only error {@code eval} merely contained, and this silent no-op — and the exit status
|
||||
* plays no part in the decision at all.
|
||||
*/
|
||||
public final class EnvAllowListScrub {
|
||||
|
||||
@@ -132,10 +162,16 @@ public final class EnvAllowListScrub {
|
||||
}
|
||||
|
||||
/**
|
||||
* A parsed {@code scrub-report.txt}: how many exported variables existed when the scrub ran,
|
||||
* how many were left untouched (allowed), and the NAMES that were blanked. Values never appear.
|
||||
* A parsed {@code scrub-report.txt}: how many exported variables existed when the scrub ran
|
||||
* ({@code total}), how many were left untouched ({@code allowed}), how many the scrub attempted
|
||||
* to blank but could not ({@code failed} — fleetd #394: a zsh read-only/special parameter such
|
||||
* as {@code UID} fatally errors on plain {@code export NAME=}, so those attempts go through
|
||||
* {@code eval} instead so the loop keeps going and the failure is counted rather than left
|
||||
* invisible), and the NAMES in each of the latter two categories. {@code allowed +
|
||||
* blanked.size() + unblankable.size() == total}, and {@code unblankable.size() == failed}.
|
||||
* Values never appear.
|
||||
*/
|
||||
record ScrubReport(int allowed, int total, List<String> blanked) {
|
||||
record ScrubReport(int allowed, int total, int failed, List<String> blanked, List<String> unblankable) {
|
||||
}
|
||||
|
||||
/**
|
||||
@@ -334,38 +370,57 @@ public final class EnvAllowListScrub {
|
||||
_cb633_blank+=("$_cb633_n")
|
||||
done
|
||||
|
||||
# `export UID=` is not a failed command: it is a FATAL zsh parameter error
|
||||
# ("failed to change user ID") that aborts this whole sourced file mid-loop,
|
||||
# leaving every later name unscrubbed and the report below unwritten — silently,
|
||||
# because of the 2>/dev/null. Neither `|| true` nor a `${(t)n}` type guard
|
||||
# contains it; only `eval` does. `eval` is safe here precisely because the loop
|
||||
# above already rejected every name that is not [A-Za-z_][A-Za-z0-9_]*, so
|
||||
# nothing but a bare identifier can reach it.
|
||||
# fleetd #394: plain `export "$n="` is FATAL for a zsh read-only/special parameter
|
||||
# (e.g. UID) and aborts this whole sourced file — every name still to come is never
|
||||
# blanked, and the report below is never written, silently. `eval` contains that
|
||||
# error to the single iteration instead: it still fails for that one name, but the
|
||||
# loop continues and we can tell allowed / blanked / unblankable apart afterwards.
|
||||
# This is not a skip-list of known-bad names (that would miss the next one nobody
|
||||
# thought of) — every name in _cb633_blank is still attempted, unconditionally.
|
||||
# Every name reaching this loop already passed the identical identifier check in the
|
||||
# enumeration loop above — but that guard is 20 lines away in a different loop, and
|
||||
# this line is about to splice the name into a string handed to `eval`. Before this
|
||||
# fix the name only ever reached `export` quoted ("$n="), which is inert on a
|
||||
# non-identifier string either way; `eval` makes THIS line the only thing standing
|
||||
# between such a string and code execution in the member's pane, so it re-asserts the
|
||||
# same check on its own rather than trusting a guard it does not own. Under normal
|
||||
# operation this can never fire (the enumeration guard already filtered everything
|
||||
# reaching _cb633_blank), so a name caught here is counted as unblankable rather than
|
||||
# silently dropped — it is real evidence that the upstream guard was bypassed.
|
||||
#
|
||||
# Enumerating the special names instead (UID|EUID|GID|EGID|PPID|LINENO) also
|
||||
# works, but only for the ones enumerated: a special that turns up exported on
|
||||
# some other host brings the abort straight back. `eval` contains all of them.
|
||||
#
|
||||
# Then VERIFY. A contained failure is still a failure, so a name that did not
|
||||
# actually blank must not be reported as blanked. It currently falls into the
|
||||
# "allowed" count, which is imprecise in the safe direction; the honest third
|
||||
# count ("tried and could not blank") needs a report-format change and belongs
|
||||
# with fleetd #394, not here.
|
||||
typeset -a _cb633_done
|
||||
_cb633_done=()
|
||||
# fleetd #400: the attempt's own exit status is NOT proof of its effect. zsh coerces
|
||||
# a bare `NAME=` assignment on an integer special parameter (SECONDS, RANDOM, SHLVL,
|
||||
# HISTSIZE, COLUMNS, LINES, USERNAME on this host) to a number instead of failing —
|
||||
# `eval` returns 0, the value is untouched, and the old exit-status check reported it
|
||||
# as blanked when it was not. Classify on the observed effect instead: attempt the
|
||||
# export, then read the name's value back with the `(P)` indirection flag and decide
|
||||
# from whether it is now empty. One check then covers all three shapes a name can
|
||||
# take here — a genuine blank, a fatal read-only error `eval` merely contained, and
|
||||
# this silent no-op — without the exit status entering the decision at all.
|
||||
typeset -a _cb633_ok _cb633_unblankable
|
||||
_cb633_ok=()
|
||||
_cb633_unblankable=()
|
||||
for _cb633_n in "${_cb633_blank[@]}"; do
|
||||
if [[ ! "$_cb633_n" =~ ^[A-Za-z_][A-Za-z0-9_]*$ ]]; then
|
||||
_cb633_unblankable+=("$_cb633_n")
|
||||
continue
|
||||
fi
|
||||
eval "export ${_cb633_n}=" 2>/dev/null
|
||||
[[ -z "${(P)_cb633_n}" ]] && _cb633_done+=("$_cb633_n")
|
||||
if [[ -z "${(P)_cb633_n}" ]]; then
|
||||
_cb633_ok+=("$_cb633_n")
|
||||
else
|
||||
_cb633_unblankable+=("$_cb633_n")
|
||||
fi
|
||||
done
|
||||
_cb633_blank=("${_cb633_done[@]}")
|
||||
|
||||
integer _cb633_kept=$(( _cb633_total - ${#_cb633_blank} ))
|
||||
{
|
||||
print -r -- "allowed $_cb633_kept of $_cb633_total"
|
||||
for _cb633_n in "${_cb633_blank[@]}"; do print -r -- "$_cb633_n"; done
|
||||
print -r -- "allowed $_cb633_kept of $_cb633_total failed ${#_cb633_unblankable}"
|
||||
for _cb633_n in "${_cb633_ok[@]}"; do print -r -- "$_cb633_n"; done
|
||||
for _cb633_n in "${_cb633_unblankable[@]}"; do print -r -- "!$_cb633_n"; done
|
||||
} > "$ZDOTDIR/%s" 2>/dev/null
|
||||
|
||||
unset _cb633_done _cb633_allowed _cb633_names _cb633_blank _cb633_n _cb633_total _cb633_kept
|
||||
unset _cb633_allowed _cb633_names _cb633_blank _cb633_ok _cb633_unblankable _cb633_n _cb633_total _cb633_kept
|
||||
""".formatted(names, MemberEnvAllowList.zshCasePattern(), REPORT_FILE);
|
||||
}
|
||||
|
||||
@@ -379,6 +434,11 @@ public final class EnvAllowListScrub {
|
||||
* Read and parse {@link #REPORT_FILE} out of a generated ZDOTDIR directory. Returns {@code null}
|
||||
* when absent or unreadable (the pane may have been torn down before its login shell ever got to
|
||||
* the scrub) — callers treat that as "no measurement available", never as success.
|
||||
*
|
||||
* <p>First line is {@code "allowed <N> of <M> failed <F>"} (fleetd #394 added the trailing
|
||||
* {@code failed <F>} — a count of names the scrub attempted to blank but could not, e.g. a zsh
|
||||
* read-only/special parameter). Every following non-blank line is a name: a bare name was
|
||||
* blanked, a {@code !}-prefixed name was attempted and failed. Values never appear on either.
|
||||
*/
|
||||
static ScrubReport readReport(Path zdotdir) {
|
||||
Path report = zdotdir.resolve(REPORT_FILE);
|
||||
@@ -391,17 +451,24 @@ public final class EnvAllowListScrub {
|
||||
return null;
|
||||
}
|
||||
String[] parts = lines.getFirst().substring("allowed ".length()).trim().split("\\s+");
|
||||
if (parts.length != 3 || !"of".equals(parts[1])) {
|
||||
if (parts.length != 5 || !"of".equals(parts[1]) || !"failed".equals(parts[3])) {
|
||||
return null;
|
||||
}
|
||||
List<String> blanked = new ArrayList<>();
|
||||
List<String> unblankable = new ArrayList<>();
|
||||
for (int i = 1; i < lines.size(); i++) {
|
||||
if (!lines.get(i).isBlank()) {
|
||||
blanked.add(lines.get(i));
|
||||
String line = lines.get(i);
|
||||
if (line.isBlank()) {
|
||||
continue;
|
||||
}
|
||||
if (line.startsWith("!")) {
|
||||
unblankable.add(line.substring(1));
|
||||
} else {
|
||||
blanked.add(line);
|
||||
}
|
||||
}
|
||||
return new ScrubReport(Integer.parseInt(parts[0]), Integer.parseInt(parts[2]),
|
||||
List.copyOf(blanked));
|
||||
Integer.parseInt(parts[4]), List.copyOf(blanked), List.copyOf(unblankable));
|
||||
} catch (IOException | NumberFormatException e) {
|
||||
return null;
|
||||
}
|
||||
|
||||
@@ -17,6 +17,7 @@ import dev.ltms.fleet.peer.PeerHandle;
|
||||
import dev.ltms.fleet.peer.PeerLauncher;
|
||||
import dev.ltms.fleet.peer.PeerUnreachableException;
|
||||
import dev.ltms.fleet.peer.SpawnRequest;
|
||||
import dev.ltms.fleet.placement.PlacementDecision;
|
||||
import org.slf4j.Logger;
|
||||
import org.slf4j.LoggerFactory;
|
||||
|
||||
@@ -594,6 +595,23 @@ public abstract class HerdrPeerLauncher implements PeerLauncher {
|
||||
req.sessionName(), spawned.agentSessionId(), spawned.receipt());
|
||||
}
|
||||
|
||||
/**
|
||||
* {@inheritDoc}
|
||||
*
|
||||
* <p>fleetd #450: re-enters {@link #spawn(SpawnRequest)} with {@code decision}'s profile named
|
||||
* explicitly. This is the re-entering form the interface javadoc describes for a launcher with
|
||||
* no placement concept of its own — an instance of this class spawns a single adapter's own
|
||||
* profile set by explicit name only ({@link #place}/{@link #defaultProfileFor} are unoverridden
|
||||
* here and just wrap {@link #defaultProfile()}); it does no quarantine/cool-off/maxLoad/model-off
|
||||
* filtering of its own to re-apply. That filtering lives one layer up, in {@code
|
||||
* CompositePeerLauncher}, which is the launcher that routes across more than one profile and
|
||||
* therefore overrides this method with the routing form instead.
|
||||
*/
|
||||
@Override
|
||||
public PeerHandle spawn(SpawnRequest req, PlacementDecision decision) {
|
||||
return spawn(req.withProfile(decision.profile()));
|
||||
}
|
||||
|
||||
/** The herdr daemon that owns this launcher's pane coordinates. */
|
||||
public HerdrClient herdr() {
|
||||
return agents.herdr();
|
||||
@@ -1643,11 +1661,19 @@ public abstract class HerdrPeerLauncher implements PeerLauncher {
|
||||
log.warn("memberCredentials allow-list: pane {} left no scrub report in {} — the "
|
||||
+ "environment scrub cannot be confirmed to have run. Either the pane ended "
|
||||
+ "before its shell finished starting, or its shell never read our generated "
|
||||
+ "startup files, in which case that member saw the full host environment.",
|
||||
+ "startup files. Either way, we cannot tell from here whether the scrub ran, "
|
||||
+ "so we do not know what that member's environment contained.",
|
||||
paneId, dir);
|
||||
} else {
|
||||
log.info("memberCredentials allow-list: pane {} allowed {} of {} environment variables",
|
||||
paneId, report.allowed(), report.total());
|
||||
if (report.failed() > 0) {
|
||||
log.warn("memberCredentials allow-list: pane {} could not blank {} environment "
|
||||
+ "variable(s) — {} (likely a zsh read-only/special parameter) — those "
|
||||
+ "names were left in the member's environment. Confirm none of them is a "
|
||||
+ "credential.",
|
||||
paneId, report.failed(), report.unblankable());
|
||||
}
|
||||
List<String> shaped = report.blanked().stream()
|
||||
.filter(name -> CREDENTIAL_SHAPED_NAME.matcher(name).matches())
|
||||
.toList();
|
||||
|
||||
@@ -331,11 +331,22 @@ public final class OpenCodeLauncher extends HerdrPeerLauncher {
|
||||
+ "form — opencode's per-model context limit could not be applied for this profile",
|
||||
cfg.profile(), cfg.model());
|
||||
}
|
||||
// fleetd #393: memberSkills seeding (GitWorktrees#seedSkills) copies skill folders into
|
||||
// EVERY provisioned worktree's .claude/skills/ regardless of which kind ultimately spawns
|
||||
// into it — that copy step cannot know the kind, only the caller of GitWorktrees#add does
|
||||
// (see that method's own javadoc). .claude/skills/ is a Claude Code CLI convention the CLI
|
||||
// discovers on its own; opencode has no such discovery, so without this, a seeded skill
|
||||
// never reaches an opencode member even though GitWorktrees logged it as seeded. Read
|
||||
// whatever landed under <cwd>/.claude/skills/ here — the one place in this launcher that
|
||||
// knows both the kind (opencode, by construction: this IS OpenCodeLauncher) and the cwd.
|
||||
List<Path> skillInstructionFiles = skillInstructionFiles(spec.cwd());
|
||||
// A config file is needed for the bridge MCP mount, a member charter, the IDE MCP (+ its
|
||||
// guidance overlay, CB-634), a pinned endpoint (CB-508), or a resolvable autoCompactWindow.
|
||||
// guidance overlay, CB-634), a pinned endpoint (CB-508), a resolvable autoCompactWindow, or
|
||||
// at least one seeded skill to deliver via instructions[] (fleetd #393).
|
||||
if (cfg.hasMcp() || cfg.hasIdeMcp() || spec.charter() != null || hasCustomProvider(cfg)
|
||||
|| wantsContextLimit) {
|
||||
workerEnv.put("OPENCODE_CONFIG", writeConfig(cfg, spec.charter(), spec.cwd()).toString());
|
||||
|| wantsContextLimit || !skillInstructionFiles.isEmpty()) {
|
||||
workerEnv.put("OPENCODE_CONFIG",
|
||||
writeConfig(cfg, spec.charter(), spec.cwd(), skillInstructionFiles).toString());
|
||||
}
|
||||
applyGitToken(workerEnv, cfg);
|
||||
List<String> argv = argvWithResume(argvWithModel(argvWithAuto(cfg), cfg), spec.resumeSessionId());
|
||||
@@ -358,6 +369,65 @@ public final class OpenCodeLauncher extends HerdrPeerLauncher {
|
||||
return withAgent;
|
||||
}
|
||||
|
||||
/**
|
||||
* fleetd #393: the {@code SKILL.md} paths under {@code <cwd>/.claude/skills/} this launcher can
|
||||
* turn into {@code instructions[]} entries, plus the honest log this ticket asks for — emitted
|
||||
* here, at the one point a skill's fate for THIS spawn is actually known, rather than trusting
|
||||
* {@code GitWorktrees#seedSkills}'s kind-blind "N of M" line to mean "and it will be read."
|
||||
*
|
||||
* <p>Every non-hidden subdirectory of {@code .claude/skills/} is a candidate, whether it got
|
||||
* there via {@code memberSkills:} seeding or because the target repo ships its own copy — this
|
||||
* launcher does not care which; it only cares what it can find at spawn time. A candidate with
|
||||
* a {@code SKILL.md} at its top level (the same shape {@link #writeConfig} already requires for
|
||||
* the charter and IDE-rules instructions entries) is delivered; anything else is a directory
|
||||
* this launcher cannot turn into a flat instructions entry, named explicitly in the log rather
|
||||
* than silently dropped, so a caller sees a real "cannot consume" reason and not just a smaller
|
||||
* number than {@code GitWorktrees}' own count.
|
||||
*
|
||||
* <p>No candidates at all (directory absent or empty) logs nothing — the same
|
||||
* no-log-when-nothing-to-say shape {@link #hasCustomProvider} and friends already follow, and
|
||||
* the shape {@code GitWorktrees#seedSkills} itself uses when {@code memberSkills:} is unset.
|
||||
* A failure to even list the directory is logged and treated as "nothing delivered" — best
|
||||
* effort, must never fail the spawn, matching {@code GitWorktrees#seedSkills}'s own contract.
|
||||
*/
|
||||
private List<Path> skillInstructionFiles(String cwd) {
|
||||
if (cwd == null || cwd.isBlank()) {
|
||||
return List.of();
|
||||
}
|
||||
Path skillsDir = Path.of(cwd, ".claude", "skills");
|
||||
if (!Files.isDirectory(skillsDir)) {
|
||||
return List.of();
|
||||
}
|
||||
List<Path> candidates;
|
||||
try (var listing = Files.list(skillsDir)) {
|
||||
candidates = listing.filter(Files::isDirectory)
|
||||
.filter(p -> !p.getFileName().toString().startsWith("."))
|
||||
.sorted()
|
||||
.toList();
|
||||
} catch (IOException e) {
|
||||
log.warn("could not scan {} for skill folders to deliver to this opencode member: {}",
|
||||
skillsDir, e.getMessage());
|
||||
return List.of();
|
||||
}
|
||||
if (candidates.isEmpty()) {
|
||||
return List.of();
|
||||
}
|
||||
List<Path> delivered = candidates.stream()
|
||||
.map(dir -> dir.resolve("SKILL.md"))
|
||||
.filter(Files::isRegularFile)
|
||||
.toList();
|
||||
List<String> undeliverable = candidates.stream()
|
||||
.filter(dir -> !Files.isRegularFile(dir.resolve("SKILL.md")))
|
||||
.map(dir -> dir.getFileName().toString())
|
||||
.toList();
|
||||
log.info("skill delivery: {} of {} skill folder(s) under {} reached this opencode member via "
|
||||
+ "instructions[] (opencode does not read .claude/skills/ natively, unlike "
|
||||
+ "Claude Code){}",
|
||||
delivered.size(), candidates.size(), skillsDir,
|
||||
undeliverable.isEmpty() ? "" : "; no SKILL.md, could not be delivered: " + undeliverable);
|
||||
return delivered;
|
||||
}
|
||||
|
||||
/**
|
||||
* True when this profile pins its own OpenAI-compatible endpoint (CB-508) rather than using
|
||||
* whatever provider opencode resolves by default.
|
||||
@@ -418,8 +488,15 @@ public final class OpenCodeLauncher extends HerdrPeerLauncher {
|
||||
* fresh per-spawn directory under {@link #configRoot}, and return the config file's path for
|
||||
* {@code OPENCODE_CONFIG}. The dir is unique per spawn so concurrent workers never race on it;
|
||||
* it is best-effort cleaned on JVM exit (worker config is disposable — regenerated every spawn).
|
||||
*
|
||||
* @param skillInstructionFiles fleetd #393: absolute {@code SKILL.md} paths from
|
||||
* {@link #skillInstructionFiles(String)}, appended to
|
||||
* {@code instructions[]} so a {@code memberSkills:}-seeded skill
|
||||
* reaches this opencode member the same way the charter and IDE
|
||||
* rules already do.
|
||||
*/
|
||||
private Path writeConfig(FleetConfig.Profile cfg, String charterText, String cwd) {
|
||||
private Path writeConfig(FleetConfig.Profile cfg, String charterText, String cwd,
|
||||
List<Path> skillInstructionFiles) {
|
||||
try {
|
||||
Path dir = Files.createTempDirectory(configParentDir(), "fleetd-opencode-");
|
||||
dir.toFile().deleteOnExit();
|
||||
@@ -447,7 +524,28 @@ public final class OpenCodeLauncher extends HerdrPeerLauncher {
|
||||
Files.writeString(charter, charterText);
|
||||
charter.toFile().deleteOnExit();
|
||||
|
||||
root.putArray("instructions").add(charter.toAbsolutePath().toString());
|
||||
// fleetd #393 follow-up: withArray, not putArray. putArray REPLACES whatever node
|
||||
// is already at "instructions" — harmless only as long as this block runs first
|
||||
// against a still-empty root, which is an ordering constraint nothing declared or
|
||||
// tested. The skills writer just below, and the IDE-rules writer further down,
|
||||
// both already use withArray (get-or-create) for exactly this reason; this was the
|
||||
// one straggler. Proven load-bearing on the fleetd #393 merge: flipping this one
|
||||
// call back to putArray left the whole suite green while silently deleting the
|
||||
// charter entry whenever skills or IDE rules ran after it — an opencode member
|
||||
// would launch with no role contract at all, worse than the bug #393 fixed, and
|
||||
// nothing caught it. See OpenCodeLauncherTest's
|
||||
// instructionsArrayHoldsCharterThenIdeRulesInOrder and
|
||||
// instructionsArrayHoldsCharterThenSkillsThenIdeRulesInOrder.
|
||||
root.withArray("instructions").add(charter.toAbsolutePath().toString());
|
||||
}
|
||||
|
||||
// fleetd #393: each seeded skill's SKILL.md, delivered as a plain instructions[] entry
|
||||
// — the only mechanism opencode has for static guidance text. Unlike Claude Code's
|
||||
// Skill tool, opencode cannot load one of these on demand by name; the content is just
|
||||
// always part of the system prompt from spawn. That is a real difference in HOW the
|
||||
// content reaches the member, not a reason to withhold it.
|
||||
for (Path skillFile : skillInstructionFiles) {
|
||||
root.withArray("instructions").add(skillFile.toAbsolutePath().toString());
|
||||
}
|
||||
|
||||
if (cfg.hasMcp() || cfg.hasIdeMcp()) {
|
||||
|
||||
@@ -394,17 +394,17 @@ public final class AmqpReplyInbox implements ReplyInbox, AutoCloseable {
|
||||
}
|
||||
|
||||
@Override
|
||||
public void ack(String target, String msgId) {
|
||||
public boolean ack(String target, String msgId) {
|
||||
var perTarget = held.get(target);
|
||||
if (perTarget == null || perTarget == RELEASED) {
|
||||
return;
|
||||
return false;
|
||||
}
|
||||
Held h;
|
||||
synchronized (perTarget) {
|
||||
h = perTarget.remove(msgId);
|
||||
}
|
||||
if (h == null) {
|
||||
return; // never held (or already acked) — no-op
|
||||
return false; // never held (or already acked) — no-op
|
||||
}
|
||||
try {
|
||||
synchronized (channelLock) {
|
||||
@@ -418,6 +418,7 @@ public final class AmqpReplyInbox implements ReplyInbox, AutoCloseable {
|
||||
}
|
||||
throw new IllegalStateException("cannot ack reply " + msgId + " on " + queueName(target), e);
|
||||
}
|
||||
return true;
|
||||
}
|
||||
|
||||
private DeliverCallback deliverCallback(String target) {
|
||||
|
||||
@@ -59,16 +59,17 @@ public final class InMemoryReplyInbox implements ReplyInbox {
|
||||
}
|
||||
|
||||
@Override
|
||||
public void ack(String target, String msgId) {
|
||||
public boolean ack(String target, String msgId) {
|
||||
if (!owned.contains(target)) {
|
||||
return;
|
||||
return false;
|
||||
}
|
||||
var perTarget = store.get(target);
|
||||
if (perTarget != null) {
|
||||
//noinspection SynchronizationOnLocalVariableOrMethodParameter
|
||||
synchronized (perTarget) {
|
||||
perTarget.remove(msgId);
|
||||
}
|
||||
if (perTarget == null) {
|
||||
return false;
|
||||
}
|
||||
//noinspection SynchronizationOnLocalVariableOrMethodParameter
|
||||
synchronized (perTarget) {
|
||||
return perTarget.remove(msgId) != null;
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
@@ -42,6 +42,18 @@ public interface LeadChannel {
|
||||
/** This daemon's own lead coordination id — the mailbox it owns, and the {@code from} it sends as. */
|
||||
String selfCoordId();
|
||||
|
||||
/**
|
||||
* Whether a message sitting in {@link #peek}'s held set (fetched but not yet {@link #ack}ed) is
|
||||
* still safe if this daemon crashes or restarts right now — the conclusion of two independent
|
||||
* facts about how this channel owns its own queue: the queue was declared <em>durable</em>, and
|
||||
* the consumer that filled {@code held} uses <em>manual ack</em>, so an unacked delivery is still
|
||||
* owned by the broker rather than only in this process's memory. Both must hold for {@code true};
|
||||
* an implementation must derive this from what it actually did when it declared and consumed its
|
||||
* queue, never return a literal — fleetd #440 found {@code FleetMcp}'s {@code heldDurable} field
|
||||
* doing exactly that, unable to ever report {@code false} even after the fact stopped being true.
|
||||
*/
|
||||
boolean heldDurable();
|
||||
|
||||
/**
|
||||
* A non-destructive look at {@code coordId}'s mailbox — does it exist, how many messages are
|
||||
* waiting on it, and how many consumers are attached — without owning, consuming, or otherwise
|
||||
|
||||
@@ -86,6 +86,12 @@ public final class LeadMailbox implements LeadChannel, AutoCloseable {
|
||||
private final Object channelLock = new Object();
|
||||
/** msgId → held delivery, for this mailbox's own queue only (there is exactly one). */
|
||||
private final LinkedHashMap<String, Held> held = new LinkedHashMap<>();
|
||||
/**
|
||||
* fleetd #440: the answer to {@link #heldDurable()}, set once by {@link #own()} from the exact
|
||||
* booleans it passed to {@code queueDeclare}/{@code basicConsume} — never a separate literal that
|
||||
* could drift from what those calls actually did.
|
||||
*/
|
||||
private boolean heldDurable;
|
||||
/** Successful broker acks on this connection, retained only to make a repeated caller ack quiet. */
|
||||
private final LinkedHashMap<String, Boolean> recentlyAcked = new LinkedHashMap<>();
|
||||
/** Bounds {@link #recentlyAcked}: it is only an idempotency aid, never delivery state. */
|
||||
@@ -191,13 +197,22 @@ public final class LeadMailbox implements LeadChannel, AutoCloseable {
|
||||
/** Declare + consume this daemon's own {@code lead.<selfCoordId>.inbox}. Called once, at construction. */
|
||||
private void own() throws IOException {
|
||||
String queue = queueName(selfCoordId);
|
||||
boolean durableQueue = true; // durable, non-exclusive, keep on idle
|
||||
boolean autoAck = false; // manual ack
|
||||
synchronized (channelLock) {
|
||||
channel.queueDeclare(queue, true, false, false, null); // durable, non-exclusive, keep on idle
|
||||
channel.basicConsume(queue, false, deliverCallback(), _ -> { }); // autoAck=false: manual ack
|
||||
channel.queueDeclare(queue, durableQueue, false, false, null);
|
||||
channel.basicConsume(queue, autoAck, deliverCallback(), _ -> { });
|
||||
}
|
||||
// fleetd #440: held mail is durable only while both hold — a durable queue AND manual ack.
|
||||
this.heldDurable = durableQueue && !autoAck;
|
||||
log.debug("lead mailbox owns queue {} for coord-id {}", queue, selfCoordId);
|
||||
}
|
||||
|
||||
@Override
|
||||
public boolean heldDurable() {
|
||||
return heldDurable;
|
||||
}
|
||||
|
||||
/**
|
||||
* Publish {@code msg} to {@code toCoordId}'s mailbox and block until the broker's publisher
|
||||
* confirm for it lands. Does <em>not</em> imply owning or consuming {@code toCoordId}'s queue.
|
||||
|
||||
@@ -826,9 +826,13 @@ public final class MessageService {
|
||||
/**
|
||||
* Acknowledge a specific reply by {@code msgId} for {@code target}. Removes it from the inbox
|
||||
* so that a subsequent drain or peek no longer returns it.
|
||||
*
|
||||
* @return {@code true} if an entry was actually removed, {@code false} if {@code msgId} was not
|
||||
* in {@code target}'s inbox (wrong id, wrong target, or already acked). The caller —
|
||||
* {@link dev.ltms.fleet.mcp.FleetMcp#ack} — must not report success on {@code false}.
|
||||
*/
|
||||
public void ackReply(String target, String msgId) {
|
||||
inbox.ack(target, msgId);
|
||||
public boolean ackReply(String target, String msgId) {
|
||||
return inbox.ack(target, msgId);
|
||||
}
|
||||
|
||||
/**
|
||||
@@ -1317,6 +1321,27 @@ public final class MessageService {
|
||||
return new TaskView(ticket, Phase.FAILED, null, null, detail, null);
|
||||
}
|
||||
|
||||
/**
|
||||
* Test seam only — carries no production behaviour, and nothing in this class calls it;
|
||||
* {@link #pruneTerminalTickets} still reads {@link Task#completedNanos} directly.
|
||||
*
|
||||
* <p>Reports whether {@code ticket}'s completion hook (the {@code whenComplete} registered in
|
||||
* {@link Task}'s constructor) has actually run yet. Exists because {@link #poll} can report
|
||||
* {@link Phase#DONE} for a ticket before that hook fires: {@code CompletableFuture.complete()}
|
||||
* publishes its result and only afterwards runs dependent actions such as {@code whenComplete}
|
||||
* (fleetd #399), so a caller that observes the future done via {@link #poll} is not thereby
|
||||
* guaranteed to also observe {@link Task#completedNanos} stamped. A test that must order both
|
||||
* events — e.g. before advancing an injected clock past the TTL, to avoid stamping the
|
||||
* *advanced* time and masking a real eviction bug — waits on this instead of on
|
||||
* {@link Phase#DONE}.
|
||||
*
|
||||
* @return {@code false} for an unknown ticket or one whose completion hook has not run yet
|
||||
*/
|
||||
boolean isCompletionStampedForTest(String ticket) {
|
||||
Task task = tasks.get(ticket);
|
||||
return task != null && task.completedNanos != null;
|
||||
}
|
||||
|
||||
/** Best-effort live worker status for a pending poll; never throws (a lookup error is just noise). */
|
||||
private String liveStatus(String target) {
|
||||
try {
|
||||
|
||||
@@ -46,6 +46,13 @@ public interface ReplyInbox {
|
||||
/** Non-destructive snapshot of pending replies for {@code target} (FIFO), empty list if none. */
|
||||
List<InboxMessage> peek(String target);
|
||||
|
||||
/** Remove the reply {@code msgId} for {@code target} once the primary has taken it. No-op if absent. */
|
||||
void ack(String target, String msgId);
|
||||
/**
|
||||
* Remove the reply {@code msgId} for {@code target} once the primary has taken it.
|
||||
*
|
||||
* @return {@code true} if an entry was actually removed, {@code false} if there was nothing to
|
||||
* remove (unknown {@code target}, unowned {@code target}, or a {@code msgId} not held for
|
||||
* it). A {@code false} is not an error — acking a {@code target} this daemon does not own is
|
||||
* part of the normal contract, not a failure.
|
||||
*/
|
||||
boolean ack(String target, String msgId);
|
||||
}
|
||||
|
||||
@@ -230,6 +230,26 @@ public final class ReplyPushLoop {
|
||||
.collect(Collectors.toUnmodifiableSet());
|
||||
}
|
||||
|
||||
/**
|
||||
* Test seam only (fleetd #418): carries no production behaviour, and nothing in this class
|
||||
* calls it. Exposes {@link #pendingQuestionTurnIdsFor} — the exact state {@link #decide} reads
|
||||
* to decide whether a question keeps a lead's schedule alive.
|
||||
*
|
||||
* <p>{@code MessageService.ask()} does three things in order before a question is fully open to
|
||||
* this loop: it flips the ticket's {@code poll()} phase to {@code Phase.ASKING}, then resolves
|
||||
* the reverse-rendezvous waiter, then calls {@link #onQuestionOpened}, which is what actually
|
||||
* populates {@link #pendingQuestions}. A test that barriers on {@code Phase.ASKING} observes only
|
||||
* the first of those three steps — under load the asker thread can be descheduled between steps
|
||||
* one and three, so the barrier releases before this method's underlying map is populated, and
|
||||
* {@link #decide} correctly reports nothing pending yet. A test that must order itself after the
|
||||
* state {@link #decide} actually reads waits on this instead of on the phase.
|
||||
*
|
||||
* @return an unmodifiable snapshot; empty for a lead with no open questions
|
||||
*/
|
||||
Set<String> pendingQuestionTurnIdsForTest(String lead) {
|
||||
return pendingQuestionTurnIdsFor(lead);
|
||||
}
|
||||
|
||||
private List<PendingIncident> pendingIncidentsFor(String lead) {
|
||||
return pendingIncidents.values().stream().filter(i -> lead.equals(i.key().lead())).toList();
|
||||
}
|
||||
|
||||
@@ -1,5 +1,7 @@
|
||||
package dev.ltms.fleet.peer;
|
||||
|
||||
import dev.ltms.fleet.placement.PlacementDecision;
|
||||
|
||||
import java.nio.file.Path;
|
||||
import java.util.List;
|
||||
import java.util.Set;
|
||||
@@ -143,9 +145,160 @@ public interface PeerLauncher {
|
||||
|
||||
/**
|
||||
* The profile a no-argument {@link #spawn(SpawnRequest)} uses, or {@code null} if none is configured.
|
||||
*
|
||||
* <p>fleetd #425: for an implementation with role pools (a no-argument spawn is read as {@link
|
||||
* MemberRole#DEV}, see {@link SpawnRequest}), this must be the profile a live spawn of that role
|
||||
* would actually be placed on right now, not a value captured once at startup — a caller such as
|
||||
* {@code fleet_profiles} relies on this to report a live, not frozen, fact.
|
||||
*/
|
||||
String defaultProfile();
|
||||
|
||||
/**
|
||||
* The profile an unqualified spawn of {@code role} would resolve to right now — the role-aware,
|
||||
* live counterpart of {@link #defaultProfile()} (fleetd #425).
|
||||
*
|
||||
* <p>A caller that must provision something profile-specific (working directory, parity overlay
|
||||
* files) <em>before</em> the actual spawn — {@code SessionManager.acquireWithWorktree} is the one
|
||||
* that exists today — needs the exact profile that spawn will use, for the caller's real role,
|
||||
* not a role-agnostic guess. Calling {@link #defaultProfile()} for that purpose reads {@code
|
||||
* MemberRole#DEV}'s answer regardless of the caller's actual role, which is wrong for any other
|
||||
* role and can provision for a profile the spawn never lands on.
|
||||
*
|
||||
* <p>Default implementation returns {@link #defaultProfile()}, ignoring {@code role} — the right
|
||||
* answer for a launcher with no role-pool concept of its own (e.g. a single {@code
|
||||
* HerdrPeerLauncher} adapter, which is never reached this way in production: {@code
|
||||
* CompositePeerLauncher} always fronts it and resolves roles itself).
|
||||
*
|
||||
* <p>fleetd #453: this default is deliberately <em>not</em> abstract — unlike {@link
|
||||
* #spawn(SpawnRequest, PlacementDecision)} (fleetd #450), there is no live defect in inheriting
|
||||
* it today, and the only current single-adapter implementer ({@code HerdrPeerLauncher}) is
|
||||
* correct to do so. But it stays correct only as long as that holds: <strong>if a launcher ever
|
||||
* routes more than one profile per role, it MUST override this method</strong>, or every role
|
||||
* silently resolves to {@link #defaultProfile()} with no error and no log line. {@code
|
||||
* HerdrPeerLauncher.spawn(SpawnRequest, PlacementDecision)} — the override in {@code
|
||||
* dev.ltms.fleet.member}, not the declaration below — names this method and {@link #place}
|
||||
* explicitly as "unoverridden here" for exactly this reason. Read it before adding role-pool
|
||||
* routing to any {@code HerdrPeerLauncher} subclass.
|
||||
*
|
||||
* <p>Who is forced to read which paragraph, because it is not symmetric. A new class that
|
||||
* implements this interface directly must write a body for {@link #spawn(SpawnRequest,
|
||||
* PlacementDecision)}, which is abstract here, so it lands on this javadoc. A subclass of
|
||||
* {@code HerdrPeerLauncher} does not: that class already implements the method, and the
|
||||
* subclass inherits the body. So for a subclass this paragraph is advice, not a gate.
|
||||
*/
|
||||
default String defaultProfileFor(MemberRole role) {
|
||||
return defaultProfile();
|
||||
}
|
||||
|
||||
/**
|
||||
* The profile an <em>unqualified</em> spawn of {@code role} would actually be routed to right
|
||||
* now — the same candidate list, the same {@code quarantined}/{@code coolingOff}/{@code
|
||||
* modelOff} filtering, and the same {@code PlacementPolicy} that {@link #spawn} itself
|
||||
* consults for a blank-profile request (fleetd #425 rework).
|
||||
*
|
||||
* <p>This is <em>not</em> {@link #defaultProfileFor}: that method answers "what is first in
|
||||
* {@code role}'s pool", blind to quarantine, cool-off, and the model on/off gate — the right
|
||||
* answer for a role-agnostic, best-effort report ({@code fleet_profiles}' {@code "default"}
|
||||
* field), but the wrong one for a caller that needs the profile a spawn will actually land on.
|
||||
* A quarantined or model-off pool-first profile makes {@link #defaultProfileFor} return a name
|
||||
* an unqualified spawn will never be routed to.
|
||||
*
|
||||
* <p>Just the resolved name, not the full {@link PlacementDecision} — a caller that only wants
|
||||
* to know the answer (a status report, a log line) can call this; a caller that will later
|
||||
* <em>act</em> on the answer by spawning — provisioning a worktree for a specific profile
|
||||
* before the peer exists is the one that matters — must call {@link #place} and carry the
|
||||
* {@link PlacementDecision} itself through to {@link #spawn(SpawnRequest, PlacementDecision)}
|
||||
* instead of calling this method and feeding the string back in as an explicit profile. Doing
|
||||
* that re-enters {@link #spawn(SpawnRequest)}'s explicit-profile branch, which disagrees with
|
||||
* the routing branch on purpose about what an excluded profile means: the routing branch (and
|
||||
* {@link #place}) falls through a quarantined/cooling-off/at-cap/model-off/unreachable/weight-0
|
||||
* profile to the next candidate, while the explicit branch refuses outright — correct for an
|
||||
* operator who named that profile on purpose, wrong for a name that only ever came from placement
|
||||
* itself. That accidental refusal is exactly the regression fleetd #425 rework round 2 fixes:
|
||||
* the default implementation below delegates to {@link #place}, so the two can never drift apart,
|
||||
* but a caller that resolves through this method alone and spawns separately can still recreate
|
||||
* the round-1 defect for itself. (Before fleetd #435, this accident was also reachable through
|
||||
* {@code maxLoad} specifically, because {@code FixedPlacementPolicy} — the default policy — did
|
||||
* not evaluate it at all for automatic selection; #435 closed that gap, so a placement decision
|
||||
* can no longer be at cap in the first place. The refusal-vs-fall-through disagreement above is
|
||||
* the part that was never about {@code maxLoad} and is still real.)
|
||||
*
|
||||
* @throws RuntimeException (implementation-specific, typically a placement exception) if no
|
||||
* candidate in {@code role}'s pool is currently placeable
|
||||
*/
|
||||
default String routedProfileFor(MemberRole role) {
|
||||
return place(role).profile();
|
||||
}
|
||||
|
||||
/**
|
||||
* Resolve, <em>without spawning</em>, the {@link PlacementDecision} an unqualified spawn of
|
||||
* {@code role} would make right now — the same candidate list, the same {@code
|
||||
* quarantined}/{@code coolingOff}/{@code modelOff} filtering, and the same {@code
|
||||
* PlacementPolicy} {@link #spawn(SpawnRequest)}'s blank-profile branch itself consults (fleetd
|
||||
* #425 rework).
|
||||
*
|
||||
* <p>Pair this with {@link #spawn(SpawnRequest, PlacementDecision)}, never with {@link
|
||||
* #spawn(SpawnRequest)} fed the decision's profile as an explicit name — see {@link
|
||||
* PlacementDecision}'s own javadoc for why that second form regressed.
|
||||
*
|
||||
* <p>Default implementation wraps {@link #defaultProfile()}, ignoring {@code role} and every
|
||||
* placement condition — the right answer for a launcher with no pool or placement-policy
|
||||
* concept of its own, matching {@link #defaultProfileFor}'s own default.
|
||||
*
|
||||
* <p>fleetd #453: same reasoning as {@link #defaultProfileFor}'s own #453 note — this default
|
||||
* is deliberately not abstract (no live defect today, correct for the sole single-adapter
|
||||
* implementer), but <strong>a launcher that ever routes more than one profile per role MUST
|
||||
* override this method too</strong>, or placement silently ignores {@code role} for it. See
|
||||
* {@code HerdrPeerLauncher.spawn(SpawnRequest, PlacementDecision)}'s javadoc, which names this
|
||||
* method as "unoverridden here" and why that is correct only for a single-profile adapter.
|
||||
*
|
||||
* @throws RuntimeException (implementation-specific, typically a placement exception) if no
|
||||
* candidate in {@code role}'s pool is currently placeable
|
||||
*/
|
||||
default PlacementDecision place(MemberRole role) {
|
||||
return new PlacementDecision(defaultProfile());
|
||||
}
|
||||
|
||||
/**
|
||||
* Spawn against an already-resolved {@link PlacementDecision} from {@link #place}, honoring it
|
||||
* completely: none of the conditions {@link #place} already applied — quarantine, cooling off,
|
||||
* {@code maxLoad} (evaluated by every placement policy including the default {@code fixed},
|
||||
* since fleetd #435), model-off — are re-evaluated here; {@code decision} already reflects them.
|
||||
* This is not skipping a check {@code place} left undone; it is not repeating one {@code place}
|
||||
* already did, and not re-opening the window between that decision and this spawn in which the
|
||||
* underlying state could otherwise move. This is what lets a resolve-then-spawn caller
|
||||
* ({@code SessionManager.acquireWithWorktree}, which must know the profile before it can
|
||||
* provision a worktree for it) and a plain blank-profile {@link #spawn(SpawnRequest)} caller
|
||||
* land on the exact same outcome for the exact same placement state (fleetd #425 rework,
|
||||
* round 2).
|
||||
*
|
||||
* <p>{@code req}'s own {@link SpawnRequest#profileName()} is ignored in favor of {@code
|
||||
* decision.profile()} — the caller is expected to have built {@code req} with a blank or
|
||||
* matching profile; passing a request that names a <em>different</em>, explicit profile than
|
||||
* the decision it is paired with is a caller bug this method does not attempt to detect.
|
||||
*
|
||||
* <p>No default implementation (fleetd #450): the two correct bodies disagree on purpose, so an
|
||||
* implementer must choose one rather than silently inherit whichever this interface happened to
|
||||
* provide. An implementer with no placement concept of its own — spawns a single profile, e.g.
|
||||
* {@code HerdrPeerLauncher} — should delegate to {@link #spawn(SpawnRequest)} with the decision's
|
||||
* profile named explicitly, since there is no separate routing path to honor there: the
|
||||
* explicit-profile branch it re-enters and the routing branch {@link #place} would have used are
|
||||
* the same thing. <strong>A launcher that routes across more than one profile — the way {@code
|
||||
* CompositePeerLauncher} routes across every configured adapter — MUST NOT re-enter {@link
|
||||
* #spawn(SpawnRequest)}.</strong> Doing so re-applies that single-argument method's
|
||||
* explicit-profile checks ({@code enforceNotQuarantined}, {@code enforceNotCoolingOff}, {@code
|
||||
* enforceMaxLoad}, {@code enforceModelEnabled} in {@code CompositePeerLauncher}), which can
|
||||
* refuse the very profile {@link #place} just chose, if the underlying placement state moved in
|
||||
* the window between the {@link #place} call and this one — the exact window this method and
|
||||
* {@link PlacementDecision} exist to close (fleetd #444). Before #450 this was a {@code default}
|
||||
* method that only {@code CompositePeerLauncher} overrode; a future placement-doing launcher
|
||||
* could have inherited the re-entering body silently and never known. Making it abstract turns
|
||||
* that silent inheritance into a compile error.
|
||||
*
|
||||
* @throws IllegalArgumentException if the decision names an unknown profile
|
||||
*/
|
||||
PeerHandle spawn(SpawnRequest req, PlacementDecision decision);
|
||||
|
||||
/**
|
||||
* Resolve the effective working directory for a spawn {@code req} without actually spawning.
|
||||
* Resolution order: requestedCwd → profile cwd → callerCwd → daemon cwd.
|
||||
@@ -192,4 +345,66 @@ public interface PeerLauncher {
|
||||
* @return {@code true} when a reset was sent and its status transition must settle before reuse
|
||||
*/
|
||||
boolean clearContext(String id);
|
||||
|
||||
/**
|
||||
* Model ids the operator has currently turned off in the central {@code models.allow:} list
|
||||
* (fleetd #422) — empty for a launcher with nothing to gate against. {@code fleet_profiles}/
|
||||
* {@code GET /profiles} (via {@code FleetMcp.profilesView}) call this to report which models
|
||||
* are off, and MUST read this exact accessor rather than deriving their own answer: the fleetd
|
||||
* #404 lesson is that a status field reading a different source than the behaviour it describes
|
||||
* can drift from what the gate ({@code CompositePeerLauncher.enforceModelEnabled} and its
|
||||
* candidate filter) actually enforces. A default of {@code Set.of()} keeps every other {@link
|
||||
* PeerLauncher} implementer (the herdr adapters, and the two test-fake implementers) unchanged.
|
||||
*
|
||||
* <p>fleetd #422 follow-up: this alone cannot tell "no {@code models:} block at all" from "a
|
||||
* {@code models:} block where nothing is currently off" — both report an empty set here. Delegates
|
||||
* to {@link #modelGateState()} so the two facts always come from the one read {@link
|
||||
* #modelGateState()}'s implementer makes; do not override this method separately from that one.
|
||||
*/
|
||||
default Set<String> disabledModels() {
|
||||
return modelGateState().off();
|
||||
}
|
||||
|
||||
/**
|
||||
* Whether the central {@code models.allow:} gate (fleetd #422) is armed at all, together with
|
||||
* which model ids are currently off — fleetd #422 follow-up. {@link #disabledModels()} alone
|
||||
* cannot distinguish two states that both report an empty set: a host with no {@code models:}
|
||||
* block (nothing is gated, and nothing can be) and a host WITH a {@code models:} block where
|
||||
* nothing is currently turned off (the gate is armed and reporting zero). This method exists so
|
||||
* a caller — the startup log, {@code fleet_profiles}/{@code GET /profiles} — can tell the two
|
||||
* apart, the same reason {@code CompletionResolver.UnsetMeaning} exists: an accessor that can
|
||||
* legitimately report "empty" must never let a caller guess why.
|
||||
*
|
||||
* <p>Default {@link ModelGateState#notConfigured()} — every launcher without a {@code models:}
|
||||
* block to read from (the herdr adapters, and the two test-fake implementers), matching {@link
|
||||
* #disabledModels()}'s own default of an empty set.
|
||||
*/
|
||||
default ModelGateState modelGateState() {
|
||||
return ModelGateState.notConfigured();
|
||||
}
|
||||
|
||||
/**
|
||||
* fleetd #422 follow-up: the result of {@link #modelGateState()} — see that method's javadoc
|
||||
* for why "armed" and "off" must be reported together from one read rather than as two
|
||||
* separately-derived facts that a reload landing between them could make disagree.
|
||||
*
|
||||
* @param configured {@code true} when a {@code models:} block exists at all (armed), regardless
|
||||
* of whether anything in it is currently turned off; {@code false} when there
|
||||
* is no block to gate against
|
||||
* @param off the model ids currently turned off; always empty when {@code configured} is
|
||||
* {@code false}
|
||||
*/
|
||||
record ModelGateState(boolean configured, Set<String> off) {
|
||||
public ModelGateState {
|
||||
off = Set.copyOf(off);
|
||||
}
|
||||
|
||||
public static ModelGateState notConfigured() {
|
||||
return new ModelGateState(false, Set.of());
|
||||
}
|
||||
|
||||
public static ModelGateState armed(Set<String> off) {
|
||||
return new ModelGateState(true, off);
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
@@ -5,14 +5,12 @@ import java.util.List;
|
||||
|
||||
/**
|
||||
* Backward-compatible placement: an unqualified spawn always resolves to the configured default
|
||||
* profile, exactly as {@code CompositePeerLauncher} did before CB-518. This ignores caps
|
||||
* ({@code maxLoad}) so that a pre-existing config behaves identically after upgrade — capacity
|
||||
* gating for automatic placement is deliberately out of scope for {@code fixed}, exactly as it
|
||||
* always has been. Reachability is a narrower exception (fleetd #315, below): a profile is never
|
||||
* checked for reachability up front, only skipped once it has already failed in <em>this same</em>
|
||||
* spawn call's retry loop — see the unreachable case below.
|
||||
* profile, exactly as {@code CompositePeerLauncher} did before CB-518. Reachability is a narrower
|
||||
* exception (fleetd #315, below): a profile is never checked for reachability up front, only
|
||||
* skipped once it has already failed in <em>this same</em> spawn call's retry loop — see the
|
||||
* unreachable case below.
|
||||
*
|
||||
* <p>Four exceptions walk past the default instead of returning it unconditionally:
|
||||
* <p>Six exceptions walk past the default instead of returning it unconditionally:
|
||||
* <ul>
|
||||
* <li>Quarantine (CB-578 stage B): a quarantined default is a credential that just refused on
|
||||
* a usage limit, not a transient capacity or reachability concern.
|
||||
@@ -20,6 +18,24 @@ import java.util.List;
|
||||
* ({@code BackendOutagePolicy}) — a separate, shorter-lived source from quarantine. When a
|
||||
* profile is both quarantined and cooling off, only the quarantine reason is reported
|
||||
* (exhaustion takes priority), matching {@code CompositePeerLauncher}'s explicit-spawn order.
|
||||
* <li>At cap (fleetd #435): a profile whose live count has reached its {@code maxLoad}
|
||||
* ({@link PlacementPolicyUtil#atCap}) — a documented, unconditional capacity limit (see
|
||||
* {@code FleetConfig.Profile#maxLoad}), so {@code fixed} must gate on it exactly as {@code
|
||||
* weighted}/{@code round-robin} already do via {@link PlacementPolicyUtil#available}. Before
|
||||
* this fix {@code fixed} built its own {@link PlacementCandidate} for the default with {@code
|
||||
* maxLoad} forced to {@code null}, so a capped default was chosen anyway on every unqualified
|
||||
* spawn — the cap was advisory, not enforced, for the one placement policy every config uses
|
||||
* by default. Reported only when quarantine and cooling off are both absent, matching {@code
|
||||
* CompositePeerLauncher}'s explicit-spawn check order (quarantine, then cooling off, then max
|
||||
* load, then model-off).
|
||||
* <li>Model off (fleetd #422): a profile whose {@code model:} the operator has turned off in
|
||||
* {@code models.allow:} — an operator decision, never a backend-reported outage, so it is a
|
||||
* fifth, independent source from quarantine, cooling off, and at-cap (never merged with any
|
||||
* of them), exactly as {@code CompositePeerLauncher.enforceModelEnabled} and {@link
|
||||
* PlacementPolicyUtil#available} treat it. When a profile is model-off <em>and</em> quarantined,
|
||||
* cooling off, or at cap, only the higher-priority reason is reported, matching {@code
|
||||
* CompositePeerLauncher}'s explicit-spawn check order (quarantine, then cooling off, then max
|
||||
* load, then model-off).
|
||||
* <li>Unreachable (fleetd #315): {@code CompositePeerLauncher.spawn} retries a failed candidate
|
||||
* on the next one and rebuilds the {@link PlacementContext} so {@code ctx.unreachable()}
|
||||
* names every profile that already failed with {@code PeerUnreachableException} in this same
|
||||
@@ -32,9 +48,10 @@ import java.util.List;
|
||||
* {@code weighted}/{@code round-robin} skip it — an explicit {@code fleet_spawn} naming
|
||||
* the profile is unaffected, only this automatic fallback walk.
|
||||
* </ul>
|
||||
* A fleet where nothing is ever quarantined, cooling off, unreachable, or weight-0 never exercises
|
||||
* any of these paths, so today's behaviour is unchanged — in particular, the very first selection
|
||||
* of a spawn call always sees an empty {@code unreachable} set, so the first choice is untouched.
|
||||
* A fleet where nothing is ever quarantined, cooling off, at cap, model-off, unreachable, or
|
||||
* weight-0 never exercises any of these paths, so today's behaviour is unchanged — in particular,
|
||||
* the very first selection of a spawn call always sees an empty {@code unreachable} set, so the
|
||||
* first choice is untouched.
|
||||
*/
|
||||
final class FixedPlacementPolicy implements PlacementPolicy {
|
||||
|
||||
@@ -42,12 +59,14 @@ final class FixedPlacementPolicy implements PlacementPolicy {
|
||||
public PlacementCandidate select(PlacementContext ctx) {
|
||||
String d = ctx.defaultProfile();
|
||||
if (d != null && !d.isBlank() && !ctx.quarantined().contains(d) && !ctx.coolingOff().contains(d)
|
||||
&& !ctx.unreachable().contains(d) && !weightExcluded(ctx, d)) {
|
||||
&& !ctx.modelOff().contains(d) && !ctx.unreachable().contains(d) && !weightExcluded(ctx, d)
|
||||
&& !capExcluded(ctx, d)) {
|
||||
return new PlacementCandidate(d, null, 1.0f, null);
|
||||
}
|
||||
for (PlacementCandidate c : ctx.candidates()) {
|
||||
if (!ctx.quarantined().contains(c.profile()) && !ctx.coolingOff().contains(c.profile())
|
||||
&& !ctx.unreachable().contains(c.profile()) && !c.excluded()) {
|
||||
&& !ctx.modelOff().contains(c.profile()) && !ctx.unreachable().contains(c.profile())
|
||||
&& !c.excluded() && !PlacementPolicyUtil.atCap(ctx, c)) {
|
||||
return new PlacementCandidate(c.profile(), null, c.weight(), c.maxLoad());
|
||||
}
|
||||
}
|
||||
@@ -56,9 +75,18 @@ final class FixedPlacementPolicy implements PlacementPolicy {
|
||||
// Exhaustion quarantine takes priority: reported only when quarantine is absent, so the
|
||||
// message never claims "cooling off" for a profile that is really backend-exhausted.
|
||||
boolean dCoolingOff = !dQuarantined && ctx.coolingOff().contains(d);
|
||||
// fleetd #435: at-cap sits between cooling off and model-off, matching
|
||||
// CompositePeerLauncher's explicit-spawn check order (quarantine, cooling off, max load,
|
||||
// then model-off) — reported only when quarantine/cooling-off are both absent.
|
||||
boolean dAtCap = !dQuarantined && !dCoolingOff && capExcluded(ctx, d);
|
||||
// fleetd #422: model-off is a fifth, independent source (an operator decision) — but
|
||||
// quarantine/cooling-off/at-cap still take priority when more than one applies, matching
|
||||
// CompositePeerLauncher's explicit-spawn check order (quarantine, cooling off, max load,
|
||||
// then model-off).
|
||||
boolean dModelOff = !dQuarantined && !dCoolingOff && !dAtCap && ctx.modelOff().contains(d);
|
||||
boolean dUnreachable = ctx.unreachable().contains(d);
|
||||
boolean dWeightExcluded = weightExcluded(ctx, d);
|
||||
if (dQuarantined || dCoolingOff || dUnreachable || dWeightExcluded) {
|
||||
if (dQuarantined || dCoolingOff || dAtCap || dModelOff || dUnreachable || dWeightExcluded) {
|
||||
List<String> reasons = new ArrayList<>();
|
||||
if (dQuarantined) {
|
||||
reasons.add("is quarantined (backend exhausted)");
|
||||
@@ -66,6 +94,14 @@ final class FixedPlacementPolicy implements PlacementPolicy {
|
||||
if (dCoolingOff) {
|
||||
reasons.add("is cooling off after repeated backend errors");
|
||||
}
|
||||
if (dAtCap) {
|
||||
PlacementCandidate c = candidateFor(ctx, d);
|
||||
int live = ctx.liveCount().apply(d);
|
||||
reasons.add("is at maxLoad (" + live + " live >= " + c.maxLoad() + " cap)");
|
||||
}
|
||||
if (dModelOff) {
|
||||
reasons.add("names a model the operator has turned off in models.allow");
|
||||
}
|
||||
if (dUnreachable) {
|
||||
reasons.add("is unreachable");
|
||||
}
|
||||
@@ -78,18 +114,38 @@ final class FixedPlacementPolicy implements PlacementPolicy {
|
||||
}
|
||||
if (!ctx.candidates().isEmpty()) {
|
||||
throw new PlacementException("all worker profiles are excluded from automatic "
|
||||
+ "selection (quarantined, cooling off, unreachable, or weight-0)");
|
||||
+ "selection (quarantined, cooling off, at cap, model-off, unreachable, or weight-0)");
|
||||
}
|
||||
throw new PlacementException("no worker profiles configured");
|
||||
}
|
||||
|
||||
/** Whether {@code profile} carries {@code weight <= 0} (CB-554) among {@code ctx}'s candidates. */
|
||||
private static boolean weightExcluded(PlacementContext ctx, String profile) {
|
||||
PlacementCandidate c = candidateFor(ctx, profile);
|
||||
return c != null && c.excluded();
|
||||
}
|
||||
|
||||
/**
|
||||
* Whether {@code profile} has reached its {@code maxLoad} cap (fleetd #435), using the shared
|
||||
* {@link PlacementPolicyUtil#atCap} definition — the same one {@code weighted}/{@code
|
||||
* round-robin} already consult via {@link PlacementPolicyUtil#available}. Looked up by name,
|
||||
* the same way {@link #weightExcluded} is: the default fast path above builds its own {@link
|
||||
* PlacementCandidate} with {@code maxLoad} forced to {@code null} (it carries no cap of its
|
||||
* own), so the candidate actually configured for {@code profile} has to be found in {@code
|
||||
* ctx.candidates()} first.
|
||||
*/
|
||||
private static boolean capExcluded(PlacementContext ctx, String profile) {
|
||||
PlacementCandidate c = candidateFor(ctx, profile);
|
||||
return c != null && PlacementPolicyUtil.atCap(ctx, c);
|
||||
}
|
||||
|
||||
/** The configured candidate named {@code profile} in {@code ctx}, or {@code null} if none. */
|
||||
private static PlacementCandidate candidateFor(PlacementContext ctx, String profile) {
|
||||
for (PlacementCandidate c : ctx.candidates()) {
|
||||
if (c.profile().equals(profile)) {
|
||||
return c.excluded();
|
||||
return c;
|
||||
}
|
||||
}
|
||||
return false;
|
||||
return null;
|
||||
}
|
||||
}
|
||||
|
||||
@@ -22,11 +22,29 @@ import java.util.function.Function;
|
||||
* the two apart so its refusal message says "cooling off", not "exhausted",
|
||||
* when only this one is active. A profile can be in both sets at once; when it
|
||||
* is, exhaustion quarantine is reported (it takes priority).
|
||||
* @param modelOff profiles whose {@code model:} is currently turned off in {@code
|
||||
* models.allow:} (fleetd #422) — an operator decision, not a backend-reported
|
||||
* outage, so a SEPARATE, independent source from both {@code quarantined} and
|
||||
* {@code coolingOff}. A profile can be in this set together with either (or
|
||||
* both) of the others; {@link PlacementPolicyUtil} counts it into its own
|
||||
* bucket rather than merging it into theirs, the same reason
|
||||
* {@code coolingOff} is kept apart from {@code quarantined}.
|
||||
*/
|
||||
public record PlacementContext(String defaultProfile,
|
||||
List<PlacementCandidate> candidates,
|
||||
Function<String, Integer> liveCount,
|
||||
Set<String> unreachable,
|
||||
Set<String> quarantined,
|
||||
Set<String> coolingOff) {
|
||||
Set<String> coolingOff,
|
||||
Set<String> modelOff) {
|
||||
|
||||
/**
|
||||
* Back-compat form before the fleetd #422 model on/off gate was added — no candidate's model
|
||||
* is off. Keeps pre-#422 call sites (tests included) compiling and behaving identically.
|
||||
*/
|
||||
public PlacementContext(String defaultProfile, List<PlacementCandidate> candidates,
|
||||
Function<String, Integer> liveCount, Set<String> unreachable,
|
||||
Set<String> quarantined, Set<String> coolingOff) {
|
||||
this(defaultProfile, candidates, liveCount, unreachable, quarantined, coolingOff, Set.of());
|
||||
}
|
||||
}
|
||||
|
||||
@@ -0,0 +1,49 @@
|
||||
package dev.ltms.fleet.placement;
|
||||
|
||||
import dev.ltms.fleet.peer.MemberRole;
|
||||
import dev.ltms.fleet.peer.PeerLauncher;
|
||||
import dev.ltms.fleet.peer.SpawnRequest;
|
||||
|
||||
/**
|
||||
* An already-completed placement choice — the outcome of one {@link PeerLauncher#place} call,
|
||||
* carried forward so a later {@link PeerLauncher#spawn(SpawnRequest, PlacementDecision)} can honor
|
||||
* it directly instead of re-resolving the profile a second time (fleetd #425 rework, round 2).
|
||||
*
|
||||
* <p>The problem this exists to close: a caller that must know the profile <em>before</em> it can
|
||||
* spawn — {@code SessionManager.acquireWithWorktree} provisions a worktree's {@code repoRoot} and
|
||||
* parity overlay for a specific profile before the peer process exists — used to resolve that name
|
||||
* with {@code PeerLauncher.routedProfileFor(role)} and then hand the SAME string back to {@link
|
||||
* PeerLauncher#spawn(SpawnRequest)} as an EXPLICIT profile. That re-resolution is not free: naming
|
||||
* a profile explicitly makes {@code CompositePeerLauncher.spawn} take its THROWING branch
|
||||
* ({@code enforceNotQuarantined}/{@code enforceNotCoolingOff}/{@code enforceMaxLoad}/{@code
|
||||
* enforceModelEnabled}), while an unqualified spawn's ROUTING branch never runs those checks at
|
||||
* all — it instead FALLS THROUGH to the next candidate on exactly the same conditions the throwing
|
||||
* branch refuses on. That disagreement is deliberate: an operator who names a profile should get a
|
||||
* refusal, not a silent substitution. The bug is turning the fall-through into a refusal by
|
||||
* accident — resolving a name through the routing side and then re-entering the refusing side with
|
||||
* it, for a decision the routing side had already approved by walking past everything else.
|
||||
* Before fleetd #435, this accident was also reachable through {@code maxLoad} specifically: the
|
||||
* default {@code fixed} placement policy did not evaluate {@code maxLoad} at all for automatic
|
||||
* selection, so a profile placement itself just approved could still die at {@code enforceMaxLoad}
|
||||
* one call later, purely because the caller's route to the spawn passed through an explicit
|
||||
* profile name instead of the routing branch — a failure a worktree-less unqualified spawn would
|
||||
* never hit. fleetd #435 closed that specific gap ({@code fixed} now evaluates {@code maxLoad}
|
||||
* exactly like every other placement policy), so a {@link PlacementDecision} can no longer be
|
||||
* at-cap in the first place — but the refusal-vs-fall-through disagreement above was never about
|
||||
* {@code maxLoad}, and resolving a name and re-entering the refusing branch with it is still wrong
|
||||
* for every OTHER condition placement filters on.
|
||||
*
|
||||
* <p>{@link PeerLauncher#spawn(SpawnRequest, PlacementDecision)} closes that by spawning through
|
||||
* the identical code path the routing branch itself uses, keyed off the SAME decision {@link
|
||||
* PeerLauncher#place} returned — no re-checking of any condition placement already evaluated. A
|
||||
* resolve-then-spawn caller and a blank-profile {@link PeerLauncher#spawn(SpawnRequest)} caller can
|
||||
* then never disagree about which conditions apply to the same placement state, and neither one
|
||||
* re-opens the window between the placement decision and the spawn in which the underlying state
|
||||
* could otherwise move.
|
||||
*
|
||||
* @param profile the profile this decision resolved to (may be {@code null} only when no profile is
|
||||
* configured at all — the same corner case {@link PeerLauncher#defaultProfile()}
|
||||
* already tolerates)
|
||||
*/
|
||||
public record PlacementDecision(String profile) {
|
||||
}
|
||||
@@ -11,28 +11,40 @@ final class PlacementPolicyUtil {
|
||||
private PlacementPolicyUtil() {
|
||||
}
|
||||
|
||||
/**
|
||||
* True when {@code c} has reached its {@code maxLoad} cap: {@code liveCount(c.profile()) >=
|
||||
* c.maxLoad()}. A {@code null} maxLoad means unlimited, so it is never at cap.
|
||||
*
|
||||
* <p>Extracted as the single shared definition of "at cap" (fleetd #435): before this fix it
|
||||
* was computed inline in both {@link #available} and {@link #emptyException}, and {@code
|
||||
* FixedPlacementPolicy} — not a caller of either — quietly kept its own {@code select} free of
|
||||
* any cap check at all, so a capped default profile was chosen anyway under the default
|
||||
* placement policy. Every automatic policy must call this, not re-derive it.
|
||||
*/
|
||||
static boolean atCap(PlacementContext ctx, PlacementCandidate c) {
|
||||
Integer cap = c.maxLoad();
|
||||
return cap != null && ctx.liveCount().apply(c.profile()) >= cap;
|
||||
}
|
||||
|
||||
/**
|
||||
* Candidates that are not weight-excluded (CB-554: explicit {@code weight <= 0}, checked
|
||||
* first because it is a static config choice rather than transient state), not
|
||||
* known-unreachable, not quarantined (CB-578 stage B), not cooling off after repeated backend
|
||||
* errors (fleetd #201 Unit 5 — a separate, shorter-lived source from quarantine), and have not
|
||||
* reached their maxLoad. A {@code null} maxLoad means unlimited.
|
||||
* errors (fleetd #201 Unit 5 — a separate, shorter-lived source from quarantine), not naming a
|
||||
* model the operator has turned off (fleetd #422 — a third, independent source: an operator
|
||||
* decision, never a backend-reported outage), and have not reached their maxLoad (see {@link
|
||||
* #atCap}). A {@code null} maxLoad means unlimited.
|
||||
*/
|
||||
static List<PlacementCandidate> available(PlacementContext ctx) {
|
||||
List<PlacementCandidate> out = new ArrayList<>();
|
||||
for (PlacementCandidate c : ctx.candidates()) {
|
||||
if (c.excluded() || ctx.unreachable().contains(c.profile())
|
||||
|| ctx.quarantined().contains(c.profile())
|
||||
|| ctx.coolingOff().contains(c.profile())) {
|
||||
|| ctx.coolingOff().contains(c.profile())
|
||||
|| ctx.modelOff().contains(c.profile())
|
||||
|| atCap(ctx, c)) {
|
||||
continue;
|
||||
}
|
||||
Integer cap = c.maxLoad();
|
||||
if (cap != null) {
|
||||
int live = ctx.liveCount().apply(c.profile());
|
||||
if (live >= cap) {
|
||||
continue;
|
||||
}
|
||||
}
|
||||
out.add(c);
|
||||
}
|
||||
return out;
|
||||
@@ -40,12 +52,13 @@ final class PlacementPolicyUtil {
|
||||
|
||||
/**
|
||||
* Build a clear exception describing why every candidate was dropped: all weight-0, all
|
||||
* quarantined, all cooling off, all at capacity, all unreachable, or a mix. Each candidate is
|
||||
* counted into exactly one bucket (weight-excluded first, then quarantined, then cooling off)
|
||||
* so a candidate excluded for more than one reason is never double-counted — a candidate that is
|
||||
* both quarantined (CB-578 stage B, backend exhausted) and cooling off (fleetd #201 Unit 5,
|
||||
* repeated backend errors) counts only as quarantined, matching {@code CompositePeerLauncher}'s
|
||||
* explicit-spawn ordering: exhaustion quarantine takes priority when both are active.
|
||||
* quarantined, all cooling off, all model-off, all at capacity, all unreachable, or a mix. Each
|
||||
* candidate is counted into exactly one bucket (weight-excluded first, then quarantined, then
|
||||
* cooling off, then model-off) so a candidate excluded for more than one reason is never
|
||||
* double-counted — a candidate that is both quarantined (CB-578 stage B, backend exhausted) and
|
||||
* cooling off (fleetd #201 Unit 5, repeated backend errors) counts only as quarantined, matching
|
||||
* {@code CompositePeerLauncher}'s explicit-spawn ordering: exhaustion quarantine takes priority
|
||||
* when more than one applies.
|
||||
*/
|
||||
static PlacementException emptyException(PlacementContext ctx) {
|
||||
int weightExcluded = 0;
|
||||
@@ -53,17 +66,19 @@ final class PlacementPolicyUtil {
|
||||
int unreachable = 0;
|
||||
int quarantined = 0;
|
||||
int coolingOff = 0;
|
||||
int modelOff = 0;
|
||||
for (PlacementCandidate c : ctx.candidates()) {
|
||||
Integer cap = c.maxLoad();
|
||||
if (c.excluded()) {
|
||||
weightExcluded++;
|
||||
} else if (ctx.quarantined().contains(c.profile())) {
|
||||
quarantined++;
|
||||
} else if (ctx.coolingOff().contains(c.profile())) {
|
||||
coolingOff++;
|
||||
} else if (ctx.modelOff().contains(c.profile())) {
|
||||
modelOff++;
|
||||
} else if (ctx.unreachable().contains(c.profile())) {
|
||||
unreachable++;
|
||||
} else if (cap != null && ctx.liveCount().apply(c.profile()) >= cap) {
|
||||
} else if (atCap(ctx, c)) {
|
||||
atCap++;
|
||||
}
|
||||
}
|
||||
@@ -83,6 +98,10 @@ final class PlacementPolicyUtil {
|
||||
return new PlacementException(
|
||||
"all worker profiles are cooling off after repeated backend errors");
|
||||
}
|
||||
if (modelOff == total) {
|
||||
return new PlacementException(
|
||||
"all worker profiles name a model the operator has turned off");
|
||||
}
|
||||
if (atCap == total) {
|
||||
return new PlacementException("all worker profiles are at maxLoad");
|
||||
}
|
||||
@@ -92,8 +111,9 @@ final class PlacementPolicyUtil {
|
||||
return new PlacementException("no worker profile available: " + atCap + " at maxLoad, "
|
||||
+ unreachable + " unreachable, " + quarantined + " quarantined, "
|
||||
+ coolingOff + " cooling off, "
|
||||
+ modelOff + " model-off, "
|
||||
+ weightExcluded + " weight-0, "
|
||||
+ (total - atCap - unreachable - quarantined - coolingOff - weightExcluded)
|
||||
+ (total - atCap - unreachable - quarantined - coolingOff - modelOff - weightExcluded)
|
||||
+ " remaining");
|
||||
}
|
||||
}
|
||||
|
||||
@@ -690,14 +690,18 @@ public final class GitWorktrees implements Worktrees {
|
||||
* {@code fleet.seededSkillsNote}, readable with {@code git config --worktree --get-all
|
||||
* fleet.seededSkills}.
|
||||
*
|
||||
* <p><b>Claude Code specific by construction, not by a backend check here.</b> Only {@code
|
||||
* .claude/skills/<name>/SKILL.md} is a path any launcher reads today (opencode's equivalent is a
|
||||
* different shape under {@code .opencode/agent}, out of scope — see issue #362). This method
|
||||
* only copies files; like {@link #isolateToolSurface} — which neutralizes BOTH {@code .mcp.json}
|
||||
* and {@code opencode.json} unconditionally — it runs the same for every worktree regardless of
|
||||
* which backend ultimately spawns into it, because the backend is not yet chosen at {@link #add}
|
||||
* time. A seeded {@code .claude/skills/} directory in an opencode member's worktree is simply
|
||||
* never read by that launcher.
|
||||
* <p><b>Kind-blind by construction, not by a backend check here — this used to be a real gap
|
||||
* (fleetd #393).</b> This method only copies files; like {@link #isolateToolSurface} — which
|
||||
* neutralizes BOTH {@code .mcp.json} and {@code opencode.json} unconditionally — it runs the
|
||||
* same for every worktree regardless of which backend ultimately spawns into it, because no
|
||||
* caller of {@link #add} hands this class a kind to consult. Before fleetd #393, that made the
|
||||
* log line below a false claim of success for a {@code kind: opencode} member: opencode has no
|
||||
* built-in discovery of {@code .claude/skills/}, unlike the Claude Code CLI, so a seeded skill
|
||||
* never reached one. It now does — {@code OpenCodeLauncher#skillInstructionFiles} reads
|
||||
* whatever this method copied into {@code .claude/skills/} and appends each {@code SKILL.md} to
|
||||
* the generated {@code instructions[]} — but that delivery, and the log line that honestly
|
||||
* claims it (kind-aware, unlike this one), happens at the launcher, once the kind is actually
|
||||
* known, not here.
|
||||
*/
|
||||
private void seedSkills(String worktreePath) {
|
||||
if (memberSkillsSource == null) {
|
||||
@@ -739,7 +743,15 @@ public final class GitWorktrees implements Worktrees {
|
||||
if (detail.isEmpty()) {
|
||||
detail = "no skill folders found under " + source;
|
||||
}
|
||||
log.info("skill seeding: {} of {} candidate(s) from {} into {}/.claude/skills — {}",
|
||||
// fleetd #393: this only claims the copy step, deliberately — it cannot know the member
|
||||
// kind that will spawn into this worktree (see this method's own javadoc), so it must not
|
||||
// read as "and the member will act on it." Whether that is true depends on the kind: the
|
||||
// Claude Code CLI discovers .claude/skills/ on its own; OpenCodeLauncher logs its own
|
||||
// "skill delivery" line, once the kind is known, naming what it could and could not turn
|
||||
// into instructions[].
|
||||
log.info("skill seeding: {} of {} candidate(s) from {} into {}/.claude/skills — {} "
|
||||
+ "(whether the spawned member can act on this depends on its kind — see "
|
||||
+ "the launcher's own log for that)",
|
||||
seeded.size(), seeded.size() + kept.size(), source, worktreePath, detail);
|
||||
if (seeded.isEmpty()) {
|
||||
return;
|
||||
|
||||
@@ -10,6 +10,7 @@ import dev.ltms.fleet.peer.MemberRole;
|
||||
import dev.ltms.fleet.peer.PeerHandle;
|
||||
import dev.ltms.fleet.peer.PeerLauncher;
|
||||
import dev.ltms.fleet.peer.SpawnRequest;
|
||||
import dev.ltms.fleet.placement.PlacementDecision;
|
||||
import org.slf4j.Logger;
|
||||
import org.slf4j.LoggerFactory;
|
||||
|
||||
@@ -584,8 +585,65 @@ public final class SessionManager implements TurnListener {
|
||||
String ownerTerminal, WorktreeRequest wt,
|
||||
String sessionName, String resumeSessionId,
|
||||
MemberLifecycle.SlotReservation reservation) {
|
||||
String preResolvedProfile = (profile == null || profile.isBlank())
|
||||
? launcher.defaultProfile() : profile;
|
||||
// fleetd #425 rework (round 2): resolved through launcher.place(memberRole) — the same
|
||||
// candidate list, quarantine/cool-off/model-off filtering, and PlacementPolicy an unqualified
|
||||
// spawn of this role is actually judged against right now — never launcher.defaultProfile()
|
||||
// (MemberRole.DEV only, wrong for any other role) and never launcher.defaultProfileFor()
|
||||
// (the role's pool FIRST entry, blind to quarantine/cool-off/model-off: a first-round fix
|
||||
// used exactly this and regressed fleetd #429's "the fleet keeps working when a model is
|
||||
// turned off" guarantee — a quarantined or model-off pool-first profile made this throw
|
||||
// instead of routing around it, which an unqualified spawn is supposed to do). This same
|
||||
// resolved decision is reused below for repoRoot, parityOverlay, AND the spawn itself so the
|
||||
// worktree is always provisioned for the profile the member actually runs on — the two could
|
||||
// disagree before fleetd #425: this name picked repoRoot/overlay, but the spawn below passed
|
||||
// the ORIGINAL (blank) profile through to placement, which re-resolves live and can pick a
|
||||
// different profile if the pool changed between the two reads, or a genuinely different one
|
||||
// under weighted/round-robin placement.
|
||||
//
|
||||
// Round 1 of this rework fed the resolved name back into launcher.spawn(SpawnRequest) as an
|
||||
// EXPLICIT profile. That was a mistake this round corrects, and the mistake is not that the
|
||||
// two branches apply different checks — they are SUPPOSED to disagree: the routing branch a
|
||||
// blank spawn takes treats a quarantined/cooling-off/at-cap/model-off/unreachable/weight-0
|
||||
// profile as a reason to fall through to the next candidate, while CompositePeerLauncher's
|
||||
// THROWING branch (enforceNotQuarantined/enforceNotCoolingOff/enforceMaxLoad/
|
||||
// enforceModelEnabled) treats naming that same profile explicitly as a reason to refuse
|
||||
// outright. That is correct: an operator who names a profile should get a refusal, not a
|
||||
// silent substitution onto a different backend. The mistake was turning a fall-through into
|
||||
// a refusal by accident — resolving a name via the routing side and then re-entering the
|
||||
// refusing side with it, for a placement the routing side had already approved by walking
|
||||
// past everything else.
|
||||
//
|
||||
// Before fleetd #435, this accident was reachable through maxLoad specifically: the default
|
||||
// `fixed` placement policy did not evaluate maxLoad at all for automatic selection, so an
|
||||
// at-cap pool-first profile that placement itself would have picked for a plain unqualified
|
||||
// spawn could die at enforceMaxLoad one call later, purely because this method's route to
|
||||
// the spawn passed through an explicit profile name — a failure a worktree-less unqualified
|
||||
// spawn never hit. fleetd #435 closed that gap (`fixed` now evaluates maxLoad exactly like
|
||||
// every other placement policy), so that specific failure can no longer happen — a
|
||||
// PlacementDecision this method resolves can no longer be at-cap in the first place. What
|
||||
// this round's fix still buys, now that maxLoad can no longer cause the accident: it keeps
|
||||
// the PlacementDecision from place() and hands it to launcher.spawn(SpawnRequest,
|
||||
// PlacementDecision) for an unqualified request, which spawns through the SAME routing
|
||||
// branch a blank spawn uses — no enforce* check is newly applied, and the window between the
|
||||
// placement decision and the spawn (in which the pool, a config reload, or another spawn
|
||||
// landing on the same profile could otherwise move the state) never reopens. An
|
||||
// explicitly-named profile still goes through launcher.spawn(SpawnRequest) and its throwing
|
||||
// branch, unchanged — that caller asked for one profile by name and still gets everything
|
||||
// enforceNotQuarantined/enforceNotCoolingOff/enforceMaxLoad/enforceModelEnabled decide about
|
||||
// it, refusal included.
|
||||
//
|
||||
// The one cost that remains, unchanged from round 1: an unqualified worktree-provisioned
|
||||
// spawn does not get CompositePeerLauncher's cross-candidate retry on a live
|
||||
// PeerUnreachableException raised by the backend itself at spawn time (a transport-level
|
||||
// failure placement cannot see in advance) — spawn(req, decision) commits to the one profile
|
||||
// place() already chose, the same way an explicit-profile spawn commits to its one name. That
|
||||
// trade is deliberate: a worktree provisioned for the wrong backend (the #425 hazard) is worse
|
||||
// than a spawn that fails cleanly and can be retried by the caller. Nothing else is lost:
|
||||
// maxLoad, quarantine, cool-off and model-off all behave identically whether or not a
|
||||
// worktree was requested — that agreement is the invariant this rework exists to hold.
|
||||
boolean unqualifiedProfile = profile == null || profile.isBlank();
|
||||
PlacementDecision decision = unqualifiedProfile ? launcher.place(memberRole) : new PlacementDecision(profile);
|
||||
String preResolvedProfile = decision.profile();
|
||||
// CB-507: resolve through the launcher's CB-112 chain (requested → profile cwd → caller →
|
||||
// daemon cwd → "."), never the raw args. A plain REST spawn supplies neither a requested
|
||||
// nor a caller cwd, so taking the first non-blank of those two yielded null and put
|
||||
@@ -608,7 +666,15 @@ public final class SessionManager implements TurnListener {
|
||||
// copies more files into the worktree after add() returns, so sharing the group any earlier
|
||||
// leaves those overlay files operator-owned and read-only for a different-uid member.
|
||||
worktrees.shareWithGroup(repoRoot, path);
|
||||
handle = launcher.spawn(new SpawnRequest(profile, path, callerCwd, sessionName, resumeSessionId, memberRole));
|
||||
// fleetd #425: preResolvedProfile, not the original (possibly blank) profile — see the
|
||||
// comment above where it is resolved. The overlay/repoRoot above and the spawn here must
|
||||
// name the same profile. An unqualified request stays unqualified here and is honored via
|
||||
// the PlacementDecision already captured above (spawn(req, decision) — the routing branch,
|
||||
// no enforce* re-check); an explicitly-named profile still goes through the single-arg
|
||||
// spawn(req) and its throwing branch, exactly as before this rework.
|
||||
SpawnRequest spawnReq = new SpawnRequest(unqualifiedProfile ? null : preResolvedProfile,
|
||||
path, callerCwd, sessionName, resumeSessionId, memberRole);
|
||||
handle = unqualifiedProfile ? launcher.spawn(spawnReq, decision) : launcher.spawn(spawnReq);
|
||||
} catch (RuntimeException e) {
|
||||
log.warn("spawn failed for profile={} role={} branch={} path={}: {}",
|
||||
preResolvedProfile, memberRole, branch, path, e.getMessage());
|
||||
@@ -636,7 +702,7 @@ public final class SessionManager implements TurnListener {
|
||||
}
|
||||
throw e;
|
||||
}
|
||||
String resolvedProfile = resolveProfile(handle, profile);
|
||||
String resolvedProfile = resolveProfile(handle, preResolvedProfile);
|
||||
String cwd = launcher.effectiveCwd(new SpawnRequest(resolvedProfile, path, callerCwd));
|
||||
long now = nowNanos.getAsLong();
|
||||
// CB-619: see the no-worktree path above — bind before recording, and store the returned
|
||||
|
||||
@@ -0,0 +1,150 @@
|
||||
package dev.ltms.fleet;
|
||||
|
||||
import ch.qos.logback.classic.Level;
|
||||
import ch.qos.logback.classic.Logger;
|
||||
import ch.qos.logback.classic.spi.ILoggingEvent;
|
||||
import ch.qos.logback.core.read.ListAppender;
|
||||
import dev.ltms.fleet.config.FleetConfig;
|
||||
import org.junit.jupiter.api.Test;
|
||||
import org.junit.jupiter.api.io.TempDir;
|
||||
import org.slf4j.LoggerFactory;
|
||||
|
||||
import java.nio.file.Files;
|
||||
import java.nio.file.Path;
|
||||
import java.util.List;
|
||||
|
||||
import static org.junit.jupiter.api.Assertions.assertFalse;
|
||||
import static org.junit.jupiter.api.Assertions.assertTrue;
|
||||
|
||||
/**
|
||||
* fleetd #395: {@code exhaustedPattern} (see {@link FleetConfig.Profile#exhaustedPattern}) is
|
||||
* deliberately opt-in — an unset one leaves usage-limit detection silently OFF for that profile,
|
||||
* and nothing quarantines its credential. {@link Fleetd#reportExhaustedPatternGap} must say so at
|
||||
* startup, naming every unarmed profile, and must never fire when every profile is armed. Mirrors
|
||||
* {@link MemberCredentialsGapReportTest}'s pattern, capturing the real log via a
|
||||
* {@link ListAppender}.
|
||||
*
|
||||
* <p>The 8-profile shape in {@link #theLiveEightProfileShapeWarnsExactlyTheSixUnarmedProfiles} is
|
||||
* the live {@code fleetd.yaml} shape measured 2026-09-10 (fleetd #395's own ticket): 6 of 8
|
||||
* profiles unarmed, 2 of those 6 ({@code opus}, {@code sonnet}) running on the operator's Claude
|
||||
* subscription. {@code fleetd.yaml} itself is gitignored and unavailable to this test, so the
|
||||
* shape is reproduced as a throwaway config in a {@code @TempDir} rather than read off disk.
|
||||
*/
|
||||
class ExhaustedPatternGapReportTest {
|
||||
|
||||
private static FleetConfig load(Path dir, String yaml) throws Exception {
|
||||
Path f = dir.resolve("fleetd.yaml");
|
||||
Files.writeString(f, yaml);
|
||||
return FleetConfig.load(f);
|
||||
}
|
||||
|
||||
private static ListAppender<ILoggingEvent> attach() {
|
||||
Logger logger = (Logger) LoggerFactory.getLogger(Fleetd.class);
|
||||
ListAppender<ILoggingEvent> appender = new ListAppender<>();
|
||||
appender.start();
|
||||
logger.addAppender(appender);
|
||||
return appender;
|
||||
}
|
||||
|
||||
private static void detach(ListAppender<ILoggingEvent> appender) {
|
||||
((Logger) LoggerFactory.getLogger(Fleetd.class)).detachAppender(appender);
|
||||
}
|
||||
|
||||
/** Every profile name mentioned by a WARN-level log line, across every WARN this call produced. */
|
||||
private static List<String> warnMessages(ListAppender<ILoggingEvent> appender) {
|
||||
return appender.list.stream()
|
||||
.filter(e -> e.getLevel() == Level.WARN)
|
||||
.map(ILoggingEvent::getFormattedMessage)
|
||||
.toList();
|
||||
}
|
||||
|
||||
@Test
|
||||
void theLiveEightProfileShapeWarnsExactlyTheSixUnarmedProfiles(@TempDir Path dir) throws Exception {
|
||||
// Reproduces the live shape measured 2026-09-10: 8 profiles, 2 armed (sol, terra), 6
|
||||
// unarmed (local, local-direct, gx, opus, sonnet, xf) — 2 of the unarmed 6 (opus, sonnet)
|
||||
// are subscription: true.
|
||||
FleetConfig cfg = load(dir, """
|
||||
profiles:
|
||||
local:
|
||||
baseUrl: http://gx00.gw:8000
|
||||
local-direct:
|
||||
baseUrl: http://gx01.gw:8000
|
||||
gx:
|
||||
kind: opencode
|
||||
baseUrl: https://llm.ltms.dev/v1
|
||||
opus:
|
||||
subscription: true
|
||||
model: claude-opus-5
|
||||
sonnet:
|
||||
subscription: true
|
||||
model: claude-sonnet-5
|
||||
sol:
|
||||
baseUrl: https://llm.ltms.dev/v1
|
||||
exhaustedPattern: "The usage limit has been reached"
|
||||
terra:
|
||||
baseUrl: https://llm.ltms.dev/v1
|
||||
exhaustedPattern: "The usage limit has been reached"
|
||||
xf:
|
||||
baseUrl: https://llm.ltms.dev/v1
|
||||
""");
|
||||
|
||||
ListAppender<ILoggingEvent> appender = attach();
|
||||
try {
|
||||
Fleetd.reportExhaustedPatternGap(cfg);
|
||||
} finally {
|
||||
detach(appender);
|
||||
}
|
||||
|
||||
List<String> warns = warnMessages(appender);
|
||||
assertFalse(warns.isEmpty(), "6 of 8 profiles are unarmed — at least one WARN must fire");
|
||||
|
||||
// Exactly one WARN aggregates every unarmed profile, naming all 6 and none of the 2 armed.
|
||||
String aggregate = warns.stream()
|
||||
.filter(m -> m.contains("no exhaustedPattern configured"))
|
||||
.findFirst()
|
||||
.orElseThrow(() -> new AssertionError("expected an aggregate unarmed-profiles WARN: " + warns));
|
||||
for (String unarmed : List.of("local", "local-direct", "gx", "opus", "sonnet", "xf")) {
|
||||
assertTrue(aggregate.contains(unarmed), "aggregate WARN must name '" + unarmed + "': " + aggregate);
|
||||
}
|
||||
for (String armed : List.of("sol", "terra")) {
|
||||
assertFalse(aggregate.contains(armed), "aggregate WARN must NOT name armed profile '" + armed + "': " + aggregate);
|
||||
}
|
||||
|
||||
// A second, louder WARN calls out the subscription profiles specifically.
|
||||
String subscriptionWarn = warns.stream()
|
||||
.filter(m -> m.contains("metered Claude plan"))
|
||||
.findFirst()
|
||||
.orElseThrow(() -> new AssertionError("expected a subscription-specific WARN: " + warns));
|
||||
assertTrue(subscriptionWarn.contains("opus"), subscriptionWarn);
|
||||
assertTrue(subscriptionWarn.contains("sonnet"), subscriptionWarn);
|
||||
assertFalse(subscriptionWarn.contains("local-direct"),
|
||||
"the subscription WARN must not name a non-subscription profile: " + subscriptionWarn);
|
||||
}
|
||||
|
||||
@Test
|
||||
void everyProfileArmedProducesNoWarningAtAll(@TempDir Path dir) throws Exception {
|
||||
FleetConfig cfg = load(dir, """
|
||||
profiles:
|
||||
sol:
|
||||
baseUrl: https://llm.ltms.dev/v1
|
||||
exhaustedPattern: "The usage limit has been reached"
|
||||
terra:
|
||||
baseUrl: https://llm.ltms.dev/v1
|
||||
exhaustedPattern: "The usage limit has been reached"
|
||||
opus:
|
||||
subscription: true
|
||||
model: claude-opus-5
|
||||
exhaustedPattern: "5-hour limit reached"
|
||||
""");
|
||||
|
||||
ListAppender<ILoggingEvent> appender = attach();
|
||||
try {
|
||||
Fleetd.reportExhaustedPatternGap(cfg);
|
||||
} finally {
|
||||
detach(appender);
|
||||
}
|
||||
|
||||
assertTrue(warnMessages(appender).isEmpty(),
|
||||
"every profile is armed — a checker that warns anyway always fires: " + warnMessages(appender));
|
||||
}
|
||||
}
|
||||
@@ -20,6 +20,7 @@ import dev.ltms.fleet.peer.PeerLauncher;
|
||||
import dev.ltms.fleet.peer.SpawnRequest;
|
||||
import dev.ltms.fleet.placement.BackendOutagePolicy;
|
||||
import dev.ltms.fleet.placement.BackendQuarantine;
|
||||
import dev.ltms.fleet.placement.PlacementDecision;
|
||||
import dev.ltms.fleet.placement.PlacementPolicies;
|
||||
import dev.ltms.fleet.session.MemberSession;
|
||||
import dev.ltms.fleet.session.SessionManager;
|
||||
@@ -123,6 +124,11 @@ class FleetdBackendErrorSinkTest {
|
||||
throw new UnsupportedOperationException("not reachable — this test never acquires a session");
|
||||
}
|
||||
|
||||
@Override
|
||||
public PeerHandle spawn(SpawnRequest req, PlacementDecision decision) {
|
||||
throw new UnsupportedOperationException("not reachable — this test never acquires a session");
|
||||
}
|
||||
|
||||
@Override
|
||||
public Set<String> profiles() {
|
||||
return Set.of();
|
||||
|
||||
@@ -0,0 +1,155 @@
|
||||
package dev.ltms.fleet;
|
||||
|
||||
import dev.ltms.fleet.config.ConfigRef;
|
||||
import dev.ltms.fleet.config.FleetConfig;
|
||||
import dev.ltms.fleet.mcp.FleetMcp;
|
||||
import org.junit.jupiter.api.DisplayName;
|
||||
import org.junit.jupiter.api.Test;
|
||||
import org.junit.jupiter.api.io.TempDir;
|
||||
|
||||
import java.nio.file.Files;
|
||||
import java.nio.file.Path;
|
||||
|
||||
import static org.junit.jupiter.api.Assertions.assertFalse;
|
||||
import static org.junit.jupiter.api.Assertions.assertTrue;
|
||||
import static org.junit.jupiter.api.Assertions.assertEquals;
|
||||
|
||||
/**
|
||||
* fleetd #416: {@code fleet_list}'s {@code CapacitySource.configuredProfiles} must enumerate the
|
||||
* <em>startup</em> profile set, not the live, hot-reloaded one.
|
||||
*
|
||||
* <p>{@code profiles} as a whole is a {@code DEFERRED} key ({@link ConfigRef#DEFERRED_KEYS}):
|
||||
* {@code HerdrPeerLauncher} takes {@code Map.copyOf(profiles)} once at construction, so a profile
|
||||
* only added to the hot-reloaded map can never actually be spawned. Before this fix, {@code
|
||||
* Fleetd.main} built {@code CapacitySource} with {@code () -> config.get().profiles().keySet()} —
|
||||
* the live map — so {@code fleet_list} would report a freshly hot-reloaded profile as available
|
||||
* ({@code free > 0}) while {@code fleet_spawn} on that same profile failed with
|
||||
* {@code unknown worker profile}. Measured on another host: adding a throwaway profile and letting
|
||||
* it hot-reload gave {@code fleet_list} -> {@code free: 3} and {@code fleet_spawn} ->
|
||||
* {@code error: unknown worker profile}.
|
||||
*
|
||||
* <p>This test needs a reload, the same reason {@link FleetdExhaustionDetectionArmedWiringTest}
|
||||
* does: at startup the two snapshots agree, so a test of only a newly started daemon would not
|
||||
* detect a live {@code config.get()} lookup for the set.
|
||||
*
|
||||
* <p><b>maxLoad must stay hot.</b> It is read off {@code config.get()} exactly like
|
||||
* {@code credentialId} ({@link ConfigRef} documents both as hot, "read live off the config
|
||||
* supplier ... exactly like weight/maxLoad"), so a reload that only changes an existing profile's
|
||||
* {@code maxLoad} — no add/remove — must still change what {@code fleet_list} reports without a
|
||||
* restart. A fix that freezes the whole {@code CapacitySource} against {@code cfg} (rather than
|
||||
* only its {@code configuredProfiles} set) would trade this bug for its mirror image and is pinned
|
||||
* wrong by {@link #reloadedMaxLoadStillChangesWhatFleetListReports}.
|
||||
*/
|
||||
class FleetdCapacitySourceWiringTest {
|
||||
|
||||
private static final String STARTUP = """
|
||||
bind:
|
||||
host: 127.0.0.1
|
||||
port: 8765
|
||||
herdrSocket: ~/.config/herdr/herdr.sock
|
||||
profiles:
|
||||
terra:
|
||||
baseUrl: http://gx00.gw:8000
|
||||
model: terra
|
||||
maxLoad: 3
|
||||
guard:
|
||||
offSubscriptionHosts:
|
||||
- gx00.gw
|
||||
""";
|
||||
|
||||
private static final String WITH_NEW_PROFILE = """
|
||||
bind:
|
||||
host: 127.0.0.1
|
||||
port: 8765
|
||||
herdrSocket: ~/.config/herdr/herdr.sock
|
||||
profiles:
|
||||
terra:
|
||||
baseUrl: http://gx00.gw:8000
|
||||
model: terra
|
||||
maxLoad: 3
|
||||
ghost404:
|
||||
baseUrl: http://gx00.gw:8001
|
||||
model: ghost404
|
||||
maxLoad: 3
|
||||
guard:
|
||||
offSubscriptionHosts:
|
||||
- gx00.gw
|
||||
""";
|
||||
|
||||
private static final String WITH_CHANGED_MAX_LOAD = """
|
||||
bind:
|
||||
host: 127.0.0.1
|
||||
port: 8765
|
||||
herdrSocket: ~/.config/herdr/herdr.sock
|
||||
profiles:
|
||||
terra:
|
||||
baseUrl: http://gx00.gw:8000
|
||||
model: terra
|
||||
maxLoad: 9
|
||||
guard:
|
||||
offSubscriptionHosts:
|
||||
- gx00.gw
|
||||
""";
|
||||
|
||||
@Test
|
||||
@DisplayName("a profile present only in the live (hot-reloaded) config is NOT listed")
|
||||
void liveOnlyProfileIsNotListed(@TempDir Path dir) throws Exception {
|
||||
Path file = dir.resolve("fleetd.yaml");
|
||||
Files.writeString(file, STARTUP);
|
||||
FleetConfig cfg = FleetConfig.load(file);
|
||||
ConfigRef config = new ConfigRef(file, cfg);
|
||||
|
||||
Files.writeString(file, WITH_NEW_PROFILE);
|
||||
assertTrue(config.reload().applied());
|
||||
// The live snapshot now has the new profile — proves the reload really happened and this
|
||||
// test is not accidentally passing because nothing changed.
|
||||
assertTrue(config.get().profiles().containsKey("ghost404"));
|
||||
|
||||
FleetMcp.CapacitySource source = Fleetd.capacitySource(config, cfg, _ -> 0);
|
||||
|
||||
assertFalse(source.configuredProfiles().get().contains("ghost404"),
|
||||
"a profile added only to the hot-reloaded config must not be listed by fleet_list — "
|
||||
+ "HerdrPeerLauncher never learns about it until a restart, so fleet_spawn on it "
|
||||
+ "would fail with 'unknown worker profile' while fleet_list claimed it free");
|
||||
}
|
||||
|
||||
@Test
|
||||
@DisplayName("a profile present in the startup set IS listed")
|
||||
void startupProfileIsListed(@TempDir Path dir) throws Exception {
|
||||
Path file = dir.resolve("fleetd.yaml");
|
||||
Files.writeString(file, STARTUP);
|
||||
FleetConfig cfg = FleetConfig.load(file);
|
||||
ConfigRef config = new ConfigRef(file, cfg);
|
||||
|
||||
FleetMcp.CapacitySource source = Fleetd.capacitySource(config, cfg, _ -> 0);
|
||||
|
||||
// fleetd #416, both-directions requirement: a test that only ever passes an empty/absent
|
||||
// startup set (the case above) cannot tell a correct lookup from one that is permanently
|
||||
// empty (e.g. a mutation replacing the supplier with Set::of). This is the direction that
|
||||
// fails if the fix regresses to reporting nothing at all.
|
||||
assertTrue(source.configuredProfiles().get().contains("terra"),
|
||||
"a profile present in the startup snapshot must still be listed by fleet_list");
|
||||
}
|
||||
|
||||
@Test
|
||||
@DisplayName("a hot maxLoad edit still changes what fleet_list reports")
|
||||
void reloadedMaxLoadStillChangesWhatFleetListReports(@TempDir Path dir) throws Exception {
|
||||
Path file = dir.resolve("fleetd.yaml");
|
||||
Files.writeString(file, STARTUP);
|
||||
FleetConfig cfg = FleetConfig.load(file);
|
||||
ConfigRef config = new ConfigRef(file, cfg);
|
||||
|
||||
FleetMcp.CapacitySource source = Fleetd.capacitySource(config, cfg, _ -> 0);
|
||||
assertEquals(3, source.maxLoad().apply("terra"),
|
||||
"sanity: maxLoad reads 3 from the startup config before any reload");
|
||||
|
||||
Files.writeString(file, WITH_CHANGED_MAX_LOAD);
|
||||
assertTrue(config.reload().applied());
|
||||
|
||||
assertEquals(9, source.maxLoad().apply("terra"),
|
||||
"maxLoad must stay hot — the SAME CapacitySource instance must reflect a reloaded "
|
||||
+ "maxLoad without a restart, exactly like credentialId. Freezing the whole "
|
||||
+ "CapacitySource against the startup snapshot (rather than only its "
|
||||
+ "configuredProfiles set) would trade fleetd #416 for its mirror image.");
|
||||
}
|
||||
}
|
||||
@@ -0,0 +1,137 @@
|
||||
package dev.ltms.fleet;
|
||||
|
||||
import dev.ltms.fleet.config.ConfigRef;
|
||||
import dev.ltms.fleet.config.FleetConfig;
|
||||
import dev.ltms.fleet.inject.LiveExhaustedPatterns;
|
||||
import dev.ltms.fleet.mcp.FleetMcp;
|
||||
import dev.ltms.fleet.placement.BackendQuarantine;
|
||||
import org.junit.jupiter.api.DisplayName;
|
||||
import org.junit.jupiter.api.Test;
|
||||
import org.junit.jupiter.api.io.TempDir;
|
||||
|
||||
import java.nio.file.Files;
|
||||
import java.nio.file.Path;
|
||||
import java.util.Map;
|
||||
|
||||
import static org.junit.jupiter.api.Assertions.assertFalse;
|
||||
import static org.junit.jupiter.api.Assertions.assertTrue;
|
||||
|
||||
/**
|
||||
* fleetd #446: {@code exhaustionDetectionArmed} must describe the LIVE config, not a startup
|
||||
* snapshot — the opposite of what fleetd #404 asked for. #404 pinned {@code exhaustedPattern} as
|
||||
* compiled once at startup (matching {@code CompletionResolver}'s then-frozen behaviour); #446's
|
||||
* whole point is that both sides — the classification {@code CompletionResolver} enforces and the
|
||||
* {@code exhaustionDetectionArmed} field this reports — now read the SAME live {@link
|
||||
* LiveExhaustedPatterns} instance, so a reload arms or disarms detection with no restart, and the
|
||||
* report can never disagree with the behaviour (the fleetd #404 rule, now upheld for real instead
|
||||
* of by freezing both sides).
|
||||
*
|
||||
* <p>This test needs a reload. At startup the two snapshots agree, so a test of only a newly
|
||||
* started daemon would not detect a live {@code config.get()} lookup in the report field.
|
||||
*
|
||||
* <p><b>The hotness proof criterion 1 demands:</b> run this against the fixed code (green) and
|
||||
* against the pre-#446 compile site — {@code Fleetd.main}'s {@code exhaustedPatternsByProfile}
|
||||
* compiled once into a {@code Map<String, Pattern>} at startup, handed to {@code quarantineSource}
|
||||
* — and show it fails there. See the class doc history above: that shape is exactly what
|
||||
* {@code reloadedPatternArmsDetectionWithNoRestart} below is written to catch, and it is the test
|
||||
* that used to assert the opposite (see git history for this file's pre-#446 version, which
|
||||
* asserted {@code assertFalse(...)} on the identical scenario this now asserts {@code assertTrue}
|
||||
* on).
|
||||
*/
|
||||
class FleetdExhaustionDetectionArmedWiringTest {
|
||||
|
||||
private static final String NO_PATTERN = """
|
||||
bind:
|
||||
host: 127.0.0.1
|
||||
port: 8765
|
||||
herdrSocket: ~/.config/herdr/herdr.sock
|
||||
profiles:
|
||||
terra:
|
||||
baseUrl: http://gx00.gw:8000
|
||||
model: terra
|
||||
guard:
|
||||
offSubscriptionHosts:
|
||||
- gx00.gw
|
||||
""";
|
||||
|
||||
private static final String WITH_PATTERN = """
|
||||
bind:
|
||||
host: 127.0.0.1
|
||||
port: 8765
|
||||
herdrSocket: ~/.config/herdr/herdr.sock
|
||||
profiles:
|
||||
terra:
|
||||
baseUrl: http://gx00.gw:8000
|
||||
model: terra
|
||||
exhaustedPattern: "usage limit"
|
||||
guard:
|
||||
offSubscriptionHosts:
|
||||
- gx00.gw
|
||||
""";
|
||||
|
||||
@Test
|
||||
@DisplayName("reloading a profile's exhaustedPattern IN arms exhaustionDetectionArmed, no restart")
|
||||
void reloadedPatternArmsDetectionWithNoRestart(@TempDir Path dir) throws Exception {
|
||||
Path file = dir.resolve("fleetd.yaml");
|
||||
Files.writeString(file, NO_PATTERN);
|
||||
ConfigRef config = new ConfigRef(file, FleetConfig.load(file));
|
||||
|
||||
FleetMcp.QuarantineSource before = Fleetd.quarantineSource(config, BackendQuarantine.none(),
|
||||
new LiveExhaustedPatterns(() -> config.get().profiles()), Map.of());
|
||||
assertFalse(before.exhaustedPatternArmed().apply("terra"),
|
||||
"no exhaustedPattern configured yet — must report unarmed");
|
||||
|
||||
Files.writeString(file, WITH_PATTERN);
|
||||
assertTrue(config.reload().applied());
|
||||
assertTrue(config.get().profiles().get("terra").hasExhaustedPattern());
|
||||
|
||||
// Re-read exhaustedPatternArmed off the SAME QuarantineSource built BEFORE the reload — no
|
||||
// new object, no new wiring — to prove the field itself is live, not merely that a freshly
|
||||
// built source would be. This is the exact axis the pre-#446 code fails on: the same
|
||||
// Function<String,Boolean> lambda, called again after a reload, sees the new answer only if
|
||||
// it re-reads config.get() on every call rather than a value captured earlier.
|
||||
assertTrue(before.exhaustedPatternArmed().apply("terra"),
|
||||
"exhaustionDetectionArmed must read the LIVE config on every call — fleetd #446 made "
|
||||
+ "this hot; it must reflect a reload with no restart and no new QuarantineSource");
|
||||
}
|
||||
|
||||
@Test
|
||||
@DisplayName("reloading exhaustedPattern OUT disarms exhaustionDetectionArmed, no restart")
|
||||
void reloadedPatternRemovalDisarmsDetectionWithNoRestart(@TempDir Path dir) throws Exception {
|
||||
Path file = dir.resolve("fleetd.yaml");
|
||||
Files.writeString(file, WITH_PATTERN);
|
||||
ConfigRef config = new ConfigRef(file, FleetConfig.load(file));
|
||||
|
||||
FleetMcp.QuarantineSource source = Fleetd.quarantineSource(config, BackendQuarantine.none(),
|
||||
new LiveExhaustedPatterns(() -> config.get().profiles()), Map.of());
|
||||
assertTrue(source.exhaustedPatternArmed().apply("terra"),
|
||||
"exhaustedPattern is configured from the start — must report armed");
|
||||
|
||||
Files.writeString(file, NO_PATTERN);
|
||||
assertTrue(config.reload().applied());
|
||||
|
||||
assertFalse(source.exhaustedPatternArmed().apply("terra"),
|
||||
"removing exhaustedPattern and reloading must disarm detection with no restart — an "
|
||||
+ "operator turning detection off (e.g. while debugging a false positive) "
|
||||
+ "must see that reflected immediately, exactly like arming it is");
|
||||
}
|
||||
|
||||
@Test
|
||||
@DisplayName("a profile with no exhaustedPattern configured is reported unarmed")
|
||||
void aProfileWithNoPatternIsUnarmed(@TempDir Path dir) throws Exception {
|
||||
// fleetd #404's second direction, carried forward: this test alone fails if
|
||||
// exhaustedPatternArmed degenerates to a constant `profile -> true` — it needs at least one
|
||||
// profile that is genuinely unarmed to catch that.
|
||||
Path file = dir.resolve("fleetd.yaml");
|
||||
Files.writeString(file, WITH_PATTERN);
|
||||
ConfigRef config = new ConfigRef(file, FleetConfig.load(file));
|
||||
|
||||
FleetMcp.QuarantineSource source = Fleetd.quarantineSource(config, BackendQuarantine.none(),
|
||||
new LiveExhaustedPatterns(() -> config.get().profiles()), Map.of());
|
||||
|
||||
assertTrue(source.exhaustedPatternArmed().apply("terra"),
|
||||
"a profile whose exhaustedPattern is configured must report armed");
|
||||
assertFalse(source.exhaustedPatternArmed().apply("sonnet"),
|
||||
"a profile absent from config entirely must not report armed");
|
||||
}
|
||||
}
|
||||
@@ -0,0 +1,175 @@
|
||||
package dev.ltms.fleet;
|
||||
|
||||
import ch.qos.logback.classic.Level;
|
||||
import ch.qos.logback.classic.Logger;
|
||||
import ch.qos.logback.classic.spi.ILoggingEvent;
|
||||
import ch.qos.logback.core.read.ListAppender;
|
||||
import dev.ltms.fleet.config.ConfigRef;
|
||||
import dev.ltms.fleet.config.FleetConfig;
|
||||
import dev.ltms.fleet.guard.SubscriptionGuard;
|
||||
import dev.ltms.fleet.herdr.AgentControl;
|
||||
import dev.ltms.fleet.herdr.FakeHerdr;
|
||||
import dev.ltms.fleet.herdr.WorkspaceControl;
|
||||
import dev.ltms.fleet.inject.ExhaustionSink;
|
||||
import dev.ltms.fleet.member.ClaudeCodeLauncher;
|
||||
import dev.ltms.fleet.placement.BackendQuarantine;
|
||||
import dev.ltms.fleet.session.SessionManager;
|
||||
import org.junit.jupiter.api.DisplayName;
|
||||
import org.junit.jupiter.api.Test;
|
||||
import org.junit.jupiter.api.io.TempDir;
|
||||
import org.slf4j.LoggerFactory;
|
||||
|
||||
import java.nio.file.Files;
|
||||
import java.nio.file.Path;
|
||||
import java.util.HashMap;
|
||||
import java.util.List;
|
||||
import java.util.Map;
|
||||
import java.util.Set;
|
||||
import java.util.concurrent.TimeUnit;
|
||||
|
||||
import static org.junit.jupiter.api.Assertions.assertFalse;
|
||||
import static org.junit.jupiter.api.Assertions.assertTrue;
|
||||
|
||||
/**
|
||||
* fleetd #446 follow-up (round 3): round 2's {@link FleetdUsageLimitFixWarningTest} calls {@link
|
||||
* Fleetd#usageLimitFixWarning}/{@link Fleetd#usageLimitFixWarningNoModel} directly, which pins
|
||||
* the two methods' TEXT but cannot see whether {@link Fleetd#exhaustionSink}'s {@code log.warn}
|
||||
* call actually invokes either one, or invokes the right one. A mutation battery run against
|
||||
* round 2's merge (1592 green tests) proved both gaps real:
|
||||
*
|
||||
* <ul>
|
||||
* <li><b>Cell A</b> — replacing the whole ternary result with a literal string
|
||||
* ({@code "usage-limit fix: MUTANT"}) left every test green;</li>
|
||||
* <li><b>Cell B</b> — swapping the ternary's two branches, so a profile WITH a {@code model:}
|
||||
* gets the no-model fallback message and vice versa, was measured unmeasured by the lead
|
||||
* but the existing test file shows no call that would notice it either.</li>
|
||||
* </ul>
|
||||
*
|
||||
* <p>This class drives {@link Fleetd#exhaustionSink} — the actual factory {@code main} wires,
|
||||
* since round 3's extraction pulled it out of the inline lambda for exactly this reason — and
|
||||
* asserts on the REAL text a {@link ListAppender} attached to {@code Fleetd}'s own logger
|
||||
* captures, mirroring {@code ExhaustedPatternGapReportTest}'s idiom. Chosen over a source-text
|
||||
* assertion (the {@code FleetMcpAuthzTest} {@code everyRegisteredToolHasItsHandlerActionPinned}
|
||||
* idiom) because that route can prove Cell A (the ternary is still there, still built from the
|
||||
* two named methods) but not Cell B (which branch a given profile actually reaches) — this route
|
||||
* answers both from one mechanism, since it reads what the sink actually logged for each shape of
|
||||
* profile.
|
||||
*
|
||||
* <p>Each test filters for the WARN line starting {@code "usage-limit fix:"} specifically — {@link
|
||||
* Fleetd#exhaustionSink} also logs a separate {@code "credential '...' quarantined for..."} WARN
|
||||
* on every call, and asserting against the wrong one would pass or fail for the wrong reason. That
|
||||
* prefix survives even under the M5 mutation above (the mutant keeps the tag, only the body
|
||||
* becomes {@code "MUTANT"}), so the filter itself is not what either mutation defeats.
|
||||
*/
|
||||
class FleetdExhaustionSinkWarningTest {
|
||||
|
||||
private static final String YAML = """
|
||||
bind:
|
||||
host: 127.0.0.1
|
||||
port: 8765
|
||||
herdrSocket: ~/.config/herdr/herdr.sock
|
||||
profiles:
|
||||
terra:
|
||||
baseUrl: http://gx00.gw:8000
|
||||
model: claude-opus-5
|
||||
gx:
|
||||
baseUrl: http://gx00.gw:8000
|
||||
guard:
|
||||
offSubscriptionHosts:
|
||||
- gx00.gw
|
||||
""";
|
||||
|
||||
private static SessionManager emptyRosterSessions() {
|
||||
FakeHerdr h = new FakeHerdr();
|
||||
FleetConfig.Profile dummy = new FleetConfig.Profile(
|
||||
"dummy", "http://gx00.gw:8000", "coder", null, "FLEETD_WORKER_TOKEN", null,
|
||||
"tab", "fleetd-workers", "worker: {profile} #{n}", null, null, null);
|
||||
ClaudeCodeLauncher launcher = new ClaudeCodeLauncher(new AgentControl(h), new WorkspaceControl(h),
|
||||
new SubscriptionGuard(Set.of("gx00.gw")), Map.of(dummy.profile(), dummy), dummy.profile(), _ -> "tok");
|
||||
// Never acquires a session — exhaustionSink's roster lookup is expected to miss and fall
|
||||
// back to the profileHint argument, exactly like OpenCodeLauncher's real call site does
|
||||
// (fleetd #234) — so the sink under test never needs a populated roster.
|
||||
return new SessionManager(launcher);
|
||||
}
|
||||
|
||||
private static ListAppender<ILoggingEvent> attach() {
|
||||
Logger logger = (Logger) LoggerFactory.getLogger(Fleetd.class);
|
||||
ListAppender<ILoggingEvent> appender = new ListAppender<>();
|
||||
appender.start();
|
||||
logger.addAppender(appender);
|
||||
return appender;
|
||||
}
|
||||
|
||||
private static void detach(ListAppender<ILoggingEvent> appender) {
|
||||
((Logger) LoggerFactory.getLogger(Fleetd.class)).detachAppender(appender);
|
||||
}
|
||||
|
||||
/** The one WARN line {@code exhaustionSink} builds from the ternary under test, or {@code null}. */
|
||||
private static String fixWarning(ListAppender<ILoggingEvent> appender) {
|
||||
List<String> warns = appender.list.stream()
|
||||
.filter(e -> e.getLevel() == Level.WARN)
|
||||
.map(ILoggingEvent::getFormattedMessage)
|
||||
.filter(m -> m.startsWith("usage-limit fix:"))
|
||||
.toList();
|
||||
assertTrue(warns.size() == 1,
|
||||
"expected exactly one 'usage-limit fix:' WARN per onExhausted call, got " + warns.size()
|
||||
+ ": " + warns);
|
||||
return warns.getFirst();
|
||||
}
|
||||
|
||||
@Test
|
||||
@DisplayName("a profile WITH a configured model gets the actionable models.allow fix, not the fallback")
|
||||
void profileWithModelGetsTheActionableFix(@TempDir Path dir) throws Exception {
|
||||
Path file = dir.resolve("fleetd.yaml");
|
||||
Files.writeString(file, YAML);
|
||||
FleetConfig cfg = FleetConfig.load(file);
|
||||
ConfigRef config = ConfigRef.fixed(cfg);
|
||||
BackendQuarantine quarantine = new BackendQuarantine(() -> 0L, TimeUnit.MINUTES.toNanos(30));
|
||||
Map<String, String> reasonByCredential = new HashMap<>();
|
||||
ExhaustionSink sink = Fleetd.exhaustionSink(emptyRosterSessions(), config, quarantine,
|
||||
reasonByCredential, cfg);
|
||||
|
||||
ListAppender<ILoggingEvent> appender = attach();
|
||||
try {
|
||||
sink.onExhausted("term_x", "The usage limit has been reached", "terra");
|
||||
} finally {
|
||||
detach(appender);
|
||||
}
|
||||
|
||||
String warning = fixWarning(appender);
|
||||
assertTrue(warning.contains("terra"), "must name the profile: " + warning);
|
||||
assertTrue(warning.contains("claude-opus-5"), "must name the model: " + warning);
|
||||
assertTrue(warning.contains("enabled: false"), "must name the action: " + warning);
|
||||
assertTrue(warning.contains("models.allow"), "must name where the action goes: " + warning);
|
||||
assertFalse(warning.contains("has no model: configured"),
|
||||
"a profile WITH a model must not get the no-model fallback text: " + warning);
|
||||
}
|
||||
|
||||
@Test
|
||||
@DisplayName("a profile with NO configured model gets the fallback, never a fix that does not exist")
|
||||
void profileWithNoModelGetsTheFallbackNotAFix(@TempDir Path dir) throws Exception {
|
||||
Path file = dir.resolve("fleetd.yaml");
|
||||
Files.writeString(file, YAML);
|
||||
FleetConfig cfg = FleetConfig.load(file);
|
||||
ConfigRef config = ConfigRef.fixed(cfg);
|
||||
BackendQuarantine quarantine = new BackendQuarantine(() -> 0L, TimeUnit.MINUTES.toNanos(30));
|
||||
Map<String, String> reasonByCredential = new HashMap<>();
|
||||
ExhaustionSink sink = Fleetd.exhaustionSink(emptyRosterSessions(), config, quarantine,
|
||||
reasonByCredential, cfg);
|
||||
|
||||
ListAppender<ILoggingEvent> appender = attach();
|
||||
try {
|
||||
sink.onExhausted("term_y", "The usage limit has been reached", "gx");
|
||||
} finally {
|
||||
detach(appender);
|
||||
}
|
||||
|
||||
String warning = fixWarning(appender);
|
||||
assertTrue(warning.contains("gx"), "must name the profile: " + warning);
|
||||
assertTrue(warning.contains("has no model: configured"),
|
||||
"must say there is no model to gate by name: " + warning);
|
||||
assertFalse(warning.contains("enabled: false"),
|
||||
"a profile with no model: configured has no models.allow entry to flip — must "
|
||||
+ "not claim one exists: " + warning);
|
||||
}
|
||||
}
|
||||
@@ -0,0 +1,62 @@
|
||||
package dev.ltms.fleet;
|
||||
|
||||
import dev.ltms.fleet.peer.PeerLauncher;
|
||||
import org.junit.jupiter.api.DisplayName;
|
||||
import org.junit.jupiter.api.Test;
|
||||
|
||||
import java.util.Set;
|
||||
|
||||
import static org.junit.jupiter.api.Assertions.assertEquals;
|
||||
import static org.junit.jupiter.api.Assertions.assertNotEquals;
|
||||
|
||||
/**
|
||||
* fleetd #422 follow-up: {@code Fleetd.modelGateCoverageLine} is the startup-log counterpart of
|
||||
* {@code exhaustedPatternCoverageLine}/{@code errorPatternCoverageLine} — see {@code
|
||||
* FleetdPatternCoverageLineTest} for the identical shape this follows — except here there is no
|
||||
* {@code UnsetMeaning} choice for a caller to get backwards: {@link
|
||||
* PeerLauncher.ModelGateState#configured()} already states, unambiguously, whether an empty
|
||||
* {@link PeerLauncher.ModelGateState#off()} means "no {@code models:} block to gate with at all"
|
||||
* or "a block armed and currently reporting zero off". This class proves {@code
|
||||
* modelGateCoverageLine} words those two states — plus the third, N off — distinctly, so a
|
||||
* mutation that made it ignore {@code configured()} either way is caught here.
|
||||
*/
|
||||
class FleetdModelGateCoverageLineTest {
|
||||
|
||||
@Test
|
||||
@DisplayName("no models: block reports not configured, distinct from armed-with-zero")
|
||||
void noModelsBlockReportsNotConfigured() {
|
||||
String line = Fleetd.modelGateCoverageLine(PeerLauncher.ModelGateState.notConfigured());
|
||||
assertEquals("not configured (no models: block — nothing is gated, and nothing can be)", line);
|
||||
}
|
||||
|
||||
@Test
|
||||
@DisplayName("a models: block armed with nothing off reports armed, distinct from not configured")
|
||||
void armedWithNothingOffReportsArmed() {
|
||||
String line = Fleetd.modelGateCoverageLine(PeerLauncher.ModelGateState.armed(Set.of()));
|
||||
assertEquals("armed (models: block present; 0 models currently turned off)", line);
|
||||
}
|
||||
|
||||
@Test
|
||||
@DisplayName("a models: block with N off names the off models")
|
||||
void armedWithModelsOffNamesThem() {
|
||||
String line = Fleetd.modelGateCoverageLine(
|
||||
PeerLauncher.ModelGateState.armed(Set.of("deepseek-v4-flash", "claude-opus-9000")));
|
||||
assertEquals("armed (2 model(s) turned off: [claude-opus-9000, deepseek-v4-flash])", line);
|
||||
}
|
||||
|
||||
@Test
|
||||
@DisplayName("the three states produce pairwise-distinct wording for the same empty-looking input")
|
||||
void theThreeStatesProduceDistinctWording() {
|
||||
String notConfigured = Fleetd.modelGateCoverageLine(PeerLauncher.ModelGateState.notConfigured());
|
||||
String armedZero = Fleetd.modelGateCoverageLine(PeerLauncher.ModelGateState.armed(Set.of()));
|
||||
String armedOne = Fleetd.modelGateCoverageLine(PeerLauncher.ModelGateState.armed(Set.of("x")));
|
||||
|
||||
// Pinned individually above; restated here so this test alone still catches a regression
|
||||
// even if one of the three tests above were ever deleted — the exact FleetdPatternCoverageLineTest
|
||||
// pattern, adapted from "two keys" to "three states of one gate".
|
||||
assertNotEquals(notConfigured, armedZero,
|
||||
"collapsing 'no models: block' into 'armed, zero off' is fleetd #422 follow-up's exact defect");
|
||||
assertNotEquals(armedZero, armedOne);
|
||||
assertNotEquals(notConfigured, armedOne);
|
||||
}
|
||||
}
|
||||
@@ -0,0 +1,67 @@
|
||||
package dev.ltms.fleet;
|
||||
|
||||
import org.junit.jupiter.api.DisplayName;
|
||||
import org.junit.jupiter.api.Test;
|
||||
|
||||
import java.util.Set;
|
||||
|
||||
import static org.junit.jupiter.api.Assertions.assertEquals;
|
||||
import static org.junit.jupiter.api.Assertions.assertNotEquals;
|
||||
|
||||
/**
|
||||
* fleetd #415 (review follow-up): {@code CompletionResolverTest} proves {@code coverage()} words
|
||||
* {@code UnsetMeaning.OFF} and {@code UnsetMeaning.BUILT_IN_DEFAULT} correctly — but every one of
|
||||
* those tests supplies the meaning itself. That proves the enum's wording, never that {@code
|
||||
* Fleetd} pairs the right meaning with the right pattern key. That pairing is #415's actual
|
||||
* defect: {@code coverage()} had no way to know what unset meant for its key, so the fix moved
|
||||
* the fact to the caller — and nothing yet proved the caller states it correctly.
|
||||
*
|
||||
* <p><b>Measured or it didn't happen:</b> swapping the two {@code UnsetMeaning} arguments at
|
||||
* {@code Fleetd}'s two coverage call sites — giving {@code exhaustedPattern} the built-in-default
|
||||
* wording and {@code errorPattern} the off wording, #415's exact defect with the keys exchanged —
|
||||
* compiled with 0 errors and left all 1506 existing tests green. This class exists to turn that
|
||||
* swap red.
|
||||
*
|
||||
* <p>It calls {@link Fleetd#exhaustedPatternCoverageLine} and {@link Fleetd#errorPatternCoverageLine}
|
||||
* directly rather than reading {@code Fleetd.java} as source text (the shape {@code
|
||||
* FleetdCompletionResolverWiringTest} uses for a different wiring gap): those two methods are the
|
||||
* extracted call sites {@code main} actually invokes, following the same {@code static} factory +
|
||||
* dedicated-test pattern as {@link Fleetd#capacitySource} and {@link Fleetd#worktreeBranchLookup}.
|
||||
*/
|
||||
class FleetdPatternCoverageLineTest {
|
||||
|
||||
private static final Set<String> PROFILES = Set.of("terra", "gx10");
|
||||
|
||||
@Test
|
||||
@DisplayName("exhaustedPatternCoverageLine says off when no profile configures exhaustedPattern")
|
||||
void exhaustedPatternCoverageLineSaysOffWhenNoProfileConfiguresIt() {
|
||||
assertEquals("off (no profile has an exhaustedPattern configured; profiles: [gx10, terra])",
|
||||
Fleetd.exhaustedPatternCoverageLine(PROFILES, Set.of()));
|
||||
}
|
||||
|
||||
@Test
|
||||
@DisplayName("errorPatternCoverageLine says built-in default when no profile configures errorPattern")
|
||||
void errorPatternCoverageLineSaysBuiltInDefaultWhenNoProfileConfiguresIt() {
|
||||
assertEquals("built-in default for all profiles (no profile customises errorPattern; "
|
||||
+ "profiles: [gx10, terra])",
|
||||
Fleetd.errorPatternCoverageLine(PROFILES, Set.of()));
|
||||
}
|
||||
|
||||
@Test
|
||||
@DisplayName("the two keys produce different wording for the identical empty-coverage input")
|
||||
void theTwoKeysProduceDifferentWordingForTheSameEmptyInput() {
|
||||
String exhaustedLine = Fleetd.exhaustedPatternCoverageLine(PROFILES, Set.of());
|
||||
String errorLine = Fleetd.errorPatternCoverageLine(PROFILES, Set.of());
|
||||
|
||||
// Pinned individually above; restated here so this test alone still catches a swap even
|
||||
// if one of the two tests above were ever deleted.
|
||||
assertEquals("off (no profile has an exhaustedPattern configured; profiles: [gx10, terra])",
|
||||
exhaustedLine);
|
||||
assertEquals("built-in default for all profiles (no profile customises errorPattern; "
|
||||
+ "profiles: [gx10, terra])", errorLine);
|
||||
assertNotEquals(exhaustedLine, errorLine,
|
||||
"swapping which UnsetMeaning pairs with which pattern key at Fleetd's call sites "
|
||||
+ "must be caught here — that pairing, not coverage()'s own wording in isolation, "
|
||||
+ "is fleetd #415's actual defect");
|
||||
}
|
||||
}
|
||||
@@ -0,0 +1,79 @@
|
||||
package dev.ltms.fleet;
|
||||
|
||||
import ch.qos.logback.classic.Logger;
|
||||
import ch.qos.logback.classic.Level;
|
||||
import ch.qos.logback.classic.spi.ILoggingEvent;
|
||||
import ch.qos.logback.core.read.ListAppender;
|
||||
import dev.ltms.fleet.config.FleetConfig;
|
||||
import org.junit.jupiter.api.Test;
|
||||
import org.junit.jupiter.api.io.TempDir;
|
||||
import org.slf4j.LoggerFactory;
|
||||
|
||||
import java.nio.file.Files;
|
||||
import java.nio.file.Path;
|
||||
|
||||
import static org.junit.jupiter.api.Assertions.assertThrows;
|
||||
import static org.junit.jupiter.api.Assertions.assertTrue;
|
||||
|
||||
/**
|
||||
* Proves {@link Fleetd#main(String[])} calls every startup report before validation aborts startup.
|
||||
* The invalid non-loopback bind makes {@link FleetConfig#validateAll()} throw before {@code main}
|
||||
* can open the herdr socket or bind a port. The fixture also triggers every report, so removing any
|
||||
* one call from {@code main} leaves its expected log line absent.
|
||||
*/
|
||||
class FleetdStartupReportTest {
|
||||
|
||||
private static Level originalLevel;
|
||||
|
||||
private static ListAppender<ILoggingEvent> attach() {
|
||||
Logger logger = (Logger) LoggerFactory.getLogger(Fleetd.class);
|
||||
originalLevel = logger.getLevel();
|
||||
logger.setLevel(Level.INFO);
|
||||
ListAppender<ILoggingEvent> appender = new ListAppender<>();
|
||||
appender.start();
|
||||
logger.addAppender(appender);
|
||||
return appender;
|
||||
}
|
||||
|
||||
private static void detach(ListAppender<ILoggingEvent> appender) {
|
||||
Logger logger = (Logger) LoggerFactory.getLogger(Fleetd.class);
|
||||
logger.detachAppender(appender);
|
||||
logger.setLevel(originalLevel);
|
||||
}
|
||||
|
||||
private static boolean contains(ListAppender<ILoggingEvent> appender, String fragment) {
|
||||
return appender.list.stream()
|
||||
.map(ILoggingEvent::getFormattedMessage)
|
||||
.anyMatch(message -> message.contains(fragment));
|
||||
}
|
||||
|
||||
@Test
|
||||
void mainReportsEveryStartupGapBeforeValidationAborts(@TempDir Path dir) throws Exception {
|
||||
Path config = dir.resolve("fleetd.yaml");
|
||||
Files.writeString(config, """
|
||||
bind:
|
||||
host: 0.0.0.0
|
||||
port: 8765
|
||||
profiles:
|
||||
worker:
|
||||
baseUrl: https://llm.ltms.dev/v1
|
||||
gitTokenEnv: GITEA_TOKEN
|
||||
""");
|
||||
|
||||
ListAppender<ILoggingEvent> appender = attach();
|
||||
try {
|
||||
assertThrows(IllegalStateException.class, () -> Fleetd.main(new String[]{config.toString()}));
|
||||
} finally {
|
||||
detach(appender);
|
||||
}
|
||||
|
||||
assertTrue(contains(appender, "startup git host GITEA_HOST:"),
|
||||
"Fleetd.main must report the git host shape");
|
||||
assertTrue(contains(appender, "member trust model: members run as the same OS user"),
|
||||
"Fleetd.main must report the member trust model");
|
||||
assertTrue(contains(appender, "memberCredentials: absent or empty"),
|
||||
"Fleetd.main must report an absent memberCredentials policy");
|
||||
assertTrue(contains(appender, "exhaustedPattern: profile(s) [worker] have no exhaustedPattern configured"),
|
||||
"Fleetd.main must report profiles without exhaustedPattern");
|
||||
}
|
||||
}
|
||||
@@ -0,0 +1,132 @@
|
||||
package dev.ltms.fleet;
|
||||
|
||||
import org.junit.jupiter.api.Test;
|
||||
import org.junit.jupiter.api.io.TempDir;
|
||||
|
||||
import java.nio.file.Files;
|
||||
import java.nio.file.Path;
|
||||
|
||||
import static org.junit.jupiter.api.Assertions.assertThrows;
|
||||
import static org.junit.jupiter.api.Assertions.assertTrue;
|
||||
|
||||
/**
|
||||
* Proves the one thing no test proved before this ticket's follow-up: that {@code Fleetd.main}
|
||||
* ITSELF — not a copy of its logic, not the validator called directly — still refuses to start on
|
||||
* a bad config. Mutation testing found that deleting {@code cfg.validateAll();} (née six
|
||||
* individual {@code cfg.validateXxx();} calls) from {@code Fleetd.main} left the full 1478-test
|
||||
* suite green; every existing test called a validator directly and none exercised {@code
|
||||
* Fleetd.main} as the caller. See {@code FleetConfigValidateAllTest} for why the fix collapses
|
||||
* those six calls into one reflective {@link FleetConfig#validateAll()}, and {@code
|
||||
* ConfigRefTest} for the equivalent proof on the {@link
|
||||
* dev.ltms.fleet.config.ConfigRef#reload()} path.
|
||||
*
|
||||
* <p>This is deliberately the real {@code static void main(String[] args)} — package-private, so
|
||||
* only a test in this package can call it, which is exactly what makes this proof strong: it is
|
||||
* not a helper extracted for testability, it is the literal method {@code java -jar fleetd.jar}
|
||||
* invokes. Every fixture below is otherwise-valid and fails exactly one validator, and — because
|
||||
* {@link FleetConfig#validateAll()} runs immediately after {@code SubscriptionGuard.
|
||||
* assertPrimaryClean}, before {@code main} opens the herdr socket, binds Javalin, or touches
|
||||
* anything else with a real side effect — calling {@code Fleetd.main} with one of these configs is
|
||||
* safe: it is guaranteed to throw before reaching any of that, precisely because the config is
|
||||
* deliberately invalid.
|
||||
*/
|
||||
class FleetdStartupValidationTest {
|
||||
|
||||
private static void assertMainRefuses(Path dir, String fileName, String yaml,
|
||||
String mustContain) throws Exception {
|
||||
Path f = dir.resolve(fileName);
|
||||
Files.writeString(f, yaml);
|
||||
IllegalStateException e = assertThrows(IllegalStateException.class,
|
||||
() -> Fleetd.main(new String[]{f.toString()}),
|
||||
fileName + ": Fleetd.main must refuse this config before doing anything else");
|
||||
assertTrue(e.getMessage().contains(mustContain),
|
||||
fileName + ": expected message to contain \"" + mustContain + "\" but was: "
|
||||
+ e.getMessage());
|
||||
}
|
||||
|
||||
@Test
|
||||
void mainRefusesANonLoopbackBindWithoutTokenMode(@TempDir Path dir) throws Exception {
|
||||
assertMainRefuses(dir, "auth-exposure.yaml", """
|
||||
bind:
|
||||
host: 0.0.0.0
|
||||
port: 8765
|
||||
""", "auth.mode: token");
|
||||
}
|
||||
|
||||
@Test
|
||||
void mainRefusesALeadTabPrefixCollision(@TempDir Path dir) throws Exception {
|
||||
assertMainRefuses(dir, "lead-tab-prefixes.yaml", """
|
||||
bind:
|
||||
host: 127.0.0.1
|
||||
port: 8765
|
||||
fleet:
|
||||
tabLabel: "lead: {role} {profile}"
|
||||
leaders:
|
||||
opus:
|
||||
tab: "lead: opus"
|
||||
""", "fleet.tabLabel");
|
||||
}
|
||||
|
||||
@Test
|
||||
void mainRefusesASubscriptionProfileThatReseatsAnthropicBaseUrl(@TempDir Path dir)
|
||||
throws Exception {
|
||||
assertMainRefuses(dir, "subscription-profiles.yaml", """
|
||||
bind:
|
||||
host: 127.0.0.1
|
||||
port: 8765
|
||||
profiles:
|
||||
sonnet:
|
||||
subscription: true
|
||||
argv: ["ccs", "sonnet"]
|
||||
env:
|
||||
ANTHROPIC_BASE_URL: http://anything-not-on-the-allowlist
|
||||
""", "ANTHROPIC_BASE_URL");
|
||||
}
|
||||
|
||||
@Test
|
||||
void mainRefusesAnUnknownCharterKey(@TempDir Path dir) throws Exception {
|
||||
assertMainRefuses(dir, "charters.yaml", """
|
||||
bind:
|
||||
host: 127.0.0.1
|
||||
port: 8765
|
||||
fleet:
|
||||
charters:
|
||||
architetc: text
|
||||
""", "architetc");
|
||||
}
|
||||
|
||||
@Test
|
||||
void mainRefusesAnArchitectSlotNamingAnUnconfiguredProfile(@TempDir Path dir) throws Exception {
|
||||
assertMainRefuses(dir, "members.yaml", """
|
||||
bind:
|
||||
host: 127.0.0.1
|
||||
port: 8765
|
||||
profiles:
|
||||
gx10:
|
||||
baseUrl: http://gx10.gw:8000
|
||||
fleet:
|
||||
architects:
|
||||
lead-designer:
|
||||
profile: sonnet
|
||||
""", "lead-designer");
|
||||
}
|
||||
|
||||
@Test
|
||||
void mainRefusesAProfileNamingAModelOutsideTheAllowList(@TempDir Path dir) throws Exception {
|
||||
assertMainRefuses(dir, "models.yaml", """
|
||||
bind:
|
||||
host: 127.0.0.1
|
||||
port: 8765
|
||||
profiles:
|
||||
sonnet:
|
||||
baseUrl: http://gx10.gw:8000
|
||||
model: claude-sonnet-5
|
||||
rogue:
|
||||
baseUrl: http://gx11.gw:8000
|
||||
model: claude-opus-9000
|
||||
models:
|
||||
allow:
|
||||
- model: claude-sonnet-5
|
||||
""", "rogue");
|
||||
}
|
||||
}
|
||||
@@ -0,0 +1,75 @@
|
||||
package dev.ltms.fleet;
|
||||
|
||||
import org.junit.jupiter.api.DisplayName;
|
||||
import org.junit.jupiter.api.Test;
|
||||
|
||||
import static org.junit.jupiter.api.Assertions.assertTrue;
|
||||
|
||||
/**
|
||||
* fleetd #446 follow-up (criterion 2): {@code exhaustionSink} used to build the "name the fix"
|
||||
* WARNING inline with two SLF4J {@code {}}-placeholder {@code log.warn} calls. Nothing pinned
|
||||
* that text — a mutation battery run against the merged PR #457 proved it: renaming the leading
|
||||
* {@code "usage-limit fix:"} tag to {@code "MUTANT no fix named:"} at both call sites left all
|
||||
* 1586 existing tests green. This class is what makes that mutation red.
|
||||
*
|
||||
* <p>{@link Fleetd#usageLimitFixWarning} and {@link Fleetd#usageLimitFixWarningNoModel} are the
|
||||
* extracted call sites {@code exhaustionSink} actually invokes — {@code log.warn} is called with
|
||||
* each method's return value as a single, already-formatted argument, so the string these tests
|
||||
* assert on is byte-identical to what {@code fleetd.out} receives. Same extracted-static-method +
|
||||
* dedicated-test idiom as {@link Fleetd#exhaustedPatternCoverageLine} /
|
||||
* {@link FleetdPatternCoverageLineTest}, chosen over the {@link ch.qos.logback.core.read.ListAppender}
|
||||
* idiom {@code ExhaustedPatternGapReportTest} uses — this warning is built once, at one call site,
|
||||
* from a single method with no branching inside the message itself, so there is no real log
|
||||
* plumbing left to prove; capturing an appender would only add setup/teardown around the same
|
||||
* assertion.
|
||||
*
|
||||
* <p>Each assertion below checks for a specific substring — the leading {@code "usage-limit fix:"}
|
||||
* tag an operator would grep the log for, the profile name, the model name, and the literal
|
||||
* action ({@code enabled: false} under {@code models.allow}) — rather than the whole sentence.
|
||||
* Criterion 2 promised those facts, not a particular wording of the prose that carries them; a
|
||||
* whole-sentence {@code assertEquals} breaks on every copy-edit of that surrounding prose and
|
||||
* teaches the next person to delete the test rather than read why it failed. The leading tag is
|
||||
* pinned on purpose, unlike the rest of the sentence: it is the identifying label the fleetd #446
|
||||
* follow-up's own mutation battery renamed ({@code "usage-limit fix:"} → {@code "MUTANT no fix
|
||||
* named:"}) to prove this class did not yet catch a tag rename — only the three-fact substrings
|
||||
* above would have missed it, since none of them mention the tag itself.
|
||||
*/
|
||||
class FleetdUsageLimitFixWarningTest {
|
||||
|
||||
@Test
|
||||
@DisplayName("usageLimitFixWarning names the profile, the model, and the models.allow fix")
|
||||
void usageLimitFixWarningNamesProfileModelAndFix() {
|
||||
String warning = Fleetd.usageLimitFixWarning("terra", "claude-opus-5");
|
||||
|
||||
assertTrue(warning.startsWith("usage-limit fix:"),
|
||||
"must carry the grep-able identifying tag: " + warning);
|
||||
assertTrue(warning.contains("terra"), "must name the profile: " + warning);
|
||||
assertTrue(warning.contains("claude-opus-5"), "must name the model: " + warning);
|
||||
assertTrue(warning.contains("enabled: false"), "must name the action: " + warning);
|
||||
assertTrue(warning.contains("models.allow"), "must name where the action goes: " + warning);
|
||||
}
|
||||
|
||||
@Test
|
||||
@DisplayName("usageLimitFixWarning does not cross-name a different profile or model")
|
||||
void usageLimitFixWarningDoesNotCrossName() {
|
||||
String warning = Fleetd.usageLimitFixWarning("terra", "claude-opus-5");
|
||||
|
||||
// Guards against the mutation of swapping the two profile/model arguments at the call
|
||||
// site — a test that only checks "some name appears" cannot catch that.
|
||||
assertTrue(!warning.contains("sonnet"), "must not name an unrelated model: " + warning);
|
||||
}
|
||||
|
||||
@Test
|
||||
@DisplayName("usageLimitFixWarningNoModel names the profile but no fix, since none exists")
|
||||
void usageLimitFixWarningNoModelNamesProfileNotAFix() {
|
||||
String warning = Fleetd.usageLimitFixWarningNoModel("gx", 1800);
|
||||
|
||||
assertTrue(warning.startsWith("usage-limit fix:"),
|
||||
"must carry the grep-able identifying tag: " + warning);
|
||||
assertTrue(warning.contains("gx"), "must name the profile: " + warning);
|
||||
assertTrue(warning.contains("1800"), "must name the quarantine duration it falls back to: " + warning);
|
||||
assertTrue(!warning.contains("enabled: false"),
|
||||
"a profile with no model: configured has no models.allow entry to flip — the "
|
||||
+ "fallback message must not claim one exists: " + warning);
|
||||
}
|
||||
}
|
||||
@@ -113,6 +113,20 @@ class AuthzTest {
|
||||
assertTrue(Authz.permits(WORKER_A, METRICS, null));
|
||||
}
|
||||
|
||||
/**
|
||||
* fleetd #421 — unlike READ (the case above), COORD_READ (a lead's own held peer mail) is the
|
||||
* primary's alone. A worker or an architect reading it would disclose lead-to-lead
|
||||
* coordination bodies, not the secret-free roster READ is open about.
|
||||
*/
|
||||
@Test
|
||||
void coordReadIsThePrimarysAloneNotAWidenedRead() {
|
||||
assertTrue(Authz.permits(PRIMARY, COORD_READ, null));
|
||||
assertFalse(Authz.permits(WORKER_A, COORD_READ, null),
|
||||
"a worker must not read held lead-to-lead mail");
|
||||
assertFalse(Authz.permits(ARCH_DESIGN, COORD_READ, null),
|
||||
"an architect holds READ today, but not-primary must mean not-architect here too");
|
||||
}
|
||||
|
||||
@Test
|
||||
void unauthenticatedIsDistinguishedFromMerelyForbidden() {
|
||||
// Drives the 401-vs-403 split: a missing credential is fixable by the caller, a wrong role
|
||||
|
||||
@@ -0,0 +1,360 @@
|
||||
package dev.ltms.fleet.auth;
|
||||
|
||||
import dev.ltms.fleet.config.ConfigRef;
|
||||
import dev.ltms.fleet.config.FleetConfig;
|
||||
import dev.ltms.fleet.herdr.FakeHerdr;
|
||||
import dev.ltms.fleet.herdr.PaneLocator;
|
||||
import dev.ltms.fleet.mcp.ConnectionIdentity;
|
||||
import dev.ltms.fleet.peer.MemberRole;
|
||||
import org.junit.jupiter.api.Test;
|
||||
import org.junit.jupiter.api.io.TempDir;
|
||||
|
||||
import java.nio.file.Files;
|
||||
import java.nio.file.Path;
|
||||
import java.util.Map;
|
||||
|
||||
import static org.junit.jupiter.api.Assertions.*;
|
||||
|
||||
/**
|
||||
* fleetd #424 — revoking (or granting) an architect slot must take effect on the next spawn with
|
||||
* no restart. A session already bound to a slot keeps its <em>binding</em> (the {@code
|
||||
* terminalToSlot} occupancy) even after that slot drops out of config, but NOT the ARCHITECT
|
||||
* <em>privilege</em> the slot used to grant — that is revoked on the bound session's very next
|
||||
* request. See {@link MemberRegistry}'s class doc for the exact rule: "config governs what a
|
||||
* bound slot still grants, as well as what may be bound next."
|
||||
*
|
||||
* <p>Every test here drives a REAL {@link ConfigRef#reload()} against a {@code @TempDir} file and
|
||||
* asserts {@link ConfigRef.Outcome#applied()}, rather than comparing two frozen
|
||||
* {@code MemberRegistry} instances in memory — the defect this ticket fixes is specifically that
|
||||
* {@link MemberRegistry} used to ignore a live reload, so a test that never reloads cannot tell the
|
||||
* fixed registry from the broken one. {@code requireSlotFor} and {@code reserve} are pinned in
|
||||
* separate tests, in both directions (removed and added), so a registry that simply refuses (or
|
||||
* simply allows) everything cannot pass by accident — see {@link MemberRegistryTest} for the
|
||||
* registry's other invariants (bind/unbind cardinality, thread-safety), which are unaffected by
|
||||
* this ticket and still exercised against the frozen constructor.
|
||||
*/
|
||||
class MemberRegistryLiveTest {
|
||||
|
||||
private static String yaml(String fleetBlock) {
|
||||
return """
|
||||
bind:
|
||||
host: 127.0.0.1
|
||||
port: 8765
|
||||
herdrSocket: ~/.config/herdr/herdr.sock
|
||||
profiles:
|
||||
sonnet:
|
||||
baseUrl: http://gx00.gw:8000
|
||||
model: sonnet
|
||||
opus:
|
||||
baseUrl: http://gx00.gw:8001
|
||||
model: opus
|
||||
guard:
|
||||
offSubscriptionHosts:
|
||||
- gx00.gw
|
||||
""" + fleetBlock;
|
||||
}
|
||||
|
||||
private static final String WITH_SONNET_SLOT = """
|
||||
fleet:
|
||||
architects:
|
||||
designer:
|
||||
profile: sonnet
|
||||
""";
|
||||
|
||||
/** No architect pool at all — developers is unrelated dead data for this registry (#424 out of scope). */
|
||||
private static final String WITHOUT_ARCHITECT_SLOTS = """
|
||||
fleet:
|
||||
developers:
|
||||
dev1:
|
||||
profile: sonnet
|
||||
""";
|
||||
|
||||
/** Same slot name ({@code designer}) as {@link #WITH_SONNET_SLOT}, repointed to a different profile. */
|
||||
private static final String WITH_OPUS_SLOT = """
|
||||
fleet:
|
||||
architects:
|
||||
designer:
|
||||
profile: opus
|
||||
""";
|
||||
|
||||
/** Same profile ({@code sonnet}) as {@link #WITH_SONNET_SLOT}, but the pool key is renamed. */
|
||||
private static final String WITH_RENAMED_SLOT = """
|
||||
fleet:
|
||||
architects:
|
||||
architect-lead:
|
||||
profile: sonnet
|
||||
""";
|
||||
|
||||
private static ConfigRef refFor(Path f) {
|
||||
return new ConfigRef(f, FleetConfig.load(f));
|
||||
}
|
||||
|
||||
// ── requireSlotFor is live (criteria 1, 2, 4) ──────────────────────────────────────────────
|
||||
|
||||
@Test
|
||||
void requireSlotForRefusesAProfileWhoseSlotWasRemovedByReload(@TempDir Path dir) throws Exception {
|
||||
Path f = dir.resolve("fleetd.yaml");
|
||||
Files.writeString(f, yaml(WITH_SONNET_SLOT));
|
||||
ConfigRef ref = refFor(f);
|
||||
MemberRegistry registry = MemberRegistry.live(() -> ref.get().fleet());
|
||||
|
||||
assertDoesNotThrow(() -> registry.requireSlotFor(MemberRole.ARCHITECT, "sonnet"),
|
||||
"the slot is configured before the reload");
|
||||
|
||||
Files.writeString(f, yaml(WITHOUT_ARCHITECT_SLOTS));
|
||||
ConfigRef.Outcome out = ref.reload();
|
||||
assertTrue(out.applied(), "the reload must actually take effect: " + out.summary());
|
||||
|
||||
assertThrows(IllegalArgumentException.class,
|
||||
() -> registry.requireSlotFor(MemberRole.ARCHITECT, "sonnet"),
|
||||
"revoking the slot must refuse the NEXT spawn that names it");
|
||||
}
|
||||
|
||||
@Test
|
||||
void requireSlotForAllowsAProfileWhoseSlotWasAddedByReload(@TempDir Path dir) throws Exception {
|
||||
Path f = dir.resolve("fleetd.yaml");
|
||||
Files.writeString(f, yaml(WITHOUT_ARCHITECT_SLOTS));
|
||||
ConfigRef ref = refFor(f);
|
||||
MemberRegistry registry = MemberRegistry.live(() -> ref.get().fleet());
|
||||
|
||||
assertThrows(IllegalArgumentException.class,
|
||||
() -> registry.requireSlotFor(MemberRole.ARCHITECT, "sonnet"),
|
||||
"no architect slot is configured yet");
|
||||
|
||||
Files.writeString(f, yaml(WITH_SONNET_SLOT));
|
||||
ConfigRef.Outcome out = ref.reload();
|
||||
assertTrue(out.applied(), "the reload must actually take effect: " + out.summary());
|
||||
|
||||
assertDoesNotThrow(() -> registry.requireSlotFor(MemberRole.ARCHITECT, "sonnet"),
|
||||
"a slot added by reload must be usable with no restart");
|
||||
}
|
||||
|
||||
// ── reserve is live too — tested separately from requireSlotFor (criteria 1, 2, 4) ────────
|
||||
|
||||
@Test
|
||||
void reserveRefusesAProfileWhoseSlotWasRemovedByReload(@TempDir Path dir) throws Exception {
|
||||
Path f = dir.resolve("fleetd.yaml");
|
||||
Files.writeString(f, yaml(WITH_SONNET_SLOT));
|
||||
ConfigRef ref = refFor(f);
|
||||
MemberRegistry registry = MemberRegistry.live(() -> ref.get().fleet());
|
||||
|
||||
MemberLifecycle.SlotReservation before = registry.reserve(MemberRole.ARCHITECT, "sonnet");
|
||||
assertEquals("architect:designer", before.slot());
|
||||
registry.release(before); // free it back up so the reload-side reserve below starts clean
|
||||
|
||||
Files.writeString(f, yaml(WITHOUT_ARCHITECT_SLOTS));
|
||||
ConfigRef.Outcome out = ref.reload();
|
||||
assertTrue(out.applied(), "the reload must actually take effect: " + out.summary());
|
||||
|
||||
assertThrows(IllegalArgumentException.class,
|
||||
() -> registry.reserve(MemberRole.ARCHITECT, "sonnet"),
|
||||
"revoking the slot must refuse the NEXT reservation for it");
|
||||
}
|
||||
|
||||
@Test
|
||||
void reserveAllowsAProfileWhoseSlotWasAddedByReload(@TempDir Path dir) throws Exception {
|
||||
Path f = dir.resolve("fleetd.yaml");
|
||||
Files.writeString(f, yaml(WITHOUT_ARCHITECT_SLOTS));
|
||||
ConfigRef ref = refFor(f);
|
||||
MemberRegistry registry = MemberRegistry.live(() -> ref.get().fleet());
|
||||
|
||||
assertThrows(IllegalArgumentException.class,
|
||||
() -> registry.reserve(MemberRole.ARCHITECT, "sonnet"),
|
||||
"no architect slot is configured yet");
|
||||
|
||||
Files.writeString(f, yaml(WITH_SONNET_SLOT));
|
||||
ConfigRef.Outcome out = ref.reload();
|
||||
assertTrue(out.applied(), "the reload must actually take effect: " + out.summary());
|
||||
|
||||
MemberLifecycle.SlotReservation after = registry.reserve(MemberRole.ARCHITECT, "sonnet");
|
||||
assertEquals("architect:designer", after.slot(),
|
||||
"a slot added by reload must be reservable with no restart");
|
||||
}
|
||||
|
||||
// ── profileForSlot, isSlot and nameForSlot are live too (fleetd #431) ─────────────────────────
|
||||
// #424 pinned roleForSlot against a live reload but left these three untested — proved by
|
||||
// mutating each to read a snapshot flattened once at construction: the full suite stayed green
|
||||
// for all three.
|
||||
//
|
||||
// The three differ in how much production behaviour depends on them, and the ticket first got
|
||||
// this ranking wrong. nameForSlot is wired: CallerResolver passes members::nameForSlot, next to
|
||||
// members::roleForSlot. isSlot is reached through bind, which calls it to refuse an unknown
|
||||
// slot. profileForSlot has NO caller in src/main at all — grepped both the ".profileForSlot("
|
||||
// and the "::profileForSlot" form — so there is no seam to drive its test through beyond the
|
||||
// accessor itself, and its own javadoc ("what the spawn lifecycle reads") describes a caller
|
||||
// that does not exist. These tests pin the accessors as they are; whether profileForSlot should
|
||||
// be wired or deleted is a separate question.
|
||||
|
||||
@Test
|
||||
void profileForSlotReflectsAProfileChangedByReload(@TempDir Path dir) throws Exception {
|
||||
Path f = dir.resolve("fleetd.yaml");
|
||||
Files.writeString(f, yaml(WITH_SONNET_SLOT));
|
||||
ConfigRef ref = refFor(f);
|
||||
MemberRegistry registry = MemberRegistry.live(() -> ref.get().fleet());
|
||||
|
||||
assertEquals("sonnet", registry.profileForSlot("architect:designer"),
|
||||
"the slot's profile before the reload");
|
||||
|
||||
Files.writeString(f, yaml(WITH_OPUS_SLOT));
|
||||
ConfigRef.Outcome out = ref.reload();
|
||||
assertTrue(out.applied(), "the reload must actually take effect: " + out.summary());
|
||||
|
||||
assertEquals("opus", registry.profileForSlot("architect:designer"),
|
||||
"repointing the slot to a different profile must take effect with no restart");
|
||||
}
|
||||
|
||||
@Test
|
||||
void isSlotStopsReportingASlotRemovedByReload(@TempDir Path dir) throws Exception {
|
||||
Path f = dir.resolve("fleetd.yaml");
|
||||
Files.writeString(f, yaml(WITH_SONNET_SLOT));
|
||||
ConfigRef ref = refFor(f);
|
||||
MemberRegistry registry = MemberRegistry.live(() -> ref.get().fleet());
|
||||
|
||||
assertTrue(registry.isSlot("architect:designer"), "the slot is configured before the reload");
|
||||
|
||||
Files.writeString(f, yaml(WITHOUT_ARCHITECT_SLOTS));
|
||||
ConfigRef.Outcome out = ref.reload();
|
||||
assertTrue(out.applied(), "the reload must actually take effect: " + out.summary());
|
||||
|
||||
assertFalse(registry.isSlot("architect:designer"),
|
||||
"removing the slot from config must make isSlot say so on the very next call, "
|
||||
+ "with no restart");
|
||||
}
|
||||
|
||||
@Test
|
||||
void isSlotStartsReportingASlotAddedByReload(@TempDir Path dir) throws Exception {
|
||||
Path f = dir.resolve("fleetd.yaml");
|
||||
Files.writeString(f, yaml(WITHOUT_ARCHITECT_SLOTS));
|
||||
ConfigRef ref = refFor(f);
|
||||
MemberRegistry registry = MemberRegistry.live(() -> ref.get().fleet());
|
||||
|
||||
assertFalse(registry.isSlot("architect:designer"), "no architect slot is configured yet");
|
||||
|
||||
Files.writeString(f, yaml(WITH_SONNET_SLOT));
|
||||
ConfigRef.Outcome out = ref.reload();
|
||||
assertTrue(out.applied(), "the reload must actually take effect: " + out.summary());
|
||||
|
||||
assertTrue(registry.isSlot("architect:designer"),
|
||||
"a slot added by reload must be visible to isSlot with no restart");
|
||||
}
|
||||
|
||||
@Test
|
||||
void nameForSlotReflectsANameChangedByReload(@TempDir Path dir) throws Exception {
|
||||
Path f = dir.resolve("fleetd.yaml");
|
||||
Files.writeString(f, yaml(WITH_SONNET_SLOT));
|
||||
ConfigRef ref = refFor(f);
|
||||
MemberRegistry registry = MemberRegistry.live(() -> ref.get().fleet());
|
||||
|
||||
assertEquals("designer", registry.nameForSlot("architect:designer"),
|
||||
"the configured name before the reload");
|
||||
|
||||
Files.writeString(f, yaml(WITH_RENAMED_SLOT));
|
||||
ConfigRef.Outcome out = ref.reload();
|
||||
assertTrue(out.applied(), "the reload must actually take effect: " + out.summary());
|
||||
|
||||
assertNull(registry.nameForSlot("architect:designer"),
|
||||
"the old key no longer names a configured slot — it was renamed away by the reload");
|
||||
assertEquals("architect-lead", registry.nameForSlot("architect:architect-lead"),
|
||||
"the new name must be visible under its new qualified key with no restart");
|
||||
}
|
||||
|
||||
// ── a bound architect is demoted, but the binding itself is not touched (fleetd #424) ───────
|
||||
// The lead's corrected ruling: the PRIVILEGE a slot grants is revoked on the bound session's
|
||||
// very next request, but the terminalToSlot BINDING itself is untouched by a reload — dropping
|
||||
// it would double-book the slot key and break unbind's compare-safe contract. See the class
|
||||
// doc's binding rule.
|
||||
|
||||
/** A caller identity resolving the one canned pane (terminal {@code term_a}) in {@link FakeHerdr}. */
|
||||
private static ConnectionIdentity boundPaneIdentity() {
|
||||
return new ConnectionIdentity(new PaneLocator(new FakeHerdr()), _ -> FakeHerdr.WORKER_PID);
|
||||
}
|
||||
|
||||
@Test
|
||||
void anArchitectAlreadyBoundToASlotIsDemotedByReload(@TempDir Path dir) throws Exception {
|
||||
Path f = dir.resolve("fleetd.yaml");
|
||||
Files.writeString(f, yaml(WITH_SONNET_SLOT));
|
||||
ConfigRef ref = refFor(f);
|
||||
MemberRegistry registry = MemberRegistry.live(() -> ref.get().fleet());
|
||||
|
||||
MemberLifecycle.SlotReservation reservation = registry.reserve(MemberRole.ARCHITECT, "sonnet");
|
||||
assertTrue(registry.bind(reservation, "term_a"));
|
||||
assertEquals("architect:designer", registry.slotForTerminal("term_a"));
|
||||
|
||||
// Drive the real caller path, not the roleForSlot seam directly: CallerResolver.resolve is
|
||||
// what a live request actually goes through (CallerResolver.java:220), and a resolver that
|
||||
// ignored roleForSlot entirely would still pass a test that only checked the seam.
|
||||
CallerResolver resolver = CallerResolver.withLeadsAndMembers(
|
||||
boundPaneIdentity(), false, null, Map::of, registry);
|
||||
|
||||
Principal before = resolver.resolve("127.0.0.1", 42, null);
|
||||
assertEquals(Role.ARCHITECT, before.role(), "sanity check: the harness binds term_a as an architect");
|
||||
assertEquals("designer", before.name());
|
||||
|
||||
Files.writeString(f, yaml(WITHOUT_ARCHITECT_SLOTS));
|
||||
ConfigRef.Outcome out = ref.reload();
|
||||
assertTrue(out.applied(), "the reload must actually take effect: " + out.summary());
|
||||
|
||||
Principal after = resolver.resolve("127.0.0.1", 42, null);
|
||||
assertEquals(Role.WORKER, after.role(),
|
||||
"removing the slot from config must demote the bound session to worker on its "
|
||||
+ "NEXT request — this is the ticket's whole point");
|
||||
assertEquals("term_a", after.terminal(), "same pane, same terminal — only the role changed");
|
||||
}
|
||||
|
||||
@Test
|
||||
void theOriginalBindingStillOccupiesTheRemovedSlotSoASecondTerminalCannotClaimIt(@TempDir Path dir)
|
||||
throws Exception {
|
||||
Path f = dir.resolve("fleetd.yaml");
|
||||
Files.writeString(f, yaml(WITH_SONNET_SLOT));
|
||||
ConfigRef ref = refFor(f);
|
||||
MemberRegistry registry = MemberRegistry.live(() -> ref.get().fleet());
|
||||
|
||||
MemberLifecycle.SlotReservation reservation = registry.reserve(MemberRole.ARCHITECT, "sonnet");
|
||||
assertTrue(registry.bind(reservation, "term_a"));
|
||||
|
||||
Files.writeString(f, yaml(WITHOUT_ARCHITECT_SLOTS));
|
||||
ConfigRef.Outcome removed = ref.reload();
|
||||
assertTrue(removed.applied(), "the reload must actually take effect: " + removed.summary());
|
||||
|
||||
// The binding survives the removal untouched.
|
||||
assertEquals("architect:designer", registry.slotForTerminal("term_a"),
|
||||
"a live binding must never be retroactively unbound by a config edit");
|
||||
assertEquals(Map.of("term_a", "architect:designer"), registry.snapshot());
|
||||
|
||||
// Bring the slot back into config. If the binding had been silently dropped by the removal
|
||||
// (rather than merely losing the privilege it grants), a second terminal could now claim
|
||||
// the "freed" key — the exact double-booking the class doc's binding rule rules out.
|
||||
Files.writeString(f, yaml(WITH_SONNET_SLOT));
|
||||
ConfigRef.Outcome restored = ref.reload();
|
||||
assertTrue(restored.applied(), "the reload must actually take effect: " + restored.summary());
|
||||
|
||||
assertFalse(registry.bind("architect:designer", "term_b"),
|
||||
"the slot is still occupied by term_a — a second terminal must not bind to it");
|
||||
assertThrows(IllegalArgumentException.class,
|
||||
() -> registry.reserve(MemberRole.ARCHITECT, "sonnet"),
|
||||
"the slot is still occupied by term_a — a fresh reservation must not find it free");
|
||||
assertEquals("architect:designer", registry.slotForTerminal("term_a"),
|
||||
"the original binding is unchanged throughout");
|
||||
}
|
||||
|
||||
@Test
|
||||
void unbindStillSucceedsForTheOriginalTerminalAfterItsSlotIsRemoved(@TempDir Path dir) throws Exception {
|
||||
Path f = dir.resolve("fleetd.yaml");
|
||||
Files.writeString(f, yaml(WITH_SONNET_SLOT));
|
||||
ConfigRef ref = refFor(f);
|
||||
MemberRegistry registry = MemberRegistry.live(() -> ref.get().fleet());
|
||||
|
||||
MemberLifecycle.SlotReservation reservation = registry.reserve(MemberRole.ARCHITECT, "sonnet");
|
||||
assertTrue(registry.bind(reservation, "term_a"));
|
||||
|
||||
Files.writeString(f, yaml(WITHOUT_ARCHITECT_SLOTS));
|
||||
ConfigRef.Outcome out = ref.reload();
|
||||
assertTrue(out.applied(), "the reload must actually take effect: " + out.summary());
|
||||
|
||||
assertTrue(registry.unbind("architect:designer", "term_a"),
|
||||
"unbind must still work for a slot that config has since removed, or a session "
|
||||
+ "that outlives its slot's removal could never release it");
|
||||
assertNull(registry.slotForTerminal("term_a"));
|
||||
assertEquals(Map.of(), registry.snapshot());
|
||||
}
|
||||
}
|
||||
@@ -161,7 +161,15 @@ class ConfigRefProfileCoverageTest {
|
||||
// the #323 bug. Only ideProjectDir and worktreeGroup were saved by a behavioural test in
|
||||
// ConfigRefTest; the other two had none. So: growing this set now requires editing this
|
||||
// line as well, which is a visible, deliberate diff rather than a quiet one.
|
||||
assertEquals(Set.of("weight", "maxLoad", "credentialId"), excluded,
|
||||
//
|
||||
// fleetd #446: "exhaustedPattern" joined the set legitimately — it moved from compiled
|
||||
// once at Fleetd.main startup to read live, cached by profile name, through
|
||||
// LiveExhaustedPatterns (see that class's doc and ConfigRef's class doc). Unlike the #323
|
||||
// hazard this comment warns about, it is pinned by its OWN behavioural tests:
|
||||
// FleetdExhaustionDetectionArmedWiringTest proves the reload takes effect with no restart,
|
||||
// and the same test rerun against the pre-#446 compile site (see that ticket's report) is
|
||||
// what proves the assertion would have failed before the fix.
|
||||
assertEquals(Set.of("weight", "maxLoad", "credentialId", "exhaustedPattern"), excluded,
|
||||
"ConfigRef.LAUNCH_SETTINGS_EXCLUDED changed. A component belongs in it ONLY if it "
|
||||
+ "is read live off the config supplier, not baked into a launcher at "
|
||||
+ "startup. If you are adding one to silence this test, that is fleetd #323 "
|
||||
|
||||
@@ -339,12 +339,18 @@ class ConfigRefTest {
|
||||
}
|
||||
|
||||
/**
|
||||
* CB-578 stage B: exhaustedPattern is compiled once into Fleetd.main's pattern map at startup
|
||||
* (see ExhaustedPatternLookup), so a reload never re-reads it — a changed pattern must be
|
||||
* reported deferred exactly like model/baseUrl, not silently claimed as applied.
|
||||
* fleetd #446: exhaustedPattern moved from deferred to hot — {@code LiveExhaustedPatterns}
|
||||
* reads {@code config.get()} fresh on every lookup, cached by profile name (not baked into a
|
||||
* frozen map inside {@code Fleetd.main} any more, the way it was before this ticket, and the
|
||||
* way its sibling {@code errorPattern} still is — see
|
||||
* {@code changingAProfilesErrorPatternIsReportedAsDeferred} right below for the contrast).
|
||||
* So a reload that changes ONLY exhaustedPattern must report a plain "config reloaded", with
|
||||
* nothing deferred — the opposite of what this test asserted before fleetd #446, when
|
||||
* ConfigRef.sameLaunchSettings still compared exhaustedPattern and this reload reported it as
|
||||
* one of the changes needing a restart.
|
||||
*/
|
||||
@Test
|
||||
void changingAProfilesExhaustedPatternIsReportedAsDeferred(@TempDir Path dir) throws Exception {
|
||||
void changingAProfilesExhaustedPatternTakesEffectWithNoRestart(@TempDir Path dir) throws Exception {
|
||||
Path f = dir.resolve("fleetd.yaml");
|
||||
Files.writeString(f, """
|
||||
bind:
|
||||
@@ -379,17 +385,21 @@ class ConfigRefTest {
|
||||
ConfigRef.Outcome out = ref.reload();
|
||||
|
||||
assertTrue(out.applied());
|
||||
assertEquals(1, out.deferred().size(), out.deferred().toString());
|
||||
assertTrue(out.deferred().getFirst().contains("sonnet"), out.deferred().toString());
|
||||
assertTrue(out.deferred().getFirst().contains("launch settings"), out.deferred().toString());
|
||||
// The snapshot still carries the new value — a restart is what makes it take effect.
|
||||
assertEquals(0, out.deferred().size(),
|
||||
"exhaustedPattern is hot since fleetd #446 — changing only this key must not be "
|
||||
+ "reported as needing a restart: " + out.deferred());
|
||||
assertEquals(0, out.split().size(), out.split().toString());
|
||||
assertEquals("config reloaded", out.summary());
|
||||
// The snapshot carries the new value immediately — this IS the live value now, not
|
||||
// merely what a restart would eventually pick up.
|
||||
assertEquals("rate limit exceeded", ref.get().profiles().get("sonnet").exhaustedPattern());
|
||||
}
|
||||
|
||||
/**
|
||||
* fleetd #201 Unit 5: errorPattern is compiled once into Fleetd.main's backend-error pattern
|
||||
* map at startup (see BackendErrorPatternLookup), the same way exhaustedPattern is above — a
|
||||
* reload never re-reads it either, so a changed value must be reported deferred.
|
||||
* map at startup (see BackendErrorPatternLookup) — unlike its sibling exhaustedPattern above,
|
||||
* which fleetd #446 made hot, errorPattern was scoped OUT of that ticket on purpose and stays
|
||||
* deferred: a reload still never re-reads it, so a changed value must be reported deferred.
|
||||
*/
|
||||
@Test
|
||||
void changingAProfilesErrorPatternIsReportedAsDeferred(@TempDir Path dir) throws Exception {
|
||||
|
||||
@@ -79,6 +79,14 @@ class ConfigRefTopLevelCoverageTest {
|
||||
* <li>{@code memberCredentials} — read live at {@code Fleetd.java:198, 205, 729}.</li>
|
||||
* <li>{@code memberLoginShell} — read live at
|
||||
* {@code HerdrPeerLauncher#configuredMemberLoginShell}.</li>
|
||||
* <li>{@code models} (fleetd #422) — membership is re-validated in full against the fresh
|
||||
* config on every {@code ConfigRef.reload()} (via {@code FleetConfig#validateAll()}), so
|
||||
* a bad edit is refused, never cached stale; the on/off half is read live by {@code
|
||||
* CompositePeerLauncher.enforceModelEnabled} and its candidate filter, and by {@code
|
||||
* fleet_profiles}/{@code GET /profiles} through {@code PeerLauncher.disabledModels()}.
|
||||
* Unlike {@code fleet} below, nothing about {@code models} is baked into a startup-built
|
||||
* object anywhere — there is no frozen half, so it belongs here whole rather than in
|
||||
* {@code SPLIT_KEYS}.</li>
|
||||
* </ul>
|
||||
*
|
||||
* <p>{@code fleet} used to sit here too, on the strength of most of it (role pools, charters,
|
||||
@@ -93,7 +101,7 @@ class ConfigRefTopLevelCoverageTest {
|
||||
* while this test stayed green throughout.</p>
|
||||
*/
|
||||
private static final Set<String> HOT_EXCLUDED_TOP_LEVEL_KEYS =
|
||||
Set.of("placement", "memberCredentials", "memberLoginShell");
|
||||
Set.of("placement", "memberCredentials", "memberLoginShell", "models");
|
||||
|
||||
@Test
|
||||
void everyTopLevelComponentIsAccountedForInExactlyOneClass() {
|
||||
@@ -120,7 +128,7 @@ class ConfigRefTopLevelCoverageTest {
|
||||
|
||||
// The escape hatch is pinned. Growing it requires editing this line — a visible, deliberate
|
||||
// diff, not a quiet one. See the field javadoc above for what "belongs here" actually means.
|
||||
assertEquals(Set.of("placement", "memberCredentials", "memberLoginShell"), hot,
|
||||
assertEquals(Set.of("placement", "memberCredentials", "memberLoginShell", "models"), hot,
|
||||
"HOT_EXCLUDED_TOP_LEVEL_KEYS changed. A component belongs here ONLY if it is read "
|
||||
+ "live off the config supplier, never because adding it makes this test "
|
||||
+ "pass. If you are adding one to silence this test, that is fleetd #323 "
|
||||
|
||||
@@ -109,6 +109,8 @@ class ConfigRefTopLevelReportingCoverageTest {
|
||||
v.put("memberLoginShell", null);
|
||||
v.put("memberSkills", "/skills/a");
|
||||
v.put("idleSleepGuard", new FleetConfig.IdleSleepGuard(true));
|
||||
v.put("models", new FleetConfig.Models(
|
||||
List.of(new FleetConfig.Models.ModelEntry("model-a"))));
|
||||
assertNamesMatchComponents(v);
|
||||
return v;
|
||||
}
|
||||
@@ -151,6 +153,8 @@ class ConfigRefTopLevelReportingCoverageTest {
|
||||
v.put("memberLoginShell", null);
|
||||
v.put("memberSkills", "/skills/b");
|
||||
v.put("idleSleepGuard", new FleetConfig.IdleSleepGuard(false));
|
||||
v.put("models", new FleetConfig.Models(
|
||||
List.of(new FleetConfig.Models.ModelEntry("model-b"))));
|
||||
assertNamesMatchComponents(v);
|
||||
return v;
|
||||
}
|
||||
|
||||
@@ -2683,4 +2683,337 @@ class FleetConfigTest {
|
||||
assertTrue(cfg.idleSleepGuard().isEnabled(),
|
||||
"unlike ConfigReload/Health, this block defaults to ON even when present but empty");
|
||||
}
|
||||
|
||||
// --- models: central allow-list of usable models -----------------------------------------
|
||||
|
||||
/**
|
||||
* An absent {@code models:} block is today's behaviour exactly: no profile's {@code model:} is
|
||||
* checked against anything, whatever it says. This is the "existing config keeps working"
|
||||
* invariant — an operator on a gitignored {@code fleetd.yaml} that predates this feature must
|
||||
* not be broken by upgrading the daemon.
|
||||
*/
|
||||
@Test
|
||||
void absentModelsBlockValidatesNothing(@TempDir Path dir) throws Exception {
|
||||
Path f = dir.resolve("fleetd.yaml");
|
||||
Files.writeString(f, """
|
||||
bind:
|
||||
port: 8080
|
||||
profiles:
|
||||
gx10:
|
||||
baseUrl: http://gx10.gw:8000
|
||||
model: totally-unheard-of-model-xyz
|
||||
""");
|
||||
|
||||
FleetConfig cfg = FleetConfig.load(f);
|
||||
assertDoesNotThrow(cfg::validateModels);
|
||||
}
|
||||
|
||||
/**
|
||||
* {@code models:} present but with an empty (or absent) {@code allow:} must behave exactly like
|
||||
* the block being absent — an operator adding the block for the first time with nothing in it
|
||||
* yet must not be surprised by every profile suddenly refusing to start.
|
||||
*/
|
||||
@Test
|
||||
void modelsBlockPresentButEmptyValidatesNothing(@TempDir Path dir) throws Exception {
|
||||
Path f = dir.resolve("fleetd.yaml");
|
||||
Files.writeString(f, """
|
||||
bind:
|
||||
port: 8080
|
||||
profiles:
|
||||
gx10:
|
||||
baseUrl: http://gx10.gw:8000
|
||||
model: whatever-the-operator-typed
|
||||
models:
|
||||
allow: []
|
||||
""");
|
||||
|
||||
FleetConfig cfg = FleetConfig.load(f);
|
||||
assertDoesNotThrow(cfg::validateModels);
|
||||
}
|
||||
|
||||
/**
|
||||
* The core of the ticket: a profile naming a model outside the configured allow-list fails
|
||||
* config load, naming both the model and the profile that wanted it.
|
||||
*/
|
||||
@Test
|
||||
void aProfileNamingAModelOutsideTheAllowListRefusesToStart(@TempDir Path dir) throws Exception {
|
||||
Path f = dir.resolve("fleetd.yaml");
|
||||
Files.writeString(f, """
|
||||
bind:
|
||||
port: 8080
|
||||
profiles:
|
||||
sonnet:
|
||||
baseUrl: http://gx10.gw:8000
|
||||
model: claude-sonnet-5
|
||||
rogue:
|
||||
baseUrl: http://gx11.gw:8000
|
||||
model: claude-opus-9000
|
||||
models:
|
||||
allow:
|
||||
- model: claude-sonnet-5
|
||||
- model: claude-haiku-5
|
||||
""");
|
||||
FleetConfig cfg = FleetConfig.load(f);
|
||||
|
||||
IllegalStateException e = assertThrows(IllegalStateException.class, cfg::validateModels);
|
||||
assertTrue(e.getMessage().contains("rogue"), "the refusal names the offending profile");
|
||||
assertTrue(e.getMessage().contains("claude-opus-9000"), "the refusal names the offending model");
|
||||
assertFalse(e.getMessage().contains("'sonnet'"),
|
||||
"the profile whose model IS allowed must not be reported");
|
||||
}
|
||||
|
||||
/** A profile whose {@code model:} is on the allow-list loads fine. */
|
||||
@Test
|
||||
void aProfileNamingAModelOnTheAllowListLoadsFine(@TempDir Path dir) throws Exception {
|
||||
Path f = dir.resolve("fleetd.yaml");
|
||||
Files.writeString(f, """
|
||||
bind:
|
||||
port: 8080
|
||||
profiles:
|
||||
sonnet:
|
||||
baseUrl: http://gx10.gw:8000
|
||||
model: claude-sonnet-5
|
||||
models:
|
||||
allow:
|
||||
- model: claude-sonnet-5
|
||||
""");
|
||||
|
||||
FleetConfig cfg = FleetConfig.load(f);
|
||||
assertDoesNotThrow(cfg::validateModels);
|
||||
}
|
||||
|
||||
/**
|
||||
* A profile that never sets {@code model:} (a {@code subscription: true} profile relying on the
|
||||
* account's own default is the live example) must not be refused just because an allow-list is
|
||||
* active — there is nothing to check it against.
|
||||
*/
|
||||
@Test
|
||||
void aProfileWithNoModelPassesEvenWithAnActiveAllowList(@TempDir Path dir) throws Exception {
|
||||
Path f = dir.resolve("fleetd.yaml");
|
||||
Files.writeString(f, """
|
||||
bind:
|
||||
port: 8080
|
||||
profiles:
|
||||
opus:
|
||||
subscription: true
|
||||
models:
|
||||
allow:
|
||||
- model: claude-sonnet-5
|
||||
""");
|
||||
|
||||
FleetConfig cfg = FleetConfig.load(f);
|
||||
assertDoesNotThrow(cfg::validateModels);
|
||||
}
|
||||
|
||||
/**
|
||||
* Reproduces the live config shape this ticket measured: a mix of {@code amazon-bedrock},
|
||||
* {@code opencode} and {@code openai}-backed profiles, some naming a bare Claude id and some a
|
||||
* provider-prefixed opencode id, in ONE allow-list. Both forms are just opaque strings compared
|
||||
* for exact equality — proves the shape decision holds against the shape actually seen live,
|
||||
* not just against a synthetic single-provider example.
|
||||
*/
|
||||
@Test
|
||||
void aBareClaudeIdAndAnOpencodeProviderPrefixedIdBothFitOneAllowList(@TempDir Path dir) throws Exception {
|
||||
Path f = dir.resolve("fleetd.yaml");
|
||||
Files.writeString(f, """
|
||||
bind:
|
||||
port: 8080
|
||||
profiles:
|
||||
sonnet:
|
||||
baseUrl: http://gx10.gw:8000
|
||||
model: claude-sonnet-5
|
||||
terra:
|
||||
kind: opencode
|
||||
model: openai/gpt-5.6-terra
|
||||
nova:
|
||||
kind: opencode
|
||||
model: amazon-bedrock/amazon.nova-pro-v1:0
|
||||
models:
|
||||
allow:
|
||||
- model: claude-sonnet-5
|
||||
- model: openai/gpt-5.6-terra
|
||||
- model: amazon-bedrock/amazon.nova-pro-v1:0
|
||||
""");
|
||||
|
||||
FleetConfig cfg = FleetConfig.load(f);
|
||||
assertDoesNotThrow(cfg::validateModels);
|
||||
}
|
||||
|
||||
/** The same live shape, but one opencode profile's model is missing from the allow-list. */
|
||||
@Test
|
||||
void anUnlistedOpencodeProviderPrefixedModelRefusesToStart(@TempDir Path dir) throws Exception {
|
||||
Path f = dir.resolve("fleetd.yaml");
|
||||
Files.writeString(f, """
|
||||
bind:
|
||||
port: 8080
|
||||
profiles:
|
||||
sonnet:
|
||||
baseUrl: http://gx10.gw:8000
|
||||
model: claude-sonnet-5
|
||||
terra:
|
||||
kind: opencode
|
||||
model: openai/gpt-5.6-terra-withdrawn
|
||||
models:
|
||||
allow:
|
||||
- model: claude-sonnet-5
|
||||
- model: openai/gpt-5.6-terra
|
||||
""");
|
||||
FleetConfig cfg = FleetConfig.load(f);
|
||||
|
||||
IllegalStateException e = assertThrows(IllegalStateException.class, cfg::validateModels);
|
||||
assertTrue(e.getMessage().contains("terra"), "the refusal names the offending profile");
|
||||
assertTrue(e.getMessage().contains("openai/gpt-5.6-terra-withdrawn"),
|
||||
"the refusal names the offending model, with its provider prefix intact");
|
||||
}
|
||||
|
||||
/**
|
||||
* Editing a {@code profiles:} entry alone must never be able to widen what is permitted — only
|
||||
* editing {@code models.allow:} itself can. This is the invariant the ticket states explicitly;
|
||||
* this test pins it by giving a profile a plausible-looking model that was never added to the
|
||||
* allow-list and confirming it is still refused.
|
||||
*/
|
||||
@Test
|
||||
void addingAProfileCannotWidenTheAllowListByItself(@TempDir Path dir) throws Exception {
|
||||
Path f = dir.resolve("fleetd.yaml");
|
||||
Files.writeString(f, """
|
||||
bind:
|
||||
port: 8080
|
||||
profiles:
|
||||
sonnet:
|
||||
baseUrl: http://gx10.gw:8000
|
||||
model: claude-sonnet-5
|
||||
brand-new:
|
||||
baseUrl: http://gx12.gw:8000
|
||||
model: claude-sonnet-6-preview
|
||||
models:
|
||||
allow:
|
||||
- model: claude-sonnet-5
|
||||
""");
|
||||
FleetConfig cfg = FleetConfig.load(f);
|
||||
|
||||
IllegalStateException e = assertThrows(IllegalStateException.class, cfg::validateModels);
|
||||
assertTrue(e.getMessage().contains("brand-new"));
|
||||
assertTrue(e.getMessage().contains("claude-sonnet-6-preview"));
|
||||
}
|
||||
|
||||
/** {@code models} is a brand-new top-level key and must be recognized, not WARN-ed as unknown. */
|
||||
@Test
|
||||
void modelsIsAKnownTopLevelKey() {
|
||||
assertTrue(FleetConfig.KNOWN_TOP_LEVEL_KEYS.contains("models"));
|
||||
}
|
||||
|
||||
// --- fleetd #422: model on/off (spawn-time gate + hot reload) -----------------------------
|
||||
|
||||
/**
|
||||
* Criterion 3: an {@code allow:} entry written before {@code enabled:} existed — no such key in
|
||||
* the YAML at all — must behave exactly as it always did: on, and absent from {@link
|
||||
* FleetConfig.Models#offIds()}. This is the old-style fixture the ticket asks for, proven
|
||||
* through a real YAML load rather than only through the {@code ModelEntry(String)} back-compat
|
||||
* constructor.
|
||||
*/
|
||||
@Test
|
||||
void anOldStyleAllowEntryWithNoEnabledFieldStaysOn(@TempDir Path dir) throws Exception {
|
||||
Path f = dir.resolve("fleetd.yaml");
|
||||
Files.writeString(f, """
|
||||
bind:
|
||||
port: 8080
|
||||
profiles:
|
||||
sonnet:
|
||||
baseUrl: http://gx10.gw:8000
|
||||
model: claude-sonnet-5
|
||||
models:
|
||||
allow:
|
||||
- model: claude-sonnet-5
|
||||
""");
|
||||
|
||||
FleetConfig cfg = FleetConfig.load(f);
|
||||
assertDoesNotThrow(cfg::validateModels);
|
||||
assertTrue(cfg.models().ids().contains("claude-sonnet-5"), "membership is unaffected");
|
||||
assertTrue(cfg.models().offIds().isEmpty(), "no enabled: field ⇒ nothing is off");
|
||||
}
|
||||
|
||||
/**
|
||||
* Criterion 7 (the whole point of the ticket, mirroring criterion 3's shape): a profile naming a
|
||||
* model that is turned off ({@code enabled: false}) must still be VALID config —
|
||||
* {@code validateModels()} checks membership only, never the on/off state, so it must not throw.
|
||||
* Collapsing "off" into "removed from allow:" would make this test fail, which is exactly the
|
||||
* design trap the ticket calls out.
|
||||
*/
|
||||
@Test
|
||||
void anOffModelIsStillValidConfigForValidateModels(@TempDir Path dir) throws Exception {
|
||||
Path f = dir.resolve("fleetd.yaml");
|
||||
Files.writeString(f, """
|
||||
bind:
|
||||
port: 8080
|
||||
profiles:
|
||||
local:
|
||||
baseUrl: http://local.gw:8000
|
||||
model: deepseek-v4-flash
|
||||
models:
|
||||
allow:
|
||||
- model: deepseek-v4-flash
|
||||
enabled: false
|
||||
""");
|
||||
|
||||
FleetConfig cfg = FleetConfig.load(f);
|
||||
assertDoesNotThrow(cfg::validateModels,
|
||||
"an off model must stay a valid allow-list member — only the spawn gate reads enabled");
|
||||
assertTrue(cfg.models().ids().contains("deepseek-v4-flash"));
|
||||
assertTrue(cfg.models().offIds().contains("deepseek-v4-flash"));
|
||||
}
|
||||
|
||||
/**
|
||||
* Criterion 8: one {@code enabled: false} entry names a model, not a profile — every profile
|
||||
* naming that model is off, on purpose (the live example: {@code deepseek-v4-flash} backs both
|
||||
* {@code local} and {@code local-direct}). Proven here at the {@code Models}/config-load level;
|
||||
* {@code CompositePeerLauncherTest} proves the spawn-time consequence for both profiles.
|
||||
*/
|
||||
@Test
|
||||
void turningOffOneModelIsIndependentOfHowManyProfilesNameIt(@TempDir Path dir) throws Exception {
|
||||
Path f = dir.resolve("fleetd.yaml");
|
||||
Files.writeString(f, """
|
||||
bind:
|
||||
port: 8080
|
||||
profiles:
|
||||
local:
|
||||
baseUrl: http://local.gw:8000
|
||||
model: deepseek-v4-flash
|
||||
local-direct:
|
||||
baseUrl: http://local.gw:8001
|
||||
model: deepseek-v4-flash
|
||||
models:
|
||||
allow:
|
||||
- model: deepseek-v4-flash
|
||||
enabled: false
|
||||
""");
|
||||
|
||||
FleetConfig cfg = FleetConfig.load(f);
|
||||
assertDoesNotThrow(cfg::validateModels);
|
||||
// One allow-list entry, one off id — the fan-out to both profiles happens at the reader
|
||||
// (CompositePeerLauncher), not by duplicating the entry per profile.
|
||||
assertEquals(Set.of("deepseek-v4-flash"), cfg.models().offIds());
|
||||
assertEquals("deepseek-v4-flash", cfg.profiles().get("local").model());
|
||||
assertEquals("deepseek-v4-flash", cfg.profiles().get("local-direct").model());
|
||||
}
|
||||
|
||||
/** An {@code enabled: true} entry (explicit, not just absent) also stays on — not just null. */
|
||||
@Test
|
||||
void explicitlyEnabledTrueStaysOn(@TempDir Path dir) throws Exception {
|
||||
Path f = dir.resolve("fleetd.yaml");
|
||||
Files.writeString(f, """
|
||||
bind:
|
||||
port: 8080
|
||||
profiles:
|
||||
sonnet:
|
||||
baseUrl: http://gx10.gw:8000
|
||||
model: claude-sonnet-5
|
||||
models:
|
||||
allow:
|
||||
- model: claude-sonnet-5
|
||||
enabled: true
|
||||
""");
|
||||
|
||||
FleetConfig cfg = FleetConfig.load(f);
|
||||
assertTrue(cfg.models().offIds().isEmpty());
|
||||
}
|
||||
}
|
||||
|
||||
@@ -0,0 +1,358 @@
|
||||
package dev.ltms.fleet.config;
|
||||
|
||||
import org.junit.jupiter.api.Test;
|
||||
import org.junit.jupiter.api.io.TempDir;
|
||||
|
||||
import java.lang.reflect.Method;
|
||||
import java.nio.file.Files;
|
||||
import java.nio.file.Path;
|
||||
import java.util.ArrayList;
|
||||
import java.util.List;
|
||||
import java.util.Set;
|
||||
import java.util.TreeSet;
|
||||
|
||||
import static org.junit.jupiter.api.Assertions.assertDoesNotThrow;
|
||||
import static org.junit.jupiter.api.Assertions.assertEquals;
|
||||
import static org.junit.jupiter.api.Assertions.assertThrows;
|
||||
import static org.junit.jupiter.api.Assertions.assertTrue;
|
||||
|
||||
/**
|
||||
* The gap this class exists to close: mutation testing on the fleetd ticket "central allow-list
|
||||
* of usable models" found that although {@link FleetConfig#validateModels()}'s own logic was well
|
||||
* pinned, nothing proved either real caller ({@code Fleetd.main} and {@link ConfigRef#reload()})
|
||||
* still invoked it — deleting the call site left the full suite green (1478/0/0/0). A follow-up
|
||||
* measurement (same technique — remove one call site, run the suite, not read the code) found the
|
||||
* SAME gap for all five of {@link FleetConfig}'s other validators at startup, and for four of the
|
||||
* six inside {@link ConfigRef#reload()}. This is a class of gap, not one line's mistake: every one
|
||||
* of those thirteen tests called the validator itself directly, never the real caller that was
|
||||
* supposed to.
|
||||
*
|
||||
* <p>The fix replaces the six individual {@code cfg.validateXxx()} calls at each of the two real
|
||||
* call sites with one {@link FleetConfig#validateAll()}, which reaches every validator by
|
||||
* reflection rather than by a hand-maintained list of names. A hand-maintained list of six names
|
||||
* would have exactly the defect it replaces: the seventh validator someone adds next month has no
|
||||
* reason to be added to it, and nothing would say so. This class proves TWO separate claims, and
|
||||
* keeps them separate on purpose:
|
||||
*
|
||||
* <ol>
|
||||
* <li>{@link #theSweepMechanismIsGenericNotHardcodedToFleetConfigsSixNames()} and its neighbours
|
||||
* prove the reflective sweep itself ({@link FleetConfig#invokeAllValidators}) is a general
|
||||
* mechanism — it runs whatever public, no-arg, void {@code validateXxx()} methods a class
|
||||
* happens to declare today, including a class with more of them than {@link FleetConfig}
|
||||
* has right now. This is the proof that a future, real seventh validator on {@link
|
||||
* FleetConfig} would be swept automatically, without needing to add a real (unwanted)
|
||||
* seventh validator just to exercise the claim.</li>
|
||||
* <li>{@link #validateAllReachesEveryOneOfTodaysSixValidators()} proves {@link
|
||||
* FleetConfig#validateAll()} itself is wired to that same generic mechanism and genuinely
|
||||
* reaches each of today's six real validators — reusing the exact minimal failing
|
||||
* configurations {@code FleetConfigTest} already established for each one directly, so a
|
||||
* single call to {@code validateAll()} is shown to reproduce every one of those six
|
||||
* failures.</li>
|
||||
* </ol>
|
||||
*
|
||||
* <p>Together with the direct-{@code Fleetd.main}-invocation tests in {@code
|
||||
* FleetdStartupValidationTest} (which prove the real startup call site still calls {@code
|
||||
* validateAll()}) and the {@code ConfigRefTest} reload tests (which prove the same for {@link
|
||||
* ConfigRef#reload()}), removing {@code cfg.validateAll();} from either real call site now fails
|
||||
* a test in this module.
|
||||
*
|
||||
* <p><b>What is NOT pinned, measured rather than assumed.</b> Reverting {@link
|
||||
* FleetConfig#validateAll()} to a hardcoded list of today's six method calls leaves the whole
|
||||
* suite green (measured at review: 1491 tests, 0 failures). Nothing ties {@code validateAll()} to
|
||||
* the generic sweep — claim 1 proves {@link FleetConfig#invokeAllValidators} is generic, and claim
|
||||
* 2 proves {@code validateAll()} reaches today's six, and a hardcoded list satisfies both. So the
|
||||
* reflective sweep is a convenience, not the guarantee. The guarantee is {@link
|
||||
* #fleetConfigDeclaresExactlyTheseSixValidatorsToday()}: it fails the moment a seventh validator
|
||||
* is declared, which forces whoever adds it to look at this file.
|
||||
*/
|
||||
class FleetConfigValidateAllTest {
|
||||
|
||||
// ── Claim 1: the reflective sweep is a general mechanism, not six names in disguise ──────────
|
||||
|
||||
/**
|
||||
* A throwaway fixture class, unrelated to {@link FleetConfig} in every way except shape: three
|
||||
* public, no-arg, void methods named {@code validateXxx}. Proves the sweep works on ANY class
|
||||
* with this shape, not on something special-cased to {@link FleetConfig}.
|
||||
*/
|
||||
static class ThreeValidators {
|
||||
final List<String> ran = new ArrayList<>();
|
||||
|
||||
public void validateAlpha() {
|
||||
ran.add("validateAlpha");
|
||||
}
|
||||
|
||||
public void validateBeta() {
|
||||
ran.add("validateBeta");
|
||||
}
|
||||
|
||||
public void validateGamma() {
|
||||
ran.add("validateGamma");
|
||||
}
|
||||
}
|
||||
|
||||
@Test
|
||||
void theSweepMechanismIsGenericNotHardcodedToFleetConfigsSixNames() {
|
||||
ThreeValidators target = new ThreeValidators();
|
||||
FleetConfig.invokeAllValidators(target);
|
||||
assertEquals(List.of("validateAlpha", "validateBeta", "validateGamma"), target.ran,
|
||||
"every validateXxx() method on this unrelated class must run, in alphabetical "
|
||||
+ "order — the sweep reads the class's own shape, not a name FleetConfig "
|
||||
+ "happens to know about");
|
||||
}
|
||||
|
||||
/**
|
||||
* The core of the "self-maintaining" requirement: the exact same class shape as {@link
|
||||
* ThreeValidators}, plus one more method — standing in for "a developer adds a validator next
|
||||
* month". Nothing about the sweep changes to pick it up; the new method is invoked purely
|
||||
* because it exists and matches the shape. This is what makes adding a seventh real validator
|
||||
* to {@link FleetConfig} safe without touching {@link FleetConfig#validateAll()} or either
|
||||
* call site — there is no "wire it in" step left to forget.
|
||||
*/
|
||||
static class FourValidators {
|
||||
final List<String> ran = new ArrayList<>();
|
||||
|
||||
public void validateAlpha() {
|
||||
ran.add("validateAlpha");
|
||||
}
|
||||
|
||||
public void validateBeta() {
|
||||
ran.add("validateBeta");
|
||||
}
|
||||
|
||||
public void validateGamma() {
|
||||
ran.add("validateGamma");
|
||||
}
|
||||
|
||||
public void validateDelta() {
|
||||
ran.add("validateDelta");
|
||||
}
|
||||
}
|
||||
|
||||
@Test
|
||||
void addingAFourthValidatorMethodGetsSweptWithNoOtherChange() {
|
||||
FourValidators target = new FourValidators();
|
||||
FleetConfig.invokeAllValidators(target);
|
||||
assertEquals(List.of("validateAlpha", "validateBeta", "validateDelta", "validateGamma"),
|
||||
sorted(target.ran),
|
||||
"the fourth method must be reached automatically — proving a class can grow the "
|
||||
+ "set of things it validates with no change to the sweep itself");
|
||||
}
|
||||
|
||||
private static List<String> sorted(List<String> in) {
|
||||
List<String> copy = new ArrayList<>(in);
|
||||
copy.sort(String::compareTo);
|
||||
return copy;
|
||||
}
|
||||
|
||||
@Test
|
||||
void aFailingValidatorStopsTheSweepAndPropagatesUnchanged() {
|
||||
class OneFails {
|
||||
@SuppressWarnings("unused")
|
||||
public void validateOk() {
|
||||
// passes
|
||||
}
|
||||
|
||||
public void validateBoom() {
|
||||
throw new IllegalStateException("refusing to start: boom");
|
||||
}
|
||||
}
|
||||
IllegalStateException e = assertThrows(IllegalStateException.class,
|
||||
() -> FleetConfig.invokeAllValidators(new OneFails()));
|
||||
assertEquals("refusing to start: boom", e.getMessage(),
|
||||
"the real exception must propagate unchanged, not be wrapped or swallowed");
|
||||
}
|
||||
|
||||
/**
|
||||
* Every rule the sweep's filter applies, proven independently: only public, no-arg, void
|
||||
* methods whose name starts with {@code "validate"} run, {@code validateAll} itself is
|
||||
* excluded (so a class that declares one of its own — as {@link FleetConfig} does — cannot
|
||||
* recurse into itself), and a same-shaped-but-wrongly-named or wrongly-shaped method never
|
||||
* runs. A reader who "simplifies" the filter in {@link FleetConfig#invokeAllValidators} in a
|
||||
* way that widens or narrows it breaks one of these.
|
||||
*/
|
||||
static class FilterEdgeCases {
|
||||
final List<String> ran = new ArrayList<>();
|
||||
|
||||
public void validateReal() {
|
||||
ran.add("validateReal");
|
||||
}
|
||||
|
||||
/** Wrong name — must not run. */
|
||||
public void checkSomething() {
|
||||
ran.add("checkSomething");
|
||||
}
|
||||
|
||||
/** Wrong shape — takes an argument. */
|
||||
public void validateWithArg(String ignored) {
|
||||
ran.add("validateWithArg");
|
||||
}
|
||||
|
||||
/** Wrong shape — returns something. */
|
||||
public boolean validateReturnsBoolean() {
|
||||
ran.add("validateReturnsBoolean");
|
||||
return true;
|
||||
}
|
||||
|
||||
/** Excluded by name on purpose, so the sweep cannot call itself. */
|
||||
public void validateAll() {
|
||||
ran.add("validateAll");
|
||||
}
|
||||
}
|
||||
|
||||
@Test
|
||||
void onlyPublicNoArgVoidValidateNamedMethodsRun() {
|
||||
FilterEdgeCases target = new FilterEdgeCases();
|
||||
FleetConfig.invokeAllValidators(target);
|
||||
assertEquals(List.of("validateReal"), target.ran,
|
||||
"checkSomething (wrong name), validateWithArg (wrong shape), "
|
||||
+ "validateReturnsBoolean (wrong shape), and validateAll (excluded by "
|
||||
+ "name) must all be skipped");
|
||||
}
|
||||
|
||||
// ── Claim 2: FleetConfig.validateAll() is wired to that mechanism and reaches all six today ──
|
||||
|
||||
/**
|
||||
* Reflectively enumerates {@link FleetConfig}'s own public, no-arg, void {@code validateXxx()}
|
||||
* methods (excluding {@code validateAll} itself) — the exact same filter {@link
|
||||
* FleetConfig#invokeAllValidators} applies. This is not the mechanism proof (that is claim 1,
|
||||
* above, on an unrelated class) — it is a visible denominator: today there are six, named
|
||||
* here, so a reader adding a seventh sees this assertion name the new count rather than a
|
||||
* silent pass at the old one.
|
||||
*/
|
||||
@Test
|
||||
void fleetConfigDeclaresExactlyTheseSixValidatorsToday() {
|
||||
Set<String> names = new TreeSet<>();
|
||||
for (Method m : FleetConfig.class.getMethods()) {
|
||||
if (java.lang.reflect.Modifier.isPublic(m.getModifiers())
|
||||
&& m.getParameterCount() == 0
|
||||
&& m.getReturnType() == void.class
|
||||
&& m.getName().startsWith("validate")
|
||||
&& !m.getName().equals("validateAll")) {
|
||||
names.add(m.getName());
|
||||
}
|
||||
}
|
||||
assertEquals(new TreeSet<>(Set.of("validateAuthExposure", "validateLeadTabPrefixes",
|
||||
"validateSubscriptionProfiles", "validateCharters", "validateMembers",
|
||||
"validateModels")), names,
|
||||
"FleetConfig's public validate*() methods changed. Do TWO things, in this "
|
||||
+ "order. First confirm validateAll() still delegates to "
|
||||
+ "invokeAllValidators(this) — a hardcoded list there passes every other "
|
||||
+ "test in this class, so this assertion is the only place that will ever "
|
||||
+ "make you check. Only then update the expected set to match.");
|
||||
}
|
||||
|
||||
/** A minimal, otherwise-valid file — same shape FleetConfigTest and ConfigRefTest use. */
|
||||
private static String minimalValidYaml() {
|
||||
return """
|
||||
bind:
|
||||
host: 127.0.0.1
|
||||
port: 8765
|
||||
profiles:
|
||||
sonnet:
|
||||
baseUrl: http://gx00.gw:8000
|
||||
model: sonnet
|
||||
""";
|
||||
}
|
||||
|
||||
@Test
|
||||
void aFullyValidConfigPassesValidateAll(@TempDir Path dir) throws Exception {
|
||||
Path f = dir.resolve("fleetd.yaml");
|
||||
Files.writeString(f, minimalValidYaml());
|
||||
assertDoesNotThrow(() -> FleetConfig.load(f).validateAll());
|
||||
}
|
||||
|
||||
/**
|
||||
* The heart of claim 2: for each of today's six real validators, a minimal file that fails
|
||||
* ONLY that one — the exact fixtures {@code FleetConfigTest} uses to test each validator
|
||||
* directly — must also fail through {@link FleetConfig#validateAll()}. If a future edit to
|
||||
* {@code validateAll()} silently dropped one validator from the sweep (e.g. a typo'd name
|
||||
* filter), exactly one of these six would start passing when it must not.
|
||||
*/
|
||||
@Test
|
||||
void validateAllReachesEveryOneOfTodaysSixValidators(@TempDir Path dir) throws Exception {
|
||||
// validateAuthExposure: a non-loopback bind without token mode.
|
||||
assertValidateAllRefuses(dir, "auth-exposure.yaml", """
|
||||
bind:
|
||||
host: 0.0.0.0
|
||||
port: 8765
|
||||
""", "auth.mode: token");
|
||||
|
||||
// validateLeadTabPrefixes: a fleet-wide tabLabel that starts with a lead's own tabPrefix.
|
||||
assertValidateAllRefuses(dir, "lead-tab-prefixes.yaml", """
|
||||
bind:
|
||||
host: 127.0.0.1
|
||||
port: 8765
|
||||
fleet:
|
||||
tabLabel: "lead: {role} {profile}"
|
||||
leaders:
|
||||
opus:
|
||||
tab: "lead: opus"
|
||||
""", "fleet.tabLabel");
|
||||
|
||||
// validateSubscriptionProfiles: subscription: true with env: reseating ANTHROPIC_BASE_URL.
|
||||
assertValidateAllRefuses(dir, "subscription-profiles.yaml", """
|
||||
bind:
|
||||
host: 127.0.0.1
|
||||
port: 8765
|
||||
profiles:
|
||||
sonnet:
|
||||
subscription: true
|
||||
argv: ["ccs", "sonnet"]
|
||||
env:
|
||||
ANTHROPIC_BASE_URL: http://anything-not-on-the-allowlist
|
||||
""", "ANTHROPIC_BASE_URL");
|
||||
|
||||
// validateCharters: a charter key that is not a role wire name.
|
||||
assertValidateAllRefuses(dir, "charters.yaml", """
|
||||
bind:
|
||||
host: 127.0.0.1
|
||||
port: 8765
|
||||
fleet:
|
||||
charters:
|
||||
architetc: text
|
||||
""", "architetc");
|
||||
|
||||
// validateMembers: an architect slot naming an unconfigured profile.
|
||||
assertValidateAllRefuses(dir, "members.yaml", """
|
||||
bind:
|
||||
host: 127.0.0.1
|
||||
port: 8765
|
||||
profiles:
|
||||
gx10:
|
||||
baseUrl: http://gx10.gw:8000
|
||||
fleet:
|
||||
architects:
|
||||
lead-designer:
|
||||
profile: sonnet
|
||||
""", "lead-designer");
|
||||
|
||||
// validateModels: a profile naming a model outside the configured allow-list.
|
||||
assertValidateAllRefuses(dir, "models.yaml", """
|
||||
bind:
|
||||
host: 127.0.0.1
|
||||
port: 8765
|
||||
profiles:
|
||||
sonnet:
|
||||
baseUrl: http://gx10.gw:8000
|
||||
model: claude-sonnet-5
|
||||
rogue:
|
||||
baseUrl: http://gx11.gw:8000
|
||||
model: claude-opus-9000
|
||||
models:
|
||||
allow:
|
||||
- model: claude-sonnet-5
|
||||
""", "rogue");
|
||||
}
|
||||
|
||||
private static void assertValidateAllRefuses(Path dir, String fileName, String yaml,
|
||||
String mustContain) throws Exception {
|
||||
Path f = dir.resolve(fileName);
|
||||
Files.writeString(f, yaml);
|
||||
FleetConfig cfg = FleetConfig.load(f);
|
||||
IllegalStateException e = assertThrows(IllegalStateException.class, cfg::validateAll,
|
||||
fileName + ": validateAll() must refuse this config");
|
||||
assertTrue(e.getMessage().contains(mustContain),
|
||||
fileName + ": expected message to contain \"" + mustContain + "\" but was: "
|
||||
+ e.getMessage());
|
||||
}
|
||||
}
|
||||
+2
@@ -97,6 +97,8 @@ class FleetConfigWithDefaultsPreservesEveryComponentTest {
|
||||
v.put("memberLoginShell", "/bin/zsh");
|
||||
v.put("memberSkills", "/skills/guard");
|
||||
v.put("idleSleepGuard", new FleetConfig.IdleSleepGuard(true));
|
||||
v.put("models", new FleetConfig.Models(
|
||||
List.of(new FleetConfig.Models.ModelEntry("model-guard"))));
|
||||
assertNamesMatchComponents(v);
|
||||
return v;
|
||||
}
|
||||
|
||||
@@ -17,15 +17,105 @@ import static org.junit.jupiter.api.Assumptions.assumeTrue;
|
||||
* SHELL directly (never {@code claude}, so no subscription/token involvement) and always tears
|
||||
* the throwaway space down.
|
||||
*
|
||||
* <p>The seed shell's own startup (restoring its session, printing its banner) is asynchronous
|
||||
* and its length is not a fleetd contract — measured here at ~2.5s on one host (fleetd #449).
|
||||
* The old version used two fixed sleeps: 1000ms before typing, then 800ms before reading. It
|
||||
* failed, and the pane showed the typed line followed by the startup banner with no command
|
||||
* output at all — which looks, at a glance, exactly like the env map never reaching the shell.
|
||||
* So this polls for a real signal instead of guessing a sleep length.
|
||||
*
|
||||
* <p><strong>What was measured, and what was not.</strong> Polling fixes it: 5 standalone runs
|
||||
* green. The load-bearing half is {@link #waitForText}, and one cell proves it. Keep the old
|
||||
* 1000ms write sleep and change only the read — the 800ms fixed sleep becomes a 5s poll — and
|
||||
* the test goes from 0 of 3 passing to 3 of 3. Removing the write wait instead
|
||||
* ({@link #SHELL_READY_TIMEOUT_MS} set to 0, so input is typed at once) also passes 3 of 3. So
|
||||
* the cause is the 800ms READ deadline, not the 1000ms write delay. The old version fails every
|
||||
* time, not sometimes, so "race" is the wrong word for it. The earlier explanation — that input
|
||||
* typed before the prompt is swallowed by the shell's startup — is not supported by anything
|
||||
* measured here. Please do not repeat it: if it were true, typing at 0ms would be worse than
|
||||
* typing at 1000ms, and it is not.
|
||||
*
|
||||
* <p><strong>Where this was measured.</strong> A 12-core macOS host, load average 2.6 to 5.9,
|
||||
* on commit 20c1094. The same four cells were also run under load, with 8 spinners on 12 cores.
|
||||
* The failure and the passes from that run are not worth the same. The cells ran in a fixed
|
||||
* order while the load climbed from 7 to 50. The old version ran last, at the top of that climb,
|
||||
* so it has a free explanation for failing and its 0 of 3 is discarded. A pass has no such free
|
||||
* explanation: a cell that survives a worse condition than a fair order would have given it is
|
||||
* evidence in the safe direction. So keep the three passes, each with the load it ran at: the
|
||||
* fixed version 3 of 3 at load 7.42 to 18.42, the 0ms-write cell 3 of 3 at 18.42 to 23.65, the
|
||||
* 1000ms-write cell 3 of 3 at 23.65 to 46.40. Above about load 20 everything here is slow for
|
||||
* reasons that have nothing to do with this seam, so read the positive claim — that the read
|
||||
* deadline was the whole cause — as "measured near idle on a 12-core host", and nothing
|
||||
* stronger. Do not carry that raw load average to another host either: load average counts
|
||||
* differently per core and per operating system, so only load per core compares. If this test
|
||||
* fails on a smaller or busier machine, raise {@link #OUTPUT_TIMEOUT_MS} before you suspect the
|
||||
* seam.
|
||||
*
|
||||
* <p>{@link #waitUntilSettled} is kept as cheap insurance against that swallow case, not because
|
||||
* anyone showed it was needed. The cell that tests the swallow case head-on is the one with
|
||||
* {@link #SHELL_READY_TIMEOUT_MS} at 0: input is typed at once, which is the worst case for
|
||||
* "typed before the prompt is ready". It passed 3 of 3 at load 18.42 to 23.65. The wider read
|
||||
* window cannot explain that pass away, because a swallowed keystroke is LOST, not late — the
|
||||
* command never runs, so no amount of polling makes its output appear. So the swallow mechanism
|
||||
* was tested and did not show up. If you want to delete this call, that is the cell to re-run.
|
||||
*
|
||||
* <p>Tagged {@code contract}; run with {@code mvn test -Pcontract}.
|
||||
*/
|
||||
@Tag("contract")
|
||||
class AgentControlContractTest {
|
||||
|
||||
private static final long POLL_INTERVAL_MS = 150;
|
||||
/** Bound for the seed shell to settle: observed ~2.5s three times running; this leaves headroom. */
|
||||
private static final long SHELL_READY_TIMEOUT_MS = 8_000;
|
||||
/** Bound for the typed command's output to appear once the shell is ready: observed ~0.2s. */
|
||||
private static final long OUTPUT_TIMEOUT_MS = 5_000;
|
||||
|
||||
private boolean noSocket() {
|
||||
return !Files.exists(UnixSocketHerdrClient.defaultSocketPath());
|
||||
}
|
||||
|
||||
private static String readPane(UnixSocketHerdrClient herdr, String paneId) {
|
||||
return herdr.call("pane.read", Map.of("pane_id", paneId, "source", "visible"))
|
||||
.path("read").path("text").asText("");
|
||||
}
|
||||
|
||||
/**
|
||||
* Poll {@code pane.read} until two consecutive reads come back identical — the shell's own
|
||||
* startup output (restore banner, prompt) has stopped changing — or {@code timeoutMs} elapses.
|
||||
* Never asserts by itself; the caller's own assertion is what actually verifies the outcome.
|
||||
* Setting {@code timeoutMs} to 0 skips the wait entirely and the test still passes here, so
|
||||
* treat this as insurance rather than as the fix — see the class javadoc.
|
||||
*/
|
||||
private static String waitUntilSettled(UnixSocketHerdrClient herdr, String paneId, long timeoutMs)
|
||||
throws InterruptedException {
|
||||
long deadline = System.currentTimeMillis() + timeoutMs;
|
||||
String previous = null;
|
||||
while (System.currentTimeMillis() < deadline) {
|
||||
Thread.sleep(POLL_INTERVAL_MS);
|
||||
String current = readPane(herdr, paneId);
|
||||
if (current.equals(previous) && !current.isBlank()) {
|
||||
return current;
|
||||
}
|
||||
previous = current;
|
||||
}
|
||||
return previous == null ? "" : previous;
|
||||
}
|
||||
|
||||
/** Poll {@code pane.read} until {@code needle} appears or {@code timeoutMs} elapses. */
|
||||
private static String waitForText(UnixSocketHerdrClient herdr, String paneId, String needle, long timeoutMs)
|
||||
throws InterruptedException {
|
||||
long deadline = System.currentTimeMillis() + timeoutMs;
|
||||
String last = "";
|
||||
while (System.currentTimeMillis() < deadline) {
|
||||
last = readPane(herdr, paneId);
|
||||
if (last.contains(needle)) {
|
||||
return last;
|
||||
}
|
||||
Thread.sleep(POLL_INTERVAL_MS);
|
||||
}
|
||||
return last;
|
||||
}
|
||||
|
||||
@Test
|
||||
void tabCreateInjectsEnvIntoTheSeedShell() throws Exception {
|
||||
assumeTrue(!noSocket(), "no herdr socket — skipping");
|
||||
@@ -36,15 +126,13 @@ class AgentControlContractTest {
|
||||
Map.of("ANTHROPIC_BASE_URL", "http://gx00.gw:8000"));
|
||||
try {
|
||||
assertNotNull(tab.rootPaneId(), "tab.create must return the seed pane");
|
||||
Thread.sleep(1000); // let the seed shell reach its prompt
|
||||
waitUntilSettled(herdr, tab.rootPaneId(), SHELL_READY_TIMEOUT_MS);
|
||||
herdr.call("pane.send_input", Map.of(
|
||||
"pane_id", tab.rootPaneId(),
|
||||
"text", "printf 'PROBE_BASE=[%s]\\n' \"$ANTHROPIC_BASE_URL\"",
|
||||
"keys", List.of("enter")));
|
||||
Thread.sleep(800);
|
||||
String visible = herdr.call("pane.read",
|
||||
Map.of("pane_id", tab.rootPaneId(), "source", "visible"))
|
||||
.path("read").path("text").asText("");
|
||||
String visible = waitForText(herdr, tab.rootPaneId(),
|
||||
"PROBE_BASE=[http://gx00.gw:8000]", OUTPUT_TIMEOUT_MS);
|
||||
assertTrue(visible.contains("PROBE_BASE=[http://gx00.gw:8000]"),
|
||||
"env map must reach the seed shell; saw: " + visible);
|
||||
} finally {
|
||||
|
||||
@@ -14,7 +14,7 @@ import static org.junit.jupiter.api.Assumptions.assumeTrue;
|
||||
* Contract test against a REAL running herdr. Tagged {@code contract} so it is
|
||||
* excluded from {@code mvn test}; run it with {@code mvn test -Pcontract}. It fails
|
||||
* loudly if herdr drifts from the protocol {@code fleetd} was built against
|
||||
* (0.7.0, protocol 14) — catching breakage that unit tests with canned frames cannot.
|
||||
* (0.8.0, protocol 19) — catching breakage that unit tests with canned frames cannot.
|
||||
*/
|
||||
@Tag("contract")
|
||||
class HerdrContractTest {
|
||||
@@ -24,13 +24,13 @@ class HerdrContractTest {
|
||||
}
|
||||
|
||||
@Test
|
||||
void pingReturnsProtocol14() {
|
||||
void pingReturnsProtocol19() {
|
||||
assumeTrue(Files.exists(socket()), "no herdr socket at " + socket() + " — skipping");
|
||||
try (UnixSocketHerdrClient herdr = UnixSocketHerdrClient.connect()) {
|
||||
JsonNode pong = herdr.call("ping");
|
||||
assertEquals("pong", pong.get("type").asText());
|
||||
assertEquals(14, pong.get("protocol").asInt(),
|
||||
"fleetd is built against herdr protocol 14");
|
||||
assertEquals(19, pong.get("protocol").asInt(),
|
||||
"fleetd is built against herdr protocol 19");
|
||||
assertFalse(pong.get("version").asText().isBlank());
|
||||
}
|
||||
}
|
||||
|
||||
@@ -20,6 +20,7 @@ import java.util.regex.Pattern;
|
||||
|
||||
import static org.junit.jupiter.api.Assertions.assertEquals;
|
||||
import static org.junit.jupiter.api.Assertions.assertFalse;
|
||||
import static org.junit.jupiter.api.Assertions.assertNotEquals;
|
||||
import static org.junit.jupiter.api.Assertions.assertTrue;
|
||||
|
||||
/** Unit behaviour of the CB-106 completion resolver in isolation from the injector. */
|
||||
@@ -884,19 +885,22 @@ class CompletionResolverTest {
|
||||
@Test
|
||||
void coverageIsOffWhenNoProfileHasAPatternConfigured() {
|
||||
assertEquals("off (no profile has an exhaustedPattern configured; profiles: [terra])",
|
||||
CompletionResolver.coverage("exhaustedPattern", Set.of("terra"), Set.of()));
|
||||
CompletionResolver.coverage("exhaustedPattern", CompletionResolver.UnsetMeaning.OFF,
|
||||
Set.of("terra"), Set.of()));
|
||||
}
|
||||
|
||||
@Test
|
||||
void coverageIsFullWhenEveryProfileHasAPatternConfigured() {
|
||||
assertEquals("full (all profiles configured: [gx10, terra])",
|
||||
CompletionResolver.coverage("exhaustedPattern", Set.of("terra", "gx10"), Set.of("terra", "gx10")));
|
||||
CompletionResolver.coverage("exhaustedPattern", CompletionResolver.UnsetMeaning.OFF,
|
||||
Set.of("terra", "gx10"), Set.of("terra", "gx10")));
|
||||
}
|
||||
|
||||
@Test
|
||||
void coverageIsPartialAndNamesWhichProfilesAreConfigured() {
|
||||
assertEquals("partial (configured: [terra]; not configured: [gx10])",
|
||||
CompletionResolver.coverage("exhaustedPattern", Set.of("terra", "gx10"), Set.of("terra")));
|
||||
CompletionResolver.coverage("exhaustedPattern", CompletionResolver.UnsetMeaning.OFF,
|
||||
Set.of("terra", "gx10"), Set.of("terra")));
|
||||
}
|
||||
|
||||
/**
|
||||
@@ -910,11 +914,42 @@ class CompletionResolverTest {
|
||||
*
|
||||
* <p>Every earlier test here passed the exhaustion case only, so none of them could see it. This
|
||||
* one pins that the message names the key the caller actually meant.
|
||||
*
|
||||
* <p>fleetd#415: the expected wording changed here too. {@code errorPattern} has a built-in
|
||||
* fallback ({@link CompletionResolver#BACKEND_ERROR}), so an empty {@code configuredProfiles}
|
||||
* for it is not "off" — see {@link #coverageDistinguishesOffFromBuiltInDefaultForTheSameEmptyInput}
|
||||
* for the test built specifically to pin that distinction.
|
||||
*/
|
||||
@Test
|
||||
void coverageNamesTheConfigKeyItsCallerMeansRatherThanAlwaysSayingExhaustedPattern() {
|
||||
assertEquals("off (no profile has an errorPattern configured; profiles: [gx10, terra])",
|
||||
CompletionResolver.coverage("errorPattern", Set.of("terra", "gx10"), Set.of()));
|
||||
assertEquals("built-in default for all profiles (no profile customises errorPattern; "
|
||||
+ "profiles: [gx10, terra])",
|
||||
CompletionResolver.coverage("errorPattern", CompletionResolver.UnsetMeaning.BUILT_IN_DEFAULT,
|
||||
Set.of("terra", "gx10"), Set.of()));
|
||||
}
|
||||
|
||||
/**
|
||||
* fleetd#415: {@code coverage()} measures pattern coverage (how many profiles set the key), but
|
||||
* for {@code errorPattern} the empty case is not the feature-off state — a profile with no
|
||||
* configured {@code errorPattern} still runs the classification against
|
||||
* {@link CompletionResolver#BACKEND_ERROR}. For {@code exhaustedPattern} there is no fallback,
|
||||
* so empty really is off. Same shape of input (empty {@code configuredProfiles}, one profile),
|
||||
* different {@link CompletionResolver.UnsetMeaning} — the wording must differ, or this method is
|
||||
* back to conflating pattern coverage with feature state for the one key where they disagree.
|
||||
*/
|
||||
@Test
|
||||
void coverageDistinguishesOffFromBuiltInDefaultForTheSameEmptyInput() {
|
||||
String exhaustedLine = CompletionResolver.coverage("exhaustedPattern",
|
||||
CompletionResolver.UnsetMeaning.OFF, Set.of("gx10", "terra"), Set.of());
|
||||
String errorLine = CompletionResolver.coverage("errorPattern",
|
||||
CompletionResolver.UnsetMeaning.BUILT_IN_DEFAULT, Set.of("gx10", "terra"), Set.of());
|
||||
|
||||
assertEquals("off (no profile has an exhaustedPattern configured; profiles: [gx10, terra])",
|
||||
exhaustedLine);
|
||||
assertEquals("built-in default for all profiles (no profile customises errorPattern; "
|
||||
+ "profiles: [gx10, terra])", errorLine);
|
||||
assertNotEquals(exhaustedLine, errorLine,
|
||||
"the same empty-coverage input must not read as the same feature state for both keys");
|
||||
}
|
||||
|
||||
// --- fleetd#201 Unit 1: target-keyed backend-error pattern + typed sink ----------------------
|
||||
|
||||
@@ -0,0 +1,138 @@
|
||||
package dev.ltms.fleet.inject;
|
||||
|
||||
import dev.ltms.fleet.config.FleetConfig;
|
||||
import org.junit.jupiter.api.DisplayName;
|
||||
import org.junit.jupiter.api.Test;
|
||||
|
||||
import java.util.Map;
|
||||
import java.util.regex.Pattern;
|
||||
|
||||
import static org.junit.jupiter.api.Assertions.assertFalse;
|
||||
import static org.junit.jupiter.api.Assertions.assertNotSame;
|
||||
import static org.junit.jupiter.api.Assertions.assertNull;
|
||||
import static org.junit.jupiter.api.Assertions.assertSame;
|
||||
import static org.junit.jupiter.api.Assertions.assertTrue;
|
||||
|
||||
/**
|
||||
* fleetd #446: {@code LiveExhaustedPatterns} is the mechanism that makes {@code exhaustedPattern}
|
||||
* hot — read fresh off a config supplier per lookup, cached by profile name. This is the direct,
|
||||
* minimal unit test of that mechanism: read the pattern once, mutate the backing "config", read
|
||||
* again, and assert the second read saw the new value — the exact hotness proof shape.
|
||||
*
|
||||
* <p>This class alone does not prove {@code Fleetd.main} actually wires the live source into
|
||||
* production — {@link dev.ltms.fleet.FleetdExhaustionDetectionArmedWiringTest} and {@code
|
||||
* CompletionResolverTest}'s {@code ExhaustedPatternLookup} tests do that at the wiring level. What
|
||||
* this class proves is that the mechanism itself is genuinely live and genuinely cached, not that
|
||||
* it is used.
|
||||
*/
|
||||
class LiveExhaustedPatternsTest {
|
||||
|
||||
private static FleetConfig.Profile profileWithPattern(String pattern) {
|
||||
return new FleetConfig.Profile("terra", "http://gx00.gw:8000", "terra",
|
||||
null, null, null, "tab", "fleetd-workers", "w #{n}", null, null, null,
|
||||
null, null, null, null, null, null, null, pattern);
|
||||
}
|
||||
|
||||
/** Simple mutable holder standing in for {@code ConfigRef} — swapped, never mutated in place. */
|
||||
private static final class MutableProfiles {
|
||||
private volatile Map<String, FleetConfig.Profile> profiles;
|
||||
|
||||
MutableProfiles(Map<String, FleetConfig.Profile> initial) {
|
||||
this.profiles = initial;
|
||||
}
|
||||
|
||||
Map<String, FleetConfig.Profile> get() {
|
||||
return profiles;
|
||||
}
|
||||
|
||||
void set(Map<String, FleetConfig.Profile> fresh) {
|
||||
this.profiles = fresh;
|
||||
}
|
||||
}
|
||||
|
||||
@Test
|
||||
@DisplayName("patternFor reads the live config: read, change, read again, see the change — no restart")
|
||||
void patternForIsHotAcrossAConfigChange() {
|
||||
MutableProfiles live = new MutableProfiles(Map.of("terra", profileWithPattern("usage limit")));
|
||||
LiveExhaustedPatterns patterns = new LiveExhaustedPatterns(live::get);
|
||||
|
||||
Pattern before = patterns.patternFor("terra");
|
||||
assertTrue(before.matcher("your usage limit has been reached").find());
|
||||
|
||||
// Change the "config" — the exact thing a reload does to ConfigRef in production.
|
||||
live.set(Map.of("terra", profileWithPattern("rate limit exceeded")));
|
||||
|
||||
Pattern after = patterns.patternFor("terra");
|
||||
assertTrue(after.matcher("429: rate limit exceeded, try later").find(),
|
||||
"the second read must see the NEW pattern text with no restart");
|
||||
assertFalse(after.matcher("your usage limit has been reached").find(),
|
||||
"the second read must have actually recompiled — not just reused a stale match");
|
||||
}
|
||||
|
||||
@Test
|
||||
@DisplayName("armed follows the same live read: on, then off, after a config change — no restart")
|
||||
void armedIsHotAcrossAConfigChange() {
|
||||
MutableProfiles live = new MutableProfiles(Map.of("terra", profileWithPattern("usage limit")));
|
||||
LiveExhaustedPatterns patterns = new LiveExhaustedPatterns(live::get);
|
||||
|
||||
assertTrue(patterns.armed("terra"));
|
||||
|
||||
live.set(Map.of("terra", profileWithPattern(null)));
|
||||
|
||||
assertFalse(patterns.armed("terra"),
|
||||
"removing exhaustedPattern and 'reloading' must disarm detection with no restart");
|
||||
}
|
||||
|
||||
@Test
|
||||
@DisplayName("an unchanged pattern string is not recompiled — the cache is reused")
|
||||
void unchangedPatternTextReusesTheCompiledInstance() {
|
||||
MutableProfiles live = new MutableProfiles(Map.of("terra", profileWithPattern("usage limit")));
|
||||
LiveExhaustedPatterns patterns = new LiveExhaustedPatterns(live::get);
|
||||
|
||||
Pattern first = patterns.patternFor("terra");
|
||||
// A NEW Profile object, same pattern text — a reload of an unrelated key produces a fresh
|
||||
// FleetConfig even though this profile's own text did not change.
|
||||
live.set(Map.of("terra", profileWithPattern("usage limit")));
|
||||
Pattern second = patterns.patternFor("terra");
|
||||
|
||||
assertSame(first, second,
|
||||
"identical pattern text must reuse the cached compiled Pattern, not recompile it — "
|
||||
+ "recompiling a regex on every check is exactly the cost the cache exists to avoid");
|
||||
}
|
||||
|
||||
@Test
|
||||
@DisplayName("a changed pattern string is recompiled, not silently reused from the cache")
|
||||
void changedPatternTextIsRecompiled() {
|
||||
MutableProfiles live = new MutableProfiles(Map.of("terra", profileWithPattern("usage limit")));
|
||||
LiveExhaustedPatterns patterns = new LiveExhaustedPatterns(live::get);
|
||||
|
||||
Pattern first = patterns.patternFor("terra");
|
||||
live.set(Map.of("terra", profileWithPattern("rate limit")));
|
||||
Pattern second = patterns.patternFor("terra");
|
||||
|
||||
assertNotSame(first, second, "changed pattern text must produce a freshly compiled Pattern");
|
||||
}
|
||||
|
||||
@Test
|
||||
void unknownProfileReturnsNullAndUnarmed() {
|
||||
LiveExhaustedPatterns patterns = new LiveExhaustedPatterns(Map::of);
|
||||
assertNull(patterns.patternFor("nope"));
|
||||
assertFalse(patterns.armed("nope"));
|
||||
}
|
||||
|
||||
@Test
|
||||
void nullProfileNameReturnsNullAndUnarmed() {
|
||||
LiveExhaustedPatterns patterns = new LiveExhaustedPatterns(
|
||||
() -> Map.of("terra", profileWithPattern("usage limit")));
|
||||
assertNull(patterns.patternFor(null));
|
||||
assertFalse(patterns.armed(null));
|
||||
}
|
||||
|
||||
@Test
|
||||
void profileWithNoPatternConfiguredReturnsNull() {
|
||||
LiveExhaustedPatterns patterns = new LiveExhaustedPatterns(
|
||||
() -> Map.of("terra", profileWithPattern(null)));
|
||||
assertNull(patterns.patternFor("terra"));
|
||||
assertFalse(patterns.armed("terra"));
|
||||
}
|
||||
}
|
||||
@@ -0,0 +1,72 @@
|
||||
package dev.ltms.fleet.mcp;
|
||||
|
||||
import dev.ltms.fleet.config.FleetConfig;
|
||||
import java.nio.file.Files;
|
||||
import java.nio.file.Path;
|
||||
import java.util.LinkedHashSet;
|
||||
import java.util.Set;
|
||||
import java.util.regex.Matcher;
|
||||
import java.util.regex.Pattern;
|
||||
import org.junit.jupiter.api.DisplayName;
|
||||
import org.junit.jupiter.api.Test;
|
||||
import org.junit.jupiter.api.io.TempDir;
|
||||
|
||||
import static org.junit.jupiter.api.Assertions.assertTrue;
|
||||
|
||||
/** fleetd #464: launch charters must not name MCP tools the server does not register. */
|
||||
class CharterToolSurfaceTest {
|
||||
|
||||
private static final Path MCP_SOURCE = Path.of("src/main/java/dev/ltms/fleet/mcp/FleetMcp.java");
|
||||
|
||||
private static Set<String> matches(String text, String regex) {
|
||||
Matcher m = Pattern.compile(regex).matcher(text);
|
||||
Set<String> found = new LinkedHashSet<>();
|
||||
while (m.find()) {
|
||||
found.add(m.group(1));
|
||||
}
|
||||
return found;
|
||||
}
|
||||
|
||||
/** Every {@code fleet_*} or legacy {@code bridge_*} token in the configured launch charters. */
|
||||
private static Set<String> toolsNamedIn(FleetConfig config) {
|
||||
return matches(String.join("\n", config.fleet().charters().values()),
|
||||
"(fleet_[a-z_]+|bridge_[a-z_]+)");
|
||||
}
|
||||
|
||||
/** Every tool {@link FleetMcp} registers, read from its {@code tool("…")} calls. */
|
||||
private static Set<String> toolsTheServerRegisters() throws Exception {
|
||||
return matches(Files.readString(MCP_SOURCE), "tool\\(\\\"(fleet_[a-z_]+)\\\"");
|
||||
}
|
||||
|
||||
@Test
|
||||
@DisplayName("[SOURCE TEXT] every tool named in a configured charter is registered by the server")
|
||||
void configuredChartersNameOnlyRegisteredTools(@TempDir Path dir) throws Exception {
|
||||
Path configFile = dir.resolve("charters.yaml");
|
||||
Files.writeString(configFile, """
|
||||
fleet:
|
||||
charters:
|
||||
dev: |
|
||||
Send the final handoff through fleet_reply.
|
||||
reviewer: |
|
||||
Use fleet_ask only for the lead's decision.
|
||||
""");
|
||||
|
||||
FleetConfig config = FleetConfig.load(configFile);
|
||||
Set<String> named = toolsNamedIn(config);
|
||||
Set<String> registered = toolsTheServerRegisters();
|
||||
|
||||
assertTrue(!named.isEmpty(),
|
||||
"the charter fixture named no fleet_* or bridge_* tool. This test would check nothing; "
|
||||
+ "add charter text that names a tool before changing the extraction.");
|
||||
assertTrue(!registered.isEmpty(),
|
||||
"the FleetMcp registration scrape found no tools. This test would check nothing; "
|
||||
+ "repair the tool(\"…\") extraction before changing the assertion.");
|
||||
|
||||
Set<String> unknown = new LinkedHashSet<>(named);
|
||||
unknown.removeAll(registered);
|
||||
assertTrue(unknown.isEmpty(),
|
||||
"configured charter text names " + unknown + ", but FleetMcp does not register it. "
|
||||
+ "Checked " + named + " against " + registered + ". Fix the charter text or "
|
||||
+ "register the tool; do NOT weaken this test.");
|
||||
}
|
||||
}
|
||||
@@ -187,6 +187,71 @@ class FleetMcpAuthzTest {
|
||||
"no CallerResolver supplied ⇒ authorization not enforced (legacy behaviour)");
|
||||
}
|
||||
|
||||
// --- fleetd #439: who may see fleet_list's coordinator row ----------------------------------
|
||||
|
||||
/**
|
||||
* fleetd #439: {@link FleetMcp#coordinatorVisibleTo} is the whole policy decision for
|
||||
* {@code fleet_list}'s {@code coordinator} row — lead-to-lead coordination state, not roster
|
||||
* observation. Only the primary may see it; a worker, an architect, and (the case the previous
|
||||
* pass of this ticket did not cover) an anonymous caller must all be refused.
|
||||
*/
|
||||
@Test
|
||||
void onlyThePrimaryMaySeeTheCoordinatorRow() {
|
||||
assertTrue(FleetMcp.coordinatorVisibleTo(PRIMARY), "the primary must see its own coordination state");
|
||||
assertFalse(FleetMcp.coordinatorVisibleTo(WORKER_A), "a worker must not see lead-to-lead coordination state");
|
||||
assertFalse(FleetMcp.coordinatorVisibleTo(ARCH_DESIGN),
|
||||
"an architect holds READ today, but that must not extend to coordinator");
|
||||
assertFalse(FleetMcp.coordinatorVisibleTo(ANON), "authenticated as nothing must not see it either");
|
||||
}
|
||||
|
||||
/**
|
||||
* fleetd #439 / PR #462 review finding M2: the predicate above can be perfectly correct while
|
||||
* the one production call site (the {@code fleet_list} MCP handler) never actually asks it —
|
||||
* a literal {@code true} compiles, and the whole suite stayed green under that mutation because
|
||||
* every existing test drives {@link FleetMcp#listFleet} directly and supplies the boolean
|
||||
* itself. This test reads {@code FleetMcp.java}'s own source (same idiom as {@link
|
||||
* #toolsTheServerRegisters()} / {@link #everyRegisteredToolHasItsHandlerActionPinned()}) and
|
||||
* asserts the handler's call passes {@code coordinatorVisibleTo(principal(exchange))} — not a
|
||||
* literal {@code true} or {@code false} — as {@code listFleet}'s trailing argument.
|
||||
*
|
||||
* <p>Anchored on argument position, not a bare substring search: {@code true} appears many
|
||||
* times elsewhere in this file for unrelated reasons, so a plain {@code contains("true")}
|
||||
* check would prove nothing. The pattern requires the literal text immediately before the
|
||||
* closing {@code );} of the {@code listFleet(} call inside the handler block to be exactly
|
||||
* {@code coordinatorVisibleTo(principal(exchange))}.
|
||||
*/
|
||||
@Test
|
||||
void theFleetListHandlerActuallyConsultsCoordinatorVisibleTo() throws Exception {
|
||||
String source = Files.readString(MCP_SOURCE);
|
||||
|
||||
// Isolate the fleet_list handler block: from its declaration up to the next handler's
|
||||
// declaration. A change to variable naming would break this scrape loudly (see the control
|
||||
// assertion just below), rather than silently reporting "no violation found".
|
||||
int start = source.indexOf("listHandler =");
|
||||
assertTrue(start >= 0, "could not find the fleet_list handler (listHandler) in " + MCP_SOURCE
|
||||
+ " -- the scrape has stopped matching, fix the anchor before trusting this test");
|
||||
int end = source.indexOf("stopHandler =", start);
|
||||
assertTrue(end > start, "could not find the handler declared after listHandler to bound the scrape");
|
||||
String handlerBlock = source.substring(start, end);
|
||||
|
||||
// CONTROL: the block we scraped really does contain a call to listFleet(...) -- if this
|
||||
// fails, the anchors above moved and the assertion below would otherwise pass on nothing.
|
||||
assertTrue(handlerBlock.contains("listFleet("),
|
||||
"control failed: the scraped listHandler block contains no listFleet( call at all -- "
|
||||
+ "the anchors have drifted, this test is not testing what it claims to");
|
||||
|
||||
Pattern trailingArg = Pattern.compile(
|
||||
"listFleet\\([^;]*?,\\s*(coordinatorVisibleTo\\(principal\\(exchange\\)\\)|true|false)\\s*\\)\\s*;",
|
||||
Pattern.DOTALL);
|
||||
Matcher m = trailingArg.matcher(handlerBlock);
|
||||
assertTrue(m.find(), "could not locate listFleet(...)'s trailing boolean argument in the "
|
||||
+ "listHandler block -- the call shape changed, update this test's anchor: " + handlerBlock);
|
||||
String trailing = m.group(1);
|
||||
assertEquals("coordinatorVisibleTo(principal(exchange))", trailing,
|
||||
"the fleet_list handler must ask coordinatorVisibleTo(principal(exchange)) who is "
|
||||
+ "calling, not pass a literal boolean -- found: " + trailing);
|
||||
}
|
||||
|
||||
// --- which action each tool hands the gate (fleetd #272) ------------------------------------
|
||||
|
||||
/**
|
||||
@@ -201,14 +266,34 @@ class FleetMcpAuthzTest {
|
||||
*/
|
||||
@Test
|
||||
void pollingByTargetIsADrainAndPollingByTicketIsARead() {
|
||||
assertEquals(Authz.Action.DRAIN, FleetMcp.pollAction("term_b"),
|
||||
assertEquals(Authz.Action.DRAIN, FleetMcp.pollAction("term_b", null),
|
||||
"poll by target removes the replies — that is a drain, not an observation");
|
||||
assertEquals(Authz.Action.READ, FleetMcp.pollAction(null),
|
||||
assertEquals(Authz.Action.READ, FleetMcp.pollAction(null, null),
|
||||
"poll by ticket changes nothing");
|
||||
assertEquals(Authz.Action.READ, FleetMcp.pollAction(" "),
|
||||
assertEquals(Authz.Action.READ, FleetMcp.pollAction(" ", null),
|
||||
"a blank target is an absent target");
|
||||
}
|
||||
|
||||
/**
|
||||
* fleetd #421: a coordId branch is a THIRD operation behind fleet_poll's one name, and it must
|
||||
* map to {@link Authz.Action#COORD_READ} — never {@link Authz.Action#READ}, even though this
|
||||
* branch also consumes nothing. READ's grant is open to every authenticated role on the premise
|
||||
* that the roster carries no secrets; a lead-to-lead body is not the roster, so folding this
|
||||
* branch into READ would let any worker read every peer lead's mail in full. coordId also takes
|
||||
* priority over target when both happen to be present — it addresses a different inbox entirely.
|
||||
*/
|
||||
@Test
|
||||
void pollingByCoordIdIsACoordReadNeverAPlainRead() {
|
||||
assertEquals(Authz.Action.COORD_READ, FleetMcp.pollAction(null, "mac-opus"),
|
||||
"reading held peer mail must not be mapped to the everyone-readable READ action");
|
||||
assertEquals(Authz.Action.COORD_READ, FleetMcp.pollAction(" ", "mac-opus"),
|
||||
"a blank target must not fall through to READ/DRAIN when coordId is present");
|
||||
assertEquals(Authz.Action.READ, FleetMcp.pollAction(null, " "),
|
||||
"a blank coordId is an absent coordId, same as target/ticket");
|
||||
assertEquals(Authz.Action.COORD_READ, FleetMcp.pollAction("term_b", "mac-opus"),
|
||||
"coordId takes priority over target — this is a different inbox, not a drain");
|
||||
}
|
||||
|
||||
@Test
|
||||
void everyRegisteredToolHasItsHandlerActionPinned() {
|
||||
Set<String> registered = toolsTheServerRegisters();
|
||||
@@ -230,6 +315,8 @@ class FleetMcpAuthzTest {
|
||||
assertEquals(Authz.Action.READ, FleetMcp.toolAction("fleet_whoami", Map.of()));
|
||||
assertEquals(Authz.Action.READ, FleetMcp.toolAction("fleet_poll", Map.of("ticket", "task")));
|
||||
assertEquals(Authz.Action.DRAIN, FleetMcp.toolAction("fleet_poll", Map.of("target", "term_b")));
|
||||
assertEquals(Authz.Action.COORD_READ,
|
||||
FleetMcp.toolAction("fleet_poll", Map.of("coordId", "mac-opus")));
|
||||
}
|
||||
|
||||
private static Set<String> toolsTheServerRegisters() {
|
||||
@@ -249,11 +336,11 @@ class FleetMcpAuthzTest {
|
||||
void aWorkerMayNotDrainAnotherSessionsInboxByPolling() {
|
||||
FleetMcp m = mcp(true);
|
||||
|
||||
assertNotNull(m.denyFor(WORKER_A, FleetMcp.pollAction("term_b"), "term_b"),
|
||||
assertNotNull(m.denyFor(WORKER_A, FleetMcp.pollAction("term_b", null), "term_b"),
|
||||
"a worker draining a peer's inbox would destroy replies queued for the primary");
|
||||
assertNotNull(m.denyFor(ARCH_DESIGN, FleetMcp.pollAction("term_b"), "term_b"),
|
||||
assertNotNull(m.denyFor(ARCH_DESIGN, FleetMcp.pollAction("term_b", null), "term_b"),
|
||||
"an architect has no lifecycle rights either — same gate as fleet_ack");
|
||||
assertNull(m.denyFor(PRIMARY, FleetMcp.pollAction("term_b"), "term_b"),
|
||||
assertNull(m.denyFor(PRIMARY, FleetMcp.pollAction("term_b", null), "term_b"),
|
||||
"collecting a held reply is the primary's job");
|
||||
}
|
||||
|
||||
@@ -265,11 +352,33 @@ class FleetMcpAuthzTest {
|
||||
void pollingAnOwnTicketStaysOpenToWorkersAndArchitects() {
|
||||
FleetMcp m = mcp(true);
|
||||
|
||||
assertNull(m.denyFor(WORKER_A, FleetMcp.pollAction(null), null));
|
||||
assertNull(m.denyFor(ARCH_DESIGN, FleetMcp.pollAction(null), null),
|
||||
assertNull(m.denyFor(WORKER_A, FleetMcp.pollAction(null, null), null));
|
||||
assertNull(m.denyFor(ARCH_DESIGN, FleetMcp.pollAction(null, null), null),
|
||||
"an architect delegates with wait:false, so it must be able to poll its ticket");
|
||||
}
|
||||
|
||||
/**
|
||||
* fleetd #421 acceptance: the whole ticket, pinned at the policy-table layer. A worker must be
|
||||
* refused the coordId branch, the primary must be allowed, and (CB-548) an architect — which
|
||||
* holds READ today — must be refused too, because "not primary" means not-architect here.
|
||||
*/
|
||||
@Test
|
||||
void aWorkerAndAnArchitectMayNotReadHeldPeerMailOnlyThePrimaryMay() {
|
||||
FleetMcp m = mcp(true);
|
||||
Authz.Action coordRead = FleetMcp.pollAction(null, "mac-opus");
|
||||
|
||||
McpSchema.CallToolResult workerDenied = m.denyFor(WORKER_A, coordRead, null);
|
||||
assertNotNull(workerDenied, "a worker must not read held lead-to-lead mail");
|
||||
assertTrue(workerDenied.isError());
|
||||
|
||||
McpSchema.CallToolResult archDenied = m.denyFor(ARCH_DESIGN, coordRead, null);
|
||||
assertNotNull(archDenied, "an architect holds READ today, but not-primary must mean "
|
||||
+ "not-architect here too");
|
||||
assertTrue(archDenied.isError());
|
||||
|
||||
assertNull(m.denyFor(PRIMARY, coordRead, null), "reading its own held mail is the primary's job");
|
||||
}
|
||||
|
||||
// --- identity reconstruction from the transport context ------------------------------------
|
||||
|
||||
@Test
|
||||
|
||||
@@ -3,6 +3,7 @@ package dev.ltms.fleet.mcp;
|
||||
import dev.ltms.fleet.auth.CallerResolver;
|
||||
import dev.ltms.fleet.auth.MemberRegistry;
|
||||
import dev.ltms.fleet.auth.Principal;
|
||||
import dev.ltms.fleet.auth.Role;
|
||||
import dev.ltms.fleet.config.FleetConfig;
|
||||
import dev.ltms.fleet.guard.SubscriptionGuard;
|
||||
import dev.ltms.fleet.herdr.AgentControl;
|
||||
@@ -28,6 +29,7 @@ import dev.ltms.fleet.placement.BackendQuarantine;
|
||||
import dev.ltms.fleet.placement.PlacementPolicies;
|
||||
import io.modelcontextprotocol.spec.McpSchema;
|
||||
import dev.ltms.fleet.msg.InMemoryReplyInbox;
|
||||
import dev.ltms.fleet.msg.ReplyInbox;
|
||||
import org.junit.jupiter.api.BeforeEach;
|
||||
import org.junit.jupiter.api.Test;
|
||||
|
||||
@@ -77,7 +79,7 @@ class FleetMcpTest {
|
||||
Thread.sleep(5);
|
||||
}
|
||||
assertTrue(rendezvous.isWaiting(target), "send should be accepted for " + target);
|
||||
FleetMcp.reply(messages, target, "received");
|
||||
FleetMcp.reply(messages, target, Role.WORKER, "received");
|
||||
assertEquals("received", textOf(send.get(6, TimeUnit.SECONDS)));
|
||||
}
|
||||
|
||||
@@ -110,7 +112,7 @@ class FleetMcpTest {
|
||||
|
||||
// fleetd #365: a resolved live send must read distinctly from a merely-queued reply —
|
||||
// see replyWithNoPendingSendIsQueuedNotError below for the other case.
|
||||
McpSchema.CallToolResult reply = FleetMcp.reply(messages, "term_a", "LGTM");
|
||||
McpSchema.CallToolResult reply = FleetMcp.reply(messages, "term_a", Role.WORKER, "LGTM");
|
||||
assertEquals(MessageService.ReplyOutcome.RESOLVED_SEND.description(), textOf(reply));
|
||||
|
||||
McpSchema.CallToolResult res = send.get(6, TimeUnit.SECONDS);
|
||||
@@ -136,7 +138,7 @@ class FleetMcpTest {
|
||||
}
|
||||
assertTrue(rendezvous.isWaiting("term_a"), "send should have opened its waiter");
|
||||
|
||||
McpSchema.CallToolResult reply = FleetMcp.reply(messages, "term_a", "async LGTM");
|
||||
McpSchema.CallToolResult reply = FleetMcp.reply(messages, "term_a", Role.WORKER, "async LGTM");
|
||||
assertEquals(MessageService.ReplyOutcome.RESOLVED_SEND.description(), textOf(reply));
|
||||
|
||||
// Poll until the async send completes and reports the reply.
|
||||
@@ -183,7 +185,7 @@ class FleetMcpTest {
|
||||
Thread.sleep(5);
|
||||
}
|
||||
assertTrue(rendezvous.isWaiting("term_a"));
|
||||
FleetMcp.reply(messages, "term_a", "done");
|
||||
FleetMcp.reply(messages, "term_a", Role.WORKER, "done");
|
||||
assertEquals("done", textOf(answer.get(6, TimeUnit.SECONDS)));
|
||||
|
||||
McpSchema.CallToolResult done = FleetMcp.poll(messages, ticket, null);
|
||||
@@ -215,7 +217,7 @@ class FleetMcpTest {
|
||||
// fleet_poll{ticket} stuck PENDING forever and later force-failed with a false "session
|
||||
// released before it replied" reason. This used to land in the inbox instead (see the old
|
||||
// assertion this replaced: messages.drainReplies("term_a").getFirst()...) — that was the bug.
|
||||
FleetMcp.reply(messages, "term_a", "finished after timeout");
|
||||
FleetMcp.reply(messages, "term_a", Role.WORKER, "finished after timeout");
|
||||
assertEquals("finished after timeout", textOf(FleetMcp.poll(messages, ticket, null)));
|
||||
assertTrue(messages.drainReplies("term_a").isEmpty(),
|
||||
"the reply completed its own ticket directly and never touched the inbox");
|
||||
@@ -269,7 +271,7 @@ class FleetMcpTest {
|
||||
}
|
||||
assertEquals(MessageService.Phase.FAILED, second.phase());
|
||||
|
||||
FleetMcp.reply(messages, "term_a", "late reply");
|
||||
FleetMcp.reply(messages, "term_a", Role.WORKER, "late reply");
|
||||
assertEquals("late reply", messages.drainReplies("term_a").getFirst().content());
|
||||
|
||||
CompletableFuture<McpSchema.CallToolResult> answer = CompletableFuture.supplyAsync(
|
||||
@@ -278,7 +280,7 @@ class FleetMcpTest {
|
||||
while (!rendezvous.isWaiting("term_a") && System.currentTimeMillis() < deadline) {
|
||||
Thread.sleep(5);
|
||||
}
|
||||
FleetMcp.reply(messages, "term_a", "done");
|
||||
FleetMcp.reply(messages, "term_a", Role.WORKER, "done");
|
||||
assertEquals("done", textOf(answer.get(6, TimeUnit.SECONDS)));
|
||||
}
|
||||
|
||||
@@ -331,7 +333,7 @@ class FleetMcpTest {
|
||||
void replyWithNoPendingSendIsQueuedNotError() {
|
||||
// CB-307: a reply with no open send is now queued in the inbox, not an error.
|
||||
// fleetd #365: it must also no longer claim "delivered" — nothing was waiting for it.
|
||||
McpSchema.CallToolResult res = FleetMcp.reply(messages, "term_a", "orphan");
|
||||
McpSchema.CallToolResult res = FleetMcp.reply(messages, "term_a", Role.WORKER, "orphan");
|
||||
assertNotEquals(Boolean.TRUE, res.isError(), "a queued reply is not an error");
|
||||
assertEquals(MessageService.ReplyOutcome.QUEUED.description(), textOf(res));
|
||||
|
||||
@@ -341,6 +343,40 @@ class FleetMcpTest {
|
||||
assertEquals("orphan", drained.getFirst().content());
|
||||
}
|
||||
|
||||
@Test
|
||||
void replyFromLeadIsRefusedBeforeItCanPublishToTheWorkerInbox() {
|
||||
ReplyInbox inboxThatRejectsPublishes = new ReplyInbox() {
|
||||
@Override public void own(String target) { }
|
||||
@Override public void release(String target) { }
|
||||
@Override public void publish(String target, String msgId, String content) {
|
||||
fail("a lead fleet_reply must not publish to the worker inbox");
|
||||
}
|
||||
@Override public List<InboxMessage> peek(String target) { return List.of(); }
|
||||
@Override public boolean ack(String target, String msgId) { return false; }
|
||||
};
|
||||
MessageService leadMessages = new MessageService(agents, new Injector(agents), new Rendezvous(),
|
||||
inboxThatRejectsPublishes);
|
||||
|
||||
McpSchema.CallToolResult res = assertDoesNotThrow(
|
||||
() -> FleetMcp.reply(leadMessages, "term_lead", Role.PRIMARY, "peer reply"));
|
||||
|
||||
assertTrue(res.isError());
|
||||
assertEquals("fleet_reply has no route to a peer lead. Use fleet_send{coordId: ...} for a peer on another "
|
||||
+ "daemon or fleet_send{sessionId: ...} for a peer on this host. fleet_reply resolves a member's "
|
||||
+ "blocked fleet_send, and a peer's coord-id message is durable and non-blocking, so there is "
|
||||
+ "nothing for it to resolve.",
|
||||
textOf(res));
|
||||
}
|
||||
|
||||
@Test
|
||||
void replyFromUnidentifiedCallerKeepsItsOwnError() {
|
||||
McpSchema.CallToolResult res = FleetMcp.reply(messages, null, Role.PRIMARY, "reply");
|
||||
|
||||
assertTrue(res.isError());
|
||||
assertEquals("fleet_reply is for workers only — could not identify the calling worker from the connection",
|
||||
textOf(res));
|
||||
}
|
||||
|
||||
@Test
|
||||
void replyWithBlankContentIsACleanToolErrorNotAnUncaughtException() {
|
||||
// fleetd #302: MessageService.reply now REJECTS blank content by throwing. fleet_reply's
|
||||
@@ -351,7 +387,7 @@ class FleetMcpTest {
|
||||
// has always used isBlank for exactly this reason.
|
||||
for (String blank : new String[] {null, "", " ", "\n\t"}) {
|
||||
McpSchema.CallToolResult res = assertDoesNotThrow(
|
||||
() -> FleetMcp.reply(messages, "term_a", blank),
|
||||
() -> FleetMcp.reply(messages, "term_a", Role.WORKER, blank),
|
||||
"blank content must be refused as a tool error, never thrown out of the handler");
|
||||
assertEquals(Boolean.TRUE, res.isError(), "blank content is an error result");
|
||||
assertTrue(textOf(res).contains("content is required"),
|
||||
@@ -364,7 +400,7 @@ class FleetMcpTest {
|
||||
@Test
|
||||
void bridgePollWithTargetDrainsReplies() {
|
||||
// A reply with no open send queues it in the inbox.
|
||||
FleetMcp.reply(messages, "term_a", "queued-msg");
|
||||
FleetMcp.reply(messages, "term_a", Role.WORKER, "queued-msg");
|
||||
|
||||
// fleet_poll with target drains the inbox.
|
||||
McpSchema.CallToolResult res = FleetMcp.poll(messages, null, "term_a");
|
||||
@@ -415,7 +451,7 @@ class FleetMcpTest {
|
||||
Thread.sleep(5);
|
||||
}
|
||||
assertTrue(rendezvous.isWaiting("term_a"), "the answer should have reopened a waiter");
|
||||
McpSchema.CallToolResult reply = FleetMcp.reply(messages, "term_a", "done");
|
||||
McpSchema.CallToolResult reply = FleetMcp.reply(messages, "term_a", Role.WORKER, "done");
|
||||
assertEquals(MessageService.ReplyOutcome.RESOLVED_SEND.description(), textOf(reply));
|
||||
assertEquals("done", textOf(answer.get(6, TimeUnit.SECONDS)));
|
||||
}
|
||||
@@ -590,8 +626,8 @@ class FleetMcpTest {
|
||||
McpSchema.CallToolResult res = FleetMcp.listFleet(
|
||||
workerService(h, "http://gx00.gw:8000", Set.of("gx00.gw")), sessions, null,
|
||||
FleetMcp.CapacitySource.none(), new FleetMcp.HealthCoverageSource(() -> "off"),
|
||||
FleetMcp.QuarantineSource.none(), Map.of(), "",
|
||||
new FleetMcp.CoordinationSource(channel, List.of()));
|
||||
FleetMcp.QuarantineSource.none(), FleetMcp.OutageSource.none(), FleetMcp.LeadSeatSource.none(),
|
||||
Map.of(), "", new FleetMcp.CoordinationSource(channel, List.of()), true);
|
||||
|
||||
String out = textOf(res);
|
||||
// fleetd #361: reports both which coord-id a peer must use to reach ME, and this daemon's
|
||||
@@ -619,8 +655,8 @@ class FleetMcpTest {
|
||||
McpSchema.CallToolResult res = FleetMcp.listFleet(
|
||||
workerService(h, "http://gx00.gw:8000", Set.of("gx00.gw")), sessions, null,
|
||||
FleetMcp.CapacitySource.none(), new FleetMcp.HealthCoverageSource(() -> "off"),
|
||||
FleetMcp.QuarantineSource.none(), Map.of(), "",
|
||||
new FleetMcp.CoordinationSource(channel, List.of()));
|
||||
FleetMcp.QuarantineSource.none(), FleetMcp.OutageSource.none(), FleetMcp.LeadSeatSource.none(),
|
||||
Map.of(), "", new FleetMcp.CoordinationSource(channel, List.of()), true);
|
||||
|
||||
String out = textOf(res);
|
||||
assertTrue(out.contains("\"mailbox\":{\"status\":\"unknown\"}"), out);
|
||||
@@ -640,13 +676,20 @@ class FleetMcpTest {
|
||||
"an ordinary fleet's output must be unchanged by this feature");
|
||||
}
|
||||
|
||||
/**
|
||||
* fleetd #463: a compat overload called with no {@code callerIsPrimary} argument at all must
|
||||
* fail closed, not open. Before this fix the hidden default was {@code true}, so a caller that
|
||||
* forgot the argument silently got lead-to-lead coordination state. Lead coordination is fully
|
||||
* configured here (a real channel, a real mailbox) specifically so this is not conflated with
|
||||
* {@link #listOmitsTheCoordinatorRowWhenLeadCoordinationIsOff} -- the row is capable of being
|
||||
* assembled, and the missing argument is the only reason it is not.
|
||||
*/
|
||||
@Test
|
||||
void listReportsHeldMessagesWithATruncatedPreviewNeverTheFullBody() {
|
||||
void listCompatOverloadWithNoCallerIsPrimaryArgumentOmitsTheCoordinatorKey() {
|
||||
FakeHerdr h = new FakeHerdr();
|
||||
SessionManager sessions = new SessionManager(workerService(h, "http://gx00.gw:8000", Set.of("gx00.gw")));
|
||||
String longContent = "x".repeat(200);
|
||||
FakeLeadChannel channel = new FakeLeadChannel("mac-opus")
|
||||
.hold(new LeadMessage("m1", "fleet01-lead", "mac-opus", longContent));
|
||||
.withMailbox("mac-opus", LeadChannel.MailboxState.exists("mac-opus", 0, 1));
|
||||
|
||||
McpSchema.CallToolResult res = FleetMcp.listFleet(
|
||||
workerService(h, "http://gx00.gw:8000", Set.of("gx00.gw")), sessions, null,
|
||||
@@ -654,13 +697,226 @@ class FleetMcpTest {
|
||||
FleetMcp.QuarantineSource.none(), Map.of(), "",
|
||||
new FleetMcp.CoordinationSource(channel, List.of()));
|
||||
|
||||
String out = textOf(res);
|
||||
assertFalse(out.contains("\"coordinator\""),
|
||||
"no callerIsPrimary argument must fail closed (absent), not open (present): " + out);
|
||||
}
|
||||
|
||||
/**
|
||||
* fleetd #439: a worker calling {@code fleet_list} must get a result with the {@code
|
||||
* coordinator} key <strong>absent</strong> -- not an empty object, not a redacted one -- even
|
||||
* though lead coordination is fully configured and would otherwise report a row. This drives
|
||||
* the same {@code callerIsPrimary} value the MCP handler computes ({@code
|
||||
* Principal.worker(...).isPrimary()}), so it pins the real production boolean, not a literal.
|
||||
*/
|
||||
@Test
|
||||
void listOmitsTheCoordinatorKeyEntirelyForAWorkerEvenWhenLeadCoordinationIsOn() {
|
||||
FakeHerdr h = new FakeHerdr();
|
||||
SessionManager sessions = new SessionManager(workerService(h, "http://gx00.gw:8000", Set.of("gx00.gw")));
|
||||
FakeLeadChannel channel = new FakeLeadChannel("mac-opus")
|
||||
.withMailbox("mac-opus", LeadChannel.MailboxState.exists("mac-opus", 0, 1));
|
||||
boolean callerIsPrimary = Principal.worker("term_a", 1).isPrimary();
|
||||
|
||||
McpSchema.CallToolResult res = FleetMcp.listFleet(
|
||||
workerService(h, "http://gx00.gw:8000", Set.of("gx00.gw")), sessions, null,
|
||||
FleetMcp.CapacitySource.none(), new FleetMcp.HealthCoverageSource(() -> "off"),
|
||||
FleetMcp.QuarantineSource.none(), FleetMcp.OutageSource.none(), FleetMcp.LeadSeatSource.none(),
|
||||
Map.of(), "", new FleetMcp.CoordinationSource(channel, List.of()), callerIsPrimary);
|
||||
|
||||
String out = textOf(res);
|
||||
assertFalse(out.contains("\"coordinator\""), "a worker must never see the coordinator key at all: " + out);
|
||||
assertFalse(out.contains("mac-opus"), "no fragment of the coordinator row may leak either: " + out);
|
||||
assertTrue(out.contains("\"leads\""), "the rest of the result must still be present: " + out);
|
||||
assertTrue(out.contains("\"members\""), out);
|
||||
assertTrue(out.contains("\"healthCoverage\""), out);
|
||||
}
|
||||
|
||||
/**
|
||||
* fleetd #439 acceptance criterion 2: an architect gets exactly the same treatment as a worker.
|
||||
* This is a real, executed test (not just reasoning by analogy) -- it drives the actual
|
||||
* {@code Principal.architect(...).isPrimary()} value the production handler would compute for
|
||||
* an architect caller, through the same gate a worker's call goes through.
|
||||
*/
|
||||
@Test
|
||||
void listOmitsTheCoordinatorKeyEntirelyForAnArchitectToo() {
|
||||
FakeHerdr h = new FakeHerdr();
|
||||
SessionManager sessions = new SessionManager(workerService(h, "http://gx00.gw:8000", Set.of("gx00.gw")));
|
||||
FakeLeadChannel channel = new FakeLeadChannel("mac-opus")
|
||||
.withMailbox("mac-opus", LeadChannel.MailboxState.exists("mac-opus", 0, 1));
|
||||
boolean callerIsPrimary = Principal.architect("lead-designer", "term_design", 400).isPrimary();
|
||||
|
||||
McpSchema.CallToolResult res = FleetMcp.listFleet(
|
||||
workerService(h, "http://gx00.gw:8000", Set.of("gx00.gw")), sessions, null,
|
||||
FleetMcp.CapacitySource.none(), new FleetMcp.HealthCoverageSource(() -> "off"),
|
||||
FleetMcp.QuarantineSource.none(), FleetMcp.OutageSource.none(), FleetMcp.LeadSeatSource.none(),
|
||||
Map.of(), "", new FleetMcp.CoordinationSource(channel, List.of()), callerIsPrimary);
|
||||
|
||||
String out = textOf(res);
|
||||
assertFalse(out.contains("\"coordinator\""), "an architect must never see the coordinator key either: " + out);
|
||||
}
|
||||
|
||||
/**
|
||||
* fleetd #439 acceptance criterion 3: an explicitly-primary caller sees the coordinator row
|
||||
* fully assembled, with the same content #439 always produced for a primary.
|
||||
*
|
||||
* <p>fleetd #463 flipped the compat overloads' hidden default from {@code true} to
|
||||
* {@code false} (fail closed), so the old "pre-#439 overload" this test used to compare
|
||||
* against no longer stands in for a primary caller -- it is now exactly the implicit-default
|
||||
* path #463 closes. Verifying the primary path means calling the canonical overload with an
|
||||
* explicit {@code callerIsPrimary=true} directly, as the production {@code fleet_list} handler
|
||||
* does.
|
||||
*/
|
||||
@Test
|
||||
void listIsByteForByteUnchangedForThePrimaryCaller() {
|
||||
FakeHerdr h = new FakeHerdr();
|
||||
SessionManager sessions = new SessionManager(workerService(h, "http://gx00.gw:8000", Set.of("gx00.gw")));
|
||||
FakeLeadChannel channel = new FakeLeadChannel("mac-opus")
|
||||
.withMailbox("mac-opus", LeadChannel.MailboxState.exists("mac-opus", 0, 1));
|
||||
|
||||
String gatedAsPrimary = textOf(FleetMcp.listFleet(
|
||||
workerService(h, "http://gx00.gw:8000", Set.of("gx00.gw")), sessions, null,
|
||||
FleetMcp.CapacitySource.none(), new FleetMcp.HealthCoverageSource(() -> "off"),
|
||||
FleetMcp.QuarantineSource.none(), FleetMcp.OutageSource.none(), FleetMcp.LeadSeatSource.none(),
|
||||
Map.of(), "", new FleetMcp.CoordinationSource(channel, List.of()), true));
|
||||
|
||||
assertTrue(gatedAsPrimary.contains("\"coordinator\""), gatedAsPrimary);
|
||||
assertTrue(gatedAsPrimary.contains("\"selfId\":\"mac-opus\""), gatedAsPrimary);
|
||||
assertTrue(gatedAsPrimary.contains("\"mailbox\""), gatedAsPrimary);
|
||||
assertTrue(gatedAsPrimary.contains("\"heldCount\""), gatedAsPrimary);
|
||||
assertTrue(gatedAsPrimary.contains("\"heldDurable\""), gatedAsPrimary);
|
||||
assertTrue(gatedAsPrimary.contains("\"held\""), gatedAsPrimary);
|
||||
assertTrue(gatedAsPrimary.contains("\"peers\""), gatedAsPrimary);
|
||||
}
|
||||
|
||||
@Test
|
||||
void listReportsHeldMessagesWithATruncatedPreviewNeverTheFullBody() {
|
||||
FakeHerdr h = new FakeHerdr();
|
||||
SessionManager sessions = new SessionManager(workerService(h, "http://gx00.gw:8000", Set.of("gx00.gw")));
|
||||
// A homogeneous "x".repeat(200) body would not pin the exact 80-char boundary: a preview
|
||||
// widened by one (81 chars) still CONTAINS "x".repeat(80) + "…" as a substring one position
|
||||
// later, because every character is 'x'. Put a sentinel ("Y") exactly at index 80 — the
|
||||
// first character a widened cap would leak — so any preview past 80 chars is caught.
|
||||
String longContent = "x".repeat(80) + "Y" + "z".repeat(119);
|
||||
FakeLeadChannel channel = new FakeLeadChannel("mac-opus")
|
||||
.hold(new LeadMessage("m1", "fleet01-lead", "mac-opus", longContent));
|
||||
|
||||
McpSchema.CallToolResult res = FleetMcp.listFleet(
|
||||
workerService(h, "http://gx00.gw:8000", Set.of("gx00.gw")), sessions, null,
|
||||
FleetMcp.CapacitySource.none(), new FleetMcp.HealthCoverageSource(() -> "off"),
|
||||
FleetMcp.QuarantineSource.none(), FleetMcp.OutageSource.none(), FleetMcp.LeadSeatSource.none(),
|
||||
Map.of(), "", new FleetMcp.CoordinationSource(channel, List.of()), true);
|
||||
|
||||
String out = textOf(res);
|
||||
assertTrue(out.contains("\"msgId\":\"m1\""), out);
|
||||
assertTrue(out.contains("\"from\":\"fleet01-lead\""), out);
|
||||
assertFalse(out.contains(longContent), "fleet_list must never dump a held message's full body: " + out);
|
||||
assertFalse(out.contains("Y"),
|
||||
"the sentinel at index 80 must never appear — a preview past 80 chars leaked it: " + out);
|
||||
assertTrue(out.contains("x".repeat(80) + "…"), "expected an 80-char preview with an ellipsis: " + out);
|
||||
}
|
||||
|
||||
/**
|
||||
* fleetd #421: {@code mailbox.pending} counts only broker-ready messages, so a blocked lead's
|
||||
* normal, healthy state is {@code "pending": 0} next to a non-empty {@code held[]} — which
|
||||
* invites the false reading that held mail is in-memory-only. {@code heldCount}/{@code
|
||||
* heldDurable} put an honest second number and fact beside {@code pending} instead of leaving it
|
||||
* as the only one.
|
||||
*/
|
||||
@Test
|
||||
void listReportsAnHonestHeldCountAndDurabilityNotJustPendingZero() {
|
||||
FakeHerdr h = new FakeHerdr();
|
||||
SessionManager sessions = new SessionManager(workerService(h, "http://gx00.gw:8000", Set.of("gx00.gw")));
|
||||
FakeLeadChannel channel = new FakeLeadChannel("mac-opus")
|
||||
.withMailbox("mac-opus", LeadChannel.MailboxState.exists("mac-opus", 0, 1))
|
||||
.hold(new LeadMessage("m1", "fleet01-lead", "mac-opus", "one"))
|
||||
.hold(new LeadMessage("m2", "fleet01-lead", "mac-opus", "two"))
|
||||
.hold(new LeadMessage("m3", "fleet01-lead", "mac-opus", "three"));
|
||||
|
||||
McpSchema.CallToolResult res = FleetMcp.listFleet(
|
||||
workerService(h, "http://gx00.gw:8000", Set.of("gx00.gw")), sessions, null,
|
||||
FleetMcp.CapacitySource.none(), new FleetMcp.HealthCoverageSource(() -> "off"),
|
||||
FleetMcp.QuarantineSource.none(), FleetMcp.OutageSource.none(), FleetMcp.LeadSeatSource.none(),
|
||||
Map.of(), "", new FleetMcp.CoordinationSource(channel, List.of()), true);
|
||||
|
||||
String out = textOf(res);
|
||||
assertTrue(out.contains("\"pending\":0"), out);
|
||||
assertTrue(out.contains("\"heldCount\":3"),
|
||||
"the honest count beside pending: 0 — three messages really are held: " + out);
|
||||
assertTrue(out.contains("\"heldDurable\":true"),
|
||||
"must state the durability fact, not leave pending as the only number next to held[]: " + out);
|
||||
}
|
||||
|
||||
/**
|
||||
* fleetd #440: {@code heldDurable} must be a derived fact, not a literal — so it can report
|
||||
* {@code false} when the channel behind it says held mail is not durable (a non-durable queue,
|
||||
* or a consumer running with {@code autoAck=true}). A test that only ever asserts {@code true}
|
||||
* repeats the defect this ticket fixes.
|
||||
*/
|
||||
@Test
|
||||
void listReportsHeldDurableFalseWhenTheChannelSaysMailIsNotDurable() {
|
||||
FakeHerdr h = new FakeHerdr();
|
||||
SessionManager sessions = new SessionManager(workerService(h, "http://gx00.gw:8000", Set.of("gx00.gw")));
|
||||
FakeLeadChannel channel = new FakeLeadChannel("mac-opus")
|
||||
.withMailbox("mac-opus", LeadChannel.MailboxState.exists("mac-opus", 0, 1))
|
||||
.withHeldDurable(false)
|
||||
.hold(new LeadMessage("m1", "fleet01-lead", "mac-opus", "one"));
|
||||
|
||||
McpSchema.CallToolResult res = FleetMcp.listFleet(
|
||||
workerService(h, "http://gx00.gw:8000", Set.of("gx00.gw")), sessions, null,
|
||||
FleetMcp.CapacitySource.none(), new FleetMcp.HealthCoverageSource(() -> "off"),
|
||||
FleetMcp.QuarantineSource.none(), FleetMcp.OutageSource.none(), FleetMcp.LeadSeatSource.none(),
|
||||
Map.of(), "", new FleetMcp.CoordinationSource(channel, List.of()), true);
|
||||
|
||||
String out = textOf(res);
|
||||
assertTrue(out.contains("\"heldDurable\":false"),
|
||||
"heldDurable must follow the channel, not a hardcoded true: " + out);
|
||||
}
|
||||
|
||||
// ── fleetd #421: a lead reads (never consumes) its own held peer mail ──────────────────────
|
||||
|
||||
@Test
|
||||
void pollWithCoordIdReturnsTheFullBodyWithoutAckingAndLeavesItHeld() {
|
||||
FakeLeadChannel channel = new FakeLeadChannel("mac-opus")
|
||||
.hold(new LeadMessage("m1", "fleet01-lead", "mac-opus", "x".repeat(200)));
|
||||
|
||||
McpSchema.CallToolResult first = FleetMcp.poll(messages, channel, null, null, "mac-opus");
|
||||
assertNotEquals(Boolean.TRUE, first.isError(), textOf(first));
|
||||
String out1 = textOf(first);
|
||||
assertTrue(out1.contains("\"msgId\":\"m1\""), out1);
|
||||
assertTrue(out1.contains("\"from\":\"fleet01-lead\""), out1);
|
||||
assertTrue(out1.contains("\"content\":\"" + "x".repeat(200) + "\""),
|
||||
"the coordId route must return the FULL body, unlike fleet_list's preview: " + out1);
|
||||
|
||||
// Read again: identical bodies, and nothing was acked — peek() still holds it.
|
||||
McpSchema.CallToolResult second = FleetMcp.poll(messages, channel, null, null, "mac-opus");
|
||||
assertEquals(out1, textOf(second), "peek is non-destructive — reading twice must return the same bodies");
|
||||
assertTrue(channel.acked().isEmpty(), "a read must never ack — that is the whole point of the ticket");
|
||||
assertEquals(1, channel.peek().size(), "the message must still be held after being read");
|
||||
}
|
||||
|
||||
@Test
|
||||
void pollWithCoordIdRefusesAPeersCoordIdInsteadOfReturningTheWrongMailOrNothing() {
|
||||
FakeLeadChannel channel = new FakeLeadChannel("mac-opus")
|
||||
.hold(new LeadMessage("m1", "fleet01-lead", "mac-opus", "secret coordination body"));
|
||||
|
||||
// The ORIGINAL fleetd #421 confusion: passing a PEER's id where a self-address was meant.
|
||||
McpSchema.CallToolResult res = FleetMcp.poll(messages, channel, null, null, "fleet01-lead");
|
||||
|
||||
assertEquals(Boolean.TRUE, res.isError());
|
||||
String out = textOf(res);
|
||||
assertFalse(out.contains("secret coordination body"),
|
||||
"a refused read must never leak the body it refused to return: " + out);
|
||||
assertTrue(out.contains("mac-opus"), "the error must name this daemon's own coordId: " + out);
|
||||
}
|
||||
|
||||
@Test
|
||||
void pollWithCoordIdErrorsHonestlyWhenLeadCoordinationIsNotConfigured() {
|
||||
McpSchema.CallToolResult res = FleetMcp.poll(messages, null, null, null, "mac-opus");
|
||||
|
||||
assertEquals(Boolean.TRUE, res.isError());
|
||||
assertTrue(textOf(res).toLowerCase().contains("coordinat"), textOf(res));
|
||||
}
|
||||
|
||||
@Test
|
||||
void listReportsEachDeclaredPeersLiveReachability() {
|
||||
FakeHerdr h = new FakeHerdr();
|
||||
@@ -676,8 +932,9 @@ class FleetMcpTest {
|
||||
McpSchema.CallToolResult res = FleetMcp.listFleet(
|
||||
workerService(h, "http://gx00.gw:8000", Set.of("gx00.gw")), sessions, null,
|
||||
FleetMcp.CapacitySource.none(), new FleetMcp.HealthCoverageSource(() -> "off"),
|
||||
FleetMcp.QuarantineSource.none(), Map.of(), "",
|
||||
new FleetMcp.CoordinationSource(channel, List.of("fleet01-lead", "fleet02-lead", "fleet03-lead")));
|
||||
FleetMcp.QuarantineSource.none(), FleetMcp.OutageSource.none(), FleetMcp.LeadSeatSource.none(),
|
||||
Map.of(), "",
|
||||
new FleetMcp.CoordinationSource(channel, List.of("fleet01-lead", "fleet02-lead", "fleet03-lead")), true);
|
||||
|
||||
String out = textOf(res);
|
||||
assertTrue(out.contains("\"coordId\":\"fleet01-lead\",\"status\":\"exists\",\"pending\":2,\"consumers\":1"), out);
|
||||
@@ -716,6 +973,9 @@ class FleetMcpTest {
|
||||
@Override
|
||||
public String selfCoordId() { return "mac-opus"; }
|
||||
|
||||
@Override
|
||||
public boolean heldDurable() { return true; }
|
||||
|
||||
@Override
|
||||
public MailboxState inspect(String coordId) {
|
||||
started.countDown();
|
||||
@@ -1239,6 +1499,10 @@ class FleetMcpTest {
|
||||
|
||||
@Test
|
||||
void bridgeAckReturnsConfirmationForValidArgs() {
|
||||
// fleet_ack only reports success for a msgId actually queued in the target's inbox
|
||||
// (fleetd #437) — publish one via the inbox directly rather than asserting on a
|
||||
// fabricated id nothing ever queued.
|
||||
inbox.publish("term_a", "msg-1", "queued reply");
|
||||
McpSchema.CallToolResult res = FleetMcp.ack(messages, "term_a", "msg-1");
|
||||
assertNotEquals(Boolean.TRUE, res.isError());
|
||||
assertTrue(textOf(res).contains("msg-1"), "response should mention the msgId");
|
||||
@@ -1251,21 +1515,51 @@ class FleetMcpTest {
|
||||
assertTrue(FleetMcp.ack(messages, " ", "msg-1").isError());
|
||||
}
|
||||
|
||||
@Test
|
||||
void bridgeAckOfAnIdInNoInboxIsAnError() {
|
||||
// fleetd #437: fleet_ack used to say "acknowledged <msgId>" for a message it never
|
||||
// touched, because nothing in the chain reported hit vs. miss. "never-queued" is in no
|
||||
// inbox at all, so this must error rather than claim success.
|
||||
McpSchema.CallToolResult res = FleetMcp.ack(messages, "term_a", "never-queued");
|
||||
assertTrue(res.isError());
|
||||
assertTrue(textOf(res).contains("never-queued"), textOf(res));
|
||||
}
|
||||
|
||||
@Test
|
||||
void bridgeAckOfACoordIdTargetIsAnErrorNamingFleetPoll() {
|
||||
// A coord-id names a peer lead's held mailbox (LeadChannel/LeadMailbox), never a
|
||||
// worker's ReplyInbox — fleet_ack has no route to it and must say so, pointing at
|
||||
// fleet_poll{coordId} instead of reporting a false "acknowledged".
|
||||
McpSchema.CallToolResult res = FleetMcp.ack(messages, "coord-some-peer", "msg-1");
|
||||
assertTrue(res.isError());
|
||||
assertTrue(textOf(res).contains("fleet_poll{coordId}"), textOf(res));
|
||||
}
|
||||
|
||||
@Test
|
||||
void bridgeAckRemovesSpecificReply() {
|
||||
// Queue a reply and capture its msgId.
|
||||
FleetMcp.reply(messages, "term_a", "orphan");
|
||||
FleetMcp.reply(messages, "term_a", Role.WORKER, "orphan");
|
||||
var before = messages.drainReplies("term_a");
|
||||
assertEquals(1, before.size(), "one reply in the inbox");
|
||||
String msgId = before.getFirst().msgId();
|
||||
|
||||
// Publish the same reply again and ack it via fleet_ack surface.
|
||||
FleetMcp.reply(messages, "term_a", "orphan-again");
|
||||
FleetMcp.reply(messages, "term_a", Role.WORKER, "orphan-again");
|
||||
var peeked = messages.drainReplies("term_a");
|
||||
assertEquals(1, peeked.size(), "one fresh reply in the inbox");
|
||||
|
||||
// ackReply works (no-op since published with a different UUID, but callable).
|
||||
assertDoesNotThrow(() -> messages.ackReply("term_a", msgId));
|
||||
// fleetd #437: msgId was already drained above (a fresh UUID each publish), so it is no
|
||||
// longer in the inbox — ackReply must now report that miss instead of pretending to ack.
|
||||
assertFalse(messages.ackReply("term_a", msgId));
|
||||
}
|
||||
|
||||
@Test
|
||||
void bridgeAckRemovingARealQueuedReplyReportsSuccessAndRemovesIt() {
|
||||
// The worker path must not change behaviour: acking a reply that IS still in the inbox
|
||||
// still succeeds and still removes it (fleetd #437).
|
||||
inbox.publish("term_a", "real-1", "still queued");
|
||||
assertTrue(messages.ackReply("term_a", "real-1"), "ack of a real queued reply must report true");
|
||||
assertTrue(inbox.peek("term_a").isEmpty(), "the acked reply must be gone from the inbox");
|
||||
}
|
||||
|
||||
@Test
|
||||
|
||||
@@ -0,0 +1,62 @@
|
||||
package dev.ltms.fleet.mcp;
|
||||
|
||||
import dev.ltms.fleet.config.FleetConfig;
|
||||
import dev.ltms.fleet.guard.SubscriptionGuard;
|
||||
import dev.ltms.fleet.herdr.AgentControl;
|
||||
import dev.ltms.fleet.herdr.FakeHerdr;
|
||||
import dev.ltms.fleet.herdr.WorkspaceControl;
|
||||
import dev.ltms.fleet.member.ClaudeCodeLauncher;
|
||||
import dev.ltms.fleet.peer.PeerLauncher;
|
||||
import io.modelcontextprotocol.spec.McpSchema;
|
||||
import org.junit.jupiter.api.Test;
|
||||
|
||||
import java.util.LinkedHashMap;
|
||||
import java.util.Map;
|
||||
import java.util.Set;
|
||||
|
||||
import static org.junit.jupiter.api.Assertions.assertNotEquals;
|
||||
import static org.junit.jupiter.api.Assertions.assertTrue;
|
||||
|
||||
/**
|
||||
* fleetd #395: {@code fleet_profiles} must let an operator tell "this profile's usage-limit
|
||||
* detection is armed" from "nothing can ever quarantine this profile" — see {@link
|
||||
* FleetMcp.QuarantineSource#exhaustedPatternArmed()} and {@link
|
||||
* FleetMcp#profilesView(PeerLauncher, FleetMcp.QuarantineSource, FleetMcp.OutageSource)}.
|
||||
*/
|
||||
class FleetProfilesArmedFieldTest {
|
||||
|
||||
private static PeerLauncher twoProfileLauncher(FakeHerdr h) {
|
||||
FleetConfig.Profile armed = new FleetConfig.Profile(
|
||||
"armed-profile", "http://gx00.gw:8000", "coder", null, "FLEETD_WORKER_TOKEN", null,
|
||||
"tab", "fleetd-workers", "worker: {profile} #{n}", null, null, null);
|
||||
FleetConfig.Profile unarmed = new FleetConfig.Profile(
|
||||
"unarmed-profile", "http://gx01.gw:8000", "coder", null, "FLEETD_WORKER_TOKEN", null,
|
||||
"tab", "fleetd-workers", "worker: {profile} #{n}", null, null, null);
|
||||
Map<String, FleetConfig.Profile> profiles = new LinkedHashMap<>();
|
||||
profiles.put(armed.profile(), armed);
|
||||
profiles.put(unarmed.profile(), unarmed);
|
||||
return new ClaudeCodeLauncher(new AgentControl(h), new WorkspaceControl(h),
|
||||
new SubscriptionGuard(Set.of("gx00.gw", "gx01.gw")), profiles, armed.profile(), _ -> "tok");
|
||||
}
|
||||
|
||||
private static String textOf(McpSchema.CallToolResult r) {
|
||||
return ((McpSchema.TextContent) r.content().getFirst()).text();
|
||||
}
|
||||
|
||||
@Test
|
||||
void armedProfileReportsArmedAndUnarmedReportsUnarmed() {
|
||||
FakeHerdr h = new FakeHerdr();
|
||||
PeerLauncher workers = twoProfileLauncher(h);
|
||||
FleetMcp.QuarantineSource source = new FleetMcp.QuarantineSource(
|
||||
_ -> null, dev.ltms.fleet.placement.BackendQuarantine.none(),
|
||||
profile -> "armed-profile".equals(profile));
|
||||
|
||||
McpSchema.CallToolResult res = FleetMcp.profiles(workers, source);
|
||||
assertNotEquals(Boolean.TRUE, res.isError());
|
||||
String out = textOf(res);
|
||||
|
||||
assertTrue(out.contains("\"exhaustionDetectionArmed\""), out);
|
||||
assertTrue(out.contains("\"armed-profile\":true"), out);
|
||||
assertTrue(out.contains("\"unarmed-profile\":false"), out);
|
||||
}
|
||||
}
|
||||
@@ -0,0 +1,114 @@
|
||||
package dev.ltms.fleet.mcp;
|
||||
|
||||
import dev.ltms.fleet.config.ConfigRef;
|
||||
import dev.ltms.fleet.config.FleetConfig;
|
||||
import dev.ltms.fleet.guard.SubscriptionGuard;
|
||||
import dev.ltms.fleet.herdr.AgentControl;
|
||||
import dev.ltms.fleet.herdr.FakeHerdr;
|
||||
import dev.ltms.fleet.herdr.WorkspaceControl;
|
||||
import dev.ltms.fleet.member.ClaudeCodeLauncher;
|
||||
import dev.ltms.fleet.member.CompositePeerLauncher;
|
||||
import dev.ltms.fleet.peer.MemberRole;
|
||||
import dev.ltms.fleet.peer.PeerLauncher;
|
||||
import dev.ltms.fleet.peer.SpawnRequest;
|
||||
import dev.ltms.fleet.placement.BackendQuarantine;
|
||||
import org.junit.jupiter.api.Test;
|
||||
import org.junit.jupiter.api.io.TempDir;
|
||||
|
||||
import java.nio.file.Files;
|
||||
import java.nio.file.Path;
|
||||
import java.util.List;
|
||||
import java.util.Map;
|
||||
import java.util.Set;
|
||||
|
||||
import static org.junit.jupiter.api.Assertions.assertEquals;
|
||||
import static org.junit.jupiter.api.Assertions.assertTrue;
|
||||
|
||||
/**
|
||||
* fleetd #425: {@code fleet_profiles}' {@code "default"} field was captured once at boot
|
||||
* ({@code cfg.effectiveDefaultProfile()}, frozen into {@code CompositePeerLauncher.defaultProfile}
|
||||
* at construction) while an unqualified spawn resolves the same underlying key
|
||||
* ({@code fleet.developers}' first entry) live, on every call. Reordering {@code fleet.developers}
|
||||
* and reloading changed where a spawn landed without ever changing what {@code fleet_profiles}
|
||||
* reported — a lead following {@code CLAUDE.md}'s "check {@code fleet_profiles} once per session"
|
||||
* instruction was told a stale answer.
|
||||
*
|
||||
* <p>This test drives the exact caller {@code fleet_profiles} uses —
|
||||
* {@link FleetMcp#profilesView(PeerLauncher, FleetMcp.QuarantineSource, FleetMcp.OutageSource)} —
|
||||
* against a real, reloadable {@link ConfigRef}, so it fails if the reporting path is ever recoupled
|
||||
* to a frozen value instead of {@link CompositePeerLauncher#defaultProfile()}'s live answer.
|
||||
*/
|
||||
class FleetProfilesLiveDefaultTest {
|
||||
|
||||
/** A minimal fleetd.yaml whose dev pool is {@code profilesInOrder}, in that definition order. */
|
||||
private static String yamlWithDevPool(String... profilesInOrder) {
|
||||
StringBuilder devPool = new StringBuilder();
|
||||
for (int i = 0; i < profilesInOrder.length; i++) {
|
||||
devPool.append(" slot").append(i).append(":\n profile: ")
|
||||
.append(profilesInOrder[i]).append('\n');
|
||||
}
|
||||
return """
|
||||
bind:
|
||||
host: 127.0.0.1
|
||||
port: 8765
|
||||
herdrSocket: ~/.config/herdr/herdr.sock
|
||||
profiles:
|
||||
opus:
|
||||
baseUrl: http://gx00.gw:8000
|
||||
model: opus-coder
|
||||
sonnet:
|
||||
baseUrl: http://gx00.gw:8000
|
||||
model: sonnet-coder
|
||||
guard:
|
||||
offSubscriptionHosts:
|
||||
- gx00.gw
|
||||
fleet:
|
||||
developers:
|
||||
""" + devPool;
|
||||
}
|
||||
|
||||
@Test
|
||||
void fleetProfilesDefaultTracksALiveDevPoolReorderAfterReload(@TempDir Path dir) throws Exception {
|
||||
Path f = dir.resolve("fleetd.yaml");
|
||||
Files.writeString(f, yamlWithDevPool("opus", "sonnet"));
|
||||
ConfigRef ref = new ConfigRef(f, FleetConfig.load(f));
|
||||
|
||||
Map<String, FleetConfig.Profile> profiles = Map.of(
|
||||
"opus", new FleetConfig.Profile("opus", "http://gx00.gw:8000", "opus-coder", null,
|
||||
"FLEETD_WORKER_TOKEN", List.of("claude"), "tab", "fleetd-workers",
|
||||
"worker: {profile} #{n}", null, null, null),
|
||||
"sonnet", new FleetConfig.Profile("sonnet", "http://gx00.gw:8000", "sonnet-coder", null,
|
||||
"FLEETD_WORKER_TOKEN", List.of("claude"), "tab", "fleetd-workers",
|
||||
"worker: {profile} #{n}", null, null, null));
|
||||
FakeHerdr herdr = new FakeHerdr();
|
||||
ClaudeCodeLauncher adapter = new ClaudeCodeLauncher(new AgentControl(herdr), new WorkspaceControl(herdr),
|
||||
new SubscriptionGuard(Set.of("gx00.gw")), profiles, "opus", _ -> null);
|
||||
PeerLauncher workers = new CompositePeerLauncher(
|
||||
List.of(adapter), "opus", ref, _ -> 0, BackendQuarantine.none());
|
||||
|
||||
assertReportedDefaultMatchesAnUnqualifiedSpawn(workers, "opus");
|
||||
|
||||
Files.writeString(f, yamlWithDevPool("sonnet", "opus"));
|
||||
ConfigRef.Outcome out = ref.reload();
|
||||
assertTrue(out.applied(), () -> "reload should apply cleanly: " + out.error());
|
||||
|
||||
assertReportedDefaultMatchesAnUnqualifiedSpawn(workers, "sonnet");
|
||||
}
|
||||
|
||||
/**
|
||||
* Asserts BOTH that {@code fleet_profiles}' {@code "default"} equals {@code expected}, AND that
|
||||
* it equals what a real unqualified {@code MemberRole#DEV} spawn actually gets placed on right
|
||||
* now — the two facts fleetd #425 found disagreeing.
|
||||
*/
|
||||
private static void assertReportedDefaultMatchesAnUnqualifiedSpawn(PeerLauncher workers, String expected) {
|
||||
Map<String, Object> view = FleetMcp.profilesView(
|
||||
workers, FleetMcp.QuarantineSource.none(), FleetMcp.OutageSource.none());
|
||||
assertEquals(expected, view.get("default"),
|
||||
"fleet_profiles' \"default\" must be the live dev-pool answer, not a boot-time snapshot");
|
||||
|
||||
String placed = workers.spawn(
|
||||
new SpawnRequest(null, null, null, null, null, MemberRole.DEV)).profile();
|
||||
assertEquals(expected, placed,
|
||||
"sanity: the profile an unqualified dev spawn actually lands on");
|
||||
}
|
||||
}
|
||||
@@ -0,0 +1,109 @@
|
||||
package dev.ltms.fleet.mcp;
|
||||
|
||||
import dev.ltms.fleet.config.FleetConfig;
|
||||
import dev.ltms.fleet.guard.SubscriptionGuard;
|
||||
import dev.ltms.fleet.herdr.AgentControl;
|
||||
import dev.ltms.fleet.herdr.FakeHerdr;
|
||||
import dev.ltms.fleet.herdr.WorkspaceControl;
|
||||
import dev.ltms.fleet.member.ClaudeCodeLauncher;
|
||||
import dev.ltms.fleet.member.CompositePeerLauncher;
|
||||
import dev.ltms.fleet.member.HerdrPeerLauncher;
|
||||
import dev.ltms.fleet.peer.PeerLauncher;
|
||||
import dev.ltms.fleet.placement.BackendOutagePolicy;
|
||||
import dev.ltms.fleet.placement.BackendQuarantine;
|
||||
import dev.ltms.fleet.placement.PlacementPolicies;
|
||||
import org.junit.jupiter.api.Test;
|
||||
|
||||
import java.util.List;
|
||||
import java.util.Map;
|
||||
import java.util.Set;
|
||||
|
||||
import static org.junit.jupiter.api.Assertions.assertEquals;
|
||||
import static org.junit.jupiter.api.Assertions.assertFalse;
|
||||
|
||||
/**
|
||||
* fleetd #422 follow-up: {@code fleet_profiles}/{@code GET /profiles} — the reporting surface a
|
||||
* lead actually reads — must let it tell apart the three states {@link
|
||||
* dev.ltms.fleet.peer.PeerLauncher#disabledModels()} alone collapses into one empty set: no
|
||||
* {@code models:} block at all, a block armed with nothing currently off, and a block with N
|
||||
* models off. See {@link dev.ltms.fleet.peer.PeerLauncher.ModelGateState}'s javadoc for why a
|
||||
* bare {@code disabledModels()} read cannot make this distinction, and {@link
|
||||
* FleetMcp#profilesView} for where {@code modelGateArmed} is added alongside the existing {@code
|
||||
* modelsOff} key.
|
||||
*
|
||||
* <p>Every assertion here goes through {@link FleetMcp#profilesView}, never {@code
|
||||
* PeerLauncher.modelGateState()} directly — {@code CompositePeerLauncherTest} already proves the
|
||||
* accessor itself; this class proves the surface a lead reads (fleet_profiles / GET /profiles)
|
||||
* renders what that accessor reports.
|
||||
*/
|
||||
class FleetProfilesModelGateStateTest {
|
||||
|
||||
private static FleetConfig.Profile profile(String name, String model) {
|
||||
return new FleetConfig.Profile(name, "http://gx00.gw:8000", model, null, "FLEETD_WORKER_TOKEN",
|
||||
null, "tab", "fleetd-workers", "worker: {profile} #{n}", null, null, null);
|
||||
}
|
||||
|
||||
private static FleetMcp.QuarantineSource noQuarantine() {
|
||||
return new FleetMcp.QuarantineSource(_ -> null, BackendQuarantine.none(), _ -> false);
|
||||
}
|
||||
|
||||
private static HerdrPeerLauncher claudeAdapter(FakeHerdr h, Map<String, FleetConfig.Profile> profiles) {
|
||||
return new ClaudeCodeLauncher(new AgentControl(h), new WorkspaceControl(h),
|
||||
new SubscriptionGuard(Set.of("gx00.gw")), profiles, "local", _ -> "tok");
|
||||
}
|
||||
|
||||
/**
|
||||
* State 1: no {@code models:} block at all — a plain {@code ClaudeCodeLauncher} (no {@code
|
||||
* models:} supplier exists for it to read) has nothing to gate against, matching the fleet01
|
||||
* host measured for this ticket: {@code grep -c '^models:' fleetd.yaml} returns 0 there.
|
||||
*/
|
||||
@Test
|
||||
void noModelsBlockReportsGateNotArmed() {
|
||||
FakeHerdr h = new FakeHerdr();
|
||||
Map<String, FleetConfig.Profile> profiles = Map.of("local", profile("local", "deepseek-v4-flash"));
|
||||
PeerLauncher workers = claudeAdapter(h, profiles);
|
||||
|
||||
Map<String, Object> view = FleetMcp.profilesView(workers, noQuarantine(), FleetMcp.OutageSource.none());
|
||||
|
||||
assertEquals(Boolean.FALSE, view.get("modelGateArmed"),
|
||||
"no models: block to read from — nothing is gated, and nothing can be");
|
||||
assertFalse(view.containsKey("modelsOff"), "nothing configured, so no off set to report either");
|
||||
}
|
||||
|
||||
/** State 2: a {@code models:} block is present, but nothing in it is currently turned off. */
|
||||
@Test
|
||||
void modelsBlockWithNothingOffReportsGateArmedAndZeroOff() {
|
||||
FakeHerdr h = new FakeHerdr();
|
||||
Map<String, FleetConfig.Profile> profiles = Map.of("local", profile("local", "deepseek-v4-flash"));
|
||||
FleetConfig.Models models = new FleetConfig.Models(
|
||||
List.of(new FleetConfig.Models.ModelEntry("deepseek-v4-flash", true)));
|
||||
PeerLauncher workers = new CompositePeerLauncher(List.of(claudeAdapter(h, profiles)), "local", profiles,
|
||||
PlacementPolicies.fixed(), _ -> 0, null, BackendQuarantine.none(),
|
||||
new BackendOutagePolicy(System::nanoTime), () -> models);
|
||||
|
||||
Map<String, Object> view = FleetMcp.profilesView(workers, noQuarantine(), FleetMcp.OutageSource.none());
|
||||
|
||||
assertEquals(Boolean.TRUE, view.get("modelGateArmed"),
|
||||
"a models: block is present, so the gate is armed even though nothing is off yet");
|
||||
assertFalse(view.containsKey("modelsOff"),
|
||||
"nothing is off, so the key stays absent — an empty list here would be indistinguishable "
|
||||
+ "from today's modelsOff omission, exactly the ambiguity modelGateArmed exists to remove");
|
||||
}
|
||||
|
||||
/** State 3: a {@code models:} block is present with one model currently turned off. */
|
||||
@Test
|
||||
void modelsBlockWithModelsOffReportsGateArmedAndTheOffSet() {
|
||||
FakeHerdr h = new FakeHerdr();
|
||||
Map<String, FleetConfig.Profile> profiles = Map.of("local", profile("local", "deepseek-v4-flash"));
|
||||
FleetConfig.Models models = new FleetConfig.Models(
|
||||
List.of(new FleetConfig.Models.ModelEntry("deepseek-v4-flash", false)));
|
||||
PeerLauncher workers = new CompositePeerLauncher(List.of(claudeAdapter(h, profiles)), "local", profiles,
|
||||
PlacementPolicies.fixed(), _ -> 0, null, BackendQuarantine.none(),
|
||||
new BackendOutagePolicy(System::nanoTime), () -> models);
|
||||
|
||||
Map<String, Object> view = FleetMcp.profilesView(workers, noQuarantine(), FleetMcp.OutageSource.none());
|
||||
|
||||
assertEquals(Boolean.TRUE, view.get("modelGateArmed"));
|
||||
assertEquals(List.of("deepseek-v4-flash"), view.get("modelsOff"));
|
||||
}
|
||||
}
|
||||
+111
@@ -0,0 +1,111 @@
|
||||
package dev.ltms.fleet.mcp;
|
||||
|
||||
import dev.ltms.fleet.config.FleetConfig;
|
||||
import dev.ltms.fleet.guard.SubscriptionGuard;
|
||||
import dev.ltms.fleet.herdr.AgentControl;
|
||||
import dev.ltms.fleet.herdr.FakeHerdr;
|
||||
import dev.ltms.fleet.herdr.WorkspaceControl;
|
||||
import dev.ltms.fleet.member.ClaudeCodeLauncher;
|
||||
import dev.ltms.fleet.peer.PeerLauncher;
|
||||
import dev.ltms.fleet.placement.BackendQuarantine;
|
||||
import io.modelcontextprotocol.spec.McpSchema;
|
||||
import org.junit.jupiter.api.Test;
|
||||
|
||||
import java.util.Map;
|
||||
import java.util.Set;
|
||||
import java.util.concurrent.TimeUnit;
|
||||
|
||||
import static org.junit.jupiter.api.Assertions.assertFalse;
|
||||
import static org.junit.jupiter.api.Assertions.assertNotEquals;
|
||||
import static org.junit.jupiter.api.Assertions.assertTrue;
|
||||
|
||||
/**
|
||||
* fleetd #446 follow-up (criterion 3): a mutation battery run against the merged PR #457 proved
|
||||
* that deleting either {@code row.put("model", model);} or {@code row.put("reason", reason);}
|
||||
* from {@link FleetMcp#profilesView} left all 1586 existing tests green — nothing exercised a
|
||||
* {@link FleetMcp.QuarantineSource} whose {@code modelFor}/{@code reasonFor} actually return a
|
||||
* value. This class is what makes both deletions red, and also pins the two "must be omitted, not
|
||||
* emitted as null/blank" branches those two {@code if} guards exist for — {@link
|
||||
* FleetMcpTest#profilesReportsAQuarantinedCredential} covers the credential/quarantine-duration
|
||||
* shape but supplies neither function, so it cannot distinguish "the field is missing" from "the
|
||||
* field was never asked for".
|
||||
*/
|
||||
class FleetProfilesQuarantineModelReasonFieldsTest {
|
||||
|
||||
private static final String PROFILE = "terra";
|
||||
private static final String CREDENTIAL = "cred-terra";
|
||||
|
||||
private static PeerLauncher launcher(FakeHerdr h) {
|
||||
FleetConfig.Profile profile = new FleetConfig.Profile(
|
||||
PROFILE, "http://gx00.gw:8000", "claude-opus-5", null, "FLEETD_WORKER_TOKEN", null,
|
||||
"tab", "fleetd-workers", "worker: {profile} #{n}", null, null, null);
|
||||
return new ClaudeCodeLauncher(new AgentControl(h), new WorkspaceControl(h),
|
||||
new SubscriptionGuard(Set.of("gx00.gw")), Map.of(PROFILE, profile), PROFILE, _ -> "tok");
|
||||
}
|
||||
|
||||
private static BackendQuarantine quarantined(String credentialId) {
|
||||
BackendQuarantine quarantine = new BackendQuarantine(() -> 0L, TimeUnit.MINUTES.toNanos(30));
|
||||
quarantine.quarantine(credentialId);
|
||||
return quarantine;
|
||||
}
|
||||
|
||||
private static String textOf(McpSchema.CallToolResult r) {
|
||||
return ((McpSchema.TextContent) r.content().getFirst()).text();
|
||||
}
|
||||
|
||||
@Test
|
||||
void quarantinedRowNamesTheModelAndTheReason() {
|
||||
FakeHerdr h = new FakeHerdr();
|
||||
FleetMcp.QuarantineSource source = new FleetMcp.QuarantineSource(
|
||||
_ -> CREDENTIAL, quarantined(CREDENTIAL), _ -> true,
|
||||
_ -> "claude-opus-5", _ -> "The usage limit has been reached");
|
||||
|
||||
McpSchema.CallToolResult res = FleetMcp.profiles(launcher(h), source);
|
||||
assertNotEquals(Boolean.TRUE, res.isError());
|
||||
String out = textOf(res);
|
||||
|
||||
assertTrue(out.contains("\"quarantined\""), out);
|
||||
assertTrue(out.contains("\"model\":\"claude-opus-5\""),
|
||||
"the quarantined row must name the model the fix applies to: " + out);
|
||||
assertTrue(out.contains("\"reason\":\"The usage limit has been reached\""),
|
||||
"the quarantined row must name why it was quarantined: " + out);
|
||||
}
|
||||
|
||||
@Test
|
||||
void quarantinedRowOmitsModelWhenNoneIsConfigured() {
|
||||
FakeHerdr h = new FakeHerdr();
|
||||
// null and "" (blank) both exercise the same production guard (`!model.isBlank()`) —
|
||||
// covered together since both must produce the identical outcome: the key absent.
|
||||
for (String noModel : new String[] {null, "", " "}) {
|
||||
FleetMcp.QuarantineSource source = new FleetMcp.QuarantineSource(
|
||||
_ -> CREDENTIAL, quarantined(CREDENTIAL), _ -> true,
|
||||
_ -> noModel, _ -> "some reason");
|
||||
|
||||
McpSchema.CallToolResult res = FleetMcp.profiles(launcher(h), source);
|
||||
String out = textOf(res);
|
||||
|
||||
assertTrue(out.contains("\"quarantined\""), out);
|
||||
assertFalse(out.contains("\"model\""),
|
||||
"modelFor returned " + (noModel == null ? "null" : "'" + noModel + "'")
|
||||
+ " — the key must be ABSENT, not emitted as null or an empty string: " + out);
|
||||
}
|
||||
}
|
||||
|
||||
@Test
|
||||
void quarantinedRowOmitsReasonWhenNoneIsKnown() {
|
||||
FakeHerdr h = new FakeHerdr();
|
||||
for (String noReason : new String[] {null, "", " "}) {
|
||||
FleetMcp.QuarantineSource source = new FleetMcp.QuarantineSource(
|
||||
_ -> CREDENTIAL, quarantined(CREDENTIAL), _ -> true,
|
||||
_ -> "claude-opus-5", _ -> noReason);
|
||||
|
||||
McpSchema.CallToolResult res = FleetMcp.profiles(launcher(h), source);
|
||||
String out = textOf(res);
|
||||
|
||||
assertTrue(out.contains("\"quarantined\""), out);
|
||||
assertFalse(out.contains("\"reason\""),
|
||||
"reasonFor returned " + (noReason == null ? "null" : "'" + noReason + "'")
|
||||
+ " — the key must be ABSENT, not emitted as null or an empty string: " + out);
|
||||
}
|
||||
}
|
||||
}
|
||||
@@ -3,6 +3,7 @@ package dev.ltms.fleet.member;
|
||||
import ch.qos.logback.classic.Logger;
|
||||
import ch.qos.logback.classic.spi.ILoggingEvent;
|
||||
import ch.qos.logback.core.read.ListAppender;
|
||||
import dev.ltms.fleet.config.ConfigRef;
|
||||
import dev.ltms.fleet.config.FleetConfig;
|
||||
import dev.ltms.fleet.guard.SubscriptionGuard;
|
||||
import dev.ltms.fleet.herdr.Agent;
|
||||
@@ -19,11 +20,15 @@ import dev.ltms.fleet.peer.PeerUnreachableException;
|
||||
import dev.ltms.fleet.peer.SpawnRequest;
|
||||
import dev.ltms.fleet.placement.BackendOutagePolicy;
|
||||
import dev.ltms.fleet.placement.BackendQuarantine;
|
||||
import dev.ltms.fleet.placement.PlacementDecision;
|
||||
import dev.ltms.fleet.placement.PlacementException;
|
||||
import dev.ltms.fleet.placement.PlacementPolicies;
|
||||
import org.junit.jupiter.api.Test;
|
||||
import org.junit.jupiter.api.io.TempDir;
|
||||
import org.slf4j.LoggerFactory;
|
||||
|
||||
import java.nio.file.Files;
|
||||
import java.nio.file.Path;
|
||||
import java.util.EnumSet;
|
||||
import java.util.HashMap;
|
||||
import java.util.LinkedHashMap;
|
||||
@@ -141,6 +146,13 @@ class CompositePeerLauncherTest {
|
||||
null, null, credentialId, null);
|
||||
}
|
||||
|
||||
/** Like {@link #stubWorker(String)}, but with an explicit {@code model:} for fleetd #422 tests. */
|
||||
private static FleetConfig.Profile stubWorkerModel(String profile, String model) {
|
||||
return new FleetConfig.Profile(profile, "http://gx00.gw:8000", model,
|
||||
null, "FLEETD_WORKER_TOKEN", List.of("claude"), "tab", "fleetd-workers",
|
||||
"w #{n}", null, null, null, null, null, null, null, null, null);
|
||||
}
|
||||
|
||||
/**
|
||||
* An <em>order-preserving</em> profile map. Never {@code Map.of} here: its iteration order is
|
||||
* salted per JVM run, and the weighted policy breaks an exact-weight tie on candidate order —
|
||||
@@ -525,6 +537,31 @@ class CompositePeerLauncherTest {
|
||||
assertEquals("claude", h.profile(), "the returned handle carries the resolved default profile");
|
||||
}
|
||||
|
||||
/**
|
||||
* fleetd #435: {@code FixedPlacementPolicy} — the default placement policy every config uses
|
||||
* unless {@code placement:} is set — never consulted {@code maxLoad}, so an unqualified spawn
|
||||
* (a blank profile, the normal delegation path) landed on a capped default anyway. Measured on
|
||||
* 7667727: a single dev profile at {@code maxLoad: 1} with 1 live, under {@code fixed()},
|
||||
* returned "SPAWNED on profile=a". This test goes through {@code CompositePeerLauncher.spawn}
|
||||
* with a blank profile, not the policy in isolation, so it proves the caller actually reaches
|
||||
* the fixed default's new cap check rather than only the {@code select} method.
|
||||
*/
|
||||
@Test
|
||||
void fixedPolicyGatesDefaultProfileAtMaxLoadOnUnqualifiedSpawn() {
|
||||
FakeHerdr herdr = new FakeHerdr();
|
||||
Map<String, FleetConfig.Profile> profiles = ordered(
|
||||
"a", stubWorker("a", 1.0f, 1),
|
||||
"b", stubWorker("b", 1.0f, null));
|
||||
StubLauncher adapter = new StubLauncher("claude", herdr, profiles, "a", Set.of());
|
||||
CompositePeerLauncher composite = new CompositePeerLauncher(
|
||||
List.of(adapter), "a", profiles, PlacementPolicies.fixed(), name -> "a".equals(name) ? 1 : 0);
|
||||
|
||||
PeerHandle h = composite.spawn(new SpawnRequest(null, null, null));
|
||||
assertEquals("b", h.profile(),
|
||||
"the default profile a is at maxLoad, so fixed placement must fall through to b");
|
||||
assertEquals(0, adapter.spawnCount("a"), "a is never spawned — it is already at cap");
|
||||
}
|
||||
|
||||
@Test
|
||||
void weightedPolicyGatesProfileAtMaxLoad() {
|
||||
FakeHerdr herdr = new FakeHerdr();
|
||||
@@ -908,6 +945,89 @@ class CompositePeerLauncherTest {
|
||||
assertTrue(e.getMessage().contains("maxLoad"), e.getMessage());
|
||||
}
|
||||
|
||||
// ── fleetd #425: defaultProfile()/defaultProfileFor() must track a live reload ─────────────
|
||||
|
||||
/** A minimal fleetd.yaml whose dev pool is {@code profilesInOrder}, in that definition order. */
|
||||
private static String yamlWithDevPool(String... profilesInOrder) {
|
||||
StringBuilder devPool = new StringBuilder();
|
||||
for (int i = 0; i < profilesInOrder.length; i++) {
|
||||
devPool.append(" slot").append(i).append(":\n profile: ")
|
||||
.append(profilesInOrder[i]).append('\n');
|
||||
}
|
||||
return """
|
||||
bind:
|
||||
host: 127.0.0.1
|
||||
port: 8765
|
||||
herdrSocket: ~/.config/herdr/herdr.sock
|
||||
profiles:
|
||||
opus:
|
||||
baseUrl: http://gx00.gw:8000
|
||||
model: opus-coder
|
||||
sonnet:
|
||||
baseUrl: http://gx00.gw:8000
|
||||
model: sonnet-coder
|
||||
guard:
|
||||
offSubscriptionHosts:
|
||||
- gx00.gw
|
||||
fleet:
|
||||
developers:
|
||||
""" + devPool;
|
||||
}
|
||||
|
||||
/**
|
||||
* Criterion 1 (fleetd #425): reorder {@code fleet.developers}, reload, and assert the reported
|
||||
* default ({@link CompositePeerLauncher#defaultProfile()} — what {@code fleet_profiles}' {@code
|
||||
* "default"} is built from, see {@code FleetMcp.profilesView}) matches what an unqualified
|
||||
* {@code MemberRole#DEV} spawn is actually placed on, both before and after the reorder. Asserts
|
||||
* {@code applied()} so the test proves the reload actually took, not that nothing changed.
|
||||
*/
|
||||
@Test
|
||||
void defaultProfileTracksALiveDevPoolReorderAfterReload(@TempDir Path dir) throws Exception {
|
||||
Path f = dir.resolve("fleetd.yaml");
|
||||
Files.writeString(f, yamlWithDevPool("opus", "sonnet"));
|
||||
ConfigRef ref = new ConfigRef(f, FleetConfig.load(f));
|
||||
|
||||
FakeHerdr herdr = new FakeHerdr();
|
||||
StubLauncher adapter = new StubLauncher("claude", herdr, threeProfiles(), "opus", Set.of());
|
||||
CompositePeerLauncher composite = new CompositePeerLauncher(
|
||||
List.of(adapter), "opus", ref, _ -> 0, BackendQuarantine.none());
|
||||
|
||||
assertEquals("opus", composite.defaultProfile(),
|
||||
"reported default starts at the dev pool's first entry");
|
||||
assertEquals("opus", composite.spawn(
|
||||
new SpawnRequest(null, null, null, null, null, MemberRole.DEV)).profile(),
|
||||
"an unqualified dev spawn must land on the same profile that was just reported");
|
||||
|
||||
Files.writeString(f, yamlWithDevPool("sonnet", "opus"));
|
||||
ConfigRef.Outcome out = ref.reload();
|
||||
assertTrue(out.applied(), () -> "reload should apply cleanly: " + out.error());
|
||||
|
||||
assertEquals("sonnet", composite.defaultProfile(),
|
||||
"the reported default must follow the reorder with no daemon restart");
|
||||
assertEquals("sonnet", composite.spawn(
|
||||
new SpawnRequest(null, null, null, null, null, MemberRole.DEV)).profile(),
|
||||
"and it must still be exactly what an unqualified spawn actually gets");
|
||||
}
|
||||
|
||||
/**
|
||||
* Criterion 2 — the mirror, and the load-bearing half (fleetd #425): with NOTHING configured (no
|
||||
* profiles at all, hence an empty pool for every role), the frozen {@code defaultProfile} field
|
||||
* is still what gets reported. A fix that always returns {@code poolFor(role).getFirst()} with no
|
||||
* empty-pool fallback throws or returns the wrong thing here even though criterion 1 above still
|
||||
* passes — this is the test that catches it.
|
||||
*/
|
||||
@Test
|
||||
void defaultProfileFallsBackToTheFrozenFieldWhenNothingIsConfiguredAtAll() {
|
||||
FakeHerdr herdr = new FakeHerdr();
|
||||
StubLauncher adapter = new StubLauncher("claude", herdr, Map.of(), "opus", Set.of());
|
||||
CompositePeerLauncher composite = new CompositePeerLauncher(
|
||||
List.of(adapter), "opus", Map.of(), PlacementPolicies.fixed(), _ -> 0);
|
||||
|
||||
assertEquals("opus", composite.defaultProfile(),
|
||||
"with no profiles configured at all, the frozen field is the only answer available");
|
||||
assertEquals("opus", composite.defaultProfileFor(MemberRole.DEV));
|
||||
}
|
||||
|
||||
// ── CB-578 stage B: a BACKEND_EXHAUSTED classification quarantines the credential ──────────
|
||||
|
||||
@Test
|
||||
@@ -974,6 +1094,98 @@ class CompositePeerLauncherTest {
|
||||
assertEquals(0, adapter.spawnCount("sol"));
|
||||
}
|
||||
|
||||
/**
|
||||
* fleetd #425 rework, acceptance 1: {@link CompositePeerLauncher#routedProfileFor} must apply
|
||||
* the SAME quarantine filtering {@link CompositePeerLauncher#spawn} does, under {@code fixed()}
|
||||
* — the DEFAULT placement policy, deliberately not {@code weighted()} (which the regressed
|
||||
* round's own tests all used, and which never exercises {@code FixedPlacementPolicy}'s own
|
||||
* inline filter). This is the exact defect: the previous round's {@code defaultProfileFor}
|
||||
* blindly returns the pool's first entry ("sol", quarantined here) with no awareness of
|
||||
* quarantine at all, which is what turned a routine unqualified spawn into a hard throw once
|
||||
* {@code acquireWithWorktree} pre-resolved through it.
|
||||
*/
|
||||
@Test
|
||||
void routedProfileForSkipsAQuarantinedPoolFirstProfileUnderFixedPolicy() {
|
||||
FakeHerdr herdr = new FakeHerdr();
|
||||
Map<String, FleetConfig.Profile> profiles = ordered(
|
||||
"sol", stubWorker("sol", "shared-openai"),
|
||||
"b", stubWorker("b"));
|
||||
StubLauncher adapter = new StubLauncher("claude", herdr, profiles, "sol", Set.of());
|
||||
BackendQuarantine quarantine = new BackendQuarantine(() -> 0L, TimeUnit.MINUTES.toNanos(30));
|
||||
quarantine.quarantine("shared-openai");
|
||||
CompositePeerLauncher composite = new CompositePeerLauncher(List.of(adapter), "sol", profiles,
|
||||
PlacementPolicies.fixed(), _ -> 0, null, quarantine);
|
||||
|
||||
assertEquals("b", composite.routedProfileFor(MemberRole.DEV),
|
||||
"sol (the pool's first entry) is quarantined, so the routed answer must be b");
|
||||
assertEquals("sol", composite.defaultProfileFor(MemberRole.DEV),
|
||||
"sanity: defaultProfileFor stays blind to quarantine — that's the gap routedProfileFor closes");
|
||||
assertEquals(0, adapter.spawnCount("sol"), "routedProfileFor never spawns anything");
|
||||
assertEquals(0, adapter.spawnCount("b"), "routedProfileFor never spawns anything");
|
||||
}
|
||||
|
||||
/**
|
||||
* fleetd #444: {@link PlacementDecision} exists to close the window between {@link
|
||||
* CompositePeerLauncher#place} and {@link CompositePeerLauncher#spawn(SpawnRequest,
|
||||
* PlacementDecision)} — the placement state must be free to move in that window without the
|
||||
* held decision being re-checked against the new state. Every quarantine test above resolves
|
||||
* and spawns in one call, so none of them ever open that window; this test is the one that
|
||||
* does: "sol" is placed FIRST, while nothing is quarantined yet, and only THEN is its
|
||||
* credential quarantined, before the held decision is spawned.
|
||||
*
|
||||
* <p>This is the test that tells the real override apart from the alternative body the ticket
|
||||
* measured: routing {@code decision.profile()} straight to its adapter (the real override)
|
||||
* never re-runs {@code enforceNotQuarantined}, so the spawn against the held decision still
|
||||
* succeeds on sol. Re-entering {@code spawn(req.withProfile(decision.profile()))} instead
|
||||
* lands in the explicit-profile branch, which refuses a now-quarantined sol outright — before
|
||||
* this test existed, replacing the real override's body with that re-entering call left the
|
||||
* whole suite green.
|
||||
*/
|
||||
@Test
|
||||
void spawnHonorsAPlacementDecisionEvenAfterItsProfileIsQuarantinedInTheWindowAfterPlace() {
|
||||
FakeHerdr herdr = new FakeHerdr();
|
||||
Map<String, FleetConfig.Profile> profiles = ordered(
|
||||
"sol", stubWorker("sol", "shared-openai"),
|
||||
"b", stubWorker("b"));
|
||||
// The adapter's OWN fallback default is "b", deliberately different from the profile place()
|
||||
// decides ("sol") — see the note below on why this must not be "sol" too.
|
||||
StubLauncher adapter = new StubLauncher("claude", herdr, profiles, "b", Set.of());
|
||||
BackendQuarantine quarantine = new BackendQuarantine(() -> 0L, TimeUnit.MINUTES.toNanos(30));
|
||||
CompositePeerLauncher composite = new CompositePeerLauncher(List.of(adapter), "sol", profiles,
|
||||
PlacementPolicies.fixed(), _ -> 0, null, quarantine);
|
||||
|
||||
// 1. Resolve BEFORE anything is quarantined — sol (definition order first, fixed policy) wins.
|
||||
// composite's own defaultProfile ("sol", the constructor arg above) never enters this: the
|
||||
// pool poolFor(DEV) resolves to is never empty here, so place() only ever reads that field as
|
||||
// a fallback for an empty pool, which this test does not exercise.
|
||||
PlacementDecision decision = composite.place(MemberRole.DEV);
|
||||
assertEquals("sol", decision.profile(), "sanity: nothing is quarantined yet, so sol is placed");
|
||||
|
||||
// 2. Move the placement state IN THE WINDOW between place() and spawn() — sol's credential
|
||||
// is now quarantined. A fresh place()/spawn(req) pair would fall through to b instead; the
|
||||
// held decision must not be re-evaluated against this new state at all.
|
||||
quarantine.quarantine("shared-openai");
|
||||
|
||||
// 3. Spawn against the HELD decision, not a fresh resolve.
|
||||
SpawnRequest req = new SpawnRequest(null, null, null, null, null, MemberRole.DEV);
|
||||
PeerHandle handle = composite.spawn(req, decision);
|
||||
|
||||
assertEquals("sol", handle.profile(),
|
||||
"the decision from place() is honored even though sol is now quarantined");
|
||||
// A fixture whose adapter falls back to "sol" too would let an UNSTAMPED request (one
|
||||
// routed but never given req.withProfile("sol")) land on spawnCount("sol") == 1 by
|
||||
// COINCIDENCE, since StubLauncher.spawn falls back to its own defaultProfile whenever
|
||||
// req.profileName() is blank. Giving the adapter "b" as its fallback instead means only an
|
||||
// actually-stamped request can produce this count — an unstamped one would count against
|
||||
// "b" and this assertion would fail.
|
||||
assertEquals(1, adapter.spawnCount("sol"),
|
||||
"the request that reached the delegate actually carried sol as its profile "
|
||||
+ "(the adapter's own fallback default is 'b', so this can't happen by accident)");
|
||||
assertEquals(0, adapter.spawnCount("b"),
|
||||
"b must never be touched — neither as the decision's profile nor as an unstamped "
|
||||
+ "request's accidental fallback");
|
||||
}
|
||||
|
||||
@Test
|
||||
void aQuarantineLiftsOnTheInjectedClockAndTheProfileBecomesSpawnableAgain() {
|
||||
FakeHerdr herdr = new FakeHerdr();
|
||||
@@ -1189,4 +1401,415 @@ class CompositePeerLauncherTest {
|
||||
assertDoesNotThrow(() -> composite.spawn(new SpawnRequest("claude", null, null)));
|
||||
assertDoesNotThrow(() -> composite.spawn(new SpawnRequest("gemini", null, null)));
|
||||
}
|
||||
|
||||
// ── fleetd #422: the model on/off gate — a FOURTH, independent reason to refuse a spawn ──────
|
||||
// ── (an operator decision, never a backend-reported outage) — never merged with quarantine ────
|
||||
// ── or cool-off above ─────────────────────────────────────────────────────────────────────────
|
||||
|
||||
private static final BackendOutagePolicy NO_OUTAGE = new BackendOutagePolicy(() -> 0L);
|
||||
|
||||
/** Criterion 1: an explicit spawn onto a profile whose model is off is refused. */
|
||||
@Test
|
||||
void explicitSpawnOntoAnOffModelProfileIsRefused() {
|
||||
FakeHerdr herdr = new FakeHerdr();
|
||||
Map<String, FleetConfig.Profile> profiles = ordered(
|
||||
"local", stubWorkerModel("local", "deepseek-v4-flash"),
|
||||
"sonnet", stubWorkerModel("sonnet", "claude-sonnet-5"));
|
||||
StubLauncher adapter = new StubLauncher("claude", herdr, profiles, "local", Set.of());
|
||||
FleetConfig.Models models = new FleetConfig.Models(List.of(
|
||||
new FleetConfig.Models.ModelEntry("deepseek-v4-flash", false)));
|
||||
CompositePeerLauncher composite = new CompositePeerLauncher(List.of(adapter), "local", profiles,
|
||||
PlacementPolicies.fixed(), _ -> 0, null, BackendQuarantine.none(), NO_OUTAGE,
|
||||
() -> models);
|
||||
|
||||
PlacementException e = assertThrows(PlacementException.class,
|
||||
() -> composite.spawn(new SpawnRequest("local", null, null)));
|
||||
assertTrue(e.getMessage().contains("local"), "message names the profile: " + e.getMessage());
|
||||
assertTrue(e.getMessage().contains("deepseek-v4-flash"),
|
||||
"message names the model: " + e.getMessage());
|
||||
assertTrue(e.getMessage().contains("operator") && e.getMessage().contains("turned off"),
|
||||
"wording says the OPERATOR turned it off: " + e.getMessage());
|
||||
// Distinct from quarantine/cool-off wording (criterion 1's explicit requirement).
|
||||
assertFalse(e.getMessage().contains("quarantined"), "must not read like quarantine: " + e.getMessage());
|
||||
assertFalse(e.getMessage().contains("cooling off"), "must not read like cool-off: " + e.getMessage());
|
||||
assertEquals(0, adapter.spawnCount("local"), "the off-model profile is never delegated to");
|
||||
}
|
||||
|
||||
/** A profile whose model is NOT off spawns normally even while another model is off. */
|
||||
@Test
|
||||
void explicitSpawnOntoAnEnabledModelProfileSucceedsWhileAnotherModelIsOff() {
|
||||
FakeHerdr herdr = new FakeHerdr();
|
||||
Map<String, FleetConfig.Profile> profiles = ordered(
|
||||
"local", stubWorkerModel("local", "deepseek-v4-flash"),
|
||||
"sonnet", stubWorkerModel("sonnet", "claude-sonnet-5"));
|
||||
StubLauncher adapter = new StubLauncher("claude", herdr, profiles, "local", Set.of());
|
||||
FleetConfig.Models models = new FleetConfig.Models(List.of(
|
||||
new FleetConfig.Models.ModelEntry("deepseek-v4-flash", false)));
|
||||
CompositePeerLauncher composite = new CompositePeerLauncher(List.of(adapter), "local", profiles,
|
||||
PlacementPolicies.fixed(), _ -> 0, null, BackendQuarantine.none(), NO_OUTAGE,
|
||||
() -> models);
|
||||
|
||||
assertDoesNotThrow(() -> composite.spawn(new SpawnRequest("sonnet", null, null)));
|
||||
assertEquals(1, adapter.spawnCount("sonnet"));
|
||||
}
|
||||
|
||||
/**
|
||||
* Criterion 8: one {@code enabled: false} entry disables EVERY profile naming that model — the
|
||||
* live example, {@code deepseek-v4-flash} backing both {@code local} and {@code local-direct}.
|
||||
*/
|
||||
@Test
|
||||
void turningOffOneModelRefusesEveryProfileThatNamesIt() {
|
||||
FakeHerdr herdr = new FakeHerdr();
|
||||
Map<String, FleetConfig.Profile> profiles = ordered(
|
||||
"local", stubWorkerModel("local", "deepseek-v4-flash"),
|
||||
"local-direct", stubWorkerModel("local-direct", "deepseek-v4-flash"));
|
||||
StubLauncher adapter = new StubLauncher("claude", herdr, profiles, "local", Set.of());
|
||||
FleetConfig.Models models = new FleetConfig.Models(List.of(
|
||||
new FleetConfig.Models.ModelEntry("deepseek-v4-flash", false)));
|
||||
CompositePeerLauncher composite = new CompositePeerLauncher(List.of(adapter), "local", profiles,
|
||||
PlacementPolicies.fixed(), _ -> 0, null, BackendQuarantine.none(), NO_OUTAGE,
|
||||
() -> models);
|
||||
|
||||
assertThrows(PlacementException.class, () -> composite.spawn(new SpawnRequest("local", null, null)));
|
||||
assertThrows(PlacementException.class,
|
||||
() -> composite.spawn(new SpawnRequest("local-direct", null, null)),
|
||||
"local-direct shares local's model, so it must be refused too");
|
||||
assertEquals(0, adapter.spawnCount("local"));
|
||||
assertEquals(0, adapter.spawnCount("local-direct"));
|
||||
}
|
||||
|
||||
/** Criterion 3: an old-style entry with no {@code enabled} field never refuses a spawn. */
|
||||
@Test
|
||||
void anEntryWithNoEnabledFieldNeverRefusesASpawn() {
|
||||
FakeHerdr herdr = new FakeHerdr();
|
||||
Map<String, FleetConfig.Profile> profiles = Map.of("local", stubWorkerModel("local", "deepseek-v4-flash"));
|
||||
StubLauncher adapter = new StubLauncher("claude", herdr, profiles, "local", Set.of());
|
||||
// The back-compat single-arg ModelEntry constructor — no enabled field in the shape at all.
|
||||
FleetConfig.Models models = new FleetConfig.Models(
|
||||
List.of(new FleetConfig.Models.ModelEntry("deepseek-v4-flash")));
|
||||
CompositePeerLauncher composite = new CompositePeerLauncher(List.of(adapter), "local", profiles,
|
||||
PlacementPolicies.fixed(), _ -> 0, null, BackendQuarantine.none(), NO_OUTAGE,
|
||||
() -> models);
|
||||
|
||||
assertDoesNotThrow(() -> composite.spawn(new SpawnRequest("local", null, null)));
|
||||
assertEquals(1, adapter.spawnCount("local"));
|
||||
}
|
||||
|
||||
/** A model absent from {@code models.allow:} entirely (nothing to gate against) is never refused. */
|
||||
@Test
|
||||
void aModelNotConfiguredInTheAllowListIsNeverGated() {
|
||||
FakeHerdr herdr = new FakeHerdr();
|
||||
Map<String, FleetConfig.Profile> profiles = Map.of("local", stubWorkerModel("local", "unlisted-model"));
|
||||
StubLauncher adapter = new StubLauncher("claude", herdr, profiles, "local", Set.of());
|
||||
FleetConfig.Models models = new FleetConfig.Models(List.of(
|
||||
new FleetConfig.Models.ModelEntry("deepseek-v4-flash", false)));
|
||||
CompositePeerLauncher composite = new CompositePeerLauncher(List.of(adapter), "local", profiles,
|
||||
PlacementPolicies.fixed(), _ -> 0, null, BackendQuarantine.none(), NO_OUTAGE,
|
||||
() -> models);
|
||||
|
||||
assertDoesNotThrow(() -> composite.spawn(new SpawnRequest("local", null, null)));
|
||||
}
|
||||
|
||||
/** Criterion 2: an unqualified spawn skips an off-model candidate and lands on another one. */
|
||||
@Test
|
||||
void placementSkipsAnOffModelProfileAndRoutesToAnotherOne() {
|
||||
FakeHerdr herdr = new FakeHerdr();
|
||||
Map<String, FleetConfig.Profile> profiles = ordered(
|
||||
"local", stubWorkerModel("local", "deepseek-v4-flash"),
|
||||
"sonnet", stubWorkerModel("sonnet", "claude-sonnet-5"));
|
||||
StubLauncher adapter = new StubLauncher("claude", herdr, profiles, "local", Set.of());
|
||||
FleetConfig.Models models = new FleetConfig.Models(List.of(
|
||||
new FleetConfig.Models.ModelEntry("deepseek-v4-flash", false)));
|
||||
CompositePeerLauncher composite = new CompositePeerLauncher(List.of(adapter), "local", profiles,
|
||||
PlacementPolicies.weighted(), _ -> 0, null, BackendQuarantine.none(), NO_OUTAGE,
|
||||
() -> models);
|
||||
|
||||
PeerHandle h = composite.spawn(new SpawnRequest(null, null, null));
|
||||
assertEquals("sonnet", h.profile(), "local's model is off, so an unqualified spawn must land on sonnet");
|
||||
assertEquals(0, adapter.spawnCount("local"));
|
||||
}
|
||||
|
||||
/**
|
||||
* Criterion 2's second half: when EVERY candidate's model is off, the placement exception must
|
||||
* name that as the cause — not a generic "no candidates" message.
|
||||
*/
|
||||
@Test
|
||||
void automaticPlacementNamesModelOffWhenEveryCandidateIsOffModel() {
|
||||
FakeHerdr herdr = new FakeHerdr();
|
||||
Map<String, FleetConfig.Profile> profiles = ordered(
|
||||
"local", stubWorkerModel("local", "deepseek-v4-flash"),
|
||||
"local-direct", stubWorkerModel("local-direct", "deepseek-v4-flash"));
|
||||
StubLauncher adapter = new StubLauncher("claude", herdr, profiles, "local", Set.of());
|
||||
FleetConfig.Models models = new FleetConfig.Models(List.of(
|
||||
new FleetConfig.Models.ModelEntry("deepseek-v4-flash", false)));
|
||||
CompositePeerLauncher composite = new CompositePeerLauncher(List.of(adapter), "local", profiles,
|
||||
PlacementPolicies.weighted(), _ -> 0, null, BackendQuarantine.none(), NO_OUTAGE,
|
||||
() -> models);
|
||||
|
||||
PlacementException e = assertThrows(PlacementException.class,
|
||||
() -> composite.spawn(new SpawnRequest(null, null, null)),
|
||||
"both candidates share the off model — nothing is available");
|
||||
assertTrue(e.getMessage().contains("model") && e.getMessage().contains("turned off"),
|
||||
"message names model-off as the cause, not a generic no-candidates message: "
|
||||
+ e.getMessage());
|
||||
}
|
||||
|
||||
/**
|
||||
* fleetd #422 follow-up: {@code fixed} is the DEFAULT placement policy ({@code
|
||||
* PlacementPolicies.fromName} returns it for an absent/blank name), and it built its own inline
|
||||
* candidate filter instead of calling {@code PlacementPolicyUtil.available()} — so it never
|
||||
* checked {@code modelOff()}. Mirrors {@code placementSkipsAnOffModelProfileAndRoutesToAnotherOne}
|
||||
* above with only the policy swapped, to prove the gate now fires on the path most fleets use.
|
||||
*/
|
||||
@Test
|
||||
void fixedPlacementSkipsAnOffModelProfileToo() {
|
||||
FakeHerdr herdr = new FakeHerdr();
|
||||
Map<String, FleetConfig.Profile> profiles = ordered(
|
||||
"local", stubWorkerModel("local", "deepseek-v4-flash"),
|
||||
"sonnet", stubWorkerModel("sonnet", "claude-sonnet-5"));
|
||||
StubLauncher adapter = new StubLauncher("claude", herdr, profiles, "local", Set.of());
|
||||
FleetConfig.Models models = new FleetConfig.Models(List.of(
|
||||
new FleetConfig.Models.ModelEntry("deepseek-v4-flash", false)));
|
||||
CompositePeerLauncher composite = new CompositePeerLauncher(List.of(adapter), "local", profiles,
|
||||
PlacementPolicies.fixed(), _ -> 0, null, BackendQuarantine.none(), NO_OUTAGE,
|
||||
() -> models);
|
||||
|
||||
PeerHandle h = composite.spawn(new SpawnRequest(null, null, null));
|
||||
assertEquals("sonnet", h.profile(), "local's model is off, so an unqualified spawn must land on sonnet");
|
||||
assertEquals(0, adapter.spawnCount("local"));
|
||||
}
|
||||
|
||||
/**
|
||||
* fleetd #425 rework, acceptance 2: same shape as {@link #fixedPlacementSkipsAnOffModelProfileToo}
|
||||
* above, but through {@link CompositePeerLauncher#routedProfileFor} rather than an actual
|
||||
* {@link CompositePeerLauncher#spawn} — the exact call {@code SessionManager.acquireWithWorktree}
|
||||
* makes to pre-resolve a profile for provisioning. This is the fleetd #429 case named in the
|
||||
* ticket: an operator turns a model off, and an unqualified worktree spawn must still route
|
||||
* around it instead of throwing "names model, which the operator has turned off" — the throw
|
||||
* {@link CompositePeerLauncher#enforceModelEnabled} raises only on the EXPLICIT-profile branch,
|
||||
* which is exactly the branch the regressed round accidentally routed every worktree spawn onto.
|
||||
*/
|
||||
@Test
|
||||
void routedProfileForSkipsAModelOffPoolFirstProfileUnderFixedPolicy() {
|
||||
FakeHerdr herdr = new FakeHerdr();
|
||||
Map<String, FleetConfig.Profile> profiles = ordered(
|
||||
"local", stubWorkerModel("local", "deepseek-v4-flash"),
|
||||
"sonnet", stubWorkerModel("sonnet", "claude-sonnet-5"));
|
||||
StubLauncher adapter = new StubLauncher("claude", herdr, profiles, "local", Set.of());
|
||||
FleetConfig.Models models = new FleetConfig.Models(List.of(
|
||||
new FleetConfig.Models.ModelEntry("deepseek-v4-flash", false)));
|
||||
CompositePeerLauncher composite = new CompositePeerLauncher(List.of(adapter), "local", profiles,
|
||||
PlacementPolicies.fixed(), _ -> 0, null, BackendQuarantine.none(), NO_OUTAGE,
|
||||
() -> models);
|
||||
|
||||
assertEquals("sonnet", composite.routedProfileFor(MemberRole.DEV),
|
||||
"local (the pool's first entry) names an off model, so the routed answer must be sonnet");
|
||||
assertEquals("local", composite.defaultProfileFor(MemberRole.DEV),
|
||||
"sanity: defaultProfileFor stays blind to model-off — that's the gap routedProfileFor closes");
|
||||
assertEquals(0, adapter.spawnCount("local"), "routedProfileFor never spawns anything");
|
||||
assertEquals(0, adapter.spawnCount("sonnet"), "routedProfileFor never spawns anything");
|
||||
}
|
||||
|
||||
/**
|
||||
* Criterion 4: turning a model off/on is HOT — no restart — proven through a REAL
|
||||
* {@code ConfigRef.reload()}, not a hand-rolled supplier swap. Also proves {@code models} is
|
||||
* correctly reclassified: the reload's {@link ConfigRef.Outcome#applied()} is {@code true} and
|
||||
* {@code "models"} never appears in {@link ConfigRef.Outcome#deferred()}.
|
||||
*/
|
||||
@Test
|
||||
void modelOnOffIsHotReloadedThroughARealConfigRef(@TempDir Path dir) throws Exception {
|
||||
Path yaml = dir.resolve("fleetd.yaml");
|
||||
Files.writeString(yaml, """
|
||||
bind:
|
||||
port: 8080
|
||||
profiles:
|
||||
local:
|
||||
baseUrl: http://local.gw:8000
|
||||
model: deepseek-v4-flash
|
||||
models:
|
||||
allow:
|
||||
- model: deepseek-v4-flash
|
||||
""");
|
||||
FleetConfig initial = FleetConfig.load(yaml);
|
||||
ConfigRef configRef = new ConfigRef(yaml, initial);
|
||||
|
||||
FakeHerdr herdr = new FakeHerdr();
|
||||
StubLauncher adapter = new StubLauncher("claude", herdr,
|
||||
Map.of("local", stubWorker("local")), "local", Set.of());
|
||||
CompositePeerLauncher composite = new CompositePeerLauncher(
|
||||
List.of(adapter), "local", configRef, _ -> 0, BackendQuarantine.none(), NO_OUTAGE);
|
||||
|
||||
assertDoesNotThrow(() -> composite.spawn(new SpawnRequest("local", null, null)),
|
||||
"the model starts enabled");
|
||||
assertEquals(1, adapter.spawnCount("local"));
|
||||
|
||||
Files.writeString(yaml, """
|
||||
bind:
|
||||
port: 8080
|
||||
profiles:
|
||||
local:
|
||||
baseUrl: http://local.gw:8000
|
||||
model: deepseek-v4-flash
|
||||
models:
|
||||
allow:
|
||||
- model: deepseek-v4-flash
|
||||
enabled: false
|
||||
""");
|
||||
ConfigRef.Outcome outcome = configRef.reload();
|
||||
assertTrue(outcome.applied(), "a models.allow on/off edit must apply live, never be refused");
|
||||
assertFalse(outcome.deferred().contains("models"),
|
||||
"models is hot-excluded now — it must never be reported as a deferred key");
|
||||
|
||||
PlacementException e = assertThrows(PlacementException.class,
|
||||
() -> composite.spawn(new SpawnRequest("local", null, null)),
|
||||
"the very next spawn must see the reload, with no restart");
|
||||
assertTrue(e.getMessage().contains("deepseek-v4-flash"));
|
||||
assertEquals(1, adapter.spawnCount("local"), "still just the one successful spawn from before");
|
||||
|
||||
// And back on, still hot, still no restart.
|
||||
Files.writeString(yaml, """
|
||||
bind:
|
||||
port: 8080
|
||||
profiles:
|
||||
local:
|
||||
baseUrl: http://local.gw:8000
|
||||
model: deepseek-v4-flash
|
||||
models:
|
||||
allow:
|
||||
- model: deepseek-v4-flash
|
||||
enabled: true
|
||||
""");
|
||||
ConfigRef.Outcome reEnabled = configRef.reload();
|
||||
assertTrue(reEnabled.applied());
|
||||
assertDoesNotThrow(() -> composite.spawn(new SpawnRequest("local", null, null)));
|
||||
assertEquals(2, adapter.spawnCount("local"));
|
||||
}
|
||||
|
||||
/**
|
||||
* fleetd #422 follow-up, acceptance criterion 2: {@link CompositePeerLauncher#modelGateState()}
|
||||
* is LIVE — no restart — proven through a REAL {@link ConfigRef#reload()}, exactly like {@link
|
||||
* #modelOnOffIsHotReloadedThroughARealConfigRef} above proves for the on/off gate itself. This
|
||||
* single reload sequence walks through all three states the ticket asks for: no {@code models:}
|
||||
* block, a block armed with nothing off, and a block with one model off — so a reload that flips
|
||||
* between any of the three is proven live, not just the on/off edit within an already-armed block.
|
||||
*/
|
||||
@Test
|
||||
void modelGateStateIsHotReloadedThroughARealConfigRef(@TempDir Path dir) throws Exception {
|
||||
Path yaml = dir.resolve("fleetd.yaml");
|
||||
Files.writeString(yaml, """
|
||||
bind:
|
||||
port: 8080
|
||||
profiles:
|
||||
local:
|
||||
baseUrl: http://local.gw:8000
|
||||
model: deepseek-v4-flash
|
||||
""");
|
||||
FleetConfig initial = FleetConfig.load(yaml);
|
||||
ConfigRef configRef = new ConfigRef(yaml, initial);
|
||||
|
||||
FakeHerdr herdr = new FakeHerdr();
|
||||
StubLauncher adapter = new StubLauncher("claude", herdr,
|
||||
Map.of("local", stubWorker("local")), "local", Set.of());
|
||||
CompositePeerLauncher composite = new CompositePeerLauncher(
|
||||
List.of(adapter), "local", configRef, _ -> 0, BackendQuarantine.none(), NO_OUTAGE);
|
||||
|
||||
PeerLauncher.ModelGateState notConfigured = composite.modelGateState();
|
||||
assertFalse(notConfigured.configured(), "no models: block in the config at all");
|
||||
assertEquals(Set.of(), notConfigured.off());
|
||||
|
||||
Files.writeString(yaml, """
|
||||
bind:
|
||||
port: 8080
|
||||
profiles:
|
||||
local:
|
||||
baseUrl: http://local.gw:8000
|
||||
model: deepseek-v4-flash
|
||||
models:
|
||||
allow:
|
||||
- model: deepseek-v4-flash
|
||||
enabled: false
|
||||
""");
|
||||
assertTrue(configRef.reload().applied(), "adding a models: block must apply live, no restart");
|
||||
PeerLauncher.ModelGateState armedWithOneOff = composite.modelGateState();
|
||||
assertTrue(armedWithOneOff.configured(), "a models: block now exists — the gate is armed");
|
||||
assertEquals(Set.of("deepseek-v4-flash"), armedWithOneOff.off());
|
||||
|
||||
Files.writeString(yaml, """
|
||||
bind:
|
||||
port: 8080
|
||||
profiles:
|
||||
local:
|
||||
baseUrl: http://local.gw:8000
|
||||
model: deepseek-v4-flash
|
||||
models:
|
||||
allow:
|
||||
- model: deepseek-v4-flash
|
||||
enabled: true
|
||||
""");
|
||||
assertTrue(configRef.reload().applied(), "flipping the entry back on must apply live too");
|
||||
PeerLauncher.ModelGateState armedWithZeroOff = composite.modelGateState();
|
||||
assertTrue(armedWithZeroOff.configured(),
|
||||
"the block is still present — armed and reporting zero, not the same as no block at all");
|
||||
assertEquals(Set.of(), armedWithZeroOff.off());
|
||||
}
|
||||
|
||||
/**
|
||||
* fleetd #422 follow-up, acceptance criterion 3: the invariant is that an absent {@code
|
||||
* models:} block stays permitted and must never be fatal. Proved, not assumed — a config
|
||||
* without one loads, validates, reports the gate as not configured, AND still spawns normally
|
||||
* (no {@link PlacementException} from a gate that has nothing to check against), using the same
|
||||
* production-shaped {@code Supplier<FleetConfig>} wiring {@code Fleetd.main} actually uses.
|
||||
*/
|
||||
@Test
|
||||
void noModelsBlockConfigStillLoadsAndSpawnsNormally(@TempDir Path dir) throws Exception {
|
||||
Path yaml = dir.resolve("fleetd.yaml");
|
||||
Files.writeString(yaml, """
|
||||
bind:
|
||||
port: 8080
|
||||
profiles:
|
||||
local:
|
||||
baseUrl: http://local.gw:8000
|
||||
model: deepseek-v4-flash
|
||||
""");
|
||||
FleetConfig cfg = FleetConfig.load(yaml);
|
||||
assertDoesNotThrow(cfg::validateAll, "a config with no models: block must load and validate cleanly");
|
||||
ConfigRef configRef = new ConfigRef(yaml, cfg);
|
||||
|
||||
FakeHerdr herdr = new FakeHerdr();
|
||||
StubLauncher adapter = new StubLauncher("claude", herdr,
|
||||
Map.of("local", stubWorker("local")), "local", Set.of());
|
||||
CompositePeerLauncher composite = new CompositePeerLauncher(
|
||||
List.of(adapter), "local", configRef, _ -> 0, BackendQuarantine.none(), NO_OUTAGE);
|
||||
|
||||
assertFalse(composite.modelGateState().configured());
|
||||
assertDoesNotThrow(() -> composite.spawn(new SpawnRequest("local", null, null)),
|
||||
"no models: block means nothing to gate against — the spawn must go through");
|
||||
assertEquals(1, adapter.spawnCount("local"));
|
||||
}
|
||||
|
||||
/** {@code fleet_profiles}/{@code GET /profiles} must read the exact same live source the gate reads. */
|
||||
@Test
|
||||
void disabledModelsReportsWhatTheGateActuallyEnforces() {
|
||||
FakeHerdr herdr = new FakeHerdr();
|
||||
Map<String, FleetConfig.Profile> profiles = ordered(
|
||||
"local", stubWorkerModel("local", "deepseek-v4-flash"),
|
||||
"sonnet", stubWorkerModel("sonnet", "claude-sonnet-5"));
|
||||
StubLauncher adapter = new StubLauncher("claude", herdr, profiles, "local", Set.of());
|
||||
FleetConfig.Models models = new FleetConfig.Models(List.of(
|
||||
new FleetConfig.Models.ModelEntry("deepseek-v4-flash", false),
|
||||
new FleetConfig.Models.ModelEntry("claude-sonnet-5", true)));
|
||||
CompositePeerLauncher composite = new CompositePeerLauncher(List.of(adapter), "local", profiles,
|
||||
PlacementPolicies.fixed(), _ -> 0, null, BackendQuarantine.none(), NO_OUTAGE,
|
||||
() -> models);
|
||||
|
||||
assertEquals(Set.of("deepseek-v4-flash"), composite.disabledModels());
|
||||
}
|
||||
|
||||
/** A caller with no {@code models:} block (the map-based constructors) reports nothing off. */
|
||||
@Test
|
||||
void disabledModelsIsEmptyWithNoModelsConfigured() {
|
||||
FakeHerdr herdr = new FakeHerdr();
|
||||
PeerLauncher composite = composite(herdr);
|
||||
assertEquals(Set.of(), composite.disabledModels());
|
||||
}
|
||||
}
|
||||
|
||||
@@ -120,40 +120,202 @@ class EnvAllowListScrubTest {
|
||||
}
|
||||
|
||||
/**
|
||||
* A pane inherits {@code UID}; a cleared test parent does not. The scrub must survive it.
|
||||
* fleetd #394: the actual defect. Plain {@code export "$n="} is FATAL for a zsh read-only or
|
||||
* special parameter (e.g. {@code UID}) and aborts the whole sourced file — every name still to
|
||||
* come is never blanked, and the {@code scrub-report.txt} below is never written at all,
|
||||
* silently ({@code 2>/dev/null} swallows the error). This plants an unblankable, exported,
|
||||
* read-only variable in the MIDDLE of the names the scrub attempts to blank, with two more
|
||||
* names after it, and asserts that both of those later names are STILL blanked and the report
|
||||
* is STILL written with the failure counted — a test that only checked names BEFORE the failure
|
||||
* point would pass today and prove nothing.
|
||||
*
|
||||
* <p>Every other test here starts zsh from a CLEARED environment, so {@code UID} is never an
|
||||
* exported name and never reaches the blanking loop. In a real member pane it is exported and
|
||||
* it IS reached — and {@code export UID=} is a fatal zsh parameter error that aborts the whole
|
||||
* sourced file, leaving every later name unscrubbed and writing no report at all. The abort is
|
||||
* silent: the loop is wrapped in {@code 2>/dev/null}.
|
||||
* <p>The planted name is a made-up one ({@code FLEETD_TEST_UNBLANKABLE}), not {@code UID} or
|
||||
* any other name a skip-list might already know about — invariant 1 is that the loop survives
|
||||
* ANY unblankable name, not a known one, so the test must not lean on one either.
|
||||
*
|
||||
* <p>The assertion is deliberately "a report exists" rather than "the canary is blanked". The
|
||||
* report is written by the last statement in the file, so its presence proves the script ran
|
||||
* to completion; the canary alone would depend on where it happens to sit in {@code env} order.
|
||||
* Both are checked, but only the first one fails deterministically without the fix.
|
||||
* <p>Exercises the real artefact: {@link EnvAllowListScrub#scrubScript} is run verbatim under a
|
||||
* real {@code /bin/zsh}, not just asserted on as a Java string. The four planted names are
|
||||
* exported one at a time via {@code typeset -x}/{@code typeset -rx} immediately before the
|
||||
* script runs, in a fixed order — zsh's {@code export}/{@code typeset -x} appends to the
|
||||
* process's environment table in call order (verified empirically: a freshly-exported name
|
||||
* always sorts after every inherited one and after every earlier freshly-exported name in
|
||||
* {@code command env}'s own output), which is what makes the "middle" position deterministic
|
||||
* here, unlike relying on the OS's own inherited-environment order.
|
||||
*/
|
||||
@Test
|
||||
void scrubSurvivesAnInheritedUidTheWayARealPaneHasIt(@TempDir Path tmp) throws Exception {
|
||||
void unblankableNameInTheMiddleDoesNotAbortNamesAfterIt(@TempDir Path tmp) throws Exception {
|
||||
assumeTrue(Files.isExecutable(ZSH), "/bin/zsh not present — nothing to prove here");
|
||||
Set<String> allowed = MemberEnvAllowList.derive(List.of());
|
||||
Path zdotdir = EnvAllowListScrub.generate(tmp, allowed);
|
||||
|
||||
// The production shape: UID present and exported, as every pane shell inherits it.
|
||||
Map<String, String> paneLikeParent = Map.of(
|
||||
"HOME", System.getProperty("user.home"),
|
||||
"PATH", "/usr/bin:/bin",
|
||||
"SHELL", "/bin/zsh",
|
||||
"UID", "1000",
|
||||
"CB633_CANARY", "must-not-survive-the-scrub");
|
||||
Set<String> survivors = exportedNamesFromCleanParent(paneLikeParent, zdotdir);
|
||||
// Only ZDOTDIR is allowed — it must survive the scrub itself, since the report is written
|
||||
// to "$ZDOTDIR/..." AFTER the blanking loop runs; if ZDOTDIR were blanked as a side effect,
|
||||
// the report write would silently go to the wrong place instead of testing anything.
|
||||
String script = EnvAllowListScrub.scrubScript(Set.of("ZDOTDIR"));
|
||||
String setup = """
|
||||
typeset -x FLEETD_TEST_BEFORE=1
|
||||
typeset -rx FLEETD_TEST_UNBLANKABLE=1
|
||||
typeset -x FLEETD_TEST_AFTER_A=1
|
||||
typeset -x FLEETD_TEST_AFTER_B=1
|
||||
""";
|
||||
|
||||
EnvAllowListScrub.ScrubReport report = EnvAllowListScrub.readReport(zdotdir);
|
||||
assertNotNull(report,
|
||||
"an inherited UID must not abort the scrub — no report means the file died mid-loop "
|
||||
+ "and every name after UID in `env` order was left unscrubbed");
|
||||
assertFalse(survivors.contains("CB633_CANARY"),
|
||||
"a non-allow-listed name must still be blanked when UID is in the environment");
|
||||
ProcessBuilder pb = new ProcessBuilder("/bin/zsh");
|
||||
pb.environment().clear();
|
||||
pb.environment().put("PATH", "/usr/bin:/bin");
|
||||
pb.environment().put("ZDOTDIR", tmp.toAbsolutePath().toString());
|
||||
pb.redirectError(ProcessBuilder.Redirect.DISCARD);
|
||||
Process zsh = pb.start();
|
||||
zsh.getOutputStream().write((setup + script).getBytes(StandardCharsets.UTF_8));
|
||||
zsh.getOutputStream().flush();
|
||||
zsh.getOutputStream().close();
|
||||
assertTrue(zsh.waitFor(60, java.util.concurrent.TimeUnit.SECONDS),
|
||||
"the scrub script did not exit within 60s");
|
||||
assertEquals(0, zsh.exitValue(),
|
||||
"the scrub script itself must never abort — an unblankable name must not kill the "
|
||||
+ "sourced file");
|
||||
|
||||
EnvAllowListScrub.ScrubReport report = EnvAllowListScrub.readReport(tmp);
|
||||
assertNotNull(report, "the report must still be written even though one name could not be "
|
||||
+ "blanked — a report that silently never appears is the #394 bug");
|
||||
assertTrue(report.blanked().contains("FLEETD_TEST_BEFORE"),
|
||||
"sanity: the name before the unblankable one must be blanked");
|
||||
assertTrue(report.blanked().contains("FLEETD_TEST_AFTER_A"),
|
||||
"the FIRST name AFTER the unblankable one must still be blanked — before the fix, "
|
||||
+ "the whole loop aborted at the unblankable name and every later name was "
|
||||
+ "silently left untouched");
|
||||
assertTrue(report.blanked().contains("FLEETD_TEST_AFTER_B"),
|
||||
"the SECOND name after the unblankable one must also still be blanked");
|
||||
assertTrue(report.unblankable().contains("FLEETD_TEST_UNBLANKABLE"),
|
||||
"the unblankable name is reported by name, not silently dropped");
|
||||
// fleetd #400 note: zsh itself auto-exports SHLVL on every shell start (measured: it appears
|
||||
// in `command env` even from a fully cleared parent), and a bare assignment to it is coerced
|
||||
// rather than failing — exactly the shape #400 fixes. Before that fix, eval's exit status
|
||||
// alone silently misclassified SHLVL as blanked, so this test's old "exactly one" assertion
|
||||
// passed by accident: it never actually proved SHLVL was absent from the candidates, only
|
||||
// that the old bug hid it. Now that classification reads the value back, SHLVL and
|
||||
// FLEETD_TEST_UNBLANKABLE both correctly land in unblankable() — real, unplanted evidence
|
||||
// the #400 fix works, not just the synthetic case in the dedicated #400 test above.
|
||||
assertTrue(report.unblankable().contains("SHLVL"),
|
||||
"fleetd #400: zsh's own auto-exported SHLVL must also be reported unblankable, not "
|
||||
+ "silently miscounted as blanked");
|
||||
assertEquals(report.unblankable().size(), report.failed(),
|
||||
"the failed count must equal the number of names actually reported unblankable");
|
||||
}
|
||||
|
||||
/**
|
||||
* fleetd #400: {@code eval}'s exit status is not proof that a name was actually blanked. zsh
|
||||
* coerces a bare {@code NAME=} assignment on an integer special parameter to a number instead of
|
||||
* failing, so {@code eval} reports success while the value stays non-empty — a status-based
|
||||
* classification calls that "blanked" when it was not. This drives all three shapes a name can
|
||||
* take through the real {@code scrubScript} in ONE run: a normal, genuinely blankable name; a
|
||||
* fatal one ({@code LINENO} — deliberately not {@code UID}, so this test does not depend on the
|
||||
* harness's uid); and the silent-no-op one the ticket is about ({@code SECONDS}, rc 0 but
|
||||
* unchanged). Under the pre-#400 exit-status check, {@code SECONDS} would land in
|
||||
* {@code blanked()} — that is the exact false receipt this fix removes.
|
||||
*
|
||||
* <p>Criterion 3: a cleared {@code ProcessBuilder} parent does not, by itself, give the child
|
||||
* zsh any of these names — {@code SECONDS}/{@code LINENO} are zsh's own built-in parameters and
|
||||
* only become CANDIDATES the enumeration loop can see (i.e. show up in {@code command env}) when
|
||||
* they arrive via the process's own environment table, not merely by existing as zsh parameters
|
||||
* inside the shell. So each is put into {@code pb.environment()} explicitly, after
|
||||
* {@code clear()} — confirmed empirically first (a throwaway probe piping
|
||||
* {@code env -i PATH=... SECONDS=999 LINENO=999 FLEETD_TEST_NORMAL=1 zsh -c 'command env | cut
|
||||
* -d= -f1'}) that all three names really appear in {@code command env}'s output under exactly
|
||||
* this construction, not relying on whatever the test-runner's own ambient environment happens
|
||||
* to contain.
|
||||
*/
|
||||
@Test
|
||||
void classifiesByObservedValueNotExitStatusAcrossAllThreeShapes(@TempDir Path tmp) throws Exception {
|
||||
assumeTrue(Files.isExecutable(ZSH), "/bin/zsh not present — nothing to prove here");
|
||||
|
||||
String script = EnvAllowListScrub.scrubScript(Set.of("ZDOTDIR"));
|
||||
|
||||
ProcessBuilder pb = new ProcessBuilder("/bin/zsh");
|
||||
pb.environment().clear();
|
||||
pb.environment().put("PATH", "/usr/bin:/bin");
|
||||
pb.environment().put("ZDOTDIR", tmp.toAbsolutePath().toString());
|
||||
// Explicitly placed in the child's environment table — see the javadoc above on why a
|
||||
// cleared parent alone does not put these on the enumeration loop's candidate list.
|
||||
pb.environment().put("SECONDS", "999"); // rc 0, value coerced/unchanged — the #400 bug
|
||||
pb.environment().put("LINENO", "999"); // fatal on assignment, eval rc != 0, contained
|
||||
pb.environment().put("FLEETD_TEST_NORMAL", "1"); // genuinely blankable, the control case
|
||||
pb.redirectError(ProcessBuilder.Redirect.DISCARD);
|
||||
Process zsh = pb.start();
|
||||
zsh.getOutputStream().write(script.getBytes(StandardCharsets.UTF_8));
|
||||
zsh.getOutputStream().flush();
|
||||
zsh.getOutputStream().close();
|
||||
assertTrue(zsh.waitFor(60, java.util.concurrent.TimeUnit.SECONDS),
|
||||
"the scrub script did not exit within 60s");
|
||||
assertEquals(0, zsh.exitValue(),
|
||||
"the scrub script must still reach its end with both a fatal name and a silent "
|
||||
+ "no-op name among the candidates");
|
||||
|
||||
EnvAllowListScrub.ScrubReport report = EnvAllowListScrub.readReport(tmp);
|
||||
assertNotNull(report, "the report must still be written");
|
||||
assertTrue(report.blanked().contains("FLEETD_TEST_NORMAL"),
|
||||
"the control case: an ordinary name is genuinely blankable and must be reported so");
|
||||
assertTrue(report.unblankable().contains("LINENO"),
|
||||
"a name fatal to assign to must be reported unblankable — sanity check that "
|
||||
+ "containment still works under the new classification");
|
||||
assertTrue(report.unblankable().contains("SECONDS"),
|
||||
"the #400 defect: eval returns rc 0 for SECONDS (zsh coerces the assignment instead "
|
||||
+ "of failing) but the value is left non-empty — classifying on the observed "
|
||||
+ "value catches this; classifying on eval's exit status would have called "
|
||||
+ "this \"blanked\" and produced a false receipt");
|
||||
assertFalse(report.blanked().contains("SECONDS"),
|
||||
"SECONDS must never appear as blanked — it was never actually emptied");
|
||||
}
|
||||
|
||||
/**
|
||||
* fleetd #394 follow-up: the blanking loop's {@code eval "export ${n}="} splices {@code n} into
|
||||
* a string that zsh then interprets as shell syntax. That is only safe because every name
|
||||
* reaching {@code _cb633_blank} already passed an identifier check in the ENUMERATION loop
|
||||
* (20 lines away, in a different loop) — so the fix re-asserts the identical check immediately
|
||||
* before the {@code eval} call, rather than trusting that distant guard to keep holding.
|
||||
*
|
||||
* <p>This test plants a value with an embedded newline, exploiting the exact "junk from
|
||||
* multi-line values" gap the enumeration loop's own comment already documents: {@code command
|
||||
* env}'s text output is read line-by-line, so a value's second line becomes a spurious extra
|
||||
* "name" that was never a real exported variable. The fragment used here ({@code
|
||||
* junk.fragment}) is merely non-conforming (it contains a dot) — never command-shaped; this
|
||||
* test must never demonstrate command execution and plants no command-shaped payload.
|
||||
*
|
||||
* <p>Exercises the real artefact end-to-end: {@link EnvAllowListScrub#scrubScript} runs
|
||||
* verbatim under a real {@code /bin/zsh}, exactly as {@code generate()} would produce it — this
|
||||
* is not a synthetic call into just the blanking loop.
|
||||
*/
|
||||
@Test
|
||||
void nonIdentifierJunkFromAMultilineValueIsSkippedNotBlankedOrUnblankable(@TempDir Path tmp)
|
||||
throws Exception {
|
||||
assumeTrue(Files.isExecutable(ZSH), "/bin/zsh not present — nothing to prove here");
|
||||
|
||||
String script = EnvAllowListScrub.scrubScript(Set.of("ZDOTDIR"));
|
||||
|
||||
ProcessBuilder pb = new ProcessBuilder("/bin/zsh");
|
||||
pb.environment().clear();
|
||||
pb.environment().put("PATH", "/usr/bin:/bin");
|
||||
pb.environment().put("ZDOTDIR", tmp.toAbsolutePath().toString());
|
||||
// Embedded newline: `command env`'s own text output splits this into two lines, and the
|
||||
// second ("junk.fragment") has no "=" at all, so `cut -d= -f1` returns it unchanged as a
|
||||
// spurious candidate "name" — it was never an actual exported variable by that name.
|
||||
pb.environment().put("FLEETD_TEST_MULTILINE", "keep\njunk.fragment");
|
||||
pb.redirectError(ProcessBuilder.Redirect.DISCARD);
|
||||
Process zsh = pb.start();
|
||||
zsh.getOutputStream().write(script.getBytes(StandardCharsets.UTF_8));
|
||||
zsh.getOutputStream().flush();
|
||||
zsh.getOutputStream().close();
|
||||
assertTrue(zsh.waitFor(60, java.util.concurrent.TimeUnit.SECONDS),
|
||||
"the scrub script did not exit within 60s");
|
||||
assertEquals(0, zsh.exitValue(), "the scrub script must reach its end");
|
||||
|
||||
EnvAllowListScrub.ScrubReport report = EnvAllowListScrub.readReport(tmp);
|
||||
assertNotNull(report, "the report must still be written");
|
||||
assertTrue(report.blanked().contains("FLEETD_TEST_MULTILINE"),
|
||||
"sanity: the real, identifier-shaped variable must still be blanked normally");
|
||||
assertFalse(report.blanked().contains("junk.fragment"),
|
||||
"a non-identifier fragment is not a real variable and must never be blanked");
|
||||
assertFalse(report.unblankable().contains("junk.fragment"),
|
||||
"a non-identifier fragment must never even become a candidate the blanking loop "
|
||||
+ "attempts — it must be filtered before either guard has to catch it, so "
|
||||
+ "it is neither blanked nor counted as a failed attempt");
|
||||
}
|
||||
|
||||
/** A group-shared ZDOTDIR still lets the member truncate and write its pre-created receipt. */
|
||||
|
||||
@@ -762,6 +762,250 @@ class OpenCodeLauncherTest {
|
||||
"no IDE server when ideMcpUrl is unset");
|
||||
}
|
||||
|
||||
// --- fleetd #393: memberSkills seeding must actually reach an opencode member -------------------
|
||||
//
|
||||
// Before this fix, GitWorktrees#seedSkills copied skill folders into EVERY provisioned
|
||||
// worktree's .claude/skills/ and logged "skill seeding: N of M" regardless of which kind ended
|
||||
// up spawning into that worktree — a claim that held for kind: claude-code (the CLI discovers
|
||||
// that directory on its own) but was a guaranteed no-op for kind: opencode, which has no such
|
||||
// discovery. These tests drive the REAL GitWorktrees#add seeding path (not a hand-built
|
||||
// .claude/skills/ fixture), then spawn an opencode-kind member against the seeded worktree and
|
||||
// assert on what the member can actually consume — an instructions[] entry — not on the
|
||||
// seeding log alone. A minimal, non-hermetic git repo is enough here: unlike
|
||||
// GitWorktreesTest's own seeding tests, nothing in this file cares about core.excludesFile
|
||||
// composition, only about what lands in .claude/skills/ and whether OpenCodeLauncher reads it.
|
||||
|
||||
private static void git(Path cwd, String... args) throws Exception {
|
||||
Process p = new ProcessBuilder(prepend("git", args)).directory(cwd.toFile())
|
||||
.redirectErrorStream(true).start();
|
||||
String out = new String(p.getInputStream().readAllBytes());
|
||||
assertTrue(p.waitFor(30, TimeUnit.SECONDS), "git timed out: git " + String.join(" ", args));
|
||||
assertEquals(0, p.exitValue(), "git " + String.join(" ", args) + " failed:\n" + out);
|
||||
}
|
||||
|
||||
private static List<String> prepend(String head, String... rest) {
|
||||
List<String> cmd = new ArrayList<>();
|
||||
cmd.add(head);
|
||||
cmd.addAll(List.of(rest));
|
||||
return cmd;
|
||||
}
|
||||
|
||||
private static Path initRepo(Path dir) throws Exception {
|
||||
Files.createDirectories(dir);
|
||||
git(dir, "init", "-q", "-b", "main");
|
||||
git(dir, "config", "user.email", "test@example.invalid");
|
||||
git(dir, "config", "user.name", "Test");
|
||||
Files.writeString(dir.resolve("README.md"), "seed\n");
|
||||
git(dir, "add", "README.md");
|
||||
git(dir, "commit", "-q", "-m", "seed");
|
||||
return dir;
|
||||
}
|
||||
|
||||
@Test
|
||||
void aSeededSkillReachesTheOpencodeMembersInstructionsArray(@TempDir Path tmp) throws Exception {
|
||||
Path repo = initRepo(tmp.resolve("repo"));
|
||||
Path skillsSource = tmp.resolve("skills-src");
|
||||
Path skillFile = skillsSource.resolve("implementer").resolve("SKILL.md");
|
||||
Files.createDirectories(skillFile.getParent());
|
||||
Files.writeString(skillFile, "IMPLEMENTER PROCEDURE\n");
|
||||
|
||||
// The real seeding path (fleetd #362), not a hand-built .claude/skills/ fixture — proves
|
||||
// OpenCodeLauncher reads what GitWorktrees#add actually produced.
|
||||
dev.ltms.fleet.session.GitWorktrees worktrees =
|
||||
new dev.ltms.fleet.session.GitWorktrees(tmp.resolve("wts").toString(), null, skillsSource.toString());
|
||||
String wt = worktrees.add(repo.toString(), "cb-393-opencode", "HEAD");
|
||||
Path seededSkillMd = Path.of(wt, ".claude", "skills", "implementer", "SKILL.md");
|
||||
assertTrue(Files.exists(seededSkillMd),
|
||||
"sanity: the real seeding step must have copied the skill into the worktree");
|
||||
|
||||
Logger logger = (Logger) LoggerFactory.getLogger(OpenCodeLauncher.class);
|
||||
Level original = logger.getLevel();
|
||||
logger.setLevel(Level.INFO);
|
||||
ListAppender<ILoggingEvent> appender = new ListAppender<>();
|
||||
appender.start();
|
||||
logger.addAppender(appender);
|
||||
|
||||
String cfgPath;
|
||||
try {
|
||||
// No mcpUrl, no ideUrl, no fleet (no charter): the seeded skill alone must be enough to
|
||||
// trigger OPENCODE_CONFIG — proves the gate itself was updated, not only writeConfig's body.
|
||||
FakeHerdr herdr = new FakeHerdr();
|
||||
Path configRoot = Files.createDirectory(tmp.resolve("configs"));
|
||||
service(herdr, configRoot, opencodeIdeCfg(null, null, wt)).spawn();
|
||||
cfgPath = startEnv(herdr).get("OPENCODE_CONFIG");
|
||||
} finally {
|
||||
logger.detachAppender(appender);
|
||||
logger.setLevel(original);
|
||||
}
|
||||
|
||||
assertNotNull(cfgPath, "a seeded skill with nothing else configured must still write a config");
|
||||
JsonNode json = new ObjectMapper().readTree(Path.of(cfgPath).toFile());
|
||||
List<String> instructions = new ArrayList<>();
|
||||
json.path("instructions").forEach(n -> instructions.add(n.asText()));
|
||||
assertTrue(instructions.contains(seededSkillMd.toAbsolutePath().toString()),
|
||||
"the seeded skill's SKILL.md must be an instructions[] entry — got: " + instructions);
|
||||
|
||||
List<String> infos = appender.list.stream()
|
||||
.filter(e -> e.getLevel() == Level.INFO)
|
||||
.map(ILoggingEvent::getFormattedMessage)
|
||||
.toList();
|
||||
assertTrue(infos.stream().anyMatch(m -> m.contains("skill delivery") && m.contains("1 of 1")),
|
||||
"the launcher must log, kind-aware, that it delivered the skill — got:\n" + infos);
|
||||
}
|
||||
|
||||
@Test
|
||||
void aSkillFolderWithoutSkillMdIsNeverDeliveredAndTheLogNamesIt(@TempDir Path tmp) throws Exception {
|
||||
Path wt = Files.createDirectories(tmp.resolve("wt"));
|
||||
Path goodSkill = wt.resolve(".claude").resolve("skills").resolve("implementer");
|
||||
Files.createDirectories(goodSkill);
|
||||
Files.writeString(goodSkill.resolve("SKILL.md"), "GOOD\n");
|
||||
Path halfShipped = wt.resolve(".claude").resolve("skills").resolve("half-shipped");
|
||||
Files.createDirectories(halfShipped);
|
||||
Files.writeString(halfShipped.resolve("README.md"), "no SKILL.md here\n");
|
||||
|
||||
Logger logger = (Logger) LoggerFactory.getLogger(OpenCodeLauncher.class);
|
||||
Level original = logger.getLevel();
|
||||
logger.setLevel(Level.INFO);
|
||||
ListAppender<ILoggingEvent> appender = new ListAppender<>();
|
||||
appender.start();
|
||||
logger.addAppender(appender);
|
||||
|
||||
String cfgPath;
|
||||
try {
|
||||
FakeHerdr herdr = new FakeHerdr();
|
||||
Path configRoot = Files.createDirectory(tmp.resolve("configs"));
|
||||
service(herdr, configRoot, opencodeIdeCfg(null, null, wt.toString())).spawn();
|
||||
cfgPath = startEnv(herdr).get("OPENCODE_CONFIG");
|
||||
} finally {
|
||||
logger.detachAppender(appender);
|
||||
logger.setLevel(original);
|
||||
}
|
||||
|
||||
JsonNode json = new ObjectMapper().readTree(Path.of(cfgPath).toFile());
|
||||
List<String> instructions = new ArrayList<>();
|
||||
json.path("instructions").forEach(n -> instructions.add(n.asText()));
|
||||
assertTrue(instructions.contains(goodSkill.resolve("SKILL.md").toAbsolutePath().toString()),
|
||||
"the well-formed skill is still delivered alongside the malformed one");
|
||||
assertFalse(instructions.stream().anyMatch(i -> i.contains("half-shipped")),
|
||||
"a skill folder with no SKILL.md can never become an instructions[] entry");
|
||||
|
||||
List<String> infos = appender.list.stream()
|
||||
.filter(e -> e.getLevel() == Level.INFO)
|
||||
.map(ILoggingEvent::getFormattedMessage)
|
||||
.toList();
|
||||
assertTrue(infos.stream().anyMatch(m -> m.contains("skill delivery") && m.contains("1 of 2")
|
||||
&& m.contains("half-shipped") && m.contains("could not be delivered")),
|
||||
"the log must say plainly which folder could not be consumed and why — got:\n" + infos);
|
||||
}
|
||||
|
||||
// --- fleetd #393 follow-up: instructions[] has three writers (charter, seeded skills, IDE
|
||||
// rules), and no test above ever exercises more than one or two of them together. A writer
|
||||
// that flips from withArray (get-or-create) to putArray (create-or-REPLACE) silently deletes
|
||||
// every entry written before it — proven live on this branch's merge: switching just the
|
||||
// skills writer to putArray left the entire suite (1603 tests) green while deleting the
|
||||
// charter entry an opencode member needs for its role contract. That hazard was found by the
|
||||
// fleet01 lead and independently verified against this branch; it is not a defect in the
|
||||
// skills-delivery or logging tests above, which both hold up under their own mutations — the
|
||||
// gap is that none of them combine all three writers in one config.
|
||||
//
|
||||
// These three tests assert instructions[] CONTENT as an exact, ordered list, not a size or a
|
||||
// "contains" check: a putArray mutation can replace N entries with a different N entries of
|
||||
// the same count, so only a content comparison can tell "all three paths present" apart from
|
||||
// "two paths present that replaced the earlier ones".
|
||||
|
||||
@Test
|
||||
void instructionsArrayHoldsExactlyTheCharterWhenNothingElseWritesToIt(@TempDir Path root) throws Exception {
|
||||
FakeHerdr herdr = new FakeHerdr();
|
||||
FleetConfig.Fleet fleet = new FleetConfig.Fleet(Map.of(), Map.of(), Map.of(), Map.of(),
|
||||
Map.of("dev", "role rule"), null);
|
||||
Path cwd = Files.createDirectory(root.resolve("checkout"));
|
||||
service(herdr, root, opencodeIdeCfg(null, null, cwd.toString()), () -> fleet).spawn();
|
||||
|
||||
String cfgPath = startEnv(herdr).get("OPENCODE_CONFIG");
|
||||
assertNotNull(cfgPath, "a role charter alone still writes a config");
|
||||
JsonNode json = new ObjectMapper().readTree(Path.of(cfgPath).toFile());
|
||||
Path charter = Path.of(cfgPath).resolveSibling("member-charter.md");
|
||||
assertTrue(Files.exists(charter), "the charter file was written");
|
||||
|
||||
List<String> instructions = new ArrayList<>();
|
||||
json.path("instructions").forEach(n -> instructions.add(n.asText()));
|
||||
assertEquals(List.of(charter.toAbsolutePath().toString()), instructions,
|
||||
"with only the charter writer active, instructions[] holds exactly one entry: the "
|
||||
+ "charter — got: " + instructions);
|
||||
}
|
||||
|
||||
@Test
|
||||
void instructionsArrayHoldsCharterThenIdeRulesInOrder(@TempDir Path root) throws Exception {
|
||||
FakeHerdr herdr = new FakeHerdr();
|
||||
FleetConfig.Fleet fleet = new FleetConfig.Fleet(Map.of(), Map.of(), Map.of(), Map.of(),
|
||||
Map.of("dev", "role rule"), null);
|
||||
Path cwd = Files.createDirectory(root.resolve("checkout"));
|
||||
service(herdr, root, opencodeIdeCfg(null,
|
||||
"http://127.0.0.1:29170/index-mcp/streamable-http", cwd.toString()), () -> fleet).spawn();
|
||||
|
||||
String cfgPath = startEnv(herdr).get("OPENCODE_CONFIG");
|
||||
assertNotNull(cfgPath, "charter + IDE rules still writes a config");
|
||||
JsonNode json = new ObjectMapper().readTree(Path.of(cfgPath).toFile());
|
||||
Path charter = Path.of(cfgPath).resolveSibling("member-charter.md");
|
||||
Path rules = Path.of(cfgPath).resolveSibling("ide-rules.md");
|
||||
assertTrue(Files.exists(charter), "the charter file was written");
|
||||
assertTrue(Files.exists(rules), "the ide-rules file was written");
|
||||
|
||||
List<String> instructions = new ArrayList<>();
|
||||
json.path("instructions").forEach(n -> instructions.add(n.asText()));
|
||||
assertEquals(List.of(charter.toAbsolutePath().toString(), rules.toAbsolutePath().toString()),
|
||||
instructions,
|
||||
"with charter + IDE-rules writers active, instructions[] holds both, charter first — "
|
||||
+ "got: " + instructions);
|
||||
}
|
||||
|
||||
@Test
|
||||
void instructionsArrayHoldsCharterThenSkillsThenIdeRulesInOrder(@TempDir Path tmp) throws Exception {
|
||||
Path repo = initRepo(tmp.resolve("repo"));
|
||||
Path skillsSource = tmp.resolve("skills-src");
|
||||
Path skillFile = skillsSource.resolve("implementer").resolve("SKILL.md");
|
||||
Files.createDirectories(skillFile.getParent());
|
||||
Files.writeString(skillFile, "IMPLEMENTER PROCEDURE\n");
|
||||
|
||||
dev.ltms.fleet.session.GitWorktrees worktrees = new dev.ltms.fleet.session.GitWorktrees(
|
||||
tmp.resolve("wts").toString(), null, skillsSource.toString());
|
||||
String wt = worktrees.add(repo.toString(), "cb-393-follow-up", "HEAD");
|
||||
Path seededSkillMd = Path.of(wt, ".claude", "skills", "implementer", "SKILL.md");
|
||||
assertTrue(Files.exists(seededSkillMd), "sanity: the real seeding step copied the skill");
|
||||
|
||||
FakeHerdr herdr = new FakeHerdr();
|
||||
FleetConfig.Fleet fleet = new FleetConfig.Fleet(Map.of(), Map.of(), Map.of(), Map.of(),
|
||||
Map.of("dev", "role rule"), null);
|
||||
Path configRoot = Files.createDirectory(tmp.resolve("configs"));
|
||||
service(herdr, configRoot, opencodeIdeCfg(null,
|
||||
"http://127.0.0.1:29170/index-mcp/streamable-http", wt), () -> fleet).spawn();
|
||||
|
||||
String cfgPath = startEnv(herdr).get("OPENCODE_CONFIG");
|
||||
assertNotNull(cfgPath, "charter + skills + IDE rules still writes a config");
|
||||
JsonNode json = new ObjectMapper().readTree(Path.of(cfgPath).toFile());
|
||||
Path charter = Path.of(cfgPath).resolveSibling("member-charter.md");
|
||||
Path rules = Path.of(cfgPath).resolveSibling("ide-rules.md");
|
||||
assertTrue(Files.exists(charter), "the charter file was written");
|
||||
assertTrue(Files.exists(rules), "the ide-rules file was written");
|
||||
|
||||
List<String> instructions = new ArrayList<>();
|
||||
json.path("instructions").forEach(n -> instructions.add(n.asText()));
|
||||
// Exact ordered list, not size or "contains": a putArray mutation on any writer after the
|
||||
// charter replaces every entry written before it, and the replacement can still be a
|
||||
// plausible-looking array of a different shape. This is the one combination all three
|
||||
// writers are active for — and per fleet01 the realistic shape on a host where weighted
|
||||
// placement makes opencode the default for most members.
|
||||
assertEquals(List.of(charter.toAbsolutePath().toString(),
|
||||
seededSkillMd.toAbsolutePath().toString(),
|
||||
rules.toAbsolutePath().toString()),
|
||||
instructions,
|
||||
"with all three writers active, instructions[] must hold charter, then the seeded "
|
||||
+ "skill, then IDE rules — in that order and with nothing replaced. A "
|
||||
+ "putArray mutation on any writer after the charter would silently drop "
|
||||
+ "earlier entries here while still producing a same-shaped array — got: "
|
||||
+ instructions);
|
||||
}
|
||||
|
||||
// --- fleetd #219: config root + discovery root under memberHerdrSocket ------------------------
|
||||
|
||||
/** A config with {@code memberHerdrSocket:} set, and optionally {@code worktreeRoot:}/{@code worktreeGroup:}. */
|
||||
|
||||
@@ -16,6 +16,7 @@ import java.util.List;
|
||||
import java.util.concurrent.TimeUnit;
|
||||
|
||||
import static org.junit.jupiter.api.Assertions.assertEquals;
|
||||
import static org.junit.jupiter.api.Assertions.assertFalse;
|
||||
import static org.junit.jupiter.api.Assertions.assertThrows;
|
||||
import static org.junit.jupiter.api.Assertions.assertTrue;
|
||||
|
||||
@@ -221,6 +222,31 @@ class AmqpReplyInboxContractTest {
|
||||
}
|
||||
}
|
||||
|
||||
@Test
|
||||
void ackReportsHitVsMissAgainstARealBroker() throws Exception {
|
||||
// fleetd #437: fleet_ack said "acknowledged <msgId>" for a message it never touched,
|
||||
// because ReplyInbox.ack() (void) could not tell a hit from a miss. Pin the fixed
|
||||
// boolean contract against a real broker — the adapter fleetd actually runs live.
|
||||
String target = "worker-ack-contract-" + System.nanoTime();
|
||||
try (AmqpReplyInbox inbox = AmqpReplyInbox.open(uri())) {
|
||||
inbox.own(target);
|
||||
|
||||
// Never held for this target at all: must report false, not throw.
|
||||
assertFalse(inbox.ack(target, "never-held"),
|
||||
"acking a msgId never held for an owned target must report false");
|
||||
|
||||
// A real message: first ack removes it and reports true...
|
||||
inbox.publish(target, "m1", "ack me");
|
||||
assertEquals(1, awaitPeek(inbox, target).size(), "the published reply should be held");
|
||||
assertTrue(inbox.ack(target, "m1"), "acking a held reply must report true");
|
||||
assertTrue(inbox.peek(target).isEmpty(), "an acked reply is dropped");
|
||||
|
||||
// ...and the second ack of the SAME msgId has nothing left to remove: false.
|
||||
assertFalse(inbox.ack(target, "m1"),
|
||||
"acking the same msgId twice must report false the second time");
|
||||
}
|
||||
}
|
||||
|
||||
@Test
|
||||
void confirmedPublishDeliversNormally() throws Exception {
|
||||
String target = "worker-confirm-" + System.nanoTime();
|
||||
|
||||
@@ -28,11 +28,19 @@ public final class FakeLeadChannel implements LeadChannel {
|
||||
private volatile IllegalStateException publishFailure;
|
||||
/** Canned {@link #inspect} results by coord-id — absent for any coord-id not configured here. */
|
||||
private final Map<String, MailboxState> mailboxes = new ConcurrentHashMap<>();
|
||||
/** fleetd #440: matches {@link LeadMailbox}'s real default (durable queue + manual ack) unless overridden. */
|
||||
private volatile boolean heldDurable = true;
|
||||
|
||||
public FakeLeadChannel(String selfCoordId) {
|
||||
this.selfCoordId = selfCoordId;
|
||||
}
|
||||
|
||||
/** Make {@link #heldDurable()} report {@code durable} — the fleetd #440 seam for the false case. */
|
||||
public FakeLeadChannel withHeldDurable(boolean durable) {
|
||||
this.heldDurable = durable;
|
||||
return this;
|
||||
}
|
||||
|
||||
/** Make {@link #inspect(String)} return {@code state} for {@code coordId} instead of "absent". */
|
||||
public FakeLeadChannel withMailbox(String coordId, MailboxState state) {
|
||||
mailboxes.put(coordId, state);
|
||||
@@ -80,6 +88,11 @@ public final class FakeLeadChannel implements LeadChannel {
|
||||
return mailboxes.getOrDefault(coordId, MailboxState.absent(coordId));
|
||||
}
|
||||
|
||||
@Override
|
||||
public boolean heldDurable() {
|
||||
return heldDurable;
|
||||
}
|
||||
|
||||
public List<LeadMessage> published() {
|
||||
return List.copyOf(published);
|
||||
}
|
||||
|
||||
@@ -41,20 +41,20 @@ class InMemoryReplyInboxTest {
|
||||
@Test
|
||||
void ackRemovesTheMessage() {
|
||||
inbox.publish("term_a", "m1", "hello");
|
||||
inbox.ack("term_a", "m1");
|
||||
assertTrue(inbox.ack("term_a", "m1"), "fleetd #437: ack of a real entry must report true");
|
||||
assertTrue(inbox.peek("term_a").isEmpty(), "after ack, the message is gone");
|
||||
}
|
||||
|
||||
@Test
|
||||
void ackForUnknownMsgIdIsNoOp() {
|
||||
inbox.publish("term_a", "m1", "hello");
|
||||
inbox.ack("term_a", "no-such-id"); // no-op
|
||||
assertFalse(inbox.ack("term_a", "no-such-id"), "fleetd #437: a miss must report false"); // no-op
|
||||
assertEquals(1, inbox.peek("term_a").size(), "the published message is still there");
|
||||
}
|
||||
|
||||
@Test
|
||||
void ackForUnknownTargetIsNoOp() {
|
||||
inbox.ack("no-such-target", "m1"); // no-op, should not throw
|
||||
assertFalse(inbox.ack("no-such-target", "m1"), "fleetd #437: a miss must report false"); // no-op, should not throw
|
||||
}
|
||||
|
||||
@Test
|
||||
@@ -167,7 +167,7 @@ class InMemoryReplyInboxTest {
|
||||
@Test
|
||||
void peekAndAckAreNoOpsForUnownedTarget() {
|
||||
assertTrue(inbox.peek("term_not_owned").isEmpty());
|
||||
inbox.ack("term_not_owned", "m1"); // no-op, should not throw
|
||||
assertFalse(inbox.ack("term_not_owned", "m1"), "fleetd #437: a miss must report false"); // no-op, should not throw
|
||||
}
|
||||
|
||||
@Test
|
||||
|
||||
@@ -207,6 +207,21 @@ class LeadMailboxTest {
|
||||
}
|
||||
}
|
||||
|
||||
/**
|
||||
* fleetd #440: {@code heldDurable()} must be derived from what {@link LeadMailbox#own} actually
|
||||
* did against the real broker — a durable queue declare plus a manual-ack consumer — not a
|
||||
* hardcoded literal. This is the mutation-sensitive test: flip {@code own()}'s {@code autoAck}
|
||||
* local to {@code true} (or its {@code durableQueue} local to {@code false}) and this must fail.
|
||||
*/
|
||||
@Test
|
||||
void heldDurableReportsTrueBecauseTheQueueIsDurableAndTheConsumeIsManualAck() throws Exception {
|
||||
String self = coordId("lead-held-durable");
|
||||
try (LeadMailbox mailbox = LeadMailbox.open(uri(), self)) {
|
||||
assertTrue(mailbox.heldDurable(),
|
||||
"own() declares a durable queue and consumes with autoAck=false, so held mail is durable");
|
||||
}
|
||||
}
|
||||
|
||||
@Test
|
||||
void inspectReportsAMissingMailboxAsAbsentRatherThanThrowing() throws Exception {
|
||||
String nobody = coordId("lead-inspect-nobody");
|
||||
|
||||
@@ -1802,7 +1802,11 @@ class MessageServiceTest {
|
||||
Thread asker = new Thread(() -> assertThrows(IllegalStateException.class,
|
||||
() -> service.ask(T, "which config file?", 30_000)));
|
||||
asker.start();
|
||||
awaitTicketPhaseOn(service, ticket, MessageService.Phase.ASKING);
|
||||
// Phase.ASKING (markAsyncQuestion) is only the FIRST of ask()'s three steps; the assertion
|
||||
// below depends on the THIRD (pushLoop.onQuestionOpened). Barrier on the push loop's own
|
||||
// pending-question state instead of the phase — see ReplyPushLoop#pendingQuestionTurnIdsForTest.
|
||||
MessageService.TaskView asking = awaitTicketPhaseOn(service, ticket, MessageService.Phase.ASKING);
|
||||
awaitQuestionPendingOn(pushLoop, LEAD, asking.turnId());
|
||||
assertEquals(ReplyPushLoop.Action.INJECT, pushLoop.decide(LEAD, 0, 0, 0),
|
||||
"the open question should be the one thing keeping this lead's schedule alive");
|
||||
|
||||
@@ -1917,6 +1921,12 @@ class MessageServiceTest {
|
||||
// Only now does it reply. Under the old clock this reply was born already expired.
|
||||
assertTrue(rendezvous.resolve(T, "the long report"));
|
||||
awaitTicketPhaseOn(wiring.service(), slow, MessageService.Phase.DONE);
|
||||
// #399: same barrier fix as the sibling eviction test, for consistency — this test's
|
||||
// own assertion happens to survive a late stamp today (the clock is not advanced any
|
||||
// further between here and the sweep below, so cutoff cannot move past whichever
|
||||
// value completedNanos ends up stamped with), but DONE is still the wrong thing to
|
||||
// order on before a sweep that matters for the TTL.
|
||||
awaitCompletionStamped(wiring.service(), slow);
|
||||
|
||||
// A second delegation runs pruneTerminalTickets before it returns.
|
||||
wiring.service().sendAsync(T, "an unrelated second task");
|
||||
@@ -1946,6 +1956,11 @@ class MessageServiceTest {
|
||||
injectDelivery();
|
||||
assertTrue(rendezvous.resolve(T, "quick result"));
|
||||
awaitTicketPhaseOn(wiring.service(), done, MessageService.Phase.DONE);
|
||||
// #399: DONE can be observed before the completion hook stamps completedNanos. Wait
|
||||
// for the real stamp before advancing the clock, or the hook can run late and stamp
|
||||
// the ADVANCED time — making cutoff = advanced - TTL unreachable and hiding the very
|
||||
// eviction this test exists to pin.
|
||||
awaitCompletionStamped(wiring.service(), done);
|
||||
|
||||
// Nobody collected it, and the TTL has now passed since it FINISHED.
|
||||
clock.addAndGet(MessageService.TICKET_TTL_NANOS + TimeUnit.SECONDS.toNanos(1));
|
||||
@@ -1956,6 +1971,80 @@ class MessageServiceTest {
|
||||
}
|
||||
}
|
||||
|
||||
/**
|
||||
* fleetd #409: pins the #399 ordering invariant deterministically, on the first run, without
|
||||
* relying on host load.
|
||||
*
|
||||
* <p>fleetd #399 was itself only reproducible probabilistically: the real race window between
|
||||
* {@code CompletableFuture.complete()} making {@link MessageService.Phase#DONE} observable and
|
||||
* the constructor's {@code whenComplete} hook actually stamping {@code completedNanos} (see
|
||||
* {@link MessageService#isCompletionStampedForTest}) is normally a handful of instructions wide,
|
||||
* and needed heavy background load on the host to show up in a run at all. This test does not
|
||||
* try to hit that narrow window by chance — it widens it on purpose: {@link #STAMP_DELAY_MILLIS}
|
||||
* is injected into the clock itself, so the completion hook's one read of {@code nowNanos} for
|
||||
* this ticket sleeps before returning, and everything the test does in the meantime (advance the
|
||||
* clock, run a sweep, assert) happens for certain inside that window, on any host.
|
||||
*
|
||||
* <p>Removing the {@link #awaitCompletionStamped} call below reproduces the pre-#399 ordering:
|
||||
* the test then advances the clock and runs its sweep while the hook is still asleep, so the
|
||||
* hook wakes up and stamps {@code completedNanos} with the clock's ALREADY-ADVANCED value
|
||||
* instead of the real completion time. {@code cutoff = advanced - TICKET_TTL_NANOS} can then
|
||||
* never exceed that stamp (the gap between them is fixed at exactly {@code TICKET_TTL_NANOS}),
|
||||
* so the sweep that already ran never evicts the ticket and the very next assertion — expecting
|
||||
* eviction — fails immediately. That is a deterministic, first-run RED failure, not a flaky one
|
||||
* and not a false pass: I ran the test with the barrier call removed and confirmed
|
||||
* {@code assertNull} fails because {@code poll} still returns the ticket's DONE view, which is
|
||||
* exactly this masked-eviction mechanism and not some unrelated defect in the test's own wiring.
|
||||
*/
|
||||
@Test
|
||||
void aTicketOrderedOnDoneInsteadOfTheCompletionStampSurvivesAnEvictionItMustNotSurvive() throws Exception {
|
||||
java.util.concurrent.atomic.AtomicLong clock = new java.util.concurrent.atomic.AtomicLong(1_000_000_000L);
|
||||
// Armed for exactly one call: MessageService's only three nowNanos() call sites are the Task
|
||||
// constructor's createdNanos, this constructor's whenComplete hook's completedNanos, and
|
||||
// pruneTerminalTickets' cutoff — none of which run between arming this (right before
|
||||
// triggering the reply below) and the completion hook firing, so the delayed call is
|
||||
// unambiguously that hook's stamp for `ticket`, never a createdNanos or cutoff read.
|
||||
java.util.concurrent.atomic.AtomicBoolean delayArmed = new java.util.concurrent.atomic.AtomicBoolean(false);
|
||||
java.util.function.LongSupplier gatedClock = () -> {
|
||||
if (delayArmed.compareAndSet(true, false)) {
|
||||
try {
|
||||
Thread.sleep(STAMP_DELAY_MILLIS);
|
||||
} catch (InterruptedException e) {
|
||||
Thread.currentThread().interrupt();
|
||||
}
|
||||
}
|
||||
return clock.get();
|
||||
};
|
||||
MessageService service = new MessageService(agents, injector, rendezvous, inbox, null, null, gatedClock);
|
||||
try {
|
||||
String ticket = service.sendAsync(T, "a quick task");
|
||||
awaitWaiting();
|
||||
injectDelivery();
|
||||
|
||||
delayArmed.set(true);
|
||||
assertTrue(rendezvous.resolve(T, "quick result"));
|
||||
|
||||
awaitTicketPhaseOn(service, ticket, MessageService.Phase.DONE);
|
||||
// The barrier under test (fleetd #409, same fix as fleetd #399): remove this one call to
|
||||
// reproduce the pre-#399 ordering — see the class-level note above for what happens then.
|
||||
awaitCompletionStamped(service, ticket);
|
||||
|
||||
// A real clock would need the whole TTL to pass; the injected one does it instantly, and
|
||||
// by now completedNanos already holds the REAL (small, unadvanced) completion time.
|
||||
clock.addAndGet(MessageService.TICKET_TTL_NANOS + TimeUnit.SECONDS.toNanos(1));
|
||||
service.sendAsync(T, "an unrelated second task"); // runs pruneTerminalTickets before returning
|
||||
|
||||
assertNull(service.poll(ticket),
|
||||
"a ticket whose real completion time is long past the advanced cutoff must be "
|
||||
+ "evicted, whatever the completion hook's clock read was delayed by");
|
||||
} finally {
|
||||
service.close();
|
||||
}
|
||||
}
|
||||
|
||||
/** Bounded delay the injected clock sleeps for in {@link #aTicketOrderedOnDoneInsteadOfTheCompletionStampSurvivesAnEvictionItMustNotSurvive}. */
|
||||
private static final long STAMP_DELAY_MILLIS = 300;
|
||||
|
||||
private MessageService.TaskView awaitTicketPhaseOn(MessageService svc, String ticket,
|
||||
MessageService.Phase phase) throws Exception {
|
||||
long deadline = System.currentTimeMillis() + 3000;
|
||||
@@ -1990,6 +2079,45 @@ class MessageServiceTest {
|
||||
return view;
|
||||
}
|
||||
|
||||
/**
|
||||
* fleetd #399: waits until {@code ticket}'s completion hook has actually stamped
|
||||
* {@code completedNanos}, not just until {@link MessageService#poll} reports
|
||||
* {@link MessageService.Phase#DONE} for it. {@code poll} can observe {@code DONE} the instant
|
||||
* the task's future resolves, before the {@code whenComplete} hook that stamps the completion
|
||||
* time has run — {@code CompletableFuture.complete()} publishes its result and only then runs
|
||||
* dependents. A test that is about to advance an injected clock past the TTL must order itself
|
||||
* after the stamp, not after {@code DONE}: winning the race the other way stamps the
|
||||
* *advanced* clock value and can hide a real eviction bug behind a false pass.
|
||||
*/
|
||||
private void awaitCompletionStamped(MessageService svc, String ticket) throws Exception {
|
||||
long deadline = System.currentTimeMillis() + 3000;
|
||||
while (!svc.isCompletionStampedForTest(ticket)) {
|
||||
assertTrue(System.currentTimeMillis() < deadline,
|
||||
"completedNanos for " + ticket + " was never stamped");
|
||||
Thread.sleep(5);
|
||||
}
|
||||
}
|
||||
|
||||
/**
|
||||
* fleetd #418: waits until {@code turnId} actually appears in {@code pushLoop}'s own open-question
|
||||
* state for {@code lead} — {@link ReplyPushLoop#pendingQuestionTurnIdsForTest} — rather than until
|
||||
* {@link MessageService#poll} reports {@link MessageService.Phase#ASKING}. {@code Phase.ASKING} is
|
||||
* set by {@code markAsyncQuestion}, the FIRST of three steps {@code MessageService.ask()} performs;
|
||||
* {@code pushLoop.onQuestionOpened} (the one that actually publishes to {@code pendingQuestions})
|
||||
* is the THIRD. Under load the asker thread can be descheduled between those two steps, so a
|
||||
* barrier on the phase alone can release before the push loop has anything pending — a test that
|
||||
* then asserts on {@link ReplyPushLoop#decide} is asserting on state that has not been published
|
||||
* yet, not on the throwing path it is named for.
|
||||
*/
|
||||
private void awaitQuestionPendingOn(ReplyPushLoop pushLoop, String lead, String turnId) throws Exception {
|
||||
long deadline = System.currentTimeMillis() + 3000;
|
||||
while (!pushLoop.pendingQuestionTurnIdsForTest(lead).contains(turnId)) {
|
||||
assertTrue(System.currentTimeMillis() < deadline,
|
||||
"turnId " + turnId + " never appeared in the push loop's pending questions for " + lead);
|
||||
Thread.sleep(5);
|
||||
}
|
||||
}
|
||||
|
||||
// --- CB-640: fleet health evidence accessors --------------------------------------------
|
||||
|
||||
@Test
|
||||
|
||||
@@ -90,6 +90,63 @@ class PlacementPolicyTest {
|
||||
assertTrue(e.getMessage().contains("quarantined"), e.getMessage());
|
||||
}
|
||||
|
||||
// --- fleetd #422: FixedPlacementPolicy must consult modelOff too, at BOTH filter sites ------
|
||||
|
||||
/** The default-profile fast path (:44) must skip a default whose model is off. */
|
||||
@Test
|
||||
void fixedSkipsModelOffDefault() {
|
||||
PlacementPolicy policy = PlacementPolicies.fixed();
|
||||
PlacementContext ctx = new PlacementContext("b",
|
||||
List.of(PlacementCandidate.profile("a"), PlacementCandidate.profile("b")),
|
||||
noSessions(), Set.of(), Set.of(), Set.of(), Set.of("b"));
|
||||
assertEquals("a", policy.select(ctx).profile(),
|
||||
"the default 'b' names an off model, so fixed falls through to the first available candidate");
|
||||
}
|
||||
|
||||
/**
|
||||
* The fallback walk (:48-53) must skip an off-model candidate too — exercised independently of
|
||||
* the default-profile fast path by using no default at all, so this is the only filter that runs.
|
||||
*/
|
||||
@Test
|
||||
void fixedFallbackWalkSkipsModelOffCandidate() {
|
||||
PlacementPolicy policy = PlacementPolicies.fixed();
|
||||
PlacementContext ctx = new PlacementContext(null,
|
||||
List.of(PlacementCandidate.profile("a"), PlacementCandidate.profile("b")),
|
||||
noSessions(), Set.of(), Set.of(), Set.of(), Set.of("a"));
|
||||
assertEquals("b", policy.select(ctx).profile(),
|
||||
"candidate 'a' names an off model, so the fallback walk skips it and picks 'b'");
|
||||
}
|
||||
|
||||
@Test
|
||||
void fixedThrowsWhenDefaultAndEveryCandidateModelOff() {
|
||||
PlacementPolicy policy = PlacementPolicies.fixed();
|
||||
PlacementContext ctx = new PlacementContext("b",
|
||||
List.of(PlacementCandidate.profile("a"), PlacementCandidate.profile("b")),
|
||||
noSessions(), Set.of(), Set.of(), Set.of(), Set.of("a", "b"));
|
||||
PlacementException e = assertThrows(PlacementException.class, () -> policy.select(ctx));
|
||||
assertTrue(e.getMessage().contains("turned off"),
|
||||
"message names model-off as the cause: " + e.getMessage());
|
||||
assertFalse(e.getMessage().contains("quarantined"), "must not read like quarantine: " + e.getMessage());
|
||||
assertFalse(e.getMessage().contains("cooling off"), "must not read like cool-off: " + e.getMessage());
|
||||
assertFalse(e.getMessage().contains("weight 0"), "must not read like weight-0: " + e.getMessage());
|
||||
}
|
||||
|
||||
/**
|
||||
* Quarantine still wins when a profile is both quarantined and model-off (mirrors {@code
|
||||
* fixedThrowsWhenDefaultAndEveryCandidateQuarantined}'s priority over cooling off).
|
||||
*/
|
||||
@Test
|
||||
void fixedReportsQuarantineNotModelOffWhenBothApply() {
|
||||
PlacementPolicy policy = PlacementPolicies.fixed();
|
||||
PlacementContext ctx = new PlacementContext("b",
|
||||
List.of(PlacementCandidate.profile("a"), PlacementCandidate.profile("b")),
|
||||
noSessions(), Set.of(), Set.of("a", "b"), Set.of(), Set.of("a", "b"));
|
||||
PlacementException e = assertThrows(PlacementException.class, () -> policy.select(ctx));
|
||||
assertTrue(e.getMessage().contains("quarantined"), e.getMessage());
|
||||
assertFalse(e.getMessage().contains("turned off"),
|
||||
"quarantine takes priority over model-off in the message: " + e.getMessage());
|
||||
}
|
||||
|
||||
@Test
|
||||
void roundRobinCyclesThroughAvailableProfiles() {
|
||||
PlacementPolicy policy = PlacementPolicies.roundRobin();
|
||||
@@ -404,6 +461,93 @@ class PlacementPolicyTest {
|
||||
assertTrue(e.getMessage().contains("weight 0"), e.getMessage());
|
||||
}
|
||||
|
||||
// --- fleetd #435: FixedPlacementPolicy must consult maxLoad too, at BOTH filter sites -------
|
||||
|
||||
/**
|
||||
* The default-profile fast path must skip a capped default. Measured on 7667727 before this
|
||||
* fix: a single dev profile at {@code maxLoad: 1} with 1 live, under {@code fixed()}, still
|
||||
* returned "SPAWNED on profile=a" — the cap was advisory for every unqualified spawn.
|
||||
*/
|
||||
@Test
|
||||
void fixedSkipsCappedDefault() {
|
||||
PlacementPolicy policy = PlacementPolicies.fixed();
|
||||
PlacementContext ctx = new PlacementContext("b",
|
||||
List.of(PlacementCandidate.profile("a", 1.0f, null),
|
||||
PlacementCandidate.profile("b", 1.0f, 1)),
|
||||
name -> "b".equals(name) ? 1 : 0, Set.of(), Set.of(), Set.of());
|
||||
assertEquals("a", policy.select(ctx).profile(),
|
||||
"the default 'b' is at its maxLoad cap, so fixed falls through to the free candidate 'a'");
|
||||
}
|
||||
|
||||
/**
|
||||
* The fallback walk must skip a capped candidate too — exercised independently of the
|
||||
* default-profile fast path by using no default at all, so this is the only filter that runs.
|
||||
*/
|
||||
@Test
|
||||
void fixedFallbackWalkSkipsCappedCandidate() {
|
||||
PlacementPolicy policy = PlacementPolicies.fixed();
|
||||
PlacementContext ctx = new PlacementContext(null,
|
||||
List.of(PlacementCandidate.profile("a", 1.0f, 1),
|
||||
PlacementCandidate.profile("b", 1.0f, null)),
|
||||
name -> "a".equals(name) ? 1 : 0, Set.of(), Set.of(), Set.of());
|
||||
assertEquals("b", policy.select(ctx).profile(),
|
||||
"candidate 'a' is at its maxLoad cap, so the fallback walk skips it and picks 'b'");
|
||||
}
|
||||
|
||||
@Test
|
||||
void fixedThrowsWhenDefaultAndEveryCandidateAtCap() {
|
||||
PlacementPolicy policy = PlacementPolicies.fixed();
|
||||
PlacementContext ctx = new PlacementContext("b",
|
||||
List.of(PlacementCandidate.profile("a", 1.0f, 1),
|
||||
PlacementCandidate.profile("b", 1.0f, 1)),
|
||||
_ -> 1, Set.of(), Set.of(), Set.of());
|
||||
PlacementException e = assertThrows(PlacementException.class, () -> policy.select(ctx));
|
||||
assertTrue(e.getMessage().contains("maxLoad"), "message names the cap: " + e.getMessage());
|
||||
assertTrue(e.getMessage().contains("1 cap"), "message names the cap value: " + e.getMessage());
|
||||
assertFalse(e.getMessage().contains("quarantined"), "must not read like quarantine: " + e.getMessage());
|
||||
assertFalse(e.getMessage().contains("turned off"), "must not read like model-off: " + e.getMessage());
|
||||
}
|
||||
|
||||
/** The mirror: an uncapped default is still chosen, so the new term cannot exclude everything. */
|
||||
@Test
|
||||
void fixedStillReturnsUncappedDefault() {
|
||||
PlacementPolicy policy = PlacementPolicies.fixed();
|
||||
PlacementContext ctx = new PlacementContext("b",
|
||||
List.of(PlacementCandidate.profile("a", 1.0f, null),
|
||||
PlacementCandidate.profile("b", 1.0f, null)),
|
||||
noSessions(), Set.of(), Set.of(), Set.of());
|
||||
assertEquals("b", policy.select(ctx).profile(), "an uncapped default is returned unconditionally");
|
||||
}
|
||||
|
||||
/** CB-585: an explicit {@code maxLoad: 0} on the default caps it at zero live members. */
|
||||
@Test
|
||||
void fixedSkipsMaxLoadZeroDefaultEvenWithZeroLiveWorkers() {
|
||||
PlacementPolicy policy = PlacementPolicies.fixed();
|
||||
PlacementContext ctx = new PlacementContext("b",
|
||||
List.of(PlacementCandidate.profile("a", 1.0f, null),
|
||||
PlacementCandidate.profile("b", 1.0f, 0)),
|
||||
_ -> 0, Set.of(), Set.of(), Set.of());
|
||||
assertEquals("a", policy.select(ctx).profile(),
|
||||
"the default 'b' has maxLoad 0, so it is already at its cap with nobody live");
|
||||
}
|
||||
|
||||
/**
|
||||
* Quarantine still wins when a profile is both quarantined and at cap (mirrors {@code
|
||||
* fixedReportsQuarantineNotModelOffWhenBothApply}'s priority over model-off).
|
||||
*/
|
||||
@Test
|
||||
void fixedReportsQuarantineNotAtCapWhenBothApply() {
|
||||
PlacementPolicy policy = PlacementPolicies.fixed();
|
||||
PlacementContext ctx = new PlacementContext("b",
|
||||
List.of(PlacementCandidate.profile("a", 1.0f, 1),
|
||||
PlacementCandidate.profile("b", 1.0f, 1)),
|
||||
_ -> 1, Set.of(), Set.of("a", "b"), Set.of());
|
||||
PlacementException e = assertThrows(PlacementException.class, () -> policy.select(ctx));
|
||||
assertTrue(e.getMessage().contains("quarantined"), e.getMessage());
|
||||
assertFalse(e.getMessage().contains("maxLoad"),
|
||||
"quarantine takes priority over at-cap in the message: " + e.getMessage());
|
||||
}
|
||||
|
||||
@Test
|
||||
void unknownPolicyNameThrows() {
|
||||
assertThrows(IllegalArgumentException.class, () -> PlacementPolicies.fromName("random"));
|
||||
|
||||
@@ -6,12 +6,14 @@ import ch.qos.logback.classic.spi.ILoggingEvent;
|
||||
import ch.qos.logback.core.read.ListAppender;
|
||||
import dev.ltms.fleet.auth.MemberRegistry;
|
||||
import dev.ltms.fleet.auth.MemberLifecycle;
|
||||
import dev.ltms.fleet.config.ConfigRef;
|
||||
import dev.ltms.fleet.config.FleetConfig;
|
||||
import dev.ltms.fleet.guard.SubscriptionGuard;
|
||||
import dev.ltms.fleet.herdr.AgentControl;
|
||||
import dev.ltms.fleet.herdr.FakeHerdr;
|
||||
import dev.ltms.fleet.herdr.WorkspaceControl;
|
||||
import dev.ltms.fleet.member.ClaudeCodeLauncher;
|
||||
import dev.ltms.fleet.member.CompositePeerLauncher;
|
||||
import dev.ltms.fleet.msg.TestTurnTokens;
|
||||
import dev.ltms.fleet.peer.Capability;
|
||||
import dev.ltms.fleet.peer.CharterReceipt;
|
||||
@@ -20,9 +22,16 @@ import dev.ltms.fleet.peer.PeerHandle;
|
||||
import dev.ltms.fleet.peer.PeerLauncher;
|
||||
import dev.ltms.fleet.peer.PeerUnreachableException;
|
||||
import dev.ltms.fleet.peer.SpawnRequest;
|
||||
import dev.ltms.fleet.placement.BackendQuarantine;
|
||||
import dev.ltms.fleet.placement.PlacementDecision;
|
||||
import dev.ltms.fleet.placement.PlacementPolicies;
|
||||
import org.junit.jupiter.api.Test;
|
||||
import org.junit.jupiter.api.io.TempDir;
|
||||
import org.slf4j.LoggerFactory;
|
||||
|
||||
import java.nio.file.Files;
|
||||
import java.nio.file.Path;
|
||||
import java.util.LinkedHashMap;
|
||||
import java.util.List;
|
||||
import java.util.Map;
|
||||
import java.util.Set;
|
||||
@@ -33,6 +42,7 @@ import java.util.concurrent.Future;
|
||||
import java.util.concurrent.TimeUnit;
|
||||
import java.util.concurrent.atomic.AtomicInteger;
|
||||
import java.util.function.LongSupplier;
|
||||
import java.util.function.Supplier;
|
||||
|
||||
import static org.junit.jupiter.api.Assertions.*;
|
||||
|
||||
@@ -1027,6 +1037,11 @@ class SessionManagerTest {
|
||||
return delegate.spawn(req);
|
||||
}
|
||||
|
||||
@Override
|
||||
public PeerHandle spawn(SpawnRequest req, PlacementDecision decision) {
|
||||
return delegate.spawn(req, decision);
|
||||
}
|
||||
|
||||
@Override
|
||||
public Set<String> profiles() {
|
||||
return delegate.profiles();
|
||||
@@ -1588,6 +1603,11 @@ class SessionManagerTest {
|
||||
throw new UnsupportedOperationException("not reachable — the capability check refuses first");
|
||||
}
|
||||
|
||||
@Override
|
||||
public PeerHandle spawn(SpawnRequest req, PlacementDecision decision) {
|
||||
throw new UnsupportedOperationException("not reachable — the capability check refuses first");
|
||||
}
|
||||
|
||||
@Override
|
||||
public Set<String> profiles() {
|
||||
return Set.of("stub-profile");
|
||||
@@ -1652,6 +1672,11 @@ class SessionManagerTest {
|
||||
return delegate.spawn(req);
|
||||
}
|
||||
|
||||
@Override
|
||||
public PeerHandle spawn(SpawnRequest req, PlacementDecision decision) {
|
||||
return delegate.spawn(req, decision);
|
||||
}
|
||||
|
||||
@Override
|
||||
public Set<String> profiles() {
|
||||
return delegate.profiles();
|
||||
@@ -1858,6 +1883,11 @@ class SessionManagerTest {
|
||||
return handle;
|
||||
}
|
||||
|
||||
@Override
|
||||
public PeerHandle spawn(SpawnRequest req, PlacementDecision decision) {
|
||||
return spawn(req.withProfile(decision.profile()));
|
||||
}
|
||||
|
||||
@Override
|
||||
public Set<String> profiles() {
|
||||
return Set.of("lazy");
|
||||
@@ -1991,4 +2021,296 @@ class SessionManagerTest {
|
||||
sessions.rosterResolved();
|
||||
assertEquals(2, handle.callCount(), "once resolved, the id must not be looked up again");
|
||||
}
|
||||
|
||||
// ── fleetd #425 criterion 3: acquireWithWorktree must provision for the profile it actually
|
||||
// spawns, never a name resolved before a live pool change is accounted for ────────────────────
|
||||
|
||||
/** Two profiles with distinct {@code cwd}/{@code parityOverlay}, and a dev pool of {@code first,second}. */
|
||||
private static String worktreeReorderYaml(String first, String second) {
|
||||
return """
|
||||
bind:
|
||||
host: 127.0.0.1
|
||||
port: 8765
|
||||
herdrSocket: ~/.config/herdr/herdr.sock
|
||||
profiles:
|
||||
a:
|
||||
baseUrl: http://gx00.gw:8000
|
||||
model: coder-a
|
||||
b:
|
||||
baseUrl: http://gx00.gw:8000
|
||||
model: coder-b
|
||||
guard:
|
||||
offSubscriptionHosts:
|
||||
- gx00.gw
|
||||
fleet:
|
||||
developers:
|
||||
slot0:
|
||||
profile: %s
|
||||
slot1:
|
||||
profile: %s
|
||||
""".formatted(first, second);
|
||||
}
|
||||
|
||||
@Test
|
||||
void acquireWithWorktreeProvisionsTheOverlayForTheProfileActuallySpawned(
|
||||
@TempDir Path dir) throws Exception {
|
||||
Path f = dir.resolve("fleetd.yaml");
|
||||
Files.writeString(f, worktreeReorderYaml("a", "b"));
|
||||
ConfigRef ref = new ConfigRef(f, FleetConfig.load(f));
|
||||
|
||||
Map<String, FleetConfig.Profile> profiles = Map.of(
|
||||
"a", new FleetConfig.Profile("a", "http://gx00.gw:8000", "coder-a", null,
|
||||
"FLEETD_WORKER_TOKEN", List.of("claude"), "tab", "fleetd-workers",
|
||||
"worker: {profile} #{n}", null, "/repo/a", List.of("a.mcp.json")),
|
||||
"b", new FleetConfig.Profile("b", "http://gx00.gw:8000", "coder-b", null,
|
||||
"FLEETD_WORKER_TOKEN", List.of("claude"), "tab", "fleetd-workers",
|
||||
"worker: {profile} #{n}", null, "/repo/b", List.of("b.mcp.json")));
|
||||
FakeHerdr herdr = new FakeHerdr();
|
||||
ClaudeCodeLauncher adapter = new ClaudeCodeLauncher(new AgentControl(herdr), new WorkspaceControl(herdr),
|
||||
new SubscriptionGuard(Set.of("gx00.gw")), profiles, "a", _ -> null);
|
||||
PeerLauncher launcher = new CompositePeerLauncher(
|
||||
List.of(adapter), "a", ref, _ -> 0, BackendQuarantine.none());
|
||||
|
||||
// The pool changes AFTER the composite/launcher is built, and BEFORE the unqualified
|
||||
// worktree spawn — exactly the fleetd #425 scenario: the live pool's first entry is "b" by
|
||||
// the time acquireWithWorktree runs, even though nothing here was rebuilt.
|
||||
Files.writeString(f, worktreeReorderYaml("b", "a"));
|
||||
ConfigRef.Outcome out = ref.reload();
|
||||
assertTrue(out.applied(), () -> "reload should apply cleanly: " + out.error());
|
||||
|
||||
FakeWorktrees worktrees = new FakeWorktrees();
|
||||
SessionManager sessions = new SessionManager(launcher, worktrees, () -> 0L);
|
||||
|
||||
MemberSession s = sessions.acquire(null, null, "/caller",
|
||||
null, new WorktreeRequest("fleetd-425", null));
|
||||
|
||||
assertEquals("b", s.profile(),
|
||||
"the live dev pool now starts at b, so the unqualified spawn must land there");
|
||||
FakeWorktrees.OverlayCall overlay = worktrees.lastOverlay();
|
||||
assertNotNull(overlay, "overlayParity must have been called");
|
||||
assertEquals(List.of("b.mcp.json"), overlay.requested(),
|
||||
"the worktree must be provisioned with profile b's overlay — the one actually "
|
||||
+ "spawned — never a's, the pool's stale first entry");
|
||||
}
|
||||
|
||||
/**
|
||||
* The deterministic, mutation-pinning half of criterion 3: {@code launcher.defaultProfile()}
|
||||
* only ever answers for {@link MemberRole#DEV} (see {@link CompositePeerLauncher#defaultProfile()}),
|
||||
* so resolving a worktree spawn's profile through it — instead of through {@link
|
||||
* PeerLauncher#defaultProfileFor(MemberRole)}, resolved against the CALLER's actual role — picks
|
||||
* the wrong pool's answer for any role other than DEV. No reload or race is needed to see it: an
|
||||
* ARCHITECT pool and a DEV pool that simply disagree, held constant, are enough.
|
||||
*/
|
||||
@Test
|
||||
void acquireWithWorktreeForANonDevRoleUsesThatRolesPoolNotTheDevPool() {
|
||||
Map<String, FleetConfig.Profile> profiles = Map.of(
|
||||
"a", new FleetConfig.Profile("a", "http://gx00.gw:8000", "coder-a", null,
|
||||
"FLEETD_WORKER_TOKEN", List.of("claude"), "tab", "fleetd-workers",
|
||||
"worker: {profile} #{n}", null, "/repo/a", List.of("a.mcp.json")),
|
||||
"b", new FleetConfig.Profile("b", "http://gx00.gw:8000", "coder-b", null,
|
||||
"FLEETD_WORKER_TOKEN", List.of("claude"), "tab", "fleetd-workers",
|
||||
"worker: {profile} #{n}", null, "/repo/b", List.of("b.mcp.json")));
|
||||
FakeHerdr herdr = new FakeHerdr();
|
||||
ClaudeCodeLauncher adapter = new ClaudeCodeLauncher(new AgentControl(herdr), new WorkspaceControl(herdr),
|
||||
new SubscriptionGuard(Set.of("gx00.gw")), profiles, "a", _ -> null);
|
||||
// developers -> a (first/only entry); architects -> b (first/only entry). The two pools
|
||||
// disagree on purpose, so a role-blind resolution (DEV's answer, "a") is visibly wrong for
|
||||
// an ARCHITECT spawn, which must land on "b".
|
||||
FleetConfig.Fleet fleet = new FleetConfig.Fleet(Map.of(),
|
||||
Map.of("s0", new FleetConfig.Slot("b")),
|
||||
Map.of("s0", new FleetConfig.Slot("a")),
|
||||
Map.of(), null);
|
||||
PeerLauncher launcher = new CompositePeerLauncher(List.of(adapter), "a", profiles,
|
||||
PlacementPolicies.fixed(), _ -> 0, fleet);
|
||||
|
||||
FakeWorktrees worktrees = new FakeWorktrees();
|
||||
SessionManager sessions = new SessionManager(launcher, worktrees, () -> 0L);
|
||||
|
||||
MemberSession s = sessions.acquire(null, MemberRole.ARCHITECT, null, "/caller",
|
||||
null, new WorktreeRequest("fleetd-425b", null));
|
||||
|
||||
assertEquals("b", s.profile(),
|
||||
"an unqualified ARCHITECT worktree spawn must land on the architect pool's profile");
|
||||
FakeWorktrees.OverlayCall overlay = worktrees.lastOverlay();
|
||||
assertNotNull(overlay, "overlayParity must have been called");
|
||||
assertEquals(List.of("b.mcp.json"), overlay.requested(),
|
||||
"the worktree must be provisioned with profile b's overlay — the ARCHITECT pool's "
|
||||
+ "answer, the one actually spawned — never a's, the DEV pool's answer that "
|
||||
+ "launcher.defaultProfile() alone would have given");
|
||||
}
|
||||
|
||||
/**
|
||||
* fleetd #425 rework, acceptance 3: repoRoot, parityOverlay, AND the actual spawn must all name
|
||||
* the SAME routed profile, proven on the ROUTED path — a quarantine skips the pool's first entry
|
||||
* — not just the "pool reordered by a live reload" path the two tests above already cover.
|
||||
*
|
||||
* <p>This is the exact regression the rework fixes: the first round resolved
|
||||
* {@code acquireWithWorktree}'s profile through {@code launcher.defaultProfileFor(memberRole)},
|
||||
* which is blind to quarantine and just returns the pool's first entry ("a" here, quarantined).
|
||||
* That name went on to provision repoRoot/overlay for "a", and then the spawn itself — now an
|
||||
* EXPLICIT-profile spawn naming "a" — hit {@code CompositePeerLauncher.enforceNotQuarantined}
|
||||
* and threw, where the pre-fix code (a blank-profile spawn) would have routed around "a" onto
|
||||
* "b" without any trouble. {@code launcher.routedProfileFor(memberRole)} closes that gap by
|
||||
* running the SAME quarantine-aware selection {@code spawn} itself uses, so all three — repoRoot,
|
||||
* overlay, and the spawn — land on "b" together.
|
||||
*/
|
||||
@Test
|
||||
void acquireWithWorktreeRoutesAroundAQuarantinedPoolFirstProfile() {
|
||||
Map<String, FleetConfig.Profile> profiles = new LinkedHashMap<>();
|
||||
profiles.put("a", new FleetConfig.Profile("a", "http://gx00.gw:8000", "coder-a", null,
|
||||
"FLEETD_WORKER_TOKEN", List.of("claude"), "tab", "fleetd-workers",
|
||||
"worker: {profile} #{n}", null, "/repo/a", List.of("a.mcp.json"),
|
||||
null, null, null, null, null, null, null, null, "shared-cred", null));
|
||||
profiles.put("b", new FleetConfig.Profile("b", "http://gx00.gw:8000", "coder-b", null,
|
||||
"FLEETD_WORKER_TOKEN", List.of("claude"), "tab", "fleetd-workers",
|
||||
"worker: {profile} #{n}", null, "/repo/b", List.of("b.mcp.json")));
|
||||
FakeHerdr herdr = new FakeHerdr();
|
||||
ClaudeCodeLauncher adapter = new ClaudeCodeLauncher(new AgentControl(herdr), new WorkspaceControl(herdr),
|
||||
new SubscriptionGuard(Set.of("gx00.gw")), profiles, "a", _ -> null);
|
||||
BackendQuarantine quarantine = new BackendQuarantine(() -> 0L, TimeUnit.MINUTES.toNanos(30));
|
||||
quarantine.quarantine("shared-cred");
|
||||
PeerLauncher launcher = new CompositePeerLauncher(List.of(adapter), "a", profiles,
|
||||
PlacementPolicies.fixed(), _ -> 0, null, quarantine);
|
||||
|
||||
FakeWorktrees worktrees = new FakeWorktrees();
|
||||
SessionManager sessions = new SessionManager(launcher, worktrees, () -> 0L);
|
||||
|
||||
MemberSession s = sessions.acquire(null, null, "/caller",
|
||||
null, new WorktreeRequest("fleetd-425c", null));
|
||||
|
||||
assertEquals("b", s.profile(),
|
||||
"a is quarantined, so the unqualified worktree spawn must route to b");
|
||||
FakeWorktrees.OverlayCall overlay = worktrees.lastOverlay();
|
||||
assertNotNull(overlay, "overlayParity must have been called");
|
||||
assertEquals(List.of("b.mcp.json"), overlay.requested(),
|
||||
"parityOverlay must be provisioned for b — the profile actually spawned, never a's, "
|
||||
+ "the quarantined pool-first entry");
|
||||
FakeWorktrees.RepoRootCall repoRootCall = worktrees.repoRootCalls().getLast();
|
||||
assertTrue(repoRootCall.cwd().contains("/repo/b"),
|
||||
"repoRoot must be resolved through b's effectiveCwd, not a's: " + repoRootCall.cwd());
|
||||
}
|
||||
|
||||
/**
|
||||
* fleetd #425 rework, round 2: this is the exact probe that found round 1's maxLoad
|
||||
* regression. One dev profile ("a") is configured with {@code maxLoad: 1} and a liveCount
|
||||
* pinned at 1 — permanently at cap — under {@code PlacementPolicies.fixed()}, the default
|
||||
* policy, which deliberately never evaluates {@code maxLoad} during automatic selection (see
|
||||
* {@code CompositePeerLauncher}'s own javadoc on {@code place}/{@code FixedPlacementPolicy}).
|
||||
*
|
||||
* <p>Round 1 resolved {@code acquireWithWorktree}'s profile through
|
||||
* {@code launcher.routedProfileFor(memberRole)} and then fed that name back into
|
||||
* {@code launcher.spawn(SpawnRequest)} as an EXPLICIT profile. Naming a profile explicitly
|
||||
* takes {@code CompositePeerLauncher.spawn}'s THROWING branch, which calls
|
||||
* {@code enforceMaxLoad} — so the worktree path died with a {@code PlacementException} while
|
||||
* the exact same unqualified request, with no worktree, still spawned cleanly through the
|
||||
* routing branch that never checks {@code maxLoad} at all. One intent, two different answers,
|
||||
* depending only on whether a worktree was asked for — the #425 shape, moved to a different
|
||||
* filter instead of closed.
|
||||
*
|
||||
* <p>This test does not hardcode which of the two outcomes is correct — whether an unqualified
|
||||
* spawn SHOULD respect {@code maxLoad} is fleetd #435, a separate ticket. It only asserts that
|
||||
* the WITH-worktree and WITHOUT-worktree paths agree: both spawn on the same profile, or both
|
||||
* fail with the same exception type and message. That way this test stays correct however
|
||||
* #435 is eventually resolved, and only breaks if the two paths disagree again.
|
||||
*/
|
||||
@Test
|
||||
void unqualifiedAcquireAgreesWithAndWithoutAWorktreeWhenTheOnlyProfileIsAtMaxLoad() {
|
||||
Map<String, FleetConfig.Profile> profiles = new LinkedHashMap<>();
|
||||
profiles.put("a", new FleetConfig.Profile("a", "http://gx00.gw:8000", "coder-a", null,
|
||||
"FLEETD_WORKER_TOKEN", List.of("claude"), "tab", "fleetd-workers",
|
||||
"worker: {profile} #{n}", null, "/repo/a", List.of("a.mcp.json"),
|
||||
null, null, null, null, 1.0f, 1));
|
||||
FakeHerdr herdr = new FakeHerdr();
|
||||
ClaudeCodeLauncher adapter = new ClaudeCodeLauncher(new AgentControl(herdr), new WorkspaceControl(herdr),
|
||||
new SubscriptionGuard(Set.of("gx00.gw")), profiles, "a", _ -> null);
|
||||
// liveCount pinned at 1 for "a", exactly matching maxLoad — "a" is permanently at cap,
|
||||
// regardless of how many times either branch below actually spawns.
|
||||
PeerLauncher launcher = new CompositePeerLauncher(List.of(adapter), "a", profiles,
|
||||
PlacementPolicies.fixed(), name -> "a".equals(name) ? 1 : 0);
|
||||
|
||||
Object without = attemptAcquire(() ->
|
||||
new SessionManager(launcher, new FakeWorktrees(), () -> 0L)
|
||||
.acquire(null, null, "/caller", null));
|
||||
Object with = attemptAcquire(() ->
|
||||
new SessionManager(launcher, new FakeWorktrees(), () -> 0L)
|
||||
.acquire(null, null, "/caller", null, new WorktreeRequest("fleetd-425-maxload", null)));
|
||||
|
||||
assertEquals(without, with, "an unqualified spawn on a profile at maxLoad must agree "
|
||||
+ "whether or not a worktree was requested — no condition may become newly fatal "
|
||||
+ "on the worktree path alone (fleetd #425 rework, round 2)");
|
||||
}
|
||||
|
||||
/**
|
||||
* fleetd #425 rework, round 4: the exact regression a mutation test found that 186 green tests
|
||||
* missed — {@code acquireWithWorktree} dropping the {@link
|
||||
* dev.ltms.fleet.placement.PlacementDecision} it already resolved via {@code launcher.place},
|
||||
* and letting the unqualified spawn re-run placement a second time (a blank-profile {@code
|
||||
* launcher.spawn(spawnReq)}) instead of carrying that decision forward via {@code
|
||||
* launcher.spawn(spawnReq, decision)}. Every earlier test in this file uses {@code
|
||||
* PlacementPolicies.fixed()}, which returns the same answer on every {@code select()} call, so
|
||||
* dropping the decision is invisible under it — two {@code select()} calls simply agree by
|
||||
* accident. {@code PlacementPolicies.roundRobin()} is deterministic AND stateful: its {@code
|
||||
* select()} advances an internal index on every call, so two consecutive calls for the SAME
|
||||
* spawn (one from {@code place()} to provision the worktree, a second from a dropped-decision
|
||||
* blank-profile {@code spawn(spawnReq)}) land on DIFFERENT profiles from a two-profile pool —
|
||||
* index 0 ("a"), then index 1 ("b").
|
||||
*
|
||||
* <p>This test does not hardcode which profile wins — asserting one specific name would pass
|
||||
* for the wrong reason the moment the rotation order changes (round-4 brief invariant 3). It
|
||||
* asserts AGREEMENT instead: whichever profile the worktree's parity overlay was provisioned
|
||||
* for must be the SAME profile the member actually spawned on. Each profile's overlay list is
|
||||
* named after the profile itself ({@code "a.mcp.json"}/{@code "b.mcp.json"}), so comparing the
|
||||
* recorded overlay against {@code s.profile() + ".mcp.json"} checks agreement without ever
|
||||
* naming an expected winner.
|
||||
*/
|
||||
@Test
|
||||
void acquireWithWorktreeSpawnsOnTheSameProfileItProvisionedTheWorktreeForUnderARotatingPolicy() {
|
||||
Map<String, FleetConfig.Profile> profiles = new LinkedHashMap<>();
|
||||
profiles.put("a", new FleetConfig.Profile("a", "http://gx00.gw:8000", "coder-a", null,
|
||||
"FLEETD_WORKER_TOKEN", List.of("claude"), "tab", "fleetd-workers",
|
||||
"worker: {profile} #{n}", null, "/repo/a", List.of("a.mcp.json")));
|
||||
profiles.put("b", new FleetConfig.Profile("b", "http://gx00.gw:8000", "coder-b", null,
|
||||
"FLEETD_WORKER_TOKEN", List.of("claude"), "tab", "fleetd-workers",
|
||||
"worker: {profile} #{n}", null, "/repo/b", List.of("b.mcp.json")));
|
||||
FakeHerdr herdr = new FakeHerdr();
|
||||
ClaudeCodeLauncher adapter = new ClaudeCodeLauncher(new AgentControl(herdr), new WorkspaceControl(herdr),
|
||||
new SubscriptionGuard(Set.of("gx00.gw")), profiles, "a", _ -> null);
|
||||
// roundRobin is deterministic AND stateful: the first select() call picks index 0 ("a"),
|
||||
// and the SAME policy instance's second select() call (reached only if the
|
||||
// PlacementDecision is dropped) picks index 1 ("b") — the two-call disagreement this test
|
||||
// needs to make a dropped decision observable, rather than merely probable.
|
||||
PeerLauncher launcher = new CompositePeerLauncher(List.of(adapter), "a", profiles,
|
||||
PlacementPolicies.roundRobin(), _ -> 0);
|
||||
|
||||
FakeWorktrees worktrees = new FakeWorktrees();
|
||||
SessionManager sessions = new SessionManager(launcher, worktrees, () -> 0L);
|
||||
|
||||
MemberSession s = sessions.acquire(null, null, "/caller",
|
||||
null, new WorktreeRequest("fleetd-425-round4", null));
|
||||
|
||||
FakeWorktrees.OverlayCall overlay = worktrees.lastOverlay();
|
||||
assertNotNull(overlay, "overlayParity must have been called");
|
||||
assertEquals(List.of(s.profile() + ".mcp.json"), overlay.requested(),
|
||||
"the worktree must be provisioned for the SAME profile the member actually spawned "
|
||||
+ "on — under a rotating policy, dropping the PlacementDecision makes the "
|
||||
+ "second, spawn-time select() call disagree with the first, place()-time "
|
||||
+ "call, so the member ends up on a profile whose worktree (repoRoot/parity "
|
||||
+ "overlay) was built for a DIFFERENT profile (fleetd #425 rework, round 4)");
|
||||
}
|
||||
|
||||
/**
|
||||
* Reduce one {@code acquire(...)} attempt to a value comparable across the with-worktree and
|
||||
* without-worktree paths: the spawned profile name on success, or the thrown exception's class
|
||||
* and message on failure. Comparing THIS — instead of asserting "both spawn" or "both throw" as
|
||||
* a hardcoded direction — is what keeps {@link
|
||||
* #unqualifiedAcquireAgreesWithAndWithoutAWorktreeWhenTheOnlyProfileIsAtMaxLoad} valid whichever
|
||||
* way fleetd #435 eventually resolves whether an unqualified spawn should respect maxLoad.
|
||||
*/
|
||||
private static Object attemptAcquire(Supplier<MemberSession> call) {
|
||||
try {
|
||||
return "spawned:" + call.get().profile();
|
||||
} catch (RuntimeException e) {
|
||||
return "threw:" + e.getClass().getName() + ":" + e.getMessage();
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
@@ -0,0 +1,369 @@
|
||||
# Plan — move the fleet back onto fleet01
|
||||
|
||||
**Goal.** Stop the fleet depending on a laptop that sleeps.
|
||||
|
||||
**Written** 2026-09-05. **Rewritten the same day** after the operator pointed out that fleet01 is a VM
|
||||
and already has the repo. They were right, and my first draft was wrong in an important way: this is
|
||||
not a stand-up. **The whole fleet already ran on fleet01 in August.** It was abandoned, not attempted.
|
||||
|
||||
Every fact below was measured on 2026-09-05. Where I did not measure something, the text says so.
|
||||
|
||||
---
|
||||
|
||||
## Status, re-measured 2026-09-10 — phases 1 to 3 are DONE, and §9 is out of date
|
||||
|
||||
This plan is prior art now, not a to-do list. Every number in this section came from a read-only
|
||||
survey over `ssh fleet01` on 2026-09-10. **Delete this section and the plan once fleet01 is the
|
||||
fleet's only daemon** — at that point the plan has been executed and stops being useful.
|
||||
|
||||
| Plan item | State on 2026-09-10 | Command that re-measures it |
|
||||
|---|---|---|
|
||||
| Phase 1, refresh + build | **done, then drifted.** Checkout is on `main` at `4887731`, **77 commits behind** `origin/main`, 0 ahead. `fleetd/target/fleetd.jar` exists, built 2026-09-10 02:10 UTC. So it was rebuilt, and main has moved since. | `git -C ~/LTMS/fleetd rev-list --count HEAD..origin/main` |
|
||||
| Phase 2, port the config | **done.** `fleetd/fleetd.yaml` exists on the host, with `bind`, `profiles`, `configReload`, `fleet`, `health`, `lifecycle`, `guard`, `memberCredentials`, `broker` and `coordinator` all present. | `test -f ~/LTMS/fleetd/fleetd/fleetd.yaml` |
|
||||
| Phase 3, supervision | **done.** `~/.config/systemd/user/fleetd.service` and `herdr.service` both exist, both `active` and `enabled`. `loginctl show-user ltms -p Linger` prints `yes`. Exactly 1 java process, so the double-daemon problem in the old `restart.sh` is not present. | `systemctl --user is-active fleetd herdr` |
|
||||
| Phase 4-5, reachable and a member proven | **partly.** `/healthz` answers 200 on loopback and reports `{"protocol":19,"version":"0.8.0"}`, matching the herdr pin. 1 herdr socket present. I did not spawn a member from here, so "a member on fleet01 opens a PR" is still unproven by me. | `curl -s http://127.0.0.1:8765/healthz` on the host |
|
||||
| Phase 6-7, cutover and reboot proof | **not done.** The Mac still runs its own daemon and is still this fleet's lead. Uptime on fleet01 is 2 weeks 2 days, so no reboot proof has been taken since the units were installed. | `uptime -p` on the host |
|
||||
| §9 "not moving the lead yet" | **out of date.** `claude` is installed at `/home/ltms/.local/bin/claude` and a fleet01 lead is live — it reaches this session over the coordination channel. So the headless-login blocker named in §9 is solved. | `ssh fleet01 'command -v claude'` |
|
||||
|
||||
### The profiles fleet01 actually offers a member
|
||||
|
||||
Measured from the live `fleetd.yaml` on the host, with `placement: weighted`:
|
||||
|
||||
| profile | kind | model | weight | maxLoad |
|
||||
|---|---|---|---|---|
|
||||
| `gx` | opencode | `gx/deepseek-v4-flash` | 100 | 2 |
|
||||
| `xf` | opencode | `opencode/mimo-v2.5-free` | 80 | 5 |
|
||||
| `local` | claude-code | `deepseek-v4-flash` | 10 | 2 |
|
||||
| `opus` | claude-code | `claude-opus-5` | 0 | 1 |
|
||||
|
||||
`weighted` spreads by ratio across every profile with a free slot, so an **unqualified** spawn on
|
||||
fleet01 lands on `gx` or `xf` almost every time. Pass `profile` explicitly there, as the canonical
|
||||
block already says.
|
||||
|
||||
### The PATH split on fleet01, which is not the one I expected
|
||||
|
||||
I went looking for a defect and found the opposite, so this is written down to stop the next
|
||||
session repeating the search.
|
||||
|
||||
`opencode` is installed at `/home/ltms/.opencode/bin/opencode`. Whether a shell can see it depends
|
||||
on which kind of shell it is:
|
||||
|
||||
```
|
||||
zsh -ic 'command -v opencode' -> /home/ltms/.opencode/bin/opencode (1 PATH entry)
|
||||
zsh -lc 'command -v opencode' -> nothing (0 PATH entries)
|
||||
```
|
||||
|
||||
So it is on the **interactive** PATH (`.zshrc`), not the login one. Two consequences, and they
|
||||
point in opposite directions:
|
||||
|
||||
- A herdr pane on Linux is a plain non-login interactive zsh, so a pane **can** launch `opencode`.
|
||||
fleetd types the launch command into the pane rather than exec'ing it, so `gx` and `xf` are not
|
||||
broken by this. I have not spawned one to confirm, so that is inference from the shell
|
||||
measurement plus the typing behaviour, not an end-to-end result.
|
||||
- The daemon itself runs `ExecStart=/bin/zsh -lc "exec java -jar target/fleetd.jar fleetd.yaml"` —
|
||||
a **login** shell, on purpose, because credentials live in `.zprofile`. Its own PATH has 10
|
||||
entries and none contains `opencode`.
|
||||
|
||||
**The general shape: the login shell and the interactive shell see different PATHs, and which one
|
||||
matters depends on whether fleetd types a command or execs it.** Credentials live on the login
|
||||
side; `~/.opencode/bin` lives on the interactive side. Anything fleetd must exec itself is
|
||||
invisible to it if it lives only on the interactive PATH — and that is the mirror of the trap the
|
||||
login-shell `ExecStart` was added to fix.
|
||||
|
||||
---
|
||||
|
||||
---
|
||||
|
||||
## 1. My first draft was wrong — read this before the rest
|
||||
|
||||
I wrote "fleetd is not deployed on fleet01". That was wrong, and I got there by looking in one place
|
||||
and concluding about all of them. Three times:
|
||||
|
||||
| I checked | I concluded | What is actually true |
|
||||
|---|---|---|
|
||||
| `~/LTMS/claude-bridge` | "no repo checkout" | the repo is at `~/LTMS/fleetd` — the directory follows the **renamed** repo, and my own notes record that rename |
|
||||
| `~/.config/herdr/herdr.sock` | "herdr never ran" | the socket lives at `~/.config/herdr/sessions/fleet01/herdr.sock` — a **named session**, and its server log runs to Aug 28 |
|
||||
| both of the above | "fleetd is not deployed" | `~/LTMS/fleetd/fleetd-run/` holds `restart.sh`, `start-herdr.sh` and a `fleetd.out` from **Aug 24** |
|
||||
|
||||
The lesson is the one already written down here: enumerate one channel, conclude about all of them.
|
||||
A single-path check is not a survey.
|
||||
|
||||
---
|
||||
|
||||
## 2. What already worked on fleet01, proven from its own log
|
||||
|
||||
`~/LTMS/fleetd/fleetd-run/fleetd.out` covers 06:07 to 16:48 on 2026-08-24. It shows:
|
||||
|
||||
```
|
||||
4 distinct panes pane=c2504101-... w1:pC w1:pE
|
||||
4 worktrees created and removed /home/ltms/LTMS/.fleet-worktrees/{05f2a7-4,dd9f51-1,f498dd-2,fdb522-3}
|
||||
4 allow-list decisions "memberCredentials allow-list: pane w1:pC allowed 16 of 32 environment variables"
|
||||
AMQP on 127.0.0.1:5672 "AMQP connection recovered; cleared held replies for fresh redelivery"
|
||||
the full turn machinery SessionManager transitions, ReplyPushLoop, CompletionResolver fallback
|
||||
```
|
||||
|
||||
So on fleet01, already: fleetd listened, herdr made panes, members spawned into git worktrees, the
|
||||
credential allow-list fed them their environment, and the broker link was **loopback**.
|
||||
|
||||
That last point is the whole reason for this move. The AMQP resets in section 3 cannot happen to a
|
||||
loopback connection.
|
||||
|
||||
### The two traps I was going to design around are already solved there
|
||||
|
||||
My first draft named these as the biggest risks. Both were already handled in August:
|
||||
|
||||
**Trap A — systemd sources no login shell.** `restart.sh` already starts through one, and says why in
|
||||
its own comment:
|
||||
|
||||
```sh
|
||||
# 1. Start java from a LOGIN shell (zsh -lc). ~/.zprofile is where the credentials live, and a
|
||||
# non-login shell starts the daemon fine with an empty AI_GATEWAY_TOKEN -- a failure that
|
||||
# stays invisible until a member actually needs it.
|
||||
setsid zsh -lc "exec java -jar target/bridged.jar fleetd.yaml" < /dev/null > "$RUN/fleetd.out" 2>&1 &
|
||||
```
|
||||
|
||||
**Trap B — a Linux herdr pane is a plain zsh, so members get no credentials.** Not a problem, and not
|
||||
for the reason I assumed. Members do not inherit from the pane's shell profile — fleetd hands them an
|
||||
allow-list. The log proves it ran: `allowed 16 of 32 environment variables`, four times. The old
|
||||
`fleetd.yaml` has a `memberCredentials:` block that configures it.
|
||||
|
||||
**The headless pty trap is solved too.** `start-herdr.sh` carries the fix and the explanation:
|
||||
|
||||
```sh
|
||||
# Why the size matters: herdr creates each pane sized to the attached client's view. Started
|
||||
# under a pty with no winsize, the client reports 0x0, and every pane.split / workspace.create
|
||||
# then fails with "ghostty error -2" -- libghostty refusing a 0x0 surface.
|
||||
cat > /tmp/herdr-inner.sh <<'INNER'
|
||||
stty rows 50 cols 200 2>/dev/null || true
|
||||
exec herdr --session fleet01
|
||||
INNER
|
||||
setsid script -qfec /tmp/herdr-inner.sh /dev/null < /dev/null > /dev/null 2>&1 &
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
## 3. Why we are moving
|
||||
|
||||
The daemon runs on a Mac laptop. On battery it idle-sleeps after **one minute**:
|
||||
|
||||
```
|
||||
pmset -g custom -> Battery Power: sleep 1
|
||||
AC Power: sleep 0
|
||||
```
|
||||
|
||||
Since the last restart: 16 `Connection reset` events and 6 ERROR lines, all on the AMQP link. I
|
||||
matched every one of the 16 to the nearest sleep or wake event in `pmset -g log`. Largest gap **55
|
||||
seconds**; most under 20. Not one reset lacked a nearby sleep or wake. The broker never sees a network
|
||||
fault — it sees the client stop sending heartbeats, then closes.
|
||||
|
||||
Every reconnect worked, so no message was lost. **The lost messages are not the problem.** The problem
|
||||
is that a member mid-turn freezes with the host, and a long worker turn with nobody typing is exactly
|
||||
the case that goes idle.
|
||||
|
||||
---
|
||||
|
||||
## 4. How the two hosts relate today
|
||||
|
||||
This answers "why is fleet01 related to this Mac at all?"
|
||||
|
||||
```
|
||||
Mac laptop fleet01 (KVM/QEMU guest)
|
||||
┌────────────────────────────┐ ┌──────────────────────┐
|
||||
│ Claude Code lead │ │ LavinMQ :5672 │
|
||||
│ fleetd 127.0.0.1:8765 │ AMQP over │ vhost /mac │
|
||||
│ herdr │──Tailscale────>│ vhost /fleet01 │
|
||||
│ members + worktrees │ utun4, 1280 │ │
|
||||
└────────────────────────────┘ │ (fleetd idle since │
|
||||
│ Aug 24) │
|
||||
└──────────────────────┘
|
||||
```
|
||||
|
||||
**Today fleet01 runs only the broker.** Everything else — daemon, herdr, members — is on the Mac. The
|
||||
single link between them is the Mac's fleetd opening AMQP to `10.10.20.13:5672` across Tailscale.
|
||||
|
||||
So the errors I reported were **the Mac's client dying when the Mac slept**, not fleet01 failing.
|
||||
fleet01 was healthy throughout: the container is up 11 days and its log shows a clean heartbeat
|
||||
timeout each time, which is what a broker sees when a client vanishes.
|
||||
|
||||
After the move that arrow becomes loopback and the whole class of problem is gone.
|
||||
|
||||
---
|
||||
|
||||
## 5. The shape we are restoring
|
||||
|
||||
```mermaid
|
||||
flowchart LR
|
||||
subgraph MAC["Mac laptop (free to sleep)"]
|
||||
LEAD["Claude Code lead"]
|
||||
TUN["ssh -N -L"]
|
||||
end
|
||||
subgraph F01["fleet01 (KVM guest, always on)"]
|
||||
FD["fleetd<br/>127.0.0.1:8765"]
|
||||
HD["herdr --session fleet01"]
|
||||
WT["members<br/>~/LTMS/.fleet-worktrees"]
|
||||
MQ["LavinMQ<br/>127.0.0.1:5672"]
|
||||
end
|
||||
LEAD --> TUN
|
||||
TUN -->|"ssh over Tailscale"| FD
|
||||
FD -->|"unix socket"| HD
|
||||
HD --> WT
|
||||
FD -->|"AMQP, loopback"| MQ
|
||||
WT -->|"AMQP, loopback"| MQ
|
||||
```
|
||||
|
||||
*The lead stays on the Mac. Everything that must survive a sleep is already able to run on fleet01.*
|
||||
|
||||
Two constraints fix this shape:
|
||||
|
||||
1. **fleetd, herdr and the worktrees must share one filesystem.** The herdr link is a Unix socket plus
|
||||
absolute path strings. `herdr --remote` is terminal attach, not a transport.
|
||||
2. **fleetd fails fast on a non-loopback bind without token auth**, by its own design. So we do not
|
||||
expose `:8765` to the `10.10.20.0/24` LAN. An SSH tunnel keeps the bind on loopback and needs no
|
||||
new secret.
|
||||
|
||||
**Why the lead stays on the Mac for now.** A session not in a herdr pane resolves as `primary`, so a
|
||||
Mac-side session over the tunnel works with nothing new. What we give up: async ticket nudges type
|
||||
into the lead's pane, and fleetd cannot type into a pane on another host — so `wait:false` tickets
|
||||
stop nudging and I poll instead. Moving the lead as well is section 9; its hard part is
|
||||
authenticating Claude Code on a headless box, which has nothing to do with sleep and must not block
|
||||
this.
|
||||
|
||||
---
|
||||
|
||||
## 6. The actual gap
|
||||
|
||||
Everything below is what stands between "it ran in August" and "it runs supervised today".
|
||||
|
||||
| # | Gap | Measured state |
|
||||
|---|---|---|
|
||||
| 1 | **Checkout is stale** | branch `cb-634-ide-mcp`, HEAD `7655f1b` (2026-08-24), **290 commits behind** `origin/main`, 0 ahead |
|
||||
| 2 | **Never rebuilt after the rename** | no `target/` anywhere; the scripts still say `bridged.jar` and `cd bridged`, but the tree is now `fleetd/` and `fleetd.jar` |
|
||||
| 3 | **No live `fleetd.yaml`** | gitignored, so not in git. Three backups exist under `bridged/` — the newest is `fleetd.yaml.bak-cb634-pin`, 5512 bytes, and it is **clean of inline secrets** (0 inline passwords, 6 uses of `uriEnv`/`tokenEnv`) |
|
||||
| 4 | **No supervision** | `Linger=no`; **zero** systemd user unit files. Only the hand-rolled `restart.sh` / `start-herdr.sh` |
|
||||
| 5 | **Nothing running now** | no herdr process, no fleetd, no answer on `:8765/healthz` |
|
||||
|
||||
Untracked files in the checkout: `.idea/`, `fleetd-run/`, `docs/CB-634-Worker-IDE-Worktree.md`, and the
|
||||
three yaml backups. All are **untracked, none modified**, and `7655f1b` is already an ancestor of
|
||||
`origin/main` — so nothing is lost by updating the branch. Keep `fleetd-run/` and the backups; they
|
||||
are the prior art this plan is built on.
|
||||
|
||||
### The old config's keys, which tell us what to port
|
||||
|
||||
```
|
||||
bind: herdrSocket: /home/ltms/.config/herdr/sessions/fleet01/herdr.sock
|
||||
profiles: placement: weighted
|
||||
configReload: fleet: health:
|
||||
lifecycle: guard: worktreeRoot: /home/ltms/LTMS/.fleet-worktrees
|
||||
memberCredentials: broker:
|
||||
```
|
||||
|
||||
`herdrSocket` already points at the **named-session** path, and `memberCredentials` is already
|
||||
configured. Those two are what made members work.
|
||||
|
||||
### What the repo already has for this
|
||||
|
||||
`deploy/fleetd.service` exists and is written for Linux. Three lines need fleet01's real paths:
|
||||
`ExecStart` names `/usr/lib/jvm/temurin-25-jdk/bin/java` (fleet01 has `/usr/bin/java`), the `PATH`
|
||||
names `/usr/share/maven/bin` (fleet01 has `/usr/bin/mvn`), and `WorkingDirectory` assumes
|
||||
`%h/src/claude-bridge`. It also declares `After=herdr.service` — **and no `herdr.service` exists in
|
||||
`deploy/`**. Writing that unit, from `start-herdr.sh`, is the one genuinely new piece of code here.
|
||||
|
||||
---
|
||||
|
||||
## 7. Phases
|
||||
|
||||
```mermaid
|
||||
flowchart TD
|
||||
P1["1. Refresh<br/>update + build"]
|
||||
P2["2. Config<br/>port fleetd.yaml"]
|
||||
P3["3. Supervise<br/>linger + 2 units"]
|
||||
P4["4. Reachability<br/>tunnel, primary"]
|
||||
P5["5. Prove a member"]
|
||||
P6["6. Cutover"]
|
||||
P7["7. Reboot proof"]
|
||||
P1 --> P2 --> P3 --> P4 --> P5 --> P6 --> P7
|
||||
```
|
||||
|
||||
**Phase 1 — refresh the checkout.** Update to `origin/main` (290 commits). Keep the untracked
|
||||
`fleetd-run/` and the yaml backups. Then `mvn clean install`, **unpiped** — a pipe hides a failure
|
||||
behind a zero exit. *Check:* `fleetd/target/fleetd.jar` exists and the suite is green.
|
||||
|
||||
**Phase 2 — port the config.** Write `fleetd/fleetd.yaml` from `bridged/fleetd.yaml.bak-cb634-pin`,
|
||||
updating it for the rename and 290 commits of config changes. Diff its keys against
|
||||
`fleetd/fleetd.example.yaml` on current main, key by key, and say what changed. *Check:* the daemon
|
||||
starts and `journalctl ... | grep 'startup secret'` reports **no** `MISSING`.
|
||||
|
||||
**Phase 3 — supervision.** This is the part that never existed. `loginctl enable-linger ltms`; write
|
||||
`deploy/herdr.service` from `start-herdr.sh`, keeping the `stty` sizing; fix the three path lines in
|
||||
`deploy/fleetd.service` and keep the login-shell `ExecStart`; install both under
|
||||
`~/.config/systemd/user/`. Secrets go in a `systemctl --user edit` drop-in or a 0600
|
||||
`EnvironmentFile`, never in the committed unit. *Check:* log out of every ssh session, log back in,
|
||||
and confirm the socket and the daemon are still there. That is what lingering is for, and it is the
|
||||
check people skip.
|
||||
|
||||
**Phase 4 — reachability.** `ssh -N -L <port>:127.0.0.1:8765 fleet01` from the Mac. Use a **different
|
||||
local port** for the first test so the Mac's own daemon on `:8765` is untouched and the whole test is
|
||||
reversible. Point the lead's `.mcp.json` at it — that file is `--skip-worktree` and must never be
|
||||
committed. Wrap the tunnel in `autossh` or a launchd `KeepAlive`, because the Mac still sleeps.
|
||||
*Check:* `fleet_whoami` answers `primary`. If it answers `worker`, the `fleet.leaders.*.tab` pin does
|
||||
not match — a known demotion, not a network fault.
|
||||
|
||||
**Phase 5 — prove a member.** `healthz` can be green while every spawn fails, so only a real spawn
|
||||
proves the herdr link. Spawn one member, then give it a real unit ending in a pushed PR — the push is
|
||||
what proves `WORKER_GITEA_TOKEN` resolved. Confirm the allow-list line still appears
|
||||
(`allowed N of M environment variables`) and that N is what you expect. *Check:* a member on fleet01
|
||||
opens a PR.
|
||||
|
||||
**Phase 6 — cutover.** Drain the Mac fleet properly first: `fleet_list`, `fleet_poll` anything still
|
||||
wanted, then `fleet_stop` each member — a restart drops in-flight tickets and a member's report is
|
||||
gone with its ticket. Then stop the Mac's launchd agent. This is the migration itself, not a change to
|
||||
the Mac's settings, and it is reversible in one command.
|
||||
|
||||
**Phase 7 — reboot proof.** Reboot fleet01. Without touching anything: socket present, healthz
|
||||
answering, `fleet_whoami` still `primary`, one spawn works. Until this passes, "supervised" is a claim.
|
||||
|
||||
---
|
||||
|
||||
## 8. Risks
|
||||
|
||||
| Risk | Why it bites | What this plan does |
|
||||
|---|---|---|
|
||||
| **290 commits of config drift** | the old yaml predates the rename and much else; a silently defaulted key turns a feature off with no error | phase 2 diffs key-by-key against current `fleetd.example.yaml` |
|
||||
| **A new config key gets silently dropped** | `FleetConfig`'s back-compat constructor ladder can absorb an arity change, so a new key compiles and is defaulted away | separate ticket already in flight; matters most here because fleet01 gets a hand-edited yaml |
|
||||
| **Nothing supervises herdr** | `deploy/fleetd.service` depends on a unit that does not exist | phase 3 writes it from the working script; phase 7 proves it |
|
||||
| **Linger left off** | everything dies at logout and looks fine until then | phase 3, checked by logging out |
|
||||
| **Wrong JDK/Maven path in the unit** | fleetd propagates its PATH to every member, so a bad PATH means no member can build | three lines fixed in phase 3, proven by phase 5 |
|
||||
| **Headless pty with no winsize** | `ghostty error -2`, reported three steps later as a spawn failure | the `stty` fix is carried into `herdr.service` |
|
||||
| **Port 8765 collides during the test** | both daemons want the same local port | phase 4 uses a different local port first |
|
||||
| **Tunnel dies when the Mac sleeps** | same sleep, far smaller blast radius — it interrupts my session, not members | `autossh`/launchd `KeepAlive` |
|
||||
| **Lead demoted to worker** | the tab pin no longer matches | phase 4's check is `fleet_whoami` |
|
||||
| **Upgrading herdr** | 0.8.0 is **protocol 19**, pinned on purpose — 0.8.2 is protocol 20 and fleetd has **no version handshake** | do not upgrade herdr during this work; both hosts measured at 0.8.0 today |
|
||||
| **`placement: tab` headless** | fails for the same 0x0 reason as `ghostty error -2` | fleet01 must keep `placement: pane`, as its August config did |
|
||||
| **Profile launch settings are deferred** | editing `placement:` and waiting for the 10s config watch does nothing — the launcher holds a startup snapshot | restart the daemon after those keys, do not wait for the reload |
|
||||
|
||||
---
|
||||
|
||||
## 9. Deliberately not doing
|
||||
|
||||
- **No changes to the Mac's power settings or host config.** The point is to stop depending on it.
|
||||
- **Not touching the leftover `bridged-lavinmq` container** on the Mac. It is unused and harmless.
|
||||
- **Not moving the broker.** It is already on fleet01 and already the durable one. After the move its
|
||||
connection becomes loopback, which is the fix.
|
||||
- **Not building a second fleet.** This is a move. Two daemons on one herdr session kill each other's
|
||||
members.
|
||||
- **Not moving the lead yet.** That needs Claude Code authenticated on a headless Ubuntu box and a
|
||||
`fleet.leaders.*.tab` pin on its pane. It buys back pane nudges. It has nothing to do with sleep, so
|
||||
it must not hold up phases 1–7. Note that fleet01 already carries an `opus` profile defined purely so the lead slot resolves; it cannot spawn until someone runs `claude` and completes `/login` on the host.
|
||||
|
||||
---
|
||||
|
||||
## 10. Open questions for the operator
|
||||
|
||||
1. **vhost** — keep `/mac`, or rename now the fleet is not on the Mac? Renaming loses the existing
|
||||
queues. (The old fleet01 config used its own; phase 2 must settle which this fleet owns.)
|
||||
2. **Fallback week** after cutover, or stop the Mac daemon for good?
|
||||
3. **Delete or keep the stale `cb-634-ide-mcp` branch** on fleet01 once the checkout is updated? Its
|
||||
tip is already in main, so nothing is lost either way.
|
||||
|
||||
The repo path question from the first draft is answered: **`/home/ltms/LTMS/fleetd`**, which already
|
||||
exists. `deploy/fleetd.service` should be pointed there rather than the reverse.
|
||||
Reference in New Issue
Block a user