Compare commits
102 Commits
| Author | SHA1 | Date | |
|---|---|---|---|
| 2c467c2553 | |||
| bfee23acc3 | |||
| 7dec74f1b4 | |||
| 482598e2a6 | |||
| 5f5d16fbd4 | |||
| 7df7985a16 | |||
| 724b35b46e | |||
| 2eb2d6112e | |||
| 7c458e8bf2 | |||
| 133f03e428 | |||
| 804279175d | |||
| 9dea289975 | |||
| cb4a6869b9 | |||
| 28b45d97e5 | |||
| 1a397e962e | |||
| 209e1231ea | |||
| 5d8b9d365c | |||
| a6aeda39e7 | |||
| 31b3c24caa | |||
| d105da978d | |||
| 5051a06443 | |||
| a134eccc57 | |||
| 6f275227d2 | |||
| 283ccf8423 | |||
| b96fba4a03 | |||
| 6d97d210b4 | |||
| 37b23cd704 | |||
| 367facf6a6 | |||
| 7f9a9c09f9 | |||
| e854957247 | |||
| b4b7cf5155 | |||
| cbb35ad947 | |||
| 656588f597 | |||
| e33377b2ca | |||
| 4b4a8688c2 | |||
| 7e48d4b86c | |||
| 41cc785534 | |||
| 60fa86a107 | |||
| f288cee2bb | |||
| a5d6ce1a37 | |||
| a52ca35d34 | |||
| 9417de1123 | |||
| 03d92be751 | |||
| 136bec8e28 | |||
| ae94d511d7 | |||
| 436b026696 | |||
| 905fa3a454 | |||
| 011ee80067 | |||
| 3fab743152 | |||
| 52eb9c2277 | |||
| 6794fd8200 | |||
| 9b5c1cdcff | |||
| 3b69e0103b | |||
| c97b1bba5a | |||
| a42253f597 | |||
| e69eafcc9f | |||
| a42b12440c | |||
| a06426c33c | |||
| 28ea0de575 | |||
| 91792e11fc | |||
| 6539efe9fa | |||
| 68397f78d5 | |||
| af901ff1d2 | |||
| dac5f88812 | |||
| 2e663e5968 | |||
| 8c14ed2846 | |||
| 141ae3b04d | |||
| faefea14c4 | |||
| ea6896f2ef | |||
| d7f94cafa2 | |||
| 096f08c866 | |||
| 4eb720029c | |||
| 0db6d31dc2 | |||
| 158a2a84b5 | |||
| 41ebc9cf69 | |||
| 619769bb52 | |||
| 180c953c42 | |||
| 4af919af0c | |||
| 09c37061c1 | |||
| 2be287ea03 | |||
| c3b0406826 | |||
| e76fa1660b | |||
| 9dca376604 | |||
| 91be5079d6 | |||
| 26f198675c | |||
| 72f46d7c0c | |||
| 640f4d5f23 | |||
| 6edeb70bc4 | |||
| bc49d87cb8 | |||
| 14b169c410 | |||
| 2b52324d9a | |||
| b6147a39f6 | |||
| dbfc34cb6d | |||
| a20cb96730 | |||
| cda1a6a917 | |||
| 608e4496be | |||
| fa61dc587c | |||
| 526e3b7459 | |||
| b6b006c651 | |||
| 63eec8a0da | |||
| 7d9a807243 | |||
| 3f7bc3815e |
@@ -59,7 +59,7 @@ as `matches HEAD`, `drift`, or `unknown`; do not turn an unclear timestamp into
|
||||
Report the process identifier (PID) and uptime too:
|
||||
|
||||
```bash
|
||||
PIDS="$(pgrep -f 'target/fleetd.jar' || true)"
|
||||
PIDS="$(pgrep -f 'fleetd.jar' || true)"
|
||||
if [ -z "$PIDS" ]; then
|
||||
printf '%s\n' 'fleetd: not running'
|
||||
else
|
||||
|
||||
@@ -164,22 +164,69 @@ fails.
|
||||
- **`accepted` does not mean your pane has been cleared.** It means every gate passed and the roll
|
||||
is scheduled to run once your current turn ends. Say your goodbye in the same turn — you will not
|
||||
get another one.
|
||||
- **If you are still running after that turn, the roll did not happen.** A roll that works clears
|
||||
you, so surviving your own goodbye is itself the signal that it refused. Check with
|
||||
`fleet_handover{action: "status", token}`, using the token you confirmed. `TURN_NEVER_SETTLED`
|
||||
means your turn ran past `leadRollover.turnSettleSeconds` and **no `/clear` was ever sent**: your
|
||||
context is intact and nothing was lost. Open a fresh request and retry. Never assume the roll
|
||||
succeeded because `confirm` answered `accepted` — by the time it refuses, there is no caller left
|
||||
to tell, so this check is the only thing that closes that gap.
|
||||
- **There is no terminal or session parameter, on purpose.** The pane is always your own, resolved
|
||||
from your connection, so you can only ever roll yourself.
|
||||
- **`operatorConfirmed` is your report of what a human told you.** Do not pass `true` because you
|
||||
are confident. Ask, wait for the answer, then pass what they said. `requireOperatorConfirm`
|
||||
defaults to `true` and this is the only thing standing between a judgement call and a wiped
|
||||
session.
|
||||
are confident. Ask, wait for the answer, then pass what they said.
|
||||
|
||||
- **Whether you must ask at all depends on `leadRollover.requireOperatorConfirm`. Check it; do not
|
||||
assume.** The default is `true` (`FleetConfig.java:1426`), and then `confirm` refuses unless you
|
||||
also pass `operatorConfirmed: true`. **This host set it to `false` on 2026-09-22**, on the
|
||||
operator's explicit grant, because they do not want to approve routine context rolls. Where it is
|
||||
`false`, the three handover-file checks are the whole gate: the file must exist, be fresher than
|
||||
`maxDocAgeSeconds`, and have been modified after the open request.
|
||||
|
||||
Read the live value rather than trusting this line:
|
||||
|
||||
```bash
|
||||
grep -A1 'requireOperatorConfirm' fleetd/fleetd.yaml
|
||||
```
|
||||
|
||||
No match means the key is unset, so the default `true` applies and you must ask. The key is
|
||||
**deferred, not hot** — it is read once at boot, so an edit does nothing until the daemon is
|
||||
redeployed.
|
||||
|
||||
**Until fleetd #621 merges, the nudge text will tell you to ask the operator even where the
|
||||
daemon no longer requires it.** `LeadHeartbeatLoop.contextNotice()` hardcodes "ask the operator"
|
||||
and takes no config, so it cannot know. Trust the config value over the nudge text. Once #621 is
|
||||
merged and deployed, the nudge matches the config and this warning can be deleted.
|
||||
|
||||
- **The roll can still refuse after `confirm` returns**, and by then there is no caller to tell.
|
||||
Those outcomes are logged only, as `lead-rollover:` lines in the daemon log.
|
||||
- **The bootstrap prompt has never yet landed, and the fix is unproven (fleetd #489).** The first
|
||||
real rollover, on 2026-09-12, joined `/clear` and the bootstrap text into one line and Claude Code
|
||||
refused it as `Unknown command: /clearFresh`. The pane was never cleared and no context was lost,
|
||||
so the failure was safe — the roll simply did nothing. PR #490 fixed the cause and is deployed,
|
||||
but no roll has bootstrapped a fresh session end to end yet. **Assume it may still fail, and tell
|
||||
the operator so before you confirm.** The recovery is the same either way: the file is already
|
||||
written, so the operator starts a session and points it at the file. That is why you write the
|
||||
file before you confirm, and never the other way round.
|
||||
- **The bootstrap prompt works end to end. Measured 2026-09-22.** This used to say the fix was
|
||||
unproven (fleetd #489) and told you to expect a failure. That is no longer true. The daemon log
|
||||
now holds four `lead-rollover: rolled` lines, and three of them ran on 2026-09-22 at 10:01:43,
|
||||
10:38:28 and 11:15:47. Each one cleared the old lead and started a fresh session against the
|
||||
handover file, with the configured `bootstrapText` arriving as its first message. No context was
|
||||
lost. The old `Unknown command: /clearFresh` failure from 2026-09-12 does not appear in the log
|
||||
at all. Re-measure both numbers with:
|
||||
|
||||
```bash
|
||||
grep -c "lead-rollover: rolled" fleetd/fleetd.out # successful rolls
|
||||
grep -c "lead-rollover:" fleetd/fleetd.out # positive control: must be larger
|
||||
grep -c "Unknown command" fleetd/fleetd.out # the old failure: expect 0
|
||||
```
|
||||
|
||||
Run the control line too. A broken pattern returns a clean `0` that reads exactly like good news.
|
||||
If the first number stops growing across rolls, or `Unknown command` returns anything above 0,
|
||||
the bootstrap has regressed and this paragraph is stale again.
|
||||
|
||||
**You still write the file before you confirm, and never the other way round.** That order is not
|
||||
about the bootstrap being unreliable. It is what the daemon checks: the handover file must have
|
||||
been modified *after* the open request, or `confirm` refuses it as stale.
|
||||
|
||||
- **One warning in the log is normal and is not a failure.** Every one of the three rolls above also
|
||||
logged `/clear on term_… was never observed as WORKING after 8 consecutive IDLE/DONE polls —
|
||||
releasing rather than wedging the roll`. The daemon could not see the pane go WORKING after
|
||||
`/clear`, so it released instead of hanging. The roll then succeeded anyway. That is the safe
|
||||
branch behaving correctly. Do not report it as a broken roll.
|
||||
|
||||
## Writing style
|
||||
|
||||
|
||||
@@ -22,13 +22,22 @@ scripts/redeploy-fleetd.sh --yes # skip the drain prompt (fleet already chec
|
||||
scripts/redeploy-fleetd.sh --no-build # restart the jar already on disk
|
||||
```
|
||||
|
||||
`--no-build` skips the build and restarts whatever jar is at `fleetd/target/fleetd.jar`. Use it only
|
||||
when you just built and nothing changed since. It gives up the protection in the next paragraph: no
|
||||
build runs, so a stale or missing jar is not caught early. The script still checks the file is there
|
||||
and dies with `no jar at … — run without --no-build` if it is not, but it cannot tell you the jar is
|
||||
old. A `mvn clean` in the tree deletes that jar while the daemon keeps running on it, and nothing
|
||||
degrades until the next restart. Run `--check` first: it prints the jar's hash and its modification
|
||||
time, so you can see for yourself whether the jar is missing or older than the code you mean to ship.
|
||||
`--no-build` skips the build and restarts whatever jar is at `fleetd/run/fleetd.jar` — the runtime
|
||||
path, not Maven's output path. Use it only when you just built and nothing changed since. It gives
|
||||
up the protection in the next paragraph: no build runs, so a stale or missing jar is not caught
|
||||
early. The script still checks the file is there and dies with `no jar at … — run without
|
||||
--no-build` if it is not, but it cannot tell you the jar is old.
|
||||
|
||||
The daemon runs from `fleetd/run/fleetd.jar`, not from `fleetd/target/fleetd.jar` where Maven
|
||||
writes its output (fleetd #664). That split is what makes a bare `mvn install`/`mvn clean` in the
|
||||
main clone harmless now: neither can reach the file the running daemon holds open, because that
|
||||
file no longer lives under `target/` at all. Verify a merge by building in a throwaway git
|
||||
worktree anyway — a build still produces nothing the fleet runs until this script's own `mv` of
|
||||
`target/fleetd.jar` onto `run/fleetd.jar`, performed only after the old daemon is confirmed gone.
|
||||
Let only `scripts/redeploy-fleetd.sh` touch `fleetd/run/fleetd.jar`. Run `--check` first: it prints
|
||||
the BUILT jar (`target/fleetd.jar`) and the RUNNING jar (`run/fleetd.jar`) as two separately
|
||||
labelled hash-and-mtime facts, so a mismatch between them — a build sitting unswapped, or a stale
|
||||
runtime jar — is visible before you decide anything.
|
||||
|
||||
It builds before it stops anything, so a failed build never leaves the fleet down; it waits for the
|
||||
old process to exit rather than assuming; it polls `/healthz`; and it anchors its log checks to a
|
||||
|
||||
@@ -20,6 +20,13 @@ fleetd.out
|
||||
fleetd/fleetd.out
|
||||
logs/
|
||||
|
||||
# fleetd #635 follow-up — scripts/config-edit.sh's backup directory. No leading slash, so this is
|
||||
# ignored at every depth: the real one lives under fleetd/ (also named in fleetd/.gitignore, next
|
||||
# to the config it backs up), and scripts/test-config-edit.sh's own throwaway fixtures build one
|
||||
# under the repo root while the suite runs. --config can point anywhere, so the directory name is
|
||||
# ignored everywhere rather than only where the live daemon happens to use it.
|
||||
.config-backups/
|
||||
|
||||
# fleetd #480: the lead rollover handover file. `leadRollover.handoverPath` points here, and the
|
||||
# outgoing lead rewrites it on every rollover. It is a snapshot of one moment's live state —
|
||||
# unpushed branches, running builds, open questions — so it is stale the moment it is written and
|
||||
|
||||
@@ -323,6 +323,13 @@ must obey belongs in the charter, not here.
|
||||
reference**, with the intent→tool table above as the short form. `McpContractDocTest` fails if
|
||||
that page names a `fleet_*` tool the server does not register. The flows are kept out of this
|
||||
file because this file loads into every session's context.
|
||||
- **The daemon runs from `fleetd/run/fleetd.jar`, not `fleetd/target/fleetd.jar`** (fleetd #664).
|
||||
Maven's own output still lands at `fleetd/target/fleetd.jar` — that part of the build is
|
||||
unchanged — but the running daemon never has that file open, so a bare `mvn install`/`mvn clean`
|
||||
in the main clone no longer corrupts anything a live process is reading. Verify merges in a
|
||||
throwaway git worktree anyway: a build in the main clone still ships nothing until
|
||||
`scripts/redeploy-fleetd.sh` moves it into place with its own atomic `mv`, performed only after
|
||||
the old daemon is confirmed gone. Let only that script touch `fleetd/run/fleetd.jar`.
|
||||
|
||||
### Redeploying the daemon — the lead may do this (primary only)
|
||||
|
||||
|
||||
@@ -47,7 +47,7 @@
|
||||
<string>/Users/dai.ha/LTMS/claude-bridge/scripts/fleetd-launchd-wrapper.sh</string>
|
||||
<string>/Users/dai.ha/Softwares/jdks/jdk-25.0.3.jdk/Contents/Home/bin/java</string>
|
||||
<string>-jar</string>
|
||||
<string>/Users/dai.ha/LTMS/claude-bridge/fleetd/target/fleetd.jar</string>
|
||||
<string>/Users/dai.ha/LTMS/claude-bridge/fleetd/run/fleetd.jar</string>
|
||||
<string>fleetd.yaml</string>
|
||||
</array>
|
||||
|
||||
|
||||
@@ -50,7 +50,7 @@ WorkingDirectory=%h/LTMS/fleetd/fleetd
|
||||
# and looks healthy, and the failure appears hours later as a member that cannot open a pull
|
||||
# request. exec keeps it one process, so systemd tracks the right PID.
|
||||
# This also avoids a SECOND copy of the secrets in a systemd drop-in: one source of truth.
|
||||
ExecStart=/bin/zsh -lc "exec java -jar target/fleetd.jar fleetd.yaml"
|
||||
ExecStart=/bin/zsh -lc "exec java -jar run/fleetd.jar fleetd.yaml"
|
||||
|
||||
# PrivateTmp MUST stay false -- see herdr.service. fleetd creates the member ZDOTDIR scrub dir and
|
||||
# the opencode config dir under java.io.tmpdir, and the member pane (a herdr child, a different
|
||||
|
||||
@@ -2,11 +2,22 @@
|
||||
target/
|
||||
dependency-reduced-pom.xml
|
||||
|
||||
# The daemon's runtime jar (fleetd #664). scripts/redeploy-fleetd.sh moves the built jar here
|
||||
# with a same-filesystem rename; this is never Maven's output path and never belongs in git.
|
||||
run/
|
||||
|
||||
# Local runtime config (copy from fleetd.example.yaml). Both names are ignored: fleetd.yaml is
|
||||
# the current name, and bridged.yaml is the legacy name Fleetd still falls back to.
|
||||
fleetd.yaml
|
||||
bridged.yaml
|
||||
|
||||
# fleetd #635 follow-up — scripts/config-edit.sh's backups of fleetd.yaml. A backup of a file
|
||||
# that must never be committed inherits that requirement. The directory is the real protection
|
||||
# (it keeps working even if the backup naming changes); the glob is a backstop for a stray
|
||||
# backup written the old way, directly beside fleetd.yaml, or by an older copy of the script.
|
||||
.config-backups/
|
||||
fleetd.yaml.bak.*
|
||||
|
||||
# CB-505 audit trail + daemon stdout/stderr — runtime records, never source
|
||||
logs/
|
||||
|
||||
|
||||
+4
-2
@@ -189,8 +189,10 @@
|
||||
|
||||
<build>
|
||||
<!-- CB-634: the cutover renamed the module dir (bridged/ -> fleetd/), the jar, and the
|
||||
launchd plist together. The installed plist names fleetd/target/fleetd.jar and
|
||||
KeepAlive is armed, so this name, the plist, and the wrapper must move as one. -->
|
||||
launchd plist together. fleetd #664: the installed plist and the systemd unit now name
|
||||
fleetd/run/fleetd.jar, not this plugin's own output path — see
|
||||
scripts/redeploy-fleetd.sh for the mv that gets a build from here to there. KeepAlive is
|
||||
armed, so this name, the plist, and the wrapper must still move as one. -->
|
||||
<finalName>fleetd</finalName>
|
||||
<plugins>
|
||||
<plugin>
|
||||
|
||||
@@ -0,0 +1,22 @@
|
||||
package dev.ltms.fleet;
|
||||
|
||||
import dev.ltms.fleet.config.ConfigRef;
|
||||
import dev.ltms.fleet.config.FleetConfig;
|
||||
import dev.ltms.fleet.guard.SubscriptionGuard;
|
||||
|
||||
/**
|
||||
* fleetd #612 Unit A: everything {@code Fleetd.main} has ready once config is loaded, reported
|
||||
* and validated — the exact point {@code main} used to keep going straight into socket and broker
|
||||
* work. Building this (and handing it to {@link FleetdAssembly#assembleAndStart}) is the seam a
|
||||
* test now has to drive the real boot composition without being {@code main} itself.
|
||||
*
|
||||
* @param cfg the boot-time {@link FleetConfig} snapshot every one-time wiring decision reads —
|
||||
* see {@code Fleetd.main}'s own comment on why this must never be swapped for a live
|
||||
* reference once loaded
|
||||
* @param config the live {@link ConfigRef} the hot-reloadable paths read per use
|
||||
* @param guard the same {@link SubscriptionGuard} {@code Fleetd.main} already used to assert the
|
||||
* launching environment is clean, reused rather than rebuilt so the assembled
|
||||
* launchers see the identical instance {@code main} already validated with
|
||||
*/
|
||||
record AssemblyInputs(FleetConfig cfg, ConfigRef config, SubscriptionGuard guard) {
|
||||
}
|
||||
@@ -42,6 +42,7 @@ import dev.ltms.fleet.mcp.LsofProcessCwdLookup;
|
||||
import dev.ltms.fleet.msg.AmqpReplyInbox;
|
||||
import dev.ltms.fleet.msg.InMemoryReplyInbox;
|
||||
import dev.ltms.fleet.msg.LeadChannel;
|
||||
import dev.ltms.fleet.msg.LeadChannelHandle;
|
||||
import dev.ltms.fleet.msg.LeadCoordLoop;
|
||||
import dev.ltms.fleet.msg.LeadMailbox;
|
||||
import dev.ltms.fleet.msg.MessageService;
|
||||
@@ -130,6 +131,27 @@ public final class Fleetd {
|
||||
}
|
||||
|
||||
static void main(String[] args) {
|
||||
main(args, ResourcePorts.system());
|
||||
}
|
||||
|
||||
/**
|
||||
* fleetd #625: package-private so a test can drive the literal startup sequence — not a copy
|
||||
* of it — against a non-production {@link ResourcePorts} whose {@link
|
||||
* ResourcePorts#environment()} is tainted, without ever touching the real process environment.
|
||||
* The real process environment is exactly what a test cannot taint from inside the JVM, which
|
||||
* is why nothing could pin {@link SubscriptionGuard#assertPrimaryClean} running at this call
|
||||
* site before this ticket.
|
||||
*
|
||||
* <p>{@link #main(String[])} is the one production caller, passing {@link
|
||||
* ResourcePorts#system()}. Every statement below, in the same order, is otherwise unchanged
|
||||
* from before this ticket — in particular the guard still runs here, before {@code
|
||||
* cfg.validateAll()} and before {@link FleetdAssembly#assembleAndStart} ever touches a socket,
|
||||
* a broker, or HTTP, exactly as it always has. {@code ports} is reused for the guard check and
|
||||
* then handed on to the assembly, rather than a second instance being constructed there, so a
|
||||
* test's fake backs the whole boot path with one consistent view — see {@code
|
||||
* FleetdSubscriptionGuardOrderingTest}.
|
||||
*/
|
||||
static void main(String[] args, ResourcePorts ports) {
|
||||
Path configPath = args.length > 0 ? Path.of(args[0]) : chooseDefaultConfigFile(Path.of(""));
|
||||
// The config file was renamed bridged.yaml -> fleetd.yaml. Name the file we actually
|
||||
// loaded, whichever of the two names it carries.
|
||||
@@ -165,7 +187,7 @@ public final class Fleetd {
|
||||
|
||||
// The primary/host env that launched fleetd must not be tainted.
|
||||
SubscriptionGuard guard = new SubscriptionGuard(cfg.guard().hostSet());
|
||||
guard.assertPrimaryClean(System.getenv());
|
||||
guard.assertPrimaryClean(ports.environment());
|
||||
|
||||
// Every FleetConfig.validateXxx() the operator's config can fail — CB-501's auth-exposure
|
||||
// check, CB-531's lead-tab-prefix check, CB-542's subscription-profile check, the charter
|
||||
@@ -177,618 +199,20 @@ public final class Fleetd {
|
||||
// and the Fleetd-startup tests actually pin — see FleetConfig#validateAll's javadoc for
|
||||
// why a name-by-name list here would have the same defect it replaces.
|
||||
cfg.validateAll();
|
||||
// fleetd #613: validateAll() (validateMembers() inside it) only refuses a slot that names a
|
||||
// bad role or profile — it says nothing about a role that has NO pool or NO charter at all,
|
||||
// because both are legitimate ("unconstrained") states, not errors. Report them here, right
|
||||
// after validation passes, so an operator sees the gap once per restart instead of finding
|
||||
// it later in a roster row (see reportRoleFallbackGaps' javadoc for the measured cause).
|
||||
reportRoleFallbackGaps(cfg);
|
||||
// fleetd #469, follow-up to #464: validateAll() (and validateCharters() inside it) only
|
||||
// checks that a charter's KEY is a role wire name and its text is non-blank — it never
|
||||
// looks at what the text actually names. This is the separate check that does: it asks
|
||||
// dev.ltms.fleet.mcp.FleetTool (the canonical registered-tool set) whether every fleet_*/
|
||||
// bridge_* token a charter names is a tool this server actually registers. It cannot live
|
||||
// inside FleetConfig#validateCharters() — config loads before the MCP server exists, and
|
||||
// must not gain a dependency on the mcp package — so it runs here instead, at the one seam
|
||||
// that already holds both a loaded FleetConfig and the mcp package, before anything below
|
||||
// opens a socket or spawns a member. fleetd #474: the same check is also wired into `config`
|
||||
// above as ConfigRef's extraValidation, so a reload refuses what this line refuses at startup.
|
||||
assertChartersNameOnlyRegisteredTools(cfg);
|
||||
|
||||
Path socket = cfg.herdrSocket() != null && !cfg.herdrSocket().isBlank()
|
||||
? Path.of(cfg.herdrSocket())
|
||||
: UnixSocketHerdrClient.defaultSocketPath();
|
||||
|
||||
UnixSocketHerdrClient herdr = UnixSocketHerdrClient.connect(socket, new com.fasterxml.jackson.databind.ObjectMapper());
|
||||
UnixSocketHerdrClient memberHerdr = cfg.memberHerdrSocket() != null && !cfg.memberHerdrSocket().isBlank()
|
||||
? UnixSocketHerdrClient.connect(Path.of(cfg.memberHerdrSocket()), new com.fasterxml.jackson.databind.ObjectMapper())
|
||||
: herdr;
|
||||
AtomicReference<Supplier<Map<String, String>>> leadsRef = new AtomicReference<>(Map::of);
|
||||
HerdrRouter router = new HerdrRouter(herdr, memberHerdr,
|
||||
target -> leadsRef.get().get().containsKey(target));
|
||||
// CB-402: one adapter per configured peer kind, fronted by a composite router. A profile's
|
||||
// `kind:` selects its adapter — claude-code (the default) and opencode partition the profile
|
||||
// set — and the composite dispatches each SPI call to the adapter that owns the profile/pane.
|
||||
Map<String, FleetConfig.Profile> claudeProfiles = new LinkedHashMap<>();
|
||||
Map<String, FleetConfig.Profile> opencodeProfiles = new LinkedHashMap<>();
|
||||
cfg.profiles().forEach((name, w) -> {
|
||||
if (w.isOpenCode()) {
|
||||
opencodeProfiles.put(name, w);
|
||||
} else {
|
||||
claudeProfiles.put(name, w);
|
||||
}
|
||||
});
|
||||
List<HerdrPeerLauncher> adapters = new ArrayList<>();
|
||||
// fleetd #175: the daemon's real ExhaustionSink can only be built once `sessions` exists
|
||||
// (below), but `sessions` needs `workers`, which needs the adapters built right here — a
|
||||
// genuine cycle. Break it exactly like liveCountRef below: a forwarding sink built now,
|
||||
// pointed at the real one once it exists.
|
||||
AtomicReference<ExhaustionSink> exhaustionSinkRef = new AtomicReference<>(ExhaustionSink.none());
|
||||
// fleetd #234, round 4: routed through the shared ExhaustionSink.forwardingTo factory
|
||||
// rather than written inline here — not because a lambda at this call site is unsafe
|
||||
// anymore (it is not: the 3-arg overload is now the interface's single abstract method, so
|
||||
// there is no 2-arg overload left for any lambda to silently bind to instead), but so a
|
||||
// test can call the exact same object this line builds, instead of asserting a copy of its
|
||||
// shape (round 3's lesson).
|
||||
// fleetd #589 Group 1: extracted to forwardingExhaustionSink(...) below (see that method's
|
||||
// javadoc) so a dedicated test can prove this factory keeps reading the reference live,
|
||||
// rather than a rebuilt copy of its shape.
|
||||
ExhaustionSink forwardingExhaustionSink = forwardingExhaustionSink(exhaustionSinkRef);
|
||||
// The claude-code adapter is the always-present default; keep it even with no profiles (so a
|
||||
// bridge configured with no workers, or opencode-only, still has a well-defined base adapter)
|
||||
// unless opencode is the only kind configured.
|
||||
// fleetd #589 Group 2: extracted to claudeCodeLauncher(...)/openCodeLauncher(...) below (see
|
||||
// those methods' javadoc) so a dedicated test can prove the CB-596 memberCredentials policy
|
||||
// supplier is actually wired to each adapter, not silently replaced with `() -> null`.
|
||||
if (!claudeProfiles.isEmpty() || opencodeProfiles.isEmpty()) {
|
||||
adapters.add(claudeCodeLauncher(router.memberAgents(), router.memberSpaces(), guard,
|
||||
claudeProfiles, cfg, config));
|
||||
}
|
||||
if (!opencodeProfiles.isEmpty()) {
|
||||
adapters.add(openCodeLauncher(router.memberAgents(), router.memberSpaces(),
|
||||
opencodeProfiles, cfg, config, forwardingExhaustionSink));
|
||||
}
|
||||
AtomicReference<Function<String, Integer>> liveCountRef = new AtomicReference<>(_ -> 0);
|
||||
// CB-578 stage B: one quarantine tracker for the whole daemon, shared between the launcher
|
||||
// (checked at spawn) and the exhaustion sink wired in below (written on BACKEND_EXHAUSTED).
|
||||
// The cooldown is deferred (see FleetConfig#quarantineCooldownSeconds): it is read once
|
||||
// here, at startup, and a config reload only changes it for a daemon restart.
|
||||
// fleetd #466: escalating, not flat — a credential that keeps reporting exhaustion (e.g. a
|
||||
// weekly subscription limit, which would otherwise be retried on every ~30-minute cooldown,
|
||||
// about 336 times across the week) backs off further each consecutive time, capped at
|
||||
// BackendQuarantine.DEFAULT_MAX_COOLDOWN_MULTIPLE x the base cooldown. See BackendQuarantine's
|
||||
// class doc for the mechanism, why this never fires on cooling-off (a separate, unescalated
|
||||
// mechanism — BackendOutagePolicy below), and the reset.
|
||||
BackendQuarantine quarantine = BackendQuarantine.withEscalation(System::nanoTime,
|
||||
TimeUnit.SECONDS.toNanos(cfg.quarantineCooldownSeconds()));
|
||||
// fleetd #201 Unit 5: one outage-cool-off tracker for the whole daemon, shared between the
|
||||
// launcher (checked at spawn, like `quarantine` above) and the backend-error sink wired in
|
||||
// below (written on a classified backend error). A SEPARATE, shorter-lived mechanism from
|
||||
// `quarantine` — see BackendOutagePolicy's class doc — never merged with it.
|
||||
BackendOutagePolicy outagePolicy = new BackendOutagePolicy(System::nanoTime);
|
||||
PeerLauncher workers = new CompositePeerLauncher(
|
||||
adapters,
|
||||
cfg.effectiveDefaultProfile(),
|
||||
config,
|
||||
profileName -> liveCountRef.get().apply(profileName),
|
||||
quarantine,
|
||||
outagePolicy);
|
||||
// fleetd #422 follow-up: say which of the three model-gate states the daemon booted into —
|
||||
// no models: block at all, a block armed with nothing off, or a block with N off — the same
|
||||
// way exhaustedPatternCoverageLine/errorPatternCoverageLine report CB-578 stage A/fleetd
|
||||
// #201 Unit 5 coverage just below. Read from workers.modelGateState() (never a separate
|
||||
// config.get().models() here) so this line and fleet_profiles' modelGateArmed can never
|
||||
// disagree about what CompositePeerLauncher's spawn gate actually enforces.
|
||||
log.info("model gate (fleetd #422): {}", modelGateCoverageLine(workers.modelGateState()));
|
||||
// CB-504: under supervision (launchd/systemd) fleetd can start before herdr's socket
|
||||
// exists. The client itself is lazy — it connects per call — but the orphan reap below is
|
||||
// the first thing that actually talks to herdr, so without this wait a boot-order race
|
||||
// would crash the daemon into a restart loop. Wait, then degrade rather than die: serving
|
||||
// with /healthz reporting "degraded" is strictly more useful than exiting.
|
||||
HerdrAwaitOutcome herdrOutcome = awaitHerdr(herdr, System::nanoTime, Fleetd::sleepHerdrPoll);
|
||||
boolean herdrUp = logHerdrWaitOutcomeAndShouldReap(herdrOutcome);
|
||||
if (herdrUp) {
|
||||
// CB-117: herdr keeps worker panes alive across a daemon restart, and their ids died
|
||||
// with the previous process — reap those leaked orphans now, before we start serving.
|
||||
workers.reapOrphanWorkers();
|
||||
}
|
||||
|
||||
// CB-301: authoritative session registry + lifecycle FSM on top of ClaudeCodeLauncher.
|
||||
// CB-301-ext: worktree provisioning seam, optionally rooted at a configured directory.
|
||||
// CB-303 part 2: context cap is opt-in and disabled (0) when absent/null.
|
||||
int contextCap = 0;
|
||||
if (cfg.lifecycle() != null && cfg.lifecycle().contextCap() != null
|
||||
&& cfg.lifecycle().contextCap() > 0) {
|
||||
contextCap = cfg.lifecycle().contextCap();
|
||||
}
|
||||
boolean clearAfterTurn = cfg.lifecycle() != null && cfg.lifecycle().clearAfterTurn();
|
||||
SessionManager sessions = new SessionManager(workers,
|
||||
new GitWorktrees(cfg.worktreeRoot(), cfg.worktreeGroup(), cfg.memberSkills()),
|
||||
System::nanoTime, contextCap, clearAfterTurn);
|
||||
liveCountRef.set(profileName -> liveSessionCount(sessions.roster(), profileName));
|
||||
|
||||
// Idle-sleep guard: hold an OS-level assertion against idle sleep while at least one
|
||||
// member is live, so an unattended host does not idle-sleep out from under a member's
|
||||
// long turn (see FleetConfig.IdleSleepGuard / dev.ltms.fleet.power.IdleSleepGuard for the
|
||||
// measurement that motivated this). Opt-out via idleSleepGuard.enabled: false; on by
|
||||
// default. Hangs off SessionManager's own onAcquire/onRelease hooks (CB-520/CB-516,
|
||||
// previously wired only to the reply inbox) and SessionManager#size() — the exact registry
|
||||
// fleet_list's live/capacity numbers are themselves computed from — rather than tracking
|
||||
// members a second way. No-op (never constructed) off macOS or when idleSleepGuard.enabled
|
||||
// is explicitly false; the mechanism itself is additionally a no-op if 'caffeinate' cannot
|
||||
// be started, so this can never fail a spawn, a release, or startup.
|
||||
boolean idleSleepGuardEnabled = cfg.idleSleepGuard() == null || cfg.idleSleepGuard().isEnabled();
|
||||
final IdleSleepGuard idleSleepGuard;
|
||||
if (idleSleepGuardEnabled) {
|
||||
idleSleepGuard = new IdleSleepGuard(new CaffeinateSleepAssertionMechanism(), sessions::size);
|
||||
sessions.onAcquire(_ -> idleSleepGuard.recheck());
|
||||
sessions.onRelease(_ -> idleSleepGuard.recheck());
|
||||
} else {
|
||||
idleSleepGuard = null;
|
||||
}
|
||||
|
||||
// CB-303 part 1: idle-ttl reaper — only when configured, defaults to disabled.
|
||||
final SessionReaper reaper;
|
||||
if (cfg.lifecycle() != null
|
||||
&& cfg.lifecycle().idleTtlSeconds() != null
|
||||
&& cfg.lifecycle().idleTtlSeconds() > 0) {
|
||||
reaper = new SessionReaper(sessions, cfg.lifecycle().idleTtlSeconds());
|
||||
reaper.start();
|
||||
} else {
|
||||
reaper = null;
|
||||
}
|
||||
|
||||
// CB-530: every pane the config names as a lead, merged from `leaders:` and the legacy
|
||||
// singular pin. PrimaryRegistry below still tracks ONE terminal — it addresses the push
|
||||
// loop's nudges, which need a single destination — so it keeps the legacy pin.
|
||||
Map<String, String> leadTerminals = cfg.leaderTerminals();
|
||||
if (leadTerminals.size() > 1) {
|
||||
log.info("leads: {} panes recognised {}", leadTerminals.size(), leadTerminals.values());
|
||||
}
|
||||
// CB-531: on top of the legacy primary.terminal pin, discover leads by the tab labels the
|
||||
// operator writes. CB-557 moved the settings onto the lead they describe, so scanning is on
|
||||
// whenever a `fleet.leaders:` entry exists — with no leads configured the supplier is a
|
||||
// constant and never touches herdr, exactly as a missing `leadScan:` block used to behave.
|
||||
// CB-579: each lead now names its own exact `tab:` label, so one scanner discovers every
|
||||
// configured lead regardless of how differently their tabs are labelled — the old
|
||||
// single-shared-tabPrefix limitation (and its warning) is gone.
|
||||
final Supplier<Map<String, String>> leads;
|
||||
var leaders = cfg.fleet().leaders();
|
||||
if (!leaders.isEmpty()) {
|
||||
Map<String, String> tabToName = new LinkedHashMap<>();
|
||||
leaders.forEach((name, leader) -> {
|
||||
if (leader != null && leader.tab() != null && !leader.tab().isBlank()) {
|
||||
tabToName.put(leader.tab(), name);
|
||||
}
|
||||
});
|
||||
// The lead and the members now share ONE workspace (the operator asked for a single
|
||||
// "session" with many tabs), so no workspace can be excluded — the lead lives in the
|
||||
// members' space by design. A lead is told from a member by its exact tab label alone:
|
||||
// a lead carries its configured `tab`, a member its `worker: {profile} #{n}` template,
|
||||
// and the two never collide. (The scanner still supports an exclusion set for a split
|
||||
// layout; the fleet's policy here is simply not to use one.)
|
||||
//
|
||||
// One shared rescan cadence: still taken from the first entry, as before — it is an
|
||||
// operational cadence, not identity, so there is no correctness reason to give every
|
||||
// lead its own scanner.
|
||||
int scanIntervalSeconds = leaders.values().iterator().next().scanIntervalSeconds();
|
||||
// This must use the lead daemon: scanning member tabs would demote the lead to a worker.
|
||||
leads = new LeadTabScanner(herdr, tabToName, Set.of(),
|
||||
TimeUnit.SECONDS.toNanos(scanIntervalSeconds), System::nanoTime);
|
||||
log.info("lead scan: tabs {} host a lead (rescan every {}s, shared fleet space)",
|
||||
tabToName.keySet(), scanIntervalSeconds);
|
||||
} else {
|
||||
leads = () -> leadTerminals;
|
||||
}
|
||||
leadsRef.set(leads);
|
||||
|
||||
// CB-558: start any declared lead that is not already running. After the scanner is built,
|
||||
// because both read the same tab labels and the ordering makes that dependency visible; and
|
||||
// only when herdr answered, because the launcher's whole safety property is that it can
|
||||
// count live leads first — it must never guess and risk a second orchestrator.
|
||||
if (herdrUp && !leaders.isEmpty()) {
|
||||
int launched = new LeadLauncher(router.leadAgents(), router.leadSpaces(), cfg).ensureLeads();
|
||||
if (launched > 0) {
|
||||
log.info("lead auto-launch: {} lead(s) started", launched);
|
||||
}
|
||||
}
|
||||
|
||||
// CB-548: config-declared architect slots. Config supplies only the stable name → profile
|
||||
// map; the terminal → slot binding is owned by the registry and is empty at startup, so no
|
||||
// pane resolves to an architect until the later spawn lifecycle binds one. The registry is
|
||||
// what CallerResolver resolves against and what that lifecycle will read profiles from;
|
||||
// nothing here spawns a slot.
|
||||
// fleetd #424: MemberRegistry.live re-reads fleet.architects through `config` on every
|
||||
// reserve/requireSlotFor call, so a reload that removes or adds an architect slot governs
|
||||
// the next spawn with no restart — the frozen `new MemberRegistry(cfg.fleet())` this used
|
||||
// to be let a "revoked" slot keep granting new architect spawns forever.
|
||||
MemberRegistry members = MemberRegistry.live(() -> config.get().fleet());
|
||||
sessions.setMemberLifecycle(members);
|
||||
if (!members.slots().isEmpty()) {
|
||||
log.info("member slots: {} configured {} — none bound yet (a slot is idle until the "
|
||||
+ "spawn lifecycle binds a live terminal to it)",
|
||||
members.slots().size(), members.slots().keySet());
|
||||
}
|
||||
|
||||
// Status-gated injector (CB-103): the single writer into workers, fed by a poller.
|
||||
// The blocking message endpoint (CB-104) is the producer; the poller is inert until then.
|
||||
// CB-106: a confirmed turn completion resolves a blocked send whose worker never replied.
|
||||
Rendezvous rendezvous = new Rendezvous();
|
||||
// CB-578 stage A: classify a completion-fallback scrape that matches a profile's configured
|
||||
// usage-limit refusal as BACKEND_EXHAUSTED rather than handing it back as a real answer.
|
||||
// fleetd #446: read live off `config` per lookup, cached by profile name — see
|
||||
// LiveExhaustedPatterns's class doc for why this replaces the old compiled-once-at-startup
|
||||
// map. A profile with no exhaustedPattern simply returns null here, so its workers keep
|
||||
// today's completion-fallback behaviour unchanged.
|
||||
// fleetd #589 Group 1: both extracted to liveExhaustedPatterns(...)/
|
||||
// exhaustedPatternLookup(...) below (see those methods' javadoc) — this is the worst
|
||||
// consequence in the whole #589 sweep: silently losing either wiring means a genuine
|
||||
// usage-limit refusal is handed back as a real completion instead of BACKEND_EXHAUSTED.
|
||||
LiveExhaustedPatterns liveExhaustedPatterns = liveExhaustedPatterns(config);
|
||||
ExhaustedPatternLookup exhaustedPatterns = exhaustedPatternLookup(sessions::roster, liveExhaustedPatterns);
|
||||
// The startup coverage line still reports the boot-time snapshot only — it is printed once,
|
||||
// here, and a reload no longer needs to change what it said; exhaustionDetectionArmed (via
|
||||
// liveExhaustedPatterns.armed, wired into quarantineSource below) is what stays live.
|
||||
Set<String> exhaustedConfiguredAtStartup = cfg.profiles().entrySet().stream()
|
||||
.filter(e -> e.getValue().hasExhaustedPattern())
|
||||
.map(Map.Entry::getKey)
|
||||
.collect(Collectors.toCollection(LinkedHashSet::new));
|
||||
log.info("backend-exhausted classification (CB-578 stage A): {}",
|
||||
exhaustedPatternCoverageLine(cfg.profiles().keySet(), exhaustedConfiguredAtStartup));
|
||||
// fleetd #201 Unit 5: classify a completion-fallback scrape that matches a profile's
|
||||
// configured backend-error refusal (a credential outage, a provider 5xx) as a backend error
|
||||
// rather than handing it back as a real answer. Still compiled once at startup, keyed by
|
||||
// profile name — unlike exhaustedPattern above (fleetd #446), errorPattern was left DEFERRED
|
||||
// on purpose: the ticket that made exhaustedPattern hot scoped errorPattern/cooling-off out
|
||||
// explicitly. A profile with no configured errorPattern is simply absent here, so
|
||||
// CompletionResolver falls back to its built-in narrow {@code (?i)\bAPI Error\s*:}
|
||||
// compatibility pattern for that profile's targets (BackendErrorPatternLookup#legacy's
|
||||
// contract — see backendErrorPatterns below).
|
||||
Map<String, Pattern> errorPatternsByProfile = new LinkedHashMap<>();
|
||||
cfg.profiles().forEach((name, profile) -> {
|
||||
if (profile.hasErrorPattern()) {
|
||||
errorPatternsByProfile.put(name, Pattern.compile(profile.errorPattern()));
|
||||
}
|
||||
});
|
||||
// fleetd #248: extracted to a static factory (see backendErrorPatternLookup below) so a
|
||||
// test can prove main() actually PASSES this into CompletionResolver, not only that the
|
||||
// lookup itself behaves correctly — the exact gap fleetd #248 exists to close.
|
||||
BackendErrorPatternLookup backendErrorPatterns = backendErrorPatternLookup(sessions::roster,
|
||||
errorPatternsByProfile);
|
||||
log.info("backend-error classification (fleetd #201 Unit 5): {}",
|
||||
errorPatternCoverageLine(cfg.profiles().keySet(), errorPatternsByProfile.keySet()));
|
||||
// CB-578 stage B: on a classification that actually wins, quarantine the exhausted profile's
|
||||
// CREDENTIAL — not the profile name — so a profile sharing that credential (e.g. two models
|
||||
// on one OpenAI account) is refused too, not just the one that happened to report it. Reads
|
||||
// the profile config live off `config`, so a credentialId edit is hot: no restart needed.
|
||||
// fleetd #446 criterion 3: the backend text that triggered the most recent quarantine of
|
||||
// each credential, so fleet_profiles can report WHY a limit was hit, not only that it was.
|
||||
// Keyed by credential id — the same key BackendQuarantine's own remainingSeconds uses —
|
||||
// and written at the one call site that actually quarantines (inside exhaustionSink below),
|
||||
// so a reason can never be reported for a quarantine that never happened. Bounded the same
|
||||
// way BackendQuarantine's own internal map is documented to be: by the number of distinct
|
||||
// credentials ever exhausted, not by how often the config is edited.
|
||||
Map<String, String> quarantineReasonByCredential = new ConcurrentHashMap<>();
|
||||
// fleetd #446 follow-up (round 3): extracted into exhaustionSink(...) below — see that
|
||||
// method's javadoc for the full fleetd #175/#234/#446 history this used to carry inline —
|
||||
// so a dedicated test can drive the exact ExhaustionSink main() builds, not a hand-rebuilt
|
||||
// copy of its shape.
|
||||
// fleetd #175: point the forwarding sink handed to OpenCodeLauncher above at the real one,
|
||||
// now that `sessions` exists to resolve target -> session -> profile.
|
||||
// fleetd #589 Group 1: both statements (build + set) folded into publishExhaustionSink(...)
|
||||
// below (see that method's javadoc), so a test can prove the reference is actually
|
||||
// repointed at the real sink, not silently left at ExhaustionSink.none().
|
||||
ExhaustionSink exhaustionSink = publishExhaustionSink(exhaustionSinkRef, sessions, config,
|
||||
quarantine, quarantineReasonByCredential, cfg);
|
||||
// fleetd #201 Unit 5: the production BackendErrorSink needs `pushLoop` (built further below,
|
||||
// after `sessions`) to tell a lead about an incident or an unmapped target — the same
|
||||
// construction-order cycle `exhaustionSinkRef` breaks above, broken the same way: a mutable
|
||||
// holder set once `pushLoop` exists, read lazily from inside the lambda built here.
|
||||
AtomicReference<ReplyPushLoop> pushLoopRef = new AtomicReference<>();
|
||||
// fleetd #248: extracted to a static factory (see backendErrorSink below), public rather
|
||||
// than package-private like the other two factories here, so
|
||||
// dev.ltms.fleet.inject.BackendOutageFlowTest can exercise the REAL production sink
|
||||
// directly instead of a hand-mirrored copy of this lambda — that copy was precisely the
|
||||
// gap fleetd #248 exists to close (see that test's class doc for the history).
|
||||
BackendErrorSink backendErrorSink = backendErrorSink(sessions, () -> config.get().profiles(),
|
||||
outagePolicy, pushLoopRef::get);
|
||||
AgentControl agents = router.memberAgents();
|
||||
// Both fleetd#201 Unit 5 (backend-error patterns + sink) and fleetd#241 (the worktree/branch
|
||||
// lookup the fallback report names) land on this one call. The full constructor takes both,
|
||||
// so neither feature is dropped; nowNanos must be passed explicitly to reach it.
|
||||
//
|
||||
// fleetd #248: every argument built specifically for this call (backendErrorPatterns and
|
||||
// backendErrorSink above, and the worktree/branch lookup right here) now comes from a
|
||||
// static factory tested on its own; FleetdCompletionResolverWiringTest source-asserts that
|
||||
// THIS call actually passes them, which is the coverage that was missing before.
|
||||
CompletionResolver completion = new CompletionResolver(agents, rendezvous, exhaustedPatterns,
|
||||
exhaustionSink, backendErrorPatterns, backendErrorSink, System::nanoTime,
|
||||
worktreeBranchLookup(sessions::roster));
|
||||
// CB-113: deliver only to an available worker (its MCP is connected), never its boot window.
|
||||
// CB-301: the manager's presence bridge records availability and drives SPAWNING → READY.
|
||||
MemberPresence presence = sessions.asPresence();
|
||||
// fleetd #561: extracted to a static factory (see turnListener below) — the anonymous class
|
||||
// this replaced had two bare, unguarded statements per callback, and nothing enforced that
|
||||
// the completion half went first beyond call order in the source.
|
||||
TurnListener turnListener = turnListener(completion, sessions);
|
||||
Predicate<String> deliverable = deliverableTo(presence, leads);
|
||||
// fleetd #556: registration is wired directly to `completion`, not folded into the
|
||||
// `turnListener` fan-out above — so it survives `sessions.onDelivered` (or any future
|
||||
// listener) throwing, regardless of call order. See TurnRegistrar's javadoc.
|
||||
// fleetd #589 Group 3 (:505): extracted to turnRegistrar(...) — see FleetdTurnRegistrarWiringTest.
|
||||
Injector injector = new Injector(router, turnListener, deliverable,
|
||||
presence::forget, turnRegistrar(completion));
|
||||
StatusPoller poller = new StatusPoller(router, injector, Injector.POLL_INTERVAL_MILLIS);
|
||||
poller.start();
|
||||
|
||||
// CB-307: reply inbox. A broker: block selects the AMQP-backed durable adapter; absent (or
|
||||
// unusable), fleetd stays soft-state on the in-memory inbox. The AMQP inbox owns a broker
|
||||
// connection, so keep the reference to close it in the ordered shutdown hook.
|
||||
// fleetd #589 Group 3 (:512): opener extracted to replyInboxOpener() — see
|
||||
// FleetdReplyInboxOpenerWiringTest.
|
||||
final ReplyInbox replyInbox = selectReplyInbox(cfg.broker(), System.getenv(), replyInboxOpener());
|
||||
// CB-637: this daemon's lead-to-lead mailbox on the SHARED coordination vhost — a separate
|
||||
// broker from the reply inbox by design (see FleetConfig.Coordinator). Absent a coordinator:
|
||||
// block this is null and every lead path below is simply not wired, which is exactly the
|
||||
// behaviour before this ticket. It owns a broker connection, so keep the reference for the
|
||||
// ordered shutdown hook.
|
||||
// fleetd #589 Group 3 (:518): opener extracted to leadMailboxOpener() — see
|
||||
// FleetdLeadMailboxOpenerWiringTest.
|
||||
final LeadMailbox leadMailbox = openLeadMailbox(cfg.coordinator(), System.getenv(), leadMailboxOpener());
|
||||
// CB-307: learn the primary's terminal from orchestration tool calls (or pin from config).
|
||||
// The pin also feeds CallerResolver below: a primary running inside a herdr pane would
|
||||
// otherwise resolve as a worker and be refused every orchestration tool.
|
||||
String pinnedPrimaryTerminal = cfg.primary() != null ? cfg.primary().terminal() : null;
|
||||
PrimaryRegistry primaryRegistry = new PrimaryRegistry(pinnedPrimaryTerminal);
|
||||
// CB-532: `primary.terminal` is superseded and no longer needed for either of its jobs —
|
||||
// identity comes from `leaders:`/`leadScan:`, and reply nudges now follow the delegating
|
||||
// lead. Say so once at startup rather than leaving a redundant pin to look load-bearing.
|
||||
if (pinnedPrimaryTerminal != null && !pinnedPrimaryTerminal.isBlank()) {
|
||||
log.warn("primary.terminal is DEPRECATED (CB-532) and can be deleted: identity now comes "
|
||||
+ "from leaders:/leadScan:, and reply nudges follow the lead that delegated. "
|
||||
+ "It still works, and is still the fallback nudge destination when a restart "
|
||||
+ "has lost the delegation map. Its pushReminders/pushBackoffMs stay valid.");
|
||||
}
|
||||
// CB-307: active push-to-primary loop — nudge the primary when replies land without an
|
||||
// open fleet_send. Uses its own lightweight scheduled executor, separate from the injector.
|
||||
int maxReminders = cfg.primary() != null ? cfg.primary().remindersOrDefault() : 5;
|
||||
long backoffMs = cfg.primary() != null ? cfg.primary().backoffMsOrDefault() : 15_000L;
|
||||
var pushScheduler = Executors.newSingleThreadScheduledExecutor(r ->
|
||||
Thread.ofVirtual().name("bridge-push-").unstarted(r));
|
||||
// CB-502: the registry is built before the service and the push loop so send/reply outcomes
|
||||
// are counted at their single funnel rather than at each of the two caller-facing surfaces.
|
||||
// CB-512: the push loop takes it too, so nudge outcomes (delivered|exhausted) are counted.
|
||||
Metrics metrics = FleetMetrics.create(sessions, replyInbox);
|
||||
var pushLoop = new ReplyPushLoop(primaryRegistry, router.leadAgents(), replyInbox,
|
||||
pushScheduler, maxReminders, backoffMs, metrics);
|
||||
// fleetd #201 Unit 5: point the forwarding holder captured by the backendErrorSink lambda
|
||||
// above at the real push loop, now that it exists.
|
||||
pushLoopRef.set(pushLoop);
|
||||
// CB-551: idle-lead heartbeat. Opt-in; absent `leadHeartbeat:` this is never constructed, so
|
||||
// an upgraded daemon cannot silently start spending subscription on nudging an idle lead.
|
||||
// It has its own single-thread scheduler and holds its own scheduler shutdown via close().
|
||||
final LeadHeartbeatLoop heartbeat;
|
||||
var heartbeatScheduler = Executors.newSingleThreadScheduledExecutor(r ->
|
||||
Thread.ofVirtual().name("bridge-heartbeat-").unstarted(r));
|
||||
if (cfg.leadHeartbeat() != null) {
|
||||
var hb = cfg.leadHeartbeat();
|
||||
// fleetd #609: own LeadContextGauge instance for the heartbeat loop — separate from the
|
||||
// one FleetMcp builds internally for fleet_list's context row. Each caches independently
|
||||
// (keyed by configDir+sessionId), so this costs at most one extra bounded tail read per
|
||||
// TTL window, never a shared-mutable-state hazard between the two callers.
|
||||
var leadContextGauge = new LeadContextGauge();
|
||||
// fleetd #621: the context-high notice's own wording must track this same effective
|
||||
// value — LeadRollover.confirm(...) already gates the roll on it (LeadRollover.java:480),
|
||||
// and absent `leadRollover:` entirely the roll is unusable regardless (NOT_CONFIGURED),
|
||||
// so `true` (the FleetConfig.LeadRollover default) is the safe, byte-identical fallback.
|
||||
boolean requireOperatorConfirm = cfg.leadRollover() == null || cfg.leadRollover().requireOperatorConfirm();
|
||||
heartbeat = new LeadHeartbeatLoop(primaryRegistry, router.leadAgents(), replyInbox, sessions::roster,
|
||||
pushLoop, heartbeatScheduler, System::nanoTime,
|
||||
TimeUnit.SECONDS.toNanos(hb.idleAfterSeconds()), hb.backoffMs(), hb.quietNudgeCap(),
|
||||
metrics,
|
||||
leadContextSource(leadContextGauge, router.leadAgents(), leads,
|
||||
leadConfigDirLookup(() -> config.get().profiles(), leaders)),
|
||||
Boolean.TRUE.equals(hb.contextHighNudge()), requireOperatorConfirm);
|
||||
heartbeat.start();
|
||||
} else {
|
||||
heartbeat = null;
|
||||
heartbeatScheduler.shutdownNow();
|
||||
}
|
||||
// fleetd #480: lead rollover. Opt-in; absent `leadRollover:` this is never constructed, so
|
||||
// an upgraded daemon cannot silently acquire the ability to clear the lead's own pane.
|
||||
// Unlike heartbeat above, this has no recurring scheduler of its own — nothing but an
|
||||
// explicit confirm() call (wired to an MCP tool by a later ticket; nothing calls it yet)
|
||||
// that passes every gate can ever schedule a roll. It does not take primaryRegistry: the
|
||||
// lead terminal to roll comes from the caller of open()/confirm() (resolved by the MCP
|
||||
// layer from the connection, the same way auth/CallerResolver#resolve builds a
|
||||
// Principal.leader(...)), never from a single-slot lookup — see LeadRollover's class
|
||||
// javadoc, fleetd #480 correction 2.
|
||||
LeadRollover leadRollover = leadRollover(cfg, router.leadAgents(), config, leads);
|
||||
MessageService messages = new MessageService(router, injector, rendezvous, replyInbox,
|
||||
pushLoop, metrics);
|
||||
|
||||
// Health is a slow whole-fleet observer. Keep it separate from the 250ms delivery poller.
|
||||
final FleetHealthMonitor healthMonitor;
|
||||
var healthScheduler = Executors.newSingleThreadScheduledExecutor(r ->
|
||||
Thread.ofVirtual().name("bridge-health-").unstarted(r));
|
||||
if (cfg.health() != null && cfg.health().isEnabled()) {
|
||||
// CB-580: a member found GONE/NEVER_READY must fail whatever ticket is waiting on it,
|
||||
// through the same idempotent target-wide operation CB-516 already uses on release.
|
||||
// fleetd #386: System::nanoTime freezes across a macOS sleep, so the stall check also
|
||||
// gets a wall-clock source to detect and correct for that freeze. Every other decision
|
||||
// in FleetHealthMonitor stays on the monotonic clock, unchanged.
|
||||
// fleetd #589 Group 3 (:591): failTarget extracted to healthFailTarget(...) — see
|
||||
// FleetdHealthFailTargetWiringTest.
|
||||
healthMonitor = new FleetHealthMonitor(agents, sessions::roster, messages, healthScheduler,
|
||||
System::nanoTime, () -> TimeUnit.MILLISECONDS.toNanos(System.currentTimeMillis()),
|
||||
cfg.health().intervalOrDefault(),
|
||||
cfg.health().workingSuspectAfterOrDefault(), healthFailTarget(messages));
|
||||
String coverage = FleetHealthMonitor.coverage(true,
|
||||
cfg.health().notifications() != null && cfg.health().notifications().configured());
|
||||
if ("detection-only".equals(coverage)) {
|
||||
log.warn("fleet health: {} (no notification sink configured)", coverage);
|
||||
} else {
|
||||
log.info("fleet health: {}", coverage);
|
||||
}
|
||||
healthMonitor.start();
|
||||
} else {
|
||||
healthMonitor = null;
|
||||
healthScheduler.shutdownNow();
|
||||
}
|
||||
|
||||
// CB-520: the reply inbox only consumes for agents this gateway owns. own on acquire,
|
||||
// release on teardown. Do this before CB-516 so the inbox is owned before any reply can land.
|
||||
sessions.onAcquire(replyInbox::own);
|
||||
// CB-516: releasing a worker must fail whatever send was waiting on it. Without this a
|
||||
// torn-down delegation kept reporting PENDING until the 30-minute async timeout, and never
|
||||
// reached /metrics — the delegation was unresolvable and nothing said so.
|
||||
// fleetd #589 Group 3 (:611-631): the whole cleanup lambda extracted to releaseCleanup(...)
|
||||
// — see FleetdReleaseCleanupWiringTest, and MessageService.abandon's javadoc for the
|
||||
// documented incident (a torn-down worker's rendezvous waiter left open) this lambda exists
|
||||
// to prevent.
|
||||
sessions.onRelease(releaseCleanup(messages, replyInbox, primaryRegistry));
|
||||
|
||||
// MCP server face (CB-105): fleet_send/fleet_reply/fleet_status, mounted at /mcp.
|
||||
// Caller identity is resolved from the connection (peer PID → herdr pane), not arguments.
|
||||
// CB-185: a caller's pane can live on either daemon (a lead's on the lead daemon, a
|
||||
// member's on the member daemon) — search both, lead first. Collapses to one scan when
|
||||
// memberHerdrSocket is unset (herdr == memberHerdr).
|
||||
ConnectionIdentity identity = new ConnectionIdentity(
|
||||
new PaneLocator(herdr, memberHerdr), new LsofPeerPidLookup(), new LsofProcessCwdLookup());
|
||||
|
||||
// CB-501: one resolver behind both entry paths. Worker identity still comes from the
|
||||
// connection and is never token-gated, so enabling token mode cannot lock the fleet out.
|
||||
final CallerResolver callers;
|
||||
if (cfg.auth().tokenMode()) {
|
||||
String token = System.getenv(cfg.auth().tokenEnv());
|
||||
if (token == null || token.isBlank()) {
|
||||
throw new IllegalStateException("auth.mode=token but env var " + cfg.auth().tokenEnv()
|
||||
+ " is unset or empty — export it before starting fleetd");
|
||||
}
|
||||
callers = CallerResolver.withLeadsAndMembers(identity, true, token, leads, members);
|
||||
log.info("auth: token mode (bearer required for non-worker callers, env {})",
|
||||
cfg.auth().tokenEnv());
|
||||
} else {
|
||||
callers = CallerResolver.withLeadsAndMembers(identity, false, null, leads, members);
|
||||
log.info("auth: loopback-trust (any loopback non-worker caller is the primary)");
|
||||
}
|
||||
|
||||
// fleetd #297: named once and reused verbatim below for FleetApp's GET /profiles, rather than
|
||||
// built a second time — two independently-constructed sources reading the SAME BackendQuarantine
|
||||
// / BackendOutagePolicy would still be able to drift (e.g. a future edit to the credentialIdFor
|
||||
// closure in only one of the two places), exactly the shape #284 was.
|
||||
FleetMcp.QuarantineSource quarantineSource = quarantineSource(config, quarantine,
|
||||
liveExhaustedPatterns, quarantineReasonByCredential);
|
||||
FleetMcp.OutageSource outageSource = new FleetMcp.OutageSource(profile -> {
|
||||
var configured = config.get().profiles().get(profile);
|
||||
return configured == null ? null : configured.effectiveCredentialId();
|
||||
}, outagePolicy);
|
||||
|
||||
FleetMcp.LoopHealthSource loopHealth = loopHealthSource(poller, reaper);
|
||||
FleetMcp mcp = new FleetMcp(messages, workers, sessions, identity, presence,
|
||||
primaryRegistry, callers, FleetMcp.AuthorizationMode.ENFORCED, metrics,
|
||||
capacitySource(config, cfg, profile -> liveCountRef.get().apply(profile)),
|
||||
healthCoverageSource(config),
|
||||
loopHealth,
|
||||
quarantineSource,
|
||||
leadMailbox,
|
||||
outageSource,
|
||||
new FleetMcp.LeadSeatSource(leadSeatLookup(() -> config.get().profiles(), leaders, leads)),
|
||||
// fleetd #602 gauge-wiring: threads each lead's configured configDir into the
|
||||
// context gauge — see leadConfigDirSource's own doc for why this, not a hardcoded
|
||||
// null, is what fleet_list's context row now reads.
|
||||
leadConfigDirSource(() -> config.get().profiles(), leaders),
|
||||
// fleetd #361: the operator-declared peers this daemon's fleet_list should try to
|
||||
// reach. Read from the SAME snapshot leadMailbox itself opened from (cfg.coordinator()),
|
||||
// not the live config.get() — coordinator wiring is already boot-time-fixed (see
|
||||
// leadMailbox above), so peers follows the same rule rather than half hot-reloading.
|
||||
cfg.coordinator() == null ? List.of() : cfg.coordinator().peers(),
|
||||
// fleetd #480 Unit C: the executor behind fleet_handover — null whenever
|
||||
// leadRollover: is not configured (see the leadRollover local above).
|
||||
leadRollover);
|
||||
|
||||
// CB-637: the receive half. Only constructed when a lead mailbox actually opened — with no
|
||||
// coordinator (or an unreachable one) there is nothing to deliver, so no scheduler is
|
||||
// created and no thread runs. It reads the SAME live lead supplier the injector's
|
||||
// deliverability gate does, so a lead found by the tab scan after startup is reachable
|
||||
// without a restart.
|
||||
final LeadCoordLoop leadCoordLoop;
|
||||
final ScheduledExecutorService leadCoordSchedulerRef;
|
||||
if (leadMailbox != null) {
|
||||
var leadCoordScheduler = Executors.newSingleThreadScheduledExecutor(r ->
|
||||
Thread.ofVirtual().name("bridge-leadcoord-").unstarted(r));
|
||||
leadCoordLoop = new LeadCoordLoop(leadMailbox, router.leadAgents(), leads, leadCoordScheduler,
|
||||
LEAD_COORD_INTERVAL_MS);
|
||||
leadCoordLoop.start();
|
||||
leadCoordSchedulerRef = leadCoordScheduler;
|
||||
} else {
|
||||
leadCoordLoop = null;
|
||||
leadCoordSchedulerRef = null;
|
||||
}
|
||||
|
||||
// CB-559: opt-in config reload. With no `configReload:` block nothing is constructed, so an
|
||||
// upgraded daemon behaves exactly as before — the file is read once at boot and never again.
|
||||
final ConfigWatcher configWatcher;
|
||||
if (cfg.configReload() != null && cfg.configReload().isEnabled()) {
|
||||
configWatcher = new ConfigWatcher(config, cfg.configReload().intervalSeconds());
|
||||
configWatcher.start();
|
||||
} else {
|
||||
configWatcher = null;
|
||||
}
|
||||
|
||||
// CB-303 part 3: single ordered shutdown hook. Drain sessions first while herdr is still
|
||||
// open (so releases reach the daemon), then stop poller/message/mcp/reaper, and close herdr
|
||||
// last. This replaces the earlier independent hooks that could race and close herdr early.
|
||||
Runtime.getRuntime().addShutdownHook(new Thread(() -> {
|
||||
sessions.close(cfg.lifecycle() != null ? cfg.lifecycle().drainTimeoutSeconds() : null);
|
||||
poller.stop();
|
||||
messages.close();
|
||||
pushLoop.close();
|
||||
if (heartbeat != null) heartbeat.close(); // CB-551: stop the idle-lead heartbeat scheduler
|
||||
if (leadCoordLoop != null) leadCoordLoop.close(); // CB-637: stop delivering peer-lead messages
|
||||
if (leadCoordSchedulerRef != null) leadCoordSchedulerRef.shutdownNow();
|
||||
if (healthMonitor != null) healthMonitor.stop();
|
||||
if (configWatcher != null) configWatcher.stop(); // CB-559: stop polling the config file
|
||||
mcp.close();
|
||||
if (reaper != null) reaper.stop();
|
||||
// Idle-sleep guard: release unconditionally, even though sessions.close() above already
|
||||
// drained every session (and each release already drove the live count to 0, which
|
||||
// releases the guard's assertion on its own) — this is the backstop for a drain that was
|
||||
// itself interrupted or threw, so no caffeinate child ever outlives the daemon.
|
||||
if (idleSleepGuard != null) idleSleepGuard.close();
|
||||
// Release the broker connection last among message resources (no-op for the in-memory inbox).
|
||||
if (replyInbox instanceof AutoCloseable closeable) {
|
||||
try {
|
||||
closeable.close();
|
||||
} catch (Exception e) {
|
||||
log.debug("reply inbox close: {}", e.toString());
|
||||
}
|
||||
}
|
||||
// CB-637: the coordination connection goes with it — after the loop that reads it has
|
||||
// stopped, so no tick can be mid-ack against a closed channel.
|
||||
if (leadMailbox != null) {
|
||||
try {
|
||||
leadMailbox.close();
|
||||
} catch (Exception e) {
|
||||
log.debug("lead mailbox close: {}", e.toString());
|
||||
}
|
||||
}
|
||||
router.close();
|
||||
}));
|
||||
|
||||
// CB-185: give FleetApp both daemons — /healthz must require both to answer and
|
||||
// GET /sessions must merge across both, or a down/unpolled member daemon is invisible.
|
||||
// fleetd #111: live (re-read-per-request) memberCredentials view for GET /member-credentials —
|
||||
// same hot-reload shape as the memberCredentials supplier passed to ClaudeCodeLauncher above.
|
||||
// fleetd #297: quarantineSource/outageSource are the SAME instances passed to FleetMcp above —
|
||||
// GET /profiles must report the identical quarantine/cool-off facts as fleet_profiles.
|
||||
Javalin app = new FleetApp(herdr, memberHerdr, workers, sessions, messages, presence, mcp.servlet(),
|
||||
callers, metrics, deliverable,
|
||||
() -> MemberCredentialPolicyView.of(config.get().memberCredentials()),
|
||||
quarantineSource, outageSource, loopHealth).build();
|
||||
app.start(cfg.bind().host(), cfg.bind().port());
|
||||
log.info("fleetd listening on {}:{}, herdr socket {}",
|
||||
cfg.bind().host(), cfg.bind().port(), socket);
|
||||
// fleetd #612 A-gaps (gap 2): everything from here on — including the two post-validation
|
||||
// reports that used to run inline right here (reportRoleFallbackGaps,
|
||||
// assertChartersNameOnlyRegisteredTools) — now lives in FleetdAssembly.assembleAndStart,
|
||||
// built against a real ResourcePorts. Unit A originally moved only the socket/broker/HTTP
|
||||
// composition and left those two calls here, between validateAll() and the assembly call —
|
||||
// outside the boundary FleetdAssemblyLifecycleTest drives, so deleting either call
|
||||
// compiled clean and left the whole suite green. Moving the boundary to start immediately
|
||||
// after validateAll() (this line) puts both back under test, in the same relative order,
|
||||
// before either one does any I/O — see FleetdAssembly's javadoc for the full boot-order
|
||||
// contract this preserves exactly.
|
||||
// fleetd #625: the same `ports` the guard check above just used, not a second
|
||||
// ResourcePorts.system() instance — see this method's own javadoc.
|
||||
FleetdAssembly.assembleAndStart(new AssemblyInputs(cfg, config, guard), ports);
|
||||
}
|
||||
|
||||
/**
|
||||
@@ -1651,7 +1075,33 @@ public final class Fleetd {
|
||||
*/
|
||||
static FleetMcp.LeadConfigDirSource leadConfigDirSource(Supplier<Map<String, FleetConfig.Profile>> profiles,
|
||||
Map<String, FleetConfig.Leader> leaders) {
|
||||
return new FleetMcp.LeadConfigDirSource(leadConfigDirLookup(profiles, leaders));
|
||||
return new FleetMcp.LeadConfigDirSource(leadConfigDirLookup(profiles, leaders),
|
||||
leadContextWindowLookup(profiles, leaders));
|
||||
}
|
||||
|
||||
/**
|
||||
* Per-lead-name factory for the effective auto-compact window {@link LeadContextGauge} scales
|
||||
* its HIGH threshold against — the same {@code fleet.leaders.<name>.profile} link {@link
|
||||
* #leadConfigDirLookup} already follows, one step further to that profile's own {@link
|
||||
* FleetConfig.Profile#effectiveAutoCompactWindow()}. A lead entry that names no
|
||||
* {@code profile:}, or whose named profile is not configured, or whose profile resolves no
|
||||
* window at all, returns {@code null} — {@link LeadContextGauge} then falls back to its own
|
||||
* fixed HIGH threshold.
|
||||
*/
|
||||
static Function<String, Long> leadContextWindowLookup(Supplier<Map<String, FleetConfig.Profile>> profiles,
|
||||
Map<String, FleetConfig.Leader> leaders) {
|
||||
return leadName -> {
|
||||
FleetConfig.Leader lead = leaders.get(leadName);
|
||||
if (lead == null || lead.profile() == null || lead.profile().isBlank()) {
|
||||
return null;
|
||||
}
|
||||
FleetConfig.Profile leadProfile = profiles.get().get(lead.profile());
|
||||
if (leadProfile == null) {
|
||||
return null;
|
||||
}
|
||||
Integer window = leadProfile.effectiveAutoCompactWindow();
|
||||
return window == null ? null : window.longValue();
|
||||
};
|
||||
}
|
||||
|
||||
/**
|
||||
@@ -1673,22 +1123,30 @@ public final class Fleetd {
|
||||
* @param liveLeadTerminals terminal id → lead name for every CURRENTLY recognised lead
|
||||
* @param configDirForLeadName lead name → {@code configDir}, normally {@link
|
||||
* #leadConfigDirLookup}'s return
|
||||
* @param windowForLeadName lead name → that lead's profile's effective auto-compact window,
|
||||
* normally {@link #leadContextWindowLookup}'s return, and passed
|
||||
* through to {@link LeadContextGauge#read} so the heartbeat's own
|
||||
* HIGH reading scales with that lead's real window, or {@code null}
|
||||
* when it cannot be resolved — either way {@link LeadContextGauge}
|
||||
* falls back to its own fixed HIGH threshold
|
||||
*/
|
||||
static Function<String, LeadContextGauge.Reading> leadContextLookup(LeadContextGauge gauge, AgentControl agents,
|
||||
Supplier<Map<String, String>> liveLeadTerminals, Function<String, String> configDirForLeadName) {
|
||||
Supplier<Map<String, String>> liveLeadTerminals, Function<String, String> configDirForLeadName,
|
||||
Function<String, Long> windowForLeadName) {
|
||||
return terminal -> {
|
||||
String leadName = liveLeadTerminals.get().get(terminal);
|
||||
if (leadName == null) {
|
||||
return LeadContextGauge.Reading.unknown();
|
||||
}
|
||||
String configDir = configDirForLeadName.apply(leadName);
|
||||
Long effectiveWindow = windowForLeadName.apply(leadName);
|
||||
Agent live;
|
||||
try {
|
||||
live = agents.get(terminal);
|
||||
} catch (RuntimeException e) {
|
||||
return LeadContextGauge.Reading.unknown();
|
||||
}
|
||||
return gauge.read(configDir, live.sessionId(), live.agentType());
|
||||
return gauge.read(configDir, live.sessionId(), live.agentType(), effectiveWindow);
|
||||
};
|
||||
}
|
||||
|
||||
@@ -1700,9 +1158,10 @@ public final class Fleetd {
|
||||
* (see {@code FleetdLeadConfigDirSourceWiringTest}'s javadoc for the measured gap this shape closes).
|
||||
*/
|
||||
static LeadHeartbeatLoop.LeadContextSource leadContextSource(LeadContextGauge gauge, AgentControl agents,
|
||||
Supplier<Map<String, String>> liveLeadTerminals, Function<String, String> configDirForLeadName) {
|
||||
Supplier<Map<String, String>> liveLeadTerminals, Function<String, String> configDirForLeadName,
|
||||
Function<String, Long> windowForLeadName) {
|
||||
return new LeadHeartbeatLoop.LeadContextSource(
|
||||
leadContextLookup(gauge, agents, liveLeadTerminals, configDirForLeadName));
|
||||
leadContextLookup(gauge, agents, liveLeadTerminals, configDirForLeadName, windowForLeadName));
|
||||
}
|
||||
|
||||
/**
|
||||
@@ -1814,10 +1273,18 @@ public final class Fleetd {
|
||||
ReplyInbox open(String uri, int prefetch);
|
||||
}
|
||||
|
||||
/** Injection seam for {@link #openLeadMailbox}: production binds {@link LeadMailbox#open}. */
|
||||
/**
|
||||
* Injection seam for {@link #openLeadMailbox}: production binds {@link LeadMailbox#open}.
|
||||
*
|
||||
* <p>fleetd #612 A-gaps (gap 1): returns {@link LeadChannelHandle}, not the concrete {@link
|
||||
* LeadMailbox}, so a test can supply a fake closeable channel instead of a real broker
|
||||
* connection — see {@link LeadChannelHandle}'s own javadoc for why the narrower type existed
|
||||
* and what widening it to add {@code close()} costs (nothing: {@code LeadMailbox} already
|
||||
* implements it).
|
||||
*/
|
||||
@FunctionalInterface
|
||||
interface LeadMailboxOpener {
|
||||
LeadMailbox open(String uri, String selfCoordId, int prefetch);
|
||||
LeadChannelHandle open(String uri, String selfCoordId, int prefetch);
|
||||
}
|
||||
|
||||
/**
|
||||
@@ -1840,7 +1307,7 @@ public final class Fleetd {
|
||||
* fleet still works exactly as it did before this feature existed.</li>
|
||||
* </ul>
|
||||
*/
|
||||
static LeadMailbox openLeadMailbox(FleetConfig.Coordinator coordinator, Map<String, String> env,
|
||||
static LeadChannelHandle openLeadMailbox(FleetConfig.Coordinator coordinator, Map<String, String> env,
|
||||
LeadMailboxOpener opener) {
|
||||
if (coordinator == null) {
|
||||
return null; // opt-in: nothing configured, nothing to say
|
||||
@@ -1860,7 +1327,7 @@ public final class Fleetd {
|
||||
return null;
|
||||
}
|
||||
try {
|
||||
LeadMailbox mailbox = opener.open(uri, coordinator.selfId(), coordinator.prefetchOrDefault());
|
||||
LeadChannelHandle mailbox = opener.open(uri, coordinator.selfId(), coordinator.prefetchOrDefault());
|
||||
log.info("lead coordination: ON as coord-id {} (prefetch={})",
|
||||
coordinator.selfId(), coordinator.prefetchOrDefault());
|
||||
return mailbox;
|
||||
@@ -2240,16 +1707,20 @@ public final class Fleetd {
|
||||
}
|
||||
|
||||
/**
|
||||
* fleetd #474: the one place both the startup call (right after {@code cfg.validateAll()} in
|
||||
* {@link #main}) and the reload call (wired into {@code config}'s {@code extraValidation} above,
|
||||
* via a method reference to this method) go through, so the two can never drift into checking
|
||||
* different things. Extracted only to give {@link ConfigRef}'s {@code Consumer<FleetConfig>}
|
||||
* hook a {@code FleetConfig -> void} shape to bind to — {@link CharterToolSurface} itself still
|
||||
* takes the raw charter map and knows nothing about {@code ConfigRef} or {@code Fleetd}.
|
||||
* fleetd #474: the one place both the startup call (fleetd #612 A-gaps gap 2: right after
|
||||
* {@code cfg.validateAll()} in {@link FleetdAssembly#assembleAndStart}, immediately after
|
||||
* {@code main} hands off to it — see that method's javadoc) and the reload call (wired into
|
||||
* {@code config}'s {@code extraValidation} above, via a method reference to this method) go
|
||||
* through, so the two can never drift into checking different things. Extracted only to give
|
||||
* {@link ConfigRef}'s {@code Consumer<FleetConfig>} hook a {@code FleetConfig -> void} shape to
|
||||
* bind to — {@link CharterToolSurface} itself still takes the raw charter map and knows nothing
|
||||
* about {@code ConfigRef} or {@code Fleetd}.
|
||||
*
|
||||
* <p>Package-private so a test can call it directly the same way the other startup-report
|
||||
* helpers above are tested, without needing to drive {@link #main} for a unit-level check;
|
||||
* {@code FleetdStartupValidationTest} proves the startup call site, and {@code
|
||||
* {@code FleetdStartupValidationTest} proves the startup call site by driving {@code main}
|
||||
* itself end to end (the throw still happens before {@code main} reaches any real socket or
|
||||
* broker work, since the assembly runs this before either), and {@code
|
||||
* FleetdConfigRefCharterToolSurfaceWiringTest} — by constructing {@code ConfigRef} with this
|
||||
* exact method reference, the same way {@code main} does above — proves the reload call site.
|
||||
* {@code dev.ltms.fleet.config.ConfigRefTest} pins the same reload behaviour too, through an
|
||||
@@ -2291,13 +1762,16 @@ public final class Fleetd {
|
||||
record HerdrAwaitOutcome(HerdrWaitResult result, long elapsedNanos) {}
|
||||
|
||||
/**
|
||||
* The real per-poll wait {@link #main} passes to {@link #awaitHerdr}: sleep
|
||||
* {@link #HERDR_WAIT_POLL_MILLIS}, and on interruption re-set the thread's interrupt flag
|
||||
* The real per-poll wait {@link FleetdAssembly#assembleAndStart} passes to {@link #awaitHerdr}:
|
||||
* sleep {@link #HERDR_WAIT_POLL_MILLIS}, and on interruption re-set the thread's interrupt flag
|
||||
* rather than throwing — {@link #awaitHerdr} detects an interruption by checking {@link
|
||||
* Thread#isInterrupted()} right after this returns, so a poller that swallowed the flag
|
||||
* instead of restoring it would make that check silently miss the interruption.
|
||||
*
|
||||
* <p>Package-private (fleetd #612 Unit A), not {@code private}: {@code FleetdAssembly}, a
|
||||
* different class in this package, needs a {@code Runnable} reference to this exact method.
|
||||
*/
|
||||
private static void sleepHerdrPoll() {
|
||||
static void sleepHerdrPoll() {
|
||||
try {
|
||||
Thread.sleep(HERDR_WAIT_POLL_MILLIS);
|
||||
} catch (InterruptedException ie) {
|
||||
|
||||
@@ -0,0 +1,537 @@
|
||||
package dev.ltms.fleet;
|
||||
|
||||
import dev.ltms.fleet.auth.CallerResolver;
|
||||
import dev.ltms.fleet.auth.MemberRegistry;
|
||||
import dev.ltms.fleet.config.ConfigRef;
|
||||
import dev.ltms.fleet.config.ConfigWatcher;
|
||||
import dev.ltms.fleet.config.FleetConfig;
|
||||
import dev.ltms.fleet.guard.SubscriptionGuard;
|
||||
import dev.ltms.fleet.health.FleetHealthMonitor;
|
||||
import dev.ltms.fleet.herdr.AgentControl;
|
||||
import dev.ltms.fleet.herdr.HerdrClient;
|
||||
import dev.ltms.fleet.herdr.HerdrRouter;
|
||||
import dev.ltms.fleet.herdr.LeadTabScanner;
|
||||
import dev.ltms.fleet.herdr.PaneLocator;
|
||||
import dev.ltms.fleet.herdr.UnixSocketHerdrClient;
|
||||
import dev.ltms.fleet.inject.BackendErrorPatternLookup;
|
||||
import dev.ltms.fleet.inject.BackendErrorSink;
|
||||
import dev.ltms.fleet.inject.CompletionResolver;
|
||||
import dev.ltms.fleet.inject.ExhaustedPatternLookup;
|
||||
import dev.ltms.fleet.inject.ExhaustionSink;
|
||||
import dev.ltms.fleet.inject.Injector;
|
||||
import dev.ltms.fleet.inject.LiveExhaustedPatterns;
|
||||
import dev.ltms.fleet.inject.MemberPresence;
|
||||
import dev.ltms.fleet.inject.StatusPoller;
|
||||
import dev.ltms.fleet.inject.TurnListener;
|
||||
import dev.ltms.fleet.lead.LeadContextGauge;
|
||||
import dev.ltms.fleet.lead.LeadLauncher;
|
||||
import dev.ltms.fleet.lead.LeadRollover;
|
||||
import dev.ltms.fleet.mcp.ConnectionIdentity;
|
||||
import dev.ltms.fleet.mcp.FleetMcp;
|
||||
import dev.ltms.fleet.mcp.LsofPeerPidLookup;
|
||||
import dev.ltms.fleet.mcp.LsofProcessCwdLookup;
|
||||
import dev.ltms.fleet.mcp.PrimaryRegistry;
|
||||
import dev.ltms.fleet.member.CompositePeerLauncher;
|
||||
import dev.ltms.fleet.member.HerdrPeerLauncher;
|
||||
import dev.ltms.fleet.member.MemberCredentialPolicyView;
|
||||
import dev.ltms.fleet.metrics.FleetMetrics;
|
||||
import dev.ltms.fleet.metrics.Metrics;
|
||||
import dev.ltms.fleet.msg.LeadChannelHandle;
|
||||
import dev.ltms.fleet.msg.LeadCoordLoop;
|
||||
import dev.ltms.fleet.msg.LeadHeartbeatLoop;
|
||||
import dev.ltms.fleet.msg.MessageService;
|
||||
import dev.ltms.fleet.msg.Rendezvous;
|
||||
import dev.ltms.fleet.msg.ReplyInbox;
|
||||
import dev.ltms.fleet.msg.ReplyPushLoop;
|
||||
import dev.ltms.fleet.peer.PeerLauncher;
|
||||
import dev.ltms.fleet.placement.BackendOutagePolicy;
|
||||
import dev.ltms.fleet.placement.BackendQuarantine;
|
||||
import dev.ltms.fleet.power.CaffeinateSleepAssertionMechanism;
|
||||
import dev.ltms.fleet.power.IdleSleepGuard;
|
||||
import dev.ltms.fleet.rest.FleetApp;
|
||||
import dev.ltms.fleet.session.GitWorktrees;
|
||||
import dev.ltms.fleet.session.SessionManager;
|
||||
import dev.ltms.fleet.session.SessionReaper;
|
||||
import io.javalin.Javalin;
|
||||
import org.slf4j.Logger;
|
||||
import org.slf4j.LoggerFactory;
|
||||
|
||||
import java.nio.file.Path;
|
||||
import java.util.ArrayList;
|
||||
import java.util.LinkedHashMap;
|
||||
import java.util.LinkedHashSet;
|
||||
import java.util.List;
|
||||
import java.util.Map;
|
||||
import java.util.Set;
|
||||
import java.util.concurrent.ConcurrentHashMap;
|
||||
import java.util.concurrent.ScheduledExecutorService;
|
||||
import java.util.concurrent.TimeUnit;
|
||||
import java.util.concurrent.atomic.AtomicReference;
|
||||
import java.util.function.Function;
|
||||
import java.util.function.Predicate;
|
||||
import java.util.function.Supplier;
|
||||
import java.util.regex.Pattern;
|
||||
import java.util.stream.Collectors;
|
||||
|
||||
/**
|
||||
* fleetd #612 Unit A: the real boot assembly, extracted out of {@code Fleetd.main} so a test can
|
||||
* drive it directly. {@link #assembleAndStart} is <em>the same statements {@code main} used to run
|
||||
* inline</em>, in the same order, against a real {@link ResourcePorts} in production and a fake one
|
||||
* in a test — see {@code FleetdAssemblyLifecycleTest}. {@code Fleetd.main} keeps config loading and
|
||||
* {@code cfg.validateAll()}; everything from immediately after that call onward moved here —
|
||||
* including, since fleetd #612 A-gaps (gap 2), the two post-validation reports ({@code
|
||||
* reportRoleFallbackGaps}, {@code assertChartersNameOnlyRegisteredTools}) that Unit A originally
|
||||
* left behind in {@code main}. Those two calls do no I/O themselves, but leaving them outside this
|
||||
* method meant deleting either one compiled clean and left the whole suite green — nothing drove
|
||||
* {@code main} itself, so nothing could notice. They run first here, in the same relative order,
|
||||
* before the herdr socket or anything else that touches the outside world.
|
||||
*
|
||||
* <p><strong>Construction and start order is preserved exactly, on purpose.</strong> This is not
|
||||
* rebuilt into "construct everything, then start everything" — that would change boot timing. The
|
||||
* order recorded before any code moved (see the ticket and {@code FleetdAssemblyLifecycleTest}):
|
||||
* {@code SessionReaper} starts first (if {@code lifecycle.idleTtlSeconds} is configured), then
|
||||
* {@link StatusPoller}, then the optional {@link LeadHeartbeatLoop} and {@link FleetHealthMonitor},
|
||||
* then the optional {@link LeadCoordLoop} and {@link ConfigWatcher}, and the HTTP server starts
|
||||
* last of all. The close order (see {@link FleetdRuntime#close()}) is the mirror the original
|
||||
* shutdown hook always used.
|
||||
*
|
||||
* <p><strong>One statement could not move without reordering startup.</strong> {@code Fleetd.main}
|
||||
* registered its shutdown hook <em>before</em> building the Javalin {@code FleetApp} — the hook
|
||||
* itself never touched {@code app} (it still doesn't; see {@link FleetdRuntime#close()}), but the
|
||||
* hook needs a {@link FleetdRuntime} to close over, and the ticket asks that runtime to also own
|
||||
* {@code FleetApp}. Building the runtime before the app exists and mutating it afterward (via
|
||||
* {@link FleetdRuntime#attachApp}) preserves the exact original order — hook registered, then app
|
||||
* built, then HTTP started — without moving the app's construction earlier or the hook's
|
||||
* registration later. That is the one seam this ticket did not get to pin any other way.
|
||||
*
|
||||
* <p><strong>No inert variant.</strong> Deliberately, there is no overload of this method that
|
||||
* accepts a smaller/optional {@link ResourcePorts} or defaults one internally. A future edit that
|
||||
* wants to skip {@code FleetdAssembly} entirely and build its own graph is still possible — no
|
||||
* static analysis stops that — but it cannot do so by quietly swapping this call for an inert
|
||||
* substitute that still compiles, because none exists.
|
||||
*/
|
||||
final class FleetdAssembly {
|
||||
|
||||
private static final Logger log = LoggerFactory.getLogger(Fleetd.class);
|
||||
|
||||
/** CB-637: how often the lead coordination loop looks for peer messages — see {@code Fleetd}'s own constant. */
|
||||
private static final long LEAD_COORD_INTERVAL_MS = 3_000L;
|
||||
|
||||
private FleetdAssembly() {
|
||||
}
|
||||
|
||||
static FleetdRuntime assembleAndStart(AssemblyInputs inputs, ResourcePorts ports) {
|
||||
FleetConfig cfg = inputs.cfg();
|
||||
ConfigRef config = inputs.config();
|
||||
SubscriptionGuard guard = inputs.guard();
|
||||
|
||||
// fleetd #612 A-gaps (gap 2): moved in from Fleetd.main, immediately after cfg.validateAll()
|
||||
// there — the exact point main used to call these two, and still the first thing this
|
||||
// method does, before any socket or broker work below. See this class's javadoc and each
|
||||
// method's own for why they run here rather than in FleetConfig#validateAll() itself.
|
||||
Fleetd.reportRoleFallbackGaps(cfg);
|
||||
Fleetd.assertChartersNameOnlyRegisteredTools(cfg);
|
||||
|
||||
Path socket = cfg.herdrSocket() != null && !cfg.herdrSocket().isBlank()
|
||||
? Path.of(cfg.herdrSocket())
|
||||
: UnixSocketHerdrClient.defaultSocketPath();
|
||||
|
||||
HerdrClient herdr = ports.connectHerdr(socket);
|
||||
HerdrClient memberHerdr = cfg.memberHerdrSocket() != null && !cfg.memberHerdrSocket().isBlank()
|
||||
? ports.connectHerdr(Path.of(cfg.memberHerdrSocket()))
|
||||
: herdr;
|
||||
AtomicReference<Supplier<Map<String, String>>> leadsRef = new AtomicReference<>(Map::of);
|
||||
HerdrRouter router = new HerdrRouter(herdr, memberHerdr,
|
||||
target -> leadsRef.get().get().containsKey(target));
|
||||
// CB-402: one adapter per configured peer kind, fronted by a composite router. A profile's
|
||||
// `kind:` selects its adapter — claude-code (the default) and opencode partition the profile
|
||||
// set — and the composite dispatches each SPI call to the adapter that owns the profile/pane.
|
||||
Map<String, FleetConfig.Profile> claudeProfiles = new LinkedHashMap<>();
|
||||
Map<String, FleetConfig.Profile> opencodeProfiles = new LinkedHashMap<>();
|
||||
cfg.profiles().forEach((name, w) -> {
|
||||
if (w.isOpenCode()) {
|
||||
opencodeProfiles.put(name, w);
|
||||
} else {
|
||||
claudeProfiles.put(name, w);
|
||||
}
|
||||
});
|
||||
List<HerdrPeerLauncher> adapters = new ArrayList<>();
|
||||
// fleetd #175: the daemon's real ExhaustionSink can only be built once `sessions` exists
|
||||
// (below), but `sessions` needs `workers`, which needs the adapters built right here — a
|
||||
// genuine cycle. Break it exactly like liveCountRef below: a forwarding sink built now,
|
||||
// pointed at the real one once it exists.
|
||||
AtomicReference<ExhaustionSink> exhaustionSinkRef = new AtomicReference<>(ExhaustionSink.none());
|
||||
ExhaustionSink forwardingExhaustionSink = Fleetd.forwardingExhaustionSink(exhaustionSinkRef);
|
||||
// The claude-code adapter is the always-present default; keep it even with no profiles (so a
|
||||
// bridge configured with no workers, or opencode-only, still has a well-defined base adapter)
|
||||
// unless opencode is the only kind configured.
|
||||
if (!claudeProfiles.isEmpty() || opencodeProfiles.isEmpty()) {
|
||||
adapters.add(Fleetd.claudeCodeLauncher(router.memberAgents(), router.memberSpaces(), guard,
|
||||
claudeProfiles, cfg, config));
|
||||
}
|
||||
if (!opencodeProfiles.isEmpty()) {
|
||||
adapters.add(Fleetd.openCodeLauncher(router.memberAgents(), router.memberSpaces(),
|
||||
opencodeProfiles, cfg, config, forwardingExhaustionSink));
|
||||
}
|
||||
AtomicReference<Function<String, Integer>> liveCountRef = new AtomicReference<>(_ -> 0);
|
||||
// CB-578 stage B: one quarantine tracker for the whole daemon, shared between the launcher
|
||||
// (checked at spawn) and the exhaustion sink wired in below (written on BACKEND_EXHAUSTED).
|
||||
BackendQuarantine quarantine = BackendQuarantine.withEscalation(ports.nanoClock(),
|
||||
TimeUnit.SECONDS.toNanos(cfg.quarantineCooldownSeconds()));
|
||||
// fleetd #201 Unit 5: one outage-cool-off tracker for the whole daemon, shared between the
|
||||
// launcher (checked at spawn, like `quarantine` above) and the backend-error sink wired in
|
||||
// below (written on a classified backend error).
|
||||
BackendOutagePolicy outagePolicy = new BackendOutagePolicy(ports.nanoClock());
|
||||
PeerLauncher workers = new CompositePeerLauncher(
|
||||
adapters,
|
||||
cfg.effectiveDefaultProfile(),
|
||||
config,
|
||||
profileName -> liveCountRef.get().apply(profileName),
|
||||
quarantine,
|
||||
outagePolicy);
|
||||
// fleetd #422 follow-up: say which of the three model-gate states the daemon booted into.
|
||||
log.info("model gate (fleetd #422): {}", Fleetd.modelGateCoverageLine(workers.modelGateState()));
|
||||
// CB-504: under supervision (launchd/systemd) fleetd can start before herdr's socket
|
||||
// exists. Wait, then degrade rather than die: serving with /healthz reporting "degraded" is
|
||||
// strictly more useful than exiting.
|
||||
Fleetd.HerdrAwaitOutcome herdrOutcome = Fleetd.awaitHerdr(herdr, ports.nanoClock(), ports.herdrPollWait());
|
||||
boolean herdrUp = Fleetd.logHerdrWaitOutcomeAndShouldReap(herdrOutcome);
|
||||
if (herdrUp) {
|
||||
// CB-117: herdr keeps worker panes alive across a daemon restart, and their ids died
|
||||
// with the previous process — reap those leaked orphans now, before we start serving.
|
||||
workers.reapOrphanWorkers();
|
||||
}
|
||||
|
||||
// CB-301: authoritative session registry + lifecycle FSM on top of ClaudeCodeLauncher.
|
||||
// CB-301-ext: worktree provisioning seam, optionally rooted at a configured directory.
|
||||
// CB-303 part 2: context cap is opt-in and disabled (0) when absent/null.
|
||||
int contextCap = 0;
|
||||
if (cfg.lifecycle() != null && cfg.lifecycle().contextCap() != null
|
||||
&& cfg.lifecycle().contextCap() > 0) {
|
||||
contextCap = cfg.lifecycle().contextCap();
|
||||
}
|
||||
boolean clearAfterTurn = cfg.lifecycle() != null && cfg.lifecycle().clearAfterTurn();
|
||||
SessionManager sessions = new SessionManager(workers,
|
||||
new GitWorktrees(cfg.worktreeRoot(), cfg.worktreeGroup(), cfg.memberSkills()),
|
||||
ports.nanoClock(), contextCap, clearAfterTurn);
|
||||
liveCountRef.set(profileName -> Fleetd.liveSessionCount(sessions.roster(), profileName));
|
||||
|
||||
// Idle-sleep guard: hold an OS-level assertion against idle sleep while at least one
|
||||
// member is live. No-op (never constructed) off macOS or when idleSleepGuard.enabled is
|
||||
// explicitly false; the mechanism itself is additionally a no-op if 'caffeinate' cannot be
|
||||
// started, so this can never fail a spawn, a release, or startup.
|
||||
boolean idleSleepGuardEnabled = cfg.idleSleepGuard() == null || cfg.idleSleepGuard().isEnabled();
|
||||
final IdleSleepGuard idleSleepGuard;
|
||||
if (idleSleepGuardEnabled) {
|
||||
idleSleepGuard = new IdleSleepGuard(new CaffeinateSleepAssertionMechanism(), sessions::size);
|
||||
sessions.onAcquire(_ -> idleSleepGuard.recheck());
|
||||
sessions.onRelease(_ -> idleSleepGuard.recheck());
|
||||
} else {
|
||||
idleSleepGuard = null;
|
||||
}
|
||||
|
||||
// CB-303 part 1: idle-ttl reaper — only when configured, defaults to disabled. FIRST of the
|
||||
// recurring background loops to start — see this class's own javadoc for the full order.
|
||||
final SessionReaper reaper;
|
||||
if (cfg.lifecycle() != null
|
||||
&& cfg.lifecycle().idleTtlSeconds() != null
|
||||
&& cfg.lifecycle().idleTtlSeconds() > 0) {
|
||||
reaper = new SessionReaper(sessions, cfg.lifecycle().idleTtlSeconds());
|
||||
reaper.start();
|
||||
} else {
|
||||
reaper = null;
|
||||
}
|
||||
|
||||
// CB-530: every pane the config names as a lead, merged from `leaders:` and the legacy
|
||||
// singular pin. PrimaryRegistry below still tracks ONE terminal — it addresses the push
|
||||
// loop's nudges, which need a single destination — so it keeps the legacy pin.
|
||||
Map<String, String> leadTerminals = cfg.leaderTerminals();
|
||||
if (leadTerminals.size() > 1) {
|
||||
log.info("leads: {} panes recognised {}", leadTerminals.size(), leadTerminals.values());
|
||||
}
|
||||
// CB-531/CB-579: discover leads by the tab labels the operator writes, one scanner per
|
||||
// configured lead's own exact `tab:` label.
|
||||
final Supplier<Map<String, String>> leads;
|
||||
var leaders = cfg.fleet().leaders();
|
||||
if (!leaders.isEmpty()) {
|
||||
Map<String, String> tabToName = new LinkedHashMap<>();
|
||||
leaders.forEach((name, leader) -> {
|
||||
if (leader != null && leader.tab() != null && !leader.tab().isBlank()) {
|
||||
tabToName.put(leader.tab(), name);
|
||||
}
|
||||
});
|
||||
int scanIntervalSeconds = leaders.values().iterator().next().scanIntervalSeconds();
|
||||
// This must use the lead daemon: scanning member tabs would demote the lead to a worker.
|
||||
leads = new LeadTabScanner(herdr, tabToName, Set.of(),
|
||||
TimeUnit.SECONDS.toNanos(scanIntervalSeconds), ports.nanoClock());
|
||||
log.info("lead scan: tabs {} host a lead (rescan every {}s, shared fleet space)",
|
||||
tabToName.keySet(), scanIntervalSeconds);
|
||||
} else {
|
||||
leads = () -> leadTerminals;
|
||||
}
|
||||
leadsRef.set(leads);
|
||||
|
||||
// CB-558: start any declared lead that is not already running. After the scanner is built,
|
||||
// and only when herdr answered — the launcher's whole safety property is that it can count
|
||||
// live leads first, and must never guess and risk a second orchestrator.
|
||||
if (herdrUp && !leaders.isEmpty()) {
|
||||
int launched = new LeadLauncher(router.leadAgents(), router.leadSpaces(), cfg).ensureLeads();
|
||||
if (launched > 0) {
|
||||
log.info("lead auto-launch: {} lead(s) started", launched);
|
||||
}
|
||||
}
|
||||
|
||||
// CB-548: config-declared architect slots. Nothing here spawns a slot; the terminal → slot
|
||||
// binding is owned by the registry and empty at startup.
|
||||
MemberRegistry members = MemberRegistry.live(() -> config.get().fleet());
|
||||
sessions.setMemberLifecycle(members);
|
||||
if (!members.slots().isEmpty()) {
|
||||
log.info("member slots: {} configured {} — none bound yet (a slot is idle until the "
|
||||
+ "spawn lifecycle binds a live terminal to it)",
|
||||
members.slots().size(), members.slots().keySet());
|
||||
}
|
||||
|
||||
// Status-gated injector (CB-103): the single writer into workers, fed by a poller.
|
||||
// CB-106: a confirmed turn completion resolves a blocked send whose worker never replied.
|
||||
Rendezvous rendezvous = new Rendezvous();
|
||||
// CB-578 stage A: classify a completion-fallback scrape that matches a profile's configured
|
||||
// usage-limit refusal as BACKEND_EXHAUSTED rather than handing it back as a real answer.
|
||||
LiveExhaustedPatterns liveExhaustedPatterns = Fleetd.liveExhaustedPatterns(config);
|
||||
ExhaustedPatternLookup exhaustedPatterns = Fleetd.exhaustedPatternLookup(sessions::roster, liveExhaustedPatterns);
|
||||
// The startup coverage line still reports the boot-time snapshot only.
|
||||
Set<String> exhaustedConfiguredAtStartup = cfg.profiles().entrySet().stream()
|
||||
.filter(e -> e.getValue().hasExhaustedPattern())
|
||||
.map(Map.Entry::getKey)
|
||||
.collect(Collectors.toCollection(LinkedHashSet::new));
|
||||
log.info("backend-exhausted classification (CB-578 stage A): {}",
|
||||
Fleetd.exhaustedPatternCoverageLine(cfg.profiles().keySet(), exhaustedConfiguredAtStartup));
|
||||
// fleetd #201 Unit 5: classify a completion-fallback scrape that matches a profile's
|
||||
// configured backend-error refusal (a credential outage, a provider 5xx) as a backend error
|
||||
// rather than handing it back as a real answer.
|
||||
Map<String, Pattern> errorPatternsByProfile = new LinkedHashMap<>();
|
||||
cfg.profiles().forEach((name, profile) -> {
|
||||
if (profile.hasErrorPattern()) {
|
||||
errorPatternsByProfile.put(name, Pattern.compile(profile.errorPattern()));
|
||||
}
|
||||
});
|
||||
BackendErrorPatternLookup backendErrorPatterns = Fleetd.backendErrorPatternLookup(sessions::roster,
|
||||
errorPatternsByProfile);
|
||||
log.info("backend-error classification (fleetd #201 Unit 5): {}",
|
||||
Fleetd.errorPatternCoverageLine(cfg.profiles().keySet(), errorPatternsByProfile.keySet()));
|
||||
// CB-578 stage B: on a classification that actually wins, quarantine the exhausted profile's
|
||||
// CREDENTIAL, so a profile sharing that credential is refused too, not just the one that
|
||||
// happened to report it.
|
||||
Map<String, String> quarantineReasonByCredential = new ConcurrentHashMap<>();
|
||||
// fleetd #175: point the forwarding sink handed to OpenCodeLauncher above at the real one,
|
||||
// now that `sessions` exists to resolve target -> session -> profile.
|
||||
ExhaustionSink exhaustionSink = Fleetd.publishExhaustionSink(exhaustionSinkRef, sessions, config,
|
||||
quarantine, quarantineReasonByCredential, cfg);
|
||||
// fleetd #201 Unit 5: the production BackendErrorSink needs `pushLoop` (built further below,
|
||||
// after `sessions`) to tell a lead about an incident or an unmapped target — the same
|
||||
// construction-order cycle `exhaustionSinkRef` breaks above, broken the same way: a mutable
|
||||
// holder set once `pushLoop` exists, read lazily from inside the lambda built here.
|
||||
AtomicReference<ReplyPushLoop> pushLoopRef = new AtomicReference<>();
|
||||
BackendErrorSink backendErrorSink = Fleetd.backendErrorSink(sessions, () -> config.get().profiles(),
|
||||
outagePolicy, pushLoopRef::get);
|
||||
AgentControl agents = router.memberAgents();
|
||||
CompletionResolver completion = new CompletionResolver(agents, rendezvous, exhaustedPatterns,
|
||||
exhaustionSink, backendErrorPatterns, backendErrorSink, ports.nanoClock(),
|
||||
Fleetd.worktreeBranchLookup(sessions::roster));
|
||||
// CB-113: deliver only to an available worker (its MCP is connected), never its boot window.
|
||||
// CB-301: the manager's presence bridge records availability and drives SPAWNING → READY.
|
||||
MemberPresence presence = sessions.asPresence();
|
||||
TurnListener turnListener = Fleetd.turnListener(completion, sessions);
|
||||
Predicate<String> deliverable = Fleetd.deliverableTo(presence, leads);
|
||||
// fleetd #556: registration is wired directly to `completion`, not folded into the
|
||||
// `turnListener` fan-out above — so it survives `sessions.onDelivered` (or any future
|
||||
// listener) throwing, regardless of call order.
|
||||
Injector injector = new Injector(router, turnListener, deliverable,
|
||||
presence::forget, Fleetd.turnRegistrar(completion));
|
||||
StatusPoller poller = new StatusPoller(router, injector, Injector.POLL_INTERVAL_MILLIS);
|
||||
poller.start(); // SECOND of the recurring background loops to start, after the reaper.
|
||||
|
||||
// CB-307: reply inbox. A broker: block selects the AMQP-backed durable adapter; absent (or
|
||||
// unusable), fleetd stays soft-state on the in-memory inbox. The AMQP inbox owns a broker
|
||||
// connection, so keep the reference to close it in the ordered shutdown hook.
|
||||
final ReplyInbox replyInbox = Fleetd.selectReplyInbox(cfg.broker(), ports.environment(),
|
||||
ports.replyInboxOpener());
|
||||
// CB-637: this daemon's lead-to-lead mailbox on the SHARED coordination vhost — a separate
|
||||
// broker from the reply inbox by design. Absent a coordinator: block this is null and every
|
||||
// lead path below is simply not wired, exactly the behaviour before this ticket. It owns a
|
||||
// broker connection, so keep the reference for the ordered shutdown hook.
|
||||
final LeadChannelHandle leadMailbox = Fleetd.openLeadMailbox(cfg.coordinator(), ports.environment(),
|
||||
ports.leadMailboxOpener());
|
||||
// CB-307: learn the primary's terminal from orchestration tool calls (or pin from config).
|
||||
String pinnedPrimaryTerminal = cfg.primary() != null ? cfg.primary().terminal() : null;
|
||||
PrimaryRegistry primaryRegistry = new PrimaryRegistry(pinnedPrimaryTerminal);
|
||||
// CB-532: `primary.terminal` is superseded and no longer needed for either of its jobs.
|
||||
if (pinnedPrimaryTerminal != null && !pinnedPrimaryTerminal.isBlank()) {
|
||||
log.warn("primary.terminal is DEPRECATED (CB-532) and can be deleted: identity now comes "
|
||||
+ "from leaders:/leadScan:, and reply nudges follow the lead that delegated. "
|
||||
+ "It still works, and is still the fallback nudge destination when a restart "
|
||||
+ "has lost the delegation map. Its pushReminders/pushBackoffMs stay valid.");
|
||||
}
|
||||
// CB-307: active push-to-primary loop — nudge the primary when replies land without an
|
||||
// open fleet_send. Uses its own lightweight scheduled executor, separate from the injector.
|
||||
int maxReminders = cfg.primary() != null ? cfg.primary().remindersOrDefault() : 5;
|
||||
long backoffMs = cfg.primary() != null ? cfg.primary().backoffMsOrDefault() : 15_000L;
|
||||
var pushScheduler = ports.newScheduler("bridge-push-");
|
||||
// CB-502: the registry is built before the service and the push loop so send/reply outcomes
|
||||
// are counted at their single funnel rather than at each of the two caller-facing surfaces.
|
||||
Metrics metrics = FleetMetrics.create(sessions, replyInbox);
|
||||
var pushLoop = new ReplyPushLoop(primaryRegistry, router.leadAgents(), replyInbox,
|
||||
pushScheduler, maxReminders, backoffMs, metrics);
|
||||
// fleetd #201 Unit 5: point the forwarding holder captured by the backendErrorSink lambda
|
||||
// above at the real push loop, now that it exists.
|
||||
pushLoopRef.set(pushLoop);
|
||||
// CB-551: idle-lead heartbeat. Opt-in; absent `leadHeartbeat:` this is never constructed, so
|
||||
// an upgraded daemon cannot silently start spending subscription on nudging an idle lead.
|
||||
// THIRD of the recurring background loops to start (optional).
|
||||
final LeadHeartbeatLoop heartbeat;
|
||||
var heartbeatScheduler = ports.newScheduler("bridge-heartbeat-");
|
||||
if (cfg.leadHeartbeat() != null) {
|
||||
var hb = cfg.leadHeartbeat();
|
||||
var leadContextGauge = new LeadContextGauge();
|
||||
// fleetd #621: the context-high notice's own wording must track this same effective
|
||||
// value — LeadRollover.confirm(...) already gates the roll on it (LeadRollover.java:480),
|
||||
// and absent `leadRollover:` entirely the roll is unusable regardless (NOT_CONFIGURED),
|
||||
// so `true` (the FleetConfig.LeadRollover default) is the safe, byte-identical fallback.
|
||||
// Carried in from Fleetd.main when #612 Unit A merged main: #622 added this line to the
|
||||
// block Unit A had already moved here, so the merge would otherwise have silently
|
||||
// dropped it — with a fully green suite, because nothing pins it (see the follow-up issue).
|
||||
boolean requireOperatorConfirm = cfg.leadRollover() == null || cfg.leadRollover().requireOperatorConfirm();
|
||||
heartbeat = new LeadHeartbeatLoop(primaryRegistry, router.leadAgents(), replyInbox, sessions::roster,
|
||||
pushLoop, heartbeatScheduler, ports.nanoClock(),
|
||||
TimeUnit.SECONDS.toNanos(hb.idleAfterSeconds()), hb.backoffMs(), hb.quietNudgeCap(),
|
||||
metrics,
|
||||
Fleetd.leadContextSource(leadContextGauge, router.leadAgents(), leads,
|
||||
Fleetd.leadConfigDirLookup(() -> config.get().profiles(), leaders),
|
||||
Fleetd.leadContextWindowLookup(() -> config.get().profiles(), leaders)),
|
||||
Boolean.TRUE.equals(hb.contextHighNudge()), requireOperatorConfirm);
|
||||
heartbeat.start();
|
||||
} else {
|
||||
heartbeat = null;
|
||||
heartbeatScheduler.shutdownNow();
|
||||
}
|
||||
// fleetd #480: lead rollover. Opt-in; absent `leadRollover:` this is never constructed.
|
||||
LeadRollover leadRollover = Fleetd.leadRollover(cfg, router.leadAgents(), config, leads);
|
||||
MessageService messages = new MessageService(router, injector, rendezvous, replyInbox,
|
||||
pushLoop, metrics);
|
||||
|
||||
// Health is a slow whole-fleet observer. Keep it separate from the 250ms delivery poller.
|
||||
// FOURTH of the recurring background loops to start (optional).
|
||||
final FleetHealthMonitor healthMonitor;
|
||||
var healthScheduler = ports.newScheduler("bridge-health-");
|
||||
if (cfg.health() != null && cfg.health().isEnabled()) {
|
||||
// CB-580: a member found GONE/NEVER_READY must fail whatever ticket is waiting on it.
|
||||
healthMonitor = new FleetHealthMonitor(agents, sessions::roster, messages, healthScheduler,
|
||||
ports.nanoClock(), ports.wallClockNanos(),
|
||||
cfg.health().intervalOrDefault(),
|
||||
cfg.health().workingSuspectAfterOrDefault(), Fleetd.healthFailTarget(messages));
|
||||
String coverage = FleetHealthMonitor.coverage(true,
|
||||
cfg.health().notifications() != null && cfg.health().notifications().configured());
|
||||
if ("detection-only".equals(coverage)) {
|
||||
log.warn("fleet health: {} (no notification sink configured)", coverage);
|
||||
} else {
|
||||
log.info("fleet health: {}", coverage);
|
||||
}
|
||||
healthMonitor.start();
|
||||
} else {
|
||||
healthMonitor = null;
|
||||
healthScheduler.shutdownNow();
|
||||
}
|
||||
|
||||
// CB-520: the reply inbox only consumes for agents this gateway owns. own on acquire,
|
||||
// release on teardown. Do this before CB-516 so the inbox is owned before any reply can land.
|
||||
sessions.onAcquire(replyInbox::own);
|
||||
// CB-516: releasing a worker must fail whatever send was waiting on it.
|
||||
sessions.onRelease(Fleetd.releaseCleanup(messages, replyInbox, primaryRegistry));
|
||||
|
||||
// MCP server face (CB-105): fleet_send/fleet_reply/fleet_status, mounted at /mcp.
|
||||
// Caller identity is resolved from the connection (peer PID → herdr pane), not arguments.
|
||||
ConnectionIdentity identity = new ConnectionIdentity(
|
||||
new PaneLocator(herdr, memberHerdr), new LsofPeerPidLookup(), new LsofProcessCwdLookup());
|
||||
|
||||
// CB-501: one resolver behind both entry paths. Worker identity still comes from the
|
||||
// connection and is never token-gated, so enabling token mode cannot lock the fleet out.
|
||||
final CallerResolver callers;
|
||||
if (cfg.auth().tokenMode()) {
|
||||
String token = ports.environment().get(cfg.auth().tokenEnv());
|
||||
if (token == null || token.isBlank()) {
|
||||
throw new IllegalStateException("auth.mode=token but env var " + cfg.auth().tokenEnv()
|
||||
+ " is unset or empty — export it before starting fleetd");
|
||||
}
|
||||
callers = CallerResolver.withLeadsAndMembers(identity, true, token, leads, members);
|
||||
log.info("auth: token mode (bearer required for non-worker callers, env {})",
|
||||
cfg.auth().tokenEnv());
|
||||
} else {
|
||||
callers = CallerResolver.withLeadsAndMembers(identity, false, null, leads, members);
|
||||
log.info("auth: loopback-trust (any loopback non-worker caller is the primary)");
|
||||
}
|
||||
|
||||
FleetMcp.QuarantineSource quarantineSource = Fleetd.quarantineSource(config, quarantine,
|
||||
liveExhaustedPatterns, quarantineReasonByCredential);
|
||||
FleetMcp.OutageSource outageSource = new FleetMcp.OutageSource(profile -> {
|
||||
var configured = config.get().profiles().get(profile);
|
||||
return configured == null ? null : configured.effectiveCredentialId();
|
||||
}, outagePolicy);
|
||||
|
||||
FleetMcp.LoopHealthSource loopHealth = Fleetd.loopHealthSource(poller, reaper);
|
||||
FleetMcp mcp = new FleetMcp(messages, workers, sessions, identity, presence,
|
||||
primaryRegistry, callers, FleetMcp.AuthorizationMode.ENFORCED, metrics,
|
||||
Fleetd.capacitySource(config, cfg, profile -> liveCountRef.get().apply(profile)),
|
||||
Fleetd.healthCoverageSource(config),
|
||||
loopHealth,
|
||||
quarantineSource,
|
||||
leadMailbox,
|
||||
outageSource,
|
||||
new FleetMcp.LeadSeatSource(Fleetd.leadSeatLookup(() -> config.get().profiles(), leaders, leads)),
|
||||
Fleetd.leadConfigDirSource(() -> config.get().profiles(), leaders),
|
||||
cfg.coordinator() == null ? List.of() : cfg.coordinator().peers(),
|
||||
leadRollover);
|
||||
|
||||
// CB-637: the receive half. Only constructed when a lead mailbox actually opened.
|
||||
final LeadCoordLoop leadCoordLoop;
|
||||
final ScheduledExecutorService leadCoordSchedulerRef;
|
||||
if (leadMailbox != null) {
|
||||
var leadCoordScheduler = ports.newScheduler("bridge-leadcoord-");
|
||||
leadCoordLoop = new LeadCoordLoop(leadMailbox, router.leadAgents(), leads, leadCoordScheduler,
|
||||
LEAD_COORD_INTERVAL_MS);
|
||||
leadCoordLoop.start(); // FIFTH of the recurring background loops to start (optional).
|
||||
leadCoordSchedulerRef = leadCoordScheduler;
|
||||
} else {
|
||||
leadCoordLoop = null;
|
||||
leadCoordSchedulerRef = null;
|
||||
}
|
||||
|
||||
// CB-559: opt-in config reload. With no `configReload:` block nothing is constructed.
|
||||
// SIXTH of the recurring background loops to start (optional).
|
||||
final ConfigWatcher configWatcher;
|
||||
if (cfg.configReload() != null && cfg.configReload().isEnabled()) {
|
||||
configWatcher = new ConfigWatcher(config, cfg.configReload().intervalSeconds());
|
||||
configWatcher.start();
|
||||
} else {
|
||||
configWatcher = null;
|
||||
}
|
||||
|
||||
// CB-303 part 3: single ordered shutdown hook — see FleetdRuntime#close() for the statements
|
||||
// this used to be. Registered here, at the exact point `main` used to register it: after
|
||||
// configWatcher, before the Javalin app exists (see this class's own javadoc for why).
|
||||
FleetdRuntime runtime = new FleetdRuntime(cfg, sessions, router, poller, messages, pushLoop, heartbeat,
|
||||
leadCoordLoop, leadCoordSchedulerRef, healthMonitor, configWatcher, mcp, reaper, idleSleepGuard,
|
||||
replyInbox, leadMailbox, completion, injector);
|
||||
ports.addShutdownHook(runtime::close);
|
||||
|
||||
// CB-185: give FleetApp both daemons — /healthz must require both to answer and
|
||||
// GET /sessions must merge across both, or a down/unpolled member daemon is invisible.
|
||||
Javalin app = new FleetApp(herdr, memberHerdr, workers, sessions, messages, presence, mcp.servlet(),
|
||||
callers, metrics, deliverable,
|
||||
() -> MemberCredentialPolicyView.of(config.get().memberCredentials()),
|
||||
quarantineSource, outageSource, loopHealth).build();
|
||||
runtime.attachApp(app);
|
||||
ports.startHttp(app, cfg.bind().host(), cfg.bind().port()); // HTTP starts LAST, always.
|
||||
log.info("fleetd listening on {}:{}, herdr socket {}",
|
||||
cfg.bind().host(), cfg.bind().port(), socket);
|
||||
return runtime;
|
||||
}
|
||||
}
|
||||
@@ -0,0 +1,166 @@
|
||||
package dev.ltms.fleet;
|
||||
|
||||
import dev.ltms.fleet.config.ConfigWatcher;
|
||||
import dev.ltms.fleet.config.FleetConfig;
|
||||
import dev.ltms.fleet.health.FleetHealthMonitor;
|
||||
import dev.ltms.fleet.herdr.HerdrRouter;
|
||||
import dev.ltms.fleet.inject.CompletionResolver;
|
||||
import dev.ltms.fleet.inject.Injector;
|
||||
import dev.ltms.fleet.inject.StatusPoller;
|
||||
import dev.ltms.fleet.mcp.FleetMcp;
|
||||
import dev.ltms.fleet.msg.LeadChannelHandle;
|
||||
import dev.ltms.fleet.msg.LeadCoordLoop;
|
||||
import dev.ltms.fleet.msg.LeadHeartbeatLoop;
|
||||
import dev.ltms.fleet.msg.MessageService;
|
||||
import dev.ltms.fleet.msg.ReplyInbox;
|
||||
import dev.ltms.fleet.msg.ReplyPushLoop;
|
||||
import dev.ltms.fleet.power.IdleSleepGuard;
|
||||
import dev.ltms.fleet.session.SessionManager;
|
||||
import dev.ltms.fleet.session.SessionReaper;
|
||||
import io.javalin.Javalin;
|
||||
import org.slf4j.Logger;
|
||||
import org.slf4j.LoggerFactory;
|
||||
|
||||
import java.util.concurrent.ScheduledExecutorService;
|
||||
|
||||
/**
|
||||
* fleetd #612 Unit A: the assembled daemon. {@link FleetdAssembly#assembleAndStart} builds exactly
|
||||
* one of these and hands it to {@code ResourcePorts.addShutdownHook}; {@link #close()} is the
|
||||
* single ordered shutdown, moved verbatim out of {@code Fleetd.main}'s old shutdown-hook
|
||||
* {@code Thread} body — same statements, same order, see that method's javadoc.
|
||||
*
|
||||
* <p><strong>Owns the real production objects, never a copy.</strong> Every package-private
|
||||
* accessor below returns the identical instance the running daemon is using. That is the entire
|
||||
* point of this class existing (see fleetd #612's problem statement): a test that inspected a
|
||||
* snapshot built alongside the real objects could pass while production silently received
|
||||
* something else — the exact shape of the #602/#606 defect this ticket exists to stop from
|
||||
* recurring one call site at a time. Nothing here is rebuilt or copied for a test's benefit.
|
||||
*/
|
||||
final class FleetdRuntime implements AutoCloseable {
|
||||
|
||||
private static final Logger log = LoggerFactory.getLogger(FleetdRuntime.class);
|
||||
|
||||
private final FleetConfig cfg;
|
||||
private final SessionManager sessions;
|
||||
private final HerdrRouter router;
|
||||
private final StatusPoller poller;
|
||||
private final MessageService messages;
|
||||
private final ReplyPushLoop pushLoop;
|
||||
private final LeadHeartbeatLoop heartbeat; // nullable — leadHeartbeat: opt-in
|
||||
private final LeadCoordLoop leadCoordLoop; // nullable — coordinator: opt-in
|
||||
private final ScheduledExecutorService leadCoordScheduler; // nullable, paired with leadCoordLoop
|
||||
private final FleetHealthMonitor healthMonitor; // nullable — health.enabled opt-in
|
||||
private final ConfigWatcher configWatcher; // nullable — configReload.enabled opt-in
|
||||
private final FleetMcp mcp;
|
||||
private final SessionReaper reaper; // nullable — lifecycle.idleTtlSeconds opt-in
|
||||
private final IdleSleepGuard idleSleepGuard; // nullable — idleSleepGuard.enabled: false
|
||||
private final ReplyInbox replyInbox;
|
||||
private final LeadChannelHandle leadMailbox; // nullable — coordinator: opt-in
|
||||
private final CompletionResolver completion;
|
||||
private final Injector injector;
|
||||
/**
|
||||
* Not final: {@code Fleetd.main}'s shutdown hook was registered <em>before</em> the Javalin
|
||||
* {@code FleetApp} was built and started — see {@link FleetdAssembly#assembleAndStart}'s javadoc
|
||||
* for why that order could not be preserved AND have this constructor take {@code app}. {@link
|
||||
* #attachApp} is called immediately after the real app is built, still before HTTP starts
|
||||
* listening, so this is set long before any test or caller could observe it unset.
|
||||
*/
|
||||
private Javalin app;
|
||||
|
||||
FleetdRuntime(FleetConfig cfg, SessionManager sessions, HerdrRouter router, StatusPoller poller,
|
||||
MessageService messages, ReplyPushLoop pushLoop, LeadHeartbeatLoop heartbeat,
|
||||
LeadCoordLoop leadCoordLoop, ScheduledExecutorService leadCoordScheduler,
|
||||
FleetHealthMonitor healthMonitor, ConfigWatcher configWatcher, FleetMcp mcp,
|
||||
SessionReaper reaper, IdleSleepGuard idleSleepGuard, ReplyInbox replyInbox,
|
||||
LeadChannelHandle leadMailbox, CompletionResolver completion, Injector injector) {
|
||||
this.cfg = cfg;
|
||||
this.sessions = sessions;
|
||||
this.router = router;
|
||||
this.poller = poller;
|
||||
this.messages = messages;
|
||||
this.pushLoop = pushLoop;
|
||||
this.heartbeat = heartbeat;
|
||||
this.leadCoordLoop = leadCoordLoop;
|
||||
this.leadCoordScheduler = leadCoordScheduler;
|
||||
this.healthMonitor = healthMonitor;
|
||||
this.configWatcher = configWatcher;
|
||||
this.mcp = mcp;
|
||||
this.reaper = reaper;
|
||||
this.idleSleepGuard = idleSleepGuard;
|
||||
this.replyInbox = replyInbox;
|
||||
this.leadMailbox = leadMailbox;
|
||||
this.completion = completion;
|
||||
this.injector = injector;
|
||||
}
|
||||
|
||||
/** See the {@link #app} field doc for why this is a late-bound setter rather than a constructor arg. */
|
||||
void attachApp(Javalin app) {
|
||||
this.app = app;
|
||||
}
|
||||
|
||||
// --- package-private accessors: the SAME instances this runtime owns, never a copy ---------
|
||||
|
||||
SessionManager sessions() { return sessions; }
|
||||
HerdrRouter router() { return router; }
|
||||
StatusPoller poller() { return poller; }
|
||||
MessageService messages() { return messages; }
|
||||
ReplyPushLoop pushLoop() { return pushLoop; }
|
||||
LeadHeartbeatLoop heartbeat() { return heartbeat; }
|
||||
LeadCoordLoop leadCoordLoop() { return leadCoordLoop; }
|
||||
FleetHealthMonitor healthMonitor() { return healthMonitor; }
|
||||
ConfigWatcher configWatcher() { return configWatcher; }
|
||||
FleetMcp mcp() { return mcp; }
|
||||
SessionReaper reaper() { return reaper; }
|
||||
IdleSleepGuard idleSleepGuard() { return idleSleepGuard; }
|
||||
ReplyInbox replyInbox() { return replyInbox; }
|
||||
LeadChannelHandle leadMailbox() { return leadMailbox; }
|
||||
CompletionResolver completion() { return completion; }
|
||||
Injector injector() { return injector; }
|
||||
Javalin app() { return app; }
|
||||
|
||||
/**
|
||||
* CB-303 part 3: the single ordered shutdown. Moved verbatim out of {@code Fleetd.main}'s
|
||||
* shutdown-hook {@code Thread} body (fleetd #612 Unit A) — drain sessions first while herdr is
|
||||
* still open (so releases reach the daemon), then stop poller/message/mcp/reaper, and close
|
||||
* herdr last, exactly as before. {@code Fleetd.main} never calls this directly; it hands the
|
||||
* reference to {@code ResourcePorts.addShutdownHook} the moment this runtime exists, the same
|
||||
* point it used to register the hook {@code Thread} itself.
|
||||
*/
|
||||
@Override
|
||||
public void close() {
|
||||
sessions.close(cfg.lifecycle() != null ? cfg.lifecycle().drainTimeoutSeconds() : null);
|
||||
poller.stop();
|
||||
messages.close();
|
||||
pushLoop.close();
|
||||
if (heartbeat != null) heartbeat.close(); // CB-551: stop the idle-lead heartbeat scheduler
|
||||
if (leadCoordLoop != null) leadCoordLoop.close(); // CB-637: stop delivering peer-lead messages
|
||||
if (leadCoordScheduler != null) leadCoordScheduler.shutdownNow();
|
||||
if (healthMonitor != null) healthMonitor.stop();
|
||||
if (configWatcher != null) configWatcher.stop(); // CB-559: stop polling the config file
|
||||
mcp.close();
|
||||
if (reaper != null) reaper.stop();
|
||||
// Idle-sleep guard: release unconditionally, even though sessions.close() above already
|
||||
// drained every session (and each release already drove the live count to 0, which
|
||||
// releases the guard's assertion on its own) — this is the backstop for a drain that was
|
||||
// itself interrupted or threw, so no caffeinate child ever outlives the daemon.
|
||||
if (idleSleepGuard != null) idleSleepGuard.close();
|
||||
// Release the broker connection last among message resources (no-op for the in-memory inbox).
|
||||
if (replyInbox instanceof AutoCloseable closeable) {
|
||||
try {
|
||||
closeable.close();
|
||||
} catch (Exception e) {
|
||||
log.debug("reply inbox close: {}", e.toString());
|
||||
}
|
||||
}
|
||||
// CB-637: the coordination connection goes with it — after the loop that reads it has
|
||||
// stopped, so no tick can be mid-ack against a closed channel.
|
||||
if (leadMailbox != null) {
|
||||
try {
|
||||
leadMailbox.close();
|
||||
} catch (Exception e) {
|
||||
log.debug("lead mailbox close: {}", e.toString());
|
||||
}
|
||||
}
|
||||
router.close();
|
||||
}
|
||||
}
|
||||
@@ -0,0 +1,80 @@
|
||||
package dev.ltms.fleet;
|
||||
|
||||
import dev.ltms.fleet.herdr.HerdrClient;
|
||||
import io.javalin.Javalin;
|
||||
|
||||
import java.nio.file.Path;
|
||||
import java.util.Map;
|
||||
import java.util.concurrent.ScheduledExecutorService;
|
||||
import java.util.function.LongSupplier;
|
||||
|
||||
/**
|
||||
* fleetd #612 Unit A: every boot-time side effect {@link FleetdAssembly#assembleAndStart} performs
|
||||
* that a real daemon must do for real, and a test must not — read the process environment, connect
|
||||
* a herdr client, open a broker (the reply inbox, the lead mailbox), read a clock, start a
|
||||
* background scheduler, register the JVM shutdown hook, and bind the HTTP server.
|
||||
*
|
||||
* <p>{@link #system()} is the one production implementation ({@code SystemResourcePorts}), wired
|
||||
* verbatim from what {@code Fleetd.main} used to call directly at each of these call sites. A test
|
||||
* builds its own implementation instead of receiving an inert default from this interface —
|
||||
* deliberately, there is no {@code ResourcePorts.none()}. fleetd #612's whole problem is a call
|
||||
* site quietly swapped for an inert variant that still compiles; adding one here, even for tests,
|
||||
* would hand a future edit to {@code FleetdAssembly} the exact compiling substitute this ticket
|
||||
* exists to rule out. A test that wants an inert resource writes its own fake and owns that
|
||||
* decision explicitly.
|
||||
*/
|
||||
public interface ResourcePorts {
|
||||
|
||||
/** The process environment. Production: {@link System#getenv()}. */
|
||||
Map<String, String> environment();
|
||||
|
||||
/**
|
||||
* Connect a herdr client bound to {@code socketPath}. Production returns a real
|
||||
* {@code UnixSocketHerdrClient} — connection-per-call, so this itself never touches the socket.
|
||||
*/
|
||||
HerdrClient connectHerdr(Path socketPath);
|
||||
|
||||
/** The reply-inbox AMQP opener (CB-307). Production: {@link Fleetd#replyInboxOpener()}. */
|
||||
Fleetd.AmqpOpener replyInboxOpener();
|
||||
|
||||
/** The lead-mailbox AMQP opener (CB-637). Production: {@link Fleetd#leadMailboxOpener()}. */
|
||||
Fleetd.LeadMailboxOpener leadMailboxOpener();
|
||||
|
||||
/** A monotonic elapsed-time clock. Production: {@link System#nanoTime()}. */
|
||||
LongSupplier nanoClock();
|
||||
|
||||
/**
|
||||
* fleetd #629: the per-poll wait {@code FleetdAssembly#assembleAndStart} passes to {@code
|
||||
* Fleetd#awaitHerdr} while polling for herdr's socket. Production: {@link
|
||||
* Fleetd#sleepHerdrPoll()} — a real {@code Thread.sleep}. {@link #nanoClock()} alone is not
|
||||
* enough to make {@code awaitHerdr}'s deadline controllable: the old call site passed {@code
|
||||
* Fleetd::sleepHerdrPoll} directly, hardcoded, so a test that injected a fake clock still had
|
||||
* to wait out the real sleep between each poll to ever reach the deadline — the clock looked
|
||||
* injected and was not actually controllable. A test supplies a no-op that advances its own
|
||||
* injected {@link #nanoClock()} instead, so the deadline becomes reachable without any real
|
||||
* wall-clock time passing.
|
||||
*/
|
||||
Runnable herdrPollWait();
|
||||
|
||||
/**
|
||||
* A wall-clock reading, in nanoseconds. Production: {@code System.currentTimeMillis()}
|
||||
* converted to nanoseconds. Kept separate from {@link #nanoClock()} because {@link
|
||||
* dev.ltms.fleet.health.FleetHealthMonitor} needs both — one monotonic clock for elapsed-time
|
||||
* decisions, one wall clock to detect and correct for a macOS sleep freezing the monotonic one.
|
||||
*/
|
||||
LongSupplier wallClockNanos();
|
||||
|
||||
/** A dedicated single-thread scheduler; production names its (virtual) thread {@code purpose}. */
|
||||
ScheduledExecutorService newScheduler(String purpose);
|
||||
|
||||
/** Register a JVM shutdown hook that runs {@code hook} on JVM exit. */
|
||||
void addShutdownHook(Runnable hook);
|
||||
|
||||
/** Bind and start the HTTP server. */
|
||||
void startHttp(Javalin app, String host, int port);
|
||||
|
||||
/** The real production ports: a live herdr socket, a live broker, real threads, a real bind. */
|
||||
static ResourcePorts system() {
|
||||
return new SystemResourcePorts();
|
||||
}
|
||||
}
|
||||
@@ -0,0 +1,71 @@
|
||||
package dev.ltms.fleet;
|
||||
|
||||
import com.fasterxml.jackson.databind.ObjectMapper;
|
||||
import dev.ltms.fleet.herdr.HerdrClient;
|
||||
import dev.ltms.fleet.herdr.UnixSocketHerdrClient;
|
||||
import io.javalin.Javalin;
|
||||
|
||||
import java.nio.file.Path;
|
||||
import java.util.Map;
|
||||
import java.util.concurrent.Executors;
|
||||
import java.util.concurrent.ScheduledExecutorService;
|
||||
import java.util.concurrent.TimeUnit;
|
||||
import java.util.function.LongSupplier;
|
||||
|
||||
/**
|
||||
* fleetd #612 Unit A: the one production {@link ResourcePorts} — every method here is the exact
|
||||
* call {@code Fleetd.main} used to make directly at each of these sites before this ticket.
|
||||
* Package-private: obtained only through {@link ResourcePorts#system()}.
|
||||
*/
|
||||
final class SystemResourcePorts implements ResourcePorts {
|
||||
|
||||
@Override
|
||||
public Map<String, String> environment() {
|
||||
return System.getenv();
|
||||
}
|
||||
|
||||
@Override
|
||||
public HerdrClient connectHerdr(Path socketPath) {
|
||||
return UnixSocketHerdrClient.connect(socketPath, new ObjectMapper());
|
||||
}
|
||||
|
||||
@Override
|
||||
public Fleetd.AmqpOpener replyInboxOpener() {
|
||||
return Fleetd.replyInboxOpener();
|
||||
}
|
||||
|
||||
@Override
|
||||
public Fleetd.LeadMailboxOpener leadMailboxOpener() {
|
||||
return Fleetd.leadMailboxOpener();
|
||||
}
|
||||
|
||||
@Override
|
||||
public LongSupplier nanoClock() {
|
||||
return System::nanoTime;
|
||||
}
|
||||
|
||||
@Override
|
||||
public Runnable herdrPollWait() {
|
||||
return Fleetd::sleepHerdrPoll;
|
||||
}
|
||||
|
||||
@Override
|
||||
public LongSupplier wallClockNanos() {
|
||||
return () -> TimeUnit.MILLISECONDS.toNanos(System.currentTimeMillis());
|
||||
}
|
||||
|
||||
@Override
|
||||
public ScheduledExecutorService newScheduler(String purpose) {
|
||||
return Executors.newSingleThreadScheduledExecutor(r -> Thread.ofVirtual().name(purpose).unstarted(r));
|
||||
}
|
||||
|
||||
@Override
|
||||
public void addShutdownHook(Runnable hook) {
|
||||
Runtime.getRuntime().addShutdownHook(new Thread(hook));
|
||||
}
|
||||
|
||||
@Override
|
||||
public void startHttp(Javalin app, String host, int port) {
|
||||
app.start(host, port);
|
||||
}
|
||||
}
|
||||
@@ -20,16 +20,22 @@ public final class Authz {
|
||||
SPAWN,
|
||||
/** Tear a worker peer down. */
|
||||
STOP,
|
||||
/** Deliver a turn to a session (or answer a worker's question). */
|
||||
/** Deliver a turn to a local session, addressed by {@code sessionId}. */
|
||||
SEND,
|
||||
/** Resolve a worker's blocked question and resume its turn, addressed by {@code turnId}. */
|
||||
ANSWER,
|
||||
/** Address a peer lead on another daemon over the coordination broker, by {@code coordId}. */
|
||||
COORD_SEND,
|
||||
/** A worker's terminal reply for its own turn. */
|
||||
REPLY,
|
||||
/** A worker's mid-turn question to the primary. */
|
||||
ASK,
|
||||
/** Collect held replies from a session's inbox. */
|
||||
DRAIN,
|
||||
/** Read-only observation: status, roster, profiles, task polling. */
|
||||
/** Read-only roster, profile, and identity observation: no ticket, task, or turn state. */
|
||||
READ,
|
||||
/** Poll a ticket, or read a session's status. */
|
||||
TASK_READ,
|
||||
/**
|
||||
* Read (never ack) this daemon's own held lead-to-lead coordination mail (fleetd #421).
|
||||
*
|
||||
@@ -71,11 +77,19 @@ public final class Authz {
|
||||
// escalating into the orchestrator role.
|
||||
case SPAWN, STOP, DRAIN, HANDOVER -> caller.isPrimary();
|
||||
|
||||
// Delivering a turn is open to the primary and the architect: an architect delegates
|
||||
// to workers (that is the role's point) but still has no lifecycle rights. A worker is
|
||||
// excluded — sending would be it escalating.
|
||||
// Delivering a turn to a local session is open to the primary and the architect: an
|
||||
// architect delegates to workers (that is the role's point) but still has no lifecycle
|
||||
// rights. A worker is excluded — sending would be it escalating.
|
||||
case SEND -> caller.isPrimary() || caller.isArchitect();
|
||||
|
||||
// Same grant as SEND. Resolving a worker's blocked question is part of delegating to
|
||||
// it, not a separate capability.
|
||||
case ANSWER -> caller.isPrimary() || caller.isArchitect();
|
||||
|
||||
// Same grant as SEND. This leaves the daemon over the coordination broker rather than
|
||||
// addressing a local session, but the caller who may do one may do the other.
|
||||
case COORD_SEND -> caller.isPrimary() || caller.isArchitect();
|
||||
|
||||
// The load-bearing rule: a caller acts only as the pane it occupies. CB-532 widened who
|
||||
// that can be — a lead answering another lead is replying for its OWN terminal, which
|
||||
// this already permits — while the rule itself is unchanged, and is what stops anyone
|
||||
@@ -84,10 +98,17 @@ public final class Authz {
|
||||
// unnamed primary (token/loopback, no pane) owns nothing and is still excluded.
|
||||
case REPLY, ASK -> caller.ownsSession(targetSession);
|
||||
|
||||
// Observation is open to every authenticated role: a worker legitimately polls its own
|
||||
// status, and the roster carries no secrets.
|
||||
// READ is roster, profile, and identity observation — fleet_list, fleet_profiles, and
|
||||
// fleet_whoami — and carries no secrets: no ticket reply, no pending question, and no
|
||||
// other session's turn state. Those live under TASK_READ. METRICS is the separate
|
||||
// Prometheus scrape. Both stay open to every authenticated role.
|
||||
case READ, METRICS -> caller.isPrimary() || caller.isWorker() || caller.isArchitect();
|
||||
|
||||
// Ticket polling and session status, open to every authenticated role the same as READ.
|
||||
// Unlike READ, a holder may poll a ticket it did not create, or read another session's
|
||||
// pending question and the turnId that answers it.
|
||||
case TASK_READ -> caller.isPrimary() || caller.isWorker() || caller.isArchitect();
|
||||
|
||||
// fleetd #421: reading held lead-to-lead mail is the primary's alone. An architect
|
||||
// holds READ today (CB-548), so "not primary" must mean not-architect here too — this
|
||||
// is coordination between leads, not observation of the roster.
|
||||
|
||||
@@ -798,6 +798,25 @@ public record FleetConfig(
|
||||
return isSubscription() ? SUBSCRIPTION_CREDENTIAL_ID : profile;
|
||||
}
|
||||
|
||||
/**
|
||||
* The auto-compaction window a launched Claude Code session actually runs on: {@code env:
|
||||
* CLAUDE_CODE_AUTO_COMPACT_WINDOW} when it parses as an integer, since that environment
|
||||
* variable wins over the {@code --autocompact} flag {@link #autoCompactWindow} produces (see
|
||||
* {@code ClaudeCodeArguments}); {@link #autoCompactWindow} otherwise. {@code null} when
|
||||
* neither resolves to a usable number.
|
||||
*/
|
||||
public Integer effectiveAutoCompactWindow() {
|
||||
String envValue = env.get(CLAUDE_CODE_AUTO_COMPACT_WINDOW_ENV);
|
||||
if (envValue == null) {
|
||||
return autoCompactWindow;
|
||||
}
|
||||
try {
|
||||
return Integer.valueOf(envValue.trim());
|
||||
} catch (NumberFormatException e) {
|
||||
return autoCompactWindow;
|
||||
}
|
||||
}
|
||||
|
||||
/** True when this profile's workers are granted a forge token to open their own PR (CB-302). */
|
||||
public boolean hasGitToken() {
|
||||
return gitTokenEnv != null && !gitTokenEnv.isBlank();
|
||||
@@ -1110,11 +1129,8 @@ public record FleetConfig(
|
||||
* {@code tab} can never be discovered, launched or not
|
||||
* @param instances how many of this lead should be live (default 1). The daemon
|
||||
* launches only the shortfall, so a restart adopts rather than doubles
|
||||
* @param tabPrefix no longer used to find a lead's tab — {@code tab} is matched
|
||||
* exactly. Its only remaining job is the startup collision guard
|
||||
* ({@link #validateLeadTabPrefixes()}), which still uses it to refuse
|
||||
* a worker {@code tabLabel} template that could be misread as a lead.
|
||||
* Default {@code "lead:"}
|
||||
* @param tabPrefix lead-tab naming convention checked against member labels. Lead
|
||||
* identity uses {@code tab}. Default {@code "lead:"}
|
||||
* @param scanIntervalSeconds how long a tab scan is cached before herdr is asked again; also the
|
||||
* worst case before a newly-labelled tab is recognised. Default 10
|
||||
* @param kind which agent runs there ({@code claude}, {@code opencode}, …)
|
||||
@@ -1217,10 +1233,7 @@ public record FleetConfig(
|
||||
String tabLabel) {
|
||||
|
||||
/**
|
||||
* Role first, so the tab bar reads as the fleet and so the label shares a namespace with a
|
||||
* lead's {@code tabPrefix}. Because {@code {role}} comes from a closed enum, a generated
|
||||
* member label can never begin with {@code "lead:"} — the clash that
|
||||
* {@link #validateLeadTabPrefixes()} used to have to check for is unrepresentable here.
|
||||
* Role first, so the tab bar identifies the member's fleet role.
|
||||
*/
|
||||
public static final String DEFAULT_TAB_LABEL = "{role}: {profile} #{n}";
|
||||
|
||||
@@ -1380,11 +1393,13 @@ public record FleetConfig(
|
||||
* called FROM the calling lead's own turn, so its pane is still {@code WORKING} the instant
|
||||
* {@code confirm()} validates every gate and schedules the roll. {@code
|
||||
* dev.ltms.fleet.lead.LeadRollover}'s deferred continuation waits up to this many seconds for
|
||||
* that SAME pane to report an injectable state again — i.e. for the calling turn to actually
|
||||
* end — before it sends {@code /clear} at all. If that wait times out, no {@code /clear} is
|
||||
* that SAME pane to report {@code IDLE} or {@code DONE} — i.e. for the calling turn to actually
|
||||
* end — before it sends {@code /clear} at all. {@code BLOCKED} does not count: that is a live
|
||||
* turn merely paused, not one that has finished. If that wait times out, no {@code /clear} is
|
||||
* ever sent: a lead that never goes idle is still doing real work, and clearing it would
|
||||
* destroy live context. This is a separate wait from {@code clearSettleSeconds} below, which
|
||||
* bounds the SECOND wait, for the pane to re-settle AFTER {@code /clear} has already gone out.
|
||||
* bounds the SECOND wait, for the pane to reach {@code IDLE} or {@code DONE} again AFTER
|
||||
* {@code /clear} has already gone out.
|
||||
*
|
||||
* @param handoverPath required when this block is present — where the handover file a fresh
|
||||
* lead session reads must live. There is no sane non-null default for an
|
||||
@@ -1403,12 +1418,13 @@ public record FleetConfig(
|
||||
* @param maxDocAgeSeconds default 3600 — refuse a handover file whose modified time is older
|
||||
* than this many seconds, so a stale leftover from an earlier rollover
|
||||
* attempt can never be mistaken for a fresh one.
|
||||
* @param turnSettleSeconds default 20 — bound on how long the deferred roll waits for the
|
||||
* CALLING lead's own turn to end (its pane to report injectable again)
|
||||
* before sending {@code /clear} at all. See the paragraph above.
|
||||
* @param turnSettleSeconds default 300 — bound on how long the deferred roll waits for the
|
||||
* CALLING lead's own turn to end (its pane to report {@code IDLE} or
|
||||
* {@code DONE}) before sending {@code /clear} at all. See the paragraph
|
||||
* above.
|
||||
* @param clearSettleSeconds default 20 — bound on how long to wait for the lead's pane to
|
||||
* report an injectable state again after {@code /clear} before giving up. A
|
||||
* roll that times out here never sends {@code bootstrapText}.
|
||||
* report {@code IDLE} or {@code DONE} again after {@code /clear} before
|
||||
* giving up. A roll that times out here never sends {@code bootstrapText}.
|
||||
* @param bootstrapText default a sentence naming the RESOLVED handover path — sent to the
|
||||
* lead's pane once it settles after {@code /clear}, telling the fresh
|
||||
* session where to read the handover and carry on. Left {@code null} here
|
||||
@@ -1425,7 +1441,7 @@ public record FleetConfig(
|
||||
public LeadRollover {
|
||||
requireOperatorConfirm = requireOperatorConfirm == null || requireOperatorConfirm;
|
||||
maxDocAgeSeconds = (maxDocAgeSeconds == null || maxDocAgeSeconds <= 0) ? 3600 : maxDocAgeSeconds;
|
||||
turnSettleSeconds = (turnSettleSeconds == null || turnSettleSeconds <= 0) ? 20 : turnSettleSeconds;
|
||||
turnSettleSeconds = (turnSettleSeconds == null || turnSettleSeconds <= 0) ? 300 : turnSettleSeconds;
|
||||
clearSettleSeconds = (clearSettleSeconds == null || clearSettleSeconds <= 0) ? 20 : clearSettleSeconds;
|
||||
bootstrapText = (bootstrapText == null || bootstrapText.isBlank()) ? null : bootstrapText;
|
||||
}
|
||||
@@ -2184,6 +2200,12 @@ public record FleetConfig(
|
||||
static final int AUTO_COMPACT_WINDOW_MIN = 100_000;
|
||||
/** Highest {@code autoCompactWindow} Claude Code's {@code --autocompact <tokens>} flag accepts. */
|
||||
static final int AUTO_COMPACT_WINDOW_MAX = 1_000_000;
|
||||
/**
|
||||
* The {@code env:} key a launched Claude Code session reads for its auto-compaction window,
|
||||
* ahead of the {@code --autocompact} launch flag {@code autoCompactWindow} produces (see
|
||||
* {@link Profile#effectiveAutoCompactWindow()}).
|
||||
*/
|
||||
static final String CLAUDE_CODE_AUTO_COMPACT_WINDOW_ENV = "CLAUDE_CODE_AUTO_COMPACT_WINDOW";
|
||||
|
||||
/**
|
||||
* Reject a profile whose {@code autoCompactWindow:} is set but outside the token band Claude
|
||||
@@ -2265,13 +2287,13 @@ public record FleetConfig(
|
||||
if (!(entry.getValue() instanceof Map<?, ?> profile)
|
||||
|| !(profile.get("autoCompactWindow") instanceof Number window)
|
||||
|| !(profile.get("env") instanceof Map<?, ?> env)
|
||||
|| !env.containsKey("CLAUDE_CODE_AUTO_COMPACT_WINDOW")) {
|
||||
|| !env.containsKey(CLAUDE_CODE_AUTO_COMPACT_WINDOW_ENV)) {
|
||||
continue;
|
||||
}
|
||||
Object kind = profile.get("kind");
|
||||
boolean claudeCode = kind == null || String.valueOf(kind).isBlank()
|
||||
|| Profile.KIND_CLAUDE_CODE.equalsIgnoreCase(String.valueOf(kind));
|
||||
Object envValue = env.get("CLAUDE_CODE_AUTO_COMPACT_WINDOW");
|
||||
Object envValue = env.get(CLAUDE_CODE_AUTO_COMPACT_WINDOW_ENV);
|
||||
if (claudeCode && !String.valueOf(window).equals(String.valueOf(envValue))) {
|
||||
String name = String.valueOf(entry.getKey());
|
||||
names.add(name);
|
||||
@@ -2657,27 +2679,14 @@ public record FleetConfig(
|
||||
}
|
||||
|
||||
/**
|
||||
* Reject a lead-scan convention that a worker tab would also satisfy (CB-531).
|
||||
* Reject a member tab-label template that could render as a configured lead tab or match a
|
||||
* lead-tab naming convention, and reject two {@code fleet.leaders} entries that share one exact
|
||||
* tab.
|
||||
*
|
||||
* <p>The scan reads a tab label and concludes "a lead lives here". fleetd also <em>writes</em>
|
||||
* tab labels — every member gets one rendered into its tab. Choose a lead {@code tabPrefix} that
|
||||
* a member template matches and the daemon starts labelling its own members as leads, promoting
|
||||
* the entire fleet to {@link dev.ltms.fleet.auth.Role#PRIMARY} with no message and no diff.
|
||||
* The member-space exclusion in {@link dev.ltms.fleet.herdr.LeadTabScanner} already blocks the
|
||||
* realistic path, but defence that depends on one workspace label holding is not defence enough
|
||||
* for a privilege boundary.
|
||||
*
|
||||
* <p>CB-557 shrank this check rather than removing it. The default template is
|
||||
* {@code "{role}: {profile} #{n}"} and {@code {role}} comes from a closed enum, so a
|
||||
* <em>generated</em> label can no longer collide by construction. What remains checkable is what
|
||||
* an operator still writes by hand: the {@code fleet.tabLabel} template and any per-profile
|
||||
* {@code tabLabel} override.
|
||||
*
|
||||
* <p>Fatal rather than a warning, unlike {@link #warnUnknownTopLevelKeys}: an unknown key means
|
||||
* a feature does nothing, while this means a feature does the opposite of what it says.
|
||||
*
|
||||
* @throws IllegalStateException when the fleet template or any profile's {@code tabLabel}
|
||||
* override starts with a configured lead prefix
|
||||
* @throws IllegalStateException when the fleet template or a profile {@code tabLabel} override
|
||||
* can render as a configured lead tab or match a lead-tab prefix,
|
||||
* or when two {@code fleet.leaders} entries carry the same exact
|
||||
* {@code tab} (case-insensitively)
|
||||
*/
|
||||
public void validateLeadTabPrefixes() {
|
||||
if (fleet == null || fleet.leaders().isEmpty()) {
|
||||
@@ -2688,28 +2697,129 @@ public record FleetConfig(
|
||||
if (leader == null) {
|
||||
return;
|
||||
}
|
||||
String tab = leader.tab();
|
||||
String prefix = leader.tabPrefix();
|
||||
// The fleet-wide template is checked once per prefix: it labels every member that has no
|
||||
// override, so one bad template promotes the entire fleet, not one profile.
|
||||
if (startsWithIgnoreCase(fleet.tabLabel(), prefix)) {
|
||||
if (templateCanRenderAs(fleet.tabLabel(), tab)) {
|
||||
bad.add("fleet.tabLabel=\"" + fleet.tabLabel() + "\" can render as the tab of "
|
||||
+ "lead '" + leadName + "' (\"" + tab + "\")");
|
||||
} else if (startsWithIgnoreCase(fleet.tabLabel(), prefix)) {
|
||||
bad.add("fleet.tabLabel=\"" + fleet.tabLabel() + "\" starts with the tabPrefix of "
|
||||
+ "lead '" + leadName + "' (\"" + prefix + "\")");
|
||||
}
|
||||
profiles().entrySet().stream()
|
||||
.filter(e -> startsWithIgnoreCase(e.getValue().tabLabel(), prefix))
|
||||
.map(Map.Entry::getKey)
|
||||
.sorted()
|
||||
.forEach(p -> bad.add("profile '" + p + "' overrides tabLabel with \""
|
||||
+ profiles().get(p).tabLabel() + "\", which starts with the tabPrefix of "
|
||||
+ "lead '" + leadName + "' (\"" + prefix + "\")"));
|
||||
.forEach(p -> {
|
||||
String label = profiles().get(p).tabLabel();
|
||||
if (templateCanRenderAs(label, tab)) {
|
||||
bad.add("profile '" + p + "' overrides tabLabel with \"" + label
|
||||
+ "\", which can render as the tab of lead '" + leadName
|
||||
+ "' (\"" + tab + "\")");
|
||||
} else if (startsWithIgnoreCase(label, prefix)) {
|
||||
bad.add("profile '" + p + "' overrides tabLabel with \"" + label
|
||||
+ "\", which starts with the tabPrefix of lead '" + leadName
|
||||
+ "' (\"" + prefix + "\")");
|
||||
}
|
||||
});
|
||||
});
|
||||
if (!bad.isEmpty()) {
|
||||
throw new IllegalStateException("refusing to start: " + String.join("; ", bad)
|
||||
+ ". Every member labelled that way would be read back as a lead and granted "
|
||||
+ "spawn/stop/send on the whole fleet. Change one of the two so member tabs "
|
||||
+ "and lead tabs cannot be confused.");
|
||||
}
|
||||
|
||||
List<String> collisions = new ArrayList<>();
|
||||
List<String> leadNames = fleet.leaders().keySet().stream().sorted().toList();
|
||||
for (int i = 0; i < leadNames.size(); i++) {
|
||||
String nameA = leadNames.get(i);
|
||||
Leader a = fleet.leaders().get(nameA);
|
||||
if (a == null || a.tab() == null || a.tab().isBlank()) {
|
||||
continue;
|
||||
}
|
||||
for (int j = i + 1; j < leadNames.size(); j++) {
|
||||
String nameB = leadNames.get(j);
|
||||
Leader b = fleet.leaders().get(nameB);
|
||||
if (b == null || b.tab() == null || b.tab().isBlank()) {
|
||||
continue;
|
||||
}
|
||||
if (a.tab().equalsIgnoreCase(b.tab())) {
|
||||
collisions.add("lead '" + nameA + "' and lead '" + nameB + "' both use tab \""
|
||||
+ a.tab() + "\"");
|
||||
}
|
||||
}
|
||||
}
|
||||
if (collisions.isEmpty()) {
|
||||
return;
|
||||
}
|
||||
throw new IllegalStateException("refusing to start: " + String.join("; ", collisions)
|
||||
+ ". Tab identity is matched exactly, so only one of two leads sharing a tab can "
|
||||
+ "ever be found — the other is silently unreachable. Give each lead its own "
|
||||
+ "exact tab.");
|
||||
}
|
||||
|
||||
private static boolean templateCanRenderAs(String template, String tab) {
|
||||
if (template == null || template.isBlank() || tab == null || tab.isBlank()) {
|
||||
return false;
|
||||
}
|
||||
var placeholders = Pattern.compile("\\{(?:role|profile|model|n)}").matcher(template);
|
||||
StringBuilder expression = new StringBuilder("^");
|
||||
int literalStart = 0;
|
||||
while (placeholders.find()) {
|
||||
expression.append(Pattern.quote(template.substring(literalStart, placeholders.start())));
|
||||
expression.append(".*");
|
||||
literalStart = placeholders.end();
|
||||
}
|
||||
expression.append(Pattern.quote(template.substring(literalStart))).append("$");
|
||||
return Pattern.compile(expression.toString(), Pattern.CASE_INSENSITIVE).matcher(tab).matches();
|
||||
}
|
||||
|
||||
/** Case-insensitive prefix test that tolerates a null or blank label. */
|
||||
private static boolean startsWithIgnoreCase(String label, String prefix) {
|
||||
if (label == null || prefix == null || prefix.isBlank()) {
|
||||
return false;
|
||||
}
|
||||
String stripped = label.strip();
|
||||
return stripped.regionMatches(true, 0, prefix, 0, prefix.length());
|
||||
}
|
||||
|
||||
/**
|
||||
* Reject a profile that places its members by {@code "pane"} while any {@code fleet.leaders}
|
||||
* entry names a {@code tab}. A pane-placed member lands inside the focused tab rather than its
|
||||
* own, so it can land inside a lead's own labelled tab. {@link
|
||||
* dev.ltms.fleet.herdr.LeadTabScanner} identifies a lead purely by that tab's label — it does
|
||||
* not exclude the member space — so a member that ends up there would be read back as the lead
|
||||
* and granted spawn/stop/send on the whole fleet.
|
||||
*
|
||||
* <p>Only a leader with a non-blank {@code tab} is in scope: one with no {@code tab} feeds
|
||||
* nothing into {@link dev.ltms.fleet.herdr.LeadTabScanner}, so it creates no hazard here.
|
||||
*
|
||||
* @throws IllegalStateException when any {@code profiles:} entry is pane-placed while any
|
||||
* {@code fleet.leaders} entry names a non-blank {@code tab}
|
||||
*/
|
||||
public void validatePanePlacementAgainstLeadTabs() {
|
||||
if (fleet == null || fleet.leaders().isEmpty()) {
|
||||
return;
|
||||
}
|
||||
boolean anyLeaderHasTab = fleet.leaders().values().stream()
|
||||
.anyMatch(leader -> leader != null && leader.tab() != null && !leader.tab().isBlank());
|
||||
if (!anyLeaderHasTab) {
|
||||
return;
|
||||
}
|
||||
List<String> bad = new ArrayList<>();
|
||||
profiles().entrySet().stream()
|
||||
.filter(e -> !e.getValue().tabPlacement())
|
||||
.map(Map.Entry::getKey)
|
||||
.sorted()
|
||||
.forEach(bad::add);
|
||||
if (bad.isEmpty()) {
|
||||
return;
|
||||
}
|
||||
throw new IllegalStateException("refusing to start: " + String.join("; ", bad)
|
||||
+ ". Every member labelled that way would be read back as a lead and granted "
|
||||
+ "spawn/stop/send on the whole fleet. Change one of the two so member tabs and "
|
||||
+ "lead tabs cannot be confused.");
|
||||
throw new IllegalStateException("refusing to start: profile(s) " + bad
|
||||
+ " use placement: pane while fleet.leaders names a tab. A pane-placed member can "
|
||||
+ "land inside a lead's labelled tab and be read back as the lead, granted "
|
||||
+ "spawn/stop/send on the whole fleet. Set placement: tab for each named profile, "
|
||||
+ "or remove the tab from every fleet.leaders entry.");
|
||||
}
|
||||
|
||||
/**
|
||||
@@ -2732,15 +2842,6 @@ public record FleetConfig(
|
||||
}
|
||||
}
|
||||
|
||||
/** Case-insensitive prefix test that tolerates a null/blank label. */
|
||||
private static boolean startsWithIgnoreCase(String label, String prefix) {
|
||||
if (label == null || prefix == null || prefix.isBlank()) {
|
||||
return false;
|
||||
}
|
||||
String stripped = label.strip();
|
||||
return stripped.regionMatches(true, 0, prefix, 0, prefix.length());
|
||||
}
|
||||
|
||||
/**
|
||||
* Reject a subscription profile whose {@code env:} block tries to reseat the Anthropic binding
|
||||
* (CB-542).
|
||||
@@ -2912,11 +3013,11 @@ public record FleetConfig(
|
||||
* Runs every validator this class declares — found by reflection, not by name.
|
||||
*
|
||||
* <p>fleetd ticket "central allow-list of usable models", follow-up: mutation testing found
|
||||
* that although each of the six validators above was well pinned on its own, nothing proved
|
||||
* that although each validator above was well pinned on its own, nothing proved
|
||||
* either real caller ({@code Fleetd.main} and {@link ConfigRef#reload()}) still
|
||||
* invoked it — deleting a call site left the full suite green. The fix is not a seventh test
|
||||
* per caller; a hand-maintained list of six names here would have the exact same defect its
|
||||
* own javadoc would warn against: the seventh validator someone adds next month has no reason
|
||||
* invoked it — deleting a call site left the full suite green. The fix is not one more test
|
||||
* per caller; a hand-maintained list of names here would have the exact same defect its
|
||||
* own javadoc would warn against: the next validator someone adds has no reason
|
||||
* to be added to it. So this method does not name any validator. It sweeps {@link
|
||||
* #getClass()}'s own public, no-argument, {@code void} methods whose name starts with {@code
|
||||
* "validate"} (excluding itself) and invokes every one it finds, via {@link
|
||||
@@ -2925,7 +3026,7 @@ public record FleetConfig(
|
||||
* which it silently never runs.
|
||||
*
|
||||
* <p>{@code Fleetd.main} and {@link ConfigRef#reload()} each call this one method instead of
|
||||
* the six individually — see the comments at those two call sites for why
|
||||
* each validator individually — see the comments at those two call sites for why
|
||||
* each must run it.
|
||||
*
|
||||
* <p>Methods run in a fixed (alphabetical) order, so a config with more than one violation
|
||||
@@ -2942,9 +3043,9 @@ public record FleetConfig(
|
||||
/**
|
||||
* The reflective sweep behind {@link #validateAll()}, kept as its own method — taking any
|
||||
* {@code target}, not just {@code this} — so a test can prove the MECHANISM is generic (it
|
||||
* would sweep a seventh {@code validateXxx()} method added to any class, not just something
|
||||
* special-cased to today's six on {@link FleetConfig}) without needing to add a real, unwanted
|
||||
* seventh validator to this class just to exercise that claim. See {@code
|
||||
* would sweep any new {@code validateXxx()} method added to any class, not just something
|
||||
* special-cased to the set {@link FleetConfig} declares today) without needing to add a real,
|
||||
* unwanted extra validator to this class just to exercise that claim. See {@code
|
||||
* FleetConfigValidateAllTest} for that proof.
|
||||
*
|
||||
* @param target an object whose public, no-argument, {@code void} methods named {@code
|
||||
|
||||
@@ -37,8 +37,11 @@ import java.util.function.Supplier;
|
||||
* <p><strong>Direction of trust.</strong> The label names the lead; it never <em>grants</em>
|
||||
* anything a pane could take for itself. Three properties keep that honest:
|
||||
* <ol>
|
||||
* <li>Worker spaces are excluded wholesale ({@code excludedWorkspaceLabels}), so a worker cannot
|
||||
* become a lead by being placed — as a split, say — inside a matching tab.</li>
|
||||
* <li>{@code excludedWorkspaceLabels} can filter a workspace out of the scan, but this class does
|
||||
* not by itself stop a worker from landing inside a matching tab — a caller may pass an empty
|
||||
* set, and the daemon does. The guard against that is {@code
|
||||
* FleetConfig.validatePanePlacementAgainstLeadTabs}: it refuses, at startup, any profile that
|
||||
* places members by pane while a lead names a tab.</li>
|
||||
* <li>A worker cannot rename a tab: {@code tab.rename} is reachable only through
|
||||
* {@link WorkspaceControl}, which no {@code fleet_*} tool exposes. The label is writable by
|
||||
* the human at the terminal and by nobody the bridge is defending against.</li>
|
||||
|
||||
@@ -68,8 +68,10 @@ import java.util.function.LongSupplier;
|
||||
* (never the whole 52 MB a long-lived transcript reaches on the host this was measured on), and
|
||||
* {@link #DEFAULT_CACHE_TTL_MILLIS} bounds how often that bounded read actually happens — a burst
|
||||
* of {@code fleet_list} calls inside one TTL window reads the file once. One instance's cache is
|
||||
* keyed by {@code (configDir, sessionId)}, so it is safe to share across every lead a single
|
||||
* {@code fleet_list} call reports on.
|
||||
* keyed by {@code (configDir, sessionId, highThreshold)}, so it is safe to share across every lead
|
||||
* a single {@code fleet_list} call reports on, and a call that resolves a different effective
|
||||
* window for the same lead never reads back a state computed against the other window's
|
||||
* threshold.
|
||||
*/
|
||||
public final class LeadContextGauge {
|
||||
|
||||
@@ -97,14 +99,19 @@ public final class LeadContextGauge {
|
||||
static final long DEFAULT_CACHE_TTL_MILLIS = 5_000;
|
||||
|
||||
/**
|
||||
* Live tokens at or above this count report {@link State#HIGH}. On the host this was measured
|
||||
* on, auto-compaction actually fires around 267,000–270,000 tokens, but the point of a HIGH
|
||||
* state is to warn before that happens, not at it — 200,000 is the standard Claude context
|
||||
* window size and a sensible built-in default: no config key is required to pick it, and a
|
||||
* lead crossing it is already deep enough into its window that a compaction is foreseeable.
|
||||
* Fallback HIGH threshold used when a caller resolves no effective auto-compact window for the
|
||||
* lead being read (see {@link #read(String, String, String, Long)}) — the built-in default so
|
||||
* no config key is required to get a warning at all.
|
||||
*/
|
||||
static final long HIGH_THRESHOLD_TOKENS = 200_000;
|
||||
|
||||
/**
|
||||
* The fraction of a resolved effective auto-compact window that HIGH warns at, so the warning
|
||||
* margin scales with the window instead of only ever meaning something against the fixed
|
||||
* {@link #HIGH_THRESHOLD_TOKENS} fallback.
|
||||
*/
|
||||
static final double HIGH_THRESHOLD_FRACTION = 2.0 / 3.0;
|
||||
|
||||
/** The only peer kind this reader understands ({@code Agent.agentType()}'s wire value). */
|
||||
private static final String CLAUDE_AGENT_TYPE = "claude";
|
||||
|
||||
@@ -165,8 +172,13 @@ public final class LeadContextGauge {
|
||||
* {@code "claude"} (including {@code null}, meaning undetected) reports
|
||||
* {@link State#UNKNOWN} — this reader only understands Claude Code's own
|
||||
* transcript format
|
||||
* @param effectiveWindowTokens the caller's resolved effective auto-compact window for this
|
||||
* lead's own profile, or {@code null} when it cannot be resolved.
|
||||
* HIGH fires at {@link #HIGH_THRESHOLD_FRACTION} of this value;
|
||||
* {@code null} (or a non-positive value) falls back to the fixed
|
||||
* {@link #HIGH_THRESHOLD_TOKENS}
|
||||
*/
|
||||
public Reading read(String configDir, String sessionId, String agentType) {
|
||||
public Reading read(String configDir, String sessionId, String agentType, Long effectiveWindowTokens) {
|
||||
if (sessionId == null || sessionId.isBlank()) {
|
||||
return Reading.unknown();
|
||||
}
|
||||
@@ -176,18 +188,27 @@ public final class LeadContextGauge {
|
||||
String base = (configDir == null || configDir.isBlank())
|
||||
? System.getProperty("user.home") + "/.claude"
|
||||
: configDir;
|
||||
String cacheKey = base + '\u0000' + sessionId;
|
||||
long highThreshold = highThreshold(effectiveWindowTokens);
|
||||
String cacheKey = base + '\u0000' + sessionId + '\u0000' + highThreshold;
|
||||
long now = clock.getAsLong();
|
||||
CacheEntry cached = cache.get(cacheKey);
|
||||
if (cached != null && now - cached.readAtMillis() < ttlMillis) {
|
||||
return cached.reading();
|
||||
}
|
||||
Reading fresh = readUncached(base, sessionId);
|
||||
Reading fresh = readUncached(base, sessionId, highThreshold);
|
||||
cache.put(cacheKey, new CacheEntry(fresh, now));
|
||||
return fresh;
|
||||
}
|
||||
|
||||
private Reading readUncached(String base, String sessionId) {
|
||||
/** {@link #HIGH_THRESHOLD_FRACTION} of {@code effectiveWindowTokens}, or the fixed fallback. */
|
||||
private static long highThreshold(Long effectiveWindowTokens) {
|
||||
if (effectiveWindowTokens == null || effectiveWindowTokens <= 0) {
|
||||
return HIGH_THRESHOLD_TOKENS;
|
||||
}
|
||||
return (long) (effectiveWindowTokens * HIGH_THRESHOLD_FRACTION);
|
||||
}
|
||||
|
||||
private Reading readUncached(String base, String sessionId, long highThreshold) {
|
||||
diskReads.incrementAndGet();
|
||||
Path file = findTranscript(base, sessionId);
|
||||
if (file == null) {
|
||||
@@ -199,7 +220,7 @@ public final class LeadContextGauge {
|
||||
} catch (IOException e) {
|
||||
return Reading.unknown();
|
||||
}
|
||||
return parse(tail);
|
||||
return parse(tail, highThreshold);
|
||||
}
|
||||
|
||||
/**
|
||||
@@ -258,7 +279,7 @@ public final class LeadContextGauge {
|
||||
* read raced — see the "torn final line" section of the class javadoc) is skipped, not fatal.
|
||||
* Only when none of the remaining lines parse does this report {@link State#UNKNOWN}.
|
||||
*/
|
||||
private Reading parse(TailRead tail) {
|
||||
private Reading parse(TailRead tail, long highThreshold) {
|
||||
String text = new String(tail.bytes(), StandardCharsets.UTF_8);
|
||||
List<String> lines = new ArrayList<>(List.of(text.split("\n", -1)));
|
||||
if (!lines.isEmpty() && lines.get(lines.size() - 1).isEmpty()) {
|
||||
@@ -299,7 +320,7 @@ public final class LeadContextGauge {
|
||||
if (tokens == null) {
|
||||
return new Reading(State.UNKNOWN, null, compactions);
|
||||
}
|
||||
State state = tokens >= HIGH_THRESHOLD_TOKENS ? State.HIGH : State.OK;
|
||||
State state = tokens >= highThreshold ? State.HIGH : State.OK;
|
||||
return new Reading(state, tokens, compactions);
|
||||
}
|
||||
|
||||
|
||||
@@ -201,8 +201,8 @@ public final class LeadLauncher {
|
||||
/**
|
||||
* How many live leads exist per configured name, and which of that name's labelled tabs are
|
||||
* <em>not</em> live: a running agent in a tab labelled with that lead's exact {@code tab}
|
||||
* (CB-579). Member workspaces are excluded, exactly as the scanner excludes them: a member must
|
||||
* not be counted as a lead because it happens to sit in a matching tab.
|
||||
* (CB-579). A member sitting in the same shared workspace is not counted as a lead because its
|
||||
* tab carries a different label, not because any workspace is excluded from this count.
|
||||
*
|
||||
* <p>There used to be a second path here — a running agent on the terminal a
|
||||
* {@code fleet.leaders.<name>.terminal} pin named, for a lead opened and pinned by hand. That
|
||||
|
||||
@@ -31,6 +31,20 @@ public final class ConnectionIdentity {
|
||||
this.cwds = cwds;
|
||||
}
|
||||
|
||||
/**
|
||||
* The {@link PaneLocator} this identity resolves callers against — fleetd #612 CB-185: lets a
|
||||
* test drive the exact {@link PaneLocator} a real assembly wired up (e.g. {@code
|
||||
* FleetdAssembly}'s {@code new ConnectionIdentity(new PaneLocator(herdr, memberHerdr), ...)})
|
||||
* directly with a chosen pid, bypassing the OS-dependent {@link PeerPidLookup} that {@link
|
||||
* #resolve} otherwise goes through. A full HTTP round trip cannot exercise this: {@code
|
||||
* LsofPeerPidLookup} excludes its own pid, and an in-process test client and server share one
|
||||
* JVM pid, so {@code pidForLocalPort} always returns {@code -1} and {@link PaneLocator} never
|
||||
* gets called at all.
|
||||
*/
|
||||
public PaneLocator panes() {
|
||||
return panes;
|
||||
}
|
||||
|
||||
/**
|
||||
* The caller resolved from the connection: its worker {@code terminal} (or {@code null} for the
|
||||
* primary / an off-host client), its {@code pid} (or {@code -1} if not resolvable), and whether
|
||||
|
||||
@@ -107,6 +107,13 @@ public final class FleetMcp {
|
||||
* untested identity heuristic.
|
||||
*/
|
||||
private final boolean authorizationEnforced;
|
||||
/**
|
||||
* fleetd #612 CB-185: kept as a field (rather than only captured by the {@code
|
||||
* contextExtractor} closure built in the constructor) so a test can reach the exact {@link
|
||||
* ConnectionIdentity} — and, through {@link ConnectionIdentity#panes()}, the exact {@link
|
||||
* dev.ltms.fleet.herdr.PaneLocator} — that a real assembly wired up. See {@link #identity()}.
|
||||
*/
|
||||
private final ConnectionIdentity identity;
|
||||
private final Metrics metrics; // CB-502: null → auth failures not counted
|
||||
private final CapacitySource capacity;
|
||||
private final HealthCoverageSource healthCoverage;
|
||||
@@ -277,10 +284,15 @@ public final class FleetMcp {
|
||||
* no profile, or that profile sets no {@code configDir} override — either
|
||||
* way {@link LeadContextGauge} then falls back to its own built-in default,
|
||||
* exactly as before this ticket
|
||||
* @param windowFor lead name → that lead's profile's effective auto-compact window (see
|
||||
* {@code dev.ltms.fleet.config.FleetConfig.Profile
|
||||
* #effectiveAutoCompactWindow()}), or {@code null} when it cannot be
|
||||
* resolved — either way {@link LeadContextGauge} falls back to its own
|
||||
* fixed HIGH threshold
|
||||
*/
|
||||
public record LeadConfigDirSource(Function<String, String> configDirFor) {
|
||||
/** Inert source — every lead reads {@link LeadContextGauge}'s built-in default {@code configDir}. */
|
||||
public static LeadConfigDirSource none() { return new LeadConfigDirSource(_ -> null); }
|
||||
public record LeadConfigDirSource(Function<String, String> configDirFor, Function<String, Long> windowFor) {
|
||||
/** Inert source — every lead reads {@link LeadContextGauge}'s built-in defaults. */
|
||||
public static LeadConfigDirSource none() { return new LeadConfigDirSource(_ -> null, _ -> null); }
|
||||
}
|
||||
|
||||
/**
|
||||
@@ -396,6 +408,7 @@ public final class FleetMcp {
|
||||
Objects.requireNonNull(callers, "callers");
|
||||
this.authorizationEnforced = Objects.requireNonNull(authorizationMode, "authorizationMode")
|
||||
== AuthorizationMode.ENFORCED;
|
||||
this.identity = identity;
|
||||
this.leadChannel = leadChannel;
|
||||
this.peers = peers == null ? List.of() : List.copyOf(peers);
|
||||
this.capacity = capacity;
|
||||
@@ -677,7 +690,7 @@ public final class FleetMcp {
|
||||
return null; // AuthorizationMode.UNENFORCED: authorization not enforced (fleetd #518)
|
||||
}
|
||||
if (Authz.permits(caller, action, target)) {
|
||||
if (action != Authz.Action.READ) {
|
||||
if (action != Authz.Action.READ && action != Authz.Action.TASK_READ) {
|
||||
AuditLog.allowed(caller, action, target); // reads would drown the trail
|
||||
}
|
||||
return null;
|
||||
@@ -755,6 +768,15 @@ public final class FleetMcp {
|
||||
return transport;
|
||||
}
|
||||
|
||||
/**
|
||||
* The {@link ConnectionIdentity} this server resolves every caller against — fleetd #612
|
||||
* CB-185: lets a test reach the exact {@link dev.ltms.fleet.herdr.PaneLocator} a real assembly
|
||||
* wired up (via {@link ConnectionIdentity#panes()}), rather than a copy built for the test.
|
||||
*/
|
||||
public ConnectionIdentity identity() {
|
||||
return identity;
|
||||
}
|
||||
|
||||
/** Mark a connected spawned member available for the injector readiness gate. */
|
||||
static void markSpawnedMemberPresent(Principal caller, MemberPresence presence) {
|
||||
if (caller.isSpawnedMember()) {
|
||||
@@ -780,6 +802,46 @@ public final class FleetMcp {
|
||||
return server.listTools();
|
||||
}
|
||||
|
||||
/**
|
||||
* fleetd #612 B3 — same reason as {@link #registeredTools()}: a test that must drive the REAL
|
||||
* {@link QuarantineSource} (and the real {@link BackendQuarantine} it wraps) this daemon was
|
||||
* assembled with, rather than scraping {@code FleetdAssembly.java}'s source text for the
|
||||
* constructor call that built it. Unlike {@link #registeredTools()}'s callers, that test cannot
|
||||
* live in this package: it also builds the {@code ResourcePorts} that drives
|
||||
* {@code FleetdAssembly.assembleAndStart}, and {@code ResourcePorts}' methods return
|
||||
* {@code Fleetd}-nested types that are only visible from package {@code dev.ltms.fleet} — so
|
||||
* this accessor is {@code public}, not package-private, to stay reachable from there. {@code
|
||||
* FleetdBackendQuarantineAssemblyTest} quarantines a credential twice through this exact
|
||||
* instance and checks the second cooldown is longer than the first — the one behavioural
|
||||
* difference {@link BackendQuarantine#withEscalation} and the flat two-argument constructor
|
||||
* actually produce.
|
||||
*/
|
||||
public QuarantineSource quarantineSource() {
|
||||
return quarantine;
|
||||
}
|
||||
|
||||
/**
|
||||
* fleetd #612 B3 — as {@link #quarantineSource()}, {@code public} for the same cross-package
|
||||
* reason, for the real {@link LeadSeatSource} this daemon was assembled with. {@code
|
||||
* FleetdLeadSeatAssemblyTest} calls {@code seatsFor} on this exact instance and checks it
|
||||
* reports a live lead's seat, which {@link LeadSeatSource#none()} can never do (it is a
|
||||
* constant-zero function regardless of input).
|
||||
*/
|
||||
public LeadSeatSource leadSeatSource() {
|
||||
return leadSeats;
|
||||
}
|
||||
|
||||
/**
|
||||
* fleetd #612 B3 — as {@link #quarantineSource()}, {@code public} for the same cross-package
|
||||
* reason, for the real {@link LeadRollover} (or {@code null}) this daemon was assembled with.
|
||||
* {@code FleetdLeadRolloverAssemblyTest} drives {@code open}/{@code confirm} on this exact
|
||||
* instance and waits for the real continuation to send {@code /clear} and {@code bootstrapText}
|
||||
* through the real {@code router.leadAgents()}.
|
||||
*/
|
||||
public LeadRollover leadRollover() {
|
||||
return leadRollover;
|
||||
}
|
||||
|
||||
// --- tool logic (thin adapters over the services; unit-testable) ---------------------------
|
||||
|
||||
/**
|
||||
@@ -973,31 +1035,21 @@ public final class FleetMcp {
|
||||
}
|
||||
|
||||
/**
|
||||
* Which authorization action a {@code fleet_poll} call needs, decided by its arguments
|
||||
* (fleetd #272, widened by fleetd #421).
|
||||
* Which authorization action a {@code fleet_poll} call needs, decided by its arguments.
|
||||
*
|
||||
* <p>{@code fleet_poll} is now <strong>three operations behind one tool name</strong>. With
|
||||
* <p>{@code fleet_poll} is <strong>three operations behind one tool name</strong>. With
|
||||
* {@code ticket} it observes an async delegation and changes nothing, which is a {@link
|
||||
* Authz.Action#READ}. With {@code target} it calls {@link MessageService#drainReplies} on that
|
||||
* session -- the replies are removed from the inbox and a second call returns nothing -- so it
|
||||
* is a {@link Authz.Action#DRAIN}, the same gate {@code fleet_ack} already uses for removing a
|
||||
* single message, and the same one the REST path uses at {@code FleetApp.drainReplies}. With
|
||||
* {@code coordId} it reads (never acks) this daemon's own held lead-to-lead mail, which is a
|
||||
* {@link Authz.Action#COORD_READ} -- <strong>not</strong> {@code READ}, even though nothing is
|
||||
* consumed: {@code READ}'s grant is open to every authenticated role on the premise that the
|
||||
* roster carries no secrets, and a lead-to-lead body is not the roster. Mapping a non-destructive
|
||||
* peer-mail read to {@code READ} would let any worker read every peer lead's mail in full.
|
||||
*
|
||||
* <p>Before this method existed (fleetd #272) the handler passed a constant {@code READ} for
|
||||
* both of the original branches. {@code READ} is open to every authenticated role, so any
|
||||
* worker could read a peer's id out of {@code fleet_list} and destroy the replies that peer had
|
||||
* queued for the primary. The gate failed open, and it did so because the required action is a
|
||||
* function of the arguments while the handler chose it before looking at them.
|
||||
* Authz.Action#TASK_READ}. With {@code target} it calls {@link MessageService#drainReplies} on
|
||||
* that session -- the replies are removed from the inbox and a second call returns nothing --
|
||||
* so it is a {@link Authz.Action#DRAIN}, the same gate {@code fleet_ack} already uses for
|
||||
* removing a single message, and the same one the REST path uses at
|
||||
* {@code FleetApp.drainReplies}. With {@code coordId} it reads (never acks) this daemon's own
|
||||
* held lead-to-lead mail, which is a {@link Authz.Action#COORD_READ} -- <strong>not</strong>
|
||||
* {@code TASK_READ} or {@code READ}: a lead-to-lead body is a different inbox from either, and
|
||||
* folding it into either would let any worker or architect read every peer lead's mail in full.
|
||||
*
|
||||
* <p>The choice lives in this method, and not inline in the handler, so that a test can assert
|
||||
* the mapping the handler actually uses. {@code FleetMcpAuthzTest} already checked every
|
||||
* {@link Authz.Action} against every {@link Role} and passed throughout -- it tested the policy
|
||||
* table, which was correct, while the defect was in which action the caller handed it.
|
||||
* the mapping the handler actually uses.
|
||||
*
|
||||
* <p>Checked first, and exclusively of {@code target}: a call naming {@code coordId} is reading
|
||||
* a different inbox entirely (this daemon's own lead channel, never a worker's), so it takes
|
||||
@@ -1010,7 +1062,30 @@ public final class FleetMcp {
|
||||
if (!isBlank(coordId)) {
|
||||
return Authz.Action.COORD_READ;
|
||||
}
|
||||
return isBlank(target) ? Authz.Action.READ : Authz.Action.DRAIN;
|
||||
return isBlank(target) ? Authz.Action.TASK_READ : Authz.Action.DRAIN;
|
||||
}
|
||||
|
||||
/**
|
||||
* Which authorization action a {@code fleet_send} call needs, decided by its arguments.
|
||||
*
|
||||
* <p>{@code fleet_send} is three call shapes behind one tool name, mirroring {@link
|
||||
* #pollAction}. With {@code coordId} it addresses a peer lead on another daemon over the
|
||||
* coordination broker, which is {@link Authz.Action#COORD_SEND}. With {@code turnId} it
|
||||
* resolves a worker's blocked {@code fleet_ask} and resumes that turn, which is {@link
|
||||
* Authz.Action#ANSWER}. Otherwise it delivers to a local session by {@code sessionId}, which is
|
||||
* the plain {@link Authz.Action#SEND}.
|
||||
*
|
||||
* <p>Checked in the same order the handler branches: {@code coordId} first and exclusively of
|
||||
* {@code turnId}, matching {@link #sendToLead}'s own mutual-exclusion check.
|
||||
*
|
||||
* @param coordId the {@code coordId} argument of the call, or {@code null}/blank when absent
|
||||
* @param turnId the {@code turnId} argument of the call, or {@code null}/blank when absent
|
||||
*/
|
||||
static Authz.Action sendAction(String coordId, String turnId) {
|
||||
if (!isBlank(coordId)) {
|
||||
return Authz.Action.COORD_SEND;
|
||||
}
|
||||
return isBlank(turnId) ? Authz.Action.SEND : Authz.Action.ANSWER;
|
||||
}
|
||||
|
||||
/**
|
||||
@@ -1035,10 +1110,11 @@ public final class FleetMcp {
|
||||
*/
|
||||
private static Authz.Action authzAction(FleetTool tool, Map<String, Object> arguments) {
|
||||
return switch (tool) {
|
||||
case SEND -> Authz.Action.SEND;
|
||||
case SEND -> sendAction(str(arguments, "coordId"), str(arguments, "turnId"));
|
||||
case REPLY -> Authz.Action.REPLY;
|
||||
case ASK -> Authz.Action.ASK;
|
||||
case STATUS, LIST, PROFILES, WHOAMI -> Authz.Action.READ;
|
||||
case STATUS -> Authz.Action.TASK_READ;
|
||||
case LIST, PROFILES, WHOAMI -> Authz.Action.READ;
|
||||
case POLL -> pollAction(str(arguments, "target"), str(arguments, "coordId"));
|
||||
case ACK -> Authz.Action.DRAIN;
|
||||
case SPAWN -> Authz.Action.SPAWN;
|
||||
@@ -2096,7 +2172,8 @@ public final class FleetMcp {
|
||||
m.put("self", true);
|
||||
}
|
||||
String configDir = leadConfigDirs.configDirFor().apply(name);
|
||||
m.put("context", contextView(contextGauge, live, configDir));
|
||||
Long effectiveWindow = leadConfigDirs.windowFor().apply(name);
|
||||
m.put("context", contextView(contextGauge, live, configDir, effectiveWindow));
|
||||
return m;
|
||||
}
|
||||
|
||||
@@ -2106,12 +2183,15 @@ public final class FleetMcp {
|
||||
* {@code fleet.leaders.<name>.profile} → that profile's own {@code configDir:} — or {@code null}
|
||||
* when the lead's entry names no profile, or that profile sets no override, in which case
|
||||
* {@link LeadContextGauge#read} falls back to its own built-in default
|
||||
* ({@code <user.home>/.claude}).
|
||||
* ({@code <user.home>/.claude}). {@code effectiveWindowTokens} is the same lead's resolved
|
||||
* auto-compact window, or {@code null} when it cannot be resolved, in which case the gauge
|
||||
* falls back to its own fixed HIGH threshold instead.
|
||||
*/
|
||||
private static Map<String, Object> contextView(LeadContextGauge contextGauge, Agent live, String configDir) {
|
||||
private static Map<String, Object> contextView(LeadContextGauge contextGauge, Agent live, String configDir,
|
||||
Long effectiveWindowTokens) {
|
||||
String sessionId = live == null ? null : live.sessionId();
|
||||
String agentType = live == null ? null : live.agentType();
|
||||
LeadContextGauge.Reading reading = contextGauge.read(configDir, sessionId, agentType);
|
||||
LeadContextGauge.Reading reading = contextGauge.read(configDir, sessionId, agentType, effectiveWindowTokens);
|
||||
Map<String, Object> c = new LinkedHashMap<>();
|
||||
c.put("state", reading.state().name().toLowerCase());
|
||||
if (reading.tokens() != null) {
|
||||
|
||||
@@ -0,0 +1,24 @@
|
||||
package dev.ltms.fleet.msg;
|
||||
|
||||
/**
|
||||
* fleetd #612 A-gaps (gap 1): a {@link LeadChannel} that its owner can also close.
|
||||
*
|
||||
* <p>{@link LeadChannel}'s own javadoc says plainly that {@code close()} is deliberately left out
|
||||
* of that interface — draining is a caller convenience nobody uses, and closing is the
|
||||
* <em>owner's</em> job. This interface is that owner's own, wider view: whoever opens the
|
||||
* coordination mailbox (the assembly that builds the daemon) also needs to close it from the
|
||||
* shutdown path, and a test standing in for a real broker connection needs a fake it can mark
|
||||
* closed, without ever holding a live connection. Every ordinary consumer ({@code FleetMcp},
|
||||
* {@link LeadCoordLoop}) keeps taking the narrower {@link LeadChannel} exactly as before — only
|
||||
* the owner speaks this wider one.
|
||||
*
|
||||
* <p>{@link LeadMailbox} is still the only production implementation. This only generalises the
|
||||
* TYPE its owner holds it as (previously the concrete class), so a test can substitute a fake
|
||||
* closeable channel instead of a real AMQP connection.
|
||||
*/
|
||||
public interface LeadChannelHandle extends LeadChannel, AutoCloseable {
|
||||
|
||||
/** Release the underlying connection. Declared with no checked exception, unlike the plain {@link AutoCloseable#close()}. */
|
||||
@Override
|
||||
void close();
|
||||
}
|
||||
@@ -461,9 +461,9 @@ public final class LeadHeartbeatLoop {
|
||||
.append(reading.compactions()).append(' ').append(compactionWord).append(" so far.");
|
||||
} else {
|
||||
// A HIGH reading always carries a non-null token count today: LeadContextGauge only
|
||||
// reaches HIGH by comparing a number against HIGH_THRESHOLD_TOKENS. That invariant
|
||||
// lives in another class and nothing asserts it, so this branch does not rely on it —
|
||||
// it drops the token clause rather than printing "null tokens".
|
||||
// reaches HIGH by comparing a number against a threshold. That invariant lives in
|
||||
// another class and nothing asserts it, so this branch does not rely on it — it drops
|
||||
// the token clause rather than printing "null tokens".
|
||||
sb.append(" (").append(reading.compactions()).append(' ').append(compactionWord)
|
||||
.append(" so far).");
|
||||
}
|
||||
|
||||
@@ -63,7 +63,7 @@ import java.util.concurrent.TimeoutException;
|
||||
* which messages reached a lead. Any publish still awaiting its confirm is failed rather than left to idle out
|
||||
* the confirm timeout against a sequence number that means nothing on the new channel.
|
||||
*/
|
||||
public final class LeadMailbox implements LeadChannel, AutoCloseable {
|
||||
public final class LeadMailbox implements LeadChannelHandle {
|
||||
|
||||
private static final Logger log = LoggerFactory.getLogger(LeadMailbox.class);
|
||||
|
||||
|
||||
@@ -49,18 +49,36 @@ import java.util.stream.Collectors;
|
||||
*/
|
||||
public final class FleetApp {
|
||||
|
||||
/** The authorization action the matching route handler hands to {@link #allow}. */
|
||||
/**
|
||||
* The authorization action the matching route handler hands to {@link #allow}, for a route
|
||||
* whose action does not depend on the request body.
|
||||
*/
|
||||
static Authz.Action routeAction(String route) {
|
||||
return routeAction(route, null);
|
||||
}
|
||||
|
||||
/**
|
||||
* As above, plus the one route whose action depends on the body: {@code POST
|
||||
* /sessions/{id}/message} carries a {@code turnId} (the answer-a-blocked-worker shape) or not
|
||||
* (a plain delivery), mirroring {@code FleetMcp#sendAction}'s split of the same two call
|
||||
* shapes over MCP. {@code turnId} is ignored by every other route.
|
||||
*
|
||||
* @param turnId the request body's {@code turnId}, or {@code null}/blank when absent or not
|
||||
* applicable to this route
|
||||
*/
|
||||
static Authz.Action routeAction(String route, String turnId) {
|
||||
return switch (route) {
|
||||
case "GET /metrics" -> Authz.Action.METRICS;
|
||||
case "POST /members" -> Authz.Action.SPAWN;
|
||||
case "DELETE /members/{paneId}" -> Authz.Action.STOP;
|
||||
case "POST /sessions/{id}/message" -> Authz.Action.SEND;
|
||||
case "POST /sessions/{id}/message" -> turnId == null || turnId.isBlank()
|
||||
? Authz.Action.SEND : Authz.Action.ANSWER;
|
||||
case "POST /sessions/{id}/reply" -> Authz.Action.REPLY;
|
||||
case "GET /sessions/{id}/replies" -> Authz.Action.DRAIN;
|
||||
case "POST /sessions/{id}/ask" -> Authz.Action.ASK;
|
||||
case "GET /sessions", "GET /agents", "GET /members", "GET /profiles",
|
||||
"GET /member-credentials", "GET /sessions/{id}/status", "GET /tasks/{ticket}" -> Authz.Action.READ;
|
||||
"GET /member-credentials" -> Authz.Action.READ;
|
||||
case "GET /sessions/{id}/status", "GET /tasks/{ticket}" -> Authz.Action.TASK_READ;
|
||||
default -> throw new IllegalArgumentException("route has no authorization gate: " + route);
|
||||
};
|
||||
}
|
||||
@@ -250,7 +268,8 @@ public final class FleetApp {
|
||||
}
|
||||
Principal caller = ctx.attribute(CALLER);
|
||||
if (Authz.permits(caller, action, target)) {
|
||||
if (action != Authz.Action.READ && action != Authz.Action.METRICS) {
|
||||
if (action != Authz.Action.READ && action != Authz.Action.METRICS
|
||||
&& action != Authz.Action.TASK_READ) {
|
||||
AuditLog.allowed(caller, action, target); // reads would drown the trail
|
||||
}
|
||||
return true;
|
||||
@@ -603,26 +622,33 @@ public final class FleetApp {
|
||||
* status-gated injector and block until the worker returns a structured {@code fleet_reply}.
|
||||
* Times out with a typed 202 (working / queued / busy) rather than an error — the message may
|
||||
* still land.
|
||||
*
|
||||
* <p>Two call shapes share this route, exactly as {@code fleet_send} does over MCP (see
|
||||
* {@code FleetMcp#sendAction}): a plain delivery to {@code id}, and -- when the body carries
|
||||
* {@code turnId} -- resolving a worker's blocked question. The body is parsed before the
|
||||
* authorization check so the right one of {@link Authz.Action#SEND}/{@link Authz.Action#ANSWER}
|
||||
* reaches the gate; a body that fails to parse is treated as the plain shape for that check
|
||||
* alone, and is rejected afterward exactly as before.
|
||||
*/
|
||||
private void sendMessage(Context ctx) {
|
||||
String id = ctx.pathParam("id");
|
||||
if (!allow(ctx, routeAction("POST /sessions/{id}/message"), id)) {
|
||||
JsonNode body;
|
||||
try {
|
||||
body = mapper.readTree(ctx.body());
|
||||
} catch (Exception e) {
|
||||
body = null;
|
||||
}
|
||||
String turnId = body == null ? null : body.path("turnId").asText(null);
|
||||
if (!allow(ctx, routeAction("POST /sessions/{id}/message", turnId), id)) {
|
||||
return;
|
||||
}
|
||||
String content;
|
||||
String turnId;
|
||||
long timeout;
|
||||
boolean wait;
|
||||
try {
|
||||
JsonNode body = mapper.readTree(ctx.body());
|
||||
content = body.path("content").asText("");
|
||||
turnId = body.path("turnId").asText(null);
|
||||
timeout = body.path("timeoutMs").asLong(DEFAULT_MESSAGE_TIMEOUT_MS);
|
||||
wait = body.path("wait").asBoolean(true); // default: block for the reply (CB-104)
|
||||
} catch (Exception e) {
|
||||
if (body == null) {
|
||||
ctx.status(400).json(Map.of("error", "bad_request", "detail", "body must be JSON"));
|
||||
return;
|
||||
}
|
||||
String content = body.path("content").asText("");
|
||||
long timeout = body.path("timeoutMs").asLong(DEFAULT_MESSAGE_TIMEOUT_MS);
|
||||
boolean wait = body.path("wait").asBoolean(true); // default: block for the reply (CB-104)
|
||||
if (content.isBlank()) {
|
||||
ctx.status(400).json(Map.of("error", "bad_request", "detail", "content is required"));
|
||||
return;
|
||||
|
||||
@@ -0,0 +1,183 @@
|
||||
package dev.ltms.fleet;
|
||||
|
||||
import ch.qos.logback.classic.Level;
|
||||
import ch.qos.logback.classic.Logger;
|
||||
import ch.qos.logback.classic.spi.ILoggingEvent;
|
||||
import ch.qos.logback.core.read.ListAppender;
|
||||
import dev.ltms.fleet.config.ConfigRef;
|
||||
import dev.ltms.fleet.config.FleetConfig;
|
||||
import dev.ltms.fleet.guard.SubscriptionGuard;
|
||||
import dev.ltms.fleet.herdr.FakeHerdr;
|
||||
import dev.ltms.fleet.herdr.HerdrClient;
|
||||
import dev.ltms.fleet.msg.InMemoryReplyInbox;
|
||||
import dev.ltms.fleet.msg.LeadChannel;
|
||||
import dev.ltms.fleet.msg.LeadChannelHandle;
|
||||
import dev.ltms.fleet.msg.LeadMessage;
|
||||
import dev.ltms.fleet.msg.ReplyInbox;
|
||||
import io.javalin.Javalin;
|
||||
import org.junit.jupiter.api.Test;
|
||||
import org.junit.jupiter.api.io.TempDir;
|
||||
import org.slf4j.LoggerFactory;
|
||||
|
||||
import java.nio.file.Files;
|
||||
import java.nio.file.Path;
|
||||
import java.util.List;
|
||||
import java.util.Map;
|
||||
import java.util.concurrent.Executors;
|
||||
import java.util.concurrent.ScheduledExecutorService;
|
||||
import java.util.concurrent.atomic.AtomicInteger;
|
||||
import java.util.function.LongSupplier;
|
||||
|
||||
import static org.junit.jupiter.api.Assertions.assertEquals;
|
||||
import static org.junit.jupiter.api.Assertions.assertFalse;
|
||||
import static org.junit.jupiter.api.Assertions.assertNotNull;
|
||||
import static org.junit.jupiter.api.Assertions.assertNull;
|
||||
import static org.junit.jupiter.api.Assertions.assertSame;
|
||||
import static org.junit.jupiter.api.Assertions.assertTrue;
|
||||
|
||||
/**
|
||||
* fleetd #612 step 4 ranks 3 and 8: the assembled daemon must use the AMQP openers from {@link
|
||||
* ResourcePorts}, and its startup report must describe the object the runtime actually owns. These
|
||||
* fakes never open a socket.
|
||||
*/
|
||||
class FleetdAssemblyAmqpOpenersTest {
|
||||
|
||||
private static final String COORD_ID = "assembly-test";
|
||||
|
||||
private static final class DurableReplyInbox implements ReplyInbox {
|
||||
@Override public void own(String target) { }
|
||||
@Override public void release(String target) { }
|
||||
@Override public void publish(String target, String msgId, String content) { }
|
||||
@Override public List<InboxMessage> peek(String target) { return List.of(); }
|
||||
@Override public boolean ack(String target, String msgId) { return false; }
|
||||
}
|
||||
|
||||
private static final class DurableLeadMailbox implements LeadChannelHandle {
|
||||
@Override public void publish(String toCoordId, LeadMessage message) { }
|
||||
@Override public List<LeadMessage> peek() { return List.of(); }
|
||||
@Override public void ack(String msgId) { }
|
||||
@Override public String selfCoordId() { return COORD_ID; }
|
||||
@Override public boolean heldDurable() { return true; }
|
||||
@Override public MailboxState inspect(String coordId) { return MailboxState.unknown(coordId); }
|
||||
@Override public void close() { }
|
||||
}
|
||||
|
||||
private static final class RecordingPorts implements ResourcePorts {
|
||||
final FakeHerdr herdr = new FakeHerdr();
|
||||
final DurableReplyInbox replyInbox = new DurableReplyInbox();
|
||||
final DurableLeadMailbox leadMailbox = new DurableLeadMailbox();
|
||||
final AtomicInteger replyOpenCalls = new AtomicInteger();
|
||||
final AtomicInteger mailboxOpenCalls = new AtomicInteger();
|
||||
final boolean openSucceeds;
|
||||
|
||||
RecordingPorts(boolean openSucceeds) {
|
||||
this.openSucceeds = openSucceeds;
|
||||
}
|
||||
|
||||
@Override public Map<String, String> environment() { return Map.of(); }
|
||||
@Override public HerdrClient connectHerdr(Path socketPath) { return herdr; }
|
||||
@Override public Fleetd.AmqpOpener replyInboxOpener() {
|
||||
return (uri, prefetch) -> {
|
||||
replyOpenCalls.incrementAndGet();
|
||||
if (!openSucceeds) throw new IllegalStateException("fake reply broker is down");
|
||||
return replyInbox;
|
||||
};
|
||||
}
|
||||
@Override public Fleetd.LeadMailboxOpener leadMailboxOpener() {
|
||||
return (uri, selfId, prefetch) -> {
|
||||
mailboxOpenCalls.incrementAndGet();
|
||||
if (!openSucceeds) throw new IllegalStateException("fake coordination broker is down");
|
||||
return leadMailbox;
|
||||
};
|
||||
}
|
||||
@Override public LongSupplier nanoClock() { return System::nanoTime; }
|
||||
@Override
|
||||
public Runnable herdrPollWait() {
|
||||
// Never invoked: this test's FakeHerdr answers immediately, so awaitHerdr never polls.
|
||||
return () -> {
|
||||
throw new UnsupportedOperationException("herdrPollWait must not be called — herdr is healthy");
|
||||
};
|
||||
}
|
||||
@Override public LongSupplier wallClockNanos() { return System::nanoTime; }
|
||||
@Override public ScheduledExecutorService newScheduler(String purpose) {
|
||||
return Executors.newSingleThreadScheduledExecutor();
|
||||
}
|
||||
@Override public void addShutdownHook(Runnable hook) { }
|
||||
@Override public void startHttp(Javalin app, String host, int port) { }
|
||||
}
|
||||
|
||||
private static FleetConfig writeConfig(Path dir) throws Exception {
|
||||
Files.createDirectories(dir);
|
||||
Path config = dir.resolve("fleetd.yaml");
|
||||
Files.writeString(config, """
|
||||
bind:
|
||||
host: 127.0.0.1
|
||||
port: 8765
|
||||
idleSleepGuard:
|
||||
enabled: false
|
||||
broker:
|
||||
uri: "amqp://fake-reply-broker/vh"
|
||||
coordinator:
|
||||
uri: "amqp://fake-coordination-broker/vh"
|
||||
selfId: "assembly-test"
|
||||
""");
|
||||
return FleetConfig.load(config);
|
||||
}
|
||||
|
||||
private static FleetdRuntime assemble(Path dir, RecordingPorts ports) throws Exception {
|
||||
FleetConfig cfg = writeConfig(dir);
|
||||
return FleetdAssembly.assembleAndStart(new AssemblyInputs(cfg, new ConfigRef(dir.resolve("fleetd.yaml"), cfg),
|
||||
new SubscriptionGuard(cfg.guard().hostSet())), ports);
|
||||
}
|
||||
|
||||
private static boolean reportContains(ListAppender<ILoggingEvent> appender, String text) {
|
||||
return appender.list.stream().map(ILoggingEvent::getFormattedMessage).anyMatch(message -> message.contains(text));
|
||||
}
|
||||
|
||||
@Test
|
||||
void assembledAmqpOpenersAndTheirReportsAgreeOnDurableAndFallbackStates(@TempDir Path dir) throws Exception {
|
||||
Logger logger = (Logger) LoggerFactory.getLogger(Fleetd.class);
|
||||
Level oldLevel = logger.getLevel();
|
||||
ListAppender<ILoggingEvent> reports = new ListAppender<>();
|
||||
reports.start();
|
||||
logger.setLevel(Level.INFO);
|
||||
logger.addAppender(reports);
|
||||
try {
|
||||
RecordingPorts durablePorts = new RecordingPorts(true);
|
||||
FleetdRuntime durable = assemble(dir.resolve("durable"), durablePorts);
|
||||
try {
|
||||
// Control: this fails loudly if the assembly did not run or used an inert opener.
|
||||
assertEquals(1, durablePorts.replyOpenCalls.get(), "assembly must call replyInboxOpener once");
|
||||
assertEquals(1, durablePorts.mailboxOpenCalls.get(), "assembly must call leadMailboxOpener once");
|
||||
assertSame(durablePorts.replyInbox, durable.replyInbox(),
|
||||
"the durable reply report must describe the exact inbox the runtime owns");
|
||||
assertSame(durablePorts.leadMailbox, durable.leadMailbox(),
|
||||
"the coordination-on report must describe the exact mailbox the runtime owns");
|
||||
assertNotNull(durable.leadCoordLoop(), "a durable mailbox must start lead coordination");
|
||||
assertTrue(reportContains(reports, "reply inbox: AMQP broker (durable)"));
|
||||
assertTrue(reportContains(reports, "lead coordination: ON as coord-id " + COORD_ID));
|
||||
} finally {
|
||||
durable.close();
|
||||
}
|
||||
|
||||
reports.list.clear();
|
||||
RecordingPorts fallbackPorts = new RecordingPorts(false);
|
||||
FleetdRuntime fallback = assemble(dir.resolve("fallback"), fallbackPorts);
|
||||
try {
|
||||
assertEquals(1, fallbackPorts.replyOpenCalls.get(), "assembly must call the failing reply opener once");
|
||||
assertEquals(1, fallbackPorts.mailboxOpenCalls.get(), "assembly must call the failing mailbox opener once");
|
||||
assertTrue(fallback.replyInbox() instanceof InMemoryReplyInbox,
|
||||
"a failed reply opener must make the runtime own the in-memory fallback");
|
||||
assertNull(fallback.leadMailbox(), "a failed mailbox opener must leave coordination off");
|
||||
assertNull(fallback.leadCoordLoop(), "coordination must not start without a mailbox");
|
||||
assertTrue(reportContains(reports, "reply inbox: in-memory (soft-state)"));
|
||||
assertTrue(reportContains(reports, "lead-to-lead messaging is OFF"));
|
||||
} finally {
|
||||
fallback.close();
|
||||
}
|
||||
} finally {
|
||||
logger.detachAppender(reports);
|
||||
logger.setLevel(oldLevel);
|
||||
}
|
||||
}
|
||||
}
|
||||
@@ -0,0 +1,166 @@
|
||||
package dev.ltms.fleet;
|
||||
|
||||
import dev.ltms.fleet.auth.Authz;
|
||||
import dev.ltms.fleet.auth.Principal;
|
||||
import dev.ltms.fleet.config.ConfigRef;
|
||||
import dev.ltms.fleet.config.FleetConfig;
|
||||
import dev.ltms.fleet.guard.SubscriptionGuard;
|
||||
import dev.ltms.fleet.herdr.FakeHerdr;
|
||||
import dev.ltms.fleet.herdr.HerdrClient;
|
||||
import dev.ltms.fleet.mcp.FleetMcp;
|
||||
import dev.ltms.fleet.msg.ReplyInbox;
|
||||
import io.javalin.Javalin;
|
||||
import io.modelcontextprotocol.spec.McpSchema;
|
||||
import org.junit.jupiter.api.AfterEach;
|
||||
import org.junit.jupiter.api.Test;
|
||||
import org.junit.jupiter.api.io.TempDir;
|
||||
|
||||
import java.lang.reflect.Method;
|
||||
import java.nio.file.Files;
|
||||
import java.nio.file.Path;
|
||||
import java.util.List;
|
||||
import java.util.Map;
|
||||
import java.util.concurrent.Executors;
|
||||
import java.util.concurrent.ScheduledExecutorService;
|
||||
import java.util.function.LongSupplier;
|
||||
|
||||
import static org.junit.jupiter.api.Assertions.assertNotNull;
|
||||
import static org.junit.jupiter.api.Assertions.assertNull;
|
||||
import static org.junit.jupiter.api.Assertions.assertTrue;
|
||||
|
||||
/**
|
||||
* Asserts that the {@link FleetMcp} built by {@link FleetdAssembly#assembleAndStart} applies the
|
||||
* authorization table: a worker is refused {@code SPAWN}, and the primary is allowed it.
|
||||
*/
|
||||
class FleetdAssemblyAuthorizationModeTest {
|
||||
|
||||
private static final class TestResourcePorts implements ResourcePorts {
|
||||
final FakeHerdr herdr = new FakeHerdr();
|
||||
Runnable shutdownHook;
|
||||
|
||||
@Override
|
||||
public Map<String, String> environment() {
|
||||
return Map.of();
|
||||
}
|
||||
|
||||
@Override
|
||||
public HerdrClient connectHerdr(Path socketPath) {
|
||||
return herdr;
|
||||
}
|
||||
|
||||
@Override
|
||||
public Fleetd.AmqpOpener replyInboxOpener() {
|
||||
return (uri, prefetch) -> new ReplyInbox() {
|
||||
@Override public void own(String target) { }
|
||||
@Override public void release(String target) { }
|
||||
@Override public void publish(String target, String msgId, String content) { }
|
||||
@Override public List<InboxMessage> peek(String target) { return List.of(); }
|
||||
@Override public boolean ack(String target, String msgId) { return false; }
|
||||
};
|
||||
}
|
||||
|
||||
@Override
|
||||
public Fleetd.LeadMailboxOpener leadMailboxOpener() {
|
||||
return (uri, selfCoordId, prefetch) -> {
|
||||
throw new UnsupportedOperationException(
|
||||
"leadMailboxOpener must not be called — no coordinator: block is configured");
|
||||
};
|
||||
}
|
||||
|
||||
@Override
|
||||
public LongSupplier nanoClock() {
|
||||
return System::nanoTime;
|
||||
}
|
||||
|
||||
@Override
|
||||
public LongSupplier wallClockNanos() {
|
||||
return System::nanoTime;
|
||||
}
|
||||
|
||||
@Override
|
||||
public ScheduledExecutorService newScheduler(String purpose) {
|
||||
return Executors.newSingleThreadScheduledExecutor();
|
||||
}
|
||||
|
||||
@Override
|
||||
public void addShutdownHook(Runnable hook) {
|
||||
shutdownHook = hook;
|
||||
}
|
||||
|
||||
@Override
|
||||
public void startHttp(Javalin app, String host, int port) {
|
||||
// Binding a real port would clash with any daemon already listening on it.
|
||||
}
|
||||
|
||||
@Override
|
||||
public Runnable herdrPollWait() {
|
||||
return () -> {
|
||||
throw new UnsupportedOperationException("FakeHerdr is healthy; no poll wait is expected");
|
||||
};
|
||||
}
|
||||
}
|
||||
|
||||
private TestResourcePorts ports;
|
||||
|
||||
@AfterEach
|
||||
void tearDown() {
|
||||
if (ports != null && ports.shutdownHook != null) {
|
||||
ports.shutdownHook.run();
|
||||
}
|
||||
}
|
||||
|
||||
private static FleetConfig writeConfig(Path dir) throws Exception {
|
||||
Path file = dir.resolve("fleetd.yaml");
|
||||
Files.writeString(file, """
|
||||
bind:
|
||||
host: 127.0.0.1
|
||||
port: 8765
|
||||
idleSleepGuard:
|
||||
enabled: false
|
||||
health:
|
||||
enabled: false
|
||||
broker:
|
||||
uri: "amqp://fake-test-broker/vh"
|
||||
""");
|
||||
return FleetConfig.load(file);
|
||||
}
|
||||
|
||||
private FleetMcp assemble(Path dir) throws Exception {
|
||||
FleetConfig cfg = writeConfig(dir);
|
||||
ports = new TestResourcePorts();
|
||||
FleetdRuntime runtime = FleetdAssembly.assembleAndStart(new AssemblyInputs(cfg,
|
||||
new ConfigRef(dir.resolve("fleetd.yaml"), cfg), new SubscriptionGuard(cfg.guard().hostSet())), ports);
|
||||
return runtime.mcp();
|
||||
}
|
||||
|
||||
/**
|
||||
* Invokes {@code FleetMcp#denyFor}, which is package-private to {@code dev.ltms.fleet.mcp}
|
||||
* while this test is in {@code dev.ltms.fleet}. Nothing here catches a missing method: if
|
||||
* {@code denyFor} is renamed or removed, {@link NoSuchMethodException} propagates and the
|
||||
* test fails.
|
||||
*/
|
||||
private static McpSchema.CallToolResult denyFor(FleetMcp mcp, Principal caller, Authz.Action action,
|
||||
String target) throws Exception {
|
||||
Method m = FleetMcp.class.getDeclaredMethod("denyFor", Principal.class, Authz.Action.class, String.class);
|
||||
m.setAccessible(true);
|
||||
return (McpSchema.CallToolResult) m.invoke(mcp, caller, action, target);
|
||||
}
|
||||
|
||||
@Test
|
||||
void productionBootPathRefusesAnUnauthorizedCallerThroughTheAssembledFleetMcp(@TempDir Path dir)
|
||||
throws Exception {
|
||||
FleetMcp mcp = assemble(dir);
|
||||
|
||||
McpSchema.CallToolResult deniedForWorker = denyFor(mcp, Principal.worker("term_a", 200),
|
||||
Authz.Action.SPAWN, "term_a");
|
||||
assertNotNull(deniedForWorker,
|
||||
"a worker must not be able to fleet_spawn through the assembled FleetMcp");
|
||||
assertTrue(deniedForWorker.isError(), "a refusal is returned as an MCP tool error");
|
||||
|
||||
McpSchema.CallToolResult allowedForPrimary = denyFor(mcp, Principal.primary(100),
|
||||
Authz.Action.SPAWN, "term_a");
|
||||
assertNull(allowedForPrimary,
|
||||
"control: the primary must still be allowed to fleet_spawn — otherwise the worker "
|
||||
+ "refusal above would pass even with the gate wired backwards");
|
||||
}
|
||||
}
|
||||
@@ -0,0 +1,230 @@
|
||||
package dev.ltms.fleet;
|
||||
|
||||
import dev.ltms.fleet.config.ConfigRef;
|
||||
import dev.ltms.fleet.config.FleetConfig;
|
||||
import dev.ltms.fleet.guard.SubscriptionGuard;
|
||||
import dev.ltms.fleet.herdr.FakeHerdr;
|
||||
import dev.ltms.fleet.herdr.HerdrClient;
|
||||
import dev.ltms.fleet.herdr.PaneLocator;
|
||||
import io.javalin.Javalin;
|
||||
import org.junit.jupiter.api.AfterEach;
|
||||
import org.junit.jupiter.api.Test;
|
||||
import org.junit.jupiter.api.io.TempDir;
|
||||
|
||||
import java.nio.file.Files;
|
||||
import java.nio.file.Path;
|
||||
import java.util.LinkedHashMap;
|
||||
import java.util.Map;
|
||||
import java.util.concurrent.CopyOnWriteArrayList;
|
||||
import java.util.concurrent.Executors;
|
||||
import java.util.concurrent.ScheduledExecutorService;
|
||||
import java.util.function.LongSupplier;
|
||||
|
||||
import static org.junit.jupiter.api.Assertions.assertEquals;
|
||||
import static org.junit.jupiter.api.Assertions.assertNull;
|
||||
|
||||
/**
|
||||
* fleetd #612 step 2, unit B2 (CB-185, identity half). Replaces the deleted
|
||||
* {@code FleetdConnectionIdentityConstructionTest}, which pinned this claim by reading {@code
|
||||
* Fleetd.java}'s source text for {@code "new PaneLocator(herdr, memberHerdr)"}. That claim moved
|
||||
* to {@code FleetdAssembly.java} (fleetd #612 Unit A) and is pinned here instead, by driving the
|
||||
* real {@link ConnectionIdentity} — via {@code runtime.mcp().identity()}, not a copy — that {@link
|
||||
* FleetdAssembly#assembleAndStart} built.
|
||||
*
|
||||
* <p><strong>What this guards against</strong> (from the deleted test's own javadoc): pinning
|
||||
* {@code PaneLocator} to {@code memberHerdr} alone leaves every LEAD's own MCP connection
|
||||
* unresolvable ({@code callerTerminal == null}) the moment {@code memberHerdrSocket} names a
|
||||
* second daemon, which breaks {@code fleet_reply}/{@code fleet_ask}/{@code fleet_whoami} for a
|
||||
* lead. {@code PaneLocatorTest} already proves {@link PaneLocator} itself can search two clients
|
||||
* given two — the gap this pins is that the assembly actually passes it two, and in the right
|
||||
* order (lead first).
|
||||
*
|
||||
* <p><strong>Why this cannot be driven through a real MCP/HTTP round trip.</strong> The natural
|
||||
* way to observe {@code ConnectionIdentity} would be a real {@code fleet_whoami} call over the
|
||||
* built {@code FleetMcp}, the way {@code FleetMcpContextExtractorTest} drives its own
|
||||
* hand-built one. That does not work for the REAL assembly, because {@code FleetdAssembly} wires
|
||||
* {@code ConnectionIdentity} with a hardcoded {@code new LsofPeerPidLookup()} (see {@code
|
||||
* FleetdAssembly.java:444}), and {@code LsofPeerPidLookup} explicitly excludes its own PID — see
|
||||
* its javadoc: "we exclude our own PID and take the other end". In a JUnit test the HTTP client
|
||||
* and the daemon under test run in the very same JVM, so the "client" and "server" ends of the
|
||||
* loopback connection ARE the same PID, and {@code pidForLocalPort} always returns {@code -1}
|
||||
* before {@link PaneLocator} is ever reached — proving nothing about which daemon(s) got searched.
|
||||
* This test instead reaches the real {@link PaneLocator} the assembly built (through {@link
|
||||
* ConnectionIdentity#panes()}, added for exactly this) and drives it with a chosen pid directly,
|
||||
* bypassing the OS-dependent PID lookup entirely — a legitimate substitute, since the pid lookup
|
||||
* is not what CB-185 is about.
|
||||
*/
|
||||
class FleetdAssemblyConnectionIdentityTest {
|
||||
|
||||
/** Same shape as {@code FleetdAssemblyLifecycleTest}'s fake, but keys {@code connectHerdr} by
|
||||
* socket path so the lead and member daemons can be two DIFFERENT {@link FakeHerdr}s. */
|
||||
private static final class TwoHerdrResourcePorts implements ResourcePorts {
|
||||
|
||||
final Map<Path, HerdrClient> herdrsBySocket = new LinkedHashMap<>();
|
||||
final CopyOnWriteArrayList<ScheduledExecutorService> schedulers = new CopyOnWriteArrayList<>();
|
||||
Runnable shutdownHook;
|
||||
|
||||
@Override
|
||||
public Map<String, String> environment() {
|
||||
return Map.of();
|
||||
}
|
||||
|
||||
@Override
|
||||
public HerdrClient connectHerdr(Path socketPath) {
|
||||
HerdrClient client = herdrsBySocket.get(socketPath);
|
||||
if (client == null) {
|
||||
throw new IllegalStateException("no fake herdr registered for socket " + socketPath);
|
||||
}
|
||||
return client;
|
||||
}
|
||||
|
||||
@Override
|
||||
public Fleetd.AmqpOpener replyInboxOpener() {
|
||||
return (uri, prefetch) -> new dev.ltms.fleet.msg.InMemoryReplyInbox();
|
||||
}
|
||||
|
||||
@Override
|
||||
public Fleetd.LeadMailboxOpener leadMailboxOpener() {
|
||||
return (uri, selfCoordId, prefetch) -> {
|
||||
throw new UnsupportedOperationException(
|
||||
"leadMailboxOpener must not be called — no coordinator: block is configured");
|
||||
};
|
||||
}
|
||||
|
||||
@Override
|
||||
public LongSupplier nanoClock() {
|
||||
return System::nanoTime;
|
||||
}
|
||||
|
||||
@Override
|
||||
public LongSupplier wallClockNanos() {
|
||||
return System::nanoTime;
|
||||
}
|
||||
|
||||
@Override
|
||||
public ScheduledExecutorService newScheduler(String purpose) {
|
||||
ScheduledExecutorService scheduler = Executors.newSingleThreadScheduledExecutor();
|
||||
schedulers.add(scheduler);
|
||||
return scheduler;
|
||||
}
|
||||
|
||||
@Override
|
||||
public void addShutdownHook(Runnable hook) {
|
||||
this.shutdownHook = hook;
|
||||
}
|
||||
|
||||
@Override
|
||||
public void startHttp(Javalin app, String host, int port) {
|
||||
// Deliberately never bind — this test never issues a real HTTP request.
|
||||
}
|
||||
|
||||
@Override
|
||||
public Runnable herdrPollWait() {
|
||||
// Never invoked: this test's FakeHerdr answers immediately, so awaitHerdr never polls.
|
||||
return () -> {
|
||||
throw new UnsupportedOperationException("herdrPollWait must not be called — herdr is healthy");
|
||||
};
|
||||
}
|
||||
}
|
||||
|
||||
private FleetdRuntime runtime;
|
||||
private TwoHerdrResourcePorts ports;
|
||||
|
||||
@AfterEach
|
||||
void tearDown() {
|
||||
if (ports != null && ports.shutdownHook != null) {
|
||||
ports.shutdownHook.run();
|
||||
}
|
||||
}
|
||||
|
||||
private static final Path LEAD_SOCKET = Path.of("/fake/lead-herdr.sock");
|
||||
private static final Path MEMBER_SOCKET = Path.of("/fake/member-herdr.sock");
|
||||
|
||||
private static FleetConfig writeConfig(Path dir) throws Exception {
|
||||
Path f = dir.resolve("fleetd.yaml");
|
||||
Files.writeString(f, """
|
||||
bind:
|
||||
host: 127.0.0.1
|
||||
port: 8765
|
||||
herdrSocket: "%s"
|
||||
memberHerdrSocket: "%s"
|
||||
lifecycle:
|
||||
idleTtlSeconds: 600
|
||||
health:
|
||||
enabled: false
|
||||
broker:
|
||||
uri: "amqp://fake-test-broker/vh"
|
||||
""".formatted(LEAD_SOCKET, MEMBER_SOCKET));
|
||||
return FleetConfig.load(f);
|
||||
}
|
||||
|
||||
private FleetdRuntime assemble(Path dir, FakeHerdr lead, FakeHerdr member) throws Exception {
|
||||
FleetConfig cfg = writeConfig(dir);
|
||||
ConfigRef config = new ConfigRef(dir.resolve("fleetd.yaml"), cfg);
|
||||
SubscriptionGuard guard = new SubscriptionGuard(cfg.guard().hostSet());
|
||||
ports = new TwoHerdrResourcePorts();
|
||||
ports.herdrsBySocket.put(LEAD_SOCKET, lead);
|
||||
ports.herdrsBySocket.put(MEMBER_SOCKET, member);
|
||||
runtime = FleetdAssembly.assembleAndStart(new AssemblyInputs(cfg, config, guard), ports);
|
||||
return runtime;
|
||||
}
|
||||
|
||||
/**
|
||||
* The pin. {@code lead} carries the one pane {@link FakeHerdr}'s canned {@code
|
||||
* pane.process_info} ties to {@link FakeHerdr#WORKER_PID} (pane {@code w2:p7}); {@code member}
|
||||
* reports NO panes at all ({@link FakeHerdr#withNoPanes()}) — modelling a second daemon that
|
||||
* simply does not host the caller's pane, exactly the CB-185 javadoc's scenario for a lead's
|
||||
* own connection. If {@code PaneLocator} only ever searches the member daemon (the bug), this
|
||||
* pid resolves to nothing, because the pane that owns it lives on the LEAD daemon the bug
|
||||
* skips.
|
||||
*/
|
||||
@Test
|
||||
void connectionIdentitySearchesTheLeadDaemonNotJustTheMemberOne(@TempDir Path dir) throws Exception {
|
||||
FakeHerdr lead = new FakeHerdr();
|
||||
FakeHerdr member = new FakeHerdr().withNoPanes();
|
||||
|
||||
assemble(dir, lead, member);
|
||||
|
||||
PaneLocator panes = runtime.mcp().identity().panes();
|
||||
PaneLocator.Lookup lookup = panes.terminalForPid(FakeHerdr.WORKER_PID);
|
||||
|
||||
assertEquals("term_a", lookup.terminal(),
|
||||
"the pane owning WORKER_PID lives on the LEAD daemon only (the member fake reports "
|
||||
+ "no panes) — PaneLocator must still find it, which is only possible if it "
|
||||
+ "searches the lead client and not just the member one");
|
||||
}
|
||||
|
||||
/**
|
||||
* The mirror control: when the pane instead lives ONLY on the member daemon (the lead reports
|
||||
* no panes), the lookup must still find it — proving the member client is genuinely searched
|
||||
* too, not merely tolerated as a second, always-losing argument.
|
||||
*/
|
||||
@Test
|
||||
void connectionIdentityAlsoSearchesTheMemberDaemon(@TempDir Path dir) throws Exception {
|
||||
FakeHerdr lead = new FakeHerdr().withNoPanes();
|
||||
FakeHerdr member = new FakeHerdr();
|
||||
|
||||
assemble(dir, lead, member);
|
||||
|
||||
PaneLocator panes = runtime.mcp().identity().panes();
|
||||
PaneLocator.Lookup lookup = panes.terminalForPid(FakeHerdr.WORKER_PID);
|
||||
|
||||
assertEquals("term_a", lookup.terminal(),
|
||||
"the pane owning WORKER_PID lives on the MEMBER daemon only — PaneLocator must "
|
||||
+ "find it there too");
|
||||
}
|
||||
|
||||
/** Sanity control: a pid nobody owns resolves to nothing on either daemon. */
|
||||
@Test
|
||||
void aPidNoPaneOwnsResolvesToNoTerminalOnEitherDaemon(@TempDir Path dir) throws Exception {
|
||||
FakeHerdr lead = new FakeHerdr();
|
||||
FakeHerdr member = new FakeHerdr();
|
||||
|
||||
assemble(dir, lead, member);
|
||||
|
||||
PaneLocator panes = runtime.mcp().identity().panes();
|
||||
PaneLocator.Lookup lookup = panes.terminalForPid(999_999L);
|
||||
|
||||
assertNull(lookup.terminal());
|
||||
}
|
||||
}
|
||||
@@ -0,0 +1,219 @@
|
||||
package dev.ltms.fleet;
|
||||
|
||||
import dev.ltms.fleet.config.ConfigRef;
|
||||
import dev.ltms.fleet.config.FleetConfig;
|
||||
import dev.ltms.fleet.guard.SubscriptionGuard;
|
||||
import dev.ltms.fleet.herdr.FakeHerdr;
|
||||
import dev.ltms.fleet.herdr.HerdrClient;
|
||||
import dev.ltms.fleet.msg.LeadChannel;
|
||||
import dev.ltms.fleet.msg.LeadChannelHandle;
|
||||
import dev.ltms.fleet.msg.LeadMessage;
|
||||
import dev.ltms.fleet.msg.ReplyInbox;
|
||||
import io.javalin.Javalin;
|
||||
import org.junit.jupiter.api.Test;
|
||||
import org.junit.jupiter.api.io.TempDir;
|
||||
|
||||
import java.nio.file.Files;
|
||||
import java.nio.file.Path;
|
||||
import java.util.List;
|
||||
import java.util.Map;
|
||||
import java.util.concurrent.Executors;
|
||||
import java.util.concurrent.ScheduledExecutorService;
|
||||
import java.util.function.LongSupplier;
|
||||
|
||||
import static org.junit.jupiter.api.Assertions.assertEquals;
|
||||
import static org.junit.jupiter.api.Assertions.assertFalse;
|
||||
import static org.junit.jupiter.api.Assertions.assertNotNull;
|
||||
import static org.junit.jupiter.api.Assertions.assertSame;
|
||||
import static org.junit.jupiter.api.Assertions.assertTrue;
|
||||
|
||||
/**
|
||||
* fleetd #612 A-gaps (gap 1): {@code FleetdAssemblyLifecycleTest}'s own class javadoc says plainly
|
||||
* that it leaves {@code coordinator:} unset, so {@code leadMailbox} and {@code leadCoordLoop} stay
|
||||
* {@code null} throughout — the configured-coordinator path is never exercised by Unit A's own
|
||||
* test. This class drives that path instead: a real {@code coordinator:} block, a fake {@link
|
||||
* Fleetd.LeadMailboxOpener} returning a fake closeable channel (never a real broker connection),
|
||||
* and proof that {@link FleetdAssembly#assembleAndStart} both builds it and, on shutdown, closes it.
|
||||
*
|
||||
* <p>Made possible by generalising {@code Fleetd.LeadMailboxOpener}'s return type (and {@code
|
||||
* FleetdRuntime}'s field) from the concrete {@code LeadMailbox} to {@link LeadChannelHandle} — a
|
||||
* {@link LeadChannel} its owner can also close. {@code FleetMcp} and {@code LeadCoordLoop} already
|
||||
* consumed the narrower {@link LeadChannel}; this only widens the one seam that owns and closes it.
|
||||
*/
|
||||
class FleetdAssemblyCoordinatorLifecycleTest {
|
||||
|
||||
private static final String SELF_COORD_ID = "test-lead";
|
||||
|
||||
/** A fake {@link LeadChannelHandle}: never touches a broker, and records whether it was closed. */
|
||||
private static final class FakeLeadChannel implements LeadChannelHandle {
|
||||
volatile boolean closed = false;
|
||||
|
||||
@Override
|
||||
public void publish(String toCoordId, LeadMessage m) {
|
||||
}
|
||||
|
||||
@Override
|
||||
public List<LeadMessage> peek() {
|
||||
return List.of();
|
||||
}
|
||||
|
||||
@Override
|
||||
public void ack(String msgId) {
|
||||
}
|
||||
|
||||
@Override
|
||||
public String selfCoordId() {
|
||||
return SELF_COORD_ID;
|
||||
}
|
||||
|
||||
@Override
|
||||
public boolean heldDurable() {
|
||||
return true;
|
||||
}
|
||||
|
||||
@Override
|
||||
public MailboxState inspect(String coordId) {
|
||||
return MailboxState.unknown(coordId);
|
||||
}
|
||||
|
||||
@Override
|
||||
public void close() {
|
||||
closed = true;
|
||||
}
|
||||
}
|
||||
|
||||
/** Minimal fake {@link ResourcePorts}: a real herdr fake, a fake reply inbox, and a real, offered fake lead channel. */
|
||||
private static final class RecordingResourcePorts implements ResourcePorts {
|
||||
|
||||
final FakeHerdr herdr = new FakeHerdr();
|
||||
final SentinelReplyInbox replyInbox = new SentinelReplyInbox();
|
||||
final FakeLeadChannel leadChannel = new FakeLeadChannel();
|
||||
String offeredUri;
|
||||
String offeredSelfId;
|
||||
Runnable shutdownHook;
|
||||
|
||||
@Override
|
||||
public Map<String, String> environment() {
|
||||
return Map.of();
|
||||
}
|
||||
|
||||
@Override
|
||||
public HerdrClient connectHerdr(Path socketPath) {
|
||||
return herdr;
|
||||
}
|
||||
|
||||
@Override
|
||||
public Fleetd.AmqpOpener replyInboxOpener() {
|
||||
return (uri, prefetch) -> replyInbox;
|
||||
}
|
||||
|
||||
@Override
|
||||
public Fleetd.LeadMailboxOpener leadMailboxOpener() {
|
||||
return (uri, selfCoordId, prefetch) -> {
|
||||
offeredUri = uri;
|
||||
offeredSelfId = selfCoordId;
|
||||
return leadChannel;
|
||||
};
|
||||
}
|
||||
|
||||
@Override
|
||||
public LongSupplier nanoClock() {
|
||||
return System::nanoTime;
|
||||
}
|
||||
|
||||
@Override
|
||||
public LongSupplier wallClockNanos() {
|
||||
return System::nanoTime;
|
||||
}
|
||||
|
||||
@Override
|
||||
public ScheduledExecutorService newScheduler(String purpose) {
|
||||
return Executors.newSingleThreadScheduledExecutor();
|
||||
}
|
||||
|
||||
@Override
|
||||
public void addShutdownHook(Runnable hook) {
|
||||
this.shutdownHook = hook;
|
||||
}
|
||||
|
||||
@Override
|
||||
public void startHttp(Javalin app, String host, int port) {
|
||||
// No real HTTP bind in a unit test.
|
||||
}
|
||||
|
||||
@Override
|
||||
public Runnable herdrPollWait() {
|
||||
// Never invoked: this test's FakeHerdr answers immediately, so awaitHerdr never polls.
|
||||
return () -> {
|
||||
throw new UnsupportedOperationException("herdrPollWait must not be called — herdr is healthy");
|
||||
};
|
||||
}
|
||||
}
|
||||
|
||||
private static final class SentinelReplyInbox implements ReplyInbox {
|
||||
@Override
|
||||
public void own(String target) {
|
||||
}
|
||||
|
||||
@Override
|
||||
public void release(String target) {
|
||||
}
|
||||
|
||||
@Override
|
||||
public void publish(String target, String msgId, String content) {
|
||||
}
|
||||
|
||||
@Override
|
||||
public List<InboxMessage> peek(String target) {
|
||||
return List.of();
|
||||
}
|
||||
|
||||
@Override
|
||||
public boolean ack(String target, String msgId) {
|
||||
return false;
|
||||
}
|
||||
}
|
||||
|
||||
private static FleetConfig writeConfig(Path dir) throws Exception {
|
||||
Path f = dir.resolve("fleetd.yaml");
|
||||
Files.writeString(f, """
|
||||
bind:
|
||||
host: 127.0.0.1
|
||||
port: 8765
|
||||
idleSleepGuard:
|
||||
enabled: false
|
||||
coordinator:
|
||||
uri: "amqp://fake-lead-broker/vh"
|
||||
selfId: "test-lead"
|
||||
""");
|
||||
return FleetConfig.load(f);
|
||||
}
|
||||
|
||||
@Test
|
||||
void configuredCoordinatorIsBuiltByTheAssemblyAndClosedOnShutdown(@TempDir Path dir) throws Exception {
|
||||
FleetConfig cfg = writeConfig(dir);
|
||||
ConfigRef config = new ConfigRef(dir.resolve("fleetd.yaml"), cfg);
|
||||
SubscriptionGuard guard = new SubscriptionGuard(cfg.guard().hostSet());
|
||||
RecordingResourcePorts ports = new RecordingResourcePorts();
|
||||
|
||||
FleetdRuntime runtime = FleetdAssembly.assembleAndStart(new AssemblyInputs(cfg, config, guard), ports);
|
||||
|
||||
// --- the assembly actually calls the configured opener and builds the coordinator path ---
|
||||
assertEquals("amqp://fake-lead-broker/vh", ports.offeredUri,
|
||||
"the assembly must open the mailbox at the configured broker uri");
|
||||
assertEquals(SELF_COORD_ID, ports.offeredSelfId,
|
||||
"the assembly must open the mailbox under the configured selfId");
|
||||
assertSame(ports.leadChannel, runtime.leadMailbox(),
|
||||
"FleetdRuntime must own the exact LeadChannelHandle the opener returned, not a copy");
|
||||
assertNotNull(runtime.leadCoordLoop(),
|
||||
"a configured coordinator: block must build the receiving LeadCoordLoop too");
|
||||
assertFalse(ports.leadChannel.closed, "the channel must still be open while the daemon is running");
|
||||
|
||||
// --- shutting the assembly down closes it -------------------------------------------------
|
||||
assertNotNull(ports.shutdownHook, "FleetdAssembly must have registered a shutdown hook");
|
||||
ports.shutdownHook.run();
|
||||
|
||||
assertTrue(ports.leadChannel.closed,
|
||||
"FleetdRuntime.close() must close the configured LeadChannelHandle");
|
||||
}
|
||||
}
|
||||
@@ -0,0 +1,278 @@
|
||||
package dev.ltms.fleet;
|
||||
|
||||
import dev.ltms.fleet.config.ConfigRef;
|
||||
import dev.ltms.fleet.config.FleetConfig;
|
||||
import dev.ltms.fleet.guard.SubscriptionGuard;
|
||||
import dev.ltms.fleet.herdr.FakeHerdr;
|
||||
import dev.ltms.fleet.herdr.HerdrClient;
|
||||
import io.javalin.Javalin;
|
||||
import org.junit.jupiter.api.AfterEach;
|
||||
import org.junit.jupiter.api.Test;
|
||||
import org.junit.jupiter.api.Timeout;
|
||||
import org.junit.jupiter.api.io.TempDir;
|
||||
|
||||
import java.net.URI;
|
||||
import java.net.http.HttpClient;
|
||||
import java.net.http.HttpRequest;
|
||||
import java.net.http.HttpResponse;
|
||||
import java.nio.file.Files;
|
||||
import java.nio.file.Path;
|
||||
import java.util.LinkedHashMap;
|
||||
import java.util.Map;
|
||||
import java.util.concurrent.CopyOnWriteArrayList;
|
||||
import java.util.concurrent.Executors;
|
||||
import java.util.concurrent.ScheduledExecutorService;
|
||||
import java.util.concurrent.TimeUnit;
|
||||
import java.util.concurrent.atomic.AtomicLong;
|
||||
import java.util.function.LongSupplier;
|
||||
|
||||
import static org.junit.jupiter.api.Assertions.assertEquals;
|
||||
|
||||
/**
|
||||
* fleetd #612 step 2, unit B2 (CB-185, {@code FleetApp} half). Replaces the deleted {@code
|
||||
* FleetdFleetAppConstructionTest}, which pinned this claim by reading {@code Fleetd.java}'s
|
||||
* source text for {@code "new FleetApp(herdr, memberHerdr, workers,"}. That claim moved to {@code
|
||||
* FleetdAssembly.java} (fleetd #612 Unit A) and is pinned here instead, by driving the real {@code
|
||||
* Javalin} app — via {@code runtime.app()}, not a copy — that {@link
|
||||
* FleetdAssembly#assembleAndStart} built and handed to {@link FleetdRuntime}.
|
||||
*
|
||||
* <p><strong>What this guards against</strong> (from the deleted test's own javadoc): constructing
|
||||
* {@code FleetApp} with the lead-only {@code herdr} client (dropping {@code memberHerdr}) makes
|
||||
* {@code GET /healthz} report green while the MEMBER daemon is down — so every spawn fails
|
||||
* invisibly — and silently drops every member workspace from {@code GET /sessions}. {@code
|
||||
* FleetAppTwoDaemonTest} already proves {@code FleetApp} itself merges/gates correctly given two
|
||||
* clients; the gap this pins is that the assembly actually passes it two.
|
||||
*
|
||||
* <p><strong>Both directions, not just one</strong> (fleetd #612 issue comment 17525): the deleted
|
||||
* guard's positive assertion required the exact pair {@code "new FleetApp(herdr, memberHerdr,
|
||||
* workers,"}, which does not survive EITHER daemon being dropped. An earlier version of this class
|
||||
* only proved the member-dropped direction, which left {@code new FleetApp(memberHerdr,
|
||||
* memberHerdr, ...)} — the symmetric bug, {@code /healthz} green while the LEAD daemon is down —
|
||||
* an undetected regression. {@link #healthzGoesRedWhenTheLeadDaemonIsDownEvenThoughTheMemberIsUp}
|
||||
* closes that.
|
||||
*
|
||||
* <p>Unlike the {@code ConnectionIdentity} half of CB-185 ({@code
|
||||
* FleetdAssemblyConnectionIdentityTest}), {@code /healthz} needs no caller identity at all, so
|
||||
* this test can bind {@link FleetdRuntime#app()} to a REAL ephemeral port (exactly {@code
|
||||
* FleetAppTwoDaemonTest} does for its own hand-built {@code FleetApp}) and drive it with a real
|
||||
* {@code HttpClient} — no accessor needed for this half.
|
||||
*
|
||||
* <p><strong>{@code GET /sessions} could not be driven the same way</strong>, so this class does
|
||||
* not pin the merge half of the deleted test's javadoc. This class configures no {@code auth:}
|
||||
* block, so it runs under the default {@code loopback-trust} mode ({@code FleetConfig}). Under
|
||||
* that mode, {@code /sessions} requires {@code Authz.Action.READ}, which — through the REAL
|
||||
* assembly's real {@code CallerResolver}/{@code ConnectionIdentity} (built with a hardcoded
|
||||
* {@code new LsofPeerPidLookup()}) — needs {@code Caller.resolved()}, i.e. a real positive pid
|
||||
* from {@code lsof}. {@code LsofPeerPidLookup} excludes its own pid (see its javadoc), and a
|
||||
* JUnit test's HTTP client and the daemon under test share one JVM pid, so the resolved pid is
|
||||
* always {@code -1} and every such request is refused as {@code ANONYMOUS} (fleetd #317's
|
||||
* fail-closed rule) before the route handler — and its {@code memberHerdr} merge — is ever
|
||||
* reached. Verified directly: driving {@code GET /sessions} here returns {@code 401
|
||||
* unauthenticated}, not the merged body. {@code FleetAppTwoDaemonTest} avoids this because it
|
||||
* builds {@code FleetApp} with {@code callers: null}, which is not what the real assembly
|
||||
* passes. The {@code /healthz} pin below is what this class relies on for CB-185's {@code
|
||||
* FleetApp} half; {@code FleetAppTwoDaemonTest} remains the full behavioural proof that
|
||||
* {@code FleetApp} itself merges {@code /sessions} correctly once handed two clients.
|
||||
*
|
||||
* <p><strong>This refusal is {@code loopback-trust}-specific, not a property of {@code
|
||||
* CallerResolver} in general.</strong> Under {@code auth.mode: token}, {@code
|
||||
* CallerResolver#resolve} returns before ever consulting {@code Caller.resolved()} or {@code
|
||||
* Caller.scanComplete()}: a request carrying a valid bearer token in its {@code Authorization}
|
||||
* header resolves to {@code Role#PRIMARY} with no pid lookup at all, so the same-JVM-pid
|
||||
* exclusion above never comes into play. {@code FleetdQuarantineOutageDualWindowAssemblyTest}
|
||||
* and {@code FleetdListReportingSourcesAssemblyTest} both drive {@code Authz.Action.READ} this
|
||||
* way, over a real {@code McpSyncClient}/{@code HttpClient} against a real {@code
|
||||
* FleetdAssembly#assembleAndStart}, and both get the real response rather than a refusal.
|
||||
*
|
||||
* <p><strong>fleetd #629 follow-up.</strong> The fix below (see {@link TwoHerdrResourcePorts})
|
||||
* makes {@link #healthzGoesRedWhenTheLeadDaemonIsDownEvenThoughTheMemberIsUp}'s fake {@code
|
||||
* nanoClock()} frozen unless {@code herdrPollWait()} itself advances it. That is a sharper pin
|
||||
* than an assertion — if a future edit to {@code FleetdAssembly} ever bypasses {@code
|
||||
* ports.herdrPollWait()} again (e.g. reverting to a hardcoded {@code Thread.sleep}), the clock
|
||||
* never advances, {@code Fleetd#awaitHerdr}'s deadline is never reached, and this test hangs
|
||||
* forever instead of failing — proven by deliberately reintroducing that exact regression while
|
||||
* fixing this ticket. {@code @Timeout} turns that silent hang into a bounded, named test failure:
|
||||
* {@code SEPARATE_THREAD} so JUnit's timeout governor can actually interrupt a thread stuck in a
|
||||
* real {@code Thread.sleep} loop (the default {@code SAME_THREAD} mode cannot — it only measures
|
||||
* elapsed time after the test method returns on its own, which never happens here). 10 seconds is
|
||||
* roughly 150x the real passing times measured here (~0.06s), so a slow CI machine has no reason
|
||||
* to flake, and it is still 3x faster than discovering the regression by burning a CI job's whole
|
||||
* wall-clock budget.
|
||||
*/
|
||||
@Timeout(value = 10, unit = TimeUnit.SECONDS, threadMode = Timeout.ThreadMode.SEPARATE_THREAD)
|
||||
class FleetdAssemblyFleetAppTest {
|
||||
|
||||
private static final class TwoHerdrResourcePorts implements ResourcePorts {
|
||||
|
||||
final Map<Path, HerdrClient> herdrsBySocket = new LinkedHashMap<>();
|
||||
final CopyOnWriteArrayList<ScheduledExecutorService> schedulers = new CopyOnWriteArrayList<>();
|
||||
// fleetd #629: a fake, advanceable clock — NOT System::nanoTime. awaitHerdr's poll wait
|
||||
// (herdrPollWait() below) advances this on every poll instead of sleeping for real, so the
|
||||
// down-lead test below reaches awaitHerdr's deadline without burning real wall-clock time.
|
||||
final AtomicLong nowNanos = new AtomicLong(1_000_000_000L); // arbitrary non-zero start
|
||||
Runnable shutdownHook;
|
||||
|
||||
@Override
|
||||
public Map<String, String> environment() {
|
||||
return Map.of();
|
||||
}
|
||||
|
||||
@Override
|
||||
public HerdrClient connectHerdr(Path socketPath) {
|
||||
HerdrClient client = herdrsBySocket.get(socketPath);
|
||||
if (client == null) {
|
||||
throw new IllegalStateException("no fake herdr registered for socket " + socketPath);
|
||||
}
|
||||
return client;
|
||||
}
|
||||
|
||||
@Override
|
||||
public Fleetd.AmqpOpener replyInboxOpener() {
|
||||
return (uri, prefetch) -> new dev.ltms.fleet.msg.InMemoryReplyInbox();
|
||||
}
|
||||
|
||||
@Override
|
||||
public Fleetd.LeadMailboxOpener leadMailboxOpener() {
|
||||
return (uri, selfCoordId, prefetch) -> {
|
||||
throw new UnsupportedOperationException(
|
||||
"leadMailboxOpener must not be called — no coordinator: block is configured");
|
||||
};
|
||||
}
|
||||
|
||||
@Override
|
||||
public LongSupplier nanoClock() {
|
||||
return nowNanos::get;
|
||||
}
|
||||
|
||||
@Override
|
||||
public LongSupplier wallClockNanos() {
|
||||
return System::nanoTime;
|
||||
}
|
||||
|
||||
@Override
|
||||
public Runnable herdrPollWait() {
|
||||
// fleetd #629: advance the fake clock instead of a real Thread.sleep, so awaitHerdr's
|
||||
// deadline is reached in real time regardless of the configured poll interval.
|
||||
return () -> nowNanos.addAndGet(TimeUnit.SECONDS.toNanos(1));
|
||||
}
|
||||
|
||||
@Override
|
||||
public ScheduledExecutorService newScheduler(String purpose) {
|
||||
ScheduledExecutorService scheduler = Executors.newSingleThreadScheduledExecutor();
|
||||
schedulers.add(scheduler);
|
||||
return scheduler;
|
||||
}
|
||||
|
||||
@Override
|
||||
public void addShutdownHook(Runnable hook) {
|
||||
this.shutdownHook = hook;
|
||||
}
|
||||
|
||||
@Override
|
||||
public void startHttp(Javalin app, String host, int port) {
|
||||
// Deliberately never bind here — this test binds runtime.app() itself, for real, below.
|
||||
}
|
||||
}
|
||||
|
||||
private static final Path LEAD_SOCKET = Path.of("/fake/lead-herdr.sock");
|
||||
private static final Path MEMBER_SOCKET = Path.of("/fake/member-herdr.sock");
|
||||
|
||||
private final HttpClient http = HttpClient.newHttpClient();
|
||||
private FleetdRuntime runtime;
|
||||
private TwoHerdrResourcePorts ports;
|
||||
private Javalin boundApp;
|
||||
|
||||
@AfterEach
|
||||
void tearDown() {
|
||||
if (boundApp != null) {
|
||||
boundApp.stop();
|
||||
}
|
||||
if (ports != null && ports.shutdownHook != null) {
|
||||
ports.shutdownHook.run();
|
||||
}
|
||||
}
|
||||
|
||||
private static FleetConfig writeConfig(Path dir) throws Exception {
|
||||
Path f = dir.resolve("fleetd.yaml");
|
||||
Files.writeString(f, """
|
||||
bind:
|
||||
host: 127.0.0.1
|
||||
port: 8765
|
||||
herdrSocket: "%s"
|
||||
memberHerdrSocket: "%s"
|
||||
lifecycle:
|
||||
idleTtlSeconds: 600
|
||||
health:
|
||||
enabled: false
|
||||
broker:
|
||||
uri: "amqp://fake-test-broker/vh"
|
||||
""".formatted(LEAD_SOCKET, MEMBER_SOCKET));
|
||||
return FleetConfig.load(f);
|
||||
}
|
||||
|
||||
/** Assembles the real graph, then binds the real {@code Javalin app} to an ephemeral port. */
|
||||
private int assembleAndBind(Path dir, FakeHerdr lead, FakeHerdr member) throws Exception {
|
||||
FleetConfig cfg = writeConfig(dir);
|
||||
ConfigRef config = new ConfigRef(dir.resolve("fleetd.yaml"), cfg);
|
||||
SubscriptionGuard guard = new SubscriptionGuard(cfg.guard().hostSet());
|
||||
ports = new TwoHerdrResourcePorts();
|
||||
ports.herdrsBySocket.put(LEAD_SOCKET, lead);
|
||||
ports.herdrsBySocket.put(MEMBER_SOCKET, member);
|
||||
runtime = FleetdAssembly.assembleAndStart(new AssemblyInputs(cfg, config, guard), ports);
|
||||
boundApp = runtime.app().start("127.0.0.1", 0);
|
||||
return boundApp.port();
|
||||
}
|
||||
|
||||
private HttpResponse<String> get(int port, String path) throws Exception {
|
||||
HttpRequest req = HttpRequest.newBuilder(URI.create("http://127.0.0.1:" + port + path)).GET().build();
|
||||
return http.send(req, HttpResponse.BodyHandlers.ofString());
|
||||
}
|
||||
|
||||
/**
|
||||
* The pin. The MEMBER daemon is down; the LEAD daemon is healthy. If the assembly built
|
||||
* {@code FleetApp} with only the lead client (the bug: passing {@code herdr} where {@code
|
||||
* memberHerdr} is expected), the down member is invisible and {@code /healthz} stays 200.
|
||||
*/
|
||||
@Test
|
||||
void healthzGoesRedWhenTheMemberDaemonIsDownEvenThoughTheLeadIsUp(@TempDir Path dir) throws Exception {
|
||||
FakeHerdr lead = new FakeHerdr();
|
||||
FakeHerdr member = new FakeHerdr().healthy(false);
|
||||
|
||||
int port = assembleAndBind(dir, lead, member);
|
||||
|
||||
HttpResponse<String> res = get(port, "/healthz");
|
||||
assertEquals(503, res.statusCode(),
|
||||
"a down MEMBER daemon must not be masked by a healthy lead: " + res.body());
|
||||
}
|
||||
|
||||
/** Sanity control: both daemons healthy must still be green through the real assembly. */
|
||||
@Test
|
||||
void healthzIsGreenWhenBothDaemonsAreUp(@TempDir Path dir) throws Exception {
|
||||
FakeHerdr lead = new FakeHerdr();
|
||||
FakeHerdr member = new FakeHerdr();
|
||||
|
||||
int port = assembleAndBind(dir, lead, member);
|
||||
|
||||
assertEquals(200, get(port, "/healthz").statusCode());
|
||||
}
|
||||
|
||||
/**
|
||||
* The symmetric pin (fleetd #612 issue comment 17525): the LEAD daemon is down; the MEMBER
|
||||
* daemon is healthy. If the assembly built {@code FleetApp} with only the member client
|
||||
* (dropping {@code herdr} — the mirror of the bug above, {@code new FleetApp(memberHerdr,
|
||||
* memberHerdr, ...)}), the down LEAD is invisible and {@code /healthz} stays 200. Without this
|
||||
* case the pair above is one-directional and does not cover the deleted guard's positive
|
||||
* assertion (it required BOTH {@code herdr,} and {@code memberHerdr,} in that order).
|
||||
*/
|
||||
@Test
|
||||
void healthzGoesRedWhenTheLeadDaemonIsDownEvenThoughTheMemberIsUp(@TempDir Path dir) throws Exception {
|
||||
FakeHerdr lead = new FakeHerdr().healthy(false);
|
||||
FakeHerdr member = new FakeHerdr();
|
||||
|
||||
int port = assembleAndBind(dir, lead, member);
|
||||
|
||||
HttpResponse<String> res = get(port, "/healthz");
|
||||
assertEquals(503, res.statusCode(),
|
||||
"a down LEAD daemon must not be masked by a healthy member: " + res.body());
|
||||
}
|
||||
}
|
||||
+198
@@ -0,0 +1,198 @@
|
||||
package dev.ltms.fleet;
|
||||
|
||||
import dev.ltms.fleet.config.ConfigRef;
|
||||
import dev.ltms.fleet.config.FleetConfig;
|
||||
import dev.ltms.fleet.guard.SubscriptionGuard;
|
||||
import dev.ltms.fleet.health.FleetHealthMonitor;
|
||||
import dev.ltms.fleet.herdr.FakeHerdr;
|
||||
import dev.ltms.fleet.herdr.HerdrClient;
|
||||
import dev.ltms.fleet.msg.MessageService;
|
||||
import dev.ltms.fleet.msg.ReplyInbox;
|
||||
import io.javalin.Javalin;
|
||||
import org.junit.jupiter.api.DisplayName;
|
||||
import org.junit.jupiter.api.Test;
|
||||
import org.junit.jupiter.api.io.TempDir;
|
||||
|
||||
import java.lang.reflect.Field;
|
||||
import java.nio.file.Files;
|
||||
import java.nio.file.Path;
|
||||
import java.util.List;
|
||||
import java.util.Map;
|
||||
import java.util.concurrent.Executors;
|
||||
import java.util.concurrent.ScheduledExecutorService;
|
||||
import java.util.function.BiConsumer;
|
||||
import java.util.function.LongSupplier;
|
||||
|
||||
import static org.junit.jupiter.api.Assertions.assertEquals;
|
||||
import static org.junit.jupiter.api.Assertions.assertNotNull;
|
||||
import static org.junit.jupiter.api.Assertions.assertTrue;
|
||||
|
||||
/**
|
||||
* fleetd #612 rank 7 — {@code FleetdAssembly.java:429} wires {@link FleetHealthMonitor}'s {@code
|
||||
* failTarget} callback with {@code Fleetd.healthFailTarget(messages)}. {@link
|
||||
* FleetdHealthFailTargetWiringTest} already pins that the FACTORY itself delegates to {@code
|
||||
* messages::abandon}, but it calls {@code Fleetd.healthFailTarget} directly — it never drives {@code
|
||||
* FleetdAssembly.assembleAndStart} and so cannot see whether the real call site at {@code :429}
|
||||
* still passes it the real, assembled {@link MessageService}. Swapping that argument for a no-op
|
||||
* {@code (a, b) -> {}} compiles clean and leaves the whole suite — including the factory-level test
|
||||
* — green: a dead member's waiting ticket then sits {@code PENDING} for the full 30-minute async
|
||||
* timeout instead of failing immediately.
|
||||
*
|
||||
* <p>This test assembles the real daemon with {@code health.enabled: true}, pulls the REAL {@code
|
||||
* failTarget} {@link BiConsumer} out of the REAL, assembled {@link FleetHealthMonitor} (via
|
||||
* reflection — the field is package-private to {@code dev.ltms.fleet.health}, and nothing public
|
||||
* exposes it; {@code StatusPollerResilienceTest} already uses the same technique in this suite), and
|
||||
* invokes it directly against the REAL {@link MessageService} {@link FleetdRuntime#messages()}
|
||||
* returns. A no-op lambda swapped in at the call site leaves the ticket {@code PENDING} forever,
|
||||
* which this test catches; the real one fails it.
|
||||
*/
|
||||
class FleetdAssemblyHealthFailTargetBehaviouralTest {
|
||||
|
||||
private static final String TARGET = "term_a";
|
||||
|
||||
private static final class RecordingResourcePorts implements ResourcePorts {
|
||||
final FakeHerdr herdr = new FakeHerdr();
|
||||
Runnable shutdownHook;
|
||||
|
||||
@Override
|
||||
public Map<String, String> environment() {
|
||||
return Map.of();
|
||||
}
|
||||
|
||||
@Override
|
||||
public HerdrClient connectHerdr(Path socketPath) {
|
||||
return herdr;
|
||||
}
|
||||
|
||||
@Override
|
||||
public Fleetd.AmqpOpener replyInboxOpener() {
|
||||
return (uri, prefetch) -> {
|
||||
throw new UnsupportedOperationException("no broker: block is configured");
|
||||
};
|
||||
}
|
||||
|
||||
@Override
|
||||
public Fleetd.LeadMailboxOpener leadMailboxOpener() {
|
||||
return (uri, selfCoordId, prefetch) -> {
|
||||
throw new UnsupportedOperationException("no coordinator: block is configured");
|
||||
};
|
||||
}
|
||||
|
||||
@Override
|
||||
public Runnable herdrPollWait() {
|
||||
// Never invoked: this test's FakeHerdr answers immediately, so awaitHerdr never polls.
|
||||
return () -> {
|
||||
throw new UnsupportedOperationException("herdrPollWait must not be called — herdr is healthy");
|
||||
};
|
||||
}
|
||||
@Override
|
||||
public LongSupplier nanoClock() {
|
||||
return System::nanoTime;
|
||||
}
|
||||
|
||||
@Override
|
||||
public LongSupplier wallClockNanos() {
|
||||
return System::nanoTime;
|
||||
}
|
||||
|
||||
@Override
|
||||
public ScheduledExecutorService newScheduler(String purpose) {
|
||||
return Executors.newSingleThreadScheduledExecutor();
|
||||
}
|
||||
|
||||
@Override
|
||||
public void addShutdownHook(Runnable hook) {
|
||||
this.shutdownHook = hook;
|
||||
}
|
||||
|
||||
@Override
|
||||
public void startHttp(Javalin app, String host, int port) {
|
||||
}
|
||||
}
|
||||
|
||||
private static FleetConfig writeConfig(Path dir) throws Exception {
|
||||
Path f = dir.resolve("fleetd.yaml");
|
||||
Files.writeString(f, """
|
||||
bind:
|
||||
host: 127.0.0.1
|
||||
port: 8765
|
||||
idleSleepGuard:
|
||||
enabled: false
|
||||
health:
|
||||
enabled: true
|
||||
intervalSeconds: 30
|
||||
""");
|
||||
return FleetConfig.load(f);
|
||||
}
|
||||
|
||||
@SuppressWarnings("unchecked")
|
||||
@Test
|
||||
@DisplayName("[BEHAVIOURAL] the real assembled FleetHealthMonitor's failTarget reaches the real "
|
||||
+ "MessageService.abandon, not a no-op")
|
||||
void assembledHealthFailTargetReachesRealMessages(@TempDir Path dir) throws Exception {
|
||||
FleetConfig cfg = writeConfig(dir);
|
||||
ConfigRef config = new ConfigRef(dir.resolve("fleetd.yaml"), cfg);
|
||||
SubscriptionGuard guard = new SubscriptionGuard(cfg.guard().hostSet());
|
||||
RecordingResourcePorts ports = new RecordingResourcePorts();
|
||||
|
||||
FleetdRuntime runtime = FleetdAssembly.assembleAndStart(new AssemblyInputs(cfg, config, guard), ports);
|
||||
assertNotNull(ports.shutdownHook, "FleetdAssembly must have registered a shutdown hook");
|
||||
|
||||
// Surefire runs the whole suite in one JVM fork, so the scheduler/loops this assembly starts
|
||||
// (SessionReaper, StatusPoller, the health monitor) must be torn down here, on the failure
|
||||
// path too — hence the try/finally, not just a statement at the end of the happy path.
|
||||
try {
|
||||
FleetHealthMonitor healthMonitor = runtime.healthMonitor();
|
||||
assertNotNull(healthMonitor, "health.enabled: true in this test's config, so "
|
||||
+ "FleetdAssembly.assembleAndStart must have built a real FleetHealthMonitor");
|
||||
|
||||
Field field = FleetHealthMonitor.class.getDeclaredField("failTarget");
|
||||
field.setAccessible(true);
|
||||
BiConsumer<String, String> failTarget = (BiConsumer<String, String>) field.get(healthMonitor);
|
||||
assertNotNull(failTarget, "FleetHealthMonitor's failTarget must never be null — the "
|
||||
+ "constructor itself requires it");
|
||||
|
||||
MessageService messages = runtime.messages();
|
||||
|
||||
// --- loud control: prove the assembled MessageService is actually wired up and a ticket is
|
||||
// genuinely PENDING before failTarget ever runs. If this fails, the test below would pass
|
||||
// vacuously on a MessageService that never got a ticket in the first place. TARGET has no
|
||||
// live agent behind it (no session was ever acquired), so nothing resolves this ticket on
|
||||
// its own — it stays PENDING until failTarget (or a timeout) ends it.
|
||||
String ticket = messages.sendAsync(TARGET, "long task");
|
||||
MessageService.TaskView before = messages.poll(ticket);
|
||||
assertEquals(MessageService.Phase.PENDING, before.phase(),
|
||||
"control: the async ticket must be PENDING before failTarget runs");
|
||||
|
||||
failTarget.accept(TARGET, "member unreachable (health monitor)");
|
||||
|
||||
MessageService.TaskView after = awaitTerminal(messages, ticket);
|
||||
assertEquals(MessageService.Phase.FAILED, after.phase(),
|
||||
"FleetdAssembly.java:429 must pass Fleetd.healthFailTarget(messages) built from the "
|
||||
+ "SAME assembled MessageService — a no-op BiConsumer at that call site leaves "
|
||||
+ "this ticket PENDING for the full 30-minute async timeout instead of failing it");
|
||||
assertTrue(after.detail() != null && after.detail().contains("member unreachable"),
|
||||
"the failure reason passed to failTarget.accept must reach MessageService.abandon and "
|
||||
+ "end up in the ticket's detail");
|
||||
} finally {
|
||||
// Proof the teardown actually ran, not just an assurance that a finally was added: the
|
||||
// captured shutdown hook's close order (FleetdAssemblyLifecycleTest) closes the herdr
|
||||
// client last, so ports.herdr.closed flips to true only if this hook really executed.
|
||||
ports.shutdownHook.run();
|
||||
assertTrue(ports.herdr.closed, "the captured shutdown hook must have run and closed herdr — "
|
||||
+ "proof this test's assembled background loops/scheduler were torn down");
|
||||
}
|
||||
}
|
||||
|
||||
private static MessageService.TaskView awaitTerminal(MessageService messages, String ticket)
|
||||
throws InterruptedException {
|
||||
long deadline = System.currentTimeMillis() + 5000;
|
||||
MessageService.TaskView view = messages.poll(ticket);
|
||||
while (view.phase() == MessageService.Phase.PENDING && System.currentTimeMillis() < deadline) {
|
||||
//noinspection BusyWait
|
||||
Thread.sleep(10);
|
||||
view = messages.poll(ticket);
|
||||
}
|
||||
return view;
|
||||
}
|
||||
}
|
||||
@@ -0,0 +1,201 @@
|
||||
package dev.ltms.fleet;
|
||||
|
||||
import dev.ltms.fleet.config.ConfigRef;
|
||||
import dev.ltms.fleet.config.FleetConfig;
|
||||
import dev.ltms.fleet.guard.SubscriptionGuard;
|
||||
import dev.ltms.fleet.herdr.FakeHerdr;
|
||||
import dev.ltms.fleet.herdr.HerdrClient;
|
||||
import dev.ltms.fleet.herdr.LeadTabScanner;
|
||||
import dev.ltms.fleet.msg.LeadChannelHandle;
|
||||
import dev.ltms.fleet.msg.LeadCoordLoop;
|
||||
import dev.ltms.fleet.msg.LeadMessage;
|
||||
import dev.ltms.fleet.msg.ReplyInbox;
|
||||
import io.javalin.Javalin;
|
||||
import org.junit.jupiter.api.Test;
|
||||
import org.junit.jupiter.api.io.TempDir;
|
||||
|
||||
import java.lang.reflect.Field;
|
||||
import java.nio.file.Files;
|
||||
import java.nio.file.Path;
|
||||
import java.util.List;
|
||||
import java.util.Map;
|
||||
import java.util.Set;
|
||||
import java.util.concurrent.Executors;
|
||||
import java.util.concurrent.ScheduledExecutorService;
|
||||
import java.util.function.LongSupplier;
|
||||
import java.util.function.Supplier;
|
||||
|
||||
import static org.junit.jupiter.api.Assertions.assertInstanceOf;
|
||||
import static org.junit.jupiter.api.Assertions.assertNotNull;
|
||||
import static org.junit.jupiter.api.Assertions.assertTrue;
|
||||
|
||||
/**
|
||||
* fleetd #670 — pins the {@code excludedWorkspaceLabels} argument {@link FleetdAssembly}'s
|
||||
* production boot path passes to {@link LeadTabScanner} at {@code FleetdAssembly.java:265}
|
||||
* ({@code Set.of()}).
|
||||
*
|
||||
* <p>{@code LeadTabScannerTest} already covers this constructor parameter, but it builds its own
|
||||
* {@link LeadTabScanner} with its own set, so it tests the seam and proves nothing about the
|
||||
* producer. This test instead reaches the exact object {@link FleetdAssembly#assembleAndStart}
|
||||
* builds: a {@code fleet.leaders:} block makes the assembly construct a real
|
||||
* {@link LeadTabScanner} for its local {@code leads} supplier, and a {@code coordinator:} block
|
||||
* makes it hand that same supplier instance to {@link LeadCoordLoop} (fleetd #637), which stores
|
||||
* it as a field. Reflection recovers it from there, and then from the scanner itself, so the
|
||||
* assertion is against the real production argument rather than a copy built for this test.
|
||||
*/
|
||||
class FleetdAssemblyLeadTabScannerExclusionTest {
|
||||
|
||||
private static final class FakeLeadChannel implements LeadChannelHandle {
|
||||
@Override
|
||||
public void publish(String toCoordId, LeadMessage message) {
|
||||
}
|
||||
|
||||
@Override
|
||||
public List<LeadMessage> peek() {
|
||||
return List.of();
|
||||
}
|
||||
|
||||
@Override
|
||||
public void ack(String msgId) {
|
||||
}
|
||||
|
||||
@Override
|
||||
public String selfCoordId() {
|
||||
return "test-lead";
|
||||
}
|
||||
|
||||
@Override
|
||||
public boolean heldDurable() {
|
||||
return true;
|
||||
}
|
||||
|
||||
@Override
|
||||
public MailboxState inspect(String coordId) {
|
||||
return MailboxState.unknown(coordId);
|
||||
}
|
||||
|
||||
@Override
|
||||
public void close() {
|
||||
}
|
||||
}
|
||||
|
||||
private static final class TestResourcePorts implements ResourcePorts {
|
||||
final FakeHerdr herdr = new FakeHerdr();
|
||||
Runnable shutdownHook;
|
||||
|
||||
@Override
|
||||
public Map<String, String> environment() {
|
||||
return Map.of();
|
||||
}
|
||||
|
||||
@Override
|
||||
public HerdrClient connectHerdr(Path socketPath) {
|
||||
return herdr;
|
||||
}
|
||||
|
||||
@Override
|
||||
public Fleetd.AmqpOpener replyInboxOpener() {
|
||||
return (uri, prefetch) -> new ReplyInbox() {
|
||||
@Override public void own(String target) { }
|
||||
@Override public void release(String target) { }
|
||||
@Override public void publish(String target, String msgId, String content) { }
|
||||
@Override public List<InboxMessage> peek(String target) { return List.of(); }
|
||||
@Override public boolean ack(String target, String msgId) { return false; }
|
||||
};
|
||||
}
|
||||
|
||||
@Override
|
||||
public Fleetd.LeadMailboxOpener leadMailboxOpener() {
|
||||
return (uri, selfCoordId, prefetch) -> new FakeLeadChannel();
|
||||
}
|
||||
|
||||
@Override
|
||||
public LongSupplier nanoClock() {
|
||||
return System::nanoTime;
|
||||
}
|
||||
|
||||
@Override
|
||||
public LongSupplier wallClockNanos() {
|
||||
return System::nanoTime;
|
||||
}
|
||||
|
||||
@Override
|
||||
public ScheduledExecutorService newScheduler(String purpose) {
|
||||
return Executors.newSingleThreadScheduledExecutor();
|
||||
}
|
||||
|
||||
@Override
|
||||
public void addShutdownHook(Runnable hook) {
|
||||
shutdownHook = hook;
|
||||
}
|
||||
|
||||
@Override
|
||||
public void startHttp(Javalin app, String host, int port) {
|
||||
// Do not bind a real port in this assembly test.
|
||||
}
|
||||
|
||||
@Override
|
||||
public Runnable herdrPollWait() {
|
||||
return () -> {
|
||||
throw new UnsupportedOperationException("FakeHerdr is healthy; no poll wait is expected");
|
||||
};
|
||||
}
|
||||
}
|
||||
|
||||
private static FleetConfig writeConfig(Path dir) throws Exception {
|
||||
Path file = dir.resolve("fleetd.yaml");
|
||||
Files.writeString(file, """
|
||||
bind:
|
||||
host: 127.0.0.1
|
||||
port: 8765
|
||||
idleSleepGuard:
|
||||
enabled: false
|
||||
coordinator:
|
||||
uri: "amqp://fake-lead-broker/vh"
|
||||
selfId: "test-lead"
|
||||
fleet:
|
||||
leaders:
|
||||
primary:
|
||||
tab: "lead: primary"
|
||||
profile: sonnet
|
||||
profiles:
|
||||
sonnet:
|
||||
subscription: true
|
||||
argv: ["ccs", "sonnet"]
|
||||
""");
|
||||
return FleetConfig.load(file);
|
||||
}
|
||||
|
||||
@Test
|
||||
void productionBootPathPassesNoExcludedWorkspaceLabels(@TempDir Path dir) throws Exception {
|
||||
FleetConfig cfg = writeConfig(dir);
|
||||
TestResourcePorts ports = new TestResourcePorts();
|
||||
FleetdRuntime runtime = FleetdAssembly.assembleAndStart(new AssemblyInputs(cfg,
|
||||
new ConfigRef(dir.resolve("fleetd.yaml"), cfg), new SubscriptionGuard(cfg.guard().hostSet())), ports);
|
||||
try {
|
||||
LeadCoordLoop coordLoop = runtime.leadCoordLoop();
|
||||
assertNotNull(coordLoop, "control: a configured coordinator: block must build LeadCoordLoop");
|
||||
|
||||
Field leadsField = LeadCoordLoop.class.getDeclaredField("leads");
|
||||
leadsField.setAccessible(true);
|
||||
@SuppressWarnings("unchecked")
|
||||
Supplier<Map<String, String>> leads = (Supplier<Map<String, String>>) leadsField.get(coordLoop);
|
||||
|
||||
assertInstanceOf(LeadTabScanner.class, leads,
|
||||
"control: a non-empty fleet.leaders: block must make FleetdAssembly build a real "
|
||||
+ "LeadTabScanner for its `leads` supplier, not the Map::of fallback — "
|
||||
+ "otherwise this test would pass for the wrong reason");
|
||||
|
||||
Field excludedField = LeadTabScanner.class.getDeclaredField("excludedWorkspaceLabels");
|
||||
excludedField.setAccessible(true);
|
||||
Set<?> excluded = (Set<?>) excludedField.get(leads);
|
||||
|
||||
assertTrue(excluded.isEmpty(),
|
||||
"FleetdAssembly.java:265 must pass an empty excludedWorkspaceLabels to "
|
||||
+ "LeadTabScanner — scanning member tabs would demote the lead to a worker");
|
||||
} finally {
|
||||
assertNotNull(ports.shutdownHook, "control: assembly must capture its shutdown hook");
|
||||
ports.shutdownHook.run();
|
||||
}
|
||||
}
|
||||
}
|
||||
@@ -0,0 +1,266 @@
|
||||
package dev.ltms.fleet;
|
||||
|
||||
import dev.ltms.fleet.config.ConfigRef;
|
||||
import dev.ltms.fleet.config.FleetConfig;
|
||||
import dev.ltms.fleet.guard.SubscriptionGuard;
|
||||
import dev.ltms.fleet.herdr.FakeHerdr;
|
||||
import dev.ltms.fleet.herdr.HerdrClient;
|
||||
import dev.ltms.fleet.inject.LoopWatchdog;
|
||||
import dev.ltms.fleet.msg.ReplyInbox;
|
||||
import io.javalin.Javalin;
|
||||
import org.junit.jupiter.api.Test;
|
||||
import org.junit.jupiter.api.io.TempDir;
|
||||
|
||||
import java.nio.file.Files;
|
||||
import java.nio.file.Path;
|
||||
import java.util.ArrayList;
|
||||
import java.util.List;
|
||||
import java.util.Map;
|
||||
import java.util.concurrent.CopyOnWriteArrayList;
|
||||
import java.util.concurrent.Executors;
|
||||
import java.util.concurrent.ScheduledExecutorService;
|
||||
import java.util.function.LongSupplier;
|
||||
|
||||
import static org.junit.jupiter.api.Assertions.assertEquals;
|
||||
import static org.junit.jupiter.api.Assertions.assertFalse;
|
||||
import static org.junit.jupiter.api.Assertions.assertNotNull;
|
||||
import static org.junit.jupiter.api.Assertions.assertTrue;
|
||||
|
||||
/**
|
||||
* fleetd #612 Unit A: proves {@link FleetdAssembly#assembleAndStart} — not a copy of its logic —
|
||||
* against a real {@link FleetConfig}, a {@link FakeHerdr} and a fake {@link ResourcePorts}, with no
|
||||
* real herdr socket, no real broker, and no real HTTP bind.
|
||||
*
|
||||
* <p><strong>The canonical order this test asserts against was recorded from {@code Fleetd.java}
|
||||
* BEFORE any code moved</strong> (fleetd #612 Unit A's mandated order of work), by reading the
|
||||
* original {@code main}'s body and its shutdown-hook {@code Thread}:
|
||||
*
|
||||
* <p>Start order: {@code SessionReaper.start()} → {@code StatusPoller.start()} →
|
||||
* {@code LeadHeartbeatLoop.start()} (opt-in) → {@code FleetHealthMonitor.start()} (opt-in) →
|
||||
* {@code LeadCoordLoop.start()} (opt-in) → {@code ConfigWatcher.start()} (opt-in) →
|
||||
* {@code app.start()} (HTTP), always last.
|
||||
*
|
||||
* <p>Close order (from the original shutdown hook body): {@code sessions.close(drainTimeoutSeconds)}
|
||||
* → {@code poller.stop()} → {@code messages.close()} → {@code pushLoop.close()} →
|
||||
* {@code heartbeat.close()} (if present) → {@code leadCoordLoop.close()} (if present) →
|
||||
* {@code leadCoordScheduler.shutdownNow()} (if present) → {@code healthMonitor.stop()} (if present)
|
||||
* → {@code configWatcher.stop()} (if present) → {@code mcp.close()} → {@code reaper.stop()} (if
|
||||
* present) → {@code idleSleepGuard.close()} (if present) → {@code replyInbox.close()} (if
|
||||
* {@code AutoCloseable}) → {@code leadMailbox.close()} (if present) → {@code router.close()}.
|
||||
*
|
||||
* <p>This test's config deliberately leaves {@code coordinator:} unset, so {@code leadMailbox} and
|
||||
* {@code leadCoordLoop} stay {@code null} throughout — the lead-mailbox resource-ledger criterion is
|
||||
* NOT exercised here; see the class-level caveat in the implementer's hand-off. {@code
|
||||
* idleSleepGuard.enabled: false} is set for the same kind of reason: it would otherwise try to spawn
|
||||
* a real {@code caffeinate} subprocess, which is not one of the resources the ticket's acceptance
|
||||
* criteria names (scheduler/inbox/mailbox/loop/MCP server/herdr router).
|
||||
*/
|
||||
class FleetdAssemblyLifecycleTest {
|
||||
|
||||
/** The one {@link HerdrClient} both {@code herdrSocket} and {@code memberHerdrSocket} resolve to. */
|
||||
private static final class RecordingResourcePorts implements ResourcePorts {
|
||||
|
||||
final FakeHerdr herdr = new FakeHerdr();
|
||||
final SentinelReplyInbox replyInbox = new SentinelReplyInbox();
|
||||
final List<String> ledger = new CopyOnWriteArrayList<>();
|
||||
final List<ScheduledExecutorService> schedulers = new CopyOnWriteArrayList<>();
|
||||
final List<String> schedulerPurposes = new ArrayList<>();
|
||||
Runnable shutdownHook;
|
||||
Javalin startedApp;
|
||||
|
||||
@Override
|
||||
public Map<String, String> environment() {
|
||||
return Map.of();
|
||||
}
|
||||
|
||||
@Override
|
||||
public HerdrClient connectHerdr(Path socketPath) {
|
||||
ledger.add("connectHerdr");
|
||||
return herdr;
|
||||
}
|
||||
|
||||
@Override
|
||||
public Fleetd.AmqpOpener replyInboxOpener() {
|
||||
ledger.add("replyInboxOpener");
|
||||
return (uri, prefetch) -> replyInbox;
|
||||
}
|
||||
|
||||
@Override
|
||||
public Fleetd.LeadMailboxOpener leadMailboxOpener() {
|
||||
// Never invoked: this test's config has no `coordinator:` block, so
|
||||
// Fleetd.openLeadMailbox returns null before calling the opener at all.
|
||||
return (uri, selfCoordId, prefetch) -> {
|
||||
throw new UnsupportedOperationException(
|
||||
"leadMailboxOpener must not be called — no coordinator: block is configured");
|
||||
};
|
||||
}
|
||||
|
||||
@Override
|
||||
public LongSupplier nanoClock() {
|
||||
return System::nanoTime;
|
||||
}
|
||||
|
||||
@Override
|
||||
public LongSupplier wallClockNanos() {
|
||||
return System::nanoTime;
|
||||
}
|
||||
|
||||
@Override
|
||||
public ScheduledExecutorService newScheduler(String purpose) {
|
||||
ledger.add("newScheduler:" + purpose);
|
||||
schedulerPurposes.add(purpose);
|
||||
ScheduledExecutorService scheduler = Executors.newSingleThreadScheduledExecutor();
|
||||
schedulers.add(scheduler);
|
||||
return scheduler;
|
||||
}
|
||||
|
||||
@Override
|
||||
public void addShutdownHook(Runnable hook) {
|
||||
ledger.add("addShutdownHook");
|
||||
this.shutdownHook = hook;
|
||||
}
|
||||
|
||||
@Override
|
||||
public void startHttp(Javalin app, String host, int port) {
|
||||
// Deliberately never call app.start(host, port): no real HTTP bind in a unit test.
|
||||
ledger.add("startHttp");
|
||||
this.startedApp = app;
|
||||
}
|
||||
|
||||
@Override
|
||||
public Runnable herdrPollWait() {
|
||||
// Never invoked: this test's FakeHerdr answers immediately, so awaitHerdr never polls.
|
||||
return () -> {
|
||||
throw new UnsupportedOperationException("herdrPollWait must not be called — herdr is healthy");
|
||||
};
|
||||
}
|
||||
}
|
||||
|
||||
/** A fake {@link ReplyInbox} that is also {@link AutoCloseable}, so the ledger can prove it closes. */
|
||||
private static final class SentinelReplyInbox implements ReplyInbox, AutoCloseable {
|
||||
volatile boolean closed = false;
|
||||
|
||||
@Override
|
||||
public void own(String target) {
|
||||
}
|
||||
|
||||
@Override
|
||||
public void release(String target) {
|
||||
}
|
||||
|
||||
@Override
|
||||
public void publish(String target, String msgId, String content) {
|
||||
}
|
||||
|
||||
@Override
|
||||
public List<InboxMessage> peek(String target) {
|
||||
return List.of();
|
||||
}
|
||||
|
||||
@Override
|
||||
public boolean ack(String target, String msgId) {
|
||||
return false;
|
||||
}
|
||||
|
||||
@Override
|
||||
public void close() {
|
||||
closed = true;
|
||||
}
|
||||
}
|
||||
|
||||
private static FleetConfig writeConfig(Path dir) throws Exception {
|
||||
Path f = dir.resolve("fleetd.yaml");
|
||||
Files.writeString(f, """
|
||||
bind:
|
||||
host: 127.0.0.1
|
||||
port: 8765
|
||||
lifecycle:
|
||||
idleTtlSeconds: 600
|
||||
health:
|
||||
enabled: true
|
||||
intervalSeconds: 30
|
||||
leadHeartbeat:
|
||||
idleAfterSeconds: 600
|
||||
backoffMs: 15000
|
||||
quietNudgeCap: 5
|
||||
configReload:
|
||||
enabled: true
|
||||
intervalSeconds: 30
|
||||
idleSleepGuard:
|
||||
enabled: false
|
||||
broker:
|
||||
uri: "amqp://fake-test-broker/vh"
|
||||
""");
|
||||
return FleetConfig.load(f);
|
||||
}
|
||||
|
||||
@Test
|
||||
void assemblesTheRealBootGraphWithoutTouchingAnyRealSocketBrokerOrPort(@TempDir Path dir)
|
||||
throws Exception {
|
||||
FleetConfig cfg = writeConfig(dir);
|
||||
ConfigRef config = new ConfigRef(dir.resolve("fleetd.yaml"), cfg);
|
||||
SubscriptionGuard guard = new SubscriptionGuard(cfg.guard().hostSet());
|
||||
RecordingResourcePorts ports = new RecordingResourcePorts();
|
||||
|
||||
FleetdRuntime runtime = FleetdAssembly.assembleAndStart(new AssemblyInputs(cfg, config, guard), ports);
|
||||
|
||||
// --- start order: the herdr connect, then the recurring background loops, in the recorded
|
||||
// order, then the shutdown hook is registered, then (finally) HTTP "starts" -------------
|
||||
assertTrue(ports.ledger.indexOf("connectHerdr") < ports.ledger.indexOf("newScheduler:bridge-push-"),
|
||||
"herdr must connect before the push scheduler is created: " + ports.ledger);
|
||||
assertEquals(List.of("bridge-push-", "bridge-heartbeat-", "bridge-health-"), ports.schedulerPurposes,
|
||||
"the three always-created schedulers must be requested in exactly this order: "
|
||||
+ ports.schedulerPurposes);
|
||||
assertTrue(ports.ledger.indexOf("addShutdownHook") < ports.ledger.indexOf("startHttp"),
|
||||
"the shutdown hook must be registered before HTTP starts — the one statement that "
|
||||
+ "could not be reordered without changing FleetdRuntime's constructor shape, "
|
||||
+ "see FleetdAssembly's javadoc: " + ports.ledger);
|
||||
assertEquals(ports.ledger.size() - 1, ports.ledger.indexOf("startHttp"),
|
||||
"HTTP must start LAST of everything this fake observes: " + ports.ledger);
|
||||
assertNotNull(ports.startedApp, "FleetdAssembly must have built and handed off a real FleetApp");
|
||||
|
||||
// No real HTTP bind and no real herdr socket: this call returning at all, plus the ledger
|
||||
// above, is the proof — a real bind or a real UnixSocketHerdrClient.connect would have
|
||||
// thrown or hung against the sockets/ports this test never opened.
|
||||
assertNotNull(runtime.app(), "FleetdRuntime must own the same Javalin app that was built");
|
||||
assertEquals(ports.startedApp, runtime.app(), "attachApp must hand FleetdRuntime the SAME instance");
|
||||
|
||||
// --- the runtime owns the real, live objects — not a copy --------------------------------
|
||||
assertEquals(LoopWatchdog.State.RUNNING, runtime.reaper().health(),
|
||||
"SessionReaper must be running: lifecycle.idleTtlSeconds is configured");
|
||||
assertEquals(LoopWatchdog.State.RUNNING, runtime.poller().health(), "StatusPoller must be running");
|
||||
assertNotNull(runtime.heartbeat(), "leadHeartbeat: is configured, so the loop must be built and started");
|
||||
assertNotNull(runtime.healthMonitor(), "health.enabled: true, so the monitor must be built and started");
|
||||
assertNotNull(runtime.configWatcher(), "configReload.enabled: true, so the watcher must be built and started");
|
||||
// Documented gap (see class javadoc): no coordinator: block, so these stay null.
|
||||
assertEquals(null, runtime.leadCoordLoop(), "no coordinator: block — leadCoordLoop must stay unbuilt");
|
||||
assertEquals(null, runtime.leadMailbox(), "no coordinator: block — leadMailbox must stay unbuilt");
|
||||
assertEquals(runtime.replyInbox(), ports.replyInbox,
|
||||
"FleetdRuntime must own the exact ReplyInbox instance this fake's AmqpOpener returned");
|
||||
assertFalse(ports.herdr.closed, "herdr must still be open while the daemon is running");
|
||||
assertFalse(ports.replyInbox.closed, "the reply inbox must still be open while the daemon is running");
|
||||
for (ScheduledExecutorService scheduler : ports.schedulers) {
|
||||
assertFalse(scheduler.isShutdown(), "a scheduler must still be running while the daemon is up");
|
||||
}
|
||||
|
||||
// --- close order: invoke the captured shutdown-hook Runnable directly (no real JVM shutdown
|
||||
// happens in a unit test) and prove every resource this fake can observe is released --------
|
||||
assertNotNull(ports.shutdownHook, "FleetdAssembly must have registered a shutdown hook");
|
||||
ports.shutdownHook.run();
|
||||
|
||||
assertEquals(LoopWatchdog.State.STOPPED, runtime.reaper().health(), "SessionReaper must stop on close");
|
||||
assertEquals(LoopWatchdog.State.STOPPED, runtime.poller().health(), "StatusPoller must stop on close");
|
||||
assertTrue(ports.herdr.closed, "router.close() must close the herdr client last");
|
||||
assertTrue(ports.replyInbox.closed, "the AutoCloseable reply inbox must be closed");
|
||||
for (int i = 0; i < ports.schedulers.size(); i++) {
|
||||
assertTrue(ports.schedulers.get(i).isShutdown(),
|
||||
"scheduler for '" + ports.schedulerPurposes.get(i) + "' must be shut down by close(): "
|
||||
+ "LeadHeartbeatLoop/FleetHealthMonitor/ReplyPushLoop each call "
|
||||
+ "scheduler.shutdownNow() on the exact instance ports.newScheduler(...) handed them");
|
||||
}
|
||||
|
||||
// Calling the captured hook a second time must never happen for a real JVM shutdown hook,
|
||||
// but nothing above should have thrown — that already proves every accessed field's close()
|
||||
// tolerated running once, in the recorded order, without an exception escaping.
|
||||
}
|
||||
}
|
||||
@@ -0,0 +1,180 @@
|
||||
package dev.ltms.fleet;
|
||||
|
||||
import dev.ltms.fleet.config.ConfigRef;
|
||||
import dev.ltms.fleet.config.FleetConfig;
|
||||
import dev.ltms.fleet.guard.SubscriptionGuard;
|
||||
import dev.ltms.fleet.herdr.FakeHerdr;
|
||||
import dev.ltms.fleet.herdr.HerdrClient;
|
||||
import dev.ltms.fleet.inject.LoopWatchdog;
|
||||
import dev.ltms.fleet.mcp.FleetMcp;
|
||||
import io.javalin.Javalin;
|
||||
import org.junit.jupiter.api.AfterEach;
|
||||
import org.junit.jupiter.api.Test;
|
||||
import org.junit.jupiter.api.io.TempDir;
|
||||
|
||||
import java.lang.reflect.Field;
|
||||
import java.net.URI;
|
||||
import java.net.http.HttpClient;
|
||||
import java.net.http.HttpRequest;
|
||||
import java.net.http.HttpResponse;
|
||||
import java.nio.file.Files;
|
||||
import java.nio.file.Path;
|
||||
import java.util.Map;
|
||||
import java.util.concurrent.Executors;
|
||||
import java.util.concurrent.ScheduledExecutorService;
|
||||
import java.util.function.LongSupplier;
|
||||
|
||||
import static org.junit.jupiter.api.Assertions.assertEquals;
|
||||
import static org.junit.jupiter.api.Assertions.assertNotNull;
|
||||
import static org.junit.jupiter.api.Assertions.assertTrue;
|
||||
|
||||
/**
|
||||
* fleetd #612 Shape A, rank 10: {@link FleetdAssembly} creates one {@link FleetMcp.LoopHealthSource}
|
||||
* from the real started {@code StatusPoller} and {@code SessionReaper}, then gives it to two operator
|
||||
* windows. These tests reach the real assembled objects through {@link FleetdRuntime}, rather than
|
||||
* building a second source beside them. A hardcoded {@code RUNNING} source would pass a simple
|
||||
* "running" test, so the mutation proof also mis-wires the source to never-started loops: both
|
||||
* windows must then report {@code STOPPED} and these assertions go red.
|
||||
*/
|
||||
class FleetdAssemblyLoopHealthTest {
|
||||
|
||||
private static final class RecordingResourcePorts implements ResourcePorts {
|
||||
final FakeHerdr herdr = new FakeHerdr();
|
||||
Runnable shutdownHook;
|
||||
|
||||
@Override
|
||||
public Map<String, String> environment() {
|
||||
return Map.of();
|
||||
}
|
||||
|
||||
@Override
|
||||
public HerdrClient connectHerdr(Path socketPath) {
|
||||
return herdr;
|
||||
}
|
||||
|
||||
@Override
|
||||
public Fleetd.AmqpOpener replyInboxOpener() {
|
||||
return (uri, prefetch) -> new dev.ltms.fleet.msg.InMemoryReplyInbox();
|
||||
}
|
||||
|
||||
@Override
|
||||
public Fleetd.LeadMailboxOpener leadMailboxOpener() {
|
||||
return (uri, selfCoordId, prefetch) -> {
|
||||
throw new UnsupportedOperationException("no coordinator is configured");
|
||||
};
|
||||
}
|
||||
|
||||
@Override
|
||||
public Runnable herdrPollWait() {
|
||||
return () -> {
|
||||
throw new UnsupportedOperationException("healthy FakeHerdr must not be polled");
|
||||
};
|
||||
}
|
||||
|
||||
@Override
|
||||
public LongSupplier nanoClock() {
|
||||
return System::nanoTime;
|
||||
}
|
||||
|
||||
@Override
|
||||
public LongSupplier wallClockNanos() {
|
||||
return System::nanoTime;
|
||||
}
|
||||
|
||||
@Override
|
||||
public ScheduledExecutorService newScheduler(String purpose) {
|
||||
return Executors.newSingleThreadScheduledExecutor();
|
||||
}
|
||||
|
||||
@Override
|
||||
public void addShutdownHook(Runnable hook) {
|
||||
shutdownHook = hook;
|
||||
}
|
||||
|
||||
@Override
|
||||
public void startHttp(Javalin app, String host, int port) {
|
||||
// Bind runtime.app() to an ephemeral port only in the REST assertion below.
|
||||
}
|
||||
}
|
||||
|
||||
private final HttpClient http = HttpClient.newHttpClient();
|
||||
private RecordingResourcePorts ports;
|
||||
private FleetdRuntime runtime;
|
||||
private Javalin boundApp;
|
||||
|
||||
@AfterEach
|
||||
void tearDown() {
|
||||
if (boundApp != null) {
|
||||
boundApp.stop();
|
||||
}
|
||||
if (runtime != null) {
|
||||
runtime.close();
|
||||
}
|
||||
if (ports != null) {
|
||||
assertNotNull(ports.shutdownHook, "the assembly must register its shutdown hook");
|
||||
assertTrue(ports.herdr.closed, "FleetdRuntime.close must close the real assembled herdr client");
|
||||
}
|
||||
}
|
||||
|
||||
private static FleetConfig writeConfig(Path dir) throws Exception {
|
||||
Path file = dir.resolve("fleetd.yaml");
|
||||
Files.writeString(file, """
|
||||
bind:
|
||||
host: 127.0.0.1
|
||||
port: 8765
|
||||
lifecycle:
|
||||
idleTtlSeconds: 600
|
||||
idleSleepGuard:
|
||||
enabled: false
|
||||
health:
|
||||
enabled: false
|
||||
broker:
|
||||
uri: "amqp://fake-test-broker/vh"
|
||||
""");
|
||||
return FleetConfig.load(file);
|
||||
}
|
||||
|
||||
private void assemble(Path dir) throws Exception {
|
||||
FleetConfig cfg = writeConfig(dir);
|
||||
ports = new RecordingResourcePorts();
|
||||
runtime = FleetdAssembly.assembleAndStart(
|
||||
new AssemblyInputs(cfg, new ConfigRef(dir.resolve("fleetd.yaml"), cfg),
|
||||
new SubscriptionGuard(cfg.guard().hostSet())), ports);
|
||||
}
|
||||
|
||||
@Test
|
||||
void fleetListUsesTheRunningLoopsInTheRealAssembledMcp(@TempDir Path dir) throws Exception {
|
||||
assemble(dir);
|
||||
|
||||
// The field is the exact source captured by FleetMcp's fleet_list handler. Reflection is
|
||||
// necessary because FleetMcp has no public source accessor; it is not a source-text check.
|
||||
FleetMcp.LoopHealthSource loopHealth = loopHealthOf(runtime.mcp());
|
||||
assertRunning(loopHealth, "FleetMcp's real fleet_list source");
|
||||
}
|
||||
|
||||
@Test
|
||||
void healthzUsesTheRunningLoopsInTheRealAssembledApp(@TempDir Path dir) throws Exception {
|
||||
assemble(dir);
|
||||
boundApp = runtime.app().start("127.0.0.1", 0);
|
||||
|
||||
HttpRequest request = HttpRequest.newBuilder(
|
||||
URI.create("http://127.0.0.1:" + boundApp.port() + "/healthz")).GET().build();
|
||||
HttpResponse<String> response = http.send(request, HttpResponse.BodyHandlers.ofString());
|
||||
assertEquals(200, response.statusCode(), response.body());
|
||||
assertTrue(response.body().contains("\"statusPoller\":\"RUNNING\""), response.body());
|
||||
assertTrue(response.body().contains("\"sessionReaper\":\"RUNNING\""), response.body());
|
||||
}
|
||||
|
||||
private static FleetMcp.LoopHealthSource loopHealthOf(FleetMcp mcp) throws Exception {
|
||||
Field field = FleetMcp.class.getDeclaredField("loopHealth");
|
||||
field.setAccessible(true);
|
||||
return (FleetMcp.LoopHealthSource) field.get(mcp);
|
||||
}
|
||||
|
||||
private static void assertRunning(FleetMcp.LoopHealthSource source, String consumer) {
|
||||
assertEquals(LoopWatchdog.State.RUNNING, source.statusPoller().get(),
|
||||
consumer + " must report the started real StatusPoller as RUNNING");
|
||||
assertEquals(LoopWatchdog.State.RUNNING, source.sessionReaper().get(),
|
||||
consumer + " must report the started real SessionReaper as RUNNING");
|
||||
}
|
||||
}
|
||||
@@ -0,0 +1,232 @@
|
||||
package dev.ltms.fleet;
|
||||
|
||||
import dev.ltms.fleet.config.ConfigRef;
|
||||
import dev.ltms.fleet.config.FleetConfig;
|
||||
import dev.ltms.fleet.guard.SubscriptionGuard;
|
||||
import dev.ltms.fleet.herdr.FakeHerdr;
|
||||
import dev.ltms.fleet.herdr.HerdrClient;
|
||||
import dev.ltms.fleet.mcp.PrimaryRegistry;
|
||||
import dev.ltms.fleet.msg.MessageService;
|
||||
import dev.ltms.fleet.msg.ReplyInbox;
|
||||
import dev.ltms.fleet.msg.ReplyPushLoop;
|
||||
import dev.ltms.fleet.session.SessionManager;
|
||||
import io.javalin.Javalin;
|
||||
import org.junit.jupiter.api.DisplayName;
|
||||
import org.junit.jupiter.api.Test;
|
||||
import org.junit.jupiter.api.io.TempDir;
|
||||
|
||||
import java.lang.reflect.Field;
|
||||
import java.nio.file.Files;
|
||||
import java.nio.file.Path;
|
||||
import java.util.List;
|
||||
import java.util.Map;
|
||||
import java.util.concurrent.Executors;
|
||||
import java.util.concurrent.ScheduledExecutorService;
|
||||
import java.util.function.Consumer;
|
||||
import java.util.function.LongSupplier;
|
||||
|
||||
import static org.junit.jupiter.api.Assertions.assertEquals;
|
||||
import static org.junit.jupiter.api.Assertions.assertNotNull;
|
||||
import static org.junit.jupiter.api.Assertions.assertTrue;
|
||||
|
||||
/**
|
||||
* fleetd #612 rank 6 — {@code FleetdAssembly.java:447} wires {@code
|
||||
* sessions.onRelease(Fleetd.releaseCleanup(messages, replyInbox, primaryRegistry))}. {@link
|
||||
* FleetdReleaseCleanupWiringTest} already pins that the FACTORY {@code Fleetd.releaseCleanup}
|
||||
* itself reaches all three collaborators — but it calls the factory directly, never {@code
|
||||
* FleetdAssembly.assembleAndStart}, so it cannot see whether the real call site at {@code :447}
|
||||
* still registers it (as opposed to a no-op {@code detail -> { }}) or still passes it the REAL,
|
||||
* assembled {@code messages}/{@code replyInbox}/{@code primaryRegistry}. Swapping the registered
|
||||
* listener for a no-op at that call site compiles clean and leaves the whole suite — including the
|
||||
* factory-level test — green: EVERY teardown then leaks a stuck rendezvous waiter, an unreleased
|
||||
* reply-inbox consumer, and a stale lead binding, all three at once.
|
||||
*
|
||||
* <p>This test assembles the real daemon with {@code idleSleepGuard.enabled: false} — the ONLY
|
||||
* other {@code onRelease} registration in {@code FleetdAssembly} (see {@code
|
||||
* dev.ltms.fleet.power.IdleSleepGuard}'s own wiring at {@code FleetdAssembly.java:228}) — so the
|
||||
* real {@link SessionManager}'s release-listener list holds exactly the one listener this call site
|
||||
* registers. It pulls that REAL listener out via reflection (the list itself is private, like
|
||||
* {@code StatusPollerResilienceTest}'s use of the same technique elsewhere in this suite), invokes
|
||||
* it directly, and asserts all three collaborator effects against the REAL, assembled {@link
|
||||
* MessageService} ({@link FleetdRuntime#messages()}), the REAL {@link ReplyInbox} ({@link
|
||||
* FleetdRuntime#replyInbox()}), and the REAL {@link PrimaryRegistry} — reached through {@link
|
||||
* FleetdRuntime#pushLoop()}, the only other accessor that was handed the same {@code
|
||||
* primaryRegistry} instance ({@code FleetdAssembly.java:382}), since {@code FleetMcp} never exposes
|
||||
* it.
|
||||
*/
|
||||
class FleetdAssemblyReleaseCleanupBehaviouralTest {
|
||||
|
||||
private static final String TARGET = "term_a";
|
||||
|
||||
private static final class RecordingResourcePorts implements ResourcePorts {
|
||||
final FakeHerdr herdr = new FakeHerdr();
|
||||
Runnable shutdownHook;
|
||||
|
||||
@Override
|
||||
public Map<String, String> environment() {
|
||||
return Map.of();
|
||||
}
|
||||
|
||||
@Override
|
||||
public HerdrClient connectHerdr(Path socketPath) {
|
||||
return herdr;
|
||||
}
|
||||
|
||||
@Override
|
||||
public Fleetd.AmqpOpener replyInboxOpener() {
|
||||
return (uri, prefetch) -> {
|
||||
throw new UnsupportedOperationException("no broker: block is configured");
|
||||
};
|
||||
}
|
||||
|
||||
@Override
|
||||
public Fleetd.LeadMailboxOpener leadMailboxOpener() {
|
||||
return (uri, selfCoordId, prefetch) -> {
|
||||
throw new UnsupportedOperationException("no coordinator: block is configured");
|
||||
};
|
||||
}
|
||||
|
||||
@Override
|
||||
public Runnable herdrPollWait() {
|
||||
// Never invoked: this test's FakeHerdr answers immediately, so awaitHerdr never polls.
|
||||
return () -> {
|
||||
throw new UnsupportedOperationException("herdrPollWait must not be called — herdr is healthy");
|
||||
};
|
||||
}
|
||||
@Override
|
||||
public LongSupplier nanoClock() {
|
||||
return System::nanoTime;
|
||||
}
|
||||
|
||||
@Override
|
||||
public LongSupplier wallClockNanos() {
|
||||
return System::nanoTime;
|
||||
}
|
||||
|
||||
@Override
|
||||
public ScheduledExecutorService newScheduler(String purpose) {
|
||||
return Executors.newSingleThreadScheduledExecutor();
|
||||
}
|
||||
|
||||
@Override
|
||||
public void addShutdownHook(Runnable hook) {
|
||||
this.shutdownHook = hook;
|
||||
}
|
||||
|
||||
@Override
|
||||
public void startHttp(Javalin app, String host, int port) {
|
||||
}
|
||||
}
|
||||
|
||||
private static FleetConfig writeConfig(Path dir) throws Exception {
|
||||
Path f = dir.resolve("fleetd.yaml");
|
||||
Files.writeString(f, """
|
||||
bind:
|
||||
host: 127.0.0.1
|
||||
port: 8765
|
||||
idleSleepGuard:
|
||||
enabled: false
|
||||
""");
|
||||
return FleetConfig.load(f);
|
||||
}
|
||||
|
||||
@SuppressWarnings("unchecked")
|
||||
@Test
|
||||
@DisplayName("[BEHAVIOURAL] the real assembled release listener reaches messages.abandon, "
|
||||
+ "replyInbox.release, AND primaryRegistry.forgetDelegation — all three leaks at once")
|
||||
void assembledReleaseListenerReachesAllThreeCollaborators(@TempDir Path dir) throws Exception {
|
||||
FleetConfig cfg = writeConfig(dir);
|
||||
ConfigRef config = new ConfigRef(dir.resolve("fleetd.yaml"), cfg);
|
||||
SubscriptionGuard guard = new SubscriptionGuard(cfg.guard().hostSet());
|
||||
RecordingResourcePorts ports = new RecordingResourcePorts();
|
||||
|
||||
FleetdRuntime runtime = FleetdAssembly.assembleAndStart(new AssemblyInputs(cfg, config, guard), ports);
|
||||
assertNotNull(ports.shutdownHook, "FleetdAssembly must have registered a shutdown hook");
|
||||
|
||||
// Surefire runs the whole suite in one JVM fork, so the scheduler/loops this assembly starts
|
||||
// must be torn down here, on the failure path too — hence the try/finally, not just a
|
||||
// statement at the end of the happy path.
|
||||
try {
|
||||
// --- reach into SessionManager's private release-listener list. idleSleepGuard.enabled:
|
||||
// false above means FleetdAssembly.java:228 never registers, so this list must hold EXACTLY
|
||||
// the one listener :447 registers.
|
||||
Field listenersField = SessionManager.class.getDeclaredField("releaseListeners");
|
||||
listenersField.setAccessible(true);
|
||||
List<Consumer<SessionManager.ReleaseDetail>> releaseListeners =
|
||||
(List<Consumer<SessionManager.ReleaseDetail>>) listenersField.get(runtime.sessions());
|
||||
assertEquals(1, releaseListeners.size(), "control: with idleSleepGuard.enabled: false, "
|
||||
+ "FleetdAssembly.java:447 must be the ONLY onRelease registration — a different "
|
||||
+ "count means this test is no longer isolating the call site it claims to pin");
|
||||
Consumer<SessionManager.ReleaseDetail> releaseListener = releaseListeners.get(0);
|
||||
|
||||
MessageService messages = runtime.messages();
|
||||
ReplyInbox replyInbox = runtime.replyInbox();
|
||||
|
||||
// primaryRegistry is never exposed by FleetdRuntime directly — ReplyPushLoop is the other
|
||||
// collaborator FleetdAssembly.java:382 hands the SAME instance to, so reach it from there.
|
||||
Field primaryRegistryField = ReplyPushLoop.class.getDeclaredField("primaryRegistry");
|
||||
primaryRegistryField.setAccessible(true);
|
||||
PrimaryRegistry primaryRegistry = (PrimaryRegistry) primaryRegistryField.get(runtime.pushLoop());
|
||||
assertNotNull(primaryRegistry, "control: the assembled ReplyPushLoop must hold a real "
|
||||
+ "PrimaryRegistry instance");
|
||||
|
||||
// --- loud controls: set up the "before" state each collaborator's effect is measured
|
||||
// against, against the REAL assembled objects. If any of these three fails, the test below
|
||||
// would pass vacuously because the subject it claims to observe never existed in the first
|
||||
// place.
|
||||
String ticket = messages.sendAsync(TARGET, "long task");
|
||||
assertEquals(MessageService.Phase.PENDING, messages.poll(ticket).phase(),
|
||||
"control: the async ticket must be PENDING before the release listener runs");
|
||||
|
||||
replyInbox.own(TARGET);
|
||||
replyInbox.publish(TARGET, "msg-1", "hello");
|
||||
assertEquals(1, replyInbox.peek(TARGET).size(),
|
||||
"control: the reply inbox must own TARGET and hold one message before the release "
|
||||
+ "listener runs");
|
||||
|
||||
primaryRegistry.recordDelegation(TARGET, "lead-1");
|
||||
assertEquals("lead-1", primaryRegistry.nudgeTargetFor(TARGET).orElse(null),
|
||||
"control: the delegation must be recorded before the release listener runs");
|
||||
|
||||
// --- the one call under test: invoke the REAL, assembled release listener directly, the
|
||||
// same way SessionManager.release(...) would on a real teardown.
|
||||
releaseListener.accept(new SessionManager.ReleaseDetail(TARGET, null, null, null, null));
|
||||
|
||||
MessageService.TaskView after = awaitTerminal(messages, ticket);
|
||||
assertEquals(MessageService.Phase.FAILED, after.phase(),
|
||||
"FleetdAssembly.java:447 must register a listener that calls messages.abandon(...) "
|
||||
+ "on the SAME assembled MessageService — an inert listener leaves this "
|
||||
+ "ticket PENDING for the full 30-minute async timeout");
|
||||
assertTrue(after.detail() != null && after.detail().contains("released"),
|
||||
"the abandon reason must say the worker session was released");
|
||||
|
||||
assertTrue(replyInbox.peek(TARGET).isEmpty(),
|
||||
"FleetdAssembly.java:447 must register a listener that calls replyInbox.release(...) "
|
||||
+ "— an inert listener leaves the inbox still owning TARGET with its message");
|
||||
|
||||
assertTrue(primaryRegistry.nudgeTargetFor(TARGET).isEmpty(),
|
||||
"FleetdAssembly.java:447 must register a listener that calls "
|
||||
+ "primaryRegistry.forgetDelegation(...) — an inert listener leaves the stale "
|
||||
+ "delegation in place");
|
||||
} finally {
|
||||
// Proof the teardown actually ran, not just an assurance that a finally was added: the
|
||||
// captured shutdown hook's close order (FleetdAssemblyLifecycleTest) closes the herdr
|
||||
// client last, so ports.herdr.closed flips to true only if this hook really executed.
|
||||
ports.shutdownHook.run();
|
||||
assertTrue(ports.herdr.closed, "the captured shutdown hook must have run and closed herdr — "
|
||||
+ "proof this test's assembled background loops/scheduler were torn down");
|
||||
}
|
||||
}
|
||||
|
||||
private static MessageService.TaskView awaitTerminal(MessageService messages, String ticket)
|
||||
throws InterruptedException {
|
||||
long deadline = System.currentTimeMillis() + 5000;
|
||||
MessageService.TaskView view = messages.poll(ticket);
|
||||
while (view.phase() == MessageService.Phase.PENDING && System.currentTimeMillis() < deadline) {
|
||||
//noinspection BusyWait
|
||||
Thread.sleep(10);
|
||||
view = messages.poll(ticket);
|
||||
}
|
||||
return view;
|
||||
}
|
||||
}
|
||||
+241
@@ -0,0 +1,241 @@
|
||||
package dev.ltms.fleet;
|
||||
|
||||
import dev.ltms.fleet.config.ConfigRef;
|
||||
import dev.ltms.fleet.config.FleetConfig;
|
||||
import dev.ltms.fleet.guard.SubscriptionGuard;
|
||||
import dev.ltms.fleet.herdr.FakeHerdr;
|
||||
import dev.ltms.fleet.herdr.HerdrClient;
|
||||
import dev.ltms.fleet.lead.LeadContextGauge;
|
||||
import dev.ltms.fleet.msg.LeadHeartbeatLoop;
|
||||
import io.javalin.Javalin;
|
||||
import org.junit.jupiter.api.DisplayName;
|
||||
import org.junit.jupiter.api.Test;
|
||||
import org.junit.jupiter.api.io.TempDir;
|
||||
|
||||
import java.lang.reflect.Field;
|
||||
import java.lang.reflect.Method;
|
||||
import java.nio.file.Files;
|
||||
import java.nio.file.Path;
|
||||
import java.util.Map;
|
||||
import java.util.concurrent.Executors;
|
||||
import java.util.concurrent.ScheduledExecutorService;
|
||||
import java.util.function.LongSupplier;
|
||||
|
||||
import static org.junit.jupiter.api.Assertions.assertFalse;
|
||||
import static org.junit.jupiter.api.Assertions.assertNotNull;
|
||||
import static org.junit.jupiter.api.Assertions.assertTrue;
|
||||
|
||||
/**
|
||||
* fleetd #630, and fleetd #612's ranks for this call site. {@code FleetdAssembly.java:402}
|
||||
* computes {@code requireOperatorConfirm} from the effective {@code leadRollover.requireOperatorConfirm}
|
||||
* config, and {@code :409} threads it as the 14th argument into the full {@link LeadHeartbeatLoop}
|
||||
* constructor. Measured on 26f1986: dropping that one argument so the 13-argument overload is
|
||||
* selected instead (it delegates with {@code true} hardcoded — see that overload's own javadoc,
|
||||
* fleetd #621) compiles with 0 errors and leaves all 1883 tests green, both with and without the
|
||||
* argument. In production this means the daemon keeps starting and keeps nudging, but the
|
||||
* context-high notice silently goes back to telling EVERY lead to ask the operator before a
|
||||
* context roll — on a host that set {@code requireOperatorConfirm: false} specifically so it would
|
||||
* not have to. That is the operator's own fix silently reverting, with a fully green suite.
|
||||
*
|
||||
* <p>{@code LeadHeartbeatLoopTest} already proves {@link LeadHeartbeatLoop}'s package-private
|
||||
* {@code contextNotice(boolean, LeadContextGauge.Reading, boolean, boolean)} branches correctly on
|
||||
* its own {@code requireOperatorConfirm} argument — that the METHOD works. It says nothing about
|
||||
* which value {@code FleetdAssembly} actually passes into the constructed loop, so it is not
|
||||
* reused here as coverage for the call site.
|
||||
*
|
||||
* <p>This test assembles the real daemon TWICE — once with {@code leadRollover.requireOperatorConfirm:
|
||||
* false}, once with {@code true} — pulls the REAL {@code requireOperatorConfirm} field out of the
|
||||
* REAL, assembled {@link LeadHeartbeatLoop} each time (reflection: the field, and {@code
|
||||
* contextNotice} itself, are package-private to {@code dev.ltms.fleet.msg}, and nothing public
|
||||
* exposes either — the same technique {@code StatusPollerResilienceTest} already uses in this
|
||||
* suite), and calls the REAL {@code contextNotice} method with that field's value to produce the
|
||||
* actual notice text the assembled loop would append to a nudge. Both directions are asserted: a
|
||||
* one-directional test here would pass on a constant.
|
||||
*/
|
||||
class FleetdAssemblyRequireOperatorConfirmBehaviouralTest {
|
||||
|
||||
private static final class RecordingResourcePorts implements ResourcePorts {
|
||||
final FakeHerdr herdr = new FakeHerdr();
|
||||
Runnable shutdownHook;
|
||||
|
||||
@Override
|
||||
public Map<String, String> environment() {
|
||||
return Map.of();
|
||||
}
|
||||
|
||||
@Override
|
||||
public HerdrClient connectHerdr(Path socketPath) {
|
||||
return herdr;
|
||||
}
|
||||
|
||||
@Override
|
||||
public Fleetd.AmqpOpener replyInboxOpener() {
|
||||
return (uri, prefetch) -> {
|
||||
throw new UnsupportedOperationException("no broker: block is configured");
|
||||
};
|
||||
}
|
||||
|
||||
@Override
|
||||
public Fleetd.LeadMailboxOpener leadMailboxOpener() {
|
||||
return (uri, selfCoordId, prefetch) -> {
|
||||
throw new UnsupportedOperationException("no coordinator: block is configured");
|
||||
};
|
||||
}
|
||||
|
||||
@Override
|
||||
public Runnable herdrPollWait() {
|
||||
// Never invoked: this test's FakeHerdr answers immediately, so awaitHerdr never polls.
|
||||
return () -> {
|
||||
throw new UnsupportedOperationException("herdrPollWait must not be called — herdr is healthy");
|
||||
};
|
||||
}
|
||||
@Override
|
||||
public LongSupplier nanoClock() {
|
||||
return System::nanoTime;
|
||||
}
|
||||
|
||||
@Override
|
||||
public LongSupplier wallClockNanos() {
|
||||
return System::nanoTime;
|
||||
}
|
||||
|
||||
@Override
|
||||
public ScheduledExecutorService newScheduler(String purpose) {
|
||||
return Executors.newSingleThreadScheduledExecutor();
|
||||
}
|
||||
|
||||
@Override
|
||||
public void addShutdownHook(Runnable hook) {
|
||||
this.shutdownHook = hook;
|
||||
}
|
||||
|
||||
@Override
|
||||
public void startHttp(Javalin app, String host, int port) {
|
||||
}
|
||||
}
|
||||
|
||||
private static FleetConfig writeConfig(Path dir, boolean requireOperatorConfirm) throws Exception {
|
||||
Path f = dir.resolve("fleetd.yaml");
|
||||
Files.writeString(f, """
|
||||
bind:
|
||||
host: 127.0.0.1
|
||||
port: 8765
|
||||
idleSleepGuard:
|
||||
enabled: false
|
||||
leadHeartbeat:
|
||||
idleAfterSeconds: 600
|
||||
backoffMs: 15000
|
||||
quietNudgeCap: 5
|
||||
leadRollover:
|
||||
handoverPath: handover.md
|
||||
requireOperatorConfirm: %s
|
||||
""".formatted(requireOperatorConfirm));
|
||||
return FleetConfig.load(f);
|
||||
}
|
||||
|
||||
/** Carries both the assembled loop under test AND its {@link RecordingResourcePorts}, so the
|
||||
* caller can tear the assembly down (this test assembles the real daemon TWICE — see the class
|
||||
* javadoc — and each assembly needs its own teardown, not just the last one). */
|
||||
private record Assembled(LeadHeartbeatLoop heartbeat, RecordingResourcePorts ports) {
|
||||
}
|
||||
|
||||
private static Assembled assembleHeartbeat(Path dir, boolean requireOperatorConfirm) throws Exception {
|
||||
FleetConfig cfg = writeConfig(dir, requireOperatorConfirm);
|
||||
ConfigRef config = new ConfigRef(dir.resolve("fleetd.yaml"), cfg);
|
||||
SubscriptionGuard guard = new SubscriptionGuard(cfg.guard().hostSet());
|
||||
RecordingResourcePorts ports = new RecordingResourcePorts();
|
||||
|
||||
FleetdRuntime runtime = FleetdAssembly.assembleAndStart(new AssemblyInputs(cfg, config, guard), ports);
|
||||
assertNotNull(ports.shutdownHook, "FleetdAssembly must have registered a shutdown hook");
|
||||
|
||||
LeadHeartbeatLoop heartbeat = runtime.heartbeat();
|
||||
assertNotNull(heartbeat, "control: leadHeartbeat: is configured, so FleetdAssembly.assembleAndStart "
|
||||
+ "must have built a real LeadHeartbeatLoop");
|
||||
return new Assembled(heartbeat, ports);
|
||||
}
|
||||
|
||||
/** Pulls the REAL {@code requireOperatorConfirm} field off the REAL, assembled loop. */
|
||||
private static boolean assembledRequireOperatorConfirm(LeadHeartbeatLoop heartbeat) throws Exception {
|
||||
Field field = LeadHeartbeatLoop.class.getDeclaredField("requireOperatorConfirm");
|
||||
field.setAccessible(true);
|
||||
return field.getBoolean(heartbeat);
|
||||
}
|
||||
|
||||
/** Calls the REAL, package-private {@code contextNotice(boolean, Reading, boolean, boolean)} via reflection. */
|
||||
private static String contextNotice(boolean enabled, LeadContextGauge.Reading reading, boolean alreadyNotified,
|
||||
boolean requireOperatorConfirm) throws Exception {
|
||||
Method method = LeadHeartbeatLoop.class.getDeclaredMethod("contextNotice", boolean.class,
|
||||
LeadContextGauge.Reading.class, boolean.class, boolean.class);
|
||||
method.setAccessible(true);
|
||||
return (String) method.invoke(null, enabled, reading, alreadyNotified, requireOperatorConfirm);
|
||||
}
|
||||
|
||||
@Test
|
||||
@DisplayName("[BEHAVIOURAL] the real assembled LeadHeartbeatLoop's context-high notice tracks "
|
||||
+ "leadRollover.requireOperatorConfirm — BOTH directions")
|
||||
void assembledRequireOperatorConfirmControlsNoticeWording(@TempDir Path dir) throws Exception {
|
||||
LeadContextGauge.Reading highReading = new LeadContextGauge.Reading(LeadContextGauge.State.HIGH,
|
||||
250_000L, 2);
|
||||
|
||||
String noticeFalse;
|
||||
String noticeTrue;
|
||||
|
||||
// --- direction 1: requireOperatorConfirm: false -----------------------------------------
|
||||
Path falseDir = dir.resolve("false");
|
||||
Files.createDirectories(falseDir);
|
||||
Assembled assembledFalse = assembleHeartbeat(falseDir, false);
|
||||
// Surefire runs the whole suite in one JVM fork, so each assembly's scheduler/loops must be
|
||||
// torn down here, on the failure path too — hence try/finally per assembly (this test
|
||||
// assembles TWICE, so both need their own teardown, not just the last one).
|
||||
try {
|
||||
boolean fieldFalse = assembledRequireOperatorConfirm(assembledFalse.heartbeat());
|
||||
assertFalse(fieldFalse, "FleetdAssembly.java:402/:409 must thread leadRollover."
|
||||
+ "requireOperatorConfirm: false into the assembled LeadHeartbeatLoop's own field — "
|
||||
+ "dropping the 14th constructor argument selects the 13-argument overload, which "
|
||||
+ "hardcodes true regardless of config (fleetd #621), and this would read true instead");
|
||||
|
||||
noticeFalse = contextNotice(true, highReading, false, fieldFalse);
|
||||
assertTrue(noticeFalse.contains("Decide for yourself when to confirm"),
|
||||
"with requireOperatorConfirm: false, the assembled loop's own notice must tell the "
|
||||
+ "lead it can decide for itself — got: " + noticeFalse);
|
||||
assertFalse(noticeFalse.contains("ask the operator") || noticeFalse.contains("Only the operator"),
|
||||
"with requireOperatorConfirm: false, the assembled loop's own notice must NOT ask the "
|
||||
+ "operator — got: " + noticeFalse);
|
||||
} finally {
|
||||
// Proof the teardown actually ran, not just an assurance that a finally was added: the
|
||||
// captured shutdown hook's close order (FleetdAssemblyLifecycleTest) closes the herdr
|
||||
// client last, so ports.herdr.closed flips to true only if this hook really executed.
|
||||
assembledFalse.ports().shutdownHook.run();
|
||||
assertTrue(assembledFalse.ports().herdr.closed, "the captured shutdown hook must have run "
|
||||
+ "and closed herdr — proof this assembly's background loops/scheduler were torn down");
|
||||
}
|
||||
|
||||
// --- direction 2: requireOperatorConfirm: true -------------------------------------------
|
||||
Path trueDir = dir.resolve("true");
|
||||
Files.createDirectories(trueDir);
|
||||
Assembled assembledTrue = assembleHeartbeat(trueDir, true);
|
||||
try {
|
||||
boolean fieldTrue = assembledRequireOperatorConfirm(assembledTrue.heartbeat());
|
||||
assertTrue(fieldTrue, "FleetdAssembly.java:402/:409 must thread leadRollover."
|
||||
+ "requireOperatorConfirm: true into the assembled LeadHeartbeatLoop's own field");
|
||||
|
||||
noticeTrue = contextNotice(true, highReading, false, fieldTrue);
|
||||
assertTrue(noticeTrue.contains("ask the operator") && noticeTrue.contains("Only the operator can approve the roll"),
|
||||
"with requireOperatorConfirm: true, the assembled loop's own notice must ask the "
|
||||
+ "operator — got: " + noticeTrue);
|
||||
assertFalse(noticeTrue.contains("Decide for yourself when to confirm"),
|
||||
"with requireOperatorConfirm: true, the assembled loop's own notice must NOT tell "
|
||||
+ "the lead it can decide for itself — got: " + noticeTrue);
|
||||
} finally {
|
||||
assembledTrue.ports().shutdownHook.run();
|
||||
assertTrue(assembledTrue.ports().herdr.closed, "the captured shutdown hook must have run "
|
||||
+ "and closed herdr — proof this assembly's background loops/scheduler were torn down");
|
||||
}
|
||||
|
||||
// --- the two directions must actually differ: a constant return would pass both assertion
|
||||
// blocks above vacuously if they happened to share wording, so compare them directly too.
|
||||
assertTrue(!noticeFalse.equals(noticeTrue),
|
||||
"the two directions must produce genuinely different notice text — got the same "
|
||||
+ "text for both: " + noticeFalse);
|
||||
}
|
||||
}
|
||||
@@ -0,0 +1,173 @@
|
||||
package dev.ltms.fleet;
|
||||
|
||||
import ch.qos.logback.classic.Level;
|
||||
import ch.qos.logback.classic.spi.ILoggingEvent;
|
||||
import dev.ltms.fleet.config.ConfigRef;
|
||||
import dev.ltms.fleet.config.FleetConfig;
|
||||
import dev.ltms.fleet.guard.SubscriptionGuard;
|
||||
import dev.ltms.fleet.herdr.FakeHerdr;
|
||||
import dev.ltms.fleet.herdr.HerdrClient;
|
||||
import dev.ltms.fleet.msg.ReplyInbox;
|
||||
import dev.ltms.fleet.testing.CapturedLog;
|
||||
import io.javalin.Javalin;
|
||||
import org.junit.jupiter.api.Test;
|
||||
import org.junit.jupiter.api.io.TempDir;
|
||||
|
||||
import java.nio.file.Files;
|
||||
import java.nio.file.Path;
|
||||
import java.util.List;
|
||||
import java.util.Map;
|
||||
import java.util.concurrent.Executors;
|
||||
import java.util.concurrent.ScheduledExecutorService;
|
||||
import java.util.function.LongSupplier;
|
||||
|
||||
import static org.junit.jupiter.api.Assertions.assertTrue;
|
||||
|
||||
/**
|
||||
* fleetd #612 A-gaps (gap 2): {@code Fleetd.java:185} used to call {@code
|
||||
* reportRoleFallbackGaps(cfg)} <em>outside</em> the boundary {@code FleetdAssembly.assembleAndStart}
|
||||
* — the assembly call itself sat at line 201, after it — so nothing that drives the assembly (the
|
||||
* seam {@code FleetdAssemblyLifecycleTest} exercises) could ever notice the call being deleted.
|
||||
* {@code Fleetd.main} still refuses a bad config end to end (see {@code
|
||||
* FleetdStartupValidationTest}), but that only pins {@code validateAll()} and {@code
|
||||
* assertChartersNameOnlyRegisteredTools} — a validator that <em>throws</em>. {@code
|
||||
* reportRoleFallbackGaps} only logs; nothing about {@code main} throwing or not throwing can
|
||||
* observe whether that particular call ran.
|
||||
*
|
||||
* <p>This drives {@link FleetdAssembly#assembleAndStart} directly — never a copy of its logic —
|
||||
* with a config that has no {@code fleet:} pools or charters configured for any role, so every role
|
||||
* trips both of {@code reportRoleFallbackGaps}' log branches, and asserts on the real log line a
|
||||
* {@code ListAppender} attached to the shared {@code Fleetd}/{@code FleetdAssembly} logger
|
||||
* captures. Deleting the call from {@code FleetdAssembly} (verified by hand, see the ticket) turns
|
||||
* this test red; deleting it from {@code Fleetd.main} instead (its old location) would not, which
|
||||
* is exactly the gap this test closes.
|
||||
*/
|
||||
class FleetdAssemblyRoleFallbackBoundaryTest {
|
||||
|
||||
/** Minimal fake {@link ResourcePorts}: enough for {@code assembleAndStart} to run with no real I/O. */
|
||||
private static final class RecordingResourcePorts implements ResourcePorts {
|
||||
|
||||
final FakeHerdr herdr = new FakeHerdr();
|
||||
final SentinelReplyInbox replyInbox = new SentinelReplyInbox();
|
||||
Runnable shutdownHook;
|
||||
|
||||
@Override
|
||||
public Map<String, String> environment() {
|
||||
return Map.of();
|
||||
}
|
||||
|
||||
@Override
|
||||
public HerdrClient connectHerdr(Path socketPath) {
|
||||
return herdr;
|
||||
}
|
||||
|
||||
@Override
|
||||
public Fleetd.AmqpOpener replyInboxOpener() {
|
||||
return (uri, prefetch) -> replyInbox;
|
||||
}
|
||||
|
||||
@Override
|
||||
public Fleetd.LeadMailboxOpener leadMailboxOpener() {
|
||||
return (uri, selfCoordId, prefetch) -> {
|
||||
throw new UnsupportedOperationException(
|
||||
"leadMailboxOpener must not be called — no coordinator: block is configured");
|
||||
};
|
||||
}
|
||||
|
||||
@Override
|
||||
public LongSupplier nanoClock() {
|
||||
return System::nanoTime;
|
||||
}
|
||||
|
||||
@Override
|
||||
public LongSupplier wallClockNanos() {
|
||||
return System::nanoTime;
|
||||
}
|
||||
|
||||
@Override
|
||||
public ScheduledExecutorService newScheduler(String purpose) {
|
||||
return Executors.newSingleThreadScheduledExecutor();
|
||||
}
|
||||
|
||||
@Override
|
||||
public void addShutdownHook(Runnable hook) {
|
||||
this.shutdownHook = hook;
|
||||
}
|
||||
|
||||
@Override
|
||||
public void startHttp(Javalin app, String host, int port) {
|
||||
// No real HTTP bind in a unit test.
|
||||
}
|
||||
|
||||
@Override
|
||||
public Runnable herdrPollWait() {
|
||||
// Never invoked: this test's FakeHerdr answers immediately, so awaitHerdr never polls.
|
||||
return () -> {
|
||||
throw new UnsupportedOperationException("herdrPollWait must not be called — herdr is healthy");
|
||||
};
|
||||
}
|
||||
}
|
||||
|
||||
private static final class SentinelReplyInbox implements ReplyInbox {
|
||||
@Override
|
||||
public void own(String target) {
|
||||
}
|
||||
|
||||
@Override
|
||||
public void release(String target) {
|
||||
}
|
||||
|
||||
@Override
|
||||
public void publish(String target, String msgId, String content) {
|
||||
}
|
||||
|
||||
@Override
|
||||
public List<InboxMessage> peek(String target) {
|
||||
return List.of();
|
||||
}
|
||||
|
||||
@Override
|
||||
public boolean ack(String target, String msgId) {
|
||||
return false;
|
||||
}
|
||||
}
|
||||
|
||||
private static FleetConfig writeConfig(Path dir) throws Exception {
|
||||
Path f = dir.resolve("fleetd.yaml");
|
||||
// Deliberately no `fleet:` block at all: every MemberRole has neither a pool nor a
|
||||
// charter, so reportRoleFallbackGaps' "role fallback: no fleet.<role>s: pool for ..."
|
||||
// branch is guaranteed to log something to assert on.
|
||||
Files.writeString(f, """
|
||||
bind:
|
||||
host: 127.0.0.1
|
||||
port: 8765
|
||||
idleSleepGuard:
|
||||
enabled: false
|
||||
""");
|
||||
return FleetConfig.load(f);
|
||||
}
|
||||
|
||||
@Test
|
||||
void assembleAndStartItselfReportsTheRoleFallbackGaps(@TempDir Path dir) throws Exception {
|
||||
FleetConfig cfg = writeConfig(dir);
|
||||
ConfigRef config = new ConfigRef(dir.resolve("fleetd.yaml"), cfg);
|
||||
SubscriptionGuard guard = new SubscriptionGuard(cfg.guard().hostSet());
|
||||
RecordingResourcePorts ports = new RecordingResourcePorts();
|
||||
|
||||
try (var captured = CapturedLog.at(Fleetd.class, Level.INFO)) {
|
||||
FleetdAssembly.assembleAndStart(new AssemblyInputs(cfg, config, guard), ports);
|
||||
|
||||
String infoLines = captured.events().stream()
|
||||
.filter(e -> e.getLevel() == Level.INFO)
|
||||
.map(ILoggingEvent::getFormattedMessage)
|
||||
.reduce("", (a, b) -> a + "\n" + b);
|
||||
assertTrue(infoLines.contains("role fallback: no fleet.<role>s: pool for"),
|
||||
() -> "FleetdAssembly.assembleAndStart itself must call reportRoleFallbackGaps "
|
||||
+ "(fleetd #612 A-gaps gap 2) — captured INFO lines: " + infoLines);
|
||||
} finally {
|
||||
if (ports.shutdownHook != null) {
|
||||
ports.shutdownHook.run();
|
||||
}
|
||||
}
|
||||
}
|
||||
}
|
||||
@@ -0,0 +1,190 @@
|
||||
package dev.ltms.fleet;
|
||||
|
||||
import dev.ltms.fleet.config.ConfigRef;
|
||||
import dev.ltms.fleet.config.FleetConfig;
|
||||
import dev.ltms.fleet.guard.SubscriptionGuard;
|
||||
import dev.ltms.fleet.herdr.FakeHerdr;
|
||||
import dev.ltms.fleet.herdr.HerdrClient;
|
||||
import dev.ltms.fleet.inject.Injector;
|
||||
import dev.ltms.fleet.inject.StatusPoller;
|
||||
import dev.ltms.fleet.msg.LeadChannelHandle;
|
||||
import dev.ltms.fleet.msg.LeadCoordLoop;
|
||||
import dev.ltms.fleet.msg.LeadMessage;
|
||||
import dev.ltms.fleet.msg.ReplyInbox;
|
||||
import dev.ltms.fleet.msg.ReplyPushLoop;
|
||||
import io.javalin.Javalin;
|
||||
import org.junit.jupiter.api.AfterEach;
|
||||
import org.junit.jupiter.api.Test;
|
||||
import org.junit.jupiter.api.io.TempDir;
|
||||
|
||||
import java.lang.reflect.Field;
|
||||
import java.nio.file.Files;
|
||||
import java.nio.file.Path;
|
||||
import java.util.List;
|
||||
import java.util.Map;
|
||||
import java.util.concurrent.Executors;
|
||||
import java.util.concurrent.ScheduledExecutorService;
|
||||
import java.util.function.LongSupplier;
|
||||
|
||||
import static org.junit.jupiter.api.Assertions.assertEquals;
|
||||
import static org.junit.jupiter.api.Assertions.assertNotNull;
|
||||
|
||||
/**
|
||||
* Asserts that the assembled loops use the production reminder, coordination, and delivery timing
|
||||
* defaults when no {@code primary:} block configures the reply-push values.
|
||||
*/
|
||||
class FleetdAssemblyTimingDefaultsTest {
|
||||
|
||||
private static final class FakeLeadChannel implements LeadChannelHandle {
|
||||
@Override
|
||||
public void publish(String toCoordId, LeadMessage message) {
|
||||
}
|
||||
|
||||
@Override
|
||||
public List<LeadMessage> peek() {
|
||||
return List.of();
|
||||
}
|
||||
|
||||
@Override
|
||||
public void ack(String msgId) {
|
||||
}
|
||||
|
||||
@Override
|
||||
public String selfCoordId() {
|
||||
return "test-lead";
|
||||
}
|
||||
|
||||
@Override
|
||||
public boolean heldDurable() {
|
||||
return true;
|
||||
}
|
||||
|
||||
@Override
|
||||
public MailboxState inspect(String coordId) {
|
||||
return MailboxState.unknown(coordId);
|
||||
}
|
||||
|
||||
@Override
|
||||
public void close() {
|
||||
}
|
||||
}
|
||||
|
||||
private static final class TestResourcePorts implements ResourcePorts {
|
||||
final FakeHerdr herdr = new FakeHerdr();
|
||||
Runnable shutdownHook;
|
||||
|
||||
@Override
|
||||
public Map<String, String> environment() {
|
||||
return Map.of();
|
||||
}
|
||||
|
||||
@Override
|
||||
public HerdrClient connectHerdr(Path socketPath) {
|
||||
return herdr;
|
||||
}
|
||||
|
||||
@Override
|
||||
public Fleetd.AmqpOpener replyInboxOpener() {
|
||||
return (uri, prefetch) -> new ReplyInbox() {
|
||||
@Override public void own(String target) { }
|
||||
@Override public void release(String target) { }
|
||||
@Override public void publish(String target, String msgId, String content) { }
|
||||
@Override public List<InboxMessage> peek(String target) { return List.of(); }
|
||||
@Override public boolean ack(String target, String msgId) { return false; }
|
||||
};
|
||||
}
|
||||
|
||||
@Override
|
||||
public Fleetd.LeadMailboxOpener leadMailboxOpener() {
|
||||
return (uri, selfCoordId, prefetch) -> new FakeLeadChannel();
|
||||
}
|
||||
|
||||
@Override
|
||||
public LongSupplier nanoClock() {
|
||||
return System::nanoTime;
|
||||
}
|
||||
|
||||
@Override
|
||||
public LongSupplier wallClockNanos() {
|
||||
return System::nanoTime;
|
||||
}
|
||||
|
||||
@Override
|
||||
public ScheduledExecutorService newScheduler(String purpose) {
|
||||
return Executors.newSingleThreadScheduledExecutor();
|
||||
}
|
||||
|
||||
@Override
|
||||
public void addShutdownHook(Runnable hook) {
|
||||
shutdownHook = hook;
|
||||
}
|
||||
|
||||
@Override
|
||||
public void startHttp(Javalin app, String host, int port) {
|
||||
}
|
||||
|
||||
@Override
|
||||
public Runnable herdrPollWait() {
|
||||
return () -> {
|
||||
throw new UnsupportedOperationException("FakeHerdr is healthy; no poll wait is expected");
|
||||
};
|
||||
}
|
||||
}
|
||||
|
||||
private TestResourcePorts ports;
|
||||
|
||||
@AfterEach
|
||||
void tearDown() {
|
||||
if (ports != null && ports.shutdownHook != null) {
|
||||
ports.shutdownHook.run();
|
||||
}
|
||||
}
|
||||
|
||||
private static FleetConfig writeConfig(Path dir) throws Exception {
|
||||
Path file = dir.resolve("fleetd.yaml");
|
||||
Files.writeString(file, """
|
||||
bind:
|
||||
host: 127.0.0.1
|
||||
port: 8765
|
||||
idleSleepGuard:
|
||||
enabled: false
|
||||
coordinator:
|
||||
uri: "amqp://fake-lead-broker/vh"
|
||||
selfId: "test-lead"
|
||||
""");
|
||||
return FleetConfig.load(file);
|
||||
}
|
||||
|
||||
private FleetdRuntime assemble(Path dir) throws Exception {
|
||||
FleetConfig cfg = writeConfig(dir);
|
||||
ports = new TestResourcePorts();
|
||||
return FleetdAssembly.assembleAndStart(new AssemblyInputs(cfg,
|
||||
new ConfigRef(dir.resolve("fleetd.yaml"), cfg), new SubscriptionGuard(cfg.guard().hostSet())), ports);
|
||||
}
|
||||
|
||||
private static long longField(Object target, String name) throws Exception {
|
||||
Field field = target.getClass().getDeclaredField(name);
|
||||
field.setAccessible(true);
|
||||
return field.getLong(target);
|
||||
}
|
||||
|
||||
@Test
|
||||
void productionBootPathUsesTheExpectedLoopTimingDefaults(@TempDir Path dir) throws Exception {
|
||||
FleetdRuntime runtime = assemble(dir);
|
||||
|
||||
ReplyPushLoop pushLoop = runtime.pushLoop();
|
||||
assertEquals(5, longField(pushLoop, "maxReminders"),
|
||||
"without primary:, ReplyPushLoop must stop after five reminder attempts");
|
||||
assertEquals(15_000L, longField(pushLoop, "backoffMs"),
|
||||
"without primary:, ReplyPushLoop must wait fifteen seconds before the next reminder");
|
||||
|
||||
LeadCoordLoop leadCoordLoop = runtime.leadCoordLoop();
|
||||
assertNotNull(leadCoordLoop, "control: coordinator: must build LeadCoordLoop");
|
||||
assertEquals(3_000L, longField(leadCoordLoop, "intervalMs"),
|
||||
"LeadCoordLoop must poll for peer-lead mail every three seconds");
|
||||
|
||||
StatusPoller poller = runtime.poller();
|
||||
assertEquals(Injector.POLL_INTERVAL_MILLIS, longField(poller, "intervalMillis"),
|
||||
"StatusPoller must use Injector's delivery poll interval");
|
||||
}
|
||||
}
|
||||
@@ -0,0 +1,184 @@
|
||||
package dev.ltms.fleet;
|
||||
|
||||
import dev.ltms.fleet.config.ConfigRef;
|
||||
import dev.ltms.fleet.config.FleetConfig;
|
||||
import dev.ltms.fleet.guard.SubscriptionGuard;
|
||||
import dev.ltms.fleet.herdr.AgentStatus;
|
||||
import dev.ltms.fleet.herdr.FakeHerdr;
|
||||
import dev.ltms.fleet.herdr.HerdrClient;
|
||||
import dev.ltms.fleet.inject.Injector;
|
||||
import dev.ltms.fleet.inject.TurnListener;
|
||||
import dev.ltms.fleet.msg.Rendezvous;
|
||||
import dev.ltms.fleet.msg.TurnToken;
|
||||
import io.javalin.Javalin;
|
||||
import org.junit.jupiter.api.Test;
|
||||
import org.junit.jupiter.api.io.TempDir;
|
||||
|
||||
import java.lang.reflect.Field;
|
||||
import java.nio.file.Files;
|
||||
import java.nio.file.Path;
|
||||
import java.util.Map;
|
||||
import java.util.concurrent.CompletableFuture;
|
||||
import java.util.concurrent.Executors;
|
||||
import java.util.concurrent.ScheduledExecutorService;
|
||||
import java.util.concurrent.TimeUnit;
|
||||
import java.util.concurrent.atomic.AtomicLong;
|
||||
import java.util.function.LongSupplier;
|
||||
|
||||
import static org.junit.jupiter.api.Assertions.assertThrows;
|
||||
import static org.junit.jupiter.api.Assertions.assertTrue;
|
||||
|
||||
/**
|
||||
* fleetd #612 rank 12 — {@code FleetdAssembly} supplies the {@link Injector}'s {@code TurnRegistrar}
|
||||
* with {@code Fleetd.turnRegistrar(completion)}. The normal delivery path cannot distinguish that
|
||||
* registrar from {@code TurnRegistrar.NOOP}: {@link dev.ltms.fleet.inject.CompletionResolver#onDelivered}
|
||||
* registers the same turn shortly afterwards. The distinction matters when a delivery listener throws
|
||||
* after the pane received the message but before normal completion runs. The registrar must already have
|
||||
* registered the waiter with the REAL assembled resolver, so a later completion can still resolve it.
|
||||
*
|
||||
* <p>The test installs a throwing wrapper around the real assembled listener. It does not replace the
|
||||
* registrar or resolver. This constructs the narrow failure condition without a sleep, then observes the
|
||||
* real resolver through {@link FleetdRuntime#completion()}. An inert registrar, or a registrar wired to a
|
||||
* throwaway resolver, leaves this waiter's turn absent from the real resolver and makes the final assertion
|
||||
* fail.
|
||||
*/
|
||||
class FleetdAssemblyTurnRegistrarBehaviouralTest {
|
||||
|
||||
private static final String TARGET = "term_a";
|
||||
|
||||
private static final class ControllableResourcePorts implements ResourcePorts {
|
||||
final FakeHerdr herdr = new FakeHerdr();
|
||||
final AtomicLong nowNanos = new AtomicLong(1_000_000_000L);
|
||||
Runnable shutdownHook;
|
||||
|
||||
void advanceSeconds(long seconds) {
|
||||
nowNanos.addAndGet(TimeUnit.SECONDS.toNanos(seconds));
|
||||
}
|
||||
|
||||
@Override
|
||||
public Map<String, String> environment() {
|
||||
return Map.of();
|
||||
}
|
||||
|
||||
@Override
|
||||
public HerdrClient connectHerdr(Path socketPath) {
|
||||
return herdr;
|
||||
}
|
||||
|
||||
@Override
|
||||
public Fleetd.AmqpOpener replyInboxOpener() {
|
||||
return (uri, prefetch) -> {
|
||||
throw new UnsupportedOperationException("no broker: block is configured");
|
||||
};
|
||||
}
|
||||
|
||||
@Override
|
||||
public Fleetd.LeadMailboxOpener leadMailboxOpener() {
|
||||
return (uri, selfCoordId, prefetch) -> {
|
||||
throw new UnsupportedOperationException("no coordinator: block is configured");
|
||||
};
|
||||
}
|
||||
|
||||
@Override
|
||||
public LongSupplier nanoClock() {
|
||||
return nowNanos::get;
|
||||
}
|
||||
|
||||
@Override
|
||||
public LongSupplier wallClockNanos() {
|
||||
return nowNanos::get;
|
||||
}
|
||||
|
||||
@Override
|
||||
public ScheduledExecutorService newScheduler(String purpose) {
|
||||
return Executors.newSingleThreadScheduledExecutor();
|
||||
}
|
||||
|
||||
@Override
|
||||
public void addShutdownHook(Runnable hook) {
|
||||
shutdownHook = hook;
|
||||
}
|
||||
|
||||
@Override
|
||||
public void startHttp(Javalin app, String host, int port) {
|
||||
// Do not bind a real port in this assembly test.
|
||||
}
|
||||
|
||||
@Override
|
||||
public Runnable herdrPollWait() {
|
||||
return () -> {
|
||||
throw new UnsupportedOperationException("FakeHerdr is healthy; no poll wait is expected");
|
||||
};
|
||||
}
|
||||
}
|
||||
|
||||
private static FleetConfig writeConfig(Path dir) throws Exception {
|
||||
Path file = dir.resolve("fleetd.yaml");
|
||||
Files.writeString(file, """
|
||||
bind:
|
||||
host: 127.0.0.1
|
||||
port: 8765
|
||||
idleSleepGuard:
|
||||
enabled: false
|
||||
fleet:
|
||||
leaders:
|
||||
primary:
|
||||
tab: "lead: primary"
|
||||
profile: sonnet
|
||||
profiles:
|
||||
sonnet:
|
||||
subscription: true
|
||||
argv: ["ccs", "sonnet"]
|
||||
""");
|
||||
return FleetConfig.load(file);
|
||||
}
|
||||
|
||||
@Test
|
||||
void realAssembledResolverStillResolvesAfterDeliveredListenerThrows(@TempDir Path dir) throws Exception {
|
||||
FleetConfig cfg = writeConfig(dir);
|
||||
ControllableResourcePorts ports = new ControllableResourcePorts();
|
||||
ports.herdr.withTab("w2", "w2:t7", "lead: primary");
|
||||
FleetdRuntime runtime = FleetdAssembly.assembleAndStart(new AssemblyInputs(cfg,
|
||||
new ConfigRef(dir.resolve("fleetd.yaml"), cfg), new SubscriptionGuard(cfg.guard().hostSet())), ports);
|
||||
try {
|
||||
Injector injector = runtime.injector();
|
||||
installThrowingDeliveredListener(injector);
|
||||
|
||||
CompletableFuture<Rendezvous.Resolution> waiter = new CompletableFuture<>();
|
||||
injector.enqueue(TARGET, "brief", new TurnToken(TARGET, waiter));
|
||||
|
||||
IllegalStateException thrown = assertThrows(IllegalStateException.class,
|
||||
() -> injector.onStatus(TARGET, AgentStatus.IDLE),
|
||||
"control: delivery must reach the installed listener and it must throw after delivery");
|
||||
assertTrue(thrown.getMessage().contains("listener failure"), thrown::getMessage);
|
||||
|
||||
ports.herdr.readText("worker report after the listener failure");
|
||||
ports.advanceSeconds(3);
|
||||
runtime.completion().resolveBeforePostAction(TARGET);
|
||||
|
||||
assertTrue(waiter.isDone(),
|
||||
"FleetdAssembly must wire the Injector registrar to this runtime's real CompletionResolver: "
|
||||
+ "after a delivered listener throws, resolveBeforePostAction must still find and "
|
||||
+ "resolve the registered waiter");
|
||||
} finally {
|
||||
assertTrue(ports.shutdownHook != null, "control: assembly must capture its shutdown hook");
|
||||
ports.shutdownHook.run();
|
||||
assertTrue(ports.herdr.closed, "teardown control: the captured shutdown hook must close herdr");
|
||||
}
|
||||
}
|
||||
|
||||
private static void installThrowingDeliveredListener(Injector injector) throws Exception {
|
||||
Field field = Injector.class.getDeclaredField("turnListener");
|
||||
field.setAccessible(true);
|
||||
field.set(injector, new TurnListener() {
|
||||
@Override
|
||||
public void onTurnComplete(String target) {
|
||||
}
|
||||
|
||||
@Override
|
||||
public void onDelivered(String target, TurnToken token) {
|
||||
throw new IllegalStateException("listener failure after delivery");
|
||||
}
|
||||
});
|
||||
}
|
||||
}
|
||||
@@ -0,0 +1,207 @@
|
||||
package dev.ltms.fleet;
|
||||
|
||||
import dev.ltms.fleet.config.ConfigRef;
|
||||
import dev.ltms.fleet.config.FleetConfig;
|
||||
import dev.ltms.fleet.guard.SubscriptionGuard;
|
||||
import dev.ltms.fleet.herdr.FakeHerdr;
|
||||
import dev.ltms.fleet.herdr.HerdrClient;
|
||||
import dev.ltms.fleet.msg.ReplyInbox;
|
||||
import dev.ltms.fleet.placement.BackendQuarantine;
|
||||
import io.javalin.Javalin;
|
||||
import org.junit.jupiter.api.DisplayName;
|
||||
import org.junit.jupiter.api.Test;
|
||||
import org.junit.jupiter.api.io.TempDir;
|
||||
|
||||
import java.nio.file.Files;
|
||||
import java.nio.file.Path;
|
||||
import java.util.List;
|
||||
import java.util.Map;
|
||||
import java.util.OptionalLong;
|
||||
import java.util.concurrent.Executors;
|
||||
import java.util.concurrent.ScheduledExecutorService;
|
||||
import java.util.concurrent.atomic.AtomicLong;
|
||||
import java.util.function.LongSupplier;
|
||||
|
||||
import static org.junit.jupiter.api.Assertions.assertEquals;
|
||||
import static org.junit.jupiter.api.Assertions.assertTrue;
|
||||
|
||||
/**
|
||||
* fleetd #612 B3 — replaces {@code FleetdBackendQuarantineWiringTest} (fleetd #466), a source-text
|
||||
* test that scraped {@code Fleetd.java} (now {@code FleetdAssembly.java}, moved there by fleetd #612
|
||||
* Unit A) for the {@code BackendQuarantine.withEscalation(...)} call, and separately asserted the
|
||||
* flat two-argument constructor's text was ABSENT. That proves the right method NAME appears in
|
||||
* source; it proves nothing about what the constructed object actually DOES.
|
||||
*
|
||||
* <p>This test instead drives the REAL {@link BackendQuarantine} the real {@link
|
||||
* FleetdAssembly#assembleAndStart} builds — reached through {@link
|
||||
* dev.ltms.fleet.mcp.FleetMcp#quarantineSource()} on the real, live {@code FleetMcp} {@code
|
||||
* FleetdRuntime} owns — and asserts the ONE behavioural difference {@code withEscalation} and the
|
||||
* flat constructor actually produce (see {@link BackendQuarantine}'s own class doc, "Mechanism"):
|
||||
* quarantining the same credential twice in a row, within one base cooldown of the first deadline,
|
||||
* must escalate the second cooldown past the first. A flat instance reports the identical cooldown
|
||||
* both times.
|
||||
*/
|
||||
class FleetdBackendQuarantineAssemblyTest {
|
||||
|
||||
/** Base cooldown used throughout — long enough that rounding never blurs the 2x escalation. */
|
||||
private static final int COOLDOWN_SECONDS = 100;
|
||||
|
||||
private static final class RecordingResourcePorts implements ResourcePorts {
|
||||
|
||||
final FakeHerdr herdr = new FakeHerdr();
|
||||
final SentinelReplyInbox replyInbox = new SentinelReplyInbox();
|
||||
final AtomicLong clockNanos = new AtomicLong(0L);
|
||||
|
||||
@Override
|
||||
public Map<String, String> environment() {
|
||||
return Map.of();
|
||||
}
|
||||
|
||||
@Override
|
||||
public HerdrClient connectHerdr(Path socketPath) {
|
||||
return herdr;
|
||||
}
|
||||
|
||||
@Override
|
||||
public Fleetd.AmqpOpener replyInboxOpener() {
|
||||
return (uri, prefetch) -> replyInbox;
|
||||
}
|
||||
|
||||
@Override
|
||||
public Fleetd.LeadMailboxOpener leadMailboxOpener() {
|
||||
return (uri, selfCoordId, prefetch) -> {
|
||||
throw new UnsupportedOperationException(
|
||||
"leadMailboxOpener must not be called — no coordinator: block is configured");
|
||||
};
|
||||
}
|
||||
|
||||
@Override
|
||||
public LongSupplier nanoClock() {
|
||||
// Controllable: the SAME LongSupplier instance BackendQuarantine.withEscalation(...) is
|
||||
// built with, so advancing clockNanos after assembly moves the quarantine tracker's own
|
||||
// clock, with no real sleep needed to observe escalation.
|
||||
return clockNanos::get;
|
||||
}
|
||||
|
||||
@Override
|
||||
public LongSupplier wallClockNanos() {
|
||||
return clockNanos::get;
|
||||
}
|
||||
|
||||
@Override
|
||||
public ScheduledExecutorService newScheduler(String purpose) {
|
||||
return Executors.newSingleThreadScheduledExecutor();
|
||||
}
|
||||
|
||||
@Override
|
||||
public void addShutdownHook(Runnable hook) {
|
||||
}
|
||||
|
||||
@Override
|
||||
public void startHttp(Javalin app, String host, int port) {
|
||||
}
|
||||
|
||||
@Override
|
||||
public Runnable herdrPollWait() {
|
||||
// Never invoked: this test's FakeHerdr answers immediately, so awaitHerdr never polls.
|
||||
return () -> {
|
||||
throw new UnsupportedOperationException("herdrPollWait must not be called — herdr is healthy");
|
||||
};
|
||||
}
|
||||
}
|
||||
|
||||
private static final class SentinelReplyInbox implements ReplyInbox, AutoCloseable {
|
||||
@Override
|
||||
public void own(String target) {
|
||||
}
|
||||
|
||||
@Override
|
||||
public void release(String target) {
|
||||
}
|
||||
|
||||
@Override
|
||||
public void publish(String target, String msgId, String content) {
|
||||
}
|
||||
|
||||
@Override
|
||||
public List<InboxMessage> peek(String target) {
|
||||
return List.of();
|
||||
}
|
||||
|
||||
@Override
|
||||
public boolean ack(String target, String msgId) {
|
||||
return false;
|
||||
}
|
||||
|
||||
@Override
|
||||
public void close() {
|
||||
}
|
||||
}
|
||||
|
||||
private static FleetConfig writeConfig(Path dir) throws Exception {
|
||||
Path f = dir.resolve("fleetd.yaml");
|
||||
Files.writeString(f, """
|
||||
bind:
|
||||
host: 127.0.0.1
|
||||
port: 8765
|
||||
idleSleepGuard:
|
||||
enabled: false
|
||||
broker:
|
||||
uri: "amqp://fake-test-broker/vh"
|
||||
quarantineCooldownSeconds: %d
|
||||
""".formatted(COOLDOWN_SECONDS));
|
||||
return FleetConfig.load(f);
|
||||
}
|
||||
|
||||
@Test
|
||||
@DisplayName("[BEHAVIOURAL] the real assembled BackendQuarantine escalates a repeated exhaustion, "
|
||||
+ "which the flat two-argument constructor can never do")
|
||||
void assembledQuarantineEscalatesOnARepeatedExhaustion(@TempDir Path dir) throws Exception {
|
||||
FleetConfig cfg = writeConfig(dir);
|
||||
ConfigRef config = new ConfigRef(dir.resolve("fleetd.yaml"), cfg);
|
||||
SubscriptionGuard guard = new SubscriptionGuard(cfg.guard().hostSet());
|
||||
RecordingResourcePorts ports = new RecordingResourcePorts();
|
||||
|
||||
FleetdRuntime runtime = FleetdAssembly.assembleAndStart(new AssemblyInputs(cfg, config, guard), ports);
|
||||
|
||||
BackendQuarantine quarantine = runtime.mcp().quarantineSource().quarantine();
|
||||
|
||||
// First exhaustion, at clock=0: a fresh occurrence, blocked for exactly the base cooldown.
|
||||
quarantine.quarantine("cred-x");
|
||||
BackendQuarantine.Status first = quarantine.status("cred-x").orElseThrow(
|
||||
() -> new AssertionError("credential must be quarantined immediately after quarantine()"));
|
||||
assertEquals(1, first.repeatCount(), "the first call is repeat #1");
|
||||
assertEquals(COOLDOWN_SECONDS, first.remainingSeconds(),
|
||||
"a fresh quarantine blocks for exactly the base cooldown");
|
||||
|
||||
// Second exhaustion, arriving just after the first deadline — well within one base cooldown
|
||||
// of it, so this is a CONTINUATION of the same streak (repeat #2), not a fresh occurrence.
|
||||
long firstDeadlineNanos = COOLDOWN_SECONDS * 1_000_000_000L;
|
||||
ports.clockNanos.set(firstDeadlineNanos + 1);
|
||||
quarantine.quarantine("cred-x");
|
||||
BackendQuarantine.Status second = quarantine.status("cred-x").orElseThrow(
|
||||
() -> new AssertionError("credential must be quarantined immediately after the second "
|
||||
+ "quarantine() call"));
|
||||
assertEquals(2, second.repeatCount(), "the second call, arriving within one base cooldown of "
|
||||
+ "the first deadline, continues the streak as repeat #2");
|
||||
|
||||
// The one behavioural difference: withEscalation doubles the cooldown on repeat #2 (capped
|
||||
// well above this at 12x base), the flat two-argument constructor never grows past the base
|
||||
// cooldown no matter how many times quarantine() is called in a row.
|
||||
assertEquals(2 * COOLDOWN_SECONDS, second.remainingSeconds(),
|
||||
"withEscalation's default backoff doubles the cooldown on the second consecutive "
|
||||
+ "exhaustion — this is the exact call FleetdAssembly.java makes at the "
|
||||
+ "BackendQuarantine.withEscalation(...) call site");
|
||||
assertTrue(second.remainingSeconds() > first.remainingSeconds(),
|
||||
"the flat two-argument BackendQuarantine constructor would report the SAME remaining "
|
||||
+ "seconds both times — this inequality is what a mutation to the flat "
|
||||
+ "constructor at that call site must fail");
|
||||
|
||||
// Also confirm isQuarantined/remainingSeconds agree, exercising the accessors a real caller
|
||||
// (fleet_profiles / fleet_list, per BackendQuarantine's own class doc) actually reads.
|
||||
assertTrue(quarantine.isQuarantined("cred-x"));
|
||||
OptionalLong remaining = quarantine.remainingSeconds("cred-x");
|
||||
assertTrue(remaining.isPresent());
|
||||
assertEquals(2 * COOLDOWN_SECONDS, remaining.getAsLong());
|
||||
}
|
||||
}
|
||||
@@ -1,65 +0,0 @@
|
||||
package dev.ltms.fleet;
|
||||
|
||||
import java.nio.file.Files;
|
||||
import java.nio.file.Path;
|
||||
import org.junit.jupiter.api.DisplayName;
|
||||
import org.junit.jupiter.api.Test;
|
||||
|
||||
import static org.junit.jupiter.api.Assertions.assertFalse;
|
||||
import static org.junit.jupiter.api.Assertions.assertTrue;
|
||||
|
||||
/**
|
||||
* fleetd #466 follow-up: {@code Fleetd.main} builds the daemon's one {@code BackendQuarantine}
|
||||
* from {@link dev.ltms.fleet.placement.BackendQuarantine#withEscalation(java.util.function.LongSupplier,
|
||||
* long)} — the escalating factory — rather than the plain two-argument constructor, which is still a
|
||||
* flat cooldown (kept for backward compatibility, see that class's doc). {@code
|
||||
* BackendQuarantineTest} proves {@code withEscalation} itself escalates, is ceilinged, and resets;
|
||||
* it says nothing about which one {@code main} actually calls.
|
||||
*
|
||||
* <p>Measured directly: reverting {@code main} to {@code new BackendQuarantine(System::nanoTime,
|
||||
* TimeUnit.SECONDS.toNanos(cfg.quarantineCooldownSeconds()))} — the pre-#466 flat call — compiles
|
||||
* with 0 errors and leaves the entire 1608-test suite (including every {@code BackendQuarantineTest}
|
||||
* case) green, because no other test constructs its {@code BackendQuarantine} through {@code main};
|
||||
* every one of them builds its own instance directly. That silent regression is exactly the shape
|
||||
* {@link FleetdLeadSeatWiringTest} and {@link FleetdCompletionResolverWiringTest} already guard
|
||||
* against for their own constructor arguments — this is the same class of gap for fleetd #466's
|
||||
* factory choice, following their approach.
|
||||
*
|
||||
* <p><b>This test checks source text, not runtime behaviour.</b> It never constructs a {@code
|
||||
* BackendQuarantine} and never runs {@code main} — a green result here proves only that the exact
|
||||
* text {@code main} calls {@code BackendQuarantine.withEscalation(...)} rather than the flat
|
||||
* constructor. It does not prove that call actually executes at startup (no test here starts the
|
||||
* daemon), and it does not prove the escalation reaches a real backend or credential — only
|
||||
* {@code BackendQuarantineTest} proves the factory's own behaviour, and only a live daemon proves
|
||||
* the wiring runs.
|
||||
*/
|
||||
class FleetdBackendQuarantineWiringTest {
|
||||
|
||||
private static String fleetdSource() throws Exception {
|
||||
return Files.readString(Path.of("src/main/java/dev/ltms/fleet/Fleetd.java"));
|
||||
}
|
||||
|
||||
@Test
|
||||
@DisplayName("[SOURCE TEXT] main's BackendQuarantine local is still built from BackendQuarantine.withEscalation(...)")
|
||||
void mainStillWiresTheEscalatingQuarantineFactory() throws Exception {
|
||||
String source = fleetdSource();
|
||||
assertTrue(source.contains(
|
||||
"BackendQuarantine quarantine = BackendQuarantine.withEscalation(System::nanoTime,\n"
|
||||
+ " TimeUnit.SECONDS.toNanos(cfg.quarantineCooldownSeconds()));"),
|
||||
"Fleetd.main's BackendQuarantine local must still be built from "
|
||||
+ "BackendQuarantine.withEscalation(System::nanoTime, "
|
||||
+ "TimeUnit.SECONDS.toNanos(cfg.quarantineCooldownSeconds())). Reverting to the flat "
|
||||
+ "two-argument constructor (fleetd #466's measured regression) compiles with 0 errors "
|
||||
+ "and leaves the whole suite green, including every BackendQuarantineTest case that "
|
||||
+ "proves the escalation itself works — this source check is what must go red instead. "
|
||||
+ "A reverted daemon would go back to retrying a weekly subscription limit on every "
|
||||
+ "flat ~30-minute cooldown, about 336 times across the week.");
|
||||
|
||||
// Negative form of the same check: the pre-#466 flat call, if it ever reappears at this
|
||||
// declaration, must not be mistaken for the escalating one by a looser positive-only check.
|
||||
assertFalse(source.contains(
|
||||
"BackendQuarantine quarantine = new BackendQuarantine(System::nanoTime,\n"
|
||||
+ " TimeUnit.SECONDS.toNanos(cfg.quarantineCooldownSeconds()));"),
|
||||
"main's BackendQuarantine local must never regress to the flat two-argument constructor");
|
||||
}
|
||||
}
|
||||
@@ -0,0 +1,315 @@
|
||||
package dev.ltms.fleet;
|
||||
|
||||
import dev.ltms.fleet.config.ConfigRef;
|
||||
import dev.ltms.fleet.config.FleetConfig;
|
||||
import dev.ltms.fleet.guard.SubscriptionGuard;
|
||||
import dev.ltms.fleet.herdr.FakeHerdr;
|
||||
import dev.ltms.fleet.herdr.HerdrClient;
|
||||
import dev.ltms.fleet.inject.CompletionResolver;
|
||||
import dev.ltms.fleet.msg.Rendezvous;
|
||||
import dev.ltms.fleet.msg.TurnToken;
|
||||
import dev.ltms.fleet.placement.PlacementException;
|
||||
import dev.ltms.fleet.session.MemberSession;
|
||||
import dev.ltms.fleet.session.WorktreeRequest;
|
||||
import io.javalin.Javalin;
|
||||
import org.junit.jupiter.api.Test;
|
||||
import org.junit.jupiter.api.io.TempDir;
|
||||
|
||||
import java.nio.file.Files;
|
||||
import java.nio.file.Path;
|
||||
import java.util.List;
|
||||
import java.util.Map;
|
||||
import java.util.concurrent.CompletableFuture;
|
||||
import java.util.concurrent.Executors;
|
||||
import java.util.concurrent.ScheduledExecutorService;
|
||||
import java.util.concurrent.TimeUnit;
|
||||
import java.util.concurrent.atomic.AtomicLong;
|
||||
import java.util.function.LongSupplier;
|
||||
|
||||
import static org.junit.jupiter.api.Assertions.assertEquals;
|
||||
import static org.junit.jupiter.api.Assertions.assertThrows;
|
||||
import static org.junit.jupiter.api.Assertions.assertTrue;
|
||||
|
||||
/**
|
||||
* fleetd #612 step 2, Unit B1 — replaces {@code FleetdCompletionResolverWiringTest} (deleted in
|
||||
* this same commit), whose four tests read {@code Fleetd.java}'s source text and asserted the
|
||||
* {@code CompletionResolver} construction call still named the right arguments. That proved the
|
||||
* call site's spelling, never that the assembled resolver actually behaves differently when an
|
||||
* argument is dropped.
|
||||
*
|
||||
* <p>These tests drive {@link FleetdAssembly#assembleAndStart} — the real boot composition,
|
||||
* fleetd #612 Unit A — and read {@link FleetdRuntime#completion()}: the exact {@link
|
||||
* CompletionResolver} instance the assembled daemon uses, never a copy built alongside it for the
|
||||
* test's benefit. Two behaviours are pinned, matching the ticket's own two measured mutations:
|
||||
*
|
||||
* <ul>
|
||||
* <li>the 8th constructor argument ({@code Fleetd.worktreeBranchLookup(sessions::roster)}) —
|
||||
* {@link #assembledResolverReportsTheMembersWorktreeAndBranchInAFallbackReport}; and</li>
|
||||
* <li>the 5th/6th arguments ({@code backendErrorPatterns}, {@code backendErrorSink}, both
|
||||
* assigned from {@code Fleetd}'s extracted factories rather than an inline lambda) —
|
||||
* {@link #assembledResolverClassifiesAndCoolsOffOnAConfiguredBackendErrorPattern}.</li>
|
||||
* </ul>
|
||||
*
|
||||
* <p>Both tests bypass {@link dev.ltms.fleet.inject.StatusPoller} and drive {@link
|
||||
* CompletionResolver#onDelivered} / {@link CompletionResolver#resolveBeforePostAction} directly —
|
||||
* the same public, synchronous entry points {@code CompletionResolverTest} uses — with a
|
||||
* hand-built {@link CompletableFuture} waiter, so no real poller loop or herdr status poll is
|
||||
* needed. The pane scrape comes from {@link FakeHerdr#readText}; the elapsed-time floor
|
||||
* ({@code CompletionResolver.MIN_TURN_NANOS}) is controlled via a fake, advanceable {@link
|
||||
* ResourcePorts#nanoClock()} rather than a real sleep.
|
||||
*/
|
||||
class FleetdCompletionResolverAssemblyTest {
|
||||
|
||||
/** Same shape as {@code FleetdAssemblyLifecycleTest}'s fake, plus a nanoClock this test can advance. */
|
||||
private static final class ControllableResourcePorts implements ResourcePorts {
|
||||
|
||||
final FakeHerdr herdr;
|
||||
final AtomicLong nowNanos = new AtomicLong(1_000_000_000L); // arbitrary non-zero start
|
||||
Runnable shutdownHook;
|
||||
|
||||
ControllableResourcePorts(FakeHerdr herdr) {
|
||||
this.herdr = herdr;
|
||||
}
|
||||
|
||||
void advanceSeconds(long seconds) {
|
||||
nowNanos.addAndGet(TimeUnit.SECONDS.toNanos(seconds));
|
||||
}
|
||||
|
||||
@Override
|
||||
public Map<String, String> environment() {
|
||||
return Map.of();
|
||||
}
|
||||
|
||||
@Override
|
||||
public HerdrClient connectHerdr(Path socketPath) {
|
||||
return herdr;
|
||||
}
|
||||
|
||||
@Override
|
||||
public Fleetd.AmqpOpener replyInboxOpener() {
|
||||
// Never invoked: this test's config has no `broker:` block, so Fleetd.selectReplyInbox
|
||||
// returns the in-memory inbox before calling the opener at all.
|
||||
return (uri, prefetch) -> {
|
||||
throw new UnsupportedOperationException("replyInboxOpener must not be called — no broker: block");
|
||||
};
|
||||
}
|
||||
|
||||
@Override
|
||||
public Fleetd.LeadMailboxOpener leadMailboxOpener() {
|
||||
// Never invoked: no `coordinator:` block configured either.
|
||||
return (uri, selfCoordId, prefetch) -> {
|
||||
throw new UnsupportedOperationException("leadMailboxOpener must not be called — no coordinator: block");
|
||||
};
|
||||
}
|
||||
|
||||
@Override
|
||||
public LongSupplier nanoClock() {
|
||||
return nowNanos::get;
|
||||
}
|
||||
|
||||
@Override
|
||||
public LongSupplier wallClockNanos() {
|
||||
return nowNanos::get;
|
||||
}
|
||||
|
||||
@Override
|
||||
public ScheduledExecutorService newScheduler(String purpose) {
|
||||
return Executors.newSingleThreadScheduledExecutor();
|
||||
}
|
||||
|
||||
@Override
|
||||
public void addShutdownHook(Runnable hook) {
|
||||
this.shutdownHook = hook;
|
||||
}
|
||||
|
||||
@Override
|
||||
public void startHttp(Javalin app, String host, int port) {
|
||||
// Deliberately never bind a real port.
|
||||
}
|
||||
|
||||
@Override
|
||||
public Runnable herdrPollWait() {
|
||||
// Never invoked: this test's FakeHerdr answers immediately, so awaitHerdr never polls.
|
||||
return () -> {
|
||||
throw new UnsupportedOperationException("herdrPollWait must not be called — herdr is healthy");
|
||||
};
|
||||
}
|
||||
}
|
||||
|
||||
private static FleetConfig writeConfig(Path dir, String profilesYaml, String extraGuardHost,
|
||||
String worktreeRootYamlLine) throws Exception {
|
||||
Path f = dir.resolve("fleetd.yaml");
|
||||
Files.writeString(f, """
|
||||
bind:
|
||||
host: 127.0.0.1
|
||||
port: 8765
|
||||
idleSleepGuard:
|
||||
enabled: false
|
||||
%s
|
||||
profiles:
|
||||
%s
|
||||
guard:
|
||||
offSubscriptionHosts:
|
||||
- %s
|
||||
""".formatted(worktreeRootYamlLine == null ? "" : worktreeRootYamlLine, profilesYaml, extraGuardHost));
|
||||
return FleetConfig.load(f);
|
||||
}
|
||||
|
||||
private static void gitQuiet(Path cwd, String... args) throws Exception {
|
||||
List<String> cmd = new java.util.ArrayList<>(List.of("git"));
|
||||
cmd.addAll(List.of(args));
|
||||
Process p = new ProcessBuilder(cmd).directory(cwd.toFile()).redirectErrorStream(true).start();
|
||||
String out = new String(p.getInputStream().readAllBytes());
|
||||
assertTrue(p.waitFor(30, TimeUnit.SECONDS), "git timed out: git " + String.join(" ", args));
|
||||
assertEquals(0, p.exitValue(), "git " + String.join(" ", args) + " failed:\n" + out);
|
||||
}
|
||||
|
||||
private static Path initRepo(Path dir) throws Exception {
|
||||
Files.createDirectories(dir);
|
||||
gitQuiet(dir, "init", "-q", "-b", "main");
|
||||
gitQuiet(dir, "config", "user.email", "test@example.invalid");
|
||||
gitQuiet(dir, "config", "user.name", "Test");
|
||||
Files.writeString(dir.resolve("README.md"), "seed\n");
|
||||
gitQuiet(dir, "add", "README.md");
|
||||
gitQuiet(dir, "commit", "-q", "-m", "seed");
|
||||
return dir;
|
||||
}
|
||||
|
||||
/**
|
||||
* fleetd #248's first measured mutation: replacing {@code CompletionResolver}'s 8th constructor
|
||||
* argument with the inert {@code _ -> null} compiles clean and leaves every existing test green
|
||||
* — it silently drops fleetd #241's fallback-report location. This drives the real assembled
|
||||
* resolver through a member echoing its own injected brief back (no {@code fleet_reply}), which
|
||||
* resolves via {@code noReportMessage(target)}, and proves the real member's {@code branch} —
|
||||
* only obtainable via {@code Fleetd.worktreeBranchLookup(sessions::roster)} reading the real,
|
||||
* worktree-provisioned {@link MemberSession} — appears in the reported text.
|
||||
*/
|
||||
@Test
|
||||
void assembledResolverReportsTheMembersWorktreeAndBranchInAFallbackReport(@TempDir Path dir) throws Exception {
|
||||
Path repo = initRepo(dir.resolve("repo"));
|
||||
FleetConfig cfg = writeConfig(dir, """
|
||||
wtprofile:
|
||||
baseUrl: http://wthost.local:8000
|
||||
model: sonnet
|
||||
""", "wthost.local", "worktreeRoot: " + dir.resolve("wts"));
|
||||
ConfigRef config = new ConfigRef(dir.resolve("fleetd.yaml"), cfg);
|
||||
SubscriptionGuard guard = new SubscriptionGuard(cfg.guard().hostSet());
|
||||
ControllableResourcePorts ports = new ControllableResourcePorts(new FakeHerdr());
|
||||
|
||||
FleetdRuntime runtime = FleetdAssembly.assembleAndStart(new AssemblyInputs(cfg, config, guard), ports);
|
||||
try {
|
||||
MemberSession session = runtime.sessions().acquire("wtprofile", repo.toString(), repo.toString(),
|
||||
null, new WorktreeRequest("fleetd-612-b1", null));
|
||||
String target = session.terminalId();
|
||||
String branch = session.branch();
|
||||
assertTrue(branch != null && branch.startsWith("worker/"),
|
||||
"sanity: a worktree-provisioned session must carry a real branch, got: " + branch);
|
||||
|
||||
CompletionResolver completion = runtime.completion();
|
||||
CompletableFuture<Rendezvous.Resolution> waiter = new CompletableFuture<>();
|
||||
String echoedBrief = "z".repeat(450); // >= CompletionResolver.ECHO_MIN_CHARS normalised chars
|
||||
|
||||
ports.herdr.readText("idle, nothing yet");
|
||||
completion.onDelivered(target, new TurnToken(target, waiter, echoedBrief));
|
||||
|
||||
ports.herdr.readText(echoedBrief); // the pane just echoes the injected brief back — no real report
|
||||
ports.advanceSeconds(3); // clear CompletionResolver.MIN_TURN_NANOS (2s) without a real sleep
|
||||
completion.resolveBeforePostAction(target);
|
||||
|
||||
Rendezvous.Resolution resolution = waiter.getNow(null);
|
||||
assertTrue(resolution != null, "the waiter must have resolved synchronously");
|
||||
assertEquals(Rendezvous.Kind.COMPLETION, resolution.kind());
|
||||
assertTrue(resolution.text().contains(CompletionResolver.NO_REPORT_PREFIX),
|
||||
"sanity: must have gone down the noReportMessage sub-path: " + resolution.text());
|
||||
assertTrue(resolution.text().contains("branch=" + branch),
|
||||
"the assembled resolver must report the member's real branch (fleetd #241 via "
|
||||
+ "fleetd #248's worktreeBranchLookup wiring); got: " + resolution.text());
|
||||
} finally {
|
||||
if (ports.shutdownHook != null) ports.shutdownHook.run();
|
||||
}
|
||||
}
|
||||
|
||||
/**
|
||||
* fleetd #248's second measured mutation, and fleetd#201 Unit 5's own gap: replacing {@code
|
||||
* backendErrorPatterns}/{@code backendErrorSink} with {@code BackendErrorPatternLookup.legacy()}
|
||||
* / {@code BackendErrorSink.none()} compiles clean and leaves every existing behavioural test
|
||||
* green.
|
||||
*
|
||||
* <p>Classification proof: this test's profile configures {@code errorPattern: "credential
|
||||
* outage"} — text the built-in {@code (?i)\bAPI Error\s*:} fallback ({@code legacy()}'s only
|
||||
* behaviour) never matches. So a real {@code Fleetd.backendErrorPatternLookup(...)} wiring
|
||||
* classifies the send as {@code FAILED}; {@code legacy()} would fall through to the plain
|
||||
* completion path instead ({@code Kind.COMPLETION}).
|
||||
*
|
||||
* <p>Cool-off proof: two distinct targets on the same profile/credential each classified as a
|
||||
* backend error inside the 60s window must cool the credential off ({@link
|
||||
* dev.ltms.fleet.placement.BackendOutagePolicy}, fleetd#201 Unit 5) — observable two ways: (1)
|
||||
* the real {@code Fleetd.backendErrorSink(...)} marks each session {@code BACKEND_ERROR} (only
|
||||
* the real sink calls {@code sessions.onBackendError}; {@code BackendErrorSink.none()} never
|
||||
* does), and (2) a third explicit-profile spawn attempt is refused with a {@link
|
||||
* PlacementException} naming the cool-off — only reachable because the real sink's {@code
|
||||
* outagePolicy.record(...)} call actually ran.
|
||||
*/
|
||||
@Test
|
||||
void assembledResolverClassifiesAndCoolsOffOnAConfiguredBackendErrorPattern(@TempDir Path dir) throws Exception {
|
||||
FleetConfig cfg = writeConfig(dir, """
|
||||
coolprofile:
|
||||
baseUrl: http://coolhost.local:8000
|
||||
model: sonnet
|
||||
errorPattern: "credential outage"
|
||||
""", "coolhost.local", null);
|
||||
ConfigRef config = new ConfigRef(dir.resolve("fleetd.yaml"), cfg);
|
||||
SubscriptionGuard guard = new SubscriptionGuard(cfg.guard().hostSet());
|
||||
ControllableResourcePorts ports = new ControllableResourcePorts(new FakeHerdr());
|
||||
|
||||
FleetdRuntime runtime = FleetdAssembly.assembleAndStart(new AssemblyInputs(cfg, config, guard), ports);
|
||||
try {
|
||||
MemberSession session1 = runtime.sessions().acquire("coolprofile", null, dir.toString(), null);
|
||||
MemberSession session2 = runtime.sessions().acquire("coolprofile", null, dir.toString(), null);
|
||||
String target1 = session1.terminalId();
|
||||
String target2 = session2.terminalId();
|
||||
assertTrue(!target1.equals(target2), "sanity: the two spawns must be distinct targets");
|
||||
|
||||
CompletionResolver completion = runtime.completion();
|
||||
|
||||
CompletableFuture<Rendezvous.Resolution> waiter1 = new CompletableFuture<>();
|
||||
ports.herdr.readText("idle 1");
|
||||
completion.onDelivered(target1, new TurnToken(target1, waiter1, null));
|
||||
ports.herdr.readText("credential outage: upstream 503");
|
||||
ports.advanceSeconds(3);
|
||||
completion.resolveBeforePostAction(target1);
|
||||
Rendezvous.Resolution resolution1 = waiter1.getNow(null);
|
||||
assertTrue(resolution1 != null, "target1's waiter must have resolved synchronously");
|
||||
assertEquals(Rendezvous.Kind.FAILED, resolution1.kind(),
|
||||
"a configured errorPattern the built-in fallback never matches must classify as "
|
||||
+ "a backend error, not a plain completion; got: " + resolution1);
|
||||
assertTrue(resolution1.text().contains("credential outage: upstream 503"), resolution1.text());
|
||||
|
||||
CompletableFuture<Rendezvous.Resolution> waiter2 = new CompletableFuture<>();
|
||||
ports.herdr.readText("idle 2");
|
||||
completion.onDelivered(target2, new TurnToken(target2, waiter2, null));
|
||||
ports.herdr.readText("credential outage: upstream 503 again");
|
||||
ports.advanceSeconds(3);
|
||||
completion.resolveBeforePostAction(target2);
|
||||
Rendezvous.Resolution resolution2 = waiter2.getNow(null);
|
||||
assertTrue(resolution2 != null, "target2's waiter must have resolved synchronously");
|
||||
assertEquals(Rendezvous.Kind.FAILED, resolution2.kind());
|
||||
|
||||
List<MemberSession> roster = runtime.sessions().roster();
|
||||
assertTrue(roster.stream().anyMatch(s -> target1.equals(s.terminalId())
|
||||
&& s.state() == MemberSession.State.BACKEND_ERROR),
|
||||
"the real backendErrorSink must have transitioned target1 to BACKEND_ERROR: " + roster);
|
||||
assertTrue(roster.stream().anyMatch(s -> target2.equals(s.terminalId())
|
||||
&& s.state() == MemberSession.State.BACKEND_ERROR),
|
||||
"the real backendErrorSink must have transitioned target2 to BACKEND_ERROR: " + roster);
|
||||
|
||||
PlacementException coolOff = assertThrows(PlacementException.class,
|
||||
() -> runtime.sessions().acquire("coolprofile", null, dir.toString(), null),
|
||||
"two distinct targets classified within the 60s window must cool the credential "
|
||||
+ "off (BackendOutagePolicy), refusing a third explicit-profile spawn");
|
||||
assertTrue(coolOff.getMessage().contains("cooling off"), coolOff.getMessage());
|
||||
} finally {
|
||||
if (ports.shutdownHook != null) ports.shutdownHook.run();
|
||||
}
|
||||
}
|
||||
}
|
||||
@@ -1,89 +0,0 @@
|
||||
package dev.ltms.fleet;
|
||||
|
||||
import java.nio.file.Files;
|
||||
import java.nio.file.Path;
|
||||
import org.junit.jupiter.api.DisplayName;
|
||||
import org.junit.jupiter.api.Test;
|
||||
|
||||
import static org.junit.jupiter.api.Assertions.assertFalse;
|
||||
import static org.junit.jupiter.api.Assertions.assertTrue;
|
||||
|
||||
/**
|
||||
* fleetd #248: this is the test that was actually missing. {@code Fleetd.main} builds its {@code
|
||||
* CompletionResolver} from an 8-argument constructor, and the ticket's own measurement proved two
|
||||
* ways to silently unwire it — both compiled with 0 errors and left every existing test green:
|
||||
*
|
||||
* <ul>
|
||||
* <li>replacing the worktree/branch argument (the 8th) with {@code _ -> null} — drops
|
||||
* fleetd#241's fallback-report location entirely;</li>
|
||||
* <li>replacing {@code backendErrorPatterns, backendErrorSink} (5th/6th) with {@code
|
||||
* BackendErrorPatternLookup.legacy(), BackendErrorSink.none()} — drops fleetd#201 Unit 5's
|
||||
* backend-error classification and cool-off entirely.</li>
|
||||
* </ul>
|
||||
*
|
||||
* <p>Neither mutation could be caught by any test that constructs its own {@code
|
||||
* CompletionResolver} (every test before this one did exactly that) or by a test of {@link
|
||||
* Fleetd#worktreeBranchLookup}, {@link Fleetd#backendErrorPatternLookup}, or {@link
|
||||
* Fleetd#backendErrorSink} in isolation (see {@code FleetdWorktreeBranchLookupTest}, {@code
|
||||
* FleetdBackendErrorPatternLookupTest}, {@code FleetdBackendErrorSinkTest}) — those prove the
|
||||
* factories work, never that {@code main} still calls them. This class is a plain source-text
|
||||
* assertion on {@code Fleetd.java} — crude, but honest about what it checks, and it turns red the
|
||||
* instant the wiring is dropped, mirroring the same fallback shape {@link
|
||||
* FleetdFleetAppConstructionTest} already uses for a different constructor argument.
|
||||
*
|
||||
* <p><b>This test checks source text, not runtime behaviour.</b> It never constructs a {@code
|
||||
* CompletionResolver} and never runs {@code main}.
|
||||
*/
|
||||
class FleetdCompletionResolverWiringTest {
|
||||
|
||||
private static String fleetdSource() throws Exception {
|
||||
return Files.readString(Path.of("src/main/java/dev/ltms/fleet/Fleetd.java"));
|
||||
}
|
||||
|
||||
@Test
|
||||
@DisplayName("[SOURCE TEXT] CompletionResolver's construction call still names backendErrorPatterns and backendErrorSink")
|
||||
void backendErrorArgumentsAreStillNamedAtTheCallSite() throws Exception {
|
||||
String source = fleetdSource();
|
||||
assertTrue(source.contains(
|
||||
"exhaustionSink, backendErrorPatterns, backendErrorSink, System::nanoTime,"),
|
||||
"CompletionResolver's construction call must still pass backendErrorPatterns and "
|
||||
+ "backendErrorSink as its 5th/6th arguments. Replacing them with "
|
||||
+ "BackendErrorPatternLookup.legacy()/BackendErrorSink.none() (fleetd #248's measured "
|
||||
+ "mutation) compiles with 0 errors and leaves every behavioural test green — this "
|
||||
+ "source check is what must go red instead.");
|
||||
}
|
||||
|
||||
@Test
|
||||
@DisplayName("[SOURCE TEXT] CompletionResolver's construction call still passes worktreeBranchLookup(sessions::roster)")
|
||||
void worktreeBranchLookupIsStillPassedAtTheCallSite() throws Exception {
|
||||
String source = fleetdSource();
|
||||
assertTrue(source.contains("worktreeBranchLookup(sessions::roster)"),
|
||||
"CompletionResolver's construction call must still pass worktreeBranchLookup(sessions::roster) "
|
||||
+ "as its 8th (last) argument. Replacing it with the inert `_ -> null` (fleetd #248's "
|
||||
+ "other measured mutation) compiles with 0 errors and leaves every behavioural test "
|
||||
+ "green — this source check is what must go red instead.");
|
||||
assertFalse(source.contains("System::nanoTime,\n _ -> null"),
|
||||
"the worktree/branch argument must never regress to the inert `_ -> null` literal");
|
||||
}
|
||||
|
||||
@Test
|
||||
@DisplayName("[SOURCE TEXT] backendErrorPatterns is assigned from the extracted backendErrorPatternLookup(...) factory")
|
||||
void backendErrorPatternsComesFromTheFactory() throws Exception {
|
||||
String source = fleetdSource();
|
||||
assertTrue(source.contains(
|
||||
"BackendErrorPatternLookup backendErrorPatterns = backendErrorPatternLookup(sessions::roster,"),
|
||||
"backendErrorPatterns must be assigned from Fleetd.backendErrorPatternLookup(...), not an "
|
||||
+ "inline lambda that a source check on the CompletionResolver call alone cannot see "
|
||||
+ "through");
|
||||
}
|
||||
|
||||
@Test
|
||||
@DisplayName("[SOURCE TEXT] backendErrorSink is assigned from the extracted backendErrorSink(...) factory")
|
||||
void backendErrorSinkComesFromTheFactory() throws Exception {
|
||||
String source = fleetdSource();
|
||||
assertTrue(source.contains(
|
||||
"BackendErrorSink backendErrorSink = backendErrorSink(sessions, () -> config.get().profiles(),"),
|
||||
"backendErrorSink must be assigned from Fleetd.backendErrorSink(...), not an inline lambda "
|
||||
+ "that a source check on the CompletionResolver call alone cannot see through");
|
||||
}
|
||||
}
|
||||
@@ -24,8 +24,8 @@ import static org.junit.jupiter.api.Assertions.assertTrue;
|
||||
* {@code ConfigRefTest} and {@code FleetdConfigRefCharterToolSurfaceWiringTest} case — because
|
||||
* neither of those tests constructs its {@code ConfigRef} through {@code main}; both build their own
|
||||
* instance directly, wired with the check by hand. That silent regression is exactly the shape
|
||||
* {@link FleetdBackendQuarantineWiringTest}, {@link FleetdLeadSeatWiringTest} and {@link
|
||||
* FleetdCompletionResolverWiringTest} already guard against for their own constructor arguments —
|
||||
* {@link FleetdBackendQuarantineAssemblyTest}, {@link FleetdLeadSeatAssemblyTest} and {@link
|
||||
* FleetdCompletionResolverAssemblyTest} already guard against for their own constructor arguments —
|
||||
* this class is the same class of gap for fleetd #474's {@code extraValidation} argument, following
|
||||
* their approach.
|
||||
*
|
||||
|
||||
@@ -1,31 +0,0 @@
|
||||
package dev.ltms.fleet;
|
||||
|
||||
import java.nio.file.Files;
|
||||
import java.nio.file.Path;
|
||||
import org.junit.jupiter.api.Test;
|
||||
|
||||
import static org.junit.jupiter.api.Assertions.assertFalse;
|
||||
import static org.junit.jupiter.api.Assertions.assertTrue;
|
||||
|
||||
/**
|
||||
* CB-185: {@code ConnectionIdentity} must resolve a caller's pane on EITHER herdr daemon (a
|
||||
* lead's MCP connection resolves against the lead daemon; a member's against the member daemon).
|
||||
* Pinning {@code PaneLocator} to {@code memberHerdr} alone — the bug this guards against — leaves
|
||||
* every lead's own connection unresolvable ({@code callerTerminal == null}) the moment
|
||||
* {@code memberHerdrSocket} names a second daemon, which breaks {@code fleet_reply}/{@code
|
||||
* fleet_ask} and {@code fleet_whoami} for a lead. A unit test on {@link
|
||||
* dev.ltms.fleet.herdr.PaneLocator} alone (see {@code PaneLocatorTest}) proves the class CAN
|
||||
* search two clients, but not that {@code Fleetd.main} actually wires it that way — hence this
|
||||
* source-level assertion, the same technique {@code FleetdHerdrControlConstructionTest} uses.
|
||||
*/
|
||||
class FleetdConnectionIdentityConstructionTest {
|
||||
@Test
|
||||
void connectionIdentitySearchesBothDaemonsNotJustTheMemberOne() throws Exception {
|
||||
String source = Files.readString(Path.of("src/main/java/dev/ltms/fleet/Fleetd.java"));
|
||||
assertFalse(source.contains("new PaneLocator(memberHerdr)"),
|
||||
"PaneLocator must not be pinned to the member daemon alone — a lead's own "
|
||||
+ "connection resolves against the LEAD daemon and would never be found");
|
||||
assertTrue(source.contains("new PaneLocator(herdr, memberHerdr)"),
|
||||
"PaneLocator must search the lead daemon first, then the member daemon");
|
||||
}
|
||||
}
|
||||
@@ -0,0 +1,218 @@
|
||||
package dev.ltms.fleet;
|
||||
|
||||
import dev.ltms.fleet.config.ConfigRef;
|
||||
import dev.ltms.fleet.config.FleetConfig;
|
||||
import dev.ltms.fleet.guard.SubscriptionGuard;
|
||||
import dev.ltms.fleet.herdr.FakeHerdr;
|
||||
import dev.ltms.fleet.herdr.HerdrClient;
|
||||
import dev.ltms.fleet.inject.CompletionResolver;
|
||||
import dev.ltms.fleet.msg.Rendezvous;
|
||||
import dev.ltms.fleet.msg.TurnToken;
|
||||
import dev.ltms.fleet.placement.BackendQuarantine;
|
||||
import dev.ltms.fleet.session.MemberSession;
|
||||
import io.javalin.Javalin;
|
||||
import org.junit.jupiter.api.DisplayName;
|
||||
import org.junit.jupiter.api.Test;
|
||||
import org.junit.jupiter.api.io.TempDir;
|
||||
|
||||
import java.nio.file.Files;
|
||||
import java.nio.file.Path;
|
||||
import java.util.Map;
|
||||
import java.util.OptionalLong;
|
||||
import java.util.concurrent.CompletableFuture;
|
||||
import java.util.concurrent.Executors;
|
||||
import java.util.concurrent.ScheduledExecutorService;
|
||||
import java.util.concurrent.TimeUnit;
|
||||
import java.util.concurrent.atomic.AtomicLong;
|
||||
import java.util.function.LongSupplier;
|
||||
|
||||
import static org.junit.jupiter.api.Assertions.assertEquals;
|
||||
import static org.junit.jupiter.api.Assertions.assertTrue;
|
||||
|
||||
/**
|
||||
* fleetd #612 step 4, ranks 1 and 2 (publish side) — {@link FleetdAssembly} lines
|
||||
* {@code liveExhaustedPatterns}/{@code exhaustedPatterns} (CB-578 stage A, the ticket's own "worst
|
||||
* consequence in the whole sweep": a genuine usage-limit refusal handed back to a waiting caller
|
||||
* AS REAL COMPLETED WORK) and {@code Fleetd.publishExhaustionSink(...)} (CB-578 stage B: the
|
||||
* credential that hit the limit is never quarantined). None of these three lines is driven by an
|
||||
* existing test through the real assembly: {@code FleetdExhaustedPatternLookupWiringTest} and
|
||||
* {@code FleetdLiveExhaustedPatternsWiringTest} (fleetd #589) call {@code Fleetd.liveExhaustedPatterns}
|
||||
* / {@code Fleetd.exhaustedPatternLookup} directly as factories, never through {@link
|
||||
* FleetdAssembly#assembleAndStart} — they prove the FACTORY classifies correctly, never that THIS
|
||||
* call site is the one that actually got wired into the running {@link CompletionResolver}. {@link
|
||||
* FleetdBackendQuarantineAssemblyTest} drives {@code BackendQuarantine.withEscalation(...)}
|
||||
* directly, a different call site from {@code publishExhaustionSink} here.
|
||||
*
|
||||
* <p>This test drives the REAL assembled {@link CompletionResolver} ({@link
|
||||
* FleetdRuntime#completion()}) with a profile carrying a configured {@code exhaustedPattern},
|
||||
* through a pane scrape that matches it, and asserts both halves of the production consequence:
|
||||
* (1) the resolution is {@link Rendezvous.Kind#BACKEND_EXHAUSTED}, never a plain completion handed
|
||||
* back as real work, and (2) the profile's credential is actually quarantined afterward, through
|
||||
* the REAL {@link BackendQuarantine} the same assembly built ({@link
|
||||
* FleetdRuntime#mcp()}{@code .quarantineSource().quarantine()}) — never a copy.
|
||||
*
|
||||
* <p>Same {@code ControllableResourcePorts} shape as {@code FleetdCompletionResolverAssemblyTest}:
|
||||
* a fake, advanceable {@code nanoClock} so {@code CompletionResolver.MIN_TURN_NANOS} clears without
|
||||
* a real sleep, and {@link FakeHerdr#readText} to drive the pane scrape.
|
||||
*/
|
||||
class FleetdExhaustedPatternAssemblyTest {
|
||||
|
||||
private static final class ControllableResourcePorts implements ResourcePorts {
|
||||
|
||||
final FakeHerdr herdr;
|
||||
final AtomicLong nowNanos = new AtomicLong(1_000_000_000L); // arbitrary non-zero start
|
||||
Runnable shutdownHook;
|
||||
|
||||
ControllableResourcePorts(FakeHerdr herdr) {
|
||||
this.herdr = herdr;
|
||||
}
|
||||
|
||||
void advanceSeconds(long seconds) {
|
||||
nowNanos.addAndGet(TimeUnit.SECONDS.toNanos(seconds));
|
||||
}
|
||||
|
||||
@Override
|
||||
public Map<String, String> environment() {
|
||||
return Map.of();
|
||||
}
|
||||
|
||||
@Override
|
||||
public HerdrClient connectHerdr(Path socketPath) {
|
||||
return herdr;
|
||||
}
|
||||
|
||||
@Override
|
||||
public Fleetd.AmqpOpener replyInboxOpener() {
|
||||
return (uri, prefetch) -> {
|
||||
throw new UnsupportedOperationException("replyInboxOpener must not be called — no broker: block");
|
||||
};
|
||||
}
|
||||
|
||||
@Override
|
||||
public Fleetd.LeadMailboxOpener leadMailboxOpener() {
|
||||
return (uri, selfCoordId, prefetch) -> {
|
||||
throw new UnsupportedOperationException("leadMailboxOpener must not be called — no coordinator: block");
|
||||
};
|
||||
}
|
||||
|
||||
@Override
|
||||
public Runnable herdrPollWait() {
|
||||
// Never invoked: this test's FakeHerdr answers immediately, so awaitHerdr never polls.
|
||||
return () -> {
|
||||
throw new UnsupportedOperationException("herdrPollWait must not be called — herdr is healthy");
|
||||
};
|
||||
}
|
||||
@Override
|
||||
public LongSupplier nanoClock() {
|
||||
return nowNanos::get;
|
||||
}
|
||||
|
||||
@Override
|
||||
public LongSupplier wallClockNanos() {
|
||||
return nowNanos::get;
|
||||
}
|
||||
|
||||
@Override
|
||||
public ScheduledExecutorService newScheduler(String purpose) {
|
||||
return Executors.newSingleThreadScheduledExecutor();
|
||||
}
|
||||
|
||||
@Override
|
||||
public void addShutdownHook(Runnable hook) {
|
||||
this.shutdownHook = hook;
|
||||
}
|
||||
|
||||
@Override
|
||||
public void startHttp(Javalin app, String host, int port) {
|
||||
// Deliberately never bind a real port.
|
||||
}
|
||||
}
|
||||
|
||||
private static FleetConfig writeConfig(Path dir, int cooldownSeconds) throws Exception {
|
||||
Path f = dir.resolve("fleetd.yaml");
|
||||
Files.writeString(f, """
|
||||
bind:
|
||||
host: 127.0.0.1
|
||||
port: 8765
|
||||
idleSleepGuard:
|
||||
enabled: false
|
||||
quarantineCooldownSeconds: %d
|
||||
profiles:
|
||||
exhaustprofile:
|
||||
baseUrl: http://exhausthost.local:8000
|
||||
model: sonnet
|
||||
exhaustedPattern: "usage limit reached"
|
||||
guard:
|
||||
offSubscriptionHosts:
|
||||
- exhausthost.local
|
||||
""".formatted(cooldownSeconds));
|
||||
return FleetConfig.load(f);
|
||||
}
|
||||
|
||||
/**
|
||||
* fleetd #589's own description of this gap ({@code Fleetd#exhaustedPatternLookup}'s javadoc):
|
||||
* "the worst consequence in the whole #589 sweep" — a genuine usage-limit refusal stops being
|
||||
* classified as {@code BACKEND_EXHAUSTED} and is handed back to a waiting {@code fleet_send} as
|
||||
* if it were real completed work. Pins {@code FleetdAssembly}'s {@code liveExhaustedPatterns}
|
||||
* AND {@code exhaustedPatterns} lines (rank 1) together with {@code publishExhaustionSink}
|
||||
* (rank 2, the non-OpenCode half) in one flow: classify, then quarantine.
|
||||
*/
|
||||
@Test
|
||||
@DisplayName("[BEHAVIOURAL] a scrape matching the profile's exhaustedPattern resolves "
|
||||
+ "BACKEND_EXHAUSTED (never a plain completion) and quarantines the credential")
|
||||
void assembledResolverClassifiesExhaustionAndQuarantinesTheCredential(@TempDir Path dir) throws Exception {
|
||||
int cooldownSeconds = 120;
|
||||
FleetConfig cfg = writeConfig(dir, cooldownSeconds);
|
||||
ConfigRef config = new ConfigRef(dir.resolve("fleetd.yaml"), cfg);
|
||||
SubscriptionGuard guard = new SubscriptionGuard(cfg.guard().hostSet());
|
||||
ControllableResourcePorts ports = new ControllableResourcePorts(new FakeHerdr());
|
||||
|
||||
FleetdRuntime runtime = FleetdAssembly.assembleAndStart(new AssemblyInputs(cfg, config, guard), ports);
|
||||
try {
|
||||
MemberSession session = runtime.sessions().acquire("exhaustprofile", null, dir.toString(), null);
|
||||
String target = session.terminalId();
|
||||
|
||||
CompletionResolver completion = runtime.completion();
|
||||
CompletableFuture<Rendezvous.Resolution> waiter = new CompletableFuture<>();
|
||||
|
||||
ports.herdr.readText("idle, nothing yet");
|
||||
completion.onDelivered(target, new TurnToken(target, waiter, null));
|
||||
// The matched text must START the pane line (CompletionResolver.startsWithExhaustion) —
|
||||
// no preceding sentence — for the quarantine side-effect to fire, same as production.
|
||||
ports.herdr.readText("usage limit reached: try again in a few hours");
|
||||
ports.advanceSeconds(3); // clear CompletionResolver.MIN_TURN_NANOS (2s), no real sleep
|
||||
completion.resolveBeforePostAction(target);
|
||||
|
||||
// CONTROL: the waiter must have resolved synchronously at all — if the assembled
|
||||
// CompletionResolver were never actually driven (e.g. a wiring break upstream silently
|
||||
// left the resolver unreachable), this fails loudly before the real assertions below
|
||||
// ever run, rather than passing on an untouched waiter.
|
||||
Rendezvous.Resolution resolution = waiter.getNow(null);
|
||||
assertTrue(resolution != null, "CONTROL: the waiter must have resolved synchronously — "
|
||||
+ "if this is null, the assembled resolver was never actually exercised");
|
||||
|
||||
assertEquals(Rendezvous.Kind.BACKEND_EXHAUSTED, resolution.kind(),
|
||||
"a scrape matching the profile's configured exhaustedPattern must classify as "
|
||||
+ "BACKEND_EXHAUSTED, not a plain completion handed back as real work — "
|
||||
+ "replacing FleetdAssembly's liveExhaustedPatterns/exhaustedPatterns "
|
||||
+ "lines with their inert forms (Map.of() / target -> null) must fail "
|
||||
+ "this assertion; got: " + resolution);
|
||||
assertTrue(resolution.text().contains("usage limit reached"), resolution.text());
|
||||
|
||||
BackendQuarantine quarantine = runtime.mcp().quarantineSource().quarantine();
|
||||
assertTrue(quarantine.isQuarantined("exhaustprofile"),
|
||||
"the real publishExhaustionSink-built sink must have quarantined the profile's "
|
||||
+ "credential (effectiveCredentialId() == the profile name here, no "
|
||||
+ "credentialId configured) — replacing FleetdAssembly's "
|
||||
+ "publishExhaustionSink call site with a hardcoded ExhaustionSink.none() "
|
||||
+ "must fail this assertion, since nothing would ever call "
|
||||
+ "quarantine.quarantine(...)");
|
||||
OptionalLong remaining = quarantine.remainingSeconds("exhaustprofile");
|
||||
assertTrue(remaining.isPresent() && remaining.getAsLong() > 0
|
||||
&& remaining.getAsLong() <= cooldownSeconds,
|
||||
"a fresh quarantine must block for at most the configured base cooldown: " + remaining);
|
||||
} finally {
|
||||
if (ports.shutdownHook != null) ports.shutdownHook.run();
|
||||
}
|
||||
}
|
||||
}
|
||||
@@ -1,29 +0,0 @@
|
||||
package dev.ltms.fleet;
|
||||
|
||||
import java.nio.file.Files;
|
||||
import java.nio.file.Path;
|
||||
import org.junit.jupiter.api.Test;
|
||||
|
||||
import static org.junit.jupiter.api.Assertions.assertFalse;
|
||||
import static org.junit.jupiter.api.Assertions.assertTrue;
|
||||
|
||||
/**
|
||||
* CB-185: {@code FleetApp} must be constructed with BOTH herdr clients (the lead's and the
|
||||
* member's), never the raw lead-only {@code herdr}. Passing only {@code herdr} — the bug this
|
||||
* guards against — makes {@code GET /healthz} green while the member daemon is down (so every
|
||||
* spawn fails invisibly) and silently drops every member workspace from {@code GET /sessions}.
|
||||
* A behavioural test on {@code FleetApp} alone (see {@code FleetAppTwoDaemonTest}) proves the
|
||||
* class merges/gates correctly when given two clients, but not that {@code Fleetd.main} actually
|
||||
* passes it two — hence this source-level assertion, mirroring
|
||||
* {@code FleetdHerdrControlConstructionTest}.
|
||||
*/
|
||||
class FleetdFleetAppConstructionTest {
|
||||
@Test
|
||||
void fleetAppIsConstructedWithBothHerdrDaemons() throws Exception {
|
||||
String source = Files.readString(Path.of("src/main/java/dev/ltms/fleet/Fleetd.java"));
|
||||
assertFalse(source.contains("new FleetApp(herdr, workers,"),
|
||||
"FleetApp must not be constructed with the lead-only herdr client");
|
||||
assertTrue(source.contains("new FleetApp(herdr, memberHerdr, workers,"),
|
||||
"FleetApp must be constructed with both the lead and the member herdr client");
|
||||
}
|
||||
}
|
||||
@@ -2,16 +2,67 @@ package dev.ltms.fleet;
|
||||
|
||||
import java.nio.file.Files;
|
||||
import java.nio.file.Path;
|
||||
import org.junit.jupiter.api.DisplayName;
|
||||
import org.junit.jupiter.api.Test;
|
||||
|
||||
import static org.junit.jupiter.api.Assertions.assertFalse;
|
||||
import static org.junit.jupiter.api.Assertions.assertTrue;
|
||||
|
||||
/**
|
||||
* {@code AgentControl} caches {@code paneByTerminal}, so {@code HerdrRouter} must be its only
|
||||
* production factory — a second instance means a second cache; the same reasoning applies to
|
||||
* {@code WorkspaceControl}. {@code HerdrRouter}'s constructor is the one place both are built.
|
||||
*
|
||||
* <p><b>This test checks source text, not runtime behaviour.</b> It never constructs a {@code
|
||||
* HerdrRouter} and never runs {@code FleetdAssembly.assembleAndStart} — a green result proves only
|
||||
* that neither watched file's text contains {@code new AgentControl(} or {@code new
|
||||
* WorkspaceControl(}. It does not prove the instances {@code HerdrRouter} does build are the ones
|
||||
* actually wired through the rest of the daemon, and it does not cover a bypass written into a
|
||||
* production file other than the two this test reads.
|
||||
*/
|
||||
class FleetdHerdrControlConstructionTest {
|
||||
|
||||
private static String source(String relativePath) throws Exception {
|
||||
return Files.readString(Path.of(relativePath));
|
||||
}
|
||||
|
||||
@Test
|
||||
@DisplayName("[SOURCE TEXT] Fleetd.java never constructs AgentControl or WorkspaceControl directly")
|
||||
void fleetdDelegatesStatefulControlsToTheRouter() throws Exception {
|
||||
// AgentControl caches paneByTerminal, so the router must be its only production factory.
|
||||
String source = Files.readString(Path.of("src/main/java/dev/ltms/fleet/Fleetd.java"));
|
||||
assertFalse(source.contains("new AgentControl("));
|
||||
assertFalse(source.contains("new WorkspaceControl("));
|
||||
String source = source("src/main/java/dev/ltms/fleet/Fleetd.java");
|
||||
|
||||
// A broken read (wrong working directory, wrong path, a file that came back empty) would
|
||||
// make the assertFalse checks below pass vacuously — a "clean" negative check that actually
|
||||
// checked nothing. Guard against that first, with an anchor that has nothing to do with
|
||||
// this mutation, so a bad read fails loudly here instead of silently proving nothing below.
|
||||
assertTrue(source.contains("public final class Fleetd"),
|
||||
"the read of Fleetd.java did not come back containing its own class declaration — "
|
||||
+ "the assertFalse checks below would pass vacuously on a broken read; fix the "
|
||||
+ "read before trusting this test.");
|
||||
|
||||
assertFalse(source.contains("new AgentControl("),
|
||||
"Fleetd.java must not construct AgentControl directly — HerdrRouter is its only "
|
||||
+ "production factory");
|
||||
assertFalse(source.contains("new WorkspaceControl("),
|
||||
"Fleetd.java must not construct WorkspaceControl directly — HerdrRouter is its only "
|
||||
+ "production factory");
|
||||
}
|
||||
|
||||
@Test
|
||||
@DisplayName("[SOURCE TEXT] FleetdAssembly.java never constructs AgentControl or WorkspaceControl directly")
|
||||
void fleetdAssemblyDelegatesStatefulControlsToTheRouter() throws Exception {
|
||||
String source = source("src/main/java/dev/ltms/fleet/FleetdAssembly.java");
|
||||
|
||||
assertTrue(source.contains("final class FleetdAssembly"),
|
||||
"the read of FleetdAssembly.java did not come back containing its own class "
|
||||
+ "declaration — the assertFalse checks below would pass vacuously on a broken "
|
||||
+ "read; fix the read before trusting this test.");
|
||||
|
||||
assertFalse(source.contains("new AgentControl("),
|
||||
"FleetdAssembly.java must not construct AgentControl directly — HerdrRouter is its "
|
||||
+ "only production factory");
|
||||
assertFalse(source.contains("new WorkspaceControl("),
|
||||
"FleetdAssembly.java must not construct WorkspaceControl directly — HerdrRouter is "
|
||||
+ "its only production factory");
|
||||
}
|
||||
}
|
||||
|
||||
@@ -0,0 +1,213 @@
|
||||
package dev.ltms.fleet;
|
||||
|
||||
import dev.ltms.fleet.config.ConfigRef;
|
||||
import dev.ltms.fleet.config.FleetConfig;
|
||||
import dev.ltms.fleet.guard.SubscriptionGuard;
|
||||
import dev.ltms.fleet.herdr.FakeHerdr;
|
||||
import dev.ltms.fleet.herdr.HerdrClient;
|
||||
import dev.ltms.fleet.mcp.FleetMcp;
|
||||
import dev.ltms.fleet.msg.ReplyInbox;
|
||||
import io.javalin.Javalin;
|
||||
import org.junit.jupiter.api.DisplayName;
|
||||
import org.junit.jupiter.api.Test;
|
||||
import org.junit.jupiter.api.io.TempDir;
|
||||
|
||||
import java.lang.reflect.Field;
|
||||
import java.nio.file.Files;
|
||||
import java.nio.file.Path;
|
||||
import java.util.List;
|
||||
import java.util.Map;
|
||||
import java.util.concurrent.Executors;
|
||||
import java.util.concurrent.ScheduledExecutorService;
|
||||
import java.util.function.LongSupplier;
|
||||
|
||||
import static org.junit.jupiter.api.Assertions.assertEquals;
|
||||
import static org.junit.jupiter.api.Assertions.assertNull;
|
||||
|
||||
/**
|
||||
* fleetd #612 Shape A, unit r5 — {@code FleetdAssembly.java:488} wires {@link
|
||||
* FleetMcp.LeadConfigDirSource} with {@code Fleetd.leadConfigDirSource(() -> config.get().profiles(),
|
||||
* leaders)}. {@link FleetdLeadConfigDirSourceWiringTest} already pins that the FACTORY itself
|
||||
* delegates to the real {@link Fleetd#leadConfigDirLookup} — but, by its own javadoc, it "does not
|
||||
* and structurally cannot cover" whether the real call site in {@code FleetdAssembly} still calls
|
||||
* that factory at all. Measured there: swapping that one-line call for a bare {@code
|
||||
* FleetMcp.LeadConfigDirSource.none()} compiles with 0 errors and leaves the full suite green.
|
||||
*
|
||||
* <p>This is the literal fleetd #602/#606 defect, one call site away from its own fix: {@code main}
|
||||
* (now {@code FleetdAssembly}) used to build {@code LeadConfigDirSource.none()} inline, the whole
|
||||
* suite passed, and the live daemon reported {@code "state":"unknown"} for every lead's context,
|
||||
* forever, with no test noticing. The fix extracted the factory; this test is the one that proves
|
||||
* {@code FleetdAssembly}'s own call site still reaches it.
|
||||
*
|
||||
* <p>This test drives the REAL {@link FleetMcp} the real {@link FleetdAssembly#assembleAndStart}
|
||||
* builds, reached through {@link FleetdRuntime#mcp()}, and reads the {@code leadConfigDirs} field it
|
||||
* was constructed with via reflection — {@code FleetMcp} exposes no public accessor for it (unlike
|
||||
* {@code quarantineSource()}/{@code leadSeatSource()}), so there is no non-reflective route to the
|
||||
* live instance. The assertion resolves a REAL lead name against a REAL configured {@code
|
||||
* configDir:}: {@link FleetMcp.LeadConfigDirSource#none()} (the historical defect, and the
|
||||
* mis-wire this test's mutation cycles reintroduce) always returns {@code null} regardless of the
|
||||
* input, so a non-null, config-matching answer is a property {@code none()} can never produce by
|
||||
* accident.
|
||||
*/
|
||||
class FleetdLeadConfigDirSourceAssemblyTest {
|
||||
|
||||
private static final String LEAD_NAME = "opus";
|
||||
private static final String LEAD_TAB = "lead: opus";
|
||||
private static final String LEAD_PROFILE = "sonnet";
|
||||
|
||||
private static final class RecordingResourcePorts implements ResourcePorts {
|
||||
|
||||
final FakeHerdr herdr = new FakeHerdr();
|
||||
final SentinelReplyInbox replyInbox = new SentinelReplyInbox();
|
||||
|
||||
@Override
|
||||
public Map<String, String> environment() {
|
||||
return Map.of();
|
||||
}
|
||||
|
||||
@Override
|
||||
public HerdrClient connectHerdr(Path socketPath) {
|
||||
return herdr;
|
||||
}
|
||||
|
||||
@Override
|
||||
public Fleetd.AmqpOpener replyInboxOpener() {
|
||||
return (uri, prefetch) -> replyInbox;
|
||||
}
|
||||
|
||||
@Override
|
||||
public Fleetd.LeadMailboxOpener leadMailboxOpener() {
|
||||
return (uri, selfCoordId, prefetch) -> {
|
||||
throw new UnsupportedOperationException(
|
||||
"leadMailboxOpener must not be called — no coordinator: block is configured");
|
||||
};
|
||||
}
|
||||
|
||||
@Override
|
||||
public LongSupplier nanoClock() {
|
||||
return System::nanoTime;
|
||||
}
|
||||
|
||||
@Override
|
||||
public LongSupplier wallClockNanos() {
|
||||
return System::nanoTime;
|
||||
}
|
||||
|
||||
@Override
|
||||
public ScheduledExecutorService newScheduler(String purpose) {
|
||||
return Executors.newSingleThreadScheduledExecutor();
|
||||
}
|
||||
|
||||
@Override
|
||||
public void addShutdownHook(Runnable hook) {
|
||||
}
|
||||
|
||||
@Override
|
||||
public void startHttp(Javalin app, String host, int port) {
|
||||
}
|
||||
|
||||
@Override
|
||||
public Runnable herdrPollWait() {
|
||||
// Never invoked: this test's FakeHerdr answers immediately, so awaitHerdr never polls.
|
||||
return () -> {
|
||||
throw new UnsupportedOperationException("herdrPollWait must not be called — herdr is healthy");
|
||||
};
|
||||
}
|
||||
}
|
||||
|
||||
private static final class SentinelReplyInbox implements ReplyInbox, AutoCloseable {
|
||||
@Override
|
||||
public void own(String target) {
|
||||
}
|
||||
|
||||
@Override
|
||||
public void release(String target) {
|
||||
}
|
||||
|
||||
@Override
|
||||
public void publish(String target, String msgId, String content) {
|
||||
}
|
||||
|
||||
@Override
|
||||
public List<InboxMessage> peek(String target) {
|
||||
return List.of();
|
||||
}
|
||||
|
||||
@Override
|
||||
public boolean ack(String target, String msgId) {
|
||||
return false;
|
||||
}
|
||||
|
||||
@Override
|
||||
public void close() {
|
||||
}
|
||||
}
|
||||
|
||||
private static FleetConfig writeConfig(Path dir, String configDir) throws Exception {
|
||||
Path f = dir.resolve("fleetd.yaml");
|
||||
Files.writeString(f, """
|
||||
bind:
|
||||
host: 127.0.0.1
|
||||
port: 8765
|
||||
idleSleepGuard:
|
||||
enabled: false
|
||||
broker:
|
||||
uri: "amqp://fake-test-broker/vh"
|
||||
fleet:
|
||||
leaders:
|
||||
%s:
|
||||
tab: "%s"
|
||||
profile: %s
|
||||
profiles:
|
||||
%s:
|
||||
subscription: true
|
||||
argv: ["ccs", "sonnet"]
|
||||
configDir: "%s"
|
||||
""".formatted(LEAD_NAME, LEAD_TAB, LEAD_PROFILE, LEAD_PROFILE, configDir));
|
||||
return FleetConfig.load(f);
|
||||
}
|
||||
|
||||
@SuppressWarnings("unchecked")
|
||||
private static FleetMcp.LeadConfigDirSource leadConfigDirSourceOf(FleetMcp mcp) throws Exception {
|
||||
Field field = FleetMcp.class.getDeclaredField("leadConfigDirs");
|
||||
field.setAccessible(true);
|
||||
return (FleetMcp.LeadConfigDirSource) field.get(mcp);
|
||||
}
|
||||
|
||||
@Test
|
||||
@DisplayName("[BEHAVIOURAL] the real assembled LeadConfigDirSource resolves a lead's REAL "
|
||||
+ "configured configDir, not the none() stand-in's hardcoded null")
|
||||
void assembledLeadConfigDirSourceResolvesTheRealConfiguredConfigDir(@TempDir Path dir) throws Exception {
|
||||
String configuredConfigDir = "/mnt/fake-lead-configdir";
|
||||
FleetConfig cfg = writeConfig(dir, configuredConfigDir);
|
||||
ConfigRef config = new ConfigRef(dir.resolve("fleetd.yaml"), cfg);
|
||||
SubscriptionGuard guard = new SubscriptionGuard(cfg.guard().hostSet());
|
||||
RecordingResourcePorts ports = new RecordingResourcePorts();
|
||||
// Label FakeHerdr's own default pane's tab (term_a / w2:p7 / w2:t7, already carrying a live
|
||||
// agent) to match fleet.leaders.opus.tab exactly, so LeadLauncher.ensureLeads() sees the
|
||||
// lead as already live and does not try to auto-launch a second one.
|
||||
ports.herdr.withTab("w2", "w2:t7", LEAD_TAB);
|
||||
|
||||
FleetdRuntime runtime = FleetdAssembly.assembleAndStart(new AssemblyInputs(cfg, config, guard), ports);
|
||||
try {
|
||||
FleetMcp.LeadConfigDirSource source = leadConfigDirSourceOf(runtime.mcp());
|
||||
|
||||
assertEquals(configuredConfigDir, source.configDirFor().apply(LEAD_NAME),
|
||||
"fleet.leaders." + LEAD_NAME + ".profile (" + LEAD_PROFILE + ") configures "
|
||||
+ "configDir: " + configuredConfigDir + " — the real assembled source must "
|
||||
+ "resolve it. FleetMcp.LeadConfigDirSource.none() (the inert stand-in "
|
||||
+ "this test's mutation cycles swap the call site for, and the historical "
|
||||
+ "fleetd #602/#606 defect) always reports null here, whatever the input");
|
||||
|
||||
// A lead name the config does not recognise still resolves to null, not a crash — the
|
||||
// same source, applied to an input that must stay at the inert answer even on the real,
|
||||
// non-inert instance.
|
||||
assertNull(source.configDirFor().apply("no-such-lead"));
|
||||
} finally {
|
||||
// Surefire runs the whole suite in one JVM fork (fleetd/pom.xml sets no forkCount /
|
||||
// reuseForks), so the scheduler/loops this assembly starts must be torn down here, on the
|
||||
// failure path too — hence try/finally rather than a bare statement at the end.
|
||||
runtime.close();
|
||||
}
|
||||
}
|
||||
}
|
||||
@@ -0,0 +1,63 @@
|
||||
package dev.ltms.fleet;
|
||||
|
||||
import dev.ltms.fleet.config.FleetConfig;
|
||||
import dev.ltms.fleet.mcp.FleetMcp;
|
||||
import org.junit.jupiter.api.DisplayName;
|
||||
import org.junit.jupiter.api.Test;
|
||||
|
||||
import java.util.Map;
|
||||
|
||||
import static org.junit.jupiter.api.Assertions.assertEquals;
|
||||
import static org.junit.jupiter.api.Assertions.assertNull;
|
||||
|
||||
/**
|
||||
* Pins {@link Fleetd#leadConfigDirSource}'s own wiring of the window lookup into the returned
|
||||
* {@link FleetMcp.LeadConfigDirSource}, not only the detached {@link Fleetd#leadContextWindowLookup}
|
||||
* factory it delegates to. Calls the producer directly, with real {@link FleetConfig.Profile}/
|
||||
* {@link FleetConfig.Leader} fixtures, and asserts on {@code windowFor()} — the companion of
|
||||
* {@link FleetdLeadConfigDirSourceWiringTest}, which pins the same factory's {@code configDirFor()}.
|
||||
*/
|
||||
class FleetdLeadConfigDirSourceWindowWiringTest {
|
||||
|
||||
private static FleetConfig.Profile profileWithWindow(String name, Integer autoCompactWindow) {
|
||||
return new FleetConfig.Profile(name, null, "claude-sonnet-5", null, null, null,
|
||||
"tab", "fleet", "w #{n}", null, null, null, null, null, null, null,
|
||||
null, null, true, null, null, null, null, null, autoCompactWindow, null);
|
||||
}
|
||||
|
||||
private static FleetConfig.Leader leadOnProfile(String profile) {
|
||||
return new FleetConfig.Leader(profile, "lead: primary", 1, "lead:", 10, "claude", "claude-sonnet-5");
|
||||
}
|
||||
|
||||
@Test
|
||||
@DisplayName("the returned source resolves the lead's REAL configured effective window, not a hardcoded null")
|
||||
void resolvesTheRealConfiguredWindow() {
|
||||
Map<String, FleetConfig.Profile> profiles = Map.of("opus", profileWithWindow("opus", 250_000));
|
||||
Map<String, FleetConfig.Leader> leaders = Map.of("primary", leadOnProfile("opus"));
|
||||
|
||||
FleetMcp.LeadConfigDirSource source = Fleetd.leadConfigDirSource(() -> profiles, leaders);
|
||||
|
||||
assertEquals(250_000L, source.windowFor().apply("primary"),
|
||||
"windowFor must delegate to the real leadContextWindowLookup, not a stub that always "
|
||||
+ "returns null");
|
||||
}
|
||||
|
||||
@Test
|
||||
@DisplayName("a lead on a profile with no window configured still resolves to null, not a crash")
|
||||
void leadWithNoWindowConfiguredResolvesToNull() {
|
||||
Map<String, FleetConfig.Profile> profiles = Map.of("opus", profileWithWindow("opus", null));
|
||||
Map<String, FleetConfig.Leader> leaders = Map.of("primary", leadOnProfile("opus"));
|
||||
|
||||
FleetMcp.LeadConfigDirSource source = Fleetd.leadConfigDirSource(() -> profiles, leaders);
|
||||
|
||||
assertNull(source.windowFor().apply("primary"));
|
||||
}
|
||||
|
||||
@Test
|
||||
@DisplayName("an unrecognised lead name resolves to null, not a thrown exception")
|
||||
void unrecognisedLeadNameResolvesToNull() {
|
||||
FleetMcp.LeadConfigDirSource source = Fleetd.leadConfigDirSource(Map::of, Map.of());
|
||||
|
||||
assertNull(source.windowFor().apply("ghost-lead"));
|
||||
}
|
||||
}
|
||||
@@ -99,7 +99,7 @@ class FleetdLeadContextLookupTest {
|
||||
@DisplayName("an unrecognised terminal resolves to UNKNOWN, not a thrown exception")
|
||||
void unrecognisedTerminalResolvesToUnknown() {
|
||||
Function<String, LeadContextGauge.Reading> lookup = Fleetd.leadContextLookup(
|
||||
new LeadContextGauge(), throwingAgentControl(), Map::of, name -> null);
|
||||
new LeadContextGauge(), throwingAgentControl(), Map::of, name -> null, name -> null);
|
||||
|
||||
LeadContextGauge.Reading reading = assertDoesNotThrow(() -> lookup.apply("ghost-terminal"));
|
||||
|
||||
@@ -112,7 +112,7 @@ class FleetdLeadContextLookupTest {
|
||||
void agentsGetThrowingDegradesToUnknown() {
|
||||
Map<String, String> liveLeadTerminals = Map.of(LEAD_TERMINAL, LEAD_NAME);
|
||||
Function<String, LeadContextGauge.Reading> lookup = Fleetd.leadContextLookup(
|
||||
new LeadContextGauge(), throwingAgentControl(), () -> liveLeadTerminals, name -> null);
|
||||
new LeadContextGauge(), throwingAgentControl(), () -> liveLeadTerminals, name -> null, name -> null);
|
||||
|
||||
LeadContextGauge.Reading reading = assertDoesNotThrow(() -> lookup.apply(LEAD_TERMINAL));
|
||||
|
||||
@@ -128,7 +128,7 @@ class FleetdLeadContextLookupTest {
|
||||
AgentControl agents = agentControlStub(SESSION_ID, "claude", "idle");
|
||||
Function<String, LeadContextGauge.Reading> lookup = Fleetd.leadContextLookup(
|
||||
new LeadContextGauge(), agents, () -> liveLeadTerminals,
|
||||
name -> LEAD_NAME.equals(name) ? tmp.toString() : null);
|
||||
name -> LEAD_NAME.equals(name) ? tmp.toString() : null, name -> null);
|
||||
|
||||
LeadContextGauge.Reading reading = lookup.apply(LEAD_TERMINAL);
|
||||
|
||||
@@ -146,7 +146,7 @@ class FleetdLeadContextLookupTest {
|
||||
AgentControl agents = agentControlStub(SESSION_ID, "opencode", "idle");
|
||||
Function<String, LeadContextGauge.Reading> lookup = Fleetd.leadContextLookup(
|
||||
new LeadContextGauge(), agents, () -> liveLeadTerminals,
|
||||
name -> LEAD_NAME.equals(name) ? tmp.toString() : null);
|
||||
name -> LEAD_NAME.equals(name) ? tmp.toString() : null, name -> null);
|
||||
|
||||
LeadContextGauge.Reading reading = lookup.apply(LEAD_TERMINAL);
|
||||
|
||||
|
||||
@@ -0,0 +1,223 @@
|
||||
package dev.ltms.fleet;
|
||||
|
||||
import dev.ltms.fleet.config.ConfigRef;
|
||||
import dev.ltms.fleet.config.FleetConfig;
|
||||
import dev.ltms.fleet.guard.SubscriptionGuard;
|
||||
import dev.ltms.fleet.herdr.FakeHerdr;
|
||||
import dev.ltms.fleet.herdr.HerdrClient;
|
||||
import dev.ltms.fleet.lead.LeadContextGauge;
|
||||
import dev.ltms.fleet.msg.LeadHeartbeatLoop;
|
||||
import dev.ltms.fleet.msg.ReplyInbox;
|
||||
import io.javalin.Javalin;
|
||||
import org.junit.jupiter.api.DisplayName;
|
||||
import org.junit.jupiter.api.Test;
|
||||
import org.junit.jupiter.api.io.TempDir;
|
||||
|
||||
import java.io.IOException;
|
||||
import java.lang.reflect.Field;
|
||||
import java.nio.charset.StandardCharsets;
|
||||
import java.nio.file.Files;
|
||||
import java.nio.file.Path;
|
||||
import java.util.List;
|
||||
import java.util.Map;
|
||||
import java.util.concurrent.Executors;
|
||||
import java.util.concurrent.ScheduledExecutorService;
|
||||
import java.util.function.LongSupplier;
|
||||
|
||||
import static org.junit.jupiter.api.Assertions.assertEquals;
|
||||
|
||||
/**
|
||||
* fleetd #659: {@code FleetdAssembly.java} wires {@code Fleetd.leadContextSource}'s window-lookup
|
||||
* argument with {@code Fleetd.leadContextWindowLookup(() -> config.get().profiles(), leaders)} —
|
||||
* but nothing called the real assembled {@link LeadHeartbeatLoop} far enough to prove that
|
||||
* argument is the one the live heartbeat reads through. Measured: swapping that one call-site
|
||||
* argument for {@code _ -> null} compiles with 0 errors and leaves the full suite green.
|
||||
*
|
||||
* <p>This test drives the REAL {@link LeadHeartbeatLoop} the real {@link
|
||||
* FleetdAssembly#assembleAndStart} builds, reached through {@link FleetdRuntime#heartbeat()}, and
|
||||
* reads its private {@code contextSource} field via reflection — the loop exposes no public
|
||||
* accessor for it, the same reason {@link FleetdLeadConfigDirSourceAssemblyTest} reflects on
|
||||
* {@code FleetMcp.leadConfigDirs}. The configured profile's {@code autoCompactWindow: 100000}
|
||||
* resolves a HIGH threshold of {@code 66666} ({@link LeadContextGauge}'s {@code 2/3} fraction) —
|
||||
* far below the fixed {@code 200000} fallback a lost window argument would silently revert to.
|
||||
* {@code 90000} live tokens sits between the two: HIGH under the real window, OK under the
|
||||
* fallback — a property the fallback can never produce by accident.
|
||||
*/
|
||||
class FleetdLeadContextSourceWindowAssemblyTest {
|
||||
|
||||
private static final String LEAD_NAME = "opus";
|
||||
private static final String LEAD_TAB = "lead: opus";
|
||||
private static final String LEAD_PROFILE = "sonnet";
|
||||
/** {@code FakeHerdr}'s own default {@code agent.list} entry: terminal {@code term_a}, session {@code sess-1111}. */
|
||||
private static final String LEAD_TERMINAL = "term_a";
|
||||
private static final String LEAD_SESSION_ID = "sess-1111";
|
||||
|
||||
private static final class RecordingResourcePorts implements ResourcePorts {
|
||||
|
||||
final FakeHerdr herdr = new FakeHerdr();
|
||||
final SentinelReplyInbox replyInbox = new SentinelReplyInbox();
|
||||
|
||||
@Override
|
||||
public Map<String, String> environment() {
|
||||
return Map.of();
|
||||
}
|
||||
|
||||
@Override
|
||||
public HerdrClient connectHerdr(Path socketPath) {
|
||||
return herdr;
|
||||
}
|
||||
|
||||
@Override
|
||||
public Fleetd.AmqpOpener replyInboxOpener() {
|
||||
return (uri, prefetch) -> replyInbox;
|
||||
}
|
||||
|
||||
@Override
|
||||
public Fleetd.LeadMailboxOpener leadMailboxOpener() {
|
||||
return (uri, selfCoordId, prefetch) -> {
|
||||
throw new UnsupportedOperationException(
|
||||
"leadMailboxOpener must not be called — no coordinator: block is configured");
|
||||
};
|
||||
}
|
||||
|
||||
@Override
|
||||
public LongSupplier nanoClock() {
|
||||
return System::nanoTime;
|
||||
}
|
||||
|
||||
@Override
|
||||
public LongSupplier wallClockNanos() {
|
||||
return System::nanoTime;
|
||||
}
|
||||
|
||||
@Override
|
||||
public ScheduledExecutorService newScheduler(String purpose) {
|
||||
return Executors.newSingleThreadScheduledExecutor();
|
||||
}
|
||||
|
||||
@Override
|
||||
public void addShutdownHook(Runnable hook) {
|
||||
}
|
||||
|
||||
@Override
|
||||
public void startHttp(Javalin app, String host, int port) {
|
||||
}
|
||||
|
||||
@Override
|
||||
public Runnable herdrPollWait() {
|
||||
// Never invoked: this test's FakeHerdr answers immediately, so awaitHerdr never polls.
|
||||
return () -> {
|
||||
throw new UnsupportedOperationException("herdrPollWait must not be called — herdr is healthy");
|
||||
};
|
||||
}
|
||||
}
|
||||
|
||||
private static final class SentinelReplyInbox implements ReplyInbox, AutoCloseable {
|
||||
@Override
|
||||
public void own(String target) {
|
||||
}
|
||||
|
||||
@Override
|
||||
public void release(String target) {
|
||||
}
|
||||
|
||||
@Override
|
||||
public void publish(String target, String msgId, String content) {
|
||||
}
|
||||
|
||||
@Override
|
||||
public List<InboxMessage> peek(String target) {
|
||||
return List.of();
|
||||
}
|
||||
|
||||
@Override
|
||||
public boolean ack(String target, String msgId) {
|
||||
return false;
|
||||
}
|
||||
|
||||
@Override
|
||||
public void close() {
|
||||
}
|
||||
}
|
||||
|
||||
private static String usageLine(long tokens) {
|
||||
return "{\"type\":\"assistant\",\"message\":{\"role\":\"assistant\",\"usage\":{"
|
||||
+ "\"input_tokens\":" + tokens + ",\"cache_read_input_tokens\":0,\"cache_creation_input_tokens\":0}}}";
|
||||
}
|
||||
|
||||
/** Lays out {@code <configDir>/projects/<anySlug>/<sessionId>.jsonl} carrying one usage record. */
|
||||
private static void writeTranscript(Path configDir, String sessionId, long tokens) throws IOException {
|
||||
Path projectDir = configDir.resolve("projects").resolve("some-project-slug");
|
||||
Files.createDirectories(projectDir);
|
||||
Files.writeString(projectDir.resolve(sessionId + ".jsonl"), usageLine(tokens) + "\n", StandardCharsets.UTF_8);
|
||||
}
|
||||
|
||||
private static FleetConfig writeConfig(Path dir, String configDir) throws Exception {
|
||||
Path f = dir.resolve("fleetd.yaml");
|
||||
Files.writeString(f, """
|
||||
bind:
|
||||
host: 127.0.0.1
|
||||
port: 8765
|
||||
idleSleepGuard:
|
||||
enabled: false
|
||||
broker:
|
||||
uri: "amqp://fake-test-broker/vh"
|
||||
leadHeartbeat:
|
||||
idleAfterSeconds: 600
|
||||
backoffMs: 15000
|
||||
quietNudgeCap: 5
|
||||
fleet:
|
||||
leaders:
|
||||
%s:
|
||||
tab: "%s"
|
||||
profile: %s
|
||||
profiles:
|
||||
%s:
|
||||
subscription: true
|
||||
argv: ["ccs", "sonnet"]
|
||||
configDir: "%s"
|
||||
autoCompactWindow: 100000
|
||||
""".formatted(LEAD_NAME, LEAD_TAB, LEAD_PROFILE, LEAD_PROFILE, configDir));
|
||||
return FleetConfig.load(f);
|
||||
}
|
||||
|
||||
@SuppressWarnings("unchecked")
|
||||
private static LeadHeartbeatLoop.LeadContextSource contextSourceOf(LeadHeartbeatLoop heartbeat) throws Exception {
|
||||
Field field = LeadHeartbeatLoop.class.getDeclaredField("contextSource");
|
||||
field.setAccessible(true);
|
||||
return (LeadHeartbeatLoop.LeadContextSource) field.get(heartbeat);
|
||||
}
|
||||
|
||||
@Test
|
||||
@DisplayName("[BEHAVIOURAL] the real assembled heartbeat loop resolves HIGH against the lead's "
|
||||
+ "REAL configured window, not the fixed 200000 fallback a lost window argument reverts to")
|
||||
void assembledHeartbeatContextSourceResolvesTheRealConfiguredWindow(@TempDir Path dir) throws Exception {
|
||||
FleetConfig cfg = writeConfig(dir, dir.toString());
|
||||
writeTranscript(dir, LEAD_SESSION_ID, 90_000);
|
||||
ConfigRef config = new ConfigRef(dir.resolve("fleetd.yaml"), cfg);
|
||||
SubscriptionGuard guard = new SubscriptionGuard(cfg.guard().hostSet());
|
||||
RecordingResourcePorts ports = new RecordingResourcePorts();
|
||||
// Label FakeHerdr's own default pane's tab (term_a / w2:p7 / w2:t7, already carrying a live
|
||||
// agent on session sess-1111) to match fleet.leaders.opus.tab exactly, so LeadTabScanner
|
||||
// recognises it as the live "opus" lead without a second auto-launched pane.
|
||||
ports.herdr.withTab("w2", "w2:t7", LEAD_TAB).agentSessionId(LEAD_SESSION_ID);
|
||||
|
||||
FleetdRuntime runtime = FleetdAssembly.assembleAndStart(new AssemblyInputs(cfg, config, guard), ports);
|
||||
try {
|
||||
LeadHeartbeatLoop.LeadContextSource source = contextSourceOf(runtime.heartbeat());
|
||||
|
||||
LeadContextGauge.Reading reading = source.readingFor().apply(LEAD_TERMINAL);
|
||||
|
||||
assertEquals(LeadContextGauge.State.HIGH, reading.state(),
|
||||
"profiles." + LEAD_PROFILE + ".autoCompactWindow: 100000 resolves a HIGH threshold "
|
||||
+ "of 66666 tokens — 90000 live tokens must read HIGH against it. Mutating "
|
||||
+ "FleetdAssembly's window-lookup argument to `_ -> null` falls back to the "
|
||||
+ "fixed 200000 threshold, under which 90000 reads OK instead: " + reading);
|
||||
} finally {
|
||||
// Surefire runs the whole suite in one JVM fork (fleetd/pom.xml sets no forkCount /
|
||||
// reuseForks), so the scheduler/loops this assembly starts must be torn down here, on the
|
||||
// failure path too — hence try/finally rather than a bare statement at the end.
|
||||
runtime.close();
|
||||
}
|
||||
}
|
||||
}
|
||||
@@ -87,7 +87,8 @@ class FleetdLeadContextSourceWiringTest {
|
||||
Map<String, String> liveLeadTerminals = Map.of(LEAD_TERMINAL, LEAD_NAME);
|
||||
|
||||
LeadHeartbeatLoop.LeadContextSource source = Fleetd.leadContextSource(new LeadContextGauge(),
|
||||
agentControlStub(), () -> liveLeadTerminals, name -> LEAD_NAME.equals(name) ? tmp.toString() : null);
|
||||
agentControlStub(), () -> liveLeadTerminals,
|
||||
name -> LEAD_NAME.equals(name) ? tmp.toString() : null, name -> null);
|
||||
|
||||
LeadContextGauge.Reading reading = source.readingFor().apply(LEAD_TERMINAL);
|
||||
|
||||
@@ -101,7 +102,7 @@ class FleetdLeadContextSourceWiringTest {
|
||||
@DisplayName("an unrecognised lead terminal resolves to UNKNOWN, not a thrown exception")
|
||||
void unrecognisedTerminalResolvesToUnknown() {
|
||||
LeadHeartbeatLoop.LeadContextSource source = Fleetd.leadContextSource(new LeadContextGauge(),
|
||||
agentControlStub(), Map::of, name -> null);
|
||||
agentControlStub(), Map::of, name -> null, name -> null);
|
||||
|
||||
assertEquals(LeadContextGauge.State.UNKNOWN, source.readingFor().apply("ghost-terminal").state());
|
||||
}
|
||||
|
||||
@@ -0,0 +1,96 @@
|
||||
package dev.ltms.fleet;
|
||||
|
||||
import dev.ltms.fleet.config.FleetConfig;
|
||||
import org.junit.jupiter.api.DisplayName;
|
||||
import org.junit.jupiter.api.Test;
|
||||
|
||||
import java.util.Map;
|
||||
import java.util.function.Function;
|
||||
|
||||
import static org.junit.jupiter.api.Assertions.assertEquals;
|
||||
import static org.junit.jupiter.api.Assertions.assertNull;
|
||||
|
||||
/**
|
||||
* {@link Fleetd#leadContextWindowLookup} is the factory wired into {@code
|
||||
* FleetMcp.LeadConfigDirSource} and {@code LeadHeartbeatLoop.LeadContextSource} so {@link
|
||||
* dev.ltms.fleet.lead.LeadContextGauge} scales its HIGH threshold against a lead's own profile's
|
||||
* effective auto-compact window instead of always the gauge's fixed fallback — the same {@code
|
||||
* fleet.leaders.<name>.profile} link {@link Fleetd#leadConfigDirLookup} already follows, one step
|
||||
* further to {@link FleetConfig.Profile#effectiveAutoCompactWindow()}.
|
||||
*/
|
||||
class FleetdLeadContextWindowLookupTest {
|
||||
|
||||
private static FleetConfig.Profile profileWithWindow(String name, Integer autoCompactWindow,
|
||||
Map<String, String> env) {
|
||||
return new FleetConfig.Profile(name, null, "claude-sonnet-5", null, null, null,
|
||||
"tab", "fleet", "w #{n}", null, null, null, null, null, null, env,
|
||||
null, null, true, null, null, null, null, null, autoCompactWindow, null);
|
||||
}
|
||||
|
||||
private static FleetConfig.Leader leadOnProfile(String profile) {
|
||||
return new FleetConfig.Leader(profile, "lead: primary", 1, "lead:", 10, "claude", "claude-sonnet-5");
|
||||
}
|
||||
|
||||
@Test
|
||||
@DisplayName("a lead on a profile that sets autoCompactWindow resolves to that window")
|
||||
void leadOnAProfileWithAutoCompactWindowResolvesToIt() {
|
||||
Map<String, FleetConfig.Profile> profiles =
|
||||
Map.of("opus", profileWithWindow("opus", 250_000, Map.of()));
|
||||
Map<String, FleetConfig.Leader> leaders = Map.of("primary", leadOnProfile("opus"));
|
||||
Function<String, Long> lookup = Fleetd.leadContextWindowLookup(() -> profiles, leaders);
|
||||
|
||||
assertEquals(250_000L, lookup.apply("primary"));
|
||||
}
|
||||
|
||||
@Test
|
||||
@DisplayName("yaml and env disagree: the lookup resolves the env value, not the yaml one")
|
||||
void yamlAndEnvDisagreeLookupResolvesTheEnvValue() {
|
||||
Map<String, FleetConfig.Profile> profiles = Map.of("opus",
|
||||
profileWithWindow("opus", 250_000, Map.of("CLAUDE_CODE_AUTO_COMPACT_WINDOW", "150000")));
|
||||
Map<String, FleetConfig.Leader> leaders = Map.of("primary", leadOnProfile("opus"));
|
||||
Function<String, Long> lookup = Fleetd.leadContextWindowLookup(() -> profiles, leaders);
|
||||
|
||||
assertEquals(150_000L, lookup.apply("primary"));
|
||||
}
|
||||
|
||||
@Test
|
||||
@DisplayName("a lead entry with no `profile:` resolves to null, not a thrown exception")
|
||||
void recogniseOnlyLeadWithNoProfileResolvesToNull() {
|
||||
Map<String, FleetConfig.Profile> profiles =
|
||||
Map.of("opus", profileWithWindow("opus", 250_000, Map.of()));
|
||||
FleetConfig.Leader recogniseOnly = new FleetConfig.Leader(null, "lead: primary", 1, "lead:", 10,
|
||||
"claude", "claude-sonnet-5");
|
||||
Map<String, FleetConfig.Leader> leaders = Map.of("primary", recogniseOnly);
|
||||
Function<String, Long> lookup = Fleetd.leadContextWindowLookup(() -> profiles, leaders);
|
||||
|
||||
assertNull(lookup.apply("primary"));
|
||||
}
|
||||
|
||||
@Test
|
||||
@DisplayName("a lead naming a profile that is not configured resolves to null, not a thrown exception")
|
||||
void leadOnAnUnconfiguredProfileResolvesToNull() {
|
||||
Map<String, FleetConfig.Leader> leaders = Map.of("primary", leadOnProfile("ghost-profile"));
|
||||
Function<String, Long> lookup = Fleetd.leadContextWindowLookup(Map::of, leaders);
|
||||
|
||||
assertNull(lookup.apply("primary"));
|
||||
}
|
||||
|
||||
@Test
|
||||
@DisplayName("a lead on a profile that resolves no window at all resolves to null")
|
||||
void leadOnAProfileWithNoWindowResolvesToNull() {
|
||||
Map<String, FleetConfig.Profile> profiles =
|
||||
Map.of("opus", profileWithWindow("opus", null, Map.of()));
|
||||
Map<String, FleetConfig.Leader> leaders = Map.of("primary", leadOnProfile("opus"));
|
||||
Function<String, Long> lookup = Fleetd.leadContextWindowLookup(() -> profiles, leaders);
|
||||
|
||||
assertNull(lookup.apply("primary"));
|
||||
}
|
||||
|
||||
@Test
|
||||
@DisplayName("an unrecognised lead name resolves to null, not a thrown exception")
|
||||
void unrecognisedLeadNameResolvesToNull() {
|
||||
Function<String, Long> lookup = Fleetd.leadContextWindowLookup(Map::of, Map.of());
|
||||
|
||||
assertNull(lookup.apply("ghost-lead"));
|
||||
}
|
||||
}
|
||||
@@ -3,6 +3,7 @@ package dev.ltms.fleet;
|
||||
import ch.qos.logback.classic.Level;
|
||||
import ch.qos.logback.classic.spi.ILoggingEvent;
|
||||
import dev.ltms.fleet.config.FleetConfig;
|
||||
import dev.ltms.fleet.msg.LeadChannelHandle;
|
||||
import dev.ltms.fleet.msg.LeadMailbox;
|
||||
import dev.ltms.fleet.testing.CapturedLog;
|
||||
import org.junit.jupiter.api.Test;
|
||||
@@ -35,7 +36,7 @@ class FleetdLeadMailboxSelectionTest {
|
||||
boolean unreachable;
|
||||
|
||||
@Override
|
||||
public LeadMailbox open(String uri, String selfCoordId, int prefetch) {
|
||||
public LeadChannelHandle open(String uri, String selfCoordId, int prefetch) {
|
||||
this.offeredUri = uri;
|
||||
this.offeredSelfId = selfCoordId;
|
||||
this.offeredPrefetch = prefetch;
|
||||
|
||||
@@ -0,0 +1,293 @@
|
||||
package dev.ltms.fleet;
|
||||
|
||||
import dev.ltms.fleet.config.ConfigRef;
|
||||
import dev.ltms.fleet.config.FleetConfig;
|
||||
import dev.ltms.fleet.guard.SubscriptionGuard;
|
||||
import dev.ltms.fleet.herdr.AgentControl;
|
||||
import dev.ltms.fleet.herdr.FakeHerdr;
|
||||
import dev.ltms.fleet.herdr.HerdrClient;
|
||||
import dev.ltms.fleet.lead.LeadRollover;
|
||||
import dev.ltms.fleet.msg.ReplyInbox;
|
||||
import io.javalin.Javalin;
|
||||
import org.junit.jupiter.api.DisplayName;
|
||||
import org.junit.jupiter.api.Test;
|
||||
import org.junit.jupiter.api.io.TempDir;
|
||||
|
||||
import java.nio.file.Files;
|
||||
import java.nio.file.Path;
|
||||
import java.util.LinkedHashMap;
|
||||
import java.util.List;
|
||||
import java.util.Map;
|
||||
import java.util.concurrent.Executors;
|
||||
import java.util.concurrent.ScheduledExecutorService;
|
||||
import java.util.function.LongSupplier;
|
||||
|
||||
import static org.junit.jupiter.api.Assertions.assertEquals;
|
||||
import static org.junit.jupiter.api.Assertions.assertNotNull;
|
||||
import static org.junit.jupiter.api.Assertions.assertNull;
|
||||
import static org.junit.jupiter.api.Assertions.assertTrue;
|
||||
import static org.junit.jupiter.api.Assertions.fail;
|
||||
|
||||
/**
|
||||
* fleetd #612 B3 — replaces {@code FleetdLeadRolloverWiringTest} (fleetd #480). That class was a
|
||||
* source-text test scraping {@code Fleetd.java} (now {@code FleetdAssembly.java}, moved there by
|
||||
* fleetd #612 Unit A) with three methods: {@code unrelatedAnchorStillPresent} (a scaffold anchor,
|
||||
* not an independent claim — needs no replacement of its own), {@code
|
||||
* mainStillCallsTheLeadRolloverFactory} (the call-site pin replaced by {@link
|
||||
* #assembledLeadRolloverRunsTheRealClearAndBootstrapSequence}), and {@code
|
||||
* factoryGatesOnConfigPresence} (the absent-config claim replaced by {@link
|
||||
* #absentLeadRolloverConfigMeansNoRolloverIsBuilt} — a claim this ticket found was NOT actually
|
||||
* covered behaviourally anywhere else: {@code LeadRolloverTest}'s only related assertion is
|
||||
* vacuous, {@code assertNull(null)}, and never calls the real factory).
|
||||
*
|
||||
* <p><strong>fleetd #612 B3 correction (ticket comment 17553):</strong> the first version of this
|
||||
* test configured a single shared {@link FakeHerdr} for both the lead and member herdr sockets.
|
||||
* {@code FleetdAssembly.java:140-142} falls back to {@code memberHerdr = herdr} whenever no
|
||||
* distinct {@code memberHerdrSocket} is configured, so with one fake, {@code
|
||||
* router.leadAgents()} and {@code router.memberAgents()} wrapped the identical client — a
|
||||
* mutation swapping {@code Fleetd.leadRollover(cfg, router.leadAgents(), config, leads)} for
|
||||
* {@code ..., router.memberAgents(), ...} at {@code FleetdAssembly.java:408} was therefore
|
||||
* invisible to this test, even though the two are genuinely different daemons in production. This
|
||||
* version configures two distinct sockets and two distinct {@link FakeHerdr} instances (the same
|
||||
* pattern {@code FleetdAssemblyConnectionIdentityTest}, fleetd #612 B2, already uses to separate
|
||||
* lead from member) and asserts the roll's {@code /clear}/bootstrap sends land on the LEAD fake
|
||||
* and never on the MEMBER one.
|
||||
*/
|
||||
class FleetdLeadRolloverAssemblyTest {
|
||||
|
||||
private static final Path LEAD_SOCKET = Path.of("/fake/lead-herdr.sock");
|
||||
private static final Path MEMBER_SOCKET = Path.of("/fake/member-herdr.sock");
|
||||
|
||||
/** Keys {@code connectHerdr} by socket path so the lead and member daemons can be two
|
||||
* DIFFERENT {@link FakeHerdr}s — same shape as B2's {@code FleetdAssemblyConnectionIdentityTest
|
||||
* .TwoHerdrResourcePorts}. */
|
||||
private static final class RecordingResourcePorts implements ResourcePorts {
|
||||
|
||||
final Map<Path, HerdrClient> herdrsBySocket = new LinkedHashMap<>();
|
||||
final SentinelReplyInbox replyInbox = new SentinelReplyInbox();
|
||||
|
||||
@Override
|
||||
public Map<String, String> environment() {
|
||||
return Map.of();
|
||||
}
|
||||
|
||||
@Override
|
||||
public HerdrClient connectHerdr(Path socketPath) {
|
||||
HerdrClient client = herdrsBySocket.get(socketPath);
|
||||
if (client == null) {
|
||||
throw new IllegalStateException("no fake herdr registered for socket " + socketPath);
|
||||
}
|
||||
return client;
|
||||
}
|
||||
|
||||
@Override
|
||||
public Fleetd.AmqpOpener replyInboxOpener() {
|
||||
return (uri, prefetch) -> replyInbox;
|
||||
}
|
||||
|
||||
@Override
|
||||
public Fleetd.LeadMailboxOpener leadMailboxOpener() {
|
||||
return (uri, selfCoordId, prefetch) -> {
|
||||
throw new UnsupportedOperationException(
|
||||
"leadMailboxOpener must not be called — no coordinator: block is configured");
|
||||
};
|
||||
}
|
||||
|
||||
@Override
|
||||
public LongSupplier nanoClock() {
|
||||
return System::nanoTime;
|
||||
}
|
||||
|
||||
@Override
|
||||
public LongSupplier wallClockNanos() {
|
||||
return System::nanoTime;
|
||||
}
|
||||
|
||||
@Override
|
||||
public ScheduledExecutorService newScheduler(String purpose) {
|
||||
return Executors.newSingleThreadScheduledExecutor();
|
||||
}
|
||||
|
||||
@Override
|
||||
public void addShutdownHook(Runnable hook) {
|
||||
}
|
||||
|
||||
@Override
|
||||
public void startHttp(Javalin app, String host, int port) {
|
||||
}
|
||||
|
||||
@Override
|
||||
public Runnable herdrPollWait() {
|
||||
// Never invoked: this test's FakeHerdr answers immediately, so awaitHerdr never polls.
|
||||
return () -> {
|
||||
throw new UnsupportedOperationException("herdrPollWait must not be called — herdr is healthy");
|
||||
};
|
||||
}
|
||||
}
|
||||
|
||||
private static final class SentinelReplyInbox implements ReplyInbox, AutoCloseable {
|
||||
@Override
|
||||
public void own(String target) {
|
||||
}
|
||||
|
||||
@Override
|
||||
public void release(String target) {
|
||||
}
|
||||
|
||||
@Override
|
||||
public void publish(String target, String msgId, String content) {
|
||||
}
|
||||
|
||||
@Override
|
||||
public List<InboxMessage> peek(String target) {
|
||||
return List.of();
|
||||
}
|
||||
|
||||
@Override
|
||||
public boolean ack(String target, String msgId) {
|
||||
return false;
|
||||
}
|
||||
|
||||
@Override
|
||||
public void close() {
|
||||
}
|
||||
}
|
||||
|
||||
private static FleetConfig writeConfig(Path dir, Path leadCwd) throws Exception {
|
||||
Path f = dir.resolve("fleetd.yaml");
|
||||
Files.writeString(f, """
|
||||
bind:
|
||||
host: 127.0.0.1
|
||||
port: 8765
|
||||
herdrSocket: "%s"
|
||||
memberHerdrSocket: "%s"
|
||||
idleSleepGuard:
|
||||
enabled: false
|
||||
broker:
|
||||
uri: "amqp://fake-test-broker/vh"
|
||||
fleet:
|
||||
leaders:
|
||||
opus:
|
||||
tab: "lead: opus"
|
||||
cwd: "%s"
|
||||
leadRollover:
|
||||
handoverPath: handover.md
|
||||
requireOperatorConfirm: false
|
||||
""".formatted(LEAD_SOCKET, MEMBER_SOCKET, leadCwd.toString()));
|
||||
return FleetConfig.load(f);
|
||||
}
|
||||
|
||||
@SuppressWarnings("unchecked")
|
||||
@Test
|
||||
@DisplayName("[BEHAVIOURAL] the real assembled LeadRollover runs the full open/confirm/continuation "
|
||||
+ "sequence — /clear, then bootstrapText — through the real herdr router")
|
||||
void assembledLeadRolloverRunsTheRealClearAndBootstrapSequence(@TempDir Path dir) throws Exception {
|
||||
Path leadCwd = dir.resolve("lead-workspace");
|
||||
Files.createDirectories(leadCwd);
|
||||
FleetConfig cfg = writeConfig(dir, leadCwd);
|
||||
ConfigRef config = new ConfigRef(dir.resolve("fleetd.yaml"), cfg);
|
||||
SubscriptionGuard guard = new SubscriptionGuard(cfg.guard().hostSet());
|
||||
RecordingResourcePorts ports = new RecordingResourcePorts();
|
||||
// Two DISTINCT fakes — one per configured socket — so leadAgents()/memberAgents() wrap
|
||||
// genuinely different clients, exactly like production when memberHerdrSocket is set.
|
||||
FakeHerdr lead = new FakeHerdr();
|
||||
lead.withTab("w2", "w2:t7", "lead: opus");
|
||||
FakeHerdr member = new FakeHerdr();
|
||||
ports.herdrsBySocket.put(LEAD_SOCKET, lead);
|
||||
ports.herdrsBySocket.put(MEMBER_SOCKET, member);
|
||||
|
||||
FleetdRuntime runtime = FleetdAssembly.assembleAndStart(new AssemblyInputs(cfg, config, guard), ports);
|
||||
|
||||
LeadRollover rollover = runtime.mcp().leadRollover();
|
||||
assertNotNull(rollover, "leadRollover: is present in this test's config, so "
|
||||
+ "FleetdAssembly.assembleAndStart must have built a real LeadRollover through the "
|
||||
+ "Fleetd.leadRollover(...) call site — a mutation to `LeadRollover leadRollover = "
|
||||
+ "null;` at that call site can never pass this");
|
||||
|
||||
LeadRollover.PendingRollover pending = rollover.open("term_a", "fleetd #612 B3 test");
|
||||
String expectedHandoverPath = leadCwd.resolve("handover.md").normalize().toString();
|
||||
assertEquals(expectedHandoverPath, pending.handoverPath());
|
||||
|
||||
// Ensure the handover file's mtime lands strictly AFTER open()'s requestedAtMillis —
|
||||
// LeadRollover.checkHandover refuses on mtime <= requestedAt (HANDOVER_STALE).
|
||||
Thread.sleep(50);
|
||||
Files.writeString(Path.of(pending.handoverPath()), "handover content for fleetd #612 B3");
|
||||
|
||||
LeadRollover.RollDecision decision = rollover.confirm("term_a", pending.token(), true);
|
||||
assertTrue(decision.accepted(), "confirm() must approve: requireOperatorConfirm is false, "
|
||||
+ "the caller terminal matches open()'s, and the handover file exists, is non-empty "
|
||||
+ "and fresh — got: " + decision);
|
||||
|
||||
// The production LeadRollover constructor always runs the post-confirm continuation on a
|
||||
// real virtual thread (see Fleetd.leadRollover, which never passes the package-private test
|
||||
// constructor), so this polls the real FleetMcp.leadRollover() instance's status(token)
|
||||
// until the real continuation finishes.
|
||||
LeadRollover.RollStatus status = pollUntilTerminal(rollover, pending.token());
|
||||
assertEquals(LeadRollover.RollState.ROLLED, status.state(),
|
||||
"the full happy path must complete: FakeHerdr's default agent status is 'idle', so "
|
||||
+ "the turn-boundary wait settles immediately and the post-/clear wait "
|
||||
+ "releases via its pickup-grace path — detail: " + status.detail());
|
||||
|
||||
// Prove the real herdr router actually sent BOTH messages, in order, to the real LEAD
|
||||
// pane — this is the one thing a source-text pin on the call site could never show.
|
||||
List<FakeHerdr.Call> prompts = lead.calls.stream()
|
||||
.filter(c -> c.method().equals("agent.prompt"))
|
||||
.toList();
|
||||
assertTrue(prompts.size() >= 2, "expected at least a /clear send and a bootstrapText send "
|
||||
+ "on the LEAD daemon, got " + prompts.size() + " agent.prompt calls: " + prompts);
|
||||
assertEquals("/clear", ((Map<String, Object>) prompts.get(0).params()).get("text"),
|
||||
"the first send must be the literal /clear housekeeping command");
|
||||
Object secondText = ((Map<String, Object>) prompts.get(1).params()).get("text");
|
||||
assertTrue(secondText instanceof String && ((String) secondText).contains(expectedHandoverPath),
|
||||
"the second send must be the default bootstrapText naming the resolved handover "
|
||||
+ "path, got: " + secondText);
|
||||
|
||||
// fleetd #612 B3 correction: prove the roll never touches the MEMBER daemon. A mutation
|
||||
// swapping router.leadAgents() for router.memberAgents() at the real call site would move
|
||||
// both sends above onto `member` instead, which this assertion catches — the thing the
|
||||
// single-fake version of this test could never see, because both wrapped the same client.
|
||||
List<FakeHerdr.Call> memberPrompts = member.calls.stream()
|
||||
.filter(c -> c.method().equals("agent.prompt"))
|
||||
.toList();
|
||||
assertTrue(memberPrompts.isEmpty(), "the roll must be wired to the LEAD daemon only — got "
|
||||
+ memberPrompts.size() + " agent.prompt call(s) on the MEMBER daemon instead: "
|
||||
+ memberPrompts);
|
||||
}
|
||||
|
||||
private static LeadRollover.RollStatus pollUntilTerminal(LeadRollover rollover, String token)
|
||||
throws InterruptedException {
|
||||
long deadline = System.nanoTime() + java.util.concurrent.TimeUnit.SECONDS.toNanos(10);
|
||||
while (System.nanoTime() < deadline) {
|
||||
LeadRollover.RollStatus status = rollover.status(token);
|
||||
if (status.state() != LeadRollover.RollState.PENDING
|
||||
&& status.state() != LeadRollover.RollState.IN_PROGRESS) {
|
||||
return status;
|
||||
}
|
||||
Thread.sleep(50);
|
||||
}
|
||||
fail("the real continuation did not reach a terminal state within 10s — last status: "
|
||||
+ rollover.status(token));
|
||||
throw new AssertionError("unreachable");
|
||||
}
|
||||
|
||||
@Test
|
||||
@DisplayName("[BEHAVIOURAL] Fleetd.leadRollover(...) returns null when leadRollover: is absent "
|
||||
+ "from config — the opt-in gate FleetdLeadRolloverWiringTest's "
|
||||
+ "factoryGatesOnConfigPresence pinned by source text alone")
|
||||
void absentLeadRolloverConfigMeansNoRolloverIsBuilt(@TempDir Path dir) throws Exception {
|
||||
Path yaml = dir.resolve("fleetd.yaml");
|
||||
Files.writeString(yaml, """
|
||||
bind:
|
||||
host: 127.0.0.1
|
||||
port: 8765
|
||||
""");
|
||||
ConfigRef config = new ConfigRef(yaml, FleetConfig.load(yaml));
|
||||
AgentControl agents = new AgentControl(new FakeHerdr());
|
||||
|
||||
LeadRollover rollover = Fleetd.leadRollover(config.get(), agents, config, Map::of);
|
||||
|
||||
assertNull(rollover, "leadRollover: is absent from this config, so the factory's opt-in "
|
||||
+ "gate (`if (cfg.leadRollover() == null) return null;`) must fire and no "
|
||||
+ "LeadRollover must be constructed at all");
|
||||
}
|
||||
}
|
||||
@@ -1,89 +0,0 @@
|
||||
package dev.ltms.fleet;
|
||||
|
||||
import java.nio.file.Files;
|
||||
import java.nio.file.Path;
|
||||
import org.junit.jupiter.api.DisplayName;
|
||||
import org.junit.jupiter.api.Test;
|
||||
|
||||
import static org.junit.jupiter.api.Assertions.assertTrue;
|
||||
|
||||
/**
|
||||
* fleetd #480 Unit A, hard requirement 6: pin {@code Fleetd.main}'s construction of {@link
|
||||
* dev.ltms.fleet.lead.LeadRollover} with a source-text assertion, mirroring {@code
|
||||
* FleetdCompletionResolverWiringTest}'s pattern — five log-only reporters in {@code Fleetd.main}
|
||||
* already survived mutation batteries this exact way (fleetd #415's extraction antidote note).
|
||||
*
|
||||
* <p>What this class still covers, and what it never claimed to. {@code LeadRolloverTest}
|
||||
* constructs its own {@code LeadRollover} directly (as every prior test of an extracted factory
|
||||
* does) with a hand-built lookup, so a mutation that deletes the {@code leadRollover(...)} call
|
||||
* from {@code main} — or replaces one of its arguments with something that still compiles, e.g.
|
||||
* {@code router.leadAgents()} swapped for {@code null}, or the whole assignment swapped for a bare
|
||||
* {@code null} literal — leaves every behavioural test green. This is a plain string read, guarded
|
||||
* by an unrelated anchor assertion so a broken or empty file read cannot pass as a real change.
|
||||
*
|
||||
* <p><b>This test checks source text, not runtime behaviour.</b> It never constructs a {@code
|
||||
* LeadRollover} and never runs {@code main}. It pins the {@code leadRollover(...)} CALL SITE's
|
||||
* argument list — that {@code main} still passes {@code leads} at all — never what the factory
|
||||
* DOES with that argument once inside its own body.
|
||||
*
|
||||
* <p><b>Correction (fleetd #480 relative-handover-path follow-up): that gap used to be real, and
|
||||
* now is not — but not here.</b> This class's javadoc previously claimed "no behavioural test can
|
||||
* catch this wiring dropping out" for the whole factory, including the lambda {@code
|
||||
* leadRollover(...)} builds internally (terminal → lead name → {@code Leader.cwd()}). That claim
|
||||
* was proven true at the time — mutating that lambda's body to {@code String leadName = null;}
|
||||
* (always "no lead found", which silently reintroduces the daemon-cwd bug this ticket fixes) left
|
||||
* the full suite green, {@code Tests run: 1669, Failures: 0}. It is no longer true: {@code
|
||||
* FleetdLeadRolloverWorkspaceLookupTest} now calls {@code Fleetd.leadRollover(...)} directly with a
|
||||
* real {@link dev.ltms.fleet.config.ConfigRef} built from a temp {@code fleetd.yaml}, and fails
|
||||
* against that exact one-line mutation. So: THIS class still covers only the call site's argument
|
||||
* list; {@code FleetdLeadRolloverWorkspaceLookupTest} is what now covers the lambda's body. Neither
|
||||
* one subsumes the other — keep both.
|
||||
*/
|
||||
class FleetdLeadRolloverWiringTest {
|
||||
|
||||
private static String fleetdSource() throws Exception {
|
||||
return Files.readString(Path.of("src/main/java/dev/ltms/fleet/Fleetd.java"));
|
||||
}
|
||||
|
||||
@Test
|
||||
@DisplayName("[SOURCE TEXT] unrelated anchor: Fleetd.java still declares the Fleetd class")
|
||||
void unrelatedAnchorStillPresent() throws Exception {
|
||||
// Guards the two assertions below: without this, a bad read (empty string, wrong file,
|
||||
// truncated file) could vacuously fail to contain the leadRollover(...) call too, and a
|
||||
// test that only asserts "contains X" would report a false pass for the wrong reason if X
|
||||
// happened to match. Asserting an unrelated, structurally distant string first proves the
|
||||
// read actually pulled real file content.
|
||||
String source = fleetdSource();
|
||||
assertTrue(source.contains("public final class Fleetd"),
|
||||
"sanity anchor failed — the file read did not return real Fleetd.java source; the "
|
||||
+ "leadRollover(...) wiring assertions below cannot be trusted until this passes");
|
||||
}
|
||||
|
||||
@Test
|
||||
@DisplayName("[SOURCE TEXT] main still constructs LeadRollover via the leadRollover(...) factory, exactly as heartbeat is constructed")
|
||||
void mainStillCallsTheLeadRolloverFactory() throws Exception {
|
||||
String source = fleetdSource();
|
||||
assertTrue(source.contains(
|
||||
"LeadRollover leadRollover = leadRollover(cfg, router.leadAgents(), config, leads);"),
|
||||
"Fleetd.main must still assign `LeadRollover leadRollover = leadRollover(cfg, "
|
||||
+ "router.leadAgents(), config, leads);`. Dropping this call, or swapping one of "
|
||||
+ "its arguments for something that still compiles (e.g. null in place of "
|
||||
+ "router.leadAgents()), leaves every behavioural test green — this source check is "
|
||||
+ "what must go red instead. fleetd #480 correction 2 deliberately dropped "
|
||||
+ "primaryRegistry from this call — see LeadRollover's class javadoc for why a "
|
||||
+ "single-slot lookup was wrong here. The fleetd #480 relative-handover-path "
|
||||
+ "follow-up added `leads` (terminal → lead name) so the factory can resolve a "
|
||||
+ "relative handoverPath against the calling lead's own workspace.");
|
||||
}
|
||||
|
||||
@Test
|
||||
@DisplayName("[SOURCE TEXT] the leadRollover(...) factory itself gates construction on cfg.leadRollover() != null")
|
||||
void factoryGatesOnConfigPresence() throws Exception {
|
||||
String source = fleetdSource();
|
||||
assertTrue(source.contains("if (cfg.leadRollover() == null) {"),
|
||||
"Fleetd.leadRollover(...) must refuse to construct a LeadRollover when the "
|
||||
+ "leadRollover: block is absent — an upgraded daemon must never silently acquire "
|
||||
+ "the ability to clear the lead's own pane. See LeadHeartbeatLoop's construction "
|
||||
+ "gate (cfg.leadHeartbeat() != null) for the pattern this mirrors.");
|
||||
}
|
||||
}
|
||||
@@ -0,0 +1,180 @@
|
||||
package dev.ltms.fleet;
|
||||
|
||||
import dev.ltms.fleet.config.ConfigRef;
|
||||
import dev.ltms.fleet.config.FleetConfig;
|
||||
import dev.ltms.fleet.guard.SubscriptionGuard;
|
||||
import dev.ltms.fleet.herdr.FakeHerdr;
|
||||
import dev.ltms.fleet.herdr.HerdrClient;
|
||||
import dev.ltms.fleet.mcp.FleetMcp;
|
||||
import dev.ltms.fleet.msg.ReplyInbox;
|
||||
import io.javalin.Javalin;
|
||||
import org.junit.jupiter.api.DisplayName;
|
||||
import org.junit.jupiter.api.Test;
|
||||
import org.junit.jupiter.api.io.TempDir;
|
||||
|
||||
import java.nio.file.Files;
|
||||
import java.nio.file.Path;
|
||||
import java.util.List;
|
||||
import java.util.Map;
|
||||
import java.util.concurrent.Executors;
|
||||
import java.util.concurrent.ScheduledExecutorService;
|
||||
import java.util.function.LongSupplier;
|
||||
|
||||
import static org.junit.jupiter.api.Assertions.assertEquals;
|
||||
|
||||
/**
|
||||
* fleetd #612 B3 — replaces {@code FleetdLeadSeatWiringTest} (fleetd #176), a source-text test that
|
||||
* scraped {@code Fleetd.java} (now {@code FleetdAssembly.java}, moved there by fleetd #612 Unit A)
|
||||
* for the exact {@code new FleetMcp.LeadSeatSource(Fleetd.leadSeatLookup(...))} constructor-call
|
||||
* text. That proves the right symbols appear in source; it proves nothing about what the daemon's
|
||||
* live {@code fleet_list} actually reports.
|
||||
*
|
||||
* <p>This test instead drives the REAL {@link FleetMcp.LeadSeatSource} the real {@link
|
||||
* FleetdAssembly#assembleAndStart} builds — including the REAL {@code LeadTabScanner} it wires
|
||||
* {@code Fleetd.leadSeatLookup} through — reached via {@link FleetMcp#leadSeatSource()} on the
|
||||
* live {@code FleetMcp} {@code FleetdRuntime} owns. It seeds one FakeHerdr tab labelled to match a
|
||||
* configured {@code fleet.leaders.opus.tab}, with a live agent already in it (FakeHerdr's own
|
||||
* default {@code agent.list}/{@code pane.list} entries for {@code term_a}/{@code w2:p7}/{@code
|
||||
* w2:t7} — no FakeHerdr change needed), and asserts the assembled seat source reports exactly the
|
||||
* seat {@link FleetMcp.LeadSeatSource#none()} (the inert stand-in) could never produce: 1, not 0.
|
||||
*/
|
||||
class FleetdLeadSeatAssemblyTest {
|
||||
|
||||
private static final class RecordingResourcePorts implements ResourcePorts {
|
||||
|
||||
final FakeHerdr herdr = new FakeHerdr();
|
||||
final SentinelReplyInbox replyInbox = new SentinelReplyInbox();
|
||||
|
||||
@Override
|
||||
public Map<String, String> environment() {
|
||||
return Map.of();
|
||||
}
|
||||
|
||||
@Override
|
||||
public HerdrClient connectHerdr(Path socketPath) {
|
||||
return herdr;
|
||||
}
|
||||
|
||||
@Override
|
||||
public Fleetd.AmqpOpener replyInboxOpener() {
|
||||
return (uri, prefetch) -> replyInbox;
|
||||
}
|
||||
|
||||
@Override
|
||||
public Fleetd.LeadMailboxOpener leadMailboxOpener() {
|
||||
return (uri, selfCoordId, prefetch) -> {
|
||||
throw new UnsupportedOperationException(
|
||||
"leadMailboxOpener must not be called — no coordinator: block is configured");
|
||||
};
|
||||
}
|
||||
|
||||
@Override
|
||||
public LongSupplier nanoClock() {
|
||||
return System::nanoTime;
|
||||
}
|
||||
|
||||
@Override
|
||||
public LongSupplier wallClockNanos() {
|
||||
return System::nanoTime;
|
||||
}
|
||||
|
||||
@Override
|
||||
public ScheduledExecutorService newScheduler(String purpose) {
|
||||
return Executors.newSingleThreadScheduledExecutor();
|
||||
}
|
||||
|
||||
@Override
|
||||
public void addShutdownHook(Runnable hook) {
|
||||
}
|
||||
|
||||
@Override
|
||||
public void startHttp(Javalin app, String host, int port) {
|
||||
}
|
||||
|
||||
@Override
|
||||
public Runnable herdrPollWait() {
|
||||
// Never invoked: this test's FakeHerdr answers immediately, so awaitHerdr never polls.
|
||||
return () -> {
|
||||
throw new UnsupportedOperationException("herdrPollWait must not be called — herdr is healthy");
|
||||
};
|
||||
}
|
||||
}
|
||||
|
||||
private static final class SentinelReplyInbox implements ReplyInbox, AutoCloseable {
|
||||
@Override
|
||||
public void own(String target) {
|
||||
}
|
||||
|
||||
@Override
|
||||
public void release(String target) {
|
||||
}
|
||||
|
||||
@Override
|
||||
public void publish(String target, String msgId, String content) {
|
||||
}
|
||||
|
||||
@Override
|
||||
public List<InboxMessage> peek(String target) {
|
||||
return List.of();
|
||||
}
|
||||
|
||||
@Override
|
||||
public boolean ack(String target, String msgId) {
|
||||
return false;
|
||||
}
|
||||
|
||||
@Override
|
||||
public void close() {
|
||||
}
|
||||
}
|
||||
|
||||
private static FleetConfig writeConfig(Path dir) throws Exception {
|
||||
Path f = dir.resolve("fleetd.yaml");
|
||||
Files.writeString(f, """
|
||||
bind:
|
||||
host: 127.0.0.1
|
||||
port: 8765
|
||||
idleSleepGuard:
|
||||
enabled: false
|
||||
broker:
|
||||
uri: "amqp://fake-test-broker/vh"
|
||||
fleet:
|
||||
leaders:
|
||||
opus:
|
||||
tab: "lead: opus"
|
||||
profile: sonnet
|
||||
profiles:
|
||||
sonnet:
|
||||
subscription: true
|
||||
argv: ["ccs", "sonnet"]
|
||||
""");
|
||||
return FleetConfig.load(f);
|
||||
}
|
||||
|
||||
@Test
|
||||
@DisplayName("[BEHAVIOURAL] the real assembled LeadSeatSource, backed by the real LeadTabScanner, "
|
||||
+ "reports a live lead's seat against its own subscription profile")
|
||||
void assembledLeadSeatSourceReportsALiveLeadsSeat(@TempDir Path dir) throws Exception {
|
||||
FleetConfig cfg = writeConfig(dir);
|
||||
ConfigRef config = new ConfigRef(dir.resolve("fleetd.yaml"), cfg);
|
||||
SubscriptionGuard guard = new SubscriptionGuard(cfg.guard().hostSet());
|
||||
RecordingResourcePorts ports = new RecordingResourcePorts();
|
||||
// Label FakeHerdr's own default pane's tab (term_a / w2:p7 / w2:t7, already carrying a live
|
||||
// agent) to match fleet.leaders.opus.tab exactly — no FakeHerdr change needed at all.
|
||||
ports.herdr.withTab("w2", "w2:t7", "lead: opus");
|
||||
|
||||
FleetdRuntime runtime = FleetdAssembly.assembleAndStart(new AssemblyInputs(cfg, config, guard), ports);
|
||||
|
||||
FleetMcp.LeadSeatSource seatSource = runtime.mcp().leadSeatSource();
|
||||
assertEquals(1, seatSource.seatsFor().apply("sonnet"),
|
||||
"the real LeadTabScanner recognises the labelled tab as a live 'opus' lead on "
|
||||
+ "profile 'sonnet' (same credential, matched by Fleetd.leadSeatLookup), so "
|
||||
+ "subscription profile 'sonnet' must be charged one seat — "
|
||||
+ "FleetMcp.LeadSeatSource.none() (the inert stand-in this test's mutation "
|
||||
+ "swaps the call site for) always reports 0, whatever the input");
|
||||
|
||||
// A profile no lead is running on gets no seat charged — the same seat source, applied to
|
||||
// an input that must stay at the inert answer even on the real, non-inert instance.
|
||||
assertEquals(0, seatSource.seatsFor().apply("no-such-profile"));
|
||||
}
|
||||
}
|
||||
@@ -1,43 +0,0 @@
|
||||
package dev.ltms.fleet;
|
||||
|
||||
import java.nio.file.Files;
|
||||
import java.nio.file.Path;
|
||||
import org.junit.jupiter.api.DisplayName;
|
||||
import org.junit.jupiter.api.Test;
|
||||
|
||||
import static org.junit.jupiter.api.Assertions.assertTrue;
|
||||
|
||||
/**
|
||||
* fleetd #176: {@code Fleetd.main} builds its {@code FleetMcp} from a 14-argument constructor whose
|
||||
* last argument is a {@code FleetMcp.LeadSeatSource} wrapping {@link Fleetd#leadSeatLookup}. That
|
||||
* argument is exactly the kind of wiring fleetd #248 warned about: dropping it (or swapping it for
|
||||
* the inert {@code FleetMcp.LeadSeatSource.none()}) compiles with 0 errors and leaves every test
|
||||
* that builds its own {@code FleetMcp}/{@code CapacitySource} directly — every test that predates
|
||||
* this ticket — green, because none of them go through {@code main} at all.
|
||||
*
|
||||
* <p>{@link FleetdLeadSeatLookupTest} proves the factory's own matching logic; this class is the
|
||||
* plain source-text assertion that proves {@code main} still passes its result in, mirroring
|
||||
* {@code FleetdCompletionResolverWiringTest}'s approach for the same class of gap.
|
||||
*
|
||||
* <p><b>This test checks source text, not runtime behaviour.</b> It never constructs a
|
||||
* {@code FleetMcp} and never runs {@code main}.
|
||||
*/
|
||||
class FleetdLeadSeatWiringTest {
|
||||
|
||||
private static String fleetdSource() throws Exception {
|
||||
return Files.readString(Path.of("src/main/java/dev/ltms/fleet/Fleetd.java"));
|
||||
}
|
||||
|
||||
@Test
|
||||
@DisplayName("[SOURCE TEXT] FleetMcp's construction call still passes a LeadSeatSource built from leadSeatLookup(...)")
|
||||
void fleetMcpConstructionStillWiresLeadSeatLookup() throws Exception {
|
||||
String source = fleetdSource();
|
||||
assertTrue(source.contains("new FleetMcp.LeadSeatSource(leadSeatLookup(() -> config.get().profiles(), "
|
||||
+ "leaders, leads))"),
|
||||
"FleetMcp's construction call must still pass a LeadSeatSource built from "
|
||||
+ "Fleetd.leadSeatLookup(...). Dropping it or swapping in "
|
||||
+ "FleetMcp.LeadSeatSource.none() (fleetd #176's would-be silent regression, the same "
|
||||
+ "shape as fleetd #248's measured mutations) compiles with 0 errors and leaves every "
|
||||
+ "existing behavioural test green — this source check is what must go red instead.");
|
||||
}
|
||||
}
|
||||
@@ -0,0 +1,243 @@
|
||||
package dev.ltms.fleet;
|
||||
|
||||
import dev.ltms.fleet.config.ConfigRef;
|
||||
import dev.ltms.fleet.config.FleetConfig;
|
||||
import dev.ltms.fleet.guard.SubscriptionGuard;
|
||||
import dev.ltms.fleet.herdr.FakeHerdr;
|
||||
import dev.ltms.fleet.herdr.HerdrClient;
|
||||
import dev.ltms.fleet.msg.LeadChannelHandle;
|
||||
import dev.ltms.fleet.msg.LeadMessage;
|
||||
import dev.ltms.fleet.msg.ReplyInbox;
|
||||
import io.javalin.Javalin;
|
||||
import io.modelcontextprotocol.client.McpClient;
|
||||
import io.modelcontextprotocol.client.McpSyncClient;
|
||||
import io.modelcontextprotocol.client.transport.HttpClientStreamableHttpTransport;
|
||||
import io.modelcontextprotocol.spec.McpClientTransport;
|
||||
import io.modelcontextprotocol.spec.McpSchema;
|
||||
import org.junit.jupiter.api.AfterEach;
|
||||
import org.junit.jupiter.api.Test;
|
||||
import org.junit.jupiter.api.Timeout;
|
||||
import org.junit.jupiter.api.io.TempDir;
|
||||
|
||||
import java.net.http.HttpRequest;
|
||||
import java.nio.file.Files;
|
||||
import java.nio.file.Path;
|
||||
import java.util.List;
|
||||
import java.util.Map;
|
||||
import java.util.concurrent.Executors;
|
||||
import java.util.concurrent.ScheduledExecutorService;
|
||||
import java.util.concurrent.TimeUnit;
|
||||
import java.util.function.LongSupplier;
|
||||
|
||||
import static org.junit.jupiter.api.Assertions.assertTrue;
|
||||
|
||||
/**
|
||||
* fleetd #612 Shape A ranks 9 and 11: drives the real {@code fleet_list} MCP route through a
|
||||
* token-authenticated primary caller. The assertions read the response from the exact {@link
|
||||
* dev.ltms.fleet.mcp.FleetMcp} instance assembled by {@link FleetdAssembly}, rather than a source
|
||||
* scrape or a separately built reporting source.
|
||||
*
|
||||
* <p>The coordinator fixture contains a peer on purpose. Both a missing coordinator and an
|
||||
* incorrectly wired peers argument can render as an empty list, so the non-empty peer assertion
|
||||
* distinguishes the real call-site value from that inert result. Its fake mailbox never contacts a
|
||||
* broker.
|
||||
*/
|
||||
class FleetdListReportingSourcesAssemblyTest {
|
||||
|
||||
private static final String TOKEN = "fleetd-r9-r11-test-token";
|
||||
private static final String TOKEN_ENV = "FLEETD_R9_TEST_TOKEN";
|
||||
private static final String PROFILE = "capacity-profile";
|
||||
private static final String PEER = "peer-fleet";
|
||||
|
||||
private static final class FakeLeadChannel implements LeadChannelHandle {
|
||||
@Override
|
||||
public void publish(String toCoordId, LeadMessage message) {
|
||||
}
|
||||
|
||||
@Override
|
||||
public List<LeadMessage> peek() {
|
||||
return List.of();
|
||||
}
|
||||
|
||||
@Override
|
||||
public void ack(String msgId) {
|
||||
}
|
||||
|
||||
@Override
|
||||
public String selfCoordId() {
|
||||
return "test-lead";
|
||||
}
|
||||
|
||||
@Override
|
||||
public boolean heldDurable() {
|
||||
return true;
|
||||
}
|
||||
|
||||
@Override
|
||||
public MailboxState inspect(String coordId) {
|
||||
return MailboxState.unknown(coordId);
|
||||
}
|
||||
|
||||
@Override
|
||||
public void close() {
|
||||
}
|
||||
}
|
||||
|
||||
private static final class TestResourcePorts implements ResourcePorts {
|
||||
private final FakeHerdr herdr = new FakeHerdr();
|
||||
private Runnable shutdownHook;
|
||||
|
||||
@Override
|
||||
public Map<String, String> environment() {
|
||||
return Map.of(TOKEN_ENV, TOKEN);
|
||||
}
|
||||
|
||||
@Override
|
||||
public HerdrClient connectHerdr(Path socketPath) {
|
||||
return herdr;
|
||||
}
|
||||
|
||||
@Override
|
||||
public Fleetd.AmqpOpener replyInboxOpener() {
|
||||
return (uri, prefetch) -> new ReplyInbox() {
|
||||
@Override public void own(String target) { }
|
||||
@Override public void release(String target) { }
|
||||
@Override public void publish(String target, String msgId, String content) { }
|
||||
@Override public List<InboxMessage> peek(String target) { return List.of(); }
|
||||
@Override public boolean ack(String target, String msgId) { return false; }
|
||||
};
|
||||
}
|
||||
|
||||
@Override
|
||||
public Fleetd.LeadMailboxOpener leadMailboxOpener() {
|
||||
return (uri, selfCoordId, prefetch) -> new FakeLeadChannel();
|
||||
}
|
||||
|
||||
@Override
|
||||
public LongSupplier nanoClock() {
|
||||
return System::nanoTime;
|
||||
}
|
||||
|
||||
@Override
|
||||
public LongSupplier wallClockNanos() {
|
||||
return System::nanoTime;
|
||||
}
|
||||
|
||||
@Override
|
||||
public ScheduledExecutorService newScheduler(String purpose) {
|
||||
return Executors.newSingleThreadScheduledExecutor();
|
||||
}
|
||||
|
||||
@Override
|
||||
public void addShutdownHook(Runnable hook) {
|
||||
shutdownHook = hook;
|
||||
}
|
||||
|
||||
@Override
|
||||
public void startHttp(Javalin app, String host, int port) {
|
||||
// The test binds runtime.app() itself, below.
|
||||
}
|
||||
|
||||
@Override
|
||||
public Runnable herdrPollWait() {
|
||||
return () -> {
|
||||
throw new UnsupportedOperationException("healthy FakeHerdr must not poll");
|
||||
};
|
||||
}
|
||||
}
|
||||
|
||||
private FleetdRuntime runtime;
|
||||
private TestResourcePorts ports;
|
||||
private Javalin boundApp;
|
||||
|
||||
@AfterEach
|
||||
void tearDown() {
|
||||
if (boundApp != null) {
|
||||
boundApp.stop();
|
||||
}
|
||||
if (ports != null && ports.shutdownHook != null) {
|
||||
ports.shutdownHook.run();
|
||||
}
|
||||
}
|
||||
|
||||
private static FleetConfig writeConfig(Path dir) throws Exception {
|
||||
Path file = dir.resolve("fleetd.yaml");
|
||||
Files.writeString(file, """
|
||||
bind:
|
||||
host: 127.0.0.1
|
||||
port: 8765
|
||||
auth:
|
||||
mode: token
|
||||
tokenEnv: %s
|
||||
idleSleepGuard:
|
||||
enabled: false
|
||||
health:
|
||||
enabled: true
|
||||
notifications:
|
||||
mode: webhook
|
||||
profiles:
|
||||
%s:
|
||||
baseUrl: http://capacity.test:8000
|
||||
model: test-model
|
||||
maxLoad: 7
|
||||
coordinator:
|
||||
uri: amqp://fake-coordinator/vh
|
||||
selfId: test-lead
|
||||
peers:
|
||||
- %s
|
||||
""".formatted(TOKEN_ENV, PROFILE, PEER));
|
||||
return FleetConfig.load(file);
|
||||
}
|
||||
|
||||
private int assembleAndBind(Path dir) throws Exception {
|
||||
FleetConfig cfg = writeConfig(dir);
|
||||
ConfigRef config = new ConfigRef(dir.resolve("fleetd.yaml"), cfg);
|
||||
ports = new TestResourcePorts();
|
||||
runtime = FleetdAssembly.assembleAndStart(new AssemblyInputs(cfg, config,
|
||||
new SubscriptionGuard(cfg.guard().hostSet())), ports);
|
||||
boundApp = runtime.app().start("127.0.0.1", 0);
|
||||
return boundApp.port();
|
||||
}
|
||||
|
||||
private static String fleetList(int port) {
|
||||
HttpRequest.Builder requestTemplate = HttpRequest.newBuilder()
|
||||
.header("Authorization", "Bearer " + TOKEN);
|
||||
McpClientTransport transport = HttpClientStreamableHttpTransport.builder("http://127.0.0.1:" + port)
|
||||
.endpoint("/mcp")
|
||||
.requestBuilder(requestTemplate)
|
||||
.build();
|
||||
try (McpSyncClient client = McpClient.sync(transport).build()) {
|
||||
client.initialize();
|
||||
McpSchema.CallToolResult response = client.callTool(McpSchema.CallToolRequest.builder("fleet_list")
|
||||
.arguments(Map.of()).build());
|
||||
return ((McpSchema.TextContent) response.content().getFirst()).text();
|
||||
}
|
||||
}
|
||||
|
||||
@Test
|
||||
@Timeout(value = 15, unit = TimeUnit.SECONDS, threadMode = Timeout.ThreadMode.SEPARATE_THREAD)
|
||||
void fleetListReportsTheAssembledCapacitySource(@TempDir Path dir) throws Exception {
|
||||
String response = fleetList(assembleAndBind(dir));
|
||||
|
||||
assertTrue(response.contains("\"capacity\""), response);
|
||||
assertTrue(response.contains("\"" + PROFILE + "\""), response);
|
||||
assertTrue(response.contains("\"maxLoad\":7"), response);
|
||||
}
|
||||
|
||||
@Test
|
||||
@Timeout(value = 15, unit = TimeUnit.SECONDS, threadMode = Timeout.ThreadMode.SEPARATE_THREAD)
|
||||
void fleetListReportsTheAssembledHealthCoverageSource(@TempDir Path dir) throws Exception {
|
||||
String response = fleetList(assembleAndBind(dir));
|
||||
|
||||
assertTrue(response.contains("\"healthCoverage\":\"full\""), response);
|
||||
}
|
||||
|
||||
@Test
|
||||
@Timeout(value = 15, unit = TimeUnit.SECONDS, threadMode = Timeout.ThreadMode.SEPARATE_THREAD)
|
||||
void fleetListReportsTheAssembledCoordinatorPeers(@TempDir Path dir) throws Exception {
|
||||
String response = fleetList(assembleAndBind(dir));
|
||||
|
||||
assertTrue(response.contains("\"coordinator\""), response);
|
||||
assertTrue(response.contains("\"coordId\":\"" + PEER + "\""), response);
|
||||
}
|
||||
}
|
||||
+197
@@ -0,0 +1,197 @@
|
||||
package dev.ltms.fleet;
|
||||
|
||||
import dev.ltms.fleet.config.ConfigRef;
|
||||
import dev.ltms.fleet.config.FleetConfig;
|
||||
import dev.ltms.fleet.guard.SubscriptionGuard;
|
||||
import dev.ltms.fleet.herdr.FakeHerdr;
|
||||
import dev.ltms.fleet.herdr.HerdrClient;
|
||||
import dev.ltms.fleet.inject.ExhaustionSink;
|
||||
import dev.ltms.fleet.member.CompositePeerLauncher;
|
||||
import dev.ltms.fleet.member.HerdrPeerLauncher;
|
||||
import dev.ltms.fleet.peer.PeerLauncher;
|
||||
import dev.ltms.fleet.placement.BackendQuarantine;
|
||||
import io.javalin.Javalin;
|
||||
import org.junit.jupiter.api.DisplayName;
|
||||
import org.junit.jupiter.api.Test;
|
||||
import org.junit.jupiter.api.io.TempDir;
|
||||
|
||||
import java.lang.reflect.Field;
|
||||
import java.nio.file.Files;
|
||||
import java.nio.file.Path;
|
||||
import java.util.Map;
|
||||
import java.util.concurrent.Executors;
|
||||
import java.util.concurrent.ScheduledExecutorService;
|
||||
import java.util.concurrent.atomic.AtomicLong;
|
||||
import java.util.function.LongSupplier;
|
||||
|
||||
import static org.junit.jupiter.api.Assertions.assertEquals;
|
||||
import static org.junit.jupiter.api.Assertions.assertTrue;
|
||||
|
||||
/**
|
||||
* fleetd #612 step 4, rank 2 (OpenCode half) — {@link FleetdAssembly}'s {@code
|
||||
* forwardingExhaustionSink} line ({@code Fleetd.forwardingExhaustionSink(exhaustionSinkRef)}),
|
||||
* handed to {@link dev.ltms.fleet.member.OpenCodeLauncher} so its fleetd #175 model-mismatch check
|
||||
* can quarantine a credential before {@code sessions} exists to build the real sink (the
|
||||
* construction-order cycle documented at that call site). The ticket calls this independent from
|
||||
* {@code publishExhaustionSink} (pinned by {@link FleetdExhaustedPatternAssemblyTest}): a credential
|
||||
* that hits a usage limit through THIS path is never quarantined if {@code forwardingExhaustionSink}
|
||||
* is swapped for a hardcoded {@link ExhaustionSink#none()} at that call site — the OpenCode
|
||||
* launcher's own quarantine check keeps compiling and keeps "running", but it permanently talks to
|
||||
* a sink that does nothing, independent of whatever {@code publishExhaustionSink} does later.
|
||||
*
|
||||
* <p>{@code FleetdExhaustionSinkForwardingWiringTest} (fleetd #589) already proves {@code
|
||||
* Fleetd.forwardingExhaustionSink(ref)} forwards to whatever {@code ref} holds — as a bare factory
|
||||
* call, never through {@link FleetdAssembly#assembleAndStart}. It proves nothing about whether
|
||||
* THIS call site is the one FleetdAssembly actually wires into the real {@code OpenCodeLauncher}
|
||||
* it builds, which is exactly the #602/#606-shaped gap this ticket exists to close.
|
||||
*
|
||||
* <p>No accessor on {@link FleetdRuntime} reaches the adapter instances (by design — see that
|
||||
* class's own javadoc: only the final collaborators it owns directly are exposed), so this test
|
||||
* reaches the REAL, assembled {@code OpenCodeLauncher}'s {@code exhaustionSink} field the same way
|
||||
* {@code SessionManager}/{@code CompositePeerLauncher} wire it internally: a short, targeted
|
||||
* reflective walk ({@code SessionManager.launcher} → {@code CompositePeerLauncher.byProfile} →
|
||||
* {@code OpenCodeLauncher.exhaustionSink}) onto the exact object the assembly built — never a copy,
|
||||
* and never a read of the source text. Reflection is used the same way elsewhere in this suite
|
||||
* (e.g. {@code StatusPollerWatchdogTest}) to reach a private collaborator a production constructor
|
||||
* intentionally does not expose a public accessor for.
|
||||
*/
|
||||
class FleetdOpenCodeExhaustionForwardingAssemblyTest {
|
||||
|
||||
private static final class ControllableResourcePorts implements ResourcePorts {
|
||||
|
||||
final FakeHerdr herdr = new FakeHerdr();
|
||||
final AtomicLong nowNanos = new AtomicLong(1_000_000_000L);
|
||||
Runnable shutdownHook;
|
||||
|
||||
@Override
|
||||
public Map<String, String> environment() {
|
||||
return Map.of();
|
||||
}
|
||||
|
||||
@Override
|
||||
public HerdrClient connectHerdr(Path socketPath) {
|
||||
return herdr;
|
||||
}
|
||||
|
||||
@Override
|
||||
public Fleetd.AmqpOpener replyInboxOpener() {
|
||||
return (uri, prefetch) -> {
|
||||
throw new UnsupportedOperationException("replyInboxOpener must not be called — no broker: block");
|
||||
};
|
||||
}
|
||||
|
||||
@Override
|
||||
public Fleetd.LeadMailboxOpener leadMailboxOpener() {
|
||||
return (uri, selfCoordId, prefetch) -> {
|
||||
throw new UnsupportedOperationException("leadMailboxOpener must not be called — no coordinator: block");
|
||||
};
|
||||
}
|
||||
|
||||
@Override
|
||||
public Runnable herdrPollWait() {
|
||||
// Never invoked: this test's FakeHerdr answers immediately, so awaitHerdr never polls.
|
||||
return () -> {
|
||||
throw new UnsupportedOperationException("herdrPollWait must not be called — herdr is healthy");
|
||||
};
|
||||
}
|
||||
@Override
|
||||
public LongSupplier nanoClock() {
|
||||
return nowNanos::get;
|
||||
}
|
||||
|
||||
@Override
|
||||
public LongSupplier wallClockNanos() {
|
||||
return nowNanos::get;
|
||||
}
|
||||
|
||||
@Override
|
||||
public ScheduledExecutorService newScheduler(String purpose) {
|
||||
return Executors.newSingleThreadScheduledExecutor();
|
||||
}
|
||||
|
||||
@Override
|
||||
public void addShutdownHook(Runnable hook) {
|
||||
this.shutdownHook = hook;
|
||||
}
|
||||
|
||||
@Override
|
||||
public void startHttp(Javalin app, String host, int port) {
|
||||
// Deliberately never bind a real port.
|
||||
}
|
||||
}
|
||||
|
||||
private static FleetConfig writeConfig(Path dir, int cooldownSeconds) throws Exception {
|
||||
Path f = dir.resolve("fleetd.yaml");
|
||||
Files.writeString(f, """
|
||||
bind:
|
||||
host: 127.0.0.1
|
||||
port: 8765
|
||||
idleSleepGuard:
|
||||
enabled: false
|
||||
quarantineCooldownSeconds: %d
|
||||
profiles:
|
||||
gemini:
|
||||
kind: opencode
|
||||
model: google/gemini-2.5-pro
|
||||
""".formatted(cooldownSeconds));
|
||||
return FleetConfig.load(f);
|
||||
}
|
||||
|
||||
/** Reach a declared field by name on {@code target}'s runtime class, bypassing the access check. */
|
||||
private static Object readField(Object target, Class<?> declaringClass, String fieldName) throws Exception {
|
||||
Field field = declaringClass.getDeclaredField(fieldName);
|
||||
field.setAccessible(true);
|
||||
return field.get(target);
|
||||
}
|
||||
|
||||
@Test
|
||||
@DisplayName("[BEHAVIOURAL] the real assembled OpenCodeLauncher's exhaustionSink field forwards "
|
||||
+ "an onExhausted call into the real daemon's BackendQuarantine")
|
||||
void assembledOpenCodeLauncherExhaustionSinkQuarantinesTheCredential(@TempDir Path dir) throws Exception {
|
||||
int cooldownSeconds = 90;
|
||||
FleetConfig cfg = writeConfig(dir, cooldownSeconds);
|
||||
ConfigRef config = new ConfigRef(dir.resolve("fleetd.yaml"), cfg);
|
||||
SubscriptionGuard guard = new SubscriptionGuard(cfg.guard().hostSet());
|
||||
ControllableResourcePorts ports = new ControllableResourcePorts();
|
||||
|
||||
FleetdRuntime runtime = FleetdAssembly.assembleAndStart(new AssemblyInputs(cfg, config, guard), ports);
|
||||
try {
|
||||
PeerLauncher launcherField = (PeerLauncher) readField(runtime.sessions(),
|
||||
runtime.sessions().getClass(), "launcher");
|
||||
// CONTROL: the composite launcher must actually be the real production type with a
|
||||
// "gemini" -> OpenCodeLauncher entry — if this fails, nothing below exercised the real
|
||||
// assembly at all, rather than silently passing on an empty/wrong object.
|
||||
assertTrue(launcherField instanceof CompositePeerLauncher,
|
||||
"CONTROL: SessionManager.launcher must be the real CompositePeerLauncher the "
|
||||
+ "assembly built, got: " + launcherField);
|
||||
@SuppressWarnings("unchecked")
|
||||
Map<String, HerdrPeerLauncher> byProfile = (Map<String, HerdrPeerLauncher>)
|
||||
readField(launcherField, CompositePeerLauncher.class, "byProfile");
|
||||
HerdrPeerLauncher adapter = byProfile.get("gemini");
|
||||
assertTrue(adapter != null && adapter.getClass().getSimpleName().equals("OpenCodeLauncher"),
|
||||
"CONTROL: the 'gemini' profile must resolve to a real OpenCodeLauncher adapter, "
|
||||
+ "got: " + adapter);
|
||||
|
||||
ExhaustionSink sink = (ExhaustionSink) readField(adapter, adapter.getClass(), "exhaustionSink");
|
||||
assertTrue(sink != null, "CONTROL: OpenCodeLauncher.exhaustionSink must never be null");
|
||||
|
||||
// The exact call OpenCodeLauncher.SessionAwareHandle#checkModelMatch makes on a real
|
||||
// model mismatch (fleetd #175): target, reason, and its own already-known profile name.
|
||||
sink.onExhausted("term_gemini_1", "opencode model mismatch (test)", "gemini");
|
||||
|
||||
BackendQuarantine quarantine = runtime.mcp().quarantineSource().quarantine();
|
||||
assertTrue(quarantine.isQuarantined("gemini"),
|
||||
"the real forwardingExhaustionSink-wired field must have delegated into the "
|
||||
+ "published production sink, which quarantines the profile's credential "
|
||||
+ "('gemini' here — no credentialId configured) — replacing "
|
||||
+ "FleetdAssembly's forwardingExhaustionSink call site with a hardcoded "
|
||||
+ "ExhaustionSink.none() must fail this assertion, since the field read "
|
||||
+ "above would then BE the inert no-op and nothing would ever reach "
|
||||
+ "quarantine.quarantine(...)");
|
||||
assertEquals(cooldownSeconds, quarantine.remainingSeconds("gemini").orElseThrow(
|
||||
() -> new AssertionError("credential must report a remaining cooldown")));
|
||||
} finally {
|
||||
if (ports.shutdownHook != null) ports.shutdownHook.run();
|
||||
}
|
||||
}
|
||||
}
|
||||
+368
@@ -0,0 +1,368 @@
|
||||
package dev.ltms.fleet;
|
||||
|
||||
import dev.ltms.fleet.config.ConfigRef;
|
||||
import dev.ltms.fleet.config.FleetConfig;
|
||||
import dev.ltms.fleet.guard.SubscriptionGuard;
|
||||
import dev.ltms.fleet.herdr.FakeHerdr;
|
||||
import dev.ltms.fleet.herdr.HerdrClient;
|
||||
import dev.ltms.fleet.inject.CompletionResolver;
|
||||
import dev.ltms.fleet.msg.Rendezvous;
|
||||
import dev.ltms.fleet.msg.TurnToken;
|
||||
import dev.ltms.fleet.session.MemberSession;
|
||||
import io.javalin.Javalin;
|
||||
import io.modelcontextprotocol.client.McpClient;
|
||||
import io.modelcontextprotocol.client.McpSyncClient;
|
||||
import io.modelcontextprotocol.client.transport.HttpClientStreamableHttpTransport;
|
||||
import io.modelcontextprotocol.spec.McpClientTransport;
|
||||
import io.modelcontextprotocol.spec.McpSchema;
|
||||
import org.junit.jupiter.api.AfterEach;
|
||||
import org.junit.jupiter.api.Test;
|
||||
import org.junit.jupiter.api.Timeout;
|
||||
import org.junit.jupiter.api.io.TempDir;
|
||||
|
||||
import java.net.URI;
|
||||
import java.net.http.HttpClient;
|
||||
import java.net.http.HttpRequest;
|
||||
import java.net.http.HttpResponse;
|
||||
import java.nio.file.Files;
|
||||
import java.nio.file.Path;
|
||||
import java.util.Map;
|
||||
import java.util.concurrent.CompletableFuture;
|
||||
import java.util.concurrent.Executors;
|
||||
import java.util.concurrent.ScheduledExecutorService;
|
||||
import java.util.concurrent.TimeUnit;
|
||||
import java.util.concurrent.atomic.AtomicLong;
|
||||
import java.util.function.LongSupplier;
|
||||
|
||||
import static org.junit.jupiter.api.Assertions.assertEquals;
|
||||
import static org.junit.jupiter.api.Assertions.assertTrue;
|
||||
|
||||
/**
|
||||
* fleetd #612 Shape A, unit r4 — {@link FleetdAssembly}'s {@code quarantineSource} (lines 471-472)
|
||||
* and {@code outageSource} (lines 473-476 at {@code main} = {@code 141ae3b}), EACH of which feeds
|
||||
* two separate consumers: {@code FleetMcp} ({@code fleet_profiles}) and {@code FleetApp}
|
||||
* ({@code GET /profiles}), one call site ({@code :484}/{@code :486}) into the MCP constructor and
|
||||
* the SAME shared local again ({@code :529}) into the REST constructor.
|
||||
*
|
||||
* <p><strong>The property pinned here</strong> (from the ticket): when the real assembly has built
|
||||
* a real {@link dev.ltms.fleet.placement.BackendQuarantine} holding a quarantined credential, and a
|
||||
* real {@link dev.ltms.fleet.placement.BackendOutagePolicy} holding a cooling-off credential, BOTH
|
||||
* operator windows must report that state — the real assembled {@code FleetMcp} (reached through
|
||||
* {@link FleetdRuntime#mcp()}) and the real assembled REST surface (reached through {@link
|
||||
* FleetdRuntime#app()}). A mutation that starves one consumer while leaving the other wired must
|
||||
* make only that consumer's assertion go red.
|
||||
*
|
||||
* <p><strong>No source-text assertion anywhere in this file.</strong> Both windows are read off the
|
||||
* REAL running objects: {@code fleet_profiles} is called through a real MCP client over a real
|
||||
* HTTP connection to the servlet {@link FleetdAssembly} actually mounted, and {@code GET /profiles}
|
||||
* is called through a real {@link java.net.http.HttpClient} against the real bound {@link
|
||||
* FleetdRuntime#app()}. Neither is a copy built alongside the assembly for this test's benefit.
|
||||
*
|
||||
* <p><strong>How this gets past CB-185's own pid-resolution dead end.</strong> {@code
|
||||
* FleetdAssemblyFleetAppTest}'s class javadoc explains that {@code GET /sessions} cannot be driven
|
||||
* over real HTTP here because {@code LsofPeerPidLookup} excludes its own pid and an in-process test
|
||||
* client/server share one JVM pid — every such request resolves {@code ANONYMOUS} and is refused
|
||||
* before the handler runs. {@code GET /profiles} and {@code fleet_profiles} sit behind the exact
|
||||
* same {@code Authz.Action.READ} gate. This test sidesteps the dead end instead of hitting it:
|
||||
* {@code auth.mode: token} (see {@link dev.ltms.fleet.auth.CallerResolver#resolve}) resolves a
|
||||
* caller to {@code PRIMARY} from a valid {@code Authorization: Bearer} header ALONE, with no pid
|
||||
* resolution involved at all — the same technique {@code FleetMcpContextExtractorTest} already uses
|
||||
* to drive a real {@code fleet_whoami} call through the real transport.
|
||||
*
|
||||
* <p><strong>How the quarantined/cooling-off state is set up.</strong> Both {@code BackendQuarantine}
|
||||
* and {@code BackendOutagePolicy} are private to the collaborators the assembly wires them into, and
|
||||
* (unlike {@code quarantineSource()}) neither {@code FleetMcp} nor {@code FleetApp} exposes a public
|
||||
* accessor for the live {@code BackendOutagePolicy} instance. Rather than add one (acceptance
|
||||
* criterion 1: no production change), this test drives the REAL production classification path —
|
||||
* exactly the recipe {@code FleetdExhaustedPatternAssemblyTest} (quarantine) and {@code
|
||||
* FleetdCompletionResolverAssemblyTest} (cool-off) already proved works end to end against this same
|
||||
* {@link FleetdAssembly#assembleAndStart}: acquire a real {@link MemberSession}, feed the real {@link
|
||||
* CompletionResolver} a pane scrape matching the profile's configured {@code exhaustedPattern} /
|
||||
* {@code errorPattern}, and let the real {@code exhaustionSink}/{@code backendErrorSink} write into
|
||||
* the real, shared tracker. Each setup step asserts its own {@link Rendezvous.Kind} as a CONTROL —
|
||||
* if the resolver were never actually exercised, the setup itself fails loudly before either window
|
||||
* is ever read.
|
||||
*/
|
||||
class FleetdQuarantineOutageDualWindowAssemblyTest {
|
||||
|
||||
private static final String TOKEN = "s3cret-r4-token";
|
||||
private static final String TOKEN_ENV = "FLEETD_R4_TEST_TOKEN";
|
||||
|
||||
private static final class ControllableResourcePorts implements ResourcePorts {
|
||||
|
||||
final FakeHerdr herdr;
|
||||
final AtomicLong nowNanos = new AtomicLong(1_000_000_000L); // arbitrary non-zero start
|
||||
Runnable shutdownHook;
|
||||
|
||||
ControllableResourcePorts(FakeHerdr herdr) {
|
||||
this.herdr = herdr;
|
||||
}
|
||||
|
||||
void advanceSeconds(long seconds) {
|
||||
nowNanos.addAndGet(TimeUnit.SECONDS.toNanos(seconds));
|
||||
}
|
||||
|
||||
@Override
|
||||
public Map<String, String> environment() {
|
||||
return Map.of(TOKEN_ENV, TOKEN);
|
||||
}
|
||||
|
||||
@Override
|
||||
public HerdrClient connectHerdr(Path socketPath) {
|
||||
return herdr;
|
||||
}
|
||||
|
||||
@Override
|
||||
public Fleetd.AmqpOpener replyInboxOpener() {
|
||||
// Never invoked: this test's config has no `broker:` block.
|
||||
return (uri, prefetch) -> {
|
||||
throw new UnsupportedOperationException("replyInboxOpener must not be called — no broker: block");
|
||||
};
|
||||
}
|
||||
|
||||
@Override
|
||||
public Fleetd.LeadMailboxOpener leadMailboxOpener() {
|
||||
// Never invoked: this test's config has no `coordinator:` block.
|
||||
return (uri, selfCoordId, prefetch) -> {
|
||||
throw new UnsupportedOperationException(
|
||||
"leadMailboxOpener must not be called — no coordinator: block is configured");
|
||||
};
|
||||
}
|
||||
|
||||
@Override
|
||||
public LongSupplier nanoClock() {
|
||||
return nowNanos::get;
|
||||
}
|
||||
|
||||
@Override
|
||||
public LongSupplier wallClockNanos() {
|
||||
return nowNanos::get;
|
||||
}
|
||||
|
||||
@Override
|
||||
public ScheduledExecutorService newScheduler(String purpose) {
|
||||
return Executors.newSingleThreadScheduledExecutor();
|
||||
}
|
||||
|
||||
@Override
|
||||
public void addShutdownHook(Runnable hook) {
|
||||
this.shutdownHook = hook;
|
||||
}
|
||||
|
||||
@Override
|
||||
public void startHttp(Javalin app, String host, int port) {
|
||||
// Deliberately never bind here — this test binds runtime.app() itself, for real, below.
|
||||
}
|
||||
|
||||
@Override
|
||||
public Runnable herdrPollWait() {
|
||||
// Never invoked: this test's FakeHerdr answers immediately, so awaitHerdr never polls.
|
||||
return () -> {
|
||||
throw new UnsupportedOperationException("herdrPollWait must not be called — herdr is healthy");
|
||||
};
|
||||
}
|
||||
}
|
||||
|
||||
private FleetdRuntime runtime;
|
||||
private ControllableResourcePorts ports;
|
||||
private Javalin boundApp;
|
||||
|
||||
@AfterEach
|
||||
void tearDown() {
|
||||
if (boundApp != null) {
|
||||
boundApp.stop();
|
||||
}
|
||||
if (ports != null && ports.shutdownHook != null) {
|
||||
ports.shutdownHook.run();
|
||||
}
|
||||
}
|
||||
|
||||
private static FleetConfig writeConfig(Path dir, int cooldownSeconds) throws Exception {
|
||||
Path f = dir.resolve("fleetd.yaml");
|
||||
Files.writeString(f, """
|
||||
bind:
|
||||
host: 127.0.0.1
|
||||
port: 8765
|
||||
auth:
|
||||
mode: token
|
||||
tokenEnv: %s
|
||||
idleSleepGuard:
|
||||
enabled: false
|
||||
quarantineCooldownSeconds: %d
|
||||
profiles:
|
||||
exhaustprofile:
|
||||
baseUrl: http://exhausthost.local:8000
|
||||
model: sonnet
|
||||
exhaustedPattern: "usage limit reached"
|
||||
coolprofile:
|
||||
baseUrl: http://coolhost.local:8000
|
||||
model: sonnet
|
||||
errorPattern: "credential outage"
|
||||
guard:
|
||||
offSubscriptionHosts:
|
||||
- exhausthost.local
|
||||
- coolhost.local
|
||||
""".formatted(TOKEN_ENV, cooldownSeconds));
|
||||
return FleetConfig.load(f);
|
||||
}
|
||||
|
||||
/** Assembles the real graph, then binds the real {@code Javalin app} to an ephemeral port. */
|
||||
private int assembleAndBind(Path dir) throws Exception {
|
||||
FleetConfig cfg = writeConfig(dir, 120);
|
||||
ConfigRef config = new ConfigRef(dir.resolve("fleetd.yaml"), cfg);
|
||||
SubscriptionGuard guard = new SubscriptionGuard(cfg.guard().hostSet());
|
||||
ports = new ControllableResourcePorts(new FakeHerdr());
|
||||
runtime = FleetdAssembly.assembleAndStart(new AssemblyInputs(cfg, config, guard), ports);
|
||||
boundApp = runtime.app().start("127.0.0.1", 0);
|
||||
return boundApp.port();
|
||||
}
|
||||
|
||||
/**
|
||||
* Drives the real assembled {@link CompletionResolver} through a scrape matching {@code
|
||||
* exhaustprofile}'s configured {@code exhaustedPattern}, exactly {@code
|
||||
* FleetdExhaustedPatternAssemblyTest}'s own recipe, so the real {@code exhaustionSink} quarantines
|
||||
* the credential ({@code effectiveCredentialId() == "exhaustprofile"}, no explicit credentialId
|
||||
* configured).
|
||||
*/
|
||||
private void quarantineExhaustProfile(Path dir) {
|
||||
MemberSession session = runtime.sessions().acquire("exhaustprofile", null, dir.toString(), null);
|
||||
String target = session.terminalId();
|
||||
|
||||
CompletionResolver completion = runtime.completion();
|
||||
CompletableFuture<Rendezvous.Resolution> waiter = new CompletableFuture<>();
|
||||
|
||||
ports.herdr.readText("idle, nothing yet");
|
||||
completion.onDelivered(target, new TurnToken(target, waiter, null));
|
||||
ports.herdr.readText("usage limit reached: try again in a few hours");
|
||||
ports.advanceSeconds(3); // clear CompletionResolver.MIN_TURN_NANOS (2s), no real sleep
|
||||
completion.resolveBeforePostAction(target);
|
||||
|
||||
Rendezvous.Resolution resolution = waiter.getNow(null);
|
||||
assertTrue(resolution != null && resolution.kind() == Rendezvous.Kind.BACKEND_EXHAUSTED,
|
||||
"CONTROL: setup must classify as BACKEND_EXHAUSTED before either window is read — "
|
||||
+ "if this fails, the assembled resolver was never actually exercised: " + resolution);
|
||||
}
|
||||
|
||||
/**
|
||||
* Drives the real assembled {@link CompletionResolver} with TWO distinct targets on {@code
|
||||
* coolprofile}, each matching its configured {@code errorPattern}, exactly {@code
|
||||
* FleetdCompletionResolverAssemblyTest}'s own recipe, so the real {@code backendErrorSink} cools
|
||||
* the credential off ({@code effectiveCredentialId() == "coolprofile"}).
|
||||
*/
|
||||
private void coolOffCoolProfile(Path dir) {
|
||||
MemberSession s1 = runtime.sessions().acquire("coolprofile", null, dir.toString(), null);
|
||||
MemberSession s2 = runtime.sessions().acquire("coolprofile", null, dir.toString(), null);
|
||||
CompletionResolver completion = runtime.completion();
|
||||
|
||||
String t1 = s1.terminalId();
|
||||
CompletableFuture<Rendezvous.Resolution> w1 = new CompletableFuture<>();
|
||||
ports.herdr.readText("idle 1");
|
||||
completion.onDelivered(t1, new TurnToken(t1, w1, null));
|
||||
ports.herdr.readText("credential outage: upstream 503");
|
||||
ports.advanceSeconds(3);
|
||||
completion.resolveBeforePostAction(t1);
|
||||
Rendezvous.Resolution r1 = w1.getNow(null);
|
||||
assertTrue(r1 != null && r1.kind() == Rendezvous.Kind.FAILED,
|
||||
"CONTROL: target1's setup must classify FAILED (backend error): " + r1);
|
||||
|
||||
String t2 = s2.terminalId();
|
||||
CompletableFuture<Rendezvous.Resolution> w2 = new CompletableFuture<>();
|
||||
ports.herdr.readText("idle 2");
|
||||
completion.onDelivered(t2, new TurnToken(t2, w2, null));
|
||||
ports.herdr.readText("credential outage: upstream 503 again");
|
||||
ports.advanceSeconds(3);
|
||||
completion.resolveBeforePostAction(t2);
|
||||
Rendezvous.Resolution r2 = w2.getNow(null);
|
||||
assertTrue(r2 != null && r2.kind() == Rendezvous.Kind.FAILED,
|
||||
"CONTROL: target2's setup must classify FAILED (backend error) — two distinct "
|
||||
+ "targets are required to cross BackendOutagePolicy.THRESHOLD: " + r2);
|
||||
}
|
||||
|
||||
/** Calls the real {@code fleet_profiles} tool over a real MCP client, token-authenticated as PRIMARY. */
|
||||
private static McpSchema.CallToolResult callProfilesViaMcp(int port) {
|
||||
HttpRequest.Builder requestTemplate = HttpRequest.newBuilder().header("Authorization", "Bearer " + TOKEN);
|
||||
McpClientTransport transport = HttpClientStreamableHttpTransport.builder("http://127.0.0.1:" + port)
|
||||
.endpoint("/mcp")
|
||||
.requestBuilder(requestTemplate)
|
||||
.build();
|
||||
try (McpSyncClient client = McpClient.sync(transport).build()) {
|
||||
client.initialize();
|
||||
return client.callTool(McpSchema.CallToolRequest.builder("fleet_profiles").arguments(Map.of()).build());
|
||||
}
|
||||
}
|
||||
|
||||
private static String textOf(McpSchema.CallToolResult r) {
|
||||
return ((McpSchema.TextContent) r.content().getFirst()).text();
|
||||
}
|
||||
|
||||
/** Calls the real {@code GET /profiles} route over a real {@link HttpClient}, same token. */
|
||||
private static String getProfilesViaRest(int port) throws Exception {
|
||||
HttpClient http = HttpClient.newHttpClient();
|
||||
HttpRequest req = HttpRequest.newBuilder(URI.create("http://127.0.0.1:" + port + "/profiles"))
|
||||
.header("Authorization", "Bearer " + TOKEN)
|
||||
.GET().build();
|
||||
HttpResponse<String> res = http.send(req, HttpResponse.BodyHandlers.ofString());
|
||||
assertEquals(200, res.statusCode(), "GET /profiles must succeed with the real token: " + res.body());
|
||||
return res.body();
|
||||
}
|
||||
|
||||
// --- quarantineSource (FleetdAssembly.java :471-472) --------------------------------------
|
||||
|
||||
@Test
|
||||
@Timeout(value = 15, unit = TimeUnit.SECONDS, threadMode = Timeout.ThreadMode.SEPARATE_THREAD)
|
||||
void quarantinedCredentialIsReportedByTheRealAssembledFleetMcp(@TempDir Path dir) throws Exception {
|
||||
int port = assembleAndBind(dir);
|
||||
quarantineExhaustProfile(dir);
|
||||
|
||||
String out = textOf(callProfilesViaMcp(port));
|
||||
|
||||
assertTrue(out.contains("\"quarantined\""), "fleet_profiles must report a quarantined "
|
||||
+ "section once the real BackendQuarantine holds a quarantined credential: " + out);
|
||||
assertTrue(out.contains("\"exhaustprofile\""), out);
|
||||
assertTrue(out.contains("\"quarantinedForSeconds\""), out);
|
||||
}
|
||||
|
||||
@Test
|
||||
@Timeout(value = 15, unit = TimeUnit.SECONDS, threadMode = Timeout.ThreadMode.SEPARATE_THREAD)
|
||||
void quarantinedCredentialIsReportedByTheRealAssembledFleetApp(@TempDir Path dir) throws Exception {
|
||||
int port = assembleAndBind(dir);
|
||||
quarantineExhaustProfile(dir);
|
||||
|
||||
String out = getProfilesViaRest(port);
|
||||
|
||||
assertTrue(out.contains("\"quarantined\""), "GET /profiles must report a quarantined "
|
||||
+ "section once the real BackendQuarantine holds a quarantined credential: " + out);
|
||||
assertTrue(out.contains("\"exhaustprofile\""), out);
|
||||
assertTrue(out.contains("\"quarantinedForSeconds\""), out);
|
||||
}
|
||||
|
||||
// --- outageSource (FleetdAssembly.java :473-476) -------------------------------------------
|
||||
|
||||
@Test
|
||||
@Timeout(value = 15, unit = TimeUnit.SECONDS, threadMode = Timeout.ThreadMode.SEPARATE_THREAD)
|
||||
void coolingOffCredentialIsReportedByTheRealAssembledFleetMcp(@TempDir Path dir) throws Exception {
|
||||
int port = assembleAndBind(dir);
|
||||
coolOffCoolProfile(dir);
|
||||
|
||||
String out = textOf(callProfilesViaMcp(port));
|
||||
|
||||
assertTrue(out.contains("\"coolingOff\""), "fleet_profiles must report a coolingOff "
|
||||
+ "section once the real BackendOutagePolicy holds a cooling-off credential: " + out);
|
||||
assertTrue(out.contains("\"coolprofile\""), out);
|
||||
assertTrue(out.contains("\"coolingOffForSeconds\""), out);
|
||||
}
|
||||
|
||||
@Test
|
||||
@Timeout(value = 15, unit = TimeUnit.SECONDS, threadMode = Timeout.ThreadMode.SEPARATE_THREAD)
|
||||
void coolingOffCredentialIsReportedByTheRealAssembledFleetApp(@TempDir Path dir) throws Exception {
|
||||
int port = assembleAndBind(dir);
|
||||
coolOffCoolProfile(dir);
|
||||
|
||||
String out = getProfilesViaRest(port);
|
||||
|
||||
assertTrue(out.contains("\"coolingOff\""), "GET /profiles must report a coolingOff "
|
||||
+ "section once the real BackendOutagePolicy holds a cooling-off credential: " + out);
|
||||
assertTrue(out.contains("\"coolprofile\""), out);
|
||||
assertTrue(out.contains("\"coolingOffForSeconds\""), out);
|
||||
}
|
||||
}
|
||||
@@ -0,0 +1,202 @@
|
||||
package dev.ltms.fleet;
|
||||
|
||||
import dev.ltms.fleet.guard.GuardException;
|
||||
import dev.ltms.fleet.herdr.HerdrClient;
|
||||
import io.javalin.Javalin;
|
||||
import org.junit.jupiter.api.Test;
|
||||
import org.junit.jupiter.api.io.TempDir;
|
||||
|
||||
import java.nio.file.Files;
|
||||
import java.nio.file.Path;
|
||||
import java.util.Map;
|
||||
import java.util.concurrent.ScheduledExecutorService;
|
||||
import java.util.function.LongSupplier;
|
||||
|
||||
import static org.junit.jupiter.api.Assertions.assertThrows;
|
||||
import static org.junit.jupiter.api.Assertions.assertTrue;
|
||||
|
||||
/**
|
||||
* fleetd #625: pins {@link dev.ltms.fleet.guard.SubscriptionGuard#assertPrimaryClean}'s call site
|
||||
* in {@link Fleetd#main(String[])} — the ONE place it runs at startup, and the check behind the
|
||||
* bridge charter's invariant 1 (never let the primary carry {@code ANTHROPIC_BASE_URL}). Nothing
|
||||
* pinned it before this ticket: deleting {@code guard.assertPrimaryClean(...)} from {@code main}
|
||||
* left the full suite green, because the call site read the real process environment ({@code
|
||||
* System.getenv()}), which a test cannot taint from inside the JVM.
|
||||
*
|
||||
* <p>{@link Fleetd#main(String[], ResourcePorts)} (added by this ticket) is the literal production
|
||||
* sequence — not a copy of it — driven here with a {@link ResourcePorts} whose {@link
|
||||
* ResourcePorts#environment()} is a plain {@code Map} a test controls. The guard itself was
|
||||
* already pinned by {@code SubscriptionGuardTest}, directly, with a {@code Map} — that proves the
|
||||
* method's behaviour, not that {@code main} still calls it at the right point. This class pins the
|
||||
* call site and, separately, the ORDER: the guard must still run before {@code cfg.validateAll()}
|
||||
* and before {@code FleetdAssembly.assembleAndStart} touches a socket, a broker, or HTTP — not just
|
||||
* be present somewhere in {@code main}.
|
||||
*
|
||||
* <p>Presence alone is not enough (a fix that pins only presence trades an invisible deletion for
|
||||
* an invisible reordering), so each test below is built so that EITHER deleting the guard call OR
|
||||
* moving it later makes the <em>same</em> test fail — with a different exception type than the one
|
||||
* asserted, not a vacuous pass. See each test's own javadoc for how.
|
||||
*/
|
||||
class FleetdSubscriptionGuardOrderingTest {
|
||||
|
||||
private static final Map<String, String> TAINTED_ENV =
|
||||
Map.of("ANTHROPIC_BASE_URL", "http://tainted.example");
|
||||
private static final Map<String, String> CLEAN_ENV = Map.of("PATH", "/usr/bin");
|
||||
|
||||
/**
|
||||
* A {@link ResourcePorts} whose {@link #environment()} is fixed to whatever the test hands it,
|
||||
* and whose every other method refuses to be called at all. That refusal is the ordering pin:
|
||||
* if {@code main} ever reaches {@link FleetdAssembly#assembleAndStart} before the guard has had
|
||||
* a chance to throw, the very first thing the assembly does with {@code ports} is {@link
|
||||
* #connectHerdr} — so a test that expects {@link GuardException} and instead observes {@link
|
||||
* UnsupportedOperationException} has just caught the guard running too late (or not at all).
|
||||
*/
|
||||
private static final class FixedEnvPorts implements ResourcePorts {
|
||||
private final Map<String, String> env;
|
||||
|
||||
FixedEnvPorts(Map<String, String> env) {
|
||||
this.env = env;
|
||||
}
|
||||
|
||||
@Override
|
||||
public Map<String, String> environment() {
|
||||
return env;
|
||||
}
|
||||
|
||||
@Override
|
||||
public HerdrClient connectHerdr(Path socketPath) {
|
||||
throw new UnsupportedOperationException("connectHerdr must not be called before the guard runs");
|
||||
}
|
||||
|
||||
@Override
|
||||
public Fleetd.AmqpOpener replyInboxOpener() {
|
||||
throw new UnsupportedOperationException("replyInboxOpener must not be called before the guard runs");
|
||||
}
|
||||
|
||||
@Override
|
||||
public Fleetd.LeadMailboxOpener leadMailboxOpener() {
|
||||
throw new UnsupportedOperationException("leadMailboxOpener must not be called before the guard runs");
|
||||
}
|
||||
|
||||
@Override
|
||||
public LongSupplier nanoClock() {
|
||||
throw new UnsupportedOperationException("nanoClock must not be called before the guard runs");
|
||||
}
|
||||
|
||||
@Override
|
||||
public LongSupplier wallClockNanos() {
|
||||
throw new UnsupportedOperationException("wallClockNanos must not be called before the guard runs");
|
||||
}
|
||||
|
||||
@Override
|
||||
public ScheduledExecutorService newScheduler(String purpose) {
|
||||
throw new UnsupportedOperationException("newScheduler must not be called before the guard runs");
|
||||
}
|
||||
|
||||
@Override
|
||||
public void addShutdownHook(Runnable hook) {
|
||||
throw new UnsupportedOperationException("addShutdownHook must not be called before the guard runs");
|
||||
}
|
||||
|
||||
@Override
|
||||
public void startHttp(Javalin app, String host, int port) {
|
||||
throw new UnsupportedOperationException("startHttp must not be called before the guard runs");
|
||||
}
|
||||
|
||||
@Override
|
||||
public Runnable herdrPollWait() {
|
||||
throw new UnsupportedOperationException("herdrPollWait must not be called before the guard runs");
|
||||
}
|
||||
}
|
||||
|
||||
/**
|
||||
* Otherwise-invalid: {@code bind.host: 0.0.0.0} with no {@code auth.mode: token} fails {@code
|
||||
* cfg.validateAll()} (CB-501's auth-exposure check — the same fixture {@code
|
||||
* FleetdStartupValidationTest#mainRefusesANonLoopbackBindWithoutTokenMode} uses), with an
|
||||
* {@link IllegalStateException}. That is deliberate: it is what {@code main} would throw INSTEAD
|
||||
* of {@link GuardException} if the guard call were deleted, or moved to run after {@code
|
||||
* validateAll()} — a different, distinguishable exception type.
|
||||
*/
|
||||
private static Path invalidConfig(Path dir) throws Exception {
|
||||
Path f = dir.resolve("fleetd.yaml");
|
||||
Files.writeString(f, """
|
||||
bind:
|
||||
host: 0.0.0.0
|
||||
port: 8765
|
||||
""");
|
||||
return f;
|
||||
}
|
||||
|
||||
/** Passes {@code cfg.validateAll()} cleanly — nothing here trips any of its checks. */
|
||||
private static Path validConfig(Path dir) throws Exception {
|
||||
Path f = dir.resolve("fleetd.yaml");
|
||||
Files.writeString(f, """
|
||||
bind:
|
||||
host: 127.0.0.1
|
||||
port: 8765
|
||||
""");
|
||||
return f;
|
||||
}
|
||||
|
||||
/**
|
||||
* Pins the order against {@code cfg.validateAll()}. The environment is tainted and the config
|
||||
* is otherwise invalid (see {@link #invalidConfig}). If the guard runs first (the required
|
||||
* order), {@code main} throws {@link GuardException} before {@code validateAll()} is ever
|
||||
* reached. If the guard were deleted, or reordered to run after {@code validateAll()}, {@code
|
||||
* validateAll()} throws {@link IllegalStateException} instead and this assertion fails on the
|
||||
* wrong exception type.
|
||||
*/
|
||||
@Test
|
||||
void mainRefusesATaintedEnvironmentBeforeValidatingTheConfig(@TempDir Path dir) throws Exception {
|
||||
Path config = invalidConfig(dir);
|
||||
FixedEnvPorts ports = new FixedEnvPorts(TAINTED_ENV);
|
||||
|
||||
GuardException ex = assertThrows(GuardException.class,
|
||||
() -> Fleetd.main(new String[]{config.toString()}, ports));
|
||||
assertTrue(ex.getMessage().contains("tainted"),
|
||||
"expected the primary-taint message, got: " + ex.getMessage());
|
||||
}
|
||||
|
||||
/**
|
||||
* Pins the order against {@code FleetdAssembly.assembleAndStart}. The environment is tainted
|
||||
* and the config is otherwise VALID (see {@link #validConfig}), so {@code cfg.validateAll()}
|
||||
* passes silently and the next thing that could possibly run is the assembly's first socket
|
||||
* call. If the guard runs first (the required order), {@code main} throws {@link
|
||||
* GuardException} before assembly starts. If the guard were deleted, or reordered to run after
|
||||
* assembly begins touching {@code ports}, {@link FixedEnvPorts#connectHerdr} throws {@link
|
||||
* UnsupportedOperationException} instead and this assertion fails on the wrong exception type.
|
||||
*/
|
||||
@Test
|
||||
void mainRefusesATaintedEnvironmentBeforeAssemblyTouchesAnyPort(@TempDir Path dir) throws Exception {
|
||||
Path config = validConfig(dir);
|
||||
FixedEnvPorts ports = new FixedEnvPorts(TAINTED_ENV);
|
||||
|
||||
GuardException ex = assertThrows(GuardException.class,
|
||||
() -> Fleetd.main(new String[]{config.toString()}, ports));
|
||||
assertTrue(ex.getMessage().contains("tainted"),
|
||||
"expected the primary-taint message, got: " + ex.getMessage());
|
||||
}
|
||||
|
||||
/**
|
||||
* The CONTROL for the two tests above. Same otherwise-valid config, same {@link FixedEnvPorts}
|
||||
* whose every method but {@code environment()} refuses to be called — but a CLEAN environment.
|
||||
* Without this, a guard that always threw {@link GuardException} regardless of input (the
|
||||
* opposite bug — e.g. the check inverted) would make the two tests above pass for the wrong
|
||||
* reason: not because they actually drove a real taint through a real guard, but because
|
||||
* anything would have thrown {@code GuardException}. Here, with nothing to taint, the guard
|
||||
* must let {@code main} proceed into {@code cfg.validateAll()} and on into the real assembly,
|
||||
* which reaches {@code ports.connectHerdr} — and THAT throws. A loud, positive assertion: if
|
||||
* the boot path never actually ran this far, there is no {@link UnsupportedOperationException}
|
||||
* to catch, only a quiet, unexpected hang or an unrelated early failure.
|
||||
*/
|
||||
@Test
|
||||
void mainProceedsPastTheGuardOnACleanEnvironment(@TempDir Path dir) throws Exception {
|
||||
Path config = validConfig(dir);
|
||||
FixedEnvPorts ports = new FixedEnvPorts(CLEAN_ENV);
|
||||
|
||||
UnsupportedOperationException ex = assertThrows(UnsupportedOperationException.class,
|
||||
() -> Fleetd.main(new String[]{config.toString()}, ports));
|
||||
assertTrue(ex.getMessage().contains("connectHerdr"),
|
||||
"expected forward progress to reach the assembly's first port call, got: " + ex.getMessage());
|
||||
}
|
||||
}
|
||||
@@ -38,6 +38,37 @@ class AuthzTest {
|
||||
}
|
||||
}
|
||||
|
||||
/**
|
||||
* {@code fleet_send} is three call shapes behind one action name until {@code
|
||||
* FleetMcp#sendAction} picks one: a plain local {@link Authz.Action#SEND}, the {@code coordId}
|
||||
* route ({@link Authz.Action#COORD_SEND}), and the {@code turnId} answer form ({@link
|
||||
* Authz.Action#ANSWER}). All three carry the same grant as the undivided action did — a worker
|
||||
* is excluded from every one, exactly as it was excluded from the one combined action before.
|
||||
*/
|
||||
@Test
|
||||
void theThreeSendShapesCarryTheSameGrantAsTheOldUndividedAction() {
|
||||
for (Authz.Action a : new Authz.Action[]{SEND, COORD_SEND, ANSWER}) {
|
||||
assertTrue(Authz.permits(PRIMARY, a, "term_a"), "the primary may " + a);
|
||||
assertTrue(Authz.permits(ARCH_DESIGN, a, "term_a"), "an architect may " + a);
|
||||
assertFalse(Authz.permits(WORKER_A, a, "term_a"),
|
||||
"a worker performing " + a + " would be escalating into the orchestrator role");
|
||||
assertFalse(Authz.permits(ANON, a, "term_a"));
|
||||
}
|
||||
}
|
||||
|
||||
/**
|
||||
* {@code fleet_poll{ticket}} and {@code fleet_status} are {@link Authz.Action#TASK_READ}, split
|
||||
* out of the roster-only {@link Authz.Action#READ} (fleetd #678). The grant is unchanged from
|
||||
* what the undivided {@code READ} action gave every one of these callers.
|
||||
*/
|
||||
@Test
|
||||
void taskReadCarriesTheSameGrantReadDidBeforeTheSplit() {
|
||||
assertTrue(Authz.permits(PRIMARY, TASK_READ, null));
|
||||
assertTrue(Authz.permits(WORKER_A, TASK_READ, null));
|
||||
assertTrue(Authz.permits(ARCH_DESIGN, TASK_READ, null));
|
||||
assertFalse(Authz.permits(ANON, TASK_READ, null));
|
||||
}
|
||||
|
||||
@Test
|
||||
void aWorkerMayReplyAndAskOnlyAsItself() {
|
||||
assertTrue(Authz.permits(WORKER_A, REPLY, "term_a"));
|
||||
|
||||
+65
@@ -0,0 +1,65 @@
|
||||
package dev.ltms.fleet.config;
|
||||
|
||||
import org.junit.jupiter.api.DisplayName;
|
||||
import org.junit.jupiter.api.Test;
|
||||
|
||||
import java.util.Map;
|
||||
|
||||
import static org.junit.jupiter.api.Assertions.assertEquals;
|
||||
import static org.junit.jupiter.api.Assertions.assertNull;
|
||||
|
||||
/**
|
||||
* {@link FleetConfig.Profile#effectiveAutoCompactWindow()} resolves the window a launched Claude
|
||||
* Code session actually runs on, not just the {@code autoCompactWindow:} launch flag — {@code env:
|
||||
* CLAUDE_CODE_AUTO_COMPACT_WINDOW} overrides that flag, so a profile setting both resolves from the
|
||||
* environment variable.
|
||||
*/
|
||||
class FleetConfigProfileEffectiveAutoCompactWindowTest {
|
||||
|
||||
private static FleetConfig.Profile profile(Integer autoCompactWindow, Map<String, String> env) {
|
||||
return new FleetConfig.Profile("sonnet", null, "claude-sonnet-5", null, null, null,
|
||||
"tab", "fleet", "w #{n}", null, null, null, null, null, null, env,
|
||||
null, null, true, null, null, null, null, null, autoCompactWindow, null);
|
||||
}
|
||||
|
||||
@Test
|
||||
@DisplayName("yaml and env disagree: the env var wins, not the yaml key")
|
||||
void yamlAndEnvDisagreeEnvWins() {
|
||||
FleetConfig.Profile p = profile(250_000, Map.of("CLAUDE_CODE_AUTO_COMPACT_WINDOW", "150000"));
|
||||
|
||||
assertEquals(150_000, p.effectiveAutoCompactWindow(),
|
||||
"the two inputs must give DIFFERENT thresholds (150,000 vs 250,000) and the env value must win");
|
||||
}
|
||||
|
||||
@Test
|
||||
@DisplayName("only autoCompactWindow set: that value resolves")
|
||||
void onlyAutoCompactWindowSetResolvesToIt() {
|
||||
FleetConfig.Profile p = profile(250_000, Map.of());
|
||||
|
||||
assertEquals(250_000, p.effectiveAutoCompactWindow());
|
||||
}
|
||||
|
||||
@Test
|
||||
@DisplayName("only the env var set: that value resolves")
|
||||
void onlyEnvVarSetResolvesToIt() {
|
||||
FleetConfig.Profile p = profile(null, Map.of("CLAUDE_CODE_AUTO_COMPACT_WINDOW", "150000"));
|
||||
|
||||
assertEquals(150_000, p.effectiveAutoCompactWindow());
|
||||
}
|
||||
|
||||
@Test
|
||||
@DisplayName("neither set: resolves to null")
|
||||
void neitherSetResolvesToNull() {
|
||||
FleetConfig.Profile p = profile(null, Map.of());
|
||||
|
||||
assertNull(p.effectiveAutoCompactWindow());
|
||||
}
|
||||
|
||||
@Test
|
||||
@DisplayName("an unparseable env value falls back to autoCompactWindow rather than throwing")
|
||||
void unparseableEnvValueFallsBackToAutoCompactWindow() {
|
||||
FleetConfig.Profile p = profile(250_000, Map.of("CLAUDE_CODE_AUTO_COMPACT_WINDOW", "not-a-number"));
|
||||
|
||||
assertEquals(250_000, p.effectiveAutoCompactWindow());
|
||||
}
|
||||
}
|
||||
@@ -676,12 +676,8 @@ class FleetConfigTest {
|
||||
assertEquals(5, hb.quietNudgeCap());
|
||||
}
|
||||
|
||||
/**
|
||||
* The hazard the guard exists for: fleetd writes worker tab labels and reads lead tab labels.
|
||||
* Overlap the two and every worker it spawns is read back as a lead.
|
||||
*/
|
||||
@Test
|
||||
void aLeadPrefixThatAProfileTabLabelOverrideAlsoMatchesRefusesToStart(@TempDir Path dir)
|
||||
void aProfileTabLabelOverrideMatchingALeadTabRefusesToStart(@TempDir Path dir)
|
||||
throws Exception {
|
||||
Path f = dir.resolve("collide.yaml");
|
||||
Files.writeString(f, """
|
||||
@@ -689,32 +685,33 @@ class FleetConfigTest {
|
||||
port: 8080
|
||||
profiles:
|
||||
gx10:
|
||||
tabLabel: "lead: {profile} #{n}"
|
||||
tabLabel: "alpha"
|
||||
fleet:
|
||||
leaders:
|
||||
opus:
|
||||
tab: "lead: opus"
|
||||
tabPrefix: "lead:"
|
||||
tab: "alpha"
|
||||
""");
|
||||
FleetConfig cfg = FleetConfig.load(f);
|
||||
|
||||
IllegalStateException e =
|
||||
assertThrows(IllegalStateException.class, cfg::validateLeadTabPrefixes);
|
||||
assertTrue(e.getMessage().contains("gx10"), "the message must name the offending profile");
|
||||
assertTrue(e.getMessage().contains("alpha"), "the message must name the offending label");
|
||||
}
|
||||
|
||||
/** A bad fleet-wide template promotes every member, not one profile — so it is checked too. */
|
||||
@Test
|
||||
void aFleetTabLabelThatMatchesALeadPrefixRefusesToStart(@TempDir Path dir) throws Exception {
|
||||
void aFleetTabLabelTemplateThatCanRenderAsALeadTabRefusesToStart(@TempDir Path dir) throws Exception {
|
||||
Path f = dir.resolve("collide-template.yaml");
|
||||
Files.writeString(f, """
|
||||
bind:
|
||||
port: 8080
|
||||
profiles:
|
||||
pha: {}
|
||||
fleet:
|
||||
tabLabel: "lead: {role} {profile}"
|
||||
tabLabel: "al{profile}"
|
||||
leaders:
|
||||
opus:
|
||||
tab: "lead: opus"
|
||||
tab: "alpha"
|
||||
""");
|
||||
FleetConfig cfg = FleetConfig.load(f);
|
||||
|
||||
@@ -723,12 +720,28 @@ class FleetConfigTest {
|
||||
assertTrue(e.getMessage().contains("fleet.tabLabel"));
|
||||
}
|
||||
|
||||
/**
|
||||
* The point of making role the label's first field: {@code {role}} comes from a closed enum, so
|
||||
* a generated label cannot begin with {@code "lead:"} however the fleet is configured.
|
||||
*/
|
||||
@Test
|
||||
void theDefaultTabLabelCannotCollideWithTheDefaultLeadPrefix(@TempDir Path dir) throws Exception {
|
||||
void anExactFleetTabLabelCollisionRefusesToStart(@TempDir Path dir) throws Exception {
|
||||
Path f = dir.resolve("exact-tab-collision.yaml");
|
||||
Files.writeString(f, """
|
||||
bind:
|
||||
port: 8080
|
||||
fleet:
|
||||
tabLabel: "alpha"
|
||||
leaders:
|
||||
alpha:
|
||||
tab: "alpha"
|
||||
""");
|
||||
|
||||
IllegalStateException e = assertThrows(IllegalStateException.class,
|
||||
() -> FleetConfig.load(f).validateAll());
|
||||
assertTrue(e.getMessage().contains("fleet.tabLabel"),
|
||||
"the message must name the offending label");
|
||||
assertTrue(e.getMessage().contains("alpha"), "the message must name the colliding lead tab");
|
||||
}
|
||||
|
||||
@Test
|
||||
void aFleetTabLabelTemplateThatCannotRenderAsALeadTabIsAllowed(@TempDir Path dir) throws Exception {
|
||||
Path f = dir.resolve("ok.yaml");
|
||||
Files.writeString(f, """
|
||||
bind:
|
||||
@@ -737,17 +750,13 @@ class FleetConfigTest {
|
||||
gx10:
|
||||
baseUrl: http://gx00.gw:8000
|
||||
fleet:
|
||||
tabLabel: "worker-{profile}"
|
||||
leaders:
|
||||
opus:
|
||||
tab: "lead: opus"
|
||||
tab: "alpha"
|
||||
""");
|
||||
|
||||
assertDoesNotThrow(() -> FleetConfig.load(f).validateLeadTabPrefixes());
|
||||
for (MemberRole role : MemberRole.values()) {
|
||||
assertFalse(FleetConfig.Fleet.DEFAULT_TAB_LABEL
|
||||
.replace("{role}", role.wireName()).startsWith("lead:"),
|
||||
"no role renders a label that reads as a lead");
|
||||
}
|
||||
assertDoesNotThrow(() -> FleetConfig.load(f).validateAll());
|
||||
}
|
||||
|
||||
@Test
|
||||
@@ -765,6 +774,164 @@ class FleetConfigTest {
|
||||
"a label that collides with a convention nobody reads is not a problem");
|
||||
}
|
||||
|
||||
/**
|
||||
* fleetd #677: identity is matched on a lead's exact {@code tab} alone, so two leads sharing
|
||||
* one tab means only one of them is ever found — the guard must catch this independently of
|
||||
* the member-template checks above.
|
||||
*/
|
||||
@Test
|
||||
void twoLeadsSharingTheSameExactTabRefusesToStart(@TempDir Path dir) throws Exception {
|
||||
Path f = dir.resolve("shared-tab.yaml");
|
||||
Files.writeString(f, """
|
||||
bind:
|
||||
port: 8080
|
||||
fleet:
|
||||
leaders:
|
||||
opus:
|
||||
tab: "shared tab"
|
||||
sonnet:
|
||||
tab: "shared tab"
|
||||
""");
|
||||
FleetConfig cfg = FleetConfig.load(f);
|
||||
|
||||
IllegalStateException e =
|
||||
assertThrows(IllegalStateException.class, cfg::validateLeadTabPrefixes);
|
||||
assertTrue(e.getMessage().contains("opus"), "the message must name one offending lead");
|
||||
assertTrue(e.getMessage().contains("sonnet"), "the message must name the other offending lead");
|
||||
}
|
||||
|
||||
/**
|
||||
* fleetd #693: the guard matches tabs case-insensitively, because
|
||||
* {@code LeadTabScanner} keys its tab map on a lowercased label — two tabs differing only in
|
||||
* case collide there too, and the guard must catch that independently of the exact-match case
|
||||
* above.
|
||||
*/
|
||||
@Test
|
||||
void twoLeadsSharingTheSameTabInDifferentCaseRefusesToStart(@TempDir Path dir) throws Exception {
|
||||
Path f = dir.resolve("shared-tab-case.yaml");
|
||||
Files.writeString(f, """
|
||||
bind:
|
||||
port: 8080
|
||||
fleet:
|
||||
leaders:
|
||||
opus:
|
||||
tab: "Shared Tab"
|
||||
sonnet:
|
||||
tab: "shared tab"
|
||||
""");
|
||||
FleetConfig cfg = FleetConfig.load(f);
|
||||
|
||||
IllegalStateException e =
|
||||
assertThrows(IllegalStateException.class, cfg::validateLeadTabPrefixes);
|
||||
assertTrue(e.getMessage().contains("opus"), "the message must name one offending lead");
|
||||
assertTrue(e.getMessage().contains("sonnet"), "the message must name the other offending lead");
|
||||
}
|
||||
|
||||
/** Control for {@link #twoLeadsSharingTheSameExactTabRefusesToStart}: distinct tabs load cleanly. */
|
||||
@Test
|
||||
void twoLeadsWithDistinctExactTabsAreAllowed(@TempDir Path dir) throws Exception {
|
||||
Path f = dir.resolve("distinct-tabs.yaml");
|
||||
Files.writeString(f, """
|
||||
bind:
|
||||
port: 8080
|
||||
fleet:
|
||||
leaders:
|
||||
opus:
|
||||
tab: "opus tab"
|
||||
sonnet:
|
||||
tab: "sonnet tab"
|
||||
""");
|
||||
FleetConfig cfg = FleetConfig.load(f);
|
||||
|
||||
assertDoesNotThrow(cfg::validateLeadTabPrefixes);
|
||||
}
|
||||
|
||||
// ── validatePanePlacementAgainstLeadTabs ────────────────────────────────────────────────────
|
||||
|
||||
/**
|
||||
* The hazard this guard closes: a pane-placed member lands inside the focused tab rather than
|
||||
* its own, so it can land inside a lead's labelled tab and be read back as that lead.
|
||||
*/
|
||||
@Test
|
||||
void aPanePlacedProfileWithALeadTabRefusesToStart(@TempDir Path dir) throws Exception {
|
||||
Path f = dir.resolve("pane-hazard.yaml");
|
||||
Files.writeString(f, """
|
||||
bind:
|
||||
port: 8080
|
||||
profiles:
|
||||
gx10:
|
||||
placement: pane
|
||||
fleet:
|
||||
leaders:
|
||||
opus:
|
||||
tab: "lead: opus"
|
||||
""");
|
||||
FleetConfig cfg = FleetConfig.load(f);
|
||||
|
||||
IllegalStateException e = assertThrows(IllegalStateException.class,
|
||||
cfg::validatePanePlacementAgainstLeadTabs);
|
||||
assertTrue(e.getMessage().contains("gx10"), "the message must name the offending profile");
|
||||
}
|
||||
|
||||
@Test
|
||||
void aPanePlacedProfileWithNoLeadTabIsAllowed(@TempDir Path dir) throws Exception {
|
||||
Path f = dir.resolve("pane-no-tab.yaml");
|
||||
Files.writeString(f, """
|
||||
bind:
|
||||
port: 8080
|
||||
profiles:
|
||||
gx10:
|
||||
placement: pane
|
||||
fleet:
|
||||
leaders:
|
||||
opus:
|
||||
profile: gx10
|
||||
""");
|
||||
|
||||
assertDoesNotThrow(() -> FleetConfig.load(f).validatePanePlacementAgainstLeadTabs(),
|
||||
"a leader with no tab feeds nothing into the scanner, so pane placement is safe");
|
||||
}
|
||||
|
||||
@Test
|
||||
void aTabPlacedProfileWithALeadTabIsAllowed(@TempDir Path dir) throws Exception {
|
||||
Path f = dir.resolve("tab-safe.yaml");
|
||||
Files.writeString(f, """
|
||||
bind:
|
||||
port: 8080
|
||||
profiles:
|
||||
gx10:
|
||||
placement: tab
|
||||
fleet:
|
||||
leaders:
|
||||
opus:
|
||||
tab: "lead: opus"
|
||||
""");
|
||||
|
||||
assertDoesNotThrow(() -> FleetConfig.load(f).validatePanePlacementAgainstLeadTabs(),
|
||||
"a member in its own tab cannot land inside a lead's tab");
|
||||
}
|
||||
|
||||
/** Proves the reflective sweep behind {@code validateAll} really reaches this validator. */
|
||||
@Test
|
||||
void validateAllAlsoRefusesPanePlacementAgainstALeadTab(@TempDir Path dir) throws Exception {
|
||||
Path f = dir.resolve("pane-hazard-sweep.yaml");
|
||||
Files.writeString(f, """
|
||||
bind:
|
||||
port: 8080
|
||||
profiles:
|
||||
gx10:
|
||||
placement: pane
|
||||
fleet:
|
||||
leaders:
|
||||
opus:
|
||||
tab: "lead: opus"
|
||||
""");
|
||||
FleetConfig cfg = FleetConfig.load(f);
|
||||
|
||||
IllegalStateException e = assertThrows(IllegalStateException.class, cfg::validateAll);
|
||||
assertTrue(e.getMessage().contains("gx10"), "the message must name the offending profile");
|
||||
}
|
||||
|
||||
// ── CB-530/CB-579: the leaders registry ─────────────────────────────────────────────────────
|
||||
|
||||
@Test
|
||||
@@ -3089,4 +3256,41 @@ class FleetConfigTest {
|
||||
FleetConfig cfg = FleetConfig.load(f);
|
||||
assertTrue(cfg.models().offIds().isEmpty());
|
||||
}
|
||||
|
||||
// ── fleetd #651: leadRollover.turnSettleSeconds default resolution ─────────────────────────
|
||||
|
||||
@Test
|
||||
void turnSettleSecondsDefaultsTo300WhenUnset(@TempDir Path dir) throws Exception {
|
||||
Path f = dir.resolve("bare-rollover.yaml");
|
||||
Files.writeString(f, "bind:\n port: 8080\nleadRollover: {}\n");
|
||||
|
||||
FleetConfig.LeadRollover rollover = FleetConfig.load(f).leadRollover();
|
||||
assertNotNull(rollover);
|
||||
assertEquals(300, rollover.turnSettleSeconds());
|
||||
}
|
||||
|
||||
@Test
|
||||
void turnSettleSecondsUsesAnExplicitPositiveValue(@TempDir Path dir) throws Exception {
|
||||
Path f = dir.resolve("rollover.yaml");
|
||||
Files.writeString(f, """
|
||||
bind:
|
||||
port: 8080
|
||||
leadRollover:
|
||||
turnSettleSeconds: 45
|
||||
""");
|
||||
|
||||
FleetConfig.LeadRollover rollover = FleetConfig.load(f).leadRollover();
|
||||
assertEquals(45, rollover.turnSettleSeconds());
|
||||
}
|
||||
|
||||
@Test
|
||||
void turnSettleSecondsFallsBackTo300WhenZeroOrNegative(@TempDir Path dir) throws Exception {
|
||||
Path zero = dir.resolve("zero.yaml");
|
||||
Files.writeString(zero, "bind:\n port: 8080\nleadRollover:\n turnSettleSeconds: 0\n");
|
||||
assertEquals(300, FleetConfig.load(zero).leadRollover().turnSettleSeconds());
|
||||
|
||||
Path negative = dir.resolve("negative.yaml");
|
||||
Files.writeString(negative, "bind:\n port: 8080\nleadRollover:\n turnSettleSeconds: -5\n");
|
||||
assertEquals(300, FleetConfig.load(negative).leadRollover().turnSettleSeconds());
|
||||
}
|
||||
}
|
||||
|
||||
@@ -17,62 +17,21 @@ import static org.junit.jupiter.api.Assertions.assertThrows;
|
||||
import static org.junit.jupiter.api.Assertions.assertTrue;
|
||||
|
||||
/**
|
||||
* The gap this class exists to close: mutation testing on the fleetd ticket "central allow-list
|
||||
* of usable models" found that although {@link FleetConfig#validateModels()}'s own logic was well
|
||||
* pinned, nothing proved either real caller ({@code Fleetd.main} and {@link ConfigRef#reload()})
|
||||
* still invoked it — deleting the call site left the full suite green (1478/0/0/0). A follow-up
|
||||
* measurement (same technique — remove one call site, run the suite, not read the code) found the
|
||||
* SAME gap for all five of {@link FleetConfig}'s other validators at startup, and for four of the
|
||||
* six inside {@link ConfigRef#reload()}. This is a class of gap, not one line's mistake: every one
|
||||
* of those thirteen tests called the validator itself directly, never the real caller that was
|
||||
* supposed to.
|
||||
* Tests the reflective validator sweep and {@link FleetConfig#validateAll()} reachability.
|
||||
*
|
||||
* <p>The fix replaces the six individual {@code cfg.validateXxx()} calls at each of the two real
|
||||
* call sites with one {@link FleetConfig#validateAll()}, which reaches every validator by
|
||||
* reflection rather than by a hand-maintained list of names. A hand-maintained list of six names
|
||||
* would have exactly the defect it replaces: the seventh validator someone adds next month has no
|
||||
* reason to be added to it, and nothing would say so. This class proves TWO separate claims, and
|
||||
* keeps them separate on purpose:
|
||||
*
|
||||
* <ol>
|
||||
* <li>{@link #theSweepMechanismIsGenericNotHardcodedToFleetConfigsSixNames()} and its neighbours
|
||||
* prove the reflective sweep itself ({@link FleetConfig#invokeAllValidators}) is a general
|
||||
* mechanism — it runs whatever public, no-arg, void {@code validateXxx()} methods a class
|
||||
* happens to declare today, including a class with more of them than {@link FleetConfig}
|
||||
* has right now. This is the proof that a future, real seventh validator on {@link
|
||||
* FleetConfig} would be swept automatically, without needing to add a real (unwanted)
|
||||
* seventh validator just to exercise the claim.</li>
|
||||
* <li>{@link #validateAllReachesEveryOneOfTodaysSixValidators()} proves {@link
|
||||
* FleetConfig#validateAll()} itself is wired to that same generic mechanism and genuinely
|
||||
* reaches each of today's six real validators — reusing the exact minimal failing
|
||||
* configurations {@code FleetConfigTest} already established for each one directly, so a
|
||||
* single call to {@code validateAll()} is shown to reproduce every one of those six
|
||||
* failures.</li>
|
||||
* </ol>
|
||||
*
|
||||
* <p>Together with the direct-{@code Fleetd.main}-invocation tests in {@code
|
||||
* FleetdStartupValidationTest} (which prove the real startup call site still calls {@code
|
||||
* validateAll()}) and the {@code ConfigRefTest} reload tests (which prove the same for {@link
|
||||
* ConfigRef#reload()}), removing {@code cfg.validateAll();} from either real call site now fails
|
||||
* a test in this module.
|
||||
*
|
||||
* <p><b>What is NOT pinned, measured rather than assumed.</b> Reverting {@link
|
||||
* FleetConfig#validateAll()} to a hardcoded list of today's six method calls leaves the whole
|
||||
* suite green (measured at review: 1491 tests, 0 failures). Nothing ties {@code validateAll()} to
|
||||
* the generic sweep — claim 1 proves {@link FleetConfig#invokeAllValidators} is generic, and claim
|
||||
* 2 proves {@code validateAll()} reaches today's six, and a hardcoded list satisfies both. So the
|
||||
* reflective sweep is a convenience, not the guarantee. The guarantee is {@link
|
||||
* #fleetConfigDeclaresExactlyTheseSixValidatorsToday()}: it fails the moment a seventh validator
|
||||
* is declared, which forces whoever adds it to look at this file.
|
||||
* <p>{@link #theSweepRunsEveryValidateMethodOnAnUnrelatedClass()} and its neighbours
|
||||
* prove that {@link FleetConfig#invokeAllValidators} runs each public, no-arg, void
|
||||
* {@code validateXxx()} method on its target. {@link #fleetConfigDeclaresExactlyTheseValidatorsToday()}
|
||||
* is the canary for the validator set. {@link #validateAllReachesEveryOneOfTodaysRealValidators()}
|
||||
* is the reachability check for that set.
|
||||
*/
|
||||
class FleetConfigValidateAllTest {
|
||||
|
||||
// ── Claim 1: the reflective sweep is a general mechanism, not six names in disguise ──────────
|
||||
// ── Claim 1: the reflective sweep is a general mechanism ─────────────────────────────────────
|
||||
|
||||
/**
|
||||
* A throwaway fixture class, unrelated to {@link FleetConfig} in every way except shape: three
|
||||
* public, no-arg, void methods named {@code validateXxx}. Proves the sweep works on ANY class
|
||||
* with this shape, not on something special-cased to {@link FleetConfig}.
|
||||
* Fixture with public, no-arg, void methods named {@code validateXxx}. It proves the sweep uses
|
||||
* the target's method shape rather than special handling for {@link FleetConfig}.
|
||||
*/
|
||||
static class ThreeValidators {
|
||||
final List<String> ran = new ArrayList<>();
|
||||
@@ -91,7 +50,7 @@ class FleetConfigValidateAllTest {
|
||||
}
|
||||
|
||||
@Test
|
||||
void theSweepMechanismIsGenericNotHardcodedToFleetConfigsSixNames() {
|
||||
void theSweepRunsEveryValidateMethodOnAnUnrelatedClass() {
|
||||
ThreeValidators target = new ThreeValidators();
|
||||
FleetConfig.invokeAllValidators(target);
|
||||
assertEquals(List.of("validateAlpha", "validateBeta", "validateGamma"), target.ran,
|
||||
@@ -101,12 +60,8 @@ class FleetConfigValidateAllTest {
|
||||
}
|
||||
|
||||
/**
|
||||
* The core of the "self-maintaining" requirement: the exact same class shape as {@link
|
||||
* ThreeValidators}, plus one more method — standing in for "a developer adds a validator next
|
||||
* month". Nothing about the sweep changes to pick it up; the new method is invoked purely
|
||||
* because it exists and matches the shape. This is what makes adding a seventh real validator
|
||||
* to {@link FleetConfig} safe without touching {@link FleetConfig#validateAll()} or either
|
||||
* call site — there is no "wire it in" step left to forget.
|
||||
* Fixture with an added valid method. It proves the sweep reaches a method because it matches
|
||||
* the validator shape.
|
||||
*/
|
||||
static class FourValidators {
|
||||
final List<String> ran = new ArrayList<>();
|
||||
@@ -134,7 +89,7 @@ class FleetConfigValidateAllTest {
|
||||
FleetConfig.invokeAllValidators(target);
|
||||
assertEquals(List.of("validateAlpha", "validateBeta", "validateDelta", "validateGamma"),
|
||||
sorted(target.ran),
|
||||
"the fourth method must be reached automatically — proving a class can grow the "
|
||||
"the added method must be reached automatically — proving a class can grow the "
|
||||
+ "set of things it validates with no change to the sweep itself");
|
||||
}
|
||||
|
||||
@@ -209,18 +164,24 @@ class FleetConfigValidateAllTest {
|
||||
+ "name) must all be skipped");
|
||||
}
|
||||
|
||||
// ── Claim 2: FleetConfig.validateAll() is wired to that mechanism and reaches all six today ──
|
||||
// ── Claim 2: FleetConfig.validateAll() is wired to that mechanism and reaches every validator ──
|
||||
|
||||
/**
|
||||
* Reflectively enumerates {@link FleetConfig}'s own public, no-arg, void {@code validateXxx()}
|
||||
* methods (excluding {@code validateAll} itself) — the exact same filter {@link
|
||||
* FleetConfig#invokeAllValidators} applies. This is not the mechanism proof (that is claim 1,
|
||||
* above, on an unrelated class) — it is a visible denominator: today there are six, named
|
||||
* here, so a reader adding a seventh sees this assertion name the new count rather than a
|
||||
* silent pass at the old one.
|
||||
* above, on an unrelated class) — it is a visible denominator: today there are eight, named in
|
||||
* the {@code Set.of} below, so a reader adding or removing one sees this assertion name the new
|
||||
* count rather than a silent pass at the old one. The count lives only in that set, not in this
|
||||
* method's name, so the two cannot drift apart.
|
||||
*
|
||||
* <p>This assertion alone proves only that the validator exists with the right shape — it
|
||||
* cannot prove {@code validateAll()} actually reaches it. Only {@link
|
||||
* #validateAllReachesEveryOneOfTodaysRealValidators()} proves reachability, which is why this
|
||||
* method's failure message sends the reader there too.
|
||||
*/
|
||||
@Test
|
||||
void fleetConfigDeclaresExactlyTheseSixValidatorsToday() {
|
||||
void fleetConfigDeclaresExactlyTheseValidatorsToday() {
|
||||
Set<String> names = new TreeSet<>();
|
||||
for (Method m : FleetConfig.class.getMethods()) {
|
||||
if (java.lang.reflect.Modifier.isPublic(m.getModifiers())
|
||||
@@ -233,12 +194,18 @@ class FleetConfigValidateAllTest {
|
||||
}
|
||||
assertEquals(new TreeSet<>(Set.of("validateAuthExposure", "validateLeadTabPrefixes",
|
||||
"validateSubscriptionProfiles", "validateCharters", "validateMembers",
|
||||
"validateModels", "validateLeadRollover")), names,
|
||||
"FleetConfig's public validate*() methods changed. Do TWO things, in this "
|
||||
"validateModels", "validateLeadRollover", "validatePanePlacementAgainstLeadTabs")),
|
||||
names,
|
||||
"FleetConfig's public validate*() methods changed. Do THREE things, in this "
|
||||
+ "order. First confirm validateAll() still delegates to "
|
||||
+ "invokeAllValidators(this) — a hardcoded list there passes every other "
|
||||
+ "test in this class, so this assertion is the only place that will ever "
|
||||
+ "make you check. Only then update the expected set to match.");
|
||||
+ "make you check. Second, update the expected set below to match. Third, "
|
||||
+ "add or remove a case for that validator in "
|
||||
+ "validateAllReachesEveryOneOfTodaysRealValidators() below — this "
|
||||
+ "assertion proves only that the validator exists with the right shape, "
|
||||
+ "never that validateAll() reaches it; that enumeration is the test that "
|
||||
+ "does.");
|
||||
}
|
||||
|
||||
/** A minimal, otherwise-valid file — same shape FleetConfigTest and ConfigRefTest use. */
|
||||
@@ -262,14 +229,21 @@ class FleetConfigValidateAllTest {
|
||||
}
|
||||
|
||||
/**
|
||||
* The heart of claim 2: for each of today's six real validators, a minimal file that fails
|
||||
* ONLY that one — the exact fixtures {@code FleetConfigTest} uses to test each validator
|
||||
* directly — must also fail through {@link FleetConfig#validateAll()}. If a future edit to
|
||||
* {@code validateAll()} silently dropped one validator from the sweep (e.g. a typo'd name
|
||||
* filter), exactly one of these six would start passing when it must not.
|
||||
* The heart of claim 2: for every one of today's real validators, a minimal file that
|
||||
* fails ONLY that one — the exact fixtures {@code FleetConfigTest} uses to test each validator
|
||||
* directly, or a dedicated minimal fixture where no other test drives that validator through
|
||||
* {@code validateAll()} — must also fail through {@link FleetConfig#validateAll()}. If a
|
||||
* future edit to {@code validateAll()} silently dropped one of these from the sweep
|
||||
* (e.g. a typo'd name filter), exactly one of them would start passing when it must not.
|
||||
*
|
||||
* <p>This is the single place that proves {@code validateAll()} reaches a given validator.
|
||||
* Adding or removing a validator on {@link FleetConfig} must add or remove a case here, not
|
||||
* only an updated name in {@link #fleetConfigDeclaresExactlyTheseValidatorsToday()}'s expected
|
||||
* set — that assertion proves the validator's shape, never that {@code validateAll()} reaches
|
||||
* it.
|
||||
*/
|
||||
@Test
|
||||
void validateAllReachesEveryOneOfTodaysSixValidators(@TempDir Path dir) throws Exception {
|
||||
void validateAllReachesEveryOneOfTodaysRealValidators(@TempDir Path dir) throws Exception {
|
||||
// validateAuthExposure: a non-loopback bind without token mode.
|
||||
assertValidateAllRefuses(dir, "auth-exposure.yaml", """
|
||||
bind:
|
||||
@@ -277,16 +251,16 @@ class FleetConfigValidateAllTest {
|
||||
port: 8765
|
||||
""", "auth.mode: token");
|
||||
|
||||
// validateLeadTabPrefixes: a fleet-wide tabLabel that starts with a lead's own tabPrefix.
|
||||
// validateLeadTabPrefixes: a fleet-wide tabLabel that equals a lead tab.
|
||||
assertValidateAllRefuses(dir, "lead-tab-prefixes.yaml", """
|
||||
bind:
|
||||
host: 127.0.0.1
|
||||
port: 8765
|
||||
fleet:
|
||||
tabLabel: "lead: {role} {profile}"
|
||||
tabLabel: "alpha"
|
||||
leaders:
|
||||
opus:
|
||||
tab: "lead: opus"
|
||||
tab: "alpha"
|
||||
""", "fleet.tabLabel");
|
||||
|
||||
// validateSubscriptionProfiles: subscription: true with env: reseating ANTHROPIC_BASE_URL.
|
||||
@@ -342,6 +316,29 @@ class FleetConfigValidateAllTest {
|
||||
allow:
|
||||
- model: claude-sonnet-5
|
||||
""", "rogue");
|
||||
|
||||
// validatePanePlacementAgainstLeadTabs: a pane-placed profile while a lead names a tab.
|
||||
assertValidateAllRefuses(dir, "pane-placement.yaml", """
|
||||
bind:
|
||||
host: 127.0.0.1
|
||||
port: 8765
|
||||
profiles:
|
||||
gx10:
|
||||
placement: pane
|
||||
fleet:
|
||||
leaders:
|
||||
opus:
|
||||
tab: "lead: opus"
|
||||
""", "gx10");
|
||||
|
||||
// validateLeadRollover: a leadRollover: block present with no handoverPath.
|
||||
assertValidateAllRefuses(dir, "lead-rollover.yaml", """
|
||||
bind:
|
||||
host: 127.0.0.1
|
||||
port: 8765
|
||||
leadRollover:
|
||||
requireOperatorConfirm: false
|
||||
""", "handoverPath");
|
||||
}
|
||||
|
||||
private static void assertValidateAllRefuses(Path dir, String fileName, String yaml,
|
||||
|
||||
@@ -51,6 +51,7 @@ public final class FakeHerdr implements HerdrClient {
|
||||
private boolean noPanes = false;
|
||||
private volatile String agentStatus = "idle"; // steady-state agent.get status
|
||||
private volatile String agentType = "claude"; // detected agent kind on agent.get; null = undetected
|
||||
private volatile String agentSessionId = null; // agent_session.value on agent.get; null = omitted
|
||||
private volatile String readText = "worker transcript tail"; // canned agent.read output
|
||||
private int pinnedStarts = 0; // how many upcoming agent.start calls report a fixed pane
|
||||
private String pinnedStartTerminal;
|
||||
@@ -173,6 +174,16 @@ public final class FakeHerdr implements HerdrClient {
|
||||
return this;
|
||||
}
|
||||
|
||||
/**
|
||||
* Set the {@code agent_session.value} that {@code agent.get} reports for {@code term_a} — the
|
||||
* default omits the field entirely (herdr not yet having resolved one), matching the real
|
||||
* daemon's own "not resolved yet" shape.
|
||||
*/
|
||||
public FakeHerdr agentSessionId(String sessionId) {
|
||||
this.agentSessionId = sessionId;
|
||||
return this;
|
||||
}
|
||||
|
||||
/**
|
||||
* Make {@code agent.get} succeed normally for its first {@code okCalls} invocations, then fail
|
||||
* every call after that with {@code code} — fleetd #176 fix 1's "backend exited mid-wait"
|
||||
@@ -311,10 +322,12 @@ public final class FakeHerdr implements HerdrClient {
|
||||
}
|
||||
}
|
||||
String agentField = agentType == null ? "null" : "\"" + agentType + "\"";
|
||||
String sessionField = agentSessionId == null ? ""
|
||||
: ",\"agent_session\":{\"kind\":\"id\",\"value\":\"" + agentSessionId + "\"}";
|
||||
yield mapper.readTree(("""
|
||||
{"type":"agent_info","agent":{"terminal_id":"term_a","agent":%s,
|
||||
"agent_status":"%s","workspace_id":"w2","tab_id":"w2:t7","pane_id":"w2:p7"}}""")
|
||||
.formatted(agentField, agentStatus));
|
||||
"agent_status":"%s","workspace_id":"w2","tab_id":"w2:t7","pane_id":"w2:p7"%s}}""")
|
||||
.formatted(agentField, agentStatus, sessionField));
|
||||
}
|
||||
case "agent.read" -> mapper.readTree(mapper.writeValueAsString(
|
||||
java.util.Map.of("type", "agent_read", "read", java.util.Map.of("text", readText))));
|
||||
@@ -456,7 +469,11 @@ public final class FakeHerdr implements HerdrClient {
|
||||
}
|
||||
}
|
||||
|
||||
/** fleetd #612 Unit A: set by {@link #close()} so a test can prove a router/client actually closed this. */
|
||||
public volatile boolean closed = false;
|
||||
|
||||
@Override
|
||||
public void close() {
|
||||
closed = true;
|
||||
}
|
||||
}
|
||||
|
||||
@@ -232,8 +232,9 @@ class LeadTabScannerTest {
|
||||
|
||||
@Test
|
||||
void everyPaneInALeadTabResolvesAsThatLead() {
|
||||
// A human may split their own lead tab. Both panes are theirs, so both are that lead —
|
||||
// nothing fleetd placed can land here (see the worker-space test above).
|
||||
// A human may split their own lead tab. Both panes are theirs, so both are that lead.
|
||||
// A pane-placed member landing here instead is refused at startup by
|
||||
// FleetConfig.validatePanePlacementAgainstLeadTabs, not by this scanner.
|
||||
TopologyHerdr herdr = twoLeads().pane("w1:p1b", "w1:t1", "term_opus_split");
|
||||
|
||||
assertEquals("opus-5.0",
|
||||
|
||||
@@ -0,0 +1,104 @@
|
||||
package dev.ltms.fleet.lead;
|
||||
|
||||
import org.junit.jupiter.api.DisplayName;
|
||||
import org.junit.jupiter.api.Test;
|
||||
import org.junit.jupiter.api.io.TempDir;
|
||||
|
||||
import java.io.IOException;
|
||||
import java.nio.charset.StandardCharsets;
|
||||
import java.nio.file.Files;
|
||||
import java.nio.file.Path;
|
||||
|
||||
import static org.junit.jupiter.api.Assertions.assertEquals;
|
||||
|
||||
/**
|
||||
* {@code LeadContextGauge}'s HIGH threshold scales with the caller's resolved effective
|
||||
* auto-compact window, so the margin it warns at makes sense against a window that can legally sit
|
||||
* as low as {@code 100_000}, not only against the fixed fallback. These properties pin that
|
||||
* scaling, its boundary, and the fallback used when no window is resolvable.
|
||||
*/
|
||||
class LeadContextGaugeHighThresholdTest {
|
||||
|
||||
private static final String SESSION_ID = "55555555-5555-5555-5555-555555555555";
|
||||
|
||||
private static String usageLine(long tokens) {
|
||||
return "{\"type\":\"assistant\",\"message\":{\"role\":\"assistant\",\"usage\":{"
|
||||
+ "\"input_tokens\":" + tokens + ",\"cache_read_input_tokens\":0,\"cache_creation_input_tokens\":0}}}";
|
||||
}
|
||||
|
||||
private static String writeTranscript(Path configDir, String sessionId, long tokens) throws IOException {
|
||||
Path projectDir = configDir.resolve("projects").resolve("some-project-slug");
|
||||
Files.createDirectories(projectDir);
|
||||
Files.writeString(projectDir.resolve(sessionId + ".jsonl"), usageLine(tokens) + "\n", StandardCharsets.UTF_8);
|
||||
return configDir.toString();
|
||||
}
|
||||
|
||||
@Test
|
||||
@DisplayName("an effective window of 100000 reports HIGH strictly below 100000")
|
||||
void effectiveWindowOf100000ReportsHighStrictlyBelow100000(@TempDir Path tmp) throws IOException {
|
||||
// 90,000 is below the 100,000 window itself, but above a fixed 200,000 fallback would ever
|
||||
// reach — only a threshold that scales with the window can report HIGH here.
|
||||
String configDir = writeTranscript(tmp, SESSION_ID, 90_000);
|
||||
LeadContextGauge gauge = new LeadContextGauge();
|
||||
|
||||
LeadContextGauge.Reading reading = gauge.read(configDir, SESSION_ID, "claude", 100_000L);
|
||||
|
||||
assertEquals(LeadContextGauge.State.HIGH, reading.state(),
|
||||
"90,000 tokens against a 100,000 effective window must already be HIGH, with margin to spare "
|
||||
+ "before a compaction at the window itself");
|
||||
}
|
||||
|
||||
@Test
|
||||
@DisplayName("the boundary sits strictly between OK and HIGH on both sides")
|
||||
void boundarySitsStrictlyBetweenOkAndHighOnBothSides(@TempDir Path tmp) throws IOException {
|
||||
long window = 100_000L;
|
||||
long threshold = (long) (window * (2.0 / 3.0)); // 66,666
|
||||
|
||||
String belowConfigDir = writeTranscript(tmp.resolve("below"), SESSION_ID, threshold - 1);
|
||||
LeadContextGauge belowGauge = new LeadContextGauge();
|
||||
assertEquals(LeadContextGauge.State.OK,
|
||||
belowGauge.read(belowConfigDir, SESSION_ID, "claude", window).state(),
|
||||
"one token short of the threshold must stay OK");
|
||||
|
||||
String atConfigDir = writeTranscript(tmp.resolve("at"), SESSION_ID, threshold);
|
||||
LeadContextGauge atGauge = new LeadContextGauge();
|
||||
assertEquals(LeadContextGauge.State.HIGH,
|
||||
atGauge.read(atConfigDir, SESSION_ID, "claude", window).state(),
|
||||
"exactly at the threshold must already be HIGH");
|
||||
}
|
||||
|
||||
@Test
|
||||
@DisplayName("with no effective window resolvable, the threshold is still the fixed 200000 default")
|
||||
void noEffectiveWindowFallsBackToTheFixed200000Default(@TempDir Path tmp) throws IOException {
|
||||
String belowConfigDir = writeTranscript(tmp.resolve("below"), SESSION_ID, 199_999);
|
||||
LeadContextGauge belowGauge = new LeadContextGauge();
|
||||
assertEquals(LeadContextGauge.State.OK,
|
||||
belowGauge.read(belowConfigDir, SESSION_ID, "claude", null).state(),
|
||||
"one token short of the fixed default must stay OK when no window is resolvable");
|
||||
|
||||
String atConfigDir = writeTranscript(tmp.resolve("at"), SESSION_ID, 200_000);
|
||||
LeadContextGauge atGauge = new LeadContextGauge();
|
||||
assertEquals(LeadContextGauge.State.HIGH,
|
||||
atGauge.read(atConfigDir, SESSION_ID, "claude", null).state(),
|
||||
"the fixed default must still be 200,000 when no window is resolvable");
|
||||
}
|
||||
|
||||
@Test
|
||||
@DisplayName("a second read with a different window, within the TTL, reports against its own window, not the first call's cached state")
|
||||
void aSecondReadWithADifferentWindowWithinTheTtlReportsAgainstItsOwnWindow(@TempDir Path tmp) throws IOException {
|
||||
String configDir = writeTranscript(tmp, SESSION_ID, 90_000);
|
||||
long[] now = {0L};
|
||||
LeadContextGauge gauge = new LeadContextGauge(() -> now[0], 5_000);
|
||||
|
||||
LeadContextGauge.Reading first = gauge.read(configDir, SESSION_ID, "claude", 100_000L);
|
||||
assertEquals(LeadContextGauge.State.HIGH, first.state(),
|
||||
"90,000 tokens against a 100,000 window is HIGH");
|
||||
|
||||
now[0] += 1_000; // stays inside the 5,000ms TTL — the cache key must still vary with the window
|
||||
LeadContextGauge.Reading second = gauge.read(configDir, SESSION_ID, "claude", 1_000_000L);
|
||||
|
||||
assertEquals(LeadContextGauge.State.OK, second.state(),
|
||||
"90,000 tokens against a 1,000,000 window must report OK regardless of the previous call's "
|
||||
+ "window, even while that call's cache entry is still within its TTL");
|
||||
}
|
||||
}
|
||||
@@ -85,7 +85,7 @@ class LeadContextGaugeTest {
|
||||
usageLine(40_000, 5_000, 3_000)); // last record: 48,000
|
||||
LeadContextGauge gauge = new LeadContextGauge();
|
||||
|
||||
LeadContextGauge.Reading first = gauge.read(configDir, SESSION_ID, "claude");
|
||||
LeadContextGauge.Reading first = gauge.read(configDir, SESSION_ID, "claude", null);
|
||||
assertEquals(48_000L, first.tokens(), "must total input+cache_read+cache_creation of the LAST usage record");
|
||||
assertEquals(LeadContextGauge.State.OK, first.state());
|
||||
|
||||
@@ -93,7 +93,7 @@ class LeadContextGaugeTest {
|
||||
// must change with it, not stay pinned to the first fixture's total.
|
||||
String otherSession = "22222222-2222-2222-2222-222222222222";
|
||||
writeTranscript(tmp, otherSession, usageLine(100_000, 50_000, 50_000)); // last record: 200,000
|
||||
LeadContextGauge.Reading second = gauge.read(configDir, otherSession, "claude");
|
||||
LeadContextGauge.Reading second = gauge.read(configDir, otherSession, "claude", null);
|
||||
assertEquals(200_000L, second.tokens());
|
||||
assertTrue(second.tokens() != first.tokens(), "changing N in the fixture must change the reported number");
|
||||
}
|
||||
@@ -111,12 +111,12 @@ class LeadContextGaugeTest {
|
||||
compactionLine(),
|
||||
usageLine(3_000, 0, 0));
|
||||
LeadContextGauge gauge = new LeadContextGauge();
|
||||
LeadContextGauge.Reading twoCompactions = gauge.read(tmp.toString(), sessionTwoCompactions, "claude");
|
||||
LeadContextGauge.Reading twoCompactions = gauge.read(tmp.toString(), sessionTwoCompactions, "claude", null);
|
||||
assertEquals(2, twoCompactions.compactions());
|
||||
|
||||
String sessionZeroCompactions = "44444444-4444-4444-4444-444444444444";
|
||||
writeTranscript(tmp, sessionZeroCompactions, usageLine(3_000, 0, 0));
|
||||
LeadContextGauge.Reading zeroCompactions = gauge.read(tmp.toString(), sessionZeroCompactions, "claude");
|
||||
LeadContextGauge.Reading zeroCompactions = gauge.read(tmp.toString(), sessionZeroCompactions, "claude", null);
|
||||
assertEquals(0, zeroCompactions.compactions(), "changing K in the fixture must change the reported count");
|
||||
}
|
||||
|
||||
@@ -126,7 +126,7 @@ class LeadContextGaugeTest {
|
||||
@DisplayName("a missing transcript file reports UNKNOWN with no token number")
|
||||
void missingFileIsUnknown(@TempDir Path tmp) {
|
||||
LeadContextGauge gauge = new LeadContextGauge();
|
||||
LeadContextGauge.Reading reading = gauge.read(tmp.toString(), SESSION_ID, "claude");
|
||||
LeadContextGauge.Reading reading = gauge.read(tmp.toString(), SESSION_ID, "claude", null);
|
||||
assertEquals(LeadContextGauge.State.UNKNOWN, reading.state());
|
||||
assertNull(reading.tokens());
|
||||
}
|
||||
@@ -148,7 +148,7 @@ class LeadContextGaugeTest {
|
||||
assumeFalse(Files.isReadable(file),
|
||||
"runs as root (CI container): the read bit does not stop root, so this case cannot be set up here");
|
||||
LeadContextGauge gauge = new LeadContextGauge();
|
||||
LeadContextGauge.Reading reading = gauge.read(configDir, SESSION_ID, "claude");
|
||||
LeadContextGauge.Reading reading = gauge.read(configDir, SESSION_ID, "claude", null);
|
||||
assertEquals(LeadContextGauge.State.UNKNOWN, reading.state());
|
||||
assertNull(reading.tokens());
|
||||
} finally {
|
||||
@@ -171,7 +171,7 @@ class LeadContextGaugeTest {
|
||||
Files.writeString(file, lastCompleteLine + "\n" + tornLine, StandardCharsets.UTF_8);
|
||||
|
||||
LeadContextGauge gauge = new LeadContextGauge();
|
||||
LeadContextGauge.Reading reading = gauge.read(tmp.toString(), SESSION_ID, "claude");
|
||||
LeadContextGauge.Reading reading = gauge.read(tmp.toString(), SESSION_ID, "claude", null);
|
||||
assertEquals(LeadContextGauge.State.OK, reading.state(),
|
||||
"a torn final line must not turn a good earlier reading into UNKNOWN");
|
||||
assertEquals(6_000L, reading.tokens(),
|
||||
@@ -187,7 +187,7 @@ class LeadContextGaugeTest {
|
||||
"{this is not json at all",
|
||||
"neither is this{{{");
|
||||
LeadContextGauge gauge = new LeadContextGauge();
|
||||
LeadContextGauge.Reading reading = gauge.read(configDir, SESSION_ID, "claude");
|
||||
LeadContextGauge.Reading reading = gauge.read(configDir, SESSION_ID, "claude", null);
|
||||
assertEquals(LeadContextGauge.State.UNKNOWN, reading.state(),
|
||||
"every line unparseable is the real format-change signal and must still report UNKNOWN");
|
||||
assertNull(reading.tokens());
|
||||
@@ -224,12 +224,12 @@ class LeadContextGaugeTest {
|
||||
AtomicLong now = new AtomicLong(0);
|
||||
LeadContextGauge gauge = new LeadContextGauge(now::get, 5_000);
|
||||
|
||||
gauge.read(configDir, SESSION_ID, "claude");
|
||||
gauge.read(configDir, SESSION_ID, "claude"); // still inside the TTL window
|
||||
gauge.read(configDir, SESSION_ID, "claude", null);
|
||||
gauge.read(configDir, SESSION_ID, "claude", null); // still inside the TTL window
|
||||
assertEquals(1, gauge.diskReadCount(), "two reads inside the TTL must touch disk once");
|
||||
|
||||
now.set(6_000); // past the TTL
|
||||
gauge.read(configDir, SESSION_ID, "claude");
|
||||
gauge.read(configDir, SESSION_ID, "claude", null);
|
||||
assertEquals(2, gauge.diskReadCount(), "a read past the TTL must touch disk again");
|
||||
}
|
||||
|
||||
@@ -241,8 +241,8 @@ class LeadContextGaugeTest {
|
||||
String configDir = writeTranscript(tmp, SESSION_ID, usageLine(1_000, 0, 0));
|
||||
LeadContextGauge gauge = new LeadContextGauge();
|
||||
|
||||
assertEquals(LeadContextGauge.State.UNKNOWN, gauge.read(configDir, SESSION_ID, "opencode").state());
|
||||
assertEquals(LeadContextGauge.State.UNKNOWN, gauge.read(configDir, SESSION_ID, null).state());
|
||||
assertEquals(LeadContextGauge.State.UNKNOWN, gauge.read(configDir, null, "claude").state());
|
||||
assertEquals(LeadContextGauge.State.UNKNOWN, gauge.read(configDir, SESSION_ID, "opencode", null).state());
|
||||
assertEquals(LeadContextGauge.State.UNKNOWN, gauge.read(configDir, SESSION_ID, null, null).state());
|
||||
assertEquals(LeadContextGauge.State.UNKNOWN, gauge.read(configDir, null, "claude", null).state());
|
||||
}
|
||||
}
|
||||
|
||||
@@ -156,6 +156,45 @@ class FleetMcpAuthzTest {
|
||||
}
|
||||
}
|
||||
|
||||
/**
|
||||
* fleetd #669 Unit A: {@code SEND} is split into three actions ({@link Authz.Action#SEND},
|
||||
* {@link Authz.Action#COORD_SEND}, {@link Authz.Action#ANSWER}), each carrying the same grant
|
||||
* the one undivided action gave. An architect holds all three, exactly as it held the one.
|
||||
*/
|
||||
@Test
|
||||
void anArchitectMayUseAllThreeSendShapesOverMcp() {
|
||||
FleetMcp m = mcp(true);
|
||||
for (Authz.Action a : new Authz.Action[]{Authz.Action.SEND, Authz.Action.COORD_SEND,
|
||||
Authz.Action.ANSWER}) {
|
||||
assertNull(m.denyFor(ARCH_DESIGN, a, "term_a"),
|
||||
a + " carries the same grant the undivided SEND action gave an architect");
|
||||
}
|
||||
}
|
||||
|
||||
/** The other half of the same split: a worker is excluded from all three, as it was from one. */
|
||||
@Test
|
||||
void aWorkerMayNotUseAnySendShapeOverMcp() {
|
||||
FleetMcp m = mcp(true);
|
||||
for (Authz.Action a : new Authz.Action[]{Authz.Action.SEND, Authz.Action.COORD_SEND,
|
||||
Authz.Action.ANSWER}) {
|
||||
McpSchema.CallToolResult denied = m.denyFor(WORKER_A, a, "term_a");
|
||||
assertNotNull(denied, a + " must stay refused to a worker");
|
||||
assertTrue(denied.isError(), "a refusal is returned as an MCP tool error");
|
||||
}
|
||||
}
|
||||
|
||||
/**
|
||||
* fleetd #669 Unit A / #678: {@code TASK_READ} (ticket polling, session status) is split out of
|
||||
* the roster-only {@code READ}, carrying forward the grant the undivided action gave. A worker
|
||||
* still has both — it never gained or lost anything by the split.
|
||||
*/
|
||||
@Test
|
||||
void aWorkerKeepsBothReadActionsAfterTheSplit() {
|
||||
FleetMcp m = mcp(true);
|
||||
assertNull(m.denyFor(WORKER_A, Authz.Action.READ, null));
|
||||
assertNull(m.denyFor(WORKER_A, Authz.Action.TASK_READ, null));
|
||||
}
|
||||
|
||||
@Test
|
||||
void anArchitectMayReplyAndAskOnlyAsItsOwnPaneOverMcp() {
|
||||
FleetMcp m = mcp(true);
|
||||
@@ -297,35 +336,53 @@ class FleetMcpAuthzTest {
|
||||
* the whole time the defect was live -- the table was right, the action fed to it was wrong.
|
||||
*/
|
||||
@Test
|
||||
void pollingByTargetIsADrainAndPollingByTicketIsARead() {
|
||||
void pollingByTargetIsADrainAndPollingByTicketIsATaskRead() {
|
||||
assertEquals(Authz.Action.DRAIN, FleetMcp.pollAction("term_b", null),
|
||||
"poll by target removes the replies — that is a drain, not an observation");
|
||||
assertEquals(Authz.Action.READ, FleetMcp.pollAction(null, null),
|
||||
"poll by ticket changes nothing");
|
||||
assertEquals(Authz.Action.READ, FleetMcp.pollAction(" ", null),
|
||||
assertEquals(Authz.Action.TASK_READ, FleetMcp.pollAction(null, null),
|
||||
"poll by ticket changes nothing, but is not the roster-only READ action");
|
||||
assertEquals(Authz.Action.TASK_READ, FleetMcp.pollAction(" ", null),
|
||||
"a blank target is an absent target");
|
||||
}
|
||||
|
||||
/**
|
||||
* fleetd #421: a coordId branch is a THIRD operation behind fleet_poll's one name, and it must
|
||||
* map to {@link Authz.Action#COORD_READ} — never {@link Authz.Action#READ}, even though this
|
||||
* branch also consumes nothing. READ's grant is open to every authenticated role on the premise
|
||||
* that the roster carries no secrets; a lead-to-lead body is not the roster, so folding this
|
||||
* branch into READ would let any worker read every peer lead's mail in full. coordId also takes
|
||||
* priority over target when both happen to be present — it addresses a different inbox entirely.
|
||||
* map to {@link Authz.Action#COORD_READ} — never {@link Authz.Action#READ} or {@link
|
||||
* Authz.Action#TASK_READ}, even though this branch also consumes nothing. A lead-to-lead body
|
||||
* is not the roster and not a ticket/status read, so folding this branch into either would let
|
||||
* any worker or architect read every peer lead's mail in full. coordId also takes priority over
|
||||
* target when both happen to be present — it addresses a different inbox entirely.
|
||||
*/
|
||||
@Test
|
||||
void pollingByCoordIdIsACoordReadNeverAPlainRead() {
|
||||
void pollingByCoordIdIsACoordReadNeverAPlainOrTaskRead() {
|
||||
assertEquals(Authz.Action.COORD_READ, FleetMcp.pollAction(null, "mac-opus"),
|
||||
"reading held peer mail must not be mapped to the everyone-readable READ action");
|
||||
"reading held peer mail must not be mapped to a widely-readable action");
|
||||
assertEquals(Authz.Action.COORD_READ, FleetMcp.pollAction(" ", "mac-opus"),
|
||||
"a blank target must not fall through to READ/DRAIN when coordId is present");
|
||||
assertEquals(Authz.Action.READ, FleetMcp.pollAction(null, " "),
|
||||
assertEquals(Authz.Action.TASK_READ, FleetMcp.pollAction(null, " "),
|
||||
"a blank coordId is an absent coordId, same as target/ticket");
|
||||
assertEquals(Authz.Action.COORD_READ, FleetMcp.pollAction("term_b", "mac-opus"),
|
||||
"coordId takes priority over target — this is a different inbox, not a drain");
|
||||
}
|
||||
|
||||
/**
|
||||
* {@code fleet_send} is three call shapes behind one tool name, exactly as {@code fleet_poll}
|
||||
* is (fleetd #669 Unit A). {@link FleetMcp#sendAction} picks the action from the arguments, not
|
||||
* the handler, for the same reason {@link FleetMcp#pollAction} does: a test can assert the
|
||||
* mapping the handler actually uses.
|
||||
*/
|
||||
@Test
|
||||
void sendMapsToThreeDifferentActionsByItsArguments() {
|
||||
assertEquals(Authz.Action.SEND, FleetMcp.sendAction(null, null),
|
||||
"a plain delivery, with neither coordId nor turnId, is a local SEND");
|
||||
assertEquals(Authz.Action.COORD_SEND, FleetMcp.sendAction("mac-opus", null),
|
||||
"coordId addresses a peer lead over the coordination broker");
|
||||
assertEquals(Authz.Action.ANSWER, FleetMcp.sendAction(null, "turn-1"),
|
||||
"turnId resolves a worker's blocked question");
|
||||
assertEquals(Authz.Action.COORD_SEND, FleetMcp.sendAction("mac-opus", "turn-1"),
|
||||
"coordId takes priority over turnId, mirroring sendToLead's own mutual-exclusion check");
|
||||
}
|
||||
|
||||
@Test
|
||||
void everyRegisteredToolHasItsHandlerActionPinned() {
|
||||
// fleetd #469: this used to scrape FleetMcp.java's tool("…") calls for the registered set —
|
||||
@@ -344,16 +401,22 @@ class FleetMcpAuthzTest {
|
||||
() -> tool + " is registered but has no pinned authorization action"));
|
||||
|
||||
assertEquals(Authz.Action.SEND, FleetMcp.toolAction("fleet_send", Map.of()));
|
||||
assertEquals(Authz.Action.SEND,
|
||||
FleetMcp.toolAction("fleet_send", Map.of("sessionId", "term_a", "content", "hi")));
|
||||
assertEquals(Authz.Action.COORD_SEND,
|
||||
FleetMcp.toolAction("fleet_send", Map.of("coordId", "mac-opus", "content", "hi")));
|
||||
assertEquals(Authz.Action.ANSWER,
|
||||
FleetMcp.toolAction("fleet_send", Map.of("turnId", "turn-1", "content", "hi")));
|
||||
assertEquals(Authz.Action.REPLY, FleetMcp.toolAction("fleet_reply", Map.of()));
|
||||
assertEquals(Authz.Action.ASK, FleetMcp.toolAction("fleet_ask", Map.of()));
|
||||
assertEquals(Authz.Action.READ, FleetMcp.toolAction("fleet_status", Map.of()));
|
||||
assertEquals(Authz.Action.TASK_READ, FleetMcp.toolAction("fleet_status", Map.of()));
|
||||
assertEquals(Authz.Action.DRAIN, FleetMcp.toolAction("fleet_ack", Map.of()));
|
||||
assertEquals(Authz.Action.SPAWN, FleetMcp.toolAction("fleet_spawn", Map.of()));
|
||||
assertEquals(Authz.Action.READ, FleetMcp.toolAction("fleet_list", Map.of()));
|
||||
assertEquals(Authz.Action.STOP, FleetMcp.toolAction("fleet_stop", Map.of()));
|
||||
assertEquals(Authz.Action.READ, FleetMcp.toolAction("fleet_profiles", Map.of()));
|
||||
assertEquals(Authz.Action.READ, FleetMcp.toolAction("fleet_whoami", Map.of()));
|
||||
assertEquals(Authz.Action.READ, FleetMcp.toolAction("fleet_poll", Map.of("ticket", "task")));
|
||||
assertEquals(Authz.Action.TASK_READ, FleetMcp.toolAction("fleet_poll", Map.of("ticket", "task")));
|
||||
assertEquals(Authz.Action.DRAIN, FleetMcp.toolAction("fleet_poll", Map.of("target", "term_b")));
|
||||
assertEquals(Authz.Action.COORD_READ,
|
||||
FleetMcp.toolAction("fleet_poll", Map.of("coordId", "mac-opus")));
|
||||
|
||||
@@ -86,7 +86,7 @@ class FleetMcpLeadContextGaugeWiringTest {
|
||||
LeadContextGauge contextGauge = new LeadContextGauge();
|
||||
AtomicReference<String> configuredDir = new AtomicReference<>(dirA.toString());
|
||||
FleetMcp.LeadConfigDirSource source = new FleetMcp.LeadConfigDirSource(name ->
|
||||
LEAD_NAME.equals(name) ? configuredDir.get() : null);
|
||||
LEAD_NAME.equals(name) ? configuredDir.get() : null, _ -> null);
|
||||
|
||||
String firstRead = textOf(listFleet(herdr, sessions, contextGauge, source));
|
||||
assertTrue(firstRead.contains("\"tokens\":11000"),
|
||||
|
||||
@@ -60,13 +60,25 @@ class MessageServiceTest {
|
||||
inbox.own(T);
|
||||
}
|
||||
|
||||
/**
|
||||
* A send budget large enough that a test's own setup — {@link #awaitWaiting()} plus whatever
|
||||
* status transitions it drives afterward — can never compete with it for the same clock. A test
|
||||
* that needs {@code send.get(...)}'s own window to be the only timing bound it depends on uses
|
||||
* {@link #sendAsync(String, long)} with this value instead of the default 5000 ms.
|
||||
*/
|
||||
private static final long GENEROUS_SEND_BUDGET_MILLIS = 30_000;
|
||||
|
||||
/** Run {@code send} on a background thread; the current thread drives the worker's turn. */
|
||||
private CompletableFuture<MessageService.Reply> sendAsync() {
|
||||
return sendAsync("do the task");
|
||||
}
|
||||
|
||||
private CompletableFuture<MessageService.Reply> sendAsync(String content) {
|
||||
return CompletableFuture.supplyAsync(() -> messages.send(T, content, 5000));
|
||||
return sendAsync(content, 5000);
|
||||
}
|
||||
|
||||
private CompletableFuture<MessageService.Reply> sendAsync(String content, long timeoutMillis) {
|
||||
return CompletableFuture.supplyAsync(() -> messages.send(T, content, timeoutMillis));
|
||||
}
|
||||
|
||||
private void awaitWaiting() throws InterruptedException {
|
||||
@@ -80,7 +92,7 @@ class MessageServiceTest {
|
||||
|
||||
@Test
|
||||
void completionFallbackResolvesATurnThatNeverCalledFleetReply() throws Exception {
|
||||
CompletableFuture<MessageService.Reply> send = sendAsync();
|
||||
CompletableFuture<MessageService.Reply> send = sendAsync("do the task", GENEROUS_SEND_BUDGET_MILLIS);
|
||||
awaitWaiting();
|
||||
|
||||
herdr.readText("$ prompt"); // pre-turn pane: no answer yet (baseline reference)
|
||||
@@ -96,6 +108,32 @@ class MessageServiceTest {
|
||||
assertTrue(reply.completed(), "a scraped completion still counts as completed");
|
||||
}
|
||||
|
||||
/**
|
||||
* Pins {@link #GENEROUS_SEND_BUDGET_MILLIS} as the budget {@link
|
||||
* #completionFallbackResolvesATurnThatNeverCalledFleetReply} depends on. A 5500 ms delay between
|
||||
* {@link #awaitWaiting()} and the status transitions that drive completion stands in for a loaded
|
||||
* machine's setup overhead — comfortably past the 5000 ms budget this send no longer uses, and
|
||||
* still well inside this method's own 30 000 ms budget. The only clock this test depends on is
|
||||
* {@code send.get}'s own 10 s window.
|
||||
*/
|
||||
@Test
|
||||
void completionFallbackSurvivesASlowHarnessBecauseItsSendBudgetIsNotTheBindingClock() throws Exception {
|
||||
CompletableFuture<MessageService.Reply> send = sendAsync("do the task", GENEROUS_SEND_BUDGET_MILLIS);
|
||||
awaitWaiting();
|
||||
|
||||
Thread.sleep(5500);
|
||||
|
||||
herdr.readText("$ prompt");
|
||||
injector.onStatus(T, AgentStatus.IDLE);
|
||||
injector.onStatus(T, AgentStatus.WORKING);
|
||||
herdr.readText("BUILD GREEN: 391 files");
|
||||
injector.onStatus(T, AgentStatus.IDLE);
|
||||
|
||||
MessageService.Reply reply = send.get(10, TimeUnit.SECONDS);
|
||||
assertEquals(MessageService.Outcome.COMPLETED_UNREPLIED, reply.outcome(),
|
||||
"a slow harness must not be mistaken for a timed-out delivery");
|
||||
}
|
||||
|
||||
@Test
|
||||
void completionFallbackReplacesAnEchoedInjectedBriefWithNoReportOutcome() throws Exception {
|
||||
String brief = "Implement the requested change. ".repeat(20);
|
||||
@@ -1846,7 +1884,8 @@ class MessageServiceTest {
|
||||
* but backed by {@link ManualScheduler} instead of a real timer (fleetd #608): its tick never
|
||||
* fires on its own — a test drives it explicitly via {@link ManualScheduler#runDueTasks()}.
|
||||
*/
|
||||
private record ManualPushWiring(MessageService service, FakeHerdr leadHerdr, ManualScheduler scheduler)
|
||||
private record ManualPushWiring(MessageService service, FakeHerdr leadHerdr, ManualScheduler scheduler,
|
||||
ReplyPushLoop pushLoop)
|
||||
implements AutoCloseable {
|
||||
@Override
|
||||
public void close() {
|
||||
@@ -1861,14 +1900,20 @@ class MessageServiceTest {
|
||||
* {@code anAlreadyCollectedTicketProducesNoNudge}, which sets it to 1 to prove exactly that.
|
||||
*/
|
||||
private ManualPushWiring wireWithManualScheduler(int maxReminders, long backoffMs) {
|
||||
return wireWithManualScheduler(maxReminders, backoffMs, System::nanoTime);
|
||||
}
|
||||
|
||||
/** As above, with an injectable clock for tests that exercise terminal-ticket pruning. */
|
||||
private ManualPushWiring wireWithManualScheduler(int maxReminders, long backoffMs,
|
||||
java.util.function.LongSupplier nowNanos) {
|
||||
PrimaryRegistry registry = new PrimaryRegistry(null);
|
||||
registry.recordDelegation(T, LEAD);
|
||||
FakeHerdr leadHerdr = new FakeHerdr();
|
||||
AgentControl leadAgents = new AgentControl(leadHerdr);
|
||||
ManualScheduler scheduler = new ManualScheduler();
|
||||
ReplyPushLoop pushLoop = new ReplyPushLoop(registry, leadAgents, inbox, scheduler, maxReminders, backoffMs);
|
||||
MessageService service = new MessageService(agents, injector, rendezvous, inbox, pushLoop, null, System::nanoTime);
|
||||
return new ManualPushWiring(service, leadHerdr, scheduler);
|
||||
MessageService service = new MessageService(agents, injector, rendezvous, inbox, pushLoop, null, nowNanos);
|
||||
return new ManualPushWiring(service, leadHerdr, scheduler, pushLoop);
|
||||
}
|
||||
|
||||
private void awaitNudge(FakeHerdr leadHerdr) throws InterruptedException {
|
||||
@@ -1959,24 +2004,35 @@ class MessageServiceTest {
|
||||
|
||||
@Test
|
||||
void severalAsyncTicketsFinishingTogetherProduceOneCoalescedNudge() throws Exception {
|
||||
try (var wiring = wireWithPushLoop(1, 300)) { // wide backoff: both tickets land before the tick fires
|
||||
String first = wiring.service().sendAsync(T, "first task");
|
||||
awaitWaiting();
|
||||
injectDelivery();
|
||||
assertTrue(rendezvous.resolve(T, "first done"));
|
||||
// Settle without polling: poll() itself marks a ticket collected (that's the point of
|
||||
// anAlreadyCollectedTicketProducesNoNudge above) — using it here to detect completion
|
||||
// would collect the ticket before the coalescing this test checks ever gets a chance.
|
||||
Thread.sleep(100);
|
||||
// A 1ms backoff is due immediately. ManualScheduler still cannot run it until this test
|
||||
// explicitly calls runDueTasks(), so both terminal tickets join one scheduled tick.
|
||||
try (var wiring = wireWithManualScheduler(1, 1)) {
|
||||
String first;
|
||||
String second;
|
||||
java.util.concurrent.CountDownLatch firstTerminalReached = new java.util.concurrent.CountDownLatch(1);
|
||||
wiring.service().setAfterFinishAsyncTaskCompleteHookForTest(firstTerminalReached::countDown);
|
||||
try {
|
||||
first = wiring.service().sendAsync(T, "first task");
|
||||
awaitWaiting();
|
||||
injectDelivery();
|
||||
assertTrue(rendezvous.resolve(T, "first done"));
|
||||
assertTrue(firstTerminalReached.await(5, TimeUnit.SECONDS),
|
||||
"the first ticket never reached its terminal phase");
|
||||
|
||||
String second = wiring.service().sendAsync(T, "second task");
|
||||
awaitWaiting();
|
||||
injectDelivery();
|
||||
assertTrue(rendezvous.resolve(T, "second done"));
|
||||
Thread.sleep(100);
|
||||
java.util.concurrent.CountDownLatch secondTerminalReached = new java.util.concurrent.CountDownLatch(1);
|
||||
wiring.service().setAfterFinishAsyncTaskCompleteHookForTest(secondTerminalReached::countDown);
|
||||
second = wiring.service().sendAsync(T, "second task");
|
||||
awaitWaiting();
|
||||
injectDelivery();
|
||||
assertTrue(rendezvous.resolve(T, "second done"));
|
||||
assertTrue(secondTerminalReached.await(5, TimeUnit.SECONDS),
|
||||
"the second ticket never reached its terminal phase");
|
||||
} finally {
|
||||
wiring.service().setAfterFinishAsyncTaskCompleteHookForTest(null);
|
||||
}
|
||||
|
||||
awaitNudge(wiring.leadHerdr());
|
||||
Thread.sleep(200); // settle — nothing more should arrive beyond the one coalesced nudge
|
||||
assertEquals(1, wiring.scheduler().runDueTasks(),
|
||||
"both terminal tickets must coalesce onto one scheduled tick");
|
||||
long nudgeCount = wiring.leadHerdr().calls.stream()
|
||||
.filter(c -> c.method().equals("agent.prompt")).count();
|
||||
assertEquals(1, nudgeCount, "two tickets finishing together must produce ONE nudge, not two");
|
||||
@@ -2055,7 +2111,8 @@ class MessageServiceTest {
|
||||
|
||||
@Test
|
||||
void answeringAQuestionStopsFurtherNudgesAboutIt() throws Exception {
|
||||
try (var wiring = wireWithPushLoop(5, 50)) {
|
||||
// A 1ms backoff is due immediately, but ManualScheduler only ticks when this test asks it to.
|
||||
try (var wiring = wireWithManualScheduler(5, 1)) {
|
||||
String ticket = wiring.service().sendAsync(T, "task that asks");
|
||||
awaitWaiting();
|
||||
injectDelivery();
|
||||
@@ -2063,8 +2120,9 @@ class MessageServiceTest {
|
||||
CompletableFuture<MessageService.AskResult> ask = CompletableFuture.supplyAsync(
|
||||
() -> wiring.service().ask(T, "which config file?", 5000));
|
||||
MessageService.TaskView asking = awaitTicketPhaseOn(wiring.service(), ticket, MessageService.Phase.ASKING);
|
||||
awaitQuestionPendingOn(wiring.pushLoop(), LEAD, asking.turnId());
|
||||
|
||||
awaitNudge(wiring.leadHerdr());
|
||||
assertEquals(1, wiring.scheduler().runDueTasks(), "the open question must have one scheduled tick");
|
||||
long callsBeforeAnswer = wiring.leadHerdr().calls.stream()
|
||||
.filter(c -> c.method().equals("agent.prompt")).count();
|
||||
|
||||
@@ -2072,12 +2130,22 @@ class MessageServiceTest {
|
||||
() -> wiring.service().answer(asking.turnId(), "config.yaml", 5000));
|
||||
assertEquals("config.yaml", ask.get(5, TimeUnit.SECONDS).answer());
|
||||
awaitWaiting();
|
||||
java.util.concurrent.CountDownLatch terminalReached = new java.util.concurrent.CountDownLatch(1);
|
||||
wiring.service().setAfterFinishAsyncTaskCompleteHookForTest(terminalReached::countDown);
|
||||
assertTrue(rendezvous.resolve(T, "done"));
|
||||
try {
|
||||
assertTrue(terminalReached.await(5, TimeUnit.SECONDS),
|
||||
"the answered ticket never reached its terminal phase");
|
||||
} finally {
|
||||
wiring.service().setAfterFinishAsyncTaskCompleteHookForTest(null);
|
||||
}
|
||||
answer.get(5, TimeUnit.SECONDS);
|
||||
|
||||
// Let several more ticks (and the ticket's own now-legitimate terminal nudge) fire —
|
||||
// none of them may still name the question's turnId, which is closed.
|
||||
Thread.sleep(300);
|
||||
// Run the ticket's legitimate terminal nudge and every later scheduled tick through its
|
||||
// reminder cap. None may still name the closed question.
|
||||
for (int tick = 0; tick < 6; tick++) {
|
||||
assertEquals(1, wiring.scheduler().runDueTasks(), "expected one scheduled reminder tick");
|
||||
}
|
||||
boolean anyNamesClosedQuestion = wiring.leadHerdr().calls.stream()
|
||||
.filter(c -> c.method().equals("agent.prompt"))
|
||||
.skip(callsBeforeAnswer)
|
||||
@@ -2160,35 +2228,45 @@ class MessageServiceTest {
|
||||
// decideTickets hits the cap and STOPs — activeLeads drops the lead, but (before the fix)
|
||||
// pendingTickets never drops the ticket. That is the exact "cap already STOPped" branch of
|
||||
// the bug report, reached deterministically rather than by timing it against a live tick.
|
||||
try (var wiring = wireWithPushLoop(1, 50, clock::get)) {
|
||||
String stale = wiring.service().sendAsync(T, "first task");
|
||||
awaitWaiting();
|
||||
injectDelivery();
|
||||
assertTrue(rendezvous.resolve(T, "stale result"));
|
||||
// A 1ms backoff is due immediately, but ManualScheduler runs only the ticks below.
|
||||
try (var wiring = wireWithManualScheduler(1, 1, clock::get)) {
|
||||
String stale;
|
||||
String fresh;
|
||||
java.util.concurrent.CountDownLatch staleTerminalReached = new java.util.concurrent.CountDownLatch(1);
|
||||
wiring.service().setAfterFinishAsyncTaskCompleteHookForTest(staleTerminalReached::countDown);
|
||||
try {
|
||||
stale = wiring.service().sendAsync(T, "first task");
|
||||
awaitWaiting();
|
||||
injectDelivery();
|
||||
assertTrue(rendezvous.resolve(T, "stale result"));
|
||||
assertTrue(staleTerminalReached.await(5, TimeUnit.SECONDS),
|
||||
"the stale ticket never reached its terminal phase");
|
||||
|
||||
// Let the reminder loop fire its one nudge and hit the cap (STOP removes it from
|
||||
// activeLeads; pendingTickets is untouched either way — that asymmetry is the bug).
|
||||
awaitNudge(wiring.leadHerdr());
|
||||
Thread.sleep(300);
|
||||
assertTrue(wiring.leadHerdr().lastCall("agent.prompt").params().toString().contains(stale),
|
||||
"sanity: the stale ticket's own reminder must have fired first");
|
||||
// Fire the stale ticket's one nudge, then its cap tick (STOP removes it from activeLeads;
|
||||
// pendingTickets is untouched either way — that asymmetry is the bug).
|
||||
assertEquals(1, wiring.scheduler().runDueTasks(), "the stale ticket must have one nudge tick");
|
||||
assertEquals(1, wiring.scheduler().runDueTasks(), "the stale ticket must have one cap tick");
|
||||
assertTrue(wiring.leadHerdr().lastCall("agent.prompt").params().toString().contains(stale),
|
||||
"sanity: the stale ticket's own reminder must have fired first");
|
||||
|
||||
// Cross the TTL — a real clock would need 10 minutes; the injected one does it instantly.
|
||||
clock.addAndGet(MessageService.TICKET_TTL_NANOS + TimeUnit.SECONDS.toNanos(1));
|
||||
// Cross the TTL — a real clock would need 10 minutes; the injected one does it instantly.
|
||||
clock.addAndGet(MessageService.TICKET_TTL_NANOS + TimeUnit.SECONDS.toNanos(1));
|
||||
|
||||
// A second, unrelated ticket to the same target/lead reaches sendAsync, which prunes.
|
||||
String fresh = wiring.service().sendAsync(T, "second task");
|
||||
awaitWaiting();
|
||||
injectDelivery();
|
||||
assertTrue(rendezvous.resolve(T, "fresh result"));
|
||||
|
||||
// The fresh ticket restarts the (now-dormant) reminder loop with its own nudge.
|
||||
long before = wiring.leadHerdr().calls.stream().filter(c -> c.method().equals("agent.prompt")).count();
|
||||
long deadline = System.currentTimeMillis() + 3000;
|
||||
while (wiring.leadHerdr().calls.stream().filter(c -> c.method().equals("agent.prompt")).count() <= before
|
||||
&& System.currentTimeMillis() < deadline) {
|
||||
Thread.sleep(10);
|
||||
// A second, unrelated ticket to the same target/lead reaches sendAsync, which prunes.
|
||||
java.util.concurrent.CountDownLatch freshTerminalReached = new java.util.concurrent.CountDownLatch(1);
|
||||
wiring.service().setAfterFinishAsyncTaskCompleteHookForTest(freshTerminalReached::countDown);
|
||||
fresh = wiring.service().sendAsync(T, "second task");
|
||||
awaitWaiting();
|
||||
injectDelivery();
|
||||
assertTrue(rendezvous.resolve(T, "fresh result"));
|
||||
assertTrue(freshTerminalReached.await(5, TimeUnit.SECONDS),
|
||||
"the fresh ticket never reached its terminal phase");
|
||||
} finally {
|
||||
wiring.service().setAfterFinishAsyncTaskCompleteHookForTest(null);
|
||||
}
|
||||
|
||||
// The fresh ticket restarts the now-dormant reminder loop with its own nudge.
|
||||
assertEquals(1, wiring.scheduler().runDueTasks(), "the fresh ticket must have one nudge tick");
|
||||
String latestNudge = wiring.leadHerdr().lastCall("agent.prompt").params().toString();
|
||||
assertTrue(latestNudge.contains(fresh), "the fresh ticket's nudge must still arrive: " + latestNudge);
|
||||
assertFalse(latestNudge.contains(stale),
|
||||
|
||||
@@ -122,12 +122,28 @@ class FleetAppAuthTest {
|
||||
assertEquals(Authz.Action.DRAIN, FleetApp.routeAction("GET /sessions/{id}/replies"));
|
||||
assertEquals(Authz.Action.ASK, FleetApp.routeAction("POST /sessions/{id}/ask"));
|
||||
for (String route : Set.of("GET /sessions", "GET /agents", "GET /members", "GET /profiles",
|
||||
"GET /member-credentials", "GET /sessions/{id}/status", "GET /tasks/{ticket}")) {
|
||||
"GET /member-credentials")) {
|
||||
assertEquals(Authz.Action.READ, FleetApp.routeAction(route), route);
|
||||
}
|
||||
for (String route : Set.of("GET /sessions/{id}/status", "GET /tasks/{ticket}")) {
|
||||
assertEquals(Authz.Action.TASK_READ, FleetApp.routeAction(route), route);
|
||||
}
|
||||
assertThrows(IllegalArgumentException.class, () -> FleetApp.routeAction("GET /healthz"));
|
||||
}
|
||||
|
||||
/**
|
||||
* fleetd #669 Unit A: {@code POST /sessions/{id}/message} is two call shapes behind one route,
|
||||
* mirroring {@code fleet_send}'s MCP-side split into {@link Authz.Action#SEND} and {@link
|
||||
* Authz.Action#ANSWER} ({@code FleetMcp#sendAction}). The route never carries a {@code coordId}
|
||||
* shape — that peer-lead route is MCP-only — so only these two apply here.
|
||||
*/
|
||||
@Test
|
||||
void theMessageRouteIsASendWithNoTurnIdAndAnAnswerWithOne() {
|
||||
assertEquals(Authz.Action.SEND, FleetApp.routeAction("POST /sessions/{id}/message", null));
|
||||
assertEquals(Authz.Action.SEND, FleetApp.routeAction("POST /sessions/{id}/message", " "));
|
||||
assertEquals(Authz.Action.ANSWER, FleetApp.routeAction("POST /sessions/{id}/message", "turn-1"));
|
||||
}
|
||||
|
||||
private static Set<String> routesTheServerRegisters() {
|
||||
try {
|
||||
String source = Files.readString(REST_SOURCE).lines()
|
||||
@@ -192,6 +208,21 @@ class FleetAppAuthTest {
|
||||
"draining an inbox is the primary's collection step");
|
||||
}
|
||||
|
||||
/**
|
||||
* fleetd #669 Unit A: the {@code turnId} shape of {@code POST /sessions/{id}/message} maps to
|
||||
* {@link Authz.Action#ANSWER}, not the plain {@link Authz.Action#SEND} the test above drives —
|
||||
* a worker must stay refused on this shape too, exactly as it was refused on the one undivided
|
||||
* action before the split.
|
||||
*/
|
||||
@Test
|
||||
void aWorkerMayNotAnswerAnotherSessionsBlockedQuestionOverRest() throws Exception {
|
||||
int port = start(FakeHerdr.WORKER_PID, false, null);
|
||||
|
||||
assertEquals(403, send(port, "POST", "/sessions/term_b/message",
|
||||
"{\"turnId\":\"turn-1\",\"content\":\"hi\"}", null).statusCode(),
|
||||
"resolving another session's blocked question would be a worker escalating too");
|
||||
}
|
||||
|
||||
// --- token mode ---------------------------------------------------------------------------
|
||||
|
||||
@Test
|
||||
|
||||
Executable
+855
@@ -0,0 +1,855 @@
|
||||
#!/usr/bin/env bash
|
||||
#
|
||||
# The one auditable way to edit the live fleetd.yaml.
|
||||
#
|
||||
# `--set` uses yq and rewrites the whole YAML document in yq's output style. Use `--from` for a
|
||||
# candidate whose comment alignment or other formatting carries meaning: it copies that file
|
||||
# verbatim while keeping this script's backup, parse check, atomic install, and verdict read-back.
|
||||
#
|
||||
# fleetd ticket #635 — why this exists at all: fleetd.yaml is gitignored and holds the live
|
||||
# fleet's settings. A bad raw edit reaches a daemon that is already serving, so a direct `Edit`
|
||||
# on it is refused by policy. This script is the allow-listed alternative, and it is not just
|
||||
# convenience — it is the thing a raw file write can never give you: a backup, a parse check
|
||||
# BEFORE the file is installed, and the daemon's own reload verdict read back afterwards. An
|
||||
# edit to a live config is not finished when the bytes are written. It is finished when the
|
||||
# daemon has said what it did with them.
|
||||
#
|
||||
# What the daemon says, and how this script finds it — measured against `ConfigRef.java` on
|
||||
# fleetd commit 158a2a8, 2026-10-01:
|
||||
#
|
||||
# 1. `ConfigRef` re-reads fleetd.yaml only when the WATCHER sees the mtime move (every 10s by
|
||||
# default — read the real interval out of the daemon's own startup line, "config watch: ...
|
||||
# re-read when it changes (every Ns)"). So a verdict never appears before the next tick.
|
||||
# 2. `ConfigRef.Outcome.summary()` logs exactly one of five strings (ConfigRef.java:371-391):
|
||||
# config reload refused — <error message>
|
||||
# config reload refused — these keys cannot change under a running daemon: <keys>. ...
|
||||
# config reloaded
|
||||
# config reloaded; these changes need a restart to take effect: <keys>
|
||||
# config reloaded; partially live — <key: detail | ...>
|
||||
# A parse/validation failure logs a DIFFERENT line instead, before any summary ever runs
|
||||
# (ConfigRef.java:425): "config reload from <path> refused, keeping the running config:
|
||||
# <message>". This script recognises both shapes of refusal.
|
||||
# 3. The em dash in those strings is a real multi-byte character — match the stable prefix
|
||||
# "config reload refused" (or "...refused, keeping the running config" for the parse-failure
|
||||
# shape), never the dash itself.
|
||||
# 4. A cold-key change (bind/herdrSocket/memberHerdrSocket/broker/auth) throws away the WHOLE
|
||||
# reload — the running config keeps every old value, not only the cold one.
|
||||
# 5. A deferred/split change IS applied (current.set(fresh) runs) — "needs a restart" is a
|
||||
# SUCCESS with a follow-up, never a failure.
|
||||
#
|
||||
# Four outcomes, and they stay four (see the exit code table below). The one most likely to be
|
||||
# gotten wrong is "cannot tell" (exit 5): the daemon may be down, or the watcher may be stalled,
|
||||
# and folding that into either "refused" or "applied" is worse than never checking at all,
|
||||
# because a caller then acts on a verdict nobody actually read. So exit 5 never restores — a
|
||||
# visible, recoverable edit beats an invisible revert of a GOOD edit.
|
||||
#
|
||||
# Usage:
|
||||
# scripts/config-edit.sh --check
|
||||
# scripts/config-edit.sh --set <yq-path>=<value> [--set ...]
|
||||
# scripts/config-edit.sh --from <candidate.yaml>
|
||||
# scripts/config-edit.sh --dry-run --set <yq-path>=<value>
|
||||
# scripts/config-edit.sh --restore
|
||||
#
|
||||
# `--set` rewrites the whole file in yq's output style, not only the requested keys. The script
|
||||
# warns before installation when the candidate changes more lines than its number of --set pairs.
|
||||
# Use `--from <candidate.yaml>` when comment alignment or other formatting is meaningful: --from
|
||||
# copies the candidate verbatim, with no yq round-trip.
|
||||
#
|
||||
# `--set .a.b=` (an empty value — a forgotten typo) is REFUSED, not accepted as "clear the
|
||||
# field": a null value falls back to its default rather than erroring, which is silent, not
|
||||
# safe. To clear a key on purpose, write a literal null: `--set .a.b=null`. Every other value
|
||||
# is always written as a YAML string (via yq's strenv(), never spliced into the expression), so
|
||||
# there is currently no --set spelling for the literal three-character STRING "null" itself — use
|
||||
# --from for that rare case.
|
||||
#
|
||||
# Overrides (so this is drivable with no daemon — see scripts/test-config-edit.sh):
|
||||
# --config <path> default: fleetd/fleetd.yaml
|
||||
# --log <path> default: fleetd/fleetd.out
|
||||
# --wait-seconds <n> default: 4x the watch interval this script reads out of --log (10 -> 40)
|
||||
#
|
||||
# Exit codes (the --check/--restore/usage-error paths are reported separately, see below):
|
||||
# 0 applied; verdict read; clean
|
||||
# 3 applied; verdict read; needs a restart (deferred or split keys named)
|
||||
# 4 REFUSED by the daemon; backup restored (and the restore's own verdict reported if seen)
|
||||
# 5 CANNOT TELL — no verdict line inside the wait window. Nothing is restored.
|
||||
#
|
||||
# Never prints a secret. fleetd.yaml keeps credentials out by indirection (broker.uriEnv,
|
||||
# gitTokenEnv) but this script does not rely on that staying true: every diff it prints is piped
|
||||
# through `redact`, which (a) blanks the userinfo of any `scheme://user:pass@host` and (b) masks
|
||||
# the whole value on any line whose key looks like a credential. See `redact` below.
|
||||
|
||||
set -euo pipefail
|
||||
|
||||
REPO="$(cd "$(dirname "${BASH_SOURCE[0]}")/.." && pwd)"
|
||||
SELF="$REPO/scripts/config-edit.sh"
|
||||
|
||||
CONFIG="$REPO/fleetd/fleetd.yaml"
|
||||
LOG="$REPO/fleetd/fleetd.out"
|
||||
WAIT_SECONDS_OVERRIDE=""
|
||||
FALLBACK_PORT=8765
|
||||
|
||||
MODE=""
|
||||
DRY_RUN=0
|
||||
SETS=()
|
||||
FROM_FILE=""
|
||||
|
||||
# fleetd #635 follow-up — a signal (or any early exit while a candidate is still uninstalled) must
|
||||
# not leave a `.config-edit.XXXXXX` file sitting beside the live config forever. CAND is global
|
||||
# (never a function-local) on purpose: this ONE trap, set once, covers every path that ever
|
||||
# creates a candidate — run_edit and dry_run_diff both assign it, and clear it back to "" once the
|
||||
# file is consumed (installed, or explicitly removed), so a later, unrelated exit never retries a
|
||||
# path that already served its purpose.
|
||||
CAND=""
|
||||
cleanup_candidate() { [ -n "$CAND" ] && rm -f "$CAND" 2>/dev/null; return 0; }
|
||||
trap cleanup_candidate EXIT INT TERM
|
||||
|
||||
say() { printf '\n\033[1m== %s\033[0m\n' "$*"; }
|
||||
ok() { printf ' ok %s\n' "$*"; }
|
||||
warn() { printf ' WARN %s\n' "$*"; }
|
||||
die() { printf '\n FAIL %s\n\n' "$*" >&2; exit 1; }
|
||||
|
||||
set_mode() {
|
||||
local new="$1"
|
||||
if [ -n "$MODE" ] && [ "$MODE" != "$new" ]; then
|
||||
die "cannot combine --$MODE and --$new in one invocation"
|
||||
fi
|
||||
MODE="$new"
|
||||
}
|
||||
|
||||
while [ $# -gt 0 ]; do
|
||||
case "$1" in
|
||||
--check) set_mode check; shift ;;
|
||||
--restore) set_mode restore; shift ;;
|
||||
--set)
|
||||
[ $# -ge 2 ] || die "--set requires <yq-path>=<value>"
|
||||
set_mode set
|
||||
SETS+=("$2")
|
||||
shift 2 ;;
|
||||
--from)
|
||||
[ $# -ge 2 ] || die "--from requires a candidate file path"
|
||||
set_mode from
|
||||
FROM_FILE="$2"
|
||||
shift 2 ;;
|
||||
--dry-run) DRY_RUN=1; shift ;;
|
||||
--config)
|
||||
[ $# -ge 2 ] || die "--config requires a path"
|
||||
CONFIG="$2"; shift 2 ;;
|
||||
--log)
|
||||
[ $# -ge 2 ] || die "--log requires a path"
|
||||
LOG="$2"; shift 2 ;;
|
||||
--wait-seconds)
|
||||
[ $# -ge 2 ] || die "--wait-seconds requires a number of seconds"
|
||||
WAIT_SECONDS_OVERRIDE="$2"; shift 2 ;;
|
||||
-h|--help) sed -n '3,70p' "$SELF"; exit 0 ;;
|
||||
*) echo "unknown option: $1 (try --help)" >&2; exit 2 ;;
|
||||
esac
|
||||
done
|
||||
|
||||
[ -n "$MODE" ] || die "no action given — use --check, --set, --from, or --restore (see --help)"
|
||||
|
||||
# ------------------------------------------------------------------------------------- redaction
|
||||
#
|
||||
# Two independent passes, applied to every diff this script ever prints:
|
||||
# 1. `scheme://user:pass@host` -> `scheme://<redacted>@host`, globally (the `g` flag matters —
|
||||
# a line can carry more than one URI).
|
||||
# 2. Any line whose key looks like TOKEN|SECRET|PASSWORD|PASSWD|PASSPHRASE|CREDENTIAL|URI|_KEY,
|
||||
# matched case-insensitively against the key text (uriEnv, gitTokenEnv, ... are camelCase,
|
||||
# not SCREAMING_CASE) has its whole value blanked, diff marker and indentation kept so the
|
||||
# shape of the change is still visible. Deliberately conservative: a false-positive
|
||||
# redaction on an unrelated line costs nothing, an unredacted secret is a security defect
|
||||
# (acceptance criterion 7).
|
||||
#
|
||||
# fleetd #635 follow-up (ticket comment 17670, defect 7) — a masked key line is not the whole
|
||||
# story: a YAML block scalar (`|`, `|-`, `>`, `>-`, ...) puts the VALUE on the lines that follow
|
||||
# the key, each indented deeper than it. The key-name match above only ever sees the key line
|
||||
# itself, so those continuation lines used to flow straight through unredacted while the key line
|
||||
# right above them printed a reassuring "<redacted>". The redactor maps masked continuation lines
|
||||
# from each complete file before it reads the diff. It then masks a printed line when that file
|
||||
# line is inside a masked key's value. This covers block-scalar bodies even when the key line is
|
||||
# outside the printed hunk, and it keeps blank lines inside the value masked.
|
||||
#
|
||||
# `redact` is always fed `diff -u` output, and every line of a unified diff starts with exactly
|
||||
# one of ' ', '+', '-' (the three body markers; '@'/'-'/'+' for the three header-line kinds too).
|
||||
# That one leading character is NOT part of the YAML indentation, and must be stripped before
|
||||
# indentation is measured or a key is matched — otherwise a changed ('+' or '-') line reads one
|
||||
# column shallower than it really is, and either wrongly escapes a continuation mask or wrongly
|
||||
# ends one early. Tabs are out of scope: YAML forbids them for indentation, and this is a bounded
|
||||
# fix, not a YAML parser.
|
||||
map_masked_lines() {
|
||||
local file="$1" side="$2" line content indent lead key line_number=0
|
||||
local masked=0 masked_indent=0
|
||||
|
||||
case "$side" in
|
||||
old) OLD_MASKED_LINES=() ;;
|
||||
new) NEW_MASKED_LINES=() ;;
|
||||
*) die "internal error: unknown redaction map side $side" ;;
|
||||
esac
|
||||
|
||||
while IFS= read -r line || [ -n "$line" ]; do
|
||||
line_number=$((line_number + 1))
|
||||
content="$line"
|
||||
indent=0
|
||||
while [ "${content:$indent:1}" = " " ]; do indent=$((indent + 1)); done
|
||||
|
||||
if [ "$masked" = 1 ]; then
|
||||
if [ -z "${content// /}" ] || [ "$indent" -gt "$masked_indent" ]; then
|
||||
case "$side" in
|
||||
old) OLD_MASKED_LINES[$line_number]=1 ;;
|
||||
new) NEW_MASKED_LINES[$line_number]=1 ;;
|
||||
esac
|
||||
continue
|
||||
fi
|
||||
masked=0
|
||||
fi
|
||||
|
||||
if [[ "$content" =~ ^([[:space:]]*)([A-Za-z0-9_.-]+:) ]]; then
|
||||
lead="${BASH_REMATCH[1]}"
|
||||
key="${BASH_REMATCH[2]}"
|
||||
if [[ "$key" =~ (TOKEN|SECRET|PASSWORD|PASSWD|PASSPHRASE|CREDENTIAL|URI|_KEY) ]]; then
|
||||
masked=1
|
||||
masked_indent="$indent"
|
||||
fi
|
||||
fi
|
||||
done < "$file"
|
||||
}
|
||||
|
||||
redact() {
|
||||
local old_file="$1" new_file="$2"
|
||||
local line prefix content indent lead key old_line=0 new_line=0 in_hunk=0
|
||||
local old_masked new_masked saved_nocasematch=0
|
||||
shopt -q nocasematch && saved_nocasematch=1
|
||||
shopt -s nocasematch
|
||||
|
||||
map_masked_lines "$old_file" old
|
||||
map_masked_lines "$new_file" new
|
||||
|
||||
while IFS= read -r line || [ -n "$line" ]; do
|
||||
if [[ "$line" =~ ^@@\ -([0-9]+)(,([0-9]+))?\ \+([0-9]+)(,([0-9]+))?\ @@ ]]; then
|
||||
old_line="${BASH_REMATCH[1]}"
|
||||
new_line="${BASH_REMATCH[4]}"
|
||||
in_hunk=1
|
||||
printf '%s\n' "$line"
|
||||
continue
|
||||
fi
|
||||
|
||||
case "$line" in
|
||||
[\ +-]*) prefix="${line:0:1}"; content="${line:1}" ;;
|
||||
*) prefix=""; content="$line" ;;
|
||||
esac
|
||||
|
||||
old_masked=0
|
||||
new_masked=0
|
||||
if [ "$in_hunk" = 1 ]; then
|
||||
case "$prefix" in
|
||||
' ')
|
||||
[ "${OLD_MASKED_LINES[$old_line]:-}" = 1 ] && old_masked=1
|
||||
[ "${NEW_MASKED_LINES[$new_line]:-}" = 1 ] && new_masked=1
|
||||
old_line=$((old_line + 1)); new_line=$((new_line + 1)) ;;
|
||||
-)
|
||||
[ "${OLD_MASKED_LINES[$old_line]:-}" = 1 ] && old_masked=1
|
||||
old_line=$((old_line + 1)) ;;
|
||||
+)
|
||||
[ "${NEW_MASKED_LINES[$new_line]:-}" = 1 ] && new_masked=1
|
||||
new_line=$((new_line + 1)) ;;
|
||||
esac
|
||||
fi
|
||||
|
||||
indent=0
|
||||
while [ "${content:$indent:1}" = " " ]; do indent=$((indent + 1)); done
|
||||
|
||||
if [ "$old_masked" = 1 ] || [ "$new_masked" = 1 ]; then
|
||||
printf '%s%*s<redacted>\n' "$prefix" "$indent" ""
|
||||
continue
|
||||
fi
|
||||
|
||||
if [[ "$content" =~ ^([[:space:]]*)([A-Za-z0-9_.-]+:) ]]; then
|
||||
lead="${BASH_REMATCH[1]}"
|
||||
key="${BASH_REMATCH[2]}"
|
||||
if [[ "$key" =~ (TOKEN|SECRET|PASSWORD|PASSWD|PASSPHRASE|CREDENTIAL|URI|_KEY) ]]; then
|
||||
printf '%s%s%s <redacted>\n' "$prefix" "$lead" "$key"
|
||||
continue
|
||||
fi
|
||||
fi
|
||||
printf '%s\n' "$line" | sed -E 's#://[^@]*@#://<redacted>@#g'
|
||||
done
|
||||
[ "$saved_nocasematch" = 1 ] || shopt -u nocasematch
|
||||
}
|
||||
|
||||
# ------------------------------------------------------------------------------------- the probe
|
||||
#
|
||||
# Probe the SOCKET, never `pgrep`/`ps -f` — both print argv, and argv holds `NAME=value`, making
|
||||
# either a credential channel. The port comes from the config's own `bind.port`; 8765 is only a
|
||||
# fallback when that key is absent or the file does not parse yet.
|
||||
resolve_port() {
|
||||
local file="$1" port
|
||||
if [ -f "$file" ] && command -v yq >/dev/null 2>&1; then
|
||||
port="$(yq eval '.bind.port' "$file" 2>/dev/null || true)"
|
||||
else
|
||||
port=""
|
||||
fi
|
||||
case "$port" in
|
||||
''|null) echo "$FALLBACK_PORT" ;;
|
||||
*) echo "$port" ;;
|
||||
esac
|
||||
}
|
||||
|
||||
daemon_listening() {
|
||||
local port="$1"
|
||||
if command -v nc >/dev/null 2>&1; then
|
||||
nc -z -w1 127.0.0.1 "$port" 2>/dev/null
|
||||
else
|
||||
( exec 3<>"/dev/tcp/127.0.0.1/$port" ) 2>/dev/null
|
||||
fi
|
||||
}
|
||||
|
||||
# --------------------------------------------------------------------------------- the log marker
|
||||
#
|
||||
# Take the log's line count BEFORE touching anything. Every later read of "what did the daemon
|
||||
# say" starts strictly after this mark, so a refusal from hours ago can never be mistaken for
|
||||
# this edit's verdict. Same approach as scripts/redeploy-fleetd.sh's RESTART_MARK.
|
||||
log_mark() {
|
||||
local file="$1"
|
||||
if [ -f "$file" ]; then
|
||||
wc -l < "$file" 2>/dev/null || echo 0
|
||||
else
|
||||
echo 0
|
||||
fi
|
||||
}
|
||||
|
||||
read_verdict_after_marker() {
|
||||
local file="$1" mark="$2"
|
||||
[ -f "$file" ] || return 0
|
||||
tail -n "+$((mark + 1))" "$file" 2>/dev/null || true
|
||||
}
|
||||
|
||||
# Classifies one log LINE. Echoes one of: refused | clean | needs-restart | none. Always
|
||||
# succeeds (every branch ends in `echo`), so it is safe to call from inside `$( )`.
|
||||
classify_verdict_line() {
|
||||
local line="$1"
|
||||
case "$line" in
|
||||
*'config reload refused'*) echo refused ;;
|
||||
*'config reload from '*'refused, keeping the running config'*) echo refused ;;
|
||||
*'config reloaded'*)
|
||||
case "$line" in
|
||||
*'need a restart'*|*'partially live'*) echo needs-restart ;;
|
||||
*) echo clean ;;
|
||||
esac ;;
|
||||
*) echo none ;;
|
||||
esac
|
||||
return 0
|
||||
}
|
||||
|
||||
scan_region_for_verdict() {
|
||||
local region="$1" line kind
|
||||
[ -n "$region" ] || return 1
|
||||
while IFS= read -r line || [ -n "$line" ]; do
|
||||
kind="$(classify_verdict_line "$line")"
|
||||
if [ "$kind" != "none" ]; then
|
||||
VERDICT_KIND="$kind"
|
||||
VERDICT_LINE="$line"
|
||||
return 0
|
||||
fi
|
||||
done <<< "$region"
|
||||
return 1
|
||||
}
|
||||
|
||||
# Sets VERDICT_KIND/VERDICT_LINE and returns 0 on the first verdict line found after $mark;
|
||||
# returns 1 (VERDICT_KIND=none) if none appeared inside $wait_s seconds. Checks once before each
|
||||
# sleep AND once more after the last sleep, the same boundary idiom
|
||||
# scripts/redeploy-fleetd.sh's wait_for_daemon_exit/wait_for_new_pid already use.
|
||||
wait_for_verdict() {
|
||||
local log="$1" mark="$2" wait_s="$3" _i region
|
||||
VERDICT_KIND="none"
|
||||
VERDICT_LINE=""
|
||||
for _i in $(seq "$wait_s"); do
|
||||
region="$(read_verdict_after_marker "$log" "$mark")"
|
||||
scan_region_for_verdict "$region" && return 0
|
||||
sleep 1
|
||||
done
|
||||
region="$(read_verdict_after_marker "$log" "$mark")"
|
||||
scan_region_for_verdict "$region" && return 0
|
||||
return 1
|
||||
}
|
||||
|
||||
last_verdict_line() {
|
||||
local file="$1" line out=""
|
||||
[ -f "$file" ] || return 0
|
||||
while IFS= read -r line || [ -n "$line" ]; do
|
||||
if [ "$(classify_verdict_line "$line")" != "none" ]; then
|
||||
out="$line"
|
||||
fi
|
||||
done < "$file"
|
||||
printf '%s' "$out"
|
||||
}
|
||||
|
||||
default_wait_seconds() {
|
||||
local log="$1" interval=""
|
||||
if [ -f "$log" ]; then
|
||||
interval="$(grep -F 'config watch:' "$log" 2>/dev/null | tail -1 \
|
||||
| sed -E 's/.*\(every ([0-9]+)s\).*/\1/' || true)"
|
||||
fi
|
||||
case "$interval" in
|
||||
''|*[!0-9]*) interval=10 ;;
|
||||
esac
|
||||
echo $((interval * 4))
|
||||
}
|
||||
|
||||
# ----------------------------------------------------------------------------------- the backup
|
||||
#
|
||||
# Timestamped, never pruned — "keep backups" per the ticket. A pid suffix avoids a same-second
|
||||
# collision between two invocations.
|
||||
#
|
||||
# fleetd #635 follow-up — lands under a DEDICATED, gitignored directory beside the config
|
||||
# (<dir>/.config-backups/), never beside the config file itself. The whole reason fleetd.yaml is
|
||||
# gitignored is that it must never be committed, and a backup of it inherits that requirement — a
|
||||
# bare `fleetd.yaml.bak.*` next to a tracked directory is one `git add -A`/`git add .` away from
|
||||
# committing the live config. A directory beats a glob on its own: the glob only protects today's
|
||||
# naming, a location keeps working even if the naming changes later. (See .gitignore for the glob
|
||||
# kept anyway, as a backstop for a stray backup written the old way.)
|
||||
BACKUP_DIRNAME=".config-backups"
|
||||
|
||||
backup_dir_for() {
|
||||
local src="$1"
|
||||
printf '%s/%s' "$(dirname "$src")" "$BACKUP_DIRNAME"
|
||||
}
|
||||
|
||||
backup_config() {
|
||||
local src="$1" ts backup dir base
|
||||
dir="$(backup_dir_for "$src")"
|
||||
mkdir -p "$dir" \
|
||||
|| die "could not create the backup directory $dir — refusing to edit without a backup. The live config at $src was NOT touched."
|
||||
ts="$(date -u +%Y%m%dT%H%M%S)Z"
|
||||
base="$(basename "$src")"
|
||||
backup="${dir}/${base}.bak.${ts}.$$"
|
||||
cp "$src" "$backup" \
|
||||
|| die "could not create a backup at $backup — refusing to edit without one. The live config at $src was NOT touched."
|
||||
printf '%s' "$backup"
|
||||
}
|
||||
|
||||
newest_backup() {
|
||||
local cfg="$1" dir base
|
||||
dir="$(backup_dir_for "$cfg")"
|
||||
base="$(basename "$cfg")"
|
||||
ls -t "${dir}/${base}".bak.* 2>/dev/null | head -1 || true
|
||||
}
|
||||
|
||||
# ------------------------------------------------------------------------------------ the file mode
|
||||
#
|
||||
# fleetd #635 follow-up — `mv` from a mktemp candidate carries mktemp's 0600 onto the live path
|
||||
# forever (measured: 644 -> 600 after one --set), and a restore does not undo it either, because
|
||||
# `cp` onto an EXISTING file keeps the DESTINATION's mode, not the source's. Capture the live
|
||||
# file's mode before anything touches it, and reapply it to whatever lands on that path
|
||||
# afterwards — the candidate before install, and the config again after a restore — so an edit
|
||||
# changes the file's CONTENT only, never its permissions. BSD `stat -f '%Lp'` first (matches this
|
||||
# project's dev machine), GNU `stat -c '%a'` as the fallback. Prints nothing when the file does
|
||||
# not exist yet, so apply_mode then does nothing and a first-ever edit falls back to the normal
|
||||
# umask default rather than inventing a number.
|
||||
file_mode() {
|
||||
local file="$1"
|
||||
[ -f "$file" ] || return 0
|
||||
stat -f '%Lp' "$file" 2>/dev/null || stat -c '%a' "$file" 2>/dev/null || true
|
||||
}
|
||||
|
||||
apply_mode() {
|
||||
local file="$1" mode="$2"
|
||||
[ -n "$mode" ] || return 0
|
||||
chmod "$mode" "$file" 2>/dev/null || true
|
||||
}
|
||||
|
||||
# --------------------------------------------------------------------------- candidate builders
|
||||
#
|
||||
# Never edit the live file in place. Each builder fills $1 (a temp file already sitting in the
|
||||
# SAME directory as the live config, so the later `mv` install is a rename, not a cross-device
|
||||
# copy — see run_edit).
|
||||
# fleetd #635 follow-up — a forgotten value (`--set .a.b=`, a plausible typo) must never be
|
||||
# accepted as "clear the field". `*=*` alone cannot tell "--set .a.b=" from "--set .a.b=7" apart
|
||||
# — both contain an `=` — so the guard has to look at the VALUE, not the shape of the argument.
|
||||
# An empty value refuses outright: nothing is installed, and the message names the likely cause
|
||||
# AND the two ways to actually mean it (clear on purpose, or an intentional empty string via
|
||||
# --from). Measured against the real daemon loader: a quoted empty string reads back as a null
|
||||
# field (`quoted empty -> OK int=null`), and a null numeric field FALLS BACK TO ITS DEFAULT rather
|
||||
# than erroring — so this is not a cosmetic nit, it is the one shape of edit that widens capacity
|
||||
# silently instead of failing loudly, which is exactly what this script exists to catch.
|
||||
#
|
||||
# A deliberate clear needs its own spelling, because `""` and YAML `null` are NOT the same value
|
||||
# to the loader (`""` is a valid empty String; `null` means absent, and an Integer field reads
|
||||
# either the same way — null — but a String field would keep `""` as a real value). `--set
|
||||
# .a.b=null` is that spelling: it writes a literal, unquoted `null` via yq, never the string
|
||||
# "null" through strenv(). One consequence worth knowing: there is currently no --set spelling
|
||||
# for the three-character STRING "null" itself (it collides with the clear spelling) — use
|
||||
# --from for that rare case.
|
||||
apply_set_pairs() {
|
||||
local cand="$1" kv path value
|
||||
shift
|
||||
for kv in "$@"; do
|
||||
case "$kv" in
|
||||
*=*) : ;;
|
||||
*) die "--set expects <yq-path>=<value>, got: '$kv'" ;;
|
||||
esac
|
||||
path="${kv%%=*}"
|
||||
path="${path#.}"
|
||||
value="${kv#*=}"
|
||||
if [ -z "$value" ]; then
|
||||
die "--set '$kv' has an EMPTY value — refusing. Nothing was installed. A forgotten value
|
||||
would NULL the field, and a null value falls back to its default rather than erroring —
|
||||
silent, not safe. Did you mean --set .${path}=null to clear it on purpose, or --from a
|
||||
file if you need a genuinely empty string?"
|
||||
fi
|
||||
# fleetd #635 follow-up (ticket comment 17673, defect 8) — these two failure messages used to
|
||||
# echo the full "$kv" (path=value, exactly as the operator typed it), unredacted. The operator
|
||||
# already has the value, so a terminal is not where this leaks — the risk is where the output
|
||||
# goes NEXT: this fleet pastes command output into tickets, PRs and fleet_reply bodies, and a
|
||||
# failure is exactly when someone copies it to ask for help. Print the PATH, which is what's
|
||||
# needed to fix the command, and never the value. $kv is not key:value-shaped YAML, so piping
|
||||
# it through redact would just pass it straight through — a false sense of coverage, the same
|
||||
# mistake as defect 7.
|
||||
if [ "$value" = "null" ]; then
|
||||
yq eval -i ".${path} = null" "$cand" \
|
||||
|| die "yq could not clear --set '.${path}=null' — nothing was installed. The live config is unchanged."
|
||||
continue
|
||||
fi
|
||||
CONFIG_EDIT_SET_VALUE="$value" yq eval -i ".${path} = strenv(CONFIG_EDIT_SET_VALUE)" "$cand" \
|
||||
|| die "yq could not apply --set '.${path}=<value>' — nothing was installed. The live config is unchanged."
|
||||
done
|
||||
}
|
||||
|
||||
build_from_set() {
|
||||
local cand="$1"
|
||||
cp "$CONFIG" "$cand"
|
||||
apply_set_pairs "$cand" "${SETS[@]}"
|
||||
}
|
||||
|
||||
build_from_file() {
|
||||
local cand="$1"
|
||||
[ -f "$FROM_FILE" ] || die "--from file not found: $FROM_FILE"
|
||||
cp "$FROM_FILE" "$cand"
|
||||
}
|
||||
|
||||
parse_check() {
|
||||
yq eval '.' "$1" >/dev/null 2>&1
|
||||
}
|
||||
|
||||
# Count logical changed lines in a unified diff. A replacement counts once, while an added or
|
||||
# deleted line also counts once. One changed `--set` value normally produces one changed line.
|
||||
changed_line_count() {
|
||||
local before="$1" after="$2" line count=0 old_count=0 new_count=0
|
||||
while IFS= read -r line || [ -n "$line" ]; do
|
||||
case "$line" in
|
||||
---\ *|+++\ *|@@\ *)
|
||||
if [ "$old_count" -gt "$new_count" ]; then
|
||||
count=$((count + old_count))
|
||||
else
|
||||
count=$((count + new_count))
|
||||
fi
|
||||
old_count=0
|
||||
new_count=0
|
||||
;;
|
||||
-*) old_count=$((old_count + 1)) ;;
|
||||
+*) new_count=$((new_count + 1)) ;;
|
||||
*)
|
||||
if [ "$old_count" -gt "$new_count" ]; then
|
||||
count=$((count + old_count))
|
||||
else
|
||||
count=$((count + new_count))
|
||||
fi
|
||||
old_count=0
|
||||
new_count=0
|
||||
;;
|
||||
esac
|
||||
done < <(diff -u "$before" "$after" || true)
|
||||
if [ "$old_count" -gt "$new_count" ]; then
|
||||
count=$((count + old_count))
|
||||
else
|
||||
count=$((count + new_count))
|
||||
fi
|
||||
printf '%s' "$count"
|
||||
}
|
||||
|
||||
warn_set_reformat() {
|
||||
local before="$1" after="$2" changed
|
||||
changed="$(changed_line_count "$before" "$after")"
|
||||
if [ "$changed" -gt "${#SETS[@]}" ]; then
|
||||
warn "--set changed $changed candidate lines for ${#SETS[@]} pair(s); yq reformatted the whole file. Use --from for meaningful comment alignment or formatting."
|
||||
fi
|
||||
}
|
||||
|
||||
install_candidate() {
|
||||
local cand="$1" live="$2"
|
||||
mv -f "$cand" "$live"
|
||||
}
|
||||
|
||||
# -------------------------------------------------------------------------------- the report path
|
||||
#
|
||||
# Masks basic-auth userinfo (scheme://user:pass@host) in a daemon verdict line before it reaches
|
||||
# the terminal.
|
||||
mask_verdict_userinfo() {
|
||||
printf '%s\n' "$1" | sed -E 's#://[^@/[:space:]]*@#://<redacted>@#g'
|
||||
}
|
||||
|
||||
# Prints the literal command the operator (or a test) can run to restore the backup by hand — the
|
||||
# absolute path to THIS script plus the overrides actually in force, so it works from any cwd.
|
||||
restore_command_line() {
|
||||
printf '%q --restore --config %q --log %q --wait-seconds %q' "$SELF" "$CONFIG" "$LOG" "$WAIT_SECONDS"
|
||||
}
|
||||
|
||||
# State 4 only: restore the pre-edit backup, then wait for a SECOND verdict confirming the
|
||||
# restore itself reloaded cleanly. Never claims a restore it did not observe — if the second wait
|
||||
# also times out, it says so plainly rather than reporting "restored" as though confirmed.
|
||||
restore_and_confirm() {
|
||||
local backup="$1" mark2 orig_mode
|
||||
orig_mode="$(file_mode "$CONFIG")"
|
||||
mark2="$(log_mark "$LOG")"
|
||||
cp "$backup" "$CONFIG" \
|
||||
|| die "could not restore $backup onto $CONFIG — the live config is left as the REFUSED edit. Fix this by hand immediately: cp \"$backup\" \"$CONFIG\""
|
||||
apply_mode "$CONFIG" "$orig_mode"
|
||||
ok "restored from $backup"
|
||||
if wait_for_verdict "$LOG" "$mark2" "$WAIT_SECONDS"; then
|
||||
case "$VERDICT_KIND" in
|
||||
refused) warn "the RESTORE was also refused by the daemon: $(mask_verdict_userinfo "$VERDICT_LINE")" ;;
|
||||
*) ok "restore confirmed: $(mask_verdict_userinfo "$VERDICT_LINE")" ;;
|
||||
esac
|
||||
else
|
||||
warn "the restore is on disk, but no confirming verdict line appeared within ${WAIT_SECONDS}s"
|
||||
warn "cannot confirm the restore reloaded cleanly — check $LOG by hand"
|
||||
fi
|
||||
return 0
|
||||
}
|
||||
|
||||
# The four-outcome decision. Echoed as a function so run_edit/restore_mode share one place that
|
||||
# can return 0/3/4/5 — never duplicated, never re-worded between the two callers.
|
||||
report_outcome() {
|
||||
local mark="$1" backup="$2" kind line
|
||||
|
||||
say "waiting for the daemon's verdict (up to ${WAIT_SECONDS}s)"
|
||||
if wait_for_verdict "$LOG" "$mark" "$WAIT_SECONDS"; then
|
||||
kind="$VERDICT_KIND"; line="$(mask_verdict_userinfo "$VERDICT_LINE")"
|
||||
else
|
||||
kind="none"
|
||||
fi
|
||||
|
||||
case "$kind" in
|
||||
clean)
|
||||
ok "daemon verdict: $line"
|
||||
say "result: applied cleanly"
|
||||
return 0 ;;
|
||||
needs-restart)
|
||||
ok "daemon verdict: $line"
|
||||
say "result: applied — a restart is needed for the change(s) named above"
|
||||
return 3 ;;
|
||||
refused)
|
||||
warn "daemon verdict: $line"
|
||||
say "result: REFUSED — restoring the backup"
|
||||
restore_and_confirm "$backup"
|
||||
return 4 ;;
|
||||
none)
|
||||
warn "no verdict line appeared within ${WAIT_SECONDS}s after $LOG line $mark"
|
||||
warn "CANNOT TELL whether the daemon applied this edit, refused it, or is simply down."
|
||||
warn "Nothing was restored — the edit is still on disk at $CONFIG."
|
||||
echo
|
||||
echo " backup: $backup"
|
||||
echo " to restore it by hand:"
|
||||
echo " $(restore_command_line)"
|
||||
return 5 ;;
|
||||
esac
|
||||
}
|
||||
|
||||
# ------------------------------------------------------------------------------------- the modes
|
||||
check_mode() {
|
||||
say "config-edit --check"
|
||||
if [ -f "$CONFIG" ]; then
|
||||
if parse_check "$CONFIG"; then
|
||||
ok "config parses: $CONFIG"
|
||||
else
|
||||
warn "config does NOT parse as valid YAML: $CONFIG"
|
||||
fi
|
||||
else
|
||||
warn "no config file at $CONFIG"
|
||||
fi
|
||||
|
||||
local port
|
||||
port="$(resolve_port "$CONFIG")"
|
||||
if daemon_listening "$port"; then
|
||||
ok "daemon is listening on 127.0.0.1:$port"
|
||||
else
|
||||
warn "no daemon detected listening on 127.0.0.1:$port"
|
||||
fi
|
||||
|
||||
ok "watch interval assumed: $(( $(default_wait_seconds "$LOG") / 4 ))s (derives --wait-seconds default of $(default_wait_seconds "$LOG")s)"
|
||||
|
||||
local verdict
|
||||
verdict="$(last_verdict_line "$LOG")"
|
||||
if [ -n "$verdict" ]; then
|
||||
ok "last verdict in log: $(mask_verdict_userinfo "$verdict")"
|
||||
else
|
||||
warn "no reload verdict line found in $LOG"
|
||||
fi
|
||||
|
||||
local backup
|
||||
backup="$(newest_backup "$CONFIG")"
|
||||
if [ -n "$backup" ]; then
|
||||
ok "newest backup: $backup"
|
||||
else
|
||||
warn "no backups found for $CONFIG"
|
||||
fi
|
||||
|
||||
if command -v yq >/dev/null 2>&1; then
|
||||
ok "yq: $(yq --version 2>&1)"
|
||||
else
|
||||
warn "yq not found on PATH"
|
||||
fi
|
||||
|
||||
return 0
|
||||
}
|
||||
|
||||
# Shared by --set and --from: backup, build, parse-check, redacted diff, install, await verdict.
|
||||
run_edit() {
|
||||
local builder="$1"
|
||||
[ -f "$CONFIG" ] || die "no config at $CONFIG — nothing to edit"
|
||||
|
||||
local mark orig_mode
|
||||
mark="$(log_mark "$LOG")"
|
||||
orig_mode="$(file_mode "$CONFIG")"
|
||||
|
||||
say "probe"
|
||||
local port
|
||||
port="$(resolve_port "$CONFIG")"
|
||||
if daemon_listening "$port"; then
|
||||
ok "daemon appears to be listening on 127.0.0.1:$port"
|
||||
else
|
||||
warn "no daemon detected listening on 127.0.0.1:$port — a verdict may never appear"
|
||||
fi
|
||||
|
||||
say "backup"
|
||||
local backup
|
||||
backup="$(backup_config "$CONFIG")"
|
||||
ok "backup: $backup"
|
||||
|
||||
say "candidate"
|
||||
local cand
|
||||
CAND="$(mktemp "$(dirname "$CONFIG")/.config-edit.XXXXXX")" \
|
||||
|| die "could not create a candidate temp file next to $CONFIG"
|
||||
cand="$CAND"
|
||||
if ! "$builder" "$cand"; then
|
||||
rm -f "$cand"; CAND=""
|
||||
die "could not build the candidate — nothing was installed. The live config at $CONFIG is unchanged."
|
||||
fi
|
||||
|
||||
if ! parse_check "$cand"; then
|
||||
rm -f "$cand"; CAND=""
|
||||
die "candidate does not parse as valid YAML — nothing was installed. The live config at $CONFIG is unchanged."
|
||||
fi
|
||||
ok "candidate parses"
|
||||
|
||||
if [ "$MODE" = "set" ]; then
|
||||
warn_set_reformat "$backup" "$cand"
|
||||
fi
|
||||
|
||||
apply_mode "$cand" "$orig_mode"
|
||||
|
||||
say "change (redacted)"
|
||||
diff -u "$backup" "$cand" | redact "$backup" "$cand" || true
|
||||
|
||||
say "install"
|
||||
install_candidate "$cand" "$CONFIG" \
|
||||
|| die "could not install the candidate onto $CONFIG — the live config was NOT changed. The validated candidate is sitting at $cand; investigate before retrying."
|
||||
CAND=""
|
||||
ok "installed: $CONFIG"
|
||||
|
||||
local rc=0
|
||||
report_outcome "$mark" "$backup" || rc=$?
|
||||
return "$rc"
|
||||
}
|
||||
|
||||
dry_run_diff() {
|
||||
local builder="$1"
|
||||
[ -f "$CONFIG" ] || die "no config at $CONFIG — nothing to diff against"
|
||||
local cand
|
||||
CAND="$(mktemp "$(dirname "$CONFIG")/.config-edit.XXXXXX")" \
|
||||
|| die "could not create a candidate temp file next to $CONFIG"
|
||||
cand="$CAND"
|
||||
if ! "$builder" "$cand"; then
|
||||
rm -f "$cand"; CAND=""
|
||||
die "could not build the candidate — this was a --dry-run, nothing would have been installed either"
|
||||
fi
|
||||
if ! parse_check "$cand"; then
|
||||
rm -f "$cand"; CAND=""
|
||||
die "candidate does not parse as valid YAML — this was a --dry-run, nothing would have been installed either"
|
||||
fi
|
||||
if [ "$MODE" = "set" ]; then
|
||||
warn_set_reformat "$CONFIG" "$cand"
|
||||
fi
|
||||
say "dry run — diff (redacted), nothing installed"
|
||||
diff -u "$CONFIG" "$cand" | redact "$CONFIG" "$cand" || true
|
||||
rm -f "$cand"; CAND=""
|
||||
return 0
|
||||
}
|
||||
|
||||
restore_mode() {
|
||||
[ -f "$CONFIG" ] || die "no config at $CONFIG to restore onto"
|
||||
local backup dir base
|
||||
backup="$(newest_backup "$CONFIG")"
|
||||
if [ -z "$backup" ]; then
|
||||
# fleetd #635 follow-up (ticket comment 17664) — this message must name the directory the
|
||||
# code actually searches (backup_dir_for, same as newest_backup), not the old beside-the-
|
||||
# config glob. A backup written the OLD way is real and NOT searched any more — say so and
|
||||
# give the one-line recovery command — but do NOT make the search itself look there; that
|
||||
# would be a behaviour change nobody asked for. The message is the only thing being fixed.
|
||||
dir="$(backup_dir_for "$CONFIG")"
|
||||
base="$(basename "$CONFIG")"
|
||||
die "no backup found matching ${dir}/${base}.bak.* — nothing to restore.
|
||||
A backup written the OLD way, directly beside the config (${CONFIG}.bak.*), is NOT
|
||||
searched — that location was retired so a backup of a file that must never be committed
|
||||
cannot sit next to a tracked directory. If one exists there, recover it by hand:
|
||||
cp ${CONFIG}.bak.<timestamp>.<pid> $CONFIG"
|
||||
fi
|
||||
[ -f "$backup" ] || die "backup candidate $backup vanished"
|
||||
|
||||
say "restore"
|
||||
ok "restoring $backup onto $CONFIG"
|
||||
local mark orig_mode
|
||||
mark="$(log_mark "$LOG")"
|
||||
orig_mode="$(file_mode "$CONFIG")"
|
||||
cp "$backup" "$CONFIG" || die "could not copy $backup onto $CONFIG"
|
||||
apply_mode "$CONFIG" "$orig_mode"
|
||||
ok "installed: $CONFIG"
|
||||
|
||||
local rc=0
|
||||
report_outcome "$mark" "$backup" || rc=$?
|
||||
return "$rc"
|
||||
}
|
||||
|
||||
# -------------------------------------------------------------------------------------- dispatch
|
||||
|
||||
if [ -n "$WAIT_SECONDS_OVERRIDE" ]; then
|
||||
WAIT_SECONDS="$WAIT_SECONDS_OVERRIDE"
|
||||
else
|
||||
WAIT_SECONDS="$(default_wait_seconds "$LOG")"
|
||||
fi
|
||||
|
||||
RC=0
|
||||
case "$MODE" in
|
||||
check)
|
||||
check_mode || RC=$?
|
||||
;;
|
||||
set)
|
||||
[ "${#SETS[@]}" -gt 0 ] || die "--set requires at least one <yq-path>=<value>"
|
||||
if [ "$DRY_RUN" = 1 ]; then
|
||||
dry_run_diff build_from_set || RC=$?
|
||||
else
|
||||
run_edit build_from_set || RC=$?
|
||||
fi
|
||||
;;
|
||||
from)
|
||||
[ -n "$FROM_FILE" ] || die "--from requires a candidate file path"
|
||||
if [ "$DRY_RUN" = 1 ]; then
|
||||
dry_run_diff build_from_file || RC=$?
|
||||
else
|
||||
run_edit build_from_file || RC=$?
|
||||
fi
|
||||
;;
|
||||
restore)
|
||||
restore_mode || RC=$?
|
||||
;;
|
||||
esac
|
||||
|
||||
exit "$RC"
|
||||
+131
-66
@@ -71,20 +71,22 @@ set -euo pipefail
|
||||
|
||||
REPO="$(cd "$(dirname "${BASH_SOURCE[0]}")/.." && pwd)"
|
||||
MODULE="$REPO/fleetd"
|
||||
JAR="$MODULE/target/fleetd.jar"
|
||||
# fleetd #493: never build into the path a running process holds. The build writes here first
|
||||
# (Maven's shade plugin has finalName=fleetd, so `clean install` still lands its output at
|
||||
# target/fleetd.jar — that part is unchanged and out of this script's control), but this script
|
||||
# now moves it out to JAR_STAGED immediately, and only swaps it back to JAR (a plain `mv`, so a
|
||||
# rename, never a byte-by-byte overwrite) after the OLD daemon has been confirmed exited. See
|
||||
# stage_built_jar/swap_staged_jar below.
|
||||
JAR_STAGED="$MODULE/target/fleetd-new.jar"
|
||||
# fleetd #664: the runtime path and Maven's output path are no longer the same file. Maven's
|
||||
# shade plugin (finalName=fleetd) always lands a fresh build at target/fleetd.jar — that is
|
||||
# Maven's own output directory and this script does not change it — but the daemon is launched
|
||||
# from $JAR instead, outside target/ entirely. That split is the whole fix: neither `mvn install`
|
||||
# nor `mvn clean` can ever reach the file a running daemon holds open, because that file no
|
||||
# longer lives under target/ at all. See swap_if_built/swap_staged_jar below for the one `mv`
|
||||
# that moves a build from one path to the other, and only after the old daemon is confirmed gone.
|
||||
BUILD_JAR="$MODULE/target/fleetd.jar"
|
||||
JAR="$MODULE/run/fleetd.jar"
|
||||
OUT="$MODULE/fleetd.out"
|
||||
# Matches BOTH the absolute form and the relative `java -jar target/fleetd.jar` a hand-start
|
||||
# produces from inside fleetd/. Anchoring on the absolute path alone was a real bug: the daemon
|
||||
# restarted correctly and the script still reported "no process appeared", because it launched with
|
||||
# a relative path and then looked for an absolute one.
|
||||
PATTERN='target/fleetd.jar'
|
||||
# Matches a fleetd daemon's command line wherever its jar sits — absolute or relative, under
|
||||
# run/, under target/, or anywhere else a build or a hand-start might point it. Detecting a
|
||||
# daemon this script did not start, including one running from a jar outside $JAR's own
|
||||
# directory, is this pattern's whole job; running_pid()'s `comm = java` allowlist below is what
|
||||
# keeps that breadth from counting a shell that merely types the pattern as literal text.
|
||||
PATTERN='fleetd.jar'
|
||||
HEALTH='http://127.0.0.1:8765/healthz'
|
||||
STOP_WAIT=30 # seconds to wait for a clean exit before reporting failure
|
||||
HEALTH_WAIT=60 # seconds to wait for /healthz to answer after start — fleetd #603: also the pid-
|
||||
@@ -169,10 +171,10 @@ hash256() {
|
||||
fi
|
||||
}
|
||||
|
||||
# Reports the hash of $JAR by default, or of whatever path is passed — used to report the STAGED
|
||||
# jar right after a build (before it has been swapped in) without ever changing what a bare
|
||||
# `jar_id` (no args) means: the live path, $JAR. --check and the final "pid ..., jar ..." line
|
||||
# both call it with no args on purpose, so neither can ever be fooled by a leftover staged file.
|
||||
# Reports the hash of $JAR by default, or of whatever path is passed — used to report the BUILT
|
||||
# jar at $BUILD_JAR (before it has been swapped in) without ever changing what a bare `jar_id`
|
||||
# (no args) means: the live path, $JAR. --check and the final "pid ..., jar ..." line both call
|
||||
# it with no args on purpose, so neither can ever be fooled by a leftover build output.
|
||||
# fleetd #550 — THREE distinct answers now, not two: `[ -f "$f" ]` already separates "the jar is
|
||||
# not there" (-> "absent") from "the jar is there"; for the second case, hash256 itself separates
|
||||
# "hashed it" (a 12-char hex string) from "could not hash it" (-> "unhashable", when no hasher is
|
||||
@@ -180,15 +182,35 @@ hash256() {
|
||||
# was the whole defect this ticket fixes.
|
||||
jar_id() { local f="${1:-$JAR}"; [ -f "$f" ] && hash256 "$f" || echo "absent"; }
|
||||
|
||||
# fleetd #664 — under the old layout $JAR and the build output were the same file, so "jar on
|
||||
# disk" was one fact. Now they are two: $BUILD_JAR (target/fleetd.jar, whatever Maven last wrote,
|
||||
# by this script or by a bare `mvn install` run by hand) and $JAR (run/fleetd.jar, whatever the
|
||||
# daemon actually has open). Printing one label for both was the trap this ticket exists to close
|
||||
# — during the incident it would have shown the NEW jar's hash while the JVM ran the OLD one.
|
||||
# Pure (reads jar_id/date, never mutates), so a test can call it directly without reaching the
|
||||
# main flow — the same shape swap_if_built/drain_gate_refusal already use.
|
||||
report_jar_state() {
|
||||
local built_hash running_hash built_mtime running_mtime
|
||||
built_hash="$(jar_id "$BUILD_JAR")"
|
||||
running_hash="$(jar_id "$JAR")"
|
||||
built_mtime="$([ -f "$BUILD_JAR" ] && date -r "$BUILD_JAR" '+%Y-%m-%d %H:%M:%S' || echo 'none')"
|
||||
running_mtime="$([ -f "$JAR" ] && date -r "$JAR" '+%Y-%m-%d %H:%M:%S' || echo 'none')"
|
||||
ok "built jar (target/fleetd.jar): $built_hash ($built_mtime)"
|
||||
ok "running jar (run/fleetd.jar): $running_hash ($running_mtime)"
|
||||
if [ "$built_hash" != "absent" ] && [ "$running_hash" != "absent" ] && [ "$built_hash" != "$running_hash" ]; then
|
||||
warn "built jar and running jar differ — target/fleetd.jar was rebuilt since the running daemon last started and is not yet live"
|
||||
fi
|
||||
}
|
||||
|
||||
# fleetd #593 — `pgrep -f "$PATTERN"` matches ANY process whose full command line CONTAINS the
|
||||
# pattern text, and that is not the same thing as "is the daemon". A shell that merely embeds the
|
||||
# pattern as literal text — a human typing this exact investigation by hand, an ssh-shaped
|
||||
# `sh -c '...; ...'`, a pipeline, or any other non-exec'ing shell that never replaced itself with
|
||||
# the pattern-holding command — still shows up in that match, and it is the INSTRUMENT, not the
|
||||
# daemon. Measured live on this Mac: `sh -c 'echo "target/fleetd.jar" >/dev/null; sleep 30' &`
|
||||
# daemon. Measured live on this Mac: `sh -c 'echo "run/fleetd.jar" >/dev/null; sleep 30' &`
|
||||
# leaves a real `sh` process alive (it forks for the `sleep`, it does not exec into it) whose own
|
||||
# `ps -o args` is `sh -c echo "target/fleetd.jar" >/dev/null; sleep 30` — `pgrep -f "$PATTERN"`
|
||||
# matches that line right alongside the real `java -jar target/fleetd.jar` process. `pgrep -c`
|
||||
# `ps -o args` is `sh -c echo "run/fleetd.jar" >/dev/null; sleep 30` — `pgrep -f "$PATTERN"`
|
||||
# matches that line right alongside the real `java -jar run/fleetd.jar` process. `pgrep -c`
|
||||
# (an in-one-call count) does not exist on BSD/macOS at all, so this cannot be fixed by switching
|
||||
# pgrep flags — it has to filter what pgrep already found, after the fact, in a way that still
|
||||
# runs on BSD.
|
||||
@@ -210,7 +232,7 @@ jar_id() { local f="${1:-$JAR}"; [ -f "$f" ] && hash256 "$f" || echo "absent"; }
|
||||
# launched as `java -jar ...` — a native image, a renamed launcher — `running_pid()` silently
|
||||
# returns nothing and `assert_single_daemon` stops noticing a second daemon at all. For a guard,
|
||||
# that false-negative direction is the worse one to be wrong in. This is not a new assumption,
|
||||
# though: `PATTERN='target/fleetd.jar'` two lines up already assumes the daemon is a jar, which
|
||||
# though: `PATTERN='fleetd.jar'` two lines up already assumes the daemon is a jar, which
|
||||
# is only ever run by `java`. If that launch method changes, `PATTERN` stops matching anything
|
||||
# before this allowlist would ever get the chance to be wrong — the allowlist rides on the same
|
||||
# assumption that is already load-bearing, it does not add a new one. Whoever changes the launch
|
||||
@@ -229,16 +251,12 @@ running_pid() {
|
||||
printf '%s' "$out"
|
||||
}
|
||||
|
||||
# fleetd #493 — three small, independently testable pieces of "never build into the path a
|
||||
# running process holds":
|
||||
# fleetd #493/#664 — the independently testable pieces of "never build into the path a running
|
||||
# process holds":
|
||||
#
|
||||
# stage_built_jar moves the jar Maven just produced OUT of the live path and onto the staging
|
||||
# path, immediately after a successful build. Dies (leaving the OLD daemon
|
||||
# untouched — this runs before the stop step) if Maven reported success but
|
||||
# left no jar behind, or if the move itself fails.
|
||||
# require_no_build_jar the --no-build path never builds or stages anything: it must find a
|
||||
# jar already sitting at the live path from an earlier successful run, and
|
||||
# die with the same truthful message this script has always used if not.
|
||||
# require_no_build_jar the --no-build path never builds anything: it must find a jar already
|
||||
# sitting at the live path ($JAR, under run/) from an earlier successful run,
|
||||
# and die with the same truthful message this script has always used if not.
|
||||
# wait_for_daemon_exit polls running_pid() for up to $1 seconds and reports whether the OLD
|
||||
# daemon actually exited — extracted to its own function so the main flow
|
||||
# can be relied on to call swap_staged_jar only AFTER this returns success,
|
||||
@@ -248,14 +266,10 @@ running_pid() {
|
||||
# so this is never a write into a path a running process holds — by the time
|
||||
# it runs, nothing holds that path anymore. If it fails, the caller must not
|
||||
# start a new daemon: die() below already refuses that by exiting the script.
|
||||
stage_built_jar() {
|
||||
[ -f "$JAR" ] || die "build succeeded but produced no jar at $JAR — cannot stage it for restart.
|
||||
The running daemon was NOT touched."
|
||||
mv -f "$JAR" "$JAR_STAGED" \
|
||||
|| die "could not move the freshly built jar from $JAR to the staging path $JAR_STAGED.
|
||||
The running daemon was NOT touched."
|
||||
}
|
||||
|
||||
# fleetd #664: the "staged" jar swap_if_built passes in is now $BUILD_JAR
|
||||
# itself (target/fleetd.jar, Maven's own output) — a build no longer needs to
|
||||
# be moved off the live path right after compiling, because target/ was never
|
||||
# the live path to begin with.
|
||||
require_no_build_jar() {
|
||||
[ -f "$JAR" ] || die "no jar at $JAR — run without --no-build"
|
||||
}
|
||||
@@ -282,8 +296,7 @@ swap_staged_jar() {
|
||||
#
|
||||
# The defect: the swap step used to be guarded inline by `if [ "$DO_BUILD" = 1 ]` in the main flow.
|
||||
# Changing that to `if false` left the suite green and the swap never ran, so a redeploy reported
|
||||
# every step succeeding while the daemon started on no jar at all (stage_built_jar has already moved
|
||||
# the freshly built one to $JAR_STAGED by then) or on a stale one.
|
||||
# every step succeeding while the daemon started on no jar at all or on a stale one.
|
||||
# test_swap_ordered_after_wait_and_before_start could not catch it: it reads this script's own text
|
||||
# and compares line positions, and a same-line edit moves no line.
|
||||
#
|
||||
@@ -312,7 +325,12 @@ swap_if_built() {
|
||||
local do_build="$1"
|
||||
should_swap "$do_build" || return 0
|
||||
say "swap"
|
||||
swap_staged_jar "$JAR_STAGED" "$JAR"
|
||||
# fleetd #664: $JAR now lives under run/, a directory target/ never created. mkdir -p here,
|
||||
# not inside swap_staged_jar itself — that function's own contract is tested on a missing
|
||||
# parent directory (a failing mv), and widening it to auto-create one would change what that
|
||||
# test proves.
|
||||
mkdir -p "$(dirname "$JAR")"
|
||||
swap_staged_jar "$BUILD_JAR" "$JAR"
|
||||
ok "jar in place: $(jar_id)"
|
||||
}
|
||||
|
||||
@@ -610,6 +628,45 @@ check_log_path_matches_plist() {
|
||||
ok "log path check: script and plist agree ($resolved_out)"
|
||||
}
|
||||
|
||||
# Reads the launchd plist's ProgramArguments for the argument that follows "-jar", resolves it
|
||||
# alongside $jar_path, and dies when the two differ. Call it only when the agent is loaded; it
|
||||
# never touches launchd or the daemon itself.
|
||||
check_jar_path_matches_plist() {
|
||||
local jar_path="$1" plist_path="$2"
|
||||
local plist_args plist_jar resolved_jar resolved_plist_jar
|
||||
if ! plist_args="$(/usr/libexec/PlistBuddy -c 'Print :ProgramArguments' "$plist_path" 2>/dev/null)"; then
|
||||
die "launchd agent is loaded but PlistBuddy could not read ProgramArguments from
|
||||
$plist_path
|
||||
— cannot verify which jar the supervised daemon launches. Fix the plist before
|
||||
redeploying supervised."
|
||||
fi
|
||||
plist_jar="$(printf '%s\n' "$plist_args" | awk '
|
||||
{ gsub(/^[ \t]+|[ \t]+$/, "") }
|
||||
prev == "-jar" { print; exit }
|
||||
{ prev = $0 }
|
||||
')"
|
||||
if [ -z "$plist_jar" ]; then
|
||||
die "launchd agent is loaded but its ProgramArguments at
|
||||
$plist_path
|
||||
do not contain a '-jar <path>' pair — cannot verify which jar the supervised daemon
|
||||
launches. Fix the plist before redeploying supervised."
|
||||
fi
|
||||
resolved_jar="$(cd "$(dirname "$jar_path")" 2>/dev/null && pwd -P)/$(basename "$jar_path")" || true
|
||||
resolved_plist_jar="$(cd "$(dirname "$plist_jar")" 2>/dev/null && pwd -P)/$(basename "$plist_jar")" || true
|
||||
if [ -z "$resolved_jar" ] || [ -z "$resolved_plist_jar" ] || [ "$resolved_jar" != "$resolved_plist_jar" ]; then
|
||||
die "jar path mismatch — this script deploys to
|
||||
$jar_path (resolved: ${resolved_jar:-<directory does not exist>})
|
||||
but the loaded plist's ProgramArguments names
|
||||
$plist_jar (resolved: ${resolved_plist_jar:-<directory does not exist>})
|
||||
The swap renames the built jar into place, so the old path stops existing after a redeploy;
|
||||
a launchd-initiated start from this plist (a reboot, or KeepAlive after a crash) would then
|
||||
run java against a missing file. Reinstall the plist at
|
||||
$plist_path
|
||||
so its ProgramArguments names $jar_path before redeploying supervised."
|
||||
fi
|
||||
ok "jar path check: script and plist agree ($resolved_jar)"
|
||||
}
|
||||
|
||||
# fleetd #552: the post-restart fresh-log capture, pulled out of the main flow so it is testable by
|
||||
# sourcing (the same reason systemd_installed/systemd_loaded above guard their OWN mktemp inline
|
||||
# instead of leaving it bare) even though its only caller sits below the SOURCED guard. By the time
|
||||
@@ -826,16 +883,25 @@ report_shutdown_drain() {
|
||||
# --no-build, staged jar present -> ALSO "nothing changed", deliberately: --no-build itself builds
|
||||
# and stages nothing (see require_no_build_jar above), so a staged jar found here is a leftover
|
||||
# from an earlier, unrelated run. THIS run truly changed nothing, and the next DO_BUILD=1 run
|
||||
# wipes that leftover before it builds (`rm -f "$JAR_STAGED"` in the build section above) — so
|
||||
# there is nothing here for the operator to lose track of.
|
||||
# wipes that leftover for free — `mvn clean` deletes all of target/, $BUILD_JAR included,
|
||||
# before the build even starts — so there is nothing here for the operator to lose track of.
|
||||
# --no-build, staged jar absent -> "nothing changed"
|
||||
drain_gate_refusal() {
|
||||
local do_build="$1" staged_path="$2"
|
||||
if [ "$do_build" = 1 ] && [ -f "$staged_path" ]; then
|
||||
# fleetd #664: under the old layout a build emptied the live path ($JAR) immediately, so
|
||||
# --no-build's own check ("no jar at $JAR") was the thing that refused a rerun here. That is
|
||||
# no longer true: a build never touches $JAR at all now, so $JAR still holds whatever was
|
||||
# already running before this gate fired (reaching this message at all requires OLD_PID to
|
||||
# have been set, which means a daemon was running from $JAR already) — a --no-build rerun
|
||||
# would NOT refuse, it would just restart that same old jar and silently throw away the one
|
||||
# sitting at %s.
|
||||
printf 'aborted — the running daemon was NOT touched, but the freshly built jar is sitting at
|
||||
%s, not yet swapped into %s. Rerun WITHOUT --no-build to finish the restart —
|
||||
the freshly built jar is no longer at the live path that --no-build requires — or
|
||||
remove %s by hand if you want to discard this build.' "$staged_path" "$JAR" "$staged_path"
|
||||
%s, not yet swapped into %s. Rerun WITHOUT --no-build to finish the restart — a
|
||||
--no-build rerun would NOT refuse here: %s already exists from before this run, so it
|
||||
would restart the daemon on that OLD jar and silently discard the one you just built — or
|
||||
remove %s by hand if you want to discard this build instead.' \
|
||||
"$staged_path" "$JAR" "$JAR" "$staged_path"
|
||||
else
|
||||
printf 'aborted — nothing changed'
|
||||
fi
|
||||
@@ -934,6 +1000,7 @@ report_supervisor_state() {
|
||||
SUPERVISED=1
|
||||
ok "launchd agent loaded ($LAUNCHD_LABEL) — launchd supervises this daemon"
|
||||
check_log_path_matches_plist "$OUT" "$LAUNCHD_PLIST"
|
||||
check_jar_path_matches_plist "$JAR" "$LAUNCHD_PLIST"
|
||||
;;
|
||||
systemd)
|
||||
SUPERVISED=1
|
||||
@@ -1170,7 +1237,7 @@ if [ -n "$OLD_PID" ]; then
|
||||
else
|
||||
warn "no daemon running — this will be a cold start"
|
||||
fi
|
||||
ok "jar on disk: $(jar_id) ($([ -f "$JAR" ] && date -r "$JAR" '+%Y-%m-%d %H:%M:%S' || echo 'none'))"
|
||||
report_jar_state
|
||||
ok "HEAD: $(git -C "$REPO" log --oneline -1)"
|
||||
|
||||
# CB-594 / fleetd #492: supervision state. Installed and loaded are different facts — a
|
||||
@@ -1252,9 +1319,6 @@ stop_if_check_only "$CHECK_ONLY"
|
||||
|
||||
if [ "$DO_BUILD" = 1 ]; then
|
||||
say "build"
|
||||
# fleetd #493: wipe a leftover staged jar from a previous failed/interrupted run BEFORE doing
|
||||
# anything else, so that run's leftovers can never be mistaken for this run's output.
|
||||
rm -f "$JAR_STAGED"
|
||||
BUILD_LOG="$(mktemp -t fleetd-build.XXXXXX)"
|
||||
echo " log: $BUILD_LOG"
|
||||
if ! mvn -f "$MODULE/pom.xml" clean install > "$BUILD_LOG" 2>&1; then
|
||||
@@ -1264,16 +1328,16 @@ if [ "$DO_BUILD" = 1 ]; then
|
||||
fi
|
||||
grep -E '^\[INFO\] Tests run:.*Failures' "$BUILD_LOG" | tail -1 | sed 's/^\[INFO\] / /' || true
|
||||
ok "BUILD SUCCESS"
|
||||
# fleetd #493: move the freshly built jar off the live path immediately — the running (OLD)
|
||||
# daemon, if any, is still up at this point (build always runs before stop). From here until the
|
||||
# swap step below (after the OLD daemon is confirmed gone), $JAR_STAGED is the only artefact this
|
||||
# script treats as "the new jar" — $JAR itself is not touched again until the swap.
|
||||
stage_built_jar
|
||||
ok "jar now: $(jar_id "$JAR_STAGED")"
|
||||
# fleetd #664: nothing to stage — $BUILD_JAR (target/fleetd.jar) is Maven's own output path and
|
||||
# was never the live path, so the running (OLD) daemon, if any, was never at risk from this build
|
||||
# at all. From here until the swap step below (after the OLD daemon is confirmed gone),
|
||||
# $BUILD_JAR is the artefact this script treats as "the new jar" — $JAR itself is not touched
|
||||
# again until the swap.
|
||||
ok "jar now: $(jar_id "$BUILD_JAR")"
|
||||
else
|
||||
say "build skipped (--no-build)"
|
||||
# fleetd #493: --no-build never builds or stages anything — it restarts whatever jar is already
|
||||
# sitting at the live path from an earlier successful run. Same check, same message as before.
|
||||
# fleetd #493: --no-build never builds anything — it restarts whatever jar is already sitting at
|
||||
# the live path from an earlier successful run. Same check, same message as before.
|
||||
require_no_build_jar
|
||||
fi
|
||||
|
||||
@@ -1282,8 +1346,8 @@ fi
|
||||
# fleetd #555: run_drain_gate above is drain_gate_required + the prompt + drain_confirmed, called
|
||||
# unconditionally — it returns immediately when the gate is not required, and composes/dies through
|
||||
# refuse_drain_gate itself when the reply does not confirm. See #493/#517/#528 for why "nothing
|
||||
# changed" would be a lie once a build has staged a jar.
|
||||
run_drain_gate "$OLD_PID" "$ASSUME_YES" "$DO_BUILD" "$JAR_STAGED"
|
||||
# changed" would be a lie once a build has produced a jar not yet swapped in.
|
||||
run_drain_gate "$OLD_PID" "$ASSUME_YES" "$DO_BUILD" "$BUILD_JAR"
|
||||
|
||||
# ------------------------------------------------------------------ stop
|
||||
#
|
||||
@@ -1334,21 +1398,22 @@ fi
|
||||
|
||||
# ------------------------------------------------------------------ swap
|
||||
#
|
||||
# fleetd #493: every branch above has now either confirmed the OLD daemon actually exited
|
||||
# (wait_for_daemon_exit, above) or established there was never one running to begin with. Only
|
||||
# NOW is it safe to put the freshly built jar at the path the NEXT `java -jar` (direct, or via
|
||||
# launchd/systemd's ExecStart) will read from — this mv is the one and only write to $JAR anywhere
|
||||
# fleetd #493/#664: every branch above has now either confirmed the OLD daemon actually exited
|
||||
# (wait_for_daemon_exit, above) or established there was never one running to begin with. Only NOW
|
||||
# is it safe to put the freshly built jar at the path the NEXT `java -jar` (direct, or via
|
||||
# launchd/systemd's ExecStart) will read from — a rename from $BUILD_JAR (target/) to $JAR (run/),
|
||||
# both under $MODULE and so on one filesystem. This mv is the one and only write to $JAR anywhere
|
||||
# in this script's mutating flow. If it fails, do not start: die() below exits before "start" runs.
|
||||
swap_if_built "$DO_BUILD"
|
||||
|
||||
# ------------------------------------------------------------------ start
|
||||
# Unsupervised: login shell (zsh -l) is what puts the secrets on the daemon's environment, and cwd
|
||||
# must be fleetd/ because the daemon resolves fleetd.yaml, logs/ and target/ relative to it.
|
||||
# must be fleetd/ because the daemon resolves fleetd.yaml, logs/ and run/ relative to it.
|
||||
# Supervised (launchd): launchd does both — deploy/dev.ltms.fleetd.plist points ProgramArguments at
|
||||
# scripts/fleetd-launchd-wrapper.sh (CB-594), which is what execs the login shell in launchd's
|
||||
# place, and WorkingDirectory in the plist already pins fleetd/.
|
||||
# Supervised (systemd --user): the unit does both too — measured on the second host, ExecStart is
|
||||
# `/bin/zsh -lc "exec java -jar target/fleetd.jar fleetd.yaml"` (a login shell, same reason as
|
||||
# `/bin/zsh -lc "exec java -jar run/fleetd.jar fleetd.yaml"` (a login shell, same reason as
|
||||
# above) and WorkingDirectory is already pinned to fleetd/.
|
||||
|
||||
say "start"
|
||||
|
||||
Executable
+793
@@ -0,0 +1,793 @@
|
||||
#!/usr/bin/env bash
|
||||
# Self-contained checks for scripts/config-edit.sh — fleetd ticket #635.
|
||||
#
|
||||
# Drives the REAL config-edit.sh as a subprocess against a FIXTURE config and a FIXTURE log in a
|
||||
# throwaway temp directory this file creates and removes. Never touches fleetd/fleetd.yaml or
|
||||
# fleetd/fleetd.out, and never starts, stops, or contacts a daemon — there is no daemon here, so
|
||||
# each test PLAYS the daemon: it starts config-edit.sh in the background (it is waiting on the
|
||||
# log), appends the verdict line it wants, then collects the real exit code.
|
||||
|
||||
set -euo pipefail
|
||||
|
||||
ROOT="$(cd "$(dirname "${BASH_SOURCE[0]}")/.." && pwd)"
|
||||
EDIT="$ROOT/scripts/config-edit.sh"
|
||||
TMP="$(mktemp -d "$ROOT/.config-edit-test.XXXXXX")"
|
||||
trap 'rm -rf "$TMP"' EXIT
|
||||
|
||||
fail() {
|
||||
printf 'FAIL: %s\n' "$*" >&2
|
||||
return 1
|
||||
}
|
||||
|
||||
assert_equals() {
|
||||
local expected="$1" actual="$2" description="$3"
|
||||
[ "$expected" = "$actual" ] || fail "$description: expected $expected, got $actual"
|
||||
}
|
||||
|
||||
assert_contains() {
|
||||
local needle="$1" text="$2" description="$3"
|
||||
printf '%s' "$text" | grep -qF -- "$needle" || fail "$description: missing [$needle]"
|
||||
}
|
||||
|
||||
assert_not_contains() {
|
||||
local needle="$1" text="$2" description="$3"
|
||||
if printf '%s' "$text" | grep -qF -- "$needle"; then
|
||||
fail "$description: must NOT contain [$needle], but it does"
|
||||
fi
|
||||
return 0
|
||||
}
|
||||
|
||||
# A fresh fixture pair per test: $1/fleetd.yaml (the config) and $1/fleetd.out (the log), plus a
|
||||
# small wait-seconds budget so no test takes long. Returns the fixture dir via stdout.
|
||||
new_fixture() {
|
||||
local dir
|
||||
dir="$(mktemp -d "$TMP/fixture.XXXXXX")"
|
||||
cat > "$dir/fleetd.yaml" <<'YAML'
|
||||
bind:
|
||||
host: 127.0.0.1
|
||||
port: 19999
|
||||
broker:
|
||||
uri: amqp://user:hunter2@host/vhost
|
||||
profiles:
|
||||
sonnet:
|
||||
weight: 3
|
||||
maxLoad: 5
|
||||
YAML
|
||||
: > "$dir/fleetd.out"
|
||||
printf '%s' "$dir"
|
||||
}
|
||||
|
||||
# Runs config-edit.sh in the background against $dir's fixtures, with the given extra args, and
|
||||
# a short --wait-seconds. Sets RUN_PID. Caller appends to $dir/fleetd.out (or not, for the
|
||||
# silence test) and then calls collect_run to block for the exit code.
|
||||
start_run() {
|
||||
local dir="$1" wait_s="$2"; shift 2
|
||||
(
|
||||
# config-edit.sh deliberately exits 3/4/5 on several of these tests. This subshell inherits
|
||||
# the parent's `set -e`, and without disabling it here the FIRST nonzero exit would kill the
|
||||
# subshell before the `echo $? > rc` line ever ran — the real code would never reach the file.
|
||||
set +e
|
||||
"$EDIT" --config "$dir/fleetd.yaml" --log "$dir/fleetd.out" --wait-seconds "$wait_s" "$@" \
|
||||
> "$dir/stdout.log" 2>&1
|
||||
echo $? > "$dir/rc"
|
||||
) &
|
||||
RUN_PID=$!
|
||||
}
|
||||
|
||||
collect_run() {
|
||||
local dir="$1"
|
||||
# wait echoes back the backgrounded subshell's own exit status (here, deliberately 3/4/5 on
|
||||
# several tests) — under `set -e` a bare nonzero `wait` would abort this whole test script, so
|
||||
# it is neutralized with `|| true`; the real code is read from $dir/rc right after.
|
||||
wait "$RUN_PID" || true
|
||||
RUN_OUTPUT="$(cat "$dir/stdout.log")"
|
||||
RUN_RC="$(cat "$dir/rc")"
|
||||
}
|
||||
|
||||
# -------------------------------------------------------------- acceptance criterion 1: refusal
|
||||
test_refusal_restores_byte_for_byte() {
|
||||
local dir
|
||||
dir="$(new_fixture)"
|
||||
cp "$dir/fleetd.yaml" "$dir/pre-edit.yaml"
|
||||
|
||||
start_run "$dir" 5 --set '.broker.uri=amqp://changed@host/x'
|
||||
sleep 1
|
||||
printf 'config reload refused — these keys cannot change under a running daemon: broker. Restart fleetd to apply them.\n' >> "$dir/fleetd.out"
|
||||
collect_run "$dir"
|
||||
|
||||
assert_equals 4 "$RUN_RC" "refusal exit code"
|
||||
cmp -s "$dir/fleetd.yaml" "$dir/pre-edit.yaml" \
|
||||
|| fail "refusal must restore the config byte for byte onto the pre-edit backup"
|
||||
}
|
||||
|
||||
# -------------------------------------------------------------- acceptance criterion 2: clean
|
||||
test_clean_reload_keeps_the_edit() {
|
||||
local dir
|
||||
dir="$(new_fixture)"
|
||||
|
||||
start_run "$dir" 5 --set '.profiles.sonnet.weight=7'
|
||||
sleep 1
|
||||
printf 'config reloaded\n' >> "$dir/fleetd.out"
|
||||
collect_run "$dir"
|
||||
|
||||
assert_equals 0 "$RUN_RC" "clean reload exit code"
|
||||
assert_equals "7" "$(yq eval '.profiles.sonnet.weight' "$dir/fleetd.yaml")" "clean reload live value"
|
||||
}
|
||||
|
||||
# ----------------------------------------------------- acceptance criterion 3: deferred != clean
|
||||
test_deferred_reload_is_told_apart_from_clean() {
|
||||
local dir
|
||||
dir="$(new_fixture)"
|
||||
|
||||
start_run "$dir" 5 --set '.profiles.sonnet.weight=9'
|
||||
sleep 1
|
||||
printf 'config reloaded; these changes need a restart to take effect: profiles\n' >> "$dir/fleetd.out"
|
||||
collect_run "$dir"
|
||||
|
||||
assert_equals 3 "$RUN_RC" "deferred reload exit code"
|
||||
[ "$RUN_RC" != 0 ] || fail "deferred reload must not report exit 0"
|
||||
assert_contains "profiles" "$RUN_OUTPUT" "deferred reload names the key"
|
||||
assert_contains "restart" "$RUN_OUTPUT" "deferred reload says a restart is needed"
|
||||
}
|
||||
|
||||
# -------------------------------------------------------------- acceptance criterion 4: silence
|
||||
test_silence_is_its_own_answer() {
|
||||
local dir
|
||||
dir="$(new_fixture)"
|
||||
|
||||
start_run "$dir" 2 --set '.profiles.sonnet.weight=11'
|
||||
# Feed the log nothing.
|
||||
collect_run "$dir"
|
||||
|
||||
assert_equals 5 "$RUN_RC" "silence exit code"
|
||||
assert_equals "11" "$(yq eval '.profiles.sonnet.weight' "$dir/fleetd.yaml")" "the edited value must still be on disk"
|
||||
assert_contains '--restore' "$RUN_OUTPUT" "silence prints the --restore command"
|
||||
|
||||
local backup restore_cmd
|
||||
backup="$(ls -t "$dir"/.config-backups/fleetd.yaml.bak.* | head -1)"
|
||||
[ -n "$backup" ] || fail "silence must still have taken a backup"
|
||||
|
||||
restore_cmd="$(printf '%s\n' "$RUN_OUTPUT" | grep -F -- '--restore --config' | sed -E 's/^[[:space:]]*//')"
|
||||
[ -n "$restore_cmd" ] || fail "could not find the printed --restore invocation in the output"
|
||||
# Running this --restore invocation installs the backup, then itself waits for a confirming
|
||||
# verdict that this fixture never feeds — so it legitimately exits 5 ("cannot tell") here, same
|
||||
# as any edit with no daemon on the other end. Only a usage/internal error (1 or 2) is a real
|
||||
# failure of the command itself; the actual assertion is the byte-for-byte cmp below.
|
||||
local restore_rc=0
|
||||
eval "$restore_cmd" > "$dir/restore.log" 2>&1 || restore_rc=$?
|
||||
case "$restore_rc" in
|
||||
0|3|4|5) : ;;
|
||||
*) fail "the printed --restore command errored out (exit $restore_rc): $(cat "$dir/restore.log")" ;;
|
||||
esac
|
||||
|
||||
cmp -s "$dir/fleetd.yaml" "$backup" \
|
||||
|| fail "running the printed --restore command must put the file back to the original backup"
|
||||
}
|
||||
|
||||
# --------------------------------------------------------- acceptance criterion 5: bad candidate
|
||||
test_broken_candidate_never_reaches_live_path() {
|
||||
local dir rc=0
|
||||
dir="$(new_fixture)"
|
||||
printf 'foo: [unclosed\n' > "$dir/broken.yaml"
|
||||
|
||||
"$EDIT" --from "$dir/broken.yaml" --config "$dir/fleetd.yaml" --log "$dir/fleetd.out" --wait-seconds 2 \
|
||||
> "$dir/stdout.log" 2>&1 || rc=$?
|
||||
|
||||
[ "$rc" -ne 0 ] || fail "a broken --from candidate must exit non-zero"
|
||||
cmp -s "$dir/fleetd.yaml" <(new_fixture_yaml) \
|
||||
|| fail "the broken candidate must never reach the live fixture config"
|
||||
}
|
||||
new_fixture_yaml() {
|
||||
cat <<'YAML'
|
||||
bind:
|
||||
host: 127.0.0.1
|
||||
port: 19999
|
||||
broker:
|
||||
uri: amqp://user:hunter2@host/vhost
|
||||
profiles:
|
||||
sonnet:
|
||||
weight: 3
|
||||
maxLoad: 5
|
||||
YAML
|
||||
}
|
||||
|
||||
# -------------------------------------------------------------------- acceptance criterion 6
|
||||
test_marker_skips_lines_before_it() {
|
||||
local dir
|
||||
dir="$(new_fixture)"
|
||||
printf 'config reload refused — something ancient\n' > "$dir/fleetd.out"
|
||||
|
||||
start_run "$dir" 5 --set '.profiles.sonnet.weight=5'
|
||||
sleep 1
|
||||
printf 'config reloaded\n' >> "$dir/fleetd.out"
|
||||
collect_run "$dir"
|
||||
|
||||
assert_equals 0 "$RUN_RC" "a stale refusal before the marker must not be read as this edit's verdict"
|
||||
}
|
||||
|
||||
# ------------------------------------------------------------------- acceptance criterion 7 (+13)
|
||||
# fleetd #635 follow-up (ticket comment 17659) — the two assertions below this comment were the
|
||||
# WHOLE test before the follow-up, and both are negative-only: they pass just as happily when the
|
||||
# diff is never printed at all as when it is printed and correctly redacted. A mutant that deletes
|
||||
# `diff -u "$backup" "$cand" | redact` from the edit path survives them, because an absent output
|
||||
# contains neither "hunter2" nor "user:" either — see the mutation-and-revert proof in the reply.
|
||||
# Criterion 13 is the fix: a LOUD positive control that only passes when a diff was demonstrably
|
||||
# printed AND the redaction demonstrably ran on real content, not merely that nothing leaked.
|
||||
test_redaction_holds() {
|
||||
local dir
|
||||
dir="$(new_fixture)"
|
||||
|
||||
start_run "$dir" 5 --set '.profiles.sonnet.weight=4'
|
||||
sleep 1
|
||||
printf 'config reloaded\n' >> "$dir/fleetd.out"
|
||||
collect_run "$dir"
|
||||
|
||||
assert_equals 0 "$RUN_RC" "redaction-case reload exit code"
|
||||
assert_not_contains "hunter2" "$RUN_OUTPUT" "full output must never contain the password"
|
||||
assert_not_contains "user:" "$RUN_OUTPUT" "full output must never contain the userinfo"
|
||||
# acceptance criterion 13 — positive control: the diff's default 3-line context around the
|
||||
# changed "weight" key also covers the fixture's "uri:" line, so a genuinely-printed, genuinely-
|
||||
# redacted diff must contain BOTH the redaction marker and the changed key's name. A test that
|
||||
# only ever asserts absence cannot tell "redacted" from "never printed" apart; this can.
|
||||
assert_contains "<redacted>" "$RUN_OUTPUT" "the redaction must be PROVEN to have run on real content, not merely absent"
|
||||
assert_contains "weight" "$RUN_OUTPUT" "a diff must have been demonstrably printed at all"
|
||||
}
|
||||
|
||||
# ------------------------------------------------------- acceptance criterion 9: forgotten value
|
||||
# `--set .a.b=` is a plausible typo (the value simply forgotten), and it must be refused outright
|
||||
# rather than silently nulling the field — a null numeric field falls back to its default, which
|
||||
# widens capacity instead of failing loudly. No background verdict feeder here: a refused --set
|
||||
# must never even reach the daemon, so this never starts a background run at all.
|
||||
test_forgotten_value_refuses_and_installs_nothing() {
|
||||
local dir rc=0
|
||||
dir="$(new_fixture)"
|
||||
cp "$dir/fleetd.yaml" "$dir/pre-edit.yaml"
|
||||
|
||||
"$EDIT" --set '.profiles.sonnet.maxLoad=' \
|
||||
--config "$dir/fleetd.yaml" --log "$dir/fleetd.out" --wait-seconds 2 \
|
||||
> "$dir/stdout.log" 2>&1 || rc=$?
|
||||
RUN_OUTPUT="$(cat "$dir/stdout.log")"
|
||||
|
||||
[ "$rc" -ne 0 ] || fail "an empty --set value must exit non-zero, got 0"
|
||||
cmp -s "$dir/fleetd.yaml" "$dir/pre-edit.yaml" \
|
||||
|| fail "an empty --set value must install nothing — the live fixture changed"
|
||||
assert_contains "EMPTY value" "$RUN_OUTPUT" "the refusal must name the empty value"
|
||||
}
|
||||
|
||||
# ---------------------------------------------------------- acceptance criterion 10: explicit null
|
||||
# `--set .a.b=null` is the deliberate-clear spelling, and it must write a REAL yaml null, never
|
||||
# the string "''" — those are different values to the daemon's loader (fleetd ticket #635's
|
||||
# follow-up comment measured `""` reading back as a null field anyway, which is exactly why the
|
||||
# two forms must not collapse onto each other: `--set path=` refuses instead of silently reaching
|
||||
# this same null outcome through the back door). Read the RAW line with grep, never only through
|
||||
# `yq` — `yq eval` reports `null` for both an actual null and a missing/absent key, so it cannot
|
||||
# tell "wrote null" apart from "wrote nothing"; only the literal line on disk can.
|
||||
test_explicit_null_writes_bare_null_not_empty_string() {
|
||||
local dir
|
||||
dir="$(new_fixture)"
|
||||
|
||||
start_run "$dir" 5 --set '.profiles.sonnet.maxLoad=null'
|
||||
sleep 1
|
||||
printf 'config reloaded\n' >> "$dir/fleetd.out"
|
||||
collect_run "$dir"
|
||||
|
||||
assert_equals 0 "$RUN_RC" "explicit null clear exit code"
|
||||
local raw_line
|
||||
raw_line="$(grep -E 'maxLoad' "$dir/fleetd.yaml")"
|
||||
assert_contains "null" "$raw_line" "the installed line must spell a bare null"
|
||||
assert_not_contains '""' "$raw_line" "the installed line must NOT be a quoted empty string"
|
||||
}
|
||||
|
||||
# ------------------------------------------------------- acceptance criterion 11: backup never committable
|
||||
# A backup of fleetd.yaml inherits fleetd.yaml's own "never commit this" requirement (fleetd #635
|
||||
# follow-up, ticket comment 17655). Proves two things: the backup lands somewhere `git
|
||||
# check-ignore` reports as ignored (equivalently, a path `git status --porcelain` never lists as
|
||||
# untracked), AND that --restore still finds and uses it from that location.
|
||||
test_backup_is_never_committable() {
|
||||
local dir backup
|
||||
dir="$(new_fixture)"
|
||||
cp "$dir/fleetd.yaml" "$dir/pre-edit.yaml"
|
||||
|
||||
start_run "$dir" 5 --set '.profiles.sonnet.weight=55'
|
||||
sleep 1
|
||||
printf 'config reloaded\n' >> "$dir/fleetd.out"
|
||||
collect_run "$dir"
|
||||
assert_equals 0 "$RUN_RC" "setup edit exit code"
|
||||
|
||||
backup="$(ls -t "$dir"/.config-backups/fleetd.yaml.bak.* 2>/dev/null | head -1)"
|
||||
[ -n "$backup" ] || fail "no backup found under .config-backups/ — did the location change?"
|
||||
|
||||
git -C "$ROOT" check-ignore -q -- "$backup" \
|
||||
|| fail "the backup at $backup is NOT gitignored — it would survive a git add -A"
|
||||
if git -C "$ROOT" status --porcelain -- "$backup" 2>/dev/null | grep -q '^??'; then
|
||||
fail "git status still lists the backup as untracked: $backup"
|
||||
fi
|
||||
|
||||
start_run "$dir" 5 --restore
|
||||
sleep 1
|
||||
printf 'config reloaded\n' >> "$dir/fleetd.out"
|
||||
collect_run "$dir"
|
||||
assert_equals 0 "$RUN_RC" "--restore after the backup-location change exit code"
|
||||
cmp -s "$dir/fleetd.yaml" "$dir/pre-edit.yaml" \
|
||||
|| fail "--restore from the new backup location must still put the file back byte for byte"
|
||||
}
|
||||
|
||||
# --------------------------------------------------------- acceptance criterion 12: file mode
|
||||
# `mv` from a mktemp candidate carries mktemp's 0600 forever, and a plain `cp` onto an existing
|
||||
# file keeps the DESTINATION's mode rather than the source's, so a restore does not undo the
|
||||
# narrowing either (fleetd #635 follow-up, ticket comment 17657). Proves the mode survives an edit
|
||||
# AND a subsequent restore, from two different starting points — 644 is the common case, 600
|
||||
# proves the fix PRESERVES whatever mode was there rather than hardcoding 644.
|
||||
test_file_mode_survives_edit_and_restore() {
|
||||
local dir want got
|
||||
for want in 644 600; do
|
||||
dir="$(new_fixture)"
|
||||
chmod "$want" "$dir/fleetd.yaml"
|
||||
|
||||
start_run "$dir" 5 --set ".profiles.sonnet.weight=${want}"
|
||||
sleep 1
|
||||
printf 'config reloaded\n' >> "$dir/fleetd.out"
|
||||
collect_run "$dir"
|
||||
assert_equals 0 "$RUN_RC" "mode-preservation setup edit exit code ($want)"
|
||||
got="$(stat -f '%Lp' "$dir/fleetd.yaml" 2>/dev/null || stat -c '%a' "$dir/fleetd.yaml")"
|
||||
assert_equals "$want" "$got" "mode must survive a --set ($want)"
|
||||
|
||||
start_run "$dir" 5 --restore
|
||||
sleep 1
|
||||
printf 'config reloaded\n' >> "$dir/fleetd.out"
|
||||
collect_run "$dir"
|
||||
assert_equals 0 "$RUN_RC" "mode-preservation restore exit code ($want)"
|
||||
got="$(stat -f '%Lp' "$dir/fleetd.yaml" 2>/dev/null || stat -c '%a' "$dir/fleetd.yaml")"
|
||||
assert_equals "$want" "$got" "mode must survive a --restore ($want)"
|
||||
done
|
||||
}
|
||||
|
||||
# ----------------------------------------- acceptance criterion 14: restore message names the real directory
|
||||
# fleetd #635 follow-up (ticket comment 17664, defect 6) — the --restore "no backup found"
|
||||
# message used to print the OLD beside-the-config glob even though newest_backup had already
|
||||
# moved to searching the managed directory. Proves BOTH directions: the not-found message names
|
||||
# the directory actually searched (not merely that it says SOMETHING), and that a real backup
|
||||
# sitting in that directory still lets --restore succeed — otherwise the fix could regress into
|
||||
# a message that is always printed regardless of whether a backup exists.
|
||||
test_restore_message_names_the_searched_directory() {
|
||||
local dir rc=0
|
||||
|
||||
# Direction 1: no backup anywhere — the message must name .config-backups/, not the bare
|
||||
# beside-the-config glob the OLD code printed.
|
||||
dir="$(new_fixture)"
|
||||
"$EDIT" --restore --config "$dir/fleetd.yaml" --log "$dir/fleetd.out" --wait-seconds 2 \
|
||||
> "$dir/stdout.log" 2>&1 || rc=$?
|
||||
RUN_OUTPUT="$(cat "$dir/stdout.log")"
|
||||
|
||||
assert_equals 1 "$rc" "--restore with no backup anywhere exit code"
|
||||
assert_contains ".config-backups/fleetd.yaml.bak.*" "$RUN_OUTPUT" \
|
||||
"the not-found message must name the directory actually searched, not the old beside-the-config glob"
|
||||
|
||||
# Direction 2: a real backup IS present in .config-backups/ — --restore must still succeed, so
|
||||
# the message fix cannot have turned into one that prints regardless of whether a backup exists.
|
||||
dir="$(new_fixture)"
|
||||
cp "$dir/fleetd.yaml" "$dir/pre-edit.yaml"
|
||||
start_run "$dir" 5 --set '.profiles.sonnet.weight=77'
|
||||
sleep 1
|
||||
printf 'config reloaded\n' >> "$dir/fleetd.out"
|
||||
collect_run "$dir"
|
||||
assert_equals 0 "$RUN_RC" "setup edit exit code for criterion 14's second half"
|
||||
|
||||
start_run "$dir" 5 --restore
|
||||
sleep 1
|
||||
printf 'config reloaded\n' >> "$dir/fleetd.out"
|
||||
collect_run "$dir"
|
||||
assert_equals 0 "$RUN_RC" "--restore with a real backup present must still succeed"
|
||||
cmp -s "$dir/fleetd.yaml" "$dir/pre-edit.yaml" \
|
||||
|| fail "--restore with a real backup present must put the file back byte for byte"
|
||||
}
|
||||
|
||||
# ------------------------- acceptance criterion 15a: block-scalar continuation lines are redacted
|
||||
# fleetd #635 follow-up (ticket comment 17670, defect 7) — redact() used to look only AT the key
|
||||
# line. A YAML block scalar (`|`) puts its value on the lines that FOLLOW the key, each indented
|
||||
# deeper than it, so the real secret flowed through untouched while the key line right above it
|
||||
# printed a reassuring "<redacted>" — worse than no redaction, because the marker stops a reader
|
||||
# from looking further. The edited key here ("retries") sits directly next to the block scalar,
|
||||
# well inside diff -u's default 3-line context window, so the printed hunk is GUARANTEED to
|
||||
# include the secret's lines — placing the edit further away would let this pass today even
|
||||
# without the fix, proving nothing (the ticket comment's own warning, from the lead's first
|
||||
# reproduction attempt). The positive control runs FIRST: without it, "the secret never entered
|
||||
# the diff at all" would pass identically to "it entered and was correctly redacted".
|
||||
new_fixture_block_scalar() {
|
||||
local dir
|
||||
dir="$(mktemp -d "$TMP/fixture.XXXXXX")"
|
||||
cat > "$dir/fleetd.yaml" <<'YAML'
|
||||
bind:
|
||||
host: 127.0.0.1
|
||||
port: 19999
|
||||
broker:
|
||||
uri: amqp://user:hunter2@host/vhost
|
||||
auth:
|
||||
token: |
|
||||
FAKELEAK-BLOCK-SCALAR
|
||||
retries: 1
|
||||
profiles:
|
||||
sonnet:
|
||||
weight: 3
|
||||
maxLoad: 5
|
||||
YAML
|
||||
: > "$dir/fleetd.out"
|
||||
printf '%s' "$dir"
|
||||
}
|
||||
|
||||
test_block_scalar_continuation_is_redacted() {
|
||||
local dir
|
||||
dir="$(new_fixture_block_scalar)"
|
||||
|
||||
start_run "$dir" 5 --set '.auth.retries=2'
|
||||
sleep 1
|
||||
printf 'config reloaded\n' >> "$dir/fleetd.out"
|
||||
collect_run "$dir"
|
||||
|
||||
assert_equals 0 "$RUN_RC" "block-scalar case reload exit code"
|
||||
# Positive control FIRST: the key's own (masked) line must really be in the printed diff, or the
|
||||
# negative assertion right after proves nothing — see the comment above this test.
|
||||
assert_contains "token:" "$RUN_OUTPUT" "block-scalar case: the key's line must be in the printed diff"
|
||||
assert_contains "<redacted>" "$RUN_OUTPUT" "block-scalar case: redaction must be proven to have run on real content"
|
||||
assert_not_contains "FAKELEAK-BLOCK-SCALAR" "$RUN_OUTPUT" "block-scalar case: the block scalar's VALUE must never leak"
|
||||
}
|
||||
|
||||
# ------------------------------------- acceptance criterion 15b: "passphrase" is also recognised
|
||||
# "passphrase" was in none of TOKEN|SECRET|PASSWORD|PASSWD|CREDENTIAL|URI|_KEY (ticket comment
|
||||
# 17670). This is a plain key:value line, not a block scalar — kept in its OWN fixture and OWN
|
||||
# function, separate from criterion 15a, so that a failure in one case can never mask a failure in
|
||||
# the other (a single combined test would abort under `set -e` at its first failing assertion,
|
||||
# and the second case would then never even run).
|
||||
new_fixture_passphrase() {
|
||||
local dir
|
||||
dir="$(mktemp -d "$TMP/fixture.XXXXXX")"
|
||||
cat > "$dir/fleetd.yaml" <<'YAML'
|
||||
bind:
|
||||
host: 127.0.0.1
|
||||
port: 19999
|
||||
broker:
|
||||
uri: amqp://user:hunter2@host/vhost
|
||||
auth:
|
||||
passphrase: FAKELEAK-PASSPHRASE
|
||||
retries: 1
|
||||
profiles:
|
||||
sonnet:
|
||||
weight: 3
|
||||
maxLoad: 5
|
||||
YAML
|
||||
: > "$dir/fleetd.out"
|
||||
printf '%s' "$dir"
|
||||
}
|
||||
|
||||
test_passphrase_key_is_redacted() {
|
||||
local dir
|
||||
dir="$(new_fixture_passphrase)"
|
||||
|
||||
start_run "$dir" 5 --set '.auth.retries=2'
|
||||
sleep 1
|
||||
printf 'config reloaded\n' >> "$dir/fleetd.out"
|
||||
collect_run "$dir"
|
||||
|
||||
assert_equals 0 "$RUN_RC" "passphrase case reload exit code"
|
||||
assert_contains "passphrase:" "$RUN_OUTPUT" "passphrase case: the key's line must be in the printed diff"
|
||||
assert_contains "<redacted>" "$RUN_OUTPUT" "passphrase case: redaction must be proven to have run on real content"
|
||||
assert_not_contains "FAKELEAK-PASSPHRASE" "$RUN_OUTPUT" "passphrase case: the passphrase VALUE must never leak"
|
||||
}
|
||||
|
||||
# ------------------- acceptance criterion 19: the key line falls outside the printed hunk
|
||||
# fleetd #656 — criteria 15a/15b both put the edit right next to the key line, so the key line is
|
||||
# always inside diff -u's default 3-line context. Neither covers the actual case #639 fixed: an
|
||||
# 8-line block-scalar body with only its SIXTH line changed, so the printed hunk (3 lines of
|
||||
# context on each side of the change) covers body lines 3-8 and never includes the "token:" key
|
||||
# line at all. The old, line-by-line redact() only ever masks after it has SEEN the key line go
|
||||
# past; with the key line outside the hunk it never sets its mask, and the whole body — the
|
||||
# changed line included — passes through raw. The control key sits right after the body, inside
|
||||
# the same hunk, so the positive control below proves the fix is not simply printing nothing.
|
||||
new_fixture_hunk_without_key_line() {
|
||||
local dir
|
||||
dir="$(mktemp -d "$TMP/fixture.XXXXXX")"
|
||||
cat > "$dir/fleetd.yaml" <<'YAML'
|
||||
bind:
|
||||
host: 127.0.0.1
|
||||
port: 19999
|
||||
broker:
|
||||
uri: amqp://user:hunter2@host/vhost
|
||||
auth:
|
||||
token: |
|
||||
SECRET-LINE-1
|
||||
SECRET-LINE-2
|
||||
SECRET-LINE-3
|
||||
SECRET-LINE-4
|
||||
SECRET-LINE-5
|
||||
SECRET-LINE-6
|
||||
SECRET-LINE-7
|
||||
SECRET-LINE-8
|
||||
control: CTRL-MUST-APPEAR
|
||||
profiles:
|
||||
sonnet:
|
||||
weight: 3
|
||||
maxLoad: 5
|
||||
YAML
|
||||
: > "$dir/fleetd.out"
|
||||
printf '%s' "$dir"
|
||||
}
|
||||
|
||||
test_key_line_outside_hunk_is_still_redacted() {
|
||||
local dir
|
||||
dir="$(new_fixture_hunk_without_key_line)"
|
||||
sed 's/SECRET-LINE-6$/SECRET-LINE-6-CHANGED/' "$dir/fleetd.yaml" > "$dir/candidate.yaml"
|
||||
|
||||
start_run "$dir" 5 --from "$dir/candidate.yaml"
|
||||
sleep 1
|
||||
printf 'config reloaded\n' >> "$dir/fleetd.out"
|
||||
collect_run "$dir"
|
||||
|
||||
assert_equals 0 "$RUN_RC" "hunk-without-key-line case reload exit code"
|
||||
# Positive control FIRST: without this, a diff that printed nothing at all would pass the
|
||||
# negative assertion right below identically to a correctly redacted one.
|
||||
assert_contains "CTRL-MUST-APPEAR" "$RUN_OUTPUT" "hunk-without-key-line case: the non-secret control line must still print unmasked"
|
||||
assert_not_contains "SECRET-LINE-6-CHANGED" "$RUN_OUTPUT" "hunk-without-key-line case: the changed body line must never leak, even with the key line outside the printed hunk"
|
||||
}
|
||||
|
||||
# ------------------------------- acceptance criterion 20: a blank line inside the value
|
||||
# fleetd #656 — the old, line-by-line redact() reset its mask on any line whose indentation was
|
||||
# not STRICTLY greater than the key's, and a wholly blank line has indentation 0, so it reset the
|
||||
# mask exactly like the "control:" line that legitimately ends the block scalar. Everything after
|
||||
# the blank line then printed raw. The current fix tracks masked lines by FILE line number instead
|
||||
# of by indentation seen so far, so a blank line inside the value stays masked.
|
||||
new_fixture_blank_line_in_value() {
|
||||
local dir
|
||||
dir="$(mktemp -d "$TMP/fixture.XXXXXX")"
|
||||
printf 'bind:\n host: 127.0.0.1\n port: 19999\nbroker:\n uri: amqp://user:hunter2@host/vhost\nauth:\n token: |\n LEAK-BEFORE-BLANK\n\n LEAK-AFTER-BLANK\n control: CTRL-MUST-APPEAR\nprofiles:\n sonnet:\n weight: 3\n maxLoad: 5\n' > "$dir/fleetd.yaml"
|
||||
: > "$dir/fleetd.out"
|
||||
printf '%s' "$dir"
|
||||
}
|
||||
|
||||
test_blank_line_inside_value_is_still_redacted() {
|
||||
local dir
|
||||
dir="$(new_fixture_blank_line_in_value)"
|
||||
sed 's/LEAK-AFTER-BLANK$/LEAK-AFTER-BLANK-CHANGED/' "$dir/fleetd.yaml" > "$dir/candidate.yaml"
|
||||
|
||||
start_run "$dir" 5 --from "$dir/candidate.yaml"
|
||||
sleep 1
|
||||
printf 'config reloaded\n' >> "$dir/fleetd.out"
|
||||
collect_run "$dir"
|
||||
|
||||
assert_equals 0 "$RUN_RC" "blank-line-in-value case reload exit code"
|
||||
assert_contains "CTRL-MUST-APPEAR" "$RUN_OUTPUT" "blank-line-in-value case: the non-secret control line must still print unmasked"
|
||||
assert_not_contains "LEAK-AFTER-BLANK-CHANGED" "$RUN_OUTPUT" "blank-line-in-value case: the line after the blank must never leak"
|
||||
}
|
||||
|
||||
# ----------------------------------- acceptance criterion 16: a failing --set must not echo value
|
||||
# fleetd #635 follow-up (ticket comment 17673, defect 8) — apply_set_pairs used to echo the FULL
|
||||
# "$kv" (path=value, exactly as typed) in its yq-failure messages, so a broken --set with a
|
||||
# secret-looking value printed that value right back out. The path alone is what the positive
|
||||
# control proves is still there — it is what the operator needs to fix their command — and the
|
||||
# negative assertion proves the value itself never appears. Kept to exactly this one failure
|
||||
# shape (an invalid yq path/expression), matching the ticket's own reproduction.
|
||||
test_failing_set_does_not_echo_its_value() {
|
||||
local dir rc=0
|
||||
dir="$(new_fixture)"
|
||||
cp "$dir/fleetd.yaml" "$dir/pre-edit.yaml"
|
||||
|
||||
"$EDIT" --dry-run --set '.broker.["bad=FAKELEAK-SETVALUE' \
|
||||
--config "$dir/fleetd.yaml" --log "$dir/fleetd.out" --wait-seconds 2 \
|
||||
> "$dir/stdout.log" 2>&1 || rc=$?
|
||||
RUN_OUTPUT="$(cat "$dir/stdout.log")"
|
||||
|
||||
[ "$rc" -ne 0 ] || fail "a --set with an invalid yq expression must exit non-zero, got 0"
|
||||
cmp -s "$dir/fleetd.yaml" "$dir/pre-edit.yaml" \
|
||||
|| fail "a failing --set must install nothing — the live fixture changed"
|
||||
# Positive control FIRST: the path must still be in the message, or the negative assertion right
|
||||
# after proves nothing (the message could simply have disappeared entirely).
|
||||
assert_contains '.broker.["bad' "$RUN_OUTPUT" "the failure message must still name the PATH"
|
||||
assert_not_contains "FAKELEAK-SETVALUE" "$RUN_OUTPUT" "the failure message must NEVER echo the VALUE"
|
||||
}
|
||||
|
||||
# dry-run must never touch the live file and must still redact.
|
||||
test_dry_run_never_installs_and_redacts() {
|
||||
local dir before
|
||||
dir="$(new_fixture)"
|
||||
before="$(cat "$dir/fleetd.yaml")"
|
||||
|
||||
"$EDIT" --dry-run --set '.profiles.sonnet.weight=99' \
|
||||
--config "$dir/fleetd.yaml" --log "$dir/fleetd.out" --wait-seconds 2 \
|
||||
> "$dir/stdout.log" 2>&1
|
||||
local rc=$?
|
||||
RUN_OUTPUT="$(cat "$dir/stdout.log")"
|
||||
|
||||
assert_equals 0 "$rc" "dry-run exit code"
|
||||
assert_equals "$before" "$(cat "$dir/fleetd.yaml")" "dry-run must never write the live config"
|
||||
assert_not_contains "hunter2" "$RUN_OUTPUT" "dry-run diff must also be redacted"
|
||||
assert_contains "99" "$RUN_OUTPUT" "dry-run diff must show the candidate value"
|
||||
# Same positive-control reasoning as acceptance criterion 13, applied to the dry-run diff path.
|
||||
assert_contains "<redacted>" "$RUN_OUTPUT" "the dry-run diff's redaction must be PROVEN to have run, not merely absent"
|
||||
}
|
||||
|
||||
# --check is read-only and always exits 0, even against a dead "daemon".
|
||||
test_check_is_read_only_and_exits_zero() {
|
||||
local dir before rc=0
|
||||
dir="$(new_fixture)"
|
||||
before="$(cat "$dir/fleetd.yaml")"
|
||||
|
||||
"$EDIT" --check --config "$dir/fleetd.yaml" --log "$dir/fleetd.out" \
|
||||
> "$dir/stdout.log" 2>&1 || rc=$?
|
||||
|
||||
assert_equals 0 "$rc" "--check exit code"
|
||||
assert_equals "$before" "$(cat "$dir/fleetd.yaml")" "--check must never modify the config"
|
||||
}
|
||||
|
||||
test_refusal_shape_from_parse_failure_wording_is_recognised() {
|
||||
local dir
|
||||
dir="$(new_fixture)"
|
||||
|
||||
start_run "$dir" 5 --set '.profiles.sonnet.weight=6'
|
||||
sleep 1
|
||||
printf 'config reload from %s refused, keeping the running config: boom\n' "$dir/fleetd.yaml" >> "$dir/fleetd.out"
|
||||
collect_run "$dir"
|
||||
|
||||
assert_equals 4 "$RUN_RC" "the parse-failure refusal shape must also exit 4, not be read as silence"
|
||||
}
|
||||
|
||||
# A verdict line carrying a credentialed URI has its userinfo masked, with a positive control
|
||||
# proving the rest of the line still reaches the output unchanged.
|
||||
test_verdict_userinfo_is_masked_with_positive_control() {
|
||||
local dir
|
||||
dir="$(new_fixture)"
|
||||
|
||||
start_run "$dir" 5 --set '.profiles.sonnet.weight=4'
|
||||
sleep 1
|
||||
printf 'config reload from %s refused, keeping the running config: refusing to start: malformed pattern — profiles.local.errorPattern ("amqp://user:hunter2@host/vhost"): Unclosed character class near index 8\n' \
|
||||
"$dir/fleetd.yaml" >> "$dir/fleetd.out"
|
||||
collect_run "$dir"
|
||||
|
||||
assert_equals 4 "$RUN_RC" "refusal-with-userinfo exit code"
|
||||
assert_not_contains "user:hunter2" "$RUN_OUTPUT" "the userinfo must never reach the output"
|
||||
assert_contains "amqp://<redacted>@host/vhost" "$RUN_OUTPUT" \
|
||||
"the userinfo must be MASKED, not deleted — the rest of the quoted value must survive"
|
||||
# Positive control: the diagnostic prose on both sides of the userinfo must still reach the
|
||||
# output. Without this, a mutant that drops the whole verdict line would pass identically.
|
||||
assert_contains "malformed pattern" "$RUN_OUTPUT" "prose BEFORE the userinfo must still reach the output"
|
||||
assert_contains "Unclosed character class near index 8" "$RUN_OUTPUT" \
|
||||
"prose AFTER the userinfo must still reach the output"
|
||||
}
|
||||
|
||||
# An ordinary refusal line quotes the offending pattern, not a credential, and must survive byte
|
||||
# for byte: the rewrite is scoped to userinfo only, and the quoted pattern is the detail an
|
||||
# operator needs to fix the refusal.
|
||||
test_ordinary_refusal_line_passes_through_unchanged() {
|
||||
local dir real_line
|
||||
dir="$(new_fixture)"
|
||||
real_line="config reload from $dir/fleetd.yaml refused, keeping the running config: refusing to start: malformed pattern — profiles.local.errorPattern (\"[unclosed\"): Unclosed character class near index 8"
|
||||
|
||||
start_run "$dir" 5 --set '.profiles.sonnet.weight=4'
|
||||
sleep 1
|
||||
printf '%s\n' "$real_line" >> "$dir/fleetd.out"
|
||||
collect_run "$dir"
|
||||
|
||||
assert_equals 4 "$RUN_RC" "ordinary refusal exit code"
|
||||
assert_contains "$real_line" "$RUN_OUTPUT" \
|
||||
"an ordinary refusal with no userinfo must pass through byte for byte, unchanged"
|
||||
}
|
||||
|
||||
# A verdict line can hold a URI with NO userinfo and a later, unrelated @ further on in the same
|
||||
# line (an email address in diagnostic prose, for example). The rewrite must stop at the end of
|
||||
# the URI and must not treat the later @ as a second userinfo delimiter.
|
||||
test_uri_without_userinfo_survives_a_later_at_sign() {
|
||||
local dir real_line
|
||||
dir="$(new_fixture)"
|
||||
real_line="config reload from $dir/fleetd.yaml refused, keeping the running config: broker.uri amqp://broker.local/vhost unreachable, contact ops@example.com"
|
||||
|
||||
start_run "$dir" 5 --set '.profiles.sonnet.weight=4'
|
||||
sleep 1
|
||||
printf '%s\n' "$real_line" >> "$dir/fleetd.out"
|
||||
collect_run "$dir"
|
||||
|
||||
assert_equals 4 "$RUN_RC" "no-userinfo-with-later-at-sign exit code"
|
||||
assert_contains "$real_line" "$RUN_OUTPUT" \
|
||||
"a URI with no userinfo plus a later @ in the same line must pass through byte for byte"
|
||||
}
|
||||
|
||||
# --set runs yq over the whole candidate. It warns when that changes more lines than the requested
|
||||
# pairs, but a simple file with only the intended changed line must stay quiet.
|
||||
new_fixture_reformat_sensitive() {
|
||||
local dir
|
||||
dir="$(mktemp -d "$TMP/fixture.XXXXXX")"
|
||||
cat > "$dir/fleetd.yaml" <<'YAML'
|
||||
# A section comment that documents the next block.
|
||||
bind:
|
||||
host: 127.0.0.1 # Keep this aligned with the port note.
|
||||
port: 19999 # A fixture port.
|
||||
|
||||
# These comments use their placement as documentation.
|
||||
profiles:
|
||||
sonnet:
|
||||
weight: 3
|
||||
bootstrapText: >-
|
||||
First line.
|
||||
Second line.
|
||||
YAML
|
||||
: > "$dir/fleetd.out"
|
||||
printf '%s' "$dir"
|
||||
}
|
||||
|
||||
test_set_warns_when_yq_reformats_extra_lines() {
|
||||
local dir
|
||||
dir="$(new_fixture_reformat_sensitive)"
|
||||
|
||||
start_run "$dir" 5 --set '.profiles.sonnet.weight=4'
|
||||
sleep 1
|
||||
printf 'config reloaded\n' >> "$dir/fleetd.out"
|
||||
collect_run "$dir"
|
||||
|
||||
assert_equals 0 "$RUN_RC" "reformat warning case reload exit code"
|
||||
assert_contains "yq reformatted the whole file" "$RUN_OUTPUT" \
|
||||
"a --set that changes extra candidate lines must warn before installation"
|
||||
}
|
||||
|
||||
test_set_stays_quiet_without_formatting_churn() {
|
||||
local dir
|
||||
dir="$(new_fixture)"
|
||||
|
||||
start_run "$dir" 5 --set '.profiles.sonnet.weight=4'
|
||||
sleep 1
|
||||
printf 'config reloaded\n' >> "$dir/fleetd.out"
|
||||
collect_run "$dir"
|
||||
|
||||
assert_equals 0 "$RUN_RC" "no-reformat warning case reload exit code"
|
||||
assert_not_contains "yq reformatted the whole file" "$RUN_OUTPUT" \
|
||||
"a --set that changes only its requested candidate line must not warn"
|
||||
}
|
||||
|
||||
echo "== acceptance criterion 1: refusal restores byte for byte =="
|
||||
test_refusal_restores_byte_for_byte
|
||||
echo "== acceptance criterion 2: clean reload keeps the edit =="
|
||||
test_clean_reload_keeps_the_edit
|
||||
echo "== acceptance criterion 3: deferred reload told apart from clean =="
|
||||
test_deferred_reload_is_told_apart_from_clean
|
||||
echo "== acceptance criterion 4: silence is its own answer =="
|
||||
test_silence_is_its_own_answer
|
||||
echo "== acceptance criterion 5: broken candidate never reaches the live path =="
|
||||
test_broken_candidate_never_reaches_live_path
|
||||
echo "== acceptance criterion 6: the marker works =="
|
||||
test_marker_skips_lines_before_it
|
||||
echo "== acceptance criterion 7 (+13: redaction is proven to have run) =="
|
||||
test_redaction_holds
|
||||
echo "== acceptance criterion 9: a forgotten value refuses and installs nothing =="
|
||||
test_forgotten_value_refuses_and_installs_nothing
|
||||
echo "== acceptance criterion 10: an explicit clear writes a bare null =="
|
||||
test_explicit_null_writes_bare_null_not_empty_string
|
||||
echo "== acceptance criterion 11: a backup is never committable =="
|
||||
test_backup_is_never_committable
|
||||
echo "== acceptance criterion 12: the file mode survives an edit and a restore =="
|
||||
test_file_mode_survives_edit_and_restore
|
||||
echo "== acceptance criterion 14: the restore message names the directory actually searched =="
|
||||
test_restore_message_names_the_searched_directory
|
||||
echo "== acceptance criterion 15a: a block scalar's continuation lines are redacted =="
|
||||
test_block_scalar_continuation_is_redacted
|
||||
echo "== acceptance criterion 15b: a passphrase key is also recognised =="
|
||||
test_passphrase_key_is_redacted
|
||||
echo "== acceptance criterion 19: the key line falls outside the printed hunk =="
|
||||
test_key_line_outside_hunk_is_still_redacted
|
||||
echo "== acceptance criterion 20: a blank line inside the value =="
|
||||
test_blank_line_inside_value_is_still_redacted
|
||||
echo "== acceptance criterion 16: a failing --set must not echo its value =="
|
||||
test_failing_set_does_not_echo_its_value
|
||||
echo "== extra: dry-run never installs, and redacts =="
|
||||
test_dry_run_never_installs_and_redacts
|
||||
echo "== extra: --check is read-only and always exits 0 =="
|
||||
test_check_is_read_only_and_exits_zero
|
||||
echo "== extra: the parse-failure refusal shape is also recognised =="
|
||||
test_refusal_shape_from_parse_failure_wording_is_recognised
|
||||
echo "== verdict-redaction criteria 2+3: verdict userinfo is masked, rest of line survives =="
|
||||
test_verdict_userinfo_is_masked_with_positive_control
|
||||
echo "== verdict-redaction criterion 4: an ordinary refusal passes through unchanged =="
|
||||
test_ordinary_refusal_line_passes_through_unchanged
|
||||
echo "== fleetd #638: a URI with no userinfo survives a later @ in the same line =="
|
||||
test_uri_without_userinfo_survives_a_later_at_sign
|
||||
echo "== acceptance criterion 17: --set warns about yq formatting churn =="
|
||||
test_set_warns_when_yq_reformats_extra_lines
|
||||
echo "== acceptance criterion 18: --set stays quiet without formatting churn =="
|
||||
test_set_stays_quiet_without_formatting_churn
|
||||
|
||||
printf 'PASS: config-edit acceptance criteria\n'
|
||||
+213
-63
@@ -447,7 +447,7 @@ test_assert_single_daemon_rejects_two_pids() {
|
||||
test_running_pid_excludes_self_matching_wrapper_shell() {
|
||||
local before after wrapper_pid
|
||||
before="$(running_pid)"
|
||||
sh -c 'echo "target/fleetd.jar" >/dev/null; sleep 20' &
|
||||
sh -c 'echo "run/fleetd.jar" >/dev/null; sleep 20' &
|
||||
wrapper_pid=$!
|
||||
sleep 0.3
|
||||
after="$(running_pid)"
|
||||
@@ -475,11 +475,26 @@ test_running_pid_excludes_self_matching_wrapper_shell() {
|
||||
# `comm` from the actually-executed binary's own path, not from `exec -a`'s argv[0] override (BSD
|
||||
# ties `comm` to argv[0], which is what makes this technique work here) — so on Linux this
|
||||
# specific fixture might report `comm=sh`, not `comm=java`, even though the REAL daemon (a literal
|
||||
# `java -jar target/fleetd.jar` process, never fabricated) is unaffected either way. I could not
|
||||
# `java -jar run/fleetd.jar` process, never fabricated) is unaffected either way. I could not
|
||||
# verify this fixture's behavior on Linux, so test_running_pid_counts_a_pid_whose_comm_is_java
|
||||
# below backstops the same claim (the allowlist admits a pid whose comm is `java`) with a stubbed
|
||||
# `ps`, which is identical bash on every platform and carries no such platform question.
|
||||
test_running_pid_finds_a_real_java_named_second_process() {
|
||||
local before after standin_pid
|
||||
before="$(running_pid)"
|
||||
( exec -a java sh -c 'echo "run/fleetd.jar" >/dev/null; sleep 20' ) &
|
||||
standin_pid=$!
|
||||
sleep 0.3
|
||||
after="$(running_pid)"
|
||||
kill "$standin_pid" 2>/dev/null || true
|
||||
wait "$standin_pid" 2>/dev/null || true
|
||||
printf '%s\n' "$after" | grep -qxF "$standin_pid" \
|
||||
|| fail "running_pid() did not find a real second process (pid $standin_pid, comm forced to 'java' via exec -a) whose own argv holds the pattern: before=[$before] after=[$after]"
|
||||
}
|
||||
|
||||
# PATTERN matches a fleetd jar in either build layout, not only the run/ one: a process whose
|
||||
# argv names a jar under target/ must be found too, the same way the run/ case above is.
|
||||
test_running_pid_finds_a_real_java_named_process_from_target_dir() {
|
||||
local before after standin_pid
|
||||
before="$(running_pid)"
|
||||
( exec -a java sh -c 'echo "target/fleetd.jar" >/dev/null; sleep 20' ) &
|
||||
@@ -489,7 +504,24 @@ test_running_pid_finds_a_real_java_named_second_process() {
|
||||
kill "$standin_pid" 2>/dev/null || true
|
||||
wait "$standin_pid" 2>/dev/null || true
|
||||
printf '%s\n' "$after" | grep -qxF "$standin_pid" \
|
||||
|| fail "running_pid() did not find a real second process (pid $standin_pid, comm forced to 'java' via exec -a) whose own argv holds the pattern: before=[$before] after=[$after]"
|
||||
|| fail "running_pid() did not find a real second process (pid $standin_pid, comm forced to 'java' via exec -a) naming a jar under target/: before=[$before] after=[$after]"
|
||||
}
|
||||
|
||||
# Broadening PATTERN to match both build layouts must not also broaden it into matching a
|
||||
# non-exec'ing shell that merely holds the target/ text as a literal argument, the same
|
||||
# self-matching shape test_running_pid_excludes_self_matching_wrapper_shell above excludes for
|
||||
# the run/ text.
|
||||
test_running_pid_excludes_self_matching_wrapper_shell_naming_target_dir() {
|
||||
local before after wrapper_pid
|
||||
before="$(running_pid)"
|
||||
sh -c 'echo "target/fleetd.jar" >/dev/null; sleep 20' &
|
||||
wrapper_pid=$!
|
||||
sleep 0.3
|
||||
after="$(running_pid)"
|
||||
kill "$wrapper_pid" 2>/dev/null || true
|
||||
wait "$wrapper_pid" 2>/dev/null || true
|
||||
[ "$after" = "$before" ] \
|
||||
|| fail "running_pid() counted a self-matching wrapper shell (pid $wrapper_pid, holding 'target/fleetd.jar' as literal text in its own argv, not the daemon): before=[$before] after=[$after]"
|
||||
}
|
||||
|
||||
# fleetd #593 CORRECTION 1, hole 2 — the round-1 filter denied known shell names (sh/bash/zsh/
|
||||
@@ -557,29 +589,30 @@ test_die_message_does_not_recommend_bare_pgrep_as_remediation() {
|
||||
|| fail "assert_single_daemon's die message does not say in words that a pattern can match the caller (fleetd #593)"
|
||||
}
|
||||
|
||||
# fleetd #511 — jar_id()'s no-argument default was unpinned by any test: nothing proved it reports
|
||||
# $JAR (the live path) rather than $JAR_STAGED. Both halves matter, so this pins both: the bare call
|
||||
# must hash the live jar, and an explicit path argument must hash THAT file, not fall back to $JAR.
|
||||
# Two files with different content, so a default pointed at the wrong one reports the wrong hash
|
||||
# rather than accidentally matching.
|
||||
# fleetd #511/#664 — jar_id()'s no-argument default was unpinned by any test: nothing proved it
|
||||
# reports $JAR (the live path) rather than whatever explicit path a caller passes it (e.g.
|
||||
# $BUILD_JAR). Both halves matter, so this pins both: the bare call must hash the live jar, and an
|
||||
# explicit path argument must hash THAT file, not fall back to $JAR. Two files with different
|
||||
# content, so a default pointed at the wrong one reports the wrong hash rather than accidentally
|
||||
# matching.
|
||||
test_jar_id_defaults_to_live_and_reports_explicit_path() {
|
||||
local dir saved_jar="$JAR" saved_staged="$JAR_STAGED"
|
||||
local live_hash staged_hash default_result explicit_result
|
||||
local dir saved_jar="$JAR"
|
||||
local live_hash other_hash default_result explicit_result other_path
|
||||
dir="$TMP/jar-id"; mkdir -p "$dir"
|
||||
JAR="$dir/fleetd.jar"; JAR_STAGED="$dir/fleetd-new.jar"
|
||||
JAR="$dir/fleetd.jar"; other_path="$dir/other.jar"
|
||||
printf 'live jar bytes' > "$JAR"
|
||||
printf 'staged jar bytes, not the same content' > "$JAR_STAGED"
|
||||
printf 'other jar bytes, not the same content' > "$other_path"
|
||||
# fleetd #550: this reference hash must be computed the same portable way jar_id() itself now
|
||||
# computes one — a bare, unguarded call to the macOS-only hasher here was exactly the item-2
|
||||
# defect, dying with "command not found" on any Linux runner that has no such hasher at all.
|
||||
live_hash="$(hash256 "$JAR")"
|
||||
staged_hash="$(hash256 "$JAR_STAGED")"
|
||||
other_hash="$(hash256 "$other_path")"
|
||||
default_result="$(jar_id)"
|
||||
explicit_result="$(jar_id "$JAR_STAGED")"
|
||||
JAR="$saved_jar"; JAR_STAGED="$saved_staged"
|
||||
[ "$live_hash" != "$staged_hash" ] || fail "test fixture error: live and staged jars hashed the same"
|
||||
explicit_result="$(jar_id "$other_path")"
|
||||
JAR="$saved_jar"
|
||||
[ "$live_hash" != "$other_hash" ] || fail "test fixture error: live and other jars hashed the same"
|
||||
assert_equals "$live_hash" "$default_result" "jar_id with no arguments must report the hash of \$JAR"
|
||||
assert_equals "$staged_hash" "$explicit_result" "jar_id \"\$JAR_STAGED\" must report the hash of the staged jar, not fall back to \$JAR"
|
||||
assert_equals "$other_hash" "$explicit_result" "jar_id with an explicit path must report the hash of that path, not fall back to \$JAR"
|
||||
}
|
||||
|
||||
# fleetd #550 — closes a gap the test above leaves open. That test's own reference hash is now ALSO
|
||||
@@ -651,34 +684,77 @@ test_jar_id_reports_unhashable_when_no_hasher_on_path() {
|
||||
assert_equals "unhashable" "$explicit_result" "jar_id (explicit path) with no hasher on PATH must report the same third state"
|
||||
}
|
||||
|
||||
# fleetd #493 — never build into the path a running process holds. stage_built_jar/swap_staged_jar
|
||||
# are exercised directly against real files on disk (not stubs), because the whole point is file
|
||||
# fleetd #664 — report_jar_state is the --check fix: under the old layout $JAR and the build
|
||||
# output were the same file, so a single "jar on disk" fact covered both. Now they can disagree,
|
||||
# and this is the function that is supposed to show that. Agreeing case: two files with IDENTICAL
|
||||
# content must print both labels and never warn.
|
||||
test_report_jar_state_agrees_when_hashes_match() {
|
||||
local dir saved_build="$BUILD_JAR" saved_jar="$JAR" output
|
||||
dir="$TMP/report-jar-agree"; mkdir -p "$dir"
|
||||
BUILD_JAR="$dir/target-fleetd.jar"; JAR="$dir/run-fleetd.jar"
|
||||
printf 'identical jar bytes' > "$BUILD_JAR"
|
||||
printf 'identical jar bytes' > "$JAR"
|
||||
output="$(report_jar_state)"
|
||||
BUILD_JAR="$saved_build"; JAR="$saved_jar"
|
||||
printf '%s' "$output" | grep -qF 'built jar' \
|
||||
|| fail "report_jar_state did not label the built jar"
|
||||
printf '%s' "$output" | grep -qF 'running jar' \
|
||||
|| fail "report_jar_state did not label the running jar"
|
||||
if printf '%s' "$output" | grep -qF 'differ'; then
|
||||
fail "report_jar_state warned about a mismatch when both jars have identical content"
|
||||
fi
|
||||
}
|
||||
|
||||
# The disagreeing case: this is the whole point of the ticket — a built jar that is NOT the
|
||||
# running jar must be visibly flagged, not silently printed as two unremarkable facts.
|
||||
test_report_jar_state_warns_when_hashes_differ() {
|
||||
local dir saved_build="$BUILD_JAR" saved_jar="$JAR" output
|
||||
dir="$TMP/report-jar-differ"; mkdir -p "$dir"
|
||||
BUILD_JAR="$dir/target-fleetd.jar"; JAR="$dir/run-fleetd.jar"
|
||||
printf 'freshly built jar bytes' > "$BUILD_JAR"
|
||||
printf 'older running jar bytes' > "$JAR"
|
||||
output="$(report_jar_state)"
|
||||
BUILD_JAR="$saved_build"; JAR="$saved_jar"
|
||||
printf '%s' "$output" | grep -qF 'differ' \
|
||||
|| fail "report_jar_state did not warn when the built jar and running jar disagree"
|
||||
}
|
||||
|
||||
# Neither file existing (a fresh checkout, never built or deployed) must report two "absent"
|
||||
# facts and never a false mismatch warning — "absent" vs "absent" is agreement, not a diff.
|
||||
test_report_jar_state_both_absent_is_not_a_mismatch() {
|
||||
local dir saved_build="$BUILD_JAR" saved_jar="$JAR" output
|
||||
dir="$TMP/report-jar-absent"; mkdir -p "$dir"
|
||||
BUILD_JAR="$dir/no-such-target.jar"; JAR="$dir/no-such-run.jar"
|
||||
output="$(report_jar_state)"
|
||||
BUILD_JAR="$saved_build"; JAR="$saved_jar"
|
||||
printf '%s' "$output" | grep -qF 'absent' \
|
||||
|| fail "report_jar_state did not report absent for a missing built/running jar"
|
||||
if printf '%s' "$output" | grep -qF 'differ'; then
|
||||
fail "report_jar_state warned about a mismatch when both jars are simply absent"
|
||||
fi
|
||||
}
|
||||
|
||||
# Reads JAR and BUILD_JAR exactly as the script sources them, with nothing here assigning
|
||||
# either first. JAR must resolve outside $MODULE/target/, and JAR must differ from BUILD_JAR:
|
||||
# the daemon's live path and Maven's own build output are never the same file.
|
||||
test_jar_and_build_jar_are_sourced_outside_target_and_differ() {
|
||||
source "$ROOT/scripts/redeploy-fleetd.sh"
|
||||
case "$JAR" in
|
||||
"$MODULE"/target/*)
|
||||
fail "\$JAR must not live under \$MODULE/target/ — got $JAR" ;;
|
||||
esac
|
||||
[ "$JAR" != "$BUILD_JAR" ] \
|
||||
|| fail "\$JAR and \$BUILD_JAR must not be the same path — got $JAR"
|
||||
}
|
||||
|
||||
# fleetd #493/#664 — never build into the path a running process holds. swap_staged_jar is
|
||||
# exercised directly against real files on disk (not stubs), because the whole point is file
|
||||
# behavior (does the content move, does the source disappear, does a failure leave both sides
|
||||
# intact) that a stubbed function cannot prove.
|
||||
test_stage_built_jar_moves_off_live_path() {
|
||||
local dir jar staged saved_jar="$JAR" saved_staged="$JAR_STAGED"
|
||||
dir="$TMP/stage-ok"; mkdir -p "$dir"
|
||||
jar="$dir/fleetd.jar"; staged="$dir/fleetd-new.jar"
|
||||
printf 'built jar bytes' > "$jar"
|
||||
JAR="$jar"; JAR_STAGED="$staged"
|
||||
stage_built_jar || fail "stage_built_jar rejected a real build output"
|
||||
JAR="$saved_jar"; JAR_STAGED="$saved_staged"
|
||||
[ ! -f "$jar" ] || fail "stage_built_jar left the jar behind at the live path $jar"
|
||||
[ -f "$staged" ] || fail "stage_built_jar did not create the staged jar at $staged"
|
||||
grep -qF 'built jar bytes' "$staged" || fail "staged jar does not carry the built content"
|
||||
}
|
||||
|
||||
test_stage_built_jar_dies_when_build_produced_nothing() {
|
||||
local dir output rc=0 saved_jar="$JAR" saved_staged="$JAR_STAGED"
|
||||
dir="$TMP/stage-missing"; mkdir -p "$dir"
|
||||
JAR="$dir/fleetd.jar"; JAR_STAGED="$dir/fleetd-new.jar"
|
||||
output="$(stage_built_jar 2>&1)" || rc=$?
|
||||
JAR="$saved_jar"; JAR_STAGED="$saved_staged"
|
||||
[ "$rc" -ne 0 ] || fail "stage_built_jar accepted a missing build output"
|
||||
printf '%s' "$output" | grep -qF "$dir/fleetd.jar" \
|
||||
|| fail "refusal message does not name the missing jar path"
|
||||
}
|
||||
|
||||
# intact) that a stubbed function cannot prove. stage_built_jar no longer exists: under the #664
|
||||
# layout $BUILD_JAR (target/fleetd.jar) was never the live path, so a build has nothing to be
|
||||
# staged OUT of — swap_staged_jar is called directly against $BUILD_JAR/$JAR (see
|
||||
# test_swap_ordered_after_wait_and_before_start and the report_jar_state tests below for the rest
|
||||
# of that seam).
|
||||
test_swap_staged_jar_moves_staged_onto_live() {
|
||||
local dir staged live
|
||||
dir="$TMP/swap-ok"; mkdir -p "$dir"
|
||||
@@ -989,24 +1065,26 @@ test_swap_ordered_after_wait_and_before_start() {
|
||||
|| fail "swap_if_built (line $swap_line) is not before the start section (line $start_line)"
|
||||
}
|
||||
|
||||
# fleetd #511: the drain-gate abort message (fired when a build has staged a jar but the operator
|
||||
# declines the drain confirmation) used to tell the operator to "Rerun (with or without --no-build)"
|
||||
# to finish the restart. That is wrong — by the time this message can fire, stage_built_jar has
|
||||
# already moved the jar off $JAR, so a rerun WITH --no-build hits require_no_build_jar's own refusal
|
||||
# ("no jar at $JAR — run without --no-build"). Like test_swap_ordered_after_wait_and_before_start
|
||||
# above, this code path is never reached by sourcing (the SOURCED guard stops before the main flow),
|
||||
# so the only way to pin its exact wording is to read the source.
|
||||
# fleetd #511: the drain-gate abort message (fired when a build has produced a jar but the
|
||||
# operator declines the drain confirmation) used to tell the operator to "Rerun (with or without
|
||||
# --no-build)" to finish the restart. That is wrong. fleetd #664 changed WHY it is wrong: under
|
||||
# the old layout a build emptied the live path immediately, so --no-build's own check refused a
|
||||
# rerun for you; now a build never touches the live path at all, so --no-build would NOT refuse —
|
||||
# it would quietly restart the daemon on the OLD jar and throw away the one just built. Like
|
||||
# test_swap_ordered_after_wait_and_before_start above, this code path is never reached by sourcing
|
||||
# (the SOURCED guard stops before the main flow), so the only way to pin its exact wording is to
|
||||
# read the source.
|
||||
test_drain_gate_abort_message_says_no_no_build() {
|
||||
local src="$ROOT/scripts/redeploy-fleetd.sh" msg
|
||||
msg="$(grep -A3 -F 'aborted — the running daemon was NOT touched, but the freshly built jar is sitting at' "$src")"
|
||||
msg="$(grep -A6 -F 'aborted — the running daemon was NOT touched, but the freshly built jar is sitting at' "$src")"
|
||||
[ -n "$msg" ] || fail "could not find the drain-gate staged-jar abort message in redeploy-fleetd.sh"
|
||||
if printf '%s' "$msg" | grep -qF 'with or without --no-build'; then
|
||||
fail "abort message still claims a rerun WITH --no-build can finish the restart"
|
||||
fi
|
||||
printf '%s' "$msg" | grep -qF 'WITHOUT --no-build' \
|
||||
|| fail "abort message does not tell the operator to rerun without --no-build"
|
||||
printf '%s' "$msg" | grep -qF 'no longer at the live path' \
|
||||
|| fail "abort message does not say why --no-build cannot finish the restart"
|
||||
printf '%s' "$msg" | grep -qF 'silently discard' \
|
||||
|| fail "abort message does not say that --no-build would silently discard the build just made"
|
||||
}
|
||||
|
||||
# fleetd #517 — the drain-gate abort branch itself. Before this, the only test of this message was
|
||||
@@ -1149,19 +1227,22 @@ test_refuse_drain_gate_no_build_staged_absent() {
|
||||
#
|
||||
# fleetd #555 — the main flow's own call site moved: it used to read
|
||||
# `refuse_drain_gate "$DO_BUILD" "$JAR_STAGED"` directly; it now reads
|
||||
# `run_drain_gate "$OLD_PID" "$ASSUME_YES" "$DO_BUILD" "$JAR_STAGED"`, and run_drain_gate (tested
|
||||
# directly below by test_run_drain_gate_*) is what calls refuse_drain_gate with its own local names.
|
||||
# This grep now pins THAT call site — the thing that would go missing if a future edit deleted the
|
||||
# main flow's call to run_drain_gate altogether, the same residual gap #521/#528 already accepted for
|
||||
# swap_if_built/refuse_drain_gate (sourcing stops before the main flow runs, so no test in this file
|
||||
# can do better than reading the source for this one specific gap).
|
||||
# `run_drain_gate "$OLD_PID" "$ASSUME_YES" "$DO_BUILD" "$BUILD_JAR"` (fleetd #664 renamed the
|
||||
# fourth argument from $JAR_STAGED to $BUILD_JAR — same role, the not-yet-swapped-in jar — when
|
||||
# that path stopped being a separate staging file and became target/fleetd.jar itself), and
|
||||
# run_drain_gate (tested directly below by test_run_drain_gate_*) is what calls refuse_drain_gate
|
||||
# with its own local names. This grep now pins THAT call site — the thing that would go missing if
|
||||
# a future edit deleted the main flow's call to run_drain_gate altogether, the same residual gap
|
||||
# #521/#528 already accepted for swap_if_built/refuse_drain_gate (sourcing stops before the main
|
||||
# flow runs, so no test in this file can do better than reading the source for this one specific
|
||||
# gap).
|
||||
#
|
||||
# The grep ends `|| true`: this file runs under `set -euo pipefail`, so an ABSENT needle would fail
|
||||
# the assignment and `set -e` would kill the whole suite before the `[ -n ... ] || fail` guard below
|
||||
# ever ran — the exact dead-check shape fleetd #528 also flags as a sweep finding (see the PR body).
|
||||
test_run_drain_gate_call_site_present() {
|
||||
local src="$ROOT/scripts/redeploy-fleetd.sh" call_line
|
||||
call_line="$(grep -Fn 'run_drain_gate "$OLD_PID" "$ASSUME_YES" "$DO_BUILD" "$JAR_STAGED"' "$src" | head -1 | cut -d: -f1 || true)"
|
||||
call_line="$(grep -Fn 'run_drain_gate "$OLD_PID" "$ASSUME_YES" "$DO_BUILD" "$BUILD_JAR"' "$src" | head -1 | cut -d: -f1 || true)"
|
||||
[ -n "$call_line" ] \
|
||||
|| fail "could not find the main flow's run_drain_gate call site in redeploy-fleetd.sh"
|
||||
}
|
||||
@@ -1269,12 +1350,71 @@ test_run_drain_gate_declined_reply_refuses() {
|
||||
source "$ROOT/scripts/redeploy-fleetd.sh"
|
||||
}
|
||||
|
||||
# Writes a launchd-plist fixture naming jar_path as the ProgramArguments entry after "-jar", so
|
||||
# check_jar_path_matches_plist has something real to read back.
|
||||
write_launchd_plist_fixture() {
|
||||
local path="$1" jar_path="$2"
|
||||
cat > "$path" <<PLIST
|
||||
<?xml version="1.0" encoding="UTF-8"?>
|
||||
<!DOCTYPE plist PUBLIC "-//Apple//DTD PLIST 1.0//EN" "http://www.apple.com/DTDs/PropertyList-1.0.dtd">
|
||||
<plist version="1.0">
|
||||
<dict>
|
||||
<key>Label</key>
|
||||
<string>test.fixture</string>
|
||||
<key>ProgramArguments</key>
|
||||
<array>
|
||||
<string>/usr/bin/java</string>
|
||||
<string>-jar</string>
|
||||
<string>$jar_path</string>
|
||||
<string>fleetd.yaml</string>
|
||||
</array>
|
||||
</dict>
|
||||
</plist>
|
||||
PLIST
|
||||
}
|
||||
|
||||
# Agreeing case: a plist whose ProgramArguments names the same jar, resolved, must proceed and
|
||||
# say so, never die.
|
||||
test_check_jar_path_matches_plist_agrees_ok() {
|
||||
local dir jar plist output rc=0
|
||||
dir="$TMP/jar-path-agree"; mkdir -p "$dir/run"
|
||||
jar="$dir/run/fleetd.jar"
|
||||
plist="$dir/agree.plist"
|
||||
write_launchd_plist_fixture "$plist" "$jar"
|
||||
output="$(check_jar_path_matches_plist "$jar" "$plist" 2>&1)" || rc=$?
|
||||
[ "$rc" -eq 0 ] \
|
||||
|| fail "check_jar_path_matches_plist must succeed when the plist names the same jar: $output"
|
||||
printf '%s' "$output" | grep -qF 'jar path check' \
|
||||
|| fail "check_jar_path_matches_plist did not print the agreement line: $output"
|
||||
}
|
||||
|
||||
# The disagreeing case, and the positive control this check exists for: a plist naming a
|
||||
# different jar must die, naming both paths.
|
||||
test_check_jar_path_matches_plist_mismatch_dies() {
|
||||
local dir jar other_jar plist output rc=0
|
||||
dir="$TMP/jar-path-mismatch"; mkdir -p "$dir/run" "$dir/target"
|
||||
jar="$dir/run/fleetd.jar"
|
||||
other_jar="$dir/target/fleetd.jar"
|
||||
plist="$dir/mismatch.plist"
|
||||
write_launchd_plist_fixture "$plist" "$other_jar"
|
||||
output="$(check_jar_path_matches_plist "$jar" "$plist" 2>&1)" || rc=$?
|
||||
[ "$rc" -ne 0 ] \
|
||||
|| fail "check_jar_path_matches_plist must die when the plist names a different jar"
|
||||
printf '%s' "$output" | grep -qF "$jar" \
|
||||
|| fail "die message does not name this script's jar path: $output"
|
||||
printf '%s' "$output" | grep -qF "$other_jar" \
|
||||
|| fail "die message does not name the plist's jar path: $output"
|
||||
}
|
||||
|
||||
# fleetd #555 item 4 — the report-state dispatch on $SUPERVISOR_KIND. Inverting this used to report
|
||||
# the wrong supervisor and, for the launchd arm specifically, skip check_log_path_matches_plist.
|
||||
CHECK_LOG_PATH_CALLED=0
|
||||
CHECK_JAR_PATH_CALLED=0
|
||||
stub_check_log_path_recorder() {
|
||||
CHECK_LOG_PATH_CALLED=0
|
||||
CHECK_JAR_PATH_CALLED=0
|
||||
check_log_path_matches_plist() { CHECK_LOG_PATH_CALLED=1; }
|
||||
check_jar_path_matches_plist() { CHECK_JAR_PATH_CALLED=1; }
|
||||
}
|
||||
|
||||
test_report_supervisor_state_launchd_sets_supervised_and_checks_log_path() {
|
||||
@@ -1285,6 +1425,8 @@ test_report_supervisor_state_launchd_sets_supervised_and_checks_log_path() {
|
||||
[ "$SUPERVISED" = 1 ] || fail "report_supervisor_state launchd must set SUPERVISED=1"
|
||||
[ "$CHECK_LOG_PATH_CALLED" = 1 ] \
|
||||
|| fail "report_supervisor_state launchd must call check_log_path_matches_plist"
|
||||
[ "$CHECK_JAR_PATH_CALLED" = 1 ] \
|
||||
|| fail "report_supervisor_state launchd must call check_jar_path_matches_plist"
|
||||
source "$ROOT/scripts/redeploy-fleetd.sh"
|
||||
}
|
||||
|
||||
@@ -1296,6 +1438,8 @@ test_report_supervisor_state_systemd_sets_supervised_without_log_path_check() {
|
||||
[ "$SUPERVISED" = 1 ] || fail "report_supervisor_state systemd must set SUPERVISED=1"
|
||||
[ "$CHECK_LOG_PATH_CALLED" = 0 ] \
|
||||
|| fail "report_supervisor_state systemd must NOT call check_log_path_matches_plist"
|
||||
[ "$CHECK_JAR_PATH_CALLED" = 0 ] \
|
||||
|| fail "report_supervisor_state systemd must NOT call check_jar_path_matches_plist"
|
||||
source "$ROOT/scripts/redeploy-fleetd.sh"
|
||||
}
|
||||
|
||||
@@ -2153,6 +2297,8 @@ test_assert_single_daemon_accepts_one_pid
|
||||
test_assert_single_daemon_rejects_two_pids
|
||||
test_running_pid_excludes_self_matching_wrapper_shell
|
||||
test_running_pid_finds_a_real_java_named_second_process
|
||||
test_running_pid_finds_a_real_java_named_process_from_target_dir
|
||||
test_running_pid_excludes_self_matching_wrapper_shell_naming_target_dir
|
||||
test_running_pid_drops_a_pid_whose_comm_is_not_java
|
||||
test_running_pid_drops_a_pid_that_exited_before_the_comm_lookup
|
||||
test_running_pid_counts_a_pid_whose_comm_is_java
|
||||
@@ -2161,8 +2307,10 @@ test_jar_id_defaults_to_live_and_reports_explicit_path
|
||||
test_hash256_computes_a_real_sha256
|
||||
test_jar_id_reports_absent_for_missing_file
|
||||
test_jar_id_reports_unhashable_when_no_hasher_on_path
|
||||
test_stage_built_jar_moves_off_live_path
|
||||
test_stage_built_jar_dies_when_build_produced_nothing
|
||||
test_report_jar_state_agrees_when_hashes_match
|
||||
test_report_jar_state_warns_when_hashes_differ
|
||||
test_report_jar_state_both_absent_is_not_a_mismatch
|
||||
test_jar_and_build_jar_are_sourced_outside_target_and_differ
|
||||
test_swap_staged_jar_moves_staged_onto_live
|
||||
test_swap_staged_jar_dies_without_staged_file
|
||||
test_swap_staged_jar_dies_when_mv_fails
|
||||
@@ -2204,6 +2352,8 @@ test_drain_confirmed_false_on_anything_else
|
||||
test_run_drain_gate_skips_prompt_when_not_required
|
||||
test_run_drain_gate_confirmed_reply_does_not_refuse
|
||||
test_run_drain_gate_declined_reply_refuses
|
||||
test_check_jar_path_matches_plist_agrees_ok
|
||||
test_check_jar_path_matches_plist_mismatch_dies
|
||||
test_report_supervisor_state_launchd_sets_supervised_and_checks_log_path
|
||||
test_report_supervisor_state_systemd_sets_supervised_without_log_path_check
|
||||
test_report_supervisor_state_none_leaves_supervised_zero
|
||||
|
||||
Reference in New Issue
Block a user