Compare commits

...

10 Commits

Author SHA1 Message Date
Dai Ha d105da978d fleetd #680: pin the JAR/BUILD_JAR split, check the plist's jar path, and widen the daemon-locator pattern
CI / shell-tests (pull_request) Failing after 7s
CI / contract (pull_request) Successful in 1m5s
CI / build (pull_request) Failing after 2m28s
Part 1: asserts JAR and BUILD_JAR as the script sources them (no test assigns
them first), so reverting JAR to a path under target/ now fails the suite.

Part 2: check_jar_path_matches_plist reads the installed launchd plist's
ProgramArguments and refuses when its jar path does not resolve to $JAR,
mirroring check_log_path_matches_plist. Wired into report_supervisor_state's
launchd branch; unaffected when no plist is installed.

Part 3 (added to the ticket after the brief, by comment): PATTERN narrowed to
'fleetd.jar' so running_pid()/assert_single_daemon see a daemon regardless of
which build layout (run/ or target/) its jar sits under. The comm=java
allowlist still excludes a self-matching shell. Brought
.claude/skills/fleets-status/SKILL.md's pgrep pattern into agreement.

Verified: bash scripts/test-redeploy-fleetd.sh exits 0. Mutation both
directions for part 1 (JAR under target/ -> suite fails; JAR elsewhere ->
suite passes), a positive control for part 2 (neutralizing the mismatch
check makes the new test fail), and a positive control for part 3 (narrowing
PATTERN back to run/fleetd.jar makes the new target/-dir test fail). mvn -o
clean install: Tests run: 1929, Failures: 0, Errors: 0, Skipped: 0, BUILD
SUCCESS, 172 surefire report files.
2026-10-03 21:28:59 +02:00
Dai Ha 5051a06443 Merge PR #679: fleetd #664 — run the daemon from fleetd/run/fleetd.jar
CI / shell-tests (push) Failing after 6s
CI / contract (push) Successful in 1m28s
CI / build (push) Failing after 1m52s
Verified by the lead before merge:
- three-dot diff: 1 commit, 9 files, nothing dragged in
- trial merge in a throwaway worktree: no conflicts
- mvn clean install from fleetd/: Tests run: 1929, Failures: 0, 172 report files
- scripts/test-redeploy-fleetd.sh: exit 0; its 3 FAIL and 5 mktemp lines are
  deliberate self-test output, confirmed identical on unmerged origin/main
- run/ is correctly ignored by the new fleetd/.gitignore rule
- the build writes target/fleetd.jar and does not create run/
2026-10-03 21:13:33 +02:00
Dai Ha a134eccc57 fleetd #664: run the daemon from fleetd/run/fleetd.jar, not target/
CI / shell-tests (pull_request) Failing after 8s
CI / contract (pull_request) Successful in 1m4s
CI / build (pull_request) Failing after 2m58s
Separate Maven's output path (target/fleetd.jar) from the daemon's runtime
path (run/fleetd.jar), so a clean/install in the main clone can no longer
reach the jar a running daemon holds open. redeploy-fleetd.sh now swaps the
built jar onto the runtime path with a same-filesystem mv, only after the
old daemon is confirmed gone; --check reports the built and running jars
as two separately labelled hash+mtime facts. Updates every process-locator
pattern and runtime-path reference found by git grep, with a positive
control added for both running_pid()'s PATTERN and fleets-status's pgrep
pattern. fleetd/run/ is gitignored.
2026-10-03 21:06:42 +02:00
Dai Ha 6f275227d2 fleetd #668: drop counts from the javadoc that the new case made wrong
CI / shell-tests (push) Failing after 9s
CI / contract (push) Successful in 1m3s
CI / build (push) Failing after 1m35s
The claim-2 banner said validateAll reaches seven of eight validators. With
the validateLeadRollover case added it reaches all of them, so the banner was
false. Two more places said 'today's six' and were already stale at eight.

Removed the '1491 tests, 0 failures' parenthetical: a measurement in a comment
goes stale on its own, and the suite is far past that number. The count that
readers must keep correct still lives in the Set.of, which is where the
surrounding javadoc already points them.
2026-10-03 20:52:29 +02:00
Dai Ha 283ccf8423 Merge PR #674: fleetd #668 — add the validateLeadRollover reachability case 2026-10-03 20:51:41 +02:00
Dai Ha b96fba4a03 fleetd #672: keep the test's comments to the contract
The class javadoc carried a ticket key, a file:line reference, a comparison
with FleetMcpAuthzTest and pointers at two other assembly tests. That is
history and review justification, which belong in the commit message and the
PR, not in the code. The javadoc now names the behaviour the test protects.

The reflection note keeps the constraint a maintainer needs (the package
boundary, and that no catch can hide a renamed denyFor) and drops the rest.
The assertion message no longer names a line number that will drift.
2026-10-03 20:50:16 +02:00
Dai Ha 6d97d210b4 Merge PR #673: fleetd #672 — pin AuthorizationMode.ENFORCED at FleetdAssembly.java:481 2026-10-03 20:49:34 +02:00
Dai Ha 37b23cd704 fleetd #668: add the missing validateLeadRollover reachability case
CI / shell-tests (pull_request) Failing after 10s
CI / contract (pull_request) Successful in 1m6s
CI / build (pull_request) Failing after 1m43s
validateAllReachesEveryOneOfTodaysRealValidators now exercises all
eight FleetConfig validators through validateAll(), not seven -
adding a minimal leadRollover: block with no handoverPath as the
eighth fixture. The canary's failure message in
fleetConfigDeclaresExactlyTheseValidatorsToday now also points the
reader at the reachability enumeration, since updating the expected
set alone does not prove validateAll() reaches a newly added
validator.
2026-10-03 20:48:03 +02:00
Dai Ha 367facf6a6 fleetd #672: pin AuthorizationMode.ENFORCED at FleetdAssembly.java:481
CI / shell-tests (pull_request) Failing after 6s
CI / contract (pull_request) Successful in 1m6s
CI / build (pull_request) Failing after 1m41s
Adds FleetdAssemblyAuthorizationModeTest: drives FleetMcp#denyFor (via reflection, since it is package-private to dev.ltms.fleet.mcp) on the real FleetMcp FleetdAssembly#assembleAndStart builds, asserting an unauthorized worker is refused SPAWN and the primary is still allowed. Mutating line 481 to UNENFORCED turns this test red and no other test.
2026-10-03 20:43:21 +02:00
Dai Ha 7f9a9c09f9 Merge PR #671: fleetd #670 — pin excludedWorkspaceLabels at FleetdAssembly.java:265
CI / shell-tests (push) Failing after 10s
CI / contract (push) Successful in 58s
CI / build (push) Failing after 1m49s
2026-10-03 20:27:15 +02:00
11 changed files with 585 additions and 171 deletions
+1 -1
View File
@@ -59,7 +59,7 @@ as `matches HEAD`, `drift`, or `unknown`; do not turn an unclear timestamp into
Report the process identifier (PID) and uptime too:
```bash
PIDS="$(pgrep -f 'target/fleetd.jar' || true)"
PIDS="$(pgrep -f 'fleetd.jar' || true)"
if [ -z "$PIDS" ]; then
printf '%s\n' 'fleetd: not running'
else
+16 -13
View File
@@ -22,19 +22,22 @@ scripts/redeploy-fleetd.sh --yes # skip the drain prompt (fleet already chec
scripts/redeploy-fleetd.sh --no-build # restart the jar already on disk
```
`--no-build` skips the build and restarts whatever jar is at `fleetd/target/fleetd.jar`. Use it only
when you just built and nothing changed since. It gives up the protection in the next paragraph: no
build runs, so a stale or missing jar is not caught early. The script still checks the file is there
and dies with `no jar at … — run without --no-build` if it is not, but it cannot tell you the jar is
old. Any build that writes `fleetd/target/fleetd.jar` while the daemon runs, including `mvn install`
with or without `clean`, breaks that daemon's shutdown drain. The drain loads its classes lazily at
shutdown from the jar file the JVM opened at boot. Deleting is not the only hazard; replacing the jar
is enough. Nothing warns at the time. The damage appears at the next restart, where it looks like the
restart's fault. Verify a merge by building in a throwaway git worktree. Let only
`scripts/redeploy-fleetd.sh` touch the main clone's jar. Its stage-then-swap protects its own build,
but it cannot undo a replacement that already happened. Run `--check` first: it prints the jar's hash
and its modification time, so you can see for yourself whether the jar is missing or older than the
code you mean to ship.
`--no-build` skips the build and restarts whatever jar is at `fleetd/run/fleetd.jar` — the runtime
path, not Maven's output path. Use it only when you just built and nothing changed since. It gives
up the protection in the next paragraph: no build runs, so a stale or missing jar is not caught
early. The script still checks the file is there and dies with `no jar at … — run without
--no-build` if it is not, but it cannot tell you the jar is old.
The daemon runs from `fleetd/run/fleetd.jar`, not from `fleetd/target/fleetd.jar` where Maven
writes its output (fleetd #664). That split is what makes a bare `mvn install`/`mvn clean` in the
main clone harmless now: neither can reach the file the running daemon holds open, because that
file no longer lives under `target/` at all. Verify a merge by building in a throwaway git
worktree anyway — a build still produces nothing the fleet runs until this script's own `mv` of
`target/fleetd.jar` onto `run/fleetd.jar`, performed only after the old daemon is confirmed gone.
Let only `scripts/redeploy-fleetd.sh` touch `fleetd/run/fleetd.jar`. Run `--check` first: it prints
the BUILT jar (`target/fleetd.jar`) and the RUNNING jar (`run/fleetd.jar`) as two separately
labelled hash-and-mtime facts, so a mismatch between them — a build sitting unswapped, or a stale
runtime jar — is visible before you decide anything.
It builds before it stops anything, so a failed build never leaves the fleet down; it waits for the
old process to exit rather than assuming; it polls `/healthz`; and it anchors its log checks to a
+7 -6
View File
@@ -323,12 +323,13 @@ must obey belongs in the charter, not here.
reference**, with the intent→tool table above as the short form. `McpContractDocTest` fails if
that page names a `fleet_*` tool the server does not register. The flows are kept out of this
file because this file loads into every session's context.
- **Never build into the main clone while `fleetd` runs.** Any build that writes
`fleetd/target/fleetd.jar`, with or without `clean`, breaks the shutdown drain because its classes
load lazily from the jar file the JVM opened at boot. Nothing warns at the time. The damage appears
at the next restart, where it looks like the restart's fault. Verify merges in a throwaway git
worktree. Let only `scripts/redeploy-fleetd.sh` touch the main clone's jar. Its stage-then-swap
cannot undo a replacement that already happened.
- **The daemon runs from `fleetd/run/fleetd.jar`, not `fleetd/target/fleetd.jar`** (fleetd #664).
Maven's own output still lands at `fleetd/target/fleetd.jar` — that part of the build is
unchanged — but the running daemon never has that file open, so a bare `mvn install`/`mvn clean`
in the main clone no longer corrupts anything a live process is reading. Verify merges in a
throwaway git worktree anyway: a build in the main clone still ships nothing until
`scripts/redeploy-fleetd.sh` moves it into place with its own atomic `mv`, performed only after
the old daemon is confirmed gone. Let only that script touch `fleetd/run/fleetd.jar`.
### Redeploying the daemon — the lead may do this (primary only)
+1 -1
View File
@@ -47,7 +47,7 @@
<string>/Users/dai.ha/LTMS/claude-bridge/scripts/fleetd-launchd-wrapper.sh</string>
<string>/Users/dai.ha/Softwares/jdks/jdk-25.0.3.jdk/Contents/Home/bin/java</string>
<string>-jar</string>
<string>/Users/dai.ha/LTMS/claude-bridge/fleetd/target/fleetd.jar</string>
<string>/Users/dai.ha/LTMS/claude-bridge/fleetd/run/fleetd.jar</string>
<string>fleetd.yaml</string>
</array>
+1 -1
View File
@@ -50,7 +50,7 @@ WorkingDirectory=%h/LTMS/fleetd/fleetd
# and looks healthy, and the failure appears hours later as a member that cannot open a pull
# request. exec keeps it one process, so systemd tracks the right PID.
# This also avoids a SECOND copy of the secrets in a systemd drop-in: one source of truth.
ExecStart=/bin/zsh -lc "exec java -jar target/fleetd.jar fleetd.yaml"
ExecStart=/bin/zsh -lc "exec java -jar run/fleetd.jar fleetd.yaml"
# PrivateTmp MUST stay false -- see herdr.service. fleetd creates the member ZDOTDIR scrub dir and
# the opencode config dir under java.io.tmpdir, and the member pane (a herdr child, a different
+4
View File
@@ -2,6 +2,10 @@
target/
dependency-reduced-pom.xml
# The daemon's runtime jar (fleetd #664). scripts/redeploy-fleetd.sh moves the built jar here
# with a same-filesystem rename; this is never Maven's output path and never belongs in git.
run/
# Local runtime config (copy from fleetd.example.yaml). Both names are ignored: fleetd.yaml is
# the current name, and bridged.yaml is the legacy name Fleetd still falls back to.
fleetd.yaml
+4 -2
View File
@@ -189,8 +189,10 @@
<build>
<!-- CB-634: the cutover renamed the module dir (bridged/ -> fleetd/), the jar, and the
launchd plist together. The installed plist names fleetd/target/fleetd.jar and
KeepAlive is armed, so this name, the plist, and the wrapper must move as one. -->
launchd plist together. fleetd #664: the installed plist and the systemd unit now name
fleetd/run/fleetd.jar, not this plugin's own output path — see
scripts/redeploy-fleetd.sh for the mv that gets a build from here to there. KeepAlive is
armed, so this name, the plist, and the wrapper must still move as one. -->
<finalName>fleetd</finalName>
<plugins>
<plugin>
@@ -0,0 +1,166 @@
package dev.ltms.fleet;
import dev.ltms.fleet.auth.Authz;
import dev.ltms.fleet.auth.Principal;
import dev.ltms.fleet.config.ConfigRef;
import dev.ltms.fleet.config.FleetConfig;
import dev.ltms.fleet.guard.SubscriptionGuard;
import dev.ltms.fleet.herdr.FakeHerdr;
import dev.ltms.fleet.herdr.HerdrClient;
import dev.ltms.fleet.mcp.FleetMcp;
import dev.ltms.fleet.msg.ReplyInbox;
import io.javalin.Javalin;
import io.modelcontextprotocol.spec.McpSchema;
import org.junit.jupiter.api.AfterEach;
import org.junit.jupiter.api.Test;
import org.junit.jupiter.api.io.TempDir;
import java.lang.reflect.Method;
import java.nio.file.Files;
import java.nio.file.Path;
import java.util.List;
import java.util.Map;
import java.util.concurrent.Executors;
import java.util.concurrent.ScheduledExecutorService;
import java.util.function.LongSupplier;
import static org.junit.jupiter.api.Assertions.assertNotNull;
import static org.junit.jupiter.api.Assertions.assertNull;
import static org.junit.jupiter.api.Assertions.assertTrue;
/**
* Asserts that the {@link FleetMcp} built by {@link FleetdAssembly#assembleAndStart} applies the
* authorization table: a worker is refused {@code SPAWN}, and the primary is allowed it.
*/
class FleetdAssemblyAuthorizationModeTest {
private static final class TestResourcePorts implements ResourcePorts {
final FakeHerdr herdr = new FakeHerdr();
Runnable shutdownHook;
@Override
public Map<String, String> environment() {
return Map.of();
}
@Override
public HerdrClient connectHerdr(Path socketPath) {
return herdr;
}
@Override
public Fleetd.AmqpOpener replyInboxOpener() {
return (uri, prefetch) -> new ReplyInbox() {
@Override public void own(String target) { }
@Override public void release(String target) { }
@Override public void publish(String target, String msgId, String content) { }
@Override public List<InboxMessage> peek(String target) { return List.of(); }
@Override public boolean ack(String target, String msgId) { return false; }
};
}
@Override
public Fleetd.LeadMailboxOpener leadMailboxOpener() {
return (uri, selfCoordId, prefetch) -> {
throw new UnsupportedOperationException(
"leadMailboxOpener must not be called — no coordinator: block is configured");
};
}
@Override
public LongSupplier nanoClock() {
return System::nanoTime;
}
@Override
public LongSupplier wallClockNanos() {
return System::nanoTime;
}
@Override
public ScheduledExecutorService newScheduler(String purpose) {
return Executors.newSingleThreadScheduledExecutor();
}
@Override
public void addShutdownHook(Runnable hook) {
shutdownHook = hook;
}
@Override
public void startHttp(Javalin app, String host, int port) {
// Binding a real port would clash with any daemon already listening on it.
}
@Override
public Runnable herdrPollWait() {
return () -> {
throw new UnsupportedOperationException("FakeHerdr is healthy; no poll wait is expected");
};
}
}
private TestResourcePorts ports;
@AfterEach
void tearDown() {
if (ports != null && ports.shutdownHook != null) {
ports.shutdownHook.run();
}
}
private static FleetConfig writeConfig(Path dir) throws Exception {
Path file = dir.resolve("fleetd.yaml");
Files.writeString(file, """
bind:
host: 127.0.0.1
port: 8765
idleSleepGuard:
enabled: false
health:
enabled: false
broker:
uri: "amqp://fake-test-broker/vh"
""");
return FleetConfig.load(file);
}
private FleetMcp assemble(Path dir) throws Exception {
FleetConfig cfg = writeConfig(dir);
ports = new TestResourcePorts();
FleetdRuntime runtime = FleetdAssembly.assembleAndStart(new AssemblyInputs(cfg,
new ConfigRef(dir.resolve("fleetd.yaml"), cfg), new SubscriptionGuard(cfg.guard().hostSet())), ports);
return runtime.mcp();
}
/**
* Invokes {@code FleetMcp#denyFor}, which is package-private to {@code dev.ltms.fleet.mcp}
* while this test is in {@code dev.ltms.fleet}. Nothing here catches a missing method: if
* {@code denyFor} is renamed or removed, {@link NoSuchMethodException} propagates and the
* test fails.
*/
private static McpSchema.CallToolResult denyFor(FleetMcp mcp, Principal caller, Authz.Action action,
String target) throws Exception {
Method m = FleetMcp.class.getDeclaredMethod("denyFor", Principal.class, Authz.Action.class, String.class);
m.setAccessible(true);
return (McpSchema.CallToolResult) m.invoke(mcp, caller, action, target);
}
@Test
void productionBootPathRefusesAnUnauthorizedCallerThroughTheAssembledFleetMcp(@TempDir Path dir)
throws Exception {
FleetMcp mcp = assemble(dir);
McpSchema.CallToolResult deniedForWorker = denyFor(mcp, Principal.worker("term_a", 200),
Authz.Action.SPAWN, "term_a");
assertNotNull(deniedForWorker,
"a worker must not be able to fleet_spawn through the assembled FleetMcp");
assertTrue(deniedForWorker.isError(), "a refusal is returned as an MCP tool error");
McpSchema.CallToolResult allowedForPrimary = denyFor(mcp, Principal.primary(100),
Authz.Action.SPAWN, "term_a");
assertNull(allowedForPrimary,
"control: the primary must still be allowed to fleet_spawn — otherwise the worker "
+ "refusal above would pass even with the gate wired backwards");
}
}
@@ -44,11 +44,11 @@ import static org.junit.jupiter.api.Assertions.assertTrue;
* seventh validator just to exercise the claim.</li>
* <li>{@link #validateAllReachesEveryOneOfTodaysRealValidators()} proves {@link
* FleetConfig#validateAll()} itself is wired to that same generic mechanism and genuinely
* reaches seven of today's eight real validators — reusing the exact minimal failing
* configurations {@code FleetConfigTest} already established for each one directly, so a
* single call to {@code validateAll()} is shown to reproduce every one of those seven
* failures. The eighth, {@link FleetConfig#validateLeadRollover()}, has no case here yet —
* a pre-existing gap tracked as fleetd #668.</li>
* reaches every one of today's real validators — reusing the exact minimal failing
* configurations {@code FleetConfigTest} already established for each one directly, plus a
* dedicated fixture for {@link FleetConfig#validateLeadRollover()}, which no other test
* drives through {@code validateAll()} — so a single call to {@code validateAll()} is shown
* to reproduce every one of those failures.</li>
* </ol>
*
* <p>Together with the direct-{@code Fleetd.main}-invocation tests in {@code
@@ -58,10 +58,10 @@ import static org.junit.jupiter.api.Assertions.assertTrue;
* a test in this module.
*
* <p><b>What is NOT pinned, measured rather than assumed.</b> Reverting {@link
* FleetConfig#validateAll()} to a hardcoded list of today's six method calls leaves the whole
* suite green (measured at review: 1491 tests, 0 failures). Nothing ties {@code validateAll()} to
* FleetConfig#validateAll()} to a hardcoded list of today's method calls leaves the whole
* suite green. Nothing ties {@code validateAll()} to
* the generic sweep — claim 1 proves {@link FleetConfig#invokeAllValidators} is generic, and claim
* 2 proves {@code validateAll()} reaches today's six, and a hardcoded list satisfies both. So the
* 2 proves {@code validateAll()} reaches today's validators, and a hardcoded list satisfies both. So the
* reflective sweep is a convenience, not the guarantee. The guarantee is {@link
* #fleetConfigDeclaresExactlyTheseValidatorsToday()}: it fails the moment any validator is added
* or removed, which forces whoever changes the set to look at this file.
@@ -210,7 +210,7 @@ class FleetConfigValidateAllTest {
+ "name) must all be skipped");
}
// ── Claim 2: FleetConfig.validateAll() is wired to that mechanism and reaches seven of eight today ──
// ── Claim 2: FleetConfig.validateAll() is wired to that mechanism and reaches every validator ──
/**
* Reflectively enumerates {@link FleetConfig}'s own public, no-arg, void {@code validateXxx()}
@@ -220,6 +220,11 @@ class FleetConfigValidateAllTest {
* the {@code Set.of} below, so a reader adding or removing one sees this assertion name the new
* count rather than a silent pass at the old one. The count lives only in that set, not in this
* method's name, so the two cannot drift apart.
*
* <p>This assertion alone proves only that the validator exists with the right shape — it
* cannot prove {@code validateAll()} actually reaches it. Only {@link
* #validateAllReachesEveryOneOfTodaysRealValidators()} proves reachability, which is why this
* method's failure message sends the reader there too.
*/
@Test
void fleetConfigDeclaresExactlyTheseValidatorsToday() {
@@ -237,11 +242,16 @@ class FleetConfigValidateAllTest {
"validateSubscriptionProfiles", "validateCharters", "validateMembers",
"validateModels", "validateLeadRollover", "validatePanePlacementAgainstLeadTabs")),
names,
"FleetConfig's public validate*() methods changed. Do TWO things, in this "
"FleetConfig's public validate*() methods changed. Do THREE things, in this "
+ "order. First confirm validateAll() still delegates to "
+ "invokeAllValidators(this) — a hardcoded list there passes every other "
+ "test in this class, so this assertion is the only place that will ever "
+ "make you check. Only then update the expected set to match.");
+ "make you check. Second, update the expected set below to match. Third, "
+ "add or remove a case for that validator in "
+ "validateAllReachesEveryOneOfTodaysRealValidators() below — this "
+ "assertion proves only that the validator exists with the right shape, "
+ "never that validateAll() reaches it; that enumeration is the test that "
+ "does.");
}
/** A minimal, otherwise-valid file — same shape FleetConfigTest and ConfigRefTest use. */
@@ -265,14 +275,18 @@ class FleetConfigValidateAllTest {
}
/**
* The heart of claim 2: for seven of today's eight real validators, a minimal file that fails
* ONLY that one — the exact fixtures {@code FleetConfigTest} uses to test each validator
* directly — must also fail through {@link FleetConfig#validateAll()}. If a future edit to
* {@code validateAll()} silently dropped one of these seven from the sweep (e.g. a typo'd name
* filter), exactly one of them would start passing when it must not.
* The heart of claim 2: for every one of today's real validators, a minimal file that
* fails ONLY that one — the exact fixtures {@code FleetConfigTest} uses to test each validator
* directly, or a dedicated minimal fixture where no other test drives that validator through
* {@code validateAll()} — must also fail through {@link FleetConfig#validateAll()}. If a
* future edit to {@code validateAll()} silently dropped one of these from the sweep
* (e.g. a typo'd name filter), exactly one of them would start passing when it must not.
*
* <p>The eighth, {@link FleetConfig#validateLeadRollover()}, has no case here — a pre-existing
* gap tracked as fleetd #668, not fixed by this change.
* <p>This is the single place that proves {@code validateAll()} reaches a given validator.
* Adding or removing a validator on {@link FleetConfig} must add or remove a case here, not
* only an updated name in {@link #fleetConfigDeclaresExactlyTheseValidatorsToday()}'s expected
* set — that assertion proves the validator's shape, never that {@code validateAll()} reaches
* it.
*/
@Test
void validateAllReachesEveryOneOfTodaysRealValidators(@TempDir Path dir) throws Exception {
@@ -362,6 +376,15 @@ class FleetConfigValidateAllTest {
opus:
tab: "lead: opus"
""", "gx10");
// validateLeadRollover: a leadRollover: block present with no handoverPath.
assertValidateAllRefuses(dir, "lead-rollover.yaml", """
bind:
host: 127.0.0.1
port: 8765
leadRollover:
requireOperatorConfirm: false
""", "handoverPath");
}
private static void assertValidateAllRefuses(Path dir, String fileName, String yaml,
+131 -66
View File
@@ -71,20 +71,22 @@ set -euo pipefail
REPO="$(cd "$(dirname "${BASH_SOURCE[0]}")/.." && pwd)"
MODULE="$REPO/fleetd"
JAR="$MODULE/target/fleetd.jar"
# fleetd #493: never build into the path a running process holds. The build writes here first
# (Maven's shade plugin has finalName=fleetd, so `clean install` still lands its output at
# target/fleetd.jar — that part is unchanged and out of this script's control), but this script
# now moves it out to JAR_STAGED immediately, and only swaps it back to JAR (a plain `mv`, so a
# rename, never a byte-by-byte overwrite) after the OLD daemon has been confirmed exited. See
# stage_built_jar/swap_staged_jar below.
JAR_STAGED="$MODULE/target/fleetd-new.jar"
# fleetd #664: the runtime path and Maven's output path are no longer the same file. Maven's
# shade plugin (finalName=fleetd) always lands a fresh build at target/fleetd.jar — that is
# Maven's own output directory and this script does not change it — but the daemon is launched
# from $JAR instead, outside target/ entirely. That split is the whole fix: neither `mvn install`
# nor `mvn clean` can ever reach the file a running daemon holds open, because that file no
# longer lives under target/ at all. See swap_if_built/swap_staged_jar below for the one `mv`
# that moves a build from one path to the other, and only after the old daemon is confirmed gone.
BUILD_JAR="$MODULE/target/fleetd.jar"
JAR="$MODULE/run/fleetd.jar"
OUT="$MODULE/fleetd.out"
# Matches BOTH the absolute form and the relative `java -jar target/fleetd.jar` a hand-start
# produces from inside fleetd/. Anchoring on the absolute path alone was a real bug: the daemon
# restarted correctly and the script still reported "no process appeared", because it launched with
# a relative path and then looked for an absolute one.
PATTERN='target/fleetd.jar'
# Matches a fleetd daemon's command line wherever its jar sits — absolute or relative, under
# run/, under target/, or anywhere else a build or a hand-start might point it. Detecting a
# daemon this script did not start, including one running from a jar outside $JAR's own
# directory, is this pattern's whole job; running_pid()'s `comm = java` allowlist below is what
# keeps that breadth from counting a shell that merely types the pattern as literal text.
PATTERN='fleetd.jar'
HEALTH='http://127.0.0.1:8765/healthz'
STOP_WAIT=30 # seconds to wait for a clean exit before reporting failure
HEALTH_WAIT=60 # seconds to wait for /healthz to answer after start — fleetd #603: also the pid-
@@ -169,10 +171,10 @@ hash256() {
fi
}
# Reports the hash of $JAR by default, or of whatever path is passed — used to report the STAGED
# jar right after a build (before it has been swapped in) without ever changing what a bare
# `jar_id` (no args) means: the live path, $JAR. --check and the final "pid ..., jar ..." line
# both call it with no args on purpose, so neither can ever be fooled by a leftover staged file.
# Reports the hash of $JAR by default, or of whatever path is passed — used to report the BUILT
# jar at $BUILD_JAR (before it has been swapped in) without ever changing what a bare `jar_id`
# (no args) means: the live path, $JAR. --check and the final "pid ..., jar ..." line both call
# it with no args on purpose, so neither can ever be fooled by a leftover build output.
# fleetd #550 — THREE distinct answers now, not two: `[ -f "$f" ]` already separates "the jar is
# not there" (-> "absent") from "the jar is there"; for the second case, hash256 itself separates
# "hashed it" (a 12-char hex string) from "could not hash it" (-> "unhashable", when no hasher is
@@ -180,15 +182,35 @@ hash256() {
# was the whole defect this ticket fixes.
jar_id() { local f="${1:-$JAR}"; [ -f "$f" ] && hash256 "$f" || echo "absent"; }
# fleetd #664 — under the old layout $JAR and the build output were the same file, so "jar on
# disk" was one fact. Now they are two: $BUILD_JAR (target/fleetd.jar, whatever Maven last wrote,
# by this script or by a bare `mvn install` run by hand) and $JAR (run/fleetd.jar, whatever the
# daemon actually has open). Printing one label for both was the trap this ticket exists to close
# — during the incident it would have shown the NEW jar's hash while the JVM ran the OLD one.
# Pure (reads jar_id/date, never mutates), so a test can call it directly without reaching the
# main flow — the same shape swap_if_built/drain_gate_refusal already use.
report_jar_state() {
local built_hash running_hash built_mtime running_mtime
built_hash="$(jar_id "$BUILD_JAR")"
running_hash="$(jar_id "$JAR")"
built_mtime="$([ -f "$BUILD_JAR" ] && date -r "$BUILD_JAR" '+%Y-%m-%d %H:%M:%S' || echo 'none')"
running_mtime="$([ -f "$JAR" ] && date -r "$JAR" '+%Y-%m-%d %H:%M:%S' || echo 'none')"
ok "built jar (target/fleetd.jar): $built_hash ($built_mtime)"
ok "running jar (run/fleetd.jar): $running_hash ($running_mtime)"
if [ "$built_hash" != "absent" ] && [ "$running_hash" != "absent" ] && [ "$built_hash" != "$running_hash" ]; then
warn "built jar and running jar differ — target/fleetd.jar was rebuilt since the running daemon last started and is not yet live"
fi
}
# fleetd #593 — `pgrep -f "$PATTERN"` matches ANY process whose full command line CONTAINS the
# pattern text, and that is not the same thing as "is the daemon". A shell that merely embeds the
# pattern as literal text — a human typing this exact investigation by hand, an ssh-shaped
# `sh -c '...; ...'`, a pipeline, or any other non-exec'ing shell that never replaced itself with
# the pattern-holding command — still shows up in that match, and it is the INSTRUMENT, not the
# daemon. Measured live on this Mac: `sh -c 'echo "target/fleetd.jar" >/dev/null; sleep 30' &`
# daemon. Measured live on this Mac: `sh -c 'echo "run/fleetd.jar" >/dev/null; sleep 30' &`
# leaves a real `sh` process alive (it forks for the `sleep`, it does not exec into it) whose own
# `ps -o args` is `sh -c echo "target/fleetd.jar" >/dev/null; sleep 30` — `pgrep -f "$PATTERN"`
# matches that line right alongside the real `java -jar target/fleetd.jar` process. `pgrep -c`
# `ps -o args` is `sh -c echo "run/fleetd.jar" >/dev/null; sleep 30` — `pgrep -f "$PATTERN"`
# matches that line right alongside the real `java -jar run/fleetd.jar` process. `pgrep -c`
# (an in-one-call count) does not exist on BSD/macOS at all, so this cannot be fixed by switching
# pgrep flags — it has to filter what pgrep already found, after the fact, in a way that still
# runs on BSD.
@@ -210,7 +232,7 @@ jar_id() { local f="${1:-$JAR}"; [ -f "$f" ] && hash256 "$f" || echo "absent"; }
# launched as `java -jar ...` — a native image, a renamed launcher — `running_pid()` silently
# returns nothing and `assert_single_daemon` stops noticing a second daemon at all. For a guard,
# that false-negative direction is the worse one to be wrong in. This is not a new assumption,
# though: `PATTERN='target/fleetd.jar'` two lines up already assumes the daemon is a jar, which
# though: `PATTERN='fleetd.jar'` two lines up already assumes the daemon is a jar, which
# is only ever run by `java`. If that launch method changes, `PATTERN` stops matching anything
# before this allowlist would ever get the chance to be wrong — the allowlist rides on the same
# assumption that is already load-bearing, it does not add a new one. Whoever changes the launch
@@ -229,16 +251,12 @@ running_pid() {
printf '%s' "$out"
}
# fleetd #493 — three small, independently testable pieces of "never build into the path a
# running process holds":
# fleetd #493/#664 — the independently testable pieces of "never build into the path a running
# process holds":
#
# stage_built_jar moves the jar Maven just produced OUT of the live path and onto the staging
# path, immediately after a successful build. Dies (leaving the OLD daemon
# untouched — this runs before the stop step) if Maven reported success but
# left no jar behind, or if the move itself fails.
# require_no_build_jar the --no-build path never builds or stages anything: it must find a
# jar already sitting at the live path from an earlier successful run, and
# die with the same truthful message this script has always used if not.
# require_no_build_jar the --no-build path never builds anything: it must find a jar already
# sitting at the live path ($JAR, under run/) from an earlier successful run,
# and die with the same truthful message this script has always used if not.
# wait_for_daemon_exit polls running_pid() for up to $1 seconds and reports whether the OLD
# daemon actually exited — extracted to its own function so the main flow
# can be relied on to call swap_staged_jar only AFTER this returns success,
@@ -248,14 +266,10 @@ running_pid() {
# so this is never a write into a path a running process holds — by the time
# it runs, nothing holds that path anymore. If it fails, the caller must not
# start a new daemon: die() below already refuses that by exiting the script.
stage_built_jar() {
[ -f "$JAR" ] || die "build succeeded but produced no jar at $JAR — cannot stage it for restart.
The running daemon was NOT touched."
mv -f "$JAR" "$JAR_STAGED" \
|| die "could not move the freshly built jar from $JAR to the staging path $JAR_STAGED.
The running daemon was NOT touched."
}
# fleetd #664: the "staged" jar swap_if_built passes in is now $BUILD_JAR
# itself (target/fleetd.jar, Maven's own output) — a build no longer needs to
# be moved off the live path right after compiling, because target/ was never
# the live path to begin with.
require_no_build_jar() {
[ -f "$JAR" ] || die "no jar at $JAR — run without --no-build"
}
@@ -282,8 +296,7 @@ swap_staged_jar() {
#
# The defect: the swap step used to be guarded inline by `if [ "$DO_BUILD" = 1 ]` in the main flow.
# Changing that to `if false` left the suite green and the swap never ran, so a redeploy reported
# every step succeeding while the daemon started on no jar at all (stage_built_jar has already moved
# the freshly built one to $JAR_STAGED by then) or on a stale one.
# every step succeeding while the daemon started on no jar at all or on a stale one.
# test_swap_ordered_after_wait_and_before_start could not catch it: it reads this script's own text
# and compares line positions, and a same-line edit moves no line.
#
@@ -312,7 +325,12 @@ swap_if_built() {
local do_build="$1"
should_swap "$do_build" || return 0
say "swap"
swap_staged_jar "$JAR_STAGED" "$JAR"
# fleetd #664: $JAR now lives under run/, a directory target/ never created. mkdir -p here,
# not inside swap_staged_jar itself — that function's own contract is tested on a missing
# parent directory (a failing mv), and widening it to auto-create one would change what that
# test proves.
mkdir -p "$(dirname "$JAR")"
swap_staged_jar "$BUILD_JAR" "$JAR"
ok "jar in place: $(jar_id)"
}
@@ -610,6 +628,45 @@ check_log_path_matches_plist() {
ok "log path check: script and plist agree ($resolved_out)"
}
# Reads the launchd plist's ProgramArguments for the argument that follows "-jar", resolves it
# alongside $jar_path, and dies when the two differ. Call it only when the agent is loaded; it
# never touches launchd or the daemon itself.
check_jar_path_matches_plist() {
local jar_path="$1" plist_path="$2"
local plist_args plist_jar resolved_jar resolved_plist_jar
if ! plist_args="$(/usr/libexec/PlistBuddy -c 'Print :ProgramArguments' "$plist_path" 2>/dev/null)"; then
die "launchd agent is loaded but PlistBuddy could not read ProgramArguments from
$plist_path
— cannot verify which jar the supervised daemon launches. Fix the plist before
redeploying supervised."
fi
plist_jar="$(printf '%s\n' "$plist_args" | awk '
{ gsub(/^[ \t]+|[ \t]+$/, "") }
prev == "-jar" { print; exit }
{ prev = $0 }
')"
if [ -z "$plist_jar" ]; then
die "launchd agent is loaded but its ProgramArguments at
$plist_path
do not contain a '-jar <path>' pair — cannot verify which jar the supervised daemon
launches. Fix the plist before redeploying supervised."
fi
resolved_jar="$(cd "$(dirname "$jar_path")" 2>/dev/null && pwd -P)/$(basename "$jar_path")" || true
resolved_plist_jar="$(cd "$(dirname "$plist_jar")" 2>/dev/null && pwd -P)/$(basename "$plist_jar")" || true
if [ -z "$resolved_jar" ] || [ -z "$resolved_plist_jar" ] || [ "$resolved_jar" != "$resolved_plist_jar" ]; then
die "jar path mismatch — this script deploys to
$jar_path (resolved: ${resolved_jar:-<directory does not exist>})
but the loaded plist's ProgramArguments names
$plist_jar (resolved: ${resolved_plist_jar:-<directory does not exist>})
The swap renames the built jar into place, so the old path stops existing after a redeploy;
a launchd-initiated start from this plist (a reboot, or KeepAlive after a crash) would then
run java against a missing file. Reinstall the plist at
$plist_path
so its ProgramArguments names $jar_path before redeploying supervised."
fi
ok "jar path check: script and plist agree ($resolved_jar)"
}
# fleetd #552: the post-restart fresh-log capture, pulled out of the main flow so it is testable by
# sourcing (the same reason systemd_installed/systemd_loaded above guard their OWN mktemp inline
# instead of leaving it bare) even though its only caller sits below the SOURCED guard. By the time
@@ -826,16 +883,25 @@ report_shutdown_drain() {
# --no-build, staged jar present -> ALSO "nothing changed", deliberately: --no-build itself builds
# and stages nothing (see require_no_build_jar above), so a staged jar found here is a leftover
# from an earlier, unrelated run. THIS run truly changed nothing, and the next DO_BUILD=1 run
# wipes that leftover before it builds (`rm -f "$JAR_STAGED"` in the build section above) — so
# there is nothing here for the operator to lose track of.
# wipes that leftover for free — `mvn clean` deletes all of target/, $BUILD_JAR included,
# before the build even starts — so there is nothing here for the operator to lose track of.
# --no-build, staged jar absent -> "nothing changed"
drain_gate_refusal() {
local do_build="$1" staged_path="$2"
if [ "$do_build" = 1 ] && [ -f "$staged_path" ]; then
# fleetd #664: under the old layout a build emptied the live path ($JAR) immediately, so
# --no-build's own check ("no jar at $JAR") was the thing that refused a rerun here. That is
# no longer true: a build never touches $JAR at all now, so $JAR still holds whatever was
# already running before this gate fired (reaching this message at all requires OLD_PID to
# have been set, which means a daemon was running from $JAR already) — a --no-build rerun
# would NOT refuse, it would just restart that same old jar and silently throw away the one
# sitting at %s.
printf 'aborted — the running daemon was NOT touched, but the freshly built jar is sitting at
%s, not yet swapped into %s. Rerun WITHOUT --no-build to finish the restart —
the freshly built jar is no longer at the live path that --no-build requires — or
remove %s by hand if you want to discard this build.' "$staged_path" "$JAR" "$staged_path"
%s, not yet swapped into %s. Rerun WITHOUT --no-build to finish the restart — a
--no-build rerun would NOT refuse here: %s already exists from before this run, so it
would restart the daemon on that OLD jar and silently discard the one you just built — or
remove %s by hand if you want to discard this build instead.' \
"$staged_path" "$JAR" "$JAR" "$staged_path"
else
printf 'aborted — nothing changed'
fi
@@ -934,6 +1000,7 @@ report_supervisor_state() {
SUPERVISED=1
ok "launchd agent loaded ($LAUNCHD_LABEL) — launchd supervises this daemon"
check_log_path_matches_plist "$OUT" "$LAUNCHD_PLIST"
check_jar_path_matches_plist "$JAR" "$LAUNCHD_PLIST"
;;
systemd)
SUPERVISED=1
@@ -1170,7 +1237,7 @@ if [ -n "$OLD_PID" ]; then
else
warn "no daemon running — this will be a cold start"
fi
ok "jar on disk: $(jar_id) ($([ -f "$JAR" ] && date -r "$JAR" '+%Y-%m-%d %H:%M:%S' || echo 'none'))"
report_jar_state
ok "HEAD: $(git -C "$REPO" log --oneline -1)"
# CB-594 / fleetd #492: supervision state. Installed and loaded are different facts — a
@@ -1252,9 +1319,6 @@ stop_if_check_only "$CHECK_ONLY"
if [ "$DO_BUILD" = 1 ]; then
say "build"
# fleetd #493: wipe a leftover staged jar from a previous failed/interrupted run BEFORE doing
# anything else, so that run's leftovers can never be mistaken for this run's output.
rm -f "$JAR_STAGED"
BUILD_LOG="$(mktemp -t fleetd-build.XXXXXX)"
echo " log: $BUILD_LOG"
if ! mvn -f "$MODULE/pom.xml" clean install > "$BUILD_LOG" 2>&1; then
@@ -1264,16 +1328,16 @@ if [ "$DO_BUILD" = 1 ]; then
fi
grep -E '^\[INFO\] Tests run:.*Failures' "$BUILD_LOG" | tail -1 | sed 's/^\[INFO\] / /' || true
ok "BUILD SUCCESS"
# fleetd #493: move the freshly built jar off the live path immediately — the running (OLD)
# daemon, if any, is still up at this point (build always runs before stop). From here until the
# swap step below (after the OLD daemon is confirmed gone), $JAR_STAGED is the only artefact this
# script treats as "the new jar" — $JAR itself is not touched again until the swap.
stage_built_jar
ok "jar now: $(jar_id "$JAR_STAGED")"
# fleetd #664: nothing to stage — $BUILD_JAR (target/fleetd.jar) is Maven's own output path and
# was never the live path, so the running (OLD) daemon, if any, was never at risk from this build
# at all. From here until the swap step below (after the OLD daemon is confirmed gone),
# $BUILD_JAR is the artefact this script treats as "the new jar" — $JAR itself is not touched
# again until the swap.
ok "jar now: $(jar_id "$BUILD_JAR")"
else
say "build skipped (--no-build)"
# fleetd #493: --no-build never builds or stages anything — it restarts whatever jar is already
# sitting at the live path from an earlier successful run. Same check, same message as before.
# fleetd #493: --no-build never builds anything — it restarts whatever jar is already sitting at
# the live path from an earlier successful run. Same check, same message as before.
require_no_build_jar
fi
@@ -1282,8 +1346,8 @@ fi
# fleetd #555: run_drain_gate above is drain_gate_required + the prompt + drain_confirmed, called
# unconditionally — it returns immediately when the gate is not required, and composes/dies through
# refuse_drain_gate itself when the reply does not confirm. See #493/#517/#528 for why "nothing
# changed" would be a lie once a build has staged a jar.
run_drain_gate "$OLD_PID" "$ASSUME_YES" "$DO_BUILD" "$JAR_STAGED"
# changed" would be a lie once a build has produced a jar not yet swapped in.
run_drain_gate "$OLD_PID" "$ASSUME_YES" "$DO_BUILD" "$BUILD_JAR"
# ------------------------------------------------------------------ stop
#
@@ -1334,21 +1398,22 @@ fi
# ------------------------------------------------------------------ swap
#
# fleetd #493: every branch above has now either confirmed the OLD daemon actually exited
# (wait_for_daemon_exit, above) or established there was never one running to begin with. Only
# NOW is it safe to put the freshly built jar at the path the NEXT `java -jar` (direct, or via
# launchd/systemd's ExecStart) will read from — this mv is the one and only write to $JAR anywhere
# fleetd #493/#664: every branch above has now either confirmed the OLD daemon actually exited
# (wait_for_daemon_exit, above) or established there was never one running to begin with. Only NOW
# is it safe to put the freshly built jar at the path the NEXT `java -jar` (direct, or via
# launchd/systemd's ExecStart) will read from — a rename from $BUILD_JAR (target/) to $JAR (run/),
# both under $MODULE and so on one filesystem. This mv is the one and only write to $JAR anywhere
# in this script's mutating flow. If it fails, do not start: die() below exits before "start" runs.
swap_if_built "$DO_BUILD"
# ------------------------------------------------------------------ start
# Unsupervised: login shell (zsh -l) is what puts the secrets on the daemon's environment, and cwd
# must be fleetd/ because the daemon resolves fleetd.yaml, logs/ and target/ relative to it.
# must be fleetd/ because the daemon resolves fleetd.yaml, logs/ and run/ relative to it.
# Supervised (launchd): launchd does both — deploy/dev.ltms.fleetd.plist points ProgramArguments at
# scripts/fleetd-launchd-wrapper.sh (CB-594), which is what execs the login shell in launchd's
# place, and WorkingDirectory in the plist already pins fleetd/.
# Supervised (systemd --user): the unit does both too — measured on the second host, ExecStart is
# `/bin/zsh -lc "exec java -jar target/fleetd.jar fleetd.yaml"` (a login shell, same reason as
# `/bin/zsh -lc "exec java -jar run/fleetd.jar fleetd.yaml"` (a login shell, same reason as
# above) and WorkingDirectory is already pinned to fleetd/.
say "start"
+213 -63
View File
@@ -447,7 +447,7 @@ test_assert_single_daemon_rejects_two_pids() {
test_running_pid_excludes_self_matching_wrapper_shell() {
local before after wrapper_pid
before="$(running_pid)"
sh -c 'echo "target/fleetd.jar" >/dev/null; sleep 20' &
sh -c 'echo "run/fleetd.jar" >/dev/null; sleep 20' &
wrapper_pid=$!
sleep 0.3
after="$(running_pid)"
@@ -475,11 +475,26 @@ test_running_pid_excludes_self_matching_wrapper_shell() {
# `comm` from the actually-executed binary's own path, not from `exec -a`'s argv[0] override (BSD
# ties `comm` to argv[0], which is what makes this technique work here) — so on Linux this
# specific fixture might report `comm=sh`, not `comm=java`, even though the REAL daemon (a literal
# `java -jar target/fleetd.jar` process, never fabricated) is unaffected either way. I could not
# `java -jar run/fleetd.jar` process, never fabricated) is unaffected either way. I could not
# verify this fixture's behavior on Linux, so test_running_pid_counts_a_pid_whose_comm_is_java
# below backstops the same claim (the allowlist admits a pid whose comm is `java`) with a stubbed
# `ps`, which is identical bash on every platform and carries no such platform question.
test_running_pid_finds_a_real_java_named_second_process() {
local before after standin_pid
before="$(running_pid)"
( exec -a java sh -c 'echo "run/fleetd.jar" >/dev/null; sleep 20' ) &
standin_pid=$!
sleep 0.3
after="$(running_pid)"
kill "$standin_pid" 2>/dev/null || true
wait "$standin_pid" 2>/dev/null || true
printf '%s\n' "$after" | grep -qxF "$standin_pid" \
|| fail "running_pid() did not find a real second process (pid $standin_pid, comm forced to 'java' via exec -a) whose own argv holds the pattern: before=[$before] after=[$after]"
}
# PATTERN matches a fleetd jar in either build layout, not only the run/ one: a process whose
# argv names a jar under target/ must be found too, the same way the run/ case above is.
test_running_pid_finds_a_real_java_named_process_from_target_dir() {
local before after standin_pid
before="$(running_pid)"
( exec -a java sh -c 'echo "target/fleetd.jar" >/dev/null; sleep 20' ) &
@@ -489,7 +504,24 @@ test_running_pid_finds_a_real_java_named_second_process() {
kill "$standin_pid" 2>/dev/null || true
wait "$standin_pid" 2>/dev/null || true
printf '%s\n' "$after" | grep -qxF "$standin_pid" \
|| fail "running_pid() did not find a real second process (pid $standin_pid, comm forced to 'java' via exec -a) whose own argv holds the pattern: before=[$before] after=[$after]"
|| fail "running_pid() did not find a real second process (pid $standin_pid, comm forced to 'java' via exec -a) naming a jar under target/: before=[$before] after=[$after]"
}
# Broadening PATTERN to match both build layouts must not also broaden it into matching a
# non-exec'ing shell that merely holds the target/ text as a literal argument, the same
# self-matching shape test_running_pid_excludes_self_matching_wrapper_shell above excludes for
# the run/ text.
test_running_pid_excludes_self_matching_wrapper_shell_naming_target_dir() {
local before after wrapper_pid
before="$(running_pid)"
sh -c 'echo "target/fleetd.jar" >/dev/null; sleep 20' &
wrapper_pid=$!
sleep 0.3
after="$(running_pid)"
kill "$wrapper_pid" 2>/dev/null || true
wait "$wrapper_pid" 2>/dev/null || true
[ "$after" = "$before" ] \
|| fail "running_pid() counted a self-matching wrapper shell (pid $wrapper_pid, holding 'target/fleetd.jar' as literal text in its own argv, not the daemon): before=[$before] after=[$after]"
}
# fleetd #593 CORRECTION 1, hole 2 — the round-1 filter denied known shell names (sh/bash/zsh/
@@ -557,29 +589,30 @@ test_die_message_does_not_recommend_bare_pgrep_as_remediation() {
|| fail "assert_single_daemon's die message does not say in words that a pattern can match the caller (fleetd #593)"
}
# fleetd #511 — jar_id()'s no-argument default was unpinned by any test: nothing proved it reports
# $JAR (the live path) rather than $JAR_STAGED. Both halves matter, so this pins both: the bare call
# must hash the live jar, and an explicit path argument must hash THAT file, not fall back to $JAR.
# Two files with different content, so a default pointed at the wrong one reports the wrong hash
# rather than accidentally matching.
# fleetd #511/#664 — jar_id()'s no-argument default was unpinned by any test: nothing proved it
# reports $JAR (the live path) rather than whatever explicit path a caller passes it (e.g.
# $BUILD_JAR). Both halves matter, so this pins both: the bare call must hash the live jar, and an
# explicit path argument must hash THAT file, not fall back to $JAR. Two files with different
# content, so a default pointed at the wrong one reports the wrong hash rather than accidentally
# matching.
test_jar_id_defaults_to_live_and_reports_explicit_path() {
local dir saved_jar="$JAR" saved_staged="$JAR_STAGED"
local live_hash staged_hash default_result explicit_result
local dir saved_jar="$JAR"
local live_hash other_hash default_result explicit_result other_path
dir="$TMP/jar-id"; mkdir -p "$dir"
JAR="$dir/fleetd.jar"; JAR_STAGED="$dir/fleetd-new.jar"
JAR="$dir/fleetd.jar"; other_path="$dir/other.jar"
printf 'live jar bytes' > "$JAR"
printf 'staged jar bytes, not the same content' > "$JAR_STAGED"
printf 'other jar bytes, not the same content' > "$other_path"
# fleetd #550: this reference hash must be computed the same portable way jar_id() itself now
# computes one — a bare, unguarded call to the macOS-only hasher here was exactly the item-2
# defect, dying with "command not found" on any Linux runner that has no such hasher at all.
live_hash="$(hash256 "$JAR")"
staged_hash="$(hash256 "$JAR_STAGED")"
other_hash="$(hash256 "$other_path")"
default_result="$(jar_id)"
explicit_result="$(jar_id "$JAR_STAGED")"
JAR="$saved_jar"; JAR_STAGED="$saved_staged"
[ "$live_hash" != "$staged_hash" ] || fail "test fixture error: live and staged jars hashed the same"
explicit_result="$(jar_id "$other_path")"
JAR="$saved_jar"
[ "$live_hash" != "$other_hash" ] || fail "test fixture error: live and other jars hashed the same"
assert_equals "$live_hash" "$default_result" "jar_id with no arguments must report the hash of \$JAR"
assert_equals "$staged_hash" "$explicit_result" "jar_id \"\$JAR_STAGED\" must report the hash of the staged jar, not fall back to \$JAR"
assert_equals "$other_hash" "$explicit_result" "jar_id with an explicit path must report the hash of that path, not fall back to \$JAR"
}
# fleetd #550 — closes a gap the test above leaves open. That test's own reference hash is now ALSO
@@ -651,34 +684,77 @@ test_jar_id_reports_unhashable_when_no_hasher_on_path() {
assert_equals "unhashable" "$explicit_result" "jar_id (explicit path) with no hasher on PATH must report the same third state"
}
# fleetd #493 — never build into the path a running process holds. stage_built_jar/swap_staged_jar
# are exercised directly against real files on disk (not stubs), because the whole point is file
# fleetd #664 — report_jar_state is the --check fix: under the old layout $JAR and the build
# output were the same file, so a single "jar on disk" fact covered both. Now they can disagree,
# and this is the function that is supposed to show that. Agreeing case: two files with IDENTICAL
# content must print both labels and never warn.
test_report_jar_state_agrees_when_hashes_match() {
local dir saved_build="$BUILD_JAR" saved_jar="$JAR" output
dir="$TMP/report-jar-agree"; mkdir -p "$dir"
BUILD_JAR="$dir/target-fleetd.jar"; JAR="$dir/run-fleetd.jar"
printf 'identical jar bytes' > "$BUILD_JAR"
printf 'identical jar bytes' > "$JAR"
output="$(report_jar_state)"
BUILD_JAR="$saved_build"; JAR="$saved_jar"
printf '%s' "$output" | grep -qF 'built jar' \
|| fail "report_jar_state did not label the built jar"
printf '%s' "$output" | grep -qF 'running jar' \
|| fail "report_jar_state did not label the running jar"
if printf '%s' "$output" | grep -qF 'differ'; then
fail "report_jar_state warned about a mismatch when both jars have identical content"
fi
}
# The disagreeing case: this is the whole point of the ticket — a built jar that is NOT the
# running jar must be visibly flagged, not silently printed as two unremarkable facts.
test_report_jar_state_warns_when_hashes_differ() {
local dir saved_build="$BUILD_JAR" saved_jar="$JAR" output
dir="$TMP/report-jar-differ"; mkdir -p "$dir"
BUILD_JAR="$dir/target-fleetd.jar"; JAR="$dir/run-fleetd.jar"
printf 'freshly built jar bytes' > "$BUILD_JAR"
printf 'older running jar bytes' > "$JAR"
output="$(report_jar_state)"
BUILD_JAR="$saved_build"; JAR="$saved_jar"
printf '%s' "$output" | grep -qF 'differ' \
|| fail "report_jar_state did not warn when the built jar and running jar disagree"
}
# Neither file existing (a fresh checkout, never built or deployed) must report two "absent"
# facts and never a false mismatch warning — "absent" vs "absent" is agreement, not a diff.
test_report_jar_state_both_absent_is_not_a_mismatch() {
local dir saved_build="$BUILD_JAR" saved_jar="$JAR" output
dir="$TMP/report-jar-absent"; mkdir -p "$dir"
BUILD_JAR="$dir/no-such-target.jar"; JAR="$dir/no-such-run.jar"
output="$(report_jar_state)"
BUILD_JAR="$saved_build"; JAR="$saved_jar"
printf '%s' "$output" | grep -qF 'absent' \
|| fail "report_jar_state did not report absent for a missing built/running jar"
if printf '%s' "$output" | grep -qF 'differ'; then
fail "report_jar_state warned about a mismatch when both jars are simply absent"
fi
}
# Reads JAR and BUILD_JAR exactly as the script sources them, with nothing here assigning
# either first. JAR must resolve outside $MODULE/target/, and JAR must differ from BUILD_JAR:
# the daemon's live path and Maven's own build output are never the same file.
test_jar_and_build_jar_are_sourced_outside_target_and_differ() {
source "$ROOT/scripts/redeploy-fleetd.sh"
case "$JAR" in
"$MODULE"/target/*)
fail "\$JAR must not live under \$MODULE/target/ — got $JAR" ;;
esac
[ "$JAR" != "$BUILD_JAR" ] \
|| fail "\$JAR and \$BUILD_JAR must not be the same path — got $JAR"
}
# fleetd #493/#664 — never build into the path a running process holds. swap_staged_jar is
# exercised directly against real files on disk (not stubs), because the whole point is file
# behavior (does the content move, does the source disappear, does a failure leave both sides
# intact) that a stubbed function cannot prove.
test_stage_built_jar_moves_off_live_path() {
local dir jar staged saved_jar="$JAR" saved_staged="$JAR_STAGED"
dir="$TMP/stage-ok"; mkdir -p "$dir"
jar="$dir/fleetd.jar"; staged="$dir/fleetd-new.jar"
printf 'built jar bytes' > "$jar"
JAR="$jar"; JAR_STAGED="$staged"
stage_built_jar || fail "stage_built_jar rejected a real build output"
JAR="$saved_jar"; JAR_STAGED="$saved_staged"
[ ! -f "$jar" ] || fail "stage_built_jar left the jar behind at the live path $jar"
[ -f "$staged" ] || fail "stage_built_jar did not create the staged jar at $staged"
grep -qF 'built jar bytes' "$staged" || fail "staged jar does not carry the built content"
}
test_stage_built_jar_dies_when_build_produced_nothing() {
local dir output rc=0 saved_jar="$JAR" saved_staged="$JAR_STAGED"
dir="$TMP/stage-missing"; mkdir -p "$dir"
JAR="$dir/fleetd.jar"; JAR_STAGED="$dir/fleetd-new.jar"
output="$(stage_built_jar 2>&1)" || rc=$?
JAR="$saved_jar"; JAR_STAGED="$saved_staged"
[ "$rc" -ne 0 ] || fail "stage_built_jar accepted a missing build output"
printf '%s' "$output" | grep -qF "$dir/fleetd.jar" \
|| fail "refusal message does not name the missing jar path"
}
# intact) that a stubbed function cannot prove. stage_built_jar no longer exists: under the #664
# layout $BUILD_JAR (target/fleetd.jar) was never the live path, so a build has nothing to be
# staged OUT of — swap_staged_jar is called directly against $BUILD_JAR/$JAR (see
# test_swap_ordered_after_wait_and_before_start and the report_jar_state tests below for the rest
# of that seam).
test_swap_staged_jar_moves_staged_onto_live() {
local dir staged live
dir="$TMP/swap-ok"; mkdir -p "$dir"
@@ -989,24 +1065,26 @@ test_swap_ordered_after_wait_and_before_start() {
|| fail "swap_if_built (line $swap_line) is not before the start section (line $start_line)"
}
# fleetd #511: the drain-gate abort message (fired when a build has staged a jar but the operator
# declines the drain confirmation) used to tell the operator to "Rerun (with or without --no-build)"
# to finish the restart. That is wrong — by the time this message can fire, stage_built_jar has
# already moved the jar off $JAR, so a rerun WITH --no-build hits require_no_build_jar's own refusal
# ("no jar at $JAR — run without --no-build"). Like test_swap_ordered_after_wait_and_before_start
# above, this code path is never reached by sourcing (the SOURCED guard stops before the main flow),
# so the only way to pin its exact wording is to read the source.
# fleetd #511: the drain-gate abort message (fired when a build has produced a jar but the
# operator declines the drain confirmation) used to tell the operator to "Rerun (with or without
# --no-build)" to finish the restart. That is wrong. fleetd #664 changed WHY it is wrong: under
# the old layout a build emptied the live path immediately, so --no-build's own check refused a
# rerun for you; now a build never touches the live path at all, so --no-build would NOT refuse —
# it would quietly restart the daemon on the OLD jar and throw away the one just built. Like
# test_swap_ordered_after_wait_and_before_start above, this code path is never reached by sourcing
# (the SOURCED guard stops before the main flow), so the only way to pin its exact wording is to
# read the source.
test_drain_gate_abort_message_says_no_no_build() {
local src="$ROOT/scripts/redeploy-fleetd.sh" msg
msg="$(grep -A3 -F 'aborted — the running daemon was NOT touched, but the freshly built jar is sitting at' "$src")"
msg="$(grep -A6 -F 'aborted — the running daemon was NOT touched, but the freshly built jar is sitting at' "$src")"
[ -n "$msg" ] || fail "could not find the drain-gate staged-jar abort message in redeploy-fleetd.sh"
if printf '%s' "$msg" | grep -qF 'with or without --no-build'; then
fail "abort message still claims a rerun WITH --no-build can finish the restart"
fi
printf '%s' "$msg" | grep -qF 'WITHOUT --no-build' \
|| fail "abort message does not tell the operator to rerun without --no-build"
printf '%s' "$msg" | grep -qF 'no longer at the live path' \
|| fail "abort message does not say why --no-build cannot finish the restart"
printf '%s' "$msg" | grep -qF 'silently discard' \
|| fail "abort message does not say that --no-build would silently discard the build just made"
}
# fleetd #517 — the drain-gate abort branch itself. Before this, the only test of this message was
@@ -1149,19 +1227,22 @@ test_refuse_drain_gate_no_build_staged_absent() {
#
# fleetd #555 — the main flow's own call site moved: it used to read
# `refuse_drain_gate "$DO_BUILD" "$JAR_STAGED"` directly; it now reads
# `run_drain_gate "$OLD_PID" "$ASSUME_YES" "$DO_BUILD" "$JAR_STAGED"`, and run_drain_gate (tested
# directly below by test_run_drain_gate_*) is what calls refuse_drain_gate with its own local names.
# This grep now pins THAT call site — the thing that would go missing if a future edit deleted the
# main flow's call to run_drain_gate altogether, the same residual gap #521/#528 already accepted for
# swap_if_built/refuse_drain_gate (sourcing stops before the main flow runs, so no test in this file
# can do better than reading the source for this one specific gap).
# `run_drain_gate "$OLD_PID" "$ASSUME_YES" "$DO_BUILD" "$BUILD_JAR"` (fleetd #664 renamed the
# fourth argument from $JAR_STAGED to $BUILD_JAR — same role, the not-yet-swapped-in jar — when
# that path stopped being a separate staging file and became target/fleetd.jar itself), and
# run_drain_gate (tested directly below by test_run_drain_gate_*) is what calls refuse_drain_gate
# with its own local names. This grep now pins THAT call site — the thing that would go missing if
# a future edit deleted the main flow's call to run_drain_gate altogether, the same residual gap
# #521/#528 already accepted for swap_if_built/refuse_drain_gate (sourcing stops before the main
# flow runs, so no test in this file can do better than reading the source for this one specific
# gap).
#
# The grep ends `|| true`: this file runs under `set -euo pipefail`, so an ABSENT needle would fail
# the assignment and `set -e` would kill the whole suite before the `[ -n ... ] || fail` guard below
# ever ran — the exact dead-check shape fleetd #528 also flags as a sweep finding (see the PR body).
test_run_drain_gate_call_site_present() {
local src="$ROOT/scripts/redeploy-fleetd.sh" call_line
call_line="$(grep -Fn 'run_drain_gate "$OLD_PID" "$ASSUME_YES" "$DO_BUILD" "$JAR_STAGED"' "$src" | head -1 | cut -d: -f1 || true)"
call_line="$(grep -Fn 'run_drain_gate "$OLD_PID" "$ASSUME_YES" "$DO_BUILD" "$BUILD_JAR"' "$src" | head -1 | cut -d: -f1 || true)"
[ -n "$call_line" ] \
|| fail "could not find the main flow's run_drain_gate call site in redeploy-fleetd.sh"
}
@@ -1269,12 +1350,71 @@ test_run_drain_gate_declined_reply_refuses() {
source "$ROOT/scripts/redeploy-fleetd.sh"
}
# Writes a launchd-plist fixture naming jar_path as the ProgramArguments entry after "-jar", so
# check_jar_path_matches_plist has something real to read back.
write_launchd_plist_fixture() {
local path="$1" jar_path="$2"
cat > "$path" <<PLIST
<?xml version="1.0" encoding="UTF-8"?>
<!DOCTYPE plist PUBLIC "-//Apple//DTD PLIST 1.0//EN" "http://www.apple.com/DTDs/PropertyList-1.0.dtd">
<plist version="1.0">
<dict>
<key>Label</key>
<string>test.fixture</string>
<key>ProgramArguments</key>
<array>
<string>/usr/bin/java</string>
<string>-jar</string>
<string>$jar_path</string>
<string>fleetd.yaml</string>
</array>
</dict>
</plist>
PLIST
}
# Agreeing case: a plist whose ProgramArguments names the same jar, resolved, must proceed and
# say so, never die.
test_check_jar_path_matches_plist_agrees_ok() {
local dir jar plist output rc=0
dir="$TMP/jar-path-agree"; mkdir -p "$dir/run"
jar="$dir/run/fleetd.jar"
plist="$dir/agree.plist"
write_launchd_plist_fixture "$plist" "$jar"
output="$(check_jar_path_matches_plist "$jar" "$plist" 2>&1)" || rc=$?
[ "$rc" -eq 0 ] \
|| fail "check_jar_path_matches_plist must succeed when the plist names the same jar: $output"
printf '%s' "$output" | grep -qF 'jar path check' \
|| fail "check_jar_path_matches_plist did not print the agreement line: $output"
}
# The disagreeing case, and the positive control this check exists for: a plist naming a
# different jar must die, naming both paths.
test_check_jar_path_matches_plist_mismatch_dies() {
local dir jar other_jar plist output rc=0
dir="$TMP/jar-path-mismatch"; mkdir -p "$dir/run" "$dir/target"
jar="$dir/run/fleetd.jar"
other_jar="$dir/target/fleetd.jar"
plist="$dir/mismatch.plist"
write_launchd_plist_fixture "$plist" "$other_jar"
output="$(check_jar_path_matches_plist "$jar" "$plist" 2>&1)" || rc=$?
[ "$rc" -ne 0 ] \
|| fail "check_jar_path_matches_plist must die when the plist names a different jar"
printf '%s' "$output" | grep -qF "$jar" \
|| fail "die message does not name this script's jar path: $output"
printf '%s' "$output" | grep -qF "$other_jar" \
|| fail "die message does not name the plist's jar path: $output"
}
# fleetd #555 item 4 — the report-state dispatch on $SUPERVISOR_KIND. Inverting this used to report
# the wrong supervisor and, for the launchd arm specifically, skip check_log_path_matches_plist.
CHECK_LOG_PATH_CALLED=0
CHECK_JAR_PATH_CALLED=0
stub_check_log_path_recorder() {
CHECK_LOG_PATH_CALLED=0
CHECK_JAR_PATH_CALLED=0
check_log_path_matches_plist() { CHECK_LOG_PATH_CALLED=1; }
check_jar_path_matches_plist() { CHECK_JAR_PATH_CALLED=1; }
}
test_report_supervisor_state_launchd_sets_supervised_and_checks_log_path() {
@@ -1285,6 +1425,8 @@ test_report_supervisor_state_launchd_sets_supervised_and_checks_log_path() {
[ "$SUPERVISED" = 1 ] || fail "report_supervisor_state launchd must set SUPERVISED=1"
[ "$CHECK_LOG_PATH_CALLED" = 1 ] \
|| fail "report_supervisor_state launchd must call check_log_path_matches_plist"
[ "$CHECK_JAR_PATH_CALLED" = 1 ] \
|| fail "report_supervisor_state launchd must call check_jar_path_matches_plist"
source "$ROOT/scripts/redeploy-fleetd.sh"
}
@@ -1296,6 +1438,8 @@ test_report_supervisor_state_systemd_sets_supervised_without_log_path_check() {
[ "$SUPERVISED" = 1 ] || fail "report_supervisor_state systemd must set SUPERVISED=1"
[ "$CHECK_LOG_PATH_CALLED" = 0 ] \
|| fail "report_supervisor_state systemd must NOT call check_log_path_matches_plist"
[ "$CHECK_JAR_PATH_CALLED" = 0 ] \
|| fail "report_supervisor_state systemd must NOT call check_jar_path_matches_plist"
source "$ROOT/scripts/redeploy-fleetd.sh"
}
@@ -2153,6 +2297,8 @@ test_assert_single_daemon_accepts_one_pid
test_assert_single_daemon_rejects_two_pids
test_running_pid_excludes_self_matching_wrapper_shell
test_running_pid_finds_a_real_java_named_second_process
test_running_pid_finds_a_real_java_named_process_from_target_dir
test_running_pid_excludes_self_matching_wrapper_shell_naming_target_dir
test_running_pid_drops_a_pid_whose_comm_is_not_java
test_running_pid_drops_a_pid_that_exited_before_the_comm_lookup
test_running_pid_counts_a_pid_whose_comm_is_java
@@ -2161,8 +2307,10 @@ test_jar_id_defaults_to_live_and_reports_explicit_path
test_hash256_computes_a_real_sha256
test_jar_id_reports_absent_for_missing_file
test_jar_id_reports_unhashable_when_no_hasher_on_path
test_stage_built_jar_moves_off_live_path
test_stage_built_jar_dies_when_build_produced_nothing
test_report_jar_state_agrees_when_hashes_match
test_report_jar_state_warns_when_hashes_differ
test_report_jar_state_both_absent_is_not_a_mismatch
test_jar_and_build_jar_are_sourced_outside_target_and_differ
test_swap_staged_jar_moves_staged_onto_live
test_swap_staged_jar_dies_without_staged_file
test_swap_staged_jar_dies_when_mv_fails
@@ -2204,6 +2352,8 @@ test_drain_confirmed_false_on_anything_else
test_run_drain_gate_skips_prompt_when_not_required
test_run_drain_gate_confirmed_reply_does_not_refuse
test_run_drain_gate_declined_reply_refuses
test_check_jar_path_matches_plist_agrees_ok
test_check_jar_path_matches_plist_mismatch_dies
test_report_supervisor_state_launchd_sets_supervised_and_checks_log_path
test_report_supervisor_state_systemd_sets_supervised_without_log_path_check
test_report_supervisor_state_none_leaves_supervised_zero