Compare commits

..

15 Commits

Author SHA1 Message Date
Dai Ha a134eccc57 fleetd #664: run the daemon from fleetd/run/fleetd.jar, not target/
CI / shell-tests (pull_request) Failing after 8s
CI / contract (pull_request) Successful in 1m4s
CI / build (pull_request) Failing after 2m58s
Separate Maven's output path (target/fleetd.jar) from the daemon's runtime
path (run/fleetd.jar), so a clean/install in the main clone can no longer
reach the jar a running daemon holds open. redeploy-fleetd.sh now swaps the
built jar onto the runtime path with a same-filesystem mv, only after the
old daemon is confirmed gone; --check reports the built and running jars
as two separately labelled hash+mtime facts. Updates every process-locator
pattern and runtime-path reference found by git grep, with a positive
control added for both running_pid()'s PATTERN and fleets-status's pgrep
pattern. fleetd/run/ is gitignored.
2026-10-03 21:06:42 +02:00
Dai Ha 7f9a9c09f9 Merge PR #671: fleetd #670 — pin excludedWorkspaceLabels at FleetdAssembly.java:265
CI / shell-tests (push) Failing after 10s
CI / contract (push) Successful in 58s
CI / build (push) Failing after 1m49s
2026-10-03 20:27:15 +02:00
Dai Ha e854957247 fleetd #670: pin excludedWorkspaceLabels at FleetdAssembly.java:265
CI / shell-tests (pull_request) Failing after 9s
CI / contract (pull_request) Successful in 1m1s
CI / build (pull_request) Failing after 1m58s
Adds a test that reaches the real LeadTabScanner FleetdAssembly's
production boot path builds (via the LeadCoordLoop field that stores
the same leads supplier instance) and asserts, by reflection, that
the excludedWorkspaceLabels field is empty. Mutating line 265 to any
non-empty set now turns this test red.
2026-10-03 20:23:06 +02:00
Dai Ha b4b7cf5155 fleetd #661: drop the validator counts from validateAll's javadoc
CI / shell-tests (push) Failing after 8s
CI / contract (push) Successful in 51s
CI / build (push) Failing after 1m46s
FleetConfig declares eight public no-arg void validate* methods, at lines
2671, 2708, 2755, 2790, 2827, 2856, 2896 and 2958. The sweep's javadoc still
named a count. The wording is now count-free, so it cannot drift again.
2026-10-03 19:59:07 +02:00
Dai Ha cbb35ad947 Merge PR #667: fleetd #661 — refuse pane placement when a lead tab is configured 2026-10-03 19:54:10 +02:00
Dai Ha 656588f597 fleetd #661: add the pane-placement case to the validateAll reachability enumeration
CI / shell-tests (pull_request) Failing after 7s
CI / contract (pull_request) Successful in 53s
CI / build (pull_request) Failing after 1m40s
validateAllReachesEveryOneOfTodaysSixValidators only covered six of
the eight real validators; the new validator was reachability-tested
only from FleetConfigTest, in a different file from the one whose job
is to enumerate every validateAll-reachability case.

Add the pane-placement case to the enumeration, rename the method to
drop the hardcoded count (validateAllReachesEveryOneOfTodaysRealValidators),
and correct the surrounding claims to say seven of eight, naming
validateLeadRollover as the one case still missing (fleetd #668, not
fixed here).
2026-10-03 19:50:16 +02:00
Dai Ha e33377b2ca fleetd #661: fix LeadCount javadoc, dangling @link, and the validator-count word
CI / shell-tests (pull_request) Failing after 7s
CI / contract (pull_request) Successful in 1m3s
CI / build (pull_request) Failing after 1m50s
LeadLauncher.LeadCount's javadoc carried the same false member-space-
exclusion claim as the three comments fixed earlier in this ticket;
its neighbouring body comment in countLeads was already correct and
is unchanged.

FleetConfigValidateAllTest's canary test is renamed to drop the
number from its name (the count now lives only in the Set.of literal
and the javadoc, so the two cannot drift), which also fixes the
dangling {@link} to the old name and the stale 'seventh' wording.
2026-10-03 19:42:50 +02:00
Dai Ha 4b4a8688c2 Merge PR #666: fleetd #664 — correct the main-clone build guidance
CI / shell-tests (push) Failing after 7s
CI / contract (push) Successful in 48s
CI / build (push) Failing after 1m44s
The old text warned only that a 'mvn clean' deletes the running daemon's jar.
Replacement is enough: the shutdown drain loads its classes lazily, at
shutdown, from the jar file the JVM opened at boot. So any build that writes
fleetd/target/fleetd.jar under a live daemon breaks its drain, nothing warns
at the time, and the damage surfaces at the next restart where it looks like
the restart's fault.

Both files now state the real rule and the positive one: verify a merge by
building in a throwaway git worktree, and let only
scripts/redeploy-fleetd.sh touch the main clone's jar.

The canonical block is untouched (20938 bytes, identical to main) and the
wiki sync check passes against wiki/7-Use-Cases.md.
2026-10-03 19:40:35 +02:00
Dai Ha 7e48d4b86c Merge PR #665: fleetd #663 — remove LeadContextGauge.read's 3-arg overload
CI / build (push) Failing after 1m35s
CI / shell-tests (push) Failing after 8s
CI / contract (push) Successful in 1m0s
The 3-arg form delegated to the 4-arg one with a null window, so any caller
reaching for it silently got the fixed HIGH_THRESHOLD_TOKENS back instead of
the profile's effective auto-compact window. It had no production callers.
All 15 test call sites move to the 4-arg form.

Verified at 60fa86a in a throwaway worktree: Tests run: 1923, Failures: 0,
Errors: 0, Skipped: 0, BUILD SUCCESS, 170 surefire report files. Mutating
HIGH_THRESHOLD_TOKENS to 200_000 * 2 turns
noEffectiveWindowFallsBackToTheFixed200000Default red at line 81.
2026-10-03 19:38:13 +02:00
Dai Ha 41cc785534 fleetd #664: explain delayed jar failure
CI / shell-tests (pull_request) Failing after 8s
CI / contract (pull_request) Successful in 55s
CI / build (pull_request) Failing after 1m43s
2026-10-03 19:36:08 +02:00
Dai Ha 60fa86a107 fleetd #663: remove the dead legacy assertion review flagged on PR 665
CI / shell-tests (pull_request) Failing after 9s
CI / contract (pull_request) Successful in 57s
CI / build (pull_request) Failing after 2m6s
noEffectiveWindowFallsBackToTheFixed200000Default's third assertion
(legacyConfigDir/legacyGauge) used to call the 3-arg read() to prove
it behaved like a null window. With the 3-arg form gone, it is the
same call, same input (200_000) and same expectation as the atGauge
assertion above it, so it cannot fail unless that one already failed,
and its message named a method that no longer exists. Deleted it; the
first two assertions (199_999 -> OK, 200_000 -> HIGH) are unchanged.
2026-10-03 19:34:47 +02:00
Dai Ha f288cee2bb fleetd #661: refuse pane placement when a lead tab is configured
CI / shell-tests (pull_request) Failing after 6s
CI / contract (pull_request) Successful in 45s
CI / build (pull_request) Failing after 1m55s
A pane-placed member lands inside the focused tab rather than its own,
so it can land inside a lead's labelled tab and be read back as that
lead by LeadTabScanner, which does not exclude the member space in
production. Add FleetConfig.validatePanePlacementAgainstLeadTabs(),
wired automatically into validateAll() by the existing reflective
sweep, to refuse that combination at startup.

Also correct three stale comments that claimed a member-space
exclusion already blocked this path, in LeadTabScanner, FleetConfig's
validateLeadTabPrefixes javadoc, and LeadTabScannerTest.
2026-10-03 19:34:34 +02:00
Dai Ha a5d6ce1a37 fleetd #664: protect the running daemon jar
CI / shell-tests (pull_request) Failing after 7s
CI / contract (pull_request) Successful in 1m25s
CI / build (pull_request) Failing after 1m53s
2026-10-03 19:29:25 +02:00
Dai Ha a52ca35d34 fleetd #663: remove LeadContextGauge.read's 3-arg overload
CI / shell-tests (pull_request) Failing after 7s
CI / contract (pull_request) Successful in 58s
CI / build (pull_request) Failing after 1m34s
The 3-arg read(configDir, sessionId, agentType) delegated to the 4-arg
form with a null window, silently restoring the fixed 200_000 fallback
that #637 moved away from. It had zero production callers; both
production call sites already use the 4-arg form.

Migrate all test call sites to the 4-arg form. For the generic property
tests in LeadContextGaugeTest, the window is irrelevant and null is
filler. In LeadContextGaugeHighThresholdTest's
noEffectiveWindowFallsBackToTheFixed200000Default, null is the
meaningful value under test, not filler; the assertion is unchanged.
2026-10-03 19:26:11 +02:00
Dai Ha 9417de1123 Merge PR #662: fleetd #659 — remove the dead back-compat forms around LeadConfigDirSource/leadContextLookup
CI / shell-tests (push) Failing after 6s
CI / contract (push) Successful in 1m33s
CI / build (push) Failing after 1m58s
2026-10-03 19:11:18 +02:00
19 changed files with 619 additions and 197 deletions
+1 -1
View File
@@ -59,7 +59,7 @@ as `matches HEAD`, `drift`, or `unknown`; do not turn an unclear timestamp into
Report the process identifier (PID) and uptime too:
```bash
PIDS="$(pgrep -f 'target/fleetd.jar' || true)"
PIDS="$(pgrep -f 'run/fleetd.jar' || true)"
if [ -z "$PIDS" ]; then
printf '%s\n' 'fleetd: not running'
else
+16 -7
View File
@@ -22,13 +22,22 @@ scripts/redeploy-fleetd.sh --yes # skip the drain prompt (fleet already chec
scripts/redeploy-fleetd.sh --no-build # restart the jar already on disk
```
`--no-build` skips the build and restarts whatever jar is at `fleetd/target/fleetd.jar`. Use it only
when you just built and nothing changed since. It gives up the protection in the next paragraph: no
build runs, so a stale or missing jar is not caught early. The script still checks the file is there
and dies with `no jar at … — run without --no-build` if it is not, but it cannot tell you the jar is
old. A `mvn clean` in the tree deletes that jar while the daemon keeps running on it, and nothing
degrades until the next restart. Run `--check` first: it prints the jar's hash and its modification
time, so you can see for yourself whether the jar is missing or older than the code you mean to ship.
`--no-build` skips the build and restarts whatever jar is at `fleetd/run/fleetd.jar` — the runtime
path, not Maven's output path. Use it only when you just built and nothing changed since. It gives
up the protection in the next paragraph: no build runs, so a stale or missing jar is not caught
early. The script still checks the file is there and dies with `no jar at … — run without
--no-build` if it is not, but it cannot tell you the jar is old.
The daemon runs from `fleetd/run/fleetd.jar`, not from `fleetd/target/fleetd.jar` where Maven
writes its output (fleetd #664). That split is what makes a bare `mvn install`/`mvn clean` in the
main clone harmless now: neither can reach the file the running daemon holds open, because that
file no longer lives under `target/` at all. Verify a merge by building in a throwaway git
worktree anyway — a build still produces nothing the fleet runs until this script's own `mv` of
`target/fleetd.jar` onto `run/fleetd.jar`, performed only after the old daemon is confirmed gone.
Let only `scripts/redeploy-fleetd.sh` touch `fleetd/run/fleetd.jar`. Run `--check` first: it prints
the BUILT jar (`target/fleetd.jar`) and the RUNNING jar (`run/fleetd.jar`) as two separately
labelled hash-and-mtime facts, so a mismatch between them — a build sitting unswapped, or a stale
runtime jar — is visible before you decide anything.
It builds before it stops anything, so a failed build never leaves the fleet down; it waits for the
old process to exit rather than assuming; it polls `/healthz`; and it anchors its log checks to a
+7
View File
@@ -323,6 +323,13 @@ must obey belongs in the charter, not here.
reference**, with the intent→tool table above as the short form. `McpContractDocTest` fails if
that page names a `fleet_*` tool the server does not register. The flows are kept out of this
file because this file loads into every session's context.
- **The daemon runs from `fleetd/run/fleetd.jar`, not `fleetd/target/fleetd.jar`** (fleetd #664).
Maven's own output still lands at `fleetd/target/fleetd.jar` — that part of the build is
unchanged — but the running daemon never has that file open, so a bare `mvn install`/`mvn clean`
in the main clone no longer corrupts anything a live process is reading. Verify merges in a
throwaway git worktree anyway: a build in the main clone still ships nothing until
`scripts/redeploy-fleetd.sh` moves it into place with its own atomic `mv`, performed only after
the old daemon is confirmed gone. Let only that script touch `fleetd/run/fleetd.jar`.
### Redeploying the daemon — the lead may do this (primary only)
+1 -1
View File
@@ -47,7 +47,7 @@
<string>/Users/dai.ha/LTMS/claude-bridge/scripts/fleetd-launchd-wrapper.sh</string>
<string>/Users/dai.ha/Softwares/jdks/jdk-25.0.3.jdk/Contents/Home/bin/java</string>
<string>-jar</string>
<string>/Users/dai.ha/LTMS/claude-bridge/fleetd/target/fleetd.jar</string>
<string>/Users/dai.ha/LTMS/claude-bridge/fleetd/run/fleetd.jar</string>
<string>fleetd.yaml</string>
</array>
+1 -1
View File
@@ -50,7 +50,7 @@ WorkingDirectory=%h/LTMS/fleetd/fleetd
# and looks healthy, and the failure appears hours later as a member that cannot open a pull
# request. exec keeps it one process, so systemd tracks the right PID.
# This also avoids a SECOND copy of the secrets in a systemd drop-in: one source of truth.
ExecStart=/bin/zsh -lc "exec java -jar target/fleetd.jar fleetd.yaml"
ExecStart=/bin/zsh -lc "exec java -jar run/fleetd.jar fleetd.yaml"
# PrivateTmp MUST stay false -- see herdr.service. fleetd creates the member ZDOTDIR scrub dir and
# the opencode config dir under java.io.tmpdir, and the member pane (a herdr child, a different
+4
View File
@@ -2,6 +2,10 @@
target/
dependency-reduced-pom.xml
# The daemon's runtime jar (fleetd #664). scripts/redeploy-fleetd.sh moves the built jar here
# with a same-filesystem rename; this is never Maven's output path and never belongs in git.
run/
# Local runtime config (copy from fleetd.example.yaml). Both names are ignored: fleetd.yaml is
# the current name, and bridged.yaml is the legacy name Fleetd still falls back to.
fleetd.yaml
+4 -2
View File
@@ -189,8 +189,10 @@
<build>
<!-- CB-634: the cutover renamed the module dir (bridged/ -> fleetd/), the jar, and the
launchd plist together. The installed plist names fleetd/target/fleetd.jar and
KeepAlive is armed, so this name, the plist, and the wrapper must move as one. -->
launchd plist together. fleetd #664: the installed plist and the systemd unit now name
fleetd/run/fleetd.jar, not this plugin's own output path — see
scripts/redeploy-fleetd.sh for the mv that gets a build from here to there. KeepAlive is
armed, so this name, the plist, and the wrapper must still move as one. -->
<finalName>fleetd</finalName>
<plugins>
<plugin>
@@ -2688,9 +2688,10 @@ public record FleetConfig(
* tab labels — every member gets one rendered into its tab. Choose a lead {@code tabPrefix} that
* a member template matches and the daemon starts labelling its own members as leads, promoting
* the entire fleet to {@link dev.ltms.fleet.auth.Role#PRIMARY} with no message and no diff.
* The member-space exclusion in {@link dev.ltms.fleet.herdr.LeadTabScanner} already blocks the
* realistic path, but defence that depends on one workspace label holding is not defence enough
* for a privilege boundary.
* {@link #validatePanePlacementAgainstLeadTabs()} is the check that stops a pane-placed member
* from landing inside a lead's tab in the first place; this check is a second, independent
* guard that catches the hazard even when every profile places members correctly, by refusing
* a label that a scan would still misread as a lead.
*
* <p>CB-557 shrank this check rather than removing it. The default template is
* {@code "{role}: {profile} #{n}"} and {@code {role}} comes from a closed enum, so a
@@ -2737,6 +2738,45 @@ public record FleetConfig(
+ "lead tabs cannot be confused.");
}
/**
* Reject a profile that places its members by {@code "pane"} while any {@code fleet.leaders}
* entry names a {@code tab}. A pane-placed member lands inside the focused tab rather than its
* own, so it can land inside a lead's own labelled tab. {@link
* dev.ltms.fleet.herdr.LeadTabScanner} identifies a lead purely by that tab's label — it does
* not exclude the member space — so a member that ends up there would be read back as the lead
* and granted spawn/stop/send on the whole fleet.
*
* <p>Only a leader with a non-blank {@code tab} is in scope: one with no {@code tab} feeds
* nothing into {@link dev.ltms.fleet.herdr.LeadTabScanner}, so it creates no hazard here.
*
* @throws IllegalStateException when any {@code profiles:} entry is pane-placed while any
* {@code fleet.leaders} entry names a non-blank {@code tab}
*/
public void validatePanePlacementAgainstLeadTabs() {
if (fleet == null || fleet.leaders().isEmpty()) {
return;
}
boolean anyLeaderHasTab = fleet.leaders().values().stream()
.anyMatch(leader -> leader != null && leader.tab() != null && !leader.tab().isBlank());
if (!anyLeaderHasTab) {
return;
}
List<String> bad = new ArrayList<>();
profiles().entrySet().stream()
.filter(e -> !e.getValue().tabPlacement())
.map(Map.Entry::getKey)
.sorted()
.forEach(bad::add);
if (bad.isEmpty()) {
return;
}
throw new IllegalStateException("refusing to start: profile(s) " + bad
+ " use placement: pane while fleet.leaders names a tab. A pane-placed member can "
+ "land inside a lead's labelled tab and be read back as the lead, granted "
+ "spawn/stop/send on the whole fleet. Set placement: tab for each named profile, "
+ "or remove the tab from every fleet.leaders entry.");
}
/**
* Reject a present {@code leadRollover:} block with no (or a blank) {@code handoverPath}
* (fleetd #480). There is no sane non-null default for an operator-specific file path, unlike
@@ -2937,11 +2977,11 @@ public record FleetConfig(
* Runs every validator this class declares — found by reflection, not by name.
*
* <p>fleetd ticket "central allow-list of usable models", follow-up: mutation testing found
* that although each of the six validators above was well pinned on its own, nothing proved
* that although each validator above was well pinned on its own, nothing proved
* either real caller ({@code Fleetd.main} and {@link ConfigRef#reload()}) still
* invoked it — deleting a call site left the full suite green. The fix is not a seventh test
* per caller; a hand-maintained list of six names here would have the exact same defect its
* own javadoc would warn against: the seventh validator someone adds next month has no reason
* invoked it — deleting a call site left the full suite green. The fix is not one more test
* per caller; a hand-maintained list of names here would have the exact same defect its
* own javadoc would warn against: the next validator someone adds has no reason
* to be added to it. So this method does not name any validator. It sweeps {@link
* #getClass()}'s own public, no-argument, {@code void} methods whose name starts with {@code
* "validate"} (excluding itself) and invokes every one it finds, via {@link
@@ -2950,7 +2990,7 @@ public record FleetConfig(
* which it silently never runs.
*
* <p>{@code Fleetd.main} and {@link ConfigRef#reload()} each call this one method instead of
* the six individually — see the comments at those two call sites for why
* each validator individually — see the comments at those two call sites for why
* each must run it.
*
* <p>Methods run in a fixed (alphabetical) order, so a config with more than one violation
@@ -2967,9 +3007,9 @@ public record FleetConfig(
/**
* The reflective sweep behind {@link #validateAll()}, kept as its own method — taking any
* {@code target}, not just {@code this} — so a test can prove the MECHANISM is generic (it
* would sweep a seventh {@code validateXxx()} method added to any class, not just something
* special-cased to today's six on {@link FleetConfig}) without needing to add a real, unwanted
* seventh validator to this class just to exercise that claim. See {@code
* would sweep any new {@code validateXxx()} method added to any class, not just something
* special-cased to the set {@link FleetConfig} declares today) without needing to add a real,
* unwanted extra validator to this class just to exercise that claim. See {@code
* FleetConfigValidateAllTest} for that proof.
*
* @param target an object whose public, no-argument, {@code void} methods named {@code
@@ -37,8 +37,11 @@ import java.util.function.Supplier;
* <p><strong>Direction of trust.</strong> The label names the lead; it never <em>grants</em>
* anything a pane could take for itself. Three properties keep that honest:
* <ol>
* <li>Worker spaces are excluded wholesale ({@code excludedWorkspaceLabels}), so a worker cannot
* become a lead by being placed — as a split, say — inside a matching tab.</li>
* <li>{@code excludedWorkspaceLabels} can filter a workspace out of the scan, but this class does
* not by itself stop a worker from landing inside a matching tab — a caller may pass an empty
* set, and the daemon does. The guard against that is {@code
* FleetConfig.validatePanePlacementAgainstLeadTabs}: it refuses, at startup, any profile that
* places members by pane while a lead names a tab.</li>
* <li>A worker cannot rename a tab: {@code tab.rename} is reachable only through
* {@link WorkspaceControl}, which no {@code fleet_*} tool exposes. The label is writable by
* the human at the terminal and by nobody the bridge is defending against.</li>
@@ -172,12 +172,6 @@ public final class LeadContextGauge {
* {@code "claude"} (including {@code null}, meaning undetected) reports
* {@link State#UNKNOWN} — this reader only understands Claude Code's own
* transcript format
*/
public Reading read(String configDir, String sessionId, String agentType) {
return read(configDir, sessionId, agentType, null);
}
/**
* @param effectiveWindowTokens the caller's resolved effective auto-compact window for this
* lead's own profile, or {@code null} when it cannot be resolved.
* HIGH fires at {@link #HIGH_THRESHOLD_FRACTION} of this value;
@@ -201,8 +201,8 @@ public final class LeadLauncher {
/**
* How many live leads exist per configured name, and which of that name's labelled tabs are
* <em>not</em> live: a running agent in a tab labelled with that lead's exact {@code tab}
* (CB-579). Member workspaces are excluded, exactly as the scanner excludes them: a member must
* not be counted as a lead because it happens to sit in a matching tab.
* (CB-579). A member sitting in the same shared workspace is not counted as a lead because its
* tab carries a different label, not because any workspace is excluded from this count.
*
* <p>There used to be a second path here — a running agent on the terminal a
* {@code fleet.leaders.<name>.terminal} pin named, for a lead opened and pinned by hand. That
@@ -0,0 +1,201 @@
package dev.ltms.fleet;
import dev.ltms.fleet.config.ConfigRef;
import dev.ltms.fleet.config.FleetConfig;
import dev.ltms.fleet.guard.SubscriptionGuard;
import dev.ltms.fleet.herdr.FakeHerdr;
import dev.ltms.fleet.herdr.HerdrClient;
import dev.ltms.fleet.herdr.LeadTabScanner;
import dev.ltms.fleet.msg.LeadChannelHandle;
import dev.ltms.fleet.msg.LeadCoordLoop;
import dev.ltms.fleet.msg.LeadMessage;
import dev.ltms.fleet.msg.ReplyInbox;
import io.javalin.Javalin;
import org.junit.jupiter.api.Test;
import org.junit.jupiter.api.io.TempDir;
import java.lang.reflect.Field;
import java.nio.file.Files;
import java.nio.file.Path;
import java.util.List;
import java.util.Map;
import java.util.Set;
import java.util.concurrent.Executors;
import java.util.concurrent.ScheduledExecutorService;
import java.util.function.LongSupplier;
import java.util.function.Supplier;
import static org.junit.jupiter.api.Assertions.assertInstanceOf;
import static org.junit.jupiter.api.Assertions.assertNotNull;
import static org.junit.jupiter.api.Assertions.assertTrue;
/**
* fleetd #670 — pins the {@code excludedWorkspaceLabels} argument {@link FleetdAssembly}'s
* production boot path passes to {@link LeadTabScanner} at {@code FleetdAssembly.java:265}
* ({@code Set.of()}).
*
* <p>{@code LeadTabScannerTest} already covers this constructor parameter, but it builds its own
* {@link LeadTabScanner} with its own set, so it tests the seam and proves nothing about the
* producer. This test instead reaches the exact object {@link FleetdAssembly#assembleAndStart}
* builds: a {@code fleet.leaders:} block makes the assembly construct a real
* {@link LeadTabScanner} for its local {@code leads} supplier, and a {@code coordinator:} block
* makes it hand that same supplier instance to {@link LeadCoordLoop} (fleetd #637), which stores
* it as a field. Reflection recovers it from there, and then from the scanner itself, so the
* assertion is against the real production argument rather than a copy built for this test.
*/
class FleetdAssemblyLeadTabScannerExclusionTest {
private static final class FakeLeadChannel implements LeadChannelHandle {
@Override
public void publish(String toCoordId, LeadMessage message) {
}
@Override
public List<LeadMessage> peek() {
return List.of();
}
@Override
public void ack(String msgId) {
}
@Override
public String selfCoordId() {
return "test-lead";
}
@Override
public boolean heldDurable() {
return true;
}
@Override
public MailboxState inspect(String coordId) {
return MailboxState.unknown(coordId);
}
@Override
public void close() {
}
}
private static final class TestResourcePorts implements ResourcePorts {
final FakeHerdr herdr = new FakeHerdr();
Runnable shutdownHook;
@Override
public Map<String, String> environment() {
return Map.of();
}
@Override
public HerdrClient connectHerdr(Path socketPath) {
return herdr;
}
@Override
public Fleetd.AmqpOpener replyInboxOpener() {
return (uri, prefetch) -> new ReplyInbox() {
@Override public void own(String target) { }
@Override public void release(String target) { }
@Override public void publish(String target, String msgId, String content) { }
@Override public List<InboxMessage> peek(String target) { return List.of(); }
@Override public boolean ack(String target, String msgId) { return false; }
};
}
@Override
public Fleetd.LeadMailboxOpener leadMailboxOpener() {
return (uri, selfCoordId, prefetch) -> new FakeLeadChannel();
}
@Override
public LongSupplier nanoClock() {
return System::nanoTime;
}
@Override
public LongSupplier wallClockNanos() {
return System::nanoTime;
}
@Override
public ScheduledExecutorService newScheduler(String purpose) {
return Executors.newSingleThreadScheduledExecutor();
}
@Override
public void addShutdownHook(Runnable hook) {
shutdownHook = hook;
}
@Override
public void startHttp(Javalin app, String host, int port) {
// Do not bind a real port in this assembly test.
}
@Override
public Runnable herdrPollWait() {
return () -> {
throw new UnsupportedOperationException("FakeHerdr is healthy; no poll wait is expected");
};
}
}
private static FleetConfig writeConfig(Path dir) throws Exception {
Path file = dir.resolve("fleetd.yaml");
Files.writeString(file, """
bind:
host: 127.0.0.1
port: 8765
idleSleepGuard:
enabled: false
coordinator:
uri: "amqp://fake-lead-broker/vh"
selfId: "test-lead"
fleet:
leaders:
primary:
tab: "lead: primary"
profile: sonnet
profiles:
sonnet:
subscription: true
argv: ["ccs", "sonnet"]
""");
return FleetConfig.load(file);
}
@Test
void productionBootPathPassesNoExcludedWorkspaceLabels(@TempDir Path dir) throws Exception {
FleetConfig cfg = writeConfig(dir);
TestResourcePorts ports = new TestResourcePorts();
FleetdRuntime runtime = FleetdAssembly.assembleAndStart(new AssemblyInputs(cfg,
new ConfigRef(dir.resolve("fleetd.yaml"), cfg), new SubscriptionGuard(cfg.guard().hostSet())), ports);
try {
LeadCoordLoop coordLoop = runtime.leadCoordLoop();
assertNotNull(coordLoop, "control: a configured coordinator: block must build LeadCoordLoop");
Field leadsField = LeadCoordLoop.class.getDeclaredField("leads");
leadsField.setAccessible(true);
@SuppressWarnings("unchecked")
Supplier<Map<String, String>> leads = (Supplier<Map<String, String>>) leadsField.get(coordLoop);
assertInstanceOf(LeadTabScanner.class, leads,
"control: a non-empty fleet.leaders: block must make FleetdAssembly build a real "
+ "LeadTabScanner for its `leads` supplier, not the Map::of fallback — "
+ "otherwise this test would pass for the wrong reason");
Field excludedField = LeadTabScanner.class.getDeclaredField("excludedWorkspaceLabels");
excludedField.setAccessible(true);
Set<?> excluded = (Set<?>) excludedField.get(leads);
assertTrue(excluded.isEmpty(),
"FleetdAssembly.java:265 must pass an empty excludedWorkspaceLabels to "
+ "LeadTabScanner — scanning member tabs would demote the lead to a worker");
} finally {
assertNotNull(ports.shutdownHook, "control: assembly must capture its shutdown hook");
ports.shutdownHook.run();
}
}
}
@@ -765,6 +765,92 @@ class FleetConfigTest {
"a label that collides with a convention nobody reads is not a problem");
}
// ── validatePanePlacementAgainstLeadTabs ────────────────────────────────────────────────────
/**
* The hazard this guard closes: a pane-placed member lands inside the focused tab rather than
* its own, so it can land inside a lead's labelled tab and be read back as that lead.
*/
@Test
void aPanePlacedProfileWithALeadTabRefusesToStart(@TempDir Path dir) throws Exception {
Path f = dir.resolve("pane-hazard.yaml");
Files.writeString(f, """
bind:
port: 8080
profiles:
gx10:
placement: pane
fleet:
leaders:
opus:
tab: "lead: opus"
""");
FleetConfig cfg = FleetConfig.load(f);
IllegalStateException e = assertThrows(IllegalStateException.class,
cfg::validatePanePlacementAgainstLeadTabs);
assertTrue(e.getMessage().contains("gx10"), "the message must name the offending profile");
}
@Test
void aPanePlacedProfileWithNoLeadTabIsAllowed(@TempDir Path dir) throws Exception {
Path f = dir.resolve("pane-no-tab.yaml");
Files.writeString(f, """
bind:
port: 8080
profiles:
gx10:
placement: pane
fleet:
leaders:
opus:
profile: gx10
""");
assertDoesNotThrow(() -> FleetConfig.load(f).validatePanePlacementAgainstLeadTabs(),
"a leader with no tab feeds nothing into the scanner, so pane placement is safe");
}
@Test
void aTabPlacedProfileWithALeadTabIsAllowed(@TempDir Path dir) throws Exception {
Path f = dir.resolve("tab-safe.yaml");
Files.writeString(f, """
bind:
port: 8080
profiles:
gx10:
placement: tab
fleet:
leaders:
opus:
tab: "lead: opus"
""");
assertDoesNotThrow(() -> FleetConfig.load(f).validatePanePlacementAgainstLeadTabs(),
"a member in its own tab cannot land inside a lead's tab");
}
/** Proves the reflective sweep behind {@code validateAll} really reaches this validator. */
@Test
void validateAllAlsoRefusesPanePlacementAgainstALeadTab(@TempDir Path dir) throws Exception {
Path f = dir.resolve("pane-hazard-sweep.yaml");
Files.writeString(f, """
bind:
port: 8080
profiles:
gx10:
placement: pane
fleet:
leaders:
opus:
tab: "lead: opus"
""");
FleetConfig cfg = FleetConfig.load(f);
IllegalStateException e = assertThrows(IllegalStateException.class, cfg::validateAll);
assertTrue(e.getMessage().contains("gx10"), "the message must name the offending profile");
}
// ── CB-530/CB-579: the leaders registry ─────────────────────────────────────────────────────
@Test
@@ -42,12 +42,13 @@ import static org.junit.jupiter.api.Assertions.assertTrue;
* has right now. This is the proof that a future, real seventh validator on {@link
* FleetConfig} would be swept automatically, without needing to add a real (unwanted)
* seventh validator just to exercise the claim.</li>
* <li>{@link #validateAllReachesEveryOneOfTodaysSixValidators()} proves {@link
* <li>{@link #validateAllReachesEveryOneOfTodaysRealValidators()} proves {@link
* FleetConfig#validateAll()} itself is wired to that same generic mechanism and genuinely
* reaches each of today's six real validators — reusing the exact minimal failing
* reaches seven of today's eight real validators — reusing the exact minimal failing
* configurations {@code FleetConfigTest} already established for each one directly, so a
* single call to {@code validateAll()} is shown to reproduce every one of those six
* failures.</li>
* single call to {@code validateAll()} is shown to reproduce every one of those seven
* failures. The eighth, {@link FleetConfig#validateLeadRollover()}, has no case here yet —
* a pre-existing gap tracked as fleetd #668.</li>
* </ol>
*
* <p>Together with the direct-{@code Fleetd.main}-invocation tests in {@code
@@ -62,8 +63,8 @@ import static org.junit.jupiter.api.Assertions.assertTrue;
* the generic sweep — claim 1 proves {@link FleetConfig#invokeAllValidators} is generic, and claim
* 2 proves {@code validateAll()} reaches today's six, and a hardcoded list satisfies both. So the
* reflective sweep is a convenience, not the guarantee. The guarantee is {@link
* #fleetConfigDeclaresExactlyTheseSixValidatorsToday()}: it fails the moment a seventh validator
* is declared, which forces whoever adds it to look at this file.
* #fleetConfigDeclaresExactlyTheseValidatorsToday()}: it fails the moment any validator is added
* or removed, which forces whoever changes the set to look at this file.
*/
class FleetConfigValidateAllTest {
@@ -209,18 +210,19 @@ class FleetConfigValidateAllTest {
+ "name) must all be skipped");
}
// ── Claim 2: FleetConfig.validateAll() is wired to that mechanism and reaches all six today ──
// ── Claim 2: FleetConfig.validateAll() is wired to that mechanism and reaches seven of eight today ──
/**
* Reflectively enumerates {@link FleetConfig}'s own public, no-arg, void {@code validateXxx()}
* methods (excluding {@code validateAll} itself) — the exact same filter {@link
* FleetConfig#invokeAllValidators} applies. This is not the mechanism proof (that is claim 1,
* above, on an unrelated class) — it is a visible denominator: today there are six, named
* here, so a reader adding a seventh sees this assertion name the new count rather than a
* silent pass at the old one.
* above, on an unrelated class) — it is a visible denominator: today there are eight, named in
* the {@code Set.of} below, so a reader adding or removing one sees this assertion name the new
* count rather than a silent pass at the old one. The count lives only in that set, not in this
* method's name, so the two cannot drift apart.
*/
@Test
void fleetConfigDeclaresExactlyTheseSixValidatorsToday() {
void fleetConfigDeclaresExactlyTheseValidatorsToday() {
Set<String> names = new TreeSet<>();
for (Method m : FleetConfig.class.getMethods()) {
if (java.lang.reflect.Modifier.isPublic(m.getModifiers())
@@ -233,7 +235,8 @@ class FleetConfigValidateAllTest {
}
assertEquals(new TreeSet<>(Set.of("validateAuthExposure", "validateLeadTabPrefixes",
"validateSubscriptionProfiles", "validateCharters", "validateMembers",
"validateModels", "validateLeadRollover")), names,
"validateModels", "validateLeadRollover", "validatePanePlacementAgainstLeadTabs")),
names,
"FleetConfig's public validate*() methods changed. Do TWO things, in this "
+ "order. First confirm validateAll() still delegates to "
+ "invokeAllValidators(this) — a hardcoded list there passes every other "
@@ -262,14 +265,17 @@ class FleetConfigValidateAllTest {
}
/**
* The heart of claim 2: for each of today's six real validators, a minimal file that fails
* The heart of claim 2: for seven of today's eight real validators, a minimal file that fails
* ONLY that one — the exact fixtures {@code FleetConfigTest} uses to test each validator
* directly — must also fail through {@link FleetConfig#validateAll()}. If a future edit to
* {@code validateAll()} silently dropped one validator from the sweep (e.g. a typo'd name
* filter), exactly one of these six would start passing when it must not.
* {@code validateAll()} silently dropped one of these seven from the sweep (e.g. a typo'd name
* filter), exactly one of them would start passing when it must not.
*
* <p>The eighth, {@link FleetConfig#validateLeadRollover()}, has no case here — a pre-existing
* gap tracked as fleetd #668, not fixed by this change.
*/
@Test
void validateAllReachesEveryOneOfTodaysSixValidators(@TempDir Path dir) throws Exception {
void validateAllReachesEveryOneOfTodaysRealValidators(@TempDir Path dir) throws Exception {
// validateAuthExposure: a non-loopback bind without token mode.
assertValidateAllRefuses(dir, "auth-exposure.yaml", """
bind:
@@ -342,6 +348,20 @@ class FleetConfigValidateAllTest {
allow:
- model: claude-sonnet-5
""", "rogue");
// validatePanePlacementAgainstLeadTabs: a pane-placed profile while a lead names a tab.
assertValidateAllRefuses(dir, "pane-placement.yaml", """
bind:
host: 127.0.0.1
port: 8765
profiles:
gx10:
placement: pane
fleet:
leaders:
opus:
tab: "lead: opus"
""", "gx10");
}
private static void assertValidateAllRefuses(Path dir, String fileName, String yaml,
@@ -232,8 +232,9 @@ class LeadTabScannerTest {
@Test
void everyPaneInALeadTabResolvesAsThatLead() {
// A human may split their own lead tab. Both panes are theirs, so both are that lead —
// nothing fleetd placed can land here (see the worker-space test above).
// A human may split their own lead tab. Both panes are theirs, so both are that lead.
// A pane-placed member landing here instead is refused at startup by
// FleetConfig.validatePanePlacementAgainstLeadTabs, not by this scanner.
TopologyHerdr herdr = twoLeads().pane("w1:p1b", "w1:t1", "term_opus_split");
assertEquals("opus-5.0",
@@ -81,12 +81,6 @@ class LeadContextGaugeHighThresholdTest {
assertEquals(LeadContextGauge.State.HIGH,
atGauge.read(atConfigDir, SESSION_ID, "claude", null).state(),
"the fixed default must still be 200,000 when no window is resolvable");
String legacyConfigDir = writeTranscript(tmp.resolve("legacy"), SESSION_ID, 200_000);
LeadContextGauge legacyGauge = new LeadContextGauge();
assertEquals(LeadContextGauge.State.HIGH,
legacyGauge.read(legacyConfigDir, SESSION_ID, "claude").state(),
"the 3-arg read() (no window argument at all) must behave exactly like passing a null window");
}
@Test
@@ -85,7 +85,7 @@ class LeadContextGaugeTest {
usageLine(40_000, 5_000, 3_000)); // last record: 48,000
LeadContextGauge gauge = new LeadContextGauge();
LeadContextGauge.Reading first = gauge.read(configDir, SESSION_ID, "claude");
LeadContextGauge.Reading first = gauge.read(configDir, SESSION_ID, "claude", null);
assertEquals(48_000L, first.tokens(), "must total input+cache_read+cache_creation of the LAST usage record");
assertEquals(LeadContextGauge.State.OK, first.state());
@@ -93,7 +93,7 @@ class LeadContextGaugeTest {
// must change with it, not stay pinned to the first fixture's total.
String otherSession = "22222222-2222-2222-2222-222222222222";
writeTranscript(tmp, otherSession, usageLine(100_000, 50_000, 50_000)); // last record: 200,000
LeadContextGauge.Reading second = gauge.read(configDir, otherSession, "claude");
LeadContextGauge.Reading second = gauge.read(configDir, otherSession, "claude", null);
assertEquals(200_000L, second.tokens());
assertTrue(second.tokens() != first.tokens(), "changing N in the fixture must change the reported number");
}
@@ -111,12 +111,12 @@ class LeadContextGaugeTest {
compactionLine(),
usageLine(3_000, 0, 0));
LeadContextGauge gauge = new LeadContextGauge();
LeadContextGauge.Reading twoCompactions = gauge.read(tmp.toString(), sessionTwoCompactions, "claude");
LeadContextGauge.Reading twoCompactions = gauge.read(tmp.toString(), sessionTwoCompactions, "claude", null);
assertEquals(2, twoCompactions.compactions());
String sessionZeroCompactions = "44444444-4444-4444-4444-444444444444";
writeTranscript(tmp, sessionZeroCompactions, usageLine(3_000, 0, 0));
LeadContextGauge.Reading zeroCompactions = gauge.read(tmp.toString(), sessionZeroCompactions, "claude");
LeadContextGauge.Reading zeroCompactions = gauge.read(tmp.toString(), sessionZeroCompactions, "claude", null);
assertEquals(0, zeroCompactions.compactions(), "changing K in the fixture must change the reported count");
}
@@ -126,7 +126,7 @@ class LeadContextGaugeTest {
@DisplayName("a missing transcript file reports UNKNOWN with no token number")
void missingFileIsUnknown(@TempDir Path tmp) {
LeadContextGauge gauge = new LeadContextGauge();
LeadContextGauge.Reading reading = gauge.read(tmp.toString(), SESSION_ID, "claude");
LeadContextGauge.Reading reading = gauge.read(tmp.toString(), SESSION_ID, "claude", null);
assertEquals(LeadContextGauge.State.UNKNOWN, reading.state());
assertNull(reading.tokens());
}
@@ -148,7 +148,7 @@ class LeadContextGaugeTest {
assumeFalse(Files.isReadable(file),
"runs as root (CI container): the read bit does not stop root, so this case cannot be set up here");
LeadContextGauge gauge = new LeadContextGauge();
LeadContextGauge.Reading reading = gauge.read(configDir, SESSION_ID, "claude");
LeadContextGauge.Reading reading = gauge.read(configDir, SESSION_ID, "claude", null);
assertEquals(LeadContextGauge.State.UNKNOWN, reading.state());
assertNull(reading.tokens());
} finally {
@@ -171,7 +171,7 @@ class LeadContextGaugeTest {
Files.writeString(file, lastCompleteLine + "\n" + tornLine, StandardCharsets.UTF_8);
LeadContextGauge gauge = new LeadContextGauge();
LeadContextGauge.Reading reading = gauge.read(tmp.toString(), SESSION_ID, "claude");
LeadContextGauge.Reading reading = gauge.read(tmp.toString(), SESSION_ID, "claude", null);
assertEquals(LeadContextGauge.State.OK, reading.state(),
"a torn final line must not turn a good earlier reading into UNKNOWN");
assertEquals(6_000L, reading.tokens(),
@@ -187,7 +187,7 @@ class LeadContextGaugeTest {
"{this is not json at all",
"neither is this{{{");
LeadContextGauge gauge = new LeadContextGauge();
LeadContextGauge.Reading reading = gauge.read(configDir, SESSION_ID, "claude");
LeadContextGauge.Reading reading = gauge.read(configDir, SESSION_ID, "claude", null);
assertEquals(LeadContextGauge.State.UNKNOWN, reading.state(),
"every line unparseable is the real format-change signal and must still report UNKNOWN");
assertNull(reading.tokens());
@@ -224,12 +224,12 @@ class LeadContextGaugeTest {
AtomicLong now = new AtomicLong(0);
LeadContextGauge gauge = new LeadContextGauge(now::get, 5_000);
gauge.read(configDir, SESSION_ID, "claude");
gauge.read(configDir, SESSION_ID, "claude"); // still inside the TTL window
gauge.read(configDir, SESSION_ID, "claude", null);
gauge.read(configDir, SESSION_ID, "claude", null); // still inside the TTL window
assertEquals(1, gauge.diskReadCount(), "two reads inside the TTL must touch disk once");
now.set(6_000); // past the TTL
gauge.read(configDir, SESSION_ID, "claude");
gauge.read(configDir, SESSION_ID, "claude", null);
assertEquals(2, gauge.diskReadCount(), "a read past the TTL must touch disk again");
}
@@ -241,8 +241,8 @@ class LeadContextGaugeTest {
String configDir = writeTranscript(tmp, SESSION_ID, usageLine(1_000, 0, 0));
LeadContextGauge gauge = new LeadContextGauge();
assertEquals(LeadContextGauge.State.UNKNOWN, gauge.read(configDir, SESSION_ID, "opencode").state());
assertEquals(LeadContextGauge.State.UNKNOWN, gauge.read(configDir, SESSION_ID, null).state());
assertEquals(LeadContextGauge.State.UNKNOWN, gauge.read(configDir, null, "claude").state());
assertEquals(LeadContextGauge.State.UNKNOWN, gauge.read(configDir, SESSION_ID, "opencode", null).state());
assertEquals(LeadContextGauge.State.UNKNOWN, gauge.read(configDir, SESSION_ID, null, null).state());
assertEquals(LeadContextGauge.State.UNKNOWN, gauge.read(configDir, null, "claude", null).state());
}
}
+87 -63
View File
@@ -71,20 +71,21 @@ set -euo pipefail
REPO="$(cd "$(dirname "${BASH_SOURCE[0]}")/.." && pwd)"
MODULE="$REPO/fleetd"
JAR="$MODULE/target/fleetd.jar"
# fleetd #493: never build into the path a running process holds. The build writes here first
# (Maven's shade plugin has finalName=fleetd, so `clean install` still lands its output at
# target/fleetd.jar — that part is unchanged and out of this script's control), but this script
# now moves it out to JAR_STAGED immediately, and only swaps it back to JAR (a plain `mv`, so a
# rename, never a byte-by-byte overwrite) after the OLD daemon has been confirmed exited. See
# stage_built_jar/swap_staged_jar below.
JAR_STAGED="$MODULE/target/fleetd-new.jar"
# fleetd #664: the runtime path and Maven's output path are no longer the same file. Maven's
# shade plugin (finalName=fleetd) always lands a fresh build at target/fleetd.jar — that is
# Maven's own output directory and this script does not change it — but the daemon is launched
# from $JAR instead, outside target/ entirely. That split is the whole fix: neither `mvn install`
# nor `mvn clean` can ever reach the file a running daemon holds open, because that file no
# longer lives under target/ at all. See swap_if_built/swap_staged_jar below for the one `mv`
# that moves a build from one path to the other, and only after the old daemon is confirmed gone.
BUILD_JAR="$MODULE/target/fleetd.jar"
JAR="$MODULE/run/fleetd.jar"
OUT="$MODULE/fleetd.out"
# Matches BOTH the absolute form and the relative `java -jar target/fleetd.jar` a hand-start
# Matches BOTH the absolute form and the relative `java -jar run/fleetd.jar` a hand-start
# produces from inside fleetd/. Anchoring on the absolute path alone was a real bug: the daemon
# restarted correctly and the script still reported "no process appeared", because it launched with
# a relative path and then looked for an absolute one.
PATTERN='target/fleetd.jar'
PATTERN='run/fleetd.jar'
HEALTH='http://127.0.0.1:8765/healthz'
STOP_WAIT=30 # seconds to wait for a clean exit before reporting failure
HEALTH_WAIT=60 # seconds to wait for /healthz to answer after start — fleetd #603: also the pid-
@@ -169,10 +170,10 @@ hash256() {
fi
}
# Reports the hash of $JAR by default, or of whatever path is passed — used to report the STAGED
# jar right after a build (before it has been swapped in) without ever changing what a bare
# `jar_id` (no args) means: the live path, $JAR. --check and the final "pid ..., jar ..." line
# both call it with no args on purpose, so neither can ever be fooled by a leftover staged file.
# Reports the hash of $JAR by default, or of whatever path is passed — used to report the BUILT
# jar at $BUILD_JAR (before it has been swapped in) without ever changing what a bare `jar_id`
# (no args) means: the live path, $JAR. --check and the final "pid ..., jar ..." line both call
# it with no args on purpose, so neither can ever be fooled by a leftover build output.
# fleetd #550 — THREE distinct answers now, not two: `[ -f "$f" ]` already separates "the jar is
# not there" (-> "absent") from "the jar is there"; for the second case, hash256 itself separates
# "hashed it" (a 12-char hex string) from "could not hash it" (-> "unhashable", when no hasher is
@@ -180,15 +181,35 @@ hash256() {
# was the whole defect this ticket fixes.
jar_id() { local f="${1:-$JAR}"; [ -f "$f" ] && hash256 "$f" || echo "absent"; }
# fleetd #664 — under the old layout $JAR and the build output were the same file, so "jar on
# disk" was one fact. Now they are two: $BUILD_JAR (target/fleetd.jar, whatever Maven last wrote,
# by this script or by a bare `mvn install` run by hand) and $JAR (run/fleetd.jar, whatever the
# daemon actually has open). Printing one label for both was the trap this ticket exists to close
# — during the incident it would have shown the NEW jar's hash while the JVM ran the OLD one.
# Pure (reads jar_id/date, never mutates), so a test can call it directly without reaching the
# main flow — the same shape swap_if_built/drain_gate_refusal already use.
report_jar_state() {
local built_hash running_hash built_mtime running_mtime
built_hash="$(jar_id "$BUILD_JAR")"
running_hash="$(jar_id "$JAR")"
built_mtime="$([ -f "$BUILD_JAR" ] && date -r "$BUILD_JAR" '+%Y-%m-%d %H:%M:%S' || echo 'none')"
running_mtime="$([ -f "$JAR" ] && date -r "$JAR" '+%Y-%m-%d %H:%M:%S' || echo 'none')"
ok "built jar (target/fleetd.jar): $built_hash ($built_mtime)"
ok "running jar (run/fleetd.jar): $running_hash ($running_mtime)"
if [ "$built_hash" != "absent" ] && [ "$running_hash" != "absent" ] && [ "$built_hash" != "$running_hash" ]; then
warn "built jar and running jar differ — target/fleetd.jar was rebuilt since the running daemon last started and is not yet live"
fi
}
# fleetd #593 — `pgrep -f "$PATTERN"` matches ANY process whose full command line CONTAINS the
# pattern text, and that is not the same thing as "is the daemon". A shell that merely embeds the
# pattern as literal text — a human typing this exact investigation by hand, an ssh-shaped
# `sh -c '...; ...'`, a pipeline, or any other non-exec'ing shell that never replaced itself with
# the pattern-holding command — still shows up in that match, and it is the INSTRUMENT, not the
# daemon. Measured live on this Mac: `sh -c 'echo "target/fleetd.jar" >/dev/null; sleep 30' &`
# daemon. Measured live on this Mac: `sh -c 'echo "run/fleetd.jar" >/dev/null; sleep 30' &`
# leaves a real `sh` process alive (it forks for the `sleep`, it does not exec into it) whose own
# `ps -o args` is `sh -c echo "target/fleetd.jar" >/dev/null; sleep 30` — `pgrep -f "$PATTERN"`
# matches that line right alongside the real `java -jar target/fleetd.jar` process. `pgrep -c`
# `ps -o args` is `sh -c echo "run/fleetd.jar" >/dev/null; sleep 30` — `pgrep -f "$PATTERN"`
# matches that line right alongside the real `java -jar run/fleetd.jar` process. `pgrep -c`
# (an in-one-call count) does not exist on BSD/macOS at all, so this cannot be fixed by switching
# pgrep flags — it has to filter what pgrep already found, after the fact, in a way that still
# runs on BSD.
@@ -210,7 +231,7 @@ jar_id() { local f="${1:-$JAR}"; [ -f "$f" ] && hash256 "$f" || echo "absent"; }
# launched as `java -jar ...` — a native image, a renamed launcher — `running_pid()` silently
# returns nothing and `assert_single_daemon` stops noticing a second daemon at all. For a guard,
# that false-negative direction is the worse one to be wrong in. This is not a new assumption,
# though: `PATTERN='target/fleetd.jar'` two lines up already assumes the daemon is a jar, which
# though: `PATTERN='run/fleetd.jar'` two lines up already assumes the daemon is a jar, which
# is only ever run by `java`. If that launch method changes, `PATTERN` stops matching anything
# before this allowlist would ever get the chance to be wrong — the allowlist rides on the same
# assumption that is already load-bearing, it does not add a new one. Whoever changes the launch
@@ -229,16 +250,12 @@ running_pid() {
printf '%s' "$out"
}
# fleetd #493 — three small, independently testable pieces of "never build into the path a
# running process holds":
# fleetd #493/#664 — the independently testable pieces of "never build into the path a running
# process holds":
#
# stage_built_jar moves the jar Maven just produced OUT of the live path and onto the staging
# path, immediately after a successful build. Dies (leaving the OLD daemon
# untouched — this runs before the stop step) if Maven reported success but
# left no jar behind, or if the move itself fails.
# require_no_build_jar the --no-build path never builds or stages anything: it must find a
# jar already sitting at the live path from an earlier successful run, and
# die with the same truthful message this script has always used if not.
# require_no_build_jar the --no-build path never builds anything: it must find a jar already
# sitting at the live path ($JAR, under run/) from an earlier successful run,
# and die with the same truthful message this script has always used if not.
# wait_for_daemon_exit polls running_pid() for up to $1 seconds and reports whether the OLD
# daemon actually exited — extracted to its own function so the main flow
# can be relied on to call swap_staged_jar only AFTER this returns success,
@@ -248,14 +265,10 @@ running_pid() {
# so this is never a write into a path a running process holds — by the time
# it runs, nothing holds that path anymore. If it fails, the caller must not
# start a new daemon: die() below already refuses that by exiting the script.
stage_built_jar() {
[ -f "$JAR" ] || die "build succeeded but produced no jar at $JAR — cannot stage it for restart.
The running daemon was NOT touched."
mv -f "$JAR" "$JAR_STAGED" \
|| die "could not move the freshly built jar from $JAR to the staging path $JAR_STAGED.
The running daemon was NOT touched."
}
# fleetd #664: the "staged" jar swap_if_built passes in is now $BUILD_JAR
# itself (target/fleetd.jar, Maven's own output) — a build no longer needs to
# be moved off the live path right after compiling, because target/ was never
# the live path to begin with.
require_no_build_jar() {
[ -f "$JAR" ] || die "no jar at $JAR — run without --no-build"
}
@@ -282,8 +295,7 @@ swap_staged_jar() {
#
# The defect: the swap step used to be guarded inline by `if [ "$DO_BUILD" = 1 ]` in the main flow.
# Changing that to `if false` left the suite green and the swap never ran, so a redeploy reported
# every step succeeding while the daemon started on no jar at all (stage_built_jar has already moved
# the freshly built one to $JAR_STAGED by then) or on a stale one.
# every step succeeding while the daemon started on no jar at all or on a stale one.
# test_swap_ordered_after_wait_and_before_start could not catch it: it reads this script's own text
# and compares line positions, and a same-line edit moves no line.
#
@@ -312,7 +324,12 @@ swap_if_built() {
local do_build="$1"
should_swap "$do_build" || return 0
say "swap"
swap_staged_jar "$JAR_STAGED" "$JAR"
# fleetd #664: $JAR now lives under run/, a directory target/ never created. mkdir -p here,
# not inside swap_staged_jar itself — that function's own contract is tested on a missing
# parent directory (a failing mv), and widening it to auto-create one would change what that
# test proves.
mkdir -p "$(dirname "$JAR")"
swap_staged_jar "$BUILD_JAR" "$JAR"
ok "jar in place: $(jar_id)"
}
@@ -826,16 +843,25 @@ report_shutdown_drain() {
# --no-build, staged jar present -> ALSO "nothing changed", deliberately: --no-build itself builds
# and stages nothing (see require_no_build_jar above), so a staged jar found here is a leftover
# from an earlier, unrelated run. THIS run truly changed nothing, and the next DO_BUILD=1 run
# wipes that leftover before it builds (`rm -f "$JAR_STAGED"` in the build section above) — so
# there is nothing here for the operator to lose track of.
# wipes that leftover for free — `mvn clean` deletes all of target/, $BUILD_JAR included,
# before the build even starts — so there is nothing here for the operator to lose track of.
# --no-build, staged jar absent -> "nothing changed"
drain_gate_refusal() {
local do_build="$1" staged_path="$2"
if [ "$do_build" = 1 ] && [ -f "$staged_path" ]; then
# fleetd #664: under the old layout a build emptied the live path ($JAR) immediately, so
# --no-build's own check ("no jar at $JAR") was the thing that refused a rerun here. That is
# no longer true: a build never touches $JAR at all now, so $JAR still holds whatever was
# already running before this gate fired (reaching this message at all requires OLD_PID to
# have been set, which means a daemon was running from $JAR already) — a --no-build rerun
# would NOT refuse, it would just restart that same old jar and silently throw away the one
# sitting at %s.
printf 'aborted — the running daemon was NOT touched, but the freshly built jar is sitting at
%s, not yet swapped into %s. Rerun WITHOUT --no-build to finish the restart —
the freshly built jar is no longer at the live path that --no-build requires — or
remove %s by hand if you want to discard this build.' "$staged_path" "$JAR" "$staged_path"
%s, not yet swapped into %s. Rerun WITHOUT --no-build to finish the restart — a
--no-build rerun would NOT refuse here: %s already exists from before this run, so it
would restart the daemon on that OLD jar and silently discard the one you just built — or
remove %s by hand if you want to discard this build instead.' \
"$staged_path" "$JAR" "$JAR" "$staged_path"
else
printf 'aborted — nothing changed'
fi
@@ -1170,7 +1196,7 @@ if [ -n "$OLD_PID" ]; then
else
warn "no daemon running — this will be a cold start"
fi
ok "jar on disk: $(jar_id) ($([ -f "$JAR" ] && date -r "$JAR" '+%Y-%m-%d %H:%M:%S' || echo 'none'))"
report_jar_state
ok "HEAD: $(git -C "$REPO" log --oneline -1)"
# CB-594 / fleetd #492: supervision state. Installed and loaded are different facts — a
@@ -1252,9 +1278,6 @@ stop_if_check_only "$CHECK_ONLY"
if [ "$DO_BUILD" = 1 ]; then
say "build"
# fleetd #493: wipe a leftover staged jar from a previous failed/interrupted run BEFORE doing
# anything else, so that run's leftovers can never be mistaken for this run's output.
rm -f "$JAR_STAGED"
BUILD_LOG="$(mktemp -t fleetd-build.XXXXXX)"
echo " log: $BUILD_LOG"
if ! mvn -f "$MODULE/pom.xml" clean install > "$BUILD_LOG" 2>&1; then
@@ -1264,16 +1287,16 @@ if [ "$DO_BUILD" = 1 ]; then
fi
grep -E '^\[INFO\] Tests run:.*Failures' "$BUILD_LOG" | tail -1 | sed 's/^\[INFO\] / /' || true
ok "BUILD SUCCESS"
# fleetd #493: move the freshly built jar off the live path immediately — the running (OLD)
# daemon, if any, is still up at this point (build always runs before stop). From here until the
# swap step below (after the OLD daemon is confirmed gone), $JAR_STAGED is the only artefact this
# script treats as "the new jar" — $JAR itself is not touched again until the swap.
stage_built_jar
ok "jar now: $(jar_id "$JAR_STAGED")"
# fleetd #664: nothing to stage — $BUILD_JAR (target/fleetd.jar) is Maven's own output path and
# was never the live path, so the running (OLD) daemon, if any, was never at risk from this build
# at all. From here until the swap step below (after the OLD daemon is confirmed gone),
# $BUILD_JAR is the artefact this script treats as "the new jar" — $JAR itself is not touched
# again until the swap.
ok "jar now: $(jar_id "$BUILD_JAR")"
else
say "build skipped (--no-build)"
# fleetd #493: --no-build never builds or stages anything — it restarts whatever jar is already
# sitting at the live path from an earlier successful run. Same check, same message as before.
# fleetd #493: --no-build never builds anything — it restarts whatever jar is already sitting at
# the live path from an earlier successful run. Same check, same message as before.
require_no_build_jar
fi
@@ -1282,8 +1305,8 @@ fi
# fleetd #555: run_drain_gate above is drain_gate_required + the prompt + drain_confirmed, called
# unconditionally — it returns immediately when the gate is not required, and composes/dies through
# refuse_drain_gate itself when the reply does not confirm. See #493/#517/#528 for why "nothing
# changed" would be a lie once a build has staged a jar.
run_drain_gate "$OLD_PID" "$ASSUME_YES" "$DO_BUILD" "$JAR_STAGED"
# changed" would be a lie once a build has produced a jar not yet swapped in.
run_drain_gate "$OLD_PID" "$ASSUME_YES" "$DO_BUILD" "$BUILD_JAR"
# ------------------------------------------------------------------ stop
#
@@ -1334,21 +1357,22 @@ fi
# ------------------------------------------------------------------ swap
#
# fleetd #493: every branch above has now either confirmed the OLD daemon actually exited
# (wait_for_daemon_exit, above) or established there was never one running to begin with. Only
# NOW is it safe to put the freshly built jar at the path the NEXT `java -jar` (direct, or via
# launchd/systemd's ExecStart) will read from — this mv is the one and only write to $JAR anywhere
# fleetd #493/#664: every branch above has now either confirmed the OLD daemon actually exited
# (wait_for_daemon_exit, above) or established there was never one running to begin with. Only NOW
# is it safe to put the freshly built jar at the path the NEXT `java -jar` (direct, or via
# launchd/systemd's ExecStart) will read from — a rename from $BUILD_JAR (target/) to $JAR (run/),
# both under $MODULE and so on one filesystem. This mv is the one and only write to $JAR anywhere
# in this script's mutating flow. If it fails, do not start: die() below exits before "start" runs.
swap_if_built "$DO_BUILD"
# ------------------------------------------------------------------ start
# Unsupervised: login shell (zsh -l) is what puts the secrets on the daemon's environment, and cwd
# must be fleetd/ because the daemon resolves fleetd.yaml, logs/ and target/ relative to it.
# must be fleetd/ because the daemon resolves fleetd.yaml, logs/ and run/ relative to it.
# Supervised (launchd): launchd does both — deploy/dev.ltms.fleetd.plist points ProgramArguments at
# scripts/fleetd-launchd-wrapper.sh (CB-594), which is what execs the login shell in launchd's
# place, and WorkingDirectory in the plist already pins fleetd/.
# Supervised (systemd --user): the unit does both too — measured on the second host, ExecStart is
# `/bin/zsh -lc "exec java -jar target/fleetd.jar fleetd.yaml"` (a login shell, same reason as
# `/bin/zsh -lc "exec java -jar run/fleetd.jar fleetd.yaml"` (a login shell, same reason as
# above) and WorkingDirectory is already pinned to fleetd/.
say "start"
+100 -63
View File
@@ -447,7 +447,7 @@ test_assert_single_daemon_rejects_two_pids() {
test_running_pid_excludes_self_matching_wrapper_shell() {
local before after wrapper_pid
before="$(running_pid)"
sh -c 'echo "target/fleetd.jar" >/dev/null; sleep 20' &
sh -c 'echo "run/fleetd.jar" >/dev/null; sleep 20' &
wrapper_pid=$!
sleep 0.3
after="$(running_pid)"
@@ -475,14 +475,14 @@ test_running_pid_excludes_self_matching_wrapper_shell() {
# `comm` from the actually-executed binary's own path, not from `exec -a`'s argv[0] override (BSD
# ties `comm` to argv[0], which is what makes this technique work here) — so on Linux this
# specific fixture might report `comm=sh`, not `comm=java`, even though the REAL daemon (a literal
# `java -jar target/fleetd.jar` process, never fabricated) is unaffected either way. I could not
# `java -jar run/fleetd.jar` process, never fabricated) is unaffected either way. I could not
# verify this fixture's behavior on Linux, so test_running_pid_counts_a_pid_whose_comm_is_java
# below backstops the same claim (the allowlist admits a pid whose comm is `java`) with a stubbed
# `ps`, which is identical bash on every platform and carries no such platform question.
test_running_pid_finds_a_real_java_named_second_process() {
local before after standin_pid
before="$(running_pid)"
( exec -a java sh -c 'echo "target/fleetd.jar" >/dev/null; sleep 20' ) &
( exec -a java sh -c 'echo "run/fleetd.jar" >/dev/null; sleep 20' ) &
standin_pid=$!
sleep 0.3
after="$(running_pid)"
@@ -557,29 +557,30 @@ test_die_message_does_not_recommend_bare_pgrep_as_remediation() {
|| fail "assert_single_daemon's die message does not say in words that a pattern can match the caller (fleetd #593)"
}
# fleetd #511 — jar_id()'s no-argument default was unpinned by any test: nothing proved it reports
# $JAR (the live path) rather than $JAR_STAGED. Both halves matter, so this pins both: the bare call
# must hash the live jar, and an explicit path argument must hash THAT file, not fall back to $JAR.
# Two files with different content, so a default pointed at the wrong one reports the wrong hash
# rather than accidentally matching.
# fleetd #511/#664 — jar_id()'s no-argument default was unpinned by any test: nothing proved it
# reports $JAR (the live path) rather than whatever explicit path a caller passes it (e.g.
# $BUILD_JAR). Both halves matter, so this pins both: the bare call must hash the live jar, and an
# explicit path argument must hash THAT file, not fall back to $JAR. Two files with different
# content, so a default pointed at the wrong one reports the wrong hash rather than accidentally
# matching.
test_jar_id_defaults_to_live_and_reports_explicit_path() {
local dir saved_jar="$JAR" saved_staged="$JAR_STAGED"
local live_hash staged_hash default_result explicit_result
local dir saved_jar="$JAR"
local live_hash other_hash default_result explicit_result other_path
dir="$TMP/jar-id"; mkdir -p "$dir"
JAR="$dir/fleetd.jar"; JAR_STAGED="$dir/fleetd-new.jar"
JAR="$dir/fleetd.jar"; other_path="$dir/other.jar"
printf 'live jar bytes' > "$JAR"
printf 'staged jar bytes, not the same content' > "$JAR_STAGED"
printf 'other jar bytes, not the same content' > "$other_path"
# fleetd #550: this reference hash must be computed the same portable way jar_id() itself now
# computes one — a bare, unguarded call to the macOS-only hasher here was exactly the item-2
# defect, dying with "command not found" on any Linux runner that has no such hasher at all.
live_hash="$(hash256 "$JAR")"
staged_hash="$(hash256 "$JAR_STAGED")"
other_hash="$(hash256 "$other_path")"
default_result="$(jar_id)"
explicit_result="$(jar_id "$JAR_STAGED")"
JAR="$saved_jar"; JAR_STAGED="$saved_staged"
[ "$live_hash" != "$staged_hash" ] || fail "test fixture error: live and staged jars hashed the same"
explicit_result="$(jar_id "$other_path")"
JAR="$saved_jar"
[ "$live_hash" != "$other_hash" ] || fail "test fixture error: live and other jars hashed the same"
assert_equals "$live_hash" "$default_result" "jar_id with no arguments must report the hash of \$JAR"
assert_equals "$staged_hash" "$explicit_result" "jar_id \"\$JAR_STAGED\" must report the hash of the staged jar, not fall back to \$JAR"
assert_equals "$other_hash" "$explicit_result" "jar_id with an explicit path must report the hash of that path, not fall back to \$JAR"
}
# fleetd #550 — closes a gap the test above leaves open. That test's own reference hash is now ALSO
@@ -651,34 +652,64 @@ test_jar_id_reports_unhashable_when_no_hasher_on_path() {
assert_equals "unhashable" "$explicit_result" "jar_id (explicit path) with no hasher on PATH must report the same third state"
}
# fleetd #493 — never build into the path a running process holds. stage_built_jar/swap_staged_jar
# are exercised directly against real files on disk (not stubs), because the whole point is file
# fleetd #664 — report_jar_state is the --check fix: under the old layout $JAR and the build
# output were the same file, so a single "jar on disk" fact covered both. Now they can disagree,
# and this is the function that is supposed to show that. Agreeing case: two files with IDENTICAL
# content must print both labels and never warn.
test_report_jar_state_agrees_when_hashes_match() {
local dir saved_build="$BUILD_JAR" saved_jar="$JAR" output
dir="$TMP/report-jar-agree"; mkdir -p "$dir"
BUILD_JAR="$dir/target-fleetd.jar"; JAR="$dir/run-fleetd.jar"
printf 'identical jar bytes' > "$BUILD_JAR"
printf 'identical jar bytes' > "$JAR"
output="$(report_jar_state)"
BUILD_JAR="$saved_build"; JAR="$saved_jar"
printf '%s' "$output" | grep -qF 'built jar' \
|| fail "report_jar_state did not label the built jar"
printf '%s' "$output" | grep -qF 'running jar' \
|| fail "report_jar_state did not label the running jar"
if printf '%s' "$output" | grep -qF 'differ'; then
fail "report_jar_state warned about a mismatch when both jars have identical content"
fi
}
# The disagreeing case: this is the whole point of the ticket — a built jar that is NOT the
# running jar must be visibly flagged, not silently printed as two unremarkable facts.
test_report_jar_state_warns_when_hashes_differ() {
local dir saved_build="$BUILD_JAR" saved_jar="$JAR" output
dir="$TMP/report-jar-differ"; mkdir -p "$dir"
BUILD_JAR="$dir/target-fleetd.jar"; JAR="$dir/run-fleetd.jar"
printf 'freshly built jar bytes' > "$BUILD_JAR"
printf 'older running jar bytes' > "$JAR"
output="$(report_jar_state)"
BUILD_JAR="$saved_build"; JAR="$saved_jar"
printf '%s' "$output" | grep -qF 'differ' \
|| fail "report_jar_state did not warn when the built jar and running jar disagree"
}
# Neither file existing (a fresh checkout, never built or deployed) must report two "absent"
# facts and never a false mismatch warning — "absent" vs "absent" is agreement, not a diff.
test_report_jar_state_both_absent_is_not_a_mismatch() {
local dir saved_build="$BUILD_JAR" saved_jar="$JAR" output
dir="$TMP/report-jar-absent"; mkdir -p "$dir"
BUILD_JAR="$dir/no-such-target.jar"; JAR="$dir/no-such-run.jar"
output="$(report_jar_state)"
BUILD_JAR="$saved_build"; JAR="$saved_jar"
printf '%s' "$output" | grep -qF 'absent' \
|| fail "report_jar_state did not report absent for a missing built/running jar"
if printf '%s' "$output" | grep -qF 'differ'; then
fail "report_jar_state warned about a mismatch when both jars are simply absent"
fi
}
# fleetd #493/#664 — never build into the path a running process holds. swap_staged_jar is
# exercised directly against real files on disk (not stubs), because the whole point is file
# behavior (does the content move, does the source disappear, does a failure leave both sides
# intact) that a stubbed function cannot prove.
test_stage_built_jar_moves_off_live_path() {
local dir jar staged saved_jar="$JAR" saved_staged="$JAR_STAGED"
dir="$TMP/stage-ok"; mkdir -p "$dir"
jar="$dir/fleetd.jar"; staged="$dir/fleetd-new.jar"
printf 'built jar bytes' > "$jar"
JAR="$jar"; JAR_STAGED="$staged"
stage_built_jar || fail "stage_built_jar rejected a real build output"
JAR="$saved_jar"; JAR_STAGED="$saved_staged"
[ ! -f "$jar" ] || fail "stage_built_jar left the jar behind at the live path $jar"
[ -f "$staged" ] || fail "stage_built_jar did not create the staged jar at $staged"
grep -qF 'built jar bytes' "$staged" || fail "staged jar does not carry the built content"
}
test_stage_built_jar_dies_when_build_produced_nothing() {
local dir output rc=0 saved_jar="$JAR" saved_staged="$JAR_STAGED"
dir="$TMP/stage-missing"; mkdir -p "$dir"
JAR="$dir/fleetd.jar"; JAR_STAGED="$dir/fleetd-new.jar"
output="$(stage_built_jar 2>&1)" || rc=$?
JAR="$saved_jar"; JAR_STAGED="$saved_staged"
[ "$rc" -ne 0 ] || fail "stage_built_jar accepted a missing build output"
printf '%s' "$output" | grep -qF "$dir/fleetd.jar" \
|| fail "refusal message does not name the missing jar path"
}
# intact) that a stubbed function cannot prove. stage_built_jar no longer exists: under the #664
# layout $BUILD_JAR (target/fleetd.jar) was never the live path, so a build has nothing to be
# staged OUT of — swap_staged_jar is called directly against $BUILD_JAR/$JAR (see
# test_swap_ordered_after_wait_and_before_start and the report_jar_state tests below for the rest
# of that seam).
test_swap_staged_jar_moves_staged_onto_live() {
local dir staged live
dir="$TMP/swap-ok"; mkdir -p "$dir"
@@ -989,24 +1020,26 @@ test_swap_ordered_after_wait_and_before_start() {
|| fail "swap_if_built (line $swap_line) is not before the start section (line $start_line)"
}
# fleetd #511: the drain-gate abort message (fired when a build has staged a jar but the operator
# declines the drain confirmation) used to tell the operator to "Rerun (with or without --no-build)"
# to finish the restart. That is wrong — by the time this message can fire, stage_built_jar has
# already moved the jar off $JAR, so a rerun WITH --no-build hits require_no_build_jar's own refusal
# ("no jar at $JAR — run without --no-build"). Like test_swap_ordered_after_wait_and_before_start
# above, this code path is never reached by sourcing (the SOURCED guard stops before the main flow),
# so the only way to pin its exact wording is to read the source.
# fleetd #511: the drain-gate abort message (fired when a build has produced a jar but the
# operator declines the drain confirmation) used to tell the operator to "Rerun (with or without
# --no-build)" to finish the restart. That is wrong. fleetd #664 changed WHY it is wrong: under
# the old layout a build emptied the live path immediately, so --no-build's own check refused a
# rerun for you; now a build never touches the live path at all, so --no-build would NOT refuse —
# it would quietly restart the daemon on the OLD jar and throw away the one just built. Like
# test_swap_ordered_after_wait_and_before_start above, this code path is never reached by sourcing
# (the SOURCED guard stops before the main flow), so the only way to pin its exact wording is to
# read the source.
test_drain_gate_abort_message_says_no_no_build() {
local src="$ROOT/scripts/redeploy-fleetd.sh" msg
msg="$(grep -A3 -F 'aborted — the running daemon was NOT touched, but the freshly built jar is sitting at' "$src")"
msg="$(grep -A6 -F 'aborted — the running daemon was NOT touched, but the freshly built jar is sitting at' "$src")"
[ -n "$msg" ] || fail "could not find the drain-gate staged-jar abort message in redeploy-fleetd.sh"
if printf '%s' "$msg" | grep -qF 'with or without --no-build'; then
fail "abort message still claims a rerun WITH --no-build can finish the restart"
fi
printf '%s' "$msg" | grep -qF 'WITHOUT --no-build' \
|| fail "abort message does not tell the operator to rerun without --no-build"
printf '%s' "$msg" | grep -qF 'no longer at the live path' \
|| fail "abort message does not say why --no-build cannot finish the restart"
printf '%s' "$msg" | grep -qF 'silently discard' \
|| fail "abort message does not say that --no-build would silently discard the build just made"
}
# fleetd #517 — the drain-gate abort branch itself. Before this, the only test of this message was
@@ -1149,19 +1182,22 @@ test_refuse_drain_gate_no_build_staged_absent() {
#
# fleetd #555 — the main flow's own call site moved: it used to read
# `refuse_drain_gate "$DO_BUILD" "$JAR_STAGED"` directly; it now reads
# `run_drain_gate "$OLD_PID" "$ASSUME_YES" "$DO_BUILD" "$JAR_STAGED"`, and run_drain_gate (tested
# directly below by test_run_drain_gate_*) is what calls refuse_drain_gate with its own local names.
# This grep now pins THAT call site — the thing that would go missing if a future edit deleted the
# main flow's call to run_drain_gate altogether, the same residual gap #521/#528 already accepted for
# swap_if_built/refuse_drain_gate (sourcing stops before the main flow runs, so no test in this file
# can do better than reading the source for this one specific gap).
# `run_drain_gate "$OLD_PID" "$ASSUME_YES" "$DO_BUILD" "$BUILD_JAR"` (fleetd #664 renamed the
# fourth argument from $JAR_STAGED to $BUILD_JAR — same role, the not-yet-swapped-in jar — when
# that path stopped being a separate staging file and became target/fleetd.jar itself), and
# run_drain_gate (tested directly below by test_run_drain_gate_*) is what calls refuse_drain_gate
# with its own local names. This grep now pins THAT call site — the thing that would go missing if
# a future edit deleted the main flow's call to run_drain_gate altogether, the same residual gap
# #521/#528 already accepted for swap_if_built/refuse_drain_gate (sourcing stops before the main
# flow runs, so no test in this file can do better than reading the source for this one specific
# gap).
#
# The grep ends `|| true`: this file runs under `set -euo pipefail`, so an ABSENT needle would fail
# the assignment and `set -e` would kill the whole suite before the `[ -n ... ] || fail` guard below
# ever ran — the exact dead-check shape fleetd #528 also flags as a sweep finding (see the PR body).
test_run_drain_gate_call_site_present() {
local src="$ROOT/scripts/redeploy-fleetd.sh" call_line
call_line="$(grep -Fn 'run_drain_gate "$OLD_PID" "$ASSUME_YES" "$DO_BUILD" "$JAR_STAGED"' "$src" | head -1 | cut -d: -f1 || true)"
call_line="$(grep -Fn 'run_drain_gate "$OLD_PID" "$ASSUME_YES" "$DO_BUILD" "$BUILD_JAR"' "$src" | head -1 | cut -d: -f1 || true)"
[ -n "$call_line" ] \
|| fail "could not find the main flow's run_drain_gate call site in redeploy-fleetd.sh"
}
@@ -2161,8 +2197,9 @@ test_jar_id_defaults_to_live_and_reports_explicit_path
test_hash256_computes_a_real_sha256
test_jar_id_reports_absent_for_missing_file
test_jar_id_reports_unhashable_when_no_hasher_on_path
test_stage_built_jar_moves_off_live_path
test_stage_built_jar_dies_when_build_produced_nothing
test_report_jar_state_agrees_when_hashes_match
test_report_jar_state_warns_when_hashes_differ
test_report_jar_state_both_absent_is_not_a_mismatch
test_swap_staged_jar_moves_staged_onto_live
test_swap_staged_jar_dies_without_staged_file
test_swap_staged_jar_dies_when_mv_fails