Compare commits

..

1 Commits

Author SHA1 Message Date
Dai Ha fe465cc425 fleetd #637: scale LeadContextGauge's HIGH threshold with the effective auto-compact window
CI / shell-tests (pull_request) Failing after 8s
CI / contract (pull_request) Successful in 1m3s
CI / build (pull_request) Failing after 2m2s
HIGH_THRESHOLD_TOKENS was a hardcoded 200_000, while the event it warns about
(auto-compaction) is configured per profile via autoCompactWindow and can legally go
as low as 100_000 — making HIGH unreachable before a compaction on such a profile.

LeadContextGauge.read now takes an optional effective window and fires HIGH at 2/3 of
it, falling back to the fixed 200_000 when no window is resolvable (unresolved callers,
including the pre-existing 3-arg read(), keep today's behaviour exactly).

FleetConfig.Profile.effectiveAutoCompactWindow() resolves that window the way a
launched Claude Code session actually reads it: env.CLAUDE_CODE_AUTO_COMPACT_WINDOW
wins over the autoCompactWindow launch flag when both are set.

Wired into both real consumers: fleet_list's context row (FleetMcp.LeadConfigDirSource,
widened with a back-compat constructor so no unrelated call site changes) and the lead
heartbeat's context-high nudge (Fleetd.leadContextLookup/leadContextSource, widened the
same way).
2026-10-03 15:54:33 +02:00
32 changed files with 277 additions and 1323 deletions
+1 -1
View File
@@ -59,7 +59,7 @@ as `matches HEAD`, `drift`, or `unknown`; do not turn an unclear timestamp into
Report the process identifier (PID) and uptime too:
```bash
PIDS="$(pgrep -f 'run/fleetd.jar' || true)"
PIDS="$(pgrep -f 'target/fleetd.jar' || true)"
if [ -z "$PIDS" ]; then
printf '%s\n' 'fleetd: not running'
else
+7 -16
View File
@@ -22,22 +22,13 @@ scripts/redeploy-fleetd.sh --yes # skip the drain prompt (fleet already chec
scripts/redeploy-fleetd.sh --no-build # restart the jar already on disk
```
`--no-build` skips the build and restarts whatever jar is at `fleetd/run/fleetd.jar` — the runtime
path, not Maven's output path. Use it only when you just built and nothing changed since. It gives
up the protection in the next paragraph: no build runs, so a stale or missing jar is not caught
early. The script still checks the file is there and dies with `no jar at … — run without
--no-build` if it is not, but it cannot tell you the jar is old.
The daemon runs from `fleetd/run/fleetd.jar`, not from `fleetd/target/fleetd.jar` where Maven
writes its output (fleetd #664). That split is what makes a bare `mvn install`/`mvn clean` in the
main clone harmless now: neither can reach the file the running daemon holds open, because that
file no longer lives under `target/` at all. Verify a merge by building in a throwaway git
worktree anyway — a build still produces nothing the fleet runs until this script's own `mv` of
`target/fleetd.jar` onto `run/fleetd.jar`, performed only after the old daemon is confirmed gone.
Let only `scripts/redeploy-fleetd.sh` touch `fleetd/run/fleetd.jar`. Run `--check` first: it prints
the BUILT jar (`target/fleetd.jar`) and the RUNNING jar (`run/fleetd.jar`) as two separately
labelled hash-and-mtime facts, so a mismatch between them — a build sitting unswapped, or a stale
runtime jar — is visible before you decide anything.
`--no-build` skips the build and restarts whatever jar is at `fleetd/target/fleetd.jar`. Use it only
when you just built and nothing changed since. It gives up the protection in the next paragraph: no
build runs, so a stale or missing jar is not caught early. The script still checks the file is there
and dies with `no jar at … — run without --no-build` if it is not, but it cannot tell you the jar is
old. A `mvn clean` in the tree deletes that jar while the daemon keeps running on it, and nothing
degrades until the next restart. Run `--check` first: it prints the jar's hash and its modification
time, so you can see for yourself whether the jar is missing or older than the code you mean to ship.
It builds before it stops anything, so a failed build never leaves the fleet down; it waits for the
old process to exit rather than assuming; it polls `/healthz`; and it anchors its log checks to a
-7
View File
@@ -323,13 +323,6 @@ must obey belongs in the charter, not here.
reference**, with the intent→tool table above as the short form. `McpContractDocTest` fails if
that page names a `fleet_*` tool the server does not register. The flows are kept out of this
file because this file loads into every session's context.
- **The daemon runs from `fleetd/run/fleetd.jar`, not `fleetd/target/fleetd.jar`** (fleetd #664).
Maven's own output still lands at `fleetd/target/fleetd.jar` — that part of the build is
unchanged — but the running daemon never has that file open, so a bare `mvn install`/`mvn clean`
in the main clone no longer corrupts anything a live process is reading. Verify merges in a
throwaway git worktree anyway: a build in the main clone still ships nothing until
`scripts/redeploy-fleetd.sh` moves it into place with its own atomic `mv`, performed only after
the old daemon is confirmed gone. Let only that script touch `fleetd/run/fleetd.jar`.
### Redeploying the daemon — the lead may do this (primary only)
+1 -1
View File
@@ -47,7 +47,7 @@
<string>/Users/dai.ha/LTMS/claude-bridge/scripts/fleetd-launchd-wrapper.sh</string>
<string>/Users/dai.ha/Softwares/jdks/jdk-25.0.3.jdk/Contents/Home/bin/java</string>
<string>-jar</string>
<string>/Users/dai.ha/LTMS/claude-bridge/fleetd/run/fleetd.jar</string>
<string>/Users/dai.ha/LTMS/claude-bridge/fleetd/target/fleetd.jar</string>
<string>fleetd.yaml</string>
</array>
+1 -1
View File
@@ -50,7 +50,7 @@ WorkingDirectory=%h/LTMS/fleetd/fleetd
# and looks healthy, and the failure appears hours later as a member that cannot open a pull
# request. exec keeps it one process, so systemd tracks the right PID.
# This also avoids a SECOND copy of the secrets in a systemd drop-in: one source of truth.
ExecStart=/bin/zsh -lc "exec java -jar run/fleetd.jar fleetd.yaml"
ExecStart=/bin/zsh -lc "exec java -jar target/fleetd.jar fleetd.yaml"
# PrivateTmp MUST stay false -- see herdr.service. fleetd creates the member ZDOTDIR scrub dir and
# the opencode config dir under java.io.tmpdir, and the member pane (a herdr child, a different
-4
View File
@@ -2,10 +2,6 @@
target/
dependency-reduced-pom.xml
# The daemon's runtime jar (fleetd #664). scripts/redeploy-fleetd.sh moves the built jar here
# with a same-filesystem rename; this is never Maven's output path and never belongs in git.
run/
# Local runtime config (copy from fleetd.example.yaml). Both names are ignored: fleetd.yaml is
# the current name, and bridged.yaml is the legacy name Fleetd still falls back to.
fleetd.yaml
+2 -4
View File
@@ -189,10 +189,8 @@
<build>
<!-- CB-634: the cutover renamed the module dir (bridged/ -> fleetd/), the jar, and the
launchd plist together. fleetd #664: the installed plist and the systemd unit now name
fleetd/run/fleetd.jar, not this plugin's own output path — see
scripts/redeploy-fleetd.sh for the mv that gets a build from here to there. KeepAlive is
armed, so this name, the plist, and the wrapper must still move as one. -->
launchd plist together. The installed plist names fleetd/target/fleetd.jar and
KeepAlive is armed, so this name, the plist, and the wrapper must move as one. -->
<finalName>fleetd</finalName>
<plugins>
<plugin>
@@ -1123,12 +1123,17 @@ public final class Fleetd {
* @param liveLeadTerminals terminal id → lead name for every CURRENTLY recognised lead
* @param configDirForLeadName lead name → {@code configDir}, normally {@link
* #leadConfigDirLookup}'s return
* @param windowForLeadName lead name → that lead's profile's effective auto-compact window,
* normally {@link #leadContextWindowLookup}'s return, and passed
* through to {@link LeadContextGauge#read} so the heartbeat's own
* HIGH reading scales with that lead's real window, or {@code null}
* when it cannot be resolved — either way {@link LeadContextGauge}
* falls back to its own fixed HIGH threshold
*/
static Function<String, LeadContextGauge.Reading> leadContextLookup(LeadContextGauge gauge, AgentControl agents,
Supplier<Map<String, String>> liveLeadTerminals, Function<String, String> configDirForLeadName) {
return leadContextLookup(gauge, agents, liveLeadTerminals, configDirForLeadName, _ -> null);
}
/**
* As above, additionally resolving each lead's effective auto-compact window (normally {@link
* #leadContextWindowLookup}'s return) and passing it through to {@link LeadContextGauge#read},
* so the heartbeat's own HIGH reading scales with that lead's real window instead of always the
* gauge's fixed fallback.
*/
static Function<String, LeadContextGauge.Reading> leadContextLookup(LeadContextGauge gauge, AgentControl agents,
Supplier<Map<String, String>> liveLeadTerminals, Function<String, String> configDirForLeadName,
@@ -1157,6 +1162,13 @@ public final class Fleetd {
* factory {@code main} calls, rather than only a lookup nothing in {@code main} is proven to use
* (see {@code FleetdLeadConfigDirSourceWiringTest}'s javadoc for the measured gap this shape closes).
*/
static LeadHeartbeatLoop.LeadContextSource leadContextSource(LeadContextGauge gauge, AgentControl agents,
Supplier<Map<String, String>> liveLeadTerminals, Function<String, String> configDirForLeadName) {
return new LeadHeartbeatLoop.LeadContextSource(
leadContextLookup(gauge, agents, liveLeadTerminals, configDirForLeadName));
}
/** As above, additionally threading the effective-window lookup through. */
static LeadHeartbeatLoop.LeadContextSource leadContextSource(LeadContextGauge gauge, AgentControl agents,
Supplier<Map<String, String>> liveLeadTerminals, Function<String, String> configDirForLeadName,
Function<String, Long> windowForLeadName) {
@@ -2688,10 +2688,9 @@ public record FleetConfig(
* tab labels — every member gets one rendered into its tab. Choose a lead {@code tabPrefix} that
* a member template matches and the daemon starts labelling its own members as leads, promoting
* the entire fleet to {@link dev.ltms.fleet.auth.Role#PRIMARY} with no message and no diff.
* {@link #validatePanePlacementAgainstLeadTabs()} is the check that stops a pane-placed member
* from landing inside a lead's tab in the first place; this check is a second, independent
* guard that catches the hazard even when every profile places members correctly, by refusing
* a label that a scan would still misread as a lead.
* The member-space exclusion in {@link dev.ltms.fleet.herdr.LeadTabScanner} already blocks the
* realistic path, but defence that depends on one workspace label holding is not defence enough
* for a privilege boundary.
*
* <p>CB-557 shrank this check rather than removing it. The default template is
* {@code "{role}: {profile} #{n}"} and {@code {role}} comes from a closed enum, so a
@@ -2738,45 +2737,6 @@ public record FleetConfig(
+ "lead tabs cannot be confused.");
}
/**
* Reject a profile that places its members by {@code "pane"} while any {@code fleet.leaders}
* entry names a {@code tab}. A pane-placed member lands inside the focused tab rather than its
* own, so it can land inside a lead's own labelled tab. {@link
* dev.ltms.fleet.herdr.LeadTabScanner} identifies a lead purely by that tab's label — it does
* not exclude the member space — so a member that ends up there would be read back as the lead
* and granted spawn/stop/send on the whole fleet.
*
* <p>Only a leader with a non-blank {@code tab} is in scope: one with no {@code tab} feeds
* nothing into {@link dev.ltms.fleet.herdr.LeadTabScanner}, so it creates no hazard here.
*
* @throws IllegalStateException when any {@code profiles:} entry is pane-placed while any
* {@code fleet.leaders} entry names a non-blank {@code tab}
*/
public void validatePanePlacementAgainstLeadTabs() {
if (fleet == null || fleet.leaders().isEmpty()) {
return;
}
boolean anyLeaderHasTab = fleet.leaders().values().stream()
.anyMatch(leader -> leader != null && leader.tab() != null && !leader.tab().isBlank());
if (!anyLeaderHasTab) {
return;
}
List<String> bad = new ArrayList<>();
profiles().entrySet().stream()
.filter(e -> !e.getValue().tabPlacement())
.map(Map.Entry::getKey)
.sorted()
.forEach(bad::add);
if (bad.isEmpty()) {
return;
}
throw new IllegalStateException("refusing to start: profile(s) " + bad
+ " use placement: pane while fleet.leaders names a tab. A pane-placed member can "
+ "land inside a lead's labelled tab and be read back as the lead, granted "
+ "spawn/stop/send on the whole fleet. Set placement: tab for each named profile, "
+ "or remove the tab from every fleet.leaders entry.");
}
/**
* Reject a present {@code leadRollover:} block with no (or a blank) {@code handoverPath}
* (fleetd #480). There is no sane non-null default for an operator-specific file path, unlike
@@ -2977,11 +2937,11 @@ public record FleetConfig(
* Runs every validator this class declares — found by reflection, not by name.
*
* <p>fleetd ticket "central allow-list of usable models", follow-up: mutation testing found
* that although each validator above was well pinned on its own, nothing proved
* that although each of the six validators above was well pinned on its own, nothing proved
* either real caller ({@code Fleetd.main} and {@link ConfigRef#reload()}) still
* invoked it — deleting a call site left the full suite green. The fix is not one more test
* per caller; a hand-maintained list of names here would have the exact same defect its
* own javadoc would warn against: the next validator someone adds has no reason
* invoked it — deleting a call site left the full suite green. The fix is not a seventh test
* per caller; a hand-maintained list of six names here would have the exact same defect its
* own javadoc would warn against: the seventh validator someone adds next month has no reason
* to be added to it. So this method does not name any validator. It sweeps {@link
* #getClass()}'s own public, no-argument, {@code void} methods whose name starts with {@code
* "validate"} (excluding itself) and invokes every one it finds, via {@link
@@ -2990,7 +2950,7 @@ public record FleetConfig(
* which it silently never runs.
*
* <p>{@code Fleetd.main} and {@link ConfigRef#reload()} each call this one method instead of
* each validator individually — see the comments at those two call sites for why
* the six individually — see the comments at those two call sites for why
* each must run it.
*
* <p>Methods run in a fixed (alphabetical) order, so a config with more than one violation
@@ -3007,9 +2967,9 @@ public record FleetConfig(
/**
* The reflective sweep behind {@link #validateAll()}, kept as its own method — taking any
* {@code target}, not just {@code this} — so a test can prove the MECHANISM is generic (it
* would sweep any new {@code validateXxx()} method added to any class, not just something
* special-cased to the set {@link FleetConfig} declares today) without needing to add a real,
* unwanted extra validator to this class just to exercise that claim. See {@code
* would sweep a seventh {@code validateXxx()} method added to any class, not just something
* special-cased to today's six on {@link FleetConfig}) without needing to add a real, unwanted
* seventh validator to this class just to exercise that claim. See {@code
* FleetConfigValidateAllTest} for that proof.
*
* @param target an object whose public, no-argument, {@code void} methods named {@code
@@ -37,11 +37,8 @@ import java.util.function.Supplier;
* <p><strong>Direction of trust.</strong> The label names the lead; it never <em>grants</em>
* anything a pane could take for itself. Three properties keep that honest:
* <ol>
* <li>{@code excludedWorkspaceLabels} can filter a workspace out of the scan, but this class does
* not by itself stop a worker from landing inside a matching tab — a caller may pass an empty
* set, and the daemon does. The guard against that is {@code
* FleetConfig.validatePanePlacementAgainstLeadTabs}: it refuses, at startup, any profile that
* places members by pane while a lead names a tab.</li>
* <li>Worker spaces are excluded wholesale ({@code excludedWorkspaceLabels}), so a worker cannot
* become a lead by being placed — as a split, say — inside a matching tab.</li>
* <li>A worker cannot rename a tab: {@code tab.rename} is reachable only through
* {@link WorkspaceControl}, which no {@code fleet_*} tool exposes. The label is writable by
* the human at the terminal and by nobody the bridge is defending against.</li>
@@ -68,10 +68,8 @@ import java.util.function.LongSupplier;
* (never the whole 52 MB a long-lived transcript reaches on the host this was measured on), and
* {@link #DEFAULT_CACHE_TTL_MILLIS} bounds how often that bounded read actually happens — a burst
* of {@code fleet_list} calls inside one TTL window reads the file once. One instance's cache is
* keyed by {@code (configDir, sessionId, highThreshold)}, so it is safe to share across every lead
* a single {@code fleet_list} call reports on, and a call that resolves a different effective
* window for the same lead never reads back a state computed against the other window's
* threshold.
* keyed by {@code (configDir, sessionId)}, so it is safe to share across every lead a single
* {@code fleet_list} call reports on.
*/
public final class LeadContextGauge {
@@ -172,6 +170,12 @@ public final class LeadContextGauge {
* {@code "claude"} (including {@code null}, meaning undetected) reports
* {@link State#UNKNOWN} — this reader only understands Claude Code's own
* transcript format
*/
public Reading read(String configDir, String sessionId, String agentType) {
return read(configDir, sessionId, agentType, null);
}
/**
* @param effectiveWindowTokens the caller's resolved effective auto-compact window for this
* lead's own profile, or {@code null} when it cannot be resolved.
* HIGH fires at {@link #HIGH_THRESHOLD_FRACTION} of this value;
@@ -188,14 +192,13 @@ public final class LeadContextGauge {
String base = (configDir == null || configDir.isBlank())
? System.getProperty("user.home") + "/.claude"
: configDir;
long highThreshold = highThreshold(effectiveWindowTokens);
String cacheKey = base + '\u0000' + sessionId + '\u0000' + highThreshold;
String cacheKey = base + '\u0000' + sessionId;
long now = clock.getAsLong();
CacheEntry cached = cache.get(cacheKey);
if (cached != null && now - cached.readAtMillis() < ttlMillis) {
return cached.reading();
}
Reading fresh = readUncached(base, sessionId, highThreshold);
Reading fresh = readUncached(base, sessionId, highThreshold(effectiveWindowTokens));
cache.put(cacheKey, new CacheEntry(fresh, now));
return fresh;
}
@@ -201,8 +201,8 @@ public final class LeadLauncher {
/**
* How many live leads exist per configured name, and which of that name's labelled tabs are
* <em>not</em> live: a running agent in a tab labelled with that lead's exact {@code tab}
* (CB-579). A member sitting in the same shared workspace is not counted as a lead because its
* tab carries a different label, not because any workspace is excluded from this count.
* (CB-579). Member workspaces are excluded, exactly as the scanner excludes them: a member must
* not be counted as a lead because it happens to sit in a matching tab.
*
* <p>There used to be a second path here — a running agent on the terminal a
* {@code fleet.leaders.<name>.terminal} pin named, for a lead opened and pinned by hand. That
@@ -293,6 +293,14 @@ public final class FleetMcp {
public record LeadConfigDirSource(Function<String, String> configDirFor, Function<String, Long> windowFor) {
/** Inert source — every lead reads {@link LeadContextGauge}'s built-in defaults. */
public static LeadConfigDirSource none() { return new LeadConfigDirSource(_ -> null, _ -> null); }
/**
* Constructor that resolves no window — every lead this source answers for keeps
* {@link LeadContextGauge}'s fixed HIGH threshold.
*/
public LeadConfigDirSource(Function<String, String> configDirFor) {
this(configDirFor, _ -> null);
}
}
/**
@@ -58,31 +58,20 @@ import static org.junit.jupiter.api.Assertions.assertEquals;
* {@code HttpClient} — no accessor needed for this half.
*
* <p><strong>{@code GET /sessions} could not be driven the same way</strong>, so this class does
* not pin the merge half of the deleted test's javadoc. This class configures no {@code auth:}
* block, so it runs under the default {@code loopback-trust} mode ({@code FleetConfig}). Under
* that mode, {@code /sessions} requires {@code Authz.Action.READ}, which — through the REAL
* assembly's real {@code CallerResolver}/{@code ConnectionIdentity} (built with a hardcoded
* {@code new LsofPeerPidLookup()}) — needs {@code Caller.resolved()}, i.e. a real positive pid
* from {@code lsof}. {@code LsofPeerPidLookup} excludes its own pid (see its javadoc), and a
* JUnit test's HTTP client and the daemon under test share one JVM pid, so the resolved pid is
* always {@code -1} and every such request is refused as {@code ANONYMOUS} (fleetd #317's
* fail-closed rule) before the route handler — and its {@code memberHerdr} merge — is ever
* reached. Verified directly: driving {@code GET /sessions} here returns {@code 401
* unauthenticated}, not the merged body. {@code FleetAppTwoDaemonTest} avoids this because it
* builds {@code FleetApp} with {@code callers: null}, which is not what the real assembly
* passes. The {@code /healthz} pin below is what this class relies on for CB-185's {@code
* FleetApp} half; {@code FleetAppTwoDaemonTest} remains the full behavioural proof that
* {@code FleetApp} itself merges {@code /sessions} correctly once handed two clients.
*
* <p><strong>This refusal is {@code loopback-trust}-specific, not a property of {@code
* CallerResolver} in general.</strong> Under {@code auth.mode: token}, {@code
* CallerResolver#resolve} returns before ever consulting {@code Caller.resolved()} or {@code
* Caller.scanComplete()}: a request carrying a valid bearer token in its {@code Authorization}
* header resolves to {@code Role#PRIMARY} with no pid lookup at all, so the same-JVM-pid
* exclusion above never comes into play. {@code FleetdQuarantineOutageDualWindowAssemblyTest}
* and {@code FleetdListReportingSourcesAssemblyTest} both drive {@code Authz.Action.READ} this
* way, over a real {@code McpSyncClient}/{@code HttpClient} against a real {@code
* FleetdAssembly#assembleAndStart}, and both get the real response rather than a refusal.
* not pin the merge half of the deleted test's javadoc. {@code /sessions} requires
* {@code Authz.Action.READ}, which — through the REAL assembly's real {@code
* CallerResolver}/{@code ConnectionIdentity} (built with a hardcoded {@code
* new LsofPeerPidLookup()}) — needs {@code Caller.resolved()}, i.e. a real positive pid from
* {@code lsof}. {@code LsofPeerPidLookup} excludes its own pid (see its javadoc), and a JUnit
* test's HTTP client and the daemon under test share one JVM pid, so the resolved pid is always
* {@code -1} and every such request is refused as {@code ANONYMOUS} (fleetd #317's fail-closed
* rule) before the route handler — and its {@code memberHerdr} merge — is ever reached. Verified
* directly: driving {@code GET /sessions} here returns {@code 401 unauthenticated}, not the
* merged body. {@code FleetAppTwoDaemonTest} avoids this because it builds {@code FleetApp} with
* {@code callers: null}, which is not what the real assembly passes. The {@code /healthz} pin
* below is what this class relies on for CB-185's {@code FleetApp} half; {@code
* FleetAppTwoDaemonTest} remains the full behavioural proof that {@code FleetApp} itself merges
* {@code /sessions} correctly once handed two clients.
*
* <p><strong>fleetd #629 follow-up.</strong> The fix below (see {@link TwoHerdrResourcePorts})
* makes {@link #healthzGoesRedWhenTheLeadDaemonIsDownEvenThoughTheMemberIsUp}'s fake {@code
@@ -1,201 +0,0 @@
package dev.ltms.fleet;
import dev.ltms.fleet.config.ConfigRef;
import dev.ltms.fleet.config.FleetConfig;
import dev.ltms.fleet.guard.SubscriptionGuard;
import dev.ltms.fleet.herdr.FakeHerdr;
import dev.ltms.fleet.herdr.HerdrClient;
import dev.ltms.fleet.herdr.LeadTabScanner;
import dev.ltms.fleet.msg.LeadChannelHandle;
import dev.ltms.fleet.msg.LeadCoordLoop;
import dev.ltms.fleet.msg.LeadMessage;
import dev.ltms.fleet.msg.ReplyInbox;
import io.javalin.Javalin;
import org.junit.jupiter.api.Test;
import org.junit.jupiter.api.io.TempDir;
import java.lang.reflect.Field;
import java.nio.file.Files;
import java.nio.file.Path;
import java.util.List;
import java.util.Map;
import java.util.Set;
import java.util.concurrent.Executors;
import java.util.concurrent.ScheduledExecutorService;
import java.util.function.LongSupplier;
import java.util.function.Supplier;
import static org.junit.jupiter.api.Assertions.assertInstanceOf;
import static org.junit.jupiter.api.Assertions.assertNotNull;
import static org.junit.jupiter.api.Assertions.assertTrue;
/**
* fleetd #670 — pins the {@code excludedWorkspaceLabels} argument {@link FleetdAssembly}'s
* production boot path passes to {@link LeadTabScanner} at {@code FleetdAssembly.java:265}
* ({@code Set.of()}).
*
* <p>{@code LeadTabScannerTest} already covers this constructor parameter, but it builds its own
* {@link LeadTabScanner} with its own set, so it tests the seam and proves nothing about the
* producer. This test instead reaches the exact object {@link FleetdAssembly#assembleAndStart}
* builds: a {@code fleet.leaders:} block makes the assembly construct a real
* {@link LeadTabScanner} for its local {@code leads} supplier, and a {@code coordinator:} block
* makes it hand that same supplier instance to {@link LeadCoordLoop} (fleetd #637), which stores
* it as a field. Reflection recovers it from there, and then from the scanner itself, so the
* assertion is against the real production argument rather than a copy built for this test.
*/
class FleetdAssemblyLeadTabScannerExclusionTest {
private static final class FakeLeadChannel implements LeadChannelHandle {
@Override
public void publish(String toCoordId, LeadMessage message) {
}
@Override
public List<LeadMessage> peek() {
return List.of();
}
@Override
public void ack(String msgId) {
}
@Override
public String selfCoordId() {
return "test-lead";
}
@Override
public boolean heldDurable() {
return true;
}
@Override
public MailboxState inspect(String coordId) {
return MailboxState.unknown(coordId);
}
@Override
public void close() {
}
}
private static final class TestResourcePorts implements ResourcePorts {
final FakeHerdr herdr = new FakeHerdr();
Runnable shutdownHook;
@Override
public Map<String, String> environment() {
return Map.of();
}
@Override
public HerdrClient connectHerdr(Path socketPath) {
return herdr;
}
@Override
public Fleetd.AmqpOpener replyInboxOpener() {
return (uri, prefetch) -> new ReplyInbox() {
@Override public void own(String target) { }
@Override public void release(String target) { }
@Override public void publish(String target, String msgId, String content) { }
@Override public List<InboxMessage> peek(String target) { return List.of(); }
@Override public boolean ack(String target, String msgId) { return false; }
};
}
@Override
public Fleetd.LeadMailboxOpener leadMailboxOpener() {
return (uri, selfCoordId, prefetch) -> new FakeLeadChannel();
}
@Override
public LongSupplier nanoClock() {
return System::nanoTime;
}
@Override
public LongSupplier wallClockNanos() {
return System::nanoTime;
}
@Override
public ScheduledExecutorService newScheduler(String purpose) {
return Executors.newSingleThreadScheduledExecutor();
}
@Override
public void addShutdownHook(Runnable hook) {
shutdownHook = hook;
}
@Override
public void startHttp(Javalin app, String host, int port) {
// Do not bind a real port in this assembly test.
}
@Override
public Runnable herdrPollWait() {
return () -> {
throw new UnsupportedOperationException("FakeHerdr is healthy; no poll wait is expected");
};
}
}
private static FleetConfig writeConfig(Path dir) throws Exception {
Path file = dir.resolve("fleetd.yaml");
Files.writeString(file, """
bind:
host: 127.0.0.1
port: 8765
idleSleepGuard:
enabled: false
coordinator:
uri: "amqp://fake-lead-broker/vh"
selfId: "test-lead"
fleet:
leaders:
primary:
tab: "lead: primary"
profile: sonnet
profiles:
sonnet:
subscription: true
argv: ["ccs", "sonnet"]
""");
return FleetConfig.load(file);
}
@Test
void productionBootPathPassesNoExcludedWorkspaceLabels(@TempDir Path dir) throws Exception {
FleetConfig cfg = writeConfig(dir);
TestResourcePorts ports = new TestResourcePorts();
FleetdRuntime runtime = FleetdAssembly.assembleAndStart(new AssemblyInputs(cfg,
new ConfigRef(dir.resolve("fleetd.yaml"), cfg), new SubscriptionGuard(cfg.guard().hostSet())), ports);
try {
LeadCoordLoop coordLoop = runtime.leadCoordLoop();
assertNotNull(coordLoop, "control: a configured coordinator: block must build LeadCoordLoop");
Field leadsField = LeadCoordLoop.class.getDeclaredField("leads");
leadsField.setAccessible(true);
@SuppressWarnings("unchecked")
Supplier<Map<String, String>> leads = (Supplier<Map<String, String>>) leadsField.get(coordLoop);
assertInstanceOf(LeadTabScanner.class, leads,
"control: a non-empty fleet.leaders: block must make FleetdAssembly build a real "
+ "LeadTabScanner for its `leads` supplier, not the Map::of fallback — "
+ "otherwise this test would pass for the wrong reason");
Field excludedField = LeadTabScanner.class.getDeclaredField("excludedWorkspaceLabels");
excludedField.setAccessible(true);
Set<?> excluded = (Set<?>) excludedField.get(leads);
assertTrue(excluded.isEmpty(),
"FleetdAssembly.java:265 must pass an empty excludedWorkspaceLabels to "
+ "LeadTabScanner — scanning member tabs would demote the lead to a worker");
} finally {
assertNotNull(ports.shutdownHook, "control: assembly must capture its shutdown hook");
ports.shutdownHook.run();
}
}
}
@@ -24,8 +24,8 @@ import static org.junit.jupiter.api.Assertions.assertTrue;
* {@code ConfigRefTest} and {@code FleetdConfigRefCharterToolSurfaceWiringTest} case — because
* neither of those tests constructs its {@code ConfigRef} through {@code main}; both build their own
* instance directly, wired with the check by hand. That silent regression is exactly the shape
* {@link FleetdBackendQuarantineAssemblyTest}, {@link FleetdLeadSeatAssemblyTest} and {@link
* FleetdCompletionResolverAssemblyTest} already guard against for their own constructor arguments —
* {@link FleetdBackendQuarantineWiringTest}, {@link FleetdLeadSeatWiringTest} and {@link
* FleetdCompletionResolverWiringTest} already guard against for their own constructor arguments —
* this class is the same class of gap for fleetd #474's {@code extraValidation} argument, following
* their approach.
*
@@ -2,67 +2,16 @@ package dev.ltms.fleet;
import java.nio.file.Files;
import java.nio.file.Path;
import org.junit.jupiter.api.DisplayName;
import org.junit.jupiter.api.Test;
import static org.junit.jupiter.api.Assertions.assertFalse;
import static org.junit.jupiter.api.Assertions.assertTrue;
/**
* {@code AgentControl} caches {@code paneByTerminal}, so {@code HerdrRouter} must be its only
* production factory — a second instance means a second cache; the same reasoning applies to
* {@code WorkspaceControl}. {@code HerdrRouter}'s constructor is the one place both are built.
*
* <p><b>This test checks source text, not runtime behaviour.</b> It never constructs a {@code
* HerdrRouter} and never runs {@code FleetdAssembly.assembleAndStart} — a green result proves only
* that neither watched file's text contains {@code new AgentControl(} or {@code new
* WorkspaceControl(}. It does not prove the instances {@code HerdrRouter} does build are the ones
* actually wired through the rest of the daemon, and it does not cover a bypass written into a
* production file other than the two this test reads.
*/
class FleetdHerdrControlConstructionTest {
private static String source(String relativePath) throws Exception {
return Files.readString(Path.of(relativePath));
}
@Test
@DisplayName("[SOURCE TEXT] Fleetd.java never constructs AgentControl or WorkspaceControl directly")
void fleetdDelegatesStatefulControlsToTheRouter() throws Exception {
String source = source("src/main/java/dev/ltms/fleet/Fleetd.java");
// A broken read (wrong working directory, wrong path, a file that came back empty) would
// make the assertFalse checks below pass vacuously — a "clean" negative check that actually
// checked nothing. Guard against that first, with an anchor that has nothing to do with
// this mutation, so a bad read fails loudly here instead of silently proving nothing below.
assertTrue(source.contains("public final class Fleetd"),
"the read of Fleetd.java did not come back containing its own class declaration — "
+ "the assertFalse checks below would pass vacuously on a broken read; fix the "
+ "read before trusting this test.");
assertFalse(source.contains("new AgentControl("),
"Fleetd.java must not construct AgentControl directly — HerdrRouter is its only "
+ "production factory");
assertFalse(source.contains("new WorkspaceControl("),
"Fleetd.java must not construct WorkspaceControl directly — HerdrRouter is its only "
+ "production factory");
}
@Test
@DisplayName("[SOURCE TEXT] FleetdAssembly.java never constructs AgentControl or WorkspaceControl directly")
void fleetdAssemblyDelegatesStatefulControlsToTheRouter() throws Exception {
String source = source("src/main/java/dev/ltms/fleet/FleetdAssembly.java");
assertTrue(source.contains("final class FleetdAssembly"),
"the read of FleetdAssembly.java did not come back containing its own class "
+ "declaration — the assertFalse checks below would pass vacuously on a broken "
+ "read; fix the read before trusting this test.");
assertFalse(source.contains("new AgentControl("),
"FleetdAssembly.java must not construct AgentControl directly — HerdrRouter is its "
+ "only production factory");
assertFalse(source.contains("new WorkspaceControl("),
"FleetdAssembly.java must not construct WorkspaceControl directly — HerdrRouter is "
+ "its only production factory");
// AgentControl caches paneByTerminal, so the router must be its only production factory.
String source = Files.readString(Path.of("src/main/java/dev/ltms/fleet/Fleetd.java"));
assertFalse(source.contains("new AgentControl("));
assertFalse(source.contains("new WorkspaceControl("));
}
}
@@ -1,63 +0,0 @@
package dev.ltms.fleet;
import dev.ltms.fleet.config.FleetConfig;
import dev.ltms.fleet.mcp.FleetMcp;
import org.junit.jupiter.api.DisplayName;
import org.junit.jupiter.api.Test;
import java.util.Map;
import static org.junit.jupiter.api.Assertions.assertEquals;
import static org.junit.jupiter.api.Assertions.assertNull;
/**
* Pins {@link Fleetd#leadConfigDirSource}'s own wiring of the window lookup into the returned
* {@link FleetMcp.LeadConfigDirSource}, not only the detached {@link Fleetd#leadContextWindowLookup}
* factory it delegates to. Calls the producer directly, with real {@link FleetConfig.Profile}/
* {@link FleetConfig.Leader} fixtures, and asserts on {@code windowFor()} — the companion of
* {@link FleetdLeadConfigDirSourceWiringTest}, which pins the same factory's {@code configDirFor()}.
*/
class FleetdLeadConfigDirSourceWindowWiringTest {
private static FleetConfig.Profile profileWithWindow(String name, Integer autoCompactWindow) {
return new FleetConfig.Profile(name, null, "claude-sonnet-5", null, null, null,
"tab", "fleet", "w #{n}", null, null, null, null, null, null, null,
null, null, true, null, null, null, null, null, autoCompactWindow, null);
}
private static FleetConfig.Leader leadOnProfile(String profile) {
return new FleetConfig.Leader(profile, "lead: primary", 1, "lead:", 10, "claude", "claude-sonnet-5");
}
@Test
@DisplayName("the returned source resolves the lead's REAL configured effective window, not a hardcoded null")
void resolvesTheRealConfiguredWindow() {
Map<String, FleetConfig.Profile> profiles = Map.of("opus", profileWithWindow("opus", 250_000));
Map<String, FleetConfig.Leader> leaders = Map.of("primary", leadOnProfile("opus"));
FleetMcp.LeadConfigDirSource source = Fleetd.leadConfigDirSource(() -> profiles, leaders);
assertEquals(250_000L, source.windowFor().apply("primary"),
"windowFor must delegate to the real leadContextWindowLookup, not a stub that always "
+ "returns null");
}
@Test
@DisplayName("a lead on a profile with no window configured still resolves to null, not a crash")
void leadWithNoWindowConfiguredResolvesToNull() {
Map<String, FleetConfig.Profile> profiles = Map.of("opus", profileWithWindow("opus", null));
Map<String, FleetConfig.Leader> leaders = Map.of("primary", leadOnProfile("opus"));
FleetMcp.LeadConfigDirSource source = Fleetd.leadConfigDirSource(() -> profiles, leaders);
assertNull(source.windowFor().apply("primary"));
}
@Test
@DisplayName("an unrecognised lead name resolves to null, not a thrown exception")
void unrecognisedLeadNameResolvesToNull() {
FleetMcp.LeadConfigDirSource source = Fleetd.leadConfigDirSource(Map::of, Map.of());
assertNull(source.windowFor().apply("ghost-lead"));
}
}
@@ -99,7 +99,7 @@ class FleetdLeadContextLookupTest {
@DisplayName("an unrecognised terminal resolves to UNKNOWN, not a thrown exception")
void unrecognisedTerminalResolvesToUnknown() {
Function<String, LeadContextGauge.Reading> lookup = Fleetd.leadContextLookup(
new LeadContextGauge(), throwingAgentControl(), Map::of, name -> null, name -> null);
new LeadContextGauge(), throwingAgentControl(), Map::of, name -> null);
LeadContextGauge.Reading reading = assertDoesNotThrow(() -> lookup.apply("ghost-terminal"));
@@ -112,7 +112,7 @@ class FleetdLeadContextLookupTest {
void agentsGetThrowingDegradesToUnknown() {
Map<String, String> liveLeadTerminals = Map.of(LEAD_TERMINAL, LEAD_NAME);
Function<String, LeadContextGauge.Reading> lookup = Fleetd.leadContextLookup(
new LeadContextGauge(), throwingAgentControl(), () -> liveLeadTerminals, name -> null, name -> null);
new LeadContextGauge(), throwingAgentControl(), () -> liveLeadTerminals, name -> null);
LeadContextGauge.Reading reading = assertDoesNotThrow(() -> lookup.apply(LEAD_TERMINAL));
@@ -128,7 +128,7 @@ class FleetdLeadContextLookupTest {
AgentControl agents = agentControlStub(SESSION_ID, "claude", "idle");
Function<String, LeadContextGauge.Reading> lookup = Fleetd.leadContextLookup(
new LeadContextGauge(), agents, () -> liveLeadTerminals,
name -> LEAD_NAME.equals(name) ? tmp.toString() : null, name -> null);
name -> LEAD_NAME.equals(name) ? tmp.toString() : null);
LeadContextGauge.Reading reading = lookup.apply(LEAD_TERMINAL);
@@ -146,7 +146,7 @@ class FleetdLeadContextLookupTest {
AgentControl agents = agentControlStub(SESSION_ID, "opencode", "idle");
Function<String, LeadContextGauge.Reading> lookup = Fleetd.leadContextLookup(
new LeadContextGauge(), agents, () -> liveLeadTerminals,
name -> LEAD_NAME.equals(name) ? tmp.toString() : null, name -> null);
name -> LEAD_NAME.equals(name) ? tmp.toString() : null);
LeadContextGauge.Reading reading = lookup.apply(LEAD_TERMINAL);
@@ -1,223 +0,0 @@
package dev.ltms.fleet;
import dev.ltms.fleet.config.ConfigRef;
import dev.ltms.fleet.config.FleetConfig;
import dev.ltms.fleet.guard.SubscriptionGuard;
import dev.ltms.fleet.herdr.FakeHerdr;
import dev.ltms.fleet.herdr.HerdrClient;
import dev.ltms.fleet.lead.LeadContextGauge;
import dev.ltms.fleet.msg.LeadHeartbeatLoop;
import dev.ltms.fleet.msg.ReplyInbox;
import io.javalin.Javalin;
import org.junit.jupiter.api.DisplayName;
import org.junit.jupiter.api.Test;
import org.junit.jupiter.api.io.TempDir;
import java.io.IOException;
import java.lang.reflect.Field;
import java.nio.charset.StandardCharsets;
import java.nio.file.Files;
import java.nio.file.Path;
import java.util.List;
import java.util.Map;
import java.util.concurrent.Executors;
import java.util.concurrent.ScheduledExecutorService;
import java.util.function.LongSupplier;
import static org.junit.jupiter.api.Assertions.assertEquals;
/**
* fleetd #659: {@code FleetdAssembly.java} wires {@code Fleetd.leadContextSource}'s window-lookup
* argument with {@code Fleetd.leadContextWindowLookup(() -> config.get().profiles(), leaders)} —
* but nothing called the real assembled {@link LeadHeartbeatLoop} far enough to prove that
* argument is the one the live heartbeat reads through. Measured: swapping that one call-site
* argument for {@code _ -> null} compiles with 0 errors and leaves the full suite green.
*
* <p>This test drives the REAL {@link LeadHeartbeatLoop} the real {@link
* FleetdAssembly#assembleAndStart} builds, reached through {@link FleetdRuntime#heartbeat()}, and
* reads its private {@code contextSource} field via reflection — the loop exposes no public
* accessor for it, the same reason {@link FleetdLeadConfigDirSourceAssemblyTest} reflects on
* {@code FleetMcp.leadConfigDirs}. The configured profile's {@code autoCompactWindow: 100000}
* resolves a HIGH threshold of {@code 66666} ({@link LeadContextGauge}'s {@code 2/3} fraction) —
* far below the fixed {@code 200000} fallback a lost window argument would silently revert to.
* {@code 90000} live tokens sits between the two: HIGH under the real window, OK under the
* fallback — a property the fallback can never produce by accident.
*/
class FleetdLeadContextSourceWindowAssemblyTest {
private static final String LEAD_NAME = "opus";
private static final String LEAD_TAB = "lead: opus";
private static final String LEAD_PROFILE = "sonnet";
/** {@code FakeHerdr}'s own default {@code agent.list} entry: terminal {@code term_a}, session {@code sess-1111}. */
private static final String LEAD_TERMINAL = "term_a";
private static final String LEAD_SESSION_ID = "sess-1111";
private static final class RecordingResourcePorts implements ResourcePorts {
final FakeHerdr herdr = new FakeHerdr();
final SentinelReplyInbox replyInbox = new SentinelReplyInbox();
@Override
public Map<String, String> environment() {
return Map.of();
}
@Override
public HerdrClient connectHerdr(Path socketPath) {
return herdr;
}
@Override
public Fleetd.AmqpOpener replyInboxOpener() {
return (uri, prefetch) -> replyInbox;
}
@Override
public Fleetd.LeadMailboxOpener leadMailboxOpener() {
return (uri, selfCoordId, prefetch) -> {
throw new UnsupportedOperationException(
"leadMailboxOpener must not be called — no coordinator: block is configured");
};
}
@Override
public LongSupplier nanoClock() {
return System::nanoTime;
}
@Override
public LongSupplier wallClockNanos() {
return System::nanoTime;
}
@Override
public ScheduledExecutorService newScheduler(String purpose) {
return Executors.newSingleThreadScheduledExecutor();
}
@Override
public void addShutdownHook(Runnable hook) {
}
@Override
public void startHttp(Javalin app, String host, int port) {
}
@Override
public Runnable herdrPollWait() {
// Never invoked: this test's FakeHerdr answers immediately, so awaitHerdr never polls.
return () -> {
throw new UnsupportedOperationException("herdrPollWait must not be called — herdr is healthy");
};
}
}
private static final class SentinelReplyInbox implements ReplyInbox, AutoCloseable {
@Override
public void own(String target) {
}
@Override
public void release(String target) {
}
@Override
public void publish(String target, String msgId, String content) {
}
@Override
public List<InboxMessage> peek(String target) {
return List.of();
}
@Override
public boolean ack(String target, String msgId) {
return false;
}
@Override
public void close() {
}
}
private static String usageLine(long tokens) {
return "{\"type\":\"assistant\",\"message\":{\"role\":\"assistant\",\"usage\":{"
+ "\"input_tokens\":" + tokens + ",\"cache_read_input_tokens\":0,\"cache_creation_input_tokens\":0}}}";
}
/** Lays out {@code <configDir>/projects/<anySlug>/<sessionId>.jsonl} carrying one usage record. */
private static void writeTranscript(Path configDir, String sessionId, long tokens) throws IOException {
Path projectDir = configDir.resolve("projects").resolve("some-project-slug");
Files.createDirectories(projectDir);
Files.writeString(projectDir.resolve(sessionId + ".jsonl"), usageLine(tokens) + "\n", StandardCharsets.UTF_8);
}
private static FleetConfig writeConfig(Path dir, String configDir) throws Exception {
Path f = dir.resolve("fleetd.yaml");
Files.writeString(f, """
bind:
host: 127.0.0.1
port: 8765
idleSleepGuard:
enabled: false
broker:
uri: "amqp://fake-test-broker/vh"
leadHeartbeat:
idleAfterSeconds: 600
backoffMs: 15000
quietNudgeCap: 5
fleet:
leaders:
%s:
tab: "%s"
profile: %s
profiles:
%s:
subscription: true
argv: ["ccs", "sonnet"]
configDir: "%s"
autoCompactWindow: 100000
""".formatted(LEAD_NAME, LEAD_TAB, LEAD_PROFILE, LEAD_PROFILE, configDir));
return FleetConfig.load(f);
}
@SuppressWarnings("unchecked")
private static LeadHeartbeatLoop.LeadContextSource contextSourceOf(LeadHeartbeatLoop heartbeat) throws Exception {
Field field = LeadHeartbeatLoop.class.getDeclaredField("contextSource");
field.setAccessible(true);
return (LeadHeartbeatLoop.LeadContextSource) field.get(heartbeat);
}
@Test
@DisplayName("[BEHAVIOURAL] the real assembled heartbeat loop resolves HIGH against the lead's "
+ "REAL configured window, not the fixed 200000 fallback a lost window argument reverts to")
void assembledHeartbeatContextSourceResolvesTheRealConfiguredWindow(@TempDir Path dir) throws Exception {
FleetConfig cfg = writeConfig(dir, dir.toString());
writeTranscript(dir, LEAD_SESSION_ID, 90_000);
ConfigRef config = new ConfigRef(dir.resolve("fleetd.yaml"), cfg);
SubscriptionGuard guard = new SubscriptionGuard(cfg.guard().hostSet());
RecordingResourcePorts ports = new RecordingResourcePorts();
// Label FakeHerdr's own default pane's tab (term_a / w2:p7 / w2:t7, already carrying a live
// agent on session sess-1111) to match fleet.leaders.opus.tab exactly, so LeadTabScanner
// recognises it as the live "opus" lead without a second auto-launched pane.
ports.herdr.withTab("w2", "w2:t7", LEAD_TAB).agentSessionId(LEAD_SESSION_ID);
FleetdRuntime runtime = FleetdAssembly.assembleAndStart(new AssemblyInputs(cfg, config, guard), ports);
try {
LeadHeartbeatLoop.LeadContextSource source = contextSourceOf(runtime.heartbeat());
LeadContextGauge.Reading reading = source.readingFor().apply(LEAD_TERMINAL);
assertEquals(LeadContextGauge.State.HIGH, reading.state(),
"profiles." + LEAD_PROFILE + ".autoCompactWindow: 100000 resolves a HIGH threshold "
+ "of 66666 tokens — 90000 live tokens must read HIGH against it. Mutating "
+ "FleetdAssembly's window-lookup argument to `_ -> null` falls back to the "
+ "fixed 200000 threshold, under which 90000 reads OK instead: " + reading);
} finally {
// Surefire runs the whole suite in one JVM fork (fleetd/pom.xml sets no forkCount /
// reuseForks), so the scheduler/loops this assembly starts must be torn down here, on the
// failure path too — hence try/finally rather than a bare statement at the end.
runtime.close();
}
}
}
@@ -87,8 +87,7 @@ class FleetdLeadContextSourceWiringTest {
Map<String, String> liveLeadTerminals = Map.of(LEAD_TERMINAL, LEAD_NAME);
LeadHeartbeatLoop.LeadContextSource source = Fleetd.leadContextSource(new LeadContextGauge(),
agentControlStub(), () -> liveLeadTerminals,
name -> LEAD_NAME.equals(name) ? tmp.toString() : null, name -> null);
agentControlStub(), () -> liveLeadTerminals, name -> LEAD_NAME.equals(name) ? tmp.toString() : null);
LeadContextGauge.Reading reading = source.readingFor().apply(LEAD_TERMINAL);
@@ -102,7 +101,7 @@ class FleetdLeadContextSourceWiringTest {
@DisplayName("an unrecognised lead terminal resolves to UNKNOWN, not a thrown exception")
void unrecognisedTerminalResolvesToUnknown() {
LeadHeartbeatLoop.LeadContextSource source = Fleetd.leadContextSource(new LeadContextGauge(),
agentControlStub(), Map::of, name -> null, name -> null);
agentControlStub(), Map::of, name -> null);
assertEquals(LeadContextGauge.State.UNKNOWN, source.readingFor().apply("ghost-terminal").state());
}
@@ -765,92 +765,6 @@ class FleetConfigTest {
"a label that collides with a convention nobody reads is not a problem");
}
// ── validatePanePlacementAgainstLeadTabs ────────────────────────────────────────────────────
/**
* The hazard this guard closes: a pane-placed member lands inside the focused tab rather than
* its own, so it can land inside a lead's labelled tab and be read back as that lead.
*/
@Test
void aPanePlacedProfileWithALeadTabRefusesToStart(@TempDir Path dir) throws Exception {
Path f = dir.resolve("pane-hazard.yaml");
Files.writeString(f, """
bind:
port: 8080
profiles:
gx10:
placement: pane
fleet:
leaders:
opus:
tab: "lead: opus"
""");
FleetConfig cfg = FleetConfig.load(f);
IllegalStateException e = assertThrows(IllegalStateException.class,
cfg::validatePanePlacementAgainstLeadTabs);
assertTrue(e.getMessage().contains("gx10"), "the message must name the offending profile");
}
@Test
void aPanePlacedProfileWithNoLeadTabIsAllowed(@TempDir Path dir) throws Exception {
Path f = dir.resolve("pane-no-tab.yaml");
Files.writeString(f, """
bind:
port: 8080
profiles:
gx10:
placement: pane
fleet:
leaders:
opus:
profile: gx10
""");
assertDoesNotThrow(() -> FleetConfig.load(f).validatePanePlacementAgainstLeadTabs(),
"a leader with no tab feeds nothing into the scanner, so pane placement is safe");
}
@Test
void aTabPlacedProfileWithALeadTabIsAllowed(@TempDir Path dir) throws Exception {
Path f = dir.resolve("tab-safe.yaml");
Files.writeString(f, """
bind:
port: 8080
profiles:
gx10:
placement: tab
fleet:
leaders:
opus:
tab: "lead: opus"
""");
assertDoesNotThrow(() -> FleetConfig.load(f).validatePanePlacementAgainstLeadTabs(),
"a member in its own tab cannot land inside a lead's tab");
}
/** Proves the reflective sweep behind {@code validateAll} really reaches this validator. */
@Test
void validateAllAlsoRefusesPanePlacementAgainstALeadTab(@TempDir Path dir) throws Exception {
Path f = dir.resolve("pane-hazard-sweep.yaml");
Files.writeString(f, """
bind:
port: 8080
profiles:
gx10:
placement: pane
fleet:
leaders:
opus:
tab: "lead: opus"
""");
FleetConfig cfg = FleetConfig.load(f);
IllegalStateException e = assertThrows(IllegalStateException.class, cfg::validateAll);
assertTrue(e.getMessage().contains("gx10"), "the message must name the offending profile");
}
// ── CB-530/CB-579: the leaders registry ─────────────────────────────────────────────────────
@Test
@@ -42,13 +42,12 @@ import static org.junit.jupiter.api.Assertions.assertTrue;
* has right now. This is the proof that a future, real seventh validator on {@link
* FleetConfig} would be swept automatically, without needing to add a real (unwanted)
* seventh validator just to exercise the claim.</li>
* <li>{@link #validateAllReachesEveryOneOfTodaysRealValidators()} proves {@link
* <li>{@link #validateAllReachesEveryOneOfTodaysSixValidators()} proves {@link
* FleetConfig#validateAll()} itself is wired to that same generic mechanism and genuinely
* reaches seven of today's eight real validators — reusing the exact minimal failing
* reaches each of today's six real validators — reusing the exact minimal failing
* configurations {@code FleetConfigTest} already established for each one directly, so a
* single call to {@code validateAll()} is shown to reproduce every one of those seven
* failures. The eighth, {@link FleetConfig#validateLeadRollover()}, has no case here yet —
* a pre-existing gap tracked as fleetd #668.</li>
* single call to {@code validateAll()} is shown to reproduce every one of those six
* failures.</li>
* </ol>
*
* <p>Together with the direct-{@code Fleetd.main}-invocation tests in {@code
@@ -63,8 +62,8 @@ import static org.junit.jupiter.api.Assertions.assertTrue;
* the generic sweep — claim 1 proves {@link FleetConfig#invokeAllValidators} is generic, and claim
* 2 proves {@code validateAll()} reaches today's six, and a hardcoded list satisfies both. So the
* reflective sweep is a convenience, not the guarantee. The guarantee is {@link
* #fleetConfigDeclaresExactlyTheseValidatorsToday()}: it fails the moment any validator is added
* or removed, which forces whoever changes the set to look at this file.
* #fleetConfigDeclaresExactlyTheseSixValidatorsToday()}: it fails the moment a seventh validator
* is declared, which forces whoever adds it to look at this file.
*/
class FleetConfigValidateAllTest {
@@ -210,19 +209,18 @@ class FleetConfigValidateAllTest {
+ "name) must all be skipped");
}
// ── Claim 2: FleetConfig.validateAll() is wired to that mechanism and reaches seven of eight today ──
// ── Claim 2: FleetConfig.validateAll() is wired to that mechanism and reaches all six today ──
/**
* Reflectively enumerates {@link FleetConfig}'s own public, no-arg, void {@code validateXxx()}
* methods (excluding {@code validateAll} itself) — the exact same filter {@link
* FleetConfig#invokeAllValidators} applies. This is not the mechanism proof (that is claim 1,
* above, on an unrelated class) — it is a visible denominator: today there are eight, named in
* the {@code Set.of} below, so a reader adding or removing one sees this assertion name the new
* count rather than a silent pass at the old one. The count lives only in that set, not in this
* method's name, so the two cannot drift apart.
* above, on an unrelated class) — it is a visible denominator: today there are six, named
* here, so a reader adding a seventh sees this assertion name the new count rather than a
* silent pass at the old one.
*/
@Test
void fleetConfigDeclaresExactlyTheseValidatorsToday() {
void fleetConfigDeclaresExactlyTheseSixValidatorsToday() {
Set<String> names = new TreeSet<>();
for (Method m : FleetConfig.class.getMethods()) {
if (java.lang.reflect.Modifier.isPublic(m.getModifiers())
@@ -235,8 +233,7 @@ class FleetConfigValidateAllTest {
}
assertEquals(new TreeSet<>(Set.of("validateAuthExposure", "validateLeadTabPrefixes",
"validateSubscriptionProfiles", "validateCharters", "validateMembers",
"validateModels", "validateLeadRollover", "validatePanePlacementAgainstLeadTabs")),
names,
"validateModels", "validateLeadRollover")), names,
"FleetConfig's public validate*() methods changed. Do TWO things, in this "
+ "order. First confirm validateAll() still delegates to "
+ "invokeAllValidators(this) — a hardcoded list there passes every other "
@@ -265,17 +262,14 @@ class FleetConfigValidateAllTest {
}
/**
* The heart of claim 2: for seven of today's eight real validators, a minimal file that fails
* The heart of claim 2: for each of today's six real validators, a minimal file that fails
* ONLY that one — the exact fixtures {@code FleetConfigTest} uses to test each validator
* directly — must also fail through {@link FleetConfig#validateAll()}. If a future edit to
* {@code validateAll()} silently dropped one of these seven from the sweep (e.g. a typo'd name
* filter), exactly one of them would start passing when it must not.
*
* <p>The eighth, {@link FleetConfig#validateLeadRollover()}, has no case here — a pre-existing
* gap tracked as fleetd #668, not fixed by this change.
* {@code validateAll()} silently dropped one validator from the sweep (e.g. a typo'd name
* filter), exactly one of these six would start passing when it must not.
*/
@Test
void validateAllReachesEveryOneOfTodaysRealValidators(@TempDir Path dir) throws Exception {
void validateAllReachesEveryOneOfTodaysSixValidators(@TempDir Path dir) throws Exception {
// validateAuthExposure: a non-loopback bind without token mode.
assertValidateAllRefuses(dir, "auth-exposure.yaml", """
bind:
@@ -348,20 +342,6 @@ class FleetConfigValidateAllTest {
allow:
- model: claude-sonnet-5
""", "rogue");
// validatePanePlacementAgainstLeadTabs: a pane-placed profile while a lead names a tab.
assertValidateAllRefuses(dir, "pane-placement.yaml", """
bind:
host: 127.0.0.1
port: 8765
profiles:
gx10:
placement: pane
fleet:
leaders:
opus:
tab: "lead: opus"
""", "gx10");
}
private static void assertValidateAllRefuses(Path dir, String fileName, String yaml,
@@ -51,7 +51,6 @@ public final class FakeHerdr implements HerdrClient {
private boolean noPanes = false;
private volatile String agentStatus = "idle"; // steady-state agent.get status
private volatile String agentType = "claude"; // detected agent kind on agent.get; null = undetected
private volatile String agentSessionId = null; // agent_session.value on agent.get; null = omitted
private volatile String readText = "worker transcript tail"; // canned agent.read output
private int pinnedStarts = 0; // how many upcoming agent.start calls report a fixed pane
private String pinnedStartTerminal;
@@ -174,16 +173,6 @@ public final class FakeHerdr implements HerdrClient {
return this;
}
/**
* Set the {@code agent_session.value} that {@code agent.get} reports for {@code term_a} — the
* default omits the field entirely (herdr not yet having resolved one), matching the real
* daemon's own "not resolved yet" shape.
*/
public FakeHerdr agentSessionId(String sessionId) {
this.agentSessionId = sessionId;
return this;
}
/**
* Make {@code agent.get} succeed normally for its first {@code okCalls} invocations, then fail
* every call after that with {@code code} — fleetd #176 fix 1's "backend exited mid-wait"
@@ -322,12 +311,10 @@ public final class FakeHerdr implements HerdrClient {
}
}
String agentField = agentType == null ? "null" : "\"" + agentType + "\"";
String sessionField = agentSessionId == null ? ""
: ",\"agent_session\":{\"kind\":\"id\",\"value\":\"" + agentSessionId + "\"}";
yield mapper.readTree(("""
{"type":"agent_info","agent":{"terminal_id":"term_a","agent":%s,
"agent_status":"%s","workspace_id":"w2","tab_id":"w2:t7","pane_id":"w2:p7"%s}}""")
.formatted(agentField, agentStatus, sessionField));
"agent_status":"%s","workspace_id":"w2","tab_id":"w2:t7","pane_id":"w2:p7"}}""")
.formatted(agentField, agentStatus));
}
case "agent.read" -> mapper.readTree(mapper.writeValueAsString(
java.util.Map.of("type", "agent_read", "read", java.util.Map.of("text", readText))));
@@ -232,9 +232,8 @@ class LeadTabScannerTest {
@Test
void everyPaneInALeadTabResolvesAsThatLead() {
// A human may split their own lead tab. Both panes are theirs, so both are that lead.
// A pane-placed member landing here instead is refused at startup by
// FleetConfig.validatePanePlacementAgainstLeadTabs, not by this scanner.
// A human may split their own lead tab. Both panes are theirs, so both are that lead —
// nothing fleetd placed can land here (see the worker-space test above).
TopologyHerdr herdr = twoLeads().pane("w1:p1b", "w1:t1", "term_opus_split");
assertEquals("opus-5.0",
@@ -81,24 +81,11 @@ class LeadContextGaugeHighThresholdTest {
assertEquals(LeadContextGauge.State.HIGH,
atGauge.read(atConfigDir, SESSION_ID, "claude", null).state(),
"the fixed default must still be 200,000 when no window is resolvable");
}
@Test
@DisplayName("a second read with a different window, within the TTL, reports against its own window, not the first call's cached state")
void aSecondReadWithADifferentWindowWithinTheTtlReportsAgainstItsOwnWindow(@TempDir Path tmp) throws IOException {
String configDir = writeTranscript(tmp, SESSION_ID, 90_000);
long[] now = {0L};
LeadContextGauge gauge = new LeadContextGauge(() -> now[0], 5_000);
LeadContextGauge.Reading first = gauge.read(configDir, SESSION_ID, "claude", 100_000L);
assertEquals(LeadContextGauge.State.HIGH, first.state(),
"90,000 tokens against a 100,000 window is HIGH");
now[0] += 1_000; // stays inside the 5,000ms TTL — the cache key must still vary with the window
LeadContextGauge.Reading second = gauge.read(configDir, SESSION_ID, "claude", 1_000_000L);
assertEquals(LeadContextGauge.State.OK, second.state(),
"90,000 tokens against a 1,000,000 window must report OK regardless of the previous call's "
+ "window, even while that call's cache entry is still within its TTL");
String legacyConfigDir = writeTranscript(tmp.resolve("legacy"), SESSION_ID, 200_000);
LeadContextGauge legacyGauge = new LeadContextGauge();
assertEquals(LeadContextGauge.State.HIGH,
legacyGauge.read(legacyConfigDir, SESSION_ID, "claude").state(),
"the 3-arg read() (no window argument at all) must behave exactly like passing a null window");
}
}
@@ -85,7 +85,7 @@ class LeadContextGaugeTest {
usageLine(40_000, 5_000, 3_000)); // last record: 48,000
LeadContextGauge gauge = new LeadContextGauge();
LeadContextGauge.Reading first = gauge.read(configDir, SESSION_ID, "claude", null);
LeadContextGauge.Reading first = gauge.read(configDir, SESSION_ID, "claude");
assertEquals(48_000L, first.tokens(), "must total input+cache_read+cache_creation of the LAST usage record");
assertEquals(LeadContextGauge.State.OK, first.state());
@@ -93,7 +93,7 @@ class LeadContextGaugeTest {
// must change with it, not stay pinned to the first fixture's total.
String otherSession = "22222222-2222-2222-2222-222222222222";
writeTranscript(tmp, otherSession, usageLine(100_000, 50_000, 50_000)); // last record: 200,000
LeadContextGauge.Reading second = gauge.read(configDir, otherSession, "claude", null);
LeadContextGauge.Reading second = gauge.read(configDir, otherSession, "claude");
assertEquals(200_000L, second.tokens());
assertTrue(second.tokens() != first.tokens(), "changing N in the fixture must change the reported number");
}
@@ -111,12 +111,12 @@ class LeadContextGaugeTest {
compactionLine(),
usageLine(3_000, 0, 0));
LeadContextGauge gauge = new LeadContextGauge();
LeadContextGauge.Reading twoCompactions = gauge.read(tmp.toString(), sessionTwoCompactions, "claude", null);
LeadContextGauge.Reading twoCompactions = gauge.read(tmp.toString(), sessionTwoCompactions, "claude");
assertEquals(2, twoCompactions.compactions());
String sessionZeroCompactions = "44444444-4444-4444-4444-444444444444";
writeTranscript(tmp, sessionZeroCompactions, usageLine(3_000, 0, 0));
LeadContextGauge.Reading zeroCompactions = gauge.read(tmp.toString(), sessionZeroCompactions, "claude", null);
LeadContextGauge.Reading zeroCompactions = gauge.read(tmp.toString(), sessionZeroCompactions, "claude");
assertEquals(0, zeroCompactions.compactions(), "changing K in the fixture must change the reported count");
}
@@ -126,7 +126,7 @@ class LeadContextGaugeTest {
@DisplayName("a missing transcript file reports UNKNOWN with no token number")
void missingFileIsUnknown(@TempDir Path tmp) {
LeadContextGauge gauge = new LeadContextGauge();
LeadContextGauge.Reading reading = gauge.read(tmp.toString(), SESSION_ID, "claude", null);
LeadContextGauge.Reading reading = gauge.read(tmp.toString(), SESSION_ID, "claude");
assertEquals(LeadContextGauge.State.UNKNOWN, reading.state());
assertNull(reading.tokens());
}
@@ -148,7 +148,7 @@ class LeadContextGaugeTest {
assumeFalse(Files.isReadable(file),
"runs as root (CI container): the read bit does not stop root, so this case cannot be set up here");
LeadContextGauge gauge = new LeadContextGauge();
LeadContextGauge.Reading reading = gauge.read(configDir, SESSION_ID, "claude", null);
LeadContextGauge.Reading reading = gauge.read(configDir, SESSION_ID, "claude");
assertEquals(LeadContextGauge.State.UNKNOWN, reading.state());
assertNull(reading.tokens());
} finally {
@@ -171,7 +171,7 @@ class LeadContextGaugeTest {
Files.writeString(file, lastCompleteLine + "\n" + tornLine, StandardCharsets.UTF_8);
LeadContextGauge gauge = new LeadContextGauge();
LeadContextGauge.Reading reading = gauge.read(tmp.toString(), SESSION_ID, "claude", null);
LeadContextGauge.Reading reading = gauge.read(tmp.toString(), SESSION_ID, "claude");
assertEquals(LeadContextGauge.State.OK, reading.state(),
"a torn final line must not turn a good earlier reading into UNKNOWN");
assertEquals(6_000L, reading.tokens(),
@@ -187,7 +187,7 @@ class LeadContextGaugeTest {
"{this is not json at all",
"neither is this{{{");
LeadContextGauge gauge = new LeadContextGauge();
LeadContextGauge.Reading reading = gauge.read(configDir, SESSION_ID, "claude", null);
LeadContextGauge.Reading reading = gauge.read(configDir, SESSION_ID, "claude");
assertEquals(LeadContextGauge.State.UNKNOWN, reading.state(),
"every line unparseable is the real format-change signal and must still report UNKNOWN");
assertNull(reading.tokens());
@@ -224,12 +224,12 @@ class LeadContextGaugeTest {
AtomicLong now = new AtomicLong(0);
LeadContextGauge gauge = new LeadContextGauge(now::get, 5_000);
gauge.read(configDir, SESSION_ID, "claude", null);
gauge.read(configDir, SESSION_ID, "claude", null); // still inside the TTL window
gauge.read(configDir, SESSION_ID, "claude");
gauge.read(configDir, SESSION_ID, "claude"); // still inside the TTL window
assertEquals(1, gauge.diskReadCount(), "two reads inside the TTL must touch disk once");
now.set(6_000); // past the TTL
gauge.read(configDir, SESSION_ID, "claude", null);
gauge.read(configDir, SESSION_ID, "claude");
assertEquals(2, gauge.diskReadCount(), "a read past the TTL must touch disk again");
}
@@ -241,8 +241,8 @@ class LeadContextGaugeTest {
String configDir = writeTranscript(tmp, SESSION_ID, usageLine(1_000, 0, 0));
LeadContextGauge gauge = new LeadContextGauge();
assertEquals(LeadContextGauge.State.UNKNOWN, gauge.read(configDir, SESSION_ID, "opencode", null).state());
assertEquals(LeadContextGauge.State.UNKNOWN, gauge.read(configDir, SESSION_ID, null, null).state());
assertEquals(LeadContextGauge.State.UNKNOWN, gauge.read(configDir, null, "claude", null).state());
assertEquals(LeadContextGauge.State.UNKNOWN, gauge.read(configDir, SESSION_ID, "opencode").state());
assertEquals(LeadContextGauge.State.UNKNOWN, gauge.read(configDir, SESSION_ID, null).state());
assertEquals(LeadContextGauge.State.UNKNOWN, gauge.read(configDir, null, "claude").state());
}
}
@@ -86,7 +86,7 @@ class FleetMcpLeadContextGaugeWiringTest {
LeadContextGauge contextGauge = new LeadContextGauge();
AtomicReference<String> configuredDir = new AtomicReference<>(dirA.toString());
FleetMcp.LeadConfigDirSource source = new FleetMcp.LeadConfigDirSource(name ->
LEAD_NAME.equals(name) ? configuredDir.get() : null, _ -> null);
LEAD_NAME.equals(name) ? configuredDir.get() : null);
String firstRead = textOf(listFleet(herdr, sessions, contextGauge, source));
assertTrue(firstRead.contains("\"tokens\":11000"),
+22 -139
View File
@@ -2,10 +2,6 @@
#
# The one auditable way to edit the live fleetd.yaml.
#
# `--set` uses yq and rewrites the whole YAML document in yq's output style. Use `--from` for a
# candidate whose comment alignment or other formatting carries meaning: it copies that file
# verbatim while keeping this script's backup, parse check, atomic install, and verdict read-back.
#
# fleetd ticket #635 — why this exists at all: fleetd.yaml is gitignored and holds the live
# fleet's settings. A bad raw edit reaches a daemon that is already serving, so a direct `Edit`
# on it is refused by policy. This script is the allow-listed alternative, and it is not just
@@ -50,11 +46,6 @@
# scripts/config-edit.sh --dry-run --set <yq-path>=<value>
# scripts/config-edit.sh --restore
#
# `--set` rewrites the whole file in yq's output style, not only the requested keys. The script
# warns before installation when the candidate changes more lines than its number of --set pairs.
# Use `--from <candidate.yaml>` when comment alignment or other formatting is meaningful: --from
# copies the candidate verbatim, with no yq round-trip.
#
# `--set .a.b=` (an empty value — a forgotten typo) is REFUSED, not accepted as "clear the
# field": a null value falls back to its default rather than erroring, which is silent, not
# safe. To clear a key on purpose, write a literal null: `--set .a.b=null`. Every other value
@@ -163,10 +154,18 @@ done
# story: a YAML block scalar (`|`, `|-`, `>`, `>-`, ...) puts the VALUE on the lines that follow
# the key, each indented deeper than it. The key-name match above only ever sees the key line
# itself, so those continuation lines used to flow straight through unredacted while the key line
# right above them printed a reassuring "<redacted>". The redactor maps masked continuation lines
# from each complete file before it reads the diff. It then masks a printed line when that file
# line is inside a masked key's value. This covers block-scalar bodies even when the key line is
# outside the printed hunk, and it keeps blank lines inside the value masked.
# right above them printed a reassuring "<redacted>" — an incomplete redactor that looks complete
# is worse than one that visibly does nothing, because it stops a reviewer from looking further.
# The fix is structural, not another name to match: once a key line is masked, every following
# line indented STRICTLY DEEPER than that key is masked too, by indentation alone, until the
# indentation returns to the key's own level or shallower. This needs no knowledge of the key's
# name, so it covers a block scalar under any masked key — but ONLY while that key's own line is
# itself inside the hunk being printed. `diff -u` prints just three lines of context, so a block
# scalar's body often reaches this function with its key line left out; there is then nothing to
# anchor to, `masked` is never set, and the body prints in full. A blank line inside a block
# scalar loses the anchor the same way, because a blank diff line measures as indent 0. Both are
# measured and filed as fleetd #639 — do not read this paragraph as a guarantee that a masked
# key's value can never be printed.
#
# `redact` is always fed `diff -u` output, and every line of a unified diff starts with exactly
# one of ' ', '+', '-' (the three body markers; '@'/'-'/'+' for the three header-line kinds too).
@@ -175,102 +174,37 @@ done
# column shallower than it really is, and either wrongly escapes a continuation mask or wrongly
# ends one early. Tabs are out of scope: YAML forbids them for indentation, and this is a bounded
# fix, not a YAML parser.
map_masked_lines() {
local file="$1" side="$2" line content indent lead key line_number=0
local masked=0 masked_indent=0
case "$side" in
old) OLD_MASKED_LINES=() ;;
new) NEW_MASKED_LINES=() ;;
*) die "internal error: unknown redaction map side $side" ;;
esac
while IFS= read -r line || [ -n "$line" ]; do
line_number=$((line_number + 1))
content="$line"
indent=0
while [ "${content:$indent:1}" = " " ]; do indent=$((indent + 1)); done
if [ "$masked" = 1 ]; then
if [ -z "${content// /}" ] || [ "$indent" -gt "$masked_indent" ]; then
case "$side" in
old) OLD_MASKED_LINES[$line_number]=1 ;;
new) NEW_MASKED_LINES[$line_number]=1 ;;
esac
continue
fi
masked=0
fi
if [[ "$content" =~ ^([[:space:]]*)([A-Za-z0-9_.-]+:) ]]; then
lead="${BASH_REMATCH[1]}"
key="${BASH_REMATCH[2]}"
if [[ "$key" =~ (TOKEN|SECRET|PASSWORD|PASSWD|PASSPHRASE|CREDENTIAL|URI|_KEY) ]]; then
masked=1
masked_indent="$indent"
fi
fi
done < "$file"
}
redact() {
local old_file="$1" new_file="$2"
local line prefix content indent lead key old_line=0 new_line=0 in_hunk=0
local old_masked new_masked saved_nocasematch=0
local line prefix content indent lead key
local masked=0 masked_indent=0 saved_nocasematch=0
shopt -q nocasematch && saved_nocasematch=1
shopt -s nocasematch
map_masked_lines "$old_file" old
map_masked_lines "$new_file" new
while IFS= read -r line || [ -n "$line" ]; do
if [[ "$line" =~ ^@@\ -([0-9]+)(,([0-9]+))?\ \+([0-9]+)(,([0-9]+))?\ @@ ]]; then
old_line="${BASH_REMATCH[1]}"
new_line="${BASH_REMATCH[4]}"
in_hunk=1
printf '%s\n' "$line"
continue
fi
sed -E 's#://[^@]*@#://<redacted>@#g' | while IFS= read -r line || [ -n "$line" ]; do
case "$line" in
[\ +-]*) prefix="${line:0:1}"; content="${line:1}" ;;
*) prefix=""; content="$line" ;;
esac
old_masked=0
new_masked=0
if [ "$in_hunk" = 1 ]; then
case "$prefix" in
' ')
[ "${OLD_MASKED_LINES[$old_line]:-}" = 1 ] && old_masked=1
[ "${NEW_MASKED_LINES[$new_line]:-}" = 1 ] && new_masked=1
old_line=$((old_line + 1)); new_line=$((new_line + 1)) ;;
-)
[ "${OLD_MASKED_LINES[$old_line]:-}" = 1 ] && old_masked=1
old_line=$((old_line + 1)) ;;
+)
[ "${NEW_MASKED_LINES[$new_line]:-}" = 1 ] && new_masked=1
new_line=$((new_line + 1)) ;;
esac
fi
indent=0
while [ "${content:$indent:1}" = " " ]; do indent=$((indent + 1)); done
if [ "$old_masked" = 1 ] || [ "$new_masked" = 1 ]; then
if [ "$masked" = 1 ] && [ "$indent" -gt "$masked_indent" ]; then
printf '%s%*s<redacted>\n' "$prefix" "$indent" ""
continue
fi
masked=0
if [[ "$content" =~ ^([[:space:]]*)([A-Za-z0-9_.-]+:) ]]; then
lead="${BASH_REMATCH[1]}"
key="${BASH_REMATCH[2]}"
if [[ "$key" =~ (TOKEN|SECRET|PASSWORD|PASSWD|PASSPHRASE|CREDENTIAL|URI|_KEY) ]]; then
printf '%s%s%s <redacted>\n' "$prefix" "$lead" "$key"
masked=1
masked_indent="$indent"
continue
fi
fi
printf '%s\n' "$line" | sed -E 's#://[^@]*@#://<redacted>@#g'
printf '%s\n' "$line"
done
[ "$saved_nocasematch" = 1 ] || shopt -u nocasematch
}
@@ -529,50 +463,6 @@ parse_check() {
yq eval '.' "$1" >/dev/null 2>&1
}
# Count logical changed lines in a unified diff. A replacement counts once, while an added or
# deleted line also counts once. One changed `--set` value normally produces one changed line.
changed_line_count() {
local before="$1" after="$2" line count=0 old_count=0 new_count=0
while IFS= read -r line || [ -n "$line" ]; do
case "$line" in
---\ *|+++\ *|@@\ *)
if [ "$old_count" -gt "$new_count" ]; then
count=$((count + old_count))
else
count=$((count + new_count))
fi
old_count=0
new_count=0
;;
-*) old_count=$((old_count + 1)) ;;
+*) new_count=$((new_count + 1)) ;;
*)
if [ "$old_count" -gt "$new_count" ]; then
count=$((count + old_count))
else
count=$((count + new_count))
fi
old_count=0
new_count=0
;;
esac
done < <(diff -u "$before" "$after" || true)
if [ "$old_count" -gt "$new_count" ]; then
count=$((count + old_count))
else
count=$((count + new_count))
fi
printf '%s' "$count"
}
warn_set_reformat() {
local before="$1" after="$2" changed
changed="$(changed_line_count "$before" "$after")"
if [ "$changed" -gt "${#SETS[@]}" ]; then
warn "--set changed $changed candidate lines for ${#SETS[@]} pair(s); yq reformatted the whole file. Use --from for meaningful comment alignment or formatting."
fi
}
install_candidate() {
local cand="$1" live="$2"
mv -f "$cand" "$live"
@@ -734,14 +624,10 @@ run_edit() {
fi
ok "candidate parses"
if [ "$MODE" = "set" ]; then
warn_set_reformat "$backup" "$cand"
fi
apply_mode "$cand" "$orig_mode"
say "change (redacted)"
diff -u "$backup" "$cand" | redact "$backup" "$cand" || true
diff -u "$backup" "$cand" | redact || true
say "install"
install_candidate "$cand" "$CONFIG" \
@@ -769,11 +655,8 @@ dry_run_diff() {
rm -f "$cand"; CAND=""
die "candidate does not parse as valid YAML — this was a --dry-run, nothing would have been installed either"
fi
if [ "$MODE" = "set" ]; then
warn_set_reformat "$CONFIG" "$cand"
fi
say "dry run — diff (redacted), nothing installed"
diff -u "$CONFIG" "$cand" | redact "$CONFIG" "$cand" || true
diff -u "$CONFIG" "$cand" | redact || true
rm -f "$cand"; CAND=""
return 0
}
+63 -87
View File
@@ -71,21 +71,20 @@ set -euo pipefail
REPO="$(cd "$(dirname "${BASH_SOURCE[0]}")/.." && pwd)"
MODULE="$REPO/fleetd"
# fleetd #664: the runtime path and Maven's output path are no longer the same file. Maven's
# shade plugin (finalName=fleetd) always lands a fresh build at target/fleetd.jar — that is
# Maven's own output directory and this script does not change it — but the daemon is launched
# from $JAR instead, outside target/ entirely. That split is the whole fix: neither `mvn install`
# nor `mvn clean` can ever reach the file a running daemon holds open, because that file no
# longer lives under target/ at all. See swap_if_built/swap_staged_jar below for the one `mv`
# that moves a build from one path to the other, and only after the old daemon is confirmed gone.
BUILD_JAR="$MODULE/target/fleetd.jar"
JAR="$MODULE/run/fleetd.jar"
JAR="$MODULE/target/fleetd.jar"
# fleetd #493: never build into the path a running process holds. The build writes here first
# (Maven's shade plugin has finalName=fleetd, so `clean install` still lands its output at
# target/fleetd.jar — that part is unchanged and out of this script's control), but this script
# now moves it out to JAR_STAGED immediately, and only swaps it back to JAR (a plain `mv`, so a
# rename, never a byte-by-byte overwrite) after the OLD daemon has been confirmed exited. See
# stage_built_jar/swap_staged_jar below.
JAR_STAGED="$MODULE/target/fleetd-new.jar"
OUT="$MODULE/fleetd.out"
# Matches BOTH the absolute form and the relative `java -jar run/fleetd.jar` a hand-start
# Matches BOTH the absolute form and the relative `java -jar target/fleetd.jar` a hand-start
# produces from inside fleetd/. Anchoring on the absolute path alone was a real bug: the daemon
# restarted correctly and the script still reported "no process appeared", because it launched with
# a relative path and then looked for an absolute one.
PATTERN='run/fleetd.jar'
PATTERN='target/fleetd.jar'
HEALTH='http://127.0.0.1:8765/healthz'
STOP_WAIT=30 # seconds to wait for a clean exit before reporting failure
HEALTH_WAIT=60 # seconds to wait for /healthz to answer after start — fleetd #603: also the pid-
@@ -170,10 +169,10 @@ hash256() {
fi
}
# Reports the hash of $JAR by default, or of whatever path is passed — used to report the BUILT
# jar at $BUILD_JAR (before it has been swapped in) without ever changing what a bare `jar_id`
# (no args) means: the live path, $JAR. --check and the final "pid ..., jar ..." line both call
# it with no args on purpose, so neither can ever be fooled by a leftover build output.
# Reports the hash of $JAR by default, or of whatever path is passed — used to report the STAGED
# jar right after a build (before it has been swapped in) without ever changing what a bare
# `jar_id` (no args) means: the live path, $JAR. --check and the final "pid ..., jar ..." line
# both call it with no args on purpose, so neither can ever be fooled by a leftover staged file.
# fleetd #550 — THREE distinct answers now, not two: `[ -f "$f" ]` already separates "the jar is
# not there" (-> "absent") from "the jar is there"; for the second case, hash256 itself separates
# "hashed it" (a 12-char hex string) from "could not hash it" (-> "unhashable", when no hasher is
@@ -181,35 +180,15 @@ hash256() {
# was the whole defect this ticket fixes.
jar_id() { local f="${1:-$JAR}"; [ -f "$f" ] && hash256 "$f" || echo "absent"; }
# fleetd #664 — under the old layout $JAR and the build output were the same file, so "jar on
# disk" was one fact. Now they are two: $BUILD_JAR (target/fleetd.jar, whatever Maven last wrote,
# by this script or by a bare `mvn install` run by hand) and $JAR (run/fleetd.jar, whatever the
# daemon actually has open). Printing one label for both was the trap this ticket exists to close
# — during the incident it would have shown the NEW jar's hash while the JVM ran the OLD one.
# Pure (reads jar_id/date, never mutates), so a test can call it directly without reaching the
# main flow — the same shape swap_if_built/drain_gate_refusal already use.
report_jar_state() {
local built_hash running_hash built_mtime running_mtime
built_hash="$(jar_id "$BUILD_JAR")"
running_hash="$(jar_id "$JAR")"
built_mtime="$([ -f "$BUILD_JAR" ] && date -r "$BUILD_JAR" '+%Y-%m-%d %H:%M:%S' || echo 'none')"
running_mtime="$([ -f "$JAR" ] && date -r "$JAR" '+%Y-%m-%d %H:%M:%S' || echo 'none')"
ok "built jar (target/fleetd.jar): $built_hash ($built_mtime)"
ok "running jar (run/fleetd.jar): $running_hash ($running_mtime)"
if [ "$built_hash" != "absent" ] && [ "$running_hash" != "absent" ] && [ "$built_hash" != "$running_hash" ]; then
warn "built jar and running jar differ — target/fleetd.jar was rebuilt since the running daemon last started and is not yet live"
fi
}
# fleetd #593 — `pgrep -f "$PATTERN"` matches ANY process whose full command line CONTAINS the
# pattern text, and that is not the same thing as "is the daemon". A shell that merely embeds the
# pattern as literal text — a human typing this exact investigation by hand, an ssh-shaped
# `sh -c '...; ...'`, a pipeline, or any other non-exec'ing shell that never replaced itself with
# the pattern-holding command — still shows up in that match, and it is the INSTRUMENT, not the
# daemon. Measured live on this Mac: `sh -c 'echo "run/fleetd.jar" >/dev/null; sleep 30' &`
# daemon. Measured live on this Mac: `sh -c 'echo "target/fleetd.jar" >/dev/null; sleep 30' &`
# leaves a real `sh` process alive (it forks for the `sleep`, it does not exec into it) whose own
# `ps -o args` is `sh -c echo "run/fleetd.jar" >/dev/null; sleep 30` — `pgrep -f "$PATTERN"`
# matches that line right alongside the real `java -jar run/fleetd.jar` process. `pgrep -c`
# `ps -o args` is `sh -c echo "target/fleetd.jar" >/dev/null; sleep 30` — `pgrep -f "$PATTERN"`
# matches that line right alongside the real `java -jar target/fleetd.jar` process. `pgrep -c`
# (an in-one-call count) does not exist on BSD/macOS at all, so this cannot be fixed by switching
# pgrep flags — it has to filter what pgrep already found, after the fact, in a way that still
# runs on BSD.
@@ -231,7 +210,7 @@ report_jar_state() {
# launched as `java -jar ...` — a native image, a renamed launcher — `running_pid()` silently
# returns nothing and `assert_single_daemon` stops noticing a second daemon at all. For a guard,
# that false-negative direction is the worse one to be wrong in. This is not a new assumption,
# though: `PATTERN='run/fleetd.jar'` two lines up already assumes the daemon is a jar, which
# though: `PATTERN='target/fleetd.jar'` two lines up already assumes the daemon is a jar, which
# is only ever run by `java`. If that launch method changes, `PATTERN` stops matching anything
# before this allowlist would ever get the chance to be wrong — the allowlist rides on the same
# assumption that is already load-bearing, it does not add a new one. Whoever changes the launch
@@ -250,12 +229,16 @@ running_pid() {
printf '%s' "$out"
}
# fleetd #493/#664 — the independently testable pieces of "never build into the path a running
# process holds":
# fleetd #493 — three small, independently testable pieces of "never build into the path a
# running process holds":
#
# require_no_build_jar the --no-build path never builds anything: it must find a jar already
# sitting at the live path ($JAR, under run/) from an earlier successful run,
# and die with the same truthful message this script has always used if not.
# stage_built_jar moves the jar Maven just produced OUT of the live path and onto the staging
# path, immediately after a successful build. Dies (leaving the OLD daemon
# untouched — this runs before the stop step) if Maven reported success but
# left no jar behind, or if the move itself fails.
# require_no_build_jar the --no-build path never builds or stages anything: it must find a
# jar already sitting at the live path from an earlier successful run, and
# die with the same truthful message this script has always used if not.
# wait_for_daemon_exit polls running_pid() for up to $1 seconds and reports whether the OLD
# daemon actually exited — extracted to its own function so the main flow
# can be relied on to call swap_staged_jar only AFTER this returns success,
@@ -265,10 +248,14 @@ running_pid() {
# so this is never a write into a path a running process holds — by the time
# it runs, nothing holds that path anymore. If it fails, the caller must not
# start a new daemon: die() below already refuses that by exiting the script.
# fleetd #664: the "staged" jar swap_if_built passes in is now $BUILD_JAR
# itself (target/fleetd.jar, Maven's own output) — a build no longer needs to
# be moved off the live path right after compiling, because target/ was never
# the live path to begin with.
stage_built_jar() {
[ -f "$JAR" ] || die "build succeeded but produced no jar at $JAR — cannot stage it for restart.
The running daemon was NOT touched."
mv -f "$JAR" "$JAR_STAGED" \
|| die "could not move the freshly built jar from $JAR to the staging path $JAR_STAGED.
The running daemon was NOT touched."
}
require_no_build_jar() {
[ -f "$JAR" ] || die "no jar at $JAR — run without --no-build"
}
@@ -295,7 +282,8 @@ swap_staged_jar() {
#
# The defect: the swap step used to be guarded inline by `if [ "$DO_BUILD" = 1 ]` in the main flow.
# Changing that to `if false` left the suite green and the swap never ran, so a redeploy reported
# every step succeeding while the daemon started on no jar at all or on a stale one.
# every step succeeding while the daemon started on no jar at all (stage_built_jar has already moved
# the freshly built one to $JAR_STAGED by then) or on a stale one.
# test_swap_ordered_after_wait_and_before_start could not catch it: it reads this script's own text
# and compares line positions, and a same-line edit moves no line.
#
@@ -324,12 +312,7 @@ swap_if_built() {
local do_build="$1"
should_swap "$do_build" || return 0
say "swap"
# fleetd #664: $JAR now lives under run/, a directory target/ never created. mkdir -p here,
# not inside swap_staged_jar itself — that function's own contract is tested on a missing
# parent directory (a failing mv), and widening it to auto-create one would change what that
# test proves.
mkdir -p "$(dirname "$JAR")"
swap_staged_jar "$BUILD_JAR" "$JAR"
swap_staged_jar "$JAR_STAGED" "$JAR"
ok "jar in place: $(jar_id)"
}
@@ -843,25 +826,16 @@ report_shutdown_drain() {
# --no-build, staged jar present -> ALSO "nothing changed", deliberately: --no-build itself builds
# and stages nothing (see require_no_build_jar above), so a staged jar found here is a leftover
# from an earlier, unrelated run. THIS run truly changed nothing, and the next DO_BUILD=1 run
# wipes that leftover for free — `mvn clean` deletes all of target/, $BUILD_JAR included,
# before the build even starts — so there is nothing here for the operator to lose track of.
# wipes that leftover before it builds (`rm -f "$JAR_STAGED"` in the build section above) — so
# there is nothing here for the operator to lose track of.
# --no-build, staged jar absent -> "nothing changed"
drain_gate_refusal() {
local do_build="$1" staged_path="$2"
if [ "$do_build" = 1 ] && [ -f "$staged_path" ]; then
# fleetd #664: under the old layout a build emptied the live path ($JAR) immediately, so
# --no-build's own check ("no jar at $JAR") was the thing that refused a rerun here. That is
# no longer true: a build never touches $JAR at all now, so $JAR still holds whatever was
# already running before this gate fired (reaching this message at all requires OLD_PID to
# have been set, which means a daemon was running from $JAR already) — a --no-build rerun
# would NOT refuse, it would just restart that same old jar and silently throw away the one
# sitting at %s.
printf 'aborted — the running daemon was NOT touched, but the freshly built jar is sitting at
%s, not yet swapped into %s. Rerun WITHOUT --no-build to finish the restart — a
--no-build rerun would NOT refuse here: %s already exists from before this run, so it
would restart the daemon on that OLD jar and silently discard the one you just built — or
remove %s by hand if you want to discard this build instead.' \
"$staged_path" "$JAR" "$JAR" "$staged_path"
%s, not yet swapped into %s. Rerun WITHOUT --no-build to finish the restart —
the freshly built jar is no longer at the live path that --no-build requires — or
remove %s by hand if you want to discard this build.' "$staged_path" "$JAR" "$staged_path"
else
printf 'aborted — nothing changed'
fi
@@ -1196,7 +1170,7 @@ if [ -n "$OLD_PID" ]; then
else
warn "no daemon running — this will be a cold start"
fi
report_jar_state
ok "jar on disk: $(jar_id) ($([ -f "$JAR" ] && date -r "$JAR" '+%Y-%m-%d %H:%M:%S' || echo 'none'))"
ok "HEAD: $(git -C "$REPO" log --oneline -1)"
# CB-594 / fleetd #492: supervision state. Installed and loaded are different facts — a
@@ -1278,6 +1252,9 @@ stop_if_check_only "$CHECK_ONLY"
if [ "$DO_BUILD" = 1 ]; then
say "build"
# fleetd #493: wipe a leftover staged jar from a previous failed/interrupted run BEFORE doing
# anything else, so that run's leftovers can never be mistaken for this run's output.
rm -f "$JAR_STAGED"
BUILD_LOG="$(mktemp -t fleetd-build.XXXXXX)"
echo " log: $BUILD_LOG"
if ! mvn -f "$MODULE/pom.xml" clean install > "$BUILD_LOG" 2>&1; then
@@ -1287,16 +1264,16 @@ if [ "$DO_BUILD" = 1 ]; then
fi
grep -E '^\[INFO\] Tests run:.*Failures' "$BUILD_LOG" | tail -1 | sed 's/^\[INFO\] / /' || true
ok "BUILD SUCCESS"
# fleetd #664: nothing to stage — $BUILD_JAR (target/fleetd.jar) is Maven's own output path and
# was never the live path, so the running (OLD) daemon, if any, was never at risk from this build
# at all. From here until the swap step below (after the OLD daemon is confirmed gone),
# $BUILD_JAR is the artefact this script treats as "the new jar" — $JAR itself is not touched
# again until the swap.
ok "jar now: $(jar_id "$BUILD_JAR")"
# fleetd #493: move the freshly built jar off the live path immediately — the running (OLD)
# daemon, if any, is still up at this point (build always runs before stop). From here until the
# swap step below (after the OLD daemon is confirmed gone), $JAR_STAGED is the only artefact this
# script treats as "the new jar" — $JAR itself is not touched again until the swap.
stage_built_jar
ok "jar now: $(jar_id "$JAR_STAGED")"
else
say "build skipped (--no-build)"
# fleetd #493: --no-build never builds anything — it restarts whatever jar is already sitting at
# the live path from an earlier successful run. Same check, same message as before.
# fleetd #493: --no-build never builds or stages anything — it restarts whatever jar is already
# sitting at the live path from an earlier successful run. Same check, same message as before.
require_no_build_jar
fi
@@ -1305,8 +1282,8 @@ fi
# fleetd #555: run_drain_gate above is drain_gate_required + the prompt + drain_confirmed, called
# unconditionally — it returns immediately when the gate is not required, and composes/dies through
# refuse_drain_gate itself when the reply does not confirm. See #493/#517/#528 for why "nothing
# changed" would be a lie once a build has produced a jar not yet swapped in.
run_drain_gate "$OLD_PID" "$ASSUME_YES" "$DO_BUILD" "$BUILD_JAR"
# changed" would be a lie once a build has staged a jar.
run_drain_gate "$OLD_PID" "$ASSUME_YES" "$DO_BUILD" "$JAR_STAGED"
# ------------------------------------------------------------------ stop
#
@@ -1357,22 +1334,21 @@ fi
# ------------------------------------------------------------------ swap
#
# fleetd #493/#664: every branch above has now either confirmed the OLD daemon actually exited
# (wait_for_daemon_exit, above) or established there was never one running to begin with. Only NOW
# is it safe to put the freshly built jar at the path the NEXT `java -jar` (direct, or via
# launchd/systemd's ExecStart) will read from — a rename from $BUILD_JAR (target/) to $JAR (run/),
# both under $MODULE and so on one filesystem. This mv is the one and only write to $JAR anywhere
# fleetd #493: every branch above has now either confirmed the OLD daemon actually exited
# (wait_for_daemon_exit, above) or established there was never one running to begin with. Only
# NOW is it safe to put the freshly built jar at the path the NEXT `java -jar` (direct, or via
# launchd/systemd's ExecStart) will read from — this mv is the one and only write to $JAR anywhere
# in this script's mutating flow. If it fails, do not start: die() below exits before "start" runs.
swap_if_built "$DO_BUILD"
# ------------------------------------------------------------------ start
# Unsupervised: login shell (zsh -l) is what puts the secrets on the daemon's environment, and cwd
# must be fleetd/ because the daemon resolves fleetd.yaml, logs/ and run/ relative to it.
# must be fleetd/ because the daemon resolves fleetd.yaml, logs/ and target/ relative to it.
# Supervised (launchd): launchd does both — deploy/dev.ltms.fleetd.plist points ProgramArguments at
# scripts/fleetd-launchd-wrapper.sh (CB-594), which is what execs the login shell in launchd's
# place, and WorkingDirectory in the plist already pins fleetd/.
# Supervised (systemd --user): the unit does both too — measured on the second host, ExecStart is
# `/bin/zsh -lc "exec java -jar run/fleetd.jar fleetd.yaml"` (a login shell, same reason as
# `/bin/zsh -lc "exec java -jar target/fleetd.jar fleetd.yaml"` (a login shell, same reason as
# above) and WorkingDirectory is already pinned to fleetd/.
say "start"
-143
View File
@@ -474,90 +474,6 @@ test_passphrase_key_is_redacted() {
assert_not_contains "FAKELEAK-PASSPHRASE" "$RUN_OUTPUT" "passphrase case: the passphrase VALUE must never leak"
}
# ------------------- acceptance criterion 19: the key line falls outside the printed hunk
# fleetd #656 — criteria 15a/15b both put the edit right next to the key line, so the key line is
# always inside diff -u's default 3-line context. Neither covers the actual case #639 fixed: an
# 8-line block-scalar body with only its SIXTH line changed, so the printed hunk (3 lines of
# context on each side of the change) covers body lines 3-8 and never includes the "token:" key
# line at all. The old, line-by-line redact() only ever masks after it has SEEN the key line go
# past; with the key line outside the hunk it never sets its mask, and the whole body — the
# changed line included — passes through raw. The control key sits right after the body, inside
# the same hunk, so the positive control below proves the fix is not simply printing nothing.
new_fixture_hunk_without_key_line() {
local dir
dir="$(mktemp -d "$TMP/fixture.XXXXXX")"
cat > "$dir/fleetd.yaml" <<'YAML'
bind:
host: 127.0.0.1
port: 19999
broker:
uri: amqp://user:hunter2@host/vhost
auth:
token: |
SECRET-LINE-1
SECRET-LINE-2
SECRET-LINE-3
SECRET-LINE-4
SECRET-LINE-5
SECRET-LINE-6
SECRET-LINE-7
SECRET-LINE-8
control: CTRL-MUST-APPEAR
profiles:
sonnet:
weight: 3
maxLoad: 5
YAML
: > "$dir/fleetd.out"
printf '%s' "$dir"
}
test_key_line_outside_hunk_is_still_redacted() {
local dir
dir="$(new_fixture_hunk_without_key_line)"
sed 's/SECRET-LINE-6$/SECRET-LINE-6-CHANGED/' "$dir/fleetd.yaml" > "$dir/candidate.yaml"
start_run "$dir" 5 --from "$dir/candidate.yaml"
sleep 1
printf 'config reloaded\n' >> "$dir/fleetd.out"
collect_run "$dir"
assert_equals 0 "$RUN_RC" "hunk-without-key-line case reload exit code"
# Positive control FIRST: without this, a diff that printed nothing at all would pass the
# negative assertion right below identically to a correctly redacted one.
assert_contains "CTRL-MUST-APPEAR" "$RUN_OUTPUT" "hunk-without-key-line case: the non-secret control line must still print unmasked"
assert_not_contains "SECRET-LINE-6-CHANGED" "$RUN_OUTPUT" "hunk-without-key-line case: the changed body line must never leak, even with the key line outside the printed hunk"
}
# ------------------------------- acceptance criterion 20: a blank line inside the value
# fleetd #656 — the old, line-by-line redact() reset its mask on any line whose indentation was
# not STRICTLY greater than the key's, and a wholly blank line has indentation 0, so it reset the
# mask exactly like the "control:" line that legitimately ends the block scalar. Everything after
# the blank line then printed raw. The current fix tracks masked lines by FILE line number instead
# of by indentation seen so far, so a blank line inside the value stays masked.
new_fixture_blank_line_in_value() {
local dir
dir="$(mktemp -d "$TMP/fixture.XXXXXX")"
printf 'bind:\n host: 127.0.0.1\n port: 19999\nbroker:\n uri: amqp://user:hunter2@host/vhost\nauth:\n token: |\n LEAK-BEFORE-BLANK\n\n LEAK-AFTER-BLANK\n control: CTRL-MUST-APPEAR\nprofiles:\n sonnet:\n weight: 3\n maxLoad: 5\n' > "$dir/fleetd.yaml"
: > "$dir/fleetd.out"
printf '%s' "$dir"
}
test_blank_line_inside_value_is_still_redacted() {
local dir
dir="$(new_fixture_blank_line_in_value)"
sed 's/LEAK-AFTER-BLANK$/LEAK-AFTER-BLANK-CHANGED/' "$dir/fleetd.yaml" > "$dir/candidate.yaml"
start_run "$dir" 5 --from "$dir/candidate.yaml"
sleep 1
printf 'config reloaded\n' >> "$dir/fleetd.out"
collect_run "$dir"
assert_equals 0 "$RUN_RC" "blank-line-in-value case reload exit code"
assert_contains "CTRL-MUST-APPEAR" "$RUN_OUTPUT" "blank-line-in-value case: the non-secret control line must still print unmasked"
assert_not_contains "LEAK-AFTER-BLANK-CHANGED" "$RUN_OUTPUT" "blank-line-in-value case: the line after the blank must never leak"
}
# ----------------------------------- acceptance criterion 16: a failing --set must not echo value
# fleetd #635 follow-up (ticket comment 17673, defect 8) — apply_set_pairs used to echo the FULL
# "$kv" (path=value, exactly as typed) in its yq-failure messages, so a broken --set with a
@@ -629,57 +545,6 @@ test_refusal_shape_from_parse_failure_wording_is_recognised() {
assert_equals 4 "$RUN_RC" "the parse-failure refusal shape must also exit 4, not be read as silence"
}
# --set runs yq over the whole candidate. It warns when that changes more lines than the requested
# pairs, but a simple file with only the intended changed line must stay quiet.
new_fixture_reformat_sensitive() {
local dir
dir="$(mktemp -d "$TMP/fixture.XXXXXX")"
cat > "$dir/fleetd.yaml" <<'YAML'
# A section comment that documents the next block.
bind:
host: 127.0.0.1 # Keep this aligned with the port note.
port: 19999 # A fixture port.
# These comments use their placement as documentation.
profiles:
sonnet:
weight: 3
bootstrapText: >-
First line.
Second line.
YAML
: > "$dir/fleetd.out"
printf '%s' "$dir"
}
test_set_warns_when_yq_reformats_extra_lines() {
local dir
dir="$(new_fixture_reformat_sensitive)"
start_run "$dir" 5 --set '.profiles.sonnet.weight=4'
sleep 1
printf 'config reloaded\n' >> "$dir/fleetd.out"
collect_run "$dir"
assert_equals 0 "$RUN_RC" "reformat warning case reload exit code"
assert_contains "yq reformatted the whole file" "$RUN_OUTPUT" \
"a --set that changes extra candidate lines must warn before installation"
}
test_set_stays_quiet_without_formatting_churn() {
local dir
dir="$(new_fixture)"
start_run "$dir" 5 --set '.profiles.sonnet.weight=4'
sleep 1
printf 'config reloaded\n' >> "$dir/fleetd.out"
collect_run "$dir"
assert_equals 0 "$RUN_RC" "no-reformat warning case reload exit code"
assert_not_contains "yq reformatted the whole file" "$RUN_OUTPUT" \
"a --set that changes only its requested candidate line must not warn"
}
echo "== acceptance criterion 1: refusal restores byte for byte =="
test_refusal_restores_byte_for_byte
echo "== acceptance criterion 2: clean reload keeps the edit =="
@@ -708,10 +573,6 @@ echo "== acceptance criterion 15a: a block scalar's continuation lines are redac
test_block_scalar_continuation_is_redacted
echo "== acceptance criterion 15b: a passphrase key is also recognised =="
test_passphrase_key_is_redacted
echo "== acceptance criterion 19: the key line falls outside the printed hunk =="
test_key_line_outside_hunk_is_still_redacted
echo "== acceptance criterion 20: a blank line inside the value =="
test_blank_line_inside_value_is_still_redacted
echo "== acceptance criterion 16: a failing --set must not echo its value =="
test_failing_set_does_not_echo_its_value
echo "== extra: dry-run never installs, and redacts =="
@@ -720,9 +581,5 @@ echo "== extra: --check is read-only and always exits 0 =="
test_check_is_read_only_and_exits_zero
echo "== extra: the parse-failure refusal shape is also recognised =="
test_refusal_shape_from_parse_failure_wording_is_recognised
echo "== acceptance criterion 17: --set warns about yq formatting churn =="
test_set_warns_when_yq_reformats_extra_lines
echo "== acceptance criterion 18: --set stays quiet without formatting churn =="
test_set_stays_quiet_without_formatting_churn
printf 'PASS: config-edit acceptance criteria\n'
+63 -100
View File
@@ -447,7 +447,7 @@ test_assert_single_daemon_rejects_two_pids() {
test_running_pid_excludes_self_matching_wrapper_shell() {
local before after wrapper_pid
before="$(running_pid)"
sh -c 'echo "run/fleetd.jar" >/dev/null; sleep 20' &
sh -c 'echo "target/fleetd.jar" >/dev/null; sleep 20' &
wrapper_pid=$!
sleep 0.3
after="$(running_pid)"
@@ -475,14 +475,14 @@ test_running_pid_excludes_self_matching_wrapper_shell() {
# `comm` from the actually-executed binary's own path, not from `exec -a`'s argv[0] override (BSD
# ties `comm` to argv[0], which is what makes this technique work here) — so on Linux this
# specific fixture might report `comm=sh`, not `comm=java`, even though the REAL daemon (a literal
# `java -jar run/fleetd.jar` process, never fabricated) is unaffected either way. I could not
# `java -jar target/fleetd.jar` process, never fabricated) is unaffected either way. I could not
# verify this fixture's behavior on Linux, so test_running_pid_counts_a_pid_whose_comm_is_java
# below backstops the same claim (the allowlist admits a pid whose comm is `java`) with a stubbed
# `ps`, which is identical bash on every platform and carries no such platform question.
test_running_pid_finds_a_real_java_named_second_process() {
local before after standin_pid
before="$(running_pid)"
( exec -a java sh -c 'echo "run/fleetd.jar" >/dev/null; sleep 20' ) &
( exec -a java sh -c 'echo "target/fleetd.jar" >/dev/null; sleep 20' ) &
standin_pid=$!
sleep 0.3
after="$(running_pid)"
@@ -557,30 +557,29 @@ test_die_message_does_not_recommend_bare_pgrep_as_remediation() {
|| fail "assert_single_daemon's die message does not say in words that a pattern can match the caller (fleetd #593)"
}
# fleetd #511/#664 — jar_id()'s no-argument default was unpinned by any test: nothing proved it
# reports $JAR (the live path) rather than whatever explicit path a caller passes it (e.g.
# $BUILD_JAR). Both halves matter, so this pins both: the bare call must hash the live jar, and an
# explicit path argument must hash THAT file, not fall back to $JAR. Two files with different
# content, so a default pointed at the wrong one reports the wrong hash rather than accidentally
# matching.
# fleetd #511 — jar_id()'s no-argument default was unpinned by any test: nothing proved it reports
# $JAR (the live path) rather than $JAR_STAGED. Both halves matter, so this pins both: the bare call
# must hash the live jar, and an explicit path argument must hash THAT file, not fall back to $JAR.
# Two files with different content, so a default pointed at the wrong one reports the wrong hash
# rather than accidentally matching.
test_jar_id_defaults_to_live_and_reports_explicit_path() {
local dir saved_jar="$JAR"
local live_hash other_hash default_result explicit_result other_path
local dir saved_jar="$JAR" saved_staged="$JAR_STAGED"
local live_hash staged_hash default_result explicit_result
dir="$TMP/jar-id"; mkdir -p "$dir"
JAR="$dir/fleetd.jar"; other_path="$dir/other.jar"
JAR="$dir/fleetd.jar"; JAR_STAGED="$dir/fleetd-new.jar"
printf 'live jar bytes' > "$JAR"
printf 'other jar bytes, not the same content' > "$other_path"
printf 'staged jar bytes, not the same content' > "$JAR_STAGED"
# fleetd #550: this reference hash must be computed the same portable way jar_id() itself now
# computes one — a bare, unguarded call to the macOS-only hasher here was exactly the item-2
# defect, dying with "command not found" on any Linux runner that has no such hasher at all.
live_hash="$(hash256 "$JAR")"
other_hash="$(hash256 "$other_path")"
staged_hash="$(hash256 "$JAR_STAGED")"
default_result="$(jar_id)"
explicit_result="$(jar_id "$other_path")"
JAR="$saved_jar"
[ "$live_hash" != "$other_hash" ] || fail "test fixture error: live and other jars hashed the same"
explicit_result="$(jar_id "$JAR_STAGED")"
JAR="$saved_jar"; JAR_STAGED="$saved_staged"
[ "$live_hash" != "$staged_hash" ] || fail "test fixture error: live and staged jars hashed the same"
assert_equals "$live_hash" "$default_result" "jar_id with no arguments must report the hash of \$JAR"
assert_equals "$other_hash" "$explicit_result" "jar_id with an explicit path must report the hash of that path, not fall back to \$JAR"
assert_equals "$staged_hash" "$explicit_result" "jar_id \"\$JAR_STAGED\" must report the hash of the staged jar, not fall back to \$JAR"
}
# fleetd #550 — closes a gap the test above leaves open. That test's own reference hash is now ALSO
@@ -652,64 +651,34 @@ test_jar_id_reports_unhashable_when_no_hasher_on_path() {
assert_equals "unhashable" "$explicit_result" "jar_id (explicit path) with no hasher on PATH must report the same third state"
}
# fleetd #664 — report_jar_state is the --check fix: under the old layout $JAR and the build
# output were the same file, so a single "jar on disk" fact covered both. Now they can disagree,
# and this is the function that is supposed to show that. Agreeing case: two files with IDENTICAL
# content must print both labels and never warn.
test_report_jar_state_agrees_when_hashes_match() {
local dir saved_build="$BUILD_JAR" saved_jar="$JAR" output
dir="$TMP/report-jar-agree"; mkdir -p "$dir"
BUILD_JAR="$dir/target-fleetd.jar"; JAR="$dir/run-fleetd.jar"
printf 'identical jar bytes' > "$BUILD_JAR"
printf 'identical jar bytes' > "$JAR"
output="$(report_jar_state)"
BUILD_JAR="$saved_build"; JAR="$saved_jar"
printf '%s' "$output" | grep -qF 'built jar' \
|| fail "report_jar_state did not label the built jar"
printf '%s' "$output" | grep -qF 'running jar' \
|| fail "report_jar_state did not label the running jar"
if printf '%s' "$output" | grep -qF 'differ'; then
fail "report_jar_state warned about a mismatch when both jars have identical content"
fi
}
# The disagreeing case: this is the whole point of the ticket — a built jar that is NOT the
# running jar must be visibly flagged, not silently printed as two unremarkable facts.
test_report_jar_state_warns_when_hashes_differ() {
local dir saved_build="$BUILD_JAR" saved_jar="$JAR" output
dir="$TMP/report-jar-differ"; mkdir -p "$dir"
BUILD_JAR="$dir/target-fleetd.jar"; JAR="$dir/run-fleetd.jar"
printf 'freshly built jar bytes' > "$BUILD_JAR"
printf 'older running jar bytes' > "$JAR"
output="$(report_jar_state)"
BUILD_JAR="$saved_build"; JAR="$saved_jar"
printf '%s' "$output" | grep -qF 'differ' \
|| fail "report_jar_state did not warn when the built jar and running jar disagree"
}
# Neither file existing (a fresh checkout, never built or deployed) must report two "absent"
# facts and never a false mismatch warning — "absent" vs "absent" is agreement, not a diff.
test_report_jar_state_both_absent_is_not_a_mismatch() {
local dir saved_build="$BUILD_JAR" saved_jar="$JAR" output
dir="$TMP/report-jar-absent"; mkdir -p "$dir"
BUILD_JAR="$dir/no-such-target.jar"; JAR="$dir/no-such-run.jar"
output="$(report_jar_state)"
BUILD_JAR="$saved_build"; JAR="$saved_jar"
printf '%s' "$output" | grep -qF 'absent' \
|| fail "report_jar_state did not report absent for a missing built/running jar"
if printf '%s' "$output" | grep -qF 'differ'; then
fail "report_jar_state warned about a mismatch when both jars are simply absent"
fi
}
# fleetd #493/#664 — never build into the path a running process holds. swap_staged_jar is
# exercised directly against real files on disk (not stubs), because the whole point is file
# fleetd #493 — never build into the path a running process holds. stage_built_jar/swap_staged_jar
# are exercised directly against real files on disk (not stubs), because the whole point is file
# behavior (does the content move, does the source disappear, does a failure leave both sides
# intact) that a stubbed function cannot prove. stage_built_jar no longer exists: under the #664
# layout $BUILD_JAR (target/fleetd.jar) was never the live path, so a build has nothing to be
# staged OUT of — swap_staged_jar is called directly against $BUILD_JAR/$JAR (see
# test_swap_ordered_after_wait_and_before_start and the report_jar_state tests below for the rest
# of that seam).
# intact) that a stubbed function cannot prove.
test_stage_built_jar_moves_off_live_path() {
local dir jar staged saved_jar="$JAR" saved_staged="$JAR_STAGED"
dir="$TMP/stage-ok"; mkdir -p "$dir"
jar="$dir/fleetd.jar"; staged="$dir/fleetd-new.jar"
printf 'built jar bytes' > "$jar"
JAR="$jar"; JAR_STAGED="$staged"
stage_built_jar || fail "stage_built_jar rejected a real build output"
JAR="$saved_jar"; JAR_STAGED="$saved_staged"
[ ! -f "$jar" ] || fail "stage_built_jar left the jar behind at the live path $jar"
[ -f "$staged" ] || fail "stage_built_jar did not create the staged jar at $staged"
grep -qF 'built jar bytes' "$staged" || fail "staged jar does not carry the built content"
}
test_stage_built_jar_dies_when_build_produced_nothing() {
local dir output rc=0 saved_jar="$JAR" saved_staged="$JAR_STAGED"
dir="$TMP/stage-missing"; mkdir -p "$dir"
JAR="$dir/fleetd.jar"; JAR_STAGED="$dir/fleetd-new.jar"
output="$(stage_built_jar 2>&1)" || rc=$?
JAR="$saved_jar"; JAR_STAGED="$saved_staged"
[ "$rc" -ne 0 ] || fail "stage_built_jar accepted a missing build output"
printf '%s' "$output" | grep -qF "$dir/fleetd.jar" \
|| fail "refusal message does not name the missing jar path"
}
test_swap_staged_jar_moves_staged_onto_live() {
local dir staged live
dir="$TMP/swap-ok"; mkdir -p "$dir"
@@ -1020,26 +989,24 @@ test_swap_ordered_after_wait_and_before_start() {
|| fail "swap_if_built (line $swap_line) is not before the start section (line $start_line)"
}
# fleetd #511: the drain-gate abort message (fired when a build has produced a jar but the
# operator declines the drain confirmation) used to tell the operator to "Rerun (with or without
# --no-build)" to finish the restart. That is wrong. fleetd #664 changed WHY it is wrong: under
# the old layout a build emptied the live path immediately, so --no-build's own check refused a
# rerun for you; now a build never touches the live path at all, so --no-build would NOT refuse —
# it would quietly restart the daemon on the OLD jar and throw away the one just built. Like
# test_swap_ordered_after_wait_and_before_start above, this code path is never reached by sourcing
# (the SOURCED guard stops before the main flow), so the only way to pin its exact wording is to
# read the source.
# fleetd #511: the drain-gate abort message (fired when a build has staged a jar but the operator
# declines the drain confirmation) used to tell the operator to "Rerun (with or without --no-build)"
# to finish the restart. That is wrong — by the time this message can fire, stage_built_jar has
# already moved the jar off $JAR, so a rerun WITH --no-build hits require_no_build_jar's own refusal
# ("no jar at $JAR — run without --no-build"). Like test_swap_ordered_after_wait_and_before_start
# above, this code path is never reached by sourcing (the SOURCED guard stops before the main flow),
# so the only way to pin its exact wording is to read the source.
test_drain_gate_abort_message_says_no_no_build() {
local src="$ROOT/scripts/redeploy-fleetd.sh" msg
msg="$(grep -A6 -F 'aborted — the running daemon was NOT touched, but the freshly built jar is sitting at' "$src")"
msg="$(grep -A3 -F 'aborted — the running daemon was NOT touched, but the freshly built jar is sitting at' "$src")"
[ -n "$msg" ] || fail "could not find the drain-gate staged-jar abort message in redeploy-fleetd.sh"
if printf '%s' "$msg" | grep -qF 'with or without --no-build'; then
fail "abort message still claims a rerun WITH --no-build can finish the restart"
fi
printf '%s' "$msg" | grep -qF 'WITHOUT --no-build' \
|| fail "abort message does not tell the operator to rerun without --no-build"
printf '%s' "$msg" | grep -qF 'silently discard' \
|| fail "abort message does not say that --no-build would silently discard the build just made"
printf '%s' "$msg" | grep -qF 'no longer at the live path' \
|| fail "abort message does not say why --no-build cannot finish the restart"
}
# fleetd #517 — the drain-gate abort branch itself. Before this, the only test of this message was
@@ -1182,22 +1149,19 @@ test_refuse_drain_gate_no_build_staged_absent() {
#
# fleetd #555 — the main flow's own call site moved: it used to read
# `refuse_drain_gate "$DO_BUILD" "$JAR_STAGED"` directly; it now reads
# `run_drain_gate "$OLD_PID" "$ASSUME_YES" "$DO_BUILD" "$BUILD_JAR"` (fleetd #664 renamed the
# fourth argument from $JAR_STAGED to $BUILD_JAR — same role, the not-yet-swapped-in jar — when
# that path stopped being a separate staging file and became target/fleetd.jar itself), and
# run_drain_gate (tested directly below by test_run_drain_gate_*) is what calls refuse_drain_gate
# with its own local names. This grep now pins THAT call site — the thing that would go missing if
# a future edit deleted the main flow's call to run_drain_gate altogether, the same residual gap
# #521/#528 already accepted for swap_if_built/refuse_drain_gate (sourcing stops before the main
# flow runs, so no test in this file can do better than reading the source for this one specific
# gap).
# `run_drain_gate "$OLD_PID" "$ASSUME_YES" "$DO_BUILD" "$JAR_STAGED"`, and run_drain_gate (tested
# directly below by test_run_drain_gate_*) is what calls refuse_drain_gate with its own local names.
# This grep now pins THAT call site — the thing that would go missing if a future edit deleted the
# main flow's call to run_drain_gate altogether, the same residual gap #521/#528 already accepted for
# swap_if_built/refuse_drain_gate (sourcing stops before the main flow runs, so no test in this file
# can do better than reading the source for this one specific gap).
#
# The grep ends `|| true`: this file runs under `set -euo pipefail`, so an ABSENT needle would fail
# the assignment and `set -e` would kill the whole suite before the `[ -n ... ] || fail` guard below
# ever ran — the exact dead-check shape fleetd #528 also flags as a sweep finding (see the PR body).
test_run_drain_gate_call_site_present() {
local src="$ROOT/scripts/redeploy-fleetd.sh" call_line
call_line="$(grep -Fn 'run_drain_gate "$OLD_PID" "$ASSUME_YES" "$DO_BUILD" "$BUILD_JAR"' "$src" | head -1 | cut -d: -f1 || true)"
call_line="$(grep -Fn 'run_drain_gate "$OLD_PID" "$ASSUME_YES" "$DO_BUILD" "$JAR_STAGED"' "$src" | head -1 | cut -d: -f1 || true)"
[ -n "$call_line" ] \
|| fail "could not find the main flow's run_drain_gate call site in redeploy-fleetd.sh"
}
@@ -2197,9 +2161,8 @@ test_jar_id_defaults_to_live_and_reports_explicit_path
test_hash256_computes_a_real_sha256
test_jar_id_reports_absent_for_missing_file
test_jar_id_reports_unhashable_when_no_hasher_on_path
test_report_jar_state_agrees_when_hashes_match
test_report_jar_state_warns_when_hashes_differ
test_report_jar_state_both_absent_is_not_a_mismatch
test_stage_built_jar_moves_off_live_path
test_stage_built_jar_dies_when_build_produced_nothing
test_swap_staged_jar_moves_staged_onto_live
test_swap_staged_jar_dies_without_staged_file
test_swap_staged_jar_dies_when_mv_fails