Compare commits

..

2 Commits

Author SHA1 Message Date
Dai Ha 51f7b0a3ca fleetd #111: probe reads the live memberCredentials policy, no hardcoded name list
CI / contract (pull_request) Successful in 1m24s
CI / build (pull_request) Successful in 2m22s
scripts/probe-member-credentials.sh carried its own hand-maintained NAMES array
(31 names, recorded 2026-08-16), so a name added later to fleetd.yaml's
memberCredentials.known was never checked and the probe still exited 0 with a
clean-looking table. Same drift shape as #114's tool catalogue.

- New dev.ltms.fleet.member.MemberCredentialPolicyView: the single place that
  turns a MemberCredentials policy into names + counts (never a value). Reused
  by Fleetd.reportMemberCredentialsGap (startup log line) and by the new
  GET /member-credentials REST endpoint (FleetApp), so the two can no longer
  drift apart the way the probe and the policy did.
- FleetApp gains one route + handler + a Supplier<MemberCredentialPolicyView>
  constructor param (legacy constructors default to ::absent, so existing call
  sites are unaffected).
- probe-member-credentials.sh now fetches its name list from
  GET /member-credentials instead of carrying one. No local fallback: an
  unreachable daemon, an empty/absent policy, or a knownCount/known[] length
  mismatch all refuse with a non-zero exit rather than silently checking zero
  names. Prints "policy contains N; this run checked N" so the two numbers are
  visibly equal.
2026-09-03 13:21:35 +07:00
Dai Ha 01a840cc14 #248 follow-up: drive the real backendErrorSink, not a copy of it
CI / contract (push) Successful in 1m16s
CI / build (push) Successful in 1m29s
BackendOutageFlowTest held a ~30-line hand-copy of the lambda in
Fleetd.main, under a comment promising it mirrored production "EXACTLY".
That promise was the defect. The test proved the copy, so any change to
the real sink left the flow test green.

#248 made Fleetd.backendErrorSink(...) public for exactly this reason.
The test now calls it.

Measured, same mutation in the real sink (an early return after
sessions.onBackendError, dropping the cool-off and the lead nudge):

  old test (hand-copy):  Tests run: 5, Failures: 0  -- blind
  new test (real sink):  Tests run: 5, Failures: 4  -- catches it

0 compile errors in both runs, so both are real results. Production
reverted and confirmed with diff -q.

Full build: 1234 tests, 0 failures, 0 compile errors.
2026-09-03 13:06:11 +07:00
10 changed files with 470 additions and 210 deletions
@@ -51,6 +51,7 @@ import dev.ltms.fleet.session.SessionReaper;
import dev.ltms.fleet.member.ClaudeCodeLauncher;
import dev.ltms.fleet.member.CompositePeerLauncher;
import dev.ltms.fleet.member.HerdrPeerLauncher;
import dev.ltms.fleet.member.MemberCredentialPolicyView;
import dev.ltms.fleet.member.OpenCodeLauncher;
import dev.ltms.fleet.placement.BackendOutagePolicy;
import dev.ltms.fleet.placement.BackendQuarantine;
@@ -708,8 +709,11 @@ public final class Fleetd {
// CB-185: give FleetApp both daemons — /healthz must require both to answer and
// GET /sessions must merge across both, or a down/unpolled member daemon is invisible.
// fleetd #111: live (re-read-per-request) memberCredentials view for GET /member-credentials —
// same hot-reload shape as the memberCredentials supplier passed to ClaudeCodeLauncher above.
Javalin app = new FleetApp(herdr, memberHerdr, workers, sessions, messages, presence, mcp.servlet(),
callers, metrics, deliverable).build();
callers, metrics, deliverable,
() -> MemberCredentialPolicyView.of(config.get().memberCredentials())).build();
app.start(cfg.bind().host(), cfg.bind().port());
log.info("fleetd listening on {}:{}, herdr socket {}",
cfg.bind().host(), cfg.bind().port(), socket);
@@ -1082,12 +1086,14 @@ public final class Fleetd {
* #requiredSecretEnvVars} is exposed for {@link #reportRequiredSecrets}'s own test.
*/
static void reportMemberCredentialsGap(FleetConfig cfg) {
FleetConfig.MemberCredentials creds = cfg.memberCredentials();
if (creds != null && !creds.known().isEmpty()) {
// fleetd #111: the counts below come from MemberCredentialPolicyView, the same class the
// live GET /member-credentials endpoint reads — one place computes them, not two.
MemberCredentialPolicyView view = MemberCredentialPolicyView.of(cfg.memberCredentials());
if (view.present()) {
log.info("memberCredentials: policy={}, {} known name(s), {} allowed — blocking {} on "
+ "every spawn{}",
creds.policy(), creds.known().size(), creds.allow().size(), creds.blockedSet().size(),
creds.isAllowList()
view.policy(), view.knownCount(), view.allowedCount(), view.blockedCount(),
cfg.memberCredentials().isAllowList()
? " (allow-list: known/allow are reporting only — the control is the derived ZDOTDIR scrub)"
: "");
return;
@@ -92,15 +92,9 @@ import java.util.regex.PatternSyntaxException;
* fleetd's own process, so fleetd's own {@code $SHELL} says nothing about what
* that pane runs. There is no channel to ask herdr for another user's shell, so
* this must be told, never guessed. {@code null}/blank (or a value not ending
* in {@code zsh}) is treated the same as "not zsh". fleetd #155: what that means
* now depends on {@code memberCredentials.policy}. Under {@code deny-by-default}
* it stays a degraded control, never a refusal to spawn — the pane-creation
* overlay is unaffected by shell type, so the spawn proceeds with a WARN naming
* the shell. Under {@code allow-list} the spawn is REFUSED instead: that policy's
* whole point is a control a sourced file cannot undo, so silently falling back
* to the weaker overlay would be the same "control silently does nothing" defect
* #155 exists to remove — configure this field (or move the account to zsh, or
* switch policy back to {@code deny-by-default}) to unblock the spawn.
* in {@code zsh}) is treated the same as "not zsh": the {@code
* memberCredentials.policy: allow-list} ZDOTDIR scrub is skipped in favour of
* the CB-596 sentinel overlay — a degraded control, never a refusal to spawn.
* When {@code memberHerdrSocket} is NOT configured this field is never
* consulted at all; fleetd keeps reading its own {@code $SHELL}, exactly as
* before this field existed.
@@ -184,11 +184,7 @@ public abstract class HerdrPeerLauncher implements PeerLauncher {
*/
private final ConcurrentMap<String, Path> zdotdirByPane = new ConcurrentHashMap<>();
/**
* Guards {@link #warnNonZsh} to one WARN per launcher instance, not one per spawn. fleetd #155:
* only the {@code policy: deny-by-default} path still warns on a non-zsh shell — the {@code
* policy: allow-list} path refuses the spawn instead (see {@link #applyEnvironmentAllowListPolicy}).
*/
/** Guards {@link #warnNonZsh} to one WARN per launcher instance, not one per spawn. */
private final AtomicBoolean nonZshShellWarned = new AtomicBoolean();
/** Live config provides URI environment names that must never enter member panes. */
private final Supplier<FleetConfig> config;
@@ -1259,11 +1255,6 @@ public abstract class HerdrPeerLauncher implements PeerLauncher {
if (!creds.isAllowList()) {
overlayBlockedCredentials(workerEnv, creds);
logCredentialGap(creds, null);
// fleetd #155: this overlay itself does not depend on the shell (it lands in the
// pane-creation env map before any shell runs), so the spawn is never refused here —
// only policy=allow-list's ZDOTDIR scrub needs a login shell to run at all. Still worth
// telling the operator: the stronger post-shell control is unavailable on this shell.
warnNonZsh(memberLoginShell());
}
}
@@ -1316,16 +1307,10 @@ public abstract class HerdrPeerLauncher implements PeerLauncher {
* dev.ltms.fleet.session.Worktrees#shareWithGroup} already uses, reused rather than
* inventing a second group key. Either one missing means the scrub cannot be guaranteed
* reachable by the member, which is the same "cannot guarantee the scrub runs" case as a
* non-zsh shell — that one still gets the fallback (see fleetd #155's javadoc note below
* on why a missing worktreeRoot/worktreeGroup is a different defect, out of that ticket's
* scope).</li>
* non-zsh shell, so it gets the identical fallback.</li>
* </ul>
* fleetd #155: the one exception to "never refuse to spawn" is a non-zsh login shell, right
* below. The operator asked for {@code policy: allow-list} specifically — its whole point is a
* control a sourced file cannot undo — so silently degrading to the weaker overlay is exactly
* the "silently does nothing" failure this ticket exists to remove. Every other branch in this
* method (worktreeRoot/worktreeGroup missing) keeps the old "degrade, never refuse" behaviour;
* that gap is real but is fleetd #213's scope, not this one.
* In every branch: never refuse to spawn. A degraded credential control must not become an
* outage for an opt-in feature.
*/
private Path applyEnvironmentAllowListPolicy(FleetConfig.Profile cfg, Launch launch) {
FleetConfig.MemberCredentials creds = memberCredentials == null ? null : memberCredentials.get();
@@ -1339,21 +1324,20 @@ public abstract class HerdrPeerLauncher implements PeerLauncher {
// must not even be called on this path, only the explicit memberLoginShell: config can
// answer it. With memberHerdrSocket absent, nothing here changes: fleetd's own $SHELL is
// still the input, exactly as before this fix.
String loginShell = memberLoginShell();
String loginShell = memberHerdrSocket ? configuredMemberLoginShell() : resolveEnv("SHELL");
boolean zsh = isZshShell(loginShell);
if (!zsh) {
// fleetd #155: a non-zsh login shell ignores ZDOTDIR entirely — NO scrub would run. The
// operator explicitly asked for policy=allow-list's blocking control, so degrading to
// the weaker CB-596 overlay and spawning anyway would be the same "control silently does
// nothing" defect this ticket exists to close. Refuse instead — the caller (FleetMcp.spawn)
// catches IllegalArgumentException and surfaces the message to the operator.
throw new IllegalArgumentException(
"memberCredentials policy=allow-list requires the member's login shell to be "
+ "zsh, so the ZDOTDIR scrub can run after it — refusing to spawn under "
+ "login shell '" + (loginShell == null ? "<unset>" : loginShell)
+ "'. Configure memberLoginShell: as a zsh path (only read when "
+ "memberHerdrSocket is set), move the member's OS account onto zsh, or "
+ "set memberCredentials.policy: deny-by-default instead.");
// A non-zsh login shell ignores ZDOTDIR entirely: NO scrub would run, so pretending
// otherwise would be worse than saying so. Warn loudly and fall back to the CB-596
// sentinel overlay over the enumerated known: names — weaker (a sourced file can undo
// it), but strictly better than nothing. Deliberately no "allowed N of M" line here: the
// scrub this count describes does not run on this path, so printing it would tell an
// operator that a fraction of names were blocked when the real number blocked is zero.
// logCredentialGap's WARN (below) is the only signal for this path.
warnNonZsh(loginShell);
overlayBlockedCredentials(launch.env(), creds);
logCredentialGap(creds, null);
return null;
}
Path parentDir;
String group = null;
@@ -1408,20 +1392,6 @@ public abstract class HerdrPeerLauncher implements PeerLauncher {
return cfg == null ? null : cfg.memberLoginShell();
}
/**
* fleetd #155: the member pane's login shell, using the same {@code memberHerdrSocket} routing
* {@link #applyEnvironmentAllowListPolicy} already used — fleetd's own {@code $SHELL} decides it
* when {@code memberHerdrSocket} is absent (today's only mode); the configured {@code
* memberLoginShell:} decides it otherwise, since fleetd's own {@code $SHELL} names a different
* user's shell once member panes run under a different OS user. Shared by both {@link
* #applyEnvironmentAllowListPolicy} (allow-list: refuses on non-zsh) and {@link
* #applyMemberCredentialPolicy} (deny-by-default: warns on non-zsh) so the two policies agree on
* what "the member's shell" means.
*/
private String memberLoginShell() {
return memberHerdrSocketConfigured() ? configuredMemberLoginShell() : resolveEnv("SHELL");
}
/**
* fleetd #213: {@code worktreeRoot}, as the ZDOTDIR scrub's parent directory under {@code
* memberHerdrSocket}, or {@code null} when unconfigured — the same "cannot guarantee the scrub
@@ -1506,23 +1476,15 @@ public abstract class HerdrPeerLauncher implements PeerLauncher {
private static final String SSH_AUTH_SOCK = MemberEnvAllowList.SSH_AUTH_SOCK;
/**
* fleetd #155: called ONLY from the {@code policy: deny-by-default} branch of {@link
* #applyMemberCredentialPolicy} — under {@code policy: allow-list}, a non-zsh shell now REFUSES
* the spawn instead (see {@link #applyEnvironmentAllowListPolicy}), since that policy's whole
* point is a control a sourced file cannot undo. deny-by-default's own overlay does not depend
* on the shell (it lands in the pane-creation env map before any shell runs), so this is a
* heads-up, not a defect report: say once per launcher instance, naming the shell, that the
* stronger allow-list control is unavailable here — never silently.
* CB-633: a non-zsh login shell means the allow-list control CANNOT run — say so once per
* launcher instance, naming the shell, instead of failing silently.
*/
private void warnNonZsh(String shell) {
if (isZshShell(shell)) {
return;
}
if (nonZshShellWarned.compareAndSet(false, true)) {
log.warn("memberCredentials policy=deny-by-default: member login shell '{}' is NOT zsh — "
+ "the pane-creation credential overlay still applies here (it does not "
+ "depend on the shell), but policy=allow-list's stronger post-shell "
+ "ZDOTDIR scrub is unavailable on this shell.",
log.warn("memberCredentials policy=allow-list: member login shell '{}' is NOT zsh — "
+ "ZDOTDIR scrubbing cannot run, so members' inherited environment is "
+ "UNPROTECTED beyond the enumerated known: fallback. Move herdr onto a "
+ "zsh account or switch policy back to deny-by-default.",
shell == null ? "<unset>" : shell);
}
}
@@ -0,0 +1,75 @@
package dev.ltms.fleet.member;
import dev.ltms.fleet.config.FleetConfig;
import java.util.List;
/**
* fleetd #111 (CB-608): a single, testable read of the {@code memberCredentials:} policy — names
* and counts only, never a value. The daemon never holds a credential's <em>value</em> in the
* first place (only the names configured under {@code known:}/{@code allow:}), so there is
* nothing here to redact by construction; the point of this class is that it is the ONE place
* that turns a policy into names-and-counts, so nothing else hand-counts a second time.
*
* <p>Before this class, {@link dev.ltms.fleet.Fleetd#reportMemberCredentialsGap} computed these
* same counts inline for the startup log line, and {@code scripts/probe-member-credentials.sh}
* carried its own hardcoded {@code NAMES} array that the live policy could grow past silently
* (#111) — the exact "hand-maintained second copy drifts" shape #114 fixed for the tool
* catalogue. Both now read this class: the startup log via {@link
* dev.ltms.fleet.Fleetd#reportMemberCredentialsGap}, and a live daemon via the {@code
* GET /member-credentials} REST endpoint ({@link dev.ltms.fleet.rest.FleetApp}), which the probe
* script fetches instead of carrying its own list.
*
* @param present policy configured with at least one {@code known} name. {@code false} for an
* absent or empty {@code memberCredentials:} block — represented honestly as "no
* policy", never as "nothing blocked" (an empty {@link #blocked} could otherwise be
* misread as a clean bill of health).
* @param policy the normalized policy mode ({@link FleetConfig.MemberCredentials#policy()}), or
* {@code null} when {@link #present} is {@code false}.
* @param known every name the policy declares, in configured order. Names only, never a value.
* @param allowed the subset of {@link #known} explicitly let through. Names only.
* @param blocked {@link #known} minus {@link #allowed} — the names an actual spawn shadows. Names
* only.
*/
public record MemberCredentialPolicyView(boolean present, String policy, List<String> known,
List<String> allowed, List<String> blocked) {
private static final MemberCredentialPolicyView ABSENT =
new MemberCredentialPolicyView(false, null, List.of(), List.of(), List.of());
public MemberCredentialPolicyView {
known = known == null ? List.of() : List.copyOf(known);
allowed = allowed == null ? List.of() : List.copyOf(allowed);
blocked = blocked == null ? List.of() : List.copyOf(blocked);
}
/** The honest "no policy configured" view. */
public static MemberCredentialPolicyView absent() {
return ABSENT;
}
/**
* Build the view straight from the live config. {@code creds} may be {@code null} (no {@code
* memberCredentials:} block at all) — treated the same as a present-but-empty block, exactly
* like {@link dev.ltms.fleet.Fleetd#reportMemberCredentialsGap} already did.
*/
public static MemberCredentialPolicyView of(FleetConfig.MemberCredentials creds) {
if (creds == null || creds.known().isEmpty()) {
return ABSENT;
}
return new MemberCredentialPolicyView(true, creds.policy(), creds.known(), creds.allow(),
List.copyOf(creds.blockedSet()));
}
public int knownCount() {
return known.size();
}
public int allowedCount() {
return allowed.size();
}
public int blockedCount() {
return blocked.size();
}
}
@@ -13,6 +13,7 @@ import dev.ltms.fleet.herdr.Agent;
import dev.ltms.fleet.herdr.HerdrClient;
import dev.ltms.fleet.herdr.HerdrException;
import dev.ltms.fleet.inject.MemberPresence;
import dev.ltms.fleet.member.MemberCredentialPolicyView;
import dev.ltms.fleet.peer.PeerUnreachableException;
import dev.ltms.fleet.placement.PlacementException;
import dev.ltms.fleet.msg.MessageService;
@@ -32,6 +33,7 @@ import java.util.List;
import java.util.Map;
import java.util.function.Function;
import java.util.function.Predicate;
import java.util.function.Supplier;
import java.util.stream.Collectors;
/**
@@ -64,6 +66,10 @@ public final class FleetApp {
private final HttpServlet mcpServlet; // MCP Streamable-HTTP endpoint, mounted at /mcp (nullable)
private final CallerResolver auth; // CB-501: null → authz not enforced (legacy behaviour)
private final Metrics metrics; // CB-502: null → /metrics not exposed
// fleetd #111: re-read per request, same hot-reload shape as every other live config read —
// absent() (the honest "no policy configured" view) for every constructor that does not wire
// a real one, so existing legacy call sites keep building without knowing this field exists.
private final Supplier<MemberCredentialPolicyView> memberCredentials;
private final ObjectMapper mapper = new ObjectMapper();
/**
@@ -111,6 +117,19 @@ public final class FleetApp {
MessageService messages, MemberPresence presence,
HttpServlet mcpServlet, CallerResolver auth, Metrics metrics,
Predicate<String> deliverable) {
this(herdr, memberHerdr, workers, sessions, messages, presence, mcpServlet, auth, metrics,
deliverable, MemberCredentialPolicyView::absent);
}
/**
* @param memberCredentials live {@code memberCredentials:} policy view (fleetd #111), re-read
* per request for {@code GET /member-credentials}; production wiring
* passes the same hot-reload shape as every other live config read
*/
public FleetApp(HerdrClient herdr, HerdrClient memberHerdr, PeerLauncher workers, SessionManager sessions,
MessageService messages, MemberPresence presence,
HttpServlet mcpServlet, CallerResolver auth, Metrics metrics,
Predicate<String> deliverable, Supplier<MemberCredentialPolicyView> memberCredentials) {
this.herdr = herdr;
this.memberHerdr = memberHerdr != null ? memberHerdr : herdr;
this.workers = workers;
@@ -120,6 +139,7 @@ public final class FleetApp {
this.mcpServlet = mcpServlet;
this.auth = auth;
this.metrics = metrics;
this.memberCredentials = memberCredentials != null ? memberCredentials : MemberCredentialPolicyView::absent;
}
/** Wire routes onto a fresh, unstarted Javalin instance. Caller starts it. */
@@ -148,6 +168,7 @@ public final class FleetApp {
app.get("/agents", this::agents);
app.get("/members", this::listMembers); // CB-304: registry roster + live herdr status
app.get("/profiles", this::profiles); // configured backend profiles
app.get("/member-credentials", this::memberCredentials); // fleetd #111: policy names + counts, never a value
app.post("/members", this::spawnMember); // optional ?role=&profile= or {"role":…,"profile":…}
app.delete("/members/{paneId}", this::stopMember);
app.post("/sessions/{id}/message", this::sendMessage); // fleet_send (primary; blocking, wait:false, or answer via turnId)
@@ -346,6 +367,29 @@ public final class FleetApp {
"default", workers.defaultProfile() == null ? "" : workers.defaultProfile()));
}
/**
* fleetd #111 (CB-608): the live {@code memberCredentials:} policy as names and counts —
* NEVER a value. The daemon does not hold a credential's value in the first place (only the
* name it is configured under), so there is nothing to redact here beyond what {@link
* MemberCredentialPolicyView} already omits by construction. This is the source
* {@code scripts/probe-member-credentials.sh} reads instead of carrying its own hardcoded
* name list, which is exactly what let the list drift silently behind the real policy.
*/
private void memberCredentials(Context ctx) {
if (!allow(ctx, Authz.Action.READ, null)) {
return;
}
MemberCredentialPolicyView view = memberCredentials.get();
ctx.status(200).json(Map.of(
"present", view.present(),
"policy", view.policy() == null ? "" : view.policy(),
"known", view.known(),
"allowed", view.allowed(),
"knownCount", view.knownCount(),
"allowedCount", view.allowedCount(),
"blockedCount", view.blockedCount()));
}
/**
* Spawn a guard-checked worker. An optional {@code profile} (query param or {@code {"profile":…}}
* body) picks which configured profile; omitted → the default. 403 if the base_url would breach
@@ -2,6 +2,7 @@ package dev.ltms.fleet.inject;
import com.fasterxml.jackson.databind.JsonNode;
import com.fasterxml.jackson.databind.ObjectMapper;
import dev.ltms.fleet.Fleetd;
import dev.ltms.fleet.config.FleetConfig;
import dev.ltms.fleet.guard.SubscriptionGuard;
import dev.ltms.fleet.herdr.AgentControl;
@@ -146,40 +147,16 @@ class BackendOutageFlowTest {
pushLoop = new ReplyPushLoop(registry, new AgentControl(leadClient), inbox, scheduler, 3, 50);
AtomicReference<ReplyPushLoop> pushLoopRef = new AtomicReference<>(pushLoop);
// --- mirrors Fleetd.main's backendErrorSink lambda EXACTLY: (1) mark BACKEND_ERROR,
// (2) resolve profile/credential via the roster, fail-loud + notify unmapped-target,
// (3) record in BackendOutagePolicy, (4) on a NEW incident, notify the lead. ------------
BackendErrorSink backendErrorSink = (target, matchedLine, reason) -> {
sessions.onBackendError(target, reason);
// fleetd #248 follow-up: this used to be a 30-line hand-copy of Fleetd.main's
// backendErrorSink lambda, with a comment promising it mirrored production "EXACTLY".
// That promise is exactly the problem: a copy proves the copy. Editing or deleting the
// real sink left this whole flow test green, because it never touched the real sink.
// #248 made Fleetd.backendErrorSink public precisely so a cross-package test could
// drive the real object, so this now calls it. Every assertion below is about
// production code again.
BackendErrorSink backendErrorSink =
Fleetd.backendErrorSink(sessions, () -> profiles, outagePolicy, pushLoopRef::get);
String profileName = sessions.roster().stream()
.filter(session -> target.equals(session.terminalId()))
.findFirst()
.map(MemberSession::profile)
.orElse(null);
FleetConfig.Profile profile = profileName == null ? null : profiles.get(profileName);
if (profile == null) {
ReplyPushLoop loop = pushLoopRef.get();
if (loop != null) {
loop.onBackendTargetUnmapped(target, reason);
}
return;
}
String credentialId = profile.effectiveCredentialId();
Optional<BackendOutagePolicy.Incident> incident = outagePolicy.record(credentialId, target, reason);
incident.ifPresent(inc -> {
List<String> affectedProfiles = profiles.values().stream()
.filter(p -> credentialId.equals(p.effectiveCredentialId()))
.map(FleetConfig.Profile::profile)
.sorted()
.toList();
ReplyPushLoop loop = pushLoopRef.get();
if (loop != null) {
loop.onBackendIncident(inc.id(), inc.targets(), credentialId, affectedProfiles,
(int) inc.remainingCoolOffSeconds());
}
});
};
BackendErrorPatternLookup patterns = target -> Pattern.compile("(?i)503 Service Unavailable");
resolver = new CompletionResolver(new AgentControl(herdr), rendezvous, ExhaustedPatternLookup.none(),
ExhaustionSink.none(), patterns, backendErrorSink);
@@ -1409,16 +1409,14 @@ class ClaudeCodeLauncherTest {
}
/**
* fleetd #155 (was CB-633 follow-up (#192) "allowListPolicyOnNonZshKeepsTheWarnWording"): on a
* non-zsh login shell {@code ZDOTDIR} is ignored, so no scrub could ever run there. Before #155
* the launcher degraded to the weaker overlay and spawned anyway; that is exactly the "control
* silently does nothing" failure #155 exists to close, since {@code policy: allow-list} is the
* operator explicitly asking for a control a sourced file cannot undo. Now the spawn is REFUSED
* instead — this test asserts the refusal, naming the shell, on the real spawn path ({@link
* ClaudeCodeLauncher#spawn}), not merely on the launcher method that computes it.
* CB-633 follow-up (#192): the mirror of the zsh test above. On a non-zsh login shell {@code
* ZDOTDIR} is ignored, so no scrub ever runs — the report must keep the WARN wording (a name here
* really is inherited unblocked) rather than claiming a scrub protects it. This is the trap PR
* #174 fell into the other direction: keying the wording on the shell, not on {@code
* creds.isAllowList()}, is what keeps this branch correct.
*/
@Test
void allowListPolicyOnNonZshRefusesTheSpawn() {
void allowListPolicyOnNonZshKeepsTheWarnWording() {
FakeHerdr herdr = new FakeHerdr();
FleetConfig.Profile cfg = new FleetConfig.Profile(
"ltms-local", "http://gx00.gw:8000", "coder", null, "FLEETD_WORKER_TOKEN",
@@ -1432,12 +1430,27 @@ class ClaudeCodeLauncherTest {
0, System::currentTimeMillis, () -> {}, null, () -> creds,
() -> Set.of("AI_GATEWAY_TOKEN", "A_BRAND_NEW_SECRET_TOKEN"));
IllegalArgumentException e = assertThrows(IllegalArgumentException.class, svc::spawn,
"policy=allow-list on a non-zsh shell must refuse the spawn, not silently degrade "
+ "to the weaker overlay");
Logger logger = (Logger) LoggerFactory.getLogger(HerdrPeerLauncher.class);
ListAppender<ILoggingEvent> appender = new ListAppender<>();
appender.start();
logger.addAppender(appender);
try {
svc.spawn();
} finally {
logger.detachAppender(appender);
}
assertTrue(e.getMessage().contains("allow-list") && e.getMessage().contains("/bin/bash"),
"the refusal must name the policy and the actual shell, got: " + e.getMessage());
assertTrue(appender.list.stream().anyMatch(e ->
e.getFormattedMessage().contains("memberCredentials gap")
&& e.getFormattedMessage().contains("UNBLOCKED")
&& e.getFormattedMessage().contains("A_BRAND_NEW_SECRET_TOKEN")),
"no scrub runs on a non-zsh shell, so the WARN wording must be kept — got: "
+ appender.list.stream().map(ILoggingEvent::getFormattedMessage).toList());
assertFalse(appender.list.stream().anyMatch(e ->
e.getFormattedMessage().contains("memberCredentials gap")
&& e.getFormattedMessage().toLowerCase(java.util.Locale.ROOT).contains("scrub")),
"nothing is scrubbed on this path, so the report must not claim otherwise — got: "
+ appender.list.stream().map(ILoggingEvent::getFormattedMessage).toList());
}
/**
@@ -29,7 +29,6 @@ import java.util.function.Supplier;
import static org.junit.jupiter.api.Assertions.assertEquals;
import static org.junit.jupiter.api.Assertions.assertFalse;
import static org.junit.jupiter.api.Assertions.assertNotNull;
import static org.junit.jupiter.api.Assertions.assertThrows;
import static org.junit.jupiter.api.Assertions.assertTrue;
import static org.junit.jupiter.api.Assumptions.assumeTrue;
@@ -100,26 +99,18 @@ class HerdrPeerLauncherAllowListWiringTest {
}
/**
* fleetd #155: a non-zsh shell cannot read {@code ZDOTDIR} at all, so under {@code
* policy: allow-list} — the policy the operator picked specifically for a control a sourced file
* cannot undo — the launcher must REFUSE the spawn rather than silently generate a directory
* nothing will ever read (protection theatre) or fall back to the weaker overlay (exactly the
* "control silently does nothing" defect this ticket exists to close). Real path: through {@link
* HerdrPeerLauncher#spawn}, the method the daemon actually calls.
* A non-zsh shell cannot read {@code ZDOTDIR} at all. The launcher must fall back rather than
* generate a directory nothing will ever read — a directory that would look like protection.
*/
@Test
void aNonZshShellUnderAllowListPolicyRefusesTheSpawn() {
void aNonZshShellGeneratesNothingAndFallsBack() {
FakeHerdr herdr = new FakeHerdr();
WiringLauncher launcher = new WiringLauncher(herdr, allowList(), "/bin/bash");
IllegalArgumentException e = assertThrows(IllegalArgumentException.class,
() -> launcher.spawn(new SpawnRequest("test", null, null, null, null, MemberRole.DEV)),
"policy=allow-list on a non-zsh shell must refuse the spawn, not silently degrade");
launcher.spawn(new SpawnRequest("test", null, null, null, null, MemberRole.DEV));
assertTrue(e.getMessage().contains("allow-list") && e.getMessage().contains("/bin/bash"),
"the refusal must name the policy and the actual shell, got: " + e.getMessage());
assertFalse(launcher.env.containsKey("ZDOTDIR"),
"a refused spawn must not have generated (or wired in) a scrub directory: " + launcher.env);
"bash ignores ZDOTDIR; setting it would be protection theatre");
}
private static Supplier<FleetConfig.MemberCredentials> allowList() {
@@ -255,14 +246,11 @@ class HerdrPeerLauncherAllowListWiringTest {
* "allowed N of M" line — which describes what the scrub does — must not be printed there either.
* Before this fix the line was logged BEFORE the zsh gate, so a non-zsh host printed e.g.
* "allowed 1 of 3" while blocking nothing at all, telling an operator a control ran when it did
* not. fleetd #155: the spawn itself is now refused on this path (see {@code
* aNonZshShellUnderAllowListPolicyRefusesTheSpawn}) rather than falling back — this test's own
* concern still holds under the refusal: the "allowed N of M" line describes a scrub that never
* ran here, so it must still never appear. Real path: goes through {@link
* HerdrPeerLauncher#spawn}, same as the sibling test above, with the shell fixed to bash.
* not. Real path: goes through {@link HerdrPeerLauncher#spawn}, same as the sibling test above,
* with the shell fixed to bash so the fallback branch is the one exercised.
*/
@Test
void noAllowedCountLineIsEmittedOnTheNonZshRefusalPath() {
void noAllowedCountLineIsEmittedOnTheNonZshFallbackPath() {
FakeHerdr herdr = new FakeHerdr();
Set<String> hostEnvNames = Set.of(INJECTED, "SOME_UNRELATED_NAME", "ANOTHER_UNRELATED_NAME");
WiringLauncher launcher = new WiringLauncher(herdr, allowList(), "/bin/bash", () -> hostEnvNames);
@@ -274,9 +262,7 @@ class HerdrPeerLauncherAllowListWiringTest {
appender.start();
logger.addAppender(appender);
try {
assertThrows(IllegalArgumentException.class,
() -> launcher.spawn(new SpawnRequest("test", null, null, null, null, MemberRole.DEV)),
"policy=allow-list on a non-zsh shell must refuse the spawn");
launcher.spawn(new SpawnRequest("test", null, null, null, null, MemberRole.DEV));
} finally {
logger.detachAppender(appender);
logger.setLevel(original);
@@ -324,20 +310,13 @@ class HerdrPeerLauncherAllowListWiringTest {
* different OS user — {@link HerdrPeerLauncher#hostEnvNames} describes fleetd's own process, not
* that user's. Neither "inherits them UNBLOCKED" nor "scrub blanks them" is evidence-backed
* there, so neither may print; the single unknown-environment WARN must, naming the config key.
*
* <p>fleetd #155: {@code memberLoginShell} is now given explicitly as zsh, so the spawn reaches
* this WARN through the (still-degrading, not refusing) missing-{@code worktreeRoot}/{@code
* worktreeGroup} fallback rather than through the zsh gate, which now refuses instead of falling
* back — see {@code aNonZshShellUnderAllowListPolicyRefusesTheSpawn}. This test's own concern
* (the unknown-environment WARN) is orthogonal to which fallback reached {@code
* logCredentialGap}, so it still holds.
*/
@Test
void gapDetectorReportsUnknownInsteadOfAConclusionWhenMemberHerdrSocketIsConfigured() {
FakeHerdr herdr = new FakeHerdr();
Set<String> hostEnvNames = Set.of("FLEETD_WORKER_TOKEN", "SOME_UNKNOWN_SECRET_TOKEN");
WiringLauncher launcher = new WiringLauncher(herdr, allowList(), "/bin/zsh", () -> hostEnvNames,
() -> configWithMemberHerdrSocket("/tmp/other-user-herdr.sock", "/bin/zsh"));
() -> configWithMemberHerdrSocket("/tmp/other-user-herdr.sock"));
List<String> messages = spawnAndCaptureLogs(launcher);
@@ -361,9 +340,8 @@ class HerdrPeerLauncherAllowListWiringTest {
void theUnknownEnvironmentWarnFiresOnceNotOncePerSpawn() {
FakeHerdr herdr = new FakeHerdr();
Set<String> hostEnvNames = Set.of("FLEETD_WORKER_TOKEN", "SOME_UNKNOWN_SECRET_TOKEN");
// fleetd #155: memberLoginShell given explicitly as zsh — see the sibling test above for why.
WiringLauncher launcher = new WiringLauncher(herdr, allowList(), "/bin/zsh", () -> hostEnvNames,
() -> configWithMemberHerdrSocket("/tmp/other-user-herdr.sock", "/bin/zsh"));
() -> configWithMemberHerdrSocket("/tmp/other-user-herdr.sock"));
Logger logger = (Logger) LoggerFactory.getLogger(HerdrPeerLauncher.class);
Level original = logger.getLevel();
@@ -411,44 +389,42 @@ class HerdrPeerLauncherAllowListWiringTest {
}
/**
* fleetd #213 defect 1, acceptance criterion 1 — updated by fleetd #155: {@code
* memberHerdrSocket} configured and {@code memberLoginShell} configured as non-zsh must now
* REFUSE the spawn (not fall back to the sentinel overlay — see {@code
* aNonZshShellUnderAllowListPolicyRefusesTheSpawn}'s javadoc for why). The WiringLauncher's own
* {@code env("SHELL")} is deliberately set to {@code /bin/zsh} — the OPPOSITE of what {@code
* memberLoginShell} says — so a launcher that (incorrectly) fell back to fleetd's own {@code
* $SHELL} here would wrongly pass the gate and fail this test.
* fleetd #213 defect 1, acceptance criterion 1: {@code memberHerdrSocket} configured and {@code
* memberLoginShell} configured as non-zsh must fall back to the sentinel overlay exactly like a
* non-zsh {@code $SHELL} does today — and the "generated ZDOTDIR" INFO must not appear, since no
* scrub actually runs. The WiringLauncher's own {@code env("SHELL")} is deliberately set to
* {@code /bin/zsh} — the OPPOSITE of what {@code memberLoginShell} says — so a launcher that
* (incorrectly) fell back to fleetd's own {@code $SHELL} here would wrongly pass the gate and
* fail this test.
*/
@Test
void memberHerdrSocketWithNonZshMemberLoginShellRefusesTheSpawn() {
void memberHerdrSocketWithNonZshMemberLoginShellFallsBackToTheOverlay() {
FakeHerdr herdr = new FakeHerdr();
WiringLauncher launcher = new WiringLauncher(herdr, allowListWithKnown(List.of("SOME_TOKEN")),
"/bin/zsh", null,
() -> configWithMemberHerdrSocket("/tmp/other-user-herdr.sock", "/bin/bash"));
IllegalArgumentException e = assertThrows(IllegalArgumentException.class,
() -> launcher.spawn(new SpawnRequest("test", null, null, null, null, MemberRole.DEV)),
"a non-zsh configured memberLoginShell under policy=allow-list must refuse the spawn");
List<String> messages = spawnAndCaptureLogs(launcher);
assertTrue(e.getMessage().contains("/bin/bash"),
"the refusal must name the configured memberLoginShell, got: " + e.getMessage());
assertFalse(launcher.env.containsKey("SOME_TOKEN"),
"a refused spawn must not have touched the pane env at all: " + launcher.env);
assertEquals("blocked-by-fleetd-cb596-see-gitea-issue-82", launcher.env.get("SOME_TOKEN"),
"a non-zsh memberLoginShell must fall back to the CB-596 sentinel overlay, exactly "
+ "like a non-zsh $SHELL does when memberHerdrSocket is absent");
assertFalse(launcher.env.containsKey("ZDOTDIR"),
"no scrub directory may be generated when the configured member login shell is not zsh");
assertFalse(messages.stream().anyMatch(m -> m.contains("generated ZDOTDIR")),
"the 'generated ZDOTDIR' INFO must not appear when the scrub never runs — got: " + messages);
}
/**
* fleetd #213 defect 1, acceptance criterion 2 — updated by fleetd #155: {@code
* memberHerdrSocket} configured and NO {@code memberLoginShell} configured must now REFUSE the
* spawn exactly like criterion 1 above — AND fleetd's own {@code $SHELL} must never even be
* consulted (not merely "not decisive"). The fixture's {@code env} function reports {@code
* /bin/zsh} for {@code SHELL} — a value that would WRONGLY pass the zsh gate if the fix
* regressed to reading it — while flagging whether it was ever asked for at all, so this test
* fails loudly on either kind of regression.
* fleetd #213 defect 1, acceptance criterion 2: {@code memberHerdrSocket} configured and NO
* {@code memberLoginShell} configured must fall back exactly like criterion 1 above — AND
* fleetd's own {@code $SHELL} must never even be consulted (not merely "not decisive"). The
* fixture's {@code env} function reports {@code /bin/zsh} for {@code SHELL} — a value that would
* WRONGLY pass the zsh gate if the fix regressed to reading it — while flagging whether it was
* ever asked for at all, so this test fails loudly on either kind of regression.
*/
@Test
void memberHerdrSocketWithNoMemberLoginShellRefusesAndNeverConsultsFleetdsOwnShell() {
void memberHerdrSocketWithNoMemberLoginShellFallsBackAndNeverConsultsFleetdsOwnShell() {
FakeHerdr herdr = new FakeHerdr();
AtomicBoolean shellQueried = new AtomicBoolean(false);
Function<String, String> env = name -> {
@@ -461,18 +437,15 @@ class HerdrPeerLauncherAllowListWiringTest {
WiringLauncher launcher = new WiringLauncher(herdr, allowListWithKnown(List.of("SOME_TOKEN")), env,
() -> configWithMemberHerdrSocket("/tmp/other-user-herdr.sock", null));
assertThrows(IllegalArgumentException.class,
() -> launcher.spawn(new SpawnRequest("test", null, null, null, null, MemberRole.DEV)),
"no configured memberLoginShell under policy=allow-list must refuse the spawn, same "
+ "as an explicit non-zsh one");
List<String> messages = spawnAndCaptureLogs(launcher);
assertFalse(shellQueried.get(), "fleetd's own $SHELL must never be consulted once "
+ "memberHerdrSocket is configured — only memberLoginShell: may decide the gate");
assertFalse(launcher.env.containsKey("SOME_TOKEN"),
"a refused spawn must not have touched the pane env at all, same as a "
assertEquals("blocked-by-fleetd-cb596-see-gitea-issue-82", launcher.env.get("SOME_TOKEN"),
"no memberLoginShell configured must fall back to the sentinel overlay, same as a "
+ "configured non-zsh shell");
assertFalse(launcher.env.containsKey("ZDOTDIR"),
"no scrub may run without a configured memberLoginShell");
assertFalse(messages.stream().anyMatch(m -> m.contains("generated ZDOTDIR")),
"no scrub may run without a configured memberLoginShell — got: " + messages);
}
/**
@@ -727,8 +700,7 @@ class HerdrPeerLauncherAllowListWiringTest {
/**
* fleetd #213: as above, plus {@code worktreeRoot:}/{@code worktreeGroup:} — both required for
* the ZDOTDIR scrub to run at all once {@code memberHerdrSocket} is configured; either missing
* falls back to the sentinel overlay (fleetd #155 left this branch alone — it is not a shell
* problem, so it is not this ticket's refusal).
* falls back to the sentinel overlay, same as a non-zsh {@code memberLoginShell}.
*/
private static FleetConfig configWithMemberHerdrSocketRootAndGroup(String memberHerdrSocket,
String memberLoginShell, String worktreeRoot, String worktreeGroup) {
@@ -0,0 +1,99 @@
package dev.ltms.fleet.member;
import dev.ltms.fleet.config.FleetConfig;
import org.junit.jupiter.api.Test;
import java.util.List;
import static org.junit.jupiter.api.Assertions.assertEquals;
import static org.junit.jupiter.api.Assertions.assertFalse;
import static org.junit.jupiter.api.Assertions.assertTrue;
/**
* fleetd #111 (CB-608): {@link MemberCredentialPolicyView} is the one place that turns a {@code
* memberCredentials:} policy into names-and-counts, so the startup log line and {@code
* GET /member-credentials} cannot drift apart. These tests pin: the counts always match the
* policy that produced them, an absent/empty policy is represented honestly (never as "nothing
* blocked"), and the view carries names only — no value ever flows through it, because it is
* built only from {@link FleetConfig.MemberCredentials}, which itself never holds a value.
*/
class MemberCredentialPolicyViewTest {
@Test
void nullPolicyIsAbsentNotClean() {
MemberCredentialPolicyView view = MemberCredentialPolicyView.of(null);
assertFalse(view.present(), "a null policy must be reported as absent");
assertEquals(0, view.knownCount());
assertEquals(0, view.allowedCount());
assertEquals(0, view.blockedCount());
assertTrue(view.known().isEmpty());
assertTrue(view.allowed().isEmpty());
assertTrue(view.blocked().isEmpty());
}
@Test
void emptyKnownListIsAbsentEvenWithAPolicyModeSet() {
// A memberCredentials: block can be present in YAML with policy: set but known: empty —
// that must still read as "no policy configured", the same as a fully absent block,
// because zero known names means the daemon blocks nothing either way.
FleetConfig.MemberCredentials creds =
new FleetConfig.MemberCredentials("deny-by-default", List.of(), List.of());
MemberCredentialPolicyView view = MemberCredentialPolicyView.of(creds);
assertFalse(view.present());
assertEquals(0, view.knownCount());
}
@Test
void countsMatchARealPolicyExactly() {
FleetConfig.MemberCredentials creds = new FleetConfig.MemberCredentials(
"deny-by-default",
List.of("AI_GATEWAY_TOKEN"),
List.of("AI_GATEWAY_TOKEN", "GITEA_ACCESS_TOKEN", "WORKER_GITEA_TOKEN"));
MemberCredentialPolicyView view = MemberCredentialPolicyView.of(creds);
assertTrue(view.present());
assertEquals("deny-by-default", view.policy());
assertEquals(List.of("AI_GATEWAY_TOKEN", "GITEA_ACCESS_TOKEN", "WORKER_GITEA_TOKEN"), view.known());
assertEquals(List.of("AI_GATEWAY_TOKEN"), view.allowed());
assertEquals(3, view.knownCount());
assertEquals(1, view.allowedCount());
// known minus allowed — the two names actually shadowed on a spawn.
assertEquals(2, view.blockedCount());
assertTrue(view.blocked().containsAll(List.of("GITEA_ACCESS_TOKEN", "WORKER_GITEA_TOKEN")));
}
@Test
void presentPolicyThatBlocksNothingIsStillDistinctFromAbsent() {
// known == allow => blockedCount is 0, exactly like an absent policy's blockedCount — the
// two must still be told apart by `present`, or a reader cannot tell "policy configured,
// nothing currently blocked" from "no policy at all".
FleetConfig.MemberCredentials creds = new FleetConfig.MemberCredentials(
"deny-by-default", List.of("X"), List.of("X"));
MemberCredentialPolicyView view = MemberCredentialPolicyView.of(creds);
assertTrue(view.present());
assertEquals(1, view.knownCount());
assertEquals(0, view.blockedCount());
assertFalse(MemberCredentialPolicyView.absent().present());
}
@Test
void namesPassThroughUnchangedNeverAValue() {
// The view is built only from FleetConfig.MemberCredentials, which itself carries names,
// never values (see its javadoc) — so there is no code path here that could substitute a
// secret's value for its name. This pins the identity: what goes into `known`/`allow` is
// exactly what comes out, character for character.
List<String> known = List.of("SOME_TOKEN_NAME", "ANOTHER_NAME");
FleetConfig.MemberCredentials creds =
new FleetConfig.MemberCredentials("deny-by-default", List.of(), known);
MemberCredentialPolicyView view = MemberCredentialPolicyView.of(creds);
assertEquals(known, view.known());
}
}
+142 -24
View File
@@ -22,8 +22,34 @@
# A prefix of a short secret is most of the secret, and it would end up pasted into a ticket. The
# hash answers every question the prefix was for — is it set, is it the same value as over there,
# is it the CB-592 sentinel — and answers none of the ones it should not.
# * It never writes anywhere, never contacts the network, and never touches secrets.sh, which is
# the operator's file.
# * It never writes anywhere, never contacts the network except the daemon's own REST port (see
# below), and never touches secrets.sh, which is the operator's file.
#
# WHERE THE NAME LIST COMES FROM (fleetd #111 / CB-608)
#
# Earlier versions of this script carried their own hardcoded NAMES array, recorded by hand on
# 2026-08-16. The live memberCredentials: policy in fleetd.yaml grew past that list, and this probe
# never noticed — it kept checking the same 31 names, printed a clean-looking table, and exited 0.
# A verification tool that silently under-reports the thing it verifies is worse than no tool at
# all, because its "clean" output gets taken as proof rather than treated with the suspicion an
# absent tool would get.
#
# The fix is the same one #114 used for the drifted tool catalogue: delete the hand-maintained copy
# rather than update it. This script now fetches the policy's name list from the daemon itself, at
# `GET /member-credentials` (dev.ltms.fleet.member.MemberCredentialPolicyView via FleetApp) — names
# and counts only, the same way the daemon's own startup log line is computed, from the SAME class.
# If fleetd adds a name to memberCredentials.known tomorrow, this probe checks it tomorrow too,
# with no edit here required. There is no local fallback list. See fetch_policy() below for what
# happens when the daemon cannot be reached — it is a hard failure, on purpose (see next section).
#
# WHY AN UNREACHABLE DAEMON IS A HARD FAILURE, NOT A DEGRADED RUN
#
# An empty (or short) name list passes every subset check trivially — a probe that checked zero
# names would print "0 of 0 names are set" and look identical to a clean bill of health. That trap
# has bitten this project twice in one week (see docs/memory — "silent defaults disable features"
# and "a test on the seam does not prove the caller"). So the denominator is guarded explicitly:
# this script refuses to proceed unless it got a policy with at least one known name, and it refuses
# just as hard if the count it fetched does not match the count it is about to check.
#
# HOW TO RUN IT
#
@@ -32,34 +58,22 @@
# 2. For the comparison row, in your OWN shell — a lead, not a member:
# bash scripts/probe-member-credentials.sh --allow-outside-member
#
# Both readings need the daemon's REST port reachable (default http://127.0.0.1:8765; override with
# FLEETD_HOST). That is normally true in every pane this script is meant to run in.
#
# The two outputs side by side are the finding: any name whose hash matches between them is a
# credential the member holds in full.
#
set -uo pipefail
# The names ${SHARED_ENV}/tools/secrets.sh exports, recorded on 2026-08-16 (issue #82). Names only —
# this list contains no values and never should. If secrets.sh gains a name, this list goes stale and
# the probe silently stops asking about it; that staleness is itself part of what #82's criterion 4
# has to solve, so it is called out in the summary rather than hidden.
NAMES=(
AI_GATEWAY_TOKEN BESZEL_ADMIN_EMAIL BESZEL_ADMIN_PASSWORD
BESZEL_HUB_URL BESZEL_KEY BESZEL_UNIVERSAL_TOKEN
BRAIN_MCP_TOKEN CF_ACCOUNT_ID CF_API_TOKEN
CF_USER_TOKEN CONFLUENCE_API_TOKEN CONFLUENCE_USERNAME
CONTEXT7_TOKEN GITEA_HOST GITLAB_OAUTH_CLIENT_SECRET
GITLAB_PERSONAL_ACCESS_TOKEN GRAFANA_ADMIN_PASSWORD GRAFANA_ADMIN_USER
HASS_TOKEN HW_PASSWORD HW_USER
LTMS_API_KEY MEMORY_MCP_TOKEN METRICS_PUSH_TOKEN
OPENCODE_AUTOMODE_MODEL TELEGRAM_BOT_TOKEN TELEGRAM_CHAT_ID
TS_API_KEY TS_AUTHKEY WORKER_GITEA_TOKEN
GITEA_ACCESS_TOKEN
)
FLEETD_HOST="${FLEETD_HOST:-http://127.0.0.1:8765}"
POLICY_URL="${FLEETD_HOST%/}/member-credentials"
allow_outside=0
for arg in "$@"; do
case "$arg" in
--allow-outside-member) allow_outside=1 ;;
-h|--help) sed -n '2,40p' "$0"; exit 0 ;;
-h|--help) sed -n '2,60p' "$0"; exit 0 ;;
*) echo "unknown argument: $arg" >&2; exit 2 ;;
esac
done
@@ -75,6 +89,107 @@ EOF
exit 1
fi
# --- fetch the policy from the daemon (fleetd #111) — no local fallback, ever ------------------
#
# Prefer jq (a real JSON parser); fall back to python3 (present on every host this has run on so
# far); if neither exists, fail loudly rather than guess at the JSON with grep/sed, which is exactly
# the kind of "looks like it worked" degradation this ticket exists to remove.
#
# NOTE: jq's `//` alternative operator treats `false` AND `0` as "missing" and substitutes the
# default — so `.present // empty` silently turns a real `"present": false` into an empty string
# ("unknown"), not the false it actually is. Every extraction below reads its field directly
# instead, so a genuine false/0 is reported as exactly that, not swallowed into "unknown".
if ! command -v jq >/dev/null 2>&1 && ! command -v python3 >/dev/null 2>&1; then
echo "refusing to run: neither jq nor python3 is on PATH, and this probe will not guess at JSON" \
"with grep/sed. Install one of them, or run from a shell that has one." >&2
exit 1
fi
POLICY_JSON="$(curl -fsS --max-time 5 "$POLICY_URL" 2>/dev/null)"
CURL_STATUS=$?
if [ "$CURL_STATUS" -ne 0 ] || [ -z "$POLICY_JSON" ]; then
cat >&2 <<EOF
refusing to run: could not fetch the memberCredentials policy from $POLICY_URL (curl exit $CURL_STATUS).
This probe has NO built-in name list any more (fleetd #111) — it only checks what the live daemon
reports, so an unreachable daemon means it cannot check anything at all. It will not fall back to a
guessed or empty list, because an empty list would pass every check trivially and look clean.
Fix: confirm fleetd is up (curl \${FLEETD_HOST:-http://127.0.0.1:8765}/healthz) and that
FLEETD_HOST (if set) points at it, then re-run.
EOF
exit 1
fi
# One parse pass: line 1 = present (true/false/null), line 2 = policy mode (possibly blank),
# lines 3-5 = knownCount/allowedCount/blockedCount, remaining lines = the known[] names. A single
# pass avoids re-parsing (and re-risking a truthiness bug) five separate times.
if command -v jq >/dev/null 2>&1; then
mapfile -t _FIELDS < <(printf '%s' "$POLICY_JSON" | jq -r '
(.present | tostring),
(.policy // ""),
(.knownCount // 0 | tostring),
(.allowedCount // 0 | tostring),
(.blockedCount // 0 | tostring),
(.known[]? // empty)')
else
mapfile -t _FIELDS < <(printf '%s' "$POLICY_JSON" | python3 - <<'PY'
import json, sys
data = json.load(sys.stdin)
print(str(data.get("present")))
print(data.get("policy") or "")
print(data.get("knownCount") if data.get("knownCount") is not None else 0)
print(data.get("allowedCount") if data.get("allowedCount") is not None else 0)
print(data.get("blockedCount") if data.get("blockedCount") is not None else 0)
for n in (data.get("known") or []):
print(n)
PY
)
fi
PRESENT="${_FIELDS[0]:-null}"
POLICY_MODE="${_FIELDS[1]:-}"
KNOWN_COUNT_REPORTED="${_FIELDS[2]:-0}"
ALLOWED_COUNT_REPORTED="${_FIELDS[3]:-0}"
BLOCKED_COUNT_REPORTED="${_FIELDS[4]:-0}"
NAMES=("${_FIELDS[@]:5}")
# knownCount must be a plain non-negative integer for the arithmetic guard below — a malformed or
# unparseable response must fail loudly, not be coerced into a number that happens to compare true.
case "$KNOWN_COUNT_REPORTED" in
''|*[!0-9]*)
echo "refusing to run: knownCount in the response ('$KNOWN_COUNT_REPORTED') is not a plain" \
"non-negative integer — the response could not be parsed as expected." >&2
exit 1
;;
esac
# --- guard the denominator explicitly — never proceed on a zero/short count ---------------------
#
# This is the exact trap named in the ticket: an empty (or truncated) NAMES array passes every
# subsequent "is it set" check vacuously and prints a table that LOOKS complete. So this is checked
# before anything else runs, with a message that says why, not just that it failed.
if [ "${#NAMES[@]}" -eq 0 ] || [ "$KNOWN_COUNT_REPORTED" -eq 0 ]; then
cat >&2 <<EOF
refusing to run: the policy fetched from $POLICY_URL contains 0 known names (present=${PRESENT:-unknown}).
Either memberCredentials: is absent/empty on the running daemon (nothing is protected — see fleetd's
own startup warning), or the response could not be parsed. Either way, checking zero names would
print a clean-looking table for a policy that protects nothing, or for a probe that read nothing.
This is refused rather than reported as a pass.
EOF
exit 1
fi
if [ "${#NAMES[@]}" -ne "$KNOWN_COUNT_REPORTED" ]; then
cat >&2 <<EOF
refusing to run: the policy reports knownCount=$KNOWN_COUNT_REPORTED but the known[] array this probe
parsed has ${#NAMES[@]} entries. That mismatch means the JSON was not parsed correctly, and this
probe will not check a name list it cannot trust to be complete.
EOF
exit 1
fi
# Prefer sha256sum (Linux), fall back to shasum (macOS). If neither exists, report presence and
# length only — degraded, but never a value.
hasher=""
@@ -95,12 +210,13 @@ else
where="NOT a member — comparison reading only"
fi
echo "CB-596 credential probe"
echo "CB-596 credential probe (fleetd #111: names sourced live from $POLICY_URL)"
echo "reading from : $where"
echo "shell : ${SHELL:-unknown}"
echo "hash : ${hasher:-none available — lengths only}"
# Only printed so the two readings can be told apart when they are pasted side by side.
echo "host : $(hostname 2>/dev/null || echo unknown)"
echo "policy : mode=${POLICY_MODE:-unknown} known=$KNOWN_COUNT_REPORTED allowed=${ALLOWED_COUNT_REPORTED:-?} blocked=${BLOCKED_COUNT_REPORTED:-?}"
echo
printf '%-30s %-7s %6s %s\n' "NAME" "STATE" "LEN" "SHA256-12"
printf '%-30s %-7s %6s %s\n' "------------------------------" "-------" "------" "------------"
@@ -118,6 +234,7 @@ done
echo
echo "$set_count of ${#NAMES[@]} names are set in this shell."
echo "policy contains $KNOWN_COUNT_REPORTED name(s); this run checked ${#NAMES[@]} — they match."
echo
cat <<'EOF'
How to read this:
@@ -129,7 +246,8 @@ How to read this:
most urgent thing on this page.
* AI_GATEWAY_TOKEN matching is expected and correct, not a leak: fleetd.yaml names it in
`tokenEnv:` for the local and gx profiles, so a member reaching the gateway is by design.
* A name that is set here but is NOT in the list above will not appear at all. The list was
recorded on 2026-08-16 and does not update itself. Anything added to secrets.sh since then is
invisible to this probe — which is the same gap issue #82 criterion 4 asks to close properly.
* The name list above is fetched live from the running daemon's memberCredentials: policy
(fleetd #111) — it is never hand-maintained here, so it cannot go stale the way the old
hardcoded list did. If the daemon's policy changes, the next run of this script reflects it
with no edit to this file.
EOF