Files
fleetd/scripts/probe-member-credentials.sh
Dai Ha 51f7b0a3ca
CI / contract (pull_request) Successful in 1m24s
CI / build (pull_request) Successful in 2m22s
fleetd #111: probe reads the live memberCredentials policy, no hardcoded name list
scripts/probe-member-credentials.sh carried its own hand-maintained NAMES array
(31 names, recorded 2026-08-16), so a name added later to fleetd.yaml's
memberCredentials.known was never checked and the probe still exited 0 with a
clean-looking table. Same drift shape as #114's tool catalogue.

- New dev.ltms.fleet.member.MemberCredentialPolicyView: the single place that
  turns a MemberCredentials policy into names + counts (never a value). Reused
  by Fleetd.reportMemberCredentialsGap (startup log line) and by the new
  GET /member-credentials REST endpoint (FleetApp), so the two can no longer
  drift apart the way the probe and the policy did.
- FleetApp gains one route + handler + a Supplier<MemberCredentialPolicyView>
  constructor param (legacy constructors default to ::absent, so existing call
  sites are unaffected).
- probe-member-credentials.sh now fetches its name list from
  GET /member-credentials instead of carrying one. No local fallback: an
  unreachable daemon, an empty/absent policy, or a knownCount/known[] length
  mismatch all refuse with a non-zero exit rather than silently checking zero
  names. Prints "policy contains N; this run checked N" so the two numbers are
  visibly equal.
2026-09-03 13:21:35 +07:00

254 lines
12 KiB
Bash
Executable File

#!/usr/bin/env bash
#
# CB-596 step 1: measure which credentials a live member actually holds.
#
# The claim under test is that a member's herdr pane starts a LOGIN shell, that shell sources
# ${SHARED_ENV}/tools/secrets.sh, and so a member inherits every name that file exports — while
# CB-592 blocks exactly one of them (GITEA_ACCESS_TOKEN). That is an inference from the code, not a
# measurement, and issue #82 says plainly: do not build a fix on the inference. This is the
# measurement.
#
# WHY THIS IS A SCRIPT AND NOT A COMMAND SOMEONE TYPES
#
# Enumerating credential names inside a member is exactly the action that should need the operator's
# explicit approval, and the command classifier refuses it. That refusal is correct. This script is
# the seam: it is one auditable file the operator can read once, top to bottom, and then run — rather
# than approving an ad-hoc shell pipeline whose behaviour they have to take on trust.
#
# WHAT IT WILL NOT DO
#
# * It never prints a credential value, and never any prefix or suffix of one. Not one character.
# Issue #82's criterion 1 asked for a 6-character prefix; this prints a truncated SHA-256 instead.
# A prefix of a short secret is most of the secret, and it would end up pasted into a ticket. The
# hash answers every question the prefix was for — is it set, is it the same value as over there,
# is it the CB-592 sentinel — and answers none of the ones it should not.
# * It never writes anywhere, never contacts the network except the daemon's own REST port (see
# below), and never touches secrets.sh, which is the operator's file.
#
# WHERE THE NAME LIST COMES FROM (fleetd #111 / CB-608)
#
# Earlier versions of this script carried their own hardcoded NAMES array, recorded by hand on
# 2026-08-16. The live memberCredentials: policy in fleetd.yaml grew past that list, and this probe
# never noticed — it kept checking the same 31 names, printed a clean-looking table, and exited 0.
# A verification tool that silently under-reports the thing it verifies is worse than no tool at
# all, because its "clean" output gets taken as proof rather than treated with the suspicion an
# absent tool would get.
#
# The fix is the same one #114 used for the drifted tool catalogue: delete the hand-maintained copy
# rather than update it. This script now fetches the policy's name list from the daemon itself, at
# `GET /member-credentials` (dev.ltms.fleet.member.MemberCredentialPolicyView via FleetApp) — names
# and counts only, the same way the daemon's own startup log line is computed, from the SAME class.
# If fleetd adds a name to memberCredentials.known tomorrow, this probe checks it tomorrow too,
# with no edit here required. There is no local fallback list. See fetch_policy() below for what
# happens when the daemon cannot be reached — it is a hard failure, on purpose (see next section).
#
# WHY AN UNREACHABLE DAEMON IS A HARD FAILURE, NOT A DEGRADED RUN
#
# An empty (or short) name list passes every subset check trivially — a probe that checked zero
# names would print "0 of 0 names are set" and look identical to a clean bill of health. That trap
# has bitten this project twice in one week (see docs/memory — "silent defaults disable features"
# and "a test on the seam does not prove the caller"). So the denominator is guarded explicitly:
# this script refuses to proceed unless it got a policy with at least one known name, and it refuses
# just as hard if the count it fetched does not match the count it is about to check.
#
# HOW TO RUN IT
#
# 1. As the operator, in a member's pane (a spawned worker's terminal):
# bash scripts/probe-member-credentials.sh
# 2. For the comparison row, in your OWN shell — a lead, not a member:
# bash scripts/probe-member-credentials.sh --allow-outside-member
#
# Both readings need the daemon's REST port reachable (default http://127.0.0.1:8765; override with
# FLEETD_HOST). That is normally true in every pane this script is meant to run in.
#
# The two outputs side by side are the finding: any name whose hash matches between them is a
# credential the member holds in full.
#
set -uo pipefail
FLEETD_HOST="${FLEETD_HOST:-http://127.0.0.1:8765}"
POLICY_URL="${FLEETD_HOST%/}/member-credentials"
allow_outside=0
for arg in "$@"; do
case "$arg" in
--allow-outside-member) allow_outside=1 ;;
-h|--help) sed -n '2,60p' "$0"; exit 0 ;;
*) echo "unknown argument: $arg" >&2; exit 2 ;;
esac
done
if [ "${BRIDGED_MEMBER:-}" != "1" ] && [ "$allow_outside" -eq 0 ]; then
cat >&2 <<'EOF'
refusing to run: BRIDGED_MEMBER is not 1, so this is not a member's shell.
The finding this probe exists for is what a MEMBER holds. Run it in a spawned worker's pane. If you
meant to take the comparison reading from your own shell, pass --allow-outside-member and the output
will be labelled as such.
EOF
exit 1
fi
# --- fetch the policy from the daemon (fleetd #111) — no local fallback, ever ------------------
#
# Prefer jq (a real JSON parser); fall back to python3 (present on every host this has run on so
# far); if neither exists, fail loudly rather than guess at the JSON with grep/sed, which is exactly
# the kind of "looks like it worked" degradation this ticket exists to remove.
#
# NOTE: jq's `//` alternative operator treats `false` AND `0` as "missing" and substitutes the
# default — so `.present // empty` silently turns a real `"present": false` into an empty string
# ("unknown"), not the false it actually is. Every extraction below reads its field directly
# instead, so a genuine false/0 is reported as exactly that, not swallowed into "unknown".
if ! command -v jq >/dev/null 2>&1 && ! command -v python3 >/dev/null 2>&1; then
echo "refusing to run: neither jq nor python3 is on PATH, and this probe will not guess at JSON" \
"with grep/sed. Install one of them, or run from a shell that has one." >&2
exit 1
fi
POLICY_JSON="$(curl -fsS --max-time 5 "$POLICY_URL" 2>/dev/null)"
CURL_STATUS=$?
if [ "$CURL_STATUS" -ne 0 ] || [ -z "$POLICY_JSON" ]; then
cat >&2 <<EOF
refusing to run: could not fetch the memberCredentials policy from $POLICY_URL (curl exit $CURL_STATUS).
This probe has NO built-in name list any more (fleetd #111) — it only checks what the live daemon
reports, so an unreachable daemon means it cannot check anything at all. It will not fall back to a
guessed or empty list, because an empty list would pass every check trivially and look clean.
Fix: confirm fleetd is up (curl \${FLEETD_HOST:-http://127.0.0.1:8765}/healthz) and that
FLEETD_HOST (if set) points at it, then re-run.
EOF
exit 1
fi
# One parse pass: line 1 = present (true/false/null), line 2 = policy mode (possibly blank),
# lines 3-5 = knownCount/allowedCount/blockedCount, remaining lines = the known[] names. A single
# pass avoids re-parsing (and re-risking a truthiness bug) five separate times.
if command -v jq >/dev/null 2>&1; then
mapfile -t _FIELDS < <(printf '%s' "$POLICY_JSON" | jq -r '
(.present | tostring),
(.policy // ""),
(.knownCount // 0 | tostring),
(.allowedCount // 0 | tostring),
(.blockedCount // 0 | tostring),
(.known[]? // empty)')
else
mapfile -t _FIELDS < <(printf '%s' "$POLICY_JSON" | python3 - <<'PY'
import json, sys
data = json.load(sys.stdin)
print(str(data.get("present")))
print(data.get("policy") or "")
print(data.get("knownCount") if data.get("knownCount") is not None else 0)
print(data.get("allowedCount") if data.get("allowedCount") is not None else 0)
print(data.get("blockedCount") if data.get("blockedCount") is not None else 0)
for n in (data.get("known") or []):
print(n)
PY
)
fi
PRESENT="${_FIELDS[0]:-null}"
POLICY_MODE="${_FIELDS[1]:-}"
KNOWN_COUNT_REPORTED="${_FIELDS[2]:-0}"
ALLOWED_COUNT_REPORTED="${_FIELDS[3]:-0}"
BLOCKED_COUNT_REPORTED="${_FIELDS[4]:-0}"
NAMES=("${_FIELDS[@]:5}")
# knownCount must be a plain non-negative integer for the arithmetic guard below — a malformed or
# unparseable response must fail loudly, not be coerced into a number that happens to compare true.
case "$KNOWN_COUNT_REPORTED" in
''|*[!0-9]*)
echo "refusing to run: knownCount in the response ('$KNOWN_COUNT_REPORTED') is not a plain" \
"non-negative integer — the response could not be parsed as expected." >&2
exit 1
;;
esac
# --- guard the denominator explicitly — never proceed on a zero/short count ---------------------
#
# This is the exact trap named in the ticket: an empty (or truncated) NAMES array passes every
# subsequent "is it set" check vacuously and prints a table that LOOKS complete. So this is checked
# before anything else runs, with a message that says why, not just that it failed.
if [ "${#NAMES[@]}" -eq 0 ] || [ "$KNOWN_COUNT_REPORTED" -eq 0 ]; then
cat >&2 <<EOF
refusing to run: the policy fetched from $POLICY_URL contains 0 known names (present=${PRESENT:-unknown}).
Either memberCredentials: is absent/empty on the running daemon (nothing is protected — see fleetd's
own startup warning), or the response could not be parsed. Either way, checking zero names would
print a clean-looking table for a policy that protects nothing, or for a probe that read nothing.
This is refused rather than reported as a pass.
EOF
exit 1
fi
if [ "${#NAMES[@]}" -ne "$KNOWN_COUNT_REPORTED" ]; then
cat >&2 <<EOF
refusing to run: the policy reports knownCount=$KNOWN_COUNT_REPORTED but the known[] array this probe
parsed has ${#NAMES[@]} entries. That mismatch means the JSON was not parsed correctly, and this
probe will not check a name list it cannot trust to be complete.
EOF
exit 1
fi
# Prefer sha256sum (Linux), fall back to shasum (macOS). If neither exists, report presence and
# length only — degraded, but never a value.
hasher=""
if command -v sha256sum >/dev/null 2>&1; then
hasher="sha256sum"
elif command -v shasum >/dev/null 2>&1; then
hasher="shasum -a 256"
fi
digest() { # value -> first 12 hex chars of its sha256, or "-" when no hasher is available
[ -z "$hasher" ] && { printf '%s' "-"; return; }
printf '%s' "$1" | $hasher | cut -c1-12
}
if [ "${BRIDGED_MEMBER:-}" = "1" ]; then
where="MEMBER (BRIDGED_MEMBER=1)"
else
where="NOT a member — comparison reading only"
fi
echo "CB-596 credential probe (fleetd #111: names sourced live from $POLICY_URL)"
echo "reading from : $where"
echo "shell : ${SHELL:-unknown}"
echo "hash : ${hasher:-none available — lengths only}"
# Only printed so the two readings can be told apart when they are pasted side by side.
echo "host : $(hostname 2>/dev/null || echo unknown)"
echo "policy : mode=${POLICY_MODE:-unknown} known=$KNOWN_COUNT_REPORTED allowed=${ALLOWED_COUNT_REPORTED:-?} blocked=${BLOCKED_COUNT_REPORTED:-?}"
echo
printf '%-30s %-7s %6s %s\n' "NAME" "STATE" "LEN" "SHA256-12"
printf '%-30s %-7s %6s %s\n' "------------------------------" "-------" "------" "------------"
set_count=0
for name in "${NAMES[@]}"; do
value="${!name:-}"
if [ -z "$value" ]; then
printf '%-30s %-7s %6s %s\n' "$name" "unset" "-" "-"
else
set_count=$((set_count + 1))
printf '%-30s %-7s %6s %s\n' "$name" "SET" "${#value}" "$(digest "$value")"
fi
done
echo
echo "$set_count of ${#NAMES[@]} names are set in this shell."
echo "policy contains $KNOWN_COUNT_REPORTED name(s); this run checked ${#NAMES[@]} — they match."
echo
cat <<'EOF'
How to read this:
* Take the MEMBER reading and the comparison reading, and line them up. A name whose SHA256-12
matches on both sides is a credential the member holds in full. That is the finding.
* GITEA_ACCESS_TOKEN is the control. CB-592 replaces it with a blocked sentinel, so its hash
should DIFFER between the two readings. If it matches, CB-592 is not working and that is the
most urgent thing on this page.
* AI_GATEWAY_TOKEN matching is expected and correct, not a leak: fleetd.yaml names it in
`tokenEnv:` for the local and gx profiles, so a member reaching the gateway is by design.
* The name list above is fetched live from the running daemon's memberCredentials: policy
(fleetd #111) — it is never hand-maintained here, so it cannot go stale the way the old
hardcoded list did. If the daemon's policy changes, the next run of this script reflects it
with no edit to this file.
EOF