d59ece6dec
probe-member-credentials.sh used mapfile < <(producer) to parse the fetched policy. That
hides a producer failure three ways: mapfile is bash 4+ and missing on macOS's /bin/bash
3.2, a process substitution's exit status is never propagated to mapfile, and the
downstream reads (":-" defaults and a slice) never fire set -u on a short or unset array.
All three converge on the same "0 known names" refusal, which blames the policy for a
failure that is actually the interpreter or the parser.
Three distinct guards, each closing one cause with its own message:
- a BASH_VERSINFO gate at the top refuses outright on bash < 4 (exit 3)
- the parser's output is captured via command substitution instead of mapfile < <(...),
so a non-zero jq/python3 exit is caught at the call while the fact still exists (exit 4)
- an arity check before the field slice refuses a parse that exits 0 but returns fewer
than 5 fields (exit 5)
The existing "0 known names" guard is now honest: by the time it fires, the three causes
above are already ruled out, so it really does mean the policy has 0 known names.
320 lines
16 KiB
Bash
Executable File
320 lines
16 KiB
Bash
Executable File
#!/usr/bin/env bash
|
|
#
|
|
# CB-596 step 1: measure which credentials a live member actually holds.
|
|
#
|
|
# The claim under test is that a member's herdr pane starts a LOGIN shell, that shell sources
|
|
# ${SHARED_ENV}/tools/secrets.sh, and so a member inherits every name that file exports — while
|
|
# CB-592 blocks exactly one of them (GITEA_ACCESS_TOKEN). That is an inference from the code, not a
|
|
# measurement, and issue #82 says plainly: do not build a fix on the inference. This is the
|
|
# measurement.
|
|
#
|
|
# WHY THIS IS A SCRIPT AND NOT A COMMAND SOMEONE TYPES
|
|
#
|
|
# Enumerating credential names inside a member is exactly the action that should need the operator's
|
|
# explicit approval, and the command classifier refuses it. That refusal is correct. This script is
|
|
# the seam: it is one auditable file the operator can read once, top to bottom, and then run — rather
|
|
# than approving an ad-hoc shell pipeline whose behaviour they have to take on trust.
|
|
#
|
|
# WHAT IT WILL NOT DO
|
|
#
|
|
# * It never prints a credential value, and never any prefix or suffix of one. Not one character.
|
|
# Issue #82's criterion 1 asked for a 6-character prefix; this prints a truncated SHA-256 instead.
|
|
# A prefix of a short secret is most of the secret, and it would end up pasted into a ticket. The
|
|
# hash answers every question the prefix was for — is it set, is it the same value as over there,
|
|
# is it the CB-592 sentinel — and answers none of the ones it should not.
|
|
# * It never writes anywhere, never contacts the network except the daemon's own REST port (see
|
|
# below), and never touches secrets.sh, which is the operator's file.
|
|
#
|
|
# WHERE THE NAME LIST COMES FROM (fleetd #111 / CB-608)
|
|
#
|
|
# Earlier versions of this script carried their own hardcoded NAMES array, recorded by hand on
|
|
# 2026-08-16. The live memberCredentials: policy in fleetd.yaml grew past that list, and this probe
|
|
# never noticed — it kept checking the same 31 names, printed a clean-looking table, and exited 0.
|
|
# A verification tool that silently under-reports the thing it verifies is worse than no tool at
|
|
# all, because its "clean" output gets taken as proof rather than treated with the suspicion an
|
|
# absent tool would get.
|
|
#
|
|
# The fix is the same one #114 used for the drifted tool catalogue: delete the hand-maintained copy
|
|
# rather than update it. This script now fetches the policy's name list from the daemon itself, at
|
|
# `GET /member-credentials` (dev.ltms.fleet.member.MemberCredentialPolicyView via FleetApp) — names
|
|
# and counts only, the same way the daemon's own startup log line is computed, from the SAME class.
|
|
# If fleetd adds a name to memberCredentials.known tomorrow, this probe checks it tomorrow too,
|
|
# with no edit here required. There is no local fallback list. See fetch_policy() below for what
|
|
# happens when the daemon cannot be reached — it is a hard failure, on purpose (see next section).
|
|
#
|
|
# WHY AN UNREACHABLE DAEMON IS A HARD FAILURE, NOT A DEGRADED RUN
|
|
#
|
|
# An empty (or short) name list passes every subset check trivially — a probe that checked zero
|
|
# names would print "0 of 0 names are set" and look identical to a clean bill of health. That trap
|
|
# has bitten this project twice in one week (see docs/memory — "silent defaults disable features"
|
|
# and "a test on the seam does not prove the caller"). So the denominator is guarded explicitly:
|
|
# this script refuses to proceed unless it got a policy with at least one known name, and it refuses
|
|
# just as hard if the count it fetched does not match the count it is about to check.
|
|
#
|
|
# HOW TO RUN IT
|
|
#
|
|
# 1. As the operator, in a member's pane (a spawned worker's terminal):
|
|
# bash scripts/probe-member-credentials.sh
|
|
# 2. For the comparison row, in your OWN shell — a lead, not a member:
|
|
# bash scripts/probe-member-credentials.sh --allow-outside-member
|
|
#
|
|
# Both readings need the daemon's REST port reachable (default http://127.0.0.1:8765; override with
|
|
# FLEETD_HOST). That is normally true in every pane this script is meant to run in.
|
|
#
|
|
# The two outputs side by side are the finding: any name whose hash matches between them is a
|
|
# credential the member holds in full.
|
|
#
|
|
set -uo pipefail
|
|
# `pipefail` is not what catches the parser failure handled below (fleetd #500): in
|
|
# `printf '%s' "$POLICY_JSON" | jq -r '...'`, jq is the LAST element of the pipe, so the pipeline's
|
|
# own exit status is already jq's status, with or without pipefail. It is kept as insurance for if
|
|
# a post-processing stage is ever appended after the parser (e.g. `| tail -n +2`) — at that point
|
|
# the parser would sit upstream and pipefail becomes the only thing that still reports its status.
|
|
|
|
# --- refuse on an interpreter that cannot run this script (fleetd #500) -------------------------
|
|
#
|
|
# mapfile, used below to parse the policy response, was added in bash 4.0. macOS ships bash 3.2.57
|
|
# at /bin/bash, which predates it. This script's own `set -uo pipefail` does not catch a missing
|
|
# mapfile: the builtin just fails with "command not found" on stderr, and every line below that
|
|
# reads the array it would have filled uses a `:-` default or a slice, neither of which `set -u`
|
|
# catches on an unset array. Left unguarded, that chain ends in the "0 known names" refusal further
|
|
# down — a claim about the POLICY, for a failure that is actually about the INTERPRETER. So the
|
|
# interpreter is checked once, explicitly, before it is asked to do anything mapfile depends on.
|
|
if (( ${BASH_VERSINFO[0]} < 4 )); then
|
|
echo "refusing to run: this script uses mapfile, which needs bash 4 or newer. This shell is bash" \
|
|
"${BASH_VERSION:-<unknown, no \$BASH_VERSION>}. Re-run it under a newer bash, for example:" \
|
|
"\"\$(command -v bash)\" \"$0\"" "$@" >&2
|
|
exit 3
|
|
fi
|
|
|
|
FLEETD_HOST="${FLEETD_HOST:-http://127.0.0.1:8765}"
|
|
POLICY_URL="${FLEETD_HOST%/}/member-credentials"
|
|
|
|
allow_outside=0
|
|
for arg in "$@"; do
|
|
case "$arg" in
|
|
--allow-outside-member) allow_outside=1 ;;
|
|
-h|--help) sed -n '2,60p' "$0"; exit 0 ;;
|
|
*) echo "unknown argument: $arg" >&2; exit 2 ;;
|
|
esac
|
|
done
|
|
|
|
if [ "${BRIDGED_MEMBER:-}" != "1" ] && [ "$allow_outside" -eq 0 ]; then
|
|
cat >&2 <<'EOF'
|
|
refusing to run: BRIDGED_MEMBER is not 1, so this is not a member's shell.
|
|
|
|
The finding this probe exists for is what a MEMBER holds. Run it in a spawned worker's pane. If you
|
|
meant to take the comparison reading from your own shell, pass --allow-outside-member and the output
|
|
will be labelled as such.
|
|
EOF
|
|
exit 1
|
|
fi
|
|
|
|
# --- fetch the policy from the daemon (fleetd #111) — no local fallback, ever ------------------
|
|
#
|
|
# Prefer jq (a real JSON parser); fall back to python3 (present on every host this has run on so
|
|
# far); if neither exists, fail loudly rather than guess at the JSON with grep/sed, which is exactly
|
|
# the kind of "looks like it worked" degradation this ticket exists to remove.
|
|
#
|
|
# NOTE: jq's `//` alternative operator treats `false` AND `0` as "missing" and substitutes the
|
|
# default — so `.present // empty` silently turns a real `"present": false` into an empty string
|
|
# ("unknown"), not the false it actually is. Every extraction below reads its field directly
|
|
# instead, so a genuine false/0 is reported as exactly that, not swallowed into "unknown".
|
|
if ! command -v jq >/dev/null 2>&1 && ! command -v python3 >/dev/null 2>&1; then
|
|
echo "refusing to run: neither jq nor python3 is on PATH, and this probe will not guess at JSON" \
|
|
"with grep/sed. Install one of them, or run from a shell that has one." >&2
|
|
exit 1
|
|
fi
|
|
|
|
POLICY_JSON="$(curl -fsS --max-time 5 "$POLICY_URL" 2>/dev/null)"
|
|
CURL_STATUS=$?
|
|
if [ "$CURL_STATUS" -ne 0 ] || [ -z "$POLICY_JSON" ]; then
|
|
cat >&2 <<EOF
|
|
refusing to run: could not fetch the memberCredentials policy from $POLICY_URL (curl exit $CURL_STATUS).
|
|
|
|
This probe has NO built-in name list any more (fleetd #111) — it only checks what the live daemon
|
|
reports, so an unreachable daemon means it cannot check anything at all. It will not fall back to a
|
|
guessed or empty list, because an empty list would pass every check trivially and look clean.
|
|
|
|
Fix: confirm fleetd is up (curl \${FLEETD_HOST:-http://127.0.0.1:8765}/healthz) and that
|
|
FLEETD_HOST (if set) points at it, then re-run.
|
|
EOF
|
|
exit 1
|
|
fi
|
|
|
|
# One parse pass: line 1 = present (true/false/null), line 2 = policy mode (possibly blank),
|
|
# lines 3-5 = knownCount/allowedCount/blockedCount, remaining lines = the known[] names. A single
|
|
# pass avoids re-parsing (and re-risking a truthiness bug) five separate times.
|
|
#
|
|
# This used to feed the parser straight into `mapfile -t _FIELDS < <(producer)`. That form cannot
|
|
# see the producer fail: `<` `<(...)` is a process substitution, not a pipeline, so `set -o
|
|
# pipefail` does not reach inside it, and mapfile's own exit status reports whether the BUILTIN
|
|
# ran, not whether the command substituted into it succeeded — a failing jq or python3 there still
|
|
# leaves mapfile at rc=0 with an empty array, read as a parse that genuinely found nothing (fleetd
|
|
# #500). Capturing the parser's output with command substitution first, and checking ITS exit
|
|
# status, reports the producer's real failure while the fact still exists — before it is handed to
|
|
# mapfile at all.
|
|
#
|
|
# mapfile then reads from that captured string with `<<<` (a herestring), not `< <(...)`: `<<<`
|
|
# materialises the whole string in memory first, where `< <(...)` would stream it. That only
|
|
# matters for a large producer; this one is a short credential-name policy response, so the
|
|
# tradeoff is irrelevant here — noted because it would not be for every producer.
|
|
if command -v jq >/dev/null 2>&1; then
|
|
_FIELDS_RAW="$(printf '%s' "$POLICY_JSON" | jq -r '
|
|
(.present | tostring),
|
|
(.policy // ""),
|
|
(.knownCount // 0 | tostring),
|
|
(.allowedCount // 0 | tostring),
|
|
(.blockedCount // 0 | tostring),
|
|
(.known[]? // empty)')"
|
|
_PARSE_STATUS=$?
|
|
_PARSER_NAME="jq"
|
|
else
|
|
_FIELDS_RAW="$(printf '%s' "$POLICY_JSON" | python3 - <<'PY'
|
|
import json, sys
|
|
data = json.load(sys.stdin)
|
|
print(str(data.get("present")))
|
|
print(data.get("policy") or "")
|
|
print(data.get("knownCount") if data.get("knownCount") is not None else 0)
|
|
print(data.get("allowedCount") if data.get("allowedCount") is not None else 0)
|
|
print(data.get("blockedCount") if data.get("blockedCount") is not None else 0)
|
|
for n in (data.get("known") or []):
|
|
print(n)
|
|
PY
|
|
)"
|
|
_PARSE_STATUS=$?
|
|
_PARSER_NAME="python3"
|
|
fi
|
|
|
|
if [ "$_PARSE_STATUS" -ne 0 ]; then
|
|
echo "refusing to run: could not parse the policy fetched from $POLICY_URL — $_PARSER_NAME exited" \
|
|
"non-zero (status $_PARSE_STATUS). That is a parser failure, not a claim about the policy" \
|
|
"itself; the policy response has not been read." >&2
|
|
exit 4
|
|
fi
|
|
|
|
mapfile -t _FIELDS <<< "$_FIELDS_RAW"
|
|
|
|
# Arity check — the CORRECTNESS fix (fleetd #500). A parser that exits 0 can still return fewer
|
|
# than the 5 fixed fields (present, policy mode, 3 counts) that every line below this expects,
|
|
# whatever the reason: a producer that printed nothing, malformed JSON that jq/python3 still
|
|
# accepted, or a schema change upstream. The slice just below this (`_FIELDS[@]:5`) does not fire
|
|
# `set -u` on an unset OR a short array, and every fixed-field read above used a `:-` default, so
|
|
# without this check a short `_FIELDS` reaches the "0 known names" guard further down with the
|
|
# same look as a policy that genuinely has 0 names. Check the count here, at the one point the
|
|
# fact is still present, before the slice consumes it.
|
|
if (( ${#_FIELDS[@]} < 5 )); then
|
|
echo "refusing to run: the policy parser ($_PARSER_NAME) returned ${#_FIELDS[@]} field(s); at" \
|
|
"least 5 are required (present, policy mode, knownCount, allowedCount, blockedCount). The" \
|
|
"parse ran but its shape is wrong — this is not a claim about how many names the policy" \
|
|
"knows." >&2
|
|
exit 5
|
|
fi
|
|
|
|
PRESENT="${_FIELDS[0]:-null}"
|
|
POLICY_MODE="${_FIELDS[1]:-}"
|
|
KNOWN_COUNT_REPORTED="${_FIELDS[2]:-0}"
|
|
ALLOWED_COUNT_REPORTED="${_FIELDS[3]:-0}"
|
|
BLOCKED_COUNT_REPORTED="${_FIELDS[4]:-0}"
|
|
NAMES=("${_FIELDS[@]:5}")
|
|
|
|
# knownCount must be a plain non-negative integer for the arithmetic guard below — a malformed or
|
|
# unparseable response must fail loudly, not be coerced into a number that happens to compare true.
|
|
case "$KNOWN_COUNT_REPORTED" in
|
|
''|*[!0-9]*)
|
|
echo "refusing to run: knownCount in the response ('$KNOWN_COUNT_REPORTED') is not a plain" \
|
|
"non-negative integer — the response could not be parsed as expected." >&2
|
|
exit 1
|
|
;;
|
|
esac
|
|
|
|
# --- guard the denominator explicitly — never proceed on a zero count ---------------------------
|
|
#
|
|
# This is the exact trap named in the ticket: an empty NAMES array passes every subsequent "is it
|
|
# set" check vacuously and prints a table that LOOKS complete. By this point the interpreter gate,
|
|
# the parser-exit-status check, and the arity check above have already ruled out "the interpreter
|
|
# couldn't run mapfile", "the parser failed", and "the parser returned the wrong shape" — so a zero
|
|
# count reaching here really does mean the policy itself reports 0 known names, not a swallowed
|
|
# failure upstream. That is still checked before anything else runs, with a message that says so.
|
|
if [ "${#NAMES[@]}" -eq 0 ] || [ "$KNOWN_COUNT_REPORTED" -eq 0 ]; then
|
|
cat >&2 <<EOF
|
|
refusing to run: the policy fetched from $POLICY_URL contains 0 known names (present=${PRESENT:-unknown}).
|
|
|
|
memberCredentials: is absent or empty on the running daemon — nothing is protected (see fleetd's own
|
|
startup warning). Checking zero names would print a clean-looking table for a policy that protects
|
|
nothing. This is refused rather than reported as a pass.
|
|
EOF
|
|
exit 1
|
|
fi
|
|
|
|
if [ "${#NAMES[@]}" -ne "$KNOWN_COUNT_REPORTED" ]; then
|
|
cat >&2 <<EOF
|
|
refusing to run: the policy reports knownCount=$KNOWN_COUNT_REPORTED but the known[] array this probe
|
|
parsed has ${#NAMES[@]} entries. That mismatch means the JSON was not parsed correctly, and this
|
|
probe will not check a name list it cannot trust to be complete.
|
|
EOF
|
|
exit 1
|
|
fi
|
|
|
|
# Prefer sha256sum (Linux), fall back to shasum (macOS). If neither exists, report presence and
|
|
# length only — degraded, but never a value.
|
|
hasher=""
|
|
if command -v sha256sum >/dev/null 2>&1; then
|
|
hasher="sha256sum"
|
|
elif command -v shasum >/dev/null 2>&1; then
|
|
hasher="shasum -a 256"
|
|
fi
|
|
|
|
digest() { # value -> first 12 hex chars of its sha256, or "-" when no hasher is available
|
|
[ -z "$hasher" ] && { printf '%s' "-"; return; }
|
|
printf '%s' "$1" | $hasher | cut -c1-12
|
|
}
|
|
|
|
if [ "${BRIDGED_MEMBER:-}" = "1" ]; then
|
|
where="MEMBER (BRIDGED_MEMBER=1)"
|
|
else
|
|
where="NOT a member — comparison reading only"
|
|
fi
|
|
|
|
echo "CB-596 credential probe (fleetd #111: names sourced live from $POLICY_URL)"
|
|
echo "reading from : $where"
|
|
echo "shell : ${SHELL:-unknown}"
|
|
echo "hash : ${hasher:-none available — lengths only}"
|
|
# Only printed so the two readings can be told apart when they are pasted side by side.
|
|
echo "host : $(hostname 2>/dev/null || echo unknown)"
|
|
echo "policy : mode=${POLICY_MODE:-unknown} known=$KNOWN_COUNT_REPORTED allowed=${ALLOWED_COUNT_REPORTED:-?} blocked=${BLOCKED_COUNT_REPORTED:-?}"
|
|
echo
|
|
printf '%-30s %-7s %6s %s\n' "NAME" "STATE" "LEN" "SHA256-12"
|
|
printf '%-30s %-7s %6s %s\n' "------------------------------" "-------" "------" "------------"
|
|
|
|
set_count=0
|
|
for name in "${NAMES[@]}"; do
|
|
value="${!name:-}"
|
|
if [ -z "$value" ]; then
|
|
printf '%-30s %-7s %6s %s\n' "$name" "unset" "-" "-"
|
|
else
|
|
set_count=$((set_count + 1))
|
|
printf '%-30s %-7s %6s %s\n' "$name" "SET" "${#value}" "$(digest "$value")"
|
|
fi
|
|
done
|
|
|
|
echo
|
|
echo "$set_count of ${#NAMES[@]} names are set in this shell."
|
|
echo "policy contains $KNOWN_COUNT_REPORTED name(s); this run checked ${#NAMES[@]} — they match."
|
|
echo
|
|
cat <<'EOF'
|
|
How to read this:
|
|
|
|
* Take the MEMBER reading and the comparison reading, and line them up. A name whose SHA256-12
|
|
matches on both sides is a credential the member holds in full. That is the finding.
|
|
* GITEA_ACCESS_TOKEN is the control. CB-592 replaces it with a blocked sentinel, so its hash
|
|
should DIFFER between the two readings. If it matches, CB-592 is not working and that is the
|
|
most urgent thing on this page.
|
|
* AI_GATEWAY_TOKEN matching is expected and correct, not a leak: fleetd.yaml names it in
|
|
`tokenEnv:` for the local and gx profiles, so a member reaching the gateway is by design.
|
|
* The name list above is fetched live from the running daemon's memberCredentials: policy
|
|
(fleetd #111) — it is never hand-maintained here, so it cannot go stale the way the old
|
|
hardcoded list did. If the daemon's policy changes, the next run of this script reflects it
|
|
with no edit to this file.
|
|
EOF
|