shasum is macOS-only: on Linux redeploy-fleetd.sh reports an existing jar as "absent" with exit 0, and the shell suite dies at 127 looking green #550

Closed
opened 2026-09-12 09:34:36 +02:00 by ltms · 3 comments
Owner

Found while verifying #545 (PR #548) under GNU coreutils. Same root cause as #545 — a macOS-only
tool used with no fallback — but a different tool, a different failure shape, and one of the two
hits the production script, not the tests.

All numbers below measured today on main at 0b032f5, in a debian:bookworm-slim container
(GNU coreutils 9.1). Debian slim ships sha256sum and does not ship shasum.

Item 1 (production script, and the worst of the three): jar_id says "absent" for a jar that exists

scripts/redeploy-fleetd.sh:147:

jar_id() { local f="${1:-$JAR}"; [ -f "$f" ] && shasum -a 256 "$f" | cut -c1-12 || echo "absent"; }

Run on Linux against a file that really is there:

shasum present?    NO
sha256sum present? /usr/bin/sha256sum
file exists?       yes
jar_id => 'absent'  rc=0
correct answer would be: 5891b5b522d5

It does not error. It answers "absent", and it exits 0. shasum is not found, cut
succeeds on empty input, pipefail makes the pipeline non-zero, and the || echo "absent" arm
then prints a confident wrong answer.

This is the a-sentinel-that-conflates-no-with-cannot-tell shape: one word, "absent", is carrying
two states that need opposite handling — the jar is not there and I could not hash it. The
caller reads the confident one.

Why it matters more than a cosmetic wrong string: the redeploy procedure tells the lead to trust
exactly this output. From the redeploy-fleetd skill:

Run --check first: it prints the jar's hash and its modification time, so you can see for
yourself whether the jar is missing or older than the code you mean to ship.

and protection 5:

Prove the new jar is the one running.

On a Linux host, --check would tell a lead the jar is missing when it is present, and the lead
has been told to believe it. The likely next action is a rebuild, or worse, a report that the
deployment is broken when it is fine.

The repo already contains the correct idiom. scripts/probe-member-credentials.sh:273-279:

# Prefer sha256sum (Linux), fall back to shasum (macOS). If neither exists, report presence and
...
elif command -v shasum >/dev/null 2>&1; then
  hasher="shasum -a 256"

So this is not a new design problem — it is one file not using the helper shape another file
already worked out. Note that file also gets the third state right ("report presence" when neither
hasher exists), which is what jar_id must do too: a distinct value for could not hash, never
"absent".

Item 2: the shell suite is dead on Linux and reports zero failures

scripts/test-redeploy-fleetd.sh:298-299 also call shasum. Under set -euo pipefail the suite
dies there:

scripts/test-redeploy-fleetd.sh: line 298: shasum: command not found
suite exit=127
anchored ^FAIL: 0

Exit 127, and grep -c '^FAIL:' reads 0 — byte-identical to a green run on the channel anyone
would check. A counter of bad things reads zero when the counter never ran. The exit code is the
pass signal here; the failure count is a detail that means nothing on its own.

I measured this on both arms of the #545 comparison, so it is not something #545 introduced or
fixed. It is independent and older.

Item 3: nothing runs this suite in CI, on any platform

$ grep -rn 'test-redeploy' .gitea/workflows/
(no output)
$ grep -n 'run:' .gitea/workflows/ci.yml
40:        run: mvn -B clean install
44:        run: mvn -B -DskipTests javadoc:javadoc -Ddoclint=reference
105:        run: mvn -B -Pcontract test -Dgroups=contract

Both CI jobs are runs-on: ubuntu-latest. The shell suite — 67 tests as of #548 — runs only when a
lead runs it by hand, and every lead who has run it has run it on macOS. That is why items 1 and 2
survived: the only platform that exercises this code is the one platform where both bugs are
invisible.

Item 3 is also why fixing items 1 and 2 without it would be temporary. The next macOS-only idiom
goes in the same way.

Suggested fix

  1. One hash256 <file> helper in scripts/redeploy-fleetd.sh, shaped like
    probe-member-credentials.sh:273-279: prefer sha256sum, fall back to shasum -a 256, and
    return a third, distinct answer when neither exists. jar_id must never say "absent" for a
    file it could not hash — [ -f "$f" ] already answered the existence question, so the two
    states are separable with no new probe.
  2. Use the same helper in scripts/test-redeploy-fleetd.sh.
  3. Add a CI job that runs bash scripts/test-redeploy-fleetd.sh on ubuntu-latest, and fail the
    job on a non-zero exit — not on a FAIL: count. Item 2 is the proof that the count alone cannot
    be the gate.

Acceptance

  • A test that fails if any script under scripts/ calls shasum without a sha256sum fallback —
    the shape, not the named lines. #545 taught this: the idiom had propagated from two sites to
    six across 91 commits on the fleet01 lead's tree, so an acceptance that names lines just invites
    a seventh.
  • A test that jar_id returns a value distinct from both a real hash and absent when no hasher
    is on PATH, driven through a stubbed PATH (the same stub technique #548 used for mktemp).
  • The CI job must be shown failing on a deliberately broken suite, not only passing on a good one.
    A green CI job that would stay green with the suite deleted is worth nothing here — that is item
    2's whole lesson, one level out.
  • For each new test: apply a mutation, show it red with that test's own message, restore, confirm
    byte-identical with the full shasum -a 256 (on macOS), run a green control, and run the proof
    cell against the un-mutated tree first to confirm it reports not-applied.

Credit

The fleet01 lead pushed the generalisation that makes item 2 a defect rather than a curiosity:

A COUNT OF BAD THINGS IS A PASS SIGNAL ONLY ONCE YOU HAVE SEPARATELY ESTABLISHED THAT THE
PRODUCER RAN. Exit 127 with zero FAIL lines, and my rc=0 from head, are the same shape — a
status channel reporting on the wrong subject while stderr carried the truth.

Related: #545 (the same class, mktemp), #497, #513.

Found while verifying #545 (PR #548) under GNU coreutils. Same root cause as #545 — a macOS-only tool used with no fallback — but a different tool, a different failure shape, and one of the two hits the **production** script, not the tests. All numbers below measured today on `main` at `0b032f5`, in a `debian:bookworm-slim` container (GNU coreutils 9.1). Debian slim ships `sha256sum` and does **not** ship `shasum`. ## Item 1 (production script, and the worst of the three): `jar_id` says "absent" for a jar that exists `scripts/redeploy-fleetd.sh:147`: ```bash jar_id() { local f="${1:-$JAR}"; [ -f "$f" ] && shasum -a 256 "$f" | cut -c1-12 || echo "absent"; } ``` Run on Linux against a file that really is there: ``` shasum present? NO sha256sum present? /usr/bin/sha256sum file exists? yes jar_id => 'absent' rc=0 correct answer would be: 5891b5b522d5 ``` It does not error. It answers **"absent"**, and it exits **0**. `shasum` is not found, `cut` succeeds on empty input, `pipefail` makes the pipeline non-zero, and the `|| echo "absent"` arm then prints a confident wrong answer. This is the `a-sentinel-that-conflates-no-with-cannot-tell` shape: one word, "absent", is carrying two states that need opposite handling — *the jar is not there* and *I could not hash it*. The caller reads the confident one. Why it matters more than a cosmetic wrong string: the redeploy procedure tells the lead to trust exactly this output. From the `redeploy-fleetd` skill: > Run `--check` first: it prints the jar's hash and its modification time, so you can see for > yourself whether the jar is missing or older than the code you mean to ship. and protection 5: > **Prove the new jar is the one running.** On a Linux host, `--check` would tell a lead the jar is missing when it is present, and the lead has been told to believe it. The likely next action is a rebuild, or worse, a report that the deployment is broken when it is fine. **The repo already contains the correct idiom.** `scripts/probe-member-credentials.sh:273-279`: ```bash # Prefer sha256sum (Linux), fall back to shasum (macOS). If neither exists, report presence and ... elif command -v shasum >/dev/null 2>&1; then hasher="shasum -a 256" ``` So this is not a new design problem — it is one file not using the helper shape another file already worked out. Note that file also gets the third state right ("report presence" when neither hasher exists), which is what `jar_id` must do too: a distinct value for *could not hash*, never "absent". ## Item 2: the shell suite is dead on Linux and reports zero failures `scripts/test-redeploy-fleetd.sh:298-299` also call `shasum`. Under `set -euo pipefail` the suite dies there: ``` scripts/test-redeploy-fleetd.sh: line 298: shasum: command not found suite exit=127 anchored ^FAIL: 0 ``` **Exit 127, and `grep -c '^FAIL:'` reads 0 — byte-identical to a green run** on the channel anyone would check. A counter of bad things reads zero when the counter never ran. The exit code is the pass signal here; the failure count is a detail that means nothing on its own. I measured this on both arms of the #545 comparison, so it is not something #545 introduced or fixed. It is independent and older. ## Item 3: nothing runs this suite in CI, on any platform ``` $ grep -rn 'test-redeploy' .gitea/workflows/ (no output) $ grep -n 'run:' .gitea/workflows/ci.yml 40: run: mvn -B clean install 44: run: mvn -B -DskipTests javadoc:javadoc -Ddoclint=reference 105: run: mvn -B -Pcontract test -Dgroups=contract ``` Both CI jobs are `runs-on: ubuntu-latest`. The shell suite — 67 tests as of #548 — runs only when a lead runs it by hand, and every lead who has run it has run it on macOS. That is why items 1 and 2 survived: the only platform that exercises this code is the one platform where both bugs are invisible. Item 3 is also why fixing items 1 and 2 without it would be temporary. The next macOS-only idiom goes in the same way. ## Suggested fix 1. One `hash256 <file>` helper in `scripts/redeploy-fleetd.sh`, shaped like `probe-member-credentials.sh:273-279`: prefer `sha256sum`, fall back to `shasum -a 256`, and return a **third, distinct** answer when neither exists. `jar_id` must never say "absent" for a file it could not hash — `[ -f "$f" ]` already answered the existence question, so the two states are separable with no new probe. 2. Use the same helper in `scripts/test-redeploy-fleetd.sh`. 3. Add a CI job that runs `bash scripts/test-redeploy-fleetd.sh` on `ubuntu-latest`, and fail the job on a non-zero exit — not on a `FAIL:` count. Item 2 is the proof that the count alone cannot be the gate. ## Acceptance - A test that fails if any script under `scripts/` calls `shasum` without a `sha256sum` fallback — the **shape**, not the named lines. #545 taught this: the idiom had propagated from two sites to six across 91 commits on the fleet01 lead's tree, so an acceptance that names lines just invites a seventh. - A test that `jar_id` returns a value distinct from both a real hash and `absent` when no hasher is on `PATH`, driven through a stubbed `PATH` (the same stub technique #548 used for `mktemp`). - The CI job must be shown failing on a deliberately broken suite, not only passing on a good one. A green CI job that would stay green with the suite deleted is worth nothing here — that is item 2's whole lesson, one level out. - For each new test: apply a mutation, show it red with that test's own message, restore, confirm byte-identical with the full `shasum -a 256` (on macOS), run a green control, and run the proof cell against the un-mutated tree first to confirm it reports not-applied. ## Credit The fleet01 lead pushed the generalisation that makes item 2 a defect rather than a curiosity: > A COUNT OF BAD THINGS IS A PASS SIGNAL ONLY ONCE YOU HAVE SEPARATELY ESTABLISHED THAT THE > PRODUCER RAN. Exit 127 with zero FAIL lines, and my rc=0 from `head`, are the same shape — a > status channel reporting on the wrong subject while stderr carried the truth. Related: #545 (the same class, `mktemp`), #497, #513.
Author
Owner

Correction from the fleet01 lead, re-measured by me. The defect stands; the root cause named in
the ticket body is the trigger, not the cause
, and the real cause is one line further in.

pipefail is what makes it lie

I re-ran the exact jar_id shape in debian:bookworm-slim with the two set variants:

set -eu              -> out=[]        rc=0
set -euo pipefail    -> out=[absent]  rc=0
(shasum present? NO; file exists? yes)

Without pipefail, the pipeline's status is cut's. cut succeeds on empty input, so the && arm
"succeeds" and jar_id returns an empty string — visibly broken, and whoever reads it knows
something is wrong.

With pipefail the status becomes 127, the && arm fails, and the || echo "absent" fallback
fires and produces a confident, in-domain, false answer.

redeploy-fleetd.sh:62 is set -euo pipefail, so it is the second row.

That inverts pipefail's usual role. It normally reveals a failure. Here it hands one to a ||
default that launders it into a valid domain value.

So the root cause is the ||, not shasum

jar_id() { local f="${1:-$JAR}"; [ -f "$f" ] && shasum -a 256 "$f" | cut -c1-12 || echo "absent"; }

The defect is a default at the read site that catches every failure and maps them all to one value
that means something specific
. "The file is not there" and "I could not compute the hash" are
different facts, and this collapses them.

shasum is merely the failure that happens to be reachable today. Swapping in sha256sum with a
fallback fixes today's trigger and leaves the shape: the next missing tool in that pipeline
reproduces the bug exactly. The X-less-mktemp lesson from #545 says that is not hypothetical — the
idiom had already spread from two sites to six.

The fix must split the existence test from the hash, and make a hash failure an error rather than
absent.
That is what this ticket's item 1 already asks for with its third state; this comment is
why the third state is the point and the tool swap is not.

Severity on fleet01, measured by them, not by me

The fleet01 lead reports both /usr/bin/shasum and /usr/bin/sha256sum present on their Ubuntu
host (perl is installed there), so jar_id is latent on that machine, not live. I have not
checked this myself — it is their measurement of their host.

That is not a reason to downgrade this. shasum is a perl script, so its presence is incidental to
whether perl happens to be installed; a slimmer image or a perl-less base loses it, and the third
state costs nothing. But it does mean this is not blocking their redeploy, which is worth knowing
before anyone sizes it.

What this changes for the fix

Nothing in the acceptance criteria — item 1 already required a third state distinct from both
absent and a hash, and that is the correct fix for the cause named here. What changes is the
reason: the third state is not a nicety on top of a tool swap, it is the entire fix. A PR that
only swaps shasum for sha256sum should not be accepted.

Correction from the fleet01 lead, re-measured by me. The defect stands; the **root cause named in the ticket body is the trigger, not the cause**, and the real cause is one line further in. ## `pipefail` is what makes it lie I re-ran the exact `jar_id` shape in `debian:bookworm-slim` with the two `set` variants: ``` set -eu -> out=[] rc=0 set -euo pipefail -> out=[absent] rc=0 (shasum present? NO; file exists? yes) ``` Without `pipefail`, the pipeline's status is `cut`'s. `cut` succeeds on empty input, so the `&&` arm "succeeds" and `jar_id` returns an **empty string** — visibly broken, and whoever reads it knows something is wrong. With `pipefail` the status becomes 127, the `&&` arm fails, and the `|| echo "absent"` fallback fires and produces a confident, in-domain, **false** answer. `redeploy-fleetd.sh:62` is `set -euo pipefail`, so it is the second row. That inverts `pipefail`'s usual role. It normally reveals a failure. Here it hands one to a `||` default that launders it into a valid domain value. ## So the root cause is the `||`, not `shasum` ```bash jar_id() { local f="${1:-$JAR}"; [ -f "$f" ] && shasum -a 256 "$f" | cut -c1-12 || echo "absent"; } ``` The defect is **a default at the read site that catches every failure and maps them all to one value that means something specific**. "The file is not there" and "I could not compute the hash" are different facts, and this collapses them. `shasum` is merely the failure that happens to be reachable today. Swapping in `sha256sum` with a fallback fixes today's trigger and leaves the shape: the next missing tool in that pipeline reproduces the bug exactly. The X-less-`mktemp` lesson from #545 says that is not hypothetical — the idiom had already spread from two sites to six. **The fix must split the existence test from the hash, and make a hash failure an error rather than `absent`.** That is what this ticket's item 1 already asks for with its third state; this comment is why the third state is the point and the tool swap is not. ## Severity on fleet01, measured by them, not by me The fleet01 lead reports both `/usr/bin/shasum` and `/usr/bin/sha256sum` present on their Ubuntu host (perl is installed there), so `jar_id` is **latent** on that machine, not live. I have not checked this myself — it is their measurement of their host. That is not a reason to downgrade this. `shasum` is a perl script, so its presence is incidental to whether perl happens to be installed; a slimmer image or a perl-less base loses it, and the third state costs nothing. But it does mean this is not blocking their redeploy, which is worth knowing before anyone sizes it. ## What this changes for the fix Nothing in the acceptance criteria — item 1 already required a third state distinct from both `absent` and a hash, and that is the correct fix for the cause named here. What changes is the **reason**: the third state is not a nicety on top of a tool swap, it is the entire fix. A PR that only swaps `shasum` for `sha256sum` should not be accepted.
Author
Owner

Verification of PR #554, measured by me on the branch at 3da44ee in a scratch worktree — not taken from the implementer's report.

The decisive measurement, which macOS cannot make

Both runs in ubuntu:latest, where shasum is absent and sha256sum is present:

main   (93a9ed3):  SUITE exit=127   anchored ^FAIL: count 0
                   scripts/test-redeploy-fleetd.sh: line 298: shasum: command not found
branch (3da44ee):  SUITE exit=0     anchored ^FAIL: count 0

Identical FAIL counts, opposite exit codes. That is this ticket's item 2 and item 3 in one pair of runs: the suite on main dies before running a single test, and a FAIL: count cannot tell that apart from a clean pass. It is the reason the new shell-tests CI job gates on the step's own exit status and deliberately does not grep for FAIL:.

Also confirmed by me: macOS exit 0 / anchored count 0; 69 tests defined and 69 invoked with an empty comm -3; bash -n exit 0 under both /bin/bash 3.2.57 and bash 5.3.9; .gitea/workflows/ci.yml parses with jobs build, shell-tests, contract, and shell-tests is ubuntu-latest + actions/checkout@v4 + bash scripts/test-redeploy-fleetd.sh.

Mutations I ran

mutation pristine anchor result
drop the sha256sum branch from hash256 command -v sha256sum >/dev/null 1 → 0 killed — Linux exit 1, test_no_unguarded_macos_only_hasher_calls with its own message
echo "unhashable" → echo "absent" echo "unhashable" 1 → 0 killed — FAIL: jar_id reported absent for a file that exists, only because no hasher was on PATH
both hash256 arms → shasum -a 1 (wrong algorithm, still content-dependent) sha256sum "$f" 1 → 0 and shasum -a 256 "$f" 1 → 0 SURVIVED — exit 0, anchored count 0

All restores verified byte-identical against 515d929bb53c9ec2c95042e47d3e4d611d60171345227b07feb0de0d20643d3a.

The survivor, and why it is worth one more test

Fixing item 2 meant rewiring the reference hashes in test_jar_id_defaults_to_live_and_reports_explicit_path off the bare macOS-only call and onto hash256:

live_hash="$(hash256 "$JAR")"
staged_hash="$(hash256 "$JAR_STAGED")"

That was the correct fix for the crash. It also made the test's reference and its subject the same instrument. They agree whatever hash256 computes, and agreement between two readings of one instrument carries no information about that instrument. The [ "$live_hash" != "$staged_hash" ] fixture guard catches a constant return; it cannot catch a wrong algorithm.

The old line proved "jar_id returns a sha256 prefix". The new line proves only "jar_id delegates to hash256 and picks the right file". That lost property matters here more than it usually would, because which hasher runs is this ticket's entire subject.

There is no live defect — jar_id's value is only ever compared across runs, so any content-dependent function works operationally. It is a test-strength regression, not a behaviour one.

Asked for before merge: a direct hash256 test pinning a literal expected constant written into the test, computed from a fixed input rather than from any hasher — otherwise the same shared instrument simply moves one level out. With that test in place the third mutation above must go red.

A note on running the mutation, because it nearly fooled me

My first attempt at the third mutation used perl -0pi -e 's/\Q...$f...\E/.../'. \Q escapes regex metacharacters after interpolation; it does not stop Perl interpolating $f, which was undefined and expanded to nothing. The pattern never matched, the file was untouched, and the suite returned exit 0 — which reads exactly like a survivor.

The pristine-anchor counts caught it: both still read 1 after the edit instead of 0. Redone with a line-anchored sed, the mutation applied (1 → 0 on both) and then survived, which is the real result reported above.

A mutation that did not apply is not a surviving mutant. Count the pristine anchor before and after, every time — it is the only string whose count you know in advance.

Verification of PR #554, measured by me on the branch at `3da44ee` in a scratch worktree — not taken from the implementer's report. ## The decisive measurement, which macOS cannot make Both runs in `ubuntu:latest`, where `shasum` is absent and `sha256sum` is present: ``` main (93a9ed3): SUITE exit=127 anchored ^FAIL: count 0 scripts/test-redeploy-fleetd.sh: line 298: shasum: command not found branch (3da44ee): SUITE exit=0 anchored ^FAIL: count 0 ``` **Identical FAIL counts, opposite exit codes.** That is this ticket's item 2 and item 3 in one pair of runs: the suite on `main` dies before running a single test, and a `FAIL:` count cannot tell that apart from a clean pass. It is the reason the new `shell-tests` CI job gates on the step's own exit status and deliberately does not grep for `FAIL:`. Also confirmed by me: macOS exit 0 / anchored count 0; 69 tests defined and 69 invoked with an empty `comm -3`; `bash -n` exit 0 under both `/bin/bash` 3.2.57 and `bash` 5.3.9; `.gitea/workflows/ci.yml` parses with jobs `build`, `shell-tests`, `contract`, and `shell-tests` is `ubuntu-latest` + `actions/checkout@v4` + `bash scripts/test-redeploy-fleetd.sh`. ## Mutations I ran | mutation | pristine anchor | result | |---|---|---| | drop the `sha256sum` branch from `hash256` | `command -v sha256sum >/dev/null` 1 → 0 | **killed** — Linux exit 1, `test_no_unguarded_macos_only_hasher_calls` with its own message | | `echo "unhashable"` → `echo "absent"` | `echo "unhashable"` 1 → 0 | **killed** — `FAIL: jar_id reported absent for a file that exists, only because no hasher was on PATH` | | **both `hash256` arms → `shasum -a 1`** (wrong algorithm, still content-dependent) | `sha256sum "$f"` 1 → 0 **and** `shasum -a 256 "$f"` 1 → 0 | **SURVIVED** — exit 0, anchored count 0 | All restores verified byte-identical against `515d929bb53c9ec2c95042e47d3e4d611d60171345227b07feb0de0d20643d3a`. ## The survivor, and why it is worth one more test Fixing item 2 meant rewiring the reference hashes in `test_jar_id_defaults_to_live_and_reports_explicit_path` off the bare macOS-only call and onto `hash256`: ```bash live_hash="$(hash256 "$JAR")" staged_hash="$(hash256 "$JAR_STAGED")" ``` That was the correct fix for the crash. It also made the test's **reference** and its **subject** the same instrument. They agree whatever `hash256` computes, and **agreement between two readings of one instrument carries no information about that instrument.** The `[ "$live_hash" != "$staged_hash" ]` fixture guard catches a *constant* return; it cannot catch a *wrong algorithm*. The old line proved "`jar_id` returns a sha256 prefix". The new line proves only "`jar_id` delegates to `hash256` and picks the right file". That lost property matters here more than it usually would, because *which hasher runs* is this ticket's entire subject. There is no live defect — `jar_id`'s value is only ever compared across runs, so any content-dependent function works operationally. It is a test-strength regression, not a behaviour one. **Asked for before merge:** a direct `hash256` test pinning a **literal expected constant written into the test**, computed from a fixed input rather than from any hasher — otherwise the same shared instrument simply moves one level out. With that test in place the third mutation above must go red. ## A note on running the mutation, because it nearly fooled me My first attempt at the third mutation used `perl -0pi -e 's/\Q...$f...\E/.../'`. `\Q` escapes regex metacharacters **after** interpolation; it does not stop Perl interpolating `$f`, which was undefined and expanded to nothing. The pattern never matched, the file was untouched, and the suite returned exit 0 — which reads exactly like a survivor. The pristine-anchor counts caught it: both still read 1 after the edit instead of 0. Redone with a line-anchored `sed`, the mutation applied (1 → 0 on both) and *then* survived, which is the real result reported above. **A mutation that did not apply is not a surviving mutant.** Count the pristine anchor before and after, every time — it is the only string whose count you know in advance.
Author
Owner

Closing. PR #554 merged as 26f380a. All three items done.

Verified on the merged tree, not on the branch

merged main (26f380a), macOS          : SUITE exit=0   anchored ^FAIL: count 0
merged main (26f380a), ubuntu:latest  : SUITE exit=0   anchored ^FAIL: count 0

For contrast, the same suite on the pre-merge main (93a9ed3) in the same image: exit=127, anchored count 0, dying at scripts/test-redeploy-fleetd.sh: line 298: shasum: command not found.

What landed

  • Item 1 — hash256() prefers sha256sum, falls back to shasum -a 256, and answers unhashable when neither is on PATH. jar_id now has three distinct answers: absent only for a file that is not there, a 12-char hash when it could be hashed, unhashable when it exists and could not be. absent can no longer mean "I could not tell".
  • Item 2 — the test harness's own two shasum calls (line ~298) go through hash256 too. That was the line killing the suite on Linux.
  • Item 3 — a shell-tests job in .gitea/workflows/ci.yml, ubuntu-latest, gated on the step's own exit status and deliberately not on a FAIL: count.

Plus three new tests: test_no_unguarded_macos_only_hasher_calls (the shape, so a seventh site cannot be written next month — the #545 lesson), test_jar_id_reports_unhashable_when_no_hasher_on_path, and test_hash256_computes_a_real_sha256. 70 defined, 70 invoked, empty comm -3.

The two things worth carrying out of this ticket

1. The root cause was the default, not the missing tool. Under set -euo pipefail a missing hasher makes the pipeline status 127, the || echo "absent" fires, and jar_id returns a confident false answer. Without pipefail the same function returns an empty string and is visibly broken. So pipefail — which normally reveals a failure — is what handed one to a default that laundered it into a valid domain value. A default at the read site that maps every failure onto one value which already means something specific is the shape to look for. shasum was only the trigger.

2. Fixing item 2 briefly cost a test property, and it was worth catching. Rewiring the reference hashes in test_jar_id_defaults_to_live_and_reports_explicit_path onto hash256 made the test's reference and its subject the same instrument. They then agree whatever hash256 computes. A mutation setting both arms to shasum -a 1 survived on the first head, 3da44ee. test_hash256_computes_a_real_sha256 closes it by pinning the FIPS 180 vector for abc as a literal constant written into the test — not read from any hasher, or the shared instrument just moves one level out. The mutation now reports expected ba7816bf8f01, got a9993e364706, and a9993e364706 is SHA-1(abc), which proves the mutated code really ran.

Follow-up filed as #555: the shell suite covers the helper functions well and barely touches the main flow — eight decisions there have no test at all.

Closing. PR #554 merged as `26f380a`. All three items done. ## Verified on the merged tree, not on the branch ``` merged main (26f380a), macOS : SUITE exit=0 anchored ^FAIL: count 0 merged main (26f380a), ubuntu:latest : SUITE exit=0 anchored ^FAIL: count 0 ``` For contrast, the same suite on the pre-merge `main` (`93a9ed3`) in the same image: **exit=127**, anchored count **0**, dying at `scripts/test-redeploy-fleetd.sh: line 298: shasum: command not found`. ## What landed - **Item 1** — `hash256()` prefers `sha256sum`, falls back to `shasum -a 256`, and answers `unhashable` when neither is on PATH. `jar_id` now has three distinct answers: `absent` only for a file that is not there, a 12-char hash when it could be hashed, `unhashable` when it exists and could not be. `absent` can no longer mean "I could not tell". - **Item 2** — the test harness's own two `shasum` calls (line ~298) go through `hash256` too. That was the line killing the suite on Linux. - **Item 3** — a `shell-tests` job in `.gitea/workflows/ci.yml`, `ubuntu-latest`, gated on the step's own exit status and deliberately **not** on a `FAIL:` count. Plus three new tests: `test_no_unguarded_macos_only_hasher_calls` (the shape, so a seventh site cannot be written next month — the #545 lesson), `test_jar_id_reports_unhashable_when_no_hasher_on_path`, and `test_hash256_computes_a_real_sha256`. 70 defined, 70 invoked, empty `comm -3`. ## The two things worth carrying out of this ticket **1. The root cause was the default, not the missing tool.** Under `set -euo pipefail` a missing hasher makes the pipeline status 127, the `|| echo "absent"` fires, and `jar_id` returns a confident false answer. Without `pipefail` the same function returns an empty string and is *visibly* broken. So `pipefail` — which normally reveals a failure — is what handed one to a default that laundered it into a valid domain value. **A default at the read site that maps every failure onto one value which already means something specific** is the shape to look for. `shasum` was only the trigger. **2. Fixing item 2 briefly cost a test property, and it was worth catching.** Rewiring the reference hashes in `test_jar_id_defaults_to_live_and_reports_explicit_path` onto `hash256` made the test's reference and its subject the same instrument. They then agree whatever `hash256` computes. A mutation setting both arms to `shasum -a 1` survived on the first head, `3da44ee`. `test_hash256_computes_a_real_sha256` closes it by pinning the FIPS 180 vector for `abc` as a **literal constant written into the test** — not read from any hasher, or the shared instrument just moves one level out. The mutation now reports `expected ba7816bf8f01, got a9993e364706`, and `a9993e364706` is SHA-1(`abc`), which proves the mutated code really ran. Follow-up filed as #555: the shell suite covers the helper functions well and barely touches the main flow — eight decisions there have no test at all.
ltms closed this issue 2026-09-12 10:21:23 +02:00
Sign in to join this conversation.
1 Participants
Notifications
Due Date
No due date set.
Dependencies

No dependencies set.

Reference: fleet/fleetd#550