redeploy-fleetd.sh cannot run on Linux at all: every mktemp -t NAME fails under GNU coreutils, and the script blames the systemd bus for it #545

Closed
opened 2026-09-12 09:11:45 +02:00 by ltms · 1 comment
Owner

Found while verifying PR #541 (#504 item 1). Not caused by that PR — four of the six sites
predate it. I merged #541 on its own terms and filed this separately.

Measured on main at 7611b69.

The defect

redeploy-fleetd.sh calls mktemp -t NAME six times, and none of the six templates contains any
X placeholders:

$ grep -n "mktemp" scripts/redeploy-fleetd.sh
254:  if ! err_file="$(mktemp -t systemd-installed-err)"; then
278:  if ! err_file="$(mktemp -t systemd-loaded-err)"; then
305:  if ! err_file="$(mktemp -t launchd-unload-err)"; then     # added by #541
320:  if ! err_file="$(mktemp -t systemd-stop-err)"; then       # added by #541
833:  BUILD_LOG="$(mktemp -t fleetd-build)"
1063: FRESH_LOG="$(mktemp -t fleetd-fresh-log)"

BSD mktemp (macOS) accepts that. GNU mktemp (every Linux distribution) does not. BSD appends
its own random suffix; GNU treats the argument as a literal template and requires at least three
Xs.

Measured in debian:stable-slim, mktemp (GNU coreutils) 9.7, bash 5.2.37, aarch64:

$ for t in systemd-installed-err systemd-loaded-err launchd-unload-err \
           systemd-stop-err fleetd-build fleetd-fresh-log; do
    if out=$(mktemp -t "$t" 2>&1); then echo "  $t -> OK $out"; else echo "  $t -> FAIL: $out"; fi
  done
  systemd-installed-err -> FAIL: mktemp: too few X's in template 'systemd-installed-err'
  systemd-loaded-err    -> FAIL: mktemp: too few X's in template 'systemd-loaded-err'
  launchd-unload-err    -> FAIL: mktemp: too few X's in template 'launchd-unload-err'
  systemd-stop-err      -> FAIL: mktemp: too few X's in template 'systemd-stop-err'
  fleetd-build          -> FAIL: mktemp: too few X's in template 'fleetd-build'
  fleetd-fresh-log      -> FAIL: mktemp: too few X's in template 'fleetd-fresh-log'

For comparison, the same call on this Mac:

$ mktemp -t systemd-loaded-err
/var/folders/wf/.../T/systemd-loaded-err.gFVGGhiB98

All six fail on Linux. All six succeed on macOS. That is why this has never been seen here.

What it does to the script — measured, not reasoned

I sourced the real script in the container with a healthy stub systemctl on PATH that gives
a clean negative (exit 3, nothing on stderr) — that is, a Linux host where systemd is perfectly
fine and the unit simply is not active:

$ source /work/scripts/redeploy-fleetd.sh
$ detect_supervisor
mktemp: too few X's in template 'systemd-loaded-err'
mktemp: too few X's in template 'systemd-installed-err'
unclear<SEP>the systemd --user probe for 'fleetd' could not answer cleanly (systemctl exited
non-zero and reported an error on stderr, not a clean negative — e.g. it cannot reach the user bus)

Two things go wrong, and the second is worse than the first.

1. The script refuses to run. detect_supervisor returns unclear, and
require_drivable_supervisor turns anything that is not exactly launchd or systemd into a
refusal. So on any Linux host with systemctl on PATH — which is fleet01 — redeploy-fleetd.sh
stops before it does anything. It is loud and it is safe. It is also completely unusable there.

2. The explanation names a cause that was never measured. The detail string says systemctl
"exited non-zero and reported an error on stderr … e.g. it cannot reach the user bus". In this run
systemctl was never executed at all. mktemp failed first, at redeploy-fleetd.sh:278, and
the if ! guard sets SYSTEMD_LOADED_ERRORED=1 and returns before the probe runs.

An operator reading that message goes and investigates systemd --user, DBUS_SESSION_BUS_ADDRESS
and their user bus. All of it healthy. The real cause is one missing XXXXXX.

This is the same shape as #404 and #425: the third state ("cannot tell") did its job and stopped the
script from guessing, but the reason attached to the third state is a guess, and it is wrong. A
"cannot tell" answer must say what it could not do, not what it assumes went wrong.

The suite does catch it — with a caveat that matters more

Running the repo's own suite in the same container, with a shasum stand-in built on sha256sum
(Debian slim has no perl):

suite exit=1
anchored ^FAIL: = 1
FAIL: unload_launchd_if_loaded must tolerate a clean already-unloaded answer
      (non-zero exit, empty stderr): mktemp: too few X's in template 'launchd-unload-err'

So #541's new tests are what turn this from a silent degradation into a visible failure on
Linux. That is worth saying plainly, because it looks at first like #541 introduced the problem.

The caveat. Without the shasum stand-in the same suite gives:

suite exit=127
anchored ^FAIL: = 0
scripts/test-redeploy-fleetd.sh: line 221: shasum: command not found

Exit 127, and zero anchored ^FAIL: lines. The ^FAIL: count is the number we have been
quoting as the pass signal, and here it reads exactly the same as a fully green run. A suite that
dies before running its tests reports no failures. The exit code is the pass signal; the FAIL
count is only a detail.
Any future acceptance criterion that quotes one must quote both.

What is wanted

  1. Give every mktemp template XXXXXX, at all six sites. mktemp -t fleetd-build.XXXXXX works on
    both platforms. This is the whole fix for item 1.
  2. Fix detect_supervisor's unclear detail so it distinguishes "the systemd probe ran and
    answered badly" from "the systemd probe could not be set up". Right now one detail string covers
    both, and it asserts the first. The SYSTEMD_*_ERRORED flags already carry the distinction at
    the point where it is known — systemd_loaded:279 (setup failed) versus :284 (probe answered
    with stderr) — so the information exists and is thrown away one frame later.
  3. Add a Linux leg to the suite, or at minimum a test that fails when a mktemp template has no
    Xs. A source-text check is enough and costs nothing:
    grep -n 'mktemp -t [^ ]*' | grep -v 'XXX' must find nothing.

Acceptance

  • All six sites fixed; grep 'mktemp -t' scripts/redeploy-fleetd.sh shows XXXXXX on every line.
  • A test that fails if any mktemp template in the script lacks Xs. Show it going red against the
    current text and green after the fix.
  • detect_supervisor reports a different detail for "could not create the temp file" than for
    "systemctl answered with stderr". A test for each, so swapping the two fails.
  • The suite still exits 0 on macOS with 0 anchored ^FAIL: lines, under both /bin/bash 3.2.57 and
    env bash 5.x.
  • Report the suite's exit code alongside the ^FAIL: count every time. Do not report the count
    alone.
  • Never run scripts/redeploy-fleetd.sh against the live daemon, with any flag.

Not measured by me — do not treat as fact

  • Whether fleet01's host has shasum. Debian slim does not, because it has no perl. fleet01's
    host probably does. The shasum finding above is a fact about the container, not about fleet01,
    and the ticket does not depend on it.
  • Whether anyone has ever run this script on Linux. The fleet01 roll test was deferred by their
    operator, so quite possibly not. That would explain how six sites survived.

Related: #504 (items 2-4 still open), #541 (the PR whose tests exposed this), #492 (the
supervisor-detection work this sits on), #404 / #425 (a status answer that reports a cause it did
not measure).

Found while verifying PR #541 (#504 item 1). **Not caused by that PR** — four of the six sites predate it. I merged #541 on its own terms and filed this separately. Measured on `main` at `7611b69`. ## The defect `redeploy-fleetd.sh` calls `mktemp -t NAME` six times, and none of the six templates contains any `X` placeholders: ``` $ grep -n "mktemp" scripts/redeploy-fleetd.sh 254: if ! err_file="$(mktemp -t systemd-installed-err)"; then 278: if ! err_file="$(mktemp -t systemd-loaded-err)"; then 305: if ! err_file="$(mktemp -t launchd-unload-err)"; then # added by #541 320: if ! err_file="$(mktemp -t systemd-stop-err)"; then # added by #541 833: BUILD_LOG="$(mktemp -t fleetd-build)" 1063: FRESH_LOG="$(mktemp -t fleetd-fresh-log)" ``` **BSD `mktemp` (macOS) accepts that. GNU `mktemp` (every Linux distribution) does not.** BSD appends its own random suffix; GNU treats the argument as a literal template and requires at least three `X`s. Measured in `debian:stable-slim`, `mktemp (GNU coreutils) 9.7`, bash 5.2.37, aarch64: ``` $ for t in systemd-installed-err systemd-loaded-err launchd-unload-err \ systemd-stop-err fleetd-build fleetd-fresh-log; do if out=$(mktemp -t "$t" 2>&1); then echo " $t -> OK $out"; else echo " $t -> FAIL: $out"; fi done systemd-installed-err -> FAIL: mktemp: too few X's in template 'systemd-installed-err' systemd-loaded-err -> FAIL: mktemp: too few X's in template 'systemd-loaded-err' launchd-unload-err -> FAIL: mktemp: too few X's in template 'launchd-unload-err' systemd-stop-err -> FAIL: mktemp: too few X's in template 'systemd-stop-err' fleetd-build -> FAIL: mktemp: too few X's in template 'fleetd-build' fleetd-fresh-log -> FAIL: mktemp: too few X's in template 'fleetd-fresh-log' ``` For comparison, the same call on this Mac: ``` $ mktemp -t systemd-loaded-err /var/folders/wf/.../T/systemd-loaded-err.gFVGGhiB98 ``` **All six fail on Linux. All six succeed on macOS.** That is why this has never been seen here. ## What it does to the script — measured, not reasoned I sourced the real script in the container with a **healthy** stub `systemctl` on `PATH` that gives a clean negative (exit 3, nothing on stderr) — that is, a Linux host where systemd is perfectly fine and the unit simply is not active: ``` $ source /work/scripts/redeploy-fleetd.sh $ detect_supervisor mktemp: too few X's in template 'systemd-loaded-err' mktemp: too few X's in template 'systemd-installed-err' unclear<SEP>the systemd --user probe for 'fleetd' could not answer cleanly (systemctl exited non-zero and reported an error on stderr, not a clean negative — e.g. it cannot reach the user bus) ``` Two things go wrong, and the second is worse than the first. **1. The script refuses to run.** `detect_supervisor` returns `unclear`, and `require_drivable_supervisor` turns anything that is not exactly `launchd` or `systemd` into a refusal. So on any Linux host with `systemctl` on `PATH` — which is fleet01 — `redeploy-fleetd.sh` stops before it does anything. It is loud and it is safe. It is also completely unusable there. **2. The explanation names a cause that was never measured.** The detail string says systemctl "exited non-zero and reported an error on stderr … e.g. it cannot reach the user bus". In this run **systemctl was never executed at all**. `mktemp` failed first, at `redeploy-fleetd.sh:278`, and the `if !` guard sets `SYSTEMD_LOADED_ERRORED=1` and returns before the probe runs. An operator reading that message goes and investigates `systemd --user`, `DBUS_SESSION_BUS_ADDRESS` and their user bus. All of it healthy. The real cause is one missing `XXXXXX`. This is the same shape as #404 and #425: the third state ("cannot tell") did its job and stopped the script from guessing, but the *reason attached to the third state* is a guess, and it is wrong. A "cannot tell" answer must say what it could not do, not what it assumes went wrong. ## The suite does catch it — with a caveat that matters more Running the repo's own suite in the same container, with a `shasum` stand-in built on `sha256sum` (Debian slim has no perl): ``` suite exit=1 anchored ^FAIL: = 1 FAIL: unload_launchd_if_loaded must tolerate a clean already-unloaded answer (non-zero exit, empty stderr): mktemp: too few X's in template 'launchd-unload-err' ``` So **#541's new tests are what turn this from a silent degradation into a visible failure** on Linux. That is worth saying plainly, because it looks at first like #541 introduced the problem. **The caveat.** Without the `shasum` stand-in the same suite gives: ``` suite exit=127 anchored ^FAIL: = 0 scripts/test-redeploy-fleetd.sh: line 221: shasum: command not found ``` **Exit 127, and zero anchored `^FAIL:` lines.** The `^FAIL:` count is the number we have been quoting as the pass signal, and here it reads exactly the same as a fully green run. A suite that dies before running its tests reports no failures. **The exit code is the pass signal; the FAIL count is only a detail.** Any future acceptance criterion that quotes one must quote both. ## What is wanted 1. Give every `mktemp` template `XXXXXX`, at all six sites. `mktemp -t fleetd-build.XXXXXX` works on both platforms. This is the whole fix for item 1. 2. Fix `detect_supervisor`'s `unclear` detail so it distinguishes "the systemd probe ran and answered badly" from "the systemd probe could not be set up". Right now one detail string covers both, and it asserts the first. The `SYSTEMD_*_ERRORED` flags already carry the distinction at the point where it is known — `systemd_loaded:279` (setup failed) versus `:284` (probe answered with stderr) — so the information exists and is thrown away one frame later. 3. Add a Linux leg to the suite, or at minimum a test that fails when a `mktemp` template has no `X`s. A source-text check is enough and costs nothing: `grep -n 'mktemp -t [^ ]*' | grep -v 'XXX'` must find nothing. ## Acceptance - All six sites fixed; `grep 'mktemp -t' scripts/redeploy-fleetd.sh` shows `XXXXXX` on every line. - A test that fails if any `mktemp` template in the script lacks `X`s. Show it going red against the current text and green after the fix. - `detect_supervisor` reports a *different* detail for "could not create the temp file" than for "systemctl answered with stderr". A test for each, so swapping the two fails. - The suite still exits 0 on macOS with 0 anchored `^FAIL:` lines, under both `/bin/bash` 3.2.57 and `env bash` 5.x. - **Report the suite's exit code alongside the `^FAIL:` count every time.** Do not report the count alone. - Never run `scripts/redeploy-fleetd.sh` against the live daemon, with any flag. ## Not measured by me — do not treat as fact - **Whether fleet01's host has `shasum`.** Debian slim does not, because it has no perl. fleet01's host probably does. The `shasum` finding above is a fact about the container, not about fleet01, and the ticket does not depend on it. - **Whether anyone has ever run this script on Linux.** The fleet01 roll test was deferred by their operator, so quite possibly not. That would explain how six sites survived. Related: #504 (items 2-4 still open), #541 (the PR whose tests exposed this), #492 (the supervisor-detection work this sits on), #404 / #425 (a status answer that reports a cause it did not measure).
Author
Owner

Fixed by PR #548, merged into main at 0b032f5. (Gitea did not auto-close this — the PR wrote
"Fixes fleetd #545", and the fleetd prefix stops the keyword matching.)

Items 1 and 2 are done and I verified them myself rather than from the worker's report:

  • All six mktemp -t sites carry XXXXXX. The only remaining mktemp -t text with no X anywhere
    under scripts/ is a prose comment in the test file.
  • detect_supervisor now gives the setup failure its own detail text, on a third value (2) of the
    existing SYSTEMD_*_ERRORED flags rather than a second flag.

The measurement that actually closes this, under GNU coreutils 9.1 in debian:bookworm-slim, with a
systemctl stub that exits non-zero and writes nothing to stderr — a clean negative:

main:        mktemp: too few X's in template 'systemd-loaded-err'
             SYSTEMD_LOADED_ERRORED=1
             detect_supervisor => unclear | "systemctl exited non-zero and reported an error on
                                            stderr, not a clean negative — e.g. it cannot reach
                                            the user bus"
this fix:    SYSTEMD_LOADED_ERRORED=0
             detect_supervisor => none

So the wrong-cause message is gone, and the script no longer refuses on a healthy Linux host.

Item 3 is NOT done and is now #550. This ticket asked for "a Linux leg to the suite, or at
minimum a test that fails when a mktemp template has no Xs". The second half shipped — that
source-text test exists, and I proved it by stripping the X from a seventh site the worker had
never touched. The first half did not, and while verifying this I found out why it matters more than
I thought when I wrote it:

  • the suite still cannot run on Linux at all — it dies at the first shasum call with exit 127 and
    zero anchored ^FAIL: lines, which is what a green run looks like on the channel anyone checks,
  • redeploy-fleetd.sh:147's jar_id() uses shasum too, so on Linux it reports an existing jar as
    absent with exit 0 — the exact output the redeploy procedure tells a lead to trust,
  • and nothing in .gitea/workflows/ runs this suite on any platform.

Those three are #550. This ticket's own text already flagged the exit-127 shape as a caveat; #550 is
that caveat promoted to a defect, with the production-script half attached.

Also filed from this work: #552, for the unguarded mktemp at :1089 that sits after the
restart — the fix here removed its trigger but not its position.

Fixed by PR #548, merged into `main` at `0b032f5`. (Gitea did not auto-close this — the PR wrote "Fixes fleetd #545", and the `fleetd ` prefix stops the keyword matching.) Items 1 and 2 are done and I verified them myself rather than from the worker's report: - All six `mktemp -t` sites carry `XXXXXX`. The only remaining `mktemp -t` text with no `X` anywhere under `scripts/` is a prose comment in the test file. - `detect_supervisor` now gives the setup failure its own detail text, on a third value (`2`) of the existing `SYSTEMD_*_ERRORED` flags rather than a second flag. The measurement that actually closes this, under GNU coreutils 9.1 in `debian:bookworm-slim`, with a `systemctl` stub that exits non-zero and writes **nothing** to stderr — a clean negative: ``` main: mktemp: too few X's in template 'systemd-loaded-err' SYSTEMD_LOADED_ERRORED=1 detect_supervisor => unclear | "systemctl exited non-zero and reported an error on stderr, not a clean negative — e.g. it cannot reach the user bus" this fix: SYSTEMD_LOADED_ERRORED=0 detect_supervisor => none ``` So the wrong-cause message is gone, and the script no longer refuses on a healthy Linux host. **Item 3 is NOT done and is now #550.** This ticket asked for "a Linux leg to the suite, or at minimum a test that fails when a `mktemp` template has no `X`s". The second half shipped — that source-text test exists, and I proved it by stripping the `X` from a seventh site the worker had never touched. The first half did not, and while verifying this I found out why it matters more than I thought when I wrote it: - the suite still cannot run on Linux at all — it dies at the first `shasum` call with exit 127 and **zero** anchored `^FAIL:` lines, which is what a green run looks like on the channel anyone checks, - `redeploy-fleetd.sh:147`'s `jar_id()` uses `shasum` too, so on Linux it reports an existing jar as `absent` with exit 0 — the exact output the redeploy procedure tells a lead to trust, - and nothing in `.gitea/workflows/` runs this suite on any platform. Those three are #550. This ticket's own text already flagged the exit-127 shape as a caveat; #550 is that caveat promoted to a defect, with the production-script half attached. Also filed from this work: #552, for the unguarded `mktemp` at `:1089` that sits **after** the restart — the fix here removed its trigger but not its position.
ltms closed this issue 2026-09-12 09:42:33 +02:00
Sign in to join this conversation.
1 Participants
Notifications
Due Date
No due date set.
Dependencies

No dependencies set.

Reference: fleet/fleetd#545