Files
fleetd/scripts/config-edit.sh
T
Dai Ha faefea14c4
CI / shell-tests (pull_request) Failing after 6s
CI / contract (pull_request) Successful in 1m8s
CI / build (pull_request) Failing after 2m11s
fleetd #639: redact()'s comment claimed more than the code does
The paragraph added in d7f94ca ended with "This needs no knowledge of the key's
name and so protects a block scalar under any masked key, present or future."
The continuation masking is real and it is an improvement, but that sentence is
too strong: the masking only holds while the masked key's own line is inside the
hunk being printed.

redact() is fed `diff -u` output, which prints three lines of context. A block
scalar's body therefore often arrives with its key line left out. With no key
line, `masked` is never set and the body prints in full, with no "<redacted>"
anywhere. A blank line inside a block scalar loses the anchor the same way: a
blank diff line measures as indent 0, so `indent > masked_indent` is false and
the mask ends early — this time directly under a "<redacted>" marker.

Both were reproduced through the real script with --dry-run, each with a
positive control run first to prove the secret's lines actually reached the
output (without that control, "the secret never entered the diff" and "it
entered and was redacted" are indistinguishable). Filed as fleetd #639, which
also records that this is latent rather than live: today's fleetd.yaml holds 5
block scalars and all 5 sit under non-secret keys.

Comment text only. No change to redact() or to any other function, and no
change to the test suite. scripts/test-config-edit.sh still passes in a clean
copy (16 criteria + 3 extras, exit 0); bash -n clean.

The reason this is worth its own commit: #635 exists because an incomplete
redactor that looks complete is worse than one that visibly does nothing. A
comment that overstates the guarantee is the same defect in prose, and the next
session to read it has no other source.
2026-10-01 19:00:29 +02:00

733 lines
29 KiB
Bash
Executable File

#!/usr/bin/env bash
#
# The one auditable way to edit the live fleetd.yaml.
#
# fleetd ticket #635 — why this exists at all: fleetd.yaml is gitignored and holds the live
# fleet's settings. A bad raw edit reaches a daemon that is already serving, so a direct `Edit`
# on it is refused by policy. This script is the allow-listed alternative, and it is not just
# convenience — it is the thing a raw file write can never give you: a backup, a parse check
# BEFORE the file is installed, and the daemon's own reload verdict read back afterwards. An
# edit to a live config is not finished when the bytes are written. It is finished when the
# daemon has said what it did with them.
#
# What the daemon says, and how this script finds it — measured against `ConfigRef.java` on
# fleetd commit 158a2a8, 2026-10-01:
#
# 1. `ConfigRef` re-reads fleetd.yaml only when the WATCHER sees the mtime move (every 10s by
# default — read the real interval out of the daemon's own startup line, "config watch: ...
# re-read when it changes (every Ns)"). So a verdict never appears before the next tick.
# 2. `ConfigRef.Outcome.summary()` logs exactly one of five strings (ConfigRef.java:371-391):
# config reload refused — <error message>
# config reload refused — these keys cannot change under a running daemon: <keys>. ...
# config reloaded
# config reloaded; these changes need a restart to take effect: <keys>
# config reloaded; partially live — <key: detail | ...>
# A parse/validation failure logs a DIFFERENT line instead, before any summary ever runs
# (ConfigRef.java:425): "config reload from <path> refused, keeping the running config:
# <message>". This script recognises both shapes of refusal.
# 3. The em dash in those strings is a real multi-byte character — match the stable prefix
# "config reload refused" (or "...refused, keeping the running config" for the parse-failure
# shape), never the dash itself.
# 4. A cold-key change (bind/herdrSocket/memberHerdrSocket/broker/auth) throws away the WHOLE
# reload — the running config keeps every old value, not only the cold one.
# 5. A deferred/split change IS applied (current.set(fresh) runs) — "needs a restart" is a
# SUCCESS with a follow-up, never a failure.
#
# Four outcomes, and they stay four (see the exit code table below). The one most likely to be
# gotten wrong is "cannot tell" (exit 5): the daemon may be down, or the watcher may be stalled,
# and folding that into either "refused" or "applied" is worse than never checking at all,
# because a caller then acts on a verdict nobody actually read. So exit 5 never restores — a
# visible, recoverable edit beats an invisible revert of a GOOD edit.
#
# Usage:
# scripts/config-edit.sh --check
# scripts/config-edit.sh --set <yq-path>=<value> [--set ...]
# scripts/config-edit.sh --from <candidate.yaml>
# scripts/config-edit.sh --dry-run --set <yq-path>=<value>
# scripts/config-edit.sh --restore
#
# `--set .a.b=` (an empty value — a forgotten typo) is REFUSED, not accepted as "clear the
# field": a null value falls back to its default rather than erroring, which is silent, not
# safe. To clear a key on purpose, write a literal null: `--set .a.b=null`. Every other value
# is always written as a YAML string (via yq's strenv(), never spliced into the expression), so
# there is currently no --set spelling for the literal three-character STRING "null" itself — use
# --from for that rare case.
#
# Overrides (so this is drivable with no daemon — see scripts/test-config-edit.sh):
# --config <path> default: fleetd/fleetd.yaml
# --log <path> default: fleetd/fleetd.out
# --wait-seconds <n> default: 4x the watch interval this script reads out of --log (10 -> 40)
#
# Exit codes (the --check/--restore/usage-error paths are reported separately, see below):
# 0 applied; verdict read; clean
# 3 applied; verdict read; needs a restart (deferred or split keys named)
# 4 REFUSED by the daemon; backup restored (and the restore's own verdict reported if seen)
# 5 CANNOT TELL — no verdict line inside the wait window. Nothing is restored.
#
# Never prints a secret. fleetd.yaml keeps credentials out by indirection (broker.uriEnv,
# gitTokenEnv) but this script does not rely on that staying true: every diff it prints is piped
# through `redact`, which (a) blanks the userinfo of any `scheme://user:pass@host` and (b) masks
# the whole value on any line whose key looks like a credential. See `redact` below.
set -euo pipefail
REPO="$(cd "$(dirname "${BASH_SOURCE[0]}")/.." && pwd)"
SELF="$REPO/scripts/config-edit.sh"
CONFIG="$REPO/fleetd/fleetd.yaml"
LOG="$REPO/fleetd/fleetd.out"
WAIT_SECONDS_OVERRIDE=""
FALLBACK_PORT=8765
MODE=""
DRY_RUN=0
SETS=()
FROM_FILE=""
# fleetd #635 follow-up — a signal (or any early exit while a candidate is still uninstalled) must
# not leave a `.config-edit.XXXXXX` file sitting beside the live config forever. CAND is global
# (never a function-local) on purpose: this ONE trap, set once, covers every path that ever
# creates a candidate — run_edit and dry_run_diff both assign it, and clear it back to "" once the
# file is consumed (installed, or explicitly removed), so a later, unrelated exit never retries a
# path that already served its purpose.
CAND=""
cleanup_candidate() { [ -n "$CAND" ] && rm -f "$CAND" 2>/dev/null; return 0; }
trap cleanup_candidate EXIT INT TERM
say() { printf '\n\033[1m== %s\033[0m\n' "$*"; }
ok() { printf ' ok %s\n' "$*"; }
warn() { printf ' WARN %s\n' "$*"; }
die() { printf '\n FAIL %s\n\n' "$*" >&2; exit 1; }
set_mode() {
local new="$1"
if [ -n "$MODE" ] && [ "$MODE" != "$new" ]; then
die "cannot combine --$MODE and --$new in one invocation"
fi
MODE="$new"
}
while [ $# -gt 0 ]; do
case "$1" in
--check) set_mode check; shift ;;
--restore) set_mode restore; shift ;;
--set)
[ $# -ge 2 ] || die "--set requires <yq-path>=<value>"
set_mode set
SETS+=("$2")
shift 2 ;;
--from)
[ $# -ge 2 ] || die "--from requires a candidate file path"
set_mode from
FROM_FILE="$2"
shift 2 ;;
--dry-run) DRY_RUN=1; shift ;;
--config)
[ $# -ge 2 ] || die "--config requires a path"
CONFIG="$2"; shift 2 ;;
--log)
[ $# -ge 2 ] || die "--log requires a path"
LOG="$2"; shift 2 ;;
--wait-seconds)
[ $# -ge 2 ] || die "--wait-seconds requires a number of seconds"
WAIT_SECONDS_OVERRIDE="$2"; shift 2 ;;
-h|--help) sed -n '3,70p' "$SELF"; exit 0 ;;
*) echo "unknown option: $1 (try --help)" >&2; exit 2 ;;
esac
done
[ -n "$MODE" ] || die "no action given — use --check, --set, --from, or --restore (see --help)"
# ------------------------------------------------------------------------------------- redaction
#
# Two independent passes, applied to every diff this script ever prints:
# 1. `scheme://user:pass@host` -> `scheme://<redacted>@host`, globally (the `g` flag matters —
# a line can carry more than one URI).
# 2. Any line whose key looks like TOKEN|SECRET|PASSWORD|PASSWD|PASSPHRASE|CREDENTIAL|URI|_KEY,
# matched case-insensitively against the key text (uriEnv, gitTokenEnv, ... are camelCase,
# not SCREAMING_CASE) has its whole value blanked, diff marker and indentation kept so the
# shape of the change is still visible. Deliberately conservative: a false-positive
# redaction on an unrelated line costs nothing, an unredacted secret is a security defect
# (acceptance criterion 7).
#
# fleetd #635 follow-up (ticket comment 17670, defect 7) — a masked key line is not the whole
# story: a YAML block scalar (`|`, `|-`, `>`, `>-`, ...) puts the VALUE on the lines that follow
# the key, each indented deeper than it. The key-name match above only ever sees the key line
# itself, so those continuation lines used to flow straight through unredacted while the key line
# right above them printed a reassuring "<redacted>" — an incomplete redactor that looks complete
# is worse than one that visibly does nothing, because it stops a reviewer from looking further.
# The fix is structural, not another name to match: once a key line is masked, every following
# line indented STRICTLY DEEPER than that key is masked too, by indentation alone, until the
# indentation returns to the key's own level or shallower. This needs no knowledge of the key's
# name, so it covers a block scalar under any masked key — but ONLY while that key's own line is
# itself inside the hunk being printed. `diff -u` prints just three lines of context, so a block
# scalar's body often reaches this function with its key line left out; there is then nothing to
# anchor to, `masked` is never set, and the body prints in full. A blank line inside a block
# scalar loses the anchor the same way, because a blank diff line measures as indent 0. Both are
# measured and filed as fleetd #639 — do not read this paragraph as a guarantee that a masked
# key's value can never be printed.
#
# `redact` is always fed `diff -u` output, and every line of a unified diff starts with exactly
# one of ' ', '+', '-' (the three body markers; '@'/'-'/'+' for the three header-line kinds too).
# That one leading character is NOT part of the YAML indentation, and must be stripped before
# indentation is measured or a key is matched — otherwise a changed ('+' or '-') line reads one
# column shallower than it really is, and either wrongly escapes a continuation mask or wrongly
# ends one early. Tabs are out of scope: YAML forbids them for indentation, and this is a bounded
# fix, not a YAML parser.
redact() {
local line prefix content indent lead key
local masked=0 masked_indent=0 saved_nocasematch=0
shopt -q nocasematch && saved_nocasematch=1
shopt -s nocasematch
sed -E 's#://[^@]*@#://<redacted>@#g' | while IFS= read -r line || [ -n "$line" ]; do
case "$line" in
[\ +-]*) prefix="${line:0:1}"; content="${line:1}" ;;
*) prefix=""; content="$line" ;;
esac
indent=0
while [ "${content:$indent:1}" = " " ]; do indent=$((indent + 1)); done
if [ "$masked" = 1 ] && [ "$indent" -gt "$masked_indent" ]; then
printf '%s%*s<redacted>\n' "$prefix" "$indent" ""
continue
fi
masked=0
if [[ "$content" =~ ^([[:space:]]*)([A-Za-z0-9_.-]+:) ]]; then
lead="${BASH_REMATCH[1]}"
key="${BASH_REMATCH[2]}"
if [[ "$key" =~ (TOKEN|SECRET|PASSWORD|PASSWD|PASSPHRASE|CREDENTIAL|URI|_KEY) ]]; then
printf '%s%s%s <redacted>\n' "$prefix" "$lead" "$key"
masked=1
masked_indent="$indent"
continue
fi
fi
printf '%s\n' "$line"
done
[ "$saved_nocasematch" = 1 ] || shopt -u nocasematch
}
# ------------------------------------------------------------------------------------- the probe
#
# Probe the SOCKET, never `pgrep`/`ps -f` — both print argv, and argv holds `NAME=value`, making
# either a credential channel. The port comes from the config's own `bind.port`; 8765 is only a
# fallback when that key is absent or the file does not parse yet.
resolve_port() {
local file="$1" port
if [ -f "$file" ] && command -v yq >/dev/null 2>&1; then
port="$(yq eval '.bind.port' "$file" 2>/dev/null || true)"
else
port=""
fi
case "$port" in
''|null) echo "$FALLBACK_PORT" ;;
*) echo "$port" ;;
esac
}
daemon_listening() {
local port="$1"
if command -v nc >/dev/null 2>&1; then
nc -z -w1 127.0.0.1 "$port" 2>/dev/null
else
( exec 3<>"/dev/tcp/127.0.0.1/$port" ) 2>/dev/null
fi
}
# --------------------------------------------------------------------------------- the log marker
#
# Take the log's line count BEFORE touching anything. Every later read of "what did the daemon
# say" starts strictly after this mark, so a refusal from hours ago can never be mistaken for
# this edit's verdict. Same approach as scripts/redeploy-fleetd.sh's RESTART_MARK.
log_mark() {
local file="$1"
if [ -f "$file" ]; then
wc -l < "$file" 2>/dev/null || echo 0
else
echo 0
fi
}
read_verdict_after_marker() {
local file="$1" mark="$2"
[ -f "$file" ] || return 0
tail -n "+$((mark + 1))" "$file" 2>/dev/null || true
}
# Classifies one log LINE. Echoes one of: refused | clean | needs-restart | none. Always
# succeeds (every branch ends in `echo`), so it is safe to call from inside `$( )`.
classify_verdict_line() {
local line="$1"
case "$line" in
*'config reload refused'*) echo refused ;;
*'config reload from '*'refused, keeping the running config'*) echo refused ;;
*'config reloaded'*)
case "$line" in
*'need a restart'*|*'partially live'*) echo needs-restart ;;
*) echo clean ;;
esac ;;
*) echo none ;;
esac
return 0
}
scan_region_for_verdict() {
local region="$1" line kind
[ -n "$region" ] || return 1
while IFS= read -r line || [ -n "$line" ]; do
kind="$(classify_verdict_line "$line")"
if [ "$kind" != "none" ]; then
VERDICT_KIND="$kind"
VERDICT_LINE="$line"
return 0
fi
done <<< "$region"
return 1
}
# Sets VERDICT_KIND/VERDICT_LINE and returns 0 on the first verdict line found after $mark;
# returns 1 (VERDICT_KIND=none) if none appeared inside $wait_s seconds. Checks once before each
# sleep AND once more after the last sleep, the same boundary idiom
# scripts/redeploy-fleetd.sh's wait_for_daemon_exit/wait_for_new_pid already use.
wait_for_verdict() {
local log="$1" mark="$2" wait_s="$3" _i region
VERDICT_KIND="none"
VERDICT_LINE=""
for _i in $(seq "$wait_s"); do
region="$(read_verdict_after_marker "$log" "$mark")"
scan_region_for_verdict "$region" && return 0
sleep 1
done
region="$(read_verdict_after_marker "$log" "$mark")"
scan_region_for_verdict "$region" && return 0
return 1
}
last_verdict_line() {
local file="$1" line out=""
[ -f "$file" ] || return 0
while IFS= read -r line || [ -n "$line" ]; do
if [ "$(classify_verdict_line "$line")" != "none" ]; then
out="$line"
fi
done < "$file"
printf '%s' "$out"
}
default_wait_seconds() {
local log="$1" interval=""
if [ -f "$log" ]; then
interval="$(grep -F 'config watch:' "$log" 2>/dev/null | tail -1 \
| sed -E 's/.*\(every ([0-9]+)s\).*/\1/' || true)"
fi
case "$interval" in
''|*[!0-9]*) interval=10 ;;
esac
echo $((interval * 4))
}
# ----------------------------------------------------------------------------------- the backup
#
# Timestamped, never pruned — "keep backups" per the ticket. A pid suffix avoids a same-second
# collision between two invocations.
#
# fleetd #635 follow-up — lands under a DEDICATED, gitignored directory beside the config
# (<dir>/.config-backups/), never beside the config file itself. The whole reason fleetd.yaml is
# gitignored is that it must never be committed, and a backup of it inherits that requirement — a
# bare `fleetd.yaml.bak.*` next to a tracked directory is one `git add -A`/`git add .` away from
# committing the live config. A directory beats a glob on its own: the glob only protects today's
# naming, a location keeps working even if the naming changes later. (See .gitignore for the glob
# kept anyway, as a backstop for a stray backup written the old way.)
BACKUP_DIRNAME=".config-backups"
backup_dir_for() {
local src="$1"
printf '%s/%s' "$(dirname "$src")" "$BACKUP_DIRNAME"
}
backup_config() {
local src="$1" ts backup dir base
dir="$(backup_dir_for "$src")"
mkdir -p "$dir" \
|| die "could not create the backup directory $dir — refusing to edit without a backup. The live config at $src was NOT touched."
ts="$(date -u +%Y%m%dT%H%M%S)Z"
base="$(basename "$src")"
backup="${dir}/${base}.bak.${ts}.$$"
cp "$src" "$backup" \
|| die "could not create a backup at $backup — refusing to edit without one. The live config at $src was NOT touched."
printf '%s' "$backup"
}
newest_backup() {
local cfg="$1" dir base
dir="$(backup_dir_for "$cfg")"
base="$(basename "$cfg")"
ls -t "${dir}/${base}".bak.* 2>/dev/null | head -1 || true
}
# ------------------------------------------------------------------------------------ the file mode
#
# fleetd #635 follow-up — `mv` from a mktemp candidate carries mktemp's 0600 onto the live path
# forever (measured: 644 -> 600 after one --set), and a restore does not undo it either, because
# `cp` onto an EXISTING file keeps the DESTINATION's mode, not the source's. Capture the live
# file's mode before anything touches it, and reapply it to whatever lands on that path
# afterwards — the candidate before install, and the config again after a restore — so an edit
# changes the file's CONTENT only, never its permissions. BSD `stat -f '%Lp'` first (matches this
# project's dev machine), GNU `stat -c '%a'` as the fallback. Prints nothing when the file does
# not exist yet, so apply_mode then does nothing and a first-ever edit falls back to the normal
# umask default rather than inventing a number.
file_mode() {
local file="$1"
[ -f "$file" ] || return 0
stat -f '%Lp' "$file" 2>/dev/null || stat -c '%a' "$file" 2>/dev/null || true
}
apply_mode() {
local file="$1" mode="$2"
[ -n "$mode" ] || return 0
chmod "$mode" "$file" 2>/dev/null || true
}
# --------------------------------------------------------------------------- candidate builders
#
# Never edit the live file in place. Each builder fills $1 (a temp file already sitting in the
# SAME directory as the live config, so the later `mv` install is a rename, not a cross-device
# copy — see run_edit).
# fleetd #635 follow-up — a forgotten value (`--set .a.b=`, a plausible typo) must never be
# accepted as "clear the field". `*=*` alone cannot tell "--set .a.b=" from "--set .a.b=7" apart
# — both contain an `=` — so the guard has to look at the VALUE, not the shape of the argument.
# An empty value refuses outright: nothing is installed, and the message names the likely cause
# AND the two ways to actually mean it (clear on purpose, or an intentional empty string via
# --from). Measured against the real daemon loader: a quoted empty string reads back as a null
# field (`quoted empty -> OK int=null`), and a null numeric field FALLS BACK TO ITS DEFAULT rather
# than erroring — so this is not a cosmetic nit, it is the one shape of edit that widens capacity
# silently instead of failing loudly, which is exactly what this script exists to catch.
#
# A deliberate clear needs its own spelling, because `""` and YAML `null` are NOT the same value
# to the loader (`""` is a valid empty String; `null` means absent, and an Integer field reads
# either the same way — null — but a String field would keep `""` as a real value). `--set
# .a.b=null` is that spelling: it writes a literal, unquoted `null` via yq, never the string
# "null" through strenv(). One consequence worth knowing: there is currently no --set spelling
# for the three-character STRING "null" itself (it collides with the clear spelling) — use
# --from for that rare case.
apply_set_pairs() {
local cand="$1" kv path value
shift
for kv in "$@"; do
case "$kv" in
*=*) : ;;
*) die "--set expects <yq-path>=<value>, got: '$kv'" ;;
esac
path="${kv%%=*}"
path="${path#.}"
value="${kv#*=}"
if [ -z "$value" ]; then
die "--set '$kv' has an EMPTY value — refusing. Nothing was installed. A forgotten value
would NULL the field, and a null value falls back to its default rather than erroring —
silent, not safe. Did you mean --set .${path}=null to clear it on purpose, or --from a
file if you need a genuinely empty string?"
fi
# fleetd #635 follow-up (ticket comment 17673, defect 8) — these two failure messages used to
# echo the full "$kv" (path=value, exactly as the operator typed it), unredacted. The operator
# already has the value, so a terminal is not where this leaks — the risk is where the output
# goes NEXT: this fleet pastes command output into tickets, PRs and fleet_reply bodies, and a
# failure is exactly when someone copies it to ask for help. Print the PATH, which is what's
# needed to fix the command, and never the value. $kv is not key:value-shaped YAML, so piping
# it through redact would just pass it straight through — a false sense of coverage, the same
# mistake as defect 7.
if [ "$value" = "null" ]; then
yq eval -i ".${path} = null" "$cand" \
|| die "yq could not clear --set '.${path}=null' — nothing was installed. The live config is unchanged."
continue
fi
CONFIG_EDIT_SET_VALUE="$value" yq eval -i ".${path} = strenv(CONFIG_EDIT_SET_VALUE)" "$cand" \
|| die "yq could not apply --set '.${path}=<value>' — nothing was installed. The live config is unchanged."
done
}
build_from_set() {
local cand="$1"
cp "$CONFIG" "$cand"
apply_set_pairs "$cand" "${SETS[@]}"
}
build_from_file() {
local cand="$1"
[ -f "$FROM_FILE" ] || die "--from file not found: $FROM_FILE"
cp "$FROM_FILE" "$cand"
}
parse_check() {
yq eval '.' "$1" >/dev/null 2>&1
}
install_candidate() {
local cand="$1" live="$2"
mv -f "$cand" "$live"
}
# -------------------------------------------------------------------------------- the report path
#
# Prints the literal command the operator (or a test) can run to restore the backup by hand — the
# absolute path to THIS script plus the overrides actually in force, so it works from any cwd.
restore_command_line() {
printf '%q --restore --config %q --log %q --wait-seconds %q' "$SELF" "$CONFIG" "$LOG" "$WAIT_SECONDS"
}
# State 4 only: restore the pre-edit backup, then wait for a SECOND verdict confirming the
# restore itself reloaded cleanly. Never claims a restore it did not observe — if the second wait
# also times out, it says so plainly rather than reporting "restored" as though confirmed.
restore_and_confirm() {
local backup="$1" mark2 orig_mode
orig_mode="$(file_mode "$CONFIG")"
mark2="$(log_mark "$LOG")"
cp "$backup" "$CONFIG" \
|| die "could not restore $backup onto $CONFIG — the live config is left as the REFUSED edit. Fix this by hand immediately: cp \"$backup\" \"$CONFIG\""
apply_mode "$CONFIG" "$orig_mode"
ok "restored from $backup"
if wait_for_verdict "$LOG" "$mark2" "$WAIT_SECONDS"; then
case "$VERDICT_KIND" in
refused) warn "the RESTORE was also refused by the daemon: $VERDICT_LINE" ;;
*) ok "restore confirmed: $VERDICT_LINE" ;;
esac
else
warn "the restore is on disk, but no confirming verdict line appeared within ${WAIT_SECONDS}s"
warn "cannot confirm the restore reloaded cleanly — check $LOG by hand"
fi
return 0
}
# The four-outcome decision. Echoed as a function so run_edit/restore_mode share one place that
# can return 0/3/4/5 — never duplicated, never re-worded between the two callers.
report_outcome() {
local mark="$1" backup="$2" kind line
say "waiting for the daemon's verdict (up to ${WAIT_SECONDS}s)"
if wait_for_verdict "$LOG" "$mark" "$WAIT_SECONDS"; then
kind="$VERDICT_KIND"; line="$VERDICT_LINE"
else
kind="none"
fi
case "$kind" in
clean)
ok "daemon verdict: $line"
say "result: applied cleanly"
return 0 ;;
needs-restart)
ok "daemon verdict: $line"
say "result: applied — a restart is needed for the change(s) named above"
return 3 ;;
refused)
warn "daemon verdict: $line"
say "result: REFUSED — restoring the backup"
restore_and_confirm "$backup"
return 4 ;;
none)
warn "no verdict line appeared within ${WAIT_SECONDS}s after $LOG line $mark"
warn "CANNOT TELL whether the daemon applied this edit, refused it, or is simply down."
warn "Nothing was restored — the edit is still on disk at $CONFIG."
echo
echo " backup: $backup"
echo " to restore it by hand:"
echo " $(restore_command_line)"
return 5 ;;
esac
}
# ------------------------------------------------------------------------------------- the modes
check_mode() {
say "config-edit --check"
if [ -f "$CONFIG" ]; then
if parse_check "$CONFIG"; then
ok "config parses: $CONFIG"
else
warn "config does NOT parse as valid YAML: $CONFIG"
fi
else
warn "no config file at $CONFIG"
fi
local port
port="$(resolve_port "$CONFIG")"
if daemon_listening "$port"; then
ok "daemon is listening on 127.0.0.1:$port"
else
warn "no daemon detected listening on 127.0.0.1:$port"
fi
ok "watch interval assumed: $(( $(default_wait_seconds "$LOG") / 4 ))s (derives --wait-seconds default of $(default_wait_seconds "$LOG")s)"
local verdict
verdict="$(last_verdict_line "$LOG")"
if [ -n "$verdict" ]; then
ok "last verdict in log: $verdict"
else
warn "no reload verdict line found in $LOG"
fi
local backup
backup="$(newest_backup "$CONFIG")"
if [ -n "$backup" ]; then
ok "newest backup: $backup"
else
warn "no backups found for $CONFIG"
fi
if command -v yq >/dev/null 2>&1; then
ok "yq: $(yq --version 2>&1)"
else
warn "yq not found on PATH"
fi
return 0
}
# Shared by --set and --from: backup, build, parse-check, redacted diff, install, await verdict.
run_edit() {
local builder="$1"
[ -f "$CONFIG" ] || die "no config at $CONFIG — nothing to edit"
local mark orig_mode
mark="$(log_mark "$LOG")"
orig_mode="$(file_mode "$CONFIG")"
say "probe"
local port
port="$(resolve_port "$CONFIG")"
if daemon_listening "$port"; then
ok "daemon appears to be listening on 127.0.0.1:$port"
else
warn "no daemon detected listening on 127.0.0.1:$port — a verdict may never appear"
fi
say "backup"
local backup
backup="$(backup_config "$CONFIG")"
ok "backup: $backup"
say "candidate"
local cand
CAND="$(mktemp "$(dirname "$CONFIG")/.config-edit.XXXXXX")" \
|| die "could not create a candidate temp file next to $CONFIG"
cand="$CAND"
if ! "$builder" "$cand"; then
rm -f "$cand"; CAND=""
die "could not build the candidate — nothing was installed. The live config at $CONFIG is unchanged."
fi
if ! parse_check "$cand"; then
rm -f "$cand"; CAND=""
die "candidate does not parse as valid YAML — nothing was installed. The live config at $CONFIG is unchanged."
fi
ok "candidate parses"
apply_mode "$cand" "$orig_mode"
say "change (redacted)"
diff -u "$backup" "$cand" | redact || true
say "install"
install_candidate "$cand" "$CONFIG" \
|| die "could not install the candidate onto $CONFIG — the live config was NOT changed. The validated candidate is sitting at $cand; investigate before retrying."
CAND=""
ok "installed: $CONFIG"
local rc=0
report_outcome "$mark" "$backup" || rc=$?
return "$rc"
}
dry_run_diff() {
local builder="$1"
[ -f "$CONFIG" ] || die "no config at $CONFIG — nothing to diff against"
local cand
CAND="$(mktemp "$(dirname "$CONFIG")/.config-edit.XXXXXX")" \
|| die "could not create a candidate temp file next to $CONFIG"
cand="$CAND"
if ! "$builder" "$cand"; then
rm -f "$cand"; CAND=""
die "could not build the candidate — this was a --dry-run, nothing would have been installed either"
fi
if ! parse_check "$cand"; then
rm -f "$cand"; CAND=""
die "candidate does not parse as valid YAML — this was a --dry-run, nothing would have been installed either"
fi
say "dry run — diff (redacted), nothing installed"
diff -u "$CONFIG" "$cand" | redact || true
rm -f "$cand"; CAND=""
return 0
}
restore_mode() {
[ -f "$CONFIG" ] || die "no config at $CONFIG to restore onto"
local backup dir base
backup="$(newest_backup "$CONFIG")"
if [ -z "$backup" ]; then
# fleetd #635 follow-up (ticket comment 17664) — this message must name the directory the
# code actually searches (backup_dir_for, same as newest_backup), not the old beside-the-
# config glob. A backup written the OLD way is real and NOT searched any more — say so and
# give the one-line recovery command — but do NOT make the search itself look there; that
# would be a behaviour change nobody asked for. The message is the only thing being fixed.
dir="$(backup_dir_for "$CONFIG")"
base="$(basename "$CONFIG")"
die "no backup found matching ${dir}/${base}.bak.* — nothing to restore.
A backup written the OLD way, directly beside the config (${CONFIG}.bak.*), is NOT
searched — that location was retired so a backup of a file that must never be committed
cannot sit next to a tracked directory. If one exists there, recover it by hand:
cp ${CONFIG}.bak.<timestamp>.<pid> $CONFIG"
fi
[ -f "$backup" ] || die "backup candidate $backup vanished"
say "restore"
ok "restoring $backup onto $CONFIG"
local mark orig_mode
mark="$(log_mark "$LOG")"
orig_mode="$(file_mode "$CONFIG")"
cp "$backup" "$CONFIG" || die "could not copy $backup onto $CONFIG"
apply_mode "$CONFIG" "$orig_mode"
ok "installed: $CONFIG"
local rc=0
report_outcome "$mark" "$backup" || rc=$?
return "$rc"
}
# -------------------------------------------------------------------------------------- dispatch
if [ -n "$WAIT_SECONDS_OVERRIDE" ]; then
WAIT_SECONDS="$WAIT_SECONDS_OVERRIDE"
else
WAIT_SECONDS="$(default_wait_seconds "$LOG")"
fi
RC=0
case "$MODE" in
check)
check_mode || RC=$?
;;
set)
[ "${#SETS[@]}" -gt 0 ] || die "--set requires at least one <yq-path>=<value>"
if [ "$DRY_RUN" = 1 ]; then
dry_run_diff build_from_set || RC=$?
else
run_edit build_from_set || RC=$?
fi
;;
from)
[ -n "$FROM_FILE" ] || die "--from requires a candidate file path"
if [ "$DRY_RUN" = 1 ]; then
dry_run_diff build_from_file || RC=$?
else
run_edit build_from_file || RC=$?
fi
;;
restore)
restore_mode || RC=$?
;;
esac
exit "$RC"