fleetd #593 (pid-count half): running_pid() no longer matches the caller #597
Reference in New Issue
Block a user
Delete Branch "worker/593-1a8025-5"
Deleting a branch is permanent. Although the deleted branch may continue to exist for a short time before it actually gets removed, it CANNOT be undone in most cases. Continue?
Fixes the pid-count half of #593 only (instance 2 + instance 3). Instance 1 (the fleetd.out log source, systemd-only) is left for a Linux host per the ticket's own scope split.
Instance 2 -- the pid count can match the caller.
running_pid()was a barepgrep -f "$PATTERN", which matches any process whose full command line contains the pattern text, including a shell that merely embeds it as literal text (a hand-typed investigation, an ssh-shapedsh -c '...; ...', or a pipeline) rather than being the daemon.pgrep -cdoes not exist on BSD/macOS, so this can't be fixed by switching flags.running_pid()now keeps pgrep to find candidates (portable across BSD and Linux), then drops any candidate whose process name (comm) names a shell (sh/bash/zsh/dash/ksh) -- the daemon is alwaysjava, so a self-matching wrapper of this shape is excluded while a genuine second daemon-shaped process still counts.assert_single_daemon(the >1 check) is unchanged and still dies on a real race.Instance 3 -- the remediation text reproduced the defect. The die message told the operator to "Investigate with 'pgrep -f "$PATTERN"'" -- typed by hand or over ssh, exactly the self-matching invocation. It now points at
ps -eo pid,comm,args | grep -F "$PATTERN"plus checking the COMM column, and says in words that a bare pgrep can match the caller.Tests added (scripts/test-redeploy-fleetd.sh): a self-matching wrapper shell (
sh -c 'echo "target/fleetd.jar" >/dev/null; sleep 20', a non-exec'ing shell that forks for its second statement and stays alive holding the pattern in its own argv) must be excluded fromrunning_pid()'s output; a real second daemon-shaped process (simulated safely viaexec -aon a harmlesssleep, never a real daemon) must still be found; and the die message must not recommend the self-matchingpgrep -fcommand. I confirmed the first test fails against the pre-fix implementation (reverted running_pid() locally, reran, watched it fail, then restored the fix) -- the regression is genuinely caught, not vacuous.Verification run in the worktree:
bash scripts/test-redeploy-fleetd.sh-> exit 0,PASS: redeploy log classifier(the interspersedmktemp/mutation-test FAIL lines are expected noise from existing mutation tests, present in the baseline run too, unrelated to this change).mvn clean installin fleetd/ ->Tests run: 1789, Failures: 0, Errors: 0, Skipped: 0/BUILD SUCCESS(this bash test script is not wired into the Maven build; it's run separately, matching .gitea/workflows/ci.yml).Caveat for review: everything above was verified on macOS/BSD only, per the ticket's own note that instance 2 is latent (not reproducible) on this Mac without a deliberately constructed wrapper. I could not verify on a real Linux/systemd host -- that's instance 1's territory anyway, which this PR does not touch. The shell-name exclusion list (sh/bash/zsh/dash/ksh) covers the common shells; an unusual shell not on that list would still be excluded from nothing (fail open, matching the pre-fix behavior for anything not on the list), not a new failure mode.
Never ran scripts/redeploy-fleetd.sh itself, and never touched the live daemon.