fleetd #492: systemd --user as a third supervisor in redeploy-fleetd.sh #495
Closed
agent
wants to merge 0 commits from
worker/492-209647-1 into main
pull from: worker/492-209647-1
merge into: fleet:main
fleet:main
fleet:worker/fleetd-612-unita-87807e-1
fleet:worker/612-b3-mcpwirings-da2b58-3
fleet:worker/612-b2-cb185-176d3a-2
fleet:worker/612-b1-completion-457459-1
fleet:worker/612-agaps-73a926-2
fleet:worker/608-sleeps-3a64ff-3
fleet:worker/621-b4520b-1
fleet:worker/618-b83894-2
fleet:worker/fleetd-615-e05481-5
fleet:worker/lead-autocompact-5f1ab2-3
fleet:worker/fleetd-613-f85deb-3
fleet:worker/fleetd-608-flaky-nudge-test-d0c2d1-3
fleet:worker/lead-context-gauge-ad404f-1
fleet:worker/gauge-wiring-9158c1-4
fleet:worker/redeploy-slowstart-ead0e5-5
fleet:worker/charter-bytes-13668c-6
fleet:worker/rollover-outcome-291483-2
fleet:worker/589-f64303-2
fleet:worker/593-1a8025-5
fleet:worker/589-fcd2aa-1
fleet:worker/568-9fdaa2-3
fleet:worker/571-attempted-outcome-5739f7-2
fleet:worker/581-completionresolver-cas-sites-0542b7-6
fleet:worker/562-loop-health-wiring-test-99611c-5
fleet:worker/562-surface-loop-health-7df5cc-4
fleet:worker/575-waiter-cleanup-sites-62ad80-1
fleet:worker/572-answer-lock-release-46a9ae-5
fleet:worker/567-probe-channel-leak-a38fc5-6
fleet:worker/551-record-before-send-7cbf56-1
fleet:worker/561-listener-fanout-survives-a-throw-61d538-2
fleet:worker/555-redeploy-main-flow-seam-65c2f5-2
fleet:worker/556-injector-owns-registration-e027a5-1
fleet:worker/552-post-restart-mktemp-abort-bc2672-4
fleet:worker/553-onstatus-completion-leak-0da881-2
fleet:worker/550-shasum-linux-196132-1
fleet:worker/538-loop-dies-on-error-4a5eeb-6
fleet:worker/426-health-coverage-ef1fd4-4
fleet:worker/504-failed-reported-clean-3cfd66-3
fleet:worker/537-capturedlog-close-e4c437-2
fleet:worker/459-broken-link-targets-cadc17-5
fleet:worker/535-appender-leak-fe74c1-1
fleet:worker/512-part2-shutdown-detection-434701-9
fleet:worker/529-logger-level-sweep-2a5533-8
fleet:worker/528-drain-gate-call-site-5de83d-7
fleet:charter/forge-mcp-vs-token
fleet:worker/521-swap-guard-unpinned-28e931-5
fleet:worker/519-probe-test-harness-d25ab8-4
fleet:worker/525-logger-level-leak-1b4eb0-6
fleet:worker/518-fleetmcp-resolver-wiring-8ef96c-1
fleet:worker/512-drain-complete-line-7edd71-3
fleet:worker/517-abort-branch-and-jar-id-41b641-2
fleet:worker/500-9e52c9-3
fleet:worker/509-4912f4-2
fleet:worker/511-9a4b23-1
fleet:worker/493-479f45-2
fleet:worker/505-03f8b2-1
fleet:worker/492-followup-detect-unclear
fleet:worker/501-a31fa0-7
fleet:worker/498-451d1c-5
fleet:worker/494-1015ce-2
fleet:worker/489-001902-2
fleet:worker/480-relative-handover-path-906323-1
fleet:worker/480-b-handover-skill-45bf1f-5
fleet:worker/474-followup-source-pin-f54a55-17
fleet:worker/474-charter-check-on-reload-f54a55-17
fleet:worker/466-quarantine-repeatcount-report
fleet:worker/393-opencode-skill-seeding-71854b-13
fleet:worker/469-canonical-tool-names-2a472a-16
fleet:worker/466-quarantine-escalation-5ae9c1-15
fleet:worker/446-hot-exhausted-pattern-0af580-6
fleet:worker/464-charter-tool-name-guard-a85635-12
fleet:worker/463-listfleet-default-fails-open-f1c76c-11
fleet:worker/458-invariant-5-by-purpose-862f9a-10
fleet:worker/439-coordinator-row-gate-bc032a-8
fleet:worker/449-herdr-protocol-576015-4
fleet:worker/450-abstract-spawn-599e1c-5
fleet:worker/437-ack-refuses-177d91-1
fleet:worker/444-placement-window-feb56a-2
fleet:worker/440-helddurable-derived-d462d7-13
fleet:worker/425-rework-placement-resolve-c58ba1-9
fleet:worker/421-lead-peek-held-msgs-cdbad2-10
fleet:worker/435-fixed-policy-cap-fe11de-12
fleet:worker/422-gate-state-observability-9e79d6-11
fleet:worker/431-memberregistry-live-readers-cdbad2-10
fleet:worker/424-architect-slot-hot-038b41-7
fleet:worker/422-model-gate-spawn-c29f48-6
fleet:worker/425-default-profile-live-f55534-8
fleet:worker/415-coverage-wording-2cbf9c-5
fleet:worker/416-3ad1da-1
fleet:worker/418-588283-3
fleet:worker/deterministic-stamp-race-409-3cb7b6-10
fleet:worker/armed-reads-live-config-404-ed931f-9
fleet:worker/reply-peer-refusal-391-5a34bd-7
fleet:worker/models-allowlist-aa9e9b-3
fleet:worker/ttl-stamp-race-399-f1122f-8
fleet:worker/scrub-receipt-400-316b3e-5
fleet:worker/exhaustion-detection-395-105105-6
fleet:worker/scrub-abort-394-316b3e-5
fleet:fix/scrub-uid-abort
fleet:worker/task-scrub-517574-2
fleet:worker/t386-clock-bd5b78-4
fleet:worker/t384-scrub-813790-5
fleet:worker/t381-cc-748314-2
fleet:worker/t373-336973-2
fleet:worker/t365-3920c5-3
fleet:worker/t358-6e989b-1
fleet:worker/t355-8b321c-1
fleet:worker/fleetd-369-hermetic-git-tests-e8b19a-3
fleet:worker/fleetd-368-stale-lead-binding-f5682e-2
fleet:worker/fleetd-360-deploy-units-0d3793-1
fleet:worker/359-dead-lead-tabs-f1253b-4
fleet:worker/362-worktree-skills-c03e51-3
fleet:worker/361-coord-visibility-655144-1
fleet:362-plugin-visibility-and-drift
fleet:worker/errscan-bed2ca-2
fleet:worker/amqp-log-identity-bed2ca-2
fleet:worker/withdefaults-guard-561704
fleet:worker/sleepguard-82076d-1
fleet:worker/fd334-9ee1b6-5
fleet:worker/fd348-f1ab27-4
fleet:worker/fd335-a71c35-1
fleet:worker/fd342-174a17-2
fleet:worker/fd345-490d0f-3
fleet:worker/fleetd-337-5ec7d4-21
fleet:worker/fleetd-341-af5a6b-24
fleet:worker/fleetd-339-5ca0a2-23
fleet:worker/fleetd-338-83a4a1-22
fleet:worker/fleetd-333-281f46-18
fleet:worker/fleetd-329-11bdbb-16
fleet:worker/fleetd-330-2770fb-17
fleet:worker/fix-326-50506e-15
fleet:worker/fix-324-3e9bbf-14
fleet:worker/fix-323-b8287d-13
fleet:worker/fix-316b-bd0860-11
fleet:worker/fix-318-76ca36-9
fleet:worker/fix-317-486aec-8
fleet:worker/fix-315-ce47c5-6
fleet:worker/fix-307-275890-6
fleet:worker/fix-308-b4f664-7
fleet:worker/fix-309-ec3939-8
fleet:worker/fix-310-7a3974-9
fleet:worker/fix-302-52ad0e-9
fleet:worker/fix-298-ce1acb-8
fleet:worker/fix-297-66bd11-7
fleet:worker/fix-296-104622-6
fleet:worker/fix-293-bare-closetab-eb22b5-3
fleet:worker/fix-280-gone-ask-lapse-bca98e-2
fleet:worker/fix-290-reapidle-guard-coverage-9b0dd1-1
fleet:worker/fix-285-trust-seed-8f3565-10
fleet:worker/fix-284-backend-error-seat-85912c-11
fleet:worker/fix-282-chained-ask-e6d0bb-8
fleet:worker/fix-283-teardown-leaks-f40dfa-9
fleet:worker/fix-281-pin-handler-actions-4921ac-7
fleet:worker/audit-rendezvous-lifecycle-d072ae-2
fleet:worker/audit-health-placement-1a2476-6
fleet:worker/audit-teardown-exits-e207a5-3
fleet:worker/audit-launcher-asymmetry-27e370-4
fleet:worker/audit-rest-authz-6ca53c-5
fleet:worker/investigate-275-abandon-asking-fdef52-8
fleet:worker/fix-274-worktree-leak-b0095d-7
fleet:worker/fix-273-exhausted-pattern-9665b5-6
fleet:worker/fleetd-267-model-check-bd8068-1
fleet:worker/fleetd-131-archunit-18b834-7
fleet:worker/fleetd-266-sshagent-rename-a014ff-6
fleet:worker/fleetd-184-uid-claim-8e1f31-4
fleet:worker/fleetd-184-warn-b381ee-10
fleet:worker/fleetd-184-docs-be1d12-9
fleet:worker/fleetd-257-9bf010-7
fleet:worker/fleetd-103-23a113-6
fleet:worker/fleetd-247-342356-5
fleet:worker/fleetd-116-04dea8-4
fleet:worker/fleetd-252-a830e0-3
fleet:worker/fleetd-111-7e8673-9
fleet:worker/fleetd-155c-f8ef4b-8
fleet:worker/fleetd-176-b928ca-3
fleet:worker/fleetd-249-7a7878-2
fleet:worker/cb248-composition-root-b-9acdf7-15
fleet:worker/cb148-envrc-default-fa6c82-12
fleet:worker/cb201-unit5-wiring-6c12e6-8
fleet:worker/cb241-fallback-echo-1175e9-11
fleet:worker/cb149-trust-dialog-2392a5-9
fleet:worker/cb134-148-overlay-visible-c9b986-10
fleet:worker/cb234-session-id-keyed-04e1fc-1
fleet:worker/cb201-unit3-nudge-abdf5c-6
fleet:worker/cb201-unit2-policy-c1102c-5
fleet:worker/cb201-unit4-outcome-a13bfa-7
fleet:worker/cb201-unit1-classifier-91b9b1-4
fleet:worker/cb201-227-refine-831980-3
fleet:worker/cb175-model-readback-0f085f-1
fleet:worker/cb222-charter-tmpdir-17f013-1
fleet:worker/cb226-architect-slot-race-cd3aa8-3
fleet:worker/cb224-worktree-root-group-024523-2
fleet:worker/cb-123-role-demotion-c600f7-2
fleet:worker/cb-219-opencode-roots-1f677e-1
fleet:worker/cb214-claude-session-id-b9eab4-4
fleet:worker/cb213-zdotdir-wrong-process-dd6de4-3
fleet:worker/cb211-exhaustion-classification-9546e0-2
fleet:worker/cb137-ambiguous-task-4df3d8-4
fleet:worker/cb209-agentsessionid-4dfdb6-2
fleet:worker/cb185-hostenvnames-2692b5-3
fleet:worker/cb206-opencode-sqlite-128718-2
fleet:worker/cb185-worktree-group-fc0c99-1
fleet:worker/cb-137-ask-ticket-e7760c-2
fleet:worker/cb-172-broker-uri-d36ae4-4
fleet:worker/cb-175-model-readback-76ead6-3
fleet:worker/cb-161-pane-ancestry-293510-1
fleet:worker/cb-164-rebase-885863-8
fleet:worker/cb-164-empty-scrape-false-success-1a80af-3
fleet:fix/cb-197-ticket-ttl-from-completion
fleet:worker/cb-189-remote-url-coverage-4692f3-1
fleet:worker/cb-185-blockers-027756-4
fleet:worker/cb-192-gap-log-11b631-2
fleet:worker/cb-633-fix-5f4396-3
fleet:worker/cb185-router-d6436d-3
fleet:worker/cb185-router-routing-gaps-9e9d33-3
fleet:worker/cb185-paneids-992586-2
fleet:worker/cb-633-allow-list-union-ed374b-1
fleet:worker/cb-157-credential-in-remote-url-496e44-2
fleet:worker/cb-641-health-herdr-evidence-8f1f54-6
fleet:worker/cb-640-health-msg-evidence-99c9cd-1
fleet:worker/cb-642-fleets-status-skill-bbbc40-5
fleet:cb-634-ide-mcp
fleet:worker/lead-comms-wiring-c014b9-7
fleet:worker/lead-mailbox-c19577-6
fleet:worker/autocompact-window-82bc2f-5
fleet:worker/cb-634-probe-18056f-4
fleet:worker/cb635-broker-urienv
fleet:worker/cb-632-config-retry-8e0efa-7
fleet:lead/cb-622e-claude-md
fleet:lead/cb-622-followup
fleet:worker/cb-622a-165dff-1
fleet:lead/cb-622d-opencode-mount
fleet:worker/cb-622b-717c67-2
fleet:worker/cb-622c-ab7759-3
fleet:worker/cb-617b2-20ca4b-3
fleet:worker/cb-617a-5c2f4a-1
fleet:worker/cb596-4e49ef-3
fleet:worker/cb586-10500c-1
fleet:worker/cb-606-b9343a-25
fleet:worker/cb604-1445f8-24
fleet:worker/cb582-477374-21
fleet:worker/cb584-8c2281-22
fleet:worker/cb600-e6b9a9-20
fleet:worker/cb602-ce257f-19
fleet:worker/cb601-b42837-18
fleet:worker/cb598-6c7ba7-17
fleet:worker/cb599-740fe4-16
fleet:worker/cb597-282224-15
fleet:worker/cb590fix-185e9a-10
fleet:worker/cb528-recovery-race
fleet:worker/cb594-96bead-8
fleet:worker/cb590-916766-2
fleet:worker/cb527-997d99-3
fleet:worker/cb592-env-leak-3cbf9c-1
fleet:worker/cb588-async-ticket-nudge-3218f7-5
fleet:worker/cb578b-9dcb13-6
fleet:worker/cb581-d24826-5
fleet:worker/m2-u5-ef8c42-15
fleet:worker/cb578a-516499-2
fleet:worker/cb576-01a04b-17
fleet:worker/cb579-lead-tab-acba06-20
fleet:worker/cb580-terminal-health-ed6058-21
fleet:worker/cb577-f36fdc-18
fleet:worker/cb573b-3db06f-16
fleet:worker/cb568c-f36fdc-18
fleet:worker/cb568-drop-cause-c3ac1c
fleet:worker/cb575-cancelled-notification-c3ac1c
fleet:worker/m4-sol-a2cbec-3
fleet:worker/cb574-async-ask-c3ac1c
fleet:worker/cb573-health-model-8ca857-14
fleet:worker/cb572-unknown-target-7f2e35-13
fleet:worker/u4-700706-9
fleet:worker/u3-b9fcb6-6
fleet:worker/u2-ef5b68-4
fleet:worker/u1-469dce-1-clean
fleet:worker/u1-469dce-1
fleet:worker/cb-564-health-events-70cf7e-2
fleet:worker/cb-565-recycle-drops-role-98e58f-3
fleet:worker/cb-563-missing-reply-df2866-1
fleet:worker/cb-562-readiness-gate-silent-6c23c9-3
fleet:worker/cb-560-architect-presence-da8155-1
fleet:worker/cb-561-architect-silent-off-a71cab-2
fleet:worker/cb-548-bind-architect-slot-fe1b8c-1
fleet:worker/parity-overlay-settings-5fb711-1
fleet:secrets-central-store
fleet:cb-559-hot-key-correction
fleet:cb-557-fleet-role-pools
fleet:worker/cb-553-maxload-explicit-spawn-305ee3-6
fleet:worker/cb-551-idle-lead-heartbeat-f1633c-1
fleet:worker/cb-544-drain-preserves-worktree-925fad-3
fleet:worker/cb-552-docs-sync-1cb9cf-4
fleet:worker/cb-548-rendezvous-guard-rebased
fleet:worker/cb-548-rendezvous-guard-116b53-10
fleet:worker/cb-548-authz-v2-586df6-8
fleet:worker/cb-548-authz-264363-5
fleet:salvage/cb-528b-codex-home
fleet:salvage/cb-528a-codex-launcher
fleet:CB-518-primary-flow
fleet:feature/peer-launcher-spi
fleet:cb-103-injector
No Reviewers
Labels
Clear labels
blocked
needs-live-proof
ready-to-delegate
silent-default
Cannot start until something else lands. The body says what.
Merged and green, but never shown working on the running daemon. Not the same as done.
Scope, files and acceptance criteria are written. A worker can be briefed from the body alone.
A feature that compiles, passes tests, and ships turned off. Nine recurrences and counting.
No Label
Milestone
No items
No Milestone
Projects
Clear projects
No project
Notifications
Due Date
No due date set.
Dependencies
No dependencies set.
Reference: fleet/fleetd#495
Reference in New Issue
Block a user
Blocking a user prevents them from interacting with repositories, such as opening or commenting on pull requests or issues. Learn more about blocking a user.
Delete Branch "worker/492-209647-1"
Deleting a branch is permanent. Although the deleted branch may continue to exist for a short time before it actually gets removed, it CANNOT be undone in most cases. Continue?
fleetd #492 — systemd --user as a third supervisor
Problem:
scripts/redeploy-fleetd.shonly knew launchd. On a Linux host running fleetd under asystemd --userunit withRestart=on-failure, the script fell into its unsupervised branch:kill "$OLD_PID". Since a SIGTERM'd JVM exits 143 even when its shutdown hook completes cleanly (CB-594), systemd treats that as a failure and revives the OLD jar — racing the script's own start of the NEW jar. Every check the script ran (healthz 200, jar id, fresh "listening" line) was satisfied by EITHER daemon, so the script reportedokwhile two daemons were running against one herdr session.Fix:
detect_supervisor()now returns one of four answers:launchd,systemd,none, orambiguous(both signals fire at once — the one case this script genuinely cannot tell apart).require_drivable_supervisor()die()s onambiguous(or any value detect_supervisor did not return) — it never falls through tokill. This runs in the state-report section, so--checkreports the refusal too, without touching anything.systemctl --user stop|start "$SYSTEMD_UNIT", mirroring the existing launchd branch's structure. Unlike launchd's KeepAlive, systemd'sRestart=on-failuredoes not fire on a deliberatestop, so no unload/disable dance is needed there.assert_single_daemon()runs after start+verify anddie()s ifpgrep -f "$PATTERN"finds more than one pid — the exact symptom this ticket exists to catch, which none of the existing checks can see.--checkstays read-only.Testing seam:
detect_supervisor/require_drivable_supervisor/count_daemon_pids/assert_single_daemonare pure functions.launchd_installed/launchd_loaded/systemd_installed/systemd_loadedare simple overridable functions (same pattern as the existing launchd probes) soscripts/test-redeploy-fleetd.shcan simulate launchd-only, systemd-only, neither, and both-at-once without a real launchd or systemd on this (macOS) dev box.Verification (measured in-worktree, not derived by reading):
bash scripts/test-redeploy-fleetd.sh→ exit 0, prints the 3 existing mutation-proof lines +PASS: redeploy log classifier(11 test invocations).require_drivable_supervisoracceptsambiguousinstead of refusing — proved applied via two greps: mutant pattern present, original pattern absent): tests → 1 failure, exit 1:FAIL: require_drivable_supervisor accepted an ambiguous (undrivable) supervisor.assert_single_daemonaccepts 2 pids, i.e.-gt 1→-gt 2, proved applied the same way): tests → 1 failure, exit 1:FAIL: assert_single_daemon accepted two simultaneously running pids.cd fleetd && mvn clean install:BUILD SUCCESS,Tests run: 1677, Failures: 0, Errors: 0, Skipped: 0. (Nopom.xmlat the repo root — only underfleetd/— so this ran from the module root per the implementer skill, not literally "the repo root".)Out of scope, spotted not fixed:
dev.ltms.fleet.deploy.SystemdUnitSafetyTest(Java, 8 tests, all passing) already exists in the main build — unrelated to this shell-script ticket, not touched.redeploy-fleetd.shbuilding into the LIVEtarget/fleetd.jarpath before stopping the daemon — a separate, real defect in this same script, out of scope here and not touched.No IDE MCP tools were available to me as a worker; verification is
bash -n, the shell test suite above, andmvn clean install.Not merging yet.
nonemeans two different things, and only one of them is safe.Good work overall — the overridable-probe seam,
require_drivable_supervisor, andassert_single_daemonare the right shapes, and the two mutations you ran are proper proofs. The problem is one level up, indetect_supervisor.detect_supervisor(scripts/redeploy-fleetd.sh:123) reads only the two*_loadedprobes. Sononecurrently means both:kill+nohupfallback is correct for it, andrequire_drivable_supervisoracceptsnonewithout question, so the refuse-don't-guess gate never sees the second case.Input 1 — unit installed, not active at this instant. I measured this.
Run in a detached worktree at
dcd5052, sourcing the script behind its ownSOURCEDguard and overriding the probes, exactly as your harness does:systemctl --user is-activeexits non-zero foractivating,deactivating,failed, and while an auto-restart is pending. Every one of those is a host that is under systemd. The script then starts an unsupervised daemon withnohupwhile systemd is about to start its own. That is the two-daemon race this ticket exists to prevent — it ran for 20 hours on fleet01 and only an unrelated consumer count revealed it.systemd_installedalready knows the answer. It is computed at:294for awarnand then never consulted by the decision.Input 2 —
systemctlerrors instead of answering. Found by the reviewer; I could not test it here.Both probes are
command -v systemctl && systemctl … 2>/dev/null. Asystemctlthat runs but cannot reach the user bus — a headless ssh session with no lingering, which is exactly the second-host case — fails the same way a real "no" does. Same landing:none, same fallback. This is the reviewer's finding, reasoned from the code; I have no systemd on this host, so I did not measure it.Two different inputs, one defect. That makes them two data points, not one.
Asked for
detect_supervisormust return a fourth undrivable answer — call itunclear— for installed but not loaded, on both supervisors.require_drivable_supervisormustdie()on it, with a message that names which supervisor looked present and what to check.2>/dev/null; a probe that errors must land inunclear, never innone.nonemust keep meaning only proven unsupervised: no plist, no unit file, and both probes answered cleanly.return 0/return 1bodies, so no test ever exercises a probe that errors. At least one test must drive the real function body with asystemctlstub that exits non-zero and writes to stderr.Keep everything else. The mutation discipline in your report is exactly right and I want the same on the new tests: prove each mutant applied with two different search strings, and run a control after restoring.
Note for whoever picks this up: #493 also changes this file (it builds into the live
target/fleetd.jarbefore stopping the daemon). #493 lands after this PR, not before.Two more requirements, to apply at review. Both found by the fleet01 lead.
A worker is already on the
unclearstate. These two go on top, and the first one is a trap that the fix itself creates.1. Adding a state without a default arm makes the script worse, not better
bash has no exhaustiveness check. A
casewith no*)arm silently matches nothing and carries on. Measured ondcd5052:require_drivable_supervisorruns once, at:305. The three switches that actually drive the daemon each enumerate exactly three values. So if a new state reaches them — because a future edit adds one, or because the guard is moved, weakened, or skipped on some path — nothing matches in the stop switch, the daemon is never stopped, nothing matches in the start switch, and the script reports no error at any point.Requirement:
:308,:419and:477each get a*) die …arm naming the unhandled value. Then a fourth state fails loudly at every site instead of silently skipping the stop or the start.The guard staying is necessary but not sufficient. One gate at the top of a script is the same shape as a one-way gate: it closes the direction you came from.
2. Two different tests, and the easy one is not the useful one
That is the same gap as asserting the nudge happens versus asserting the wait waits on #489, and it is the gap that let this defect into
dcd5052in the first place — every existing test stubssystemd_loadedwith a cleanreturn 0/return 1, so no test ever drives the real body.Both are required. The brief already asks for a
systemctlstub that exits non-zero and writes to stderr; that is the second kind. Do not let it be replaced by a direct construction of the state.Filed separately, on purpose
The underlying shape is now #497 — a sentinel that conflates "measured: no" with "could not measure", with three instances across three subsystems. It is explicitly not the same family as #494, and the fleet01 lead was firm about not merging the two: #494's fix is compute the value honestly, and here the value already is honest — the vocabulary is one symbol short. A worker who reads #494's rule against this code will confirm
noneis genuinely what the probe returned and close it.Amending my own requirement twice. One arm was wrong, and one worry does not apply here.
1.
*) dieis the wrong arm for the reporting switch. Use*) echo.I asked for
*) die …on all three switches. That is wrong for:308.The fleet01 lead's point, and I agree on sight:
:308is the diagnostic path. A reporting channel that aborts on a value it does not recognise goes silent exactly when the state is novel — which is #497's instance 1 one level up, the failure this whole line of work is ranking against. Stopping the report was never the guard's purpose.Choose the arm by what the caller does with the value:
:308*) echo "unknown supervisor state: $SUPERVISOR_KIND"— keep the operator informed:419*) die …:477*) die …A blanket "every switch carries
*) die" gets the two drivable ones right and turns the diagnostic one into the defect.2. Their
set -ewarning is real bash, and it does not reach this file. Both halves measured.They warned that
*) dieis defeated by the calling context, since adieinside a command substitution exits only the subshell. I reproduced all of it on this host — bash 5.3.9 via/usr/bin/env bash, not macOS's system 3.2 — with the guard byte-identical in every case and only the call site varying:A fires, so the guard is correctly written in all six; the difference is entirely context. Their table is right, and F shows
pipefailis what closes E.But the preconditions are not present in this script, and I checked each one rather than assuming:
So all three guarded switches are case A, both substitutions are case C, and there is no case D anywhere. The
*) diearms will fire.3. What to carry forward anyway — as a standing constraint, not a fix
The mechanism is real even though it does not bite today, and two single-line edits would make it bite:
set -euo pipefailfrom:50. Without it,SUPERVISOR_KIND="$(detect_supervisor)"at:304silently yields an empty string and the script continues.if VAR="$(detect_supervisor)"; thenor pipe these functions. Case D suppressesset -eeven when it is set, and no convention about the switch can catch it, because the defect is at the caller.detect_supervisor,:304becomes the load-bearing site. It is safe today only becauseset -eis on.Add those three as a comment block above
detect_supervisorso the next contributor reads them before editing.Closing: this branch's work is already on
main, merged as part of PR #499, not abandoned.PR #499 was built on top of this branch rather than on
main— its own description says"based on
origin/worker/492-209647-1@dcd5052, NOT main". So merging #499 carried every commithere in with it.
Measured just now, in the main clone:
mainis at136312f. Nothing from this PR is lost, and there is nothing left to merge here.Thanks for the work — the systemd-as-third-supervisor change is live in
scripts/redeploy-fleetd.sh,and
--checkruns clean on this Mac against the merged script (supervisor detected: launchd).Pull request closed