#664 follow-up: the runtime-vs-build path split is unpinned by tests, and nothing checks the installed plist's jar path #680
Closed
opened 2026-10-03 21:14:36 +02:00 by ltms
·
4 comments
No Branch/Tag Specified
main
worker/702-4f5c7f-2
worker/715-5c43fc-1
worker/721-70f9ea-5
worker/718-99362b-2
worker/task-15-af0d10-12
worker/task-16-50a702-13
worker/task-12-4d0479-9
worker/task-13-823ce2-10
worker/705-ticket-owner-af9928-8
worker/703-list-collaborators-9c06c2-7
worker/669-example-truth-0b303d-6
worker/669-collab-deliverability-9ba859-3
worker/669-collab-reload-report-2a21bd-4
worker/669-7e80a6-1
worker/669-unit-d-efbbd7-1
worker/669-1b786a-1
worker/669-1d1d9f-1
worker/692-4afb9d-2
worker/689-02fced-13
worker/693-cf23fa-14
worker/677-fix-lead-collision-f69073-12
worker/638-fix-overmask-dbb1bf-11
worker/675-5b7478-4
worker/669-unit-a-70cc8f-3
worker/677-8cdaaf-5
worker/638-a7b391-1
worker/683-4536d6-2
worker/651-a75bbe-8
worker/680-20607d-7
worker/664-c12e95-3
worker/668-08534d-4
worker/672-0f2469-2
worker/670-7d1022-1
worker/661-ac7c28-2
worker/664-37fb9b-3
worker/663-remove-3arg-read-3f6783-1
worker/659-remove-dead-backcompat-ba5e6f-1
worker/637-revision-60a488-23
worker/656-redact-regression-tests-892903-19
worker/637-context-gauge-threshold-466eb5-16
worker/639-redact-line-numbers-de4ac4-17
worker/641-set-reformat-guard-6f96a4-18
worker/642-herdr-guard-scope-5de0e4-15
worker/650-javadoc-scope-95f3b3-14
worker/612-01e9f7-13
worker/612-a-r4-quarantine-outage-7ab0e8-5
worker/612-a-r9-r11-capacity-coverage-peers-cfcc79-7
worker/612-a-r10-loophealth-ccc872-8
worker/612-a-r12-turnregistrar-9e3bb7-9
worker/612-a-r5-leadconfigdir-9e70cf-6
lead/config-edit-redact-anchor-wording
worker/config-edit-seam-ca8dc1-1
worker/612-r67-630-lifecycle-290b8d-3
worker/629-625-ports-seams-da7d5d-4
worker/612-r12-exhaustion-f37cd7-1
worker/612-r38-amqp-24b083-2
worker/fleetd-612-unita-87807e-1
worker/612-b3-mcpwirings-da2b58-3
worker/612-b2-cb185-176d3a-2
worker/612-b1-completion-457459-1
worker/612-agaps-73a926-2
worker/608-sleeps-3a64ff-3
worker/621-b4520b-1
worker/618-b83894-2
worker/fleetd-615-e05481-5
worker/lead-autocompact-5f1ab2-3
worker/fleetd-613-f85deb-3
worker/fleetd-608-flaky-nudge-test-d0c2d1-3
worker/lead-context-gauge-ad404f-1
worker/gauge-wiring-9158c1-4
worker/redeploy-slowstart-ead0e5-5
worker/charter-bytes-13668c-6
worker/rollover-outcome-291483-2
worker/589-f64303-2
worker/593-1a8025-5
worker/589-fcd2aa-1
worker/568-9fdaa2-3
worker/571-attempted-outcome-5739f7-2
worker/581-completionresolver-cas-sites-0542b7-6
worker/562-loop-health-wiring-test-99611c-5
worker/562-surface-loop-health-7df5cc-4
worker/575-waiter-cleanup-sites-62ad80-1
worker/572-answer-lock-release-46a9ae-5
worker/567-probe-channel-leak-a38fc5-6
worker/551-record-before-send-7cbf56-1
worker/561-listener-fanout-survives-a-throw-61d538-2
worker/555-redeploy-main-flow-seam-65c2f5-2
worker/556-injector-owns-registration-e027a5-1
worker/552-post-restart-mktemp-abort-bc2672-4
worker/553-onstatus-completion-leak-0da881-2
worker/550-shasum-linux-196132-1
worker/538-loop-dies-on-error-4a5eeb-6
worker/426-health-coverage-ef1fd4-4
worker/504-failed-reported-clean-3cfd66-3
worker/537-capturedlog-close-e4c437-2
worker/459-broken-link-targets-cadc17-5
worker/535-appender-leak-fe74c1-1
worker/512-part2-shutdown-detection-434701-9
worker/529-logger-level-sweep-2a5533-8
worker/528-drain-gate-call-site-5de83d-7
charter/forge-mcp-vs-token
worker/521-swap-guard-unpinned-28e931-5
worker/519-probe-test-harness-d25ab8-4
worker/525-logger-level-leak-1b4eb0-6
worker/518-fleetmcp-resolver-wiring-8ef96c-1
worker/512-drain-complete-line-7edd71-3
worker/517-abort-branch-and-jar-id-41b641-2
worker/500-9e52c9-3
worker/509-4912f4-2
worker/511-9a4b23-1
worker/493-479f45-2
worker/505-03f8b2-1
worker/492-followup-detect-unclear
worker/501-a31fa0-7
worker/498-451d1c-5
worker/494-1015ce-2
worker/492-209647-1
worker/489-001902-2
worker/480-relative-handover-path-906323-1
worker/480-b-handover-skill-45bf1f-5
worker/474-followup-source-pin-f54a55-17
worker/474-charter-check-on-reload-f54a55-17
worker/466-quarantine-repeatcount-report
worker/393-opencode-skill-seeding-71854b-13
worker/469-canonical-tool-names-2a472a-16
worker/466-quarantine-escalation-5ae9c1-15
worker/446-hot-exhausted-pattern-0af580-6
worker/464-charter-tool-name-guard-a85635-12
worker/463-listfleet-default-fails-open-f1c76c-11
worker/458-invariant-5-by-purpose-862f9a-10
worker/439-coordinator-row-gate-bc032a-8
worker/449-herdr-protocol-576015-4
worker/450-abstract-spawn-599e1c-5
worker/437-ack-refuses-177d91-1
worker/444-placement-window-feb56a-2
worker/440-helddurable-derived-d462d7-13
worker/425-rework-placement-resolve-c58ba1-9
worker/421-lead-peek-held-msgs-cdbad2-10
worker/435-fixed-policy-cap-fe11de-12
worker/422-gate-state-observability-9e79d6-11
worker/431-memberregistry-live-readers-cdbad2-10
worker/424-architect-slot-hot-038b41-7
worker/422-model-gate-spawn-c29f48-6
worker/425-default-profile-live-f55534-8
worker/415-coverage-wording-2cbf9c-5
worker/416-3ad1da-1
worker/418-588283-3
worker/deterministic-stamp-race-409-3cb7b6-10
worker/armed-reads-live-config-404-ed931f-9
worker/reply-peer-refusal-391-5a34bd-7
worker/models-allowlist-aa9e9b-3
worker/ttl-stamp-race-399-f1122f-8
worker/scrub-receipt-400-316b3e-5
worker/exhaustion-detection-395-105105-6
worker/scrub-abort-394-316b3e-5
fix/scrub-uid-abort
worker/task-scrub-517574-2
worker/t386-clock-bd5b78-4
worker/t384-scrub-813790-5
worker/t381-cc-748314-2
worker/t373-336973-2
worker/t365-3920c5-3
worker/t358-6e989b-1
worker/t355-8b321c-1
worker/fleetd-369-hermetic-git-tests-e8b19a-3
worker/fleetd-368-stale-lead-binding-f5682e-2
worker/fleetd-360-deploy-units-0d3793-1
worker/359-dead-lead-tabs-f1253b-4
worker/362-worktree-skills-c03e51-3
worker/361-coord-visibility-655144-1
362-plugin-visibility-and-drift
worker/errscan-bed2ca-2
worker/amqp-log-identity-bed2ca-2
worker/withdefaults-guard-561704
worker/sleepguard-82076d-1
worker/fd334-9ee1b6-5
worker/fd348-f1ab27-4
worker/fd335-a71c35-1
worker/fd342-174a17-2
worker/fd345-490d0f-3
worker/fleetd-337-5ec7d4-21
worker/fleetd-341-af5a6b-24
worker/fleetd-339-5ca0a2-23
worker/fleetd-338-83a4a1-22
worker/fleetd-333-281f46-18
worker/fleetd-329-11bdbb-16
worker/fleetd-330-2770fb-17
worker/fix-326-50506e-15
worker/fix-324-3e9bbf-14
worker/fix-323-b8287d-13
worker/fix-316b-bd0860-11
worker/fix-318-76ca36-9
worker/fix-317-486aec-8
worker/fix-315-ce47c5-6
worker/fix-307-275890-6
worker/fix-308-b4f664-7
worker/fix-309-ec3939-8
worker/fix-310-7a3974-9
worker/fix-302-52ad0e-9
worker/fix-298-ce1acb-8
worker/fix-297-66bd11-7
worker/fix-296-104622-6
worker/fix-293-bare-closetab-eb22b5-3
worker/fix-280-gone-ask-lapse-bca98e-2
worker/fix-290-reapidle-guard-coverage-9b0dd1-1
worker/fix-285-trust-seed-8f3565-10
worker/fix-284-backend-error-seat-85912c-11
worker/fix-282-chained-ask-e6d0bb-8
worker/fix-283-teardown-leaks-f40dfa-9
worker/fix-281-pin-handler-actions-4921ac-7
worker/audit-rendezvous-lifecycle-d072ae-2
worker/audit-health-placement-1a2476-6
worker/audit-teardown-exits-e207a5-3
worker/audit-launcher-asymmetry-27e370-4
worker/audit-rest-authz-6ca53c-5
worker/investigate-275-abandon-asking-fdef52-8
worker/fix-274-worktree-leak-b0095d-7
worker/fix-273-exhausted-pattern-9665b5-6
worker/fleetd-267-model-check-bd8068-1
worker/fleetd-131-archunit-18b834-7
worker/fleetd-266-sshagent-rename-a014ff-6
worker/fleetd-184-uid-claim-8e1f31-4
worker/fleetd-184-warn-b381ee-10
worker/fleetd-184-docs-be1d12-9
worker/fleetd-257-9bf010-7
worker/fleetd-103-23a113-6
worker/fleetd-247-342356-5
worker/fleetd-116-04dea8-4
worker/fleetd-252-a830e0-3
worker/fleetd-111-7e8673-9
worker/fleetd-155c-f8ef4b-8
worker/fleetd-176-b928ca-3
worker/fleetd-249-7a7878-2
worker/cb248-composition-root-b-9acdf7-15
worker/cb148-envrc-default-fa6c82-12
worker/cb201-unit5-wiring-6c12e6-8
worker/cb241-fallback-echo-1175e9-11
worker/cb149-trust-dialog-2392a5-9
worker/cb134-148-overlay-visible-c9b986-10
worker/cb234-session-id-keyed-04e1fc-1
worker/cb201-unit3-nudge-abdf5c-6
worker/cb201-unit2-policy-c1102c-5
worker/cb201-unit4-outcome-a13bfa-7
worker/cb201-unit1-classifier-91b9b1-4
worker/cb201-227-refine-831980-3
worker/cb175-model-readback-0f085f-1
worker/cb222-charter-tmpdir-17f013-1
worker/cb226-architect-slot-race-cd3aa8-3
worker/cb224-worktree-root-group-024523-2
worker/cb-123-role-demotion-c600f7-2
worker/cb-219-opencode-roots-1f677e-1
worker/cb214-claude-session-id-b9eab4-4
worker/cb213-zdotdir-wrong-process-dd6de4-3
worker/cb211-exhaustion-classification-9546e0-2
worker/cb137-ambiguous-task-4df3d8-4
worker/cb209-agentsessionid-4dfdb6-2
worker/cb185-hostenvnames-2692b5-3
worker/cb206-opencode-sqlite-128718-2
worker/cb185-worktree-group-fc0c99-1
worker/cb-137-ask-ticket-e7760c-2
worker/cb-172-broker-uri-d36ae4-4
worker/cb-175-model-readback-76ead6-3
worker/cb-161-pane-ancestry-293510-1
worker/cb-164-rebase-885863-8
worker/cb-164-empty-scrape-false-success-1a80af-3
fix/cb-197-ticket-ttl-from-completion
worker/cb-189-remote-url-coverage-4692f3-1
worker/cb-185-blockers-027756-4
worker/cb-192-gap-log-11b631-2
worker/cb-633-fix-5f4396-3
worker/cb185-router-d6436d-3
worker/cb185-router-routing-gaps-9e9d33-3
worker/cb185-paneids-992586-2
worker/cb-633-allow-list-union-ed374b-1
worker/cb-157-credential-in-remote-url-496e44-2
worker/cb-641-health-herdr-evidence-8f1f54-6
worker/cb-640-health-msg-evidence-99c9cd-1
worker/cb-642-fleets-status-skill-bbbc40-5
cb-634-ide-mcp
worker/lead-comms-wiring-c014b9-7
worker/lead-mailbox-c19577-6
worker/autocompact-window-82bc2f-5
worker/cb-634-probe-18056f-4
worker/cb635-broker-urienv
worker/cb-632-config-retry-8e0efa-7
lead/cb-622e-claude-md
lead/cb-622-followup
worker/cb-622a-165dff-1
lead/cb-622d-opencode-mount
worker/cb-622b-717c67-2
worker/cb-622c-ab7759-3
worker/cb-617b2-20ca4b-3
worker/cb-617a-5c2f4a-1
worker/cb596-4e49ef-3
worker/cb586-10500c-1
worker/cb-606-b9343a-25
worker/cb604-1445f8-24
worker/cb582-477374-21
worker/cb584-8c2281-22
worker/cb600-e6b9a9-20
worker/cb602-ce257f-19
worker/cb601-b42837-18
worker/cb598-6c7ba7-17
worker/cb599-740fe4-16
worker/cb597-282224-15
worker/cb590fix-185e9a-10
worker/cb528-recovery-race
worker/cb594-96bead-8
worker/cb590-916766-2
worker/cb527-997d99-3
worker/cb592-env-leak-3cbf9c-1
worker/cb588-async-ticket-nudge-3218f7-5
worker/cb578b-9dcb13-6
worker/cb581-d24826-5
worker/m2-u5-ef8c42-15
worker/cb578a-516499-2
worker/cb576-01a04b-17
worker/cb579-lead-tab-acba06-20
worker/cb580-terminal-health-ed6058-21
worker/cb577-f36fdc-18
worker/cb573b-3db06f-16
worker/cb568c-f36fdc-18
worker/cb568-drop-cause-c3ac1c
worker/cb575-cancelled-notification-c3ac1c
worker/m4-sol-a2cbec-3
worker/cb574-async-ask-c3ac1c
worker/cb573-health-model-8ca857-14
worker/cb572-unknown-target-7f2e35-13
worker/u4-700706-9
worker/u3-b9fcb6-6
worker/u2-ef5b68-4
worker/u1-469dce-1-clean
worker/u1-469dce-1
worker/cb-564-health-events-70cf7e-2
worker/cb-565-recycle-drops-role-98e58f-3
worker/cb-563-missing-reply-df2866-1
worker/cb-562-readiness-gate-silent-6c23c9-3
worker/cb-560-architect-presence-da8155-1
worker/cb-561-architect-silent-off-a71cab-2
worker/cb-548-bind-architect-slot-fe1b8c-1
worker/parity-overlay-settings-5fb711-1
secrets-central-store
cb-559-hot-key-correction
cb-557-fleet-role-pools
worker/cb-553-maxload-explicit-spawn-305ee3-6
worker/cb-551-idle-lead-heartbeat-f1633c-1
worker/cb-544-drain-preserves-worktree-925fad-3
worker/cb-552-docs-sync-1cb9cf-4
worker/cb-548-rendezvous-guard-rebased
worker/cb-548-rendezvous-guard-116b53-10
worker/cb-548-authz-v2-586df6-8
worker/cb-548-authz-264363-5
salvage/cb-528b-codex-home
salvage/cb-528a-codex-launcher
CB-518-primary-flow
feature/peer-launcher-spi
cb-103-injector
v1.1.0
v1.0.0
Labels
Clear labels
blocked
needs-live-proof
ready-to-delegate
silent-default
Cannot start until something else lands. The body says what.
Merged and green, but never shown working on the running daemon. Not the same as done.
Scope, files and acceptance criteria are written. A worker can be briefed from the body alone.
A feature that compiles, passes tests, and ships turned off. Nine recurrences and counting.
No Label
Milestone
No items
No Milestone
Projects
Clear projects
No project
Notifications
Due Date
No due date set.
Dependencies
No dependencies set.
Reference: fleet/fleetd#680
Reference in New Issue
Block a user
Blocking a user prevents them from interacting with repositories, such as opening or commenting on pull requests or issues. Learn more about blocking a user.
Delete Branch "%!s()"
Deleting a branch is permanent. Although the deleted branch may continue to exist for a short time before it actually gets removed, it CANNOT be undone in most cases. Continue?
Two gaps I found while verifying PR #679 before merging it (merged as
5051a06). The implementation is correct — these are guards it did not bring with it. Both live inscripts/redeploy-fleetd.shand its suite.Part 1 — reverting the core change leaves the whole suite green
I mutated the merged tree and ran
scripts/test-redeploy-fleetd.sh:report_jar_state's mismatch condition (!=→=)JAR="$MODULE/run/fleetd.jar"→JAR="$MODULE/target/fleetd.jar"The second mutation reverts the entire point of #664, and all 129 test functions still pass.
Why it survives, measured — this is "no assertion", not "cannot see". Mutation 1 killing proves the suite observes sourced globals, so the mechanism works. The cause is that every test supplies its own paths before calling anything:
No test ever observes the script's real
JAR, and nothing anywhere assertsJARis outsidetarget/. A test that supplies its own dependency says nothing about the producer.The property was checked — PR #679's acceptance criterion 1 ran
source scripts/redeploy-fleetd.sh; echo $BUILD_JAR; echo $JARand reportedDIFFER: yes. That is a correct check made once by hand. It is not a guard, and the hand-check is exactly what the mutation shows is missing.Part 2 — the installed launchd plist still names the old path, and nothing notices
Measured on this host 2026-10-03 21:10 CEST:
The daemon is launchd-supervised right now (
launchctl list→42543 0 dev.ltms.fleetd, ppid 1). The repo'sdeploy/dev.ltms.fleetd.plistis a template; the installed copy in~/Library/LaunchAgents/is a separate file, last written Aug 25, and editing the template does not touch it.This matters because the swap is
mv "$BUILD_JAR" "$JAR"— a rename, so after a redeploytarget/fleetd.jarno longer exists. Any launchd-initiated start from the stale installed plist (a reboot, orKeepAliveafter a crash) then runsjava -jar …/target/fleetd.jaragainst a missing file.The script already has the precedent and the seam.
check_log_path_matches_plist()exists for precisely this class of drift, and its own comment says "Nothing forced the two to agree." It readsStandardOutPathonly. It never readsProgramArguments:One comment, no code. This is a one-way gate: a guard written after the log-path incident closes only the direction that incident came from.
I am handling the plist reinstall by hand for today's redeploy, so this ticket is about the guard, not about unblocking me.
Acceptance criteria
Properties under a change, not names of constructs.
Part 1
JARto any path under$MODULE/target/makesscripts/test-redeploy-fleetd.shexit non-zero. Changing it to a different path outsidetarget/still passes — both directions, so the test is not satisfied by refusing every edit.JAR/BUILD_JARas sourced, without the test assigning them first. A test that sets them and then checks what it set is the defect being fixed, not the fix.JAR != BUILD_JARis asserted.Part 2
4. When the launchd agent is installed and its
ProgramArgumentsjar path does not resolve to$JAR, the script refuses with a message naming both paths and saying the installed plist needs reinstalling. Make it fail the same waycheck_log_path_matches_plistdoes.5. When the two agree, the script proceeds and says so, in the style of the existing
ok "log path check: …"line.6. A host with no installed plist is unaffected — no new failure on the unsupervised path. The existing
launchd_installed()seam already answers this.7. A positive control: a test proves the new check can actually fail, by pointing it at a plist whose jar path differs.
Notes for whoever takes this
report_jar_stateprintsabsentfor the built jar right after a successful redeploy, because the swap movestarget/fleetd.jarrather than copying it. That is correct behaviour and not part of this ticket, but do not "fix" it into a copy — the rename is what makes the swap atomic on one filesystem.Same shape, found and not fixed
PR #679's author reported
scripts/rename-checkout.sh:74(PATTERN="target/$JAR_NAME") and:428(JAR="$NEW_PATH/$MODULE_SUB/target/$JAR_NAME") as the same build-output-path-doubling-as-locator shape. It is a one-off historical migration tool. Worth a look while in here, but not in scope unless it is trivially the same fix.PART 3 — added after the brief. This comment is newer than your brief, so it wins where they differ. Same scope (
scripts/redeploy-fleetd.sh+ its suite), and it is the most urgent of the three.The merged script cannot see a daemon started under the old layout
#664 changed
PATTERNfromtarget/fleetd.jartorun/fleetd.jar.running_pid()is built on that pattern, and so isassert_single_daemon. Measured on this host just now, against the real live daemon plus a simulated process, with a positive control:So on the next redeploy,
OLD_PIDcomes back empty, the script printsno daemon running — this will be a cold start, skips the drain gate and the stop-and-wait entirely, and starts a second daemon while the first is still up. Two daemons against one herdr session take each other's members down.This is a transition hazard, but the transition is the next thing that happens, and it is also permanent for a hand-run: anyone who runs
java -jar fleetd/target/fleetd.jarby hand is now invisible toassert_single_daemon, which is the guard whose entire job is noticing a second daemon.What to change
Make the daemon-detection pattern match a fleetd daemon whatever directory its jar sits in, while
JARstays the single runtime path the script deploys to and starts from. Those are two different jobs that #664 accidentally merged:JAR/BUILD_JAR— where this script puts things and starts from. Keep exactly as merged.Narrowing a safety guard's pattern to the happy path is how the guard stops guarding.
Mind the existing
running_pid()comm = javaallowlist — it already filters out shells that merely mention the pattern, so broadening the pattern does not reintroduce thesh -cfalse positive that #593 fixed. Keep that allowlist.Acceptance criteria for part 3
target/,running_pid()finds it. With one naming a jar underrun/,running_pid()finds it. Both asserted — a test that only checks the second direction is the bug.#593false-positive case still does not match: a non-exec'ingsh -cthat merely contains the pattern text is still excluded. Keep a positive control so this test can actually fail.assert_single_daemonnotices two daemons when one runs fromtarget/and one fromrun/. This is the case that bites today.Note on the other two parts
Part 1's mutation (
JAR→ a path undertarget/) must still make the suite red. Part 3 broadens only the detection pattern, notJAR, so the two do not conflict. If you find they do, say so in your reply rather than quietly relaxing part 1.I am holding today's redeploy until this lands, so take the time to get the test directions right. Do not run the redeploy script, and do not touch the live daemon.
Scope addendum to part 3 — one more file. My brief said "
scripts/redeploy-fleetd.shandscripts/test-redeploy-fleetd.shonly". That was too narrow. Part 3 also covers one line in.claude/skills/fleets-status/SKILL.md.There are exactly two daemon-locator sites in the repo, and #664 narrowed both:
(The third hit is a comment inside the script; update it only if your change makes it untrue.)
Against the live daemon today, the
fleets-statusline printsfleetd: not running— the daemon is up as pid 42543, it just runs fromtarget/. A status tool that reports a running daemon as down is worse than one that says nothing, because the reader acts on it.Fix that line the same way you fix
PATTERN, so the two agree. The skill is a Markdown playbook, not shell the suite can source, so it gets no test — just make the two patterns identical, and say in your reply that you checked they match.This is the shape worth naming:
fleets-statusandredeploy-fleetd.sheach keep their own copy of "how to find the daemon". One fact, two places, and #664 updated both to the same wrong value in one pass. Do not add a third copy to fix it. If you see a clean way to make the skill reference the script's value rather than restate it, say so in your reply — but do not build it in this ticket.Correction to part 3. I wrote that an empty
OLD_PIDmakes the script "start a second daemon beside pid 42543". That is wrong. I had not read the stop dispatch when I wrote it. The blindness is real and part 3 still stands; the consequence is different, and in one respect worse.What actually happens with an empty
OLD_PIDThe stop section branches on
$OLD_PIDfirst, then on the supervisor:Supervisor detection does not use
PATTERN—launchd_loaded()askslaunchctl list. SoSUPERVISOR_KINDis stilllaunchd, the second branch runs, andlaunchctl unload -wdoes stop the running daemon. There is no second daemon. I was wrong.The real consequences, each read off the code
drain_gate_required()is[ -n "$old_pid" ] && [ "$assume_yes" = 0 ]. WithOLD_PIDempty it returns false andrun_drain_gatereturns immediately. No prompt, no "a restart drops every in-flight ticket" warning. Live members lose their reports with nothing asked and nothing printed.ok "launchd agent unloaded (was already not running)"— while it was running. Anyone reading the transcript afterwards is told the opposite of what happened.wait_for_daemon_exitis never called on that branch, so the script proceeds without confirming the old process is gone.HAD_OLD_PID="$(compute_had_old_pid "$OLD_PID")"is 0, soreport_shutdown_drainis told no previous daemon was stopped. That report is exactly the signal #664's live probe needs, so the blindness also hides the evidence that would settle #664.One thing I over-stated the risk of: the swap is
mvon one filesystem, which is a rename. An already-open file descriptor follows the inode, so moving the jar does not disturb a JVM that still has it open. Overwriting would. So point 3 is a missing confirmation, not jar corruption.What changes for you
Nothing in parts 1, 2, or 3's actual work, and nothing about criteria 8 and 9 — a detection pattern that only finds daemons in the directory you expect is still the defect.
Criterion 10 restated. I wrote "
assert_single_daemonnotices two daemons when one runs fromtarget/and one fromrun/— this is the case that bites today." The property is still worth pinning, but the "bites today" framing was wrong. Treat it as: a daemon running from a jar outside$JAR's directory must still be visible torunning_pid()and therefore countable byassert_single_daemon. Drop it if it fights the other criteria, and say so.One criterion added, because it is the consequence that actually costs something:
$JAR's directory, the script must not reach the stop step with an emptyOLD_PID. Assert thatrunning_pid()finds it — that is enough, because the drain gate,wait_for_daemon_exitandHAD_OLD_PIDall key offOLD_PID, and fixing detection fixes all four at once. Do not patch the four call sites separately.Sorry for the churn. The blindness is measured and real; my account of the damage was not.
PR #681 merged as
a6aeda3. All three parts in. Closing.I re-ran every mutation myself rather than trusting the report
On the merged tree, baseline
exit=0, 134 test functions (5 new):JAR→ under$MODULE/target/FAIL: $JAR must not live under $MODULE/target/JAR→ elsewhere outsidetarget/PATTERN→ back torun/fleetd.jarrunning_pid() did not find a real second process … naming a jar under target/dieneutralisedcheck_jar_path_matches_plist must die when the plist names a different jarcomm = javaallowlist removedThe second row is the one that matters most. A test that fires on any edit to
JARwould have passed row 1 while pinning nothing; it had to stay green here. It did.The last row answers the question widening
PATTERNraises: broadening it did not weaken the#593protection against a shell that merely holds the pattern as literal text.mvn -o clean installfromfleetd/:Tests run: 1929, Failures: 0,BUILD SUCCESS, 172 report files, 0 failure lines.Blocker 1 is closed, verified against the real process
The script now finds the live daemon it was blind to an hour ago. Both locator sites agree (
redeploy-fleetd.sh:89andfleets-status/SKILL.md:62, bothfleetd.jar).I checked the risk that widening
PATTERNcreates, because the implementer could notMy worry was
assert_single_daemonfalse-positiving while a worker builds: Maven producestarget/fleetd.jar, and surefire forks havecomm=java. A count taken on an idle host would have proved nothing, so I ran a fullmvn -o clean installin a throwaway worktree and sampled every 5 seconds through it:Only the live daemon. Maven and surefire do not put
fleetd.jarin their argv — they use classpath directories, not the shaded jar. So the widened pattern does not block a redeploy while a worker builds.Credit where it is due
The implementer picked up all three ticket comments — including my own correction walking back the "two daemons" framing — and implemented the restated criterion 11 rather than the superseded criterion 10. It also flagged in its report that part 3 arrived after its brief. That is the ticket-beats-brief rule working as intended.
What this does to the next redeploy
check_jar_path_matches_plistis live, and the installed plist still names…/fleetd/target/fleetd.jar. So my next redeploy will now refuse with the mismatch message instead of starting a daemon against a path the swap is about to empty. That refusal is the guard doing its job; reinstalling the plist is my next step, and it is tracked on #664.