Building in the main clone while the daemon runs breaks its shutdown drain — NoClassDefFoundError on a drain-only class #664
Open
opened 2026-10-03 19:19:30 +02:00 by ltms
·
7 comments
No Branch/Tag Specified
main
worker/664-c12e95-3
worker/668-08534d-4
worker/672-0f2469-2
worker/670-7d1022-1
worker/661-ac7c28-2
worker/664-37fb9b-3
worker/663-remove-3arg-read-3f6783-1
worker/659-remove-dead-backcompat-ba5e6f-1
worker/637-revision-60a488-23
worker/656-redact-regression-tests-892903-19
worker/637-context-gauge-threshold-466eb5-16
worker/639-redact-line-numbers-de4ac4-17
worker/641-set-reformat-guard-6f96a4-18
worker/642-herdr-guard-scope-5de0e4-15
worker/650-javadoc-scope-95f3b3-14
worker/612-01e9f7-13
worker/612-a-r4-quarantine-outage-7ab0e8-5
worker/612-a-r9-r11-capacity-coverage-peers-cfcc79-7
worker/612-a-r10-loophealth-ccc872-8
worker/612-a-r12-turnregistrar-9e3bb7-9
worker/612-a-r5-leadconfigdir-9e70cf-6
lead/config-edit-redact-anchor-wording
worker/config-edit-seam-ca8dc1-1
worker/612-r67-630-lifecycle-290b8d-3
worker/629-625-ports-seams-da7d5d-4
worker/612-r12-exhaustion-f37cd7-1
worker/612-r38-amqp-24b083-2
worker/fleetd-612-unita-87807e-1
worker/612-b3-mcpwirings-da2b58-3
worker/612-b2-cb185-176d3a-2
worker/612-b1-completion-457459-1
worker/612-agaps-73a926-2
worker/608-sleeps-3a64ff-3
worker/621-b4520b-1
worker/618-b83894-2
worker/fleetd-615-e05481-5
worker/lead-autocompact-5f1ab2-3
worker/fleetd-613-f85deb-3
worker/fleetd-608-flaky-nudge-test-d0c2d1-3
worker/lead-context-gauge-ad404f-1
worker/gauge-wiring-9158c1-4
worker/redeploy-slowstart-ead0e5-5
worker/charter-bytes-13668c-6
worker/rollover-outcome-291483-2
worker/589-f64303-2
worker/593-1a8025-5
worker/589-fcd2aa-1
worker/568-9fdaa2-3
worker/571-attempted-outcome-5739f7-2
worker/581-completionresolver-cas-sites-0542b7-6
worker/562-loop-health-wiring-test-99611c-5
worker/562-surface-loop-health-7df5cc-4
worker/575-waiter-cleanup-sites-62ad80-1
worker/572-answer-lock-release-46a9ae-5
worker/567-probe-channel-leak-a38fc5-6
worker/551-record-before-send-7cbf56-1
worker/561-listener-fanout-survives-a-throw-61d538-2
worker/555-redeploy-main-flow-seam-65c2f5-2
worker/556-injector-owns-registration-e027a5-1
worker/552-post-restart-mktemp-abort-bc2672-4
worker/553-onstatus-completion-leak-0da881-2
worker/550-shasum-linux-196132-1
worker/538-loop-dies-on-error-4a5eeb-6
worker/426-health-coverage-ef1fd4-4
worker/504-failed-reported-clean-3cfd66-3
worker/537-capturedlog-close-e4c437-2
worker/459-broken-link-targets-cadc17-5
worker/535-appender-leak-fe74c1-1
worker/512-part2-shutdown-detection-434701-9
worker/529-logger-level-sweep-2a5533-8
worker/528-drain-gate-call-site-5de83d-7
charter/forge-mcp-vs-token
worker/521-swap-guard-unpinned-28e931-5
worker/519-probe-test-harness-d25ab8-4
worker/525-logger-level-leak-1b4eb0-6
worker/518-fleetmcp-resolver-wiring-8ef96c-1
worker/512-drain-complete-line-7edd71-3
worker/517-abort-branch-and-jar-id-41b641-2
worker/500-9e52c9-3
worker/509-4912f4-2
worker/511-9a4b23-1
worker/493-479f45-2
worker/505-03f8b2-1
worker/492-followup-detect-unclear
worker/501-a31fa0-7
worker/498-451d1c-5
worker/494-1015ce-2
worker/492-209647-1
worker/489-001902-2
worker/480-relative-handover-path-906323-1
worker/480-b-handover-skill-45bf1f-5
worker/474-followup-source-pin-f54a55-17
worker/474-charter-check-on-reload-f54a55-17
worker/466-quarantine-repeatcount-report
worker/393-opencode-skill-seeding-71854b-13
worker/469-canonical-tool-names-2a472a-16
worker/466-quarantine-escalation-5ae9c1-15
worker/446-hot-exhausted-pattern-0af580-6
worker/464-charter-tool-name-guard-a85635-12
worker/463-listfleet-default-fails-open-f1c76c-11
worker/458-invariant-5-by-purpose-862f9a-10
worker/439-coordinator-row-gate-bc032a-8
worker/449-herdr-protocol-576015-4
worker/450-abstract-spawn-599e1c-5
worker/437-ack-refuses-177d91-1
worker/444-placement-window-feb56a-2
worker/440-helddurable-derived-d462d7-13
worker/425-rework-placement-resolve-c58ba1-9
worker/421-lead-peek-held-msgs-cdbad2-10
worker/435-fixed-policy-cap-fe11de-12
worker/422-gate-state-observability-9e79d6-11
worker/431-memberregistry-live-readers-cdbad2-10
worker/424-architect-slot-hot-038b41-7
worker/422-model-gate-spawn-c29f48-6
worker/425-default-profile-live-f55534-8
worker/415-coverage-wording-2cbf9c-5
worker/416-3ad1da-1
worker/418-588283-3
worker/deterministic-stamp-race-409-3cb7b6-10
worker/armed-reads-live-config-404-ed931f-9
worker/reply-peer-refusal-391-5a34bd-7
worker/models-allowlist-aa9e9b-3
worker/ttl-stamp-race-399-f1122f-8
worker/scrub-receipt-400-316b3e-5
worker/exhaustion-detection-395-105105-6
worker/scrub-abort-394-316b3e-5
fix/scrub-uid-abort
worker/task-scrub-517574-2
worker/t386-clock-bd5b78-4
worker/t384-scrub-813790-5
worker/t381-cc-748314-2
worker/t373-336973-2
worker/t365-3920c5-3
worker/t358-6e989b-1
worker/t355-8b321c-1
worker/fleetd-369-hermetic-git-tests-e8b19a-3
worker/fleetd-368-stale-lead-binding-f5682e-2
worker/fleetd-360-deploy-units-0d3793-1
worker/359-dead-lead-tabs-f1253b-4
worker/362-worktree-skills-c03e51-3
worker/361-coord-visibility-655144-1
362-plugin-visibility-and-drift
worker/errscan-bed2ca-2
worker/amqp-log-identity-bed2ca-2
worker/withdefaults-guard-561704
worker/sleepguard-82076d-1
worker/fd334-9ee1b6-5
worker/fd348-f1ab27-4
worker/fd335-a71c35-1
worker/fd342-174a17-2
worker/fd345-490d0f-3
worker/fleetd-337-5ec7d4-21
worker/fleetd-341-af5a6b-24
worker/fleetd-339-5ca0a2-23
worker/fleetd-338-83a4a1-22
worker/fleetd-333-281f46-18
worker/fleetd-329-11bdbb-16
worker/fleetd-330-2770fb-17
worker/fix-326-50506e-15
worker/fix-324-3e9bbf-14
worker/fix-323-b8287d-13
worker/fix-316b-bd0860-11
worker/fix-318-76ca36-9
worker/fix-317-486aec-8
worker/fix-315-ce47c5-6
worker/fix-307-275890-6
worker/fix-308-b4f664-7
worker/fix-309-ec3939-8
worker/fix-310-7a3974-9
worker/fix-302-52ad0e-9
worker/fix-298-ce1acb-8
worker/fix-297-66bd11-7
worker/fix-296-104622-6
worker/fix-293-bare-closetab-eb22b5-3
worker/fix-280-gone-ask-lapse-bca98e-2
worker/fix-290-reapidle-guard-coverage-9b0dd1-1
worker/fix-285-trust-seed-8f3565-10
worker/fix-284-backend-error-seat-85912c-11
worker/fix-282-chained-ask-e6d0bb-8
worker/fix-283-teardown-leaks-f40dfa-9
worker/fix-281-pin-handler-actions-4921ac-7
worker/audit-rendezvous-lifecycle-d072ae-2
worker/audit-health-placement-1a2476-6
worker/audit-teardown-exits-e207a5-3
worker/audit-launcher-asymmetry-27e370-4
worker/audit-rest-authz-6ca53c-5
worker/investigate-275-abandon-asking-fdef52-8
worker/fix-274-worktree-leak-b0095d-7
worker/fix-273-exhausted-pattern-9665b5-6
worker/fleetd-267-model-check-bd8068-1
worker/fleetd-131-archunit-18b834-7
worker/fleetd-266-sshagent-rename-a014ff-6
worker/fleetd-184-uid-claim-8e1f31-4
worker/fleetd-184-warn-b381ee-10
worker/fleetd-184-docs-be1d12-9
worker/fleetd-257-9bf010-7
worker/fleetd-103-23a113-6
worker/fleetd-247-342356-5
worker/fleetd-116-04dea8-4
worker/fleetd-252-a830e0-3
worker/fleetd-111-7e8673-9
worker/fleetd-155c-f8ef4b-8
worker/fleetd-176-b928ca-3
worker/fleetd-249-7a7878-2
worker/cb248-composition-root-b-9acdf7-15
worker/cb148-envrc-default-fa6c82-12
worker/cb201-unit5-wiring-6c12e6-8
worker/cb241-fallback-echo-1175e9-11
worker/cb149-trust-dialog-2392a5-9
worker/cb134-148-overlay-visible-c9b986-10
worker/cb234-session-id-keyed-04e1fc-1
worker/cb201-unit3-nudge-abdf5c-6
worker/cb201-unit2-policy-c1102c-5
worker/cb201-unit4-outcome-a13bfa-7
worker/cb201-unit1-classifier-91b9b1-4
worker/cb201-227-refine-831980-3
worker/cb175-model-readback-0f085f-1
worker/cb222-charter-tmpdir-17f013-1
worker/cb226-architect-slot-race-cd3aa8-3
worker/cb224-worktree-root-group-024523-2
worker/cb-123-role-demotion-c600f7-2
worker/cb-219-opencode-roots-1f677e-1
worker/cb214-claude-session-id-b9eab4-4
worker/cb213-zdotdir-wrong-process-dd6de4-3
worker/cb211-exhaustion-classification-9546e0-2
worker/cb137-ambiguous-task-4df3d8-4
worker/cb209-agentsessionid-4dfdb6-2
worker/cb185-hostenvnames-2692b5-3
worker/cb206-opencode-sqlite-128718-2
worker/cb185-worktree-group-fc0c99-1
worker/cb-137-ask-ticket-e7760c-2
worker/cb-172-broker-uri-d36ae4-4
worker/cb-175-model-readback-76ead6-3
worker/cb-161-pane-ancestry-293510-1
worker/cb-164-rebase-885863-8
worker/cb-164-empty-scrape-false-success-1a80af-3
fix/cb-197-ticket-ttl-from-completion
worker/cb-189-remote-url-coverage-4692f3-1
worker/cb-185-blockers-027756-4
worker/cb-192-gap-log-11b631-2
worker/cb-633-fix-5f4396-3
worker/cb185-router-d6436d-3
worker/cb185-router-routing-gaps-9e9d33-3
worker/cb185-paneids-992586-2
worker/cb-633-allow-list-union-ed374b-1
worker/cb-157-credential-in-remote-url-496e44-2
worker/cb-641-health-herdr-evidence-8f1f54-6
worker/cb-640-health-msg-evidence-99c9cd-1
worker/cb-642-fleets-status-skill-bbbc40-5
cb-634-ide-mcp
worker/lead-comms-wiring-c014b9-7
worker/lead-mailbox-c19577-6
worker/autocompact-window-82bc2f-5
worker/cb-634-probe-18056f-4
worker/cb635-broker-urienv
worker/cb-632-config-retry-8e0efa-7
lead/cb-622e-claude-md
lead/cb-622-followup
worker/cb-622a-165dff-1
lead/cb-622d-opencode-mount
worker/cb-622b-717c67-2
worker/cb-622c-ab7759-3
worker/cb-617b2-20ca4b-3
worker/cb-617a-5c2f4a-1
worker/cb596-4e49ef-3
worker/cb586-10500c-1
worker/cb-606-b9343a-25
worker/cb604-1445f8-24
worker/cb582-477374-21
worker/cb584-8c2281-22
worker/cb600-e6b9a9-20
worker/cb602-ce257f-19
worker/cb601-b42837-18
worker/cb598-6c7ba7-17
worker/cb599-740fe4-16
worker/cb597-282224-15
worker/cb590fix-185e9a-10
worker/cb528-recovery-race
worker/cb594-96bead-8
worker/cb590-916766-2
worker/cb527-997d99-3
worker/cb592-env-leak-3cbf9c-1
worker/cb588-async-ticket-nudge-3218f7-5
worker/cb578b-9dcb13-6
worker/cb581-d24826-5
worker/m2-u5-ef8c42-15
worker/cb578a-516499-2
worker/cb576-01a04b-17
worker/cb579-lead-tab-acba06-20
worker/cb580-terminal-health-ed6058-21
worker/cb577-f36fdc-18
worker/cb573b-3db06f-16
worker/cb568c-f36fdc-18
worker/cb568-drop-cause-c3ac1c
worker/cb575-cancelled-notification-c3ac1c
worker/m4-sol-a2cbec-3
worker/cb574-async-ask-c3ac1c
worker/cb573-health-model-8ca857-14
worker/cb572-unknown-target-7f2e35-13
worker/u4-700706-9
worker/u3-b9fcb6-6
worker/u2-ef5b68-4
worker/u1-469dce-1-clean
worker/u1-469dce-1
worker/cb-564-health-events-70cf7e-2
worker/cb-565-recycle-drops-role-98e58f-3
worker/cb-563-missing-reply-df2866-1
worker/cb-562-readiness-gate-silent-6c23c9-3
worker/cb-560-architect-presence-da8155-1
worker/cb-561-architect-silent-off-a71cab-2
worker/cb-548-bind-architect-slot-fe1b8c-1
worker/parity-overlay-settings-5fb711-1
secrets-central-store
cb-559-hot-key-correction
cb-557-fleet-role-pools
worker/cb-553-maxload-explicit-spawn-305ee3-6
worker/cb-551-idle-lead-heartbeat-f1633c-1
worker/cb-544-drain-preserves-worktree-925fad-3
worker/cb-552-docs-sync-1cb9cf-4
worker/cb-548-rendezvous-guard-rebased
worker/cb-548-rendezvous-guard-116b53-10
worker/cb-548-authz-v2-586df6-8
worker/cb-548-authz-264363-5
salvage/cb-528b-codex-home
salvage/cb-528a-codex-launcher
CB-518-primary-flow
feature/peer-launcher-spi
cb-103-injector
v1.1.0
v1.0.0
Labels
Clear labels
blocked
needs-live-proof
ready-to-delegate
silent-default
Cannot start until something else lands. The body says what.
Merged and green, but never shown working on the running daemon. Not the same as done.
Scope, files and acceptance criteria are written. A worker can be briefed from the body alone.
A feature that compiles, passes tests, and ships turned off. Nine recurrences and counting.
No Label
Milestone
No items
No Milestone
Projects
Clear projects
No project
Notifications
Due Date
No due date set.
Dependencies
No dependencies set.
Reference: fleet/fleetd#664
Reference in New Issue
Block a user
Blocking a user prevents them from interacting with repositories, such as opening or commenting on pull requests or issues. Learn more about blocking a user.
Delete Branch "%!s()"
Deleting a branch is permanent. Although the deleted branch may continue to exist for a short time before it actually gets removed, it CANNOT be undone in most cases. Continue?
Found during the redeploy after PR #662 merged, on 2026-10-03. The redeploy itself succeeded; this
is about the previous daemon's shutdown.
What happened
scripts/redeploy-fleetd.sh --yesreported this about the daemon it replaced:The class is not missing — the jar was swapped under a running JVM
SessionManager$DrainTally.classis present in the jar:(I first ran this with
grep -i "SessionManager\$DrainTally"inside double quotes, where\$collapses to a regex end-anchor, and got a clean
0that looked like "the class is gone". Pairevery zero with a positive control — the control here was
unzip -l … | grep -cF 'SessionManager'→ 5.)Cause, and it was my own command
The daemon loads classes lazily from the jar it opened at boot. The drain path's classes are, by
definition, only loaded at shutdown. So if the jar file is replaced between boot and shutdown,
the shutdown hook tries to load a class out of a file whose content is no longer the one it opened.
The chain, measured:
5e68eb6e3d38(recorded in the previous lead's hand-off).mvn -o installin the main clone at 19:12 to verify the merged tree.--checkthenreported the daemon's own jar path as
a37441b9e9bc (2026-10-03 19:12:30)— a different fileat the path pid 42930 was running from.
f3b541ba677aand swapped it in after stopping the daemon.DrainTally.So the jar at the running daemon's path was overwritten while it ran, by my build, before the
script was ever involved. The script's own stage-then-swap (#493) protected its build; it could not
undo a replacement that had already happened.
Impact
Low this time, and that is luck rather than design. I had already stopped all five members by hand
and
fleet_listreportedmembers: [], so the drain had nothing to release. With live members atredeploy time, their sessions would not be released cleanly.
The general shape is worse than this one incident: any
mvn installin the main clone while thedaemon runs leaves that daemon unable to execute its shutdown drain. Nothing warns at the time. The
damage only shows up at the next restart, attributed to the restart rather than to the build.
The project instruction is part of the problem
CLAUDE.mdcurrently says, about verifying a merge in the main clone:That guidance stops the jar being deleted but not being replaced, and replacement is enough
to break lazy class loading. A lead following the instruction exactly still breaks the drain. The
same text appears in the lead hand-off template.
Suggested fix
Pick one, or argue for another:
CLAUDE.mdand the hand-off template to say: verify a merge bybuilding in a throwaway worktree, and let only
scripts/redeploy-fleetd.shtouch the mainclone's jar. Cheapest, and it matches what the worktree advice already half says — but it
relies on every future lead reading it.
jar to a private path at startup and run from the copy, so a rebuild cannot reach it. Removes
the hazard rather than documenting it.
maven output path — build to a staged name always, and have the script be the only thing that
publishes to the runtime path.
Option 2 is the only one that survives someone not reading the instruction. Option 1 should happen
regardless, because it is true and cheap.
Not verified by me
a marker taken before the restart, so I only have this one shutdown window, not a history.
launchctl unloadversus a plain signal changes the shutdown hook's behaviour here.DrainTallyis the one that threw; theremay be more behind it.
Lead review of PR #666 (the instruction half) — two small changes before merge
First, two corrections to this issue's own text
I measured both in the main clone. The issue body is wrong on them, so nobody should re-derive from it:
CLAUDE.md. It lived in exactly one place:.claude/skills/redeploy-fleetd/SKILL.md:29. Command:grep -rn "mvn clean" . --exclude-dir=.git --exclude-dir=wiki --exclude-dir=target.CLAUDE.md's ownmvn cleanhits are about IDE validation and CVE checks, a different topic.grep -nE "mvn|jar|build" .claude/skills/handover/SKILL.mdfinds no build guidance at all. The quoted sentence was probably read out of a previous lead's own.handover/HANDOVER.md, which is a transient file, not a template.Verified good on PR #666
git diff --name-only origin/main..refs/pull/666/head | grep -E '\.(java|xml|sh|yaml|yml|json)$'returns 0.CLAUDE.mdis identical onmainand on the PR head — 20938 bytes both sides.CLAUDE.mdagainstwiki/7-Use-Cases.md(submodule ata61f729):in sync: True. The worker correctly reported it could not run this and did not claim it as passed.### Redeploying the daemon. Right section.Change 1 — a true sentence was dropped, and it is the load-bearing one
The old text ended with "and nothing degrades until the next restart". The new text has no equivalent, in either file. I checked with a positive control:
That fact is why the rule gets obeyed. Without it the hazard reads as theory. With it, a lead understands that the build looks completely fine at the time, and the breakage surfaces at the next restart where it is naturally blamed on the restart instead of on the build. That misattribution is exactly what happened here on 2026-10-03.
Add it back to both files, in plain words: nothing warns at the time, and the damage appears at the next restart, where it looks like the restart's fault.
Change 2 — restore the file's final newline
The diff also deletes the trailing empty line at the end of
.claude/skills/redeploy-fleetd/SKILL.md. The worker flagged this itself, which is the right call. It is unrelated to the fix, so put it back and keep the diff to the two places that needed changing.Nothing else. The rest of the PR is correct and I will merge it once these two land.
The instruction half is merged as
4b4a868and pushed. The code half (option 2) is still open.PR #666 merged locally and pushed; I closed the PR by hand, since a local merge never closes it. This issue stays open — option 1 was always "should happen regardless", not the fix.
What I verified myself on the merged tree
git diff --name-only origin/main...<head>lists exactly two Markdown files; the code-file filter (.java|.xml|.sh|.yaml|.yml|.json) returns 0.mainand the PR head.wiki/7-Use-Cases.md:in sync: True.SKILL.mdare0a0a, identical tomain.SKILL.md) and 464 (CLAUDE.md).fleetd/src/test/javareferenceCLAUDE.mdor.claude/skills(OpenCodeLauncherTest,FleetProfilesLiveDefaultTest,ClaudeCodeLauncherTest,McpContractDocTest,RestRouteInventoryTest,GitWorktreesTest), so a doc edit can break the build. Built the merged tree in a throwaway worktree:Tests run: 1923, Failures: 0, Errors: 0, Skipped: 0,BUILD SUCCESS, exit 0, 170 report files.The review round
The first version dropped a true sentence that the old text had carried — "nothing degrades until the next restart". I caught it with a positive control rather than by eye:
That fact is the load-bearing one. Without it the hazard reads as theory; with it a reader understands the build looks completely fine at the time and the breakage surfaces at the next restart, where it gets blamed on the restart. Both files now say so.
Still owed on this issue — option 2
The two architects were asked which immunity fix to take: 2a preload the drain path's classes at boot, or 2b run from a private copy of the jar. The
solarchitect has reported and chose 2b. I am waiting on the second position before I settle it, and I will write the decision here.One thing from
sol's report is worth recording now because it bears on this issue's own "Not verified by me" list, and I checked it myself:DrainTallyis loaded unconditionally —drainAllassigns it atSessionManager.java:1117and even an empty snapshot constructs it. ButReleaseCause,ReleaseDetailandShuttingDownExceptionload only on particular branches, andFleetdRuntime.close()continues through many more components after the drain (FleetdRuntime.java:130-164). So the set depends on live state: listeners wired, backend kind, broker kind, worktree state, and which error branches fire. That is an argument against 2a on its merits, not merely a gap in testing.grep -rn -E 'CodeSource|getProtectionDomain|getLocation\(' fleetd/src/main/javareturns 0, and my positive control on the same tree (getProperty("user.home")→ 6 hits) shows the search works rather than silently matching nothing. So 2b needs no change to Java self-location code; its cost is all in the deployment surface.Decision: 2b, in the "separate runtime path, atomic rename" variant
Both architects formed positions independently and both chose 2b. They agree, so this is settled at the fleet level and does not go to the operator.
solargued 2b on the grounds that a preload list cannot be completed.opusargued the same and went further, specifying the variant below. I checked the load-bearing claims myself before accepting.2a is rejected. Do not implement it, and do not add it alongside 2b "for belt and braces."
What to build
The daemon runs from
fleetd/run/fleetd.jar— outsidetarget/, so neithermvn installnormvn cleancan reach it. The deploy handoff is a singlemvfromtarget/fleetd.jartorun/fleetd.jar, performed only after the old daemon is confirmed gone, reusing the gateswap_staged_jaralready has.Why a
mvand not a copy at startup:mvwithin one filesystem isrename(2), which is atomic. There is never a half-written runtime jar. Acphas a truncation window, and that window is the failure we are fixing.deploy/fleetd.service:53, and thenohupatscripts/redeploy-fleetd.sh:1025. Three places to forget, and forgetting is silent.target/, somvn installbecomes harmless from any clone, by any lead, whether or not they read the instructions. That is the actual goal: the current rule is enforced by prose alone.What I verified myself in the code
I did not take the inventory on trust. Every line below is from a command I ran in the main clone at
b4b7cf5:fleetd/pom.xml:194—<finalName>fleetd</finalName>, with no<directory>or<outputDirectory>override. This is why Maven's output path is the runtime path. Root cause confirmed.scripts/redeploy-fleetd.sh:74JAR="$MODULE/target/fleetd.jar",:81JAR_STAGED=,:87PATTERN='target/fleetd.jar'.deploy/fleetd.service:53anddeploy/dev.ltms.fleetd.plist:50..claude/skills/fleets-status/SKILL.md:62runspgrep -f 'target/fleetd.jar'.getProtectionDomain|getCodeSource|java.class.path|ProcessHandleoverfleetd/src/main/javareturns onlyLsofPeerPidLookup.java:20(own pid) andParentResolver.java:17-18(parent walk). So the Java side needs no change, and 2b cannot be done by the daemon — the JVM opens the jar beforemainruns. It belongs in the launch path.target/fleetd.jaroutsidetarget/and.git.The two silent breakages to handle, not discover
PATTERNatscripts/redeploy-fleetd.sh:87.running_pid()ispgrep -f "$PATTERN". Move the jar and leave the pattern, andrunning_pidreturns nothing — the script then believes no daemon is running and starts a second one. The comment at:209-217already warns about this..claude/skills/fleets-status/SKILL.md:62. Same pattern, and it would report no fleet at all. A skill is not covered by any test, so nothing catches this.Both fail by reporting absence, which reads as good news. Pair every such check with a positive control.
--checkmust stop lyingscripts/redeploy-fleetd.sh:1173printsjar_idanddate -r "$JAR"under the label "jar on disk". With one jar path that label is one name for two different facts, and during this incident it would have printed the new jar's hash while the JVM ran the old one. Under 2b there are two facts, so print both: built hash/mtime and running hash/mtime. A visible mismatch is the whole value.Accept this limit explicitly
2b does not remove the stale-jar risk; it adds one more place to get the handoff wrong, and a daemon booted from a stale
run/fleetd.jaranswers/healthzand looks healthy. That risk already exists —--no-buildrestarts whatever jar is at the path today. The reason 2b still wins: it makes the instrument honest instead of breaking it, and 2a's failure mode is both silent and worse, landing at shutdown with live members.What settles it
One live probe, lead-only: boot from
fleetd/run/fleetd.jar, runmvn -o installin the main clone while it runs, then redeploy and checkfleetd.outfor thedrain complete: released=… abandoned=…line fromSessionManager.java:1126. That line appearing is the pass. Neither architect ran it — correctly, since it would cut their own channel.Claims I am passing on without checking
rename(2)leaves a running JVM's open inode intact while an in-place rewrite does not. This is POSIX semantics and the incident fits it, but neither architect tested it on this host, and nor have I.maven-shade-plugintruncates in place or renames. The incident proves the write reached the inode the JVM held, whichever did it.fleetd/fleetd.yaml(gitignored) andwiki/(a submodule) were not searched for jar paths by the architect, whose worktree lacks both. I have not searched them either, so the 11-file inventory may be short.The docs-only fix for this ticket already shipped in
4b4a868. This decision covers the code change, which is not yet implemented.Addendum: I closed the two gaps the architect could not reach from a worker worktree, and the inventory was indeed short.
fleetd/fleetd.yaml(gitignored, so absent from every member's worktree): no jar path. Nothing to change there.wiki/(a submodule, uninitialized in every member's worktree —git submodule statusprints+a61f729, 19 entries in my clone): two more sites.wiki/13-User-Guide.md:93—java -jar fleetd/target/fleetd.jar fleetd/fleetd.yamlwiki/11-Features.md:2202— describes the launchd job as runningjava -jar target/fleetd.jar fleetd.yamlSo the inventory is 13 files, not 11. Both new sites are documentation that tells a human how to start the daemon, so leaving them stale would send an operator to boot from
target/— the one path 2b exists to stop anyone running from. They must be updated in the same change, and the wiki edits are lead-only.This is the pattern worth naming: the architect listed every site it could see and said plainly which two surfaces it could not. That caveat is what made the gap findable. A report that had simply said "11 files" would have been read as complete.
Before-state, measured on the live daemon 2026-10-03 21:05 CEST. Recording it now because the redeploy that closes this ticket destroys the reading.
Three things this pins:
fleetd/target/fleetd.jar— the exact pathmvn installin the main clone writes. This is the premise the whole ticket rests on, and it is now measured rather than assumed.fleetd/run/does not exist yet, so after the change lands, its appearance is itself a check that the new path is being used.I used
ps -o lstart=for the start time, not the log —fleetd.outtimestamps carry no date and drift timezone, so two boots in one file can read hours apart.The live probe that settles this ticket is mine, not the implementer's. Booting from the new path and then running
mvn -o installin the main clone while it runs would cut a worker's own channel mid-turn. The pass condition is thedrain complete: released=… abandoned=…line fromSessionManager.java:1126appearing infleetd.outafter that sequence — today that drain is what breaks, and nothing warns at the time.PR #679 merged as
5051a06(pushed,origin/mainand local HEAD both5051a06). This ticket stays open — the code is in, the live probe that settles it is not done, and I am holding the redeploy. Reasons below.What I verified myself before merging
Not taking the implementer's word for any of it:
origin/main...pr/679mvn -o clean installfromfleetd/Tests run: 1929, Failures: 0,BUILD SUCCESS, 172 report files, 0 failure linestarget/fleetd.jar, does not createrun/run/ignored by the new rulefleetd/.gitignore:7:run/scripts/test-redeploy-fleetd.shThe implementer reported 1928 tests / 171 reports; I measured 1929 / 172. The gap is explained: it built on merge-base
7f9a9c0, before the test that landed in6f27522today.Its claim that the suite's 3
FAIL:lines and 5mktemp:lines are deliberate self-test output is correct — I ran the same suite on unmergedorigin/mainand diffed: identical. I did not take that on trust, because a suite that printsFAILand exits 0 is exactly the shape worth checking.My brief had a wrong number
I briefed "13 files name
target/fleetd.jar, 11 non-wiki". The implementer measured 9 non-wiki and said so plainly instead of bending its result to match me. It was right:The real total is 9 non-wiki + 2 wiki = 11, and one of the 9 (
plans/fleet01-standup/plan.md) is a dated historical snapshot deliberately left alone. My number was unmeasured. Noted so the next reader does not re-derive it from my brief.Why the redeploy is held — two blockers, both filed as #680
1. The script cannot see the daemon that is running right now.
PATTERNmoved fromtarget/fleetd.jartorun/fleetd.jar, and the live daemon runs fromtarget/. Measured with a positive control:So a redeploy today prints
no daemon running — this will be a cold start, skips the drain gate and the stop-and-wait, and starts a second daemon beside pid 42543. Two daemons on one herdr session take each other's members down..claude/skills/fleets-status/SKILL.md:62has the same narrowed pattern and would report the live daemon as not running.2. The installed launchd plist still names the old path.
~/Library/LaunchAgents/dev.ltms.fleetd.plist(written Aug 25) pointsProgramArgumentsat…/fleetd/target/fleetd.jar. The repo file is a template; editing it does not touch the installed copy. Since the swap ismv,target/fleetd.jarstops existing after a redeploy, so a reboot or aKeepAliverestart would launch a missing jar. The script checks the plist'sStandardOutPathand never itsProgramArguments.Neither is a defect in the merged change's own logic — the path split does what it says. They are guards the change needed to bring with it, and #680 covers all of it.
Also found: the core invariant is unpinned
Mutation on the merged tree:
report_jar_state's mismatch condition → killed, suite exits 1. Those new tests are real.JAR="$MODULE/run/fleetd.jar"→"$MODULE/target/fleetd.jar"→ survived, exit 0, output byte-identical.The second reverts the whole ticket and all 129 tests still pass. Cause measured, and it is "no assertion", not "cannot see": every test assigns its own
JAR/BUILD_JARbefore calling anything, so none observes the script's real value. The first mutation killing proves the sourcing mechanism works.The property was checked once, by hand, in the PR's own acceptance criterion 1 (
DIFFER: yes). A correct hand-check is not a guard.What remains before this closes
fleetd/run/fleetd.jar, runmvn -o installin the main clone while it runs, redeploy, and confirm thedrain complete: released=… abandoned=…line fromSessionManager.java:1126appears infleetd.out. That line is the pass.Current state is safe: the live daemon is still pid 42543 on the old jar, and I built only in throwaway worktrees, so the main clone's
target/fleetd.jaris untouched (still 20:05:53, 28785114 bytes).Correction to blocker 1 in my comment above. I wrote that a redeploy today would "start a second daemon beside pid 42543". That is wrong — I had not read the stop dispatch when I wrote it. Full reasoning in #680; the short version:
SUPERVISOR_KINDis resolved bylaunchd_loaded()vialaunchctl list, not byPATTERN. So with an emptyOLD_PIDthe script still takes theelif [ "$SUPERVISOR_KIND" = "launchd" ]branch and runslaunchctl unload -w, which does stop the running daemon. There is no second daemon.The blindness is still real and still blocks the redeploy. The damage is different:
drain_gate_required()needs[ -n "$old_pid" ], so live members lose in-flight reports with no prompt at all;ok "launchd agent unloaded (was already not running)"while it was running;wait_for_daemon_exitis never called on that branch;HAD_OLD_PIDis 0, soreport_shutdown_drainis suppressed — which is the signal this ticket's live probe depends on. The blindness hides the evidence that would close #664.I also over-stated one risk: the swap is
mvon one filesystem, so an open descriptor follows the inode and a running JVM is undisturbed. Overwriting would break it; moving does not.The hold on the redeploy is unchanged, and point 4 is now a stronger reason for it than what I originally wrote.