Building in the main clone while the daemon runs breaks its shutdown drain — NoClassDefFoundError on a drain-only class #664

Open
opened 2026-10-03 19:19:30 +02:00 by ltms · 7 comments
Owner

Found during the redeploy after PR #662 merged, on 2026-10-03. The redeploy itself succeeded; this
is about the previous daemon's shutdown.

What happened

scripts/redeploy-fleetd.sh --yes reported this about the daemon it replaced:

== previous daemon's shutdown drain
   WARN  previous daemon's shutdown drain DIED — no drain-complete line, and an uncaught exception
   WARN  was found in its shutdown window (1 line(s)):
   WARN      Exception in thread "Thread-0" java.lang.NoClassDefFoundError: dev/ltms/fleet/session/SessionManager$DrainTally
   WARN  Some sessions from the PREVIOUS daemon may not have been released.

The class is not missing — the jar was swapped under a running JVM

SessionManager$DrainTally.class is present in the jar:

$ unzip -l fleetd/target/fleetd.jar | grep -F 'DrainTally'
     1901  10-03-2026 19:15   dev/ltms/fleet/session/SessionManager$DrainTally.class

(I first ran this with grep -i "SessionManager\$DrainTally" inside double quotes, where \$
collapses to a regex end-anchor, and got a clean 0 that looked like "the class is gone". Pair
every zero with a positive control — the control here was
unzip -l … | grep -cF 'SessionManager' → 5.)

Cause, and it was my own command

The daemon loads classes lazily from the jar it opened at boot. The drain path's classes are, by
definition, only loaded at shutdown. So if the jar file is replaced between boot and shutdown,
the shutdown hook tries to load a class out of a file whose content is no longer the one it opened.

The chain, measured:

  • Daemon pid 42930 started 16:33 on jar 5e68eb6e3d38 (recorded in the previous lead's hand-off).
  • I ran mvn -o install in the main clone at 19:12 to verify the merged tree. --check then
    reported the daemon's own jar path as a37441b9e9bc (2026-10-03 19:12:30) — a different file
    at the path pid 42930 was running from.
  • The redeploy script then built f3b541ba677a and swapped it in after stopping the daemon.
  • pid 42930's shutdown hook at ~19:17 could no longer load DrainTally.

So the jar at the running daemon's path was overwritten while it ran, by my build, before the
script was ever involved. The script's own stage-then-swap (#493) protected its build; it could not
undo a replacement that had already happened.

Impact

Low this time, and that is luck rather than design. I had already stopped all five members by hand
and fleet_list reported members: [], so the drain had nothing to release. With live members at
redeploy time, their sessions would not be released cleanly.

The general shape is worse than this one incident: any mvn install in the main clone while the
daemon runs leaves that daemon unable to execute its shutdown drain. Nothing warns at the time. The
damage only shows up at the next restart, attributed to the restart rather than to the build.

The project instruction is part of the problem

CLAUDE.md currently says, about verifying a merge in the main clone:

Never mvn clean in the main clone — it deletes the running daemon's jar. Use mvn -o install
there, or clean install in a throwaway worktree.

That guidance stops the jar being deleted but not being replaced, and replacement is enough
to break lazy class loading. A lead following the instruction exactly still breaks the drain. The
same text appears in the lead hand-off template.

Suggested fix

Pick one, or argue for another:

  1. Instruction-only. Change CLAUDE.md and the hand-off template to say: verify a merge by
    building in a throwaway worktree, and let only scripts/redeploy-fleetd.sh touch the main
    clone's jar. Cheapest, and it matches what the worktree advice already half says — but it
    relies on every future lead reading it.
  2. Make the daemon immune. Have fleetd preload the drain path's classes at boot, or copy its
    jar to a private path at startup and run from the copy, so a rebuild cannot reach it. Removes
    the hazard rather than documenting it.
  3. Make the build safe. Configure the build so the daemon's runtime jar path is never the
    maven output path — build to a staged name always, and have the script be the only thing that
    publishes to the runtime path.

Option 2 is the only one that survives someone not reading the instruction. Option 1 should happen
regardless, because it is true and cheap.

Not verified by me

  • Whether the drain had previously ever completed on this host. The script anchors its log checks to
    a marker taken before the restart, so I only have this one shutdown window, not a history.
  • Whether launchctl unload versus a plain signal changes the shutdown hook's behaviour here.
  • The exact set of classes the drain path loads lazily. DrainTally is the one that threw; there
    may be more behind it.
Found during the redeploy after PR #662 merged, on 2026-10-03. The redeploy itself succeeded; this is about the **previous** daemon's shutdown. ## What happened `scripts/redeploy-fleetd.sh --yes` reported this about the daemon it replaced: ``` == previous daemon's shutdown drain WARN previous daemon's shutdown drain DIED — no drain-complete line, and an uncaught exception WARN was found in its shutdown window (1 line(s)): WARN Exception in thread "Thread-0" java.lang.NoClassDefFoundError: dev/ltms/fleet/session/SessionManager$DrainTally WARN Some sessions from the PREVIOUS daemon may not have been released. ``` ## The class is not missing — the jar was swapped under a running JVM `SessionManager$DrainTally.class` is present in the jar: ``` $ unzip -l fleetd/target/fleetd.jar | grep -F 'DrainTally' 1901 10-03-2026 19:15 dev/ltms/fleet/session/SessionManager$DrainTally.class ``` (I first ran this with `grep -i "SessionManager\$DrainTally"` inside double quotes, where `\$` collapses to a regex end-anchor, and got a clean `0` that looked like "the class is gone". Pair every zero with a positive control — the control here was `unzip -l … | grep -cF 'SessionManager'` → 5.) ## Cause, and it was my own command The daemon loads classes lazily from the jar it opened at boot. The drain path's classes are, by definition, only loaded at shutdown. So if the jar file is **replaced** between boot and shutdown, the shutdown hook tries to load a class out of a file whose content is no longer the one it opened. The chain, measured: - Daemon pid 42930 started 16:33 on jar `5e68eb6e3d38` (recorded in the previous lead's hand-off). - I ran `mvn -o install` in the main clone at 19:12 to verify the merged tree. `--check` then reported the daemon's own jar path as `a37441b9e9bc (2026-10-03 19:12:30)` — a different file at the path pid 42930 was running from. - The redeploy script then built `f3b541ba677a` and swapped it in after stopping the daemon. - pid 42930's shutdown hook at ~19:17 could no longer load `DrainTally`. So the jar at the running daemon's path was overwritten while it ran, by **my** build, before the script was ever involved. The script's own stage-then-swap (#493) protected its build; it could not undo a replacement that had already happened. ## Impact Low this time, and that is luck rather than design. I had already stopped all five members by hand and `fleet_list` reported `members: []`, so the drain had nothing to release. With live members at redeploy time, their sessions would not be released cleanly. The general shape is worse than this one incident: **any** `mvn install` in the main clone while the daemon runs leaves that daemon unable to execute its shutdown drain. Nothing warns at the time. The damage only shows up at the next restart, attributed to the restart rather than to the build. ## The project instruction is part of the problem `CLAUDE.md` currently says, about verifying a merge in the main clone: > Never `mvn clean` in the main clone — it deletes the running daemon's jar. Use `mvn -o install` > there, or `clean install` in a throwaway worktree. That guidance stops the jar being **deleted** but not being **replaced**, and replacement is enough to break lazy class loading. A lead following the instruction exactly still breaks the drain. The same text appears in the lead hand-off template. ## Suggested fix Pick one, or argue for another: 1. **Instruction-only.** Change `CLAUDE.md` and the hand-off template to say: verify a merge by building in a throwaway worktree, and let only `scripts/redeploy-fleetd.sh` touch the main clone's jar. Cheapest, and it matches what the worktree advice already half says — but it relies on every future lead reading it. 2. **Make the daemon immune.** Have fleetd preload the drain path's classes at boot, or copy its jar to a private path at startup and run from the copy, so a rebuild cannot reach it. Removes the hazard rather than documenting it. 3. **Make the build safe.** Configure the build so the daemon's runtime jar path is never the maven output path — build to a staged name always, and have the script be the only thing that publishes to the runtime path. Option 2 is the only one that survives someone not reading the instruction. Option 1 should happen regardless, because it is true and cheap. ## Not verified by me - Whether the drain had previously ever completed on this host. The script anchors its log checks to a marker taken before the restart, so I only have this one shutdown window, not a history. - Whether `launchctl unload` versus a plain signal changes the shutdown hook's behaviour here. - The exact set of classes the drain path loads lazily. `DrainTally` is the one that threw; there may be more behind it.
Author
Owner

Lead review of PR #666 (the instruction half) — two small changes before merge

First, two corrections to this issue's own text

I measured both in the main clone. The issue body is wrong on them, so nobody should re-derive from it:

  1. The stale guidance was never in CLAUDE.md. It lived in exactly one place: .claude/skills/redeploy-fleetd/SKILL.md:29. Command: grep -rn "mvn clean" . --exclude-dir=.git --exclude-dir=wiki --exclude-dir=target. CLAUDE.md's own mvn clean hits are about IDE validation and CVE checks, a different topic.
  2. It is not in the lead hand-off template either. grep -nE "mvn|jar|build" .claude/skills/handover/SKILL.md finds no build guidance at all. The quoted sentence was probably read out of a previous lead's own .handover/HANDOVER.md, which is a transient file, not a template.

Verified good on PR #666

  • Docs-only. git diff --name-only origin/main..refs/pull/666/head | grep -E '\.(java|xml|sh|yaml|yml|json)$' returns 0.
  • The canonical block is untouched. Its byte range in CLAUDE.md is identical on main and on the PR head — 20938 bytes both sides.
  • The wiki sync check passes on the PR head. I ran the real check from CLAUDE.md against wiki/7-Use-Cases.md (submodule at a61f729): in sync: True. The worker correctly reported it could not run this and did not claim it as passed.
  • The new bullet sits inside the project addendum, after the flows bullet and before ### Redeploying the daemon. Right section.
  • The technical content is accurate: replacement and not only deletion, lazy class loading at shutdown from the jar opened at boot, throwaway worktree for merge checks, and the script's stage-then-swap not undoing an earlier replacement.

Change 1 — a true sentence was dropped, and it is the load-bearing one

The old text ended with "and nothing degrades until the next restart". The new text has no equivalent, in either file. I checked with a positive control:

$ git show origin/main:.claude/skills/redeploy-fleetd/SKILL.md | grep -niE "next restart|nothing degrades"
30:degrades until the next restart. Run `--check` first: ...

$ git show refs/pull/666/head:.claude/skills/redeploy-fleetd/SKILL.md | grep -niE "next restart|nothing warns|nothing degrades"
(no matching line; the only "until the" hit is line 9, about a different topic)

That fact is why the rule gets obeyed. Without it the hazard reads as theory. With it, a lead understands that the build looks completely fine at the time, and the breakage surfaces at the next restart where it is naturally blamed on the restart instead of on the build. That misattribution is exactly what happened here on 2026-10-03.

Add it back to both files, in plain words: nothing warns at the time, and the damage appears at the next restart, where it looks like the restart's fault.

Change 2 — restore the file's final newline

The diff also deletes the trailing empty line at the end of .claude/skills/redeploy-fleetd/SKILL.md. The worker flagged this itself, which is the right call. It is unrelated to the fix, so put it back and keep the diff to the two places that needed changing.

Nothing else. The rest of the PR is correct and I will merge it once these two land.

## Lead review of PR #666 (the instruction half) — two small changes before merge ### First, two corrections to this issue's own text I measured both in the main clone. The issue body is wrong on them, so nobody should re-derive from it: 1. **The stale guidance was never in `CLAUDE.md`.** It lived in exactly one place: `.claude/skills/redeploy-fleetd/SKILL.md:29`. Command: `grep -rn "mvn clean" . --exclude-dir=.git --exclude-dir=wiki --exclude-dir=target`. `CLAUDE.md`'s own `mvn clean` hits are about IDE validation and CVE checks, a different topic. 2. **It is not in the lead hand-off template either.** `grep -nE "mvn|jar|build" .claude/skills/handover/SKILL.md` finds no build guidance at all. The quoted sentence was probably read out of a previous lead's own `.handover/HANDOVER.md`, which is a transient file, not a template. ### Verified good on PR #666 - Docs-only. `git diff --name-only origin/main..refs/pull/666/head | grep -E '\.(java|xml|sh|yaml|yml|json)$'` returns **0**. - **The canonical block is untouched.** Its byte range in `CLAUDE.md` is identical on `main` and on the PR head — 20938 bytes both sides. - **The wiki sync check passes on the PR head.** I ran the real check from `CLAUDE.md` against `wiki/7-Use-Cases.md` (submodule at `a61f729`): `in sync: True`. The worker correctly reported it could not run this and did not claim it as passed. - The new bullet sits inside the project addendum, after the flows bullet and before `### Redeploying the daemon`. Right section. - The technical content is accurate: replacement and not only deletion, lazy class loading at shutdown from the jar opened at boot, throwaway worktree for merge checks, and the script's stage-then-swap not undoing an earlier replacement. ### Change 1 — a true sentence was dropped, and it is the load-bearing one The old text ended with "and nothing degrades **until the next restart**". The new text has no equivalent, in either file. I checked with a positive control: ``` $ git show origin/main:.claude/skills/redeploy-fleetd/SKILL.md | grep -niE "next restart|nothing degrades" 30:degrades until the next restart. Run `--check` first: ... $ git show refs/pull/666/head:.claude/skills/redeploy-fleetd/SKILL.md | grep -niE "next restart|nothing warns|nothing degrades" (no matching line; the only "until the" hit is line 9, about a different topic) ``` That fact is why the rule gets obeyed. Without it the hazard reads as theory. With it, a lead understands that the build looks completely fine at the time, and the breakage surfaces at the next restart where it is naturally blamed on the restart instead of on the build. That misattribution is exactly what happened here on 2026-10-03. Add it back to both files, in plain words: nothing warns at the time, and the damage appears at the next restart, where it looks like the restart's fault. ### Change 2 — restore the file's final newline The diff also deletes the trailing empty line at the end of `.claude/skills/redeploy-fleetd/SKILL.md`. The worker flagged this itself, which is the right call. It is unrelated to the fix, so put it back and keep the diff to the two places that needed changing. Nothing else. The rest of the PR is correct and I will merge it once these two land.
Author
Owner

The instruction half is merged as 4b4a868 and pushed. The code half (option 2) is still open.

PR #666 merged locally and pushed; I closed the PR by hand, since a local merge never closes it. This issue stays open — option 1 was always "should happen regardless", not the fix.

What I verified myself on the merged tree

  • Docs-only. git diff --name-only origin/main...<head> lists exactly two Markdown files; the code-file filter (.java|.xml|.sh|.yaml|.yml|.json) returns 0.
  • The canonical block is untouched, byte for byte: 20938 bytes on both main and the PR head.
  • The wiki sync check passes on the merged tree, run against wiki/7-Use-Cases.md: in sync: True.
  • The trailing newline is back. Last two bytes of SKILL.md are 0a0a, identical to main.
  • No line got longer than the file already allowed. Longest line unchanged at 271 (SKILL.md) and 464 (CLAUDE.md).
  • I built it anyway, and it is green. "Docs-only" is not a free pass here: six test files under fleetd/src/test/java reference CLAUDE.md or .claude/skills (OpenCodeLauncherTest, FleetProfilesLiveDefaultTest, ClaudeCodeLauncherTest, McpContractDocTest, RestRouteInventoryTest, GitWorktreesTest), so a doc edit can break the build. Built the merged tree in a throwaway worktree: Tests run: 1923, Failures: 0, Errors: 0, Skipped: 0, BUILD SUCCESS, exit 0, 170 report files.

The review round

The first version dropped a true sentence that the old text had carried — "nothing degrades until the next restart". I caught it with a positive control rather than by eye:

$ git show origin/main:.claude/skills/redeploy-fleetd/SKILL.md | grep -niE "next restart|nothing degrades"
30:degrades until the next restart. ...

$ git show <pr-head>:.claude/skills/redeploy-fleetd/SKILL.md | grep -niE "next restart|nothing warns|nothing degrades"
(nothing)

That fact is the load-bearing one. Without it the hazard reads as theory; with it a reader understands the build looks completely fine at the time and the breakage surfaces at the next restart, where it gets blamed on the restart. Both files now say so.

Still owed on this issue — option 2

The two architects were asked which immunity fix to take: 2a preload the drain path's classes at boot, or 2b run from a private copy of the jar. The sol architect has reported and chose 2b. I am waiting on the second position before I settle it, and I will write the decision here.

One thing from sol's report is worth recording now because it bears on this issue's own "Not verified by me" list, and I checked it myself:

  • The question "what is the exact set of classes the drain path loads lazily?" has no stable answer. DrainTally is loaded unconditionally — drainAll assigns it at SessionManager.java:1117 and even an empty snapshot constructs it. But ReleaseCause, ReleaseDetail and ShuttingDownException load only on particular branches, and FleetdRuntime.close() continues through many more components after the drain (FleetdRuntime.java:130-164). So the set depends on live state: listeners wired, backend kind, broker kind, worktree state, and which error branches fire. That is an argument against 2a on its merits, not merely a gap in testing.
  • Nothing in production Java resolves its own jar location. grep -rn -E 'CodeSource|getProtectionDomain|getLocation\(' fleetd/src/main/java returns 0, and my positive control on the same tree (getProperty("user.home") → 6 hits) shows the search works rather than silently matching nothing. So 2b needs no change to Java self-location code; its cost is all in the deployment surface.
## The instruction half is merged as `4b4a868` and pushed. The code half (option 2) is still open. PR #666 merged locally and pushed; I closed the PR by hand, since a local merge never closes it. **This issue stays open** — option 1 was always "should happen regardless", not the fix. ### What I verified myself on the merged tree - **Docs-only.** `git diff --name-only origin/main...<head>` lists exactly two Markdown files; the code-file filter (`.java|.xml|.sh|.yaml|.yml|.json`) returns **0**. - **The canonical block is untouched**, byte for byte: 20938 bytes on both `main` and the PR head. - **The wiki sync check passes** on the merged tree, run against `wiki/7-Use-Cases.md`: `in sync: True`. - **The trailing newline is back.** Last two bytes of `SKILL.md` are `0a0a`, identical to `main`. - **No line got longer than the file already allowed.** Longest line unchanged at 271 (`SKILL.md`) and 464 (`CLAUDE.md`). - **I built it anyway, and it is green.** "Docs-only" is not a free pass here: six test files under `fleetd/src/test/java` reference `CLAUDE.md` or `.claude/skills` (`OpenCodeLauncherTest`, `FleetProfilesLiveDefaultTest`, `ClaudeCodeLauncherTest`, `McpContractDocTest`, `RestRouteInventoryTest`, `GitWorktreesTest`), so a doc edit can break the build. Built the merged tree in a throwaway worktree: `Tests run: 1923, Failures: 0, Errors: 0, Skipped: 0`, `BUILD SUCCESS`, exit 0, 170 report files. ### The review round The first version dropped a true sentence that the old text had carried — "nothing degrades until the next restart". I caught it with a positive control rather than by eye: ``` $ git show origin/main:.claude/skills/redeploy-fleetd/SKILL.md | grep -niE "next restart|nothing degrades" 30:degrades until the next restart. ... $ git show <pr-head>:.claude/skills/redeploy-fleetd/SKILL.md | grep -niE "next restart|nothing warns|nothing degrades" (nothing) ``` That fact is the load-bearing one. Without it the hazard reads as theory; with it a reader understands the build looks completely fine at the time and the breakage surfaces at the next restart, where it gets blamed on the restart. Both files now say so. ### Still owed on this issue — option 2 The two architects were asked which immunity fix to take: **2a** preload the drain path's classes at boot, or **2b** run from a private copy of the jar. The `sol` architect has reported and chose **2b**. I am waiting on the second position before I settle it, and I will write the decision here. One thing from `sol`'s report is worth recording now because it bears on this issue's own "Not verified by me" list, and I checked it myself: - **The question "what is the exact set of classes the drain path loads lazily?" has no stable answer.** `DrainTally` is loaded unconditionally — `drainAll` assigns it at `SessionManager.java:1117` and even an empty snapshot constructs it. But `ReleaseCause`, `ReleaseDetail` and `ShuttingDownException` load only on particular branches, and `FleetdRuntime.close()` continues through many more components after the drain (`FleetdRuntime.java:130-164`). So the set depends on live state: listeners wired, backend kind, broker kind, worktree state, and which error branches fire. That is an argument against 2a on its merits, not merely a gap in testing. - **Nothing in production Java resolves its own jar location.** `grep -rn -E 'CodeSource|getProtectionDomain|getLocation\(' fleetd/src/main/java` returns **0**, and my positive control on the same tree (`getProperty("user.home")` → 6 hits) shows the search works rather than silently matching nothing. So 2b needs no change to Java self-location code; its cost is all in the deployment surface.
Author
Owner

Decision: 2b, in the "separate runtime path, atomic rename" variant

Both architects formed positions independently and both chose 2b. They agree, so this is settled at the fleet level and does not go to the operator. sol argued 2b on the grounds that a preload list cannot be completed. opus argued the same and went further, specifying the variant below. I checked the load-bearing claims myself before accepting.

2a is rejected. Do not implement it, and do not add it alongside 2b "for belt and braces."

What to build

The daemon runs from fleetd/run/fleetd.jar — outside target/, so neither mvn install nor mvn clean can reach it. The deploy handoff is a single mv from target/fleetd.jar to run/fleetd.jar, performed only after the old daemon is confirmed gone, reusing the gate swap_staged_jar already has.

Why a mv and not a copy at startup:

  • A mv within one filesystem is rename(2), which is atomic. There is never a half-written runtime jar. A cp has a truncation window, and that window is the failure we are fixing.
  • A copy at startup would have to live in three launch routes — the launchd plist, deploy/fleetd.service:53, and the nohup at scripts/redeploy-fleetd.sh:1025. Three places to forget, and forgetting is silent.
  • Maven's only output path stays target/, so mvn install becomes harmless from any clone, by any lead, whether or not they read the instructions. That is the actual goal: the current rule is enforced by prose alone.

What I verified myself in the code

I did not take the inventory on trust. Every line below is from a command I ran in the main clone at b4b7cf5:

  • fleetd/pom.xml:194 — <finalName>fleetd</finalName>, with no <directory> or <outputDirectory> override. This is why Maven's output path is the runtime path. Root cause confirmed.
  • scripts/redeploy-fleetd.sh:74 JAR="$MODULE/target/fleetd.jar", :81 JAR_STAGED=, :87 PATTERN='target/fleetd.jar'.
  • Both launch routes hardcode the path: deploy/fleetd.service:53 and deploy/dev.ltms.fleetd.plist:50.
  • .claude/skills/fleets-status/SKILL.md:62 runs pgrep -f 'target/fleetd.jar'.
  • No Java resolves its own jar. A grep for getProtectionDomain|getCodeSource|java.class.path|ProcessHandle over fleetd/src/main/java returns only LsofPeerPidLookup.java:20 (own pid) and ParentResolver.java:17-18 (parent walk). So the Java side needs no change, and 2b cannot be done by the daemon — the JVM opens the jar before main runs. It belongs in the launch path.
  • 11 files name target/fleetd.jar outside target/ and .git.

The two silent breakages to handle, not discover

  1. PATTERN at scripts/redeploy-fleetd.sh:87. running_pid() is pgrep -f "$PATTERN". Move the jar and leave the pattern, and running_pid returns nothing — the script then believes no daemon is running and starts a second one. The comment at :209-217 already warns about this.
  2. .claude/skills/fleets-status/SKILL.md:62. Same pattern, and it would report no fleet at all. A skill is not covered by any test, so nothing catches this.

Both fail by reporting absence, which reads as good news. Pair every such check with a positive control.

--check must stop lying

scripts/redeploy-fleetd.sh:1173 prints jar_id and date -r "$JAR" under the label "jar on disk". With one jar path that label is one name for two different facts, and during this incident it would have printed the new jar's hash while the JVM ran the old one. Under 2b there are two facts, so print both: built hash/mtime and running hash/mtime. A visible mismatch is the whole value.

Accept this limit explicitly

2b does not remove the stale-jar risk; it adds one more place to get the handoff wrong, and a daemon booted from a stale run/fleetd.jar answers /healthz and looks healthy. That risk already exists — --no-build restarts whatever jar is at the path today. The reason 2b still wins: it makes the instrument honest instead of breaking it, and 2a's failure mode is both silent and worse, landing at shutdown with live members.

What settles it

One live probe, lead-only: boot from fleetd/run/fleetd.jar, run mvn -o install in the main clone while it runs, then redeploy and check fleetd.out for the drain complete: released=… abandoned=… line from SessionManager.java:1126. That line appearing is the pass. Neither architect ran it — correctly, since it would cut their own channel.

Claims I am passing on without checking

  • That rename(2) leaves a running JVM's open inode intact while an in-place rewrite does not. This is POSIX semantics and the incident fits it, but neither architect tested it on this host, and nor have I.
  • Whether maven-shade-plugin truncates in place or renames. The incident proves the write reached the inode the JVM held, whichever did it.
  • fleetd/fleetd.yaml (gitignored) and wiki/ (a submodule) were not searched for jar paths by the architect, whose worktree lacks both. I have not searched them either, so the 11-file inventory may be short.

The docs-only fix for this ticket already shipped in 4b4a868. This decision covers the code change, which is not yet implemented.

## Decision: 2b, in the "separate runtime path, atomic rename" variant Both architects formed positions independently and both chose **2b**. They agree, so this is settled at the fleet level and does not go to the operator. `sol` argued 2b on the grounds that a preload list cannot be completed. `opus` argued the same and went further, specifying the variant below. I checked the load-bearing claims myself before accepting. **2a is rejected. Do not implement it, and do not add it alongside 2b "for belt and braces."** ### What to build The daemon runs from **`fleetd/run/fleetd.jar`** — outside `target/`, so neither `mvn install` nor `mvn clean` can reach it. The deploy handoff is a single `mv` from `target/fleetd.jar` to `run/fleetd.jar`, performed only after the old daemon is confirmed gone, reusing the gate `swap_staged_jar` already has. Why a `mv` and not a copy at startup: - A `mv` within one filesystem is `rename(2)`, which is atomic. There is never a half-written runtime jar. A `cp` has a truncation window, and that window is the failure we are fixing. - A copy at startup would have to live in three launch routes — the launchd plist, `deploy/fleetd.service:53`, and the `nohup` at `scripts/redeploy-fleetd.sh:1025`. Three places to forget, and forgetting is silent. - Maven's only output path stays `target/`, so `mvn install` becomes harmless from any clone, by any lead, whether or not they read the instructions. That is the actual goal: the current rule is enforced by prose alone. ### What I verified myself in the code I did not take the inventory on trust. Every line below is from a command I ran in the main clone at `b4b7cf5`: - `fleetd/pom.xml:194` — `<finalName>fleetd</finalName>`, with no `<directory>` or `<outputDirectory>` override. This is why Maven's output path *is* the runtime path. Root cause confirmed. - `scripts/redeploy-fleetd.sh:74` `JAR="$MODULE/target/fleetd.jar"`, `:81` `JAR_STAGED=`, `:87` `PATTERN='target/fleetd.jar'`. - Both launch routes hardcode the path: `deploy/fleetd.service:53` and `deploy/dev.ltms.fleetd.plist:50`. - `.claude/skills/fleets-status/SKILL.md:62` runs `pgrep -f 'target/fleetd.jar'`. - No Java resolves its own jar. A grep for `getProtectionDomain|getCodeSource|java.class.path|ProcessHandle` over `fleetd/src/main/java` returns only `LsofPeerPidLookup.java:20` (own pid) and `ParentResolver.java:17-18` (parent walk). **So the Java side needs no change, and 2b cannot be done by the daemon** — the JVM opens the jar before `main` runs. It belongs in the launch path. - **11** files name `target/fleetd.jar` outside `target/` and `.git`. ### The two silent breakages to handle, not discover 1. **`PATTERN` at `scripts/redeploy-fleetd.sh:87`.** `running_pid()` is `pgrep -f "$PATTERN"`. Move the jar and leave the pattern, and `running_pid` returns nothing — the script then believes no daemon is running and starts a second one. The comment at `:209-217` already warns about this. 2. **`.claude/skills/fleets-status/SKILL.md:62`.** Same pattern, and it would report no fleet at all. A skill is not covered by any test, so nothing catches this. Both fail by reporting *absence*, which reads as good news. Pair every such check with a positive control. ### `--check` must stop lying `scripts/redeploy-fleetd.sh:1173` prints `jar_id` and `date -r "$JAR"` under the label "jar on disk". With one jar path that label is one name for two different facts, and during this incident it would have printed the **new** jar's hash while the JVM ran the old one. Under 2b there are two facts, so print both: built hash/mtime and running hash/mtime. A visible mismatch is the whole value. ### Accept this limit explicitly 2b does not remove the stale-jar risk; it adds one more place to get the handoff wrong, and a daemon booted from a stale `run/fleetd.jar` answers `/healthz` and looks healthy. That risk already exists — `--no-build` restarts whatever jar is at the path today. The reason 2b still wins: it makes the instrument honest instead of breaking it, and 2a's failure mode is both silent *and* worse, landing at shutdown with live members. ### What settles it One live probe, lead-only: boot from `fleetd/run/fleetd.jar`, run `mvn -o install` in the main clone while it runs, then redeploy and check `fleetd.out` for the `drain complete: released=… abandoned=…` line from `SessionManager.java:1126`. That line appearing is the pass. Neither architect ran it — correctly, since it would cut their own channel. ### Claims I am passing on without checking - That `rename(2)` leaves a running JVM's open inode intact while an in-place rewrite does not. This is POSIX semantics and the incident fits it, but **neither architect tested it on this host**, and nor have I. - Whether `maven-shade-plugin` truncates in place or renames. The incident proves the write reached the inode the JVM held, whichever did it. - `fleetd/fleetd.yaml` (gitignored) and `wiki/` (a submodule) were not searched for jar paths by the architect, whose worktree lacks both. I have not searched them either, so the 11-file inventory may be short. The docs-only fix for this ticket already shipped in **4b4a868**. This decision covers the code change, which is not yet implemented.
Author
Owner

Addendum: I closed the two gaps the architect could not reach from a worker worktree, and the inventory was indeed short.

  • fleetd/fleetd.yaml (gitignored, so absent from every member's worktree): no jar path. Nothing to change there.
  • wiki/ (a submodule, uninitialized in every member's worktree — git submodule status prints +a61f729, 19 entries in my clone): two more sites.
    • wiki/13-User-Guide.md:93 — java -jar fleetd/target/fleetd.jar fleetd/fleetd.yaml
    • wiki/11-Features.md:2202 — describes the launchd job as running java -jar target/fleetd.jar fleetd.yaml

So the inventory is 13 files, not 11. Both new sites are documentation that tells a human how to start the daemon, so leaving them stale would send an operator to boot from target/ — the one path 2b exists to stop anyone running from. They must be updated in the same change, and the wiki edits are lead-only.

This is the pattern worth naming: the architect listed every site it could see and said plainly which two surfaces it could not. That caveat is what made the gap findable. A report that had simply said "11 files" would have been read as complete.

Addendum: I closed the two gaps the architect could not reach from a worker worktree, and the inventory was indeed short. - **`fleetd/fleetd.yaml`** (gitignored, so absent from every member's worktree): **no jar path**. Nothing to change there. - **`wiki/`** (a submodule, uninitialized in every member's worktree — `git submodule status` prints `+a61f729`, 19 entries in my clone): **two more sites**. - `wiki/13-User-Guide.md:93` — `java -jar fleetd/target/fleetd.jar fleetd/fleetd.yaml` - `wiki/11-Features.md:2202` — describes the launchd job as running `java -jar target/fleetd.jar fleetd.yaml` So the inventory is **13** files, not 11. Both new sites are documentation that tells a human how to start the daemon, so leaving them stale would send an operator to boot from `target/` — the one path 2b exists to stop anyone running from. They must be updated in the same change, and the wiki edits are lead-only. This is the pattern worth naming: the architect listed every site it *could* see and said plainly which two surfaces it could not. That caveat is what made the gap findable. A report that had simply said "11 files" would have been read as complete.
Author
Owner

Before-state, measured on the live daemon 2026-10-03 21:05 CEST. Recording it now because the redeploy that closes this ticket destroys the reading.

$ PID=$(lsof -nP -iTCP:8765 -sTCP:LISTEN -t | head -1); echo $PID
42543

$ ps -o lstart= -p 42543
Sat Oct  3 20:05:53 2026

$ lsof -p 42543 | grep -i '\.jar' | awk '{print $NF}' | sort -u
/Users/dai.ha/LTMS/claude-bridge/fleetd/target/fleetd.jar

$ ls -la fleetd/run/
ls: fleetd/run/: No such file or directory

$ stat -f '%Sm %z bytes %N' -t '%Y-%m-%d %H:%M:%S' fleetd/target/fleetd.jar
2026-10-03 20:05:53 28785114 bytes .../fleetd/target/fleetd.jar

Three things this pins:

  1. The hazard is real right now, not theoretical. The running JVM holds an open descriptor on fleetd/target/fleetd.jar — the exact path mvn install in the main clone writes. This is the premise the whole ticket rests on, and it is now measured rather than assumed.
  2. fleetd/run/ does not exist yet, so after the change lands, its appearance is itself a check that the new path is being used.
  3. The jar mtime equals the process start time to the second, which is consistent with the last redeploy and means nothing has overwritten it since boot.

I used ps -o lstart= for the start time, not the log — fleetd.out timestamps carry no date and drift timezone, so two boots in one file can read hours apart.

The live probe that settles this ticket is mine, not the implementer's. Booting from the new path and then running mvn -o install in the main clone while it runs would cut a worker's own channel mid-turn. The pass condition is the drain complete: released=… abandoned=… line from SessionManager.java:1126 appearing in fleetd.out after that sequence — today that drain is what breaks, and nothing warns at the time.

**Before-state, measured on the live daemon 2026-10-03 21:05 CEST.** Recording it now because the redeploy that closes this ticket destroys the reading. ``` $ PID=$(lsof -nP -iTCP:8765 -sTCP:LISTEN -t | head -1); echo $PID 42543 $ ps -o lstart= -p 42543 Sat Oct 3 20:05:53 2026 $ lsof -p 42543 | grep -i '\.jar' | awk '{print $NF}' | sort -u /Users/dai.ha/LTMS/claude-bridge/fleetd/target/fleetd.jar $ ls -la fleetd/run/ ls: fleetd/run/: No such file or directory $ stat -f '%Sm %z bytes %N' -t '%Y-%m-%d %H:%M:%S' fleetd/target/fleetd.jar 2026-10-03 20:05:53 28785114 bytes .../fleetd/target/fleetd.jar ``` Three things this pins: 1. **The hazard is real right now, not theoretical.** The running JVM holds an open descriptor on `fleetd/target/fleetd.jar` — the exact path `mvn install` in the main clone writes. This is the premise the whole ticket rests on, and it is now measured rather than assumed. 2. **`fleetd/run/` does not exist yet**, so after the change lands, its appearance is itself a check that the new path is being used. 3. The jar mtime equals the process start time to the second, which is consistent with the last redeploy and means nothing has overwritten it since boot. I used `ps -o lstart=` for the start time, not the log — `fleetd.out` timestamps carry no date and drift timezone, so two boots in one file can read hours apart. **The live probe that settles this ticket is mine, not the implementer's.** Booting from the new path and then running `mvn -o install` in the main clone while it runs would cut a worker's own channel mid-turn. The pass condition is the `drain complete: released=… abandoned=…` line from `SessionManager.java:1126` appearing in `fleetd.out` after that sequence — today that drain is what breaks, and nothing warns at the time.
Author
Owner

PR #679 merged as 5051a06 (pushed, origin/main and local HEAD both 5051a06). This ticket stays open — the code is in, the live probe that settles it is not done, and I am holding the redeploy. Reasons below.

What I verified myself before merging

Not taking the implementer's word for any of it:

check result
three-dot diff origin/main...pr/679 1 commit, 9 files, nothing dragged in
trial merge in a throwaway worktree no conflicts
mvn -o clean install from fleetd/ Tests run: 1929, Failures: 0, BUILD SUCCESS, 172 report files, 0 failure lines
build writes target/fleetd.jar, does not create run/ confirmed
run/ ignored by the new rule fleetd/.gitignore:7:run/
scripts/test-redeploy-fleetd.sh exit 0

The implementer reported 1928 tests / 171 reports; I measured 1929 / 172. The gap is explained: it built on merge-base 7f9a9c0, before the test that landed in 6f27522 today.

Its claim that the suite's 3 FAIL: lines and 5 mktemp: lines are deliberate self-test output is correct — I ran the same suite on unmerged origin/main and diffed: identical. I did not take that on trust, because a suite that prints FAIL and exits 0 is exactly the shape worth checking.

My brief had a wrong number

I briefed "13 files name target/fleetd.jar, 11 non-wiki". The implementer measured 9 non-wiki and said so plainly instead of bending its result to match me. It was right:

$ git grep -l 'target/fleetd\.jar' origin/main -- . ':!*/target/*' | wc -l
9

The real total is 9 non-wiki + 2 wiki = 11, and one of the 9 (plans/fleet01-standup/plan.md) is a dated historical snapshot deliberately left alone. My number was unmeasured. Noted so the next reader does not re-derive it from my brief.

Why the redeploy is held — two blockers, both filed as #680

1. The script cannot see the daemon that is running right now. PATTERN moved from target/fleetd.jar to run/fleetd.jar, and the live daemon runs from target/. Measured with a positive control:

$ pgrep -f 'run/fleetd.jar'            # new pattern
                                        <-- empty
$ pgrep -f 'target/fleetd.jar'         # old pattern, control
42543 94037                             <-- live daemon + a simulated hand-run

So a redeploy today prints no daemon running — this will be a cold start, skips the drain gate and the stop-and-wait, and starts a second daemon beside pid 42543. Two daemons on one herdr session take each other's members down. .claude/skills/fleets-status/SKILL.md:62 has the same narrowed pattern and would report the live daemon as not running.

2. The installed launchd plist still names the old path. ~/Library/LaunchAgents/dev.ltms.fleetd.plist (written Aug 25) points ProgramArguments at …/fleetd/target/fleetd.jar. The repo file is a template; editing it does not touch the installed copy. Since the swap is mv, target/fleetd.jar stops existing after a redeploy, so a reboot or a KeepAlive restart would launch a missing jar. The script checks the plist's StandardOutPath and never its ProgramArguments.

Neither is a defect in the merged change's own logic — the path split does what it says. They are guards the change needed to bring with it, and #680 covers all of it.

Also found: the core invariant is unpinned

Mutation on the merged tree:

  • invert report_jar_state's mismatch condition → killed, suite exits 1. Those new tests are real.
  • JAR="$MODULE/run/fleetd.jar" → "$MODULE/target/fleetd.jar" → survived, exit 0, output byte-identical.

The second reverts the whole ticket and all 129 tests still pass. Cause measured, and it is "no assertion", not "cannot see": every test assigns its own JAR/BUILD_JAR before calling anything, so none observes the script's real value. The first mutation killing proves the sourcing mechanism works.

The property was checked once, by hand, in the PR's own acceptance criterion 1 (DIFFER: yes). A correct hand-check is not a guard.

What remains before this closes

  1. #680 lands (all three parts).
  2. I reinstall the launchd plist — lead-only, outside the repo.
  3. The live probe: boot from fleetd/run/fleetd.jar, run mvn -o install in the main clone while it runs, redeploy, and confirm the drain complete: released=… abandoned=… line from SessionManager.java:1126 appears in fleetd.out. That line is the pass.
  4. The two wiki pages, plus a Features entry — lead-only.

Current state is safe: the live daemon is still pid 42543 on the old jar, and I built only in throwaway worktrees, so the main clone's target/fleetd.jar is untouched (still 20:05:53, 28785114 bytes).

**PR #679 merged as `5051a06`** (pushed, `origin/main` and local HEAD both `5051a06`). **This ticket stays open** — the code is in, the live probe that settles it is not done, and I am holding the redeploy. Reasons below. ## What I verified myself before merging Not taking the implementer's word for any of it: | check | result | |---|---| | three-dot diff `origin/main...pr/679` | 1 commit, 9 files, nothing dragged in | | trial merge in a throwaway worktree | no conflicts | | `mvn -o clean install` from `fleetd/` | `Tests run: 1929, Failures: 0`, `BUILD SUCCESS`, 172 report files, 0 failure lines | | build writes `target/fleetd.jar`, does **not** create `run/` | confirmed | | `run/` ignored by the new rule | `fleetd/.gitignore:7:run/` | | `scripts/test-redeploy-fleetd.sh` | exit 0 | The implementer reported 1928 tests / 171 reports; I measured 1929 / 172. The gap is explained: it built on merge-base `7f9a9c0`, before the test that landed in `6f27522` today. Its claim that the suite's 3 `FAIL:` lines and 5 `mktemp:` lines are deliberate self-test output **is correct** — I ran the same suite on unmerged `origin/main` and diffed: identical. I did not take that on trust, because a suite that prints `FAIL` and exits 0 is exactly the shape worth checking. ## My brief had a wrong number I briefed "13 files name `target/fleetd.jar`, 11 non-wiki". The implementer measured 9 non-wiki and said so plainly instead of bending its result to match me. It was right: ``` $ git grep -l 'target/fleetd\.jar' origin/main -- . ':!*/target/*' | wc -l 9 ``` The real total is 9 non-wiki + 2 wiki = **11**, and one of the 9 (`plans/fleet01-standup/plan.md`) is a dated historical snapshot deliberately left alone. My number was unmeasured. Noted so the next reader does not re-derive it from my brief. ## Why the redeploy is held — two blockers, both filed as #680 **1. The script cannot see the daemon that is running right now.** `PATTERN` moved from `target/fleetd.jar` to `run/fleetd.jar`, and the live daemon runs from `target/`. Measured with a positive control: ``` $ pgrep -f 'run/fleetd.jar' # new pattern <-- empty $ pgrep -f 'target/fleetd.jar' # old pattern, control 42543 94037 <-- live daemon + a simulated hand-run ``` So a redeploy today prints `no daemon running — this will be a cold start`, skips the drain gate and the stop-and-wait, and starts a **second** daemon beside pid 42543. Two daemons on one herdr session take each other's members down. `.claude/skills/fleets-status/SKILL.md:62` has the same narrowed pattern and would report the live daemon as not running. **2. The installed launchd plist still names the old path.** `~/Library/LaunchAgents/dev.ltms.fleetd.plist` (written Aug 25) points `ProgramArguments` at `…/fleetd/target/fleetd.jar`. The repo file is a template; editing it does not touch the installed copy. Since the swap is `mv`, `target/fleetd.jar` stops existing after a redeploy, so a reboot or a `KeepAlive` restart would launch a missing jar. The script checks the plist's `StandardOutPath` and never its `ProgramArguments`. Neither is a defect in the merged change's own logic — the path split does what it says. They are guards the change needed to bring with it, and #680 covers all of it. ## Also found: the core invariant is unpinned Mutation on the merged tree: - invert `report_jar_state`'s mismatch condition → **killed**, suite exits 1. Those new tests are real. - `JAR="$MODULE/run/fleetd.jar"` → `"$MODULE/target/fleetd.jar"` → **survived**, exit 0, output byte-identical. The second reverts the whole ticket and all 129 tests still pass. Cause measured, and it is "no assertion", not "cannot see": every test assigns its own `JAR`/`BUILD_JAR` before calling anything, so none observes the script's real value. The first mutation killing proves the sourcing mechanism works. The property *was* checked once, by hand, in the PR's own acceptance criterion 1 (`DIFFER: yes`). A correct hand-check is not a guard. ## What remains before this closes 1. #680 lands (all three parts). 2. I reinstall the launchd plist — lead-only, outside the repo. 3. The live probe: boot from `fleetd/run/fleetd.jar`, run `mvn -o install` in the main clone while it runs, redeploy, and confirm the `drain complete: released=… abandoned=…` line from `SessionManager.java:1126` appears in `fleetd.out`. That line is the pass. 4. The two wiki pages, plus a Features entry — lead-only. Current state is safe: the live daemon is still pid 42543 on the old jar, and I built only in throwaway worktrees, so the main clone's `target/fleetd.jar` is untouched (still 20:05:53, 28785114 bytes).
Author
Owner

Correction to blocker 1 in my comment above. I wrote that a redeploy today would "start a second daemon beside pid 42543". That is wrong — I had not read the stop dispatch when I wrote it. Full reasoning in #680; the short version:

SUPERVISOR_KIND is resolved by launchd_loaded() via launchctl list, not by PATTERN. So with an empty OLD_PID the script still takes the elif [ "$SUPERVISOR_KIND" = "launchd" ] branch and runs launchctl unload -w, which does stop the running daemon. There is no second daemon.

The blindness is still real and still blocks the redeploy. The damage is different:

  • the drain gate is skipped silently — drain_gate_required() needs [ -n "$old_pid" ], so live members lose in-flight reports with no prompt at all;
  • the script prints ok "launchd agent unloaded (was already not running)" while it was running;
  • wait_for_daemon_exit is never called on that branch;
  • HAD_OLD_PID is 0, so report_shutdown_drain is suppressed — which is the signal this ticket's live probe depends on. The blindness hides the evidence that would close #664.

I also over-stated one risk: the swap is mv on one filesystem, so an open descriptor follows the inode and a running JVM is undisturbed. Overwriting would break it; moving does not.

The hold on the redeploy is unchanged, and point 4 is now a stronger reason for it than what I originally wrote.

**Correction to blocker 1 in my comment above.** I wrote that a redeploy today would "start a **second** daemon beside pid 42543". That is wrong — I had not read the stop dispatch when I wrote it. Full reasoning in #680; the short version: `SUPERVISOR_KIND` is resolved by `launchd_loaded()` via `launchctl list`, not by `PATTERN`. So with an empty `OLD_PID` the script still takes the `elif [ "$SUPERVISOR_KIND" = "launchd" ]` branch and runs `launchctl unload -w`, which does stop the running daemon. There is no second daemon. The blindness is still real and still blocks the redeploy. The damage is different: - the drain gate is skipped silently — `drain_gate_required()` needs `[ -n "$old_pid" ]`, so live members lose in-flight reports with no prompt at all; - the script prints `ok "launchd agent unloaded (was already not running)"` while it *was* running; - `wait_for_daemon_exit` is never called on that branch; - `HAD_OLD_PID` is 0, so `report_shutdown_drain` is suppressed — **which is the signal this ticket's live probe depends on**. The blindness hides the evidence that would close #664. I also over-stated one risk: the swap is `mv` on one filesystem, so an open descriptor follows the inode and a running JVM is undisturbed. Overwriting would break it; moving does not. The hold on the redeploy is unchanged, and point 4 is now a stronger reason for it than what I originally wrote.
Sign in to join this conversation.
1 Participants
Notifications
Due Date
No due date set.
Dependencies

No dependencies set.

Reference: fleet/fleetd#664