From a5d6ce1a377dbfdc11f99b0d7b37718898700564 Mon Sep 17 00:00:00 2001 From: Dai Ha Date: Sat, 3 Oct 2026 19:29:25 +0200 Subject: [PATCH 1/2] fleetd #664: protect the running daemon jar --- .claude/skills/redeploy-fleetd/SKILL.md | 12 ++++++++---- CLAUDE.md | 5 +++++ 2 files changed, 13 insertions(+), 4 deletions(-) diff --git a/.claude/skills/redeploy-fleetd/SKILL.md b/.claude/skills/redeploy-fleetd/SKILL.md index 18b0e17..01eea89 100644 --- a/.claude/skills/redeploy-fleetd/SKILL.md +++ b/.claude/skills/redeploy-fleetd/SKILL.md @@ -26,9 +26,14 @@ scripts/redeploy-fleetd.sh --no-build # restart the jar already on disk when you just built and nothing changed since. It gives up the protection in the next paragraph: no build runs, so a stale or missing jar is not caught early. The script still checks the file is there and dies with `no jar at … — run without --no-build` if it is not, but it cannot tell you the jar is -old. A `mvn clean` in the tree deletes that jar while the daemon keeps running on it, and nothing -degrades until the next restart. Run `--check` first: it prints the jar's hash and its modification -time, so you can see for yourself whether the jar is missing or older than the code you mean to ship. +old. Any build that writes `fleetd/target/fleetd.jar` while the daemon runs, including `mvn install` +with or without `clean`, breaks that daemon's shutdown drain. The drain loads its classes lazily at +shutdown from the jar file the JVM opened at boot. Deleting is not the only hazard; replacing the jar +is enough. Verify a merge by building in a throwaway git worktree. Let only +`scripts/redeploy-fleetd.sh` touch the main clone's jar. Its stage-then-swap protects its own build, +but it cannot undo a replacement that already happened. Run `--check` first: it prints the jar's hash +and its modification time, so you can see for yourself whether the jar is missing or older than the +code you mean to ship. It builds before it stops anything, so a failed build never leaves the fleet down; it waits for the old process to exit rather than assuming; it polls `/healthz`; and it anchors its log checks to a @@ -69,4 +74,3 @@ start every time. The rule lives in the operator's Claude Code settings: Granted by the operator on 2026-08-15. If a call is still refused, do **not** route around it by running the stop and start as separate commands — that is exactly the approval the script replaced. Say what you were going to run and why, and let the operator decide. - diff --git a/CLAUDE.md b/CLAUDE.md index 73675fb..99feb10 100644 --- a/CLAUDE.md +++ b/CLAUDE.md @@ -323,6 +323,11 @@ must obey belongs in the charter, not here. reference**, with the intent→tool table above as the short form. `McpContractDocTest` fails if that page names a `fleet_*` tool the server does not register. The flows are kept out of this file because this file loads into every session's context. +- **Never build into the main clone while `fleetd` runs.** Any build that writes + `fleetd/target/fleetd.jar`, with or without `clean`, breaks the shutdown drain because its classes + load lazily from the jar file the JVM opened at boot. Verify merges in a throwaway git worktree. + Let only `scripts/redeploy-fleetd.sh` touch the main clone's jar. Its stage-then-swap cannot undo + a replacement that already happened. ### Redeploying the daemon — the lead may do this (primary only) From 41cc785534ff01fbb3083ca7c5ab607e211611b4 Mon Sep 17 00:00:00 2001 From: Dai Ha Date: Sat, 3 Oct 2026 19:36:08 +0200 Subject: [PATCH 2/2] fleetd #664: explain delayed jar failure --- .claude/skills/redeploy-fleetd/SKILL.md | 4 +++- CLAUDE.md | 7 ++++--- 2 files changed, 7 insertions(+), 4 deletions(-) diff --git a/.claude/skills/redeploy-fleetd/SKILL.md b/.claude/skills/redeploy-fleetd/SKILL.md index 01eea89..7f98d6e 100644 --- a/.claude/skills/redeploy-fleetd/SKILL.md +++ b/.claude/skills/redeploy-fleetd/SKILL.md @@ -29,7 +29,8 @@ and dies with `no jar at … — run without --no-build` if it is not, but it ca old. Any build that writes `fleetd/target/fleetd.jar` while the daemon runs, including `mvn install` with or without `clean`, breaks that daemon's shutdown drain. The drain loads its classes lazily at shutdown from the jar file the JVM opened at boot. Deleting is not the only hazard; replacing the jar -is enough. Verify a merge by building in a throwaway git worktree. Let only +is enough. Nothing warns at the time. The damage appears at the next restart, where it looks like the +restart's fault. Verify a merge by building in a throwaway git worktree. Let only `scripts/redeploy-fleetd.sh` touch the main clone's jar. Its stage-then-swap protects its own build, but it cannot undo a replacement that already happened. Run `--check` first: it prints the jar's hash and its modification time, so you can see for yourself whether the jar is missing or older than the @@ -74,3 +75,4 @@ start every time. The rule lives in the operator's Claude Code settings: Granted by the operator on 2026-08-15. If a call is still refused, do **not** route around it by running the stop and start as separate commands — that is exactly the approval the script replaced. Say what you were going to run and why, and let the operator decide. + diff --git a/CLAUDE.md b/CLAUDE.md index 99feb10..4d4a20e 100644 --- a/CLAUDE.md +++ b/CLAUDE.md @@ -325,9 +325,10 @@ must obey belongs in the charter, not here. file because this file loads into every session's context. - **Never build into the main clone while `fleetd` runs.** Any build that writes `fleetd/target/fleetd.jar`, with or without `clean`, breaks the shutdown drain because its classes - load lazily from the jar file the JVM opened at boot. Verify merges in a throwaway git worktree. - Let only `scripts/redeploy-fleetd.sh` touch the main clone's jar. Its stage-then-swap cannot undo - a replacement that already happened. + load lazily from the jar file the JVM opened at boot. Nothing warns at the time. The damage appears + at the next restart, where it looks like the restart's fault. Verify merges in a throwaway git + worktree. Let only `scripts/redeploy-fleetd.sh` touch the main clone's jar. Its stage-then-swap + cannot undo a replacement that already happened. ### Redeploying the daemon — the lead may do this (primary only)