Compare commits

...

3 Commits

Author SHA1 Message Date
Dai Ha 41cc785534 fleetd #664: explain delayed jar failure
CI / shell-tests (pull_request) Failing after 8s
CI / contract (pull_request) Successful in 55s
CI / build (pull_request) Failing after 1m43s
2026-10-03 19:36:08 +02:00
Dai Ha a5d6ce1a37 fleetd #664: protect the running daemon jar
CI / shell-tests (pull_request) Failing after 7s
CI / contract (pull_request) Successful in 1m25s
CI / build (pull_request) Failing after 1m53s
2026-10-03 19:29:25 +02:00
Dai Ha 9417de1123 Merge PR #662: fleetd #659 — remove the dead back-compat forms around LeadConfigDirSource/leadContextLookup
CI / shell-tests (push) Failing after 6s
CI / contract (push) Successful in 1m33s
CI / build (push) Failing after 1m58s
2026-10-03 19:11:18 +02:00
2 changed files with 15 additions and 3 deletions
+9 -3
View File
@@ -26,9 +26,15 @@ scripts/redeploy-fleetd.sh --no-build # restart the jar already on disk
when you just built and nothing changed since. It gives up the protection in the next paragraph: no
build runs, so a stale or missing jar is not caught early. The script still checks the file is there
and dies with `no jar at … — run without --no-build` if it is not, but it cannot tell you the jar is
old. A `mvn clean` in the tree deletes that jar while the daemon keeps running on it, and nothing
degrades until the next restart. Run `--check` first: it prints the jar's hash and its modification
time, so you can see for yourself whether the jar is missing or older than the code you mean to ship.
old. Any build that writes `fleetd/target/fleetd.jar` while the daemon runs, including `mvn install`
with or without `clean`, breaks that daemon's shutdown drain. The drain loads its classes lazily at
shutdown from the jar file the JVM opened at boot. Deleting is not the only hazard; replacing the jar
is enough. Nothing warns at the time. The damage appears at the next restart, where it looks like the
restart's fault. Verify a merge by building in a throwaway git worktree. Let only
`scripts/redeploy-fleetd.sh` touch the main clone's jar. Its stage-then-swap protects its own build,
but it cannot undo a replacement that already happened. Run `--check` first: it prints the jar's hash
and its modification time, so you can see for yourself whether the jar is missing or older than the
code you mean to ship.
It builds before it stops anything, so a failed build never leaves the fleet down; it waits for the
old process to exit rather than assuming; it polls `/healthz`; and it anchors its log checks to a
+6
View File
@@ -323,6 +323,12 @@ must obey belongs in the charter, not here.
reference**, with the intent→tool table above as the short form. `McpContractDocTest` fails if
that page names a `fleet_*` tool the server does not register. The flows are kept out of this
file because this file loads into every session's context.
- **Never build into the main clone while `fleetd` runs.** Any build that writes
`fleetd/target/fleetd.jar`, with or without `clean`, breaks the shutdown drain because its classes
load lazily from the jar file the JVM opened at boot. Nothing warns at the time. The damage appears
at the next restart, where it looks like the restart's fault. Verify merges in a throwaway git
worktree. Let only `scripts/redeploy-fleetd.sh` touch the main clone's jar. Its stage-then-swap
cannot undo a replacement that already happened.
### Redeploying the daemon — the lead may do this (primary only)