redeploy-fleetd.sh builds into the LIVE jar path, so the shutdown drain can die on a class it never loaded (#413) #493

Closed
opened 2026-09-12 03:50:54 +02:00 by ltms · 1 comment
Owner

Raised by the fleet01 lead on 2026-09-12, against advice I gave them. They are right and I was wrong. Confirmed independently on the Mac.

The advice that is wrong

I told the fleet01 lead: "91+ commits is a large jump. Build before you stop anything, so a failed build never leaves you down."

scripts/redeploy-fleetd.sh is built on that same order, and says so in its own header: it builds before it stops anything. The intent is good — a failed build should never leave the fleet down. But the build writes into fleetd/target/fleetd.jar, the path the live daemon is still running from, and that is the one option that quietly loses in-flight work.

Why it loses work

A JVM resolves a class the first time it is actually used. A class only reachable on a failure path — a shutdown drain, an abandon, an error branch — has never been loaded while the daemon ran normally. When the shutdown hook finally walks that path, the class is resolved from whatever is at the jar path now, which is no longer the jar the process started with.

The thread that dies is the shutdown drain itself, so sessions are not released and waiting askers never get their rendezvous resolved.

And you cannot see it happen. The supervisor records status=143, which is indistinguishable from a clean SIGTERM stop — the script's own CB-594 comment already establishes that a SIGTERM'd JVM reports 143 even when its hook completes. The only evidence is a stack trace in the log.

Two hosts, two different classes, one mechanism

fleet01, measured and reported by its lead, 2026-09-10:

jar rewritten           02:10:41
old daemon pid 1610855  died 02:11:17
  NoClassDefFoundError: dev/ltms/fleet/msg/Rendezvous$Resolution
    at msg/Rendezvous.resolveFailure:231
    <- msg/MessageService.abandon:734
    <- session/SessionManager release/drainSnapshot/drainAll/close
  systemd recorded status=143

The Mac, measured by me today. I went looking because of their report:

grep -c NoClassDefFoundError fleetd/fleetd.out   -> 1
grep -c 'fleetd listening' fleetd/fleetd.out     -> 58     (control, must be non-zero)

line 49902, in the boot that begins at line 49890:
  Exception in thread "Thread-0" java.lang.NoClassDefFoundError: reactor/core/Exceptions
    at reactor.core.publisher.Operators.onOperatorError(Operators.java:756)

I cannot date that line: this log carries no per-line dates, so I am giving line numbers instead of a timestamp. A different class from fleet01's, on a different OS, in a shutdown thread — the same mechanism.

Two hosts agreeing would be one data point if they shared an instrument. They do not: different operating systems, different supervisors, different classes, and the two were found by different people looking for different things.

Today's redeploy on the Mac did not hit it — 0 matches between the current boot line and EOF. So this is intermittent, and its probability rises with the fraction of class files that change. On a 91-commit jump it stops being a risk and becomes near-certain, which is exactly the case I was advising on.

The fix

Small, and the choice should be deliberate:

  1. Build to a staging path and swap — build to target/fleetd-new.jar, stop the daemon, then move it into place. Keeps the "never leave the fleet down on a failed build" property, which is worth keeping, and removes the swap-under-a-live-process window.
  2. Or stop the daemon before the jar is replaced, and accept the downtime the current order exists to avoid.

Option 1 preserves both properties and is what I would do. Either way the script must stop claiming the current order is safe, because its header currently presents build-first as a protection with no cost.

A detection step is also cheap and worth having: after a restart, scan the shutdown window of the previous boot for NoClassDefFoundError and report it loudly. Right now nothing does, the exit code lies, and the failure is silent by construction.

Related

  • #413 — the mechanism (mvn clean deleting the running daemon's jar; the jar's build time, not HEAD, defines "the running tree"). This ticket is the same root cause reached by a different route: not deleting the jar, but rewriting it.
  • #492 — the other defect in this script, found in the same pass. Both touch scripts/redeploy-fleetd.sh, so they should land in order, not in parallel.
Raised by the fleet01 lead on 2026-09-12, against advice I gave them. They are right and I was wrong. Confirmed independently on the Mac. ## The advice that is wrong I told the fleet01 lead: *"91+ commits is a large jump. Build before you stop anything, so a failed build never leaves you down."* `scripts/redeploy-fleetd.sh` is built on that same order, and says so in its own header: it builds before it stops anything. The intent is good — a failed build should never leave the fleet down. But the build writes **into `fleetd/target/fleetd.jar`, the path the live daemon is still running from**, and that is the one option that quietly loses in-flight work. ## Why it loses work A JVM resolves a class the first time it is actually used. A class only reachable on a failure path — a shutdown drain, an abandon, an error branch — has never been loaded while the daemon ran normally. When the shutdown hook finally walks that path, the class is resolved **from whatever is at the jar path now**, which is no longer the jar the process started with. The thread that dies is the shutdown drain itself, so sessions are not released and waiting askers never get their rendezvous resolved. And you cannot see it happen. The supervisor records `status=143`, which is indistinguishable from a clean SIGTERM stop — the script's own CB-594 comment already establishes that a SIGTERM'd JVM reports 143 even when its hook completes. The only evidence is a stack trace in the log. ## Two hosts, two different classes, one mechanism **fleet01**, measured and reported by its lead, 2026-09-10: ``` jar rewritten 02:10:41 old daemon pid 1610855 died 02:11:17 NoClassDefFoundError: dev/ltms/fleet/msg/Rendezvous$Resolution at msg/Rendezvous.resolveFailure:231 <- msg/MessageService.abandon:734 <- session/SessionManager release/drainSnapshot/drainAll/close systemd recorded status=143 ``` **The Mac**, measured by me today. I went looking because of their report: ``` grep -c NoClassDefFoundError fleetd/fleetd.out -> 1 grep -c 'fleetd listening' fleetd/fleetd.out -> 58 (control, must be non-zero) line 49902, in the boot that begins at line 49890: Exception in thread "Thread-0" java.lang.NoClassDefFoundError: reactor/core/Exceptions at reactor.core.publisher.Operators.onOperatorError(Operators.java:756) ``` I cannot date that line: this log carries no per-line dates, so I am giving line numbers instead of a timestamp. A different class from fleet01's, on a different OS, in a shutdown thread — the same mechanism. Two hosts agreeing would be one data point if they shared an instrument. They do not: different operating systems, different supervisors, different classes, and the two were found by different people looking for different things. Today's redeploy on the Mac did **not** hit it — 0 matches between the current boot line and EOF. So this is intermittent, and its probability rises with the fraction of class files that change. On a 91-commit jump it stops being a risk and becomes near-certain, which is exactly the case I was advising on. ## The fix Small, and the choice should be deliberate: 1. **Build to a staging path and swap** — build to `target/fleetd-new.jar`, stop the daemon, then move it into place. Keeps the "never leave the fleet down on a failed build" property, which is worth keeping, and removes the swap-under-a-live-process window. 2. **Or stop the daemon before the jar is replaced**, and accept the downtime the current order exists to avoid. Option 1 preserves both properties and is what I would do. Either way the script must stop claiming the current order is safe, because its header currently presents build-first as a protection with no cost. A detection step is also cheap and worth having: after a restart, scan the shutdown window of the previous boot for `NoClassDefFoundError` and report it loudly. Right now nothing does, the exit code lies, and the failure is silent by construction. ## Related - #413 — the mechanism (`mvn clean` deleting the running daemon's jar; the jar's build time, not HEAD, defines "the running tree"). This ticket is the same root cause reached by a different route: not deleting the jar, but rewriting it. - #492 — the other defect in this script, found in the same pass. Both touch `scripts/redeploy-fleetd.sh`, so they should land in order, not in parallel.
Author
Owner

Fixed and proven on a live redeploy. Closing.

Merged as #510 (aa4c0b8). The script no longer builds into the path the running daemon holds: it stages to target/fleetd-new.jar right after a successful build, and swaps to target/fleetd.jar only after wait_for_daemon_exit confirms the old pid is gone. That is option 1 from the ticket, so both properties are kept — a failed build still never leaves the fleet down.

The acceptance test was the real redeploy, and I ran it

The worker could not prove the swap end-to-end, because that needs a live daemon and workers must never restart the one they talk through. They said so in their report rather than claiming it. I ran it myself on this host, just now:

== build
   Tests run: 1694, Failures: 0, Errors: 0, Skipped: 0
   ok    BUILD SUCCESS
   ok    jar now: 46368e1786d3          <- the STAGED jar
== stop
   supervision is ON (launchd): using 'launchctl unload' (not kill)
   ok    pid 61055 exited               <- old pid confirmed gone FIRST
== swap
   ok    jar in place: 46368e1786d3     <- same hash, now at the live path
== start
   ok    started, pid 35106
== verify
   ok    /healthz 200 — {"status":"ok","herdr":{"protocol":19,"version":"0.8.0"}}
   ok    10:47:00.306 INFO [main] dev.ltms.fleet.Fleetd - fleetd listening on 127.0.0.1:8765
== result
   ok    pid 35106, jar 46368e1786d3

The order in the output is the fix: jar now (staged) → pid exited → swap → start. The jar hash is identical on both sides, so the file that was built is the file that is running.

After it: ls fleetd/target/fleetd-new.jar → no such file (nothing left behind), one daemon under pgrep -f 'fleetd.jar', fleet_whoami still primary, and a real fleet_spawn succeeded, so the herdr protocol survived the restart.

The launchd race this could have opened, checked

Staging creates a new window: between the stop and the swap, the live path has no jar. If a supervisor restarted the daemon in that window it would find nothing. It cannot here — the launchd branch uses launchctl unload -w, which deregisters the job before the process stops, so no KeepAlive is left armed. The systemd branch uses systemctl --user stop, and Restart= does not fire on a deliberate stop. Both were already correct for the CB-594 reason and they cover this too.

I reviewed the diff and mutated the half the worker did not

Two survivors, both now filed as #511: the new drain-gate abort message tells the operator to "rerun (with or without --no-build)" and --no-build cannot finish that restart, which I reproduced with a control; and jar_id()'s no-argument default is pinned by no test, so the PR's own claim that --check cannot be fooled is unverified.

The detection half of this ticket is NOT done

This ticket asked for two things, and #510 did only the first. The second was:

after a restart, scan the shutdown window of the previous boot for NoClassDefFoundError and report it loudly

Filed as #512, with a measurement that makes it worse than "not done": the script's no ERROR lines since restart check is structurally blind to this failure, because an uncaught exception in a shutdown thread never goes through the logger and so never carries an ERROR token. On this host, the one real NoClassDefFoundError line matches grep -c 'ERROR' → 0, against a control of 1977 lines that do carry the token. So the script currently prints a reassuring line over exactly the failure this ticket is about.

Closing this one: the defect it names is fixed and proven. #512 carries the detection, #511 carries the two follow-up defects.

## Fixed and proven on a live redeploy. Closing. Merged as #510 (`aa4c0b8`). The script no longer builds into the path the running daemon holds: it stages to `target/fleetd-new.jar` right after a successful build, and swaps to `target/fleetd.jar` only after `wait_for_daemon_exit` confirms the old pid is gone. That is option 1 from the ticket, so both properties are kept — a failed build still never leaves the fleet down. ### The acceptance test was the real redeploy, and I ran it The worker could not prove the swap end-to-end, because that needs a live daemon and workers must never restart the one they talk through. They said so in their report rather than claiming it. I ran it myself on this host, just now: ``` == build Tests run: 1694, Failures: 0, Errors: 0, Skipped: 0 ok BUILD SUCCESS ok jar now: 46368e1786d3 <- the STAGED jar == stop supervision is ON (launchd): using 'launchctl unload' (not kill) ok pid 61055 exited <- old pid confirmed gone FIRST == swap ok jar in place: 46368e1786d3 <- same hash, now at the live path == start ok started, pid 35106 == verify ok /healthz 200 — {"status":"ok","herdr":{"protocol":19,"version":"0.8.0"}} ok 10:47:00.306 INFO [main] dev.ltms.fleet.Fleetd - fleetd listening on 127.0.0.1:8765 == result ok pid 35106, jar 46368e1786d3 ``` The order in the output is the fix: `jar now` (staged) → `pid exited` → `swap` → `start`. The jar hash is identical on both sides, so the file that was built is the file that is running. After it: `ls fleetd/target/fleetd-new.jar` → no such file (nothing left behind), one daemon under `pgrep -f 'fleetd.jar'`, `fleet_whoami` still `primary`, and a real `fleet_spawn` succeeded, so the herdr protocol survived the restart. ### The launchd race this could have opened, checked Staging creates a new window: between the stop and the swap, the live path has no jar. If a supervisor restarted the daemon in that window it would find nothing. It cannot here — the launchd branch uses `launchctl unload -w`, which deregisters the job before the process stops, so no `KeepAlive` is left armed. The systemd branch uses `systemctl --user stop`, and `Restart=` does not fire on a deliberate stop. Both were already correct for the CB-594 reason and they cover this too. ### I reviewed the diff and mutated the half the worker did not Two survivors, both now filed as **#511**: the new drain-gate abort message tells the operator to "rerun (with or without `--no-build`)" and `--no-build` cannot finish that restart, which I reproduced with a control; and `jar_id()`'s no-argument default is pinned by no test, so the PR's own claim that `--check` cannot be fooled is unverified. ### The detection half of this ticket is NOT done This ticket asked for two things, and #510 did only the first. The second was: > after a restart, scan the shutdown window of the previous boot for `NoClassDefFoundError` and report it loudly Filed as **#512**, with a measurement that makes it worse than "not done": the script's `no ERROR lines since restart` check is structurally blind to this failure, because an uncaught exception in a shutdown thread never goes through the logger and so never carries an `ERROR` token. On this host, the one real `NoClassDefFoundError` line matches `grep -c 'ERROR'` → 0, against a control of 1977 lines that do carry the token. So the script currently prints a reassuring line over exactly the failure this ticket is about. Closing this one: the defect it names is fixed and proven. #512 carries the detection, #511 carries the two follow-up defects.
ltms closed this issue 2026-09-12 05:53:55 +02:00
Sign in to join this conversation.
1 Participants
Notifications
Due Date
No due date set.
Dependencies

No dependencies set.

Reference: fleet/fleetd#493