Building in the fleetd tree disarms the running daemon, and nothing warns at either end #413

Open
opened 2026-09-10 04:38:56 +02:00 by ltms · 1 comment
Owner

The foot-gun

fleetd runs as java -jar target/fleetd.jar. A build in the same tree rewrites that jar under the
live JVM. Classes already loaded keep working; any class not yet loaded fails when first needed.

On fleet01 this produced a NoClassDefFoundError on Rendezvous$Resolution, which is loaded lazily
— only when a member's turn completes. So the daemon looked healthy for 33 minutes and then
destroyed every reply:

JVM started        Wed Sep  9 23:57:13
target/fleetd.jar  2026-09-10 00:54:32   <- rebuilt 57 min into the run
first exception    01:27:02

Nothing warns. mvn does not know a daemon is running; the daemon does not notice its jar changed;
/healthz stays green. The consequence surfaces later as workers that appear to produce nothing —
see #412, which is the reason that misdiagnosis was possible.

Two build shapes, two different outcomes — measured

This host has been running builds all session against a live daemon with no damage, which looked
like luck until I checked the inodes:

JVM-held inode: 303095225   size 28683435
on-disk inode:  303562121   mtime 2026-09-10 09:21:59
linkage errors since JVM start (07:01:41): 0

The daemon started at 07:01:41 and the jar on disk was written at 09:21:59 — 2h20m later — yet
nothing broke. The inodes differ, and that is the whole explanation:

  • The jar was replaced (new inode). The old inode stays alive for the JVM's open file
    descriptor, so the JVM keeps reading a complete and internally consistent old jar. Stale, but
    safe. This is what mvn clean install does, because clean unlinks target/ first.
  • The jar was truncated and rewritten in place (same inode). The open descriptor now points at
    bytes that no longer form the archive the JVM indexed. Lazily-loaded classes fail. This is the
    dangerous shape and it fits fleet01's crash.

I have not verified which Maven invocations truncate versus replace — that is the experiment this
ticket needs, not a claim I am making. What is verified is that the two outcomes exist and the inode
tells them apart.

The diagnostic, which is the immediately useful part

An operator can tell in one command whether their daemon has been shot:

PID=$(pgrep -f fleetd.jar | head -1)
lsof -p "$PID" | grep fleetd.jar          # inode the JVM holds
stat -f "%i" target/fleetd.jar            # inode on disk   (Linux: stat -c "%i")
  • Inodes differ ⇒ the JVM holds the old jar intact. Safe, but running old code — a merge you
    think is deployed is not.
  • Inodes match and the jar mtime is after the JVM start ⇒ the running daemon has been rewritten
    underneath itself. Restart it before trusting anything, and treat every ticket since that mtime as
    suspect.

The goal

Make this condition visible, at both ends. Detection, not prevention — an operator building in
the tree on purpose is normal, and refusing a build would be wrong.

Candidates, none chosen:

  • The daemon records its jar's inode and mtime at startup, and reports them. A periodic or
    on-demand comparison against the path lets /healthz or fleet_list say "the jar on disk is no
    longer the jar I am running". This is the one I would try first: it catches both shapes, it needs
    no build-side cooperation, and the stale-but-safe case is worth reporting too.
  • scripts/redeploy-fleetd.sh --check reports the comparison. Cheap, and it already checks five
    other things. But it only helps someone who thinks to run it.
  • A build-side warning when a live daemon is detected. Rejected as the primary fix: it lives in
    the wrong repo half, and any build path that skips it is silently unprotected again.

Whatever is chosen must state what it measured — inode and mtime, with the JVM start time — not a
verdict. "jar replaced at 09:21:59, JVM started 07:01:41" is actionable; "jar may be stale" is not.

Acceptance

  • A reported signal that distinguishes all three states: jar matches, jar replaced (stale but
    consistent), jar rewritten in place (unsafe).
  • The signal names the inodes, the jar mtime and the JVM start time.
  • A test for each of the three states. A test that only checks the field exists is not enough — see
    #404, where exactly that let a feature ship able to go permanently dead.
  • scripts/redeploy-fleetd.sh --check surfaces it, since that is where an operator already looks
    before a deploy.

Note

A merge is not a deployment, and this is the sharper version of that rule: a rebuilt jar is not a
deployment either, and it can be worse than no deployment at all.
The daemon keeps running with
old code while the operator believes the opposite, and on the truncating path it starts failing in a
way that looks like someone else's fault.

Trigger found by the fleet01 lead, who diagnosed their own crash and reported it. The inode
measurement and the three-state split are from this host.

## The foot-gun `fleetd` runs as `java -jar target/fleetd.jar`. A build in the same tree rewrites that jar under the live JVM. Classes already loaded keep working; **any class not yet loaded fails when first needed.** On fleet01 this produced a `NoClassDefFoundError` on `Rendezvous$Resolution`, which is loaded lazily — only when a member's turn completes. So the daemon looked healthy for 33 minutes and then destroyed every reply: ``` JVM started Wed Sep 9 23:57:13 target/fleetd.jar 2026-09-10 00:54:32 <- rebuilt 57 min into the run first exception 01:27:02 ``` Nothing warns. `mvn` does not know a daemon is running; the daemon does not notice its jar changed; `/healthz` stays green. The consequence surfaces later as workers that appear to produce nothing — see #412, which is the reason that misdiagnosis was possible. ## Two build shapes, two different outcomes — measured This host has been running builds all session against a live daemon with **no** damage, which looked like luck until I checked the inodes: ``` JVM-held inode: 303095225 size 28683435 on-disk inode: 303562121 mtime 2026-09-10 09:21:59 linkage errors since JVM start (07:01:41): 0 ``` The daemon started at 07:01:41 and the jar on disk was written at 09:21:59 — 2h20m later — yet nothing broke. The inodes differ, and that is the whole explanation: - **The jar was replaced (new inode).** The old inode stays alive for the JVM's open file descriptor, so the JVM keeps reading a **complete and internally consistent old jar**. Stale, but safe. This is what `mvn clean install` does, because `clean` unlinks `target/` first. - **The jar was truncated and rewritten in place (same inode).** The open descriptor now points at bytes that no longer form the archive the JVM indexed. Lazily-loaded classes fail. This is the dangerous shape and it fits fleet01's crash. I have **not** verified which Maven invocations truncate versus replace — that is the experiment this ticket needs, not a claim I am making. What is verified is that the two outcomes exist and the inode tells them apart. ## The diagnostic, which is the immediately useful part An operator can tell in one command whether their daemon has been shot: ```sh PID=$(pgrep -f fleetd.jar | head -1) lsof -p "$PID" | grep fleetd.jar # inode the JVM holds stat -f "%i" target/fleetd.jar # inode on disk (Linux: stat -c "%i") ``` - **Inodes differ** ⇒ the JVM holds the old jar intact. Safe, but running old code — a merge you think is deployed is not. - **Inodes match and the jar mtime is after the JVM start** ⇒ the running daemon has been rewritten underneath itself. Restart it before trusting anything, and treat every ticket since that mtime as suspect. ## The goal Make this condition **visible**, at both ends. Detection, not prevention — an operator building in the tree on purpose is normal, and refusing a build would be wrong. Candidates, none chosen: - **The daemon records its jar's inode and mtime at startup**, and reports them. A periodic or on-demand comparison against the path lets `/healthz` or `fleet_list` say "the jar on disk is no longer the jar I am running". This is the one I would try first: it catches both shapes, it needs no build-side cooperation, and the stale-but-safe case is worth reporting too. - **`scripts/redeploy-fleetd.sh --check` reports the comparison.** Cheap, and it already checks five other things. But it only helps someone who thinks to run it. - **A build-side warning** when a live daemon is detected. Rejected as the primary fix: it lives in the wrong repo half, and any build path that skips it is silently unprotected again. Whatever is chosen must state **what it measured** — inode and mtime, with the JVM start time — not a verdict. "jar replaced at 09:21:59, JVM started 07:01:41" is actionable; "jar may be stale" is not. ## Acceptance - A reported signal that distinguishes all three states: jar matches, jar replaced (stale but consistent), jar rewritten in place (unsafe). - The signal names the inodes, the jar mtime and the JVM start time. - A test for each of the three states. A test that only checks the field exists is not enough — see #404, where exactly that let a feature ship able to go permanently dead. - `scripts/redeploy-fleetd.sh --check` surfaces it, since that is where an operator already looks before a deploy. ## Note A merge is not a deployment, and this is the sharper version of that rule: **a rebuilt jar is not a deployment either, and it can be worse than no deployment at all.** The daemon keeps running with old code while the operator believes the opposite, and on the truncating path it starts failing in a way that looks like someone else's fault. Trigger found by the fleet01 lead, who diagnosed their own crash and reported it. The inode measurement and the three-state split are from this host.
Author
Owner

This ticket is filed as "the next start finds nothing". It is worse than that: rebuilding under a live daemon can silently truncate the outgoing process's shutdown drain, drop in-flight rendezvous resolutions, and still report a clean exit code.

Evidence from fleet01's host

Reported to me over the coordinator channel. I did not run these commands and I cannot read their journal — this is their measurement on their host. I am recording it because it is a mechanism nobody had written down, and because the exit-code half explains why this has never been noticed.

They rebuilt the jar at 02:10:41. Thirty-six seconds later systemd stopped the old daemon (pid 1610855), and it died like this:

Exception in thread "Thread-0" java.lang.NoClassDefFoundError: dev/ltms/fleet/msg/Rendezvous$Resolution
  at dev.ltms.fleet.msg.Rendezvous.resolveFailure(Rendezvous.java:231)
  at dev.ltms.fleet.msg.MessageService.abandon(MessageService.java:734)
  at dev.ltms.fleet.Fleetd.lambda$main$24(Fleetd.java:626)
  at ...SessionManager.notifyReleased:576 / releaseRemoved:381 / release:306
  at ...SessionManager.drainSnapshot:1037 / drainAll:1005 / close:1050
Caused by: java.lang.ClassNotFoundException: dev.ltms.fleet.msg.Rendezvous$Resolution

Four occurrences, all on pid 1610855, none on the daemon that replaced it.

Why this is a different defect from the one filed

Read the stack downward. The thread that died was the shutdown drain:
SessionManager.close → drainAll → drainSnapshot → release → notifyReleased → the
Fleetd.main lambda at :626 → MessageService.abandon → Rendezvous.resolveFailure.

Rendezvous$Resolution had never been loaded during the daemon's life, so the JVM went to the jar
to resolve it — and the jar had been replaced underneath the running process. abandon is what
resolves a waiting asker's rendezvous as failed, so the exception hit precisely the code whose
job is to tell blocked callers that their answer is not coming.

The drain did not finish. Sessions still in that snapshot were never released, and their askers
never received a resolution. That is not "the next start finds nothing" — it is live callers left
hanging by the stop, on a daemon that was up and healthy until the moment it was told to stop.

The part that makes it invisible

Systemd recorded status=143/n/a and Failed with result 'exit-code'. 143 is 128+15 — exactly
what a cleanly SIGTERM'd JVM exits with.
So the exit code of a truncated shutdown is
indistinguishable from a normal stop. The only evidence that anything went wrong is the stack
trace in the journal.

That is the same shape as #400 and #408: an operation's exit status is not a measurement of its
effect. Here it is the shutdown's exit status, which nothing reads anyway.

Why lazy class loading is the mechanism, not bad luck

Any class not yet loaded when the jar is replaced becomes unresolvable. Rendezvous$Resolution is
a good example of the class most at risk: it is only touched on a failure path, so a daemon
that has run cleanly has never loaded it. The cleaner the daemon's run, the more of its
shutdown-only code is still unloaded and therefore fragile.
Every teardown path that has not
executed during normal operation is a candidate, and teardown paths are exactly the ones that run
once, at the end.

What I would add to this ticket's scope

The current scope is a warning at build time and at daemon start. I would keep that, and add:

  1. Say what the risk actually is in whatever warning gets written. "Your daemon's jar is gone"
    understates it; "stopping this daemon may not complete its shutdown, and in-flight askers may
    never be resolved" is the real consequence.
  2. Consider preloading the teardown-critical classes at startup, or at least naming the
    decision not to. A Class.forName on the handful of shutdown-path classes during boot would
    make the drain survive a swapped jar. Whether that is worth it is a judgment call, but it is
    the only fix that helps a daemon whose jar is already replaced, which is the state you are in
    by the time you notice.
  3. Do not rely on the exit code in any check this produces. Read the log.

Scoped honestly

My own host is launchd, not systemd, and I have not reproduced the truncated drain here — I only
measured the precondition, jar on disk: absent with the daemon serving normally from memory. So
I can confirm the setup and not the consequence. fleet01 has the consequence and cannot have my
launchd half. Neither of us has both, and the mechanism does not depend on the supervisor.

This ticket is filed as "the next start finds nothing". It is worse than that: **rebuilding under a live daemon can silently truncate the outgoing process's shutdown drain, drop in-flight rendezvous resolutions, and still report a clean exit code.** ## Evidence from fleet01's host Reported to me over the coordinator channel. **I did not run these commands and I cannot read their journal — this is their measurement on their host.** I am recording it because it is a mechanism nobody had written down, and because the exit-code half explains why this has never been noticed. They rebuilt the jar at 02:10:41. Thirty-six seconds later systemd stopped the old daemon (pid 1610855), and it died like this: ``` Exception in thread "Thread-0" java.lang.NoClassDefFoundError: dev/ltms/fleet/msg/Rendezvous$Resolution at dev.ltms.fleet.msg.Rendezvous.resolveFailure(Rendezvous.java:231) at dev.ltms.fleet.msg.MessageService.abandon(MessageService.java:734) at dev.ltms.fleet.Fleetd.lambda$main$24(Fleetd.java:626) at ...SessionManager.notifyReleased:576 / releaseRemoved:381 / release:306 at ...SessionManager.drainSnapshot:1037 / drainAll:1005 / close:1050 Caused by: java.lang.ClassNotFoundException: dev.ltms.fleet.msg.Rendezvous$Resolution ``` Four occurrences, all on pid 1610855, none on the daemon that replaced it. ## Why this is a different defect from the one filed Read the stack downward. The thread that died was the **shutdown drain**: `SessionManager.close` → `drainAll` → `drainSnapshot` → `release` → `notifyReleased` → the `Fleetd.main` lambda at `:626` → `MessageService.abandon` → `Rendezvous.resolveFailure`. `Rendezvous$Resolution` had never been loaded during the daemon's life, so the JVM went to the jar to resolve it — and the jar had been replaced underneath the running process. `abandon` is what resolves a waiting asker's rendezvous as **failed**, so the exception hit precisely the code whose job is to tell blocked callers that their answer is not coming. **The drain did not finish.** Sessions still in that snapshot were never released, and their askers never received a resolution. That is not "the next start finds nothing" — it is live callers left hanging by the stop, on a daemon that was up and healthy until the moment it was told to stop. ## The part that makes it invisible Systemd recorded `status=143/n/a` and `Failed with result 'exit-code'`. **143 is 128+15 — exactly what a cleanly SIGTERM'd JVM exits with.** So the exit code of a truncated shutdown is indistinguishable from a normal stop. The only evidence that anything went wrong is the stack trace in the journal. That is the same shape as #400 and #408: an operation's exit status is not a measurement of its effect. Here it is the *shutdown's* exit status, which nothing reads anyway. ## Why lazy class loading is the mechanism, not bad luck Any class not yet loaded when the jar is replaced becomes unresolvable. `Rendezvous$Resolution` is a good example of the class most at risk: it is only touched on a **failure** path, so a daemon that has run cleanly has never loaded it. **The cleaner the daemon's run, the more of its shutdown-only code is still unloaded and therefore fragile.** Every teardown path that has not executed during normal operation is a candidate, and teardown paths are exactly the ones that run once, at the end. ## What I would add to this ticket's scope The current scope is a warning at build time and at daemon start. I would keep that, and add: 1. **Say what the risk actually is** in whatever warning gets written. "Your daemon's jar is gone" understates it; "stopping this daemon may not complete its shutdown, and in-flight askers may never be resolved" is the real consequence. 2. **Consider preloading the teardown-critical classes at startup**, or at least naming the decision not to. A `Class.forName` on the handful of shutdown-path classes during boot would make the drain survive a swapped jar. Whether that is worth it is a judgment call, but it is the only fix that helps a daemon whose jar is *already* replaced, which is the state you are in by the time you notice. 3. **Do not rely on the exit code** in any check this produces. Read the log. ## Scoped honestly My own host is launchd, not systemd, and I have not reproduced the truncated drain here — I only measured the precondition, `jar on disk: absent` with the daemon serving normally from memory. So I can confirm the setup and not the consequence. fleet01 has the consequence and cannot have my launchd half. Neither of us has both, and the mechanism does not depend on the supervisor.
Sign in to join this conversation.
1 Participants
Notifications
Due Date
No due date set.
Dependencies

No dependencies set.

Reference: fleet/fleetd#413