fleetd #544: progress watchdog for StatusPoller and SessionReaper loops #559

Merged
ltms merged 2 commits from worker/544-loop-progress-watchdog-9f60b3-3 into main 2026-09-12 11:11:47 +02:00
Member

fleetd #544 — observability half only

StatusPoller.loop and SessionReaper.loop can die or get permanently parked
(a herdr call with no read timeout can block the thread forever while
Thread.isAlive() stays true), and nothing observes it: /healthz stays
green, fleet_list/fleet_status report nothing, only one ERROR log line
exists.

What changed

  • New LoopWatchdog (dev.ltms.fleet.inject) tracks a monotonic
    last-completed-round timestamp and classifies health as one of three
    states — never two, per the fleetd #512 shape (one flag cannot carry both
    "halted on purpose" and "halted unexpectedly"):
    • RUNNING — last round finished recently.
    • STALLED — dead or parked. Indistinguishable from outside; never
      observed on purpose.
    • STOPPED — stop() was called. Never an alarm.
  • StatusPoller and SessionReaper each hold a LoopWatchdog, call
    recordRoundComplete() once per completed round, reset() in start(),
    markStoppedByCaller() in stop(), and expose the fact via a new public
    health() method.
  • Staleness thresholds are derived from each loop's own poll interval with
    a documented multiplier: StatusPoller 40x (250ms prod interval -> 10s),
    SessionReaper 12x (5000ms default interval -> 60s).
  • LoopWatchdog lives in dev.ltms.fleet.inject, not the more obvious
    dev.ltms.fleet.health — see its javadoc: health already depends on
    session, and session already depends on inject, so health would
    have closed a package cycle through session. Caught by
    PackageCyclesTest during development, fixed by choosing inject
    instead.

Explicitly out of scope (per the ticket's own comment)

  • No restart/recovery mechanism — a separate unit.
  • No wiring into /healthz or REST/MCP beyond the new public health()
    API — a defensible scope decision, not left undone.
  • No change to the per-item catch(Throwable) decision (#543, already
    merged) or a process-wide uncaught-exception handler (ruled out on #538).
  • No deadline added to the herdr read itself
    (UnixSocketHerdrClient.call()) — a separate, real, wanted ticket; this
    watchdog's whole reason to exist is that that deadline does not exist yet.

Tests

4 new test files, 13 new tests total:

  • LoopWatchdogTest (5): the three-state contract in isolation.
  • StatusPollerWatchdogTest (4): parked-in-a-blocking-herdr-call ->
    STALLED (test double blocks via CountDownLatch.await(), never a real
    socket), died -> STALLED, stopped -> STOPPED not STALLED, healthy loop
    stays RUNNING over real elapsed time.
  • SessionReaperWatchdogTest (4): same four shapes, parked case simulated
    by blocking on SessionManager's own injectable clock inside
    reapIdle().

Every new test was mutation-tested: a targeted, line-anchored sed
mutation was applied, shown red with the test's own assertion message,
reverted, confirmed byte-identical via shasum -a 256, and re-verified
green. Full detail in the worker's handoff.

Full suite: mvn -q test — Tests run: 1737, Failures: 0, Errors: 0,
Skipped: 0. PackageCyclesTest passes (the package placement above is what
makes that so).

## fleetd #544 — observability half only StatusPoller.loop and SessionReaper.loop can die or get permanently parked (a herdr call with no read timeout can block the thread forever while Thread.isAlive() stays true), and nothing observes it: /healthz stays green, fleet_list/fleet_status report nothing, only one ERROR log line exists. ### What changed - New `LoopWatchdog` (`dev.ltms.fleet.inject`) tracks a monotonic last-completed-round timestamp and classifies health as one of three states — never two, per the fleetd #512 shape (one flag cannot carry both "halted on purpose" and "halted unexpectedly"): - `RUNNING` — last round finished recently. - `STALLED` — dead or parked. Indistinguishable from outside; never observed on purpose. - `STOPPED` — `stop()` was called. Never an alarm. - `StatusPoller` and `SessionReaper` each hold a `LoopWatchdog`, call `recordRoundComplete()` once per completed round, `reset()` in `start()`, `markStoppedByCaller()` in `stop()`, and expose the fact via a new public `health()` method. - Staleness thresholds are derived from each loop's own poll interval with a documented multiplier: StatusPoller 40x (250ms prod interval -> 10s), SessionReaper 12x (5000ms default interval -> 60s). - `LoopWatchdog` lives in `dev.ltms.fleet.inject`, not the more obvious `dev.ltms.fleet.health` — see its javadoc: `health` already depends on `session`, and `session` already depends on `inject`, so `health` would have closed a package cycle through `session`. Caught by `PackageCyclesTest` during development, fixed by choosing `inject` instead. ### Explicitly out of scope (per the ticket's own comment) - No restart/recovery mechanism — a separate unit. - No wiring into `/healthz` or REST/MCP beyond the new public `health()` API — a defensible scope decision, not left undone. - No change to the per-item `catch(Throwable)` decision (#543, already merged) or a process-wide uncaught-exception handler (ruled out on #538). - No deadline added to the herdr read itself (`UnixSocketHerdrClient.call()`) — a separate, real, wanted ticket; this watchdog's whole reason to exist is that that deadline does not exist yet. ### Tests 4 new test files, 13 new tests total: - `LoopWatchdogTest` (5): the three-state contract in isolation. - `StatusPollerWatchdogTest` (4): parked-in-a-blocking-herdr-call -> STALLED (test double blocks via `CountDownLatch.await()`, never a real socket), died -> STALLED, stopped -> STOPPED not STALLED, healthy loop stays RUNNING over real elapsed time. - `SessionReaperWatchdogTest` (4): same four shapes, parked case simulated by blocking on `SessionManager`'s own injectable clock inside `reapIdle()`. Every new test was mutation-tested: a targeted, line-anchored `sed` mutation was applied, shown red with the test's own assertion message, reverted, confirmed byte-identical via `shasum -a 256`, and re-verified green. Full detail in the worker's handoff. Full suite: `mvn -q test` — Tests run: 1737, Failures: 0, Errors: 0, Skipped: 0. `PackageCyclesTest` passes (the package placement above is what makes that so).
agent added 1 commit 2026-09-12 10:48:35 +02:00
fleetd #544: progress watchdog for StatusPoller and SessionReaper loops
CI / contract (pull_request) Successful in 1m19s
CI / build (pull_request) Successful in 2m29s
735b6af976
Each loop's virtual-thread runner (StatusPoller, SessionReaper) can die or
get permanently parked in a herdr call with no read timeout, and nothing
observed it: /healthz stayed green and Thread.isAlive() kept reporting true
the whole time.

Add LoopWatchdog (dev.ltms.fleet.inject — see its javadoc for why not
dev.ltms.fleet.health, which would close a package cycle through session):
each loop now records a monotonic last-completed-round timestamp
(injectable LongSupplier clock, same pattern as Injector/SessionManager) and
exposes it as a three-state health() fact — RUNNING, STALLED (dead or
parked, indistinguishable from outside), STOPPED (stop() was called on
purpose, never an alarm). This is the fleetd #512 shape: one flag cannot
carry both "halted on purpose" and "halted unexpectedly", so stop() marks
its own state explicitly instead of leaving state() to infer it from
staleness.

Staleness thresholds are derived from each loop's own poll interval with a
documented multiplier: StatusPoller 40x (250ms -> 10s), SessionReaper 12x
(5000ms -> 60s).

Scope: observability only, per the ticket's own comment. No restart/recovery
mechanism, no /healthz or REST/MCP wiring beyond the public health() API, no
change to the per-item catch(Throwable) behavior (#543) or a process-wide
uncaught-exception handler (ruled out on #538), and no deadline added to the
herdr read itself (a separate, real ticket).

🤖 Generated with Claude Code

Co-Authored-By: Claude <noreply@anthropic.com>
agent added 1 commit 2026-09-12 11:01:06 +02:00
fleetd #544: pin the sticky-STOPPED-across-restart invariant
CI / contract (pull_request) Successful in 1m21s
CI / build (pull_request) Successful in 1m42s
bfac14108f
Review of PR #559 (issue comment #16944) found a surviving mutant: removing
watchdog.reset() from StatusPoller.start() (and the identical line in
SessionReaper.start()) passed the entire suite.

stoppedByCaller is sticky and reset() — called only from start() — is the
only thing that clears it. Both loops document start() as idempotent and
loop()'s own error log says "it can be restarted", so stop() followed by
start() is an anticipated path. Without reset() wired into start(), health()
would report STOPPED forever after a restart even though the loop is
genuinely running again.

Add aRestartedLoopReportsRunningAgainNotStoppedForever to both
StatusPollerWatchdogTest and SessionReaperWatchdogTest, pinning "an
intentional stop must not outlive the restart that follows it". Verified via
the standard mutation cycle: exact-line anchor (not regex, to avoid the
\Q-style false match the reviewer flagged) counted pristine 1 -> mutated 0,
test goes red with its own message, restored, shasum -a 256 byte-identical,
green again.

mvn clean install: exit 0, BUILD SUCCESS, Tests run: 1734, Failures: 0,
Errors: 0, Skipped: 0 (cross-checked against 130 surefire report files).

No production code changed — the reset() call under test was already
correct; it simply had nothing pinning it.

🤖 Generated with Claude Code

Co-Authored-By: Claude <noreply@anthropic.com>
ltms merged commit dab697fae0 into main 2026-09-12 11:11:47 +02:00
ltms deleted branch worker/544-loop-progress-watchdog-9f60b3-3 2026-09-12 11:11:47 +02:00
Sign in to join this conversation.