A dead StatusPoller or SessionReaper can now be restarted, but nothing restarts it and nothing notices #544
Closed
opened 2026-09-12 09:03:48 +02:00 by ltms
·
5 comments
No Branch/Tag Specified
main
worker/fleetd-612-unita-87807e-1
worker/612-b3-mcpwirings-da2b58-3
worker/612-b2-cb185-176d3a-2
worker/612-b1-completion-457459-1
worker/612-agaps-73a926-2
worker/608-sleeps-3a64ff-3
worker/621-b4520b-1
worker/618-b83894-2
worker/fleetd-615-e05481-5
worker/lead-autocompact-5f1ab2-3
worker/fleetd-613-f85deb-3
worker/fleetd-608-flaky-nudge-test-d0c2d1-3
worker/lead-context-gauge-ad404f-1
worker/gauge-wiring-9158c1-4
worker/redeploy-slowstart-ead0e5-5
worker/charter-bytes-13668c-6
worker/rollover-outcome-291483-2
worker/589-f64303-2
worker/593-1a8025-5
worker/589-fcd2aa-1
worker/568-9fdaa2-3
worker/571-attempted-outcome-5739f7-2
worker/581-completionresolver-cas-sites-0542b7-6
worker/562-loop-health-wiring-test-99611c-5
worker/562-surface-loop-health-7df5cc-4
worker/575-waiter-cleanup-sites-62ad80-1
worker/572-answer-lock-release-46a9ae-5
worker/567-probe-channel-leak-a38fc5-6
worker/551-record-before-send-7cbf56-1
worker/561-listener-fanout-survives-a-throw-61d538-2
worker/555-redeploy-main-flow-seam-65c2f5-2
worker/556-injector-owns-registration-e027a5-1
worker/552-post-restart-mktemp-abort-bc2672-4
worker/553-onstatus-completion-leak-0da881-2
worker/550-shasum-linux-196132-1
worker/538-loop-dies-on-error-4a5eeb-6
worker/426-health-coverage-ef1fd4-4
worker/504-failed-reported-clean-3cfd66-3
worker/537-capturedlog-close-e4c437-2
worker/459-broken-link-targets-cadc17-5
worker/535-appender-leak-fe74c1-1
worker/512-part2-shutdown-detection-434701-9
worker/529-logger-level-sweep-2a5533-8
worker/528-drain-gate-call-site-5de83d-7
charter/forge-mcp-vs-token
worker/521-swap-guard-unpinned-28e931-5
worker/519-probe-test-harness-d25ab8-4
worker/525-logger-level-leak-1b4eb0-6
worker/518-fleetmcp-resolver-wiring-8ef96c-1
worker/512-drain-complete-line-7edd71-3
worker/517-abort-branch-and-jar-id-41b641-2
worker/500-9e52c9-3
worker/509-4912f4-2
worker/511-9a4b23-1
worker/493-479f45-2
worker/505-03f8b2-1
worker/492-followup-detect-unclear
worker/501-a31fa0-7
worker/498-451d1c-5
worker/494-1015ce-2
worker/492-209647-1
worker/489-001902-2
worker/480-relative-handover-path-906323-1
worker/480-b-handover-skill-45bf1f-5
worker/474-followup-source-pin-f54a55-17
worker/474-charter-check-on-reload-f54a55-17
worker/466-quarantine-repeatcount-report
worker/393-opencode-skill-seeding-71854b-13
worker/469-canonical-tool-names-2a472a-16
worker/466-quarantine-escalation-5ae9c1-15
worker/446-hot-exhausted-pattern-0af580-6
worker/464-charter-tool-name-guard-a85635-12
worker/463-listfleet-default-fails-open-f1c76c-11
worker/458-invariant-5-by-purpose-862f9a-10
worker/439-coordinator-row-gate-bc032a-8
worker/449-herdr-protocol-576015-4
worker/450-abstract-spawn-599e1c-5
worker/437-ack-refuses-177d91-1
worker/444-placement-window-feb56a-2
worker/440-helddurable-derived-d462d7-13
worker/425-rework-placement-resolve-c58ba1-9
worker/421-lead-peek-held-msgs-cdbad2-10
worker/435-fixed-policy-cap-fe11de-12
worker/422-gate-state-observability-9e79d6-11
worker/431-memberregistry-live-readers-cdbad2-10
worker/424-architect-slot-hot-038b41-7
worker/422-model-gate-spawn-c29f48-6
worker/425-default-profile-live-f55534-8
worker/415-coverage-wording-2cbf9c-5
worker/416-3ad1da-1
worker/418-588283-3
worker/deterministic-stamp-race-409-3cb7b6-10
worker/armed-reads-live-config-404-ed931f-9
worker/reply-peer-refusal-391-5a34bd-7
worker/models-allowlist-aa9e9b-3
worker/ttl-stamp-race-399-f1122f-8
worker/scrub-receipt-400-316b3e-5
worker/exhaustion-detection-395-105105-6
worker/scrub-abort-394-316b3e-5
fix/scrub-uid-abort
worker/task-scrub-517574-2
worker/t386-clock-bd5b78-4
worker/t384-scrub-813790-5
worker/t381-cc-748314-2
worker/t373-336973-2
worker/t365-3920c5-3
worker/t358-6e989b-1
worker/t355-8b321c-1
worker/fleetd-369-hermetic-git-tests-e8b19a-3
worker/fleetd-368-stale-lead-binding-f5682e-2
worker/fleetd-360-deploy-units-0d3793-1
worker/359-dead-lead-tabs-f1253b-4
worker/362-worktree-skills-c03e51-3
worker/361-coord-visibility-655144-1
362-plugin-visibility-and-drift
worker/errscan-bed2ca-2
worker/amqp-log-identity-bed2ca-2
worker/withdefaults-guard-561704
worker/sleepguard-82076d-1
worker/fd334-9ee1b6-5
worker/fd348-f1ab27-4
worker/fd335-a71c35-1
worker/fd342-174a17-2
worker/fd345-490d0f-3
worker/fleetd-337-5ec7d4-21
worker/fleetd-341-af5a6b-24
worker/fleetd-339-5ca0a2-23
worker/fleetd-338-83a4a1-22
worker/fleetd-333-281f46-18
worker/fleetd-329-11bdbb-16
worker/fleetd-330-2770fb-17
worker/fix-326-50506e-15
worker/fix-324-3e9bbf-14
worker/fix-323-b8287d-13
worker/fix-316b-bd0860-11
worker/fix-318-76ca36-9
worker/fix-317-486aec-8
worker/fix-315-ce47c5-6
worker/fix-307-275890-6
worker/fix-308-b4f664-7
worker/fix-309-ec3939-8
worker/fix-310-7a3974-9
worker/fix-302-52ad0e-9
worker/fix-298-ce1acb-8
worker/fix-297-66bd11-7
worker/fix-296-104622-6
worker/fix-293-bare-closetab-eb22b5-3
worker/fix-280-gone-ask-lapse-bca98e-2
worker/fix-290-reapidle-guard-coverage-9b0dd1-1
worker/fix-285-trust-seed-8f3565-10
worker/fix-284-backend-error-seat-85912c-11
worker/fix-282-chained-ask-e6d0bb-8
worker/fix-283-teardown-leaks-f40dfa-9
worker/fix-281-pin-handler-actions-4921ac-7
worker/audit-rendezvous-lifecycle-d072ae-2
worker/audit-health-placement-1a2476-6
worker/audit-teardown-exits-e207a5-3
worker/audit-launcher-asymmetry-27e370-4
worker/audit-rest-authz-6ca53c-5
worker/investigate-275-abandon-asking-fdef52-8
worker/fix-274-worktree-leak-b0095d-7
worker/fix-273-exhausted-pattern-9665b5-6
worker/fleetd-267-model-check-bd8068-1
worker/fleetd-131-archunit-18b834-7
worker/fleetd-266-sshagent-rename-a014ff-6
worker/fleetd-184-uid-claim-8e1f31-4
worker/fleetd-184-warn-b381ee-10
worker/fleetd-184-docs-be1d12-9
worker/fleetd-257-9bf010-7
worker/fleetd-103-23a113-6
worker/fleetd-247-342356-5
worker/fleetd-116-04dea8-4
worker/fleetd-252-a830e0-3
worker/fleetd-111-7e8673-9
worker/fleetd-155c-f8ef4b-8
worker/fleetd-176-b928ca-3
worker/fleetd-249-7a7878-2
worker/cb248-composition-root-b-9acdf7-15
worker/cb148-envrc-default-fa6c82-12
worker/cb201-unit5-wiring-6c12e6-8
worker/cb241-fallback-echo-1175e9-11
worker/cb149-trust-dialog-2392a5-9
worker/cb134-148-overlay-visible-c9b986-10
worker/cb234-session-id-keyed-04e1fc-1
worker/cb201-unit3-nudge-abdf5c-6
worker/cb201-unit2-policy-c1102c-5
worker/cb201-unit4-outcome-a13bfa-7
worker/cb201-unit1-classifier-91b9b1-4
worker/cb201-227-refine-831980-3
worker/cb175-model-readback-0f085f-1
worker/cb222-charter-tmpdir-17f013-1
worker/cb226-architect-slot-race-cd3aa8-3
worker/cb224-worktree-root-group-024523-2
worker/cb-123-role-demotion-c600f7-2
worker/cb-219-opencode-roots-1f677e-1
worker/cb214-claude-session-id-b9eab4-4
worker/cb213-zdotdir-wrong-process-dd6de4-3
worker/cb211-exhaustion-classification-9546e0-2
worker/cb137-ambiguous-task-4df3d8-4
worker/cb209-agentsessionid-4dfdb6-2
worker/cb185-hostenvnames-2692b5-3
worker/cb206-opencode-sqlite-128718-2
worker/cb185-worktree-group-fc0c99-1
worker/cb-137-ask-ticket-e7760c-2
worker/cb-172-broker-uri-d36ae4-4
worker/cb-175-model-readback-76ead6-3
worker/cb-161-pane-ancestry-293510-1
worker/cb-164-rebase-885863-8
worker/cb-164-empty-scrape-false-success-1a80af-3
fix/cb-197-ticket-ttl-from-completion
worker/cb-189-remote-url-coverage-4692f3-1
worker/cb-185-blockers-027756-4
worker/cb-192-gap-log-11b631-2
worker/cb-633-fix-5f4396-3
worker/cb185-router-d6436d-3
worker/cb185-router-routing-gaps-9e9d33-3
worker/cb185-paneids-992586-2
worker/cb-633-allow-list-union-ed374b-1
worker/cb-157-credential-in-remote-url-496e44-2
worker/cb-641-health-herdr-evidence-8f1f54-6
worker/cb-640-health-msg-evidence-99c9cd-1
worker/cb-642-fleets-status-skill-bbbc40-5
cb-634-ide-mcp
worker/lead-comms-wiring-c014b9-7
worker/lead-mailbox-c19577-6
worker/autocompact-window-82bc2f-5
worker/cb-634-probe-18056f-4
worker/cb635-broker-urienv
worker/cb-632-config-retry-8e0efa-7
lead/cb-622e-claude-md
lead/cb-622-followup
worker/cb-622a-165dff-1
lead/cb-622d-opencode-mount
worker/cb-622b-717c67-2
worker/cb-622c-ab7759-3
worker/cb-617b2-20ca4b-3
worker/cb-617a-5c2f4a-1
worker/cb596-4e49ef-3
worker/cb586-10500c-1
worker/cb-606-b9343a-25
worker/cb604-1445f8-24
worker/cb582-477374-21
worker/cb584-8c2281-22
worker/cb600-e6b9a9-20
worker/cb602-ce257f-19
worker/cb601-b42837-18
worker/cb598-6c7ba7-17
worker/cb599-740fe4-16
worker/cb597-282224-15
worker/cb590fix-185e9a-10
worker/cb528-recovery-race
worker/cb594-96bead-8
worker/cb590-916766-2
worker/cb527-997d99-3
worker/cb592-env-leak-3cbf9c-1
worker/cb588-async-ticket-nudge-3218f7-5
worker/cb578b-9dcb13-6
worker/cb581-d24826-5
worker/m2-u5-ef8c42-15
worker/cb578a-516499-2
worker/cb576-01a04b-17
worker/cb579-lead-tab-acba06-20
worker/cb580-terminal-health-ed6058-21
worker/cb577-f36fdc-18
worker/cb573b-3db06f-16
worker/cb568c-f36fdc-18
worker/cb568-drop-cause-c3ac1c
worker/cb575-cancelled-notification-c3ac1c
worker/m4-sol-a2cbec-3
worker/cb574-async-ask-c3ac1c
worker/cb573-health-model-8ca857-14
worker/cb572-unknown-target-7f2e35-13
worker/u4-700706-9
worker/u3-b9fcb6-6
worker/u2-ef5b68-4
worker/u1-469dce-1-clean
worker/u1-469dce-1
worker/cb-564-health-events-70cf7e-2
worker/cb-565-recycle-drops-role-98e58f-3
worker/cb-563-missing-reply-df2866-1
worker/cb-562-readiness-gate-silent-6c23c9-3
worker/cb-560-architect-presence-da8155-1
worker/cb-561-architect-silent-off-a71cab-2
worker/cb-548-bind-architect-slot-fe1b8c-1
worker/parity-overlay-settings-5fb711-1
secrets-central-store
cb-559-hot-key-correction
cb-557-fleet-role-pools
worker/cb-553-maxload-explicit-spawn-305ee3-6
worker/cb-551-idle-lead-heartbeat-f1633c-1
worker/cb-544-drain-preserves-worktree-925fad-3
worker/cb-552-docs-sync-1cb9cf-4
worker/cb-548-rendezvous-guard-rebased
worker/cb-548-rendezvous-guard-116b53-10
worker/cb-548-authz-v2-586df6-8
worker/cb-548-authz-264363-5
salvage/cb-528b-codex-home
salvage/cb-528a-codex-launcher
CB-518-primary-flow
feature/peer-launcher-spi
cb-103-injector
v1.1.0
v1.0.0
Labels
Clear labels
blocked
needs-live-proof
ready-to-delegate
silent-default
Cannot start until something else lands. The body says what.
Merged and green, but never shown working on the running daemon. Not the same as done.
Scope, files and acceptance criteria are written. A worker can be briefed from the body alone.
A feature that compiles, passes tests, and ships turned off. Nine recurrences and counting.
No Label
Milestone
No items
No Milestone
Projects
Clear projects
No project
Notifications
Due Date
No due date set.
Dependencies
No dependencies set.
Reference: fleet/fleetd#544
Reference in New Issue
Block a user
Blocking a user prevents them from interacting with repositories, such as opening or commenting on pull requests or issues. Learn more about blocking a user.
Delete Branch "%!s()"
Deleting a branch is permanent. Although the deleted branch may continue to exist for a short time before it actually gets removed, it CANNOT be undone in most cases. Continue?
Follow-up to #538, which merged as PR #543. Filed by me while verifying that merge.
#538 fixed the part that made the failure permanent. This is the part it deliberately left open.
What #543 changed
Both
StatusPoller.loopandSessionReaper.loopnow catchThrowableper work item and continue,and both wrap the whole loop in a
finallythat logs at ERROR and setsrunning = false:Before this, a dead loop left
running == true, andstart()opens withif (running) return;,so a restart was impossible even in principle. Now it is possible.
What is still missing — measured on main at
4a8a780+ PR #543Nothing calls
start()a second time.Each is constructed once and started once. The remaining hits in that grep are javadoc mentions in
StatusRefiner,InjectorandHerdrPeerLauncher, not call sites.So
"it can be restarted"is true and nobody does it. The operational picture is better thanbefore #538 but still bad:
/healthzfleet_list/fleet_statusstart()callThe single ERROR line is a real improvement over #538's silence. It is still a line in a log nobody
is tailing, on a daemon that keeps answering
/healthzand keeps acceptingfleet_sendwhiledelivering nothing.
What is wanted
Two separate things. Either is useful alone; do not let the harder one block the easier one.
snapshot is the obvious home — a dead injection poller is more urgent than anything else that
snapshot currently carries. #538 listed this as "consider surfacing it, out of scope unless
cheap", and it was correctly left out.
running == falseon a loop that was never stopped,and calls
start()again. Decide and write down what happens if it dies immediately andrepeatedly — an unbounded restart loop hammering a broken herdr is its own outage. A bounded
count, or a backoff, with the give-up state reported through item 1.
Item 2 without item 1 is a trap: a supervisor that silently restarts a loop that keeps dying hides
exactly the failure that #538 was filed about.
Please read before choosing a design
stop()also setsrunning = false. A supervisor must be able to tell "stopped on purpose" from"died", or it will fight
stop()during shutdown. That is the samecannot-tell-no-from-cannot-tell shape that has bitten this repo before: if one flag has to carry
two states that need opposite handling, it needs a third state, not a cleverer reader.
Fleetd.main. Whatever supervises them lives there oralongside; it is not a per-target concern.
Acceptance
the fix.
stop()is not restarted. This is the one that catches thetwo-states-one-flag defect, so it is not optional.
the reported value fails a test.
message, restore, confirm byte-identical with the full
shasum -a 256, and run a green control.@TempDir.Out of scope
catch (Throwable)decision from #538. It is merged and verified.same reason it still applies: it hides per-loop recovery behind a global net.
Related: #538 (the fix this follows), #413 (the route that makes
NoClassDefFoundErrorreachablehere), #412 (the same family in the completion thread).
Correction to my own ticket, before anyone starts it. As filed above, item 2 asks for a
supervisor that "notices
running == false". That design is blind to the more likely failure.The fleet01 lead measured the herdr call path and sent me the result. I re-measured it here rather
than take their word for it, because their tree is 91 commits behind mine and they said so
themselves.
Measured on
mainat7611b69UnixSocketHerdrClient.call()opens aSocketChannel, writes, and then reads with no deadline:No timeout mechanism exists anywhere under
herdr/:The control is there because a zero match is a fact about the pattern until something proves the
search could hit. It could.
The call path has no wrapper timeout either:
AgentControl.status→get(target)→agentCall("agent.get", …)→herdr.call(…), andgrep -c TimeUnitoverStatusPoller.java,AgentControl.javaandUnixSocketHerdrClient.javais 0.What that means for this ticket
StatusPoller.loopcallscontrol.status(target)on the poller thread. A herdrd that accepts theconnection and never writes a line parks that thread forever.
Compare the two ways the loop stops working:
runningisisAlive()saysErrorescapes the loop (#538)false(since #543)falsetruetrueSo a watchdog built to item 2 as I wrote it closes the door #538 came through and leaves this one
open, while saying everything is fine. That is a negative check over a channel that cannot carry the
failure — the same shape as #512.
Item 2 must be a PROGRESS watchdog, not a liveness watchdog. Record a last-completed-round
timestamp from inside each loop, and alarm on the timestamp being stale. That detects both cases: a
dead loop stops updating it, and a parked loop stops updating it.
The hang needs no
Errorat all, so it does not depend on #413 or on any of #538's reachabilityargument. A herdrd that is up but wedged is an ordinary operational state.
One thing to check before designing the recovery
SocketChannelis an interruptible channel, soThread.interrupt()on a thread blocked inch.readshould close the channel and raiseClosedByInterruptException, which becomes anIOExceptionand then aHerdrException. If that holds,stop()followed bystart()is aworking recovery lever for the parked case, not only the dead one.
I have not measured this. It is the documented JDK contract, not something I ran. Verify it with
a real test before building recovery on top of it.
Revised acceptance for item 2
Replacing the item 2 bullets above:
reports it. This is the test the original wording would have let a worker skip.
stop()is reported as stopped, not as stalled.stop()also setsrunning = false, so one flag carries two states that need opposite handling; that needs a thirdstate, not a cleverer reader.
One note for whoever implements this, and it is an argument for the obvious choice rather than against it — so it belongs in a comment in the code, not just here.
System.nanoTime()is the right clock for this, and it is worth saying whyA progress watchdog compares "now" against a last-completed-round timestamp. The natural worry is that
nanoTimeis not wall-clock, and someone will later "fix" it toSystem.currentTimeMillis()or anInstant. That would be a regression.nanoTimedoes not advance while this host is asleep. Measured here previously:ps etimereported 60,541s for the fleetd process whilejcmd VM.uptimereported 2,966s — a 16-hour gap, all of it sleep. Every fleetd TTL and stall threshold already freezes with the host for that reason.For this watchdog that behaviour is correct, not a bug:
nowandlastRoundNanosfreeze together, so the delta does not grow and the state staysRUNNING. The loop genuinely was not running, and alarming about it would be a false positive — nobody was being served either.A wall-clock implementation would do the opposite: every laptop lid-close longer than the threshold would report
STALLEDon wake for a perfectly healthy daemon. That is the alarm-fatigue failure, and it would very likely get "fixed" by raising the threshold until the watchdog stops detecting anything.So: use a monotonic
LongSupplier(System::nanoTimelive, a controllable stub in tests), and put the reason in a javadoc line on the field. The next person to look at it will otherwise see a non-wall-clock timestamp and assume it was an oversight.Two smaller things while in there
start()runs, the first staleness judgement is measured from construction. Areset()called fromstart()— clearing any previous stop mark and re-stamping the timestamp — gives the loop a fresh grace period before its first round completes.Injector.POLL_INTERVAL_MILLISfor the poller, the reaper's own) with generous headroom, and say in a comment how it was derived. A bare constant invites exactly the "just raise it" fix described above.Follow-up, measured on
mainat26f380a. This comment is newer than any brief you were sent — where it disagrees, this wins.A watchdog nothing reads is a watchdog that cannot raise an alarm.
state()has to reach a surface. Below is what I measured about where it can go and what it must not do when it gets there.1. It does not belong in
HealthSnapshotEvery field there is per member, built once per member per tick at
FleetHealthMonitor.java:192. A loop's progress is one daemon-wide fact. Putting it in a per-member record makes N copies of one fact and invites a reader to attribute a daemon-wide stall to whichever member it happened to read.2.
/healthzis the right surface — and 503 is the only safe non-200 codeFleetApp.java:293-330reports herdr reachability and nothing else. A parkedStatusPollerleaves it fully green today.Before changing it, I measured both consumers the javadoc at
:280-291names:So the javadoc's "both only check the HTTP status code" is understated: both special-case 503 as non-fatal and treat every other non-200 as fatal.
That gives a hard rule:
redeploy-fleetd.shdieat:1087and report a successful redeploy as a failure. That is precisely the #552 defect, re-created one layer up. Do not introduce one.If you report 503, extend the
detail/herdrwording so a reader can tell "herdr unreachable" from "a loop stalled". Both scripts print the body verbatim, so the distinction reaches a human for free. I would also add the loop states to the 200 body, so aRUNNINGreading is positively visible and not merely the absence of an alarm.3. One claim I checked and had wrong — do not re-derive it
I expected a parked
StatusPollerto blindFleetHealthMonitor. It does not:liveStatuscomes from the monitor's ownagents.list(), independent of the poller.The real coupling is
session.state(). Session states move oninjector.onStatus(...), which only the poller calls. So a parked poller freezes every session atBUSY, and then:turns true for every member at once. The failure is not silence — it is a fleet-wide false-positive storm with no field anywhere naming the single common cause. That is the strongest argument for surfacing
state(): it converts N misleading per-member alarms into one true daemon-level fact.4. Two smaller things
LoopWatchdog.java:38says "a monotonic elapsed-time clock", which says what it is and not why it must stay that way.nanoTimedoes not advance while this host sleeps — measured here:ps etime60,541s againstjcmd VM.uptime2,966s. For this watchdog that is correct, because a frozen delta during sleep is a trueRUNNING(nobody was being served either), while a wall clock would reportSTALLEDafter every lid-close. Without that sentence someone "fixes" it to anInstantand the alarm becomes noise.StatusPoller.java:30-38andSessionReaper.java:20-28are right and I am not asking you to change them. Deriving from each loop's own interval with the reasoning written down is exactly what this needed.Lead verification of PR #559 at
735b6af. This comment is newer than any brief — where it disagrees, this wins.The PR is good. Verified on the merged tree (
origin/main+735b6af), not on the branch:The package move to
dev.ltms.fleet.injectis right and I confirmed the reason independently before reading the worker's note:healthimportssession.MemberSession, andsessionimportsinject.TurnListener/inject.MemberPresence, whileinjectimports neither. SoLoopWatchdoginhealthwould have closedinject → health → session → inject. Good catch, andPackageCyclesTestearning its keep is worth recording.The threshold derivations (40x / 12x, each from its own loop's interval, with the reasoning written down) are exactly right and I am not asking for changes there.
One surviving mutant
I ran a mutation the worker did not: remove
watchdog.reset()fromStatusPoller.start()(:105).(My first anchor was a regex on
watchdog\.reset\(\), which still matched inside the comment I had inserted and read1— a false not-applied. Re-counted with exact full-line equality against a pristine control. Recording that because it is the same trap the suite's own\Qincident was.)Result — the whole suite survives it:
Harness is live:
StatusPollerWatchdogTestdoes go red for the worker's own M4 onstop(), so this is a real survivor and not a dead harness.What the survivor costs
stoppedByCalleris sticky, andreset()is the only thing that clears it:start()isreset()'s only caller. So without that one line,stop()thenstart()leaveshealth()reportingSTOPPEDforever while the loop is genuinely running. Both classes documentstart()as idempotent, andloop()'s own error line says "it can be restarted" — so restart is an anticipated path, and the watchdog would quietly lie about a live daemon on exactly that path.The
reset()call is already there and already correct. It simply has nothing pinning it.The missing test — written and verified, not proposed
Restored,
shasum -a 256 -cOK.SessionReaperhas the identical sticky-flag shape and needs the twin test.The shape, which is the reason this is worth a comment rather than a silent fix
This is the invariant-not-idiom case the fleet01 lead named, in its other direction. The
reset()call is load-bearing, its absence has no symptom in any existing test, and the invariant it maintains is nameable in one sentence: an intentional stop must not outlive the restart that follows it. That sentence is the test. Had it been unnameable, the call would have been the suspicious thing instead.Not in this PR, and that is fine
health()still has no production caller — I grepped the diff, and nothing touchesFleetAppor/healthz. The worker scoped this as "the observability half only" and I agree that is a clean cut, so I am not holding the PR for it. The surfacing has real consumer constraints (measured in my earlier comment: both script consumers tolerate 503 specifically anddieon any other non-200) and deserves its own unit. I will file it.Merged as part of #559. Verified by the lead, on a merged tree — not on the branch.
Setup
mainat7a3b2bb+ branchbfac141, merged in a throwaway worktree =5de807f.Baseline:
mvn clean install→MVN_EXIT=0,Tests run: 1744, Failures: 0, Errors: 0, Skipped: 0, 130 surefire report files.Mutations — three, all killed
Each one deleted or changed one line, counted the pristine full line by exact string equality (
awk -v p="$PRISTINE" '$0==p'), kept a pristine copy as a control, and restored to a byte-identical file before the next run. This matters here: an earlier round of this same ticket produced a false "not applied" because the anchor regex matched inside the comment the mutation itself had inserted (comment 16944). Exact full-line equality has no such failure mode.StatusPoller.java:105— deletewatchdog.reset();fromstart()StatusPollerWatchdogTest.aRestartedLoopReportsRunningAgainNotStoppedForeveran intentional stop must not outlive the restart that follows it ==> expected: <RUNNING> but was: <STOPPED>LoopWatchdog.java:73— deletelastRoundNanos = nowNanos.getAsLong();fromreset()LoopWatchdogTest.resetClearsAPreviousStopAndTheStaleClockreset() must clear both the stop mark and the stale timestamp ==> expected: <RUNNING> but was: <STALLED>LoopWatchdog.java:89—>=→>LoopWatchdogTest.reportsStalledOnceTheLastRoundAgesPastTheThresholdexpected: <STALLED> but was: <RUNNING>M1 is the surviving mutant this rework was sent back for. It is now dead, and it dies with the invariant named in the failure message rather than a bare
expected/was.M2 and M3 are mutations the worker did not run. M2 in particular is a line whose count is 2, not 1 —
recordRoundComplete()at:63holds the identical text — so a naive 1 → 0 check would have read as "not applied" and scored a false survivor. It is 2 → 1, with the control still reading 2.M2 was also a prediction I got wrong, and that is the useful part. I picked it expecting a survivor: the javadoc on
reset()claims it "resets the clock, so the loop gets a fresh grace period", and the new restart test freezes its clock at0, so that test cannot possibly see the difference. The harness caught it anyway, from a different test file. A comment claiming an invariant is a free test case — and here the assertion the sentence describes already existed.Control build after all three restores:
MVN_EXIT=0,Tests run: 1744, Failures: 0.git status --shortempty in the verification worktree.Scope, unchanged
Still the observability half only.
health()is public and returns the three-state fact; nothing reads it yet. The surfacing unit — wiringhealth()into/healthz,fleet_listorfleet_status— is not filed yet and carries one measured constraint worth writing down before anyone designs it:Closing.