LoopWatchdog.health() is public and nothing reads it — surface it without breaking the two scripts that parse /healthz #562
Closed
opened 2026-09-12 11:15:10 +02:00 by ltms
·
3 comments
No Branch/Tag Specified
main
worker/fleetd-612-unita-87807e-1
worker/612-b3-mcpwirings-da2b58-3
worker/612-b2-cb185-176d3a-2
worker/612-b1-completion-457459-1
worker/612-agaps-73a926-2
worker/608-sleeps-3a64ff-3
worker/621-b4520b-1
worker/618-b83894-2
worker/fleetd-615-e05481-5
worker/lead-autocompact-5f1ab2-3
worker/fleetd-613-f85deb-3
worker/fleetd-608-flaky-nudge-test-d0c2d1-3
worker/lead-context-gauge-ad404f-1
worker/gauge-wiring-9158c1-4
worker/redeploy-slowstart-ead0e5-5
worker/charter-bytes-13668c-6
worker/rollover-outcome-291483-2
worker/589-f64303-2
worker/593-1a8025-5
worker/589-fcd2aa-1
worker/568-9fdaa2-3
worker/571-attempted-outcome-5739f7-2
worker/581-completionresolver-cas-sites-0542b7-6
worker/562-loop-health-wiring-test-99611c-5
worker/562-surface-loop-health-7df5cc-4
worker/575-waiter-cleanup-sites-62ad80-1
worker/572-answer-lock-release-46a9ae-5
worker/567-probe-channel-leak-a38fc5-6
worker/551-record-before-send-7cbf56-1
worker/561-listener-fanout-survives-a-throw-61d538-2
worker/555-redeploy-main-flow-seam-65c2f5-2
worker/556-injector-owns-registration-e027a5-1
worker/552-post-restart-mktemp-abort-bc2672-4
worker/553-onstatus-completion-leak-0da881-2
worker/550-shasum-linux-196132-1
worker/538-loop-dies-on-error-4a5eeb-6
worker/426-health-coverage-ef1fd4-4
worker/504-failed-reported-clean-3cfd66-3
worker/537-capturedlog-close-e4c437-2
worker/459-broken-link-targets-cadc17-5
worker/535-appender-leak-fe74c1-1
worker/512-part2-shutdown-detection-434701-9
worker/529-logger-level-sweep-2a5533-8
worker/528-drain-gate-call-site-5de83d-7
charter/forge-mcp-vs-token
worker/521-swap-guard-unpinned-28e931-5
worker/519-probe-test-harness-d25ab8-4
worker/525-logger-level-leak-1b4eb0-6
worker/518-fleetmcp-resolver-wiring-8ef96c-1
worker/512-drain-complete-line-7edd71-3
worker/517-abort-branch-and-jar-id-41b641-2
worker/500-9e52c9-3
worker/509-4912f4-2
worker/511-9a4b23-1
worker/493-479f45-2
worker/505-03f8b2-1
worker/492-followup-detect-unclear
worker/501-a31fa0-7
worker/498-451d1c-5
worker/494-1015ce-2
worker/492-209647-1
worker/489-001902-2
worker/480-relative-handover-path-906323-1
worker/480-b-handover-skill-45bf1f-5
worker/474-followup-source-pin-f54a55-17
worker/474-charter-check-on-reload-f54a55-17
worker/466-quarantine-repeatcount-report
worker/393-opencode-skill-seeding-71854b-13
worker/469-canonical-tool-names-2a472a-16
worker/466-quarantine-escalation-5ae9c1-15
worker/446-hot-exhausted-pattern-0af580-6
worker/464-charter-tool-name-guard-a85635-12
worker/463-listfleet-default-fails-open-f1c76c-11
worker/458-invariant-5-by-purpose-862f9a-10
worker/439-coordinator-row-gate-bc032a-8
worker/449-herdr-protocol-576015-4
worker/450-abstract-spawn-599e1c-5
worker/437-ack-refuses-177d91-1
worker/444-placement-window-feb56a-2
worker/440-helddurable-derived-d462d7-13
worker/425-rework-placement-resolve-c58ba1-9
worker/421-lead-peek-held-msgs-cdbad2-10
worker/435-fixed-policy-cap-fe11de-12
worker/422-gate-state-observability-9e79d6-11
worker/431-memberregistry-live-readers-cdbad2-10
worker/424-architect-slot-hot-038b41-7
worker/422-model-gate-spawn-c29f48-6
worker/425-default-profile-live-f55534-8
worker/415-coverage-wording-2cbf9c-5
worker/416-3ad1da-1
worker/418-588283-3
worker/deterministic-stamp-race-409-3cb7b6-10
worker/armed-reads-live-config-404-ed931f-9
worker/reply-peer-refusal-391-5a34bd-7
worker/models-allowlist-aa9e9b-3
worker/ttl-stamp-race-399-f1122f-8
worker/scrub-receipt-400-316b3e-5
worker/exhaustion-detection-395-105105-6
worker/scrub-abort-394-316b3e-5
fix/scrub-uid-abort
worker/task-scrub-517574-2
worker/t386-clock-bd5b78-4
worker/t384-scrub-813790-5
worker/t381-cc-748314-2
worker/t373-336973-2
worker/t365-3920c5-3
worker/t358-6e989b-1
worker/t355-8b321c-1
worker/fleetd-369-hermetic-git-tests-e8b19a-3
worker/fleetd-368-stale-lead-binding-f5682e-2
worker/fleetd-360-deploy-units-0d3793-1
worker/359-dead-lead-tabs-f1253b-4
worker/362-worktree-skills-c03e51-3
worker/361-coord-visibility-655144-1
362-plugin-visibility-and-drift
worker/errscan-bed2ca-2
worker/amqp-log-identity-bed2ca-2
worker/withdefaults-guard-561704
worker/sleepguard-82076d-1
worker/fd334-9ee1b6-5
worker/fd348-f1ab27-4
worker/fd335-a71c35-1
worker/fd342-174a17-2
worker/fd345-490d0f-3
worker/fleetd-337-5ec7d4-21
worker/fleetd-341-af5a6b-24
worker/fleetd-339-5ca0a2-23
worker/fleetd-338-83a4a1-22
worker/fleetd-333-281f46-18
worker/fleetd-329-11bdbb-16
worker/fleetd-330-2770fb-17
worker/fix-326-50506e-15
worker/fix-324-3e9bbf-14
worker/fix-323-b8287d-13
worker/fix-316b-bd0860-11
worker/fix-318-76ca36-9
worker/fix-317-486aec-8
worker/fix-315-ce47c5-6
worker/fix-307-275890-6
worker/fix-308-b4f664-7
worker/fix-309-ec3939-8
worker/fix-310-7a3974-9
worker/fix-302-52ad0e-9
worker/fix-298-ce1acb-8
worker/fix-297-66bd11-7
worker/fix-296-104622-6
worker/fix-293-bare-closetab-eb22b5-3
worker/fix-280-gone-ask-lapse-bca98e-2
worker/fix-290-reapidle-guard-coverage-9b0dd1-1
worker/fix-285-trust-seed-8f3565-10
worker/fix-284-backend-error-seat-85912c-11
worker/fix-282-chained-ask-e6d0bb-8
worker/fix-283-teardown-leaks-f40dfa-9
worker/fix-281-pin-handler-actions-4921ac-7
worker/audit-rendezvous-lifecycle-d072ae-2
worker/audit-health-placement-1a2476-6
worker/audit-teardown-exits-e207a5-3
worker/audit-launcher-asymmetry-27e370-4
worker/audit-rest-authz-6ca53c-5
worker/investigate-275-abandon-asking-fdef52-8
worker/fix-274-worktree-leak-b0095d-7
worker/fix-273-exhausted-pattern-9665b5-6
worker/fleetd-267-model-check-bd8068-1
worker/fleetd-131-archunit-18b834-7
worker/fleetd-266-sshagent-rename-a014ff-6
worker/fleetd-184-uid-claim-8e1f31-4
worker/fleetd-184-warn-b381ee-10
worker/fleetd-184-docs-be1d12-9
worker/fleetd-257-9bf010-7
worker/fleetd-103-23a113-6
worker/fleetd-247-342356-5
worker/fleetd-116-04dea8-4
worker/fleetd-252-a830e0-3
worker/fleetd-111-7e8673-9
worker/fleetd-155c-f8ef4b-8
worker/fleetd-176-b928ca-3
worker/fleetd-249-7a7878-2
worker/cb248-composition-root-b-9acdf7-15
worker/cb148-envrc-default-fa6c82-12
worker/cb201-unit5-wiring-6c12e6-8
worker/cb241-fallback-echo-1175e9-11
worker/cb149-trust-dialog-2392a5-9
worker/cb134-148-overlay-visible-c9b986-10
worker/cb234-session-id-keyed-04e1fc-1
worker/cb201-unit3-nudge-abdf5c-6
worker/cb201-unit2-policy-c1102c-5
worker/cb201-unit4-outcome-a13bfa-7
worker/cb201-unit1-classifier-91b9b1-4
worker/cb201-227-refine-831980-3
worker/cb175-model-readback-0f085f-1
worker/cb222-charter-tmpdir-17f013-1
worker/cb226-architect-slot-race-cd3aa8-3
worker/cb224-worktree-root-group-024523-2
worker/cb-123-role-demotion-c600f7-2
worker/cb-219-opencode-roots-1f677e-1
worker/cb214-claude-session-id-b9eab4-4
worker/cb213-zdotdir-wrong-process-dd6de4-3
worker/cb211-exhaustion-classification-9546e0-2
worker/cb137-ambiguous-task-4df3d8-4
worker/cb209-agentsessionid-4dfdb6-2
worker/cb185-hostenvnames-2692b5-3
worker/cb206-opencode-sqlite-128718-2
worker/cb185-worktree-group-fc0c99-1
worker/cb-137-ask-ticket-e7760c-2
worker/cb-172-broker-uri-d36ae4-4
worker/cb-175-model-readback-76ead6-3
worker/cb-161-pane-ancestry-293510-1
worker/cb-164-rebase-885863-8
worker/cb-164-empty-scrape-false-success-1a80af-3
fix/cb-197-ticket-ttl-from-completion
worker/cb-189-remote-url-coverage-4692f3-1
worker/cb-185-blockers-027756-4
worker/cb-192-gap-log-11b631-2
worker/cb-633-fix-5f4396-3
worker/cb185-router-d6436d-3
worker/cb185-router-routing-gaps-9e9d33-3
worker/cb185-paneids-992586-2
worker/cb-633-allow-list-union-ed374b-1
worker/cb-157-credential-in-remote-url-496e44-2
worker/cb-641-health-herdr-evidence-8f1f54-6
worker/cb-640-health-msg-evidence-99c9cd-1
worker/cb-642-fleets-status-skill-bbbc40-5
cb-634-ide-mcp
worker/lead-comms-wiring-c014b9-7
worker/lead-mailbox-c19577-6
worker/autocompact-window-82bc2f-5
worker/cb-634-probe-18056f-4
worker/cb635-broker-urienv
worker/cb-632-config-retry-8e0efa-7
lead/cb-622e-claude-md
lead/cb-622-followup
worker/cb-622a-165dff-1
lead/cb-622d-opencode-mount
worker/cb-622b-717c67-2
worker/cb-622c-ab7759-3
worker/cb-617b2-20ca4b-3
worker/cb-617a-5c2f4a-1
worker/cb596-4e49ef-3
worker/cb586-10500c-1
worker/cb-606-b9343a-25
worker/cb604-1445f8-24
worker/cb582-477374-21
worker/cb584-8c2281-22
worker/cb600-e6b9a9-20
worker/cb602-ce257f-19
worker/cb601-b42837-18
worker/cb598-6c7ba7-17
worker/cb599-740fe4-16
worker/cb597-282224-15
worker/cb590fix-185e9a-10
worker/cb528-recovery-race
worker/cb594-96bead-8
worker/cb590-916766-2
worker/cb527-997d99-3
worker/cb592-env-leak-3cbf9c-1
worker/cb588-async-ticket-nudge-3218f7-5
worker/cb578b-9dcb13-6
worker/cb581-d24826-5
worker/m2-u5-ef8c42-15
worker/cb578a-516499-2
worker/cb576-01a04b-17
worker/cb579-lead-tab-acba06-20
worker/cb580-terminal-health-ed6058-21
worker/cb577-f36fdc-18
worker/cb573b-3db06f-16
worker/cb568c-f36fdc-18
worker/cb568-drop-cause-c3ac1c
worker/cb575-cancelled-notification-c3ac1c
worker/m4-sol-a2cbec-3
worker/cb574-async-ask-c3ac1c
worker/cb573-health-model-8ca857-14
worker/cb572-unknown-target-7f2e35-13
worker/u4-700706-9
worker/u3-b9fcb6-6
worker/u2-ef5b68-4
worker/u1-469dce-1-clean
worker/u1-469dce-1
worker/cb-564-health-events-70cf7e-2
worker/cb-565-recycle-drops-role-98e58f-3
worker/cb-563-missing-reply-df2866-1
worker/cb-562-readiness-gate-silent-6c23c9-3
worker/cb-560-architect-presence-da8155-1
worker/cb-561-architect-silent-off-a71cab-2
worker/cb-548-bind-architect-slot-fe1b8c-1
worker/parity-overlay-settings-5fb711-1
secrets-central-store
cb-559-hot-key-correction
cb-557-fleet-role-pools
worker/cb-553-maxload-explicit-spawn-305ee3-6
worker/cb-551-idle-lead-heartbeat-f1633c-1
worker/cb-544-drain-preserves-worktree-925fad-3
worker/cb-552-docs-sync-1cb9cf-4
worker/cb-548-rendezvous-guard-rebased
worker/cb-548-rendezvous-guard-116b53-10
worker/cb-548-authz-v2-586df6-8
worker/cb-548-authz-264363-5
salvage/cb-528b-codex-home
salvage/cb-528a-codex-launcher
CB-518-primary-flow
feature/peer-launcher-spi
cb-103-injector
v1.1.0
v1.0.0
Labels
Clear labels
blocked
needs-live-proof
ready-to-delegate
silent-default
Cannot start until something else lands. The body says what.
Merged and green, but never shown working on the running daemon. Not the same as done.
Scope, files and acceptance criteria are written. A worker can be briefed from the body alone.
A feature that compiles, passes tests, and ships turned off. Nine recurrences and counting.
No Label
Milestone
No items
No Milestone
Projects
Clear projects
No project
Notifications
Due Date
No due date set.
Dependencies
No dependencies set.
Reference: fleet/fleetd#562
Reference in New Issue
Block a user
Blocking a user prevents them from interacting with repositories, such as opening or commenting on pull requests or issues. Learn more about blocking a user.
Delete Branch "%!s()"
Deleting a branch is permanent. Although the deleted branch may continue to exist for a short time before it actually gets removed, it CANNOT be undone in most cases. Continue?
Follow-up to #544, which merged as PR #559. Filed by me while verifying that merge.
#544 shipped the observability half:
StatusPoller.health()andSessionReaper.health()each return aLoopWatchdog.StateofRUNNING/STALLED/STOPPED, and both are proven by tests. Nothing in production calls either one.Measured on the merged tree (
main7a3b2bb+bfac141, identical tomainatdab697f):A zero match proves nothing on its own, so the same pattern against the test sources, as a positive control that the search really runs:
And the declarations themselves, so the zero above is a zero about callers and not about the method existing:
Do not use the pattern
\.health()for this. It returns 7 hits infleetd/src/main/java, every one of themcfg.health()/config.get().health()— the config accessor inFleetd.javaandConfigRef.java, an unrelated method with the same name. Anyone re-measuring with the loose pattern will read those hits as production callers of the watchdog and conclude this ticket is already done.So a
StatusPollerloop that dies or parks now has a correct three-state fact about it, sitting in memory, that no operator and no tool can read. That is better than #538's silence and it is not yet observability.What is wanted
Make
health()reachable from outside the process. At least one of these; they are independent and the easiest one alone is worth shipping:fleet_listand/orfleet_status— a field per loop, alongside the existing capacity and coordinator rows. Probably the cheapest, and it reaches a lead directly./healthz— see the hard constraint below before designing anything here.HealthSnapshot/FleetHealthMonitor) — notehealthCoveragecurrently reportsdetection-only, so check what that promise already means before adding to it.The constraint that decides the
/healthzdesign — measured, 2026-09-12503 is the only safe non-200 status code
/healthzcan return. Both scripts that parse it treat every other non-200 as fatal:scripts/redeploy-fleetd.sh:1082:1087—diescripts/rename-checkout.sh:448:452—dieRe-measure before relying on this:
If both files still branch on
503with adieon the else, this section still applies. If either stops parsing the status code, delete this section.That leaves a real dilemma, and whoever takes this should decide it deliberately rather than discover it:
StatusPollerindistinguishable from an unreachable herdr, which is the exact two-states-one-symbol shape (#512) thatLoopWatchdogexists to avoid. Putting it back one layer out would be a poor trade./healthzat 200 and putting the state in the body costs nothing for either script, because neither reads the body. This is probably the right answer, but it means the loop state does not reach anything that only checks the status code.FleetApp.java:293-330is the endpoint. Today it reports herdr reachability and nothing else.A design warning from the fleet01 lead
Worth writing down before anyone chooses an alerting shape, because the failure direction here is the unusual one:
That specific bug is fixed and pinned —
aRestartedLoopReportsRunningAgainNotStoppedForeverin bothStatusPollerWatchdogTestandSessionReaperWatchdogTest, each proven by deletingwatchdog.reset()fromstart(). The general point survives the fix: whatever surfaces this must be at least as trustworthy as the thing it reports on, because a monitor that cries wolf is worse than no monitor. A surfacedSTALLEDthat turns out to be a bug in the surfacing is the failure to design against.Acceptance
STALLEDbyhealth()is reported as stalled through the new surface, and that inverting the reported value fails a test.STOPPEDis notSTALLED; collapsing the two at the boundary throws away the whole point of the three-state design./healthzis touched: a test pinning that the status code for a stalled loop is one the two scripts above survive, with the decision above stated in the PR.fleet_list/fleet_statusis touched: the intent→tool table inCLAUDE.mdand the wiki copy stay in sync, andMcpContractDocTestpasses. A visible new behaviour also earns an entry inwiki/11-Features.md— what it does, the knob, why it exists, the gotcha.sedonly, the pristine full line counted by exact string equality (awk '$0==p', not a regex — a regex can match inside the comment the mutation inserted, which produced a false "not applied" on #544), red with that test's own assertion message, restored byte-identical under a fullshasum -a 256, green control re-run.@TempDir.Out of scope
UnixSocketHerdrClient.call(). Separate, real, wanted — and this watchdog exists because that deadline does not exist yet.LoopWatchdog. It lives ininjectto avoid aninject → health → session → injectcycle;PackageCyclesTestenforces it.Related: #544 (the half that merged), #538 / #543 (the
catch (Throwable)work underneath), #512 (one symbol carrying two states).Decision, and a re-measurement of every number this ticket rests on
Measured on
mainat204da67, 2026-09-12. I re-ran the ticket's own commands rather than trust them, and two of the line numbers had already drifted.The zero still holds
Positive control, same idea against the test sources — non-empty, so the search really runs:
And the declarations, so the zero is about callers and not about the method existing:
StatusPoller.java:97andSessionReaper.java:82. Still no production caller.The 503 constraint still holds, but the line numbers moved
The ticket says
scripts/redeploy-fleetd.sh:1082/:1087. They are now:1009/:1014.scripts/rename-checkout.sh:448is still:448. Locate them fresh; do not trust either number, including mine.The constraint itself is intact, and the second script's half is worse than the ticket describes:
redeploy-fleetd.shreport_health():503→ warn and continue (:1010). Every other non-200 →die(:1014).rename-checkout.sh:445-452: acaseon the code with arms for200and503only. An unknown code does not die immediately — it keeps polling untilHEALTH_WAITruns out and then dies with "/healthz never answered within Ns". That message is wrong: the daemon answered every time. So a new status code here does not just break the script, it breaks it with a misleading diagnosis, on the day the fleet is already degraded.The body is free, and I measured why
The ticket assumes neither script reads the body. Half true, and the true half is better than assumed.
redeploy-fleetd.sh's green path is a body predicate, not a code check:Body non-empty. That is all. So adding fields to the
/healthzbody costs nothing for that script, andrename-checkout.shreads only the status code. The body really is free.The decision
1.
/healthzkeeps its status codes exactly as they are. 200 for ok, 503 for herdr unreachable. The loop states go in the body only. No new status code, and do not reuse 503 for a stalled loop — that is the one-symbol-two-states shape (#512) thatLoopWatchdogwas built to avoid, and putting it back one layer out would be a bad trade.2.
fleet_listis the primary surface. It is the cheapest of the three, it reaches a lead directly, and it carries no script risk at all. Do this one properly; it is the deliverable.3.
STOPPEDis never reported as an alarm. CollapsingSTOPPEDintoSTALLEDat the boundary throws away the whole reason the design has three states. A loop stopped on purpose is a fact, not a fault.4. Option 3 —
HealthSnapshot/FleetHealthMonitor— is out of scope for this unit.healthCoveragestill reportsdetection-only; changing what that word promises is its own decision and does not belong in the same PR.On the fleet01 lead's warning
Keeping it, and it shapes the acceptance rather than just sitting in the ticket. A monitor whose failure mode is a false positive gets muted, and then the real alarm is muted too. So the test that a deliberately stopped loop is not surfaced as an alarm is not a nice-to-have — it is the test that decides whether anyone will still be listening the first time this fires for real.
One thing the member cannot do, so it must not be asked of them
The ticket's acceptance says the
CLAUDE.mdblock and the wiki copy must stay in sync, and that a visible new behaviour earns awiki/11-Features.mdentry. A member's provisioned worktree haswiki/uninitialized — measured here in three worker worktrees — so the sync script dies withFileNotFoundErrorand the Features file does not exist for them. Asking a worker for either is asking them to invent a pass.So the split is: the member edits
CLAUDE.mdif the tool's described output changes, writes the Features entry as text in the PR body, and says plainly that it could not run the sync check. I run the sync check in the main clone and push the wiki myself before merging.Delegating now.
HOLD on PR #579 — I mutated the production wiring and it survived. One more test is required before merge.
Everything the worker reported checks out. This is not a correction of their work; it is a gap their tests do not cover, found by the mutation I run on every PR against the half the worker did not change.
What I verified first (all of it good)
Tests run: 1771, Failures: 0, Errors: 0, Skipped: 0, confirmed twice — from the MavenResults:block, and from an independent sum over 131 surefire report files. That is the 1766 baseline plus the worker's 5 new tests.FleetMcp.javaandFleetApp.javarestored byte-identical, as reported.The mutation that survived
Fleetd.java:669— located fresh, not from the ticket, because line numbers drift. Pristine anchor counted withgrep -Fxc= 1.Anchor count 1 → 0, marker present. Result:
The mutation survived.
Fleetd.javahas been restored;shasum -a 256matches the pristine75262d31753ea626.What that means in production terms
The daemon can be changed to always report the
StatusPollerasRUNNING— so the watchdog can never fire and a stalled poller is invisible — and all 1771 tests stay green.That is precisely the false-negative this ticket exists to prevent. Worth naming the direction: the fleet01 lead's warning already on this ticket was about a monitoring component whose failure mode is a false positive getting muted. This is the mirror. A false negative is worse, because there is no noise for anyone to notice and then silence.
Why the tests miss it (I ruled out the alternatives)
Not a dead line, and not an excluded test group. The cause is instantiation:
LoopHealthSourcewith fixed lambdas, e.g.new FleetMcp.LoopHealthSource(() -> statusPoller, () -> sessionReaper).grep -rln 'LoopHealthSource'across the branch's test tree returns onlyFleetMcpTest.javaandFleetAppTest.java.LoopHealthSourcereports what it is given — and prove nothing about whatFleetd.maingives it. This is the hand-built vs config-wired shape, the same one as #561.What is required to merge
Add
FleetdLoopHealthSourceWiringTest, following the 12 sibling wiring tests already in this project — specificallyFleetdHealthCoverageSourceWiringTest, which is the direct analogue forhealthCoverageSource(config).Acceptance, and please read the last point before starting:
poller::healthatFleetd.java:669is replaced by a constant() -> LoopWatchdog.State.RUNNING. Run that mutation yourself, paste the failure, then restore and paste the matchingshasum -a 256. A test that is not red under that exact mutation is not the test being asked for.reaper.health()is replaced by a constant. Two halves means two assertions — one invariant wired at two places needs one assertion per place. A single combined assertion whose total is non-zero is how a gap at one half hides.reaper == nullbranch working. That null check is real behaviour (STOPPEDwhen there is no reaper), so do not delete it to make the test easier — assert it instead, as a third case.loopHealthis built inline, and the sibling factories are not.capacitySource(...)atFleetd.java:1003andhealthCoverageSource(...)at:1037are package-private named factory methods, which is exactly what makes their wiring tests able to call them directly (Fleetd.healthCoverageSource(config)).loopHealthis an inlinenewinsidemain, so there is nothing for a test to call. Extract it to a package-private factory in the same style as those two, then test the factory. That extraction is the substance of this unit — without it there is no seam to pin, and the same mutation stays green whatever test you write.The rest of PR #579 is good and stays as it is. This adds one test and one small extraction; do not widen it further.
Merged. The survivor is dead — I re-ran my own mutation and it is now caught.
Merged to
mainas part of634d33b, via PR #584, which supersedes #579 (its branch is merged in as the base).My own verification:
Fleetd.java:669now readsloopHealthSource(poller, reaper)— the extracted factory is genuinely wired intomain, not a parallel copy sitting beside an unchanged inlinenew. That was the thing most worth checking, because an extraction that leaves the original call site alone passes its own test and changes nothing.Tests run: 1784, Failures: 0, Errors: 0, confirmed from theResults:block and an independent sum over 132 report files.poller::healthwith() -> LoopWatchdog.State.RUNNINGinside the factory is killed byFleetdLoopHealthSourceWiringTest.statusPollerHalfReflectsThePollersRealHealth:61. One test, not a crowd.The worker improved on the brief and said so. I asked for "replace
reaper.health()with a constant". They first did the literal thing — collapsing the whole ternary — found it killed both the reaper test and the null-reaper test, and recognised that as a less isolating mutation. They then did a surgical version that replaces only thereaper.health()sub-expression, keeps thereaper == nullbranch intact, and kills exactly one test, which additionally proves the third assertion is independent. They reported both, and chose the discriminating one. That is the right instinct: a kill that takes a crowd with it proves less than a kill of exactly one.The framing I got wrong, corrected by the fleet01 lead. I described the five original tests as weak coverage of the wiring question. They are not weak — they are structurally zero. Every one of them built its own
LoopHealthSourcewith fixed lambdas, and a test that supplies its own dependency is a test of the consumer that can never be evidence about the producer. So the survivor was not a puzzle to explain; it was the expected result. That is now written into the Features entry, so the next reader does not try to strengthen the five.Context worth keeping. The same peer measured that a tree 91 commits back has only 2
Fleetd*WiringTestclasses against 12 onmaintoday. Ten arrived in those 91 commits. So this convention is new and actively being swept — #562 did not miss an old rule, it landed wiring while the sweep was in progress and did not join it. A clean sweep ofFleetd.java's constructor-arg method refs finds roughly 17 wiring sites against 12 test classes whose names do not map one-to-one. Worth its own ticket; not yet filed.CLAUDE.md↔wiki/7-Use-Cases.mdsync check run in the main clone: in sync after re-syncing the block, and the wiki is pushed and verified by ref. Features entry written.Closing.