Lead review of PR #196 found two issues in CompositePeerLauncher.probeOwner:
1. The "more than one daemon claims this pane" throw kept the old
pre-fix message ("no owning herdr daemon was recorded"), which was
only true of the code it replaced. Reworded to say what actually
happened: N configured herdr daemons report this pane, so it is
genuinely ambiguous. Updated the one test pinning the old string.
2. probeOwner let list() propagate straight out of the probe loop, so
one unreachable daemon aborted the whole probe and made a pane on a
DIFFERENT, healthy daemon un-stoppable too — resurrecting the exact
bug blocker 1 fixes. Now catches HerdrException per daemon, logs the
exception class only, and treats that daemon as not knowing the pane
so probing continues. New test proves this: verified it fails with
the try/catch removed (HerdrException propagates and the stop that
should succeed via the healthy daemon throws instead), then restored.
Full mvn clean install: 1019 tests, 0 failures, 0 errors.
1. CompositePeerLauncher.stop() was permanently un-stoppable for any
member that survived a daemon restart, because spawnedBy is in-memory
only. On a cache miss with more than one configured herdr daemon, probe
each distinct daemon's agent.list() for the pane instead of refusing
outright: exactly one owner routes and caches; zero owners is treated
as already-stopped (a no-op, matching the tolerance HerdrPeerLauncher
already gives an already-gone pane); more than one owner is the
genuine per-daemon-pane-id ambiguity and still throws.
2. FleetApp#healthz always reported the LEAD daemon's herdr version/
protocol even when a second (member) daemon was configured, so a
member-daemon protocol mismatch was invisible behind a green
/healthz while every spawn silently failed. Added a separate "member"
key alongside the unchanged "herdr" key, and a "protocolMismatch"
flag when the two differ. Verified scripts/redeploy-fleetd.sh and
scripts/rename-checkout.sh only check the HTTP status code and print
the body verbatim — neither parses a specific field — so adding a key
is safe.
Both fixes are covered by tests written to fail without the fix
(verified by reverting each fix and watching the new tests fail, then
restoring). Full `mvn clean install`: 1018 tests, 0 failures, 0 errors.