fleetd #498: awaitHerdr distinguishes deadline-passed from interrupted, with measured elapsed time #502

Merged
ltms merged 1 commits from worker/498-451d1c-5 into main 2026-09-12 05:01:02 +02:00
Member

Fixes fleetd #498.

awaitHerdr used to return a bare boolean, collapsing two different facts onto the same false: the configured wait budget genuinely running out, and the waiting thread being interrupted possibly milliseconds in. The caller's log line also printed only the configured budget, never the measured elapsed time.

  • awaitHerdr now returns HerdrAwaitOutcome(HerdrWaitResult, elapsedNanos) with three named states: ANSWERED, DEADLINE_PASSED, INTERRUPTED.
  • awaitHerdr takes the clock (LongSupplier) and the per-poll sleep (Runnable) as required parameters, no defaulted overload (fleetd #415's shape), so it is testable. The one production call site passes System::nanoTime and a real sleep helper.
  • The startup call site's logging decision is extracted into logHerdrWaitOutcomeAndShouldReap (main() itself cannot be driven from a unit test), which logs a distinct message per outcome, always printing configured=Ns elapsed=Xms together, never the configured budget alone.
  • Added FleetdAwaitHerdrTest: covers the seam (all three outcomes with an injected clock/stub HerdrClient, including asserting the interrupt flag is preserved) and the call site (the three distinct log messages via ListAppender, same pattern as AmqpConnectionFailureLoggerTest/LeadRolloverTest).

Tests: mvn clean install — Tests run: 1683, Failures: 0, Errors: 0, Skipped: 0 — BUILD SUCCESS.

Mutation testing (see PR discussion / worker report for the full table): 3 mutations run, one per outcome/branch, each caught by a distinct failing test; control run afterward matched the baseline (6/6 green); harness-proof (broken string) also caught by the suite before any mutation was applied.

Out of scope, noted per the ticket, not fixed: selectLeadMailbox (~Fleetd.java:1330-1358) returns null for three distinct causes (not configured, blank coordinator.selfId, broker unreachable) — a weaker instance of the same shape, though each cause is already logged distinctly at its own return site. No other boolean-collapsing or configured-constant-only log line found elsewhere in Fleetd.java.

Fixes fleetd #498. `awaitHerdr` used to return a bare `boolean`, collapsing two different facts onto the same `false`: the configured wait budget genuinely running out, and the waiting thread being interrupted possibly milliseconds in. The caller's log line also printed only the configured budget, never the measured elapsed time. - `awaitHerdr` now returns `HerdrAwaitOutcome(HerdrWaitResult, elapsedNanos)` with three named states: `ANSWERED`, `DEADLINE_PASSED`, `INTERRUPTED`. - `awaitHerdr` takes the clock (`LongSupplier`) and the per-poll sleep (`Runnable`) as required parameters, no defaulted overload (fleetd #415's shape), so it is testable. The one production call site passes `System::nanoTime` and a real sleep helper. - The startup call site's logging decision is extracted into `logHerdrWaitOutcomeAndShouldReap` (main() itself cannot be driven from a unit test), which logs a distinct message per outcome, always printing `configured=Ns elapsed=Xms` together, never the configured budget alone. - Added `FleetdAwaitHerdrTest`: covers the seam (all three outcomes with an injected clock/stub HerdrClient, including asserting the interrupt flag is preserved) and the call site (the three distinct log messages via `ListAppender`, same pattern as `AmqpConnectionFailureLoggerTest`/`LeadRolloverTest`). **Tests:** `mvn clean install` — `Tests run: 1683, Failures: 0, Errors: 0, Skipped: 0` — `BUILD SUCCESS`. **Mutation testing** (see PR discussion / worker report for the full table): 3 mutations run, one per outcome/branch, each caught by a distinct failing test; control run afterward matched the baseline (6/6 green); harness-proof (broken string) also caught by the suite before any mutation was applied. **Out of scope, noted per the ticket, not fixed:** `selectLeadMailbox` (~Fleetd.java:1330-1358) returns `null` for three distinct causes (not configured, blank `coordinator.selfId`, broker unreachable) — a weaker instance of the same shape, though each cause is already logged distinctly at its own return site. No other `boolean`-collapsing or configured-constant-only log line found elsewhere in `Fleetd.java`.
agent added 1 commit 2026-09-12 04:55:28 +02:00
fleetd #498: awaitHerdr distinguishes deadline-passed from interrupted, with measured elapsed time
CI / contract (pull_request) Successful in 51s
CI / build (pull_request) Successful in 1m51s
274afafde6
- awaitHerdr now returns a HerdrAwaitOutcome(HerdrWaitResult, elapsedNanos) instead of a bare
  boolean, so 'the wait budget genuinely ran out' and 'the waiting thread was interrupted' are
  two distinct, named states instead of the same false (fleetd #497's shape).
- awaitHerdr takes the clock (LongSupplier) and the per-poll sleep (Runnable) as required
  parameters, with no defaulted overload (fleetd #415), so a test can drive it.
- The startup call site is extracted into logHerdrWaitOutcomeAndShouldReap, since main() itself
  cannot be driven from a unit test; it logs a distinct message per outcome, always printing the
  measured elapsed time next to the configured budget, never the budget alone.
- Adds FleetdAwaitHerdrTest covering the seam (all three outcomes, plus the preserved interrupt
  flag) and the call site (the three distinct log messages), using ListAppender.
ltms merged commit 708f1795ad into main 2026-09-12 05:01:02 +02:00
Sign in to join this conversation.