aa517ae0ec92f711871cbd9b0e97e06faa8bb69a
1020 Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
034e17bb32 |
fleetd #561 follow-up: pin the session half of bothMustRunKeepingSecondResult
Two helpers maintain one invariant (the second callback half always runs, even when the first throws): bothMustRun and bothMustRunKeepingSecondResult. Only bothMustRun's "session half still runs" direction was asserted (sessionHalfStillRunsWhenTheCompletionHalfThrowsSynchronously, via onTurnComplete). bothMustRunKeepingSecondResult — the helper onTurnCompleteWithPostAction uses — could be reverted to the pre-#561 broken shape and the suite stayed green. Adds two tests to FleetdTurnListenerCompositionTest: - sessionHalfStillRunsWhenTheCompletionHalfThrowsSynchronouslyForPostAction: mirrors the existing onTurnComplete case for onTurnCompleteWithPostAction/ bothMustRunKeepingSecondResult. - bothFailuresEscapeWhenBothHalvesThrowDistinctExceptions: proves a second, distinct failure from the session half is preserved via addSuppressed rather than silently dropped when both halves of bothMustRun throw. Also rewords the onDelivered comment in Fleetd.turnListener: it previously said this pair is safe because registration survives a throw via #556's Injector wiring, which is true but is not why THIS pair is unguarded. CompletionResolver.captureBaseline already catches RuntimeException around its scrape read and fails open, so completion.onDelivered does not realistically throw. Comment text only, no logic change. |
||
|
|
68b428c484 |
fleetd #551 rework (comment 17058): fix TIMED_OUT_QUEUED javadoc for ATTEMPTED
Javadoc-only change. #551 added Injector.Cancellation.ATTEMPTED, which MessageService.send's timeout path folds into Outcome.TIMED_OUT_QUEUED alongside CANCELLED and NOT_DELIVERED (the collapse itself is unchanged behaviour and is being tracked as a separate follow-up ticket). TIMED_OUT_QUEUED's javadoc — landed by #513 to state the routes that reach it — named only two routes and said "on every route it will not arrive later", with a "may already hold a partial paste" caveat. Both claims are now stale: ATTEMPTED is a third route, and because agent.prompt pastes AND submits in one call, that route may mean the target holds a complete, already-submitted turn and is working on it right now. Names all three routes, says which one is uncertain, and drops the now-false blanket claim. No behaviour change. |
||
|
|
e20ccab1eb |
fleetd #561: harden the completion/session TurnListener fan-out
Fleetd's turnListener composition had four callbacks (onTurnComplete, onTurnCompleteWithPostAction, and both onTurnFailed overloads) built from two bare, unguarded statements each. onDelivered's registration was already fixed structurally by #556; these four had the identical fragility and were still untested: nothing enforced that the completion resolver's half ran before the session half beyond call order in the source, so a future reorder (or a throwing session listener sequenced first) could silently skip the completion resolver's effect and strand a caller for its full timeout. Extracted the composition to a package-private static factory, Fleetd.turnListener(completion, sessions), and hardened it with bothMustRun/bothMustRunKeepingSecondResult: both callback halves are always attempted regardless of whether the other throws, and whatever escapes is rethrown afterward (never swallowed) so it still reaches StatusPoller's catch (Throwable) and logs at ERROR. FleetdTurnListenerCompositionTest builds this real composition from a real CompletionResolver and a throwing fake sessions half, and asserts the completion resolver's effect (the waiter resolving) survives the session half throwing, for all four callbacks, plus a mirror case showing the session half still runs when the completion half throws first. onTurnCompleteWithPostAction keeps completion-before-session as a functional requirement (resolveBeforePostAction must run before the context-reset housekeeping can erase the pane), not just fault tolerance, so it is not reorder-symmetric like the other three — documented in Fleetd.turnListener's javadoc. |
||
|
|
d83821bbce |
fleetd #551: record the delivery attempt before the irreversible send
The Injector wrote the delivery outcome AFTER calling AgentControl.send(), so a HerdrException thrown from the response half of that call (herdr already replied, or may have) was recorded as a confident NOT_DELIVERED for text that may already be sitting in the worker's pane. Poll the queue entry and mark it Pending.State.ATTEMPTED before send() is called, not after. On success it is upgraded to DELIVERED; on an ordinary failure it stays ATTEMPTED (honest uncertainty), except a herdr *_not_found error, which the rest of this codebase already treats as a confirmed absence and which now still writes NOT_DELIVERED. cancellationOf gets a matching third answer (Cancellation.ATTEMPTED) instead of folding the new state into NOT_DELIVERED, so a caller that cancels an already-attempted delivery is told the truth too. |
||
|
|
ba2f4d16f8 |
Merge #555: the redeploy main flow is lifted into tested predicates, and the guard now catches functions below the SOURCED line
fleetd #555. Eight main-flow decisions in scripts/redeploy-fleetd.sh move into predicate and dispatch functions the suite can source and test. The guard test test_no_untested_main_flow_conditionals stops new bare conditionals reappearing. The rework closes a hole the lead found (comment 17012): a conditional wrapped in a function defined BELOW the SOURCED guard was invisible to the guard, and such a function can never be sourced, so it can never be tested. The guard now fails on any function definition after the guard line, on its own. Verified by the lead on a tree with main merged in, redeploy-fleetd.sh at its pristine sha 4ffacc5185807d39720a3484d85b922413806eb5347318265bd8897dfd61e8d9: bash 5.3.9 -> exit 0, 0 lines matching ^FAIL: /bin/bash 3.2.57 -> exit 0, 0 lines matching ^FAIL: Three mutations, each restored to the pristine sha afterwards: a bare conditional appended to the main flow -> exit 1, reported by line the same conditional wrapped in a function below the SOURCED guard (the found hole) -> exit 1, "function defined after the SOURCED guard (line 1038) - it cannot be sourced, so it cannot be tested" the new FUNC emission deleted from the guard -> that same case returns to exit 0, so the new assertion is what catches it. Anchor count 1 -> 0, test-redeploy-fleetd.sh restored to sha 0d713a3092a0c0ea8c05663ffa0cb1595e21b9d73872e662eb99b5035fd38151. |
||
|
|
db4c98ac60 |
Merge #556: the Injector owns turn registration, and the #553 backstop is pinned too
fleetd #556. Registration moves off the TurnListener fan-out onto its own narrow
TurnRegistrar seam, wired directly to CompletionResolver::register in Fleetd.java,
so it survives any listener throwing regardless of call order.
Verified by the lead on a tree with main (
|
||
|
|
8f80d267a0 |
fleetd #555 rework: catch function definitions after the SOURCED guard
Comment 17012 on #555 found a hole in test_no_untested_main_flow_conditionals: the guard's function-body detection treats anything inside a function as "fine, out of scope for this scan" — but a function DEFINED after the SOURCED guard line can never be reached by sourcing this script (sourcing stops before the main flow runs), so its body is untestable by construction while still reading to the guard as safely inside a function. mainflow_bare_conditionals now also emits a FUNC record for every function opened after the guard line (reusing the same open-brace detection already used for depth tracking), and test_no_untested_main_flow_conditionals treats any such record as a violation on its own, independent of what the function's body contains or whether the allowlist would otherwise excuse a bare conditional inside it. Proof (redeploy-fleetd.sh restored to 4ffacc5185807d39720a3484d85b922413806eb5347318265bd8897dfd61e8d9 after each): - CONTROL — a bare conditional appended to the main flow is still caught: EXIT=1, "found 1 untested main-flow if/elif/case line(s) ... line 1341: if [ "$MY_CONTROL_BARE" = 1 ]; then :; fi" - CANDIDATE — the same conditional wrapped in a function defined after the boundary, previously invisible (EXIT=0), is now caught: EXIT=1, "line 1341: function defined after the SOURCED guard (line 1038) — it cannot be sourced, so it cannot be tested: newfunc_below_the_boundary() {" Full suite re-run green on both bash 5.3.9 and /bin/bash 3.2.57 (macOS system bash): exit 0, 0 FAIL lines, reached the final PASS line, on both. No change to redeploy-fleetd.sh; the 8 lifted decisions, their mutation proofs, and the allowlist all stand as before. |
||
|
|
738d34a609 |
fleetd #556 rework: pin registration on the #553 finally backstop path
Comment 17009: there are TWO registrar.register(target, sent.token()) call sites in Injector's delivery method — the ordinary path inside `if (sent != null)`, and the fleetd #553 finally backstop, reached only when an earlier block throws before the ordinary path ever runs. The lead's mutation on Injector.java:703 (the backstop call) survived the full suite: the existing #553 regression test for this exact scenario (aRuntimeExceptionFromOnTurnCompleteStillCompletesTheNextDelivery) asserts only that the delivered future completes, never that the turn is registered with CompletionResolver — so a redesign that dropped registration from the backstop would reopen this ticket's own defect on precisely the path #553 exists for, with every existing test green. Adds aRuntimeExceptionFromOnTurnCompleteStillLeavesTheNextDeliveryRegisteredOnTheRecoveryPath: drives the same construction as the existing #553 test (onTurnComplete throws for a previous turn, forcing the next turn's delivery down the finally backstop) and additionally asserts the new turn is registered with CompletionResolver and carries the correct waiter — the same assertion the ordinary-path test makes, now made on the recovery path. Proven by mutation: removing Injector.java:703 alone (exact-line anchor 1 -> 0) turns the new test red with its own assertion message; restored and confirmed byte-identical (sha256 8fcb698afccc254b0c99d3a4bf9e960c0e7c85e542024e870c1e815dce6d62a9, matching the pre-mutation tree); re-run green as a control. Full suite after restore: 1750 tests, 0 failures, 0 errors, 0 skipped (Maven's own summary and an independent sum over surefire-reports/*.txt agree), BUILD SUCCESS. |
||
|
|
a46e4058ac |
fleetd #556: make turn registration structural, independent of any TurnListener
The Injector owns the invariant "every delivered turn has a registered waiter," but before this the only thing that satisfied it was CompletionResolver.captureBaseline, called from inside a TurnListener callback wired in Fleetd.java. Any TurnListener that throws (from onDelivered or elsewhere) could break the invariant with no way for the Injector to detect it. #553 only made the one reachable listener behave via a try/finally backstop; it did not remove this structural dependency. Add a narrow TurnRegistrar functional interface, decoupled from TurnListener, whose only job is registering a delivered turn's waiter. CompletionResolver now implements it via a new register() method (extracted from captureBaseline's registration half; captureBaseline keeps its own full body unchanged, so existing direct callers/tests are untouched). Injector gets an explicit registrar field/constructor family (auto-derived from the TurnListener via instanceof where that still works, explicit where Fleetd's anonymous fan-out listener can't implement two interfaces at once) and calls registrar.register(...) directly and unconditionally in both the ordinary delivery path and the #553 finally backstop, before turnListener.onDelivered(...) — so registration no longer depends on that notification callback succeeding. Fleetd.java wires completion::register explicitly as the registrar, bypassing the turnListener fan-out for registration purposes. CB-116 ordering (onTurnComplete reads the PREVIOUS turn's inFlight entry before the new turn's registrar.register() runs) and the two-arg inFlight.remove(target, turn) vs one-arg distinction on the completion path are both preserved unchanged. #561's order-dependent test (asserting Fleetd.java's completion.onDelivered -> sessions.onDelivered call order) does not exist anywhere in this repo at this branch point — nothing to delete. Adds 5 tests: a TurnListener that throws from every callback still leaves the delivered turn registered and resolvable; the new turn's registration still runs after the previous turn's completion is read (CB-116 guard, pinned as a call-order assertion); and three tests naming the two-arg-remove invariant directly (a superseded turn's terminal handling must not evict its successor's registration) across resolve()'s plain-completion branch, its echoed-noReportMessage sub-path, and fail(). |
||
|
|
f4f5f3106e |
Merge #513: TIMED_OUT_QUEUED has four routes, and the javadoc now says which
The old javadoc said a TIMED_OUT_QUEUED message was "still sitting in the injector's
per-target queue". It is not — send() has already given up on it and it will never
arrive. That wrong claim told an operator to wait for a message that was never coming.
The first fix replaced it with a narrower wrong claim: that Injector.cancel() cancelled
the entry and "the target never saw a word of it". That describes one of four routes.
TIMED_OUT_QUEUED is returned whenever injector.cancel() returns anything but DELIVERED:
CANCELLED the Pending was still queued and this call removed it
NOT_DELIVERED Injector.java:400 - the herdr agent.prompt call threw
NOT_DELIVERED Injector.java:419 - readiness grace expired, never attempted
NOT_DELIVERED Injector.java:691 - drop(), the target is gone
On the last three, cancel() cancels nothing: it reads a state another path already set
(cancellationOf, Injector.java:288-290). And on the herdr-threw route, agent.prompt
pastes and submits in one call, so a throw does not prove the pane stayed clean - an
operator told "never saw a word" will not go and look at the one place the evidence is.
The javadoc now states the two facts that hold on every route - the message will not
arrive later, and it is not in any queue - and attaches "the target saw nothing" only to
the CANCELLED case. Five blocks: the enum constant, queuedDeliveries, hasQueuedDelivery,
hasOrphanedDelegation, and the inline comment in the TimeoutException branch that seeded
the wording.
Comment-only; no logic changed.
Verified on a tree merged with main (fast-forward to
|
||
|
|
8d79d229ff |
fleetd #555: lift 8 main-flow decisions into tested predicate/dispatch functions
redeploy-fleetd.sh's main flow had 8 bare if/case decisions (CHECK_ONLY
short-circuit, drain-gate entry+confirm, supervisor report/stop/start
dispatch, health-poll decision, HAD_OLD_PID computation) that lived outside
any function, so the 67-test suite could not reach them and any one could be
silently inverted with the whole suite green.
Follows the existing swap_if_built/refuse_drain_gate pattern: each bare
guard becomes a small predicate or dispatch function (should_stop_for_check,
drain_gate_required/drain_confirmed/run_drain_gate, report_supervisor_state,
dispatch_stop, dispatch_start, health_is_up/report_health,
compute_had_old_pid), called unconditionally by the main flow so the
decision itself is unit-testable in isolation.
Adds a structural guard, test_no_untested_main_flow_conditionals, that scans
the main flow (everything after the SOURCED guard) for bare if/elif/case
lines outside any function body, tracking function boundaries via this
file's one consistent name() { / } convention. It fails on any new bare
conditional not covered by MAIN_FLOW_ALLOWED_CONDITIONALS, an explicit
exact-text allowlist of the report-only/display conditionals and the two
#504-family supervisor elif branches that stay out of scope for this
ticket. This is the "shape, not the eight sites" guard the ticket asked
for: a ninth bare decision fails immediately, naming its line.
Out of scope, not touched: #504 items 2/3/4 and #528 item 2 (same
untested-main-flow family) — the seam here generalizes to make them
testable too, but lifting them was left for their own tickets.
|
||
|
|
cb64bc8157 |
fleetd #513: rework — TIMED_OUT_QUEUED has four routes, not one
Comment 16984 on the ticket showed my first pass (
|
||
|
|
ed28b51f12 |
fleetd #513: fix TIMED_OUT_QUEUED javadoc — cancelled, not queued
Two javadoc blocks (queuedDeliveries field, hasQueuedDelivery) said a timed-out message is still sitting in the injector's per-target queue. CB-640 made send() cancel it via Injector.cancel() instead, so the message is gone and will never arrive. Rewrote both to describe cancellation. Also fixed a third instance of the same stale claim in hasOrphanedDelegation's javadoc, and added a line to the TIMED_OUT_QUEUED enum constant's own comment clarifying the name is kept but no longer means the message stays queued. Per the ticket's follow-up comment: no rename (TIMED_OUT_QUEUED reaches FleetApp.java REST mapping and FleetMcp.java — out of scope here) and no behavior change; comments only. |
||
|
|
4f9aba40e7 |
Merge #558: the ticket is the pull channel, and the brief is write-once
Docs only, 16 additions / 3 deletions — read by the lead in full, per the under-50-lines self-review rule in CLAUDE.md. Verified on the merged tree (main |
||
|
|
dab697fae0 |
Merge #559: progress watchdog for StatusPoller and SessionReaper loops (fleetd #544)
Verified by the lead on a merged tree (main |
||
|
|
bfac14108f |
fleetd #544: pin the sticky-STOPPED-across-restart invariant
Review of PR #559 (issue comment #16944) found a surviving mutant: removing watchdog.reset() from StatusPoller.start() (and the identical line in SessionReaper.start()) passed the entire suite. stoppedByCaller is sticky and reset() — called only from start() — is the only thing that clears it. Both loops document start() as idempotent and loop()'s own error log says "it can be restarted", so stop() followed by start() is an anticipated path. Without reset() wired into start(), health() would report STOPPED forever after a restart even though the loop is genuinely running again. Add aRestartedLoopReportsRunningAgainNotStoppedForever to both StatusPollerWatchdogTest and SessionReaperWatchdogTest, pinning "an intentional stop must not outlive the restart that follows it". Verified via the standard mutation cycle: exact-line anchor (not regex, to avoid the \Q-style false match the reviewer flagged) counted pristine 1 -> mutated 0, test goes red with its own message, restored, shasum -a 256 byte-identical, green again. mvn clean install: exit 0, BUILD SUCCESS, Tests run: 1734, Failures: 0, Errors: 0, Skipped: 0 (cross-checked against 130 surefire report files). No production code changed — the reset() call under test was already correct; it simply had nothing pinning it. 🤖 Generated with Claude Code Co-Authored-By: Claude <noreply@anthropic.com> |
||
|
|
7a3b2bb7ee | Merge pull request 'fleetd #552: warn instead of aborting when the post-restart mktemp fails' (#560) from worker/552-post-restart-mktemp-abort-bc2672-4 into main | ||
|
|
f188947750 |
fleetd #552: warn instead of aborting when the post-restart mktemp fails
By the time the fresh-log mktemp ran, the daemon had already been stopped, the jar swapped, and the new daemon started — an unguarded mktemp failure there aborted the whole script anyway, so a caller read the resulting non-zero exit as "the redeploy failed" and would restart an already-correctly-restarted daemon. Extract the mktemp into capture_fresh_log_region, guarded the same way unload_launchd_if_loaded/stop_systemd_if_loaded guard their own, but warn instead of die: there is nothing left to protect by refusing after a successful restart. The trap is now installed before the assignment it cleans up, using an FRESH_LOG="" sentinel readers can check. classify_amqp_connection_errors and report_shutdown_drain both gain a new state (REDEPLOY_AMQP_CHECK_SKIPPED / REDEPLOY_DRAIN_STATE=skipped) for an uncapturable log region, distinct from "captured a region with nothing in it" — and the result section gains a matching branch, so a skipped capture can never read as a clean bill of health. |
||
|
|
f606fccf7f | Merge pull request 'fleetd #553: register the rendezvous waiter in onStatus's finally backstop' (#557) from worker/553-onstatus-completion-leak-0da881-2 into main | ||
|
|
735b6af976 |
fleetd #544: progress watchdog for StatusPoller and SessionReaper loops
Each loop's virtual-thread runner (StatusPoller, SessionReaper) can die or get permanently parked in a herdr call with no read timeout, and nothing observed it: /healthz stayed green and Thread.isAlive() kept reporting true the whole time. Add LoopWatchdog (dev.ltms.fleet.inject — see its javadoc for why not dev.ltms.fleet.health, which would close a package cycle through session): each loop now records a monotonic last-completed-round timestamp (injectable LongSupplier clock, same pattern as Injector/SessionManager) and exposes it as a three-state health() fact — RUNNING, STALLED (dead or parked, indistinguishable from outside), STOPPED (stop() was called on purpose, never an alarm). This is the fleetd #512 shape: one flag cannot carry both "halted on purpose" and "halted unexpectedly", so stop() marks its own state explicitly instead of leaving state() to infer it from staleness. Staleness thresholds are derived from each loop's own poll interval with a documented multiplier: StatusPoller 40x (250ms -> 10s), SessionReaper 12x (5000ms -> 60s). Scope: observability only, per the ticket's own comment. No restart/recovery mechanism, no /healthz or REST/MCP wiring beyond the public health() API, no change to the per-item catch(Throwable) behavior (#543) or a process-wide uncaught-exception handler (ruled out on #538), and no deadline added to the herdr read itself (a separate, real ticket). 🤖 Generated with Claude Code Co-Authored-By: Claude <noreply@anthropic.com> |
||
|
|
8b4320ed24 |
fleetd #553: split sentHandled's two meanings so onDelivered's own throw still completes the future
Lead review of PR #557 (ticket comment 16916) found one path left open: sentHandled is set to true BEFORE onDelivered() runs (correctly, per the earlier fix), so when onDelivered() itself throws on the normal path, the finally's 'if (sent != null && !sentHandled)' guard skipped the whole recovery -- completion included -- and left sent.delivered() pending forever for a message that really was delivered. sentHandled must guard only the onDelivered RE-CALL (the permanent-suppression hazard), never the future completion, since CompletableFuture.complete/ completeExceptionally are idempotent and a no-op on the already-handled path. Split the one flag's two jobs: the outer 'if (sent != null)' now always runs the recovery block, and '!sentHandled' moved onto just the onDelivered call inside it. Added anOnDeliveredThrowOnTheNormalPathStillCompletesTheDeliveryFuture, proven with the lead's own mutation (reverting !sentHandled onto the outer if): the new test goes red while anOnDeliveredThrowAfterItsOwnRegistrationDoesNotRunASecondTime stays green, showing the two concerns are genuinely separate. |
||
|
|
351ee1ea6d |
CLAUDE.md: the ticket is the pull channel, and the brief is write-once
A send to a working member is accepted and returns a ticket, then is never delivered. That happened three times in one session here, and the member was released still executing a brief that had been retracted twice. The receipt is true — it is a fact about the mailbox, when what was needed was a fact about the pane. The fleet01 lead named the mechanism: a push delivery needs the recipient free at send time, while a pull channel needs only that they look before acting. So the ticket is not more reliable than the mailbox, it is a different direction, and its success depends on the member's procedure rather than on the timing of the send. Both halves have to be written down, because each is useless alone: - Member (turn contract, new item 4): re-read the ticket before acting on anything told earlier, and again before committing. A ticket comment that contradicts the brief is newer and wins. - Lead (step 5): all corrections go to the ticket, and the brief is write-once. The member cannot check which source is newer — it just always prefers the ticket — so revising a brief in place makes it obey the rule and do the wrong thing. A first brief for a unit not yet running is not a correction. wiki/7-Use-Cases.md is updated to keep the canonical block byte-identical; the sync check passes. The wiki submodule pointer is deliberately left unstaged. |
||
|
|
d4a51c6274 |
fleetd #553: register the rendezvous waiter in onStatus's finally backstop, not just the delivery future
The previous try/finally around onStatus's post-monitor region completed sent.delivered() but never registered sent.token().waiter() when an earlier listener threw. That waiter is registered only by turnListener.onDelivered(), inside the very if (sent != null) block the finally backstops, so a caller was told its send landed and then waited out its full timeout for an answer that could never resolve (worse than a plain hang). The finally now does that block's whole job on the unhandled path: it calls onDelivered() (when sendError == null) before completing the future, guarded by its own try/catch(Throwable) so a failure there cannot mask the original throwable. A sentHandled flag, set true at the START of the normal block (before any side effect), tells the finally whether that already ran, so a throw partway through onDelivered cannot trigger a second, late captureBaseline that would permanently suppress the turn's completion. Also removed a leftover duplicated forget.accept(target) call (with a stray 'MUTATION-TEST-3' comment) in the notReady block — residue from the previous worker's own mutation testing that was not fully reverted. |
||
|
|
26f380a00b |
Merge #554: portable hash256, a third state for an unhashable jar, and the shell suite in CI (#550)
Closes fleetd #550, all three items. Verified by me on the branch at |
||
|
|
b8182c96c2 |
fleetd #550: pin hash256's algorithm against a literal SHA-256 test vector
test_jar_id_defaults_to_live_and_reports_explicit_path's reference hash is computed by calling hash256 itself (needed so it doesn't call the Linux-crashing bare shasum directly). That made subject and reference the same instrument: they agree no matter which algorithm hash256 actually runs, so a mutation swapping both of hash256's arms for the wrong algorithm was invisible to the suite. Adds test_hash256_computes_a_real_sha256, pinned against the published SHA-256 test vector for the 3-byte input "abc" (ba7816bf8f01...), written as a literal constant rather than computed by any hasher at test time. Verified the constant myself both ways (sha256sum and shasum -a 256) before writing it in. |
||
|
|
3da44eed63 |
fleetd #550: replace shasum with a portable hash256 helper, add a Linux CI job for the shell suite
jar_id() in redeploy-fleetd.sh called shasum directly, which does not exist on GNU coreutils Linux (Debian/Ubuntu/etc.) — there it silently reported an existing jar as "absent" with exit 0, because the missing command made `cut` succeed on empty input and pipefail's failure was then swallowed by the `|| echo "absent"` fallback. The shell test suite hit the same tool at test-redeploy-fleetd.sh:298-299 and died at exit 127 with zero FAIL lines printed — the same shape as a clean pass on the one channel anyone would check. Adds one hash256() helper (prefer sha256sum, fall back to shasum -a 256, same idiom already used in probe-member-credentials.sh) and points jar_id and the test suite's own reference hash at it. jar_id now has three distinct answers instead of two: absent, a hash, or "unhashable" when neither hasher is on PATH — "absent" is never used for a file that exists. Adds a CI job (shell-tests) that runs scripts/test-redeploy-fleetd.sh on ubuntu-latest, gated on the step's own exit code rather than a FAIL-line count, since a suite that dies before running is exactly what a green run also looks like by that count. New tests: test_jar_id_reports_unhashable_when_no_hasher_on_path (stubbed PATH with neither hasher) and test_no_unguarded_macos_only_hasher_calls (a shape check across every script under scripts/, not named lines — #545 already showed this idiom spreading from two sites to six). |
||
|
|
93a9ed3f83 |
Merge #549: widen Injector's delivery catch to Throwable (#546)
Closes the re-delivery window that merging #543 opened. I caused that; this closes it the same day. Verified by me on the branch at |
||
|
|
0b032f5a1a |
Merge #548: fix mktemp -t templates for GNU coreutils, split the unclear-supervisor detail (#545)
Verified by me on the branch, not on the worker's report.
Source audit, on a scratch worktree at
|
||
|
|
fad99c4c5e |
Merge #547: record the finally non-goal on drainAll's completion line
Comment only. No behaviour change.
Verified by me before merging:
- `mvn -B clean install` in a scratch worktree: exit 0, Tests run: 1716, Failures: 0, Errors: 0,
Skipped: 0. Same count as main, as expected for a javadoc-only change.
- #459's javadoc reference gate: exit 0, 0 reference errors.
- Gitea CI run 1801 on
|
||
|
|
87871eaefb |
fleetd #546: widen Injector's delivery catch to Throwable, stop re-delivery on Error
Injector.java:391 caught only RuntimeException around the herdr send seam. PR #543 (fleetd #538) widened StatusPoller's per-target catch to Throwable so the polling loop now survives an Error there, which means it comes back round — and Injector's narrower catch let the poisoned message stay QUEUED (the loop peeks, not polls), so the next round re-sent the same text into the member's pane. Widen the catch to Throwable, matching #543 one layer down. sendError's declared type widens from RuntimeException to Throwable to keep compiling; its only consumer (CompletableFuture.completeExceptionally(Throwable)) already accepts that type, so no other caller-visible behavior changes. The ordinary HerdrException/RuntimeException path is unchanged. Adds three tests: an Error at the send seam is dropped and marked NOT_DELIVERED, a second onStatus round does not re-send it, and a HerdrException control proves the ordinary path is untouched. |
||
|
|
a476a14f1c |
fleetd #545: fix mktemp -t templates for GNU coreutils, split unclear-supervisor detail
Every mktemp -t template in redeploy-fleetd.sh lacked an X placeholder. BSD mktemp (macOS) tolerates that and appends its own suffix; GNU mktemp (every Linux distribution) refuses it and exits non-zero. All six sites now use .XXXXXX. detect_supervisor's 'unclear' detail used to cover two different facts with one message that always named 'systemctl exited non-zero and reported an error on stderr' — even when systemctl was never run, because mktemp failed first. The SYSTEMD_LOADED_ERRORED/SYSTEMD_INSTALLED_ERRORED flags now carry a third value (2 = the probe's own mktemp setup failed) alongside the existing 1 (systemctl ran and answered badly on stderr), and detect_supervisor gives each its own detail text. kind stays 'unclear' in both cases; require_drivable_supervisor is unchanged. Tests added to scripts/test-redeploy-fleetd.sh: - test_mktemp_dash_t_templates_have_x_placeholders: source-text check, fails if any mktemp -t template lacks an X. - test_detect_supervisor_systemd_probe_setup_failure_is_unclear: proves the SET-UP-FAILED detail when mktemp itself fails (systemctl never runs). - test_detect_supervisor_systemd_probe_error_is_unclear: extended with assertions that the PROBE-ANSWERED-WITH-STDERR detail is present and the SET-UP-FAILED wording is absent, so swapping the two messages fails a test in both directions. |
||
|
|
5eb4267a4a |
fleetd #512 follow-up: record why the drain-complete line must not move into a finally
The javadoc said the log.info fires "every time", which reads as an invitation to the exact edit that destroys it. The absence of the line is the signal that the drain died, so a finally would remove the signal and print partial counts in the same change. Raised by the fleet01 lead from their 2026-09-10 incident: that drain is known to have died only because it threw and left a stack trace. A drain that hung, or returned early on a condition, leaves no trace, no ERROR token and no priority -- only a missing line. Comment only. No behaviour change. |
||
|
|
7611b69667 | Merge pull request 'fleetd #504 item 1: stop the false ok on the loaded-but-not-running stop path' (#541) from worker/504-failed-reported-clean-3cfd66-3 into main | ||
|
|
cc302fe4af | Merge pull request 'fleetd #538: recover polling loops after errors' (#543) from worker/538-loop-dies-on-error-4a5eeb-6 into main | ||
|
|
4a8a780274 | Merge pull request 'fleetd #426: pin FleetHealthMonitor.coverage and its HealthCoverageSource call site' (#542) from worker/426-health-coverage-ef1fd4-4 into main | ||
|
|
343ce0f4c0 | fleetd #538: recover polling loops after errors | ||
|
|
cec3e191d4 | Merge pull request 'fleetd #459: lint Javadoc references in CI' (#539) from worker/459-broken-link-targets-cadc17-5 into main | ||
|
|
1850a5f324 |
fleetd #426: pin FleetHealthMonitor.coverage and its HealthCoverageSource call site
FleetHealthMonitor.coverage had zero references in the test tree — not the method, not either output string, not the field it populates. Inverting `enabled`, swapping "full"/"detection-only", or breaking the argument pairing at the HealthCoverageSource call site in Fleetd.java all shipped a green build. Extract the HealthCoverageSource lambda out of Fleetd.main into a package-private static factory (healthCoverageSource(ConfigRef)), the same shape capacitySource/quarantineSource already use for the identical argument-pairing risk (fleetd #415). #407's "keep the config invalid, assert on the log line before validateAll() throws" option does not apply here: this call site is built well after validateAll() and after a real herdr socket connect, so driving it through a real Fleetd.main would require the socket I/O this ticket's tests must not do. Add FleetHealthMonitorCoverageTest (the three-branch method itself) and FleetdHealthCoverageSourceWiringTest (the call site, via a real FleetConfig.load + ConfigRef against @TempDir fixtures, including a hot notifications-reload case). Output strings are unchanged — "detection-only" is still what a live fleet_list reports today. Measured: all three mutations killed by the new tests. |
||
|
|
e4eb3dbed4 |
fleetd #504 item 1: stop swallowing real launchctl/systemctl failures on the 'loaded but not running' path
The two 'loaded but not currently running' branches in the stop step (launchd/systemd, reached when $OLD_PID is empty) ran 'launchctl unload'/'systemctl --user stop' with '2>/dev/null || true' and printed 'ok' unconditionally. That swallowed a real supervisor failure (e.g. launchd or the systemd user bus unreachable) exactly like a harmless already-stopped answer, and let the script proceed to start a new daemon believing nothing was loaded -- the two-daemons failure fleetd #492 exists to prevent. Adds unload_launchd_if_loaded/stop_systemd_if_loaded, applying systemd_loaded's own pattern (capture stderr separately; a non-zero exit WITH stderr is a real failure, a non-zero exit with empty stderr is a clean already-stopped answer) to the write side. The two call sites now use these functions instead of the bare '|| true'. Adds 5 tests: dies-on-real-failure and tolerates-clean-negative for each function, plus a source-text check that the main flow calls the new functions instead of the original bare '2>/dev/null || true'. All 5 verified by mutation (reintroducing the swallow, and separately over-correcting to die unconditionally) -- each goes red with its own message, restores byte-identical (full sha256), and passes a green control. |
||
|
|
57cd96f5e6 | Merge pull request 'fleetd #537: pin CapturedLog.close()'s appender-detach and setLevel-immunity contracts' (#540) from worker/537-capturedlog-close-e4c437-2 into main | ||
|
|
202e37e3b3 |
fleetd #537: pin CapturedLog.close()'s appender-detach and setLevel-immunity contracts
Only the level-restore half of close() was pinned before this (WorktreeSessionManagerTest). Deleting logger.detachAppender(appender) from close() left mvn clean install green (1701 tests, 0 failures) -- the appender-detach half of the contract was unmeasured. Adds CapturedLogTest with three tests, each using a logger name no production class uses: - closeDetachesTheAppenderSoALaterLogIsNotCaptured: an event logged after close() must not land in events(). - closeRestoresTheLevelCapturedAtOpen: the helper's headline contract in one place, independent of any production class. - setLevelDuringCaptureDoesNotChangeWhatCloseRestores: setLevel()'s own javadoc claim that close() always restores the level captured at construction, never a value set through setLevel() mid-capture. Test-only change; CapturedLog.java itself is untouched. |
||
|
|
90253f832d | fleetd #459: lint Javadoc references in CI | ||
|
|
f1640f5dcc | Merge pull request 'fleetd #535: convert FleetdLeadMailboxSelectionTest to CapturedLog' (#536) from worker/535-appender-leak-fe74c1-1 into main | ||
|
|
c7903c1efe |
fleetd #535: convert FleetdLeadMailboxSelectionTest to CapturedLog
Three call sites (captureFleetdLogs at :55) attached a ListAppender to the Fleetd.class logger with addAppender and never detached it, and never called appender.setContext(...) either. Logback Logger instances are cached per class and shared for the whole JVM, and surefire reuses forks, so all three appenders stayed attached for every later test in the fork. Convert all three call sites to CapturedLog.of(Fleetd.class) (added in #533) via try-with-resources, which detaches the appender and sets the context for free. Delete captureFleetdLogs(); nothing calls it now. |
||
|
|
7d711942fe |
Merge #534: detect a died shutdown drain the ERROR count is blind to (fleetd #512 part 2)
Verified independently. The branch is based on |
||
|
|
bec87f987c |
scripts: name the mechanism in detect_supervisor's constraint 2, not a line number
Constraint 2 read "This script runs under `set -euo pipefail` (line 50), so an unset variable is a loud failure." Two problems, both small and both the same family as fleetd #494 — a comment that states the wrong reason. The line number was stale: the `set` line is at 54, not 50. It was the only line-number citation in the file, and a citation like that goes stale on the next insert above it, silently, with nothing to catch it. The mechanism was also misattributed. What makes an unset variable a loud failure is `set -u`. Naming the whole `-euo pipefail` string invites the reader to credit pipefail for it, which is the mistake fleet01 flagged on a different cell this week: pipefail is insurance against a future pipeline stage, not what catches the current shape. Now names `set -u` and says where it is without a number, and records why the number is gone so nobody adds one back. Comment only. bash -n exit 0 under /bin/bash 3.2.57 and env bash 5.3.9; scripts/test-redeploy-fleetd.sh exit 0 with 0 lines matching ^FAIL:. |
||
|
|
7f8a8829f9 |
fleetd #529 follow-up: the new helper's javadoc claimed a reach it does not have
CapturedLog's class javadoc said it is "the one way to pin or capture a logger's
level and output in this test tree". Measured on main at
|
||
|
|
af9589783e |
Merge #533: promote CapturedLog to a shared test helper and close the logger-level leak (fleetd #529)
Verified independently in a scratch worktree at |
||
|
|
190436c9cf |
fleetd #512 part 2: detect a died shutdown drain the ERROR count is blind to
The previous daemon's dead shutdown drain (an uncaught exception in a shutdown thread) never passes through the logger, so it never carries an ERROR/SEVERE token, so redeploy-fleetd.sh's existing ERROR-count classifier is structurally blind to it and prints a confident "no ERROR lines since restart" while the drain actually died. Add scan_uncaught_exceptions (greps the shutdown window for the failure's real shape: `Exception in thread`, `NoClassDefFoundError`) and find_drain_complete_line (checks for #522's SessionManager.drainAll completion line). Compose both in report_shutdown_drain, a single decision+action function the main flow calls unconditionally (same shape as swap_if_built/refuse_drain_gate from #521/#528), which resolves to one of four outcomes: complete, died, unknown ("cannot tell" — the line is absent for either of two reasons that need opposite handling: the previous daemon predates #522, or its drain failed without throwing), or n/a (no previous daemon was actually stopped this run). Never fails the redeploy; warns loudly instead. Gate the "no ERROR lines since restart" summary line on the new outcome so it never reads as reassurance when the drain died or the outcome is "cannot tell" (item 4 of the ticket). Tests: 11 new test functions (60 defined/invoked, was 49), covering both pure classifiers, all four report_shutdown_drain outcomes, a source-grep proof of the main-flow call site (sourcing stops before the main flow runs), an ordering check, and the item-4 gating. Full suite green (exit 0, 0 anchored FAIL lines). Five mutations applied and killed by hand during review, each restored to a byte-identical file afterward. |
||
|
|
8ea5c2bb1f |
fleetd #529: promote CapturedLog to a shared test helper, close the logger-level leak
ch.qos.logback.classic.Logger instances are cached per class and shared for the whole JVM, and surefire reuses forks. A test that pins a shared logger's level and restores only the appender leaves that level pinned for every test that runs after it, in the same class or a different one in the same fork. Move CapturedLog (merged in #527 for #525) out of SessionManagerTest into dev.ltms.fleet.testing.CapturedLog, and convert all 19 unrestored setLevel pins across 9 files to it, so there is exactly one way to capture and pin a logger in this test tree: - FleetdAwaitHerdrTest, FleetdReplyInboxSelectionTest, AuditLogTest, CompletionResolverTest, InjectorTest (4), LeadRolloverTest, AmqpConnectionFailureLoggerTest, GitWorktreesTest (8), WorktreeSessionManagerTest. AuditLog logs through a named "audit" logger rather than a class, so CapturedLog gains String-named at()/of() overloads alongside the existing Class-based ones, plus a setLevel() method so a fixture that already pinned a coarser baseline (GitWorktreesTest's @BeforeEach) can re-pin further for one test without losing what close() restores. Adds an ordered proving test to WorktreeSessionManagerTest asserting the SessionManager logger level is back to a known baseline after the dirty- worktree release test runs; this proves the within-class case only, since JUnit does not guarantee cross-class ordering. Every existing intentional pin (SessionManagerTest's two explicit INFO pins and its @BeforeAll DEBUG baseline) is left untouched, per the ticket. |