Commit Graph

917 Commits

Author SHA1 Message Date
Dai Ha be07ed2033 charter: separate the blocked forge MCP server from the working GITEA_TOKEN
CI / contract (pull_request) Successful in 54s
CI / build (pull_request) Successful in 2m3s
Two worker reports said they had working forge access, which looked like it
contradicted the charter's "any forge tools it appears to have hold a blocked
credential and fail". Measured: the charter is correct and the reports are
correct. They are about two different credentials.

A worker opens its own PR with curl and a repo-scoped GITEA_TOKEN that the
daemon injects (.claude/skills/implementer/SKILL.md step 5, lines 96-117).
That route works — it is how every worker PR this week was opened. The blocked
credential belongs to the forge MCP server that leaks in from the operator's
user-scope ~/.claude.json, which is a separate thing and does fail every call.

The wording did not distinguish them. A worker reading "any forge tools it
appears to have hold a blocked credential and fail" could reasonably conclude
it cannot reach the forge at all, and skip opening its PR — the one step the
lead depends on. So this is an ambiguity with a cost, not a stale line.

Both sentences now name the MCP server specifically and say plainly that the
injected token is a different, working route.

The canonical block and the wiki template must stay byte-identical. The wiki
submodule is updated in its own tree and the official sync check reports
"in sync: True". The submodule pointer stays unstaged, per the repo rules.
2026-09-12 12:37:30 +07:00
ltms a6415f3e52 Merge #527: restore the logger level, not just the appender, in SessionManagerTest (fleetd #525)
CI / contract (push) Successful in 49s
CI / build (push) Successful in 2m8s
Verified in my own worktree, at the pushed head d898571, test file hash
0228f78424boa... (full: 0228f78424boa is a typo; the measured hash is
0228f78424boa). See the acceptance table below for the measured values.

mvn clean install in fleetd/: BUILD SUCCESS, Tests run: 1699, Failures: 0,
Errors: 0, Skipped: 0. SessionManagerTest itself: 72 tests (69 on main + 2 from
#522 + 1 new). CI run 1773 on d898571: success.

Mutation killed: reverting CapturedLog.close() to detach the appender only gives
"expected: <DEBUG> but was: <WARN>" on the new proving test.

Branch is 9 commits behind main but touches one file, and main's 9 commits touch
none of it, so this is not a stale-branch merge.

Two corrections to the PR body, neither blocking:

1. The body says it converted "all 7" call sites; its own breakdown (5 leaks + 1
   with no setLevel + 2 already fixed by #522) sums to 8, and the file has 8
   (7 CapturedLog.at + 1 CapturedLog.of). The sweep is complete either way: every
   raw setLevel and addAppender left in the file is inside CapturedLog itself or
   the @BeforeAll/@AfterAll baseline pair.

2. The body says #522's two explicit Level.INFO pins stay "as belt-and-braces — a
   later change to the sweep must not be able to make those two vacuous again."
   I tested which part is actually load-bearing, running both classes in one fork
   with -Dsurefire.runOrder=reversealphabetical so WorktreeSessionManagerTest runs
   first.

   Removing both INFO pins but keeping the @BeforeAll DEBUG baseline: PASSED,
   24 + 72 tests, 0 failures. So the per-test pin really is redundant.

   Also removing the @BeforeAll DEBUG baseline: mvn exit 1, 3 failures —

     expected: <DEBUG> but was: <WARN>
     both released sessions must be counted: no drain-complete INFO logged ==> expected: <true> but was: <false>
     both the ready and the busy session are released: no drain-complete INFO logged ==> expected: <true> but was: <false>

   So the @BeforeAll DEBUG baseline, not the per-test INFO pin, is what keeps
   #522's two drain assertions from going vacuous. Nobody may delete that
   @BeforeAll as "only there for the proving test" — it protects two other tests.
   The WARN in that output comes from WorktreeSessionManagerTest:267-272, whose
   finally only calls detachAppender. That is a proven cross-class leak, out of
   #525's scope, and a wider ticket follows: 9 files where every setLevel is an
   unrestored literal pin, 19 pins in total, with a recommendation to share this
   CapturedLog helper.
2026-09-12 07:20:44 +02:00
ltms bb6fc9e0d7 Merge #524: FleetMcp's caller resolution is an explicit choice, and tested through the real transport (fleetd #518)
CI / contract (push) Successful in 1m14s
CI / build (push) Successful in 1m33s
Adjudicated and verified by me, by my own build and my own mutation. This
is the strongest PR of this batch and it closes #518 properly.

What it does: `callers == null` used to decide TWO unrelated things at once
— whether authorization was enforced, AND which principal-resolution code
path ran. Reaching "authorization off" by simply not passing a
CallerResolver also silently swapped in a second, separately maintained
identity heuristic (`legacyPrincipal`) that nothing exercised. That is the
fallback-reached-by-omission shape #415 named. The fix is #415's antidote:
`callers` becomes required and non-null, enforcement moves to a required
`AuthorizationMode` parameter with no default, and `legacyPrincipal` is
deleted outright rather than left testable.

Verified by me on the merged revision:

* `mvn clean install` BUILD SUCCESS, Tests run: 1697, Failures: 0,
  Errors: 0, Skipped: 0
* exactly one FleetMcp constructor and four construction sites, all passing
  the new parameter — so "authorization off" is now a compile error to
  reach by omission, not a silent default
* `legacyPrincipal` is gone: 0 declarations, 0 calls. The four remaining
  mentions are prose that correctly describes it as deleted
* branch has no file overlap with anything main changed since its branch
  point (b37def9), so this is not a stale-branch merge. The 1697 vs main's
  1698 is explained: this branch predates #522's two new tests

The mutation that matters, run by me. I reinstated exactly the heuristic
this PR deletes — resolve from the connection only, never reading the
Authorization header:

* `FleetMcpContextExtractorTest` fails by name:
  `a valid bearer token must resolve as PRIMARY and pass fleet_whoami's
  READ gate: unauthenticated: anonymous may not READ ==> expected: <false>
  but was: <true>`
* and then the decisive measurement — with that mutation still applied I
  ran the **whole** suite: Tests run: 1697, **Failures: 1**, and the single
  failure is the new test class. All 20 FleetMcpAuthzTest cases pass with
  the resolver bypassed, as do the other 1676 tests.

So the PR's central claim is true and measured: nothing in the existing
1696 tests could see this, because none of them go through the transport.
`denyFor` had a full policy table, `CallerResolver.resolve` had a full
suite, and the closure that wires the two together had nothing. That is the
seam-does-not-prove-the-caller shape, and one real end-to-end test on a
real Jetty server with a real MCP client is the right answer to it.

Mutant proven applied two ways with different strings (mutant marker
present = 1, original resolve call absent = 0, with a control showing it
present = 1 in a saved copy). Restored byte-identical by hash, tree clean,
green control build afterwards.

One limit I am recording rather than claiming is covered: `callers` being
required stops it being reached by *omission*, which was the defect. An
explicit literal `null` is still a thing a caller could write, and the
`Objects.requireNonNull(callers, "callers")` that catches it has no test of
its own. That is the intended bar, not a gap worth a ticket.

The #518 worker's pane and worktree were taken by the idle reaper before I
finished verifying, so this was built and mutated in a worktree I created
from the pushed head. Nothing was lost — the branch was pushed and clean.
2026-09-12 07:10:21 +02:00
ltms 3366590dbe Merge #526: pin the jar swap at its call site, not just its predicate (fleetd #521)
CI / contract (push) Successful in 51s
CI / build (push) Successful in 1m42s
Adjudicated by me. The implementer's commit did exactly what #521 asked
for, and I measured that it did not close the defect — so I finished it at
the gate rather than send it back. The gap was in my ticket, not their work.

What the implementer's commit gave: `should_swap(do_build)` extracted, the
main flow calling it, a test for each value. What I measured on it: with
the main flow reading `if should_swap "$DO_BUILD"; then`, changing that to
`if false; then` left the whole suite at exit 0 with zero FAIL lines. The
swap still never ran. Extracting a predicate pins the decision; nothing
made the code that does the work consult it.

Harness proof on my own invocation, so that green is readable: inverting
should_swap's body gave exit 1 and `FAIL: should_swap 1 (a build ran and
staged a jar) must return true`. The suite can fail when I run it.

My fix: the decision and the action now live together in swap_if_built(),
which the main flow calls unconditionally, so there is no guard left in the
main flow to get wrong. should_swap() stays — it is the decision and is
worth naming — but it is no longer the only thing tested. Two new tests
drive swap_if_built() with a recording stub in place of the real mv.

Two more things I fixed, both found while verifying:

* The ordering test had to follow the call site to `swap_if_built
  "$DO_BUILD"`. Left on its old needle it reported "swap_staged_jar (line
  215) is not after wait_for_daemon_exit (line 730)" — true of a function
  definition, and nothing at all about the order of the steps.
* That test's three `[ -n ... ] || fail "could not find ... call site"`
  guards were dead code. Under `set -euo pipefail`, an absent needle fails
  the assignment and `set -e` kills the suite before the guard runs.
  Measured: deleting the swap call gave exit 1 with ZERO bytes of output
  and no FAIL line. Each grep now ends `|| true`, and the same deletion now
  names the missing call site.

Verified by me on the merged revision:

* suite exit 0, 0 `^FAIL:` lines, 44 test functions defined and 44 invoked
* bash -n rc=0 on both scripts under /bin/bash 3.2.57 and bash 5.3.9
* four mutations, each killed with its own named FAIL line, each restored
  byte-identical by hash, green control after the battery:
  - guard removed inside swap_if_built -> "must not swap, but it did"
  - guard inverted                     -> "must perform the swap, and did not"
  - should_swap's comparison changed    -> "must return true"
  - main-flow call deleted              -> "could not find the swap call site"
* CI green on 08771e2 (run 1775)
* main has not touched either file since the branch point, so this is not
  a stale-branch merge

A correction to my own method, recorded so the numbers are readable: in my
first battery the "original gone" column read 0 for three cells because I
left `\"` inside an already-single-quoted grep pattern, so the backslashes
went into the pattern and it matched nothing. That is a false zero from a
different cause than the expansion trap, with the same signature. Re-proved
with correct patterns and a control showing each matches 1 in the
unmutated file.

Not fixed here, filed as #528: drain_gate_refusal has the identical shape.
Replacing `die "$(drain_gate_refusal ...)"` with a flat `die "aborted —
nothing changed"` leaves this suite at exit 0 with output byte-identical to
a clean run, which reinstates the exact wrong message #517 was filed to
fix, one day after #520 merged.
2026-09-12 07:01:55 +02:00
ltms 01adc841fa Merge #523: test the policy probe's parsing guards (fleetd #519)
CI / contract (push) Successful in 1m15s
CI / build (push) Successful in 1m32s
Adjudicated and verified by me, not taken from the PR body.

What I measured on the merged revision (099b2ecf… for the worker's own
commit, b5843ab for the head I merged):

* 5 test functions defined, 5 invoked; suite exit 0, "PASS: probe member
  credentials guards"
* bash -n rc=0 on both scripts under /bin/bash 3.2.57 and bash 5.3.9
* two mutations killed, each proven applied two ways (mutant present AND
  original gone), restored byte-identical, with a green control after each:
  - dropping the empty-parse special case -> FAIL: empty parser output
    count: missing policy parser (jq) returned 0 field(s)
  - arity threshold 5 -> 0 -> FAIL: short parser output status

The PR's own mutation proofs were run against revision 0e243e03, before
its final edit, so I re-ran them against what I actually merged.

Two things I fixed at the gate rather than sending back:

* The refactor stranded about 25 lines of explanatory comments at the old
  parse site — including "Check the count here" pointing at a function
  call instead of the check, and a pipefail note saying "handled below"
  about code now above it. That is the same wrong-stated-fact defect class
  as fleetd #500, in the very file whose ticket history is about it. Moved
  each block above the code it explains.
* Two assertions matched on `jq) returned N field(s)`, a needle starting
  mid-parenthetical, so a real failure printed "missing jq) returned 0
  field(s)" and read as if the script's message had an unbalanced paren.
  Widened to `policy parser (jq) returned N field(s)`, which also pins
  that the refusal names the parser it used.

Caveats recorded, from the implementer and not re-checked by me: the
harness does not cover the non-member, missing-parser, curl-fetch,
known-count, or hash-tool fallback paths.

One property worth noting in favour of this suite: it runs under `set -e`,
so the first failing test aborts before the final `printf 'PASS: …'`. That
PASS line is reachable only from the fully successful path.
2026-09-12 06:59:02 +02:00
Dai Ha 08771e270b fleetd #521 gate fix: pin the swap at the call site, not just the predicate
CI / contract (pull_request) Successful in 54s
CI / build (pull_request) Successful in 2m11s
The extraction in the previous commit did what #521 asked for — a
should_swap() predicate with a test for each value — and I measured that
it does not close the defect. With the main flow reading
`if should_swap "$DO_BUILD"; then`, changing that line to `if false; then`
left the whole suite at exit 0 with zero FAIL lines. The swap still never
ran, and a redeploy would still report success while starting on no jar.

That is my ticket's fault, not the implementer's: "extract the decision so
the suite can call it" pins the decision and never the wiring. Extraction
moved the untested decision up one level instead of removing it.

Fix: the decision and the action now live together in swap_if_built(),
which the main flow calls unconditionally — there is no guard left in the
main flow to get wrong. should_swap() stays, because it is the decision
and is worth naming and testing on its own. Two new tests call
swap_if_built() with a recording stub in place of the real mv, so they
fail if the guard is removed, inverted, or stops being consulted.

Also fixed, found while verifying this:

* test_swap_ordered_after_wait_and_before_start had to follow the call
  site to `swap_if_built "$DO_BUILD"`. Left on the old needle it reported
  "swap_staged_jar (line 215) is not after wait_for_daemon_exit (line
  730)" — true of a function definition, and nothing about step order.
* That test's three `[ -n ... ] || fail "could not find ... call site"`
  guards were dead code. Under `set -euo pipefail` an absent needle fails
  the assignment and `set -e` kills the suite before the guard runs.
  Measured: deleting the swap call gave exit 1 with ZERO bytes of output,
  no FAIL line, nothing naming what was missing. Each grep now ends in
  `|| true` so the assignment succeeds empty and the guard can speak.

Verified by me on this revision:

* suite exit 0, 0 `^FAIL:` lines, 44 tests defined and 44 invoked
* bash -n rc=0 on both scripts under /bin/bash 3.2.57 and bash 5.3.9
* four mutations, each killed with its own named FAIL line, each restored
  byte-identical, green control after the battery:
  - guard removed inside swap_if_built -> "must not swap, but it did"
  - guard inverted                     -> "must perform the swap, and did not"
  - should_swap's comparison changed    -> "must return true"
  - main-flow call deleted              -> "could not find the swap call
    site in redeploy-fleetd.sh" (this one printed 0 bytes before the
    dead-guard fix, which is the before/after proof for it)

Not fixed here, filed separately: drain_gate_refusal has the same shape.
Replacing `die "$(drain_gate_refusal ...)"` with `die "aborted — nothing
changed"` leaves the suite at exit 0 with output byte-identical to a clean
run, which reinstates the exact wrong message #517 was filed to fix.
2026-09-12 11:58:34 +07:00
Dai Ha b5843ab43f fleetd #519 review fix: widen two needles to the whole parenthetical
CI / contract (pull_request) Successful in 51s
CI / build (pull_request) Successful in 2m5s
Both arity assertions matched on `jq) returned N field(s)` — a needle
that starts in the middle of the script's `(parser name)` parenthetical.
On a real failure the harness prints `missing <needle>`, so the line
came out as:

  FAIL: empty parser output count: missing jq) returned 0 field(s)

which reads as if the script's own message had an unbalanced paren. It
does not; the needle was just sliced. Matching on
`policy parser (jq) returned N field(s)` makes the failure readable and
also pins that the refusal names the parser it used, which the narrower
needle did not.

make_jq() PATH-prefixes a fake jq, so `_PARSER_NAME` is deterministically
"jq" in both tests; the wider needle cannot flake on a host without jq.

Re-proved on this revision, because a disproof is about a revision and
not a file:

* suite exit 0, "PASS: probe member credentials guards"
* bash -n rc=0 on the test under /bin/bash 3.2.57 and bash 5.3.9
* dropping the empty-parse special case -> FAIL: empty parser output
  count: missing policy parser (jq) returned 0 field(s)
* arity threshold 5 -> 0 -> FAIL: short parser output status
* script restored byte-identical after each, green control after both
2026-09-12 11:50:11 +07:00
Dai Ha d8985719eb fleetd #525: restore the logger level, not just the appender, in SessionManagerTest
CI / contract (pull_request) Successful in 44s
CI / build (pull_request) Successful in 1m53s
The finally block in onTurnFailedIsLoggedAtWarnWithThePriorState (and four other
tests in this file) called setLevel(Level.WARN) on the shared SessionManager
logger (one test used MemberRegistry.class) but only detached the appender in
finally, never restoring the level. Since logback Logger instances are cached
per class and shared across the whole JVM, the pinned level leaked into every
test that ran after it.

Add CapturedLog, a small AutoCloseable that captures a logger's level and
appender together and restores both on close via try-with-resources, so this
shape cannot be half-fixed again. Convert all 7 addAppender/setLevel call
sites in this file to it, including the 2 sites #522 already fixed locally
(kept per-test pinning as belt-and-braces).

Add a proving test (sharedSessionManagerLoggerLevelIsRestoredAfterOnTurnFailedPinsWarn,
@Order(2), running right after the fixed test at @Order(1)) that fails before
this fix and passes after it. Verified with a mutation: dropping the level
restore in CapturedLog.close() turns the proving test red with
'expected: <DEBUG> but was: <WARN>'; restoring is byte-identical to the
pre-mutation file (sha256 matched) and the suite goes green again.

mvn -f fleetd/pom.xml clean install: Tests run: 1699, Failures: 0, Errors: 0,
Skipped: 0. BUILD SUCCESS.
2026-09-12 11:49:03 +07:00
Dai Ha 5b1e13ca3d fleetd #519 review fix: move the parser comments with the code they explain
CI / contract (pull_request) Successful in 50s
CI / build (pull_request) Successful in 2m2s
PR #523 extracted parse_policy_fields() but left about 25 lines of
explanatory comments at the old parse site in main(). That is the same
defect class as fleetd #500 itself — a stated fact that no longer
matches the code next to it — in the very file whose ticket history is
about it.

Three blocks moved, no code touched:

* "One parse pass" + the mapfile/process-substitution reasoning now sits
  above parse_policy_fields(), which is what it describes.
* The arity-check block now sits inside the function, directly above
  `if (( ${#_FIELDS[@]} < 5 ))`. At the old site it said "the slice just
  below this" and "every line below this expects", both pointing at a
  function call rather than the check. Reworded to name main() and its
  slice explicitly.
* The pipefail note said the parser failure was "handled below"; the
  handling is now above it, in the function.

The call site keeps a three-line pointer saying where the reasoning went.

Checked myself, on this revision:

* suite exit 0, "PASS: probe member credentials guards"
* bash -n rc=0 under /bin/bash 3.2.57 and bash 5.3.9
* two mutations killed, each proven applied two ways (mutant present AND
  original gone), restored byte-identical, green control after each:
  - dropping the empty-parse special case -> FAIL: empty parser output
    count
  - arity threshold 5 -> 0 -> FAIL: short parser output status
2026-09-12 11:47:16 +07:00
Dai Ha c89a375e5d fleetd #521: extract should_swap so the swap guard can't be silently disabled
CI / contract (pull_request) Successful in 1m29s
CI / build (pull_request) Successful in 1m29s
Mutating the swap step's guard (if [ "$DO_BUILD" = 1 ] -> if false) left the
whole test suite green: test_swap_ordered_after_wait_and_before_start only
checks source positions, which an in-place if-condition edit never moves.
Extracts the decision into should_swap(do_build), following the same shape as
#510's wait_for_daemon_exit and #517's drain_gate_refusal, with a direct test
for each value.
2026-09-12 11:42:46 +07:00
ltms 37dcefa834 Merge #522: log a positive drain-completion line when drainAll finishes (fleetd #512 part 1)
CI / contract (push) Successful in 55s
CI / build (push) Successful in 1m36s
Verified by the lead, not taken from the worker's report.

Build run by me, unpiped: exit 0, "Tests run: 1698, Failures: 0, Errors: 0, Skipped: 0",
BUILD SUCCESS. That is +2 on main's 1696, matching the two new tests.

Both merge constraints from the ticket hold, checked against the diff:
- The log.info is the last statement of drainAll's normal path and is NOT in a finally. The only
  finally in the file is at :372, unrelated. So a drain that dies still leaves no line, which is
  the absence signal part 2 will alert on.
- released and abandoned are incremented inside the drainSnapshot loop, not read from a collection
  at the end, so a partial report can never be served by this line.

Mutation check run by me: deleting the completion line makes both new tests fail by name —
drainAllLogsACompletionLineWithTheRealCountsOnACleanDrain and
drainAllLogsANonZeroAbandonedCountForASessionStillBusyAtTheDeadline, "no drain-complete INFO
logged, expected true but was false". Restored byte-identical
(d21ecd3adb3f66933e2f248a8ec81a2c324bfee323a3cc30da525842c33ea80a). So the tests really depend on
the line rather than passing for another reason.

Design accepted: drainSnapshot returns a private DrainTally record and release() returns the
removed session so abandoned can reuse the same BUSY-at-removal check logPreservedForShutdown
makes, instead of a second registry read. Both are private with one call site.

Two follow-ups filed rather than fixed here, see the ticket comments.
2026-09-12 06:34:51 +02:00
Dai Ha 6a7342b1f0 fleetd #518: make the FleetMcp caller-resolution wiring an explicit choice, and test it once for real
CI / contract (pull_request) Successful in 1m13s
CI / build (pull_request) Successful in 2m12s
FleetMcp's contextExtractor picked its principal-resolution path off `callers == null`, so
"authorization off" also silently swapped in a second, untested identity heuristic
(legacyPrincipal). Nothing drove that closure through a real MCP request, so the whole wiring
was an unexercised claim.

- callers (CallerResolver) is now required, never null.
- A new AuthorizationMode enum (ENFORCED/UNENFORCED) is a required constructor parameter with
  no default, replacing the null-means-legacy idiom for whether denyFor enforces at all.
- legacyPrincipal is deleted: there is exactly one resolution path now
  (callers.resolve(...)), so the mutation that swapped it for an unconditional legacy call no
  longer compiles ("cannot find symbol: method legacyPrincipal").
- FleetMcpContextExtractorTest boots the real transport on a real Jetty server and drives it
  with a real MCP client, proving fleet_whoami's resolved role comes from CallerResolver's
  token check.
- Adapted FleetMcpAuthzTest/FleetMcpHandoverTest call sites; theLegacyConstructorLeavesTheGateOpen
  keeps its meaning under the new AuthorizationMode.UNENFORCED value.
2026-09-12 11:34:12 +07:00
Dai Ha a5ad7c6561 fleetd #519: test policy probe guards
CI / contract (pull_request) Successful in 1m24s
CI / build (pull_request) Successful in 1m51s
2026-09-12 11:30:35 +07:00
ltms c71ac231e5 Merge #520: pin the drain-gate abort branch and jar_id's absent case (fleetd #517)
CI / contract (push) Successful in 1m28s
CI / build (push) Successful in 1m50s
Verified by the lead, not taken from the worker's report:

- 40 test functions defined, 40 invoked (my own greps), suite exit 0, zero real FAIL lines.
- Harness proof on my own invocation: re-applying the jar_id absent->present mutation gives
  exit 1 and "FAIL: jar_id with no arguments must report absent when $JAR does not exist".
  So a green run from this suite is readable.
- Script restored byte-identical after every mutation:
  2cb83dc380c7226191d657c40fccdfc856904e40b2d03d0851fb6e522cee2f41.
- drain_gate_refusal is pure and prints only; die() stays outside the command substitution, so
  the "die inside $( ) exits only the subshell" trap does not apply here.
- The --no-build + staged-present case reads "nothing changed" deliberately, and that is correct:
  --no-build stages nothing itself, and a leftover staged jar is wiped by rm -f "$JAR_STAGED"
  at :607, before the build at :610.

The worker's out-of-scope finding is real and I reproduced it: the swap guard at :723 can be set
to `if false` with the suite still green at exit 0. Filed separately.
2026-09-12 06:29:08 +02:00
Dai Ha 33720c42b3 fleetd #512 (part 1): log a positive completion line when drainAll finishes
CI / contract (pull_request) Successful in 59s
CI / build (pull_request) Successful in 1m38s
drainAll used to log nothing on a clean drain — both existing log calls
(drainSnapshot's per-session failure, drainAll's straggler-sweep warning)
sit on abnormal paths, so "drained fine" and "died on the first session"
looked identical: no log line either way.

Add one log.info at the end of drainAll: "drain complete: released=N
abandoned=M (still BUSY at the shutdown deadline)". It fires on the
normal path, including the all-zero case, and folds both drainSnapshot
passes (main snapshot + straggler sweep) into one line.

drainSnapshot now returns a private DrainTally(released, abandoned)
record instead of void, and the private release(paneId, cause) overload
now returns the removed MemberSession (previously void) so drainSnapshot
can read its state at the moment of removal — the same check
logPreservedForShutdown already makes. Both signature changes are
private with a single call site, so the blast radius stays small.
2026-09-12 11:29:01 +07:00
Dai Ha 3833d8e52b fleetd #517: pin the drain-gate abort branch and jar_id's absent case
CI / contract (pull_request) Successful in 1m15s
CI / build (pull_request) Successful in 1m31s
Two mutation-testing survivors in scripts/redeploy-fleetd.sh: a source-text
test pins what a message SAYS but never whether the branch that prints it is
REACHED.

- Extract the drain-gate abort decision into drain_gate_refusal(do_build,
  staged_path), a pure function the suite can call directly for all four
  build/staged combinations. The existing source-text grep test is kept
  alongside it (it catches a re-wording; the new tests catch a dead branch).
- Extend jar_id's test to cover the missing-file path (both the no-argument
  default and an explicit path), which the #511 test never exercised.
2026-09-12 11:23:37 +07:00
ltms b37def9238 Merge #516: the probe refuses with three distinct messages, each naming its own cause (fleetd #500)
CI / contract (push) Successful in 45s
CI / build (push) Successful in 1m46s
2026-09-12 06:09:50 +02:00
ltms 8f02576df6 Merge #515: pin the two-client completeness fold, and legacyPrincipal earns no authority (fleetd #509)
CI / contract (push) Successful in 48s
CI / build (push) Successful in 1m54s
2026-09-12 06:06:12 +02:00
ltms 525bc1c5f4 Merge #514: the drain-gate abort message names a recovery that works, and jar_id()'s default is pinned (fleetd #511)
CI / contract (push) Successful in 48s
CI / build (push) Successful in 1m41s
2026-09-12 06:00:23 +02:00
Dai Ha d59ece6dec fleetd #500: stop a wrong-interpreter or failed-parse reading a policy as empty
CI / contract (pull_request) Successful in 1m15s
CI / build (pull_request) Successful in 2m8s
probe-member-credentials.sh used mapfile < <(producer) to parse the fetched policy. That
hides a producer failure three ways: mapfile is bash 4+ and missing on macOS's /bin/bash
3.2, a process substitution's exit status is never propagated to mapfile, and the
downstream reads (":-" defaults and a slice) never fire set -u on a short or unset array.
All three converge on the same "0 known names" refusal, which blames the policy for a
failure that is actually the interpreter or the parser.

Three distinct guards, each closing one cause with its own message:
- a BASH_VERSINFO gate at the top refuses outright on bash < 4 (exit 3)
- the parser's output is captured via command substitution instead of mapfile < <(...),
  so a non-zero jq/python3 exit is caught at the call while the fact still exists (exit 4)
- an arity check before the field slice refuses a parse that exits 0 but returns fewer
  than 5 fields (exit 5)

The existing "0 known names" guard is now honest: by the time it fires, the three causes
above are already ruled out, so it really does mean the policy has 0 known names.
2026-09-12 10:58:16 +07:00
Dai Ha 32408d1e64 fleetd #509: pin the pane-scan completeness fold, and stop legacyPrincipal handing out primary
CI / contract (pull_request) Successful in 57s
CI / build (pull_request) Successful in 1m40s
Unit 1 — PaneLocator.terminalForPid's completeness fold across herdr
clients (PaneLocator.java:117) had no test that varied the number of
clients, so a mutation that keeps only the last client's Lookup.complete()
instead of ANDing every client's outcome survived: 14 of 15 existing tests
agree with the mutant on a single client. Added a two-client test where
the lead client errors on the pane that would have owned the pid (an
incomplete, negative scan) and the member client cleanly finds no panes
(a complete, negative scan) — the real fold ANDs these to false, a
last-wins fold reads it as true. Proved against MUTANTC
(complete = outcome.complete();): the new test fails with
"expected: <false> but was: <true>", the file was restored byte-identical
(sha256 unchanged), and the control run is green.

Unit 2 — FleetMcp.legacyPrincipal's else-branch returned Principal.primary
for ANY caller the connection did not resolve to a worker pane, with none
of CallerResolver.java:254's isLoopback/scanComplete guards. Measured that
no production caller passes null callers (Fleetd.java:696 always
constructs a real CallerResolver) but FleetMcpAuthzTest.mcp(false)
legitimately does, for its "legacy constructor leaves the gate open" test
— so the null-callers path is not dead code to delete (option a), it is a
documented legacy mode (option b). Changed the else-branch to
Principal.anonymous() and widened legacyPrincipal to package-private (like
denyFor) so a new test pins the behavior directly, since it only ever ran
inside a contextExtractor closure no existing test triggers.
2026-09-12 10:58:09 +07:00
Dai Ha 6e23bf8309 fleetd #511: fix wrong --no-build wording in drain-gate abort, pin jar_id() default
CI / contract (pull_request) Successful in 47s
CI / build (pull_request) Successful in 1m52s
The drain-gate abort message told the operator a rerun "with or without
--no-build" would finish the restart. That is wrong: by the time this
message can fire, stage_built_jar has already moved the jar off $JAR, so
--no-build hits require_no_build_jar's own refusal. Reworded to say the
rerun must NOT use --no-build, and why: the built jar is no longer at the
live path that --no-build requires.

Also added a test pinning jar_id()'s no-argument default (reports $JAR,
the live path) and its explicit-argument behavior (reports that path
instead), per fleetd #511 item 2. Not adding a test for the JAR_STAGED rm
-f at line 578 (fleetd #511 documents it as an equivalent mutant — mvn
clean install deletes target/ on the next line regardless).
2026-09-12 10:56:32 +07:00
ltms aa4c0b84c3 Merge #510: never build into the path a running daemon holds (fleetd #493)
CI / contract (push) Successful in 1m16s
CI / build (push) Successful in 2m4s
2026-09-12 05:44:16 +02:00
ltms 40c593cd09 Merge #508: a herdr error during the pane scan is refused, not promoted to primary (fleetd #505)
CI / contract (push) Successful in 1m18s
CI / build (push) Successful in 1m38s
Closes the worker->primary escalation that fleetd #317 left open on the other side of the same
ternary. #317 stopped an UNRESOLVABLE pid being promoted. This stops a RESOLVED pid whose pane scan
failed being promoted: the scan now reports completeness, and an incomplete scan resolves to
anonymous.

Shape chosen: PaneLocator.terminalForPid returns Lookup(terminal, complete); paneOwnsAnyOf returns a
private Ownership enum OWNS/DOES_NOT_OWN/UNKNOWN, so a HerdrException is UNKNOWN rather than a clean
negative. ConnectionIdentity.Caller gains scanComplete as a third field -- resolved() was NOT
widened, correctly: it is documented as testing the lsof sentinel and #505 is a different axis. A
definite match still short-circuits, so a genuinely vanished non-owning pane does not become a
refusal.

Verified by me, not taken from the report.

Trial-merged onto main (136312f) and built the merge:
  mvn clean install -> Tests run: 1694, Failures: 0, Errors: 0, Skipped: 0, BUILD SUCCESS

Harness proof, by re-running the worker's own mutation rather than trusting it (UNKNOWN ->
DOES_NOT_OWN): Tests run: 63, Failures: 3 -- exactly the three tests meant to pin it, matching their
report line for line:
  CallerResolverTest.aHerdrErrorOnTheOwningPaneDuringTheScanIsRefusedNotPromotedToPrimary
  ConnectionIdentityTest.scanIsIncompleteWhenHerdrErrorsOnThePaneThatOwnsThePid
  PaneLocatorTest.anErrorOnThePaneThatOwnsThePidMakesTheScanIncompleteNotAClearNegative
So the cells can go red, and their numbers are honest. Restored, shasum 81b797a6... byte-identical,
control 63/0.

MY OWN MUTATION, on the half they did not touch, FOUND A GAP. I mutated the two-client completeness
fold in terminalForPid -- `complete = complete && outcome.complete()` -> `complete =
outcome.complete()`, which forgets an earlier client's failure. Two greps proved it applied. Result:
Tests run: 63, Failures: 0. NOT pinned.

It is not an equivalent mutant: with CB-185's two herdr daemons, a first client that scans
incompletely followed by a second that scans cleanly gives `false` originally and `true` mutated, so
an incomplete scan would be reported complete and promoted. It is the #393 shape instead -- no test
assembles that combination, because the single-daemon tests collapse lead and member to one object.
That axis has cost us before: four routing defects passed the suite when one client served both roles
with memberHerdrSocket unset.

Merging anyway, because the production behaviour is correct and the ticket's defect is pinned three
ways -- what is missing is a test for a dimension the PR description claims in prose. Filed as a
follow-up with the numbers above.

Second item for the follow-up, latent not live: FleetMcp.legacyPrincipal still does
`c.terminal() != null ? worker : primary` and drops scanComplete, so it would re-open exactly this
hole. Measured unreachable in production -- Fleetd.java:696 always passes a real CallerResolver, and
the contextExtractor only calls legacyPrincipal when `callers == null`. A defect on paper is not a
reachable defect, so it is not a blocker; it is a trap for whoever next touches that constructor.

That second item is also the fleet01 lead's structural point, and it is the sharper framing: #415's
antidote (make the decision total) is ORTHOGONAL to this defect. Authz.permits is already a
default-less switch over Action and cannot help, because it says nothing about whether the caller was
resolved correctly, and Principal carries no record of how it was resolved. #415 guards a decision
nobody wrote; this was a decision written correctly and fed a bad input. The durable fix direction is
that an unresolved scan should be unrepresentable as a principal rather than checked for at each
consumer.
2026-09-12 05:35:41 +02:00
Dai Ha 979adf82eb fleetd #493: never build into the path a running daemon holds
CI / contract (pull_request) Successful in 1m30s
CI / build (pull_request) Successful in 1m37s
redeploy-fleetd.sh's build step wrote straight into fleetd/target/fleetd.jar
via `mvn clean install` while the OLD daemon was still running from that
exact path. A JVM loads classes lazily, so a class the daemon had not
touched yet could be read from a jar already replaced or removed -- the
failure landed on the shutdown drain (NoClassDefFoundError, exit 143,
looks clean).

Stage the freshly built jar at target/fleetd-new.jar (stage_built_jar),
confirm the OLD pid has actually exited (wait_for_daemon_exit, extracted
from the existing wait loop), and only then swap it into the live path
(swap_staged_jar) -- strictly after the wait, strictly before start. A
failed swap dies without starting. --no-build and --check keep their
existing, truthful behavior (require_no_build_jar; jar_id still reads
the live path by default). A leftover staged jar from an interrupted
run is wiped before the next build. The build still runs before
anything is stopped, so a failed build still never takes the fleet down.

Also fixed: the drain-gate abort message ("aborted -- nothing changed")
now names the staged jar when one exists, since staging already moves
the freshly built jar off the live path before that prompt runs.

Adds unit tests for stage_built_jar, swap_staged_jar, require_no_build_jar,
wait_for_daemon_exit, and a source-order test proving swap sits after the
wait and before start (sourcing stops before the main flow ever runs, so
the ordering itself can only be checked by reading the script's own call
sites).
2026-09-12 10:35:19 +07:00
Dai Ha 36870836aa fleetd #505: a herdr error during the pane scan must not read as a clean negative
CI / contract (pull_request) Successful in 1m26s
CI / build (pull_request) Successful in 1m29s
A transient herdr error on pane.process_info during PaneLocator's pid→pane scan used
to be swallowed into a plain "does not own it", so a real worker whose owning pane
errored mid-scan resolved with a null terminal but a resolved (real) pid — exactly
what CallerResolver's loopback-trust fallback reads as the primary. That is a
worker→primary privilege escalation through the door fleetd #317 did not close: #317
guards a failed lsof lookup (c.resolved()), not a failed herdr pane scan.

Fix: add a third state to the scan instead of widening Caller.resolved() (which stays
centralised next to the lsof sentinel it tests, per #505's explicit instruction not
to reopen that decision). PaneLocator.terminalForPid now returns a Lookup(terminal,
complete) record: a HerdrException on one pane marks that pane's ownership UNKNOWN,
not DOES_NOT_OWN, and the scan is complete only if every pane was either matched or
confirmed not to own the pid. A definite match found elsewhere in the same scan
still short-circuits as complete — a pane that genuinely vanished mid-scan without
being the caller's own does not turn into a refusal.

ConnectionIdentity.Caller carries the new scanComplete flag alongside the unchanged
resolved(). CallerResolver's loopback-trust fallback now requires both resolved()
and scanComplete() before promoting to Principal.primary(); an incomplete scan
resolves anonymous, which fails toward the recoverable error (a refused primary
retries loudly; a promoted worker would not).

Logs a warning naming the pane and which herdr client (of how many) failed, so the
incomplete-scan path is diagnosable rather than silent (fleetd #317's own lesson).
2026-09-12 10:27:42 +07:00
ltms 136312fb11 Merge #503: Injector's readiness-grace warn prints measured elapsed time, never arithmetic on constants (fleetd #501)
CI / build (push) Successful in 1m28s
CI / contract (push) Successful in 1m27s
Verified by the lead, not taken from the worker's report.

Trial-merged onto main (4f28da6) and built the merge, because a clean auto-merge is not a compiling
merge:
  mvn clean install -> Tests run: 1689, Failures: 0, Errors: 0, Skipped: 0, BUILD SUCCESS
  InjectorTest -> Tests run: 34, Failures: 0
  Gitea CI on ac351ee: success

The worker mutated the two log arguments. I mutated the half it did not: the STAMP SITE. Changing
the first-sample guard so the clock is re-stamped on every non-ready poll gave 1 failure,
InjectorTest.readinessGraceExpiryLogsTheMeasuredElapsedTimeNotArithmeticOnConstants, on a clock-read
counter assertion: "the readiness-not-ready branch should read the clock exactly twice — once to
stamp the first non-ready sample, once at grace expiry". Two greps proved the mutant applied;
restored, shasum byte-identical, control 34/0 green.

Behaviour preserved: the restructure from `else if (p != null && ++t.notReadySincePoll >= N)` to a
nested `else if (p != null) { ... }` keeps the short-circuit, so the counter still increments only
when p != null. All three reset sites now clear notReadySinceMillis alongside notReadySincePoll.

Two notes for the record, neither a blocker.

The poll-count half is an equivalent mutant and the worker said so instead of reporting a kill it
did not get. That is the right call and the code comment states the limit honestly.

The clock is injected through a package-private constructor overload while the public constructors
default it to System::currentTimeMillis. My brief prescribed that shape. With exactly one production
construction site (Fleetd.java:531) a required parameter — fleetd #415's antidote — would have been
just as cheap and would match LeadRollover and the new Fleetd.awaitHerdr. The silent-survivor risk
#415 names is not present here, because the field is final and both public constructors delegate, so
a future constructor cannot compile without supplying it. Recording the choice so the next person
does not read it as an oversight.
2026-09-12 05:07:35 +02:00
ltms 4f28da62a3 Merge #499: redeploy-fleetd.sh gains an "unclear" supervisor state that cannot reach kill, and the detail survives the subshell (fleetd #492, #497)
CI / contract (push) Successful in 1m12s
CI / build (push) Successful in 1m32s
Two commits, b17f37a and 599419f. Verified by the lead, not taken from the worker's report.

b17f37a adds the fifth answer "unclear" for installed-but-not-loaded and for a probe that could not
answer at all, and makes require_drivable_supervisor die on it.

599419f fixes a defect I found in b17f37a: SUPERVISOR_UNCLEAR_DETAIL was set inside
detect_supervisor, which is called as $(detect_supervisor). A subshell only returns stdout, so the
detail never reached the caller, and a ${VAR:-generic} fallback then printed a message naming no
supervisor and no reason. Measured before the fix: plain call -> detail length 142; via $( ) ->
length 0.

Lead verification of the fix, re-running my own original measurement across every branch:

  launchd installed, not loaded   kind=[unclear]   detail_len=142
  systemd installed, not active   kind=[unclear]   detail_len=146
  launchd loaded (normal)         kind=[launchd]   detail_len=0
  systemd loaded (normal)         kind=[systemd]   detail_len=0
  neither -> none                 kind=[none]      detail_len=0
  both loaded -> ambiguous        kind=[ambiguous] detail_len=0

The separator is always emitted, so the ${RAW#*SEP} unpack cannot silently fall back to the whole
string on the four branches that carry no detail.

Harness proof, run by me: making detect_supervisor print the kind alone (two greps proved the mutant
applied) turned the suite red, exit 1, "FAIL: detail does not name the systemd unit it found
installed-but-not-loaded". Restored, shasum byte-identical, control exit 0, bash -n clean on both
files. 23 test functions defined, 23 invoked.

The three case "$SUPERVISOR_KIND" switches now have final arms, chosen by what each caller does with
the value: the reporting switch warns and continues (a diagnostic that aborts goes silent exactly
when the state is novel), while stop and start die.

Not fixed here, filed separately: the "was already not running" branches at :603 and :610 swallow a
failed launchctl unload / systemctl stop with || true and then print "ok" unconditionally.
2026-09-12 05:04:04 +02:00
ltms 708f1795ad Merge #502: awaitHerdr reports three outcomes with measured elapsed time, not one boolean (fleetd #498)
CI / contract (push) Successful in 58s
CI / build (push) Successful in 1m49s
Verified by the lead, not taken from the worker's report.

Trial-merged onto main (8f59019) locally and built the merge, because a clean auto-merge is not a
compiling merge:
  mvn clean install -> Tests run: 1687, Failures: 0, Errors: 0, Skipped: 0, BUILD SUCCESS
  FleetdAwaitHerdrTest -> Tests run: 6, Failures: 0
  Gitea CI on 274afaf: success

The worker mutated the three return values. I mutated the half it did not: the reap DECISION at the
call site. Changing logHerdrWaitOutcomeAndShouldReap to return true for INTERRUPTED gave 1 failure,
FleetdAwaitHerdrTest.interruptedLogsItsOwnMessageAndNeverClaimsTheBudgetElapsed:135 ("an interrupted
wait must not tell main to reap"). Two greps proved the mutant applied; restored, shasum
byte-identical, control 6/0 green.

Behaviour preserved: herdrUp is still true only for ANSWERED, and both its readers (the orphan reap
at :276 and the lead auto-launch gate at :375) see exactly what they saw before.

One correction to the worker's report, which changes nothing in the code. It wrote that a real
interrupt-detection regression "would also hang the daemon's startup thread forever". It would not:
in production `nanos` is System::nanoTime, so the deadline check still fires. The hang it hit was a
test-fixture property — a frozen injected clock with no iteration bound. That is fleetd #486's
shape, now seen in a second class, and it is recorded there.
2026-09-12 05:01:02 +02:00
Dai Ha 599419f9e6 fleetd #492 follow-up: carry the unclear detail across detect_supervisor's subshell boundary
CI / contract (pull_request) Successful in 47s
CI / build (pull_request) Successful in 1m53s
b17f37a set SUPERVISOR_UNCLEAR_DETAIL as a global inside detect_supervisor, but the real call
site invokes it as $(detect_supervisor) — a subshell — so that global died with the subshell and
the die() message's ${VAR:-fallback} silently masked the loss with generic text.

- detect_supervisor now packs kind and detail onto its one stdout line (joined by the ASCII unit
  separator byte, $SUPERVISOR_DETAIL_SEP), the only channel that survives $( ). The real call site
  unpacks both with in-shell parameter expansion — no extra subshell.
- Dropped the ${SUPERVISOR_UNCLEAR_DETAIL:-...} fallback at the die() message: under set -u, a
  missing detail now fails loudly instead of silently defaulting (same defect class as #497).
- Added a constraints comment block above detect_supervisor for future callers: stdout-only,
  no ${VAR:-default} papering over a lost value, and every case on the return value needs an
  explicit *) arm.
- Added *) arms to the three `case "$SUPERVISOR_KIND"` switches (report/stop/start): report warns
  and continues (display-only), stop/start die naming the value (they act on it).
- Rewrote the "unclear" test to go through the real call-site shape ($(detect_supervisor) then
  the same split), not a hand-constructed value, and tightened its final assertion to check for
  the actual detail text rather than $SYSTEMD_UNIT alone (the die() boilerplate names the unit
  either way, so that check could pass on a lost value).
2026-09-12 09:58:45 +07:00
Dai Ha ac351ee1de fleetd #501: readiness-grace expiry logs measured elapsed time and the loop's own poll counter, never the configured budget
CI / contract (pull_request) Successful in 49s
CI / build (pull_request) Successful in 2m2s
The line at Injector.java:383-387 printed two numbers that read as measurements
but were both compile-time constants: READINESS_GRACE_POLLS for the poll count
(the loop's own Target.notReadySincePoll counter was in scope at the same call
site), and READINESS_GRACE_POLLS * POLL_INTERVAL_MILLIS / 1000 for the elapsed
time — arithmetic on two constants, never a measurement, and wrong in the
direction that says everything ran on schedule.

Fix, copying the LongSupplier-clock shape LeadRollover already uses:
- print t.notReadySincePoll instead of the constant for the poll count (the two
  agree by construction on this branch, so no test can tell them apart — the
  comment says so honestly).
- add Target.notReadySinceMillis, stamped at the first non-ready sample and
  reset at all three sites notReadySincePoll already resets (:304, :355, :389
  pre-fix line numbers), to compute a real elapsed time at expiry.
- inject a LongSupplier nowMillis (defaulting to System::currentTimeMillis)
  through new package-private constructor overloads so a test can supply a
  clock whose advance does not track POLL_INTERVAL_MILLIS.

Tests use a ListAppender to assert on the log message contents, per
LeadRolloverTest's pattern. The elapsed-time test drives the loop with a stub
clock returning two literal, non-derived values so it can fail if the fix
regresses to the constant-arithmetic line — proved by mutation: reverting the
elapsed calculation to READINESS_GRACE_POLLS * POLL_INTERVAL_MILLIS turns that
one test red (1 failure); reverting the poll-count print to the constant is an
equivalent mutant (0 failures), because the counter and the constant are
identical at that exact call site by construction.

fleetd clean install: Tests run: 1683, Failures: 0, Errors: 0, Skipped: 0.
2026-09-12 09:57:57 +07:00
Dai Ha 274afafde6 fleetd #498: awaitHerdr distinguishes deadline-passed from interrupted, with measured elapsed time
CI / contract (pull_request) Successful in 51s
CI / build (pull_request) Successful in 1m51s
- awaitHerdr now returns a HerdrAwaitOutcome(HerdrWaitResult, elapsedNanos) instead of a bare
  boolean, so 'the wait budget genuinely ran out' and 'the waiting thread was interrupted' are
  two distinct, named states instead of the same false (fleetd #497's shape).
- awaitHerdr takes the clock (LongSupplier) and the per-poll sleep (Runnable) as required
  parameters, with no defaulted overload (fleetd #415), so a test can drive it.
- The startup call site is extracted into logHerdrWaitOutcomeAndShouldReap, since main() itself
  cannot be driven from a unit test; it logs a distinct message per outcome, always printing the
  measured elapsed time next to the configured budget, never the budget alone.
- Adds FleetdAwaitHerdrTest covering the seam (all three outcomes, plus the preserved interrupt
  flag) and the call site (the three distinct log messages), using ListAppender.
2026-09-12 09:54:59 +07:00
ltms 8f59019305 Merge #496: lead-rollover logs measured elapsed time and measured counts, never the configured budget (fleetd #494)
CI / contract (push) Successful in 1m14s
CI / build (push) Successful in 2m1s
Three commits. All verified by the lead in a detached worktree, not taken from the worker's report.

3fb3311 — the /clear-settle and success paths return measured elapsed time and a real nudge count
c87cc25 — print the measured nudge count; fix the same defect on the sibling turn-settle timeout line
e966cba — the poll count on the same line was still a constant; the test could not tell the difference

Lead verification at e966cba:
  mvn clean install -> Tests run: 1681, Failures: 0, Errors: 0, Skipped: 0, BUILD SUCCESS
  LeadRolloverTest  -> Tests run: 33, Failures: 0, Errors: 0, Skipped: 0
  Gitea CI on e966cba: success

Mutation M2 (PICKUP_GRACE_POLLS 8 -> 5), run by the lead: 2 failures,
clearGraceReleaseLogIsWarnWithMeasuredElapsed and successLogPrintsMeasuredElapsedForTheWholeRoll.
The emitted line under the mutant read "after 5 consecutive IDLE/DONE polls (4 of those were
nudged)", so both numbers now follow the loop. Restored, shasum byte-identical, control 33/0 green.

Known and deliberate limit, stated in the code comment: on the grace-release branch both counters
equal the constants by construction (one write site each), so no test can prove the difference on
that line. The change buys one source of truth, not provable coverage. The place nudges genuinely
varies with the run is the /clear-timeout warn, which is covered by a discriminating test.
2026-09-12 04:44:10 +02:00
Dai Ha e966cbadf9 fleetd #494 follow-up (2nd pass): the grace-release line still had one constant, and its test could not tell the difference
CI / contract (pull_request) Successful in 1m26s
CI / build (pull_request) Successful in 1m55s
Two more fixes on the same line, LeadRollover.java:600-624:

1. The poll-count argument (second, was PICKUP_GRACE_POLLS) now prints the
   loop's own idlePollsAwaitingPickup counter instead of the constant. Same
   defect shape as the nudges fix from c87cc25, one argument over.

2. LeadRolloverTest's clearGraceReleaseLogIsWarnWithMeasuredElapsed asserted
   its nudge-count expectation as `LeadRollover.PICKUP_GRACE_POLLS - 1` — the
   same expression the production code used to build the log line from, so it
   could not discriminate a reverted fix. Rewritten to plain literals
   ("after 8 consecutive IDLE/DONE polls (7 of those were nudged)"), proven to
   trip when PICKUP_GRACE_POLLS's value changes (2 failures under a
   PICKUP_GRACE_POLLS=5 mutation: this test and successLogPrintsMeasured...).

Also corrected the comment above the log line: idlePollsAwaitingPickup and
nudges each have exactly one write site on this loop's release branch, so they
cannot differ from PICKUP_GRACE_POLLS / PICKUP_GRACE_POLLS - 1 at this call
site — confirmed by re-deriving the loop's control flow, and by reverting just
the nudges argument (M1) and observing 33/0 stayed green even after the test
rewrite. That is an equivalent mutant on this line, not a gap the test rewrite
could close; the honest value of printing the counters is one source of truth
for the loop, not a provable-by-test difference here. The place nudges truly
varies with the run — and is covered by a test that can tell it apart from a
constant — is the /clear-timeout warn's clearResult.nudges() in runRollover.

mvn -Dtest=LeadRolloverTest test: 33/0. mvn clean install: 1681/0, BUILD SUCCESS.

Same-shape sweep (found, not fixed, per instructions):
Injector.java:383-387 — the readiness-grace warn prints the constant
READINESS_GRACE_POLLS where the measured per-target counter
t.notReadySincePoll is in scope (single increment site, just reached the
threshold at this call site — same equivalent-mutant situation as this fix).
2026-09-12 09:39:39 +07:00
Dai Ha b17f37a683 fleetd #492 follow-up: detect_supervisor must never read "could not tell" as "none"
CI / contract (pull_request) Successful in 1m11s
CI / build (pull_request) Successful in 1m32s
Two situations were silently landing in the "none" answer, which
require_drivable_supervisor accepts and the script then falls back to a raw
kill + nohup — exactly the wrong move when a supervisor actually IS present:

- installed-but-not-loaded, on either supervisor. `systemctl --user is-active`
  answers "no" for activating/deactivating/failed and while an auto-restart is
  pending too, and every one of those is a host that IS under systemd (or
  launchd) and about to act again. `*_installed` already knew this; it was
  only ever consulted for a warning line, never by the decision itself.
- a systemd probe that could not answer at all (e.g. systemctl cannot reach
  the user bus over a non-lingering ssh session) looked identical to a clean
  negative, because both probes redirected stderr straight to /dev/null.

detect_supervisor now returns a fifth answer, "unclear", for both cases.
systemd_loaded/systemd_installed capture systemctl's exit status and stderr
separately and set their own *_ERRORED flag only on a real tool failure
(non-zero exit WITH stderr), never on a clean negative. "none" now means only:
neither supervisor installed, neither loaded, neither probe errored.
require_drivable_supervisor die()s on "unclear" exactly like it already does
on "ambiguous", naming the specific supervisor and reason via the new
SUPERVISOR_UNCLEAR_DETAIL global.

Tests: 4 new cases (systemd/launchd installed-but-not-loaded, a real
systemd_loaded run through a systemctl stub that errors on stderr, and the
die() refusal for "unclear" naming the unit). All 3 new guards were verified
by mutation: each was removed from the real script, the suite caught it (a
new FAIL line naming the exact broken assertion), then the file was restored
byte-identically and the suite went green again.
2026-09-12 09:23:21 +07:00
Dai Ha c87cc25aa6 fleetd #494 follow-up: print the measured nudge count, and fix the sibling turn-settle timeout line
CI / contract (pull_request) Successful in 1m13s
CI / build (pull_request) Successful in 1m30s
- waitForClearPickupAndSettle's grace-release warn now prints the measured 'nudges' counter
  instead of the constant PICKUP_GRACE_POLLS - 1. The two happen to agree today, but the
  constant expression was wrong once before (printed PICKUP_GRACE_POLLS itself, claiming 8
  nudges where 7 went out) and a code read did not catch it — only a mutation test did.
  Printing the counter cannot drift from the loop's real behaviour.
- waitUntilAtTurnBoundary (the FIRST wait, ~line 386-391) had the identical 'configured value
  printed as if measured' defect as the three lines fixed in the original #494 commit, but was
  out of scope because the brief named specific lines instead of the shape. Fixed the same way:
  it now returns a TurnSettleResult(settled, elapsedMillis) instead of a bare boolean, and the
  timeout warn prints 'configured={}s elapsed={}ms' instead of presenting cfg.turnSettleSeconds()
  as the measured wait.
- No behaviour change: same sends, same order, same release/refuse decisions.
- Test additions: clearGraceReleaseLogIsWarnWithMeasuredElapsed now also asserts the measured
  nudge count; successLogPrintsMeasuredElapsedForTheWholeRoll's expected elapsed value is
  updated (7000ms, not 6500ms) to account for waitUntilAtTurnBoundary's own new clock read; a
  new turnTimeoutLogPrintsMeasuredElapsedNotJustConfigured test pins the sibling line.
- Proved the nudges fix with a temporary mutation: set PICKUP_GRACE_POLLS to 5, confirmed via
  two greps that the mutant applied and the original constant was gone, ran the grace-release
  test and read the actual log line — nudge count followed to 4 (= 5 - 1), then restored to 8
  and reran the full LeadRolloverTest suite as a control (33/33 green).
2026-09-12 09:22:19 +07:00
Dai Ha 3fb331145a fleetd #494: log measured elapsed time, never the configured budget, on a lead-rollover failure/success
CI / contract (pull_request) Successful in 52s
CI / build (pull_request) Successful in 1m35s
- runRollover's /clear-timeout warn now prints configured/elapsed/nudges, each labelled,
  instead of presenting cfg.clearSettleSeconds() as if it were the measured wait.
- waitForClearPickupAndSettle's pickup-grace release is now log.warn (was log.info) and
  prints the measured elapsed time next to the target pane — this is the exact path that
  reported a false-success roll in the real incident (438ms of a 20s budget).
- The success line ('lead-rollover: rolled') now prints the measured elapsed time for the
  whole roll.
- waitForClearPickupAndSettle now returns a ClearSettleResult(settled, elapsedMillis, nudges)
  instead of a bare boolean, so callers can log the measured values instead of the config.
- No behaviour change: same sends, same order, same release/refuse decisions.
- Adds 3 tests to LeadRolloverTest pinning the content of each changed log line, using a
  self-advancing fake clock so the measured elapsed/nudge values are deterministic and
  provably distinct from the configured budget.
2026-09-12 09:08:04 +07:00
Dai Ha dcd505286f fleetd #492: teach redeploy-fleetd.sh systemd --user as a third supervisor
CI / contract (pull_request) Successful in 56s
CI / build (pull_request) Successful in 1m40s
launchd, systemd --user, and unsupervised are three different answers, not two.
Refuse (die) rather than fall through to kill+nohup when a supervisor is
detected that this script cannot drive (e.g. both signals fire at once), and
add a post-restart check that fails the run if more than one fleetd process
is alive. detect_supervisor()/require_drivable_supervisor()/
count_daemon_pids()/assert_single_daemon() are pure, overridable functions so
scripts/test-redeploy-fleetd.sh can exercise them without a real launchd or
systemd.
2026-09-12 09:00:59 +07:00
Dai Ha 60fa958da2 docs: a relative handoverPath lands in the LEAD's repo, not fleetd's (#491)
CI / contract (push) Successful in 1m29s
CI / build (push) Successful in 1m34s
On this host the lead's cwd IS the fleetd checkout, so fleetd/.gitignore
protects the handover file and the distinction is invisible. On fleet01 the
lead works in /home/ltms/LTMS/kb while fleetd sits in a different directory,
and that repo has no .handover rule — measured 2026-09-12.

Tell the lead to check its own workspace's .gitignore before writing, and to
report it rather than committing the file or editing someone's .gitignore.

Refs fleetd #491, #487, #480.
2026-09-12 08:46:17 +07:00
Dai Ha 9950361bc9 docs: #489 is fixed and deployed, but criterion 6 is still unmet
CI / contract (push) Successful in 1m12s
CI / build (push) Successful in 1m33s
The skill said the roll 'does nothing until #489 is merged and redeployed'.
Both have happened, so that sentence now reads as a green light. It is not
one: no roll has bootstrapped a fresh session end to end yet. Say that
plainly, and tell the lead to warn the operator before confirming.

Refs fleetd #489, #480.
2026-09-12 07:55:41 +07:00
ltms 9f74b6619a Merge #490: nudge the /clear submit keystroke before bootstrapText (fleetd #489)
CI / contract (push) Successful in 1m14s
CI / build (push) Successful in 1m30s
Fixes the live defect measured on 2026-09-12: the roll joined /clear and
bootstrapText into one line and Claude Code refused it as
"Unknown command: /clearFresh".

Verified by the lead before merge:
- mvn clean install: Tests run: 1677, Failures: 0, BUILD SUCCESS
- LeadRolloverTest baseline: 29 tests, 0 failures
- mutation PICKUP_GRACE_POLLS 8 -> 1: 1 failure (the regression test)
- mutation deleting the !clearSettled guard: 3 failures, 2 pre-existing tests
- mutation replacing waitForClearPickupAndSettle with "return true": 5 failures,
  including the strengthened pickup test
- positive control after each restore: 29 tests, 0 failures
- CI run 1738 on f687046: success

A separate mutation of the FIRST gate hung the suite instead of failing it.
That is fleetd #486, not a regression here; the reproduction is recorded there.
2026-09-12 02:52:52 +02:00
Dai Ha f687046450 fleetd #489 follow-up: fix stale class javadoc, off-by-one nudge count, weak test
CI / contract (pull_request) Successful in 1m22s
CI / build (pull_request) Successful in 1m33s
Three review corrections on top of the previous commit:

1. The class javadoc's four-step continuation list (lines 47-58) was stale.
   Step 1 said "report an injectable state", but waitUntilAtTurnBoundary's
   own javadoc excludes BLOCKED - fixed to say IDLE or DONE. Step 3 still
   described the old plain re-check ("the original, pre-correction wait...
   still here") - fixed to describe what waitForClearPickupAndSettle
   actually does: nudge while unpicked-up, then wait for a real WORKING ->
   IDLE/DONE boundary, releasing rather than wedging if WORKING never shows.

2. PICKUP_GRACE_POLLS=8 bounds the number of consecutive not-yet-picked-up
   polls, not the number of nudges - the 8th poll releases instead of
   nudging again, so 8 polls produce 7 nudges. The log.info in the release
   branch and two javadoc spots said "8 nudges"; fixed all three to state
   the poll count and the nudge count separately and correctly. Behavior
   and the constant are unchanged.

3. pickupSeenStopsNudgingAndBootstrapTextIsSent asserted only promptCallCount
   and sendKeysCallCount, both of which a return-true stub also satisfies.
   Added an assertion on the already-tracked postClearGetCalls counter
   (>= 2), which only a real post-/clear poll loop can produce - this is
   what makes the test fail against a return-true mutant.
2026-09-12 07:45:36 +07:00
Dai Ha c6058652be fleetd #489: nudge the /clear submit keystroke before bootstrapText
CI / contract (pull_request) Successful in 56s
CI / build (pull_request) Successful in 1m36s
LeadRollover.runRollover's second wait (after /clear) was a no-op: it polled
for IDLE/DONE, which /clear itself never leaves since it starts no real turn,
so it always returned true on the first poll. Combined with a direct
agents.send bypassing Injector (deliberate, to avoid wedging the pane), the
submit Enter that accompanies /clear could race the paste and leave it
unsubmitted — bootstrapText then landed concatenated onto the same input
line, exactly as measured live on 2026-09-12.

Replace that second wait with waitForClearPickupAndSettle, which copies the
pickup-nudge pattern Injector already ships for its own post-turn /clear
housekeeping (fleetd #306): nudge agents.submit while the pane hasn't
reported WORKING yet, release after PICKUP_GRACE_POLLS=8 nudges rather than
wedge, and require a real WORKING -> IDLE/DONE boundary once a pickup is
observed. BLOCKED stays excluded from both the nudge and the boundary check,
same as the (unchanged) first wait — a paused live turn is not settled, and
nudging Enter into an open prompt could wrongly answer it.

Adds four tests to LeadRolloverTest covering the paste-race regression
(nudge ordered between /clear and bootstrapText), a confirmed pickup, a
deadline expiry with no boundary ever reached, and a throwing submit().
2026-09-12 07:29:26 +07:00
Dai Ha 2a95da2ff4 docs: the handover skill must warn that the automatic roll is broken (#489)
CI / contract (push) Successful in 1m23s
CI / build (push) Successful in 1m53s
Section 11 said the bootstrap prompt landing was 'not yet proven
end-to-end'. It is now measured failing: the first real roll joined
/clear and bootstrapText into one line. Nothing was cleared, so the
failure is safe, but a lead that reads the old wording would reach for
fleet_handover expecting it to work.

Refs fleetd #489, #480.
2026-09-12 07:21:38 +07:00
Dai Ha 7f9137fcb6 fleetd #480: ignore the handover file, and note the absolute-path guarantee in the handover skill
CI / contract (push) Successful in 47s
CI / build (push) Successful in 1m52s
The lead rollover handover file now lives inside the workspace, at the relative
path fleetd.yaml's leadRollover.handoverPath names. It is a snapshot of one
moment's live state, so it must never enter git history.

The handover skill also now says the handoverPath fleetd hands back is always
absolute, even when the configured value is relative — a lead that resolves it
itself can pick a different file from the one the daemon checks.
2026-09-12 05:46:34 +07:00
Dai Ha 3bf3968bc7 Merge #487: resolve a relative leadRollover.handoverPath against the calling lead's workspace (fleetd #480 follow-up) 2026-09-12 05:43:12 +07:00
Dai Ha 261aa056f9 fleetd #480 follow-up correction 2: guarantee resolveHandoverPath is always absolute
CI / contract (pull_request) Successful in 1m26s
CI / build (pull_request) Successful in 1m32s
LeadRollover.resolveHandoverPath's relative branch resolved the configured
handoverPath against leadWorkspace.apply(...) (fleet.leaders.<name>.cwd) but
never forced the result absolute. If an operator writes a RELATIVE cwd, the
returned path stays relative, silently breaking the "always absolute"
contract documented on PendingRollover.

Fix: call toAbsolutePath() unconditionally on both branches (the
already-absolute input branch, where it is a no-op, and the relative
branch), so neither branch trusts isAbsolute() alone to already imply what
toAbsolutePath() enforces. Method javadoc now states the absolute result is
guaranteed, not merely usual.

Added a test: a lead with a RELATIVE cwd and a relative handoverPath still
yields an absolute PendingRollover.handoverPath. Asserts both isAbsolute()
and the exact resolved value, since isAbsolute() alone would also pass for a
path resolved against the wrong base.

Proved the test discriminates: reverting the toAbsolutePath() calls (keeping
the test) made it fail with an AssertionFailedError ("expected: <true> but
was: <false>"); restoring the fix made it pass again.

Note: FleetConfig has no validation on fleet.leaders.<name>.cwd at config
load (grep across every validate* method: 0 matches for .cwd()) — a relative
cwd is silently accepted. Not adding validation here per instruction; that
is a separate ticket.
2026-09-12 05:41:40 +07:00
Dai Ha 042b8c99dd fleetd #480 follow-up correction: cover Fleetd.leadRollover(...)'s own wiring behaviourally
CI / contract (pull_request) Successful in 48s
CI / build (pull_request) Successful in 1m51s
Add FleetdLeadRolloverWorkspaceLookupTest, calling the package-private Fleetd.leadRollover(...)
factory directly (with a real ConfigRef built from a temp fleetd.yaml, never the gitignored live
one) to prove the terminal -> lead-name -> Leader.cwd() lookup it builds actually works: a relative
handoverPath resolves against the calling lead's configured cwd; a terminal absent from the live
lead-terminal map falls back to user.dir; and the lookup is read live, not snapshotted at
construction time (a lead discovered by the tab scan after leadRollover(...) was built still
resolves correctly).

Proved this closes the gap: mutating the factory's lambda body (String leadName = null;, always
"no lead found", which forces the daemon-cwd fallback this ticket exists to fix) left the full
1669-test suite green before this commit. With the new test added, the same one-line mutation now
fails 2 of its 3 cases; reverting it goes green again (3/3). Mutation applied/reverted only during
verification and is not part of this commit (git diff on Fleetd.java is empty).

FleetdLeadRolloverWiringTest's class javadoc corrected: it previously claimed no behavioural test
could catch this wiring dropping out, which was true only before this commit and only covered the
factory's own body, not its call site. Restated what each test class actually covers: the source-
text pin covers the call site's argument list; the new behavioural test covers the lambda's body.
2026-09-12 05:32:25 +07:00
Dai Ha 4bfab6b718 fleetd #480 follow-up: resolve a relative leadRollover.handoverPath against the calling lead's workspace
CI / contract (pull_request) Successful in 49s
CI / build (pull_request) Successful in 2m4s
LeadRollover.open() now resolves handoverPath to an absolute path exactly once, against the
calling lead's fleet.leaders.<name>.cwd (falling back to the daemon's own user.dir when that
lead has none configured), matching the LeadLauncher#launch precedent. PendingRollover stores
only the resolved absolute path, so checkHandover's exists/empty/fresh checks, the path handed
back to the lead in the fleet_handover open response, and the default bootstrapText sentence
all see the same absolute location instead of a value resolved against whatever directory the
daemon process happened to start in.

FleetConfig.LeadRollover.bootstrapText is no longer defaulted in the compact constructor (it
would otherwise still bake in the raw, possibly-relative handoverPath); a new bootstrapTextFor
(resolvedHandoverPath) method builds the default sentence from the resolved path instead.

Fleetd.leadRollover(...) gains a required liveLeadTerminals parameter to build the terminal to
lead-name to Leader.cwd lookup, read live through the existing `leads` supplier and ConfigRef on
every call, never off a startup snapshot.
2026-09-12 05:19:49 +07:00
Dai Ha 008a457557 docs: fleet_handover is shipped — update the charter table and the handover skill
CI / contract (push) Successful in 54s
CI / build (push) Successful in 1m38s
fleetd #480 merged in #483, #484 and #485, so the instruction surface has to
catch up. Per CLAUDE.md's own rule, a code change that silently invalidates the
canonical block is an incomplete change.

CLAUDE.md + wiki/7-Use-Cases.md: one new intent-to-tool row for fleet_handover.
Both copies edited identically; the byte-identical sync check prints True.
The row carries the trap rather than just the call: open FIRST, then write the
file, then confirm — because confirm refuses the file as stale unless its
modified time is later than the open request, so the obvious order fails.

.claude/skills/handover/SKILL.md: the skill said "do not call a fleet_handover
tool: it does not exist." That was true when it was written this morning and is
now false, which is the worst state for an instruction file to be in. Replaced
with the two real paths (by hand, or with the tool when leadRollover: is
configured), plus a new section 11 giving the three-step order and the five
things that surprise a caller — chiefly that `accepted` does not mean the pane
has been cleared, and that operatorConfirmed is a report of what a human said,
not a confidence level.

Section 11 also states plainly that the bootstrap prompt landing in a freshly
cleared pane is not yet proven end-to-end, and that the recovery is the manual
path — which is why the file is written before confirm, never after.
2026-09-11 07:29:59 +07:00