Commit Graph

212 Commits

Author SHA1 Message Date
Dai Ha 1dbe3a03fc Merge lead-comms-wiring: wire LeadMailbox into the daemon (fleet_send{coordId} + receive loop) 2026-08-24 17:42:11 +02:00
Dai Ha 6058b8472b lead comms: wire LeadMailbox into the daemon so lead-to-lead messages flow
CI / contract (pull_request) Successful in 41s
CI / build (pull_request) Failing after 1m22s
The sibling ticket landed the mechanism (LeadMailbox, LeadMessage, the
coordinator: config block) but nothing opened it, nothing sent through it, and
nothing read it. This is the wiring.

- LeadChannel: a small interface LeadMailbox now implements (publish/peek/ack
  plus a selfCoordId() accessor). It exists so FleetMcp and the receive loop can
  be tested with a fake instead of a live broker. LeadMailbox's AMQP logic is
  untouched — the diff is the implements clause, four @Override marks and the
  accessor.

- Fleetd.openLeadMailbox: opens this daemon's mailbox after the reply inbox is
  selected, with the same env-injected seam selectReplyInbox uses. Every "off"
  path returns null and the daemon still starts: no coordinator block (silent),
  a uriEnv that does not resolve (INFO), a configured broker with no selfId
  (WARN — a mailbox is named after the coord-id that owns it), or a broker that
  refuses at boot (WARN, credentials stripped). Closed in the ordered shutdown
  hook, after the loop that reads it has stopped.

- fleet_send{coordId}: publishes a LeadMessage(from=selfCoordId, to=coordId) to
  the peer's mailbox and returns the broker-confirmed receipt. coordId is
  mutually exclusive with sessionId/turnId and is rejected by name rather than
  resolved by precedence. An unroutable/nacked/timed-out publish comes back as a
  tool error naming the coordId, never a crash. The worker send/reply path is
  not touched.

- LeadCoordLoop: the receive half. Each tick peeks the mailbox, resolves the
  local lead pane, and — only at a turn boundary — injects "[lead <from>] <text>"
  and acks. Anything not delivered stays unacked and is retried, so a message is
  never dropped; one message per tick, so every delivery is gated on a status
  read that already saw the previous one.

- fleet_list reports {selfId, configured} when coordination is on, so an
  operator can find the coord-id a peer must use to reach them. Omitted
  entirely when it is off.

Tests: 20 new hermetic tests (no broker) across routing, delivery and startup
selection. mvn clean install: Tests run: 944, Failures: 0, Errors: 0, Skipped: 0
— BUILD SUCCESS.
2026-08-24 17:39:29 +02:00
Dai Ha f29968c334 Merge autocompact-window: per-profile autoCompactWindow for claude-code + opencode members 2026-08-24 17:32:53 +02:00
Dai Ha 7b98cca967 LeadMailbox: durable leader-to-leader mailbox over a shared coordination vhost
CI / contract (pull_request) Successful in 44s
CI / build (pull_request) Failing after 1m35s
Adds the broker-side mechanism for lead-to-lead messages across daemons/hosts
(unit 1 of 2): a LeadMessage envelope carrying from/to coord-ids, an
AMQP-backed LeadMailbox modeled closely on AmqpReplyInbox (consume-and-hold,
deferred manual ack, confirm-mode publish, recovery handling), and a new
optional coordinator: config block (separate vhost from broker:, leader
traffic only). Config parsing + accessors only — FleetMcp/Fleetd/Injector/
MessageService and the send path are untouched; wiring is a separate ticket.
2026-08-24 17:23:42 +02:00
Dai Ha 2757bc7185 autoCompactWindow: per-profile bounded auto-compaction, both backends
CI / contract (pull_request) Successful in 41s
CI / build (pull_request) Failing after 1m37s
Add opt-in Integer autoCompactWindow to FleetConfig.Profile (last field,
null/unset = today's behaviour). Validated at config load to [100000,
1000000] — the band Claude Code's own --autocompact flag accepts.

Claude Code: appends --autocompact <window> to argv (mirrors --model),
so it survives the ccs <profile> wrapper.

opencode: has no absolute compact-at-N knob (only compaction.auto/prune/
reserved/tail_turns/preserve_recent_tokens), so the window is applied as
the resolved model's own limit.context (+ a required limit.output:16384
default) in the generated opencode.json, merged via get-or-create nodes
so it does not clobber a custom-provider block. Only applies when model:
resolves to "provider/model"; otherwise logs a WARN naming the profile
rather than silently doing nothing.

Docs added to fleetd.example.yaml explaining the cross-backend semantics
difference (compacts AT the window vs. WITHIN it).
2026-08-24 16:54:40 +02:00
Dai Ha 7655f1b51a CB-634: pin the IDE overlay to the module dir + best-effort auto-open
The overlay pinned project_path to the worktree root. For a repo whose Maven
module is a subdir (this repo's pom is in `bridged/`, not at the root), opening
the root imports no module and every ide_* call resolves nothing. Pin and open
the module dir instead.

Two new opt-in per-Profile keys, both read only when ideMcpUrl is set:
- ideProjectDir: repo-relative module dir the IDE opens and the overlay pins;
  blank keeps the old worktree-root behaviour.
- ideOpenCommand: host command that opens that dir in the IDE at spawn, with
  {dir} substituted and run through /bin/sh -c so env (e.g. DISPLAY) can be set
  inline. Best-effort and non-fatal — a failure never fails the spawn. Blank
  keeps the manual-open behaviour. No close half yet (deferred).

Shared helpers PeerLauncher.ideProjectPath / openInIde back both launchers.
The two Profile fields ride a back-compat constructor, so every existing call
site and YAML compiles and behaves unchanged.

Tests: overlay content pins the module dir when ideProjectDir is set;
ideProjectPath resolution; openInIde no-op on a blank command. 918 tests green.
2026-08-24 06:47:37 +02:00
Dai Ha a5efb7c676 CB-634: write the overlay exclude to the common git dir, not the per-worktree gitdir
git reads info/exclude from the common dir for a linked worktree (only
info/sparse-checkout is per-worktree), so the entry written into
<common>/worktrees/<name>/info/exclude was never honoured and CLAUDE.local.md
showed as untracked -- at risk of being swept into a worker's PR. Derive the
common dir (<common>/worktrees/<name> -> <common>) and write there. Found by
dogfooding a real spawn on fleet01; the test now uses the real worktree layout
and asserts the entry lands in the common dir, not the per-worktree gitdir.
2026-08-23 20:06:36 +02:00
Dai Ha a3fc7e4df8 CB-634: deliver IDE guidance as an on-disk overlay, not the system-prompt charter
Move the IDE guidance text to PeerLauncher.ideOverlayText (shared by both
launchers). ClaudeCodeLauncher drops it from the reply-charter file and writes
CLAUDE.local.md into a provisioned worktree instead, gated on a .git FILE
(safety: never writes into the primary's real .git-DIRECTORY checkout) and
registers it in info/exclude. OpenCodeLauncher mounts the intellij server and
adds the rules file to the instructions array.
2026-08-23 16:37:08 +02:00
Dai Ha d811b30df3 CB-634 (draft): mount the IDE Index MCP into a member, opt-in per profile
Adds `ideMcpUrl` to FleetConfig.Profile (default off). When set, the
Claude Code launcher mounts the IDE Index MCP as a second inline
--mcp-config server named `intellij`, and appends an IDE charter that
pins every ide_* call to the member's own worktree (spec.cwd()). The
charter order is role -> ide -> reply, one --append-system-prompt-file,
reply last (CB-618). The mount gate now fires on ideMcpUrl alone, not
only mcpUrl. ConfigRef treats an ideMcpUrl change as deferred, like the
other launch flags.

Never touches .mcp.json or CLAUDE.md — the mount and the rule arrive as
launch flags, so a project's own config is untouched.

Not yet done (see fleetd #162): the bridged-owned IDE lifecycle
(open on provision, close before worktree removal), and the opencode
adapter (separate ticket). fleetd.example.yaml documents ideMcpUrl and
fixes the stale parityOverlay default.

911 tests green.
2026-08-23 16:00:36 +02:00
Dai Ha 95a8dbcea9 broker: uriEnv config + non-fatal unreachable broker on boot
CI / build (pull_request) Failing after 56s
CI / contract (pull_request) Successful in 1m6s
#151: Broker gains uriEnv beside uri, taking the AMQP URI from an env var so
the password stays out of fleetd.yaml (same pattern as auth.tokenEnv). uriEnv
wins when set; isConfigured() treats a uriEnv naming an unset/blank variable
as unconfigured. A configured uriEnv is added to the startup required-secrets
report. Never logs the resolved URI (it carries the password).

#152: AmqpReplyInbox.open throwing at boot no longer stops the daemon. The
selection at the call site catches the failure and falls back to the in-memory
inbox for the process lifetime, warning loudly that durable cross-restart
delivery is off and logging the failed URI with credentials stripped.
2026-08-23 13:49:26 +02:00
Dai Ha 6f968d59a4 CB-633 round 2: the scrub ran on macOS and did nothing on Linux
CI / build (pull_request) Failing after 1m8s
CI / contract (pull_request) Successful in 1m10s
Six review findings, from the peer lead `vms` and a reviewer worker. The first
one is a real defect that would have shipped as a dead control.

1. The scrub only ran in a login shell. It lived in the generated `.zlogin`,
   and zsh reads `.zlogin` only for a login shell. herdr does not open the same
   kind of shell everywhere: measured on herdr 0.8.0, a macOS pane runs `-zsh`
   (login) while a Linux pane runs a plain `/usr/bin/zsh`. So on the vhost this
   was being built for, `.zlogin` never ran and every member kept the whole
   secret store, in silence.

   The scrub body now lives in a generated `scrub.zsh` that BOTH `.zshrc` and
   `.zlogin` source, each after sourcing its own `$HOME` counterpart. Linux
   runs the first, macOS runs both, and the second pass is not merely harmless
   -- it re-scrubs anything the operator's `~/.zlogin` exported after `~/.zshrc`
   had finished. Re-running is idempotent.

2. `INFRASTRUCTURE_PASSTHROUGH` listed names that are not infrastructure:
   ANTHROPIC_AUTH_TOKEN, GITEA_TOKEN, GITEA_HOST, ANTHROPIC_BASE_URL,
   ANTHROPIC_MODEL, CLAUDE_CONFIG_DIR, OPENCODE_CONFIG, BRIDGED_MEMBER. I read
   each injection point and confirmed every one of them reaches `launch.env()`
   only when actually injected, so `allowed.addAll(launch.env().keySet())`
   already covers the legitimate case. As static entries they were pure leak
   surface: a host that happened to export ANTHROPIC_AUTH_TOKEN would have had
   it passed straight through.

3. A missing `scrub-report.txt` at teardown was logged at debug. The report is
   the only evidence the scrub ran at all. Its absence has an innocent reading
   and a serious one, and we cannot tell them apart from the daemon -- so it is
   now a WARN that says exactly that. Logging it at debug is how a control that
   quietly stopped working stays unnoticed.

4. Nothing tested that the control was wired in. Deleting the single
   `applyEnvironmentAllowListPolicy(cfg, launch)` line left all 896 tests green
   while turning the feature completely off -- the CB-586/CB-611 shape again.
   `HerdrPeerLauncherAllowListWiringTest` starts a real spawn and asserts on the
   env that reached herdr. Mutation-checked: unwiring that line fails it.

5. `EnvAllowListScrubTest` now also runs `zsh -i` with no `-l`, which is the
   Linux pane shape, so finding 1 is tested from a Mac. Mutation-checked:
   putting the scrub back in `.zlogin` alone fails that test alone, while the
   login-shell test still passes -- which is exactly the blind spot that let
   the bug through.

6. Two ZDOTDIR leaks closed. A failed spawn has no pane id, so its directory
   was never keyed for teardown; it is now removed on the way out. And
   `deleteOnExit` covers a clean shutdown and nothing else, so `generate` now
   reaps sibling directories older than 24h left by a killed daemon.

Also: `policy:` is lowercased with Locale.ROOT, and the `.zlogin`-only claim is
corrected in fleetd.example.yaml, FleetConfig and HerdrPeerLauncher.

901 tests, 0 failures, `mvn clean install` green.
2026-08-23 08:17:52 +02:00
Dai Ha d432df8e5c CB-633: memberCredentials policy=allow-list — derived ZDOTDIR env scrub
CI / build (pull_request) Failing after 1m4s
CI / contract (pull_request) Successful in 1m19s
Move member environment control out of the pane-creation env overlay
(defeated by any file the login shell sources) into a per-spawn ZDOTDIR
directory whose .zlogin runs LAST, after the operator's whole chain, and
blanks every exported variable not on an allow-list DERIVED from what the
launcher itself injects (profiles' tokenEnv/gitTokenEnv/gitHostEnv/env
keys + an infrastructure set) — never hand-typed.

- memberCredentials.policy: allow-list (deny-by-default/deny-list stay
  default and unchanged); known:/allow: become reporting only under it.
- memberCredentials.sshAuthSock knob, blocked by default; allowing it is
  an explicit decision (operator ssh-agent handle).
- Non-zsh login shell: loud WARN, protection off, fallback to the old
  enumerated-name overlay.
- Scrub writes an 'allowed N of M' denominator report, read at teardown;
  credential-shaped blanked names go to WARN (names only, never values).
- Equality test against a real login zsh from a clean parent: surviving
  non-empty exports EQUAL baseline ∩ derived allow-list.
2026-08-23 07:43:38 +02:00
Dai Ha 4b822731e6 CB-632: rename the member's MCP mount bridge -> fleet
CI / contract (push) Successful in 40s
CI / build (push) Successful in 1m22s
Every launcher writes the bridge's MCP server into the config it hands its peer,
and it named that server "bridge". So a member addressed its tools as
mcp__bridge__fleet_send while the tools themselves are already fleet_*. The
mount is named "fleet" now, and a member's tools are mcp__fleet__*.

The name was a bare literal in three files: ClaudeCodeLauncher and LeadLauncher
build a --mcp-config JSON string, OpenCodeLauncher writes an opencode.json node.
Three hand-written copies of one name is how a rename lands in two of them, so
the name is now one constant, PeerLauncher.MCP_MOUNT_NAME.

The mount name is local to the peer — it is the label its own client puts on the
server, and nothing in the daemon reads it back. Renaming it changes no wire
call.

Tests. Each launcher's test now asserts the mount is named fleet AND that
nothing writes "bridge"; the second half is the part that would have caught a
half-done rename. LeadLauncherTest never checked the name at all, only the URL,
so it gained the assertion rather than had one updated.

CLAUDE.md's role-detection ladder quoted mcp__bridge__* as the marker of a
spawned member. It names mcp__fleet__* now, and says that a member spawned
before this change still reports the old prefix. The portable block stays
byte-identical with the wiki template (wiki 569a917).

Build: cd bridged && mvn clean install, then read target/surefire-reports/*.xml
directly — 884 tests, 0 failures, 0 errors.
2026-08-23 07:04:44 +02:00
Dai Ha 8101290933 CB-632: rename the exported metrics bridged_* -> fleet_*
Part of #145 (CB-632). Doing this NOW, ahead of the rest of the path
renames, for one reason: the vms lead is about to wire a monitoring
dashboard to these series. Renaming a metric after a dashboard points at
it breaks continuity and silently leaves a dead panel. Renaming it before
costs nothing, so it goes first rather than at the cutover.

Nine series renamed, all declared in FleetMetrics.

Two real defects found while doing it:

  - FleetApp had "bridged_auth_failures_total" written as a LITERAL
    instead of using FleetMetrics.AUTH_FAILURES -- a second hand-written
    copy of a name, which is how these drift. It now uses the constant,
    so there is one source for that name again.
  - Nothing guarded the prefix. One test does assert a wire name
    (FleetAppAuthTest checks the real /metrics body for fleet_sessions),
    which is good, but it covers one series out of nine. The literal
    above was covered by nothing at all.

So this adds MetricNamesTest, which reads the constants reflectively
rather than listing them -- a test that lists the nine names is itself a
second hand-written copy, and would pass while a tenth went unchecked.
It asserts its own denominator too: "no name starts with bridged_" is
true of an empty set, so a sweep that found nothing would pass loudly.
Asserting the count of 9 makes a broken sweep fail instead.

Also renamed three herdr contract-test workspace labels, __bridged_* ->
__fleet_*. Those are throwaway workspaces created by ensureWorkspace, not
metrics, but they are the same word.

Verified: mvn clean install green, 52 classes, 881 tests, 0 failures.
Test count is up by 3 -- the new prefix, denominator and uniqueness
checks. No bridged_ string remains anywhere outside wiki/.
2026-08-23 06:59:37 +02:00
Dai Ha 6e7fc12f89 CB-632 unit 3: point text references at fleetd.example.yaml
CI / contract (pull_request) Successful in 1m17s
CI / build (pull_request) Successful in 1m36s
2026-08-23 06:51:48 +02:00
Dai Ha 50df14a50f CB-632 unit 3: prefer fleetd.yaml, fall back to bridged.yaml; log the chosen file 2026-08-23 06:51:35 +02:00
Dai Ha 9b50dd69d8 CB-632 unit 1: rename the Java package and classes to fleet
Part of #145 (CB-632), under epic #125.

The product is called fleet and the daemon is called fleetd, but the code
still said bridge everywhere. This renames the Java half:

  package dev.ltms.bridged -> dev.ltms.fleet
  Bridged        -> Fleetd          (the main class)
  BridgedConfig  -> FleetConfig
  BridgeMcp      -> FleetMcp
  BridgedApp     -> FleetApp
  BridgedMetrics -> FleetMetrics

The package root is dev.ltms.fleet, not dev.ltms.fleetd. The trailing d
means daemon, which names a process, not a namespace.

What this commit deliberately does NOT change:

  - The module directory stays bridged/, and <finalName> stays bridged.
    The installed launchd plist names bridged/target/bridged.jar and its
    KeepAlive is armed, so renaming the jar on its own strands a restart.
    Both change at the cutover, together with the plist, in one step.
  - The bridge_* MCP tool aliases. CB-622 shipped both names on purpose.
    One test names a local variable viaBridge because it holds the result
    of the deprecated call; the rename collided with it and the compiler
    caught it. That variable is back.
  - BRIDGED_* env var names, and bridged.yaml. Both are operator
    contracts and need a read-both shim, which is a later unit.

Two things a plain search-and-replace would have missed:

  - logback.xml and logback-test.xml name the package twice, once as a
    turboFilter class= attribute. The compiler never checks those.
  - BSD sed does not support \b. The word-boundary expression matched
    nothing and said nothing, while the other ten in the same command
    worked. Checked the leftovers instead of trusting the exit code.

Verified: mvn clean install green, 51 test classes, 878 tests, 0 failures
-- the same count as before the rename.
2026-08-23 06:27:43 +02:00
Dai Ha 41a3114d03 CB-622 Unit A: register fleet_* MCP tools, keep bridge_* working
CI / build (pull_request) Successful in 1m2s
CI / contract (pull_request) Successful in 1m8s
fleet_* is now the documented tool name for all eleven MCP tools; each
bridge_* twin is registered against the exact same handler (no logic
duplication) and its description leads with a DEPRECATED notice. A
bridge_* call logs one WARN naming the old and new name, once per name
for the life of the process (a Set, not a numeric sentinel).

REPLY_CHARTER in HerdrPeerLauncher now tells a spawned member to call
fleet_reply — the one rule that must survive with no repo checkout.

All other bridge_* string literals across mcp/, Javadoc, and tests were
renamed to fleet_* for consistency with the new documented name.
2026-08-22 21:54:32 +02:00
Dai Ha cc47672b7c CB-618: one charter file, and agent files need a name:
CI / contract (push) Successful in 1m10s
CI / build (push) Successful in 1m27s
Two defects in CB-617 that only a live spawn could find. Both were shipped
green: every test passed because every test read the argv we built, and none
ran the binary that has to accept it.

1. Claude Code refuses to start when both --append-system-prompt and
   --append-system-prompt-file are on the command line:

     Error: Cannot use both --append-system-prompt and
     --append-system-prompt-file. Please use only one.

   CB-617 put the role charter on the file flag and left the reply charter on
   the inline flag, so every Claude-profile spawn with a role charter died at
   launch. The pane exited on its own and bridged reported it as
   spawn_timeout ("did not reach injectable state within 20000ms"), which
   hides the real cause.

   Both charters now go in the one file, role charter first and reply charter
   last — last is where the reply rule must sit, because it is the rule that
   must survive. A member with only a reply charter keeps the proven inline
   flag, which is also the only form that reaches a member with no repo
   checkout.

2. A Claude Code agent definition needs `name:` in its frontmatter. Ours had
   only `description:`, so the files were skipped and --agent architect failed
   with "not found. Available agents: claude, Explore, ...". Added to all
   three. The OpenCode files take their name from the filename and are
   unchanged.

Checked on this host, in this repo, with the real binary:

  claude --model claude-sonnet-5 --agent architect \
    --append-system-prompt-file /tmp/combined.md -p '...'
  -> ROLEOK, REPLYOK, yes

so the agent definition, the role charter and the reply charter all compose.

The updated test now asserts the constraint that actually binds: with a role
charter present, --append-system-prompt must be absent, and the file must open
with the role charter and end with the reply charter.

874 tests pass.
2026-08-22 12:32:44 +02:00
Dai Ha 7e97f5bff5 CB-617 Unit A: role charter via file, not argv
CI / build (pull_request) Successful in 1m16s
CI / contract (pull_request) Successful in 1m28s
herdr refused to shell-encode a multi-line inline --append-system-prompt
argument (invalid_agent_argument), which broke any Claude Code profile with
a configured multi-line fleet.charters.<role>. Split delivery: the role
charter now writes to a temp file and mounts via
--append-system-prompt-file; the one-line REPLY_CHARTER keeps its inline
--append-system-prompt delivery, since it must reach a member with no repo
checkout. Also pass --agent <role> when a role's agent-definition file
exists under the worker's cwd (.claude/agents/<role>.md for claude-code,
.opencode/agent/<role>.md for opencode); absent either file, nothing extra
is added and the member still spawns.

The base's LaunchSpec now carries roleCharter/replyCharter/cwd alongside
the existing composed charter field, so CharterReceipt keeps fingerprinting
the same composed text it always did -- unchanged, since the digest covers
the full logical charter content regardless of how it is delivered.
2026-08-22 12:04:38 +02:00
Dai Ha 7930a31b94 CB-596 round 2: the exec-time argv-prefix fix has no seam — stop and report
CI / contract (pull_request) Successful in 47s
CI / build (pull_request) Successful in 1m6s
herdr protocol 19's agent.start takes a fixed `kind` (herdr resolves the
executable) plus trailing CLI args for that binary; only tab.create/pane.split
accept an env map, and that IS the round-1 pane-creation overlay already
shipped. There is no argv/env control point that runs after the pane's login
shell and before the agent process starts, so the proposed `env NAME=value ...`
argv prefix cannot be implemented against this API. Documented the finding and
corrected bridged.example.yaml's round-1 comments, which had overclaimed that
the overlay survives the login shell.

Kept everything else: added a startup WARN (Bridged.reportMemberCredentialsGap)
when memberCredentials: is absent or its known: list is empty, so CB-592's
protection loss is never silent, mirroring CB-594's reportRequiredSecrets.
2026-08-17 09:13:57 +02:00
Dai Ha 850fb12807 CB-596: member credential blocking is config-driven deny-by-default, not one hardcoded name
CI / build (pull_request) Successful in 1m1s
CI / contract (pull_request) Successful in 1m26s
Replaces CB-592's single hardcoded GITEA_ACCESS_TOKEN shadow in HerdrPeerLauncher
with BridgedConfig.MemberCredentials (memberCredentials: policy/allow/known).
Every known name not also allowed is overlaid with a sentinel; an unrecognized
policy value refuses at load; a credential-shaped host env var on neither list
is logged as a gap (name only, never a value). bridged.yaml is gitignored, so
the 31 measured names + 4-name allowlist ship as a commented block in
bridged.example.yaml for the operator to apply live.
2026-08-17 09:02:21 +02:00
Dai Ha b14b66ab03 CB-586: the retention sweep never ran — Long.MIN_VALUE overflowed the gate
CI / contract (push) Successful in 1m20s
CI / build (push) Successful in 1m36s
SessionReaper.lastWipSweepNanos started at Long.MIN_VALUE as a "never swept
yet" sentinel. That sentinel cannot be compared by subtraction. nanoTime() is
positive on this platform, so `now - Long.MIN_VALUE` wraps to a large negative
number, the gate `delta < WIP_SWEEP_INTERVAL_NANOS` reads it as "swept moments
ago", and the method returns before the assignment that would have fixed the
field. The sweep never ran once, for the life of the process, and nothing in
the log said so.

Measured:
  System.nanoTime()      = 31305820625625   (positive)
  now - Long.MIN_VALUE   = -9223340731034150183
  interval (6h in nanos) = 21600000000000
  gate 'delta < interval' -> true  => returns early, every iteration, forever

Fix: a separate `sweptOnce` boolean holds "never yet", so the subtraction only
runs once both operands come from nanoTime. The first pass always sweeps — a
restart is a fine moment for it, the 24h age floor keeps it safe, and the
feature becomes observable right after a redeploy instead of six hours later.

The existing tests all passed because they call SessionManager.sweepWipRefs
directly, which walks around the gate. The new test asserts through the reaper
loop instead: it spawns a worktree session, starts the reaper, and requires a
real prune call at the seam with the 24h floor intact. Removing the fix makes
it fail with "it never reached the seam".

860 tests, mvn clean install, BUILD SUCCESS.
2026-08-16 20:13:37 +02:00
Dai Ha 15ff6bcde5 Merge CB-586: prune refs/wip snapshots whose content is already on main
A refs/wip snapshot ref is deleted only when both hold: its commit's tree is
already reachable from main, and it is older than 24h. Reachability is the safety
floor — a snapshot exists because the work was committed nowhere else, so an
unreachable one is the last copy and is never swept. Every deletion logs the ref
and the sha.

/members gains wipRefs{count,costBytes} so the growth is visible.

Verified against real git, not only the fakes: a recoverable+old ref is deleted,
a recoverable+young one survives the age floor, and the last copy survives. A repo
with no main deletes nothing, and a repo with no snapshots is a clean no-op.

Closes #67
2026-08-16 20:09:16 +02:00
Dai Ha 0efe1567c0 CB-586: prune refs/wip/* older than 24h whose tree is reachable from main
CI / build (pull_request) Successful in 1m13s
CI / contract (pull_request) Successful in 1m21s
Add the CB-586 retention rule to GitWorktrees and drive it from the reaper:
a snapshot is deleted only when its tree content is already reachable from
main AND the ref is older than 24h. Reachability keeps the last copy of a
worker's work; the age floor stops a fresh snapshot being swept while a
lead is still looking at it. Every deletion logs the ref name and commit
sha so it is recoverable from the reflog. The /members response gains a
wipRefs{count,costBytes} census the operator can read without shelling
into the repo.
2026-08-16 19:02:33 +02:00
Dai Ha bb750cdba3 CB-606: refuse an unrecognized auth.mode, per-profile placement, or top-level placement policy at config load
CI / build (pull_request) Successful in 1m7s
CI / contract (pull_request) Successful in 1m25s
An auth.mode typo (e.g. "toekn") used to silently fall back to loopback-trust with no signal
anywhere — validateAuthExposure() only checks the pairing on a non-loopback bind, so on the
common loopback bind the daemon started cleanly and authenticated nobody. Per-profile placement
had the same shape, falling back to legacy pane placement. The top-level placement policy name
was already validated by PlacementPolicies.fromName, but only lazily at first spawn through
CompositePeerLauncher's Supplier; it is now checked eagerly at load, calling fromName itself as
the single source of truth.
2026-08-16 18:54:35 +02:00
Dai Ha d56c77b368 CB-582: close the question when ask() leaves by throwing
CI / contract (push) Successful in 48s
CI / build (push) Successful in 1m38s
Found reviewing CB-582 before the merge, not by the implementer.

ask() clears its question on three paths — no-waiter, timed out, and (from
answer()) answered. It can also leave by throwing: an interrupt while blocked on
the answer, or an ExecutionException from the answer future. Those run only the
finally block, which tore down the rendezvous turn but not the push loop's copy.

The result was a question that stayed pending for good: named in every nudge
until it hit its own cap, then left in pendingQuestions with no remover at all.

Teardown now happens where the rendezvous teardown already happens, so the two
cannot drift apart again. Closing a turnId that was never pending is a no-op, so
the normal paths are unaffected.

The new test fails on the pre-fix code with expected: <STOP> but was: <INJECT>.
2026-08-16 18:47:18 +02:00
Dai Ha 83e2ff06cf Merge CB-582: nudge the lead when a worker pauses on bridge_ask
An async (wait:false) delegation opens a ~55s reverse-rendezvous window when its
worker calls bridge_ask. A lead polling on its normal minutes-long cadence never
sees that window, so the worker times out and proceeds without an answer.

The question is now a third source in the CB-588 per-lead push schedule, with its
own per-item nudge count (CB-598's shape), and it is surfaced by bridge_status and
by REST /sessions/{id}/status and /tasks/{ticket} (which previously dropped turnId
on an ASKING phase, so a REST caller could see the question but not answer it).

Closes #61
2026-08-16 18:45:14 +02:00
Dai Ha fe46311266 CB-604: reject an unrecognized profile kind at config load
CI / contract (pull_request) Successful in 53s
CI / build (pull_request) Successful in 1m10s
2026-08-16 18:37:21 +02:00
Dai Ha 08968bb1b7 CB-582: make a pending bridge_ask question visible on the lead's poll cadence
CI / build (pull_request) Successful in 1m23s
CI / contract (pull_request) Successful in 1m23s
bridge_ask blocks the worker's turn for ~55s by default (BridgedApp.java,
BridgeMcp.java) — a value deliberately kept just under the worker's own MCP
client's ~60s call cap so the daemon can return a clean timeout before the
client severs the call, not a value that can usefully be widened. A lead
following the charter's wait:false + poll cadence is minutes away, so the
window closes long before a poll would ever see the question — and until now
bridge_poll on such a ticket just read as ordinary "pending" progress.

bridge_poll(ticket) already surfaced Phase.ASKING with the question and
turnId (CB-205); this ships the two pieces that were still missing:

- The lead's own pane is now nudged the instant a question opens, reusing
  the CB-588 ReplyPushLoop push mechanism (a third source alongside queued
  replies and terminal tickets) rather than a new path. The nudge is capped
  by the loop's existing maxReminders budget, and stops the moment the
  question is answered or lapses.
- bridge_status(sessionId) and REST GET /sessions/{id}/status now also show
  an open question and how to answer it, via a new
  MessageService.pendingAsk() lookup — covering the case where a lead checks
  status directly rather than the ticket.
- The REST /tasks/{ticket} endpoint was silently missing turnId on an ASKING
  phase (only the MCP layer's formatted text carried it) — fixed as part of
  making the state genuinely visible over both surfaces.

An unanswered question still behaves as today: the worker proceeds and its
reply says the ask went unanswered — not a hard failure.
2026-08-16 18:35:45 +02:00
Dai Ha cc1df11f69 CB-584: carry agentSessionId on a failed ticket's ReleaseDetail
CI / contract (pull_request) Successful in 1m6s
CI / build (pull_request) Successful in 1m42s
Closes the one piece issue #65 deliberately left out of 5d5b3bd: a
released member's agentSessionId now rides alongside worktree, branch
and snapshotRef on ReleaseDetail, and Bridged's onRelease handler
names it in the abandon reason, so a lead can resume the member's
conversation instead of only re-dispatching a fresh one onto the same
files.

Also updates the CLAUDE.md bridge_spawn/bridge_list table row, which
5d5b3bd shipped the sessionName/resumeSessionId/agentSessionId surface
for but never updated.
2026-08-16 18:25:24 +02:00
Dai Ha d5dd5639ae Merge CB-602: guard against a config key that never reaches the example
bridged.yaml is gitignored, so bridged.example.yaml is the only
committed description of the config schema. Two tests already covered
example -> code; nothing covered code -> example, so a brand-new key
could ship undocumented and no test would notice.

A new test compares BridgedConfig.KNOWN_TOP_LEVEL_KEYS against the
example scanned as TEXT, so a key documented only as a comment counts as
documented. That is what makes the guard correct rather than annoying:
most of the example is commented on purpose.

Verified here: added an undocumented key and watched the test fail with
an actionable message naming it; then documented that key as a comment
only and watched it pass. Probe reverted, tree clean.

Closes #96
2026-08-16 18:17:11 +02:00
Dai Ha 7a120b3256 Merge CB-598: per-item reminder counts so backoff-window work is never orphaned
CI / contract (push) Successful in 45s
CI / build (push) Successful in 1m39s
The reminder count was one counter per lead per source, carried forward
across ticks. A counter carried forward has no memory of which item it
counted, so work arriving during the backoff window inherited an
already-capped count and was never named in a nudge.

tick() now recomputes each source's count fresh from the minimum count
among the items actually pending, tracked per item. A fresh item keeps
its source eligible; an older capped item still rides along in the text
without spending more budget. decide() is unchanged.

Verified here: read the diff; the bumped set is exactly the set named in
the nudge, and the empty early-return skips the bump. Trial merge onto
main builds 826 tests BUILD SUCCESS, unpiped.

Closes #87
2026-08-16 18:13:29 +02:00
Dai Ha 863d477966 CB-603: make FakeHerdr.calls thread-safe
Background loops call the fake from their own scheduler threads while a
test polls called() from the test thread. The list was a plain
ArrayList, so a nudge landing mid-stream threw
ConcurrentModificationException out of called().

It surfaced while I was verifying CB-598, which nudges more often, but
the race is on main today and is unrelated to that change.

824 tests, BUILD SUCCESS.
2026-08-16 18:12:23 +02:00
Dai Ha a36b7ccd7c Merge CB-601: make the recovery-race test deterministic
CI / build (push) Successful in 1m5s
CI / contract (push) Successful in 1m7s
The test asserted one of two interleavings that are both correct, and
steered toward it with a 5 ms Thread.sleep. Under load the other
interleaving happened and main went red on a correct implementation.

The head start is now a latch counted down from inside the sweep's
guarded loop, so the ordering is guaranteed, not likely. Test file only;
the production guard is unchanged.

Verified here: 822 tests BUILD SUCCESS; 10/10 passes while a full clean
install ran in parallel; 5/5 failures with the guard removed, so the test
still catches the bug it exists for.

Closes #95
2026-08-16 18:08:18 +02:00
Dai Ha 32bf324a1e CB-602: guard against a config key that never reaches the example
CI / contract (pull_request) Successful in 49s
CI / build (pull_request) Failing after 1m24s
BridgedConfig.KNOWN_TOP_LEVEL_KEYS is now package-private so a test can assert
every key the parser accepts appears in bridged.example.yaml — live or
commented-out, since the file is gitignored and the example is the only
committed description of the config schema. The existing tests only checked
the example->code direction; this adds code->example.
2026-08-16 18:06:36 +02:00
Dai Ha 16de9df000 CB-601: make the recovery-race test's head start deterministic, not a sleep
CI / contract (pull_request) Successful in 1m8s
CI / build (pull_request) Successful in 1m9s
2026-08-16 18:02:57 +02:00
Dai Ha 81a0cf4710 CB-598: track reminder counts per pending item, not per lead per source
CI / build (pull_request) Successful in 1m12s
CI / contract (pull_request) Successful in 1m11s
Work that arrived during the ~15s push_backoff_ms window between two
ticks landed in the pending map before the next tick's start-of-tick
snapshot, so a shared per-lead-per-source counter (carried forward via
scheduleNext(lead, count+1, ...)) already treated it as exhausted
backlog even though no nudge had ever named it. ReplyPushLoop.tick now
recomputes each source's reminder count fresh every tick as the
minimum nudge count among that source's currently pending items, so a
freshly-arrived item (count 0) keeps its source eligible regardless of
how depleted an older, still-undrained sibling's count is. decide()
itself is unchanged.
2026-08-16 17:49:21 +02:00
Dai Ha 8a837a2830 CB-599: surface a capacity refusal's reason instead of a bare 500
CI / contract (pull_request) Successful in 50s
CI / build (pull_request) Successful in 1m35s
PlacementException extends IllegalStateException, which neither BridgedApp
nor BridgeMcp's spawn catch blocks handled, so a maxLoad/quarantine/
all-exhausted refusal fell through to a blank 500 on REST and lost its
message on MCP. Catch it on both surfaces, ahead of the unrelated
IllegalArgumentException(unknown_profile) mapping, and return its message
structured: REST as {"error":"no_capacity","detail":...} with status 503,
MCP as an isError result prefixed "no capacity: ...".
2026-08-16 17:43:07 +02:00
ltms 613ece92dc Merge CB-594: make supervision and a working fleet possible at the same time
CI / contract (push) Successful in 49s
CI / build (push) Failing after 1m45s
The launchd unit was a CHANGEME template that had never been installed, and it could not have worked if it were: launchd does not source a login shell, so the daemon would have started with no forge or gateway token, and the failure would only appear much later as workers unable to open a PR.

Four parts: a wrapper that execs one login shell in place so the job inherits the secret store; a startup report naming which required secret env vars resolved and which are MISSING, by name only, never a value; the plist filled in with this host's real verified paths; and `redeploy-bridged.sh` detecting the agent and switching stop/start to `launchctl unload -w` / `load -w`, falling back to the existing kill + nohup when it is not installed.

Verified by the lead. My own unpiped build of the branch: 813 tests, BUILD SUCCESS. Trial-merged onto current main (which had moved twice) and rebuilt: Tests run: 822, Failures: 0, Errors: 0, Skipped: 0 — BUILD SUCCESS. All three host paths in the plist exist; no CHANGEME remains; the secret report leaks no values.

An independent reviewer, briefed only from the diff, verified the parts that matter today and said merge. It confirmed live that the unsupervised path is unchanged, that `launchctl list` exits 113 for an absent agent so the detection reads correctly, that the wrapper round-trips arguments containing spaces and quotes and stays a single exec chain, and that neither branch of the secret report can print a value.

Its two open findings only bite once the agent is actually loaded, which has not happened and is the operator's call. Filed as #91 — that must land before the agent is ever installed. The important one is that the script computes its log path from its own location while the plist hard-codes an absolute one; if they ever disagree, the post-restart ERROR check reads the wrong file and reports "ok" while the daemon crash-loops.

The agent is deliberately NOT installed by this merge. Nothing here changes how the daemon runs today.
2026-08-16 17:32:57 +02:00
ltms 8d4206c2b5 Merge CB-590: one nudge schedule per lead, with a reminder budget per source
CI / contract (push) Successful in 1m7s
CI / build (push) Failing after 1m15s
Carries PR #84 (its head `78ca24d` is an ancestor of this one), so this single merge delivers both rounds.

CB-590 collapses the CB-307 reply schedule and the CB-588 ticket schedule into one per lead, which closes the double-injection race. The review round found that this also collapsed the two reminder caps into one shared budget — so a busy reply stream could exhaust the cap and a ticket arriving afterwards would never be nudged at all. That was a regression, not a pre-existing wart: before this PR the two sources had independent counters.

The follow-up keeps the single schedule and gives each source its own budget. `decide()` returns INJECT while either source has pending work under its own cap, and STOP only when neither does. `stopOrRestart` and its snapshot-diff race logic are untouched.

Verified by the lead: `mvn -f bridged/pom.xml clean install` unpiped, exit code captured — Tests run: 814, Failures: 0, Errors: 0, Skipped: 0 — BUILD SUCCESS. `ReplyPushLoopTest`: 41 tests.

Known and deliberately out of scope: an item arriving during the 15s backoff is already inside the tick's "before" snapshot, so at-cap work can be abandoned rather than nudged once. That shape predates this PR in the ticket-only loop and is filed as #87 (CB-598).
2026-08-16 17:28:57 +02:00
Dai Ha 88b9503c3b CB-590 follow-up: give each nudge source its own reminder budget
CI / contract (pull_request) Successful in 43s
CI / build (pull_request) Successful in 1m46s
decide(lead, reminderCount) shared one counter across the reply and
ticket sources after PR #84 collapsed both onto a single per-lead
schedule. A reply stream that used up the whole budget could then
make decide() STOP even for a ticket that had never been nudged and
had coalesced onto the same still-active schedule — stranding it with
no live schedule left, since stopOrRestart's racedIn check does not
save work that was already present in the "before" snapshot.

decide() now tracks a per-source count (replyReminderCount,
ticketReminderCount) and returns INJECT while either source is still
under its own cap, STOP only when both are exhausted. Still exactly
one schedule per lead; stopOrRestart's snapshot-diff logic is
untouched.
2026-08-16 17:26:07 +02:00
Dai Ha c553d795d8 CB-528: close the recovery race in AmqpReplyInbox
CI / contract (pull_request) Successful in 44s
CI / build (pull_request) Failing after 1m20s
failPendingPublishesOnRecovery walked and cleared pendingBySeq/pendingByMsgId
without holding publishChannelLock, so a publish() that registered while the
sweep was still iterating could be failed even though it published on the
already-recovered channel — a successful publish reported as failed, and
since MessageService.reply mints a fresh msgId per retry, dedup can't catch
the resulting duplicate. Guard the sweep with publishChannelLock: publish()
only holds it for the seq/map-put/basicPublish, so the sweep can only ever
wait for an in-flight basicPublish to return, never a broker round trip.

Also make close() fail in-flight publishes immediately with a clear message
instead of leaving them to idle out the 10s confirm timeout, and record the
(currently unreachable) msgId-uniqueness assumption pendingByMsgId relies on.

AmqpReplyInboxRecoveryRaceTest drives the sweep and a real publish() against
each other directly (no live broker reconnect) using Proxy-backed fake AMQP
channels and a large in-flight backlog to make the race window observable;
confirmed it fails without the guard (reply "fresh" wrongly failed as
"connection recovered mid-publish") and passes with it.
2026-08-16 17:24:43 +02:00
Dai Ha 3ba6d6784c CB-594: make supervision and a working fleet possible at the same time
CI / contract (pull_request) Successful in 44s
CI / build (pull_request) Successful in 1m40s
Adds scripts/bridged-launchd-wrapper.sh so the launchd-run daemon still gets
WORKER_GITEA_TOKEN/AI_GATEWAY_TOKEN by execing through a login shell (launchd
never sources secrets.sh itself). bridged now logs at startup which required
token env vars (derived from each profile's tokenEnv/gitTokenEnv, not a
hand-written list) resolved or are MISSING, by name only. Fills in the real
paths in deploy/dev.ltms.bridged.plist for this host and points it at the
wrapper. scripts/redeploy-bridged.sh now detects a loaded launchd agent and
uses launchctl unload/load instead of a raw kill+nohup, because a bare
SIGTERM exits this JVM at 143 (measured) which KeepAlive.SuccessfulExit=false
reads as a crash and would race the script's own restart; --check reports
installed/loaded state and stays read-only.
2026-08-16 17:13:17 +02:00
Dai Ha 78ca24dc3f CB-590: collapse the CB-307 and CB-588 nudge schedules into one per lead
CI / build (pull_request) Successful in 1m0s
CI / contract (pull_request) Successful in 1m33s
Both reply-queued and ticket-terminal nudges could independently decide
to inject into the same lead pane in the same window, since they ran as
two separate schedules keyed differently (worker target vs. lead) that
never checked each other. Replace both with a single per-lead schedule
(activeLeads) that drains pending reply targets and pending tickets
together, sends at most one combined nudge per tick, and shares one
reminder cap across both sources — so two injections into the same pane
can no longer overlap, and work queued while the lead is busy is never
lost, only deferred.
2026-08-16 16:56:04 +02:00
Dai Ha 4fa6553db5 CB-527/CB-528: bound AMQP prefetch and confirm publishes before claiming durable
CI / build (pull_request) Successful in 1m3s
CI / contract (pull_request) Successful in 1m15s
CB-527: basicQos(prefetch) on the consume channel before basicConsume, configurable
via broker.prefetch (default 32), so an undrained inbox backlog stays on the broker
instead of growing the JVM heap without limit.

CB-528: publish moves to its own confirm-mode channel with mandatory=true and a
return listener, so an unroutable or unconfirmed reply now throws instead of
vanishing silently. The confirm callback checks the per-message returned flag
(set by the return listener, which the broker always fires before the matching
confirm) so an acked-but-returned publish is still reported as a failure. The ack
path stays on its own channel/lock and never waits on a publish confirm.
2026-08-16 16:55:40 +02:00
Dai Ha 0331ecd5d3 CB-592: add the BRIDGED_MEMBER marker — the sentinel alone cannot hold
CI / contract (push) Successful in 1m5s
CI / build (push) Successful in 1m39s
Live check on a member pane showed the CB-592 shadow did NOT take effect:
GITEA_ACCESS_TOKEN inside the pane was still the real admin token.

Measured cause. The overlay itself works — GITEA_TOKEN is injected the same
way, is exported by no shell file, and does reach the pane. The sentinel loses
one step later. A herdr pane runs a LOGIN shell, ~/.zprofile line 41 sources
${SHARED_ENV}/tools/secrets.sh, and that file does a plain unconditional
`export GITEA_ACCESS_TOKEN=...`. A login shell overwrites a value already in
the environment, so the real token is put back before the member starts.
Confirmed directly:

    GITEA_ACCESS_TOKEN=cb592-sentinel zsh -lc ...
    -> RESULT: sentinel was OVERWRITTEN by the login shell

This defeats any launcher-side overlay for any name secrets.sh exports. No
change in this repo can win it alone.

So this adds the half that does survive: BRIDGED_MEMBER=1, a name secrets.sh
never exports. It is a no-op until the operator guards the export:

    [ -n "${BRIDGED_MEMBER:-}" ] || export GITEA_ACCESS_TOKEN=...

Setting it now costs nothing and makes that one line the whole remaining fix.
The sentinel stays: it is correct for any peer kind whose pane does not start
a login shell, and it keeps the intent explicit where every adapter passes.

Also corrects the javadoc and the test javadoc, which both claimed a
protection that was measured not to hold.

The other reported failure was my own bad test, not a regression. The probe
called /api/v1/user, which a minimal write:repository token cannot read. Same
token on the repo endpoint answers 200, so CB-302 is intact:

    GITEA_ACCESS_TOKEN: /user=200  /repos/lms/claude-bridge=200
    WORKER_GITEA_TOKEN: /user=403  /repos/lms/claude-bridge=200

Tests 805 -> 807. Both new tests proved to discriminate by reverting the
marker: everySpawnMarksThePaneAsAMember and
aProfileEnvEntryCannotClearTheMemberMarker both fail without it.

Refs: gitea #77
2026-08-15 18:37:52 +02:00
Dai Ha 3db5277ae8 CB-592: shadow the admin GITEA_ACCESS_TOKEN in every member's herdr overlay
CI / contract (pull_request) Successful in 1m12s
CI / build (pull_request) Successful in 1m37s
herdr spawns a pane from its own login-shell process env and layers our map on
top, so any key baseEnv never mentions passes straight through — including the
admin forge token. baseEnv now puts a non-blank sentinel for
GITEA_ACCESS_TOKEN, applied after the profile's own env: so no profile can
restore it. One place, every adapter, every profile including future ones.
CB-302's GITEA_TOKEN grant (applyGitToken) is untouched.
2026-08-15 18:25:16 +02:00
Dai Ha ac044e7573 CB-588 round 3: close the STOP-vs-onTicketTerminal race, pin nudge polarity, fix comment
CI / build (pull_request) Successful in 1m19s
CI / contract (pull_request) Successful in 1m31s
- ReplyPushLoop.stopOrRestartTicketLoop: after releasing a lead's active-schedule
  slot on STOP, restart only if a ticket landed that the pre-decision snapshot
  did not already account for. A naive "restart on any pending ticket" version
  was tried first and reverted: it defeated the reminder cap by restarting
  forever on a stale, never-collected ticket (broke
  successfulTicketNudgeIncrementsDelivered and ticketNudgesSendUpToCapThenStop).
  Diffing against a pendingBefore snapshot distinguishes a genuine race arrival
  from stale cap-exhausted backlog.
- Two new ReplyPushLoopTest cases exercise stopOrRestartTicketLoop directly
  (now package-private) rather than forcing the underlying thread race:
  aTicketStillPendingWhenTheLoopStopsIsNotStranded (the race must restart) and
  aStaleUncollectedTicketAtCapDoesNotRestartTheLoop (the cap must still hold).
  Both were verified to fail against deliberately-reverted versions of the fix
  before being restored to green.
- MessageServiceTest: pin the success-path nudge test's negative direction too
  (must not contain "FAILED"), not just the failure-path test.
- MessageService: complete the whenComplete comment's list of completion paths
  (TIMED_OUT/BUSY/BACKEND_EXHAUSTED via finishAsyncTask, completeExceptionally
  on throw, and answer() -> finishAsyncTask(turnId, result)).
2026-08-15 17:02:05 +02:00
Dai Ha 6ebad2a91f CB-588 follow-up: reclaim a pruned ticket's pendingTickets entry too
CI / contract (pull_request) Successful in 45s
CI / build (pull_request) Successful in 1m0s
tasks is the sole authority on whether a ticket exists, but
pruneTerminalTickets dropped entries from it without telling
ReplyPushLoop. ticketCollected(ticket) was only ever called from
MessageService.poll's terminal branch, which pruneTerminalTickets
short-circuits past once a ticket is gone (poll returns null at the
top). A ticket the lead never polled — or one the reminder cap already
gave up on — was pruned from tasks but never collected in
ReplyPushLoop.pendingTickets, so it rode along on every later nudge to
the same lead forever, naming a ticket bridge_poll could no longer
find, and the map itself never shrank.

pruneTerminalTickets now calls pushLoop.ticketCollected for every
ticket it actually removes (guarded on pushLoop != null), so a pending
nudge entry lives exactly as long as its ticket is pollable. Reused
tasks as the only removal trigger rather than adding a second live
query back into MessageService — no new source of truth.

Added an injectable clock (LongSupplier nowNanos, defaulting to
System::nanoTime) to MessageService, mirroring the SessionManager/
SessionReaper nowNanos seam, so a test can cross the 10-minute
TICKET_TTL_NANOS deterministically instead of sleeping for real.
Confirmed the new regression test fails against the prior
pruneTerminalTickets (a stale ticket rides along on a later coalesced
nudge) before restoring the fix.
2026-08-15 16:41:23 +02:00