Compare commits

...

81 Commits

Author SHA1 Message Date
Dai Ha 0efe1567c0 CB-586: prune refs/wip/* older than 24h whose tree is reachable from main
CI / build (pull_request) Successful in 1m13s
CI / contract (pull_request) Successful in 1m21s
Add the CB-586 retention rule to GitWorktrees and drive it from the reaper:
a snapshot is deleted only when its tree content is already reachable from
main AND the ref is older than 24h. Reachability keeps the last copy of a
worker's work; the age floor stops a fresh snapshot being swept while a
lead is still looking at it. Every deletion logs the ref name and commit
sha so it is recoverable from the reflog. The /members response gains a
wipRefs{count,costBytes} census the operator can read without shelling
into the repo.
2026-08-16 19:02:33 +02:00
Dai Ha 2124e043ce CB-593: correct the member MCP claim — measured, not assumed
CI / contract (push) Successful in 43s
CI / build (push) Successful in 1m19s
CLAUDE.md told every member 'You mount only the bridge MCP' and told the lead
'a worker mounts only the bridge MCP and cannot run your other tooling'. Both
were false for Claude Code members.

Measured by spawning one member per backend and asking each what it actually has:

  opencode (gx)      11 bridge tools only                     claim TRUE
  claude-code (local) 11 bridge + 45 gitea + 2 context7       claim FALSE

Source is ~/.claude.json user-scope mcpServers; --mcp-config adds to that scope
rather than replacing it, so the worktree parity overlay (which correctly
neutralises .mcp.json and opencode.json) cannot see or stop it.

The forge tools are mounted but not usable: CB-592's blocked sentinel means
get_me and list_issues both fail with 'invalid username, password or token'.
That is defence in depth working in a path it was not designed for, so the text
now says a mounted tool is not a working tool rather than pretending the tools
are absent.

Block propagated byte-identically to wiki 7-Use-Cases.md (wiki 1d95e3f).
2026-08-16 09:46:42 +02:00
Dai Ha e4f3620acb CB-591: record the final 86400s route timeout and the request:0s trap
CI / contract (push) Successful in 44s
CI / build (push) Successful in 1m19s
systems/vms moved the LLM route timeout again, 1800s -> 86400s (24h), after
the silent-truncation risk was discussed. They tried request: 0s first: it
removes the total-duration timer, but on an AIGatewayRoute the idle timeout
is derived from the request timeout, so 0s also removed any bound on a
stalled connection.

At 86400s our own MessageService.ASYNC_TIMEOUT_MS (30 min) binds first, so a
runaway request now ends as a clean FAILED ticket we raised instead of a
silently truncated 200. While the gateway sat at 1800s the two numbers were
equal and did not nest.
2026-08-15 21:29:49 +02:00
Dai Ha 032a59a34d CB-591: fleet moved onto the gateway — both ceilings fixed and re-verified
CI / contract (push) Successful in 40s
CI / build (push) Successful in 1m19s
`local` runs on /anthropic and `gx` on /v1, both weight 100; `local-direct`
stays weight 0 as the escape hatch.

systems/vms fixed both blockers, and each was re-checked from this side rather
than taken on trust:

    listener buffer    32 KiB -> 32 Mi   ours: 1.2 MB body -> 200 (was 413)
    LLM route timeout  60s    -> 1800s   ours: 101s stream -> 200,
                                               message_stop present, 4000/4000

Neither was deliberate: 32 KiB was Envoy Gateway's default
per_connection_buffer_limit_bytes, and 60s was Envoy AI Gateway's own default.
The 60s bounded GENERATION as well as prompt size — a tiny prompt with a long
answer returned 504 at 60.05s.

Verified with real workloads, not liveness probes. A `local` member read this
document and CLAUDE.md in full — 48,344 bytes of file content, comfortably past
the old 32,768 ceiling — and answered four questions correctly, including the
document's length (said ~456, actual 455). A `gx` member did the same. The
trivial 3-question probe is what hid the 32 KiB ceiling for an afternoon, so it
no longer counts as proof here.

§7.2 is new and is the part that matters later. One risk is ACCEPTED, not
solved: on a mid-response timeout over chunked HTTP/1.1, Envoy ends the chunked
encoding cleanly instead of resetting, so a truncated answer arrives as HTTP 200
with no error and no terminator (envoyproxy/envoy#17186, acknowledged 2021,
never fixed; the Dec 2025 fix #42269 is HTTP/2 only and SSE here is HTTP/1.1).
Measured at the old 60s: 200, 61.07s, 2473 of 4000 emitted, message_stop 0,
error events 0, ending on a well-formed frame.

The recommended defence — reject a stream with no terminator — does NOT
transfer to us: Claude Code and opencode are third-party clients and we do not
own their SSE parsing. So this is acceptable because a request would have to run
1800s to trip it, not because we could detect it. If a member ever returns a
confident but truncated answer, suspect this before anything in our own code.

Also recorded, from the upstream bisection: ClientTrafficPolicy is honoured in
standalone `aigw run` but BackendTrafficPolicy is silently ignored, and nothing
external distinguishes them (envoyproxy/gateway#9513). Same silent-default shape
this repo keeps hitting.

bridged.yaml carries the same notes inline (gitignored, so not in this commit).

Refs: gitea #76
2026-08-15 20:47:41 +02:00
Dai Ha e689090024 CB-591: correct the root cause — Envoy's buffer limit, not Caddy
CI / contract (push) Successful in 1m7s
CI / build (push) Successful in 1m35s
I wrote "Caddy request_body max_size and/or Envoy's own" and marked it
unverified. The Caddy half was wrong, and an unverified guess still points the
next reader at the wrong component.

Confirmed by the systems/vms side: Envoy Gateway defaults a listener's
per_connection_buffer_limit_bytes to 32768, and aigw buffers the WHOLE request
body before it can route on the model name. So that default is not a network
tuning knob — it is a hard ceiling on prompt size. From the live config_dump:

    listener default/llm/http    per_connection_buffer_limit_bytes: 32768

Nobody chose 32 KiB; it was inherited.

Both TLS edges are innocent, and the technique that showed it is better than
mine: both 413s carry x-llm-consumer, a header their auth proxy sets only AFTER
authenticating, so the body cleared both edges and the auth. On llm.vm, aigw
413s at 39 KB while the vLLM backend answers 200 at the same size. I found the
boundary; they found the component, by reading the failure's response headers.

Consequences recorded in the doc:

  * DO NOT plan around 32 KiB. The intended ceiling is far higher, so sizing our
    profiles to it would be designing around a bug.
  * Their fix (ClientTrafficPolicy, bufferLimit: 8Mi) is written but NOT
    deployed, pending their operator's approval. We do not re-test until they
    confirm — a half-changed system gives a number neither side can trust.
  * In standalone `aigw run` a SecurityPolicy is accepted and then silently
    ignored, so "the config was accepted" proves nothing there. They will verify
    by re-reading the live config_dump and sending a large request. Same
    silent-default shape this repo keeps hitting, one layer down.

bridged.yaml carries the same correction (gitignored, so not in this commit).

Refs: gitea #76
2026-08-15 19:47:01 +02:00
Dai Ha 1cc34888fd CB-591: record the live result — blocked by a 32 KiB body limit at the gateway
CI / contract (push) Successful in 44s
CI / build (push) Successful in 55s
Deployed U1-U2c, restarted, spawned both new profiles for real, then reverted.

llm.ltms.dev answers HTTP 413 above 32 KiB (32768 bytes), on BOTH surfaces:

    /v1        32695 bytes -> 200        /anthropic  32095 bytes -> 200
    /v1        32795 bytes -> 413        /anthropic  32855 bytes -> 413

That is far below one agent turn. It is an edge limit (Caddy request_body
max_size, and/or Envoy), so the fix is in systems/vms, not here.

The part worth recording is how it nearly passed. Two members, same message,
same moment: `local` finished in 66s, `gx` never finished at all. `local`
passed only because the probe was three trivial questions in a fresh session,
so the request fit under 32 KiB — the profile looked healthy and was a
landmine set to fire on the first turn that reads a file. So §7's checklist
was not wrong, it was too easy; it now says to use a file-reading task.

opencode's failure mode is worse than a crash: it catches the 413, compacts
its context, retries, and loops. Observed 10+ minutes BUSY with no reply. From
the lead's side that is indistinguishable from a slow worker. Reproduced
outside the bridge with the launcher's own generated config, which is how it
became a one-line error instead of a hang; §7.1 records that procedure.

Everything else about the migration checked out and is recorded so it is not
re-tested: token accepted on both surfaces, unauthenticated 401 (the Caddy
proxy does gate, whatever the gateway's own fail-open policy does),
/v1/models exactly ["deepseek-v4-flash"], the guard allowlist accepted
llm.ltms.dev, and the generated opencode provider block is correct with a real
llmk- key.

Also answers §3b's open question: reasoning survives BOTH surfaces —
/anthropic returns a real "type":"thinking" block and /v1 returns a populated
reasoning_content. The feared /v1 translation loss did not happen.

Config state (bridged.yaml is gitignored, so it is described rather than
committed): `local` back on http://gx00.gw:8000, `gx` kept at weight 0,
`local-direct` kept, llm.ltms.dev left in the guard allowlist. The file
carries these numbers and the exact two-key edit to switch back.

Verified after the revert with a task that reads two large files: correct on
all three questions. Daemon pid 66745, jar f1fd659423e6.

Refs: gitea #76
2026-08-15 19:27:45 +02:00
Dai Ha 0331ecd5d3 CB-592: add the BRIDGED_MEMBER marker — the sentinel alone cannot hold
CI / contract (push) Successful in 1m5s
CI / build (push) Successful in 1m39s
Live check on a member pane showed the CB-592 shadow did NOT take effect:
GITEA_ACCESS_TOKEN inside the pane was still the real admin token.

Measured cause. The overlay itself works — GITEA_TOKEN is injected the same
way, is exported by no shell file, and does reach the pane. The sentinel loses
one step later. A herdr pane runs a LOGIN shell, ~/.zprofile line 41 sources
${SHARED_ENV}/tools/secrets.sh, and that file does a plain unconditional
`export GITEA_ACCESS_TOKEN=...`. A login shell overwrites a value already in
the environment, so the real token is put back before the member starts.
Confirmed directly:

    GITEA_ACCESS_TOKEN=cb592-sentinel zsh -lc ...
    -> RESULT: sentinel was OVERWRITTEN by the login shell

This defeats any launcher-side overlay for any name secrets.sh exports. No
change in this repo can win it alone.

So this adds the half that does survive: BRIDGED_MEMBER=1, a name secrets.sh
never exports. It is a no-op until the operator guards the export:

    [ -n "${BRIDGED_MEMBER:-}" ] || export GITEA_ACCESS_TOKEN=...

Setting it now costs nothing and makes that one line the whole remaining fix.
The sentinel stays: it is correct for any peer kind whose pane does not start
a login shell, and it keeps the intent explicit where every adapter passes.

Also corrects the javadoc and the test javadoc, which both claimed a
protection that was measured not to hold.

The other reported failure was my own bad test, not a regression. The probe
called /api/v1/user, which a minimal write:repository token cannot read. Same
token on the repo endpoint answers 200, so CB-302 is intact:

    GITEA_ACCESS_TOKEN: /user=200  /repos/lms/claude-bridge=200
    WORKER_GITEA_TOKEN: /user=403  /repos/lms/claude-bridge=200

Tests 805 -> 807. Both new tests proved to discriminate by reverting the
marker: everySpawnMarksThePaneAsAMember and
aProfileEnvEntryCannotClearTheMemberMarker both fail without it.

Refs: gitea #77
2026-08-15 18:37:52 +02:00
Dai Ha 831a918c30 Merge CB-592: shadow the admin GITEA_ACCESS_TOKEN in every member's environment
CI / contract (push) Successful in 41s
CI / build (push) Successful in 1m22s
A live probe showed every spawned member carried the admin GITEA_ACCESS_TOKEN: 108
environment variables in a member's pane against 99 in the primary's. The operator's
rule is that only the leader and architects may use it; everyone else uses
WORKER_GITEA_TOKEN. We were not enforcing that at all.

The cause is invisible from inside the launcher. baseEnv builds a fresh map holding only
PATH and the profile's env:, so a member looks like it gets a small explicit environment.
That map is an OVERLAY: WorkspaceControl.createTab/splitPane send only the keys it
contains, and herdr spawns the pane from its own login-shell environment, so every key we
never mention passes straight through — admin token included.

The fix puts a non-blank sentinel over the key in baseEnv, applied AFTER the profile's
env: so no profile, present or future, can restore the real token by naming it in config.
One place, every adapter, including peer kinds not yet written — deliberately not a
per-profile bridged.yaml entry, which is the silent-default shape this repo has shipped
nine times.

A non-blank sentinel rather than the empty string, on purpose: whether an empty overlay
value overrides an inherited variable or is skipped as blank cannot be settled from this
repo, because herdr's merge happens in an external process. baseEnv's own PATH seeding
(CB-511) already relies on a non-blank value replacing an inherited one, so this reuses
the shape that is demonstrated to work rather than the one that is merely plausible.

CB-302's repo-scoped GITEA_TOKEN grant is untouched — a worker can still open its own PR.
The subscription boundary was checked and is unaffected: the primary's pane carries no
ANTHROPIC_* at all, so nothing is inherited there.

Closes gitea #77. Live verification follows separately: the daemon must be redeployed
before this reaches any pane.
2026-08-15 18:27:05 +02:00
Dai Ha 3db5277ae8 CB-592: shadow the admin GITEA_ACCESS_TOKEN in every member's herdr overlay
CI / contract (pull_request) Successful in 1m12s
CI / build (pull_request) Successful in 1m37s
herdr spawns a pane from its own login-shell process env and layers our map on
top, so any key baseEnv never mentions passes straight through — including the
admin forge token. baseEnv now puts a non-blank sentinel for
GITEA_ACCESS_TOKEN, applied after the profile's own env: so no profile can
restore it. One place, every adapter, every profile including future ones.
CB-302's GITEA_TOKEN grant (applyGitToken) is untouched.
2026-08-15 18:25:16 +02:00
Dai Ha 6939e0cbbc CB-592: the tracked opencode.json must name the worker forge token, not the admin one
Operator's rule, 2026-08-15: only the leader and architects may use GITEA_ACCESS_TOKEN;
everyone else uses WORKER_GITEA_TOKEN.

opencode.json is TRACKED, so it ships in every worker worktree, and it mounted the gitea
MCP with {env:GITEA_ACCESS_TOKEN}. A live probe confirmed that variable actually resolves
inside a member: herdr spawns each pane from its own login-shell environment and layers
the launcher's map on top, so a member sees 108 variables rather than the small explicit
set baseEnv appears to build. That gave an opencode member admin forge TOOLS — enough to
merge its own PR, which both CLAUDE.md and the member contract forbid.

This is the narrow half of the fix: it removes the tooling. The admin token is still
present as a string in every member's environment, which is the real defect and is
tracked as CB-592 (gitea #77) — that fix belongs in the launcher, in one place, not
per-profile in bridged.yaml where a sixth profile would silently reopen it.

.mcp.json keeps GITEA_ACCESS_TOKEN and is correct to: it is skip-worktree, the primary's
own local copy, and the primary is the lead. That is the pattern this change follows —
the shared tracked file grants least privilege, and anything needing more overrides
locally.
2026-08-15 18:20:55 +02:00
Dai Ha f0095bf8b2 CB-591: plan the move onto the LLM/MCP gateway, and check AI_GATEWAY_TOKEN
CI / contract (push) Successful in 1m5s
CI / build (push) Successful in 1m38s
The gateway (llm.ltms.dev) replaced Bifrost on 2026-08-15 and serves an Anthropic
surface and an OpenAI surface, so both member kinds can point at it. The plan is in
docs/CB-591-Gateway-Migration.md; gitea #76 tracks the work.

The opencode half needs no code: OpenCodeLauncher already pins an OpenAI-compatible
endpoint (CB-508), so baseUrl + tokenEnv + provider/model is a config change. That
matters more than it looks — every opencode member today is sol or terra, and both sit
on one OpenAI account via credentialId: openai-shared, so an exhaustion on either locks
out both. A gateway-backed opencode profile is free and off that credential, which
retires a single point of failure rather than only adding capacity.

Also extends the redeploy script's --check to AI_GATEWAY_TOKEN. A profile's tokenEnv is
resolved from the DAEMON's own environment by HerdrPeerLauncher.resolveEnv, so a token
added to secrets.sh after the daemon started is simply absent: the launcher injects an
empty token and the gateway answers 401, long after the restart and with nothing tying
the two together. That is the same trap as WORKER_GITEA_TOKEN, and it gets the same
login-shell check that never prints the value.
2026-08-15 17:06:31 +02:00
Dai Ha 5206679efd Merge CB-588: nudge the lead when an async ticket goes terminal
An async delegation ticket (bridge_send wait:false) resolves on MessageService.reply's
rendezvous fast path, which returns before onReplyQueued. So CB-307's push loop only ever
heard about the durable-inbox case, and the mode CLAUDE.md tells leads to prefer never
nudged anyone. Closes gitea #72.

Adds a second, independent reminder schedule keyed by the lead terminal, so several
tickets finishing together coalesce into one nudge. The CB-307 path is untouched.

Three defects were found in review and fixed before merge:
 * a pendingTickets entry outlived the ticket it named. poll() returns null once
   pruneTerminalTickets drops a ticket, so ticketCollected was never reached and the
   entry leaked for the daemon's life, riding along on every later nudge and sending
   the lead after a ticket bridge_poll can no longer find.
 * a lost nudge: a ticket landing between decideTickets returning STOP and
   activeLeads.remove coalesced onto a schedule that was about to die. That is the
   exact failure this ticket exists to remove, reintroduced in a narrow window.
 * the success direction was unpinned in tests, and the comment listing the paths that
   complete the future was short by several.

The obvious fix for the second one was wrong: restarting on any pending ticket defeats
the reminder cap, because a never-collected ticket at cap is expected to still be there.
The fix diffs against a snapshot taken before the decision, so only a ticket that truly
arrived during the window restarts the schedule.

Verified on my own unpiped build: 802 tests, 0 failures, BUILD SUCCESS.
Two reviewers on the diff; the loop-gating finding they raised is split out as CB-590.
2026-08-15 17:06:12 +02:00
Dai Ha ac044e7573 CB-588 round 3: close the STOP-vs-onTicketTerminal race, pin nudge polarity, fix comment
CI / build (pull_request) Successful in 1m19s
CI / contract (pull_request) Successful in 1m31s
- ReplyPushLoop.stopOrRestartTicketLoop: after releasing a lead's active-schedule
  slot on STOP, restart only if a ticket landed that the pre-decision snapshot
  did not already account for. A naive "restart on any pending ticket" version
  was tried first and reverted: it defeated the reminder cap by restarting
  forever on a stale, never-collected ticket (broke
  successfulTicketNudgeIncrementsDelivered and ticketNudgesSendUpToCapThenStop).
  Diffing against a pendingBefore snapshot distinguishes a genuine race arrival
  from stale cap-exhausted backlog.
- Two new ReplyPushLoopTest cases exercise stopOrRestartTicketLoop directly
  (now package-private) rather than forcing the underlying thread race:
  aTicketStillPendingWhenTheLoopStopsIsNotStranded (the race must restart) and
  aStaleUncollectedTicketAtCapDoesNotRestartTheLoop (the cap must still hold).
  Both were verified to fail against deliberately-reverted versions of the fix
  before being restored to green.
- MessageServiceTest: pin the success-path nudge test's negative direction too
  (must not contain "FAILED"), not just the failure-path test.
- MessageService: complete the whenComplete comment's list of completion paths
  (TIMED_OUT/BUSY/BACKEND_EXHAUSTED via finishAsyncTask, completeExceptionally
  on throw, and answer() -> finishAsyncTask(turnId, result)).
2026-08-15 17:02:05 +02:00
Dai Ha 6ebad2a91f CB-588 follow-up: reclaim a pruned ticket's pendingTickets entry too
CI / contract (pull_request) Successful in 45s
CI / build (pull_request) Successful in 1m0s
tasks is the sole authority on whether a ticket exists, but
pruneTerminalTickets dropped entries from it without telling
ReplyPushLoop. ticketCollected(ticket) was only ever called from
MessageService.poll's terminal branch, which pruneTerminalTickets
short-circuits past once a ticket is gone (poll returns null at the
top). A ticket the lead never polled — or one the reminder cap already
gave up on — was pruned from tasks but never collected in
ReplyPushLoop.pendingTickets, so it rode along on every later nudge to
the same lead forever, naming a ticket bridge_poll could no longer
find, and the map itself never shrank.

pruneTerminalTickets now calls pushLoop.ticketCollected for every
ticket it actually removes (guarded on pushLoop != null), so a pending
nudge entry lives exactly as long as its ticket is pollable. Reused
tasks as the only removal trigger rather than adding a second live
query back into MessageService — no new source of truth.

Added an injectable clock (LongSupplier nowNanos, defaulting to
System::nanoTime) to MessageService, mirroring the SessionManager/
SessionReaper nowNanos seam, so a test can cross the 10-minute
TICKET_TTL_NANOS deterministically instead of sleeping for real.
Confirmed the new regression test fails against the prior
pruneTerminalTickets (a stale ticket rides along on a later coalesced
nudge) before restoring the fix.
2026-08-15 16:41:23 +02:00
Dai Ha 3d10ed385c CB-588: nudge the lead when an async ticket reaches a terminal phase
CI / contract (pull_request) Successful in 41s
CI / build (pull_request) Successful in 1m46s
An async bridge_send(wait:false) registers a rendezvous waiter, so its
reply always takes MessageService.reply's fast path and returns before
ReplyPushLoop.onReplyQueued is ever called — the exact mode the charter
tells leads to prefer never nudged.

Add ReplyPushLoop.onTicketTerminal(ticket, target, failed), a second
entry point reusing the loop's status gating, bounded/backoff reminders
and metrics, keyed by the nudge-receiving lead so several tickets
finishing together coalesce into one nudge naming the count. MessageService
wires it via task.future.whenComplete in sendAsync (covers reply,
completion fallback, wedge, and abandon() alike) and calls the new
ticketCollected(ticket) from poll() once a terminal view is handed back,
so an already-collected ticket is never nudged again. The CB-307 inbox
path (onReplyQueued/decide/NUDGE_FORMAT) is untouched.
2026-08-15 16:10:22 +02:00
Dai Ha 6d0c94dbdb Correct enforceMaxLoad's comment after CB-585
CI / contract (push) Successful in 42s
CI / build (push) Successful in 1m16s
The comment said non-positive means unlimited at load. That stopped being
true when CB-585 made an explicit maxLoad: 0 survive as a real cap of zero
and made a negative value refuse config load. The code below it was already
right — only the comment described the old normalisation. Flagged by the
CB-585 worker, which correctly stayed out of a file not on its list.
2026-08-15 16:02:35 +02:00
Dai Ha e01563a550 Merge cb584-session-resume-b62247-4
CI / build (push) Successful in 1m20s
CI / contract (push) Successful in 1m44s
2026-08-15 15:58:29 +02:00
Dai Ha 30e3225a3c Merge cb585-maxload-zero-4585c4-3 2026-08-15 15:58:29 +02:00
Dai Ha ef186a1516 Merge cb587-snapshot-index-flags-24e236-2 2026-08-15 15:58:29 +02:00
Dai Ha 5d5b3bdc76 CB-584: persist agentSessionId at acquire and expose it via bridge_spawn/bridge_list
CI / contract (pull_request) Successful in 1m5s
CI / build (pull_request) Successful in 2m18s
Wires up the two links that made session resume unreachable: MemberSession
now records agentSessionId from PeerHandle at acquire (both the plain and
worktree paths), and it survives onto the roster (rosterView, so both
bridge_list and GET /members show it). bridge_spawn accepts sessionName and
resumeSessionId, threading them into the SpawnRequest fields that already
existed but were never reachable from the MCP surface.

A resumeSessionId now requires an explicit profile (a resumed conversation
is tied to the specific backend that started it, so an unqualified spawn
routed by placement has no safe candidate to check) and is refused, naming
Capability.SESSION_RESUME, when that profile's adapter does not declare it
— PeerLauncher gains capabilitiesFor(profileName) so a mixed fleet is
checked per-adapter rather than against the fleet-wide capability union.

PeerHandle.agentSessionId() loses its default, the same fix CB-571 (c00a86b)
made for charterReceipt() one method above it — both existing adapters
already overrode it, so this only closes the landmine for a future one.

Out of scope, left for follow-up: carrying agentSessionId on a failed
ticket's ReleaseDetail (issue #65 criterion 5) — CB-578 stage C just
landed in that file and this ticket deliberately stayed out of it.
2026-08-15 15:48:06 +02:00
Dai Ha 2f8c98dac9 CB-585: maxLoad: 0 caps a profile at zero, negative refused at load
CI / contract (pull_request) Successful in 41s
CI / build (pull_request) Successful in 1m18s
2026-08-15 15:38:50 +02:00
Dai Ha 2db7189067 CB-587: seed the snapshot's temp index from the worktree's real index
CI / build (pull_request) Successful in 1m1s
CI / contract (pull_request) Successful in 1m24s
--skip-worktree is an index flag. GitWorktrees.snapshot staged into a fresh
empty temp index, which carried none of the real index's skip-worktree bits,
so add -A staged local on-disk content for files git status correctly hides
(e.g. .mcp.json). Copy the real index (resolved via git rev-parse --git-path
index, correct for linked worktrees) into the temp index before staging, so
add -A skips exactly what git status skips.
2026-08-15 15:36:06 +02:00
Dai Ha 94476ac109 CB-529: drainReplies javadoc said the ack is local — false for AMQP
CI / build (push) Successful in 55s
CI / contract (push) Successful in 1m5s
The doc claimed an in-flight failure re-surfaces the messages on a later
drain, because the ack is local. That was true when InMemoryReplyInbox was
the only inbox, and became false without anyone noticing when the AMQP
adapter landed: there the ack is a broker-side basicAck, so a crash while
writing the response loses the reply outright. Re-polling cannot recover it,
since the broker has already forgotten it.

Behaviour is unchanged and the window stays accepted — acking on the next
poll instead would double-deliver on every normal drain. The point is that
the comment sat exactly where the next implementer would read it and said
the opposite of what happens.

Closes #12.
2026-08-15 15:34:59 +02:00
Dai Ha 5fe02b7c98 Fix redeploy script reporting failure on a successful restart
CI / contract (push) Successful in 45s
CI / build (push) Successful in 59s
Found by running it. The script launched the daemon with a relative jar path
(cwd is bridged/) but detected it with an absolute one, so pgrep never matched.
The daemon restarted correctly and booted clean, and the script still failed
with 'no process appeared' — the worst shape of bug for a deploy tool, because
it invites a second restart on a daemon that is already healthy.

Detection now matches both path forms, and the launch uses the absolute path so
ps names which checkout is running.
2026-08-15 15:18:44 +02:00
Dai Ha a1052f4fd1 Add scripts/redeploy-bridged.sh so the lead can deploy in one command
CI / build (push) Successful in 57s
CI / contract (push) Successful in 1m23s
The lead already owned the redeploy, but the command classifier refuses a bare
kill on the daemon, so in practice every deploy still needed the operator to
approve a stop and a start by hand. A single script is the seam that fixes
that: the operator allow-lists one auditable command instead of two ad-hoc
ones.

It also stops the procedure from living only in a checklist people read after
things go wrong. It builds before it stops anything, so a failed build never
leaves the fleet down; waits for the old process to exit instead of assuming;
polls /healthz; and anchors its log checks to a line marker taken before the
restart, so old errors cannot be misread as new ones.

The check with no log line anywhere in the daemon is the reason --check exists:
bridged inherits WORKER_GITEA_TOKEN from the shell that starts it, and starting
from a non-login shell empties it. The daemon boots fine, healthz is green, and
the failure only appears later as workers that cannot open a PR. --check tests
whether the name resolves and never prints the value.
2026-08-15 15:14:21 +02:00
Dai Ha 8c9904a7c4 Record that the classifier still blocks the daemon kill
CI / contract (push) Successful in 45s
CI / build (push) Successful in 59s
A CLAUDE.md rule grants intent, not tool permission. Tried the redeploy right
after writing the section and the classifier refused `kill`, so the section
would have been misleading as written. States the honest split until a Bash
permission rule exists: the lead builds and verifies, the operator runs the
stop and start with `!`.
2026-08-15 13:57:18 +02:00
Dai Ha 23af5dfdff The lead may redeploy the daemon — write down the procedure and its traps
A merge is not a deployment: the running bridged holds the jar it started
with, so merged code does nothing until the daemon is rebuilt and restarted.
Calling that work shipped is a false report. This makes the redeploy the
lead's job rather than something handed back to the operator, and records the
five things that have gone wrong doing it here — chiefly that starting the
daemon from a non-login shell empties WORKER_GITEA_TOKEN, which nothing logs
and which only surfaces later as workers that cannot open a PR.

Goes in the project addendum, not the canonical block; the block is unchanged
and still byte-identical with the wiki template.
2026-08-15 13:54:38 +02:00
Dai Ha 3a10f6ad17 Note that stage C's snapshot makes the parityOverlay gitignore rule load-bearing
CI / build (push) Successful in 56s
CI / contract (push) Successful in 1m19s
CB-578 stage C commits a preserved dirty worktree to refs/wip/<branch> with
git add -A. The parity overlay copies the primary's environment files into
every worktree, so a non-gitignored overlay path now reaches a durable git
object instead of only sitting on disk. add -A respects .gitignore, which is
what stops it — so the rule documented for CB-581 is now what keeps a secret
out of a commit, not just out of a directory.
2026-08-15 13:09:01 +02:00
Dai Ha a3842c873d Merge cb578c-92885c-1 2026-08-15 13:07:50 +02:00
Dai Ha c29c3f063d Merge cb554-e8e224-3 2026-08-15 13:07:50 +02:00
Dai Ha 758d62a396 Merge cb583-b82d11-2 2026-08-15 13:07:50 +02:00
Dai Ha 6725642274 CB-578 stage C: snapshot a dirty worktree into refs/wip before it can be lost
CI / build (pull_request) Successful in 54s
CI / contract (pull_request) Successful in 1m21s
Adds Worktrees.snapshot(worktreePath, branch, message): stages into a
temporary GIT_INDEX_FILE (never the worker's real index/HEAD), writes
the tree, commit-trees it onto the worktree's current HEAD, and points
refs/wip/<branch> at the result. add -A (never -f) respects .gitignore.

SessionManager.release() now snapshots any dirty worktree before the
preserve-or-remove decision, regardless of release cause (COMPLETED or
SHUTDOWN) — preserving on disk alone is one `worktree remove --force`
away from gone. A failing snapshot never escalates: the worktree is
still preserved, the pane still stops, and the release listener is
still notified.

onRelease's listener now receives a ReleaseDetail (terminal, worktree
path, branch, snapshot ref) instead of a bare terminal id, so Bridged's
WORKER_FAILED wiring can put the same three facts into a failed
ticket's detail — a lead can re-dispatch onto the same tree instead of
starting from the base commit.
2026-08-15 12:56:12 +02:00
Dai Ha a1f4dc365a CB-554: weight <= 0 excludes a profile from automatic placement, not 1.0
CI / contract (pull_request) Successful in 44s
CI / build (pull_request) Successful in 52s
2026-08-15 12:54:02 +02:00
Dai Ha 165b62ee20 CB-583: make bridge_list's capacity view quarantine-aware
CI / contract (pull_request) Successful in 44s
CI / build (pull_request) Successful in 1m32s
A quarantined profile's capacity row now forces free:0 and names the
quarantine (credentialId, quarantinedForSeconds), reusing the same
QuarantineSource bridge_profiles already reads instead of a second
lookup. An ordinary fleet's capacity rows are unchanged (no new keys).
2026-08-15 12:51:36 +02:00
Dai Ha 72d3481de3 Merge CB-578 stage B: quarantine the exhausted credential, not the profile
CI / contract (push) Successful in 41s
CI / build (push) Successful in 1m32s
Closes issue #50. Verified by the lead: own unpiped build of the branch merged
onto main — 740 tests, BUILD SUCCESS, exit 0.

Stage A only classified a usage-limit refusal. Nothing acted on it, so the
fleet would spawn another member onto the same exhausted account and fail the
same way. Now a BACKEND_EXHAUSTED classification quarantines the CREDENTIAL
for a cooldown: an explicit spawn onto it is refused with the profile, the
credential and the seconds remaining, placement skips it under every policy,
and it lifts itself on an injected clock.

Keyed by credential, not profile name, because sol and terra are two models on
one OpenAI account. Quarantining only the profile that reported the refusal
would leave its sibling live, and the next spawn walks onto the same dead
account. A profile that sets no credentialId quarantines alone under its own
name, so an existing config behaves exactly as before.

The implementer also found a real bug outside the brief: ConfigRef's
sameLaunchSettings never compared exhaustedPattern, so a reload changing only
that key reported 'config reloaded' while the value is actually deferred —
the precise failure that file's own doc calls the worst outcome a reload can
produce. Fixed with a regression test.

Two things I changed at review. The example.yaml conflict with e09cac6 was
resolved toward the branch, which was a strict superset. And
BackendQuarantine.none() held a clock frozen at 0, so a quarantine() call on
it would have locked a credential out for the life of the daemon — two
production CompositePeerLauncher constructors default to none(). It is now a
real no-op with a test.

Left open deliberately: the implementer's bridge_ask about the credentialId
name timed out after 55s with no answer, because I was polling minutes apart.
It proceeded on its own judgment and chose well. The 55s ask window against an
async lead is a bridge problem, not a worker problem.
2026-08-15 10:33:59 +02:00
Dai Ha 29ccb747ad CB-578 stage B: make BackendQuarantine.none() actually inert
Added at merge review. none() held a clock frozen at 0 with a 1ns cooldown, so
a quarantine() call on it recorded a deadline that could never pass — the
credential would be locked out for the life of the daemon. Two production
CompositePeerLauncher constructors default to none(), so that failure would
have been silent and permanent.

The implementer documented the limitation honestly rather than hiding it, but
a stand-in named none() should not need the caveat. quarantine() is now a
no-op on that instance, with a test asserting it. An inert value must omit the
fact, never invent one.
2026-08-15 10:33:44 +02:00
Dai Ha 4749e27453 Merge remote-tracking branch 'origin/main' into HEAD
# Conflicts:
#	bridged/bridged.example.yaml
2026-08-15 10:31:41 +02:00
Dai Ha e501d39988 CB-578 stage B: quarantine the exhausted credential, not the profile
CI / build (pull_request) Successful in 55s
CI / contract (pull_request) Successful in 1m21s
A BACKEND_EXHAUSTED classification (stage A) now puts that profile's
credential into a BackendQuarantine for a configurable cooldown. A spawn
onto a quarantined profile is refused naming the credential and roughly
when it lifts; weighted/round-robin/fixed placement skip a quarantined
candidate; the quarantine lifts itself on the injected clock; and it is
visible on bridge_profiles.

Keyed by credential, not by profile name, via the new Profile.credentialId
(profiles sharing one credential quarantine together — e.g. two models on
one account) and effectiveCredentialId() (unset ⇒ quarantines alone,
today's behaviour unchanged). Fixed a related gap along the way: a reload
changing exhaustedPattern was silently reported "applied" even though it's
deferred — sameLaunchSettings() now catches it too.
2026-08-15 10:29:54 +02:00
Dai Ha 6b6cf25862 CB-581: neutralise the parityOverlay false-preserve risk, and record the rule
CI / contract (push) Successful in 1m5s
CI / build (push) Successful in 1m32s
Issue #57 criterion 5 asked for an explicit decision on this rather than
silence. It is closer to live than the issue assumed.

The DEFAULT parityOverlay is List.of(".env", ".envrc"), so it applies to every
profile, and neither path was gitignored. Since CB-576 a release preserves any
worktree that git status --porcelain calls dirty, and that deliberately counts
untracked files — the work lost in CB-576 was a file nobody had added. So one
.env at the repo root would make every COMPLETED release preserve its
worktree, and worktrees would accumulate with no error to notice.

Inert today only because neither file exists here. Both are now gitignored,
which they deserve on their own as environment files. The javadoc carries the
rule for the next overlay path: it must be gitignored, or tracked and
skip-worktree'd.
2026-08-15 10:17:26 +02:00
Dai Ha 180de840eb Merge CB-581: a throw in release() no longer orphans the pane or aborts reapIdle
CI / contract (push) Successful in 43s
CI / build (push) Successful in 52s
Verified by the lead: own build of the branch merged onto main — 717 tests,
BUILD SUCCESS, exit 0.

release() ran an unprotected sequence. hasUncommitted shells out to git status
and throws on a non-zero exit, which skipped both notifyReleased and
launcher.stop. The session was already out of the registry, so nothing retried
it: a live pane kept burning a fleet slot while absent from the roster, and any
send blocked on it was never resolved. reapIdle called release bare inside a
loop, so one such session aborted the whole pass and skipped every session
after it.

Now the dirty-check block is guarded and fails toward preserving the worktree —
'we could not tell' must not be treated as 'it is clean', because deleting on a
guess destroys work with no other copy. notifyReleased runs in a finally and
launcher.stop runs unconditionally, so the pane always stops. reapIdle catches
per session, matching the shape drainAll already used.

Checked while reviewing: notifyReleased cannot throw out of the finally — it
already guards each listener and only logs. The pane stop is genuinely
unconditional.
2026-08-15 10:14:41 +02:00
Dai Ha 8fd2d7e5e7 CB-581: fail-safe release() so a throw never orphans the pane or aborts reapIdle
CI / build (pull_request) Successful in 52s
CI / contract (pull_request) Successful in 1m1s
hasUncommitted shells out to git and can throw; release() now catches that
inside a try/finally so notifyReleased and launcher.stop always run, and
defaults to preserving the worktree on a throw (can't tell dirty vs clean,
so don't risk deleting unrecoverable work). reapIdle wraps each per-session
release in try/catch, matching drainAll, so one bad session no longer
skips the rest of the reaping pass.
2026-08-15 10:09:32 +02:00
Dai Ha 129dd4a838 M4: record what Unit 2 has actually landed, and correct criterion 15
CI / build (push) Successful in 58s
CI / contract (push) Successful in 1m16s
Unit 2 was written as one block, but CB-571/576/580 and unit 2a have since
built parts of it. A symbol survey of main source plus the merge history puts
it at 1 of 21 criteria done, with three partials.

Criterion 15 says a normal COMPLETED release removes the worktree. Since
CB-576 that is false on purpose — a dirty worktree is preserved, because
deleting it destroys work that cannot be recovered. The criterion is wrong,
not the code.
2026-08-15 10:08:38 +02:00
Dai Ha e09cac6f1f CB-578 stage A: document exhaustedPattern as a deferred reload key
CI / contract (push) Successful in 44s
CI / build (push) Successful in 1m31s
The patterns are compiled once in Bridged.main from the startup config
snapshot, so adding one to a profile does nothing until a restart. Neither
the per-key docs nor the reload-class table said so.
2026-08-15 10:06:41 +02:00
Dai Ha c00a86b32c Merge CB-571: charter receipt on the roster, with no silent null for OpenCode
CI / contract (push) Successful in 42s
CI / build (push) Successful in 1m19s
Verified by the lead: own build of this branch merged onto main — 710 tests,
BUILD SUCCESS, exit 0.

Supersedes PR #54. A reviewer found that OpenCodeLauncher.SessionAwareHandle
wrapped the base WorkerHandle but never overrode charterReceipt(), so it
inherited the interface default of null while the real receipt sat on its
delegate — sol and terra would have shown no charterSource/charterSha256 on
the roster while Claude Code members showed both.

The fix is the root one, not the one-line override: PeerHandle.charterReceipt()
is no longer a default, so the compiler forces every implementation to answer.
This repo had shipped that same class of defect — a defaulted dependency that
compiles, passes tests, and quietly turns a feature off — eight times before
this one.
2026-08-15 10:02:36 +02:00
Dai Ha d3ae0350a2 Merge remote-tracking branch 'origin/main' into fix-charter 2026-08-15 10:01:10 +02:00
Dai Ha 2c2196a1f1 Merge CB-578 stage A: classify a usage-limit refusal instead of a completed reply
CI / build (push) Successful in 55s
CI / contract (push) Successful in 1m24s
Verified by the lead: own build of the branch merged onto main — 704 tests,
BUILD SUCCESS, exit 0. No vendor wording in any Java source (grep clean); a
profile with no exhaustedPattern keeps today's completion-fallback path exactly.

Accepted the implementer's deviation from the brief. The brief asked for a
terminal HealthState BACKEND_EXHAUSTED. FleetHealth.decide() is a pure
classifier over HealthSnapshot, which carries only booleans, AgentStatus and
MemberSession.State — never pane text. FleetHealth's own javadoc already says
ERROR_ON_SCREEN 'is not decided yet because it needs a bounded pane detection
read'. A second declared-but-unproduced value would repeat a gap the class
already documents as a problem, so the signal was put where the evidence
actually lives: the completion scrape.
2026-08-15 10:00:34 +02:00
Dai Ha 35ade14630 CB-571: make PeerHandle.charterReceipt() abstract, fix OpenCode adapter's silent null
CI / build (pull_request) Successful in 50s
CI / contract (pull_request) Successful in 1m21s
SessionAwareHandle wrapped the base's WorkerHandle but never overrode
charterReceipt(), so it silently inherited the interface default (null)
while the real receipt sat on its delegate. sol/terra never got a
charterSource/charterSha256 roster row.

Deletes the default so every PeerHandle must answer explicitly; the
compiler now catches this class of gap instead of a roster field
quietly going missing.
2026-08-15 09:55:52 +02:00
Dai Ha 541df87272 CB-578 stage A: classify a usage-limit refusal instead of a completed reply
CI / build (pull_request) Failing after 59s
CI / contract (pull_request) Successful in 1m10s
A backend that refuses on a subscription usage limit leaves the pane healthy but
the turn ends with no bridge_reply; the completion fallback used to scrape and
hand that refusal back as if it were a real answer. CompletionResolver now
matches the scrape against a per-profile exhaustedPattern (config, never a
vendor string) and resolves the send as Rendezvous.Kind/Outcome.BACKEND_EXHAUSTED
with a reason carrying the matched line, kept distinct from GONE/WORKER_FAILED.
A profile with no pattern configured is unaffected. Coverage is logged at
startup via CompletionResolver.coverage(...), naming which profiles have a
pattern and which don't, following FleetHealthMonitor.coverage's pattern.
2026-08-15 09:55:06 +02:00
Dai Ha 337b6ccd6e Merge remote-tracking branch 'origin/main' into fix-charter
# Conflicts:
#	bridged/src/test/java/dev/ltms/bridged/session/SessionManagerTest.java
2026-08-15 09:52:39 +02:00
ltms 42f46dfe9a Merge CB-579: resolve a lead by its tab name, drop the terminal-id pin
CI / build (push) Successful in 58s
CI / contract (push) Successful in 1m4s
Verified by the lead: merged onto main (9088d2b) in a scratch worktree, mvn -f bridged/pom.xml
clean install unpiped — MVN_EXIT=0, Tests run: 696, Failures: 0, BUILD SUCCESS. main alone measures
692, so this adds 4 net tests. Merges cleanly; Bridged.java auto-merged against CB-580.

Reviewed by the lead reading the full production diff and the three test files. The member's own
report was lost to the idle reaper before collection, so there was no author write-up.

Closes the live bug: LeadTabScanner.scan() no longer merges the config pin over the scan result, and
the cache no longer seeds from it, so a lead disappears once its tab is gone. The ghost this fixes
had begun throwing agent_not_found from ReplyPushLoop.decide on every tick, against a dead terminal
that still owned two live members.

Beyond the brief, and correct: `tab` is required for every leader, not only non-creatable ones,
because LeadLauncher also uses it to label a tab it creates. terminal: is rejected by a raw-YAML
check rather than by record shape — the only way to beat @JsonIgnoreProperties(ignoreUnknown = true).
The silent-default trap is avoided: the back-compat constructor still takes `tab` positionally.

Behaviour change worth knowing: the scanner's initial cache is now empty instead of the config pins,
so a herdr failure on the very first scan yields no leads until a scan succeeds. That is unavoidable
once the pins are gone, it fails loudly rather than silently, and it is covered by
aFailedFirstScanReturnsEmptyRatherThanThrowing.

Operator action required: bridged.yaml must replace fleet.leaders.<name>.terminal with tab. Already
done for this deployment.
2026-08-15 09:29:00 +02:00
ltms 9088d2b2c5 Merge CB-580: fail a ticket when its member reaches a terminal health state
CI / contract (push) Successful in 41s
CI / build (push) Successful in 1m32s
Verified by the lead on 0af902e: mvn -f bridged/pom.xml clean install, unpiped, in a scratch
worktree — MVN_EXIT=0, Tests run: 689, Failures: 0, Errors: 0, BUILD SUCCESS.

Reviewed by the lead reading the diff. The member's own report was lost to the idle reaper before
it was collected, so there was no author write-up to review against.

All three defects that sank 3b2f395 are absent:
  * failTarget is required — one constructor, Objects.requireNonNull, no defaulting overload
    anywhere in the repo, and Bridged.java:365 updated to pass messages::abandon.
  * The failure fires inside reportTransition after its `if (previous == next) return;` guard, so an
    unchanged tick cannot reach it.
  * No AtomicReference; MessageService is passed directly, so there is no empty window.

abandon(String target, String reason) confirmed as CB-568's target-wide operation that resolves
waiters as a failure rather than letting them time out.

Known limitation, accepted and covered by its own test: states.put records the new state before the
bounded retries run, so if all three attempts throw, the tickets stay pending and no later tick
retries. It is logged at WARN, and the retries carry no backoff.
2026-08-15 08:57:03 +02:00
ltms 500bfa2c33 Merge CB-576: release preserves a dirty worktree instead of deleting it
CI / build (push) Successful in 51s
CI / contract (push) Successful in 1m1s
Verified by the lead on b525b0f: mvn -f bridged/pom.xml clean install, unpiped, in a scratch
worktree — MVN_EXIT=0, Tests run: 686, Failures: 0, Errors: 0, BUILD SUCCESS.

Review accepted the required-interface-method shape (no defaulting overload) and the decision to
count untracked files as dirty — the work lost in the incident was a file that was never added.

One blocking defect was found and fixed in b525b0f: hasUncommitted called git with no existence
check, so a missing worktree threw WorktreeException from inside release() after registry.remove()
but before notifyReleased() and launcher.stop(), orphaning the pane and stranding a blocked
bridge_send caller. It now mirrors remove()'s already-gone tolerance.

Waived on merge, tracked as follow-up: a SessionManager-level test that teardown completes when the
worktree is gone, and the stronger fix behind it — reapIdle calls release() with no try/catch while
drainAll wraps it, so any exception in that window aborts the whole reaping pass.
2026-08-15 08:54:33 +02:00
Dai Ha 9ca9c43dfa CB-571: retag charter-receipt references from the taken CB-575
CI / contract (pull_request) Successful in 43s
CI / build (pull_request) Successful in 54s
CB-575 already names the merged MCP-cancellation-filter change, so the
charter-receipt comments used the wrong number. Retag to CB-571, the number
this work was authored against.
2026-08-15 08:48:57 +02:00
Dai Ha b525b0f08f CB-576: hasUncommitted tolerates an already-gone worktree
CI / contract (pull_request) Successful in 1m8s
CI / build (pull_request) Successful in 1m37s
2026-08-15 08:47:12 +02:00
Dai Ha 1966c69994 CB-575: charter receipt on spawn, in the roster and in the logs
CI / contract (pull_request) Successful in 46s
CI / build (pull_request) Successful in 1m34s
Record a CharterReceipt (role, source, sha-256 digest, byte count) for every
launch, store it on the MemberSession, expose it in bridge_list and GET
/members, and log it at spawn as digest+role only. Redact the charter argv
argument in the legacy pane-placement spawn log so the charter text never
reaches the daemon log. The charter prose itself is never recorded.
2026-08-15 08:10:56 +02:00
Dai Ha 976eff8ad1 CB-579: resolve a lead by its tab name, drop the terminal-id pin
CI / build (pull_request) Successful in 52s
CI / contract (pull_request) Successful in 1m18s
Leader.terminal -> Leader.tab (exact tab label, case-insensitive match).
LeadTabScanner matches an exact tab->name map instead of stripping a
shared tabPrefix, and no longer merges configured leads into every
scan result -- a stale pin can no longer outlive its tab.
LeadLauncher.tabLabel() returns the configured tab directly; the
terminalId pinned-terminal fallback in liveLeads() is gone.
Config load now rejects a leftover fleet.leaders.*.terminal key
instead of silently ignoring it. primary.terminal is untouched.
2026-08-15 08:05:21 +02:00
Dai Ha 0af902ec43 CB-580: fail a ticket when its member reaches a terminal health state
CI / build (pull_request) Successful in 54s
CI / contract (pull_request) Successful in 1m5s
FleetHealthMonitor now requires a failTarget BiConsumer<String,String>
collaborator (no defaulting overload) and calls it exactly once when a
member transitions into GONE or NEVER_READY, via CB-568's idempotent
target-wide abandon() operation. The reason string names the real
terminal state. failTarget invocation retries up to
MAX_FAIL_TARGET_ATTEMPTS (3) within the same transition if it throws,
and never refires on a later tick where the state is unchanged.

Bridged.java wires messages::abandon as the production failTarget.
2026-08-15 07:54:28 +02:00
Dai Ha 9118ce2537 CB-576: release preserves a dirty worktree instead of deleting it
CI / contract (pull_request) Successful in 1m2s
CI / build (pull_request) Successful in 1m35s
2026-08-15 07:34:06 +02:00
Dai Ha 2f48e08f1f Merge CB-577 follow-up: drop the target-keyed async index
CI / contract (push) Successful in 43s
CI / build (push) Successful in 55s
asyncTasksByWaiter correlates an async question by the exact rendezvous
waiter, so the target-keyed set it replaced can no longer decide
anything. Keeping it meant two indexes of the same fact, one of them
ambiguous whenever a target has two accepted tickets.

The race the old test modelled by reflection is gone with it: identity
keys make 'some other task reached this target' unrepresentable, so
there is no longer a wrong task for the question to land on.

Also carries the criterion-1 doc correction, which is identical to
aac29d6 on a different parent.
2026-08-15 06:40:31 +02:00
Dai Ha fec284e7cb Merge M4 unit 2a: accepted-turn identity carried to delivery
CI / build (push) Successful in 59s
CI / contract (push) Successful in 1m17s
TurnToken, owned by MessageService, binds a target to the exact
rendezvous waiter for one accepted send. Injector.Pending carries it and
the delivery callback hands it to CompletionResolver, so the baseline is
bound to the send it belongs to by construction rather than by a lookup
that could pick a different one.

The callback signature is required, not a defaulted overload: a delivery
with no token is exactly the unbound baseline this unit forbids, so a
default would let a caller silently produce it.

The token deliberately omits the session turn number. MessageService
owns acceptance but never learns of delivery, and
CompletionResolver.onDelivered runs before SessionManager.onDelivered,
so the number does not exist yet at the only point the token could
capture it. docs/M4-Fleet-Health.md criterion 1 records this and the two
rejected alternatives.

Still open for the next slice: the missing/post-restart baseline test and
the no-replay test.
2026-08-15 06:36:46 +02:00
Dai Ha 8e2e4c5e73 M4 unit 2a: migrate test call sites to the required turn token
The delivery callback now requires a TurnToken, so 54 test call sites
had to pass one. They use an explicit TestTurnTokens.inert(target)
rather than a defaulted overload, because a delivery with no token is
the unbound baseline this unit forbids.

The first version of inert() returned a fresh CompletableFuture as the
waiter, which turned captureBaselineSkipsTheReadWhenNoSendIsWaiting red:
the resolver saw a non-null waiter, concluded a turn was in flight, and
scraped a pane no send was blocked on. An inert value must omit the
fact, not invent it, so the waiter is now null and the production skip
fires as designed.
2026-08-15 06:36:35 +02:00
Dai Ha 0edc6615fc M4: correct unit 2 criterion 1 — TurnToken cannot carry the session turn
The criterion required the token to bind the session turn number. Three
independent refusals from the implementer showed why that is not
implementable at this layer: MessageService owns acceptance but never
learns of delivery, and CompletionResolver.onDelivered runs before
SessionManager.onDelivered, so the turn number does not exist yet at the
only point the token could capture it.

Records both rejected alternatives and why, so the next reader does not
re-derive them: a target-keyed registry restores the ambiguity the token
exists to remove, and injecting a turn counter couples layers to fill a
field nothing reads yet.
2026-08-15 06:32:02 +02:00
Dai Ha b745e159de CB-573: carry accepted turn tokens on delivery 2026-08-15 06:30:41 +02:00
Dai Ha aac29d604c M4: correct unit 2 criterion 1 — TurnToken cannot carry the session turn
CI / contract (push) Successful in 1m0s
CI / build (push) Successful in 1m32s
The criterion required the token to bind the session turn number. Three
independent refusals from the implementer showed why that is not
implementable at this layer: MessageService owns acceptance but never
learns of delivery, and CompletionResolver.onDelivered runs before
SessionManager.onDelivered, so the turn number does not exist yet at the
only point the token could capture it.

Records both rejected alternatives and why, so the next reader does not
re-derive them: a target-keyed registry restores the ambiguity the token
exists to remove, and injecting a turn counter couples layers to fill a
field nothing reads yet.
2026-08-15 06:29:27 +02:00
Dai Ha c884802b13 CB-577: remove obsolete async target tracking
CI / contract (pull_request) Successful in 43s
CI / build (pull_request) Successful in 56s
2026-08-15 06:28:53 +02:00
Dai Ha 5f5573a24e Merge CB-577: correlate an async question by its exact waiter
CI / contract (push) Failing after 0s
CI / build (push) Successful in 1m13s
markAsyncQuestion picked the first not-done task out of an unordered
set, so between resolveQuestion waking the first async send and the
question being recorded, a queued second send could join the set and
take the question. A lead answering with bridge_send{turnId} would then
resume a turn it did not mean to.

Each async task is now indexed by its exact rendezvous waiter, which has
identity semantics, so no other task can hold the same key. The question
is recorded before resolveQuestion, with a rollback when no waiter is
there, which closes the window rather than narrowing it.

An unanswered async question stays PENDING — the worker resumes after
its ask times out, so the delegation is not failed — and its stale
target tracking is now cleared instead of leaking.
2026-08-15 06:27:16 +02:00
Dai Ha 74b0087ebb CB-577: model async question ownership race
CI / build (pull_request) Successful in 1m29s
CI / contract (pull_request) Failing after 0s
2026-08-15 06:24:05 +02:00
Dai Ha 927e0151d4 CB-577: test async question waiter ownership 2026-08-15 06:23:22 +02:00
Dai Ha 5275922d1d CB-577: handle questions without async waiters 2026-08-15 06:22:22 +02:00
Dai Ha e186c7945a CB-577: correlate async questions to turns 2026-08-15 06:21:53 +02:00
Dai Ha 33a6e77f0e Merge CB-573: dormant fleet health monitor (M4 unit 1)
CI / build (push) Successful in 49s
CI / contract (push) Successful in 1m14s
An opt-in whole-fleet observer, separate from the 250ms delivery poller.
One AgentControl.list and one roster snapshot per tick, joined and fed to
the FleetHealth classifier, because a fault is a disagreement between the
two views at the same instant. Absent a health: block nothing is built
and no herdr call is made.

Adds bridge_list healthCoverage: off, detection-only, or full. Detection
is deliberately separate from notification, so a single-lead setup with
no webhook still gets detection and is told its coverage is partial
rather than being refused.

Two review fixes worth naming. tick() rescheduled itself as its last
statement with no try/catch, and a ScheduledExecutorService does not
re-run a task that threw — so the first agents.list failure would have
stopped health permanently and silently, which is exactly when the
control link is down. It now catches Throwable and reschedules in a
finally. And the snapshot fields this unit cannot supply are the named
constant NOT_YET_OBSERVED rather than bare false literals, because false
means no fault to this classifier.
2026-08-15 06:19:15 +02:00
Dai Ha 4e47489d53 Merge CB-568c: fail every pending async ticket on teardown
A released target left its second async ticket pending for the full
30-minute async timeout. abandon resolved only the rendezvous waiter,
and async tickets live in a separate map that could not even represent
two tasks on one target.

abandon now sweeps every non-question async ticket for the target, and
asyncTasksByTarget holds a set. The sweep is a plain loop: the first
attempt used Stream.anyMatch, which short-circuits on the first true, so
it completed one ticket and left the rest pending — the exact bug it was
fixing. Its test passed only because the rendezvous path failed the
first ticket anyway; the test now uses three tickets so a single
completion cannot satisfy it.

An ASKING ticket is an active turn, not a pending send, so the sweep
skips it and CB-574 is unaffected.
2026-08-15 06:18:28 +02:00
Dai Ha c76b2f149e CB-573: keep health monitoring after failures
CI / contract (pull_request) Successful in 42s
CI / build (pull_request) Successful in 1m15s
2026-08-15 06:17:28 +02:00
Dai Ha 826fffe05b CB-573: add dormant fleet health monitor 2026-08-15 06:15:24 +02:00
Dai Ha 75f57cdba7 CB-568c: fail every queued async ticket
CI / contract (pull_request) Successful in 44s
CI / build (pull_request) Successful in 1m34s
2026-08-15 06:15:21 +02:00
Dai Ha f556af5d4e CB-568: fail queued async tickets on teardown 2026-08-15 06:14:28 +02:00
Dai Ha d5f33f0c6e Merge CB-568: a dropped send reports the real cause
CI / build (push) Successful in 55s
CI / contract (push) Successful in 1m0s
Injector.drop knew the precise cause (herdr agent_not_found) but the
sender was told only 'worker unreachable or stuck', so a lead could not
tell a dead pane from a stalled model.

TurnListener.onTurnFailed gains a reason, defaulting to the old one-arg
form. CompletionResolver prefers that reason, then the pane scrape, then
the old fixed text.

drop now fires onTurnFailed unconditionally. That is the substantive
fix: the sender blocks on the rendezvous waiter, not on the delivered
future, so failing delivered() alone never woke it and a queued send sat
until its timeout.
2026-08-15 06:08:41 +02:00
Dai Ha 16e17b32ad CB-568: preserve dropped turn causes
CI / contract (pull_request) Successful in 43s
CI / build (pull_request) Successful in 1m26s
2026-08-15 06:06:07 +02:00
Dai Ha 988494e18b Merge M4 fleet-health design (docs only)
The full 5-unit design behind M4: evidence model, classification
precedence, the automatic-vs-lead action boundary, worktree safety on
release, typed inbox and lead routing, capacity, and human escalation.

Two decisions worth keeping visible. Detection is split from
notification, so health works in a single-lead setup with no webhook and
reports partial coverage instead of refusing to run. And capacity stays
a view: the bridge reports free slots but never spawns, reassigns, or
stops a member to improve utilisation, because only the lead holds the
work list.

Section 13 records eleven things nobody checked, including live LavinMQ,
OpenCode pane fixtures, and multi-lead routing.
2026-08-15 06:05:26 +02:00
Dai Ha c4549a5e20 Merge CB-575: filter the routine MCP cancellation WARN
The MCP SDK 2.0.0 registers no handler for notifications/cancelled, so every
client abort logged a WARN. M4 fleet health treats WARN as action-needed, so
that noise had a cost. A Logback TurboFilter denies only that one event:
right logger, WARN level, the SDK's exact format string, and a
JSONRPCNotification whose method is notifications/cancelled. Everything else
is NEUTRAL. If a later SDK handles cancellation the filter stops matching.
2026-08-15 06:00:40 +02:00
Dai Ha 46fa4f38d5 CB-575: filter MCP cancellation warnings
CI / build (pull_request) Successful in 49s
CI / contract (pull_request) Successful in 1m1s
2026-08-15 05:57:45 +02:00
68 changed files with 6815 additions and 547 deletions
+8
View File
@@ -6,6 +6,14 @@
# Settings backups inherit the env block — and secrets with it.
.claude/settings.local.json.bak*
# The default profile parityOverlay copies these primary→worktree, so they appear in EVERY worker
# worktree. Two reasons they must be ignored. They hold environment values, which is reason enough.
# And since CB-576 a release preserves any worktree that `git status --porcelain` calls dirty —
# untracked files included, deliberately. An untracked overlay file would therefore make every
# COMPLETED release preserve its worktree, and worktrees would pile up with no error to notice.
.env
.envrc
# Daemon runtime artefacts. bridged appends its log wherever it is launched from, so both the
# repo root and bridged/ collect one; neither belongs in git.
bridged.out
+68 -6
View File
@@ -81,9 +81,9 @@ below are the procedure — run them in order, every task, not only the big ones
5. **Collect** — `bridge_poll{ticket}` → `bridge_ack{ticket, msgId}`. Answer a worker's `bridge_ask`
with `bridge_send{turnId, content}` — **not** `sessionId`. A worker gone quiet is diagnosed with
`bridge_status`, never by reading its terminal.
6. **Verify yourself.** Re-run the build and the checks. A worker mounts only the bridge MCP and
cannot run your other tooling, and a piped command (`… | tail`) hides failures behind a zero
exit — never promote a worker's "clean" to a fact.
6. **Verify yourself.** Re-run the build and the checks. A worker cannot run your IDE tooling, any
forge tools it appears to have hold a blocked credential and fail, and a piped command
(`… | tail`) hides failures behind a zero exit — never promote a worker's "clean" to a fact.
7. **Review — fan out.** Spawn reviewers against the diff, one per dimension or per file, with
`wait:false`. Never the implementer of the scope it reviews, and brief them from the diff — not
from the implementer's rationale, which carries its own blind spot. Dispatch each PR's reviewers
@@ -155,9 +155,13 @@ you.
without replying, the bridge scrapes your pane, and it can return only the last 4000 characters.
A clipped scrape is marked as partial, but the missing text is gone — your report reaches the
lead with its end cut off.
5. **Report honestly.** State only what you actually ran and its real output, including failures.
You mount **only** the bridge MCP — the primary's other servers (IDE, forge, docs) are not yours,
so never claim the result of a check you had no way to run.
5. **Report honestly.** State only what you actually ran and its real output, including failures,
and never claim the result of a check you had no way to run. **Measure your own tools; do not
assume them.** What you mount depends on your backend: an opencode member gets the bridge and
nothing else, while a Claude Code member also inherits the operator's user-scope MCP servers,
which the bridge never chose for you. Two rules follow. The primary's IDE tooling is still not
yours, whatever you see. And **a mounted tool is not a working tool** — the forge server you may
find there holds a deliberately blocked credential and fails every call, by design.
6. **Never merge.** Stage files explicitly — never `git add -A` — and leave alone anything the
project marks as not-yours-to-commit.
@@ -189,6 +193,64 @@ charter, not here.
are diagrammed in `docs/MCP-Contract.md` §6, kept out of this file because it loads into every
session's context.
### Redeploying the daemon — the lead may do this (primary only)
**A merge is not a deployment.** The running `bridged` holds the jar it was started with, so a
feature merged to `main` does nothing until the daemon is rebuilt and restarted. Saying "shipped"
about code the live daemon has never loaded is a false report. The lead **may and should** redeploy
rather than hand the job back to the operator.
Workers must never do this. A worker has no business restarting the daemon it is talking through,
and stopping it kills the worker's own channel mid-turn.
**Use the script — do not hand-roll the steps.**
```bash
scripts/redeploy-bridged.sh --check # report state, change nothing
scripts/redeploy-bridged.sh # build, confirm drain, restart, verify
scripts/redeploy-bridged.sh --yes # skip the drain prompt (fleet already checked)
```
It builds before it stops anything, so a failed build never leaves the fleet down; it waits for the
old process to exit rather than assuming; it polls `/healthz`; and it anchors its log checks to a
line marker taken before the restart, so old errors cannot be misread as new ones. Run `--check`
first — it is read-only and reports whether the forge token resolves, which nothing else tells you.
The script encodes the five things below, each of which has gone wrong here before. Read them anyway:
if the script is unavailable or a step fails, this is what it was protecting you from.
1. **Login shell, or workers silently lose their forge token.** The daemon inherits
`WORKER_GITEA_TOKEN` from the shell that starts it, and that comes from
`${SHARED_ENV}/tools/secrets.sh`. Start it from a non-login shell and the variable is empty, the
daemon starts fine, and the failure appears much later as workers that cannot open a PR. Nothing
logs this at startup — the script's `--check` is the only thing that reports it, and it checks
whether the name resolves without ever printing the value.
2. **Drain live members first.** `bridge_list`, then `bridge_stop` each member, and collect anything
you still want with `bridge_poll` before you kill anything. A restart drops in-flight tickets and
rendezvous, and a member's report is not recoverable once its ticket is gone.
3. **A restart is the only way deferred config keys take effect.** That is usually the reason to do
it. The startup log names which keys it accepted and which it deferred — read those lines rather
than assuming.
4. **Re-check identity afterwards.** Call `bridge_whoami` and confirm it still answers `primary`. The
lead is found by its tab label (`fleet.leaders.*.tab`), and a lead whose tab no longer matches is
demoted to worker, which refuses every orchestration call.
5. **Prove the new jar is the one running.** Confirm a *fresh* `bridged listening` line at the end of
`bridged/bridged.out`, dated after the restart. An old daemon that never died looks identical from
the outside.
**Permission.** A `CLAUDE.md` rule grants intent, not tool permission — the command classifier
refuses a bare `kill` on the daemon whatever this file says. The script is the seam that fixes that:
it is one auditable command, so the operator allow-lists it once instead of approving a stop and a
start every time. The rule lives in the operator's Claude Code settings:
```json
{ "permissions": { "allow": ["Bash(scripts/redeploy-bridged.sh:*)"] } }
```
Granted by the operator on 2026-08-15. If a call is still refused, do **not** route around it by
running the stop and start as separate commands — that is exactly the approval the script replaced.
Say what you were going to run and why, and let the operator decide.
### The prompt is part of the product — update it with the code (mandatory)
This repo *is* the bridge, so the canonical block above is not documentation about someone else's
+79 -24
View File
@@ -47,15 +47,18 @@ bind:
# (say a Claude lead and an opencode lead) work as peers: the second is silently demoted and refused
# every orchestration call. List each lead's pane here and all of them resolve as leads.
#
# terminal → the ONLY field identity depends on; get it from that session's bridge_whoami
# tab → the ONLY field identity depends on (CB-579); the exact label of the tab hosting the lead.
# Label the tab yourself, or let bridged label one it launches — see `fleet.leaders:` below.
# kind/model → descriptive; they document what runs in the pane and are echoed by bridge_whoami
#
# A lead is never spawned — it pre-exists, which is exactly why it must be named rather than created.
# A lead's tab must already carry its label (or be launched by bridged, which labels it) — there is
# no terminal id to paste in and nothing to re-pin when the session restarts: the tab survives, so
# the same label resolves the same lead again on the next scan.
# `bridge_whoami` reports `{"role":"primary","leader":"<name>"}`; role stays "primary" because a lead
# IS a primary for authorization, so nothing that keys on the role breaks.
#
# KEEP `primary:` when adding leads: it still addresses the CB-307 push loop, which needs a single
# destination for its nudges. If both name the same terminal, the `fleet.leaders:` entry wins.
# destination for its nudges, and is a separate mechanism from lead identity — see `fleet.leaders:`.
#
# Leads are configured under `fleet.leaders:` — see THE FLEET further down.
#
@@ -85,6 +88,16 @@ bind:
# backoffMs: 60000
# quietNudgeCap: 3
# Fleet health detection is dormant unless enabled. It reads one whole-fleet agent list per tick.
# It can run without a webhook; bridge_list then reports healthCoverage: detection-only.
# health:
# enabled: true
# intervalSeconds: 30 # minimum 15
# workingSuspectAfterSeconds: 600 # minimum 300
# paneProbeIntervalSeconds: 60 # minimum 60
# notifications:
# mode: disabled # disabled (default) or webhook
# herdr Unix socket. Omit to use the client default
# (${HERDR_SOCKET_PATH:-~/.config/herdr/herdr.sock}).
herdrSocket: ~/.config/herdr/herdr.sock
@@ -125,6 +138,23 @@ herdrSocket: ~/.config/herdr/herdr.sock
# SSH is unaffected). The token value itself is never stored in this file.
# gitHostEnv → host env var holding the forge host (default GITEA_HOST). Injected as
# GITEA_HOST *only* alongside a resolved gitTokenEnv.
# exhaustedPattern → regex matched against a completion-fallback scrape (CB-578 stage A) to
# classify a turn that ended with no bridge_reply as the backend having
# refused on a subscription usage limit, rather than a real answer. Opt-in —
# omit and this profile's completion fallback behaves exactly as before.
# Every backend words its refusal differently, so this is config, never a
# vendor string baked into bridged itself.
# DEFERRED: compiled once into a startup pattern map — editing it needs a
# daemon restart, same as this profile's model/baseUrl/argv.
# credentialId → CB-578 stage B: the credential this profile quarantines WITH when a
# BACKEND_EXHAUSTED classification fires. Two profiles that set the SAME
# credentialId share one quarantine — the case this exists for is two models
# on one account (e.g. sol and terra both billing one OpenAI credential): an
# exhaustion on either one must lock out both, or the fleet just walks onto
# the same dead account under the sibling's name. Opt-in — omit and this
# profile quarantines alone, under its own name, exactly as if the field did
# not exist. Cooldown length is the top-level quarantineCooldownSeconds below.
# HOT: read live at every spawn/exhaustion check — no restart needed.
# env → extra environment for this profile's workers, as a literal key/value map
# (CB-511). Use it to give workers a toolchain.
#
@@ -154,10 +184,21 @@ profiles:
mcpUrl: http://127.0.0.1:8765/mcp
tokenEnv: BRIDGED_WORKER_TOKEN
argv: ["ccs", "gx10"]
weight: 0.5 # relative selection weight for placement: weighted
maxLoad: 2 # max live workers on this profile (omit for unlimited)
# weight: relative selection weight for automatic placement (weighted, round-robin, and
# fixed's fallback walk). Absent defaults to 1.0. An explicit 0 or negative value means
# "never auto-select this profile" (CB-554) — it stays reachable via an explicit
# `bridge_spawn{profile:"gx10"}`, which bypasses placement entirely; only automatic
# selection skips it.
weight: 0.5
# maxLoad: max live workers on this profile. Omit for unlimited. An explicit 0 (CB-585) caps
# the profile at zero live members — it is excluded from automatic placement and an explicit
# `bridge_spawn{profile:"gx10"}` against it is refused too; a cap holds even when the profile
# is named directly. Negative is refused at config load — there is no sane meaning for it.
maxLoad: 2
# gitTokenEnv: GITEA_TOKEN # opt-in: let this profile's workers open their own PR (CB-302)
# gitHostEnv: GITEA_HOST # defaults to GITEA_HOST; injected only with gitTokenEnv
# exhaustedPattern: "usage limit has been reached" # opt-in: classify a usage-limit refusal (CB-578)
# credentialId: shared-openai # opt-in: quarantine together with every other profile sharing this id (CB-578)
# configDir: /Users/me/.ccs/instances/gx10 # CLAUDE_CONFIG_DIR — inherit that profile's skills/MCP
# cwd: /Users/me/src/myrepo # pin the working dir; omit to inherit the primary's
# parityOverlay: [".claude/settings.local.json", ".env", ".envrc"] # never add .mcp.json — see above
@@ -223,6 +264,15 @@ profiles:
# round-robin, or weighted. Omitting this key is a strict no-op for existing configs.
placement: weighted
# How long a credential sits out after a BACKEND_EXHAUSTED classification (CB-578 stage B), in
# seconds, before a spawn may land on it again. Applies to every profile's effective credential
# (its own name, or its credentialId if set above) — there is no per-profile override. Default
# 1800 (30 minutes) when omitted or non-positive.
# DEFERRED: baked once into the BackendQuarantine built at startup — a running quarantine keeps
# its original cooldown regardless; a new value only applies to a quarantine that starts after a
# restart. Editing this needs a daemon restart to take effect.
# quarantineCooldownSeconds: 1800
# Re-read this file without restarting the daemon (CB-559). Off unless you add this block, so an
# upgraded bridged keeps the old behaviour: the file is read once at boot and never again.
# enabled → turn the watch on. bridged checks the file's modified time on a timer and
@@ -232,17 +282,20 @@ placement: weighted
# Not every key can move under a running daemon, and the difference is about what already exists
# when the reload happens — not about how important the key is:
# HOT → takes effect on the next spawn: the whole `fleet:` block (every role pool,
# `charters`, and `tabLabel`), `placement:`, and an existing profile's weight / maxLoad. Those are
# hot because the placement policy reads them through a supplier — being config is
# `charters`, and `tabLabel`), `placement:`, and an existing profile's weight / maxLoad
# / credentialId. Those are hot because the placement policy (and, for credentialId,
# the CB-578 stage B quarantine check) reads them through a supplier — being config is
# not by itself enough to make a key hot.
# DEFERRED → accepted into the new config, but the wiring built at startup keeps the old value
# until you restart: `lifecycle:`, `leadHeartbeat:`, `guard:`, `worktreeRoot:`,
# `spawnReadyTimeoutMs` / `spawnReadyPollMs`, ADDING or REMOVING a profile (a new
# backend needs its own launcher, and launchers are built once), AND an existing
# profile's launch settings — model, baseUrl, argv, env, configDir, mcpUrl, tabLabel.
# The launcher takes a copy of `profiles:` at startup and resolves every spawn out of
# that copy, so those never reach a launch until you restart. The reload logs them by
# name rather than pretending they applied.
# `spawnReadyTimeoutMs` / `spawnReadyPollMs`, `quarantineCooldownSeconds` (CB-578
# stage B — baked once into the quarantine tracker built at startup), ADDING or
# REMOVING a profile (a new backend needs its own launcher, and launchers are built
# once), AND an existing profile's launch settings — model, baseUrl, argv, env,
# configDir, mcpUrl, tabLabel, exhaustedPattern. The launcher takes a copy of
# `profiles:` at startup and resolves every spawn out of that copy, so those never
# reach a launch until you restart. The reload logs them by name rather than
# pretending they applied.
# COLD → cannot change at all: `bind:`, `herdrSocket:`, `broker:` and `auth:`. The socket is
# bound, the broker connection is open, and the auth mode decides who may reach the
# port that is already listening.
@@ -299,15 +352,18 @@ fleet:
# Panes that orchestrate rather than are orchestrated. A lead may now be CREATED as well as
# recognised: give it a `profile:` and the daemon launches the shortfall when fewer than
# `instances` are live. Give it only a `terminal:` and it is recognise-only, as before.
# `instances` are live. Omit `profile:` and it is recognise-only, as before.
#
# `tabPrefix` is the naming convention that finds a lead without pasting a terminal id: label the
# tab `lead: <name>` when you open it and the pane is recognised on the next rescan. Reopen the
# tab later and the id changes; the label does not.
# `tab:` (CB-579) is REQUIRED and is the only field identity depends on — the exact label of the
# tab hosting the lead, matched case-insensitively. Label the tab yourself and put that same
# string here, and the pane is recognised on the next rescan. Reopen the tab later, or the session
# inside it restarts — the terminal id changes; the tab, and its label, do not, so no config edit
# follows a restart.
#
# A lead the daemon launches is labelled BY the daemon, using the same convention, so it is found
# A lead the daemon launches is labelled BY the daemon with this same `tab:` value, so it is found
# by the same scan. A lead counts as live only when herdr also reports a running agent in that
# tab — a label left behind by a session that died does not block the relaunch.
# tab — a label left behind by a session that died does not block the relaunch, and a tab that is
# gone entirely drops out of the next scan rather than being remembered forever.
#
# An auto-launched lead is NOT a member: it gets no worker reply charter, is never registered with
# the session lifecycle (the idle reaper would kill your orchestrator), and stays on the
@@ -316,10 +372,9 @@ fleet:
# opus-5.0:
# profile: opus # omit to never create this lead, only recognise it
# instances: 1 # desired live count; only the shortfall is launched. 0 = off
# terminal: term_0123456789abcd # optional hand-pin; usually found by tabPrefix instead.
# # A running agent on this terminal also counts as live, so a
# # lead you opened by hand is not relaunched under you.
# tabPrefix: "lead:" # `lead: opus-5.0` ⇒ a lead named opus-5.0 (case-insensitive)
# tab: "lead: opus-5.0" # REQUIRED — the exact tab label this lead lives in
# tabPrefix: "lead:" # only used to guard against a worker tabLabel colliding with
# # this convention at startup; plays no part in matching a lead
# scanIntervalSeconds: 10 # rescan cadence, and the worst case before a new tab is seen
# workspace: leads # where a launched lead's tab is created (default "leads").
# # MUST NOT be a member workspace — those are excluded from the
@@ -327,7 +382,7 @@ fleet:
# cwd: /path/to/repo # the launched lead's working directory (default: bridged's own)
# kind: claude # descriptive; reported by bridge_whoami
# gpt-sol-5.6:
# terminal: term_fedcba9876543
# tab: "lead: gpt-sol-5.6"
# kind: opencode
# model: openai/gpt-5.6-terra
@@ -13,6 +13,8 @@ import dev.ltms.bridged.herdr.PaneLocator;
import dev.ltms.bridged.herdr.UnixSocketHerdrClient;
import dev.ltms.bridged.herdr.WorkspaceControl;
import dev.ltms.bridged.inject.CompletionResolver;
import dev.ltms.bridged.inject.ExhaustedPatternLookup;
import dev.ltms.bridged.inject.ExhaustionSink;
import dev.ltms.bridged.inject.Injector;
import dev.ltms.bridged.inject.StatusPoller;
import dev.ltms.bridged.inject.TurnListener;
@@ -23,6 +25,7 @@ import dev.ltms.bridged.mcp.BridgeMcp;
import dev.ltms.bridged.mcp.ConnectionIdentity;
import dev.ltms.bridged.metrics.BridgedMetrics;
import dev.ltms.bridged.metrics.Metrics;
import dev.ltms.bridged.health.FleetHealthMonitor;
import dev.ltms.bridged.mcp.PrimaryRegistry;
import dev.ltms.bridged.mcp.LsofPeerPidLookup;
import dev.ltms.bridged.mcp.LsofProcessCwdLookup;
@@ -35,6 +38,7 @@ import dev.ltms.bridged.msg.LeadHeartbeatLoop;
import dev.ltms.bridged.msg.ReplyPushLoop;
import dev.ltms.bridged.rest.BridgedApp;
import dev.ltms.bridged.session.GitWorktrees;
import dev.ltms.bridged.session.MemberSession;
import dev.ltms.bridged.session.SessionManager;
import dev.ltms.bridged.peer.PeerLauncher;
import dev.ltms.bridged.session.SessionReaper;
@@ -42,6 +46,7 @@ import dev.ltms.bridged.member.ClaudeCodeLauncher;
import dev.ltms.bridged.member.CompositePeerLauncher;
import dev.ltms.bridged.member.HerdrPeerLauncher;
import dev.ltms.bridged.member.OpenCodeLauncher;
import dev.ltms.bridged.placement.BackendQuarantine;
import io.javalin.Javalin;
import org.slf4j.Logger;
import org.slf4j.LoggerFactory;
@@ -59,6 +64,7 @@ import java.util.concurrent.atomic.AtomicReference;
import java.util.function.Function;
import java.util.function.Predicate;
import java.util.function.Supplier;
import java.util.regex.Pattern;
import java.util.stream.Collectors;
/**
@@ -138,11 +144,18 @@ public final class Bridged {
() -> config.get().fleet()));
}
AtomicReference<Function<String, Integer>> liveCountRef = new AtomicReference<>(_ -> 0);
// CB-578 stage B: one quarantine tracker for the whole daemon, shared between the launcher
// (checked at spawn) and the exhaustion sink wired in below (written on BACKEND_EXHAUSTED).
// The cooldown is deferred (see BridgedConfig#quarantineCooldownSeconds): it is read once
// here, at startup, and a config reload only changes it for a daemon restart.
BackendQuarantine quarantine = new BackendQuarantine(System::nanoTime,
TimeUnit.SECONDS.toNanos(cfg.quarantineCooldownSeconds()));
PeerLauncher workers = new CompositePeerLauncher(
adapters,
cfg.effectiveDefaultProfile(),
config,
profileName -> liveCountRef.get().apply(profileName));
profileName -> liveCountRef.get().apply(profileName),
quarantine);
// CB-504: under supervision (launchd/systemd) bridged can start before herdr's socket
// exists. The client itself is lazy — it connects per call — but the orphan reap below is
// the first thing that actually talks to herdr, so without this wait a boot-order race
@@ -192,34 +205,34 @@ public final class Bridged {
if (leadTerminals.size() > 1) {
log.info("leads: {} panes recognised {}", leadTerminals.size(), leadTerminals.values());
}
// CB-531: on top of the static registry, discover leads by the tab labels the operator
// writes. CB-557 moved the settings onto the lead they describe, so scanning is on whenever
// a `fleet.leaders:` entry exists — with no leads configured the supplier is a constant and
// never touches herdr, exactly as a missing `leadScan:` block used to behave.
// CB-531: on top of the legacy primary.terminal pin, discover leads by the tab labels the
// operator writes. CB-557 moved the settings onto the lead they describe, so scanning is on
// whenever a `fleet.leaders:` entry exists — with no leads configured the supplier is a
// constant and never touches herdr, exactly as a missing `leadScan:` block used to behave.
// CB-579: each lead now names its own exact `tab:` label, so one scanner discovers every
// configured lead regardless of how differently their tabs are labelled — the old
// single-shared-tabPrefix limitation (and its warning) is gone.
final Supplier<Map<String, String>> leads;
var leaders = cfg.fleet().leaders();
if (!leaders.isEmpty()) {
// One scanner, so one prefix and one interval. Distinct per-lead prefixes would need a
// scanner each; until a config actually wants that, take the first entry's settings and
// say so, rather than silently honouring one lead's prefix and dropping another's.
var scan = leaders.values().iterator().next();
Set<String> memberSpaces = cfg.profiles().values().stream()
.map(BridgedConfig.Profile::workspace)
.filter(Objects::nonNull)
.collect(Collectors.toSet());
leads = new LeadTabScanner(herdr, scan.tabPrefix(), memberSpaces, leadTerminals,
TimeUnit.SECONDS.toNanos(scan.scanIntervalSeconds()), System::nanoTime);
log.info("lead scan: tabs labelled '{}…' host a lead (rescan every {}s, member spaces {} "
+ "excluded)",
scan.tabPrefix(), scan.scanIntervalSeconds(), memberSpaces);
long distinctPrefixes = leaders.values().stream()
.map(BridgedConfig.Leader::tabPrefix).distinct().count();
if (distinctPrefixes > 1) {
log.warn("fleet.leaders declares {} different tabPrefix values; only '{}' is scanned "
+ "for. Give every lead the same tabPrefix, or leads under the others "
+ "will not be discovered.",
distinctPrefixes, scan.tabPrefix());
}
Map<String, String> tabToName = new LinkedHashMap<>();
leaders.forEach((name, leader) -> {
if (leader != null && leader.tab() != null && !leader.tab().isBlank()) {
tabToName.put(leader.tab(), name);
}
});
// One shared rescan cadence: still taken from the first entry, as before — it is an
// operational cadence, not identity, so there is no correctness reason to give every
// lead its own scanner.
int scanIntervalSeconds = leaders.values().iterator().next().scanIntervalSeconds();
leads = new LeadTabScanner(herdr, tabToName, memberSpaces,
TimeUnit.SECONDS.toNanos(scanIntervalSeconds), System::nanoTime);
log.info("lead scan: tabs {} host a lead (rescan every {}s, member spaces {} excluded)",
tabToName.keySet(), scanIntervalSeconds, memberSpaces);
} else {
leads = () -> leadTerminals;
}
@@ -252,7 +265,40 @@ public final class Bridged {
// The blocking message endpoint (CB-104) is the producer; the poller is inert until then.
// CB-106: a confirmed turn completion resolves a blocked send whose worker never replied.
Rendezvous rendezvous = new Rendezvous();
CompletionResolver completion = new CompletionResolver(agents, rendezvous);
// CB-578 stage A: classify a completion-fallback scrape that matches a profile's configured
// usage-limit refusal as BACKEND_EXHAUSTED rather than handing it back as a real answer.
// Compiled once at startup, keyed by profile name; a profile with no exhaustedPattern is
// simply absent here, so its workers keep today's completion-fallback behaviour unchanged.
Map<String, Pattern> exhaustedPatternsByProfile = new LinkedHashMap<>();
cfg.profiles().forEach((name, profile) -> {
if (profile.hasExhaustedPattern()) {
exhaustedPatternsByProfile.put(name, Pattern.compile(profile.exhaustedPattern()));
}
});
ExhaustedPatternLookup exhaustedPatterns = target -> sessions.roster().stream()
.filter(session -> target.equals(session.terminalId()))
.findFirst()
.map(session -> exhaustedPatternsByProfile.get(session.profile()))
.orElse(null);
log.info("backend-exhausted classification (CB-578 stage A): {}",
CompletionResolver.coverage(cfg.profiles().keySet(), exhaustedPatternsByProfile.keySet()));
// CB-578 stage B: on a classification that actually wins, quarantine the exhausted profile's
// CREDENTIAL — not the profile name — so a profile sharing that credential (e.g. two models
// on one OpenAI account) is refused too, not just the one that happened to report it. Reads
// the profile config live off `config`, so a credentialId edit is hot: no restart needed.
ExhaustionSink exhaustionSink = (target, reason) -> sessions.roster().stream()
.filter(session -> target.equals(session.terminalId()))
.findFirst()
.map(MemberSession::profile)
.map(profileName -> config.get().profiles().get(profileName))
.ifPresent(profile -> {
String credentialId = profile.effectiveCredentialId();
quarantine.quarantine(credentialId);
log.warn("credential '{}' quarantined for {}s (profile '{}' classified "
+ "BACKEND_EXHAUSTED): {}", credentialId,
cfg.quarantineCooldownSeconds(), profile.profile(), reason);
});
CompletionResolver completion = new CompletionResolver(agents, rendezvous, exhaustedPatterns, exhaustionSink);
// CB-113: deliver only to an available worker (its MCP is connected), never its boot window.
// CB-301: the manager's presence bridge records availability and drives SPAWNING → READY.
MemberPresence presence = sessions.asPresence();
@@ -275,9 +321,9 @@ public final class Bridged {
}
@Override
public void onDelivered(String target) {
completion.onDelivered(target);
sessions.onDelivered(target);
public void onDelivered(String target, dev.ltms.bridged.msg.TurnToken token) {
completion.onDelivered(target, token);
sessions.onDelivered(target, token);
}
@Override
@@ -285,6 +331,12 @@ public final class Bridged {
completion.onTurnFailed(target);
sessions.onTurnFailed(target);
}
@Override
public void onTurnFailed(String target, String reason) {
completion.onTurnFailed(target, reason);
sessions.onTurnFailed(target);
}
};
Injector injector = new Injector(agents, turnListener, deliverableTo(presence, leads),
presence::forget);
@@ -348,16 +400,45 @@ public final class Bridged {
MessageService messages = new MessageService(agents, injector, rendezvous, replyInbox,
pushLoop, metrics);
// Health is a slow whole-fleet observer. Keep it separate from the 250ms delivery poller.
final FleetHealthMonitor healthMonitor;
var healthScheduler = Executors.newSingleThreadScheduledExecutor(r ->
Thread.ofVirtual().name("bridge-health-").unstarted(r));
if (cfg.health() != null && cfg.health().isEnabled()) {
// CB-580: a member found GONE/NEVER_READY must fail whatever ticket is waiting on it,
// through the same idempotent target-wide operation CB-516 already uses on release.
healthMonitor = new FleetHealthMonitor(agents, sessions::roster, messages, healthScheduler,
System::nanoTime, cfg.health().intervalOrDefault(), messages::abandon);
String coverage = FleetHealthMonitor.coverage(true,
cfg.health().notifications() != null && cfg.health().notifications().configured());
if ("detection-only".equals(coverage)) {
log.warn("fleet health: {} (no notification sink configured)", coverage);
} else {
log.info("fleet health: {}", coverage);
}
healthMonitor.start();
} else {
healthMonitor = null;
healthScheduler.shutdownNow();
}
// CB-520: the reply inbox only consumes for agents this gateway owns. own on acquire,
// release on teardown. Do this before CB-516 so the inbox is owned before any reply can land.
sessions.onAcquire(replyInbox::own);
// CB-516: releasing a worker must fail whatever send was waiting on it. Without this a
// torn-down delegation kept reporting PENDING until the 30-minute async timeout, and never
// reached /metrics — the delegation was unresolvable and nothing said so.
sessions.onRelease(terminal -> {
messages.abandon(terminal, "the worker session was released before it replied");
replyInbox.release(terminal);
primaryRegistry.forgetDelegation(terminal); // CB-532: don't leak the lead binding
sessions.onRelease(detail -> {
// CB-578 stage C, acceptance criterion 10: a failed ticket's detail should tell a lead
// where to re-dispatch onto the same tree, not just that the worker vanished.
String reason = "the worker session was released before it replied";
if (detail.worktreePath() != null) {
reason += "; worktree=" + detail.worktreePath() + " branch=" + detail.branch()
+ " snapshot=" + (detail.snapshotRef() != null ? detail.snapshotRef() : "none");
}
messages.abandon(detail.terminalId(), reason);
replyInbox.release(detail.terminalId());
primaryRegistry.forgetDelegation(detail.terminalId()); // CB-532: don't leak the lead binding
});
// MCP server face (CB-105): bridge_send/bridge_reply/bridge_status, mounted at /mcp.
@@ -387,7 +468,16 @@ public final class Bridged {
profile -> {
var configured = config.get().profiles().get(profile);
return configured == null ? null : configured.maxLoad();
}, () -> config.get().profiles().keySet(), System::nanoTime));
}, () -> config.get().profiles().keySet(), System::nanoTime),
new BridgeMcp.HealthCoverageSource(() -> {
var health = config.get().health();
return FleetHealthMonitor.coverage(health != null && health.isEnabled(),
health != null && health.notifications() != null && health.notifications().configured());
}),
new BridgeMcp.QuarantineSource(profile -> {
var configured = config.get().profiles().get(profile);
return configured == null ? null : configured.effectiveCredentialId();
}, quarantine));
// CB-559: opt-in config reload. With no `configReload:` block nothing is constructed, so an
// upgraded daemon behaves exactly as before — the file is read once at boot and never again.
@@ -408,6 +498,7 @@ public final class Bridged {
messages.close();
pushLoop.close();
if (heartbeat != null) heartbeat.close(); // CB-551: stop the idle-lead heartbeat scheduler
if (healthMonitor != null) healthMonitor.stop();
if (configWatcher != null) configWatcher.stop(); // CB-559: stop polling the config file
mcp.close();
if (reaper != null) reaper.stop();
@@ -58,6 +58,12 @@ import java.util.Set;
* {@code fixed} (default), {@code round-robin}, or {@code weighted}
* @param auth API authentication mode ({@code null} → {@code loopback-trust}, the
* historical behaviour), CB-501
* @param quarantineCooldownSeconds how long a credential stays quarantined after a
* {@code BACKEND_EXHAUSTED} classification (CB-578 stage B); {@code null}/{@code
* <=0} → {@link #DEFAULT_QUARANTINE_COOLDOWN_SECONDS}. Baked once into the
* {@code BackendQuarantine} built at startup, so it is DEFERRED: changing it
* needs a restart, and a quarantine already running keeps whatever cooldown was
* live when it started.
*/
@JsonIgnoreProperties(ignoreUnknown = true)
public record BridgedConfig(
@@ -73,9 +79,14 @@ public record BridgedConfig(
Primary primary,
Fleet fleet,
LeadHeartbeat leadHeartbeat,
Health health,
String placement,
Auth auth,
ConfigReload configReload) {
ConfigReload configReload,
Integer quarantineCooldownSeconds) {
/** Default cooldown (CB-578 stage B) when {@code quarantineCooldownSeconds} is absent/non-positive. */
public static final int DEFAULT_QUARANTINE_COOLDOWN_SECONDS = 1800;
/** Back-compat 14-arg form — no {@code configReload:} block, so file watching stays off. */
public BridgedConfig(Bind bind, String herdrSocket, Map<String, Profile> profiles, Guard guard,
@@ -83,7 +94,27 @@ public record BridgedConfig(
Integer spawnReadyPollMs, Broker broker, Primary primary, Fleet fleet,
LeadHeartbeat leadHeartbeat, String placement, Auth auth) {
this(bind, herdrSocket, profiles, guard, worktreeRoot, lifecycle, spawnReadyTimeoutMs,
spawnReadyPollMs, broker, primary, fleet, leadHeartbeat, placement, auth, null);
spawnReadyPollMs, broker, primary, fleet, leadHeartbeat, null, placement, auth, null, null);
}
/** Back-compat form before the optional {@code health:} block was added. */
public BridgedConfig(Bind bind, String herdrSocket, Map<String, Profile> profiles, Guard guard,
String worktreeRoot, Lifecycle lifecycle, Integer spawnReadyTimeoutMs,
Integer spawnReadyPollMs, Broker broker, Primary primary, Fleet fleet,
LeadHeartbeat leadHeartbeat, String placement, Auth auth, ConfigReload configReload) {
this(bind, herdrSocket, profiles, guard, worktreeRoot, lifecycle, spawnReadyTimeoutMs,
spawnReadyPollMs, broker, primary, fleet, leadHeartbeat, null, placement, auth, configReload, null);
}
/** Back-compat form before the CB-578 stage B {@code quarantineCooldownSeconds} field was added. */
public BridgedConfig(Bind bind, String herdrSocket, Map<String, Profile> profiles, Guard guard,
String worktreeRoot, Lifecycle lifecycle, Integer spawnReadyTimeoutMs,
Integer spawnReadyPollMs, Broker broker, Primary primary, Fleet fleet,
LeadHeartbeat leadHeartbeat, Health health, String placement, Auth auth,
ConfigReload configReload) {
this(bind, herdrSocket, profiles, guard, worktreeRoot, lifecycle, spawnReadyTimeoutMs,
spawnReadyPollMs, broker, primary, fleet, leadHeartbeat, health, placement, auth,
configReload, null);
}
/**
@@ -141,19 +172,58 @@ public record BridgedConfig(
* @param cwd fixed working directory for this profile's workers (CB-112 "told otherwise");
* {@code null}/blank → inherit the primary's cwd, else the daemon's
* @param parityOverlay repo-relative paths copied primary→worktree for config parity; null/empty
* defaults to a sensible set of local config files
* defaults to a sensible set of local config files.
* <p><strong>Every overlay path must be gitignored or tracked-and-skipped.</strong>
* CB-576 made {@code release()} preserve a worktree that {@code git status
* --porcelain} reports as dirty, and it deliberately counts untracked files —
* the work lost in CB-576 was a new file nobody had added. So an overlay path
* that is neither gitignored nor tracked lands in every worktree as an
* untracked file, makes every {@code COMPLETED} release preserve, and
* worktrees then accumulate with no error anywhere.
* <p>Checked on 2026-08-15 (CB-581): inert as configured. Tracked overlay
* files carry {@code --skip-worktree} so {@code --porcelain} cannot see them,
* {@code bridged.yaml} is gitignored, and the default pair {@code .env} /
* {@code .envrc} does not exist in this repo. Note the default applies to
* <em>every</em> profile, so creating either file at the repo root is enough
* to make it live. Add a new overlay path to {@code .gitignore} in the same
* change that adds it here.
* <p>CB-578 stage C raised the stakes: a preserved dirty worktree is now also
* committed to {@code refs/wip/<branch>} via {@code git add -A}. The overlay
* carries the primary's own environment files into the worktree, so a
* non-gitignored overlay path no longer merely sits there as an untracked
* file — it gets committed into a git object that survives the worktree's
* removal. {@code add -A} respects {@code .gitignore}, which is exactly what
* keeps that from happening, so the gitignore rule above is now what stops a
* secret from reaching a durable commit.
* @param gitTokenEnv name of the host env var holding the git-forge API token; when set, its
* value is injected as {@code GITEA_TOKEN} so the worker can open its own PR
* at checkpoint (CB-302). {@code null}/blank ⇒ no token is injected
* (minimal-grant default — push over SSH stays free, PR-create is opt-in)
* @param gitHostEnv name of the host env var holding the forge host (default {@code GITEA_HOST});
* injected as {@code GITEA_HOST} <em>only</em> when {@code gitTokenEnv} is set
* @param weight relative selection weight for {@code placement: weighted}. Absent or
* non-positive ⇒ 1.0. Weights are normalised by the policy, so they need
* not sum to 1.0.
* @param maxLoad max live workers allowed on this profile at one time; absent or
* non-positive ⇒ unlimited. Live means any session the registry still owns
* (acquired and not yet released), in any state.
* @param weight relative selection weight for automatic placement ({@code weighted},
* {@code round-robin}, and {@code fixed}'s fallback walk). Absent ⇒ 1.0.
* An explicit {@code 0} or negative value (CB-554) means "never
* auto-select this profile": it is excluded from every automatic policy's
* pool the same way a quarantined candidate is (see
* {@code PlacementPolicyUtil}). This does not make the profile
* unreachable — an explicit {@code bridge_spawn{profile:"..."}} bypasses
* placement entirely and still resolves it. Weights among the remaining
* (non-excluded) candidates need not sum to 1.0; only their ratios matter.
* @param maxLoad max live workers allowed on this profile at one time. Absent
* (unset/{@code null}) ⇒ unlimited — most profiles rely on this. An
* explicit {@code 0} (CB-585) means "cap this profile at zero live
* members": it is excluded from every automatic policy's candidate pool
* the same way a {@code weight <= 0} profile is (see
* {@code PlacementPolicyUtil.available()}, which already treats "at cap"
* and "excluded" alike), and an explicit
* {@code bridge_spawn{profile:"..."}} against it is refused too (see
* {@code CompositePeerLauncher.enforceMaxLoad}) — a cap is a capacity
* statement that does not stop being true just because the profile was
* named directly. A negative value has no sane meaning (there is no
* "excluded" to degrade to below zero) and is refused at config load
* instead, naming the profile and the key. Live means any session the
* registry still owns (acquired and not yet released), in any state.
* @param kind which peer launcher spawns this profile: {@code "claude-code"} (default —
* the {@link dev.ltms.bridged.member.ClaudeCodeLauncher}) or {@code "opencode"}.
* The {@code CompositePeerLauncher} routes {@code spawn}/reap by this value, so
@@ -171,6 +241,23 @@ public record BridgedConfig(
* opposite intents). For the same reason, an {@code env:} entry naming
* {@code ANTHROPIC_BASE_URL} or {@code ANTHROPIC_AUTH_TOKEN} is refused at
* config load (CB-542): on the subscription path no guard would vet it.
* @param exhaustedPattern regex matched against a completion-fallback scrape (CB-578 stage A) to
* classify a turn that ended with no {@code bridge_reply} as the backend
* having refused on a subscription usage limit, rather than a real answer.
* {@code null}/blank ⇒ the classification never fires for this profile and
* today's completion-fallback behaviour is unchanged. Every backend words
* its refusal differently, so this is config, never a vendor string in code.
* @param credentialId the shared account this profile authenticates as (CB-578 stage B). Two
* or more profiles setting the <em>same</em> non-blank value are quarantined
* together by one {@code exhaustedPattern} classification on any one of
* them — the case a single OpenAI (or any other) credential backing several
* profiles ({@code sol}, {@code terra}, ...) needs, so a spawn does not walk
* straight onto the other profile sharing the same exhausted account.
* {@code null}/blank ⇒ this profile's {@link #effectiveCredentialId()} is
* its own name, so it quarantines alone — today's behaviour for every
* profile that does not opt in. Read live off the current config, so it is
* HOT: a change takes effect on the next exhaustion classification / spawn,
* no restart needed.
*/
@JsonIgnoreProperties(ignoreUnknown = true)
public record Profile(String profile, String baseUrl, String model,
@@ -183,7 +270,9 @@ public record BridgedConfig(
Map<String, String> env,
Float weight,
Integer maxLoad,
Boolean subscription) {
Boolean subscription,
String exhaustedPattern,
String credentialId) {
/** Peer kind spawned by {@link dev.ltms.bridged.member.ClaudeCodeLauncher} (the default). */
public static final String KIND_CLAUDE_CODE = "claude-code";
@@ -223,9 +312,25 @@ public record BridgedConfig(
// checkpoints need only set gitTokenEnv; it is injected only alongside a resolved token.
gitHostEnv = (gitHostEnv == null || gitHostEnv.isBlank()) ? "GITEA_HOST" : gitHostEnv;
env = (env == null) ? Map.of() : Map.copyOf(env);
weight = (weight == null || weight <= 0.0f) ? 1.0f : weight;
maxLoad = (maxLoad == null || maxLoad <= 0) ? null : maxLoad;
// CB-554: absent still means 1.0, but an explicit non-positive value must survive as
// "excluded from automatic selection" (PlacementCandidate.excluded(), weight <= 0), not
// get coerced back up to 1.0 — that coercion was the bug (weight: 0 looked like "never
// pick me" and actually meant "pick me as often as anyone else").
weight = (weight == null) ? 1.0f : Math.max(weight, 0.0f);
// CB-585: absent still means unlimited (null), but an explicit maxLoad: 0 must survive
// as "capped at zero," not get coerced back up to null/unlimited — that coercion was
// the bug (maxLoad: 0 read as "never run anything here" and behaved as the opposite,
// the one throttle a subscription: true profile has against the operator's own paid
// plan). A negative value is refused earlier, at config load (rejectNegativeMaxLoad),
// so it never reaches this constructor and needs no clamping here.
subscription = (subscription != null && subscription) ? Boolean.TRUE : Boolean.FALSE;
// exhaustedPattern stays null when unset/blank (opt-in) — no defaulting, no vendor
// wording: an unconfigured profile keeps today's completion-fallback behaviour exactly.
exhaustedPattern = (exhaustedPattern == null || exhaustedPattern.isBlank()) ? null : exhaustedPattern;
// credentialId stays null when unset/blank — effectiveCredentialId() is where the
// "quarantines alone" fallback actually lives, so today's behaviour needs no defaulting
// here at all.
credentialId = (credentialId == null || credentialId.isBlank()) ? null : credentialId;
}
/**
@@ -238,7 +343,7 @@ public record BridgedConfig(
String placement, String workspace, String tabLabel, String mcpUrl,
String cwd, List<String> parityOverlay) {
this(profile, baseUrl, model, configDir, tokenEnv, argv, placement, workspace, tabLabel,
mcpUrl, cwd, parityOverlay, null, null, null, null, null, null, null);
mcpUrl, cwd, parityOverlay, null, null, null, null, null, null, null, null);
}
/**
@@ -250,7 +355,7 @@ public record BridgedConfig(
String placement, String workspace, String tabLabel, String mcpUrl,
String cwd, List<String> parityOverlay, String gitTokenEnv, String gitHostEnv) {
this(profile, baseUrl, model, configDir, tokenEnv, argv, placement, workspace, tabLabel,
mcpUrl, cwd, parityOverlay, gitTokenEnv, gitHostEnv, null, null, null, null, null);
mcpUrl, cwd, parityOverlay, gitTokenEnv, gitHostEnv, null, null, null, null, null, null);
}
/**
@@ -263,13 +368,14 @@ public record BridgedConfig(
String cwd, List<String> parityOverlay, String gitTokenEnv, String gitHostEnv,
String kind) {
this(profile, baseUrl, model, configDir, tokenEnv, argv, placement, workspace, tabLabel,
mcpUrl, cwd, parityOverlay, gitTokenEnv, gitHostEnv, kind, null, null, null, null);
mcpUrl, cwd, parityOverlay, gitTokenEnv, gitHostEnv, kind, null, null, null, null, null);
}
/** A copy with {@code profile} set — used to default a profile to its {@code workers} key. */
public Profile withProfile(String p) {
return new Profile(p, baseUrl, model, configDir, tokenEnv, argv, placement, workspace, tabLabel,
mcpUrl, cwd, parityOverlay, gitTokenEnv, gitHostEnv, kind, env, weight, maxLoad, subscription);
mcpUrl, cwd, parityOverlay, gitTokenEnv, gitHostEnv, kind, env, weight, maxLoad, subscription,
exhaustedPattern, credentialId);
}
/** True when this profile is served by the Claude Code adapter (the default kind). */
@@ -302,7 +408,39 @@ public record BridgedConfig(
String cwd, List<String> parityOverlay, String gitTokenEnv, String gitHostEnv,
String kind, Map<String, String> env, Float weight, Integer maxLoad) {
this(profile, baseUrl, model, configDir, tokenEnv, argv, placement, workspace, tabLabel,
mcpUrl, cwd, parityOverlay, gitTokenEnv, gitHostEnv, kind, env, weight, maxLoad, null);
mcpUrl, cwd, parityOverlay, gitTokenEnv, gitHostEnv, kind, env, weight, maxLoad, null, null);
}
/**
* Backward-compatible constructor without the CB-578 stage B {@code credentialId} — the
* profile quarantines alone under its own name (see {@link #effectiveCredentialId()}). Keeps
* pre-stage-B call sites (and any YAML that omits the key) compiling and behaving identically.
*/
public Profile(String profile, String baseUrl, String model,
String configDir, String tokenEnv, List<String> argv,
String placement, String workspace, String tabLabel, String mcpUrl,
String cwd, List<String> parityOverlay, String gitTokenEnv, String gitHostEnv,
String kind, Map<String, String> env, Float weight, Integer maxLoad,
Boolean subscription, String exhaustedPattern) {
this(profile, baseUrl, model, configDir, tokenEnv, argv, placement, workspace, tabLabel,
mcpUrl, cwd, parityOverlay, gitTokenEnv, gitHostEnv, kind, env, weight, maxLoad,
subscription, exhaustedPattern, null);
}
/** True when this profile's CB-578 stage A backend-exhausted classification is configured. */
public boolean hasExhaustedPattern() {
return exhaustedPattern != null;
}
/**
* The credential group this profile quarantines with (CB-578 stage B): the configured
* {@link #credentialId} when set, else this profile's own name — so an unconfigured profile
* quarantines alone, exactly as it did before this field existed. Two profiles that set the
* same non-blank {@code credentialId} share one quarantine: a {@code BACKEND_EXHAUSTED}
* classification on either one quarantines both.
*/
public String effectiveCredentialId() {
return (credentialId == null || credentialId.isBlank()) ? profile : credentialId;
}
/** True when this profile's workers are granted a forge token to open their own PR (CB-302). */
@@ -377,6 +515,17 @@ public record BridgedConfig(
boolean clearAfterTurn) {
}
/** Optional fleet detection. A missing block stays dormant. */
@JsonIgnoreProperties(ignoreUnknown = true)
public record Health(Boolean enabled, Integer intervalSeconds, Integer workingSuspectAfterSeconds,
Integer paneProbeIntervalSeconds, Notifications notifications) {
public boolean isEnabled() { return Boolean.TRUE.equals(enabled); }
public int intervalOrDefault() { return Math.max(15, intervalSeconds == null ? 30 : intervalSeconds); }
public record Notifications(String mode) {
public boolean configured() { return "webhook".equalsIgnoreCase(mode); }
}
}
/**
* External AMQP broker for durable, cross-restart reply delivery (CB-307 Stage 2). Its mere
* presence swaps the in-memory {@code ReplyInbox} for the AMQP-backed adapter; absent, bridged
@@ -437,24 +586,32 @@ public record BridgedConfig(
* pre-existed, which is why it had to be recognised by configuration rather than created. With
* {@code profile} and {@code instances} the daemon may stand one up when none is live, so the
* pane no longer has to exist before the daemon does. Recognition still comes first: a lead
* already running under {@code tabPrefix} is adopted, and only the shortfall is launched.
* already running in its configured {@code tab} is adopted, and only the shortfall is launched.
*
* <p><b>{@code tab} replaced {@code terminal} (CB-579).</b> A herdr {@code terminal_id} changes
* every time the lead's session restarts, so pinning one cost a config edit and a daemon restart
* per restart. A tab is stable: a human opens it once, it holds exactly one pane, and its label
* survives restarts of the agent inside it — so identity is now the tab label alone.
*
* @param profile the {@code profiles:} entry to launch this lead on when one must
* be created; {@code null} ⇒ recognise-only, never create
* @param terminal the lead's herdr {@code terminal_id} when pinned by hand; the only
* field identity depends on. {@code null} ⇒ found by {@code tabPrefix}
* @param tab the exact tab label hosting this lead, matched case-insensitively;
* the only field identity depends on. Required — a lead with no
* {@code tab} can never be discovered, launched or not
* @param instances how many of this lead should be live (default 1). The daemon
* launches only the shortfall, so a restart adopts rather than doubles
* @param tabPrefix label prefix marking this lead's tab, matched case-insensitively;
* the remainder is the lead's name ({@code "lead: opus"} →
* {@code opus}). Default {@code "lead:"}
* @param tabPrefix no longer used to find a lead's tab — {@code tab} is matched
* exactly. Its only remaining job is the startup collision guard
* ({@link #validateLeadTabPrefixes()}), which still uses it to refuse
* a worker {@code tabLabel} template that could be misread as a lead.
* Default {@code "lead:"}
* @param scanIntervalSeconds how long a tab scan is cached before herdr is asked again; also the
* worst case before a newly-labelled tab is recognised. Default 10
* @param kind which agent runs there ({@code claude}, {@code opencode}, …)
* @param model the model or selector it runs, for operators reading the roster
*/
@JsonIgnoreProperties(ignoreUnknown = true)
public record Leader(String profile, String terminal, Integer instances, String tabPrefix,
public record Leader(String profile, String tab, Integer instances, String tabPrefix,
Integer scanIntervalSeconds, String kind, String model,
String workspace, String cwd) {
@@ -472,12 +629,13 @@ public record BridgedConfig(
(scanIntervalSeconds == null || scanIntervalSeconds <= 0) ? 10 : scanIntervalSeconds;
workspace = (workspace == null || workspace.isBlank())
? DEFAULT_WORKSPACE : workspace.strip();
tab = (tab == null || tab.isBlank()) ? null : tab.strip();
}
/** Back-compat 7-arg form — no workspace or cwd, so both take their defaults. */
public Leader(String profile, String terminal, Integer instances, String tabPrefix,
public Leader(String profile, String tab, Integer instances, String tabPrefix,
Integer scanIntervalSeconds, String kind, String model) {
this(profile, terminal, instances, tabPrefix, scanIntervalSeconds, kind, model, null, null);
this(profile, tab, instances, tabPrefix, scanIntervalSeconds, kind, model, null, null);
}
/** True when this lead may be launched by the daemon rather than only recognised. */
@@ -485,9 +643,9 @@ public record BridgedConfig(
return profile != null && !profile.isBlank() && instances > 0;
}
/** The tab label an auto-launched instance of this lead gets — what the scanner reads back. */
public String tabLabel(String name) {
return tabPrefix + " " + name;
/** The tab label an auto-launched instance of this lead gets — its configured {@code tab}. */
public String tabLabel() {
return tab;
}
}
@@ -686,28 +844,20 @@ public record BridgedConfig(
}
/**
* The terminal → lead-name map that {@link dev.ltms.bridged.auth.CallerResolver} resolves
* against, merging the {@code leaders:} registry with the legacy singular {@code primary:} pin.
* The terminal → lead-name map seeded from the legacy singular {@code primary:} pin (CB-530).
*
* <p>Precedence: an explicit {@code leaders:} entry wins over the {@code primary:} pin for the
* same terminal. The pin is the older, less expressive spelling of the same fact, so when both
* name a pane the named entry is the one an operator meant. The pin is still honoured on its
* own — a config carrying only {@code primary:} behaves exactly as it did before CB-530.
* <p>{@code fleet.leaders} no longer carries a per-entry terminal pin (CB-579): a lead's identity
* comes from its {@code tab} alone, resolved live by {@code LeadTabScanner}. This method now
* exists only for the {@code primary.terminal} fallback — a config that never migrated off it
* still resolves that one pane as a lead named {@code "primary"}, exactly as before CB-530.
*
* @return an unmodifiable map, empty when neither block is configured (nothing is pinned, and
* every pane therefore resolves as a worker — the pre-CB-307 behaviour)
* @return an unmodifiable map, empty when {@code primary.terminal} is not configured (nothing is
* pinned, and every pane therefore resolves as a worker — the pre-CB-307 behaviour)
*/
public Map<String, String> leaderTerminals() {
Map<String, String> byTerminal = new LinkedHashMap<>();
if (fleet != null) {
fleet.leaders().forEach((name, leader) -> {
if (leader != null && leader.terminal() != null && !leader.terminal().isBlank()) {
byTerminal.put(leader.terminal(), name);
}
});
}
if (primary != null && primary.terminal() != null && !primary.terminal().isBlank()) {
byTerminal.putIfAbsent(primary.terminal(), "primary");
byTerminal.put(primary.terminal(), "primary");
}
return Collections.unmodifiableMap(byTerminal);
}
@@ -812,15 +962,17 @@ public record BridgedConfig(
private static final Set<String> KNOWN_TOP_LEVEL_KEYS = Set.of(
"bind", "herdrSocket", "profiles", "guard", "worktreeRoot",
"lifecycle", "spawnReadyTimeoutMs", "spawnReadyPollMs", "broker", "primary", "fleet",
"leadHeartbeat", "placement", "auth", "configReload");
"leadHeartbeat", "health", "placement", "auth", "configReload", "quarantineCooldownSeconds");
/** Load and validate config from {@code path}. */
public static BridgedConfig load(Path path) {
try {
String yaml = Files.readString(path);
rejectRenamedTopLevelKeys(yaml);
rejectLeaderTerminalKey(yaml);
warnUnknownTopLevelKeys(yaml, path);
rejectDuplicateMemberSlots(yaml);
rejectNegativeMaxLoad(yaml);
BridgedConfig cfg = YAML.readValue(yaml, BridgedConfig.class);
return cfg.withDefaults();
} catch (IOException e) {
@@ -1043,6 +1195,78 @@ public record BridgedConfig(
}
}
/**
* Reject a config whose {@code fleet.leaders.<name>} still carries the retired {@code terminal:}
* pin (CB-579), naming {@code tab:} as its replacement.
*
* <p>{@code Leader} is {@code @JsonIgnoreProperties(ignoreUnknown = true)}, so simply dropping
* the record component would make a leftover {@code terminal:} key silently no-op — the daemon
* would start, the pin would never take effect, and nothing would say why. Fatal and specific
* instead, exactly like {@link #rejectRenamedTopLevelKeys}, which this mirrors for a key one
* level deeper than the ones that method covers.
*
* @param yaml the raw config text
* @throws IllegalStateException when any {@code fleet.leaders.<name>.terminal} key is present
*/
static void rejectLeaderTerminalKey(String yaml) {
Map<?, ?> raw;
try {
raw = YAML.readValue(yaml, Map.class);
} catch (IOException | IllegalArgumentException e) {
return; // a malformed file is reported by the real parse, not here
}
if (raw == null || !(raw.get("fleet") instanceof Map<?, ?> fleet)
|| !(fleet.get("leaders") instanceof Map<?, ?> leaders)) {
return;
}
List<String> bad = leaders.entrySet().stream()
.filter(e -> e.getValue() instanceof Map<?, ?> leader && leader.containsKey("terminal"))
.map(e -> String.valueOf(e.getKey()))
.sorted()
.toList();
if (!bad.isEmpty()) {
throw new IllegalStateException("refusing to start: fleet.leaders entries ["
+ String.join(", ", bad) + "] still use the retired 'terminal:' key — replace it "
+ "with 'tab:', the exact tab label hosting the lead. A terminal_id changes on "
+ "every restart of the lead's session; a tab label does not.");
}
}
/**
* Reject a profile whose {@code maxLoad:} is negative, naming both the profile and the key.
*
* <p>Unlike {@code weight} (CB-554), where a negative value degrades to the same "excluded"
* meaning as an explicit {@code 0}, {@code maxLoad} has nowhere lower to degrade to — {@code 0}
* already means "capped at zero live members" (CB-585), the strictest cap there is. Silently
* normalising a negative value to something else is exactly the shape of bug this ticket fixes
* for {@code 0}, so it is refused instead, loud and specific, rather than guessed at.
*
* @param yaml the raw config text
* @throws IllegalStateException when any profile's {@code maxLoad} is negative
*/
static void rejectNegativeMaxLoad(String yaml) {
Map<?, ?> raw;
try {
raw = YAML.readValue(yaml, Map.class);
} catch (IOException | IllegalArgumentException e) {
return; // a malformed file is reported by the real parse, not here
}
if (raw == null || !(raw.get("profiles") instanceof Map<?, ?> profiles)) {
return;
}
List<String> bad = profiles.entrySet().stream()
.filter(e -> e.getValue() instanceof Map<?, ?> p
&& p.get("maxLoad") instanceof Number n && n.doubleValue() < 0)
.map(e -> String.valueOf(e.getKey()))
.sorted()
.toList();
if (!bad.isEmpty()) {
throw new IllegalStateException("refusing to start: profile(s) [" + String.join(", ", bad)
+ "] set a negative maxLoad — maxLoad caps live workers at a non-negative count;"
+ " use 0 to cap a profile at zero live members, or omit the key for unlimited.");
}
}
static List<String> unknownTopLevelKeys(String yaml) {
Map<?, ?> raw;
try {
@@ -1081,8 +1305,15 @@ public record BridgedConfig(
// configReload is left as-is: null is "off", and ConfigReload's own compact constructor
// defaults the fields of a block that IS present. Defaulting it here would start watching
// the file for every config that never asked to be watched.
// quarantineCooldownSeconds IS defaulted, unlike leadHeartbeat/configReload above: it has no
// separate on/off switch of its own (BackendQuarantine only ever quarantines a credential
// after a BACKEND_EXHAUSTED classification, which stays opt-in via exhaustedPattern), so a
// config that never mentions it should still get a sane cooldown rather than a null one.
Integer quarantineCooldown = (quarantineCooldownSeconds != null && quarantineCooldownSeconds > 0)
? quarantineCooldownSeconds : DEFAULT_QUARANTINE_COOLDOWN_SECONDS;
return new BridgedConfig(b, herdrSocket, profiles, g, worktreeRoot, l, timeout, pollMs,
broker, primary, f, leadHeartbeat, placementOrDefault, a, configReload);
broker, primary, f, leadHeartbeat, health, placementOrDefault, a, configReload,
quarantineCooldown);
}
/**
@@ -1292,10 +1523,9 @@ public record BridgedConfig(
+ "', which is not a configured profiles: entry (have: " + profiles.keySet()
+ ").");
}
if (!leader.isCreatable() && (leader.terminal() == null || leader.terminal().isBlank())) {
bad.add("fleet.leaders." + name + " can neither be found nor created — it pins no "
+ "terminal: and names no profile: to launch one on. Give it one or the "
+ "other, or drop the entry.");
if (leader.tab() == null || leader.tab().isBlank()) {
bad.add("fleet.leaders." + name + " has no tab: — a lead is now found (and, if "
+ "auto-launched, labelled) purely by its tab, so every entry must name one.");
}
});
if (!bad.isEmpty()) {
@@ -30,13 +30,22 @@ import java.util.function.Supplier;
* {@code fleet:} (every role pool, {@code charters}, and {@code tabLabel}),
* {@code placement:}, and an existing profile's {@code weight} / {@code maxLoad}. Those
* three are read through a supplier on {@code CompositePeerLauncher}, which is what makes
* them hot — not the fact that they are config.</li>
* them hot — not the fact that they are config. <strong>This does NOT include
* {@code fleet.leaders}</strong>: {@code Bridged.main} reads {@code cfg.fleet().leaders()}
* once at startup to build the {@code LeadTabScanner} and the {@code LeadLauncher}, and
* neither is reconstructed on reload — so a lead added, removed, or re-{@code tab}'d under
* {@code fleet.leaders} needs a restart, the same as any deferred key below.</li>
* <li><strong>Deferred</strong> — accepted into the new snapshot, but the wiring built at startup
* keeps the old value until a restart: {@code lifecycle:}, {@code leadHeartbeat:},
* {@code spawnReadyTimeoutMs} / {@code spawnReadyPollMs}, {@code guard:},
* {@code worktreeRoot:}, adding or removing a profile (a new backend needs its own launcher,
* {@code spawnReadyTimeoutMs} / {@code spawnReadyPollMs}, {@code quarantineCooldownSeconds}
* (CB-578 stage B — baked once into the {@code BackendQuarantine} built at startup),
* {@code guard:}, {@code worktreeRoot:}, adding or removing a profile (a new backend needs its own launcher,
* which is constructed once), <em>and an existing profile's launch settings</em> —
* {@code model}, {@code baseUrl}, {@code argv}, {@code env}, {@code mcpUrl} and the rest.
* {@code model}, {@code baseUrl}, {@code argv}, {@code env}, {@code mcpUrl},
* {@code exhaustedPattern} (CB-578 stage A — compiled once into {@code Bridged.main}'s
* pattern map at startup), and the rest. {@code credentialId} (CB-578 stage B) is NOT on
* this list — it is read live off the config supplier at every quarantine check and
* exhaustion event, exactly like {@code weight} / {@code maxLoad}, so it is hot instead.
* {@code HerdrPeerLauncher} takes {@code Map.copyOf(profiles)} at construction and resolves
* each spawn out of that copy, so those never reach a launch until the daemon restarts. A
* reload logs these rather than pretending they applied.</li>
@@ -211,6 +220,12 @@ public final class ConfigRef implements Supplier<BridgedConfig> {
|| !Objects.equals(old.spawnReadyPollMs(), fresh.spawnReadyPollMs())) {
changed.add("spawnReady*");
}
// CB-578 stage B: baked once into the BackendQuarantine built at startup — a running
// quarantine keeps its original cooldown regardless, and a new cooldown only applies to a
// quarantine that starts after a restart.
if (!Objects.equals(old.quarantineCooldownSeconds(), fresh.quarantineCooldownSeconds())) {
changed.add("quarantineCooldownSeconds");
}
Map<String, BridgedConfig.Profile> before =
old.profiles() == null ? Map.of() : old.profiles();
Map<String, BridgedConfig.Profile> after =
@@ -246,8 +261,10 @@ public final class ConfigRef implements Supplier<BridgedConfig> {
/**
* Whether two versions of a profile would launch a peer identically. Compares every component
* the launcher reads at spawn; {@code weight} and {@code maxLoad} are excluded because those are
* read live by the placement policy and really do take effect on the next spawn.
* the launcher reads at spawn; {@code weight}, {@code maxLoad} and {@code credentialId} are
* excluded because those are read live (by the placement policy and, for credentialId, by
* {@code CompositePeerLauncher}/the CB-578 stage B exhaustion sink) and really do take effect on
* the next spawn.
*/
private static boolean sameLaunchSettings(BridgedConfig.Profile a, BridgedConfig.Profile b) {
return Objects.equals(a.baseUrl(), b.baseUrl())
@@ -265,6 +282,10 @@ public final class ConfigRef implements Supplier<BridgedConfig> {
&& Objects.equals(a.gitHostEnv(), b.gitHostEnv())
&& Objects.equals(a.kind(), b.kind())
&& Objects.equals(a.env(), b.env())
&& Objects.equals(a.subscription(), b.subscription());
&& Objects.equals(a.subscription(), b.subscription())
// CB-578 stage B: exhaustedPattern is compiled once into Bridged.main's pattern map
// at startup (see ExhaustedPatternLookup wiring) — a reload never re-reads it, so a
// changed pattern must be reported as deferred, exactly like model/baseUrl/argv.
&& Objects.equals(a.exhaustedPattern(), b.exhaustedPattern());
}
}
@@ -0,0 +1,151 @@
package dev.ltms.bridged.health;
import dev.ltms.bridged.herdr.Agent;
import dev.ltms.bridged.herdr.AgentControl;
import dev.ltms.bridged.herdr.AgentStatus;
import dev.ltms.bridged.msg.MessageService;
import dev.ltms.bridged.session.MemberSession;
import org.slf4j.Logger;
import org.slf4j.LoggerFactory;
import java.util.HashMap;
import java.util.HashSet;
import java.util.List;
import java.util.Map;
import java.util.Objects;
import java.util.concurrent.ScheduledExecutorService;
import java.util.concurrent.TimeUnit;
import java.util.function.BiConsumer;
import java.util.function.LongSupplier;
import java.util.function.Supplier;
/** Slow whole-fleet evidence collection. It is deliberately separate from the delivery poller. */
public final class FleetHealthMonitor {
private static final Logger log = LoggerFactory.getLogger(FleetHealthMonitor.class);
/** Bounded attempts to run {@link #failTarget} for one transition. Never retried tick-to-tick (CB-580). */
static final int MAX_FAIL_TARGET_ATTEMPTS = 3;
private final AgentControl agents;
private final Supplier<List<MemberSession>> roster;
private final MessageService messages;
private final ScheduledExecutorService scheduler;
private final LongSupplier clock;
private final long intervalSeconds;
private final BiConsumer<String, String> failTarget;
private final Map<String, HealthPrior> priors = new HashMap<>();
private final Map<String, HealthState> states = new HashMap<>();
// These facts need the evidence publishers introduced by later M4 units. They are not negatives.
private static final boolean NOT_YET_OBSERVED = false;
/**
* @param failTarget CB-568's idempotent target-wide failure operation (e.g. {@code messages::abandon}),
* invoked once when a member transitions into a terminal health state. Required —
* there is deliberately no defaulting overload; a caller that does not want the
* fail-tickets-on-terminal-health behavior must pass an explicit inert value (see
* {@code TestTurnTokens.inert} / {@code BridgeMcp.CapacitySource.none()} for the pattern).
*/
public FleetHealthMonitor(AgentControl agents, Supplier<List<MemberSession>> roster, MessageService messages,
ScheduledExecutorService scheduler, LongSupplier clock, long intervalSeconds,
BiConsumer<String, String> failTarget) {
this.agents = agents;
this.roster = roster;
this.messages = messages;
this.scheduler = scheduler;
this.clock = clock;
this.intervalSeconds = intervalSeconds;
this.failTarget = Objects.requireNonNull(failTarget, "failTarget");
}
/** Pure per-member decision seam. */
static HealthDecision decide(HealthSnapshot snapshot, HealthPrior prior, long nowNanos) {
return FleetHealth.decide(snapshot, prior, nowNanos);
}
public void start() { scheduler.schedule(this::tick, intervalSeconds, TimeUnit.SECONDS); }
public void stop() { scheduler.shutdownNow(); }
// Package-private so tests can run one tick without waiting.
void tick() {
try {
List<Agent> agentsNow = agents.list(); // Exactly one list call for this complete observation.
List<MemberSession> rosterNow = roster.get(); // One in-memory roster snapshot for this tick.
Map<String, Agent> live = new HashMap<>();
for (Agent agent : agentsNow) live.put(agent.terminalId(), agent);
HashSet<String> current = new HashSet<>();
for (MemberSession session : rosterNow) {
current.add(session.terminalId());
Agent agent = live.get(session.terminalId());
AgentStatus status = agent == null ? AgentStatus.UNKNOWN : agent.status();
boolean accepted = messages.hasAcceptedDelivery(session.terminalId());
HealthSnapshot snapshot = new HealthSnapshot(session.state(), status, accepted, NOT_YET_OBSERVED,
messages.hasInboxMessage(session.terminalId()), agent != null, NOT_YET_OBSERVED,
NOT_YET_OBSERVED, NOT_YET_OBSERVED, NOT_YET_OBSERVED, NOT_YET_OBSERVED, NOT_YET_OBSERVED);
HealthDecision decision = decide(snapshot, priors.getOrDefault(session.terminalId(), HealthPrior.NONE),
clock.getAsLong());
priors.put(session.terminalId(), decision.prior());
reportTransition(session.terminalId(), decision.state());
}
priors.keySet().retainAll(current);
states.keySet().retainAll(current);
} catch (Throwable error) {
// A list failure is health evidence, and must never kill the monitor's only scheduler task.
log.warn("fleet health collection failed; will retry next tick", error);
} finally {
if (!scheduler.isShutdown()) {
scheduler.schedule(this::tick, intervalSeconds, TimeUnit.SECONDS);
}
}
}
void reportTransition(String target, HealthState next) {
HealthState previous = states.put(target, next);
if (previous == next) return;
if (fault(next)) {
log.warn("fleet health member={} state={} previous={}", target, next, previous);
} else if (previous != null && fault(previous)) {
log.info("fleet health member={} recovered state={} previous={}", target, next, previous);
}
// CB-580: a member entering GONE/NEVER_READY must not leave its waiting tickets pending
// forever. Fire exactly once per transition — never on a tick where the state is unchanged,
// which is what made the rejected commit call abandon() once per tick for as long as a
// member stayed terminal.
if (terminal(next)) {
failTerminalTarget(target, next);
}
}
private void failTerminalTarget(String target, HealthState state) {
String reason = "fleet health: member reached terminal state " + state.name();
RuntimeException last = null;
for (int attempt = 1; attempt <= MAX_FAIL_TARGET_ATTEMPTS; attempt++) {
try {
failTarget.accept(target, reason);
return;
} catch (RuntimeException error) {
last = error;
log.warn("fleet health: failTarget attempt {}/{} failed for member={} state={}",
attempt, MAX_FAIL_TARGET_ATTEMPTS, target, state, error);
}
}
log.warn("fleet health: giving up on failTarget for member={} state={} after {} attempts",
target, state, MAX_FAIL_TARGET_ATTEMPTS, last);
}
private static boolean terminal(HealthState state) {
return state == HealthState.GONE || state == HealthState.NEVER_READY;
}
private static boolean fault(HealthState state) {
return switch (state) {
case NEVER_READY, GONE, TURN_BOUNDARY_LOST, ERROR_ON_SCREEN, STALL_SUSPECTED,
MUTE, REPLY_STRANDED, DELEGATION_ORPHANED, CONTROL_LINK_DOWN -> true;
default -> false;
};
}
public static String coverage(boolean enabled, boolean notificationConfigured) {
return !enabled ? "off" : notificationConfigured ? "full" : "detection-only";
}
}
@@ -6,6 +6,7 @@ import org.slf4j.LoggerFactory;
import java.util.Collections;
import java.util.LinkedHashMap;
import java.util.Locale;
import java.util.Map;
import java.util.Set;
import java.util.function.LongSupplier;
@@ -23,6 +24,15 @@ import java.util.function.Supplier;
* by first starting the session and asking it. Scanning closes that loop: label the tab, and the
* pane is recognised on the next resolve.
*
* <p><strong>CB-579 — matched by name, not prefix.</strong> This used to strip one shared
* {@code tabPrefix} off a label to derive the lead's name, and merged a config-supplied
* {@code terminal_id} pin over every scan result so the pin could never expire. Both are gone: each
* lead now configures its own exact {@code tab} label ({@code fleet.leaders.<name>.tab}), so this
* class is handed a {@code tab → name} map up front and matches labels against it exactly
* (case-insensitively). There is no merge step — a scan result is the whole answer. That is the
* fix for the bug this replaces: a {@code terminal_id} pin surviving in config after the pane it
* named was gone, so the daemon kept treating a dead session as a live lead forever.
*
* <p><strong>Direction of trust.</strong> The label names the lead; it never <em>grants</em>
* anything a pane could take for itself. Three properties keep that honest:
* <ol>
@@ -48,7 +58,7 @@ import java.util.function.Supplier;
* ever make a decision that <em>removes</em> something based on this map, add the same check.
* The remaining hazard is an <em>operator</em> one — a worker {@code tabLabel} template that
* happens to start with the same prefix would promote the whole fleet — and that is refused at
* startup by {@code BridgedConfig.validateLeadScan} rather than documented here.
* startup by {@code BridgedConfig.validateLeadTabPrefixes} rather than documented here.
*
* <p><strong>Caching.</strong> {@link #get()} is on the request path (every resolve), so the scan
* is TTL-cached and a stale-but-valid map is preferred to a herdr round-trip. A failed scan keeps
@@ -60,38 +70,46 @@ public final class LeadTabScanner implements Supplier<Map<String, String>> {
private static final Logger log = LoggerFactory.getLogger(LeadTabScanner.class);
private final HerdrClient herdr;
private final String tabPrefix;
private final Map<String, String> tabToName;
private final Set<String> excludedWorkspaceLabels;
private final Map<String, String> configuredLeads;
private final long ttlNanos;
private final LongSupplier clock;
private Map<String, String> cached;
private Map<String, String> cached = Map.of();
private long scannedAtNanos;
private boolean everScanned;
/**
* @param herdr the herdr client to query ({@code workspace.list},
* {@code tab.list}, {@code pane.list} — all read-only)
* @param tabPrefix a tab whose label starts with this (case-insensitively) hosts a
* lead; the rest of the label, trimmed, is the lead's name
* @param tabToName every configured lead's exact tab label → its name
* ({@code fleet.leaders.<name>.tab}), matched case-insensitively
* @param excludedWorkspaceLabels workspaces never scanned — the configured worker spaces
* @param configuredLeads the static {@code leaders:}/{@code primary:} registry, merged
* over every scan result. Explicit config outranks the
* convention, and survives a scan that cannot run at all
* @param ttlNanos how long a scan result is reused before the next one
* @param clock nanosecond time source ({@code System::nanoTime} in production)
*/
public LeadTabScanner(HerdrClient herdr, String tabPrefix, Set<String> excludedWorkspaceLabels,
Map<String, String> configuredLeads, long ttlNanos, LongSupplier clock) {
public LeadTabScanner(HerdrClient herdr, Map<String, String> tabToName,
Set<String> excludedWorkspaceLabels, long ttlNanos, LongSupplier clock) {
this.herdr = herdr;
this.tabPrefix = tabPrefix == null || tabPrefix.isBlank() ? "lead:" : tabPrefix.strip();
this.tabToName = normalize(tabToName);
this.excludedWorkspaceLabels = excludedWorkspaceLabels == null
? Set.of() : Set.copyOf(excludedWorkspaceLabels);
this.configuredLeads = configuredLeads == null ? Map.of() : Map.copyOf(configuredLeads);
this.ttlNanos = ttlNanos;
this.clock = clock;
this.cached = this.configuredLeads;
}
/** Keys stripped and lower-cased once, so every lookup is a plain map hit. */
private static Map<String, String> normalize(Map<String, String> tabToName) {
if (tabToName == null || tabToName.isEmpty()) {
return Map.of();
}
Map<String, String> out = new LinkedHashMap<>();
tabToName.forEach((tab, name) -> {
if (tab != null && !tab.isBlank() && name != null && !name.isBlank()) {
out.put(tab.strip().toLowerCase(Locale.ROOT), name);
}
});
return Collections.unmodifiableMap(out);
}
/**
@@ -151,25 +169,20 @@ public final class LeadTabScanner implements Supplier<Map<String, String>> {
}
}
}
byTerminal.putAll(configuredLeads); // an explicit pin outranks a label
return Collections.unmodifiableMap(byTerminal);
}
/**
* The lead name a tab label declares, or {@code null} if it declares none.
* The lead name a tab label declares, or {@code null} if it names none of the configured leads.
*
* <p>{@code "lead: opus-5.0"} → {@code "opus-5.0"}. A bare {@code "lead:"} names nobody and is
* rejected: an unnamed lead would resolve as {@code PRIMARY} with nothing to attribute it to.
* <p>Exact match (case-insensitive, ends stripped) against {@link #tabToName} — no prefix
* stripping, so an operator's {@code "lead: something-else"} tab is never mistaken for a
* configured lead just because it shares a prefix.
*/
private String leadNameOf(String label) {
if (label == null) {
return null;
}
String l = label.strip();
if (!l.regionMatches(true, 0, tabPrefix, 0, tabPrefix.length())) {
return null;
}
String name = l.substring(tabPrefix.length()).strip();
return name.isEmpty() ? null : name;
return tabToName.get(label.strip().toLowerCase(Locale.ROOT));
}
}
@@ -2,11 +2,17 @@ package dev.ltms.bridged.inject;
import dev.ltms.bridged.herdr.AgentControl;
import dev.ltms.bridged.msg.Rendezvous;
import dev.ltms.bridged.msg.TurnToken;
import org.slf4j.Logger;
import org.slf4j.LoggerFactory;
import java.util.List;
import java.util.Objects;
import java.util.Set;
import java.util.TreeSet;
import java.util.concurrent.CompletableFuture;
import java.util.concurrent.ConcurrentHashMap;
import java.util.regex.Pattern;
/**
* The CB-106 completion fallback: bridges the {@link Injector}'s turn-completion signal to the
@@ -60,6 +66,8 @@ public final class CompletionResolver implements TurnListener {
private final AgentControl agents;
private final Rendezvous rendezvous;
private final ExhaustedPatternLookup exhaustedPatterns;
private final ExhaustionSink exhaustionSink;
/**
* Per-target record of the turn currently in flight: the exact {@link Rendezvous} waiter its
@@ -77,23 +85,36 @@ public final class CompletionResolver implements TurnListener {
private final ConcurrentHashMap<String, InFlight> inFlight = new ConcurrentHashMap<>();
public CompletionResolver(AgentControl agents, Rendezvous rendezvous) {
/**
* @param exhaustedPatterns CB-578 stage A: per-target lookup for a profile's configured
* usage-limit refusal pattern. Required — there is deliberately no
* defaulting overload; a caller that does not want the classification
* must pass an explicit inert value ({@link ExhaustedPatternLookup#none()}).
* @param exhaustionSink CB-578 stage B: notified when a {@code BACKEND_EXHAUSTED}
* classification actually resolves a waiter. Required for the same
* reason as {@code exhaustedPatterns} — pass {@link ExhaustionSink#none()}
* to opt out.
*/
public CompletionResolver(AgentControl agents, Rendezvous rendezvous, ExhaustedPatternLookup exhaustedPatterns,
ExhaustionSink exhaustionSink) {
this.agents = agents;
this.rendezvous = rendezvous;
this.exhaustedPatterns = Objects.requireNonNull(exhaustedPatterns, "exhaustedPatterns");
this.exhaustionSink = Objects.requireNonNull(exhaustionSink, "exhaustionSink");
}
@Override
public void onDelivered(String target) {
public void onDelivered(String target, TurnToken token) {
// Capture the exact waiter this turn belongs to (CB-116) and snapshot the pane's pre-turn
// content — what it shows *before* the just-delivered turn produces output — as the staleness
// reference (CB-115). Done synchronously (like the delivering send itself) so both are in
// place before this turn's completion can fire.
captureBaseline(target);
captureBaseline(target, token);
}
/** Capture the in-flight turn: its waiter and pre-turn baseline (the testable core of {@link #onDelivered}). */
void captureBaseline(String target) {
CompletableFuture<Rendezvous.Resolution> waiter = rendezvous.currentWaiter(target);
void captureBaseline(String target, TurnToken token) {
CompletableFuture<Rendezvous.Resolution> waiter = token.waiter();
if (waiter == null) {
inFlight.remove(target); // no send is waiting on this delivery — nothing to resolve later
return;
@@ -137,7 +158,13 @@ public final class CompletionResolver implements TurnListener {
@Override
public void onTurnFailed(String target) {
InFlight turn = inFlight.get(target);
Thread.ofVirtual().name("turn-failed-" + target).start(() -> fail(target, turn));
Thread.ofVirtual().name("turn-failed-" + target).start(() -> fail(target, turn, null));
}
@Override
public void onTurnFailed(String target, String reason) {
InFlight turn = inFlight.get(target);
Thread.ofVirtual().name("turn-failed-" + target).start(() -> fail(target, turn, reason));
}
/** Synchronous resolve (the unit-testable core of {@link #onTurnComplete}). */
@@ -150,11 +177,12 @@ public final class CompletionResolver implements TurnListener {
return;
}
String tail;
String assistantBlock = null;
int originalLength = 0;
boolean clipped = false;
boolean scrapeFailed = false;
try {
String assistantBlock = lastAssistantBlock(agents.read(target, SCRAPE_SOURCE));
assistantBlock = lastAssistantBlock(agents.read(target, SCRAPE_SOURCE));
originalLength = assistantBlock.strip().length();
clipped = originalLength > MAX_SCRAPE_CHARS;
tail = clip(assistantBlock);
@@ -177,6 +205,25 @@ public final class CompletionResolver implements TurnListener {
target);
return; // keep the in-flight record: a later genuine completion still needs it
}
// CB-578 stage A: a turn that ended with no bridge_reply AND whose scrape matches the
// backend's configured usage-limit pattern is a refusal, not an answer. Classify it as
// BACKEND_EXHAUSTED rather than handing the caller a scrape that reads like a real reply.
if (!scrapeFailed) {
Pattern exhausted = exhaustedPatterns.patternFor(target);
String matchedLine = exhausted == null ? null : firstMatchingLine(assistantBlock, exhausted);
if (matchedLine != null) {
String reason = "backend exhausted (usage limit): " + matchedLine;
if (rendezvous.resolveExhausted(waiter, reason)) {
inFlight.remove(target, turn);
log.warn("completion for {} classified BACKEND_EXHAUSTED (no bridge_reply; scrape "
+ "matched the profile's exhausted pattern): {}", target, reason);
// CB-578 stage B: only on the resolution that actually won the race — a late
// duplicate must never quarantine a credential twice for one refusal.
exhaustionSink.onExhausted(target, reason);
}
return;
}
}
String completion = clipped ? tail + "\n" + CLIPPED_PANE_TAIL_MARKER : tail;
if (rendezvous.resolveCompletion(waiter, completion)) {
inFlight.remove(target, turn);
@@ -192,6 +239,11 @@ public final class CompletionResolver implements TurnListener {
/** Synchronous fail (the unit-testable core of {@link #onTurnFailed}). */
void fail(String target, InFlight turn) {
fail(target, turn, null);
}
/** Synchronous fail with an optional reason supplied by a dropped worker queue. */
void fail(String target, InFlight turn, String explicitReason) {
// A never-delivered readiness failure has no in-flight record but still has a blocked send;
// fall back to the currently-registered waiter (unambiguous — that send never completed, so
// no next turn exists to confuse it with).
@@ -201,16 +253,18 @@ public final class CompletionResolver implements TurnListener {
inFlight.remove(target, turn); // nobody blocked on this worker — nothing to fail
return;
}
String reason;
try {
reason = clip(agents.read(target, SCRAPE_SOURCE));
} catch (RuntimeException e) {
reason = "";
}
if (reason.isBlank()) {
// No screen to scrape — either the worker is stuck (CB-109) or gone (CB-110).
reason = "worker did not reply; its turn ended in an unrecoverable state "
+ "(worker unreachable or stuck)";
String reason = explicitReason;
if (reason == null || reason.isBlank()) {
try {
reason = clip(agents.read(target, SCRAPE_SOURCE));
} catch (RuntimeException e) {
reason = "";
}
if (reason.isBlank()) {
// No screen to scrape — either the worker is stuck (CB-109) or gone (CB-110).
reason = "worker did not reply; its turn ended in an unrecoverable state "
+ "(worker unreachable or stuck)";
}
}
if (rendezvous.resolveFailure(waiter, reason)) {
inFlight.remove(target, turn);
@@ -218,6 +272,45 @@ public final class CompletionResolver implements TurnListener {
}
}
/**
* The first line of {@code text} matching {@code pattern}, stripped — the CB-578 stage A
* evidence carried in a {@code BACKEND_EXHAUSTED} reason so the operator sees the real refusal
* text, never a generic label. {@code null} if no line matches.
*/
static String firstMatchingLine(String text, Pattern pattern) {
if (text == null || text.isEmpty()) return null;
for (String line : text.split("\n", -1)) {
if (pattern.matcher(line).find()) {
return line.strip();
}
}
return null;
}
/**
* Coverage summary for the CB-578 stage A exhausted-pattern classification, logged at startup
* the way {@link dev.ltms.bridged.health.FleetHealthMonitor#coverage} is — so an operator can
* see whether the classification is on, and for which profiles, without reading every
* profile's config by hand.
*
* @param allProfiles every configured profile name
* @param configuredProfiles the subset of {@code allProfiles} that carry an exhausted pattern
*/
public static String coverage(Set<String> allProfiles, Set<String> configuredProfiles) {
if (configuredProfiles.isEmpty()) {
return "off (no profile has an exhaustedPattern configured; profiles: " + sorted(allProfiles) + ")";
}
Set<String> unconfigured = new TreeSet<>(allProfiles);
unconfigured.removeAll(configuredProfiles);
return unconfigured.isEmpty()
? "full (all profiles configured: " + sorted(allProfiles) + ")"
: "partial (configured: " + sorted(configuredProfiles) + "; not configured: " + sorted(unconfigured) + ")";
}
private static List<String> sorted(Set<String> names) {
return names.stream().sorted().toList();
}
private static String clip(String s) {
if (s == null) return "";
String trimmed = s.strip();
@@ -0,0 +1,27 @@
package dev.ltms.bridged.inject;
import java.util.regex.Pattern;
/**
* Per-target lookup for a profile's configured usage-limit refusal pattern (CB-578 stage A): how
* {@link CompletionResolver} tells a backend that refused on a subscription usage limit — the
* worker's pane stays healthy, but the account is exhausted — apart from a genuine completion.
*
* <p>The pattern is always profile config, never a vendor string in Java source: every backend
* words its refusal differently, so a hardcoded sentence would only ever match one of them.
*/
@FunctionalInterface
public interface ExhaustedPatternLookup {
/** The compiled pattern configured for {@code target}'s profile, or {@code null} if none. */
Pattern patternFor(String target);
/**
* Inert lookup — no profile has a pattern configured, so the classification never fires and
* the completion fallback behaves exactly as before CB-578 stage A. The explicit stand-in a
* caller (or a test not exercising this feature) passes instead of a defaulting overload.
*/
static ExhaustedPatternLookup none() {
return target -> null;
}
}
@@ -0,0 +1,31 @@
package dev.ltms.bridged.inject;
/**
* Notified when {@link CompletionResolver} actually delivers a {@code BACKEND_EXHAUSTED}
* classification to a waiting send (CB-578 stage B) — never on a race that lost (see
* {@link CompletionResolver#resolve}, which only calls this after
* {@code Rendezvous.resolveExhausted} returns {@code true}).
*
* <p>{@link CompletionResolver} knows only {@code target} (a herdr terminal id); it has no notion of
* profiles or credentials, so mapping {@code target} to whatever should be quarantined is entirely
* the sink's job — see {@code Bridged.main}'s wiring, which resolves target → session → profile →
* {@code effectiveCredentialId()} and calls {@code BackendQuarantine.quarantine} on it.
*/
@FunctionalInterface
public interface ExhaustionSink {
/**
* @param target the herdr terminal id whose turn was classified {@code BACKEND_EXHAUSTED}
* @param reason the matched-line reason carried by the classification
*/
void onExhausted(String target, String reason);
/**
* Inert sink — nothing happens on exhaustion. The explicit stand-in a caller (or a test not
* exercising this feature) passes instead of a defaulting overload, exactly like
* {@link ExhaustedPatternLookup#none()}.
*/
static ExhaustionSink none() {
return (target, reason) -> { };
}
}
@@ -2,6 +2,7 @@ package dev.ltms.bridged.inject;
import dev.ltms.bridged.herdr.AgentControl;
import dev.ltms.bridged.herdr.AgentStatus;
import dev.ltms.bridged.msg.TurnToken;
import org.slf4j.Logger;
import org.slf4j.LoggerFactory;
@@ -129,7 +130,7 @@ public final class Injector {
}
/** A pending message and the future that completes when it has been delivered. */
private record Pending(String text, CompletableFuture<Void> delivered) {
private record Pending(String text, TurnToken token, CompletableFuture<Void> delivered) {
}
/** Per-worker delivery state, guarded by its own monitor (single writer per worker). */
@@ -159,9 +160,9 @@ public final class Injector {
* <p>Uses an atomic map update so a concurrent {@link #drop} cannot slip between "find the
* target" and "queue the message" and orphan it in a target it just removed.
*/
public CompletableFuture<Void> enqueue(String target, String text) {
public CompletableFuture<Void> enqueue(String target, String text, TurnToken token) {
CompletableFuture<Void> delivered = new CompletableFuture<>();
Pending p = new Pending(text, delivered);
Pending p = new Pending(text, token, delivered);
targets.compute(target, (_, existing) -> {
Target t = (existing != null) ? existing : new Target();
t.add(p); // synchronized on the Target monitor — atomic with a concurrent drop
@@ -356,7 +357,7 @@ public final class Injector {
} else {
// Baseline the pane's pre-turn content so a misattributed completion (no new output)
// can't resolve this send with the previous turn's stale answer (CB-115).
turnListener.onDelivered(target);
turnListener.onDelivered(target, sent.token());
sent.delivered().complete(null);
}
}
@@ -410,8 +411,8 @@ public final class Injector {
for (Pending p : pending) {
p.delivered().completeExceptionally(cause);
}
if (hadDeliveredTurn) {
turnListener.onTurnFailed(target);
}
// A queued send has no in-flight record, while a delivered turn does. CompletionResolver
// handles both forms and resolves its waiter at most once.
turnListener.onTurnFailed(target, cause.getMessage());
}
}
@@ -1,5 +1,7 @@
package dev.ltms.bridged.inject;
import dev.ltms.bridged.msg.TurnToken;
/**
* Notified when a worker's delegated turn is observed to complete — a confirmed
* {@code WORKING → IDLE} transition after a delivery. This is the CB-106 completion signal the
@@ -41,6 +43,14 @@ public interface TurnListener {
default void onTurnFailed(String target) {
}
/**
* As {@link #onTurnFailed(String)}, carrying the reason a worker became unreachable. The default
* keeps existing listeners working while allowing the completion resolver to report a useful cause.
*/
default void onTurnFailed(String target, String reason) {
onTurnFailed(target);
}
/**
* A message was just delivered into {@code target}'s pane (CB-115). Fired so the completion
* resolver can snapshot the pane's pre-turn content: a later {@link #onTurnComplete} whose
@@ -49,7 +59,7 @@ public interface TurnListener {
* resolve the send with the previous turn's stale answer. A default no-op keeps the interface
* functional for callers that don't scrape.
*/
default void onDelivered(String target) {
default void onDelivered(String target, TurnToken token) {
}
/** No-op default for callers that only need delivery, not completion signalling. */
@@ -59,7 +59,7 @@ public final class LeadLauncher {
/**
* @param agents herdr agent control (start, list)
* @param spaces workspace / tab control (ensure, create, label, list)
* @param cfg the loaded config — {@code fleet.leaders}, {@code profiles} and the lead pins
* @param cfg the loaded config — {@code fleet.leaders}, {@code profiles} and each lead's tab
*/
public LeadLauncher(AgentControl agents, WorkspaceControl spaces, BridgedConfig cfg) {
this.agents = agents;
@@ -102,8 +102,8 @@ public final class LeadLauncher {
continue;
}
if (!lead.isCreatable()) {
// A lead with a `terminal:` pin and no `profile:` is recognise-only by design: the
// operator opens it by hand. Say so once rather than looking like a silent failure.
// A lead with a `tab:` but no `profile:` is recognise-only by design: the operator
// opens it by hand. Say so once rather than looking like a silent failure.
log.info("lead '{}' is not live, and names no profile — it can be recognised but not "
+ "launched. Add `profile:` under fleet.leaders.{} to have bridged start it.",
name, name);
@@ -127,18 +127,16 @@ public final class LeadLauncher {
}
/**
* How many live leads exist per configured name.
* How many live leads exist per configured name: a running agent in a tab labelled with that
* lead's exact {@code tab} (CB-579). Member workspaces are excluded, exactly as the scanner
* excludes them: a member must not be counted as a lead because it happens to sit in a matching
* tab.
*
* <p>Two independent pieces of evidence, because either alone double-spawns:
* <ul>
* <li>a running agent in a tab labelled {@code "<tabPrefix> <name>"} — how an auto-launched
* lead, or an operator following the labelling convention, is found;</li>
* <li>a running agent on a terminal the config pins in {@code fleet.leaders.<name>.terminal} —
* how a lead the operator opened and pinned by hand is found. Without this, a pinned lead
* whose tab carries no matching label would be relaunched on every boot.</li>
* </ul>
* Member workspaces are excluded, exactly as the scanner excludes them: a member must not be
* counted as a lead because it happens to sit in a matching tab.
* <p>There used to be a second path here — a running agent on the terminal a
* {@code fleet.leaders.<name>.terminal} pin named, for a lead opened and pinned by hand. That
* pin is retired: {@code tab} is now the only field identity depends on, and {@link Agent}
* already carries {@link Agent#tabId()} directly, so a hand-opened lead is found the same way an
* auto-launched one is — by labelling its tab to match.
*/
private Map<String, Integer> liveLeads(Map<String, BridgedConfig.Leader> leaders) {
Set<String> memberSpaces = cfg.profiles().values().stream()
@@ -160,20 +158,9 @@ public final class LeadLauncher {
}
}
// terminalId → the lead name the config pins it to.
Map<String, String> nameByPinnedTerminal = new LinkedHashMap<>();
leaders.forEach((name, lead) -> {
if (lead.terminal() != null && !lead.terminal().isBlank()) {
nameByPinnedTerminal.put(lead.terminal().strip(), name);
}
});
Map<String, Integer> counts = new LinkedHashMap<>();
for (Agent a : agents.list()) {
String name = nameByTab.get(a.tabId());
if (name == null) {
name = nameByPinnedTerminal.get(a.terminalId());
}
if (name != null) {
counts.merge(name, 1, Integer::sum);
}
@@ -184,7 +171,7 @@ public final class LeadLauncher {
/**
* The configured lead a tab label names, or {@code null} for a label that names none.
*
* <p>Matched against the declared lead names rather than by splitting on the prefix, so an
* <p>Matched exactly (case-insensitively) against each lead's configured {@code tab}, so an
* operator's {@code "lead: something-else"} tab is not mistaken for a configured lead.
*/
private String leadNameOf(String label, Map<String, BridgedConfig.Leader> leaders) {
@@ -193,7 +180,8 @@ public final class LeadLauncher {
}
String l = label.strip();
for (Map.Entry<String, BridgedConfig.Leader> e : leaders.entrySet()) {
if (l.equalsIgnoreCase(e.getValue().tabLabel(e.getKey()).strip())) {
String tab = e.getValue().tabLabel();
if (tab != null && l.equalsIgnoreCase(tab.strip())) {
return e.getKey();
}
}
@@ -202,7 +190,7 @@ public final class LeadLauncher {
/** Start one lead. Returns false (having logged) rather than throwing on any failure. */
private boolean launch(String name, BridgedConfig.Leader lead, BridgedConfig.Profile profile) {
String label = lead.tabLabel(name);
String label = lead.tabLabel();
String cwd = (lead.cwd() == null || lead.cwd().isBlank())
? System.getProperty("user.dir") : lead.cwd();
@@ -0,0 +1,30 @@
package dev.ltms.bridged.logging;
import ch.qos.logback.classic.Level;
import ch.qos.logback.classic.Logger;
import ch.qos.logback.classic.turbo.TurboFilter;
import ch.qos.logback.core.spi.FilterReply;
import io.modelcontextprotocol.spec.McpSchema;
import org.slf4j.Marker;
/** Suppresses only the SDK warning for the normal MCP cancellation notification. */
public final class McpCancelledNotificationFilter extends TurboFilter {
static final String LOGGER = "io.modelcontextprotocol.spec.McpStreamableServerSession";
static final String UNHANDLED_NOTIFICATION = "No handler registered for notification method: {}";
@Override
public FilterReply decide(Marker marker, Logger logger, Level level, String format, Object[] params,
Throwable throwable) {
if (level == Level.WARN
&& LOGGER.equals(logger.getName())
&& UNHANDLED_NOTIFICATION.equals(format)
&& params != null
&& params.length == 1
&& params[0] instanceof McpSchema.JSONRPCNotification notification
&& "notifications/cancelled".equals(notification.method())) {
return FilterReply.DENY;
}
return FilterReply.NEUTRAL;
}
}
@@ -14,6 +14,7 @@ import dev.ltms.bridged.herdr.HerdrException;
import dev.ltms.bridged.msg.MessageService;
import dev.ltms.bridged.msg.Rendezvous;
import dev.ltms.bridged.peer.PeerUnreachableException;
import dev.ltms.bridged.placement.BackendQuarantine;
import dev.ltms.bridged.session.SessionManager;
import dev.ltms.bridged.session.MemberSession;
import dev.ltms.bridged.session.WorktreeRequest;
@@ -33,6 +34,7 @@ import jakarta.servlet.http.HttpServlet;
import java.util.LinkedHashMap;
import java.util.List;
import java.util.Map;
import java.util.Objects;
import java.util.Set;
import java.util.function.Function;
import java.util.function.LongSupplier;
@@ -77,6 +79,8 @@ public final class BridgeMcp {
private final CallerResolver authz; // CB-501: null → authorization not enforced (legacy)
private final Metrics metrics; // CB-502: null → auth failures not counted
private final CapacitySource capacity;
private final HealthCoverageSource healthCoverage;
private final QuarantineSource quarantine;
/** Capacity facts used by {@code bridge_list}; production must supply the placement live count. */
public record CapacitySource(Function<String, Integer> liveCount, Function<String, Integer> maxLoad,
@@ -86,17 +90,34 @@ public final class BridgeMcp {
boolean available() { return !configuredProfiles.get().isEmpty(); }
}
/** Coverage is supplied by the health wiring, not inferred from a missing dependency. */
public record HealthCoverageSource(Supplier<String> value) { }
/**
* CB-578 stage B quarantine facts used by {@code bridge_profiles}: a profile → credential id
* lookup, plus the shared {@link BackendQuarantine} to read remaining cooldowns off.
*/
public record QuarantineSource(Function<String, String> credentialIdFor, BackendQuarantine quarantine) {
/** Inert source — no profile is ever reported quarantined. Explicit stand-in, not a default. */
public static QuarantineSource none() { return new QuarantineSource(_ -> null, BackendQuarantine.none()); }
}
/**
* @param callers resolves each call's {@link Principal}; {@code null} disables authorization.
* This surface needs its own enforcement: {@code /mcp} is a raw servlet on
* Jetty's context handler and never passes through Javalin's {@code before}
* filter, so the REST guard does not cover it.
* @param metrics registry for auth-failure counting; may be {@code null}
* @param quarantine CB-578 stage B facts for {@code bridge_profiles}; required — pass
* {@link QuarantineSource#none()} for a caller that does not want the feature
*/
public BridgeMcp(MessageService messages, PeerLauncher workers, SessionManager sessions,
ConnectionIdentity identity, MemberPresence presence, PrimaryRegistry primaryRegistry,
CallerResolver callers, Metrics metrics, CapacitySource capacity) {
CallerResolver callers, Metrics metrics, CapacitySource capacity, HealthCoverageSource healthCoverage,
QuarantineSource quarantine) {
this.capacity = capacity;
this.quarantine = Objects.requireNonNull(quarantine, "quarantine");
this.healthCoverage = healthCoverage;
McpJsonMapper json = new JacksonMcpJsonMapperSupplier().get();
this.transport = HttpServletStreamableServerTransportProvider.builder()
.jsonMapper(json)
@@ -205,12 +226,13 @@ public final class BridgeMcp {
// CB-301-ext: optional isolated worktree for parallel implementers.
String callerCwd = identity.cwdForPid(callerPid(exchange));
return spawn(sessions, str(a, "profile"), str(a, "role"), str(a, "cwd"), callerCwd,
callerTerminal(exchange), worktreeRequest(a));
callerTerminal(exchange), worktreeRequest(a),
str(a, "sessionName"), str(a, "resumeSessionId"));
})
.toolCall(listTool(), (exchange, _) -> {
McpSchema.CallToolResult denied = deny(exchange, Authz.Action.READ, null);
if (denied != null) return denied;
return listFleet(workers, sessions, messages, capacity,
return listFleet(workers, sessions, messages, capacity, healthCoverage, quarantine,
callers == null ? Map.of() : callers.leads(),
callerTerminal(exchange));
})
@@ -223,7 +245,7 @@ public final class BridgeMcp {
.toolCall(profilesTool(), (exchange, _) -> {
McpSchema.CallToolResult denied = deny(exchange, Authz.Action.READ, null);
if (denied != null) return denied;
return profiles(workers);
return profiles(workers, quarantine);
})
.toolCall(whoamiTool(), (exchange, _) -> {
McpSchema.CallToolResult denied = deny(exchange, Authz.Action.READ, null);
@@ -453,6 +475,11 @@ public final class BridgeMcp {
"[worker finished without a structured bridge_reply — transcript tail follows]\n" + r.text());
// The worker ran the turn then wedged (CB-109) — surface the error context.
case WORKER_FAILED -> text("[worker failed — turn ended in an unrecoverable state]\n" + r.text());
// The backend refused on a subscription usage limit (CB-578 stage A) — the worker's
// pane stayed healthy, but its account is exhausted. Distinct from WORKER_FAILED so the
// primary gets the real cause, not a generic wedge.
case BACKEND_EXHAUSTED -> text("[backend exhausted — the worker's account refused on a "
+ "usage limit]\n" + r.text());
// The worker paused mid-turn to ask (CB-205) — tell the primary how to answer in-turn.
case QUESTION -> text("[question] the worker paused to ask before it can finish:\n" + r.text()
+ "\n\nAnswer it by calling bridge_send again with turnId=\"" + r.turnId()
@@ -630,7 +657,7 @@ public final class BridgeMcp {
/** {@code bridge_spawn} without cwd/caller context (default resolution). */
static McpSchema.CallToolResult spawn(SessionManager sessions, String profile) {
return spawn(sessions, profile, null, null, null, null, null);
return spawn(sessions, profile, null, null, null, null, null, null, null);
}
/**
@@ -640,13 +667,18 @@ public final class BridgeMcp {
* {@code callerCwd} (the primary's directory), else the daemon's.
* CB-301: the session is registered with {@code ownerTerminal} as its owner.
* CB-301-ext: {@code worktreeRequest} non-null provisions an isolated git worktree.
* CB-584: {@code sessionName}/{@code resumeSessionId} request agent session identity — a resumed
* conversation requires an explicit {@code profile} whose adapter declares
* {@link dev.ltms.bridged.peer.Capability#SESSION_RESUME}, or the spawn is refused rather than
* silently starting a cold session.
*
* <p>{@code role} and {@code profile} are independent: the role picks the contract, the profile
* picks the backend. A reviewer on the same profile as the dev it reviews is a normal spawn.
*/
static McpSchema.CallToolResult spawn(SessionManager sessions, String profile, String role,
String requestedCwd, String callerCwd,
String ownerTerminal, WorktreeRequest worktreeRequest) {
String ownerTerminal, WorktreeRequest worktreeRequest,
String sessionName, String resumeSessionId) {
MemberRole memberRole;
try {
memberRole = isBlank(role) ? MemberRole.DEV : MemberRole.parse(role);
@@ -655,12 +687,13 @@ public final class BridgeMcp {
}
try {
MemberSession member = sessions.acquire(isBlank(profile) ? null : profile, memberRole,
requestedCwd, callerCwd, ownerTerminal, worktreeRequest);
requestedCwd, callerCwd, ownerTerminal, worktreeRequest,
isBlank(sessionName) ? null : sessionName, isBlank(resumeSessionId) ? null : resumeSessionId);
return text(json(memberView(member)));
} catch (GuardException e) {
return error("subscription boundary: " + e.getMessage());
} catch (IllegalArgumentException e) {
return error(e.getMessage()); // unknown / no-default profile
return error(e.getMessage()); // unknown / no-default profile, or a refused resumeSessionId
} catch (PeerUnreachableException e) {
return error("spawn timed out — worker pane never reached injectable state: " + e.getMessage());
} catch (HerdrException e) {
@@ -696,11 +729,33 @@ public final class BridgeMcp {
return null;
}
/** {@code bridge_profiles}: the configured worker profiles and the default. */
static McpSchema.CallToolResult profiles(PeerLauncher workers) {
return text(json(Map.of(
"profiles", workers.profiles(),
"default", workers.defaultProfile() == null ? "" : workers.defaultProfile())));
/**
* {@code bridge_profiles}: the configured worker profiles, the default, and — CB-578 stage B —
* which of them are currently quarantined (backend exhausted) and for how much longer. The
* {@code quarantined} key is present only when at least one profile is, so a fleet where nothing
* has ever been quarantined gets exactly the pre-stage-B shape.
*/
static McpSchema.CallToolResult profiles(PeerLauncher workers, QuarantineSource quarantine) {
Map<String, Object> result = new LinkedHashMap<>();
result.put("profiles", workers.profiles());
result.put("default", workers.defaultProfile() == null ? "" : workers.defaultProfile());
Map<String, Object> quarantined = new LinkedHashMap<>();
for (String profile : workers.profiles()) {
String credentialId = quarantine.credentialIdFor().apply(profile);
if (credentialId == null) {
continue;
}
quarantine.quarantine().remainingSeconds(credentialId).ifPresent(remaining -> {
Map<String, Object> row = new LinkedHashMap<>();
row.put("credentialId", credentialId);
row.put("quarantinedForSeconds", remaining);
quarantined.put(profile, row);
});
}
if (!quarantined.isEmpty()) {
result.put("quarantined", quarantined);
}
return text(json(result));
}
/**
@@ -719,17 +774,22 @@ public final class BridgeMcp {
* caller's own row is flagged {@code "self": true}: a peer needs to tell its own pane apart from
* a peer's, and the alternative is every lead calling {@code bridge_whoami} to subtract itself.
*
* <p>CB-583: the {@code capacity} rows reuse {@code quarantine} (the same {@link QuarantineSource}
* {@code bridge_profiles} reads) so the two surfaces cannot disagree about which profile is
* quarantined — see {@link #capacityView}.
*
* @param leads terminal_id → lead name, live from the resolver
* @param selfTerm the calling pane's terminal id, or blank for a caller with no pane
*/
static McpSchema.CallToolResult listFleet(PeerLauncher workers, SessionManager sessions,
Map<String, String> leads, String selfTerm) {
return listFleet(workers, sessions, null, CapacitySource.none(), leads, selfTerm);
return listFleet(workers, sessions, null, CapacitySource.none(), new HealthCoverageSource(() -> "off"),
QuarantineSource.none(), leads, selfTerm);
}
static McpSchema.CallToolResult listFleet(PeerLauncher workers, SessionManager sessions, MessageService messages,
CapacitySource capacity,
Map<String, String> leads, String selfTerm) {
CapacitySource capacity, HealthCoverageSource healthCoverage,
QuarantineSource quarantine, Map<String, String> leads, String selfTerm) {
try {
Map<String, Agent> live = workers.list().stream()
.map(Agent.class::cast)
@@ -747,9 +807,10 @@ public final class BridgeMcp {
roster.stream().map(MemberSession::profile).forEach(profiles::add);
Map<String, Object> result = new LinkedHashMap<>();
result.put("leads", leadRows); result.put("members", out);
result.put("healthCoverage", healthCoverage.value().get());
if (capacity.available()) result.put("capacity", profiles.stream()
.map(profile -> capacityView(profile, capacity.liveCount(), capacity.maxLoad(), roster, messages,
capacity.clock().getAsLong())).toList());
capacity.clock().getAsLong(), quarantine)).toList());
return text(json(result));
} catch (HerdrException e) {
return error("herdr error listing the fleet: " + e.getMessage());
@@ -774,9 +835,18 @@ public final class BridgeMcp {
return row;
}
/**
* CB-583: {@code free} alone cannot tell a lead "busy, will free up" from "refusing, and
* nothing changes for N seconds" — those need different decisions. So a quarantined profile
* forces {@code free} to 0, whatever its {@code maxLoad}/{@code live} say, and the row carries
* the same {@code credentialId}/{@code quarantinedForSeconds} facts {@code bridge_profiles}
* reports, reusing {@link QuarantineSource} rather than a second lookup. Both new keys are
* added only when the profile is actually quarantined, so an ordinary fleet's rows are
* byte-identical to before this change.
*/
private static Map<String, Object> capacityView(String profile, Function<String, Integer> liveCount,
Function<String, Integer> maxLoad, List<MemberSession> roster,
MessageService messages, long nowNanos) {
MessageService messages, long nowNanos, QuarantineSource quarantine) {
Integer cap = maxLoad.apply(profile);
int live = liveCount.apply(profile);
int reclaimable = (int) roster.stream().filter(s -> profile.equals(s.profile()))
@@ -786,6 +856,14 @@ public final class BridgeMcp {
Map<String, Object> row = new LinkedHashMap<>();
row.put("profile", profile); row.put("maxLoad", cap); row.put("live", live);
row.put("free", cap == null ? null : Math.max(0, cap - live)); row.put("reclaimable", reclaimable);
String credentialId = quarantine.credentialIdFor().apply(profile);
if (credentialId != null) {
quarantine.quarantine().remainingSeconds(credentialId).ifPresent(remaining -> {
row.put("free", 0);
row.put("credentialId", credentialId);
row.put("quarantinedForSeconds", remaining);
});
}
return row;
}
@@ -916,20 +994,30 @@ public final class BridgeMcp {
+ "for the default. The two are independent: a reviewer may run on the same profile "
+ "as the dev it reviews. The member opens your current directory by default; pass "
+ "cwd to pin a different one. Pass worktree:true (with ticket) or "
+ "worktree:<ticket-slug> to provision an isolated git worktree. Returns the member's "
+ "sessionId (use with bridge_send) and paneId (use with bridge_stop).",
+ "worktree:<ticket-slug> to provision an isolated git worktree. Pass resumeSessionId "
+ "to relaunch onto a prior conversation instead of starting cold — this requires an "
+ "explicit profile whose backend supports it (bridge_list shows agentSessionId for "
+ "resumable members), and is refused otherwise rather than silently starting fresh. "
+ "sessionName gives the member a display name in its own UI when the backend supports "
+ "one. Returns the member's sessionId (use with bridge_send) and paneId (use with "
+ "bridge_stop).",
objectSchema(Map.of(
"role", stringProp("What the member is for: architect, dev or reviewer (default dev)"),
"profile", stringProp("Which backend to run it on (omit for the default profile)"),
"cwd", stringProp("Working directory for the member (omit to inherit yours)"),
"worktree", Map.of("type", "string", "description", "'true' or a ticket slug — requests an isolated git worktree"),
"ticket", stringProp("Ticket slug when worktree:true")),
"ticket", stringProp("Ticket slug when worktree:true"),
"sessionName", stringProp("Logical display name for the member's own session, when its backend supports one"),
"resumeSessionId", stringProp("A prior member's agentSessionId (from bridge_list) to resume — requires an explicit profile that supports it")),
List.of()));
}
private static McpSchema.Tool profilesTool() {
return tool("bridge_profiles",
"List the configured worker profiles (backends) and which one bridge_spawn uses by default.",
"List the configured worker profiles (backends) and which one bridge_spawn uses by "
+ "default. A 'quarantined' map is present when a backend-exhausted refusal put "
+ "a profile's credential on cooldown — bridge_spawn onto it is refused until "
+ "quarantinedForSeconds elapses; a profile sharing that credential is listed too.",
objectSchema(Map.of(), List.of()));
}
@@ -941,8 +1029,15 @@ public final class BridgeMcp {
+ "discover a peer lead without being told its address. 'members' are the "
+ "sessions delegated to — each with sessionId, paneId, role (architect/dev/"
+ "reviewer), profile (the backend it runs on), state, optional "
+ "worktree/branch/owner, and live herdr status. An empty 'members' "
+ "means no members are spawned; it says nothing about peers.",
+ "worktree/branch/owner/agentSessionId (the id to pass as bridge_spawn's "
+ "resumeSessionId to relaunch onto that same conversation, when the backend "
+ "supports it), and live herdr status. An empty 'members' "
+ "means no members are spawned; it says nothing about peers. When capacity "
+ "facts are configured, a 'capacity' row per profile also reports free: 0 for "
+ "a quarantined profile's credential (see bridge_profiles), whatever its "
+ "maxLoad/live — with credentialId and quarantinedForSeconds naming the "
+ "quarantine, so 'free: 0, busy' can be told apart from 'free: 0, refusing "
+ "for N seconds'.",
objectSchema(Map.of(), List.of()));
}
@@ -8,6 +8,7 @@ import dev.ltms.bridged.peer.PeerHandle;
import dev.ltms.bridged.peer.PeerLauncher;
import dev.ltms.bridged.peer.PeerUnreachableException;
import dev.ltms.bridged.peer.SpawnRequest;
import dev.ltms.bridged.placement.BackendQuarantine;
import dev.ltms.bridged.placement.PlacementCandidate;
import dev.ltms.bridged.placement.PlacementContext;
import dev.ltms.bridged.placement.PlacementException;
@@ -23,10 +24,12 @@ import java.util.HashSet;
import java.util.LinkedHashMap;
import java.util.List;
import java.util.Map;
import java.util.Objects;
import java.util.Set;
import java.util.concurrent.ConcurrentHashMap;
import java.util.function.Function;
import java.util.function.Supplier;
import java.util.stream.Collectors;
/**
* The {@link PeerLauncher} the core actually holds when more than one adapter is configured — a thin
@@ -78,6 +81,9 @@ public final class CompositePeerLauncher implements PeerLauncher {
private final Supplier<Map<String, BridgedConfig.Profile>> profileConfigs;
private final Supplier<PlacementPolicy> placementPolicy;
/** CB-578 stage B: credential cooldown, checked before an explicit spawn and filtered into placement. */
private final BackendQuarantine quarantine;
/**
* CB-557: the role pools an unqualified spawn draws its candidates from. A supplier that yields
* {@code null}, and an empty pool for a role, both fall back to every configured profile — the
@@ -99,7 +105,10 @@ public final class CompositePeerLauncher implements PeerLauncher {
}
/**
* Production constructor with a placement policy and live-worker counter.
* Production constructor with a placement policy and live-worker counter. Quarantine (CB-578
* stage B) is off for this constructor — {@link BackendQuarantine#none()} — since it predates
* the feature and existing callers of this exact overload never exercised it; use the 7-arg
* overload below to wire a real {@link BackendQuarantine}.
*
* @param delegates one adapter per configured peer kind; must be non-empty and declare
* disjoint profile-name sets
@@ -113,13 +122,14 @@ public final class CompositePeerLauncher implements PeerLauncher {
Map<String, BridgedConfig.Profile> profileConfigs,
PlacementPolicy placementPolicy,
Function<String, Integer> liveCount) {
this(delegates, defaultProfile, profileConfigs, placementPolicy, liveCount, null);
this(delegates, defaultProfile, profileConfigs, placementPolicy, liveCount, null, BackendQuarantine.none());
}
/**
* Production constructor with role pools (CB-557). An unqualified spawn draws its candidates from
* {@code fleet.<role>} instead of from every configured profile, so a reviewer is placed on a
* reviewer backend and never on, say, the architect-only one.
* reviewer backend and never on, say, the architect-only one. Quarantine is off for this
* constructor too, for the same reason as the 5-arg overload above.
*
* @param fleet the configured role pools; {@code null} ⇒ every profile is a candidate for every
* role, which is the pre-CB-557 behaviour
@@ -130,29 +140,50 @@ public final class CompositePeerLauncher implements PeerLauncher {
PlacementPolicy placementPolicy,
Function<String, Integer> liveCount,
BridgedConfig.Fleet fleet) {
this(delegates, defaultProfile, profileConfigs, placementPolicy, liveCount, fleet, BackendQuarantine.none());
}
/**
* Production constructor with role pools and quarantine (CB-578 stage B). The full-featured
* non-reloading form; {@link #CompositePeerLauncher(List, String, Supplier, Function, BackendQuarantine)}
* is what {@code Bridged.main} actually wires up.
*
* @param quarantine required — pass {@link BackendQuarantine#none()} for a caller that does not
* want the feature, never a defaulting overload (CB-578 stage B's own rule).
*/
public CompositePeerLauncher(List<HerdrPeerLauncher> delegates,
String defaultProfile,
Map<String, BridgedConfig.Profile> profileConfigs,
PlacementPolicy placementPolicy,
Function<String, Integer> liveCount,
BridgedConfig.Fleet fleet,
BackendQuarantine quarantine) {
// LinkedHashMap, not Map.copyOf: candidates() promises definition order and the weighted
// policy breaks exact-weight ties on it, so a salted iteration order would make placement
// differ from one JVM run to the next.
this(delegates, defaultProfile,
constant(Collections.unmodifiableMap(new LinkedHashMap<>(profileConfigs))),
constant(placementPolicy), liveCount, constant(fleet));
constant(placementPolicy), liveCount, constant(fleet), quarantine);
}
/**
* Production constructor that re-reads its placement inputs per spawn (CB-559), so a config
* reload retargets the next member without a restart.
*
* @param config the live configuration — read at every spawn, never captured
* @param config the live configuration — read at every spawn, never captured
* @param quarantine required — CB-578 stage B; pass {@link BackendQuarantine#none()} to opt out
*/
public CompositePeerLauncher(List<HerdrPeerLauncher> delegates,
String defaultProfile,
Supplier<BridgedConfig> config,
Function<String, Integer> liveCount) {
Function<String, Integer> liveCount,
BackendQuarantine quarantine) {
this(delegates, defaultProfile,
() -> config.get().profiles(),
() -> PlacementPolicies.fromName(config.get().placement()),
liveCount,
() -> config.get().fleet());
() -> config.get().fleet(),
quarantine);
}
/** The all-suppliers form every other constructor funnels into. */
@@ -161,8 +192,10 @@ public final class CompositePeerLauncher implements PeerLauncher {
Supplier<Map<String, BridgedConfig.Profile>> profileConfigs,
Supplier<PlacementPolicy> placementPolicy,
Function<String, Integer> liveCount,
Supplier<BridgedConfig.Fleet> fleet) {
Supplier<BridgedConfig.Fleet> fleet,
BackendQuarantine quarantine) {
this.fleet = fleet;
this.quarantine = Objects.requireNonNull(quarantine, "quarantine");
if (delegates.isEmpty()) {
throw new IllegalArgumentException("at least one peer adapter must be configured");
}
@@ -225,6 +258,7 @@ public final class CompositePeerLauncher implements PeerLauncher {
// the charter makes explicit-profile spawns the normal path — so skipping the check
// here would leave the cap dead config in real operation.
HerdrPeerLauncher d = route(requestedProfile);
enforceNotQuarantined(requestedProfile);
enforceMaxLoad(requestedProfile);
PeerHandle handle = d.spawn(req);
spawnedBy.put(handle.id(), d);
@@ -238,7 +272,10 @@ public final class CompositePeerLauncher implements PeerLauncher {
List<PlacementCandidate> candidates = candidates(req.role());
String roleDefault = defaultProfileFor(req.role());
Set<String> unreachable = new HashSet<>();
PlacementContext ctx = new PlacementContext(roleDefault, candidates, liveCount, unreachable);
// CB-578 stage B: computed once up front — a quarantine's expiry cannot pass within one spawn
// call, so re-deriving it per retry would only cost work, never change the answer.
Set<String> quarantined = quarantinedProfiles(candidates);
PlacementContext ctx = new PlacementContext(roleDefault, candidates, liveCount, unreachable, quarantined);
int maxAttempts = candidates.isEmpty() ? 1 : candidates.size();
for (int attempt = 0; attempt < maxAttempts; attempt++) {
@@ -268,7 +305,7 @@ public final class CompositePeerLauncher implements PeerLauncher {
chosen.profile(), e.getMessage());
unreachable.add(chosen.profile());
// Update the context for the next selection so the policy excludes this profile.
ctx = new PlacementContext(roleDefault, candidates, liveCount, unreachable);
ctx = new PlacementContext(roleDefault, candidates, liveCount, unreachable, quarantined);
}
}
@@ -298,9 +335,48 @@ public final class CompositePeerLauncher implements PeerLauncher {
* @param profile the profile the caller explicitly named
* @throws PlacementException when the profile is at capacity
*/
/**
* Refuse an explicit-profile spawn whose credential is quarantined (CB-578 stage B): a prior
* {@code BACKEND_EXHAUSTED} classification on this profile, or on another profile sharing its
* {@code credentialId}, is still on cooldown.
*
* <p>Checked before {@link #enforceMaxLoad}, and for the same reason that check exists: an
* explicit profile bypasses the placement policy's filtering entirely, so without this the cap
* (there, quarantine here) would be dead config on the very path the charter calls normal.
* Deliberately no fallback to another profile, matching {@link #enforceMaxLoad}'s own reasoning —
* the caller named this profile for a cost/model reason.
*
* @throws PlacementException naming the profile, its credential, and the remaining cooldown
*/
private void enforceNotQuarantined(String profile) {
String credentialId = credentialIdFor(profile);
quarantine.remainingSeconds(credentialId).ifPresent(remaining -> {
throw new PlacementException("worker profile '" + profile + "' is quarantined "
+ "(credential '" + credentialId + "' exhausted; ~" + remaining
+ "s remaining) — refusing spawn");
});
}
/** {@code profile}'s credential group (CB-578 stage B), or the profile's own name if unconfigured. */
private String credentialIdFor(String profile) {
BridgedConfig.Profile cfg = profiles0().get(profile);
return cfg == null ? profile : cfg.effectiveCredentialId();
}
/** The subset of {@code candidates} whose credential is currently quarantined (CB-578 stage B). */
private Set<String> quarantinedProfiles(List<PlacementCandidate> candidates) {
return candidates.stream()
.map(PlacementCandidate::profile)
.filter(p -> quarantine.isQuarantined(credentialIdFor(p)))
.collect(Collectors.toSet());
}
private void enforceMaxLoad(String profile) {
// Absent config, or a config whose maxLoad normalized to null (non-positive ⇒ unlimited at
// load), means no cap — never cap what wasn't configured.
// Absent config, or a config whose maxLoad normalized to null (ABSENT ⇒ unlimited at load),
// means no cap — never cap what wasn't configured. Note "non-positive ⇒ unlimited" was true
// until CB-585: an explicit `maxLoad: 0` now survives as 0 and is a real cap of zero, so the
// check below refuses every spawn on that profile, and a negative value is refused at config
// load rather than normalized away.
BridgedConfig.Profile cfg = profiles0().get(profile);
Integer cap = (cfg == null) ? null : cfg.maxLoad();
if (cap == null) {
@@ -387,6 +463,18 @@ public final class CompositePeerLauncher implements PeerLauncher {
return defaultProfile;
}
/**
* {@inheritDoc}
*
* <p>Routes to the specific delegate {@code profileName} resolves to, not the fleet-wide
* union {@link #capabilities()} returns — the whole reason this method exists (CB-584): in a
* mixed fleet, one adapter's capability must never be read as every profile's.
*/
@Override
public Set<Capability> capabilitiesFor(String profileName) {
return route(profileName).capabilities();
}
/** Every herdr agent, deduplicated by pane id (all delegates share one herdr and list globally). */
@Override
public List<Agent> list() {
@@ -7,6 +7,8 @@ import dev.ltms.bridged.herdr.HerdrException;
import dev.ltms.bridged.herdr.Tab;
import dev.ltms.bridged.herdr.Workspace;
import dev.ltms.bridged.herdr.WorkspaceControl;
import dev.ltms.bridged.peer.Capability;
import dev.ltms.bridged.peer.CharterReceipt;
import dev.ltms.bridged.peer.MemberRole;
import dev.ltms.bridged.peer.PeerHandle;
import dev.ltms.bridged.peer.PeerLauncher;
@@ -247,6 +249,19 @@ public abstract class HerdrPeerLauncher implements PeerLauncher {
return defaultProfile;
}
/**
* {@inheritDoc}
*
* <p>One {@link HerdrPeerLauncher} instance always serves exactly one adapter kind, so every
* profile it owns shares that adapter's {@link #capabilities()} — {@code profileName} only
* needs validating (throwing on an unknown profile, same as {@link #spawn}), not routing.
*/
@Override
public Set<Capability> capabilitiesFor(String profileName) {
requireProfile(profileName);
return capabilities();
}
/** The configured profiles, for adapter capability decisions (e.g. any git-token grant). */
protected Collection<BridgedConfig.Profile> profileConfigs() {
return profiles.values();
@@ -269,8 +284,11 @@ public abstract class HerdrPeerLauncher implements PeerLauncher {
// --- spawn ---------------------------------------------------------------------------------
/** A started peer plus the launch's agent-session id (the resume handle, or null). */
private record Spawned(Agent agent, String agentSessionId) {
/**
* A started peer plus the launch's agent-session id (the resume handle, or null) and the
* charter receipt (CB-571) the base composed for it.
*/
private record Spawned(Agent agent, String agentSessionId, CharterReceipt receipt) {
}
/**
@@ -305,12 +323,41 @@ public abstract class HerdrPeerLauncher implements PeerLauncher {
String replyCharter = cfg.hasMcp() ? REPLY_CHARTER : null;
String charter = roleCharter == null ? replyCharter
: replyCharter == null ? roleCharter : roleCharter + "\n\n" + replyCharter;
Launch launch = buildLaunch(cfg, new LaunchSpec(sessionName, resumeSessionId, role, charter));
String cwd = resolveCwd(requestedCwd, cfg, callerCwd);
Agent agent = cfg.tabPlacement()
? spawnInTab(cfg, launch.env(), launch.argv(), cwd, role, liveFleet)
: spawnAsPane(cfg, launch.env(), launch.argv(), cwd);
return new Spawned(agent, launch.agentSessionId());
// CB-571: fingerprint the exact composed charter bytes once, here in the base, before the
// string leaves for an adapter — so Claude and OpenCode derive the same digest. A failed
// start has no bridge_spawn result and no roster row, so the failure log below is the only
// surface the byte count can appear on. The charter text itself is never logged.
CharterReceipt receipt = CharterReceipt.compose(role, cfg.profile(), roleCharter, charter);
try {
Launch launch = buildLaunch(cfg, new LaunchSpec(sessionName, resumeSessionId, role, charter));
String cwd = resolveCwd(requestedCwd, cfg, callerCwd);
Agent agent = cfg.tabPlacement()
? spawnInTab(cfg, launch.env(), launch.argv(), cwd, role, liveFleet)
: spawnAsPane(cfg, launch.env(), launch.argv(), cwd, charter);
logCharterReceipt(receipt, true);
return new Spawned(agent, launch.agentSessionId(), receipt);
} catch (RuntimeException e) {
logCharterReceipt(receipt, false);
throw e;
}
}
/**
* The one place the charter's size and digest appear in the logs. {@code success} true after a
* start, false from the failure path of {@link #spawnInternal} where no handle or roster row
* exists to carry the receipt. Always metadata only — never the charter text.
*/
private static void logCharterReceipt(CharterReceipt receipt, boolean success) {
String role = receipt.role() == null ? "" : receipt.role().wireName();
if (success) {
log.info("spawned role={} profile={} charterSource={} charterSha256={} charterBytes={}",
role, receipt.profile(), receipt.charterSource(),
receipt.charterSha256(), receipt.charterBytes());
} else {
log.warn("spawn failed; charter role={} profile={} charterSource={} charterSha256={} charterBytes={}",
role, receipt.profile(), receipt.charterSource(),
receipt.charterSha256(), receipt.charterBytes());
}
}
/**
@@ -349,7 +396,7 @@ public abstract class HerdrPeerLauncher implements PeerLauncher {
String id = UUID.randomUUID().toString();
paneByAgentId.put(id, paneId);
return new WorkerHandle(id, agent.terminalId(), requireProfile(req.profileName()).profile(),
req.sessionName(), spawned.agentSessionId());
req.sessionName(), spawned.agentSessionId(), spawned.receipt());
}
@Override
@@ -433,11 +480,19 @@ public abstract class HerdrPeerLauncher implements PeerLauncher {
}
}
/** Legacy placement: split the currently-focused tab; the peer still starts in {@code cwd}. */
/**
* Legacy placement: split the currently-focused tab; the peer still starts in {@code cwd}.
*
* <p>CB-571: this is the one legacy log that printed the full argv, and the charter travels
* inside argv — so the charter text went to the daemon log on every pane-placement spawn. The
* {@code spawnInTab} path never logs argv, so only this site is fixed. {@code charter} is the
* composed charter, if any; its argv element is replaced by its digest so the log still shows
* which args were passed without exposing the charter prose.
*/
private Agent spawnAsPane(BridgedConfig.Profile cfg, Map<String, String> workerEnv,
List<String> argv, String cwd) {
List<String> argv, String cwd, String charter) {
log.info("spawning {} (pane placement) profile={} cwd={} argv={}",
namePrefix, cfg.profile(), cwd, argv);
namePrefix, cfg.profile(), cwd, redactCharter(argv, charter));
String paneId = spaces.splitPane(cwd, workerEnv);
if (paneId == null) {
throw new IllegalStateException("pane.split returned no pane — cannot start a peer");
@@ -447,6 +502,21 @@ public abstract class HerdrPeerLauncher implements PeerLauncher {
return peer;
}
/**
* A copy of {@code argv} with an element equal to {@code charter} replaced by its digest, so
* the pane log never prints the charter prose. The charter is handed to an adapter as one argv
* element, so exact-equality is the right match; every other argument passes through unchanged.
*/
private static List<String> redactCharter(List<String> argv, String charter) {
if (charter == null || charter.isBlank() || argv == null || argv.isEmpty()) {
return argv;
}
String digest = CharterReceipt.digestOf(charter);
return argv.stream()
.map(a -> a.equals(charter) ? "<charter sha256=" + digest + ">" : a)
.toList();
}
/** A started peer together with the sequence its unique name/label used. */
private record Started(Agent agent, long seq) {
}
@@ -651,11 +721,18 @@ public abstract class HerdrPeerLauncher implements PeerLauncher {
/**
* A concrete {@link PeerHandle} wrapping herdr agent coordinates, the profile that spawned it,
* and the session identity the launch resolved (CB-547a): the bridge's logical name and the
* peer's own session id, both null when the spawn carried no identity.
* the session identity the launch resolved (CB-547a): the bridge's logical name and the peer's
* own session id, both null when the spawn carried no identity — and the charter receipt
* (CB-571) the base computed for this launch.
*/
private record WorkerHandle(String id, String terminalId, String profile,
String sessionName, String agentSessionId) implements PeerHandle {
String sessionName, String agentSessionId,
CharterReceipt receipt) implements PeerHandle {
@Override
public CharterReceipt charterReceipt() {
return receipt;
}
}
// --- shared helpers ------------------------------------------------------------------------
@@ -689,10 +766,61 @@ public abstract class HerdrPeerLauncher implements PeerLauncher {
}
}
/** A fresh mutable env map — the conventional starting point for {@link #buildLaunch}. */
/**
* CB-592: overlay value that shadows the admin {@code GITEA_ACCESS_TOKEN} a herdr pane
* otherwise inherits from herdr's own login-shell process environment (gitea issue #77).
* herdr spawns a pane from its <em>own</em> process environment and layers our map on top —
* {@link dev.ltms.bridged.herdr.WorkspaceControl#createTab} and {@code #splitPane} send only
* the keys we put in that map, so any key we never mention passes straight through from
* herdr's own shell, admin token included.
*
* <p>Deliberately a non-blank sentinel, not {@code ""}. Whether an empty-string overlay value
* overrides an inherited variable or is skipped as blank could not be settled by reading this
* codebase — herdr's server-side merge is an external process, not something in this repo.
* A non-blank replacement sidesteps that ambiguity entirely: {@link #baseEnv}'s own {@code
* PATH} seeding already depends on the overlay reliably replacing an inherited value (see its
* javadoc), and that is only demonstrated for a non-blank value, so this reuses the same,
* proven-reliable shape rather than the unverified one.
*
* <p><b>MEASURED ON A LIVE PANE, 2026-08-15: this sentinel alone does NOT hold.</b> The overlay
* itself works — {@code GITEA_TOKEN} is injected here, is exported by no shell file, and does
* reach the pane. The sentinel loses one step later. A herdr pane runs a <em>login</em> shell,
* {@code ~/.zprofile} sources {@code ${SHARED_ENV}/tools/secrets.sh}, and that file does a plain
* unconditional {@code export GITEA_ACCESS_TOKEN=...}. A login shell overwrites a value already
* in the environment, so the real admin token is put back over this sentinel before the member
* process ever starts. That defeat applies to <em>every</em> name {@code secrets.sh} exports,
* and no launcher-side overlay can win against it.
*
* <p>So this constant is not the control on its own — {@link #MEMBER_MARKER} is the other half.
* Keeping the sentinel is still worth it: it is correct for any peer kind whose pane does not
* start a login shell, and it makes the intent explicit at the one place every adapter passes.
*/
private static final String BLOCKED_GITEA_ACCESS_TOKEN =
"blocked-by-bridged-cb592-see-gitea-issue-77";
/**
* CB-592: marks a pane as a bridged member so a shell startup file can decline to export
* operator-only credentials into it (gitea issue #77).
*
* <p>This name is deliberately one that {@code secrets.sh} never exports, which is exactly why
* it survives the login shell that wipes {@link #BLOCKED_GITEA_ACCESS_TOKEN}. The mechanism is
* measured, not assumed: {@code GITEA_TOKEN} is injected the same way, is absent from a login
* shell of its own, and was observed set inside a live member pane.
*
* <p>It is a no-op until the operator guards the export, which is a one-line change in a file
* this repo does not own and must not edit unasked:
*
* <pre>{@code
* [ -n "${BRIDGED_MEMBER:-}" ] || export GITEA_ACCESS_TOKEN=...
* }</pre>
*
* <p>Setting the marker now costs nothing and means that edit is the whole remaining fix.
*/
static final String MEMBER_MARKER = "BRIDGED_MEMBER";
/**
* Seed a worker's environment (CB-511): the daemon's own {@code PATH}, then the profile's
* {@code env:} entries.
* {@code env:} entries, then the CB-592 admin-token shadow.
*
* <p>Why this exists: bridged passes herdr an explicit env map, and herdr merges it into
* <em>its own</em> process environment. So before this, a worker inherited whatever PATH the
@@ -706,6 +834,12 @@ public abstract class HerdrPeerLauncher implements PeerLauncher {
* overriding {@code ANTHROPIC_BASE_URL} and slipping past {@link
* dev.ltms.bridged.guard.SubscriptionGuard}, which is checked against the profile's
* {@code baseUrl} and nothing else.
*
* <p>The CB-592 shadow and marker are put in <em>last</em>, after the profile's own
* {@code env:}, so no profile — present or future — can restore the admin token, or hide that
* the pane is a member, by naming either in config. This is the one place both are applied:
* every {@code buildLaunch} in every adapter calls this first, so a new profile, and a peer
* kind not yet written, gets them for free.
*/
protected Map<String, String> baseEnv(BridgedConfig.Profile cfg) {
Map<String, String> workerEnv = new LinkedHashMap<>();
@@ -716,6 +850,8 @@ public abstract class HerdrPeerLauncher implements PeerLauncher {
if (cfg != null && cfg.env() != null) {
workerEnv.putAll(cfg.env());
}
workerEnv.put("GITEA_ACCESS_TOKEN", BLOCKED_GITEA_ACCESS_TOKEN);
workerEnv.put(MEMBER_MARKER, "1");
return workerEnv;
}
@@ -7,6 +7,7 @@ import dev.ltms.bridged.herdr.Agent;
import dev.ltms.bridged.herdr.AgentControl;
import dev.ltms.bridged.herdr.WorkspaceControl;
import dev.ltms.bridged.peer.Capability;
import dev.ltms.bridged.peer.CharterReceipt;
import dev.ltms.bridged.peer.PeerHandle;
import dev.ltms.bridged.peer.SpawnRequest;
@@ -406,6 +407,11 @@ public final class OpenCodeLauncher extends HerdrPeerLauncher {
// appeared).
return discovery.sessionIdForDirectory(cwd);
}
@Override
public CharterReceipt charterReceipt() {
return delegate.charterReceipt();
}
}
// --- Agent-returning convenience spawns (used by callers/tests that want the herdr Agent) ---
@@ -20,6 +20,7 @@ import java.util.concurrent.TimeUnit;
import java.util.concurrent.TimeoutException;
import java.util.concurrent.atomic.AtomicLong;
import java.util.concurrent.locks.ReentrantLock;
import java.util.function.LongSupplier;
/**
* The blocking delegation feature (CB-104): deliver {@code content} into a worker and block until
@@ -52,8 +53,12 @@ public final class MessageService {
*/
private static final long ASYNC_TIMEOUT_MS = 30 * 60 * 1_000L;
/** How long a finished (terminal) ticket is retained for polling before it is pruned. */
private static final long TICKET_TTL_NANOS = 10 * 60 * 1_000_000_000L;
/**
* How long a finished (terminal) ticket is retained for polling before it is pruned. Package-
* private (not {@code private}) so a test can advance an injected clock past it deterministically
* instead of duplicating the magic number or sleeping for real.
*/
static final long TICKET_TTL_NANOS = 10 * 60 * 1_000_000_000L;
/** Outcome of a blocking send. */
public enum Outcome {
@@ -69,6 +74,14 @@ public final class MessageService {
* failure context (e.g. the error screen). Terminal, but not a successful completion.
*/
WORKER_FAILED,
/**
* The turn finished without a {@code bridge_reply} and the scrape matched the backend's
* configured usage-limit refusal pattern (CB-578 stage A); {@code text} is the reason,
* carrying the matched line. The worker's pane is healthy — only its account is refusing —
* so this is never reported as a completed reply, and is kept distinct from
* {@link #WORKER_FAILED} (a wedged worker) and a session simply going {@code GONE}.
*/
BACKEND_EXHAUSTED,
/**
* The worker paused mid-turn to ask the primary a question (CB-205); {@code text} is the
* question and {@code turnId} correlates the answer. Not terminal — the primary answers with
@@ -153,12 +166,13 @@ public final class MessageService {
private static final class Task {
private final String target;
private final CompletableFuture<Reply> future = new CompletableFuture<>();
private final long createdNanos = System.nanoTime();
private final long createdNanos;
private volatile Reply question;
private volatile String turnId;
private Task(String target) {
private Task(String target, long createdNanos) {
this.target = target;
this.createdNanos = createdNanos;
}
}
@@ -168,10 +182,14 @@ public final class MessageService {
private final ReplyInbox inbox;
private final ReplyPushLoop pushLoop;
private final Metrics metrics; // CB-502: nullable — no registry in unit tests
// CB-588: injectable so pruneTerminalTickets' 10-minute TICKET_TTL_NANOS can be exercised in a
// test without a real wait — same seam SessionManager already uses for its idle reaper (nowNanos).
private final LongSupplier nowNanos;
private final ConcurrentHashMap<String, ReentrantLock> sessionLocks = new ConcurrentHashMap<>();
private final ConcurrentHashMap<String, Task> tasks = new ConcurrentHashMap<>();
/** The async task that currently owns a target's send lock. */
private final ConcurrentHashMap<String, Task> asyncTasksByTarget = new ConcurrentHashMap<>();
/** Async task that owns each exact forward rendezvous waiter. */
private final ConcurrentHashMap<CompletableFuture<Rendezvous.Resolution>, Task> asyncTasksByWaiter =
new ConcurrentHashMap<>();
/** Async tickets paused on a specific {@code bridge_ask} turn. */
private final ConcurrentHashMap<String, Task> asyncTasksByTurn = new ConcurrentHashMap<>();
private final AtomicLong ticketSeq = new AtomicLong();
@@ -182,7 +200,9 @@ public final class MessageService {
* Create with an explicit {@link ReplyInbox} and optional {@link ReplyPushLoop}.
*
* @param pushLoop nullable — when non-null, the push loop is notified on the no-waiter reply
* branch ({@link #reply}) so it can nudge the primary to drain the inbox
* branch ({@link #reply}) so it can nudge the primary to drain the inbox, and
* (CB-588) whenever an async ticket started by {@link #sendAsync} reaches a
* terminal phase, and whenever {@link #poll} hands a terminal ticket to its caller
*/
public MessageService(AgentControl agents, Injector injector, Rendezvous rendezvous,
ReplyInbox inbox, ReplyPushLoop pushLoop) {
@@ -197,12 +217,19 @@ public final class MessageService {
*/
public MessageService(AgentControl agents, Injector injector, Rendezvous rendezvous,
ReplyInbox inbox, ReplyPushLoop pushLoop, Metrics metrics) {
this(agents, injector, rendezvous, inbox, pushLoop, metrics, System::nanoTime);
}
/** Test constructor with an injectable clock (CB-588: exercise the ticket-prune TTL without a real wait). */
MessageService(AgentControl agents, Injector injector, Rendezvous rendezvous, ReplyInbox inbox,
ReplyPushLoop pushLoop, Metrics metrics, LongSupplier nowNanos) {
this.agents = agents;
this.injector = injector;
this.rendezvous = rendezvous;
this.inbox = inbox;
this.pushLoop = pushLoop;
this.metrics = metrics;
this.nowNanos = nowNanos;
}
/** Create with an explicit {@link ReplyInbox} and no push loop. */
@@ -279,6 +306,7 @@ public final class MessageService {
case COMPLETED_UNREPLIED -> "completion_fallback";
case TIMED_OUT_WORKING, TIMED_OUT_QUEUED, BUSY -> "timeout";
case WORKER_FAILED -> "failed";
case BACKEND_EXHAUSTED -> "backend_exhausted";
case STALE_TURN, QUESTION -> null; // not a completed delegation
};
}
@@ -300,14 +328,18 @@ public final class MessageService {
*/
public boolean abandon(String target, String reason) {
CompletableFuture<Rendezvous.Resolution> waiter = rendezvous.currentWaiter(target);
if (waiter == null || waiter.isDone()) {
return false; // nobody is blocked on this worker — nothing to abandon
boolean failed = waiter != null && !waiter.isDone() && rendezvous.resolveFailure(waiter, reason);
boolean asyncFailed = false;
for (Task task : tasks.values()) {
if (target.equals(task.target) && task.question == null
&& task.future.complete(new Reply(Outcome.WORKER_FAILED, reason))) {
asyncFailed = true;
}
}
boolean failed = rendezvous.resolveFailure(waiter, reason);
if (failed) {
log.warn("abandoning the blocked send to {}: {}", target, reason);
}
return failed;
return failed || asyncFailed;
}
/**
@@ -319,9 +351,30 @@ public final class MessageService {
}
/**
* Drain (peek + ack) all pending inbox replies for {@code target}. At-least-once: returns the
* messages and acknowledges them; an in-flight failure between returning and the caller
* processing them re-surfaces them on a subsequent drain (the ack is local).
* Drain (peek + ack) all pending inbox replies for {@code target}.
*
* <p><strong>The ack happens here, before the caller has the messages</strong> — before the MCP
* or REST response carrying them has been written, and long before the client has processed
* them. That ordering is what the two adapters disagree about, so do not read this method as
* "at-least-once" without qualifying which inbox is behind it (CB-529):
*
* <ul>
* <li>{@code InMemoryReplyInbox} — the ack only drops an entry from a local map. The messages
* are already in the returned list, so nothing can be lost after this point.
* <li>{@code AmqpReplyInbox} — the ack is a broker-side {@code basicAck}. Once it lands the
* broker has forgotten the message. If the daemon dies while writing the response, the
* reply is gone from the broker <em>and</em> the client never received it. Re-polling
* cannot recover it, because there is nothing left to re-deliver.
* </ul>
*
* <p>So the loss window is the response write, and it is a genuine loss rather than a
* redelivery. This is accepted, not overlooked: the alternative — ack on the next poll — turns
* every normal drain into a double delivery, which costs more than the window it closes. A
* caller that needs certainty re-polls; that is idempotent for every case except this one.
*
* <p>Any change here must be checked against <em>both</em> adapters. The previous version of
* this javadoc claimed "the ack is local", which was true when only the in-memory inbox existed
* and silently became false when the AMQP adapter landed.
*
* @return the drained messages, newest last (FIFO); empty list if none
*/
@@ -354,6 +407,11 @@ public final class MessageService {
* never earned. {@code null} disables the hook.
*/
public Reply send(String target, String content, long timeoutMillis, Runnable onAccepted) {
return send(target, content, timeoutMillis, onAccepted, null);
}
/** Run a send, optionally stopping an async task that teardown already failed before acceptance. */
private Reply send(String target, String content, long timeoutMillis, Runnable onAccepted, Task task) {
long deadlineNanos = System.nanoTime() + timeoutMillis * 1_000_000L;
ReentrantLock lock = sessionLocks.computeIfAbsent(target, _ -> new ReentrantLock());
@@ -361,6 +419,9 @@ public final class MessageService {
return new Reply(Outcome.BUSY, null); // another send held the session the whole window
}
try {
if (task != null && task.future.isDone()) {
return task.future.getNow(null);
}
if (hasAsyncQuestion(target)) {
return new Reply(Outcome.BUSY, null); // the worker's current turn is paused for its lead
}
@@ -372,13 +433,17 @@ public final class MessageService {
// failed send leaves no stale waiter behind.
CompletableFuture<Rendezvous.Resolution> reply = rendezvous.open(target);
try {
if (task != null) {
asyncTasksByWaiter.put(reply, task);
}
TurnToken token = new TurnToken(target, reply);
// The send has won the lock; the accepted-delivery hook records delegator ownership
// here (CB-548). It runs BEFORE enqueue so a throwing hook — onAccepted is now a
// public callback — fails the send without queuing a message that would orphan.
if (onAccepted != null) {
onAccepted.run();
}
CompletableFuture<Void> delivered = injector.enqueue(target, content);
CompletableFuture<Void> delivered = injector.enqueue(target, content, token);
try {
Rendezvous.Resolution r = reply.get(remainingMillis(deadlineNanos), TimeUnit.MILLISECONDS);
return recorded(new Reply(outcomeOf(r.kind()), r.text(), r.turnId()));
@@ -395,6 +460,7 @@ public final class MessageService {
throw new IllegalStateException("interrupted awaiting reply from " + target, e);
}
} finally {
asyncTasksByWaiter.remove(reply);
rendezvous.close(target, reply);
}
} finally {
@@ -420,11 +486,15 @@ public final class MessageService {
if (ticket.fresh()) {
// Register the reverse waiter first, then surface the question — so the answer, which can
// arrive the instant the primary reacts, always finds an open waiter to resolve.
CompletableFuture<Rendezvous.Resolution> waiter = rendezvous.currentWaiter(workerSession);
Task task = markAsyncQuestion(waiter, question, ticket.turnId());
if (!rendezvous.resolveQuestion(workerSession, question, ticket.turnId())) {
if (task != null) {
clearAsyncQuestion(ticket.turnId(), true);
}
rendezvous.closeAsk(ticket.turnId());
return new AskResult(AskOutcome.NO_WAITER, null); // no primary is blocked on this worker
}
markAsyncQuestion(workerSession, question, ticket.turnId());
}
try {
String answer = ticket.answer().get(timeoutMillis, TimeUnit.MILLISECONDS);
@@ -519,25 +589,35 @@ public final class MessageService {
*/
public String sendAsync(String target, String content, Runnable onAccepted) {
String ticket = "task-" + ticketSeq.incrementAndGet();
Task task = new Task(target);
Task task = new Task(target, nowNanos.getAsLong());
tasks.put(ticket, task);
if (pushLoop != null) {
// CB-588: task.future only ever completes on a terminal phase (DONE or a failure) — a
// worker paused in bridge_ask leaves it running, per finishAsyncTask's own contract — so
// this fires exactly once, from whichever path completes it: finishAsyncTask(task, result)
// below on any non-QUESTION outcome of send() — a worker's bridge_reply, the CB-106
// completion fallback, a CB-109 wedge, TIMED_OUT, BUSY, or BACKEND_EXHAUSTED — the same
// finishAsyncTask reached via answer()'s finishAsyncTask(turnId, result) once a QUESTION
// is resolved, completeExceptionally(t) just below when send() itself throws, or a CB-516
// abandon() on teardown. Without this, MessageService.reply's rendezvous fast path (the
// one an async ticket always takes) never told the push loop anything happened — see the
// class javadoc on sendAsync/CB-107.
task.future.whenComplete((reply, ex) -> {
boolean failed = ex != null || reply == null || !reply.completed();
pushLoop.onTicketTerminal(ticket, target, failed);
});
}
asyncExecutor.submit(() -> {
try {
Runnable trackingAccepted = () -> {
if (onAccepted != null) {
onAccepted.run();
}
asyncTasksByTarget.put(target, task);
};
Reply result = send(target, content, ASYNC_TIMEOUT_MS, trackingAccepted);
Reply result = send(target, content, ASYNC_TIMEOUT_MS, onAccepted, task);
if (result.outcome() == Outcome.QUESTION) {
asyncTasksByTarget.remove(target, task);
// Keep the accepted owner until answer() finishes it. markAsyncQuestion may run
// just after resolveQuestion wakes this thread.
} else {
finishAsyncTask(task, result);
}
} catch (Throwable t) {
task.future.completeExceptionally(t);
asyncTasksByTarget.remove(target, task);
}
});
pruneTerminalTickets();
@@ -564,6 +644,11 @@ public final class MessageService {
}
return new TaskView(ticket, Phase.PENDING, null, null, "worker " + liveStatus(task.target), null);
}
// CB-588: the ticket is terminal and being handed to the caller right here — tell the push
// loop it is collected so a later tick's nudge never names a ticket the lead already has.
if (pushLoop != null) {
pushLoop.ticketCollected(ticket);
}
Reply r;
try {
r = f.getNow(null);
@@ -575,9 +660,11 @@ public final class MessageService {
String source = r.outcome() == Outcome.REPLIED ? "reply" : "transcript";
return new TaskView(ticket, Phase.DONE, r.text(), source, null, null);
}
// A wedged worker (CB-109) carries the error context as its reason; the timeout/busy
// outcomes carry none, so fall back to the outcome name.
String detail = r.outcome() == Outcome.WORKER_FAILED && r.text() != null
// A wedged worker (CB-109) or a backend-exhausted classification (CB-578 stage A) carries
// the real cause as its reason; the timeout/busy outcomes carry none, so fall back to the
// outcome name.
boolean carriesReason = r.outcome() == Outcome.WORKER_FAILED || r.outcome() == Outcome.BACKEND_EXHAUSTED;
String detail = carriesReason && r.text() != null
? r.text()
: "no reply — " + r.outcome().name().toLowerCase();
return new TaskView(ticket, Phase.FAILED, null, null, detail, null);
@@ -592,20 +679,38 @@ public final class MessageService {
}
}
/** Drop finished tickets older than the TTL so the registry cannot grow without bound. */
/**
* Drop finished tickets older than the TTL so {@link #tasks} cannot grow without bound.
*
* <p>{@code tasks} is the sole authority on whether a ticket still exists — {@link #poll} returns
* {@code null} the instant a ticket is gone from here, before it ever reaches the terminal branch
* that calls {@link ReplyPushLoop#ticketCollected}. Without telling the push loop about a prune
* too, its own {@code pendingTickets} entry would outlive the ticket it names: an unpolled ticket
* (or one the reminder cap already gave up on) is pruned here but never collected there, so it
* lingers in {@code pendingTickets} forever and rides along on every later nudge to the same lead
* — naming a ticket {@code bridge_poll} can no longer find (CB-588 follow-up).
*/
private void pruneTerminalTickets() {
long cutoff = System.nanoTime() - TICKET_TTL_NANOS;
tasks.values().removeIf(t -> t.future.isDone() && t.createdNanos < cutoff);
long cutoff = nowNanos.getAsLong() - TICKET_TTL_NANOS;
tasks.entrySet().removeIf(e -> {
Task t = e.getValue();
boolean expired = t.future.isDone() && t.createdNanos < cutoff;
if (expired && pushLoop != null) {
pushLoop.ticketCollected(e.getKey());
}
return expired;
});
}
/** Record the active question for an async ticket; blocking sends have no entry and stay unchanged. */
private void markAsyncQuestion(String target, String text, String turnId) {
Task task = asyncTasksByTarget.get(target);
private Task markAsyncQuestion(CompletableFuture<Rendezvous.Resolution> waiter, String text, String turnId) {
Task task = waiter == null ? null : asyncTasksByWaiter.get(waiter);
if (task != null) {
task.question = new Reply(Outcome.QUESTION, text, turnId);
task.turnId = turnId;
asyncTasksByTurn.put(turnId, task);
}
return task;
}
/** Clear an answered or lapsed question, but only when it matches the ticket's current turn. */
@@ -623,7 +728,6 @@ public final class MessageService {
/** Complete and detach an async ticket after its worker's actual terminal reply. */
private void finishAsyncTask(Task task, Reply result) {
task.future.complete(result);
asyncTasksByTarget.remove(task.target, task);
if (task.turnId != null) {
asyncTasksByTurn.remove(task.turnId, task);
}
@@ -653,6 +757,7 @@ public final class MessageService {
case REPLY -> Outcome.REPLIED;
case COMPLETION -> Outcome.COMPLETED_UNREPLIED;
case FAILED -> Outcome.WORKER_FAILED;
case BACKEND_EXHAUSTED -> Outcome.BACKEND_EXHAUSTED;
case QUESTION -> Outcome.QUESTION;
};
}
@@ -33,6 +33,13 @@ public final class Rendezvous {
COMPLETION,
/** The worker ran the turn then wedged (CB-109); {@code text} is the failure context. */
FAILED,
/**
* The turn finished without a {@code bridge_reply}, and the scrape matched the backend's
* configured usage-limit refusal pattern (CB-578 stage A); {@code text} is the reason,
* carrying the matched line. The pane is healthy — only the account is refusing — so this
* is kept separate from a session simply going {@code GONE}.
*/
BACKEND_EXHAUSTED,
/**
* The worker paused mid-turn to ask the primary a question (CB-205 reverse rendezvous);
* {@code text} is the question and {@code turnId} correlates the primary's answer back to
@@ -224,6 +231,19 @@ public final class Rendezvous {
return waiter != null && waiter.complete(new Resolution(Kind.FAILED, reason));
}
/**
* Resolve a specific captured {@code waiter} as {@link Kind#BACKEND_EXHAUSTED} (CB-578 stage A):
* the turn finished with no {@code bridge_reply} and the scrape matched the backend's configured
* usage-limit pattern; {@code reason} carries the matched line. Like
* {@link #resolveCompletion(CompletableFuture, String)} it targets the exact captured send
* (CB-116). A no-op if that waiter was already resolved — first resolution wins.
*
* @return {@code true} if this call resolved the waiter, {@code false} if it was null or already resolved
*/
public boolean resolveExhausted(CompletableFuture<Resolution> waiter, String reason) {
return waiter != null && waiter.complete(new Resolution(Kind.BACKEND_EXHAUSTED, reason));
}
private boolean complete(String session, Resolution resolution) {
CompletableFuture<Resolution> waiter = waiters.get(session);
return waiter != null && waiter.complete(resolution);
@@ -8,9 +8,12 @@ import dev.ltms.bridged.metrics.Metrics;
import org.slf4j.Logger;
import org.slf4j.LoggerFactory;
import java.util.List;
import java.util.Set;
import java.util.concurrent.ConcurrentHashMap;
import java.util.concurrent.ScheduledExecutorService;
import java.util.concurrent.TimeUnit;
import java.util.stream.Collectors;
/**
* Mechanism (b) of CB-307: a dedicated, status-gated push loop that nudges the primary's own
@@ -23,11 +26,27 @@ import java.util.concurrent.TimeUnit;
*
* <p>Bounded: at most {@link #maxReminders} nudges per target, with a configurable backoff
* between them. The reply is never lost — the durable inbox is the backstop.
*
* <p><strong>CB-588 ticket nudges.</strong> {@link #onTicketTerminal(String, String, boolean)} is a
* second, independent entry point for an async delegation ticket ({@code bridge_send(wait:false)})
* reaching a terminal phase. That path takes the rendezvous fast path in {@link MessageService#reply}
* and never reaches {@link #onReplyQueued}, so without this the lead's own charter — prefer
* {@code wait:false} for anything non-trivial — was exactly the mode this loop failed to cover. It
* reuses the same status gating, bounded/backoff reminders, and metrics, keyed by the nudge-receiving
* lead terminal rather than the worker target so several tickets finishing together coalesce into one
* nudge. The two entry points do not interact: {@link #onReplyQueued} / {@link #decide} / their nudge
* text and bound are unchanged.
*/
public final class ReplyPushLoop {
private static final Logger log = LoggerFactory.getLogger(ReplyPushLoop.class);
static final String NUDGE_FORMAT = "Worker %s returned a reply — run bridge_poll(target=%s) to collect it";
/** CB-588: singular form, one uncollected ticket. */
static final String TICKET_NUDGE_FORMAT =
"Ticket %s finished%s — run bridge_poll(ticket=%s) to collect it";
/** CB-588: coalesced form, several uncollected tickets for the same lead. */
static final String TICKETS_NUDGE_FORMAT =
"%d tickets finished%s — run bridge_poll(ticket=...) for each to collect them: %s";
private final PrimaryRegistry primaryRegistry;
private final AgentControl agents;
@@ -39,6 +58,10 @@ public final class ReplyPushLoop {
/** Track targets that have an active schedule. */
private final ConcurrentHashMap<String, Boolean> activeTargets = new ConcurrentHashMap<>();
/** CB-588: tickets that have gone terminal but not yet been polled, keyed by ticket. */
private final ConcurrentHashMap<String, PendingTicket> pendingTickets = new ConcurrentHashMap<>();
/** CB-588: leads with an active ticket-reminder schedule. */
private final ConcurrentHashMap<String, Boolean> activeLeads = new ConcurrentHashMap<>();
public ReplyPushLoop(PrimaryRegistry primaryRegistry, AgentControl agents, ReplyInbox inbox,
ScheduledExecutorService scheduler,
@@ -173,23 +196,211 @@ public final class ReplyPushLoop {
scheduler.schedule(() -> tick(target, nextReminderCount), backoffMs, TimeUnit.MILLISECONDS);
}
// --- CB-588: async ticket terminal nudges ---------------------------------------------------
/** A ticket awaiting collection: which lead to nudge, and whether it ended in failure. */
private record PendingTicket(String ticket, String lead, boolean failed) {
}
/**
* Called when an async delegation ticket ({@code bridge_send(wait:false)}, CB-107) reaches a
* terminal phase — DONE or a failure. Unlike {@link #onReplyQueued}, which nudges about the
* durable-inbox no-waiter path, this covers the path {@code MessageService.reply} takes when a
* fire-and-poll send's own rendezvous waiter resolves the reply directly: that path returns
* before {@link #onReplyQueued} is ever called, so without this entry point a ticket finishing
* that way never nudged anyone (CB-588 / gitea #72).
*
* <p>Idempotent per lead: several tickets going terminal for the same lead while its schedule is
* already active coalesce onto that schedule's next tick rather than firing a nudge each.
*
* @param ticket the ticket to nudge about
* @param target the worker session the ticket was sent to — resolves which lead delegated it
* @param failed whether the ticket ended in a failure phase rather than {@code DONE}
*/
public void onTicketTerminal(String ticket, String target, boolean failed) {
var lead = primaryRegistry.nudgeTargetFor(target);
if (lead.isEmpty()) {
log.debug("push: no lead is known to be waiting on ticket {} (target {}), skipping nudge",
ticket, target);
return;
}
pendingTickets.put(ticket, new PendingTicket(ticket, lead.get(), failed));
if (activeLeads.putIfAbsent(lead.get(), Boolean.TRUE) != null) {
log.debug("push: ticket reminder loop already active for lead {}, {} coalesced in",
lead.get(), ticket);
return;
}
log.debug("push: starting ticket reminder loop for lead {}", lead.get());
scheduleTicketTick(lead.get(), 0);
}
/**
* Called when a ticket's terminal state has been collected via {@code bridge_poll}. Removes it
* from the pending set so a scheduled tick — and any nudge it sends — never names a ticket the
* lead already has (CB-588 acceptance #5). A ticket that was never pending (unknown ticket, or
* one nudged with no push loop configured) is a no-op.
*/
public void ticketCollected(String ticket) {
pendingTickets.remove(ticket);
}
/** Tickets still pending for {@code lead}, snapshotted fresh for one tick. */
private List<PendingTicket> pendingFor(String lead) {
return pendingTickets.values().stream().filter(t -> lead.equals(t.lead())).toList();
}
/**
* Pure decision function for ticket nudges, mirroring {@link #decide(String, int)} but keyed by
* the nudge-receiving lead terminal rather than the worker session — several tickets from
* different workers delegated by the same lead coalesce onto it.
*/
Action decideTickets(String lead, int reminderCount) {
if (pendingFor(lead).isEmpty()) {
log.debug("push: nothing pending for lead {}, stopping ticket reminder", lead);
return Action.STOP;
}
if (reminderCount >= maxReminders) {
log.debug("push: ticket reminder cap ({}) reached for lead {}, stopping", maxReminders, lead);
countNudge("exhausted");
return Action.STOP;
}
AgentStatus status;
try {
status = agents.status(lead);
} catch (RuntimeException e) {
log.debug("push: status check failed for lead {}, will retry", lead, e);
return Action.WAIT_BUSY;
}
if (status.injectable()) {
return Action.INJECT;
}
log.debug("push: lead {} is {} (not injectable), waiting", lead, status);
return Action.WAIT_BUSY;
}
/** Execute one ticket-loop tick — called on the scheduler thread. */
private void ticketTick(String lead, int reminderCount) {
Set<String> pendingBefore = pendingIdsFor(lead);
var action = decideTickets(lead, reminderCount);
switch (action) {
case INJECT -> {
injectTicketNudge(lead, reminderCount);
scheduleTicketTick(lead, reminderCount + 1);
}
case WAIT_BUSY -> scheduleTicketTick(lead, reminderCount);
case STOP -> stopOrRestartTicketLoop(lead, pendingBefore);
}
}
/** Ticket IDs pending for {@code lead} right now, as a plain snapshot for race comparison. */
private Set<String> pendingIdsFor(String lead) {
return pendingFor(lead).stream().map(PendingTicket::ticket).collect(Collectors.toUnmodifiableSet());
}
/**
* Release {@code lead}'s active-schedule slot, then restart it only if a ticket landed that
* {@code pendingBefore} — the snapshot taken just before this tick's decision — did not already
* account for. {@code onTicketTerminal} reads {@code activeLeads} to decide whether to coalesce
* onto an existing schedule or start one, so a ticket that lands between {@code decideTickets}
* returning {@link Action#STOP} and this removal running sees the (soon-to-be-stale) slot as
* occupied, coalesces onto a schedule that is about to die, and gets no nudge scheduled at all —
* a lost nudge, the exact failure CB-588 exists to remove (found in review, gitea PR #73).
*
* <p>Restarting on ANY non-empty {@code pendingFor(lead)} would be wrong: when STOP is reached
* because the reminder cap was hit rather than the backlog draining, the same never-collected
* ticket is expected to still be sitting there — that is the cap doing its job — and restarting
* would nudge about it forever, defeating the bound (the original CB-307 bounded-reminder
* guarantee, carried into CB-588 by acceptance criterion #7 — this exact regression showed up as
* two existing tests failing once a naive "any pending ticket restarts" version of this fix went
* in: {@code successfulTicketNudgeIncrementsDelivered} and {@code ticketNudgesSendUpToCapThenStop}).
* Diffing the current pending set against {@code pendingBefore} tells the two cases apart: a
* ticket present before this tick's decision is stale backlog, not a race; only a ticket absent
* from {@code pendingBefore} can only have arrived during the decision-to-release window, which is
* exactly the race this method closes.
*
* <p>Package-private so a test can drive the interleaving directly rather than trying to force a
* genuine thread race: pass the exact {@code pendingBefore} snapshot a race requires (or does
* not) and call this to prove the recheck responds correctly either way.
*
* <p>Terminates rather than spinning: this method restarts the schedule at most once per call, and
* a fresh {@link #onTicketTerminal} racing the recheck below still terminates in one of two ways —
* either it observes the slot already vacated (by the {@code activeLeads.remove} above, which
* happens-before this recheck in program order) and claims it itself, or it lands first and this
* recheck then observes its ticket in {@code pendingTickets} and reclaims the slot instead. Exactly
* one side always wins; neither can miss the other, so this never loops on its own account.
*/
void stopOrRestartTicketLoop(String lead, Set<String> pendingBefore) {
activeLeads.remove(lead);
boolean ticketRacedIn = pendingFor(lead).stream().anyMatch(t -> !pendingBefore.contains(t.ticket()));
if (ticketRacedIn && activeLeads.putIfAbsent(lead, Boolean.TRUE) == null) {
log.debug("push: a ticket for lead {} raced the reminder loop's stop — restarting", lead);
scheduleTicketTick(lead, 0);
return;
}
log.debug("push: ticket reminder loop ended for lead {}", lead);
}
/** Send the coalesced ticket nudge and log the event. */
private void injectTicketNudge(String lead, int reminderCount) {
// Re-read rather than threading it down from decideTickets(): a ticket can be collected (or
// another can arrive) between the decision and the injection.
List<PendingTicket> pending = pendingFor(lead);
if (pending.isEmpty()) {
log.debug("push: pending tickets for lead {} drained before the nudge could be sent", lead);
return;
}
String nudge = formatTicketsNudge(pending);
try {
agents.send(lead, nudge);
log.debug("push: ticket nudge {}/{} sent to lead {} for {} ticket(s)",
reminderCount + 1, maxReminders, lead, pending.size());
countNudge("delivered");
} catch (RuntimeException e) {
log.warn("push: failed to nudge lead {} for {} ticket(s) (reminder {}/{}): {}",
lead, pending.size(), reminderCount + 1, maxReminders, e.toString());
}
}
/** Schedule the next ticket-loop tick on the scheduler thread pool. */
private void scheduleTicketTick(String lead, int nextReminderCount) {
scheduler.schedule(() -> ticketTick(lead, nextReminderCount), backoffMs, TimeUnit.MILLISECONDS);
}
/** Render one or several pending tickets as a single nudge line. */
private static String formatTicketsNudge(List<PendingTicket> pending) {
if (pending.size() == 1) {
PendingTicket t = pending.get(0);
return TICKET_NUDGE_FORMAT.formatted(t.ticket(), t.failed() ? " (FAILED)" : "", t.ticket());
}
long failedCount = pending.stream().filter(PendingTicket::failed).count();
String ids = pending.stream()
.map(t -> t.failed() ? t.ticket() + " (FAILED)" : t.ticket())
.collect(Collectors.joining(", "));
String failedNote = failedCount > 0 ? " (%d failed)".formatted(failedCount) : "";
return TICKETS_NUDGE_FORMAT.formatted(pending.size(), failedNote, ids);
}
// --- lifecycle -----------------------------------------------------------------------------
/**
* Whether any reminder loop is currently active for some target (CB-551). The idle-lead heartbeat
* uses this to stand aside: while the push loop is actively nudging the lead, a concurrent
* heartbeat injection would start a second competing turn in the same pane — racing loops multiply
* turns and context burn. "Active" means a schedule exists in {@link #activeTargets}; the set is
* bounded by what has been triggered, not by any persistent state.
* turns and context burn. "Active" means a schedule exists in {@link #activeTargets} or
* {@link #activeLeads} (CB-588 ticket nudges are a second source of pane injections the heartbeat
* must equally stand aside for); the sets are bounded by what has been triggered, not by any
* persistent state.
*/
public boolean isActive() {
return !activeTargets.isEmpty();
return !activeTargets.isEmpty() || !activeLeads.isEmpty();
}
/** Shut down the scheduler. Outstanding reminders are cancelled. */
public void stop() {
scheduler.shutdownNow();
activeTargets.clear();
activeLeads.clear();
pendingTickets.clear();
}
/** @see #stop() */
@@ -0,0 +1,20 @@
package dev.ltms.bridged.msg;
import java.util.concurrent.CompletableFuture;
/**
* Identity for one accepted send. The session turn is deliberately absent: CompletionResolver's
* delivery callback runs before SessionManager.onDelivered, so binding it needs a later ordering design.
*/
public final class TurnToken {
private final String target;
private final CompletableFuture<Rendezvous.Resolution> waiter;
public TurnToken(String target, CompletableFuture<Rendezvous.Resolution> waiter) {
this.target = target;
this.waiter = waiter;
}
public String target() { return target; }
public CompletableFuture<Rendezvous.Resolution> waiter() { return waiter; }
}
@@ -0,0 +1,84 @@
package dev.ltms.bridged.peer;
import java.nio.charset.StandardCharsets;
import java.security.MessageDigest;
import java.security.NoSuchAlgorithmException;
import java.util.HexFormat;
/**
* CB-571: a fingerprint of the exact charter bytes handed to a spawned member.
*
* <p>Lets an operator prove <em>which</em> charter a member actually got, without ever logging the
* charter text. The digest covers the exact composed UTF-8 string {@code HerdrPeerLauncher} passes
* to its adapter as {@code LaunchSpec.charter()}, so every adapter that receives the same string —
* Claude inlining it, OpenCode writing it to a file — produces the same digest for the same config.
* Two spawns of the same role from the same config agree; editing the charter changes the digest.
*
* <p>Deliberately places no charter prose. A charter is operator-authored text that may name
* internal projects or unreleased plans, and logs get tailed, shipped, and pasted into tickets.
* The {@code charterSource} key is what the operator wants to confirm, and it carries no content.
*/
public record CharterReceipt(
MemberRole role,
String profile,
String charterSource,
String charterSha256,
int charterBytes) {
/** Source reported when the role has no configured charter, so the field is never omitted. */
public static final String NO_SOURCE = "none";
/**
* The config key that supplied the role's charter text, e.g. {@code fleet.charters.architect}.
*/
public static String sourceKey(MemberRole role) {
return "fleet.charters." + (role == null ? "?" : role.wireName());
}
/**
* Fingerprint the composed charter for {@code role} on {@code profile}. {@code configured} is
* the role's charter text as read from config ({@code null} when none is configured);
* {@code composed} is the exact string the launcher will pass to the adapter — the reply
* charter may be appended to {@code configured}, or stand alone when no role charter exists.
*
* <p>No composed charter at all is reported as an explicit absence — a {@code null} digest and
* a zero byte count — never a digest of the empty string, which would hide the fact that no
* text was supplied. {@code configured} being {@code null} while {@code composed} is the reply
* charter alone is a normal case, and the source says so.
*/
public static CharterReceipt compose(MemberRole role, String profile,
String configured, String composed) {
String source = (configured == null || configured.isBlank())
? NO_SOURCE : sourceKey(role);
if (composed == null) {
return new CharterReceipt(role, profile, source, null, 0);
}
byte[] bytes = composed.getBytes(StandardCharsets.UTF_8);
return new CharterReceipt(role, profile, source, digestOf(composed), bytes.length);
}
/** Whether the composed charter was absent (no text was given to the member). */
public boolean absent() {
return charterSha256 == null;
}
/**
* The stable SHA-256 hex digest of {@code text}, or {@code null} for null/blank text. Used both
* for the receipt's fingerprint and to redact a charter argument in a spawn log.
*/
public static String digestOf(String text) {
if (text == null || text.isBlank()) {
return null;
}
return sha256Hex(text.getBytes(StandardCharsets.UTF_8));
}
private static String sha256Hex(byte[] bytes) {
try {
MessageDigest md = MessageDigest.getInstance("SHA-256");
return HexFormat.of().formatHex(md.digest(bytes));
} catch (NoSuchAlgorithmException e) {
throw new IllegalStateException("SHA-256 is unavailable", e);
}
}
}
@@ -62,9 +62,27 @@ public interface PeerHandle {
* non-null for a spawn that requested session identity, because it knows the id before the
* peer has written anything.
*
* <p>Deliberately not a {@code default} (CB-584, the same fix CB-571 made for
* {@link #charterReceipt()} one method below): a decorator that forgets to override this
* silently answers {@code null} for a question it has no basis to answer, and the gap surfaces
* only as a resume that quietly starts a cold session, not a compile error. Every
* implementation must answer explicitly.
*
* @return the peer's own session id, or {@code null} when not determinable
*/
default String agentSessionId() {
return null;
}
String agentSessionId();
/**
* The charter receipt (CB-571) for this peer's launch — the fingerprint of the exact charter
* bytes it was started with. {@code null} when the launcher records none (a non-instrumented
* adapter, or a launcher before this field); the session registry stores it so the spawn result
* and the roster row can show an operator which charter a member actually got.
*
* <p>Deliberately not a {@code default}: a decorator that forgets to override this silently
* answers {@code null} for a question it has no basis to answer, and the gap surfaces only as
* a missing roster field, not a compile error. Every implementation must answer explicitly.
*
* @return the fingerprint, or {@code null} when the launcher carries none
*/
CharterReceipt charterReceipt();
}
@@ -25,6 +25,19 @@ public interface PeerLauncher {
*/
Set<Capability> capabilities();
/**
* The capabilities of the adapter that {@code profileName} resolves to (null/blank → the
* default profile, the same resolution {@link #spawn} uses). Distinct from {@link
* #capabilities()}, which unions every configured adapter: a caller that must know whether
* <em>this</em> profile's backend supports a capability — e.g. {@link Capability#SESSION_RESUME}
* before honoring {@link SpawnRequest#resumeSessionId()} — needs the per-profile answer, not
* the fleet-wide union, or a mixed fleet could OK a resume that lands on a non-supporting
* adapter (CB-584).
*
* @throws IllegalArgumentException if the profile is unknown and no default is configured
*/
Set<Capability> capabilitiesFor(String profileName);
/**
* {@code profileName}/requestedCwd null/blank → default resolution. Returns after the peer
* process is live (env + argv + placement complete). Never returns {@code null}.
@@ -0,0 +1,115 @@
package dev.ltms.bridged.placement;
import java.util.LinkedHashMap;
import java.util.Map;
import java.util.Objects;
import java.util.OptionalLong;
import java.util.concurrent.ConcurrentHashMap;
import java.util.function.LongSupplier;
/**
* Where a credential (not a profile — see {@code BridgedConfig.Profile#effectiveCredentialId()})
* sits out a cooldown after a {@code BACKEND_EXHAUSTED} classification (CB-578 stage B), so a fresh
* spawn does not walk straight back onto the account that just refused on a usage limit.
*
* <p>Keyed by credential id, never by profile name: two profiles sharing one credential (e.g. two
* models on the same OpenAI account) share one quarantine — {@link #quarantine} one credential id
* and every profile whose {@code effectiveCredentialId()} equals it is quarantined too, without this
* class knowing anything about profiles at all. That mapping is the caller's job (see
* {@code CompositePeerLauncher} and {@code dev.ltms.bridged.inject.ExhaustionSink}).
*
* <p>The clock is injected ({@link LongSupplier}, conventionally {@code System::nanoTime} like
* {@code FleetHealthMonitor}), never read inline, so a quarantine's expiry is testable without a
* real sleep.
*/
public final class BackendQuarantine {
private final ConcurrentHashMap<String, Long> quarantinedUntilNanos = new ConcurrentHashMap<>();
private final LongSupplier nowNanos;
private final long cooldownNanos;
/** True only for {@link #none()}. See {@link #quarantine} for why this exists. */
private final boolean inert;
/**
* @param nowNanos monotonic clock, injected for testability
* @param cooldownNanos how long a fresh {@link #quarantine} call blocks the credential for;
* must be positive
*/
public BackendQuarantine(LongSupplier nowNanos, long cooldownNanos) {
this(nowNanos, cooldownNanos, false);
}
private BackendQuarantine(LongSupplier nowNanos, long cooldownNanos, boolean inert) {
this.nowNanos = Objects.requireNonNull(nowNanos, "nowNanos");
if (cooldownNanos <= 0) {
throw new IllegalArgumentException("cooldownNanos must be positive: " + cooldownNanos);
}
this.cooldownNanos = cooldownNanos;
this.inert = inert;
}
/**
* Inert quarantine — {@link #quarantine} does nothing on this instance, so nothing is ever
* quarantined. The explicit stand-in a caller (or a test not exercising this feature) passes
* instead of a defaulting overload, exactly like {@code ExhaustedPatternLookup.none()}.
*/
public static BackendQuarantine none() {
return new BackendQuarantine(() -> 0L, 1, true);
}
/**
* Quarantine {@code credentialId} for the configured cooldown, starting now. A repeat call while
* already quarantined restarts the cooldown at full length — a fresh refusal is fresh evidence the
* account is still exhausted, not a reason to let an earlier, shorter wait stand.
*
* <p>On {@link #none()} this is a no-op. It has to be: that instance holds a clock frozen at 0,
* so recording a deadline would produce a quarantine that never expires — a credential locked out
* for the life of the daemon. Two production {@code CompositePeerLauncher} constructors default to
* {@code none()}, so the failure would be silent and permanent. An inert stand-in must omit the
* fact, never invent one.
*/
public void quarantine(String credentialId) {
Objects.requireNonNull(credentialId, "credentialId");
if (inert) {
return;
}
quarantinedUntilNanos.put(credentialId, nowNanos.getAsLong() + cooldownNanos);
}
/** Whether {@code credentialId} is quarantined right now. */
public boolean isQuarantined(String credentialId) {
return remainingNanos(credentialId) > 0;
}
/** Seconds left on {@code credentialId}'s quarantine, or empty when it is not quarantined. */
public OptionalLong remainingSeconds(String credentialId) {
long remaining = remainingNanos(credentialId);
return remaining > 0 ? OptionalLong.of(toSecondsRoundedUp(remaining)) : OptionalLong.empty();
}
/**
* Every currently-quarantined credential id and its remaining seconds (CB-578 stage B fleet
* reporting) — expired entries are never included. Not pruned from the backing map here: it stays
* small (bounded by the number of distinct credentials ever exhausted) and a lazily-stale entry is
* harmless, since every read already checks the deadline.
*/
public Map<String, Long> activeRemainingSeconds() {
Map<String, Long> out = new LinkedHashMap<>();
quarantinedUntilNanos.forEach((credentialId, deadline) -> {
long remaining = deadline - nowNanos.getAsLong();
if (remaining > 0) {
out.put(credentialId, toSecondsRoundedUp(remaining));
}
});
return out;
}
private long remainingNanos(String credentialId) {
Long deadline = quarantinedUntilNanos.get(credentialId);
return deadline == null ? 0L : deadline - nowNanos.getAsLong();
}
private static long toSecondsRoundedUp(long nanos) {
return (nanos + 999_999_999L) / 1_000_000_000L;
}
}
@@ -4,19 +4,63 @@ package dev.ltms.bridged.placement;
* Backward-compatible placement: an unqualified spawn always resolves to the configured default
* profile, exactly as {@code CompositePeerLauncher} did before CB-518. This ignores caps and
* reachability so that a pre-existing config behaves identically after upgrade.
*
* <p>Two exceptions walk past the default instead of returning it unconditionally:
* <ul>
* <li>Quarantine (CB-578 stage B): a quarantined default is a credential that just refused on
* a usage limit, not a transient capacity or reachability concern.
* <li>Weight 0 (CB-554): {@code fixed} is still automatic selection, so a profile the operator
* marked "never auto-select me" ({@code weight <= 0}) must be skipped here exactly as
* {@code weighted}/{@code round-robin} skip it — an explicit {@code bridge_spawn} naming
* the profile is unaffected, only this automatic fallback walk.
* </ul>
* A fleet where nothing is ever quarantined or weight-0 never exercises either path, so today's
* behaviour is unchanged.
*/
final class FixedPlacementPolicy implements PlacementPolicy {
@Override
public PlacementCandidate select(PlacementContext ctx) {
String d = ctx.defaultProfile();
if (d != null && !d.isBlank()) {
if (d != null && !d.isBlank() && !ctx.quarantined().contains(d) && !weightExcluded(ctx, d)) {
return new PlacementCandidate(d, null, 1.0f, null);
}
for (PlacementCandidate c : ctx.candidates()) {
if (!ctx.quarantined().contains(c.profile()) && !c.excluded()) {
return new PlacementCandidate(c.profile(), null, c.weight(), c.maxLoad());
}
}
if (d != null && !d.isBlank()) {
boolean dQuarantined = ctx.quarantined().contains(d);
boolean dWeightExcluded = weightExcluded(ctx, d);
if (dQuarantined && dWeightExcluded) {
throw new PlacementException("worker profile '" + d + "' is quarantined (backend "
+ "exhausted) and has weight 0 (excluded from automatic selection), and no "
+ "available candidate remains");
}
if (dWeightExcluded) {
throw new PlacementException("worker profile '" + d + "' has weight 0 (excluded "
+ "from automatic selection) and no available candidate remains");
}
if (dQuarantined) {
throw new PlacementException("worker profile '" + d + "' is quarantined (backend "
+ "exhausted) and no un-quarantined candidate is available");
}
}
if (!ctx.candidates().isEmpty()) {
PlacementCandidate first = ctx.candidates().getFirst();
return new PlacementCandidate(first.profile(), null, first.weight(), first.maxLoad());
throw new PlacementException(
"all worker profiles are excluded from automatic selection (quarantined or weight-0)");
}
throw new PlacementException("no worker profiles configured");
}
/** Whether {@code profile} carries {@code weight <= 0} (CB-554) among {@code ctx}'s candidates. */
private static boolean weightExcluded(PlacementContext ctx, String profile) {
for (PlacementCandidate c : ctx.candidates()) {
if (c.profile().equals(profile)) {
return c.excluded();
}
}
return false;
}
}
@@ -17,4 +17,18 @@ public record PlacementCandidate(String profile, String host, float weight, Inte
public static PlacementCandidate profile(String profile) {
return new PlacementCandidate(profile, null, 1.0f, null);
}
/**
* True when this candidate carries an explicit {@code weight <= 0} (CB-554) and must be
* skipped by every automatic policy — the same way a quarantined or unreachable candidate is
* skipped. {@code BridgedConfig.Profile}'s compact constructor already normalises "absent" to
* {@code 1.0} and "negative" to {@code 0.0}, so this is a plain threshold check here; it does
* not need to distinguish "explicit 0" from "absent" itself.
*
* <p>Exclusion is about <em>automatic</em> selection only — an explicit
* {@code bridge_spawn{profile:"..."}} bypasses placement entirely and is unaffected.
*/
public boolean excluded() {
return weight <= 0.0f;
}
}
@@ -11,9 +11,13 @@ import java.util.function.Function;
* @param candidates every configured candidate; the policy filters out those at cap or unreachable
* @param liveCount current live worker count per profile (from the session registry)
* @param unreachable profiles already known to have failed in this spawn attempt
* @param quarantined profiles whose credential is currently quarantined (CB-578 stage B) — a
* {@code BACKEND_EXHAUSTED} classification put it, or a profile it shares a
* credential with, on cooldown. Filtered the same way as {@code unreachable}.
*/
public record PlacementContext(String defaultProfile,
List<PlacementCandidate> candidates,
Function<String, Integer> liveCount,
Set<String> unreachable) {
Set<String> unreachable,
Set<String> quarantined) {
}
@@ -12,13 +12,16 @@ final class PlacementPolicyUtil {
}
/**
* Candidates that are not known-unreachable and have not reached their maxLoad.
* Candidates that are not weight-excluded (CB-554: explicit {@code weight <= 0}, checked
* first because it is a static config choice rather than transient state), not
* known-unreachable, not quarantined (CB-578 stage B), and have not reached their maxLoad.
* A {@code null} maxLoad means unlimited.
*/
static List<PlacementCandidate> available(PlacementContext ctx) {
List<PlacementCandidate> out = new ArrayList<>();
for (PlacementCandidate c : ctx.candidates()) {
if (ctx.unreachable().contains(c.profile())) {
if (c.excluded() || ctx.unreachable().contains(c.profile())
|| ctx.quarantined().contains(c.profile())) {
continue;
}
Integer cap = c.maxLoad();
@@ -34,15 +37,23 @@ final class PlacementPolicyUtil {
}
/**
* Build a clear exception describing why every candidate was dropped: all at capacity,
* all unreachable, or a mix.
* Build a clear exception describing why every candidate was dropped: all weight-0, all
* quarantined, all at capacity, all unreachable, or a mix. Each candidate is counted into
* exactly one bucket (weight-excluded takes priority) so a candidate excluded for more than
* one reason is never double-counted.
*/
static PlacementException emptyException(PlacementContext ctx) {
int weightExcluded = 0;
int atCap = 0;
int unreachable = 0;
int quarantined = 0;
for (PlacementCandidate c : ctx.candidates()) {
Integer cap = c.maxLoad();
if (ctx.unreachable().contains(c.profile())) {
if (c.excluded()) {
weightExcluded++;
} else if (ctx.quarantined().contains(c.profile())) {
quarantined++;
} else if (ctx.unreachable().contains(c.profile())) {
unreachable++;
} else if (cap != null && ctx.liveCount().apply(c.profile()) >= cap) {
atCap++;
@@ -53,6 +64,13 @@ final class PlacementPolicyUtil {
if (total == 0) {
return new PlacementException("no worker profiles configured");
}
if (weightExcluded == total) {
return new PlacementException(
"all worker profiles have weight 0 (excluded from automatic selection)");
}
if (quarantined == total) {
return new PlacementException("all worker profiles are quarantined (backend exhausted)");
}
if (atCap == total) {
return new PlacementException("all worker profiles are at maxLoad");
}
@@ -60,6 +78,8 @@ final class PlacementPolicyUtil {
return new PlacementException("all worker profiles are unreachable");
}
return new PlacementException("no worker profile available: " + atCap + " at maxLoad, "
+ unreachable + " unreachable, " + (total - atCap - unreachable) + " remaining");
+ unreachable + " unreachable, " + quarantined + " quarantined, "
+ weightExcluded + " weight-0, "
+ (total - atCap - unreachable - quarantined - weightExcluded) + " remaining");
}
}
@@ -234,7 +234,14 @@ public final class BridgedApp {
List<Map<String, Object>> out = sessions.roster().stream()
.map(s -> SessionManager.rosterView(s, live.get(s.terminalId())))
.toList();
ctx.status(200).json(Map.of("workers", out));
Map<String, Object> body = new LinkedHashMap<>();
body.put("workers", out);
// CB-586: operator visibility for the refs/wip snapshot store without shelling into the
// repo — how many snapshot refs exist and roughly what they cost. Present only once a
// worktree session has established the repo, so a never-snapshotted fleet reports nothing.
sessions.wipRefs().ifPresent(st -> body.put("wipRefs",
Map.of("count", st.count(), "costBytes", st.costBytes())));
ctx.status(200).json(body);
}
/** The configured worker profiles and which one a no-argument spawn uses. */
@@ -403,9 +410,12 @@ public final class BridgedApp {
case TIMED_OUT_QUEUED -> "queued";
case BUSY -> "busy";
case WORKER_FAILED -> "failed";
case BACKEND_EXHAUSTED -> "backend_exhausted";
default -> "done"; // unreachable (terminal outcomes handled above)
},
"detail", reply.outcome() == MessageService.Outcome.WORKER_FAILED && reply.text() != null
"detail", (reply.outcome() == MessageService.Outcome.WORKER_FAILED
|| reply.outcome() == MessageService.Outcome.BACKEND_EXHAUSTED)
&& reply.text() != null
? reply.text()
: "no reply within " + timeout + "ms; poll status or retry"));
}
@@ -12,7 +12,12 @@ import java.nio.file.Files;
import java.nio.file.Path;
import java.nio.file.StandardCopyOption;
import java.security.SecureRandom;
import java.util.ArrayList;
import java.util.HashSet;
import java.util.List;
import java.util.Map;
import java.util.Optional;
import java.util.Set;
import java.util.concurrent.TimeUnit;
import java.util.concurrent.atomic.AtomicLong;
import java.util.stream.Collectors;
@@ -167,6 +172,22 @@ public final class GitWorktrees implements Worktrees {
exec("git", "-C", repoRoot, "worktree", "remove", "--force", worktreePath);
}
@Override
public boolean hasUncommitted(String worktreePath) {
// A worktree that is already gone holds no work to lose, and it must not break teardown:
// git -C <missing-dir> status exits non-zero and would throw where release() is mid-way
// through stopping a pane. Mirror remove()'s already-gone tolerance by treating it as clean.
Path p = Path.of(worktreePath);
if (!Files.exists(p)) {
log.debug("worktree {} already gone — nothing can be uncommitted", worktreePath);
return false;
}
// No --untracked-files=no: the exact shape of the work lost in CB-576 was a new file
// that was never added, so an untracked-only worktree is still dirty.
String out = exec("git", "-C", worktreePath, "status", "--porcelain");
return !out.isBlank();
}
@Override
public void overlayParity(String repoRoot, String worktreePath, List<String> overlay) {
if (overlay == null || overlay.isEmpty()) {
@@ -201,6 +222,209 @@ public final class GitWorktrees implements Worktrees {
return Path.of(out.trim()).toAbsolutePath().normalize().toString();
}
/**
* CB-578 stage C, fixed by CB-587. Stages into a <em>temporary</em> index (never the worktree's
* real one, which the worker may still be writing to) — but that temp index is first <em>seeded</em>
* from the worktree's real one, rather than starting empty:
*
* <pre>
* cp $(git -C worktree rev-parse --git-path index) &lt;temp&gt;
* GIT_INDEX_FILE=&lt;temp&gt; git -C worktree add -A
* tree=$(GIT_INDEX_FILE=&lt;temp&gt; git -C worktree write-tree)
* commit=$(git -C worktree commit-tree $tree -p HEAD -m message)
* git -C worktree update-ref refs/wip/branch $commit
* </pre>
*
* A fresh empty index carries none of the real index's {@code --skip-worktree} /
* {@code --assume-unchanged} bits, so {@code add -A} into it stages a skip-worktree file's local
* on-disk content even though {@code git status --porcelain} correctly hides that file (CB-587).
* Seeding from the real index preserves those bits, so {@code add -A} then skips exactly what
* {@code git status} skips. {@code add -A} (never {@code -f}) also still respects
* {@code .gitignore} exactly as it would in the real index — a gitignored file staying ignored is
* what keeps secrets and local config out of the snapshot's tree. The temporary index file is
* removed afterwards regardless of outcome; the worker's real index is never opened for writing.
*/
@Override
public Optional<String> snapshot(String worktreePath, String branch, String message) {
if (!Files.exists(Path.of(worktreePath))) {
log.debug("worktree {} already gone — nothing to snapshot", worktreePath);
return Optional.empty();
}
Path tempIndex;
try {
tempIndex = Files.createTempFile("bridged-wip-index-", ".tmp");
} catch (IOException e) {
throw new WorktreeException("cannot create a temporary index for snapshot: " + e.getMessage(), e);
}
Map<String, String> indexEnv = Map.of("GIT_INDEX_FILE", tempIndex.toAbsolutePath().toString());
try {
Path realIndex = resolveRealIndex(worktreePath);
try {
Files.copy(realIndex, tempIndex, StandardCopyOption.REPLACE_EXISTING);
} catch (IOException e) {
throw new WorktreeException("cannot copy the worktree's real index (" + realIndex
+ ") into the temporary snapshot index: " + e.getMessage(), e);
}
exec(indexEnv, "git", "-C", worktreePath, "add", "-A");
String tree = exec(indexEnv, "git", "-C", worktreePath, "write-tree").trim();
String commit = exec("git", "-C", worktreePath, "commit-tree", tree, "-p", "HEAD", "-m", message).trim();
exec("git", "-C", worktreePath, "update-ref", "refs/wip/" + branch, commit);
log.info("snapshotted worktree {} to refs/wip/{} commit={}", worktreePath, branch, commit);
return Optional.of(commit);
} finally {
try {
Files.deleteIfExists(tempIndex);
} catch (IOException e) {
log.debug("could not delete temporary snapshot index {}: {}", tempIndex, e.getMessage());
}
}
}
/**
* Resolve the path of {@code worktreePath}'s real index. Never assume {@code <worktree>/.git/index}:
* in a linked worktree {@code .git} is a <em>file</em> pointing at the main repo's
* {@code worktrees/<name>/} directory, and that is where the real per-worktree index lives.
* {@code git rev-parse --git-path index} resolves this correctly for both a linked worktree and
* the main checkout. Throws {@link WorktreeException} — same as every other failure in this
* class — if the command fails or the resolved path does not exist, rather than silently
* snapshotting from an empty index.
*/
private Path resolveRealIndex(String worktreePath) {
String out = exec("git", "-C", worktreePath, "rev-parse", "--git-path", "index").trim();
Path index = Path.of(out);
if (!index.isAbsolute()) {
index = Path.of(worktreePath).resolve(index).normalize();
}
if (!Files.exists(index)) {
throw new WorktreeException("worktree's real index not found at resolved path " + index
+ " (git rev-parse --git-path index reported '" + out + "')");
}
return index;
}
/**
* One {@code refs/wip/<branch>} snapshot ref as read by {@link #listWipRefs}: its full ref name,
* the snapshot commit's sha, and that commit's committer time in unix millis (the age of the
* snapshot — a snapshot is written once and never rewritten, so the commit date is the ref's).
*/
private record WipRef(String refName, String sha, long committerMillis) {
String branch() {
return refName.substring("refs/wip/".length());
}
}
@Override
public WipRefStats wipRefs(String repoRoot) {
List<WipRef> refs = listWipRefs(repoRoot);
long costBytes = 0;
for (WipRef ref : refs) {
costBytes += treeSize(repoRoot, ref.sha());
}
return new WipRefStats(refs.size(), costBytes);
}
@Override
public int pruneWipRefs(String repoRoot, long minAgeMillis) {
// The rule is documented on Worktrees#pruneWipRefs: delete only a snapshot whose tree
// content is already reachable from main AND that is older than minAgeMillis. Reachability
// is the floor that keeps a worker's last copy; the age floor keeps a just-written snapshot
// from being swept while a lead may still be looking at it.
List<WipRef> refs = listWipRefs(repoRoot);
if (refs.isEmpty()) {
return 0;
}
long nowMillis = System.currentTimeMillis();
// Resolve what main carries once per sweep, not once per ref.
Set<String> mainObjects = reachableObjectsFromMain(repoRoot);
int deleted = 0;
for (WipRef ref : refs) {
long ageMillis = nowMillis - ref.committerMillis();
if (ageMillis <= minAgeMillis) {
continue; // too recent — never swept, even if it looks recoverable (CB-586)
}
String tree = exec("git", "-C", repoRoot, "rev-parse", ref.sha() + "^{tree}").trim();
if (!mainObjects.contains(tree)) {
// Last copy of the snapshot's content — the worker's work exists nowhere else.
// Never delete automatically (CB-586 criterion 2).
continue;
}
exec("git", "-C", repoRoot, "update-ref", "-d", ref.refName());
deleted++;
log.info("pruned snapshot ref refs/wip/{} commit={} (age {}h): its tree is already "
+ "reachable from main, so the work is preserved; recover from reflog via "
+ "git update-ref refs/wip/{} {}",
ref.branch(), ref.sha(), TimeUnit.MILLISECONDS.toHours(ageMillis),
ref.branch(), ref.sha());
}
return deleted;
}
/**
* Every {@code refs/wip/*} ref (see {@link WipRef}). The committer date is read as a unix
* count of seconds and converted to millis. {@code %00} (NUL) separates the fields because a
* branch name may contain spaces.
*/
private List<WipRef> listWipRefs(String repoRoot) {
String out = exec("git", "-C", repoRoot, "for-each-ref",
"--format=%(refname)%00%(objectname)%00%(committerdate:unix)", "refs/wip/");
List<WipRef> refs = new ArrayList<>();
for (String line : out.split("\\R")) {
if (line.isBlank()) {
continue;
}
String[] parts = line.split("\u0000", -1);
if (parts.length == 3 && !parts[1].isBlank()) {
refs.add(new WipRef(parts[0], parts[1], Long.parseLong(parts[2]) * 1000L));
}
}
return refs;
}
/**
* The set of object shas reachable from {@code main}, or an empty set when {@code main} cannot
* be resolved. An empty set is the safe direction: the retention sweep then concludes nothing
* is recoverable, so it deletes nothing — a repo with no {@code main} must never cause a
* worker's last copy of a snapshot to be dropped on a reachability misreading.
*/
private Set<String> reachableObjectsFromMain(String repoRoot) {
if (exitCode("git", "-C", repoRoot, "rev-parse", "--verify", "main") != 0) {
log.debug("refs/wip retention: no 'main' ref in {} — treating nothing as reachable", repoRoot);
return Set.of();
}
String out = exec("git", "-C", repoRoot, "rev-list", "--objects", "main");
Set<String> objects = new HashSet<>();
for (String line : out.split("\\R")) {
if (line.isBlank()) {
continue;
}
int sp = line.indexOf(' ');
objects.add(sp < 0 ? line : line.substring(0, sp));
}
return objects;
}
/** Approximate cost of a snapshot: the sum of every blob's size in its committed tree. */
private long treeSize(String repoRoot, String sha) {
String out = exec("git", "-C", repoRoot, "ls-tree", "-r", "-l", sha);
long total = 0;
for (String line : out.split("\\R")) {
if (line.isBlank()) {
continue;
}
// ls-tree -l row: "<mode> <type> <object> <size>\t<path>"; the size is only numeric for
// blobs (trees read "-"), so gate on the type token and take the 4th whitespace field.
String[] parts = line.split("\\s+");
if (parts.length >= 4 && "blob".equals(parts[1])) {
try {
total += Long.parseLong(parts[3]);
} catch (NumberFormatException ignored) {
// a '-' size (or any anomaly) contributes nothing to the rough figure
}
}
}
return total;
}
/** Resolve the directory that will hold per-session worktree checkouts. */
private Path resolveRoot(String repoRoot) {
if (configuredRoot != null && !configuredRoot.isBlank()) {
@@ -223,11 +447,20 @@ public final class GitWorktrees implements Worktrees {
* stdout and stderr (merged by redirectErrorStream).
*/
private String exec(String... command) {
return exec(Map.of(), command);
}
/** Same as {@link #exec(String...)}, with extra environment variables set on the child process. */
private String exec(Map<String, String> extraEnv, String... command) {
String out;
int code;
Process p;
try {
p = new ProcessBuilder(command).redirectErrorStream(true).start();
ProcessBuilder pb = new ProcessBuilder(command).redirectErrorStream(true);
if (extraEnv != null && !extraEnv.isEmpty()) {
pb.environment().putAll(extraEnv);
}
p = pb.start();
} catch (IOException e) {
throw new WorktreeException("failed to start " + command[0] + ": " + e.getMessage(), e);
}
@@ -1,5 +1,6 @@
package dev.ltms.bridged.session;
import dev.ltms.bridged.peer.CharterReceipt;
import dev.ltms.bridged.peer.MemberRole;
/**
@@ -23,6 +24,13 @@ import dev.ltms.bridged.peer.MemberRole;
* @param lastActivityAtNanos {@link System#nanoTime()} of the most recent lifecycle event
* @param turnCount number of delegated turns that have been delivered to this session
* @param state current lifecycle state in the one-shot FSM
* @param charterReceipt the fingerprint (CB-571) of the charter bytes this member was started
* with; {@code null} for a session whose launcher recorded none
* @param agentSessionId the peer's OWN session id (CB-584) — the handle a later
* {@code resumeSessionId} spawn would pass back to resume this exact
* conversation; {@code null} when the launcher could not determine one
* (an adapter that declines {@code Capability.SESSION_RESUME}, or one
* that resolves it lazily and has not yet)
*/
public record MemberSession(
String paneId,
@@ -36,7 +44,9 @@ public record MemberSession(
int turnCount,
State state,
String worktree,
String branch) {
String branch,
CharterReceipt charterReceipt,
String agentSessionId) {
/** One-shot worker lifecycle states. */
public enum State {
@@ -48,21 +58,35 @@ public record MemberSession(
RELEASED
}
/**
* Backward-compatible shape: a session with no charter receipt and no agent session id (a test
* or a launcher before CB-571 / CB-584). A separate constructor rather than a new parameter on
* the canonical one, so existing call sites that have nothing to record keep compiling
* unchanged.
*/
public MemberSession(String paneId, String terminalId, String profile, MemberRole role,
String cwd, String ownerTerminal, long spawnedAtNanos,
long lastActivityAtNanos, int turnCount, State state,
String worktree, String branch) {
this(paneId, terminalId, profile, role, cwd, ownerTerminal, spawnedAtNanos,
lastActivityAtNanos, turnCount, state, worktree, branch, null, null);
}
/** Return a copy of this session in {@code state}. */
public MemberSession withState(State state) {
return new MemberSession(paneId, terminalId, profile, role, cwd, ownerTerminal, spawnedAtNanos,
lastActivityAtNanos, turnCount, state, worktree, branch);
lastActivityAtNanos, turnCount, state, worktree, branch, charterReceipt, agentSessionId);
}
/** Return a copy with {@code lastActivityAtNanos} updated to {@code nowNanos}. */
public MemberSession withActivity(long nowNanos) {
return new MemberSession(paneId, terminalId, profile, role, cwd, ownerTerminal, spawnedAtNanos,
nowNanos, turnCount, state, worktree, branch);
nowNanos, turnCount, state, worktree, branch, charterReceipt, agentSessionId);
}
/** Return a copy with the turn count incremented and activity timestamped at {@code nowNanos}. */
public MemberSession bumpTurn(long nowNanos) {
return new MemberSession(paneId, terminalId, profile, role, cwd, ownerTerminal, spawnedAtNanos,
nowNanos, turnCount + 1, state, worktree, branch);
nowNanos, turnCount + 1, state, worktree, branch, charterReceipt, agentSessionId);
}
}
@@ -4,6 +4,8 @@ import dev.ltms.bridged.auth.MemberLifecycle;
import dev.ltms.bridged.herdr.Agent;
import dev.ltms.bridged.inject.TurnListener;
import dev.ltms.bridged.inject.MemberPresence;
import dev.ltms.bridged.msg.TurnToken;
import dev.ltms.bridged.peer.Capability;
import dev.ltms.bridged.peer.MemberRole;
import dev.ltms.bridged.peer.PeerHandle;
import dev.ltms.bridged.peer.PeerLauncher;
@@ -16,6 +18,7 @@ import java.util.LinkedHashMap;
import java.util.List;
import java.util.Map;
import java.util.Optional;
import java.util.Set;
import java.util.concurrent.ConcurrentHashMap;
import java.util.concurrent.TimeUnit;
import java.util.concurrent.atomic.AtomicLong;
@@ -50,11 +53,19 @@ public final class SessionManager implements TurnListener {
private final int contextCap;
private final boolean clearAfterTurn;
private volatile MemberLifecycle memberLifecycle = MemberLifecycle.NONE;
/**
* CB-586: the repo root the fleet actually works in, remembered the first time a worktree
* session is spawned (worktrees are checkouts of it). {@code refs/wip/*} live there, and this
* single cached value is what the snapshot retention sweep and the operator-visible census run
* against. The daemon is bridged into one project at a time, so "the first worktree's repo" is
* the repo; {@code null} until any worktree is spawned, meaning nothing to sweep or measure.
*/
private volatile String fleetRepoRoot;
/** CB-520: notified with a terminalId on every acquire; no-op until wired. */
private final List<Consumer<String>> acquireListeners = new java.util.concurrent.CopyOnWriteArrayList<>();
/** CB-516: notified with a terminalId on every release; no-op until wired. */
private final List<Consumer<String>> releaseListeners = new java.util.concurrent.CopyOnWriteArrayList<>();
/** CB-516: notified with a {@link ReleaseDetail} on every release; no-op until wired. */
private final List<Consumer<ReleaseDetail>> releaseListeners = new java.util.concurrent.CopyOnWriteArrayList<>();
/** Backward-compatible constructor: shared-tree sessions, production git seam. */
public SessionManager(PeerLauncher launcher) {
@@ -125,21 +136,49 @@ public final class SessionManager implements TurnListener {
}
/**
* Spawn a member, optionally inside a fresh git worktree. When {@code wt} is non-null the
* worktree is provisioned, parity-overlaid, and its path becomes the member's cwd. On any
* failure before registration the worktree is removed so no dangling checkout is left.
* Spawn a member, optionally inside a fresh git worktree, with no session identity requested.
* Equivalent to {@link #acquire(String, MemberRole, String, String, String, WorktreeRequest,
* String, String)} with both trailing args {@code null}.
*
* @param profile which backend to run on — a {@code profiles:} key
* @param role which contract the member runs under; never {@code null}
*/
public MemberSession acquire(String profile, MemberRole role, String requestedCwd, String callerCwd,
String ownerTerminal, WorktreeRequest wt) {
return acquire(profile, role, requestedCwd, callerCwd, ownerTerminal, wt, null, null);
}
/**
* Spawn a member, optionally inside a fresh git worktree, optionally onto a chosen or resumed
* agent session (CB-547a / CB-584). When {@code wt} is non-null the worktree is provisioned,
* parity-overlaid, and its path becomes the member's cwd. On any failure before registration the
* worktree is removed so no dangling checkout is left.
*
* <p>A non-blank {@code resumeSessionId} requires an explicit {@code profile}: a resumed
* conversation is tied to the specific backend that started it, so an unqualified spawn (whose
* backend a placement policy picks at spawn time) has no safe candidate to check the capability
* against. It also requires that profile's adapter to declare {@link Capability#SESSION_RESUME};
* refusing rather than silently delivering a cold session on a non-supporting adapter is the
* whole point of checking before spawn, not after (CB-584).
*
* @param profile which backend to run on — a {@code profiles:} key
* @param role which contract the member runs under; never {@code null}
* @param sessionName the bridge's logical name for the session, or {@code null}
* @param resumeSessionId the prior agent session to resume, or {@code null} for a fresh one
* @throws IllegalArgumentException if {@code resumeSessionId} is set with no explicit profile,
* or the resolved profile's adapter lacks
* {@link Capability#SESSION_RESUME}
*/
public MemberSession acquire(String profile, MemberRole role, String requestedCwd, String callerCwd,
String ownerTerminal, WorktreeRequest wt,
String sessionName, String resumeSessionId) {
MemberRole memberRole = (role == null) ? MemberRole.DEV : role;
requireResumeCapability(profile, resumeSessionId);
if (wt == null) {
// CB-557: the role must ride on the SpawnRequest, not stay a local. The launcher needs it
// to pick the profile out of that role's pool and to label the tab; a role kept only on
// the MemberSession is recorded after the spawn it was supposed to steer.
SpawnRequest req = new SpawnRequest(profile, requestedCwd, callerCwd, null, null, memberRole);
SpawnRequest req = new SpawnRequest(profile, requestedCwd, callerCwd, sessionName, resumeSessionId, memberRole);
PeerHandle handle;
try {
handle = launcher.spawn(req);
@@ -162,7 +201,9 @@ public final class SessionManager implements TurnListener {
0,
MemberSession.State.SPAWNING,
null,
null);
null,
handle.charterReceipt(),
handle.agentSessionId());
registry.put(handle.id(), session);
memberLifecycle.acquired(session.role(), session.profile(), session.terminalId());
log.debug("acquired session id={} terminal={} profile={} owner={}",
@@ -170,7 +211,30 @@ public final class SessionManager implements TurnListener {
notifyAcquired(session.terminalId());
return session;
}
return acquireWithWorktree(profile, memberRole, requestedCwd, callerCwd, ownerTerminal, wt);
return acquireWithWorktree(profile, memberRole, requestedCwd, callerCwd, ownerTerminal, wt,
sessionName, resumeSessionId);
}
/**
* Refuse a {@code resumeSessionId} that either names no explicit profile or names one whose
* adapter does not declare {@link Capability#SESSION_RESUME}. A no-op when
* {@code resumeSessionId} is blank — the ordinary, no-identity spawn path.
*/
private void requireResumeCapability(String profile, String resumeSessionId) {
if (resumeSessionId == null || resumeSessionId.isBlank()) {
return;
}
if (profile == null || profile.isBlank()) {
throw new IllegalArgumentException("resumeSessionId requires an explicit profile — "
+ "a resumed conversation is tied to the backend that started it, so it cannot "
+ "be left to placement to pick");
}
Set<Capability> caps = launcher.capabilitiesFor(profile);
if (!caps.contains(Capability.SESSION_RESUME)) {
throw new IllegalArgumentException("worker profile '" + profile + "' does not declare "
+ "Capability.SESSION_RESUME — refusing resumeSessionId rather than silently "
+ "starting a cold session");
}
}
/** Tear a worker down by pane id and remove it from the registry. Idempotent. */
@@ -194,24 +258,107 @@ public final class SessionManager implements TurnListener {
private void release(String paneId, ReleaseCause cause) {
MemberSession removed = registry.remove(paneId);
boolean preserveWorktree = cause == ReleaseCause.SHUTDOWN;
String snapshotRef = null;
if (removed != null) {
memberLifecycle.released(removed.terminalId());
log.debug("releasing session pane={} terminal={} state={} cause={}",
removed.paneId(), removed.terminalId(), removed.state(), cause);
if (preserveWorktree && removed.worktree() != null) {
logPreservedForShutdown(removed);
try {
memberLifecycle.released(removed.terminalId());
log.debug("releasing session pane={} terminal={} state={} cause={}",
removed.paneId(), removed.terminalId(), removed.state(), cause);
boolean dirty = removed.worktree() != null && worktrees.hasUncommitted(removed.worktree());
if (preserveWorktree && removed.worktree() != null) {
logPreservedForShutdown(removed);
} else if (dirty) {
// CB-576: a release that would otherwise remove the worktree finds it holding
// uncommitted work the bridge cannot see. A worker that ends a turn without
// committing (normally because it stopped to ask a question or refused the turn)
// has its only copy of that work in the worktree. Remove would --force-delete it,
// so preserve the directory and tell an operator where to find it.
preserveWorktree = true;
log.warn("release {} preserves dirty worktree {} for pane={} terminal={}: "
+ "the worktree holds uncommitted changes that --force remove would destroy",
cause, removed.worktree(), removed.paneId(), removed.terminalId());
}
if (dirty) {
// CB-578 stage C: preserving on disk is not saving — the directory is one
// `worktree remove --force`, or an operator tidying up, away from gone. Commit
// its full state to a ref before the preserve-or-remove decision above can be
// undone by anything else, regardless of why this release fired.
snapshotRef = trySnapshot(removed, cause);
}
} catch (RuntimeException e) {
// CB-581: hasUncommitted shells out to `git status` and can throw on a non-zero
// exit. We can no longer tell whether the worktree holds uncommitted work, so fail
// toward the safe answer and preserve it — deleting on a guess can destroy work
// that has no other copy (CB-576), while keeping it on a false alarm only costs
// disk. The exception must not propagate: the pane still has to stop below.
preserveWorktree = true;
log.warn("release {} could not tell whether worktree {} for pane={} terminal={} has "
+ "uncommitted changes; preserving it rather than risk destroying unsaved work: {}",
cause, removed.worktree(), removed.paneId(), removed.terminalId(), e.toString());
} finally {
// CB-516/CB-581: a send still waiting on this worker can never be answered now, no
// matter what happened above. Tell the listener BEFORE the pane is torn down, so a
// blocked caller fails fast with a real reason instead of sitting on a rendezvous
// nothing will ever resolve. CB-578 stage C: carry the worktree/branch/snapshot ref
// too, so a failed ticket's detail can point a lead at the same tree to re-dispatch.
notifyReleased(new ReleaseDetail(removed.terminalId(), removed.worktree(),
removed.branch(), snapshotRef));
}
// CB-516: a send still waiting on this worker can never be answered now. Tell the
// listener BEFORE the pane is torn down, so a blocked caller fails fast with a real
// reason instead of sitting on a rendezvous nothing will ever resolve.
notifyReleased(removed.terminalId());
}
// CB-581: the pane must always stop, even if the dirty check above threw. A session removed
// from the registry with no pane stop is an orphaned pane — a live terminal burning a fleet
// slot that no longer appears in the roster and can never be reclaimed.
launcher.stop(paneId);
if (removed != null && !preserveWorktree && removed.worktree() != null) {
worktrees.remove(worktrees.repoRoot(removed.cwd()), removed.worktree());
}
}
/**
* Best-effort snapshot of a dirty worktree into {@code refs/wip/<branch>} (CB-578 stage C). A
* failure here must never escalate: the caller has already decided to preserve the worktree
* regardless of whether this succeeds, so the only cost of a failed snapshot is a WARN and a
* missing ref — never a lost pane stop or a lost release notification.
*/
private String trySnapshot(MemberSession session, ReleaseCause cause) {
if (session.worktree() == null || session.branch() == null) {
return null;
}
try {
Optional<String> ref = worktrees.snapshot(session.worktree(), session.branch(),
snapshotMessage(session, cause));
ref.ifPresent(sha -> log.info(
"snapshotted dirty worktree {} to refs/wip/{} commit={} for pane={} terminal={}",
session.worktree(), session.branch(), sha, session.paneId(), session.terminalId()));
return ref.orElse(null);
} catch (RuntimeException e) {
log.warn("snapshot of dirty worktree {} failed for pane={} terminal={} branch={}: the "
+ "worktree is still preserved on disk, just not committed to refs/wip/{}: {}",
session.worktree(), session.paneId(), session.terminalId(), session.branch(),
session.branch(), e.toString());
return null;
}
}
/** Commit message for a CB-578 stage C snapshot — names the member so an operator can tell runs apart. */
private String snapshotMessage(MemberSession session, ReleaseCause cause) {
return "CB-578 stage C: snapshot of a released worker\n\n"
+ "terminal: " + session.terminalId() + "\n"
+ "profile: " + session.profile() + "\n"
+ "branch: " + session.branch() + "\n"
+ "cause: " + cause;
}
/**
* Facts about a released session that a listener needs beyond the bare terminal id — enough
* for a caller to point a lead at where to re-dispatch onto the same tree after a failed
* release (CB-578 stage C, acceptance criterion 10). {@code worktreePath} and {@code branch}
* are {@code null} for a shared-tree session; {@code snapshotRef} is {@code null} unless this
* release snapshotted a dirty worktree into {@code refs/wip/<branch>}.
*/
public record ReleaseDetail(String terminalId, String worktreePath, String branch, String snapshotRef) {
}
/**
* Why a session is being released — governs whether its worktree is preserved or removed.
* Worktree removal is reserved for the one case that is genuinely finished; everything else
@@ -260,7 +407,7 @@ public final class SessionManager implements TurnListener {
* presence view this manager exposes). Wiring it at construction would require breaking that
* cycle for one callback.
*/
public void onRelease(Consumer<String> listener) {
public void onRelease(Consumer<ReleaseDetail> listener) {
if (listener != null) {
releaseListeners.add(listener);
}
@@ -286,21 +433,22 @@ public final class SessionManager implements TurnListener {
}
/** A listener failure must never prevent the teardown it is reacting to. */
private void notifyReleased(String terminalId) {
if (terminalId == null) {
private void notifyReleased(ReleaseDetail detail) {
if (detail.terminalId() == null) {
return;
}
for (Consumer<String> listener : releaseListeners) {
for (Consumer<ReleaseDetail> listener : releaseListeners) {
try {
listener.accept(terminalId);
listener.accept(detail);
} catch (RuntimeException e) {
log.warn("release listener failed for terminal {}: {}", terminalId, e.toString());
log.warn("release listener failed for terminal {}: {}", detail.terminalId(), e.toString());
}
}
}
private MemberSession acquireWithWorktree(String profile, MemberRole memberRole, String requestedCwd, String callerCwd,
String ownerTerminal, WorktreeRequest wt) {
String ownerTerminal, WorktreeRequest wt,
String sessionName, String resumeSessionId) {
String preResolvedProfile = (profile == null || profile.isBlank())
? launcher.defaultProfile() : profile;
// CB-507: resolve through the launcher's CB-112 chain (requested → profile cwd → caller →
@@ -310,13 +458,18 @@ public final class SessionManager implements TurnListener {
// The non-worktree path always used this chain; only this branch was missed.
String repoRoot = worktrees.repoRoot(
launcher.effectiveCwd(new SpawnRequest(preResolvedProfile, requestedCwd, callerCwd)));
if (fleetRepoRoot == null) {
// CB-586: remember the repo whose worktrees the fleet spawns — its refs/wip/* are the
// snapshot store the retention sweep and the operator census operate on.
fleetRepoRoot = repoRoot;
}
String branch = "worker/" + slug(wt.ticketSlug()) + "-" + nonce();
String path = null;
PeerHandle handle;
try {
path = worktrees.add(repoRoot, branch, wt.baseRef());
worktrees.overlayParity(repoRoot, path, launcher.parityOverlay(preResolvedProfile));
handle = launcher.spawn(new SpawnRequest(profile, path, callerCwd, null, null, memberRole));
handle = launcher.spawn(new SpawnRequest(profile, path, callerCwd, sessionName, resumeSessionId, memberRole));
} catch (RuntimeException e) {
log.warn("spawn failed for profile={} role={} branch={} path={}: {}",
preResolvedProfile, memberRole, branch, path, e.getMessage());
@@ -344,7 +497,9 @@ public final class SessionManager implements TurnListener {
0,
MemberSession.State.SPAWNING,
path,
branch);
branch,
handle.charterReceipt(),
handle.agentSessionId());
registry.put(handle.id(), session);
memberLifecycle.acquired(session.role(), session.profile(), session.terminalId());
log.debug("acquired worktree session id={} terminal={} profile={} branch={} path={}",
@@ -410,6 +565,22 @@ public final class SessionManager implements TurnListener {
if (session.ownerTerminal() != null) {
m.put("owner", session.ownerTerminal());
}
// CB-584: which conversation this member holds — the id a later resumeSessionId spawn would
// pass back. Absent for an adapter that declines Capability.SESSION_RESUME, or one that
// resolves it lazily and has not yet.
if (session.agentSessionId() != null) {
m.put("agentSessionId", session.agentSessionId());
}
// CB-571: which charter this member was started with — never the charter text itself. The
// digest lets a lead tell at a glance whether all members got the same charter; the source
// records whether a role charter was configured ("fleet.charters.<role>") or only the reply
// charter was composed ("none").
if (session.charterReceipt() != null) {
m.put("charterSource", session.charterReceipt().charterSource());
if (session.charterReceipt().charterSha256() != null) {
m.put("charterSha256", session.charterReceipt().charterSha256());
}
}
m.put("liveStatus", live == null ? "unknown" : live.status().name().toLowerCase());
return m;
}
@@ -425,7 +596,7 @@ public final class SessionManager implements TurnListener {
* can be re-delivered for multi-turn reuse until it is released.
*/
@Override
public void onDelivered(String target) {
public void onDelivered(String target, TurnToken token) {
MemberSession current = findByTerminal(target);
if (current == null) return;
if (current.state() != MemberSession.State.READY && current.state() != MemberSession.State.DONE) {
@@ -516,8 +687,15 @@ public final class SessionManager implements TurnListener {
log.debug("reaping idle session terminal={} pane={}: idle {}s exceeds the {}s ttl",
s.terminalId(), s.paneId(), TimeUnit.NANOSECONDS.toSeconds(idleNanos),
TimeUnit.NANOSECONDS.toSeconds(idleTtlNanos));
release(s.paneId());
reaped++;
// CB-581: one session that fails to release must not abort the whole reaping pass —
// match drainAll's per-session try/catch so the rest of the roster still gets reaped.
try {
release(s.paneId());
reaped++;
} catch (RuntimeException e) {
log.warn("reap failed for pane={} terminal={} worktree={}; continuing with "
+ "remaining sessions", s.paneId(), s.terminalId(), s.worktree(), e);
}
}
}
return reaped;
@@ -575,6 +753,26 @@ public final class SessionManager implements TurnListener {
return registry.size();
}
/**
* CB-586: the operator-visible census of {@code refs/wip/*} in the repo the fleet works in —
* how many snapshot refs exist and roughly what they cost. Empty (no repo known) until at
* least one worktree session has been spawned, exactly so a fleet that has never snapshotted
* anything surfaces nothing new, as it did before CB-586.
*/
public Optional<Worktrees.WipRefStats> wipRefs() {
String repo = fleetRepoRoot;
return repo == null ? Optional.empty() : Optional.of(worktrees.wipRefs(repo));
}
/**
* CB-586: run the snapshot retention sweep in the fleet's repo (a no-op until a worktree has
* been spawned, which establishes the repo). Returns how many {@code refs/wip/*} it deleted.
*/
public int sweepWipRefs(long minAgeMillis) {
String repo = fleetRepoRoot;
return repo == null ? 0 : worktrees.pruneWipRefs(repo, minAgeMillis);
}
/**
* The registered session owning {@code terminalId}, or {@code null} if none does.
*
@@ -14,12 +14,20 @@ public final class SessionReaper {
private static final Logger log = LoggerFactory.getLogger(SessionReaper.class);
private static final long DEFAULT_INTERVAL_MILLIS = 5000;
/** CB-586: the refs/wip age floor — never sweep a snapshot younger than 24h (the CB-586 rule). */
private static final long WIP_MIN_AGE_MILLIS = TimeUnit.HOURS.toMillis(24);
/**
* CB-586: how often the retention sweep runs. Given the 24h age floor, running it every few
* hours means a ref is dropped within hours of becoming eligible, never within minutes.
*/
private static final long WIP_SWEEP_INTERVAL_NANOS = TimeUnit.HOURS.toNanos(6);
private final SessionManager sessions;
private final long idleTtlNanos;
private final long intervalMillis;
private volatile boolean running;
private Thread thread;
private volatile long lastWipSweepNanos = Long.MIN_VALUE;
/** Construct a reaper with the default 5-second polling interval. */
public SessionReaper(SessionManager sessions, long idleTtlSeconds) {
@@ -49,10 +57,32 @@ public final class SessionReaper {
} catch (RuntimeException e) {
log.warn("session reaper iteration failed; continuing", e);
}
maybeSweepWipRefs();
sleep();
}
}
/**
* CB-586: run the refs/wip retention sweep on a slow cadence (hours, not the per-iteration
* millisecond loop). Best-effort — a failure must never take the idle-reap loop down with it.
*/
private void maybeSweepWipRefs() {
long now = System.nanoTime();
if (now - lastWipSweepNanos < WIP_SWEEP_INTERVAL_NANOS) {
return;
}
try {
int deleted = sessions.sweepWipRefs(WIP_MIN_AGE_MILLIS);
if (deleted > 0) {
log.info("refs/wip retention sweep deleted {} snapshot ref(s) older than 24h whose "
+ "content was already reachable from main", deleted);
}
} catch (RuntimeException e) {
log.warn("refs/wip retention sweep failed; continuing", e);
}
lastWipSweepNanos = now;
}
private void sleep() {
try {
Thread.sleep(intervalMillis);
@@ -1,6 +1,7 @@
package dev.ltms.bridged.session;
import java.util.List;
import java.util.Optional;
/** Seam between {@link SessionManager} and git worktree operations. Tests use a recording fake. */
public interface Worktrees {
@@ -10,9 +11,91 @@ public interface Worktrees {
/** git -C <repoRoot> worktree remove --force <path>. Idempotent (already-gone tolerated). */
void remove(String repoRoot, String worktreePath);
/**
* True when the worktree holds uncommitted changes the bridge cannot see: tracked
* modifications, staged files, or untracked files. {@code git status --porcelain} is the
* test; an empty result means clean. Callers use this to decide whether removing the
* worktree would silently destroy a worker's only copy of its work.
*
* <p>An already-gone worktree is reported as clean (no throw), matching {@link #remove}'s
* idempotent contract: a path that does not exist holds no work to lose, and must not break
* a teardown that is mid-way through stopping the pane.
*/
boolean hasUncommitted(String worktreePath);
/** Copy each existing overlay path repoRoot→worktree; mark tracked ones --skip-worktree. */
void overlayParity(String repoRoot, String worktreePath, List<String> overlay);
/** git -C <cwd> rev-parse --show-toplevel — the repo root that owns cwd. */
String repoRoot(String cwd);
/**
* Commit the worktree's full on-disk state — tracked and untracked, respecting
* {@code .gitignore} — to {@code refs/wip/<branch>}, so a release that would otherwise leave
* the work as loose, unprotected files has a durable git object to fall back on (CB-578 stage
* C). Built on a <em>temporary</em> index: the worker's own index, working tree, and HEAD are
* never touched, since the worker may still be mid-write. The commit is parented on the
* worktree's current HEAD.
*
* <p>Never writes under {@code refs/heads/} — the ref must not appear in {@code git branch},
* must not be pushed by default, and must not be swept by a later {@code git branch -d}.
*
* <p>Callers are expected to have already confirmed {@link #hasUncommitted} before reaching
* for this; it always stages and commits whatever {@code git add -A} finds, so calling it on
* a clean worktree still produces a (harmless, tree-identical-to-HEAD) commit rather than
* detecting cleanliness itself.
*
* @param worktreePath absolute path of the worktree to snapshot
* @param branch the worktree's own branch — keys {@code refs/wip/<branch>}
* @param message the commit message; should name the member, its branch, and the release
* cause so an operator can tell which run produced it
* @return the created commit's sha, or {@link Optional#empty()} if {@code worktreePath} does
* not exist (mirrors {@link #remove} and {@link #hasUncommitted}'s already-gone
* tolerance — a worktree that is gone holds nothing to snapshot)
*/
Optional<String> snapshot(String worktreePath, String branch, String message);
/**
* CB-586: how many {@code refs/wip/*} snapshot refs exist in {@code repoRoot} and roughly what
* they cost. This is the operator-visible surface for the snapshot growth CB-578 stage C left
* behind — counts of refs alone hide that each one pins a whole tree for {@code git gc}.
*
* @param repoRoot the repository to scan
* @return count of snapshot refs, and {@code costBytes} = the approximate total working-tree
* size of every snapshot's committed content (summed per ref, so shared objects are
* counted once per ref that carries them)
*/
WipRefStats wipRefs(String repoRoot);
/**
* CB-586: run the {@code refs/wip/*} retention sweep and return how many refs it deleted.
*
* <p>The retention rule is <em>reachability plus an age floor</em>. A snapshot ref is deleted
* only when <strong>both</strong> hold:
* <ol>
* <li>its commit's <em>tree content</em> is already reachable from {@code main} — the work
* the snapshot preserved has been recovered, so dropping the ref loses nothing; and</li>
* <li>the ref is older than {@code minAgeMillis} — a very recent snapshot is never swept
* while a lead may still be looking at it.</li>
* </ol>
*
* <p>Reachability is the safety property. A snapshot exists precisely because the work was not
* committed anywhere else, so a snapshot whose content is <em>not</em> reachable from
* {@code main} is the <strong>last copy</strong> of a worker's work and must never be deleted
* automatically — that is the failure CB-576 and CB-578 stage C were built to stop. Age alone
* must never drive a deletion, because age-based sweeping is exactly how the last copy gets
* destroyed. (Both numbers and the rule are CB-586's decision; this method only implements it.)
*
* <p>Every deletion logs the ref name and the commit sha, so an operator who finds they lost
* the wrong thing can still recover it from git's reflog.
*
* @param repoRoot the repository whose {@code refs/wip/*} to sweep
* @param minAgeMillis the age floor; a ref younger than this is never touched
* @return the number of snapshot refs deleted
*/
int pruneWipRefs(String repoRoot, long minAgeMillis);
/** CB-586: the operator-visible census of {@code refs/wip/*} in one repository. */
record WipRefStats(int count, long costBytes) {
}
}
+9
View File
@@ -1,4 +1,13 @@
<configuration>
<!--
CB-575: MCP SDK 2.0.0 has no public notification registration API. It registers only
notifications/initialized and notifications/roots/list_changed, so other client notifications
still warn when unhandled. Clients may legitimately send notifications/cancelled; suppress only
that SDK WARN because it would devalue the action-needed WARN level used by M4 fleet health.
If a later SDK handles cancellation, this filter simply stops matching and can be removed.
-->
<turboFilter class="dev.ltms.bridged.logging.McpCancelledNotificationFilter"/>
<appender name="STDOUT" class="ch.qos.logback.core.ConsoleAppender">
<encoder>
<pattern>%d{HH:mm:ss.SSS} %-5level [%thread] %logger{28} - %msg%n</pattern>
@@ -194,7 +194,7 @@ class BridgedConfigTest {
fleet:
leaders:
opus:
terminal: term_opus
tab: "lead: opus"
""");
BridgedConfig.Leader lead = BridgedConfig.load(f).fleet().leaders().get("opus");
@@ -212,7 +212,7 @@ class BridgedConfigTest {
fleet:
leaders:
opus:
terminal: term_opus
tab: "drive: opus"
tabPrefix: "drive:"
scanIntervalSeconds: 30
""");
@@ -224,7 +224,9 @@ class BridgedConfigTest {
/**
* The pane no longer has to exist before the daemon does (CB-557): a lead naming a profile may
* be launched, while one that names only a terminal is recognised and never created.
* be launched, while one that names no profile is recognised and never created. Either way it
* still needs its own {@code tab:} (CB-579) — that part is unconditional, see
* {@link #aLeadWithNoTabRefusesToStart}.
*/
@Test
void aLeadIsCreatableOnlyWhenItNamesAProfile(@TempDir Path dir) throws Exception {
@@ -239,8 +241,9 @@ class BridgedConfigTest {
leaders:
launched:
profile: opus
tab: "lead: launched"
pinned:
terminal: term_opus
tab: "lead: pinned"
""");
var leaders = BridgedConfig.load(f).fleet().leaders();
@@ -249,8 +252,12 @@ class BridgedConfigTest {
"no profile to launch on ⇒ recognise-only, the pre-CB-557 behaviour");
}
/**
* CB-579: {@code tab} is the only field a lead's identity depends on now, so it is required
* whether the entry is creatable or recognise-only — without it the entry can never be found.
*/
@Test
void aLeadThatCanBeNeitherFoundNorCreatedRefusesToStart(@TempDir Path dir) throws Exception {
void aLeadWithNoTabRefusesToStart(@TempDir Path dir) throws Exception {
Path f = dir.resolve("useless-lead.yaml");
Files.writeString(f, """
bind:
@@ -264,6 +271,80 @@ class BridgedConfigTest {
IllegalStateException e = assertThrows(IllegalStateException.class, cfg::validateMembers);
assertTrue(e.getMessage().contains("ghost"), "the message must name the useless entry");
assertTrue(e.getMessage().contains("tab:"), "the message must say what is missing");
}
/**
* CB-579 acceptance (2): a config still spelling {@code fleet.leaders.<name>.terminal} must fail
* loudly at load, not be silently dropped by {@code Leader}'s {@code @JsonIgnoreProperties}.
*/
@Test
void aLeaderTerminalKeyFailsLoadAndNamesTabAsTheReplacement(@TempDir Path dir) throws Exception {
Path f = dir.resolve("stale-terminal.yaml");
Files.writeString(f, """
bind:
port: 8080
fleet:
leaders:
opus:
terminal: term_opus
""");
IllegalStateException e =
assertThrows(IllegalStateException.class, () -> BridgedConfig.load(f));
assertTrue(e.getMessage().contains("opus"), "the message must name the offending entry");
assertTrue(e.getMessage().contains("tab:"), "the message must name the replacement key");
assertTrue(e.getMessage().contains("terminal"), "the message must name the retired key");
}
/** The same refusal, and it must name every offending entry, not just the first. */
@Test
void everyLeaderStillUsingTerminalIsReportedAtOnce(@TempDir Path dir) throws Exception {
Path f = dir.resolve("stale-terminals.yaml");
Files.writeString(f, """
bind:
port: 8080
fleet:
leaders:
opus:
terminal: term_opus
sol:
terminal: term_sol
""");
IllegalStateException e =
assertThrows(IllegalStateException.class, () -> BridgedConfig.load(f));
assertTrue(e.getMessage().contains("opus"));
assertTrue(e.getMessage().contains("sol"));
}
/** A {@code terminal:} anywhere else in the document (not under a leader entry) is unaffected. */
@Test
void aTerminalKeyOutsideFleetLeadersIsNotRejected(@TempDir Path dir) throws Exception {
Path f = dir.resolve("primary-terminal-ok.yaml");
Files.writeString(f, "bind:\n port: 8080\nprimary:\n terminal: term_fixed\n");
assertDoesNotThrow(() -> BridgedConfig.load(f));
}
/** CB-579 acceptance (3): distinct `tab:` labels need no shared prefix — one scanner finds both. */
@Test
void twoLeadersWithDifferentTabsAreBothConfigured(@TempDir Path dir) throws Exception {
Path f = dir.resolve("two-tabs.yaml");
Files.writeString(f, """
bind:
port: 8080
fleet:
leaders:
opus:
tab: "lead: opus"
sol:
tab: "captain: sol"
""");
var leaders = BridgedConfig.load(f).fleet().leaders();
assertEquals("lead: opus", leaders.get("opus").tab());
assertEquals("captain: sol", leaders.get("sol").tab());
}
// ── CB-551: the idle-lead heartbeat ─────────────────────────────────────────────────────────
@@ -328,7 +409,7 @@ class BridgedConfigTest {
fleet:
leaders:
opus:
terminal: term_opus
tab: "lead: opus"
tabPrefix: "lead:"
""");
BridgedConfig cfg = BridgedConfig.load(f);
@@ -349,7 +430,7 @@ class BridgedConfigTest {
tabLabel: "lead: {role} {profile}"
leaders:
opus:
terminal: term_opus
tab: "lead: opus"
""");
BridgedConfig cfg = BridgedConfig.load(f);
@@ -374,7 +455,7 @@ class BridgedConfigTest {
fleet:
leaders:
opus:
terminal: term_opus
tab: "lead: opus"
""");
assertDoesNotThrow(() -> BridgedConfig.load(f).validateLeadTabPrefixes());
@@ -400,7 +481,7 @@ class BridgedConfigTest {
"a label that collides with a convention nobody reads is not a problem");
}
// ── CB-530: the leaders registry ────────────────────────────────────────────────────────────
// ── CB-530/CB-579: the leaders registry ─────────────────────────────────────────────────────
@Test
void leadersBlockRegistersEveryPaneByName(@TempDir Path dir) throws Exception {
@@ -411,10 +492,10 @@ class BridgedConfigTest {
fleet:
leaders:
opus-5.0:
terminal: term_opus
tab: "lead: opus-5.0"
kind: claude
gpt-sol-5.6:
terminal: term_sol
tab: "lead: gpt-sol-5.6"
kind: opencode
model: openai/gpt-5.6-terra
""");
@@ -425,9 +506,9 @@ class BridgedConfigTest {
assertEquals(Set.of("opus-5.0", "gpt-sol-5.6"), leaders.keySet());
assertEquals("opencode", leaders.get("gpt-sol-5.6").kind());
assertEquals("openai/gpt-5.6-terra", leaders.get("gpt-sol-5.6").model());
// The whole point: BOTH panes resolve as leads, so neither is demoted to worker.
assertEquals(Map.of("term_opus", "opus-5.0", "term_sol", "gpt-sol-5.6"),
cfg.leaderTerminals());
// Identity is the tab now (CB-579) — both entries carry their own, distinct label.
assertEquals("lead: opus-5.0", leaders.get("opus-5.0").tab());
assertEquals("lead: gpt-sol-5.6", leaders.get("gpt-sol-5.6").tab());
}
@Test
@@ -439,42 +520,6 @@ class BridgedConfigTest {
"configs that never migrate must behave exactly as they did before CB-530");
}
@Test
void anExplicitLeadersEntryWinsOverThePinForTheSameTerminal(@TempDir Path dir) throws Exception {
Path f = dir.resolve("both.yaml");
Files.writeString(f, """
bind:
port: 8080
primary:
terminal: term_shared
fleet:
leaders:
opus-5.0:
terminal: term_shared
""");
assertEquals(Map.of("term_shared", "opus-5.0"), BridgedConfig.load(f).leaderTerminals(),
"the pin is the older spelling of the same fact; the named entry is what was meant");
}
@Test
void bothBlocksTogetherRegisterTheUnionOfTheirTerminals(@TempDir Path dir) throws Exception {
Path f = dir.resolve("union.yaml");
Files.writeString(f, """
bind:
port: 8080
primary:
terminal: term_pinned
fleet:
leaders:
gpt-sol-5.6:
terminal: term_sol
""");
assertEquals(Map.of("term_pinned", "primary", "term_sol", "gpt-sol-5.6"),
BridgedConfig.load(f).leaderTerminals());
}
@Test
void neitherBlockLeavesNothingRegistered(@TempDir Path dir) throws Exception {
Path f = dir.resolve("none.yaml");
@@ -483,22 +528,25 @@ class BridgedConfigTest {
assertTrue(BridgedConfig.load(f).leaderTerminals().isEmpty());
}
/** A lead entry with no terminal identifies nothing — it must not register a null key. */
/**
* CB-579: {@code fleet.leaders} no longer feeds {@code leaderTerminals()} at all — a lead's
* identity comes from the live tab scan, not a config-held terminal map. This method now exists
* only for the {@code primary.terminal} fallback.
*/
@Test
void aLeadWithoutATerminalIsNotRegistered(@TempDir Path dir) throws Exception {
Path f = dir.resolve("no-terminal.yaml");
void fleetLeadersNeverContributesToLeaderTerminals(@TempDir Path dir) throws Exception {
Path f = dir.resolve("leaders-only.yaml");
Files.writeString(f, """
bind:
port: 8080
fleet:
leaders:
sketch:
kind: opencode
real:
terminal: term_real
opus-5.0:
tab: "lead: opus-5.0"
""");
assertEquals(Map.of("term_real", "real"), BridgedConfig.load(f).leaderTerminals());
assertTrue(BridgedConfig.load(f).leaderTerminals().isEmpty(),
"no primary.terminal pin ⇒ nothing registered, even with fleet.leaders configured");
}
// ── CB-548: the architects registry ────────────────────────────────────────────────────────
@@ -1092,6 +1140,7 @@ class BridgedConfigTest {
gitHostEnv: GITEA_HOST
weight: 0.5
maxLoad: 2
exhaustedPattern: "usage limit has been reached"
placement: weighted
lifecycle:
idleTtlSeconds: 300
@@ -1123,6 +1172,8 @@ class BridgedConfigTest {
assertEquals("GITEA_HOST", w.gitHostEnv());
assertEquals(0.5f, w.weight(), 0.0001f, "weight binds as a float");
assertEquals(2, w.maxLoad(), "maxLoad binds as an integer");
assertTrue(w.hasExhaustedPattern(), "exhaustedPattern binds and enables the CB-578 stage A classification");
assertEquals("usage limit has been reached", w.exhaustedPattern());
assertEquals("weighted", cfg.placement(), "placement binds at the top level");
assertEquals(300, cfg.lifecycle().idleTtlSeconds());
@@ -1170,6 +1221,80 @@ class BridgedConfigTest {
assertNull(w.maxLoad(), "absent maxLoad defaults to unlimited (null)");
}
@Test
void explicitWeightZeroStaysZeroInsteadOfCoercingToOne(@TempDir Path dir) throws Exception {
// CB-554: weight: 0 used to be normalised to 1.0 by this same compact constructor, which
// made it read as "never auto-select" while behaving as "select like anyone else."
Path f = dir.resolve("weight-zero.yaml");
Files.writeString(f, """
bind:
port: 8080
profiles:
opus:
baseUrl: http://gx10.gw:8000
weight: 0
""");
BridgedConfig cfg = BridgedConfig.load(f);
BridgedConfig.Profile w = cfg.profiles().get("opus");
assertEquals(0.0f, w.weight(), 0.0001f, "an explicit weight: 0 must stay 0, not coerce to 1.0");
}
@Test
void negativeWeightNormalisesToZeroNotOne(@TempDir Path dir) throws Exception {
Path f = dir.resolve("weight-negative.yaml");
Files.writeString(f, """
bind:
port: 8080
profiles:
gx10:
baseUrl: http://gx10.gw:8000
weight: -3.0
""");
BridgedConfig cfg = BridgedConfig.load(f);
BridgedConfig.Profile w = cfg.profiles().get("gx10");
assertEquals(0.0f, w.weight(), 0.0001f, "a negative weight behaves as excluded (0), not as an error and not as 1.0");
}
@Test
void explicitMaxLoadZeroStaysZeroInsteadOfCoercingToUnlimited(@TempDir Path dir) throws Exception {
// CB-585: maxLoad: 0 used to be normalised to null (unlimited) by this same compact
// constructor, so an operator writing it to mean "never run anything here" got the exact
// opposite. On a subscription: true profile, maxLoad is the only throttle against the
// operator's own paid plan.
Path f = dir.resolve("maxload-zero.yaml");
Files.writeString(f, """
bind:
port: 8080
profiles:
opus:
baseUrl: http://gx10.gw:8000
maxLoad: 0
""");
BridgedConfig cfg = BridgedConfig.load(f);
BridgedConfig.Profile w = cfg.profiles().get("opus");
assertEquals(0, w.maxLoad(), "an explicit maxLoad: 0 must stay 0, not coerce to unlimited (null)");
}
@Test
void negativeMaxLoadIsRefusedAtLoadNamingTheProfileAndKey(@TempDir Path dir) throws Exception {
Path f = dir.resolve("maxload-negative.yaml");
Files.writeString(f, """
bind:
port: 8080
profiles:
opus:
baseUrl: http://gx10.gw:8000
maxLoad: -2
""");
IllegalStateException e = assertThrows(IllegalStateException.class, () -> BridgedConfig.load(f));
assertTrue(e.getMessage().contains("opus"), "error names the profile: " + e.getMessage());
assertTrue(e.getMessage().contains("maxLoad"), "error names the key: " + e.getMessage());
}
@Test
void subscriptionFlagBindsAndDefaultsFalse(@TempDir Path dir) throws Exception {
Path f = dir.resolve("subscription.yaml");
@@ -338,6 +338,54 @@ class ConfigRefTest {
assertEquals("deepseek-v4-flash", ref.get().profiles().get("sonnet").model());
}
/**
* CB-578 stage B: exhaustedPattern is compiled once into Bridged.main's pattern map at startup
* (see ExhaustedPatternLookup), so a reload never re-reads it — a changed pattern must be
* reported deferred exactly like model/baseUrl, not silently claimed as applied.
*/
@Test
void changingAProfilesExhaustedPatternIsReportedAsDeferred(@TempDir Path dir) throws Exception {
Path f = dir.resolve("bridged.yaml");
Files.writeString(f, """
bind:
host: 127.0.0.1
port: 8765
herdrSocket: ~/.config/herdr/herdr.sock
profiles:
sonnet:
baseUrl: http://gx00.gw:8000
model: sonnet
exhaustedPattern: "usage limit has been reached"
guard:
offSubscriptionHosts:
- gx00.gw
""");
ConfigRef ref = refFor(f);
Files.writeString(f, """
bind:
host: 127.0.0.1
port: 8765
herdrSocket: ~/.config/herdr/herdr.sock
profiles:
sonnet:
baseUrl: http://gx00.gw:8000
model: sonnet
exhaustedPattern: "rate limit exceeded"
guard:
offSubscriptionHosts:
- gx00.gw
""");
ConfigRef.Outcome out = ref.reload();
assertTrue(out.applied());
assertEquals(1, out.deferred().size(), out.deferred().toString());
assertTrue(out.deferred().getFirst().contains("sonnet"), out.deferred().toString());
assertTrue(out.deferred().getFirst().contains("launch settings"), out.deferred().toString());
// The snapshot still carries the new value — a restart is what makes it take effect.
assertEquals("rate limit exceeded", ref.get().profiles().get("sonnet").exhaustedPattern());
}
@Test
void aFixedRefHasNoFileAndRefusesToReload() {
BridgedConfig cfg = new BridgedConfig(null, null, null, null, null, null,
@@ -0,0 +1,170 @@
package dev.ltms.bridged.health;
import ch.qos.logback.classic.Logger;
import ch.qos.logback.classic.spi.ILoggingEvent;
import ch.qos.logback.core.read.ListAppender;
import dev.ltms.bridged.herdr.AgentControl;
import dev.ltms.bridged.herdr.FakeHerdr;
import dev.ltms.bridged.inject.Injector;
import dev.ltms.bridged.msg.InMemoryReplyInbox;
import dev.ltms.bridged.msg.MessageService;
import dev.ltms.bridged.msg.Rendezvous;
import dev.ltms.bridged.session.SessionManager;
import dev.ltms.bridged.member.ClaudeCodeLauncher;
import dev.ltms.bridged.herdr.WorkspaceControl;
import dev.ltms.bridged.guard.SubscriptionGuard;
import dev.ltms.bridged.config.BridgedConfig;
import org.junit.jupiter.api.Test;
import org.slf4j.LoggerFactory;
import java.util.Map;
import java.util.Set;
import java.util.concurrent.Executors;
import java.util.function.BiConsumer;
import static org.junit.jupiter.api.Assertions.assertEquals;
import static org.junit.jupiter.api.Assertions.assertTrue;
class FleetHealthMonitorTest {
@Test void oneTickUsesOneFleetListForAnyRosterSize() {
FakeHerdr herdr = new FakeHerdr();
AgentControl agents = new AgentControl(herdr);
BridgedConfig.Profile profile = new BridgedConfig.Profile("test", "http://test:1", null,
null, null, null, null, null, null, null, null, null);
ClaudeCodeLauncher launcher = new ClaudeCodeLauncher(agents, new WorkspaceControl(herdr),
new SubscriptionGuard(Set.of("test")), Map.of("test", profile), "test", _ -> "token");
SessionManager sessions = new SessionManager(launcher);
sessions.acquire("test", null, null, null);
sessions.acquire("test", null, null, null);
herdr.calls.clear();
MessageService messages = new MessageService(agents, new Injector(agents), new Rendezvous(), new InMemoryReplyInbox());
var scheduler = Executors.newSingleThreadScheduledExecutor();
FleetHealthMonitor monitor = new FleetHealthMonitor(agents, sessions::roster, messages, scheduler, () -> 1, 60,
(_, _) -> { });
monitor.tick();
monitor.stop();
assertEquals(1, herdr.calls.stream().filter(call -> call.method().equals("agent.list")).count());
}
@Test void failedTickDoesNotStopTheNextTick() {
FakeHerdr herdr = new FakeHerdr().healthy(false);
AgentControl agents = new AgentControl(herdr);
var scheduler = Executors.newSingleThreadScheduledExecutor();
FleetHealthMonitor monitor = new FleetHealthMonitor(agents, java.util.List::of,
new MessageService(agents, new Injector(agents), new Rendezvous(), new InMemoryReplyInbox()),
scheduler, () -> 1, 60, (_, _) -> { });
monitor.tick();
herdr.healthy(true);
monitor.tick();
monitor.stop();
assertEquals(2, herdr.calls.stream().filter(call -> call.method().equals("agent.list")).count());
}
@Test void faultTransitionLogsOnlyOnceUntilItChanges() {
Logger logger = (Logger) LoggerFactory.getLogger(FleetHealthMonitor.class);
ListAppender<ILoggingEvent> appender = new ListAppender<>();
appender.start();
logger.addAppender(appender);
try {
FakeHerdr herdr = new FakeHerdr();
AgentControl agents = new AgentControl(herdr);
var scheduler = Executors.newSingleThreadScheduledExecutor();
FleetHealthMonitor monitor = new FleetHealthMonitor(agents, java.util.List::of,
new MessageService(agents, new Injector(agents), new Rendezvous(), new InMemoryReplyInbox()),
scheduler, () -> 1, 60, (_, _) -> { });
monitor.reportTransition("term_a", HealthState.TURN_BOUNDARY_LOST);
monitor.reportTransition("term_a", HealthState.TURN_BOUNDARY_LOST);
monitor.stop();
assertEquals(1, appender.list.stream().filter(event -> event.getFormattedMessage()
.contains("member=term_a state=TURN_BOUNDARY_LOST")).count());
} finally {
logger.detachAppender(appender);
}
}
// --- CB-580: a member that reaches GONE/NEVER_READY must fail its waiting tickets
private static FleetHealthMonitor monitorWith(BiConsumer<String, String> failTarget) {
FakeHerdr herdr = new FakeHerdr();
AgentControl agents = new AgentControl(herdr);
var scheduler = Executors.newSingleThreadScheduledExecutor();
return new FleetHealthMonitor(agents, java.util.List::of,
new MessageService(agents, new Injector(agents), new Rendezvous(), new InMemoryReplyInbox()),
scheduler, () -> 1, 60, failTarget);
}
@Test void terminalTransitionFailsTheTargetOnce() {
RecordingFailTarget failTarget = new RecordingFailTarget();
FleetHealthMonitor monitor = monitorWith(failTarget);
monitor.reportTransition("term_a", HealthState.GONE);
monitor.stop();
assertEquals(1, failTarget.calls.size());
assertEquals("term_a", failTarget.calls.get(0).target());
assertTrue(failTarget.calls.get(0).reason().contains("GONE"));
}
@Test void neverReadyNamesItselfAsTheReason() {
RecordingFailTarget failTarget = new RecordingFailTarget();
FleetHealthMonitor monitor = monitorWith(failTarget);
monitor.reportTransition("term_a", HealthState.NEVER_READY);
monitor.stop();
assertEquals(1, failTarget.calls.size());
assertTrue(failTarget.calls.get(0).reason().contains("NEVER_READY"));
}
@Test void stayingInATerminalStateProducesOneFailureNotN() {
RecordingFailTarget failTarget = new RecordingFailTarget();
FleetHealthMonitor monitor = monitorWith(failTarget);
monitor.reportTransition("term_a", HealthState.GONE);
monitor.reportTransition("term_a", HealthState.GONE);
monitor.reportTransition("term_a", HealthState.GONE);
monitor.reportTransition("term_a", HealthState.GONE);
monitor.stop();
assertEquals(1, failTarget.calls.size());
}
@Test void aNonTerminalFaultStateDoesNotFailTheTarget() {
RecordingFailTarget failTarget = new RecordingFailTarget();
FleetHealthMonitor monitor = monitorWith(failTarget);
monitor.reportTransition("term_a", HealthState.TURN_BOUNDARY_LOST);
monitor.stop();
assertEquals(0, failTarget.calls.size());
}
@Test void failTargetRetryIsBounded() {
AlwaysThrowingFailTarget failTarget = new AlwaysThrowingFailTarget();
FleetHealthMonitor monitor = monitorWith(failTarget);
monitor.reportTransition("term_a", HealthState.GONE);
monitor.stop();
assertEquals(FleetHealthMonitor.MAX_FAIL_TARGET_ATTEMPTS, failTarget.calls);
}
@Test void exhaustedRetryStillDoesNotRefireOnAnUnchangedTick() {
AlwaysThrowingFailTarget failTarget = new AlwaysThrowingFailTarget();
FleetHealthMonitor monitor = monitorWith(failTarget);
monitor.reportTransition("term_a", HealthState.GONE);
int afterFirstTransition = failTarget.calls;
monitor.reportTransition("term_a", HealthState.GONE);
monitor.stop();
assertEquals(afterFirstTransition, failTarget.calls);
}
private record RecordedCall(String target, String reason) { }
private static final class RecordingFailTarget implements BiConsumer<String, String> {
final java.util.List<RecordedCall> calls = new java.util.ArrayList<>();
@Override public void accept(String target, String reason) {
calls.add(new RecordedCall(target, reason));
}
}
private static final class AlwaysThrowingFailTarget implements BiConsumer<String, String> {
int calls = 0;
@Override public void accept(String target, String reason) {
calls++;
throw new RuntimeException("boom");
}
}
}
@@ -15,8 +15,9 @@ import java.util.concurrent.atomic.AtomicLong;
import static org.junit.jupiter.api.Assertions.*;
/**
* CB-531. A lead is never spawned, so the daemon has to <em>find</em> it: these assert that an
* operator-labelled tab is what makes a pane a lead, and — just as importantly — what does not.
* CB-531/CB-579. A lead is never spawned, so the daemon has to <em>find</em> it: these assert that
* an operator-labelled tab matching a configured {@code tab:} is what makes a pane a lead, and —
* just as importantly — what does not, and that a stale entry does not linger forever.
*/
class LeadTabScannerTest {
@@ -120,23 +121,48 @@ class LeadTabScannerTest {
.pane("w9:p1", "w9:t1", "term_worker");
}
private LeadTabScanner scanner(TopologyHerdr herdr, Map<String, String> configured,
/** The {@code tab:} → name map {@code twoLeads()}'s two lead tabs are configured under. */
private static Map<String, String> twoLeadsConfigured() {
return Map.of("lead: opus-5.0", "opus-5.0", "lead: gpt-sol-5.6", "gpt-sol-5.6");
}
private LeadTabScanner scanner(TopologyHerdr herdr, Map<String, String> tabToName,
AtomicLong clock) {
return new LeadTabScanner(herdr, "lead:", Set.of("bridged-workers"), configured, TTL,
clock::get);
return new LeadTabScanner(herdr, tabToName, Set.of("bridged-workers"), TTL, clock::get);
}
@Test
void everyLabelledTabBecomesALeadNamedByItsLabel() {
Map<String, String> leads = scanner(twoLeads(), Map.of(), new AtomicLong()).get();
void everyConfiguredTabBecomesALeadNamedByItsEntry() {
Map<String, String> leads = scanner(twoLeads(), twoLeadsConfigured(), new AtomicLong()).get();
assertEquals(Map.of("term_opus", "opus-5.0", "term_gpt", "gpt-sol-5.6"), leads,
"two leads discovered from labels alone — no terminal_id was ever configured");
"two leads discovered by their configured tab — no terminal_id was ever configured");
}
@Test
void anUnlabelledTabContributesNothing() {
assertFalse(scanner(twoLeads(), Map.of(), new AtomicLong()).get().containsKey("term_notes"));
void anUnconfiguredTabContributesNothing() {
assertFalse(scanner(twoLeads(), twoLeadsConfigured(), new AtomicLong())
.get().containsKey("term_notes"));
}
/**
* CB-579: matching is exact against the configured map now, not a shared prefix — two leads with
* completely different labels are both discovered by one scanner, no convention required.
*/
@Test
void twoLeadsWithCompletelyDifferentLabelsAreBothDiscovered() {
TopologyHerdr herdr = new TopologyHerdr()
.workspace("w1", "main")
.tab("w1:t1", "w1", "orchestrator: opus")
.tab("w1:t2", "w1", "captain: sol")
.pane("w1:p1", "w1:t1", "term_opus")
.pane("w1:p2", "w1:t2", "term_sol");
Map<String, String> tabToName = Map.of("orchestrator: opus", "opus", "captain: sol", "sol");
Map<String, String> leads = scanner(herdr, tabToName, new AtomicLong()).get();
assertEquals(Map.of("term_opus", "opus", "term_sol", "sol"), leads,
"no shared prefix needed — each lead is matched by its own configured tab");
}
/**
@@ -148,25 +174,28 @@ class LeadTabScannerTest {
void aTabInAWorkerSpaceIsNeverALeadEvenWhenItsLabelMatches() {
TopologyHerdr herdr = twoLeads().tab("w9:t2", "w9", "lead: impostor")
.pane("w9:p2", "w9:t2", "term_impostor");
Map<String, String> tabToName = new LinkedHashMap<>(twoLeadsConfigured());
tabToName.put("lead: impostor", "impostor");
assertFalse(scanner(herdr, Map.of(), new AtomicLong()).get().containsKey("term_impostor"));
assertFalse(scanner(herdr, tabToName, new AtomicLong()).get().containsKey("term_impostor"));
}
@Test
void aBarePrefixNamesNobodyAndIsRejected() {
void aLabelWithNoConfiguredEntryIsIgnored() {
TopologyHerdr herdr = new TopologyHerdr().workspace("w1", "main")
.tab("w1:t1", "w1", "lead:").pane("w1:p1", "w1:t1", "term_a");
.tab("w1:t1", "w1", "lead: nobody-configured").pane("w1:p1", "w1:t1", "term_a");
assertEquals(Map.of(), scanner(herdr, Map.of(), new AtomicLong()).get(),
"a lead with no name would resolve as PRIMARY with nothing to attribute it to");
assertEquals(Map.of(), scanner(herdr, twoLeadsConfigured(), new AtomicLong()).get(),
"a label that names no configured lead resolves nobody");
}
@Test
void thePrefixMatchesCaseInsensitivelyAndTheNameIsTrimmed() {
void matchingIsCaseInsensitiveAndToleratesSurroundingWhitespace() {
TopologyHerdr herdr = new TopologyHerdr().workspace("w1", "main")
.tab("w1:t1", "w1", " LEAD: opus-5.0 ").pane("w1:p1", "w1:t1", "term_a");
.tab("w1:t1", "w1", " LEAD: Opus-5.0 ").pane("w1:p1", "w1:t1", "term_a");
assertEquals(Map.of("term_a", "opus-5.0"), scanner(herdr, Map.of(), new AtomicLong()).get());
assertEquals(Map.of("term_a", "opus-5.0"),
scanner(herdr, Map.of("lead: Opus-5.0", "opus-5.0"), new AtomicLong()).get());
}
@Test
@@ -175,18 +204,51 @@ class LeadTabScannerTest {
// nothing bridged placed can land here (see the worker-space test above).
TopologyHerdr herdr = twoLeads().pane("w1:p1b", "w1:t1", "term_opus_split");
assertEquals("opus-5.0", scanner(herdr, Map.of(), new AtomicLong()).get().get("term_opus_split"));
assertEquals("opus-5.0",
scanner(herdr, twoLeadsConfigured(), new AtomicLong()).get().get("term_opus_split"));
}
/**
* CB-579 acceptance (6): this is the bug the ticket closes. A stale pin used to be merged back
* over every scan and never expire; now a scan is the whole answer, so a lead whose tab is gone
* drops out on the very next scan.
*/
@Test
void anExplicitlyConfiguredLeadIsMergedInAndOutranksALabel() {
Map<String, String> configured = Map.of("term_opus", "pinned-name", "term_extra", "from-config");
void aTabNoLongerPresentDropsTheLeadOnTheNextScan() {
TopologyHerdr herdr = twoLeads();
AtomicLong clock = new AtomicLong();
LeadTabScanner s = scanner(herdr, twoLeadsConfigured(), clock);
assertTrue(s.get().containsKey("term_opus"));
Map<String, String> leads = scanner(twoLeads(), configured, new AtomicLong()).get();
// The session behind term_opus restarted — herdr no longer reports that tab or pane at all.
herdr.tabs.remove("w1:t1");
herdr.panes.remove("w1:p1");
clock.addAndGet(TTL);
assertEquals("pinned-name", leads.get("term_opus"), "an explicit pin is the operator's last word");
assertEquals("from-config", leads.get("term_extra"), "a configured lead needs no tab at all");
assertEquals("gpt-sol-5.6", leads.get("term_gpt"));
assertFalse(s.get().containsKey("term_opus"),
"a stale entry must expire once the tab it named is gone, not be merged back forever");
}
/**
* CB-579 acceptance (5): the whole point of matching by tab instead of {@code terminal_id} — a
* restart changes the terminal, not the tab, so the lead resolves under the same name with no
* config edit.
*/
@Test
void aLeadRestartingInTheSameTabResolvesUnderTheSameName() {
TopologyHerdr herdr = twoLeads();
AtomicLong clock = new AtomicLong();
LeadTabScanner s = scanner(herdr, twoLeadsConfigured(), clock);
assertEquals("opus-5.0", s.get().get("term_opus"));
// The session restarts: herdr assigns the pane a new terminal_id, same tab (w1:t1).
herdr.panes.remove("w1:p1");
herdr.pane("w1:p1", "w1:t1", "term_opus_v2");
clock.addAndGet(TTL);
Map<String, String> leads = s.get();
assertEquals("opus-5.0", leads.get("term_opus_v2"), "the new terminal resolves immediately");
assertFalse(leads.containsKey("term_opus"), "the old terminal_id is simply gone, not carried");
}
// ── caching ─────────────────────────────────────────────────────────────────────────────────
@@ -195,7 +257,7 @@ class LeadTabScannerTest {
void aSecondLookupWithinTheTtlDoesNotTouchHerdr() {
TopologyHerdr herdr = twoLeads();
AtomicLong clock = new AtomicLong();
LeadTabScanner s = scanner(herdr, Map.of(), clock);
LeadTabScanner s = scanner(herdr, twoLeadsConfigured(), clock);
s.get();
int afterFirst = herdr.calls;
@@ -210,21 +272,23 @@ class LeadTabScannerTest {
void aTabLabelledAfterStartupIsPickedUpOnceTheTtlExpires() {
TopologyHerdr herdr = twoLeads();
AtomicLong clock = new AtomicLong();
LeadTabScanner s = scanner(herdr, Map.of(), clock);
Map<String, String> tabToName = new LinkedHashMap<>(twoLeadsConfigured());
tabToName.put("lead: late-arrival", "late-arrival");
LeadTabScanner s = scanner(herdr, tabToName, clock);
assertFalse(s.get().containsKey("term_notes"));
herdr.tab("w1:t3", "w1", "lead: late-arrival"); // the operator renames their tab
clock.addAndGet(TTL);
assertEquals("late-arrival", s.get().get("term_notes"),
"the whole point over `leaders:`: no config edit, no restart");
"the whole point over a config-held terminal_id: no config edit, no restart");
}
@Test
void aFailedScanKeepsTheLeadsAlreadyKnownRatherThanDemotingThem() {
TopologyHerdr herdr = twoLeads();
AtomicLong clock = new AtomicLong();
LeadTabScanner s = scanner(herdr, Map.of(), clock);
LeadTabScanner s = scanner(herdr, twoLeadsConfigured(), clock);
Map<String, String> before = s.get();
herdr.failing = true;
@@ -235,14 +299,14 @@ class LeadTabScannerTest {
}
@Test
void aFailedFirstScanStillHonoursTheConfiguredLeads() {
void aFailedFirstScanReturnsEmptyRatherThanThrowing() {
TopologyHerdr herdr = twoLeads();
herdr.failing = true;
Map<String, String> leads = scanner(herdr, Map.of("term_x", "opus-5.0"), new AtomicLong()).get();
Map<String, String> leads = scanner(herdr, twoLeadsConfigured(), new AtomicLong()).get();
assertEquals(Map.of("term_x", "opus-5.0"), leads,
"config-named leads must not depend on herdr answering at all");
assertEquals(Map.of(), leads,
"with nothing scanned yet and no override to fall back on, the map is simply empty");
}
@Test
@@ -250,7 +314,7 @@ class LeadTabScannerTest {
TopologyHerdr herdr = twoLeads();
herdr.failing = true;
AtomicLong clock = new AtomicLong();
LeadTabScanner s = scanner(herdr, Map.of(), clock);
LeadTabScanner s = scanner(herdr, twoLeadsConfigured(), clock);
s.get();
int afterFirst = herdr.calls;
@@ -7,9 +7,14 @@ import ch.qos.logback.core.read.ListAppender;
import dev.ltms.bridged.herdr.AgentControl;
import dev.ltms.bridged.herdr.FakeHerdr;
import dev.ltms.bridged.msg.Rendezvous;
import dev.ltms.bridged.msg.TestTurnTokens;
import dev.ltms.bridged.msg.TurnToken;
import org.junit.jupiter.api.Test;
import org.slf4j.LoggerFactory;
import java.util.Set;
import java.util.regex.Pattern;
import static org.junit.jupiter.api.Assertions.assertEquals;
import static org.junit.jupiter.api.Assertions.assertFalse;
import static org.junit.jupiter.api.Assertions.assertTrue;
@@ -21,7 +26,7 @@ class CompletionResolverTest {
void skipsTheScrapeWhenNoSendIsWaiting() {
FakeHerdr herdr = new FakeHerdr();
Rendezvous rendezvous = new Rendezvous(); // no waiter opened
CompletionResolver resolver = new CompletionResolver(new AgentControl(herdr), rendezvous);
CompletionResolver resolver = new CompletionResolver(new AgentControl(herdr), rendezvous, ExhaustedPatternLookup.none(), ExhaustionSink.none());
resolver.resolve("term_a", null); // no in-flight turn captured for this target
@@ -33,7 +38,7 @@ class CompletionResolverTest {
void failSkipsTheScrapeWhenNoSendIsWaiting() {
FakeHerdr herdr = new FakeHerdr();
Rendezvous rendezvous = new Rendezvous(); // no waiter opened
CompletionResolver resolver = new CompletionResolver(new AgentControl(herdr), rendezvous);
CompletionResolver resolver = new CompletionResolver(new AgentControl(herdr), rendezvous, ExhaustedPatternLookup.none(), ExhaustionSink.none());
resolver.fail("term_a", null); // no in-flight turn, and no registered waiter to fall back to
@@ -45,9 +50,9 @@ class CompletionResolverTest {
void captureBaselineSkipsTheReadWhenNoSendIsWaiting() {
FakeHerdr herdr = new FakeHerdr().readText("⏺ X\n❯ ");
Rendezvous rendezvous = new Rendezvous(); // no waiter opened
CompletionResolver resolver = new CompletionResolver(new AgentControl(herdr), rendezvous);
CompletionResolver resolver = new CompletionResolver(new AgentControl(herdr), rendezvous, ExhaustedPatternLookup.none(), ExhaustionSink.none());
resolver.captureBaseline("term_a"); // no send to attribute a later completion to
resolver.captureBaseline("term_a", TestTurnTokens.inert("term_a")); // no send to attribute a later completion to
assertFalse(herdr.called("agent.read"),
"with no waiting send there is no turn to baseline — skip the scrape");
@@ -134,7 +139,7 @@ class CompletionResolverTest {
// send must NOT be resolved with the stale answer.
FakeHerdr herdr = new FakeHerdr().readText("⏺ 391\n❯ ");
Rendezvous rendezvous = new Rendezvous();
CompletionResolver resolver = new CompletionResolver(new AgentControl(herdr), rendezvous);
CompletionResolver resolver = new CompletionResolver(new AgentControl(herdr), rendezvous, ExhaustedPatternLookup.none(), ExhaustionSink.none());
var waiter = rendezvous.open("term_a"); // a send is blocked on this turn
// The turn as captured at delivery: its waiter, and the previous turn's answer still on screen.
@@ -149,7 +154,7 @@ class CompletionResolverTest {
void resolvesACompletionWhoseScrapeChangedSinceDelivery() {
FakeHerdr herdr = new FakeHerdr().readText("⏺ No, 391 = 17 × 23.\n❯ "); // the worker's real answer
Rendezvous rendezvous = new Rendezvous();
CompletionResolver resolver = new CompletionResolver(new AgentControl(herdr), rendezvous);
CompletionResolver resolver = new CompletionResolver(new AgentControl(herdr), rendezvous, ExhaustedPatternLookup.none(), ExhaustionSink.none());
var waiter = rendezvous.open("term_a");
// Delivery baseline was the previous turn's "391"; the scrape now differs → resolve.
@@ -166,7 +171,7 @@ class CompletionResolverTest {
String block = "⏺ " + "x".repeat(CompletionResolver.MAX_SCRAPE_CHARS + 1) + "\n❯ ";
FakeHerdr herdr = new FakeHerdr().readText(block);
Rendezvous rendezvous = new Rendezvous();
CompletionResolver resolver = new CompletionResolver(new AgentControl(herdr), rendezvous);
CompletionResolver resolver = new CompletionResolver(new AgentControl(herdr), rendezvous, ExhaustedPatternLookup.none(), ExhaustionSink.none());
var waiter = rendezvous.open("term_a");
resolver.resolve("term_a", new CompletionResolver.InFlight(waiter, null));
@@ -180,7 +185,7 @@ class CompletionResolverTest {
void leavesAnUnclippedCompletionPaneTailUnmarked() {
FakeHerdr herdr = new FakeHerdr().readText("⏺ complete report\n❯ ");
Rendezvous rendezvous = new Rendezvous();
CompletionResolver resolver = new CompletionResolver(new AgentControl(herdr), rendezvous);
CompletionResolver resolver = new CompletionResolver(new AgentControl(herdr), rendezvous, ExhaustedPatternLookup.none(), ExhaustionSink.none());
var waiter = rendezvous.open("term_a");
resolver.resolve("term_a", new CompletionResolver.InFlight(waiter, null));
@@ -192,9 +197,9 @@ class CompletionResolverTest {
void resolvesSynchronouslyBeforePostTurnContextClearing() {
FakeHerdr herdr = new FakeHerdr().readText("⏺ previous answer\n❯ ");
Rendezvous rendezvous = new Rendezvous();
CompletionResolver resolver = new CompletionResolver(new AgentControl(herdr), rendezvous);
CompletionResolver resolver = new CompletionResolver(new AgentControl(herdr), rendezvous, ExhaustedPatternLookup.none(), ExhaustionSink.none());
var waiter = rendezvous.open("term_a");
resolver.captureBaseline("term_a");
resolver.captureBaseline("term_a", new TurnToken("term_a", waiter));
herdr.readText("⏺ answer that /clear would erase\n❯ ");
resolver.resolveBeforePostAction("term_a");
@@ -214,10 +219,10 @@ class CompletionResolverTest {
String longBlock = "⏺ " + "x".repeat(CompletionResolver.MAX_SCRAPE_CHARS + 500) + "\n❯ ";
FakeHerdr herdr = new FakeHerdr().readText(longBlock);
Rendezvous rendezvous = new Rendezvous();
CompletionResolver resolver = new CompletionResolver(new AgentControl(herdr), rendezvous);
CompletionResolver resolver = new CompletionResolver(new AgentControl(herdr), rendezvous, ExhaustedPatternLookup.none(), ExhaustionSink.none());
var waiter = rendezvous.open("term_a"); // a send is blocked on this turn
resolver.captureBaseline("term_a"); // baseline is the clipped >cap block
resolver.captureBaseline("term_a", new TurnToken("term_a", waiter)); // baseline is the clipped >cap block
var turn = resolver.inFlight("term_a");
assertEquals(CompletionResolver.MAX_SCRAPE_CHARS, turn.baseline().length(),
"the delivery baseline is clipped to the same cap resolve() applies to the tail");
@@ -234,7 +239,7 @@ class CompletionResolverTest {
// No delivery baseline (e.g. the pre-turn read failed) ⇒ never suppress; the completion resolves.
FakeHerdr herdr = new FakeHerdr().readText("⏺ hello\n❯ ");
Rendezvous rendezvous = new Rendezvous();
CompletionResolver resolver = new CompletionResolver(new AgentControl(herdr), rendezvous);
CompletionResolver resolver = new CompletionResolver(new AgentControl(herdr), rendezvous, ExhaustedPatternLookup.none(), ExhaustionSink.none());
var waiter = rendezvous.open("term_a");
resolver.resolve("term_a", new CompletionResolver.InFlight(waiter, null));
@@ -252,7 +257,7 @@ class CompletionResolverTest {
// byte-identical guard would wrongly match the empty tail and suppress.
FakeHerdr herdr = new FakeHerdr().healthy(false); // agent.read throws HerdrException
Rendezvous rendezvous = new Rendezvous();
CompletionResolver resolver = new CompletionResolver(new AgentControl(herdr), rendezvous);
CompletionResolver resolver = new CompletionResolver(new AgentControl(herdr), rendezvous, ExhaustedPatternLookup.none(), ExhaustionSink.none());
var waiter = rendezvous.open("term_a");
var turn = new CompletionResolver.InFlight(waiter, ""); // empty pane baselined at delivery
@@ -272,7 +277,7 @@ class CompletionResolverTest {
// fail must not overwrite that value, and must not even scrape the worker — nobody needs it.
FakeHerdr herdr = new FakeHerdr().readText("an error screen");
Rendezvous rendezvous = new Rendezvous();
CompletionResolver resolver = new CompletionResolver(new AgentControl(herdr), rendezvous);
CompletionResolver resolver = new CompletionResolver(new AgentControl(herdr), rendezvous, ExhaustedPatternLookup.none(), ExhaustionSink.none());
var waiter = rendezvous.open("term_a");
var turn = new CompletionResolver.InFlight(waiter, null);
@@ -293,7 +298,7 @@ class CompletionResolverTest {
// fail falls back to the waiter currently registered on the Rendezvous and fails it.
FakeHerdr herdr = new FakeHerdr().readText("stuck on an error screen");
Rendezvous rendezvous = new Rendezvous();
CompletionResolver resolver = new CompletionResolver(new AgentControl(herdr), rendezvous);
CompletionResolver resolver = new CompletionResolver(new AgentControl(herdr), rendezvous, ExhaustedPatternLookup.none(), ExhaustionSink.none());
var waiter = rendezvous.open("term_a"); // send registered, but no captureBaseline ever ran
resolver.fail("term_a", null); // no in-flight turn → fall back to the registered waiter
@@ -319,7 +324,7 @@ class CompletionResolverTest {
try {
FakeHerdr herdr = new FakeHerdr().readText("stuck on an error screen");
Rendezvous rendezvous = new Rendezvous();
CompletionResolver resolver = new CompletionResolver(new AgentControl(herdr), rendezvous);
CompletionResolver resolver = new CompletionResolver(new AgentControl(herdr), rendezvous, ExhaustedPatternLookup.none(), ExhaustionSink.none());
var waiter = rendezvous.open("term_a");
resolver.fail("term_a", null);
@@ -347,7 +352,7 @@ class CompletionResolverTest {
// scrape to turn N+1; targeting turn N's captured waiter makes the late completion a no-op.
FakeHerdr herdr = new FakeHerdr().readText("⏺ turn N answer\n❯ ");
Rendezvous rendezvous = new Rendezvous();
CompletionResolver resolver = new CompletionResolver(new AgentControl(herdr), rendezvous);
CompletionResolver resolver = new CompletionResolver(new AgentControl(herdr), rendezvous, ExhaustedPatternLookup.none(), ExhaustionSink.none());
var waiterN = rendezvous.open("term_a"); // turn N's send
// The turn as the injector captured it at delivery (waiter + pre-turn baseline).
@@ -368,4 +373,129 @@ class CompletionResolverTest {
"turn N stays resolved by its own reply");
assertTrue(rendezvous.isWaiting("term_a"), "turn N+1 is still awaiting its own resolution");
}
// --- CB-578 stage A: backend-exhausted classification ---------------------------------
@Test
void classifiesAMatchingScrapeAsBackendExhaustedInsteadOfACompletedReply() {
String block = "⏺ Working on it...\nThe usage limit has been reached. Try again later.\n❯ ";
FakeHerdr herdr = new FakeHerdr().readText(block);
Rendezvous rendezvous = new Rendezvous();
ExhaustedPatternLookup patterns = target -> Pattern.compile("usage limit has been reached");
CompletionResolver resolver = new CompletionResolver(new AgentControl(herdr), rendezvous, patterns, ExhaustionSink.none());
var waiter = rendezvous.open("term_a");
resolver.resolve("term_a", new CompletionResolver.InFlight(waiter, null));
assertTrue(waiter.isDone(), "a matching scrape still resolves the blocked send");
assertEquals(Rendezvous.Kind.BACKEND_EXHAUSTED, waiter.getNow(null).kind(),
"not reported as a completed reply — the classification is distinct");
}
@Test
void theExhaustedReasonCarriesTheMatchedLine() {
String block = "⏺ Working on it...\nThe usage limit has been reached. Try again later.\n❯ ";
FakeHerdr herdr = new FakeHerdr().readText(block);
Rendezvous rendezvous = new Rendezvous();
ExhaustedPatternLookup patterns = target -> Pattern.compile("usage limit has been reached");
CompletionResolver resolver = new CompletionResolver(new AgentControl(herdr), rendezvous, patterns, ExhaustionSink.none());
var waiter = rendezvous.open("term_a");
resolver.resolve("term_a", new CompletionResolver.InFlight(waiter, null));
assertEquals("backend exhausted (usage limit): The usage limit has been reached. Try again later.",
waiter.getNow(null).text(), "the reason names the real cause and carries the matched line");
}
@Test
void aWinningBackendExhaustedClassificationNotifiesTheExhaustionSink() {
String block = "⏺ Working on it...\nThe usage limit has been reached. Try again later.\n❯ ";
FakeHerdr herdr = new FakeHerdr().readText(block);
Rendezvous rendezvous = new Rendezvous();
ExhaustedPatternLookup patterns = target -> Pattern.compile("usage limit has been reached");
java.util.List<String> notified = new java.util.ArrayList<>();
ExhaustionSink sink = (target, reason) -> notified.add(target + ": " + reason);
CompletionResolver resolver = new CompletionResolver(new AgentControl(herdr), rendezvous, patterns, sink);
var waiter = rendezvous.open("term_a");
resolver.resolve("term_a", new CompletionResolver.InFlight(waiter, null));
assertEquals(1, notified.size(), "the sink is notified exactly once for the winning classification");
assertTrue(notified.get(0).startsWith("term_a: "), "the sink is told which target exhausted");
assertTrue(notified.get(0).contains("The usage limit has been reached"),
"the sink is told the matched reason: " + notified.get(0));
}
@Test
void aLosingBackendExhaustedClassificationNeverNotifiesTheExhaustionSink() {
// The waiter was already resolved (e.g. by the worker's own reply) before this scrape landed —
// resolveExhausted loses the race and must return false, so the sink must not fire either.
String block = "⏺ Working on it...\nThe usage limit has been reached. Try again later.\n❯ ";
FakeHerdr herdr = new FakeHerdr().readText(block);
Rendezvous rendezvous = new Rendezvous();
ExhaustedPatternLookup patterns = target -> Pattern.compile("usage limit has been reached");
java.util.List<String> notified = new java.util.ArrayList<>();
ExhaustionSink sink = (target, reason) -> notified.add(target + ": " + reason);
CompletionResolver resolver = new CompletionResolver(new AgentControl(herdr), rendezvous, patterns, sink);
var waiter = rendezvous.open("term_a");
var turn = new CompletionResolver.InFlight(waiter, null);
assertTrue(rendezvous.resolveCompletion(waiter, "already replied"));
resolver.resolve("term_a", turn);
assertTrue(notified.isEmpty(), "a classification that loses the race must not quarantine anything");
assertEquals("already replied", waiter.getNow(null).text(), "the earlier resolution stands untouched");
}
@Test
void aNonMatchingScrapeResolvesAsAnOrdinaryCompletion() {
FakeHerdr herdr = new FakeHerdr().readText("⏺ complete report\n❯ ");
Rendezvous rendezvous = new Rendezvous();
ExhaustedPatternLookup patterns = target -> Pattern.compile("usage limit has been reached");
CompletionResolver resolver = new CompletionResolver(new AgentControl(herdr), rendezvous, patterns, ExhaustionSink.none());
var waiter = rendezvous.open("term_a");
resolver.resolve("term_a", new CompletionResolver.InFlight(waiter, null));
assertEquals(Rendezvous.Kind.COMPLETION, waiter.getNow(null).kind(),
"a scrape that does not match the pattern is an ordinary completion");
assertEquals("complete report", waiter.getNow(null).text());
}
@Test
void aProfileWithNoConfiguredPatternKeepsTodaysCompletionFallbackUnchanged() {
// Even a scrape that WOULD have matched some other profile's pattern must resolve as a
// plain completion when this target's own profile has none configured (CB-578 criterion 4).
String block = "⏺ The usage limit has been reached.\n❯ ";
FakeHerdr herdr = new FakeHerdr().readText(block);
Rendezvous rendezvous = new Rendezvous();
CompletionResolver resolver =
new CompletionResolver(new AgentControl(herdr), rendezvous, ExhaustedPatternLookup.none(), ExhaustionSink.none());
var waiter = rendezvous.open("term_a");
resolver.resolve("term_a", new CompletionResolver.InFlight(waiter, null));
assertEquals(Rendezvous.Kind.COMPLETION, waiter.getNow(null).kind(),
"no pattern configured for this target's profile ⇒ unchanged completion-fallback behaviour");
assertEquals("The usage limit has been reached.", waiter.getNow(null).text());
}
@Test
void coverageIsOffWhenNoProfileHasAPatternConfigured() {
assertEquals("off (no profile has an exhaustedPattern configured; profiles: [terra])",
CompletionResolver.coverage(Set.of("terra"), Set.of()));
}
@Test
void coverageIsFullWhenEveryProfileHasAPatternConfigured() {
assertEquals("full (all profiles configured: [gx10, terra])",
CompletionResolver.coverage(Set.of("terra", "gx10"), Set.of("terra", "gx10")));
}
@Test
void coverageIsPartialAndNamesWhichProfilesAreConfigured() {
assertEquals("partial (configured: [terra]; not configured: [gx10])",
CompletionResolver.coverage(Set.of("terra", "gx10"), Set.of("terra")));
}
}
@@ -8,6 +8,7 @@ import dev.ltms.bridged.herdr.AgentControl;
import dev.ltms.bridged.herdr.AgentStatus;
import dev.ltms.bridged.herdr.FakeHerdr;
import dev.ltms.bridged.herdr.HerdrException;
import dev.ltms.bridged.msg.TestTurnTokens;
import org.junit.jupiter.api.Test;
import org.slf4j.LoggerFactory;
@@ -47,7 +48,7 @@ class InjectorTest {
@Test
void deliversWhenIdle() {
CompletableFuture<Void> f = injector.enqueue(T, "hello");
CompletableFuture<Void> f = injector.enqueue(T, "hello", TestTurnTokens.inert(T));
assertFalse(f.isDone(), "not delivered until an injectable status arrives");
injector.onStatus(T, AgentStatus.IDLE);
assertTrue(f.isDone());
@@ -59,7 +60,7 @@ class InjectorTest {
// CB-113: idle alone is not enough — hold until the worker's MCP is connected (ready).
java.util.Set<String> ready = new java.util.HashSet<>();
Injector inj = new Injector(new AgentControl(herdr), TurnListener.NOOP, ready::contains);
inj.enqueue(T, "task");
inj.enqueue(T, "task", TestTurnTokens.inert(T));
inj.onStatus(T, AgentStatus.IDLE); // idle but not yet available → held out of the boot window
assertEquals(List.of(), sent(), "must not deliver into a not-yet-available worker");
@@ -79,7 +80,7 @@ class InjectorTest {
void resubmitsEnterWhenADeliveredMessageIsNotPickedUp() {
// CB-113: the Enter at delivery can race the paste; while the worker stays idle (not picked
// up), the injector re-nudges Enter so the pending paste submits.
injector.enqueue(T, "task");
injector.enqueue(T, "task", TestTurnTokens.inert(T));
injector.onStatus(T, AgentStatus.IDLE); // deliver: paste + one Enter
long afterDeliver = enterKeystrokes();
@@ -95,7 +96,7 @@ class InjectorTest {
@Test
void holdsWhileWorkingThenDeliversOnIdle() {
injector.enqueue(T, "later");
injector.enqueue(T, "later", TestTurnTokens.inert(T));
injector.onStatus(T, AgentStatus.WORKING);
assertEquals(List.of(), sent(), "must not inject mid-turn");
injector.onStatus(T, AgentStatus.IDLE);
@@ -104,7 +105,7 @@ class InjectorTest {
@Test
void blockedIsInjectableButUnknownIsNot() {
injector.enqueue(T, "answer");
injector.enqueue(T, "answer", TestTurnTokens.inert(T));
injector.onStatus(T, AgentStatus.UNKNOWN);
assertEquals(List.of(), sent(), "unknown status is not safe to inject");
injector.onStatus(T, AgentStatus.BLOCKED);
@@ -113,8 +114,8 @@ class InjectorTest {
@Test
void twoRapidDeliveriesNeverInterleave() {
injector.enqueue(T, "m1");
injector.enqueue(T, "m2");
injector.enqueue(T, "m1", TestTurnTokens.inert(T));
injector.enqueue(T, "m2", TestTurnTokens.inert(T));
// First idle window delivers only m1, even if idle is observed twice before pickup.
injector.onStatus(T, AgentStatus.IDLE);
@@ -129,8 +130,8 @@ class InjectorTest {
@Test
void transientUnknownDoesNotReleaseThePickupLatch() {
injector.enqueue(T, "m1");
injector.enqueue(T, "m2");
injector.enqueue(T, "m1", TestTurnTokens.inert(T));
injector.enqueue(T, "m2", TestTurnTokens.inert(T));
injector.onStatus(T, AgentStatus.IDLE); // m1 sent, awaiting pickup
assertEquals(List.of("m1"), sent());
@@ -145,8 +146,8 @@ class InjectorTest {
@Test
void missedPickupEdgeIsReleasedByGraceSoTheQueueNeverWedges() {
injector.enqueue(T, "m1");
injector.enqueue(T, "m2");
injector.enqueue(T, "m1", TestTurnTokens.inert(T));
injector.enqueue(T, "m2", TestTurnTokens.inert(T));
injector.onStatus(T, AgentStatus.IDLE); // m1 sent
assertEquals(List.of("m1"), sent());
@@ -158,9 +159,9 @@ class InjectorTest {
@Test
void fifoOrderAcrossManyTurns() {
injector.enqueue(T, "a");
injector.enqueue(T, "b");
injector.enqueue(T, "c");
injector.enqueue(T, "a", TestTurnTokens.inert(T));
injector.enqueue(T, "b", TestTurnTokens.inert(T));
injector.enqueue(T, "c", TestTurnTokens.inert(T));
for (int i = 0; i < 3; i++) {
injector.onStatus(T, AgentStatus.IDLE); // deliver one
injector.onStatus(T, AgentStatus.WORKING); // pickup
@@ -172,7 +173,7 @@ class InjectorTest {
@Test
void activeWhileQueuedOrInFlightThenQuietAfterTurnCompletes() {
assertTrue(injector.activeTargets().isEmpty());
injector.enqueue(T, "x");
injector.enqueue(T, "x", TestTurnTokens.inert(T));
assertEquals(Set.of(T), injector.activeTargets(), "active while a message is queued");
injector.onStatus(T, AgentStatus.IDLE); // delivers; awaiting pickup
@@ -191,7 +192,7 @@ class InjectorTest {
void firesTurnCompleteOnAConfirmedWorkingThenIdle() {
List<String> completed = new ArrayList<>();
Injector inj = new Injector(new AgentControl(herdr), completed::add);
inj.enqueue(T, "task");
inj.enqueue(T, "task", TestTurnTokens.inert(T));
inj.onStatus(T, AgentStatus.IDLE); // deliver
inj.onStatus(T, AgentStatus.WORKING); // pickup + turn running
@@ -226,8 +227,8 @@ class InjectorTest {
}
ResetListener listener = new ResetListener();
Injector inj = new Injector(agents, listener);
inj.enqueue(T, "first");
inj.enqueue(T, "second");
inj.enqueue(T, "first", TestTurnTokens.inert(T));
inj.enqueue(T, "second", TestTurnTokens.inert(T));
inj.onStatus(T, AgentStatus.IDLE); // first delegation
inj.onStatus(T, AgentStatus.WORKING);
@@ -246,7 +247,7 @@ class InjectorTest {
void doesNotSynthesizeCompletionFromAnUnconfirmedTurn() {
List<String> completed = new ArrayList<>();
Injector inj = new Injector(new AgentControl(herdr), completed::add);
inj.enqueue(T, "task");
inj.enqueue(T, "task", TestTurnTokens.inert(T));
// Deliver, then only ever idle — a `working` sample is never seen. The pickup grace unwedges
// the queue but must NOT invent a completion: without a sampled turn there is no trustworthy
@@ -259,6 +260,7 @@ class InjectorTest {
private static final class Captor implements TurnListener {
final List<String> completed = new ArrayList<>();
final List<String> failed = new ArrayList<>();
final List<String> failureReasons = new ArrayList<>();
@Override
public void onTurnComplete(String target) {
@@ -269,6 +271,12 @@ class InjectorTest {
public void onTurnFailed(String target) {
failed.add(target);
}
@Override
public void onTurnFailed(String target, String reason) {
failed.add(target);
failureReasons.add(reason);
}
}
// ~30s of unknown at the 250ms prod poll interval; enough onStatus samples to trip the stall.
@@ -278,7 +286,7 @@ class InjectorTest {
void failsAnOutstandingDelegationWhoseWorkerWedgesInUnknown() {
Captor cap = new Captor();
Injector inj = new Injector(new AgentControl(herdr), cap);
inj.enqueue(T, "task");
inj.enqueue(T, "task", TestTurnTokens.inert(T));
inj.onStatus(T, AgentStatus.IDLE); // deliver
inj.onStatus(T, AgentStatus.WORKING); // worker starts the turn
@@ -293,7 +301,7 @@ class InjectorTest {
void aTransientUnknownGlitchNeitherFailsNorBlocksCompletion() {
Captor cap = new Captor();
Injector inj = new Injector(new AgentControl(herdr), cap);
inj.enqueue(T, "task");
inj.enqueue(T, "task", TestTurnTokens.inert(T));
inj.onStatus(T, AgentStatus.IDLE); // deliver
inj.onStatus(T, AgentStatus.WORKING); // confirmed turn
@@ -308,7 +316,7 @@ class InjectorTest {
void sendFailureDropsMessageAndFailsItsFuture() {
FakeHerdr failing = new FakeHerdr().agentSendFailsWith("send_failed");
Injector inj = new Injector(new AgentControl(failing));
CompletableFuture<Void> f = inj.enqueue(T, "boom");
CompletableFuture<Void> f = inj.enqueue(T, "boom", TestTurnTokens.inert(T));
inj.onStatus(T, AgentStatus.IDLE);
assertTrue(f.isCompletedExceptionally());
@@ -317,18 +325,35 @@ class InjectorTest {
@Test
void dropFailsPendingWaiters() {
CompletableFuture<Void> f = injector.enqueue(T, "orphan");
CompletableFuture<Void> f = injector.enqueue(T, "orphan", TestTurnTokens.inert(T));
injector.drop(T, new HerdrException("worker gone", "pane_not_found", null));
assertTrue(f.isCompletedExceptionally(), "queued waiters unblock when the worker vanishes");
}
@Test
void dropPassesTheRealCauseForQueuedAndDeliveredWork() {
Captor cap = new Captor();
Injector inj = new Injector(new AgentControl(herdr), cap);
CompletableFuture<Void> delivered = inj.enqueue(T, "delivered", TestTurnTokens.inert(T));
CompletableFuture<Void> queued = inj.enqueue(T, "queued", TestTurnTokens.inert(T));
inj.onStatus(T, AgentStatus.IDLE); // deliver the first message
inj.onStatus(T, AgentStatus.WORKING); // its turn is now in flight; one remains queued
inj.drop(T, new HerdrException("agent target sol not found", "agent_not_found", null));
assertEquals(List.of(T), cap.failed, "drop signals one turn failure for both affected states");
assertEquals(List.of("agent target sol not found"), cap.failureReasons);
assertTrue(delivered.isDone(), "the delivered future has already completed");
assertTrue(queued.isCompletedExceptionally(), "the queued future fails with the drop cause");
}
@Test
void dropFailsTheTurnOfADeliveredMessageWhenTheWorkerVanishes() {
// CB-110: the message was delivered (no longer queued), so failing queued waiters alone would
// leave its send hanging. A vanished worker must fail that in-flight turn too.
Captor cap = new Captor();
Injector inj = new Injector(new AgentControl(herdr), cap);
inj.enqueue(T, "task");
inj.enqueue(T, "task", TestTurnTokens.inert(T));
inj.onStatus(T, AgentStatus.IDLE); // deliver
inj.onStatus(T, AgentStatus.WORKING); // turn running
@@ -343,7 +368,7 @@ class InjectorTest {
// by an off-sub worker's review of CB-110, delegated through the bridge.)
Captor cap = new Captor();
Injector inj = new Injector(new AgentControl(herdr), cap);
inj.enqueue(T, "task");
inj.enqueue(T, "task", TestTurnTokens.inert(T));
inj.onStatus(T, AgentStatus.IDLE); // deliver; pickup never confirmed
inj.drop(T, new HerdrException("worker gone", "pane_not_found", null));
@@ -362,7 +387,7 @@ class InjectorTest {
Captor cap = new Captor();
List<String> forgotten = new ArrayList<>();
Injector inj = new Injector(new AgentControl(herdr), cap, _ -> false, forgotten::add);
CompletableFuture<Void> f = inj.enqueue(T, "task");
CompletableFuture<Void> f = inj.enqueue(T, "task", TestTurnTokens.inert(T));
for (int i = 0; i < READINESS_SAMPLES; i++) inj.onStatus(T, AgentStatus.IDLE);
@@ -390,7 +415,7 @@ class InjectorTest {
try {
Injector inj = new Injector(new AgentControl(herdr), TurnListener.NOOP, _ -> false, _ -> {
});
inj.enqueue(T, "task");
inj.enqueue(T, "task", TestTurnTokens.inert(T));
for (int i = 0; i < READINESS_SAMPLES; i++) inj.onStatus(T, AgentStatus.IDLE);
@@ -415,7 +440,7 @@ class InjectorTest {
Set<String> ready = new java.util.HashSet<>();
Injector inj = new Injector(new AgentControl(herdr), TurnListener.NOOP, ready::contains, _ -> {
});
inj.enqueue(T, "task");
inj.enqueue(T, "task", TestTurnTokens.inert(T));
for (int i = 0; i < 100; i++) inj.onStatus(T, AgentStatus.IDLE); // still booting, well under grace
assertEquals(List.of(), sent());
@@ -431,7 +456,7 @@ class InjectorTest {
// linger past the worker's life (MemberPresence.forget had no caller before this).
List<String> forgotten = new ArrayList<>();
Injector inj = new Injector(new AgentControl(herdr), TurnListener.NOOP, _ -> true, forgotten::add);
inj.enqueue(T, "orphan");
inj.enqueue(T, "orphan", TestTurnTokens.inert(T));
inj.drop(T, new HerdrException("worker gone", "pane_not_found", null));
assertEquals(List.of(T), forgotten, "drop clears the gone worker's presence");
}
@@ -452,7 +477,7 @@ class InjectorTest {
try {
Injector inj = new Injector(new AgentControl(herdr), TurnListener.NOOP, _ -> true, _ -> {
});
inj.enqueue(T, "orphan");
inj.enqueue(T, "orphan", TestTurnTokens.inert(T));
inj.drop(T, new HerdrException("worker gone", "pane_not_found", null));
String warn = appender.list.stream()
@@ -476,7 +501,7 @@ class InjectorTest {
StatusPoller poller = new StatusPoller(new AgentControl(idle), inj, 10);
poller.start();
try {
CompletableFuture<Void> delivered = inj.enqueue(T, "via-poller");
CompletableFuture<Void> delivered = inj.enqueue(T, "via-poller", TestTurnTokens.inert(T));
delivered.get(2, TimeUnit.SECONDS); // completes when the poller drives the send
} finally {
poller.stop();
@@ -495,7 +520,7 @@ class InjectorTest {
void deliveredFutureCarriesSendFailure() {
FakeHerdr failing = new FakeHerdr().agentSendFailsWith("send_failed");
Injector inj = new Injector(new AgentControl(failing));
CompletableFuture<Void> f = inj.enqueue(T, "boom");
CompletableFuture<Void> f = inj.enqueue(T, "boom", TestTurnTokens.inert(T));
inj.onStatus(T, AgentStatus.IDLE);
ExecutionException ex = assertThrows(ExecutionException.class, f::get);
assertInstanceOf(HerdrException.class, ex.getCause());
@@ -28,7 +28,7 @@ class LeadLauncherTest {
List.of("ccs", "ltms"), "tab", "bridged-workers", null,
"http://127.0.0.1:8765/mcp", null, null,
null, null, null,
Map.of("CLAUDE_CODE_AUTO_COMPACT_WINDOW", "300000"), null, null, true);
Map.of("CLAUDE_CODE_AUTO_COMPACT_WINDOW", "300000"), null, null, true, null);
}
private static BridgedConfig configWith(BridgedConfig.Leader lead) {
@@ -41,8 +41,8 @@ class LeadLauncherTest {
null, null, fleet, null, "fixed", null).withDefaults();
}
private static BridgedConfig.Leader lead(String profile, String terminal, int instances) {
return new BridgedConfig.Leader(profile, terminal, instances, "lead:", 10, null, null,
private static BridgedConfig.Leader lead(String profile, String tab, int instances) {
return new BridgedConfig.Leader(profile, tab, instances, "lead:", 10, null, null,
"leads", "/repo");
}
@@ -70,16 +70,16 @@ class LeadLauncherTest {
void startsTheDeclaredLeadWhenNoneIsRunning() {
FakeHerdr herdr = new FakeHerdr();
assertEquals(1, launcher(herdr, configWith(lead("opus", null, 1))).ensureLeads());
assertEquals(1, launcher(herdr, configWith(lead("opus", "lead: opus", 1))).ensureLeads());
assertTrue(herdr.called("agent.start"), "a lead must actually be started");
assertEquals("lead-opus", startedName(herdr));
}
/** The tab is labelled so the scanner finds the lead on the next resolve. */
/** The tab is labelled with the configured `tab:` so the scanner finds the lead on the next resolve. */
@Test
void labelsTheTabWithThePrefixTheScannerReadsBack() {
void labelsTheTabWithTheConfiguredTabValue() {
FakeHerdr herdr = new FakeHerdr();
launcher(herdr, configWith(lead("opus", null, 1))).ensureLeads();
launcher(herdr, configWith(lead("opus", "lead: opus", 1))).ensureLeads();
assertEquals("lead: opus",
((Map<?, ?>) herdr.lastCall("tab.rename").params()).get("label"));
@@ -90,7 +90,7 @@ class LeadLauncherTest {
void startsAsManyInstancesAsAreDeclared() {
FakeHerdr herdr = new FakeHerdr();
assertEquals(2, launcher(herdr, configWith(lead("opus", null, 2))).ensureLeads());
assertEquals(2, launcher(herdr, configWith(lead("opus", "lead: opus", 2))).ensureLeads());
assertEquals(2, herdr.calls.stream().filter(c -> c.method().equals("agent.start")).count());
}
@@ -104,7 +104,7 @@ class LeadLauncherTest {
.withTab("wL", "wL:t1", "lead: opus")
.withAgent("lead-opus", "term_lead", "wL:p1", "wL:t1");
assertEquals(0, launcher(herdr, configWith(lead("opus", null, 1))).ensureLeads());
assertEquals(0, launcher(herdr, configWith(lead("opus", "lead: opus", 1))).ensureLeads());
assertFalse(herdr.called("agent.start"), "the live lead must not be duplicated");
}
@@ -118,20 +118,23 @@ class LeadLauncherTest {
.withWorkspace("wL", "leads")
.withTab("wL", "wL:t1", "lead: opus"); // label only — nothing running in it
assertEquals(1, launcher(herdr, configWith(lead("opus", null, 1))).ensureLeads(),
assertEquals(1, launcher(herdr, configWith(lead("opus", "lead: opus", 1))).ensureLeads(),
"a stale label is not a lead; the lead must be relaunched");
}
/**
* A lead the operator opened by hand and pinned with `terminal:` is live even though its tab
* carries no matching label. Counting labels alone would relaunch it on every boot.
* A lead the operator opened by hand is live once its tab carries the configured `tab:` label —
* CB-579 retired the `terminal:` pin, so a hand-opened lead is found the same way an
* auto-launched one is, by its tab, not by a terminal id nobody wrote down in advance.
*/
@Test
void aPinnedTerminalWithARunningAgentCountsAsLive() {
void aHandOpenedLeadWithTheConfiguredTabLabelCountsAsLive() {
FakeHerdr herdr = new FakeHerdr()
.withAgent("hand-opened", "term_pinned", "wX:p1", "wX:t1");
.withWorkspace("wX", "main")
.withTab("wX", "wX:t1", "lead: opus")
.withAgent("hand-opened", "term_hand", "wX:p1", "wX:t1");
assertEquals(0, launcher(herdr, configWith(lead("opus", "term_pinned", 1))).ensureLeads());
assertEquals(0, launcher(herdr, configWith(lead("opus", "lead: opus", 1))).ensureLeads());
assertFalse(herdr.called("agent.start"));
}
@@ -143,7 +146,7 @@ class LeadLauncherTest {
.withTab("wM", "wM:t1", "lead: opus") // a member tab that looks like a lead
.withAgent("claude-opus-x", "term_m", "wM:p1", "wM:t1");
assertEquals(1, launcher(herdr, configWith(lead("opus", null, 1))).ensureLeads(),
assertEquals(1, launcher(herdr, configWith(lead("opus", "lead: opus", 1))).ensureLeads(),
"a member in a lead-labelled tab is not a lead, so the real lead is still missing");
}
@@ -152,7 +155,7 @@ class LeadLauncherTest {
void anUncountableHerdrStartsNothing() {
FakeHerdr herdr = new FakeHerdr().healthy(false);
assertEquals(0, launcher(herdr, configWith(lead("opus", null, 1))).ensureLeads());
assertEquals(0, launcher(herdr, configWith(lead("opus", "lead: opus", 1))).ensureLeads());
assertFalse(herdr.called("agent.start"));
}
@@ -166,7 +169,7 @@ class LeadLauncherTest {
@Test
void theLeadNeverReceivesTheWorkerReplyCharter() {
FakeHerdr herdr = new FakeHerdr();
launcher(herdr, configWith(lead("opus", null, 1))).ensureLeads();
launcher(herdr, configWith(lead("opus", "lead: opus", 1))).ensureLeads();
List<String> args = startedArgs(herdr);
assertFalse(args.contains("--append-system-prompt"),
@@ -178,7 +181,7 @@ class LeadLauncherTest {
@Test
void theLeadMountsTheBridgeMcpAndPinsItsModel() {
FakeHerdr herdr = new FakeHerdr();
launcher(herdr, configWith(lead("opus", null, 1))).ensureLeads();
launcher(herdr, configWith(lead("opus", "lead: opus", 1))).ensureLeads();
List<String> args = startedArgs(herdr);
assertTrue(args.contains("--mcp-config"));
@@ -192,7 +195,7 @@ class LeadLauncherTest {
@Test
void theLeadEnvCarriesNoAnthropicBinding() {
FakeHerdr herdr = new FakeHerdr();
launcher(herdr, configWith(lead("opus", null, 1))).ensureLeads();
launcher(herdr, configWith(lead("opus", "lead: opus", 1))).ensureLeads();
Map<String, String> env = tabEnv(herdr);
assertNull(env.get("ANTHROPIC_BASE_URL"));
@@ -205,7 +208,7 @@ class LeadLauncherTest {
@Test
void theLeadTabIsCreatedOutsideEveryMemberWorkspace() {
FakeHerdr herdr = new FakeHerdr();
launcher(herdr, configWith(lead("opus", null, 1))).ensureLeads();
launcher(herdr, configWith(lead("opus", "lead: opus", 1))).ensureLeads();
String label = (String) ((Map<?, ?>) herdr.lastCall("workspace.create").params()).get("label");
assertEquals("leads", label);
@@ -214,12 +217,12 @@ class LeadLauncherTest {
// ── recognise-only and misconfiguration ───────────────────────────────────────────────────
/** A lead with a pin but no profile is recognise-only by design — not an error, not a launch. */
/** A lead with a tab but no profile is recognise-only by design — not an error, not a launch. */
@Test
void aLeadThatNamesNoProfileIsRecognisedButNeverLaunched() {
FakeHerdr herdr = new FakeHerdr();
assertEquals(0, launcher(herdr, configWith(lead(null, "term_dead", 1))).ensureLeads());
assertEquals(0, launcher(herdr, configWith(lead(null, "lead: dead", 1))).ensureLeads());
assertFalse(herdr.called("agent.start"));
}
@@ -228,7 +231,7 @@ class LeadLauncherTest {
void zeroInstancesLaunchesNothing() {
FakeHerdr herdr = new FakeHerdr();
assertEquals(0, launcher(herdr, configWith(lead("opus", null, 0))).ensureLeads());
assertEquals(0, launcher(herdr, configWith(lead("opus", "lead: opus", 0))).ensureLeads());
assertFalse(herdr.called("agent.start"));
}
@@ -237,7 +240,7 @@ class LeadLauncherTest {
void anUnknownProfileIsSkippedRatherThanThrown() {
FakeHerdr herdr = new FakeHerdr();
assertEquals(0, launcher(herdr, configWith(lead("nope", null, 1))).ensureLeads());
assertEquals(0, launcher(herdr, configWith(lead("nope", "lead: opus", 1))).ensureLeads());
assertFalse(herdr.called("agent.start"));
}
@@ -0,0 +1,37 @@
package dev.ltms.bridged.logging;
import ch.qos.logback.classic.Level;
import ch.qos.logback.classic.Logger;
import ch.qos.logback.classic.spi.ILoggingEvent;
import ch.qos.logback.core.read.ListAppender;
import io.modelcontextprotocol.spec.McpSchema;
import org.junit.jupiter.api.Test;
import org.slf4j.LoggerFactory;
import java.util.Map;
import static org.junit.jupiter.api.Assertions.assertEquals;
import static org.junit.jupiter.api.Assertions.assertTrue;
class McpCancelledNotificationFilterTest {
@Test
void suppressesOnlyTheCancelledNotificationWarning() {
Logger logger = (Logger) LoggerFactory.getLogger(McpCancelledNotificationFilter.LOGGER);
ListAppender<ILoggingEvent> appender = new ListAppender<>();
appender.start();
logger.addAppender(appender);
try {
logger.warn(McpCancelledNotificationFilter.UNHANDLED_NOTIFICATION,
new McpSchema.JSONRPCNotification("notifications/cancelled", Map.of("requestId", 7)));
logger.warn(McpCancelledNotificationFilter.UNHANDLED_NOTIFICATION,
new McpSchema.JSONRPCNotification("notifications/progress", Map.of("progress", 1)));
assertEquals(1, appender.list.size());
assertEquals(Level.WARN, appender.list.getFirst().getLevel());
assertTrue(appender.list.getFirst().getFormattedMessage().contains("notifications/progress"));
} finally {
logger.detachAppender(appender);
}
}
}
@@ -72,7 +72,8 @@ class BridgeMcpAuthzTest {
new PrimaryRegistry(null),
enforce ? CallerResolver.withLeadsAndMembers(identity, false, null,
Map::of, new MemberRegistry(null)) : null,
metrics, BridgeMcp.CapacitySource.none());
metrics, BridgeMcp.CapacitySource.none(), new BridgeMcp.HealthCoverageSource(() -> "off"),
BridgeMcp.QuarantineSource.none());
return mcp;
}
@@ -16,6 +16,7 @@ import dev.ltms.bridged.inject.MemberPresence;
import dev.ltms.bridged.session.MemberSession;
import dev.ltms.bridged.session.WorktreeRequest;
import dev.ltms.bridged.member.ClaudeCodeLauncher;
import dev.ltms.bridged.placement.BackendQuarantine;
import io.modelcontextprotocol.spec.McpSchema;
import dev.ltms.bridged.msg.InMemoryReplyInbox;
import org.junit.jupiter.api.BeforeEach;
@@ -426,7 +427,8 @@ class BridgeMcpTest {
void spawnPassesTheRequestedCwdToTheWorker() {
FakeHerdr h = new FakeHerdr();
McpSchema.CallToolResult res = BridgeMcp.spawn(
sessionManager(h, "http://gx00.gw:8000", Set.of("gx00.gw")), null, null, "/req/dir", null, null, null);
sessionManager(h, "http://gx00.gw:8000", Set.of("gx00.gw")), null, null, "/req/dir", null, null, null,
null, null);
assertNotEquals(Boolean.TRUE, res.isError());
// Protocol 19: the requested cwd roots the worker's pane at creation (tab.create).
@SuppressWarnings("unchecked")
@@ -437,11 +439,29 @@ class BridgeMcpTest {
@Test
void profilesListsConfiguredProfilesAndDefault() {
FakeHerdr h = new FakeHerdr();
McpSchema.CallToolResult res = BridgeMcp.profiles(workerService(h, "http://gx00.gw:8000", Set.of("gx00.gw")));
McpSchema.CallToolResult res = BridgeMcp.profiles(
workerService(h, "http://gx00.gw:8000", Set.of("gx00.gw")), BridgeMcp.QuarantineSource.none());
assertNotEquals(Boolean.TRUE, res.isError());
String out = textOf(res);
assertTrue(out.contains("ltms-local"), out);
assertTrue(out.contains("\"default\":\"ltms-local\""), out);
assertFalse(out.contains("quarantined"), "no profile is quarantined, so the key is omitted: " + out);
}
@Test
void profilesReportsAQuarantinedCredential() {
FakeHerdr h = new FakeHerdr();
BackendQuarantine quarantine = new BackendQuarantine(() -> 0L, TimeUnit.MINUTES.toNanos(30));
quarantine.quarantine("shared-openai");
BridgeMcp.QuarantineSource source = new BridgeMcp.QuarantineSource(
profile -> "ltms-local".equals(profile) ? "shared-openai" : null, quarantine);
McpSchema.CallToolResult res = BridgeMcp.profiles(
workerService(h, "http://gx00.gw:8000", Set.of("gx00.gw")), source);
assertNotEquals(Boolean.TRUE, res.isError());
String out = textOf(res);
assertTrue(out.contains("\"quarantined\""), out);
assertTrue(out.contains("shared-openai"), out);
assertTrue(out.contains("\"quarantinedForSeconds\":1800"), out);
}
@Test
@@ -474,11 +494,14 @@ class BridgeMcpTest {
sessions.acquire("ltms-local", null, null, null);
McpSchema.CallToolResult res = BridgeMcp.listFleet(workerService(h, "http://gx00.gw:8000", Set.of("gx00.gw")),
sessions, null, new BridgeMcp.CapacitySource(profile -> 2, profile -> 2,
() -> Set.of("ltms-local"), () -> 0), Map.of(), "");
() -> Set.of("ltms-local"), () -> 0), new BridgeMcp.HealthCoverageSource(() -> "off"),
BridgeMcp.QuarantineSource.none(), Map.of(), "");
String out = textOf(res);
assertTrue(out.contains("\"maxLoad\":2"), out);
assertTrue(out.contains("\"live\":2"), out);
assertTrue(out.contains("\"free\":0"), out);
assertFalse(out.contains("credentialId"), "nothing is quarantined, so no new key appears: " + out);
assertFalse(out.contains("quarantinedForSeconds"), out);
}
@Test
@@ -487,7 +510,8 @@ class BridgeMcpTest {
SessionManager sessions = new SessionManager(workerService(h, "http://gx00.gw:8000", Set.of("gx00.gw")));
String out = textOf(BridgeMcp.listFleet(workerService(h, "http://gx00.gw:8000", Set.of("gx00.gw")),
sessions, null, new BridgeMcp.CapacitySource(profile -> 0, profile -> 2,
() -> Set.of("terra"), () -> 0), Map.of(), ""));
() -> Set.of("terra"), () -> 0), new BridgeMcp.HealthCoverageSource(() -> "off"),
BridgeMcp.QuarantineSource.none(), Map.of(), ""));
assertTrue(out.contains("\"profile\":\"terra\""), out);
assertTrue(out.contains("\"live\":0"), out);
assertTrue(out.contains("\"free\":2"), out);
@@ -499,10 +523,90 @@ class BridgeMcpTest {
FakeHerdr h = new FakeHerdr();
String out = textOf(BridgeMcp.listFleet(workerService(h, "http://gx00.gw:8000", Set.of("gx00.gw")),
new SessionManager(workerService(h, "http://gx00.gw:8000", Set.of("gx00.gw"))), null,
BridgeMcp.CapacitySource.none(), Map.of(), ""));
BridgeMcp.CapacitySource.none(), new BridgeMcp.HealthCoverageSource(() -> "off"),
BridgeMcp.QuarantineSource.none(), Map.of(), ""));
assertFalse(out.contains("\"capacity\":"), out);
}
@Test
void quarantinedProfileReportsZeroFreeRegardlessOfMaxLoadAndLive() {
FakeHerdr h = new FakeHerdr();
SessionManager sessions = new SessionManager(workerService(h, "http://gx00.gw:8000", Set.of("gx00.gw")));
BackendQuarantine quarantine = new BackendQuarantine(() -> 0L, TimeUnit.MINUTES.toNanos(20));
quarantine.quarantine("shared-openai");
BridgeMcp.QuarantineSource source = new BridgeMcp.QuarantineSource(
profile -> "terra".equals(profile) ? "shared-openai" : null, quarantine);
String out = textOf(BridgeMcp.listFleet(workerService(h, "http://gx00.gw:8000", Set.of("gx00.gw")),
sessions, null, new BridgeMcp.CapacitySource(profile -> 0, profile -> 2,
() -> Set.of("terra"), () -> 0), new BridgeMcp.HealthCoverageSource(() -> "off"),
source, Map.of(), ""));
assertTrue(out.contains("\"profile\":\"terra\""), out);
assertTrue(out.contains("\"maxLoad\":2"), out);
assertTrue(out.contains("\"live\":0"), out);
assertTrue(out.contains("\"free\":0"), out);
assertTrue(out.contains("\"credentialId\":\"shared-openai\""), out);
assertTrue(out.contains("\"quarantinedForSeconds\":1200"), out);
}
@Test
void everyProfileSharingTheQuarantinedCredentialReportsZeroFree() {
FakeHerdr h = new FakeHerdr();
SessionManager sessions = new SessionManager(workerService(h, "http://gx00.gw:8000", Set.of("gx00.gw")));
BackendQuarantine quarantine = new BackendQuarantine(() -> 0L, TimeUnit.MINUTES.toNanos(20));
quarantine.quarantine("shared-openai");
BridgeMcp.QuarantineSource source = new BridgeMcp.QuarantineSource(_ -> "shared-openai", quarantine);
String out = textOf(BridgeMcp.listFleet(workerService(h, "http://gx00.gw:8000", Set.of("gx00.gw")),
sessions, null, new BridgeMcp.CapacitySource(profile -> 0, profile -> 2,
() -> Set.of("terra", "sol"), () -> 0), new BridgeMcp.HealthCoverageSource(() -> "off"),
source, Map.of(), ""));
assertEquals(2, out.split("\"free\":0", -1).length - 1, out);
}
@Test
void listAndProfilesAgreeOnWhatIsQuarantined() {
FakeHerdr h = new FakeHerdr();
SessionManager sessions = new SessionManager(workerService(h, "http://gx00.gw:8000", Set.of("gx00.gw")));
BackendQuarantine quarantine = new BackendQuarantine(() -> 0L, TimeUnit.MINUTES.toNanos(20));
quarantine.quarantine("shared-openai");
BridgeMcp.QuarantineSource source = new BridgeMcp.QuarantineSource(
profile -> "ltms-local".equals(profile) ? "shared-openai" : null, quarantine);
var workers = workerService(h, "http://gx00.gw:8000", Set.of("gx00.gw"));
String listOut = textOf(BridgeMcp.listFleet(workers, sessions, null,
new BridgeMcp.CapacitySource(profile -> 0, profile -> 2, () -> Set.of("ltms-local"), () -> 0),
new BridgeMcp.HealthCoverageSource(() -> "off"), source, Map.of(), ""));
String profilesOut = textOf(BridgeMcp.profiles(workers, source));
assertTrue(listOut.contains("\"free\":0"), listOut);
assertTrue(listOut.contains("\"credentialId\":\"shared-openai\""), listOut);
assertTrue(profilesOut.contains("\"quarantined\""), profilesOut);
assertTrue(profilesOut.contains("shared-openai"), profilesOut);
}
@Test
void expiredQuarantineLeavesTheCapacityRowOrdinary() {
FakeHerdr h = new FakeHerdr();
SessionManager sessions = new SessionManager(workerService(h, "http://gx00.gw:8000", Set.of("gx00.gw")));
long[] clockNanos = {0L};
BackendQuarantine quarantine = new BackendQuarantine(() -> clockNanos[0], TimeUnit.MINUTES.toNanos(20));
quarantine.quarantine("shared-openai");
clockNanos[0] = TimeUnit.MINUTES.toNanos(21); // clock advances past the cooldown
BridgeMcp.QuarantineSource source = new BridgeMcp.QuarantineSource(
profile -> "terra".equals(profile) ? "shared-openai" : null, quarantine);
String out = textOf(BridgeMcp.listFleet(workerService(h, "http://gx00.gw:8000", Set.of("gx00.gw")),
sessions, null, new BridgeMcp.CapacitySource(profile -> 0, profile -> 2,
() -> Set.of("terra"), () -> 0), new BridgeMcp.HealthCoverageSource(() -> "off"),
source, Map.of(), ""));
assertTrue(out.contains("\"free\":2"), out);
assertFalse(out.contains("quarantinedForSeconds"), out);
}
@Test
void listReportsLeadsAndFlagsTheCallersOwnRow() {
FakeHerdr h = new FakeHerdr();
@@ -783,7 +887,8 @@ class BridgeMcpTest {
void spawnDefaultsTheRoleToDev() {
FakeHerdr h = new FakeHerdr();
McpSchema.CallToolResult res = BridgeMcp.spawn(
sessionManager(h, "http://gx00.gw:8000", Set.of("gx00.gw")), null, null, null, null, null, null);
sessionManager(h, "http://gx00.gw:8000", Set.of("gx00.gw")), null, null, null, null, null, null,
null, null);
assertNotEquals(Boolean.TRUE, res.isError());
assertTrue(textOf(res).contains("\"role\":\"dev\""), textOf(res));
@@ -794,7 +899,7 @@ class BridgeMcpTest {
FakeHerdr h = new FakeHerdr();
McpSchema.CallToolResult res = BridgeMcp.spawn(
sessionManager(h, "http://gx00.gw:8000", Set.of("gx00.gw")), null, "architect",
null, null, null, null);
null, null, null, null, null, null);
assertNotEquals(Boolean.TRUE, res.isError());
assertTrue(textOf(res).contains("\"role\":\"architect\""), textOf(res));
@@ -805,9 +910,36 @@ class BridgeMcpTest {
FakeHerdr h = new FakeHerdr();
McpSchema.CallToolResult res = BridgeMcp.spawn(
sessionManager(h, "http://gx00.gw:8000", Set.of("gx00.gw")), null, "worker",
null, null, null, null);
null, null, null, null, null, null);
assertEquals(Boolean.TRUE, res.isError());
assertTrue(textOf(res).contains("architect, dev, reviewer"), textOf(res));
}
// ── CB-584: bridge_spawn accepts sessionName/resumeSessionId; roster shows agentSessionId ──
@Test
void spawnWithResumeSessionIdPutsTheIdOnTheRoster() {
FakeHerdr h = new FakeHerdr();
SessionManager sessions = sessionManager(h, "http://gx00.gw:8000", Set.of("gx00.gw"));
McpSchema.CallToolResult res = BridgeMcp.spawn(sessions, "ltms-local", null, null, null, null, null,
null, "cb-resume-99");
assertNotEquals(Boolean.TRUE, res.isError(), textOf(res));
String listOut = textOf(BridgeMcp.listFleet(
workerService(h, "http://gx00.gw:8000", Set.of("gx00.gw")), sessions, Map.of(), ""));
assertTrue(listOut.contains("\"agentSessionId\":\"cb-resume-99\""), listOut);
}
@Test
void spawnRefusesResumeSessionIdWithoutAnExplicitProfile() {
FakeHerdr h = new FakeHerdr();
McpSchema.CallToolResult res = BridgeMcp.spawn(
sessionManager(h, "http://gx00.gw:8000", Set.of("gx00.gw")), null, null,
null, null, null, null, null, "cb-resume-99");
assertEquals(Boolean.TRUE, res.isError());
assertTrue(textOf(res).contains("explicit profile"), textOf(res));
}
}
@@ -676,6 +676,85 @@ class ClaudeCodeLauncherTest {
"the guard-checked baseUrl must win over any env: entry, or the boundary is bypassable");
}
// --- CB-592: the admin GITEA_ACCESS_TOKEN never reaches a member -----------------------------
/**
* herdr's env map is an overlay onto its own (login-shell) process environment, so a worker
* inherits whatever the daemon's shell carries — including the admin GITEA_ACCESS_TOKEN — for
* every key baseEnv does not explicitly shadow. This pins that the launcher DOES send an
* explicit (non-blank) GITEA_ACCESS_TOKEN to herdr on every spawn, whatever the profile is, so
* a future baseEnv refactor cannot silently drop it and reopen the leak. Asserted against what
* tab.create's params actually carry, not an internal map built in the test (gitea #77).
*
* <p>Scope, measured on a live pane 2026-08-15: this pins what the launcher SENDS, and that is
* all it can pin. It does not prove the value survives, and it does not: the pane runs a login
* shell, ~/.zprofile sources secrets.sh, and its unconditional `export GITEA_ACCESS_TOKEN=...`
* puts the real token back over this sentinel. Closing that needs the operator to guard the
* export on BRIDGED_MEMBER — see everySpawnMarksThePaneAsAMember below.
*/
@Test
void everySpawnShadowsTheAdminGiteaAccessToken() {
FakeHerdr herdr = new FakeHerdr();
service(herdr, List.of("claude"), null).spawn();
String shadowed = startEnv(herdr).get("GITEA_ACCESS_TOKEN");
assertNotNull(shadowed, "GITEA_ACCESS_TOKEN must be explicitly overlaid, not left unmentioned");
assertFalse(shadowed.isBlank(), "a blank overlay value's override behaviour is unverified — must be non-blank");
}
/**
* No profile — present or future — may restore the admin token by naming it in {@code env:}.
* The shadow is applied after the profile's own env in {@link HerdrPeerLauncher#baseEnv}
* precisely so this can never happen; this test pins that ordering.
*/
@Test
void aProfileEnvEntryCannotRestoreTheAdminGiteaAccessToken() {
FakeHerdr herdr = new FakeHerdr();
BridgedConfig.Profile cfg = new BridgedConfig.Profile(
"ltms-local", "http://gx00.gw:8000", "coder", null, "BRIDGED_WORKER_TOKEN",
List.of("claude"), "tab", "bridged-workers", "w #{n}", null, null, null, null, null,
null, Map.of("GITEA_ACCESS_TOKEN", "admin-secret-from-profile-config"), null, null);
new ClaudeCodeLauncher(new AgentControl(herdr), new WorkspaceControl(herdr),
new SubscriptionGuard(Set.of("gx00.gw")), Map.of(cfg.profile(), cfg), cfg.profile(),
_ -> null).spawn();
assertNotEquals("admin-secret-from-profile-config", startEnv(herdr).get("GITEA_ACCESS_TOKEN"),
"a profile's own env: must not be able to smuggle the admin token back in");
}
/**
* The half of CB-592 that can actually survive the pane's login shell. BRIDGED_MEMBER is a name
* secrets.sh never exports, so nothing overwrites it — measured: GITEA_TOKEN is injected the
* same way, is absent from a login shell of its own, and was observed set inside a live member
* pane. It lets the operator guard the admin export with
* `[ -n "${BRIDGED_MEMBER:-}" ] || export GITEA_ACCESS_TOKEN=...`, which is the whole fix.
* Pinned here so a refactor cannot drop the marker and quietly un-guard every member (#77).
*/
@Test
void everySpawnMarksThePaneAsAMember() {
FakeHerdr herdr = new FakeHerdr();
service(herdr, List.of("claude"), null).spawn();
assertEquals("1", startEnv(herdr).get("BRIDGED_MEMBER"),
"every member pane must be marked, or a shell file cannot tell it apart from the operator's");
}
/** A profile must not be able to hide that its pane is a member, for the same reason as above. */
@Test
void aProfileEnvEntryCannotClearTheMemberMarker() {
FakeHerdr herdr = new FakeHerdr();
BridgedConfig.Profile cfg = new BridgedConfig.Profile(
"ltms-local", "http://gx00.gw:8000", "coder", null, "BRIDGED_WORKER_TOKEN",
List.of("claude"), "tab", "bridged-workers", "w #{n}", null, null, null, null, null,
null, Map.of("BRIDGED_MEMBER", ""), null, null);
new ClaudeCodeLauncher(new AgentControl(herdr), new WorkspaceControl(herdr),
new SubscriptionGuard(Set.of("gx00.gw")), Map.of(cfg.profile(), cfg), cfg.profile(),
_ -> null).spawn();
assertEquals("1", startEnv(herdr).get("BRIDGED_MEMBER"),
"a profile's own env: must not be able to unmark its pane");
}
// ── CB-533: the model is pinned on the command line, not only in the environment ────────────
/** A launcher for a profile identical but for its {@code model:} — the only variable here. */
@@ -733,7 +812,7 @@ class ClaudeCodeLauncherTest {
return new BridgedConfig.Profile(
profile, baseUrl, "sonnet", null, "BRIDGED_WORKER_TOKEN",
List.of("ccs", profile), "tab", "bridged-workers", "w #{n}", null, null, null,
null, null, null, Map.of(), null, null, true);
null, null, null, Map.of(), null, null, true, null);
}
@Test
@@ -819,7 +898,7 @@ class ClaudeCodeLauncherTest {
null, null, null,
Map.of("ANTHROPIC_BASE_URL", "http://evil.example.com",
"ANTHROPIC_AUTH_TOKEN", "sk-ant-bad", "JAVA_HOME", "/opt/jdk"),
null, null, true);
null, null, true, null);
new ClaudeCodeLauncher(new AgentControl(herdr), new WorkspaceControl(herdr),
new SubscriptionGuard(Set.of("gx00.gw")), Map.of(cfg.profile(), cfg), cfg.profile(),
_ -> "would-be-token").spawn();
@@ -10,11 +10,13 @@ import dev.ltms.bridged.herdr.AgentControl;
import dev.ltms.bridged.herdr.FakeHerdr;
import dev.ltms.bridged.herdr.WorkspaceControl;
import dev.ltms.bridged.peer.Capability;
import dev.ltms.bridged.peer.CharterReceipt;
import dev.ltms.bridged.peer.MemberRole;
import dev.ltms.bridged.peer.PeerHandle;
import dev.ltms.bridged.peer.PeerLauncher;
import dev.ltms.bridged.peer.PeerUnreachableException;
import dev.ltms.bridged.peer.SpawnRequest;
import dev.ltms.bridged.placement.BackendQuarantine;
import dev.ltms.bridged.placement.PlacementException;
import dev.ltms.bridged.placement.PlacementPolicies;
import org.junit.jupiter.api.Test;
@@ -26,6 +28,8 @@ import java.util.LinkedHashMap;
import java.util.List;
import java.util.Map;
import java.util.Set;
import java.util.concurrent.TimeUnit;
import java.util.concurrent.atomic.AtomicLong;
import static org.junit.jupiter.api.Assertions.*;
@@ -93,6 +97,8 @@ class CompositePeerLauncherTest {
@Override public String id() { return "pane-" + p; }
@Override public String terminalId() { return "term-" + p; }
@Override public String profile() { return p; }
@Override public String agentSessionId() { return null; }
@Override public CharterReceipt charterReceipt() { return null; }
};
}
@@ -126,6 +132,13 @@ class CompositePeerLauncherTest {
weight, maxLoad);
}
private static BridgedConfig.Profile stubWorker(String profile, String credentialId) {
return new BridgedConfig.Profile(profile, "http://gx00.gw:8000", "coder",
null, "BRIDGED_WORKER_TOKEN", List.of("claude"), "tab", "bridged-workers",
"w #{n}", null, null, null, null, null, null, null, null, null,
null, null, credentialId);
}
/**
* An <em>order-preserving</em> profile map. Never {@code Map.of} here: its iteration order is
* salted per JVM run, and the weighted policy breaks an exact-weight tie on candidate order —
@@ -324,6 +337,42 @@ class CompositePeerLauncherTest {
assertEquals(10, b);
}
@Test
void weightedPolicySkipsWeightZeroProfileOnUnqualifiedSpawn() {
// CB-554: weight: 0 must exclude a profile from automatic placement, not coerce to 1.0.
FakeHerdr herdr = new FakeHerdr();
Map<String, BridgedConfig.Profile> profiles = ordered(
"a", stubWorker("a", 0.0f, null),
"b", stubWorker("b", 1.0f, null));
StubLauncher adapter = new StubLauncher("claude", herdr, profiles, "a", Set.of());
CompositePeerLauncher composite = new CompositePeerLauncher(
List.of(adapter), "a", profiles, PlacementPolicies.weighted(), _ -> 0);
for (int i = 0; i < 5; i++) {
PeerHandle h = composite.spawn(new SpawnRequest(null, null, null));
assertEquals("b", h.profile(), "profile a has weight 0, so every unqualified spawn must land on b");
}
assertEquals(0, adapter.spawnCount("a"), "a is never chosen automatically");
}
@Test
void explicitSpawnStillSucceedsOnWeightZeroProfile() {
// CB-554: weight: 0 excludes a profile from AUTOMATIC selection only — an explicit
// bridge_spawn{profile:"a"} must still work exactly as today (e.g. `opus` on the
// operator's own subscription, kept weight-0 so it is never picked automatically).
FakeHerdr herdr = new FakeHerdr();
Map<String, BridgedConfig.Profile> profiles = ordered(
"a", stubWorker("a", 0.0f, null),
"b", stubWorker("b", 1.0f, null));
StubLauncher adapter = new StubLauncher("claude", herdr, profiles, "b", Set.of());
CompositePeerLauncher composite = new CompositePeerLauncher(
List.of(adapter), "b", profiles, PlacementPolicies.weighted(), _ -> 0);
PeerHandle h = composite.spawn(new SpawnRequest("a", null, null));
assertEquals("a", h.profile(), "naming a weight-0 profile explicitly bypasses placement and still spawns it");
assertEquals(1, adapter.spawnCount("a"));
}
@Test
void failoverRetriesNextCandidateWhenProfileIsUnreachable() {
FakeHerdr herdr = new FakeHerdr();
@@ -425,6 +474,27 @@ class CompositePeerLauncherTest {
assertEquals("a", h.profile(), "a profile with no maxLoad is never capped, however many live workers");
}
@Test
void explicitSpawnOnMaxLoadZeroProfileIsRefusedEvenWithZeroLiveWorkers() {
// CB-585: before the fix, maxLoad: 0 normalised to null (unlimited) in the compact
// constructor, so this exact case — naming a zero-cap profile explicitly, with nothing
// live on it yet — would have spawned instead of refusing. "at most zero members" must
// hold even when the profile is named directly, not only against automatic placement.
FakeHerdr herdr = new FakeHerdr();
Map<String, BridgedConfig.Profile> profiles = ordered(
"a", stubWorker("a", 1.0f, 0),
"b", stubWorker("b"));
StubLauncher adapter = new StubLauncher("claude", herdr, profiles, "a", Set.of());
CompositePeerLauncher composite = new CompositePeerLauncher(
List.of(adapter), "a", profiles, PlacementPolicies.fixed(), _ -> 0);
PlacementException e = assertThrows(PlacementException.class,
() -> composite.spawn(new SpawnRequest("a", null, null)));
assertTrue(e.getMessage().contains("'a'"), "message names the profile: " + e.getMessage());
assertTrue(e.getMessage().contains("0 cap"), "message names the zero cap: " + e.getMessage());
assertEquals(0, adapter.spawnCount("a"), "a maxLoad: 0 profile accepts no explicit spawn");
}
@Test
void emptyCandidateSetThrowsClearException() {
FakeHerdr herdr = new FakeHerdr();
@@ -552,4 +622,105 @@ class CompositePeerLauncherTest {
new SpawnRequest(null, null, null, null, null, MemberRole.DEV)));
assertTrue(e.getMessage().contains("maxLoad"), e.getMessage());
}
// ── CB-578 stage B: a BACKEND_EXHAUSTED classification quarantines the credential ──────────
@Test
void explicitSpawnOntoAQuarantinedProfileIsRefusedNamingTheCredentialAndRemainingTime() {
FakeHerdr herdr = new FakeHerdr();
Map<String, BridgedConfig.Profile> profiles = ordered(
"sol", stubWorker("sol", "shared-openai"),
"terra", stubWorker("terra", "shared-openai"));
StubLauncher adapter = new StubLauncher("claude", herdr, profiles, "sol", Set.of());
BackendQuarantine quarantine = new BackendQuarantine(() -> 0L, TimeUnit.MINUTES.toNanos(30));
quarantine.quarantine("shared-openai");
CompositePeerLauncher composite = new CompositePeerLauncher(List.of(adapter), "sol", profiles,
PlacementPolicies.fixed(), _ -> 0, null, quarantine);
PlacementException e = assertThrows(PlacementException.class,
() -> composite.spawn(new SpawnRequest("sol", null, null)));
assertTrue(e.getMessage().contains("sol"), "message names the profile: " + e.getMessage());
assertTrue(e.getMessage().contains("shared-openai"), "message names the credential: " + e.getMessage());
assertTrue(e.getMessage().contains("1800"), "message names roughly when it lifts: " + e.getMessage());
assertEquals(0, adapter.spawnCount("sol"), "the quarantined profile is never delegated to");
}
/**
* The part CB-578 stage B calls out as easy to get wrong: sol and terra are two different
* profiles sharing one OpenAI credential. Quarantining because of an exhaustion classified on
* ONE of them must lock out the other too, or the fleet just walks onto the same dead account
* under the sibling's name.
*/
@Test
void twoProfilesSharingACredentialAreBothQuarantinedByOneExhaustionEvent() {
FakeHerdr herdr = new FakeHerdr();
Map<String, BridgedConfig.Profile> profiles = ordered(
"sol", stubWorker("sol", "shared-openai"),
"terra", stubWorker("terra", "shared-openai"));
StubLauncher adapter = new StubLauncher("claude", herdr, profiles, "sol", Set.of());
BackendQuarantine quarantine = new BackendQuarantine(() -> 0L, TimeUnit.MINUTES.toNanos(30));
// Only "sol" was classified BACKEND_EXHAUSTED — but the two profiles share one credential.
quarantine.quarantine("shared-openai");
CompositePeerLauncher composite = new CompositePeerLauncher(List.of(adapter), "sol", profiles,
PlacementPolicies.fixed(), _ -> 0, null, quarantine);
assertThrows(PlacementException.class, () -> composite.spawn(new SpawnRequest("sol", null, null)),
"sol was the one classified exhausted");
assertThrows(PlacementException.class, () -> composite.spawn(new SpawnRequest("terra", null, null)),
"terra shares sol's credential, so it must be locked out too");
assertEquals(0, adapter.spawnCount("sol"));
assertEquals(0, adapter.spawnCount("terra"));
}
@Test
void placementSkipsAQuarantinedProfileAndRoutesToAnUnquarantinedOne() {
FakeHerdr herdr = new FakeHerdr();
Map<String, BridgedConfig.Profile> profiles = ordered(
"sol", stubWorker("sol", "shared-openai"),
"b", stubWorker("b"));
StubLauncher adapter = new StubLauncher("claude", herdr, profiles, "sol", Set.of());
BackendQuarantine quarantine = new BackendQuarantine(() -> 0L, TimeUnit.MINUTES.toNanos(30));
quarantine.quarantine("shared-openai");
CompositePeerLauncher composite = new CompositePeerLauncher(List.of(adapter), "sol", profiles,
PlacementPolicies.weighted(), _ -> 0, null, quarantine);
PeerHandle h = composite.spawn(new SpawnRequest(null, null, null));
assertEquals("b", h.profile(), "sol is quarantined, so an unqualified spawn must land on b");
assertEquals(0, adapter.spawnCount("sol"));
}
@Test
void aQuarantineLiftsOnTheInjectedClockAndTheProfileBecomesSpawnableAgain() {
FakeHerdr herdr = new FakeHerdr();
Map<String, BridgedConfig.Profile> profiles = ordered(
"sol", stubWorker("sol", "shared-openai"),
"b", stubWorker("b"));
StubLauncher adapter = new StubLauncher("claude", herdr, profiles, "sol", Set.of());
AtomicLong nowNanos = new AtomicLong(0L);
BackendQuarantine quarantine = new BackendQuarantine(nowNanos::get, TimeUnit.MINUTES.toNanos(30));
quarantine.quarantine("shared-openai");
CompositePeerLauncher composite = new CompositePeerLauncher(List.of(adapter), "sol", profiles,
PlacementPolicies.fixed(), _ -> 0, null, quarantine);
assertThrows(PlacementException.class, () -> composite.spawn(new SpawnRequest("sol", null, null)),
"still inside the cooldown");
nowNanos.set(TimeUnit.MINUTES.toNanos(31));
PeerHandle h = composite.spawn(new SpawnRequest("sol", null, null));
assertEquals("sol", h.profile(), "the cooldown expired on the injected clock — sol is spawnable again");
assertEquals(1, adapter.spawnCount("sol"));
}
@Test
void aFleetWithNoExhaustedPatternAnywhereBehavesExactlyAsBeforeQuarantineExisted() {
// BackendQuarantine.none() is the inert stand-in every constructor already defaults to when
// no quarantine is wired — the 2/5/6-arg constructors used throughout this file all exercise
// it. This test pins that an explicit .none() also never refuses a spawn, for any profile.
FakeHerdr herdr = new FakeHerdr();
PeerLauncher composite = composite(herdr);
assertDoesNotThrow(() -> composite.spawn(new SpawnRequest("claude", null, null)));
assertDoesNotThrow(() -> composite.spawn(new SpawnRequest("gemini", null, null)));
}
}
@@ -1,13 +1,19 @@
package dev.ltms.bridged.member;
import ch.qos.logback.classic.Level;
import ch.qos.logback.classic.Logger;
import ch.qos.logback.classic.spi.ILoggingEvent;
import ch.qos.logback.core.read.ListAppender;
import dev.ltms.bridged.config.BridgedConfig;
import dev.ltms.bridged.herdr.AgentControl;
import dev.ltms.bridged.herdr.FakeHerdr;
import dev.ltms.bridged.herdr.WorkspaceControl;
import dev.ltms.bridged.peer.Capability;
import dev.ltms.bridged.peer.CharterReceipt;
import dev.ltms.bridged.peer.MemberRole;
import dev.ltms.bridged.peer.SpawnRequest;
import org.junit.jupiter.api.Test;
import org.slf4j.LoggerFactory;
import java.util.ArrayList;
import java.util.List;
@@ -17,6 +23,8 @@ import java.util.concurrent.atomic.AtomicReference;
import java.util.function.Supplier;
import static org.junit.jupiter.api.Assertions.assertEquals;
import static org.junit.jupiter.api.Assertions.assertFalse;
import static org.junit.jupiter.api.Assertions.assertTrue;
class HerdrPeerLauncherCharterTest {
@@ -38,10 +46,72 @@ class HerdrPeerLauncherCharterTest {
"a role charter does not depend on an MCP mount");
}
@Test
void panePlacementSpawnLogNeverContainsTheCharterText() {
// A pane-placement spawn used to log the whole argv (CB-571), and the charter travels
// inside argv — so the charter text leaked to the daemon log. Prove the legacy pane path
// now redacts it to its digest.
String secret = "TOP SECRET charter marker 99x"; // distinctive, so a leak is unambiguous
AtomicReference<BridgedConfig.Fleet> fleet = new AtomicReference<>(fleet(Map.of("dev", secret)));
CharterArgLauncher launcher = new CharterArgLauncher(fleet::get);
Logger logger = (Logger) LoggerFactory.getLogger(HerdrPeerLauncher.class);
Level previous = logger.getLevel();
ListAppender<ILoggingEvent> appender = new ListAppender<>();
appender.start();
logger.addAppender(appender);
logger.setLevel(Level.INFO); // the test logback sets dev.ltms.bridged to WARN; a leak lives at INFO
try {
launcher.spawn(new SpawnRequest("mcp", null, null, null, null, MemberRole.DEV));
String all = String.join("\n", appender.list.stream().map(ILoggingEvent::getFormattedMessage).toList());
assertFalse(all.contains(secret),
"the pane-placement spawn log must not contain the charter text; got:\n" + all);
// The "mcp" profile composes role + reply charter; the digest must match that composed
// string (the exact bytes the adapter receives), proving the redaction hashes and
// removes the real, full charter — not some placeholder.
String composed = secret + "\n\n" + HerdrPeerLauncher.REPLY_CHARTER;
assertTrue(all.contains("<charter sha256=" + CharterReceipt.digestOf(composed) + ">"),
"the charter argv argument should be replaced by its digest; got:\n" + all);
} finally {
logger.setLevel(previous);
logger.detachAppender(appender);
}
}
private static BridgedConfig.Fleet fleet(Map<String, String> charters) {
return new BridgedConfig.Fleet(Map.of(), Map.of(), Map.of(), Map.of(), charters, null);
}
private static BridgedConfig.Profile profile(String name, String mcpUrl) {
return new BridgedConfig.Profile(name, "http://gx00.gw:8000", null, null,
"BRIDGED_WORKER_TOKEN", List.of("test"), "pane", null, null, mcpUrl, null, null);
}
/**
* A launcher whose {@code buildLaunch} hands the composed charter to herdr as one argv element
* (what the claude-cod adapter does), so a pane-placement spawn log would print it unless the
* base redacts it.
*/
private static final class CharterArgLauncher extends HerdrPeerLauncher {
CharterArgLauncher(Supplier<BridgedConfig.Fleet> fleet) {
super("test", new AgentControl(new FakeHerdr()), new WorkspaceControl(new FakeHerdr()),
Map.of("mcp", profile("mcp", "http://bridge")),
"mcp", _ -> null, 0, () -> 0L, () -> { }, fleet);
}
@Override
protected Launch buildLaunch(BridgedConfig.Profile cfg, LaunchSpec spec) {
return new Launch(Map.of(), List.of("test", spec.charter() == null ? "none" : spec.charter()));
}
@Override
public Set<Capability> capabilities() {
return Set.of();
}
}
private static final class CapturingLauncher extends HerdrPeerLauncher {
private final List<LaunchSpec> specs = new ArrayList<>();
@@ -62,10 +132,5 @@ class HerdrPeerLauncherCharterTest {
public Set<Capability> capabilities() {
return Set.of();
}
private static BridgedConfig.Profile profile(String name, String mcpUrl) {
return new BridgedConfig.Profile(name, "http://gx00.gw:8000", null, null,
"BRIDGED_WORKER_TOKEN", List.of("test"), "pane", null, null, mcpUrl, null, null);
}
}
}
@@ -7,6 +7,7 @@ import dev.ltms.bridged.herdr.AgentControl;
import dev.ltms.bridged.herdr.FakeHerdr;
import dev.ltms.bridged.herdr.WorkspaceControl;
import dev.ltms.bridged.peer.Capability;
import dev.ltms.bridged.peer.CharterReceipt;
import dev.ltms.bridged.peer.PeerHandle;
import dev.ltms.bridged.peer.PeerUnreachableException;
import dev.ltms.bridged.peer.SpawnRequest;
@@ -207,6 +208,21 @@ class OpenCodeLauncherTest {
"a git-token profile gets the peer-neutral GITEA_TOKEN grant, same as Claude");
}
/**
* CB-592: the shadow lives in {@link HerdrPeerLauncher#baseEnv}, shared by every adapter — this
* pins that the opencode path gets it too, not just Claude's. See the matching test in
* {@code ClaudeCodeLauncherTest} for the full rationale (gitea #77).
*/
@Test
void everySpawnShadowsTheAdminGiteaAccessToken(@TempDir Path root) {
FakeHerdr herdr = new FakeHerdr();
service(herdr, root, opencodeCfg(null, null, null)).spawn();
String shadowed = startEnv(herdr).get("GITEA_ACCESS_TOKEN");
assertNotNull(shadowed, "GITEA_ACCESS_TOKEN must be explicitly overlaid, not left unmentioned");
assertFalse(shadowed.isBlank(), "a blank overlay value's override behaviour is unverified — must be non-blank");
}
@Test
void capabilitiesDeclareOrphanReapAndMcpAskAndConditionalSelfPr(@TempDir Path root) {
FakeHerdr herdr = new FakeHerdr();
@@ -322,6 +338,26 @@ class OpenCodeLauncherTest {
assertFalse(herdr.called("agent.get"), "no polling when the gate is disabled");
}
@Test
void handleCarriesTheRealCharterReceiptNotTheInterfaceDefault(@TempDir Path root) {
// The base's WorkerHandle computes a real CharterReceipt (CB-571), but the opencode adapter
// wraps it in SessionAwareHandle for lazy session discovery. Before this fix that decorator
// did not override charterReceipt(), so it silently inherited PeerHandle's `null` default
// and the real receipt sitting on its delegate was lost.
FakeHerdr herdr = new FakeHerdr();
BridgedConfig.Fleet fleet = new BridgedConfig.Fleet(Map.of(), Map.of(), Map.of(), Map.of(),
Map.of("dev", "role rule"), null);
PeerHandle handle = service(herdr, root,
opencodeCfg("google/gemini-2.5-pro", "http://127.0.0.1:8765/mcp", null), () -> fleet)
.spawn(new SpawnRequest(null, null, null));
assertNotNull(handle.charterReceipt(),
"an opencode spawn's charterReceipt() must not silently be null");
String composed = "role rule\n\n" + HerdrPeerLauncher.REPLY_CHARTER;
assertEquals(CharterReceipt.digestOf(composed), handle.charterReceipt().charterSha256(),
"the receipt on the wrapped handle must match the exact composed charter bytes");
}
// --- CB-508: pinned OpenAI-compatible endpoint (e.g. a local vLLM) ---------------------------
/** A profile with a baseUrl but no model provider prefix cannot be resolved — fail loudly. */
@@ -5,6 +5,8 @@ import dev.ltms.bridged.herdr.AgentStatus;
import dev.ltms.bridged.herdr.FakeHerdr;
import dev.ltms.bridged.herdr.HerdrException;
import dev.ltms.bridged.inject.CompletionResolver;
import dev.ltms.bridged.inject.ExhaustedPatternLookup;
import dev.ltms.bridged.inject.ExhaustionSink;
import dev.ltms.bridged.mcp.PrimaryRegistry;
import dev.ltms.bridged.inject.Injector;
import org.junit.jupiter.api.BeforeEach;
@@ -31,7 +33,8 @@ class MessageServiceTest {
private final FakeHerdr herdr = new FakeHerdr().readText("BUILD GREEN: 391 files");
private final AgentControl agents = new AgentControl(herdr);
private final Rendezvous rendezvous = new Rendezvous();
private final CompletionResolver completion = new CompletionResolver(agents, rendezvous);
private final CompletionResolver completion =
new CompletionResolver(agents, rendezvous, ExhaustedPatternLookup.none(), ExhaustionSink.none());
private final Injector injector = new Injector(agents, completion);
private final InMemoryReplyInbox inbox = new InMemoryReplyInbox();
private final MessageService messages = new MessageService(agents, injector, rendezvous, inbox);
@@ -120,6 +123,36 @@ class MessageServiceTest {
assertFalse(reply.completed());
}
@Test
void droppedQueuedAndDeliveredTurnsExposeTheRealCauseExactlyOnce() throws Exception {
CompletableFuture<MessageService.Reply> first = sendAsync();
awaitWaiting();
injector.onStatus(T, AgentStatus.IDLE); // first delivery
injector.onStatus(T, AgentStatus.WORKING); // first turn in flight
CompletableFuture<Void> queued = injector.enqueue(T, "second task", TestTurnTokens.inert(T));
CompletableFuture<Rendezvous.Resolution> waiter = rendezvous.currentWaiter(T);
injector.drop(T, new HerdrException("agent target sol not found", "agent_not_found", null));
MessageService.Reply reply = first.get(5, TimeUnit.SECONDS);
assertEquals(MessageService.Outcome.WORKER_FAILED, reply.outcome());
assertEquals("agent target sol not found", reply.text());
assertTrue(queued.isCompletedExceptionally(), "the queued delivery future also fails");
assertFalse(rendezvous.resolveFailure(waiter, "second failure"), "the waiter fails exactly once");
}
@Test
void noDropReasonKeepsTheExistingFallbackText() throws Exception {
herdr.readText("");
CompletableFuture<Rendezvous.Resolution> waiter = rendezvous.open(T);
completion.onTurnFailed(T);
Rendezvous.Resolution resolution = waiter.get(2, TimeUnit.SECONDS);
assertEquals("worker did not reply; its turn ended in an unrecoverable state "
+ "(worker unreachable or stuck)", resolution.text());
}
// --- bridge_ask reverse rendezvous (CB-205) ------------------------------------------------
@Test
@@ -609,4 +642,298 @@ class MessageServiceTest {
assertTrue(view.detail() != null && view.detail().contains("released"),
"and the detail says why, rather than 'worker unknown'");
}
@Test
void abandonFailsEveryPendingAsyncTicketForTheReleasedTarget() throws Exception {
String first = messages.sendAsync(T, "first task");
awaitWaiting(); // first task owns the target lock and rendezvous waiter
String second = messages.sendAsync(T, "second task"); // parked on the same lock, not yet queued
String third = messages.sendAsync(T, "third task"); // a second queued ticket proves the full sweep
assertTrue(messages.abandon(T, "agent target term_a not found"));
assertFailedTicket(first, "agent target term_a not found");
assertFailedTicket(second, "agent target term_a not found");
assertFailedTicket(third, "agent target term_a not found");
}
@Test
void abandonDoesNotFailAnAsyncTicketWaitingForAnAnswer() throws Exception {
String ticket = messages.sendAsync(T, "task that asks");
awaitWaiting();
injectDelivery();
CompletableFuture<MessageService.AskResult> ask =
CompletableFuture.supplyAsync(() -> messages.ask(T, "which config?", 5000));
MessageService.TaskView asking = awaitTicketPhase(ticket, MessageService.Phase.ASKING);
assertFalse(messages.abandon(T, "agent target term_a not found"),
"an asking ticket is an active turn, not a pending send to sweep");
assertEquals(MessageService.Phase.ASKING, messages.poll(ticket).phase());
CompletableFuture<MessageService.Reply> answer = CompletableFuture.supplyAsync(
() -> messages.answer(asking.turnId(), "config.yaml", 5000));
assertEquals("config.yaml", ask.get(5, TimeUnit.SECONDS).answer());
awaitWaiting();
assertTrue(rendezvous.resolve(T, "done"));
assertEquals(MessageService.Outcome.REPLIED, answer.get(5, TimeUnit.SECONDS).outcome());
}
@Test
void unansweredAsyncQuestionReturnsTheTicketToPendingAndReleasesItsTarget() throws Exception {
String ticket = messages.sendAsync(T, "task that asks");
awaitWaiting();
injectDelivery();
assertEquals(MessageService.AskOutcome.TIMED_OUT,
messages.ask(T, "which config?", 200).outcome());
assertEquals(MessageService.Phase.PENDING, messages.poll(ticket).phase(),
"only the question wait ended; the delegated turn may still finish");
String next = messages.sendAsync(T, "next task");
awaitWaiting();
assertTrue(rendezvous.resolve(T, "done"));
assertEquals(MessageService.Phase.DONE, awaitTicketPhase(next, MessageService.Phase.DONE).phase());
}
@Test
void asyncQuestionBelongsToTheTaskThatOwnsItsForwardWaiter() throws Exception {
String first = messages.sendAsync(T, "first task");
awaitWaiting();
CompletableFuture<MessageService.AskResult> ask =
CompletableFuture.supplyAsync(() -> messages.ask(T, "which config?", 5000));
assertEquals(MessageService.Phase.ASKING, awaitTicketPhase(first, MessageService.Phase.ASKING).phase());
MessageService.TaskView asking = messages.poll(first);
CompletableFuture<MessageService.Reply> answer = CompletableFuture.supplyAsync(
() -> messages.answer(asking.turnId(), "config.yaml", 5000));
assertEquals("config.yaml", ask.get(5, TimeUnit.SECONDS).answer());
awaitWaiting();
assertTrue(rendezvous.resolve(T, "done"));
assertEquals(MessageService.Outcome.REPLIED, answer.get(5, TimeUnit.SECONDS).outcome());
}
// --- CB-588: async ticket terminal nudges ---------------------------------------------------
//
// MessageService.reply's rendezvous fast path is exactly what an async ticket always takes
// (sendAsync registers a rendezvous waiter — see asyncTasksByWaiter), so it never reached
// ReplyPushLoop.onReplyQueued. These prove the ticket reaches ReplyPushLoop through the new
// onTicketTerminal entry point instead, with no bridge_poll from the lead first.
private static final String LEAD = "term_lead";
/** A MessageService wired to a real ReplyPushLoop pointed at a private herdr fake for the lead. */
private record PushWiring(MessageService service, FakeHerdr leadHerdr,
java.util.concurrent.ScheduledExecutorService scheduler) implements AutoCloseable {
@Override
public void close() {
scheduler.shutdownNow();
}
}
private PushWiring wireWithPushLoop(int maxReminders, long backoffMs) {
return wireWithPushLoop(maxReminders, backoffMs, System::nanoTime);
}
/** As above, with an injectable clock (CB-588 follow-up: exercise pruneTerminalTickets' TTL). */
private PushWiring wireWithPushLoop(int maxReminders, long backoffMs, java.util.function.LongSupplier nowNanos) {
PrimaryRegistry registry = new PrimaryRegistry(null);
registry.recordDelegation(T, LEAD);
FakeHerdr leadHerdr = new FakeHerdr();
AgentControl leadAgents = new AgentControl(leadHerdr);
var scheduler = java.util.concurrent.Executors.newSingleThreadScheduledExecutor();
ReplyPushLoop pushLoop = new ReplyPushLoop(registry, leadAgents, inbox, scheduler, maxReminders, backoffMs);
MessageService service = new MessageService(agents, injector, rendezvous, inbox, pushLoop, null, nowNanos);
return new PushWiring(service, leadHerdr, scheduler);
}
private void awaitNudge(FakeHerdr leadHerdr) throws InterruptedException {
long deadline = System.currentTimeMillis() + 3000;
while (!leadHerdr.called("agent.prompt") && System.currentTimeMillis() < deadline) {
Thread.sleep(10);
}
assertTrue(leadHerdr.called("agent.prompt"), "expected a nudge in the lead's pane");
}
@Test
void anAsyncTicketThatFinishesNudgesTheLeadWithNoPriorPollCall() throws Exception {
try (var wiring = wireWithPushLoop(1, 50)) {
String ticket = wiring.service().sendAsync(T, "long task");
awaitWaiting();
injectDelivery();
assertTrue(rendezvous.resolve(T, "async result"));
awaitNudge(wiring.leadHerdr());
String nudge = wiring.leadHerdr().lastCall("agent.prompt").params().toString();
assertTrue(nudge.contains(ticket), "the nudge should name the ticket: " + nudge);
assertTrue(nudge.contains("bridge_poll(ticket="),
"the nudge should name the exact ticket-collecting call: " + nudge);
assertFalse(nudge.toUpperCase().contains("FAILED"),
"a successfully-replied ticket's nudge must not say it failed: " + nudge);
}
}
@Test
void aFailedAsyncTicketAlsoNudgesAndSaysSo() throws Exception {
try (var wiring = wireWithPushLoop(1, 50)) {
wiring.service().sendAsync(T, "long task");
awaitWaiting();
wiring.service().abandon(T, "session released"); // a terminal failure phase
awaitNudge(wiring.leadHerdr());
String nudge = wiring.leadHerdr().lastCall("agent.prompt").params().toString();
assertTrue(nudge.toUpperCase().contains("FAILED"), "a failed ticket's nudge must say so: " + nudge);
}
}
@Test
void anAlreadyCollectedTicketProducesNoNudge() throws Exception {
try (var wiring = wireWithPushLoop(1, 300)) { // wide backoff: poll before the first tick fires
String ticket = wiring.service().sendAsync(T, "long task");
awaitWaiting();
injectDelivery();
assertTrue(rendezvous.resolve(T, "async result"));
MessageService.TaskView view = awaitTicketPhaseOn(wiring.service(), ticket, MessageService.Phase.DONE);
assertEquals(MessageService.Phase.DONE, view.phase());
Thread.sleep(400); // let the scheduled tick run — it must find nothing pending
assertFalse(wiring.leadHerdr().called("agent.prompt"),
"a ticket the lead already polled must never be nudged");
}
}
@Test
void severalAsyncTicketsFinishingTogetherProduceOneCoalescedNudge() throws Exception {
try (var wiring = wireWithPushLoop(1, 300)) { // wide backoff: both tickets land before the tick fires
String first = wiring.service().sendAsync(T, "first task");
awaitWaiting();
injectDelivery();
assertTrue(rendezvous.resolve(T, "first done"));
// Settle without polling: poll() itself marks a ticket collected (that's the point of
// anAlreadyCollectedTicketProducesNoNudge above) — using it here to detect completion
// would collect the ticket before the coalescing this test checks ever gets a chance.
Thread.sleep(100);
String second = wiring.service().sendAsync(T, "second task");
awaitWaiting();
injectDelivery();
assertTrue(rendezvous.resolve(T, "second done"));
Thread.sleep(100);
awaitNudge(wiring.leadHerdr());
Thread.sleep(200); // settle — nothing more should arrive beyond the one coalesced nudge
long nudgeCount = wiring.leadHerdr().calls.stream()
.filter(c -> c.method().equals("agent.prompt")).count();
assertEquals(1, nudgeCount, "two tickets finishing together must produce ONE nudge, not two");
String nudge = wiring.leadHerdr().lastCall("agent.prompt").params().toString();
assertTrue(nudge.contains(first) && nudge.contains(second),
"the coalesced nudge should name both tickets: " + nudge);
}
}
@Test
void aFleetWithNoPushLoopConfiguredBehavesExactlyAsToday() throws Exception {
// `messages` (the shared field) uses the no-pushLoop constructor — poll() must not throw,
// and no nudge mechanism exists to fire regardless of how the ticket resolves.
String ticket = messages.sendAsync(T, "long task");
awaitWaiting();
injectDelivery();
assertTrue(rendezvous.resolve(T, "async result"));
MessageService.TaskView view = awaitTicketPhase(ticket, MessageService.Phase.DONE);
assertEquals(MessageService.Phase.DONE, view.phase());
assertEquals("async result", view.reply());
}
/**
* CB-588 follow-up: {@code tasks} is the sole authority on whether a ticket exists, and
* {@code pruneTerminalTickets} drops entries from it once {@link MessageService#TICKET_TTL_NANOS}
* elapses. Before this test, that prune never told {@code ReplyPushLoop} — its own
* {@code pendingTickets} entry for a pruned, never-collected ticket had no remover at all, so it
* rode along on every later nudge to the same lead, naming a ticket {@code bridge_poll} could no
* longer find. Uses the injectable clock (mirroring {@code SessionManager}'s {@code nowNanos} seam
* for its idle reaper) to cross the 10-minute TTL without a real wait.
*/
@Test
void aPrunedTicketIsReclaimedFromThePushLoopNotLeakedForever() throws Exception {
java.util.concurrent.atomic.AtomicLong clock = new java.util.concurrent.atomic.AtomicLong(1_000_000_000L);
// maxReminders=1 + a short backoff: the stale ticket gets its one legitimate reminder, then
// decideTickets hits the cap and STOPs — activeLeads drops the lead, but (before the fix)
// pendingTickets never drops the ticket. That is the exact "cap already STOPped" branch of
// the bug report, reached deterministically rather than by timing it against a live tick.
try (var wiring = wireWithPushLoop(1, 50, clock::get)) {
String stale = wiring.service().sendAsync(T, "first task");
awaitWaiting();
injectDelivery();
assertTrue(rendezvous.resolve(T, "stale result"));
// Let the reminder loop fire its one nudge and hit the cap (STOP removes it from
// activeLeads; pendingTickets is untouched either way — that asymmetry is the bug).
awaitNudge(wiring.leadHerdr());
Thread.sleep(300);
assertTrue(wiring.leadHerdr().lastCall("agent.prompt").params().toString().contains(stale),
"sanity: the stale ticket's own reminder must have fired first");
// Cross the TTL — a real clock would need 10 minutes; the injected one does it instantly.
clock.addAndGet(MessageService.TICKET_TTL_NANOS + TimeUnit.SECONDS.toNanos(1));
// A second, unrelated ticket to the same target/lead reaches sendAsync, which prunes.
String fresh = wiring.service().sendAsync(T, "second task");
awaitWaiting();
injectDelivery();
assertTrue(rendezvous.resolve(T, "fresh result"));
// The fresh ticket restarts the (now-dormant) reminder loop with its own nudge.
long before = wiring.leadHerdr().calls.stream().filter(c -> c.method().equals("agent.prompt")).count();
long deadline = System.currentTimeMillis() + 3000;
while (wiring.leadHerdr().calls.stream().filter(c -> c.method().equals("agent.prompt")).count() <= before
&& System.currentTimeMillis() < deadline) {
Thread.sleep(10);
}
String latestNudge = wiring.leadHerdr().lastCall("agent.prompt").params().toString();
assertTrue(latestNudge.contains(fresh), "the fresh ticket's nudge must still arrive: " + latestNudge);
assertFalse(latestNudge.contains(stale),
"a pruned ticket must never be named in a later nudge — it is gone and bridge_poll "
+ "on it would return nothing: " + latestNudge);
// And bridge_poll(ticket=stale) really does return nothing now — the nudge would have lied.
assertNull(wiring.service().poll(stale), "the pruned ticket must actually be gone, not just unmentioned");
}
}
private MessageService.TaskView awaitTicketPhaseOn(MessageService svc, String ticket,
MessageService.Phase phase) throws Exception {
long deadline = System.currentTimeMillis() + 3000;
MessageService.TaskView view;
do {
view = svc.poll(ticket);
if (view.phase() == phase) {
return view;
}
Thread.sleep(5);
} while (System.currentTimeMillis() < deadline);
assertEquals(phase, view.phase());
return view;
}
private void assertFailedTicket(String ticket, String reason) throws Exception {
MessageService.TaskView view = awaitTicketPhase(ticket, MessageService.Phase.FAILED);
assertEquals(reason, view.detail());
}
private MessageService.TaskView awaitTicketPhase(String ticket, MessageService.Phase phase) throws Exception {
long deadline = System.currentTimeMillis() + 3000;
MessageService.TaskView view;
do {
view = messages.poll(ticket);
if (view.phase() == phase) {
return view;
}
Thread.sleep(5);
} while (System.currentTimeMillis() < deadline);
assertEquals(phase, view.phase());
return view;
}
}
@@ -114,6 +114,17 @@ class RendezvousTest {
"the first resolution wins; the stored value is unchanged");
}
@Test
void resolveExhaustedTwiceIsANoOpTheSecondTime() {
CompletableFuture<Rendezvous.Resolution> waiter = rendezvous.open(W);
assertTrue(rendezvous.resolveExhausted(waiter, "first reason"), "the first classification resolves");
assertFalse(rendezvous.resolveExhausted(waiter, "second reason"),
"a second exhausted resolution on an already-resolved waiter returns false");
assertEquals(Rendezvous.Kind.BACKEND_EXHAUSTED, waiter.getNow(null).kind());
assertEquals("first reason", waiter.getNow(null).text(),
"the first resolution wins; the stored value is unchanged");
}
@Test
void closeAskRemovesTheTurn() {
Rendezvous.AskTicket t = rendezvous.openAsk(W);
@@ -15,6 +15,7 @@ import java.util.ArrayList;
import java.util.Collections;
import java.util.List;
import java.util.Map;
import java.util.Set;
import java.util.concurrent.CountDownLatch;
import java.util.concurrent.Executors;
import java.util.concurrent.ScheduledExecutorService;
@@ -200,6 +201,210 @@ class ReplyPushLoopTest {
assertTrue(nudge.contains("bridge_poll(target=term_worker)"));
}
// --- CB-588: async ticket terminal nudges — decideTickets() logic --------------------------
@Test
void decideTicketsWithNothingPendingIsStop() {
agents = agentWithStatus("idle");
assertEquals(ReplyPushLoop.Action.STOP, loop().decideTickets(PRIMARY, 0));
}
@Test
void decideTicketsAtCapIsStop() {
agents = agentWithStatus("idle");
var loop = loop(2, 100_000);
loop.onTicketTerminal("task-1", WORKER, false);
assertEquals(ReplyPushLoop.Action.STOP, loop.decideTickets(PRIMARY, 2));
}
@Test
void decideTicketsUnderCapWithInjectableLeadIsInject() {
agents = agentWithStatus("idle");
var loop = loop(5, 100_000);
loop.onTicketTerminal("task-1", WORKER, false);
assertEquals(ReplyPushLoop.Action.INJECT, loop.decideTickets(PRIMARY, 0));
}
@Test
void decideTicketsUnderCapWithBusyLeadIsWaitBusy() {
agents = agentWithStatus("working");
var loop = loop(5, 100_000);
loop.onTicketTerminal("task-1", WORKER, false);
assertEquals(ReplyPushLoop.Action.WAIT_BUSY, loop.decideTickets(PRIMARY, 0));
}
@Test
void decideTicketsWithoutAKnownLeadIsStop() {
agents = agentWithStatus("idle");
var loop = new ReplyPushLoop(new PrimaryRegistry(null), agents, inbox, scheduler, 5, 100_000);
loop.onTicketTerminal("task-1", WORKER, false); // no lead known -> never registered as pending
assertEquals(ReplyPushLoop.Action.STOP, loop.decideTickets(PRIMARY, 0));
}
// --- CB-588: onTicketTerminal integration ---------------------------------------------------
@Test
void onTicketTerminalCausesExactlyOneNudgeNamingTheTicketAndThePollCall() throws Exception {
var rec = recordingClient();
agents = new AgentControl(rec);
loop(1, 50).onTicketTerminal("task-1", WORKER, false);
assertTrue(rec.sendLatch.await(3, TimeUnit.SECONDS), "one ticket nudge should have been sent");
assertEquals(1, rec.sendCount());
String nudge = rec.sentParams().getFirst().getValue().toString();
assertTrue(nudge.contains("task-1"), "nudge should name the ticket");
assertTrue(nudge.contains("bridge_poll(ticket="), "nudge should name the exact ticket-poll call");
assertFalse(nudge.contains("bridge_poll(target="), "a ticket nudge must not tell the lead to run the target-poll call");
}
@Test
void aFailedTicketNudgeSaysItWasAFailure() throws Exception {
var rec = recordingClient();
agents = new AgentControl(rec);
loop(1, 50).onTicketTerminal("task-1", WORKER, true);
assertTrue(rec.sendLatch.await(3, TimeUnit.SECONDS));
String nudge = rec.sentParams().getFirst().getValue().toString();
assertTrue(nudge.toUpperCase().contains("FAILED"), "a failed ticket's nudge must say so: " + nudge);
}
@Test
void severalTicketsFinishingTogetherProduceOneCoalescedNudge() throws Exception {
var rec = recordingClient();
agents = new AgentControl(rec);
var loop = loop(1, 300); // backoff wide enough that both onTicketTerminal calls land first
loop.onTicketTerminal("task-1", WORKER, false);
loop.onTicketTerminal("task-2", WORKER, true); // arrives while the schedule is already active
assertTrue(rec.sendLatch.await(3, TimeUnit.SECONDS), "one coalesced nudge should have been sent");
assertEquals(1, rec.sendCount(), "two tickets finishing together must produce ONE nudge, not two");
String nudge = rec.sentParams().getFirst().getValue().toString();
assertTrue(nudge.contains("task-1") && nudge.contains("task-2"),
"the coalesced nudge should name both tickets: " + nudge);
assertTrue(nudge.contains("2 tickets"), "the coalesced nudge should name the count: " + nudge);
}
@Test
void onTicketTerminalIsIdempotentPerLeadWhileActive() throws Exception {
var rec = recordingClient();
agents = new AgentControl(rec);
rec.sendLatch = new CountDownLatch(1);
var loop = loop(1, 100);
loop.onTicketTerminal("task-1", WORKER, false);
loop.onTicketTerminal("task-1", WORKER, false); // duplicate — should not start a second schedule
assertTrue(rec.sendLatch.await(3, TimeUnit.SECONDS));
Thread.sleep(200);
assertEquals(1, rec.sendCount(), "a duplicate onTicketTerminal for the same lead must not double-nudge");
}
@Test
void aTicketAlreadyCollectedProducesNoNudge() throws Exception {
var rec = recordingClient();
agents = new AgentControl(rec);
var loop = loop(1, 100);
loop.onTicketTerminal("task-1", WORKER, false);
loop.ticketCollected("task-1"); // the lead polled before the first tick fired
Thread.sleep(300); // let the scheduled tick run
assertEquals(0, rec.sendCount(), "an already-collected ticket must never be nudged");
}
@Test
void ticketNudgesSendUpToCapThenStop() throws Exception {
int cap = 2;
var rec = recordingClient();
agents = new AgentControl(rec);
rec.sendLatch = new CountDownLatch(cap);
loop(cap, 50).onTicketTerminal("task-1", WORKER, false);
assertTrue(rec.sendLatch.await(5, TimeUnit.SECONDS), cap + " ticket nudges should have fired");
Thread.sleep(300);
assertEquals(cap, rec.sendCount(), "exactly " + cap + " ticket nudges (cap=" + cap + ")");
}
@Test
void isActiveReflectsALiveTicketScheduleForTheHeartbeatStandDown() {
var rec = recordingClient();
agents = new AgentControl(rec);
ReplyPushLoop loop = loop(1, 100_000); // long backoff so the tick cannot fire mid-test
loop.onTicketTerminal("task-1", WORKER, false);
assertTrue(loop.isActive(), "an active ticket-reminder schedule must also stand off the heartbeat");
loop.stop();
assertFalse(loop.isActive(), "stopping clears the active ticket schedule too");
}
@Test
void aTicketStillPendingWhenTheLoopStopsIsNotStranded() {
// Regression for the race a reviewer found in gitea PR #73: onTicketTerminal's
// activeLeads.putIfAbsent can see the lead's slot as still occupied a moment before
// decideTickets' STOP releases it, so the ticket coalesces onto a schedule that is about to
// die and nothing ever nudges about it. Forcing that exact thread interleaving is not
// reliable, so this drives stopOrRestartTicketLoop — the STOP path's own release-and-recheck —
// directly, arranging the state it must not lose a ticket in: a ticket pending for the lead
// that was NOT part of the pre-decision snapshot (pendingBefore=empty), standing in for one
// that races in during the decision-to-release window.
var rec = recordingClient();
agents = new AgentControl(rec);
ReplyPushLoop loop = loop(1, 100_000); // long backoff — no natural tick fires during this test
loop.onTicketTerminal("task-1", WORKER, false); // pendingTickets={task-1}; activeLeads={PRIMARY}
// Stand in for the scheduler thread reaching decideTickets==STOP for this lead — with nothing
// pending at decide time — while task-1 races in before the release below runs.
loop.stopOrRestartTicketLoop(PRIMARY, Set.of());
assertTrue(loop.isActive(), "a ticket that raced the loop's stop must reclaim the schedule "
+ "slot, not be stranded with no schedule left to ever nudge about it");
}
@Test
void aStaleUncollectedTicketAtCapDoesNotRestartTheLoop() {
// The other direction of the same fix: restarting on ANY non-empty pendingFor(lead) would be
// wrong. When STOP is reached because the reminder cap was hit, the same never-collected
// ticket is expected to still be there — that is the cap doing its job (acceptance criterion
// #7: nudges stay bounded). task-1 here was already accounted for at decide time (it is in
// pendingBefore), so it must not restart the loop just because it is still sitting there.
var rec = recordingClient();
agents = new AgentControl(rec);
ReplyPushLoop loop = loop(1, 100_000);
loop.onTicketTerminal("task-1", WORKER, false); // pendingTickets={task-1}; activeLeads={PRIMARY}
loop.stopOrRestartTicketLoop(PRIMARY, Set.of("task-1"));
assertFalse(loop.isActive(), "a stale ticket already accounted for at decide time must not "
+ "restart the loop — that would defeat the reminder cap");
}
@Test
void ticketNudgeFormatIsCorrect() {
String single = ReplyPushLoop.TICKET_NUDGE_FORMAT.formatted("task-1", "", "task-1");
assertTrue(single.contains("Ticket task-1"));
assertTrue(single.contains("bridge_poll(ticket=task-1)"));
String multi = ReplyPushLoop.TICKETS_NUDGE_FORMAT.formatted(2, "", "task-1, task-2");
assertTrue(multi.contains("2 tickets"));
assertTrue(multi.contains("bridge_poll(ticket=...)"));
}
// --- CB-307 nudge path is unchanged (regression) --------------------------------------------
@Test
void inboxNudgeStillUsesTheOriginalTargetPollCall() {
String nudge = ReplyPushLoop.NUDGE_FORMAT.formatted(WORKER, WORKER);
assertTrue(nudge.contains("bridge_poll(target=" + WORKER + ")"),
"CB-588 must not change the CB-307 inbox nudge's call shape");
}
// --- metrics (CB-512) ----------------------------------------------------------------------
@Test
@@ -233,6 +438,33 @@ class ReplyPushLoopTest {
assertEquals(0, metrics.count(BridgedMetrics.PUSH_NUDGES, "outcome", "delivered"));
}
@Test
void successfulTicketNudgeIncrementsDelivered() throws Exception {
var rec = recordingClient();
agents = new AgentControl(rec);
Metrics metrics = new Metrics();
loop(1, 50, metrics).onTicketTerminal("task-1", WORKER, false);
assertTrue(rec.sendLatch.await(3, TimeUnit.SECONDS), "one ticket nudge should have been sent");
Thread.sleep(200);
assertEquals(1, metrics.count(BridgedMetrics.PUSH_NUDGES, "outcome", "delivered"),
"a successfully sent ticket nudge must count as delivered, same metric as CB-307");
}
@Test
void ticketReminderCapIncrementsExhausted() {
agents = agentWithStatus("idle");
Metrics metrics = new Metrics();
var loop = loop(2, 100_000, metrics);
loop.onTicketTerminal("task-1", WORKER, false);
assertEquals(ReplyPushLoop.Action.STOP, loop.decideTickets(PRIMARY, 2));
assertEquals(1, metrics.count(BridgedMetrics.PUSH_NUDGES, "outcome", "exhausted"),
"hitting the ticket reminder cap must count as exhausted");
}
// --- helpers -------------------------------------------------------------------------------
private ReplyPushLoop loop() {
@@ -0,0 +1,20 @@
package dev.ltms.bridged.msg;
/**
* Explicit unbound tokens for tests that exercise delivery without an accepted send.
*
* <p>The waiter is {@code null} on purpose. "No accepted send" is an <em>absence</em>, and a helper
* that handed back a fresh {@code CompletableFuture} would invent one — which is how the first
* version of this class turned {@code captureBaselineSkipsTheReadWhenNoSendIsWaiting} red: the
* resolver saw a non-null waiter, decided a turn was in flight, and scraped a pane that no send was
* blocked on. An inert value must omit the fact, never fabricate it.
*/
public final class TestTurnTokens {
private TestTurnTokens() {
}
/** A token for a delivery that no send is waiting on: it authorises nothing. */
public static TurnToken inert(String target) {
return new TurnToken(target, null);
}
}
@@ -0,0 +1,57 @@
package dev.ltms.bridged.peer;
import org.junit.jupiter.api.Test;
import static org.junit.jupiter.api.Assertions.assertEquals;
import static org.junit.jupiter.api.Assertions.assertFalse;
import static org.junit.jupiter.api.Assertions.assertNotEquals;
import static org.junit.jupiter.api.Assertions.assertNull;
import static org.junit.jupiter.api.Assertions.assertTrue;
class CharterReceiptTest {
@Test
void recordsNoCharterConfiguredDistinctFromCharterDelivered() {
// No role charter configured — only the reply charter is composed. Source is "none", but
// text was still delivered, so absent() is false and the digest is present.
CharterReceipt viaReply = CharterReceipt.compose(MemberRole.DEV, "s", null, "reply charter");
// A role charter was configured AND delivered.
CharterReceipt delivered = CharterReceipt.compose(MemberRole.DEV, "s", "role charter",
"role charter\n\nreply charter");
// The two cases must not collapse: the no-role-charter case reports "none", the delivered
// case reports the config key, and their digests differ.
assertEquals(CharterReceipt.NO_SOURCE, viaReply.charterSource());
assertEquals("fleet.charters.dev", delivered.charterSource());
assertNotEquals(viaReply.charterSource(), delivered.charterSource());
assertNotEquals(viaReply.charterSha256(), delivered.charterSha256());
// Both actually delivered text — the distinction is the source and digest, not absence.
assertFalse(viaReply.absent());
assertFalse(delivered.absent());
}
@Test
void recordsExplicitAbsenceWhenNoCharterIsComposed() {
CharterReceipt none = CharterReceipt.compose(MemberRole.REVIEWER, "s", null, null);
assertTrue(none.absent());
assertNull(none.charterSha256());
assertEquals(0, none.charterBytes());
assertEquals(CharterReceipt.NO_SOURCE, none.charterSource(),
"no configured charter and nothing composed still reports a source, never a gap");
}
@Test
void digestIsStableForSameTextAndDiffersForDifferentText() {
assertEquals(CharterReceipt.digestOf("charter-aaa"), CharterReceipt.digestOf("charter-aaa"),
"the same text must always produce the same digest");
assertNotEquals(CharterReceipt.digestOf("charter-aaa"), CharterReceipt.digestOf("charter-bbb"),
"different text must produce a different digest");
assertNull(CharterReceipt.digestOf(""), "blank text carries no digest");
// The record's fingerprint matches the standalone digest for the same composed string.
CharterReceipt r = CharterReceipt.compose(MemberRole.DEV, "s", "role", "the composed text");
assertEquals(CharterReceipt.digestOf("the composed text"), r.charterSha256());
assertFalse(r.absent());
}
}
@@ -0,0 +1,119 @@
package dev.ltms.bridged.placement;
import org.junit.jupiter.api.Test;
import java.util.Map;
import java.util.OptionalLong;
import java.util.concurrent.TimeUnit;
import java.util.concurrent.atomic.AtomicLong;
import static org.junit.jupiter.api.Assertions.assertEquals;
import static org.junit.jupiter.api.Assertions.assertFalse;
import static org.junit.jupiter.api.Assertions.assertThrows;
import static org.junit.jupiter.api.Assertions.assertTrue;
/**
* CB-578 stage B: the credential-keyed quarantine tracker itself, isolated from placement/spawn
* wiring (that's {@code CompositePeerLauncherTest}). The clock is a plain {@link AtomicLong} of
* nanos so expiry is exercised without a real sleep.
*/
class BackendQuarantineTest {
@Test
void aFreshCredentialIsNotQuarantined() {
BackendQuarantine q = new BackendQuarantine(() -> 0L, TimeUnit.MINUTES.toNanos(30));
assertFalse(q.isQuarantined("shared-openai"));
assertEquals(OptionalLong.empty(), q.remainingSeconds("shared-openai"));
}
@Test
void quarantineBlocksTheCredentialForTheFullCooldown() {
BackendQuarantine q = new BackendQuarantine(() -> 0L, TimeUnit.MINUTES.toNanos(30));
q.quarantine("shared-openai");
assertTrue(q.isQuarantined("shared-openai"));
assertEquals(OptionalLong.of(1800L), q.remainingSeconds("shared-openai"));
}
@Test
void onlyTheQuarantinedCredentialIsAffected() {
BackendQuarantine q = new BackendQuarantine(() -> 0L, TimeUnit.MINUTES.toNanos(30));
q.quarantine("shared-openai");
assertFalse(q.isQuarantined("some-other-credential"),
"an unrelated credential must not be swept into the quarantine");
}
@Test
void expiresOnTheInjectedClock() {
AtomicLong now = new AtomicLong(0L);
BackendQuarantine q = new BackendQuarantine(now::get, TimeUnit.MINUTES.toNanos(30));
q.quarantine("shared-openai");
assertTrue(q.isQuarantined("shared-openai"));
now.set(TimeUnit.MINUTES.toNanos(29));
assertTrue(q.isQuarantined("shared-openai"), "still inside the cooldown");
now.set(TimeUnit.MINUTES.toNanos(31));
assertFalse(q.isQuarantined("shared-openai"), "the cooldown has elapsed on the injected clock");
assertEquals(OptionalLong.empty(), q.remainingSeconds("shared-openai"));
}
@Test
void aRepeatQuarantineCallRestartsTheCooldownAtFullLength() {
AtomicLong now = new AtomicLong(0L);
BackendQuarantine q = new BackendQuarantine(now::get, TimeUnit.MINUTES.toNanos(30));
q.quarantine("shared-openai");
now.set(TimeUnit.MINUTES.toNanos(20));
q.quarantine("shared-openai");
now.set(TimeUnit.MINUTES.toNanos(45)); // 25 min after the second call, 45 after the first
assertTrue(q.isQuarantined("shared-openai"),
"a fresh exhaustion resets the cooldown to full length, not the earlier shorter wait");
}
@Test
void activeRemainingSecondsListsOnlyStillQuarantinedCredentials() {
AtomicLong now = new AtomicLong(0L);
BackendQuarantine q = new BackendQuarantine(now::get, TimeUnit.MINUTES.toNanos(30));
q.quarantine("shared-openai");
q.quarantine("another-credential");
now.set(TimeUnit.MINUTES.toNanos(31));
q.quarantine("shared-openai"); // re-quarantined after the first one expired
Map<String, Long> active = q.activeRemainingSeconds();
assertEquals(Map.of("shared-openai", 1800L), active,
"the expired credential is dropped; the re-quarantined one is reported");
}
@Test
void noneReportsNothingQuarantinedWhenNeverToldTo() {
BackendQuarantine q = BackendQuarantine.none();
assertFalse(q.isQuarantined("anything"));
assertTrue(q.activeRemainingSeconds().isEmpty());
}
@Test
void noneIgnoresAQuarantineCallInsteadOfLockingTheCredentialForever() {
// .none() holds a clock frozen at 0, so if #quarantine recorded a deadline the credential
// would never expire — locked out for the life of the daemon. Two production
// CompositePeerLauncher constructors default to none(), so that failure would be silent and
// permanent. An inert stand-in must omit the fact, never invent one.
BackendQuarantine q = BackendQuarantine.none();
q.quarantine("shared-openai");
assertFalse(q.isQuarantined("shared-openai"), "none() must not quarantine anything");
assertTrue(q.remainingSeconds("shared-openai").isEmpty());
assertTrue(q.activeRemainingSeconds().isEmpty());
}
@Test
void aNonPositiveCooldownIsRejected() {
assertThrows(IllegalArgumentException.class, () -> new BackendQuarantine(() -> 0L, 0L));
assertThrows(IllegalArgumentException.class, () -> new BackendQuarantine(() -> 0L, -1L));
}
}
@@ -24,7 +24,13 @@ class PlacementPolicyTest {
private static PlacementContext ctx(List<PlacementCandidate> candidates,
Function<String, Integer> liveCount,
Set<String> unreachable) {
return new PlacementContext("b", candidates, liveCount, unreachable);
return ctx(candidates, liveCount, unreachable, Set.of());
}
private static PlacementContext ctx(List<PlacementCandidate> candidates,
Function<String, Integer> liveCount,
Set<String> unreachable, Set<String> quarantined) {
return new PlacementContext("b", candidates, liveCount, unreachable, quarantined);
}
private static PlacementContext ctx(List<PlacementCandidate> candidates,
@@ -46,17 +52,37 @@ class PlacementPolicyTest {
PlacementPolicy policy = PlacementPolicies.fixed();
PlacementContext ctx = new PlacementContext(null,
List.of(PlacementCandidate.profile("a"), PlacementCandidate.profile("b")),
noSessions(), Set.of());
noSessions(), Set.of(), Set.of());
assertEquals("a", policy.select(ctx).profile());
}
@Test
void fixedThrowsWhenNoProfilesAndNoDefault() {
PlacementPolicy policy = PlacementPolicies.fixed();
PlacementContext ctx = new PlacementContext(null, List.of(), noSessions(), Set.of());
PlacementContext ctx = new PlacementContext(null, List.of(), noSessions(), Set.of(), Set.of());
assertThrows(PlacementException.class, () -> policy.select(ctx));
}
@Test
void fixedSkipsQuarantinedDefault() {
PlacementPolicy policy = PlacementPolicies.fixed();
PlacementContext ctx = new PlacementContext("b",
List.of(PlacementCandidate.profile("a"), PlacementCandidate.profile("b")),
noSessions(), Set.of(), Set.of("b"));
assertEquals("a", policy.select(ctx).profile(),
"the default 'b' is quarantined, so fixed falls through to the first un-quarantined candidate");
}
@Test
void fixedThrowsWhenDefaultAndEveryCandidateQuarantined() {
PlacementPolicy policy = PlacementPolicies.fixed();
PlacementContext ctx = new PlacementContext("b",
List.of(PlacementCandidate.profile("a"), PlacementCandidate.profile("b")),
noSessions(), Set.of(), Set.of("a", "b"));
PlacementException e = assertThrows(PlacementException.class, () -> policy.select(ctx));
assertTrue(e.getMessage().contains("quarantined"), e.getMessage());
}
@Test
void roundRobinCyclesThroughAvailableProfiles() {
PlacementPolicy policy = PlacementPolicies.roundRobin();
@@ -82,6 +108,18 @@ class PlacementPolicyTest {
}
}
@Test
void roundRobinSkipsQuarantinedProfiles() {
PlacementPolicy policy = PlacementPolicies.roundRobin();
List<PlacementCandidate> candidates = List.of(
PlacementCandidate.profile("a"),
PlacementCandidate.profile("b"));
PlacementContext ctx = ctx(candidates, noSessions(), Set.of(), Set.of("a"));
for (int i = 0; i < 5; i++) {
assertEquals("b", policy.select(ctx).profile(), "a is quarantined, so every pick lands on b");
}
}
@Test
void roundRobinThrowsWhenAllAtMaxLoad() {
PlacementPolicy policy = PlacementPolicies.roundRobin();
@@ -150,6 +188,29 @@ class PlacementPolicyTest {
assertTrue(e.getMessage().contains("maxLoad"), e.getMessage());
}
@Test
void weightedSkipsQuarantinedProfile() {
PlacementPolicy policy = PlacementPolicies.weighted();
List<PlacementCandidate> candidates = List.of(
PlacementCandidate.profile("a", 1.0f, null),
PlacementCandidate.profile("b", 1.0f, null));
PlacementContext ctx = ctx(candidates, noSessions(), Set.of(), Set.of("a"));
for (int i = 0; i < 5; i++) {
assertEquals("b", policy.select(ctx).profile(), "a is quarantined, so every pick lands on b");
}
}
@Test
void weightedThrowsWhenAllQuarantined() {
PlacementPolicy policy = PlacementPolicies.weighted();
List<PlacementCandidate> candidates = List.of(
PlacementCandidate.profile("a"),
PlacementCandidate.profile("b"));
PlacementException e = assertThrows(PlacementException.class,
() -> policy.select(ctx(candidates, noSessions(), Set.of(), Set.of("a", "b"))));
assertTrue(e.getMessage().contains("quarantined"), e.getMessage());
}
@Test
void weightedThrowsWhenAllUnreachable() {
PlacementPolicy policy = PlacementPolicies.weighted();
@@ -176,8 +237,134 @@ class PlacementPolicyTest {
assertTrue(e.getMessage().contains("1 unreachable"), e.getMessage());
}
// --- CB-554: weight <= 0 excludes a candidate from automatic selection ------------------
@Test
void weightedSkipsWeightZeroProfile() {
PlacementPolicy policy = PlacementPolicies.weighted();
List<PlacementCandidate> candidates = List.of(
PlacementCandidate.profile("a", 0.0f, null),
PlacementCandidate.profile("b", 1.0f, null));
for (int i = 0; i < 5; i++) {
assertEquals("b", policy.select(ctx(candidates, noSessions())).profile(),
"a has weight 0, so every pick lands on b");
}
}
@Test
void weightedThrowsWhenAllWeightZero() {
PlacementPolicy policy = PlacementPolicies.weighted();
List<PlacementCandidate> candidates = List.of(
PlacementCandidate.profile("a", 0.0f, null),
PlacementCandidate.profile("b", 0.0f, null));
PlacementException e = assertThrows(PlacementException.class,
() -> policy.select(ctx(candidates, noSessions())));
assertTrue(e.getMessage().contains("weight 0"), e.getMessage());
}
@Test
void weightedTreatsNegativeWeightAsExcludedNotError() {
PlacementPolicy policy = PlacementPolicies.weighted();
List<PlacementCandidate> candidates = List.of(
PlacementCandidate.profile("a", -1.0f, null),
PlacementCandidate.profile("b", 1.0f, null));
for (int i = 0; i < 5; i++) {
assertEquals("b", policy.select(ctx(candidates, noSessions())).profile(),
"a has a negative weight, so it is excluded like weight 0 — never picked, never an error");
}
}
@Test
void weightZeroAndQuarantinedIsNotDoubleCounted() {
PlacementPolicy policy = PlacementPolicies.weighted();
List<PlacementCandidate> candidates = List.of(
PlacementCandidate.profile("a", 0.0f, null),
PlacementCandidate.profile("b", 1.0f, null),
PlacementCandidate.profile("c", 1.0f, null));
// a is BOTH weight-0 and quarantined; b and c are quarantined only. If a were counted
// into both buckets, "quarantined" would read 3 (or the arithmetic would go negative)
// instead of the true count of 2 candidates actually quarantined.
PlacementContext mixedCtx = ctx(candidates, noSessions(), Set.of(), Set.of("a", "b", "c"));
PlacementException e = assertThrows(PlacementException.class, () -> policy.select(mixedCtx));
assertTrue(e.getMessage().contains("1 weight-0"), e.getMessage());
assertTrue(e.getMessage().contains("2 quarantined"), e.getMessage());
assertTrue(e.getMessage().contains("0 remaining"), e.getMessage());
}
@Test
void roundRobinSkipsWeightZeroProfiles() {
PlacementPolicy policy = PlacementPolicies.roundRobin();
List<PlacementCandidate> candidates = List.of(
PlacementCandidate.profile("a", 0.0f, null),
PlacementCandidate.profile("b", 1.0f, null));
for (int i = 0; i < 5; i++) {
assertEquals("b", policy.select(ctx(candidates, noSessions())).profile(),
"a has weight 0, so every pick lands on b");
}
}
@Test
void roundRobinThrowsWhenAllWeightZero() {
PlacementPolicy policy = PlacementPolicies.roundRobin();
List<PlacementCandidate> candidates = List.of(
PlacementCandidate.profile("a", 0.0f, null),
PlacementCandidate.profile("b", 0.0f, null));
PlacementException e = assertThrows(PlacementException.class,
() -> policy.select(ctx(candidates, noSessions())));
assertTrue(e.getMessage().contains("weight 0"), e.getMessage());
}
@Test
void fixedSkipsWeightZeroDefault() {
PlacementPolicy policy = PlacementPolicies.fixed();
PlacementContext ctx = new PlacementContext("b",
List.of(PlacementCandidate.profile("a", 1.0f, null),
PlacementCandidate.profile("b", 0.0f, null)),
noSessions(), Set.of(), Set.of());
assertEquals("a", policy.select(ctx).profile(),
"the default 'b' has weight 0, so fixed falls through to the first available candidate");
}
@Test
void fixedThrowsWhenDefaultAndEveryCandidateWeightZero() {
PlacementPolicy policy = PlacementPolicies.fixed();
PlacementContext ctx = new PlacementContext("b",
List.of(PlacementCandidate.profile("a", 0.0f, null),
PlacementCandidate.profile("b", 0.0f, null)),
noSessions(), Set.of(), Set.of());
PlacementException e = assertThrows(PlacementException.class, () -> policy.select(ctx));
assertTrue(e.getMessage().contains("weight 0"), e.getMessage());
}
@Test
void unknownPolicyNameThrows() {
assertThrows(IllegalArgumentException.class, () -> PlacementPolicies.fromName("random"));
}
// --- CB-585: maxLoad 0 caps a profile at zero live members, even with nothing live yet -----
@Test
void weightedSkipsMaxLoadZeroProfileEvenWithZeroLiveWorkers() {
PlacementPolicy policy = PlacementPolicies.weighted();
List<PlacementCandidate> candidates = List.of(
PlacementCandidate.profile("a", 1.0f, 0),
PlacementCandidate.profile("b", 1.0f, null));
Function<String, Integer> liveCount = _ -> 0;
for (int i = 0; i < 5; i++) {
assertEquals("b", policy.select(ctx(candidates, liveCount)).profile(),
"a has maxLoad 0, so it is already at its cap with nobody live — every pick lands on b");
}
}
@Test
void weightedThrowsWhenAllMaxLoadZero() {
PlacementPolicy policy = PlacementPolicies.weighted();
List<PlacementCandidate> candidates = List.of(
PlacementCandidate.profile("a", 1.0f, 0),
PlacementCandidate.profile("b", 1.0f, 0));
Function<String, Integer> liveCount = _ -> 0;
PlacementException e = assertThrows(PlacementException.class,
() -> policy.select(ctx(candidates, liveCount)));
assertTrue(e.getMessage().contains("maxLoad"), e.getMessage());
}
}
@@ -2,9 +2,11 @@ package dev.ltms.bridged.session;
import java.util.Collections;
import java.util.List;
import java.util.Optional;
import java.util.Set;
import java.util.concurrent.ConcurrentHashMap;
import java.util.concurrent.CopyOnWriteArrayList;
import java.util.concurrent.atomic.AtomicLong;
/** Recording fake {@link Worktrees} for CB-301-ext acceptance tests (no live git). */
public final class FakeWorktrees implements Worktrees {
@@ -22,15 +24,28 @@ public final class FakeWorktrees implements Worktrees {
public record RepoRootCall(String cwd) {
}
public record SnapshotCall(String worktreePath, String branch, String message) {
}
public record PruneCall(String repoRoot, long minAgeMillis) {
}
private final List<AddCall> addCalls = new CopyOnWriteArrayList<>();
private final List<RemoveCall> removeCalls = new CopyOnWriteArrayList<>();
private final List<OverlayCall> overlayCalls = new CopyOnWriteArrayList<>();
private final List<RepoRootCall> repoRootCalls = new CopyOnWriteArrayList<>();
private final List<SnapshotCall> snapshotCalls = new CopyOnWriteArrayList<>();
private final List<PruneCall> pruneCalls = new CopyOnWriteArrayList<>();
private final Set<String> existingPaths = ConcurrentHashMap.newKeySet();
private final Set<String> trackedPaths = ConcurrentHashMap.newKeySet();
private final AtomicLong snapshotSeq = new AtomicLong();
private volatile RuntimeException addFailure;
private volatile RuntimeException snapshotFailure;
private volatile boolean dirty = false;
private volatile String repoRoot = "/repo";
private volatile String prefix = "/worktrees";
private volatile WipRefStats wipRefs = new WipRefStats(0, 0L);
private volatile int pruneResult = 0;
public FakeWorktrees withRepoRoot(String root) {
this.repoRoot = root;
@@ -61,6 +76,30 @@ public final class FakeWorktrees implements Worktrees {
return this;
}
/** Mark the worktree dirty so {@link #hasUncommitted} reports true (simulates uncommitted work). */
public FakeWorktrees withDirty(boolean dirty) {
this.dirty = dirty;
return this;
}
/** Make subsequent {@link #snapshot} calls throw (simulates a failing git snapshot). */
public FakeWorktrees failSnapshot(String message) {
this.snapshotFailure = new WorktreeException(message);
return this;
}
/** Configure the value returned by {@link #wipRefs}. */
public FakeWorktrees withWipRefs(WipRefStats stats) {
this.wipRefs = stats;
return this;
}
/** Configure the value returned by {@link #pruneWipRefs}. */
public FakeWorktrees withPruneResult(int deleted) {
this.pruneResult = deleted;
return this;
}
@Override
public String add(String repoRoot, String branch, String baseRef) {
addCalls.add(new AddCall(repoRoot, branch, baseRef));
@@ -77,6 +116,11 @@ public final class FakeWorktrees implements Worktrees {
removeCalls.add(new RemoveCall(repoRoot, worktreePath));
}
@Override
public boolean hasUncommitted(String worktreePath) {
return dirty;
}
@Override
public void overlayParity(String repoRoot, String worktreePath, List<String> overlay) {
List<String> copied = new java.util.ArrayList<>();
@@ -100,6 +144,26 @@ public final class FakeWorktrees implements Worktrees {
return repoRoot;
}
@Override
public Optional<String> snapshot(String worktreePath, String branch, String message) {
snapshotCalls.add(new SnapshotCall(worktreePath, branch, message));
if (snapshotFailure != null) {
throw snapshotFailure;
}
return Optional.of("wip" + snapshotSeq.incrementAndGet());
}
@Override
public WipRefStats wipRefs(String repoRoot) {
return wipRefs;
}
@Override
public int pruneWipRefs(String repoRoot, long minAgeMillis) {
pruneCalls.add(new PruneCall(repoRoot, minAgeMillis));
return pruneResult;
}
public List<AddCall> addCalls() {
return List.copyOf(addCalls);
}
@@ -127,4 +191,16 @@ public final class FakeWorktrees implements Worktrees {
public OverlayCall lastOverlay() {
return overlayCalls.isEmpty() ? null : overlayCalls.getLast();
}
public List<SnapshotCall> snapshotCalls() {
return List.copyOf(snapshotCalls);
}
public List<PruneCall> pruneCalls() {
return List.copyOf(pruneCalls);
}
public SnapshotCall lastSnapshot() {
return snapshotCalls.isEmpty() ? null : snapshotCalls.getLast();
}
}
@@ -3,9 +3,13 @@ package dev.ltms.bridged.session;
import org.junit.jupiter.api.Test;
import org.junit.jupiter.api.io.TempDir;
import java.nio.charset.StandardCharsets;
import java.nio.file.Files;
import java.nio.file.Path;
import java.util.HashSet;
import java.util.List;
import java.util.Optional;
import java.util.Set;
import java.util.concurrent.TimeUnit;
import static org.junit.jupiter.api.Assertions.*;
@@ -68,6 +72,120 @@ class GitWorktreesTest {
return out;
}
/** Every pending change in {@code cwd} — the whole-tree porcelain status, unlike {@link #status}. */
private static String fullStatus(Path cwd) throws Exception {
Process p = new ProcessBuilder("git", "status", "--porcelain")
.directory(cwd.toFile()).redirectErrorStream(true).start();
String out = new String(p.getInputStream().readAllBytes());
assertTrue(p.waitFor(30, TimeUnit.SECONDS), "git status timed out");
return out;
}
private static String revParse(Path cwd, String ref) throws Exception {
Process p = new ProcessBuilder("git", "-C", cwd.toString(), "rev-parse", ref)
.redirectErrorStream(true).start();
String out = new String(p.getInputStream().readAllBytes()).trim();
assertTrue(p.waitFor(30, TimeUnit.SECONDS), "git rev-parse timed out");
assertEquals(0, p.exitValue(), "git rev-parse " + ref + " failed:\n" + out);
return out;
}
/** The recursive file list of a commit's tree — used to check what a snapshot actually committed. */
private static String lsTree(Path cwd, String ref) throws Exception {
Process p = new ProcessBuilder("git", "-C", cwd.toString(), "ls-tree", "-r", "--name-only", ref)
.redirectErrorStream(true).start();
String out = new String(p.getInputStream().readAllBytes());
assertTrue(p.waitFor(30, TimeUnit.SECONDS), "git ls-tree timed out");
assertEquals(0, p.exitValue(), "git ls-tree " + ref + " failed:\n" + out);
return out;
}
/** The set of paths in {@code git diff --name-only from..to} — used to check exactly what a
* snapshot's tree changed relative to its parent, the same shape {@code git status --porcelain}
* reports for the worktree it was taken from. */
private static Set<String> diffNameOnly(Path cwd, String from, String to) throws Exception {
Process p = new ProcessBuilder("git", "-C", cwd.toString(), "diff", "--name-only", from, to)
.redirectErrorStream(true).start();
String out = new String(p.getInputStream().readAllBytes());
assertTrue(p.waitFor(30, TimeUnit.SECONDS), "git diff timed out");
assertEquals(0, p.exitValue(), "git diff " + from + ".." + to + " failed:\n" + out);
Set<String> paths = new HashSet<>();
for (String line : out.split("\\R")) {
if (!line.isBlank()) {
paths.add(line.trim());
}
}
return paths;
}
/** The set of paths a {@code git status --porcelain} listing names, stripping the two-char status
* code prefix each line carries. */
private static Set<String> porcelainPaths(String porcelain) {
Set<String> paths = new HashSet<>();
for (String line : porcelain.split("\\R")) {
if (!line.isBlank()) {
paths.add(line.substring(3).trim());
}
}
return paths;
}
private static String forEachRef(Path cwd, String pattern) throws Exception {
Process p = new ProcessBuilder("git", "-C", cwd.toString(), "for-each-ref", pattern)
.redirectErrorStream(true).start();
String out = new String(p.getInputStream().readAllBytes());
assertTrue(p.waitFor(30, TimeUnit.SECONDS), "git for-each-ref timed out");
assertEquals(0, p.exitValue(), "git for-each-ref " + pattern + " failed:\n" + out);
return out;
}
/** Write {@code content} as a blob into the object database; returns its sha. */
private static String blobOf(Path cwd, String content) throws Exception {
Process p = new ProcessBuilder("git", "-C", cwd.toString(), "hash-object", "-w", "--stdin")
.redirectErrorStream(true).start();
p.getOutputStream().write(content.getBytes(StandardCharsets.UTF_8));
p.getOutputStream().close();
String out = new String(p.getInputStream().readAllBytes()).trim();
assertTrue(p.waitFor(30, TimeUnit.SECONDS), "git hash-object timed out");
assertEquals(0, p.exitValue(), "git hash-object failed:\n" + out);
return out;
}
/** Build a single-file tree object from {@code blob}; returns the tree's sha. */
private static String treeOf(Path cwd, String path, String blob) throws Exception {
Process p = new ProcessBuilder("git", "-C", cwd.toString(), "mktree")
.redirectErrorStream(true).start();
p.getOutputStream().write(("100644 blob " + blob + "\t" + path + "\n").getBytes(StandardCharsets.UTF_8));
p.getOutputStream().close();
String out = new String(p.getInputStream().readAllBytes()).trim();
assertTrue(p.waitFor(30, TimeUnit.SECONDS), "git mktree timed out");
assertEquals(0, p.exitValue(), "git mktree failed:\n" + out);
return out;
}
/** {@code git commit-tree} rooted at {@code tree} with a chosen committer date; returns the sha. */
private static String commitTree(Path cwd, String tree, String parent, String committerDate,
String message) throws Exception {
ProcessBuilder pb = new ProcessBuilder("git", "-C", cwd.toString(), "commit-tree",
tree, "-p", parent, "-m", message);
pb.environment().put("GIT_COMMITTER_DATE", committerDate);
Process p = pb.redirectErrorStream(true).start();
String out = new String(p.getInputStream().readAllBytes()).trim();
assertTrue(p.waitFor(30, TimeUnit.SECONDS), "git commit-tree timed out");
assertEquals(0, p.exitValue(), "git commit-tree failed:\n" + out);
return out;
}
/** {@code git update-ref <ref> <sha>} — create the snapshot ref directly. */
private static void updateRef(Path cwd, String ref, String sha) throws Exception {
git(cwd, "update-ref", ref, sha);
}
/** True when {@code ref} exists in the repo (for-each-ref on a missing ref is empty, not an error). */
private static boolean refExists(Path cwd, String ref) throws Exception {
return !forEachRef(cwd, ref).trim().isEmpty();
}
/**
* The heart of CB-525: a provisioned worktree must not inherit the primary's MCP servers. Without
* the isolation step the checked-out {@code .mcp.json} carries them in, and a worker navigating
@@ -175,6 +293,46 @@ class GitWorktreesTest {
assertTrue(Files.exists(Path.of(wt).resolve(".mcp.json")), ".mcp.json stub was dropped");
}
/**
* CB-576. {@code hasUncommitted} must treat a freshly-provisioned worktree as clean, but a
* worktree holding a brand-new, never-added file as dirty. The untracked-file-only shape is
* exactly the work lost in the incident — a worker's draft that compiled but was never
* committed because it stopped to ask its lead a question.
*/
@Test
void anUntrackedOnlyWorktreeCountsAsDirty(@TempDir Path tmp) throws Exception {
Path repo = initRepo(tmp.resolve("repo"));
GitWorktrees gitWorktrees = new GitWorktrees(tmp.resolve("wts").toString());
String wt = gitWorktrees.add(repo.toString(), "cb-576-u", "HEAD");
assertFalse(gitWorktrees.hasUncommitted(wt),
"a freshly provisioned worktree must read as clean");
Files.writeString(Path.of(wt).resolve("brand-new.txt"), "draft that was never added\n");
assertTrue(gitWorktrees.hasUncommitted(wt),
"an untracked-only file must count as dirty");
Files.writeString(Path.of(wt).resolve("README.md"), "edited tracked file\n");
assertTrue(gitWorktrees.hasUncommitted(wt),
"a tracked modification must also count as dirty");
}
/**
* CB-576 review. {@code hasUncommitted} must tolerate a missing worktree exactly like
* {@code remove}: an already-gone directory holds no work to lose, and throwing here would
* break teardown — SessionManager.release() calls it before stopping the pane, so an
* exception would orphan a live pane and skip the release notification (CB-516).
*/
@Test
void hasUncommittedOnAMissingWorktreeReturnsFalseWithoutThrowing(@TempDir Path tmp) {
GitWorktrees gitWorktrees = new GitWorktrees(tmp.resolve("wts").toString());
String gone = tmp.resolve("wts").resolve("does-not-exist").toString();
assertFalse(gitWorktrees.hasUncommitted(gone),
"a missing worktree is reported clean, not an error");
}
/** All three protected configs are covered: each one present in a worktree is neutralized and hidden. */
@Test
void allThreeConfigsAreNeutralizedWhenPresent(@TempDir Path tmp) throws Exception {
@@ -205,4 +363,263 @@ class GitWorktreesTest {
assertEquals("", status(Path.of(wt), "opencode.json"), "opencode.json still shows as modified");
assertEquals("", status(Path.of(wt), ".autoenv"), ".autoenv still shows as modified");
}
/**
* CB-578 stage C, acceptance criterion 1. A dirty worktree — a tracked edit plus a brand-new
* untracked file, exactly the shape lost in CB-576 — must land in {@code refs/wip/<branch>}'s
* tree, and that ref must live outside {@code refs/heads} so it never shows up in
* {@code git branch} or gets swept by a branch cleanup.
*/
@Test
void dirtySnapshotCreatesARefWhoseTreeContainsUntrackedAndTrackedChanges(@TempDir Path tmp) throws Exception {
Path repo = initRepo(tmp.resolve("repo"));
GitWorktrees gitWorktrees = new GitWorktrees(tmp.resolve("wts").toString());
String branch = "cb-578-c-a";
String wt = gitWorktrees.add(repo.toString(), branch, "HEAD");
Files.writeString(Path.of(wt).resolve("untracked.txt"), "draft that was never added\n");
Files.writeString(Path.of(wt).resolve("README.md"), "edited tracked file\n");
Optional<String> ref = gitWorktrees.snapshot(wt, branch, "test snapshot");
assertTrue(ref.isPresent(), "a dirty worktree snapshot returns a commit sha");
String tree = lsTree(repo, "refs/wip/" + branch);
assertTrue(tree.contains("untracked.txt"), "snapshot tree must include the untracked file:\n" + tree);
assertTrue(tree.contains("README.md"), "snapshot tree must include the tracked edit:\n" + tree);
String heads = forEachRef(repo, "refs/heads");
assertFalse(heads.contains("refs/wip/"), "the snapshot ref must not live under refs/heads:\n" + heads);
}
/**
* CB-578 stage C, acceptance criterion 2. Building the commit through a temporary
* {@code GIT_INDEX_FILE} must leave the worker's own index, working tree, and HEAD exactly as
* they were — the worker may still be mid-write, and staging into the real index would corrupt
* that.
*/
@Test
void snapshotDoesNotTouchTheWorkersOwnIndexWorkingTreeOrHead(@TempDir Path tmp) throws Exception {
Path repo = initRepo(tmp.resolve("repo"));
GitWorktrees gitWorktrees = new GitWorktrees(tmp.resolve("wts").toString());
String branch = "cb-578-c-b";
String wt = gitWorktrees.add(repo.toString(), branch, "HEAD");
Files.writeString(Path.of(wt).resolve("untracked.txt"), "draft that was never added\n");
Files.writeString(Path.of(wt).resolve("README.md"), "edited tracked file\n");
String headBefore = revParse(Path.of(wt), "HEAD");
String statusBefore = fullStatus(Path.of(wt));
gitWorktrees.snapshot(wt, branch, "test snapshot");
assertEquals(headBefore, revParse(Path.of(wt), "HEAD"), "snapshot must not move HEAD");
assertEquals(statusBefore, fullStatus(Path.of(wt)),
"snapshot must not change the worker's own index or working tree status");
}
/**
* CB-578 stage C, acceptance criterion 3. {@code add -A} (never {@code -f}) respects
* {@code .gitignore}, and that is the only thing keeping a gitignored file (secrets, local
* config) out of a snapshot commit — an untracked file that is NOT ignored must still be
* included, so this isn't just "untracked files are dropped".
*/
@Test
void gitignoredFileIsExcludedButOtherUntrackedFilesAreNot(@TempDir Path tmp) throws Exception {
Path repo = tmp.resolve("repo");
Files.createDirectories(repo);
git(repo, "init", "-q", "-b", "main");
git(repo, "config", "user.email", "test@example.invalid");
git(repo, "config", "user.name", "Test");
Files.writeString(repo.resolve(".gitignore"), ".env\n");
Files.writeString(repo.resolve("README.md"), "seed\n");
git(repo, "add", ".gitignore", "README.md");
git(repo, "commit", "-q", "-m", "seed");
GitWorktrees gitWorktrees = new GitWorktrees(tmp.resolve("wts").toString());
String branch = "cb-578-c-c";
String wt = gitWorktrees.add(repo.toString(), branch, "HEAD");
Files.writeString(Path.of(wt).resolve(".env"), "SECRET=shh\n");
Files.writeString(Path.of(wt).resolve("untracked.txt"), "draft that was never added\n");
Optional<String> ref = gitWorktrees.snapshot(wt, branch, "test snapshot");
assertTrue(ref.isPresent());
String tree = lsTree(repo, "refs/wip/" + branch);
assertFalse(tree.contains(".env"), "a gitignored file must never enter the snapshot:\n" + tree);
assertTrue(tree.contains("untracked.txt"),
"a non-ignored untracked file must still be included:\n" + tree);
}
/** CB-578 stage C. A clean worktree still produces a valid, if tree-identical, commit — the caller
* (SessionManager) is the one that decides not to call this on a clean worktree. */
@Test
void snapshotOfAMissingWorktreeReturnsEmptyWithoutThrowing(@TempDir Path tmp) {
GitWorktrees gitWorktrees = new GitWorktrees(tmp.resolve("wts").toString());
String gone = tmp.resolve("wts").resolve("does-not-exist").toString();
assertTrue(gitWorktrees.snapshot(gone, "some-branch", "msg").isEmpty(),
"a missing worktree has nothing to snapshot, and must not throw");
}
/**
* CB-587, acceptance criteria 1 and 2. A file with a REAL {@code --skip-worktree} bit set in the
* worktree's own index must never enter the snapshot, even though its on-disk content has locally
* diverged from what is committed — that is exactly the divergence {@code --skip-worktree} exists
* to hide from {@code git status}, and a snapshot built from a fresh empty temp index (the bug)
* stages that local content anyway because the fresh index carries none of the real index's flags.
* The snapshot's diff against its parent must list exactly what {@code git status --porcelain}
* reports for the worktree — no more, no less.
*/
@Test
void dirtySnapshotHonoursARealSkipWorktreeBit(@TempDir Path tmp) throws Exception {
Path repo = tmp.resolve("repo");
Files.createDirectories(repo);
git(repo, "init", "-q", "-b", "main");
git(repo, "config", "user.email", "test@example.invalid");
git(repo, "config", "user.name", "Test");
Files.writeString(repo.resolve("protected.cfg"), "committed-value\n");
Files.writeString(repo.resolve("README.md"), "seed\n");
git(repo, "add", "protected.cfg", "README.md");
git(repo, "commit", "-q", "-m", "seed");
GitWorktrees gitWorktrees = new GitWorktrees(tmp.resolve("wts").toString());
String branch = "cb-587-a";
String wt = gitWorktrees.add(repo.toString(), branch, "HEAD");
git(Path.of(wt), "update-index", "--skip-worktree", "protected.cfg");
Files.writeString(Path.of(wt).resolve("protected.cfg"), "locally-diverged-never-commit\n");
Files.writeString(Path.of(wt).resolve("untracked.txt"), "draft that was never added\n");
Files.writeString(Path.of(wt).resolve("README.md"), "edited tracked file\n");
String porcelain = fullStatus(Path.of(wt));
assertFalse(porcelain.contains("protected.cfg"),
"test setup invalid — protected.cfg must not show in git status once skip-worktree is set:\n"
+ porcelain);
Optional<String> ref = gitWorktrees.snapshot(wt, branch, "test snapshot");
assertTrue(ref.isPresent(), "a dirty worktree snapshot returns a commit sha");
Set<String> diffPaths = diffNameOnly(repo, "HEAD", "refs/wip/" + branch);
assertFalse(diffPaths.contains("protected.cfg"),
"a --skip-worktree file's local drift leaked into the snapshot:\n" + diffPaths);
assertEquals(porcelainPaths(porcelain), diffPaths,
"snapshot diff must list exactly what git status --porcelain reports, no more, no less");
}
/**
* CB-587, acceptance criterion 6. If the worktree's real index cannot be resolved/read, snapshot
* must fail loudly (throw) rather than silently falling back to an empty temp index and producing
* a wrong snapshot. SessionManager's caller already catches and WARNs on any exception here — this
* only needs to confirm the failure is not swallowed inside snapshot() itself.
*/
@Test
void snapshotThrowsWhenTheWorktreesRealIndexCannotBeResolved(@TempDir Path tmp) throws Exception {
Path repo = initRepo(tmp.resolve("repo"));
GitWorktrees gitWorktrees = new GitWorktrees(tmp.resolve("wts").toString());
String branch = "cb-587-b";
String wt = gitWorktrees.add(repo.toString(), branch, "HEAD");
// Break git's ability to resolve the worktree's real index by removing the linked worktree's
// `.git` file (which normally points at the main repo's worktrees/<name>/ directory).
Files.delete(Path.of(wt).resolve(".git"));
assertThrows(WorktreeException.class, () -> gitWorktrees.snapshot(wt, branch, "test snapshot"),
"an unresolvable real index must fail loudly, not silently snapshot from an empty index");
}
/**
* CB-586, criterion 2. A snapshot whose content is NOT reachable from {@code main} is the last
* copy of a worker's work, and must never be deleted automatically — even when it is old and
* even when the caller passes a zero age floor. Uses the real snapshot path on a dirty worktree,
* so the unreachable tree is exactly the shape CB-576/CB-578 stage C exist to protect.
*/
@Test
void anUnreachableSnapshotIsNeverPruned(@TempDir Path tmp) throws Exception {
Path repo = initRepo(tmp.resolve("repo"));
GitWorktrees gitWorktrees = new GitWorktrees(tmp.resolve("wts").toString());
String branch = "cb-586-unreachable";
String wt = gitWorktrees.add(repo.toString(), branch, "HEAD");
Files.writeString(Path.of(wt).resolve("worker-draft.txt"), "work that exists nowhere else\n");
Optional<String> ref = gitWorktrees.snapshot(wt, branch, "snapshot with unreachable content");
assertTrue(ref.isPresent());
// Age floor 0 makes age a non-issue: only reachability can save it — and it must.
assertEquals(0, gitWorktrees.pruneWipRefs(repo.toString(), 0),
"the unreachable snapshot is the last copy and must not be pruned");
assertTrue(refExists(repo, "refs/wip/" + branch),
"an unreachable snapshot must survive the sweep");
}
/**
* CB-586, criterion 1 (the reachable half). A snapshot whose tree content IS already reachable
* from {@code main} and which is older than the age floor is pure duplication — the work is
* recovered — so it must be pruned.
*/
@Test
void aReachableSnapshotOlderThanTheFloorIsPruned(@TempDir Path tmp) throws Exception {
Path repo = initRepo(tmp.resolve("repo"));
GitWorktrees gitWorktrees = new GitWorktrees(tmp.resolve("wts").toString());
// A snapshot whose tree is exactly main's current tree: fully reachable from main.
String mainTree = revParse(repo, "main^{tree}");
String old = commitTree(repo, mainTree, revParse(repo, "HEAD"), "2020-01-01T00:00:00", "snapshot");
updateRef(repo, "refs/wip/recovered", old);
assertEquals(1, gitWorktrees.pruneWipRefs(repo.toString(), TimeUnit.HOURS.toMillis(24)),
"an old, main-reachable snapshot must be pruned");
assertFalse(refExists(repo, "refs/wip/recovered"),
"the reachable snapshot's ref must be gone after the sweep");
}
/**
* CB-586, the age floor. A snapshot whose content IS reachable from {@code main} but which is
* younger than the age floor must not be swept — a lead may still be looking at it.
*/
@Test
void aReachableButRecentSnapshotIsNotPruned(@TempDir Path tmp) throws Exception {
Path repo = initRepo(tmp.resolve("repo"));
GitWorktrees gitWorktrees = new GitWorktrees(tmp.resolve("wts").toString());
// Reachable from main, but committed "now" — a fresh snapshot. The 24h floor must protect it.
String mainTree = revParse(repo, "main^{tree}");
String fresh = commitTree(repo, mainTree, revParse(repo, "HEAD"),
"2038-01-01T00:00:00", "snapshot just taken");
updateRef(repo, "refs/wip/fresh", fresh);
assertEquals(0, gitWorktrees.pruneWipRefs(repo.toString(), TimeUnit.HOURS.toMillis(24)),
"a recent snapshot must be kept even when reachable");
assertTrue(refExists(repo, "refs/wip/fresh"),
"the recent reachable snapshot must survive the sweep");
}
/**
* CB-586, criterion 5. A fleet that has never snapshotted anything has no {@code refs/wip/*},
* so a sweep is a no-op and the census reports none — identical to before CB-586 existed.
*/
@Test
void aFleetWithNoSnapshotsPrunesNothingAndReportsNothing(@TempDir Path tmp) throws Exception {
Path repo = initRepo(tmp.resolve("repo"));
GitWorktrees gitWorktrees = new GitWorktrees(tmp.resolve("wts").toString());
assertEquals(0, gitWorktrees.pruneWipRefs(repo.toString(), 0),
"no snapshot refs means nothing to prune");
Worktrees.WipRefStats stats = gitWorktrees.wipRefs(repo.toString());
assertEquals(0, stats.count(), "a never-snapshotted fleet has zero refs/wip refs");
assertEquals(0L, stats.costBytes(), "a never-snapshotted fleet costs zero bytes");
}
/** CB-586, criterion 4: the census reports how many refs exist and roughly what they cost. */
@Test
void wipRefsReportsCountAndCost(@TempDir Path tmp) throws Exception {
Path repo = initRepo(tmp.resolve("repo"));
GitWorktrees gitWorktrees = new GitWorktrees(tmp.resolve("wts").toString());
String blob = blobOf(repo, "a recoverable snapshot's worth of content");
String tree = treeOf(repo, "snapshot.txt", blob);
updateRef(repo, "refs/wip/one", commitTree(repo, tree, revParse(repo, "HEAD"),
"2020-01-01T00:00:00", "snapshot"));
updateRef(repo, "refs/wip/two", commitTree(repo, tree, revParse(repo, "HEAD"),
"2020-01-02T00:00:00", "snapshot"));
Worktrees.WipRefStats stats = gitWorktrees.wipRefs(repo.toString());
assertEquals(2, stats.count(), "two snapshot refs are reported");
assertTrue(stats.costBytes() > 0, "the cost of the snapshots is a positive byte count");
}
}
@@ -10,7 +10,14 @@ import dev.ltms.bridged.herdr.AgentControl;
import dev.ltms.bridged.herdr.FakeHerdr;
import dev.ltms.bridged.herdr.WorkspaceControl;
import dev.ltms.bridged.member.ClaudeCodeLauncher;
import dev.ltms.bridged.msg.TestTurnTokens;
import dev.ltms.bridged.peer.Capability;
import dev.ltms.bridged.peer.CharterReceipt;
import dev.ltms.bridged.peer.MemberRole;
import dev.ltms.bridged.peer.PeerHandle;
import dev.ltms.bridged.peer.PeerLauncher;
import dev.ltms.bridged.peer.PeerUnreachableException;
import dev.ltms.bridged.peer.SpawnRequest;
import org.junit.jupiter.api.Test;
import org.slf4j.LoggerFactory;
@@ -43,6 +50,113 @@ class SessionManagerTest {
return sessionManager(herdr, clock, 0);
}
private SessionManager sessionManager(FakeHerdr herdr, Worktrees worktrees) {
return sessionManager(herdr, worktrees, System::nanoTime);
}
private SessionManager sessionManager(FakeHerdr herdr, Worktrees worktrees, LongSupplier clock) {
BridgedConfig.Profile cfg = new BridgedConfig.Profile(
"ltms-local", "http://gx00.gw:8000", "coder", null, "BRIDGED_WORKER_TOKEN",
List.of("ccs", "ltms-local"), "tab", "bridged-workers",
"worker: {profile} #{n}", null, null, null);
ClaudeCodeLauncher workers = new ClaudeCodeLauncher(new AgentControl(herdr), new WorkspaceControl(herdr),
new SubscriptionGuard(Set.of("gx00.gw")), Map.of(cfg.profile(), cfg), cfg.profile(), _ -> null);
return new SessionManager(workers, worktrees, clock);
}
/**
* CB-581: a {@link Worktrees} test double whose {@code hasUncommitted} and {@code remove} can
* be told to throw, so {@link SessionManager#release} can be exercised against exactly the
* failure {@code GitWorktrees} produces when {@code git status}/{@code git worktree remove}
* exits non-zero.
*/
private static final class RecordingWorktrees implements Worktrees {
private final List<String> removeCalls = new java.util.ArrayList<>();
private final List<String> snapshotCalls = new java.util.ArrayList<>();
private final java.util.Set<String> failRemoveFor = new java.util.HashSet<>();
private volatile boolean dirty = false;
private volatile RuntimeException hasUncommittedFailure;
private volatile RuntimeException snapshotFailure;
private final java.util.concurrent.atomic.AtomicLong snapshotSeq = new java.util.concurrent.atomic.AtomicLong();
RecordingWorktrees dirty(boolean dirty) {
this.dirty = dirty;
return this;
}
RecordingWorktrees failHasUncommittedWith(RuntimeException e) {
this.hasUncommittedFailure = e;
return this;
}
RecordingWorktrees failRemoveFor(String worktreePath) {
failRemoveFor.add(worktreePath);
return this;
}
RecordingWorktrees failSnapshotWith(RuntimeException e) {
this.snapshotFailure = e;
return this;
}
@Override
public String add(String repoRoot, String branch, String baseRef) {
return "/wt/" + branch.replace('/', '_');
}
@Override
public void remove(String repoRoot, String worktreePath) {
if (failRemoveFor.contains(worktreePath)) {
throw new WorktreeException("simulated remove failure for " + worktreePath);
}
removeCalls.add(worktreePath);
}
@Override
public boolean hasUncommitted(String worktreePath) {
if (hasUncommittedFailure != null) {
throw hasUncommittedFailure;
}
return dirty;
}
@Override
public void overlayParity(String repoRoot, String worktreePath, List<String> overlay) {
}
@Override
public String repoRoot(String cwd) {
return "/repo";
}
@Override
public java.util.Optional<String> snapshot(String worktreePath, String branch, String message) {
snapshotCalls.add(worktreePath);
if (snapshotFailure != null) {
throw snapshotFailure;
}
return java.util.Optional.of("wip" + snapshotSeq.incrementAndGet());
}
@Override
public WipRefStats wipRefs(String repoRoot) {
return new WipRefStats(0, 0L);
}
@Override
public int pruneWipRefs(String repoRoot, long minAgeMillis) {
return 0;
}
List<String> removeCalls() {
return List.copyOf(removeCalls);
}
List<String> snapshotCalls() {
return List.copyOf(snapshotCalls);
}
}
private SessionManager sessionManager(FakeHerdr herdr, LongSupplier clock, int contextCap) {
return sessionManager(herdr, clock, contextCap, false);
}
@@ -90,6 +204,24 @@ class SessionManagerTest {
assertEquals(2, sessions.roster().size(), "both sessions are registered");
}
@Test
void rosterViewExposesTheCharterReceiptButNeverTheCharterText() {
// The roster (bridge_list and GET /members both render through rosterView) must let a lead
// see which charter a member got, without ever carrying the charter prose itself (CB-571).
MemberSession s = new MemberSession("p1", "term1", "prof", MemberRole.DEV, "/cwd", null,
0, 0, 0, MemberSession.State.READY, null, null,
CharterReceipt.compose(MemberRole.DEV, "prof", "role charter", "role charter\n\nreply"), null);
Map<String, Object> view = SessionManager.rosterView(s, null);
assertEquals("fleet.charters.dev", view.get("charterSource"),
"the config key that supplied the role charter is reported");
assertEquals(CharterReceipt.digestOf("role charter\n\nreply"), view.get("charterSha256"),
"the digest of the exact composed charter bytes is reported");
assertFalse(view.values().toString().contains("role charter"),
"the roster row must not embed the charter text itself");
}
@Test
void aNullTerminalFromThePrimaryIsANoOpEvenWithSessionsRegistered() {
// The primary resolves to a Principal with no terminal, and BridgeMcp's context extractor
@@ -101,7 +233,7 @@ class SessionManagerTest {
assertDoesNotThrow(() -> sessions.asPresence().markPresent(null),
"the primary's null terminal must not blow up an unrelated tool call");
assertDoesNotThrow(() -> sessions.onDelivered(null));
assertDoesNotThrow(() -> sessions.onDelivered(null, TestTurnTokens.inert(null)));
assertDoesNotThrow(() -> sessions.onTurnComplete(null));
assertDoesNotThrow(() -> sessions.onTurnFailed(null));
@@ -121,7 +253,7 @@ class SessionManagerTest {
"MCP presence moves SPAWNING → READY");
assertTrue(sessions.asPresence().isPresent(terminal), "presence is also recorded");
sessions.onDelivered(terminal);
sessions.onDelivered(terminal, TestTurnTokens.inert(terminal));
assertEquals(MemberSession.State.BUSY, sessions.get(session.paneId()).orElseThrow().state(),
"delivery moves READY → BUSY");
@@ -153,7 +285,7 @@ class SessionManagerTest {
MemberSession session = sessions.acquire("ltms-local", null, "/caller", "term_primary");
String terminal = session.terminalId();
sessions.asPresence().markPresent(terminal);
sessions.onDelivered(terminal);
sessions.onDelivered(terminal, TestTurnTokens.inert(terminal));
sessions.onTurnFailed(terminal);
@@ -181,7 +313,7 @@ class SessionManagerTest {
MemberSession session = sessions.acquire("ltms-local", null, "/caller", "term_primary");
String terminal = session.terminalId();
sessions.asPresence().markPresent(terminal);
sessions.onDelivered(terminal);
sessions.onDelivered(terminal, TestTurnTokens.inert(terminal));
sessions.onTurnFailed(terminal);
@@ -264,7 +396,7 @@ class SessionManagerTest {
MemberSession session = sessions.acquire("ltms-local", null, "/caller", "term_primary");
String terminal = session.terminalId();
sessions.asPresence().markPresent(terminal);
sessions.onDelivered(terminal);
sessions.onDelivered(terminal, TestTurnTokens.inert(terminal));
clock[0] = 100;
assertEquals(0, sessions.reapIdle(10), "BUSY session past TTL is never reaped");
@@ -281,7 +413,7 @@ class SessionManagerTest {
MemberSession session = sessions.acquire("ltms-local", null, "/caller", "term_primary");
String terminal = session.terminalId();
sessions.asPresence().markPresent(terminal);
sessions.onDelivered(terminal);
sessions.onDelivered(terminal, TestTurnTokens.inert(terminal));
sessions.onTurnComplete(terminal);
clock[0] = 21;
@@ -299,7 +431,7 @@ class SessionManagerTest {
MemberSession busy = sessions.acquire("ltms-local", "/busy", "/caller", "owner2");
sessions.asPresence().markPresent(ready.terminalId());
sessions.asPresence().markPresent(busy.terminalId());
sessions.onDelivered(busy.terminalId());
sessions.onDelivered(busy.terminalId(), TestTurnTokens.inert(busy.terminalId()));
clock[0] = 50;
assertEquals(1, sessions.reapIdle(30), "only READY past TTL is reaped");
@@ -317,9 +449,9 @@ class SessionManagerTest {
String terminal = session.terminalId();
sessions.asPresence().markPresent(terminal);
sessions.onDelivered(terminal);
sessions.onDelivered(terminal, TestTurnTokens.inert(terminal));
sessions.onTurnComplete(terminal);
sessions.onDelivered(terminal);
sessions.onDelivered(terminal, TestTurnTokens.inert(terminal));
sessions.onTurnComplete(terminal);
MemberSession updated = sessions.get(session.paneId()).orElseThrow();
@@ -337,13 +469,13 @@ class SessionManagerTest {
String terminal = session.terminalId();
sessions.asPresence().markPresent(terminal);
sessions.onDelivered(terminal);
sessions.onDelivered(terminal, TestTurnTokens.inert(terminal));
sessions.onTurnComplete(terminal);
assertEquals(MemberSession.State.DONE,
sessions.get(session.paneId()).orElseThrow().state(),
"first turn completes without release");
sessions.onDelivered(terminal);
sessions.onDelivered(terminal, TestTurnTokens.inert(terminal));
sessions.onTurnComplete(terminal);
assertTrue(sessions.get(session.paneId()).isEmpty(), "session released after cap reached");
@@ -359,7 +491,7 @@ class SessionManagerTest {
MemberSession session = sessions.acquire("ltms-local", null, "/caller", "term_primary");
sessions.asPresence().markPresent(session.terminalId());
sessions.onDelivered(session.terminalId());
sessions.onDelivered(session.terminalId(), TestTurnTokens.inert(session.terminalId()));
assertTrue(sessions.onTurnCompleteWithPostAction(session.terminalId()));
MemberSession updated = sessions.get(session.paneId()).orElseThrow();
@@ -373,7 +505,7 @@ class SessionManagerTest {
SessionManager sessions = sessionManager(herdr, () -> 0L, 1, true);
MemberSession session = sessions.acquire("ltms-local", null, "/caller", "term_primary");
sessions.asPresence().markPresent(session.terminalId());
sessions.onDelivered(session.terminalId());
sessions.onDelivered(session.terminalId(), TestTurnTokens.inert(session.terminalId()));
assertFalse(sessions.hasPostTurnAction(session.terminalId()),
"a session at its cap will be released, not reset for reuse");
@@ -388,7 +520,7 @@ class SessionManagerTest {
SessionManager sessions = sessionManager(herdr, () -> 0L, 0, false);
MemberSession session = sessions.acquire("ltms-local", null, "/caller", "term_primary");
sessions.asPresence().markPresent(session.terminalId());
sessions.onDelivered(session.terminalId());
sessions.onDelivered(session.terminalId(), TestTurnTokens.inert(session.terminalId()));
sessions.onTurnComplete(session.terminalId());
@@ -406,7 +538,7 @@ class SessionManagerTest {
MemberSession busy = sessions.acquire("ltms-local", "/busy", "/caller", "ownerB");
sessions.asPresence().markPresent(ready.terminalId());
sessions.asPresence().markPresent(busy.terminalId());
sessions.onDelivered(busy.terminalId());
sessions.onDelivered(busy.terminalId(), TestTurnTokens.inert(busy.terminalId()));
sessions.drainAll(TimeUnit.MILLISECONDS.toNanos(100));
@@ -470,7 +602,7 @@ class SessionManagerTest {
FakeHerdr herdr = new FakeHerdr();
SessionManager sessions = sessionManager(herdr);
java.util.List<String> released = new java.util.concurrent.CopyOnWriteArrayList<>();
sessions.onRelease(released::add);
sessions.onRelease(detail -> released.add(detail.terminalId()));
MemberSession s = sessions.acquire("ltms-local", null, "/caller", null);
sessions.release(s.paneId());
@@ -484,7 +616,7 @@ class SessionManagerTest {
FakeHerdr herdr = new FakeHerdr();
SessionManager sessions = sessionManager(herdr);
java.util.List<String> released = new java.util.concurrent.CopyOnWriteArrayList<>();
sessions.onRelease(released::add);
sessions.onRelease(detail -> released.add(detail.terminalId()));
sessions.release("w9:p404"); // idempotent teardown of something already gone
@@ -504,4 +636,295 @@ class SessionManagerTest {
"a listener failure must never prevent the teardown it is reacting to");
assertTrue(sessions.get(s.paneId()).isEmpty(), "and the session is still deregistered");
}
// --- CB-581: a throw inside release() must not orphan the pane or abort reapIdle -----------
@Test
void releasePreservesWorktreeWhenDirtyCheckThrows() {
FakeHerdr herdr = new FakeHerdr();
RecordingWorktrees worktrees = new RecordingWorktrees();
SessionManager sessions = sessionManager(herdr, worktrees);
MemberSession s = sessions.acquire("ltms-local", null, "/caller/proj", null,
new WorktreeRequest("cb-581a", null));
worktrees.failHasUncommittedWith(new WorktreeException("git status exited 128"));
LoggerContext ctx = (LoggerContext) LoggerFactory.getILoggerFactory();
ch.qos.logback.classic.Logger sessionLog =
(ch.qos.logback.classic.Logger) LoggerFactory.getLogger(SessionManager.class);
ListAppender<ILoggingEvent> appender = new ListAppender<>();
appender.setContext(ctx);
appender.start();
sessionLog.addAppender(appender);
sessionLog.setLevel(Level.WARN);
try {
assertDoesNotThrow(() -> sessions.release(s.paneId()),
"a throwing dirty check must not abort the release");
assertTrue(worktrees.removeCalls().isEmpty(),
"the worktree is preserved when its dirty state cannot be determined");
String warn = appender.list.stream()
.filter(e -> e.getLevel().equals(Level.WARN))
.map(ILoggingEvent::getFormattedMessage)
.filter(m -> m.contains(s.worktree()))
.findFirst()
.orElse("no warn logged naming the worktree");
assertTrue(warn.contains(s.paneId()), "the WARN names the pane: " + warn);
assertTrue(warn.contains(s.terminalId()), "the WARN names the terminal: " + warn);
} finally {
sessionLog.detachAppender(appender);
}
}
@Test
void releaseStillStopsThePaneWhenDirtyCheckThrows() {
FakeHerdr herdr = new FakeHerdr();
RecordingWorktrees worktrees = new RecordingWorktrees();
SessionManager sessions = sessionManager(herdr, worktrees);
MemberSession s = sessions.acquire("ltms-local", null, "/caller/proj", null,
new WorktreeRequest("cb-581b", null));
worktrees.failHasUncommittedWith(new WorktreeException("git status exited 128"));
sessions.release(s.paneId());
assertEquals(1, paneCloseCallsFor(herdr, "w9:pRoot_1"),
"the pane is stopped exactly once even though the dirty check threw");
}
@Test
void releaseStillNotifiesTheListenerWhenDirtyCheckThrows() {
FakeHerdr herdr = new FakeHerdr();
RecordingWorktrees worktrees = new RecordingWorktrees();
SessionManager sessions = sessionManager(herdr, worktrees);
java.util.List<String> released = new java.util.concurrent.CopyOnWriteArrayList<>();
sessions.onRelease(detail -> released.add(detail.terminalId()));
MemberSession s = sessions.acquire("ltms-local", null, "/caller/proj", null,
new WorktreeRequest("cb-581c", null));
worktrees.failHasUncommittedWith(new WorktreeException("git status exited 128"));
sessions.release(s.paneId());
assertEquals(java.util.List.of(s.terminalId()), released,
"a blocked caller must still be told the terminal was released, even though the "
+ "dirty check threw");
}
@Test
void reapIdleSurvivesOneSessionThatFailsToRelease() {
long[] clock = {0};
FakeHerdr herdr = new FakeHerdr();
RecordingWorktrees worktrees = new RecordingWorktrees();
SessionManager sessions = sessionManager(herdr, worktrees, () -> clock[0]);
MemberSession a = sessions.acquire("ltms-local", null, "/caller/proj", null,
new WorktreeRequest("cb-581d", null));
MemberSession b = sessions.acquire("ltms-local", null, "/caller/proj", null,
new WorktreeRequest("cb-581e", null));
MemberSession c = sessions.acquire("ltms-local", null, "/caller/proj", null,
new WorktreeRequest("cb-581f", null));
sessions.asPresence().markPresent(a.terminalId());
sessions.asPresence().markPresent(b.terminalId());
sessions.asPresence().markPresent(c.terminalId());
// The middle session's worktree removal fails — release() propagates that, so this is the
// one call reapIdle's per-session guard must survive without skipping the rest of the pass.
worktrees.failRemoveFor(b.worktree());
LoggerContext ctx = (LoggerContext) LoggerFactory.getILoggerFactory();
ch.qos.logback.classic.Logger sessionLog =
(ch.qos.logback.classic.Logger) LoggerFactory.getLogger(SessionManager.class);
ListAppender<ILoggingEvent> appender = new ListAppender<>();
appender.setContext(ctx);
appender.start();
sessionLog.addAppender(appender);
sessionLog.setLevel(Level.WARN);
int reaped;
try {
clock[0] = 100;
reaped = sessions.reapIdle(10);
String warn = appender.list.stream()
.filter(e -> e.getLevel().equals(Level.WARN))
.map(ILoggingEvent::getFormattedMessage)
.filter(m -> m.contains(b.paneId()))
.findFirst()
.orElse("no reap-failure WARN logged");
assertTrue(warn.contains(b.terminalId()), "the WARN names the failed session's terminal: " + warn);
assertTrue(warn.contains(b.worktree()), "the WARN names the failed session's worktree: " + warn);
} finally {
sessionLog.detachAppender(appender);
}
assertEquals(2, reaped, "the middle session's failure is logged, not counted as reaped");
assertTrue(sessions.get(a.paneId()).isEmpty(), "the first session is still released");
assertTrue(sessions.get(c.paneId()).isEmpty(), "the third session is still released");
assertTrue(sessions.get(b.paneId()).isEmpty(),
"the middle session is still deregistered even though its worktree removal threw");
assertEquals(1, paneCloseCallsFor(herdr, "w9:pRoot_1"), "the first pane is stopped");
assertEquals(1, paneCloseCallsFor(herdr, "w9:pRoot_2"),
"the middle pane is still stopped even though its worktree removal failed");
assertEquals(1, paneCloseCallsFor(herdr, "w9:pRoot_3"), "the third pane is stopped");
}
@Test
void unchangedRegressionCleanCompletedReleaseStillRemovesTheWorktree() {
FakeHerdr herdr = new FakeHerdr();
RecordingWorktrees worktrees = new RecordingWorktrees();
SessionManager sessions = sessionManager(herdr, worktrees);
MemberSession s = sessions.acquire("ltms-local", null, "/caller/proj", null,
new WorktreeRequest("cb-581g", null));
sessions.release(s.paneId());
assertEquals(List.of(s.worktree()), worktrees.removeCalls(),
"COMPLETED release of a clean worktree still removes it");
}
@Test
void unchangedRegressionDirtyCompletedReleaseStillPreservesTheWorktree() {
FakeHerdr herdr = new FakeHerdr();
RecordingWorktrees worktrees = new RecordingWorktrees().dirty(true);
SessionManager sessions = sessionManager(herdr, worktrees);
MemberSession s = sessions.acquire("ltms-local", null, "/caller/proj", null,
new WorktreeRequest("cb-581h", null));
sessions.release(s.paneId());
assertTrue(worktrees.removeCalls().isEmpty(),
"COMPLETED release of a dirty worktree still preserves it");
}
@Test
void unchangedRegressionShutdownDrainStillPreservesTheWorktree() {
FakeHerdr herdr = new FakeHerdr();
RecordingWorktrees worktrees = new RecordingWorktrees();
SessionManager sessions = sessionManager(herdr, worktrees);
MemberSession s = sessions.acquire("ltms-local", null, "/caller/proj", null,
new WorktreeRequest("cb-581i", null));
sessions.asPresence().markPresent(s.terminalId());
sessions.drainAll(TimeUnit.MILLISECONDS.toNanos(100));
assertTrue(worktrees.removeCalls().isEmpty(), "SHUTDOWN drain still preserves the worktree");
}
// ── CB-584: agentSessionId recorded at acquire, and gated by Capability.SESSION_RESUME ─────
/**
* A minimal {@link PeerLauncher} whose capabilities exclude SESSION_RESUME — the "does not
* support it" case. {@link #requireResumeCapability} throws before ever reaching {@link #spawn},
* so every method beyond {@link #capabilitiesFor} is unreachable in these tests and left unimplemented.
*/
private static final class NoResumeLauncher implements PeerLauncher {
@Override
public Set<Capability> capabilities() {
return Set.of(Capability.WORKTREE); // deliberately no SESSION_RESUME
}
@Override
public Set<Capability> capabilitiesFor(String profileName) {
return capabilities();
}
@Override
public PeerHandle spawn(SpawnRequest req) {
throw new UnsupportedOperationException("not reachable — the capability check refuses first");
}
@Override
public Set<String> profiles() {
return Set.of("stub-profile");
}
@Override
public String defaultProfile() {
return "stub-profile";
}
@Override
public String effectiveCwd(SpawnRequest req) {
throw new UnsupportedOperationException("not reachable — the capability check refuses first");
}
@Override
public List<String> parityOverlay(String profileName) {
return List.of();
}
@Override
public List<?> list() {
return List.of();
}
@Override
public int reapOrphanWorkers() {
return 0;
}
@Override
public void stop(String id) {
}
@Override
public boolean clearContext(String id) {
return false;
}
}
@Test
void acquireRecordsTheAgentSessionIdFromTheHandle() {
FakeHerdr herdr = new FakeHerdr();
SessionManager sessions = sessionManager(herdr);
MemberSession s = sessions.acquire("ltms-local", MemberRole.DEV, null, null, null, null,
"my-session-name", null);
assertNotNull(s.agentSessionId(), "a spawn that asked for session identity gets one back");
Map<String, Object> view = SessionManager.rosterView(s, null);
assertEquals(s.agentSessionId(), view.get("agentSessionId"),
"the roster exposes the same id the session recorded");
}
@Test
void acquireWithNeitherSessionFieldLeavesAgentSessionIdNull() {
FakeHerdr herdr = new FakeHerdr();
SessionManager sessions = sessionManager(herdr);
MemberSession s = sessions.acquire("ltms-local", null, null, null);
assertNull(s.agentSessionId(), "no identity requested — unchanged from before CB-584");
assertFalse(SessionManager.rosterView(s, null).containsKey("agentSessionId"),
"a null id is omitted from the roster, like charterSha256 for a receipt-less session");
}
@Test
void resumeSessionIdResumesTheSameConversationOnASupportingAdapter() {
FakeHerdr herdr = new FakeHerdr();
SessionManager sessions = sessionManager(herdr);
MemberSession s = sessions.acquire("ltms-local", MemberRole.DEV, null, null, null, null,
null, "cb-resume-77");
assertEquals("cb-resume-77", s.agentSessionId(),
"a resume adopts the prior id as its own agentSessionId");
}
@Test
void resumeSessionIdWithoutAnExplicitProfileIsRefused() {
FakeHerdr herdr = new FakeHerdr();
SessionManager sessions = sessionManager(herdr);
IllegalArgumentException e = assertThrows(IllegalArgumentException.class, () ->
sessions.acquire(null, MemberRole.DEV, null, null, null, null, null, "cb-resume-1"));
assertTrue(e.getMessage().contains("explicit profile"), e.getMessage());
}
@Test
void resumeSessionIdOnAnAdapterWithoutTheCapabilityIsRefusedNamingIt() {
SessionManager sessions = new SessionManager(new NoResumeLauncher());
IllegalArgumentException e = assertThrows(IllegalArgumentException.class, () ->
sessions.acquire("stub-profile", MemberRole.DEV, null, null, null, null, null, "cb-resume-1"));
assertTrue(e.getMessage().contains("SESSION_RESUME"), e.getMessage());
assertTrue(e.getMessage().contains("stub-profile"), e.getMessage());
}
}
@@ -6,9 +6,15 @@ import dev.ltms.bridged.guard.SubscriptionGuard;
import dev.ltms.bridged.herdr.AgentControl;
import dev.ltms.bridged.herdr.FakeHerdr;
import dev.ltms.bridged.herdr.WorkspaceControl;
import ch.qos.logback.classic.Level;
import ch.qos.logback.classic.LoggerContext;
import ch.qos.logback.classic.spi.ILoggingEvent;
import ch.qos.logback.core.read.ListAppender;
import dev.ltms.bridged.member.ClaudeCodeLauncher;
import dev.ltms.bridged.msg.TestTurnTokens;
import dev.ltms.bridged.peer.MemberRole;
import org.junit.jupiter.api.Test;
import org.slf4j.LoggerFactory;
import java.util.List;
import java.util.Map;
@@ -178,6 +184,49 @@ class WorktreeSessionManagerTest {
assertTrue(sessions.get(paneId).isEmpty(), "released session is no longer retrievable");
}
/**
* CB-576. A normal {@code COMPLETED} release whose worktree holds uncommitted work must NOT
* remove it — {@code --force} would destroy the worker's only copy. The bridge cannot see
* uncommitted files, so the worktree is preserved and the release logged at WARN naming the
* path, the session, and the cause an operator needs to find the work.
*/
@Test
void releasePreservesDirtyWorktreeAndLogsWarn() {
FakeHerdr herdr = new FakeHerdr();
FakeWorktrees worktrees = new FakeWorktrees().withRepoRoot("/repo").withPrefix("/wt")
.withDirty(true);
SessionManager sessions = new SessionManager(workerService(herdr), worktrees);
MemberSession s = sessions.acquire("ltms-local", null, "/caller/proj", null,
new WorktreeRequest("cb-576", null));
LoggerContext ctx = (LoggerContext) LoggerFactory.getILoggerFactory();
ch.qos.logback.classic.Logger sessionLog =
(ch.qos.logback.classic.Logger) LoggerFactory.getLogger(SessionManager.class);
ListAppender<ILoggingEvent> appender = new ListAppender<>();
appender.setContext(ctx);
appender.start();
sessionLog.addAppender(appender);
sessionLog.setLevel(Level.WARN);
try {
sessions.release(s.paneId());
assertTrue(herdr.called("pane.close"), "release still tears the worker pane down");
assertTrue(worktrees.removeCalls().isEmpty(),
"a dirty worktree is never removed — it holds the only copy of the work");
String warn = appender.list.stream()
.filter(e -> e.getLevel().equals(Level.WARN))
.map(ILoggingEvent::getFormattedMessage)
.filter(m -> m.contains("dirty worktree"))
.findFirst()
.orElse("no dirty-release WARN logged");
assertTrue(warn.contains(s.worktree()), "the WARN names the worktree path: " + warn);
assertTrue(warn.contains(s.terminalId()), "the WARN names the session: " + warn);
assertTrue(warn.contains("COMPLETED"), "the WARN names the release cause: " + warn);
} finally {
sessionLog.detachAppender(appender);
}
}
@Test
void drainAllPreservesWorktreeOfIdleSession() {
FakeHerdr herdr = new FakeHerdr();
@@ -204,7 +253,7 @@ class WorktreeSessionManagerTest {
new WorktreeRequest("cb-544", null));
String terminal = s.terminalId();
sessions.asPresence().markPresent(terminal);
sessions.onDelivered(terminal); // BUSY, never completes → still BUSY when the timeout hits
sessions.onDelivered(terminal, TestTurnTokens.inert(terminal)); // BUSY, never completes → still BUSY when the timeout hits
sessions.drainAll(TimeUnit.MILLISECONDS.toNanos(100));
@@ -313,4 +362,119 @@ class WorktreeSessionManagerTest {
"the profile's configured cwd must reach repoRoot, not be ignored");
}
// --- CB-578 stage C: snapshot a dirty worktree into git before the preserve-or-remove decision ---
@Test
void releaseOfACleanWorktreeNeverSnapshots() {
FakeHerdr herdr = new FakeHerdr();
FakeWorktrees worktrees = new FakeWorktrees().withRepoRoot("/repo").withPrefix("/wt");
SessionManager sessions = new SessionManager(workerService(herdr), worktrees);
MemberSession s = sessions.acquire("ltms-local", null, "/caller/proj", null,
new WorktreeRequest("cb-578c-a", null));
sessions.release(s.paneId());
assertTrue(worktrees.snapshotCalls().isEmpty(),
"acceptance criterion 4: a clean release must create no snapshot ref or commit");
assertEquals(1, worktrees.removeCalls().size(), "a clean worktree is still removed as before");
}
@Test
void releaseOfADirtyWorktreeSnapshotsBeforePreserving() {
FakeHerdr herdr = new FakeHerdr();
FakeWorktrees worktrees = new FakeWorktrees().withRepoRoot("/repo").withPrefix("/wt")
.withDirty(true);
SessionManager sessions = new SessionManager(workerService(herdr), worktrees);
MemberSession s = sessions.acquire("ltms-local", null, "/caller/proj", null,
new WorktreeRequest("cb-578c-b", null));
sessions.release(s.paneId());
assertEquals(1, worktrees.snapshotCalls().size(),
"acceptance criterion 1: a dirty release snapshots the worktree exactly once");
FakeWorktrees.SnapshotCall call = worktrees.lastSnapshot();
assertEquals(s.worktree(), call.worktreePath());
assertEquals(s.branch(), call.branch());
assertTrue(worktrees.removeCalls().isEmpty(), "the dirty worktree is still preserved, not removed");
}
@Test
void releaseNotifiesTheListenerWithWorktreeBranchAndSnapshotRef() {
FakeHerdr herdr = new FakeHerdr();
FakeWorktrees worktrees = new FakeWorktrees().withRepoRoot("/repo").withPrefix("/wt")
.withDirty(true);
SessionManager sessions = new SessionManager(workerService(herdr), worktrees);
java.util.List<SessionManager.ReleaseDetail> released = new java.util.concurrent.CopyOnWriteArrayList<>();
sessions.onRelease(released::add);
MemberSession s = sessions.acquire("ltms-local", null, "/caller/proj", null,
new WorktreeRequest("cb-578c-c", null));
sessions.release(s.paneId());
assertEquals(1, released.size());
SessionManager.ReleaseDetail detail = released.getFirst();
assertEquals(s.terminalId(), detail.terminalId());
assertEquals(s.worktree(), detail.worktreePath(),
"acceptance criterion 6: a failed ticket's detail must carry the worktree path");
assertEquals(s.branch(), detail.branch(),
"acceptance criterion 6: a failed ticket's detail must carry the branch");
assertEquals("wip1", detail.snapshotRef(),
"acceptance criterion 6: a failed ticket's detail must carry the snapshot ref");
}
@Test
void aFailingSnapshotStillPreservesTheWorktreeStopsThePaneAndNotifies() {
FakeHerdr herdr = new FakeHerdr();
FakeWorktrees worktrees = new FakeWorktrees().withRepoRoot("/repo").withPrefix("/wt")
.withDirty(true).failSnapshot("git commit-tree exited 128");
SessionManager sessions = new SessionManager(workerService(herdr), worktrees);
java.util.List<SessionManager.ReleaseDetail> released = new java.util.concurrent.CopyOnWriteArrayList<>();
sessions.onRelease(released::add);
MemberSession s = sessions.acquire("ltms-local", null, "/caller/proj", null,
new WorktreeRequest("cb-578c-d", null));
assertDoesNotThrow(() -> sessions.release(s.paneId()),
"acceptance criterion 5: a failing snapshot must not propagate out of release()");
assertTrue(herdr.called("pane.close"), "acceptance criterion 5: the pane must still stop");
assertTrue(worktrees.removeCalls().isEmpty(),
"acceptance criterion 5: the worktree must still be preserved when the snapshot fails");
assertEquals(1, released.size(), "acceptance criterion 5: the listener must still be notified");
assertNull(released.getFirst().snapshotRef(),
"a failed snapshot leaves no ref to report");
}
/**
* CB-586, criteria 4 and 1. Once a worktree session is spawned the repo is known, so the
* operator census and the retention sweep delegate to that repo's {@code refs/wip/*}.
*/
@Test
void wipRefsAndSweepDelegateToTheFleetRepoOnceKnown() {
FakeHerdr herdr = new FakeHerdr();
FakeWorktrees worktrees = new FakeWorktrees().withRepoRoot("/repo").withPrefix("/wt")
.withWipRefs(new Worktrees.WipRefStats(3, 42L)).withPruneResult(2);
SessionManager sessions = new SessionManager(workerService(herdr), worktrees);
assertTrue(sessions.wipRefs().isEmpty(),
"no worktree spawned yet means no repo is known and nothing to report");
assertEquals(0, sessions.sweepWipRefs(TimeUnit.HOURS.toMillis(24)),
"no worktree spawned yet means the sweep is a no-op");
sessions.acquire("ltms-local", null, "/caller/proj", null,
new WorktreeRequest("cb-586", null));
Worktrees.WipRefStats stats = sessions.wipRefs().orElseThrow();
assertEquals(3, stats.count(), "the census comes from the fleet repo");
assertEquals(42L, stats.costBytes(), "the cost comes from the fleet repo");
assertEquals(2, sessions.sweepWipRefs(TimeUnit.HOURS.toMillis(24)),
"the sweep runs against the fleet repo");
List<FakeWorktrees.PruneCall> prunes = worktrees.pruneCalls();
assertEquals(1, prunes.size(), "the no-op short-circuits before reaching the seam, so only "
+ "the repo-known sweep issues a call");
assertEquals("/repo", prunes.getFirst().repoRoot(), "the sweep targets the fleet repo");
assertEquals(TimeUnit.HOURS.toMillis(24), prunes.getFirst().minAgeMillis(),
"the caller's age floor is passed through");
}
}
+4 -1
View File
@@ -16,6 +16,9 @@
AuditLogTest, which attaches its own ListAppender and asserts on emitted records.
-->
<!-- Keep the test logger behaviour aligned with the production cancellation filter. -->
<turboFilter class="dev.ltms.bridged.logging.McpCancelledNotificationFilter"/>
<appender name="STDOUT" class="ch.qos.logback.core.ConsoleAppender">
<encoder>
<pattern>%d{HH:mm:ss.SSS} [%thread] %-5level %logger{36} - %msg%n</pattern>
@@ -34,4 +37,4 @@
<appender-ref ref="STDOUT"/>
</root>
</configuration>
</configuration>
+467
View File
@@ -0,0 +1,467 @@
# CB-591 — move the fleet onto the LLM and MCP gateway
**Status: DONE — the fleet is on the gateway as of 2026-08-15.** `local` runs on `/anthropic` and
`gx` on `/v1`, both at `weight: 100`; `local-direct` stays at `weight: 0` as the escape hatch. Getting
here took a revert and two upstream fixes — see §7.1, which is the useful part of this document. One
risk is **accepted rather than solved**: a stream cut by any mid-response timer arrives as HTTP 200
with no terminator, and our third-party members cannot detect it (§7.2).
· **Upstream:** [systems/vms wiki → LLM and MCP Gateway](https://git.ltms.dev/systems/vms/wiki/LLM-and-MCP-Gateway)
· **Upstream issue:** [systems/vms#31](https://git.ltms.dev/systems/vms/issues/31)
The gateway went live on 2026-08-15 and replaced Bifrost. This plan says what that means for a
**member definition** in `bridged.yaml`, because that is the part of this repo the change actually
touches.
---
## 1. What changed upstream
One front door for every LLM and MCP client: `https://llm.ltms.dev`, one token per consumer.
| Surface | URL |
|---|---|
| OpenAI chat | `https://llm.ltms.dev/v1/chat/completions` |
| OpenAI models | `https://llm.ltms.dev/v1/models` |
| **Anthropic messages** | `https://llm.ltms.dev/anthropic/v1/messages` |
| MCP, all servers multiplexed | `https://llm.ltms.dev/mcp` |
Anything outside that list returns **404 before any token is checked**, on purpose — the gateway must
never become a blanket proxy.
The model backend is unchanged: GX10 vLLM at `10.10.10.26:8000` (`gx00.gw`), model name exactly
`deepseek-v4-flash`. The direct LAN path stays open on purpose as an escape hatch.
---
## 2. Where claude-bridge sits today
We do **not** use the gateway. The `local` profile talks straight to the vLLM:
```yaml
local:
kind: claude-code
baseUrl: http://gx00.gw:8000 # direct vLLM — no auth, LAN only
model: deepseek-v4-flash
configDir: /Users/dai.ha/.ccs/instances/gx10
```
Three facts about our side that decide the shape of this work:
1. **`baseUrl` becomes `ANTHROPIC_BASE_URL`** in the member's environment, and `tokenEnv` becomes
`ANTHROPIC_AUTH_TOKEN` (the value is read from a host env var and never stored in config).
`local` sets no `tokenEnv` today, because a direct vLLM needs no token.
2. **`SubscriptionGuard` refuses any host not on an allowlist**, and that allowlist is
`guard.offSubscriptionHosts: [gx00.gw]`. It is built once in `Bridged.java:93` and handed to the
launcher, so **it is a restart-required key**, not a hot one. Changing `baseUrl` without changing
this makes every `local` spawn throw.
3. **The wiki names us as a blocker.** Under *Not done yet*: retiring the shared `legacy` token is
blocked because "kb, brain, **claude-bridge** and the workstation still share it. Each needs its
own consumer first."
Context7 is mounted twice today, both times straight at `https://ct7.ltms.dev/mcp` — once in
`.mcp.json` (the primary) and once in `opencode.json` (the `sol` and `terra` members).
```mermaid
flowchart LR
subgraph now["Today"]
M1["local member<br/>claude-code"] -->|"ANTHROPIC_BASE_URL"| V1["vLLM gx00.gw:8000<br/>no auth, LAN only"]
M2["sol / terra<br/>opencode"] --> CT1["ct7.ltms.dev/mcp"]
P1["primary"] --> CT1
end
subgraph after["Proposed"]
M3["local member"] -->|"ANTHROPIC_BASE_URL<br/>+ ANTHROPIC_AUTH_TOKEN"| G["llm.ltms.dev/anthropic<br/>consumer: claude-bridge"]
G --> V2["vLLM gx00.gw:8000"]
M4["local-direct<br/>weight 0, escape hatch"] --> V2
end
```
*The member definition is the only thing that moves. The model behind it does not.*
---
## 3. The member definition change
The gateway serves an Anthropic surface *and* an OpenAI surface, so **both member kinds can point at
it**. That is the main opportunity here, and it is bigger than the `local` profile alone.
### 3a. `local` — claude-code, on `/anthropic`
| Key | Today | After | Note |
|---|---|---|---|
| `baseUrl` | `http://gx00.gw:8000` | `https://llm.ltms.dev/anthropic` | see the schema warning below |
| `tokenEnv` | *(unset)* | `AI_GATEWAY_TOKEN` | new consumer token, `llmk-claude-bridge-<32 hex>` |
| `model` | `deepseek-v4-flash` | unchanged | must stay **exact**; a regex match returns an empty `/v1/models` while completions keep working |
| `guard.offSubscriptionHosts` | `[gx00.gw]` | `[gx00.gw, llm.ltms.dev]` | **restart required** |
### 3b. A new opencode profile on `/v1` — no code needed
`OpenCodeLauncher` already supports a pinned OpenAI-compatible endpoint (CB-508). Given `baseUrl` it
writes a custom provider block into the worker's opencode config:
- `baseUrl` → `options.baseURL`. `openAiBaseUrl` appends `/v1` to a bare host, and takes a URL that
already has a path **as-is** — so `https://llm.ltms.dev/v1` works unchanged.
- `tokenEnv` → `options.apiKey` (falls back to a placeholder when unset, since a local vLLM ignores it).
- `model:` **must** be `<provider>/<model>` when `baseUrl` is set — a bare name is rejected loudly
rather than silently falling back to opencode's default gateway.
So the profile is pure config:
```yaml
gx:
kind: opencode
baseUrl: https://llm.ltms.dev/v1
tokenEnv: AI_GATEWAY_TOKEN
model: gx/deepseek-v4-flash # provider id is ours to choose; the half after / is the model
argv: ["opencode"]
mcpUrl: http://127.0.0.1:8765/mcp
gitTokenEnv: WORKER_GITEA_TOKEN
weight: 100 # same tier as `local` — free
maxLoad: 2
# deliberately NO credentialId — this is our own box, not the shared OpenAI account
```
**Why this matters more than it looks.** Today every opencode member is `sol` or `terra`, and those
are two models on **one** OpenAI account sharing `credentialId: openai-shared` — so an exhaustion on
either locks out both, and half the fleet's opencode capacity dies at once. A gateway-backed opencode
profile is free, is not on that credential, and therefore is not in that quarantine pair. It removes
a single point of failure rather than just adding capacity.
**Note the asymmetry, it is deliberate:** `SubscriptionGuard` does not apply to opencode at all — the
guard exists to stop a *Claude* worker borrowing the operator's subscription, and opencode reads its
own provider credentials. So 3b needs **no allowlist change**; only 3a does.
**Both still need a restart, for a different reason.** `tokenEnv` is resolved by
`HerdrPeerLauncher.resolveEnv` → `env.apply(name)`, which reads the **daemon's own process
environment**. The running `bridged` inherited its environment when it started, so a variable added to
`secrets.sh` afterwards is simply not there — the launcher would inject an empty token and the
gateway would answer 401. This is the same failure as trap 1 in `scripts/redeploy-bridged.sh`
(`WORKER_GITEA_TOKEN`), and it has the same fix: **restart from a login shell**, and use
`scripts/redeploy-bridged.sh --check` to confirm the name resolves before restarting anything.
### 3c. What this does to `ccs`
Once a profile carries `baseUrl`, `tokenEnv` and `model` itself, the ccs instance stops being what
routes a member. Be precise about what is left, though: `configDir` still supplies **folder trust**
and `settings.json`, and dropping it is what produced the trust dialog and the wrong-model error
recorded in `bridged.yaml`. So ccs goes from *deciding where the tokens go* to *holding client-side
state*. Less load-bearing, not removable.
### Why `/anthropic` and never `/v1/chat/completions`
The gateway declares its Anthropic backend as `schema.name: Anthropic`, which means **no
translation** — streaming, tool use and thinking blocks pass through exactly as they do against vLLM
directly.
Declared as `OpenAI`, Envoy's translator looks for a `thinking_blocks` field that our vLLM does not
send (it sends `reasoning_content`), and **every thinking delta disappears silently**. Claude Code
speaks the Anthropic protocol, so `/anthropic` is both correct and the only safe choice.
This is the exact failure shape this repo keeps hitting: it compiles, it answers, it looks healthy,
and a capability is quietly off. Treat it as a `silent-default` risk, not a config preference.
**Open question for 3b — ANSWERED, 2026-08-15.** The worry was that the OpenAI surface might drop
reasoning the way the wiki documents for a mis-declared Anthropic backend. It does not. Checked at
the API before any profile was switched:
| surface | request | result |
|---|---|---|
| `/anthropic/v1/messages` | `deepseek-v4-flash`, 64 tokens | 200, response carries a real `"type":"thinking"` block |
| `/v1/chat/completions` | same | 200, message carries a populated `reasoning_content` (and a `reasoning` field) |
| `/v1/models` | — | 200, exactly `["deepseek-v4-flash"]` — the exact-name trap is clear |
| `/v1/models`, **no token** | — | **401** — Caddy is gating, as designed |
So reasoning survives on **both** surfaces, and the `/anthropic` choice for `local` is about protocol
correctness rather than a repair for a known loss. The last row matters on its own: the wiki warns
the gateway's own `SecurityPolicy` fails open, so it is worth knowing the proxy in front really does
refuse an unauthenticated request here.
---
## 4. Decisions
### D1 — switch, but keep the direct path as an explicit profile · **recommended**
Switching buys four things we do not have:
- **Free opencode capacity, off the shared credential.** The largest single win. See §3b — it retires
a real single point of failure, not just a cost line.
- **Per-consumer usage figures.** The cockpit counts requests per consumer. That is the first real
measurement of what the fleet consumes, and it feeds [CB-589](https://git.ltms.dev/lms/claude-bridge/issues/74) Gap 2 directly.
- **Our own revocable token.** One consumer to revoke if a worker ever leaks it, instead of a shared
`legacy` token used by four systems.
- **It works off-LAN.** `gx00.gw` resolves on the LAN only.
The cost is honest and worth stating: we add a TLS edge, an auth proxy and a gateway to the path of
every member spawn. The wiki keeps the direct route open precisely because "if the gateway breaks,
nothing that matters is blocked."
So keep it. Add a second profile `local-direct` pointing at `http://gx00.gw:8000` with **`weight: 0`**
— never auto-selected, still spawnable with an explicit `bridge_spawn{profile: "local-direct"}`.
That is exactly what CB-554 made `weight: 0` mean, and it turns the escape hatch into something the
lead can actually reach during an incident.
### D2 — do members also mount the gateway's `/mcp`? · **OPEN, operator's call**
Not a detail. `CLAUDE.md` states in two places that a member mounts **only** the bridge MCP, and a
worker's honesty rule leans on it ("never claim the result of a check you had no way to run").
- **Keep bridge-only.** The invariant stays true and simple. Workers stay cheap and narrow.
- **Add the gateway MCP.** Implementers get context7 documentation lookups, which is genuinely useful
for library work. But `mcpUrl` in `BridgedConfig.Profile` is a **single `String`**, so a
claude-code member can mount exactly one MCP — this needs a code change, not a config edit.
Note the invariant is **already inaccurate**: `opencode.json` gives `sol` and `terra` both context7
and gitea. So the choice is really "make the rule true" or "make the rule match reality". Either is
defensible; picking one is not mine to do.
### D3 — token scope
One consumer, `claude-bridge`, its token in `${SHARED_ENV}/tools/secrets.sh` as `AI_GATEWAY_TOKEN`,
referenced by name only. Never the literal value in `bridged.yaml` — `tokenEnv` exists for this.
---
## 5. Units of work
```mermaid
flowchart TB
U1["U1 · consumer token<br/>issue via cockpit, add to secrets.sh"]
U2["U2 · profile + guard<br/>bridged.yaml, restart"]
U3["U3 · verify live<br/>spawn, prove thinking survives"]
U4["U4 · context7 via gateway<br/>.mcp.json + opencode.json"]
U5["U5 · docs<br/>CLAUDE.md, wiki 11-Features"]
U1 --> U2 --> U3
U4 --> U5
U3 --> U5
```
| # | Scope | Who | Why |
|---|---|---|---|
| U1 | Issue the `claude-bridge` consumer at `auth.ltms.dev`; store as `AI_GATEWAY_TOKEN` | **operator** | touches secrets and a host we do not own |
| U2a | New `gx` opencode profile on `/v1` — pure config, no guard change | **lead** | `bridged.yaml` is gitignored, so a worker cannot see or edit it |
| U2b | `local` → `/anthropic`; add `local-direct` weight 0; add `llm.ltms.dev` to the guard allowlist | **lead** | same |
| U2c | One restart from a **login shell**, after U2a and U2b | **lead** | picks up `AI_GATEWAY_TOKEN` into the daemon env *and* the guard allowlist, in one stop |
| U3 | Live spawn on both new profiles; confirm reasoning survives on each surface | **lead** | needs real spawns and the running daemon |
| U4 | Point `.mcp.json` and `opencode.json` context7 at the gateway `/mcp`; rename pinned tools | delegatable | tracked files, self-contained |
| U5 | Fix the "members mount only the bridge" claim; add a `wiki/11-Features.md` entry | delegatable | writing, clear criteria |
U1 blocks U2a, U2b and U3. U4 and U5 do not depend on it.
**Write U2a and U2b, then restart once (U2c), then verify `gx` before `local`.** Since both profiles
need the same restart there is no reason to do two, but there is still a reason to *verify* in order:
`gx` exercises the token and the gateway with no guard involved, so if it fails the cause is upstream.
`local` adds the guard allowlist on top, so a failure there points at our config instead. Testing them
in that order separates the two causes instead of confusing them.
> **U1 status, 2026-08-15:** the operator issued the consumer and exported it as `AI_GATEWAY_TOKEN`
> (one key for every agent and MCP client behind `llm.ltms.dev`). Confirmed: it resolves in a login
> shell, is 48 characters and carries the documented `llmk-` prefix. The value was never printed.
---
## 6. Traps carried over from the wiki
Each of these cost someone real debugging time upstream. They apply to us.
1. **Rotating a token restarts the auth proxy, which drops in-flight streaming responses.** For us
that means rotating `AI_GATEWAY_TOKEN` kills every live member mid-turn, and an async ticket's
report goes with it. This is the same rule as a daemon redeploy: **drain the fleet first**
(`bridge_list` → `bridge_poll` anything wanted → `bridge_stop`), then rotate.
2. **The gateway's own `SecurityPolicy` fails open.** Standalone `aigw run` accepts it and silently
ignores it — an unauthenticated request returned **200**. Auth is the Caddy proxy in front, and
nothing else. Never reason as if the gateway authenticates.
3. **Exact model name.** A regex match routes fine but returns an **empty** `/v1/models` list while
completions keep working. A wrong name returns a bare 404 that reads exactly like a dead gateway.
4. **MCP tool names changed prefix separator.** Bifrost used one dash (`ct7-resolve-library-id`); the
gateway uses **two underscores** (`ct7__resolve-library-id`). Relevant only if U4 is done.
5. **`/v1/models` 404 vs empty list are different faults.** 404 means no route loaded at all; empty
means the model match is a regex. Do not conflate them when diagnosing.
---
## 7. Verification — what would prove this works
Merging config is not proving it. The checks, in order:
1. `bridge_spawn{profile: "gx"}` succeeds and the member completes a real turn ending in
`bridge_reply`. This is the first proof of the token, the URL and the model name, and it risks
nothing the fleet depends on.
2. `bridge_spawn{profile: "local"}` succeeds. If the guard allowlist was missed, this **throws** — a
loud, self-correcting failure, which is the good kind. If the restart was missed, it also throws,
for the same reason.
3. A `local` member completes a turn. That exercises streaming through two TLS edges, the auth proxy
and the gateway.
4. **Reasoning survives, checked separately on each surface.** For `local` on `/anthropic` this is
the check that catches the `/v1` versus `/anthropic` mistake, and it is the only one that does —
nothing else distinguishes a working passthrough from a translator quietly dropping thinking
deltas. For `gx` on `/v1`, this answers the open question in §3 rather than assuming it.
5. The cockpit at `auth.ltms.dev` shows requests counted against the `claude-bridge` consumer, not
`legacy`. That is the whole point of taking our own token.
6. `bridge_spawn{profile: "local-direct"}` still works, so the escape hatch is real rather than
theoretical.
7. `bridge_list` shows `gx` carrying no `credentialId`, so a `sol`/`terra` exhaustion cannot
quarantine it. This is the single-point-of-failure claim in §3b, checked rather than asserted.
---
## 7.1 What the live run actually found — 2026-08-15
U1–U2c were done, the daemon restarted onto them, and both new profiles were spawned for real. The
migration was then **reverted**. This section is the result, so none of it has to be re-derived.
### The blocker
`llm.ltms.dev` answers **HTTP 413 Request Entity Too Large** above **32 KiB (32768 bytes)**, on both
surfaces:
```
/v1 32695 bytes -> 200 /anthropic 32095 bytes -> 200
/v1 32795 bytes -> 413 /anthropic 32855 bytes -> 413
```
32 KiB is far below one real agent turn.
**Root cause — confirmed by the systems/vms side, 2026-08-15.** My guess that it was a Caddy
`request_body max_size` was **wrong**. It is Envoy, inside `aigw` on `llm.vm`. Envoy Gateway defaults
a listener's `per_connection_buffer_limit_bytes` to **32768**, and the AI Gateway buffers the *whole*
request body before it can route on the model name — so that default is not a network tuning knob
here, it is a hard ceiling on prompt size. Read out of the live Envoy `config_dump`:
```
listener default/llm/http per_connection_buffer_limit_bytes: 32768
```
Nobody chose 32 KiB; it was inherited from the default. Both TLS edges are innocent: the same
boundary reproduces on the LAN path and the internet path, and both 413s carry an `x-llm-consumer`
header their auth proxy sets only *after* authenticating — so the body cleared both edges and the
auth. Directly on `llm.vm`, `aigw` 413s at 39 KB while the vLLM backend accepts the same 39 KB and
answers 200.
**Do not plan around 32 KiB.** The intended ceiling is far higher. Their fix — a `ClientTrafficPolicy`
setting `bufferLimit: 8Mi` — is written but **not deployed** as of this note, pending their operator's
approval. I have not re-tested and will not until they confirm, so as not to measure a half-changed
system. Fixed in **systems/vms**, not here.
### The part worth remembering
Two members were spawned at the same moment with the same message:
| | `local` (claude-code, `/anthropic`) | `gx` (opencode, `/v1`) |
|---|---|---|
| READY → BUSY | 19:07:26 | 19:07:45 |
| BUSY → DONE | **19:08:51 (66s)** | **never — 10+ min, ticket FAILED** |
**`local` passed.** It passed only because the probe was three trivial questions in a fresh session,
so the request fit under 32 KiB. The profile looked healthy and was a landmine set to fire on the
first turn that reads a file.
So §7's checklist was not wrong, it was **too easy**. Any future run of it must use a task that reads
a real file. A liveness probe proves the token and the URL; it does not prove the path.
`gx` did not fail loudly either. Reproduced outside the bridge by running `opencode` by hand with the
launcher's own generated config:
```
Error: Request Entity Too Large
...compacts context, retries...
Error: Request Entity Too Large
```
opencode **catches the 413, compacts, and retries — indefinitely**. A member that fails loudly costs
one turn; this one costs the whole task and is indistinguishable from a slow worker.
> **Diagnosing a stuck opencode member.** Do not read its pane. The launcher writes its config to a
> temp dir and passes it as `OPENCODE_CONFIG` — find it with
> `ls -dt /var/folders/*/*/T/bridged-opencode-* | head -1`, check the provider block and the key's
> length and prefix (never its value), then reproduce with `opencode run --auto -m <provider>/<model>`
> using the same `OPENCODE_CONFIG`. That is what turned "it hangs" into a one-line error.
### What checked out, and needs no re-testing
- Token accepted on both surfaces. **Unauthenticated → 401**, so the Caddy proxy really does gate —
the wiki's "SecurityPolicy fails open" warning is about the gateway itself, not the edge.
- `/v1/models` returns exactly `["deepseek-v4-flash"]`, so trap 3 is clear.
- **Reasoning survives both surfaces** — see §3b above.
- The launcher's generated opencode provider block is correct, carrying a real 48-character `llmk-`
key rather than the `bridged-local-noauth` placeholder.
- `SubscriptionGuard` accepted `llm.ltms.dev` after the allowlist edit and the restart: `local`
spawned without throwing, which is the check that catches a missed restart.
### Resolution — both ceilings fixed, migration completed
systems/vms fixed both, and each was re-checked from this side rather than taken on trust:
| ceiling | was | now | our own check |
|---|---|---|---|
| listener buffer | 32 KiB | 32 Mi | 1.2 MB body → **200** (was 413) |
| LLM route timeout | 60s | 86400s | the request that truncated: **101s, `message_stop` present, 4000/4000** |
The timeout moved in two steps on 2026-08-15: 60s → 1800s, then 1800s → **86400s (24 hours)** after
the truncation risk below was discussed. They tried `request: 0s` first, which removes the
total-duration timer completely. It works, but on an `AIGatewayRoute` the **idle timeout is derived
from the request timeout**, so `0s` also removed any bound on a stalled connection. 86400s keeps a
reaper for dead connections while putting the truncation timer out of practical reach.
Neither was deliberate. The 32 KiB was Envoy Gateway's default `per_connection_buffer_limit_bytes`;
the 60s was Envoy AI Gateway's own documented default. The 60s bounded **generation** as well as
prompt size — a tiny prompt with a long answer returned 504 at 60.05s.
Two configuration facts worth keeping, from their bisection:
- **`ClientTrafficPolicy` is honoured in standalone `aigw run`; `BackendTrafficPolicy` is NOT.** A
`BackendTrafficPolicy` setting `requestTimeout` is accepted, logs nothing, and leaves the routes
unchanged (upstream `envoyproxy/gateway#9513`). What works is `timeouts: {request: …}` on each
`AIGatewayRoute` rule. Nothing from the outside distinguishes the two — the same silent-default
shape as their `SecurityPolicy` caveat.
- In that stack, "the config was accepted" proves nothing. Read the live `config_dump`.
## 7.2 The risk we accepted, and why we could not remove it
Raising the timeout made the failure **rare, not impossible**, and the residual failure is silent.
On a mid-response timeout over chunked HTTP/1.1, Envoy ends the chunked encoding *cleanly* instead of
resetting the connection, so the client receives what looks like a complete transfer
(`envoyproxy/envoy#17186` — acknowledged as a bug in 2021, closed by a stale bot, never fixed). The
December 2025 fix `envoyproxy/envoy#42269` changes locally-originated resets from `NO_ERROR` to
`INTERNAL_ERROR`, but it is **HTTP/2 only** and SSE clients here speak HTTP/1.1.
Measured on our side while the timeout was still 60s:
```
HTTP 200 61.07s 141992 bytes
message_stop 0 message_delta 0 error events 0
emitted 2473 of 4000, ending on a WELL-FORMED SSE frame
```
A syntactically valid stream that simply stops. Any timer firing mid-stream — route timeout, idle
timeout, `max_stream_duration` — fails this same way.
**The recommended defence does not transfer to us.** The right fix is to treat a stream with no
`message_stop` / `[DONE]` / `finish_reason` as failed. We cannot: our members are Claude Code and
opencode, third-party clients whose SSE parsing we do not own, and there is no seam to insert the
check. Whether either detects a missing terminator is unverified — and opencode's handling of the 413
(swallow, compact, retry forever, never surface an error) does not suggest it is strict.
So the honest statement of our position:
> Gateway traffic is acceptable at 86400s because a single request would have to run for 24 hours to
> trip the bug — **not** because we could detect it if it did.
At 86400s our **own** limit binds first, which is the ordering we want. `MessageService.ASYNC_TIMEOUT_MS`
caps a turn at 30 minutes, so a runaway request ends as a clean `FAILED` ticket that we raised, rather
than as a silently truncated `200` that we cannot see. While the gateway sat at 1800s the two numbers
were equal and did not nest, so a gateway-side stall could have been misread as a bug in our own ticket
handling. That ambiguity is now gone.
**If a member ever returns a confident but truncated answer, suspect this before anything in our own
code.** That is the whole reason this section exists.
---
## 8. Related
- [CB-589 / #74](https://git.ltms.dev/lms/claude-bridge/issues/74) — cost-first placement and a
gateway that reports live capacity. The per-consumer figures this migration unlocks are the first
input that ticket actually needs.
- `docs/CB-500-Multi-Tier-Coordination.md` §11 — the distributed-sandbox topology this gateway is
part of.
+59 -4
View File
@@ -1,7 +1,9 @@
# M4 - Fleet health, recovery, routing, and capacity
**Status:** Design accepted on 2026-08-15. CB-573 part 1 has shipped the classification model and
the `bridge_list` capacity view; the remaining M4 units are not yet shipped.
the `bridge_list` capacity view; the remaining M4 units are not yet shipped. See
[Unit 2 — what has landed so far](#unit-2---what-has-landed-so-far) before planning Unit 2 work:
some of its criteria were met by separate CB tickets, and one of them contradicts the unit text.
**Scope:** Fleet evidence, safe mechanical repair, lead routing, capacity reporting, and optional
human notification.
**Grounded in:** `health/FleetHealth`, `health/PaneBudget`, `inject/StatusPoller`,
@@ -237,7 +239,7 @@ state never presents stop as the only action.
| Release cause | Process action | Provisioned worktree |
|---|---|---|
| `SPAWN_ROLLBACK` before registration or delivery | Stop and clean up | Remove |
| `COMPLETED` for `READY` or `DONE` without pending work, idle TTL, or successful context-cap completion | Stop | Remove under completed policy |
| `COMPLETED` for `READY` or `DONE` without pending work, idle TTL, or successful context-cap completion | Stop | Remove only if clean; preserve a dirty worktree (CB-576) |
| `NEVER_READY` | Stop | Preserve |
| `GONE` | Best-effort stop | Preserve |
| `TURN_FAILED` or lead abort while `BUSY` or `FAILED` | Stop | Preserve |
@@ -744,8 +746,28 @@ worktree discovery.
Acceptance criteria:
1. Every accepted send receives a stable `TurnToken` tied to target, exact waiter, session turn,
delivery baseline, and task outcome.
1. Every accepted send receives a stable `TurnToken` tied to target, exact waiter, and delivery
baseline.
**Corrected during implementation (2026-08-15).** This criterion first also required the session
turn number and the task outcome. That is not implementable at this layer, and the implementer
refused it three times rather than fabricate a value — correctly. The reason is an ordering fact
that is invisible from any single class: `MessageService` owns acceptance and holds the waiter and
the async `Task`, but it learns nothing about delivery, because the delivery event goes to
`CompletionResolver` through `TurnListener.onDelivered`. And `CompletionResolver.onDelivered` runs
*before* `SessionManager.onDelivered`, so the session turn number does not exist yet at the only
point where the token could capture it.
Two ways out were rejected. A shared registry keyed by target reintroduces exactly the "whichever
send happens to be waiting" ambiguity the token exists to remove — the same weak claim
`Rendezvous.currentWaiter` warns about. Injecting a turn counter into `MessageService` adds a
required cross-layer dependency to populate a field that nothing in this slice reads, which is
speculative coupling across a boundary already shown to be fragile.
So the token identifies the **accepted send**, and `SessionManager` keeps verifying its own
delivery separately. Repair (criterion 2) does need the session turn; binding it means resolving
that acceptance-versus-delivery ordering first, and that work belongs to the repair unit, not
here. The token record carries a comment saying the field is deliberately absent.
2. Repair requires the same `BUSY` token, two raw `IDLE` or `DONE` snapshots, no conflicting
observation, exact open waiter, successful baseline, and new recognised assistant output.
3. Missing, failed, late, or post-restart baseline never authorises repair.
@@ -779,6 +801,39 @@ Acceptance criteria:
21. Tests cover both real traces, all repair refusals, clipping, explicit-reply and next-turn races,
restart without capture, concurrent send and release, and preserved discovery after restart.
#### Unit 2 - what has landed so far
Checked against `main` at `e09cac6` on 2026-08-15. Unit 2 was written as one block, but parts of it
have since been built by separate CB tickets. Read this before planning the rest, or that work gets
done twice.
The check was a symbol survey of `bridged/src/main/java` plus the merge history. It tells you whether
the machinery exists at all. It is **not** a line-by-line audit of whether each criterion is fully
met, and I did not run one.
| Criterion | Marker searched for | Found in main source | Reading |
|---|---|---|---|
| 1 | `TurnToken` | 8 files | **Done** — unit 2a, merged as `fec284e`. Criterion 1 was corrected first; see the note under it. |
| 2-5, 9 | `REPAIRED` | 0 files | Not started. The whole guarded-repair path is absent. |
| 6, 7, 10 | `reconcileLostBoundary` | 0 files | Not started. |
| 12 | CB-568 failure operation | via CB-580 | **Partial.** CB-580 (`0af902e`) routes `GONE` and `NEVER_READY` into the one idempotent target-wide failure. I did not check that release and abnormal stop go through the same call. |
| 14 | `DELEGATION_ORPHANED` | 3 files | **Partial.** The health state exists. The teardown-invariant check that creates it, and the retry rule, do not. |
| 15 | `SPAWN_ROLLBACK` | 0 files | **Contradicted — see below.** |
| 16 | — | — | Partial at best. CB-576 made release preserve a dirty worktree; whether explicit stop is state-aware is not checked. |
| 17, 18 | `preservedWorktrees` | 0 files | Not started. No manifest, and no lead-only `bridge_list` field. |
| 19 | `WORK_PRODUCT_AT_RISK` | 0 files | Not started. |
**Criterion 15 no longer matches the code, and the code is right.** It says "normal `COMPLETED`
remove worktrees". Since CB-576 (`500bfa2`) that is false on purpose: a `COMPLETED` release now
preserves the worktree when it still holds uncommitted work, because deleting it destroys work
nobody can get back. CB-576 was filed after exactly that loss. CB-581 goes further — if the
dirty-check itself fails, the worktree is preserved rather than removed, since "we could not tell"
must not be treated as "it is clean".
So criterion 15 should be rewritten as: `SPAWN_ROLLBACK` and a `COMPLETED` release with a **clean**
worktree remove it; abnormal causes, shutdown, a dirty worktree, and a failed dirty-check all
preserve it. `SPAWN_ROLLBACK` itself does not exist yet.
### Unit 3 - Typed inbox and member routing
Scope: semantic record, AMQP migration, both adapters, member routing, polling, and member health in
+1 -1
View File
@@ -26,7 +26,7 @@
],
"enabled": true,
"environment": {
"GITEA_ACCESS_TOKEN": "{env:GITEA_ACCESS_TOKEN}",
"GITEA_ACCESS_TOKEN": "{env:WORKER_GITEA_TOKEN}",
"GITEA_HOST": "{env:GITEA_HOST}"
}
}
+233
View File
@@ -0,0 +1,233 @@
#!/usr/bin/env bash
#
# Rebuild and restart the bridged daemon.
#
# A merge is not a deployment: the running daemon holds the jar it was started with, so code merged
# to main does nothing until this runs. See CLAUDE.md -> "Redeploying the daemon".
#
# This script exists to turn five remembered traps into one auditable command:
#
# 1. A piped `mvn` hides BUILD FAILURE behind a zero exit, so the build here is never piped.
# 2. The daemon must start from a LOGIN shell, or the tokens it hands to members are empty:
# WORKER_GITEA_TOKEN (workers cannot open a PR) and AI_GATEWAY_TOKEN (401 at llm.ltms.dev).
# Both are read from the DAEMON's own environment at spawn time, so a value added to
# secrets.sh after startup is absent. Nothing logs this, so the script checks and says so.
# 3. An old daemon that never actually died looks identical from the outside, so the script waits
# for the process to exit and for the port to free before it starts a new one.
# 4. "It started" is not "it works": the script polls /healthz until it answers, and reports the
# herdr protocol number, because healthz can be green while every spawn fails on a protocol
# mismatch.
# 5. Restarting under live members drops their tickets, so the script refuses unless you confirm
# the fleet is drained.
#
# Usage:
# scripts/redeploy-bridged.sh # build, confirm, restart, verify
# scripts/redeploy-bridged.sh --yes # skip the drain confirmation (fleet already checked)
# scripts/redeploy-bridged.sh --no-build # restart the jar already on disk
# scripts/redeploy-bridged.sh --check # report state and exit; changes nothing
#
# Exits non-zero on any failure. A failed build never stops the running daemon.
set -euo pipefail
REPO="$(cd "$(dirname "${BASH_SOURCE[0]}")/.." && pwd)"
BRIDGED="$REPO/bridged"
JAR="$BRIDGED/target/bridged.jar"
OUT="$BRIDGED/bridged.out"
# Matches BOTH the absolute form and the relative `java -jar target/bridged.jar` a hand-start
# produces from inside bridged/. Anchoring on the absolute path alone was a real bug: the daemon
# restarted correctly and the script still reported "no process appeared", because it launched with
# a relative path and then looked for an absolute one.
PATTERN='target/bridged.jar'
HEALTH='http://127.0.0.1:8765/healthz'
STOP_WAIT=30 # seconds to wait for a clean exit before reporting failure
HEALTH_WAIT=60 # seconds to wait for /healthz to answer after start
DO_BUILD=1; ASSUME_YES=0; CHECK_ONLY=0
for arg in "$@"; do
case "$arg" in
--yes|-y) ASSUME_YES=1 ;;
--no-build) DO_BUILD=0 ;;
--check) CHECK_ONLY=1 ;;
-h|--help) sed -n '3,30p' "${BASH_SOURCE[0]}"; exit 0 ;;
*) echo "unknown option: $arg (try --help)" >&2; exit 2 ;;
esac
done
say() { printf '\n\033[1m== %s\033[0m\n' "$*"; }
ok() { printf ' ok %s\n' "$*"; }
warn() { printf ' WARN %s\n' "$*"; }
die() { printf '\n FAIL %s\n\n' "$*" >&2; exit 1; }
jar_id() { [ -f "$JAR" ] && shasum -a 256 "$JAR" | cut -c1-12 || echo "absent"; }
running_pid() { pgrep -f "$PATTERN" || true; }
# ---------------------------------------------------------------- report state
say "current state"
OLD_PID="$(running_pid)"
if [ -n "$OLD_PID" ]; then
ok "daemon running, pid $OLD_PID"
else
warn "no daemon running — this will be a cold start"
fi
ok "jar on disk: $(jar_id) ($([ -f "$JAR" ] && date -r "$JAR" '+%Y-%m-%d %H:%M:%S' || echo 'none'))"
ok "HEAD: $(git -C "$REPO" log --oneline -1)"
# The trap with no log line. Checked in a LOGIN shell, because that is how the daemon is started
# below. Never prints the value — only whether it resolved.
if zsh -lc '[ -n "${WORKER_GITEA_TOKEN:-}" ]' 2>/dev/null; then
ok "WORKER_GITEA_TOKEN resolves in a login shell"
else
warn "WORKER_GITEA_TOKEN is EMPTY in a login shell."
warn "The daemon will start fine and workers will silently fail to open PRs."
warn "Fix \${SHARED_ENV}/tools/secrets.sh before relying on worker checkpoints."
fi
# Same trap, second variable (CB-591). A profile's `tokenEnv:` is resolved from the DAEMON's own
# process environment by HerdrPeerLauncher.resolveEnv, so a token added to secrets.sh after the
# daemon started is simply absent. The launcher then injects an empty token and llm.ltms.dev answers
# 401 — long after the restart, and with nothing tying the two together.
if zsh -lc '[ -n "${AI_GATEWAY_TOKEN:-}" ]' 2>/dev/null; then
ok "AI_GATEWAY_TOKEN resolves in a login shell"
else
warn "AI_GATEWAY_TOKEN is EMPTY in a login shell."
warn "Any profile whose tokenEnv is AI_GATEWAY_TOKEN will get an empty token and 401 at the gateway."
warn "This only matters once a profile points at llm.ltms.dev — harmless before that."
fi
if [ "$CHECK_ONLY" = 1 ]; then
say "--check: nothing changed"
exit 0
fi
# ---------------------------------------------------------------------- build
# Deliberately before the stop: a failed build must never leave the fleet down.
if [ "$DO_BUILD" = 1 ]; then
say "build"
BUILD_LOG="$(mktemp -t bridged-build)"
echo " log: $BUILD_LOG"
if ! mvn -f "$BRIDGED/pom.xml" clean install > "$BUILD_LOG" 2>&1; then
grep -E 'ERROR|BUILD FAILURE|Tests run:.*Failures: [1-9]|Tests run:.*Errors: [1-9]' "$BUILD_LOG" \
| head -20 || true
die "build failed — the running daemon was NOT touched. Full log: $BUILD_LOG"
fi
grep -E '^\[INFO\] Tests run:.*Failures' "$BUILD_LOG" | tail -1 | sed 's/^\[INFO\] / /' || true
ok "BUILD SUCCESS"
ok "jar now: $(jar_id)"
else
say "build skipped (--no-build)"
fi
[ -f "$JAR" ] || die "no jar at $JAR — run without --no-build"
# ----------------------------------------------------------------- drain gate
if [ -n "$OLD_PID" ] && [ "$ASSUME_YES" = 0 ]; then
say "drain check"
echo " A restart drops every in-flight ticket and rendezvous. A member's report"
echo " is NOT recoverable once its ticket is gone."
echo
echo " Confirm with bridge_list that no members are live, and bridge_poll anything"
echo " you still want, BEFORE continuing."
echo
read -r -p " Fleet drained? type yes to restart: " reply
[ "$reply" = "yes" ] || die "aborted — nothing changed"
fi
# ------------------------------------------------------------------ stop
if [ -n "$OLD_PID" ]; then
say "stop"
RESTART_MARK="$(wc -l < "$OUT" 2>/dev/null || echo 0)" # verify a FRESH line appears later
kill "$OLD_PID"
for _ in $(seq "$STOP_WAIT"); do
[ -z "$(running_pid)" ] && break
sleep 1
done
if [ -n "$(running_pid)" ]; then
die "pid $OLD_PID still alive after ${STOP_WAIT}s. Not escalating to kill -9 automatically:
the shutdown hook releases sessions and worktrees in order, and killing it hard can
leave worktrees and panes behind. Investigate, then kill -9 by hand if you accept that."
fi
ok "pid $OLD_PID exited"
else
RESTART_MARK="$(wc -l < "$OUT" 2>/dev/null || echo 0)"
fi
# ------------------------------------------------------------------ start
# Login shell (zsh -l) is what puts the secrets on the daemon's environment. cwd must be bridged/
# because the daemon resolves bridged.yaml, logs/ and target/ relative to it.
say "start"
# Absolute jar path so `ps` names which checkout is running; cwd still bridged/ because the daemon
# resolves bridged.yaml, logs/ and target/ relative to it.
( cd "$BRIDGED" && zsh -lc "nohup java -jar '$JAR' >> bridged.out 2>&1 &" )
for _ in $(seq 10); do
NEW_PID="$(running_pid)"
[ -n "$NEW_PID" ] && break
sleep 1
done
[ -n "${NEW_PID:-}" ] || die "no process appeared. Last lines of $OUT:
$(tail -20 "$OUT" 2>/dev/null)"
[ "$NEW_PID" != "${OLD_PID:-}" ] || die "pid unchanged ($NEW_PID) — the old daemon never died"
ok "started, pid $NEW_PID"
# ------------------------------------------------------------------ verify
say "verify"
HEALTH_BODY=""
for _ in $(seq "$HEALTH_WAIT"); do
if HEALTH_BODY="$(curl -fsS --max-time 2 "$HEALTH" 2>/dev/null)"; then break; fi
HEALTH_BODY=""
sleep 1
done
if [ -z "$HEALTH_BODY" ]; then
# 503 still means the daemon is up — it means herdr is unreachable. Say which.
CODE="$(curl -s -o /dev/null -w '%{http_code}' --max-time 2 "$HEALTH" 2>/dev/null || echo 000)"
if [ "$CODE" = "503" ]; then
warn "/healthz answers 503 degraded — the daemon is up but herdr is unreachable."
warn "Spawns will fail. Check herdr before delegating anything."
curl -s --max-time 2 "$HEALTH" 2>/dev/null | head -3 || true
else
die "/healthz never answered within ${HEALTH_WAIT}s (last code: $CODE). Last lines of $OUT:
$(tail -30 "$OUT" 2>/dev/null)"
fi
else
ok "/healthz 200 — $HEALTH_BODY"
warn "healthz green only proves herdr ANSWERS. If its protocol number changed, spawns can still"
warn "fail — prove a real spawn before trusting the fleet."
fi
# A fresh listening line, strictly after the restart mark. An old daemon that never died would
# otherwise let an old line pass for a new one.
if tail -n "+$((RESTART_MARK + 1))" "$OUT" 2>/dev/null | grep -q 'bridged listening'; then
ok "$(tail -n "+$((RESTART_MARK + 1))" "$OUT" | grep 'bridged listening' | tail -1)"
else
warn "no fresh 'bridged listening' line after the restart — check $OUT yourself"
fi
# Config keys the daemon accepted or deferred at boot. This is usually WHY you restarted.
say "config at boot"
tail -n "+$((RESTART_MARK + 1))" "$OUT" 2>/dev/null \
| grep -iE 'deferred|classification:|fleet health:|coverage' | tail -8 | sed 's/^/ /' \
|| echo " (nothing reported)"
# Errors since the restart, anchored to the marker so old noise cannot leak in.
ERRS="$(tail -n "+$((RESTART_MARK + 1))" "$OUT" 2>/dev/null | grep -cE ' (ERROR|SEVERE) ' || true)"
say "result"
ok "pid $NEW_PID, jar $(jar_id)"
if [ "${ERRS:-0}" -gt 0 ]; then
warn "$ERRS ERROR lines since restart:"
tail -n "+$((RESTART_MARK + 1))" "$OUT" | grep -E ' (ERROR|SEVERE) ' | tail -5 | sed 's/^/ /'
else
ok "no ERROR lines since restart"
fi
echo
echo " Next: call bridge_whoami and confirm it still answers 'primary'. A lead whose tab label"
echo " no longer matches fleet.leaders.*.tab is demoted to worker and refuses orchestration."
echo