Compare commits

..

195 Commits

Author SHA1 Message Date
Dai Ha 0a2b3a4a56 CB-597: fix inaccuracies the example config already had, none actually missing
CI / contract (pull_request) Successful in 1m7s
CI / build (pull_request) Failing after 1m21s
Audited bridged.example.yaml against BridgedConfig's KNOWN_TOP_LEVEL_KEYS and
found every top-level key already documented (broker, health, lifecycle,
configReload, quarantineCooldownSeconds, fleet.leaders/architects/reviewers
included) — CB-573/CB-566/CB-559/CB-579/CB-527/528 each updated the example
alongside their feature. What was actually wrong:

- health.workingSuspectAfterSeconds/paneProbeIntervalSeconds claimed enforced
  minimums (300/60) that don't exist in code — only intervalSeconds is
  clamped (floor 15); the other two are parsed but never read anywhere.
- notifications.mode: webhook was undocumented as only flipping the
  healthCoverage label bridge_list reports — no webhook is ever sent.
- lifecycle.clearAfterTurn was missing entirely.
- the HOT bullet under configReload claimed the whole fleet: block reloads
  live, but ConfigRef's own javadoc carves out fleet.leaders as needing a
  restart with no deferred-list warning — added that exception.
- fleet.leaders' demotion consequence (unmatched tab -> silent WORKER
  demotion, no startup error) is now stated inline next to the block, not
  just implied by the multi-lead rationale higher up.
2026-08-16 17:37:28 +02:00
ltms 8d4206c2b5 Merge CB-590: one nudge schedule per lead, with a reminder budget per source
CI / contract (push) Successful in 1m7s
CI / build (push) Failing after 1m15s
Carries PR #84 (its head `78ca24d` is an ancestor of this one), so this single merge delivers both rounds.

CB-590 collapses the CB-307 reply schedule and the CB-588 ticket schedule into one per lead, which closes the double-injection race. The review round found that this also collapsed the two reminder caps into one shared budget — so a busy reply stream could exhaust the cap and a ticket arriving afterwards would never be nudged at all. That was a regression, not a pre-existing wart: before this PR the two sources had independent counters.

The follow-up keeps the single schedule and gives each source its own budget. `decide()` returns INJECT while either source has pending work under its own cap, and STOP only when neither does. `stopOrRestart` and its snapshot-diff race logic are untouched.

Verified by the lead: `mvn -f bridged/pom.xml clean install` unpiped, exit code captured — Tests run: 814, Failures: 0, Errors: 0, Skipped: 0 — BUILD SUCCESS. `ReplyPushLoopTest`: 41 tests.

Known and deliberately out of scope: an item arriving during the 15s backoff is already inside the tick's "before" snapshot, so at-cap work can be abandoned rather than nudged once. That shape predates this PR in the ticket-only loop and is filed as #87 (CB-598).
2026-08-16 17:28:57 +02:00
ltms 03286a589b Merge CB-527/CB-528: bound AMQP prefetch, confirm publishes, and close the recovery race
CI / contract (push) Successful in 1m4s
CI / build (push) Failing after 1m18s
Carries PR #83 (its head is an ancestor of this one), so this single merge delivers both rounds.

CB-527 bounds prefetch. CB-528 makes a publish wait for its broker confirm, so a failed publish is never reported as success, and correlates a `mandatory` Return back to the right publish.

The review round on #83 found one real race: `failPendingPublishesOnRecovery` swept the pending maps without holding `publishChannelLock`, so a publish issued on the already-recovered channel could be failed by the sweep — the exact inversion of what CB-528 exists to prevent, and undetectable by dedup because `MessageService.reply` mints a fresh msgId per call. Fixed by taking the same lock. `close()` now fails in-flight publishes promptly instead of letting them time out after 10s.

Verified by the lead, not taken on the worker's word:
- `mvn -f bridged/pom.xml clean install` unpiped: Tests run: 809, Failures: 0, Errors: 0, Skipped: 0 — BUILD SUCCESS
- `mvn test -Pcontract -Dtest=AmqpReplyInboxContractTest` against a real broker: Tests run: 8 — BUILD SUCCESS

That second run mattered: #83's own new tests are contract-tagged, so the default suite and CI prove nothing about them.
2026-08-16 17:27:21 +02:00
Dai Ha 88b9503c3b CB-590 follow-up: give each nudge source its own reminder budget
CI / contract (pull_request) Successful in 43s
CI / build (pull_request) Successful in 1m46s
decide(lead, reminderCount) shared one counter across the reply and
ticket sources after PR #84 collapsed both onto a single per-lead
schedule. A reply stream that used up the whole budget could then
make decide() STOP even for a ticket that had never been nudged and
had coalesced onto the same still-active schedule — stranding it with
no live schedule left, since stopOrRestart's racedIn check does not
save work that was already present in the "before" snapshot.

decide() now tracks a per-source count (replyReminderCount,
ticketReminderCount) and returns INJECT while either source is still
under its own cap, STOP only when both are exhausted. Still exactly
one schedule per lead; stopOrRestart's snapshot-diff logic is
untouched.
2026-08-16 17:26:07 +02:00
Dai Ha c553d795d8 CB-528: close the recovery race in AmqpReplyInbox
CI / contract (pull_request) Successful in 44s
CI / build (pull_request) Failing after 1m20s
failPendingPublishesOnRecovery walked and cleared pendingBySeq/pendingByMsgId
without holding publishChannelLock, so a publish() that registered while the
sweep was still iterating could be failed even though it published on the
already-recovered channel — a successful publish reported as failed, and
since MessageService.reply mints a fresh msgId per retry, dedup can't catch
the resulting duplicate. Guard the sweep with publishChannelLock: publish()
only holds it for the seq/map-put/basicPublish, so the sweep can only ever
wait for an in-flight basicPublish to return, never a broker round trip.

Also make close() fail in-flight publishes immediately with a clear message
instead of leaving them to idle out the 10s confirm timeout, and record the
(currently unreachable) msgId-uniqueness assumption pendingByMsgId relies on.

AmqpReplyInboxRecoveryRaceTest drives the sweep and a real publish() against
each other directly (no live broker reconnect) using Proxy-backed fake AMQP
channels and a large in-flight backlog to make the race window observable;
confirmed it fails without the guard (reply "fresh" wrongly failed as
"connection recovered mid-publish") and passes with it.
2026-08-16 17:24:43 +02:00
Dai Ha 78ca24dc3f CB-590: collapse the CB-307 and CB-588 nudge schedules into one per lead
CI / build (pull_request) Successful in 1m0s
CI / contract (pull_request) Successful in 1m33s
Both reply-queued and ticket-terminal nudges could independently decide
to inject into the same lead pane in the same window, since they ran as
two separate schedules keyed differently (worker target vs. lead) that
never checked each other. Replace both with a single per-lead schedule
(activeLeads) that drains pending reply targets and pending tickets
together, sends at most one combined nudge per tick, and shares one
reminder cap across both sources — so two injections into the same pane
can no longer overlap, and work queued while the lead is busy is never
lost, only deferred.
2026-08-16 16:56:04 +02:00
Dai Ha 4fa6553db5 CB-527/CB-528: bound AMQP prefetch and confirm publishes before claiming durable
CI / build (pull_request) Successful in 1m3s
CI / contract (pull_request) Successful in 1m15s
CB-527: basicQos(prefetch) on the consume channel before basicConsume, configurable
via broker.prefetch (default 32), so an undrained inbox backlog stays on the broker
instead of growing the JVM heap without limit.

CB-528: publish moves to its own confirm-mode channel with mandatory=true and a
return listener, so an unroutable or unconfirmed reply now throws instead of
vanishing silently. The confirm callback checks the per-message returned flag
(set by the return listener, which the broker always fires before the matching
confirm) so an acked-but-returned publish is still reported as a failure. The ack
path stays on its own channel/lock and never waits on a publish confirm.
2026-08-16 16:55:40 +02:00
Dai Ha 2124e043ce CB-593: correct the member MCP claim — measured, not assumed
CI / contract (push) Successful in 43s
CI / build (push) Successful in 1m19s
CLAUDE.md told every member 'You mount only the bridge MCP' and told the lead
'a worker mounts only the bridge MCP and cannot run your other tooling'. Both
were false for Claude Code members.

Measured by spawning one member per backend and asking each what it actually has:

  opencode (gx)      11 bridge tools only                     claim TRUE
  claude-code (local) 11 bridge + 45 gitea + 2 context7       claim FALSE

Source is ~/.claude.json user-scope mcpServers; --mcp-config adds to that scope
rather than replacing it, so the worktree parity overlay (which correctly
neutralises .mcp.json and opencode.json) cannot see or stop it.

The forge tools are mounted but not usable: CB-592's blocked sentinel means
get_me and list_issues both fail with 'invalid username, password or token'.
That is defence in depth working in a path it was not designed for, so the text
now says a mounted tool is not a working tool rather than pretending the tools
are absent.

Block propagated byte-identically to wiki 7-Use-Cases.md (wiki 1d95e3f).
2026-08-16 09:46:42 +02:00
Dai Ha e4f3620acb CB-591: record the final 86400s route timeout and the request:0s trap
CI / contract (push) Successful in 44s
CI / build (push) Successful in 1m19s
systems/vms moved the LLM route timeout again, 1800s -> 86400s (24h), after
the silent-truncation risk was discussed. They tried request: 0s first: it
removes the total-duration timer, but on an AIGatewayRoute the idle timeout
is derived from the request timeout, so 0s also removed any bound on a
stalled connection.

At 86400s our own MessageService.ASYNC_TIMEOUT_MS (30 min) binds first, so a
runaway request now ends as a clean FAILED ticket we raised instead of a
silently truncated 200. While the gateway sat at 1800s the two numbers were
equal and did not nest.
2026-08-15 21:29:49 +02:00
Dai Ha 032a59a34d CB-591: fleet moved onto the gateway — both ceilings fixed and re-verified
CI / contract (push) Successful in 40s
CI / build (push) Successful in 1m19s
`local` runs on /anthropic and `gx` on /v1, both weight 100; `local-direct`
stays weight 0 as the escape hatch.

systems/vms fixed both blockers, and each was re-checked from this side rather
than taken on trust:

    listener buffer    32 KiB -> 32 Mi   ours: 1.2 MB body -> 200 (was 413)
    LLM route timeout  60s    -> 1800s   ours: 101s stream -> 200,
                                               message_stop present, 4000/4000

Neither was deliberate: 32 KiB was Envoy Gateway's default
per_connection_buffer_limit_bytes, and 60s was Envoy AI Gateway's own default.
The 60s bounded GENERATION as well as prompt size — a tiny prompt with a long
answer returned 504 at 60.05s.

Verified with real workloads, not liveness probes. A `local` member read this
document and CLAUDE.md in full — 48,344 bytes of file content, comfortably past
the old 32,768 ceiling — and answered four questions correctly, including the
document's length (said ~456, actual 455). A `gx` member did the same. The
trivial 3-question probe is what hid the 32 KiB ceiling for an afternoon, so it
no longer counts as proof here.

§7.2 is new and is the part that matters later. One risk is ACCEPTED, not
solved: on a mid-response timeout over chunked HTTP/1.1, Envoy ends the chunked
encoding cleanly instead of resetting, so a truncated answer arrives as HTTP 200
with no error and no terminator (envoyproxy/envoy#17186, acknowledged 2021,
never fixed; the Dec 2025 fix #42269 is HTTP/2 only and SSE here is HTTP/1.1).
Measured at the old 60s: 200, 61.07s, 2473 of 4000 emitted, message_stop 0,
error events 0, ending on a well-formed frame.

The recommended defence — reject a stream with no terminator — does NOT
transfer to us: Claude Code and opencode are third-party clients and we do not
own their SSE parsing. So this is acceptable because a request would have to run
1800s to trip it, not because we could detect it. If a member ever returns a
confident but truncated answer, suspect this before anything in our own code.

Also recorded, from the upstream bisection: ClientTrafficPolicy is honoured in
standalone `aigw run` but BackendTrafficPolicy is silently ignored, and nothing
external distinguishes them (envoyproxy/gateway#9513). Same silent-default shape
this repo keeps hitting.

bridged.yaml carries the same notes inline (gitignored, so not in this commit).

Refs: gitea #76
2026-08-15 20:47:41 +02:00
Dai Ha e689090024 CB-591: correct the root cause — Envoy's buffer limit, not Caddy
CI / contract (push) Successful in 1m7s
CI / build (push) Successful in 1m35s
I wrote "Caddy request_body max_size and/or Envoy's own" and marked it
unverified. The Caddy half was wrong, and an unverified guess still points the
next reader at the wrong component.

Confirmed by the systems/vms side: Envoy Gateway defaults a listener's
per_connection_buffer_limit_bytes to 32768, and aigw buffers the WHOLE request
body before it can route on the model name. So that default is not a network
tuning knob — it is a hard ceiling on prompt size. From the live config_dump:

    listener default/llm/http    per_connection_buffer_limit_bytes: 32768

Nobody chose 32 KiB; it was inherited.

Both TLS edges are innocent, and the technique that showed it is better than
mine: both 413s carry x-llm-consumer, a header their auth proxy sets only AFTER
authenticating, so the body cleared both edges and the auth. On llm.vm, aigw
413s at 39 KB while the vLLM backend answers 200 at the same size. I found the
boundary; they found the component, by reading the failure's response headers.

Consequences recorded in the doc:

  * DO NOT plan around 32 KiB. The intended ceiling is far higher, so sizing our
    profiles to it would be designing around a bug.
  * Their fix (ClientTrafficPolicy, bufferLimit: 8Mi) is written but NOT
    deployed, pending their operator's approval. We do not re-test until they
    confirm — a half-changed system gives a number neither side can trust.
  * In standalone `aigw run` a SecurityPolicy is accepted and then silently
    ignored, so "the config was accepted" proves nothing there. They will verify
    by re-reading the live config_dump and sending a large request. Same
    silent-default shape this repo keeps hitting, one layer down.

bridged.yaml carries the same correction (gitignored, so not in this commit).

Refs: gitea #76
2026-08-15 19:47:01 +02:00
Dai Ha 1cc34888fd CB-591: record the live result — blocked by a 32 KiB body limit at the gateway
CI / contract (push) Successful in 44s
CI / build (push) Successful in 55s
Deployed U1-U2c, restarted, spawned both new profiles for real, then reverted.

llm.ltms.dev answers HTTP 413 above 32 KiB (32768 bytes), on BOTH surfaces:

    /v1        32695 bytes -> 200        /anthropic  32095 bytes -> 200
    /v1        32795 bytes -> 413        /anthropic  32855 bytes -> 413

That is far below one agent turn. It is an edge limit (Caddy request_body
max_size, and/or Envoy), so the fix is in systems/vms, not here.

The part worth recording is how it nearly passed. Two members, same message,
same moment: `local` finished in 66s, `gx` never finished at all. `local`
passed only because the probe was three trivial questions in a fresh session,
so the request fit under 32 KiB — the profile looked healthy and was a
landmine set to fire on the first turn that reads a file. So §7's checklist
was not wrong, it was too easy; it now says to use a file-reading task.

opencode's failure mode is worse than a crash: it catches the 413, compacts
its context, retries, and loops. Observed 10+ minutes BUSY with no reply. From
the lead's side that is indistinguishable from a slow worker. Reproduced
outside the bridge with the launcher's own generated config, which is how it
became a one-line error instead of a hang; §7.1 records that procedure.

Everything else about the migration checked out and is recorded so it is not
re-tested: token accepted on both surfaces, unauthenticated 401 (the Caddy
proxy does gate, whatever the gateway's own fail-open policy does),
/v1/models exactly ["deepseek-v4-flash"], the guard allowlist accepted
llm.ltms.dev, and the generated opencode provider block is correct with a real
llmk- key.

Also answers §3b's open question: reasoning survives BOTH surfaces —
/anthropic returns a real "type":"thinking" block and /v1 returns a populated
reasoning_content. The feared /v1 translation loss did not happen.

Config state (bridged.yaml is gitignored, so it is described rather than
committed): `local` back on http://gx00.gw:8000, `gx` kept at weight 0,
`local-direct` kept, llm.ltms.dev left in the guard allowlist. The file
carries these numbers and the exact two-key edit to switch back.

Verified after the revert with a task that reads two large files: correct on
all three questions. Daemon pid 66745, jar f1fd659423e6.

Refs: gitea #76
2026-08-15 19:27:45 +02:00
Dai Ha 0331ecd5d3 CB-592: add the BRIDGED_MEMBER marker — the sentinel alone cannot hold
CI / contract (push) Successful in 1m5s
CI / build (push) Successful in 1m39s
Live check on a member pane showed the CB-592 shadow did NOT take effect:
GITEA_ACCESS_TOKEN inside the pane was still the real admin token.

Measured cause. The overlay itself works — GITEA_TOKEN is injected the same
way, is exported by no shell file, and does reach the pane. The sentinel loses
one step later. A herdr pane runs a LOGIN shell, ~/.zprofile line 41 sources
${SHARED_ENV}/tools/secrets.sh, and that file does a plain unconditional
`export GITEA_ACCESS_TOKEN=...`. A login shell overwrites a value already in
the environment, so the real token is put back before the member starts.
Confirmed directly:

    GITEA_ACCESS_TOKEN=cb592-sentinel zsh -lc ...
    -> RESULT: sentinel was OVERWRITTEN by the login shell

This defeats any launcher-side overlay for any name secrets.sh exports. No
change in this repo can win it alone.

So this adds the half that does survive: BRIDGED_MEMBER=1, a name secrets.sh
never exports. It is a no-op until the operator guards the export:

    [ -n "${BRIDGED_MEMBER:-}" ] || export GITEA_ACCESS_TOKEN=...

Setting it now costs nothing and makes that one line the whole remaining fix.
The sentinel stays: it is correct for any peer kind whose pane does not start
a login shell, and it keeps the intent explicit where every adapter passes.

Also corrects the javadoc and the test javadoc, which both claimed a
protection that was measured not to hold.

The other reported failure was my own bad test, not a regression. The probe
called /api/v1/user, which a minimal write:repository token cannot read. Same
token on the repo endpoint answers 200, so CB-302 is intact:

    GITEA_ACCESS_TOKEN: /user=200  /repos/lms/claude-bridge=200
    WORKER_GITEA_TOKEN: /user=403  /repos/lms/claude-bridge=200

Tests 805 -> 807. Both new tests proved to discriminate by reverting the
marker: everySpawnMarksThePaneAsAMember and
aProfileEnvEntryCannotClearTheMemberMarker both fail without it.

Refs: gitea #77
2026-08-15 18:37:52 +02:00
Dai Ha 831a918c30 Merge CB-592: shadow the admin GITEA_ACCESS_TOKEN in every member's environment
CI / contract (push) Successful in 41s
CI / build (push) Successful in 1m22s
A live probe showed every spawned member carried the admin GITEA_ACCESS_TOKEN: 108
environment variables in a member's pane against 99 in the primary's. The operator's
rule is that only the leader and architects may use it; everyone else uses
WORKER_GITEA_TOKEN. We were not enforcing that at all.

The cause is invisible from inside the launcher. baseEnv builds a fresh map holding only
PATH and the profile's env:, so a member looks like it gets a small explicit environment.
That map is an OVERLAY: WorkspaceControl.createTab/splitPane send only the keys it
contains, and herdr spawns the pane from its own login-shell environment, so every key we
never mention passes straight through — admin token included.

The fix puts a non-blank sentinel over the key in baseEnv, applied AFTER the profile's
env: so no profile, present or future, can restore the real token by naming it in config.
One place, every adapter, including peer kinds not yet written — deliberately not a
per-profile bridged.yaml entry, which is the silent-default shape this repo has shipped
nine times.

A non-blank sentinel rather than the empty string, on purpose: whether an empty overlay
value overrides an inherited variable or is skipped as blank cannot be settled from this
repo, because herdr's merge happens in an external process. baseEnv's own PATH seeding
(CB-511) already relies on a non-blank value replacing an inherited one, so this reuses
the shape that is demonstrated to work rather than the one that is merely plausible.

CB-302's repo-scoped GITEA_TOKEN grant is untouched — a worker can still open its own PR.
The subscription boundary was checked and is unaffected: the primary's pane carries no
ANTHROPIC_* at all, so nothing is inherited there.

Closes gitea #77. Live verification follows separately: the daemon must be redeployed
before this reaches any pane.
2026-08-15 18:27:05 +02:00
Dai Ha 3db5277ae8 CB-592: shadow the admin GITEA_ACCESS_TOKEN in every member's herdr overlay
CI / contract (pull_request) Successful in 1m12s
CI / build (pull_request) Successful in 1m37s
herdr spawns a pane from its own login-shell process env and layers our map on
top, so any key baseEnv never mentions passes straight through — including the
admin forge token. baseEnv now puts a non-blank sentinel for
GITEA_ACCESS_TOKEN, applied after the profile's own env: so no profile can
restore it. One place, every adapter, every profile including future ones.
CB-302's GITEA_TOKEN grant (applyGitToken) is untouched.
2026-08-15 18:25:16 +02:00
Dai Ha 6939e0cbbc CB-592: the tracked opencode.json must name the worker forge token, not the admin one
Operator's rule, 2026-08-15: only the leader and architects may use GITEA_ACCESS_TOKEN;
everyone else uses WORKER_GITEA_TOKEN.

opencode.json is TRACKED, so it ships in every worker worktree, and it mounted the gitea
MCP with {env:GITEA_ACCESS_TOKEN}. A live probe confirmed that variable actually resolves
inside a member: herdr spawns each pane from its own login-shell environment and layers
the launcher's map on top, so a member sees 108 variables rather than the small explicit
set baseEnv appears to build. That gave an opencode member admin forge TOOLS — enough to
merge its own PR, which both CLAUDE.md and the member contract forbid.

This is the narrow half of the fix: it removes the tooling. The admin token is still
present as a string in every member's environment, which is the real defect and is
tracked as CB-592 (gitea #77) — that fix belongs in the launcher, in one place, not
per-profile in bridged.yaml where a sixth profile would silently reopen it.

.mcp.json keeps GITEA_ACCESS_TOKEN and is correct to: it is skip-worktree, the primary's
own local copy, and the primary is the lead. That is the pattern this change follows —
the shared tracked file grants least privilege, and anything needing more overrides
locally.
2026-08-15 18:20:55 +02:00
Dai Ha f0095bf8b2 CB-591: plan the move onto the LLM/MCP gateway, and check AI_GATEWAY_TOKEN
CI / contract (push) Successful in 1m5s
CI / build (push) Successful in 1m38s
The gateway (llm.ltms.dev) replaced Bifrost on 2026-08-15 and serves an Anthropic
surface and an OpenAI surface, so both member kinds can point at it. The plan is in
docs/CB-591-Gateway-Migration.md; gitea #76 tracks the work.

The opencode half needs no code: OpenCodeLauncher already pins an OpenAI-compatible
endpoint (CB-508), so baseUrl + tokenEnv + provider/model is a config change. That
matters more than it looks — every opencode member today is sol or terra, and both sit
on one OpenAI account via credentialId: openai-shared, so an exhaustion on either locks
out both. A gateway-backed opencode profile is free and off that credential, which
retires a single point of failure rather than only adding capacity.

Also extends the redeploy script's --check to AI_GATEWAY_TOKEN. A profile's tokenEnv is
resolved from the DAEMON's own environment by HerdrPeerLauncher.resolveEnv, so a token
added to secrets.sh after the daemon started is simply absent: the launcher injects an
empty token and the gateway answers 401, long after the restart and with nothing tying
the two together. That is the same trap as WORKER_GITEA_TOKEN, and it gets the same
login-shell check that never prints the value.
2026-08-15 17:06:31 +02:00
Dai Ha 5206679efd Merge CB-588: nudge the lead when an async ticket goes terminal
An async delegation ticket (bridge_send wait:false) resolves on MessageService.reply's
rendezvous fast path, which returns before onReplyQueued. So CB-307's push loop only ever
heard about the durable-inbox case, and the mode CLAUDE.md tells leads to prefer never
nudged anyone. Closes gitea #72.

Adds a second, independent reminder schedule keyed by the lead terminal, so several
tickets finishing together coalesce into one nudge. The CB-307 path is untouched.

Three defects were found in review and fixed before merge:
 * a pendingTickets entry outlived the ticket it named. poll() returns null once
   pruneTerminalTickets drops a ticket, so ticketCollected was never reached and the
   entry leaked for the daemon's life, riding along on every later nudge and sending
   the lead after a ticket bridge_poll can no longer find.
 * a lost nudge: a ticket landing between decideTickets returning STOP and
   activeLeads.remove coalesced onto a schedule that was about to die. That is the
   exact failure this ticket exists to remove, reintroduced in a narrow window.
 * the success direction was unpinned in tests, and the comment listing the paths that
   complete the future was short by several.

The obvious fix for the second one was wrong: restarting on any pending ticket defeats
the reminder cap, because a never-collected ticket at cap is expected to still be there.
The fix diffs against a snapshot taken before the decision, so only a ticket that truly
arrived during the window restarts the schedule.

Verified on my own unpiped build: 802 tests, 0 failures, BUILD SUCCESS.
Two reviewers on the diff; the loop-gating finding they raised is split out as CB-590.
2026-08-15 17:06:12 +02:00
Dai Ha ac044e7573 CB-588 round 3: close the STOP-vs-onTicketTerminal race, pin nudge polarity, fix comment
CI / build (pull_request) Successful in 1m19s
CI / contract (pull_request) Successful in 1m31s
- ReplyPushLoop.stopOrRestartTicketLoop: after releasing a lead's active-schedule
  slot on STOP, restart only if a ticket landed that the pre-decision snapshot
  did not already account for. A naive "restart on any pending ticket" version
  was tried first and reverted: it defeated the reminder cap by restarting
  forever on a stale, never-collected ticket (broke
  successfulTicketNudgeIncrementsDelivered and ticketNudgesSendUpToCapThenStop).
  Diffing against a pendingBefore snapshot distinguishes a genuine race arrival
  from stale cap-exhausted backlog.
- Two new ReplyPushLoopTest cases exercise stopOrRestartTicketLoop directly
  (now package-private) rather than forcing the underlying thread race:
  aTicketStillPendingWhenTheLoopStopsIsNotStranded (the race must restart) and
  aStaleUncollectedTicketAtCapDoesNotRestartTheLoop (the cap must still hold).
  Both were verified to fail against deliberately-reverted versions of the fix
  before being restored to green.
- MessageServiceTest: pin the success-path nudge test's negative direction too
  (must not contain "FAILED"), not just the failure-path test.
- MessageService: complete the whenComplete comment's list of completion paths
  (TIMED_OUT/BUSY/BACKEND_EXHAUSTED via finishAsyncTask, completeExceptionally
  on throw, and answer() -> finishAsyncTask(turnId, result)).
2026-08-15 17:02:05 +02:00
Dai Ha 6ebad2a91f CB-588 follow-up: reclaim a pruned ticket's pendingTickets entry too
CI / contract (pull_request) Successful in 45s
CI / build (pull_request) Successful in 1m0s
tasks is the sole authority on whether a ticket exists, but
pruneTerminalTickets dropped entries from it without telling
ReplyPushLoop. ticketCollected(ticket) was only ever called from
MessageService.poll's terminal branch, which pruneTerminalTickets
short-circuits past once a ticket is gone (poll returns null at the
top). A ticket the lead never polled — or one the reminder cap already
gave up on — was pruned from tasks but never collected in
ReplyPushLoop.pendingTickets, so it rode along on every later nudge to
the same lead forever, naming a ticket bridge_poll could no longer
find, and the map itself never shrank.

pruneTerminalTickets now calls pushLoop.ticketCollected for every
ticket it actually removes (guarded on pushLoop != null), so a pending
nudge entry lives exactly as long as its ticket is pollable. Reused
tasks as the only removal trigger rather than adding a second live
query back into MessageService — no new source of truth.

Added an injectable clock (LongSupplier nowNanos, defaulting to
System::nanoTime) to MessageService, mirroring the SessionManager/
SessionReaper nowNanos seam, so a test can cross the 10-minute
TICKET_TTL_NANOS deterministically instead of sleeping for real.
Confirmed the new regression test fails against the prior
pruneTerminalTickets (a stale ticket rides along on a later coalesced
nudge) before restoring the fix.
2026-08-15 16:41:23 +02:00
Dai Ha 3d10ed385c CB-588: nudge the lead when an async ticket reaches a terminal phase
CI / contract (pull_request) Successful in 41s
CI / build (pull_request) Successful in 1m46s
An async bridge_send(wait:false) registers a rendezvous waiter, so its
reply always takes MessageService.reply's fast path and returns before
ReplyPushLoop.onReplyQueued is ever called — the exact mode the charter
tells leads to prefer never nudged.

Add ReplyPushLoop.onTicketTerminal(ticket, target, failed), a second
entry point reusing the loop's status gating, bounded/backoff reminders
and metrics, keyed by the nudge-receiving lead so several tickets
finishing together coalesce into one nudge naming the count. MessageService
wires it via task.future.whenComplete in sendAsync (covers reply,
completion fallback, wedge, and abandon() alike) and calls the new
ticketCollected(ticket) from poll() once a terminal view is handed back,
so an already-collected ticket is never nudged again. The CB-307 inbox
path (onReplyQueued/decide/NUDGE_FORMAT) is untouched.
2026-08-15 16:10:22 +02:00
Dai Ha 6d0c94dbdb Correct enforceMaxLoad's comment after CB-585
CI / contract (push) Successful in 42s
CI / build (push) Successful in 1m16s
The comment said non-positive means unlimited at load. That stopped being
true when CB-585 made an explicit maxLoad: 0 survive as a real cap of zero
and made a negative value refuse config load. The code below it was already
right — only the comment described the old normalisation. Flagged by the
CB-585 worker, which correctly stayed out of a file not on its list.
2026-08-15 16:02:35 +02:00
Dai Ha e01563a550 Merge cb584-session-resume-b62247-4
CI / build (push) Successful in 1m20s
CI / contract (push) Successful in 1m44s
2026-08-15 15:58:29 +02:00
Dai Ha 30e3225a3c Merge cb585-maxload-zero-4585c4-3 2026-08-15 15:58:29 +02:00
Dai Ha ef186a1516 Merge cb587-snapshot-index-flags-24e236-2 2026-08-15 15:58:29 +02:00
Dai Ha 5d5b3bdc76 CB-584: persist agentSessionId at acquire and expose it via bridge_spawn/bridge_list
CI / contract (pull_request) Successful in 1m5s
CI / build (pull_request) Successful in 2m18s
Wires up the two links that made session resume unreachable: MemberSession
now records agentSessionId from PeerHandle at acquire (both the plain and
worktree paths), and it survives onto the roster (rosterView, so both
bridge_list and GET /members show it). bridge_spawn accepts sessionName and
resumeSessionId, threading them into the SpawnRequest fields that already
existed but were never reachable from the MCP surface.

A resumeSessionId now requires an explicit profile (a resumed conversation
is tied to the specific backend that started it, so an unqualified spawn
routed by placement has no safe candidate to check) and is refused, naming
Capability.SESSION_RESUME, when that profile's adapter does not declare it
— PeerLauncher gains capabilitiesFor(profileName) so a mixed fleet is
checked per-adapter rather than against the fleet-wide capability union.

PeerHandle.agentSessionId() loses its default, the same fix CB-571 (c00a86b)
made for charterReceipt() one method above it — both existing adapters
already overrode it, so this only closes the landmine for a future one.

Out of scope, left for follow-up: carrying agentSessionId on a failed
ticket's ReleaseDetail (issue #65 criterion 5) — CB-578 stage C just
landed in that file and this ticket deliberately stayed out of it.
2026-08-15 15:48:06 +02:00
Dai Ha 2f8c98dac9 CB-585: maxLoad: 0 caps a profile at zero, negative refused at load
CI / contract (pull_request) Successful in 41s
CI / build (pull_request) Successful in 1m18s
2026-08-15 15:38:50 +02:00
Dai Ha 2db7189067 CB-587: seed the snapshot's temp index from the worktree's real index
CI / build (pull_request) Successful in 1m1s
CI / contract (pull_request) Successful in 1m24s
--skip-worktree is an index flag. GitWorktrees.snapshot staged into a fresh
empty temp index, which carried none of the real index's skip-worktree bits,
so add -A staged local on-disk content for files git status correctly hides
(e.g. .mcp.json). Copy the real index (resolved via git rev-parse --git-path
index, correct for linked worktrees) into the temp index before staging, so
add -A skips exactly what git status skips.
2026-08-15 15:36:06 +02:00
Dai Ha 94476ac109 CB-529: drainReplies javadoc said the ack is local — false for AMQP
CI / build (push) Successful in 55s
CI / contract (push) Successful in 1m5s
The doc claimed an in-flight failure re-surfaces the messages on a later
drain, because the ack is local. That was true when InMemoryReplyInbox was
the only inbox, and became false without anyone noticing when the AMQP
adapter landed: there the ack is a broker-side basicAck, so a crash while
writing the response loses the reply outright. Re-polling cannot recover it,
since the broker has already forgotten it.

Behaviour is unchanged and the window stays accepted — acking on the next
poll instead would double-deliver on every normal drain. The point is that
the comment sat exactly where the next implementer would read it and said
the opposite of what happens.

Closes #12.
2026-08-15 15:34:59 +02:00
Dai Ha 5fe02b7c98 Fix redeploy script reporting failure on a successful restart
CI / contract (push) Successful in 45s
CI / build (push) Successful in 59s
Found by running it. The script launched the daemon with a relative jar path
(cwd is bridged/) but detected it with an absolute one, so pgrep never matched.
The daemon restarted correctly and booted clean, and the script still failed
with 'no process appeared' — the worst shape of bug for a deploy tool, because
it invites a second restart on a daemon that is already healthy.

Detection now matches both path forms, and the launch uses the absolute path so
ps names which checkout is running.
2026-08-15 15:18:44 +02:00
Dai Ha a1052f4fd1 Add scripts/redeploy-bridged.sh so the lead can deploy in one command
CI / build (push) Successful in 57s
CI / contract (push) Successful in 1m23s
The lead already owned the redeploy, but the command classifier refuses a bare
kill on the daemon, so in practice every deploy still needed the operator to
approve a stop and a start by hand. A single script is the seam that fixes
that: the operator allow-lists one auditable command instead of two ad-hoc
ones.

It also stops the procedure from living only in a checklist people read after
things go wrong. It builds before it stops anything, so a failed build never
leaves the fleet down; waits for the old process to exit instead of assuming;
polls /healthz; and anchors its log checks to a line marker taken before the
restart, so old errors cannot be misread as new ones.

The check with no log line anywhere in the daemon is the reason --check exists:
bridged inherits WORKER_GITEA_TOKEN from the shell that starts it, and starting
from a non-login shell empties it. The daemon boots fine, healthz is green, and
the failure only appears later as workers that cannot open a PR. --check tests
whether the name resolves and never prints the value.
2026-08-15 15:14:21 +02:00
Dai Ha 8c9904a7c4 Record that the classifier still blocks the daemon kill
CI / contract (push) Successful in 45s
CI / build (push) Successful in 59s
A CLAUDE.md rule grants intent, not tool permission. Tried the redeploy right
after writing the section and the classifier refused `kill`, so the section
would have been misleading as written. States the honest split until a Bash
permission rule exists: the lead builds and verifies, the operator runs the
stop and start with `!`.
2026-08-15 13:57:18 +02:00
Dai Ha 23af5dfdff The lead may redeploy the daemon — write down the procedure and its traps
A merge is not a deployment: the running bridged holds the jar it started
with, so merged code does nothing until the daemon is rebuilt and restarted.
Calling that work shipped is a false report. This makes the redeploy the
lead's job rather than something handed back to the operator, and records the
five things that have gone wrong doing it here — chiefly that starting the
daemon from a non-login shell empties WORKER_GITEA_TOKEN, which nothing logs
and which only surfaces later as workers that cannot open a PR.

Goes in the project addendum, not the canonical block; the block is unchanged
and still byte-identical with the wiki template.
2026-08-15 13:54:38 +02:00
Dai Ha 3a10f6ad17 Note that stage C's snapshot makes the parityOverlay gitignore rule load-bearing
CI / build (push) Successful in 56s
CI / contract (push) Successful in 1m19s
CB-578 stage C commits a preserved dirty worktree to refs/wip/<branch> with
git add -A. The parity overlay copies the primary's environment files into
every worktree, so a non-gitignored overlay path now reaches a durable git
object instead of only sitting on disk. add -A respects .gitignore, which is
what stops it — so the rule documented for CB-581 is now what keeps a secret
out of a commit, not just out of a directory.
2026-08-15 13:09:01 +02:00
Dai Ha a3842c873d Merge cb578c-92885c-1 2026-08-15 13:07:50 +02:00
Dai Ha c29c3f063d Merge cb554-e8e224-3 2026-08-15 13:07:50 +02:00
Dai Ha 758d62a396 Merge cb583-b82d11-2 2026-08-15 13:07:50 +02:00
Dai Ha 6725642274 CB-578 stage C: snapshot a dirty worktree into refs/wip before it can be lost
CI / build (pull_request) Successful in 54s
CI / contract (pull_request) Successful in 1m21s
Adds Worktrees.snapshot(worktreePath, branch, message): stages into a
temporary GIT_INDEX_FILE (never the worker's real index/HEAD), writes
the tree, commit-trees it onto the worktree's current HEAD, and points
refs/wip/<branch> at the result. add -A (never -f) respects .gitignore.

SessionManager.release() now snapshots any dirty worktree before the
preserve-or-remove decision, regardless of release cause (COMPLETED or
SHUTDOWN) — preserving on disk alone is one `worktree remove --force`
away from gone. A failing snapshot never escalates: the worktree is
still preserved, the pane still stops, and the release listener is
still notified.

onRelease's listener now receives a ReleaseDetail (terminal, worktree
path, branch, snapshot ref) instead of a bare terminal id, so Bridged's
WORKER_FAILED wiring can put the same three facts into a failed
ticket's detail — a lead can re-dispatch onto the same tree instead of
starting from the base commit.
2026-08-15 12:56:12 +02:00
Dai Ha a1f4dc365a CB-554: weight <= 0 excludes a profile from automatic placement, not 1.0
CI / contract (pull_request) Successful in 44s
CI / build (pull_request) Successful in 52s
2026-08-15 12:54:02 +02:00
Dai Ha 165b62ee20 CB-583: make bridge_list's capacity view quarantine-aware
CI / contract (pull_request) Successful in 44s
CI / build (pull_request) Successful in 1m32s
A quarantined profile's capacity row now forces free:0 and names the
quarantine (credentialId, quarantinedForSeconds), reusing the same
QuarantineSource bridge_profiles already reads instead of a second
lookup. An ordinary fleet's capacity rows are unchanged (no new keys).
2026-08-15 12:51:36 +02:00
Dai Ha 72d3481de3 Merge CB-578 stage B: quarantine the exhausted credential, not the profile
CI / contract (push) Successful in 41s
CI / build (push) Successful in 1m32s
Closes issue #50. Verified by the lead: own unpiped build of the branch merged
onto main — 740 tests, BUILD SUCCESS, exit 0.

Stage A only classified a usage-limit refusal. Nothing acted on it, so the
fleet would spawn another member onto the same exhausted account and fail the
same way. Now a BACKEND_EXHAUSTED classification quarantines the CREDENTIAL
for a cooldown: an explicit spawn onto it is refused with the profile, the
credential and the seconds remaining, placement skips it under every policy,
and it lifts itself on an injected clock.

Keyed by credential, not profile name, because sol and terra are two models on
one OpenAI account. Quarantining only the profile that reported the refusal
would leave its sibling live, and the next spawn walks onto the same dead
account. A profile that sets no credentialId quarantines alone under its own
name, so an existing config behaves exactly as before.

The implementer also found a real bug outside the brief: ConfigRef's
sameLaunchSettings never compared exhaustedPattern, so a reload changing only
that key reported 'config reloaded' while the value is actually deferred —
the precise failure that file's own doc calls the worst outcome a reload can
produce. Fixed with a regression test.

Two things I changed at review. The example.yaml conflict with e09cac6 was
resolved toward the branch, which was a strict superset. And
BackendQuarantine.none() held a clock frozen at 0, so a quarantine() call on
it would have locked a credential out for the life of the daemon — two
production CompositePeerLauncher constructors default to none(). It is now a
real no-op with a test.

Left open deliberately: the implementer's bridge_ask about the credentialId
name timed out after 55s with no answer, because I was polling minutes apart.
It proceeded on its own judgment and chose well. The 55s ask window against an
async lead is a bridge problem, not a worker problem.
2026-08-15 10:33:59 +02:00
Dai Ha 29ccb747ad CB-578 stage B: make BackendQuarantine.none() actually inert
Added at merge review. none() held a clock frozen at 0 with a 1ns cooldown, so
a quarantine() call on it recorded a deadline that could never pass — the
credential would be locked out for the life of the daemon. Two production
CompositePeerLauncher constructors default to none(), so that failure would
have been silent and permanent.

The implementer documented the limitation honestly rather than hiding it, but
a stand-in named none() should not need the caveat. quarantine() is now a
no-op on that instance, with a test asserting it. An inert value must omit the
fact, never invent one.
2026-08-15 10:33:44 +02:00
Dai Ha 4749e27453 Merge remote-tracking branch 'origin/main' into HEAD
# Conflicts:
#	bridged/bridged.example.yaml
2026-08-15 10:31:41 +02:00
Dai Ha e501d39988 CB-578 stage B: quarantine the exhausted credential, not the profile
CI / build (pull_request) Successful in 55s
CI / contract (pull_request) Successful in 1m21s
A BACKEND_EXHAUSTED classification (stage A) now puts that profile's
credential into a BackendQuarantine for a configurable cooldown. A spawn
onto a quarantined profile is refused naming the credential and roughly
when it lifts; weighted/round-robin/fixed placement skip a quarantined
candidate; the quarantine lifts itself on the injected clock; and it is
visible on bridge_profiles.

Keyed by credential, not by profile name, via the new Profile.credentialId
(profiles sharing one credential quarantine together — e.g. two models on
one account) and effectiveCredentialId() (unset ⇒ quarantines alone,
today's behaviour unchanged). Fixed a related gap along the way: a reload
changing exhaustedPattern was silently reported "applied" even though it's
deferred — sameLaunchSettings() now catches it too.
2026-08-15 10:29:54 +02:00
Dai Ha 6b6cf25862 CB-581: neutralise the parityOverlay false-preserve risk, and record the rule
CI / contract (push) Successful in 1m5s
CI / build (push) Successful in 1m32s
Issue #57 criterion 5 asked for an explicit decision on this rather than
silence. It is closer to live than the issue assumed.

The DEFAULT parityOverlay is List.of(".env", ".envrc"), so it applies to every
profile, and neither path was gitignored. Since CB-576 a release preserves any
worktree that git status --porcelain calls dirty, and that deliberately counts
untracked files — the work lost in CB-576 was a file nobody had added. So one
.env at the repo root would make every COMPLETED release preserve its
worktree, and worktrees would accumulate with no error to notice.

Inert today only because neither file exists here. Both are now gitignored,
which they deserve on their own as environment files. The javadoc carries the
rule for the next overlay path: it must be gitignored, or tracked and
skip-worktree'd.
2026-08-15 10:17:26 +02:00
Dai Ha 180de840eb Merge CB-581: a throw in release() no longer orphans the pane or aborts reapIdle
CI / contract (push) Successful in 43s
CI / build (push) Successful in 52s
Verified by the lead: own build of the branch merged onto main — 717 tests,
BUILD SUCCESS, exit 0.

release() ran an unprotected sequence. hasUncommitted shells out to git status
and throws on a non-zero exit, which skipped both notifyReleased and
launcher.stop. The session was already out of the registry, so nothing retried
it: a live pane kept burning a fleet slot while absent from the roster, and any
send blocked on it was never resolved. reapIdle called release bare inside a
loop, so one such session aborted the whole pass and skipped every session
after it.

Now the dirty-check block is guarded and fails toward preserving the worktree —
'we could not tell' must not be treated as 'it is clean', because deleting on a
guess destroys work with no other copy. notifyReleased runs in a finally and
launcher.stop runs unconditionally, so the pane always stops. reapIdle catches
per session, matching the shape drainAll already used.

Checked while reviewing: notifyReleased cannot throw out of the finally — it
already guards each listener and only logs. The pane stop is genuinely
unconditional.
2026-08-15 10:14:41 +02:00
Dai Ha 8fd2d7e5e7 CB-581: fail-safe release() so a throw never orphans the pane or aborts reapIdle
CI / build (pull_request) Successful in 52s
CI / contract (pull_request) Successful in 1m1s
hasUncommitted shells out to git and can throw; release() now catches that
inside a try/finally so notifyReleased and launcher.stop always run, and
defaults to preserving the worktree on a throw (can't tell dirty vs clean,
so don't risk deleting unrecoverable work). reapIdle wraps each per-session
release in try/catch, matching drainAll, so one bad session no longer
skips the rest of the reaping pass.
2026-08-15 10:09:32 +02:00
Dai Ha 129dd4a838 M4: record what Unit 2 has actually landed, and correct criterion 15
CI / build (push) Successful in 58s
CI / contract (push) Successful in 1m16s
Unit 2 was written as one block, but CB-571/576/580 and unit 2a have since
built parts of it. A symbol survey of main source plus the merge history puts
it at 1 of 21 criteria done, with three partials.

Criterion 15 says a normal COMPLETED release removes the worktree. Since
CB-576 that is false on purpose — a dirty worktree is preserved, because
deleting it destroys work that cannot be recovered. The criterion is wrong,
not the code.
2026-08-15 10:08:38 +02:00
Dai Ha e09cac6f1f CB-578 stage A: document exhaustedPattern as a deferred reload key
CI / contract (push) Successful in 44s
CI / build (push) Successful in 1m31s
The patterns are compiled once in Bridged.main from the startup config
snapshot, so adding one to a profile does nothing until a restart. Neither
the per-key docs nor the reload-class table said so.
2026-08-15 10:06:41 +02:00
Dai Ha c00a86b32c Merge CB-571: charter receipt on the roster, with no silent null for OpenCode
CI / contract (push) Successful in 42s
CI / build (push) Successful in 1m19s
Verified by the lead: own build of this branch merged onto main — 710 tests,
BUILD SUCCESS, exit 0.

Supersedes PR #54. A reviewer found that OpenCodeLauncher.SessionAwareHandle
wrapped the base WorkerHandle but never overrode charterReceipt(), so it
inherited the interface default of null while the real receipt sat on its
delegate — sol and terra would have shown no charterSource/charterSha256 on
the roster while Claude Code members showed both.

The fix is the root one, not the one-line override: PeerHandle.charterReceipt()
is no longer a default, so the compiler forces every implementation to answer.
This repo had shipped that same class of defect — a defaulted dependency that
compiles, passes tests, and quietly turns a feature off — eight times before
this one.
2026-08-15 10:02:36 +02:00
Dai Ha d3ae0350a2 Merge remote-tracking branch 'origin/main' into fix-charter 2026-08-15 10:01:10 +02:00
Dai Ha 2c2196a1f1 Merge CB-578 stage A: classify a usage-limit refusal instead of a completed reply
CI / build (push) Successful in 55s
CI / contract (push) Successful in 1m24s
Verified by the lead: own build of the branch merged onto main — 704 tests,
BUILD SUCCESS, exit 0. No vendor wording in any Java source (grep clean); a
profile with no exhaustedPattern keeps today's completion-fallback path exactly.

Accepted the implementer's deviation from the brief. The brief asked for a
terminal HealthState BACKEND_EXHAUSTED. FleetHealth.decide() is a pure
classifier over HealthSnapshot, which carries only booleans, AgentStatus and
MemberSession.State — never pane text. FleetHealth's own javadoc already says
ERROR_ON_SCREEN 'is not decided yet because it needs a bounded pane detection
read'. A second declared-but-unproduced value would repeat a gap the class
already documents as a problem, so the signal was put where the evidence
actually lives: the completion scrape.
2026-08-15 10:00:34 +02:00
Dai Ha 35ade14630 CB-571: make PeerHandle.charterReceipt() abstract, fix OpenCode adapter's silent null
CI / build (pull_request) Successful in 50s
CI / contract (pull_request) Successful in 1m21s
SessionAwareHandle wrapped the base's WorkerHandle but never overrode
charterReceipt(), so it silently inherited the interface default (null)
while the real receipt sat on its delegate. sol/terra never got a
charterSource/charterSha256 roster row.

Deletes the default so every PeerHandle must answer explicitly; the
compiler now catches this class of gap instead of a roster field
quietly going missing.
2026-08-15 09:55:52 +02:00
Dai Ha 541df87272 CB-578 stage A: classify a usage-limit refusal instead of a completed reply
CI / build (pull_request) Failing after 59s
CI / contract (pull_request) Successful in 1m10s
A backend that refuses on a subscription usage limit leaves the pane healthy but
the turn ends with no bridge_reply; the completion fallback used to scrape and
hand that refusal back as if it were a real answer. CompletionResolver now
matches the scrape against a per-profile exhaustedPattern (config, never a
vendor string) and resolves the send as Rendezvous.Kind/Outcome.BACKEND_EXHAUSTED
with a reason carrying the matched line, kept distinct from GONE/WORKER_FAILED.
A profile with no pattern configured is unaffected. Coverage is logged at
startup via CompletionResolver.coverage(...), naming which profiles have a
pattern and which don't, following FleetHealthMonitor.coverage's pattern.
2026-08-15 09:55:06 +02:00
Dai Ha 337b6ccd6e Merge remote-tracking branch 'origin/main' into fix-charter
# Conflicts:
#	bridged/src/test/java/dev/ltms/bridged/session/SessionManagerTest.java
2026-08-15 09:52:39 +02:00
ltms 42f46dfe9a Merge CB-579: resolve a lead by its tab name, drop the terminal-id pin
CI / build (push) Successful in 58s
CI / contract (push) Successful in 1m4s
Verified by the lead: merged onto main (9088d2b) in a scratch worktree, mvn -f bridged/pom.xml
clean install unpiped — MVN_EXIT=0, Tests run: 696, Failures: 0, BUILD SUCCESS. main alone measures
692, so this adds 4 net tests. Merges cleanly; Bridged.java auto-merged against CB-580.

Reviewed by the lead reading the full production diff and the three test files. The member's own
report was lost to the idle reaper before collection, so there was no author write-up.

Closes the live bug: LeadTabScanner.scan() no longer merges the config pin over the scan result, and
the cache no longer seeds from it, so a lead disappears once its tab is gone. The ghost this fixes
had begun throwing agent_not_found from ReplyPushLoop.decide on every tick, against a dead terminal
that still owned two live members.

Beyond the brief, and correct: `tab` is required for every leader, not only non-creatable ones,
because LeadLauncher also uses it to label a tab it creates. terminal: is rejected by a raw-YAML
check rather than by record shape — the only way to beat @JsonIgnoreProperties(ignoreUnknown = true).
The silent-default trap is avoided: the back-compat constructor still takes `tab` positionally.

Behaviour change worth knowing: the scanner's initial cache is now empty instead of the config pins,
so a herdr failure on the very first scan yields no leads until a scan succeeds. That is unavoidable
once the pins are gone, it fails loudly rather than silently, and it is covered by
aFailedFirstScanReturnsEmptyRatherThanThrowing.

Operator action required: bridged.yaml must replace fleet.leaders.<name>.terminal with tab. Already
done for this deployment.
2026-08-15 09:29:00 +02:00
ltms 9088d2b2c5 Merge CB-580: fail a ticket when its member reaches a terminal health state
CI / contract (push) Successful in 41s
CI / build (push) Successful in 1m32s
Verified by the lead on 0af902e: mvn -f bridged/pom.xml clean install, unpiped, in a scratch
worktree — MVN_EXIT=0, Tests run: 689, Failures: 0, Errors: 0, BUILD SUCCESS.

Reviewed by the lead reading the diff. The member's own report was lost to the idle reaper before
it was collected, so there was no author write-up to review against.

All three defects that sank 3b2f395 are absent:
  * failTarget is required — one constructor, Objects.requireNonNull, no defaulting overload
    anywhere in the repo, and Bridged.java:365 updated to pass messages::abandon.
  * The failure fires inside reportTransition after its `if (previous == next) return;` guard, so an
    unchanged tick cannot reach it.
  * No AtomicReference; MessageService is passed directly, so there is no empty window.

abandon(String target, String reason) confirmed as CB-568's target-wide operation that resolves
waiters as a failure rather than letting them time out.

Known limitation, accepted and covered by its own test: states.put records the new state before the
bounded retries run, so if all three attempts throw, the tickets stay pending and no later tick
retries. It is logged at WARN, and the retries carry no backoff.
2026-08-15 08:57:03 +02:00
ltms 500bfa2c33 Merge CB-576: release preserves a dirty worktree instead of deleting it
CI / build (push) Successful in 51s
CI / contract (push) Successful in 1m1s
Verified by the lead on b525b0f: mvn -f bridged/pom.xml clean install, unpiped, in a scratch
worktree — MVN_EXIT=0, Tests run: 686, Failures: 0, Errors: 0, BUILD SUCCESS.

Review accepted the required-interface-method shape (no defaulting overload) and the decision to
count untracked files as dirty — the work lost in the incident was a file that was never added.

One blocking defect was found and fixed in b525b0f: hasUncommitted called git with no existence
check, so a missing worktree threw WorktreeException from inside release() after registry.remove()
but before notifyReleased() and launcher.stop(), orphaning the pane and stranding a blocked
bridge_send caller. It now mirrors remove()'s already-gone tolerance.

Waived on merge, tracked as follow-up: a SessionManager-level test that teardown completes when the
worktree is gone, and the stronger fix behind it — reapIdle calls release() with no try/catch while
drainAll wraps it, so any exception in that window aborts the whole reaping pass.
2026-08-15 08:54:33 +02:00
Dai Ha 9ca9c43dfa CB-571: retag charter-receipt references from the taken CB-575
CI / contract (pull_request) Successful in 43s
CI / build (pull_request) Successful in 54s
CB-575 already names the merged MCP-cancellation-filter change, so the
charter-receipt comments used the wrong number. Retag to CB-571, the number
this work was authored against.
2026-08-15 08:48:57 +02:00
Dai Ha b525b0f08f CB-576: hasUncommitted tolerates an already-gone worktree
CI / contract (pull_request) Successful in 1m8s
CI / build (pull_request) Successful in 1m37s
2026-08-15 08:47:12 +02:00
Dai Ha 1966c69994 CB-575: charter receipt on spawn, in the roster and in the logs
CI / contract (pull_request) Successful in 46s
CI / build (pull_request) Successful in 1m34s
Record a CharterReceipt (role, source, sha-256 digest, byte count) for every
launch, store it on the MemberSession, expose it in bridge_list and GET
/members, and log it at spawn as digest+role only. Redact the charter argv
argument in the legacy pane-placement spawn log so the charter text never
reaches the daemon log. The charter prose itself is never recorded.
2026-08-15 08:10:56 +02:00
Dai Ha 976eff8ad1 CB-579: resolve a lead by its tab name, drop the terminal-id pin
CI / build (pull_request) Successful in 52s
CI / contract (pull_request) Successful in 1m18s
Leader.terminal -> Leader.tab (exact tab label, case-insensitive match).
LeadTabScanner matches an exact tab->name map instead of stripping a
shared tabPrefix, and no longer merges configured leads into every
scan result -- a stale pin can no longer outlive its tab.
LeadLauncher.tabLabel() returns the configured tab directly; the
terminalId pinned-terminal fallback in liveLeads() is gone.
Config load now rejects a leftover fleet.leaders.*.terminal key
instead of silently ignoring it. primary.terminal is untouched.
2026-08-15 08:05:21 +02:00
Dai Ha 0af902ec43 CB-580: fail a ticket when its member reaches a terminal health state
CI / build (pull_request) Successful in 54s
CI / contract (pull_request) Successful in 1m5s
FleetHealthMonitor now requires a failTarget BiConsumer<String,String>
collaborator (no defaulting overload) and calls it exactly once when a
member transitions into GONE or NEVER_READY, via CB-568's idempotent
target-wide abandon() operation. The reason string names the real
terminal state. failTarget invocation retries up to
MAX_FAIL_TARGET_ATTEMPTS (3) within the same transition if it throws,
and never refires on a later tick where the state is unchanged.

Bridged.java wires messages::abandon as the production failTarget.
2026-08-15 07:54:28 +02:00
Dai Ha 9118ce2537 CB-576: release preserves a dirty worktree instead of deleting it
CI / contract (pull_request) Successful in 1m2s
CI / build (pull_request) Successful in 1m35s
2026-08-15 07:34:06 +02:00
Dai Ha 2f48e08f1f Merge CB-577 follow-up: drop the target-keyed async index
CI / contract (push) Successful in 43s
CI / build (push) Successful in 55s
asyncTasksByWaiter correlates an async question by the exact rendezvous
waiter, so the target-keyed set it replaced can no longer decide
anything. Keeping it meant two indexes of the same fact, one of them
ambiguous whenever a target has two accepted tickets.

The race the old test modelled by reflection is gone with it: identity
keys make 'some other task reached this target' unrepresentable, so
there is no longer a wrong task for the question to land on.

Also carries the criterion-1 doc correction, which is identical to
aac29d6 on a different parent.
2026-08-15 06:40:31 +02:00
Dai Ha fec284e7cb Merge M4 unit 2a: accepted-turn identity carried to delivery
CI / build (push) Successful in 59s
CI / contract (push) Successful in 1m17s
TurnToken, owned by MessageService, binds a target to the exact
rendezvous waiter for one accepted send. Injector.Pending carries it and
the delivery callback hands it to CompletionResolver, so the baseline is
bound to the send it belongs to by construction rather than by a lookup
that could pick a different one.

The callback signature is required, not a defaulted overload: a delivery
with no token is exactly the unbound baseline this unit forbids, so a
default would let a caller silently produce it.

The token deliberately omits the session turn number. MessageService
owns acceptance but never learns of delivery, and
CompletionResolver.onDelivered runs before SessionManager.onDelivered,
so the number does not exist yet at the only point the token could
capture it. docs/M4-Fleet-Health.md criterion 1 records this and the two
rejected alternatives.

Still open for the next slice: the missing/post-restart baseline test and
the no-replay test.
2026-08-15 06:36:46 +02:00
Dai Ha 8e2e4c5e73 M4 unit 2a: migrate test call sites to the required turn token
The delivery callback now requires a TurnToken, so 54 test call sites
had to pass one. They use an explicit TestTurnTokens.inert(target)
rather than a defaulted overload, because a delivery with no token is
the unbound baseline this unit forbids.

The first version of inert() returned a fresh CompletableFuture as the
waiter, which turned captureBaselineSkipsTheReadWhenNoSendIsWaiting red:
the resolver saw a non-null waiter, concluded a turn was in flight, and
scraped a pane no send was blocked on. An inert value must omit the
fact, not invent it, so the waiter is now null and the production skip
fires as designed.
2026-08-15 06:36:35 +02:00
Dai Ha 0edc6615fc M4: correct unit 2 criterion 1 — TurnToken cannot carry the session turn
The criterion required the token to bind the session turn number. Three
independent refusals from the implementer showed why that is not
implementable at this layer: MessageService owns acceptance but never
learns of delivery, and CompletionResolver.onDelivered runs before
SessionManager.onDelivered, so the turn number does not exist yet at the
only point the token could capture it.

Records both rejected alternatives and why, so the next reader does not
re-derive them: a target-keyed registry restores the ambiguity the token
exists to remove, and injecting a turn counter couples layers to fill a
field nothing reads yet.
2026-08-15 06:32:02 +02:00
Dai Ha b745e159de CB-573: carry accepted turn tokens on delivery 2026-08-15 06:30:41 +02:00
Dai Ha aac29d604c M4: correct unit 2 criterion 1 — TurnToken cannot carry the session turn
CI / contract (push) Successful in 1m0s
CI / build (push) Successful in 1m32s
The criterion required the token to bind the session turn number. Three
independent refusals from the implementer showed why that is not
implementable at this layer: MessageService owns acceptance but never
learns of delivery, and CompletionResolver.onDelivered runs before
SessionManager.onDelivered, so the turn number does not exist yet at the
only point the token could capture it.

Records both rejected alternatives and why, so the next reader does not
re-derive them: a target-keyed registry restores the ambiguity the token
exists to remove, and injecting a turn counter couples layers to fill a
field nothing reads yet.
2026-08-15 06:29:27 +02:00
Dai Ha c884802b13 CB-577: remove obsolete async target tracking
CI / contract (pull_request) Successful in 43s
CI / build (pull_request) Successful in 56s
2026-08-15 06:28:53 +02:00
Dai Ha 5f5573a24e Merge CB-577: correlate an async question by its exact waiter
CI / contract (push) Failing after 0s
CI / build (push) Successful in 1m13s
markAsyncQuestion picked the first not-done task out of an unordered
set, so between resolveQuestion waking the first async send and the
question being recorded, a queued second send could join the set and
take the question. A lead answering with bridge_send{turnId} would then
resume a turn it did not mean to.

Each async task is now indexed by its exact rendezvous waiter, which has
identity semantics, so no other task can hold the same key. The question
is recorded before resolveQuestion, with a rollback when no waiter is
there, which closes the window rather than narrowing it.

An unanswered async question stays PENDING — the worker resumes after
its ask times out, so the delegation is not failed — and its stale
target tracking is now cleared instead of leaking.
2026-08-15 06:27:16 +02:00
Dai Ha 74b0087ebb CB-577: model async question ownership race
CI / build (pull_request) Successful in 1m29s
CI / contract (pull_request) Failing after 0s
2026-08-15 06:24:05 +02:00
Dai Ha 927e0151d4 CB-577: test async question waiter ownership 2026-08-15 06:23:22 +02:00
Dai Ha 5275922d1d CB-577: handle questions without async waiters 2026-08-15 06:22:22 +02:00
Dai Ha e186c7945a CB-577: correlate async questions to turns 2026-08-15 06:21:53 +02:00
Dai Ha 33a6e77f0e Merge CB-573: dormant fleet health monitor (M4 unit 1)
CI / build (push) Successful in 49s
CI / contract (push) Successful in 1m14s
An opt-in whole-fleet observer, separate from the 250ms delivery poller.
One AgentControl.list and one roster snapshot per tick, joined and fed to
the FleetHealth classifier, because a fault is a disagreement between the
two views at the same instant. Absent a health: block nothing is built
and no herdr call is made.

Adds bridge_list healthCoverage: off, detection-only, or full. Detection
is deliberately separate from notification, so a single-lead setup with
no webhook still gets detection and is told its coverage is partial
rather than being refused.

Two review fixes worth naming. tick() rescheduled itself as its last
statement with no try/catch, and a ScheduledExecutorService does not
re-run a task that threw — so the first agents.list failure would have
stopped health permanently and silently, which is exactly when the
control link is down. It now catches Throwable and reschedules in a
finally. And the snapshot fields this unit cannot supply are the named
constant NOT_YET_OBSERVED rather than bare false literals, because false
means no fault to this classifier.
2026-08-15 06:19:15 +02:00
Dai Ha 4e47489d53 Merge CB-568c: fail every pending async ticket on teardown
A released target left its second async ticket pending for the full
30-minute async timeout. abandon resolved only the rendezvous waiter,
and async tickets live in a separate map that could not even represent
two tasks on one target.

abandon now sweeps every non-question async ticket for the target, and
asyncTasksByTarget holds a set. The sweep is a plain loop: the first
attempt used Stream.anyMatch, which short-circuits on the first true, so
it completed one ticket and left the rest pending — the exact bug it was
fixing. Its test passed only because the rendezvous path failed the
first ticket anyway; the test now uses three tickets so a single
completion cannot satisfy it.

An ASKING ticket is an active turn, not a pending send, so the sweep
skips it and CB-574 is unaffected.
2026-08-15 06:18:28 +02:00
Dai Ha c76b2f149e CB-573: keep health monitoring after failures
CI / contract (pull_request) Successful in 42s
CI / build (pull_request) Successful in 1m15s
2026-08-15 06:17:28 +02:00
Dai Ha 826fffe05b CB-573: add dormant fleet health monitor 2026-08-15 06:15:24 +02:00
Dai Ha 75f57cdba7 CB-568c: fail every queued async ticket
CI / contract (pull_request) Successful in 44s
CI / build (pull_request) Successful in 1m34s
2026-08-15 06:15:21 +02:00
Dai Ha f556af5d4e CB-568: fail queued async tickets on teardown 2026-08-15 06:14:28 +02:00
Dai Ha d5f33f0c6e Merge CB-568: a dropped send reports the real cause
CI / build (push) Successful in 55s
CI / contract (push) Successful in 1m0s
Injector.drop knew the precise cause (herdr agent_not_found) but the
sender was told only 'worker unreachable or stuck', so a lead could not
tell a dead pane from a stalled model.

TurnListener.onTurnFailed gains a reason, defaulting to the old one-arg
form. CompletionResolver prefers that reason, then the pane scrape, then
the old fixed text.

drop now fires onTurnFailed unconditionally. That is the substantive
fix: the sender blocks on the rendezvous waiter, not on the delivered
future, so failing delivered() alone never woke it and a queued send sat
until its timeout.
2026-08-15 06:08:41 +02:00
Dai Ha 16e17b32ad CB-568: preserve dropped turn causes
CI / contract (pull_request) Successful in 43s
CI / build (pull_request) Successful in 1m26s
2026-08-15 06:06:07 +02:00
Dai Ha 988494e18b Merge M4 fleet-health design (docs only)
The full 5-unit design behind M4: evidence model, classification
precedence, the automatic-vs-lead action boundary, worktree safety on
release, typed inbox and lead routing, capacity, and human escalation.

Two decisions worth keeping visible. Detection is split from
notification, so health works in a single-lead setup with no webhook and
reports partial coverage instead of refusing to run. And capacity stays
a view: the bridge reports free slots but never spawns, reassigns, or
stops a member to improve utilisation, because only the lead holds the
work list.

Section 13 records eleven things nobody checked, including live LavinMQ,
OpenCode pane fixtures, and multi-lead routing.
2026-08-15 06:05:26 +02:00
Dai Ha c4549a5e20 Merge CB-575: filter the routine MCP cancellation WARN
The MCP SDK 2.0.0 registers no handler for notifications/cancelled, so every
client abort logged a WARN. M4 fleet health treats WARN as action-needed, so
that noise had a cost. A Logback TurboFilter denies only that one event:
right logger, WARN level, the SDK's exact format string, and a
JSONRPCNotification whose method is notifications/cancelled. Everything else
is NEUTRAL. If a later SDK handles cancellation the filter stops matching.
2026-08-15 06:00:40 +02:00
Dai Ha 46fa4f38d5 CB-575: filter MCP cancellation warnings
CI / build (pull_request) Successful in 49s
CI / contract (pull_request) Successful in 1m1s
2026-08-15 05:57:45 +02:00
Dai Ha 20e0e68ad7 M4: align design status with CB-573
CI / contract (pull_request) Successful in 43s
CI / build (pull_request) Successful in 1m28s
2026-08-15 05:56:00 +02:00
Dai Ha 17468a234a M4: document fleet health design 2026-08-15 05:55:00 +02:00
Dai Ha 8a53d5bfc6 Merge CB-574: an async delegation can now receive a worker's question
CI / contract (push) Successful in 1m2s
CI / build (push) Failing after 1m40s
A worker on a wait:false delegation called bridge_ask and the lead never
saw the question. Outcome.QUESTION is deliberately non-terminal, but
taskView tested r.completed() and fell into the failure branch, so the
ticket was marked FAILED and both the question text and its turnId were
discarded. The worker blocked for 55s, gave up, and had to abandon its
task. CLAUDE.md tells leads to prefer wait:false and to answer an ask with
bridge_send{turnId, content}; those two could not both be followed.

bridge_poll now returns a non-terminal ASKING phase carrying the question
and its turnId, and the ticket stays live so the worker's real reply still
lands on it. An unanswered ask returns the ticket to PENDING, because only
the question wait ended - the delegated turn continues. The 55s/115s ask
caps are unchanged: they exist because the worker's own MCP call would time
out, so widening them would only move the failure.

Two defects found reviewing the first revision, both from replacing
supplyAsync with a manually completed future:

- an exception inside the send left the future uncompleted, so the ticket
  stayed PENDING for the life of the daemon. Now caught and completed
  exceptionally.
- correlation was keyed by target, one entry per worker, registered before
  the session lock. With two tickets outstanding on one target the second
  overwrote the first, so a late reply could resolve the wrong ticket.
  Correlation is now per turn, the target entry exists only while that send
  owns the lock, and a reply with no live waiter still goes to the durable
  inbox as before.
2026-08-15 05:50:45 +02:00
Dai Ha 695da7418e Merge CB-573 (part 1): health classification model and the bridge_list capacity view
CI / contract (push) Successful in 45s
CI / build (push) Successful in 55s
Fleet capacity was invisible. A finished member held a terra slot until a
spawn was refused with 'at maxLoad: 2 live >= 2 cap', and nothing had told
the lead the slot was still held. bridge_list now reports, per configured
profile, maxLoad / live / free / reclaimable, and per member idleForSeconds
and reclaimable.

live comes from the same liveCountRef function placement consumes, so the
advertised free slots cannot drift from what bridge_spawn will accept. The
profile list is the union of configured and roster profiles: an empty
configured profile still appears with its full capacity, and a member whose
profile was removed from config stays visible rather than vanishing.

reclaimable is advisory. The bridge never spawns, stops or retasks a member
to improve utilisation: it has capacity facts but no work list, and choosing
work needs authority it does not have.

Also lands the pure health classifier, its precedence chain, the MUTE counter
and the pane budget. The classifier never reports IDLE while an accepted
delivery is open — IDLE is a claim that nothing is outstanding, and the
capacity view reads exactly that field.

Capacity dependencies are one required CapacitySource rather than defaulted
constructor arguments. A defaulted liveCount would report free slots that do
not exist, which is the dangerous direction; CapacitySource.none() omits the
block instead of inventing zeros.
2026-08-15 05:49:51 +02:00
Dai Ha abd26c796b CB-574: retain async task correlation
CI / build (pull_request) Failing after 1m14s
CI / contract (pull_request) Successful in 1m15s
2026-08-15 05:49:27 +02:00
Dai Ha 24559d81ac CB-573: require explicit capacity source
CI / contract (pull_request) Successful in 44s
CI / build (pull_request) Successful in 55s
2026-08-15 05:49:04 +02:00
Dai Ha 6c1c2c3994 CB-573: report empty configured profile capacity
CI / build (pull_request) Successful in 59s
CI / contract (pull_request) Successful in 1m0s
2026-08-15 05:45:39 +02:00
Dai Ha 01fab15713 CB-573: add fleet capacity view
CI / contract (pull_request) Successful in 1m1s
CI / build (pull_request) Successful in 1m25s
2026-08-15 05:43:46 +02:00
Dai Ha 0cd00e71c3 CB-574: surface async worker questions
CI / contract (pull_request) Successful in 42s
CI / build (pull_request) Successful in 56s
2026-08-15 05:43:43 +02:00
Dai Ha bf0ff2adbf CB-573: keep active delegations out of idle
CI / contract (pull_request) Successful in 45s
CI / build (pull_request) Successful in 1m27s
2026-08-15 05:40:37 +02:00
Dai Ha ed4bbc1c56 CB-573: add pure health classification model
CI / build (pull_request) Successful in 59s
CI / contract (pull_request) Successful in 1m15s
2026-08-15 05:37:30 +02:00
Dai Ha c456402cc5 Merge CB-572: reject a configured profile name as a send target
CI / build (push) Successful in 59s
CI / contract (push) Successful in 1m16s
A lead sent to sessionId "sol" — a profile name, not a terminal id. The
bridge accepted it, handed out a ticket, failed 60s later inside the
injector, and still reported the ticket as pending 20 minutes on. The
sender never learned anything and a whole delegation was lost.

bridge_send now rejects a target that exactly matches a configured
profile name, on both the blocking and the wait:false path, before any
ticket is issued. The error names the value and points at bridge_list.

The check is deliberately narrow. A target absent from the member roster
may still be a peer lead's terminal or a herdr-owned pane, so only a
value the bridge can prove is a profile is refused. profiles is a
required parameter on both send methods — the earlier revision kept
overloads that defaulted it to an empty set, which is the same silent
disable shape as CB-561.
2026-08-15 05:35:12 +02:00
Dai Ha 4f0bf667b1 CB-572: require profiles for send validation
CI / build (pull_request) Successful in 1m0s
CI / contract (pull_request) Successful in 1m2s
2026-08-15 05:33:15 +02:00
Dai Ha 61af9aa574 CB-572: reject profile names as send targets
CI / build (pull_request) Successful in 59s
CI / contract (pull_request) Successful in 1m20s
2026-08-15 05:29:37 +02:00
Dai Ha 3b59b34e76 CB-570 follow-up: rewrap an over-long javadoc line
CI / build (push) Successful in 58s
CI / contract (push) Successful in 1m17s
2026-08-15 05:12:58 +02:00
Dai Ha 65b38997f7 Merge CB-570: OpenCode receives the composed charter (PR #40)
One file, member-charter.md, not two. Two files would have made the U5
digest non-comparable between the Claude adapter and this one, which is
the whole point of the receipt; and nobody verified how OpenCode merges
multiple instruction files, so array order was an unverified dependency.

The OPENCODE_CONFIG condition widens to include a charter. It used to be
hasMcp() || hasCustomProvider(cfg), so a profile with a role charter but
no MCP and no custom provider would have got no config file and therefore
no charter — the feature silently doing nothing for that profile.

The file stays in the per-spawn temp dir, never the worktree: the
worktree is removed on release, the parity overlay already writes into
it, and CB-525's lesson was that config the bridge copied into a worktree
made a worker operate on the wrong tree. Being outside the repo is also
what stops it being committed, which a .gitignore line does not.
2026-08-15 05:12:26 +02:00
Dai Ha 7510f7649c Merge CB-569: Claude receives the composed charter (PR #39)
argvWithBridge used to gate the charter on cfg.hasMcp(), because the only
charter was the reply rule and telling a peer to call a tool it was not
given is a bug. A role charter is identity, not a tool instruction, so
the two gates are now separate: the MCP mount still depends on mcpUrl,
while --append-system-prompt depends only on the base having composed
something. A profile with a role charter and no MCP now gets its charter.

A null charter adds no flag at all. An empty --append-system-prompt is
not the same as no system prompt.
2026-08-15 05:11:39 +02:00
Dai Ha 619792a81c CB-570: deliver composed OpenCode charter
CI / build (pull_request) Successful in 58s
CI / contract (pull_request) Successful in 1m13s
2026-08-15 05:10:49 +02:00
Dai Ha cea1183f75 CB-569: pass composed charter to Claude
CI / contract (pull_request) Successful in 43s
CI / build (pull_request) Successful in 50s
2026-08-15 05:09:30 +02:00
Dai Ha e2af4c5ae4 CB-567 follow-up: realign the changed constructor signatures
CI / contract (push) Successful in 44s
CI / build (push) Successful in 1m26s
The supplier swap left three parameter lists indented one column past
their siblings and one javadoc line broken mid-sentence. No behaviour
change.
2026-08-15 05:06:03 +02:00
Dai Ha dfd5f82894 Merge CB-567: read the charter per spawn, compose it once (PR #38)
HerdrPeerLauncher now takes Supplier<BridgedConfig.Fleet> instead of
Supplier<String> tabLabelTemplate, and reads it once per spawn. A field
taken at construction would have made the charter deferred, and deferred
looks exactly like working — which is why the test uses a mutable
supplier and spawns twice, rather than ConfigRef.fixed().

Composition happens once in the base, not in each adapter: two copies
drift while both adapter-local tests keep passing. Role charter first,
reply charter last, because the final instruction is the one that must
not be overridden. The reply charter stays gated on hasMcp() — telling a
peer to call a tool it was not given is a bug — while the role charter is
not, being identity rather than a tool instruction.

Both REPLY_CHARTER copies collapse into one, and it now says 'spawned
member' rather than 'off-subscription worker'. The old text made the
launch prompt contradict bridge_whoami for an architect; both architects
read it in their own prompts and reported it.

The old buildLaunch overloads are removed rather than kept as defaults: a
surviving one is the same shape as a stale snapshot, a route that drops
role and charter while looking healthy. That removal also let the
OpenCode adapter drop its ThreadLocal resume-id hack, since LaunchSpec
now carries the value down the same path.
2026-08-15 05:04:08 +02:00
Dai Ha e81944cef6 CB-567: compose charters per spawn
CI / build (pull_request) Successful in 57s
CI / contract (pull_request) Successful in 1m3s
2026-08-15 05:00:48 +02:00
Dai Ha 7bdd39ab9a CB-566 follow-up: say why the five-arg Fleet constructor is kept
CI / contract (push) Successful in 43s
CI / build (push) Successful in 1m8s
Reviewing the merge I read the five-arg constructor as dead code and
removed it. That was wrong: the tests call it as BridgedConfig.Fleet,
which my grep for 'new Fleet(' did not match, and the build failed on
eight call sites. It is restored with a javadoc that says why keeping it
is safe here even though an overload that drops a new field is normally
the shape to avoid — nothing reads a charter through a constructor, and
Jackson binds the canonical one, so it cannot swallow an operator's YAML.

Also drop a redundant java.util.Arrays qualifier (the class is already
imported) and rewrap a javadoc line the change had left over-long.
2026-08-15 04:54:09 +02:00
Dai Ha f4b38f040e Merge CB-566: the charter config surface (PR #37)
fleet.charters is a validated Map<String,String>, not a record: Fleet is
@JsonIgnoreProperties(ignoreUnknown = true), so a record field named
architetc would be dropped in silence and the operator would never learn
of the typo. A map lets validateCharters see the bad key and refuse it.

The key sits under fleet: because ConfigRef already treats that block as
hot and changedDeferredKeys does not list it. A new top-level key would
inherit nothing, and forgetting to classify it means a reload prints
'config reloaded' and does nothing.

A blank value is refused while an absent one is fine: an absent key means
the operator configured no charter, a blank one means they tried and
failed. Refusing at both startup and reload is the point — wiring only
one of the two paths is the whole bug.
2026-08-15 04:52:48 +02:00
Dai Ha 799668e129 CB-566: add fleet charter config
CI / contract (pull_request) Successful in 43s
CI / build (pull_request) Successful in 55s
2026-08-15 04:50:23 +02:00
Dai Ha 890190263e CB-560/562/563: the shipped block tells the truth again
The canonical CLAUDE.md block is the instruction surface this daemon ships to
every agent that mounts it, so a code change that silently invalidates it is an
incomplete change. CB-548 added a third principal kind and the block was never
revisited. Four statements in it were simply false:

  * whoami was documented as returning only primary or worker;
  * "spawn/stop/send/drain are lead-only" — Authz permits SEND to an architect;
  * delivery was documented as idle/blocked — injectable() is IDLE|BLOCKED|DONE,
    and a spawned member must also have mounted the bridge MCP, which is the
    exact condition that made every architect undeliverable for a day;
  * the tool table said bridge_list returns `workers` — the JSON key is
    `members`.

Also corrected: the fallback ladder claimed each one-way signal identifies a
"worker", but an architect gets the same charter, the same mount and the same
env, so those signals identify a spawned member and only bridge_whoami
separates the two. The safe default stays "act as a worker" — it is the most
restricted member role.

The turn contract now covers both member kinds, and says why the completion
fallback is not a substitute for bridge_reply: it returns at most the last 4000
characters, so a long report reaches the lead with its end cut off. That is not
hypothetical — it happened twice today.

Found by a reviewer asked whether the block still matches the code. Verified
against injectable(), Authz, BridgeMcp.listFleet and LeadLauncher before
applying. The wiki template is updated in the same shape and re-checked
byte-identical.
2026-08-15 04:37:03 +02:00
Dai Ha f8522edacd Merge CB-564: give every silent failure path a voice
CI / build (push) Successful in 1m0s
CI / contract (push) Successful in 1m15s
The bridge cannot classify what it never emits. Fleet-health monitoring —
detect a wedged member, decide, escalate — is blocked on that, so this is the
foundation rather than the feature.

A survey of Injector, CompletionResolver, StatusPoller, SessionManager,
SessionReaper and MessageService found six conditions that ended a member's
usefulness while saying nothing useful:

  * Injector.drop()             — worker gone, queue cleared: SILENT
  * CompletionResolver.fail()   — "via turn-stall fallback" at DEBUG, no reason
  * SessionManager.onFailed()   — "session marked failed" at DEBUG, no stage
  * SessionManager.reapIdle()   — indistinguishable from any other release
  * SessionManager.acquire()    — spawn failure rethrown with no log at all
  * MessageService.abandon()    — failed a caller's request at DEBUG

The first two are the exact phrases that misled the CB-560 diagnosis: both name
a symptom and neither names a cause. They are now WARN and carry the reason,
the stage, and the counts.

Conditions already loud were left alone, and so were two by-design timeouts in
MessageService — an async model exists precisely for those, and promoting them
would turn healthy operation into noise.

Behaviour is unchanged: every edit is a log statement.

Merge note: the recycle test deleted by CB-565 conflicted with a test added
here. Resolved by keeping the new onTurnFailed assertion and dropping the
recycle test, which tests a method that no longer exists.
2026-08-15 04:36:09 +02:00
Dai Ha b958747855 Merge CB-565: remove the recycle trap (PR #35)
CI / contract (push) Successful in 58s
CI / build (push) Successful in 1m27s
SessionManager.recycle called the 4-argument acquire overload, which defaults
the role to DEV and requests no worktree. So it carried profile, cwd and owner
across and silently dropped two things: the member's role, and its worktree.

A recycled architect would have come back a plain worker, never rebound to its
slot, with nothing logged. Worse, a recycled member would have come back with
no worktree at all — and a member's uncommitted work exists in exactly one
place. Role loss is recoverable; that is not.

Nothing called it. The only reference outside its own javadoc was one test, and
the context-cap path calls release, not recycle. So this was a trap waiting for
its first caller, and that caller would not have noticed either loss.

Deleted rather than repaired, on the CB-561 precedent: an API that looks correct
and silently drops a property is worse than no API. The no-reuse invariant it
documented is still true and is now stated directly.

Found by a reviewer asked to hunt the rest of the CB-548 fallout. The worktree
half was found while verifying the report.
2026-08-15 04:31:46 +02:00
Dai Ha 078bde2c02 CB-564: give a voice to silent member-failure paths
Injector.drop, CompletionResolver.fail, SessionManager.onFailed/reapIdle/
acquire spawn failures, and MessageService.abandon used to fail a member
or a caller's request with no log, a bare DEBUG, or a log that named only
the symptom ("session marked failed", "failed send via turn-stall
fallback"). Each now logs at WARN and names the real cause and the
numbers involved. Observability only — no behaviour changed.
2026-08-15 04:31:42 +02:00
Dai Ha 0b10ea987b CB-565: remove unsafe session recycle
CI / contract (pull_request) Successful in 44s
CI / build (pull_request) Successful in 55s
2026-08-15 04:29:50 +02:00
Dai Ha 1553d38182 Merge CB-563: a clipped pane scrape says it was clipped (PR #34)
CI / contract (push) Successful in 45s
CI / build (push) Successful in 1m12s
When a member ends its turn without bridge_reply, CompletionResolver scrapes
the pane and resolves the waiting send with that text. The scrape is capped at
MAX_SCRAPE_CHARS (4000), and nothing told the caller when the cap had bitten.
A delegating lead could act on a report missing its end and believe it was
complete. That happened to me today: a member's full engineering report arrived
cut at exactly 4000 characters, and the only hint was a DEBUG line reading
"(4000 chars scraped)", which reads like a size and not like a warning.

The returned text now carries a marker when, and only when, it was clipped, and
the clip is logged at WARN with the original length and the cap.

The cap itself is unchanged. The problem was silence, not the number.

The CB-115 misattribution guard still compares the unmarked clipped tail to the
unmarked baseline, and the marker is appended only afterwards. Verified in the
code, not taken on report: clip() strips before truncating and the new length
check uses the same stripped length, so there is no off-by-one either.
2026-08-15 04:26:57 +02:00
Dai Ha d04b075996 CB-563: mark clipped completion scrapes
CI / contract (pull_request) Successful in 42s
CI / build (pull_request) Successful in 56s
2026-08-15 04:25:54 +02:00
Dai Ha 644927636d Merge CB-562: the readiness gate says why it gave up (PR #33)
CI / build (push) Successful in 1m1s
CI / contract (push) Successful in 1m14s
When the injector's readiness grace expired it cleared the queue and logged
nothing. The failure then surfaced elsewhere as a turn-stall, which names the
wrong cause. Diagnosing CB-560 cost two live spawns and a wrong first
hypothesis for exactly this reason: the logs said "session marked failed" and
"failed send via turn-stall fallback", and neither says the message was never
typed into the pane at all.

The expiry now logs the target, the number of messages being failed, the grace
in polls and seconds, and the real cause in plain words.

The grace in seconds is derived, not written down twice: Bridged's own
INJECT_POLL_MILLIS is deleted and Injector.POLL_INTERVAL_MILLIS is the single
source, passed to every StatusPoller. A cadence change can no longer leave a
log line confidently stating the wrong duration.

Behaviour is unchanged. This is the first structured health event in the
daemon, and the foundation the fleet-health work will build on.
2026-08-15 04:21:29 +02:00
Dai Ha 95e45007aa CB-562: single source for the injector poll cadence; tighten count assertion
CI / build (pull_request) Successful in 59s
CI / contract (pull_request) Successful in 1m17s
2026-08-15 04:20:29 +02:00
Dai Ha 7a583c4045 CB-562: log why the readiness gate gave up on a target
CI / build (pull_request) Successful in 57s
CI / contract (pull_request) Successful in 1m4s
2026-08-14 21:52:25 +02:00
Dai Ha 65bce058bb Merge CB-561: one public way to build a CallerResolver (PR #32)
CI / build (push) Successful in 1m17s
CI / contract (push) Successful in 1m16s
CallerResolver had 5 public constructors and 4 public factories, and only one
of them could ever produce an architect. The rest defaulted memberSlotRoles to
`_ -> null`, so every architect quietly fell through to Principal.worker().
Nothing logged, nothing threw — the role was simply off.

That is the same failure shape as CB-560, so the fix is structural rather than
a warning: withLeadsAndMembers(.., MemberRegistry) is now the only public
construction path. Two overloads with no caller at all are deleted; the rest
are package-private and marked test-only. No path remains that accepts
architect bindings without a slot-role lookup, so no runtime WARN is needed.

Also checked and closed: the suspected slot leak on shutdown drain is not
real. SessionManager.release calls memberLifecycle.released() for every
removed session, and drainAll routes every session through release,
SPAWNING included. Verified in code.
2026-08-14 21:50:04 +02:00
Dai Ha 48b437083b Merge CB-560: mark spawned members present (PR #31)
Merging CB-548 made an architect resolve as Role.ARCHITECT, and BridgeMcp
marked MCP presence only for Role.WORKER. So an architect was never present.
That one line broke two things, because SessionManager.asPresence() is the
same object the injector's readiness gate reads:

  * the session never left SPAWNING;
  * Bridged.deliverableTo was false, so the injector held every delivery,
    waited out READINESS_GRACE_POLLS (~60s) and failed the send.

Confirmed live twice, on claude-code and on opencode. Both logged
"session marked failed" then "failed send via turn-stall fallback", neither
of which names the real cause.

Principal.isSpawnedMember() now means "a spawned member with a pane" —
WORKER or ARCHITECT, never PRIMARY. The Bridged.deliverableTo javadoc, which
still stated the worker-only rule as fact, is corrected: it is the clearest
description of this invariant anywhere, and leaving it stale is how the bug
comes back.
2026-08-14 21:50:04 +02:00
Dai Ha fa0612859b CB-560: document spawned member presence
CI / build (pull_request) Successful in 1m0s
CI / contract (pull_request) Successful in 1m4s
2026-08-14 21:48:58 +02:00
Dai Ha 4ffbcd0b7d CB-561: require member registry for architects
CI / contract (pull_request) Successful in 43s
CI / build (pull_request) Successful in 50s
2026-08-14 21:47:53 +02:00
Dai Ha a0cd053fd9 CB-560: mark architect members present
CI / build (pull_request) Successful in 55s
CI / contract (pull_request) Successful in 1m4s
2026-08-14 21:45:28 +02:00
Dai Ha f9a5e066b5 Merge PR #30: bind a spawned architect to its slot (CB-548's missing half)
CI / contract (push) Successful in 48s
CI / build (push) Successful in 1m18s
MemberRegistry.bind() had 18 call sites, every one in its own unit test.
Nothing in production ever bound a terminal to a slot, so CallerResolver's
architect branch was unreachable, every spawned architect resolved as WORKER,
and Authz's architect row was dead code. Since SEND is granted to
isPrimary() || isArchitect(), no architect could message anyone.

The spawn lifecycle now binds on both acquire paths and unbinds on release,
through a MemberLifecycle seam with a no-op default, so all six SessionManager
constructors are unchanged.

The escalation guard is deliberately double-sided. MemberRegistry flattens
every pool, so dev:* and reviewer:* slots sit in the same map the architect
resolver reads. The lifecycle binds only ARCHITECT roles, AND CallerResolver
independently checks roleForSlot(slot) == ARCHITECT before granting. Either
alone would be enough today; together, a regression in one cannot escalate a
worker. A reviewer confirmed a third barrier already existed: Authz gates
SPAWN on isPrimary(), so only a lead can request role=ARCHITECT at all.

Verified by the primary: mvn clean install, 640 tests, 0 failures.

Two findings recorded, neither blocking:

1. Known race, low severity. registry.put() and memberLifecycle.acquired()
   are not atomic. A release landing between them unbinds nothing (no binding
   exists yet), then acquire binds a dead terminal that no later release will
   ever clear - the slot leaks. Only the shutdown drain can reach a SPAWNING
   session (the idle reaper skips it, and an explicit stop needs a paneId the
   spawn has not returned), so the leak dies with the process. Note that
   binding before put does NOT fix it: the drain then iterates a roster the
   session is not in yet. A real fix needs atomicity.

2. The legacy Supplier-based withLeadsAndMembers overload and the Map-form
   constructors now silently disable architect resolution: memberSlotRoles
   defaults to _ -> null, so the architect branch can never be taken there.
   Production uses the registry form, and the tests were migrated, so nothing
   fails - but a future caller gets workers with no error.
2026-08-14 21:25:45 +02:00
Dai Ha e9bc192160 CB-548: bind architect sessions to member slots
CI / contract (pull_request) Successful in 45s
CI / build (pull_request) Successful in 55s
2026-08-14 21:10:48 +02:00
Dai Ha e0a57988ad Merge PR #29: drop .claude/settings.local.json from the default parityOverlay
CI / build (push) Successful in 1m21s
CI / contract (push) Successful in 1m21s
A member no longer receives the primary's pre-approved tool grants by default.
The grants were inert — GitWorktrees.isolateToolSurface already strips each
worktree's .mcp.json to an empty server map — so this is defence in depth: two
independent guards instead of one. Same reasoning as CB-525, which removed the
sibling .mcp.json and left this file behind.

An explicit parityOverlay: in config is unaffected; only the default changes.

Verified by the primary: mvn clean install, 637 tests, 0 failures.
2026-08-14 21:02:25 +02:00
Dai Ha b414a74c26 CB-559: drop .claude/settings.local.json from the default parity overlay
CI / build (pull_request) Successful in 54s
CI / contract (pull_request) Successful in 1m7s
2026-08-14 20:56:16 +02:00
Dai Ha 2c3796d598 Merge: keep every credential in one store
CI / build (push) Successful in 55s
CI / contract (push) Successful in 1m13s
2026-08-14 20:38:36 +02:00
Dai Ha 7e0ff9ab06 Keep every credential in one store, not a per-repo copy
The repo carried a gitignored .secrets/ directory with four files. Two of them
(context7-token, gitea-token) were byte-identical copies of variables the login
shell already exported. One (gitea-host) is not a secret. The fourth
(worker-gitea-token) was the only copy anywhere, and nothing exported it, so
bridged read gitTokenEnv from an environment that never had it and every worker
push got an empty token.

All four values now live in the operator's single sourced secrets file, verified by
sha256 before the copies were removed. opencode.json reads them as {env:...}, which
.mcp.json already did. A second copy of a secret is the problem: the copy you forget
is the one that leaks or goes stale.

This makes worktree isolation load-bearing rather than a workaround. opencode.json is
tracked, so it lands in every worktree. It used to fail there, because {file:.secrets/}
pointed at files a worktree never receives and OpenCode refuses to start on a dangling
reference. With {env:...} the reference resolves, and a member would silently inherit
the primary's admin-scoped GITEA_ACCESS_TOKEN. GitWorktrees already neutralizes the
file; only its stated reason changes, and it is now a confidentiality boundary.

The port-to-opencode skill taught {file:.secrets/} as the preferred pattern, so it is
rewritten to teach the central store and to say why we moved. .gitignore keeps the
.secrets/ line as a backstop against habit.

Includes the wiki pointer, which also carries the CB-559 config-reload correction.
2026-08-14 20:38:31 +02:00
Dai Ha 6d1565d8b6 Merge CB-559 correction: profile launch settings are deferred, not hot 2026-08-14 20:37:47 +02:00
Dai Ha e18e002d2f CB-559: a profile's launch settings are deferred, not hot
The shipped docs and javadoc said a profile's `model` and `tabLabel` take effect on
the next spawn. They do not, and ConfigRef did not detect the change either, so a
reload logged a clean "config reloaded" and silently did nothing. That is the worst
outcome a reload can produce: the operator has no reason to doubt it.

What makes a key hot is who reads it and when, not that it is config. Placement
reads weight and maxLoad through a supplier on CompositePeerLauncher, so those are
genuinely hot. HerdrPeerLauncher takes Map.copyOf(profiles) at construction and
resolves every spawn out of that copy, so model, baseUrl, argv, env and the rest
cannot move until the daemon restarts.

changedDeferredKeys now compares every launch component of an existing profile,
excluding weight and maxLoad, and names the profiles that need a restart. The
javadoc and bridged.example.yaml say the same thing. Two tests pin the pair:
weight/maxLoad reports nothing deferred, a changed model reports the profile by name.
2026-08-14 20:37:42 +02:00
Dai Ha c18572ea9d Bump wiki to 8b1eb68 — the CB-557/558/559 Features entries
CI / contract (push) Successful in 45s
CI / build (push) Successful in 1m14s
2026-08-14 19:52:02 +02:00
Dai Ha e9d02bc5e8 Merge CB-557/558/559: the fleet block, role pools, lead auto-launch, config reload
Four commits that together reshape how the daemon is told who it may run.

CB-557 (schema) — one `fleet:` block replaces `leaders:`, `members:`, `leadScan:`
and `defaultProfile:`. Role is the containing map key, not a `role:` field, so a
misspelled pool name declares nothing instead of producing a member with no
contract. Tab labels are role-first and counted per role+profile.

CB-557 (routing) — an unqualified spawn picks its profile from its role's pool.
Role and profile stay orthogonal: a reviewer may run on the same profile as the
dev whose diff it reads, so the two cannot be one field. An explicit
`bridge_spawn{profile:…}` stays exempt, because it carries no role and would
otherwise be refused for a profile the operator named.

CB-558 — the daemon launches a declared lead when fewer than `instances` are
running. An auto-launched lead is NOT a member: no reply charter, never
registered with SessionManager (the idle reaper would kill it), and no Anthropic
binding in its env. A lead counts as live only when herdr reports a running
agent, so one crash does not disable auto-launch forever, and an unreachable
herdr starts nothing at all.

CB-559 — `bridged.yaml` can be re-read without a restart, opt-in through
`configReload:`. Keys are hot (`fleet:`, `placement:`, an existing profile's
fields), deferred (lifecycle, guard, adding a profile) or cold (bind,
herdrSocket, broker, auth). A cold change refuses the WHOLE reload rather than
half-applying it, because a daemon matching no file on disk is worse for an
operator than no reload at all.

634 tests, IDE diagnostics clean. Verified live against the running daemon: the
lead launcher recognised the existing lead and started nothing; the reload
applied a hot change, refused a `bind:` change exactly once (not once per tick),
and reverted cleanly.
2026-08-14 19:51:34 +02:00
Dai Ha a2108a8a14 CB-559: re-read bridged.yaml without restarting the daemon
Tuning a fleet meant restarting bridged, and a restart tears down every lead
and worker it owns. Changing one pool's weight cost the whole fleet's state,
so in practice nobody changed it.

ConfigRef holds the live BridgedConfig in an AtomicReference. Consumers read
it at the point of use, so a change reaches the next spawn with nothing
rebuilt. The launchers that used to capture config into fields now take
suppliers: the fleet tabLabel template, the profile map, the placement policy
and the fleet block.

Keys fall into three classes, and the difference is what already exists when
the reload happens:

  hot       fleet: (pools + tabLabel), placement:, and an existing profile's
            weight / maxLoad / model / tabLabel — live on the next spawn.
  deferred  lifecycle:, leadHeartbeat:, guard:, worktreeRoot:, spawnReady*,
            and adding/removing a profile — accepted, but the startup wiring
            keeps the old value. The reload logs these by name.
  cold      bind:, herdrSocket:, broker:, auth: — refuses the WHOLE reload.

A cold change refuses everything rather than applying the hot half. A
half-applied reload leaves the daemon matching no file on disk, which is the
worst thing a reload can do to an operator reading that file to work out what
the daemon is doing. Refusing keeps the invariant that the live config is
always some version of the file.

A parse failure or a failed startup validator is refused the same way, and
the running config stays live: a file being saved is sometimes read
mid-write, and degrading a working daemon over a half-written file is a bad
trade. The same four validators startup runs are re-run, so a config that
could not have booted cannot slip in through a reload.

ConfigWatcher polls the modified time on a daemon thread, opt-in through
configReload.enabled (default off, so an upgraded daemon is unchanged). It
stamps the timestamp BEFORE reloading, so a refused file is not retried every
tick — the next save earns a fresh attempt. A missing file is skipped
silently, because editors unlink briefly mid-save.

MicroProfile Config was the first idea and does not fit: @ConfigMapping needs
interfaces, resolves once at bootstrap, and reload would still mean rebuild
and swap. The port would also lose the raw-YAML duplicate-key detection,
since duplicates have already collapsed once the tree is flattened to
properties.

634 tests.
2026-08-14 18:23:32 +02:00
Dai Ha 57f8fa257a CB-558: launch a declared lead at startup when none is live
`fleet.leaders.<name>.instances` was descriptive. Now the daemon reads it: a
lead that names a `profile:` is started when fewer than `instances` are running.
A lead with only a `terminal:` stays recognise-only, as before.

A lead is not a member, and LeadLauncher exists to keep it that way. Every other
spawn path goes through HerdrPeerLauncher, which does three things a lead must
never get: it appends the worker reply charter ("you are an off-subscription
worker … end every turn with bridge_reply" — the opposite of an orchestrator);
it registers the session with SessionManager, whose idle reaper would kill a
lead for being idle, which is a lead's normal state; and it can move a peer off
the subscription. So this launcher talks to AgentControl/WorkspaceControl
directly. The duplicated argv/env assembly is the cheaper half of that trade.

Not double-spawning is the safety property, so liveness needs two pieces of
evidence. A running agent in a tab labelled `lead: <name>` finds an
auto-launched lead. A running agent on a pinned `terminal:` finds one the
operator opened by hand — without it, a pinned lead whose tab carries no
matching label would be relaunched on every boot. Member workspaces are
excluded, so a member in a matching tab is never counted. If herdr cannot be
reached, nothing is started: a second orchestrator is worse than none.

Liveness deliberately requires the AGENT, not just the label. LeadTabScanner
used to promise that bridged never writes a lead label, so there was no
round-trip from the daemon's own rename back into its next decision. That is no
longer true, and its javadoc now says so. The trust direction is unaffected — a
label is a name, not a capability — but staleness becomes real: a label left by
a crashed session would otherwise read as a live lead forever and disable
auto-launch permanently.

Two new knobs. `workspace:` (default "leads") is where a launched lead's tab
goes; it must not be a member workspace, because those are excluded from the
scan and a lead placed in one would never be found again. `cwd:` defaults to
bridged's own working directory.

Also: WorkspaceControl.listTabs, and a FakeHerdr tab seeder that leaves the
canned response byte-identical when no tab is seeded.

617 tests pass (16 new), IDE-clean.
2026-08-14 16:47:57 +02:00
Dai Ha 61944fc045 CB-557: place an unqualified spawn inside its role's pool
The pools were config-only until now: the launchers still received one global
effectiveDefaultProfile and placement still ranged over every configured
profile, so a reviewer could be placed on an architect-only backend.

Three parts:

SessionManager computed the role, stored it on the MemberSession, and never put
it on the SpawnRequest. So the role reached the record that describes the spawn
but not the call that performs it — every launcher saw DEV. Both spawn paths
(plain and worktree) now carry it.

CompositePeerLauncher takes the Fleet and draws its candidates from
fleet.<role> instead of from all profiles. An absent or empty pool means
unconstrained, not blocked: a config that declares pools for some roles must
keep spawning the rest, so it falls back to every profile. A null Fleet is the
pre-CB-557 wiring and behaves exactly as before.

Bridged passes cfg.fleet() to the composite and cfg.fleet().tabLabel() to both
launchers. The tab-label knob was accepted by HerdrPeerLauncher but passed by
nobody, so it was inert — the label only looked right because the fallback
happened to match the configured template. Four tests now pin the wiring
instead of the coincidence.

An EXPLICIT profile stays exempt from the pool. `bridge_spawn{profile:"opus"}`
carries no role, so it defaults to DEV; judging it against the dev pool would
refuse a spawn the operator asked for by name. maxLoad still applies to it.

Also cleared the IDE warnings in the touched files: an immediately-rethrown
catch (the comment stays, the redundant block goes), unused lambda params, a
javadoc link to a package-private class, two unused imports.

601 tests pass.
2026-08-14 16:37:29 +02:00
Dai Ha 4b48d2d921 CB-557: fleet role pools — role is the config key, and the tab label says it
Four top-level keys (leaders:, members:, leadScan:, defaultProfile:) become one
`fleet:` block, and a member's role becomes the map key that contains it rather
than a `role:` field inside it.

Why the key and not a field: a misspelled `role: architct` used to produce a
member with no contract, which nothing rejected. A misspelled pool name declares
nothing, which is a shape the loader can see.

`fleet.architects/developers/reviewers` are pools of profiles a role MAY run on.
That replaces the single global `defaultProfile:`, so an unqualified spawn now
resolves its profile from the pool of the role it asked for. Role and profile
stay orthogonal: a reviewer may run on the same profile as the dev it reviews,
and one profile may appear in several pools.

Tab labels are role-first — `dev: sonnet #4`. The template lives on `fleet:`
because a profile cannot know the role of the member launched on it; a profile
may still override it. The `{n}` counter is scoped per role+profile, so a dev
and a reviewer on one profile each start at #1. Making {role} the first field
also turns the lead/member namespace check into a structural guarantee: roles
are a closed enum, so only hand-written templates can still collide with a lead
tabPrefix.

Removed keys are hard errors that name their successor. `defaultProfile:` has no
single successor key, so its message explains the new model instead of pointing
at a key that does not exist.

Map order is kept with LinkedHashMap, deliberately not Map.copyOf — the latter
salts iteration order per JVM run, which would destroy the YAML definition order
that `placement: fixed` selects on.

Not yet wired: SessionManager still hands the launchers one effectiveDefault-
Profile, so pools are not enforced at spawn time yet, and placement still ranges
over all profiles.

595 tests pass.
2026-08-14 16:32:54 +02:00
Dai Ha 3aa145cbef Merge CB-557: member taxonomy in the API surface
CI / contract (push) Successful in 1m0s
CI / build (push) Successful in 1m1s
2026-08-14 07:08:26 +02:00
Dai Ha 4875127daa CB-557: adopt the member taxonomy in the API surface
Every spawned peer is now a member with a role, and the role travels with it
from the spawn call to the roster.

MCP:
  bridge_spawn gains role: architect | dev | reviewer (default dev). An
    unknown role is refused with the valid spellings in the message.
  bridge_list returns "members" instead of "workers"; each row carries both
    role (what it is for) and profile (which backend it runs on).
  The spawn result echoes the role back, so a spawn that fell back to dev is
    visible rather than silent.

REST:
  GET/POST /members and DELETE /members/{paneId} replace /workers.
  POST accepts role= as a query param or a body field; an unknown role is 400.

Code:
  dev.ltms.bridged.worker package -> dev.ltms.bridged.member
  WorkerSession   -> MemberSession, plus a MemberRole role component
  WorkerPresence  -> MemberPresence
  SessionManager.acquire gains a role parameter; the existing overloads keep
    working and default to DEV, which is exactly what "worker" used to mean.

ClaudeCodeLauncher and OpenCodeLauncher keep their names on purpose — they
are named after the backend, not the role.

Not done here: the launch charter is still one string for every role, so a
member is told its role by nobody yet. That is the next ticket.

mvn clean install: 583 tests, 0 failures, 0 errors, BUILD SUCCESS.
2026-08-14 07:08:26 +02:00
Dai Ha d14a624421 Merge CB-557: member taxonomy in config
CI / contract (push) Successful in 45s
CI / build (push) Successful in 1m32s
2026-08-14 07:02:22 +02:00
Dai Ha 246f50b778 CB-557: adopt the member taxonomy in config
A member is anything a lead spawns. Every member carries two independent
attributes:

  role    — which contract: architect, dev or reviewer. It picks the launch
            charter, the role file, the playbook skill and the authz row.
  profile — which backend: model, CLI adapter, credentials, cost.

They vary on their own. A reviewer may run on the same profile as the dev
whose diff it reads, which is the case that proves the two cannot be one
field.

Config changes (breaking — we are in active development, so no aliases):

  workers:       -> profiles:        it was never a list of workers; it is a
                                     catalogue of backends
  defaultWorker: -> defaultProfile:
  architects:    -> members:         each slot now names its role

An old config is rejected at load with the new key named, rather than being
warned about once and then running with zero profiles — that failure would
surface much later, at the first spawn, pointing nowhere near the cause.

Also:
  - BridgedConfig.Worker    -> BridgedConfig.Profile
  - ArchitectRegistry       -> MemberRegistry
  - new peer.MemberRole enum, validated at startup
  - profiles map is normalized once in the compact constructor, so the raw
    map and the derived one can no longer disagree
  - the legacy singular worker: block is dropped
  - workerProfiles() -> profiles(); defaultProfile() -> effectiveDefaultProfile()
    (the record component now owns the plain name)

mvn clean install: 578 tests, 0 failures, 0 errors, BUILD SUCCESS.
2026-08-14 07:02:17 +02:00
Dai Ha 15b53c6cfa Merge CB-553: enforce maxLoad on explicit-profile spawns
CI / contract (push) Successful in 46s
CI / build (push) Successful in 1m8s
2026-08-14 06:51:21 +02:00
Dai Ha 3a7ef0adbd CB-553: enforce maxLoad on explicit-profile spawns (no cap bypass) 2026-08-14 06:50:28 +02:00
Dai Ha 293a305748 Merge CB-551: idle-lead heartbeat
CI / build (push) Successful in 1m15s
CI / contract (push) Successful in 1m21s
2026-08-13 21:21:33 +02:00
Dai Ha e5038c6d13 CB-551: idle-lead heartbeat — nudge an idle lead back to work on a timer
CI / contract (pull_request) Successful in 43s
CI / build (pull_request) Successful in 56s
2026-08-13 21:10:21 +02:00
Dai Ha 13ea79f6fd CB-523 (opencode): generated worker config pins compaction.auto
CI / build (push) Successful in 56s
CI / contract (push) Successful in 1m10s
A worker that exhausts its context dies mid-turn and its bridge_reply — the
whole point of the turn — dies with it. Auto-compaction is a condition of the
turn contract for a spawned peer, not an operator preference.

OPENCODE_CONFIG is merged over ~/.config/opencode/config.json rather than
replacing it, so workers already inherited auto:true from the home file. That
inheritance is exactly what this removes as a dependency: the home file is
outside this repo, differs per machine, and is not ours to rely on.

The trade is recorded in the code: an OPENCODE_CONFIG value overrides the home
value, so an operator cannot disable compaction for bridged workers from home.
Deliberate for peers we spawn and whose turns we must land; per-profile control
would be a profile knob, not the removal of this line.

Origin: fleet-wide auto-compaction audit by peer lead gpt-sol-5.6, which found
autoCompactEnabled:false alongside a 300k window across the Claude instance
configs. Handed over uncommitted; taken deliberately, with rationale added.
2026-08-13 21:03:52 +02:00
Dai Ha f55d3c203b Merge CB-552: supersede rejected orchestrator design; sync wiki Features
CI / build (push) Successful in 1m0s
CI / contract (push) Successful in 1m9s
2026-08-13 20:59:50 +02:00
Dai Ha 37c4d47d3a Merge CB-544: shutdown drain preserves worker worktrees
CI / build (push) Successful in 55s
CI / contract (push) Successful in 1m17s
2026-08-13 20:58:06 +02:00
Dai Ha 04e21c9243 CB-544: shutdown drain preserves worker worktrees (no data loss)
CI / contract (pull_request) Successful in 41s
CI / build (pull_request) Successful in 1m29s
2026-08-13 20:36:32 +02:00
Dai Ha 83cac07f6e CB-552: update wiki documentation pointer 2026-08-13 20:19:46 +02:00
Dai Ha f472e0f782 CB-552: supersede rejected orchestrator design 2026-08-13 20:19:46 +02:00
ltms 91c9f981c5 Merge CB-548 send ownership and rendezvous hardening
CI / build (push) Successful in 1m16s
CI / contract (push) Successful in 1m17s
2026-08-13 19:44:02 +02:00
Dai Ha dd906526c0 CB-548: guard primary singleton to PRIMARY callers; open waiter before enqueue
CI / contract (pull_request) Successful in 42s
CI / build (pull_request) Successful in 53s
Fix two handler-level bugs found in PR #21:
- Only PRIMARY callers may update PrimaryRegistry.record (the legacy singleton 'primary'
  fallback for no-delegation inbox nudges). An architect SEND previously recorded its terminal
  as the fallback; the per-target delegation map does not cure the singleton. New
  BridgeMcp.recordPrimarySingleton uses the resolved role (caller.isPrimary()) — named leads
  (PRIMARY) still record, architects never do.
- MessageService.send now opens the rendezvous waiter BEFORE queueing delivery, fixing both the
  enqueue-before-open fast-reply race (a fast reply no longer orphans into the inbox) and
  callback-failure ordering: a throwing onAccepted (public callback) fails the send cleanly with
  no stale waiter and no queued, orphanable message.

Tests: architect SEND vs lead SEND primary-singleton regression; throwing onAccepted leaves no
stale waiter or queued orphan.
2026-08-13 19:42:48 +02:00
Dai Ha ec3001796a CB-548: record delegator ownership only when a send is accepted
Record PrimaryRegistry delegator ownership via a MessageService accepted-delivery
hook (won the session lock + queued delivery), never at bridge_send request time, so
a concurrent sender that times out BUSY cannot steal a live turn's reply routing.
Make Rendezvous.open atomic fail-if-present so a double open trips loudly instead of
replacing the waiter another send is blocked on. Answering a bridge_ask keeps the same
ownership (no rewrite). Adds ownership/rendezvous regression tests.
2026-08-13 19:42:48 +02:00
ltms 86cf4c285a CB-547b: opencode session identity — resume (-s) + post-hoc discovery (#18)
CI / build (push) Successful in 48s
CI / contract (push) Successful in 1m17s
2026-08-13 19:41:06 +02:00
Dai Ha cd18887b69 CB-547b: opencode session identity — resume (-s) + post-hoc session discovery
CI / contract (pull_request) Successful in 42s
CI / build (pull_request) Successful in 1m16s
2026-08-13 19:40:17 +02:00
ltms 4637c68295 CB-543: neutralize worktree-hostile tracked configs (#22)
CI / build (push) Successful in 55s
CI / contract (push) Successful in 1m14s
2026-08-13 19:39:54 +02:00
Dai Ha 0114bd1fa7 CB-543: neutralize tracked opencode.json and .autoenv in provisioned worktrees
CI / contract (pull_request) Successful in 45s
CI / build (pull_request) Successful in 1m26s
2026-08-13 19:30:19 +02:00
ltms 6dee84ca71 Merge CB-548 architect config and authorization foundation
CI / build (push) Successful in 57s
CI / contract (push) Successful in 1m17s
2026-08-13 18:33:11 +02:00
Dai Ha f004a0c654 CB-548: make duplicate-architect-slot detection top-level-only and depth-safe
CI / contract (pull_request) Successful in 42s
CI / build (pull_request) Successful in 1m13s
2026-08-13 18:31:29 +02:00
Dai Ha 6123576c68 CB-548: correct architect premise — profile-only slots, registry-owned bindings, dup-key rejection
CI / build (pull_request) Successful in 54s
CI / contract (pull_request) Successful in 1m6s
2026-08-13 18:05:53 +02:00
Dai Ha 21cfc09f8e CB-548: config-declared architect slots + Role.ARCHITECT authz
CI / build (pull_request) Successful in 55s
CI / contract (pull_request) Successful in 1m6s
2026-08-13 17:28:14 +02:00
ltms 509530e235 CB-547a: peer-neutral session-identity contract + Claude Code adapter (#17)
CI / build (push) Successful in 1m0s
CI / contract (push) Successful in 1m2s
2026-08-13 16:56:32 +02:00
Dai Ha d146a01422 CB-547a: peer-neutral session-identity contract + Claude Code adapter
CI / contract (pull_request) Successful in 44s
CI / build (pull_request) Successful in 1m31s
2026-08-13 16:43:47 +02:00
Dai Ha 91332723db Bump wiki to 8c63db5 — the CB-539/CB-542 docs
CI / build (push) Successful in 1m18s
CI / contract (push) Successful in 1m52s
Pairs with 244fbd9. The docs half lived on the submodule's own remote, so the
pointer bump is separate by necessity, not by preference.

The substantive part is not the new 'Run a worker on the subscription' entry but
the correction beside it: 'Give workers a toolchain' asserted flatly that an env:
entry cannot repoint a worker past the SubscriptionGuard. CB-539 made that false,
and the wiki went on claiming it until CB-542 closed the hole. The entry now
states the rule and its one exception together.
2026-08-13 15:32:48 +02:00
ltms 244fbd98a5 Merge CB-539 + CB-542: subscription profiles, with the env: bypass closed
CI / contract (push) Successful in 40s
CI / build (push) Successful in 50s
Verified by the lead in a clean worktree at 5afe8e1 rather than on the worker's
report: Tests run: 474, Failures: 0, Errors: 0, Skipped: 0, BUILD SUCCESS (main
was at 464).

The invariant this lands: there is no configuration in which a worker reaches an
Anthropic endpoint that no guard vetted. Closed at two layers — a fatal, profile-
naming refusal at config load, and a launcher-side strip so it holds for profiles
built in code that never passed validation.

The wiki half is a separate branch on the submodule's own remote; the pointer bump
follows as its own commit on main.
2026-08-13 15:31:51 +02:00
Dai Ha 5afe8e14d9 CB-542: close the subscription env: bypass, and fix the env: env-carries-anthropic docs
CI / contract (pull_request) Successful in 44s
CI / build (pull_request) Successful in 51s
2026-08-13 15:26:20 +02:00
Dai Ha 73f6b12dd8 CB-539: let a worker profile run on the subscription, explicitly
Add a per-profile subscription: true opt-in that lets a claude-code worker run on
the operator's Claude subscription when there is no off-subscription endpoint for it
(e.g. sonnet on ccs). When set, the launcher injects neither ANTHROPIC_BASE_URL nor
ANTHROPIC_AUTH_TOKEN and skips SubscriptionGuard's base_url requirement for that
profile only, logging a WARN naming the profile. subscription: true alongside a
baseUrl is refused as contradictory. The default (absent/false) keeps today's hard
refusal unchanged; every other profile stays allowlist-checked and SubscriptionGuard
is untouched.
2026-08-13 15:15:14 +02:00
Dai Ha 793f2e7157 Revert "Commit .autoenv" — it wedges every worker spawn
CI / build (push) Successful in 53s
CI / contract (push) Successful in 1m24s
Regression I introduced one commit ago, caught by two consecutive spawn failures
and reproduced end-to-end.

Tracking `.autoenv` means git checks it out into every provisioned worktree. A
worktree is a new path, and autoenv authorizes by path, so the file is always
unauthorized there — autoenv prints "[autoenv] Authorize this file? (y/n/d)" and
blocks on `read` (activate.sh:211-222). The pane's shell sits at that prompt, so
`ccs <profile>` never runs, the peer never becomes injectable, and CB-306's
readiness gate closes the pane after 20s. Symptom is a bare "spawn timed out";
nothing names autoenv, which is what made it worth writing down.

Confirmed rather than inferred: the peer lead's cb-537 worker, spawned BEFORE
f8182e4, has no `.autoenv` in its worktree and is still alive; a throwaway
worktree at HEAD reproduces the prompt on entry.

The file is still worth committing — the reasoning in f8182e4 stands, and it is
recoverable from there. What is missing is the other half: GitWorktrees already
neutralizes `.mcp.json` in a provisioned worktree ("worker tool surface is
launcher-mounted only"), and `.autoenv` needs exactly the same treatment for
exactly the same reason — a worker's environment is launcher-supplied, never
repo-supplied. Re-land it with that, tracked as CB-543.

Rejected the quicker fix of exporting AUTOENV_ASSUME_YES: it auto-approves
arbitrary repo-controlled shell code in every spawned worker, which is a worse
trade than one unset variable.
2026-08-13 15:07:59 +02:00
Dai Ha cf48983046 Merge CB-538: opencode peers launch with --auto
CI / contract (push) Successful in 50s
CI / build (push) Successful in 1m30s
A spawned peer has no human at its pane, so an approval prompt is not a pause —
it is a wedge. The agent stops, looks identical to a legitimate mid-turn wait,
and can never reach its bridge_reply, so the delegation dies silently and the
lead learns nothing until the timeout.

Unconditional rather than a per-profile knob, which is the right call: there is
no configuration under which a bridge-spawned opencode worker WANTS to block on
an approval it has no way to answer. opencode's help calls --auto 'dangerous!',
and that warning is written for a human at a terminal; the blast radius here is
already bounded by the layer above — a worker runs in its own git worktree, on
its own branch, off-subscription, and cannot merge. The lead is the gate.

Reviewed by me rather than fanned out: 24 lines across one method and its two
tests, below the threshold where a reviewer pass pays for itself.

argvWithModel is re-signatured to take composed argv instead of building it,
so the two flag-appenders compose rather than each owning construction.
2026-08-13 15:03:59 +02:00
Dai Ha d038516750 Bump wiki to 0a21b49 — multi-lead arc, and opencode's secret supply
CI / build (push) Successful in 53s
CI / contract (push) Successful in 1m48s
Moves the submodule pointer 4320c1c -> 0a21b49, matching the docs to the code in
bc13b8e rather than leaving chapter 11 describing a fleet with one lead in it.

  7b5381b  11, 7: CB-530..536 — leaders:, leadScan:, lead-to-lead messaging,
           lead deliverability, leads in bridge_list; 7's copy of the portable
           block re-synced byte-identically
  0a21b49  12: where opencode's secrets actually come from, and why {env:...}
           is green under Claude Code and broken in a terminal

Both are already pushed to the wiki's own remote, so this pointer resolves for
anyone who clones — the ordering that matters, and the reason the wiki went
first.

This is the pointer only. `wiki/` is a submodule with its own remote and its
content is never committed here; bumping the recorded SHA in its own commit is
how this repo has always recorded a docs update (see "bump wiki to 4320c1c").
2026-08-13 10:44:38 +02:00
Dai Ha f8182e4514 Commit .autoenv — the loader that keeps Claude Code on one secret store
This file has always been designed to be committed and says so in its own header;
it simply never was, so every clone of this workspace has been reconstructing it
by hand or duplicating tokens instead.

It holds no secret. It reads `.secrets/` (gitignored, 0600) and exports three
variables, because Claude Code expands `${VAR}` in `.mcp.json` from the *process*
environment and cannot read a file — so without it, CONTEXT7_TOKEN and the gitea
pair must be duplicated as literals in `.claude/settings.local.json`. opencode
needs none of this: it reads `.secrets/` directly via `{file:...}`.

Verified before committing that no value appears in it, only the three names and
the paths they are read from.

Safe in a worktree by construction: a worktree receives tracked files only, so
`.secrets/` is absent there and the whole block is skipped rather than failing.
Workers are fed by the launcher's env instead — which is where the name mismatch
documented in wiki chapter 12 (GITEA_TOKEN vs GITEA_ACCESS_TOKEN) has to be
reconciled.
2026-08-13 10:44:21 +02:00
Dai Ha bc13b8e92c CB-530..536, CB-538 groundwork: leads as peers, and a fleet that can find itself
Two leads now work as peers rather than one primary plus workers. The arc:

  CB-530/531  lead identity: `leaders:` names panes, `leadScan:` discovers them by
              tab label (LeadTabScanner, TTL-cached, worker spaces excluded).
  CB-532      leads can message each other AND be answered. Principal.leader now
              carries its terminal, so ownsSession() can be true for a lead; the
              "and you must be a worker" conjunct beside it protected nothing.
              Retires `primary:` — reply nudges follow the delegating lead, a
              binding recorded at bridge_send where both halves are known.
  CB-533      ClaudeCodeLauncher passes --model. argv is usually a wrapper
              (`ccs <profile>`) that re-exports its own model family, so
              ANTHROPIC_MODEL alone was silently overruled.
  CB-534      a lead is deliverable. The CB-113 readiness gate only opened for
              terminals in WorkerPresence, which only workers ever enter, so every
              lead->lead send waited out the ~60s grace and failed having never
              been typed. The gate guards a *spawned* peer's boot window; a lead
              is never spawned.
  CB-535      bridge_list returns `leads` alongside `workers`, with `self` on the
              caller's row. An empty worker roster no longer reads as "no peers".
  CB-536      CLAUDE.md: lead<->lead is coordinate-only, never sideways delegation.
              Propagated byte-identically to wiki/7-Use-Cases.md.

MIXED PROVENANCE — recorded deliberately rather than hidden. This tree also carries
in-progress CB-537 (context separation) authored by the peer lead gpt-sol-5.6 and
its worker: Capability.CONTEXT_RESET, SessionManager.clearAfterTurn, and the
Injector/TurnListener/CompletionResolver/launcher changes around it. That work was
done in this shared working tree rather than a worktree, and is entangled with the
above in BridgedConfig.java, Bridged.java and ClaudeCodeLauncher.java, so neither
lead could stage its own half without sweeping in the other's. Committing the whole
green state is the honest resolution; the peer branches from here.

Note for whoever picks CB-537 up: the design in this commit is SUPERSEDED. Both
leads agreed to replace the global `clearAfterTurn` boolean with per-delivery
policy (inherit|fresh|thread) applied PRE-delivery, because a post-turn reset races
by construction — Injector.onStatus clears awaitingCompletion and dequeues the next
message in the same tick. `fresh` is also a correctness guarantee, so an adapter
without a reset capability must refuse it rather than log a no-op.

mvn clean install: Tests run: 464, Failures: 0, Errors: 0, Skipped: 0. BUILD SUCCESS.
2026-08-13 09:38:24 +02:00
Dai Ha fbe79258bf CB-538: opencode peers launch with --auto 2026-08-13 07:19:19 +02:00
Dai Ha ef1e014b41 CB-527: ship the bridge as an installable Claude Code plugin
CI / build (push) Successful in 1m0s
CI / contract (push) Successful in 1m21s
The orchestration contract had no distributable form. Every consuming project
hand-copied a block of CLAUDE.md and hand-wrote an .mcp.json, and we keep a
script whose only job is to notice those copies drifting apart. A plugin is
versioned, installed once, and updates in place.

Ships no credentials, deliberately: every secret is referenced by environment
variable NAME and the value never enters a file, which is what makes the
artifact safe to publish. The setup skill states the two rules that are easy to
get wrong for the right-sounding reasons — the PR token must not be able to
merge (a worker opens, the primary gates), and ANTHROPIC_BASE_URL must never be
set by setup, because mounting the bridge must not move a session off
subscription.

The plugin root is plugin/, not the repo root. An installed plugin's .mcp.json
is a committed file, while this repo's root .mcp.json is local-only and
--skip-worktree; rooting the plugin at the repo would commit the primary's IDE
servers and hand them to every worker — the exact failure CB-525 exists to
prevent.

Scope is client-side setup only. herdr and bridged stay separate services with
their own lifecycles, and the skill refuses to install them rather than guess.
It also refuses to accept /healthz as proof: health reports only that the daemon
can reach herdr, and CB-521 showed it staying green while every spawn failed, so
verification ends with a real spawn.

Both manifests pass `claude plugin validate --strict`.
2026-08-10 19:55:17 +02:00
Dai Ha e3c8393d1b CB-527: retire the ollama profile from the worker choice
The ollama backend is decommissioned, so the example config stops pointing
readers at a dead host and the guard allowlist stops carrying an entry with
no profile behind it — a stale entry there is dead permission, and that list
is the only thing keeping a worker off the primary's subscription.

The second illustrative profile survives as gx11: the example exists to show
`placement: weighted` having something to choose between, and a one-profile
example would quietly stop demonstrating that.

It also moves the CB-523 auto-compact override onto the surviving profile.
That guard had been attached to `ollama` alone, so retiring the profile would
have removed the fleet's only protection against the failure it was written
for — a worker whose prompt is rejected before auto-compact ever fires. The
window belongs on every profile, not on whichever one happened to hit it.
2026-08-10 19:55:17 +02:00
kevin 379e03f9d0 CB-308: fold in the adversarial review; bump wiki to 4320c1c
CI / build (push) Successful in 59s
CI / contract (push) Successful in 1m16s
Four new resolved decisions (§7.7-7.10): turn state split by where the signals
are, with ABANDONED explicitly belt-and-braces over the waiter timeout; the
dual ack model with spawn idempotence by construction (gid stored IN the herdr
pane — the load-bearing detail of the no-ledger position — plus an in-flight
reservation for redelivery during a slow spawn); enforced publish semantics
(confirms + mandatory on a separate channel, return-before-confirm caveat);
queue lifecycle = session lifecycle with .v2 names for the redeclare hazard.

§8 reworked: global id scheme resolved and moved up; control authorization
sharpened into the hard gate on U4 (per-host allowlist beside the peer keys);
key distribution/rotation added. New §9: implementation order, each step
verifiable single-host, U4 gated, U8 last.

Wiki pointer bumped to 4320c1c (chapter 10 same-pass changes).

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_013ZGgxLQ2VpwZhEYoru8rkf
2026-08-10 22:32:00 +07:00
kevin e13921aa8a CB-308: resolve the cross-host design review; bump wiki to 36bb865
CI / build (push) Successful in 54s
CI / contract (push) Successful in 1m5s
Six decisions recorded in §7, replacing the matching open questions: signed
messages (identity from the key, extending identity-from-connection across the
broker), target-host-owned profiles advertised via presence, repo provisioning
by pinned forge clone, live-only asks with TTL + TOO_LATE notice, spawn-id
dedup on the target, and broker-outage semantics (local unaffected, remote
fails fast, gateway stays soft-state). §8 keeps what is genuinely still open,
with control *authorization* now separated from the resolved *authenticity*.

The broker-level half (U8 broadcast, exclusive consumers, inbox caps, TLS,
schema version, trace id) lands in wiki chapter 10 §10 — pointer bumped
(also picks up 1710a77, chapter 11 Features).

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_013ZGgxLQ2VpwZhEYoru8rkf
2026-08-10 21:54:30 +07:00
kevin 8a549d8610 Release 1.0.0 — one leader, one host, complete
CI / contract (push) Successful in 1m13s
CI / build (push) Successful in 1m36s
Bump bridged to 1.0.0 and add the release notes: the single-leader,
single-host scope is closed — gateway, lifecycle, two-way delivery,
pluggable peers, auth/authz/audit, supervision, CI. Cross-host
federation (CB-308) is the next major line.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_013ZGgxLQ2VpwZhEYoru8rkf
2026-08-10 20:58:06 +07:00
Dai Ha 84081b2bd8 CB-526: make a shipped capability undocumentable-by-accident
CI / contract (push) Successful in 49s
CI / build (push) Successful in 1m23s
CLAUDE.md already carries a mandatory before-done checklist ("the prompt is part
of the product"). It covered the instruction surface but not the operator-facing
one, and the result was measurable: CB-506 through CB-525 shipped without a
single wiki mention, while the Roadmap went on claiming Stage 5 was finished.

Discipline is what already failed, so this rides the existing gate rather than
adding a new habit to remember: one more row, firing when a change touches
anything an operator can use, configure, or observe. The row names where the
other two kinds of change go too (contracts to Implementation, coverage to the
Roadmap), so "nothing to document" is a decision the table makes rather than a
default you fall into.

Addendum-only — the canonical block is untouched and still byte-identical to the
wiki template (verified).
2026-08-09 19:38:34 +02:00
Dai Ha defe3365c4 CB-519: re-point pane assertions at the protocol-19 coordinate
CI / build (push) Successful in 1m22s
CI / contract (push) Successful in 2m32s
Rebase integration only, no behaviour change. CB-519's tests named the pane the
pre-protocol-19 fake produced (w9:pW_n); upstream's herdr 0.8.0 port creates the
pane through tab.create and starts the agent into it, so the fake now reports
w9:pRoot_n. Five assertions were therefore counting closes of a pane that never
existed and reading 0.

mvn clean install: Tests run: 399, Failures: 0, Errors: 0 — BUILD SUCCESS
2026-08-09 05:59:14 +02:00
Dai Ha 4e6201ecd1 CB-525: isolate a worker's tool surface to what its launcher mounts
A worker in a provisioned worktree was inheriting the primary's MCP servers by
two independent routes: the repo commits a .mcp.json declaring the IDE servers,
so a fresh checkout mounts them, and the default parity overlay then copied the
primary's own copy over the top.

Those servers are bound to the primary's IntelliJ project, so every path they
hand back points into the primary's checkout. A CB-523 worker made all 59 of its
edits there while running `mvn -f bridged/pom.xml` against its worktree — every
build it ran was of code that did not contain its changes, and it passed. The
worker's own `ls` of the file it had "edited" returned "No such file".

GitWorktrees now neutralizes .mcp.json at provisioning: an explicitly empty
server map, --skip-worktree'd when tracked so it never reads as pending work a
worker might commit. Unconditional, because the overlay was only half the leak.
The bridge itself is unaffected — it reaches a worker through the launcher's
--mcp-config flag, not the project file, so bridge_reply still works.

- BridgedConfig: .mcp.json out of the default parity overlay
- GitWorktrees: isolateToolSurface() on add(), with the rationale in javadoc
- GitWorktreesTest: 4 real-git acceptance tests (2 fail if the call is removed)
- implementer skill: work from $PWD, and quote a green unpiped `mvn clean
  install` from the worktree as the acceptance criterion

mvn clean install: Tests run: 392, Failures: 0, Errors: 0 — BUILD SUCCESS
2026-08-09 05:57:30 +02:00
Dai Ha ccf50f950e CB-524: make worker placement order deterministic across JVM runs
The weighted policy breaks an exact-weight tie on candidate list order
(WeightedRoundRobinPolicy picks the first candidate with a strictly greater
score), and that list comes from CompositePeerLauncher.candidates(), which
iterates profileConfigs. Both that map and BridgedConfig.workerProfiles() were
built with Map.copyOf, whose iteration order is salted per JVM run — so the
"in definition order" contract candidates() documents was not held.

Two consequences. In production, a config with equal weights (ollama 0.5 /
gx10 0.5) placed its first worker on a profile chosen at random on every daemon
restart. In the suite, CompositePeerLauncherTest.failoverRetriesNextCandidate-
WhenProfileIsUnreachable failed roughly one run in four, because whether
profile "a" was tried first depended on the salt.

Preserve definition order at every layer: unmodifiable LinkedHashMap for
workerProfiles(), profileConfigs, and byProfile (which also feeds the
user-visible bridge_profiles listing). The tests build profile maps with an
ordered helper rather than Map.of, which is salted for the same reason.

Guarded by a pair of tests declaring the same two profiles in opposite order
and asserting opposite first attempts, so any order-scrambling implementation
must fail one of them. Verified by mutation: reverting profileConfigs to
Map.copyOf fails 8/8 runs (6 caught by the original test, 2 only by the new
reversed-order one); with the fix, 10/10 fresh JVMs pass, 388 tests green.
2026-08-09 05:57:30 +02:00
Dai Ha a7f0211e2f CB-518: weighted placement policy for worker spawns 2026-08-09 05:57:30 +02:00
Dai Ha 049ce4828c CB-520: split ReplyInbox into explicit own/release and publish halves 2026-08-09 05:57:12 +02:00
Dai Ha c1173346ef CB-521: make the AMQP contract test runnable locally and in CI 2026-08-09 05:56:45 +02:00
Dai Ha 7ace184fe6 CB-519: make PeerHandle.id() a host-unique opaque UUID, decoupled from the herdr pane id 2026-08-09 05:56:45 +02:00
Dai Ha 54b314ace5 CB-522: let the primary run inside a herdr pane
Caller identity resolved any loopback PID that mapped to a herdr pane as a
WORKER, and PaneLocator scans every pane -- not just bridged-spawned ones. A
primary running inside a herdr pane therefore classified itself as a worker and
was refused SPAWN/SEND/STOP, i.e. every orchestration verb it exists to call.

The failure is self-locking: PrimaryRegistry only learns the primary's terminal
from bridge_send/bridge_spawn, the exact calls being refused, so the learned
value can never bootstrap. Only an operator-set pin breaks the cycle.

CallerResolver now consults primary.terminal from config *before* the pane
lookup. Deliberately the pinned value only, never the learned one -- the learned
terminal is populated by the callers this method is itself classifying, so
trusting it would be circular. Config is operator input, never network input, so
this widens no attack surface; bridge_whoami and the authz gate still share one
resolution.

Fixing that exposed a second, older bug. BridgeMcp's context extractor forwards
the caller's terminal into markPresent on every MCP call, documented as "no-op
for the primary (null terminal)". WorkerPresence.markPresent honours that, but
PresenceBridge overrides it and forwards the same null into SessionManager.
onReady -> transitionByTerminal -> findByTerminal, which called
terminalId.equals(...) unguarded. It only reached the scan once the registry was
non-empty, so the primary's first spawn succeeded and every later call NPE'd
with an HTTP 500 -- and it would have fired for ANY primary not living in a
herdr pane, pinned or not.

findByTerminal is now total. That covers onReady, onDelivered, onTurnComplete
and onTurnFailed at once; a null id could never match a registered session
anyway, so "no match" is the honest answer rather than taking down an unrelated
tool call.

Also drops two dead pass-throughs on CallerResolver (cwdForPid, tokenMode) that
IDE inspections flagged -- callers use ConnectionIdentity and BridgedConfig.Auth
directly.

The example config now states that primary.terminal is REQUIRED, not just a
push-loop optimisation, when the primary shares a herdr pane.

mvn clean install: 360 tests, 0 failures. Verified live: daemon restarted on
this jar, bridge_whoami reports primary, and four concurrent worktree spawns --
the exact shape that NPE'd -- now all succeed.
2026-08-09 05:55:12 +02:00
kevin 224b344445 CB-522: resolve the pinned primary.terminal pane as the primary, not a worker
CI / build (push) Successful in 2m54s
A primary running INSIDE a herdr pane was resolved as a worker by the
pane-match rule and refused every orchestration tool — the exact lockout
bridge_whoami surfaced on this deployment. The CB-307 primary.terminal pin
always claimed to replace connection-derived identity but only fed the push
loop; it now short-circuits CallerResolver ahead of the pane→worker rule
(the pane mapping is as unforgeable as a worker's, so no credential needed,
even in token mode). bridged.example.yaml documents the block.

Also guard the presence bridge against the primary's null terminal: the MCP
context extractor marks presence on every request, and the first genuine
primary contact NPEd into the SPAWNING→READY transition.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01SUTLvxtRPr2iT5u5g45BEs
2026-08-08 21:52:38 +07:00
kevin 0b28b4cb0f CB-521: port the herdr adapter to protocol 19 (herdr 0.8.0)
herdr 0.8.0 redesigned the agent API out from under the daemon: agent.start
now launches a supported kind INTO an existing pane, env/cwd move to pane
creation (tab.create / pane.split — the subscription-boundary seam now),
agent.send is replaced by agent.prompt (self-submitting) plus agent.send_keys
for the Enter nudge, and terminal ids are no longer valid agent.* targets.

- AgentControl: start(name, kind, args, paneId); prompt/send_keys delivery;
  cached terminal→pane target translation (invalidated on agent_not_found).
- WorkspaceControl: tab.create carries cwd+env; pane.split for legacy placement.
- HerdrPeerLauncher: the seed pane IS the worker pane (no drop step); retry
  agent.start while the seed shell boots (agent_pane_busy).
- FakeHerdr and the test suite model protocol 19 (unique seed panes, required
  kind/pane_id, prompt-based delivery); contract tests probe the seed shell
  instead of arbitrary-command agents, which protocol 19 removed.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01SUTLvxtRPr2iT5u5g45BEs
2026-08-08 21:52:25 +07:00
kevin 83129e165c Merge CB-518: state the primary's orchestration as an explicit, ordered flow
CI / build (push) Successful in 1m19s
Turns the primary's half of the bridge charter from a bullet list of
policies into a numbered 0-8 procedure, and splits delegated review out
of the merge step it used to sit beside. Wiki template kept byte-identical
by splicing; pointer bumped to 0c896eb in the same commit.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_017Kw1FosEt3Noix5GG9wJ2r
2026-08-04 22:37:17 +07:00
145 changed files with 22787 additions and 2751 deletions
+18
View File
@@ -0,0 +1,18 @@
{
"name": "claude-bridge",
"description": "Tooling for orchestrating a fleet of delegated coding agents through the bridged MCP gateway.",
"owner": {
"name": "LTMS"
},
"plugins": [
{
"name": "claude-bridge",
"source": "./plugin",
"description": "Make a project bridge-ready: mount the bridged MCP gateway and apply standard Claude Code settings so a session can orchestrate delegated workers. Ships no credentials.",
"version": "0.1.0",
"author": {
"name": "LTMS"
}
}
]
}
+37 -9
View File
@@ -9,11 +9,13 @@ The turn contract (one `bridge_reply`, `bridge_ask` for the lead's decisions, ho
never merge, never commit `.mcp.json` or `wiki/`) is in **`CLAUDE.md` → Bridge communication →
Worker** and already applies. This skill is only the *implement-and-hand-off procedure*.
You run in an **isolated git worktree on your own branch** — a full peer of the primary (same
repo, `CLAUDE.md`, skills, MCP), differing in the model behind you and the branch you sit on.
You run in an **isolated git worktree on your own branch** — a full peer of the primary (same repo,
`CLAUDE.md`, skills), differing in the model behind you and the branch you sit on. Your MCP surface
is **only what your launcher mounted** (the bridge): the primary's IDE and forge servers are not
yours, and the worktree's `.mcp.json` is deliberately emptied so you cannot inherit them.
The worktree model is documented in [`docs/Worker-Git-Workflow.md`](../../../docs/Worker-Git-Workflow.md).
## 1. Confirm where you are
## 1. Confirm where you are — then never leave
Before touching anything:
@@ -26,13 +28,36 @@ git status # should be clean at the start
Do **all** work here, on this branch. Never `git checkout main`, never rebase onto or push to
`main`. The branch is your isolation — respect it.
**Every path you read, edit, or build is relative to that root.** Work from `$PWD`; if a tool, a
brief, or your own memory hands you an absolute path, check it starts with your worktree root
before you touch it, and stop if it doesn't. An absolute path pointing anywhere else is the
primary's checkout — editing there while building here means **every build you run is of code that
does not contain your changes**, and it passes while your work goes nowhere. This has happened:
a worker made all 59 of its edits in the primary's tree and never noticed.
```bash
test "$(git rev-parse --show-toplevel)" = "$PWD" || cd "$(git rev-parse --show-toplevel)"
```
## 2. Implement
- Implement exactly the scope the lead named. Keep the diff focused; note anything out of scope
in your reply instead of widening it.
- Match the surrounding code's style, naming, and idioms.
- Run whatever build/test you can — `mvn clean install` from the module root. Read its **full**
output; a piped `mvn ... | tail` hides failures.
**Acceptance criterion — a green build, quoted.** Your work is not done until this passes *inside
your worktree*:
```bash
cd "$(git rev-parse --show-toplevel)/bridged" && mvn clean install
echo "exit=$?"
```
Read its **full** output — never pipe it through `tail`/`head`/`grep`, which hide a failure behind
a zero exit. Then quote the real `Tests run: … Failures: … Errors: …` line and the
`BUILD SUCCESS`/`FAILURE` verbatim in your reply. If it does not go green, say so with the actual
error; a failing build honestly reported is a usable result, a claimed-green one is not. You have
no IDE MCP tools, so `mvn` is your only verification — never claim a check you had no way to run.
## 3. Commit
@@ -41,7 +66,8 @@ git add <the files you changed> # explicitly — never `git add -A` / `git a
git commit -m "<ticket>: <clear one-line summary>"
```
`.mcp.json` will show as modified. Leave it — it is `--skip-worktree` and not yours to commit.
`.mcp.json` is neutralized and `--skip-worktree` in your worktree — never `git add` it, and never
"restore" it from the primary's copy. Same for `wiki/` (a submodule with its own remote).
## 4. Push
@@ -84,8 +110,9 @@ The reply is the entire handoff; the lead cannot see your terminal.
```
PR: <html_url from step 5, or "not created: <reason>" + branch name>
branch: <your branch>
files: <the files you changed>
tests: <what you ran and its REAL result — or "not run: <why>">
root: <git rev-parse --show-toplevel — proves you worked in your own worktree>
files: <worktree-relative paths you changed>
build: <the verbatim "Tests run: …" and BUILD SUCCESS/FAILURE lines — or "not run: <why>">
summary: <2-3 lines: what you implemented and any caveat the reviewer needs>
```
@@ -97,7 +124,8 @@ sequenceDiagram
participant G as git / gitea
L->>I: delegated task (you are in a worktree on your branch)
I->>I: implement + build/test here
I->>I: "implement here — every path under $PWD"
I->>I: "mvn clean install in this worktree, unpiped, until green"
I->>G: git commit (never .mcp.json / wiki)
I->>G: git push -u origin HEAD
I->>G: POST /pulls (GITEA_TOKEN) — open PR to main
+180
View File
@@ -0,0 +1,180 @@
---
name: port-to-opencode
description: Make an OpenCode session a first-class participant in a Claude Code workspace — instructions, MCP servers, and secrets — without duplicating config. Use when onboarding opencode to a project that already has CLAUDE.md and .mcp.json, or when an opencode peer needs the same tools and rules as the Claude session.
---
# Porting a Claude Code workspace to OpenCode
**The headline: there is almost nothing to port.** OpenCode reads `CLAUDE.md` natively. The only
artifact you create is one `opencode.json` mapping MCP servers. Do not translate instructions, do
not generate a second rules file, and do not install a sync tool — every one of those makes the
workspace worse.
Everything below was verified against `opencode 1.18.16` and the OpenCode docs.
## 1. Know what you get for free
OpenCode's instruction search order:
```
1. walking up from cwd: AGENTS.md , then CLAUDE.md
2. global: ~/.config/opencode/AGENTS.md
3. Claude Code global: ~/.claude/CLAUDE.md (unless disabled)
```
*"The first matching file wins in each category."*
Consequences that decide the whole procedure:
- **A project `CLAUDE.md` is already read.** No port needed.
- **Your user-level `~/.claude/CLAUDE.md` is already read too.** Global preferences carry over.
- **An `AGENTS.md` in the repo SHADOWS `CLAUDE.md`.** If one exists from a previous Codex port,
**delete it** — otherwise opencode reads the stale translated copy instead of the real rules.
This is the single most likely way to get this wrong.
## 2. Create `opencode.json` for MCP servers only
Project config lives at `opencode.json` in the repo root; the global one is
`~/.config/opencode/opencode.json`. **Configs merge, they do not replace** — so machine-local
servers belong in the global file and shared ones in the project file.
Map each entry from `.mcp.json`:
| `.mcp.json` | `opencode.json` |
|---|---|
| `"type": "http"` / `"sse"` | `"type": "remote"`, `"url"` |
| `"type": "stdio"` | `"type": "local"`, `"command": ["bin", "arg"]` |
| `"command"` + `"args"` | single `"command"` array |
| `"env"` | `"environment"` |
| `"headers"` | `"headers"` |
```json
{
"$schema": "https://opencode.ai/config.json",
"instructions": ["CLAUDE.md"],
"mcp": {
"bridged": { "type": "remote", "url": "http://127.0.0.1:8765/mcp", "enabled": true },
"context7": { "type": "remote", "url": "https://example.dev/mcp", "enabled": true,
"headers": { "Authorization": "Bearer {env:CONTEXT7_TOKEN}" } },
"gitea": { "type": "local", "command": ["gitea-mcp", "-t", "stdio"], "enabled": true,
"environment": { "GITEA_ACCESS_TOKEN": "{env:GITEA_ACCESS_TOKEN}" } }
}
}
```
Set `instructions` explicitly even though `CLAUDE.md` is found anyway — the fallback only applies
while no `AGENTS.md` exists, and being explicit survives someone adding one later.
## 3. Reference secrets, never embed them
OpenCode substitutes at load time, in both `headers` and `environment`:
```
{env:VARIABLE_NAME} value from the environment
{file:~/.secrets/token} value read from a file
```
**No credential ever belongs in `opencode.json`.** With `{env:…}` there is no reason to write one,
which is what makes this file safe to commit — and it must be committable, because a peer running
in a git worktree receives tracked files only.
**But `{env:…}` reads OpenCode's *process* environment — and OpenCode has no env store of its
own.** A Claude Code `env` block in `~/.claude/settings.json` or `.claude/settings.local.json` does
**not** reach it: those files are Claude Code's, and the variables exist only inside processes
Claude Code spawned. Verify with a clean login shell, not the shell your agent hands you:
```bash
env -u MY_TOKEN zsh -lc 'echo "${MY_TOKEN:-NOT IN PROFILE}"'
```
So `{env:…}` only works if something puts the variable in the environment first. Pick the supply
route by who launches opencode:
| Launcher | Route |
|---|---|
| a human, from a terminal | one central store, sourced by the login shell |
| a spawner (bridge, CI, IDE) | `{env:…}`, with the spawner injecting the variable |
**For the human case, keep every credential in one file the login shell sources.** Here that file
is `${SHARED_ENV}/tools/secrets.sh`, sourced from `${SHARED_ENV}/.ltms`, kept at mode 600 and never
committed. `opencode.json` then names variables and holds no values:
```json
"headers": { "Authorization": "Bearer {env:CONTEXT7_TOKEN}" }
```
Both routes end at the same syntax, and that is the point. The file does not change when a human
launches opencode instead of the bridge.
**`{file:…}` also works, and this project moved away from it.** Opencode resolves a relative
`{file:}` path against the project root, so a gitignored `.secrets/` beside `opencode.json` needs no
shell setup at all. It did not fail; the problem is that it makes a second copy of the token. The
same secret then lives in two places, and the copy you forget is the one that leaks or goes stale.
One store with many references is easier to rotate and to audit.
**One catch survives either choice, so state it out loud:** what a human's shell exports does not
reach a spawned peer, and neither does a gitignored `.secrets/` — a git worktree receives tracked
files only. Spawned peers must be fed through `{env:…}` by whatever launches them. Check the
variable *names* match: a spawner often injects under a different name than your shell uses, and
the config has no fallback. In this repo the bridge goes further: it neutralizes a worktree's
`opencode.json`, so a member cannot inherit the primary's credentials by accident.
## 4. Do not port machine-local MCP servers
IDE indexes, language servers, editor bridges — anything bound to *your* checkout — stay out of the
project file. Put them in `~/.config/opencode/opencode.json` if you want them personally.
A committed project config reaches every worktree. A worker that mounts servers whose paths point
into the primary's checkout will edit the primary's files while building in its own — every build
passes, every change lands in the wrong tree.
Port the servers the work needs. Leave the rest.
## 5. Verify against the running agent, not the file
A file on disk proves nothing about what the agent loaded.
```bash
opencode run "In one line: state a rule from this project's instructions."
```
The answer must reflect the actual `CLAUDE.md`. If it answers generically, the instructions did not
reach the model and everything after this is built on sand.
Then confirm the tools are mounted:
```bash
opencode mcp list
```
**"connected" does not mean "working".** A stdio server with a missing credential still completes
the MCP handshake and reports green; only a real tool call reveals it. Verified: `gitea` showed
`✓ connected` with no token, then failed the first call with `token is required`. A remote server
is more honest (`⚠ needs authentication`), but do not rely on that difference — **exercise one
authenticated tool per server**:
```bash
opencode run "Call <server>'s <tool>. Report the result or the exact error. One line."
```
Check too that no server you deliberately withheld is present.
## 6. Report
State what you changed, which servers crossed and which you withheld and why, and quote the
verification answer verbatim. If any server failed to connect, say so plainly — a partially mounted
peer is worse than a missing one, because it looks configured.
## Gotchas
- **A leftover `AGENTS.md` silently wins over `CLAUDE.md`.** Check for one before anything else.
- **`opencode.json` is merged, not overridden** — a global entry and a project entry with the same
server name both matter; keep names distinct unless you intend to layer them.
- **`enabled: false`** turns a server off without deleting its config — prefer it over removal when
you may want the server back.
- **OAuth-based servers** store tokens in `~/.local/share/opencode/mcp-auth.json` after
`opencode mcp auth <server>`; that is machine state, never config to commit.
- **Skills and subagents do not port.** OpenCode uses its own agent markdown under
`.opencode/agents/`; `.claude/skills/**` is not read. If a delegation brief tells a peer to load a
skill by name, that instruction has no effect on an opencode peer — spell the procedure out in the
brief, or author the equivalent agent file.
+52
View File
@@ -53,3 +53,55 @@ jobs:
grep -qE "Failures: [1-9]|Errors: [1-9]" "$f" && { echo "===== $f ====="; cat "$f"; }
done
exit 0
# CB-521 — actually run the AMQP contract test in CI, against a REAL broker. The broker is a
# RabbitMQ SERVICE CONTAINER, not Testcontainers-with-Docker: the runner image has no Docker, so
# AmqpReplyInboxContractTest reads AMQP_URI (set below to the service's network alias) and binds
# straight to it — no Docker, no skipped tests. This separation (build job hermetic and
# Docker-free; contract job broker-provided) is deliberate — see the default-excludes/contract
# profiles in bridged/pom.xml. `setup-java` provides the JDK only; Maven is installed separately,
# exactly as in the build job above.
contract:
runs-on: ubuntu-latest
services:
rabbitmq:
image: rabbitmq:3.13 # same AMQP 0-9-1 engine the local Testcontainers fixture uses
env:
RABBITMQ_DEFAULT_USER: guest
RABBITMQ_DEFAULT_PASS: guest
ports:
- 5672:5672
env:
# Service containers are reachable from the job by their network alias on their internal port.
AMQP_URI: amqp://guest:guest@rabbitmq:5672
steps:
- uses: actions/checkout@v4
- name: Set up JDK 25
uses: actions/setup-java@v4
with:
distribution: temurin
java-version: '25'
cache: maven
- name: Install Maven
run: |
apt-get update && apt-get install -y --no-install-recommends maven
mvn -version
# The `contract` profile clears the default-excludes group, so the @Tag("contract") AMQP test
# runs against the RabbitMQ service container (AMQP_URI). Pinned to the one contract test to
# avoid re-running the unit suite already covered by the `build` job.
- name: Contract tests
working-directory: bridged
run: mvn -B -Pcontract test -Dtest=AmqpReplyInboxContractTest
- name: Failing test output
if: failure()
working-directory: bridged
run: |
for f in target/surefire-reports/*.txt; do
[ -f "$f" ] || continue
grep -qE "Failures: [1-9]|Errors: [1-9]" "$f" && { echo "===== $f ====="; cat "$f"; }
done
exit 0
+21
View File
@@ -0,0 +1,21 @@
# No secret belongs in this repo any more: every credential lives in one shell-level store
# (${SHARED_ENV}/tools/secrets.sh), and opencode.json reads it as {env:...}. This line stays as a
# backstop, so a workspace-scoped copy that someone re-creates by habit still cannot be committed.
.secrets/
# Settings backups inherit the env block — and secrets with it.
.claude/settings.local.json.bak*
# The default profile parityOverlay copies these primary→worktree, so they appear in EVERY worker
# worktree. Two reasons they must be ignored. They hold environment values, which is reason enough.
# And since CB-576 a release preserves any worktree that `git status --porcelain` calls dirty —
# untracked files included, deliberately. An untracked overlay file would therefore make every
# COMPLETED release preserve its worktree, and worktrees would pile up with no error to notice.
.env
.envrc
# Daemon runtime artefacts. bridged appends its log wherever it is launched from, so both the
# repo root and bridged/ collect one; neither belongs in git.
bridged.out
bridged/bridged.out
logs/
+141 -30
View File
@@ -11,40 +11,45 @@
If no `bridge_*` MCP tools are mounted in this session, this section does not apply — skip it.
`bridged` is the **sole communication gateway** between agents here. The orchestrating session (the
**primary**) and every delegated peer (a **worker**) mount the *same* MCP server and talk only
**primary**) and every delegated peer (a **member**) mount the *same* MCP server and talk only
through its `bridge_*` tools. No session addresses a peer, a broker, or the network directly.
### Which role am I? — settle this before acting
**Both roles read this file.** A worker runs in a git worktree of this same repo, so it inherits
**Every role reads this file.** A member runs in a git worktree of this same repo, so it inherits
this `CLAUDE.md` verbatim, and every rule below is role-conditional.
**Call `bridge_whoami`.** It returns `{"role":"primary"}` or `{"role":"worker","sessionId":…,
"profile":…,"worktree":…,"branch":…}`, resolved by the daemon from your connection — unforgeable,
and the same resolution its authorization gate uses. Don't infer what you can ask.
**Call `bridge_whoami`.** It returns `primary`, `worker`, or `architect`, resolved by the daemon from
your connection — unforgeable, and the same resolution its authorization gate uses. A worker also
carries its `sessionId`, `profile`, `worktree` and `branch`; an architect carries the slot name it
was bound to. Don't infer what you can ask.
Only if that call is unavailable, fall back to these — each is one-way, so keep reading until one
fires: the reply charter in your system prompt (*"You are an off-subscription worker in the
claude-bridge fleet"*) ⇒ **worker**; bridge tools prefixed `mcp__bridge__*` ⇒ **worker** (the
launcher fixes that mount name; a primary's mount is named by whoever wrote its `.mcp.json`, so it
varies); `ANTHROPIC_BASE_URL` set ⇒ **worker** (Claude-model workers run on a clean env, so its
*absence* proves nothing). **Still unsure ⇒ act as a worker.** The two mistakes are not symmetric: a
primary acting as a worker is refused by the authorization gate — loud and self-correcting — while a
worker acting as the primary ends its turn with no `bridge_reply`, and the sender silently receives
nothing. Fail toward the recoverable error.
claude-bridge fleet"*) ⇒ **spawned member**; bridge tools prefixed `mcp__bridge__*` ⇒ **spawned
member** (the launcher fixes that mount name; a primary's mount is named by whoever wrote its
`.mcp.json`, so it varies); `ANTHROPIC_BASE_URL` set ⇒ **spawned member** (Claude-model members run
on a clean env, so its *absence* proves nothing). None of these separate a worker from an architect —
only `bridge_whoami` does. **Still unsure ⇒ act as a worker**, the most restricted member role. The
two mistakes are not symmetric: a primary acting as a worker is refused by the authorization gate —
loud and self-correcting — while a member acting as the primary ends its turn with no `bridge_reply`,
and the sender silently receives nothing. Fail toward the recoverable error.
### Invariants — both roles, no exceptions
1. **Never set, export, or forward `ANTHROPIC_BASE_URL`** (or `ANTHROPIC_AUTH_TOKEN`). The primary
stays on subscription; only the bridge puts a worker off it, at spawn. Mounting the bridge must
stays on subscription; only the bridge puts a member off it, at spawn. Mounting the bridge must
never move a session across that boundary.
2. **The bridge is the only channel.** Text you print in your terminal reaches nobody — the other
side cannot see your screen. An answer that isn't in a `bridge_*` call is silently discarded.
3. **Identity comes from the connection, never an argument.** Workers never pass a target; you
cannot act as another session. Spawn/stop/send/drain are primary-only; reply/ask are
worker-only-and-only-as-itself. A call outside your role is refused, not queued.
cannot act as another session. Spawn/stop/drain are lead-only; **send is lead or architect**;
reply/ask are only-as-itself — any peer may answer for its own pane, and for no other. A call
outside your role is refused, not queued.
4. **Delivery is status-gated: one message per turn.** Don't busy-poll a peer's terminal and don't
re-send because a call looks slow — the bridge delivers when the peer is `idle`/`blocked`.
re-send because a call looks slow — the bridge delivers when the peer is `idle`, `blocked` or
`done`. A spawned member must **also** have mounted the bridge MCP: until it has, it is not
deliverable, and a send waits on that gate for ~60s and then fails without ever reaching its pane.
5. **Never drive the terminal multiplexer directly** (no `herdr` CLI, no socket). The bridge owns
policy; the multiplexer owns PTYs. Going around the bridge bypasses every rule above.
@@ -76,9 +81,9 @@ below are the procedure — run them in order, every task, not only the big ones
5. **Collect** — `bridge_poll{ticket}` → `bridge_ack{ticket, msgId}`. Answer a worker's `bridge_ask`
with `bridge_send{turnId, content}` — **not** `sessionId`. A worker gone quiet is diagnosed with
`bridge_status`, never by reading its terminal.
6. **Verify yourself.** Re-run the build and the checks. A worker mounts only the bridge MCP and
cannot run your other tooling, and a piped command (`… | tail`) hides failures behind a zero
exit — never promote a worker's "clean" to a fact.
6. **Verify yourself.** Re-run the build and the checks. A worker cannot run your IDE tooling, any
forge tools it appears to have hold a blocked credential and fail, and a piped command
(`… | tail`) hides failures behind a zero exit — never promote a worker's "clean" to a fact.
7. **Review — fan out.** Spawn reviewers against the diff, one per dimension or per file, with
`wait:false`. Never the implementer of the scope it reviews, and brief them from the diff — not
from the implementer's rationale, which carries its own blind spot. Dispatch each PR's reviewers
@@ -100,15 +105,44 @@ the merge — and merging on a reviewer's word is delegating it by proxy.
|---|---|
| Confirm your own role | `bridge_whoami` |
| See backends available | `bridge_profiles` |
| Start a worker | `bridge_spawn{profile?, cwd?, worktree?, ticket?}` → `sessionId` + `paneId` |
| See the fleet | `bridge_list` · one worker's state: `bridge_status{sessionId}` |
| Start a member | `bridge_spawn{role?, profile?, cwd?, worktree?, ticket?}` → `sessionId` + `paneId` |
| See the fleet | `bridge_list` → `leads` (your peers) + `members` · one peer's state: `bridge_status{sessionId}` |
| Delegate (blocking) | `bridge_send{sessionId, content}` |
| Delegate (long task) | `bridge_send{sessionId, content, wait:false}` → ticket → `bridge_poll{ticket}` |
| Answer a worker's `bridge_ask` | `bridge_send{turnId, content}` — **not** `sessionId` |
| Answer a member's `bridge_ask` | `bridge_send{turnId, content}` — **not** `sessionId` |
| Message a **peer lead** | `bridge_send{sessionId: <their terminal>, content}` — `bridge_list` → `leads` reports it. Coordination only, **never** a task |
| Answer a peer lead that messaged you | `bridge_reply{content}` — the one case a lead replies |
| Collect a held reply | `bridge_poll{target}` · then `bridge_ack{target, msgId}` |
| Tear down | `bridge_stop{paneId}` |
| Tear down a member | `bridge_stop{paneId}` |
### Worker — the turn contract
### Lead ↔ lead — coordinate, never delegate
`bridge_list` returns `leads` alongside `members`; your own row carries `self: true`. Every other row
is a peer — an orchestrator with its own context, its own members, and its own judgment. An empty
`members` array means no members are spawned; it says nothing about peers.
**A lead never assigns a task to another lead.** Work goes to members — only ever downward, never
sideways. Sending a peer a brief with acceptance criteria is a category error: a brief is a member's
artefact, and a peer is not yours to task. If a unit needs doing and it falls in your area, spawn a
member and delegate it yourself; if it falls in the peer's area, say so and let the peer assign it.
The traffic between leads is coordination and nothing else:
1. **Divide the map, not the work.** Agree who owns which area, then each of you assigns inside your
own. Split by **context ownership** — whoever already holds the context owns that area — and say
who takes what, in one message, before either of you starts. Two leads silently working the same
unit is the failure mode here, and neither notices until the merge.
2. **Share findings, hazards, and corrections.** What you have already discovered, what broke, what
the next person will trip on. This is the traffic that actually pays for the channel: it costs one
message and saves a peer a rediscovery.
3. **Verify a peer exactly as you verify yourself.** Peer status buys nothing: check the claim
against the code, and re-run the build. A peer's correction gets the same treatment — right or
wrong on the evidence, not on who said it. Neither of you merges the other's work unreviewed.
Being messaged by a peer does not make you its worker: answer with `bridge_reply`, and push back on
the substance if it is wrong. A peer that simply complies has thrown away the reason there are two of
you.
### Member (worker or architect) — the turn contract
1. **Load the playbook skill the lead named** before doing anything else.
2. **Do the assigned scope only.** Note anything you spot outside it in one line; don't go hunt it.
@@ -117,9 +151,17 @@ the merge — and merging on a reviewer's word is delegating it by proxy.
Don't ask what you could decide yourself.
4. **End the turn with exactly one `bridge_reply{content}`**, carrying your complete answer. This is
the whole handoff. No `bridge_reply` ⇒ the sender gets nothing and the exchange stalls.
5. **Report honestly.** State only what you actually ran and its real output, including failures.
You mount **only** the bridge MCP — the primary's other servers (IDE, forge, docs) are not yours,
so never claim the result of a check you had no way to run.
Do **not** lean on the completion fallback to carry your answer for you: when you end a turn
without replying, the bridge scrapes your pane, and it can return only the last 4000 characters.
A clipped scrape is marked as partial, but the missing text is gone — your report reaches the
lead with its end cut off.
5. **Report honestly.** State only what you actually ran and its real output, including failures,
and never claim the result of a check you had no way to run. **Measure your own tools; do not
assume them.** What you mount depends on your backend: an opencode member gets the bridge and
nothing else, while a Claude Code member also inherits the operator's user-scope MCP servers,
which the bridge never chose for you. Two rules follow. The primary's IDE tooling is still not
yours, whatever you see. And **a mounted tool is not a working tool** — the forge server you may
find there holds a deliberately blocked credential and fails every call, by design.
6. **Never merge.** Stage files explicitly — never `git add -A` — and leave alone anything the
project marks as not-yours-to-commit.
@@ -127,9 +169,9 @@ the merge — and merging on a reviewer's word is delegating it by proxy.
| Layer | Scope | Reaches |
|---|---|---|
| the launcher's reply charter | the one rule that must survive with no repo: *end every turn with `bridge_reply`* | every worker, at launch, every peer kind |
| **this section** | protocol + orchestration policy | primary **and** every Claude worker — tracked in git, so worktrees inherit it |
| role playbook skills | per-job procedure (commit/PR recipe, finding format) | a worker told to load one |
| the launcher's reply charter | the one rule that must survive with no repo: *end every turn with `bridge_reply`* | every spawned member, at launch, every peer kind — never a lead |
| **this section** | protocol + orchestration policy | primary **and** every member that reads the repo — tracked in git, so worktrees inherit it |
| role playbook skills | per-job procedure (commit/PR recipe, finding format) | a member told to load one |
| the bridge's own docs | design detail, flows, error model | on demand |
A rule belongs in **exactly one** layer — the outermost one that must obey it. Peers that don't read
@@ -143,12 +185,72 @@ charter, not here.
`mcp/ConnectionIdentity` (connection→role), and `worker/*Launcher` (`REPLY_CHARTER`).
- **Skills available to delegate:** `implementer` (worktree → commit → push → own PR) and
`reviewer` (scoped review → one structured finding). Name one in every delegation.
- **Primary-side skills** (not delegation playbooks — a worker cannot use them):
`port-to-opencode` (make an OpenCode session a participant in this workspace).
- **Never commit** `.mcp.json` (the primary's local copy, flagged `--skip-worktree`) or `wiki/`
(a submodule with its own remote).
- **Flows and the error model** — rendezvous, `bridge_ask`, detached delivery, turn-done fallback —
are diagrammed in `docs/MCP-Contract.md` §6, kept out of this file because it loads into every
session's context.
### Redeploying the daemon — the lead may do this (primary only)
**A merge is not a deployment.** The running `bridged` holds the jar it was started with, so a
feature merged to `main` does nothing until the daemon is rebuilt and restarted. Saying "shipped"
about code the live daemon has never loaded is a false report. The lead **may and should** redeploy
rather than hand the job back to the operator.
Workers must never do this. A worker has no business restarting the daemon it is talking through,
and stopping it kills the worker's own channel mid-turn.
**Use the script — do not hand-roll the steps.**
```bash
scripts/redeploy-bridged.sh --check # report state, change nothing
scripts/redeploy-bridged.sh # build, confirm drain, restart, verify
scripts/redeploy-bridged.sh --yes # skip the drain prompt (fleet already checked)
```
It builds before it stops anything, so a failed build never leaves the fleet down; it waits for the
old process to exit rather than assuming; it polls `/healthz`; and it anchors its log checks to a
line marker taken before the restart, so old errors cannot be misread as new ones. Run `--check`
first — it is read-only and reports whether the forge token resolves, which nothing else tells you.
The script encodes the five things below, each of which has gone wrong here before. Read them anyway:
if the script is unavailable or a step fails, this is what it was protecting you from.
1. **Login shell, or workers silently lose their forge token.** The daemon inherits
`WORKER_GITEA_TOKEN` from the shell that starts it, and that comes from
`${SHARED_ENV}/tools/secrets.sh`. Start it from a non-login shell and the variable is empty, the
daemon starts fine, and the failure appears much later as workers that cannot open a PR. Nothing
logs this at startup — the script's `--check` is the only thing that reports it, and it checks
whether the name resolves without ever printing the value.
2. **Drain live members first.** `bridge_list`, then `bridge_stop` each member, and collect anything
you still want with `bridge_poll` before you kill anything. A restart drops in-flight tickets and
rendezvous, and a member's report is not recoverable once its ticket is gone.
3. **A restart is the only way deferred config keys take effect.** That is usually the reason to do
it. The startup log names which keys it accepted and which it deferred — read those lines rather
than assuming.
4. **Re-check identity afterwards.** Call `bridge_whoami` and confirm it still answers `primary`. The
lead is found by its tab label (`fleet.leaders.*.tab`), and a lead whose tab no longer matches is
demoted to worker, which refuses every orchestration call.
5. **Prove the new jar is the one running.** Confirm a *fresh* `bridged listening` line at the end of
`bridged/bridged.out`, dated after the restart. An old daemon that never died looks identical from
the outside.
**Permission.** A `CLAUDE.md` rule grants intent, not tool permission — the command classifier
refuses a bare `kill` on the daemon whatever this file says. The script is the seam that fixes that:
it is one auditable command, so the operator allow-lists it once instead of approving a stop and a
start every time. The rule lives in the operator's Claude Code settings:
```json
{ "permissions": { "allow": ["Bash(scripts/redeploy-bridged.sh:*)"] } }
```
Granted by the operator on 2026-08-15. If a call is still refused, do **not** route around it by
running the stop and start as separate commands — that is exactly the approval the script replaced.
Say what you were going to run and why, and let the operator decide.
### The prompt is part of the product — update it with the code (mandatory)
This repo *is* the bridge, so the canonical block above is not documentation about someone else's
@@ -168,6 +270,15 @@ Before you call any work done, check the row that matches what you touched:
| worktree provisioning or the parity overlay | the "both roles read this file" premise — it rests on the worker's worktree being a checkout of this repo |
| `.claude/skills/**` | the addendum's skill list, and the "name the playbook" rule |
| a new peer kind (non-Claude adapter) | what that peer can read — anything it must obey belongs in its charter, not in the block |
| **anything an operator can use, configure, or observe** — an MCP tool, a `bridged.yaml` knob, an endpoint, a visible behaviour | **[Features](wiki/11-Features.md)** — one entry: what it does · the knob that turns it on · **why it exists** · the gotcha |
That last row is not bookkeeping. Chapters 1–10 answer *how is this built* and *why this way*;
none of them has a home for *what can it do and how do I turn it on*, so for twenty tickets a
shipped capability landed nowhere and the Roadmap went on claiming the stage was finished. The
*why* line is the one that matters — without it a decision gets re-litigated from scratch a month
later. Internal contract changes go to `wiki/9-Implementation.md` instead; test and coverage work
is a Roadmap line. A change that touches none of the three earns no entry, and that is a normal
outcome rather than an omission.
Then **propagate**: the block in this file and the template in the wiki
([Use Cases](https://git.ltms.dev/lms/claude-bridge/wiki/7-Use-Cases) → *The portable `CLAUDE.md`
+310 -16
View File
@@ -31,15 +31,97 @@ bind:
# mode: token
# tokenEnv: BRIDGED_API_TOKEN
# Optional pinned primary terminal (CB-307). Names the herdr pane the PRIMARY itself runs in:
# a caller whose connection maps to this pane resolves as the primary (no credential needed —
# the pane mapping is as unforgeable as a worker's), and reply nudges are pushed to it.
# REQUIRED when the primary runs inside a herdr pane — without it the pane match reads the
# primary as a worker and refuses spawn/send/stop. Get the id from bridge_whoami; re-pin if
# the primary moves panes.
# primary:
# terminal: term_0123456789abcd
# pushReminders: 5 # max nudges before giving up (default 5)
# pushBackoffMs: 15000 # delay between nudges (default 15000)
# CB-530: MORE THAN ONE LEAD. `primary:` above is singular by construction — every other pane
# resolves as a worker — which is right for one lead driving a fleet and wrong the moment two leads
# (say a Claude lead and an opencode lead) work as peers: the second is silently demoted and refused
# every orchestration call. List each lead's pane here and all of them resolve as leads.
#
# tab → the ONLY field identity depends on (CB-579); the exact label of the tab hosting the lead.
# Label the tab yourself, or let bridged label one it launches — see `fleet.leaders:` below.
# kind/model → descriptive; they document what runs in the pane and are echoed by bridge_whoami
#
# A lead's tab must already carry its label (or be launched by bridged, which labels it) — there is
# no terminal id to paste in and nothing to re-pin when the session restarts: the tab survives, so
# the same label resolves the same lead again on the next scan.
# `bridge_whoami` reports `{"role":"primary","leader":"<name>"}`; role stays "primary" because a lead
# IS a primary for authorization, so nothing that keys on the role breaks.
#
# KEEP `primary:` when adding leads: it still addresses the CB-307 push loop, which needs a single
# destination for its nudges, and is a separate mechanism from lead identity — see `fleet.leaders:`.
#
# Leads are configured under `fleet.leaders:` — see THE FLEET further down.
#
# Two things stop the tab-name convention from becoming a way to claim leadership: the configured
# member spaces are excluded from the scan, so nothing bridged places can land in a matching tab;
# and startup REFUSES a `tabPrefix` that the fleet tabLabel template, or any per-profile `tabLabel`
# override, also matches — so the two namespaces cannot overlap by accident. The label is a NAME,
# never a capability: what a pane may do is decided by the role the daemon resolves for it.
# CB-551: IDLE-LEAD HEARTBEAT — nudge the single lead back to work when it has been continuously
# idle (no open bridge_send driving it) past the quiet period. The fleet is one lead + architects +
# workers, so a lead that stalls is a single point of failure; the ReplyPushLoop only nudges when a
# reply lands, and this timer catches the gap where nothing lands and the lead just sits idle.
#
# Opt-in on purpose — it SPENDS the operator's subscription on its own initiative (each nudge starts
# a lead turn nobody asked for), so upgrading the daemon must never switch it on for you. Absent
# block = feature off, exactly as before.
#
# Three knobs, each with a default that errs on the side of not burning context:
# idleAfterSeconds: 300 # how long the lead must stay idle before the FIRST nudge (default 300 —
# # absorbs normal post-turn pauses; re-prompting every pause burns context)
# backoffMs: 60000 # re-check cadence / spacing between nudges past the quiet period (default 60000)
# quietNudgeCap: 3 # cap on consecutive nudges that find NOTHING pending, then it stops
# # until real state appears (default 3 — never nag an empty fleet forever)
# leadHeartbeat:
# idleAfterSeconds: 300
# backoffMs: 60000
# quietNudgeCap: 3
# Fleet health detection is dormant unless enabled (CB-573). It reads one whole-fleet agent list
# per tick.
# intervalSeconds → how often a tick runs (default 30). ENFORCED floor of 15: the code computes
# Math.max(15, intervalSeconds), so a lower value is silently raised, not
# rejected.
# workingSuspectAfterSeconds, paneProbeIntervalSeconds → accepted and parsed, but NOT YET READ by
# anything — the dormant monitor only consumes intervalSeconds today (CB-573
# shipped ahead of the evidence publishers these two knobs are for). Setting
# them changes nothing right now, and no minimum is enforced on either, because
# nothing reads them to enforce one. They exist so a later build can start
# honouring them without another config-shape change.
# notifications.mode → "webhook" flips what bridge_list REPORTS (healthCoverage: "full" instead
# of "detection-only") — it does NOT make bridged send any webhook call; no
# delivery mechanism is implemented yet. Any other value, or omitting the
# block, reports "detection-only".
# health:
# enabled: true
# intervalSeconds: 30
# workingSuspectAfterSeconds: 600
# paneProbeIntervalSeconds: 60
# notifications:
# mode: disabled
# herdr Unix socket. Omit to use the client default
# (${HERDR_SOCKET_PATH:-~/.config/herdr/herdr.sock}).
herdrSocket: ~/.config/herdr/herdr.sock
# How worker sessions are spawned. Define one or more named profiles (backends) under
# `workers`; each key is the profile name (also the ccs profile). `defaultWorker` picks
# which one a no-argument spawn uses (bridge_spawn with no profile / POST /workers).
# How member sessions are spawned. Define one or more named profiles (backends) under
# `profiles`; each key is the profile name (also the ccs profile). A profile says only WHICH
# BACKEND — model, CLI adapter, credentials, cost. It says nothing about what a member spawned on
# it is for; that is the member's role, and roles live under `fleet:` below. Which profile an
# unqualified spawn lands on comes from that role's pool, not from a global default.
#
# Shared knobs (placement/workspace/tabLabel) can be repeated per profile; they usually match.
# Shared knobs (placement/workspace) can be repeated per profile; they usually match.
# placement: tab → each worker lands in its OWN tab in a dedicated worker space (default).
# Use `pane` for the legacy behaviour (split the focused tab).
# mcpUrl → bridged mounts the bridge MCP (--mcp-config, inline) + reply charter
@@ -53,13 +135,39 @@ herdrSocket: ~/.config/herdr/herdr.sock
# skills/MCP/hooks. Omit to leave the worker on the host default.
# parityOverlay → repo-relative paths copied primary→worktree so a worker in a provisioned
# worktree sees the same local config (CB-301-ext). Omit for the default set:
# [.mcp.json, .claude/settings.local.json, .env, .envrc].
# [.claude/settings.local.json, .env, .envrc].
#
# Do NOT add .mcp.json (CB-525). A worker's tools are whatever its launcher
# mounts — the bridge, and nothing else. Replicating the primary's MCP config
# handed a worker the primary's IDE servers, which are bound to the primary's
# checkout, so its navigation returned paths OUTSIDE its own worktree: one
# worker made all 59 of its edits in the primary tree while compiling its
# worktree, and every build it ran was of code that did not contain them.
# bridged neutralizes a provisioned worktree's .mcp.json for this reason;
# listing it here would copy the primary's back over that.
# gitTokenEnv → host env var holding the git-forge API token. When set, its value is injected
# as GITEA_TOKEN so the worker can open its OWN PR at checkpoint (CB-302).
# Opt-in by design — omit and the worker gets no PR-create grant (push over
# SSH is unaffected). The token value itself is never stored in this file.
# gitHostEnv → host env var holding the forge host (default GITEA_HOST). Injected as
# GITEA_HOST *only* alongside a resolved gitTokenEnv.
# exhaustedPattern → regex matched against a completion-fallback scrape (CB-578 stage A) to
# classify a turn that ended with no bridge_reply as the backend having
# refused on a subscription usage limit, rather than a real answer. Opt-in —
# omit and this profile's completion fallback behaves exactly as before.
# Every backend words its refusal differently, so this is config, never a
# vendor string baked into bridged itself.
# DEFERRED: compiled once into a startup pattern map — editing it needs a
# daemon restart, same as this profile's model/baseUrl/argv.
# credentialId → CB-578 stage B: the credential this profile quarantines WITH when a
# BACKEND_EXHAUSTED classification fires. Two profiles that set the SAME
# credentialId share one quarantine — the case this exists for is two models
# on one account (e.g. sol and terra both billing one OpenAI credential): an
# exhaustion on either one must lock out both, or the fleet just walks onto
# the same dead account under the sibling's name. Opt-in — omit and this
# profile quarantines alone, under its own name, exactly as if the field did
# not exist. Cooldown length is the top-level quarantineCooldownSeconds below.
# HOT: read live at every spawn/exhaustion check — no restart needed.
# env → extra environment for this profile's workers, as a literal key/value map
# (CB-511). Use it to give workers a toolchain.
#
@@ -78,29 +186,50 @@ herdrSocket: ~/.config/herdr/herdr.sock
# entry cannot repoint a worker past the SubscriptionGuard — which is checked
# against `baseUrl` alone.
# Put `defaultMode: "auto"` in each ccs profile so the worker runs autonomously.
workers:
profiles:
gx10: # ccs profile name (NOT a hostname)
kind: claude-code # which adapter spawns this profile (default; may omit)
baseUrl: http://gx01.gw:8000 # the vLLM host this profile targets (gx00.gw / gx01.gw)
model: coder
placement: tab
workspace: bridged-workers
tabLabel: "worker: {profile} #{n}" # {profile}/{model}/{n} substituted; {n} keeps sibling tabs distinct
# tabLabel: an optional per-profile override; the fleet template usually covers it
mcpUrl: http://127.0.0.1:8765/mcp
tokenEnv: BRIDGED_WORKER_TOKEN
argv: ["ccs", "gx10"]
# weight: relative selection weight for automatic placement (weighted, round-robin, and
# fixed's fallback walk). Absent defaults to 1.0. An explicit 0 or negative value means
# "never auto-select this profile" (CB-554) — it stays reachable via an explicit
# `bridge_spawn{profile:"gx10"}`, which bypasses placement entirely; only automatic
# selection skips it.
weight: 0.5
# maxLoad: max live workers on this profile. Omit for unlimited. An explicit 0 (CB-585) caps
# the profile at zero live members — it is excluded from automatic placement and an explicit
# `bridge_spawn{profile:"gx10"}` against it is refused too; a cap holds even when the profile
# is named directly. Negative is refused at config load — there is no sane meaning for it.
maxLoad: 2
# gitTokenEnv: GITEA_TOKEN # opt-in: let this profile's workers open their own PR (CB-302)
# gitHostEnv: GITEA_HOST # defaults to GITEA_HOST; injected only with gitTokenEnv
# exhaustedPattern: "usage limit has been reached" # opt-in: classify a usage-limit refusal (CB-578)
# credentialId: shared-openai # opt-in: quarantine together with every other profile sharing this id (CB-578)
# configDir: /Users/me/.ccs/instances/gx10 # CLAUDE_CONFIG_DIR — inherit that profile's skills/MCP
# cwd: /Users/me/src/myrepo # pin the working dir; omit to inherit the primary's
# parityOverlay: [".mcp.json", ".claude/settings.local.json", ".env", ".envrc"]
ollama:
baseUrl: http://ollama.ltms.dev # local/self-hosted; usually no token
# parityOverlay: [".claude/settings.local.json", ".env", ".envrc"] # never add .mcp.json — see above
gx11: # a second backend, so `placement: weighted` has a choice
baseUrl: http://gx01.gw:8000 # self-hosted; ccs handles the model + token
placement: tab
workspace: bridged-workers
tabLabel: "worker: {profile} #{n}"
# tabLabel: an optional per-profile override; the fleet template usually covers it
mcpUrl: http://127.0.0.1:8765/mcp
argv: ["ccs", "ollama"]
argv: ["ccs", "gx11"]
weight: 0.5
maxLoad: 2
# Pin an auto-compact window BELOW the served model's context ceiling. The global
# ~/.claude/settings.json value is shared by every ccs instance and the primary, so the
# per-profile override belongs here. Equal to the ceiling means auto-compact never fires
# before the server rejects the prompt, which kills a worker mid-turn (CB-523).
env:
CLAUDE_CODE_AUTO_COMPACT_WINDOW: "280000"
# CB-402: a second coding-agent kind, proving the PeerLauncher SPI is provider-neutral.
# opencode is provider-agnostic and uses NONE of Claude's private seams: no ANTHROPIC_BASE_URL /
# SubscriptionGuard (so it needs no `guard` host entry), no --mcp-config / --append-system-prompt.
@@ -144,7 +273,155 @@ workers:
# tabLabel: "opencode: {profile} #{n}"
# mcpUrl: http://127.0.0.1:8765/mcp
# argv: ["opencode"]
defaultWorker: gx10
# How an unqualified spawn chooses a profile: fixed (default, reproduces pre-CB-518 behaviour),
# round-robin, or weighted. Omitting this key is a strict no-op for existing configs.
placement: weighted
# How long a credential sits out after a BACKEND_EXHAUSTED classification (CB-578 stage B), in
# seconds, before a spawn may land on it again. Applies to every profile's effective credential
# (its own name, or its credentialId if set above) — there is no per-profile override. Default
# 1800 (30 minutes) when omitted or non-positive.
# DEFERRED: baked once into the BackendQuarantine built at startup — a running quarantine keeps
# its original cooldown regardless; a new value only applies to a quarantine that starts after a
# restart. Editing this needs a daemon restart to take effect.
# quarantineCooldownSeconds: 1800
# Re-read this file without restarting the daemon (CB-559). Off unless you add this block, so an
# upgraded bridged keeps the old behaviour: the file is read once at boot and never again.
# enabled → turn the watch on. bridged checks the file's modified time on a timer and
# reloads when it moves.
# intervalSeconds → how often to check (default 10). One `stat` per tick, so this is cheap.
#
# Not every key can move under a running daemon, and the difference is about what already exists
# when the reload happens — not about how important the key is:
# HOT → takes effect on the next spawn: the whole `fleet:` block (every role pool,
# `charters`, and `tabLabel`), `placement:`, and an existing profile's weight / maxLoad
# / credentialId. Those are hot because the placement policy (and, for credentialId,
# the CB-578 stage B quarantine check) reads them through a supplier — being config is
# not by itself enough to make a key hot.
# EXCEPT `fleet.leaders`: Bridged.main reads it once at startup to build the lead tab
# scanner and launcher, and neither is rebuilt on reload. A changed/added/removed
# `fleet.leaders` entry is silently accepted — the reload reports "config reloaded"
# with nothing in the deferred list — but has NO effect until you restart. Treat it
# as deferred in practice, even though today's reload output does not say so.
# DEFERRED → accepted into the new config, but the wiring built at startup keeps the old value
# until you restart: `lifecycle:`, `leadHeartbeat:`, `guard:`, `worktreeRoot:`,
# `spawnReadyTimeoutMs` / `spawnReadyPollMs`, `quarantineCooldownSeconds` (CB-578
# stage B — baked once into the quarantine tracker built at startup), ADDING or
# REMOVING a profile (a new backend needs its own launcher, and launchers are built
# once), AND an existing profile's launch settings — model, baseUrl, argv, env,
# configDir, mcpUrl, tabLabel, exhaustedPattern. The launcher takes a copy of
# `profiles:` at startup and resolves every spawn out of that copy, so those never
# reach a launch until you restart. The reload logs them by name rather than
# pretending they applied.
# COLD → cannot change at all: `bind:`, `herdrSocket:`, `broker:` and `auth:`. The socket is
# bound, the broker connection is open, and the auth mode decides who may reach the
# port that is already listening.
#
# A changed COLD key refuses the WHOLE reload — not the hot half applied and the cold half warned
# about. A half-applied reload would leave the daemon matching no file on disk, which is the worst
# thing a reload can do to an operator debugging one. A file that fails to parse or fails a startup
# validator is refused the same way, and the running config stays live.
# configReload:
# enabled: true
# intervalSeconds: 10
# THE FLEET (CB-557) — who the daemon may run, and under which role. This one block replaced four
# older keys: `leaders:`, `members:`, `leadScan:` and `defaultProfile:`.
#
# A member is anything a lead spawns, and every member has two INDEPENDENT attributes:
# role — which contract: architect, dev or reviewer. It picks the launch charter, the role
# file, the playbook skill and the authz row.
# profile — which backend: one of the `profiles:` keys above (model, CLI adapter, cost).
# They vary on their own. A reviewer may run on the same profile as the dev whose diff it reads,
# which is why the two cannot be one field.
#
# The ROLE IS THE CONTAINING KEY, not a `role:` field. That is not only tidier: a misspelled role
# used to parse into a member with no contract at all, while a misspelled pool name here simply
# declares nothing.
#
# Each pool lists the profiles that role MAY run on — these are pools, not identities. That is also
# what replaced `defaultProfile:`: an unqualified spawn names a role, and that role's pool supplies
# the candidates, in definition order. A dev and a reviewer staying anonymous is exactly compatible
# with being listed here; the entry key just names the entry.
fleet:
# Optional launch-charter text, keyed only by the singular role wire names: architect, dev,
# reviewer. Changes are HOT and reach the next spawn without a daemon restart. Do not put secrets
# here: a later launch step writes this text to a world-readable temp file, and ${ENV} interpolation
# is deliberately not supported.
charters:
architect: |-
You are an architect in this fleet. You refine work before anyone builds it:
scope, acceptance criteria, risks, and a unit split. You read the repo and
write analysis. You never commit production code and never open a PR.
A design task is worked by two architects. Design alone first, then exchange
and say plainly where you disagree. Do not concede just to agree.
dev: |-
You implement the one unit you were given, and nothing else. You test it,
commit it, and open your own pull request. You never merge.
reviewer: |-
You review the diff you were given. You report bugs, risks and missing tests.
You do not change code.
# Optional. Template for a member tab's label; {role}, {profile}, {model} and {n} are substituted.
# {n} counts per role+profile, so `dev: sonnet #2` really is the second sonnet dev. Because {role}
# comes from a closed enum, a generated label can never begin with a lead's tabPrefix.
# tabLabel: "{role}: {profile} #{n}"
# Panes that orchestrate rather than are orchestrated. A lead may now be CREATED as well as
# recognised: give it a `profile:` and the daemon launches the shortfall when fewer than
# `instances` are live. Omit `profile:` and it is recognise-only, as before.
#
# `tab:` (CB-579) is REQUIRED and is the only field identity depends on — the exact label of the
# tab hosting the lead, matched case-insensitively. Label the tab yourself and put that same
# string here, and the pane is recognised on the next rescan. Reopen the tab later, or the session
# inside it restarts — the terminal id changes; the tab, and its label, do not, so no config edit
# follows a restart.
#
# A lead the daemon launches is labelled BY the daemon with this same `tab:` value, so it is found
# by the same scan. A lead counts as live only when herdr also reports a running agent in that
# tab — a label left behind by a session that died does not block the relaunch, and a tab that is
# gone entirely drops out of the next scan rather than being remembered forever.
#
# An auto-launched lead is NOT a member: it gets no worker reply charter, is never registered with
# the session lifecycle (the idle reaper would kill your orchestrator), and stays on the
# subscription — ANTHROPIC_BASE_URL/AUTH_TOKEN are stripped from its env whatever the profile says.
#
# GET THE `tab:` VALUE RIGHT. A pane that does not match any configured `tab:` (a typo, a renamed
# tab, a pane no entry names at all) is not recognised as a lead — it resolves as an ordinary
# WORKER instead, silently, and every orchestration call it makes (spawn/stop/send/drain) is
# refused. There is no error at startup for this: an unmatched pane is simply not a lead. If your
# primary suddenly can't spawn or send, check this section first.
# leaders:
# opus-5.0:
# profile: opus # omit to never create this lead, only recognise it
# instances: 1 # desired live count; only the shortfall is launched. 0 = off
# tab: "lead: opus-5.0" # REQUIRED — the exact tab label this lead lives in
# tabPrefix: "lead:" # only used to guard against a worker tabLabel colliding with
# # this convention at startup; plays no part in matching a lead
# scanIntervalSeconds: 10 # rescan cadence, and the worst case before a new tab is seen
# workspace: leads # where a launched lead's tab is created (default "leads").
# # MUST NOT be a member workspace — those are excluded from the
# # scan, so a lead placed in one is never found again.
# cwd: /path/to/repo # the launched lead's working directory (default: bridged's own)
# kind: claude # descriptive; reported by bridge_whoami
# gpt-sol-5.6:
# tab: "lead: gpt-sol-5.6"
# kind: opencode
# model: openai/gpt-5.6-terra
# architects:
# architect-1:
# profile: opus # a strong model, on the operator's subscription
# architect-2:
# profile: sol # a different vendor on purpose — two architects that share a
# # model share its blind spots
developers:
gx10:
profile: gx10
# reviewers:
# gx10:
# profile: gx10 # the same backend may serve two roles; that is the point
# Subscription boundary. A worker's base_url host MUST be one of these; the primary
# must carry none. Every profile above must have its host listed here.
@@ -152,7 +429,6 @@ guard:
offSubscriptionHosts:
- gx00.gw
- gx01.gw
- ollama.ltms.dev
# Spawn-readiness gate (CB-306). The launcher blocks until the worker's herdr status is
# injectable (IDLE/BLOCKED/DONE) or the timeout elapses. 0 disables the gate.
@@ -171,10 +447,15 @@ guard:
# idleTtlSeconds → reap READY/DONE sessions idle longer than this (never BUSY/SPAWNING)
# contextCap → force-release a session after this many delegated turns
# drainTimeoutSeconds → seconds to wait for BUSY sessions on shutdown before forced teardown
# clearAfterTurn → whether a reusable worker discards its conversation context after every
# completed delegated turn (default false). Works for claude-code workers
# only — any other peer kind (e.g. opencode) logs "context reset is
# unsupported for peer kind …" once and the reset is a no-op.
# lifecycle:
# idleTtlSeconds: 300
# contextCap: 10
# drainTimeoutSeconds: 5
# clearAfterTurn: false
# Durable reply delivery (CB-307 Stage 2). OMIT this block entirely to keep the default
# in-memory, soft-state reply inbox (late worker replies are held only until a daemon bounce).
@@ -182,10 +463,14 @@ guard:
# on a durable per-target queue (agent.<target>.inbox) and survive a restart — the broker
# redelivers anything the primary had not yet drained. Production default is LavinMQ; a stock
# RabbitMQ speaks the same AMQP 0-9-1, so it is a URI-only swap.
# uri → AMQP connection URI. No trailing slash ⇒ the default vhost "/"; an empty path ("/")
# is vhost "" and will NOT connect. Encode a named vhost as .../%2Fmyvhost.
# uri → AMQP connection URI. No trailing slash ⇒ the default vhost "/"; an empty path ("/")
# is vhost "" and will NOT connect. Encode a named vhost as .../%2Fmyvhost.
# prefetch → CB-527: consumer basicQos, capping how many unacked messages the inbox holds
# in-heap per owned target (the rest sits on the broker's durable queue instead of
# growing the JVM heap). Default 32 when omitted.
# broker:
# uri: amqp://guest:guest@127.0.0.1:5672
# prefetch: 32
# Active push-to-primary (CB-307 Stage 3). When a worker reply lands with no open bridge_send,
# the ReplyPushLoop injects a *drain nudge* (never the payload) into the primary's own herdr
@@ -195,6 +480,15 @@ guard:
# the first orchestration-side MCP call (the normal case). An off-host or
# non-herdr primary leaves this unresolved → the loop is a no-op and delivery
# degrades to pull; the reply is still never lost.
#
# REQUIRED (CB-522) if the primary itself runs inside a herdr pane. Caller
# identity resolves a loopback PID to its herdr pane, and PaneLocator scans
# EVERY pane — not just bridged-spawned ones — so such a primary is otherwise
# classified as a WORKER and refused SPAWN/SEND/STOP. That failure is
# self-locking: the learned terminal is populated by the very orchestration
# calls being refused, so only this pinned value can break the cycle. Read the
# id off bridge_whoami (it reports the current terminal even while
# misclassified) and re-pin whenever the primary moves panes.
# pushReminders → max nudges before giving up (default 5)
# pushBackoffMs → delay between nudges in ms (default 15000)
# primary:
+21 -1
View File
@@ -6,7 +6,7 @@
<groupId>dev.ltms</groupId>
<artifactId>bridged</artifactId>
<version>0.1.0-SNAPSHOT</version>
<version>1.0.0</version>
<packaging>jar</packaging>
<name>bridged</name>
@@ -236,6 +236,26 @@
<profile>
<id>contract</id>
<properties><excludedGroups/></properties>
<build>
<plugins>
<plugin>
<artifactId>maven-surefire-plugin</artifactId>
<configuration>
<!-- Docker-engine compat (see "Running the contract tests" in
docs/CB-307-Reliable-Delivery.md): Testcontainers 1.20.4's docker-java
client defaults to Docker API 1.32 when no version is set, but modern
engines (OrbStack on this dev host, min 1.40) reject that as too old —
which surfaces as "Could not find a valid Docker environment". Pinning
api.version=1.43 works on OrbStack and Docker 24+, and is overridable
per-host via -Dapi.version. Only active under -Pcontract, so the
default hermetic build never sets it. -->
<systemPropertyVariables>
<api.version>1.43</api.version>
</systemPropertyVariables>
</configuration>
</plugin>
</plugins>
</build>
</profile>
</profiles>
</project>
@@ -1,23 +1,31 @@
package dev.ltms.bridged;
import dev.ltms.bridged.config.BridgedConfig;
import dev.ltms.bridged.config.ConfigRef;
import dev.ltms.bridged.config.ConfigWatcher;
import dev.ltms.bridged.guard.SubscriptionGuard;
import dev.ltms.bridged.herdr.AgentControl;
import dev.ltms.bridged.herdr.HerdrClient;
import dev.ltms.bridged.herdr.HerdrException;
import dev.ltms.bridged.herdr.LeadTabScanner;
import dev.ltms.bridged.lead.LeadLauncher;
import dev.ltms.bridged.herdr.PaneLocator;
import dev.ltms.bridged.herdr.UnixSocketHerdrClient;
import dev.ltms.bridged.herdr.WorkspaceControl;
import dev.ltms.bridged.inject.CompletionResolver;
import dev.ltms.bridged.inject.ExhaustedPatternLookup;
import dev.ltms.bridged.inject.ExhaustionSink;
import dev.ltms.bridged.inject.Injector;
import dev.ltms.bridged.inject.StatusPoller;
import dev.ltms.bridged.inject.TurnListener;
import dev.ltms.bridged.inject.WorkerPresence;
import dev.ltms.bridged.inject.MemberPresence;
import dev.ltms.bridged.auth.MemberRegistry;
import dev.ltms.bridged.auth.CallerResolver;
import dev.ltms.bridged.mcp.BridgeMcp;
import dev.ltms.bridged.mcp.ConnectionIdentity;
import dev.ltms.bridged.metrics.BridgedMetrics;
import dev.ltms.bridged.metrics.Metrics;
import dev.ltms.bridged.health.FleetHealthMonitor;
import dev.ltms.bridged.mcp.PrimaryRegistry;
import dev.ltms.bridged.mcp.LsofPeerPidLookup;
import dev.ltms.bridged.mcp.LsofProcessCwdLookup;
@@ -26,16 +34,19 @@ import dev.ltms.bridged.msg.InMemoryReplyInbox;
import dev.ltms.bridged.msg.MessageService;
import dev.ltms.bridged.msg.Rendezvous;
import dev.ltms.bridged.msg.ReplyInbox;
import dev.ltms.bridged.msg.LeadHeartbeatLoop;
import dev.ltms.bridged.msg.ReplyPushLoop;
import dev.ltms.bridged.rest.BridgedApp;
import dev.ltms.bridged.session.GitWorktrees;
import dev.ltms.bridged.session.MemberSession;
import dev.ltms.bridged.session.SessionManager;
import dev.ltms.bridged.peer.PeerLauncher;
import dev.ltms.bridged.session.SessionReaper;
import dev.ltms.bridged.worker.ClaudeCodeLauncher;
import dev.ltms.bridged.worker.CompositePeerLauncher;
import dev.ltms.bridged.worker.HerdrPeerLauncher;
import dev.ltms.bridged.worker.OpenCodeLauncher;
import dev.ltms.bridged.member.ClaudeCodeLauncher;
import dev.ltms.bridged.member.CompositePeerLauncher;
import dev.ltms.bridged.member.HerdrPeerLauncher;
import dev.ltms.bridged.member.OpenCodeLauncher;
import dev.ltms.bridged.placement.BackendQuarantine;
import io.javalin.Javalin;
import org.slf4j.Logger;
import org.slf4j.LoggerFactory;
@@ -45,7 +56,16 @@ import java.util.ArrayList;
import java.util.LinkedHashMap;
import java.util.List;
import java.util.Map;
import java.util.Objects;
import java.util.Set;
import java.util.concurrent.Executors;
import java.util.concurrent.TimeUnit;
import java.util.concurrent.atomic.AtomicReference;
import java.util.function.Function;
import java.util.function.Predicate;
import java.util.function.Supplier;
import java.util.regex.Pattern;
import java.util.stream.Collectors;
/**
* {@code bridged} entry point. Wires the real herdr socket client to the REST app and
@@ -56,9 +76,6 @@ public final class Bridged {
private static final Logger log = LoggerFactory.getLogger(Bridged.class);
/** How often the injector samples a busy worker's status while it has queued work. */
private static final long INJECT_POLL_MILLIS = 250;
/** CB-504: how long to wait at startup for herdr's socket before serving degraded. */
private static final long HERDR_WAIT_SECONDS = 30;
private static final long HERDR_WAIT_POLL_MILLIS = 500;
@@ -66,6 +83,11 @@ public final class Bridged {
static void main(String[] args) {
Path configPath = Path.of(args.length > 0 ? args[0] : "bridged.yaml");
BridgedConfig cfg = BridgedConfig.load(configPath);
// CB-559: `cfg` stays the startup snapshot — every validation and every piece of one-time
// wiring below reads it, and must, because those decisions cannot be unmade. `config` is the
// live reference the hot paths read per use. Which keys can actually move is ConfigRef's
// contract; adding a reader here does not make a key reloadable by itself.
ConfigRef config = new ConfigRef(configPath, cfg);
// The primary/host env that launched bridged must not be tainted.
SubscriptionGuard guard = new SubscriptionGuard(cfg.guard().hostSet());
@@ -76,6 +98,14 @@ public final class Bridged {
// refuses remote connections to a loopback socket. This throws rather than warns so the
// dangerous configuration cannot be reached by ignoring a log line.
cfg.validateAuthExposure();
cfg.validateLeadTabPrefixes();
// CB-542: a subscription:true profile whose env: reseats ANTHROPIC_BASE_URL/AUTH_TOKEN would
// reach an unguarded endpoint (the launcher skips SubscriptionGuard for it). Refuse at load.
cfg.validateSubscriptionProfiles();
cfg.validateCharters();
// CB-548: every architect slot must name a configured workers: profile — the strong-model
// backend the future spawn lifecycle would read. A stale reference dies here, not later.
cfg.validateMembers();
Path socket = cfg.herdrSocket() != null && !cfg.herdrSocket().isBlank()
? Path.of(cfg.herdrSocket())
@@ -88,9 +118,9 @@ public final class Bridged {
// CB-402: one adapter per configured peer kind, fronted by a composite router. A profile's
// `kind:` selects its adapter — claude-code (the default) and opencode partition the profile
// set — and the composite dispatches each SPI call to the adapter that owns the profile/pane.
Map<String, BridgedConfig.Worker> claudeProfiles = new LinkedHashMap<>();
Map<String, BridgedConfig.Worker> opencodeProfiles = new LinkedHashMap<>();
cfg.workerProfiles().forEach((name, w) -> {
Map<String, BridgedConfig.Profile> claudeProfiles = new LinkedHashMap<>();
Map<String, BridgedConfig.Profile> opencodeProfiles = new LinkedHashMap<>();
cfg.profiles().forEach((name, w) -> {
if (w.isOpenCode()) {
opencodeProfiles.put(name, w);
} else {
@@ -103,21 +133,36 @@ public final class Bridged {
// unless opencode is the only kind configured.
if (!claudeProfiles.isEmpty() || opencodeProfiles.isEmpty()) {
adapters.add(new ClaudeCodeLauncher(agents, spaces, guard,
claudeProfiles, cfg.defaultProfile(), System::getenv,
cfg.spawnReadyTimeoutMs(), cfg.spawnReadyPollMs()));
claudeProfiles, cfg.effectiveDefaultProfile(), System::getenv,
cfg.spawnReadyTimeoutMs(), cfg.spawnReadyPollMs(),
() -> config.get().fleet()));
}
if (!opencodeProfiles.isEmpty()) {
adapters.add(new OpenCodeLauncher(agents, spaces,
opencodeProfiles, cfg.defaultProfile(), System::getenv,
cfg.spawnReadyTimeoutMs(), cfg.spawnReadyPollMs()));
opencodeProfiles, cfg.effectiveDefaultProfile(), System::getenv,
cfg.spawnReadyTimeoutMs(), cfg.spawnReadyPollMs(),
() -> config.get().fleet()));
}
PeerLauncher workers = new CompositePeerLauncher(adapters, cfg.defaultProfile());
AtomicReference<Function<String, Integer>> liveCountRef = new AtomicReference<>(_ -> 0);
// CB-578 stage B: one quarantine tracker for the whole daemon, shared between the launcher
// (checked at spawn) and the exhaustion sink wired in below (written on BACKEND_EXHAUSTED).
// The cooldown is deferred (see BridgedConfig#quarantineCooldownSeconds): it is read once
// here, at startup, and a config reload only changes it for a daemon restart.
BackendQuarantine quarantine = new BackendQuarantine(System::nanoTime,
TimeUnit.SECONDS.toNanos(cfg.quarantineCooldownSeconds()));
PeerLauncher workers = new CompositePeerLauncher(
adapters,
cfg.effectiveDefaultProfile(),
config,
profileName -> liveCountRef.get().apply(profileName),
quarantine);
// CB-504: under supervision (launchd/systemd) bridged can start before herdr's socket
// exists. The client itself is lazy — it connects per call — but the orphan reap below is
// the first thing that actually talks to herdr, so without this wait a boot-order race
// would crash the daemon into a restart loop. Wait, then degrade rather than die: serving
// with /healthz reporting "degraded" is strictly more useful than exiting.
if (awaitHerdr(herdr)) {
boolean herdrUp = awaitHerdr(herdr);
if (herdrUp) {
// CB-117: herdr keeps worker panes alive across a daemon restart, and their ids died
// with the previous process — reap those leaked orphans now, before we start serving.
workers.reapOrphanWorkers();
@@ -135,7 +180,12 @@ public final class Bridged {
&& cfg.lifecycle().contextCap() > 0) {
contextCap = cfg.lifecycle().contextCap();
}
SessionManager sessions = new SessionManager(workers, new GitWorktrees(cfg.worktreeRoot()), contextCap);
boolean clearAfterTurn = cfg.lifecycle() != null && cfg.lifecycle().clearAfterTurn();
SessionManager sessions = new SessionManager(workers, new GitWorktrees(cfg.worktreeRoot()),
System::nanoTime, contextCap, clearAfterTurn);
liveCountRef.set(profileName -> (int) sessions.roster().stream()
.filter(s -> profileName.equals(s.profile()))
.count());
// CB-303 part 1: idle-ttl reaper — only when configured, defaults to disabled.
final SessionReaper reaper;
@@ -148,14 +198,110 @@ public final class Bridged {
reaper = null;
}
// CB-530: every pane the config names as a lead, merged from `leaders:` and the legacy
// singular pin. PrimaryRegistry below still tracks ONE terminal — it addresses the push
// loop's nudges, which need a single destination — so it keeps the legacy pin.
Map<String, String> leadTerminals = cfg.leaderTerminals();
if (leadTerminals.size() > 1) {
log.info("leads: {} panes recognised {}", leadTerminals.size(), leadTerminals.values());
}
// CB-531: on top of the legacy primary.terminal pin, discover leads by the tab labels the
// operator writes. CB-557 moved the settings onto the lead they describe, so scanning is on
// whenever a `fleet.leaders:` entry exists — with no leads configured the supplier is a
// constant and never touches herdr, exactly as a missing `leadScan:` block used to behave.
// CB-579: each lead now names its own exact `tab:` label, so one scanner discovers every
// configured lead regardless of how differently their tabs are labelled — the old
// single-shared-tabPrefix limitation (and its warning) is gone.
final Supplier<Map<String, String>> leads;
var leaders = cfg.fleet().leaders();
if (!leaders.isEmpty()) {
Set<String> memberSpaces = cfg.profiles().values().stream()
.map(BridgedConfig.Profile::workspace)
.filter(Objects::nonNull)
.collect(Collectors.toSet());
Map<String, String> tabToName = new LinkedHashMap<>();
leaders.forEach((name, leader) -> {
if (leader != null && leader.tab() != null && !leader.tab().isBlank()) {
tabToName.put(leader.tab(), name);
}
});
// One shared rescan cadence: still taken from the first entry, as before — it is an
// operational cadence, not identity, so there is no correctness reason to give every
// lead its own scanner.
int scanIntervalSeconds = leaders.values().iterator().next().scanIntervalSeconds();
leads = new LeadTabScanner(herdr, tabToName, memberSpaces,
TimeUnit.SECONDS.toNanos(scanIntervalSeconds), System::nanoTime);
log.info("lead scan: tabs {} host a lead (rescan every {}s, member spaces {} excluded)",
tabToName.keySet(), scanIntervalSeconds, memberSpaces);
} else {
leads = () -> leadTerminals;
}
// CB-558: start any declared lead that is not already running. After the scanner is built,
// because both read the same tab labels and the ordering makes that dependency visible; and
// only when herdr answered, because the launcher's whole safety property is that it can
// count live leads first — it must never guess and risk a second orchestrator.
if (herdrUp && !leaders.isEmpty()) {
int launched = new LeadLauncher(agents, spaces, cfg).ensureLeads();
if (launched > 0) {
log.info("lead auto-launch: {} lead(s) started", launched);
}
}
// CB-548: config-declared architect slots. Config supplies only the stable name → profile
// map; the terminal → slot binding is owned by the registry and is empty at startup, so no
// pane resolves to an architect until the later spawn lifecycle binds one. The registry is
// what CallerResolver resolves against and what that lifecycle will read profiles from;
// nothing here spawns a slot.
MemberRegistry members = new MemberRegistry(cfg.fleet());
sessions.setMemberLifecycle(members);
if (!members.slots().isEmpty()) {
log.info("member slots: {} configured {} — none bound yet (a slot is idle until the "
+ "spawn lifecycle binds a live terminal to it)",
members.slots().size(), members.slots().keySet());
}
// Status-gated injector (CB-103): the single writer into workers, fed by a poller.
// The blocking message endpoint (CB-104) is the producer; the poller is inert until then.
// CB-106: a confirmed turn completion resolves a blocked send whose worker never replied.
Rendezvous rendezvous = new Rendezvous();
CompletionResolver completion = new CompletionResolver(agents, rendezvous);
// CB-578 stage A: classify a completion-fallback scrape that matches a profile's configured
// usage-limit refusal as BACKEND_EXHAUSTED rather than handing it back as a real answer.
// Compiled once at startup, keyed by profile name; a profile with no exhaustedPattern is
// simply absent here, so its workers keep today's completion-fallback behaviour unchanged.
Map<String, Pattern> exhaustedPatternsByProfile = new LinkedHashMap<>();
cfg.profiles().forEach((name, profile) -> {
if (profile.hasExhaustedPattern()) {
exhaustedPatternsByProfile.put(name, Pattern.compile(profile.exhaustedPattern()));
}
});
ExhaustedPatternLookup exhaustedPatterns = target -> sessions.roster().stream()
.filter(session -> target.equals(session.terminalId()))
.findFirst()
.map(session -> exhaustedPatternsByProfile.get(session.profile()))
.orElse(null);
log.info("backend-exhausted classification (CB-578 stage A): {}",
CompletionResolver.coverage(cfg.profiles().keySet(), exhaustedPatternsByProfile.keySet()));
// CB-578 stage B: on a classification that actually wins, quarantine the exhausted profile's
// CREDENTIAL — not the profile name — so a profile sharing that credential (e.g. two models
// on one OpenAI account) is refused too, not just the one that happened to report it. Reads
// the profile config live off `config`, so a credentialId edit is hot: no restart needed.
ExhaustionSink exhaustionSink = (target, reason) -> sessions.roster().stream()
.filter(session -> target.equals(session.terminalId()))
.findFirst()
.map(MemberSession::profile)
.map(profileName -> config.get().profiles().get(profileName))
.ifPresent(profile -> {
String credentialId = profile.effectiveCredentialId();
quarantine.quarantine(credentialId);
log.warn("credential '{}' quarantined for {}s (profile '{}' classified "
+ "BACKEND_EXHAUSTED): {}", credentialId,
cfg.quarantineCooldownSeconds(), profile.profile(), reason);
});
CompletionResolver completion = new CompletionResolver(agents, rendezvous, exhaustedPatterns, exhaustionSink);
// CB-113: deliver only to an available worker (its MCP is connected), never its boot window.
// CB-301: the manager's presence bridge records availability and drives SPAWNING → READY.
WorkerPresence presence = sessions.asPresence();
MemberPresence presence = sessions.asPresence();
TurnListener turnListener = new TurnListener() {
@Override
public void onTurnComplete(String target) {
@@ -164,9 +310,20 @@ public final class Bridged {
}
@Override
public void onDelivered(String target) {
completion.onDelivered(target);
sessions.onDelivered(target);
public boolean hasPostTurnAction(String target) {
return sessions.hasPostTurnAction(target);
}
@Override
public boolean onTurnCompleteWithPostAction(String target) {
completion.resolveBeforePostAction(target);
return sessions.onTurnCompleteWithPostAction(target);
}
@Override
public void onDelivered(String target, dev.ltms.bridged.msg.TurnToken token) {
completion.onDelivered(target, token);
sessions.onDelivered(target, token);
}
@Override
@@ -174,9 +331,16 @@ public final class Bridged {
completion.onTurnFailed(target);
sessions.onTurnFailed(target);
}
@Override
public void onTurnFailed(String target, String reason) {
completion.onTurnFailed(target, reason);
sessions.onTurnFailed(target);
}
};
Injector injector = new Injector(agents, turnListener, presence::isPresent, presence::forget);
StatusPoller poller = new StatusPoller(agents, injector, INJECT_POLL_MILLIS);
Injector injector = new Injector(agents, turnListener, deliverableTo(presence, leads),
presence::forget);
StatusPoller poller = new StatusPoller(agents, injector, Injector.POLL_INTERVAL_MILLIS);
poller.start();
// CB-307: reply inbox. A broker: block (with a uri) selects the AMQP-backed durable adapter;
@@ -184,16 +348,27 @@ public final class Bridged {
// connection, so keep the reference to close it in the ordered shutdown hook.
final ReplyInbox replyInbox;
if (cfg.broker() != null && cfg.broker().isConfigured()) {
replyInbox = AmqpReplyInbox.open(cfg.broker().uri());
log.info("reply inbox: AMQP broker (durable) at {}", cfg.broker().uri());
replyInbox = AmqpReplyInbox.open(cfg.broker().uri(), cfg.broker().prefetchOrDefault());
log.info("reply inbox: AMQP broker (durable) at {} (prefetch={})",
cfg.broker().uri(), cfg.broker().prefetchOrDefault());
} else {
replyInbox = new InMemoryReplyInbox();
log.info("reply inbox: in-memory (soft-state)");
}
// CB-307: learn the primary's terminal from orchestration tool calls (or pin from config).
PrimaryRegistry primaryRegistry = new PrimaryRegistry(
cfg.primary() != null ? cfg.primary().terminal() : null);
// The pin also feeds CallerResolver below: a primary running inside a herdr pane would
// otherwise resolve as a worker and be refused every orchestration tool.
String pinnedPrimaryTerminal = cfg.primary() != null ? cfg.primary().terminal() : null;
PrimaryRegistry primaryRegistry = new PrimaryRegistry(pinnedPrimaryTerminal);
// CB-532: `primary.terminal` is superseded and no longer needed for either of its jobs —
// identity comes from `leaders:`/`leadScan:`, and reply nudges now follow the delegating
// lead. Say so once at startup rather than leaving a redundant pin to look load-bearing.
if (pinnedPrimaryTerminal != null && !pinnedPrimaryTerminal.isBlank()) {
log.warn("primary.terminal is DEPRECATED (CB-532) and can be deleted: identity now comes "
+ "from leaders:/leadScan:, and reply nudges follow the lead that delegated. "
+ "It still works, and is still the fallback nudge destination when a restart "
+ "has lost the delegation map. Its pushReminders/pushBackoffMs stay valid.");
}
// CB-307: active push-to-primary loop — nudge the primary when replies land without an
// open bridge_send. Uses its own lightweight scheduled executor, separate from the injector.
int maxReminders = cfg.primary() != null ? cfg.primary().remindersOrDefault() : 5;
@@ -206,14 +381,66 @@ public final class Bridged {
Metrics metrics = BridgedMetrics.create(sessions, replyInbox);
var pushLoop = new ReplyPushLoop(primaryRegistry, agents, replyInbox,
pushScheduler, maxReminders, backoffMs, metrics);
// CB-551: idle-lead heartbeat. Opt-in; absent `leadHeartbeat:` this is never constructed, so
// an upgraded daemon cannot silently start spending subscription on nudging an idle lead.
// It has its own single-thread scheduler and holds its own scheduler shutdown via close().
final LeadHeartbeatLoop heartbeat;
var heartbeatScheduler = Executors.newSingleThreadScheduledExecutor(r ->
Thread.ofVirtual().name("bridge-heartbeat-").unstarted(r));
if (cfg.leadHeartbeat() != null) {
var hb = cfg.leadHeartbeat();
heartbeat = new LeadHeartbeatLoop(primaryRegistry, agents, replyInbox, sessions::roster,
pushLoop, heartbeatScheduler, System::nanoTime,
TimeUnit.SECONDS.toNanos(hb.idleAfterSeconds()), hb.backoffMs(), hb.quietNudgeCap(),
metrics);
heartbeat.start();
} else {
heartbeat = null;
heartbeatScheduler.shutdownNow();
}
MessageService messages = new MessageService(agents, injector, rendezvous, replyInbox,
pushLoop, metrics);
// Health is a slow whole-fleet observer. Keep it separate from the 250ms delivery poller.
final FleetHealthMonitor healthMonitor;
var healthScheduler = Executors.newSingleThreadScheduledExecutor(r ->
Thread.ofVirtual().name("bridge-health-").unstarted(r));
if (cfg.health() != null && cfg.health().isEnabled()) {
// CB-580: a member found GONE/NEVER_READY must fail whatever ticket is waiting on it,
// through the same idempotent target-wide operation CB-516 already uses on release.
healthMonitor = new FleetHealthMonitor(agents, sessions::roster, messages, healthScheduler,
System::nanoTime, cfg.health().intervalOrDefault(), messages::abandon);
String coverage = FleetHealthMonitor.coverage(true,
cfg.health().notifications() != null && cfg.health().notifications().configured());
if ("detection-only".equals(coverage)) {
log.warn("fleet health: {} (no notification sink configured)", coverage);
} else {
log.info("fleet health: {}", coverage);
}
healthMonitor.start();
} else {
healthMonitor = null;
healthScheduler.shutdownNow();
}
// CB-520: the reply inbox only consumes for agents this gateway owns. own on acquire,
// release on teardown. Do this before CB-516 so the inbox is owned before any reply can land.
sessions.onAcquire(replyInbox::own);
// CB-516: releasing a worker must fail whatever send was waiting on it. Without this a
// torn-down delegation kept reporting PENDING until the 30-minute async timeout, and never
// reached /metrics — the delegation was unresolvable and nothing said so.
sessions.onRelease(terminal ->
messages.abandon(terminal, "the worker session was released before it replied"));
sessions.onRelease(detail -> {
// CB-578 stage C, acceptance criterion 10: a failed ticket's detail should tell a lead
// where to re-dispatch onto the same tree, not just that the worker vanished.
String reason = "the worker session was released before it replied";
if (detail.worktreePath() != null) {
reason += "; worktree=" + detail.worktreePath() + " branch=" + detail.branch()
+ " snapshot=" + (detail.snapshotRef() != null ? detail.snapshotRef() : "none");
}
messages.abandon(detail.terminalId(), reason);
replyInbox.release(detail.terminalId());
primaryRegistry.forgetDelegation(detail.terminalId()); // CB-532: don't leak the lead binding
});
// MCP server face (CB-105): bridge_send/bridge_reply/bridge_status, mounted at /mcp.
// Caller identity is resolved from the connection (peer PID → herdr pane), not arguments.
@@ -229,16 +456,39 @@ public final class Bridged {
throw new IllegalStateException("auth.mode=token but env var " + cfg.auth().tokenEnv()
+ " is unset or empty — export it before starting bridged");
}
callers = new CallerResolver(identity, true, token);
callers = CallerResolver.withLeadsAndMembers(identity, true, token, leads, members);
log.info("auth: token mode (bearer required for non-worker callers, env {})",
cfg.auth().tokenEnv());
} else {
callers = new CallerResolver(identity);
callers = CallerResolver.withLeadsAndMembers(identity, false, null, leads, members);
log.info("auth: loopback-trust (any loopback non-worker caller is the primary)");
}
BridgeMcp mcp = new BridgeMcp(messages, workers, sessions, identity, presence,
primaryRegistry, callers, metrics);
primaryRegistry, callers, metrics, new BridgeMcp.CapacitySource(profile -> liveCountRef.get().apply(profile),
profile -> {
var configured = config.get().profiles().get(profile);
return configured == null ? null : configured.maxLoad();
}, () -> config.get().profiles().keySet(), System::nanoTime),
new BridgeMcp.HealthCoverageSource(() -> {
var health = config.get().health();
return FleetHealthMonitor.coverage(health != null && health.isEnabled(),
health != null && health.notifications() != null && health.notifications().configured());
}),
new BridgeMcp.QuarantineSource(profile -> {
var configured = config.get().profiles().get(profile);
return configured == null ? null : configured.effectiveCredentialId();
}, quarantine));
// CB-559: opt-in config reload. With no `configReload:` block nothing is constructed, so an
// upgraded daemon behaves exactly as before — the file is read once at boot and never again.
final ConfigWatcher configWatcher;
if (cfg.configReload() != null && cfg.configReload().isEnabled()) {
configWatcher = new ConfigWatcher(config, cfg.configReload().intervalSeconds());
configWatcher.start();
} else {
configWatcher = null;
}
// CB-303 part 3: single ordered shutdown hook. Drain sessions first while herdr is still
// open (so releases reach the daemon), then stop poller/message/mcp/reaper, and close herdr
@@ -248,6 +498,9 @@ public final class Bridged {
poller.stop();
messages.close();
pushLoop.close();
if (heartbeat != null) heartbeat.close(); // CB-551: stop the idle-lead heartbeat scheduler
if (healthMonitor != null) healthMonitor.stop();
if (configWatcher != null) configWatcher.stop(); // CB-559: stop polling the config file
mcp.close();
if (reaper != null) reaper.stop();
// Release the broker connection last among message resources (no-op for the in-memory inbox).
@@ -268,6 +521,29 @@ public final class Bridged {
cfg.bind().host(), cfg.bind().port(), socket);
}
/**
* The {@link Injector}'s readiness gate (CB-534): a target is deliverable if it is a spawned
* member whose agent has connected the bridge MCP, <em>or</em> a lead.
*
* <p>The gate exists for one reason — to hold a delivery out of a <em>spawned</em> member's boot
* window, where herdr already reports {@code idle} but the TUI would drop an injected paste. That
* hazard is a property of spawning. A lead is never spawned: the operator started it and named it
* (or labelled its tab) only once it was up, so there is no boot window to guard.
*
* <p>A lead is also never enrolled in {@link MemberPresence} — {@code BridgeMcp} marks presence
* for every spawned member (worker and architect), deliberately, since that map doubles as the
* member roster's availability signal and a lead counted there would show up as an available
* member. So without the second disjunct a lead is permanently un-deliverable: every
* lead→lead send sat on the gate for {@code READINESS_GRACE_POLLS} (~60s) and then failed
* having never been typed into the pane.
*
* <p>The lead set is read through the supplier on each call rather than snapshotted, so a lead
* discovered by {@code leadScan} after startup becomes deliverable without a restart.
*/
static Predicate<String> deliverableTo(MemberPresence presence, Supplier<Map<String, String>> leads) {
return target -> presence.isPresent(target) || leads.get().containsKey(target);
}
/**
* Poll herdr's {@code ping} until it answers or {@link #HERDR_WAIT_SECONDS} elapses (CB-504).
*
@@ -46,18 +46,28 @@ public final class Authz {
return false; // authenticated as nothing ⇒ authorized for nothing
}
return switch (action) {
// Orchestration is the primary's alone. A worker driving spawn/stop/send would be a
// worker escalating into the orchestrator role.
case SPAWN, STOP, SEND, DRAIN -> caller.isPrimary();
// Fleet lifecycle is the primary's alone — spawn, stop, drain. An architect
// deliberately does NOT get these (CB-548), so it cannot tear down or stand up workers
// even though it coordinates them; and a worker driving any of these would be a worker
// escalating into the orchestrator role.
case SPAWN, STOP, DRAIN -> caller.isPrimary();
// The load-bearing rule: a worker acts only as itself. The primary is deliberately
// excluded — a reply/ask is a worker's own turn output, and letting the primary forge
// one would corrupt the rendezvous correlation it is itself waiting on.
// Delivering a turn is open to the primary and the architect: an architect delegates
// to workers (that is the role's point) but still has no lifecycle rights. A worker is
// excluded — sending would be it escalating.
case SEND -> caller.isPrimary() || caller.isArchitect();
// The load-bearing rule: a caller acts only as the pane it occupies. CB-532 widened who
// that can be — a lead answering another lead is replying for its OWN terminal, which
// this already permits — while the rule itself is unchanged, and is what stops anyone
// forging a reply for a rendezvous someone else is waiting on. An architect's own pane
// passes through the same check, so it can answer a funnel that delegated to it. An
// unnamed primary (token/loopback, no pane) owns nothing and is still excluded.
case REPLY, ASK -> caller.ownsSession(targetSession);
// Observation is open to both authenticated roles: a worker legitimately polls its own
// Observation is open to every authenticated role: a worker legitimately polls its own
// status, and the roster carries no secrets.
case READ, METRICS -> caller.isPrimary() || caller.isWorker();
case READ, METRICS -> caller.isPrimary() || caller.isWorker() || caller.isArchitect();
};
}
@@ -1,9 +1,13 @@
package dev.ltms.bridged.auth;
import dev.ltms.bridged.mcp.ConnectionIdentity;
import dev.ltms.bridged.peer.MemberRole;
import java.nio.charset.StandardCharsets;
import java.security.MessageDigest;
import java.util.Map;
import java.util.function.Function;
import java.util.function.Supplier;
/**
* Resolves every caller to a {@link Principal}, for both entry paths into the core (CB-501).
@@ -15,7 +19,18 @@ import java.security.MessageDigest;
*
* <p><strong>Resolution order</strong> — connection identity first, token second, nothing third:
* <ol>
* <li>A loopback peer PID that maps to a herdr worker pane ⇒ {@link Role#WORKER}. This is
* <li>A loopback peer PID that maps to a pane named by {@code leaders:}, by the legacy
* {@code primary.terminal} pin, or by an operator-labelled lead tab (CB-307, CB-530, CB-531)
* ⇒ {@link Role#PRIMARY}, carrying that lead's
* name. The pane mapping is as unforgeable as a worker's, and the config explicitly names
* that pane as a lead's own — without this rule a lead running <em>inside</em> a herdr pane
* is misread as a worker and locked out of orchestration. More than one pane may be named,
* so two leads can work as peers rather than one being demoted.</li>
* <li>A loopback peer PID that maps to a pane bound to a CB-548 architect slot ⇒
* {@link Role#ARCHITECT}, carrying the slot name. Just unforgeable as a worker's, and
* resolved from the <em>live</em> terminal→slot binding (never a request argument), before
* the generic worker fallback.</li>
* <li>A loopback peer PID that maps to any other herdr pane ⇒ {@link Role#WORKER}. This is
* unforgeable (the OS reports the PID, herdr owns the PID→pane map) and is honoured
* regardless of auth mode, so enabling auth never breaks the fleet.</li>
* <li>Otherwise, under {@code token} mode, a valid bearer token ⇒ {@link Role#PRIMARY}.</li>
@@ -29,18 +44,122 @@ public final class CallerResolver {
private final ConnectionIdentity identity;
private final boolean tokenMode;
private final byte[] expectedToken; // null unless tokenMode
/**
* terminal_id → lead name; empty when nothing is pinned. CB-530.
*
* <p>A supplier rather than a map because the registry is no longer fixed at startup: CB-531
* discovers leads by scanning herdr for operator-labelled tabs, so a lead that opens its tab
* after the daemon booted must still be recognised. Consulted per resolve; the scanner behind
* it is TTL-cached, so this is a map lookup in the common case.
*/
private final Supplier<Map<String, String>> leadTerminals;
/**
* terminal_id → architect slot name; empty when nothing is configured. CB-548.
*
* <p>Like {@link #leadTerminals}, a supplier rather than a fixed map, so a binding injected
* after startup — when the later spawn lifecycle establishes a live architect session, or an
* operator pins one — takes effect without a restart. Consulted per resolve; today's wiring
* in {@code Bridged} reads a constant from config, which is the degenerate live case.
*/
private final Supplier<Map<String, String>> architectTerminals;
private final Function<String, MemberRole> memberSlotRoles;
private final Function<String, String> memberSlotNames;
/** Loopback-trust resolver: no token required, historical behaviour. */
public CallerResolver(ConnectionIdentity identity) {
this(identity, false, null);
/** Loopback-trust resolver: no token required, historical behaviour. Test-only. */
CallerResolver(ConnectionIdentity identity) {
this(identity, false, null, Map.of());
}
/** As {@link #CallerResolver(ConnectionIdentity, boolean, String, Map)} with no leads pinned. Test-only. */
CallerResolver(ConnectionIdentity identity, boolean tokenMode, String token) {
this(identity, tokenMode, token, Map.of());
}
/**
* @param identity connection-based worker identification
* @param tokenMode when true, a non-worker caller must present a valid bearer token
* @param token the expected bearer token; required (non-blank) when {@code tokenMode}
* Single-pin form for the legacy {@code primary.terminal}-only configuration — one lead, named
* {@code primary}.
*
* <p>A static factory rather than a fourth constructor overload on purpose: {@code String} and
* {@code Map} overloads are ambiguous for a literal {@code null} argument, which is a compile
* error at the call site and exactly the shape "unpinned" is written in.
*
* @param pinnedPrimaryTerminal the primary's own herdr {@code terminal_id}
* ({@code null}/blank = unpinned)
*/
public CallerResolver(ConnectionIdentity identity, boolean tokenMode, String token) {
static CallerResolver pinnedTo(ConnectionIdentity identity, boolean tokenMode,
String token, String pinnedPrimaryTerminal) {
return new CallerResolver(identity, tokenMode, token,
pinnedPrimaryTerminal == null || pinnedPrimaryTerminal.isBlank()
? Map.of() : Map.of(pinnedPrimaryTerminal, "primary"));
}
/**
* @param identity connection-based worker identification
* @param tokenMode when true, a non-worker caller must present a valid bearer token
* @param token the expected bearer token; required (non-blank) when {@code tokenMode}
* @param leadTerminals herdr {@code terminal_id} → lead name for every configured lead
* (CB-530). A caller resolving to one of these panes is that lead — a
* {@link Role#PRIMARY} — rather than a worker. Empty = nothing pinned,
* so every pane resolves as a worker.
*/
CallerResolver(ConnectionIdentity identity, boolean tokenMode, String token,
Map<String, String> leadTerminals) {
this(identity, tokenMode, token, fixed(leadTerminals), null);
}
/**
* Live-registry form: {@code leadTerminals} is consulted on every resolve, so leads discovered
* after startup (CB-531's tab scan) take effect without a restart.
*
* <p>A static factory rather than a fourth constructor overload, for the same reason as
* {@link #pinnedTo}: {@code Map} and {@code Supplier} overloads are ambiguous for a literal
* {@code null}.
*/
static CallerResolver withLeads(ConnectionIdentity identity, boolean tokenMode,
String token,
Supplier<Map<String, String>> leadTerminals) {
return new CallerResolver(identity, tokenMode, token, leadTerminals, null);
}
/**
* Live registry form that can confirm a bound slot is an architect slot.
*
* <p>This is the only public construction path. It keeps terminal bindings and slot roles in
* the same {@link MemberRegistry}, so a configured architect can resolve as an architect.
*/
public static CallerResolver withLeadsAndMembers(ConnectionIdentity identity,
boolean tokenMode, String token,
Supplier<Map<String, String>> leadTerminals,
MemberRegistry members) {
return new CallerResolver(identity, tokenMode, token, leadTerminals,
members == null ? null : members::snapshot,
members == null ? null : members::roleForSlot,
members == null ? null : members::nameForSlot);
}
private static Supplier<Map<String, String>> fixed(Map<String, String> leadTerminals) {
Map<String, String> snapshot = leadTerminals == null ? Map.of() : Map.copyOf(leadTerminals);
return () -> snapshot;
}
private CallerResolver(ConnectionIdentity identity, boolean tokenMode, String token,
Supplier<Map<String, String>> leadTerminals,
Supplier<Map<String, String>> architectTerminals) {
this(identity, tokenMode, token, leadTerminals, architectTerminals, null);
}
private CallerResolver(ConnectionIdentity identity, boolean tokenMode, String token,
Supplier<Map<String, String>> leadTerminals,
Supplier<Map<String, String>> architectTerminals,
Function<String, MemberRole> memberSlotRoles) {
this(identity, tokenMode, token, leadTerminals, architectTerminals, memberSlotRoles, null);
}
private CallerResolver(ConnectionIdentity identity, boolean tokenMode, String token,
Supplier<Map<String, String>> leadTerminals,
Supplier<Map<String, String>> architectTerminals,
Function<String, MemberRole> memberSlotRoles,
Function<String, String> memberSlotNames) {
if (tokenMode && (token == null || token.isBlank())) {
throw new IllegalArgumentException(
"auth.mode=token requires a non-empty token; check that the env var named by "
@@ -49,6 +168,34 @@ public final class CallerResolver {
this.identity = identity;
this.tokenMode = tokenMode;
this.expectedToken = tokenMode ? token.getBytes(StandardCharsets.UTF_8) : null;
this.leadTerminals = leadTerminals == null ? Map::of : leadTerminals;
this.architectTerminals = architectTerminals == null ? Map::of : architectTerminals;
this.memberSlotRoles = memberSlotRoles == null ? _ -> null : memberSlotRoles;
this.memberSlotNames = memberSlotNames == null ? Function.identity() : memberSlotNames;
}
/**
* The currently-recognised leads, {@code terminal_id → name} (CB-535).
*
* <p>Deliberately read from the same supplier {@link #resolve} consults, rather than from a
* second copy handed to the roster: a lead that is <em>listed</em> but would not <em>resolve</em>
* (or the reverse) is an address a peer cannot actually reach, and the two answers drifting apart
* is precisely the confusion this exists to end. Live, so a lead discovered by the tab scan after
* startup appears without a restart.
*/
public Map<String, String> leads() {
return leadTerminals.get();
}
/**
* The currently-recognised architect slots, {@code terminal_id → slot name} (CB-548).
*
* <p>Read from the same supplier {@link #resolve} consults, so a slot that is <em>listed</em>
* here but would not <em>resolve</em> (or the reverse) cannot drift apart. Live for the same
* reason as {@link #leads()}.
*/
public Map<String, String> members() {
return architectTerminals.get();
}
/**
@@ -61,6 +208,22 @@ public final class CallerResolver {
public Principal resolve(String remoteAddr, int remotePort, String authorizationHeader) {
ConnectionIdentity.Caller c = identity.resolve(remoteAddr, remotePort);
if (c.terminal() != null) {
String lead = leadTerminals.get().get(c.terminal());
if (lead != null) {
// The config names this pane as a lead's own. The pane mapping is exactly as
// unforgeable as a worker's, so it outranks the token path — no credential needed.
// Checked before the architect registry so a pane named in BOTH is still the lead
// (CB-548 preserves every existing leader behaviour).
return Principal.leader(lead, c.terminal(), c.pid());
}
String slot = architectTerminals.get().get(c.terminal());
if (slot != null && memberSlotRoles.apply(slot) == MemberRole.ARCHITECT) {
// The config/live binding names this pane as an architect slot's own. Same
// unforgeable pane mapping; the live binding, never a request argument, decides.
// Check the slot role too: this defence in depth prevents a bad lifecycle bind from
// escalating a dev or reviewer into an architect. Checked before the worker fallback.
return Principal.architect(memberSlotNames.apply(slot), c.terminal(), c.pid());
}
return Principal.worker(c.terminal(), c.pid()); // unforgeable; never token-gated
}
@@ -76,16 +239,6 @@ public final class CallerResolver {
return isLoopback(remoteAddr) ? Principal.primary(c.pid()) : Principal.anonymous();
}
/** The working directory of the calling process (CB-112 spawn cwd inheritance), or {@code null}. */
public String cwdForPid(long pid) {
return identity.cwdForPid(pid);
}
/** True when auth requires a bearer token of non-worker callers. */
public boolean tokenMode() {
return tokenMode;
}
private boolean presentedTokenMatches(String authorizationHeader) {
String presented = bearerValue(authorizationHeader);
if (presented == null) {
@@ -0,0 +1,21 @@
package dev.ltms.bridged.auth;
import dev.ltms.bridged.peer.MemberRole;
/** Optional session lifecycle hook for live member-slot bindings. */
public interface MemberLifecycle {
MemberLifecycle NONE = new MemberLifecycle() {
@Override
public void acquired(MemberRole role, String profile, String terminal) {
}
@Override
public void released(String terminal) {
}
};
void acquired(MemberRole role, String profile, String terminal);
void released(String terminal);
}
@@ -0,0 +1,232 @@
package dev.ltms.bridged.auth;
import dev.ltms.bridged.config.BridgedConfig;
import dev.ltms.bridged.peer.MemberRole;
import org.slf4j.Logger;
import org.slf4j.LoggerFactory;
import java.util.Collections;
import java.util.HashMap;
import java.util.LinkedHashMap;
import java.util.Map;
import java.util.Objects;
/**
* The architect-slot registry (CB-548): every gateway-local architect name and the strong-model
* profile it points at, plus the <em>live</em> bindings from a live architect's herdr terminal to
* its slot.
*
* <p>Two halves, split by who owns each:
* <ul>
* <li><b>slots</b> — configured once, keyed by the gateway-local unique name; each carries the
* {@code profile} reference the spawn lifecycle reads when it stands the slot up. A read-only
* snapshot taken at construction.</li>
* <li><b>terminal bindings</b> — owned by this registry and initially <em>empty</em>. Config
* declares no architect terminal, so at startup every slot is idle and nothing resolves to an
* architect; a session only becomes one when the spawn lifecycle {@linkplain #bind(String,
* String) binds} its terminal to a slot. {@link CallerResolver} reads this through
* {@link #snapshot()} to turn a pane into an {@link Role#ARCHITECT}.</li>
* </ul>
*
* <p>Spawning/lifecycle is deliberately a separate unit: this class only owns the bindings and
* exposes the map the resolver resolves against plus the profile lookup lifecycle will call.
* Nothing here creates or manages an architect session.
*/
public final class MemberRegistry implements MemberLifecycle {
private static final Logger log = LoggerFactory.getLogger(MemberRegistry.class);
/**
* One flattened {@code fleet:} entry.
*
* <p>Flattened because a slot name is unique only <em>within</em> its pool — {@code sonnet} may
* legitimately be both a developer and a reviewer — while a terminal binds to exactly one thing.
* The qualified {@link #key()} is what that binding uses.
*
* @param name the slot's key inside its pool
* @param role the pool it came from
* @param profile the backend it runs on
*/
public record Entry(String name, MemberRole role, String profile) {
/** {@code "architect:opus"} — unique across pools, unlike {@link #name()}. */
public String key() {
return role.wireName() + ":" + name;
}
}
private final Map<String, Entry> slots;
/** Live {@code terminal_id → qualified slot key}; guarded by {@code this}. */
private final Map<String, String> terminalToSlot = new HashMap<>();
/** Flatten every role pool in {@code fleet} into one registry. Leaders are not members. */
public MemberRegistry(BridgedConfig.Fleet fleet) {
Map<String, Entry> flat = new LinkedHashMap<>();
if (fleet != null) {
for (MemberRole role : MemberRole.values()) {
fleet.pool(role).forEach((name, slot) -> {
if (slot != null) {
Entry e = new Entry(name, role, slot.profile());
flat.put(e.key(), e);
}
});
}
}
this.slots = Collections.unmodifiableMap(flat);
}
/** The configured slots, keyed by qualified {@link Entry#key()}. Unmodifiable snapshot. */
public Map<String, Entry> slots() {
return slots;
}
/** The slots belonging to {@code role}, in definition order. */
public Map<String, Entry> slotsFor(MemberRole role) {
Map<String, Entry> out = new LinkedHashMap<>();
slots.forEach((key, e) -> {
if (e.role() == role) {
out.put(key, e);
}
});
return Collections.unmodifiableMap(out);
}
/**
* An immutable copy of the live {@code terminal_id → slot name} bindings.
*
* <p>Passed to {@link CallerResolver} as the source of architect identity, and what
* {@code bridge_whoami}/the roster will read to say which slot a pane hosts. Empty until the
* spawn lifecycle binds a slot.
*/
public Map<String, String> snapshot() {
synchronized (terminalToSlot) {
return Map.copyOf(terminalToSlot);
}
}
/** The slot a live terminal is bound to, or {@code null} if it is not an architect slot. */
public String slotForTerminal(String terminal) {
if (terminal == null) {
return null;
}
synchronized (terminalToSlot) {
return terminalToSlot.get(terminal);
}
}
/**
* The strong-model profile a slot runs under — what the spawn lifecycle reads.
*
* @return the slot's configured {@code profile}, or {@code null} if the slot is unknown or
* declares none
*/
public String profileForSlot(String slotName) {
Entry e = slots.get(slotName);
return (e == null || e.profile() == null) ? null : e.profile();
}
/** The role a qualified slot key belongs to, or {@code null} when the key is unknown. */
public MemberRole roleForSlot(String slotName) {
Entry e = slots.get(slotName);
return e == null ? null : e.role();
}
/** The unqualified configured name for a slot, or {@code null} if it is unknown. */
public String nameForSlot(String slotName) {
Entry e = slots.get(slotName);
return e == null ? null : e.name();
}
/** True when {@code slotName} is a configured architect slot. */
public boolean isSlot(String slotName) {
return slots.containsKey(slotName);
}
/**
* Bind {@code terminal} to {@code slot} (CB-548).
*
* <p>The spawn lifecycle calls this when it stands a slot up. The bind is atomic and preserves
* the two cardinality invariants: a terminal may occupy at most one slot, and a slot may host at
* most one terminal. Binding the same terminal to the same slot again is a harmless no-op.
*
* @param slot a configured slot name, or the bind is refused
* @param terminal the pane that will act as this architect
* @return {@code true} if the binding is now {@code terminal → slot}; {@code false} if it was
* refused — an unknown slot, a terminal already bound to a different slot, or a slot
* already hosting a different terminal
*/
public boolean bind(String slot, String terminal) {
if (slot == null || terminal == null || terminal.isBlank()) {
return false;
}
synchronized (terminalToSlot) {
if (!isSlot(slot)) {
return false; // unknown slot — nothing to bind to
}
String existingSlot = terminalToSlot.get(terminal);
if (existingSlot != null) {
return slot.equals(existingSlot); // already this slot (idempotent) or a different one
}
if (terminalToSlot.containsValue(slot)) {
return false; // slot already hosts a terminal — no second one
}
terminalToSlot.put(terminal, slot);
return true;
}
}
/**
* Compare-safe unbind of {@code expectedTerminal} from {@code slot} (CB-548).
*
* <p>The spawn lifecycle calls this when it tears a slot down. Only the exact binding
* {@code expectedTerminal → slot} is removed; if that terminal was since rebound to a different
* slot (or the slot to a different terminal), the call is a no-op returning {@code false} — a
* stale unbind must never remove a replacement.
*
* @param slot the slot the caller believes the terminal is bound to
* @param expectedTerminal the terminal it expects to be bound there
* @return {@code true} if {@code expectedTerminal → slot} was removed; {@code false} if nothing
* was (no such binding, or the binding had already moved)
*/
public boolean unbind(String slot, String expectedTerminal) {
if (slot == null || expectedTerminal == null) {
return false;
}
synchronized (terminalToSlot) {
String current = terminalToSlot.get(expectedTerminal);
if (current == null || !slot.equals(current)) {
return false; // absent, or a replacement/moved binding — leave it in place
}
terminalToSlot.remove(expectedTerminal);
return true;
}
}
/**
* Bind only architect sessions to a free slot with the resolved profile.
*
* <p>The role check is lifecycle policy. {@link CallerResolver} repeats it when resolving a
* binding, so a later lifecycle regression cannot turn a worker into an architect.
*/
@Override
public void acquired(MemberRole role, String profile, String terminal) {
if (role != MemberRole.ARCHITECT || terminal == null || terminal.isBlank()) {
return;
}
// slotsFor preserves definition order, so duplicate-profile slots use the first free one.
for (Entry entry : slotsFor(MemberRole.ARCHITECT).values()) {
if (Objects.equals(profile, entry.profile()) && bind(entry.key(), terminal)) {
return;
}
}
log.info("member slot: no free architect slot for profile={}; session remains a worker", profile);
}
/** Unbind a released terminal using the compare-safe registry operation. */
@Override
public void released(String terminal) {
String slot = slotForTerminal(terminal);
if (slot != null) {
unbind(slot, terminal);
}
}
}
@@ -5,52 +5,122 @@ package dev.ltms.bridged.auth;
* identifies which worker it is (CB-501).
*
* @param role what this caller is authorized to act as
* @param terminal the worker's herdr terminal id; {@code null} for {@code PRIMARY}/{@code ANONYMOUS}
* @param terminal the herdr {@code terminal_id} of the pane this caller occupies — a worker's, or
* (since CB-532) a named lead's; {@code null} for an unnamed primary resolved off a
* token or loopback trust, and for {@code ANONYMOUS}
* @param pid the connecting process id, or {@code -1} when not resolvable (audit context)
* @param name for a lead resolved from the CB-530 {@code leaders:} registry, which lead it is;
* for an architect resolved from the CB-548 {@code architects:} registry, which
* slot it occupies; {@code null} for every other caller, including an unnamed primary
*/
public record Principal(Role role, String terminal, long pid) {
public record Principal(Role role, String terminal, long pid, String name) {
/**
* Three-arg form for the callers that have no name to carry (workers, anonymous, and the
* token/loopback primary paths). Kept so adding CB-530's {@code name} did not churn every
* construction site — and so a reconstructed principal without a stashed name still works.
*/
public Principal(Role role, String terminal, long pid) {
this(role, terminal, pid, null);
}
/** A caller authenticated as nothing — the default when no check establishes anything else. */
public static Principal anonymous() {
return new Principal(Role.ANONYMOUS, null, -1);
}
/** The orchestrating session. */
/** The orchestrating session, unnamed (token or loopback-trust path). */
public static Principal primary(long pid) {
return new Principal(Role.PRIMARY, null, pid);
}
/**
* A named lead from the {@code leaders:} registry (CB-530).
*
* <p>Carries {@link Role#PRIMARY}: a lead <em>is</em> a primary as far as authorization goes,
* so every existing {@code isPrimary()} gate keeps working unchanged and the role table needed
* no new entry. The name is reporting only — it lets {@code bridge_whoami} say <em>which</em>
* lead is asking once more than one is configured.
*
* <p><strong>CB-532: a lead now carries the terminal it was matched by.</strong> Under CB-530 it
* deliberately did not, because {@code terminal} meant "which worker pane" everywhere and a
* non-null one would have enrolled the lead in the worker presence map. That reading was what
* made a lead unaddressable: {@link #ownsSession} could never be true for it, so
* {@code bridge_reply} was refused and one lead could send to another but never be answered.
* The terminal now means "which pane is this caller", the presence map keys on
* {@link #isSpawnedMember()} instead, and a lead is a peer that can both send and receive.
*/
public static Principal leader(String name, String terminal, long pid) {
return new Principal(Role.PRIMARY, terminal, pid, name);
}
/** A worker peer, identified by its herdr pane. */
public static Principal worker(String terminal, long pid) {
return new Principal(Role.WORKER, terminal, pid);
}
/**
* An architect (CB-548), identified by the slot it occupies and the pane bound to it.
*
* <p>Carries {@link Role#ARCHITECT}. {@code slotName} is reporting only — it lets
* {@code bridge_whoami} say <em>which</em> architect slot is asking, and it is the key the
* (future) spawn lifecycle reads a profile back from. Identity is the {@code terminal}: like a
* worker's it comes from the connection and the live terminal→slot binding, so
* {@code ownsSession} works exactly as it does for a worker — an architect acts as its own
* pane and no other.
*/
public static Principal architect(String slotName, String terminal, long pid) {
return new Principal(Role.ARCHITECT, terminal, pid, slotName);
}
public boolean isPrimary() {
return role == Role.PRIMARY;
}
public boolean isArchitect() {
return role == Role.ARCHITECT;
}
public boolean isWorker() {
return role == Role.WORKER;
}
/**
* Whether this caller is a spawned member with its own pane.
*
* <p>Both workers and architects are spawned members. A lead is excluded because recording it
* as present would count it as an available member in the roster.
*/
public boolean isSpawnedMember() {
return role == Role.WORKER || role == Role.ARCHITECT;
}
public boolean isAnonymous() {
return role == Role.ANONYMOUS;
}
/**
* Whether this caller may act <em>as</em> {@code sessionId} — the "own session only" rule that
* keeps one worker from replying or asking on another's behalf. Only a worker can own a
* session, and only its own.
* keeps one peer from replying or asking on another's behalf.
*
* <p>The rule is about <em>identity</em>, not rank: a caller may act as the pane it demonstrably
* occupies, and as no other. CB-532 dropped the extra {@code isWorker()} conjunct that used to
* be here. It was not what enforced the rule — {@code terminal.equals(sessionId)} is, and that
* terminal comes from the connection, so it cannot be forged either way. All the conjunct did
* was make a lead permanently unable to answer anyone, since a lead's terminal was null and a
* lead is not a worker. A caller with no terminal at all (an off-host or token-authenticated
* primary) still owns nothing, which is the case the null check covers.
*/
public boolean ownsSession(String sessionId) {
return isWorker() && terminal != null && terminal.equals(sessionId);
return terminal != null && terminal.equals(sessionId);
}
/** Short, non-sensitive description for audit lines and error details. */
public String describe() {
return switch (role) {
case WORKER -> "worker:" + terminal;
case PRIMARY -> "primary";
case ARCHITECT -> "architect:" + name;
case PRIMARY -> name == null ? "primary" : "leader:" + name;
case ANONYMOUS -> "anonymous";
};
}
@@ -24,6 +24,16 @@ public enum Role {
*/
WORKER,
/**
* A config-declared architect slot (CB-548): a gateway-local named session on a strong-model
* profile that coordinates and delegates turns but does not own the fleet. Unforgeable like a
* worker's — derived from the connection's pane and the live terminal→slot binding, never from
* a request argument. May {@code SEND} a turn, {@code REPLY}/{@code ASK} only as its own pane,
* and {@code READ}/{@code METRICS}; may <em>not</em> {@code SPAWN}/{@code STOP}/{@code DRAIN}
* (those stay the primary's, to keep lifecycle in one pair of hands).
*/
ARCHITECT,
/** Authenticated as nothing. Authorized for nothing but {@code /healthz}. */
ANONYMOUS
}
File diff suppressed because it is too large Load Diff
@@ -0,0 +1,291 @@
package dev.ltms.bridged.config;
import org.slf4j.Logger;
import org.slf4j.LoggerFactory;
import java.nio.file.Path;
import java.util.ArrayList;
import java.util.LinkedHashSet;
import java.util.List;
import java.util.Map;
import java.util.Objects;
import java.util.Set;
import java.util.concurrent.atomic.AtomicReference;
import java.util.function.Supplier;
/**
* The daemon's live configuration, re-readable without a restart (CB-559).
*
* <p>Consumers hold this, not a {@link BridgedConfig}, and read through {@link #get()} at the point
* of use. A component that captures {@code ref.get()} into a field at construction has opted out of
* reload — which is sometimes right (see <em>deferred</em> below), but it must then be a deliberate
* choice rather than an accident of where the field was initialised.
*
* <h2>Not every key can change under a running daemon</h2>
* Keys fall into three classes, and the difference is about what already exists when the reload
* happens — not about how important the key is.
*
* <ul>
* <li><strong>Hot</strong> — re-read per use, so a reload takes effect on the next spawn:
* {@code fleet:} (every role pool, {@code charters}, and {@code tabLabel}),
* {@code placement:}, and an existing profile's {@code weight} / {@code maxLoad}. Those
* three are read through a supplier on {@code CompositePeerLauncher}, which is what makes
* them hot — not the fact that they are config. <strong>This does NOT include
* {@code fleet.leaders}</strong>: {@code Bridged.main} reads {@code cfg.fleet().leaders()}
* once at startup to build the {@code LeadTabScanner} and the {@code LeadLauncher}, and
* neither is reconstructed on reload — so a lead added, removed, or re-{@code tab}'d under
* {@code fleet.leaders} needs a restart, the same as any deferred key below.</li>
* <li><strong>Deferred</strong> — accepted into the new snapshot, but the wiring built at startup
* keeps the old value until a restart: {@code lifecycle:}, {@code leadHeartbeat:},
* {@code spawnReadyTimeoutMs} / {@code spawnReadyPollMs}, {@code quarantineCooldownSeconds}
* (CB-578 stage B — baked once into the {@code BackendQuarantine} built at startup),
* {@code guard:}, {@code worktreeRoot:}, adding or removing a profile (a new backend needs its own launcher,
* which is constructed once), <em>and an existing profile's launch settings</em> —
* {@code model}, {@code baseUrl}, {@code argv}, {@code env}, {@code mcpUrl},
* {@code exhaustedPattern} (CB-578 stage A — compiled once into {@code Bridged.main}'s
* pattern map at startup), and the rest. {@code credentialId} (CB-578 stage B) is NOT on
* this list — it is read live off the config supplier at every quarantine check and
* exhaustion event, exactly like {@code weight} / {@code maxLoad}, so it is hot instead.
* {@code HerdrPeerLauncher} takes {@code Map.copyOf(profiles)} at construction and resolves
* each spawn out of that copy, so those never reach a launch until the daemon restarts. A
* reload logs these rather than pretending they applied.</li>
* <li><strong>Cold</strong> — cannot change at all under a running daemon: {@code bind:},
* {@code herdrSocket:}, {@code broker:} and {@code auth:}. The socket is bound, the broker
* connection is open, and the auth mode decides who may reach the port that is already
* listening.</li>
* </ul>
*
* <p><strong>A cold change refuses the whole reload.</strong> Not the hot half applied and the cold
* half warned about: that would leave the running daemon in a state matching no file on disk, which
* is the worst thing a reload can do to an operator debugging one. Refusing keeps the invariant that
* the live config is always some version of the file, and the message names the keys that must
* change through a restart.
*
* <p>A reload that fails to parse or fails validation is also refused, and the previous config keeps
* running. A config file being edited is normally read once mid-save; degrading a working daemon
* because it caught a half-written file would be a bad trade.
*/
public final class ConfigRef implements Supplier<BridgedConfig> {
private static final Logger log = LoggerFactory.getLogger(ConfigRef.class);
/** Keys that cannot change under a running daemon — see the class doc. */
private static final Set<String> COLD_KEYS =
Set.of("bind", "herdrSocket", "broker", "auth");
private final Path path;
private final AtomicReference<BridgedConfig> current;
public ConfigRef(Path path, BridgedConfig initial) {
this.path = path;
this.current = new AtomicReference<>(Objects.requireNonNull(initial, "initial config"));
}
/** A fixed reference that never reloads — for tests and for wiring built from a config in code. */
public static ConfigRef fixed(BridgedConfig cfg) {
return new ConfigRef(null, cfg);
}
/** The live configuration. Read this per use; do not cache it in a field. */
@Override
public BridgedConfig get() {
return current.get();
}
/** The file this ref reloads from, or {@code null} for a {@link #fixed} ref. */
public Path path() {
return path;
}
/**
* What a reload attempt did.
*
* @param applied true when the new config is now live
* @param coldKeys cold keys whose value changed, which is why an unapplied reload was refused
* @param deferred keys that changed and were accepted, but whose effect waits for a restart
* @param error the parse or validation failure that refused the reload, else {@code null}
*/
public record Outcome(boolean applied, List<String> coldKeys, List<String> deferred,
String error) {
public Outcome {
coldKeys = List.copyOf(coldKeys);
deferred = List.copyOf(deferred);
}
static Outcome refusedCold(List<String> keys) {
return new Outcome(false, keys, List.of(), null);
}
static Outcome failed(String error) {
return new Outcome(false, List.of(), List.of(), error);
}
/** A one-line summary for the operator — the reason, not just the verdict. */
public String summary() {
if (error != null) {
return "config reload refused — " + error;
}
if (!applied) {
return "config reload refused — these keys cannot change under a running daemon: "
+ String.join(", ", coldKeys) + ". Restart bridged to apply them.";
}
if (!deferred.isEmpty()) {
return "config reloaded; these changes need a restart to take effect: "
+ String.join(", ", deferred);
}
return "config reloaded";
}
}
/**
* Re-read the file, validate it, and swap it in when nothing cold changed.
*
* <p>Never throws: a reload is a best-effort operation on a daemon that is already serving, and
* a bad edit must not take it down. Every failure path leaves the previous config live and is
* reported through the returned {@link Outcome}.
*/
public Outcome reload() {
if (path == null) {
return Outcome.failed("this config was built in code and has no file to reload from");
}
BridgedConfig old = current.get();
BridgedConfig fresh;
try {
fresh = BridgedConfig.load(path);
// The same gate startup runs. A config that would have refused to boot must not be able
// to slip in through a reload — that is how a daemon ends up in a state it could never
// have started in, which is the hardest kind to debug.
fresh.validateAuthExposure();
fresh.validateLeadTabPrefixes();
fresh.validateSubscriptionProfiles();
fresh.validateCharters();
fresh.validateMembers();
} catch (RuntimeException e) {
String msg = e.getMessage() == null ? e.toString() : e.getMessage();
log.warn("config reload from {} refused, keeping the running config: {}", path, msg);
return Outcome.failed(msg);
}
List<String> cold = changedColdKeys(old, fresh);
if (!cold.isEmpty()) {
Outcome out = Outcome.refusedCold(cold);
log.warn(out.summary());
return out;
}
List<String> deferred = changedDeferredKeys(old, fresh);
current.set(fresh);
Outcome out = new Outcome(true, List.of(), deferred, null);
log.info(out.summary());
return out;
}
/** Cold keys whose value differs between the running config and the candidate. */
private static List<String> changedColdKeys(BridgedConfig old, BridgedConfig fresh) {
List<String> changed = new ArrayList<>();
if (!Objects.equals(old.bind(), fresh.bind())) {
changed.add("bind");
}
if (!Objects.equals(old.herdrSocket(), fresh.herdrSocket())) {
changed.add("herdrSocket");
}
if (!Objects.equals(old.broker(), fresh.broker())) {
changed.add("broker");
}
if (!Objects.equals(old.auth(), fresh.auth())) {
changed.add("auth");
}
// Kept in step with COLD_KEYS so the doc and the code cannot drift apart silently.
assert COLD_KEYS.containsAll(changed) : "a cold key was reported that COLD_KEYS omits";
return changed;
}
/** Changed keys that were accepted but whose effect waits for a restart. */
private static List<String> changedDeferredKeys(BridgedConfig old, BridgedConfig fresh) {
List<String> changed = new ArrayList<>();
if (!Objects.equals(old.lifecycle(), fresh.lifecycle())) {
changed.add("lifecycle");
}
if (!Objects.equals(old.leadHeartbeat(), fresh.leadHeartbeat())) {
changed.add("leadHeartbeat");
}
if (!Objects.equals(old.guard(), fresh.guard())) {
changed.add("guard");
}
if (!Objects.equals(old.worktreeRoot(), fresh.worktreeRoot())) {
changed.add("worktreeRoot");
}
if (!Objects.equals(old.spawnReadyTimeoutMs(), fresh.spawnReadyTimeoutMs())
|| !Objects.equals(old.spawnReadyPollMs(), fresh.spawnReadyPollMs())) {
changed.add("spawnReady*");
}
// CB-578 stage B: baked once into the BackendQuarantine built at startup — a running
// quarantine keeps its original cooldown regardless, and a new cooldown only applies to a
// quarantine that starts after a restart.
if (!Objects.equals(old.quarantineCooldownSeconds(), fresh.quarantineCooldownSeconds())) {
changed.add("quarantineCooldownSeconds");
}
Map<String, BridgedConfig.Profile> before =
old.profiles() == null ? Map.of() : old.profiles();
Map<String, BridgedConfig.Profile> after =
fresh.profiles() == null ? Map.of() : fresh.profiles();
// Adding or removing a profile is deferred: a new backend needs its own launcher, and
// launchers are built once at startup.
if (!before.keySet().equals(after.keySet())) {
Set<String> diff = new LinkedHashSet<>(before.keySet());
diff.addAll(after.keySet());
diff.removeIf(p -> before.containsKey(p) && after.containsKey(p));
changed.add("profiles (added/removed: " + String.join(", ", diff) + ")");
}
// An EXISTING profile's launch settings are deferred too, and this is easy to get wrong:
// `HerdrPeerLauncher` takes `Map.copyOf(profiles)` at construction and `spawn` resolves the
// profile out of that snapshot, so a reloaded model/baseUrl/argv/env never reaches a launch.
// Only weight and maxLoad are genuinely hot, because placement reads them through the
// supplier on the composite rather than from the adapter's copy. Without this check a
// changed model would report "config reloaded" and silently do nothing — the worst outcome
// a reload can produce, because the operator has no reason to doubt it.
List<String> relaunch = new ArrayList<>();
before.forEach((name, was) -> {
BridgedConfig.Profile now = after.get(name);
if (now != null && !sameLaunchSettings(was, now)) {
relaunch.add(name);
}
});
if (!relaunch.isEmpty()) {
changed.add("profiles." + String.join("/", relaunch) + " launch settings "
+ "(model, baseUrl, argv, env, …) — the launcher holds a startup snapshot");
}
return changed;
}
/**
* Whether two versions of a profile would launch a peer identically. Compares every component
* the launcher reads at spawn; {@code weight}, {@code maxLoad} and {@code credentialId} are
* excluded because those are read live (by the placement policy and, for credentialId, by
* {@code CompositePeerLauncher}/the CB-578 stage B exhaustion sink) and really do take effect on
* the next spawn.
*/
private static boolean sameLaunchSettings(BridgedConfig.Profile a, BridgedConfig.Profile b) {
return Objects.equals(a.baseUrl(), b.baseUrl())
&& Objects.equals(a.model(), b.model())
&& Objects.equals(a.configDir(), b.configDir())
&& Objects.equals(a.tokenEnv(), b.tokenEnv())
&& Objects.equals(a.argv(), b.argv())
&& Objects.equals(a.placement(), b.placement())
&& Objects.equals(a.workspace(), b.workspace())
&& Objects.equals(a.tabLabel(), b.tabLabel())
&& Objects.equals(a.mcpUrl(), b.mcpUrl())
&& Objects.equals(a.cwd(), b.cwd())
&& Objects.equals(a.parityOverlay(), b.parityOverlay())
&& Objects.equals(a.gitTokenEnv(), b.gitTokenEnv())
&& Objects.equals(a.gitHostEnv(), b.gitHostEnv())
&& Objects.equals(a.kind(), b.kind())
&& Objects.equals(a.env(), b.env())
&& Objects.equals(a.subscription(), b.subscription())
// CB-578 stage B: exhaustedPattern is compiled once into Bridged.main's pattern map
// at startup (see ExhaustedPatternLookup wiring) — a reload never re-reads it, so a
// changed pattern must be reported as deferred, exactly like model/baseUrl/argv.
&& Objects.equals(a.exhaustedPattern(), b.exhaustedPattern());
}
}
@@ -0,0 +1,96 @@
package dev.ltms.bridged.config;
import org.slf4j.Logger;
import org.slf4j.LoggerFactory;
import java.io.IOException;
import java.nio.file.Files;
import java.nio.file.Path;
import java.util.concurrent.Executors;
import java.util.concurrent.ScheduledExecutorService;
import java.util.concurrent.TimeUnit;
/**
* Polls {@code bridged.yaml}'s modified time and asks {@link ConfigRef} to reload when it moves
* (CB-559). Opt-in through {@code configReload.enabled}.
*
* <p><strong>Why polling and not a filesystem watch.</strong> {@code WatchService} on macOS has no
* native backend — it falls back to polling internally anyway, at an interval this code does not
* control — and editors save config files in ways that produce a different event mix per editor
* (write-in-place, write-and-rename, write-temp-and-swap). A modified-time check treats all of them
* the same and is a single {@code stat} per tick, which at a ten-second cadence costs nothing worth
* measuring.
*
* <p><strong>A missing or unreadable file is not a reason to act.</strong> Many editors briefly
* unlink the file during a save. Reloading on "it vanished" would mean reloading from a file that no
* longer exists; reporting an error every tick would bury the log. So an unreadable file is skipped
* silently and the next tick tries again — the running config stays live, which is the correct
* outcome either way.
*/
public final class ConfigWatcher {
private static final Logger log = LoggerFactory.getLogger(ConfigWatcher.class);
private final ConfigRef ref;
private final long intervalSeconds;
private final ScheduledExecutorService scheduler;
private volatile long lastSeenMillis;
public ConfigWatcher(ConfigRef ref, long intervalSeconds) {
this.ref = ref;
this.intervalSeconds = intervalSeconds;
this.lastSeenMillis = modifiedMillis(ref.path());
this.scheduler = Executors.newSingleThreadScheduledExecutor(r -> {
Thread t = new Thread(r, "config-watcher");
// A daemon thread: an operator's config watch must never be the reason the JVM refuses
// to exit after everything else has shut down.
t.setDaemon(true);
return t;
});
}
/** Begin watching. A ref with no file (a fixed one) is a no-op rather than an error. */
public void start() {
if (ref.path() == null) {
log.debug("config watch not started — this config has no file behind it");
return;
}
scheduler.scheduleWithFixedDelay(this::tick, intervalSeconds, intervalSeconds,
TimeUnit.SECONDS);
log.info("config watch: {} re-read when it changes (every {}s)", ref.path(), intervalSeconds);
}
/** One poll. Never throws — an exception here would silently cancel the schedule. */
void tick() {
try {
long now = modifiedMillis(ref.path());
if (now == 0 || now == lastSeenMillis) {
return;
}
// Stamp BEFORE reloading. A file whose reload is refused (a bad edit, or a cold key)
// must not be retried every tick — that would log the same refusal forever. The next
// save moves the timestamp again and earns a fresh attempt.
lastSeenMillis = now;
ref.reload();
} catch (RuntimeException e) {
log.warn("config watch tick failed, still watching: {}", e.getMessage());
}
}
private static long modifiedMillis(Path path) {
if (path == null) {
return 0;
}
try {
return Files.getLastModifiedTime(path).toMillis();
} catch (IOException e) {
return 0; // mid-save, or gone: say nothing and try again next tick
}
}
/** Stop polling. Called from the daemon's ordered shutdown hook, alongside the other loops. */
public void stop() {
scheduler.shutdownNow();
}
}
@@ -0,0 +1,40 @@
package dev.ltms.bridged.health;
import dev.ltms.bridged.herdr.AgentStatus;
import dev.ltms.bridged.session.MemberSession;
/**
* Pure classifier. Collection and repair are deliberately outside this package.
* {@link HealthState#ERROR_ON_SCREEN} is not decided yet because it needs a bounded pane detection
* read and an adapter-specific fatal signature; status facts alone must not guess it.
*/
public final class FleetHealth {
private FleetHealth() { }
public static HealthDecision decide(HealthSnapshot s, HealthPrior prior, long nowNanos) {
if (s.controlLinkDown()) return result(HealthState.CONTROL_LINK_DOWN, false);
if (s.targetNotFound()) return result(HealthState.GONE, false);
if (s.sessionState() == MemberSession.State.SPAWNING && !s.present() && s.readinessGraceElapsed()) {
return result(HealthState.NEVER_READY, false);
}
if (s.orphanedDelegation()) return result(HealthState.DELEGATION_ORPHANED, false);
boolean disagreement = s.sessionState() == MemberSession.State.BUSY && s.acceptedDelivery()
&& (s.liveStatus() == AgentStatus.IDLE || s.liveStatus() == AgentStatus.DONE);
if (disagreement && prior.busyButDone()) return result(HealthState.TURN_BOUNDARY_LOST, true);
if (s.stalled()) return result(HealthState.STALL_SUSPECTED, disagreement);
if (s.replyStranded()) return result(HealthState.REPLY_STRANDED, disagreement);
if (s.queuedDelivery() || s.inboxMessage()) return result(HealthState.WORK_PENDING, disagreement);
if (s.sessionState() == MemberSession.State.SPAWNING) return result(HealthState.STARTING, disagreement);
if (s.acceptedDelivery() && s.liveStatus() == AgentStatus.BLOCKED) {
return result(HealthState.BLOCKED_AMBIGUOUS, disagreement);
}
// An accepted delivery remains bridge work even when herdr is late, unknown, or has already
// reported DONE once. It cannot be IDLE until the delegation has resolved.
if (s.acceptedDelivery()) return result(HealthState.WORKING, disagreement);
return result(HealthState.IDLE, disagreement);
}
private static HealthDecision result(HealthState state, boolean disagreement) {
return new HealthDecision(state, new HealthPrior(disagreement));
}
}
@@ -0,0 +1,151 @@
package dev.ltms.bridged.health;
import dev.ltms.bridged.herdr.Agent;
import dev.ltms.bridged.herdr.AgentControl;
import dev.ltms.bridged.herdr.AgentStatus;
import dev.ltms.bridged.msg.MessageService;
import dev.ltms.bridged.session.MemberSession;
import org.slf4j.Logger;
import org.slf4j.LoggerFactory;
import java.util.HashMap;
import java.util.HashSet;
import java.util.List;
import java.util.Map;
import java.util.Objects;
import java.util.concurrent.ScheduledExecutorService;
import java.util.concurrent.TimeUnit;
import java.util.function.BiConsumer;
import java.util.function.LongSupplier;
import java.util.function.Supplier;
/** Slow whole-fleet evidence collection. It is deliberately separate from the delivery poller. */
public final class FleetHealthMonitor {
private static final Logger log = LoggerFactory.getLogger(FleetHealthMonitor.class);
/** Bounded attempts to run {@link #failTarget} for one transition. Never retried tick-to-tick (CB-580). */
static final int MAX_FAIL_TARGET_ATTEMPTS = 3;
private final AgentControl agents;
private final Supplier<List<MemberSession>> roster;
private final MessageService messages;
private final ScheduledExecutorService scheduler;
private final LongSupplier clock;
private final long intervalSeconds;
private final BiConsumer<String, String> failTarget;
private final Map<String, HealthPrior> priors = new HashMap<>();
private final Map<String, HealthState> states = new HashMap<>();
// These facts need the evidence publishers introduced by later M4 units. They are not negatives.
private static final boolean NOT_YET_OBSERVED = false;
/**
* @param failTarget CB-568's idempotent target-wide failure operation (e.g. {@code messages::abandon}),
* invoked once when a member transitions into a terminal health state. Required —
* there is deliberately no defaulting overload; a caller that does not want the
* fail-tickets-on-terminal-health behavior must pass an explicit inert value (see
* {@code TestTurnTokens.inert} / {@code BridgeMcp.CapacitySource.none()} for the pattern).
*/
public FleetHealthMonitor(AgentControl agents, Supplier<List<MemberSession>> roster, MessageService messages,
ScheduledExecutorService scheduler, LongSupplier clock, long intervalSeconds,
BiConsumer<String, String> failTarget) {
this.agents = agents;
this.roster = roster;
this.messages = messages;
this.scheduler = scheduler;
this.clock = clock;
this.intervalSeconds = intervalSeconds;
this.failTarget = Objects.requireNonNull(failTarget, "failTarget");
}
/** Pure per-member decision seam. */
static HealthDecision decide(HealthSnapshot snapshot, HealthPrior prior, long nowNanos) {
return FleetHealth.decide(snapshot, prior, nowNanos);
}
public void start() { scheduler.schedule(this::tick, intervalSeconds, TimeUnit.SECONDS); }
public void stop() { scheduler.shutdownNow(); }
// Package-private so tests can run one tick without waiting.
void tick() {
try {
List<Agent> agentsNow = agents.list(); // Exactly one list call for this complete observation.
List<MemberSession> rosterNow = roster.get(); // One in-memory roster snapshot for this tick.
Map<String, Agent> live = new HashMap<>();
for (Agent agent : agentsNow) live.put(agent.terminalId(), agent);
HashSet<String> current = new HashSet<>();
for (MemberSession session : rosterNow) {
current.add(session.terminalId());
Agent agent = live.get(session.terminalId());
AgentStatus status = agent == null ? AgentStatus.UNKNOWN : agent.status();
boolean accepted = messages.hasAcceptedDelivery(session.terminalId());
HealthSnapshot snapshot = new HealthSnapshot(session.state(), status, accepted, NOT_YET_OBSERVED,
messages.hasInboxMessage(session.terminalId()), agent != null, NOT_YET_OBSERVED,
NOT_YET_OBSERVED, NOT_YET_OBSERVED, NOT_YET_OBSERVED, NOT_YET_OBSERVED, NOT_YET_OBSERVED);
HealthDecision decision = decide(snapshot, priors.getOrDefault(session.terminalId(), HealthPrior.NONE),
clock.getAsLong());
priors.put(session.terminalId(), decision.prior());
reportTransition(session.terminalId(), decision.state());
}
priors.keySet().retainAll(current);
states.keySet().retainAll(current);
} catch (Throwable error) {
// A list failure is health evidence, and must never kill the monitor's only scheduler task.
log.warn("fleet health collection failed; will retry next tick", error);
} finally {
if (!scheduler.isShutdown()) {
scheduler.schedule(this::tick, intervalSeconds, TimeUnit.SECONDS);
}
}
}
void reportTransition(String target, HealthState next) {
HealthState previous = states.put(target, next);
if (previous == next) return;
if (fault(next)) {
log.warn("fleet health member={} state={} previous={}", target, next, previous);
} else if (previous != null && fault(previous)) {
log.info("fleet health member={} recovered state={} previous={}", target, next, previous);
}
// CB-580: a member entering GONE/NEVER_READY must not leave its waiting tickets pending
// forever. Fire exactly once per transition — never on a tick where the state is unchanged,
// which is what made the rejected commit call abandon() once per tick for as long as a
// member stayed terminal.
if (terminal(next)) {
failTerminalTarget(target, next);
}
}
private void failTerminalTarget(String target, HealthState state) {
String reason = "fleet health: member reached terminal state " + state.name();
RuntimeException last = null;
for (int attempt = 1; attempt <= MAX_FAIL_TARGET_ATTEMPTS; attempt++) {
try {
failTarget.accept(target, reason);
return;
} catch (RuntimeException error) {
last = error;
log.warn("fleet health: failTarget attempt {}/{} failed for member={} state={}",
attempt, MAX_FAIL_TARGET_ATTEMPTS, target, state, error);
}
}
log.warn("fleet health: giving up on failTarget for member={} state={} after {} attempts",
target, state, MAX_FAIL_TARGET_ATTEMPTS, last);
}
private static boolean terminal(HealthState state) {
return state == HealthState.GONE || state == HealthState.NEVER_READY;
}
private static boolean fault(HealthState state) {
return switch (state) {
case NEVER_READY, GONE, TURN_BOUNDARY_LOST, ERROR_ON_SCREEN, STALL_SUSPECTED,
MUTE, REPLY_STRANDED, DELEGATION_ORPHANED, CONTROL_LINK_DOWN -> true;
default -> false;
};
}
public static String coverage(boolean enabled, boolean notificationConfigured) {
return !enabled ? "off" : notificationConfigured ? "full" : "detection-only";
}
}
@@ -0,0 +1,4 @@
package dev.ltms.bridged.health;
/** Classification plus the private fact that the next pure decision needs. */
public record HealthDecision(HealthState state, HealthPrior prior) { }
@@ -0,0 +1,6 @@
package dev.ltms.bridged.health;
/** Private cross-tick observation. It is deliberately not a reported health value. */
public record HealthPrior(boolean busyButDone) {
public static final HealthPrior NONE = new HealthPrior(false);
}
@@ -0,0 +1,11 @@
package dev.ltms.bridged.health;
import dev.ltms.bridged.herdr.AgentStatus;
import dev.ltms.bridged.session.MemberSession;
/** Read-only facts from one fleet collection tick. */
public record HealthSnapshot(MemberSession.State sessionState, AgentStatus liveStatus,
boolean acceptedDelivery, boolean queuedDelivery, boolean inboxMessage,
boolean present, boolean targetNotFound, boolean controlLinkDown,
boolean readinessGraceElapsed, boolean orphanedDelegation,
boolean replyStranded, boolean stalled) { }
@@ -0,0 +1,8 @@
package dev.ltms.bridged.health;
/** Health classifications reported for a member. */
public enum HealthState {
STARTING, IDLE, WORKING, WORK_PENDING, BLOCKED_AMBIGUOUS,
NEVER_READY, GONE, TURN_BOUNDARY_LOST, ERROR_ON_SCREEN, STALL_SUSPECTED,
MUTE, REPLY_STRANDED, DELEGATION_ORPHANED, CONTROL_LINK_DOWN
}
@@ -0,0 +1,25 @@
package dev.ltms.bridged.health;
import dev.ltms.bridged.msg.Rendezvous;
import java.util.Map;
import java.util.concurrent.ConcurrentHashMap;
/**
* Counts turns that ended via the completion fallback instead of {@code bridge_reply}.
* MUTE is an observation by target and profile, not a classifier state and never suppresses faults.
*/
public final class MuteCounter {
private final Map<String, Integer> byTarget = new ConcurrentHashMap<>();
private final Map<String, Integer> byProfile = new ConcurrentHashMap<>();
/** Record only fallback completion; a structured reply does not make a member mute. */
public void observe(String target, String profile, Rendezvous.Kind kind) {
if (kind != Rendezvous.Kind.COMPLETION) return;
byTarget.merge(target, 1, Integer::sum);
byProfile.merge(profile, 1, Integer::sum);
}
public int forTarget(String target) { return byTarget.getOrDefault(target, 0); }
public int forProfile(String profile) { return byProfile.getOrDefault(profile, 0); }
}
@@ -0,0 +1,26 @@
package dev.ltms.bridged.health;
import java.util.ArrayList;
import java.util.HashMap;
import java.util.List;
import java.util.Map;
/** Fixed pane-probe limits. Pane content is never retained here. */
public final class PaneBudget {
public static final long COOLDOWN_NANOS = 60_000_000_000L;
public static final int MAX_PER_TICK = 2;
private final Map<String, Long> lastProbe = new HashMap<>();
private int cursor;
public List<String> choose(List<String> candidates, long nowNanos, long configuredCooldownNanos) {
long cooldown = Math.max(COOLDOWN_NANOS, configuredCooldownNanos);
List<String> out = new ArrayList<>();
for (int n = 0; n < candidates.size() && out.size() < MAX_PER_TICK; n++) {
String target = candidates.get((cursor + n) % candidates.size());
Long last = lastProbe.get(target);
if (last == null || nowNanos - last >= cooldown) { out.add(target); lastProbe.put(target, nowNanos); }
}
if (!candidates.isEmpty()) cursor = (cursor + 1) % candidates.size();
return List.copyOf(out);
}
}
@@ -6,93 +6,120 @@ import java.util.ArrayList;
import java.util.LinkedHashMap;
import java.util.List;
import java.util.Map;
import java.util.concurrent.ConcurrentHashMap;
/**
* Domain layer over herdr's native {@code agent.*} namespace — the worker south side.
* Chosen in the CB-102 spike over the pane + {@code send_text} fallback because
* {@code agent.start} takes a first-class {@code env} map (clean, guard-checked
* subscription injection) and herdr tracks each worker's Claude session UUID itself.
* Chosen in the CB-102 spike over the pane + {@code send_text} fallback because herdr
* tracks each worker's Claude session UUID itself.
*
* <p>Ported to herdr protocol 19 (herdr 0.8.0, CB-521): {@code agent.start} now starts a
* <em>supported</em> agent ({@code kind}) into an <em>existing</em> pane, so the worker's
* {@code env}/{@code cwd} move to pane creation ({@code tab.create}/{@code pane.split} — see
* {@link WorkspaceControl}), and {@code agent.send} is replaced by {@code agent.prompt}
* (which submits in one call) plus {@code agent.send_keys} for the raw Enter nudge.
*
* <p>Every method is one herdr call through the injected {@link HerdrClient}, so this
* layer is unit-testable with a fake and contract-tested against a live daemon.
*/
public final class AgentControl {
/**
* The keystroke that submits a prompt in the Claude Code TUI: a carriage return (Enter).
* It must be delivered as its <em>own</em> {@code agent.send} call — herdr delivers a message's
* text as a bracketed paste, and a {@code "\r"} appended to that same text is swallowed as
* literal newline content, not a submit. Sent as a separate keystroke event it lands outside
* the paste and submits. (A bare {@code "\n"} inserts a newline either way.) Verified live
* against Claude Code v2.1.210: an injected task stayed unsubmitted with {@code "text\r"} in
* one call, and submitted the instant a standalone {@code "\r"} was sent.
*/
static final String SUBMIT_KEY = "\r";
private final HerdrClient herdr;
/**
* Protocol 19 dropped {@code terminal_id} as an {@code agent.*} target — herdr now resolves
* targets by pane id or agent name only, while the bridge keys every session on the terminal.
* This caches the terminal→pane mapping (stable for a worker's lifetime) so callers keep
* addressing agents by terminal; entries are invalidated on {@code agent_not_found}.
*/
private final Map<String, String> paneByTerminal = new ConcurrentHashMap<>();
public AgentControl(HerdrClient herdr) {
this.herdr = herdr;
}
/** One agent-targeted call, translating a terminal id to its pane id (retrying once fresh). */
private JsonNode agentCall(String method, String target, Map<String, Object> extra) {
String resolved = resolveTarget(target);
try {
return herdr.call(method, withTarget(resolved, extra));
} catch (HerdrException e) {
if (!"agent_not_found".equals(e.code()) || resolved.equals(target)) throw e;
paneByTerminal.remove(target); // the cached pane went away — re-resolve once
String fresh = resolveTarget(target);
if (fresh.equals(resolved)) throw e;
return herdr.call(method, withTarget(fresh, extra));
}
}
private static Map<String, Object> withTarget(String target, Map<String, Object> extra) {
Map<String, Object> m = new LinkedHashMap<>();
m.put("target", target);
m.putAll(extra);
return m;
}
/** The pane id behind a terminal-id target, or the target verbatim for pane ids / names. */
private String resolveTarget(String target) {
if (target == null || !target.startsWith("term_")) {
return target;
}
String cached = paneByTerminal.get(target);
if (cached != null) {
return cached;
}
for (JsonNode a : herdr.call("agent.list").path("agents")) {
if (target.equals(a.path("terminal_id").asText(null))) {
String pane = a.path("pane_id").asText(null);
if (pane != null) {
paneByTerminal.put(target, pane);
return pane;
}
}
}
return target; // unknown terminal — let herdr report it against the original target
}
/**
* Spawn an agent. {@code env} is applied to the process environment verbatim — this
* is where a worker's {@code ANTHROPIC_BASE_URL} lives, and the ONLY place it should.
* Start an agent into {@code paneId}, which must be sitting at its interactive shell prompt —
* the seed pane of a freshly-created worker tab, or a fresh split. The pane's shell already
* carries the worker's env ({@code ANTHROPIC_BASE_URL}, token, …) and cwd from pane creation;
* herdr resolves the executable from {@code kind} and waits (its default timeout) until the
* agent is detected and ready for input.
*
* @param name label/kind for herdr status detection (e.g. {@code "claude"})
* @param argv launch command, e.g. {@code ["claude"]}
* @param env process environment additions ({@code ANTHROPIC_BASE_URL}, token, …)
* @param name unique label for this agent ({@code <kind>-<profile>-<nonce>-<seq>})
* @param kind supported agent kind and canonical executable, e.g. {@code "claude"},
* {@code "opencode"}
* @param args extra arguments after the executable, e.g. {@code --mcp-config …}
* @param paneId the pane to start the agent in
*/
public Agent start(String name, List<String> argv, Map<String, String> env) {
return start(name, argv, env, null);
}
/** Spawn an agent into {@code tabId} at herdr's default cwd. */
public Agent start(String name, List<String> argv, Map<String, String> env, String tabId) {
return start(name, argv, env, tabId, null);
}
/**
* Spawn an agent. With a non-null {@code tabId} the worker lands in that tab (the placement
* policy's dedicated worker tab); with {@code null} herdr splits the currently-focused tab
* (legacy pane placement). A non-blank {@code cwd} sets the worker process's working directory —
* {@code agent.start} honours {@code cwd} directly (an agent pane does <em>not</em> inherit the
* tab's or workspace's cwd, so this is the only way to root a worker in the primary's directory;
* CB-112).
*/
public Agent start(String name, List<String> argv, Map<String, String> env, String tabId, String cwd) {
public Agent start(String name, String kind, List<String> args, String paneId) {
Map<String, Object> params = new LinkedHashMap<>();
params.put("name", name);
params.put("argv", argv);
params.put("env", env);
if (tabId != null) {
params.put("tab_id", tabId);
}
if (cwd != null && !cwd.isBlank()) {
params.put("cwd", cwd);
}
params.put("kind", kind);
params.put("pane_id", paneId);
params.put("args", args);
JsonNode result = herdr.call("agent.start", params);
return Agent.from(result.get("agent"));
}
/**
* Deliver {@code text} to an agent as its next prompt <em>and submit it</em> — two keystroke
* events: the message (a bracketed paste, so any embedded newlines are preserved verbatim),
* then a standalone {@link #SUBMIT_KEY} (Enter) that actually submits it. Without the second
* event the text just sits in the worker's input box, never processed (see {@link #SUBMIT_KEY}).
* Deliver {@code text} to an agent as its next prompt <em>and submit it</em> — herdr's
* {@code agent.prompt} pastes the text (embedded newlines preserved verbatim) and submits it
* in the same call, replacing the pre-protocol-19 two-event {@code agent.send} dance.
*/
public void send(String target, String text) {
herdr.call("agent.send", Map.of("target", target, "text", text));
herdr.call("agent.send", Map.of("target", target, "text", SUBMIT_KEY));
agentCall("agent.prompt", target, Map.of("text", text));
}
/**
* Re-send the submit keystroke (Enter) to {@code target}. The Enter that accompanies a delivery
* can race the paste — especially right as the worker's TUI becomes interactive — leaving the
* text unsubmitted; the injector nudges it with this until the worker actually picks up (CB-113).
* Re-send the submit keystroke (Enter) to {@code target}. The submit that accompanies a
* delivery can race the paste — especially right as the worker's TUI becomes interactive —
* leaving the text unsubmitted; the injector nudges it with this until the worker actually
* picks up (CB-113).
*/
public void submit(String target) {
herdr.call("agent.send", Map.of("target", target, "text", SUBMIT_KEY));
agentCall("agent.send_keys", target, Map.of("keys", List.of("enter")));
}
/**
@@ -101,13 +128,13 @@ public final class AgentControl {
* @param source one of {@code visible|recent|recent_unwrapped|detection}
*/
public String read(String target, String source) {
JsonNode result = herdr.call("agent.read", Map.of("target", target, "source", source));
JsonNode result = agentCall("agent.read", target, Map.of("source", source));
return result.path("read").path("text").asText("");
}
/** Current agent record (status, session UUID, pane). */
public Agent get(String target) {
return Agent.from(herdr.call("agent.get", Map.of("target", target)).get("agent"));
return Agent.from(agentCall("agent.get", target, Map.of()).get("agent"));
}
/** Just the lifecycle status — what the status-gated injector checks before send. */
@@ -0,0 +1,188 @@
package dev.ltms.bridged.herdr;
import com.fasterxml.jackson.databind.JsonNode;
import org.slf4j.Logger;
import org.slf4j.LoggerFactory;
import java.util.Collections;
import java.util.LinkedHashMap;
import java.util.Locale;
import java.util.Map;
import java.util.Set;
import java.util.function.LongSupplier;
import java.util.function.Supplier;
/**
* Discovers which panes host a lead by scanning herdr for tabs the operator labelled by convention
* (CB-531), and hands {@link dev.ltms.bridged.auth.CallerResolver} the resulting
* {@code terminal_id → lead name} map.
*
* <p><strong>Why scan at all.</strong> A lead is never spawned — a human opens a tab and starts an
* agent in it — so the daemon cannot learn a lead's {@code terminal_id} at creation time the way it
* does a worker's. CB-530 solved that by having the operator paste each id into {@code leaders:},
* which works but costs a config edit and a daemon restart per lead, and the id is only obtainable
* by first starting the session and asking it. Scanning closes that loop: label the tab, and the
* pane is recognised on the next resolve.
*
* <p><strong>CB-579 — matched by name, not prefix.</strong> This used to strip one shared
* {@code tabPrefix} off a label to derive the lead's name, and merged a config-supplied
* {@code terminal_id} pin over every scan result so the pin could never expire. Both are gone: each
* lead now configures its own exact {@code tab} label ({@code fleet.leaders.<name>.tab}), so this
* class is handed a {@code tab → name} map up front and matches labels against it exactly
* (case-insensitively). There is no merge step — a scan result is the whole answer. That is the
* fix for the bug this replaces: a {@code terminal_id} pin surviving in config after the pane it
* named was gone, so the daemon kept treating a dead session as a live lead forever.
*
* <p><strong>Direction of trust.</strong> The label names the lead; it never <em>grants</em>
* anything a pane could take for itself. Three properties keep that honest:
* <ol>
* <li>Worker spaces are excluded wholesale ({@code excludedWorkspaceLabels}), so a worker cannot
* become a lead by being placed — as a split, say — inside a matching tab.</li>
* <li>A worker cannot rename a tab: {@code tab.rename} is reachable only through
* {@link WorkspaceControl}, which no {@code bridge_*} tool exposes. The label is writable by
* the human at the terminal and by nobody the bridge is defending against.</li>
* <li>The label is a <em>name</em>, not a capability. What a pane may do is decided by
* {@code Authz} against the role {@code CallerResolver} returns; a tab that calls itself a
* lead still cannot act as one unless the daemon's own registry agrees.</li>
* </ol>
*
* <p><strong>CB-558 — bridged now writes lead labels too.</strong> This class used to be able to say
* that bridged never renames a lead tab, so the label was always the human's own writing and there
* was no round-trip from the daemon's rename back into its next decision.
* {@code dev.ltms.bridged.lead.LeadLauncher} ends that: an auto-launched lead is labelled by the
* daemon and found again by this scan. The trust direction above is unaffected — bridged writing a
* name for a lead it just started is not a pane promoting itself — but <em>staleness</em> becomes
* real: a label left behind by a session that has since died would read as a live lead forever.
* This scanner does not solve that (its job is naming, and a stale name costs nothing here); the
* launcher does, by requiring a running agent in the tab before it counts the lead as live. If you
* ever make a decision that <em>removes</em> something based on this map, add the same check.
* The remaining hazard is an <em>operator</em> one — a worker {@code tabLabel} template that
* happens to start with the same prefix would promote the whole fleet — and that is refused at
* startup by {@code BridgedConfig.validateLeadTabPrefixes} rather than documented here.
*
* <p><strong>Caching.</strong> {@link #get()} is on the request path (every resolve), so the scan
* is TTL-cached and a stale-but-valid map is preferred to a herdr round-trip. A failed scan keeps
* the previous answer instead of emptying it — a herdr hiccup must not silently demote a live lead
* mid-session.
*/
public final class LeadTabScanner implements Supplier<Map<String, String>> {
private static final Logger log = LoggerFactory.getLogger(LeadTabScanner.class);
private final HerdrClient herdr;
private final Map<String, String> tabToName;
private final Set<String> excludedWorkspaceLabels;
private final long ttlNanos;
private final LongSupplier clock;
private Map<String, String> cached = Map.of();
private long scannedAtNanos;
private boolean everScanned;
/**
* @param herdr the herdr client to query ({@code workspace.list},
* {@code tab.list}, {@code pane.list} — all read-only)
* @param tabToName every configured lead's exact tab label → its name
* ({@code fleet.leaders.<name>.tab}), matched case-insensitively
* @param excludedWorkspaceLabels workspaces never scanned — the configured worker spaces
* @param ttlNanos how long a scan result is reused before the next one
* @param clock nanosecond time source ({@code System::nanoTime} in production)
*/
public LeadTabScanner(HerdrClient herdr, Map<String, String> tabToName,
Set<String> excludedWorkspaceLabels, long ttlNanos, LongSupplier clock) {
this.herdr = herdr;
this.tabToName = normalize(tabToName);
this.excludedWorkspaceLabels = excludedWorkspaceLabels == null
? Set.of() : Set.copyOf(excludedWorkspaceLabels);
this.ttlNanos = ttlNanos;
this.clock = clock;
}
/** Keys stripped and lower-cased once, so every lookup is a plain map hit. */
private static Map<String, String> normalize(Map<String, String> tabToName) {
if (tabToName == null || tabToName.isEmpty()) {
return Map.of();
}
Map<String, String> out = new LinkedHashMap<>();
tabToName.forEach((tab, name) -> {
if (tab != null && !tab.isBlank() && name != null && !name.isBlank()) {
out.put(tab.strip().toLowerCase(Locale.ROOT), name);
}
});
return Collections.unmodifiableMap(out);
}
/**
* The current {@code terminal_id → lead name} map, rescanning when the cache has expired.
*
* <p>Synchronized so a burst of concurrent calls produces one scan rather than one each; a scan
* is a handful of RPCs over a Unix socket and is rate-limited to one per TTL.
*/
@Override
public synchronized Map<String, String> get() {
long now = clock.getAsLong();
if (everScanned && now - scannedAtNanos < ttlNanos) {
return cached;
}
// Stamp before scanning, not after: a herdr that is down must cost one attempt per TTL, not
// one per request.
scannedAtNanos = now;
everScanned = true;
try {
Map<String, String> fresh = scan();
if (!fresh.equals(cached)) {
log.info("lead panes: {}", fresh);
}
cached = fresh;
} catch (HerdrException e) {
log.warn("lead-tab scan failed, keeping the {} lead(s) already known: {}",
cached.size(), e.getMessage());
}
return cached;
}
/** One full pass: labelled tabs → their panes → those panes' terminals. */
private Map<String, String> scan() {
Map<String, String> nameByTab = new LinkedHashMap<>();
for (JsonNode w : herdr.call("workspace.list").path("workspaces")) {
Workspace ws = Workspace.from(w);
if (ws.workspaceId() == null || excludedWorkspaceLabels.contains(ws.label())) {
continue;
}
for (JsonNode t : herdr.call("tab.list", Map.of("workspace_id", ws.workspaceId())).path("tabs")) {
Tab tab = Tab.from(t);
String name = leadNameOf(tab.label());
if (name != null && tab.tabId() != null) {
nameByTab.put(tab.tabId(), name);
}
}
}
Map<String, String> byTerminal = new LinkedHashMap<>();
if (!nameByTab.isEmpty()) {
// One pane.list for every tab: panes carry tab_id, so the join is local.
for (JsonNode p : herdr.call("pane.list", Map.of()).path("panes")) {
String name = nameByTab.get(p.path("tab_id").asText(null));
String terminal = p.path("terminal_id").asText(null);
if (name != null && terminal != null && !terminal.isBlank()) {
byTerminal.put(terminal, name);
}
}
}
return Collections.unmodifiableMap(byTerminal);
}
/**
* The lead name a tab label declares, or {@code null} if it names none of the configured leads.
*
* <p>Exact match (case-insensitive, ends stripped) against {@link #tabToName} — no prefix
* stripping, so an operator's {@code "lead: something-else"} tab is never mistaken for a
* configured lead just because it shares a prefix.
*/
private String leadNameOf(String label) {
if (label == null) {
return null;
}
return tabToName.get(label.strip().toLowerCase(Locale.ROOT));
}
}
@@ -23,9 +23,9 @@ public record Tab(String tabId, String workspaceId, String label, int paneCount)
}
/**
* A freshly-created tab together with the placeholder shell pane herdr seeds it with.
* The caller starts the worker into {@link #tab()} then closes {@link #rootPaneId()} so
* only the worker pane remains.
* A freshly-created tab together with the shell pane herdr seeds it with. Under protocol 19
* the caller starts the worker <em>into</em> {@link #rootPaneId()} — the seed pane's shell
* carries the worker's cwd and env from {@code tab.create}, and becomes the worker pane.
*/
public record Created(Tab tab, String rootPaneId) {
/** Project a {@code tab_created} result ({@code {tab, root_pane}}). */
@@ -5,6 +5,7 @@ import org.slf4j.Logger;
import org.slf4j.LoggerFactory;
import java.util.ArrayList;
import java.util.LinkedHashMap;
import java.util.List;
import java.util.Map;
import java.util.Optional;
@@ -40,6 +41,16 @@ public final class WorkspaceControl {
return out;
}
/** Every tab in {@code workspaceId}, in herdr's order. */
public List<Tab> listTabs(String workspaceId) {
JsonNode result = herdr.call("tab.list", Map.of("workspace_id", workspaceId));
List<Tab> out = new ArrayList<>();
for (JsonNode t : result.path("tabs")) {
out.add(Tab.from(t));
}
return out;
}
/** The first workspace with this exact label, if any. */
public Optional<Workspace> findByLabel(String label) {
return listWorkspaces().stream()
@@ -65,13 +76,37 @@ public final class WorkspaceControl {
}
/**
* A brand-new tab in {@code workspaceId} plus the placeholder shell pane herdr seeds it with.
* Start the worker into the tab, then {@code pane.close} the root pane so the tab holds only the
* worker. (The worker's own cwd is set on {@code agent.start}, not here — an {@code agent.start}
* pane does not inherit the tab's cwd; see {@code AgentControl.start}.)
* A brand-new tab in {@code workspaceId} plus the shell pane herdr seeds it with. Under
* protocol 19 that seed pane is where the worker <em>starts</em>: its shell carries
* {@code cwd} and {@code env} (the worker's {@code ANTHROPIC_BASE_URL} — this is the
* subscription-injection seam now), and {@code agent.start} launches the agent into it.
*/
public Tab.Created createTab(String workspaceId) {
return Tab.Created.from(herdr.call("tab.create", Map.of("workspace_id", workspaceId)));
public Tab.Created createTab(String workspaceId, String cwd, Map<String, String> env) {
Map<String, Object> params = new LinkedHashMap<>();
params.put("workspace_id", workspaceId);
if (cwd != null && !cwd.isBlank()) {
params.put("cwd", cwd);
}
if (env != null && !env.isEmpty()) {
params.put("env", env);
}
return Tab.Created.from(herdr.call("tab.create", params));
}
/**
* Split the currently-focused tab and return the new pane's id — the legacy pane placement's
* seed pane, carrying {@code cwd} and {@code env} exactly as {@link #createTab}'s does.
*/
public String splitPane(String cwd, Map<String, String> env) {
Map<String, Object> params = new LinkedHashMap<>();
params.put("direction", "right");
if (cwd != null && !cwd.isBlank()) {
params.put("cwd", cwd);
}
if (env != null && !env.isEmpty()) {
params.put("env", env);
}
return herdr.call("pane.split", params).path("pane").path("pane_id").asText(null);
}
/** Give a worker's tab a human label in the tab bar. */
@@ -2,11 +2,17 @@ package dev.ltms.bridged.inject;
import dev.ltms.bridged.herdr.AgentControl;
import dev.ltms.bridged.msg.Rendezvous;
import dev.ltms.bridged.msg.TurnToken;
import org.slf4j.Logger;
import org.slf4j.LoggerFactory;
import java.util.List;
import java.util.Objects;
import java.util.Set;
import java.util.TreeSet;
import java.util.concurrent.CompletableFuture;
import java.util.concurrent.ConcurrentHashMap;
import java.util.regex.Pattern;
/**
* The CB-106 completion fallback: bridges the {@link Injector}'s turn-completion signal to the
@@ -55,8 +61,13 @@ public final class CompletionResolver implements TurnListener {
/** Cap the scraped tail so a long transcript can't return an unbounded blob. */
static final int MAX_SCRAPE_CHARS = 4000;
private static final String CLIPPED_PANE_TAIL_MARKER =
"[Pane tail clipped: member did not call bridge_reply.]";
private final AgentControl agents;
private final Rendezvous rendezvous;
private final ExhaustedPatternLookup exhaustedPatterns;
private final ExhaustionSink exhaustionSink;
/**
* Per-target record of the turn currently in flight: the exact {@link Rendezvous} waiter its
@@ -74,23 +85,36 @@ public final class CompletionResolver implements TurnListener {
private final ConcurrentHashMap<String, InFlight> inFlight = new ConcurrentHashMap<>();
public CompletionResolver(AgentControl agents, Rendezvous rendezvous) {
/**
* @param exhaustedPatterns CB-578 stage A: per-target lookup for a profile's configured
* usage-limit refusal pattern. Required — there is deliberately no
* defaulting overload; a caller that does not want the classification
* must pass an explicit inert value ({@link ExhaustedPatternLookup#none()}).
* @param exhaustionSink CB-578 stage B: notified when a {@code BACKEND_EXHAUSTED}
* classification actually resolves a waiter. Required for the same
* reason as {@code exhaustedPatterns} — pass {@link ExhaustionSink#none()}
* to opt out.
*/
public CompletionResolver(AgentControl agents, Rendezvous rendezvous, ExhaustedPatternLookup exhaustedPatterns,
ExhaustionSink exhaustionSink) {
this.agents = agents;
this.rendezvous = rendezvous;
this.exhaustedPatterns = Objects.requireNonNull(exhaustedPatterns, "exhaustedPatterns");
this.exhaustionSink = Objects.requireNonNull(exhaustionSink, "exhaustionSink");
}
@Override
public void onDelivered(String target) {
public void onDelivered(String target, TurnToken token) {
// Capture the exact waiter this turn belongs to (CB-116) and snapshot the pane's pre-turn
// content — what it shows *before* the just-delivered turn produces output — as the staleness
// reference (CB-115). Done synchronously (like the delivering send itself) so both are in
// place before this turn's completion can fire.
captureBaseline(target);
captureBaseline(target, token);
}
/** Capture the in-flight turn: its waiter and pre-turn baseline (the testable core of {@link #onDelivered}). */
void captureBaseline(String target) {
CompletableFuture<Rendezvous.Resolution> waiter = rendezvous.currentWaiter(target);
void captureBaseline(String target, TurnToken token) {
CompletableFuture<Rendezvous.Resolution> waiter = token.waiter();
if (waiter == null) {
inFlight.remove(target); // no send is waiting on this delivery — nothing to resolve later
return;
@@ -122,10 +146,25 @@ public final class CompletionResolver implements TurnListener {
Thread.ofVirtual().name("completion-" + target).start(() -> resolve(target, turn));
}
/**
* Resolve the completed turn before adapter housekeeping can erase its rendered output. This is
* intentionally synchronous and used only when a post-turn context reset is enabled; the normal
* path remains off-loaded so polling is not blocked by a scrape.
*/
public void resolveBeforePostAction(String target) {
resolve(target, inFlight.get(target));
}
@Override
public void onTurnFailed(String target) {
InFlight turn = inFlight.get(target);
Thread.ofVirtual().name("turn-failed-" + target).start(() -> fail(target, turn));
Thread.ofVirtual().name("turn-failed-" + target).start(() -> fail(target, turn, null));
}
@Override
public void onTurnFailed(String target, String reason) {
InFlight turn = inFlight.get(target);
Thread.ofVirtual().name("turn-failed-" + target).start(() -> fail(target, turn, reason));
}
/** Synchronous resolve (the unit-testable core of {@link #onTurnComplete}). */
@@ -138,9 +177,15 @@ public final class CompletionResolver implements TurnListener {
return;
}
String tail;
String assistantBlock = null;
int originalLength = 0;
boolean clipped = false;
boolean scrapeFailed = false;
try {
tail = clip(lastAssistantBlock(agents.read(target, SCRAPE_SOURCE)));
assistantBlock = lastAssistantBlock(agents.read(target, SCRAPE_SOURCE));
originalLength = assistantBlock.strip().length();
clipped = originalLength > MAX_SCRAPE_CHARS;
tail = clip(assistantBlock);
} catch (RuntimeException e) {
// The worker finished but we couldn't read its screen — still resolve the send so the
// caller unblocks; an empty tail beats hanging until the caller's timeout.
@@ -160,8 +205,33 @@ public final class CompletionResolver implements TurnListener {
target);
return; // keep the in-flight record: a later genuine completion still needs it
}
if (rendezvous.resolveCompletion(waiter, tail)) {
// CB-578 stage A: a turn that ended with no bridge_reply AND whose scrape matches the
// backend's configured usage-limit pattern is a refusal, not an answer. Classify it as
// BACKEND_EXHAUSTED rather than handing the caller a scrape that reads like a real reply.
if (!scrapeFailed) {
Pattern exhausted = exhaustedPatterns.patternFor(target);
String matchedLine = exhausted == null ? null : firstMatchingLine(assistantBlock, exhausted);
if (matchedLine != null) {
String reason = "backend exhausted (usage limit): " + matchedLine;
if (rendezvous.resolveExhausted(waiter, reason)) {
inFlight.remove(target, turn);
log.warn("completion for {} classified BACKEND_EXHAUSTED (no bridge_reply; scrape "
+ "matched the profile's exhausted pattern): {}", target, reason);
// CB-578 stage B: only on the resolution that actually won the race — a late
// duplicate must never quarantine a credential twice for one refusal.
exhaustionSink.onExhausted(target, reason);
}
return;
}
}
String completion = clipped ? tail + "\n" + CLIPPED_PANE_TAIL_MARKER : tail;
if (rendezvous.resolveCompletion(waiter, completion)) {
inFlight.remove(target, turn);
if (clipped) {
log.warn("completion scrape for {} clipped from {} chars to the {} char cap; "
+ "member did not call bridge_reply, so the pane tail is partial",
target, originalLength, MAX_SCRAPE_CHARS);
}
log.debug("resolved send to {} via turn-completion fallback ({} chars scraped)",
target, tail.length());
}
@@ -169,6 +239,11 @@ public final class CompletionResolver implements TurnListener {
/** Synchronous fail (the unit-testable core of {@link #onTurnFailed}). */
void fail(String target, InFlight turn) {
fail(target, turn, null);
}
/** Synchronous fail with an optional reason supplied by a dropped worker queue. */
void fail(String target, InFlight turn, String explicitReason) {
// A never-delivered readiness failure has no in-flight record but still has a blocked send;
// fall back to the currently-registered waiter (unambiguous — that send never completed, so
// no next turn exists to confuse it with).
@@ -178,23 +253,64 @@ public final class CompletionResolver implements TurnListener {
inFlight.remove(target, turn); // nobody blocked on this worker — nothing to fail
return;
}
String reason;
try {
reason = clip(agents.read(target, SCRAPE_SOURCE));
} catch (RuntimeException e) {
reason = "";
}
if (reason.isBlank()) {
// No screen to scrape — either the worker is stuck (CB-109) or gone (CB-110).
reason = "worker did not reply; its turn ended in an unrecoverable state "
+ "(worker unreachable or stuck)";
String reason = explicitReason;
if (reason == null || reason.isBlank()) {
try {
reason = clip(agents.read(target, SCRAPE_SOURCE));
} catch (RuntimeException e) {
reason = "";
}
if (reason.isBlank()) {
// No screen to scrape — either the worker is stuck (CB-109) or gone (CB-110).
reason = "worker did not reply; its turn ended in an unrecoverable state "
+ "(worker unreachable or stuck)";
}
}
if (rendezvous.resolveFailure(waiter, reason)) {
inFlight.remove(target, turn);
log.debug("failed send to {} via turn-stall fallback", target);
log.warn("failing send to {} via turn-stall fallback: {}", target, reason);
}
}
/**
* The first line of {@code text} matching {@code pattern}, stripped — the CB-578 stage A
* evidence carried in a {@code BACKEND_EXHAUSTED} reason so the operator sees the real refusal
* text, never a generic label. {@code null} if no line matches.
*/
static String firstMatchingLine(String text, Pattern pattern) {
if (text == null || text.isEmpty()) return null;
for (String line : text.split("\n", -1)) {
if (pattern.matcher(line).find()) {
return line.strip();
}
}
return null;
}
/**
* Coverage summary for the CB-578 stage A exhausted-pattern classification, logged at startup
* the way {@link dev.ltms.bridged.health.FleetHealthMonitor#coverage} is — so an operator can
* see whether the classification is on, and for which profiles, without reading every
* profile's config by hand.
*
* @param allProfiles every configured profile name
* @param configuredProfiles the subset of {@code allProfiles} that carry an exhausted pattern
*/
public static String coverage(Set<String> allProfiles, Set<String> configuredProfiles) {
if (configuredProfiles.isEmpty()) {
return "off (no profile has an exhaustedPattern configured; profiles: " + sorted(allProfiles) + ")";
}
Set<String> unconfigured = new TreeSet<>(allProfiles);
unconfigured.removeAll(configuredProfiles);
return unconfigured.isEmpty()
? "full (all profiles configured: " + sorted(allProfiles) + ")"
: "partial (configured: " + sorted(configuredProfiles) + "; not configured: " + sorted(unconfigured) + ")";
}
private static List<String> sorted(Set<String> names) {
return names.stream().sorted().toList();
}
private static String clip(String s) {
if (s == null) return "";
String trimmed = s.strip();
@@ -0,0 +1,27 @@
package dev.ltms.bridged.inject;
import java.util.regex.Pattern;
/**
* Per-target lookup for a profile's configured usage-limit refusal pattern (CB-578 stage A): how
* {@link CompletionResolver} tells a backend that refused on a subscription usage limit — the
* worker's pane stays healthy, but the account is exhausted — apart from a genuine completion.
*
* <p>The pattern is always profile config, never a vendor string in Java source: every backend
* words its refusal differently, so a hardcoded sentence would only ever match one of them.
*/
@FunctionalInterface
public interface ExhaustedPatternLookup {
/** The compiled pattern configured for {@code target}'s profile, or {@code null} if none. */
Pattern patternFor(String target);
/**
* Inert lookup — no profile has a pattern configured, so the classification never fires and
* the completion fallback behaves exactly as before CB-578 stage A. The explicit stand-in a
* caller (or a test not exercising this feature) passes instead of a defaulting overload.
*/
static ExhaustedPatternLookup none() {
return target -> null;
}
}
@@ -0,0 +1,31 @@
package dev.ltms.bridged.inject;
/**
* Notified when {@link CompletionResolver} actually delivers a {@code BACKEND_EXHAUSTED}
* classification to a waiting send (CB-578 stage B) — never on a race that lost (see
* {@link CompletionResolver#resolve}, which only calls this after
* {@code Rendezvous.resolveExhausted} returns {@code true}).
*
* <p>{@link CompletionResolver} knows only {@code target} (a herdr terminal id); it has no notion of
* profiles or credentials, so mapping {@code target} to whatever should be quarantined is entirely
* the sink's job — see {@code Bridged.main}'s wiring, which resolves target → session → profile →
* {@code effectiveCredentialId()} and calls {@code BackendQuarantine.quarantine} on it.
*/
@FunctionalInterface
public interface ExhaustionSink {
/**
* @param target the herdr terminal id whose turn was classified {@code BACKEND_EXHAUSTED}
* @param reason the matched-line reason carried by the classification
*/
void onExhausted(String target, String reason);
/**
* Inert sink — nothing happens on exhaustion. The explicit stand-in a caller (or a test not
* exercising this feature) passes instead of a defaulting overload, exactly like
* {@link ExhaustedPatternLookup#none()}.
*/
static ExhaustionSink none() {
return (target, reason) -> { };
}
}
@@ -2,6 +2,7 @@ package dev.ltms.bridged.inject;
import dev.ltms.bridged.herdr.AgentControl;
import dev.ltms.bridged.herdr.AgentStatus;
import dev.ltms.bridged.msg.TurnToken;
import org.slf4j.Logger;
import org.slf4j.LoggerFactory;
@@ -78,6 +79,15 @@ public final class Injector {
*/
private static final int READINESS_GRACE_POLLS = 240;
/**
* The single source for the injector poll cadence — how often the {@link StatusPoller} drives
* {@link #onStatus} at. {@code Bridged} passes this to every {@link StatusPoller} it constructs,
* and this class reads it to state the readiness grace in seconds on the CB-562 expiry log
* instead of hardcoding "60s". One constant, so a cadence change cannot silently desync a log
* that claims a grace duration.
*/
public static final long POLL_INTERVAL_MILLIS = 250;
private final AgentControl agents;
private final TurnListener turnListener;
private final Predicate<String> ready; // CB-113: a target is deliverable only when available
@@ -120,7 +130,7 @@ public final class Injector {
}
/** A pending message and the future that completes when it has been delivered. */
private record Pending(String text, CompletableFuture<Void> delivered) {
private record Pending(String text, TurnToken token, CompletableFuture<Void> delivered) {
}
/** Per-worker delivery state, guarded by its own monitor (single writer per worker). */
@@ -132,6 +142,10 @@ public final class Injector {
boolean turnObserved; // saw a real `working` sample since that delivery (turn ran)
int unknownSinceTurn; // consecutive `unknown` samples while a delegation is outstanding (CB-109)
int notReadySincePoll; // consecutive injectable samples a queued message waited on the readiness gate (CB-114)
boolean postTurnPending; // completion observed; adapter housekeeping has not started yet
boolean awaitingPostTurnPickup;
boolean postTurnObserved;
int injectableSincePostTurnPickup;
synchronized void add(Pending p) {
queue.add(p);
@@ -146,9 +160,9 @@ public final class Injector {
* <p>Uses an atomic map update so a concurrent {@link #drop} cannot slip between "find the
* target" and "queue the message" and orphan it in a target it just removed.
*/
public CompletableFuture<Void> enqueue(String target, String text) {
public CompletableFuture<Void> enqueue(String target, String text, TurnToken token) {
CompletableFuture<Void> delivered = new CompletableFuture<>();
Pending p = new Pending(text, delivered);
Pending p = new Pending(text, token, delivered);
targets.compute(target, (_, existing) -> {
Target t = (existing != null) ? existing : new Target();
t.add(p); // synchronized on the Target monitor — atomic with a concurrent drop
@@ -173,9 +187,15 @@ public final class Injector {
boolean turnCompleted = false;
boolean turnFailed = false;
boolean resubmit = false;
boolean startPostTurn = false;
List<Pending> notReady = null; // queued messages failed because the worker never became ready
synchronized (t) {
if (status == AgentStatus.WORKING) {
if (t.awaitingPostTurnPickup) {
t.awaitingPostTurnPickup = false;
t.injectableSincePostTurnPickup = 0;
t.postTurnObserved = true;
}
// Definitive pickup: the worker is busy on our last message, and (if a delivery is
// outstanding) a real turn is now confirmed to be running.
t.awaitingPickup = false;
@@ -185,6 +205,16 @@ public final class Injector {
if (t.awaitingCompletion) t.turnObserved = true;
} else if (status.injectable()) { // IDLE or BLOCKED
t.unknownSinceTurn = 0;
if (t.awaitingPostTurnPickup) {
if (++t.injectableSincePostTurnPickup >= PICKUP_GRACE_POLLS) {
t.awaitingPostTurnPickup = false;
t.injectableSincePostTurnPickup = 0;
} else {
resubmit = true;
}
} else if (t.postTurnObserved) {
t.postTurnObserved = false;
}
if (t.awaitingPickup) {
if (++t.injectableSincePickup >= PICKUP_GRACE_POLLS) {
// Pickup edge was never sampled (turn faster than the poll, or status lag).
@@ -209,11 +239,16 @@ public final class Injector {
t.awaitingCompletion = false;
t.turnObserved = false;
turnCompleted = true;
if (turnListener.hasPostTurnAction(target)) {
t.postTurnPending = true;
startPostTurn = true;
}
}
// Deliver the next queued message only once the prior turn is fully settled, so a
// completion is never confused with the pickup of the following message — and only
// once the worker is available (CB-113), so we never paste into its boot window.
if (!t.awaitingCompletion) {
if (!t.awaitingCompletion && !t.postTurnPending
&& !t.awaitingPostTurnPickup && !t.postTurnObserved) {
Pending p = t.queue.peek();
if (p != null && ready.test(target)) {
t.notReadySincePoll = 0;
@@ -239,6 +274,11 @@ public final class Injector {
// fail every queued message and release the target (CB-114) instead of
// polling it indefinitely with the caller's future never completing.
notReady = new ArrayList<>(t.queue);
log.warn("readiness grace for {} expired after {} polls ({}s): target never "
+ "became deliverable, so failing {} queued message(s) that never "
+ "reached its pane",
target, READINESS_GRACE_POLLS,
READINESS_GRACE_POLLS * POLL_INTERVAL_MILLIS / 1000, notReady.size());
t.queue.clear();
t.notReadySincePoll = 0;
}
@@ -263,7 +303,8 @@ public final class Injector {
// Reclaim the entry once the worker is fully quiescent (nothing queued, no pickup or
// completion awaited), so the map cannot grow without bound across short-lived workers.
if (t.queue.isEmpty() && !t.awaitingPickup && !t.awaitingCompletion) {
if (t.queue.isEmpty() && !t.awaitingPickup && !t.awaitingCompletion
&& !t.postTurnPending && !t.awaitingPostTurnPickup && !t.postTurnObserved) {
targets.remove(target, t);
}
}
@@ -290,7 +331,21 @@ public final class Injector {
turnListener.onTurnFailed(target);
}
if (turnCompleted) {
turnListener.onTurnComplete(target);
if (startPostTurn) {
boolean started = turnListener.onTurnCompleteWithPostAction(target);
synchronized (t) {
t.postTurnPending = false;
if (started) {
t.awaitingPostTurnPickup = true;
t.injectableSincePostTurnPickup = 0;
}
if (t.queue.isEmpty() && !t.awaitingPostTurnPickup) {
targets.remove(target, t);
}
}
} else {
turnListener.onTurnComplete(target);
}
}
if (turnFailed) {
turnListener.onTurnFailed(target);
@@ -302,7 +357,7 @@ public final class Injector {
} else {
// Baseline the pane's pre-turn content so a misattributed completion (no new output)
// can't resolve this send with the previous turn's stale answer (CB-115).
turnListener.onDelivered(target);
turnListener.onDelivered(target, sent.token());
sent.delivered().complete(null);
}
}
@@ -317,7 +372,8 @@ public final class Injector {
.filter(e -> {
synchronized (e.getValue()) {
Target t = e.getValue();
return !t.queue.isEmpty() || t.awaitingPickup || t.awaitingCompletion;
return !t.queue.isEmpty() || t.awaitingPickup || t.awaitingCompletion
|| t.postTurnPending || t.awaitingPostTurnPickup || t.postTurnObserved;
}
})
.map(java.util.Map.Entry::getKey)
@@ -343,13 +399,20 @@ public final class Injector {
hadDeliveredTurn = t.awaitingCompletion;
t.awaitingCompletion = false;
t.awaitingPickup = false;
t.postTurnPending = false;
t.awaitingPostTurnPickup = false;
t.postTurnObserved = false;
}
log.warn("{} is gone, dropping its queue: {} message(s) failed{}; cause: {}", target,
pending.size(),
hadDeliveredTurn ? " (including one turn already in flight whose completion was never confirmed)" : "",
cause.getMessage());
forget.accept(target); // the worker is gone — clear its readiness/presence too (CB-114)
for (Pending p : pending) {
p.delivered().completeExceptionally(cause);
}
if (hadDeliveredTurn) {
turnListener.onTurnFailed(target);
}
// A queued send has no in-flight record, while a delivered turn does. CompletionResolver
// handles both forms and resolves its waiter at most once.
turnListener.onTurnFailed(target, cause.getMessage());
}
}
@@ -15,7 +15,7 @@ import java.util.Set;
* mounts the bridge MCP is never marked present — its sends stay queued until they time out, which is
* correct (it could not have replied anyway).
*/
public class WorkerPresence {
public class MemberPresence {
private final Set<String> present = ConcurrentHashMap.newKeySet();
@@ -1,5 +1,7 @@
package dev.ltms.bridged.inject;
import dev.ltms.bridged.msg.TurnToken;
/**
* Notified when a worker's delegated turn is observed to complete — a confirmed
* {@code WORKING → IDLE} transition after a delivery. This is the CB-106 completion signal the
@@ -13,6 +15,25 @@ public interface TurnListener {
/** A worker's delegated turn finished (worker returned to idle after visibly working). */
void onTurnComplete(String target);
/**
* Whether completion must pause ordinary delivery while adapter-specific housekeeping starts.
* This is queried before the injector considers the next queued message, closing the same-tick
* ordering gap. Implementations must not mutate state here.
*/
default boolean hasPostTurnAction(String target) {
return false;
}
/**
* Complete the delegated turn and start its non-turn housekeeping operation.
*
* @return true when the injector must observe that operation settle before the next delivery
*/
default boolean onTurnCompleteWithPostAction(String target) {
onTurnComplete(target);
return false;
}
/**
* A worker that visibly ran a delegated turn then wedged in a non-idle, non-working state
* (CB-109) — e.g. an error screen herdr classifies as {@code unknown} — so no
@@ -22,6 +43,14 @@ public interface TurnListener {
default void onTurnFailed(String target) {
}
/**
* As {@link #onTurnFailed(String)}, carrying the reason a worker became unreachable. The default
* keeps existing listeners working while allowing the completion resolver to report a useful cause.
*/
default void onTurnFailed(String target, String reason) {
onTurnFailed(target);
}
/**
* A message was just delivered into {@code target}'s pane (CB-115). Fired so the completion
* resolver can snapshot the pane's pre-turn content: a later {@link #onTurnComplete} whose
@@ -30,7 +59,7 @@ public interface TurnListener {
* resolve the send with the previous turn's stale answer. A default no-op keeps the interface
* functional for callers that don't scrape.
*/
default void onDelivered(String target) {
default void onDelivered(String target, TurnToken token) {
}
/** No-op default for callers that only need delivery, not completion signalling. */
@@ -0,0 +1,292 @@
package dev.ltms.bridged.lead;
import dev.ltms.bridged.config.BridgedConfig;
import dev.ltms.bridged.herdr.Agent;
import dev.ltms.bridged.herdr.AgentControl;
import dev.ltms.bridged.herdr.HerdrException;
import dev.ltms.bridged.herdr.Tab;
import dev.ltms.bridged.herdr.Workspace;
import dev.ltms.bridged.herdr.WorkspaceControl;
import org.slf4j.Logger;
import org.slf4j.LoggerFactory;
import java.util.ArrayList;
import java.util.LinkedHashMap;
import java.util.List;
import java.util.Map;
import java.util.Objects;
import java.util.Set;
import java.util.stream.Collectors;
/**
* Starts the leads {@code fleet.leaders:} declares, when none is already running (CB-558).
*
* <p><strong>A lead is not a member, and this class exists to keep it that way.</strong> Every other
* spawn path in the daemon goes through {@code HerdrPeerLauncher}, which does three things a lead
* must never receive:
* <ol>
* <li>it appends the <em>reply charter</em> — "you are an off-subscription worker … end every turn
* with {@code bridge_reply}". A lead is the orchestrator; telling it that it is a worker is
* exactly backwards.</li>
* <li>it registers the session with {@code SessionManager}, which subjects it to the idle reaper,
* the context cap and the shutdown drain. An idle lead is the normal state of a lead, so the
* reaper would kill the orchestrator for doing its job.</li>
* <li>it can move a peer off the subscription via {@code ANTHROPIC_BASE_URL}. A lead stays on the
* operator's subscription, always.</li>
* </ol>
* So this launcher talks to {@link AgentControl}/{@link WorkspaceControl} directly. The duplication
* with the member launchers (argv and env assembly) is deliberate and is the cheaper half of the
* trade: entangling the worker path with a not-a-worker case is how the three rules above get
* broken later, quietly.
*
* <p><strong>Liveness, and the label round-trip.</strong> {@code LeadTabScanner} used to be able to
* promise that bridged never writes a lead label. That is no longer true — an auto-launched lead is
* labelled by this class, and the scanner reads that label back. The risk this opens is not
* privilege escalation (the tab label never granted anything a pane could take for itself; see that
* class's javadoc), but <em>staleness</em>: a label left behind by a crashed session would otherwise
* read as a live lead forever, and the lead would never be relaunched. So a lead counts as live only
* when herdr also reports a <em>running agent</em> in that tab — see {@link #liveLeads}. A labelled
* tab with no agent in it is not a lead.
*/
public final class LeadLauncher {
private static final Logger log = LoggerFactory.getLogger(LeadLauncher.class);
private final AgentControl agents;
private final WorkspaceControl spaces;
private final BridgedConfig cfg;
/**
* @param agents herdr agent control (start, list)
* @param spaces workspace / tab control (ensure, create, label, list)
* @param cfg the loaded config — {@code fleet.leaders}, {@code profiles} and each lead's tab
*/
public LeadLauncher(AgentControl agents, WorkspaceControl spaces, BridgedConfig cfg) {
this.agents = agents;
this.spaces = spaces;
this.cfg = cfg;
}
/**
* Bring every declared lead up to its {@code instances} count, and return how many were started.
*
* <p>Never throws: a daemon that cannot start a lead must still serve. herdr being unreachable,
* a profile that does not exist, a failed {@code agent.start} — each is logged and skipped, and
* the remaining leads are still attempted.
*/
public int ensureLeads() {
Map<String, BridgedConfig.Leader> leaders = cfg.fleet().leaders();
if (leaders.isEmpty()) {
return 0;
}
Map<String, Integer> live;
try {
live = liveLeads(leaders);
} catch (HerdrException e) {
// Counting is the whole safety mechanism against double-spawning. If we cannot count, we
// must not guess — spawning a second orchestrator is worse than starting none.
log.warn("lead auto-launch skipped — cannot tell which leads are live: {}", e.getMessage());
return 0;
}
int started = 0;
for (Map.Entry<String, BridgedConfig.Leader> e : leaders.entrySet()) {
String name = e.getKey();
BridgedConfig.Leader lead = e.getValue();
int running = live.getOrDefault(name, 0);
int wanted = lead.instances();
if (running >= wanted) {
log.info("lead '{}': {} live, {} wanted — nothing to start", name, running, wanted);
continue;
}
if (!lead.isCreatable()) {
// A lead with a `tab:` but no `profile:` is recognise-only by design: the operator
// opens it by hand. Say so once rather than looking like a silent failure.
log.info("lead '{}' is not live, and names no profile — it can be recognised but not "
+ "launched. Add `profile:` under fleet.leaders.{} to have bridged start it.",
name, name);
continue;
}
BridgedConfig.Profile profile = cfg.profiles().get(lead.profile());
if (profile == null) {
log.warn("lead '{}' names profile '{}', which is not configured — not launching",
name, lead.profile());
continue;
}
for (int i = running; i < wanted; i++) {
if (launch(name, lead, profile)) {
started++;
}
}
}
return started;
}
/**
* How many live leads exist per configured name: a running agent in a tab labelled with that
* lead's exact {@code tab} (CB-579). Member workspaces are excluded, exactly as the scanner
* excludes them: a member must not be counted as a lead because it happens to sit in a matching
* tab.
*
* <p>There used to be a second path here — a running agent on the terminal a
* {@code fleet.leaders.<name>.terminal} pin named, for a lead opened and pinned by hand. That
* pin is retired: {@code tab} is now the only field identity depends on, and {@link Agent}
* already carries {@link Agent#tabId()} directly, so a hand-opened lead is found the same way an
* auto-launched one is — by labelling its tab to match.
*/
private Map<String, Integer> liveLeads(Map<String, BridgedConfig.Leader> leaders) {
Set<String> memberSpaces = cfg.profiles().values().stream()
.map(BridgedConfig.Profile::workspace)
.filter(w -> w != null && !w.isBlank())
.collect(Collectors.toSet());
// tabId → the lead name its label declares.
Map<String, String> nameByTab = new LinkedHashMap<>();
for (Workspace ws : spaces.listWorkspaces()) {
if (ws.workspaceId() == null || memberSpaces.contains(ws.label())) {
continue;
}
for (Tab tab : spaces.listTabs(ws.workspaceId())) {
String declared = leadNameOf(tab.label(), leaders);
if (declared != null && tab.tabId() != null) {
nameByTab.put(tab.tabId(), declared);
}
}
}
Map<String, Integer> counts = new LinkedHashMap<>();
for (Agent a : agents.list()) {
String name = nameByTab.get(a.tabId());
if (name != null) {
counts.merge(name, 1, Integer::sum);
}
}
return counts;
}
/**
* The configured lead a tab label names, or {@code null} for a label that names none.
*
* <p>Matched exactly (case-insensitively) against each lead's configured {@code tab}, so an
* operator's {@code "lead: something-else"} tab is not mistaken for a configured lead.
*/
private String leadNameOf(String label, Map<String, BridgedConfig.Leader> leaders) {
if (label == null) {
return null;
}
String l = label.strip();
for (Map.Entry<String, BridgedConfig.Leader> e : leaders.entrySet()) {
String tab = e.getValue().tabLabel();
if (tab != null && l.equalsIgnoreCase(tab.strip())) {
return e.getKey();
}
}
return null;
}
/** Start one lead. Returns false (having logged) rather than throwing on any failure. */
private boolean launch(String name, BridgedConfig.Leader lead, BridgedConfig.Profile profile) {
String label = lead.tabLabel();
String cwd = (lead.cwd() == null || lead.cwd().isBlank())
? System.getProperty("user.dir") : lead.cwd();
Tab.Created tab = null;
try {
Workspace ws = spaces.ensureWorkspace(lead.workspace());
tab = spaces.createTab(ws.workspaceId(), cwd, leadEnv(profile));
if (tab.rootPaneId() == null) {
throw new IllegalStateException("tab " + tab.tab().tabId()
+ " came back with no seed pane — nowhere to start the lead");
}
// Same shape as the member launchers: herdr resolves the executable from `kind`, so
// argv[0] (the configured launcher, e.g. `ccs`) is dropped and only the rest is passed.
List<String> argv = leadArgv(profile);
Agent started = agents.start("lead-" + name, herdrKind(profile),
argv.isEmpty() ? argv : argv.subList(1, argv.size()), tab.rootPaneId());
// Label AFTER the start succeeds. A label written before would survive a failed start
// and then read back as a live lead on the next boot, which is the exact staleness the
// agent-liveness check exists to prevent — no need to create the case ourselves.
spaces.renameTab(tab.tab().tabId(), label);
log.info("lead '{}' launched: profile={} tab={} pane={} terminal={} label='{}' cwd={}",
name, profile.profile(), tab.tab().tabId(), started.paneId(),
started.terminalId(), label, cwd);
return true;
} catch (RuntimeException e) {
log.warn("lead '{}' failed to launch on profile '{}': {}",
name, profile.profile(), e.getMessage());
if (tab != null && tab.tab() != null && tab.tab().tabId() != null) {
try {
spaces.closeTab(tab.tab().tabId());
} catch (RuntimeException cleanup) {
log.warn("could not close the orphaned lead tab {}: {}",
tab.tab().tabId(), cleanup.getMessage());
}
}
return false;
}
}
/**
* The herdr agent kind for this profile — the same value the matching member adapter passes, so
* herdr resolves the same executable for a lead as it does for a member on that backend.
*/
private static String herdrKind(BridgedConfig.Profile profile) {
return profile.isOpenCode() ? "opencode" : "claude";
}
/**
* The lead's argv: the profile's own command, the model pin, and the bridge MCP mount.
*
* <p>No {@code --append-system-prompt}. That flag carries the worker reply charter, and a lead
* is not a worker — it reads its orchestration rules from the project's {@code CLAUDE.md} like
* any other primary. This is the single most important difference from the member launchers;
* do not "unify" it back.
*/
private List<String> leadArgv(BridgedConfig.Profile profile) {
List<String> argv = new ArrayList<>(profile.argv());
if (profile.hasMcp()) {
argv.add("--mcp-config");
argv.add("{\"mcpServers\":{\"bridge\":{\"type\":\"http\",\"url\":\""
+ profile.mcpUrl() + "\"}}}");
}
// Appended last, for the same reason the member launcher does it (CB-533): the argv is
// usually a wrapper such as `ccs <profile>`, which exports its own model family over
// whatever it inherited, and --model outranks the environment.
if (profile.model() != null && !profile.model().isBlank()) {
argv.add("--model");
argv.add(profile.model());
}
return argv;
}
/**
* The lead's environment: the profile's {@code env:} block, and nothing that could move it off
* the subscription.
*
* <p>{@code ANTHROPIC_BASE_URL} and {@code ANTHROPIC_AUTH_TOKEN} are stripped unconditionally —
* not defaulted, not guarded, stripped. A lead runs on the operator's subscription by
* definition, so there is no configuration under which pointing it elsewhere is correct, and a
* profile that carries them (a member profile reused as a lead's backend) must not leak them in.
*/
private Map<String, String> leadEnv(BridgedConfig.Profile profile) {
Map<String, String> out = new LinkedHashMap<>();
if (profile.env() != null) {
out.putAll(profile.env());
}
out.remove("ANTHROPIC_BASE_URL");
out.remove("ANTHROPIC_AUTH_TOKEN");
if (profile.configDir() != null && !profile.configDir().isBlank()) {
out.put("CLAUDE_CONFIG_DIR", profile.configDir());
}
// Deliberately no git token: a lead reviews and merges through the operator's own
// credentials, and never needs the scoped write:repository token a member is granted.
out.values().removeIf(Objects::isNull);
return out;
}
}
@@ -0,0 +1,30 @@
package dev.ltms.bridged.logging;
import ch.qos.logback.classic.Level;
import ch.qos.logback.classic.Logger;
import ch.qos.logback.classic.turbo.TurboFilter;
import ch.qos.logback.core.spi.FilterReply;
import io.modelcontextprotocol.spec.McpSchema;
import org.slf4j.Marker;
/** Suppresses only the SDK warning for the normal MCP cancellation notification. */
public final class McpCancelledNotificationFilter extends TurboFilter {
static final String LOGGER = "io.modelcontextprotocol.spec.McpStreamableServerSession";
static final String UNHANDLED_NOTIFICATION = "No handler registered for notification method: {}";
@Override
public FilterReply decide(Marker marker, Logger logger, Level level, String format, Object[] params,
Throwable throwable) {
if (level == Level.WARN
&& LOGGER.equals(logger.getName())
&& UNHANDLED_NOTIFICATION.equals(format)
&& params != null
&& params.length == 1
&& params[0] instanceof McpSchema.JSONRPCNotification notification
&& "notifications/cancelled".equals(notification.method())) {
return FilterReply.DENY;
}
return FilterReply.NEUTRAL;
}
}
@@ -9,14 +9,16 @@ import dev.ltms.bridged.guard.GuardException;
import dev.ltms.bridged.herdr.Agent;
import dev.ltms.bridged.metrics.BridgedMetrics;
import dev.ltms.bridged.metrics.Metrics;
import dev.ltms.bridged.inject.WorkerPresence;
import dev.ltms.bridged.inject.MemberPresence;
import dev.ltms.bridged.herdr.HerdrException;
import dev.ltms.bridged.msg.MessageService;
import dev.ltms.bridged.msg.Rendezvous;
import dev.ltms.bridged.peer.PeerUnreachableException;
import dev.ltms.bridged.placement.BackendQuarantine;
import dev.ltms.bridged.session.SessionManager;
import dev.ltms.bridged.session.WorkerSession;
import dev.ltms.bridged.session.MemberSession;
import dev.ltms.bridged.session.WorktreeRequest;
import dev.ltms.bridged.peer.MemberRole;
import dev.ltms.bridged.peer.PeerLauncher;
import io.modelcontextprotocol.common.McpTransportContext;
import io.modelcontextprotocol.json.McpJsonMapper;
@@ -32,7 +34,11 @@ import jakarta.servlet.http.HttpServlet;
import java.util.LinkedHashMap;
import java.util.List;
import java.util.Map;
import java.util.Objects;
import java.util.Set;
import java.util.function.Function;
import java.util.function.LongSupplier;
import java.util.function.Supplier;
import java.util.stream.Collectors;
/**
@@ -65,20 +71,35 @@ public final class BridgeMcp {
static final String CALLER_PID = "callerPid";
/** Transport-context key under which the extractor stashes the resolved {@link Role} (CB-501). */
static final String CALLER_ROLE = "callerRole";
/** Transport-context key for the lead's configured name, when the caller is one (CB-530). */
static final String CALLER_NAME = "callerName";
private final HttpServletStreamableServerTransportProvider transport;
private final McpSyncServer server;
private final CallerResolver authz; // CB-501: null → authorization not enforced (legacy)
private final Metrics metrics; // CB-502: null → auth failures not counted
private final CapacitySource capacity;
private final HealthCoverageSource healthCoverage;
private final QuarantineSource quarantine;
/** Capacity facts used by {@code bridge_list}; production must supply the placement live count. */
public record CapacitySource(Function<String, Integer> liveCount, Function<String, Integer> maxLoad,
Supplier<Set<String>> configuredProfiles, LongSupplier clock) {
/** Inert test-only source. It omits capacity rather than inventing zero live counts. */
public static CapacitySource none() { return new CapacitySource(_ -> 0, _ -> null, Set::of, System::nanoTime); }
boolean available() { return !configuredProfiles.get().isEmpty(); }
}
/** Coverage is supplied by the health wiring, not inferred from a missing dependency. */
public record HealthCoverageSource(Supplier<String> value) { }
/**
* Legacy constructor — no authorization. Retained so existing tests exercise tool behaviour
* without an auth fixture.
* CB-578 stage B quarantine facts used by {@code bridge_profiles}: a profile → credential id
* lookup, plus the shared {@link BackendQuarantine} to read remaining cooldowns off.
*/
public BridgeMcp(MessageService messages, PeerLauncher workers,
SessionManager sessions, ConnectionIdentity identity, WorkerPresence presence,
PrimaryRegistry primaryRegistry) {
this(messages, workers, sessions, identity, presence, primaryRegistry, null, null);
public record QuarantineSource(Function<String, String> credentialIdFor, BackendQuarantine quarantine) {
/** Inert source — no profile is ever reported quarantined. Explicit stand-in, not a default. */
public static QuarantineSource none() { return new QuarantineSource(_ -> null, BackendQuarantine.none()); }
}
/**
@@ -87,10 +108,16 @@ public final class BridgeMcp {
* Jetty's context handler and never passes through Javalin's {@code before}
* filter, so the REST guard does not cover it.
* @param metrics registry for auth-failure counting; may be {@code null}
* @param quarantine CB-578 stage B facts for {@code bridge_profiles}; required — pass
* {@link QuarantineSource#none()} for a caller that does not want the feature
*/
public BridgeMcp(MessageService messages, PeerLauncher workers,
SessionManager sessions, ConnectionIdentity identity, WorkerPresence presence,
PrimaryRegistry primaryRegistry, CallerResolver callers, Metrics metrics) {
public BridgeMcp(MessageService messages, PeerLauncher workers, SessionManager sessions,
ConnectionIdentity identity, MemberPresence presence, PrimaryRegistry primaryRegistry,
CallerResolver callers, Metrics metrics, CapacitySource capacity, HealthCoverageSource healthCoverage,
QuarantineSource quarantine) {
this.capacity = capacity;
this.quarantine = Objects.requireNonNull(quarantine, "quarantine");
this.healthCoverage = healthCoverage;
McpJsonMapper json = new JacksonMcpJsonMapperSupplier().get();
this.transport = HttpServletStreamableServerTransportProvider.builder()
.jsonMapper(json)
@@ -106,11 +133,15 @@ public final class BridgeMcp {
? callers.resolve(req.getRemoteAddr(), req.getRemotePort(),
req.getHeader("Authorization"))
: legacyPrincipal(identity, req.getRemoteAddr(), req.getRemotePort());
presence.markPresent(p.terminal()); // no-op for the primary (null terminal)
// CB-532: guard on the ROLE, not on the terminal being null. This excludes a
// lead, which carries its pane too, while including every spawned member role.
// Enrolling a lead would count it as an available member in the roster.
markSpawnedMemberPresent(p, presence);
return McpTransportContext.create(Map.of(
CALLER_TERMINAL, orEmpty(p.terminal()),
CALLER_PID, Long.toString(p.pid()),
CALLER_ROLE, p.role().name()));
CALLER_ROLE, p.role().name(),
CALLER_NAME, orEmpty(p.name())));
})
.build();
this.server = McpServer.sync(transport)
@@ -121,18 +152,31 @@ public final class BridgeMcp {
str(req.arguments(), "sessionId"));
if (denied != null) return denied;
String caller = callerTerminal(exchange);
if (caller != null) primaryRegistry.record(caller);
// CB-548: only a PRIMARY caller may claim the legacy singleton "primary" fallback.
// An architect delegates as its own pane but must never become the fallback that
// no-delegation inbox nudges target as if it were the primary (the per-target
// delegation map does not cure the singleton).
recordPrimarySingleton(primaryRegistry, caller, principal(exchange));
Map<String, Object> a = req.arguments();
String target = str(a, "sessionId");
String content = str(a, "content");
String turnId = str(a, "turnId");
if (turnId != null && !turnId.isBlank()) {
// Answering a worker's bridge_ask (CB-205): resolve its blocked question and
// block for the worker's reply as it resumes the same turn.
return answer(messages, turnId, str(a, "content"), timeoutMs(a));
// block for the worker's reply as it resumes the same turn. This is the same
// delegation, so ownership is left untouched (CB-548) — never re-recorded.
return answer(messages, turnId, content, timeoutMs(a));
}
// CB-548: delegator ownership (which lead's reply nudge this worker routes to,
// CB-532) is recorded only once the send is ACCEPTED — MessageService has won the
// session lock and queued delivery — via the accepted-delivery callback, never at
// request time. A concurrent sender that times out BUSY therefore cannot steal a
// live turn's reply routing without ever owning the turn.
Runnable onAccepted = () -> primaryRegistry.recordDelegation(target, caller);
// wait defaults to true (block for the reply); wait:false is fire-and-poll.
return Boolean.FALSE.equals(a.get("wait"))
? sendAsync(messages, str(a, "sessionId"), str(a, "content"))
: send(messages, str(a, "sessionId"), str(a, "content"), timeoutMs(a));
? sendAsync(messages, target, content, onAccepted, workers.profiles())
: send(messages, target, content, timeoutMs(a), onAccepted, workers.profiles());
})
// bridge_reply's identity is the CONNECTION, never an argument — so the authz check
// is "is this caller a worker at all", and it can only ever reply as itself.
@@ -173,19 +217,24 @@ public final class BridgeMcp {
McpSchema.CallToolResult denied = deny(exchange, Authz.Action.SPAWN, null);
if (denied != null) return denied;
String caller = callerTerminal(exchange);
if (caller != null) primaryRegistry.record(caller);
// SPAWN is already auth-gated to PRIMARY (architects can never call it), but
// enforce the same invariant here: only a PRIMARY may claim the legacy singleton.
recordPrimarySingleton(primaryRegistry, caller, principal(exchange));
Map<String, Object> a = req.arguments();
// CB-112: worker inherits the primary's cwd unless the call pins one.
// CB-301: carry the caller's identity as the session owner (null for the primary).
// CB-301-ext: optional isolated worktree for parallel implementers.
String callerCwd = identity.cwdForPid(callerPid(exchange));
return spawn(sessions, str(a, "profile"), str(a, "cwd"), callerCwd,
callerTerminal(exchange), worktreeRequest(a));
return spawn(sessions, str(a, "profile"), str(a, "role"), str(a, "cwd"), callerCwd,
callerTerminal(exchange), worktreeRequest(a),
str(a, "sessionName"), str(a, "resumeSessionId"));
})
.toolCall(listTool(), (exchange, _) -> {
McpSchema.CallToolResult denied = deny(exchange, Authz.Action.READ, null);
if (denied != null) return denied;
return listWorkers(workers, sessions);
return listFleet(workers, sessions, messages, capacity, healthCoverage, quarantine,
callers == null ? Map.of() : callers.leads(),
callerTerminal(exchange));
})
.toolCall(stopTool(), (exchange, req) -> {
String paneId = str(req.arguments(), "paneId");
@@ -196,7 +245,7 @@ public final class BridgeMcp {
.toolCall(profilesTool(), (exchange, _) -> {
McpSchema.CallToolResult denied = deny(exchange, Authz.Action.READ, null);
if (denied != null) return denied;
return profiles(workers);
return profiles(workers, quarantine);
})
.toolCall(whoamiTool(), (exchange, _) -> {
McpSchema.CallToolResult denied = deny(exchange, Authz.Action.READ, null);
@@ -222,7 +271,7 @@ public final class BridgeMcp {
/** The caller reconstructed from the transport context. */
private static Principal principal(McpSyncServerExchange exchange) {
return principalFrom(exchange.transportContext().get(CALLER_ROLE),
callerTerminal(exchange), callerPid(exchange));
callerTerminal(exchange), callerPid(exchange), callerName(exchange));
}
/**
@@ -237,11 +286,16 @@ public final class BridgeMcp {
* @param pid the calling pid, or {@code -1}
*/
static Principal principalFrom(Object role, String terminal, long pid) {
return principalFrom(role, terminal, pid, null);
}
/** As {@link #principalFrom(Object, String, long)}, carrying a lead's name (CB-530). */
static Principal principalFrom(Object role, String terminal, long pid, String name) {
if (role == null) {
// No role stashed (legacy path): fall back to the historical interpretation.
return terminal != null ? Principal.worker(terminal, pid) : Principal.primary(pid);
}
return new Principal(Role.valueOf(role.toString()), terminal, pid);
return new Principal(Role.valueOf(role.toString()), terminal, pid, name);
}
/**
@@ -285,6 +339,24 @@ public final class BridgeMcp {
return error(reason + ": " + caller.describe() + " may not " + action);
}
/**
* Update the legacy singleton "primary" fallback used for no-delegation inbox nudges (CB-548).
*
* <p>Only {@link Role#PRIMARY} callers — the unnamed primary and named leads alike — may claim
* it. An architect delegates as its own pane but must never become the fallback: the per-target
* delegation map ({@code PrimaryRegistry#recordDelegation}) does not cure the singleton, so an
* architect left here would draw nudges that belong to a primary. The decision uses the resolved
* role, never name/kind sniffing. A null {@code caller} (legacy/no-auth path) records nothing.
*
* <p>Split out of the tool handlers so the guard is unit-testable without fabricating an SDK
* {@code McpSyncServerExchange} (same pattern as {@link #denyFor}/{@link #principalFrom}).
*/
static void recordPrimarySingleton(PrimaryRegistry registry, String callerTerminal, Principal caller) {
if (caller != null && caller.isPrimary()) {
registry.record(callerTerminal);
}
}
/** The worker identity resolved from this call's connection, or {@code null} if the primary. */
private static String callerTerminal(McpSyncServerExchange exchange) {
Object v = exchange.transportContext().get(CALLER_TERMINAL);
@@ -292,6 +364,13 @@ public final class BridgeMcp {
return (s == null || s.isBlank()) ? null : s;
}
/** The lead name resolved from this call's connection, or {@code null} (CB-530). */
private static String callerName(McpSyncServerExchange exchange) {
Object v = exchange.transportContext().get(CALLER_NAME);
String s = v == null ? null : v.toString();
return (s == null || s.isBlank()) ? null : s;
}
/** The caller's PID resolved from this call's connection, or {@code -1} if unknown. */
private static long callerPid(McpSyncServerExchange exchange) {
Object v = exchange.transportContext().get(CALLER_PID);
@@ -311,6 +390,13 @@ public final class BridgeMcp {
return transport;
}
/** Mark a connected spawned member available for the injector readiness gate. */
static void markSpawnedMemberPresent(Principal caller, MemberPresence presence) {
if (caller.isSpawnedMember()) {
presence.markPresent(caller.terminal());
}
}
/** Graceful shutdown of the MCP server. */
public void close() {
server.closeGracefully();
@@ -318,14 +404,25 @@ public final class BridgeMcp {
// --- tool logic (thin adapters over the services; unit-testable) ---------------------------
/** {@code bridge_send}: delegate {@code content} to a worker session and block for its reply. */
static McpSchema.CallToolResult send(MessageService messages, String sessionId, String content, Long timeoutMs) {
/**
* {@code bridge_send}: delegate {@code content} to a worker session and block for its reply.
* The configured profiles are required so a profile name can never bypass target validation.
*
* (CB-548): {@code onAccepted} records delegator ownership the instant the send is accepted, so
* a BUSY interloper never claims a turn it did not win. {@code null} disables recording.
*/
static McpSchema.CallToolResult send(MessageService messages, String sessionId, String content,
Long timeoutMs, Runnable onAccepted, Set<String> profiles) {
if (isBlank(sessionId) || isBlank(content)) {
return error("sessionId and content are required");
}
McpSchema.CallToolResult targetError = profileTargetError(sessionId, profiles);
if (targetError != null) {
return targetError;
}
long timeout = clamp(timeoutMs == null ? DEFAULT_TIMEOUT_MS : timeoutMs);
try {
return formatReply(messages.send(sessionId, content, timeout), timeout);
return formatReply(messages.send(sessionId, content, timeout, onAccepted), timeout);
} catch (HerdrException e) {
return error("herdr error contacting session " + sessionId + ": " + e.getMessage());
}
@@ -378,6 +475,11 @@ public final class BridgeMcp {
"[worker finished without a structured bridge_reply — transcript tail follows]\n" + r.text());
// The worker ran the turn then wedged (CB-109) — surface the error context.
case WORKER_FAILED -> text("[worker failed — turn ended in an unrecoverable state]\n" + r.text());
// The backend refused on a subscription usage limit (CB-578 stage A) — the worker's
// pane stayed healthy, but its account is exhausted. Distinct from WORKER_FAILED so the
// primary gets the real cause, not a generic wedge.
case BACKEND_EXHAUSTED -> text("[backend exhausted — the worker's account refused on a "
+ "usage limit]\n" + r.text());
// The worker paused mid-turn to ask (CB-205) — tell the primary how to answer in-turn.
case QUESTION -> text("[question] the worker paused to ask before it can finish:\n" + r.text()
+ "\n\nAnswer it by calling bridge_send again with turnId=\"" + r.turnId()
@@ -392,15 +494,33 @@ public final class BridgeMcp {
/**
* {@code bridge_send} with {@code wait:false}: delegate {@code content} and return a ticket
* immediately (fire-and-poll), so a long task isn't cut off by the caller's MCP call timeout.
* The configured profiles are required so a profile name can never bypass target validation.
*
* This wires the accepted-delivery hook
* (CB-548) so an async flooding send records delegator ownership exactly once it is accepted.
*/
static McpSchema.CallToolResult sendAsync(MessageService messages, String sessionId, String content) {
static McpSchema.CallToolResult sendAsync(MessageService messages, String sessionId, String content,
Runnable onAccepted, Set<String> profiles) {
if (isBlank(sessionId) || isBlank(content)) {
return error("sessionId and content are required");
}
String ticket = messages.sendAsync(sessionId, content);
McpSchema.CallToolResult targetError = profileTargetError(sessionId, profiles);
if (targetError != null) {
return targetError;
}
String ticket = messages.sendAsync(sessionId, content, onAccepted);
return text("accepted — task delegated. Poll bridge_poll with ticket=" + ticket);
}
/** A configured profile is never a send target; other unknown values may be herdr-owned panes. */
private static McpSchema.CallToolResult profileTargetError(String sessionId, Set<String> profiles) {
if (profiles.contains(sessionId)) {
return error("unknown send target \"" + sessionId + "\": it is a configured profile name, not a "
+ "session id. Call bridge_list to find a member or lead sessionId.");
}
return null;
}
/** {@code bridge_poll}: check an async delegation by ticket, or drain a worker's inbox by target. */
static McpSchema.CallToolResult poll(MessageService messages, String ticket, String target) {
if (!isBlank(target)) {
@@ -422,6 +542,9 @@ public final class BridgeMcp {
? "[done — worker finished without a structured bridge_reply; transcript tail follows]\n" + v.reply()
: v.reply());
case PENDING -> text("[pending — " + v.detail() + "]");
case ASKING -> text("[question — worker is waiting for your answer]\n" + v.reply()
+ "\n\nAnswer it by calling bridge_send again with turnId=\"" + v.turnId()
+ "\" and content set to your answer; the worker resumes the same turn.");
case FAILED -> text("[failed — " + v.detail() + "]");
};
}
@@ -484,7 +607,29 @@ public final class BridgeMcp {
static McpSchema.CallToolResult whoami(Principal caller, SessionManager sessions) {
Map<String, Object> m = new LinkedHashMap<>();
m.put("role", caller.role().name().toLowerCase());
if (caller.isArchitect()) {
// CB-548: the role reads "architect"; the name is the gateway-local slot the pane is
// bound to, and the pane itself so a peer knows where to reach it.
if (caller.name() != null) {
m.put("architect", caller.name());
}
if (caller.terminal() != null) {
m.put("sessionId", caller.terminal());
}
return text(json(m));
}
if (!caller.isWorker()) {
// CB-530: which lead, once more than one pane is configured as one. `role` deliberately
// still reads "primary" — the fallback ladder in CLAUDE.md keys on it, and a lead IS a
// primary for authorization; the name is additive so no existing reader breaks.
if (caller.name() != null) {
m.put("leader", caller.name());
}
// CB-532: a lead's own pane, so it can tell a peer where to reach it — and so an
// operator can read off which tab hosts which lead without going to herdr.
if (caller.terminal() != null) {
m.put("sessionId", caller.terminal());
}
return text(json(m));
}
m.put("sessionId", caller.terminal());
@@ -512,27 +657,43 @@ public final class BridgeMcp {
/** {@code bridge_spawn} without cwd/caller context (default resolution). */
static McpSchema.CallToolResult spawn(SessionManager sessions, String profile) {
return spawn(sessions, profile, null, null, null, null);
return spawn(sessions, profile, null, null, null, null, null, null, null);
}
/**
* {@code bridge_spawn}: launch a guard-checked worker for {@code profile} (blank → the default
* profile) and return its session id + pane id. The worker's cwd is {@code requestedCwd} if given,
* else the profile's config, else {@code callerCwd} (the primary's directory), else the daemon's.
* {@code bridge_spawn}: launch a guard-checked member for {@code profile} (blank → the default
* profile) under {@code role} (blank → {@code dev}), and return its session id + pane id. The
* member's cwd is {@code requestedCwd} if given, else the profile's config, else
* {@code callerCwd} (the primary's directory), else the daemon's.
* CB-301: the session is registered with {@code ownerTerminal} as its owner.
* CB-301-ext: {@code worktreeRequest} non-null provisions an isolated git worktree.
* CB-584: {@code sessionName}/{@code resumeSessionId} request agent session identity — a resumed
* conversation requires an explicit {@code profile} whose adapter declares
* {@link dev.ltms.bridged.peer.Capability#SESSION_RESUME}, or the spawn is refused rather than
* silently starting a cold session.
*
* <p>{@code role} and {@code profile} are independent: the role picks the contract, the profile
* picks the backend. A reviewer on the same profile as the dev it reviews is a normal spawn.
*/
static McpSchema.CallToolResult spawn(SessionManager sessions, String profile,
static McpSchema.CallToolResult spawn(SessionManager sessions, String profile, String role,
String requestedCwd, String callerCwd,
String ownerTerminal, WorktreeRequest worktreeRequest) {
String ownerTerminal, WorktreeRequest worktreeRequest,
String sessionName, String resumeSessionId) {
MemberRole memberRole;
try {
WorkerSession worker = sessions.acquire(isBlank(profile) ? null : profile,
requestedCwd, callerCwd, ownerTerminal, worktreeRequest);
return text(json(workerView(worker)));
memberRole = isBlank(role) ? MemberRole.DEV : MemberRole.parse(role);
} catch (IllegalArgumentException e) {
return error(e.getMessage());
}
try {
MemberSession member = sessions.acquire(isBlank(profile) ? null : profile, memberRole,
requestedCwd, callerCwd, ownerTerminal, worktreeRequest,
isBlank(sessionName) ? null : sessionName, isBlank(resumeSessionId) ? null : resumeSessionId);
return text(json(memberView(member)));
} catch (GuardException e) {
return error("subscription boundary: " + e.getMessage());
} catch (IllegalArgumentException e) {
return error(e.getMessage()); // unknown / no-default profile
return error(e.getMessage()); // unknown / no-default profile, or a refused resumeSessionId
} catch (PeerUnreachableException e) {
return error("spawn timed out — worker pane never reached injectable state: " + e.getMessage());
} catch (HerdrException e) {
@@ -568,29 +729,165 @@ public final class BridgeMcp {
return null;
}
/** {@code bridge_profiles}: the configured worker profiles and the default. */
static McpSchema.CallToolResult profiles(PeerLauncher workers) {
return text(json(Map.of(
"profiles", workers.profiles(),
"default", workers.defaultProfile() == null ? "" : workers.defaultProfile())));
/**
* {@code bridge_profiles}: the configured worker profiles, the default, and — CB-578 stage B —
* which of them are currently quarantined (backend exhausted) and for how much longer. The
* {@code quarantined} key is present only when at least one profile is, so a fleet where nothing
* has ever been quarantined gets exactly the pre-stage-B shape.
*/
static McpSchema.CallToolResult profiles(PeerLauncher workers, QuarantineSource quarantine) {
Map<String, Object> result = new LinkedHashMap<>();
result.put("profiles", workers.profiles());
result.put("default", workers.defaultProfile() == null ? "" : workers.defaultProfile());
Map<String, Object> quarantined = new LinkedHashMap<>();
for (String profile : workers.profiles()) {
String credentialId = quarantine.credentialIdFor().apply(profile);
if (credentialId == null) {
continue;
}
quarantine.quarantine().remainingSeconds(credentialId).ifPresent(remaining -> {
Map<String, Object> row = new LinkedHashMap<>();
row.put("credentialId", credentialId);
row.put("quarantinedForSeconds", remaining);
quarantined.put(profile, row);
});
}
if (!quarantined.isEmpty()) {
result.put("quarantined", quarantined);
}
return text(json(result));
}
/** {@code bridge_list}: bridge-owned roster merged with live herdr status by paneId. */
static McpSchema.CallToolResult listWorkers(PeerLauncher workers, SessionManager sessions) {
/**
* {@code bridge_list}: the whole fleet — {@code leads} and {@code workers} — each merged with
* live herdr status. CB-519 decoupled the registry key (a host-unique id) from the herdr pane
* coordinate, so the join is on the terminal id, which both the session and the live agent carry.
*
* <p>CB-535 added the {@code leads} half. Until then this listed the worker roster alone, and a
* lead asking "who else is here?" got an empty array — which reads as <em>no peers</em> but
* actually means <em>no workers spawned</em>. There was no way at all for a lead to learn a
* peer's address; it had to be carried across by a human. Both halves are reported even when a
* half is empty, so an empty {@code workers} can no longer be mistaken for an empty fleet.
*
* <p>Leads are drawn from the resolver rather than from a second registry, so an address listed
* here is one that would actually resolve as a lead — see {@link CallerResolver#leads()}. The
* caller's own row is flagged {@code "self": true}: a peer needs to tell its own pane apart from
* a peer's, and the alternative is every lead calling {@code bridge_whoami} to subtract itself.
*
* <p>CB-583: the {@code capacity} rows reuse {@code quarantine} (the same {@link QuarantineSource}
* {@code bridge_profiles} reads) so the two surfaces cannot disagree about which profile is
* quarantined — see {@link #capacityView}.
*
* @param leads terminal_id → lead name, live from the resolver
* @param selfTerm the calling pane's terminal id, or blank for a caller with no pane
*/
static McpSchema.CallToolResult listFleet(PeerLauncher workers, SessionManager sessions,
Map<String, String> leads, String selfTerm) {
return listFleet(workers, sessions, null, CapacitySource.none(), new HealthCoverageSource(() -> "off"),
QuarantineSource.none(), leads, selfTerm);
}
static McpSchema.CallToolResult listFleet(PeerLauncher workers, SessionManager sessions, MessageService messages,
CapacitySource capacity, HealthCoverageSource healthCoverage,
QuarantineSource quarantine, Map<String, String> leads, String selfTerm) {
try {
Map<String, Agent> live = workers.list().stream()
.map(Agent.class::cast)
.filter(a -> a.paneId() != null)
.collect(Collectors.toMap(Agent::paneId, Function.identity(), (_, b) -> b));
List<Map<String, Object>> out = sessions.roster().stream()
.map(s -> SessionManager.rosterView(s, live.get(s.paneId())))
.filter(a -> a.terminalId() != null)
.collect(Collectors.toMap(Agent::terminalId, Function.identity(), (_, b) -> b));
List<Map<String, Object>> leadRows = leads.entrySet().stream()
.sorted(Map.Entry.comparingByValue())
.map(e -> leadView(e.getKey(), e.getValue(), live.get(e.getKey()), selfTerm))
.toList();
return text(json(Map.of("workers", out)));
List<MemberSession> roster = sessions.roster();
List<Map<String, Object>> out = roster.stream()
.map(s -> memberCapacityView(s, live.get(s.terminalId()), messages, capacity.clock().getAsLong()))
.toList();
Set<String> profiles = new java.util.TreeSet<>(capacity.configuredProfiles().get());
roster.stream().map(MemberSession::profile).forEach(profiles::add);
Map<String, Object> result = new LinkedHashMap<>();
result.put("leads", leadRows); result.put("members", out);
result.put("healthCoverage", healthCoverage.value().get());
if (capacity.available()) result.put("capacity", profiles.stream()
.map(profile -> capacityView(profile, capacity.liveCount(), capacity.maxLoad(), roster, messages,
capacity.clock().getAsLong(), quarantine)).toList());
return text(json(result));
} catch (HerdrException e) {
return error("herdr error listing workers: " + e.getMessage());
return error("herdr error listing the fleet: " + e.getMessage());
}
}
/**
* Capacity is advisory only. {@code reclaimable} says there is no bridge work, not that bridged
* may stop the member: the bridge has capacity facts but no work list, and choosing work needs
* authority it does not have. {@code idleForSeconds} is derived from monotonic nanoTime and has
* no wall-clock meaning across a daemon restart.
*/
private static Map<String, Object> memberCapacityView(MemberSession session, Agent live,
MessageService messages, long nowNanos) {
Map<String, Object> row = SessionManager.rosterView(session, live);
boolean open = messages != null && messages.hasAcceptedDelivery(session.terminalId());
boolean inbox = messages != null && messages.hasInboxMessage(session.terminalId());
boolean reclaimable = (session.state() == MemberSession.State.READY || session.state() == MemberSession.State.DONE)
&& !open && !inbox;
row.put("reclaimable", reclaimable);
row.put("idleForSeconds", reclaimable ? Math.max(0, (nowNanos - session.lastActivityAtNanos()) / 1_000_000_000L) : null);
return row;
}
/**
* CB-583: {@code free} alone cannot tell a lead "busy, will free up" from "refusing, and
* nothing changes for N seconds" — those need different decisions. So a quarantined profile
* forces {@code free} to 0, whatever its {@code maxLoad}/{@code live} say, and the row carries
* the same {@code credentialId}/{@code quarantinedForSeconds} facts {@code bridge_profiles}
* reports, reusing {@link QuarantineSource} rather than a second lookup. Both new keys are
* added only when the profile is actually quarantined, so an ordinary fleet's rows are
* byte-identical to before this change.
*/
private static Map<String, Object> capacityView(String profile, Function<String, Integer> liveCount,
Function<String, Integer> maxLoad, List<MemberSession> roster,
MessageService messages, long nowNanos, QuarantineSource quarantine) {
Integer cap = maxLoad.apply(profile);
int live = liveCount.apply(profile);
int reclaimable = (int) roster.stream().filter(s -> profile.equals(s.profile()))
.filter(s -> (s.state() == MemberSession.State.READY || s.state() == MemberSession.State.DONE))
.filter(s -> messages == null || (!messages.hasAcceptedDelivery(s.terminalId()) && !messages.hasInboxMessage(s.terminalId())))
.count();
Map<String, Object> row = new LinkedHashMap<>();
row.put("profile", profile); row.put("maxLoad", cap); row.put("live", live);
row.put("free", cap == null ? null : Math.max(0, cap - live)); row.put("reclaimable", reclaimable);
String credentialId = quarantine.credentialIdFor().apply(profile);
if (credentialId != null) {
quarantine.quarantine().remainingSeconds(credentialId).ifPresent(remaining -> {
row.put("free", 0);
row.put("credentialId", credentialId);
row.put("quarantinedForSeconds", remaining);
});
}
return row;
}
/**
* One lead's row: its address, its name, and whether it can be reached right now.
*
* <p>{@code status} is herdr's live view, and {@code unknown} when herdr is not tracking that
* pane as an agent — the honest answer, and the one that matters: a lead whose pane herdr cannot
* see is a lead a {@code bridge_send} cannot be typed into. It is reported rather than hidden,
* because a peer that has gone unreachable is exactly what the sender needs to know.
*/
private static Map<String, Object> leadView(String terminal, String name, Agent live,
String selfTerm) {
Map<String, Object> m = new LinkedHashMap<>();
m.put("sessionId", terminal);
m.put("name", name);
m.put("status", live == null || live.status() == null
? "unknown" : live.status().name().toLowerCase());
if (terminal.equals(selfTerm)) {
m.put("self", true);
}
return m;
}
/** {@code bridge_stop}: tear a worker down by its pane id. */
static McpSchema.CallToolResult stop(SessionManager sessions, String paneId) {
if (isBlank(paneId)) {
@@ -605,10 +902,14 @@ public final class BridgeMcp {
}
/** CB-301 projection from the authoritative session registry. */
private static Map<String, Object> workerView(WorkerSession s) {
private static Map<String, Object> memberView(MemberSession s) {
Map<String, Object> m = new LinkedHashMap<>();
m.put("sessionId", s.terminalId());
m.put("paneId", s.paneId());
// Echo the role back so the caller can see what it actually got, not what it meant to ask
// for — a spawn that silently fell back to dev is otherwise invisible.
m.put("role", s.role() == null ? "dev" : s.role().wireName());
m.put("profile", s.profile());
m.put("status", s.state().name().toLowerCase());
if (s.worktree() != null) {
m.put("worktree", s.worktree());
@@ -685,29 +986,58 @@ public final class BridgeMcp {
private static McpSchema.Tool spawnTool() {
return tool("bridge_spawn",
"Spawn a new off-subscription worker session. Pass a profile (from bridge_profiles) to "
+ "pick the backend, or omit it for the default. The worker opens your current "
+ "directory by default; pass cwd to pin a different one. Pass worktree:true (with "
+ "ticket) or worktree:<ticket-slug> to provision an isolated git worktree. "
+ "Returns the worker's sessionId (use with bridge_send) and paneId (use with bridge_stop).",
"Spawn a new off-subscription member session. A member has two independent attributes: "
+ "role (what it is for) and profile (which backend it runs on). Pass role to pick "
+ "the contract — 'dev' implements a unit and opens its own PR, 'reviewer' reviews a "
+ "diff it did not write, 'architect' refines a ticket before anyone builds it; omit "
+ "it for 'dev'. Pass profile (from bridge_profiles) to pick the backend, or omit it "
+ "for the default. The two are independent: a reviewer may run on the same profile "
+ "as the dev it reviews. The member opens your current directory by default; pass "
+ "cwd to pin a different one. Pass worktree:true (with ticket) or "
+ "worktree:<ticket-slug> to provision an isolated git worktree. Pass resumeSessionId "
+ "to relaunch onto a prior conversation instead of starting cold — this requires an "
+ "explicit profile whose backend supports it (bridge_list shows agentSessionId for "
+ "resumable members), and is refused otherwise rather than silently starting fresh. "
+ "sessionName gives the member a display name in its own UI when the backend supports "
+ "one. Returns the member's sessionId (use with bridge_send) and paneId (use with "
+ "bridge_stop).",
objectSchema(Map.of(
"profile", stringProp("Worker profile to spawn (omit for the default profile)"),
"cwd", stringProp("Working directory for the worker (omit to inherit yours)"),
"role", stringProp("What the member is for: architect, dev or reviewer (default dev)"),
"profile", stringProp("Which backend to run it on (omit for the default profile)"),
"cwd", stringProp("Working directory for the member (omit to inherit yours)"),
"worktree", Map.of("type", "string", "description", "'true' or a ticket slug — requests an isolated git worktree"),
"ticket", stringProp("Ticket slug when worktree:true")),
"ticket", stringProp("Ticket slug when worktree:true"),
"sessionName", stringProp("Logical display name for the member's own session, when its backend supports one"),
"resumeSessionId", stringProp("A prior member's agentSessionId (from bridge_list) to resume — requires an explicit profile that supports it")),
List.of()));
}
private static McpSchema.Tool profilesTool() {
return tool("bridge_profiles",
"List the configured worker profiles (backends) and which one bridge_spawn uses by default.",
"List the configured worker profiles (backends) and which one bridge_spawn uses by "
+ "default. A 'quarantined' map is present when a backend-exhausted refusal put "
+ "a profile's credential on cooldown — bridge_spawn onto it is refused until "
+ "quarantinedForSeconds elapses; a profile sharing that credential is listed too.",
objectSchema(Map.of(), List.of()));
}
private static McpSchema.Tool listTool() {
return tool("bridge_list",
"List the worker sessions the bridge tracks — each with its sessionId, paneId, profile, "
+ "state, optional worktree/branch/owner, and live herdr status.",
"List the whole fleet the bridge tracks, in two parts. 'leads' are your PEERS — other "
+ "orchestrators, each with its sessionId (the address to bridge_send to), "
+ "name, live status, and 'self': true on your own row; this is how you "
+ "discover a peer lead without being told its address. 'members' are the "
+ "sessions delegated to — each with sessionId, paneId, role (architect/dev/"
+ "reviewer), profile (the backend it runs on), state, optional "
+ "worktree/branch/owner/agentSessionId (the id to pass as bridge_spawn's "
+ "resumeSessionId to relaunch onto that same conversation, when the backend "
+ "supports it), and live herdr status. An empty 'members' "
+ "means no members are spawned; it says nothing about peers. When capacity "
+ "facts are configured, a 'capacity' row per profile also reports free: 0 for "
+ "a quarantined profile's credential (see bridge_profiles), whatever its "
+ "maxLoad/live — with credentialId and quarantinedForSeconds naming the "
+ "quarantine, so 'free: 0, busy' can be told apart from 'free: 0, refusing "
+ "for N seconds'.",
objectSchema(Map.of(), List.of()));
}
@@ -720,10 +1050,12 @@ public final class BridgeMcp {
}
private static McpSchema.Tool replyTool() {
// No session/target arg — the worker's identity is resolved from the connection.
// No session/target arg — the caller's identity is resolved from the connection.
return tool("bridge_reply",
"Return your structured answer for the task you were delegated, "
+ "resolving the caller's blocked bridge_send.",
"Return your structured answer for a message you were sent, resolving the sender's "
+ "blocked bridge_send. A worker MUST end every delegated turn with exactly "
+ "one of these. A lead uses it only to answer another lead that messaged "
+ "it — never to answer a worker, whose turn it is not.",
objectSchema(Map.of(
"content", stringProp("Your reply/answer")),
List.of("content")));
@@ -741,11 +1073,14 @@ public final class BridgeMcp {
return tool("bridge_whoami",
"Report who YOU are on the bridge — your role is resolved from your connection "
+ "(unforgeable), never from anything you claim. Returns role 'primary' (you "
+ "orchestrate: spawn/send/stop, and you must never call bridge_reply) or "
+ "'worker' (you were delegated to: you must end every turn with exactly one "
+ "bridge_reply, and cannot spawn or send), plus your own sessionId, profile, "
+ "worktree and branch when you are a worker. Call this first when following "
+ "role-conditional instructions rather than guessing your role.",
+ "orchestrate: spawn/send/stop; reply ONLY to answer a peer lead that "
+ "messaged you, never to answer a worker), 'architect' (you delegate turns "
+ "and reply/ask as your own pane, but cannot spawn/stop/drain), or 'worker' "
+ "(you were delegated to: you must end every turn with exactly one "
+ "bridge_reply, and cannot spawn or send), plus 'leader'/'architect' naming "
+ "which one you are, your own sessionId, and profile/worktree/branch when "
+ "you are a worker. Call this first when following role-conditional "
+ "instructions rather than guessing.",
objectSchema(Map.of(), List.of()));
}
@@ -4,6 +4,7 @@ import org.slf4j.Logger;
import org.slf4j.LoggerFactory;
import java.util.Optional;
import java.util.concurrent.ConcurrentHashMap;
import java.util.concurrent.atomic.AtomicReference;
/**
@@ -25,6 +26,15 @@ public final class PrimaryRegistry {
private final AtomicReference<String> terminal = new AtomicReference<>();
private final boolean pinned;
/**
* CB-532: worker terminal → the lead that delegated to it. The single slot above answers "who is
* THE primary", a question with no correct answer once two leads orchestrate the same fleet:
* whichever called {@code bridge_send} first captured every nudge, including nudges for the
* other lead's delegations. This map answers the question that actually matters — "who is
* waiting on THIS worker" — and is what lets {@code primary.terminal} be retired.
*/
private final ConcurrentHashMap<String, String> leadByTarget = new ConcurrentHashMap<>();
/**
* @param pinnedTerminal an optional pinned terminal from config ({@code null}/blank = unpinned)
*/
@@ -56,6 +66,45 @@ public final class PrimaryRegistry {
}
}
/**
* Record that {@code leadTerminal} owns the accepted delegation of worker {@code target} (CB-532).
*
* <p>Called from the {@code MessageService} accepted-delivery hook — only after a send has won
* the session's send lock and queued delivery — where both halves are known (CB-548). It is
* deliberately <em>not</em> called at {@code bridge_send} request time: a concurrent sender that
* times out {@code BUSY} must not steal a live delegation's reply routing without ever owning
* the turn. Last writer wins — if a second lead's later send is accepted, replies follow the
* lead that most recently delegated to it, which is the one waiting.
*/
public void recordDelegation(String target, String leadTerminal) {
if (target == null || target.isBlank() || leadTerminal == null || leadTerminal.isBlank()) {
return;
}
leadByTarget.put(target, leadTerminal);
}
/** Forget a worker's delegating lead — call on release, so a torn-down session leaks nothing. */
public void forgetDelegation(String target) {
if (target != null) {
leadByTarget.remove(target);
}
}
/**
* Where a nudge about {@code target}'s reply should go: the lead that delegated to it, falling
* back to the single known primary.
*
* <p>The fallback matters after a daemon restart, which loses the map while the durable inbox
* keeps the reply. With one lead the fallback is unambiguous and correct. With several and no
* recorded delegation there is no right answer, so this returns empty rather than guessing —
* delivery degrades to pull, which is exactly what the durable inbox is for, instead of
* interrupting the wrong lead with someone else's result.
*/
public Optional<String> nudgeTargetFor(String target) {
String lead = target == null ? null : leadByTarget.get(target);
return lead != null ? Optional.of(lead) : Optional.ofNullable(terminal.get());
}
/** The known primary terminal, or empty if not yet learned (and not pinned). */
public Optional<String> primaryTerminal() {
return Optional.ofNullable(terminal.get());
@@ -0,0 +1,325 @@
package dev.ltms.bridged.member;
import dev.ltms.bridged.config.BridgedConfig;
import dev.ltms.bridged.guard.SubscriptionGuard;
import dev.ltms.bridged.herdr.Agent;
import dev.ltms.bridged.herdr.AgentControl;
import dev.ltms.bridged.herdr.WorkspaceControl;
import dev.ltms.bridged.peer.Capability;
import org.slf4j.Logger;
import org.slf4j.LoggerFactory;
import java.util.EnumSet;
import java.util.List;
import java.util.Map;
import java.util.Set;
import java.util.UUID;
import java.util.function.Function;
import java.util.function.LongSupplier;
import java.util.function.Supplier;
/**
* The {@link HerdrPeerLauncher} adapter for <strong>Claude Code</strong> — the safe path from a
* delegation request to a running off-subscription Claude.
*
* <p>Everything transport-related (tab/pane placement, the CB-306 spawn-readiness gate, unique
* naming, CB-117 orphan reap, teardown, listing, cwd resolution) lives in the base. This class
* supplies only the two Claude-specific seams:
* <ul>
* <li>the {@code claude} name prefix (so reap matches {@code claude-*} panes, never another
* adapter's), and</li>
* <li>{@link #buildLaunch}, which encodes the subscription boundary: build the worker env with
* {@code ANTHROPIC_BASE_URL}, assert that host is on the allowlist <em>before</em> touching
* herdr, and mount the bridge MCP + reply charter as inline launch flags. A worker's base_url
* lives in the env map handed to herdr and nowhere else; {@code bridged}'s own environment is
* never mutated, and nothing is written to the worker's profile.</li>
* </ul>
*/
public final class ClaudeCodeLauncher extends HerdrPeerLauncher {
/** Label prefix for this adapter's herdr agent names (drives naming + orphan reap). */
private static final String NAME_PREFIX = "claude";
private static final Logger log = LoggerFactory.getLogger(ClaudeCodeLauncher.class);
private final SubscriptionGuard guard;
/**
* Production constructor — disables the spawn-ready gate ({@code spawnReadyTimeoutMs == 0}) so
* existing deployments and tests keep the legacy non-blocking spawn semantics.
*/
public ClaudeCodeLauncher(AgentControl agents, WorkspaceControl spaces, SubscriptionGuard guard,
Map<String, BridgedConfig.Profile> profiles, String defaultProfile,
Function<String, String> env) {
this(agents, spaces, guard, profiles, defaultProfile, env, 0,
System::currentTimeMillis, () -> sleepUninterruptibly(300));
}
/**
* Production constructor with spawn-ready gate enabled. The gate polls {@code agents.status()}
* until the pane reports an injectable state or {@code spawnReadyTimeoutMs} elapses.
*/
public ClaudeCodeLauncher(AgentControl agents, WorkspaceControl spaces, SubscriptionGuard guard,
Map<String, BridgedConfig.Profile> profiles, String defaultProfile,
Function<String, String> env,
long spawnReadyTimeoutMs, long spawnReadyPollMs) {
this(agents, spaces, guard, profiles, defaultProfile, env,
spawnReadyTimeoutMs, spawnReadyPollMs, null);
}
/**
* Production constructor carrying the fleet-wide tab-label template (CB-557). The template comes
* from {@code fleet.tabLabel}, which a profile cannot know because it names the member's
* <em>role</em>; a profile may still override it with its own {@code tabLabel}.
*/
public ClaudeCodeLauncher(AgentControl agents, WorkspaceControl spaces, SubscriptionGuard guard,
Map<String, BridgedConfig.Profile> profiles, String defaultProfile,
Function<String, String> env,
long spawnReadyTimeoutMs, long spawnReadyPollMs,
Supplier<BridgedConfig.Fleet> fleet) {
this(agents, spaces, guard, profiles, defaultProfile, env,
spawnReadyTimeoutMs,
System::currentTimeMillis, () -> sleepUninterruptibly(spawnReadyPollMs),
fleet);
}
/**
* Full testability constructor. Every injectable collaborator is explicit so unit tests supply
* fakes for the clock ({@code nowMillis}) and poll-loop wait ({@code sleeper}). The
* {@code sleeper} is never called when the gate is disabled ({@code spawnReadyTimeoutMs == 0}).
*
* @param agents herdr agent control (start, status, close)
* @param spaces workspace / tab control (ensure, create, close)
* @param guard subscription-boundary guard (checked before spawning)
* @param profiles configured worker profiles
* @param defaultProfile profile a no-argument spawn uses (nullable)
* @param env host env lookup (injectable for tests)
* @param spawnReadyTimeoutMs max ms to wait for injectable state (0 disables the gate)
* @param nowMillis monotonic clock source (e.g. {@code System::currentTimeMillis})
* @param sleeper sleep/wait hook (e.g. {@code () -> Thread.sleep(pollMs)}); it
* already encodes the poll interval, so the 8th positional argument
* (poll ms) is accepted for API symmetry but otherwise unused here
*/
public ClaudeCodeLauncher(AgentControl agents, WorkspaceControl spaces, SubscriptionGuard guard,
Map<String, BridgedConfig.Profile> profiles, String defaultProfile,
Function<String, String> env,
long spawnReadyTimeoutMs,
LongSupplier nowMillis, Runnable sleeper) {
this(agents, spaces, guard, profiles, defaultProfile, env,
spawnReadyTimeoutMs, nowMillis, sleeper, null);
}
/**
* Full testability constructor, plus the fleet-wide tab-label template (CB-557).
*
* @param fleet live fleet config, read once for each spawn
*/
public ClaudeCodeLauncher(AgentControl agents, WorkspaceControl spaces, SubscriptionGuard guard,
Map<String, BridgedConfig.Profile> profiles, String defaultProfile,
Function<String, String> env,
long spawnReadyTimeoutMs,
LongSupplier nowMillis, Runnable sleeper,
Supplier<BridgedConfig.Fleet> fleet) {
super(NAME_PREFIX, agents, spaces, profiles, defaultProfile, env,
spawnReadyTimeoutMs, nowMillis, sleeper, fleet);
this.guard = guard;
}
/**
* {@inheritDoc}
*
* <p>The spawn sequence encodes the subscription boundary: assert the profile's base_url is on
* the allowlist <em>before</em> any herdr call, then build the worker env with
* {@code ANTHROPIC_*}, the parity-neutral git-forge grant, and the bridge MCP + reply charter
* mounted as inline launch flags. When the request carries session identity (CB-547a) it is
* applied here — see {@link #applySessionIdentity}.
*/
@Override
protected Launch buildLaunch(BridgedConfig.Profile cfg, LaunchSpec spec) {
// CB-539: a profile may deliberately opt into the subscription (subscription: true) when no
// off-subscription endpoint exists for it — e.g. `sonnet` on `ccs`. That profile gets no
// ANTHROPIC_BASE_URL/AUTH_TOKEN (there is nothing to point them at) and the guard's base_url
// requirement is skipped FOR IT ONLY. Every other profile keeps the hard boundary below.
boolean onSubscription = cfg.isSubscription();
String baseUrl = cfg.baseUrl();
if (onSubscription) {
// NO SILENT CONTRADICTION: subscription:true + a baseUrl state opposite intents; refuse
// loudly rather than pick a winner.
if (baseUrl != null && !baseUrl.isBlank()) {
throw new IllegalStateException("profile '" + cfg.profile()
+ "' sets both subscription: true and a baseUrl ('" + baseUrl + "') — the two "
+ "are contradictory: a subscription profile must not point at an endpoint. "
+ "Drop baseUrl, or drop subscription: true.");
}
// Visible without anyone going looking for it: this worker bills the subscription.
log.warn("spawning profile '{}' on the Claude subscription (subscription: true) — this "
+ "worker WILL bill the operator's subscription", cfg.profile());
} else {
guard.assertWorker(baseUrl); // hard stop before we spawn anything
}
Map<String, String> workerEnv = baseEnv(cfg);
if (onSubscription) {
// CB-542 belt-and-braces: on the subscription path no guard vets these two keys, and the
// profile's env: is layered in by baseEnv — so strip any that rode in there. Config load
// already rejects this (loudly, naming the profile); this makes the boundary hold even
// for a profile built in code that never passed through that validation.
workerEnv.remove("ANTHROPIC_BASE_URL");
workerEnv.remove("ANTHROPIC_AUTH_TOKEN");
} else {
workerEnv.put("ANTHROPIC_BASE_URL", baseUrl);
putIfPresent(workerEnv, "ANTHROPIC_AUTH_TOKEN", env.apply(cfg.tokenEnv()));
}
putIfPresent(workerEnv, "ANTHROPIC_MODEL", cfg.model());
putIfPresent(workerEnv, "CLAUDE_CONFIG_DIR", cfg.configDir());
applyGitToken(workerEnv, cfg);
// CB-547a: Claude Code can MINT its own session id, so bridged chooses it — a fresh spawn
// gets a UUID we pass as --session-id and return from agentSessionId(), so the resume
// handle is known BEFORE the agent has written anything; a resume spawn adopts its prior
// id via -r and passes no --session-id (the two conflict). Both are injected before the
// model flag so --model keeps outranking the operator's own argv.
// mutableArgv: argvWithBridge may hand back the profile's own (immutable) List.of when it
// has neither MCP nor a charter — session flags must be added into a list we own.
List<String> argv = mutableArgv(argvWithBridge(cfg, spec.charter()));
String agentSessionId = applySessionIdentity(argv, spec.sessionName(), spec.resumeSessionId());
return new Launch(workerEnv, argvWithModel(argv, cfg), agentSessionId);
}
/**
* Add the Claude-specific session-identity flags to {@code argv} and return the peer's OWN
* session id — the resume handle. A resume request passes the prior id via {@code -r} and
* returns that id; a fresh named session mints a new UUID, passes it via {@code --session-id},
* and returns the mint. The bridge's logical name rides along as {@code -n} when present. When
* <em>no</em> identity is requested (sessionName and resumeSessionId both blank) this adds
* nothing and returns {@code null}, keeping the legacy no-identity launch byte-identical.
*/
private static String applySessionIdentity(List<String> argv, String sessionName, String resumeSessionId) {
boolean resuming = resumeSessionId != null && !resumeSessionId.isBlank();
boolean named = sessionName != null && !sessionName.isBlank();
if (!resuming && !named) {
return null; // no identity requested — keep the legacy launch byte-identical
}
if (named) {
argv.add("-n");
argv.add(sessionName);
}
if (resuming) {
argv.add("-r");
argv.add(resumeSessionId);
return resumeSessionId;
}
String minted = UUID.randomUUID().toString();
argv.add("--session-id");
argv.add(minted);
return minted;
}
/**
* The launch argv, plus an inline {@code --mcp-config} when {@code worker.mcpUrl} is set and
* {@code --append-system-prompt} when the base composed a charter. Neither touches the profile's
* config; both are pure command-line flags. This inline-flag mount is Claude Code specific —
* other adapters mount MCP and instructions their own way.
*/
private List<String> argvWithBridge(BridgedConfig.Profile cfg, String charter) {
if (!cfg.hasMcp() && charter == null) {
return cfg.argv();
}
List<String> argv = mutableArgv(cfg.argv());
if (cfg.hasMcp()) {
String mcpJson = "{\"mcpServers\":{\"bridge\":{\"type\":\"http\",\"url\":\""
+ cfg.mcpUrl() + "\"}}}";
argv.add("--mcp-config");
argv.add(mcpJson);
}
if (charter != null) {
argv.add("--append-system-prompt");
argv.add(charter);
}
return argv;
}
/**
* Pin the model on the command line as well as in {@code ANTHROPIC_MODEL} (CB-533).
*
* <p>The env var alone is not a reliable pin for this adapter, because the argv is usually a
* launcher rather than {@code claude} itself — {@code ["ccs", "<profile>"]} — and {@code ccs}
* exports its profile's own model family ({@code ANTHROPIC_MODEL}, {@code DEFAULT_OPUS/SONNET/
* HAIKU}, {@code CLAUDE_CODE_SUBAGENT_MODEL}) over whatever it inherited. A worker profile that
* set {@code model:} therefore got silently overruled by its own launcher. Claude Code's
* {@code --model} flag outranks the environment, and {@code ccs <profile> [claude-args...]}
* passes trailing arguments through, so the flag survives the wrapper.
*
* <p>Appended last so it also outranks anything in the operator's own {@code argv}. Profiles
* that deliberately leave {@code model:} unset (letting {@code ccs} own model selection, as
* {@code gx10} does) are untouched — this adds nothing when there is nothing to add. This is
* the {@code kind: claude} counterpart of the opencode adapter's {@code -m provider/model}.
*/
private static List<String> argvWithModel(List<String> argv, BridgedConfig.Profile cfg) {
if (cfg.model() == null || cfg.model().isBlank()) {
return argv;
}
List<String> withModel = mutableArgv(argv);
withModel.add("--model");
withModel.add(cfg.model());
return withModel;
}
// --- Agent-returning convenience spawns (used by callers/tests that want the herdr Agent) ---
/** Spawn a worker for the default profile in the resolved default cwd. */
public Agent spawn() {
return spawnInternal(null, null, null);
}
/** Spawn a worker for a named profile (null → default) in the resolved default cwd. */
public Agent spawn(String profileName) {
return spawnInternal(profileName, null, null);
}
/** Spawn a worker for a named profile with an explicit requested/caller cwd (CB-112). */
public Agent spawn(String profileName, String requestedCwd, String callerCwd) {
return spawnInternal(profileName, requestedCwd, callerCwd);
}
// --- capabilities --------------------------------------------------------------------------
@Override
public Set<Capability> capabilities() {
Set<Capability> caps = EnumSet.of(Capability.MID_TURN_ASK, Capability.WORKTREE,
Capability.CONTEXT_RESET, Capability.ORPHAN_REAP,
Capability.SESSION_NAME, Capability.SESSION_RESUME);
if (hasGitTokenProfile()) {
caps.add(Capability.SELF_PR);
}
return Set.copyOf(caps);
}
@Override
public boolean clearContext(String id) {
String target = agentTarget(id);
if (target == null) {
return false;
}
// This deliberately bypasses Injector: /clear is housekeeping, not a delegated turn.
agents().send(target, "/clear");
return true;
}
/** Whether any configured profile opts into a git-forge token (required for {@link Capability#SELF_PR}). */
private boolean hasGitTokenProfile() {
return profileConfigs().stream().anyMatch(BridgedConfig.Profile::hasGitToken);
}
// --- CB-117 reap predicate (Claude prefix), kept for direct unit testing -------------------
/**
* Whether {@code name} is a Claude Code bridge worker started by a <em>different</em> process
* than {@code currentNonce}. A thin {@code claude}-prefix binding of
* {@link HerdrPeerLauncher#isForeignWorker(String, String, String)}.
*/
static boolean isForeignWorker(String name, String currentNonce) {
return HerdrPeerLauncher.isForeignWorker(NAME_PREFIX, name, currentNonce);
}
}
@@ -0,0 +1,510 @@
package dev.ltms.bridged.member;
import dev.ltms.bridged.config.BridgedConfig;
import dev.ltms.bridged.herdr.Agent;
import dev.ltms.bridged.peer.Capability;
import dev.ltms.bridged.peer.MemberRole;
import dev.ltms.bridged.peer.PeerHandle;
import dev.ltms.bridged.peer.PeerLauncher;
import dev.ltms.bridged.peer.PeerUnreachableException;
import dev.ltms.bridged.peer.SpawnRequest;
import dev.ltms.bridged.placement.BackendQuarantine;
import dev.ltms.bridged.placement.PlacementCandidate;
import dev.ltms.bridged.placement.PlacementContext;
import dev.ltms.bridged.placement.PlacementException;
import dev.ltms.bridged.placement.PlacementPolicies;
import dev.ltms.bridged.placement.PlacementPolicy;
import org.slf4j.Logger;
import org.slf4j.LoggerFactory;
import java.util.ArrayList;
import java.util.Collections;
import java.util.EnumSet;
import java.util.HashSet;
import java.util.LinkedHashMap;
import java.util.List;
import java.util.Map;
import java.util.Objects;
import java.util.Set;
import java.util.concurrent.ConcurrentHashMap;
import java.util.function.Function;
import java.util.function.Supplier;
import java.util.stream.Collectors;
/**
* The {@link PeerLauncher} the core actually holds when more than one adapter is configured — a thin
* router in front of one {@link HerdrPeerLauncher} per peer {@code kind} (Claude Code, opencode, …).
* It owns no transport of its own; it dispatches each SPI call to the delegate that owns the profile
* involved, and fans the fleet-wide queries (list/reap/caps/profiles) across all delegates.
*
* <p>Routing rules:
* <ul>
* <li><strong>By profile</strong> — {@link #spawn}, {@link #effectiveCwd}, {@link #parityOverlay}
* resolve the profile (a null/blank name → the global {@link #defaultProfile}) and delegate to
* the single adapter that declares it. Profiles partition cleanly across adapters: the
* constructor rejects a name claimed by two.</li>
* <li><strong>By pane id</strong> — {@link #stop} routes to the adapter that spawned that pane
* (recorded at spawn time). A pane the composite never spawned (only real for a caller that
* hand-rolls an id) falls back to the first delegate; teardown is pane-id addressed and
* tab cleanup is single-occupant guarded, so it is safe either way.</li>
* <li><strong>Fleet-wide</strong> — {@link #reapOrphanWorkers} and {@link #capabilities} fan out
* and combine. {@link #list} is deduplicated by pane id because every herdr-backed delegate
* shares one herdr connection and so reports the same global agent set.</li>
* </ul>
*
* <p>CB-518: an unqualified spawn is routed through a {@link PlacementPolicy}. The default
* {@code fixed} policy reproduces the historical default-profile behaviour; {@code weighted} uses
* smooth weighted round-robin with {@code maxLoad} gating. If a chosen profile fails with
* {@link PeerUnreachableException}, the composite advances to the next available candidate and
* retries, bounded by the number of candidates.
*/
public final class CompositePeerLauncher implements PeerLauncher {
private static final Logger log = LoggerFactory.getLogger(CompositePeerLauncher.class);
private final List<HerdrPeerLauncher> delegates;
private final Map<String, HerdrPeerLauncher> byProfile;
private final String defaultProfile;
/** paneId → the delegate that spawned it, so {@link #stop} tears down through the right adapter. */
private final Map<String, HerdrPeerLauncher> spawnedBy = new ConcurrentHashMap<>();
private final Function<String, Integer> liveCount;
/**
* CB-559: the placement inputs are read <em>per spawn</em>, not captured at construction, so a
* config reload changes where the next member lands without a restart. These are the hot keys —
* role pools, an existing profile's weight/maxLoad, and the placement policy. What cannot change
* this way is the set of adapters ({@link #byProfile}), because a new backend needs a launcher
* and launchers are built once; {@code ConfigRef} classifies that as deferred and says so.
*/
private final Supplier<Map<String, BridgedConfig.Profile>> profileConfigs;
private final Supplier<PlacementPolicy> placementPolicy;
/** CB-578 stage B: credential cooldown, checked before an explicit spawn and filtered into placement. */
private final BackendQuarantine quarantine;
/**
* CB-557: the role pools an unqualified spawn draws its candidates from. A supplier that yields
* {@code null}, and an empty pool for a role, both fall back to every configured profile — the
* pre-CB-557 behaviour.
*/
private final Supplier<BridgedConfig.Fleet> fleet;
/**
* Backward-compatible constructor: fixed placement, no live-counting. Use this for tests and
* simple wiring; it preserves the pre-CB-518 behaviour exactly.
*
* @param delegates one adapter per configured peer kind; must be non-empty and declare
* disjoint profile-name sets
* @param defaultProfile the profile a no-argument spawn resolves to (may be null)
* @throws IllegalArgumentException if {@code delegates} is empty or two adapters claim one profile
*/
public CompositePeerLauncher(List<HerdrPeerLauncher> delegates, String defaultProfile) {
this(delegates, defaultProfile, Map.of(), PlacementPolicies.fixed(), _ -> 0);
}
/**
* Production constructor with a placement policy and live-worker counter. Quarantine (CB-578
* stage B) is off for this constructor — {@link BackendQuarantine#none()} — since it predates
* the feature and existing callers of this exact overload never exercised it; use the 7-arg
* overload below to wire a real {@link BackendQuarantine}.
*
* @param delegates one adapter per configured peer kind; must be non-empty and declare
* disjoint profile-name sets
* @param defaultProfile the profile a no-argument spawn resolves to under {@code fixed} policy
* @param profileConfigs all configured worker profiles (used for candidate weights/caps)
* @param placementPolicy which policy governs unqualified spawns
* @param liveCount live worker count per profile (must never return {@code null})
*/
public CompositePeerLauncher(List<HerdrPeerLauncher> delegates,
String defaultProfile,
Map<String, BridgedConfig.Profile> profileConfigs,
PlacementPolicy placementPolicy,
Function<String, Integer> liveCount) {
this(delegates, defaultProfile, profileConfigs, placementPolicy, liveCount, null, BackendQuarantine.none());
}
/**
* Production constructor with role pools (CB-557). An unqualified spawn draws its candidates from
* {@code fleet.<role>} instead of from every configured profile, so a reviewer is placed on a
* reviewer backend and never on, say, the architect-only one. Quarantine is off for this
* constructor too, for the same reason as the 5-arg overload above.
*
* @param fleet the configured role pools; {@code null} ⇒ every profile is a candidate for every
* role, which is the pre-CB-557 behaviour
*/
public CompositePeerLauncher(List<HerdrPeerLauncher> delegates,
String defaultProfile,
Map<String, BridgedConfig.Profile> profileConfigs,
PlacementPolicy placementPolicy,
Function<String, Integer> liveCount,
BridgedConfig.Fleet fleet) {
this(delegates, defaultProfile, profileConfigs, placementPolicy, liveCount, fleet, BackendQuarantine.none());
}
/**
* Production constructor with role pools and quarantine (CB-578 stage B). The full-featured
* non-reloading form; {@link #CompositePeerLauncher(List, String, Supplier, Function, BackendQuarantine)}
* is what {@code Bridged.main} actually wires up.
*
* @param quarantine required — pass {@link BackendQuarantine#none()} for a caller that does not
* want the feature, never a defaulting overload (CB-578 stage B's own rule).
*/
public CompositePeerLauncher(List<HerdrPeerLauncher> delegates,
String defaultProfile,
Map<String, BridgedConfig.Profile> profileConfigs,
PlacementPolicy placementPolicy,
Function<String, Integer> liveCount,
BridgedConfig.Fleet fleet,
BackendQuarantine quarantine) {
// LinkedHashMap, not Map.copyOf: candidates() promises definition order and the weighted
// policy breaks exact-weight ties on it, so a salted iteration order would make placement
// differ from one JVM run to the next.
this(delegates, defaultProfile,
constant(Collections.unmodifiableMap(new LinkedHashMap<>(profileConfigs))),
constant(placementPolicy), liveCount, constant(fleet), quarantine);
}
/**
* Production constructor that re-reads its placement inputs per spawn (CB-559), so a config
* reload retargets the next member without a restart.
*
* @param config the live configuration — read at every spawn, never captured
* @param quarantine required — CB-578 stage B; pass {@link BackendQuarantine#none()} to opt out
*/
public CompositePeerLauncher(List<HerdrPeerLauncher> delegates,
String defaultProfile,
Supplier<BridgedConfig> config,
Function<String, Integer> liveCount,
BackendQuarantine quarantine) {
this(delegates, defaultProfile,
() -> config.get().profiles(),
() -> PlacementPolicies.fromName(config.get().placement()),
liveCount,
() -> config.get().fleet(),
quarantine);
}
/** The all-suppliers form every other constructor funnels into. */
private CompositePeerLauncher(List<HerdrPeerLauncher> delegates,
String defaultProfile,
Supplier<Map<String, BridgedConfig.Profile>> profileConfigs,
Supplier<PlacementPolicy> placementPolicy,
Function<String, Integer> liveCount,
Supplier<BridgedConfig.Fleet> fleet,
BackendQuarantine quarantine) {
this.fleet = fleet;
this.quarantine = Objects.requireNonNull(quarantine, "quarantine");
if (delegates.isEmpty()) {
throw new IllegalArgumentException("at least one peer adapter must be configured");
}
this.delegates = List.copyOf(delegates);
this.defaultProfile = defaultProfile;
this.profileConfigs = profileConfigs;
this.placementPolicy = placementPolicy;
this.liveCount = liveCount;
Map<String, HerdrPeerLauncher> index = new LinkedHashMap<>();
for (HerdrPeerLauncher d : this.delegates) {
for (String profile : d.profiles()) {
HerdrPeerLauncher prev = index.putIfAbsent(profile, d);
if (prev != null) {
throw new IllegalArgumentException(
"worker profile '" + profile + "' is claimed by two peer adapters");
}
}
}
// Order-preserving for the same reason, and because profiles() is user-visible (bridge_profiles).
this.byProfile = Collections.unmodifiableMap(index);
}
/** A supplier of a value fixed at construction — how the non-reloading constructors funnel in. */
private static <T> Supplier<T> constant(T value) {
return () -> value;
}
/**
* The currently-configured profiles, never null.
*
* <p>Read fresh on every call so a reload is visible; a caller that needs two consistent reads
* takes one local, as {@link #poolFor} does.
*/
private Map<String, BridgedConfig.Profile> profiles0() {
Map<String, BridgedConfig.Profile> m = profileConfigs.get();
return m == null ? Map.of() : m;
}
/** The adapter owning {@code profileName} (null/blank → the default). Throws on an unknown profile. */
private HerdrPeerLauncher route(String profileName) {
String resolved = (profileName == null || profileName.isBlank()) ? defaultProfile : profileName;
if (resolved == null) {
// No profile and no default configured — hand to the first delegate so it raises the
// same "no default" error it would on its own; keeps the SPI contract single-sourced.
return delegates.getFirst();
}
HerdrPeerLauncher d = byProfile.get(resolved);
if (d == null) {
throw new IllegalArgumentException("unknown worker profile: " + resolved);
}
return d;
}
@Override
public PeerHandle spawn(SpawnRequest req) {
String requestedProfile = req.profileName();
if (requestedProfile != null && !requestedProfile.isBlank()) {
// An explicit profile bypasses the placement policy, but not the capacity cap: maxLoad
// is documented as an unconditional limit on this profile (BridgedConfig.Profile), and
// the charter makes explicit-profile spawns the normal path — so skipping the check
// here would leave the cap dead config in real operation.
HerdrPeerLauncher d = route(requestedProfile);
enforceNotQuarantined(requestedProfile);
enforceMaxLoad(requestedProfile);
PeerHandle handle = d.spawn(req);
spawnedBy.put(handle.id(), d);
return handle;
}
// CB-557: an unqualified spawn is placed inside the pool of the role it asked for, not across
// the whole profile list. An EXPLICIT profile (above) is left alone on purpose — it is the
// operator overriding, and refusing it would break `bridge_spawn{profile:"opus"}`, which
// carries no role and so would be judged against the dev pool it was never meant for.
List<PlacementCandidate> candidates = candidates(req.role());
String roleDefault = defaultProfileFor(req.role());
Set<String> unreachable = new HashSet<>();
// CB-578 stage B: computed once up front — a quarantine's expiry cannot pass within one spawn
// call, so re-deriving it per retry would only cost work, never change the answer.
Set<String> quarantined = quarantinedProfiles(candidates);
PlacementContext ctx = new PlacementContext(roleDefault, candidates, liveCount, unreachable, quarantined);
int maxAttempts = candidates.isEmpty() ? 1 : candidates.size();
for (int attempt = 0; attempt < maxAttempts; attempt++) {
// Deliberately uncaught: when no candidate is left (all at cap, or all unreachable) the
// policy already throws a clear message. Catching it to rethrow a generic
// PeerUnreachableException would replace a precise diagnosis with a vague one.
PlacementCandidate chosen = placementPolicy.get().select(ctx);
HerdrPeerLauncher d = byProfile.get(chosen.profile());
if (d == null) {
// A configured profile with no adapter is a wiring bug; fail fast.
unreachable.add(chosen.profile());
continue;
}
// CB-547a: route the chosen profile but keep the caller's session identity — dropping it
// here would silently sever the resume handle on every policy-routed spawn. CB-557: the
// role rides along for the same reason, or a routed spawn would be labelled as a dev.
SpawnRequest routedReq = new SpawnRequest(chosen.profile(), req.requestedCwd(), req.callerCwd(),
req.sessionName(), req.resumeSessionId(), req.role());
try {
PeerHandle handle = d.spawn(routedReq);
spawnedBy.put(handle.id(), d);
return handle;
} catch (PeerUnreachableException e) {
log.warn("spawn on profile {} unreachable, will retry next candidate if any: {}",
chosen.profile(), e.getMessage());
unreachable.add(chosen.profile());
// Update the context for the next selection so the policy excludes this profile.
ctx = new PlacementContext(roleDefault, candidates, liveCount, unreachable, quarantined);
}
}
throw new PeerUnreachableException(
"no reachable worker profile available after trying " + unreachable.size()
+ " candidate(s): " + String.join(", ", unreachable));
}
/**
* Refuse an explicit-profile spawn when the profile is at its {@code maxLoad} cap.
*
* <p>maxLoad is a documented, unconditional capacity limit (see {@code BridgedConfig.Profile#maxLoad}),
* and the charter makes explicit-profile spawns the normal path — so enforcing it only in placement
* ({@code PlacementPolicyUtil}, package-private, hence not linked) would leave the cap dead config
* on every call that names a profile. Same rule as placement: {@code live >= cap} is at capacity.
*
* <p>Deliberately no fallback to another profile: the caller named {@code profile} for a cost/model
* reason, and silently re-routing a paid-tier (subscription) request elsewhere is worse than
* refusing it. A caller that wants placement should omit the profile and let the policy pick.
*
* <p>Known TOCTOU limitation — documented, not fixed. {@link #liveCount} is read outside any lock and
* {@code SessionManager} registers a session only after {@code launcher.spawn} returns, so two
* genuinely concurrent spawns can both pass this check. The race already exists on the placement
* path. Closing it needs slot reservation in the registry; serializing spawn here would block on
* the readiness gate and is a far worse trade.
*
* @param profile the profile the caller explicitly named
* @throws PlacementException when the profile is at capacity
*/
/**
* Refuse an explicit-profile spawn whose credential is quarantined (CB-578 stage B): a prior
* {@code BACKEND_EXHAUSTED} classification on this profile, or on another profile sharing its
* {@code credentialId}, is still on cooldown.
*
* <p>Checked before {@link #enforceMaxLoad}, and for the same reason that check exists: an
* explicit profile bypasses the placement policy's filtering entirely, so without this the cap
* (there, quarantine here) would be dead config on the very path the charter calls normal.
* Deliberately no fallback to another profile, matching {@link #enforceMaxLoad}'s own reasoning —
* the caller named this profile for a cost/model reason.
*
* @throws PlacementException naming the profile, its credential, and the remaining cooldown
*/
private void enforceNotQuarantined(String profile) {
String credentialId = credentialIdFor(profile);
quarantine.remainingSeconds(credentialId).ifPresent(remaining -> {
throw new PlacementException("worker profile '" + profile + "' is quarantined "
+ "(credential '" + credentialId + "' exhausted; ~" + remaining
+ "s remaining) — refusing spawn");
});
}
/** {@code profile}'s credential group (CB-578 stage B), or the profile's own name if unconfigured. */
private String credentialIdFor(String profile) {
BridgedConfig.Profile cfg = profiles0().get(profile);
return cfg == null ? profile : cfg.effectiveCredentialId();
}
/** The subset of {@code candidates} whose credential is currently quarantined (CB-578 stage B). */
private Set<String> quarantinedProfiles(List<PlacementCandidate> candidates) {
return candidates.stream()
.map(PlacementCandidate::profile)
.filter(p -> quarantine.isQuarantined(credentialIdFor(p)))
.collect(Collectors.toSet());
}
private void enforceMaxLoad(String profile) {
// Absent config, or a config whose maxLoad normalized to null (ABSENT ⇒ unlimited at load),
// means no cap — never cap what wasn't configured. Note "non-positive ⇒ unlimited" was true
// until CB-585: an explicit `maxLoad: 0` now survives as 0 and is a real cap of zero, so the
// check below refuses every spawn on that profile, and a negative value is refused at config
// load rather than normalized away.
BridgedConfig.Profile cfg = profiles0().get(profile);
Integer cap = (cfg == null) ? null : cfg.maxLoad();
if (cap == null) {
return;
}
int live = liveCount.apply(profile);
if (live >= cap) {
throw new PlacementException("worker profile '" + profile + "' is at maxLoad: " + live
+ " live >= " + cap + " cap; refusing spawn — no fallback to another profile");
}
}
/**
* The profile names {@code role} may be placed on, in definition order.
*
* <p>An empty or absent pool means "unconstrained", not "nothing allowed": a config that declares
* no pool for a role must keep spawning, so it falls back to every configured profile. Names in a
* pool that no adapter declares are dropped here rather than thrown — config load already rejects
* a pool entry with no profile, so a survivor is a profile this particular composite does not own.
*/
private List<String> poolFor(MemberRole role) {
Map<String, BridgedConfig.Profile> configured = profiles0();
BridgedConfig.Fleet f = fleet.get();
List<String> pool = (f == null) ? List.of() : f.profilesFor(role);
List<String> known = pool.stream().filter(configured::containsKey).toList();
return known.isEmpty() ? List.copyOf(configured.keySet()) : known;
}
/** The profile an unqualified spawn for {@code role} falls back to under {@code fixed} placement. */
private String defaultProfileFor(MemberRole role) {
List<String> pool = poolFor(role);
return pool.isEmpty() ? defaultProfile : pool.getFirst();
}
/** Build the candidate list from {@code role}'s pool, in definition order. */
private List<PlacementCandidate> candidates(MemberRole role) {
List<PlacementCandidate> out = new ArrayList<>();
for (String name : poolFor(role)) {
BridgedConfig.Profile w = profiles0().get(name);
if (w != null) {
out.add(new PlacementCandidate(name, null, w.weight(), w.maxLoad()));
}
}
return out;
}
@Override
public String effectiveCwd(SpawnRequest req) {
return route(req.profileName()).effectiveCwd(req);
}
@Override
public List<String> parityOverlay(String profileName) {
return route(profileName).parityOverlay(profileName);
}
@Override
public void stop(String id) {
HerdrPeerLauncher d = spawnedBy.remove(id);
if (d == null) {
log.debug("stop({}) — no recorded owner, routing to the first adapter (pane-addressed)", id);
d = delegates.getFirst();
}
d.stop(id);
}
@Override
public boolean clearContext(String id) {
HerdrPeerLauncher delegate = spawnedBy.get(id);
if (delegate == null) {
log.debug("clearContext({}) ignored — no recorded owning adapter", id);
return false;
}
return delegate.clearContext(id);
}
@Override
public Set<String> profiles() {
return byProfile.keySet();
}
@Override
public String defaultProfile() {
return defaultProfile;
}
/**
* {@inheritDoc}
*
* <p>Routes to the specific delegate {@code profileName} resolves to, not the fleet-wide
* union {@link #capabilities()} returns — the whole reason this method exists (CB-584): in a
* mixed fleet, one adapter's capability must never be read as every profile's.
*/
@Override
public Set<Capability> capabilitiesFor(String profileName) {
return route(profileName).capabilities();
}
/** Every herdr agent, deduplicated by pane id (all delegates share one herdr and list globally). */
@Override
public List<Agent> list() {
Map<String, Agent> byPane = new LinkedHashMap<>();
for (HerdrPeerLauncher d : delegates) {
for (Agent a : d.list()) {
if (a.paneId() != null) {
byPane.putIfAbsent(a.paneId(), a);
}
}
}
return List.copyOf(byPane.values());
}
@Override
public int reapOrphanWorkers() {
int reaped = 0;
for (HerdrPeerLauncher d : delegates) {
reaped += d.reapOrphanWorkers();
}
return reaped;
}
/** The union of every adapter's capabilities — a capability any adapter offers, the fleet offers. */
@Override
public Set<Capability> capabilities() {
EnumSet<Capability> caps = EnumSet.noneOf(Capability.class);
for (HerdrPeerLauncher d : delegates) {
caps.addAll(d.capabilities());
}
return Set.copyOf(caps);
}
}
@@ -0,0 +1,876 @@
package dev.ltms.bridged.member;
import dev.ltms.bridged.config.BridgedConfig;
import dev.ltms.bridged.herdr.Agent;
import dev.ltms.bridged.herdr.AgentControl;
import dev.ltms.bridged.herdr.HerdrException;
import dev.ltms.bridged.herdr.Tab;
import dev.ltms.bridged.herdr.Workspace;
import dev.ltms.bridged.herdr.WorkspaceControl;
import dev.ltms.bridged.peer.Capability;
import dev.ltms.bridged.peer.CharterReceipt;
import dev.ltms.bridged.peer.MemberRole;
import dev.ltms.bridged.peer.PeerHandle;
import dev.ltms.bridged.peer.PeerLauncher;
import dev.ltms.bridged.peer.PeerUnreachableException;
import dev.ltms.bridged.peer.SpawnRequest;
import org.slf4j.Logger;
import org.slf4j.LoggerFactory;
import java.security.SecureRandom;
import java.util.ArrayList;
import java.util.Collection;
import java.util.LinkedHashMap;
import java.util.List;
import java.util.Map;
import java.util.Set;
import java.util.UUID;
import java.util.concurrent.ConcurrentHashMap;
import java.util.concurrent.ConcurrentMap;
import java.util.concurrent.atomic.AtomicBoolean;
import java.util.concurrent.atomic.AtomicLong;
import java.util.function.Function;
import java.util.function.LongSupplier;
import java.util.function.Supplier;
import java.util.regex.Matcher;
import java.util.regex.Pattern;
/**
* Abstract base for {@link PeerLauncher} adapters that materialize a peer as a <em>herdr</em>
* agent (a CLI coding agent running in a herdr tab/pane). It owns everything that is the same
* regardless of <em>which</em> coding agent runs: tab/pane placement, the CB-306 spawn-readiness
* gate, unique naming, CB-117 orphan reap, teardown, {@link #list() listing}, and cwd resolution.
*
* <p>Two seams are peer-specific and supplied by the concrete adapter:
* <ul>
* <li>{@code namePrefix} (constructor arg) — the label prefix ({@code claude}, {@code opencode})
* that drives both unique naming and the orphan-reap pattern, so each adapter reaps only its
* own kind of pane and never another's.</li>
* <li>{@link #buildLaunch(BridgedConfig.Profile, LaunchSpec)} — the peer-specific env map + argv, including any
* subscription/guard check, MCP mount, and instruction injection. The base never sees how the
* peer is configured; it only places and starts the returned {@link Launch}.</li>
* </ul>
*
* <p>Placement: in the default {@code tab} policy a peer lands in its own tab inside a dedicated
* worker space (found-or-created once, then shared), so peers never split or clutter the user's
* real work spaces. Teardown removes the peer's pane <em>and</em> its now-empty tab, tolerating an
* already-gone peer so a repeated DELETE is harmless.
*/
public abstract class HerdrPeerLauncher implements PeerLauncher {
private static final Logger log = LoggerFactory.getLogger(HerdrPeerLauncher.class);
/** herdr rejects a duplicate agent {@code name}; we retry a bumped name this many times. */
private static final int NAME_RETRIES = 8;
/**
* Retries for {@code agent.start} against a seed pane whose shell has not reached its prompt
* yet — {@code tab.create}/{@code pane.split} return as soon as the pane exists, and herdr
* refuses to start an agent in a pane that is not "an available shell" ({@code agent_pane_busy}).
*/
private static final int SHELL_READY_RETRIES = 20;
private final String namePrefix; // label prefix: naming + reap scheme
private final AgentControl agents;
private final WorkspaceControl spaces;
private final Map<String, BridgedConfig.Profile> profiles; // profile name → spawn settings
private final String defaultProfile; // profile a no-arg spawn uses (nullable)
/** Host env lookup (injectable for tests); adapters read it in {@link #buildLaunch}. */
protected final Function<String, String> env;
private final AtomicLong nameSeq = new AtomicLong(); // per-peer counter (herdr agent names only)
/**
* Live fleet config, read once per spawn. A null supplier or value leaves tab labels at their
* default and supplies no role charter. A profile's own {@code tabLabel} still overrides it.
*
* <p>CB-559: a supplier rather than a snapshot, so a config reload affects the next launch
* without a restart. Existing tabs keep the label they were given.
*/
private final Supplier<BridgedConfig.Fleet> fleet;
/** The final instruction always requires a bridge reply when the bridge MCP is mounted. */
protected static final String REPLY_CHARTER =
"You are a spawned member in the claude-bridge fleet. Every message you receive arrives "
+ "through the bridge, and the ONLY channel back to the sender is the bridge_reply MCP tool. "
+ "Text you write in your terminal is NOT sent anywhere — the sender cannot see your screen, "
+ "so an in-terminal answer is silently discarded. Therefore you MUST end EVERY turn by calling "
+ "bridge_reply with `content` set to your complete response. This holds for every message without "
+ "exception — tasks, questions, clarifications, acknowledgements, and ordinary back-and-forth "
+ "conversation. Call bridge_reply exactly once, as the final action of your turn, with your full "
+ "answer in `content`; never wait for confirmation first. If you end a turn without calling "
+ "bridge_reply, the sender receives nothing and the exchange stalls.";
/**
* Tab numbers, counted per {@code role/profile} pair (CB-557).
*
* <p>Deliberately not {@link #nameSeq}. That counter is shared by every profile this launcher
* serves, because its job is to make herdr <em>agent names</em> unique. Reusing it for the tab
* label made the numbers global, so sibling tabs read {@code #4}, {@code #9}, {@code #17} — gaps
* that look like a member died. Counting per role+profile makes {@code dev: sonnet #2} mean the
* second sonnet dev, which is what a reader assumes it means.
*
* <p>Resets when the daemon restarts, and that is fine: the label is a human-facing hint, not an
* identity. Identity is {@link PeerHandle#id()}.
*/
private final ConcurrentMap<String, AtomicLong> labelSeq = new ConcurrentHashMap<>();
private final long spawnReadyTimeoutMs; // 0 = disable gate (legacy non-blocking spawn)
private final LongSupplier nowMillis; // monotonic clock (injectable for tests)
private final Runnable sleeper; // sleep/wait hook (injectable for tests; never real-sleep in unit tests)
// Per-process token mixed into each peer name so a fresh process (nameSeq back at 0) cannot
// collide with same-profile peers that outlived a restart. See startUniquelyNamed.
private final String nameNonce = String.format("%06x", new SecureRandom().nextInt(1 << 24));
// CB-519: PeerHandle.id() is a host-unique opaque UUID, decoupled from the herdr pane id. The
// routing/registry key is the UUID; the herdr pane id is a launcher-private placement/teardown
// coordinate. This map bridges the two so stop(id) can resolve a host-unique key back to the
// exact pane it must tear down. The pane id is launcher-private (never the routing key) — see
// PeerHandle.id().
private final ConcurrentMap<String, String> paneByAgentId = new ConcurrentHashMap<>();
private final AtomicBoolean resetUnsupportedLogged = new AtomicBoolean();
/**
* @param namePrefix label prefix for this peer kind (drives naming and reap)
* @param agents herdr agent control (start, status, close)
* @param spaces workspace / tab control (ensure, create, close)
* @param profiles configured peer profiles
* @param defaultProfile profile a no-argument spawn uses (nullable)
* @param env host env lookup (injectable for tests)
* @param spawnReadyTimeoutMs max ms to wait for injectable state (0 disables the gate)
* @param nowMillis monotonic clock source (e.g. {@code System::currentTimeMillis})
* @param sleeper sleep/wait hook (never called when the gate is disabled); the poll
* interval is baked into this hook, so the base needs no poll field
*/
protected HerdrPeerLauncher(String namePrefix, AgentControl agents, WorkspaceControl spaces,
Map<String, BridgedConfig.Profile> profiles, String defaultProfile,
Function<String, String> env,
long spawnReadyTimeoutMs,
LongSupplier nowMillis, Runnable sleeper) {
this(namePrefix, agents, spaces, profiles, defaultProfile, env, spawnReadyTimeoutMs,
nowMillis, sleeper, null);
}
/**
* As above, plus the live {@code fleet} config (CB-557).
*
* @param fleet live fleet config, read once per spawn; {@code null} ⇒ default tab label and no
* role charter. A separate constructor rather than a new parameter on the one
* above, so every existing call site keeps the default without an edit.
*/
protected HerdrPeerLauncher(String namePrefix, AgentControl agents, WorkspaceControl spaces,
Map<String, BridgedConfig.Profile> profiles, String defaultProfile,
Function<String, String> env,
long spawnReadyTimeoutMs,
LongSupplier nowMillis, Runnable sleeper,
Supplier<BridgedConfig.Fleet> fleet) {
this.fleet = fleet;
this.namePrefix = namePrefix;
this.agents = agents;
this.spaces = spaces;
this.profiles = Map.copyOf(profiles);
this.defaultProfile = defaultProfile;
this.env = env;
this.spawnReadyTimeoutMs = spawnReadyTimeoutMs;
this.nowMillis = nowMillis;
this.sleeper = sleeper;
}
// --- adapter seams -------------------------------------------------------------------------
/**
* Build the peer-specific launch for {@code cfg}: the environment map and argv handed to herdr.
* Any subscription/guard check, MCP mount, and instruction injection happen here. The env map
* and argv are adapter-private; the base only places and starts what is returned.
*/
protected abstract Launch buildLaunch(BridgedConfig.Profile cfg, LaunchSpec spec);
/** Direct transport access for peer-specific, non-turn control operations. */
protected final AgentControl agents() {
return agents;
}
/** Resolve the public peer id to the launcher's private herdr target. */
protected final String agentTarget(String id) {
return paneByAgentId.get(id);
}
@Override
public boolean clearContext(String id) {
if (resetUnsupportedLogged.compareAndSet(false, true)) {
log.warn("context reset is unsupported for peer kind {}; clearAfterTurn is a no-op",
namePrefix);
}
return false;
}
/**
* A peer-specific launch: the herdr {@code env} map and {@code argv}, plus — for an adapter
* that carries durable session identity (CB-547a) — the peer's OWN session id
* ({@link PeerHandle#agentSessionId()}), known before the peer has written anything. Null for
* a launch that carries no identity.
*/
protected record Launch(Map<String, String> env, List<String> argv, String agentSessionId) {
/** A launch without a discoverable agent session id (an adapter that carries none). */
Launch(Map<String, String> env, List<String> argv) {
this(env, argv, null);
}
}
/** All per-spawn values adapters may need, including the base-composed effective charter. */
protected record LaunchSpec(String sessionName, String resumeSessionId, MemberRole role, String charter) {
}
// --- profile surface -----------------------------------------------------------------------
/** The configured peer profile names (what {@code spawn(profile)} accepts). */
@Override
public Set<String> profiles() {
return profiles.keySet();
}
/** The parity-overlay file list for {@code profileName} (default list when unset). */
@Override
public List<String> parityOverlay(String profileName) {
String name = (profileName == null || profileName.isBlank()) ? defaultProfile : profileName;
if (name == null || name.isBlank()) {
return List.of();
}
BridgedConfig.Profile cfg = profiles.get(name);
return cfg == null ? List.of() : cfg.parityOverlay();
}
/** The profile a no-argument spawn uses, or {@code null} if none is configured. */
@Override
public String defaultProfile() {
return defaultProfile;
}
/**
* {@inheritDoc}
*
* <p>One {@link HerdrPeerLauncher} instance always serves exactly one adapter kind, so every
* profile it owns shares that adapter's {@link #capabilities()} — {@code profileName} only
* needs validating (throwing on an unknown profile, same as {@link #spawn}), not routing.
*/
@Override
public Set<Capability> capabilitiesFor(String profileName) {
requireProfile(profileName);
return capabilities();
}
/** The configured profiles, for adapter capability decisions (e.g. any git-token grant). */
protected Collection<BridgedConfig.Profile> profileConfigs() {
return profiles.values();
}
/** Resolve {@code profileName} (null/blank → default) to its config, or throw with the options. */
protected BridgedConfig.Profile requireProfile(String profileName) {
String name = (profileName == null || profileName.isBlank()) ? defaultProfile : profileName;
if (name == null || name.isBlank()) {
throw new IllegalArgumentException("no default worker profile is configured — "
+ "pass a profile; configured: " + profiles.keySet());
}
BridgedConfig.Profile cfg = profiles.get(name);
if (cfg == null) {
throw new IllegalArgumentException("unknown worker profile '" + name
+ "' — configured: " + profiles.keySet());
}
return cfg;
}
// --- spawn ---------------------------------------------------------------------------------
/**
* A started peer plus the launch's agent-session id (the resume handle, or null) and the
* charter receipt (CB-571) the base composed for it.
*/
private record Spawned(Agent agent, String agentSessionId, CharterReceipt receipt) {
}
/**
* Spawn a peer. {@code profileName} null/blank → the default profile. The working directory
* (CB-112) is resolved by {@link #resolveCwd}: an explicit {@code requestedCwd}, else the
* profile's configured {@code cwd}, else {@code callerCwd} (the primary's cwd, when the spawn
* came from the primary over MCP), else the daemon's cwd — never assumed to be {@code $HOME}.
* The adapter's {@link #buildLaunch} runs before any herdr call.
*/
protected Agent spawnInternal(String profileName, String requestedCwd, String callerCwd) {
return spawnInternal(profileName, requestedCwd, callerCwd, null, null).agent();
}
/** Pre-CB-557 shape: no explicit role, so the tab is labelled as a {@code dev}. */
protected Spawned spawnInternal(String profileName, String requestedCwd, String callerCwd,
String sessionName, String resumeSessionId) {
return spawnInternal(profileName, requestedCwd, callerCwd, sessionName, resumeSessionId,
MemberRole.DEV);
}
/**
* Spawn a peer with session identity (CB-547a). The session values, role, and charter are
* threaded from the {@link SpawnRequest} into {@link #buildLaunch(BridgedConfig.Profile,
* LaunchSpec)}, and the launch's resolved agent-session id is returned alongside the agent so
* the caller can put it on the {@link PeerHandle}.
*/
protected Spawned spawnInternal(String profileName, String requestedCwd, String callerCwd,
String sessionName, String resumeSessionId, MemberRole role) {
BridgedConfig.Profile cfg = requireProfile(profileName);
BridgedConfig.Fleet liveFleet = fleet == null ? null : fleet.get();
String roleCharter = liveFleet == null ? null : liveFleet.charterFor(role);
String replyCharter = cfg.hasMcp() ? REPLY_CHARTER : null;
String charter = roleCharter == null ? replyCharter
: replyCharter == null ? roleCharter : roleCharter + "\n\n" + replyCharter;
// CB-571: fingerprint the exact composed charter bytes once, here in the base, before the
// string leaves for an adapter — so Claude and OpenCode derive the same digest. A failed
// start has no bridge_spawn result and no roster row, so the failure log below is the only
// surface the byte count can appear on. The charter text itself is never logged.
CharterReceipt receipt = CharterReceipt.compose(role, cfg.profile(), roleCharter, charter);
try {
Launch launch = buildLaunch(cfg, new LaunchSpec(sessionName, resumeSessionId, role, charter));
String cwd = resolveCwd(requestedCwd, cfg, callerCwd);
Agent agent = cfg.tabPlacement()
? spawnInTab(cfg, launch.env(), launch.argv(), cwd, role, liveFleet)
: spawnAsPane(cfg, launch.env(), launch.argv(), cwd, charter);
logCharterReceipt(receipt, true);
return new Spawned(agent, launch.agentSessionId(), receipt);
} catch (RuntimeException e) {
logCharterReceipt(receipt, false);
throw e;
}
}
/**
* The one place the charter's size and digest appear in the logs. {@code success} true after a
* start, false from the failure path of {@link #spawnInternal} where no handle or roster row
* exists to carry the receipt. Always metadata only — never the charter text.
*/
private static void logCharterReceipt(CharterReceipt receipt, boolean success) {
String role = receipt.role() == null ? "" : receipt.role().wireName();
if (success) {
log.info("spawned role={} profile={} charterSource={} charterSha256={} charterBytes={}",
role, receipt.profile(), receipt.charterSource(),
receipt.charterSha256(), receipt.charterBytes());
} else {
log.warn("spawn failed; charter role={} profile={} charterSource={} charterSha256={} charterBytes={}",
role, receipt.profile(), receipt.charterSource(),
receipt.charterSha256(), receipt.charterBytes());
}
}
/**
* The next tab number for {@code role} on {@code profile}, starting at 1.
*
* <p>Starts at 1 rather than 0 because the number is read by a person: {@code "dev: sonnet #1"}
* is the first one, and {@code #0} invites the question of where {@code #1} went.
*/
private long nextLabelSeq(MemberRole role, String profile) {
String key = (role == null ? "" : role.wireName()) + "/" + profile;
return labelSeq.computeIfAbsent(key, _ -> new AtomicLong()).incrementAndGet();
}
/**
* {@inheritDoc}
*
* <p>Delegates to {@link #spawnInternal} and wraps the resulting herdr {@link Agent} in a
* {@link WorkerHandle} whose {@link PeerHandle#id()} is a fresh <em>host-unique</em> opaque
* UUID (CB-519), deliberately decoupled from the herdr pane id: the id is the registry/routing
* key and must never collide across daemon processes on the same host, while the herdr pane id
* stays a launcher-private placement/teardown coordinate, remembered here so {@link #stop}
* can resolve the host-unique key back to its pane. When {@code spawnReadyTimeoutMs > 0},
* blocks until the peer's herdr status is injectable or the timeout elapses; on timeout the
* pane is closed (no orphan) and a {@link PeerUnreachableException} is thrown.
*/
@Override
public PeerHandle spawn(SpawnRequest req) {
Spawned spawned = spawnInternal(req.profileName(), req.requestedCwd(), req.callerCwd(),
req.sessionName(), req.resumeSessionId(), req.role());
Agent agent = spawned.agent();
String paneId = agent.paneId();
if (spawnReadyTimeoutMs > 0) {
waitUntilInjectableOrThrow(paneId);
}
// CB-519: the handle id is a host-unique UUID; the herdr pane it maps to stays internal.
String id = UUID.randomUUID().toString();
paneByAgentId.put(id, paneId);
return new WorkerHandle(id, agent.terminalId(), requireProfile(req.profileName()).profile(),
req.sessionName(), spawned.agentSessionId(), spawned.receipt());
}
@Override
public String effectiveCwd(SpawnRequest req) {
return effectiveCwd(req.profileName(), req.requestedCwd(), req.callerCwd());
}
/**
* CB-301: the effective working directory a spawn for {@code profileName} would use, without
* actually spawning.
*/
private String effectiveCwd(String profileName, String requestedCwd, String callerCwd) {
return resolveCwd(requestedCwd, requireProfile(profileName), callerCwd);
}
/**
* CB-112 cwd resolution: spawn arg → profile config → the primary's cwd → the daemon's cwd.
* Never returns {@code null}/blank: {@code "."} (the daemon's own working directory) is the
* guaranteed last resort so a pathological environment with an unset {@code user.dir} still
* honours the "never assume {@code $HOME}" contract rather than letting herdr default the pane.
*/
private static String resolveCwd(String requestedCwd, BridgedConfig.Profile cfg, String callerCwd) {
return firstNonBlank(requestedCwd, cfg.cwd(), callerCwd, System.getProperty("user.dir"), ".");
}
private static String firstNonBlank(String... values) {
for (String v : values) {
if (v != null && !v.isBlank()) return v;
}
return null;
}
/** Dedicated worker space → own tab (carrying cwd+env) → start the peer into the seed pane. */
private Agent spawnInTab(BridgedConfig.Profile cfg, Map<String, String> workerEnv,
List<String> argv, String cwd, MemberRole role, BridgedConfig.Fleet liveFleet) {
Workspace space = spaces.ensureWorkspace(cfg.workspace());
Tab.Created tab = spaces.createTab(space.workspaceId(), cwd, workerEnv);
log.info("spawning {} profile={} space={} tab={} cwd={}",
namePrefix, cfg.profile(), space.workspaceId(), tab.tab().tabId(), cwd);
Started started;
try {
if (tab.rootPaneId() == null) {
// Protocol 19 starts the agent INTO the seed pane — without one there is nowhere
// to start, and a partial tab would be left behind.
throw new IllegalStateException("tab " + tab.tab().tabId()
+ " had no seed pane in the create response — cannot start a peer in it");
}
started = startUniquelyNamed(cfg, argv, tab.rootPaneId());
} catch (RuntimeException e) {
// The peer never started — don't leave the tab we just created orphaned.
// Best-effort cleanup; never let it mask the real spawn failure.
try {
spaces.closeTab(tab.tab().tabId());
} catch (RuntimeException cleanup) {
log.warn("failed to close orphaned tab {} after spawn error: {}",
tab.tab().tabId(), cleanup.getMessage());
}
throw e;
}
// The peer is LIVE now, in the seed pane itself (no shell pane to drop — protocol 19).
// Labelling is cosmetic: it must not fail the spawn or orphan the running peer — on error
// we log and still return it so the caller gets its paneId and can tear it down.
tidy("label tab " + tab.tab().tabId(),
() -> spaces.renameTab(tab.tab().tabId(),
cfg.renderTabLabel(
liveFleet == null ? null : liveFleet.tabLabel(),
role, nextLabelSeq(role, cfg.profile()))));
log.info("{} started pane={} tab={} terminal={}",
namePrefix, started.agent().paneId(), started.agent().tabId(), started.agent().terminalId());
return started.agent();
}
/** Run a best-effort post-start cleanup step, logging (not throwing) on failure. */
private void tidy(String what, Runnable step) {
try {
step.run();
} catch (RuntimeException e) {
log.warn("post-start step failed ({}) — peer is running regardless: {}", what, e.getMessage());
}
}
/**
* Legacy placement: split the currently-focused tab; the peer still starts in {@code cwd}.
*
* <p>CB-571: this is the one legacy log that printed the full argv, and the charter travels
* inside argv — so the charter text went to the daemon log on every pane-placement spawn. The
* {@code spawnInTab} path never logs argv, so only this site is fixed. {@code charter} is the
* composed charter, if any; its argv element is replaced by its digest so the log still shows
* which args were passed without exposing the charter prose.
*/
private Agent spawnAsPane(BridgedConfig.Profile cfg, Map<String, String> workerEnv,
List<String> argv, String cwd, String charter) {
log.info("spawning {} (pane placement) profile={} cwd={} argv={}",
namePrefix, cfg.profile(), cwd, redactCharter(argv, charter));
String paneId = spaces.splitPane(cwd, workerEnv);
if (paneId == null) {
throw new IllegalStateException("pane.split returned no pane — cannot start a peer");
}
Agent peer = startUniquelyNamed(cfg, argv, paneId).agent();
log.info("{} started pane={} terminal={}", namePrefix, peer.paneId(), peer.terminalId());
return peer;
}
/**
* A copy of {@code argv} with an element equal to {@code charter} replaced by its digest, so
* the pane log never prints the charter prose. The charter is handed to an adapter as one argv
* element, so exact-equality is the right match; every other argument passes through unchanged.
*/
private static List<String> redactCharter(List<String> argv, String charter) {
if (charter == null || charter.isBlank() || argv == null || argv.isEmpty()) {
return argv;
}
String digest = CharterReceipt.digestOf(charter);
return argv.stream()
.map(a -> a.equals(charter) ? "<charter sha256=" + digest + ">" : a)
.toList();
}
/** A started peer together with the sequence its unique name/label used. */
private record Started(Agent agent, long seq) {
}
/**
* Start the peer under a unique herdr agent name. herdr requires each running agent's
* {@code name} to be distinct (a 2nd identical {@code name} fails {@code agent_name_taken}) —
* the exact case that makes multiple peers useful. The name is
* {@code <prefix>-<profile>-<nonce>-<seq>}: {@code seq} distinguishes peers within this process,
* and the per-process {@code nonce} keeps a fresh process (whose {@code seq} restarts at 0) from
* colliding with same-profile peers that outlived a restart. The retry is a belt-and-braces
* backstop for the astronomically unlikely nonce+seq clash; the name is a label only — herdr
* detects kind and status from terminal output, not from it.
*/
private Started startUniquelyNamed(BridgedConfig.Profile cfg, List<String> argv, String paneId) {
// Protocol 19 resolves the executable from the agent kind (== namePrefix here), so
// argv[0] — the configured executable — is dropped and only the extra args are passed.
List<String> args = argv.isEmpty() ? argv : argv.subList(1, argv.size());
HerdrException last = null;
for (int attempt = 0; attempt < NAME_RETRIES; attempt++) {
long seq = nameSeq.incrementAndGet();
String name = namePrefix + "-" + cfg.profile() + "-" + nameNonce + "-" + seq;
try {
return new Started(startAwaitingShellPrompt(name, args, paneId), seq);
} catch (HerdrException e) {
if (!"agent_name_taken".equals(e.code())) throw e;
log.debug("peer name '{}' taken, retrying", name);
last = e;
}
}
throw last;
}
/** Start the agent into {@code paneId}, waiting out the seed shell's boot with the sleeper. */
private Agent startAwaitingShellPrompt(String name, List<String> args, String paneId) {
HerdrException busy = null;
for (int attempt = 0; attempt < SHELL_READY_RETRIES; attempt++) {
try {
return agents.start(name, namePrefix, args, paneId);
} catch (HerdrException e) {
if (!"agent_pane_busy".equals(e.code())) throw e;
log.debug("pane {} not at its shell prompt yet, retrying agent.start", paneId);
busy = e;
sleeper.run();
}
}
throw busy;
}
// --- discovery + reap ----------------------------------------------------------------------
/** All herdr-tracked agents — discovery for "what peers exist". */
@Override
public List<Agent> list() {
return agents.list();
}
/**
* Reap peer panes left behind by an earlier daemon process (CB-117). herdr keeps a peer's pane
* alive across a daemon restart <em>by design</em>, and that pane's id is held only by its
* spawner — so a peer whose owning process exited before issuing the matching teardown leaks
* with nothing tracking it. On boot we scan herdr for agents whose name matches our
* {@code <prefix>-<profile>-<nonce>-<seq>} scheme with a nonce <em>other</em> than this
* process's {@link #nameNonce}, and tear each one down (its pane and, via {@link #stop}, its
* now-empty dedicated tab). A current-nonce peer is ours and live, so it is left running; a
* user's own session carries no such name and is never touched. A peer from a <em>different</em>
* adapter (different prefix) is likewise never touched. Best-effort: a failed listing, or a
* failure to stop any one peer, is logged and never aborts startup.
*
* @return the number of orphaned peers reaped
*/
@Override
public int reapOrphanWorkers() {
List<Agent> all;
try {
all = agents.list();
} catch (RuntimeException e) {
log.warn("orphan-peer reap skipped — agent.list failed: {}", e.getMessage());
return 0;
}
int reaped = 0;
for (Agent a : all) {
if (!isForeignWorker(namePrefix, a.name(), nameNonce)) continue;
try {
stop(a.paneId());
reaped++;
log.info("reaped orphan {} {} (pane={} tab={}) left by a prior daemon",
namePrefix, a.name(), a.paneId(), a.tabId());
} catch (RuntimeException e) {
log.warn("could not reap orphan {} {} (pane={}): {}",
namePrefix, a.name(), a.paneId(), e.getMessage());
}
}
if (reaped > 0) {
log.info("orphan-peer reap complete — {} stale {} peer(s) removed at startup", reaped, namePrefix);
}
return reaped;
}
/** The {@code <prefix>-<profile>-<nonce>-<seq>} name pattern; group 1 captures the 6-hex nonce. */
static Pattern workerNamePattern(String prefix) {
return Pattern.compile(prefix + "-.*-([0-9a-f]{6})-\\d+");
}
/**
* Whether {@code name} is a peer of kind {@code prefix} started by a <em>different</em> process
* than {@code currentNonce} — the reap predicate (CB-117). True only for the prefix's naming
* scheme with a foreign nonce: a non-peer name, a different adapter's name, or our own live
* nonce is excluded. Pure and package-private so the decision is unit-testable without herdr.
*/
static boolean isForeignWorker(String prefix, String name, String currentNonce) {
String nonce = workerNonce(prefix, name);
return nonce != null && !nonce.equals(currentNonce);
}
/** The 6-hex nonce embedded in a {@code prefix} peer name, or {@code null} if not one. */
static String workerNonce(String prefix, String name) {
if (name == null) return null;
Matcher m = workerNamePattern(prefix).matcher(name);
return m.matches() ? m.group(1) : null;
}
/** This process's peer-name nonce (a label component only; exposed for reaper tests). */
String nameNonce() {
return nameNonce;
}
// --- teardown ------------------------------------------------------------------------------
/**
* Tear a peer down: close the pane, and close its tab <em>only</em> when the peer is that tab's
* sole occupant. The single-pane check is what makes this safe regardless of how the peer was
* placed (or a placement-config change across a restart): a pane-placement peer sitting in one
* of the user's shared tabs has siblings, so its tab is never closed — we only ever remove a
* tab we created to hold one peer.
*
* <p>{@code idOrPane} is the {@link PeerHandle#id()} of a peer this launcher spawned (CB-519's
* host-unique opaque UUID), resolved through {@link #paneByAgentId} to the pane it must tear
* down. An argument that is not one of our ids is treated as a raw herdr pane id — the
* {@link #reapOrphanWorkers() orphan-reap} and spawn-gate-timeout paths, plus any caller that
* passes a pane directly, keep working without an owning id.
*
* <p>Resolves the tab from the pane <em>before</em> closing it. An already-gone pane/tab
* (repeated DELETE, crashed peer) is treated as success; any other failure propagates so a
* genuinely failed teardown is not reported as done.
*/
@Override
public void stop(String idOrPane) {
// Teardown knows only the pane, not which profile spawned it. Attempt tab cleanup when any
// profile uses tab placement (so the bridge may have created a dedicated peer tab); the
// single-occupant check below is what actually protects the user's shared tabs.
String paneId = paneByAgentId.remove(idOrPane);
if (paneId == null) {
paneId = idOrPane; // raw-pane fallback (reap, gate timeout, pane-addressed callers)
}
WorkspaceControl.PaneLocation loc = usesTabPlacement() ? spaces.locatePane(paneId) : null;
try {
agents.close(paneId);
} catch (HerdrException e) {
if (!isAlreadyGone(e)) throw e;
log.debug("pane.close({}) ignored — already gone: {}", paneId, e.getMessage());
}
if (loc != null && loc.tabPaneCount() == 1) {
spaces.closeTab(loc.tabId());
} else if (loc != null) {
log.debug("not closing tab {} — it holds {} panes (not a dedicated peer tab)",
loc.tabId(), loc.tabPaneCount());
}
}
/** Whether any configured profile places peers in their own tab (so tabs may need cleanup). */
private boolean usesTabPlacement() {
return profiles.values().stream().anyMatch(BridgedConfig.Profile::tabPlacement);
}
/** True when a herdr error means the target is already gone (safe to treat as done). */
private static boolean isAlreadyGone(HerdrException e) {
return e.code() != null && e.code().endsWith("_not_found");
}
// --- spawn-readiness gate (CB-306) ---------------------------------------------------------
/**
* Poll {@link AgentControl#status} until the pane reports an injectable state or the configured
* timeout elapses. On timeout, close the pane (self-reap) and throw.
*/
private void waitUntilInjectableOrThrow(String paneId) {
long deadline = nowMillis.getAsLong() + spawnReadyTimeoutMs;
while (nowMillis.getAsLong() < deadline) {
if (agents.status(paneId).injectable()) {
log.debug("peer pane={} reached injectable state", paneId);
return;
}
sleeper.run();
}
log.warn("peer pane={} did not become injectable within {}ms — closing", paneId, spawnReadyTimeoutMs);
stop(paneId);
throw new PeerUnreachableException(
"worker pane " + paneId + " did not reach injectable state within "
+ spawnReadyTimeoutMs + "ms");
}
/**
* A concrete {@link PeerHandle} wrapping herdr agent coordinates, the profile that spawned it,
* the session identity the launch resolved (CB-547a): the bridge's logical name and the peer's
* own session id, both null when the spawn carried no identity — and the charter receipt
* (CB-571) the base computed for this launch.
*/
private record WorkerHandle(String id, String terminalId, String profile,
String sessionName, String agentSessionId,
CharterReceipt receipt) implements PeerHandle {
@Override
public CharterReceipt charterReceipt() {
return receipt;
}
}
// --- shared helpers ------------------------------------------------------------------------
/** Put {@code k → v} only when {@code v} is present (non-null, non-blank). */
protected static void putIfPresent(Map<String, String> m, String k, String v) {
if (v != null && !v.isBlank()) {
m.put(k, v);
}
}
/** Host env lookup that tolerates an unconfigured (null/blank) var name — returns null then. */
protected String resolveEnv(String name) {
return (name == null || name.isBlank()) ? null : env.apply(name);
}
/**
* The parity-neutral git-forge token grant (CB-302): when {@code cfg} opts in via
* {@code gitTokenEnv} and the token resolves, inject {@code GITEA_TOKEN} plus its paired
* {@code GITEA_HOST}. Push over SSH is unaffected; the only incremental grant is PR-create.
* Peer-neutral, so every herdr adapter reuses it unchanged.
*/
protected void applyGitToken(Map<String, String> workerEnv, BridgedConfig.Profile cfg) {
if (!cfg.hasGitToken()) {
return;
}
String gitToken = resolveEnv(cfg.gitTokenEnv());
if (gitToken != null) {
workerEnv.put("GITEA_TOKEN", gitToken);
putIfPresent(workerEnv, "GITEA_HOST", resolveEnv(cfg.gitHostEnv()));
}
}
/**
* CB-592: overlay value that shadows the admin {@code GITEA_ACCESS_TOKEN} a herdr pane
* otherwise inherits from herdr's own login-shell process environment (gitea issue #77).
* herdr spawns a pane from its <em>own</em> process environment and layers our map on top —
* {@link dev.ltms.bridged.herdr.WorkspaceControl#createTab} and {@code #splitPane} send only
* the keys we put in that map, so any key we never mention passes straight through from
* herdr's own shell, admin token included.
*
* <p>Deliberately a non-blank sentinel, not {@code ""}. Whether an empty-string overlay value
* overrides an inherited variable or is skipped as blank could not be settled by reading this
* codebase — herdr's server-side merge is an external process, not something in this repo.
* A non-blank replacement sidesteps that ambiguity entirely: {@link #baseEnv}'s own {@code
* PATH} seeding already depends on the overlay reliably replacing an inherited value (see its
* javadoc), and that is only demonstrated for a non-blank value, so this reuses the same,
* proven-reliable shape rather than the unverified one.
*
* <p><b>MEASURED ON A LIVE PANE, 2026-08-15: this sentinel alone does NOT hold.</b> The overlay
* itself works — {@code GITEA_TOKEN} is injected here, is exported by no shell file, and does
* reach the pane. The sentinel loses one step later. A herdr pane runs a <em>login</em> shell,
* {@code ~/.zprofile} sources {@code ${SHARED_ENV}/tools/secrets.sh}, and that file does a plain
* unconditional {@code export GITEA_ACCESS_TOKEN=...}. A login shell overwrites a value already
* in the environment, so the real admin token is put back over this sentinel before the member
* process ever starts. That defeat applies to <em>every</em> name {@code secrets.sh} exports,
* and no launcher-side overlay can win against it.
*
* <p>So this constant is not the control on its own — {@link #MEMBER_MARKER} is the other half.
* Keeping the sentinel is still worth it: it is correct for any peer kind whose pane does not
* start a login shell, and it makes the intent explicit at the one place every adapter passes.
*/
private static final String BLOCKED_GITEA_ACCESS_TOKEN =
"blocked-by-bridged-cb592-see-gitea-issue-77";
/**
* CB-592: marks a pane as a bridged member so a shell startup file can decline to export
* operator-only credentials into it (gitea issue #77).
*
* <p>This name is deliberately one that {@code secrets.sh} never exports, which is exactly why
* it survives the login shell that wipes {@link #BLOCKED_GITEA_ACCESS_TOKEN}. The mechanism is
* measured, not assumed: {@code GITEA_TOKEN} is injected the same way, is absent from a login
* shell of its own, and was observed set inside a live member pane.
*
* <p>It is a no-op until the operator guards the export, which is a one-line change in a file
* this repo does not own and must not edit unasked:
*
* <pre>{@code
* [ -n "${BRIDGED_MEMBER:-}" ] || export GITEA_ACCESS_TOKEN=...
* }</pre>
*
* <p>Setting the marker now costs nothing and means that edit is the whole remaining fix.
*/
static final String MEMBER_MARKER = "BRIDGED_MEMBER";
/**
* Seed a worker's environment (CB-511): the daemon's own {@code PATH}, then the profile's
* {@code env:} entries, then the CB-592 admin-token shadow.
*
* <p>Why this exists: bridged passes herdr an explicit env map, and herdr merges it into
* <em>its own</em> process environment. So before this, a worker inherited whatever PATH the
* herdr server happened to be started with — on this host, one from weeks earlier with no JDK
* and no Maven, which left workers unable to run the build they were being asked to run. The
* worker's toolchain must follow from configuration, not from how a long-lived daemon was
* launched.
*
* <p>Adapter-specific variables are layered on top of this by {@code buildLaunch} and therefore
* win. That ordering is deliberate and load-bearing: it stops a profile's {@code env:} from
* overriding {@code ANTHROPIC_BASE_URL} and slipping past {@link
* dev.ltms.bridged.guard.SubscriptionGuard}, which is checked against the profile's
* {@code baseUrl} and nothing else.
*
* <p>The CB-592 shadow and marker are put in <em>last</em>, after the profile's own
* {@code env:}, so no profile — present or future — can restore the admin token, or hide that
* the pane is a member, by naming either in config. This is the one place both are applied:
* every {@code buildLaunch} in every adapter calls this first, so a new profile, and a peer
* kind not yet written, gets them for free.
*/
protected Map<String, String> baseEnv(BridgedConfig.Profile cfg) {
Map<String, String> workerEnv = new LinkedHashMap<>();
String path = env.apply("PATH");
if (path != null && !path.isBlank()) {
workerEnv.put("PATH", path);
}
if (cfg != null && cfg.env() != null) {
workerEnv.putAll(cfg.env());
}
workerEnv.put("GITEA_ACCESS_TOKEN", BLOCKED_GITEA_ACCESS_TOKEN);
workerEnv.put(MEMBER_MARKER, "1");
return workerEnv;
}
/** Defensive copy of {@code argv} plus room to append launch flags. */
protected static List<String> mutableArgv(List<String> argv) {
return new ArrayList<>(argv);
}
/**
* Uninterruptible sleep — the production {@link #sleeper}. Tests supply their own no-op /
* fast-faking sleeper so they never real-sleep.
*/
protected static void sleepUninterruptibly(long ms) {
try {
Thread.sleep(ms);
} catch (InterruptedException e) {
Thread.currentThread().interrupt();
// preserve the interrupt flag but continue — poll loops should not be aborted by an
// interrupt that was not meant for them.
}
}
}
@@ -1,4 +1,4 @@
package dev.ltms.bridged.worker;
package dev.ltms.bridged.member;
import com.fasterxml.jackson.databind.ObjectMapper;
import com.fasterxml.jackson.databind.node.ObjectNode;
@@ -7,6 +7,9 @@ import dev.ltms.bridged.herdr.Agent;
import dev.ltms.bridged.herdr.AgentControl;
import dev.ltms.bridged.herdr.WorkspaceControl;
import dev.ltms.bridged.peer.Capability;
import dev.ltms.bridged.peer.CharterReceipt;
import dev.ltms.bridged.peer.PeerHandle;
import dev.ltms.bridged.peer.SpawnRequest;
import java.io.IOException;
import java.io.UncheckedIOException;
@@ -18,6 +21,7 @@ import java.util.Map;
import java.util.Set;
import java.util.function.Function;
import java.util.function.LongSupplier;
import java.util.function.Supplier;
/**
* The {@link HerdrPeerLauncher} adapter for <strong>opencode</strong> — an open-source,
@@ -35,7 +39,7 @@ import java.util.function.LongSupplier;
* <li><strong>File-based MCP mount + instructions.</strong> opencode has no inline
* {@code --mcp-config}/{@code --append-system-prompt}. Instead the bridge writes an ephemeral
* {@code opencode.json} that declares the bridge as a {@code remote} MCP server and lists a
* reply-charter file under {@code instructions}, then points the worker at it with
* member-charter file under {@code instructions}, then points the worker at it with
* {@code OPENCODE_CONFIG}. This is the one place the launcher touches disk — Claude never did.</li>
* <li><strong>Model as a flag.</strong> the {@code provider/model} selector is passed as
* {@code -m}, not an env var.</li>
@@ -51,37 +55,26 @@ public final class OpenCodeLauncher extends HerdrPeerLauncher {
/** Writer for the generated {@code opencode.json}. */
private static final ObjectMapper JSON = new ObjectMapper();
/**
* Standing instruction written to the charter file and mounted via the config's
* {@code instructions} so the worker returns its result through {@code bridge_reply}. Kept on
* disk (not a launch flag) because opencode's {@code instructions} takes file paths, not inline
* text — the file is regenerated per spawn and never touches the worker's own profile.
*/
static final String REPLY_CHARTER =
"You are an off-subscription worker in the claude-bridge fleet, running under opencode. "
+ "Every message you receive arrives through the bridge, and the ONLY channel back to the "
+ "sender is the bridge_reply MCP tool. Text you write in your terminal is NOT sent "
+ "anywhere — the sender cannot see your screen, so an in-terminal answer is silently "
+ "discarded. Therefore you MUST end EVERY turn by calling bridge_reply with `content` set "
+ "to your complete response. This holds for every message without exception — tasks, "
+ "questions, clarifications, acknowledgements, and ordinary back-and-forth conversation. "
+ "Call bridge_reply exactly once, as the final action of your turn, with your full answer "
+ "in `content`; never wait for confirmation first. If you end a turn without calling "
+ "bridge_reply, the sender receives nothing and the exchange stalls.";
/** Root under which per-spawn opencode config dirs are created (injectable for tests). */
private final Path configRoot;
/**
* Session discovery against opencode's on-disk storage ({@link OpenCodeSessionDiscovery}) —
* the one seam that knows opencode's private session-file layout. Its root is injectable for
* tests so they never touch the operator's real {@code ~/.local/share/opencode}.
*/
private final OpenCodeSessionDiscovery discovery;
/**
* Production constructor — disables the spawn-ready gate ({@code spawnReadyTimeoutMs == 0}) so it
* matches the legacy non-blocking spawn semantics. Config dirs are created under the JVM temp dir.
*/
public OpenCodeLauncher(AgentControl agents, WorkspaceControl spaces,
Map<String, BridgedConfig.Worker> profiles, String defaultProfile,
Map<String, BridgedConfig.Profile> profiles, String defaultProfile,
Function<String, String> env) {
this(agents, spaces, profiles, defaultProfile, env, 0,
System::currentTimeMillis, () -> sleepUninterruptibly(300),
defaultConfigRoot());
defaultConfigRoot(), defaultDiscoveryRoot());
}
/**
@@ -89,12 +82,26 @@ public final class OpenCodeLauncher extends HerdrPeerLauncher {
* the pane reports an injectable state or {@code spawnReadyTimeoutMs} elapses.
*/
public OpenCodeLauncher(AgentControl agents, WorkspaceControl spaces,
Map<String, BridgedConfig.Worker> profiles, String defaultProfile,
Map<String, BridgedConfig.Profile> profiles, String defaultProfile,
Function<String, String> env,
long spawnReadyTimeoutMs, long spawnReadyPollMs) {
this(agents, spaces, profiles, defaultProfile, env,
spawnReadyTimeoutMs, spawnReadyPollMs, null);
}
/**
* Production constructor carrying the fleet-wide tab-label template (CB-557). The template comes
* from {@code fleet.tabLabel}, which a profile cannot know because it names the member's
* <em>role</em>; a profile may still override it with its own {@code tabLabel}.
*/
public OpenCodeLauncher(AgentControl agents, WorkspaceControl spaces,
Map<String, BridgedConfig.Profile> profiles, String defaultProfile,
Function<String, String> env,
long spawnReadyTimeoutMs, long spawnReadyPollMs,
Supplier<BridgedConfig.Fleet> fleet) {
this(agents, spaces, profiles, defaultProfile, env, spawnReadyTimeoutMs,
System::currentTimeMillis, () -> sleepUninterruptibly(spawnReadyPollMs),
defaultConfigRoot());
defaultConfigRoot(), defaultDiscoveryRoot(), fleet);
}
/**
@@ -112,39 +119,66 @@ public final class OpenCodeLauncher extends HerdrPeerLauncher {
* @param sleeper sleep/wait hook (encodes the poll interval; never called when the
* gate is disabled)
* @param configRoot existing directory under which per-spawn config dirs are created
* @param discoveryRoot opencode's on-disk storage root to scan for session records
* (injectable for tests; opencode's layout is matched at
* {@link OpenCodeSessionDiscovery})
*/
public OpenCodeLauncher(AgentControl agents, WorkspaceControl spaces,
Map<String, BridgedConfig.Worker> profiles, String defaultProfile,
Map<String, BridgedConfig.Profile> profiles, String defaultProfile,
Function<String, String> env,
long spawnReadyTimeoutMs,
LongSupplier nowMillis, Runnable sleeper, Path configRoot) {
LongSupplier nowMillis, Runnable sleeper,
Path configRoot, Path discoveryRoot) {
this(agents, spaces, profiles, defaultProfile, env, spawnReadyTimeoutMs,
nowMillis, sleeper, configRoot, discoveryRoot, null);
}
/**
* Full testability constructor, plus the fleet-wide tab-label template (CB-557).
*
* @param fleet live fleet config, read once for each spawn
*/
public OpenCodeLauncher(AgentControl agents, WorkspaceControl spaces,
Map<String, BridgedConfig.Profile> profiles, String defaultProfile,
Function<String, String> env,
long spawnReadyTimeoutMs,
LongSupplier nowMillis, Runnable sleeper,
Path configRoot, Path discoveryRoot,
Supplier<BridgedConfig.Fleet> fleet) {
super(NAME_PREFIX, agents, spaces, profiles, defaultProfile, env,
spawnReadyTimeoutMs, nowMillis, sleeper);
spawnReadyTimeoutMs, nowMillis, sleeper, fleet);
this.configRoot = configRoot;
this.discovery = new OpenCodeSessionDiscovery(discoveryRoot);
}
private static Path defaultConfigRoot() {
return Path.of(System.getProperty("java.io.tmpdir"));
}
/** The default opencode storage root: {@code ~/.local/share/opencode} (the XDG data dir). */
private static Path defaultDiscoveryRoot() {
return Path.of(System.getProperty("user.home"), ".local", "share", "opencode");
}
/**
* {@inheritDoc}
*
* <p>Builds the opencode launch: no {@code ANTHROPIC_*} and no guard (opencode reads its own
* provider credentials); when the profile mounts the bridge MCP, generate an ephemeral
* {@code opencode.json} (remote MCP server + reply-charter instructions) and point the worker at
* it via {@code OPENCODE_CONFIG}; carry the parity-neutral git-forge grant; and select the model
* with {@code -m}.
* provider credentials); when the profile mounts the bridge MCP or has a member charter,
* generate an ephemeral {@code opencode.json} (remote MCP server + member-charter instructions)
* and point the worker at it via {@code OPENCODE_CONFIG}; carry the parity-neutral git-forge
* grant; and select the model with {@code -m}.
*/
@Override
protected Launch buildLaunch(BridgedConfig.Worker cfg) {
protected Launch buildLaunch(BridgedConfig.Profile cfg, LaunchSpec spec) {
Map<String, String> workerEnv = baseEnv(cfg);
// A config file is needed for the bridge MCP mount, for a pinned endpoint (CB-508), or both.
if (cfg.hasMcp() || hasCustomProvider(cfg)) {
workerEnv.put("OPENCODE_CONFIG", writeConfig(cfg).toString());
// A config file is needed for the bridge MCP mount, a member charter, or a pinned endpoint (CB-508).
if (cfg.hasMcp() || spec.charter() != null || hasCustomProvider(cfg)) {
workerEnv.put("OPENCODE_CONFIG", writeConfig(cfg, spec.charter()).toString());
}
applyGitToken(workerEnv, cfg);
return new Launch(workerEnv, argvWithModel(cfg));
return new Launch(workerEnv,
argvWithResume(argvWithModel(argvWithAuto(cfg), cfg), spec.resumeSessionId()));
}
/**
@@ -157,13 +191,44 @@ public final class OpenCodeLauncher extends HerdrPeerLauncher {
* primary's Anthropic subscription, and an opencode process has no Anthropic credential path
* at all. Pointing it at a local vLLM cannot leak the subscription.
*/
private static boolean hasCustomProvider(BridgedConfig.Worker cfg) {
private static boolean hasCustomProvider(BridgedConfig.Profile cfg) {
return cfg.baseUrl() != null && !cfg.baseUrl().isBlank();
}
/** The launch argv plus, when a model is configured, the opencode {@code -m provider/model} flag. */
private List<String> argvWithModel(BridgedConfig.Worker cfg) {
/**
* The launch argv plus the unconditional {@code --auto} flag, which auto-approves the
* permissions opencode does not explicitly deny. It is unconditional, not a preference: a
* spawned peer has no human at its pane — the bridge spawned it — so one that stops at an
* approval prompt is a wedged agent, indistinguishable from a legitimate mid-turn wait and
* unable to end its turn with {@code bridge_reply}. opencode's own help calls this
* "dangerous!", but the blast radius here is already bounded by design: a worker runs in its
* own git worktree on its own branch, is off-subscription, and cannot merge — the lead is the
* gate.
*/
private List<String> argvWithAuto(BridgedConfig.Profile cfg) {
List<String> argv = mutableArgv(cfg.argv());
argv.add("--auto");
return argv;
}
/**
* The launch argv plus, on a resumed spawn, opencode's {@code -s <id>} flag to continue a prior
* conversation by its session id. {@code -s, --session <id>} resumes an existing session; on a
* fresh spawn (no resume target) no flag is added, letting opencode start a brand-new session.
* The id comes from the base launch spec.
*/
private List<String> argvWithResume(List<String> argv, String id) {
if (id == null || id.isBlank()) {
return argv;
}
List<String> withResume = mutableArgv(argv);
withResume.add("-s");
withResume.add(id);
return withResume;
}
/** The launch argv plus, when a model is configured, the opencode {@code -m provider/model} flag. */
private List<String> argvWithModel(List<String> argv, BridgedConfig.Profile cfg) {
if (cfg.model() != null && !cfg.model().isBlank()) {
argv.add("-m");
argv.add(cfg.model());
@@ -172,29 +237,47 @@ public final class OpenCodeLauncher extends HerdrPeerLauncher {
}
/**
* Write an ephemeral {@code opencode.json} (and the reply-charter file it references) into a
* Write an ephemeral {@code opencode.json} (and the member-charter file it references) into a
* fresh per-spawn directory under {@link #configRoot}, and return the config file's path for
* {@code OPENCODE_CONFIG}. The dir is unique per spawn so concurrent workers never race on it;
* it is best-effort cleaned on JVM exit (worker config is disposable — regenerated every spawn).
*/
private Path writeConfig(BridgedConfig.Worker cfg) {
private Path writeConfig(BridgedConfig.Profile cfg, String charterText) {
try {
Path dir = Files.createTempDirectory(configRoot, "bridged-opencode-");
dir.toFile().deleteOnExit();
ObjectNode root = JSON.createObjectNode();
root.put("$schema", "https://opencode.ai/config.json");
// CB-523, opencode side: a worker that runs out of context dies mid-turn, and its reply
// — the entire point of the turn — is lost with it. Auto-compaction is therefore not an
// operator preference for a bridged worker, it is a condition of the turn contract.
//
// Stated deliberately even though it is redundant today: OPENCODE_CONFIG is MERGED over
// ~/.config/opencode/config.json rather than replacing it, so a worker already inherits
// an `auto: true` set at home. We do not want that inheritance to be what the guarantee
// rests on — the home file is outside this repo, differs per machine, and is not ours.
//
// Know the cost before removing it: this key WINS over the home config (verified — an
// OPENCODE_CONFIG value overrides the home value, it does not defer to it), so an
// operator who sets `compaction.auto: false` at home cannot turn it off for bridged
// workers. That is the intended trade for peers we spawn and whose turns we must land;
// if per-profile control is ever wanted, add a profile knob rather than dropping this.
root.putObject("compaction").put("auto", true);
if (cfg.hasMcp()) {
Path charter = dir.resolve("reply-charter.md");
Files.writeString(charter, REPLY_CHARTER);
if (charterText != null) {
Path charter = dir.resolve("member-charter.md");
Files.writeString(charter, charterText);
charter.toFile().deleteOnExit();
root.putArray("instructions").add(charter.toAbsolutePath().toString());
}
if (cfg.hasMcp()) {
ObjectNode bridge = root.putObject("mcp").putObject("bridge");
bridge.put("type", "remote");
bridge.put("url", cfg.mcpUrl());
bridge.put("enabled", true);
root.putArray("instructions").add(charter.toAbsolutePath().toString());
}
if (hasCustomProvider(cfg)) {
addCustomProvider(root, cfg);
@@ -219,7 +302,7 @@ public final class OpenCodeLauncher extends HerdrPeerLauncher {
* <p>The provider id comes from the {@code provider/model} selector in {@code model:}, so one
* field drives both the declaration and the {@code -m} flag and they cannot drift apart.
*/
private void addCustomProvider(ObjectNode root, BridgedConfig.Worker cfg) {
private void addCustomProvider(ObjectNode root, BridgedConfig.Profile cfg) {
String[] parts = splitModelSelector(cfg);
String providerId = parts[0];
String modelId = parts[1];
@@ -242,7 +325,7 @@ public final class OpenCodeLauncher extends HerdrPeerLauncher {
* needs both, so a bare model name is rejected loudly rather than silently falling back to the
* default gateway — a worker quietly talking to the wrong endpoint is the failure this avoids.
*/
private static String[] splitModelSelector(BridgedConfig.Worker cfg) {
private static String[] splitModelSelector(BridgedConfig.Profile cfg) {
String model = cfg.model();
int slash = model == null ? -1 : model.indexOf('/');
if (model == null || model.isBlank() || slash <= 0 || slash == model.length() - 1) {
@@ -270,6 +353,67 @@ public final class OpenCodeLauncher extends HerdrPeerLauncher {
return afterScheme.contains("/") ? trimmed : trimmed + "/v1";
}
/** Add lazy on-disk session discovery to the base handle. */
@Override
public PeerHandle spawn(SpawnRequest req) {
PeerHandle inner = super.spawn(req);
return new SessionAwareHandle(inner, discovery, effectiveCwd(req));
}
/**
* A {@link PeerHandle} that delegates everything to the base's worker handle but resolves
* {@link #agentSessionId()} lazily through opencode session discovery. Delegate-only, so the
* base's id/terminalId/profile semantics (CB-519's host-unique routing key, herdr coordinates)
* are untouched — only the opencode-specific identity answer is added. {@code sessionName()}
* stays null: opencode has no display-name seam, so the logical name lives only in the bridge's
* roster (see the SESSION_NAME capability).
*/
private static final class SessionAwareHandle implements PeerHandle {
private final PeerHandle delegate;
private final OpenCodeSessionDiscovery discovery;
private final String cwd;
SessionAwareHandle(PeerHandle delegate, OpenCodeSessionDiscovery discovery, String cwd) {
this.delegate = delegate;
this.discovery = discovery;
this.cwd = cwd;
}
@Override
public String id() {
return delegate.id();
}
@Override
public String terminalId() {
return delegate.terminalId();
}
@Override
public String profile() {
return delegate.profile();
}
@Override
public String sessionName() {
return delegate.sessionName();
}
@Override
public String agentSessionId() {
// Lazy + retried, never a spawn-time blocker: opencode writes the session record only
// when the session is first persisted, so null here is the correct interim answer and
// the caller re-calls later (each call re-scans, picking up a record that has since
// appeared).
return discovery.sessionIdForDirectory(cwd);
}
@Override
public CharterReceipt charterReceipt() {
return delegate.charterReceipt();
}
}
// --- Agent-returning convenience spawns (used by callers/tests that want the herdr Agent) ---
/** Spawn a worker for the default profile in the resolved default cwd. */
@@ -291,16 +435,21 @@ public final class OpenCodeLauncher extends HerdrPeerLauncher {
@Override
public Set<Capability> capabilities() {
Set<Capability> caps = EnumSet.of(Capability.MID_TURN_ASK, Capability.WORKTREE, Capability.ORPHAN_REAP);
Set<Capability> caps = EnumSet.of(Capability.MID_TURN_ASK, Capability.WORKTREE,
Capability.ORPHAN_REAP, Capability.SESSION_RESUME);
if (hasGitTokenProfile()) {
caps.add(Capability.SELF_PR);
}
// Deliberately NOT SESSION_NAME: opencode has no display-name flag, so the bridge's logical
// name can't surface in the peer's own UI — declaring the capability would hide that
// asymmetry rather than make it honest. For opencode the name lives only in the bridge's
// roster (see PeerHandle.sessionName() returning null).
return Set.copyOf(caps);
}
/** Whether any configured profile opts into a git-forge token (required for {@link Capability#SELF_PR}). */
private boolean hasGitTokenProfile() {
return profileConfigs().stream().anyMatch(BridgedConfig.Worker::hasGitToken);
return profileConfigs().stream().anyMatch(BridgedConfig.Profile::hasGitToken);
}
// --- CB-117 reap predicate (opencode prefix), kept for direct unit testing -----------------
@@ -0,0 +1,122 @@
package dev.ltms.bridged.member;
import com.fasterxml.jackson.databind.JsonNode;
import com.fasterxml.jackson.databind.ObjectMapper;
import java.io.IOException;
import java.nio.file.Files;
import java.nio.file.Path;
import java.util.stream.Stream;
/**
* Resolves the opencode session id for a bridged worker from opencode's on-disk storage — the
* only place this adapter touches opencode's private layout, and deliberately the <em>only</em>
* class that does.
*
* <p><strong>Why this is isolated behind one seam.</strong> The layout is version-coupled and not a
* stable contract: opencode writes one JSON file per session under
* {@code <storageRoot>/session/<projectID>/<ses_*.json>}, and each record carries a
* {@code "version"} field (e.g. {@code "1.1.31"}), so the exact directory shape, file naming, and
* field names can move between opencode releases. opencode also ships a headless HTTP server that
* may supersede file scanning entirely. Everything this adapter knows about that private storage —
* its shape, naming, and field names — lives here, so a layout change, or a switch to the HTTP
* server, changes exactly one class and nothing in {@link OpenCodeLauncher}.
*
* <p>The determinism that makes this useful is structural, not a guess: every bridged worker runs
* in its own unique git worktree, so the record's {@code directory} (its project root) equals the
* worker's cwd identifies <em>its</em> session unambiguously. We match on {@code directory} rather
* than diffing {@code opencode session list} before/after — that races under concurrent spawns, and
* the CLI listing does not even show the directory.
*
* <p>All reads are best-effort and never throw: a missing or unreadable storage root, a record that
* fails to parse, or a directory with no record yet all yield {@code null}, and the caller (the
* session handle) treats that as "identity not resolved yet" and retries later.
*/
final class OpenCodeSessionDiscovery {
private final Path storageRoot; // e.g. ~/.local/share/opencode (injectable for tests)
private final ObjectMapper json;
OpenCodeSessionDiscovery(Path storageRoot) {
this.storageRoot = storageRoot;
this.json = new ObjectMapper();
}
/**
* The opencode session id whose record references {@code directory} (the worker's cwd), or
* {@code null} when no record matches yet. When several records share the directory — e.g.
* repeated spawns into the same worktree — the <em>most recently modified</em> one wins: it is
* the session the pane most likely corresponds to.
*
* <p>Never throws: a missing {@code storageRoot}, an unreadable/malformed record, or a
* directory that has not been persisted yet all resolve to {@code null} rather than failing a
* spawn. A bridged worker's session record is written lazily (when the session is first
* persisted), so {@code null} here is the normal answer right after the pane is ready, and the
* caller retries later.
*
* @param directory the worker's cwd, as resolved for this spawn
* @return the matching session id, or {@code null} if none is known yet
*/
String sessionIdForDirectory(String directory) {
if (directory == null || directory.isBlank()) {
return null;
}
Path sessionRoot = storageRoot.resolve("session");
if (!Files.isDirectory(sessionRoot)) {
return null;
}
String best = null;
long bestMtime = Long.MIN_VALUE;
try (Stream<Path> projectDirs = Files.list(sessionRoot)) {
for (Path projectDir : projectDirs.filter(Files::isDirectory).toList()) {
try (Stream<Path> records = Files.list(projectDir)) {
for (Path record : records.toList()) {
String id = matchId(record, directory);
if (id == null) {
continue;
}
long mtime = lastModifiedEpochMillis(record);
if (mtime > bestMtime) {
bestMtime = mtime;
best = id;
}
}
} catch (IOException ignored) {
// one project dir unreadable — skip it; another may still match
}
}
} catch (IOException ignored) {
// storage root vanished or became unreadable — "no session known yet"
return null;
}
return best;
}
/**
* The record's session id when it references {@code directory}, else {@code null}. A record
* that is not JSON, lacks {@code id}/{@code directory}, or points at a different directory is
* simply not our session; a malformed one is skipped, never fatal.
*/
private String matchId(Path record, String directory) {
try {
JsonNode node = json.readTree(record.toFile());
JsonNode id = node == null ? null : node.get("id");
JsonNode dir = node == null ? null : node.get("directory");
if (id == null || dir == null || !directory.equals(dir.asText())) {
return null;
}
return id.asText();
} catch (IOException e) {
return null;
}
}
/** The record's last-modified epoch ms, or {@code Long.MIN_VALUE} if unreadable (never wins). */
private static long lastModifiedEpochMillis(Path record) {
try {
return Files.getLastModifiedTime(record).toMillis();
} catch (IOException e) {
return Long.MIN_VALUE;
}
}
}
@@ -2,7 +2,7 @@ package dev.ltms.bridged.metrics;
import dev.ltms.bridged.msg.ReplyInbox;
import dev.ltms.bridged.session.SessionManager;
import dev.ltms.bridged.session.WorkerSession;
import dev.ltms.bridged.session.MemberSession;
import java.util.LinkedHashMap;
import java.util.Map;
@@ -26,6 +26,8 @@ public final class BridgedMetrics {
public static final String REPLIES = "bridged_replies_total";
/** Counter: push-loop nudges to the primary, by outcome. */
public static final String PUSH_NUDGES = "bridged_push_nudges_total";
/** Counter: idle-lead heartbeat nudges to the lead, by outcome (CB-551). */
public static final String HEARTBEAT_NUDGES = "bridged_lead_heartbeat_nudges_total";
/** Counter: spawn attempts by peer kind and outcome. */
public static final String SPAWNS = "bridged_spawns_total";
/** Counter: herdr socket calls by method and outcome. */
@@ -55,6 +57,9 @@ public final class BridgedMetrics {
"Worker replies by delivery path (rendezvous=resolved an open send, inbox=stranded and held).");
m.describe(PUSH_NUDGES, "counter",
"CB-307 push-loop nudges to the primary (delivered|exhausted).");
m.describe(HEARTBEAT_NUDGES, "counter",
"CB-551 idle-lead heartbeat nudges (delivered|failed|exhausted). Quiet-cap exhaustion "
+ "means the lead idled with nothing pending and was told to stand down.");
m.describe(SPAWNS, "counter",
"Worker spawn attempts by peer kind and outcome (ready|timeout|guard_rejected).");
m.describe(HERDR_CALLS, "counter",
@@ -69,7 +74,7 @@ public final class BridgedMetrics {
// One gauge per state so a scrape shows the whole census even when a state is empty —
// an absent series and a zero series read very differently on a dashboard.
for (WorkerSession.State state : WorkerSession.State.values()) {
for (MemberSession.State state : MemberSession.State.values()) {
String label = state.name().toLowerCase();
m.gauge(SESSIONS, () -> countIn(sessions, state), "state", label);
}
@@ -78,7 +83,7 @@ public final class BridgedMetrics {
// port's non-destructive read — scraping metrics must never ack a reply out of the inbox.
m.collector(INBOX_DEPTH, "target", () -> {
Map<String, Number> depths = new LinkedHashMap<>();
for (WorkerSession s : sessions.roster()) {
for (MemberSession s : sessions.roster()) {
String target = s.terminalId();
if (target == null) {
continue;
@@ -90,7 +95,7 @@ public final class BridgedMetrics {
return m;
}
private static long countIn(SessionManager sessions, WorkerSession.State state) {
private static long countIn(SessionManager sessions, MemberSession.State state) {
return sessions.roster().stream().filter(s -> s.state() == state).count();
}
}
@@ -7,6 +7,7 @@ import com.rabbitmq.client.ConnectionFactory;
import com.rabbitmq.client.DeliverCallback;
import com.rabbitmq.client.Recoverable;
import com.rabbitmq.client.RecoveryListener;
import com.rabbitmq.client.Return;
import org.slf4j.Logger;
import org.slf4j.LoggerFactory;
@@ -14,25 +15,57 @@ import java.io.IOException;
import java.nio.charset.StandardCharsets;
import java.util.LinkedHashMap;
import java.util.List;
import java.util.Set;
import java.util.NavigableMap;
import java.util.concurrent.CompletableFuture;
import java.util.concurrent.ConcurrentHashMap;
import java.util.concurrent.ConcurrentSkipListMap;
import java.util.concurrent.ExecutionException;
import java.util.concurrent.TimeUnit;
import java.util.concurrent.TimeoutException;
/**
* AMQP-backed {@link ReplyInbox} (CB-307 Stage 2): genuine cross-restart durability behind the same
* port {@link InMemoryReplyInbox} implements as soft state.
*
* <p><strong>Mapping — consume-and-hold with deferred manual ack.</strong> Each target owns a durable
* queue {@code agent.<target>.inbox}. A manual-ack consumer pulls persistent messages off that queue
* into an in-memory <em>held</em> map (keyed by {@code msgId}) but does <em>not</em> ack them.
* {@link #peek} returns that snapshot; {@link #ack} acks the broker delivery-tag and drops the entry.
* Because messages stay unacked until the primary actually drains them, a crash (or a {@code java -jar}
* bounce) before caller-ack leaves them on the broker — it redelivers on reconnect. That is the
* durability the in-memory adapter cannot give, with the port contract preserved.
* <p><strong>Mapping — consume-and-hold with deferred manual ack.</strong> Each target has a durable
* queue {@code agent.<target>.inbox}. The gateway that owns the target starts a manual-ack consumer
* ({@link #own}) that pulls persistent messages off that queue into an in-memory <em>held</em> map
* (keyed by {@code msgId}) but does <em>not</em> ack them. {@link #peek} returns that snapshot;
* {@link #ack} acks the broker delivery-tag and drops the entry. Because messages stay unacked until
* the owning gateway actually drains them, a crash (or a {@code java -jar} bounce) before caller-ack
* leaves them on the broker — it redelivers on reconnect. That is the durability the in-memory
* adapter cannot give, with the port contract preserved.
*
* <p><strong>Prefetch bounds the held backlog (CB-527).</strong> The consumer channel calls
* {@code basicQos} with a configurable prefetch count ({@link #DEFAULT_PREFETCH} unless the caller
* passes another value to {@link #open(String, int)}) before starting any consumer. Without a bound,
* the broker pushes its entire queue into {@link #held} the instant a target is {@link #own owned},
* so an undrained primary grows the JVM heap without limit and any queue-level control
* ({@code x-max-length}, per-message TTL) never fires because the queue never actually holds a
* backlog. Prefetch keeps the backlog where it is visible — on the broker — until the owner drains it.
*
* <p><strong>Publishes require a confirmed, routable delivery (CB-528).</strong> {@link #publish}
* runs on a channel separate from the consume/ack channel ({@link #channel}), so a slow or blocked
* publish confirm can never hold {@link #channelLock} and stall an ack — the ack path never waits on
* a publish confirm. That publish channel is in publisher-confirm mode and every publish sets the
* {@code mandatory} flag, so an unroutable publish (queue not declared, e.g. the owner never called
* {@link #own}) is returned by the broker instead of silently dropped. The broker sends the
* <em>return</em> for an unroutable message before the <em>confirm</em> that covers it — the ack/nack
* callback checks the returned-set at confirm time rather than assuming an ack means routed — so
* "confirmed" here means "durably queued", not merely "accepted by the broker". A returned or nacked
* (or un-confirmed within the timeout) publish surfaces as an {@link IllegalStateException} on the
* caller's thread; the caller — {@link MessageService#reply} — must not report success for a
* black-holed reply.
*
* <p><strong>Ownership is explicit.</strong> {@link #own} declares the queue and starts the consumer;
* {@link #release} cancels it. {@link #publish} sends to the queue but does <em>not</em> imply ownership
* and does not attach a consumer. This split is required by CB-308 federation, where one gateway may
* publish to an agent owned by another gateway; in that case the publisher must not compete for
* deliveries.
*
* <p><strong>Dedup.</strong> The consumer keys the held map by {@code msgId}; a redelivered duplicate
* (at-least-once, or a producer double-publish) is acked-and-dropped on arrival, so it never
* double-queues. {@link #publish} additionally short-circuits an already-held {@code msgId} — a
* fast path; the consumer-side check is the real guarantee.
* double-queues.
*
* <p><strong>Visibility.</strong> Unlike the in-memory adapter, publish → broker → consumer is
* asynchronous, so a {@link #peek} immediately after {@link #publish} may not yet see the message
@@ -49,48 +82,102 @@ public final class AmqpReplyInbox implements ReplyInbox, AutoCloseable {
private static final String QUEUE_PREFIX = "agent.";
private static final String QUEUE_SUFFIX = ".inbox";
/** CB-527: the prefetch used when a caller does not pass an explicit value to {@link #open(String, int)}. */
public static final int DEFAULT_PREFETCH = 32;
/** How long {@link #publish} waits for its publisher confirm before failing the call (CB-528). */
private static final long CONFIRM_TIMEOUT_MS = 10_000L;
private final Connection connection;
private final Channel channel;
/** All channel operations (publish/declare/ack) serialize on this — a Channel is not thread-safe. */
/** All channel operations (publish/declare/ack/cancel) serialize on this — a Channel is not thread-safe. */
private final Object channelLock = new Object();
/** target → (msgId → held delivery). Per-target map is guarded by synchronizing on itself. */
private final ConcurrentHashMap<String, LinkedHashMap<String, Held>> held = new ConcurrentHashMap<>();
/** Targets whose queue is declared and consumer is running. */
private final Set<String> consuming = ConcurrentHashMap.newKeySet();
/** Targets whose queue is declared and consumer is running, mapped to their broker consumer tag. */
private final ConcurrentHashMap<String, String> consumerTags = new ConcurrentHashMap<>();
/**
* CB-528: a dedicated channel for {@link #publish}, kept separate from {@link #channel} (consume
* + ack) so a publish confirm round trip never blocks under {@link #channelLock} and stalls an ack.
*/
private final Channel publishChannel;
private final Object publishChannelLock = new Object();
/** In-flight publishes awaiting their confirm, keyed by the publish channel's sequence number. */
private final ConcurrentSkipListMap<Long, Pending> pendingBySeq = new ConcurrentSkipListMap<>();
/**
* The same in-flight publishes, keyed by {@code msgId} — a broker {@code Return} carries no delivery
* tag. Assumes {@code msgId} is unique per in-flight publish: a second {@link #publish} for a
* {@code msgId} still awaiting its confirm would overwrite this entry and misdirect
* {@link #onReturn}'s lookup. Not reachable today — {@code MessageService.reply} generates a fresh
* {@code UUID} per call — so no guard is added for it.
*/
private final ConcurrentHashMap<String, Pending> pendingByMsgId = new ConcurrentHashMap<>();
/** A message pulled off the broker but not yet acked: its delivery-tag plus the port payload. */
private record Held(long deliveryTag, InboxMessage message) {}
/** Connect to {@code uri} (e.g. {@code amqp://guest:guest@127.0.0.1:5672/}) and open the inbox. */
/** A publish awaiting its confirm; {@link #returned} records whether the broker already returned it. */
private static final class Pending {
final String msgId;
final CompletableFuture<Void> confirmed = new CompletableFuture<>();
volatile boolean returned;
Pending(String msgId) {
this.msgId = msgId;
}
}
/** Connect to {@code uri} (e.g. {@code amqp://guest:guest@127.0.0.1:5672/}) with {@link #DEFAULT_PREFETCH}. */
public static AmqpReplyInbox open(String uri) {
return open(uri, DEFAULT_PREFETCH);
}
/** As {@link #open(String)}, with an explicit consumer prefetch (CB-527: caps the held backlog per target). */
public static AmqpReplyInbox open(String uri, int prefetch) {
try {
ConnectionFactory factory = new ConnectionFactory();
factory.setUri(uri);
// Self-heal transient blips; topology recovery re-declares queues and re-attaches consumers.
factory.setAutomaticRecoveryEnabled(true);
factory.setTopologyRecoveryEnabled(true);
return new AmqpReplyInbox(factory.newConnection("bridged-reply-inbox"));
return new AmqpReplyInbox(factory.newConnection("bridged-reply-inbox"), prefetch);
} catch (Exception e) {
throw new IllegalStateException("cannot connect to AMQP broker at " + uri, e);
}
}
/** Wrap an already-open connection (injection seam for the contract test). */
/** Wrap an already-open connection with {@link #DEFAULT_PREFETCH} (injection seam for the contract test). */
AmqpReplyInbox(Connection connection) {
this(connection, DEFAULT_PREFETCH);
}
/** As above, with an explicit prefetch (injection seam for the contract test). */
AmqpReplyInbox(Connection connection, int prefetch) {
this.connection = connection;
try {
this.channel = connection.createChannel();
// CB-527: bound the held backlog per owned target — must be set before any own()/basicConsume.
this.channel.basicQos(prefetch);
this.publishChannel = connection.createChannel();
this.publishChannel.confirmSelect();
this.publishChannel.addReturnListener(this::onReturn);
this.publishChannel.addConfirmListener(this::onAck, this::onNack);
} catch (IOException e) {
throw new IllegalStateException("cannot open AMQP channel", e);
}
// On automatic recovery the broker redelivers unacked messages with FRESH delivery-tags; the
// tags we were holding are now stale. Drop the held snapshot so the re-attached consumer
// repopulates it with valid tags (dedup by msgId still prevents any double-queue).
// repopulates it with valid tags (dedup by msgId still prevents any double-queue). Any publish
// confirm still in flight when the connection dropped is equally stale — its sequence number
// meant nothing on the old channel and means nothing on the recovered one, so fail it now
// rather than let it silently ride out CONFIRM_TIMEOUT_MS.
if (connection instanceof Recoverable recoverable) {
recoverable.addRecoveryListener(new RecoveryListener() {
@Override
public void handleRecovery(Recoverable recoverable) {
held.clear();
failPendingPublishesOnRecovery();
log.info("AMQP connection recovered; cleared held replies for fresh redelivery");
}
@@ -103,33 +190,86 @@ public final class AmqpReplyInbox implements ReplyInbox, AutoCloseable {
}
@Override
public void publish(String target, String msgId, String content) {
ensureConsuming(target);
var perTarget = held.get(target);
if (perTarget != null) {
synchronized (perTarget) {
if (perTarget.containsKey(msgId)) {
return; // already held — producer-side fast dedup
}
public void own(String target) {
synchronized (channelLock) {
if (consumerTags.containsKey(target)) {
return; // already owning this target
}
String queue = queueName(target);
try {
channel.queueDeclare(queue, true, false, false, null); // durable, non-exclusive, keep on idle
String tag = channel.basicConsume(queue, false, deliverCallback(target), _ -> { });
consumerTags.put(target, tag);
log.debug("AMQP inbox owns queue {} for target {}", queue, target);
} catch (IOException e) {
throw new IllegalStateException("cannot own queue " + queue, e);
}
}
}
@Override
public void release(String target) {
synchronized (channelLock) {
String tag = consumerTags.remove(target);
held.remove(target); // stale delivery tags must not survive release
if (tag == null) {
return;
}
try {
channel.basicCancel(tag);
} catch (IOException e) {
throw new IllegalStateException("cannot cancel consumer for " + target, e);
}
}
}
/**
* Publish {@code content} and block until the broker's publisher confirm for it lands (CB-528).
* Throws {@link IllegalStateException} if the message is returned as unroutable, nacked, or not
* confirmed within {@link #CONFIRM_TIMEOUT_MS} — the caller must treat that as a failed publish,
* not a lost-and-forgotten one.
*/
@Override
public void publish(String target, String msgId, String content) {
AMQP.BasicProperties props = new AMQP.BasicProperties.Builder()
.messageId(msgId)
.deliveryMode(2) // persistent — survives a broker restart
.contentType("text/plain")
.build();
try {
synchronized (channelLock) {
channel.basicPublish("", queueName(target), props, content.getBytes(StandardCharsets.UTF_8));
Pending pending = new Pending(msgId);
long seq;
synchronized (publishChannelLock) {
seq = publishChannel.getNextPublishSeqNo();
pendingBySeq.put(seq, pending);
pendingByMsgId.put(msgId, pending);
try {
publishChannel.basicPublish("", queueName(target), true, props,
content.getBytes(StandardCharsets.UTF_8));
} catch (IOException e) {
pendingBySeq.remove(seq, pending);
pendingByMsgId.remove(msgId, pending);
throw new IllegalStateException("cannot publish reply to " + queueName(target), e);
}
} catch (IOException e) {
throw new IllegalStateException("cannot publish reply to " + queueName(target), e);
}
try {
pending.confirmed.get(CONFIRM_TIMEOUT_MS, TimeUnit.MILLISECONDS);
} catch (ExecutionException e) {
Throwable cause = e.getCause();
throw cause instanceof RuntimeException re ? re : new IllegalStateException(cause);
} catch (TimeoutException e) {
throw new IllegalStateException("publish confirm for reply " + msgId + " to " + queueName(target)
+ " timed out after " + CONFIRM_TIMEOUT_MS + "ms — broker may be unreachable or overloaded", e);
} catch (InterruptedException e) {
Thread.currentThread().interrupt();
throw new IllegalStateException("interrupted awaiting publish confirm for " + msgId, e);
} finally {
pendingBySeq.remove(seq, pending);
pendingByMsgId.remove(msgId, pending);
}
}
@Override
public List<InboxMessage> peek(String target) {
ensureConsuming(target);
var perTarget = held.get(target);
if (perTarget == null) {
return List.of();
@@ -166,26 +306,6 @@ public final class AmqpReplyInbox implements ReplyInbox, AutoCloseable {
}
}
/** Declare the durable per-target queue and start its manual-ack consumer, once per target. */
private void ensureConsuming(String target) {
if (consuming.contains(target)) {
return;
}
synchronized (channelLock) {
if (!consuming.add(target)) {
return; // another thread just set it up
}
String queue = queueName(target);
try {
channel.queueDeclare(queue, true, false, false, null); // durable, non-exclusive, keep on idle
channel.basicConsume(queue, false, deliverCallback(target), _ -> { });
} catch (IOException e) {
consuming.remove(target);
throw new IllegalStateException("cannot consume queue " + queue, e);
}
}
}
private DeliverCallback deliverCallback(String target) {
return (_, delivery) -> {
String msgId = delivery.getProperties().getMessageId();
@@ -213,17 +333,115 @@ public final class AmqpReplyInbox implements ReplyInbox, AutoCloseable {
};
}
/** Broker return for an unroutable {@code mandatory} publish — arrives BEFORE its confirm (CB-528). */
private void onReturn(Return r) {
String msgId = r.getProperties() == null ? null : r.getProperties().getMessageId();
Pending pending = msgId == null ? null : pendingByMsgId.get(msgId);
if (pending != null) {
pending.returned = true;
} else {
log.warn("AMQP return for reply {} (routingKey={}, {} {}) with no matching in-flight publish"
+ " — already resolved by a prior confirm", msgId, r.getRoutingKey(), r.getReplyCode(),
r.getReplyText());
}
}
private void onAck(long seq, boolean multiple) {
resolveConfirm(seq, multiple, true);
}
private void onNack(long seq, boolean multiple) {
resolveConfirm(seq, multiple, false);
}
/**
* Resolve every pending publish covered by this confirm (a single seq, or — {@code multiple} —
* every seq up to and including it). Checks {@link Pending#returned} at confirm time: since the
* broker's return for an unroutable message always precedes its confirm, an ack that arrives after
* a return means "confirmed but never routed", not "durably queued".
*/
private void resolveConfirm(long seq, boolean multiple, boolean ack) {
NavigableMap<Long, Pending> covered = multiple
? pendingBySeq.headMap(seq, true)
: pendingBySeq.subMap(seq, true, seq, true);
for (var it = covered.entrySet().iterator(); it.hasNext(); ) {
Pending pending = it.next().getValue();
it.remove();
pendingByMsgId.remove(pending.msgId, pending);
if (ack && !pending.returned) {
pending.confirmed.complete(null);
} else if (ack) {
pending.confirmed.completeExceptionally(new IllegalStateException(
"reply " + pending.msgId + " was returned as unroutable (queue not declared/owned)"));
} else {
pending.confirmed.completeExceptionally(new IllegalStateException(
"broker nacked publish of reply " + pending.msgId));
}
}
}
/**
* Fail every publish still awaiting its confirm — their sequence numbers are stale after recovery.
* Guarded by {@link #publishChannelLock}, the same lock {@link #publish} holds while it takes its
* sequence number and registers its {@link Pending}: without it, a {@link #publish} that starts
* after the connection has already recovered (so it publishes — and will be confirmed — on the
* <em>new</em> channel) can register between this sweep's iteration and its clear, and this sweep
* then fails a publish that actually succeeded. {@link #publish} only holds the lock for the
* seq/map-put/{@code basicPublish} — it awaits the confirm outside it — so this sweep can only ever
* wait for an in-flight {@code basicPublish} call to return, never for a broker round trip. No
* deadlock.
*
* <p>Package-private (rather than {@code private}) only so the unit test can drive it directly
* against a concurrent {@link #publish} without a live broker reconnect.
*/
void failPendingPublishesOnRecovery() {
synchronized (publishChannelLock) {
for (var it = pendingBySeq.entrySet().iterator(); it.hasNext(); ) {
Pending pending = it.next().getValue();
it.remove();
pendingByMsgId.remove(pending.msgId, pending);
pending.confirmed.completeExceptionally(new IllegalStateException(
"AMQP connection recovered mid-publish; confirm status of reply " + pending.msgId
+ " is unknown"));
}
}
}
/**
* Fail every publish still awaiting its confirm with a clear, immediate error instead of leaving it
* to time out after {@link #CONFIRM_TIMEOUT_MS} once the channels are closed underneath it. Guarded
* by {@link #publishChannelLock} for the same reason as {@link #failPendingPublishesOnRecovery}.
*/
private void failPendingPublishesOnClose() {
synchronized (publishChannelLock) {
for (var it = pendingBySeq.entrySet().iterator(); it.hasNext(); ) {
Pending pending = it.next().getValue();
it.remove();
pendingByMsgId.remove(pending.msgId, pending);
pending.confirmed.completeExceptionally(new IllegalStateException(
"AMQP reply inbox closed while publish of reply " + pending.msgId
+ " was still awaiting its confirm"));
}
}
}
private static String queueName(String target) {
return QUEUE_PREFIX + target + QUEUE_SUFFIX;
}
@Override
public void close() {
failPendingPublishesOnClose();
try {
channel.close();
} catch (Exception e) {
log.debug("AMQP channel close: {}", e.toString());
}
try {
publishChannel.close();
} catch (Exception e) {
log.debug("AMQP publish channel close: {}", e.toString());
}
try {
connection.close();
} catch (Exception e) {
@@ -2,6 +2,7 @@ package dev.ltms.bridged.msg;
import java.util.LinkedHashMap;
import java.util.List;
import java.util.Set;
import java.util.concurrent.ConcurrentHashMap;
/**
@@ -9,12 +10,29 @@ import java.util.concurrent.ConcurrentHashMap;
* Per-target FIFO ordering (insertion order via {@link LinkedHashMap}). Dedup by {@code msgId}
* within a target. Thread-safe for concurrent publish vs. drain.
*
* <p><strong>Ownership is explicit.</strong> {@link #own} marks a target as locally owned so that
* {@link #peek} and {@link #ack} operate on it; {@link #publish} works whether or not the target is
* owned. {@link #release} clears the local snapshot. This mirrors the AMQP adapter's contract so the
* non-broker path stays interchangeable.
*
* <p><strong>This is soft-state, NOT persistence.</strong> Lost on a {@code java -jar} bounce — that
* is correct and consistent with "bridged stays soft-state." The Stage-2 AMQP adapter replaces this.
*/
public final class InMemoryReplyInbox implements ReplyInbox {
private final ConcurrentHashMap<String, LinkedHashMap<String, InboxMessage>> store = new ConcurrentHashMap<>();
private final Set<String> owned = ConcurrentHashMap.newKeySet();
@Override
public void own(String target) {
owned.add(target);
}
@Override
public void release(String target) {
owned.remove(target);
store.remove(target);
}
@Override
public void publish(String target, String msgId, String content) {
@@ -27,6 +45,9 @@ public final class InMemoryReplyInbox implements ReplyInbox {
@Override
public List<InboxMessage> peek(String target) {
if (!owned.contains(target)) {
return List.of();
}
var perTarget = store.get(target);
if (perTarget == null) {
return List.of();
@@ -39,6 +60,9 @@ public final class InMemoryReplyInbox implements ReplyInbox {
@Override
public void ack(String target, String msgId) {
if (!owned.contains(target)) {
return;
}
var perTarget = store.get(target);
if (perTarget != null) {
//noinspection SynchronizationOnLocalVariableOrMethodParameter
@@ -0,0 +1,351 @@
package dev.ltms.bridged.msg;
import dev.ltms.bridged.herdr.AgentControl;
import dev.ltms.bridged.herdr.AgentStatus;
import dev.ltms.bridged.mcp.PrimaryRegistry;
import dev.ltms.bridged.metrics.BridgedMetrics;
import dev.ltms.bridged.metrics.Metrics;
import dev.ltms.bridged.session.MemberSession;
import org.slf4j.Logger;
import org.slf4j.LoggerFactory;
import java.util.ArrayList;
import java.util.List;
import java.util.concurrent.ScheduledExecutorService;
import java.util.concurrent.TimeUnit;
import java.util.function.LongSupplier;
import java.util.function.Supplier;
/**
* CB-551: an opt-in heartbeat that nudges the single idle lead back to work once it has been
* continuously idle past a quiet period with no open {@code bridge_send} driving it.
*
* <p>Why this exists: the fleet is ONE lead + architects + workers, so an idle, stalled lead is a
* single point of failure for the fleet's progress. {@link ReplyPushLoop} nudges the lead only when
* a worker reply lands; this loop is the timer that catches the gap where nothing lands and the
* lead simply sits idle with nobody to prompt it onward.
*
* <p>Mechanism mirrors {@link ReplyPushLoop}: a pure, unit-testable decision function
* ({@link #decide(long, Long, int, AgentStatus, boolean, boolean, FleetState)}) plus a thin
* scheduler around it. The loop is <em>entirely opt-in</em> ({@code leadHeartbeat:} in config); with
* no such block it is never constructed, so upgrading the daemon cannot silently acquire a behaviour
* that spends the operator's model subscription on its own initiative (constraint 1).
*
* <p>Four invariants keep it from becoming a runaway subscription burner:
* <ol>
* <li><b>Status-gated</b> — a {@code WORKING} lead is making progress and is never touched; only an
* injectable (idle/done/blocked) lead is even considered (constraint 2).</li>
* <li><b>Debounced</b> — a nudge only fires once the lead has been <em>continuously</em> idle past
* {@code idleAfterSeconds}, so a lead that just finished a turn is not re-prompted into every
* natural pause (constraint 3).</li>
* <li><b>Bounded when quiet</b> — consecutive nudges that find no pending fleet state are capped at
* {@code quietNudgeCap}, then the loop stops nudging until real state returns (constraint 4).</li>
* <li><b>Never races {@link ReplyPushLoop}</b> — while that loop is actively nudging any target this
* loop stands down, so two competing injections never start two turns in the same pane
* (constraint 6).</li>
* </ol>
*/
public final class LeadHeartbeatLoop {
private static final Logger log = LoggerFactory.getLogger(LeadHeartbeatLoop.class);
/** Sentinel for {@link #idleSinceNanos}: the lead is not currently in an idle stretch. */
private static final long NOT_IDLE = Long.MIN_VALUE;
private final PrimaryRegistry primaryRegistry;
private final AgentControl agents;
private final ReplyInbox inbox;
private final Supplier<List<MemberSession>> roster;
private final ReplyPushLoop pushLoop;
private final ScheduledExecutorService scheduler;
private final LongSupplier clock;
private final long idleAfterNanos;
private final long backoffMs;
private final int quietNudgeCap;
private final Metrics metrics; // CB-512 pattern: nullable — no registry in unit tests
/** When the current idle stretch began (nanos), or {@link #NOT_IDLE}. Single scheduler thread only. */
private long idleSinceNanos = NOT_IDLE;
/** Consecutive nudges that found no pending fleet state. Single scheduler thread only. */
private int quietCount = 0;
/** Constructor with an injectable clock and no metric registry (unit tests, or wiring that opts out). */
public LeadHeartbeatLoop(PrimaryRegistry primaryRegistry, AgentControl agents, ReplyInbox inbox,
Supplier<List<MemberSession>> roster, ReplyPushLoop pushLoop,
ScheduledExecutorService scheduler, LongSupplier clock,
long idleAfterNanos, long backoffMs, int quietNudgeCap) {
this(primaryRegistry, agents, inbox, roster, pushLoop, scheduler, clock,
idleAfterNanos, backoffMs, quietNudgeCap, null);
}
/** As above, with a metric registry (the CB-512 pattern) so nudge outcomes are counted. */
public LeadHeartbeatLoop(PrimaryRegistry primaryRegistry, AgentControl agents, ReplyInbox inbox,
Supplier<List<MemberSession>> roster, ReplyPushLoop pushLoop,
ScheduledExecutorService scheduler, LongSupplier clock,
long idleAfterNanos, long backoffMs, int quietNudgeCap, Metrics metrics) {
this.primaryRegistry = primaryRegistry;
this.agents = agents;
this.inbox = inbox;
this.roster = roster;
this.pushLoop = pushLoop;
this.scheduler = scheduler;
this.clock = clock;
this.idleAfterNanos = idleAfterNanos;
this.backoffMs = backoffMs;
this.quietNudgeCap = quietNudgeCap;
this.metrics = metrics;
}
/** Count one nudge outcome when a registry is wired; a no-op in unit tests. */
private void countNudge(String outcome) {
if (metrics != null) {
metrics.inc(BridgedMetrics.HEARTBEAT_NUDGES, "outcome", outcome);
}
}
// --- decision logic (package-private for unit-testing) --------------------------------------
/** The action the loop should take on one tick. */
enum Action {
/** Send a heartbeat nudge to the lead now. */
INJECT,
/** Lead is idle but not yet past the quiet period — keep waiting, do nothing. */
WAIT_IDLE,
/** Lead is making progress (or its status cannot be read) — reset the idle window and re-arm. */
LEAD_BUSY,
/** Idle past the quiet period with nothing pending and the cap exhausted — stop until new state appears. */
QUIET_DONE,
/** {@link ReplyPushLoop} is actively nudging — stand aside rather than start a competing turn. */
STAND_DOWN
}
/** The outcome of one decision: the action plus the state to persist for the next tick. */
record Decision(Action action, Long idleSinceNanos, int quietCount) {}
/**
* Pure decision function: given the current loop state and fleet/lead facts, return what to do
* next and the exact state to carry forward. Pure means no I/O and no mutation — the caller
* ({@link #tick}) applies {@link Decision} to its fields. This is what makes the loop unit-testable
* with a fake clock and no sleeping.
*
* @param nowNanos the current time (an injected clock in tests, {@link System#nanoTime()} live)
* @param idleSinceNanos when the current idle stretch began, or {@code null} if the lead is not idle
* @param quietCount consecutive nudges so far that found no pending fleet state
* @param status the lead's herdr status
* @param pushLoopActive whether {@link ReplyPushLoop} is currently nudging some target (constraint 6)
* @param leadKnown whether a lead terminal is known to nudge at all
* @param fleet a snapshot of the pending fleet state (constraint 5)
* @return the action to take and the state to persist
*/
Decision decide(long nowNanos, Long idleSinceNanos, int quietCount, AgentStatus status,
boolean pushLoopActive, boolean leadKnown, FleetState fleet) {
// Constraint 6: while ReplyPushLoop is actively nudging the lead, injecting a second,
// competing prompt into the same pane would start a second turn — racing loops multiply
// turns and context burn. Stand aside, and treat the active push as real state (re-arm the
// quiet counter), because the reply that drove it is exactly the kind of new state that
// should reset the cap.
if (pushLoopActive) {
return new Decision(Action.STAND_DOWN, idleSinceNanos, 0);
}
// Constraint 2: a WORKING lead is making progress and must NOT be touched; an unreadable
// status (read failure, or the agent is gone) is safest treated the same way — never inject
// into a state we cannot read. Either way, reset the idle window and the quiet counter: the
// lead was / may be active, so the next idle stretch must count its own quiet period fresh.
if (status == null || !status.injectable()) {
return new Decision(Action.LEAD_BUSY, null, 0);
}
if (idleSinceNanos == null) {
// The lead just became injectable — record the start of an idle stretch and wait out the
// debounce quiet period before ever nudging (constraint 3).
return new Decision(Action.WAIT_IDLE, nowNanos, quietCount);
}
if (nowNanos - idleSinceNanos < idleAfterNanos) {
// Still within the quiet period: the lead that just finished a turn sits momentarily idle
// and must not be re-prompted into every natural pause.
return new Decision(Action.WAIT_IDLE, idleSinceNanos, quietCount);
}
// Past the quiet period with an injectable lead: it is a genuine candidate for a nudge. Two
// gating facts decide whether and how:
if (!leadKnown) {
// No lead terminal is known yet (e.g. the registry has not learned one) — there is nobody
// to nudge. Keep waiting; the window stays open so discovery re-arms it without a fresh
// quiet period.
return new Decision(Action.WAIT_IDLE, idleSinceNanos, quietCount);
}
if (fleet.hasPending()) {
// Real fleet state is waiting — a worker reply or a DONE session. This is new state, so
// it resets the quiet counter (constraint 4) and the lead is nudged to go collect it.
return new Decision(Action.INJECT, idleSinceNanos, 0);
}
if (quietCount < quietNudgeCap) {
// Nothing is pending, but the cap is not exhausted: nudge anyway, telling the lead
// exactly that nothing is waiting so it can choose to stand down rather than hunt
// (constraint 5). Count it toward the consecutive-quiet cap.
return new Decision(Action.INJECT, idleSinceNanos, quietCount + 1);
}
// Nothing pending and the cap is exhausted: stop nudging until real state appears again
// (constraint 4). The loop still ticks on backoff so a genuinely new reply or session change
// re-arms it — QUIET_DONE stops injection, not observation.
return new Decision(Action.QUIET_DONE, idleSinceNanos, quietCount);
}
// --- loop ----------------------------------------------------------------------------------
/**
* Start the heartbeat. Schedules the first tick after one backoff so a freshly-booting daemon
* does not evaluate the lead's idle state before the fleet has settled.
*/
public void start() {
log.info("idle-lead heartbeat: on — nudge lead after {}s idle (recheck {}ms, quiet cap {})",
TimeUnit.NANOSECONDS.toSeconds(idleAfterNanos), backoffMs, quietNudgeCap);
scheduler.schedule(this::tick, backoffMs, TimeUnit.MILLISECONDS);
}
/** One loop tick, every {@link #backoffMs} — the thin scheduler around {@link #decide}. */
private void tick() {
boolean leadKnown = primaryRegistry.primaryTerminal().isPresent();
FleetState fleet = snapshot(inbox, roster);
AgentStatus status = AgentStatus.UNKNOWN;
if (leadKnown) {
try {
status = agents.status(primaryRegistry.primaryTerminal().orElseThrow());
} catch (RuntimeException e) {
// A failed status read degrades to "unknown" — decide() treats that like a busy lead
// and never injects into a state it cannot read. Retry on the next backoff.
log.debug("idle-heartbeat: status check failed for lead, will retry: {}", e.toString());
}
}
Decision d = decide(clock.getAsLong(),
idleSinceNanos == NOT_IDLE ? null : idleSinceNanos,
quietCount, status, pushLoop.isActive(), leadKnown, fleet);
applyDecision(d);
switch (d.action()) {
case INJECT -> injectNudge(fleet);
case QUIET_DONE -> countNudge("exhausted");
case WAIT_IDLE, LEAD_BUSY, STAND_DOWN -> { /* nothing to inject, nothing to count */ }
}
scheduleNext();
}
/** Persist the state a decision returned, so the next tick starts from it. */
private void applyDecision(Decision d) {
idleSinceNanos = d.idleSinceNanos() == null ? NOT_IDLE : d.idleSinceNanos();
quietCount = d.quietCount();
}
/** Send the nudge to the known lead. */
private void injectNudge(FleetState fleet) {
var lead = primaryRegistry.primaryTerminal();
if (lead.isEmpty()) {
return; // the lead disappeared between the decision and the injection
}
String leadTerminal = lead.get();
try {
agents.send(leadTerminal, fleet.nudgeText());
log.debug("idle-heartbeat: nudge sent to lead {} (quiet nudges so far in this stretch: {})",
leadTerminal, quietCount);
countNudge("delivered");
} catch (RuntimeException e) {
log.warn("idle-heartbeat: failed to nudge lead {}: {}", leadTerminal, e.toString());
countNudge("failed");
}
}
/** Schedule the next tick on the scheduler thread pool. */
private void scheduleNext() {
scheduler.schedule(this::tick, backoffMs, TimeUnit.MILLISECONDS);
}
// --- lifecycle -----------------------------------------------------------------------------
/** Shut down the scheduler; outstanding ticks are cancelled. */
public void stop() {
scheduler.shutdownNow();
}
/** @see #stop() */
public void close() {
stop();
}
/**
* A read-only snapshot of the fleet facts the heartbeat reports to an idle lead. Carries exact
* counts and names the targets that hold replies, so the nudge is actionable rather than a
* content-free "keep working" that would cost the lead a turn just to discover what there is to
* do (constraint 5).
*
* @param pendingReplies total undrained worker replies across owned targets
* @param doneSessions session count in the DONE state — finished turns awaiting teardown
* @param liveWorkers sessions still registered (acquired minus released)
* @param replyTargets the terminals whose inboxes currently hold at least one reply
*/
record FleetState(int pendingReplies, int doneSessions, int liveWorkers, List<String> replyTargets) {
/** True when the lead has something real to collect — a reply or a session awaiting teardown. */
boolean hasPending() {
return pendingReplies > 0 || doneSessions > 0;
}
/** The nudge body, phrased for the two cases the heartbeat distinguishes. */
String nudgeText() {
StringBuilder sb = new StringBuilder(
"Heartbeat: you are idle and no bridge_send is waiting on you.");
if (hasPending()) {
sb.append(" The fleet has state to collect: ").append(pendingDetail());
} else {
sb.append(" Fleet is quiet — nothing pending to collect (")
.append(pendingReplies).append(" worker replies, ")
.append(doneSessions).append(" DONE sessions); ")
.append(liveWorkers).append(" workers live. You may stand down until new state re-arms me.");
}
return sb.toString();
}
private String pendingDetail() {
StringBuilder sb = new StringBuilder();
sb.append(pendingReplies).append(" worker ")
.append(pendingReplies == 1 ? "reply" : "replies")
.append(" pending collection");
if (!replyTargets.isEmpty()) {
// Render each as the exact command so the lead can act without parsing: the nearest
// analogue to ReplyPushLoop's bridge_poll(target=...) nudge.
sb.append(" (")
.append(String.join(", ",
replyTargets.stream().map(t -> "bridge_poll(target=" + t + ")").toList()))
.append(")");
}
sb.append(", ").append(doneSessions).append(" DONE session").append(doneSessions == 1 ? "" : "s")
.append(" awaiting teardown")
.append(", ").append(liveWorkers).append(" worker").append(liveWorkers == 1 ? "" : "s")
.append(" live");
return sb.toString();
}
}
/**
* Build a {@link FleetState} snapshot over the live roster and the reply inbox — the same
* sources the metrics collector uses, so the heartbeat reports facts the operator can cross-check
* against {@code /metrics}. Package-private and pure (reads only) so it is unit-testable.
*/
static FleetState snapshot(ReplyInbox inbox, Supplier<List<MemberSession>> roster) {
List<MemberSession> sessions = roster.get();
int replies = 0;
int done = 0;
List<String> replyTargets = new ArrayList<>();
for (MemberSession s : sessions) {
if (s.terminalId() == null) {
continue;
}
int depth = inbox.peek(s.terminalId()).size();
if (depth > 0) {
replies += depth;
replyTargets.add(s.terminalId());
}
if (s.state() == MemberSession.State.DONE) {
done++;
}
}
return new FleetState(replies, done, sessions.size(), List.copyOf(replyTargets));
}
}
@@ -20,6 +20,7 @@ import java.util.concurrent.TimeUnit;
import java.util.concurrent.TimeoutException;
import java.util.concurrent.atomic.AtomicLong;
import java.util.concurrent.locks.ReentrantLock;
import java.util.function.LongSupplier;
/**
* The blocking delegation feature (CB-104): deliver {@code content} into a worker and block until
@@ -52,8 +53,12 @@ public final class MessageService {
*/
private static final long ASYNC_TIMEOUT_MS = 30 * 60 * 1_000L;
/** How long a finished (terminal) ticket is retained for polling before it is pruned. */
private static final long TICKET_TTL_NANOS = 10 * 60 * 1_000_000_000L;
/**
* How long a finished (terminal) ticket is retained for polling before it is pruned. Package-
* private (not {@code private}) so a test can advance an injected clock past it deterministically
* instead of duplicating the magic number or sleeping for real.
*/
static final long TICKET_TTL_NANOS = 10 * 60 * 1_000_000_000L;
/** Outcome of a blocking send. */
public enum Outcome {
@@ -69,6 +74,14 @@ public final class MessageService {
* failure context (e.g. the error screen). Terminal, but not a successful completion.
*/
WORKER_FAILED,
/**
* The turn finished without a {@code bridge_reply} and the scrape matched the backend's
* configured usage-limit refusal pattern (CB-578 stage A); {@code text} is the reason,
* carrying the matched line. The worker's pane is healthy — only its account is refusing —
* so this is never reported as a completed reply, and is kept distinct from
* {@link #WORKER_FAILED} (a wedged worker) and a session simply going {@code GONE}.
*/
BACKEND_EXHAUSTED,
/**
* The worker paused mid-turn to ask the primary a question (CB-205); {@code text} is the
* question and {@code turnId} correlates the answer. Not terminal — the primary answers with
@@ -127,6 +140,8 @@ public final class MessageService {
public enum Phase {
/** Delegated and in flight — queued for the worker or being worked. */
PENDING,
/** The worker is paused in {@code bridge_ask}; {@link TaskView#reply} and {@link TaskView#turnId} identify it. */
ASKING,
/** The worker's turn finished; {@link TaskView#reply} holds the answer. */
DONE,
/** The delegation could not complete (timed out, worker gone, or busy). */
@@ -136,16 +151,29 @@ public final class MessageService {
/**
* A poll snapshot of an async delegation.
*
* @param reply the answer when {@link #phase} is {@link Phase#DONE}, else {@code null}
* @param reply the answer when {@link #phase} is {@link Phase#DONE}, or the question when
* {@link #phase} is {@link Phase#ASKING}; otherwise {@code null}
* @param replySource {@code "reply"} (structured {@code bridge_reply}) or {@code "transcript"}
* (completion scrape) when {@link Phase#DONE}, else {@code null}
* @param detail a human note (live worker status while pending, or the failure reason)
* @param detail a human note (live worker status while pending, ask state, or failure reason)
* @param turnId correlation id for an {@link Phase#ASKING} ticket, else {@code null}
*/
public record TaskView(String ticket, Phase phase, String reply, String replySource, String detail) {
public record TaskView(String ticket, Phase phase, String reply, String replySource, String detail,
String turnId) {
}
/** An in-flight or finished async delegation, keyed by its ticket. */
private record Task(String target, CompletableFuture<Reply> future, long createdNanos) {
private static final class Task {
private final String target;
private final CompletableFuture<Reply> future = new CompletableFuture<>();
private final long createdNanos;
private volatile Reply question;
private volatile String turnId;
private Task(String target, long createdNanos) {
this.target = target;
this.createdNanos = createdNanos;
}
}
private final AgentControl agents;
@@ -154,8 +182,16 @@ public final class MessageService {
private final ReplyInbox inbox;
private final ReplyPushLoop pushLoop;
private final Metrics metrics; // CB-502: nullable — no registry in unit tests
// CB-588: injectable so pruneTerminalTickets' 10-minute TICKET_TTL_NANOS can be exercised in a
// test without a real wait — same seam SessionManager already uses for its idle reaper (nowNanos).
private final LongSupplier nowNanos;
private final ConcurrentHashMap<String, ReentrantLock> sessionLocks = new ConcurrentHashMap<>();
private final ConcurrentHashMap<String, Task> tasks = new ConcurrentHashMap<>();
/** Async task that owns each exact forward rendezvous waiter. */
private final ConcurrentHashMap<CompletableFuture<Rendezvous.Resolution>, Task> asyncTasksByWaiter =
new ConcurrentHashMap<>();
/** Async tickets paused on a specific {@code bridge_ask} turn. */
private final ConcurrentHashMap<String, Task> asyncTasksByTurn = new ConcurrentHashMap<>();
private final AtomicLong ticketSeq = new AtomicLong();
private final ExecutorService asyncExecutor = Executors.newThreadPerTaskExecutor(
Thread.ofVirtual().name("bridge-async-", 0).factory());
@@ -164,7 +200,9 @@ public final class MessageService {
* Create with an explicit {@link ReplyInbox} and optional {@link ReplyPushLoop}.
*
* @param pushLoop nullable — when non-null, the push loop is notified on the no-waiter reply
* branch ({@link #reply}) so it can nudge the primary to drain the inbox
* branch ({@link #reply}) so it can nudge the primary to drain the inbox, and
* (CB-588) whenever an async ticket started by {@link #sendAsync} reaches a
* terminal phase, and whenever {@link #poll} hands a terminal ticket to its caller
*/
public MessageService(AgentControl agents, Injector injector, Rendezvous rendezvous,
ReplyInbox inbox, ReplyPushLoop pushLoop) {
@@ -179,12 +217,19 @@ public final class MessageService {
*/
public MessageService(AgentControl agents, Injector injector, Rendezvous rendezvous,
ReplyInbox inbox, ReplyPushLoop pushLoop, Metrics metrics) {
this(agents, injector, rendezvous, inbox, pushLoop, metrics, System::nanoTime);
}
/** Test constructor with an injectable clock (CB-588: exercise the ticket-prune TTL without a real wait). */
MessageService(AgentControl agents, Injector injector, Rendezvous rendezvous, ReplyInbox inbox,
ReplyPushLoop pushLoop, Metrics metrics, LongSupplier nowNanos) {
this.agents = agents;
this.injector = injector;
this.rendezvous = rendezvous;
this.inbox = inbox;
this.pushLoop = pushLoop;
this.metrics = metrics;
this.nowNanos = nowNanos;
}
/** Create with an explicit {@link ReplyInbox} and no push loop. */
@@ -202,6 +247,16 @@ public final class MessageService {
return agents.status(target);
}
/** Read-only delegation fact for fleet views. */
public boolean hasAcceptedDelivery(String target) {
return rendezvous.isWaiting(target);
}
/** Read-only inbox fact for fleet views. */
public boolean hasInboxMessage(String target) {
return !inbox.peek(target).isEmpty();
}
/**
* Route a worker's explicit {@code bridge_reply}: resolve an open send, or queue it in the
* inbox if no send is currently open. Unlike the bare {@link Rendezvous#resolve}, a no-waiter
@@ -251,6 +306,7 @@ public final class MessageService {
case COMPLETED_UNREPLIED -> "completion_fallback";
case TIMED_OUT_WORKING, TIMED_OUT_QUEUED, BUSY -> "timeout";
case WORKER_FAILED -> "failed";
case BACKEND_EXHAUSTED -> "backend_exhausted";
case STALE_TURN, QUESTION -> null; // not a completed delegation
};
}
@@ -272,14 +328,18 @@ public final class MessageService {
*/
public boolean abandon(String target, String reason) {
CompletableFuture<Rendezvous.Resolution> waiter = rendezvous.currentWaiter(target);
if (waiter == null || waiter.isDone()) {
return false; // nobody is blocked on this worker — nothing to abandon
boolean failed = waiter != null && !waiter.isDone() && rendezvous.resolveFailure(waiter, reason);
boolean asyncFailed = false;
for (Task task : tasks.values()) {
if (target.equals(task.target) && task.question == null
&& task.future.complete(new Reply(Outcome.WORKER_FAILED, reason))) {
asyncFailed = true;
}
}
boolean failed = rendezvous.resolveFailure(waiter, reason);
if (failed) {
log.debug("abandoned send to {}: {}", target, reason);
log.warn("abandoning the blocked send to {}: {}", target, reason);
}
return failed;
return failed || asyncFailed;
}
/**
@@ -291,9 +351,30 @@ public final class MessageService {
}
/**
* Drain (peek + ack) all pending inbox replies for {@code target}. At-least-once: returns the
* messages and acknowledges them; an in-flight failure between returning and the caller
* processing them re-surfaces them on a subsequent drain (the ack is local).
* Drain (peek + ack) all pending inbox replies for {@code target}.
*
* <p><strong>The ack happens here, before the caller has the messages</strong> — before the MCP
* or REST response carrying them has been written, and long before the client has processed
* them. That ordering is what the two adapters disagree about, so do not read this method as
* "at-least-once" without qualifying which inbox is behind it (CB-529):
*
* <ul>
* <li>{@code InMemoryReplyInbox} — the ack only drops an entry from a local map. The messages
* are already in the returned list, so nothing can be lost after this point.
* <li>{@code AmqpReplyInbox} — the ack is a broker-side {@code basicAck}. Once it lands the
* broker has forgotten the message. If the daemon dies while writing the response, the
* reply is gone from the broker <em>and</em> the client never received it. Re-polling
* cannot recover it, because there is nothing left to re-deliver.
* </ul>
*
* <p>So the loss window is the response write, and it is a genuine loss rather than a
* redelivery. This is accepted, not overlooked: the alternative — ack on the next poll — turns
* every normal drain into a double delivery, which costs more than the window it closes. A
* caller that needs certainty re-polls; that is idempotent for every case except this one.
*
* <p>Any change here must be checked against <em>both</em> adapters. The previous version of
* this javadoc claimed "the ack is local", which was true when only the in-memory inbox existed
* and silently became false when the AMQP adapter landed.
*
* @return the drained messages, newest last (FIFO); empty list if none
*/
@@ -310,6 +391,27 @@ public final class MessageService {
* worker replies via {@link Rendezvous} or {@code timeoutMillis} elapses.
*/
public Reply send(String target, String content, long timeoutMillis) {
return send(target, content, timeoutMillis, null);
}
/**
* As {@link #send(String, String, long)}, but with an accepted-delivery hook.
*
* <p>{@code onAccepted} is invoked exactly once, once this send has won {@code target}'s send
* lock and so become the <em>accepted target turn</em> — it runs <em>before</em> delivery is
* queued, so a throwing hook fails the send cleanly (the waiter it already opened is closed and
* nothing is left queued). It is <em>not</em> invoked when the send is {@link Outcome#BUSY}
* (lock never taken). A caller uses this to record that <em>it</em> now owns the delegation's
* reply routing (CB-548: {@code PrimaryRegistry} delegator ownership) — recording only on
* acceptance means a concurrent sender that times out {@code BUSY} can never steal ownership it
* never earned. {@code null} disables the hook.
*/
public Reply send(String target, String content, long timeoutMillis, Runnable onAccepted) {
return send(target, content, timeoutMillis, onAccepted, null);
}
/** Run a send, optionally stopping an async task that teardown already failed before acceptance. */
private Reply send(String target, String content, long timeoutMillis, Runnable onAccepted, Task task) {
long deadlineNanos = System.nanoTime() + timeoutMillis * 1_000_000L;
ReentrantLock lock = sessionLocks.computeIfAbsent(target, _ -> new ReentrantLock());
@@ -317,23 +419,48 @@ public final class MessageService {
return new Reply(Outcome.BUSY, null); // another send held the session the whole window
}
try {
CompletableFuture<Void> delivered = injector.enqueue(target, content);
if (task != null && task.future.isDone()) {
return task.future.getNow(null);
}
if (hasAsyncQuestion(target)) {
return new Reply(Outcome.BUSY, null); // the worker's current turn is paused for its lead
}
// Open the waiter BEFORE queueing delivery (CB-548). A fast reply — the worker already
// injectable the instant we enqueue — otherwise arrives before the waiter is registered
// and orphans into the inbox while this send blocks to the timeout (the enqueue-before-
// open race). Opening first also means a throwing onAccepted (fired before enqueue) or an
// enqueue failure is safely closed by the finally below: nothing is left queued, and the
// failed send leaves no stale waiter behind.
CompletableFuture<Rendezvous.Resolution> reply = rendezvous.open(target);
try {
Rendezvous.Resolution r = reply.get(remainingMillis(deadlineNanos), TimeUnit.MILLISECONDS);
return recorded(new Reply(outcomeOf(r.kind()), r.text(), r.turnId()));
} catch (TimeoutException e) {
boolean wasDelivered = delivered.isDone() && !delivered.isCompletedExceptionally();
log.debug("send to {} timed out (delivered={})", target, wasDelivered);
return recorded(new Reply(
wasDelivered ? Outcome.TIMED_OUT_WORKING : Outcome.TIMED_OUT_QUEUED, null));
} catch (ExecutionException e) {
Throwable cause = e.getCause();
throw cause instanceof RuntimeException re ? re : new IllegalStateException(cause);
} catch (InterruptedException e) {
Thread.currentThread().interrupt();
throw new IllegalStateException("interrupted awaiting reply from " + target, e);
if (task != null) {
asyncTasksByWaiter.put(reply, task);
}
TurnToken token = new TurnToken(target, reply);
// The send has won the lock; the accepted-delivery hook records delegator ownership
// here (CB-548). It runs BEFORE enqueue so a throwing hook — onAccepted is now a
// public callback — fails the send without queuing a message that would orphan.
if (onAccepted != null) {
onAccepted.run();
}
CompletableFuture<Void> delivered = injector.enqueue(target, content, token);
try {
Rendezvous.Resolution r = reply.get(remainingMillis(deadlineNanos), TimeUnit.MILLISECONDS);
return recorded(new Reply(outcomeOf(r.kind()), r.text(), r.turnId()));
} catch (TimeoutException e) {
boolean wasDelivered = delivered.isDone() && !delivered.isCompletedExceptionally();
log.debug("send to {} timed out (delivered={})", target, wasDelivered);
return recorded(new Reply(
wasDelivered ? Outcome.TIMED_OUT_WORKING : Outcome.TIMED_OUT_QUEUED, null));
} catch (ExecutionException e) {
Throwable cause = e.getCause();
throw cause instanceof RuntimeException re ? re : new IllegalStateException(cause);
} catch (InterruptedException e) {
Thread.currentThread().interrupt();
throw new IllegalStateException("interrupted awaiting reply from " + target, e);
}
} finally {
asyncTasksByWaiter.remove(reply);
rendezvous.close(target, reply);
}
} finally {
@@ -359,7 +486,12 @@ public final class MessageService {
if (ticket.fresh()) {
// Register the reverse waiter first, then surface the question — so the answer, which can
// arrive the instant the primary reacts, always finds an open waiter to resolve.
CompletableFuture<Rendezvous.Resolution> waiter = rendezvous.currentWaiter(workerSession);
Task task = markAsyncQuestion(waiter, question, ticket.turnId());
if (!rendezvous.resolveQuestion(workerSession, question, ticket.turnId())) {
if (task != null) {
clearAsyncQuestion(ticket.turnId(), true);
}
rendezvous.closeAsk(ticket.turnId());
return new AskResult(AskOutcome.NO_WAITER, null); // no primary is blocked on this worker
}
@@ -369,6 +501,7 @@ public final class MessageService {
return new AskResult(AskOutcome.ANSWERED, answer);
} catch (TimeoutException e) {
log.debug("bridge_ask from {} went unanswered in {}ms", workerSession, timeoutMillis);
clearAsyncQuestion(ticket.turnId(), true);
return new AskResult(AskOutcome.TIMED_OUT, null);
} catch (ExecutionException e) {
Throwable cause = e.getCause();
@@ -411,9 +544,12 @@ public final class MessageService {
rendezvous.close(workerSession, reply);
return new Reply(Outcome.STALE_TURN, null); // lapsed between the lookup and the unblock
}
clearAsyncQuestion(turnId, false);
try {
Rendezvous.Resolution r = reply.get(remainingMillis(deadlineNanos), TimeUnit.MILLISECONDS);
return new Reply(outcomeOf(r.kind()), r.text(), r.turnId());
Reply result = new Reply(outcomeOf(r.kind()), r.text(), r.turnId());
finishAsyncTask(turnId, result);
return result;
} catch (TimeoutException e) {
// The worker resumed but hasn't replied yet — no completion fallback arms an answered
// turn (it never re-entered the injector), so a silent worker rides out the window.
@@ -440,10 +576,50 @@ public final class MessageService {
* @return the ticket to poll for the eventual result
*/
public String sendAsync(String target, String content) {
return sendAsync(target, content, null);
}
/**
* As {@link #sendAsync(String, String)}, with the accepted-delivery hook of
* {@link #send(String, String, long, Runnable)} — the running {@code send} invokes {@code onAccepted}
* the moment it becomes the accepted target turn, so async flooding records delegator ownership
* exactly as the blocking path does (CB-548).
*
* @return the ticket to poll for the eventual result
*/
public String sendAsync(String target, String content, Runnable onAccepted) {
String ticket = "task-" + ticketSeq.incrementAndGet();
CompletableFuture<Reply> future =
CompletableFuture.supplyAsync(() -> send(target, content, ASYNC_TIMEOUT_MS), asyncExecutor);
tasks.put(ticket, new Task(target, future, System.nanoTime()));
Task task = new Task(target, nowNanos.getAsLong());
tasks.put(ticket, task);
if (pushLoop != null) {
// CB-588: task.future only ever completes on a terminal phase (DONE or a failure) — a
// worker paused in bridge_ask leaves it running, per finishAsyncTask's own contract — so
// this fires exactly once, from whichever path completes it: finishAsyncTask(task, result)
// below on any non-QUESTION outcome of send() — a worker's bridge_reply, the CB-106
// completion fallback, a CB-109 wedge, TIMED_OUT, BUSY, or BACKEND_EXHAUSTED — the same
// finishAsyncTask reached via answer()'s finishAsyncTask(turnId, result) once a QUESTION
// is resolved, completeExceptionally(t) just below when send() itself throws, or a CB-516
// abandon() on teardown. Without this, MessageService.reply's rendezvous fast path (the
// one an async ticket always takes) never told the push loop anything happened — see the
// class javadoc on sendAsync/CB-107.
task.future.whenComplete((reply, ex) -> {
boolean failed = ex != null || reply == null || !reply.completed();
pushLoop.onTicketTerminal(ticket, target, failed);
});
}
asyncExecutor.submit(() -> {
try {
Reply result = send(target, content, ASYNC_TIMEOUT_MS, onAccepted, task);
if (result.outcome() == Outcome.QUESTION) {
// Keep the accepted owner until answer() finishes it. markAsyncQuestion may run
// just after resolveQuestion wakes this thread.
} else {
finishAsyncTask(task, result);
}
} catch (Throwable t) {
task.future.completeExceptionally(t);
}
});
pruneTerminalTickets();
log.debug("async send {} -> {}", ticket, target);
return ticket;
@@ -459,27 +635,39 @@ public final class MessageService {
if (task == null) {
return null;
}
CompletableFuture<Reply> f = task.future();
CompletableFuture<Reply> f = task.future;
if (!f.isDone()) {
return new TaskView(ticket, Phase.PENDING, null, null, "worker " + liveStatus(task.target()));
Reply question = task.question;
if (question != null) {
return new TaskView(ticket, Phase.ASKING, question.text(), null,
"worker is waiting for your answer", question.turnId());
}
return new TaskView(ticket, Phase.PENDING, null, null, "worker " + liveStatus(task.target), null);
}
// CB-588: the ticket is terminal and being handed to the caller right here — tell the push
// loop it is collected so a later tick's nudge never names a ticket the lead already has.
if (pushLoop != null) {
pushLoop.ticketCollected(ticket);
}
Reply r;
try {
r = f.getNow(null);
} catch (CompletionException | java.util.concurrent.CancellationException e) {
Throwable cause = (e instanceof CompletionException ce && ce.getCause() != null) ? ce.getCause() : e;
return new TaskView(ticket, Phase.FAILED, null, null, cause.getMessage());
return new TaskView(ticket, Phase.FAILED, null, null, cause.getMessage(), null);
}
if (r.completed()) {
String source = r.outcome() == Outcome.REPLIED ? "reply" : "transcript";
return new TaskView(ticket, Phase.DONE, r.text(), source, null);
return new TaskView(ticket, Phase.DONE, r.text(), source, null, null);
}
// A wedged worker (CB-109) carries the error context as its reason; the timeout/busy
// outcomes carry none, so fall back to the outcome name.
String detail = r.outcome() == Outcome.WORKER_FAILED && r.text() != null
// A wedged worker (CB-109) or a backend-exhausted classification (CB-578 stage A) carries
// the real cause as its reason; the timeout/busy outcomes carry none, so fall back to the
// outcome name.
boolean carriesReason = r.outcome() == Outcome.WORKER_FAILED || r.outcome() == Outcome.BACKEND_EXHAUSTED;
String detail = carriesReason && r.text() != null
? r.text()
: "no reply — " + r.outcome().name().toLowerCase();
return new TaskView(ticket, Phase.FAILED, null, null, detail);
return new TaskView(ticket, Phase.FAILED, null, null, detail, null);
}
/** Best-effort live worker status for a pending poll; never throws (a lookup error is just noise). */
@@ -491,10 +679,71 @@ public final class MessageService {
}
}
/** Drop finished tickets older than the TTL so the registry cannot grow without bound. */
/**
* Drop finished tickets older than the TTL so {@link #tasks} cannot grow without bound.
*
* <p>{@code tasks} is the sole authority on whether a ticket still exists — {@link #poll} returns
* {@code null} the instant a ticket is gone from here, before it ever reaches the terminal branch
* that calls {@link ReplyPushLoop#ticketCollected}. Without telling the push loop about a prune
* too, its own {@code pendingTickets} entry would outlive the ticket it names: an unpolled ticket
* (or one the reminder cap already gave up on) is pruned here but never collected there, so it
* lingers in {@code pendingTickets} forever and rides along on every later nudge to the same lead
* — naming a ticket {@code bridge_poll} can no longer find (CB-588 follow-up).
*/
private void pruneTerminalTickets() {
long cutoff = System.nanoTime() - TICKET_TTL_NANOS;
tasks.values().removeIf(t -> t.future().isDone() && t.createdNanos() < cutoff);
long cutoff = nowNanos.getAsLong() - TICKET_TTL_NANOS;
tasks.entrySet().removeIf(e -> {
Task t = e.getValue();
boolean expired = t.future.isDone() && t.createdNanos < cutoff;
if (expired && pushLoop != null) {
pushLoop.ticketCollected(e.getKey());
}
return expired;
});
}
/** Record the active question for an async ticket; blocking sends have no entry and stay unchanged. */
private Task markAsyncQuestion(CompletableFuture<Rendezvous.Resolution> waiter, String text, String turnId) {
Task task = waiter == null ? null : asyncTasksByWaiter.get(waiter);
if (task != null) {
task.question = new Reply(Outcome.QUESTION, text, turnId);
task.turnId = turnId;
asyncTasksByTurn.put(turnId, task);
}
return task;
}
/** Clear an answered or lapsed question, but only when it matches the ticket's current turn. */
private void clearAsyncQuestion(String turnId, boolean forgetTurn) {
Task task = asyncTasksByTurn.get(turnId);
if (task != null && turnId.equals(task.turnId)) {
task.question = null;
if (forgetTurn) {
asyncTasksByTurn.remove(turnId, task);
task.turnId = null;
}
}
}
/** Complete and detach an async ticket after its worker's actual terminal reply. */
private void finishAsyncTask(Task task, Reply result) {
task.future.complete(result);
if (task.turnId != null) {
asyncTasksByTurn.remove(task.turnId, task);
}
}
/** Complete the async ticket correlated to a specific answered turn. */
private void finishAsyncTask(String turnId, Reply result) {
Task task = asyncTasksByTurn.get(turnId);
if (task != null) {
finishAsyncTask(task, result);
}
}
/** A new send must not open a waiter while an async ticket owns this worker's paused turn. */
private boolean hasAsyncQuestion(String target) {
return asyncTasksByTurn.values().stream().anyMatch(task -> target.equals(task.target));
}
/** Release the async executor. */
@@ -508,6 +757,7 @@ public final class MessageService {
case REPLY -> Outcome.REPLIED;
case COMPLETION -> Outcome.COMPLETED_UNREPLIED;
case FAILED -> Outcome.WORKER_FAILED;
case BACKEND_EXHAUSTED -> Outcome.BACKEND_EXHAUSTED;
case QUESTION -> Outcome.QUESTION;
};
}
@@ -33,6 +33,13 @@ public final class Rendezvous {
COMPLETION,
/** The worker ran the turn then wedged (CB-109); {@code text} is the failure context. */
FAILED,
/**
* The turn finished without a {@code bridge_reply}, and the scrape matched the backend's
* configured usage-limit refusal pattern (CB-578 stage A); {@code text} is the reason,
* carrying the matched line. The pane is healthy — only the account is refusing — so this
* is kept separate from a session simply going {@code GONE}.
*/
BACKEND_EXHAUSTED,
/**
* The worker paused mid-turn to ask the primary a question (CB-205 reverse rendezvous);
* {@code text} is the question and {@code turnId} correlates the primary's answer back to
@@ -74,15 +81,30 @@ public final class Rendezvous {
/**
* Register a waiter for {@code session} — the await side of the public {@code resolve*} methods.
* The caller must hold that session's send lock.
*
* <p>Atomic fail-if-present (CB-548): if a waiter is already registered for {@code session}, an
* {@link IllegalStateException} is thrown rather than replacing the first — so any future
* invariant violation fails loudly instead of silently swapping the waiter another send is
* blocked on. {@code MessageService} serializes sends per session (the send lock), so in correct
* code a double open is impossible; this is a tripwire for the day that no longer holds.
*/
public CompletableFuture<Resolution> open(String session) {
CompletableFuture<Resolution> waiter = new CompletableFuture<>();
waiters.put(session, waiter);
CompletableFuture<Resolution> existing = waiters.putIfAbsent(session, waiter);
if (existing != null) {
throw new IllegalStateException(
"rendezvous double-open for session " + session + " — a waiter is already registered");
}
return waiter;
}
/** Remove {@code waiter} for {@code session} (only if it is still the registered one). */
void close(String session, CompletableFuture<Resolution> waiter) {
/**
* Remove {@code waiter} for {@code session}, only if it is still the registered one. The
* symmetric complement of {@link #open}: a terminal send deregisters its waiter so the next
* send on the session may {@link #open} a fresh one (CB-548 makes double-open an error, so a
* successful {@code open} after a finished turn requires this close to have happened first).
*/
public void close(String session, CompletableFuture<Resolution> waiter) {
waiters.remove(session, waiter);
}
@@ -209,6 +231,19 @@ public final class Rendezvous {
return waiter != null && waiter.complete(new Resolution(Kind.FAILED, reason));
}
/**
* Resolve a specific captured {@code waiter} as {@link Kind#BACKEND_EXHAUSTED} (CB-578 stage A):
* the turn finished with no {@code bridge_reply} and the scrape matched the backend's configured
* usage-limit pattern; {@code reason} carries the matched line. Like
* {@link #resolveCompletion(CompletableFuture, String)} it targets the exact captured send
* (CB-116). A no-op if that waiter was already resolved — first resolution wins.
*
* @return {@code true} if this call resolved the waiter, {@code false} if it was null or already resolved
*/
public boolean resolveExhausted(CompletableFuture<Resolution> waiter, String reason) {
return waiter != null && waiter.complete(new Resolution(Kind.BACKEND_EXHAUSTED, reason));
}
private boolean complete(String session, Resolution resolution) {
CompletableFuture<Resolution> waiter = waiters.get(session);
return waiter != null && waiter.complete(resolution);
@@ -10,16 +10,36 @@ import java.util.List;
* <p><strong>This interface is the port.</strong> {@link InMemoryReplyInbox} is the Stage-1 adapter;
* an AMQP-backed adapter (Stage 2) must implement the same contract (idempotent publish, FIFO peek,
* at-least-once ack).
*
* <p><strong>Ownership is explicit.</strong> A gateway {@link #own owns} the inbox for each agent it
* spawned; only the owner consumes and drains it. {@link #publish} sends a reply to the target's
* inbox but does <em>not</em> imply ownership or start a consumer. This separation is required by
* CB-308 federation, where one gateway may publish to an agent owned by another gateway.
*/
public interface ReplyInbox {
/** A queued reply: an idempotency id, the worker session it came from, and the reply text. */
record InboxMessage(String msgId, String target, String content) {}
/**
* Start owning (consuming) the inbox for {@code target}. Idempotent: multiple calls for the same
* target are no-ops. The owner is the only gateway that may {@link #peek} and {@link #ack} replies
* for this target.
*/
void own(String target);
/**
* Stop owning (consuming) the inbox for {@code target}. Idempotent. Any replies held locally but
* not yet acked are dropped from the local snapshot; the underlying durable queue keeps
* unacked messages for redelivery when the target is re-owned.
*/
void release(String target);
/**
* Queue {@code content} from worker {@code target} under {@code msgId}. Idempotent: publishing an
* already-present {@code msgId} for {@code target} is a no-op (dedup), so an at-least-once Stage-2
* redelivery cannot double-queue.
* redelivery cannot double-queue. Publishing does <em>not</em> imply ownership and must not start a
* consumer.
*/
void publish(String target, String msgId, String content);
@@ -8,26 +8,55 @@ import dev.ltms.bridged.metrics.Metrics;
import org.slf4j.Logger;
import org.slf4j.LoggerFactory;
import java.util.ArrayList;
import java.util.HashSet;
import java.util.List;
import java.util.Set;
import java.util.concurrent.ConcurrentHashMap;
import java.util.concurrent.ScheduledExecutorService;
import java.util.concurrent.TimeUnit;
import java.util.stream.Collectors;
/**
* Mechanism (b) of CB-307: a dedicated, status-gated push loop that nudges the primary's own
* herdr pane when a worker reply lands with no live {@code bridge_send} to resolve it.
* A status-gated push loop that nudges a lead's own herdr pane when it has uncollected work
* waiting: a worker reply queued with no live {@code bridge_send} to resolve it (CB-307), or an
* async delegation ticket ({@code bridge_send(wait:false)}) that reached a terminal phase
* (CB-588).
*
* <p>The loop is triggered by {@link #onReplyQueued(String)} (called from
* {@link MessageService#reply} after the durable inbox publish). It checks four conditions
* at each tick via {@link #decide(String, int)}, then either injects a drain nudge,
* waits for the primary to become injectable, or stops reminding.
* <p><strong>CB-590: one schedule per lead.</strong> Both kinds of work are triggered through
* their own entry point — {@link #onReplyQueued(String)} and
* {@link #onTicketTerminal(String, String, boolean)} — but both resolve the lead that should be
* nudged and coalesce onto a single per-lead reminder schedule, tracked in {@link #activeLeads}.
* Earlier this was two independent schedules (one keyed by worker target for replies, one keyed
* by lead for tickets) that could both decide to inject into the same pane in the same window —
* a race, not routine behaviour, but the expensive kind: it interrupts the lead's live turn
* twice. Collapsing to one schedule per lead makes that structurally impossible: at most one
* scheduled tick chain is ever live for a given lead (guarded by {@link #activeLeads}'
* compare-and-set), so at most one {@code agents.send} to that lead's pane is ever in flight.
*
* <p>Bounded: at most {@link #maxReminders} nudges per target, with a configurable backoff
* between them. The reply is never lost — the durable inbox is the backstop.
* <p>Each tick examines <em>everything</em> pending for that lead — reply targets whose inbox
* still holds an unacked message ({@link #pendingReplies}) and tickets not yet collected
* ({@link #pendingTickets}) — and sends at most one combined nudge per tick
* ({@link #injectNudge(String, int, int)}). Work that arrives while the lead is busy is never
* lost: it is re-read fresh on every tick until the lead is injectable or its own reminder cap
* ({@link #maxReminders}) is reached — reply and ticket work each spend from their own budget, so
* one source exhausting its cap does not stop nudges about the other (post-CB-590 regression fix;
* see {@link #decide}) — whichever the durable inbox / pending-ticket set doesn't already answer
* via {@code STOP}.
*/
public final class ReplyPushLoop {
private static final Logger log = LoggerFactory.getLogger(ReplyPushLoop.class);
static final String NUDGE_FORMAT = "Worker %s returned a reply — run bridge_poll(target=%s) to collect it";
/** Coalesced form, several uncollected replies for the same lead. */
static final String REPLIES_NUDGE_FORMAT =
"%d workers returned replies — run bridge_poll(target=...) for each to collect them: %s";
/** Singular form, one uncollected ticket. */
static final String TICKET_NUDGE_FORMAT =
"Ticket %s finished%s — run bridge_poll(ticket=%s) to collect it";
/** Coalesced form, several uncollected tickets for the same lead. */
static final String TICKETS_NUDGE_FORMAT =
"%d tickets finished%s — run bridge_poll(ticket=...) for each to collect them: %s";
private final PrimaryRegistry primaryRegistry;
private final AgentControl agents;
@@ -37,8 +66,12 @@ public final class ReplyPushLoop {
private final long backoffMs;
private final Metrics metrics; // CB-512: nullable — no registry in unit tests
/** Track targets that have an active schedule. */
private final ConcurrentHashMap<String, Boolean> activeTargets = new ConcurrentHashMap<>();
/** Worker targets with a reply queued, and the lead to nudge about it, keyed by target. */
private final ConcurrentHashMap<String, String> pendingReplies = new ConcurrentHashMap<>();
/** Tickets that have gone terminal but not yet been polled, keyed by ticket. */
private final ConcurrentHashMap<String, PendingTicket> pendingTickets = new ConcurrentHashMap<>();
/** CB-590: leads with an active combined reminder schedule (replies and/or tickets). */
private final ConcurrentHashMap<String, Boolean> activeLeads = new ConcurrentHashMap<>();
public ReplyPushLoop(PrimaryRegistry primaryRegistry, AgentControl agents, ReplyInbox inbox,
ScheduledExecutorService scheduler,
@@ -59,114 +92,324 @@ public final class ReplyPushLoop {
this.metrics = metrics;
}
/** Record a counter sample when a registry is wired; a no-op in unit tests. */
private void count(String name, String... labels) {
/** Count one nudge outcome when a registry is wired; a no-op in unit tests. */
private void countNudge(String outcome) {
if (metrics != null) {
metrics.inc(name, labels);
metrics.inc(BridgedMetrics.PUSH_NUDGES, "outcome", outcome);
}
}
// --- decision logic (package-private for unit-testing) -------------------------------------
// --- pending-work lookups (package-private for unit-testing) -------------------------------
/** The action the loop should take for a target at the given reminder count. */
/** The action the loop should take for a lead at the given reminder count. */
enum Action { INJECT, WAIT_BUSY, STOP }
/**
* Pure decision function: examine the current state and return what the loop should do.
* Reply targets still pending for {@code lead} — registered via {@link #onReplyQueued} and
* whose inbox still holds an unacked message. A target whose inbox has since drained (acked,
* or collected via a live {@code bridge_send} rendezvous instead) is dropped from
* {@link #pendingReplies} here rather than lingering forever; there is no explicit "reply
* collected" callback the way {@link #ticketCollected} exists for tickets, so the inbox itself
* is the only signal.
*/
private Set<String> pendingReplyTargetsFor(String lead) {
Set<String> result = new HashSet<>();
for (var entry : pendingReplies.entrySet()) {
String target = entry.getKey();
String owningLead = entry.getValue();
if (!lead.equals(owningLead)) continue;
if (inbox.peek(target).isEmpty()) {
pendingReplies.remove(target, owningLead);
continue;
}
result.add(target);
}
return result;
}
/** A ticket awaiting collection: which lead to nudge, and whether it ended in failure. */
private record PendingTicket(String ticket, String lead, boolean failed) {
}
/** Tickets still pending for {@code lead}, snapshotted fresh for one tick. */
private List<PendingTicket> pendingTicketsFor(String lead) {
return pendingTickets.values().stream().filter(t -> lead.equals(t.lead())).toList();
}
/** Ticket ids still pending for {@code lead} — a plain snapshot for race comparison. */
private Set<String> pendingTicketIdsFor(String lead) {
return pendingTicketsFor(lead).stream().map(PendingTicket::ticket)
.collect(Collectors.toUnmodifiableSet());
}
/**
* Pure decision function: examine everything pending for {@code lead} — reply targets and
* tickets alike — and return what the loop should do.
*
* @param target the worker session (target terminal id)
* @param reminderCount how many nudges have been sent so far for this target
* <p><strong>CB-590-fix: one schedule, two budgets.</strong> The single per-lead schedule
* (CB-590) still ticks once for both sources, but each source is capped independently —
* {@code replyReminderCount} against a reply target still pending, {@code ticketReminderCount}
* against a ticket still pending. A busy reply stream that exhausts its own cap must not stop
* the loop from nudging about a ticket that still has budget left, and vice versa: either
* source being eligible (has pending work AND is under its own cap) is enough for
* {@link Action#INJECT}. Only when neither source has eligible work does the loop
* {@link Action#STOP}.
*
* @param lead the lead terminal to nudge
* @param replyReminderCount how many nudges have covered pending reply work for this lead
* @param ticketReminderCount how many nudges have covered pending ticket work for this lead
* @return the action the caller should take
*/
Action decide(String target, int reminderCount) {
if (!primaryRegistry.isKnown()) {
log.debug("push: primary unknown, stopping reminder for {}", target);
Action decide(String lead, int replyReminderCount, int ticketReminderCount) {
boolean hasReplyWork = !pendingReplyTargetsFor(lead).isEmpty();
boolean hasTicketWork = !pendingTicketIdsFor(lead).isEmpty();
if (!hasReplyWork && !hasTicketWork) {
log.debug("push: nothing pending for lead {}, stopping reminder", lead);
return Action.STOP;
}
if (inbox.peek(target).isEmpty()) {
log.debug("push: inbox empty for {}, stopping reminder", target);
boolean replyEligible = hasReplyWork && replyReminderCount < maxReminders;
boolean ticketEligible = hasTicketWork && ticketReminderCount < maxReminders;
if (!replyEligible && !ticketEligible) {
log.debug("push: reminder cap ({}) reached for lead {} on every source with pending work, stopping",
maxReminders, lead);
countNudge("exhausted");
return Action.STOP;
}
if (reminderCount >= maxReminders) {
log.debug("push: reminder cap ({}) reached for {}, stopping", maxReminders, target);
count(BridgedMetrics.PUSH_NUDGES, "outcome", "exhausted");
return Action.STOP;
}
var primaryTerminal = primaryRegistry.primaryTerminal().orElseThrow();
AgentStatus status;
try {
status = agents.status(primaryTerminal);
status = agents.status(lead);
} catch (RuntimeException e) {
log.debug("push: status check failed for primary {}, will retry", primaryTerminal, e);
log.debug("push: status check failed for lead {}, will retry", lead, e);
return Action.WAIT_BUSY;
}
if (status.injectable()) {
return Action.INJECT;
}
log.debug("push: primary {} is {} (not injectable), waiting", primaryTerminal, status);
log.debug("push: lead {} is {} (not injectable), waiting", lead, status);
return Action.WAIT_BUSY;
}
// --- public entrypoint ---------------------------------------------------------------------
// --- public entrypoints ----------------------------------------------------------------------
/**
* Called when a reply is queued for {@code target}. Idempotent per target: a second call while
* a schedule is active is a no-op. The schedule nudges the primary, then schedules a follow-up
* check (reminder on backoff, or re-check on WAIT_BUSY), until the inbox is empty or the cap
* is reached.
* Called when a reply is queued for {@code target}. Resolves the lead delegating to
* {@code target} (CB-532) and coalesces onto that lead's single reminder schedule — starting
* one if none is active, joining an already-active one otherwise. A no-op if no lead is known
* to be waiting on {@code target}: there is nobody to nudge yet, and the durable inbox is the
* backstop until a lead is recorded.
*/
public void onReplyQueued(String target) {
if (activeTargets.putIfAbsent(target, Boolean.TRUE) != null) {
log.debug("push: already active for {}, ignoring duplicate trigger", target);
return; // already scheduled
var lead = primaryRegistry.nudgeTargetFor(target);
if (lead.isEmpty()) {
log.debug("push: no lead is known to be waiting on {}, skipping reminder", target);
return;
}
log.debug("push: starting reminder loop for {}", target);
scheduleNext(target, 0);
pendingReplies.put(target, lead.get());
startOrCoalesce(lead.get());
}
/**
* Called when an async delegation ticket ({@code bridge_send(wait:false)}, CB-107) reaches a
* terminal phase — DONE or a failure. Unlike {@link #onReplyQueued}, which nudges about the
* durable-inbox no-waiter path, this covers the path {@code MessageService.reply} takes when a
* fire-and-poll send's own rendezvous waiter resolves the reply directly: that path returns
* before {@link #onReplyQueued} is ever called, so without this entry point a ticket finishing
* that way never nudged anyone (CB-588 / gitea #72).
*
* <p>Resolves the delegating lead the same way {@link #onReplyQueued} does and coalesces onto
* the same per-lead schedule (CB-590) — several tickets, or a ticket and a reply, finishing
* for the same lead while its schedule is already active all ride the existing schedule's next
* tick rather than firing a nudge each.
*
* @param ticket the ticket to nudge about
* @param target the worker session the ticket was sent to — resolves which lead delegated it
* @param failed whether the ticket ended in a failure phase rather than {@code DONE}
*/
public void onTicketTerminal(String ticket, String target, boolean failed) {
var lead = primaryRegistry.nudgeTargetFor(target);
if (lead.isEmpty()) {
log.debug("push: no lead is known to be waiting on ticket {} (target {}), skipping nudge",
ticket, target);
return;
}
pendingTickets.put(ticket, new PendingTicket(ticket, lead.get(), failed));
startOrCoalesce(lead.get());
}
/**
* Called when a ticket's terminal state has been collected via {@code bridge_poll}. Removes it
* from the pending set so a scheduled tick — and any nudge it sends — never names a ticket the
* lead already has. A ticket that was never pending (unknown ticket, or one nudged with no push
* loop configured) is a no-op.
*/
public void ticketCollected(String ticket) {
pendingTickets.remove(ticket);
}
// --- the schedule ----------------------------------------------------------------------------
/** Start a reminder schedule for {@code lead}, or join the one already running. */
private void startOrCoalesce(String lead) {
if (activeLeads.putIfAbsent(lead, Boolean.TRUE) != null) {
log.debug("push: reminder loop already active for lead {}, work coalesced in", lead);
return;
}
log.debug("push: starting reminder loop for lead {}", lead);
scheduleNext(lead, 0, 0);
}
/** Execute one loop tick — called on the scheduler thread. */
private void tick(String target, int reminderCount) {
var action = decide(target, reminderCount);
private void tick(String lead, int replyReminderCount, int ticketReminderCount) {
Set<String> repliesBefore = pendingReplyTargetsFor(lead);
Set<String> ticketsBefore = pendingTicketIdsFor(lead);
var action = decide(lead, replyReminderCount, ticketReminderCount);
switch (action) {
case INJECT -> {
injectNudge(target, reminderCount);
scheduleNext(target, reminderCount + 1);
}
// Re-check after the configured backoff; the primary may become injectable soon.
case WAIT_BUSY -> scheduleNext(target, reminderCount);
case STOP -> {
activeTargets.remove(target);
log.debug("push: reminder loop ended for {}", target);
injectNudge(lead, replyReminderCount, ticketReminderCount);
// Only the source(s) actually eligible this tick spend a unit of their own budget —
// an exhausted source riding along in the combined message (still pending, still
// named) does not get charged again; its count stays put until it drains.
boolean replyEligible = !repliesBefore.isEmpty() && replyReminderCount < maxReminders;
boolean ticketEligible = !ticketsBefore.isEmpty() && ticketReminderCount < maxReminders;
scheduleNext(lead,
replyEligible ? replyReminderCount + 1 : replyReminderCount,
ticketEligible ? ticketReminderCount + 1 : ticketReminderCount);
}
// Re-check after the configured backoff; the lead may become injectable soon.
case WAIT_BUSY -> scheduleNext(lead, replyReminderCount, ticketReminderCount);
case STOP -> stopOrRestart(lead, repliesBefore, ticketsBefore);
}
}
/** Send the nudge and log the event. */
private void injectNudge(String target, int reminderCount) {
var primaryTerminal = primaryRegistry.primaryTerminal().orElseThrow();
String nudge = NUDGE_FORMAT.formatted(target, target);
/**
* Release {@code lead}'s active-schedule slot, then restart it only if work landed that
* {@code repliesBefore} / {@code ticketsBefore} — the snapshots taken just before this tick's
* decision — did not already account for. {@link #onReplyQueued} / {@link #onTicketTerminal}
* read {@link #activeLeads} to decide whether to coalesce onto an existing schedule or start
* one, so work that lands between {@link #decide} returning {@link Action#STOP} and this
* removal running sees the (soon-to-be-stale) slot as occupied, coalesces onto a schedule that
* is about to die, and gets no nudge scheduled at all — a lost nudge, exactly what CB-588 (and
* now CB-590) exist to remove (originally found in review, gitea PR #73, for the ticket-only
* loop; carried forward here for the unified one).
*
* <p>Restarting on ANY non-empty pending set would be wrong: when STOP is reached because the
* reminder cap was hit rather than the backlog draining, the same never-collected work is
* expected to still be sitting there — that is the cap doing its job — and restarting would
* nudge about it forever, defeating the bound. Diffing the current pending sets against the
* "before" snapshots tells the two cases apart: an item present before this tick's decision is
* stale backlog, not a race; only an item absent from the "before" snapshot can only have
* arrived during the decision-to-release window, which is exactly the race this method closes.
*
* <p>Package-private so a test can drive the interleaving directly rather than trying to force a
* genuine thread race.
*
* <p>Terminates rather than spinning: this method restarts the schedule at most once per call,
* and a fresh {@link #onReplyQueued} / {@link #onTicketTerminal} racing the recheck below still
* terminates in one of two ways — either it observes the slot already vacated (by the
* {@code activeLeads.remove} above, which happens-before this recheck in program order) and
* claims it itself, or it lands first and this recheck then observes its work in
* {@link #pendingReplies} / {@link #pendingTickets} and reclaims the slot instead. Exactly one
* side always wins; neither can miss the other, so this never loops on its own account.
*/
void stopOrRestart(String lead, Set<String> repliesBefore, Set<String> ticketsBefore) {
activeLeads.remove(lead);
boolean racedIn = pendingReplyTargetsFor(lead).stream().anyMatch(t -> !repliesBefore.contains(t))
|| pendingTicketIdsFor(lead).stream().anyMatch(t -> !ticketsBefore.contains(t));
if (racedIn && activeLeads.putIfAbsent(lead, Boolean.TRUE) == null) {
log.debug("push: new work for lead {} raced the reminder loop's stop — restarting", lead);
scheduleNext(lead, 0, 0);
return;
}
log.debug("push: reminder loop ended for lead {}", lead);
}
/** Send one combined nudge covering everything currently pending for {@code lead}. */
private void injectNudge(String lead, int replyReminderCount, int ticketReminderCount) {
// Re-read rather than threading it down from decide(): a reply can drain, or a ticket be
// collected (or another arrive), between the decision and the injection.
Set<String> replyTargets = pendingReplyTargetsFor(lead);
List<PendingTicket> tickets = pendingTicketsFor(lead);
if (replyTargets.isEmpty() && tickets.isEmpty()) {
log.debug("push: pending work for lead {} drained before the nudge could be sent", lead);
return;
}
String nudge = formatNudge(replyTargets, tickets);
try {
agents.send(primaryTerminal, nudge);
log.debug("push: nudge {}/{} sent to primary {} for target {}",
reminderCount + 1, maxReminders, primaryTerminal, target);
count(BridgedMetrics.PUSH_NUDGES, "outcome", "delivered");
agents.send(lead, nudge);
log.debug("push: nudge sent to lead {} (reply {}/{}, ticket {}/{}; {} reply target(s), {} ticket(s))",
lead, replyReminderCount + 1, maxReminders, ticketReminderCount + 1, maxReminders,
replyTargets.size(), tickets.size());
countNudge("delivered");
} catch (RuntimeException e) {
log.warn("push: failed to nudge primary {} for target {} (reminder {}/{}): {}",
primaryTerminal, target, reminderCount + 1, maxReminders, e.toString());
log.warn("push: failed to nudge lead {} (reply {}/{}, ticket {}/{}): {}",
lead, replyReminderCount + 1, maxReminders, ticketReminderCount + 1, maxReminders, e.toString());
}
}
/** Schedule the next tick on the scheduler thread pool. */
private void scheduleNext(String target, int nextReminderCount) {
scheduler.schedule(() -> tick(target, nextReminderCount), backoffMs, TimeUnit.MILLISECONDS);
private void scheduleNext(String lead, int nextReplyReminderCount, int nextTicketReminderCount) {
scheduler.schedule(() -> tick(lead, nextReplyReminderCount, nextTicketReminderCount),
backoffMs, TimeUnit.MILLISECONDS);
}
// --- nudge formatting ------------------------------------------------------------------------
/** Render everything pending for one lead as a single nudge line. */
private static String formatNudge(Set<String> replyTargets, List<PendingTicket> tickets) {
List<String> parts = new ArrayList<>();
if (!replyTargets.isEmpty()) {
parts.add(formatRepliesNudge(replyTargets));
}
if (!tickets.isEmpty()) {
parts.add(formatTicketsNudge(tickets));
}
return String.join(" | ", parts);
}
/** Render one or several pending reply targets. */
private static String formatRepliesNudge(Set<String> targets) {
if (targets.size() == 1) {
String target = targets.iterator().next();
return NUDGE_FORMAT.formatted(target, target);
}
String ids = String.join(", ", targets);
return REPLIES_NUDGE_FORMAT.formatted(targets.size(), ids);
}
/** Render one or several pending tickets. */
private static String formatTicketsNudge(List<PendingTicket> pending) {
if (pending.size() == 1) {
PendingTicket t = pending.get(0);
return TICKET_NUDGE_FORMAT.formatted(t.ticket(), t.failed() ? " (FAILED)" : "", t.ticket());
}
long failedCount = pending.stream().filter(PendingTicket::failed).count();
String ids = pending.stream()
.map(t -> t.failed() ? t.ticket() + " (FAILED)" : t.ticket())
.collect(Collectors.joining(", "));
String failedNote = failedCount > 0 ? " (%d failed)".formatted(failedCount) : "";
return TICKETS_NUDGE_FORMAT.formatted(pending.size(), failedNote, ids);
}
// --- lifecycle -----------------------------------------------------------------------------
/**
* Whether any reminder loop is currently active for some lead (CB-551). The idle-lead heartbeat
* uses this to stand aside: while the push loop is actively nudging a lead, a concurrent
* heartbeat injection would start a second competing turn in the same pane — racing loops
* multiply turns and context burn. "Active" means a schedule exists in {@link #activeLeads},
* which now covers both reply-queued (CB-307) and ticket-terminal (CB-588) work (CB-590) —
* bounded by what has been triggered, not by any persistent state.
*/
public boolean isActive() {
return !activeLeads.isEmpty();
}
/** Shut down the scheduler. Outstanding reminders are cancelled. */
public void stop() {
scheduler.shutdownNow();
activeTargets.clear();
activeLeads.clear();
pendingReplies.clear();
pendingTickets.clear();
}
/** @see #stop() */
@@ -0,0 +1,20 @@
package dev.ltms.bridged.msg;
import java.util.concurrent.CompletableFuture;
/**
* Identity for one accepted send. The session turn is deliberately absent: CompletionResolver's
* delivery callback runs before SessionManager.onDelivered, so binding it needs a later ordering design.
*/
public final class TurnToken {
private final String target;
private final CompletableFuture<Rendezvous.Resolution> waiter;
public TurnToken(String target, CompletableFuture<Rendezvous.Resolution> waiter) {
this.target = target;
this.waiter = waiter;
}
public String target() { return target; }
public CompletableFuture<Rendezvous.Resolution> waiter() { return waiter; }
}
@@ -26,10 +26,33 @@ public enum Capability {
*/
WORKTREE,
/**
* The peer can discard its current conversation context without starting a delegated bridge
* turn. The command or API used to do that is adapter-specific.
*/
CONTEXT_RESET,
/**
* The spawner can reconcile orphaned peers on boot — workers that outlived a prior daemon
* process and whose pane ids died with it (CB-117). Claude Code over herdr supports this
* via name-based matching against the herdr agent list.
*/
ORPHAN_REAP
ORPHAN_REAP,
/**
* The peer surfaces the bridge's logical name ({@link SpawnRequest#sessionName()}) in its own
* UI at spawn — e.g. Claude Code's {@code -n} display name, which shows in the prompt box, the
* {@code /resume} picker, and the terminal title. This is what lets an operator tell the
* bridge's worker apart from a user's own session in the same terminal after a restart.
*/
SESSION_NAME,
/**
* The peer can be relaunched onto a prior conversation via that conversation's own session id
* ({@link SpawnRequest#resumeSessionId()}) — e.g. Claude Code's {@code -r}, which adopts the id
* as the agent's own identity rather than starting a new conversation. A launcher that declares
* this mints/resolves that id at spawn and exposes it on the returned
* {@link PeerHandle#agentSessionId()}, so a later resume addresses the same conversation.
*/
SESSION_RESUME
}
@@ -0,0 +1,84 @@
package dev.ltms.bridged.peer;
import java.nio.charset.StandardCharsets;
import java.security.MessageDigest;
import java.security.NoSuchAlgorithmException;
import java.util.HexFormat;
/**
* CB-571: a fingerprint of the exact charter bytes handed to a spawned member.
*
* <p>Lets an operator prove <em>which</em> charter a member actually got, without ever logging the
* charter text. The digest covers the exact composed UTF-8 string {@code HerdrPeerLauncher} passes
* to its adapter as {@code LaunchSpec.charter()}, so every adapter that receives the same string —
* Claude inlining it, OpenCode writing it to a file — produces the same digest for the same config.
* Two spawns of the same role from the same config agree; editing the charter changes the digest.
*
* <p>Deliberately places no charter prose. A charter is operator-authored text that may name
* internal projects or unreleased plans, and logs get tailed, shipped, and pasted into tickets.
* The {@code charterSource} key is what the operator wants to confirm, and it carries no content.
*/
public record CharterReceipt(
MemberRole role,
String profile,
String charterSource,
String charterSha256,
int charterBytes) {
/** Source reported when the role has no configured charter, so the field is never omitted. */
public static final String NO_SOURCE = "none";
/**
* The config key that supplied the role's charter text, e.g. {@code fleet.charters.architect}.
*/
public static String sourceKey(MemberRole role) {
return "fleet.charters." + (role == null ? "?" : role.wireName());
}
/**
* Fingerprint the composed charter for {@code role} on {@code profile}. {@code configured} is
* the role's charter text as read from config ({@code null} when none is configured);
* {@code composed} is the exact string the launcher will pass to the adapter — the reply
* charter may be appended to {@code configured}, or stand alone when no role charter exists.
*
* <p>No composed charter at all is reported as an explicit absence — a {@code null} digest and
* a zero byte count — never a digest of the empty string, which would hide the fact that no
* text was supplied. {@code configured} being {@code null} while {@code composed} is the reply
* charter alone is a normal case, and the source says so.
*/
public static CharterReceipt compose(MemberRole role, String profile,
String configured, String composed) {
String source = (configured == null || configured.isBlank())
? NO_SOURCE : sourceKey(role);
if (composed == null) {
return new CharterReceipt(role, profile, source, null, 0);
}
byte[] bytes = composed.getBytes(StandardCharsets.UTF_8);
return new CharterReceipt(role, profile, source, digestOf(composed), bytes.length);
}
/** Whether the composed charter was absent (no text was given to the member). */
public boolean absent() {
return charterSha256 == null;
}
/**
* The stable SHA-256 hex digest of {@code text}, or {@code null} for null/blank text. Used both
* for the receipt's fingerprint and to redact a charter argument in a spawn log.
*/
public static String digestOf(String text) {
if (text == null || text.isBlank()) {
return null;
}
return sha256Hex(text.getBytes(StandardCharsets.UTF_8));
}
private static String sha256Hex(byte[] bytes) {
try {
MessageDigest md = MessageDigest.getInstance("SHA-256");
return HexFormat.of().formatHex(md.digest(bytes));
} catch (NoSuchAlgorithmException e) {
throw new IllegalStateException("SHA-256 is unavailable", e);
}
}
}
@@ -0,0 +1,122 @@
package dev.ltms.bridged.peer;
import java.util.Locale;
/**
* What a member is <em>for</em> — the contract it runs under.
*
* <p>A member is anything a lead spawns. Every member carries two independent attributes:
*
* <ul>
* <li><b>role</b> (this enum) — <em>which contract</em>: the launch charter it is given, the role
* file it reads, the playbook skill it loads, and its authorization row.</li>
* <li><b>profile</b> (a {@code profiles:} key) — <em>which backend</em>: model, CLI adapter,
* credentials, cost.</li>
* </ul>
*
* <p>These are separate axes on purpose. A {@code REVIEWER} may run on the same profile as the
* {@code DEV} whose diff it reviews — same backend, different contract. That case is what proves
* role and profile cannot be collapsed into one field.
*
* <p>A lead is deliberately <em>not</em> a role here. A lead is not spawned: it is a pre-existing,
* human-facing session that config recognises. Only spawned peers have a member role.
*/
public enum MemberRole {
/**
* Refines a ticket before anyone implements it: scope, acceptance criteria, risks, unit split.
*
* <p>Reads the repo and writes analysis. Never commits code and never opens a pull request —
* an architect that starts implementing has stopped doing the job that makes it useful.
*
* <p>Architects are the one member kind declared in config, because a lead addresses the same
* slots across many tickets and needs a stable name for them.
*/
ARCHITECT,
/**
* Implements one unit of work: provisions a worktree, writes the code, commits, pushes, and
* opens its own pull request.
*
* <p>Never merges. The lead is the gate, and a member that merged its own work would remove the
* only independent check in the flow.
*/
DEV,
/**
* Reviews a diff it did not write and reports one structured finding.
*
* <p>Never commits, never merges, and is never the member that wrote the scope under review.
* Briefed from the diff rather than from the implementer's rationale, because that rationale
* carries the same blind spot that produced the bug.
*/
REVIEWER;
/** The lowercase spelling used in config and on the wire ({@code architect}, {@code dev}, …). */
public String wireName() {
return name().toLowerCase(Locale.ROOT);
}
/**
* The {@code fleet:} block that holds this role's pool — {@code architects},
* {@code developers}, {@code reviewers}.
*
* <p>Plural, and not always the wire name: the pool of things a {@code dev} may run on reads
* naturally as {@code developers:}. The wire name stays the singular {@code dev}, because that
* is what a tab label and a roster row say.
*/
public String configKey() {
return switch (this) {
case ARCHITECT -> "architects";
case DEV -> "developers";
case REVIEWER -> "reviewers";
};
}
/**
* The role owning the {@code fleet:} pool named {@code key}, or {@code null} when the key is not
* a role pool ({@code leaders}, {@code tabLabel}, …). Null rather than a throw: callers use this
* to sort a fleet block's children, where a non-pool key is normal rather than an error.
*/
public static MemberRole fromConfigKey(String key) {
if (key == null) {
return null;
}
String k = key.trim().toLowerCase(Locale.ROOT);
for (MemberRole r : values()) {
if (r.configKey().equals(k)) {
return r;
}
}
return null;
}
/**
* Parse a config/wire spelling, case-insensitively.
*
* @param s the spelling to parse
* @return the matching role
* @throws IllegalArgumentException when {@code s} is null, blank, or not a known role — the
* message lists the valid spellings, because a typo'd role in
* config should fail at startup with the fix in the error
*/
public static MemberRole parse(String s) {
if (s != null && !s.isBlank()) {
String t = s.trim().toLowerCase(Locale.ROOT);
for (MemberRole r : values()) {
if (r.wireName().equals(t)) {
return r;
}
}
}
StringBuilder valid = new StringBuilder();
for (MemberRole r : values()) {
if (!valid.isEmpty()) {
valid.append(", ");
}
valid.append(r.wireName());
}
throw new IllegalArgumentException(
"unknown member role '" + s + "'; valid roles are: " + valid);
}
}
@@ -12,9 +12,13 @@ package dev.ltms.bridged.peer;
public interface PeerHandle {
/**
* The registry/routing key — an opaque, launcher-assigned identifier. For the herdr-backed
* launcher this is the herdr pane id; for other launchers it is whatever their transport
* uses. Guaranteed to be non-null and unique among live peers within a single daemon process.
* The registry/routing key — an opaque, launcher-assigned identifier (CB-519). Multiple
* daemon processes may run on one host, so the contract is <em>host-unique</em>, not merely
* process-unique: the herdr-backed launcher mints a fresh UUID per spawn, and a non-herdr
* launcher is likewise expected to return an identifier that cannot collide across processes
* on the same host. This id is the routing key and is deliberately decoupled from any launcher
* transport coordinate (e.g. a herdr pane id), which stays launcher-private. Guaranteed to be
* non-null and unique among live peers on the host.
*/
String id();
@@ -26,4 +30,59 @@ public interface PeerHandle {
default String terminalId() {
return null;
}
/**
* The worker profile that spawned this peer, if the launcher resolved one. A launcher that
* performs dynamic profile selection (e.g. CB-518 weighted placement) sets this so the
* session registry records the actual profile rather than the requested/default one.
*
* @return the profile name, or {@code null} when the launcher leaves it unspecified
*/
default String profile() {
return null;
}
/**
* The bridge's logical name for this session, as assigned at spawn
* ({@link SpawnRequest#sessionName()}). Stable across restarts and meaningful to an operator,
* unlike the transport identifiers above; a launch with no name leaves the peer's display
* identity to the launcher to derive.
*
* @return the bridge-assigned logical session name, or {@code null} if none was assigned
*/
default String sessionName() {
return null;
}
/**
* The peer's OWN session id — the handle that resumes this conversation later (the id a
* later {@link Capability#SESSION_RESUME resume} spawn would pass back). Null when the adapter
* cannot determine it — the contract for an adapter that declines
* {@link Capability#SESSION_RESUME}; an adapter that declares that capability returns this
* non-null for a spawn that requested session identity, because it knows the id before the
* peer has written anything.
*
* <p>Deliberately not a {@code default} (CB-584, the same fix CB-571 made for
* {@link #charterReceipt()} one method below): a decorator that forgets to override this
* silently answers {@code null} for a question it has no basis to answer, and the gap surfaces
* only as a resume that quietly starts a cold session, not a compile error. Every
* implementation must answer explicitly.
*
* @return the peer's own session id, or {@code null} when not determinable
*/
String agentSessionId();
/**
* The charter receipt (CB-571) for this peer's launch — the fingerprint of the exact charter
* bytes it was started with. {@code null} when the launcher records none (a non-instrumented
* adapter, or a launcher before this field); the session registry stores it so the spawn result
* and the roster row can show an operator which charter a member actually got.
*
* <p>Deliberately not a {@code default}: a decorator that forgets to override this silently
* answers {@code null} for a question it has no basis to answer, and the gap surfaces only as
* a missing roster field, not a compile error. Every implementation must answer explicitly.
*
* @return the fingerprint, or {@code null} when the launcher carries none
*/
CharterReceipt charterReceipt();
}
@@ -25,6 +25,19 @@ public interface PeerLauncher {
*/
Set<Capability> capabilities();
/**
* The capabilities of the adapter that {@code profileName} resolves to (null/blank → the
* default profile, the same resolution {@link #spawn} uses). Distinct from {@link
* #capabilities()}, which unions every configured adapter: a caller that must know whether
* <em>this</em> profile's backend supports a capability — e.g. {@link Capability#SESSION_RESUME}
* before honoring {@link SpawnRequest#resumeSessionId()} — needs the per-profile answer, not
* the fleet-wide union, or a mixed fleet could OK a resume that lands on a non-supporting
* adapter (CB-584).
*
* @throws IllegalArgumentException if the profile is unknown and no default is configured
*/
Set<Capability> capabilitiesFor(String profileName);
/**
* {@code profileName}/requestedCwd null/blank → default resolution. Returns after the peer
* process is live (env + argv + placement complete). Never returns {@code null}.
@@ -82,4 +95,13 @@ public interface PeerLauncher {
* when safe to do so.
*/
void stop(String id);
/**
* Discard the context of the peer identified by {@code id}. Implementations must bypass normal
* bridge delivery/turn accounting. Unsupported peer kinds return {@code false} without sending
* a guessed command.
*
* @return {@code true} when a reset was sent and its status transition must settle before reuse
*/
boolean clearContext(String id);
}
@@ -1,13 +1,32 @@
package dev.ltms.bridged.peer;
/**
* Parameters for a {@link PeerLauncher#spawn(SpawnRequest)} call — the peer-neutral
* aggregation of what the core knows at delegation time: which profile to use, the caller's
* requested working directory, and the caller's own cwd (to inherit when no other cwd is set).
* What a caller asks for when spawning a peer.
*
* <p>A null or blank {@code profileName} means "use the launcher's default profile."
* A null or blank {@code requestedCwd} means "inherit from config or caller."
* A null {@code callerCwd} means "the request came from the daemon itself (not a primary)."
* @param profileName the {@code profiles:} entry to spawn on; {@code null}/blank ⇒ the caller
* did not choose, and the role's pool supplies the candidates
* @param requestedCwd working directory asked for by the caller ({@code null} ⇒ unset)
* @param callerCwd the caller's own working directory, used when nothing else pins one
* @param sessionName agent session name (CB-547a); {@code null} ⇒ the launcher mints one
* @param resumeSessionId prior agent session to resume; {@code null} ⇒ a fresh session
* @param role the contract this member runs under (CB-557). Picks the tab label and,
* with the profile, the counter its tab number comes from. {@code null} is
* read as {@link MemberRole#DEV} — an unqualified spawn is a unit of work
*/
public record SpawnRequest(String profileName, String requestedCwd, String callerCwd) {
public record SpawnRequest(String profileName, String requestedCwd, String callerCwd,
String sessionName, String resumeSessionId, MemberRole role) {
public SpawnRequest {
role = (role == null) ? MemberRole.DEV : role;
}
public SpawnRequest(String profileName, String requestedCwd, String callerCwd) {
this(profileName, requestedCwd, callerCwd, null, null, null);
}
/** Pre-CB-557 shape: session identity without an explicit role (defaults to {@code dev}). */
public SpawnRequest(String profileName, String requestedCwd, String callerCwd,
String sessionName, String resumeSessionId) {
this(profileName, requestedCwd, callerCwd, sessionName, resumeSessionId, null);
}
}
@@ -0,0 +1,115 @@
package dev.ltms.bridged.placement;
import java.util.LinkedHashMap;
import java.util.Map;
import java.util.Objects;
import java.util.OptionalLong;
import java.util.concurrent.ConcurrentHashMap;
import java.util.function.LongSupplier;
/**
* Where a credential (not a profile — see {@code BridgedConfig.Profile#effectiveCredentialId()})
* sits out a cooldown after a {@code BACKEND_EXHAUSTED} classification (CB-578 stage B), so a fresh
* spawn does not walk straight back onto the account that just refused on a usage limit.
*
* <p>Keyed by credential id, never by profile name: two profiles sharing one credential (e.g. two
* models on the same OpenAI account) share one quarantine — {@link #quarantine} one credential id
* and every profile whose {@code effectiveCredentialId()} equals it is quarantined too, without this
* class knowing anything about profiles at all. That mapping is the caller's job (see
* {@code CompositePeerLauncher} and {@code dev.ltms.bridged.inject.ExhaustionSink}).
*
* <p>The clock is injected ({@link LongSupplier}, conventionally {@code System::nanoTime} like
* {@code FleetHealthMonitor}), never read inline, so a quarantine's expiry is testable without a
* real sleep.
*/
public final class BackendQuarantine {
private final ConcurrentHashMap<String, Long> quarantinedUntilNanos = new ConcurrentHashMap<>();
private final LongSupplier nowNanos;
private final long cooldownNanos;
/** True only for {@link #none()}. See {@link #quarantine} for why this exists. */
private final boolean inert;
/**
* @param nowNanos monotonic clock, injected for testability
* @param cooldownNanos how long a fresh {@link #quarantine} call blocks the credential for;
* must be positive
*/
public BackendQuarantine(LongSupplier nowNanos, long cooldownNanos) {
this(nowNanos, cooldownNanos, false);
}
private BackendQuarantine(LongSupplier nowNanos, long cooldownNanos, boolean inert) {
this.nowNanos = Objects.requireNonNull(nowNanos, "nowNanos");
if (cooldownNanos <= 0) {
throw new IllegalArgumentException("cooldownNanos must be positive: " + cooldownNanos);
}
this.cooldownNanos = cooldownNanos;
this.inert = inert;
}
/**
* Inert quarantine — {@link #quarantine} does nothing on this instance, so nothing is ever
* quarantined. The explicit stand-in a caller (or a test not exercising this feature) passes
* instead of a defaulting overload, exactly like {@code ExhaustedPatternLookup.none()}.
*/
public static BackendQuarantine none() {
return new BackendQuarantine(() -> 0L, 1, true);
}
/**
* Quarantine {@code credentialId} for the configured cooldown, starting now. A repeat call while
* already quarantined restarts the cooldown at full length — a fresh refusal is fresh evidence the
* account is still exhausted, not a reason to let an earlier, shorter wait stand.
*
* <p>On {@link #none()} this is a no-op. It has to be: that instance holds a clock frozen at 0,
* so recording a deadline would produce a quarantine that never expires — a credential locked out
* for the life of the daemon. Two production {@code CompositePeerLauncher} constructors default to
* {@code none()}, so the failure would be silent and permanent. An inert stand-in must omit the
* fact, never invent one.
*/
public void quarantine(String credentialId) {
Objects.requireNonNull(credentialId, "credentialId");
if (inert) {
return;
}
quarantinedUntilNanos.put(credentialId, nowNanos.getAsLong() + cooldownNanos);
}
/** Whether {@code credentialId} is quarantined right now. */
public boolean isQuarantined(String credentialId) {
return remainingNanos(credentialId) > 0;
}
/** Seconds left on {@code credentialId}'s quarantine, or empty when it is not quarantined. */
public OptionalLong remainingSeconds(String credentialId) {
long remaining = remainingNanos(credentialId);
return remaining > 0 ? OptionalLong.of(toSecondsRoundedUp(remaining)) : OptionalLong.empty();
}
/**
* Every currently-quarantined credential id and its remaining seconds (CB-578 stage B fleet
* reporting) — expired entries are never included. Not pruned from the backing map here: it stays
* small (bounded by the number of distinct credentials ever exhausted) and a lazily-stale entry is
* harmless, since every read already checks the deadline.
*/
public Map<String, Long> activeRemainingSeconds() {
Map<String, Long> out = new LinkedHashMap<>();
quarantinedUntilNanos.forEach((credentialId, deadline) -> {
long remaining = deadline - nowNanos.getAsLong();
if (remaining > 0) {
out.put(credentialId, toSecondsRoundedUp(remaining));
}
});
return out;
}
private long remainingNanos(String credentialId) {
Long deadline = quarantinedUntilNanos.get(credentialId);
return deadline == null ? 0L : deadline - nowNanos.getAsLong();
}
private static long toSecondsRoundedUp(long nanos) {
return (nanos + 999_999_999L) / 1_000_000_000L;
}
}
@@ -0,0 +1,66 @@
package dev.ltms.bridged.placement;
/**
* Backward-compatible placement: an unqualified spawn always resolves to the configured default
* profile, exactly as {@code CompositePeerLauncher} did before CB-518. This ignores caps and
* reachability so that a pre-existing config behaves identically after upgrade.
*
* <p>Two exceptions walk past the default instead of returning it unconditionally:
* <ul>
* <li>Quarantine (CB-578 stage B): a quarantined default is a credential that just refused on
* a usage limit, not a transient capacity or reachability concern.
* <li>Weight 0 (CB-554): {@code fixed} is still automatic selection, so a profile the operator
* marked "never auto-select me" ({@code weight <= 0}) must be skipped here exactly as
* {@code weighted}/{@code round-robin} skip it — an explicit {@code bridge_spawn} naming
* the profile is unaffected, only this automatic fallback walk.
* </ul>
* A fleet where nothing is ever quarantined or weight-0 never exercises either path, so today's
* behaviour is unchanged.
*/
final class FixedPlacementPolicy implements PlacementPolicy {
@Override
public PlacementCandidate select(PlacementContext ctx) {
String d = ctx.defaultProfile();
if (d != null && !d.isBlank() && !ctx.quarantined().contains(d) && !weightExcluded(ctx, d)) {
return new PlacementCandidate(d, null, 1.0f, null);
}
for (PlacementCandidate c : ctx.candidates()) {
if (!ctx.quarantined().contains(c.profile()) && !c.excluded()) {
return new PlacementCandidate(c.profile(), null, c.weight(), c.maxLoad());
}
}
if (d != null && !d.isBlank()) {
boolean dQuarantined = ctx.quarantined().contains(d);
boolean dWeightExcluded = weightExcluded(ctx, d);
if (dQuarantined && dWeightExcluded) {
throw new PlacementException("worker profile '" + d + "' is quarantined (backend "
+ "exhausted) and has weight 0 (excluded from automatic selection), and no "
+ "available candidate remains");
}
if (dWeightExcluded) {
throw new PlacementException("worker profile '" + d + "' has weight 0 (excluded "
+ "from automatic selection) and no available candidate remains");
}
if (dQuarantined) {
throw new PlacementException("worker profile '" + d + "' is quarantined (backend "
+ "exhausted) and no un-quarantined candidate is available");
}
}
if (!ctx.candidates().isEmpty()) {
throw new PlacementException(
"all worker profiles are excluded from automatic selection (quarantined or weight-0)");
}
throw new PlacementException("no worker profiles configured");
}
/** Whether {@code profile} carries {@code weight <= 0} (CB-554) among {@code ctx}'s candidates. */
private static boolean weightExcluded(PlacementContext ctx, String profile) {
for (PlacementCandidate c : ctx.candidates()) {
if (c.profile().equals(profile)) {
return c.excluded();
}
}
return false;
}
}
@@ -0,0 +1,34 @@
package dev.ltms.bridged.placement;
/**
* A profile (and, in CB-308, a host) that can be chosen by a {@link PlacementPolicy}.
*
* <p>Keeping this as a small descriptor rather than a bare profile name lets CB-308 widen
* selection to {@code (host, profile)} pairs without changing the policy interface.
*/
public record PlacementCandidate(String profile, String host, float weight, Integer maxLoad) {
/** A candidate with no explicit host (the single-host default) and the given weight/cap. */
public static PlacementCandidate profile(String profile, float weight, Integer maxLoad) {
return new PlacementCandidate(profile, null, weight, maxLoad);
}
/** A candidate with no explicit host, unit weight, and no cap. */
public static PlacementCandidate profile(String profile) {
return new PlacementCandidate(profile, null, 1.0f, null);
}
/**
* True when this candidate carries an explicit {@code weight <= 0} (CB-554) and must be
* skipped by every automatic policy — the same way a quarantined or unreachable candidate is
* skipped. {@code BridgedConfig.Profile}'s compact constructor already normalises "absent" to
* {@code 1.0} and "negative" to {@code 0.0}, so this is a plain threshold check here; it does
* not need to distinguish "explicit 0" from "absent" itself.
*
* <p>Exclusion is about <em>automatic</em> selection only — an explicit
* {@code bridge_spawn{profile:"..."}} bypasses placement entirely and is unaffected.
*/
public boolean excluded() {
return weight <= 0.0f;
}
}
@@ -0,0 +1,23 @@
package dev.ltms.bridged.placement;
import java.util.List;
import java.util.Set;
import java.util.function.Function;
/**
* Everything a {@link PlacementPolicy} needs to make one selection.
*
* @param defaultProfile profile a {@code fixed} policy should return (may be {@code null})
* @param candidates every configured candidate; the policy filters out those at cap or unreachable
* @param liveCount current live worker count per profile (from the session registry)
* @param unreachable profiles already known to have failed in this spawn attempt
* @param quarantined profiles whose credential is currently quarantined (CB-578 stage B) — a
* {@code BACKEND_EXHAUSTED} classification put it, or a profile it shares a
* credential with, on cooldown. Filtered the same way as {@code unreachable}.
*/
public record PlacementContext(String defaultProfile,
List<PlacementCandidate> candidates,
Function<String, Integer> liveCount,
Set<String> unreachable,
Set<String> quarantined) {
}
@@ -0,0 +1,12 @@
package dev.ltms.bridged.placement;
/**
* Thrown when a {@link PlacementPolicy} has no candidate available. Kept as a distinct type so
* callers can distinguish "no capacity" from a spawn-time transport failure.
*/
public final class PlacementException extends IllegalStateException {
public PlacementException(String message) {
super(message);
}
}
@@ -0,0 +1,41 @@
package dev.ltms.bridged.placement;
/**
* Factory for the built-in placement policies.
*/
public final class PlacementPolicies {
private PlacementPolicies() {
}
/**
* Resolve a policy name from config. Absent/blank values and {@code "fixed"} return the
* backward-compatible fixed policy; unknown names throw.
*/
public static PlacementPolicy fromName(String name) {
String n = (name == null) ? "" : name.toLowerCase();
if (n.isBlank() || "fixed".equals(n)) {
return fixed();
}
if ("weighted".equals(n)) {
return weighted();
}
if ("round-robin".equals(n)) {
return roundRobin();
}
throw new IllegalArgumentException("unknown placement policy '" + name
+ "' — must be one of: fixed, round-robin, weighted");
}
public static PlacementPolicy fixed() {
return new FixedPlacementPolicy();
}
public static PlacementPolicy weighted() {
return new WeightedRoundRobinPolicy();
}
public static PlacementPolicy roundRobin() {
return new RoundRobinPlacementPolicy();
}
}
@@ -0,0 +1,16 @@
package dev.ltms.bridged.placement;
/**
* How {@code bridged} chooses a worker profile when a spawn names none. Implementations are
* deterministic and unit-testable; the caller (the composite launcher) handles failover retries.
*/
public interface PlacementPolicy {
/**
* Pick one candidate from the configured set.
*
* @throws java.lang.IllegalStateException when no candidate is available, with a message naming
* whether every profile is at capacity or unreachable
*/
PlacementCandidate select(PlacementContext ctx);
}
@@ -0,0 +1,85 @@
package dev.ltms.bridged.placement;
import java.util.ArrayList;
import java.util.List;
/**
* Shared filtering and empty-set reporting used by the built-in placement policies.
*/
final class PlacementPolicyUtil {
private PlacementPolicyUtil() {
}
/**
* Candidates that are not weight-excluded (CB-554: explicit {@code weight <= 0}, checked
* first because it is a static config choice rather than transient state), not
* known-unreachable, not quarantined (CB-578 stage B), and have not reached their maxLoad.
* A {@code null} maxLoad means unlimited.
*/
static List<PlacementCandidate> available(PlacementContext ctx) {
List<PlacementCandidate> out = new ArrayList<>();
for (PlacementCandidate c : ctx.candidates()) {
if (c.excluded() || ctx.unreachable().contains(c.profile())
|| ctx.quarantined().contains(c.profile())) {
continue;
}
Integer cap = c.maxLoad();
if (cap != null) {
int live = ctx.liveCount().apply(c.profile());
if (live >= cap) {
continue;
}
}
out.add(c);
}
return out;
}
/**
* Build a clear exception describing why every candidate was dropped: all weight-0, all
* quarantined, all at capacity, all unreachable, or a mix. Each candidate is counted into
* exactly one bucket (weight-excluded takes priority) so a candidate excluded for more than
* one reason is never double-counted.
*/
static PlacementException emptyException(PlacementContext ctx) {
int weightExcluded = 0;
int atCap = 0;
int unreachable = 0;
int quarantined = 0;
for (PlacementCandidate c : ctx.candidates()) {
Integer cap = c.maxLoad();
if (c.excluded()) {
weightExcluded++;
} else if (ctx.quarantined().contains(c.profile())) {
quarantined++;
} else if (ctx.unreachable().contains(c.profile())) {
unreachable++;
} else if (cap != null && ctx.liveCount().apply(c.profile()) >= cap) {
atCap++;
}
}
int total = ctx.candidates().size();
if (total == 0) {
return new PlacementException("no worker profiles configured");
}
if (weightExcluded == total) {
return new PlacementException(
"all worker profiles have weight 0 (excluded from automatic selection)");
}
if (quarantined == total) {
return new PlacementException("all worker profiles are quarantined (backend exhausted)");
}
if (atCap == total) {
return new PlacementException("all worker profiles are at maxLoad");
}
if (unreachable == total) {
return new PlacementException("all worker profiles are unreachable");
}
return new PlacementException("no worker profile available: " + atCap + " at maxLoad, "
+ unreachable + " unreachable, " + quarantined + " quarantined, "
+ weightExcluded + " weight-0, "
+ (total - atCap - unreachable - quarantined - weightExcluded) + " remaining");
}
}
@@ -0,0 +1,23 @@
package dev.ltms.bridged.placement;
import java.util.List;
import java.util.concurrent.atomic.AtomicInteger;
/**
* Deterministic round-robin over the profiles that still have capacity and are not known to be
* unreachable. The index advances only on successful selections so the distribution stays even
* across spawns.
*/
final class RoundRobinPlacementPolicy implements PlacementPolicy {
private final AtomicInteger index = new AtomicInteger(0);
@Override
public synchronized PlacementCandidate select(PlacementContext ctx) {
List<PlacementCandidate> available = PlacementPolicyUtil.available(ctx);
if (available.isEmpty()) {
throw PlacementPolicyUtil.emptyException(ctx);
}
return available.get(index.getAndIncrement() % available.size());
}
}
@@ -0,0 +1,50 @@
package dev.ltms.bridged.placement;
import java.util.List;
import java.util.Map;
import java.util.concurrent.ConcurrentHashMap;
/**
* Smooth weighted round-robin (nginx-style): for each selection, add the candidate's weight to
* its current score, pick the highest score, then subtract the total weight of all available
* candidates from the winner. Weights are not required to sum to 1.0; only their ratios matter.
*
* <p>The state is per-policy instance and protected by {@code synchronized} so concurrent spawns
* see a consistent, deterministic sequence rather than interleaving updates.
*/
final class WeightedRoundRobinPolicy implements PlacementPolicy {
private final Map<String, Double> current = new ConcurrentHashMap<>();
@Override
public synchronized PlacementCandidate select(PlacementContext ctx) {
List<PlacementCandidate> available = PlacementPolicyUtil.available(ctx);
if (available.isEmpty()) {
throw PlacementPolicyUtil.emptyException(ctx);
}
double total = 0.0;
for (PlacementCandidate c : available) {
total += c.weight();
}
if (total <= 0.0) {
throw new PlacementException("all available profiles have non-positive weight");
}
PlacementCandidate best = null;
double bestScore = Double.NEGATIVE_INFINITY;
for (PlacementCandidate c : available) {
double score = current.merge(c.profile(), (double) c.weight(), (old, add) -> old + add);
if (score > bestScore) {
bestScore = score;
best = c;
}
}
if (best == null) {
throw new PlacementException("no placement candidate could be selected");
}
current.put(best.profile(), current.get(best.profile()) - total);
return best;
}
}
@@ -11,11 +11,12 @@ import dev.ltms.bridged.metrics.Metrics;
import dev.ltms.bridged.herdr.Agent;
import dev.ltms.bridged.herdr.HerdrClient;
import dev.ltms.bridged.herdr.HerdrException;
import dev.ltms.bridged.inject.WorkerPresence;
import dev.ltms.bridged.inject.MemberPresence;
import dev.ltms.bridged.peer.PeerUnreachableException;
import dev.ltms.bridged.msg.MessageService;
import dev.ltms.bridged.session.SessionManager;
import dev.ltms.bridged.session.WorkerSession;
import dev.ltms.bridged.peer.MemberRole;
import dev.ltms.bridged.session.MemberSession;
import dev.ltms.bridged.session.WorktreeRequest;
import dev.ltms.bridged.peer.PeerLauncher;
import io.javalin.Javalin;
@@ -55,7 +56,7 @@ public final class BridgedApp {
private final PeerLauncher workers;
private final SessionManager sessions; // CB-301: authoritative session registry
private final MessageService messages;
private final WorkerPresence presence; // CB-113: which workers are MCP-connected (available)
private final MemberPresence presence; // CB-113: which workers are MCP-connected (available)
private final HttpServlet mcpServlet; // MCP Streamable-HTTP endpoint, mounted at /mcp (nullable)
private final CallerResolver auth; // CB-501: null → authz not enforced (legacy behaviour)
private final Metrics metrics; // CB-502: null → /metrics not exposed
@@ -67,7 +68,7 @@ public final class BridgedApp {
* behaviour without each needing an auth fixture.
*/
public BridgedApp(HerdrClient herdr, PeerLauncher workers, SessionManager sessions,
MessageService messages, WorkerPresence presence,
MessageService messages, MemberPresence presence,
HttpServlet mcpServlet) {
this(herdr, workers, sessions, messages, presence, mcpServlet, null, null);
}
@@ -79,7 +80,7 @@ public final class BridgedApp {
* the endpoint
*/
public BridgedApp(HerdrClient herdr, PeerLauncher workers, SessionManager sessions,
MessageService messages, WorkerPresence presence,
MessageService messages, MemberPresence presence,
HttpServlet mcpServlet, CallerResolver auth, Metrics metrics) {
this.herdr = herdr;
this.workers = workers;
@@ -115,10 +116,10 @@ public final class BridgedApp {
}
app.get("/sessions", this::sessions);
app.get("/agents", this::agents);
app.get("/workers", this::listWorkers); // CB-304: registry roster + live herdr status
app.get("/profiles", this::profiles); // configured worker profiles
app.post("/workers", this::spawnWorker); // optional ?profile= or {"profile":…}
app.delete("/workers/{paneId}", this::stopWorker);
app.get("/members", this::listMembers); // CB-304: registry roster + live herdr status
app.get("/profiles", this::profiles); // configured backend profiles
app.post("/members", this::spawnMember); // optional ?role=&profile= or {"role":…,"profile":…}
app.delete("/members/{paneId}", this::stopMember);
app.post("/sessions/{id}/message", this::sendMessage); // bridge_send (primary; blocking, wait:false, or answer via turnId)
app.post("/sessions/{id}/reply", this::replyMessage); // bridge_reply (worker)
app.get("/sessions/{id}/replies", this::drainReplies); // drain reply inbox (CB-307)
@@ -221,16 +222,17 @@ public final class BridgedApp {
}
/** CB-304: bridge-owned roster merged with live herdr status by paneId. */
private void listWorkers(Context ctx) {
private void listMembers(Context ctx) {
if (!allow(ctx, Authz.Action.READ, null)) {
return;
}
// CB-519: the registry key is a host-unique id, not the pane coordinate — join on terminal.
Map<String, Agent> live = workers.list().stream()
.map(Agent.class::cast)
.filter(a -> a.paneId() != null)
.collect(Collectors.toMap(Agent::paneId, Function.identity(), (_, b) -> b));
.filter(a -> a.terminalId() != null)
.collect(Collectors.toMap(Agent::terminalId, Function.identity(), (_, b) -> b));
List<Map<String, Object>> out = sessions.roster().stream()
.map(s -> SessionManager.rosterView(s, live.get(s.paneId())))
.map(s -> SessionManager.rosterView(s, live.get(s.terminalId())))
.toList();
ctx.status(200).json(Map.of("workers", out));
}
@@ -250,10 +252,11 @@ public final class BridgedApp {
* body) picks which configured profile; omitted → the default. 403 if the base_url would breach
* the subscription boundary, 400 for an unknown profile.
*/
private void spawnWorker(Context ctx) {
private void spawnMember(Context ctx) {
if (!allow(ctx, Authz.Action.SPAWN, null)) {
return;
}
String role = ctx.queryParam("role");
String profile = ctx.queryParam("profile");
String cwd = ctx.queryParam("cwd");
String worktree = ctx.queryParam("worktree");
@@ -264,6 +267,7 @@ public final class BridgedApp {
String body = ctx.body();
if (!body.isBlank()) {
JsonNode b = mapper.readTree(body);
if (role == null || role.isBlank()) role = b.path("role").asText(null);
if (profile == null || profile.isBlank()) profile = b.path("profile").asText(null);
if (cwd == null || cwd.isBlank()) cwd = b.path("cwd").asText(null);
if (worktree == null || worktree.isBlank()) worktree = b.path("worktree").asText(null);
@@ -274,10 +278,18 @@ public final class BridgedApp {
}
}
WorktreeRequest wt = worktreeRequest(worktree, ticket);
MemberRole memberRole;
try {
memberRole = (role == null || role.isBlank()) ? MemberRole.DEV : MemberRole.parse(role);
} catch (IllegalArgumentException e) {
ctx.status(400).json(Map.of("error", "unknown_role", "detail", e.getMessage()));
return;
}
try {
// No MCP caller over REST, so callerCwd and ownerTerminal are null.
WorkerSession worker = sessions.acquire(blankToNull(profile), blankToNull(cwd), null, null, wt);
ctx.status(201).json(view(worker));
MemberSession member = sessions.acquire(blankToNull(profile), memberRole,
blankToNull(cwd), null, null, wt);
ctx.status(201).json(view(member));
} catch (GuardException e) {
ctx.status(403).json(Map.of("error", "subscription_boundary", "detail", e.getMessage()));
} catch (IllegalArgumentException e) {
@@ -305,7 +317,7 @@ public final class BridgedApp {
}
/** Tear a worker down by pane id. */
private void stopWorker(Context ctx) {
private void stopMember(Context ctx) {
String paneId = ctx.pathParam("paneId");
if (!allow(ctx, Authz.Action.STOP, paneId)) {
return;
@@ -391,9 +403,12 @@ public final class BridgedApp {
case TIMED_OUT_QUEUED -> "queued";
case BUSY -> "busy";
case WORKER_FAILED -> "failed";
case BACKEND_EXHAUSTED -> "backend_exhausted";
default -> "done"; // unreachable (terminal outcomes handled above)
},
"detail", reply.outcome() == MessageService.Outcome.WORKER_FAILED && reply.text() != null
"detail", (reply.outcome() == MessageService.Outcome.WORKER_FAILED
|| reply.outcome() == MessageService.Outcome.BACKEND_EXHAUSTED)
&& reply.text() != null
? reply.text()
: "no reply within " + timeout + "ms; poll status or retry"));
}
@@ -543,7 +558,7 @@ public final class BridgedApp {
}
/** CB-301 projection of an authoritative bridge-owned session. */
private static Map<String, Object> view(WorkerSession s) {
private static Map<String, Object> view(MemberSession s) {
Map<String, Object> m = new LinkedHashMap<>();
m.put("terminalId", s.terminalId());
m.put("paneId", s.paneId());
@@ -13,6 +13,8 @@ import java.nio.file.Path;
import java.nio.file.StandardCopyOption;
import java.security.SecureRandom;
import java.util.List;
import java.util.Map;
import java.util.Optional;
import java.util.concurrent.TimeUnit;
import java.util.concurrent.atomic.AtomicLong;
import java.util.stream.Collectors;
@@ -27,6 +29,54 @@ public final class GitWorktrees implements Worktrees {
private static final Logger log = LoggerFactory.getLogger(GitWorktrees.class);
/** Project-level MCP config. Present in the repo, so every worktree would otherwise inherit the
* primary's IDE server mounts (a CB-523 worker edited the primary checkout; see the isolation
* javadoc). Neutralized unconditionally. */
private static final String MCP_CONFIG = ".mcp.json";
/** What {@link #isolateToolSurface} writes for {@code .mcp.json}: a valid, explicitly empty server map. */
private static final String NEUTRAL_MCP_CONFIG = "{\n \"mcpServers\": {}\n}\n";
/** OpenCode's repo-level config. Tracked here, so it lands in every worktree, and it mounts the
* primary's gitea and context7 servers with the primary's credentials. Neutralized so the worker
* gets only the config its launcher writes via {@code OPENCODE_CONFIG}.
*
* <p>The reason has changed shape and is now stronger. It used to be a crash: the file carried
* {@code {file:.secrets/...}} references to gitignored files that never reached a worktree, and
* opencode refuses to start on a dangling reference (CB-543). Those credentials now live in one
* shell-level store and the file reads them as {@code {env:...}}, so in a worktree the reference
* resolves instead of failing. That is worse, not better: a member would silently inherit the
* primary's admin-scoped {@code GITEA_ACCESS_TOKEN}. A loud crash became a quiet privilege leak,
* so this entry protects a boundary now rather than papering over a startup error. */
private static final String OPENCODE_CONFIG = "opencode.json";
/** What {@link #isolateToolSurface} writes for {@code opencode.json}: a valid, empty JSON object. */
private static final String NEUTRAL_OPENCODE_CONFIG = "{}\n";
/** Autoenv's repo-level config. Not tracked today, but re-landing it must stay safe: autoenv
* authorizes by path, so a fresh worktree path is always unauthorized and its interactive prompt
* would block every spawn — neutralize it so it can never be committed. */
private static final String AUTOENV_CONFIG = ".autoenv";
/** What {@link #isolateToolSurface} writes for {@code .autoenv}: a valid, empty env file. */
private static final String NEUTRAL_AUTOENV_CONFIG = "";
/**
* A tracked project config that is hostile in a provisioned worktree, and what to replace it
* with. {@link #file} is the repo-relative path; {@link #stub} is a neutral but VALID payload for
* that file's format — a malformed stub would only trade one crash for another;
* {@link #createIfAbsent} keeps {@code .mcp.json}'s long-standing behaviour of writing its stub
* even when the repo carries no such file, whereas the others are only touched when present.
*/
private record WorktreeHostileConfig(String file, String stub, boolean createIfAbsent) {}
/** The worktree-hostile configs neutralized in every provisioned worktree, in order. */
private static final List<WorktreeHostileConfig> WORKTREE_HOSTILE_CONFIGS = List.of(
new WorktreeHostileConfig(MCP_CONFIG, NEUTRAL_MCP_CONFIG, true),
new WorktreeHostileConfig(OPENCODE_CONFIG, NEUTRAL_OPENCODE_CONFIG, false),
new WorktreeHostileConfig(AUTOENV_CONFIG, NEUTRAL_AUTOENV_CONFIG, false)
);
private final String configuredRoot;
private final SecureRandom random = new SecureRandom();
private final AtomicLong seq = new AtomicLong();
@@ -55,9 +105,59 @@ public final class GitWorktrees implements Worktrees {
String wt = path.toAbsolutePath().toString();
log.info("adding worktree branch={} path={} base={}", branch, wt, base);
exec("git", "-C", repoRoot, "worktree", "add", wt, "-b", branch, base);
isolateToolSurface(wt);
return wt;
}
/**
* Neutralize the worktree's worktree-hostile project configs so a worker inherits only the tools
* and environment its launcher mounts (the bridge via {@code --mcp-config}, the opencode config
* via {@code OPENCODE_CONFIG}) — never the primary's.
*
* <p>This is unconditional, and it is not the same job as the parity overlay. The repo's own
* committed {@code .mcp.json} declares the primary's IDE servers, so a fresh checkout mounts them
* whether or not the overlay copies anything; a worker that inherits them navigates and edits
* through tools bound to the <em>primary's</em> IntelliJ project, which silently hands it absolute
* paths outside its own worktree. That is not hypothetical: a CB-523 worker made all 59 of its
* edits in the primary checkout while compiling its worktree, so every build it ran was of code
* that did not contain its changes. {@code opencode.json} is the same trap one tool over — tracked,
* so it lands in every worktree, and it mounts gitea and context7 with the primary's own
* credentials, which a member must never hold. {@code .autoenv} extends the principle to a
* config that is not tracked today: autoenv authorizes by path, so a fresh worktree path is always
* unauthorized and its interactive prompt would block every spawn, so re-landing one must be safe.
*
* <p>Where a config exists it is replaced by a valid neutral stub (an explicitly empty
* map/object, or an empty env file — never a deletion, which would still let a later
* {@code git checkout} restore the hostile copy). The {@code --skip-worktree} bit keeps the
* neutralized copy from ever showing up as a local modification the worker might commit. A config
* the repo does not carry is skipped silently — no stub is invented for a file the repo does not
* have, and one missing file must never fail provisioning.
*/
private void isolateToolSurface(String worktreePath) {
Path root = Path.of(worktreePath).toAbsolutePath().normalize();
for (WorktreeHostileConfig cfg : WORKTREE_HOSTILE_CONFIGS) {
neutralize(root, worktreePath, cfg);
}
}
private void neutralize(Path root, String worktreePath, WorktreeHostileConfig cfg) {
Path target = root.resolve(cfg.file());
if (!Files.exists(target) && !cfg.createIfAbsent()) {
log.debug("{} absent in the worktree — skipping (repo does not carry it)", cfg.file());
return;
}
try {
Files.writeString(target, cfg.stub());
} catch (IOException e) {
throw new WorktreeException("cannot neutralize " + cfg.file() + " in the worktree: "
+ e.getMessage(), e);
}
if (isTracked(root, cfg.file())) {
exec("git", "-C", worktreePath, "update-index", "--skip-worktree", cfg.file());
}
log.debug("neutralized {} — worker tool surface is launcher-mounted only", cfg.file());
}
@Override
public void remove(String repoRoot, String worktreePath) {
Path p = Path.of(worktreePath);
@@ -69,6 +169,22 @@ public final class GitWorktrees implements Worktrees {
exec("git", "-C", repoRoot, "worktree", "remove", "--force", worktreePath);
}
@Override
public boolean hasUncommitted(String worktreePath) {
// A worktree that is already gone holds no work to lose, and it must not break teardown:
// git -C <missing-dir> status exits non-zero and would throw where release() is mid-way
// through stopping a pane. Mirror remove()'s already-gone tolerance by treating it as clean.
Path p = Path.of(worktreePath);
if (!Files.exists(p)) {
log.debug("worktree {} already gone — nothing can be uncommitted", worktreePath);
return false;
}
// No --untracked-files=no: the exact shape of the work lost in CB-576 was a new file
// that was never added, so an untracked-only worktree is still dirty.
String out = exec("git", "-C", worktreePath, "status", "--porcelain");
return !out.isBlank();
}
@Override
public void overlayParity(String repoRoot, String worktreePath, List<String> overlay) {
if (overlay == null || overlay.isEmpty()) {
@@ -103,6 +219,86 @@ public final class GitWorktrees implements Worktrees {
return Path.of(out.trim()).toAbsolutePath().normalize().toString();
}
/**
* CB-578 stage C, fixed by CB-587. Stages into a <em>temporary</em> index (never the worktree's
* real one, which the worker may still be writing to) — but that temp index is first <em>seeded</em>
* from the worktree's real one, rather than starting empty:
*
* <pre>
* cp $(git -C worktree rev-parse --git-path index) &lt;temp&gt;
* GIT_INDEX_FILE=&lt;temp&gt; git -C worktree add -A
* tree=$(GIT_INDEX_FILE=&lt;temp&gt; git -C worktree write-tree)
* commit=$(git -C worktree commit-tree $tree -p HEAD -m message)
* git -C worktree update-ref refs/wip/branch $commit
* </pre>
*
* A fresh empty index carries none of the real index's {@code --skip-worktree} /
* {@code --assume-unchanged} bits, so {@code add -A} into it stages a skip-worktree file's local
* on-disk content even though {@code git status --porcelain} correctly hides that file (CB-587).
* Seeding from the real index preserves those bits, so {@code add -A} then skips exactly what
* {@code git status} skips. {@code add -A} (never {@code -f}) also still respects
* {@code .gitignore} exactly as it would in the real index — a gitignored file staying ignored is
* what keeps secrets and local config out of the snapshot's tree. The temporary index file is
* removed afterwards regardless of outcome; the worker's real index is never opened for writing.
*/
@Override
public Optional<String> snapshot(String worktreePath, String branch, String message) {
if (!Files.exists(Path.of(worktreePath))) {
log.debug("worktree {} already gone — nothing to snapshot", worktreePath);
return Optional.empty();
}
Path tempIndex;
try {
tempIndex = Files.createTempFile("bridged-wip-index-", ".tmp");
} catch (IOException e) {
throw new WorktreeException("cannot create a temporary index for snapshot: " + e.getMessage(), e);
}
Map<String, String> indexEnv = Map.of("GIT_INDEX_FILE", tempIndex.toAbsolutePath().toString());
try {
Path realIndex = resolveRealIndex(worktreePath);
try {
Files.copy(realIndex, tempIndex, StandardCopyOption.REPLACE_EXISTING);
} catch (IOException e) {
throw new WorktreeException("cannot copy the worktree's real index (" + realIndex
+ ") into the temporary snapshot index: " + e.getMessage(), e);
}
exec(indexEnv, "git", "-C", worktreePath, "add", "-A");
String tree = exec(indexEnv, "git", "-C", worktreePath, "write-tree").trim();
String commit = exec("git", "-C", worktreePath, "commit-tree", tree, "-p", "HEAD", "-m", message).trim();
exec("git", "-C", worktreePath, "update-ref", "refs/wip/" + branch, commit);
log.info("snapshotted worktree {} to refs/wip/{} commit={}", worktreePath, branch, commit);
return Optional.of(commit);
} finally {
try {
Files.deleteIfExists(tempIndex);
} catch (IOException e) {
log.debug("could not delete temporary snapshot index {}: {}", tempIndex, e.getMessage());
}
}
}
/**
* Resolve the path of {@code worktreePath}'s real index. Never assume {@code <worktree>/.git/index}:
* in a linked worktree {@code .git} is a <em>file</em> pointing at the main repo's
* {@code worktrees/<name>/} directory, and that is where the real per-worktree index lives.
* {@code git rev-parse --git-path index} resolves this correctly for both a linked worktree and
* the main checkout. Throws {@link WorktreeException} — same as every other failure in this
* class — if the command fails or the resolved path does not exist, rather than silently
* snapshotting from an empty index.
*/
private Path resolveRealIndex(String worktreePath) {
String out = exec("git", "-C", worktreePath, "rev-parse", "--git-path", "index").trim();
Path index = Path.of(out);
if (!index.isAbsolute()) {
index = Path.of(worktreePath).resolve(index).normalize();
}
if (!Files.exists(index)) {
throw new WorktreeException("worktree's real index not found at resolved path " + index
+ " (git rev-parse --git-path index reported '" + out + "')");
}
return index;
}
/** Resolve the directory that will hold per-session worktree checkouts. */
private Path resolveRoot(String repoRoot) {
if (configuredRoot != null && !configuredRoot.isBlank()) {
@@ -125,11 +321,20 @@ public final class GitWorktrees implements Worktrees {
* stdout and stderr (merged by redirectErrorStream).
*/
private String exec(String... command) {
return exec(Map.of(), command);
}
/** Same as {@link #exec(String...)}, with extra environment variables set on the child process. */
private String exec(Map<String, String> extraEnv, String... command) {
String out;
int code;
Process p;
try {
p = new ProcessBuilder(command).redirectErrorStream(true).start();
ProcessBuilder pb = new ProcessBuilder(command).redirectErrorStream(true);
if (extraEnv != null && !extraEnv.isEmpty()) {
pb.environment().putAll(extraEnv);
}
p = pb.start();
} catch (IOException e) {
throw new WorktreeException("failed to start " + command[0] + ": " + e.getMessage(), e);
}
@@ -0,0 +1,92 @@
package dev.ltms.bridged.session;
import dev.ltms.bridged.peer.CharterReceipt;
import dev.ltms.bridged.peer.MemberRole;
/**
* A bridge-owned worker session — the authoritative in-daemon record of a worker this
* process spawned. Immutable; state transitions are performed by replacing the record in
* {@link SessionManager}'s registry.
*
* @param paneId the host-unique opaque id (CB-519) — the registry key and the argument to
* teardown. Despite the historical name this is the {@link
* dev.ltms.bridged.peer.PeerHandle#id()}, a UUID, and is distinct from the
* launcher-private herdr pane coordinate.
* @param terminalId herdr terminal handle — the {@code target} for send/read/status
* @param profile the profile name that spawned this session — WHICH BACKEND
* @param role the contract this member runs under — WHAT IT IS FOR. Independent of
* {@code profile}: a reviewer may run on the same profile as the dev
* whose diff it reads. {@code null} only for a session recorded before
* the role was known
* @param cwd the resolved working directory the worker started in
* @param ownerTerminal the caller that requested this worker ({@code null} = daemon/anon)
* @param spawnedAtNanos {@link System#nanoTime()} when the session was registered
* @param lastActivityAtNanos {@link System#nanoTime()} of the most recent lifecycle event
* @param turnCount number of delegated turns that have been delivered to this session
* @param state current lifecycle state in the one-shot FSM
* @param charterReceipt the fingerprint (CB-571) of the charter bytes this member was started
* with; {@code null} for a session whose launcher recorded none
* @param agentSessionId the peer's OWN session id (CB-584) — the handle a later
* {@code resumeSessionId} spawn would pass back to resume this exact
* conversation; {@code null} when the launcher could not determine one
* (an adapter that declines {@code Capability.SESSION_RESUME}, or one
* that resolves it lazily and has not yet)
*/
public record MemberSession(
String paneId,
String terminalId,
String profile,
MemberRole role,
String cwd,
String ownerTerminal,
long spawnedAtNanos,
long lastActivityAtNanos,
int turnCount,
State state,
String worktree,
String branch,
CharterReceipt charterReceipt,
String agentSessionId) {
/** One-shot worker lifecycle states. */
public enum State {
SPAWNING,
READY,
BUSY,
DONE,
FAILED,
RELEASED
}
/**
* Backward-compatible shape: a session with no charter receipt and no agent session id (a test
* or a launcher before CB-571 / CB-584). A separate constructor rather than a new parameter on
* the canonical one, so existing call sites that have nothing to record keep compiling
* unchanged.
*/
public MemberSession(String paneId, String terminalId, String profile, MemberRole role,
String cwd, String ownerTerminal, long spawnedAtNanos,
long lastActivityAtNanos, int turnCount, State state,
String worktree, String branch) {
this(paneId, terminalId, profile, role, cwd, ownerTerminal, spawnedAtNanos,
lastActivityAtNanos, turnCount, state, worktree, branch, null, null);
}
/** Return a copy of this session in {@code state}. */
public MemberSession withState(State state) {
return new MemberSession(paneId, terminalId, profile, role, cwd, ownerTerminal, spawnedAtNanos,
lastActivityAtNanos, turnCount, state, worktree, branch, charterReceipt, agentSessionId);
}
/** Return a copy with {@code lastActivityAtNanos} updated to {@code nowNanos}. */
public MemberSession withActivity(long nowNanos) {
return new MemberSession(paneId, terminalId, profile, role, cwd, ownerTerminal, spawnedAtNanos,
nowNanos, turnCount, state, worktree, branch, charterReceipt, agentSessionId);
}
/** Return a copy with the turn count incremented and activity timestamped at {@code nowNanos}. */
public MemberSession bumpTurn(long nowNanos) {
return new MemberSession(paneId, terminalId, profile, role, cwd, ownerTerminal, spawnedAtNanos,
nowNanos, turnCount + 1, state, worktree, branch, charterReceipt, agentSessionId);
}
}
@@ -1,8 +1,12 @@
package dev.ltms.bridged.session;
import dev.ltms.bridged.auth.MemberLifecycle;
import dev.ltms.bridged.herdr.Agent;
import dev.ltms.bridged.inject.TurnListener;
import dev.ltms.bridged.inject.WorkerPresence;
import dev.ltms.bridged.inject.MemberPresence;
import dev.ltms.bridged.msg.TurnToken;
import dev.ltms.bridged.peer.Capability;
import dev.ltms.bridged.peer.MemberRole;
import dev.ltms.bridged.peer.PeerHandle;
import dev.ltms.bridged.peer.PeerLauncher;
import dev.ltms.bridged.peer.SpawnRequest;
@@ -14,6 +18,7 @@ import java.util.LinkedHashMap;
import java.util.List;
import java.util.Map;
import java.util.Optional;
import java.util.Set;
import java.util.concurrent.ConcurrentHashMap;
import java.util.concurrent.TimeUnit;
import java.util.concurrent.atomic.AtomicLong;
@@ -27,11 +32,10 @@ import java.util.function.LongSupplier;
* teardown on top.
*
* <p>The state machine is intentionally one-shot / no-reuse: every acquired worker is fresh,
* and a finished or released worker is torn down, never pooled. {@link #recycle} is a convenience
* for {@code release + acquire} with a new distinct pane id.
* and a finished or released worker is torn down, never pooled or reused.
*
* <p>The manager implements {@link TurnListener} so the injector's turn boundaries drive
* {@code READY → BUSY → DONE} (or {@code FAILED}). It exposes a {@link WorkerPresence} view via
* {@code READY → BUSY → DONE} (or {@code FAILED}). It exposes a {@link MemberPresence} view via
* {@link #asPresence()}: any MCP contact from a worker marks it present and simultaneously
* transitions the session {@code SPAWNING → READY}.
*/
@@ -41,145 +45,403 @@ public final class SessionManager implements TurnListener {
private final PeerLauncher launcher;
private final Worktrees worktrees;
private final ConcurrentHashMap<String /*paneId*/, WorkerSession> registry = new ConcurrentHashMap<>();
private final WorkerPresence presence;
private final ConcurrentHashMap<String /*paneId*/, MemberSession> registry = new ConcurrentHashMap<>();
private final MemberPresence presence;
private final SecureRandom nonceRandom = new SecureRandom();
private final AtomicLong nonceSeq = new AtomicLong();
private final LongSupplier nowNanos;
private final int contextCap;
private final boolean clearAfterTurn;
private volatile MemberLifecycle memberLifecycle = MemberLifecycle.NONE;
/** CB-516: notified with a terminalId on every release; no-op until wired. */
private volatile Consumer<String> releaseListener = _ -> { };
/** CB-520: notified with a terminalId on every acquire; no-op until wired. */
private final List<Consumer<String>> acquireListeners = new java.util.concurrent.CopyOnWriteArrayList<>();
/** CB-516: notified with a {@link ReleaseDetail} on every release; no-op until wired. */
private final List<Consumer<ReleaseDetail>> releaseListeners = new java.util.concurrent.CopyOnWriteArrayList<>();
/** Backward-compatible constructor: shared-tree sessions, production git seam. */
public SessionManager(PeerLauncher launcher) {
this(launcher, new GitWorktrees(), System::nanoTime, 0);
this(launcher, new GitWorktrees(), System::nanoTime, 0, false);
}
/** Backward-compatible constructor with an injectable worktree seam. */
public SessionManager(PeerLauncher launcher, Worktrees worktrees) {
this(launcher, worktrees, System::nanoTime, 0);
this(launcher, worktrees, System::nanoTime, 0, false);
}
/** Test constructor with an injectable clock. */
public SessionManager(PeerLauncher launcher, Worktrees worktrees, LongSupplier nowNanos) {
this(launcher, worktrees, nowNanos, 0);
this(launcher, worktrees, nowNanos, 0, false);
}
/** Production constructor with a configured context turn cap. */
public SessionManager(PeerLauncher launcher, Worktrees worktrees, int contextCap) {
this(launcher, worktrees, System::nanoTime, contextCap);
this(launcher, worktrees, System::nanoTime, contextCap, false);
}
public SessionManager(PeerLauncher launcher, Worktrees worktrees, LongSupplier nowNanos,
int contextCap) {
int contextCap) {
this(launcher, worktrees, nowNanos, contextCap, false);
}
public SessionManager(PeerLauncher launcher, Worktrees worktrees, LongSupplier nowNanos,
int contextCap, boolean clearAfterTurn) {
this.launcher = launcher;
this.worktrees = worktrees;
this.presence = new PresenceBridge(this);
this.nowNanos = nowNanos;
this.contextCap = contextCap;
this.clearAfterTurn = clearAfterTurn;
}
/**
* The single {@link WorkerPresence} view of this manager: it records availability and forwards
* The single {@link MemberPresence} view of this manager: it records availability and forwards
* the signal to the {@code SPAWNING → READY} transition. Pass this to the {@code Injector} and
* {@code BridgeMcp} where they previously accepted a plain {@link WorkerPresence}. The same
* {@code BridgeMcp} where they previously accepted a plain {@link MemberPresence}. The same
* instance is returned every call — presence is shared state, so a fresh bridge per call would
* fragment the {@code present} set and lose signals across callers.
*/
public WorkerPresence asPresence() {
public MemberPresence asPresence() {
return presence;
}
/**
* Spawn a worker and register it as {@link WorkerSession.State#SPAWNING}. The caller's
* Spawn a worker and register it as {@link MemberSession.State#SPAWNING}. The caller's
* identity is recorded as {@code ownerTerminal} ({@code null} for daemon/anon callers).
*/
public WorkerSession acquire(String profile, String requestedCwd, String callerCwd,
public MemberSession acquire(String profile, String requestedCwd, String callerCwd,
String ownerTerminal) {
return acquire(profile, requestedCwd, callerCwd, ownerTerminal, null);
}
/**
* Spawn a worker, optionally inside a fresh git worktree. When {@code wt} is non-null the
* worktree is provisioned, parity-overlaid, and its path becomes the worker's cwd. On any
* failure before registration the worktree is removed so no dangling checkout is left.
* Spawn a member, optionally inside a fresh git worktree, defaulting the role to
* {@link MemberRole#DEV}.
*
* <p>{@code DEV} is the right default because it is exactly what the old "worker" meant: an
* unqualified spawn is a unit of implementation work. An architect or a reviewer is always
* asked for on purpose, so neither is ever what a caller silently gets.
*/
public WorkerSession acquire(String profile, String requestedCwd, String callerCwd,
public MemberSession acquire(String profile, String requestedCwd, String callerCwd,
String ownerTerminal, WorktreeRequest wt) {
return acquire(profile, MemberRole.DEV, requestedCwd, callerCwd, ownerTerminal, wt);
}
/**
* Spawn a member, optionally inside a fresh git worktree, with no session identity requested.
* Equivalent to {@link #acquire(String, MemberRole, String, String, String, WorktreeRequest,
* String, String)} with both trailing args {@code null}.
*
* @param profile which backend to run on — a {@code profiles:} key
* @param role which contract the member runs under; never {@code null}
*/
public MemberSession acquire(String profile, MemberRole role, String requestedCwd, String callerCwd,
String ownerTerminal, WorktreeRequest wt) {
return acquire(profile, role, requestedCwd, callerCwd, ownerTerminal, wt, null, null);
}
/**
* Spawn a member, optionally inside a fresh git worktree, optionally onto a chosen or resumed
* agent session (CB-547a / CB-584). When {@code wt} is non-null the worktree is provisioned,
* parity-overlaid, and its path becomes the member's cwd. On any failure before registration the
* worktree is removed so no dangling checkout is left.
*
* <p>A non-blank {@code resumeSessionId} requires an explicit {@code profile}: a resumed
* conversation is tied to the specific backend that started it, so an unqualified spawn (whose
* backend a placement policy picks at spawn time) has no safe candidate to check the capability
* against. It also requires that profile's adapter to declare {@link Capability#SESSION_RESUME};
* refusing rather than silently delivering a cold session on a non-supporting adapter is the
* whole point of checking before spawn, not after (CB-584).
*
* @param profile which backend to run on — a {@code profiles:} key
* @param role which contract the member runs under; never {@code null}
* @param sessionName the bridge's logical name for the session, or {@code null}
* @param resumeSessionId the prior agent session to resume, or {@code null} for a fresh one
* @throws IllegalArgumentException if {@code resumeSessionId} is set with no explicit profile,
* or the resolved profile's adapter lacks
* {@link Capability#SESSION_RESUME}
*/
public MemberSession acquire(String profile, MemberRole role, String requestedCwd, String callerCwd,
String ownerTerminal, WorktreeRequest wt,
String sessionName, String resumeSessionId) {
MemberRole memberRole = (role == null) ? MemberRole.DEV : role;
requireResumeCapability(profile, resumeSessionId);
if (wt == null) {
SpawnRequest req = new SpawnRequest(profile, requestedCwd, callerCwd);
PeerHandle handle = launcher.spawn(req);
String resolvedProfile = (profile == null || profile.isBlank())
? launcher.defaultProfile() : profile;
String cwd = launcher.effectiveCwd(req);
// CB-557: the role must ride on the SpawnRequest, not stay a local. The launcher needs it
// to pick the profile out of that role's pool and to label the tab; a role kept only on
// the MemberSession is recorded after the spawn it was supposed to steer.
SpawnRequest req = new SpawnRequest(profile, requestedCwd, callerCwd, sessionName, resumeSessionId, memberRole);
PeerHandle handle;
try {
handle = launcher.spawn(req);
} catch (RuntimeException e) {
log.warn("spawn failed for profile={} role={}: {}", profile, memberRole, e.getMessage());
throw e;
}
String resolvedProfile = resolveProfile(handle, profile);
String cwd = launcher.effectiveCwd(new SpawnRequest(resolvedProfile, requestedCwd, callerCwd));
long now = nowNanos.getAsLong();
WorkerSession session = new WorkerSession(
MemberSession session = new MemberSession(
handle.id(),
handle.terminalId(),
resolvedProfile,
memberRole,
cwd,
ownerTerminal,
now,
now,
0,
WorkerSession.State.SPAWNING,
MemberSession.State.SPAWNING,
null,
null);
null,
handle.charterReceipt(),
handle.agentSessionId());
registry.put(handle.id(), session);
memberLifecycle.acquired(session.role(), session.profile(), session.terminalId());
log.debug("acquired session id={} terminal={} profile={} owner={}",
handle.id(), handle.terminalId(), session.profile(), session.ownerTerminal());
notifyAcquired(session.terminalId());
return session;
}
return acquireWithWorktree(profile, requestedCwd, callerCwd, ownerTerminal, wt);
return acquireWithWorktree(profile, memberRole, requestedCwd, callerCwd, ownerTerminal, wt,
sessionName, resumeSessionId);
}
/**
* Refuse a {@code resumeSessionId} that either names no explicit profile or names one whose
* adapter does not declare {@link Capability#SESSION_RESUME}. A no-op when
* {@code resumeSessionId} is blank — the ordinary, no-identity spawn path.
*/
private void requireResumeCapability(String profile, String resumeSessionId) {
if (resumeSessionId == null || resumeSessionId.isBlank()) {
return;
}
if (profile == null || profile.isBlank()) {
throw new IllegalArgumentException("resumeSessionId requires an explicit profile — "
+ "a resumed conversation is tied to the backend that started it, so it cannot "
+ "be left to placement to pick");
}
Set<Capability> caps = launcher.capabilitiesFor(profile);
if (!caps.contains(Capability.SESSION_RESUME)) {
throw new IllegalArgumentException("worker profile '" + profile + "' does not declare "
+ "Capability.SESSION_RESUME — refusing resumeSessionId rather than silently "
+ "starting a cold session");
}
}
/** Tear a worker down by pane id and remove it from the registry. Idempotent. */
public void release(String paneId) {
WorkerSession removed = registry.remove(paneId);
release(paneId, ReleaseCause.COMPLETED);
}
/**
* Core teardown: always stops the worker pane and deregisters the session; whether the worker's
* git worktree is also removed depends on {@code cause}.
*
* <p>CB-544: these are two concerns that used to be fused. Stopping the pane is correct on every
* teardown — the worker process must end. Removing the worktree is a destructive act that is only
* correct for a deliberately-finished teardown (an explicit stop of a completed session, the
* reaper releasing a genuinely idle one, or a context-capped session). A shutdown drain
* must stop panes but preserve worktrees: a worker's uncommitted work exists in exactly one
* place — its worktree — so deleting it while the daemon simply goes down is silent data loss,
* with no copy and no error. Do NOT fuse these back together; the cost of an orphaned worktree
* is a logged path an operator can reclaim, the cost of a deleted one is unrecoverable work.
*/
private void release(String paneId, ReleaseCause cause) {
MemberSession removed = registry.remove(paneId);
boolean preserveWorktree = cause == ReleaseCause.SHUTDOWN;
String snapshotRef = null;
if (removed != null) {
log.debug("releasing session pane={} terminal={} state={}",
removed.paneId(), removed.terminalId(), removed.state());
// CB-516: a send still waiting on this worker can never be answered now. Tell the
// listener BEFORE the pane is torn down, so a blocked caller fails fast with a real
// reason instead of sitting on a rendezvous nothing will ever resolve.
notifyReleased(removed.terminalId());
try {
memberLifecycle.released(removed.terminalId());
log.debug("releasing session pane={} terminal={} state={} cause={}",
removed.paneId(), removed.terminalId(), removed.state(), cause);
boolean dirty = removed.worktree() != null && worktrees.hasUncommitted(removed.worktree());
if (preserveWorktree && removed.worktree() != null) {
logPreservedForShutdown(removed);
} else if (dirty) {
// CB-576: a release that would otherwise remove the worktree finds it holding
// uncommitted work the bridge cannot see. A worker that ends a turn without
// committing (normally because it stopped to ask a question or refused the turn)
// has its only copy of that work in the worktree. Remove would --force-delete it,
// so preserve the directory and tell an operator where to find it.
preserveWorktree = true;
log.warn("release {} preserves dirty worktree {} for pane={} terminal={}: "
+ "the worktree holds uncommitted changes that --force remove would destroy",
cause, removed.worktree(), removed.paneId(), removed.terminalId());
}
if (dirty) {
// CB-578 stage C: preserving on disk is not saving — the directory is one
// `worktree remove --force`, or an operator tidying up, away from gone. Commit
// its full state to a ref before the preserve-or-remove decision above can be
// undone by anything else, regardless of why this release fired.
snapshotRef = trySnapshot(removed, cause);
}
} catch (RuntimeException e) {
// CB-581: hasUncommitted shells out to `git status` and can throw on a non-zero
// exit. We can no longer tell whether the worktree holds uncommitted work, so fail
// toward the safe answer and preserve it — deleting on a guess can destroy work
// that has no other copy (CB-576), while keeping it on a false alarm only costs
// disk. The exception must not propagate: the pane still has to stop below.
preserveWorktree = true;
log.warn("release {} could not tell whether worktree {} for pane={} terminal={} has "
+ "uncommitted changes; preserving it rather than risk destroying unsaved work: {}",
cause, removed.worktree(), removed.paneId(), removed.terminalId(), e.toString());
} finally {
// CB-516/CB-581: a send still waiting on this worker can never be answered now, no
// matter what happened above. Tell the listener BEFORE the pane is torn down, so a
// blocked caller fails fast with a real reason instead of sitting on a rendezvous
// nothing will ever resolve. CB-578 stage C: carry the worktree/branch/snapshot ref
// too, so a failed ticket's detail can point a lead at the same tree to re-dispatch.
notifyReleased(new ReleaseDetail(removed.terminalId(), removed.worktree(),
removed.branch(), snapshotRef));
}
}
// CB-581: the pane must always stop, even if the dirty check above threw. A session removed
// from the registry with no pane stop is an orphaned pane — a live terminal burning a fleet
// slot that no longer appears in the roster and can never be reclaimed.
launcher.stop(paneId);
if (removed != null && removed.worktree() != null) {
if (removed != null && !preserveWorktree && removed.worktree() != null) {
worktrees.remove(worktrees.repoRoot(removed.cwd()), removed.worktree());
}
}
/**
* Best-effort snapshot of a dirty worktree into {@code refs/wip/<branch>} (CB-578 stage C). A
* failure here must never escalate: the caller has already decided to preserve the worktree
* regardless of whether this succeeds, so the only cost of a failed snapshot is a WARN and a
* missing ref — never a lost pane stop or a lost release notification.
*/
private String trySnapshot(MemberSession session, ReleaseCause cause) {
if (session.worktree() == null || session.branch() == null) {
return null;
}
try {
Optional<String> ref = worktrees.snapshot(session.worktree(), session.branch(),
snapshotMessage(session, cause));
ref.ifPresent(sha -> log.info(
"snapshotted dirty worktree {} to refs/wip/{} commit={} for pane={} terminal={}",
session.worktree(), session.branch(), sha, session.paneId(), session.terminalId()));
return ref.orElse(null);
} catch (RuntimeException e) {
log.warn("snapshot of dirty worktree {} failed for pane={} terminal={} branch={}: the "
+ "worktree is still preserved on disk, just not committed to refs/wip/{}: {}",
session.worktree(), session.paneId(), session.terminalId(), session.branch(),
session.branch(), e.toString());
return null;
}
}
/** Commit message for a CB-578 stage C snapshot — names the member so an operator can tell runs apart. */
private String snapshotMessage(MemberSession session, ReleaseCause cause) {
return "CB-578 stage C: snapshot of a released worker\n\n"
+ "terminal: " + session.terminalId() + "\n"
+ "profile: " + session.profile() + "\n"
+ "branch: " + session.branch() + "\n"
+ "cause: " + cause;
}
/**
* Facts about a released session that a listener needs beyond the bare terminal id — enough
* for a caller to point a lead at where to re-dispatch onto the same tree after a failed
* release (CB-578 stage C, acceptance criterion 10). {@code worktreePath} and {@code branch}
* are {@code null} for a shared-tree session; {@code snapshotRef} is {@code null} unless this
* release snapshotted a dirty worktree into {@code refs/wip/<branch>}.
*/
public record ReleaseDetail(String terminalId, String worktreePath, String branch, String snapshotRef) {
}
/**
* Why a session is being released — governs whether its worktree is preserved or removed.
* Worktree removal is reserved for the one case that is genuinely finished; everything else
* must keep the worker's only copy of its work.
*/
public enum ReleaseCause {
/** Deliberate teardown of a finished session. Stops the pane and removes the worktree. */
COMPLETED,
/** Daemon shutdown drain. Stops the pane but PRESERVES the worktree. */
SHUTDOWN
}
/**
* CB-544 shutdown drain log for a worktree we deliberately kept. A session still {@code BUSY}
* when the drain timeout expired was abandoned mid-turn — that work may be uncommitted and is
* the only copy — so the message is loud and points at the path an operator needs to reclaim.
*/
private void logPreservedForShutdown(MemberSession session) {
if (session.state() == MemberSession.State.BUSY) {
log.warn("shutdown drain abandoned BUSY session pane={} terminal={} mid-turn; "
+ "worktree preserved at {}", session.paneId(), session.terminalId(),
session.worktree());
} else {
log.info("shutdown drain preserved worktree at {} for pane={}",
session.worktree(), session.paneId());
}
}
/**
* Register a callback invoked with a session's {@code terminalId} whenever it is acquired
* (CB-520). This is the hook that lets the reply inbox {@code own} a target's queue.
*/
public void onAcquire(Consumer<String> listener) {
if (listener != null) {
acquireListeners.add(listener);
}
}
/**
* Register a callback invoked with a session's {@code terminalId} whenever it is released
* (CB-516). Every teardown path funnels through {@link #release}, so one hook covers the REST
* and MCP stop tools, the idle-TTL reaper, {@code recycle}, and shutdown drain alike.
* and MCP stop tools, the idle-TTL reaper, and shutdown drain alike.
*
* <p>Set rather than injected because {@code MessageService} — the intended listener — is
* <p>Added rather than injected because {@code MessageService} — one intended listener — is
* constructed after this manager (it needs the injector and rendezvous, which need the session
* presence view this manager exposes). Wiring it at construction would require breaking that
* cycle for one callback.
*/
public void onRelease(Consumer<String> listener) {
this.releaseListener = (listener == null) ? _ -> { } : listener;
public void onRelease(Consumer<ReleaseDetail> listener) {
if (listener != null) {
releaseListeners.add(listener);
}
}
/** A listener failure must never prevent the teardown it is reacting to. */
private void notifyReleased(String terminalId) {
/** Inject the optional member-slot lifecycle after construction without changing constructors. */
public void setMemberLifecycle(MemberLifecycle memberLifecycle) {
this.memberLifecycle = memberLifecycle == null ? MemberLifecycle.NONE : memberLifecycle;
}
/** A listener failure must never prevent the acquisition it is reacting to. */
private void notifyAcquired(String terminalId) {
if (terminalId == null) {
return;
}
try {
releaseListener.accept(terminalId);
} catch (RuntimeException e) {
log.warn("release listener failed for terminal {}: {}", terminalId, e.toString());
for (Consumer<String> listener : acquireListeners) {
try {
listener.accept(terminalId);
} catch (RuntimeException e) {
log.warn("acquire listener failed for terminal {}: {}", terminalId, e.toString());
}
}
}
private WorkerSession acquireWithWorktree(String profile, String requestedCwd, String callerCwd,
String ownerTerminal, WorktreeRequest wt) {
String resolvedProfile = (profile == null || profile.isBlank())
/** A listener failure must never prevent the teardown it is reacting to. */
private void notifyReleased(ReleaseDetail detail) {
if (detail.terminalId() == null) {
return;
}
for (Consumer<ReleaseDetail> listener : releaseListeners) {
try {
listener.accept(detail);
} catch (RuntimeException e) {
log.warn("release listener failed for terminal {}: {}", detail.terminalId(), e.toString());
}
}
}
private MemberSession acquireWithWorktree(String profile, MemberRole memberRole, String requestedCwd, String callerCwd,
String ownerTerminal, WorktreeRequest wt,
String sessionName, String resumeSessionId) {
String preResolvedProfile = (profile == null || profile.isBlank())
? launcher.defaultProfile() : profile;
// CB-507: resolve through the launcher's CB-112 chain (requested → profile cwd → caller →
// daemon cwd → "."), never the raw args. A plain REST spawn supplies neither a requested
@@ -187,15 +449,17 @@ public final class SessionManager implements TurnListener {
// `git -C null` on the command line — an NPE out of ProcessBuilder, surfacing as HTTP 500.
// The non-worktree path always used this chain; only this branch was missed.
String repoRoot = worktrees.repoRoot(
launcher.effectiveCwd(new SpawnRequest(resolvedProfile, requestedCwd, callerCwd)));
launcher.effectiveCwd(new SpawnRequest(preResolvedProfile, requestedCwd, callerCwd)));
String branch = "worker/" + slug(wt.ticketSlug()) + "-" + nonce();
String path = null;
PeerHandle handle;
try {
path = worktrees.add(repoRoot, branch, wt.baseRef());
worktrees.overlayParity(repoRoot, path, launcher.parityOverlay(resolvedProfile));
handle = launcher.spawn(new SpawnRequest(profile, path, callerCwd));
worktrees.overlayParity(repoRoot, path, launcher.parityOverlay(preResolvedProfile));
handle = launcher.spawn(new SpawnRequest(profile, path, callerCwd, sessionName, resumeSessionId, memberRole));
} catch (RuntimeException e) {
log.warn("spawn failed for profile={} role={} branch={} path={}: {}",
preResolvedProfile, memberRole, branch, path, e.getMessage());
if (path != null) {
try {
worktrees.remove(repoRoot, path);
@@ -205,22 +469,29 @@ public final class SessionManager implements TurnListener {
}
throw e;
}
String resolvedProfile = resolveProfile(handle, profile);
String cwd = launcher.effectiveCwd(new SpawnRequest(resolvedProfile, path, callerCwd));
long now = nowNanos.getAsLong();
WorkerSession session = new WorkerSession(
MemberSession session = new MemberSession(
handle.id(),
handle.terminalId(),
resolvedProfile,
resolveCwd(path, profile, callerCwd),
memberRole,
cwd,
ownerTerminal,
now,
now,
0,
WorkerSession.State.SPAWNING,
MemberSession.State.SPAWNING,
path,
branch);
branch,
handle.charterReceipt(),
handle.agentSessionId());
registry.put(handle.id(), session);
memberLifecycle.acquired(session.role(), session.profile(), session.terminalId());
log.debug("acquired worktree session id={} terminal={} profile={} branch={} path={}",
handle.id(), handle.terminalId(), session.profile(), session.branch(), session.worktree());
notifyAcquired(session.terminalId());
return session;
}
@@ -233,25 +504,28 @@ public final class SessionManager implements TurnListener {
}
/**
* Release the old session and acquire a fresh one with the same profile and working directory.
* The new session is guaranteed to have a pane id distinct from the old one (no-reuse invariant).
* The profile to record for a session. A launcher that performed dynamic selection tells us
* the actual profile via {@link PeerHandle#profile()}; otherwise fall back to what the caller
* requested (or the launcher's default for a no-profile spawn).
*/
public WorkerSession recycle(String paneId) {
WorkerSession old = registry.get(paneId);
if (old == null) {
throw new IllegalArgumentException("no session for paneId " + paneId);
private String resolveProfile(PeerHandle handle, String requestedProfile) {
String fromHandle = handle.profile();
if (fromHandle != null && !fromHandle.isBlank()) {
return fromHandle;
}
release(paneId);
return acquire(old.profile(), old.cwd(), old.cwd(), old.ownerTerminal());
if (requestedProfile != null && !requestedProfile.isBlank()) {
return requestedProfile;
}
return launcher.defaultProfile();
}
/** The session for {@code paneId}, if it is still registered and not released. */
public Optional<WorkerSession> get(String paneId) {
public Optional<MemberSession> get(String paneId) {
return Optional.ofNullable(registry.get(paneId));
}
/** Bridge-owned roster: all registered sessions (acquired minus released). */
public List<WorkerSession> roster() {
public List<MemberSession> roster() {
return List.copyOf(registry.values());
}
@@ -259,11 +533,15 @@ public final class SessionManager implements TurnListener {
* CB-304 merged roster+live view. The registry is authoritative for worktree, branch,
* profile, owner, and state; the optional live agent supplies the herdr-reported status.
*/
public static Map<String, Object> rosterView(WorkerSession session, Agent live) {
public static Map<String, Object> rosterView(MemberSession session, Agent live) {
Map<String, Object> m = new LinkedHashMap<>();
m.put("sessionId", session.terminalId());
m.put("paneId", session.paneId());
// Both axes, always: profile says which backend this member runs on, role says what it is
// for. A lead reading the roster needs both — two rows may share a profile and still be
// allowed to do entirely different things.
m.put("profile", session.profile());
m.put("role", session.role() == null ? "dev" : session.role().wireName());
m.put("state", session.state().name().toLowerCase());
if (session.worktree() != null) {
m.put("worktree", session.worktree());
@@ -274,13 +552,29 @@ public final class SessionManager implements TurnListener {
if (session.ownerTerminal() != null) {
m.put("owner", session.ownerTerminal());
}
// CB-584: which conversation this member holds — the id a later resumeSessionId spawn would
// pass back. Absent for an adapter that declines Capability.SESSION_RESUME, or one that
// resolves it lazily and has not yet.
if (session.agentSessionId() != null) {
m.put("agentSessionId", session.agentSessionId());
}
// CB-571: which charter this member was started with — never the charter text itself. The
// digest lets a lead tell at a glance whether all members got the same charter; the source
// records whether a role charter was configured ("fleet.charters.<role>") or only the reply
// charter was composed ("none").
if (session.charterReceipt() != null) {
m.put("charterSource", session.charterReceipt().charterSource());
if (session.charterReceipt().charterSha256() != null) {
m.put("charterSha256", session.charterReceipt().charterSha256());
}
}
m.put("liveStatus", live == null ? "unknown" : live.status().name().toLowerCase());
return m;
}
/** Lifecycle hook: worker became available on the bridge MCP. */
void onReady(String terminalId) {
transitionByTerminal(terminalId, WorkerSession.State.SPAWNING, WorkerSession.State.READY);
transitionByTerminal(terminalId, MemberSession.State.SPAWNING, MemberSession.State.READY);
}
/**
@@ -289,14 +583,14 @@ public final class SessionManager implements TurnListener {
* can be re-delivered for multi-turn reuse until it is released.
*/
@Override
public void onDelivered(String target) {
WorkerSession current = findByTerminal(target);
public void onDelivered(String target, TurnToken token) {
MemberSession current = findByTerminal(target);
if (current == null) return;
if (current.state() != WorkerSession.State.READY && current.state() != WorkerSession.State.DONE) {
if (current.state() != MemberSession.State.READY && current.state() != MemberSession.State.DONE) {
return;
}
long now = nowNanos.getAsLong();
WorkerSession updated = current.withState(WorkerSession.State.BUSY).bumpTurn(now);
MemberSession updated = current.withState(MemberSession.State.BUSY).bumpTurn(now);
if (replace(current, updated)) {
log.debug("session transitioned terminal={} pane={} {} -> BUSY turn={}",
target, current.paneId(), current.state(), updated.turnCount());
@@ -306,16 +600,42 @@ public final class SessionManager implements TurnListener {
/** Lifecycle hook: the worker's delegated turn completed successfully. */
@Override
public void onTurnComplete(String target) {
WorkerSession current = findByTerminal(target);
if (current == null || current.state() != WorkerSession.State.BUSY) return;
completeTurn(target, false);
}
@Override
public boolean hasPostTurnAction(String target) {
if (!clearAfterTurn) return false;
MemberSession current = findByTerminal(target);
return current != null && current.state() == MemberSession.State.BUSY
&& (contextCap <= 0 || current.turnCount() < contextCap);
}
@Override
public boolean onTurnCompleteWithPostAction(String target) {
return completeTurn(target, true);
}
private boolean completeTurn(String target, boolean startContextReset) {
MemberSession current = findByTerminal(target);
if (current == null || current.state() != MemberSession.State.BUSY) return false;
long now = nowNanos.getAsLong();
WorkerSession updated = current.withState(WorkerSession.State.DONE).withActivity(now);
MemberSession updated = current.withState(MemberSession.State.DONE).withActivity(now);
if (replace(current, updated)) {
log.debug("session transitioned terminal={} pane={} BUSY -> DONE turn={}",
target, current.paneId(), updated.turnCount());
}
if (contextCap > 0 && updated.turnCount() >= contextCap) {
release(current.paneId());
return false;
}
if (!startContextReset || !clearAfterTurn) return false;
try {
return launcher.clearContext(current.paneId());
} catch (RuntimeException e) {
log.warn("context reset failed for terminal={} pane={}; continuing without reset: {}",
target, current.paneId(), e.getMessage());
return false;
}
}
@@ -327,11 +647,13 @@ public final class SessionManager implements TurnListener {
/** Lifecycle hook: the worker vanished or was dropped mid-life. */
void onFailed(String target) {
WorkerSession current = findByTerminal(target);
MemberSession current = findByTerminal(target);
if (current == null) return;
if (current.state() == WorkerSession.State.RELEASED) return;
if (replace(current, current.withState(WorkerSession.State.FAILED))) {
log.debug("session marked failed terminal={} pane={}", target, current.paneId());
if (current.state() == MemberSession.State.RELEASED) return;
MemberSession.State priorState = current.state();
if (replace(current, current.withState(MemberSession.State.FAILED))) {
log.warn("member terminal={} pane={} can no longer be delegated to: its turn never resolved "
+ "(was {} when it failed)", target, current.paneId(), priorState);
}
}
@@ -343,32 +665,49 @@ public final class SessionManager implements TurnListener {
int reapIdle(long idleTtlNanos) {
long now = nowNanos.getAsLong();
int reaped = 0;
for (WorkerSession s : roster()) {
if (s.state() != WorkerSession.State.READY && s.state() != WorkerSession.State.DONE) {
for (MemberSession s : roster()) {
if (s.state() != MemberSession.State.READY && s.state() != MemberSession.State.DONE) {
continue;
}
if (now - s.lastActivityAtNanos() > idleTtlNanos) {
release(s.paneId());
reaped++;
long idleNanos = now - s.lastActivityAtNanos();
if (idleNanos > idleTtlNanos) {
log.debug("reaping idle session terminal={} pane={}: idle {}s exceeds the {}s ttl",
s.terminalId(), s.paneId(), TimeUnit.NANOSECONDS.toSeconds(idleNanos),
TimeUnit.NANOSECONDS.toSeconds(idleTtlNanos));
// CB-581: one session that fails to release must not abort the whole reaping pass —
// match drainAll's per-session try/catch so the rest of the roster still gets reaped.
try {
release(s.paneId());
reaped++;
} catch (RuntimeException e) {
log.warn("reap failed for pane={} terminal={} worktree={}; continuing with "
+ "remaining sessions", s.paneId(), s.terminalId(), s.worktree(), e);
}
}
}
return reaped;
}
/**
* Gracefully drain all registered sessions. For each session that is {@code BUSY}, poll up to
* {@code timeoutNanos} for it to leave {@code BUSY}, then release it regardless. Non-busy
* sessions are released immediately. A failure releasing one session is logged and does not
* abort the rest.
* Gracefully drain all registered sessions on daemon shutdown. For each session that is
* {@code BUSY}, poll up to {@code timeoutNanos} for it to leave {@code BUSY}, then release it
* regardless. Non-busy sessions are released immediately. A failure releasing one session is
* logged and does not abort the rest.
*
* <p>CB-544: this is a {@link ReleaseCause#SHUTDOWN} release — the worker's pane is stopped
* (the process must end) but its worktree is preserved and its path logged. Shutdown is never
* a reason to delete a worker's only copy of its uncommitted work. A session still {@code BUSY}
* when the timeout expired is abandoned mid-turn and logged loudly so an operator can find its
* kept worktree.
*/
void drainAll(long timeoutNanos) {
long deadline = System.nanoTime() + timeoutNanos;
for (WorkerSession s : roster()) {
for (MemberSession s : roster()) {
try {
if (s.state() == WorkerSession.State.BUSY) {
if (s.state() == MemberSession.State.BUSY) {
while (System.nanoTime() < deadline) {
WorkerSession current = registry.get(s.paneId());
if (current == null || current.state() != WorkerSession.State.BUSY) {
MemberSession current = registry.get(s.paneId());
if (current == null || current.state() != MemberSession.State.BUSY) {
break;
}
try {
@@ -380,7 +719,7 @@ public final class SessionManager implements TurnListener {
}
}
}
release(s.paneId());
release(s.paneId(), ReleaseCause.SHUTDOWN);
} catch (RuntimeException e) {
log.warn("drain failed for pane={}; continuing with remaining sessions", s.paneId(), e);
}
@@ -401,16 +740,28 @@ public final class SessionManager implements TurnListener {
return registry.size();
}
private WorkerSession findByTerminal(String terminalId) {
for (WorkerSession s : registry.values()) {
/**
* The registered session owning {@code terminalId}, or {@code null} if none does.
*
* <p>A null {@code terminalId} is a normal input, not a caller bug: every lifecycle hook here is
* fed from the MCP transport, where the <em>primary</em> resolves to a {@link
* dev.ltms.bridged.auth.Principal} with no terminal. {@code BridgeMcp} documents that contact as
* a no-op, and {@link dev.ltms.bridged.inject.MemberPresence#markPresent} honours it — but
* {@code PresenceBridge} then forwards the same null here. Matching on a null id can never
* succeed anyway (a registered session always has a terminal), so answer "no match" rather than
* throwing: an NPE on this path takes down an unrelated tool call for the primary.
*/
private MemberSession findByTerminal(String terminalId) {
if (terminalId == null) return null;
for (MemberSession s : registry.values()) {
if (terminalId.equals(s.terminalId())) return s;
}
return null;
}
private void transitionByTerminal(String terminalId, WorkerSession.State from,
WorkerSession.State to) {
WorkerSession current = findByTerminal(terminalId);
private void transitionByTerminal(String terminalId, MemberSession.State from,
MemberSession.State to) {
MemberSession current = findByTerminal(terminalId);
if (current == null || current.state() != from) return;
long now = nowNanos.getAsLong();
if (replace(current, current.withState(to).withActivity(now))) {
@@ -419,16 +770,12 @@ public final class SessionManager implements TurnListener {
}
}
private boolean replace(WorkerSession expected, WorkerSession updated) {
private boolean replace(MemberSession expected, MemberSession updated) {
return registry.replace(expected.paneId(), expected, updated);
}
private String resolveCwd(String requestedCwd, String profileName, String callerCwd) {
return launcher.effectiveCwd(new SpawnRequest(profileName, requestedCwd, callerCwd));
}
/** WorkerPresence bridge that also drives the manager's READY transition. */
private static final class PresenceBridge extends WorkerPresence {
/** MemberPresence bridge that also drives the manager's READY transition. */
private static final class PresenceBridge extends MemberPresence {
private final SessionManager sessions;
PresenceBridge(SessionManager sessions) {
@@ -437,6 +784,9 @@ public final class SessionManager implements TurnListener {
@Override
public void markPresent(String terminal) {
if (terminal == null || terminal.isBlank()) {
return; // the primary's contact carries no worker terminal — not a readiness signal
}
super.markPresent(terminal);
sessions.onReady(terminal);
}
@@ -1,58 +0,0 @@
package dev.ltms.bridged.session;
/**
* A bridge-owned worker session — the authoritative in-daemon record of a worker this
* process spawned. Immutable; state transitions are performed by replacing the record in
* {@link SessionManager}'s registry.
*
* @param paneId herdr pane handle — the registry key and the argument to teardown
* @param terminalId herdr terminal handle — the {@code target} for send/read/status
* @param profile the worker profile name that spawned this session
* @param cwd the resolved working directory the worker started in
* @param ownerTerminal the caller that requested this worker ({@code null} = daemon/anon)
* @param spawnedAtNanos {@link System#nanoTime()} when the session was registered
* @param lastActivityAtNanos {@link System#nanoTime()} of the most recent lifecycle event
* @param turnCount number of delegated turns that have been delivered to this session
* @param state current lifecycle state in the one-shot FSM
*/
public record WorkerSession(
String paneId,
String terminalId,
String profile,
String cwd,
String ownerTerminal,
long spawnedAtNanos,
long lastActivityAtNanos,
int turnCount,
State state,
String worktree,
String branch) {
/** One-shot worker lifecycle states. */
public enum State {
SPAWNING,
READY,
BUSY,
DONE,
FAILED,
RELEASED
}
/** Return a copy of this session in {@code state}. */
public WorkerSession withState(State state) {
return new WorkerSession(paneId, terminalId, profile, cwd, ownerTerminal, spawnedAtNanos,
lastActivityAtNanos, turnCount, state, worktree, branch);
}
/** Return a copy with {@code lastActivityAtNanos} updated to {@code nowNanos}. */
public WorkerSession withActivity(long nowNanos) {
return new WorkerSession(paneId, terminalId, profile, cwd, ownerTerminal, spawnedAtNanos,
nowNanos, turnCount, state, worktree, branch);
}
/** Return a copy with the turn count incremented and activity timestamped at {@code nowNanos}. */
public WorkerSession bumpTurn(long nowNanos) {
return new WorkerSession(paneId, terminalId, profile, cwd, ownerTerminal, spawnedAtNanos,
nowNanos, turnCount + 1, state, worktree, branch);
}
}
@@ -1,6 +1,7 @@
package dev.ltms.bridged.session;
import java.util.List;
import java.util.Optional;
/** Seam between {@link SessionManager} and git worktree operations. Tests use a recording fake. */
public interface Worktrees {
@@ -10,9 +11,47 @@ public interface Worktrees {
/** git -C <repoRoot> worktree remove --force <path>. Idempotent (already-gone tolerated). */
void remove(String repoRoot, String worktreePath);
/**
* True when the worktree holds uncommitted changes the bridge cannot see: tracked
* modifications, staged files, or untracked files. {@code git status --porcelain} is the
* test; an empty result means clean. Callers use this to decide whether removing the
* worktree would silently destroy a worker's only copy of its work.
*
* <p>An already-gone worktree is reported as clean (no throw), matching {@link #remove}'s
* idempotent contract: a path that does not exist holds no work to lose, and must not break
* a teardown that is mid-way through stopping the pane.
*/
boolean hasUncommitted(String worktreePath);
/** Copy each existing overlay path repoRoot→worktree; mark tracked ones --skip-worktree. */
void overlayParity(String repoRoot, String worktreePath, List<String> overlay);
/** git -C <cwd> rev-parse --show-toplevel — the repo root that owns cwd. */
String repoRoot(String cwd);
/**
* Commit the worktree's full on-disk state — tracked and untracked, respecting
* {@code .gitignore} — to {@code refs/wip/<branch>}, so a release that would otherwise leave
* the work as loose, unprotected files has a durable git object to fall back on (CB-578 stage
* C). Built on a <em>temporary</em> index: the worker's own index, working tree, and HEAD are
* never touched, since the worker may still be mid-write. The commit is parented on the
* worktree's current HEAD.
*
* <p>Never writes under {@code refs/heads/} — the ref must not appear in {@code git branch},
* must not be pushed by default, and must not be swept by a later {@code git branch -d}.
*
* <p>Callers are expected to have already confirmed {@link #hasUncommitted} before reaching
* for this; it always stages and commits whatever {@code git add -A} finds, so calling it on
* a clean worktree still produces a (harmless, tree-identical-to-HEAD) commit rather than
* detecting cleanliness itself.
*
* @param worktreePath absolute path of the worktree to snapshot
* @param branch the worktree's own branch — keys {@code refs/wip/<branch>}
* @param message the commit message; should name the member, its branch, and the release
* cause so an operator can tell which run produced it
* @return the created commit's sha, or {@link Optional#empty()} if {@code worktreePath} does
* not exist (mirrors {@link #remove} and {@link #hasUncommitted}'s already-gone
* tolerance — a worktree that is gone holds nothing to snapshot)
*/
Optional<String> snapshot(String worktreePath, String branch, String message);
}
@@ -1,195 +0,0 @@
package dev.ltms.bridged.worker;
import dev.ltms.bridged.config.BridgedConfig;
import dev.ltms.bridged.guard.SubscriptionGuard;
import dev.ltms.bridged.herdr.Agent;
import dev.ltms.bridged.herdr.AgentControl;
import dev.ltms.bridged.herdr.WorkspaceControl;
import dev.ltms.bridged.peer.Capability;
import java.util.EnumSet;
import java.util.List;
import java.util.Map;
import java.util.Set;
import java.util.function.Function;
import java.util.function.LongSupplier;
/**
* The {@link HerdrPeerLauncher} adapter for <strong>Claude Code</strong> — the safe path from a
* delegation request to a running off-subscription Claude.
*
* <p>Everything transport-related (tab/pane placement, the CB-306 spawn-readiness gate, unique
* naming, CB-117 orphan reap, teardown, listing, cwd resolution) lives in the base. This class
* supplies only the two Claude-specific seams:
* <ul>
* <li>the {@code claude} name prefix (so reap matches {@code claude-*} panes, never another
* adapter's), and</li>
* <li>{@link #buildLaunch}, which encodes the subscription boundary: build the worker env with
* {@code ANTHROPIC_BASE_URL}, assert that host is on the allowlist <em>before</em> touching
* herdr, and mount the bridge MCP + reply charter as inline launch flags. A worker's base_url
* lives in the env map handed to herdr and nowhere else; {@code bridged}'s own environment is
* never mutated, and nothing is written to the worker's profile.</li>
* </ul>
*/
public final class ClaudeCodeLauncher extends HerdrPeerLauncher {
/** Label prefix for this adapter's herdr agent names (drives naming + orphan reap). */
private static final String NAME_PREFIX = "claude";
private final SubscriptionGuard guard;
/**
* Standing instruction appended to the worker's system prompt so it returns its result via
* {@code bridge_reply}. Injected as a launch flag, so nothing is written to the worker's
* profile — it is guidance, and a worker that never replies is caught by the send's timeout.
*/
static final String REPLY_CHARTER =
"You are an off-subscription worker in the claude-bridge fleet. Every message you "
+ "receive arrives through the bridge, and the ONLY channel back to the sender is the "
+ "bridge_reply MCP tool. Text you write in your terminal is NOT sent anywhere — the "
+ "sender cannot see your screen, so an in-terminal answer is silently discarded. "
+ "Therefore you MUST end EVERY turn by calling bridge_reply with `content` set to your "
+ "complete response. This holds for every message without exception — tasks, questions, "
+ "clarifications, acknowledgements, and ordinary back-and-forth conversation. Call "
+ "bridge_reply exactly once, as the final action of your turn, with your full answer in "
+ "`content`; never wait for confirmation first. If you end a turn without calling "
+ "bridge_reply, the sender receives nothing and the exchange stalls.";
/**
* Production constructor — disables the spawn-ready gate ({@code spawnReadyTimeoutMs == 0}) so
* existing deployments and tests keep the legacy non-blocking spawn semantics.
*/
public ClaudeCodeLauncher(AgentControl agents, WorkspaceControl spaces, SubscriptionGuard guard,
Map<String, BridgedConfig.Worker> profiles, String defaultProfile,
Function<String, String> env) {
this(agents, spaces, guard, profiles, defaultProfile, env, 0,
System::currentTimeMillis, () -> sleepUninterruptibly(300));
}
/**
* Production constructor with spawn-ready gate enabled. The gate polls {@code agents.status()}
* until the pane reports an injectable state or {@code spawnReadyTimeoutMs} elapses.
*/
public ClaudeCodeLauncher(AgentControl agents, WorkspaceControl spaces, SubscriptionGuard guard,
Map<String, BridgedConfig.Worker> profiles, String defaultProfile,
Function<String, String> env,
long spawnReadyTimeoutMs, long spawnReadyPollMs) {
this(agents, spaces, guard, profiles, defaultProfile, env,
spawnReadyTimeoutMs,
System::currentTimeMillis, () -> sleepUninterruptibly(spawnReadyPollMs));
}
/**
* Full testability constructor. Every injectable collaborator is explicit so unit tests supply
* fakes for the clock ({@code nowMillis}) and poll-loop wait ({@code sleeper}). The
* {@code sleeper} is never called when the gate is disabled ({@code spawnReadyTimeoutMs == 0}).
*
* @param agents herdr agent control (start, status, close)
* @param spaces workspace / tab control (ensure, create, close)
* @param guard subscription-boundary guard (checked before spawning)
* @param profiles configured worker profiles
* @param defaultProfile profile a no-argument spawn uses (nullable)
* @param env host env lookup (injectable for tests)
* @param spawnReadyTimeoutMs max ms to wait for injectable state (0 disables the gate)
* @param nowMillis monotonic clock source (e.g. {@code System::currentTimeMillis})
* @param sleeper sleep/wait hook (e.g. {@code () -> Thread.sleep(pollMs)}); it
* already encodes the poll interval, so the 8th positional argument
* (poll ms) is accepted for API symmetry but otherwise unused here
*/
public ClaudeCodeLauncher(AgentControl agents, WorkspaceControl spaces, SubscriptionGuard guard,
Map<String, BridgedConfig.Worker> profiles, String defaultProfile,
Function<String, String> env,
long spawnReadyTimeoutMs,
LongSupplier nowMillis, Runnable sleeper) {
super(NAME_PREFIX, agents, spaces, profiles, defaultProfile, env,
spawnReadyTimeoutMs, nowMillis, sleeper);
this.guard = guard;
}
/**
* {@inheritDoc}
*
* <p>The spawn sequence encodes the subscription boundary: assert the profile's base_url is on
* the allowlist <em>before</em> any herdr call, then build the worker env with
* {@code ANTHROPIC_*}, the parity-neutral git-forge grant, and the bridge MCP + reply charter
* mounted as inline launch flags.
*/
@Override
protected Launch buildLaunch(BridgedConfig.Worker cfg) {
String baseUrl = cfg.baseUrl();
guard.assertWorker(baseUrl); // hard stop before we spawn anything
Map<String, String> workerEnv = baseEnv(cfg);
workerEnv.put("ANTHROPIC_BASE_URL", baseUrl);
putIfPresent(workerEnv, "ANTHROPIC_MODEL", cfg.model());
putIfPresent(workerEnv, "CLAUDE_CONFIG_DIR", cfg.configDir());
putIfPresent(workerEnv, "ANTHROPIC_AUTH_TOKEN", env.apply(cfg.tokenEnv()));
applyGitToken(workerEnv, cfg);
return new Launch(workerEnv, argvWithBridge(cfg));
}
/**
* The launch argv, plus — when {@code worker.mcpUrl} is set — inline {@code --mcp-config} for
* the bridge server and {@code --append-system-prompt} for the {@link #REPLY_CHARTER}. Neither
* touches the profile's config; both are pure command-line flags. This inline-flag mount is
* Claude Code specific — other adapters mount MCP and instructions their own way.
*/
private List<String> argvWithBridge(BridgedConfig.Worker cfg) {
if (!cfg.hasMcp()) {
return cfg.argv();
}
String mcpJson = "{\"mcpServers\":{\"bridge\":{\"type\":\"http\",\"url\":\""
+ cfg.mcpUrl() + "\"}}}";
List<String> argv = mutableArgv(cfg.argv());
argv.add("--mcp-config");
argv.add(mcpJson);
argv.add("--append-system-prompt");
argv.add(REPLY_CHARTER);
return argv;
}
// --- Agent-returning convenience spawns (used by callers/tests that want the herdr Agent) ---
/** Spawn a worker for the default profile in the resolved default cwd. */
public Agent spawn() {
return spawnInternal(null, null, null);
}
/** Spawn a worker for a named profile (null → default) in the resolved default cwd. */
public Agent spawn(String profileName) {
return spawnInternal(profileName, null, null);
}
/** Spawn a worker for a named profile with an explicit requested/caller cwd (CB-112). */
public Agent spawn(String profileName, String requestedCwd, String callerCwd) {
return spawnInternal(profileName, requestedCwd, callerCwd);
}
// --- capabilities --------------------------------------------------------------------------
@Override
public Set<Capability> capabilities() {
Set<Capability> caps = EnumSet.of(Capability.MID_TURN_ASK, Capability.WORKTREE, Capability.ORPHAN_REAP);
if (hasGitTokenProfile()) {
caps.add(Capability.SELF_PR);
}
return Set.copyOf(caps);
}
/** Whether any configured profile opts into a git-forge token (required for {@link Capability#SELF_PR}). */
private boolean hasGitTokenProfile() {
return profileConfigs().stream().anyMatch(BridgedConfig.Worker::hasGitToken);
}
// --- CB-117 reap predicate (Claude prefix), kept for direct unit testing -------------------
/**
* Whether {@code name} is a Claude Code bridge worker started by a <em>different</em> process
* than {@code currentNonce}. A thin {@code claude}-prefix binding of
* {@link HerdrPeerLauncher#isForeignWorker(String, String, String)}.
*/
static boolean isForeignWorker(String name, String currentNonce) {
return HerdrPeerLauncher.isForeignWorker(NAME_PREFIX, name, currentNonce);
}
}
@@ -1,160 +0,0 @@
package dev.ltms.bridged.worker;
import dev.ltms.bridged.herdr.Agent;
import dev.ltms.bridged.peer.Capability;
import dev.ltms.bridged.peer.PeerHandle;
import dev.ltms.bridged.peer.PeerLauncher;
import dev.ltms.bridged.peer.SpawnRequest;
import org.slf4j.Logger;
import org.slf4j.LoggerFactory;
import java.util.EnumSet;
import java.util.LinkedHashMap;
import java.util.List;
import java.util.Map;
import java.util.Set;
import java.util.concurrent.ConcurrentHashMap;
/**
* The {@link PeerLauncher} the core actually holds when more than one adapter is configured — a thin
* router in front of one {@link HerdrPeerLauncher} per peer {@code kind} (Claude Code, opencode, …).
* It owns no transport of its own; it dispatches each SPI call to the delegate that owns the profile
* involved, and fans the fleet-wide queries (list/reap/caps/profiles) across all delegates.
*
* <p>Routing rules:
* <ul>
* <li><strong>By profile</strong> — {@link #spawn}, {@link #effectiveCwd}, {@link #parityOverlay}
* resolve the profile (a null/blank name → the global {@link #defaultProfile}) and delegate to
* the single adapter that declares it. Profiles partition cleanly across adapters: the
* constructor rejects a name claimed by two.</li>
* <li><strong>By pane id</strong> — {@link #stop} routes to the adapter that spawned that pane
* (recorded at spawn time). A pane the composite never spawned (only real for a caller that
* hand-rolls an id) falls back to the first delegate; teardown is pane-id addressed and
* tab cleanup is single-occupant guarded, so it is safe either way.</li>
* <li><strong>Fleet-wide</strong> — {@link #reapOrphanWorkers} and {@link #capabilities} fan out
* and combine. {@link #list} is deduplicated by pane id because every herdr-backed delegate
* shares one herdr connection and so reports the same global agent set.</li>
* </ul>
*/
public final class CompositePeerLauncher implements PeerLauncher {
private static final Logger log = LoggerFactory.getLogger(CompositePeerLauncher.class);
private final List<HerdrPeerLauncher> delegates;
private final Map<String, HerdrPeerLauncher> byProfile;
private final String defaultProfile;
/** paneId → the delegate that spawned it, so {@link #stop} tears down through the right adapter. */
private final Map<String, HerdrPeerLauncher> spawnedBy = new ConcurrentHashMap<>();
/**
* @param delegates one adapter per configured peer kind; must be non-empty and declare
* disjoint profile-name sets
* @param defaultProfile the profile a no-argument spawn resolves to (may be null)
* @throws IllegalArgumentException if {@code delegates} is empty or two adapters claim one profile
*/
public CompositePeerLauncher(List<HerdrPeerLauncher> delegates, String defaultProfile) {
if (delegates.isEmpty()) {
throw new IllegalArgumentException("at least one peer adapter must be configured");
}
this.delegates = List.copyOf(delegates);
this.defaultProfile = defaultProfile;
Map<String, HerdrPeerLauncher> index = new LinkedHashMap<>();
for (HerdrPeerLauncher d : this.delegates) {
for (String profile : d.profiles()) {
HerdrPeerLauncher prev = index.putIfAbsent(profile, d);
if (prev != null) {
throw new IllegalArgumentException(
"worker profile '" + profile + "' is claimed by two peer adapters");
}
}
}
this.byProfile = Map.copyOf(index);
}
/** The adapter owning {@code profileName} (null/blank → the default). Throws on an unknown profile. */
private HerdrPeerLauncher route(String profileName) {
String resolved = (profileName == null || profileName.isBlank()) ? defaultProfile : profileName;
if (resolved == null) {
// No profile and no default configured — hand to the first delegate so it raises the
// same "no default" error it would on its own; keeps the SPI contract single-sourced.
return delegates.getFirst();
}
HerdrPeerLauncher d = byProfile.get(resolved);
if (d == null) {
throw new IllegalArgumentException("unknown worker profile: " + resolved);
}
return d;
}
@Override
public PeerHandle spawn(SpawnRequest req) {
HerdrPeerLauncher d = route(req.profileName());
PeerHandle handle = d.spawn(req);
spawnedBy.put(handle.id(), d);
return handle;
}
@Override
public String effectiveCwd(SpawnRequest req) {
return route(req.profileName()).effectiveCwd(req);
}
@Override
public List<String> parityOverlay(String profileName) {
return route(profileName).parityOverlay(profileName);
}
@Override
public void stop(String id) {
HerdrPeerLauncher d = spawnedBy.remove(id);
if (d == null) {
log.debug("stop({}) — no recorded owner, routing to the first adapter (pane-addressed)", id);
d = delegates.getFirst();
}
d.stop(id);
}
@Override
public Set<String> profiles() {
return byProfile.keySet();
}
@Override
public String defaultProfile() {
return defaultProfile;
}
/** Every herdr agent, deduplicated by pane id (all delegates share one herdr and list globally). */
@Override
public List<Agent> list() {
Map<String, Agent> byPane = new LinkedHashMap<>();
for (HerdrPeerLauncher d : delegates) {
for (Agent a : d.list()) {
if (a.paneId() != null) {
byPane.putIfAbsent(a.paneId(), a);
}
}
}
return List.copyOf(byPane.values());
}
@Override
public int reapOrphanWorkers() {
int reaped = 0;
for (HerdrPeerLauncher d : delegates) {
reaped += d.reapOrphanWorkers();
}
return reaped;
}
/** The union of every adapter's capabilities — a capability any adapter offers, the fleet offers. */
@Override
public Set<Capability> capabilities() {
EnumSet<Capability> caps = EnumSet.noneOf(Capability.class);
for (HerdrPeerLauncher d : delegates) {
caps.addAll(d.capabilities());
}
return Set.copyOf(caps);
}
}
@@ -1,546 +0,0 @@
package dev.ltms.bridged.worker;
import dev.ltms.bridged.config.BridgedConfig;
import dev.ltms.bridged.herdr.Agent;
import dev.ltms.bridged.herdr.AgentControl;
import dev.ltms.bridged.herdr.HerdrException;
import dev.ltms.bridged.herdr.Tab;
import dev.ltms.bridged.herdr.Workspace;
import dev.ltms.bridged.herdr.WorkspaceControl;
import dev.ltms.bridged.peer.PeerHandle;
import dev.ltms.bridged.peer.PeerLauncher;
import dev.ltms.bridged.peer.PeerUnreachableException;
import dev.ltms.bridged.peer.SpawnRequest;
import org.slf4j.Logger;
import org.slf4j.LoggerFactory;
import java.security.SecureRandom;
import java.util.ArrayList;
import java.util.Collection;
import java.util.LinkedHashMap;
import java.util.List;
import java.util.Map;
import java.util.Set;
import java.util.concurrent.atomic.AtomicLong;
import java.util.function.Function;
import java.util.function.LongSupplier;
import java.util.regex.Matcher;
import java.util.regex.Pattern;
/**
* Abstract base for {@link PeerLauncher} adapters that materialize a peer as a <em>herdr</em>
* agent (a CLI coding agent running in a herdr tab/pane). It owns everything that is the same
* regardless of <em>which</em> coding agent runs: tab/pane placement, the CB-306 spawn-readiness
* gate, unique naming, CB-117 orphan reap, teardown, {@link #list() listing}, and cwd resolution.
*
* <p>Two seams are peer-specific and supplied by the concrete adapter:
* <ul>
* <li>{@code namePrefix} (constructor arg) — the label prefix ({@code claude}, {@code opencode})
* that drives both unique naming and the orphan-reap pattern, so each adapter reaps only its
* own kind of pane and never another's.</li>
* <li>{@link #buildLaunch(BridgedConfig.Worker)} — the peer-specific env map + argv, including any
* subscription/guard check, MCP mount, and instruction injection. The base never sees how the
* peer is configured; it only places and starts the returned {@link Launch}.</li>
* </ul>
*
* <p>Placement: in the default {@code tab} policy a peer lands in its own tab inside a dedicated
* worker space (found-or-created once, then shared), so peers never split or clutter the user's
* real work spaces. Teardown removes the peer's pane <em>and</em> its now-empty tab, tolerating an
* already-gone peer so a repeated DELETE is harmless.
*/
public abstract class HerdrPeerLauncher implements PeerLauncher {
private static final Logger log = LoggerFactory.getLogger(HerdrPeerLauncher.class);
/** herdr rejects a duplicate agent {@code name}; we retry a bumped name this many times. */
private static final int NAME_RETRIES = 8;
private final String namePrefix; // label prefix: naming + reap scheme
private final AgentControl agents;
private final WorkspaceControl spaces;
private final Map<String, BridgedConfig.Worker> profiles; // profile name → spawn settings
private final String defaultProfile; // profile a no-arg spawn uses (nullable)
/** Host env lookup (injectable for tests); adapters read it in {@link #buildLaunch}. */
protected final Function<String, String> env;
private final AtomicLong nameSeq = new AtomicLong(); // per-peer counter (also the tab #)
private final long spawnReadyTimeoutMs; // 0 = disable gate (legacy non-blocking spawn)
private final LongSupplier nowMillis; // monotonic clock (injectable for tests)
private final Runnable sleeper; // sleep/wait hook (injectable for tests; never real-sleep in unit tests)
// Per-process token mixed into each peer name so a fresh process (nameSeq back at 0) cannot
// collide with same-profile peers that outlived a restart. See startUniquelyNamed.
private final String nameNonce = String.format("%06x", new SecureRandom().nextInt(1 << 24));
/**
* @param namePrefix label prefix for this peer kind (drives naming and reap)
* @param agents herdr agent control (start, status, close)
* @param spaces workspace / tab control (ensure, create, close)
* @param profiles configured peer profiles
* @param defaultProfile profile a no-argument spawn uses (nullable)
* @param env host env lookup (injectable for tests)
* @param spawnReadyTimeoutMs max ms to wait for injectable state (0 disables the gate)
* @param nowMillis monotonic clock source (e.g. {@code System::currentTimeMillis})
* @param sleeper sleep/wait hook (never called when the gate is disabled); the poll
* interval is baked into this hook, so the base needs no poll field
*/
protected HerdrPeerLauncher(String namePrefix, AgentControl agents, WorkspaceControl spaces,
Map<String, BridgedConfig.Worker> profiles, String defaultProfile,
Function<String, String> env,
long spawnReadyTimeoutMs,
LongSupplier nowMillis, Runnable sleeper) {
this.namePrefix = namePrefix;
this.agents = agents;
this.spaces = spaces;
this.profiles = Map.copyOf(profiles);
this.defaultProfile = defaultProfile;
this.env = env;
this.spawnReadyTimeoutMs = spawnReadyTimeoutMs;
this.nowMillis = nowMillis;
this.sleeper = sleeper;
}
// --- adapter seams -------------------------------------------------------------------------
/**
* Build the peer-specific launch for {@code cfg}: the environment map and argv handed to herdr.
* Any subscription/guard check, MCP mount, and instruction injection happen here. The env map
* and argv are adapter-private; the base only places and starts what is returned.
*/
protected abstract Launch buildLaunch(BridgedConfig.Worker cfg);
/** A peer-specific launch: the herdr {@code env} map and {@code argv}. */
protected record Launch(Map<String, String> env, List<String> argv) {
}
// --- profile surface -----------------------------------------------------------------------
/** The configured peer profile names (what {@code spawn(profile)} accepts). */
@Override
public Set<String> profiles() {
return profiles.keySet();
}
/** The parity-overlay file list for {@code profileName} (default list when unset). */
@Override
public List<String> parityOverlay(String profileName) {
String name = (profileName == null || profileName.isBlank()) ? defaultProfile : profileName;
if (name == null || name.isBlank()) {
return List.of();
}
BridgedConfig.Worker cfg = profiles.get(name);
return cfg == null ? List.of() : cfg.parityOverlay();
}
/** The profile a no-argument spawn uses, or {@code null} if none is configured. */
@Override
public String defaultProfile() {
return defaultProfile;
}
/** The configured profiles, for adapter capability decisions (e.g. any git-token grant). */
protected Collection<BridgedConfig.Worker> profileConfigs() {
return profiles.values();
}
/** Resolve {@code profileName} (null/blank → default) to its config, or throw with the options. */
protected BridgedConfig.Worker requireProfile(String profileName) {
String name = (profileName == null || profileName.isBlank()) ? defaultProfile : profileName;
if (name == null || name.isBlank()) {
throw new IllegalArgumentException("no default worker profile is configured — "
+ "pass a profile; configured: " + profiles.keySet());
}
BridgedConfig.Worker cfg = profiles.get(name);
if (cfg == null) {
throw new IllegalArgumentException("unknown worker profile '" + name
+ "' — configured: " + profiles.keySet());
}
return cfg;
}
// --- spawn ---------------------------------------------------------------------------------
/**
* Spawn a peer. {@code profileName} null/blank → the default profile. The working directory
* (CB-112) is resolved by {@link #resolveCwd}: an explicit {@code requestedCwd}, else the
* profile's configured {@code cwd}, else {@code callerCwd} (the primary's cwd, when the spawn
* came from the primary over MCP), else the daemon's cwd — never assumed to be {@code $HOME}.
* The adapter's {@link #buildLaunch} runs before any herdr call.
*/
protected Agent spawnInternal(String profileName, String requestedCwd, String callerCwd) {
BridgedConfig.Worker cfg = requireProfile(profileName);
Launch launch = buildLaunch(cfg);
String cwd = resolveCwd(requestedCwd, cfg, callerCwd);
return cfg.tabPlacement()
? spawnInTab(cfg, launch.env(), launch.argv(), cwd)
: spawnAsPane(cfg, launch.env(), launch.argv(), cwd);
}
/**
* {@inheritDoc}
*
* <p>Delegates to {@link #spawnInternal} and wraps the resulting herdr {@link Agent} in a
* {@link WorkerHandle} whose {@link PeerHandle#id()} equals the agent's paneId. When
* {@code spawnReadyTimeoutMs > 0}, blocks until the peer's herdr status is injectable or the
* timeout elapses; on timeout the pane is closed (no orphan) and a
* {@link PeerUnreachableException} is thrown.
*/
@Override
public PeerHandle spawn(SpawnRequest req) {
Agent agent = spawnInternal(req.profileName(), req.requestedCwd(), req.callerCwd());
String paneId = agent.paneId();
if (spawnReadyTimeoutMs > 0) {
waitUntilInjectableOrThrow(paneId);
}
return new WorkerHandle(paneId, agent.terminalId());
}
@Override
public String effectiveCwd(SpawnRequest req) {
return effectiveCwd(req.profileName(), req.requestedCwd(), req.callerCwd());
}
/**
* CB-301: the effective working directory a spawn for {@code profileName} would use, without
* actually spawning.
*/
private String effectiveCwd(String profileName, String requestedCwd, String callerCwd) {
return resolveCwd(requestedCwd, requireProfile(profileName), callerCwd);
}
/**
* CB-112 cwd resolution: spawn arg → profile config → the primary's cwd → the daemon's cwd.
* Never returns {@code null}/blank: {@code "."} (the daemon's own working directory) is the
* guaranteed last resort so a pathological environment with an unset {@code user.dir} still
* honours the "never assume {@code $HOME}" contract rather than letting herdr default the pane.
*/
private static String resolveCwd(String requestedCwd, BridgedConfig.Worker cfg, String callerCwd) {
return firstNonBlank(requestedCwd, cfg.cwd(), callerCwd, System.getProperty("user.dir"), ".");
}
private static String firstNonBlank(String... values) {
for (String v : values) {
if (v != null && !v.isBlank()) return v;
}
return null;
}
/** Dedicated worker space → own tab → start the peer (rooted at {@code cwd}) → drop the shell. */
private Agent spawnInTab(BridgedConfig.Worker cfg, Map<String, String> workerEnv,
List<String> argv, String cwd) {
Workspace space = spaces.ensureWorkspace(cfg.workspace());
Tab.Created tab = spaces.createTab(space.workspaceId());
log.info("spawning {} profile={} space={} tab={} cwd={}",
namePrefix, cfg.profile(), space.workspaceId(), tab.tab().tabId(), cwd);
Started started;
try {
started = startUniquelyNamed(cfg, workerEnv, argv, tab.tab().tabId(), cwd);
} catch (RuntimeException e) {
// The peer never started — don't leave the tab we just created orphaned.
// Best-effort cleanup; never let it mask the real spawn failure.
try {
spaces.closeTab(tab.tab().tabId());
} catch (RuntimeException cleanup) {
log.warn("failed to close orphaned tab {} after spawn error: {}",
tab.tab().tabId(), cleanup.getMessage());
}
throw e;
}
// The peer is LIVE now. The remaining steps are cosmetic (drop herdr's seed shell so the
// tab holds only the peer; label the tab). They must not fail the spawn or orphan the
// running peer — on error we log and still return it so the caller gets its paneId and can
// tear it down.
if (tab.rootPaneId() != null) {
tidy("close seed pane " + tab.rootPaneId(), () -> agents.close(tab.rootPaneId()));
} else {
log.warn("tab {} had no seed pane in the create response; peer tab may hold an extra pane",
tab.tab().tabId());
}
tidy("label tab " + tab.tab().tabId(),
() -> spaces.renameTab(tab.tab().tabId(), cfg.renderTabLabel(started.seq())));
log.info("{} started pane={} tab={} terminal={}",
namePrefix, started.agent().paneId(), started.agent().tabId(), started.agent().terminalId());
return started.agent();
}
/** Run a best-effort post-start cleanup step, logging (not throwing) on failure. */
private void tidy(String what, Runnable step) {
try {
step.run();
} catch (RuntimeException e) {
log.warn("post-start step failed ({}) — peer is running regardless: {}", what, e.getMessage());
}
}
/** Legacy placement: herdr splits the currently-focused tab; the peer still starts in {@code cwd}. */
private Agent spawnAsPane(BridgedConfig.Worker cfg, Map<String, String> workerEnv,
List<String> argv, String cwd) {
log.info("spawning {} (pane placement) profile={} cwd={} argv={}",
namePrefix, cfg.profile(), cwd, argv);
Agent peer = startUniquelyNamed(cfg, workerEnv, argv, null, cwd).agent();
log.info("{} started pane={} terminal={}", namePrefix, peer.paneId(), peer.terminalId());
return peer;
}
/** A started peer together with the sequence its unique name/label used. */
private record Started(Agent agent, long seq) {
}
/**
* Start the peer under a unique herdr agent name. herdr requires each running agent's
* {@code name} to be distinct (a 2nd identical {@code name} fails {@code agent_name_taken}) —
* the exact case that makes multiple peers useful. The name is
* {@code <prefix>-<profile>-<nonce>-<seq>}: {@code seq} distinguishes peers within this process,
* and the per-process {@code nonce} keeps a fresh process (whose {@code seq} restarts at 0) from
* colliding with same-profile peers that outlived a restart. The retry is a belt-and-braces
* backstop for the astronomically unlikely nonce+seq clash; the name is a label only — herdr
* detects kind and status from terminal output, not from it.
*/
private Started startUniquelyNamed(BridgedConfig.Worker cfg, Map<String, String> workerEnv,
List<String> argv, String tabId, String cwd) {
HerdrException last = null;
for (int attempt = 0; attempt < NAME_RETRIES; attempt++) {
long seq = nameSeq.incrementAndGet();
String name = namePrefix + "-" + cfg.profile() + "-" + nameNonce + "-" + seq;
try {
return new Started(agents.start(name, argv, workerEnv, tabId, cwd), seq);
} catch (HerdrException e) {
if (!"agent_name_taken".equals(e.code())) throw e;
log.debug("peer name '{}' taken, retrying", name);
last = e;
}
}
throw last;
}
// --- discovery + reap ----------------------------------------------------------------------
/** All herdr-tracked agents — discovery for "what peers exist". */
@Override
public List<Agent> list() {
return agents.list();
}
/**
* Reap peer panes left behind by an earlier daemon process (CB-117). herdr keeps a peer's pane
* alive across a daemon restart <em>by design</em>, and that pane's id is held only by its
* spawner — so a peer whose owning process exited before issuing the matching teardown leaks
* with nothing tracking it. On boot we scan herdr for agents whose name matches our
* {@code <prefix>-<profile>-<nonce>-<seq>} scheme with a nonce <em>other</em> than this
* process's {@link #nameNonce}, and tear each one down (its pane and, via {@link #stop}, its
* now-empty dedicated tab). A current-nonce peer is ours and live, so it is left running; a
* user's own session carries no such name and is never touched. A peer from a <em>different</em>
* adapter (different prefix) is likewise never touched. Best-effort: a failed listing, or a
* failure to stop any one peer, is logged and never aborts startup.
*
* @return the number of orphaned peers reaped
*/
@Override
public int reapOrphanWorkers() {
List<Agent> all;
try {
all = agents.list();
} catch (RuntimeException e) {
log.warn("orphan-peer reap skipped — agent.list failed: {}", e.getMessage());
return 0;
}
int reaped = 0;
for (Agent a : all) {
if (!isForeignWorker(namePrefix, a.name(), nameNonce)) continue;
try {
stop(a.paneId());
reaped++;
log.info("reaped orphan {} {} (pane={} tab={}) left by a prior daemon",
namePrefix, a.name(), a.paneId(), a.tabId());
} catch (RuntimeException e) {
log.warn("could not reap orphan {} {} (pane={}): {}",
namePrefix, a.name(), a.paneId(), e.getMessage());
}
}
if (reaped > 0) {
log.info("orphan-peer reap complete — {} stale {} peer(s) removed at startup", reaped, namePrefix);
}
return reaped;
}
/** The {@code <prefix>-<profile>-<nonce>-<seq>} name pattern; group 1 captures the 6-hex nonce. */
static Pattern workerNamePattern(String prefix) {
return Pattern.compile(prefix + "-.*-([0-9a-f]{6})-\\d+");
}
/**
* Whether {@code name} is a peer of kind {@code prefix} started by a <em>different</em> process
* than {@code currentNonce} — the reap predicate (CB-117). True only for the prefix's naming
* scheme with a foreign nonce: a non-peer name, a different adapter's name, or our own live
* nonce is excluded. Pure and package-private so the decision is unit-testable without herdr.
*/
static boolean isForeignWorker(String prefix, String name, String currentNonce) {
String nonce = workerNonce(prefix, name);
return nonce != null && !nonce.equals(currentNonce);
}
/** The 6-hex nonce embedded in a {@code prefix} peer name, or {@code null} if not one. */
static String workerNonce(String prefix, String name) {
if (name == null) return null;
Matcher m = workerNamePattern(prefix).matcher(name);
return m.matches() ? m.group(1) : null;
}
/** This process's peer-name nonce (a label component only; exposed for reaper tests). */
String nameNonce() {
return nameNonce;
}
// --- teardown ------------------------------------------------------------------------------
/**
* Tear a peer down by pane id: close the pane, and close its tab <em>only</em> when the peer is
* that tab's sole occupant. The single-pane check is what makes this safe regardless of how the
* peer was placed (or a placement-config change across a restart): a pane-placement peer sitting
* in one of the user's shared tabs has siblings, so its tab is never closed — we only ever
* remove a tab we created to hold one peer.
*
* <p>Resolves the tab from the pane <em>before</em> closing it. An already-gone pane/tab
* (repeated DELETE, crashed peer) is treated as success; any other failure propagates so a
* genuinely failed teardown is not reported as done.
*/
@Override
public void stop(String paneId) {
// Teardown knows only the paneId, not which profile spawned it. Attempt tab cleanup when any
// profile uses tab placement (so the bridge may have created a dedicated peer tab); the
// single-occupant check below is what actually protects the user's shared tabs.
WorkspaceControl.PaneLocation loc = usesTabPlacement() ? spaces.locatePane(paneId) : null;
try {
agents.close(paneId);
} catch (HerdrException e) {
if (!isAlreadyGone(e)) throw e;
log.debug("pane.close({}) ignored — already gone: {}", paneId, e.getMessage());
}
if (loc != null && loc.tabPaneCount() == 1) {
spaces.closeTab(loc.tabId());
} else if (loc != null) {
log.debug("not closing tab {} — it holds {} panes (not a dedicated peer tab)",
loc.tabId(), loc.tabPaneCount());
}
}
/** Whether any configured profile places peers in their own tab (so tabs may need cleanup). */
private boolean usesTabPlacement() {
return profiles.values().stream().anyMatch(BridgedConfig.Worker::tabPlacement);
}
/** True when a herdr error means the target is already gone (safe to treat as done). */
private static boolean isAlreadyGone(HerdrException e) {
return e.code() != null && e.code().endsWith("_not_found");
}
// --- spawn-readiness gate (CB-306) ---------------------------------------------------------
/**
* Poll {@link AgentControl#status} until the pane reports an injectable state or the configured
* timeout elapses. On timeout, close the pane (self-reap) and throw.
*/
private void waitUntilInjectableOrThrow(String paneId) {
long deadline = nowMillis.getAsLong() + spawnReadyTimeoutMs;
while (nowMillis.getAsLong() < deadline) {
if (agents.status(paneId).injectable()) {
log.debug("peer pane={} reached injectable state", paneId);
return;
}
sleeper.run();
}
log.warn("peer pane={} did not become injectable within {}ms — closing", paneId, spawnReadyTimeoutMs);
stop(paneId);
throw new PeerUnreachableException(
"worker pane " + paneId + " did not reach injectable state within "
+ spawnReadyTimeoutMs + "ms");
}
/** A concrete {@link PeerHandle} wrapping herdr agent coordinates. */
private record WorkerHandle(String id, String terminalId) implements PeerHandle {
}
// --- shared helpers ------------------------------------------------------------------------
/** Put {@code k → v} only when {@code v} is present (non-null, non-blank). */
protected static void putIfPresent(Map<String, String> m, String k, String v) {
if (v != null && !v.isBlank()) {
m.put(k, v);
}
}
/** Host env lookup that tolerates an unconfigured (null/blank) var name — returns null then. */
protected String resolveEnv(String name) {
return (name == null || name.isBlank()) ? null : env.apply(name);
}
/**
* The parity-neutral git-forge token grant (CB-302): when {@code cfg} opts in via
* {@code gitTokenEnv} and the token resolves, inject {@code GITEA_TOKEN} plus its paired
* {@code GITEA_HOST}. Push over SSH is unaffected; the only incremental grant is PR-create.
* Peer-neutral, so every herdr adapter reuses it unchanged.
*/
protected void applyGitToken(Map<String, String> workerEnv, BridgedConfig.Worker cfg) {
if (!cfg.hasGitToken()) {
return;
}
String gitToken = resolveEnv(cfg.gitTokenEnv());
if (gitToken != null) {
workerEnv.put("GITEA_TOKEN", gitToken);
putIfPresent(workerEnv, "GITEA_HOST", resolveEnv(cfg.gitHostEnv()));
}
}
/** A fresh mutable env map — the conventional starting point for {@link #buildLaunch}. */
/**
* Seed a worker's environment (CB-511): the daemon's own {@code PATH}, then the profile's
* {@code env:} entries.
*
* <p>Why this exists: bridged passes herdr an explicit env map, and herdr merges it into
* <em>its own</em> process environment. So before this, a worker inherited whatever PATH the
* herdr server happened to be started with — on this host, one from weeks earlier with no JDK
* and no Maven, which left workers unable to run the build they were being asked to run. The
* worker's toolchain must follow from configuration, not from how a long-lived daemon was
* launched.
*
* <p>Adapter-specific variables are layered on top of this by {@code buildLaunch} and therefore
* win. That ordering is deliberate and load-bearing: it stops a profile's {@code env:} from
* overriding {@code ANTHROPIC_BASE_URL} and slipping past {@link
* dev.ltms.bridged.guard.SubscriptionGuard}, which is checked against the profile's
* {@code baseUrl} and nothing else.
*/
protected Map<String, String> baseEnv(BridgedConfig.Worker cfg) {
Map<String, String> workerEnv = new LinkedHashMap<>();
String path = env.apply("PATH");
if (path != null && !path.isBlank()) {
workerEnv.put("PATH", path);
}
if (cfg != null && cfg.env() != null) {
workerEnv.putAll(cfg.env());
}
return workerEnv;
}
/** Defensive copy of {@code argv} plus room to append launch flags. */
protected static List<String> mutableArgv(List<String> argv) {
return new ArrayList<>(argv);
}
/**
* Uninterruptible sleep — the production {@link #sleeper}. Tests supply their own no-op /
* fast-faking sleeper so they never real-sleep.
*/
protected static void sleepUninterruptibly(long ms) {
try {
Thread.sleep(ms);
} catch (InterruptedException e) {
Thread.currentThread().interrupt();
// preserve the interrupt flag but continue — poll loops should not be aborted by an
// interrupt that was not meant for them.
}
}
}
+9
View File
@@ -1,4 +1,13 @@
<configuration>
<!--
CB-575: MCP SDK 2.0.0 has no public notification registration API. It registers only
notifications/initialized and notifications/roots/list_changed, so other client notifications
still warn when unhandled. Clients may legitimately send notifications/cancelled; suppress only
that SDK WARN because it would devalue the action-needed WARN level used by M4 fleet health.
If a later SDK handles cancellation, this filter simply stops matching and can be removed.
-->
<turboFilter class="dev.ltms.bridged.logging.McpCancelledNotificationFilter"/>
<appender name="STDOUT" class="ch.qos.logback.core.ConsoleAppender">
<encoder>
<pattern>%d{HH:mm:ss.SSS} %-5level [%thread] %logger{28} - %msg%n</pattern>
@@ -0,0 +1,86 @@
package dev.ltms.bridged;
import dev.ltms.bridged.inject.MemberPresence;
import org.junit.jupiter.api.DisplayName;
import org.junit.jupiter.api.Test;
import java.util.HashMap;
import java.util.Map;
import java.util.function.Predicate;
import java.util.function.Supplier;
import static org.junit.jupiter.api.Assertions.assertFalse;
import static org.junit.jupiter.api.Assertions.assertTrue;
/**
* CB-534: the injector's readiness gate must open for a lead as well as for a present worker.
*
* <p>The bug these cover was silent and slow: a lead was never marked present (only workers are), so
* every lead→lead delivery sat on the gate for the full readiness grace and failed ~60s later without
* a keystroke ever reaching the pane.
*/
class BridgedDeliverabilityTest {
private static Supplier<Map<String, String>> leads(Map<String, String> m) {
return () -> m;
}
@Test
@DisplayName("a worker that has connected its MCP is deliverable")
void presentWorkerIsDeliverable() {
MemberPresence presence = new MemberPresence();
presence.markPresent("term_worker");
assertTrue(Bridged.deliverableTo(presence, leads(Map.of())).test("term_worker"));
}
@Test
@DisplayName("a worker still in its boot window is held back")
void absentWorkerIsNotDeliverable() {
assertFalse(Bridged.deliverableTo(new MemberPresence(), leads(Map.of())).test("term_booting"));
}
@Test
@DisplayName("a lead is deliverable without ever being marked present")
void leadIsDeliverableWithoutPresence() {
MemberPresence presence = new MemberPresence();
Predicate<String> deliverable =
Bridged.deliverableTo(presence, leads(Map.of("term_lead", "opus-5.0")));
assertFalse(presence.isPresent("term_lead"), "a lead is never enrolled in worker presence");
assertTrue(deliverable.test("term_lead"), "…and must be deliverable anyway");
}
@Test
@DisplayName("an unknown terminal is deliverable to neither")
void strangerIsNotDeliverable() {
MemberPresence presence = new MemberPresence();
presence.markPresent("term_worker");
assertFalse(Bridged.deliverableTo(presence, leads(Map.of("term_lead", "opus-5.0")))
.test("term_stranger"));
}
@Test
@DisplayName("a lead discovered after startup becomes deliverable with no restart")
void leadSetIsReadThroughOnEveryCall() {
Map<String, String> discovered = new HashMap<>();
Predicate<String> deliverable = Bridged.deliverableTo(new MemberPresence(), leads(discovered));
assertFalse(deliverable.test("term_late"));
discovered.put("term_late", "gpt-sol-5.6"); // leadScan picks up a newly labelled tab
assertTrue(deliverable.test("term_late"), "the supplier must be re-read, not snapshotted");
}
@Test
@DisplayName("forgetting a torn-down worker does not strip a lead of its deliverability")
void forgetDoesNotDisarmALead() {
MemberPresence presence = new MemberPresence();
Predicate<String> deliverable =
Bridged.deliverableTo(presence, leads(Map.of("term_lead", "opus-5.0")));
presence.forget("term_lead"); // the injector's cleanup path runs against every target
assertTrue(deliverable.test("term_lead"));
}
}
@@ -12,6 +12,8 @@ class AuthzTest {
private static final Principal WORKER_A = Principal.worker("term_a", 200);
private static final Principal WORKER_B = Principal.worker("term_b", 300);
private static final Principal ANON = Principal.anonymous();
private static final Principal ARCH_DESIGN = Principal.architect("lead-designer", "term_design", 400);
private static final Principal ARCH_OTHER = Principal.architect("reviewer", "term_review", 500);
@Test
void anonymousIsAuthorizedForNothing() {
@@ -61,6 +63,48 @@ class AuthzTest {
"an absent session id must not satisfy the own-session rule");
}
// ── CB-548: the architect matrix ───────────────────────────────────────────────────────────
@Test
void anArchitectMaySendButNotSpawnStopOrDrain() {
assertTrue(Authz.permits(ARCH_DESIGN, SEND, "term_worker"),
"delegating a turn to a worker IS the architect's job");
assertTrue(Authz.permits(ARCH_DESIGN, SEND, null));
for (Authz.Action a : new Authz.Action[]{SPAWN, STOP, DRAIN}) {
assertFalse(Authz.permits(ARCH_DESIGN, a, null),
"an architect must not " + a + " — fleet lifecycle is the primary's alone, so "
+ "a coordinator cannot also stand up or tear down the fleet");
}
}
@Test
void anArchitectMayReplyAndAskOnlyAsItsOwnPane() {
assertTrue(Authz.permits(ARCH_DESIGN, REPLY, "term_design"), "its own pane is its own");
assertTrue(Authz.permits(ARCH_DESIGN, ASK, "term_design"));
assertFalse(Authz.permits(ARCH_DESIGN, REPLY, "term_review"),
"architect 'lead-designer' must not reply on reviewer's pane");
assertFalse(Authz.permits(ARCH_OTHER, ASK, "term_design"),
"reviewer must not ask as lead-designer — no terminal is another's");
assertFalse(Authz.permits(ARCH_DESIGN, REPLY, null),
"an absent target must not pass the own-session rule");
}
@Test
void anArchitectMayReadAndScrapeMetrics() {
assertTrue(Authz.permits(ARCH_DESIGN, READ, null));
assertTrue(Authz.permits(ARCH_DESIGN, METRICS, null));
}
@Test
void anArchitectIsNotCountedAsPrimaryOrWorker() {
assertFalse(Authz.permits(ARCH_DESIGN, SPAWN, null), "not a primary — no lifecycle");
assertFalse(ARCH_DESIGN.isPrimary());
assertFalse(ARCH_DESIGN.isWorker(), "an architect is its own role, not a widened worker");
assertTrue(ARCH_DESIGN.isArchitect());
}
@Test
void observationIsOpenToBothAuthenticatedRoles() {
assertTrue(Authz.permits(PRIMARY, READ, null));
@@ -3,8 +3,12 @@ package dev.ltms.bridged.auth;
import dev.ltms.bridged.herdr.FakeHerdr;
import dev.ltms.bridged.herdr.PaneLocator;
import dev.ltms.bridged.mcp.ConnectionIdentity;
import dev.ltms.bridged.config.BridgedConfig;
import dev.ltms.bridged.peer.MemberRole;
import org.junit.jupiter.api.Test;
import java.util.Map;
import static org.junit.jupiter.api.Assertions.*;
/**
@@ -30,6 +34,17 @@ class CallerResolverTest {
return identity(999_999);
}
private static MemberRegistry boundMembers(String slot, MemberRole role) {
Map<String, BridgedConfig.Slot> architects = role == MemberRole.ARCHITECT
? Map.of("lead-designer", new BridgedConfig.Slot("sonnet")) : Map.of();
Map<String, BridgedConfig.Slot> devs = role == MemberRole.DEV
? Map.of("builder", new BridgedConfig.Slot("sonnet")) : Map.of();
MemberRegistry members = new MemberRegistry(
new BridgedConfig.Fleet(Map.of(), architects, devs, Map.of(), null));
assertTrue(members.bind(slot, "term_a"));
return members;
}
@Test
void aLoopbackWorkerPaneResolvesToWorkerRegardlessOfAuthMode() {
Principal underTrust = new CallerResolver(workerIdentity()).resolve("127.0.0.1", 42, null);
@@ -44,6 +59,43 @@ class CallerResolverTest {
assertEquals("term_a", underToken.terminal());
}
@Test
void aPinnedPrimaryTerminalResolvesToPrimaryNotWorker() {
// The primary's own session lives in a herdr pane (term_a here). Without the pin the pane
// match wins and the primary is locked out of spawn/send/stop as a misread worker.
Principal p = CallerResolver.pinnedTo(workerIdentity(), false, null, "term_a")
.resolve("127.0.0.1", 42, null);
assertEquals(Role.PRIMARY, p.role());
}
@Test
void aPinnedPrimaryTerminalNeedsNoTokenEvenInTokenMode() {
Principal p = CallerResolver.pinnedTo(workerIdentity(), true, "s3cret", "term_a")
.resolve("127.0.0.1", 42, null);
assertEquals(Role.PRIMARY, p.role(),
"the pane mapping is as unforgeable as a worker's — the pin outranks the token path");
}
@Test
void otherPanesRemainWorkersWhenAPinIsSet() {
Principal p = CallerResolver.pinnedTo(workerIdentity(), false, null, "term_someone_else")
.resolve("127.0.0.1", 42, null);
assertEquals(Role.WORKER, p.role());
assertEquals("term_a", p.terminal());
}
/** The pin is optional config, so an absent or whitespace one must change nothing at all. */
@Test
void aBlankPinLeavesWorkerResolutionUntouched() {
assertEquals(Role.WORKER,
CallerResolver.pinnedTo(workerIdentity(), false, null, " ").resolve("127.0.0.1", 42, null).role());
assertEquals(Role.WORKER,
CallerResolver.pinnedTo(workerIdentity(), false, null, null).resolve("127.0.0.1", 42, null).role());
}
@Test
void loopbackTrustTreatsANonWorkerLoopbackCallerAsThePrimary() {
Principal p = new CallerResolver(nonWorkerIdentity()).resolve("127.0.0.1", 99, null);
@@ -96,6 +148,264 @@ class CallerResolverTest {
assertEquals(Role.ANONYMOUS, p.role());
}
// ── CB-530: the leaders registry ────────────────────────────────────────────────────────────
// The fake resolves exactly one pane (term_a) from a PID, so "two leads both resolve" is
// asserted at the config layer (BridgedConfigTest#leaderTerminals…). What matters here is that
// resolution is a REGISTRY LOOKUP rather than a single equality test against one pin.
@Test
void aRegisteredLeadPaneResolvesToPrimaryCarryingItsName() {
Principal p = new CallerResolver(workerIdentity(), false, null, Map.of("term_a", "opus-5.0"))
.resolve("127.0.0.1", 42, null);
assertEquals(Role.PRIMARY, p.role());
assertEquals("opus-5.0", p.name(), "whoami must be able to say WHICH lead is asking");
assertEquals("term_a", p.terminal(),
"CB-532: a lead carries the pane it was matched by. Without it ownsSession() can "
+ "never be true for a lead, so it can send to a peer but never answer one");
}
@Test
void aSecondLeadIsRecognisedRatherThanSilentlyDemoted() {
// The regression this feature exists for: with a singular pin, whichever lead was not the
// pin resolved as a worker and was refused every orchestration call.
Principal p = new CallerResolver(workerIdentity(), false, null,
Map.of("term_elsewhere", "gpt-sol-5.6", "term_a", "opus-5.0"))
.resolve("127.0.0.1", 42, null);
assertEquals(Role.PRIMARY, p.role());
assertEquals("opus-5.0", p.name());
}
@Test
void aPaneAbsentFromTheRegistryIsStillAWorker() {
Principal p = new CallerResolver(workerIdentity(), false, null,
Map.of("term_elsewhere", "gpt-sol-5.6"))
.resolve("127.0.0.1", 42, null);
assertEquals(Role.WORKER, p.role());
assertEquals("term_a", p.terminal());
assertNull(p.name());
}
@Test
void aRegisteredLeadNeedsNoTokenEvenInTokenMode() {
Principal p = new CallerResolver(workerIdentity(), true, "s3cret", Map.of("term_a", "opus"))
.resolve("127.0.0.1", 42, null);
assertEquals(Role.PRIMARY, p.role(),
"the pane mapping is as unforgeable as a worker's — it outranks the token path");
assertEquals("opus", p.name());
}
/** The pre-CB-530 spelling must keep working, exactly, including for configs that never migrate. */
@Test
void theLegacySinglePinBehavesAsALeadNamedPrimary() {
Principal p = CallerResolver.pinnedTo(workerIdentity(), false, null, "term_a")
.resolve("127.0.0.1", 42, null);
assertEquals(Role.PRIMARY, p.role());
assertEquals("primary", p.name());
}
@Test
void anEmptyRegistryLeavesEveryPaneAWorker() {
Map<String, String> noLeads = null;
assertEquals(Role.WORKER,
new CallerResolver(workerIdentity(), false, null, Map.of())
.resolve("127.0.0.1", 42, null).role());
assertEquals(Role.WORKER,
new CallerResolver(workerIdentity(), false, null, noLeads)
.resolve("127.0.0.1", 42, null).role());
// CB-531: and the same for the live-registry form, whose supplier may also be absent.
assertEquals(Role.WORKER,
CallerResolver.withLeads(workerIdentity(), false, null, null)
.resolve("127.0.0.1", 42, null).role());
}
/** The audit line must distinguish leads once several exist, or a log says nothing useful. */
@Test
void describeNamesTheLeadButStillReadsPrimaryWhenUnnamed() {
assertEquals("leader:opus-5.0", Principal.leader("opus-5.0", "term_a", 1).describe());
assertEquals("primary", Principal.primary(1).describe());
assertEquals("worker:term_a", Principal.worker("term_a", 1).describe());
}
// ── CB-532: a lead is an addressable peer, not only a sender ────────────────────────────────
/**
* The regression this ticket exists for: two leads could both be recognised (CB-530/531) and
* still not converse, because REPLY is gated on ownsSession() and a lead owned nothing.
*/
@Test
void aLeadOwnsItsOwnPaneSoItMayAnswerAPeer() {
Principal lead = new CallerResolver(workerIdentity(), false, null, Map.of("term_a", "opus-5.0"))
.resolve("127.0.0.1", 42, null);
assertTrue(lead.ownsSession("term_a"));
assertTrue(Authz.permits(lead, Authz.Action.REPLY, "term_a"),
"a lead answering a peer replies for its OWN terminal — the rendezvous the sender "
+ "opened is keyed on exactly that");
assertTrue(Authz.permits(lead, Authz.Action.ASK, "term_a"));
}
@Test
void aLeadStillCannotActAsAnyoneElse() {
Principal lead = new CallerResolver(workerIdentity(), false, null, Map.of("term_a", "opus-5.0"))
.resolve("127.0.0.1", 42, null);
assertFalse(lead.ownsSession("term_someone_else"));
assertFalse(Authz.permits(lead, Authz.Action.REPLY, "term_someone_else"),
"widening WHO may reply must not widen WHAT they may reply as");
}
/** A primary with no pane — token mode, or off-host — owns nothing and must stay a sender only. */
@Test
void anUnnamedPrimaryWithNoPaneOwnsNothing() {
Principal p = new CallerResolver(nonWorkerIdentity(), true, "s3cret")
.resolve("127.0.0.1", 99, "Bearer s3cret");
assertEquals(Role.PRIMARY, p.role());
assertNull(p.terminal());
assertFalse(p.ownsSession(null), "a null terminal must never match a null session id");
assertFalse(Authz.permits(p, Authz.Action.REPLY, null));
}
@Test
void aLeadKeepsEveryOrchestrationRightItAlreadyHad() {
Principal lead = new CallerResolver(workerIdentity(), false, null, Map.of("term_a", "opus-5.0"))
.resolve("127.0.0.1", 42, null);
assertTrue(Authz.permits(lead, Authz.Action.SPAWN, null));
assertTrue(Authz.permits(lead, Authz.Action.SEND, "term_worker"));
assertTrue(Authz.permits(lead, Authz.Action.STOP, null));
assertTrue(Authz.permits(lead, Authz.Action.DRAIN, null));
}
/**
* CB-531: the registry is read per resolve, not snapshotted at construction — a lead that
* labels its tab after the daemon booted is recognised without a restart.
*/
@Test
void aLeadRegisteredAfterConstructionIsHonouredWithoutRebuildingTheResolver() {
Map<String, String> live = new java.util.HashMap<>();
CallerResolver r = CallerResolver.withLeads(workerIdentity(), false, null, () -> live);
assertEquals(Role.WORKER, r.resolve("127.0.0.1", 42, null).role());
live.put("term_a", "gpt-sol-5.6"); // the scanner sees a newly-labelled tab
Principal p = r.resolve("127.0.0.1", 42, null);
assertEquals(Role.PRIMARY, p.role());
assertEquals("gpt-sol-5.6", p.name());
}
/** The map form must stay a snapshot: a caller handing over a map is not offering live state. */
@Test
void theMapFormIsCopiedSoLaterMutationCannotGrantLeadership() {
Map<String, String> mutable = new java.util.HashMap<>();
CallerResolver r = new CallerResolver(workerIdentity(), false, null, mutable);
mutable.put("term_a", "sneaky");
assertEquals(Role.WORKER, r.resolve("127.0.0.1", 42, null).role());
}
// ── CB-548: architect slots ─────────────────────────────────────────────────────────────────
@Test
void aBoundArchitectPaneResolvesToArchitectBeforeTheWorkerFallback() {
// This is the production construction path used by Bridged.
Principal p = CallerResolver.withLeadsAndMembers(workerIdentity(), false, null,
Map::of, boundMembers("architect:lead-designer", MemberRole.ARCHITECT))
.resolve("127.0.0.1", 42, null);
assertEquals(Role.ARCHITECT, p.role(),
"a terminal bound to an architect slot is an architect, NOT the generic worker it "
+ "would otherwise resolve to");
assertEquals("lead-designer", p.name(), "whoami must say WHICH slot is asking");
assertEquals("term_a", p.terminal(), "the pane identity is carried so ownsSession works");
}
@Test
void anArchitectNeedsNoTokenEvenInTokenMode() {
Principal p = CallerResolver.withLeadsAndMembers(workerIdentity(), true, "s3cret",
Map::of, boundMembers("architect:lead-designer", MemberRole.ARCHITECT))
.resolve("127.0.0.1", 42, null);
assertEquals(Role.ARCHITECT, p.role(),
"the pane mapping is as unforgeable as a worker's — it outranks the token path");
}
@Test
void anUnboundPaneStillResolvesAsAWorker() {
MemberRegistry members = new MemberRegistry(new BridgedConfig.Fleet(Map.of(),
Map.of("lead-designer", new BridgedConfig.Slot("sonnet")), Map.of(), Map.of(), null));
Principal p = CallerResolver.withLeadsAndMembers(workerIdentity(), false, null,
Map::of, members).resolve("127.0.0.1", 42, null);
assertEquals(Role.WORKER, p.role());
assertNull(p.name());
}
/** CB-548 precedence: lead > architect > worker, so a pane named in BOTH is still a lead. */
@Test
void aLeadWinsOverAnArchitectBindingForTheSamePane() {
Principal p = CallerResolver.withLeadsAndMembers(workerIdentity(), false, null,
() -> Map.of("term_a", "opus-5.0"),
boundMembers("architect:lead-designer", MemberRole.ARCHITECT))
.resolve("127.0.0.1", 42, null);
assertEquals(Role.PRIMARY, p.role(),
"a pane the config calls a lead must keep resolving as a lead — no behaviour change "
+ "when an architect binding is added to an existing fleet");
assertEquals("opus-5.0", p.name());
}
/** The registry is live, like leads: a binding injected after construction is honoured. */
@Test
void anArchitectBoundAfterConstructionIsHonouredWithoutRebuildingTheResolver() {
MemberRegistry members = new MemberRegistry(new BridgedConfig.Fleet(Map.of(),
Map.of("lead-designer", new BridgedConfig.Slot("sonnet")), Map.of(), Map.of(), null));
CallerResolver r = CallerResolver.withLeadsAndMembers(workerIdentity(), false, null,
Map::of, members);
assertEquals(Role.WORKER, r.resolve("127.0.0.1", 42, null).role());
assertTrue(members.bind("architect:lead-designer", "term_a")); // the later lifecycle binds the slot
assertEquals(Role.ARCHITECT, r.resolve("127.0.0.1", 42, null).role());
assertEquals("architect:lead-designer", r.members().get("term_a"));
}
@Test
void describeNamesTheArchitectSlot() {
assertEquals("architect:lead-designer",
Principal.architect("lead-designer", "term_a", 1).describe());
}
/** An architect acts only as its own pane — the same ownsSession rule as a worker or lead. */
@Test
void anArchitectOwnsItsOwnPaneAndNoOther() {
Principal arch = CallerResolver.withLeadsAndMembers(workerIdentity(), false, null,
Map::of, boundMembers("architect:lead-designer", MemberRole.ARCHITECT))
.resolve("127.0.0.1", 42, null);
assertTrue(arch.ownsSession("term_a"));
assertTrue(Authz.permits(arch, Authz.Action.REPLY, "term_a"));
assertFalse(arch.ownsSession("term_b"));
assertFalse(Authz.permits(arch, Authz.Action.REPLY, "term_b"));
}
@Test
void aBoundNonArchitectSlotStillResolvesAsAWorker() {
Principal p = CallerResolver.withLeadsAndMembers(workerIdentity(), false, null,
Map::of, boundMembers("dev:builder", MemberRole.DEV))
.resolve("127.0.0.1", 42, null);
assertEquals(Role.WORKER, p.role(), "a dev binding must never grant architect rights");
}
@Test
void tokenModeRequiresANonEmptyConfiguredToken() {
ConnectionIdentity id = nonWorkerIdentity();
@@ -0,0 +1,269 @@
package dev.ltms.bridged.auth;
import dev.ltms.bridged.config.BridgedConfig;
import dev.ltms.bridged.peer.MemberRole;
import org.junit.jupiter.api.Test;
import java.util.ArrayList;
import java.util.LinkedHashMap;
import java.util.List;
import java.util.Map;
import java.util.Set;
import java.util.concurrent.CountDownLatch;
import java.util.concurrent.ExecutorService;
import java.util.concurrent.Executors;
import java.util.concurrent.Future;
import static org.junit.jupiter.api.Assertions.*;
/**
* CB-548 — the architect-slot registry: the config snapshot of slot → profile, and the live
* terminal → slot bindings it owns. The role a binding produces is asserted in
* {@link CallerResolverTest}; this pins the registry object itself — its invariants and their
* thread-safety.
*/
class MemberRegistryTest {
/**
* Slot keys are qualified by role (CB-557): a bare name is unique only within its pool, so
* {@code sonnet} can be both a developer and a reviewer, while a terminal binds to exactly one.
*/
private static final String DESIGNER = "architect:lead-designer";
private static final String REVIEWER = "architect:code-reviewer";
private static BridgedConfig.Fleet fleetWith(Map<String, BridgedConfig.Slot> architects) {
return new BridgedConfig.Fleet(Map.of(), architects, Map.of(), Map.of(), null);
}
private static Map<String, BridgedConfig.Slot> architects() {
Map<String, BridgedConfig.Slot> pool = new LinkedHashMap<>();
pool.put("lead-designer", new BridgedConfig.Slot("sonnet"));
pool.put("code-reviewer", new BridgedConfig.Slot("gx10"));
return pool;
}
private final MemberRegistry registry = new MemberRegistry(fleetWith(architects()));
@Test
void exposesTheConfiguredSlotsQualifiedByRole() {
assertEquals(Set.of(DESIGNER, REVIEWER), registry.slots().keySet());
assertTrue(registry.isSlot(REVIEWER));
assertFalse(registry.isSlot("nope"));
assertFalse(registry.isSlot("lead-designer"),
"the bare name is not the key — it is unique only inside its pool");
}
/** The case the pool shape exists for: one profile serving two roles is not a duplicate. */
@Test
void oneProfileMayServeTwoRolesUnderTheSameSlotName() {
Map<String, BridgedConfig.Slot> devs = Map.of("sonnet", new BridgedConfig.Slot("sonnet"));
Map<String, BridgedConfig.Slot> revs = Map.of("sonnet", new BridgedConfig.Slot("sonnet"));
MemberRegistry r = new MemberRegistry(
new BridgedConfig.Fleet(Map.of(), Map.of(), devs, revs, null));
assertEquals(Set.of("dev:sonnet", "reviewer:sonnet"), r.slots().keySet());
assertEquals(MemberRole.DEV, r.roleForSlot("dev:sonnet"));
assertEquals(MemberRole.REVIEWER, r.roleForSlot("reviewer:sonnet"));
}
@Test
void theSpawnLifecycleReadsTheProfileBackFromASlot() {
assertEquals("sonnet", registry.profileForSlot(DESIGNER));
assertEquals("gx10", registry.profileForSlot(REVIEWER));
assertNull(registry.profileForSlot("unknown"), "an unknown slot has no profile");
}
@Test
void startsEmptySoNoTerminalResolvesToAnArchitect() {
assertTrue(registry.snapshot().isEmpty());
assertNull(registry.slotForTerminal("term_design"),
"config declares no architect terminal — nothing is recognised until a bind");
assertNull(registry.slotForTerminal(null), "no terminal ⇒ no slot");
}
// ── bind ──────────────────────────────────────────────────────────────────────────────────
@Test
void bindResolvesTheTerminalToTheSlot() {
assertTrue(registry.bind(DESIGNER, "term_design"));
assertEquals(DESIGNER, registry.slotForTerminal("term_design"));
assertEquals(Map.of("term_design", DESIGNER), registry.snapshot());
}
@Test
void bindRefusesAnUnknownSlot() {
assertFalse(registry.bind("nope", "term_x"),
"a slot that is not configured must be refused — bind is not a way to invent one");
assertNull(registry.slotForTerminal("term_x"));
}
@Test
void bindRefusesATerminalInTwoSlots() {
assertTrue(registry.bind(DESIGNER, "term_design"));
assertFalse(registry.bind(REVIEWER, "term_design"),
"a terminal may occupy at most one slot");
assertEquals(DESIGNER, registry.slotForTerminal("term_design"),
"the first binding survives the refused second");
}
@Test
void bindRefusesASlotWithTwoTerminals() {
assertTrue(registry.bind(DESIGNER, "term_design"));
assertFalse(registry.bind(DESIGNER, "term_other"),
"a slot may host at most one terminal");
assertEquals(DESIGNER, registry.slotForTerminal("term_design"),
"the first binding survives the refused second");
assertNull(registry.slotForTerminal("term_other"));
}
@Test
void rebindingTheSamePairIsAnIdempotentNoOp() {
assertTrue(registry.bind(DESIGNER, "term_design"));
assertTrue(registry.bind(DESIGNER, "term_design"),
"the same terminal → slot is harmless to repeat");
assertEquals(1, registry.snapshot().size());
}
// ── unbind ────────────────────────────────────────────────────────────────────────────────
@Test
void unbindRemovesTheExactBinding() {
assertTrue(registry.bind(DESIGNER, "term_design"));
assertTrue(registry.unbind(DESIGNER, "term_design"));
assertNull(registry.slotForTerminal("term_design"));
assertTrue(registry.snapshot().isEmpty());
}
@Test
void aStaleUnbindDoesNotRemoveAReplacement() {
// Bind, tear down, and stand the slot back up with a NEW terminal.
assertTrue(registry.bind(DESIGNER, "term_design"));
registry.unbind(DESIGNER, "term_design");
assertTrue(registry.bind(DESIGNER, "term_new"));
// A late unbind naming the OLD terminal must not remove the replacement binding.
assertFalse(registry.unbind(DESIGNER, "term_design"));
assertEquals(DESIGNER, registry.slotForTerminal("term_new"),
"the replacement terminal stays bound");
}
@Test
void aStaleUnbindForATerminalThatMovedSlotsDoesNothing() {
// term_design starts in lead-designer, is torn down, and stands back up in a FREE slot.
assertTrue(registry.bind(DESIGNER, "term_design"));
registry.unbind(DESIGNER, "term_design");
assertTrue(registry.bind(REVIEWER, "term_design"));
// Unbinding against the slot it no longer occupies is refused; the new binding is intact.
assertFalse(registry.unbind(DESIGNER, "term_design"),
"the old slot must not unbind a terminal that moved elsewhere");
assertEquals(REVIEWER, registry.slotForTerminal("term_design"));
}
@Test
void unbindOfNothingIsAFalseNoOp() {
assertFalse(registry.unbind(DESIGNER, "term_design"),
"nothing was bound, so nothing is removed");
}
// ── snapshot ─────────────────────────────────────────────────────────────────────────────
@Test
void theSnapshotIsAnImmutableCopyNotAliveState() {
assertTrue(registry.bind(DESIGNER, "term_design"));
Map<String, String> snap = registry.snapshot();
assertThrows(UnsupportedOperationException.class, () -> snap.put("x", "y"),
"a handed-out snapshot cannot be mutated in place");
// Later binds must not leak into an earlier snapshot.
assertTrue(registry.bind(REVIEWER, "term_review"));
assertFalse(snap.containsKey("term_review"),
"a snapshot is a point-in-time copy, not a live view");
}
// ── concurrency (CB-548 invariants hold under contention) ─────────────────────────────────
@Test
void concurrentBindsNeverGiveASlotTwoTerminals() throws Exception {
int n = 16;
ExecutorService pool = Executors.newFixedThreadPool(n);
try {
CountDownLatch go = new CountDownLatch(1);
List<Future<Boolean>> results = new ArrayList<>();
for (int i = 0; i < n; i++) {
final String term = "term_" + i; // every thread races for the SAME slot
results.add(pool.submit(() -> {
go.await();
return registry.bind(DESIGNER, term);
}));
}
go.countDown();
int won = 0;
for (Future<Boolean> r : results) {
if (r.get()) {
won++;
}
}
assertEquals(1, won, "exactly one terminal may win the sole slot, got " + won);
assertEquals(1, registry.snapshot().size(),
"the slot hosts at most one terminal after the race");
} finally {
pool.shutdownNow();
}
}
@Test
void concurrentBindsNeverPutOneTerminalInTwoSlots() throws Exception {
int n = 16;
ExecutorService pool = Executors.newFixedThreadPool(n);
try {
CountDownLatch go = new CountDownLatch(1);
List<Future<String>> results = new ArrayList<>();
for (int i = 0; i < n; i++) {
final String slot = (i % 2 == 0) ? DESIGNER : REVIEWER; // all race for ONE terminal
results.add(pool.submit(() -> {
go.await();
return registry.bind(slot, "shared_term")
? registry.slotForTerminal("shared_term") : null;
}));
}
go.countDown();
// Rebinding the same terminal to the same slot is a harmless idempotent true, so count
// winners is not the assertion — agreement is: every thread that reported success must
// have seen the terminal in the SAME slot, never in two at once.
String bound = null;
boolean conflict = false;
for (Future<String> r : results) {
String s = r.get();
if (s != null) {
if (bound == null) {
bound = s;
} else if (!bound.equals(s)) {
conflict = true;
}
}
}
assertFalse(conflict, "a terminal was observed in two slots at once");
assertNotNull(bound, "at least one thread bound the terminal");
assertEquals(1, registry.snapshot().size(),
"the terminal occupies exactly one slot in the final snapshot");
assertEquals(bound, registry.slotForTerminal("shared_term"));
} finally {
pool.shutdownNow();
}
}
/** A handed-over slot map is snapshotted at construction, not offered as live state. */
@Test
void theSlotSnapshotIsFixedByConstruction() {
Map<String, BridgedConfig.Slot> mutable = architects();
MemberRegistry r = new MemberRegistry(fleetWith(mutable));
mutable.put("hijack", new BridgedConfig.Slot("gx10"));
assertFalse(r.isSlot("architect:hijack"), "a handed-over map is not offered as live state");
}
}
File diff suppressed because it is too large Load Diff
@@ -0,0 +1,401 @@
package dev.ltms.bridged.config;
import org.junit.jupiter.api.Test;
import org.junit.jupiter.api.io.TempDir;
import java.nio.file.Files;
import java.nio.file.Path;
import static org.junit.jupiter.api.Assertions.*;
/**
* CB-559: re-reading {@code bridged.yaml} under a running daemon.
*
* <p>The tests that matter here are the refusals. A reload that applies a good file is the easy
* half; the half that protects an operator is the one that keeps the running config when the new
* file is bad, and the one that refuses a change the running daemon cannot honour.
*/
class ConfigRefTest {
/** A minimal file that loads and passes every startup validator. */
private static String yaml(String extra) {
return """
bind:
host: 127.0.0.1
port: 8765
herdrSocket: ~/.config/herdr/herdr.sock
profiles:
sonnet:
baseUrl: http://gx00.gw:8000
model: sonnet
guard:
offSubscriptionHosts:
- gx00.gw
""" + extra;
}
private static ConfigRef refFor(Path f) {
return new ConfigRef(f, BridgedConfig.load(f));
}
@Test
void aHotChangeIsAppliedAndReadThroughGet(@TempDir Path dir) throws Exception {
Path f = dir.resolve("bridged.yaml");
Files.writeString(f, yaml("""
fleet:
tabLabel: "{role}: {profile} #{n}"
developers:
a:
profile: sonnet
"""));
ConfigRef ref = refFor(f);
assertEquals("{role}: {profile} #{n}", ref.get().fleet().tabLabel());
Files.writeString(f, yaml("""
fleet:
tabLabel: "[{profile}] {role}"
developers:
a:
profile: sonnet
"""));
ConfigRef.Outcome out = ref.reload();
assertTrue(out.applied());
assertTrue(out.deferred().isEmpty());
assertEquals("config reloaded", out.summary());
assertEquals("[{profile}] {role}", ref.get().fleet().tabLabel());
}
@Test
void aCharterChangeIsHotAndReachesTheLiveConfig(@TempDir Path dir) throws Exception {
Path f = dir.resolve("bridged.yaml");
Files.writeString(f, yaml("""
fleet:
charters:
architect: old charter
"""));
ConfigRef ref = refFor(f);
assertEquals("old charter", ref.get().fleet().charterFor(
dev.ltms.bridged.peer.MemberRole.ARCHITECT));
Files.writeString(f, yaml("""
fleet:
charters:
architect: new charter
"""));
ConfigRef.Outcome out = ref.reload();
assertTrue(out.applied());
assertTrue(out.deferred().isEmpty());
assertEquals("new charter", ref.get().fleet().charterFor(
dev.ltms.bridged.peer.MemberRole.ARCHITECT));
}
@Test
void invalidChartersRefuseReloadAndKeepTheRunningConfig(@TempDir Path dir) throws Exception {
Path f = dir.resolve("bridged.yaml");
Files.writeString(f, yaml("""
fleet:
charters:
architect: valid charter
"""));
ConfigRef ref = refFor(f);
BridgedConfig before = ref.get();
Files.writeString(f, yaml("""
fleet:
charters:
architect: " "
"""));
ConfigRef.Outcome blank = ref.reload();
assertFalse(blank.applied());
assertTrue(blank.error().contains("fleet.charters.architect is blank"));
assertSame(before, ref.get());
Files.writeString(f, yaml("""
fleet:
charters:
architetc: valid charter
"""));
ConfigRef.Outcome unknown = ref.reload();
assertFalse(unknown.applied());
assertTrue(unknown.error().contains("architetc"));
assertTrue(unknown.error().contains("architect"));
assertSame(before, ref.get());
}
/**
* The point of the whole class: a consumer holding the ref sees the new value without being
* rebuilt. A component that captured {@code get()} into a field would still show the old one.
*/
@Test
void aConsumerHoldingTheRefSeesTheNewValue(@TempDir Path dir) throws Exception {
Path f = dir.resolve("bridged.yaml");
Files.writeString(f, yaml("placement: weighted\n"));
ConfigRef ref = refFor(f);
java.util.function.Supplier<String> reader = () -> ref.get().placement();
assertEquals("weighted", reader.get());
Files.writeString(f, yaml("placement: fixed\n"));
assertTrue(ref.reload().applied());
assertEquals("fixed", reader.get());
}
@Test
void aChangedColdKeyRefusesTheWholeReload(@TempDir Path dir) throws Exception {
Path f = dir.resolve("bridged.yaml");
Files.writeString(f, yaml("placement: weighted\n"));
ConfigRef ref = refFor(f);
// Two changes in one file: a cold one (the port) and a hot one (placement).
Files.writeString(f, yaml("placement: fixed\n").replace("port: 8765", "port: 9999"));
ConfigRef.Outcome out = ref.reload();
assertFalse(out.applied());
assertEquals(java.util.List.of("bind"), out.coldKeys());
assertTrue(out.summary().contains("Restart bridged"), out.summary());
// The hot half must NOT have leaked in. A half-applied reload leaves the daemon matching no
// file on disk, which is worse for an operator than no reload at all.
assertEquals("weighted", ref.get().placement());
assertEquals(8765, ref.get().bind().port());
}
@Test
void aFileThatNoLongerParsesKeepsTheRunningConfig(@TempDir Path dir) throws Exception {
Path f = dir.resolve("bridged.yaml");
Files.writeString(f, yaml("placement: weighted\n"));
ConfigRef ref = refFor(f);
BridgedConfig before = ref.get();
Files.writeString(f, "profiles:\n sonnet:\n baseUrl: \"unclosed\n");
ConfigRef.Outcome out = ref.reload();
assertFalse(out.applied());
assertNotNull(out.error());
assertTrue(out.summary().startsWith("config reload refused"), out.summary());
assertSame(before, ref.get());
}
/** A file that would have refused to boot must not be able to slip in through a reload. */
@Test
void aFileThatFailsAValidatorKeepsTheRunningConfig(@TempDir Path dir) throws Exception {
Path f = dir.resolve("bridged.yaml");
Files.writeString(f, yaml("placement: weighted\n"));
ConfigRef ref = refFor(f);
BridgedConfig before = ref.get();
// A member slot naming a profile that does not exist — validateMembers refuses this at
// startup, so it must refuse it here too.
Files.writeString(f, yaml("""
fleet:
developers:
a:
profile: no-such-profile
"""));
ConfigRef.Outcome out = ref.reload();
assertFalse(out.applied());
assertNotNull(out.error());
assertSame(before, ref.get());
}
@Test
void aDeletedFileIsRefusedRatherThanCrashing(@TempDir Path dir) throws Exception {
Path f = dir.resolve("bridged.yaml");
Files.writeString(f, yaml(""));
ConfigRef ref = refFor(f);
BridgedConfig before = ref.get();
Files.delete(f);
ConfigRef.Outcome out = ref.reload();
assertFalse(out.applied());
assertNotNull(out.error());
assertSame(before, ref.get());
}
/** A deferred change applies to the snapshot but the operator is told it needs a restart. */
@Test
void aDeferredChangeIsAppliedAndReported(@TempDir Path dir) throws Exception {
Path f = dir.resolve("bridged.yaml");
Files.writeString(f, yaml("""
lifecycle:
drainTimeoutSeconds: 30
"""));
ConfigRef ref = refFor(f);
Files.writeString(f, yaml("""
lifecycle:
drainTimeoutSeconds: 60
"""));
ConfigRef.Outcome out = ref.reload();
assertTrue(out.applied());
assertEquals(java.util.List.of("lifecycle"), out.deferred());
assertTrue(out.summary().contains("needs a restart") || out.summary().contains("need a restart"),
out.summary());
assertEquals(60, ref.get().lifecycle().drainTimeoutSeconds());
}
/**
* Adding a profile is deferred, not hot: a new backend needs its own launcher, and launchers are
* built once at startup. The snapshot carries it so a restart picks it up.
*/
@Test
void addingAProfileIsReportedAsDeferred(@TempDir Path dir) throws Exception {
Path f = dir.resolve("bridged.yaml");
Files.writeString(f, yaml(""));
ConfigRef ref = refFor(f);
Files.writeString(f, """
bind:
host: 127.0.0.1
port: 8765
herdrSocket: ~/.config/herdr/herdr.sock
profiles:
sonnet:
baseUrl: http://gx00.gw:8000
model: sonnet
haiku:
baseUrl: http://gx00.gw:8000
model: haiku
guard:
offSubscriptionHosts:
- gx00.gw
""");
ConfigRef.Outcome out = ref.reload();
assertTrue(out.applied());
assertEquals(1, out.deferred().size());
assertTrue(out.deferred().getFirst().contains("haiku"), out.deferred().toString());
}
/**
* Changing an existing profile's weight or maxLoad IS hot: placement reads those live through
* the supplier on the composite, so the next spawn already sees them.
*/
@Test
void changingAProfilesWeightOrMaxLoadIsHot(@TempDir Path dir) throws Exception {
Path f = dir.resolve("bridged.yaml");
Files.writeString(f, yaml(""));
ConfigRef ref = refFor(f);
Files.writeString(f, """
bind:
host: 127.0.0.1
port: 8765
herdrSocket: ~/.config/herdr/herdr.sock
profiles:
sonnet:
baseUrl: http://gx00.gw:8000
model: sonnet
maxLoad: 7
weight: 3.0
guard:
offSubscriptionHosts:
- gx00.gw
""");
ConfigRef.Outcome out = ref.reload();
assertTrue(out.applied());
assertTrue(out.deferred().isEmpty(), out.deferred().toString());
assertEquals(7, ref.get().profiles().get("sonnet").maxLoad());
}
/**
* Changing an existing profile's MODEL is deferred, not hot — and saying so is the whole point.
* HerdrPeerLauncher takes Map.copyOf(profiles) at construction and resolves every spawn out of
* that copy, so a reloaded model never reaches a launch. Reporting it as applied would be the
* worst outcome a reload can produce: the operator has no reason to doubt a clean "reloaded".
*/
@Test
void changingAProfilesLaunchSettingsIsReportedAsDeferred(@TempDir Path dir) throws Exception {
Path f = dir.resolve("bridged.yaml");
Files.writeString(f, yaml(""));
ConfigRef ref = refFor(f);
Files.writeString(f, """
bind:
host: 127.0.0.1
port: 8765
herdrSocket: ~/.config/herdr/herdr.sock
profiles:
sonnet:
baseUrl: http://gx00.gw:8000
model: deepseek-v4-flash
guard:
offSubscriptionHosts:
- gx00.gw
""");
ConfigRef.Outcome out = ref.reload();
assertTrue(out.applied());
assertEquals(1, out.deferred().size(), out.deferred().toString());
assertTrue(out.deferred().getFirst().contains("sonnet"), out.deferred().toString());
assertTrue(out.deferred().getFirst().contains("launch settings"), out.deferred().toString());
// The snapshot still carries the new value — a restart is what makes it take effect.
assertEquals("deepseek-v4-flash", ref.get().profiles().get("sonnet").model());
}
/**
* CB-578 stage B: exhaustedPattern is compiled once into Bridged.main's pattern map at startup
* (see ExhaustedPatternLookup), so a reload never re-reads it — a changed pattern must be
* reported deferred exactly like model/baseUrl, not silently claimed as applied.
*/
@Test
void changingAProfilesExhaustedPatternIsReportedAsDeferred(@TempDir Path dir) throws Exception {
Path f = dir.resolve("bridged.yaml");
Files.writeString(f, """
bind:
host: 127.0.0.1
port: 8765
herdrSocket: ~/.config/herdr/herdr.sock
profiles:
sonnet:
baseUrl: http://gx00.gw:8000
model: sonnet
exhaustedPattern: "usage limit has been reached"
guard:
offSubscriptionHosts:
- gx00.gw
""");
ConfigRef ref = refFor(f);
Files.writeString(f, """
bind:
host: 127.0.0.1
port: 8765
herdrSocket: ~/.config/herdr/herdr.sock
profiles:
sonnet:
baseUrl: http://gx00.gw:8000
model: sonnet
exhaustedPattern: "rate limit exceeded"
guard:
offSubscriptionHosts:
- gx00.gw
""");
ConfigRef.Outcome out = ref.reload();
assertTrue(out.applied());
assertEquals(1, out.deferred().size(), out.deferred().toString());
assertTrue(out.deferred().getFirst().contains("sonnet"), out.deferred().toString());
assertTrue(out.deferred().getFirst().contains("launch settings"), out.deferred().toString());
// The snapshot still carries the new value — a restart is what makes it take effect.
assertEquals("rate limit exceeded", ref.get().profiles().get("sonnet").exhaustedPattern());
}
@Test
void aFixedRefHasNoFileAndRefusesToReload() {
BridgedConfig cfg = new BridgedConfig(null, null, null, null, null, null,
null, null, null, null, null, null, null, null, null).withDefaults();
ConfigRef ref = ConfigRef.fixed(cfg);
assertNull(ref.path());
assertSame(cfg, ref.get());
ConfigRef.Outcome out = ref.reload();
assertFalse(out.applied());
assertNotNull(out.error());
}
}
@@ -0,0 +1,162 @@
package dev.ltms.bridged.config;
import org.junit.jupiter.api.Test;
import org.junit.jupiter.api.io.TempDir;
import java.nio.file.Files;
import java.nio.file.Path;
import java.nio.file.attribute.FileTime;
import static org.junit.jupiter.api.Assertions.*;
/**
* CB-559: the mtime poller behind {@code configReload:}. The tests drive {@link
* ConfigWatcher#tick()} directly rather than the scheduler, so nothing here sleeps.
*/
class ConfigWatcherTest {
private static String yaml(String placement) {
return """
bind:
host: 127.0.0.1
port: 8765
herdrSocket: ~/.config/herdr/herdr.sock
profiles:
sonnet:
baseUrl: http://gx00.gw:8000
model: sonnet
guard:
offSubscriptionHosts:
- gx00.gw
""" + "placement: " + placement + "\n";
}
/** Write and stamp an mtime, so a test never depends on the filesystem's clock resolution. */
private static void write(Path f, String content, long millis) throws Exception {
Files.writeString(f, content);
Files.setLastModifiedTime(f, FileTime.fromMillis(millis));
}
@Test
void aChangedMtimeTriggersAReload(@TempDir Path dir) throws Exception {
Path f = dir.resolve("bridged.yaml");
write(f, yaml("weighted"), 1_000L);
ConfigRef ref = new ConfigRef(f, BridgedConfig.load(f));
ConfigWatcher watcher = new ConfigWatcher(ref, 10);
try {
write(f, yaml("fixed"), 2_000L);
watcher.tick();
assertEquals("fixed", ref.get().placement());
} finally {
watcher.stop();
}
}
@Test
void anUnchangedFileIsNotReloaded(@TempDir Path dir) throws Exception {
Path f = dir.resolve("bridged.yaml");
write(f, yaml("weighted"), 1_000L);
ConfigRef ref = new ConfigRef(f, BridgedConfig.load(f));
BridgedConfig before = ref.get();
ConfigWatcher watcher = new ConfigWatcher(ref, 10);
try {
watcher.tick();
watcher.tick();
// Same instance, so no reload happened — a reload always swaps in a fresh object.
assertSame(before, ref.get());
} finally {
watcher.stop();
}
}
/**
* A refused reload must not be retried every tick. Without the stamp-before-reload order the
* same refusal would be logged forever, which buries the log an operator needs.
*/
@Test
void aRefusedReloadIsNotRetriedUntilTheFileChangesAgain(@TempDir Path dir) throws Exception {
Path f = dir.resolve("bridged.yaml");
write(f, yaml("weighted"), 1_000L);
ConfigRef ref = new ConfigRef(f, BridgedConfig.load(f));
BridgedConfig before = ref.get();
ConfigWatcher watcher = new ConfigWatcher(ref, 10);
try {
// A cold change: refused, and the running config stays.
write(f, yaml("fixed").replace("port: 8765", "port: 9999"), 2_000L);
watcher.tick();
assertSame(before, ref.get());
// The next tick sees the same mtime, so it does nothing at all.
watcher.tick();
assertSame(before, ref.get());
// A fresh save earns a fresh attempt — and this one is hot, so it applies.
write(f, yaml("fixed"), 3_000L);
watcher.tick();
assertEquals("fixed", ref.get().placement());
} finally {
watcher.stop();
}
}
/**
* Editors briefly unlink the file mid-save. A missing file is skipped, not an error and not a
* reload — the running config stays live, which is right either way.
*/
@Test
void aMissingFileIsSkippedAndTheNextTickTriesAgain(@TempDir Path dir) throws Exception {
Path f = dir.resolve("bridged.yaml");
write(f, yaml("weighted"), 1_000L);
ConfigRef ref = new ConfigRef(f, BridgedConfig.load(f));
BridgedConfig before = ref.get();
ConfigWatcher watcher = new ConfigWatcher(ref, 10);
try {
Files.delete(f);
assertDoesNotThrow(watcher::tick);
assertSame(before, ref.get());
write(f, yaml("fixed"), 2_000L);
watcher.tick();
assertEquals("fixed", ref.get().placement());
} finally {
watcher.stop();
}
}
/** A ref with no file behind it must not start a scheduler that could never do anything. */
@Test
void aFixedRefStartsNoWatch() {
BridgedConfig cfg = new BridgedConfig(null, null, null, null, null, null,
null, null, null, null, null, null, null, null, null).withDefaults();
ConfigWatcher watcher = new ConfigWatcher(ConfigRef.fixed(cfg), 10);
try {
assertDoesNotThrow(watcher::start);
assertDoesNotThrow(watcher::tick);
} finally {
watcher.stop();
}
}
@Test
void configReloadIsOffUnlessTheBlockSaysOtherwise(@TempDir Path dir) throws Exception {
Path f = dir.resolve("bridged.yaml");
write(f, yaml("weighted"), 1_000L);
assertNull(BridgedConfig.load(f).configReload());
write(f, yaml("weighted") + "configReload:\n enabled: true\n", 2_000L);
BridgedConfig.ConfigReload on = BridgedConfig.load(f).configReload();
assertTrue(on.isEnabled());
assertEquals(10, on.intervalSeconds(), "an absent interval defaults to 10s");
write(f, yaml("weighted") + "configReload:\n intervalSeconds: 30\n", 3_000L);
BridgedConfig.ConfigReload off = BridgedConfig.load(f).configReload();
assertFalse(off.isEnabled(), "a block that only sets the interval does not enable the watch");
assertEquals(30, off.intervalSeconds());
}
}
@@ -0,0 +1,170 @@
package dev.ltms.bridged.health;
import ch.qos.logback.classic.Logger;
import ch.qos.logback.classic.spi.ILoggingEvent;
import ch.qos.logback.core.read.ListAppender;
import dev.ltms.bridged.herdr.AgentControl;
import dev.ltms.bridged.herdr.FakeHerdr;
import dev.ltms.bridged.inject.Injector;
import dev.ltms.bridged.msg.InMemoryReplyInbox;
import dev.ltms.bridged.msg.MessageService;
import dev.ltms.bridged.msg.Rendezvous;
import dev.ltms.bridged.session.SessionManager;
import dev.ltms.bridged.member.ClaudeCodeLauncher;
import dev.ltms.bridged.herdr.WorkspaceControl;
import dev.ltms.bridged.guard.SubscriptionGuard;
import dev.ltms.bridged.config.BridgedConfig;
import org.junit.jupiter.api.Test;
import org.slf4j.LoggerFactory;
import java.util.Map;
import java.util.Set;
import java.util.concurrent.Executors;
import java.util.function.BiConsumer;
import static org.junit.jupiter.api.Assertions.assertEquals;
import static org.junit.jupiter.api.Assertions.assertTrue;
class FleetHealthMonitorTest {
@Test void oneTickUsesOneFleetListForAnyRosterSize() {
FakeHerdr herdr = new FakeHerdr();
AgentControl agents = new AgentControl(herdr);
BridgedConfig.Profile profile = new BridgedConfig.Profile("test", "http://test:1", null,
null, null, null, null, null, null, null, null, null);
ClaudeCodeLauncher launcher = new ClaudeCodeLauncher(agents, new WorkspaceControl(herdr),
new SubscriptionGuard(Set.of("test")), Map.of("test", profile), "test", _ -> "token");
SessionManager sessions = new SessionManager(launcher);
sessions.acquire("test", null, null, null);
sessions.acquire("test", null, null, null);
herdr.calls.clear();
MessageService messages = new MessageService(agents, new Injector(agents), new Rendezvous(), new InMemoryReplyInbox());
var scheduler = Executors.newSingleThreadScheduledExecutor();
FleetHealthMonitor monitor = new FleetHealthMonitor(agents, sessions::roster, messages, scheduler, () -> 1, 60,
(_, _) -> { });
monitor.tick();
monitor.stop();
assertEquals(1, herdr.calls.stream().filter(call -> call.method().equals("agent.list")).count());
}
@Test void failedTickDoesNotStopTheNextTick() {
FakeHerdr herdr = new FakeHerdr().healthy(false);
AgentControl agents = new AgentControl(herdr);
var scheduler = Executors.newSingleThreadScheduledExecutor();
FleetHealthMonitor monitor = new FleetHealthMonitor(agents, java.util.List::of,
new MessageService(agents, new Injector(agents), new Rendezvous(), new InMemoryReplyInbox()),
scheduler, () -> 1, 60, (_, _) -> { });
monitor.tick();
herdr.healthy(true);
monitor.tick();
monitor.stop();
assertEquals(2, herdr.calls.stream().filter(call -> call.method().equals("agent.list")).count());
}
@Test void faultTransitionLogsOnlyOnceUntilItChanges() {
Logger logger = (Logger) LoggerFactory.getLogger(FleetHealthMonitor.class);
ListAppender<ILoggingEvent> appender = new ListAppender<>();
appender.start();
logger.addAppender(appender);
try {
FakeHerdr herdr = new FakeHerdr();
AgentControl agents = new AgentControl(herdr);
var scheduler = Executors.newSingleThreadScheduledExecutor();
FleetHealthMonitor monitor = new FleetHealthMonitor(agents, java.util.List::of,
new MessageService(agents, new Injector(agents), new Rendezvous(), new InMemoryReplyInbox()),
scheduler, () -> 1, 60, (_, _) -> { });
monitor.reportTransition("term_a", HealthState.TURN_BOUNDARY_LOST);
monitor.reportTransition("term_a", HealthState.TURN_BOUNDARY_LOST);
monitor.stop();
assertEquals(1, appender.list.stream().filter(event -> event.getFormattedMessage()
.contains("member=term_a state=TURN_BOUNDARY_LOST")).count());
} finally {
logger.detachAppender(appender);
}
}
// --- CB-580: a member that reaches GONE/NEVER_READY must fail its waiting tickets
private static FleetHealthMonitor monitorWith(BiConsumer<String, String> failTarget) {
FakeHerdr herdr = new FakeHerdr();
AgentControl agents = new AgentControl(herdr);
var scheduler = Executors.newSingleThreadScheduledExecutor();
return new FleetHealthMonitor(agents, java.util.List::of,
new MessageService(agents, new Injector(agents), new Rendezvous(), new InMemoryReplyInbox()),
scheduler, () -> 1, 60, failTarget);
}
@Test void terminalTransitionFailsTheTargetOnce() {
RecordingFailTarget failTarget = new RecordingFailTarget();
FleetHealthMonitor monitor = monitorWith(failTarget);
monitor.reportTransition("term_a", HealthState.GONE);
monitor.stop();
assertEquals(1, failTarget.calls.size());
assertEquals("term_a", failTarget.calls.get(0).target());
assertTrue(failTarget.calls.get(0).reason().contains("GONE"));
}
@Test void neverReadyNamesItselfAsTheReason() {
RecordingFailTarget failTarget = new RecordingFailTarget();
FleetHealthMonitor monitor = monitorWith(failTarget);
monitor.reportTransition("term_a", HealthState.NEVER_READY);
monitor.stop();
assertEquals(1, failTarget.calls.size());
assertTrue(failTarget.calls.get(0).reason().contains("NEVER_READY"));
}
@Test void stayingInATerminalStateProducesOneFailureNotN() {
RecordingFailTarget failTarget = new RecordingFailTarget();
FleetHealthMonitor monitor = monitorWith(failTarget);
monitor.reportTransition("term_a", HealthState.GONE);
monitor.reportTransition("term_a", HealthState.GONE);
monitor.reportTransition("term_a", HealthState.GONE);
monitor.reportTransition("term_a", HealthState.GONE);
monitor.stop();
assertEquals(1, failTarget.calls.size());
}
@Test void aNonTerminalFaultStateDoesNotFailTheTarget() {
RecordingFailTarget failTarget = new RecordingFailTarget();
FleetHealthMonitor monitor = monitorWith(failTarget);
monitor.reportTransition("term_a", HealthState.TURN_BOUNDARY_LOST);
monitor.stop();
assertEquals(0, failTarget.calls.size());
}
@Test void failTargetRetryIsBounded() {
AlwaysThrowingFailTarget failTarget = new AlwaysThrowingFailTarget();
FleetHealthMonitor monitor = monitorWith(failTarget);
monitor.reportTransition("term_a", HealthState.GONE);
monitor.stop();
assertEquals(FleetHealthMonitor.MAX_FAIL_TARGET_ATTEMPTS, failTarget.calls);
}
@Test void exhaustedRetryStillDoesNotRefireOnAnUnchangedTick() {
AlwaysThrowingFailTarget failTarget = new AlwaysThrowingFailTarget();
FleetHealthMonitor monitor = monitorWith(failTarget);
monitor.reportTransition("term_a", HealthState.GONE);
int afterFirstTransition = failTarget.calls;
monitor.reportTransition("term_a", HealthState.GONE);
monitor.stop();
assertEquals(afterFirstTransition, failTarget.calls);
}
private record RecordedCall(String target, String reason) { }
private static final class RecordingFailTarget implements BiConsumer<String, String> {
final java.util.List<RecordedCall> calls = new java.util.ArrayList<>();
@Override public void accept(String target, String reason) {
calls.add(new RecordedCall(target, reason));
}
}
private static final class AlwaysThrowingFailTarget implements BiConsumer<String, String> {
int calls = 0;
@Override public void accept(String target, String reason) {
calls++;
throw new RuntimeException("boom");
}
}
}
@@ -0,0 +1,56 @@
package dev.ltms.bridged.health;
import dev.ltms.bridged.herdr.AgentStatus;
import dev.ltms.bridged.session.MemberSession;
import dev.ltms.bridged.msg.Rendezvous;
import org.junit.jupiter.api.Test;
import static org.junit.jupiter.api.Assertions.assertEquals;
class FleetHealthTest {
@Test void muteCountsOnlyCompletionFallbacks() {
MuteCounter mute = new MuteCounter();
mute.observe("target", "terra", Rendezvous.Kind.REPLY);
mute.observe("target", "terra", Rendezvous.Kind.COMPLETION);
assertEquals(1, mute.forTarget("target"));
assertEquals(1, mute.forProfile("terra"));
}
@Test void turnBoundaryNeedsTwoSnapshots() {
HealthSnapshot s = snapshot(MemberSession.State.BUSY, AgentStatus.DONE, true);
HealthDecision first = FleetHealth.decide(s, HealthPrior.NONE, 1);
assertEquals(HealthState.WORKING, first.state());
assertEquals(new HealthPrior(true), first.prior());
assertEquals(HealthState.TURN_BOUNDARY_LOST, FleetHealth.decide(s, first.prior(), 2).state());
}
@Test void unknownLiveStatusWithAcceptedDeliveryIsNotIdle() {
assertEquals(HealthState.WORKING, FleetHealth.decide(
snapshot(MemberSession.State.BUSY, AgentStatus.UNKNOWN, true), HealthPrior.NONE, 1).state());
}
@Test void acceptedDeliveryNeverReportsIdle() {
for (MemberSession.State session : MemberSession.State.values()) {
for (AgentStatus live : AgentStatus.values()) {
HealthSnapshot s = snapshot(session, live, true);
assertEquals(false, FleetHealth.decide(s, HealthPrior.NONE, 1).state() == HealthState.IDLE,
() -> "accepted delivery returned IDLE for " + session + "/" + live);
}
}
}
@Test void blockedDoesNotGuessPromptKind() {
assertEquals(HealthState.BLOCKED_AMBIGUOUS, FleetHealth.decide(
snapshot(MemberSession.State.BUSY, AgentStatus.BLOCKED, true), HealthPrior.NONE, 1).state());
}
@Test void controlLinkOutranksMemberFault() {
HealthSnapshot s = new HealthSnapshot(MemberSession.State.BUSY, AgentStatus.DONE, true, false,
false, true, true, true, false, true, true, true);
assertEquals(HealthState.CONTROL_LINK_DOWN, FleetHealth.decide(s, HealthPrior.NONE, 1).state());
}
private static HealthSnapshot snapshot(MemberSession.State state, AgentStatus live, boolean accepted) {
return new HealthSnapshot(state, live, accepted, false, false, true, false, false,
false, false, false, false);
}
}
@@ -0,0 +1,14 @@
package dev.ltms.bridged.health;
import org.junit.jupiter.api.Test;
import java.util.List;
import static org.junit.jupiter.api.Assertions.assertEquals;
class PaneBudgetTest {
@Test void fixedLimitsIgnoreWeakerConfig() {
PaneBudget budget = new PaneBudget();
List<String> targets = List.of("a", "b", "c");
assertEquals(List.of("a", "b"), budget.choose(targets, 0, 0));
assertEquals(List.of("c"), budget.choose(targets, 1, 0));
}
}
@@ -11,10 +11,11 @@ import static org.junit.jupiter.api.Assertions.*;
import static org.junit.jupiter.api.Assumptions.assumeTrue;
/**
* Contract test for the {@code agent.*} south side against a REAL herdr, locking in
* the CB-102 spike findings. It spawns a HARMLESS probe command (never {@code claude},
* so no subscription/token involvement), proves the {@code env} map reaches the process
* environment, exercises status/read, and always tears the pane down.
* Contract test for the worker env seam against a REAL herdr. Under protocol 19 (CB-521) the
* env map is injected at PANE CREATION ({@code tab.create}), not {@code agent.start} — and
* {@code agent.start} now only launches supported agent kinds, so this probes the seed pane's
* SHELL directly (never {@code claude}, so no subscription/token involvement) and always tears
* the throwaway space down.
*
* <p>Tagged {@code contract}; run with {@code mvn test -Pcontract}.
*/
@@ -26,38 +27,30 @@ class AgentControlContractTest {
}
@Test
void startInjectsEnvThenReadAndClose() throws Exception {
void tabCreateInjectsEnvIntoTheSeedShell() throws Exception {
assumeTrue(!noSocket(), "no herdr socket — skipping");
try (UnixSocketHerdrClient herdr = UnixSocketHerdrClient.connect()) {
AgentControl agents = new AgentControl(herdr);
Agent probe = agents.start(
"__contract__",
List.of("bash", "-c", "printf 'PROBE_BASE=[%s]\\n' \"$ANTHROPIC_BASE_URL\"; sleep 20"),
WorkspaceControl spaces = new WorkspaceControl(herdr);
Workspace space = spaces.ensureWorkspace("__bridged_env_contract__");
Tab.Created tab = spaces.createTab(space.workspaceId(), null,
Map.of("ANTHROPIC_BASE_URL", "http://gx00.gw:8000"));
assertNotNull(probe.terminalId());
assertNotNull(probe.paneId());
try {
// Give the shell a moment to print, then confirm env reached the process.
assertNotNull(tab.rootPaneId(), "tab.create must return the seed pane");
Thread.sleep(1000); // let the seed shell reach its prompt
herdr.call("pane.send_input", Map.of(
"pane_id", tab.rootPaneId(),
"text", "printf 'PROBE_BASE=[%s]\\n' \"$ANTHROPIC_BASE_URL\"",
"keys", List.of("enter")));
Thread.sleep(800);
String visible = agents.read(probe.terminalId(), "visible");
String visible = herdr.call("pane.read",
Map.of("pane_id", tab.rootPaneId(), "source", "visible"))
.path("read").path("text").asText("");
assertTrue(visible.contains("PROBE_BASE=[http://gx00.gw:8000]"),
"env map must reach the process; saw: " + visible);
// Status is queryable; the probe appears in the agent list.
assertNotNull(agents.status(probe.terminalId()));
assertTrue(agents.list().stream()
.anyMatch(a -> probe.terminalId().equals(a.terminalId())),
"spawned probe should appear in agent.list");
"env map must reach the seed shell; saw: " + visible);
} finally {
agents.close(probe.paneId());
spaces.closeTab(tab.tab().tabId());
herdr.call("workspace.close", Map.of("workspace_id", space.workspaceId()));
}
// After close the pane is gone.
assertFalse(agents.list().stream()
.anyMatch(a -> probe.terminalId().equals(a.terminalId())),
"closed probe should no longer be listed");
}
}
}
@@ -7,34 +7,68 @@ import java.util.Map;
import static org.junit.jupiter.api.Assertions.assertEquals;
/** Unit-level behaviour of {@link AgentControl} over a fake herdr. */
/** Unit-level behaviour of {@link AgentControl} over a fake herdr (protocol 19). */
class AgentControlTest {
/** The {@code text} of every agent.send, in call order. */
/** The {@code text} of every agent.prompt, in call order. */
@SuppressWarnings("unchecked")
private static List<String> sendTexts(FakeHerdr herdr) {
private static List<String> promptTexts(FakeHerdr herdr) {
return herdr.calls.stream()
.filter(c -> c.method().equals("agent.send"))
.filter(c -> c.method().equals("agent.prompt"))
.map(c -> ((Map<String, Object>) c.params()).get("text").toString())
.toList();
}
@Test
void sendDeliversThePayloadThenAStandaloneSubmitKey() {
void sendDeliversThePayloadAsOnePromptThatSubmitsItself() {
FakeHerdr herdr = new FakeHerdr();
new AgentControl(herdr).send("term_x", "do the thing");
// The Enter must be its own event — appended to the paste it would be swallowed as text.
assertEquals(List.of("do the thing", "\r"), sendTexts(herdr),
"payload paste first, then a separate carriage-return keystroke to submit it");
// agent.prompt pastes AND submits in one call — no separate Enter event to assert.
assertEquals(List.of("do the thing"), promptTexts(herdr),
"exactly one agent.prompt carrying the payload");
}
@Test
void sendPreservesEmbeddedNewlinesAndSubmitsOnlyOnce() {
void sendPreservesEmbeddedNewlinesVerbatim() {
FakeHerdr herdr = new FakeHerdr();
new AgentControl(herdr).send("term_x", "line1\nline2");
assertEquals(List.of("line1\nline2", "\r"), sendTexts(herdr),
"multiline content is delivered verbatim; a single trailing Enter submits it");
assertEquals(List.of("line1\nline2"), promptTexts(herdr),
"multiline content is delivered verbatim in the single prompt");
}
@Test
void submitNudgesWithAStandaloneEnterKeystroke() {
FakeHerdr herdr = new FakeHerdr();
new AgentControl(herdr).submit("term_x");
FakeHerdr.Call keys = herdr.lastCall("agent.send_keys");
assertEquals(Map.of("target", "term_x", "keys", List.of("enter")), keys.params(),
"the raced-Enter nudge is a raw send_keys, not a second prompt");
}
@Test
@SuppressWarnings("unchecked")
void aTerminalIdTargetIsTranslatedToItsPaneId() {
// Protocol 19 rejects terminal_id as an agent.* target; the fake's agent.list maps
// term_a to pane w2:p7, and the control layer must address herdr by that pane.
FakeHerdr herdr = new FakeHerdr();
new AgentControl(herdr).send("term_a", "hello");
Map<String, Object> prompt = (Map<String, Object>) herdr.lastCall("agent.prompt").params();
assertEquals("w2:p7", prompt.get("target"), "terminal target resolved to the agent's pane id");
}
@Test
@SuppressWarnings("unchecked")
void theTerminalToPaneMappingIsCachedAcrossCalls() {
FakeHerdr herdr = new FakeHerdr();
AgentControl agents = new AgentControl(herdr);
agents.send("term_a", "one");
agents.send("term_a", "two");
long lists = herdr.calls.stream().filter(c -> c.method().equals("agent.list")).count();
assertEquals(1, lists, "one agent.list resolution serves every later call to the same terminal");
}
}
@@ -4,12 +4,14 @@ import com.fasterxml.jackson.databind.JsonNode;
import com.fasterxml.jackson.databind.ObjectMapper;
import java.util.ArrayList;
import java.util.LinkedHashMap;
import java.util.List;
import java.util.Map;
/**
* Recording fake {@link HerdrClient} for unit/acceptance tests. Returns canned frames
* captured from the real herdr 0.7.0 daemon and records every call so tests can assert
* both behaviour and that guard-blocked paths never reached herdr.
* matching the real herdr 0.8.0 daemon (protocol 19) and records every call so tests can
* assert both behaviour and that guard-blocked paths never reached herdr.
*/
public final class FakeHerdr implements HerdrClient {
@@ -24,12 +26,18 @@ public final class FakeHerdr implements HerdrClient {
private boolean healthy = true;
private final List<String> extraWorkspaces = new ArrayList<>();
private final List<String> extraAgents = new ArrayList<>();
/** workspaceId → extra tabs that {@code tab.list} reports for it (CB-558 lead scans). */
private final Map<String, List<String>> extraTabs = new LinkedHashMap<>();
private int agentNameTakenFor = 0;
private int agentPaneBusyFor = 0;
private int workerTabPaneCount = 1;
private String paneCloseErrorCode = null;
private String agentSendErrorCode = null;
private volatile String agentStatus = "idle"; // steady-state agent.get status
private volatile String readText = "worker transcript tail"; // canned agent.read output
private int pinnedStarts = 0; // how many upcoming agent.start calls report a fixed pane
private String pinnedStartTerminal;
private String pinnedStartPane;
public FakeHerdr healthy(boolean h) {
this.healthy = h;
@@ -42,6 +50,12 @@ public final class FakeHerdr implements HerdrClient {
return this;
}
/** Reject the first {@code n} {@code agent.start} calls with {@code agent_pane_busy}. */
public FakeHerdr agentPaneBusyTimes(int n) {
this.agentPaneBusyFor = n;
return this;
}
/** Make the worker tab (w9:t2) report this many panes in {@code tab.list} (default 1). */
public FakeHerdr withWorkerTabPaneCount(int n) {
this.workerTabPaneCount = n;
@@ -66,12 +80,25 @@ public final class FakeHerdr implements HerdrClient {
return this;
}
/** Make {@code agent.send} fail with this herdr error code. */
/** Make delivery ({@code agent.prompt} / {@code agent.send_keys}) fail with this error code. */
public FakeHerdr agentSendFailsWith(String code) {
this.agentSendErrorCode = code;
return this;
}
/**
* Force the next {@code n} {@code agent.start} calls to report this terminal/pane coordinate,
* instead of the fake's usual incrementing {@code term_new_n}/{@code w9:pRoot_n}. Lets a test make
* two spawns report the <em>same</em> herdr pane, to prove the host-unique id (CB-519) never
* collides on that coordinate.
*/
public FakeHerdr pinNextStarts(int n, String terminalId, String paneId) {
this.pinnedStarts = n;
this.pinnedStartTerminal = terminalId;
this.pinnedStartPane = paneId;
return this;
}
/**
* Seed a named agent into {@code agent.list} (e.g. an orphaned worker for CB-117 reaper tests).
@@ -86,6 +113,18 @@ public final class FakeHerdr implements HerdrClient {
return this;
}
/**
* Seed a labelled tab into {@code tab.list} for one workspace (e.g. an existing {@code lead: x}
* tab). Pair it with {@link #withAgent} on the same {@code tabId} to make the lead <em>live</em>;
* seeding the tab alone models the stale-label case.
*/
public FakeHerdr withTab(String workspaceId, String tabId, String label) {
extraTabs.computeIfAbsent(workspaceId, _ -> new ArrayList<>())
.add(("{\"tab_id\":\"%s\",\"workspace_id\":\"%s\",\"label\":\"%s\",\"pane_count\":1}")
.formatted(tabId, workspaceId, label));
return this;
}
/** Seed an additional workspace into {@code workspace.list} (e.g. a pre-existing worker space). */
public FakeHerdr withWorkspace(String id, String label) {
extraWorkspaces.add(("{\"workspace_id\":\"%s\",\"label\":\"%s\",\"focused\":false,"
@@ -110,7 +149,7 @@ public final class FakeHerdr implements HerdrClient {
try {
return switch (method) {
case "ping" -> mapper.readTree(
"{\"type\":\"pong\",\"version\":\"0.7.0\",\"protocol\":14}");
"{\"type\":\"pong\",\"version\":\"0.8.0\",\"protocol\":19}");
case "workspace.list" -> mapper.readTree(("""
{"type":"workspace_list","workspaces":[
{"workspace_id":"w1","label":"dev-mgnl","focused":true,"pane_count":7,"agent_status":"unknown"},
@@ -122,9 +161,19 @@ public final class FakeHerdr implements HerdrClient {
"agent_session":{"kind":"id","value":"sess-1111"},
"workspace_id":"w2","tab_id":"w2:t7","pane_id":"w2:p7"}%s]}""")
.formatted(extraAgents.isEmpty() ? "" : "," + String.join(",", extraAgents)));
case "agent.send" -> {
case "agent.prompt" -> {
if (agentSendErrorCode != null) {
throw new HerdrException("herdr error [" + agentSendErrorCode + "]: agent.send failed",
throw new HerdrException("herdr error [" + agentSendErrorCode + "]: agent.prompt failed",
agentSendErrorCode, null);
}
yield mapper.readTree(("""
{"type":"agent_prompted","agent":{"terminal_id":"term_a","agent":"claude",
"agent_status":"%s","workspace_id":"w2","tab_id":"w2:t7","pane_id":"w2:p7"}}""")
.formatted(agentStatus));
}
case "agent.send_keys" -> {
if (agentSendErrorCode != null) {
throw new HerdrException("herdr error [" + agentSendErrorCode + "]: agent.send_keys failed",
agentSendErrorCode, null);
}
yield mapper.readTree("{\"type\":\"ok\"}");
@@ -136,37 +185,78 @@ public final class FakeHerdr implements HerdrClient {
case "agent.read" -> mapper.readTree(mapper.writeValueAsString(
java.util.Map.of("type", "agent_read", "read", java.util.Map.of("text", readText))));
case "agent.start" -> {
// Protocol 19: kind and pane_id are required — reject like the real daemon.
java.util.Map<?, ?> p = params instanceof java.util.Map<?, ?> m ? m : java.util.Map.of();
for (String required : new String[]{"kind", "pane_id"}) {
if (p.get(required) == null) {
throw new HerdrException(
"herdr error [invalid_request]: invalid request: missing field `"
+ required + "`", "invalid_request", null);
}
}
long starts = calls.stream().filter(c -> c.method().equals("agent.start")).count();
if (starts <= agentNameTakenFor) {
if (starts <= agentPaneBusyFor) {
throw new HerdrException(
"herdr error [agent_pane_busy]: agent target pane is not an available shell",
"agent_pane_busy", null);
}
long busyAdjusted = starts - agentPaneBusyFor;
if (busyAdjusted <= agentNameTakenFor) {
throw new HerdrException(
"herdr error [agent_name_taken]: agent name already used",
"agent_name_taken", null);
}
long n = starts - agentNameTakenFor;
long n = busyAdjusted - agentNameTakenFor;
boolean pinned = pinnedStarts > 0;
if (pinned) {
pinnedStarts--;
}
// Protocol 19: the agent starts INTO the requested pane, so its pane_id normally
// echoes the param. A pin overrides both coordinates, which is the only way to
// make two spawns report one pane — what CB-519's collision test needs.
String terminal = pinned ? pinnedStartTerminal : ("term_new_" + n);
Object pane = pinned ? pinnedStartPane : p.get("pane_id");
yield mapper.readTree(("""
{"type":"agent_started","agent":{
"terminal_id":"term_new_%d","name":"claude","agent_status":"unknown",
"workspace_id":"w9","tab_id":"w9:t2","pane_id":"w9:pW_%d"}}""")
.formatted(n, n));
"terminal_id":"%s","name":"claude","agent_status":"unknown",
"workspace_id":"w9","tab_id":"w9:t2","pane_id":"%s"}}""")
.formatted(terminal, pane));
}
case "pane.split" -> mapper.readTree("""
{"type":"pane_info","pane":{"pane_id":"w1:pSplit","workspace_id":"w1",
"tab_id":"w1:t1"}}""");
case "workspace.create" -> mapper.readTree("""
{"type":"workspace_created",
"workspace":{"workspace_id":"w9","label":"bridged-workers","focused":false,
"pane_count":1,"tab_count":1,"active_tab_id":"w9:t1","agent_status":"unknown"},
"tab":{"tab_id":"w9:t1","workspace_id":"w9","label":"1","pane_count":1},
"root_pane":{"pane_id":"w9:p1","workspace_id":"w9","tab_id":"w9:t1"}}""");
case "tab.create" -> mapper.readTree("""
case "tab.create" -> {
// Each tab gets its own seed pane — under protocol 19 that pane becomes the
// worker pane, so distinct spawns must yield distinct pane ids.
long tabs = calls.stream().filter(c -> c.method().equals("tab.create")).count();
yield mapper.readTree(("""
{"type":"tab_created",
"tab":{"tab_id":"w9:t2","workspace_id":"w9","label":"2","pane_count":1},
"root_pane":{"pane_id":"w9:pRoot","workspace_id":"w9","tab_id":"w9:t2"}}""");
"root_pane":{"pane_id":"w9:pRoot_%d","workspace_id":"w9","tab_id":"w9:t2"}}""")
.formatted(tabs));
}
case "tab.rename" -> mapper.readTree("""
{"type":"tab_info","tab":{"tab_id":"w9:t2","workspace_id":"w9",
"label":"worker: ltms-local","pane_count":1}}""");
case "tab.list" -> mapper.readTree(("""
case "tab.list" -> {
// The base pair is returned for every workspace, exactly as before. Seeded tabs
// are appended only for the workspace they were registered against, so a test
// that seeds none sees the historical response byte for byte.
Object wsId = params instanceof java.util.Map<?, ?> m ? m.get("workspace_id") : null;
List<String> seeded = extraTabs.getOrDefault(String.valueOf(wsId), List.of());
yield mapper.readTree(("""
{"type":"tab_list","tabs":[
{"tab_id":"w9:t1","workspace_id":"w9","label":"1","pane_count":1},
{"tab_id":"w9:t2","workspace_id":"w9","label":"worker: ltms-local","pane_count":%d}]}""")
.formatted(workerTabPaneCount));
{"tab_id":"w9:t2","workspace_id":"w9","label":"worker: ltms-local","pane_count":%d}%s]}""")
.formatted(workerTabPaneCount,
seeded.isEmpty() ? "" : "," + String.join(",", seeded)));
}
case "tab.close" -> mapper.readTree("{\"type\":\"ok\"}");
case "pane.get" -> mapper.readTree("""
{"type":"pane_info","pane":{"pane_id":"w9:pW","workspace_id":"w9",
@@ -0,0 +1,326 @@
package dev.ltms.bridged.herdr;
import com.fasterxml.jackson.databind.JsonNode;
import com.fasterxml.jackson.databind.ObjectMapper;
import org.junit.jupiter.api.Test;
import java.util.ArrayList;
import java.util.LinkedHashMap;
import java.util.List;
import java.util.Map;
import java.util.Set;
import java.util.concurrent.TimeUnit;
import java.util.concurrent.atomic.AtomicLong;
import static org.junit.jupiter.api.Assertions.*;
/**
* CB-531/CB-579. A lead is never spawned, so the daemon has to <em>find</em> it: these assert that
* an operator-labelled tab matching a configured {@code tab:} is what makes a pane a lead, and —
* just as importantly — what does not, and that a stale entry does not linger forever.
*/
class LeadTabScannerTest {
private static final ObjectMapper MAPPER = new ObjectMapper();
private static final long TTL = TimeUnit.SECONDS.toNanos(10);
/**
* A herdr whose workspace/tab/pane topology is declared per test. Counts calls so the caching
* contract can be asserted, and can be made to fail on demand.
*/
private static final class TopologyHerdr implements HerdrClient {
/** workspace_id → label. */
final Map<String, String> workspaces = new LinkedHashMap<>();
/** tab_id → [workspace_id, label]. */
final Map<String, String[]> tabs = new LinkedHashMap<>();
/** pane_id → [tab_id, terminal_id]. */
final Map<String, String[]> panes = new LinkedHashMap<>();
int calls;
boolean failing;
TopologyHerdr workspace(String id, String label) {
workspaces.put(id, label);
return this;
}
TopologyHerdr tab(String tabId, String workspaceId, String label) {
tabs.put(tabId, new String[]{workspaceId, label});
return this;
}
TopologyHerdr pane(String paneId, String tabId, String terminalId) {
panes.put(paneId, new String[]{tabId, terminalId});
return this;
}
@Override
public JsonNode call(String method, Object params) {
calls++;
if (failing) {
throw new HerdrException("socket closed");
}
List<String> items = new ArrayList<>();
switch (method) {
case "workspace.list" -> {
workspaces.forEach((id, label) -> items.add(
"{\"workspace_id\":\"%s\",\"label\":\"%s\"}".formatted(id, label)));
return read("{\"workspaces\":[%s]}".formatted(String.join(",", items)));
}
case "tab.list" -> {
String ws = String.valueOf(((Map<?, ?>) params).get("workspace_id"));
tabs.forEach((id, t) -> {
if (ws.equals(t[0])) {
items.add(("{\"tab_id\":\"%s\",\"workspace_id\":\"%s\",\"label\":%s,"
+ "\"pane_count\":1}").formatted(id, t[0],
t[1] == null ? "null" : "\"" + t[1] + "\""));
}
});
return read("{\"tabs\":[%s]}".formatted(String.join(",", items)));
}
case "pane.list" -> {
panes.forEach((id, p) -> items.add(
"{\"pane_id\":\"%s\",\"tab_id\":\"%s\",\"terminal_id\":\"%s\"}"
.formatted(id, p[0], p[1])));
return read("{\"panes\":[%s]}".formatted(String.join(",", items)));
}
default -> throw new AssertionError("unexpected herdr call: " + method);
}
}
private static JsonNode read(String json) {
try {
return MAPPER.readTree(json);
} catch (Exception e) {
throw new AssertionError(e);
}
}
@Override
public void close() {
}
}
/**
* The usual shape: one user space with lead tabs, one worker space bridged owns.
*
* <p>Not closed: {@code close()} is a no-op on this fake, and every test needs the handle after
* the scanner is built (to mutate the topology or read {@code calls}).
*/
@SuppressWarnings("resource")
private TopologyHerdr twoLeads() {
return new TopologyHerdr()
.workspace("w1", "main")
.workspace("w9", "bridged-workers")
.tab("w1:t1", "w1", "lead: opus-5.0")
.tab("w1:t2", "w1", "lead: gpt-sol-5.6")
.tab("w1:t3", "w1", "notes")
.tab("w9:t1", "w9", "worker: gx10 #1")
.pane("w1:p1", "w1:t1", "term_opus")
.pane("w1:p2", "w1:t2", "term_gpt")
.pane("w1:p3", "w1:t3", "term_notes")
.pane("w9:p1", "w9:t1", "term_worker");
}
/** The {@code tab:} → name map {@code twoLeads()}'s two lead tabs are configured under. */
private static Map<String, String> twoLeadsConfigured() {
return Map.of("lead: opus-5.0", "opus-5.0", "lead: gpt-sol-5.6", "gpt-sol-5.6");
}
private LeadTabScanner scanner(TopologyHerdr herdr, Map<String, String> tabToName,
AtomicLong clock) {
return new LeadTabScanner(herdr, tabToName, Set.of("bridged-workers"), TTL, clock::get);
}
@Test
void everyConfiguredTabBecomesALeadNamedByItsEntry() {
Map<String, String> leads = scanner(twoLeads(), twoLeadsConfigured(), new AtomicLong()).get();
assertEquals(Map.of("term_opus", "opus-5.0", "term_gpt", "gpt-sol-5.6"), leads,
"two leads discovered by their configured tab — no terminal_id was ever configured");
}
@Test
void anUnconfiguredTabContributesNothing() {
assertFalse(scanner(twoLeads(), twoLeadsConfigured(), new AtomicLong())
.get().containsKey("term_notes"));
}
/**
* CB-579: matching is exact against the configured map now, not a shared prefix — two leads with
* completely different labels are both discovered by one scanner, no convention required.
*/
@Test
void twoLeadsWithCompletelyDifferentLabelsAreBothDiscovered() {
TopologyHerdr herdr = new TopologyHerdr()
.workspace("w1", "main")
.tab("w1:t1", "w1", "orchestrator: opus")
.tab("w1:t2", "w1", "captain: sol")
.pane("w1:p1", "w1:t1", "term_opus")
.pane("w1:p2", "w1:t2", "term_sol");
Map<String, String> tabToName = Map.of("orchestrator: opus", "opus", "captain: sol", "sol");
Map<String, String> leads = scanner(herdr, tabToName, new AtomicLong()).get();
assertEquals(Map.of("term_opus", "opus", "term_sol", "sol"), leads,
"no shared prefix needed — each lead is matched by its own configured tab");
}
/**
* The guard that matters: bridged labels its own worker tabs, so if a worker space were scanned
* a naming accident would promote the fleet. The exclusion is by workspace, not by hoping the
* worker template never collides.
*/
@Test
void aTabInAWorkerSpaceIsNeverALeadEvenWhenItsLabelMatches() {
TopologyHerdr herdr = twoLeads().tab("w9:t2", "w9", "lead: impostor")
.pane("w9:p2", "w9:t2", "term_impostor");
Map<String, String> tabToName = new LinkedHashMap<>(twoLeadsConfigured());
tabToName.put("lead: impostor", "impostor");
assertFalse(scanner(herdr, tabToName, new AtomicLong()).get().containsKey("term_impostor"));
}
@Test
void aLabelWithNoConfiguredEntryIsIgnored() {
TopologyHerdr herdr = new TopologyHerdr().workspace("w1", "main")
.tab("w1:t1", "w1", "lead: nobody-configured").pane("w1:p1", "w1:t1", "term_a");
assertEquals(Map.of(), scanner(herdr, twoLeadsConfigured(), new AtomicLong()).get(),
"a label that names no configured lead resolves nobody");
}
@Test
void matchingIsCaseInsensitiveAndToleratesSurroundingWhitespace() {
TopologyHerdr herdr = new TopologyHerdr().workspace("w1", "main")
.tab("w1:t1", "w1", " LEAD: Opus-5.0 ").pane("w1:p1", "w1:t1", "term_a");
assertEquals(Map.of("term_a", "opus-5.0"),
scanner(herdr, Map.of("lead: Opus-5.0", "opus-5.0"), new AtomicLong()).get());
}
@Test
void everyPaneInALeadTabResolvesAsThatLead() {
// A human may split their own lead tab. Both panes are theirs, so both are that lead —
// nothing bridged placed can land here (see the worker-space test above).
TopologyHerdr herdr = twoLeads().pane("w1:p1b", "w1:t1", "term_opus_split");
assertEquals("opus-5.0",
scanner(herdr, twoLeadsConfigured(), new AtomicLong()).get().get("term_opus_split"));
}
/**
* CB-579 acceptance (6): this is the bug the ticket closes. A stale pin used to be merged back
* over every scan and never expire; now a scan is the whole answer, so a lead whose tab is gone
* drops out on the very next scan.
*/
@Test
void aTabNoLongerPresentDropsTheLeadOnTheNextScan() {
TopologyHerdr herdr = twoLeads();
AtomicLong clock = new AtomicLong();
LeadTabScanner s = scanner(herdr, twoLeadsConfigured(), clock);
assertTrue(s.get().containsKey("term_opus"));
// The session behind term_opus restarted — herdr no longer reports that tab or pane at all.
herdr.tabs.remove("w1:t1");
herdr.panes.remove("w1:p1");
clock.addAndGet(TTL);
assertFalse(s.get().containsKey("term_opus"),
"a stale entry must expire once the tab it named is gone, not be merged back forever");
}
/**
* CB-579 acceptance (5): the whole point of matching by tab instead of {@code terminal_id} — a
* restart changes the terminal, not the tab, so the lead resolves under the same name with no
* config edit.
*/
@Test
void aLeadRestartingInTheSameTabResolvesUnderTheSameName() {
TopologyHerdr herdr = twoLeads();
AtomicLong clock = new AtomicLong();
LeadTabScanner s = scanner(herdr, twoLeadsConfigured(), clock);
assertEquals("opus-5.0", s.get().get("term_opus"));
// The session restarts: herdr assigns the pane a new terminal_id, same tab (w1:t1).
herdr.panes.remove("w1:p1");
herdr.pane("w1:p1", "w1:t1", "term_opus_v2");
clock.addAndGet(TTL);
Map<String, String> leads = s.get();
assertEquals("opus-5.0", leads.get("term_opus_v2"), "the new terminal resolves immediately");
assertFalse(leads.containsKey("term_opus"), "the old terminal_id is simply gone, not carried");
}
// ── caching ─────────────────────────────────────────────────────────────────────────────────
@Test
void aSecondLookupWithinTheTtlDoesNotTouchHerdr() {
TopologyHerdr herdr = twoLeads();
AtomicLong clock = new AtomicLong();
LeadTabScanner s = scanner(herdr, twoLeadsConfigured(), clock);
s.get();
int afterFirst = herdr.calls;
clock.addAndGet(TTL - 1);
s.get();
assertEquals(afterFirst, herdr.calls,
"resolve() runs on every request — an un-cached scan would put herdr on that path");
}
@Test
void aTabLabelledAfterStartupIsPickedUpOnceTheTtlExpires() {
TopologyHerdr herdr = twoLeads();
AtomicLong clock = new AtomicLong();
Map<String, String> tabToName = new LinkedHashMap<>(twoLeadsConfigured());
tabToName.put("lead: late-arrival", "late-arrival");
LeadTabScanner s = scanner(herdr, tabToName, clock);
assertFalse(s.get().containsKey("term_notes"));
herdr.tab("w1:t3", "w1", "lead: late-arrival"); // the operator renames their tab
clock.addAndGet(TTL);
assertEquals("late-arrival", s.get().get("term_notes"),
"the whole point over a config-held terminal_id: no config edit, no restart");
}
@Test
void aFailedScanKeepsTheLeadsAlreadyKnownRatherThanDemotingThem() {
TopologyHerdr herdr = twoLeads();
AtomicLong clock = new AtomicLong();
LeadTabScanner s = scanner(herdr, twoLeadsConfigured(), clock);
Map<String, String> before = s.get();
herdr.failing = true;
clock.addAndGet(TTL);
assertEquals(before, s.get(),
"a herdr hiccup must not silently demote a live lead to a worker mid-session");
}
@Test
void aFailedFirstScanReturnsEmptyRatherThanThrowing() {
TopologyHerdr herdr = twoLeads();
herdr.failing = true;
Map<String, String> leads = scanner(herdr, twoLeadsConfigured(), new AtomicLong()).get();
assertEquals(Map.of(), leads,
"with nothing scanned yet and no override to fall back on, the map is simply empty");
}
@Test
void aDownHerdrIsRetriedOncePerTtlNotOncePerRequest() {
TopologyHerdr herdr = twoLeads();
herdr.failing = true;
AtomicLong clock = new AtomicLong();
LeadTabScanner s = scanner(herdr, twoLeadsConfigured(), clock);
s.get();
int afterFirst = herdr.calls;
s.get();
s.get();
assertEquals(afterFirst, herdr.calls, "the failure path must be rate-limited too");
}
}
@@ -5,16 +5,15 @@ import org.junit.jupiter.api.Tag;
import org.junit.jupiter.api.Test;
import java.nio.file.Files;
import java.util.List;
import java.util.Map;
import static org.junit.jupiter.api.Assertions.*;
import static org.junit.jupiter.api.Assumptions.assumeTrue;
/**
* Contract test for the herdr half of connection-based identity against a REAL herdr: spawn a
* harmless probe, read its actual {@code shell_pid} from {@code pane.process_info}, and confirm
* {@link PaneLocator} resolves that PID back to the probe's own {@code terminal_id}.
* Contract test for {@link PaneLocator} against a REAL herdr: the PID→pane mapping that
* connection identity rests on. Uses a throwaway tab's seed shell as the probe process
* (protocol 19 removed arbitrary-command agents), and always tears the space down.
*
* <p>Tagged {@code contract}; run with {@code mvn test -Pcontract}.
*/
@@ -25,18 +24,24 @@ class PaneLocatorContractTest {
void resolvesTheTerminalOwningARealProcessPid() throws Exception {
assumeTrue(Files.exists(UnixSocketHerdrClient.defaultSocketPath()), "no herdr socket — skipping");
try (UnixSocketHerdrClient herdr = UnixSocketHerdrClient.connect()) {
AgentControl agents = new AgentControl(herdr);
Agent probe = agents.start("__pidprobe__", List.of("bash", "-c", "sleep 20"), Map.of());
WorkspaceControl spaces = new WorkspaceControl(herdr);
Workspace space = spaces.ensureWorkspace("__bridged_pid_contract__");
Tab.Created tab = spaces.createTab(space.workspaceId(), null, Map.of());
try {
JsonNode info = herdr.call("pane.process_info", Map.of("pane_id", probe.paneId()))
JsonNode info = herdr.call("pane.process_info", Map.of("pane_id", tab.rootPaneId()))
.path("process_info");
long shellPid = info.path("shell_pid").asLong(-1);
assertTrue(shellPid > 0, "probe pane should report a shell pid");
assertTrue(shellPid > 0, "seed pane should report a shell pid");
assertEquals(probe.terminalId(), new PaneLocator(herdr).terminalForPid(shellPid),
String terminalId = herdr.call("pane.get", Map.of("pane_id", tab.rootPaneId()))
.path("pane").path("terminal_id").asText(null);
assertNotNull(terminalId, "seed pane should carry a terminal_id");
assertEquals(terminalId, new PaneLocator(herdr).terminalForPid(shellPid),
"a real PID must resolve back to its own pane's terminal_id");
} finally {
agents.close(probe.paneId());
spaces.closeTab(tab.tab().tabId());
herdr.call("workspace.close", Map.of("workspace_id", space.workspaceId()));
}
}
}
@@ -5,7 +5,6 @@ import org.junit.jupiter.api.Tag;
import org.junit.jupiter.api.Test;
import java.nio.file.Files;
import java.util.List;
import java.util.Map;
import static org.junit.jupiter.api.Assertions.*;
@@ -13,9 +12,9 @@ import static org.junit.jupiter.api.Assumptions.assumeTrue;
/**
* Contract test for the placement layer ({@code workspace.*}/{@code tab.*}) against a
* REAL herdr, locking in the "one clean tab per worker" recipe: find-or-create a worker
* space, give the worker its own tab, drop herdr's seed shell so the tab holds only the
* worker, and tear it all down. Uses a HARMLESS probe (never {@code claude}) in a
* REAL herdr, locking in the "one clean tab per worker" recipe under protocol 19: find-or-create
* a worker space, give the worker its own tab, and the SEED pane is where the worker starts —
* the tab holds exactly that one pane from creation. Uses no agent (never {@code claude}) in a
* throwaway space that is fully removed at the end.
*
* <p>Tagged {@code contract}; run with {@code mvn test -Pcontract}.
@@ -41,7 +40,6 @@ class WorkspacePlacementContractTest {
void workerGetsOwnCleanTabAndTearsDownCompletely() throws Exception {
assumeTrue(!noSocket(), "no herdr socket — skipping");
try (UnixSocketHerdrClient herdr = UnixSocketHerdrClient.connect()) {
AgentControl agents = new AgentControl(herdr);
WorkspaceControl spaces = new WorkspaceControl(herdr);
Workspace space = spaces.ensureWorkspace(LABEL);
@@ -49,36 +47,27 @@ class WorkspacePlacementContractTest {
// Idempotent: a second ensure finds the same space, never creates a duplicate.
assertEquals(space.workspaceId(), spaces.ensureWorkspace(LABEL).workspaceId());
Tab.Created tab = spaces.createTab(space.workspaceId());
Agent worker = agents.start(
"__contract__",
List.of("bash", "-c", "sleep 20"),
Map.of(),
tab.tab().tabId());
Tab.Created tab = spaces.createTab(space.workspaceId(), null, Map.of());
try {
// The worker landed in its dedicated tab in the worker space.
assertEquals(tab.tab().tabId(), worker.tabId());
assertEquals(space.workspaceId(), worker.workspaceId());
assertNotNull(tab.rootPaneId(), "tab.create must return the seed pane");
// Drop the seed shell; the tab now holds exactly the worker pane.
agents.close(tab.rootPaneId());
// Protocol 19: the seed pane IS the worker pane — the tab holds exactly it.
spaces.renameTab(tab.tab().tabId(), "worker: contract");
assertEquals(1, paneCount(herdr, space.workspaceId(), tab.tab().tabId()),
"worker tab must hold only the worker pane after the seed shell is dropped");
"worker tab must hold exactly the seed/worker pane");
// Teardown resolves the tab from the pane, and sees it holds exactly one pane.
WorkspaceControl.PaneLocation loc = spaces.locatePane(worker.paneId());
WorkspaceControl.PaneLocation loc = spaces.locatePane(tab.rootPaneId());
assertNotNull(loc);
assertEquals(tab.tab().tabId(), loc.tabId());
assertEquals(1, loc.tabPaneCount(), "worker is the tab's sole occupant");
assertEquals(1, loc.tabPaneCount(), "worker pane is the tab's sole occupant");
} finally {
agents.close(worker.paneId());
spaces.closeTab(tab.tab().tabId());
}
// Tolerant teardown: closing an already-gone tab / reading a gone pane is a no-op.
spaces.closeTab(tab.tab().tabId());
assertNull(spaces.locatePane(worker.paneId()), "closed worker pane must be gone");
assertNull(spaces.locatePane(tab.rootPaneId()), "closed worker pane must be gone");
// Remove the throwaway space entirely so the test leaves no residue.
herdr.call("workspace.close", Map.of("workspace_id", space.workspaceId()));
@@ -1,9 +1,19 @@
package dev.ltms.bridged.inject;
import ch.qos.logback.classic.Level;
import ch.qos.logback.classic.LoggerContext;
import ch.qos.logback.classic.spi.ILoggingEvent;
import ch.qos.logback.core.read.ListAppender;
import dev.ltms.bridged.herdr.AgentControl;
import dev.ltms.bridged.herdr.FakeHerdr;
import dev.ltms.bridged.msg.Rendezvous;
import dev.ltms.bridged.msg.TestTurnTokens;
import dev.ltms.bridged.msg.TurnToken;
import org.junit.jupiter.api.Test;
import org.slf4j.LoggerFactory;
import java.util.Set;
import java.util.regex.Pattern;
import static org.junit.jupiter.api.Assertions.assertEquals;
import static org.junit.jupiter.api.Assertions.assertFalse;
@@ -16,7 +26,7 @@ class CompletionResolverTest {
void skipsTheScrapeWhenNoSendIsWaiting() {
FakeHerdr herdr = new FakeHerdr();
Rendezvous rendezvous = new Rendezvous(); // no waiter opened
CompletionResolver resolver = new CompletionResolver(new AgentControl(herdr), rendezvous);
CompletionResolver resolver = new CompletionResolver(new AgentControl(herdr), rendezvous, ExhaustedPatternLookup.none(), ExhaustionSink.none());
resolver.resolve("term_a", null); // no in-flight turn captured for this target
@@ -28,7 +38,7 @@ class CompletionResolverTest {
void failSkipsTheScrapeWhenNoSendIsWaiting() {
FakeHerdr herdr = new FakeHerdr();
Rendezvous rendezvous = new Rendezvous(); // no waiter opened
CompletionResolver resolver = new CompletionResolver(new AgentControl(herdr), rendezvous);
CompletionResolver resolver = new CompletionResolver(new AgentControl(herdr), rendezvous, ExhaustedPatternLookup.none(), ExhaustionSink.none());
resolver.fail("term_a", null); // no in-flight turn, and no registered waiter to fall back to
@@ -40,9 +50,9 @@ class CompletionResolverTest {
void captureBaselineSkipsTheReadWhenNoSendIsWaiting() {
FakeHerdr herdr = new FakeHerdr().readText("⏺ X\n❯ ");
Rendezvous rendezvous = new Rendezvous(); // no waiter opened
CompletionResolver resolver = new CompletionResolver(new AgentControl(herdr), rendezvous);
CompletionResolver resolver = new CompletionResolver(new AgentControl(herdr), rendezvous, ExhaustedPatternLookup.none(), ExhaustionSink.none());
resolver.captureBaseline("term_a"); // no send to attribute a later completion to
resolver.captureBaseline("term_a", TestTurnTokens.inert("term_a")); // no send to attribute a later completion to
assertFalse(herdr.called("agent.read"),
"with no waiting send there is no turn to baseline — skip the scrape");
@@ -129,7 +139,7 @@ class CompletionResolverTest {
// send must NOT be resolved with the stale answer.
FakeHerdr herdr = new FakeHerdr().readText("⏺ 391\n❯ ");
Rendezvous rendezvous = new Rendezvous();
CompletionResolver resolver = new CompletionResolver(new AgentControl(herdr), rendezvous);
CompletionResolver resolver = new CompletionResolver(new AgentControl(herdr), rendezvous, ExhaustedPatternLookup.none(), ExhaustionSink.none());
var waiter = rendezvous.open("term_a"); // a send is blocked on this turn
// The turn as captured at delivery: its waiter, and the previous turn's answer still on screen.
@@ -144,7 +154,7 @@ class CompletionResolverTest {
void resolvesACompletionWhoseScrapeChangedSinceDelivery() {
FakeHerdr herdr = new FakeHerdr().readText("⏺ No, 391 = 17 × 23.\n❯ "); // the worker's real answer
Rendezvous rendezvous = new Rendezvous();
CompletionResolver resolver = new CompletionResolver(new AgentControl(herdr), rendezvous);
CompletionResolver resolver = new CompletionResolver(new AgentControl(herdr), rendezvous, ExhaustedPatternLookup.none(), ExhaustionSink.none());
var waiter = rendezvous.open("term_a");
// Delivery baseline was the previous turn's "391"; the scrape now differs → resolve.
@@ -156,20 +166,63 @@ class CompletionResolverTest {
assertEquals("No, 391 = 17 × 23.", waiter.getNow(null).text());
}
@Test
void marksAClippedCompletionPaneTail() {
String block = "⏺ " + "x".repeat(CompletionResolver.MAX_SCRAPE_CHARS + 1) + "\n❯ ";
FakeHerdr herdr = new FakeHerdr().readText(block);
Rendezvous rendezvous = new Rendezvous();
CompletionResolver resolver = new CompletionResolver(new AgentControl(herdr), rendezvous, ExhaustedPatternLookup.none(), ExhaustionSink.none());
var waiter = rendezvous.open("term_a");
resolver.resolve("term_a", new CompletionResolver.InFlight(waiter, null));
assertEquals("x".repeat(CompletionResolver.MAX_SCRAPE_CHARS)
+ "\n[Pane tail clipped: member did not call bridge_reply.]",
waiter.getNow(null).text());
}
@Test
void leavesAnUnclippedCompletionPaneTailUnmarked() {
FakeHerdr herdr = new FakeHerdr().readText("⏺ complete report\n❯ ");
Rendezvous rendezvous = new Rendezvous();
CompletionResolver resolver = new CompletionResolver(new AgentControl(herdr), rendezvous, ExhaustedPatternLookup.none(), ExhaustionSink.none());
var waiter = rendezvous.open("term_a");
resolver.resolve("term_a", new CompletionResolver.InFlight(waiter, null));
assertEquals("complete report", waiter.getNow(null).text());
}
@Test
void resolvesSynchronouslyBeforePostTurnContextClearing() {
FakeHerdr herdr = new FakeHerdr().readText("⏺ previous answer\n❯ ");
Rendezvous rendezvous = new Rendezvous();
CompletionResolver resolver = new CompletionResolver(new AgentControl(herdr), rendezvous, ExhaustedPatternLookup.none(), ExhaustionSink.none());
var waiter = rendezvous.open("term_a");
resolver.captureBaseline("term_a", new TurnToken("term_a", waiter));
herdr.readText("⏺ answer that /clear would erase\n❯ ");
resolver.resolveBeforePostAction("term_a");
assertTrue(waiter.isDone(), "the answer is captured before the adapter sends /clear");
assertEquals("answer that /clear would erase", waiter.getNow(null).text());
}
@Test
void suppressesAnUnchangedCompletionEvenWhenTheBlockExceedsTheScrapeCap() {
// The fan-out issue-hunt finding: captureBaseline once stored the RAW (unclipped) assistant
// block while resolve compares against a clip()'d tail. For a block longer than MAX_SCRAPE_CHARS
// the two capped representations differ even when the pane never changed, so the CB-115
// byte-identical guard failed to fire and a stale completion could resolve the send. Both sides
// must clip identically; here an unchanged >cap block on rapid back-to-back turns stays suppressed.
// must clip identically. The returned-text marker is added only after this comparison, so an
// unchanged >cap block on rapid back-to-back turns still stays suppressed.
String longBlock = "⏺ " + "x".repeat(CompletionResolver.MAX_SCRAPE_CHARS + 500) + "\n❯ ";
FakeHerdr herdr = new FakeHerdr().readText(longBlock);
Rendezvous rendezvous = new Rendezvous();
CompletionResolver resolver = new CompletionResolver(new AgentControl(herdr), rendezvous);
CompletionResolver resolver = new CompletionResolver(new AgentControl(herdr), rendezvous, ExhaustedPatternLookup.none(), ExhaustionSink.none());
var waiter = rendezvous.open("term_a"); // a send is blocked on this turn
resolver.captureBaseline("term_a"); // baseline is the clipped >cap block
resolver.captureBaseline("term_a", new TurnToken("term_a", waiter)); // baseline is the clipped >cap block
var turn = resolver.inFlight("term_a");
assertEquals(CompletionResolver.MAX_SCRAPE_CHARS, turn.baseline().length(),
"the delivery baseline is clipped to the same cap resolve() applies to the tail");
@@ -186,7 +239,7 @@ class CompletionResolverTest {
// No delivery baseline (e.g. the pre-turn read failed) ⇒ never suppress; the completion resolves.
FakeHerdr herdr = new FakeHerdr().readText("⏺ hello\n❯ ");
Rendezvous rendezvous = new Rendezvous();
CompletionResolver resolver = new CompletionResolver(new AgentControl(herdr), rendezvous);
CompletionResolver resolver = new CompletionResolver(new AgentControl(herdr), rendezvous, ExhaustedPatternLookup.none(), ExhaustionSink.none());
var waiter = rendezvous.open("term_a");
resolver.resolve("term_a", new CompletionResolver.InFlight(waiter, null));
@@ -204,7 +257,7 @@ class CompletionResolverTest {
// byte-identical guard would wrongly match the empty tail and suppress.
FakeHerdr herdr = new FakeHerdr().healthy(false); // agent.read throws HerdrException
Rendezvous rendezvous = new Rendezvous();
CompletionResolver resolver = new CompletionResolver(new AgentControl(herdr), rendezvous);
CompletionResolver resolver = new CompletionResolver(new AgentControl(herdr), rendezvous, ExhaustedPatternLookup.none(), ExhaustionSink.none());
var waiter = rendezvous.open("term_a");
var turn = new CompletionResolver.InFlight(waiter, ""); // empty pane baselined at delivery
@@ -224,7 +277,7 @@ class CompletionResolverTest {
// fail must not overwrite that value, and must not even scrape the worker — nobody needs it.
FakeHerdr herdr = new FakeHerdr().readText("an error screen");
Rendezvous rendezvous = new Rendezvous();
CompletionResolver resolver = new CompletionResolver(new AgentControl(herdr), rendezvous);
CompletionResolver resolver = new CompletionResolver(new AgentControl(herdr), rendezvous, ExhaustedPatternLookup.none(), ExhaustionSink.none());
var waiter = rendezvous.open("term_a");
var turn = new CompletionResolver.InFlight(waiter, null);
@@ -245,7 +298,7 @@ class CompletionResolverTest {
// fail falls back to the waiter currently registered on the Rendezvous and fails it.
FakeHerdr herdr = new FakeHerdr().readText("stuck on an error screen");
Rendezvous rendezvous = new Rendezvous();
CompletionResolver resolver = new CompletionResolver(new AgentControl(herdr), rendezvous);
CompletionResolver resolver = new CompletionResolver(new AgentControl(herdr), rendezvous, ExhaustedPatternLookup.none(), ExhaustionSink.none());
var waiter = rendezvous.open("term_a"); // send registered, but no captureBaseline ever ran
resolver.fail("term_a", null); // no in-flight turn → fall back to the registered waiter
@@ -255,6 +308,40 @@ class CompletionResolverTest {
assertEquals("stuck on an error screen", waiter.getNow(null).text());
}
@Test
void failIsLoggedAtWarnWithTheReason() {
// CB-564: this used to be a bare DEBUG "failed send to X via turn-stall fallback" — a symptom
// with no cause, and below the level anyone watching for member health would see. A fail that
// resolves a caller's blocked send is at least WARN and must carry the reason.
LoggerContext ctx = (LoggerContext) LoggerFactory.getILoggerFactory();
ch.qos.logback.classic.Logger resolverLog =
(ch.qos.logback.classic.Logger) LoggerFactory.getLogger(CompletionResolver.class);
ListAppender<ILoggingEvent> appender = new ListAppender<>();
appender.setContext(ctx);
appender.start();
resolverLog.addAppender(appender);
resolverLog.setLevel(Level.WARN);
try {
FakeHerdr herdr = new FakeHerdr().readText("stuck on an error screen");
Rendezvous rendezvous = new Rendezvous();
CompletionResolver resolver = new CompletionResolver(new AgentControl(herdr), rendezvous, ExhaustedPatternLookup.none(), ExhaustionSink.none());
var waiter = rendezvous.open("term_a");
resolver.fail("term_a", null);
String warn = appender.list.stream()
.filter(e -> e.getLevel().equals(Level.WARN))
.map(ILoggingEvent::getFormattedMessage)
.findFirst()
.orElse("no turn-stall WARN logged");
assertTrue(warn.contains("term_a"), "the log names the target: " + warn);
assertTrue(warn.contains("stuck on an error screen"), "the log carries the reason: " + warn);
assertTrue(waiter.isDone());
} finally {
resolverLog.detachAppender(appender);
}
}
// --- CB-116 waiter identity: a late completion never crosses into the next turn ---------
@Test
@@ -265,16 +352,18 @@ class CompletionResolverTest {
// scrape to turn N+1; targeting turn N's captured waiter makes the late completion a no-op.
FakeHerdr herdr = new FakeHerdr().readText("⏺ turn N answer\n❯ ");
Rendezvous rendezvous = new Rendezvous();
CompletionResolver resolver = new CompletionResolver(new AgentControl(herdr), rendezvous);
CompletionResolver resolver = new CompletionResolver(new AgentControl(herdr), rendezvous, ExhaustedPatternLookup.none(), ExhaustionSink.none());
var waiterN = rendezvous.open("term_a"); // turn N's send
// The turn as the injector captured it at delivery (waiter + pre-turn baseline).
var turnN = new CompletionResolver.InFlight(waiterN, "an earlier answer");
// Turn N is resolved by the worker's explicit reply.
// Turn N is resolved by the worker's explicit reply, and its send deregisters the waiter.
assertTrue(rendezvous.resolve("term_a", "N replied"));
rendezvous.close("term_a", waiterN); // the sender's finally, before the next turn opens
// Turn N+1's send opens its own waiter on the same session (replacing the registered one).
// Turn N+1's send opens its own waiter on the same session (CB-548: open fails if the
// previous waiter is still registered, so a clean turn deregisters it first as above).
var waiterN1 = rendezvous.open("term_a");
resolver.resolve("term_a", turnN); // turn N's completion fallback finally fires
@@ -284,4 +373,129 @@ class CompletionResolverTest {
"turn N stays resolved by its own reply");
assertTrue(rendezvous.isWaiting("term_a"), "turn N+1 is still awaiting its own resolution");
}
// --- CB-578 stage A: backend-exhausted classification ---------------------------------
@Test
void classifiesAMatchingScrapeAsBackendExhaustedInsteadOfACompletedReply() {
String block = "⏺ Working on it...\nThe usage limit has been reached. Try again later.\n❯ ";
FakeHerdr herdr = new FakeHerdr().readText(block);
Rendezvous rendezvous = new Rendezvous();
ExhaustedPatternLookup patterns = target -> Pattern.compile("usage limit has been reached");
CompletionResolver resolver = new CompletionResolver(new AgentControl(herdr), rendezvous, patterns, ExhaustionSink.none());
var waiter = rendezvous.open("term_a");
resolver.resolve("term_a", new CompletionResolver.InFlight(waiter, null));
assertTrue(waiter.isDone(), "a matching scrape still resolves the blocked send");
assertEquals(Rendezvous.Kind.BACKEND_EXHAUSTED, waiter.getNow(null).kind(),
"not reported as a completed reply — the classification is distinct");
}
@Test
void theExhaustedReasonCarriesTheMatchedLine() {
String block = "⏺ Working on it...\nThe usage limit has been reached. Try again later.\n❯ ";
FakeHerdr herdr = new FakeHerdr().readText(block);
Rendezvous rendezvous = new Rendezvous();
ExhaustedPatternLookup patterns = target -> Pattern.compile("usage limit has been reached");
CompletionResolver resolver = new CompletionResolver(new AgentControl(herdr), rendezvous, patterns, ExhaustionSink.none());
var waiter = rendezvous.open("term_a");
resolver.resolve("term_a", new CompletionResolver.InFlight(waiter, null));
assertEquals("backend exhausted (usage limit): The usage limit has been reached. Try again later.",
waiter.getNow(null).text(), "the reason names the real cause and carries the matched line");
}
@Test
void aWinningBackendExhaustedClassificationNotifiesTheExhaustionSink() {
String block = "⏺ Working on it...\nThe usage limit has been reached. Try again later.\n❯ ";
FakeHerdr herdr = new FakeHerdr().readText(block);
Rendezvous rendezvous = new Rendezvous();
ExhaustedPatternLookup patterns = target -> Pattern.compile("usage limit has been reached");
java.util.List<String> notified = new java.util.ArrayList<>();
ExhaustionSink sink = (target, reason) -> notified.add(target + ": " + reason);
CompletionResolver resolver = new CompletionResolver(new AgentControl(herdr), rendezvous, patterns, sink);
var waiter = rendezvous.open("term_a");
resolver.resolve("term_a", new CompletionResolver.InFlight(waiter, null));
assertEquals(1, notified.size(), "the sink is notified exactly once for the winning classification");
assertTrue(notified.get(0).startsWith("term_a: "), "the sink is told which target exhausted");
assertTrue(notified.get(0).contains("The usage limit has been reached"),
"the sink is told the matched reason: " + notified.get(0));
}
@Test
void aLosingBackendExhaustedClassificationNeverNotifiesTheExhaustionSink() {
// The waiter was already resolved (e.g. by the worker's own reply) before this scrape landed —
// resolveExhausted loses the race and must return false, so the sink must not fire either.
String block = "⏺ Working on it...\nThe usage limit has been reached. Try again later.\n❯ ";
FakeHerdr herdr = new FakeHerdr().readText(block);
Rendezvous rendezvous = new Rendezvous();
ExhaustedPatternLookup patterns = target -> Pattern.compile("usage limit has been reached");
java.util.List<String> notified = new java.util.ArrayList<>();
ExhaustionSink sink = (target, reason) -> notified.add(target + ": " + reason);
CompletionResolver resolver = new CompletionResolver(new AgentControl(herdr), rendezvous, patterns, sink);
var waiter = rendezvous.open("term_a");
var turn = new CompletionResolver.InFlight(waiter, null);
assertTrue(rendezvous.resolveCompletion(waiter, "already replied"));
resolver.resolve("term_a", turn);
assertTrue(notified.isEmpty(), "a classification that loses the race must not quarantine anything");
assertEquals("already replied", waiter.getNow(null).text(), "the earlier resolution stands untouched");
}
@Test
void aNonMatchingScrapeResolvesAsAnOrdinaryCompletion() {
FakeHerdr herdr = new FakeHerdr().readText("⏺ complete report\n❯ ");
Rendezvous rendezvous = new Rendezvous();
ExhaustedPatternLookup patterns = target -> Pattern.compile("usage limit has been reached");
CompletionResolver resolver = new CompletionResolver(new AgentControl(herdr), rendezvous, patterns, ExhaustionSink.none());
var waiter = rendezvous.open("term_a");
resolver.resolve("term_a", new CompletionResolver.InFlight(waiter, null));
assertEquals(Rendezvous.Kind.COMPLETION, waiter.getNow(null).kind(),
"a scrape that does not match the pattern is an ordinary completion");
assertEquals("complete report", waiter.getNow(null).text());
}
@Test
void aProfileWithNoConfiguredPatternKeepsTodaysCompletionFallbackUnchanged() {
// Even a scrape that WOULD have matched some other profile's pattern must resolve as a
// plain completion when this target's own profile has none configured (CB-578 criterion 4).
String block = "⏺ The usage limit has been reached.\n❯ ";
FakeHerdr herdr = new FakeHerdr().readText(block);
Rendezvous rendezvous = new Rendezvous();
CompletionResolver resolver =
new CompletionResolver(new AgentControl(herdr), rendezvous, ExhaustedPatternLookup.none(), ExhaustionSink.none());
var waiter = rendezvous.open("term_a");
resolver.resolve("term_a", new CompletionResolver.InFlight(waiter, null));
assertEquals(Rendezvous.Kind.COMPLETION, waiter.getNow(null).kind(),
"no pattern configured for this target's profile ⇒ unchanged completion-fallback behaviour");
assertEquals("The usage limit has been reached.", waiter.getNow(null).text());
}
@Test
void coverageIsOffWhenNoProfileHasAPatternConfigured() {
assertEquals("off (no profile has an exhaustedPattern configured; profiles: [terra])",
CompletionResolver.coverage(Set.of("terra"), Set.of()));
}
@Test
void coverageIsFullWhenEveryProfileHasAPatternConfigured() {
assertEquals("full (all profiles configured: [gx10, terra])",
CompletionResolver.coverage(Set.of("terra", "gx10"), Set.of("terra", "gx10")));
}
@Test
void coverageIsPartialAndNamesWhichProfilesAreConfigured() {
assertEquals("partial (configured: [terra]; not configured: [gx10])",
CompletionResolver.coverage(Set.of("terra", "gx10"), Set.of("terra")));
}
}
@@ -1,10 +1,16 @@
package dev.ltms.bridged.inject;
import ch.qos.logback.classic.Level;
import ch.qos.logback.classic.LoggerContext;
import ch.qos.logback.classic.spi.ILoggingEvent;
import ch.qos.logback.core.read.ListAppender;
import dev.ltms.bridged.herdr.AgentControl;
import dev.ltms.bridged.herdr.AgentStatus;
import dev.ltms.bridged.herdr.FakeHerdr;
import dev.ltms.bridged.herdr.HerdrException;
import dev.ltms.bridged.msg.TestTurnTokens;
import org.junit.jupiter.api.Test;
import org.slf4j.LoggerFactory;
import java.util.ArrayList;
import java.util.List;
@@ -28,23 +34,21 @@ class InjectorTest {
private final Injector injector = new Injector(new AgentControl(herdr));
/**
* The logical messages delivered, in order. AgentControl.send emits each delivery as two
* agent.send calls — the payload, then a standalone Enter keystroke ({@code "\r"}) to submit
* it; these tests assert delivery ordering/gating, not the submit event, so drop the bare
* carriage returns.
* The logical messages delivered, in order. Under protocol 19 each delivery is one
* {@code agent.prompt} carrying the payload (it submits itself); the Enter nudge is a
* separate {@code agent.send_keys} and never appears here.
*/
@SuppressWarnings("unchecked")
private List<String> sent() {
return herdr.calls.stream()
.filter(c -> c.method().equals("agent.send"))
.filter(c -> c.method().equals("agent.prompt"))
.map(c -> ((Map<String, Object>) c.params()).get("text").toString())
.filter(t -> !t.equals("\r"))
.toList();
}
@Test
void deliversWhenIdle() {
CompletableFuture<Void> f = injector.enqueue(T, "hello");
CompletableFuture<Void> f = injector.enqueue(T, "hello", TestTurnTokens.inert(T));
assertFalse(f.isDone(), "not delivered until an injectable status arrives");
injector.onStatus(T, AgentStatus.IDLE);
assertTrue(f.isDone());
@@ -56,7 +60,7 @@ class InjectorTest {
// CB-113: idle alone is not enough — hold until the worker's MCP is connected (ready).
java.util.Set<String> ready = new java.util.HashSet<>();
Injector inj = new Injector(new AgentControl(herdr), TurnListener.NOOP, ready::contains);
inj.enqueue(T, "task");
inj.enqueue(T, "task", TestTurnTokens.inert(T));
inj.onStatus(T, AgentStatus.IDLE); // idle but not yet available → held out of the boot window
assertEquals(List.of(), sent(), "must not deliver into a not-yet-available worker");
@@ -66,11 +70,9 @@ class InjectorTest {
assertEquals(List.of("task"), sent(), "delivers once the worker is available");
}
@SuppressWarnings("unchecked")
private long enterKeystrokes() {
return herdr.calls.stream()
.filter(c -> c.method().equals("agent.send"))
.filter(c -> "\r".equals(((Map<String, Object>) c.params()).get("text")))
.filter(c -> c.method().equals("agent.send_keys"))
.count();
}
@@ -78,7 +80,7 @@ class InjectorTest {
void resubmitsEnterWhenADeliveredMessageIsNotPickedUp() {
// CB-113: the Enter at delivery can race the paste; while the worker stays idle (not picked
// up), the injector re-nudges Enter so the pending paste submits.
injector.enqueue(T, "task");
injector.enqueue(T, "task", TestTurnTokens.inert(T));
injector.onStatus(T, AgentStatus.IDLE); // deliver: paste + one Enter
long afterDeliver = enterKeystrokes();
@@ -94,7 +96,7 @@ class InjectorTest {
@Test
void holdsWhileWorkingThenDeliversOnIdle() {
injector.enqueue(T, "later");
injector.enqueue(T, "later", TestTurnTokens.inert(T));
injector.onStatus(T, AgentStatus.WORKING);
assertEquals(List.of(), sent(), "must not inject mid-turn");
injector.onStatus(T, AgentStatus.IDLE);
@@ -103,7 +105,7 @@ class InjectorTest {
@Test
void blockedIsInjectableButUnknownIsNot() {
injector.enqueue(T, "answer");
injector.enqueue(T, "answer", TestTurnTokens.inert(T));
injector.onStatus(T, AgentStatus.UNKNOWN);
assertEquals(List.of(), sent(), "unknown status is not safe to inject");
injector.onStatus(T, AgentStatus.BLOCKED);
@@ -112,8 +114,8 @@ class InjectorTest {
@Test
void twoRapidDeliveriesNeverInterleave() {
injector.enqueue(T, "m1");
injector.enqueue(T, "m2");
injector.enqueue(T, "m1", TestTurnTokens.inert(T));
injector.enqueue(T, "m2", TestTurnTokens.inert(T));
// First idle window delivers only m1, even if idle is observed twice before pickup.
injector.onStatus(T, AgentStatus.IDLE);
@@ -128,8 +130,8 @@ class InjectorTest {
@Test
void transientUnknownDoesNotReleaseThePickupLatch() {
injector.enqueue(T, "m1");
injector.enqueue(T, "m2");
injector.enqueue(T, "m1", TestTurnTokens.inert(T));
injector.enqueue(T, "m2", TestTurnTokens.inert(T));
injector.onStatus(T, AgentStatus.IDLE); // m1 sent, awaiting pickup
assertEquals(List.of("m1"), sent());
@@ -144,8 +146,8 @@ class InjectorTest {
@Test
void missedPickupEdgeIsReleasedByGraceSoTheQueueNeverWedges() {
injector.enqueue(T, "m1");
injector.enqueue(T, "m2");
injector.enqueue(T, "m1", TestTurnTokens.inert(T));
injector.enqueue(T, "m2", TestTurnTokens.inert(T));
injector.onStatus(T, AgentStatus.IDLE); // m1 sent
assertEquals(List.of("m1"), sent());
@@ -157,9 +159,9 @@ class InjectorTest {
@Test
void fifoOrderAcrossManyTurns() {
injector.enqueue(T, "a");
injector.enqueue(T, "b");
injector.enqueue(T, "c");
injector.enqueue(T, "a", TestTurnTokens.inert(T));
injector.enqueue(T, "b", TestTurnTokens.inert(T));
injector.enqueue(T, "c", TestTurnTokens.inert(T));
for (int i = 0; i < 3; i++) {
injector.onStatus(T, AgentStatus.IDLE); // deliver one
injector.onStatus(T, AgentStatus.WORKING); // pickup
@@ -171,7 +173,7 @@ class InjectorTest {
@Test
void activeWhileQueuedOrInFlightThenQuietAfterTurnCompletes() {
assertTrue(injector.activeTargets().isEmpty());
injector.enqueue(T, "x");
injector.enqueue(T, "x", TestTurnTokens.inert(T));
assertEquals(Set.of(T), injector.activeTargets(), "active while a message is queued");
injector.onStatus(T, AgentStatus.IDLE); // delivers; awaiting pickup
@@ -190,7 +192,7 @@ class InjectorTest {
void firesTurnCompleteOnAConfirmedWorkingThenIdle() {
List<String> completed = new ArrayList<>();
Injector inj = new Injector(new AgentControl(herdr), completed::add);
inj.enqueue(T, "task");
inj.enqueue(T, "task", TestTurnTokens.inert(T));
inj.onStatus(T, AgentStatus.IDLE); // deliver
inj.onStatus(T, AgentStatus.WORKING); // pickup + turn running
@@ -200,11 +202,52 @@ class InjectorTest {
assertEquals(List.of(T), completed, "a confirmed working→idle fires exactly one completion");
}
@Test
void postTurnResetSettlesBeforeNextDelegationWithoutBecomingATurn() {
AgentControl agents = new AgentControl(herdr);
class ResetListener implements TurnListener {
int completed;
@Override
public void onTurnComplete(String target) {
completed++;
}
@Override
public boolean hasPostTurnAction(String target) {
return true;
}
@Override
public boolean onTurnCompleteWithPostAction(String target) {
completed++;
agents.send(target, "/clear"); // direct housekeeping, never Injector.enqueue
return true;
}
}
ResetListener listener = new ResetListener();
Injector inj = new Injector(agents, listener);
inj.enqueue(T, "first", TestTurnTokens.inert(T));
inj.enqueue(T, "second", TestTurnTokens.inert(T));
inj.onStatus(T, AgentStatus.IDLE); // first delegation
inj.onStatus(T, AgentStatus.WORKING);
inj.onStatus(T, AgentStatus.IDLE); // first complete; reset dispatched
assertEquals(List.of("first", "/clear"), sent(),
"same-tick completion must not let the queued delegation overtake reset");
inj.onStatus(T, AgentStatus.WORKING); // reset picked up, but this is not a bridge turn
inj.onStatus(T, AgentStatus.IDLE); // reset settled; second may now deliver
assertEquals(List.of("first", "/clear", "second"), sent());
assertEquals(1, listener.completed,
"reset settlement must not recursively emit another turn completion");
}
@Test
void doesNotSynthesizeCompletionFromAnUnconfirmedTurn() {
List<String> completed = new ArrayList<>();
Injector inj = new Injector(new AgentControl(herdr), completed::add);
inj.enqueue(T, "task");
inj.enqueue(T, "task", TestTurnTokens.inert(T));
// Deliver, then only ever idle — a `working` sample is never seen. The pickup grace unwedges
// the queue but must NOT invent a completion: without a sampled turn there is no trustworthy
@@ -217,6 +260,7 @@ class InjectorTest {
private static final class Captor implements TurnListener {
final List<String> completed = new ArrayList<>();
final List<String> failed = new ArrayList<>();
final List<String> failureReasons = new ArrayList<>();
@Override
public void onTurnComplete(String target) {
@@ -227,6 +271,12 @@ class InjectorTest {
public void onTurnFailed(String target) {
failed.add(target);
}
@Override
public void onTurnFailed(String target, String reason) {
failed.add(target);
failureReasons.add(reason);
}
}
// ~30s of unknown at the 250ms prod poll interval; enough onStatus samples to trip the stall.
@@ -236,7 +286,7 @@ class InjectorTest {
void failsAnOutstandingDelegationWhoseWorkerWedgesInUnknown() {
Captor cap = new Captor();
Injector inj = new Injector(new AgentControl(herdr), cap);
inj.enqueue(T, "task");
inj.enqueue(T, "task", TestTurnTokens.inert(T));
inj.onStatus(T, AgentStatus.IDLE); // deliver
inj.onStatus(T, AgentStatus.WORKING); // worker starts the turn
@@ -251,7 +301,7 @@ class InjectorTest {
void aTransientUnknownGlitchNeitherFailsNorBlocksCompletion() {
Captor cap = new Captor();
Injector inj = new Injector(new AgentControl(herdr), cap);
inj.enqueue(T, "task");
inj.enqueue(T, "task", TestTurnTokens.inert(T));
inj.onStatus(T, AgentStatus.IDLE); // deliver
inj.onStatus(T, AgentStatus.WORKING); // confirmed turn
@@ -266,7 +316,7 @@ class InjectorTest {
void sendFailureDropsMessageAndFailsItsFuture() {
FakeHerdr failing = new FakeHerdr().agentSendFailsWith("send_failed");
Injector inj = new Injector(new AgentControl(failing));
CompletableFuture<Void> f = inj.enqueue(T, "boom");
CompletableFuture<Void> f = inj.enqueue(T, "boom", TestTurnTokens.inert(T));
inj.onStatus(T, AgentStatus.IDLE);
assertTrue(f.isCompletedExceptionally());
@@ -275,18 +325,35 @@ class InjectorTest {
@Test
void dropFailsPendingWaiters() {
CompletableFuture<Void> f = injector.enqueue(T, "orphan");
CompletableFuture<Void> f = injector.enqueue(T, "orphan", TestTurnTokens.inert(T));
injector.drop(T, new HerdrException("worker gone", "pane_not_found", null));
assertTrue(f.isCompletedExceptionally(), "queued waiters unblock when the worker vanishes");
}
@Test
void dropPassesTheRealCauseForQueuedAndDeliveredWork() {
Captor cap = new Captor();
Injector inj = new Injector(new AgentControl(herdr), cap);
CompletableFuture<Void> delivered = inj.enqueue(T, "delivered", TestTurnTokens.inert(T));
CompletableFuture<Void> queued = inj.enqueue(T, "queued", TestTurnTokens.inert(T));
inj.onStatus(T, AgentStatus.IDLE); // deliver the first message
inj.onStatus(T, AgentStatus.WORKING); // its turn is now in flight; one remains queued
inj.drop(T, new HerdrException("agent target sol not found", "agent_not_found", null));
assertEquals(List.of(T), cap.failed, "drop signals one turn failure for both affected states");
assertEquals(List.of("agent target sol not found"), cap.failureReasons);
assertTrue(delivered.isDone(), "the delivered future has already completed");
assertTrue(queued.isCompletedExceptionally(), "the queued future fails with the drop cause");
}
@Test
void dropFailsTheTurnOfADeliveredMessageWhenTheWorkerVanishes() {
// CB-110: the message was delivered (no longer queued), so failing queued waiters alone would
// leave its send hanging. A vanished worker must fail that in-flight turn too.
Captor cap = new Captor();
Injector inj = new Injector(new AgentControl(herdr), cap);
inj.enqueue(T, "task");
inj.enqueue(T, "task", TestTurnTokens.inert(T));
inj.onStatus(T, AgentStatus.IDLE); // deliver
inj.onStatus(T, AgentStatus.WORKING); // turn running
@@ -301,7 +368,7 @@ class InjectorTest {
// by an off-sub worker's review of CB-110, delegated through the bridge.)
Captor cap = new Captor();
Injector inj = new Injector(new AgentControl(herdr), cap);
inj.enqueue(T, "task");
inj.enqueue(T, "task", TestTurnTokens.inert(T));
inj.onStatus(T, AgentStatus.IDLE); // deliver; pickup never confirmed
inj.drop(T, new HerdrException("worker gone", "pane_not_found", null));
@@ -320,7 +387,7 @@ class InjectorTest {
Captor cap = new Captor();
List<String> forgotten = new ArrayList<>();
Injector inj = new Injector(new AgentControl(herdr), cap, _ -> false, forgotten::add);
CompletableFuture<Void> f = inj.enqueue(T, "task");
CompletableFuture<Void> f = inj.enqueue(T, "task", TestTurnTokens.inert(T));
for (int i = 0; i < READINESS_SAMPLES; i++) inj.onStatus(T, AgentStatus.IDLE);
@@ -332,6 +399,40 @@ class InjectorTest {
assertTrue(inj.activeTargets().isEmpty(), "the target is reclaimed, not polled forever");
}
@Test
void readinessGraceExpiryIsLogged() {
// CB-562: the grace-expiry path used to clear the queue silently, so a message that never
// reached the worker's pane surfaced elsewhere as an unrelated turn-stall failure. Assert the
// expiry now names the real cause. (ListAppender capture pattern mirrors AuditLogTest.)
LoggerContext ctx = (LoggerContext) LoggerFactory.getILoggerFactory();
ch.qos.logback.classic.Logger injectorLog =
(ch.qos.logback.classic.Logger) LoggerFactory.getLogger(Injector.class);
ListAppender<ILoggingEvent> appender = new ListAppender<>();
appender.setContext(ctx);
appender.start();
injectorLog.addAppender(appender);
injectorLog.setLevel(Level.WARN);
try {
Injector inj = new Injector(new AgentControl(herdr), TurnListener.NOOP, _ -> false, _ -> {
});
inj.enqueue(T, "task", TestTurnTokens.inert(T));
for (int i = 0; i < READINESS_SAMPLES; i++) inj.onStatus(T, AgentStatus.IDLE);
String warn = appender.list.stream()
.filter(e -> e.getLevel().equals(Level.WARN))
.map(ILoggingEvent::getFormattedMessage)
.findFirst()
.orElse("no grace-expiry WARN logged");
assertTrue(warn.contains(T), "the log names the target terminal: " + warn);
assertTrue(warn.contains("never reached"), "the log names the real cause: " + warn);
assertTrue(warn.contains("1 queued message"),
"the log carries the failed message count: " + warn);
} finally {
injectorLog.detachAppender(appender);
}
}
@Test
void aWorkerThatBecomesReadyWithinTheGraceIsDeliveredNormally() {
// The readiness grace must not fail a worker that is merely slow to boot: once it becomes
@@ -339,7 +440,7 @@ class InjectorTest {
Set<String> ready = new java.util.HashSet<>();
Injector inj = new Injector(new AgentControl(herdr), TurnListener.NOOP, ready::contains, _ -> {
});
inj.enqueue(T, "task");
inj.enqueue(T, "task", TestTurnTokens.inert(T));
for (int i = 0; i < 100; i++) inj.onStatus(T, AgentStatus.IDLE); // still booting, well under grace
assertEquals(List.of(), sent());
@@ -352,14 +453,46 @@ class InjectorTest {
@Test
void dropClearsWorkerPresence() {
// CB-114 (finding #1): a vanished worker's readiness must be forgotten so a stale entry cannot
// linger past the worker's life (WorkerPresence.forget had no caller before this).
// linger past the worker's life (MemberPresence.forget had no caller before this).
List<String> forgotten = new ArrayList<>();
Injector inj = new Injector(new AgentControl(herdr), TurnListener.NOOP, _ -> true, forgotten::add);
inj.enqueue(T, "orphan");
inj.enqueue(T, "orphan", TestTurnTokens.inert(T));
inj.drop(T, new HerdrException("worker gone", "pane_not_found", null));
assertEquals(List.of(T), forgotten, "drop clears the gone worker's presence");
}
@Test
void dropIsLogged() {
// CB-564: a vanished worker used to drop its queue with no log at all — the only trace was
// whatever failed downstream (e.g. a caller's send timing out with no clue why). Assert the
// drop itself now names the cause and the number of messages it failed.
LoggerContext ctx = (LoggerContext) LoggerFactory.getILoggerFactory();
ch.qos.logback.classic.Logger injectorLog =
(ch.qos.logback.classic.Logger) LoggerFactory.getLogger(Injector.class);
ListAppender<ILoggingEvent> appender = new ListAppender<>();
appender.setContext(ctx);
appender.start();
injectorLog.addAppender(appender);
injectorLog.setLevel(Level.WARN);
try {
Injector inj = new Injector(new AgentControl(herdr), TurnListener.NOOP, _ -> true, _ -> {
});
inj.enqueue(T, "orphan", TestTurnTokens.inert(T));
inj.drop(T, new HerdrException("worker gone", "pane_not_found", null));
String warn = appender.list.stream()
.filter(e -> e.getLevel().equals(Level.WARN))
.map(ILoggingEvent::getFormattedMessage)
.findFirst()
.orElse("no drop WARN logged");
assertTrue(warn.contains(T), "the log names the target terminal: " + warn);
assertTrue(warn.contains("1 message"), "the log carries the failed message count: " + warn);
assertTrue(warn.contains("worker gone"), "the log carries the real cause: " + warn);
} finally {
injectorLog.detachAppender(appender);
}
}
@Test
void pollerDeliversToAnIdleWorker() throws Exception {
// End-to-end through the poller: idle worker → message delivered without manual onStatus.
@@ -368,19 +501,18 @@ class InjectorTest {
StatusPoller poller = new StatusPoller(new AgentControl(idle), inj, 10);
poller.start();
try {
CompletableFuture<Void> delivered = inj.enqueue(T, "via-poller");
CompletableFuture<Void> delivered = inj.enqueue(T, "via-poller", TestTurnTokens.inert(T));
delivered.get(2, TimeUnit.SECONDS); // completes when the poller drives the send
} finally {
poller.stop();
}
assertEquals(List.of("via-poller"), idle.calls.stream()
.filter(c -> c.method().equals("agent.send"))
.filter(c -> c.method().equals("agent.prompt"))
.map(c -> {
@SuppressWarnings("unchecked")
Map<String, Object> p = (Map<String, Object>) c.params();
return p.get("text").toString();
})
.filter(t -> !t.equals("\r")) // drop the standalone submit keystroke
.toList());
}
@@ -388,7 +520,7 @@ class InjectorTest {
void deliveredFutureCarriesSendFailure() {
FakeHerdr failing = new FakeHerdr().agentSendFailsWith("send_failed");
Injector inj = new Injector(new AgentControl(failing));
CompletableFuture<Void> f = inj.enqueue(T, "boom");
CompletableFuture<Void> f = inj.enqueue(T, "boom", TestTurnTokens.inert(T));
inj.onStatus(T, AgentStatus.IDLE);
ExecutionException ex = assertThrows(ExecutionException.class, f::get);
assertInstanceOf(HerdrException.class, ex.getCause());
@@ -11,7 +11,7 @@ class WorkerPresenceTest {
@Test
void tracksPresenceAndForgets() {
WorkerPresence p = new WorkerPresence();
MemberPresence p = new MemberPresence();
assertFalse(p.isPresent("term_a"), "unseen worker is not available");
p.markPresent("term_a");
@@ -23,7 +23,7 @@ class WorkerPresenceTest {
@Test
void nullOrBlankMarkIsANoOp() {
WorkerPresence p = new WorkerPresence();
MemberPresence p = new MemberPresence();
assertDoesNotThrow(() -> {
p.markPresent(null);
p.markPresent(" ");
@@ -0,0 +1,258 @@
package dev.ltms.bridged.lead;
import dev.ltms.bridged.config.BridgedConfig;
import dev.ltms.bridged.herdr.AgentControl;
import dev.ltms.bridged.herdr.FakeHerdr;
import dev.ltms.bridged.herdr.WorkspaceControl;
import org.junit.jupiter.api.Test;
import java.util.LinkedHashMap;
import java.util.List;
import java.util.Map;
import static org.junit.jupiter.api.Assertions.*;
/**
* CB-558 — the daemon starts a declared lead when none is running.
*
* <p>Two properties carry the whole feature. It must not double-spawn (a second orchestrator is
* worse than none), and what it starts must be a <em>lead</em> and not a member: no reply charter,
* no off-subscription env, and never registered with the session lifecycle.
*/
class LeadLauncherTest {
/** The lead's backend: a subscription profile with the bridge mounted, as `opus` really is. */
private static BridgedConfig.Profile opusProfile() {
return new BridgedConfig.Profile(
"opus", null, "claude-opus-5", null, "BRIDGED_WORKER_TOKEN",
List.of("ccs", "ltms"), "tab", "bridged-workers", null,
"http://127.0.0.1:8765/mcp", null, null,
null, null, null,
Map.of("CLAUDE_CODE_AUTO_COMPACT_WINDOW", "300000"), null, null, true, null);
}
private static BridgedConfig configWith(BridgedConfig.Leader lead) {
Map<String, BridgedConfig.Leader> leaders = new LinkedHashMap<>();
leaders.put("opus", lead);
BridgedConfig.Fleet fleet =
new BridgedConfig.Fleet(leaders, Map.of(), Map.of(), Map.of(), null);
return new BridgedConfig(
null, null, Map.of("opus", opusProfile()), null, null, null, null, null,
null, null, fleet, null, "fixed", null).withDefaults();
}
private static BridgedConfig.Leader lead(String profile, String tab, int instances) {
return new BridgedConfig.Leader(profile, tab, instances, "lead:", 10, null, null,
"leads", "/repo");
}
private static LeadLauncher launcher(FakeHerdr herdr, BridgedConfig cfg) {
return new LeadLauncher(new AgentControl(herdr), new WorkspaceControl(herdr), cfg);
}
@SuppressWarnings("unchecked")
private static List<String> startedArgs(FakeHerdr herdr) {
return (List<String>) ((Map<String, Object>) herdr.lastCall("agent.start").params()).get("args");
}
private static String startedName(FakeHerdr herdr) {
return (String) ((Map<?, ?>) herdr.lastCall("agent.start").params()).get("name");
}
@SuppressWarnings("unchecked")
private static Map<String, String> tabEnv(FakeHerdr herdr) {
return (Map<String, String>) ((Map<String, Object>) herdr.lastCall("tab.create").params()).get("env");
}
// ── it starts a lead when none is live ────────────────────────────────────────────────────
@Test
void startsTheDeclaredLeadWhenNoneIsRunning() {
FakeHerdr herdr = new FakeHerdr();
assertEquals(1, launcher(herdr, configWith(lead("opus", "lead: opus", 1))).ensureLeads());
assertTrue(herdr.called("agent.start"), "a lead must actually be started");
assertEquals("lead-opus", startedName(herdr));
}
/** The tab is labelled with the configured `tab:` so the scanner finds the lead on the next resolve. */
@Test
void labelsTheTabWithTheConfiguredTabValue() {
FakeHerdr herdr = new FakeHerdr();
launcher(herdr, configWith(lead("opus", "lead: opus", 1))).ensureLeads();
assertEquals("lead: opus",
((Map<?, ?>) herdr.lastCall("tab.rename").params()).get("label"));
}
/** `instances: 2` with none live means two starts, not one. */
@Test
void startsAsManyInstancesAsAreDeclared() {
FakeHerdr herdr = new FakeHerdr();
assertEquals(2, launcher(herdr, configWith(lead("opus", "lead: opus", 2))).ensureLeads());
assertEquals(2, herdr.calls.stream().filter(c -> c.method().equals("agent.start")).count());
}
// ── it must not double-spawn ──────────────────────────────────────────────────────────────
/** A labelled tab WITH a running agent in it is a live lead — leave it alone. */
@Test
void doesNotStartASecondLeadWhenOneIsAlreadyRunning() {
FakeHerdr herdr = new FakeHerdr()
.withWorkspace("wL", "leads")
.withTab("wL", "wL:t1", "lead: opus")
.withAgent("lead-opus", "term_lead", "wL:p1", "wL:t1");
assertEquals(0, launcher(herdr, configWith(lead("opus", "lead: opus", 1))).ensureLeads());
assertFalse(herdr.called("agent.start"), "the live lead must not be duplicated");
}
/**
* The reason liveness is not "does the label exist". A tab left labelled by a session that has
* since died must not block the relaunch, or one crash disables auto-launch permanently.
*/
@Test
void aLabelledTabWithNoRunningAgentIsNotALiveLead() {
FakeHerdr herdr = new FakeHerdr()
.withWorkspace("wL", "leads")
.withTab("wL", "wL:t1", "lead: opus"); // label only — nothing running in it
assertEquals(1, launcher(herdr, configWith(lead("opus", "lead: opus", 1))).ensureLeads(),
"a stale label is not a lead; the lead must be relaunched");
}
/**
* A lead the operator opened by hand is live once its tab carries the configured `tab:` label —
* CB-579 retired the `terminal:` pin, so a hand-opened lead is found the same way an
* auto-launched one is, by its tab, not by a terminal id nobody wrote down in advance.
*/
@Test
void aHandOpenedLeadWithTheConfiguredTabLabelCountsAsLive() {
FakeHerdr herdr = new FakeHerdr()
.withWorkspace("wX", "main")
.withTab("wX", "wX:t1", "lead: opus")
.withAgent("hand-opened", "term_hand", "wX:p1", "wX:t1");
assertEquals(0, launcher(herdr, configWith(lead("opus", "lead: opus", 1))).ensureLeads());
assertFalse(herdr.called("agent.start"));
}
/** A member sitting in a matching tab must never be counted — or spawn a lead — as one. */
@Test
void aMemberWorkspaceIsNeverScannedForLeads() {
FakeHerdr herdr = new FakeHerdr()
.withWorkspace("wM", "bridged-workers") // a configured member space
.withTab("wM", "wM:t1", "lead: opus") // a member tab that looks like a lead
.withAgent("claude-opus-x", "term_m", "wM:p1", "wM:t1");
assertEquals(1, launcher(herdr, configWith(lead("opus", "lead: opus", 1))).ensureLeads(),
"a member in a lead-labelled tab is not a lead, so the real lead is still missing");
}
/** If herdr cannot be counted, start nothing: guessing risks a second orchestrator. */
@Test
void anUncountableHerdrStartsNothing() {
FakeHerdr herdr = new FakeHerdr().healthy(false);
assertEquals(0, launcher(herdr, configWith(lead("opus", "lead: opus", 1))).ensureLeads());
assertFalse(herdr.called("agent.start"));
}
// ── what it starts is a LEAD, not a member ────────────────────────────────────────────────
/**
* The single most important assertion here. The worker charter tells its reader it is an
* off-subscription worker that must end every turn with bridge_reply — the opposite of what an
* orchestrator is. A lead must never receive it.
*/
@Test
void theLeadNeverReceivesTheWorkerReplyCharter() {
FakeHerdr herdr = new FakeHerdr();
launcher(herdr, configWith(lead("opus", "lead: opus", 1))).ensureLeads();
List<String> args = startedArgs(herdr);
assertFalse(args.contains("--append-system-prompt"),
"the reply charter is a worker contract and must not be injected into a lead");
assertTrue(args.stream().noneMatch(a -> a.contains("bridge_reply")), args.toString());
}
/** It still mounts the bridge — a lead that cannot orchestrate is pointless. */
@Test
void theLeadMountsTheBridgeMcpAndPinsItsModel() {
FakeHerdr herdr = new FakeHerdr();
launcher(herdr, configWith(lead("opus", "lead: opus", 1))).ensureLeads();
List<String> args = startedArgs(herdr);
assertTrue(args.contains("--mcp-config"));
assertTrue(args.stream().anyMatch(a -> a.contains("http://127.0.0.1:8765/mcp")), args.toString());
assertEquals("claude-opus-5", args.get(args.indexOf("--model") + 1));
assertTrue(args.indexOf("--model") > args.indexOf("--mcp-config"),
"--model is appended last so it outranks the ccs wrapper (CB-533)");
}
/** A lead runs on the operator's subscription. Nothing may move it off. */
@Test
void theLeadEnvCarriesNoAnthropicBinding() {
FakeHerdr herdr = new FakeHerdr();
launcher(herdr, configWith(lead("opus", "lead: opus", 1))).ensureLeads();
Map<String, String> env = tabEnv(herdr);
assertNull(env.get("ANTHROPIC_BASE_URL"));
assertNull(env.get("ANTHROPIC_AUTH_TOKEN"));
assertEquals("300000", env.get("CLAUDE_CODE_AUTO_COMPACT_WINDOW"),
"the profile's own env: still applies");
}
/** The lead's tab goes in its own workspace, never a member one — the scanner skips those. */
@Test
void theLeadTabIsCreatedOutsideEveryMemberWorkspace() {
FakeHerdr herdr = new FakeHerdr();
launcher(herdr, configWith(lead("opus", "lead: opus", 1))).ensureLeads();
String label = (String) ((Map<?, ?>) herdr.lastCall("workspace.create").params()).get("label");
assertEquals("leads", label);
assertNotEquals("bridged-workers", label);
}
// ── recognise-only and misconfiguration ───────────────────────────────────────────────────
/** A lead with a tab but no profile is recognise-only by design — not an error, not a launch. */
@Test
void aLeadThatNamesNoProfileIsRecognisedButNeverLaunched() {
FakeHerdr herdr = new FakeHerdr();
assertEquals(0, launcher(herdr, configWith(lead(null, "lead: dead", 1))).ensureLeads());
assertFalse(herdr.called("agent.start"));
}
/** `instances: 0` is a deliberate off switch. */
@Test
void zeroInstancesLaunchesNothing() {
FakeHerdr herdr = new FakeHerdr();
assertEquals(0, launcher(herdr, configWith(lead("opus", "lead: opus", 0))).ensureLeads());
assertFalse(herdr.called("agent.start"));
}
/** A profile name with no matching profile is logged and skipped, never a daemon crash. */
@Test
void anUnknownProfileIsSkippedRatherThanThrown() {
FakeHerdr herdr = new FakeHerdr();
assertEquals(0, launcher(herdr, configWith(lead("nope", "lead: opus", 1))).ensureLeads());
assertFalse(herdr.called("agent.start"));
}
/** No leads declared at all: not a herdr call in sight. */
@Test
void noLeadersConfiguredTouchesHerdrNotAtAll() {
FakeHerdr herdr = new FakeHerdr();
BridgedConfig cfg = new BridgedConfig(
null, null, Map.of("opus", opusProfile()), null, null, null, null, null,
null, null, null, null, "fixed", null).withDefaults();
assertEquals(0, launcher(herdr, cfg).ensureLeads());
assertTrue(herdr.calls.isEmpty(), "nothing declared ⇒ nothing scanned");
}
}

Some files were not shown because too many files have changed in this diff Show More