Compare commits

..

60 Commits

Author SHA1 Message Date
Dai Ha 74b0087ebb CB-577: model async question ownership race
CI / build (pull_request) Successful in 1m29s
CI / contract (pull_request) Failing after 0s
2026-08-15 06:24:05 +02:00
Dai Ha 927e0151d4 CB-577: test async question waiter ownership 2026-08-15 06:23:22 +02:00
Dai Ha 5275922d1d CB-577: handle questions without async waiters 2026-08-15 06:22:22 +02:00
Dai Ha e186c7945a CB-577: correlate async questions to turns 2026-08-15 06:21:53 +02:00
Dai Ha 33a6e77f0e Merge CB-573: dormant fleet health monitor (M4 unit 1)
CI / build (push) Successful in 49s
CI / contract (push) Successful in 1m14s
An opt-in whole-fleet observer, separate from the 250ms delivery poller.
One AgentControl.list and one roster snapshot per tick, joined and fed to
the FleetHealth classifier, because a fault is a disagreement between the
two views at the same instant. Absent a health: block nothing is built
and no herdr call is made.

Adds bridge_list healthCoverage: off, detection-only, or full. Detection
is deliberately separate from notification, so a single-lead setup with
no webhook still gets detection and is told its coverage is partial
rather than being refused.

Two review fixes worth naming. tick() rescheduled itself as its last
statement with no try/catch, and a ScheduledExecutorService does not
re-run a task that threw — so the first agents.list failure would have
stopped health permanently and silently, which is exactly when the
control link is down. It now catches Throwable and reschedules in a
finally. And the snapshot fields this unit cannot supply are the named
constant NOT_YET_OBSERVED rather than bare false literals, because false
means no fault to this classifier.
2026-08-15 06:19:15 +02:00
Dai Ha 4e47489d53 Merge CB-568c: fail every pending async ticket on teardown
A released target left its second async ticket pending for the full
30-minute async timeout. abandon resolved only the rendezvous waiter,
and async tickets live in a separate map that could not even represent
two tasks on one target.

abandon now sweeps every non-question async ticket for the target, and
asyncTasksByTarget holds a set. The sweep is a plain loop: the first
attempt used Stream.anyMatch, which short-circuits on the first true, so
it completed one ticket and left the rest pending — the exact bug it was
fixing. Its test passed only because the rendezvous path failed the
first ticket anyway; the test now uses three tickets so a single
completion cannot satisfy it.

An ASKING ticket is an active turn, not a pending send, so the sweep
skips it and CB-574 is unaffected.
2026-08-15 06:18:28 +02:00
Dai Ha c76b2f149e CB-573: keep health monitoring after failures
CI / contract (pull_request) Successful in 42s
CI / build (pull_request) Successful in 1m15s
2026-08-15 06:17:28 +02:00
Dai Ha 826fffe05b CB-573: add dormant fleet health monitor 2026-08-15 06:15:24 +02:00
Dai Ha 75f57cdba7 CB-568c: fail every queued async ticket
CI / contract (pull_request) Successful in 44s
CI / build (pull_request) Successful in 1m34s
2026-08-15 06:15:21 +02:00
Dai Ha f556af5d4e CB-568: fail queued async tickets on teardown 2026-08-15 06:14:28 +02:00
Dai Ha d5f33f0c6e Merge CB-568: a dropped send reports the real cause
CI / build (push) Successful in 55s
CI / contract (push) Successful in 1m0s
Injector.drop knew the precise cause (herdr agent_not_found) but the
sender was told only 'worker unreachable or stuck', so a lead could not
tell a dead pane from a stalled model.

TurnListener.onTurnFailed gains a reason, defaulting to the old one-arg
form. CompletionResolver prefers that reason, then the pane scrape, then
the old fixed text.

drop now fires onTurnFailed unconditionally. That is the substantive
fix: the sender blocks on the rendezvous waiter, not on the delivered
future, so failing delivered() alone never woke it and a queued send sat
until its timeout.
2026-08-15 06:08:41 +02:00
Dai Ha 16e17b32ad CB-568: preserve dropped turn causes
CI / contract (pull_request) Successful in 43s
CI / build (pull_request) Successful in 1m26s
2026-08-15 06:06:07 +02:00
Dai Ha 988494e18b Merge M4 fleet-health design (docs only)
The full 5-unit design behind M4: evidence model, classification
precedence, the automatic-vs-lead action boundary, worktree safety on
release, typed inbox and lead routing, capacity, and human escalation.

Two decisions worth keeping visible. Detection is split from
notification, so health works in a single-lead setup with no webhook and
reports partial coverage instead of refusing to run. And capacity stays
a view: the bridge reports free slots but never spawns, reassigns, or
stops a member to improve utilisation, because only the lead holds the
work list.

Section 13 records eleven things nobody checked, including live LavinMQ,
OpenCode pane fixtures, and multi-lead routing.
2026-08-15 06:05:26 +02:00
Dai Ha c4549a5e20 Merge CB-575: filter the routine MCP cancellation WARN
The MCP SDK 2.0.0 registers no handler for notifications/cancelled, so every
client abort logged a WARN. M4 fleet health treats WARN as action-needed, so
that noise had a cost. A Logback TurboFilter denies only that one event:
right logger, WARN level, the SDK's exact format string, and a
JSONRPCNotification whose method is notifications/cancelled. Everything else
is NEUTRAL. If a later SDK handles cancellation the filter stops matching.
2026-08-15 06:00:40 +02:00
Dai Ha 46fa4f38d5 CB-575: filter MCP cancellation warnings
CI / build (pull_request) Successful in 49s
CI / contract (pull_request) Successful in 1m1s
2026-08-15 05:57:45 +02:00
Dai Ha 20e0e68ad7 M4: align design status with CB-573
CI / contract (pull_request) Successful in 43s
CI / build (pull_request) Successful in 1m28s
2026-08-15 05:56:00 +02:00
Dai Ha 17468a234a M4: document fleet health design 2026-08-15 05:55:00 +02:00
Dai Ha 8a53d5bfc6 Merge CB-574: an async delegation can now receive a worker's question
CI / contract (push) Successful in 1m2s
CI / build (push) Failing after 1m40s
A worker on a wait:false delegation called bridge_ask and the lead never
saw the question. Outcome.QUESTION is deliberately non-terminal, but
taskView tested r.completed() and fell into the failure branch, so the
ticket was marked FAILED and both the question text and its turnId were
discarded. The worker blocked for 55s, gave up, and had to abandon its
task. CLAUDE.md tells leads to prefer wait:false and to answer an ask with
bridge_send{turnId, content}; those two could not both be followed.

bridge_poll now returns a non-terminal ASKING phase carrying the question
and its turnId, and the ticket stays live so the worker's real reply still
lands on it. An unanswered ask returns the ticket to PENDING, because only
the question wait ended - the delegated turn continues. The 55s/115s ask
caps are unchanged: they exist because the worker's own MCP call would time
out, so widening them would only move the failure.

Two defects found reviewing the first revision, both from replacing
supplyAsync with a manually completed future:

- an exception inside the send left the future uncompleted, so the ticket
  stayed PENDING for the life of the daemon. Now caught and completed
  exceptionally.
- correlation was keyed by target, one entry per worker, registered before
  the session lock. With two tickets outstanding on one target the second
  overwrote the first, so a late reply could resolve the wrong ticket.
  Correlation is now per turn, the target entry exists only while that send
  owns the lock, and a reply with no live waiter still goes to the durable
  inbox as before.
2026-08-15 05:50:45 +02:00
Dai Ha 695da7418e Merge CB-573 (part 1): health classification model and the bridge_list capacity view
CI / contract (push) Successful in 45s
CI / build (push) Successful in 55s
Fleet capacity was invisible. A finished member held a terra slot until a
spawn was refused with 'at maxLoad: 2 live >= 2 cap', and nothing had told
the lead the slot was still held. bridge_list now reports, per configured
profile, maxLoad / live / free / reclaimable, and per member idleForSeconds
and reclaimable.

live comes from the same liveCountRef function placement consumes, so the
advertised free slots cannot drift from what bridge_spawn will accept. The
profile list is the union of configured and roster profiles: an empty
configured profile still appears with its full capacity, and a member whose
profile was removed from config stays visible rather than vanishing.

reclaimable is advisory. The bridge never spawns, stops or retasks a member
to improve utilisation: it has capacity facts but no work list, and choosing
work needs authority it does not have.

Also lands the pure health classifier, its precedence chain, the MUTE counter
and the pane budget. The classifier never reports IDLE while an accepted
delivery is open — IDLE is a claim that nothing is outstanding, and the
capacity view reads exactly that field.

Capacity dependencies are one required CapacitySource rather than defaulted
constructor arguments. A defaulted liveCount would report free slots that do
not exist, which is the dangerous direction; CapacitySource.none() omits the
block instead of inventing zeros.
2026-08-15 05:49:51 +02:00
Dai Ha abd26c796b CB-574: retain async task correlation
CI / build (pull_request) Failing after 1m14s
CI / contract (pull_request) Successful in 1m15s
2026-08-15 05:49:27 +02:00
Dai Ha 24559d81ac CB-573: require explicit capacity source
CI / contract (pull_request) Successful in 44s
CI / build (pull_request) Successful in 55s
2026-08-15 05:49:04 +02:00
Dai Ha 6c1c2c3994 CB-573: report empty configured profile capacity
CI / build (pull_request) Successful in 59s
CI / contract (pull_request) Successful in 1m0s
2026-08-15 05:45:39 +02:00
Dai Ha 01fab15713 CB-573: add fleet capacity view
CI / contract (pull_request) Successful in 1m1s
CI / build (pull_request) Successful in 1m25s
2026-08-15 05:43:46 +02:00
Dai Ha 0cd00e71c3 CB-574: surface async worker questions
CI / contract (pull_request) Successful in 42s
CI / build (pull_request) Successful in 56s
2026-08-15 05:43:43 +02:00
Dai Ha bf0ff2adbf CB-573: keep active delegations out of idle
CI / contract (pull_request) Successful in 45s
CI / build (pull_request) Successful in 1m27s
2026-08-15 05:40:37 +02:00
Dai Ha ed4bbc1c56 CB-573: add pure health classification model
CI / build (pull_request) Successful in 59s
CI / contract (pull_request) Successful in 1m15s
2026-08-15 05:37:30 +02:00
Dai Ha c456402cc5 Merge CB-572: reject a configured profile name as a send target
CI / build (push) Successful in 59s
CI / contract (push) Successful in 1m16s
A lead sent to sessionId "sol" — a profile name, not a terminal id. The
bridge accepted it, handed out a ticket, failed 60s later inside the
injector, and still reported the ticket as pending 20 minutes on. The
sender never learned anything and a whole delegation was lost.

bridge_send now rejects a target that exactly matches a configured
profile name, on both the blocking and the wait:false path, before any
ticket is issued. The error names the value and points at bridge_list.

The check is deliberately narrow. A target absent from the member roster
may still be a peer lead's terminal or a herdr-owned pane, so only a
value the bridge can prove is a profile is refused. profiles is a
required parameter on both send methods — the earlier revision kept
overloads that defaulted it to an empty set, which is the same silent
disable shape as CB-561.
2026-08-15 05:35:12 +02:00
Dai Ha 4f0bf667b1 CB-572: require profiles for send validation
CI / build (pull_request) Successful in 1m0s
CI / contract (pull_request) Successful in 1m2s
2026-08-15 05:33:15 +02:00
Dai Ha 61af9aa574 CB-572: reject profile names as send targets
CI / build (pull_request) Successful in 59s
CI / contract (pull_request) Successful in 1m20s
2026-08-15 05:29:37 +02:00
Dai Ha 3b59b34e76 CB-570 follow-up: rewrap an over-long javadoc line
CI / build (push) Successful in 58s
CI / contract (push) Successful in 1m17s
2026-08-15 05:12:58 +02:00
Dai Ha 65b38997f7 Merge CB-570: OpenCode receives the composed charter (PR #40)
One file, member-charter.md, not two. Two files would have made the U5
digest non-comparable between the Claude adapter and this one, which is
the whole point of the receipt; and nobody verified how OpenCode merges
multiple instruction files, so array order was an unverified dependency.

The OPENCODE_CONFIG condition widens to include a charter. It used to be
hasMcp() || hasCustomProvider(cfg), so a profile with a role charter but
no MCP and no custom provider would have got no config file and therefore
no charter — the feature silently doing nothing for that profile.

The file stays in the per-spawn temp dir, never the worktree: the
worktree is removed on release, the parity overlay already writes into
it, and CB-525's lesson was that config the bridge copied into a worktree
made a worker operate on the wrong tree. Being outside the repo is also
what stops it being committed, which a .gitignore line does not.
2026-08-15 05:12:26 +02:00
Dai Ha 7510f7649c Merge CB-569: Claude receives the composed charter (PR #39)
argvWithBridge used to gate the charter on cfg.hasMcp(), because the only
charter was the reply rule and telling a peer to call a tool it was not
given is a bug. A role charter is identity, not a tool instruction, so
the two gates are now separate: the MCP mount still depends on mcpUrl,
while --append-system-prompt depends only on the base having composed
something. A profile with a role charter and no MCP now gets its charter.

A null charter adds no flag at all. An empty --append-system-prompt is
not the same as no system prompt.
2026-08-15 05:11:39 +02:00
Dai Ha 619792a81c CB-570: deliver composed OpenCode charter
CI / build (pull_request) Successful in 58s
CI / contract (pull_request) Successful in 1m13s
2026-08-15 05:10:49 +02:00
Dai Ha cea1183f75 CB-569: pass composed charter to Claude
CI / contract (pull_request) Successful in 43s
CI / build (pull_request) Successful in 50s
2026-08-15 05:09:30 +02:00
Dai Ha e2af4c5ae4 CB-567 follow-up: realign the changed constructor signatures
CI / contract (push) Successful in 44s
CI / build (push) Successful in 1m26s
The supplier swap left three parameter lists indented one column past
their siblings and one javadoc line broken mid-sentence. No behaviour
change.
2026-08-15 05:06:03 +02:00
Dai Ha dfd5f82894 Merge CB-567: read the charter per spawn, compose it once (PR #38)
HerdrPeerLauncher now takes Supplier<BridgedConfig.Fleet> instead of
Supplier<String> tabLabelTemplate, and reads it once per spawn. A field
taken at construction would have made the charter deferred, and deferred
looks exactly like working — which is why the test uses a mutable
supplier and spawns twice, rather than ConfigRef.fixed().

Composition happens once in the base, not in each adapter: two copies
drift while both adapter-local tests keep passing. Role charter first,
reply charter last, because the final instruction is the one that must
not be overridden. The reply charter stays gated on hasMcp() — telling a
peer to call a tool it was not given is a bug — while the role charter is
not, being identity rather than a tool instruction.

Both REPLY_CHARTER copies collapse into one, and it now says 'spawned
member' rather than 'off-subscription worker'. The old text made the
launch prompt contradict bridge_whoami for an architect; both architects
read it in their own prompts and reported it.

The old buildLaunch overloads are removed rather than kept as defaults: a
surviving one is the same shape as a stale snapshot, a route that drops
role and charter while looking healthy. That removal also let the
OpenCode adapter drop its ThreadLocal resume-id hack, since LaunchSpec
now carries the value down the same path.
2026-08-15 05:04:08 +02:00
Dai Ha e81944cef6 CB-567: compose charters per spawn
CI / build (pull_request) Successful in 57s
CI / contract (pull_request) Successful in 1m3s
2026-08-15 05:00:48 +02:00
Dai Ha 7bdd39ab9a CB-566 follow-up: say why the five-arg Fleet constructor is kept
CI / contract (push) Successful in 43s
CI / build (push) Successful in 1m8s
Reviewing the merge I read the five-arg constructor as dead code and
removed it. That was wrong: the tests call it as BridgedConfig.Fleet,
which my grep for 'new Fleet(' did not match, and the build failed on
eight call sites. It is restored with a javadoc that says why keeping it
is safe here even though an overload that drops a new field is normally
the shape to avoid — nothing reads a charter through a constructor, and
Jackson binds the canonical one, so it cannot swallow an operator's YAML.

Also drop a redundant java.util.Arrays qualifier (the class is already
imported) and rewrap a javadoc line the change had left over-long.
2026-08-15 04:54:09 +02:00
Dai Ha f4b38f040e Merge CB-566: the charter config surface (PR #37)
fleet.charters is a validated Map<String,String>, not a record: Fleet is
@JsonIgnoreProperties(ignoreUnknown = true), so a record field named
architetc would be dropped in silence and the operator would never learn
of the typo. A map lets validateCharters see the bad key and refuse it.

The key sits under fleet: because ConfigRef already treats that block as
hot and changedDeferredKeys does not list it. A new top-level key would
inherit nothing, and forgetting to classify it means a reload prints
'config reloaded' and does nothing.

A blank value is refused while an absent one is fine: an absent key means
the operator configured no charter, a blank one means they tried and
failed. Refusing at both startup and reload is the point — wiring only
one of the two paths is the whole bug.
2026-08-15 04:52:48 +02:00
Dai Ha 799668e129 CB-566: add fleet charter config
CI / contract (pull_request) Successful in 43s
CI / build (pull_request) Successful in 55s
2026-08-15 04:50:23 +02:00
Dai Ha 890190263e CB-560/562/563: the shipped block tells the truth again
The canonical CLAUDE.md block is the instruction surface this daemon ships to
every agent that mounts it, so a code change that silently invalidates it is an
incomplete change. CB-548 added a third principal kind and the block was never
revisited. Four statements in it were simply false:

  * whoami was documented as returning only primary or worker;
  * "spawn/stop/send/drain are lead-only" — Authz permits SEND to an architect;
  * delivery was documented as idle/blocked — injectable() is IDLE|BLOCKED|DONE,
    and a spawned member must also have mounted the bridge MCP, which is the
    exact condition that made every architect undeliverable for a day;
  * the tool table said bridge_list returns `workers` — the JSON key is
    `members`.

Also corrected: the fallback ladder claimed each one-way signal identifies a
"worker", but an architect gets the same charter, the same mount and the same
env, so those signals identify a spawned member and only bridge_whoami
separates the two. The safe default stays "act as a worker" — it is the most
restricted member role.

The turn contract now covers both member kinds, and says why the completion
fallback is not a substitute for bridge_reply: it returns at most the last 4000
characters, so a long report reaches the lead with its end cut off. That is not
hypothetical — it happened twice today.

Found by a reviewer asked whether the block still matches the code. Verified
against injectable(), Authz, BridgeMcp.listFleet and LeadLauncher before
applying. The wiki template is updated in the same shape and re-checked
byte-identical.
2026-08-15 04:37:03 +02:00
Dai Ha f8522edacd Merge CB-564: give every silent failure path a voice
CI / build (push) Successful in 1m0s
CI / contract (push) Successful in 1m15s
The bridge cannot classify what it never emits. Fleet-health monitoring —
detect a wedged member, decide, escalate — is blocked on that, so this is the
foundation rather than the feature.

A survey of Injector, CompletionResolver, StatusPoller, SessionManager,
SessionReaper and MessageService found six conditions that ended a member's
usefulness while saying nothing useful:

  * Injector.drop()             — worker gone, queue cleared: SILENT
  * CompletionResolver.fail()   — "via turn-stall fallback" at DEBUG, no reason
  * SessionManager.onFailed()   — "session marked failed" at DEBUG, no stage
  * SessionManager.reapIdle()   — indistinguishable from any other release
  * SessionManager.acquire()    — spawn failure rethrown with no log at all
  * MessageService.abandon()    — failed a caller's request at DEBUG

The first two are the exact phrases that misled the CB-560 diagnosis: both name
a symptom and neither names a cause. They are now WARN and carry the reason,
the stage, and the counts.

Conditions already loud were left alone, and so were two by-design timeouts in
MessageService — an async model exists precisely for those, and promoting them
would turn healthy operation into noise.

Behaviour is unchanged: every edit is a log statement.

Merge note: the recycle test deleted by CB-565 conflicted with a test added
here. Resolved by keeping the new onTurnFailed assertion and dropping the
recycle test, which tests a method that no longer exists.
2026-08-15 04:36:09 +02:00
Dai Ha b958747855 Merge CB-565: remove the recycle trap (PR #35)
CI / contract (push) Successful in 58s
CI / build (push) Successful in 1m27s
SessionManager.recycle called the 4-argument acquire overload, which defaults
the role to DEV and requests no worktree. So it carried profile, cwd and owner
across and silently dropped two things: the member's role, and its worktree.

A recycled architect would have come back a plain worker, never rebound to its
slot, with nothing logged. Worse, a recycled member would have come back with
no worktree at all — and a member's uncommitted work exists in exactly one
place. Role loss is recoverable; that is not.

Nothing called it. The only reference outside its own javadoc was one test, and
the context-cap path calls release, not recycle. So this was a trap waiting for
its first caller, and that caller would not have noticed either loss.

Deleted rather than repaired, on the CB-561 precedent: an API that looks correct
and silently drops a property is worse than no API. The no-reuse invariant it
documented is still true and is now stated directly.

Found by a reviewer asked to hunt the rest of the CB-548 fallout. The worktree
half was found while verifying the report.
2026-08-15 04:31:46 +02:00
Dai Ha 078bde2c02 CB-564: give a voice to silent member-failure paths
Injector.drop, CompletionResolver.fail, SessionManager.onFailed/reapIdle/
acquire spawn failures, and MessageService.abandon used to fail a member
or a caller's request with no log, a bare DEBUG, or a log that named only
the symptom ("session marked failed", "failed send via turn-stall
fallback"). Each now logs at WARN and names the real cause and the
numbers involved. Observability only — no behaviour changed.
2026-08-15 04:31:42 +02:00
Dai Ha 0b10ea987b CB-565: remove unsafe session recycle
CI / contract (pull_request) Successful in 44s
CI / build (pull_request) Successful in 55s
2026-08-15 04:29:50 +02:00
Dai Ha 1553d38182 Merge CB-563: a clipped pane scrape says it was clipped (PR #34)
CI / contract (push) Successful in 45s
CI / build (push) Successful in 1m12s
When a member ends its turn without bridge_reply, CompletionResolver scrapes
the pane and resolves the waiting send with that text. The scrape is capped at
MAX_SCRAPE_CHARS (4000), and nothing told the caller when the cap had bitten.
A delegating lead could act on a report missing its end and believe it was
complete. That happened to me today: a member's full engineering report arrived
cut at exactly 4000 characters, and the only hint was a DEBUG line reading
"(4000 chars scraped)", which reads like a size and not like a warning.

The returned text now carries a marker when, and only when, it was clipped, and
the clip is logged at WARN with the original length and the cap.

The cap itself is unchanged. The problem was silence, not the number.

The CB-115 misattribution guard still compares the unmarked clipped tail to the
unmarked baseline, and the marker is appended only afterwards. Verified in the
code, not taken on report: clip() strips before truncating and the new length
check uses the same stripped length, so there is no off-by-one either.
2026-08-15 04:26:57 +02:00
Dai Ha d04b075996 CB-563: mark clipped completion scrapes
CI / contract (pull_request) Successful in 42s
CI / build (pull_request) Successful in 56s
2026-08-15 04:25:54 +02:00
Dai Ha 644927636d Merge CB-562: the readiness gate says why it gave up (PR #33)
CI / build (push) Successful in 1m1s
CI / contract (push) Successful in 1m14s
When the injector's readiness grace expired it cleared the queue and logged
nothing. The failure then surfaced elsewhere as a turn-stall, which names the
wrong cause. Diagnosing CB-560 cost two live spawns and a wrong first
hypothesis for exactly this reason: the logs said "session marked failed" and
"failed send via turn-stall fallback", and neither says the message was never
typed into the pane at all.

The expiry now logs the target, the number of messages being failed, the grace
in polls and seconds, and the real cause in plain words.

The grace in seconds is derived, not written down twice: Bridged's own
INJECT_POLL_MILLIS is deleted and Injector.POLL_INTERVAL_MILLIS is the single
source, passed to every StatusPoller. A cadence change can no longer leave a
log line confidently stating the wrong duration.

Behaviour is unchanged. This is the first structured health event in the
daemon, and the foundation the fleet-health work will build on.
2026-08-15 04:21:29 +02:00
Dai Ha 95e45007aa CB-562: single source for the injector poll cadence; tighten count assertion
CI / build (pull_request) Successful in 59s
CI / contract (pull_request) Successful in 1m17s
2026-08-15 04:20:29 +02:00
Dai Ha 7a583c4045 CB-562: log why the readiness gate gave up on a target
CI / build (pull_request) Successful in 57s
CI / contract (pull_request) Successful in 1m4s
2026-08-14 21:52:25 +02:00
Dai Ha 65bce058bb Merge CB-561: one public way to build a CallerResolver (PR #32)
CI / build (push) Successful in 1m17s
CI / contract (push) Successful in 1m16s
CallerResolver had 5 public constructors and 4 public factories, and only one
of them could ever produce an architect. The rest defaulted memberSlotRoles to
`_ -> null`, so every architect quietly fell through to Principal.worker().
Nothing logged, nothing threw — the role was simply off.

That is the same failure shape as CB-560, so the fix is structural rather than
a warning: withLeadsAndMembers(.., MemberRegistry) is now the only public
construction path. Two overloads with no caller at all are deleted; the rest
are package-private and marked test-only. No path remains that accepts
architect bindings without a slot-role lookup, so no runtime WARN is needed.

Also checked and closed: the suspected slot leak on shutdown drain is not
real. SessionManager.release calls memberLifecycle.released() for every
removed session, and drainAll routes every session through release,
SPAWNING included. Verified in code.
2026-08-14 21:50:04 +02:00
Dai Ha 48b437083b Merge CB-560: mark spawned members present (PR #31)
Merging CB-548 made an architect resolve as Role.ARCHITECT, and BridgeMcp
marked MCP presence only for Role.WORKER. So an architect was never present.
That one line broke two things, because SessionManager.asPresence() is the
same object the injector's readiness gate reads:

  * the session never left SPAWNING;
  * Bridged.deliverableTo was false, so the injector held every delivery,
    waited out READINESS_GRACE_POLLS (~60s) and failed the send.

Confirmed live twice, on claude-code and on opencode. Both logged
"session marked failed" then "failed send via turn-stall fallback", neither
of which names the real cause.

Principal.isSpawnedMember() now means "a spawned member with a pane" —
WORKER or ARCHITECT, never PRIMARY. The Bridged.deliverableTo javadoc, which
still stated the worker-only rule as fact, is corrected: it is the clearest
description of this invariant anywhere, and leaving it stale is how the bug
comes back.
2026-08-14 21:50:04 +02:00
Dai Ha fa0612859b CB-560: document spawned member presence
CI / build (pull_request) Successful in 1m0s
CI / contract (pull_request) Successful in 1m4s
2026-08-14 21:48:58 +02:00
Dai Ha 4ffbcd0b7d CB-561: require member registry for architects
CI / contract (pull_request) Successful in 43s
CI / build (pull_request) Successful in 50s
2026-08-14 21:47:53 +02:00
Dai Ha a0cd053fd9 CB-560: mark architect members present
CI / build (pull_request) Successful in 55s
CI / contract (pull_request) Successful in 1m4s
2026-08-14 21:45:28 +02:00
Dai Ha f9a5e066b5 Merge PR #30: bind a spawned architect to its slot (CB-548's missing half)
CI / contract (push) Successful in 48s
CI / build (push) Successful in 1m18s
MemberRegistry.bind() had 18 call sites, every one in its own unit test.
Nothing in production ever bound a terminal to a slot, so CallerResolver's
architect branch was unreachable, every spawned architect resolved as WORKER,
and Authz's architect row was dead code. Since SEND is granted to
isPrimary() || isArchitect(), no architect could message anyone.

The spawn lifecycle now binds on both acquire paths and unbinds on release,
through a MemberLifecycle seam with a no-op default, so all six SessionManager
constructors are unchanged.

The escalation guard is deliberately double-sided. MemberRegistry flattens
every pool, so dev:* and reviewer:* slots sit in the same map the architect
resolver reads. The lifecycle binds only ARCHITECT roles, AND CallerResolver
independently checks roleForSlot(slot) == ARCHITECT before granting. Either
alone would be enough today; together, a regression in one cannot escalate a
worker. A reviewer confirmed a third barrier already existed: Authz gates
SPAWN on isPrimary(), so only a lead can request role=ARCHITECT at all.

Verified by the primary: mvn clean install, 640 tests, 0 failures.

Two findings recorded, neither blocking:

1. Known race, low severity. registry.put() and memberLifecycle.acquired()
   are not atomic. A release landing between them unbinds nothing (no binding
   exists yet), then acquire binds a dead terminal that no later release will
   ever clear - the slot leaks. Only the shutdown drain can reach a SPAWNING
   session (the idle reaper skips it, and an explicit stop needs a paneId the
   spawn has not returned), so the leak dies with the process. Note that
   binding before put does NOT fix it: the drain then iterates a roster the
   session is not in yet. A real fix needs atomicity.

2. The legacy Supplier-based withLeadsAndMembers overload and the Map-form
   constructors now silently disable architect resolution: memberSlotRoles
   defaults to _ -> null, so the architect branch can never be taken there.
   Production uses the registry form, and the tests were migrated, so nothing
   fails - but a future caller gets workers with no error.
2026-08-14 21:25:45 +02:00
Dai Ha e9bc192160 CB-548: bind architect sessions to member slots
CI / contract (pull_request) Successful in 45s
CI / build (pull_request) Successful in 55s
2026-08-14 21:10:48 +02:00
Dai Ha e0a57988ad Merge PR #29: drop .claude/settings.local.json from the default parityOverlay
CI / build (push) Successful in 1m21s
CI / contract (push) Successful in 1m21s
A member no longer receives the primary's pre-approved tool grants by default.
The grants were inert — GitWorktrees.isolateToolSurface already strips each
worktree's .mcp.json to an empty server map — so this is defence in depth: two
independent guards instead of one. Same reasoning as CB-525, which removed the
sibling .mcp.json and left this file behind.

An explicit parityOverlay: in config is unaffected; only the default changes.

Verified by the primary: mvn clean install, 637 tests, 0 failures.
2026-08-14 21:02:25 +02:00
Dai Ha b414a74c26 CB-559: drop .claude/settings.local.json from the default parity overlay
CI / build (pull_request) Successful in 54s
CI / contract (pull_request) Successful in 1m7s
2026-08-14 20:56:16 +02:00
Dai Ha 2c3796d598 Merge: keep every credential in one store
CI / build (push) Successful in 55s
CI / contract (push) Successful in 1m13s
2026-08-14 20:38:36 +02:00
50 changed files with 3167 additions and 441 deletions
+39 -31
View File
@@ -11,41 +11,45 @@
If no `bridge_*` MCP tools are mounted in this session, this section does not apply — skip it.
`bridged` is the **sole communication gateway** between agents here. The orchestrating session (the
**primary**) and every delegated peer (a **worker**) mount the *same* MCP server and talk only
**primary**) and every delegated peer (a **member**) mount the *same* MCP server and talk only
through its `bridge_*` tools. No session addresses a peer, a broker, or the network directly.
### Which role am I? — settle this before acting
**Both roles read this file.** A worker runs in a git worktree of this same repo, so it inherits
**Every role reads this file.** A member runs in a git worktree of this same repo, so it inherits
this `CLAUDE.md` verbatim, and every rule below is role-conditional.
**Call `bridge_whoami`.** It returns `{"role":"primary"}` or `{"role":"worker","sessionId":…,
"profile":…,"worktree":…,"branch":…}`, resolved by the daemon from your connection — unforgeable,
and the same resolution its authorization gate uses. Don't infer what you can ask.
**Call `bridge_whoami`.** It returns `primary`, `worker`, or `architect`, resolved by the daemon from
your connection — unforgeable, and the same resolution its authorization gate uses. A worker also
carries its `sessionId`, `profile`, `worktree` and `branch`; an architect carries the slot name it
was bound to. Don't infer what you can ask.
Only if that call is unavailable, fall back to these — each is one-way, so keep reading until one
fires: the reply charter in your system prompt (*"You are an off-subscription worker in the
claude-bridge fleet"*) ⇒ **worker**; bridge tools prefixed `mcp__bridge__*` ⇒ **worker** (the
launcher fixes that mount name; a primary's mount is named by whoever wrote its `.mcp.json`, so it
varies); `ANTHROPIC_BASE_URL` set ⇒ **worker** (Claude-model workers run on a clean env, so its
*absence* proves nothing). **Still unsure ⇒ act as a worker.** The two mistakes are not symmetric: a
primary acting as a worker is refused by the authorization gate — loud and self-correcting — while a
worker acting as the primary ends its turn with no `bridge_reply`, and the sender silently receives
nothing. Fail toward the recoverable error.
claude-bridge fleet"*) ⇒ **spawned member**; bridge tools prefixed `mcp__bridge__*` ⇒ **spawned
member** (the launcher fixes that mount name; a primary's mount is named by whoever wrote its
`.mcp.json`, so it varies); `ANTHROPIC_BASE_URL` set ⇒ **spawned member** (Claude-model members run
on a clean env, so its *absence* proves nothing). None of these separate a worker from an architect —
only `bridge_whoami` does. **Still unsure ⇒ act as a worker**, the most restricted member role. The
two mistakes are not symmetric: a primary acting as a worker is refused by the authorization gate —
loud and self-correcting — while a member acting as the primary ends its turn with no `bridge_reply`,
and the sender silently receives nothing. Fail toward the recoverable error.
### Invariants — both roles, no exceptions
1. **Never set, export, or forward `ANTHROPIC_BASE_URL`** (or `ANTHROPIC_AUTH_TOKEN`). The primary
stays on subscription; only the bridge puts a worker off it, at spawn. Mounting the bridge must
stays on subscription; only the bridge puts a member off it, at spawn. Mounting the bridge must
never move a session across that boundary.
2. **The bridge is the only channel.** Text you print in your terminal reaches nobody — the other
side cannot see your screen. An answer that isn't in a `bridge_*` call is silently discarded.
3. **Identity comes from the connection, never an argument.** Workers never pass a target; you
cannot act as another session. Spawn/stop/send/drain are lead-only; reply/ask are
only-as-itself — any peer may answer for its own pane, and for no other. A call outside your
role is refused, not queued.
cannot act as another session. Spawn/stop/drain are lead-only; **send is lead or architect**;
reply/ask are only-as-itself — any peer may answer for its own pane, and for no other. A call
outside your role is refused, not queued.
4. **Delivery is status-gated: one message per turn.** Don't busy-poll a peer's terminal and don't
re-send because a call looks slow — the bridge delivers when the peer is `idle`/`blocked`.
re-send because a call looks slow — the bridge delivers when the peer is `idle`, `blocked` or
`done`. A spawned member must **also** have mounted the bridge MCP: until it has, it is not
deliverable, and a send waits on that gate for ~60s and then fails without ever reaching its pane.
5. **Never drive the terminal multiplexer directly** (no `herdr` CLI, no socket). The bridge owns
policy; the multiplexer owns PTYs. Going around the bridge bypasses every rule above.
@@ -101,26 +105,26 @@ the merge — and merging on a reviewer's word is delegating it by proxy.
|---|---|
| Confirm your own role | `bridge_whoami` |
| See backends available | `bridge_profiles` |
| Start a worker | `bridge_spawn{profile?, cwd?, worktree?, ticket?}` → `sessionId` + `paneId` |
| See the fleet | `bridge_list` → `leads` (your peers) + `workers` · one peer's state: `bridge_status{sessionId}` |
| Start a member | `bridge_spawn{role?, profile?, cwd?, worktree?, ticket?}` → `sessionId` + `paneId` |
| See the fleet | `bridge_list` → `leads` (your peers) + `members` · one peer's state: `bridge_status{sessionId}` |
| Delegate (blocking) | `bridge_send{sessionId, content}` |
| Delegate (long task) | `bridge_send{sessionId, content, wait:false}` → ticket → `bridge_poll{ticket}` |
| Answer a worker's `bridge_ask` | `bridge_send{turnId, content}` — **not** `sessionId` |
| Answer a member's `bridge_ask` | `bridge_send{turnId, content}` — **not** `sessionId` |
| Message a **peer lead** | `bridge_send{sessionId: <their terminal>, content}` — `bridge_list` → `leads` reports it. Coordination only, **never** a task |
| Answer a peer lead that messaged you | `bridge_reply{content}` — the one case a lead replies |
| Collect a held reply | `bridge_poll{target}` · then `bridge_ack{target, msgId}` |
| Tear down | `bridge_stop{paneId}` |
| Tear down a member | `bridge_stop{paneId}` |
### Lead ↔ lead — coordinate, never delegate
`bridge_list` returns `leads` alongside `workers`; your own row carries `self: true`. Every other row
is a peer — an orchestrator with its own context, its own workers, and its own judgment. An empty
`workers` array means no workers are spawned; it says nothing about peers.
`bridge_list` returns `leads` alongside `members`; your own row carries `self: true`. Every other row
is a peer — an orchestrator with its own context, its own members, and its own judgment. An empty
`members` array means no members are spawned; it says nothing about peers.
**A lead never assigns a task to another lead.** Work goes to workers — only ever downward, never
sideways. Sending a peer a brief with acceptance criteria is a category error: a brief is a worker's
**A lead never assigns a task to another lead.** Work goes to members — only ever downward, never
sideways. Sending a peer a brief with acceptance criteria is a category error: a brief is a member's
artefact, and a peer is not yours to task. If a unit needs doing and it falls in your area, spawn a
worker and delegate it yourself; if it falls in the peer's area, say so and let the peer assign it.
member and delegate it yourself; if it falls in the peer's area, say so and let the peer assign it.
The traffic between leads is coordination and nothing else:
1. **Divide the map, not the work.** Agree who owns which area, then each of you assigns inside your
@@ -138,7 +142,7 @@ Being messaged by a peer does not make you its worker: answer with `bridge_reply
the substance if it is wrong. A peer that simply complies has thrown away the reason there are two of
you.
### Worker — the turn contract
### Member (worker or architect) — the turn contract
1. **Load the playbook skill the lead named** before doing anything else.
2. **Do the assigned scope only.** Note anything you spot outside it in one line; don't go hunt it.
@@ -147,6 +151,10 @@ you.
Don't ask what you could decide yourself.
4. **End the turn with exactly one `bridge_reply{content}`**, carrying your complete answer. This is
the whole handoff. No `bridge_reply` ⇒ the sender gets nothing and the exchange stalls.
Do **not** lean on the completion fallback to carry your answer for you: when you end a turn
without replying, the bridge scrapes your pane, and it can return only the last 4000 characters.
A clipped scrape is marked as partial, but the missing text is gone — your report reaches the
lead with its end cut off.
5. **Report honestly.** State only what you actually ran and its real output, including failures.
You mount **only** the bridge MCP — the primary's other servers (IDE, forge, docs) are not yours,
so never claim the result of a check you had no way to run.
@@ -157,9 +165,9 @@ you.
| Layer | Scope | Reaches |
|---|---|---|
| the launcher's reply charter | the one rule that must survive with no repo: *end every turn with `bridge_reply`* | every worker, at launch, every peer kind |
| **this section** | protocol + orchestration policy | primary **and** every Claude worker — tracked in git, so worktrees inherit it |
| role playbook skills | per-job procedure (commit/PR recipe, finding format) | a worker told to load one |
| the launcher's reply charter | the one rule that must survive with no repo: *end every turn with `bridge_reply`* | every spawned member, at launch, every peer kind — never a lead |
| **this section** | protocol + orchestration policy | primary **and** every member that reads the repo — tracked in git, so worktrees inherit it |
| role playbook skills | per-job procedure (commit/PR recipe, finding format) | a member told to load one |
| the bridge's own docs | design detail, flows, error model | on demand |
A rule belongs in **exactly one** layer — the outermost one that must obey it. Peers that don't read
+30 -2
View File
@@ -85,6 +85,16 @@ bind:
# backoffMs: 60000
# quietNudgeCap: 3
# Fleet health detection is dormant unless enabled. It reads one whole-fleet agent list per tick.
# It can run without a webhook; bridge_list then reports healthCoverage: detection-only.
# health:
# enabled: true
# intervalSeconds: 30 # minimum 15
# workingSuspectAfterSeconds: 600 # minimum 300
# paneProbeIntervalSeconds: 60 # minimum 60
# notifications:
# mode: disabled # disabled (default) or webhook
# herdr Unix socket. Omit to use the client default
# (${HERDR_SOCKET_PATH:-~/.config/herdr/herdr.sock}).
herdrSocket: ~/.config/herdr/herdr.sock
@@ -231,8 +241,8 @@ placement: weighted
#
# Not every key can move under a running daemon, and the difference is about what already exists
# when the reload happens — not about how important the key is:
# HOT → takes effect on the next spawn: the whole `fleet:` block (every role pool and
# `tabLabel`), `placement:`, and an existing profile's weight / maxLoad. Those are
# HOT → takes effect on the next spawn: the whole `fleet:` block (every role pool,
# `charters`, and `tabLabel`), `placement:`, and an existing profile's weight / maxLoad. Those are
# hot because the placement policy reads them through a supplier — being config is
# not by itself enough to make a key hot.
# DEFERRED → accepted into the new config, but the wiring built at startup keeps the old value
@@ -274,6 +284,24 @@ placement: weighted
# the candidates, in definition order. A dev and a reviewer staying anonymous is exactly compatible
# with being listed here; the entry key just names the entry.
fleet:
# Optional launch-charter text, keyed only by the singular role wire names: architect, dev,
# reviewer. Changes are HOT and reach the next spawn without a daemon restart. Do not put secrets
# here: a later launch step writes this text to a world-readable temp file, and ${ENV} interpolation
# is deliberately not supported.
charters:
architect: |-
You are an architect in this fleet. You refine work before anyone builds it:
scope, acceptance criteria, risks, and a unit split. You read the repo and
write analysis. You never commit production code and never open a PR.
A design task is worked by two architects. Design alone first, then exchange
and say plainly where you disagree. Do not concede just to agree.
dev: |-
You implement the one unit you were given, and nothing else. You test it,
commit it, and open your own pull request. You never merge.
reviewer: |-
You review the diff you were given. You report bugs, risks and missing tests.
You do not change code.
# Optional. Template for a member tab's label; {role}, {profile}, {model} and {n} are substituted.
# {n} counts per role+profile, so `dev: sonnet #2` really is the second sonnet dev. Because {role}
# comes from a closed enum, a generated label can never begin with a lead's tabPrefix.
@@ -23,6 +23,7 @@ import dev.ltms.bridged.mcp.BridgeMcp;
import dev.ltms.bridged.mcp.ConnectionIdentity;
import dev.ltms.bridged.metrics.BridgedMetrics;
import dev.ltms.bridged.metrics.Metrics;
import dev.ltms.bridged.health.FleetHealthMonitor;
import dev.ltms.bridged.mcp.PrimaryRegistry;
import dev.ltms.bridged.mcp.LsofPeerPidLookup;
import dev.ltms.bridged.mcp.LsofProcessCwdLookup;
@@ -70,9 +71,6 @@ public final class Bridged {
private static final Logger log = LoggerFactory.getLogger(Bridged.class);
/** How often the injector samples a busy worker's status while it has queued work. */
private static final long INJECT_POLL_MILLIS = 250;
/** CB-504: how long to wait at startup for herdr's socket before serving degraded. */
private static final long HERDR_WAIT_SECONDS = 30;
private static final long HERDR_WAIT_POLL_MILLIS = 500;
@@ -99,6 +97,7 @@ public final class Bridged {
// CB-542: a subscription:true profile whose env: reseats ANTHROPIC_BASE_URL/AUTH_TOKEN would
// reach an unguarded endpoint (the launcher skips SubscriptionGuard for it). Refuse at load.
cfg.validateSubscriptionProfiles();
cfg.validateCharters();
// CB-548: every architect slot must name a configured workers: profile — the strong-model
// backend the future spawn lifecycle would read. A stale reference dies here, not later.
cfg.validateMembers();
@@ -131,13 +130,13 @@ public final class Bridged {
adapters.add(new ClaudeCodeLauncher(agents, spaces, guard,
claudeProfiles, cfg.effectiveDefaultProfile(), System::getenv,
cfg.spawnReadyTimeoutMs(), cfg.spawnReadyPollMs(),
() -> config.get().fleet().tabLabel()));
() -> config.get().fleet()));
}
if (!opencodeProfiles.isEmpty()) {
adapters.add(new OpenCodeLauncher(agents, spaces,
opencodeProfiles, cfg.effectiveDefaultProfile(), System::getenv,
cfg.spawnReadyTimeoutMs(), cfg.spawnReadyPollMs(),
() -> config.get().fleet().tabLabel()));
() -> config.get().fleet()));
}
AtomicReference<Function<String, Integer>> liveCountRef = new AtomicReference<>(_ -> 0);
PeerLauncher workers = new CompositePeerLauncher(
@@ -243,6 +242,7 @@ public final class Bridged {
// what CallerResolver resolves against and what that lifecycle will read profiles from;
// nothing here spawns a slot.
MemberRegistry members = new MemberRegistry(cfg.fleet());
sessions.setMemberLifecycle(members);
if (!members.slots().isEmpty()) {
log.info("member slots: {} configured {} — none bound yet (a slot is idle until the "
+ "spawn lifecycle binds a live terminal to it)",
@@ -286,10 +286,16 @@ public final class Bridged {
completion.onTurnFailed(target);
sessions.onTurnFailed(target);
}
@Override
public void onTurnFailed(String target, String reason) {
completion.onTurnFailed(target, reason);
sessions.onTurnFailed(target);
}
};
Injector injector = new Injector(agents, turnListener, deliverableTo(presence, leads),
presence::forget);
StatusPoller poller = new StatusPoller(agents, injector, INJECT_POLL_MILLIS);
StatusPoller poller = new StatusPoller(agents, injector, Injector.POLL_INTERVAL_MILLIS);
poller.start();
// CB-307: reply inbox. A broker: block (with a uri) selects the AMQP-backed durable adapter;
@@ -349,6 +355,26 @@ public final class Bridged {
MessageService messages = new MessageService(agents, injector, rendezvous, replyInbox,
pushLoop, metrics);
// Health is a slow whole-fleet observer. Keep it separate from the 250ms delivery poller.
final FleetHealthMonitor healthMonitor;
var healthScheduler = Executors.newSingleThreadScheduledExecutor(r ->
Thread.ofVirtual().name("bridge-health-").unstarted(r));
if (cfg.health() != null && cfg.health().isEnabled()) {
healthMonitor = new FleetHealthMonitor(agents, sessions::roster, messages, healthScheduler,
System::nanoTime, cfg.health().intervalOrDefault());
String coverage = FleetHealthMonitor.coverage(true,
cfg.health().notifications() != null && cfg.health().notifications().configured());
if ("detection-only".equals(coverage)) {
log.warn("fleet health: {} (no notification sink configured)", coverage);
} else {
log.info("fleet health: {}", coverage);
}
healthMonitor.start();
} else {
healthMonitor = null;
healthScheduler.shutdownNow();
}
// CB-520: the reply inbox only consumes for agents this gateway owns. own on acquire,
// release on teardown. Do this before CB-516 so the inbox is owned before any reply can land.
sessions.onAcquire(replyInbox::own);
@@ -375,18 +401,25 @@ public final class Bridged {
throw new IllegalStateException("auth.mode=token but env var " + cfg.auth().tokenEnv()
+ " is unset or empty — export it before starting bridged");
}
callers = CallerResolver.withLeadsAndMembers(identity, true, token, leads,
members::snapshot);
callers = CallerResolver.withLeadsAndMembers(identity, true, token, leads, members);
log.info("auth: token mode (bearer required for non-worker callers, env {})",
cfg.auth().tokenEnv());
} else {
callers = CallerResolver.withLeadsAndMembers(identity, false, null, leads,
members::snapshot);
callers = CallerResolver.withLeadsAndMembers(identity, false, null, leads, members);
log.info("auth: loopback-trust (any loopback non-worker caller is the primary)");
}
BridgeMcp mcp = new BridgeMcp(messages, workers, sessions, identity, presence,
primaryRegistry, callers, metrics);
primaryRegistry, callers, metrics, new BridgeMcp.CapacitySource(profile -> liveCountRef.get().apply(profile),
profile -> {
var configured = config.get().profiles().get(profile);
return configured == null ? null : configured.maxLoad();
}, () -> config.get().profiles().keySet(), System::nanoTime),
new BridgeMcp.HealthCoverageSource(() -> {
var health = config.get().health();
return FleetHealthMonitor.coverage(health != null && health.isEnabled(),
health != null && health.notifications() != null && health.notifications().configured());
}));
// CB-559: opt-in config reload. With no `configReload:` block nothing is constructed, so an
// upgraded daemon behaves exactly as before — the file is read once at boot and never again.
@@ -407,6 +440,7 @@ public final class Bridged {
messages.close();
pushLoop.close();
if (heartbeat != null) heartbeat.close(); // CB-551: stop the idle-lead heartbeat scheduler
if (healthMonitor != null) healthMonitor.stop();
if (configWatcher != null) configWatcher.stop(); // CB-559: stop polling the config file
mcp.close();
if (reaper != null) reaper.stop();
@@ -429,19 +463,20 @@ public final class Bridged {
}
/**
* The {@link Injector}'s readiness gate (CB-534): a target is deliverable if it is a worker whose
* agent has connected the bridge MCP, <em>or</em> a lead.
* The {@link Injector}'s readiness gate (CB-534): a target is deliverable if it is a spawned
* member whose agent has connected the bridge MCP, <em>or</em> a lead.
*
* <p>The gate exists for one reason — to hold a delivery out of a <em>spawned</em> worker's boot
* <p>The gate exists for one reason — to hold a delivery out of a <em>spawned</em> member's boot
* window, where herdr already reports {@code idle} but the TUI would drop an injected paste. That
* hazard is a property of spawning. A lead is never spawned: the operator started it and named it
* (or labelled its tab) only once it was up, so there is no boot window to guard.
*
* <p>A lead is also never enrolled in {@link MemberPresence} — {@code BridgeMcp} marks presence
* only for a worker, deliberately, since that map doubles as the worker roster's availability
* signal and a lead counted there would show up as an available worker. So without the second
* disjunct a lead is permanently un-deliverable: every lead→lead send sat on the gate for
* {@code READINESS_GRACE_POLLS} (~60s) and then failed having never been typed into the pane.
* for every spawned member (worker and architect), deliberately, since that map doubles as the
* member roster's availability signal and a lead counted there would show up as an available
* member. So without the second disjunct a lead is permanently un-deliverable: every
* lead→lead send sat on the gate for {@code READINESS_GRACE_POLLS} (~60s) and then failed
* having never been typed into the pane.
*
* <p>The lead set is read through the supplier on each call rather than snapshotted, so a lead
* discovered by {@code leadScan} after startup becomes deliverable without a restart.
@@ -1,10 +1,12 @@
package dev.ltms.bridged.auth;
import dev.ltms.bridged.mcp.ConnectionIdentity;
import dev.ltms.bridged.peer.MemberRole;
import java.nio.charset.StandardCharsets;
import java.security.MessageDigest;
import java.util.Map;
import java.util.function.Function;
import java.util.function.Supplier;
/**
@@ -60,14 +62,16 @@ public final class CallerResolver {
* in {@code Bridged} reads a constant from config, which is the degenerate live case.
*/
private final Supplier<Map<String, String>> architectTerminals;
private final Function<String, MemberRole> memberSlotRoles;
private final Function<String, String> memberSlotNames;
/** Loopback-trust resolver: no token required, historical behaviour. */
public CallerResolver(ConnectionIdentity identity) {
/** Loopback-trust resolver: no token required, historical behaviour. Test-only. */
CallerResolver(ConnectionIdentity identity) {
this(identity, false, null, Map.of());
}
/** As {@link #CallerResolver(ConnectionIdentity, boolean, String, Map)} with no leads pinned. */
public CallerResolver(ConnectionIdentity identity, boolean tokenMode, String token) {
/** As {@link #CallerResolver(ConnectionIdentity, boolean, String, Map)} with no leads pinned. Test-only. */
CallerResolver(ConnectionIdentity identity, boolean tokenMode, String token) {
this(identity, tokenMode, token, Map.of());
}
@@ -82,8 +86,8 @@ public final class CallerResolver {
* @param pinnedPrimaryTerminal the primary's own herdr {@code terminal_id}
* ({@code null}/blank = unpinned)
*/
public static CallerResolver pinnedTo(ConnectionIdentity identity, boolean tokenMode,
String token, String pinnedPrimaryTerminal) {
static CallerResolver pinnedTo(ConnectionIdentity identity, boolean tokenMode,
String token, String pinnedPrimaryTerminal) {
return new CallerResolver(identity, tokenMode, token,
pinnedPrimaryTerminal == null || pinnedPrimaryTerminal.isBlank()
? Map.of() : Map.of(pinnedPrimaryTerminal, "primary"));
@@ -98,21 +102,11 @@ public final class CallerResolver {
* {@link Role#PRIMARY} — rather than a worker. Empty = nothing pinned,
* so every pane resolves as a worker.
*/
public CallerResolver(ConnectionIdentity identity, boolean tokenMode, String token,
Map<String, String> leadTerminals) {
CallerResolver(ConnectionIdentity identity, boolean tokenMode, String token,
Map<String, String> leadTerminals) {
this(identity, tokenMode, token, fixed(leadTerminals), null);
}
/**
* Map-form of both registries (CB-548): lead terminals and the initial architect terminal
* bindings, each snapshotted at construction (a handed-over map is not offered as live state).
*/
public CallerResolver(ConnectionIdentity identity, boolean tokenMode, String token,
Map<String, String> leadTerminals,
Map<String, String> architectTerminals) {
this(identity, tokenMode, token, fixed(leadTerminals), fixed(architectTerminals));
}
/**
* Live-registry form: {@code leadTerminals} is consulted on every resolve, so leads discovered
* after startup (CB-531's tab scan) take effect without a restart.
@@ -121,26 +115,26 @@ public final class CallerResolver {
* {@link #pinnedTo}: {@code Map} and {@code Supplier} overloads are ambiguous for a literal
* {@code null}.
*/
public static CallerResolver withLeads(ConnectionIdentity identity, boolean tokenMode,
String token,
Supplier<Map<String, String>> leadTerminals) {
static CallerResolver withLeads(ConnectionIdentity identity, boolean tokenMode,
String token,
Supplier<Map<String, String>> leadTerminals) {
return new CallerResolver(identity, tokenMode, token, leadTerminals, null);
}
/**
* Live-registry form for both {@code leadTerminals} and the CB-548 architect registry: both
* are consulted on every resolve, so a slot binding injected after startup takes effect
* without a restart.
* Live registry form that can confirm a bound slot is an architect slot.
*
* <p>A static factory rather than a constructor overload, for the same reason as
* {@link #pinnedTo}: too many {@code Map}/{@code Supplier} combinations to make {@code null}
* unambiguous.
* <p>This is the only public construction path. It keeps terminal bindings and slot roles in
* the same {@link MemberRegistry}, so a configured architect can resolve as an architect.
*/
public static CallerResolver withLeadsAndMembers(ConnectionIdentity identity,
boolean tokenMode, String token,
Supplier<Map<String, String>> leadTerminals,
Supplier<Map<String, String>> architectTerminals) {
return new CallerResolver(identity, tokenMode, token, leadTerminals, architectTerminals);
boolean tokenMode, String token,
Supplier<Map<String, String>> leadTerminals,
MemberRegistry members) {
return new CallerResolver(identity, tokenMode, token, leadTerminals,
members == null ? null : members::snapshot,
members == null ? null : members::roleForSlot,
members == null ? null : members::nameForSlot);
}
private static Supplier<Map<String, String>> fixed(Map<String, String> leadTerminals) {
@@ -151,6 +145,21 @@ public final class CallerResolver {
private CallerResolver(ConnectionIdentity identity, boolean tokenMode, String token,
Supplier<Map<String, String>> leadTerminals,
Supplier<Map<String, String>> architectTerminals) {
this(identity, tokenMode, token, leadTerminals, architectTerminals, null);
}
private CallerResolver(ConnectionIdentity identity, boolean tokenMode, String token,
Supplier<Map<String, String>> leadTerminals,
Supplier<Map<String, String>> architectTerminals,
Function<String, MemberRole> memberSlotRoles) {
this(identity, tokenMode, token, leadTerminals, architectTerminals, memberSlotRoles, null);
}
private CallerResolver(ConnectionIdentity identity, boolean tokenMode, String token,
Supplier<Map<String, String>> leadTerminals,
Supplier<Map<String, String>> architectTerminals,
Function<String, MemberRole> memberSlotRoles,
Function<String, String> memberSlotNames) {
if (tokenMode && (token == null || token.isBlank())) {
throw new IllegalArgumentException(
"auth.mode=token requires a non-empty token; check that the env var named by "
@@ -161,6 +170,8 @@ public final class CallerResolver {
this.expectedToken = tokenMode ? token.getBytes(StandardCharsets.UTF_8) : null;
this.leadTerminals = leadTerminals == null ? Map::of : leadTerminals;
this.architectTerminals = architectTerminals == null ? Map::of : architectTerminals;
this.memberSlotRoles = memberSlotRoles == null ? _ -> null : memberSlotRoles;
this.memberSlotNames = memberSlotNames == null ? Function.identity() : memberSlotNames;
}
/**
@@ -206,11 +217,12 @@ public final class CallerResolver {
return Principal.leader(lead, c.terminal(), c.pid());
}
String slot = architectTerminals.get().get(c.terminal());
if (slot != null) {
if (slot != null && memberSlotRoles.apply(slot) == MemberRole.ARCHITECT) {
// The config/live binding names this pane as an architect slot's own. Same
// unforgeable pane mapping; the live binding, never a request argument, decides.
// Checked before the generic worker fallback, per the CB-548 precedence order.
return Principal.architect(slot, c.terminal(), c.pid());
// Check the slot role too: this defence in depth prevents a bad lifecycle bind from
// escalating a dev or reviewer into an architect. Checked before the worker fallback.
return Principal.architect(memberSlotNames.apply(slot), c.terminal(), c.pid());
}
return Principal.worker(c.terminal(), c.pid()); // unforgeable; never token-gated
}
@@ -0,0 +1,21 @@
package dev.ltms.bridged.auth;
import dev.ltms.bridged.peer.MemberRole;
/** Optional session lifecycle hook for live member-slot bindings. */
public interface MemberLifecycle {
MemberLifecycle NONE = new MemberLifecycle() {
@Override
public void acquired(MemberRole role, String profile, String terminal) {
}
@Override
public void released(String terminal) {
}
};
void acquired(MemberRole role, String profile, String terminal);
void released(String terminal);
}
@@ -2,11 +2,14 @@ package dev.ltms.bridged.auth;
import dev.ltms.bridged.config.BridgedConfig;
import dev.ltms.bridged.peer.MemberRole;
import org.slf4j.Logger;
import org.slf4j.LoggerFactory;
import java.util.Collections;
import java.util.HashMap;
import java.util.LinkedHashMap;
import java.util.Map;
import java.util.Objects;
/**
* The architect-slot registry (CB-548): every gateway-local architect name and the strong-model
@@ -29,7 +32,9 @@ import java.util.Map;
* exposes the map the resolver resolves against plus the profile lookup lifecycle will call.
* Nothing here creates or manages an architect session.
*/
public final class MemberRegistry {
public final class MemberRegistry implements MemberLifecycle {
private static final Logger log = LoggerFactory.getLogger(MemberRegistry.class);
/**
* One flattened {@code fleet:} entry.
@@ -125,6 +130,12 @@ public final class MemberRegistry {
return e == null ? null : e.role();
}
/** The unqualified configured name for a slot, or {@code null} if it is unknown. */
public String nameForSlot(String slotName) {
Entry e = slots.get(slotName);
return e == null ? null : e.name();
}
/** True when {@code slotName} is a configured architect slot. */
public boolean isSlot(String slotName) {
return slots.containsKey(slotName);
@@ -189,4 +200,33 @@ public final class MemberRegistry {
return true;
}
}
/**
* Bind only architect sessions to a free slot with the resolved profile.
*
* <p>The role check is lifecycle policy. {@link CallerResolver} repeats it when resolving a
* binding, so a later lifecycle regression cannot turn a worker into an architect.
*/
@Override
public void acquired(MemberRole role, String profile, String terminal) {
if (role != MemberRole.ARCHITECT || terminal == null || terminal.isBlank()) {
return;
}
// slotsFor preserves definition order, so duplicate-profile slots use the first free one.
for (Entry entry : slotsFor(MemberRole.ARCHITECT).values()) {
if (Objects.equals(profile, entry.profile()) && bind(entry.key(), terminal)) {
return;
}
}
log.info("member slot: no free architect slot for profile={}; session remains a worker", profile);
}
/** Unbind a released terminal using the compare-safe registry operation. */
@Override
public void released(String terminal) {
String slot = slotForTerminal(terminal);
if (slot != null) {
unbind(slot, terminal);
}
}
}
@@ -48,7 +48,7 @@ public record Principal(Role role, String terminal, long pid, String name) {
* made a lead unaddressable: {@link #ownsSession} could never be true for it, so
* {@code bridge_reply} was refused and one lead could send to another but never be answered.
* The terminal now means "which pane is this caller", the presence map keys on
* {@link #isWorker()} instead, and a lead is a peer that can both send and receive.
* {@link #isSpawnedMember()} instead, and a lead is a peer that can both send and receive.
*/
public static Principal leader(String name, String terminal, long pid) {
return new Principal(Role.PRIMARY, terminal, pid, name);
@@ -85,6 +85,16 @@ public record Principal(Role role, String terminal, long pid, String name) {
return role == Role.WORKER;
}
/**
* Whether this caller is a spawned member with its own pane.
*
* <p>Both workers and architects are spawned members. A lead is excluded because recording it
* as present would count it as an available member in the roster.
*/
public boolean isSpawnedMember() {
return role == Role.WORKER || role == Role.ARCHITECT;
}
public boolean isAnonymous() {
return role == Role.ANONYMOUS;
}
@@ -73,6 +73,7 @@ public record BridgedConfig(
Primary primary,
Fleet fleet,
LeadHeartbeat leadHeartbeat,
Health health,
String placement,
Auth auth,
ConfigReload configReload) {
@@ -83,7 +84,16 @@ public record BridgedConfig(
Integer spawnReadyPollMs, Broker broker, Primary primary, Fleet fleet,
LeadHeartbeat leadHeartbeat, String placement, Auth auth) {
this(bind, herdrSocket, profiles, guard, worktreeRoot, lifecycle, spawnReadyTimeoutMs,
spawnReadyPollMs, broker, primary, fleet, leadHeartbeat, placement, auth, null);
spawnReadyPollMs, broker, primary, fleet, leadHeartbeat, null, placement, auth, null);
}
/** Back-compat form before the optional {@code health:} block was added. */
public BridgedConfig(Bind bind, String herdrSocket, Map<String, Profile> profiles, Guard guard,
String worktreeRoot, Lifecycle lifecycle, Integer spawnReadyTimeoutMs,
Integer spawnReadyPollMs, Broker broker, Primary primary, Fleet fleet,
LeadHeartbeat leadHeartbeat, String placement, Auth auth, ConfigReload configReload) {
this(bind, herdrSocket, profiles, guard, worktreeRoot, lifecycle, spawnReadyTimeoutMs,
spawnReadyPollMs, broker, primary, fleet, leadHeartbeat, null, placement, auth, configReload);
}
/**
@@ -210,8 +220,14 @@ public record BridgedConfig(
// worker the primary's IDE servers, which are bound to the primary's checkout — so its
// navigation returned paths outside its own worktree. GitWorktrees now neutralizes that
// file instead; a worker's tools are whatever its launcher mounts.
// .claude/settings.local.json is the sibling that was left behind: it pre-approves tools
// (mcp__context7__*, mcp__jetbrains, mcp__intellij-index, Workflow(code-review)) a worker
// must never hold, and enables MCP servers by name. Its grants are currently INERT because
// GitWorktrees.isolateToolSurface strips every worktree's server map to empty — the named
// servers do not exist there to be enabled. This is defence in depth, not a live fix: the
// worker stays isolated only because this separate mechanism already removes the servers.
parityOverlay = (parityOverlay == null || parityOverlay.isEmpty())
? List.of(".claude/settings.local.json", ".env", ".envrc")
? List.of(".env", ".envrc")
: List.copyOf(parityOverlay);
// gitTokenEnv stays null when unset (opt-in). gitHostEnv defaults so operators enabling
// checkpoints need only set gitTokenEnv; it is injected only alongside a resolved token.
@@ -371,6 +387,17 @@ public record BridgedConfig(
boolean clearAfterTurn) {
}
/** Optional fleet detection. A missing block stays dormant. */
@JsonIgnoreProperties(ignoreUnknown = true)
public record Health(Boolean enabled, Integer intervalSeconds, Integer workingSuspectAfterSeconds,
Integer paneProbeIntervalSeconds, Notifications notifications) {
public boolean isEnabled() { return Boolean.TRUE.equals(enabled); }
public int intervalOrDefault() { return Math.max(15, intervalSeconds == null ? 30 : intervalSeconds); }
public record Notifications(String mode) {
public boolean configured() { return "webhook".equalsIgnoreCase(mode); }
}
}
/**
* External AMQP broker for durable, cross-restart reply delivery (CB-307 Stage 2). Its mere
* presence swaps the in-memory {@code ReplyInbox} for the AMQP-backed adapter; absent, bridged
@@ -526,6 +553,7 @@ public record BridgedConfig(
* @param architects profiles the {@code architect} role may run on
* @param developers profiles the {@code dev} role may run on
* @param reviewers profiles the {@code reviewer} role may run on
* @param charters optional launch-charter text keyed by singular role wire name
* @param tabLabel template for a member tab's label; {@code {role}}, {@code {profile}},
* {@code {model}} and {@code {n}} (a per role+profile counter) are
* substituted. Default {@link #DEFAULT_TAB_LABEL}
@@ -535,6 +563,7 @@ public record BridgedConfig(
Map<String, Slot> architects,
Map<String, Slot> developers,
Map<String, Slot> reviewers,
Map<String, String> charters,
String tabLabel) {
/**
@@ -550,9 +579,24 @@ public record BridgedConfig(
architects = unmodifiableOrEmpty(architects);
developers = unmodifiableOrEmpty(developers);
reviewers = unmodifiableOrEmpty(reviewers);
charters = unmodifiableOrEmpty(charters);
tabLabel = (tabLabel == null || tabLabel.isBlank()) ? DEFAULT_TAB_LABEL : tabLabel;
}
/**
* A fleet with no configured launch charters — the shape every deployment had before
* CB-566, and what most tests want.
*
* <p>Kept deliberately, even though an overload that drops a new field is normally the
* shape to avoid. It is safe here because nothing <em>reads</em> a charter through a
* constructor: the launcher reads {@code fleet.charters()} from the live config. Jackson
* binds the canonical constructor, so this one cannot swallow an operator's YAML.
*/
public Fleet(Map<String, Leader> leaders, Map<String, Slot> architects,
Map<String, Slot> developers, Map<String, Slot> reviewers, String tabLabel) {
this(leaders, architects, developers, reviewers, null, tabLabel);
}
/**
* Deliberately not {@code Map.copyOf}: its iteration order is salted per JVM run, which
* would discard YAML definition order. The {@code fixed} placement policy answers with a
@@ -576,6 +620,11 @@ public record BridgedConfig(
};
}
/** The configured launch charter for {@code role}, or {@code null} when it is absent. */
public String charterFor(MemberRole role) {
return role == null ? null : charters.get(role.wireName());
}
/**
* The profile names {@code role} may run on, in definition order, without repeats.
*
@@ -784,7 +833,7 @@ public record BridgedConfig(
private static final Set<String> KNOWN_TOP_LEVEL_KEYS = Set.of(
"bind", "herdrSocket", "profiles", "guard", "worktreeRoot",
"lifecycle", "spawnReadyTimeoutMs", "spawnReadyPollMs", "broker", "primary", "fleet",
"leadHeartbeat", "placement", "auth", "configReload");
"leadHeartbeat", "health", "placement", "auth", "configReload");
/** Load and validate config from {@code path}. */
public static BridgedConfig load(Path path) {
@@ -1046,7 +1095,7 @@ public record BridgedConfig(
// fleet IS defaulted, unlike the leadScan: block it replaced, because an empty Fleet is not
// the same as an enabled one: every pool is empty, so no lead is scanned for or created and
// no role has a pool. Constructing it saves every reader a null check for no behaviour change.
Fleet f = (fleet != null) ? fleet : new Fleet(null, null, null, null, null);
Fleet f = (fleet != null) ? fleet : new Fleet(null, null, null, null, null, null);
// leadHeartbeat is left as-is (CB-551): null is "off", and LeadHeartbeat's own compact
// constructor defaults the fields of a block that IS present. Defaulting it here would
// switch the feature on for every config that never mentioned it.
@@ -1054,7 +1103,7 @@ public record BridgedConfig(
// defaults the fields of a block that IS present. Defaulting it here would start watching
// the file for every config that never asked to be watched.
return new BridgedConfig(b, herdrSocket, profiles, g, worktreeRoot, l, timeout, pollMs,
broker, primary, f, leadHeartbeat, placementOrDefault, a, configReload);
broker, primary, f, leadHeartbeat, health, placementOrDefault, a, configReload);
}
/**
@@ -1184,6 +1233,35 @@ public record BridgedConfig(
}
}
/**
* Reject configured charter entries that would remove a role's contract or never be read.
*
* <p>The map deliberately retains every key from {@code fleet.charters:}. A typed record would
* silently discard an unknown child because {@link Fleet} ignores unknown JSON properties, which
* would make a typo look like an accepted configuration.
*
* @throws IllegalStateException when a charter key is not a role wire name or its value is blank
*/
public void validateCharters() {
if (fleet == null || fleet.charters().isEmpty()) {
return;
}
List<String> valid = Arrays.stream(MemberRole.values())
.map(MemberRole::wireName)
.toList();
List<String> bad = new ArrayList<>();
fleet.charters().forEach((key, charter) -> {
if (!valid.contains(key)) {
bad.add("fleet.charters." + key + " is not a role wire name (valid: " + valid + ").");
} else if (charter == null || charter.isBlank()) {
bad.add("fleet.charters." + key + " is blank; a configured role needs charter text.");
}
});
if (!bad.isEmpty()) {
throw new IllegalStateException("refusing to start: " + String.join(" ", bad));
}
}
/**
* Reject a member slot whose {@code role} or {@code profile} does not resolve.
*
@@ -27,10 +27,10 @@ import java.util.function.Supplier;
*
* <ul>
* <li><strong>Hot</strong> — re-read per use, so a reload takes effect on the next spawn:
* {@code fleet:} (every role pool and {@code tabLabel}), {@code placement:}, and an existing
* profile's {@code weight} / {@code maxLoad}. Those three are read through a supplier on
* {@code CompositePeerLauncher}, which is what makes them hot — not the fact that they are
* config.</li>
* {@code fleet:} (every role pool, {@code charters}, and {@code tabLabel}),
* {@code placement:}, and an existing profile's {@code weight} / {@code maxLoad}. Those
* three are read through a supplier on {@code CompositePeerLauncher}, which is what makes
* them hot — not the fact that they are config.</li>
* <li><strong>Deferred</strong> — accepted into the new snapshot, but the wiring built at startup
* keeps the old value until a restart: {@code lifecycle:}, {@code leadHeartbeat:},
* {@code spawnReadyTimeoutMs} / {@code spawnReadyPollMs}, {@code guard:},
@@ -150,6 +150,7 @@ public final class ConfigRef implements Supplier<BridgedConfig> {
fresh.validateAuthExposure();
fresh.validateLeadTabPrefixes();
fresh.validateSubscriptionProfiles();
fresh.validateCharters();
fresh.validateMembers();
} catch (RuntimeException e) {
String msg = e.getMessage() == null ? e.toString() : e.getMessage();
@@ -0,0 +1,40 @@
package dev.ltms.bridged.health;
import dev.ltms.bridged.herdr.AgentStatus;
import dev.ltms.bridged.session.MemberSession;
/**
* Pure classifier. Collection and repair are deliberately outside this package.
* {@link HealthState#ERROR_ON_SCREEN} is not decided yet because it needs a bounded pane detection
* read and an adapter-specific fatal signature; status facts alone must not guess it.
*/
public final class FleetHealth {
private FleetHealth() { }
public static HealthDecision decide(HealthSnapshot s, HealthPrior prior, long nowNanos) {
if (s.controlLinkDown()) return result(HealthState.CONTROL_LINK_DOWN, false);
if (s.targetNotFound()) return result(HealthState.GONE, false);
if (s.sessionState() == MemberSession.State.SPAWNING && !s.present() && s.readinessGraceElapsed()) {
return result(HealthState.NEVER_READY, false);
}
if (s.orphanedDelegation()) return result(HealthState.DELEGATION_ORPHANED, false);
boolean disagreement = s.sessionState() == MemberSession.State.BUSY && s.acceptedDelivery()
&& (s.liveStatus() == AgentStatus.IDLE || s.liveStatus() == AgentStatus.DONE);
if (disagreement && prior.busyButDone()) return result(HealthState.TURN_BOUNDARY_LOST, true);
if (s.stalled()) return result(HealthState.STALL_SUSPECTED, disagreement);
if (s.replyStranded()) return result(HealthState.REPLY_STRANDED, disagreement);
if (s.queuedDelivery() || s.inboxMessage()) return result(HealthState.WORK_PENDING, disagreement);
if (s.sessionState() == MemberSession.State.SPAWNING) return result(HealthState.STARTING, disagreement);
if (s.acceptedDelivery() && s.liveStatus() == AgentStatus.BLOCKED) {
return result(HealthState.BLOCKED_AMBIGUOUS, disagreement);
}
// An accepted delivery remains bridge work even when herdr is late, unknown, or has already
// reported DONE once. It cannot be IDLE until the delegation has resolved.
if (s.acceptedDelivery()) return result(HealthState.WORKING, disagreement);
return result(HealthState.IDLE, disagreement);
}
private static HealthDecision result(HealthState state, boolean disagreement) {
return new HealthDecision(state, new HealthPrior(disagreement));
}
}
@@ -0,0 +1,107 @@
package dev.ltms.bridged.health;
import dev.ltms.bridged.herdr.Agent;
import dev.ltms.bridged.herdr.AgentControl;
import dev.ltms.bridged.herdr.AgentStatus;
import dev.ltms.bridged.msg.MessageService;
import dev.ltms.bridged.session.MemberSession;
import org.slf4j.Logger;
import org.slf4j.LoggerFactory;
import java.util.HashMap;
import java.util.HashSet;
import java.util.List;
import java.util.Map;
import java.util.concurrent.ScheduledExecutorService;
import java.util.concurrent.TimeUnit;
import java.util.function.LongSupplier;
import java.util.function.Supplier;
/** Slow whole-fleet evidence collection. It is deliberately separate from the delivery poller. */
public final class FleetHealthMonitor {
private static final Logger log = LoggerFactory.getLogger(FleetHealthMonitor.class);
private final AgentControl agents;
private final Supplier<List<MemberSession>> roster;
private final MessageService messages;
private final ScheduledExecutorService scheduler;
private final LongSupplier clock;
private final long intervalSeconds;
private final Map<String, HealthPrior> priors = new HashMap<>();
private final Map<String, HealthState> states = new HashMap<>();
// These facts need the evidence publishers introduced by later M4 units. They are not negatives.
private static final boolean NOT_YET_OBSERVED = false;
public FleetHealthMonitor(AgentControl agents, Supplier<List<MemberSession>> roster, MessageService messages,
ScheduledExecutorService scheduler, LongSupplier clock, long intervalSeconds) {
this.agents = agents;
this.roster = roster;
this.messages = messages;
this.scheduler = scheduler;
this.clock = clock;
this.intervalSeconds = intervalSeconds;
}
/** Pure per-member decision seam. */
static HealthDecision decide(HealthSnapshot snapshot, HealthPrior prior, long nowNanos) {
return FleetHealth.decide(snapshot, prior, nowNanos);
}
public void start() { scheduler.schedule(this::tick, intervalSeconds, TimeUnit.SECONDS); }
public void stop() { scheduler.shutdownNow(); }
// Package-private so tests can run one tick without waiting.
void tick() {
try {
List<Agent> agentsNow = agents.list(); // Exactly one list call for this complete observation.
List<MemberSession> rosterNow = roster.get(); // One in-memory roster snapshot for this tick.
Map<String, Agent> live = new HashMap<>();
for (Agent agent : agentsNow) live.put(agent.terminalId(), agent);
HashSet<String> current = new HashSet<>();
for (MemberSession session : rosterNow) {
current.add(session.terminalId());
Agent agent = live.get(session.terminalId());
AgentStatus status = agent == null ? AgentStatus.UNKNOWN : agent.status();
boolean accepted = messages.hasAcceptedDelivery(session.terminalId());
HealthSnapshot snapshot = new HealthSnapshot(session.state(), status, accepted, NOT_YET_OBSERVED,
messages.hasInboxMessage(session.terminalId()), agent != null, NOT_YET_OBSERVED,
NOT_YET_OBSERVED, NOT_YET_OBSERVED, NOT_YET_OBSERVED, NOT_YET_OBSERVED, NOT_YET_OBSERVED);
HealthDecision decision = decide(snapshot, priors.getOrDefault(session.terminalId(), HealthPrior.NONE),
clock.getAsLong());
priors.put(session.terminalId(), decision.prior());
reportTransition(session.terminalId(), decision.state());
}
priors.keySet().retainAll(current);
states.keySet().retainAll(current);
} catch (Throwable error) {
// A list failure is health evidence, and must never kill the monitor's only scheduler task.
log.warn("fleet health collection failed; will retry next tick", error);
} finally {
if (!scheduler.isShutdown()) {
scheduler.schedule(this::tick, intervalSeconds, TimeUnit.SECONDS);
}
}
}
void reportTransition(String target, HealthState next) {
HealthState previous = states.put(target, next);
if (previous == next) return;
if (fault(next)) {
log.warn("fleet health member={} state={} previous={}", target, next, previous);
} else if (previous != null && fault(previous)) {
log.info("fleet health member={} recovered state={} previous={}", target, next, previous);
}
}
private static boolean fault(HealthState state) {
return switch (state) {
case NEVER_READY, GONE, TURN_BOUNDARY_LOST, ERROR_ON_SCREEN, STALL_SUSPECTED,
MUTE, REPLY_STRANDED, DELEGATION_ORPHANED, CONTROL_LINK_DOWN -> true;
default -> false;
};
}
public static String coverage(boolean enabled, boolean notificationConfigured) {
return !enabled ? "off" : notificationConfigured ? "full" : "detection-only";
}
}
@@ -0,0 +1,4 @@
package dev.ltms.bridged.health;
/** Classification plus the private fact that the next pure decision needs. */
public record HealthDecision(HealthState state, HealthPrior prior) { }
@@ -0,0 +1,6 @@
package dev.ltms.bridged.health;
/** Private cross-tick observation. It is deliberately not a reported health value. */
public record HealthPrior(boolean busyButDone) {
public static final HealthPrior NONE = new HealthPrior(false);
}
@@ -0,0 +1,11 @@
package dev.ltms.bridged.health;
import dev.ltms.bridged.herdr.AgentStatus;
import dev.ltms.bridged.session.MemberSession;
/** Read-only facts from one fleet collection tick. */
public record HealthSnapshot(MemberSession.State sessionState, AgentStatus liveStatus,
boolean acceptedDelivery, boolean queuedDelivery, boolean inboxMessage,
boolean present, boolean targetNotFound, boolean controlLinkDown,
boolean readinessGraceElapsed, boolean orphanedDelegation,
boolean replyStranded, boolean stalled) { }
@@ -0,0 +1,8 @@
package dev.ltms.bridged.health;
/** Health classifications reported for a member. */
public enum HealthState {
STARTING, IDLE, WORKING, WORK_PENDING, BLOCKED_AMBIGUOUS,
NEVER_READY, GONE, TURN_BOUNDARY_LOST, ERROR_ON_SCREEN, STALL_SUSPECTED,
MUTE, REPLY_STRANDED, DELEGATION_ORPHANED, CONTROL_LINK_DOWN
}
@@ -0,0 +1,25 @@
package dev.ltms.bridged.health;
import dev.ltms.bridged.msg.Rendezvous;
import java.util.Map;
import java.util.concurrent.ConcurrentHashMap;
/**
* Counts turns that ended via the completion fallback instead of {@code bridge_reply}.
* MUTE is an observation by target and profile, not a classifier state and never suppresses faults.
*/
public final class MuteCounter {
private final Map<String, Integer> byTarget = new ConcurrentHashMap<>();
private final Map<String, Integer> byProfile = new ConcurrentHashMap<>();
/** Record only fallback completion; a structured reply does not make a member mute. */
public void observe(String target, String profile, Rendezvous.Kind kind) {
if (kind != Rendezvous.Kind.COMPLETION) return;
byTarget.merge(target, 1, Integer::sum);
byProfile.merge(profile, 1, Integer::sum);
}
public int forTarget(String target) { return byTarget.getOrDefault(target, 0); }
public int forProfile(String profile) { return byProfile.getOrDefault(profile, 0); }
}
@@ -0,0 +1,26 @@
package dev.ltms.bridged.health;
import java.util.ArrayList;
import java.util.HashMap;
import java.util.List;
import java.util.Map;
/** Fixed pane-probe limits. Pane content is never retained here. */
public final class PaneBudget {
public static final long COOLDOWN_NANOS = 60_000_000_000L;
public static final int MAX_PER_TICK = 2;
private final Map<String, Long> lastProbe = new HashMap<>();
private int cursor;
public List<String> choose(List<String> candidates, long nowNanos, long configuredCooldownNanos) {
long cooldown = Math.max(COOLDOWN_NANOS, configuredCooldownNanos);
List<String> out = new ArrayList<>();
for (int n = 0; n < candidates.size() && out.size() < MAX_PER_TICK; n++) {
String target = candidates.get((cursor + n) % candidates.size());
Long last = lastProbe.get(target);
if (last == null || nowNanos - last >= cooldown) { out.add(target); lastProbe.put(target, nowNanos); }
}
if (!candidates.isEmpty()) cursor = (cursor + 1) % candidates.size();
return List.copyOf(out);
}
}
@@ -55,6 +55,9 @@ public final class CompletionResolver implements TurnListener {
/** Cap the scraped tail so a long transcript can't return an unbounded blob. */
static final int MAX_SCRAPE_CHARS = 4000;
private static final String CLIPPED_PANE_TAIL_MARKER =
"[Pane tail clipped: member did not call bridge_reply.]";
private final AgentControl agents;
private final Rendezvous rendezvous;
@@ -134,7 +137,13 @@ public final class CompletionResolver implements TurnListener {
@Override
public void onTurnFailed(String target) {
InFlight turn = inFlight.get(target);
Thread.ofVirtual().name("turn-failed-" + target).start(() -> fail(target, turn));
Thread.ofVirtual().name("turn-failed-" + target).start(() -> fail(target, turn, null));
}
@Override
public void onTurnFailed(String target, String reason) {
InFlight turn = inFlight.get(target);
Thread.ofVirtual().name("turn-failed-" + target).start(() -> fail(target, turn, reason));
}
/** Synchronous resolve (the unit-testable core of {@link #onTurnComplete}). */
@@ -147,9 +156,14 @@ public final class CompletionResolver implements TurnListener {
return;
}
String tail;
int originalLength = 0;
boolean clipped = false;
boolean scrapeFailed = false;
try {
tail = clip(lastAssistantBlock(agents.read(target, SCRAPE_SOURCE)));
String assistantBlock = lastAssistantBlock(agents.read(target, SCRAPE_SOURCE));
originalLength = assistantBlock.strip().length();
clipped = originalLength > MAX_SCRAPE_CHARS;
tail = clip(assistantBlock);
} catch (RuntimeException e) {
// The worker finished but we couldn't read its screen — still resolve the send so the
// caller unblocks; an empty tail beats hanging until the caller's timeout.
@@ -169,8 +183,14 @@ public final class CompletionResolver implements TurnListener {
target);
return; // keep the in-flight record: a later genuine completion still needs it
}
if (rendezvous.resolveCompletion(waiter, tail)) {
String completion = clipped ? tail + "\n" + CLIPPED_PANE_TAIL_MARKER : tail;
if (rendezvous.resolveCompletion(waiter, completion)) {
inFlight.remove(target, turn);
if (clipped) {
log.warn("completion scrape for {} clipped from {} chars to the {} char cap; "
+ "member did not call bridge_reply, so the pane tail is partial",
target, originalLength, MAX_SCRAPE_CHARS);
}
log.debug("resolved send to {} via turn-completion fallback ({} chars scraped)",
target, tail.length());
}
@@ -178,6 +198,11 @@ public final class CompletionResolver implements TurnListener {
/** Synchronous fail (the unit-testable core of {@link #onTurnFailed}). */
void fail(String target, InFlight turn) {
fail(target, turn, null);
}
/** Synchronous fail with an optional reason supplied by a dropped worker queue. */
void fail(String target, InFlight turn, String explicitReason) {
// A never-delivered readiness failure has no in-flight record but still has a blocked send;
// fall back to the currently-registered waiter (unambiguous — that send never completed, so
// no next turn exists to confuse it with).
@@ -187,20 +212,22 @@ public final class CompletionResolver implements TurnListener {
inFlight.remove(target, turn); // nobody blocked on this worker — nothing to fail
return;
}
String reason;
try {
reason = clip(agents.read(target, SCRAPE_SOURCE));
} catch (RuntimeException e) {
reason = "";
}
if (reason.isBlank()) {
// No screen to scrape — either the worker is stuck (CB-109) or gone (CB-110).
reason = "worker did not reply; its turn ended in an unrecoverable state "
+ "(worker unreachable or stuck)";
String reason = explicitReason;
if (reason == null || reason.isBlank()) {
try {
reason = clip(agents.read(target, SCRAPE_SOURCE));
} catch (RuntimeException e) {
reason = "";
}
if (reason.isBlank()) {
// No screen to scrape — either the worker is stuck (CB-109) or gone (CB-110).
reason = "worker did not reply; its turn ended in an unrecoverable state "
+ "(worker unreachable or stuck)";
}
}
if (rendezvous.resolveFailure(waiter, reason)) {
inFlight.remove(target, turn);
log.debug("failed send to {} via turn-stall fallback", target);
log.warn("failing send to {} via turn-stall fallback: {}", target, reason);
}
}
@@ -78,6 +78,15 @@ public final class Injector {
*/
private static final int READINESS_GRACE_POLLS = 240;
/**
* The single source for the injector poll cadence — how often the {@link StatusPoller} drives
* {@link #onStatus} at. {@code Bridged} passes this to every {@link StatusPoller} it constructs,
* and this class reads it to state the readiness grace in seconds on the CB-562 expiry log
* instead of hardcoding "60s". One constant, so a cadence change cannot silently desync a log
* that claims a grace duration.
*/
public static final long POLL_INTERVAL_MILLIS = 250;
private final AgentControl agents;
private final TurnListener turnListener;
private final Predicate<String> ready; // CB-113: a target is deliverable only when available
@@ -264,6 +273,11 @@ public final class Injector {
// fail every queued message and release the target (CB-114) instead of
// polling it indefinitely with the caller's future never completing.
notReady = new ArrayList<>(t.queue);
log.warn("readiness grace for {} expired after {} polls ({}s): target never "
+ "became deliverable, so failing {} queued message(s) that never "
+ "reached its pane",
target, READINESS_GRACE_POLLS,
READINESS_GRACE_POLLS * POLL_INTERVAL_MILLIS / 1000, notReady.size());
t.queue.clear();
t.notReadySincePoll = 0;
}
@@ -388,12 +402,16 @@ public final class Injector {
t.awaitingPostTurnPickup = false;
t.postTurnObserved = false;
}
log.warn("{} is gone, dropping its queue: {} message(s) failed{}; cause: {}", target,
pending.size(),
hadDeliveredTurn ? " (including one turn already in flight whose completion was never confirmed)" : "",
cause.getMessage());
forget.accept(target); // the worker is gone — clear its readiness/presence too (CB-114)
for (Pending p : pending) {
p.delivered().completeExceptionally(cause);
}
if (hadDeliveredTurn) {
turnListener.onTurnFailed(target);
}
// A queued send has no in-flight record, while a delivered turn does. CompletionResolver
// handles both forms and resolves its waiter at most once.
turnListener.onTurnFailed(target, cause.getMessage());
}
}
@@ -41,6 +41,14 @@ public interface TurnListener {
default void onTurnFailed(String target) {
}
/**
* As {@link #onTurnFailed(String)}, carrying the reason a worker became unreachable. The default
* keeps existing listeners working while allowing the completion resolver to report a useful cause.
*/
default void onTurnFailed(String target, String reason) {
onTurnFailed(target);
}
/**
* A message was just delivered into {@code target}'s pane (CB-115). Fired so the completion
* resolver can snapshot the pane's pre-turn content: a later {@link #onTurnComplete} whose
@@ -0,0 +1,30 @@
package dev.ltms.bridged.logging;
import ch.qos.logback.classic.Level;
import ch.qos.logback.classic.Logger;
import ch.qos.logback.classic.turbo.TurboFilter;
import ch.qos.logback.core.spi.FilterReply;
import io.modelcontextprotocol.spec.McpSchema;
import org.slf4j.Marker;
/** Suppresses only the SDK warning for the normal MCP cancellation notification. */
public final class McpCancelledNotificationFilter extends TurboFilter {
static final String LOGGER = "io.modelcontextprotocol.spec.McpStreamableServerSession";
static final String UNHANDLED_NOTIFICATION = "No handler registered for notification method: {}";
@Override
public FilterReply decide(Marker marker, Logger logger, Level level, String format, Object[] params,
Throwable throwable) {
if (level == Level.WARN
&& LOGGER.equals(logger.getName())
&& UNHANDLED_NOTIFICATION.equals(format)
&& params != null
&& params.length == 1
&& params[0] instanceof McpSchema.JSONRPCNotification notification
&& "notifications/cancelled".equals(notification.method())) {
return FilterReply.DENY;
}
return FilterReply.NEUTRAL;
}
}
@@ -33,7 +33,10 @@ import jakarta.servlet.http.HttpServlet;
import java.util.LinkedHashMap;
import java.util.List;
import java.util.Map;
import java.util.Set;
import java.util.function.Function;
import java.util.function.LongSupplier;
import java.util.function.Supplier;
import java.util.stream.Collectors;
/**
@@ -73,17 +76,20 @@ public final class BridgeMcp {
private final McpSyncServer server;
private final CallerResolver authz; // CB-501: null → authorization not enforced (legacy)
private final Metrics metrics; // CB-502: null → auth failures not counted
private final CapacitySource capacity;
private final HealthCoverageSource healthCoverage;
/**
* Legacy constructor — no authorization. Retained so existing tests exercise tool behaviour
* without an auth fixture.
*/
public BridgeMcp(MessageService messages, PeerLauncher workers,
SessionManager sessions, ConnectionIdentity identity, MemberPresence presence,
PrimaryRegistry primaryRegistry) {
this(messages, workers, sessions, identity, presence, primaryRegistry, null, null);
/** Capacity facts used by {@code bridge_list}; production must supply the placement live count. */
public record CapacitySource(Function<String, Integer> liveCount, Function<String, Integer> maxLoad,
Supplier<Set<String>> configuredProfiles, LongSupplier clock) {
/** Inert test-only source. It omits capacity rather than inventing zero live counts. */
public static CapacitySource none() { return new CapacitySource(_ -> 0, _ -> null, Set::of, System::nanoTime); }
boolean available() { return !configuredProfiles.get().isEmpty(); }
}
/** Coverage is supplied by the health wiring, not inferred from a missing dependency. */
public record HealthCoverageSource(Supplier<String> value) { }
/**
* @param callers resolves each call's {@link Principal}; {@code null} disables authorization.
* This surface needs its own enforcement: {@code /mcp} is a raw servlet on
@@ -91,9 +97,11 @@ public final class BridgeMcp {
* filter, so the REST guard does not cover it.
* @param metrics registry for auth-failure counting; may be {@code null}
*/
public BridgeMcp(MessageService messages, PeerLauncher workers,
SessionManager sessions, ConnectionIdentity identity, MemberPresence presence,
PrimaryRegistry primaryRegistry, CallerResolver callers, Metrics metrics) {
public BridgeMcp(MessageService messages, PeerLauncher workers, SessionManager sessions,
ConnectionIdentity identity, MemberPresence presence, PrimaryRegistry primaryRegistry,
CallerResolver callers, Metrics metrics, CapacitySource capacity, HealthCoverageSource healthCoverage) {
this.capacity = capacity;
this.healthCoverage = healthCoverage;
McpJsonMapper json = new JacksonMcpJsonMapperSupplier().get();
this.transport = HttpServletStreamableServerTransportProvider.builder()
.jsonMapper(json)
@@ -109,10 +117,10 @@ public final class BridgeMcp {
? callers.resolve(req.getRemoteAddr(), req.getRemotePort(),
req.getHeader("Authorization"))
: legacyPrincipal(identity, req.getRemoteAddr(), req.getRemotePort());
// CB-532: guard on the ROLE, not on the terminal being null. A named lead now
// carries its pane too, and enrolling a lead in the worker presence map would
// have it counted as an available worker.
if (p.isWorker()) presence.markPresent(p.terminal());
// CB-532: guard on the ROLE, not on the terminal being null. This excludes a
// lead, which carries its pane too, while including every spawned member role.
// Enrolling a lead would count it as an available member in the roster.
markSpawnedMemberPresent(p, presence);
return McpTransportContext.create(Map.of(
CALLER_TERMINAL, orEmpty(p.terminal()),
CALLER_PID, Long.toString(p.pid()),
@@ -151,8 +159,8 @@ public final class BridgeMcp {
Runnable onAccepted = () -> primaryRegistry.recordDelegation(target, caller);
// wait defaults to true (block for the reply); wait:false is fire-and-poll.
return Boolean.FALSE.equals(a.get("wait"))
? sendAsync(messages, target, content, onAccepted)
: send(messages, target, content, timeoutMs(a), onAccepted);
? sendAsync(messages, target, content, onAccepted, workers.profiles())
: send(messages, target, content, timeoutMs(a), onAccepted, workers.profiles());
})
// bridge_reply's identity is the CONNECTION, never an argument — so the authz check
// is "is this caller a worker at all", and it can only ever reply as itself.
@@ -207,7 +215,7 @@ public final class BridgeMcp {
.toolCall(listTool(), (exchange, _) -> {
McpSchema.CallToolResult denied = deny(exchange, Authz.Action.READ, null);
if (denied != null) return denied;
return listFleet(workers, sessions,
return listFleet(workers, sessions, messages, capacity, healthCoverage,
callers == null ? Map.of() : callers.leads(),
callerTerminal(exchange));
})
@@ -365,6 +373,13 @@ public final class BridgeMcp {
return transport;
}
/** Mark a connected spawned member available for the injector readiness gate. */
static void markSpawnedMemberPresent(Principal caller, MemberPresence presence) {
if (caller.isSpawnedMember()) {
presence.markPresent(caller.terminal());
}
}
/** Graceful shutdown of the MCP server. */
public void close() {
server.closeGracefully();
@@ -372,21 +387,22 @@ public final class BridgeMcp {
// --- tool logic (thin adapters over the services; unit-testable) ---------------------------
/** {@code bridge_send}: delegate {@code content} to a worker session and block for its reply. */
static McpSchema.CallToolResult send(MessageService messages, String sessionId, String content, Long timeoutMs) {
return send(messages, sessionId, content, timeoutMs, null);
}
/**
* As {@link #send(MessageService, String, String, Long)}, wiring an accepted-delivery hook
* {@code bridge_send}: delegate {@code content} to a worker session and block for its reply.
* The configured profiles are required so a profile name can never bypass target validation.
*
* (CB-548): {@code onAccepted} records delegator ownership the instant the send is accepted, so
* a BUSY interloper never claims a turn it did not win. {@code null} disables recording.
*/
static McpSchema.CallToolResult send(MessageService messages, String sessionId, String content,
Long timeoutMs, Runnable onAccepted) {
Long timeoutMs, Runnable onAccepted, Set<String> profiles) {
if (isBlank(sessionId) || isBlank(content)) {
return error("sessionId and content are required");
}
McpSchema.CallToolResult targetError = profileTargetError(sessionId, profiles);
if (targetError != null) {
return targetError;
}
long timeout = clamp(timeoutMs == null ? DEFAULT_TIMEOUT_MS : timeoutMs);
try {
return formatReply(messages.send(sessionId, content, timeout, onAccepted), timeout);
@@ -456,24 +472,33 @@ public final class BridgeMcp {
/**
* {@code bridge_send} with {@code wait:false}: delegate {@code content} and return a ticket
* immediately (fire-and-poll), so a long task isn't cut off by the caller's MCP call timeout.
*/
static McpSchema.CallToolResult sendAsync(MessageService messages, String sessionId, String content) {
return sendAsync(messages, sessionId, content, null);
}
/**
* As {@link #sendAsync(MessageService, String, String)}, wiring the accepted-delivery hook
* The configured profiles are required so a profile name can never bypass target validation.
*
* This wires the accepted-delivery hook
* (CB-548) so an async flooding send records delegator ownership exactly once it is accepted.
*/
static McpSchema.CallToolResult sendAsync(MessageService messages, String sessionId, String content,
Runnable onAccepted) {
Runnable onAccepted, Set<String> profiles) {
if (isBlank(sessionId) || isBlank(content)) {
return error("sessionId and content are required");
}
McpSchema.CallToolResult targetError = profileTargetError(sessionId, profiles);
if (targetError != null) {
return targetError;
}
String ticket = messages.sendAsync(sessionId, content, onAccepted);
return text("accepted — task delegated. Poll bridge_poll with ticket=" + ticket);
}
/** A configured profile is never a send target; other unknown values may be herdr-owned panes. */
private static McpSchema.CallToolResult profileTargetError(String sessionId, Set<String> profiles) {
if (profiles.contains(sessionId)) {
return error("unknown send target \"" + sessionId + "\": it is a configured profile name, not a "
+ "session id. Call bridge_list to find a member or lead sessionId.");
}
return null;
}
/** {@code bridge_poll}: check an async delegation by ticket, or drain a worker's inbox by target. */
static McpSchema.CallToolResult poll(MessageService messages, String ticket, String target) {
if (!isBlank(target)) {
@@ -495,6 +520,9 @@ public final class BridgeMcp {
? "[done — worker finished without a structured bridge_reply; transcript tail follows]\n" + v.reply()
: v.reply());
case PENDING -> text("[pending — " + v.detail() + "]");
case ASKING -> text("[question — worker is waiting for your answer]\n" + v.reply()
+ "\n\nAnswer it by calling bridge_send again with turnId=\"" + v.turnId()
+ "\" and content set to your answer; the worker resumes the same turn.");
case FAILED -> text("[failed — " + v.detail() + "]");
};
}
@@ -701,6 +729,12 @@ public final class BridgeMcp {
*/
static McpSchema.CallToolResult listFleet(PeerLauncher workers, SessionManager sessions,
Map<String, String> leads, String selfTerm) {
return listFleet(workers, sessions, null, CapacitySource.none(), new HealthCoverageSource(() -> "off"), leads, selfTerm);
}
static McpSchema.CallToolResult listFleet(PeerLauncher workers, SessionManager sessions, MessageService messages,
CapacitySource capacity, HealthCoverageSource healthCoverage,
Map<String, String> leads, String selfTerm) {
try {
Map<String, Agent> live = workers.list().stream()
.map(Agent.class::cast)
@@ -710,15 +744,57 @@ public final class BridgeMcp {
.sorted(Map.Entry.comparingByValue())
.map(e -> leadView(e.getKey(), e.getValue(), live.get(e.getKey()), selfTerm))
.toList();
List<Map<String, Object>> out = sessions.roster().stream()
.map(s -> SessionManager.rosterView(s, live.get(s.terminalId())))
List<MemberSession> roster = sessions.roster();
List<Map<String, Object>> out = roster.stream()
.map(s -> memberCapacityView(s, live.get(s.terminalId()), messages, capacity.clock().getAsLong()))
.toList();
return text(json(Map.of("leads", leadRows, "members", out)));
Set<String> profiles = new java.util.TreeSet<>(capacity.configuredProfiles().get());
roster.stream().map(MemberSession::profile).forEach(profiles::add);
Map<String, Object> result = new LinkedHashMap<>();
result.put("leads", leadRows); result.put("members", out);
result.put("healthCoverage", healthCoverage.value().get());
if (capacity.available()) result.put("capacity", profiles.stream()
.map(profile -> capacityView(profile, capacity.liveCount(), capacity.maxLoad(), roster, messages,
capacity.clock().getAsLong())).toList());
return text(json(result));
} catch (HerdrException e) {
return error("herdr error listing the fleet: " + e.getMessage());
}
}
/**
* Capacity is advisory only. {@code reclaimable} says there is no bridge work, not that bridged
* may stop the member: the bridge has capacity facts but no work list, and choosing work needs
* authority it does not have. {@code idleForSeconds} is derived from monotonic nanoTime and has
* no wall-clock meaning across a daemon restart.
*/
private static Map<String, Object> memberCapacityView(MemberSession session, Agent live,
MessageService messages, long nowNanos) {
Map<String, Object> row = SessionManager.rosterView(session, live);
boolean open = messages != null && messages.hasAcceptedDelivery(session.terminalId());
boolean inbox = messages != null && messages.hasInboxMessage(session.terminalId());
boolean reclaimable = (session.state() == MemberSession.State.READY || session.state() == MemberSession.State.DONE)
&& !open && !inbox;
row.put("reclaimable", reclaimable);
row.put("idleForSeconds", reclaimable ? Math.max(0, (nowNanos - session.lastActivityAtNanos()) / 1_000_000_000L) : null);
return row;
}
private static Map<String, Object> capacityView(String profile, Function<String, Integer> liveCount,
Function<String, Integer> maxLoad, List<MemberSession> roster,
MessageService messages, long nowNanos) {
Integer cap = maxLoad.apply(profile);
int live = liveCount.apply(profile);
int reclaimable = (int) roster.stream().filter(s -> profile.equals(s.profile()))
.filter(s -> (s.state() == MemberSession.State.READY || s.state() == MemberSession.State.DONE))
.filter(s -> messages == null || (!messages.hasAcceptedDelivery(s.terminalId()) && !messages.hasInboxMessage(s.terminalId())))
.count();
Map<String, Object> row = new LinkedHashMap<>();
row.put("profile", profile); row.put("maxLoad", cap); row.put("live", live);
row.put("free", cap == null ? null : Math.max(0, cap - live)); row.put("reclaimable", reclaimable);
return row;
}
/**
* One lead's row: its address, its name, and whether it can be reached right now.
*
@@ -44,23 +44,6 @@ public final class ClaudeCodeLauncher extends HerdrPeerLauncher {
private final SubscriptionGuard guard;
/**
* Standing instruction appended to the worker's system prompt so it returns its result via
* {@code bridge_reply}. Injected as a launch flag, so nothing is written to the worker's
* profile — it is guidance, and a worker that never replies is caught by the send's timeout.
*/
static final String REPLY_CHARTER =
"You are an off-subscription worker in the claude-bridge fleet. Every message you "
+ "receive arrives through the bridge, and the ONLY channel back to the sender is the "
+ "bridge_reply MCP tool. Text you write in your terminal is NOT sent anywhere — the "
+ "sender cannot see your screen, so an in-terminal answer is silently discarded. "
+ "Therefore you MUST end EVERY turn by calling bridge_reply with `content` set to your "
+ "complete response. This holds for every message without exception — tasks, questions, "
+ "clarifications, acknowledgements, and ordinary back-and-forth conversation. Call "
+ "bridge_reply exactly once, as the final action of your turn, with your full answer in "
+ "`content`; never wait for confirmation first. If you end a turn without calling "
+ "bridge_reply, the sender receives nothing and the exchange stalls.";
/**
* Production constructor — disables the spawn-ready gate ({@code spawnReadyTimeoutMs == 0}) so
* existing deployments and tests keep the legacy non-blocking spawn semantics.
@@ -93,11 +76,11 @@ public final class ClaudeCodeLauncher extends HerdrPeerLauncher {
Map<String, BridgedConfig.Profile> profiles, String defaultProfile,
Function<String, String> env,
long spawnReadyTimeoutMs, long spawnReadyPollMs,
Supplier<String> tabLabelTemplate) {
Supplier<BridgedConfig.Fleet> fleet) {
this(agents, spaces, guard, profiles, defaultProfile, env,
spawnReadyTimeoutMs,
System::currentTimeMillis, () -> sleepUninterruptibly(spawnReadyPollMs),
tabLabelTemplate);
fleet);
}
/**
@@ -129,31 +112,19 @@ public final class ClaudeCodeLauncher extends HerdrPeerLauncher {
/**
* Full testability constructor, plus the fleet-wide tab-label template (CB-557).
*
* @param tabLabelTemplate {@code fleet.tabLabel}; {@code null}/blank ⇒
* {@link BridgedConfig.Fleet#DEFAULT_TAB_LABEL}
* @param fleet live fleet config, read once for each spawn
*/
public ClaudeCodeLauncher(AgentControl agents, WorkspaceControl spaces, SubscriptionGuard guard,
Map<String, BridgedConfig.Profile> profiles, String defaultProfile,
Function<String, String> env,
long spawnReadyTimeoutMs,
LongSupplier nowMillis, Runnable sleeper,
Supplier<String> tabLabelTemplate) {
Supplier<BridgedConfig.Fleet> fleet) {
super(NAME_PREFIX, agents, spaces, profiles, defaultProfile, env,
spawnReadyTimeoutMs, nowMillis, sleeper, tabLabelTemplate);
spawnReadyTimeoutMs, nowMillis, sleeper, fleet);
this.guard = guard;
}
/**
* {@inheritDoc}
*
* <p>A legacy spawn with no session identity is a fresh, launcher-derived session — delegate to
* the session-aware form with no name and no resume id.
*/
@Override
protected Launch buildLaunch(BridgedConfig.Profile cfg) {
return buildLaunch(cfg, null, null);
}
/**
* {@inheritDoc}
*
@@ -164,7 +135,7 @@ public final class ClaudeCodeLauncher extends HerdrPeerLauncher {
* applied here — see {@link #applySessionIdentity}.
*/
@Override
protected Launch buildLaunch(BridgedConfig.Profile cfg, String sessionName, String resumeSessionId) {
protected Launch buildLaunch(BridgedConfig.Profile cfg, LaunchSpec spec) {
// CB-539: a profile may deliberately opt into the subscription (subscription: true) when no
// off-subscription endpoint exists for it — e.g. `sonnet` on `ccs`. That profile gets no
// ANTHROPIC_BASE_URL/AUTH_TOKEN (there is nothing to point them at) and the guard's base_url
@@ -210,9 +181,9 @@ public final class ClaudeCodeLauncher extends HerdrPeerLauncher {
// id via -r and passes no --session-id (the two conflict). Both are injected before the
// model flag so --model keeps outranking the operator's own argv.
// mutableArgv: argvWithBridge may hand back the profile's own (immutable) List.of when it
// has no MCP — session flags must be added into a list we own.
List<String> argv = mutableArgv(argvWithBridge(cfg));
String agentSessionId = applySessionIdentity(argv, sessionName, resumeSessionId);
// has neither MCP nor a charter — session flags must be added into a list we own.
List<String> argv = mutableArgv(argvWithBridge(cfg, spec.charter()));
String agentSessionId = applySessionIdentity(argv, spec.sessionName(), spec.resumeSessionId());
return new Launch(workerEnv, argvWithModel(argv, cfg), agentSessionId);
}
@@ -246,22 +217,26 @@ public final class ClaudeCodeLauncher extends HerdrPeerLauncher {
}
/**
* The launch argv, plus — when {@code worker.mcpUrl} is set — inline {@code --mcp-config} for
* the bridge server and {@code --append-system-prompt} for the {@link #REPLY_CHARTER}. Neither
* touches the profile's config; both are pure command-line flags. This inline-flag mount is
* Claude Code specific — other adapters mount MCP and instructions their own way.
* The launch argv, plus an inline {@code --mcp-config} when {@code worker.mcpUrl} is set and
* {@code --append-system-prompt} when the base composed a charter. Neither touches the profile's
* config; both are pure command-line flags. This inline-flag mount is Claude Code specific —
* other adapters mount MCP and instructions their own way.
*/
private List<String> argvWithBridge(BridgedConfig.Profile cfg) {
if (!cfg.hasMcp()) {
private List<String> argvWithBridge(BridgedConfig.Profile cfg, String charter) {
if (!cfg.hasMcp() && charter == null) {
return cfg.argv();
}
String mcpJson = "{\"mcpServers\":{\"bridge\":{\"type\":\"http\",\"url\":\""
+ cfg.mcpUrl() + "\"}}}";
List<String> argv = mutableArgv(cfg.argv());
argv.add("--mcp-config");
argv.add(mcpJson);
argv.add("--append-system-prompt");
argv.add(REPLY_CHARTER);
if (cfg.hasMcp()) {
String mcpJson = "{\"mcpServers\":{\"bridge\":{\"type\":\"http\",\"url\":\""
+ cfg.mcpUrl() + "\"}}}";
argv.add("--mcp-config");
argv.add(mcpJson);
}
if (charter != null) {
argv.add("--append-system-prompt");
argv.add(charter);
}
return argv;
}
@@ -44,7 +44,7 @@ import java.util.regex.Pattern;
* <li>{@code namePrefix} (constructor arg) — the label prefix ({@code claude}, {@code opencode})
* that drives both unique naming and the orphan-reap pattern, so each adapter reaps only its
* own kind of pane and never another's.</li>
* <li>{@link #buildLaunch(BridgedConfig.Profile)} — the peer-specific env map + argv, including any
* <li>{@link #buildLaunch(BridgedConfig.Profile, LaunchSpec)} — the peer-specific env map + argv, including any
* subscription/guard check, MCP mount, and instruction injection. The base never sees how the
* peer is configured; it only places and starts the returned {@link Launch}.</li>
* </ul>
@@ -80,15 +80,25 @@ public abstract class HerdrPeerLauncher implements PeerLauncher {
private final AtomicLong nameSeq = new AtomicLong(); // per-peer counter (herdr agent names only)
/**
* The {@code fleet.tabLabel} template; a {@code null} supplier or a {@code null}/blank value ⇒
* {@link BridgedConfig.Fleet#DEFAULT_TAB_LABEL}. A profile's own {@code tabLabel} still
* overrides it.
* Live fleet config, read once per spawn. A null supplier or value leaves tab labels at their
* default and supplies no role charter. A profile's own {@code tabLabel} still overrides it.
*
* <p>CB-559: a supplier rather than a String, so a config reload renames the <em>next</em> tab
* without a restart. Existing tabs keep the label they were given — bridged does not rewrite a
* label it already wrote.
* <p>CB-559: a supplier rather than a snapshot, so a config reload affects the next launch
* without a restart. Existing tabs keep the label they were given.
*/
private final Supplier<String> tabLabelTemplate;
private final Supplier<BridgedConfig.Fleet> fleet;
/** The final instruction always requires a bridge reply when the bridge MCP is mounted. */
protected static final String REPLY_CHARTER =
"You are a spawned member in the claude-bridge fleet. Every message you receive arrives "
+ "through the bridge, and the ONLY channel back to the sender is the bridge_reply MCP tool. "
+ "Text you write in your terminal is NOT sent anywhere — the sender cannot see your screen, "
+ "so an in-terminal answer is silently discarded. Therefore you MUST end EVERY turn by calling "
+ "bridge_reply with `content` set to your complete response. This holds for every message without "
+ "exception — tasks, questions, clarifications, acknowledgements, and ordinary back-and-forth "
+ "conversation. Call bridge_reply exactly once, as the final action of your turn, with your full "
+ "answer in `content`; never wait for confirmation first. If you end a turn without calling "
+ "bridge_reply, the sender receives nothing and the exchange stalls.";
/**
* Tab numbers, counted per {@code role/profile} pair (CB-557).
@@ -142,21 +152,19 @@ public abstract class HerdrPeerLauncher implements PeerLauncher {
}
/**
* As above, plus the {@code fleet.tabLabel} template (CB-557).
* As above, plus the live {@code fleet} config (CB-557).
*
* @param tabLabelTemplate fleet-wide tab-label template, read per spawn (CB-559); {@code null},
* or a supplier yielding {@code null}/blank ⇒
* {@link BridgedConfig.Fleet#DEFAULT_TAB_LABEL}. A separate constructor
* rather than a new parameter on the one above, so every existing call
* site keeps the default without an edit.
* @param fleet live fleet config, read once per spawn; {@code null} ⇒ default tab label and no
* role charter. A separate constructor rather than a new parameter on the one
* above, so every existing call site keeps the default without an edit.
*/
protected HerdrPeerLauncher(String namePrefix, AgentControl agents, WorkspaceControl spaces,
Map<String, BridgedConfig.Profile> profiles, String defaultProfile,
Function<String, String> env,
long spawnReadyTimeoutMs,
LongSupplier nowMillis, Runnable sleeper,
Supplier<String> tabLabelTemplate) {
this.tabLabelTemplate = tabLabelTemplate;
Supplier<BridgedConfig.Fleet> fleet) {
this.fleet = fleet;
this.namePrefix = namePrefix;
this.agents = agents;
this.spaces = spaces;
@@ -175,22 +183,7 @@ public abstract class HerdrPeerLauncher implements PeerLauncher {
* Any subscription/guard check, MCP mount, and instruction injection happen here. The env map
* and argv are adapter-private; the base only places and starts what is returned.
*/
protected abstract Launch buildLaunch(BridgedConfig.Profile cfg);
/**
* Session-aware variant of {@link #buildLaunch(BridgedConfig.Profile)} (CB-547a). Default
* discards the session identity and delegates to the profile-only form, so an adapter that
* carries no durable peer session (opencode, say) inherits byte-identical behaviour and needs
* no change. An adapter that does (Claude Code) overrides this to mint/resume the id and to
* surface it on the returned {@link Launch#agentSessionId()}.
*
* @param cfg the resolved profile to spawn
* @param sessionName the bridge's logical session name, or null/blank for launcher-derived
* @param resumeSessionId the peer's own prior session id to resume, or null/blank for fresh
*/
protected Launch buildLaunch(BridgedConfig.Profile cfg, String sessionName, String resumeSessionId) {
return buildLaunch(cfg);
}
protected abstract Launch buildLaunch(BridgedConfig.Profile cfg, LaunchSpec spec);
/** Direct transport access for peer-specific, non-turn control operations. */
protected final AgentControl agents() {
@@ -225,6 +218,10 @@ public abstract class HerdrPeerLauncher implements PeerLauncher {
}
}
/** All per-spawn values adapters may need, including the base-composed effective charter. */
protected record LaunchSpec(String sessionName, String resumeSessionId, MemberRole role, String charter) {
}
// --- profile surface -----------------------------------------------------------------------
/** The configured peer profile names (what {@code spawn(profile)} accepts). */
@@ -295,18 +292,23 @@ public abstract class HerdrPeerLauncher implements PeerLauncher {
}
/**
* Spawn a peer with session identity (CB-547a). {@code sessionName} and {@code resumeSessionId}
* are threaded from the {@link SpawnRequest} into {@link #buildLaunch(BridgedConfig.Profile,
* String, String)}, and the launch's resolved agent-session id is returned alongside the agent
* so the caller can put it on the {@link PeerHandle}.
* Spawn a peer with session identity (CB-547a). The session values, role, and charter are
* threaded from the {@link SpawnRequest} into {@link #buildLaunch(BridgedConfig.Profile,
* LaunchSpec)}, and the launch's resolved agent-session id is returned alongside the agent so
* the caller can put it on the {@link PeerHandle}.
*/
protected Spawned spawnInternal(String profileName, String requestedCwd, String callerCwd,
String sessionName, String resumeSessionId, MemberRole role) {
BridgedConfig.Profile cfg = requireProfile(profileName);
Launch launch = buildLaunch(cfg, sessionName, resumeSessionId);
BridgedConfig.Fleet liveFleet = fleet == null ? null : fleet.get();
String roleCharter = liveFleet == null ? null : liveFleet.charterFor(role);
String replyCharter = cfg.hasMcp() ? REPLY_CHARTER : null;
String charter = roleCharter == null ? replyCharter
: replyCharter == null ? roleCharter : roleCharter + "\n\n" + replyCharter;
Launch launch = buildLaunch(cfg, new LaunchSpec(sessionName, resumeSessionId, role, charter));
String cwd = resolveCwd(requestedCwd, cfg, callerCwd);
Agent agent = cfg.tabPlacement()
? spawnInTab(cfg, launch.env(), launch.argv(), cwd, role)
? spawnInTab(cfg, launch.env(), launch.argv(), cwd, role, liveFleet)
: spawnAsPane(cfg, launch.env(), launch.argv(), cwd);
return new Spawned(agent, launch.agentSessionId());
}
@@ -382,7 +384,7 @@ public abstract class HerdrPeerLauncher implements PeerLauncher {
/** Dedicated worker space → own tab (carrying cwd+env) → start the peer into the seed pane. */
private Agent spawnInTab(BridgedConfig.Profile cfg, Map<String, String> workerEnv,
List<String> argv, String cwd, MemberRole role) {
List<String> argv, String cwd, MemberRole role, BridgedConfig.Fleet liveFleet) {
Workspace space = spaces.ensureWorkspace(cfg.workspace());
Tab.Created tab = spaces.createTab(space.workspaceId(), cwd, workerEnv);
log.info("spawning {} profile={} space={} tab={} cwd={}",
@@ -415,7 +417,7 @@ public abstract class HerdrPeerLauncher implements PeerLauncher {
tidy("label tab " + tab.tab().tabId(),
() -> spaces.renameTab(tab.tab().tabId(),
cfg.renderTabLabel(
tabLabelTemplate == null ? null : tabLabelTemplate.get(),
liveFleet == null ? null : liveFleet.tabLabel(),
role, nextLabelSeq(role, cfg.profile()))));
log.info("{} started pane={} tab={} terminal={}",
namePrefix, started.agent().paneId(), started.agent().tabId(), started.agent().terminalId());
@@ -38,7 +38,7 @@ import java.util.function.Supplier;
* <li><strong>File-based MCP mount + instructions.</strong> opencode has no inline
* {@code --mcp-config}/{@code --append-system-prompt}. Instead the bridge writes an ephemeral
* {@code opencode.json} that declares the bridge as a {@code remote} MCP server and lists a
* reply-charter file under {@code instructions}, then points the worker at it with
* member-charter file under {@code instructions}, then points the worker at it with
* {@code OPENCODE_CONFIG}. This is the one place the launcher touches disk — Claude never did.</li>
* <li><strong>Model as a flag.</strong> the {@code provider/model} selector is passed as
* {@code -m}, not an env var.</li>
@@ -54,38 +54,9 @@ public final class OpenCodeLauncher extends HerdrPeerLauncher {
/** Writer for the generated {@code opencode.json}. */
private static final ObjectMapper JSON = new ObjectMapper();
/**
* Standing instruction written to the charter file and mounted via the config's
* {@code instructions} so the worker returns its result through {@code bridge_reply}. Kept on
* disk (not a launch flag) because opencode's {@code instructions} takes file paths, not inline
* text — the file is regenerated per spawn and never touches the worker's own profile.
*/
static final String REPLY_CHARTER =
"You are an off-subscription worker in the claude-bridge fleet, running under opencode. "
+ "Every message you receive arrives through the bridge, and the ONLY channel back to the "
+ "sender is the bridge_reply MCP tool. Text you write in your terminal is NOT sent "
+ "anywhere — the sender cannot see your screen, so an in-terminal answer is silently "
+ "discarded. Therefore you MUST end EVERY turn by calling bridge_reply with `content` set "
+ "to your complete response. This holds for every message without exception — tasks, "
+ "questions, clarifications, acknowledgements, and ordinary back-and-forth conversation. "
+ "Call bridge_reply exactly once, as the final action of your turn, with your full answer "
+ "in `content`; never wait for confirmation first. If you end a turn without calling "
+ "bridge_reply, the sender receives nothing and the exchange stalls.";
/** Root under which per-spawn opencode config dirs are created (injectable for tests). */
private final Path configRoot;
/**
* The current spawn's resume-target session id, threaded from {@link #spawn(SpawnRequest)} to
* {@link #buildLaunch} across the base's {@code spawn -> spawnInternal -> buildLaunch} chain,
* which carries no request. A plain field would race under concurrent spawns (the base supports
* them), so it is thread-local: each spawn captures its own request's id on its own thread, and
* {@code buildLaunch}, synchronous and same-thread, reads exactly that one. Set only around the
* {@code super.spawn} call and cleared in {@code finally}, so a paused/leftover value can never
* bleed into the next spawn.
*/
private final ThreadLocal<String> resumeSessionId = new ThreadLocal<>();
/**
* Session discovery against opencode's on-disk storage ({@link OpenCodeSessionDiscovery}) —
* the one seam that knows opencode's private session-file layout. Its root is injectable for
@@ -126,10 +97,10 @@ public final class OpenCodeLauncher extends HerdrPeerLauncher {
Map<String, BridgedConfig.Profile> profiles, String defaultProfile,
Function<String, String> env,
long spawnReadyTimeoutMs, long spawnReadyPollMs,
Supplier<String> tabLabelTemplate) {
Supplier<BridgedConfig.Fleet> fleet) {
this(agents, spaces, profiles, defaultProfile, env, spawnReadyTimeoutMs,
System::currentTimeMillis, () -> sleepUninterruptibly(spawnReadyPollMs),
defaultConfigRoot(), defaultDiscoveryRoot(), tabLabelTemplate);
defaultConfigRoot(), defaultDiscoveryRoot(), fleet);
}
/**
@@ -164,8 +135,7 @@ public final class OpenCodeLauncher extends HerdrPeerLauncher {
/**
* Full testability constructor, plus the fleet-wide tab-label template (CB-557).
*
* @param tabLabelTemplate {@code fleet.tabLabel}; {@code null}/blank ⇒
* {@link BridgedConfig.Fleet#DEFAULT_TAB_LABEL}
* @param fleet live fleet config, read once for each spawn
*/
public OpenCodeLauncher(AgentControl agents, WorkspaceControl spaces,
Map<String, BridgedConfig.Profile> profiles, String defaultProfile,
@@ -173,9 +143,9 @@ public final class OpenCodeLauncher extends HerdrPeerLauncher {
long spawnReadyTimeoutMs,
LongSupplier nowMillis, Runnable sleeper,
Path configRoot, Path discoveryRoot,
Supplier<String> tabLabelTemplate) {
Supplier<BridgedConfig.Fleet> fleet) {
super(NAME_PREFIX, agents, spaces, profiles, defaultProfile, env,
spawnReadyTimeoutMs, nowMillis, sleeper, tabLabelTemplate);
spawnReadyTimeoutMs, nowMillis, sleeper, fleet);
this.configRoot = configRoot;
this.discovery = new OpenCodeSessionDiscovery(discoveryRoot);
}
@@ -193,20 +163,21 @@ public final class OpenCodeLauncher extends HerdrPeerLauncher {
* {@inheritDoc}
*
* <p>Builds the opencode launch: no {@code ANTHROPIC_*} and no guard (opencode reads its own
* provider credentials); when the profile mounts the bridge MCP, generate an ephemeral
* {@code opencode.json} (remote MCP server + reply-charter instructions) and point the worker at
* it via {@code OPENCODE_CONFIG}; carry the parity-neutral git-forge grant; and select the model
* with {@code -m}.
* provider credentials); when the profile mounts the bridge MCP or has a member charter,
* generate an ephemeral {@code opencode.json} (remote MCP server + member-charter instructions)
* and point the worker at it via {@code OPENCODE_CONFIG}; carry the parity-neutral git-forge
* grant; and select the model with {@code -m}.
*/
@Override
protected Launch buildLaunch(BridgedConfig.Profile cfg) {
protected Launch buildLaunch(BridgedConfig.Profile cfg, LaunchSpec spec) {
Map<String, String> workerEnv = baseEnv(cfg);
// A config file is needed for the bridge MCP mount, for a pinned endpoint (CB-508), or both.
if (cfg.hasMcp() || hasCustomProvider(cfg)) {
workerEnv.put("OPENCODE_CONFIG", writeConfig(cfg).toString());
// A config file is needed for the bridge MCP mount, a member charter, or a pinned endpoint (CB-508).
if (cfg.hasMcp() || spec.charter() != null || hasCustomProvider(cfg)) {
workerEnv.put("OPENCODE_CONFIG", writeConfig(cfg, spec.charter()).toString());
}
applyGitToken(workerEnv, cfg);
return new Launch(workerEnv, argvWithResume(argvWithModel(argvWithAuto(cfg), cfg)));
return new Launch(workerEnv,
argvWithResume(argvWithModel(argvWithAuto(cfg), cfg), spec.resumeSessionId()));
}
/**
@@ -243,11 +214,9 @@ public final class OpenCodeLauncher extends HerdrPeerLauncher {
* The launch argv plus, on a resumed spawn, opencode's {@code -s <id>} flag to continue a prior
* conversation by its session id. {@code -s, --session <id>} resumes an existing session; on a
* fresh spawn (no resume target) no flag is added, letting opencode start a brand-new session.
* The id comes from the current spawn request's {@code resumeSessionId}, threaded per-thread by
* {@link #spawn(SpawnRequest)}.
* The id comes from the base launch spec.
*/
private List<String> argvWithResume(List<String> argv) {
String id = resumeSessionId.get();
private List<String> argvWithResume(List<String> argv, String id) {
if (id == null || id.isBlank()) {
return argv;
}
@@ -267,12 +236,12 @@ public final class OpenCodeLauncher extends HerdrPeerLauncher {
}
/**
* Write an ephemeral {@code opencode.json} (and the reply-charter file it references) into a
* Write an ephemeral {@code opencode.json} (and the member-charter file it references) into a
* fresh per-spawn directory under {@link #configRoot}, and return the config file's path for
* {@code OPENCODE_CONFIG}. The dir is unique per spawn so concurrent workers never race on it;
* it is best-effort cleaned on JVM exit (worker config is disposable — regenerated every spawn).
*/
private Path writeConfig(BridgedConfig.Profile cfg) {
private Path writeConfig(BridgedConfig.Profile cfg, String charterText) {
try {
Path dir = Files.createTempDirectory(configRoot, "bridged-opencode-");
dir.toFile().deleteOnExit();
@@ -295,16 +264,19 @@ public final class OpenCodeLauncher extends HerdrPeerLauncher {
// if per-profile control is ever wanted, add a profile knob rather than dropping this.
root.putObject("compaction").put("auto", true);
if (cfg.hasMcp()) {
Path charter = dir.resolve("reply-charter.md");
Files.writeString(charter, REPLY_CHARTER);
if (charterText != null) {
Path charter = dir.resolve("member-charter.md");
Files.writeString(charter, charterText);
charter.toFile().deleteOnExit();
root.putArray("instructions").add(charter.toAbsolutePath().toString());
}
if (cfg.hasMcp()) {
ObjectNode bridge = root.putObject("mcp").putObject("bridge");
bridge.put("type", "remote");
bridge.put("url", cfg.mcpUrl());
bridge.put("enabled", true);
root.putArray("instructions").add(charter.toAbsolutePath().toString());
}
if (hasCustomProvider(cfg)) {
addCustomProvider(root, cfg);
@@ -380,31 +352,11 @@ public final class OpenCodeLauncher extends HerdrPeerLauncher {
return afterScheme.contains("/") ? trimmed : trimmed + "/v1";
}
/**
* {@inheritDoc}
*
* <p>adds this adapter's session-identity work around the base's spawn — as opencode cannot be
* told its session id at spawn (see {@link Capability#SESSION_RESUME} vs
* {@link Capability#SESSION_NAME}), identity is only ever adopted after the fact:
* <ul>
* <li>the request's {@code resumeSessionId} is remembered for {@link #buildLaunch} to turn
* into {@code -s <id>}; and</li>
* <li>the returned handle is wrapped so its
* {@link dev.ltms.bridged.peer.PeerHandle#agentSessionId()} performs lazy session
* discovery against opencode's storage (see {@link OpenCodeSessionDiscovery}) — always
* non-blocking, {@code null} until opencode has persisted the session record.</li>
* </ul>
*/
/** Add lazy on-disk session discovery to the base handle. */
@Override
public PeerHandle spawn(SpawnRequest req) {
resumeSessionId.set(req.resumeSessionId());
try {
PeerHandle inner = super.spawn(req);
return new SessionAwareHandle(inner, discovery, effectiveCwd(req));
} finally {
// Never let a paused/leftover resume id bleed into the next spawn on this thread.
resumeSessionId.remove();
}
PeerHandle inner = super.spawn(req);
return new SessionAwareHandle(inner, discovery, effectiveCwd(req));
}
/**
@@ -9,6 +9,7 @@ import org.slf4j.Logger;
import org.slf4j.LoggerFactory;
import java.util.List;
import java.util.Set;
import java.util.UUID;
import java.util.concurrent.CompletableFuture;
import java.util.concurrent.CompletionException;
@@ -127,6 +128,8 @@ public final class MessageService {
public enum Phase {
/** Delegated and in flight — queued for the worker or being worked. */
PENDING,
/** The worker is paused in {@code bridge_ask}; {@link TaskView#reply} and {@link TaskView#turnId} identify it. */
ASKING,
/** The worker's turn finished; {@link TaskView#reply} holds the answer. */
DONE,
/** The delegation could not complete (timed out, worker gone, or busy). */
@@ -136,16 +139,28 @@ public final class MessageService {
/**
* A poll snapshot of an async delegation.
*
* @param reply the answer when {@link #phase} is {@link Phase#DONE}, else {@code null}
* @param reply the answer when {@link #phase} is {@link Phase#DONE}, or the question when
* {@link #phase} is {@link Phase#ASKING}; otherwise {@code null}
* @param replySource {@code "reply"} (structured {@code bridge_reply}) or {@code "transcript"}
* (completion scrape) when {@link Phase#DONE}, else {@code null}
* @param detail a human note (live worker status while pending, or the failure reason)
* @param detail a human note (live worker status while pending, ask state, or failure reason)
* @param turnId correlation id for an {@link Phase#ASKING} ticket, else {@code null}
*/
public record TaskView(String ticket, Phase phase, String reply, String replySource, String detail) {
public record TaskView(String ticket, Phase phase, String reply, String replySource, String detail,
String turnId) {
}
/** An in-flight or finished async delegation, keyed by its ticket. */
private record Task(String target, CompletableFuture<Reply> future, long createdNanos) {
private static final class Task {
private final String target;
private final CompletableFuture<Reply> future = new CompletableFuture<>();
private final long createdNanos = System.nanoTime();
private volatile Reply question;
private volatile String turnId;
private Task(String target) {
this.target = target;
}
}
private final AgentControl agents;
@@ -156,6 +171,13 @@ public final class MessageService {
private final Metrics metrics; // CB-502: nullable — no registry in unit tests
private final ConcurrentHashMap<String, ReentrantLock> sessionLocks = new ConcurrentHashMap<>();
private final ConcurrentHashMap<String, Task> tasks = new ConcurrentHashMap<>();
/** Async tasks that have accepted delivery for a target. */
private final ConcurrentHashMap<String, Set<Task>> asyncTasksByTarget = new ConcurrentHashMap<>();
/** Async task that owns each exact forward rendezvous waiter. */
private final ConcurrentHashMap<CompletableFuture<Rendezvous.Resolution>, Task> asyncTasksByWaiter =
new ConcurrentHashMap<>();
/** Async tickets paused on a specific {@code bridge_ask} turn. */
private final ConcurrentHashMap<String, Task> asyncTasksByTurn = new ConcurrentHashMap<>();
private final AtomicLong ticketSeq = new AtomicLong();
private final ExecutorService asyncExecutor = Executors.newThreadPerTaskExecutor(
Thread.ofVirtual().name("bridge-async-", 0).factory());
@@ -202,6 +224,16 @@ public final class MessageService {
return agents.status(target);
}
/** Read-only delegation fact for fleet views. */
public boolean hasAcceptedDelivery(String target) {
return rendezvous.isWaiting(target);
}
/** Read-only inbox fact for fleet views. */
public boolean hasInboxMessage(String target) {
return !inbox.peek(target).isEmpty();
}
/**
* Route a worker's explicit {@code bridge_reply}: resolve an open send, or queue it in the
* inbox if no send is currently open. Unlike the bare {@link Rendezvous#resolve}, a no-waiter
@@ -272,14 +304,18 @@ public final class MessageService {
*/
public boolean abandon(String target, String reason) {
CompletableFuture<Rendezvous.Resolution> waiter = rendezvous.currentWaiter(target);
if (waiter == null || waiter.isDone()) {
return false; // nobody is blocked on this worker — nothing to abandon
boolean failed = waiter != null && !waiter.isDone() && rendezvous.resolveFailure(waiter, reason);
boolean asyncFailed = false;
for (Task task : tasks.values()) {
if (target.equals(task.target) && task.question == null
&& task.future.complete(new Reply(Outcome.WORKER_FAILED, reason))) {
asyncFailed = true;
}
}
boolean failed = rendezvous.resolveFailure(waiter, reason);
if (failed) {
log.debug("abandoned send to {}: {}", target, reason);
log.warn("abandoning the blocked send to {}: {}", target, reason);
}
return failed;
return failed || asyncFailed;
}
/**
@@ -326,6 +362,11 @@ public final class MessageService {
* never earned. {@code null} disables the hook.
*/
public Reply send(String target, String content, long timeoutMillis, Runnable onAccepted) {
return send(target, content, timeoutMillis, onAccepted, null);
}
/** Run a send, optionally stopping an async task that teardown already failed before acceptance. */
private Reply send(String target, String content, long timeoutMillis, Runnable onAccepted, Task task) {
long deadlineNanos = System.nanoTime() + timeoutMillis * 1_000_000L;
ReentrantLock lock = sessionLocks.computeIfAbsent(target, _ -> new ReentrantLock());
@@ -333,6 +374,12 @@ public final class MessageService {
return new Reply(Outcome.BUSY, null); // another send held the session the whole window
}
try {
if (task != null && task.future.isDone()) {
return task.future.getNow(null);
}
if (hasAsyncQuestion(target)) {
return new Reply(Outcome.BUSY, null); // the worker's current turn is paused for its lead
}
// Open the waiter BEFORE queueing delivery (CB-548). A fast reply — the worker already
// injectable the instant we enqueue — otherwise arrives before the waiter is registered
// and orphans into the inbox while this send blocks to the timeout (the enqueue-before-
@@ -341,6 +388,9 @@ public final class MessageService {
// failed send leaves no stale waiter behind.
CompletableFuture<Rendezvous.Resolution> reply = rendezvous.open(target);
try {
if (task != null) {
asyncTasksByWaiter.put(reply, task);
}
// The send has won the lock; the accepted-delivery hook records delegator ownership
// here (CB-548). It runs BEFORE enqueue so a throwing hook — onAccepted is now a
// public callback — fails the send without queuing a message that would orphan.
@@ -364,6 +414,7 @@ public final class MessageService {
throw new IllegalStateException("interrupted awaiting reply from " + target, e);
}
} finally {
asyncTasksByWaiter.remove(reply);
rendezvous.close(target, reply);
}
} finally {
@@ -389,7 +440,12 @@ public final class MessageService {
if (ticket.fresh()) {
// Register the reverse waiter first, then surface the question — so the answer, which can
// arrive the instant the primary reacts, always finds an open waiter to resolve.
CompletableFuture<Rendezvous.Resolution> waiter = rendezvous.currentWaiter(workerSession);
Task task = markAsyncQuestion(waiter, question, ticket.turnId());
if (!rendezvous.resolveQuestion(workerSession, question, ticket.turnId())) {
if (task != null) {
clearAsyncQuestion(ticket.turnId(), true);
}
rendezvous.closeAsk(ticket.turnId());
return new AskResult(AskOutcome.NO_WAITER, null); // no primary is blocked on this worker
}
@@ -399,6 +455,7 @@ public final class MessageService {
return new AskResult(AskOutcome.ANSWERED, answer);
} catch (TimeoutException e) {
log.debug("bridge_ask from {} went unanswered in {}ms", workerSession, timeoutMillis);
clearAsyncQuestion(ticket.turnId(), true);
return new AskResult(AskOutcome.TIMED_OUT, null);
} catch (ExecutionException e) {
Throwable cause = e.getCause();
@@ -441,9 +498,12 @@ public final class MessageService {
rendezvous.close(workerSession, reply);
return new Reply(Outcome.STALE_TURN, null); // lapsed between the lookup and the unblock
}
clearAsyncQuestion(turnId, false);
try {
Rendezvous.Resolution r = reply.get(remainingMillis(deadlineNanos), TimeUnit.MILLISECONDS);
return new Reply(outcomeOf(r.kind()), r.text(), r.turnId());
Reply result = new Reply(outcomeOf(r.kind()), r.text(), r.turnId());
finishAsyncTask(turnId, result);
return result;
} catch (TimeoutException e) {
// The worker resumed but hasn't replied yet — no completion fallback arms an answered
// turn (it never re-entered the injector), so a silent worker rides out the window.
@@ -483,9 +543,28 @@ public final class MessageService {
*/
public String sendAsync(String target, String content, Runnable onAccepted) {
String ticket = "task-" + ticketSeq.incrementAndGet();
CompletableFuture<Reply> future = CompletableFuture.supplyAsync(
() -> send(target, content, ASYNC_TIMEOUT_MS, onAccepted), asyncExecutor);
tasks.put(ticket, new Task(target, future, System.nanoTime()));
Task task = new Task(target);
tasks.put(ticket, task);
asyncExecutor.submit(() -> {
try {
Runnable trackingAccepted = () -> {
if (onAccepted != null) {
onAccepted.run();
}
asyncTasksByTarget.computeIfAbsent(target, _ -> ConcurrentHashMap.newKeySet()).add(task);
};
Reply result = send(target, content, ASYNC_TIMEOUT_MS, trackingAccepted, task);
if (result.outcome() == Outcome.QUESTION) {
// Keep the accepted owner until answer() finishes it. markAsyncQuestion may run
// just after resolveQuestion wakes this thread.
} else {
finishAsyncTask(task, result);
}
} catch (Throwable t) {
task.future.completeExceptionally(t);
untrackAsyncTarget(task);
}
});
pruneTerminalTickets();
log.debug("async send {} -> {}", ticket, target);
return ticket;
@@ -501,27 +580,32 @@ public final class MessageService {
if (task == null) {
return null;
}
CompletableFuture<Reply> f = task.future();
CompletableFuture<Reply> f = task.future;
if (!f.isDone()) {
return new TaskView(ticket, Phase.PENDING, null, null, "worker " + liveStatus(task.target()));
Reply question = task.question;
if (question != null) {
return new TaskView(ticket, Phase.ASKING, question.text(), null,
"worker is waiting for your answer", question.turnId());
}
return new TaskView(ticket, Phase.PENDING, null, null, "worker " + liveStatus(task.target), null);
}
Reply r;
try {
r = f.getNow(null);
} catch (CompletionException | java.util.concurrent.CancellationException e) {
Throwable cause = (e instanceof CompletionException ce && ce.getCause() != null) ? ce.getCause() : e;
return new TaskView(ticket, Phase.FAILED, null, null, cause.getMessage());
return new TaskView(ticket, Phase.FAILED, null, null, cause.getMessage(), null);
}
if (r.completed()) {
String source = r.outcome() == Outcome.REPLIED ? "reply" : "transcript";
return new TaskView(ticket, Phase.DONE, r.text(), source, null);
return new TaskView(ticket, Phase.DONE, r.text(), source, null, null);
}
// A wedged worker (CB-109) carries the error context as its reason; the timeout/busy
// outcomes carry none, so fall back to the outcome name.
String detail = r.outcome() == Outcome.WORKER_FAILED && r.text() != null
? r.text()
: "no reply — " + r.outcome().name().toLowerCase();
return new TaskView(ticket, Phase.FAILED, null, null, detail);
return new TaskView(ticket, Phase.FAILED, null, null, detail, null);
}
/** Best-effort live worker status for a pending poll; never throws (a lookup error is just noise). */
@@ -536,7 +620,64 @@ public final class MessageService {
/** Drop finished tickets older than the TTL so the registry cannot grow without bound. */
private void pruneTerminalTickets() {
long cutoff = System.nanoTime() - TICKET_TTL_NANOS;
tasks.values().removeIf(t -> t.future().isDone() && t.createdNanos() < cutoff);
tasks.values().removeIf(t -> t.future.isDone() && t.createdNanos < cutoff);
}
/** Record the active question for an async ticket; blocking sends have no entry and stay unchanged. */
private Task markAsyncQuestion(CompletableFuture<Rendezvous.Resolution> waiter, String text, String turnId) {
Task task = waiter == null ? null : asyncTasksByWaiter.get(waiter);
if (task != null) {
task.question = new Reply(Outcome.QUESTION, text, turnId);
task.turnId = turnId;
asyncTasksByTurn.put(turnId, task);
}
return task;
}
/** Clear an answered or lapsed question, but only when it matches the ticket's current turn. */
private void clearAsyncQuestion(String turnId, boolean forgetTurn) {
Task task = asyncTasksByTurn.get(turnId);
if (task != null && turnId.equals(task.turnId)) {
task.question = null;
if (forgetTurn) {
asyncTasksByTurn.remove(turnId, task);
task.turnId = null;
untrackAsyncTarget(task);
}
}
}
/** Complete and detach an async ticket after its worker's actual terminal reply. */
private void finishAsyncTask(Task task, Reply result) {
task.future.complete(result);
untrackAsyncTarget(task);
if (task.turnId != null) {
asyncTasksByTurn.remove(task.turnId, task);
}
}
/** Complete the async ticket correlated to a specific answered turn. */
private void finishAsyncTask(String turnId, Reply result) {
Task task = asyncTasksByTurn.get(turnId);
if (task != null) {
finishAsyncTask(task, result);
}
}
/** A new send must not open a waiter while an async ticket owns this worker's paused turn. */
private boolean hasAsyncQuestion(String target) {
return asyncTasksByTurn.values().stream().anyMatch(task -> target.equals(task.target));
}
/** Stop tracking a task once it no longer owns an accepted target turn. */
private void untrackAsyncTarget(Task task) {
Set<Task> targetTasks = asyncTasksByTarget.get(task.target);
if (targetTasks != null) {
targetTasks.remove(task);
if (targetTasks.isEmpty()) {
asyncTasksByTarget.remove(task.target, targetTasks);
}
}
}
/** Release the async executor. */
@@ -1,5 +1,6 @@
package dev.ltms.bridged.session;
import dev.ltms.bridged.auth.MemberLifecycle;
import dev.ltms.bridged.herdr.Agent;
import dev.ltms.bridged.inject.TurnListener;
import dev.ltms.bridged.inject.MemberPresence;
@@ -28,8 +29,7 @@ import java.util.function.LongSupplier;
* teardown on top.
*
* <p>The state machine is intentionally one-shot / no-reuse: every acquired worker is fresh,
* and a finished or released worker is torn down, never pooled. {@link #recycle} is a convenience
* for {@code release + acquire} with a new distinct pane id.
* and a finished or released worker is torn down, never pooled or reused.
*
* <p>The manager implements {@link TurnListener} so the injector's turn boundaries drive
* {@code READY → BUSY → DONE} (or {@code FAILED}). It exposes a {@link MemberPresence} view via
@@ -49,6 +49,7 @@ public final class SessionManager implements TurnListener {
private final LongSupplier nowNanos;
private final int contextCap;
private final boolean clearAfterTurn;
private volatile MemberLifecycle memberLifecycle = MemberLifecycle.NONE;
/** CB-520: notified with a terminalId on every acquire; no-op until wired. */
private final List<Consumer<String>> acquireListeners = new java.util.concurrent.CopyOnWriteArrayList<>();
@@ -139,7 +140,13 @@ public final class SessionManager implements TurnListener {
// to pick the profile out of that role's pool and to label the tab; a role kept only on
// the MemberSession is recorded after the spawn it was supposed to steer.
SpawnRequest req = new SpawnRequest(profile, requestedCwd, callerCwd, null, null, memberRole);
PeerHandle handle = launcher.spawn(req);
PeerHandle handle;
try {
handle = launcher.spawn(req);
} catch (RuntimeException e) {
log.warn("spawn failed for profile={} role={}: {}", profile, memberRole, e.getMessage());
throw e;
}
String resolvedProfile = resolveProfile(handle, profile);
String cwd = launcher.effectiveCwd(new SpawnRequest(resolvedProfile, requestedCwd, callerCwd));
long now = nowNanos.getAsLong();
@@ -157,6 +164,7 @@ public final class SessionManager implements TurnListener {
null,
null);
registry.put(handle.id(), session);
memberLifecycle.acquired(session.role(), session.profile(), session.terminalId());
log.debug("acquired session id={} terminal={} profile={} owner={}",
handle.id(), handle.terminalId(), session.profile(), session.ownerTerminal());
notifyAcquired(session.terminalId());
@@ -177,7 +185,7 @@ public final class SessionManager implements TurnListener {
* <p>CB-544: these are two concerns that used to be fused. Stopping the pane is correct on every
* teardown — the worker process must end. Removing the worktree is a destructive act that is only
* correct for a deliberately-finished teardown (an explicit stop of a completed session, the
* reaper releasing a genuinely idle one, a context-capped or recycled session). A shutdown drain
* reaper releasing a genuinely idle one, or a context-capped session). A shutdown drain
* must stop panes but preserve worktrees: a worker's uncommitted work exists in exactly one
* place — its worktree — so deleting it while the daemon simply goes down is silent data loss,
* with no copy and no error. Do NOT fuse these back together; the cost of an orphaned worktree
@@ -187,6 +195,7 @@ public final class SessionManager implements TurnListener {
MemberSession removed = registry.remove(paneId);
boolean preserveWorktree = cause == ReleaseCause.SHUTDOWN;
if (removed != null) {
memberLifecycle.released(removed.terminalId());
log.debug("releasing session pane={} terminal={} state={} cause={}",
removed.paneId(), removed.terminalId(), removed.state(), cause);
if (preserveWorktree && removed.worktree() != null) {
@@ -244,7 +253,7 @@ public final class SessionManager implements TurnListener {
/**
* Register a callback invoked with a session's {@code terminalId} whenever it is released
* (CB-516). Every teardown path funnels through {@link #release}, so one hook covers the REST
* and MCP stop tools, the idle-TTL reaper, {@code recycle}, and shutdown drain alike.
* and MCP stop tools, the idle-TTL reaper, and shutdown drain alike.
*
* <p>Added rather than injected because {@code MessageService} — one intended listener — is
* constructed after this manager (it needs the injector and rendezvous, which need the session
@@ -257,6 +266,11 @@ public final class SessionManager implements TurnListener {
}
}
/** Inject the optional member-slot lifecycle after construction without changing constructors. */
public void setMemberLifecycle(MemberLifecycle memberLifecycle) {
this.memberLifecycle = memberLifecycle == null ? MemberLifecycle.NONE : memberLifecycle;
}
/** A listener failure must never prevent the acquisition it is reacting to. */
private void notifyAcquired(String terminalId) {
if (terminalId == null) {
@@ -304,6 +318,8 @@ public final class SessionManager implements TurnListener {
worktrees.overlayParity(repoRoot, path, launcher.parityOverlay(preResolvedProfile));
handle = launcher.spawn(new SpawnRequest(profile, path, callerCwd, null, null, memberRole));
} catch (RuntimeException e) {
log.warn("spawn failed for profile={} role={} branch={} path={}: {}",
preResolvedProfile, memberRole, branch, path, e.getMessage());
if (path != null) {
try {
worktrees.remove(repoRoot, path);
@@ -330,6 +346,7 @@ public final class SessionManager implements TurnListener {
path,
branch);
registry.put(handle.id(), session);
memberLifecycle.acquired(session.role(), session.profile(), session.terminalId());
log.debug("acquired worktree session id={} terminal={} profile={} branch={} path={}",
handle.id(), handle.terminalId(), session.profile(), session.branch(), session.worktree());
notifyAcquired(session.terminalId());
@@ -360,19 +377,6 @@ public final class SessionManager implements TurnListener {
return launcher.defaultProfile();
}
/**
* Release the old session and acquire a fresh one with the same profile and working directory.
* The new session is guaranteed to have a pane id distinct from the old one (no-reuse invariant).
*/
public MemberSession recycle(String paneId) {
MemberSession old = registry.get(paneId);
if (old == null) {
throw new IllegalArgumentException("no session for paneId " + paneId);
}
release(paneId);
return acquire(old.profile(), old.cwd(), old.cwd(), old.ownerTerminal());
}
/** The session for {@code paneId}, if it is still registered and not released. */
public Optional<MemberSession> get(String paneId) {
return Optional.ofNullable(registry.get(paneId));
@@ -488,8 +492,10 @@ public final class SessionManager implements TurnListener {
MemberSession current = findByTerminal(target);
if (current == null) return;
if (current.state() == MemberSession.State.RELEASED) return;
MemberSession.State priorState = current.state();
if (replace(current, current.withState(MemberSession.State.FAILED))) {
log.debug("session marked failed terminal={} pane={}", target, current.paneId());
log.warn("member terminal={} pane={} can no longer be delegated to: its turn never resolved "
+ "(was {} when it failed)", target, current.paneId(), priorState);
}
}
@@ -505,7 +511,11 @@ public final class SessionManager implements TurnListener {
if (s.state() != MemberSession.State.READY && s.state() != MemberSession.State.DONE) {
continue;
}
if (now - s.lastActivityAtNanos() > idleTtlNanos) {
long idleNanos = now - s.lastActivityAtNanos();
if (idleNanos > idleTtlNanos) {
log.debug("reaping idle session terminal={} pane={}: idle {}s exceeds the {}s ttl",
s.terminalId(), s.paneId(), TimeUnit.NANOSECONDS.toSeconds(idleNanos),
TimeUnit.NANOSECONDS.toSeconds(idleTtlNanos));
release(s.paneId());
reaped++;
}
+9
View File
@@ -1,4 +1,13 @@
<configuration>
<!--
CB-575: MCP SDK 2.0.0 has no public notification registration API. It registers only
notifications/initialized and notifications/roots/list_changed, so other client notifications
still warn when unhandled. Clients may legitimately send notifications/cancelled; suppress only
that SDK WARN because it would devalue the action-needed WARN level used by M4 fleet health.
If a later SDK handles cancellation, this filter simply stops matching and can be removed.
-->
<turboFilter class="dev.ltms.bridged.logging.McpCancelledNotificationFilter"/>
<appender name="STDOUT" class="ch.qos.logback.core.ConsoleAppender">
<encoder>
<pattern>%d{HH:mm:ss.SSS} %-5level [%thread] %logger{28} - %msg%n</pattern>
@@ -3,6 +3,8 @@ package dev.ltms.bridged.auth;
import dev.ltms.bridged.herdr.FakeHerdr;
import dev.ltms.bridged.herdr.PaneLocator;
import dev.ltms.bridged.mcp.ConnectionIdentity;
import dev.ltms.bridged.config.BridgedConfig;
import dev.ltms.bridged.peer.MemberRole;
import org.junit.jupiter.api.Test;
import java.util.Map;
@@ -32,6 +34,17 @@ class CallerResolverTest {
return identity(999_999);
}
private static MemberRegistry boundMembers(String slot, MemberRole role) {
Map<String, BridgedConfig.Slot> architects = role == MemberRole.ARCHITECT
? Map.of("lead-designer", new BridgedConfig.Slot("sonnet")) : Map.of();
Map<String, BridgedConfig.Slot> devs = role == MemberRole.DEV
? Map.of("builder", new BridgedConfig.Slot("sonnet")) : Map.of();
MemberRegistry members = new MemberRegistry(
new BridgedConfig.Fleet(Map.of(), architects, devs, Map.of(), null));
assertTrue(members.bind(slot, "term_a"));
return members;
}
@Test
void aLoopbackWorkerPaneResolvesToWorkerRegardlessOfAuthMode() {
Principal underTrust = new CallerResolver(workerIdentity()).resolve("127.0.0.1", 42, null);
@@ -302,8 +315,9 @@ class CallerResolverTest {
@Test
void aBoundArchitectPaneResolvesToArchitectBeforeTheWorkerFallback() {
// This is the production construction path used by Bridged.
Principal p = CallerResolver.withLeadsAndMembers(workerIdentity(), false, null,
Map::of, () -> Map.of("term_a", "lead-designer"))
Map::of, boundMembers("architect:lead-designer", MemberRole.ARCHITECT))
.resolve("127.0.0.1", 42, null);
assertEquals(Role.ARCHITECT, p.role(),
@@ -316,7 +330,7 @@ class CallerResolverTest {
@Test
void anArchitectNeedsNoTokenEvenInTokenMode() {
Principal p = CallerResolver.withLeadsAndMembers(workerIdentity(), true, "s3cret",
Map::of, () -> Map.of("term_a", "lead-designer"))
Map::of, boundMembers("architect:lead-designer", MemberRole.ARCHITECT))
.resolve("127.0.0.1", 42, null);
assertEquals(Role.ARCHITECT, p.role(),
@@ -325,9 +339,10 @@ class CallerResolverTest {
@Test
void anUnboundPaneStillResolvesAsAWorker() {
Map<String, String> arch = Map.of("term_elsewhere", "reviewer");
MemberRegistry members = new MemberRegistry(new BridgedConfig.Fleet(Map.of(),
Map.of("lead-designer", new BridgedConfig.Slot("sonnet")), Map.of(), Map.of(), null));
Principal p = CallerResolver.withLeadsAndMembers(workerIdentity(), false, null,
Map::of, () -> arch).resolve("127.0.0.1", 42, null);
Map::of, members).resolve("127.0.0.1", 42, null);
assertEquals(Role.WORKER, p.role());
assertNull(p.name());
@@ -337,7 +352,8 @@ class CallerResolverTest {
@Test
void aLeadWinsOverAnArchitectBindingForTheSamePane() {
Principal p = CallerResolver.withLeadsAndMembers(workerIdentity(), false, null,
() -> Map.of("term_a", "opus-5.0"), () -> Map.of("term_a", "lead-designer"))
() -> Map.of("term_a", "opus-5.0"),
boundMembers("architect:lead-designer", MemberRole.ARCHITECT))
.resolve("127.0.0.1", 42, null);
assertEquals(Role.PRIMARY, p.role(),
@@ -349,26 +365,17 @@ class CallerResolverTest {
/** The registry is live, like leads: a binding injected after construction is honoured. */
@Test
void anArchitectBoundAfterConstructionIsHonouredWithoutRebuildingTheResolver() {
Map<String, String> live = new java.util.HashMap<>();
MemberRegistry members = new MemberRegistry(new BridgedConfig.Fleet(Map.of(),
Map.of("lead-designer", new BridgedConfig.Slot("sonnet")), Map.of(), Map.of(), null));
CallerResolver r = CallerResolver.withLeadsAndMembers(workerIdentity(), false, null,
Map::of, () -> live);
Map::of, members);
assertEquals(Role.WORKER, r.resolve("127.0.0.1", 42, null).role());
live.put("term_a", "lead-designer"); // the later lifecycle binds the slot
assertTrue(members.bind("architect:lead-designer", "term_a")); // the later lifecycle binds the slot
assertEquals(Role.ARCHITECT, r.resolve("127.0.0.1", 42, null).role());
assertEquals("lead-designer", r.members().get("term_a"));
}
@Test
void theArchitectMapFormIsCopiedSoLaterMutationCannotGrantArchitect() {
Map<String, String> mutable = new java.util.LinkedHashMap<>();
CallerResolver r = new CallerResolver(workerIdentity(), false, null, Map.of(), mutable);
mutable.put("term_a", "sneaky");
assertEquals(Role.WORKER, r.resolve("127.0.0.1", 42, null).role());
assertEquals("architect:lead-designer", r.members().get("term_a"));
}
@Test
@@ -381,7 +388,8 @@ class CallerResolverTest {
@Test
void anArchitectOwnsItsOwnPaneAndNoOther() {
Principal arch = CallerResolver.withLeadsAndMembers(workerIdentity(), false, null,
Map::of, () -> Map.of("term_a", "lead-designer")).resolve("127.0.0.1", 42, null);
Map::of, boundMembers("architect:lead-designer", MemberRole.ARCHITECT))
.resolve("127.0.0.1", 42, null);
assertTrue(arch.ownsSession("term_a"));
assertTrue(Authz.permits(arch, Authz.Action.REPLY, "term_a"));
@@ -389,6 +397,15 @@ class CallerResolverTest {
assertFalse(Authz.permits(arch, Authz.Action.REPLY, "term_b"));
}
@Test
void aBoundNonArchitectSlotStillResolvesAsAWorker() {
Principal p = CallerResolver.withLeadsAndMembers(workerIdentity(), false, null,
Map::of, boundMembers("dev:builder", MemberRole.DEV))
.resolve("127.0.0.1", 42, null);
assertEquals(Role.WORKER, p.role(), "a dev binding must never grant architect rights");
}
@Test
void tokenModeRequiresANonEmptyConfiguredToken() {
ConnectionIdentity id = nonWorkerIdentity();
@@ -111,6 +111,28 @@ class BridgedConfigTest {
assertDoesNotThrow(() -> BridgedConfig.load(f));
}
@Test
void absentChartersRemainValidAndPresentChartersUseRoleWireNames(@TempDir Path dir) throws Exception {
Path absent = dir.resolve("absent.yaml");
Files.writeString(absent, "fleet: {}\n");
BridgedConfig withoutCharters = BridgedConfig.load(absent);
assertDoesNotThrow(withoutCharters::validateCharters);
assertNull(withoutCharters.fleet().charterFor(MemberRole.ARCHITECT));
Path blank = dir.resolve("blank.yaml");
Files.writeString(blank, "fleet:\n charters:\n architect: ' '\n");
IllegalStateException blankError = assertThrows(IllegalStateException.class,
() -> BridgedConfig.load(blank).validateCharters());
assertTrue(blankError.getMessage().contains("fleet.charters.architect is blank"));
Path unknown = dir.resolve("unknown.yaml");
Files.writeString(unknown, "fleet:\n charters:\n architetc: text\n");
IllegalStateException unknownError = assertThrows(IllegalStateException.class,
() -> BridgedConfig.load(unknown).validateCharters());
assertTrue(unknownError.getMessage().contains("architetc"));
assertTrue(unknownError.getMessage().contains("[architect, dev, reviewer]"));
}
/**
* CB-530. Unknown keys stay ignored — config must be allowed to run ahead of the code — but they
* must be NAMED at load. A whole block that parses, is dropped, and is never mentioned again is
@@ -1167,6 +1189,41 @@ class BridgedConfigTest {
"a profile without the key stays off-subscription (the default)");
}
@Test
void parityOverlayDefaultsToEnvFilesOnly(@TempDir Path dir) throws Exception {
Path f = dir.resolve("no-overlay.yaml");
Files.writeString(f, """
bind:
port: 8080
profiles:
gx10:
baseUrl: http://gx10.gw:8000
""");
BridgedConfig cfg = BridgedConfig.load(f);
assertEquals(List.of(".env", ".envrc"),
cfg.profiles().get("gx10").parityOverlay(),
"the default parity overlay is the env files; settings.local.json is no longer copied by default");
}
@Test
void parityOverlayExplicitListIsPreservedVerbatim(@TempDir Path dir) throws Exception {
Path f = dir.resolve("explicit-overlay.yaml");
Files.writeString(f, """
bind:
port: 8080
profiles:
gx10:
baseUrl: http://gx10.gw:8000
parityOverlay: [".claude/settings.local.json", ".env"]
""");
BridgedConfig cfg = BridgedConfig.load(f);
assertEquals(List.of(".claude/settings.local.json", ".env"),
cfg.profiles().get("gx10").parityOverlay(),
"an operator's explicit list survives verbatim — the default only changes when unset");
}
// ── CB-542: subscription:true must not smuggle an unguarded endpoint via env: ───────────────
@Test
@@ -66,6 +66,64 @@ class ConfigRefTest {
assertEquals("[{profile}] {role}", ref.get().fleet().tabLabel());
}
@Test
void aCharterChangeIsHotAndReachesTheLiveConfig(@TempDir Path dir) throws Exception {
Path f = dir.resolve("bridged.yaml");
Files.writeString(f, yaml("""
fleet:
charters:
architect: old charter
"""));
ConfigRef ref = refFor(f);
assertEquals("old charter", ref.get().fleet().charterFor(
dev.ltms.bridged.peer.MemberRole.ARCHITECT));
Files.writeString(f, yaml("""
fleet:
charters:
architect: new charter
"""));
ConfigRef.Outcome out = ref.reload();
assertTrue(out.applied());
assertTrue(out.deferred().isEmpty());
assertEquals("new charter", ref.get().fleet().charterFor(
dev.ltms.bridged.peer.MemberRole.ARCHITECT));
}
@Test
void invalidChartersRefuseReloadAndKeepTheRunningConfig(@TempDir Path dir) throws Exception {
Path f = dir.resolve("bridged.yaml");
Files.writeString(f, yaml("""
fleet:
charters:
architect: valid charter
"""));
ConfigRef ref = refFor(f);
BridgedConfig before = ref.get();
Files.writeString(f, yaml("""
fleet:
charters:
architect: " "
"""));
ConfigRef.Outcome blank = ref.reload();
assertFalse(blank.applied());
assertTrue(blank.error().contains("fleet.charters.architect is blank"));
assertSame(before, ref.get());
Files.writeString(f, yaml("""
fleet:
charters:
architetc: valid charter
"""));
ConfigRef.Outcome unknown = ref.reload();
assertFalse(unknown.applied());
assertTrue(unknown.error().contains("architetc"));
assertTrue(unknown.error().contains("architect"));
assertSame(before, ref.get());
}
/**
* The point of the whole class: a consumer holding the ref sees the new value without being
* rebuilt. A component that captured {@code get()} into a field would still show the old one.
@@ -0,0 +1,81 @@
package dev.ltms.bridged.health;
import ch.qos.logback.classic.Logger;
import ch.qos.logback.classic.spi.ILoggingEvent;
import ch.qos.logback.core.read.ListAppender;
import dev.ltms.bridged.herdr.AgentControl;
import dev.ltms.bridged.herdr.FakeHerdr;
import dev.ltms.bridged.inject.Injector;
import dev.ltms.bridged.msg.InMemoryReplyInbox;
import dev.ltms.bridged.msg.MessageService;
import dev.ltms.bridged.msg.Rendezvous;
import dev.ltms.bridged.session.SessionManager;
import dev.ltms.bridged.member.ClaudeCodeLauncher;
import dev.ltms.bridged.herdr.WorkspaceControl;
import dev.ltms.bridged.guard.SubscriptionGuard;
import dev.ltms.bridged.config.BridgedConfig;
import org.junit.jupiter.api.Test;
import org.slf4j.LoggerFactory;
import java.util.Map;
import java.util.Set;
import java.util.concurrent.Executors;
import static org.junit.jupiter.api.Assertions.assertEquals;
class FleetHealthMonitorTest {
@Test void oneTickUsesOneFleetListForAnyRosterSize() {
FakeHerdr herdr = new FakeHerdr();
AgentControl agents = new AgentControl(herdr);
BridgedConfig.Profile profile = new BridgedConfig.Profile("test", "http://test:1", null,
null, null, null, null, null, null, null, null, null);
ClaudeCodeLauncher launcher = new ClaudeCodeLauncher(agents, new WorkspaceControl(herdr),
new SubscriptionGuard(Set.of("test")), Map.of("test", profile), "test", _ -> "token");
SessionManager sessions = new SessionManager(launcher);
sessions.acquire("test", null, null, null);
sessions.acquire("test", null, null, null);
herdr.calls.clear();
MessageService messages = new MessageService(agents, new Injector(agents), new Rendezvous(), new InMemoryReplyInbox());
var scheduler = Executors.newSingleThreadScheduledExecutor();
FleetHealthMonitor monitor = new FleetHealthMonitor(agents, sessions::roster, messages, scheduler, () -> 1, 60);
monitor.tick();
monitor.stop();
assertEquals(1, herdr.calls.stream().filter(call -> call.method().equals("agent.list")).count());
}
@Test void failedTickDoesNotStopTheNextTick() {
FakeHerdr herdr = new FakeHerdr().healthy(false);
AgentControl agents = new AgentControl(herdr);
var scheduler = Executors.newSingleThreadScheduledExecutor();
FleetHealthMonitor monitor = new FleetHealthMonitor(agents, java.util.List::of,
new MessageService(agents, new Injector(agents), new Rendezvous(), new InMemoryReplyInbox()),
scheduler, () -> 1, 60);
monitor.tick();
herdr.healthy(true);
monitor.tick();
monitor.stop();
assertEquals(2, herdr.calls.stream().filter(call -> call.method().equals("agent.list")).count());
}
@Test void faultTransitionLogsOnlyOnceUntilItChanges() {
Logger logger = (Logger) LoggerFactory.getLogger(FleetHealthMonitor.class);
ListAppender<ILoggingEvent> appender = new ListAppender<>();
appender.start();
logger.addAppender(appender);
try {
FakeHerdr herdr = new FakeHerdr();
AgentControl agents = new AgentControl(herdr);
var scheduler = Executors.newSingleThreadScheduledExecutor();
FleetHealthMonitor monitor = new FleetHealthMonitor(agents, java.util.List::of,
new MessageService(agents, new Injector(agents), new Rendezvous(), new InMemoryReplyInbox()),
scheduler, () -> 1, 60);
monitor.reportTransition("term_a", HealthState.TURN_BOUNDARY_LOST);
monitor.reportTransition("term_a", HealthState.TURN_BOUNDARY_LOST);
monitor.stop();
assertEquals(1, appender.list.stream().filter(event -> event.getFormattedMessage()
.contains("member=term_a state=TURN_BOUNDARY_LOST")).count());
} finally {
logger.detachAppender(appender);
}
}
}
@@ -0,0 +1,56 @@
package dev.ltms.bridged.health;
import dev.ltms.bridged.herdr.AgentStatus;
import dev.ltms.bridged.session.MemberSession;
import dev.ltms.bridged.msg.Rendezvous;
import org.junit.jupiter.api.Test;
import static org.junit.jupiter.api.Assertions.assertEquals;
class FleetHealthTest {
@Test void muteCountsOnlyCompletionFallbacks() {
MuteCounter mute = new MuteCounter();
mute.observe("target", "terra", Rendezvous.Kind.REPLY);
mute.observe("target", "terra", Rendezvous.Kind.COMPLETION);
assertEquals(1, mute.forTarget("target"));
assertEquals(1, mute.forProfile("terra"));
}
@Test void turnBoundaryNeedsTwoSnapshots() {
HealthSnapshot s = snapshot(MemberSession.State.BUSY, AgentStatus.DONE, true);
HealthDecision first = FleetHealth.decide(s, HealthPrior.NONE, 1);
assertEquals(HealthState.WORKING, first.state());
assertEquals(new HealthPrior(true), first.prior());
assertEquals(HealthState.TURN_BOUNDARY_LOST, FleetHealth.decide(s, first.prior(), 2).state());
}
@Test void unknownLiveStatusWithAcceptedDeliveryIsNotIdle() {
assertEquals(HealthState.WORKING, FleetHealth.decide(
snapshot(MemberSession.State.BUSY, AgentStatus.UNKNOWN, true), HealthPrior.NONE, 1).state());
}
@Test void acceptedDeliveryNeverReportsIdle() {
for (MemberSession.State session : MemberSession.State.values()) {
for (AgentStatus live : AgentStatus.values()) {
HealthSnapshot s = snapshot(session, live, true);
assertEquals(false, FleetHealth.decide(s, HealthPrior.NONE, 1).state() == HealthState.IDLE,
() -> "accepted delivery returned IDLE for " + session + "/" + live);
}
}
}
@Test void blockedDoesNotGuessPromptKind() {
assertEquals(HealthState.BLOCKED_AMBIGUOUS, FleetHealth.decide(
snapshot(MemberSession.State.BUSY, AgentStatus.BLOCKED, true), HealthPrior.NONE, 1).state());
}
@Test void controlLinkOutranksMemberFault() {
HealthSnapshot s = new HealthSnapshot(MemberSession.State.BUSY, AgentStatus.DONE, true, false,
false, true, true, true, false, true, true, true);
assertEquals(HealthState.CONTROL_LINK_DOWN, FleetHealth.decide(s, HealthPrior.NONE, 1).state());
}
private static HealthSnapshot snapshot(MemberSession.State state, AgentStatus live, boolean accepted) {
return new HealthSnapshot(state, live, accepted, false, false, true, false, false,
false, false, false, false);
}
}
@@ -0,0 +1,14 @@
package dev.ltms.bridged.health;
import org.junit.jupiter.api.Test;
import java.util.List;
import static org.junit.jupiter.api.Assertions.assertEquals;
class PaneBudgetTest {
@Test void fixedLimitsIgnoreWeakerConfig() {
PaneBudget budget = new PaneBudget();
List<String> targets = List.of("a", "b", "c");
assertEquals(List.of("a", "b"), budget.choose(targets, 0, 0));
assertEquals(List.of("c"), budget.choose(targets, 1, 0));
}
}
@@ -1,9 +1,14 @@
package dev.ltms.bridged.inject;
import ch.qos.logback.classic.Level;
import ch.qos.logback.classic.LoggerContext;
import ch.qos.logback.classic.spi.ILoggingEvent;
import ch.qos.logback.core.read.ListAppender;
import dev.ltms.bridged.herdr.AgentControl;
import dev.ltms.bridged.herdr.FakeHerdr;
import dev.ltms.bridged.msg.Rendezvous;
import org.junit.jupiter.api.Test;
import org.slf4j.LoggerFactory;
import static org.junit.jupiter.api.Assertions.assertEquals;
import static org.junit.jupiter.api.Assertions.assertFalse;
@@ -156,6 +161,33 @@ class CompletionResolverTest {
assertEquals("No, 391 = 17 × 23.", waiter.getNow(null).text());
}
@Test
void marksAClippedCompletionPaneTail() {
String block = "⏺ " + "x".repeat(CompletionResolver.MAX_SCRAPE_CHARS + 1) + "\n❯ ";
FakeHerdr herdr = new FakeHerdr().readText(block);
Rendezvous rendezvous = new Rendezvous();
CompletionResolver resolver = new CompletionResolver(new AgentControl(herdr), rendezvous);
var waiter = rendezvous.open("term_a");
resolver.resolve("term_a", new CompletionResolver.InFlight(waiter, null));
assertEquals("x".repeat(CompletionResolver.MAX_SCRAPE_CHARS)
+ "\n[Pane tail clipped: member did not call bridge_reply.]",
waiter.getNow(null).text());
}
@Test
void leavesAnUnclippedCompletionPaneTailUnmarked() {
FakeHerdr herdr = new FakeHerdr().readText("⏺ complete report\n❯ ");
Rendezvous rendezvous = new Rendezvous();
CompletionResolver resolver = new CompletionResolver(new AgentControl(herdr), rendezvous);
var waiter = rendezvous.open("term_a");
resolver.resolve("term_a", new CompletionResolver.InFlight(waiter, null));
assertEquals("complete report", waiter.getNow(null).text());
}
@Test
void resolvesSynchronouslyBeforePostTurnContextClearing() {
FakeHerdr herdr = new FakeHerdr().readText("⏺ previous answer\n❯ ");
@@ -177,7 +209,8 @@ class CompletionResolverTest {
// block while resolve compares against a clip()'d tail. For a block longer than MAX_SCRAPE_CHARS
// the two capped representations differ even when the pane never changed, so the CB-115
// byte-identical guard failed to fire and a stale completion could resolve the send. Both sides
// must clip identically; here an unchanged >cap block on rapid back-to-back turns stays suppressed.
// must clip identically. The returned-text marker is added only after this comparison, so an
// unchanged >cap block on rapid back-to-back turns still stays suppressed.
String longBlock = "⏺ " + "x".repeat(CompletionResolver.MAX_SCRAPE_CHARS + 500) + "\n❯ ";
FakeHerdr herdr = new FakeHerdr().readText(longBlock);
Rendezvous rendezvous = new Rendezvous();
@@ -270,6 +303,40 @@ class CompletionResolverTest {
assertEquals("stuck on an error screen", waiter.getNow(null).text());
}
@Test
void failIsLoggedAtWarnWithTheReason() {
// CB-564: this used to be a bare DEBUG "failed send to X via turn-stall fallback" — a symptom
// with no cause, and below the level anyone watching for member health would see. A fail that
// resolves a caller's blocked send is at least WARN and must carry the reason.
LoggerContext ctx = (LoggerContext) LoggerFactory.getILoggerFactory();
ch.qos.logback.classic.Logger resolverLog =
(ch.qos.logback.classic.Logger) LoggerFactory.getLogger(CompletionResolver.class);
ListAppender<ILoggingEvent> appender = new ListAppender<>();
appender.setContext(ctx);
appender.start();
resolverLog.addAppender(appender);
resolverLog.setLevel(Level.WARN);
try {
FakeHerdr herdr = new FakeHerdr().readText("stuck on an error screen");
Rendezvous rendezvous = new Rendezvous();
CompletionResolver resolver = new CompletionResolver(new AgentControl(herdr), rendezvous);
var waiter = rendezvous.open("term_a");
resolver.fail("term_a", null);
String warn = appender.list.stream()
.filter(e -> e.getLevel().equals(Level.WARN))
.map(ILoggingEvent::getFormattedMessage)
.findFirst()
.orElse("no turn-stall WARN logged");
assertTrue(warn.contains("term_a"), "the log names the target: " + warn);
assertTrue(warn.contains("stuck on an error screen"), "the log carries the reason: " + warn);
assertTrue(waiter.isDone());
} finally {
resolverLog.detachAppender(appender);
}
}
// --- CB-116 waiter identity: a late completion never crosses into the next turn ---------
@Test
@@ -1,10 +1,15 @@
package dev.ltms.bridged.inject;
import ch.qos.logback.classic.Level;
import ch.qos.logback.classic.LoggerContext;
import ch.qos.logback.classic.spi.ILoggingEvent;
import ch.qos.logback.core.read.ListAppender;
import dev.ltms.bridged.herdr.AgentControl;
import dev.ltms.bridged.herdr.AgentStatus;
import dev.ltms.bridged.herdr.FakeHerdr;
import dev.ltms.bridged.herdr.HerdrException;
import org.junit.jupiter.api.Test;
import org.slf4j.LoggerFactory;
import java.util.ArrayList;
import java.util.List;
@@ -254,6 +259,7 @@ class InjectorTest {
private static final class Captor implements TurnListener {
final List<String> completed = new ArrayList<>();
final List<String> failed = new ArrayList<>();
final List<String> failureReasons = new ArrayList<>();
@Override
public void onTurnComplete(String target) {
@@ -264,6 +270,12 @@ class InjectorTest {
public void onTurnFailed(String target) {
failed.add(target);
}
@Override
public void onTurnFailed(String target, String reason) {
failed.add(target);
failureReasons.add(reason);
}
}
// ~30s of unknown at the 250ms prod poll interval; enough onStatus samples to trip the stall.
@@ -317,6 +329,23 @@ class InjectorTest {
assertTrue(f.isCompletedExceptionally(), "queued waiters unblock when the worker vanishes");
}
@Test
void dropPassesTheRealCauseForQueuedAndDeliveredWork() {
Captor cap = new Captor();
Injector inj = new Injector(new AgentControl(herdr), cap);
CompletableFuture<Void> delivered = inj.enqueue(T, "delivered");
CompletableFuture<Void> queued = inj.enqueue(T, "queued");
inj.onStatus(T, AgentStatus.IDLE); // deliver the first message
inj.onStatus(T, AgentStatus.WORKING); // its turn is now in flight; one remains queued
inj.drop(T, new HerdrException("agent target sol not found", "agent_not_found", null));
assertEquals(List.of(T), cap.failed, "drop signals one turn failure for both affected states");
assertEquals(List.of("agent target sol not found"), cap.failureReasons);
assertTrue(delivered.isDone(), "the delivered future has already completed");
assertTrue(queued.isCompletedExceptionally(), "the queued future fails with the drop cause");
}
@Test
void dropFailsTheTurnOfADeliveredMessageWhenTheWorkerVanishes() {
// CB-110: the message was delivered (no longer queued), so failing queued waiters alone would
@@ -369,6 +398,40 @@ class InjectorTest {
assertTrue(inj.activeTargets().isEmpty(), "the target is reclaimed, not polled forever");
}
@Test
void readinessGraceExpiryIsLogged() {
// CB-562: the grace-expiry path used to clear the queue silently, so a message that never
// reached the worker's pane surfaced elsewhere as an unrelated turn-stall failure. Assert the
// expiry now names the real cause. (ListAppender capture pattern mirrors AuditLogTest.)
LoggerContext ctx = (LoggerContext) LoggerFactory.getILoggerFactory();
ch.qos.logback.classic.Logger injectorLog =
(ch.qos.logback.classic.Logger) LoggerFactory.getLogger(Injector.class);
ListAppender<ILoggingEvent> appender = new ListAppender<>();
appender.setContext(ctx);
appender.start();
injectorLog.addAppender(appender);
injectorLog.setLevel(Level.WARN);
try {
Injector inj = new Injector(new AgentControl(herdr), TurnListener.NOOP, _ -> false, _ -> {
});
inj.enqueue(T, "task");
for (int i = 0; i < READINESS_SAMPLES; i++) inj.onStatus(T, AgentStatus.IDLE);
String warn = appender.list.stream()
.filter(e -> e.getLevel().equals(Level.WARN))
.map(ILoggingEvent::getFormattedMessage)
.findFirst()
.orElse("no grace-expiry WARN logged");
assertTrue(warn.contains(T), "the log names the target terminal: " + warn);
assertTrue(warn.contains("never reached"), "the log names the real cause: " + warn);
assertTrue(warn.contains("1 queued message"),
"the log carries the failed message count: " + warn);
} finally {
injectorLog.detachAppender(appender);
}
}
@Test
void aWorkerThatBecomesReadyWithinTheGraceIsDeliveredNormally() {
// The readiness grace must not fail a worker that is merely slow to boot: once it becomes
@@ -397,6 +460,38 @@ class InjectorTest {
assertEquals(List.of(T), forgotten, "drop clears the gone worker's presence");
}
@Test
void dropIsLogged() {
// CB-564: a vanished worker used to drop its queue with no log at all — the only trace was
// whatever failed downstream (e.g. a caller's send timing out with no clue why). Assert the
// drop itself now names the cause and the number of messages it failed.
LoggerContext ctx = (LoggerContext) LoggerFactory.getILoggerFactory();
ch.qos.logback.classic.Logger injectorLog =
(ch.qos.logback.classic.Logger) LoggerFactory.getLogger(Injector.class);
ListAppender<ILoggingEvent> appender = new ListAppender<>();
appender.setContext(ctx);
appender.start();
injectorLog.addAppender(appender);
injectorLog.setLevel(Level.WARN);
try {
Injector inj = new Injector(new AgentControl(herdr), TurnListener.NOOP, _ -> true, _ -> {
});
inj.enqueue(T, "orphan");
inj.drop(T, new HerdrException("worker gone", "pane_not_found", null));
String warn = appender.list.stream()
.filter(e -> e.getLevel().equals(Level.WARN))
.map(ILoggingEvent::getFormattedMessage)
.findFirst()
.orElse("no drop WARN logged");
assertTrue(warn.contains(T), "the log names the target terminal: " + warn);
assertTrue(warn.contains("1 message"), "the log carries the failed message count: " + warn);
assertTrue(warn.contains("worker gone"), "the log carries the real cause: " + warn);
} finally {
injectorLog.detachAppender(appender);
}
}
@Test
void pollerDeliversToAnIdleWorker() throws Exception {
// End-to-end through the poller: idle worker → message delivered without manual onStatus.
@@ -0,0 +1,37 @@
package dev.ltms.bridged.logging;
import ch.qos.logback.classic.Level;
import ch.qos.logback.classic.Logger;
import ch.qos.logback.classic.spi.ILoggingEvent;
import ch.qos.logback.core.read.ListAppender;
import io.modelcontextprotocol.spec.McpSchema;
import org.junit.jupiter.api.Test;
import org.slf4j.LoggerFactory;
import java.util.Map;
import static org.junit.jupiter.api.Assertions.assertEquals;
import static org.junit.jupiter.api.Assertions.assertTrue;
class McpCancelledNotificationFilterTest {
@Test
void suppressesOnlyTheCancelledNotificationWarning() {
Logger logger = (Logger) LoggerFactory.getLogger(McpCancelledNotificationFilter.LOGGER);
ListAppender<ILoggingEvent> appender = new ListAppender<>();
appender.start();
logger.addAppender(appender);
try {
logger.warn(McpCancelledNotificationFilter.UNHANDLED_NOTIFICATION,
new McpSchema.JSONRPCNotification("notifications/cancelled", Map.of("requestId", 7)));
logger.warn(McpCancelledNotificationFilter.UNHANDLED_NOTIFICATION,
new McpSchema.JSONRPCNotification("notifications/progress", Map.of("progress", 1)));
assertEquals(1, appender.list.size());
assertEquals(Level.WARN, appender.list.getFirst().getLevel());
assertTrue(appender.list.getFirst().getFormattedMessage().contains("notifications/progress"));
} finally {
logger.detachAppender(appender);
}
}
}
@@ -2,6 +2,7 @@ package dev.ltms.bridged.mcp;
import dev.ltms.bridged.auth.Authz;
import dev.ltms.bridged.auth.CallerResolver;
import dev.ltms.bridged.auth.MemberRegistry;
import dev.ltms.bridged.auth.Principal;
import dev.ltms.bridged.auth.Role;
import dev.ltms.bridged.config.BridgedConfig;
@@ -69,8 +70,9 @@ class BridgeMcpAuthzTest {
mcp = new BridgeMcp(messages, workers, sessions, identity, sessions.asPresence(),
new PrimaryRegistry(null),
enforce ? new CallerResolver(identity) : null,
metrics);
enforce ? CallerResolver.withLeadsAndMembers(identity, false, null,
Map::of, new MemberRegistry(null)) : null,
metrics, BridgeMcp.CapacitySource.none(), new BridgeMcp.HealthCoverageSource(() -> "off"));
return mcp;
}
@@ -12,6 +12,7 @@ import dev.ltms.bridged.msg.Rendezvous;
import dev.ltms.bridged.session.FakeWorktrees;
import dev.ltms.bridged.session.SessionManager;
import dev.ltms.bridged.peer.MemberRole;
import dev.ltms.bridged.inject.MemberPresence;
import dev.ltms.bridged.session.MemberSession;
import dev.ltms.bridged.session.WorktreeRequest;
import dev.ltms.bridged.member.ClaudeCodeLauncher;
@@ -52,6 +53,18 @@ class BridgeMcpTest {
return ((McpSchema.TextContent) r.content().getFirst()).text();
}
private void assertSendRoundTrips(String target, Set<String> profiles) throws Exception {
CompletableFuture<McpSchema.CallToolResult> send = CompletableFuture.supplyAsync(
() -> BridgeMcp.send(messages, target, "hi", 4000L, null, profiles));
long deadline = System.currentTimeMillis() + 3000;
while (!rendezvous.isWaiting(target) && System.currentTimeMillis() < deadline) {
Thread.sleep(5);
}
assertTrue(rendezvous.isWaiting(target), "send should be accepted for " + target);
BridgeMcp.reply(messages, target, "received");
assertEquals("received", textOf(send.get(6, TimeUnit.SECONDS)));
}
private static ClaudeCodeLauncher workerService(FakeHerdr h, String baseUrl, Set<String> allow) {
BridgedConfig.Profile cfg = new BridgedConfig.Profile(
"ltms-local", baseUrl, "coder", null, "BRIDGED_WORKER_TOKEN", null,
@@ -68,7 +81,7 @@ class BridgeMcpTest {
void sendThenReplyRoundTrips() throws Exception {
// bridge_send blocks; bridge_reply resolves it with the worker's structured answer.
CompletableFuture<McpSchema.CallToolResult> send = CompletableFuture.supplyAsync(
() -> BridgeMcp.send(messages, "term_a", "review this", 4000L));
() -> BridgeMcp.send(messages, "term_a", "review this", 4000L, null, Set.of()));
// Wait until the send has opened its waiter so the reply resolves it (CB-307: reply now
// queues in the inbox if no waiter is open, which would break the round-trip).
@@ -90,7 +103,7 @@ class BridgeMcpTest {
@Test
void asyncSendReturnsATicketThenPollReportsTheReply() throws Exception {
// wait:false parity — a ticket is issued, resolved by a reply, and surfaced by bridge_poll.
McpSchema.CallToolResult accepted = BridgeMcp.sendAsync(messages, "term_a", "do it");
McpSchema.CallToolResult accepted = BridgeMcp.sendAsync(messages, "term_a", "do it", null, Set.of());
assertNotEquals(Boolean.TRUE, accepted.isError());
String out = textOf(accepted);
assertTrue(out.contains("ticket="), out);
@@ -119,6 +132,131 @@ class BridgeMcpTest {
assertEquals("async LGTM", textOf(polled));
}
@Test
void asyncSendSurfacesAnAskThenKeepsTheTicketForTheFinalReply() throws Exception {
McpSchema.CallToolResult accepted = BridgeMcp.sendAsync(messages, "term_a", "do it", null, Set.of());
String ticket = textOf(accepted).substring(textOf(accepted).indexOf("ticket=") + "ticket=".length()).trim();
long deadline = System.currentTimeMillis() + 3000;
while (!rendezvous.isWaiting("term_a") && System.currentTimeMillis() < deadline) {
Thread.sleep(5);
}
assertTrue(rendezvous.isWaiting("term_a"));
CompletableFuture<McpSchema.CallToolResult> ask = CompletableFuture.supplyAsync(
() -> BridgeMcp.ask(messages, "term_a", "which config?", 5000L));
McpSchema.CallToolResult question = BridgeMcp.poll(messages, ticket, null);
deadline = System.currentTimeMillis() + 3000;
while (!textOf(question).contains("[question") && System.currentTimeMillis() < deadline) {
Thread.sleep(5);
question = BridgeMcp.poll(messages, ticket, null);
}
assertTrue(textOf(question).contains("which config?"), textOf(question));
String questionText = textOf(question);
String afterTurnId = questionText.substring(questionText.indexOf("turnId=\"") + "turnId=\"".length());
String turnId = afterTurnId.substring(0, afterTurnId.indexOf('"'));
CompletableFuture<McpSchema.CallToolResult> answer = CompletableFuture.supplyAsync(
() -> BridgeMcp.answer(messages, turnId, "config.yaml", 5000L));
assertEquals("config.yaml", textOf(ask.get(6, TimeUnit.SECONDS)));
deadline = System.currentTimeMillis() + 3000;
while (!rendezvous.isWaiting("term_a") && System.currentTimeMillis() < deadline) {
Thread.sleep(5);
}
assertTrue(rendezvous.isWaiting("term_a"));
BridgeMcp.reply(messages, "term_a", "done");
assertEquals("done", textOf(answer.get(6, TimeUnit.SECONDS)));
McpSchema.CallToolResult done = BridgeMcp.poll(messages, ticket, null);
deadline = System.currentTimeMillis() + 3000;
while (!"done".equals(textOf(done)) && System.currentTimeMillis() < deadline) {
Thread.sleep(5);
done = BridgeMcp.poll(messages, ticket, null);
}
assertEquals("done", textOf(done));
}
@Test
void unansweredAsyncAskReturnsTheTicketToPending() throws Exception {
McpSchema.CallToolResult accepted = BridgeMcp.sendAsync(messages, "term_a", "do it", null, Set.of());
String ticket = textOf(accepted).substring(textOf(accepted).indexOf("ticket=") + "ticket=".length()).trim();
long deadline = System.currentTimeMillis() + 3000;
while (!rendezvous.isWaiting("term_a") && System.currentTimeMillis() < deadline) {
Thread.sleep(5);
}
assertTrue(rendezvous.isWaiting("term_a"));
McpSchema.CallToolResult ask = BridgeMcp.ask(messages, "term_a", "still there?", 50L);
assertTrue(textOf(ask).contains("no answer"), textOf(ask));
assertTrue(textOf(BridgeMcp.poll(messages, ticket, null)).startsWith("[pending"));
BridgeMcp.reply(messages, "term_a", "finished after timeout");
assertEquals("finished after timeout", messages.drainReplies("term_a").getFirst().content());
}
@Test
void asyncSendFailureDoesNotLeaveItsTicketPending() throws Exception {
String ticket = messages.sendAsync("term_a", "do it", () -> {
throw new IllegalStateException("accept failed");
});
long deadline = System.currentTimeMillis() + 3000;
MessageService.TaskView view = messages.poll(ticket);
while (view.phase() == MessageService.Phase.PENDING && System.currentTimeMillis() < deadline) {
Thread.sleep(5);
view = messages.poll(ticket);
}
assertEquals(MessageService.Phase.FAILED, view.phase());
assertEquals("accept failed", view.detail());
}
@Test
void anotherAsyncTicketCannotCaptureAReplyWhileTheFirstTicketIsAsking() throws Exception {
McpSchema.CallToolResult firstAccepted = BridgeMcp.sendAsync(messages, "term_a", "first", null, Set.of());
String firstTicket = textOf(firstAccepted).substring(textOf(firstAccepted).indexOf("ticket=") + "ticket=".length()).trim();
long deadline = System.currentTimeMillis() + 3000;
while (!rendezvous.isWaiting("term_a") && System.currentTimeMillis() < deadline) {
Thread.sleep(5);
}
assertTrue(rendezvous.isWaiting("term_a"));
CompletableFuture<McpSchema.CallToolResult> ask = CompletableFuture.supplyAsync(
() -> BridgeMcp.ask(messages, "term_a", "which config?", 5000L));
MessageService.TaskView first = messages.poll(firstTicket);
deadline = System.currentTimeMillis() + 3000;
while (first.phase() != MessageService.Phase.ASKING && System.currentTimeMillis() < deadline) {
Thread.sleep(5);
first = messages.poll(firstTicket);
}
assertEquals(MessageService.Phase.ASKING, first.phase());
String firstTurnId = first.turnId();
McpSchema.CallToolResult secondAccepted = BridgeMcp.sendAsync(messages, "term_a", "second", null, Set.of());
String secondTicket = textOf(secondAccepted).substring(textOf(secondAccepted).indexOf("ticket=") + "ticket=".length()).trim();
MessageService.TaskView second = messages.poll(secondTicket);
deadline = System.currentTimeMillis() + 3000;
while (second.phase() == MessageService.Phase.PENDING && System.currentTimeMillis() < deadline) {
Thread.sleep(5);
second = messages.poll(secondTicket);
}
assertEquals(MessageService.Phase.FAILED, second.phase());
BridgeMcp.reply(messages, "term_a", "late reply");
assertEquals("late reply", messages.drainReplies("term_a").getFirst().content());
CompletableFuture<McpSchema.CallToolResult> answer = CompletableFuture.supplyAsync(
() -> BridgeMcp.answer(messages, firstTurnId, "config.yaml", 5000L));
assertEquals("config.yaml", textOf(ask.get(6, TimeUnit.SECONDS)));
while (!rendezvous.isWaiting("term_a") && System.currentTimeMillis() < deadline) {
Thread.sleep(5);
}
BridgeMcp.reply(messages, "term_a", "done");
assertEquals("done", textOf(answer.get(6, TimeUnit.SECONDS)));
}
@Test
void pollUnknownTicketIsAnError() {
McpSchema.CallToolResult res = BridgeMcp.poll(messages, "task-999", null);
@@ -128,15 +266,40 @@ class BridgeMcpTest {
@Test
void sendTimesOutWithAWorkingNote() {
McpSchema.CallToolResult res = BridgeMcp.send(messages, "term_a", "hi", 120L);
McpSchema.CallToolResult res = BridgeMcp.send(messages, "term_a", "hi", 120L, null, Set.of());
assertNotEquals(Boolean.TRUE, res.isError(), "a timeout is informational, not a tool error");
assertTrue(textOf(res).contains("no reply"), "got: " + textOf(res));
}
@Test
void sendRejectsMissingArgs() {
assertTrue(BridgeMcp.send(messages, null, "hi", null).isError());
assertTrue(BridgeMcp.send(messages, "term_a", " ", null).isError());
assertTrue(BridgeMcp.send(messages, null, "hi", null, null, Set.of()).isError());
assertTrue(BridgeMcp.send(messages, "term_a", " ", null, null, Set.of()).isError());
}
@Test
void sendRejectsAConfiguredProfileNameBeforeAcceptingIt() {
McpSchema.CallToolResult blocking = BridgeMcp.send(messages, "sol", "hi", 100L, null, Set.of("sol"));
McpSchema.CallToolResult async = BridgeMcp.sendAsync(messages, "sol", "hi", null, Set.of("sol"));
assertTrue(blocking.isError());
assertTrue(async.isError());
assertTrue(textOf(blocking).contains("sol"));
assertTrue(textOf(blocking).contains("configured profile name"));
assertTrue(textOf(blocking).contains("bridge_list"));
assertFalse(textOf(async).contains("ticket="));
}
@Test
void sendAllowsPeerLeadMemberAndUnclassifiedTargets() throws Exception {
Set<String> profiles = Set.of("sol");
assertSendRoundTrips("term_peer_lead", profiles);
assertSendRoundTrips("term_live_member", profiles);
// A herdr-owned pane outside the bridge roster cannot be classified at accept time.
McpSchema.CallToolResult result = BridgeMcp.send(messages, "external-pane", "hi", 10L, null, profiles);
assertFalse(result.isError(), "an unclassified target must not be rejected at acceptance time");
}
@Test
@@ -172,7 +335,7 @@ class BridgeMcpTest {
void askThenAnswerRoundTrips() throws Exception {
// The primary delegates and blocks; wait until its waiter is open before the worker asks.
CompletableFuture<McpSchema.CallToolResult> send = CompletableFuture.supplyAsync(
() -> BridgeMcp.send(messages, "term_a", "do X", 5000L));
() -> BridgeMcp.send(messages, "term_a", "do X", 5000L, null, Set.of()));
long deadline = System.currentTimeMillis() + 3000;
while (!rendezvous.isWaiting("term_a") && System.currentTimeMillis() < deadline) {
//noinspection BusyWait
@@ -304,6 +467,42 @@ class BridgeMcpTest {
assertTrue(out.contains("\"liveStatus\":\"unknown\""), out);
}
@Test
void capacityUsesThePlacementLiveCount() {
FakeHerdr h = new FakeHerdr();
SessionManager sessions = new SessionManager(workerService(h, "http://gx00.gw:8000", Set.of("gx00.gw")));
sessions.acquire("ltms-local", null, null, null);
McpSchema.CallToolResult res = BridgeMcp.listFleet(workerService(h, "http://gx00.gw:8000", Set.of("gx00.gw")),
sessions, null, new BridgeMcp.CapacitySource(profile -> 2, profile -> 2,
() -> Set.of("ltms-local"), () -> 0), new BridgeMcp.HealthCoverageSource(() -> "off"), Map.of(), "");
String out = textOf(res);
assertTrue(out.contains("\"maxLoad\":2"), out);
assertTrue(out.contains("\"live\":2"), out);
assertTrue(out.contains("\"free\":0"), out);
}
@Test
void capacityIncludesConfiguredProfileWithoutMembers() {
FakeHerdr h = new FakeHerdr();
SessionManager sessions = new SessionManager(workerService(h, "http://gx00.gw:8000", Set.of("gx00.gw")));
String out = textOf(BridgeMcp.listFleet(workerService(h, "http://gx00.gw:8000", Set.of("gx00.gw")),
sessions, null, new BridgeMcp.CapacitySource(profile -> 0, profile -> 2,
() -> Set.of("terra"), () -> 0), new BridgeMcp.HealthCoverageSource(() -> "off"), Map.of(), ""));
assertTrue(out.contains("\"profile\":\"terra\""), out);
assertTrue(out.contains("\"live\":0"), out);
assertTrue(out.contains("\"free\":2"), out);
assertTrue(out.contains("\"reclaimable\":0"), out);
}
@Test
void inertCapacitySourceOmitsCapacityBlock() {
FakeHerdr h = new FakeHerdr();
String out = textOf(BridgeMcp.listFleet(workerService(h, "http://gx00.gw:8000", Set.of("gx00.gw")),
new SessionManager(workerService(h, "http://gx00.gw:8000", Set.of("gx00.gw"))), null,
BridgeMcp.CapacitySource.none(), new BridgeMcp.HealthCoverageSource(() -> "off"), Map.of(), ""));
assertFalse(out.contains("\"capacity\":"), out);
}
@Test
void listReportsLeadsAndFlagsTheCallersOwnRow() {
FakeHerdr h = new FakeHerdr();
@@ -404,6 +603,30 @@ class BridgeMcpTest {
assertDoesNotThrow(() -> messages.ackReply("term_a", msgId));
}
@Test
void spawnedMembersAreMarkedPresent() {
Principal worker = Principal.worker("term_worker", 200);
Principal architect = Principal.architect("lead-designer", "term_design", 400);
MemberPresence presence = new MemberPresence();
BridgeMcp.markSpawnedMemberPresent(worker, presence);
BridgeMcp.markSpawnedMemberPresent(architect, presence);
assertTrue(presence.isPresent("term_worker"));
assertTrue(presence.isPresent("term_design"));
}
@Test
void nonMembersAreNotMarkedPresent() {
Principal lead = Principal.leader("opus", "term_lead", 100);
MemberPresence presence = new MemberPresence();
BridgeMcp.markSpawnedMemberPresent(lead, presence);
BridgeMcp.markSpawnedMemberPresent(Principal.anonymous(), presence);
assertFalse(presence.isPresent("term_lead"));
}
@Test
void statusReportsLiveAgentStatus() {
FakeHerdr blocked = new FakeHerdr().agentStatus("blocked");
@@ -553,7 +776,7 @@ class BridgeMcpTest {
assertTrue(out.contains("\"role\":\"dev\""), out);
assertTrue(out.contains("\"role\":\"reviewer\""), out);
assertEquals(2, out.split("\"profile\":\"ltms-local\"", -1).length - 1,
"both members share one profile — that is the point: " + out);
"inert capacity is omitted, leaving the two member rows: " + out);
}
@Test
@@ -102,6 +102,53 @@ class ClaudeCodeLauncherTest {
"the operator's own args are preserved, in order, ahead of the model flag");
}
@Test
void appendsTheBaseComposedRoleAndReplyCharter() {
FakeHerdr herdr = new FakeHerdr();
String roleCharter = "You review changes.";
BridgedConfig.Profile cfg = new BridgedConfig.Profile(
"sonnet", "http://gx00.gw:8000", null, null, "BRIDGED_WORKER_TOKEN",
List.of("claude"), "tab", "bridged-workers", "w #{n}",
"http://127.0.0.1:8765/mcp", null, null);
ClaudeCodeLauncher svc = new ClaudeCodeLauncher(new AgentControl(herdr), new WorkspaceControl(herdr),
new SubscriptionGuard(Set.of("gx00.gw")), Map.of(cfg.profile(), cfg), cfg.profile(), _ -> null,
0, 0L, () -> fleet(Map.of("reviewer", roleCharter), null));
svc.spawn(new SpawnRequest("sonnet", null, null, null, null, MemberRole.REVIEWER));
List<String> args = spawnedArgs(herdr);
int flag = args.indexOf("--append-system-prompt");
assertEquals(1, args.stream().filter("--append-system-prompt"::equals).count(),
"the composed charter is passed once");
assertEquals(roleCharter + "\n\n" + HerdrPeerLauncher.REPLY_CHARTER, args.get(flag + 1),
"the role charter comes first and the reply rule comes last");
}
@Test
void profileWithoutMcpOrRoleCharterGetsNoSystemPrompt() {
FakeHerdr herdr = new FakeHerdr();
ClaudeCodeLauncher svc = labelService(herdr, () -> fleet(Map.of(), null));
svc.spawn(new SpawnRequest("sonnet", null, null, null, null, MemberRole.DEV));
assertFalse(spawnedArgs(herdr).contains("--append-system-prompt"));
}
@Test
void profileWithoutMcpStillGetsItsRoleCharter() {
FakeHerdr herdr = new FakeHerdr();
String roleCharter = "You design changes.";
ClaudeCodeLauncher svc = labelService(herdr, () -> fleet(Map.of("architect", roleCharter), null));
svc.spawn(new SpawnRequest("sonnet", null, null, null, null, MemberRole.ARCHITECT));
List<String> args = spawnedArgs(herdr);
int flag = args.indexOf("--append-system-prompt");
assertTrue(flag >= 0, "a role charter does not need an MCP mount");
assertEquals(roleCharter, args.get(flag + 1));
assertFalse(args.contains("--mcp-config"));
}
private ClaudeCodeLauncher multiProfile(FakeHerdr herdr) {
BridgedConfig.Profile gx10 = new BridgedConfig.Profile("gx10", "http://gx10.gw:8000", "coder",
null, "BRIDGED_WORKER_TOKEN", List.of("claude"), "tab", "bridged-workers", "w #{n}", null, null, null);
@@ -796,14 +843,18 @@ class ClaudeCodeLauncherTest {
.toList();
}
private static BridgedConfig.Fleet fleet(Map<String, String> charters, String tabLabel) {
return new BridgedConfig.Fleet(Map.of(), Map.of(), Map.of(), Map.of(), charters, tabLabel);
}
/** A profile with no {@code tabLabel:} of its own — the fleet template decides. */
private ClaudeCodeLauncher labelService(FakeHerdr herdr, Supplier<String> fleetTemplate) {
private ClaudeCodeLauncher labelService(FakeHerdr herdr, Supplier<BridgedConfig.Fleet> fleet) {
BridgedConfig.Profile cfg = new BridgedConfig.Profile(
"sonnet", "http://gx00.gw:8000", "sonnet", null, "BRIDGED_WORKER_TOKEN",
List.of("claude"), "tab", "bridged-workers", null, null, null, null);
return new ClaudeCodeLauncher(new AgentControl(herdr), new WorkspaceControl(herdr),
new SubscriptionGuard(Set.of("gx00.gw")), Map.of(cfg.profile(), cfg), cfg.profile(),
_ -> null, 0, 0L, fleetTemplate);
_ -> null, 0, 0L, fleet);
}
/**
@@ -814,7 +865,7 @@ class ClaudeCodeLauncherTest {
@Test
void theFleetTemplateNamesTheRoleTheMemberWasSpawnedFor() {
FakeHerdr herdr = new FakeHerdr();
ClaudeCodeLauncher svc = labelService(herdr, () -> "{role}: {profile} #{n}");
ClaudeCodeLauncher svc = labelService(herdr, () -> fleet(null, "{role}: {profile} #{n}"));
svc.spawn(new SpawnRequest("sonnet", null, null, null, null, MemberRole.REVIEWER));
@@ -825,7 +876,7 @@ class ClaudeCodeLauncherTest {
@Test
void theCounterRunsPerRoleAndProfileNotPerFleet() {
FakeHerdr herdr = new FakeHerdr();
ClaudeCodeLauncher svc = labelService(herdr, () -> "{role}: {profile} #{n}");
ClaudeCodeLauncher svc = labelService(herdr, () -> fleet(null, "{role}: {profile} #{n}"));
svc.spawn(new SpawnRequest("sonnet", null, null, null, null, MemberRole.DEV));
svc.spawn(new SpawnRequest("sonnet", null, null, null, null, MemberRole.REVIEWER));
@@ -839,7 +890,7 @@ class ClaudeCodeLauncherTest {
@Test
void aBlankFleetTemplateFallsBackToTheRoleFirstDefault() {
FakeHerdr herdr = new FakeHerdr();
labelService(herdr, () -> null).spawn(
labelService(herdr, () -> fleet(null, null)).spawn(
new SpawnRequest("sonnet", null, null, null, null, MemberRole.ARCHITECT));
assertEquals(List.of("architect: sonnet #1"), tabLabels(herdr));
@@ -855,7 +906,7 @@ class ClaudeCodeLauncherTest {
List.of("claude"), "tab", "bridged-workers", "pinned {profile}", null, null, null);
new ClaudeCodeLauncher(new AgentControl(herdr), new WorkspaceControl(herdr),
new SubscriptionGuard(Set.of("gx00.gw")), Map.of(cfg.profile(), cfg), cfg.profile(),
_ -> null, 0, 0L, () -> "{role}: {profile} #{n}")
_ -> null, 0, 0L, () -> fleet(null, "{role}: {profile} #{n}"))
.spawn(new SpawnRequest("sonnet", null, null, null, null, MemberRole.REVIEWER));
assertEquals(List.of("pinned sonnet"), tabLabels(herdr));
@@ -870,7 +921,7 @@ class ClaudeCodeLauncherTest {
void theTemplateIsReadOnEverySpawnSoAnEditTakesEffect() {
FakeHerdr herdr = new FakeHerdr();
AtomicReference<String> template = new AtomicReference<>("{role}: {profile} #{n}");
ClaudeCodeLauncher svc = labelService(herdr, template::get);
ClaudeCodeLauncher svc = labelService(herdr, () -> fleet(null, template.get()));
svc.spawn(new SpawnRequest("sonnet", null, null, null, null, MemberRole.DEV));
template.set("[{profile}] {role} {n}");
@@ -77,7 +77,7 @@ class CompositePeerLauncherTest {
}
@Override
protected Launch buildLaunch(BridgedConfig.Profile cfg) {
protected Launch buildLaunch(BridgedConfig.Profile cfg, LaunchSpec spec) {
return new Launch(Map.of(), List.of());
}
@@ -0,0 +1,71 @@
package dev.ltms.bridged.member;
import dev.ltms.bridged.config.BridgedConfig;
import dev.ltms.bridged.herdr.AgentControl;
import dev.ltms.bridged.herdr.FakeHerdr;
import dev.ltms.bridged.herdr.WorkspaceControl;
import dev.ltms.bridged.peer.Capability;
import dev.ltms.bridged.peer.MemberRole;
import dev.ltms.bridged.peer.SpawnRequest;
import org.junit.jupiter.api.Test;
import java.util.ArrayList;
import java.util.List;
import java.util.Map;
import java.util.Set;
import java.util.concurrent.atomic.AtomicReference;
import java.util.function.Supplier;
import static org.junit.jupiter.api.Assertions.assertEquals;
class HerdrPeerLauncherCharterTest {
@Test
void readsAndComposesTheFleetCharterForEachSpawn() {
AtomicReference<BridgedConfig.Fleet> fleet = new AtomicReference<>(fleet(Map.of()));
CapturingLauncher launcher = new CapturingLauncher(fleet::get);
launcher.spawn(new SpawnRequest("mcp", null, null, null, null, MemberRole.DEV));
fleet.set(fleet(Map.of("dev", "role charter")));
launcher.spawn(new SpawnRequest("mcp", null, null, null, null, MemberRole.DEV));
launcher.spawn(new SpawnRequest("no-mcp", null, null, null, null, MemberRole.DEV));
assertEquals(HerdrPeerLauncher.REPLY_CHARTER, launcher.specs.get(0).charter(),
"without a role charter, MCP profiles receive only the reply charter");
assertEquals("role charter\n\n" + HerdrPeerLauncher.REPLY_CHARTER, launcher.specs.get(1).charter(),
"the changed supplier value is read for the next spawn and the reply rule is last");
assertEquals("role charter", launcher.specs.get(2).charter(),
"a role charter does not depend on an MCP mount");
}
private static BridgedConfig.Fleet fleet(Map<String, String> charters) {
return new BridgedConfig.Fleet(Map.of(), Map.of(), Map.of(), Map.of(), charters, null);
}
private static final class CapturingLauncher extends HerdrPeerLauncher {
private final List<LaunchSpec> specs = new ArrayList<>();
CapturingLauncher(Supplier<BridgedConfig.Fleet> fleet) {
super("test", new AgentControl(new FakeHerdr()), new WorkspaceControl(new FakeHerdr()),
Map.of("mcp", profile("mcp", "http://bridge"),
"no-mcp", profile("no-mcp", null)),
"mcp", _ -> null, 0, () -> 0L, () -> { }, fleet);
}
@Override
protected Launch buildLaunch(BridgedConfig.Profile cfg, LaunchSpec spec) {
specs.add(spec);
return new Launch(Map.of(), List.of("test"));
}
@Override
public Set<Capability> capabilities() {
return Set.of();
}
private static BridgedConfig.Profile profile(String name, String mcpUrl) {
return new BridgedConfig.Profile(name, "http://gx00.gw:8000", null, null,
"BRIDGED_WORKER_TOKEN", List.of("test"), "pane", null, null, mcpUrl, null, null);
}
}
}
@@ -17,6 +17,10 @@ import java.nio.file.Files;
import java.nio.file.Path;
import java.util.List;
import java.util.Map;
import java.util.concurrent.ExecutorService;
import java.util.concurrent.Executors;
import java.util.concurrent.Future;
import java.util.function.Supplier;
import static org.junit.jupiter.api.Assertions.*;
@@ -34,12 +38,19 @@ class OpenCodeLauncherTest {
}
/** Gate-disabled launcher whose per-spawn config dirs land under an inspectable temp root. */
private OpenCodeLauncher service(FakeHerdr herdr, Path configRoot, BridgedConfig.Profile cfg) {
private static OpenCodeLauncher service(FakeHerdr herdr, Path configRoot, BridgedConfig.Profile cfg) {
return new OpenCodeLauncher(new AgentControl(herdr), new WorkspaceControl(herdr),
Map.of(cfg.profile(), cfg), cfg.profile(), k -> "GITEA_ACCESS_TOKEN".equals(k) ? "tok" : null,
0, System::currentTimeMillis, () -> { }, configRoot, configRoot);
}
private static OpenCodeLauncher service(FakeHerdr herdr, Path configRoot, BridgedConfig.Profile cfg,
Supplier<BridgedConfig.Fleet> fleet) {
return new OpenCodeLauncher(new AgentControl(herdr), new WorkspaceControl(herdr),
Map.of(cfg.profile(), cfg), cfg.profile(), _ -> null,
0, System::currentTimeMillis, () -> { }, configRoot, configRoot, fleet);
}
@SuppressWarnings("unchecked")
private static Map<String, Object> lastStart(FakeHerdr herdr) {
return (Map<String, Object>) herdr.lastCall("agent.start").params();
@@ -62,8 +73,10 @@ class OpenCodeLauncherTest {
@Test
void writesRemoteMcpConfigAndCharterInstructionsWhenMcpUrlSet(@TempDir Path root) throws Exception {
FakeHerdr herdr = new FakeHerdr();
service(herdr, root, opencodeCfg("google/gemini-2.5-pro", "http://127.0.0.1:8765/mcp", null))
.spawn();
BridgedConfig.Fleet fleet = new BridgedConfig.Fleet(Map.of(), Map.of(), Map.of(), Map.of(),
Map.of("dev", "role rule"), null);
service(herdr, root, opencodeCfg("google/gemini-2.5-pro", "http://127.0.0.1:8765/mcp", null),
() -> fleet).spawn();
Map<String, String> env = startEnv(herdr);
assertNull(env.get("ANTHROPIC_BASE_URL"), "opencode carries no ANTHROPIC_* / subscription boundary");
@@ -82,13 +95,15 @@ class OpenCodeLauncherTest {
"the profile's bridge MCP url is present");
assertTrue(bridge.path("enabled").asBoolean(), "the bridge server is enabled");
assertTrue(json.path("instructions").isArray() && !json.path("instructions").isEmpty(),
"the reply charter is mounted via instructions");
"the member charter is mounted via instructions");
// The instructions entry is a real file path holding the reply charter.
Path charter = Path.of(cfgPath).resolveSibling("reply-charter.md");
// The instructions entry is a real file path holding the composed member charter.
Path charter = Path.of(cfgPath).resolveSibling("member-charter.md");
assertTrue(Files.exists(charter), "the charter file the config references was written");
assertTrue(Files.readString(charter).contains("bridge_reply"),
"the charter instructs the worker to answer via bridge_reply");
assertEquals("role rule\n\n" + HerdrPeerLauncher.REPLY_CHARTER, Files.readString(charter),
"the composed charter keeps the role rule first and the reply rule last");
assertEquals(charter.toAbsolutePath().toString(), json.path("instructions").get(0).asText(),
"instructions names the charter file by its absolute path");
}
@Test
@@ -100,6 +115,69 @@ class OpenCodeLauncherTest {
"no bridge MCP url → no config file and no OPENCODE_CONFIG");
}
@Test
void roleCharterWithoutMcpOrCustomProviderStillWritesAConfig(@TempDir Path root) throws Exception {
FakeHerdr herdr = new FakeHerdr();
BridgedConfig.Fleet fleet = new BridgedConfig.Fleet(Map.of(), Map.of(), Map.of(), Map.of(),
Map.of("dev", "role rule"), null);
Path configRoot = Files.createDirectory(root.resolve("configs"));
Path checkout = Files.createDirectory(root.resolve("checkout"));
service(herdr, configRoot, opencodeCfg("google/gemini-2.5-pro", null, null), () -> fleet).spawn();
String cfgPath = startEnv(herdr).get("OPENCODE_CONFIG");
assertNotNull(cfgPath, "a role charter needs a config even without MCP or custom provider");
JsonNode json = new ObjectMapper().readTree(Path.of(cfgPath).toFile());
Path charter = Path.of(json.path("instructions").get(0).asText());
assertEquals("role rule", Files.readString(charter), "the base-composed role charter is unchanged");
assertTrue(json.path("mcp").isMissingNode(), "a charter does not add an MCP mount");
assertTrue(charter.startsWith(configRoot), "the charter is written under the temp config root");
try (var files = Files.walk(checkout)) {
assertFalse(files.anyMatch(path -> path.getFileName().toString().equals("member-charter.md")),
"the worker checkout receives no charter file");
}
}
@Test
void nullCharterWritesNoCharterFileOrInstructions(@TempDir Path root) throws Exception {
FakeHerdr herdr = new FakeHerdr();
service(herdr, root, pinnedCfg("local-vllm/model", "http://127.0.0.1:8000", null),
() -> new BridgedConfig.Fleet(Map.of(), Map.of(), Map.of(), Map.of(), Map.of(), null)).spawn();
String config = startEnv(herdr).get("OPENCODE_CONFIG");
assertNotNull(config, "the custom provider still needs a config");
JsonNode json = new ObjectMapper().readTree(Path.of(config).toFile());
assertTrue(json.path("instructions").isMissingNode(), "a null charter adds no instructions entry");
try (var files = Files.walk(root)) {
assertFalse(files.anyMatch(path -> path.getFileName().toString().equals("member-charter.md")),
"a null charter creates no charter file");
}
}
@Test
void concurrentSpawnsWriteSeparateCharterDirectories(@TempDir Path root) throws Exception {
BridgedConfig.Profile cfg = opencodeCfg("google/gemini-2.5-pro", "http://127.0.0.1:8765/mcp", null);
ExecutorService executor = Executors.newFixedThreadPool(2);
try {
Future<String> first = executor.submit(() -> spawnConfigPath(root, cfg));
Future<String> second = executor.submit(() -> spawnConfigPath(root, cfg));
Path firstCharter = Path.of(first.get()).resolveSibling("member-charter.md");
Path secondCharter = Path.of(second.get()).resolveSibling("member-charter.md");
assertNotEquals(firstCharter.getParent(), secondCharter.getParent(),
"each concurrent spawn owns a separate config directory");
assertTrue(Files.exists(firstCharter));
assertTrue(Files.exists(secondCharter));
} finally {
executor.shutdownNow();
}
}
private static String spawnConfigPath(Path root, BridgedConfig.Profile cfg) {
FakeHerdr herdr = new FakeHerdr();
service(herdr, root, cfg).spawn();
return startEnv(herdr).get("OPENCODE_CONFIG");
}
@Test
void passesTheModelAsDashMFlagAlongsideAutoApprove(@TempDir Path root) {
FakeHerdr herdr = new FakeHerdr();
@@ -10,6 +10,9 @@ import dev.ltms.bridged.inject.Injector;
import org.junit.jupiter.api.BeforeEach;
import org.junit.jupiter.api.Test;
import java.lang.reflect.Field;
import java.util.Map;
import java.util.Set;
import java.util.concurrent.CompletableFuture;
import java.util.concurrent.TimeUnit;
@@ -120,6 +123,36 @@ class MessageServiceTest {
assertFalse(reply.completed());
}
@Test
void droppedQueuedAndDeliveredTurnsExposeTheRealCauseExactlyOnce() throws Exception {
CompletableFuture<MessageService.Reply> first = sendAsync();
awaitWaiting();
injector.onStatus(T, AgentStatus.IDLE); // first delivery
injector.onStatus(T, AgentStatus.WORKING); // first turn in flight
CompletableFuture<Void> queued = injector.enqueue(T, "second task");
CompletableFuture<Rendezvous.Resolution> waiter = rendezvous.currentWaiter(T);
injector.drop(T, new HerdrException("agent target sol not found", "agent_not_found", null));
MessageService.Reply reply = first.get(5, TimeUnit.SECONDS);
assertEquals(MessageService.Outcome.WORKER_FAILED, reply.outcome());
assertEquals("agent target sol not found", reply.text());
assertTrue(queued.isCompletedExceptionally(), "the queued delivery future also fails");
assertFalse(rendezvous.resolveFailure(waiter, "second failure"), "the waiter fails exactly once");
}
@Test
void noDropReasonKeepsTheExistingFallbackText() throws Exception {
herdr.readText("");
CompletableFuture<Rendezvous.Resolution> waiter = rendezvous.open(T);
completion.onTurnFailed(T);
Rendezvous.Resolution resolution = waiter.get(2, TimeUnit.SECONDS);
assertEquals("worker did not reply; its turn ended in an unrecoverable state "
+ "(worker unreachable or stuck)", resolution.text());
}
// --- bridge_ask reverse rendezvous (CB-205) ------------------------------------------------
@Test
@@ -609,4 +642,106 @@ class MessageServiceTest {
assertTrue(view.detail() != null && view.detail().contains("released"),
"and the detail says why, rather than 'worker unknown'");
}
@Test
void abandonFailsEveryPendingAsyncTicketForTheReleasedTarget() throws Exception {
String first = messages.sendAsync(T, "first task");
awaitWaiting(); // first task owns the target lock and rendezvous waiter
String second = messages.sendAsync(T, "second task"); // parked on the same lock, not yet queued
String third = messages.sendAsync(T, "third task"); // a second queued ticket proves the full sweep
assertTrue(messages.abandon(T, "agent target term_a not found"));
assertFailedTicket(first, "agent target term_a not found");
assertFailedTicket(second, "agent target term_a not found");
assertFailedTicket(third, "agent target term_a not found");
}
@Test
void abandonDoesNotFailAnAsyncTicketWaitingForAnAnswer() throws Exception {
String ticket = messages.sendAsync(T, "task that asks");
awaitWaiting();
injectDelivery();
CompletableFuture<MessageService.AskResult> ask =
CompletableFuture.supplyAsync(() -> messages.ask(T, "which config?", 5000));
MessageService.TaskView asking = awaitTicketPhase(ticket, MessageService.Phase.ASKING);
assertFalse(messages.abandon(T, "agent target term_a not found"),
"an asking ticket is an active turn, not a pending send to sweep");
assertEquals(MessageService.Phase.ASKING, messages.poll(ticket).phase());
CompletableFuture<MessageService.Reply> answer = CompletableFuture.supplyAsync(
() -> messages.answer(asking.turnId(), "config.yaml", 5000));
assertEquals("config.yaml", ask.get(5, TimeUnit.SECONDS).answer());
awaitWaiting();
assertTrue(rendezvous.resolve(T, "done"));
assertEquals(MessageService.Outcome.REPLIED, answer.get(5, TimeUnit.SECONDS).outcome());
}
@Test
void unansweredAsyncQuestionReturnsTheTicketToPendingAndReleasesItsTarget() throws Exception {
String ticket = messages.sendAsync(T, "task that asks");
awaitWaiting();
injectDelivery();
assertEquals(MessageService.AskOutcome.TIMED_OUT,
messages.ask(T, "which config?", 200).outcome());
assertEquals(MessageService.Phase.PENDING, messages.poll(ticket).phase(),
"only the question wait ended; the delegated turn may still finish");
String next = messages.sendAsync(T, "next task");
awaitWaiting();
assertTrue(rendezvous.resolve(T, "done"));
assertEquals(MessageService.Phase.DONE, awaitTicketPhase(next, MessageService.Phase.DONE).phase());
}
@Test
@SuppressWarnings("unchecked")
void asyncQuestionBelongsToTheTaskThatOwnsItsForwardWaiter() throws Exception {
String first = messages.sendAsync(T, "first task");
awaitWaiting();
// Model the resolveQuestion/markAsyncQuestion race: another accepted task reached the target set.
Field byTargetField = MessageService.class.getDeclaredField("asyncTasksByTarget");
byTargetField.setAccessible(true);
Map<String, Set<Object>> byTarget = (Map<String, Set<Object>>) byTargetField.get(messages);
Set<Object> targetTasks = byTarget.get(T);
Class<?> taskClass = Class.forName(MessageService.class.getName() + "$Task");
var constructor = taskClass.getDeclaredConstructor(String.class);
constructor.setAccessible(true);
targetTasks.clear();
targetTasks.add(constructor.newInstance(T));
CompletableFuture<MessageService.AskResult> ask =
CompletableFuture.supplyAsync(() -> messages.ask(T, "which config?", 5000));
assertEquals(MessageService.Phase.ASKING, awaitTicketPhase(first, MessageService.Phase.ASKING).phase());
MessageService.TaskView asking = messages.poll(first);
CompletableFuture<MessageService.Reply> answer = CompletableFuture.supplyAsync(
() -> messages.answer(asking.turnId(), "config.yaml", 5000));
assertEquals("config.yaml", ask.get(5, TimeUnit.SECONDS).answer());
awaitWaiting();
assertTrue(rendezvous.resolve(T, "done"));
assertEquals(MessageService.Outcome.REPLIED, answer.get(5, TimeUnit.SECONDS).outcome());
}
private void assertFailedTicket(String ticket, String reason) throws Exception {
MessageService.TaskView view = awaitTicketPhase(ticket, MessageService.Phase.FAILED);
assertEquals(reason, view.detail());
}
private MessageService.TaskView awaitTicketPhase(String ticket, MessageService.Phase phase) throws Exception {
long deadline = System.currentTimeMillis() + 3000;
MessageService.TaskView view;
do {
view = messages.poll(ticket);
if (view.phase() == phase) {
return view;
}
Thread.sleep(5);
} while (System.currentTimeMillis() < deadline);
assertEquals(phase, view.phase());
return view;
}
}
@@ -1,6 +1,7 @@
package dev.ltms.bridged.rest;
import dev.ltms.bridged.auth.CallerResolver;
import dev.ltms.bridged.auth.MemberRegistry;
import dev.ltms.bridged.config.BridgedConfig;
import dev.ltms.bridged.guard.SubscriptionGuard;
import dev.ltms.bridged.herdr.AgentControl;
@@ -65,9 +66,8 @@ class BridgedAppAuthTest {
MessageService messages = new MessageService(agents, injector, new Rendezvous());
ConnectionIdentity identity = new ConnectionIdentity(new PaneLocator(herdr), _ -> pid);
CallerResolver callers = tokenMode
? new CallerResolver(identity, true, token)
: new CallerResolver(identity);
CallerResolver callers = CallerResolver.withLeadsAndMembers(identity, tokenMode, token,
Map::of, new MemberRegistry(null));
metrics = BridgedMetrics.create(sessions, new dev.ltms.bridged.msg.InMemoryReplyInbox());
app = new BridgedApp(herdr, workers, sessions, messages, sessions.asPresence(), null,
@@ -1,5 +1,9 @@
package dev.ltms.bridged.session;
import ch.qos.logback.classic.Level;
import ch.qos.logback.classic.LoggerContext;
import ch.qos.logback.classic.spi.ILoggingEvent;
import ch.qos.logback.core.read.ListAppender;
import dev.ltms.bridged.config.BridgedConfig;
import dev.ltms.bridged.guard.SubscriptionGuard;
import dev.ltms.bridged.herdr.AgentControl;
@@ -8,6 +12,7 @@ import dev.ltms.bridged.herdr.WorkspaceControl;
import dev.ltms.bridged.member.ClaudeCodeLauncher;
import dev.ltms.bridged.peer.PeerUnreachableException;
import org.junit.jupiter.api.Test;
import org.slf4j.LoggerFactory;
import java.util.List;
import java.util.Map;
@@ -158,32 +163,38 @@ class SessionManagerTest {
}
@Test
void recycleProducesNewPaneIdAndOldOneIsGone() {
FakeHerdr herdr = new FakeHerdr();
SessionManager sessions = sessionManager(herdr);
MemberSession oldSession = sessions.acquire("ltms-local", null, "/caller", "term_primary");
String oldPane = oldSession.paneId();
String oldTerminal = oldSession.terminalId();
void onTurnFailedIsLoggedAtWarnWithThePriorState() {
// CB-564: this transition used to be a bare DEBUG "session marked failed" — a symptom with no
// cause. A member that can no longer be delegated to must be at least WARN, and should name
// what stage it failed at (here: BUSY, i.e. a turn was in flight and never resolved).
LoggerContext ctx = (LoggerContext) LoggerFactory.getILoggerFactory();
ch.qos.logback.classic.Logger sessionLog =
(ch.qos.logback.classic.Logger) LoggerFactory.getLogger(SessionManager.class);
ListAppender<ILoggingEvent> appender = new ListAppender<>();
appender.setContext(ctx);
appender.start();
sessionLog.addAppender(appender);
sessionLog.setLevel(Level.WARN);
try {
FakeHerdr herdr = new FakeHerdr();
SessionManager sessions = sessionManager(herdr);
MemberSession session = sessions.acquire("ltms-local", null, "/caller", "term_primary");
String terminal = session.terminalId();
sessions.asPresence().markPresent(terminal);
sessions.onDelivered(terminal);
MemberSession fresh = sessions.recycle(oldPane);
sessions.onTurnFailed(terminal);
assertNotEquals(oldPane, fresh.paneId(), "recycle yields a new pane id");
assertNotEquals(oldTerminal, fresh.terminalId(), "recycle yields a new terminal id");
assertEquals(oldSession.profile(), fresh.profile(), "profile is preserved");
assertEquals(oldSession.cwd(), fresh.cwd(), "cwd is preserved");
assertEquals(oldSession.ownerTerminal(), fresh.ownerTerminal(), "owner is preserved");
assertTrue(sessions.get(oldPane).isEmpty(), "old pane is deregistered");
assertEquals(1, sessions.roster().size(), "only the fresh session remains");
assertEquals(fresh.paneId(), sessions.roster().getFirst().paneId());
// The old session was the first spawn → pane w9:pRoot_1 (CB-519: the registry key is the
// uuid id, so teardown is asserted on the real pane coordinate).
long paneCloseCount = herdr.calls.stream()
.filter(c -> "pane.close".equals(c.method()))
.filter(c -> "w9:pRoot_1".equals(((Map<?, ?>) c.params()).get("pane_id")))
.count();
assertEquals(1, paneCloseCount, "the old worker was torn down");
String warn = appender.list.stream()
.filter(e -> e.getLevel().equals(Level.WARN))
.map(ILoggingEvent::getFormattedMessage)
.findFirst()
.orElse("no turn-failed WARN logged");
assertTrue(warn.contains(terminal), "the log names the member: " + warn);
assertTrue(warn.contains("BUSY"), "the log names the stage it failed at: " + warn);
} finally {
sessionLog.detachAppender(appender);
}
}
@Test
@@ -1,11 +1,13 @@
package dev.ltms.bridged.session;
import dev.ltms.bridged.auth.MemberRegistry;
import dev.ltms.bridged.config.BridgedConfig;
import dev.ltms.bridged.guard.SubscriptionGuard;
import dev.ltms.bridged.herdr.AgentControl;
import dev.ltms.bridged.herdr.FakeHerdr;
import dev.ltms.bridged.herdr.WorkspaceControl;
import dev.ltms.bridged.member.ClaudeCodeLauncher;
import dev.ltms.bridged.peer.MemberRole;
import org.junit.jupiter.api.Test;
import java.util.List;
@@ -22,6 +24,13 @@ import static org.junit.jupiter.api.Assertions.*;
*/
class WorktreeSessionManagerTest {
private static MemberRegistry members() {
return new MemberRegistry(new BridgedConfig.Fleet(Map.of(),
Map.of("architect", new BridgedConfig.Slot("ltms-local")),
Map.of("dev", new BridgedConfig.Slot("ltms-local")),
Map.of("reviewer", new BridgedConfig.Slot("ltms-local")), null));
}
private static ClaudeCodeLauncher workerService(FakeHerdr herdr) {
BridgedConfig.Profile cfg = new BridgedConfig.Profile(
"ltms-local", "http://gx00.gw:8000", "coder", null, "BRIDGED_WORKER_TOKEN",
@@ -56,6 +65,36 @@ class WorktreeSessionManagerTest {
assertEquals("/caller/proj", startCwd(herdr), "spawn receives the caller's cwd");
}
@Test
void onlyArchitectsBindAndReleaseMakesTheirSlotReusable() {
FakeHerdr herdr = new FakeHerdr();
MemberRegistry members = members();
SessionManager sessions = new SessionManager(workerService(herdr), new FakeWorktrees());
sessions.setMemberLifecycle(members);
MemberSession architect = sessions.acquire("ltms-local", MemberRole.ARCHITECT,
null, "/caller/proj", null, null);
MemberSession dev = sessions.acquire("ltms-local", MemberRole.DEV,
null, "/caller/proj", null, null);
MemberSession reviewer = sessions.acquire("ltms-local", MemberRole.REVIEWER,
null, "/caller/proj", null, null);
assertEquals("architect:architect", members.slotForTerminal(architect.terminalId()));
assertNull(members.slotForTerminal(dev.terminalId()), "a dev must never receive architect rights");
assertNull(members.slotForTerminal(reviewer.terminalId()),
"a reviewer must never receive architect rights");
sessions.release(architect.paneId());
MemberSession replacement = sessions.acquire("ltms-local", MemberRole.ARCHITECT,
null, "/caller/proj", null, null);
assertEquals("architect:architect", members.slotForTerminal(replacement.terminalId()));
MemberSession overflow = sessions.acquire("ltms-local", MemberRole.ARCHITECT,
null, "/caller/proj", null, null);
assertNull(members.slotForTerminal(overflow.terminalId()),
"a full slot pool must not stop the architect spawn");
}
@Test
void worktreeAcquireProvisionsAndRecordsPathAndBranch() {
FakeHerdr herdr = new FakeHerdr();
@@ -79,6 +118,19 @@ class WorktreeSessionManagerTest {
assertEquals(expectedPath, s.cwd(), "session cwd is the worktree path");
}
@Test
void worktreeArchitectAcquireAlsoBindsItsSlot() {
FakeHerdr herdr = new FakeHerdr();
MemberRegistry members = members();
SessionManager sessions = new SessionManager(workerService(herdr), new FakeWorktrees());
sessions.setMemberLifecycle(members);
MemberSession architect = sessions.acquire("ltms-local", MemberRole.ARCHITECT,
null, "/caller/proj", null, new WorktreeRequest("cb-548", null));
assertEquals("architect:architect", members.slotForTerminal(architect.terminalId()));
}
@Test
void worktreeAcquireRunsParityOverlayWithProfileDefaults() {
FakeHerdr herdr = new FakeHerdr();
@@ -94,12 +146,12 @@ class WorktreeSessionManagerTest {
FakeWorktrees.OverlayCall overlay = worktrees.lastOverlay();
assertNotNull(overlay);
assertEquals("/repo", overlay.repoRoot());
assertEquals(List.of(".claude/settings.local.json", ".env", ".envrc"),
assertEquals(List.of(".env", ".envrc"),
overlay.requested(), "default parity overlay is used when unset");
assertFalse(overlay.requested().contains(".mcp.json"),
"CB-525: replicating the primary's MCP config gives a worker the primary's IDE "
+ "servers, which navigate its edits out of its own worktree");
assertEquals(List.of(".claude/settings.local.json", ".envrc"), overlay.copied(),
assertEquals(List.of(".envrc"), overlay.copied(),
"existing paths are copied; missing paths are skipped");
assertEquals(List.of(".envrc"), overlay.skipWorktree(),
"tracked copied paths are --skip-worktree'd");
+4 -1
View File
@@ -16,6 +16,9 @@
AuditLogTest, which attaches its own ListAppender and asserts on emitted records.
-->
<!-- Keep the test logger behaviour aligned with the production cancellation filter. -->
<turboFilter class="dev.ltms.bridged.logging.McpCancelledNotificationFilter"/>
<appender name="STDOUT" class="ch.qos.logback.core.ConsoleAppender">
<encoder>
<pattern>%d{HH:mm:ss.SSS} [%thread] %-5level %logger{36} - %msg%n</pattern>
@@ -34,4 +37,4 @@
<appender-ref ref="STDOUT"/>
</root>
</configuration>
</configuration>
+2 -9
View File
@@ -35,7 +35,7 @@ build on.
- **No checkpoint content.** Writing `STATE.md` + commit on teardown is CB-302; CB-301 only exposes
the release hook it will attach to.
"Recycle" under no-reuse is simply **release + fresh acquire** — a helper, not a pool operation.
Under no-reuse, a released session is terminal. A new `acquire` always creates a fresh session.
## Design
@@ -46,11 +46,6 @@ ownership on top.
**Package:** new `dev.ltms.bridged.session` — keeps the registry/lifecycle concern separate from
the `worker` spawn mechanics. Holds `SessionManager` + `WorkerSession`.
**`recycle` is IN SCOPE for CB-301** (decided): implement `recycle(paneId, …)` = `release` the old
session then `acquire` a fresh one, asserting a new distinct paneId (the no-reuse invariant). It is
a thin convenience over the two primitives, shipped now so the no-reuse teardown+respawn path is
covered by a test from day one.
### `WorkerSession` (record or small mutable holder)
| Field | Source | Notes |
@@ -88,7 +83,6 @@ SPAWNING|READY|BUSY|DONE --vanished/drop--> FAILED
final class SessionManager {
WorkerSession acquire(String profile, String requestedCwd, String callerCwd, String ownerTerminal);
void release(String paneId); // deterministic teardown + deregister
WorkerSession recycle(String paneId, ...); // release + acquire (no-reuse convenience)
Optional<WorkerSession> get(String paneId);
List<WorkerSession> roster(); // bridge-owned view (CB-304 consumes this)
// lifecycle hooks (package-private): onReady/onDelivered/onComplete/onFailed(target)
@@ -120,8 +114,7 @@ final class SessionManager {
3. `release` tears the worker down via `WorkerService.stop` and removes it from `roster()`;
a second `release` on the same paneId is a harmless no-op.
4. `onTurnFailed` / drop moves the session to `FAILED` and it is absent from the live roster.
5. `recycle` produces a new paneId and the old one is gone (no-reuse invariant).
6. `roster()` reflects exactly the sessions acquired-minus-released, joined with live status.
5. `roster()` reflects exactly the sessions acquired-minus-released, joined with live status.
## Seams left open (deliberately)
+917
View File
@@ -0,0 +1,917 @@
# M4 - Fleet health, recovery, routing, and capacity
**Status:** Design accepted on 2026-08-15. CB-573 part 1 has shipped the classification model and
the `bridge_list` capacity view; the remaining M4 units are not yet shipped.
**Scope:** Fleet evidence, safe mechanical repair, lead routing, capacity reporting, and optional
human notification.
**Grounded in:** `health/FleetHealth`, `health/PaneBudget`, `inject/StatusPoller`,
`inject/StatusRefiner`, `inject/CompletionResolver`, `inject/Injector`, `session/SessionManager`,
`msg/MessageService`, `msg/ReplyInbox`, `msg/ReplyPushLoop`, `msg/LeadHeartbeatLoop`,
`mcp/PrimaryRegistry`, and `herdr/AgentControl`.
## 1. Problem and decision boundary
The operator asked the bridge to detect idle agents, exceptions, stopped work, and broken
communication. The bridge may read an agent pane from time to time. It must notify a person when
the fleet cannot move forward.
The four operator terms are not four equal health states. `IDLE` is a normal mode. An exception is
sometimes visible only as pane text. Stopped work may look the same as slow work. Broken
communication can occur on several links.
M4 uses this boundary:
- The bridge detects facts and joins evidence.
- The bridge repairs only mechanical failures with no judgement.
- The lead decides whether to stop, retry, replace, or reassign a member.
- A human is notified only when no healthy lead can act.
- n8n may route an outbound incident. It never classifies state or chooses recovery.
An inbound n8n decider would need bridge authority. No narrow machine-decider role exists. Giving a
workflow engine lead authority is unsafe, while adding a new role is a separate authorization
design. An outbound sink needs no bridge role.
The bridge must never replay a delivered task. That task may already have changed files, pushed a
branch, opened a pull request, or changed external state. A replay can run those side effects twice.
This rule must remain true even if later code stores delivered prompt text.
## 2. Evidence model
A health state is mainly a comparison between two views:
- **herdr view:** current agents and raw live status from one `AgentControl.list()` call.
- **bridge view:** session FSM, MCP presence, accepted turns, tasks, inbox state, and lead ownership.
A strong fault often appears as a disagreement between those views. For example, `BUSY` in the
session FSM and `DONE` in herdr means the bridge missed a turn boundary. Pane reads support this
model, but they are not the main monitor.
`SessionManager.rosterView` already joins session state and live status. `AgentControl.list()`
already gets the whole live fleet in one call. M4 makes that join persistent and adds timers,
accepted-turn state, and incident state.
### 2.1 Real traces behind the design
The first trace was an architect that stopped making progress:
```text
profile=opus role=architect state=busy liveStatus=done
```
The session moved from `DONE` to `BUSY` for turn 2. Eighteen minutes later, the session still said
`BUSY`, herdr still said `DONE`, the async task still said `PENDING`, and no completion fallback had
run. This is `TURN_BOUNDARY_LOST`, not a general slow-turn guess.
The second trace had two async sends to the same pane, one second apart. The pane was then stopped.
One ticket became failed. The other stayed `pending - worker unknown`. Current
`MessageService.abandon` resolves only `Rendezvous.currentWaiter(target)`, while async tasks live in
a separate ticket map. CB-568 is intended to fix that bug. M4 still keeps an independent
post-teardown invariant so a later regression becomes `DELEGATION_ORPHANED`.
### 2.2 Corrections made during design
The first state table missed `BUSY` in bridged plus `IDLE` or `DONE` in herdr. It would have found
the real trace only through a late, weak stall timer. The final model adds
`TURN_BOUNDARY_LOST` as a strong disagreement state.
The first notification design also required a webhook before `health.enabled` could turn on. That
removed useful local detection to avoid a narrower human-notification gap. The final design splits
detection from notification. Missing human escalation is shown as partial coverage instead of
disabling health.
## 3. Classification precedence
Evidence is applied in this order. A lower rule cannot hide a higher one.
1. **Control link:** failed fleet list plus failed ping becomes `CONTROL_LINK_DOWN`.
2. **Definitive target loss:** `_not_found` becomes `GONE` or `LEAD_UNREACHABLE` when the control
link is healthy.
3. **Startup and teardown invariants:** readiness expiry becomes `NEVER_READY`; surviving tasks
after teardown become `DELEGATION_ORPHANED`.
4. **Bridge/live disagreement:** `BUSY` plus stable raw `IDLE` or `DONE` becomes
`TURN_BOUNDARY_LOST`.
5. **Known screen evidence:** a tested fatal signature becomes `ERROR_ON_SCREEN`.
6. **Timed suspicion:** unchanged sparse pane probes may become `STALL_SUSPECTED`.
7. **Communication quality:** completion fallback becomes `MUTE`; an old inbox entry becomes
`REPLY_STRANDED`.
8. **Normal mode:** `STARTING`, `IDLE`, `WORKING`, `WORK_PENDING`, or `BLOCKED_AMBIGUOUS`.
The member flow in Figure 1 shows lifecycle states and the main fault exits. Fault states are
reported beside the session FSM; most are not new FSM values.
```mermaid
flowchart TD
Registered["Member registered"] --> Starting["STARTING"]
Starting -->|"MCP presence"| Idle["IDLE"]
Starting -->|"Readiness grace expires"| NeverReady["NEVER_READY"]
Idle -->|"Accepted delivery"| Working["WORKING"]
Working -->|"Trusted turn boundary"| Idle
Working -->|"Bridge BUSY and herdr IDLE or DONE"| Lost["TURN_BOUNDARY_LOST"]
Working -->|"Known fatal screen"| Error["ERROR_ON_SCREEN"]
Working -->|"Long age and unchanged sparse probes"| Stall["STALL_SUSPECTED"]
Working -->|"Target not found"| Gone["GONE"]
Idle -->|"Inbox or queued delivery exists"| Pending["WORK_PENDING"]
Pending -->|"Delivery or collection finishes"| Idle
Idle -->|"Raw BLOCKED with an open turn"| Blocked["BLOCKED_AMBIGUOUS"]
Lost -->|"Strict guarded repair"| Repaired["DONE with reconciled completion"]
Lost -->|"Repair refused"| LeadDecision["Lead decision required"]
```
*Figure 1. The member lifecycle and the main health exits. Pane-based states never authorise an
automatic retry of the task.*
## 4. State model
### 4.1 Normal and transitional member states
| State | Exact evidence | Meaning and certainty |
|---|---|---|
| `STARTING` | Session is `SPAWNING`; MCP presence is absent | Normal inside the startup grace. MCP contact is the readiness signal. |
| `IDLE` | Session is `READY` or `DONE`; live status is `IDLE` or `DONE`; no open turn or inbox item exists | Normal. Idle is not a fault. |
| `WORKING` | Session is `BUSY`; raw live status is `WORKING`; the accepted turn is open | Certain that herdr sees work. It does not prove useful progress. |
| `WORK_PENDING` | Queued delivery or inbox content exists while the target is injectable | Transitional. Existing injector or push logic should move it. |
| `BLOCKED_AMBIGUOUS` | An open turn exists and raw live status is `BLOCKED` | The bridge cannot tell whether this is permission, input, or a settled screen. |
Idle may drive configured resource cleanup. It never opens an incident and never pages a person.
### 4.2 Member fault and quality states
| State | Exact evidence | Certainty and action |
|---|---|---|
| `NEVER_READY` | `SPAWNING`, no MCP presence, and an accepted delivery waits through the existing readiness grace | Delivery never became possible. The exact cause is unknown. Fail the send, stop the process, and preserve a provisioned worktree. |
| `GONE` | Per-target herdr call returns `_not_found` while fleet list or ping works | Certain target loss. Fail all target work. Do not replay it. |
| `TURN_BOUNDARY_LOST` | Same session turn stays `BUSY`; same accepted task stays open; two raw snapshots show `IDLE` or `DONE` | Strong disagreement. Strict reconciliation may repair it. |
| `ERROR_ON_SCREEN` | Suspicious non-working state survives grace; `detection` matches a tested adapter-specific fatal signature | Certain only for the matched signature. A bare word such as `Exception` is not enough. |
| `STALL_SUSPECTED` | Open turn is older than the configured threshold; two normalised `recent_unwrapped` digests are unchanged; no boundary or reply occurs | Not certain. A long valid API call can look the same. Lead decides. |
| `MUTE` | Turn resolves through completion fallback instead of `bridge_reply` | Certain that no structured reply won. It does not prove an MCP failure. A single event is a metric, not an incident. |
| `REPLY_STRANDED` | Typed reply or health message remains after owning-lead push reaches its cap | Collection failed. This does not explain whether the lead is busy, dead, or ignoring the nudge. |
| `DELEGATION_ORPHANED` | Target is gone, failed, or released, but one or more tasks remain `PENDING` after reconciliation grace | Certain bridge invariant failure. This is not an inbox-drain fault. |
| `WORK_PRODUCT_AT_RISK` | Provisioned branch has commits after its recorded base; member is `DONE`, `FAILED`, or preserved after release; no turn or inbox item remains; long-idle threshold passed | A warning, not proof of loss. Work may already have an open pull request or a squash merge. |
`MUTE` opens an incident only after a small fixed rate threshold for one target or profile, or when
it appears with another fault.
`WORK_PRODUCT_AT_RISK` must not become `WORK_PRODUCT_UNCOLLECTED`. The bridge does not know pull
request or merge state. If committed work appears with `REPLY_STRANDED` or
`DELEGATION_ORPHANED`, the existing incident gains `committedWorkAtRisk: true`.
### 4.3 Control-link state
| State | Exact evidence | Certainty and action |
|---|---|---|
| `CONTROL_LINK_DOWN` | Two full-fleet `agent.list` calls fail across the grace, and herdr `ping` also fails | Certain for the bridged-to-herdr link. Retry calls, record the incident, and use human escalation if no lead can be reached. |
A failed fleet list alone is not a dead-member claim. A single `_not_found` with a healthy global
link is a target fault, not a control-link fault.
### 4.4 Lead states
| State | Exact evidence | Meaning and action |
|---|---|---|
| `LEAD_IDLE` | Expected lead is present with raw injectable status; no actionable state waits | Normal. Existing heartbeat may run under its own policy. |
| `LEAD_WORKING` | Expected lead is present with raw `WORKING`; stall threshold is not met | Reachable and busy. Never inject into the live turn. |
| `LEAD_STATUS_UNKNOWN` | Expected lead is present with raw `UNKNOWN` | Neither dead nor a healthy routing target. Retain evidence and retry. |
| `LEAD_UNREACHABLE` | Expected lead is absent from two successful live-agent snapshots while ping works, or targeted lookup returns `_not_found` with a healthy control link | Route to a healthy peer. If none exists, use human escalation. |
| `LEAD_UNRESPONSIVE` | Actionable state waits; lead stays injectable; bounded nudges exhaust; inbox remains uncollected | Route to a healthy peer or a person. |
| `LEAD_STALL_SUSPECTED` | Lead stays `WORKING` past threshold; two sparse pane probes show no progress | Not certain. Never kill or restart automatically. Route to peer or person. |
The monitor retains the lead name and terminal, last successful sighting, raw status and age,
consecutive list absences, targeted errors, pane-probe facts, pending incident age, and nudge
outcomes. Current heartbeat and push loops discard much of this history.
Expected lead identity comes from the same supplier used by `CallerResolver`. It is not liveness
evidence. `LeadTabScanner` keeps cached identity after a failed scan, so the health monitor compares
that identity with a fresh successful agent list. A dynamic identity also survives a two-successful-
snapshot retirement grace. This stops a dead lead from escaping health by disappearing from one map.
### 4.5 Evidence limits
M4 cannot tell these cases apart with current evidence:
- A valid long call and a hung call may have the same status and pane digest.
- `BLOCKED` does not explain which input is needed.
- An idle prompt after failure may look like an idle prompt after success.
- A missing structured reply does not prove a broken MCP connection.
- An undrained inbox does not explain why the lead did not collect it.
- Arbitrary pane text cannot safely classify arbitrary exceptions.
- A branch ahead of its base does not prove that work was not collected.
Logs are outputs, not classifier inputs. The monitor never parses its own logs.
## 5. Automatic action and lead action
### 5.1 Actions the bridge may take
The bridge may:
- retry transient herdr status, list, ping, and pane-read failures with bounded backoff;
- re-submit Enter after the existing paste/submit race;
- fail queued delivery after `NEVER_READY`;
- stop a never-ready process while preserving its provisioned worktree;
- fail all queued, accepted, and async tasks for a gone or released target;
- reconcile one lost boundary when every strict gate in Section 8 passes;
- hold typed messages, nudge the owning lead, and stop at the configured cap;
- use the existing bounded idle-lead heartbeat;
- deduplicate, route, update, and resolve incidents.
These actions do not choose new work and do not replay old work.
### 5.2 Decisions reserved for the lead
Only the lead may:
- stop or continue `BLOCKED_AMBIGUOUS`;
- stop, inspect, or wait on `ERROR_ON_SCREEN`;
- kill or continue `STALL_SUSPECTED`;
- spawn a replacement or reassign work;
- retry a delivered task;
- choose how to use partial work in a worktree;
- restart herdr or change network, model, credentials, backend, or configuration.
Reports include literal safe tool calls such as `bridge_status(sessionId="...")`,
`bridge_poll(ticket="...")`, `bridge_list()`, and optional `bridge_stop(paneId="...")`. A judgement
state never presents stop as the only action.
### 5.3 Release causes and worktree safety
| Release cause | Process action | Provisioned worktree |
|---|---|---|
| `SPAWN_ROLLBACK` before registration or delivery | Stop and clean up | Remove |
| `COMPLETED` for `READY` or `DONE` without pending work, idle TTL, or successful context-cap completion | Stop | Remove under completed policy |
| `NEVER_READY` | Stop | Preserve |
| `GONE` | Best-effort stop | Preserve |
| `TURN_FAILED` or lead abort while `BUSY` or `FAILED` | Stop | Preserve |
| `RELEASE_WITH_PENDING_TASKS` | Stop | Preserve |
| `SHUTDOWN` | Stop | Preserve |
Explicit stop is state-aware. `SPAWNING`, `BUSY`, `FAILED`, or any target with pending tasks uses a
preserving cause.
Before abnormal release removes the live session, M4 writes an atomic manifest under the worktree
root. It records session identity, owner, role, profile, repository, path, branch, base commit,
release cause, release time, state, and pending task ids. `bridge_list.preservedWorktrees` loads these
manifests after restart. Stop output and WARN logs also name the path and cause. M4 never
auto-deletes a preserved worktree.
## 6. Fleet health monitor
Add `FleetHealthMonitor`. Do not widen `StatusPoller` into a policy loop.
`StatusPoller` has a 250 ms delivery cadence and samples only injector targets with outstanding
work. Health needs all sessions, all leads, task state, inbox age, and global control evidence. One
loop cannot serve both cadences safely.
Build the monitor like `LeadHeartbeatLoop`:
- pure `decide(snapshot, priorState, now)` logic;
- a thin scheduler;
- an injected clock;
- edge-triggered state changes;
- no network work in the pure function;
- no sleeping in tests.
Each enabled fleet tick reads:
- one `AgentControl.list()` result for the whole fleet;
- one in-memory `SessionManager.roster()` snapshot;
- accepted turns and async task state;
- typed inbox depth, kind, and age;
- push and heartbeat outcomes;
- configured and discovered leads.
Existing failure paths publish structured evidence to the monitor. The monitor does not infer events
from log text.
### 6.1 Pane budget
Healthy idle members, recent working members, and quiet leads cause no pane reads.
A pane is eligible only for a stable lost boundary, sustained `BLOCKED` or `UNKNOWN`, work older
than the suspect threshold, or one final evidence read for a confirmed fault when the pane exists.
Compiled brakes apply even if config asks for more:
- per-target pane cooldown is at least 60 seconds;
- working age before the first progress probe is at least 300 seconds;
- at most two pane reads occur in one fleet tick;
- targets rotate fairly;
- only a normalised digest and optional clipped local excerpt are stored;
- no pane excerpt leaves bridged in a human webhook.
Use `detection` for tested screen signatures. Use normalised `recent_unwrapped` only for progress
comparison.
## 7. Typed inbox and routing
### 7.1 Semantic record
The typed inbox record carries:
```text
schemaVersion
kind: reply | health
msgId, target, subjectTerminal, recipientLead
severity, state, evidence
createdAtEpochMillis, firstSeenEpochMillis, lastSeenEpochMillis
recoveryTried, suggestedToolCalls, content
```
A health message never calls `Rendezvous.resolve`. It cannot look like the member's task result.
Both inbox adapters share field preservation, first-id-wins dedup, FIFO among decoded messages,
explicit ownership, ack, and release rules. The in-memory adapter stores typed records directly. It
does not copy AMQP migration logic.
### 7.2 AMQP migration
The reader uses AMQP `content_type`, never body sniffing:
```text
Legacy v0: text/plain
Typed family: application/vnd.ltms.bridged.inbox-message+json
```
A legacy reply may begin with `{`. It remains plain text because its media type is `text/plain`.
Legacy text becomes `kind=reply` with exact UTF-8 content and absent typed metadata.
Typed JSON has required integer `schemaVersion: 1`. Version 1 ignores unknown optional fields.
Missing required fields, invalid enums, malformed UTF-8 or JSON, and property/body identity mismatch
are invalid data.
An unknown schema version is not partly decoded. It remains unacknowledged on the original queue and
creates one operator-visible `unsupported_version` failure. A newer daemon may read it later.
Invalid known-format data is copied byte-for-byte to durable queue
`agent.<target>.inbox.quarantine`. A dedicated confirm-mode publisher confirms the persistent copy
before the original is acknowledged. A failed quarantine handoff leaves the original unacknowledged.
The raw body never enters logs.
Decode failure creates a redacted WARN, metric, `bridge_list` summary, and routed health incident.
One bad entry never escapes the consumer callback and never stops later valid messages.
Safe downgrade is not supported. The previous build ignores `content_type` and would show typed JSON
as ordinary reply text. If drained, it would acknowledge the message and lose typed meaning. Typed
queues must be drained or preserved before an old jar runs.
The existing contract suite uses RabbitMQ. Production uses LavinMQ. The migration and lead-key
ownership cases must run once against production LavinMQ before release, or the release must state
that LavinMQ was not checked.
### 7.3 Member routing
A member incident first goes to the exact lead that owns its accepted delegation.
`PrimaryRegistry` needs a no-fallback `delegatingLeadFor(memberTarget)` query. Health routing must not
use the old singular-primary fallback when several leads exist.
Publish the incident under the affected member target. Trigger the existing bounded push route. The
push waits until the owning lead is injectable, so it does not interrupt a live lead turn.
### 7.4 Peer lead routing
Figure 2 shows the route from incident to lead, peer, or person.
```mermaid
flowchart TD
Incident["Open incident"] --> Member{"Member incident?"}
Member -->|"yes"| Known{"Exact delegation owner known?"}
Known -->|"no"| Sink{"Human webhook enabled and healthy?"}
Known -->|"yes"| Owner{"Owner lead healthy?"}
Owner -->|"yes"| OwnerInbox["Publish to owner lead path"]
Owner -->|"no"| PeerSet["Build healthy peer candidate set"]
Member -->|"no, lead incident"| PeerSet
PeerSet --> Peer{"Healthy peer exists?"}
Peer -->|"yes"| Select["Choose fewest assigned incidents<br/>then stable name and terminal id"]
Select --> PeerInbox["Publish to peer lead inbox<br/>and status-gated push"]
Peer -->|"no"| Sink
Sink -->|"yes"| Webhook["Send classified outbound incident"]
Sink -->|"no"| Passive["Keep incident open<br/>show partial coverage on local surfaces"]
```
*Figure 2. Routing keeps delegation ownership separate from temporary peer fallback.*
Peer candidates exclude the incident subject, failed owner, absent leads, raw-unknown leads, and
leads with an open unhealthy state. A reachable `WORKING` peer may be selected; its push waits for an
injectable window.
Choose the candidate with the fewest assigned foreign incidents. Break ties by stable lead name,
then terminal id. Pin the recipient. Reassign only if that peer becomes unhealthy or retires. A
routing generation marks a reassignment, and old pending assignments become superseded.
`bridge_list` lead rows show health, health age, assigned foreign incident count, and a bounded list
of incident id, subject, state, severity, age, and routing generation. The top-level view also shows
owner, recipient, and routing reason.
A peer incident is published under the recipient lead's inbox key, not the failed subject's key. Its
status-gated nudge names the failed lead and gives the exact
`bridge_poll(target="<recipient-terminal>")` call.
### 7.5 Lead inbox ownership
Add `LeadInboxRegistry`, driven by the same expected-lead supplier as `CallerResolver`.
It calls `replyInbox.own(leadTerminal)` at startup for configured leads, after successful discovery,
after config adds a lead, and before publication. Ownership is not an authorization side effect.
A missing lead keeps its key owned. Release happens only after confirmed retirement, all incidents
are reassigned or resolved, typed health messages move or ack, and the queue is empty. Own a
replacement terminal before moving messages from the old key. Never release a non-empty in-memory
lead key, because in-memory release clears local data.
### 7.6 Single-lead deployment
One lead and no peer is a normal mode, not an edge case.
An idle, reachable lead may receive the existing bounded nudge. An unreachable or stalled sole lead
has no safe in-loop recovery. The bridge must not restart or replace it. A new lead would not have the
failed lead's plan or context, and an uncertain relaunch could create two orchestrators.
With no webhook, only `bridge_list`, `/healthz`, metrics, WARN logs, and the incident journal remain.
These are passive surfaces. They are not a human notification.
## 8. Lost-boundary reconciliation
This is the only M4 path that reconstructs a result. It must prefer a visible stall over a fabricated
reply.
### 8.1 Why normal completion rules are not enough
Current `CompletionResolver.resolve` has two fail-open rules. It resolves when the delivery baseline
is missing. It also resolves an empty completion when the pane read fails. Those choices are valid
after a trusted `WORKING -> IDLE` boundary because the bridge knows the turn ran. They are unsafe
when health only guesses that a boundary was lost.
M4 gives each accepted send an internal `TurnToken`. It ties target, exact waiter, session turn,
delivery baseline, and task outcome together.
### 8.2 Delivery baseline
Capture the baseline immediately after prompt send and before the delivery future completes. Store:
```text
TurnToken
exact waiter identity
capture time and pane source
normalised assistant block clipped to MAX_SCRAPE_CHARS
whether a supported assistant marker was recognised
capture result: PRESENT | READ_FAILED | UNRECOGNISED
```
A failed or missing baseline never authorises repair. A late baseline is not valid evidence. After a
daemon restart, the old waiter, task, token, and baseline are gone, so the old turn cannot be
repaired.
Automatic repair is enabled only for agent kinds with tested assistant-block fixtures. Current
extraction is Claude Code-specific and falls back to arbitrary raw text without `⏺`. That raw fallback
cannot authorise repair. OpenCode repair stays disabled until live pane fixtures exist.
### 8.3 Strict gates and resolver result
Figure 3 shows the repair gates. Any failed gate keeps the waiter unchanged.
The two raw snapshots must describe the same `TurnToken` and session turn. No `WORKING`,
`BLOCKED`, `UNKNOWN`, missing-agent, reply, failure, or new-delivery observation may occur between
them.
```mermaid
flowchart TD
Candidate["TURN_BOUNDARY_LOST candidate"] --> Stable{"Same TurnToken and BUSY turn<br/>across two raw IDLE or DONE snapshots?"}
Stable -->|"no"| Resnapshot["Take a fresh snapshot"]
Stable -->|"yes"| Waiter{"Exact captured waiter<br/>still open by identity?"}
Waiter -->|"no"| Stale["STALE_TURN or ALREADY_RESOLVED"]
Waiter -->|"yes"| Baseline{"Successful recognised<br/>delivery baseline exists?"}
Baseline -->|"no"| Refuse["Refuse repair<br/>leave ticket pending"]
Baseline -->|"yes"| Read{"Fresh pane read succeeds?"}
Read -->|"no"| Refuse
Read -->|"yes"| Output{"Recognised non-blank assistant block<br/>differs from clipped baseline?"}
Output -->|"no"| Refuse
Output -->|"yes"| Resolve["Shared CompletionResolver guard core<br/>resolves exact waiter"]
Resolve -->|"won race"| Repaired["RECONCILED_COMPLETION<br/>same turn becomes DONE"]
Resolve -->|"lost race"| Resnapshot
```
*Figure 3. Repair needs stronger evidence than a normal observed turn boundary.*
Refactor the current resolver into one guard core with two policies:
```text
resolveCaptured(target, inFlight, OBSERVED_BOUNDARY)
resolveCaptured(target, inFlight, LOST_BOUNDARY_REPAIR)
```
The health monitor calls only:
```text
CompletionResolver.reconcileLostBoundary(target, expectedTurnToken)
```
It returns `REPAIRED`, `ALREADY_RESOLVED`, `REFUSED_NO_CAPTURE`, `REFUSED_NO_BASELINE`,
`REFUSED_UNREADABLE`, `REFUSED_UNCHANGED`, `REFUSED_AMBIGUOUS_OUTPUT`, `STALE_TURN`, or
`RACE_LOST`.
Only `REPAIRED` and same-turn `ALREADY_RESOLVED` may move that turn from `BUSY` to `DONE`. A
per-target reconciliation gate stops a queued second send from being accepted between waiter
resolution and the FSM transition.
### 8.4 Lead-visible marker and refusal
A repaired result uses distinct `RECONCILED_COMPLETION` values in `Rendezvous`, `MessageService`,
task poll source, and metrics. The lead sees:
```text
[repaired completion - bridged detected a lost turn boundary. The member did not call
bridge_reply; pane-derived text follows and may be partial]
```
Clipped text also keeps the existing clipped-tail marker.
A refused repair leaves `TURN_BOUNDARY_LOST` open and the ticket pending. The report states that no
reply was reconstructed and no task was replayed. `UNCHANGED`, `UNREADABLE`, and
`AMBIGUOUS_OUTPUT` get at most one delayed retry for the same token. Missing capture or baseline gets
no retry. After two refused scrapes, automatic repair stops for that token.
### 8.5 Target-wide teardown invariant
CB-568 owns the multi-ticket cancellation mechanism. M4 routes every terminal cause through that one
idempotent operation and checks this independent invariant after teardown:
- no injector entry exists for the target;
- no accepted turn or completion record exists;
- no rendezvous waiter or ask exists;
- every async task is terminal or was already terminal;
- no thread waiting for the target send lock can later accept it;
- new sends fail immediately;
- each old task has one terminal outcome and one metric count.
A violation becomes `DELEGATION_ORPHANED`. The monitor may call the same idempotent target-wide
failure operation once. It never recreates the task.
## 9. Capacity and utilisation
Capacity is a view, not a health state.
`bridge_list` adds one block per profile:
```text
profile, maxLoad, live, free, reclaimable
```
For an unlimited profile, `maxLoad` and `free` are null. `free` is
`max(0, maxLoad - live)` for a capped profile.
The view must use the exact live-count function used by placement. A second calculation could show a
free slot that placement then refuses. Member rows add `idleForSeconds` only when state is `READY` or
`DONE`, no accepted turn exists, and the inbox is empty. `reclaimable` means only that the member
holds capacity without open bridge work.
The existing idle-lead nudge gains a bounded capacity summary. It lists per-profile live, cap, free,
and reclaimable counts, plus at most three long-idle members. Capacity does not make
`FleetState.hasPending()` true. A changed capacity fingerprint may re-arm one capped heartbeat
sequence. The fingerprint excludes changing idle durations, so a static idle fleet cannot reset the
cap forever. Reply-push stand-down remains first.
The bridge must never:
- spawn a member because a slot is free;
- generate a task or acceptance criteria;
- move queued work to another member or profile;
- treat a free slot or idle member as an incident;
- stop an idle member only to improve utilisation.
The bridge knows capacity facts but has no work list. Only the lead has the plan, task context,
side-effect history, and acceptance criteria.
Capacity calculation is in memory and adds no pane reads. Work-product checks run on a terminal
session edge, not every fleet tick.
This capacity design adds no automatic stop. The accepted `NEVER_READY` cleanup can still stop a
very slow startup after the existing grace, which is a known risk. Free capacity and long idle time
never trigger that path.
## 10. Human escalation and notification
### 10.1 Escalation rule
Notify a person only when no healthy lead can act:
- `CONTROL_LINK_DOWN` survives grace;
- a lead is unhealthy and no healthy peer can receive the incident;
- a member incident has no known owning lead;
- the only owning lead becomes unreachable, unresponsive, or stalled;
- incident publication or routing itself fails.
Do not page a person for a member fault while a healthy owning lead exists. An uncollected member
incident feeds lead-health evidence. If the lead then becomes unhealthy, peer or human routing starts.
### 10.2 Detection and notification switches
`health.enabled` controls detection and bridge-local reporting. It does not require a webhook.
`health.notifications.mode` is `disabled` or `webhook`. Disabled is valid and is the default.
Webhook mode requires a resolved environment variable. Turning notification off stops outbound
attempts but keeps incidents. Turning it back on resumes still-open human incidents.
Without a sink, `bridge_list.healthCoverage` states that human escalation is unavailable. `/healthz`
keeps its existing HTTP liveness result and adds a nested `fleetHealth.status=partial` component.
Metrics and one startup or reload WARN expose the same limit.
### 10.3 Incident and delivery deduplication
One open incident uses this key:
```text
(scope, subjectStableId, state, causeFingerprint)
```
The cause fingerprint includes stable error codes, dependency names, signature ids, or invariant
names. It excludes times, ages, retry counts, pane text, and changing digests. A later recurrence
after resolution gets a new generation and incident id.
Each outbound event uses:
```text
Idempotency-Key = hash(incidentId, eventType, eventRevision)
```
Event types are `open`, `severity_changed`, `reminder`, and `resolved`. Transport retries keep the
same key.
An atomic owner-only journal beside the active config stores open incidents, routing, delivered
revisions, retry state, and resolution state. It stores no pane or task content. Journal failure does
not stop detection, but notification coverage becomes degraded.
### 10.4 Retry, reminder, and resolve
Send the first event immediately. Retry network errors, timeouts, HTTP 408, HTTP 429, and HTTP 5xx
with full-jitter exponential backoff:
```text
base: 5 seconds
factor: 3
maximum delay: 15 minutes
one outstanding attempt per event
```
Respect `Retry-After` up to 15 minutes. Other HTTP 4xx responses are permanent for that event until
config changes or a person requests replay.
Transport retry is not an incident reminder. `humanRepeatSeconds` creates a new reminder revision
for an unresolved critical incident after the last successful human event. Disabled mode does not
build an unbounded reminder queue.
Send `resolved` only if at least one human event for that incident was delivered. If an incident
resolves before its first successful delivery, cancel the pending open event and record local
resolution.
### 10.5 Outbound payload boundary
An outbound payload may contain incident id and event type, severity, state, scope, stable bridge
ids, role or profile, times, duration, structured evidence type and counts, recovery attempted,
routing reason, coverage, and safe tool calls.
It must never contain:
- raw pane text, pane excerpts, or pane digests;
- task briefs, prompts, or member reply content;
- source files, diffs, or worktree file content;
- worktree paths;
- environment values, tokens, credentials, headers, or webhook URL;
- raw exception messages or stack traces;
- arbitrary model output.
The sink response body is ignored. A webhook cannot direct recovery. n8n remains outbound-only.
### 10.6 Metrics
M4 adds bounded-label series:
```text
bridged_health_incidents{scope,state,severity}
bridged_health_incidents_total{event}
bridged_health_notifications_total{event,outcome}
bridged_health_notification_queue_depth
bridged_health_notification_last_success_seconds
bridged_health_notification_capability{mode,status}
bridged_lead_health{lead,state}
bridged_lead_assigned_incidents{lead}
```
Metric labels never include terminal ids, incident ids, URLs, or error text.
## 11. Configuration
The optional `health:` block is absent or disabled by default. The dormant monitor scheduler does no
herdr or pane work while disabled. Every listed key is hot because the monitor reads `ConfigRef` on
each tick or notification.
| Key | Class | Default and hard bound | Purpose |
|---|---|---|---|
| `health.enabled` | Hot | `false` | Enable detection and bridge-local reporting. |
| `health.snapshotIntervalSeconds` | Hot | default 30, minimum 15 | Whole-fleet comparison cadence. |
| `health.workingSuspectAfterSeconds` | Hot | default 600, minimum 300 | Age before working-pane probes. |
| `health.paneProbeIntervalSeconds` | Hot | default 60, minimum 60 | Per-target pane cooldown. |
| `health.leadUnresponsiveAfterSeconds` | Hot | default 300, minimum 120 | Delay after exhausted actionable nudges before lead fault. |
| `health.humanRepeatSeconds` | Hot | default 3600, minimum 900 | Minimum repeat period for one open human incident. |
| `health.capacityLongIdleAfterSeconds` | Hot | default 900, minimum 300 | Long-idle threshold for capacity summaries. |
| `health.includePaneExcerpt` | Hot | `false` | Allow a clipped excerpt in local lead reports only. Human payloads still exclude it. |
| `health.notifications.mode` | Hot | `disabled` | Select `disabled` or `webhook`. |
| `health.notifications.webhookUrlEnv` | Hot | required in webhook mode | Name of the environment variable that holds the sink URL. |
| `health.notifications.requestTimeoutMs` | Hot | default 10000, range 1000-30000 | Whole webhook request limit. |
Two consecutive snapshots are compiled floors for lost boundary, lead disappearance, and control
link failure. The two-pane-reads-per-tick limit is also compiled and cannot be weakened by config.
## 12. Delivery units and acceptance
### Unit 1 - Evidence model and fleet snapshot
Scope: health state model, fleet join, clocks, evidence retention, and pane budget.
Acceptance criteria:
1. One `agent.list` call covers one enabled fleet tick.
2. Pure decision tests cover every state and every evidence limit in Section 4.
3. `BUSY` plus stable raw `DONE` opens `TURN_BOUNDARY_LOST` after two snapshots.
4. Healthy fleet snapshots perform zero pane reads.
5. Pane cooldown, two-read fleet budget, and fair rotation cannot be disabled by config.
6. Logs are outputs only; no log parsing exists.
7. Fleet snapshots expose the same profile live-count calculation that placement uses.
8. Capacity rows report cap, live, free, and reclaimable values without opening incidents.
### Unit 2 - Lost boundary and task reconciliation
Scope: accepted-turn identity, guarded repair, target-wide teardown, release causes, and preserved
worktree discovery.
Acceptance criteria:
1. Every accepted send receives a stable `TurnToken` tied to target, exact waiter, session turn,
delivery baseline, and task outcome.
2. Repair requires the same `BUSY` token, two raw `IDLE` or `DONE` snapshots, no conflicting
observation, exact open waiter, successful baseline, and new recognised assistant output.
3. Missing, failed, late, or post-restart baseline never authorises repair.
4. Repair is enabled only for agent kinds with tested assistant-block extraction. Raw-text fallback
without a recognised marker refuses repair.
5. Normal completion and repair use one resolver guard core. Waiter, scrape, clipping, unchanged, and
exact-turn guards are not duplicated.
6. `reconcileLostBoundary` returns every typed result named in Section 8.3.
7. Only `REPAIRED` and same-turn `ALREADY_RESOLVED` may move the same turn to `DONE`.
8. A per-target reconciliation gate blocks a queued second send during repair and FSM update.
9. Repaired completion has distinct rendezvous kind, message outcome, poll source, lead marker, and
metric. Clipping keeps its extra marker.
10. Unchanged, unreadable, or ambiguous evidence gets at most one delayed retry. Missing capture or
baseline gets none.
11. Refusal leaves the ticket pending and tells the lead that no result was rebuilt or replayed.
12. Release, gone, never-ready, and abnormal stop use CB-568's one idempotent target-wide failure
operation.
13. The post-teardown invariant in Section 8.5 is tested independently of CB-568 internals.
14. A violated teardown invariant creates `DELEGATION_ORPHANED` and retries only the idempotent
failure operation.
15. `SPAWN_ROLLBACK` and normal `COMPLETED` remove worktrees. Abnormal and shutdown causes preserve
them.
16. Explicit stop is state-aware. Any pending task or non-terminal state preserves the worktree.
17. Atomic preserved-worktree manifests reload after restart and appear in lead-only
`bridge_list.preservedWorktrees`.
18. Manifest failure preserves the worktree and opens an operator-visible health failure.
19. Provision records the base commit. Terminal, long-idle worktrees report
`WORK_PRODUCT_AT_RISK` only under the evidence in Section 4.2 and never auto-delete work.
20. No path replays a delivered task, rebuilds its brief, or retargets it, even when prompt text is
available.
21. Tests cover both real traces, all repair refusals, clipping, explicit-reply and next-turn races,
restart without capture, concurrent send and release, and preserved discovery after restart.
### Unit 3 - Typed inbox and member routing
Scope: semantic record, AMQP migration, both adapters, member routing, polling, and member health in
`bridge_list`.
Acceptance criteria:
1. AMQP selects legacy or typed decoding only from `content_type`; it never sniffs the body.
2. Persistent `text/plain` from the old build becomes `kind=reply` with exact UTF-8 content,
including content beginning with `{`.
3. New entries use the vendor media type, `schemaVersion: 1`, UTF-8, persistent delivery, and AMQP
message ids.
4. Version 1 ignores unknown optional fields but rejects missing fields and identity mismatch.
5. Unknown versions are not decoded or acked. They remain on the original queue and create one
deduplicated failure.
6. Invalid known data never escapes the callback, appears as a reply, or blocks later valid messages.
7. Invalid data reaches durable per-target quarantine before original ack. Failed handoff leaves the
original unacked.
8. Decode failures create redacted WARN, metric, `bridge_list` summary, and routed incident without
raw content.
9. Both adapters pass one semantic contract for fields, FIFO, dedup, ownership, ack, and release.
10. Lead keys require explicit ownership. Publication never claims a queue.
11. Unit codec tests cover legacy `{`, Unicode, malformed UTF-8, typed round trip, additive fields,
malformed JSON, missing fields, identity mismatch, media type, version, and dedup.
12. A live broker contract writes old wire data and reads it with the new adapter after reconnect.
13. Live contract tests cover mixed entries, quarantine confirm-before-ack, unsupported redelivery,
later progress past poison, property persistence, lead ownership, and ack removal.
14. Safe downgrade is documented as unsupported.
15. RabbitMQ contract tests pass with `mvn test -Pcontract`. The same cases run once on production
LavinMQ, or the release states that LavinMQ was not checked.
16. Member incidents route to the exact delegating lead and never resolve a task rendezvous.
17. `bridge_list` shows compact member health and capacity without pane content. Member
`idleForSeconds` is present only when no accepted turn or inbox item exists.
### Unit 4 - Lead health and peer routing
Scope: lead evidence, exact ownership, peer selection, explicit-recipient push, and lead inbox
lifecycle.
Acceptance criteria:
1. Lead identity uses the `CallerResolver` supplier. Liveness uses successful current agent data.
2. Two successful-list absences with healthy ping become `LEAD_UNREACHABLE`; global link failure does
not.
3. Raw `WORKING`, raw `UNKNOWN`, first-seen time, failures, last success, and error class persist
across ticks.
4. Heartbeat and push publish status and nudge outcomes before safe no-injection decisions.
5. Dynamic lead identity survives a two-successful-snapshot retirement grace.
6. Member incidents first use exact delegation ownership with no singular-primary fallback.
7. Peer selection follows the exclusions, load rule, and stable tie break in Section 7.4.
8. A selected working peer is not interrupted. Its push waits for an injectable window.
9. Recipient assignment stays pinned. Reassignment increments generation and supersedes old pending
assignment.
10. `bridge_list` shows bounded foreign assignments, recipient, reason, and generation without pane
content.
11. `LeadInboxRegistry` owns configured and discovered lead keys before publication.
12. Missing leads keep ownership. Retirement needs an empty queue and handled incidents.
13. Replacement owns the new key before messages move. Non-empty in-memory keys are not released.
14. Tests cover dead versus busy, unknown, global failure, stale scan cache, disappearance, one peer,
several peers, reassignment, and no peer.
15. Adapter tests cover lead ownership, restart re-ownership, retirement, and terminal replacement.
LavinMQ is checked or named as unchecked.
16. A sole unreachable or stalled lead is never restarted or replaced. Without a sink, only passive
evidence remains and every coverage surface says so.
### Unit 5 - Human sink, hot config, metrics, and operator coverage
Scope: generic webhook, config split, incident journal, retry, resolve, metrics, example config, and
operator documentation.
Acceptance criteria:
1. `health.enabled` works without a human sink.
2. Notification mode is hot, defaults to disabled, and supports disabled or webhook.
3. Webhook mode requires a resolved environment value. Bad notification config does not disable an
already valid detector.
4. Mode changes keep open incidents. Re-enable resumes eligible incidents.
5. `bridge_list`, `/healthz`, metrics, and one WARN show partial coverage without a sink. HTTP
liveness behavior stays unchanged.
6. One-lead, no-sink coverage states that lead failure has no active notification or recovery.
7. Incident and outbound dedupe use the stable keys in Section 10.3.
8. The owner-only local journal survives restart and contains no pane or task content.
9. Journal failure keeps detection running but marks notification coverage degraded.
10. Retry tests cover network failure, timeout, 408, 429, `Retry-After`, 5xx, permanent 4xx, jitter,
delay cap, config re-arm, and one outstanding attempt.
11. Reminders and transport retries remain separate. Disabled mode does not build an unbounded queue.
12. Resolve sends only after an earlier human event succeeded. Resolve-before-delivery cancels stale
open delivery.
13. Metrics use bounded labels and exclude ids, URLs, and error text.
14. Payload tests reject every content type forbidden in Section 10.5.
15. Webhook response bodies are ignored and cannot direct recovery.
16. Tests cover disabled mode, one lead without sink, open/update/reminder/resolve, restart, dedup,
reassignment, disable/re-enable, and sink failure while local health continues.
17. `bridged.example.yaml` documents all hot keys and compiled floors.
18. The operator Features wiki is updated separately. The portable `CLAUDE.md` block is checked and
changed only if shipped tool or inbox semantics make it untrue.
19. `mvn clean install` passes.
## 13. Not checked and release gates
These limits are part of the design, not optional follow-up notes.
- **OpenCode pane status and assistant markers were not checked.** OpenCode lost-boundary repair is
disabled until live fixtures exist.
- **Permission-prompt status was not checked** for Claude Code or OpenCode. `BLOCKED` remains
ambiguous and has no automatic action.
- **`recent_unwrapped` stability was not checked** across all supported agent kinds. If normalisation
is not stable, `STALL_SUSPECTED` must say its evidence is weaker.
- **The real `BUSY + DONE` trace was not replayed against live herdr.** The design uses the observed
production trace and current poller behavior.
- **CB-568 was not present when Unit 2 was designed.** Unit 2 must inspect the landed API and keep its
independent teardown invariant.
- **Production LavinMQ was not checked.** Existing durable-inbox contracts use RabbitMQ. Migration,
quarantine, redelivery, lead ownership, and reassignment must run on LavinMQ before release or be
recorded as unchecked.
- **Live multi-lead routing was not checked.** Peer choice and reassignment are design rules backed by
fake-clock and adapter tests until a live exercise runs.
- **A live sole-lead failure with a webhook was not checked.** The no-peer path is a design result,
not a tested recovery.
- **No n8n, Slack, PagerDuty, or other receiver was checked.** The webhook remains generic and
outbound-only.
- **Deployment supervisor behavior for nested `/healthz` fields was not checked.** HTTP liveness
status stays unchanged to reduce this risk.
- **Incident-journal crash behavior was not checked** because the journal does not exist yet. Unit 5
must test atomic replacement and restart recovery.
- **Worktree merge state cannot be checked reliably** without forge or explicit collection evidence.
`WORK_PRODUCT_AT_RISK` stays a warning.
## 14. Locked exclusions
M4 does not expose `agent.read` as a bridge tool. It does not add a workflow engine, inbound n8n
authority, automatic task assignment, task replay, automatic lead replacement, or automatic member
spawn for free capacity.
The bridge remains a message bus with evidence and bounded mechanical repair. The lead remains the
place where judgement and work planning happen.