Compare commits

...

77 Commits

Author SHA1 Message Date
Dai Ha d04b075996 CB-563: mark clipped completion scrapes
CI / contract (pull_request) Successful in 42s
CI / build (pull_request) Successful in 56s
2026-08-15 04:25:54 +02:00
Dai Ha 644927636d Merge CB-562: the readiness gate says why it gave up (PR #33)
CI / build (push) Successful in 1m1s
CI / contract (push) Successful in 1m14s
When the injector's readiness grace expired it cleared the queue and logged
nothing. The failure then surfaced elsewhere as a turn-stall, which names the
wrong cause. Diagnosing CB-560 cost two live spawns and a wrong first
hypothesis for exactly this reason: the logs said "session marked failed" and
"failed send via turn-stall fallback", and neither says the message was never
typed into the pane at all.

The expiry now logs the target, the number of messages being failed, the grace
in polls and seconds, and the real cause in plain words.

The grace in seconds is derived, not written down twice: Bridged's own
INJECT_POLL_MILLIS is deleted and Injector.POLL_INTERVAL_MILLIS is the single
source, passed to every StatusPoller. A cadence change can no longer leave a
log line confidently stating the wrong duration.

Behaviour is unchanged. This is the first structured health event in the
daemon, and the foundation the fleet-health work will build on.
2026-08-15 04:21:29 +02:00
Dai Ha 95e45007aa CB-562: single source for the injector poll cadence; tighten count assertion
CI / build (pull_request) Successful in 59s
CI / contract (pull_request) Successful in 1m17s
2026-08-15 04:20:29 +02:00
Dai Ha 7a583c4045 CB-562: log why the readiness gate gave up on a target
CI / build (pull_request) Successful in 57s
CI / contract (pull_request) Successful in 1m4s
2026-08-14 21:52:25 +02:00
Dai Ha 65bce058bb Merge CB-561: one public way to build a CallerResolver (PR #32)
CI / build (push) Successful in 1m17s
CI / contract (push) Successful in 1m16s
CallerResolver had 5 public constructors and 4 public factories, and only one
of them could ever produce an architect. The rest defaulted memberSlotRoles to
`_ -> null`, so every architect quietly fell through to Principal.worker().
Nothing logged, nothing threw — the role was simply off.

That is the same failure shape as CB-560, so the fix is structural rather than
a warning: withLeadsAndMembers(.., MemberRegistry) is now the only public
construction path. Two overloads with no caller at all are deleted; the rest
are package-private and marked test-only. No path remains that accepts
architect bindings without a slot-role lookup, so no runtime WARN is needed.

Also checked and closed: the suspected slot leak on shutdown drain is not
real. SessionManager.release calls memberLifecycle.released() for every
removed session, and drainAll routes every session through release,
SPAWNING included. Verified in code.
2026-08-14 21:50:04 +02:00
Dai Ha 48b437083b Merge CB-560: mark spawned members present (PR #31)
Merging CB-548 made an architect resolve as Role.ARCHITECT, and BridgeMcp
marked MCP presence only for Role.WORKER. So an architect was never present.
That one line broke two things, because SessionManager.asPresence() is the
same object the injector's readiness gate reads:

  * the session never left SPAWNING;
  * Bridged.deliverableTo was false, so the injector held every delivery,
    waited out READINESS_GRACE_POLLS (~60s) and failed the send.

Confirmed live twice, on claude-code and on opencode. Both logged
"session marked failed" then "failed send via turn-stall fallback", neither
of which names the real cause.

Principal.isSpawnedMember() now means "a spawned member with a pane" —
WORKER or ARCHITECT, never PRIMARY. The Bridged.deliverableTo javadoc, which
still stated the worker-only rule as fact, is corrected: it is the clearest
description of this invariant anywhere, and leaving it stale is how the bug
comes back.
2026-08-14 21:50:04 +02:00
Dai Ha fa0612859b CB-560: document spawned member presence
CI / build (pull_request) Successful in 1m0s
CI / contract (pull_request) Successful in 1m4s
2026-08-14 21:48:58 +02:00
Dai Ha 4ffbcd0b7d CB-561: require member registry for architects
CI / contract (pull_request) Successful in 43s
CI / build (pull_request) Successful in 50s
2026-08-14 21:47:53 +02:00
Dai Ha a0cd053fd9 CB-560: mark architect members present
CI / build (pull_request) Successful in 55s
CI / contract (pull_request) Successful in 1m4s
2026-08-14 21:45:28 +02:00
Dai Ha f9a5e066b5 Merge PR #30: bind a spawned architect to its slot (CB-548's missing half)
CI / contract (push) Successful in 48s
CI / build (push) Successful in 1m18s
MemberRegistry.bind() had 18 call sites, every one in its own unit test.
Nothing in production ever bound a terminal to a slot, so CallerResolver's
architect branch was unreachable, every spawned architect resolved as WORKER,
and Authz's architect row was dead code. Since SEND is granted to
isPrimary() || isArchitect(), no architect could message anyone.

The spawn lifecycle now binds on both acquire paths and unbinds on release,
through a MemberLifecycle seam with a no-op default, so all six SessionManager
constructors are unchanged.

The escalation guard is deliberately double-sided. MemberRegistry flattens
every pool, so dev:* and reviewer:* slots sit in the same map the architect
resolver reads. The lifecycle binds only ARCHITECT roles, AND CallerResolver
independently checks roleForSlot(slot) == ARCHITECT before granting. Either
alone would be enough today; together, a regression in one cannot escalate a
worker. A reviewer confirmed a third barrier already existed: Authz gates
SPAWN on isPrimary(), so only a lead can request role=ARCHITECT at all.

Verified by the primary: mvn clean install, 640 tests, 0 failures.

Two findings recorded, neither blocking:

1. Known race, low severity. registry.put() and memberLifecycle.acquired()
   are not atomic. A release landing between them unbinds nothing (no binding
   exists yet), then acquire binds a dead terminal that no later release will
   ever clear - the slot leaks. Only the shutdown drain can reach a SPAWNING
   session (the idle reaper skips it, and an explicit stop needs a paneId the
   spawn has not returned), so the leak dies with the process. Note that
   binding before put does NOT fix it: the drain then iterates a roster the
   session is not in yet. A real fix needs atomicity.

2. The legacy Supplier-based withLeadsAndMembers overload and the Map-form
   constructors now silently disable architect resolution: memberSlotRoles
   defaults to _ -> null, so the architect branch can never be taken there.
   Production uses the registry form, and the tests were migrated, so nothing
   fails - but a future caller gets workers with no error.
2026-08-14 21:25:45 +02:00
Dai Ha e9bc192160 CB-548: bind architect sessions to member slots
CI / contract (pull_request) Successful in 45s
CI / build (pull_request) Successful in 55s
2026-08-14 21:10:48 +02:00
Dai Ha e0a57988ad Merge PR #29: drop .claude/settings.local.json from the default parityOverlay
CI / build (push) Successful in 1m21s
CI / contract (push) Successful in 1m21s
A member no longer receives the primary's pre-approved tool grants by default.
The grants were inert — GitWorktrees.isolateToolSurface already strips each
worktree's .mcp.json to an empty server map — so this is defence in depth: two
independent guards instead of one. Same reasoning as CB-525, which removed the
sibling .mcp.json and left this file behind.

An explicit parityOverlay: in config is unaffected; only the default changes.

Verified by the primary: mvn clean install, 637 tests, 0 failures.
2026-08-14 21:02:25 +02:00
Dai Ha b414a74c26 CB-559: drop .claude/settings.local.json from the default parity overlay
CI / build (pull_request) Successful in 54s
CI / contract (pull_request) Successful in 1m7s
2026-08-14 20:56:16 +02:00
Dai Ha 2c3796d598 Merge: keep every credential in one store
CI / build (push) Successful in 55s
CI / contract (push) Successful in 1m13s
2026-08-14 20:38:36 +02:00
Dai Ha 7e0ff9ab06 Keep every credential in one store, not a per-repo copy
The repo carried a gitignored .secrets/ directory with four files. Two of them
(context7-token, gitea-token) were byte-identical copies of variables the login
shell already exported. One (gitea-host) is not a secret. The fourth
(worker-gitea-token) was the only copy anywhere, and nothing exported it, so
bridged read gitTokenEnv from an environment that never had it and every worker
push got an empty token.

All four values now live in the operator's single sourced secrets file, verified by
sha256 before the copies were removed. opencode.json reads them as {env:...}, which
.mcp.json already did. A second copy of a secret is the problem: the copy you forget
is the one that leaks or goes stale.

This makes worktree isolation load-bearing rather than a workaround. opencode.json is
tracked, so it lands in every worktree. It used to fail there, because {file:.secrets/}
pointed at files a worktree never receives and OpenCode refuses to start on a dangling
reference. With {env:...} the reference resolves, and a member would silently inherit
the primary's admin-scoped GITEA_ACCESS_TOKEN. GitWorktrees already neutralizes the
file; only its stated reason changes, and it is now a confidentiality boundary.

The port-to-opencode skill taught {file:.secrets/} as the preferred pattern, so it is
rewritten to teach the central store and to say why we moved. .gitignore keeps the
.secrets/ line as a backstop against habit.

Includes the wiki pointer, which also carries the CB-559 config-reload correction.
2026-08-14 20:38:31 +02:00
Dai Ha 6d1565d8b6 Merge CB-559 correction: profile launch settings are deferred, not hot 2026-08-14 20:37:47 +02:00
Dai Ha e18e002d2f CB-559: a profile's launch settings are deferred, not hot
The shipped docs and javadoc said a profile's `model` and `tabLabel` take effect on
the next spawn. They do not, and ConfigRef did not detect the change either, so a
reload logged a clean "config reloaded" and silently did nothing. That is the worst
outcome a reload can produce: the operator has no reason to doubt it.

What makes a key hot is who reads it and when, not that it is config. Placement
reads weight and maxLoad through a supplier on CompositePeerLauncher, so those are
genuinely hot. HerdrPeerLauncher takes Map.copyOf(profiles) at construction and
resolves every spawn out of that copy, so model, baseUrl, argv, env and the rest
cannot move until the daemon restarts.

changedDeferredKeys now compares every launch component of an existing profile,
excluding weight and maxLoad, and names the profiles that need a restart. The
javadoc and bridged.example.yaml say the same thing. Two tests pin the pair:
weight/maxLoad reports nothing deferred, a changed model reports the profile by name.
2026-08-14 20:37:42 +02:00
Dai Ha c18572ea9d Bump wiki to 8b1eb68 — the CB-557/558/559 Features entries
CI / contract (push) Successful in 45s
CI / build (push) Successful in 1m14s
2026-08-14 19:52:02 +02:00
Dai Ha e9d02bc5e8 Merge CB-557/558/559: the fleet block, role pools, lead auto-launch, config reload
Four commits that together reshape how the daemon is told who it may run.

CB-557 (schema) — one `fleet:` block replaces `leaders:`, `members:`, `leadScan:`
and `defaultProfile:`. Role is the containing map key, not a `role:` field, so a
misspelled pool name declares nothing instead of producing a member with no
contract. Tab labels are role-first and counted per role+profile.

CB-557 (routing) — an unqualified spawn picks its profile from its role's pool.
Role and profile stay orthogonal: a reviewer may run on the same profile as the
dev whose diff it reads, so the two cannot be one field. An explicit
`bridge_spawn{profile:…}` stays exempt, because it carries no role and would
otherwise be refused for a profile the operator named.

CB-558 — the daemon launches a declared lead when fewer than `instances` are
running. An auto-launched lead is NOT a member: no reply charter, never
registered with SessionManager (the idle reaper would kill it), and no Anthropic
binding in its env. A lead counts as live only when herdr reports a running
agent, so one crash does not disable auto-launch forever, and an unreachable
herdr starts nothing at all.

CB-559 — `bridged.yaml` can be re-read without a restart, opt-in through
`configReload:`. Keys are hot (`fleet:`, `placement:`, an existing profile's
fields), deferred (lifecycle, guard, adding a profile) or cold (bind,
herdrSocket, broker, auth). A cold change refuses the WHOLE reload rather than
half-applying it, because a daemon matching no file on disk is worse for an
operator than no reload at all.

634 tests, IDE diagnostics clean. Verified live against the running daemon: the
lead launcher recognised the existing lead and started nothing; the reload
applied a hot change, refused a `bind:` change exactly once (not once per tick),
and reverted cleanly.
2026-08-14 19:51:34 +02:00
Dai Ha a2108a8a14 CB-559: re-read bridged.yaml without restarting the daemon
Tuning a fleet meant restarting bridged, and a restart tears down every lead
and worker it owns. Changing one pool's weight cost the whole fleet's state,
so in practice nobody changed it.

ConfigRef holds the live BridgedConfig in an AtomicReference. Consumers read
it at the point of use, so a change reaches the next spawn with nothing
rebuilt. The launchers that used to capture config into fields now take
suppliers: the fleet tabLabel template, the profile map, the placement policy
and the fleet block.

Keys fall into three classes, and the difference is what already exists when
the reload happens:

  hot       fleet: (pools + tabLabel), placement:, and an existing profile's
            weight / maxLoad / model / tabLabel — live on the next spawn.
  deferred  lifecycle:, leadHeartbeat:, guard:, worktreeRoot:, spawnReady*,
            and adding/removing a profile — accepted, but the startup wiring
            keeps the old value. The reload logs these by name.
  cold      bind:, herdrSocket:, broker:, auth: — refuses the WHOLE reload.

A cold change refuses everything rather than applying the hot half. A
half-applied reload leaves the daemon matching no file on disk, which is the
worst thing a reload can do to an operator reading that file to work out what
the daemon is doing. Refusing keeps the invariant that the live config is
always some version of the file.

A parse failure or a failed startup validator is refused the same way, and
the running config stays live: a file being saved is sometimes read
mid-write, and degrading a working daemon over a half-written file is a bad
trade. The same four validators startup runs are re-run, so a config that
could not have booted cannot slip in through a reload.

ConfigWatcher polls the modified time on a daemon thread, opt-in through
configReload.enabled (default off, so an upgraded daemon is unchanged). It
stamps the timestamp BEFORE reloading, so a refused file is not retried every
tick — the next save earns a fresh attempt. A missing file is skipped
silently, because editors unlink briefly mid-save.

MicroProfile Config was the first idea and does not fit: @ConfigMapping needs
interfaces, resolves once at bootstrap, and reload would still mean rebuild
and swap. The port would also lose the raw-YAML duplicate-key detection,
since duplicates have already collapsed once the tree is flattened to
properties.

634 tests.
2026-08-14 18:23:32 +02:00
Dai Ha 57f8fa257a CB-558: launch a declared lead at startup when none is live
`fleet.leaders.<name>.instances` was descriptive. Now the daemon reads it: a
lead that names a `profile:` is started when fewer than `instances` are running.
A lead with only a `terminal:` stays recognise-only, as before.

A lead is not a member, and LeadLauncher exists to keep it that way. Every other
spawn path goes through HerdrPeerLauncher, which does three things a lead must
never get: it appends the worker reply charter ("you are an off-subscription
worker … end every turn with bridge_reply" — the opposite of an orchestrator);
it registers the session with SessionManager, whose idle reaper would kill a
lead for being idle, which is a lead's normal state; and it can move a peer off
the subscription. So this launcher talks to AgentControl/WorkspaceControl
directly. The duplicated argv/env assembly is the cheaper half of that trade.

Not double-spawning is the safety property, so liveness needs two pieces of
evidence. A running agent in a tab labelled `lead: <name>` finds an
auto-launched lead. A running agent on a pinned `terminal:` finds one the
operator opened by hand — without it, a pinned lead whose tab carries no
matching label would be relaunched on every boot. Member workspaces are
excluded, so a member in a matching tab is never counted. If herdr cannot be
reached, nothing is started: a second orchestrator is worse than none.

Liveness deliberately requires the AGENT, not just the label. LeadTabScanner
used to promise that bridged never writes a lead label, so there was no
round-trip from the daemon's own rename back into its next decision. That is no
longer true, and its javadoc now says so. The trust direction is unaffected — a
label is a name, not a capability — but staleness becomes real: a label left by
a crashed session would otherwise read as a live lead forever and disable
auto-launch permanently.

Two new knobs. `workspace:` (default "leads") is where a launched lead's tab
goes; it must not be a member workspace, because those are excluded from the
scan and a lead placed in one would never be found again. `cwd:` defaults to
bridged's own working directory.

Also: WorkspaceControl.listTabs, and a FakeHerdr tab seeder that leaves the
canned response byte-identical when no tab is seeded.

617 tests pass (16 new), IDE-clean.
2026-08-14 16:47:57 +02:00
Dai Ha 61944fc045 CB-557: place an unqualified spawn inside its role's pool
The pools were config-only until now: the launchers still received one global
effectiveDefaultProfile and placement still ranged over every configured
profile, so a reviewer could be placed on an architect-only backend.

Three parts:

SessionManager computed the role, stored it on the MemberSession, and never put
it on the SpawnRequest. So the role reached the record that describes the spawn
but not the call that performs it — every launcher saw DEV. Both spawn paths
(plain and worktree) now carry it.

CompositePeerLauncher takes the Fleet and draws its candidates from
fleet.<role> instead of from all profiles. An absent or empty pool means
unconstrained, not blocked: a config that declares pools for some roles must
keep spawning the rest, so it falls back to every profile. A null Fleet is the
pre-CB-557 wiring and behaves exactly as before.

Bridged passes cfg.fleet() to the composite and cfg.fleet().tabLabel() to both
launchers. The tab-label knob was accepted by HerdrPeerLauncher but passed by
nobody, so it was inert — the label only looked right because the fallback
happened to match the configured template. Four tests now pin the wiring
instead of the coincidence.

An EXPLICIT profile stays exempt from the pool. `bridge_spawn{profile:"opus"}`
carries no role, so it defaults to DEV; judging it against the dev pool would
refuse a spawn the operator asked for by name. maxLoad still applies to it.

Also cleared the IDE warnings in the touched files: an immediately-rethrown
catch (the comment stays, the redundant block goes), unused lambda params, a
javadoc link to a package-private class, two unused imports.

601 tests pass.
2026-08-14 16:37:29 +02:00
Dai Ha 4b48d2d921 CB-557: fleet role pools — role is the config key, and the tab label says it
Four top-level keys (leaders:, members:, leadScan:, defaultProfile:) become one
`fleet:` block, and a member's role becomes the map key that contains it rather
than a `role:` field inside it.

Why the key and not a field: a misspelled `role: architct` used to produce a
member with no contract, which nothing rejected. A misspelled pool name declares
nothing, which is a shape the loader can see.

`fleet.architects/developers/reviewers` are pools of profiles a role MAY run on.
That replaces the single global `defaultProfile:`, so an unqualified spawn now
resolves its profile from the pool of the role it asked for. Role and profile
stay orthogonal: a reviewer may run on the same profile as the dev it reviews,
and one profile may appear in several pools.

Tab labels are role-first — `dev: sonnet #4`. The template lives on `fleet:`
because a profile cannot know the role of the member launched on it; a profile
may still override it. The `{n}` counter is scoped per role+profile, so a dev
and a reviewer on one profile each start at #1. Making {role} the first field
also turns the lead/member namespace check into a structural guarantee: roles
are a closed enum, so only hand-written templates can still collide with a lead
tabPrefix.

Removed keys are hard errors that name their successor. `defaultProfile:` has no
single successor key, so its message explains the new model instead of pointing
at a key that does not exist.

Map order is kept with LinkedHashMap, deliberately not Map.copyOf — the latter
salts iteration order per JVM run, which would destroy the YAML definition order
that `placement: fixed` selects on.

Not yet wired: SessionManager still hands the launchers one effectiveDefault-
Profile, so pools are not enforced at spawn time yet, and placement still ranges
over all profiles.

595 tests pass.
2026-08-14 16:32:54 +02:00
Dai Ha 3aa145cbef Merge CB-557: member taxonomy in the API surface
CI / contract (push) Successful in 1m0s
CI / build (push) Successful in 1m1s
2026-08-14 07:08:26 +02:00
Dai Ha 4875127daa CB-557: adopt the member taxonomy in the API surface
Every spawned peer is now a member with a role, and the role travels with it
from the spawn call to the roster.

MCP:
  bridge_spawn gains role: architect | dev | reviewer (default dev). An
    unknown role is refused with the valid spellings in the message.
  bridge_list returns "members" instead of "workers"; each row carries both
    role (what it is for) and profile (which backend it runs on).
  The spawn result echoes the role back, so a spawn that fell back to dev is
    visible rather than silent.

REST:
  GET/POST /members and DELETE /members/{paneId} replace /workers.
  POST accepts role= as a query param or a body field; an unknown role is 400.

Code:
  dev.ltms.bridged.worker package -> dev.ltms.bridged.member
  WorkerSession   -> MemberSession, plus a MemberRole role component
  WorkerPresence  -> MemberPresence
  SessionManager.acquire gains a role parameter; the existing overloads keep
    working and default to DEV, which is exactly what "worker" used to mean.

ClaudeCodeLauncher and OpenCodeLauncher keep their names on purpose — they
are named after the backend, not the role.

Not done here: the launch charter is still one string for every role, so a
member is told its role by nobody yet. That is the next ticket.

mvn clean install: 583 tests, 0 failures, 0 errors, BUILD SUCCESS.
2026-08-14 07:08:26 +02:00
Dai Ha d14a624421 Merge CB-557: member taxonomy in config
CI / contract (push) Successful in 45s
CI / build (push) Successful in 1m32s
2026-08-14 07:02:22 +02:00
Dai Ha 246f50b778 CB-557: adopt the member taxonomy in config
A member is anything a lead spawns. Every member carries two independent
attributes:

  role    — which contract: architect, dev or reviewer. It picks the launch
            charter, the role file, the playbook skill and the authz row.
  profile — which backend: model, CLI adapter, credentials, cost.

They vary on their own. A reviewer may run on the same profile as the dev
whose diff it reads, which is the case that proves the two cannot be one
field.

Config changes (breaking — we are in active development, so no aliases):

  workers:       -> profiles:        it was never a list of workers; it is a
                                     catalogue of backends
  defaultWorker: -> defaultProfile:
  architects:    -> members:         each slot now names its role

An old config is rejected at load with the new key named, rather than being
warned about once and then running with zero profiles — that failure would
surface much later, at the first spawn, pointing nowhere near the cause.

Also:
  - BridgedConfig.Worker    -> BridgedConfig.Profile
  - ArchitectRegistry       -> MemberRegistry
  - new peer.MemberRole enum, validated at startup
  - profiles map is normalized once in the compact constructor, so the raw
    map and the derived one can no longer disagree
  - the legacy singular worker: block is dropped
  - workerProfiles() -> profiles(); defaultProfile() -> effectiveDefaultProfile()
    (the record component now owns the plain name)

mvn clean install: 578 tests, 0 failures, 0 errors, BUILD SUCCESS.
2026-08-14 07:02:17 +02:00
Dai Ha 15b53c6cfa Merge CB-553: enforce maxLoad on explicit-profile spawns
CI / contract (push) Successful in 46s
CI / build (push) Successful in 1m8s
2026-08-14 06:51:21 +02:00
Dai Ha 3a7ef0adbd CB-553: enforce maxLoad on explicit-profile spawns (no cap bypass) 2026-08-14 06:50:28 +02:00
Dai Ha 293a305748 Merge CB-551: idle-lead heartbeat
CI / build (push) Successful in 1m15s
CI / contract (push) Successful in 1m21s
2026-08-13 21:21:33 +02:00
Dai Ha e5038c6d13 CB-551: idle-lead heartbeat — nudge an idle lead back to work on a timer
CI / contract (pull_request) Successful in 43s
CI / build (pull_request) Successful in 56s
2026-08-13 21:10:21 +02:00
Dai Ha 13ea79f6fd CB-523 (opencode): generated worker config pins compaction.auto
CI / build (push) Successful in 56s
CI / contract (push) Successful in 1m10s
A worker that exhausts its context dies mid-turn and its bridge_reply — the
whole point of the turn — dies with it. Auto-compaction is a condition of the
turn contract for a spawned peer, not an operator preference.

OPENCODE_CONFIG is merged over ~/.config/opencode/config.json rather than
replacing it, so workers already inherited auto:true from the home file. That
inheritance is exactly what this removes as a dependency: the home file is
outside this repo, differs per machine, and is not ours to rely on.

The trade is recorded in the code: an OPENCODE_CONFIG value overrides the home
value, so an operator cannot disable compaction for bridged workers from home.
Deliberate for peers we spawn and whose turns we must land; per-profile control
would be a profile knob, not the removal of this line.

Origin: fleet-wide auto-compaction audit by peer lead gpt-sol-5.6, which found
autoCompactEnabled:false alongside a 300k window across the Claude instance
configs. Handed over uncommitted; taken deliberately, with rationale added.
2026-08-13 21:03:52 +02:00
Dai Ha f55d3c203b Merge CB-552: supersede rejected orchestrator design; sync wiki Features
CI / build (push) Successful in 1m0s
CI / contract (push) Successful in 1m9s
2026-08-13 20:59:50 +02:00
Dai Ha 37c4d47d3a Merge CB-544: shutdown drain preserves worker worktrees
CI / build (push) Successful in 55s
CI / contract (push) Successful in 1m17s
2026-08-13 20:58:06 +02:00
Dai Ha 04e21c9243 CB-544: shutdown drain preserves worker worktrees (no data loss)
CI / contract (pull_request) Successful in 41s
CI / build (pull_request) Successful in 1m29s
2026-08-13 20:36:32 +02:00
Dai Ha 83cac07f6e CB-552: update wiki documentation pointer 2026-08-13 20:19:46 +02:00
Dai Ha f472e0f782 CB-552: supersede rejected orchestrator design 2026-08-13 20:19:46 +02:00
ltms 91c9f981c5 Merge CB-548 send ownership and rendezvous hardening
CI / build (push) Successful in 1m16s
CI / contract (push) Successful in 1m17s
2026-08-13 19:44:02 +02:00
Dai Ha dd906526c0 CB-548: guard primary singleton to PRIMARY callers; open waiter before enqueue
CI / contract (pull_request) Successful in 42s
CI / build (pull_request) Successful in 53s
Fix two handler-level bugs found in PR #21:
- Only PRIMARY callers may update PrimaryRegistry.record (the legacy singleton 'primary'
  fallback for no-delegation inbox nudges). An architect SEND previously recorded its terminal
  as the fallback; the per-target delegation map does not cure the singleton. New
  BridgeMcp.recordPrimarySingleton uses the resolved role (caller.isPrimary()) — named leads
  (PRIMARY) still record, architects never do.
- MessageService.send now opens the rendezvous waiter BEFORE queueing delivery, fixing both the
  enqueue-before-open fast-reply race (a fast reply no longer orphans into the inbox) and
  callback-failure ordering: a throwing onAccepted (public callback) fails the send cleanly with
  no stale waiter and no queued, orphanable message.

Tests: architect SEND vs lead SEND primary-singleton regression; throwing onAccepted leaves no
stale waiter or queued orphan.
2026-08-13 19:42:48 +02:00
Dai Ha ec3001796a CB-548: record delegator ownership only when a send is accepted
Record PrimaryRegistry delegator ownership via a MessageService accepted-delivery
hook (won the session lock + queued delivery), never at bridge_send request time, so
a concurrent sender that times out BUSY cannot steal a live turn's reply routing.
Make Rendezvous.open atomic fail-if-present so a double open trips loudly instead of
replacing the waiter another send is blocked on. Answering a bridge_ask keeps the same
ownership (no rewrite). Adds ownership/rendezvous regression tests.
2026-08-13 19:42:48 +02:00
ltms 86cf4c285a CB-547b: opencode session identity — resume (-s) + post-hoc discovery (#18)
CI / build (push) Successful in 48s
CI / contract (push) Successful in 1m17s
2026-08-13 19:41:06 +02:00
Dai Ha cd18887b69 CB-547b: opencode session identity — resume (-s) + post-hoc session discovery
CI / contract (pull_request) Successful in 42s
CI / build (pull_request) Successful in 1m16s
2026-08-13 19:40:17 +02:00
ltms 4637c68295 CB-543: neutralize worktree-hostile tracked configs (#22)
CI / build (push) Successful in 55s
CI / contract (push) Successful in 1m14s
2026-08-13 19:39:54 +02:00
Dai Ha 0114bd1fa7 CB-543: neutralize tracked opencode.json and .autoenv in provisioned worktrees
CI / contract (pull_request) Successful in 45s
CI / build (pull_request) Successful in 1m26s
2026-08-13 19:30:19 +02:00
ltms 6dee84ca71 Merge CB-548 architect config and authorization foundation
CI / build (push) Successful in 57s
CI / contract (push) Successful in 1m17s
2026-08-13 18:33:11 +02:00
Dai Ha f004a0c654 CB-548: make duplicate-architect-slot detection top-level-only and depth-safe
CI / contract (pull_request) Successful in 42s
CI / build (pull_request) Successful in 1m13s
2026-08-13 18:31:29 +02:00
Dai Ha 6123576c68 CB-548: correct architect premise — profile-only slots, registry-owned bindings, dup-key rejection
CI / build (pull_request) Successful in 54s
CI / contract (pull_request) Successful in 1m6s
2026-08-13 18:05:53 +02:00
Dai Ha 21cfc09f8e CB-548: config-declared architect slots + Role.ARCHITECT authz
CI / build (pull_request) Successful in 55s
CI / contract (pull_request) Successful in 1m6s
2026-08-13 17:28:14 +02:00
ltms 509530e235 CB-547a: peer-neutral session-identity contract + Claude Code adapter (#17)
CI / build (push) Successful in 1m0s
CI / contract (push) Successful in 1m2s
2026-08-13 16:56:32 +02:00
Dai Ha d146a01422 CB-547a: peer-neutral session-identity contract + Claude Code adapter
CI / contract (pull_request) Successful in 44s
CI / build (pull_request) Successful in 1m31s
2026-08-13 16:43:47 +02:00
Dai Ha 91332723db Bump wiki to 8c63db5 — the CB-539/CB-542 docs
CI / build (push) Successful in 1m18s
CI / contract (push) Successful in 1m52s
Pairs with 244fbd9. The docs half lived on the submodule's own remote, so the
pointer bump is separate by necessity, not by preference.

The substantive part is not the new 'Run a worker on the subscription' entry but
the correction beside it: 'Give workers a toolchain' asserted flatly that an env:
entry cannot repoint a worker past the SubscriptionGuard. CB-539 made that false,
and the wiki went on claiming it until CB-542 closed the hole. The entry now
states the rule and its one exception together.
2026-08-13 15:32:48 +02:00
ltms 244fbd98a5 Merge CB-539 + CB-542: subscription profiles, with the env: bypass closed
CI / contract (push) Successful in 40s
CI / build (push) Successful in 50s
Verified by the lead in a clean worktree at 5afe8e1 rather than on the worker's
report: Tests run: 474, Failures: 0, Errors: 0, Skipped: 0, BUILD SUCCESS (main
was at 464).

The invariant this lands: there is no configuration in which a worker reaches an
Anthropic endpoint that no guard vetted. Closed at two layers — a fatal, profile-
naming refusal at config load, and a launcher-side strip so it holds for profiles
built in code that never passed validation.

The wiki half is a separate branch on the submodule's own remote; the pointer bump
follows as its own commit on main.
2026-08-13 15:31:51 +02:00
Dai Ha 5afe8e14d9 CB-542: close the subscription env: bypass, and fix the env: env-carries-anthropic docs
CI / contract (pull_request) Successful in 44s
CI / build (pull_request) Successful in 51s
2026-08-13 15:26:20 +02:00
Dai Ha 73f6b12dd8 CB-539: let a worker profile run on the subscription, explicitly
Add a per-profile subscription: true opt-in that lets a claude-code worker run on
the operator's Claude subscription when there is no off-subscription endpoint for it
(e.g. sonnet on ccs). When set, the launcher injects neither ANTHROPIC_BASE_URL nor
ANTHROPIC_AUTH_TOKEN and skips SubscriptionGuard's base_url requirement for that
profile only, logging a WARN naming the profile. subscription: true alongside a
baseUrl is refused as contradictory. The default (absent/false) keeps today's hard
refusal unchanged; every other profile stays allowlist-checked and SubscriptionGuard
is untouched.
2026-08-13 15:15:14 +02:00
Dai Ha 793f2e7157 Revert "Commit .autoenv" — it wedges every worker spawn
CI / build (push) Successful in 53s
CI / contract (push) Successful in 1m24s
Regression I introduced one commit ago, caught by two consecutive spawn failures
and reproduced end-to-end.

Tracking `.autoenv` means git checks it out into every provisioned worktree. A
worktree is a new path, and autoenv authorizes by path, so the file is always
unauthorized there — autoenv prints "[autoenv] Authorize this file? (y/n/d)" and
blocks on `read` (activate.sh:211-222). The pane's shell sits at that prompt, so
`ccs <profile>` never runs, the peer never becomes injectable, and CB-306's
readiness gate closes the pane after 20s. Symptom is a bare "spawn timed out";
nothing names autoenv, which is what made it worth writing down.

Confirmed rather than inferred: the peer lead's cb-537 worker, spawned BEFORE
f8182e4, has no `.autoenv` in its worktree and is still alive; a throwaway
worktree at HEAD reproduces the prompt on entry.

The file is still worth committing — the reasoning in f8182e4 stands, and it is
recoverable from there. What is missing is the other half: GitWorktrees already
neutralizes `.mcp.json` in a provisioned worktree ("worker tool surface is
launcher-mounted only"), and `.autoenv` needs exactly the same treatment for
exactly the same reason — a worker's environment is launcher-supplied, never
repo-supplied. Re-land it with that, tracked as CB-543.

Rejected the quicker fix of exporting AUTOENV_ASSUME_YES: it auto-approves
arbitrary repo-controlled shell code in every spawned worker, which is a worse
trade than one unset variable.
2026-08-13 15:07:59 +02:00
Dai Ha cf48983046 Merge CB-538: opencode peers launch with --auto
CI / contract (push) Successful in 50s
CI / build (push) Successful in 1m30s
A spawned peer has no human at its pane, so an approval prompt is not a pause —
it is a wedge. The agent stops, looks identical to a legitimate mid-turn wait,
and can never reach its bridge_reply, so the delegation dies silently and the
lead learns nothing until the timeout.

Unconditional rather than a per-profile knob, which is the right call: there is
no configuration under which a bridge-spawned opencode worker WANTS to block on
an approval it has no way to answer. opencode's help calls --auto 'dangerous!',
and that warning is written for a human at a terminal; the blast radius here is
already bounded by the layer above — a worker runs in its own git worktree, on
its own branch, off-subscription, and cannot merge. The lead is the gate.

Reviewed by me rather than fanned out: 24 lines across one method and its two
tests, below the threshold where a reviewer pass pays for itself.

argvWithModel is re-signatured to take composed argv instead of building it,
so the two flag-appenders compose rather than each owning construction.
2026-08-13 15:03:59 +02:00
Dai Ha d038516750 Bump wiki to 0a21b49 — multi-lead arc, and opencode's secret supply
CI / build (push) Successful in 53s
CI / contract (push) Successful in 1m48s
Moves the submodule pointer 4320c1c -> 0a21b49, matching the docs to the code in
bc13b8e rather than leaving chapter 11 describing a fleet with one lead in it.

  7b5381b  11, 7: CB-530..536 — leaders:, leadScan:, lead-to-lead messaging,
           lead deliverability, leads in bridge_list; 7's copy of the portable
           block re-synced byte-identically
  0a21b49  12: where opencode's secrets actually come from, and why {env:...}
           is green under Claude Code and broken in a terminal

Both are already pushed to the wiki's own remote, so this pointer resolves for
anyone who clones — the ordering that matters, and the reason the wiki went
first.

This is the pointer only. `wiki/` is a submodule with its own remote and its
content is never committed here; bumping the recorded SHA in its own commit is
how this repo has always recorded a docs update (see "bump wiki to 4320c1c").
2026-08-13 10:44:38 +02:00
Dai Ha f8182e4514 Commit .autoenv — the loader that keeps Claude Code on one secret store
This file has always been designed to be committed and says so in its own header;
it simply never was, so every clone of this workspace has been reconstructing it
by hand or duplicating tokens instead.

It holds no secret. It reads `.secrets/` (gitignored, 0600) and exports three
variables, because Claude Code expands `${VAR}` in `.mcp.json` from the *process*
environment and cannot read a file — so without it, CONTEXT7_TOKEN and the gitea
pair must be duplicated as literals in `.claude/settings.local.json`. opencode
needs none of this: it reads `.secrets/` directly via `{file:...}`.

Verified before committing that no value appears in it, only the three names and
the paths they are read from.

Safe in a worktree by construction: a worktree receives tracked files only, so
`.secrets/` is absent there and the whole block is skipped rather than failing.
Workers are fed by the launcher's env instead — which is where the name mismatch
documented in wiki chapter 12 (GITEA_TOKEN vs GITEA_ACCESS_TOKEN) has to be
reconciled.
2026-08-13 10:44:21 +02:00
Dai Ha bc13b8e92c CB-530..536, CB-538 groundwork: leads as peers, and a fleet that can find itself
Two leads now work as peers rather than one primary plus workers. The arc:

  CB-530/531  lead identity: `leaders:` names panes, `leadScan:` discovers them by
              tab label (LeadTabScanner, TTL-cached, worker spaces excluded).
  CB-532      leads can message each other AND be answered. Principal.leader now
              carries its terminal, so ownsSession() can be true for a lead; the
              "and you must be a worker" conjunct beside it protected nothing.
              Retires `primary:` — reply nudges follow the delegating lead, a
              binding recorded at bridge_send where both halves are known.
  CB-533      ClaudeCodeLauncher passes --model. argv is usually a wrapper
              (`ccs <profile>`) that re-exports its own model family, so
              ANTHROPIC_MODEL alone was silently overruled.
  CB-534      a lead is deliverable. The CB-113 readiness gate only opened for
              terminals in WorkerPresence, which only workers ever enter, so every
              lead->lead send waited out the ~60s grace and failed having never
              been typed. The gate guards a *spawned* peer's boot window; a lead
              is never spawned.
  CB-535      bridge_list returns `leads` alongside `workers`, with `self` on the
              caller's row. An empty worker roster no longer reads as "no peers".
  CB-536      CLAUDE.md: lead<->lead is coordinate-only, never sideways delegation.
              Propagated byte-identically to wiki/7-Use-Cases.md.

MIXED PROVENANCE — recorded deliberately rather than hidden. This tree also carries
in-progress CB-537 (context separation) authored by the peer lead gpt-sol-5.6 and
its worker: Capability.CONTEXT_RESET, SessionManager.clearAfterTurn, and the
Injector/TurnListener/CompletionResolver/launcher changes around it. That work was
done in this shared working tree rather than a worktree, and is entangled with the
above in BridgedConfig.java, Bridged.java and ClaudeCodeLauncher.java, so neither
lead could stage its own half without sweeping in the other's. Committing the whole
green state is the honest resolution; the peer branches from here.

Note for whoever picks CB-537 up: the design in this commit is SUPERSEDED. Both
leads agreed to replace the global `clearAfterTurn` boolean with per-delivery
policy (inherit|fresh|thread) applied PRE-delivery, because a post-turn reset races
by construction — Injector.onStatus clears awaitingCompletion and dequeues the next
message in the same tick. `fresh` is also a correctness guarantee, so an adapter
without a reset capability must refuse it rather than log a no-op.

mvn clean install: Tests run: 464, Failures: 0, Errors: 0, Skipped: 0. BUILD SUCCESS.
2026-08-13 09:38:24 +02:00
Dai Ha fbe79258bf CB-538: opencode peers launch with --auto 2026-08-13 07:19:19 +02:00
Dai Ha ef1e014b41 CB-527: ship the bridge as an installable Claude Code plugin
CI / build (push) Successful in 1m0s
CI / contract (push) Successful in 1m21s
The orchestration contract had no distributable form. Every consuming project
hand-copied a block of CLAUDE.md and hand-wrote an .mcp.json, and we keep a
script whose only job is to notice those copies drifting apart. A plugin is
versioned, installed once, and updates in place.

Ships no credentials, deliberately: every secret is referenced by environment
variable NAME and the value never enters a file, which is what makes the
artifact safe to publish. The setup skill states the two rules that are easy to
get wrong for the right-sounding reasons — the PR token must not be able to
merge (a worker opens, the primary gates), and ANTHROPIC_BASE_URL must never be
set by setup, because mounting the bridge must not move a session off
subscription.

The plugin root is plugin/, not the repo root. An installed plugin's .mcp.json
is a committed file, while this repo's root .mcp.json is local-only and
--skip-worktree; rooting the plugin at the repo would commit the primary's IDE
servers and hand them to every worker — the exact failure CB-525 exists to
prevent.

Scope is client-side setup only. herdr and bridged stay separate services with
their own lifecycles, and the skill refuses to install them rather than guess.
It also refuses to accept /healthz as proof: health reports only that the daemon
can reach herdr, and CB-521 showed it staying green while every spawn failed, so
verification ends with a real spawn.

Both manifests pass `claude plugin validate --strict`.
2026-08-10 19:55:17 +02:00
Dai Ha e3c8393d1b CB-527: retire the ollama profile from the worker choice
The ollama backend is decommissioned, so the example config stops pointing
readers at a dead host and the guard allowlist stops carrying an entry with
no profile behind it — a stale entry there is dead permission, and that list
is the only thing keeping a worker off the primary's subscription.

The second illustrative profile survives as gx11: the example exists to show
`placement: weighted` having something to choose between, and a one-profile
example would quietly stop demonstrating that.

It also moves the CB-523 auto-compact override onto the surviving profile.
That guard had been attached to `ollama` alone, so retiring the profile would
have removed the fleet's only protection against the failure it was written
for — a worker whose prompt is rejected before auto-compact ever fires. The
window belongs on every profile, not on whichever one happened to hit it.
2026-08-10 19:55:17 +02:00
kevin 379e03f9d0 CB-308: fold in the adversarial review; bump wiki to 4320c1c
CI / build (push) Successful in 59s
CI / contract (push) Successful in 1m16s
Four new resolved decisions (§7.7-7.10): turn state split by where the signals
are, with ABANDONED explicitly belt-and-braces over the waiter timeout; the
dual ack model with spawn idempotence by construction (gid stored IN the herdr
pane — the load-bearing detail of the no-ledger position — plus an in-flight
reservation for redelivery during a slow spawn); enforced publish semantics
(confirms + mandatory on a separate channel, return-before-confirm caveat);
queue lifecycle = session lifecycle with .v2 names for the redeclare hazard.

§8 reworked: global id scheme resolved and moved up; control authorization
sharpened into the hard gate on U4 (per-host allowlist beside the peer keys);
key distribution/rotation added. New §9: implementation order, each step
verifiable single-host, U4 gated, U8 last.

Wiki pointer bumped to 4320c1c (chapter 10 same-pass changes).

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_013ZGgxLQ2VpwZhEYoru8rkf
2026-08-10 22:32:00 +07:00
kevin e13921aa8a CB-308: resolve the cross-host design review; bump wiki to 36bb865
CI / build (push) Successful in 54s
CI / contract (push) Successful in 1m5s
Six decisions recorded in §7, replacing the matching open questions: signed
messages (identity from the key, extending identity-from-connection across the
broker), target-host-owned profiles advertised via presence, repo provisioning
by pinned forge clone, live-only asks with TTL + TOO_LATE notice, spawn-id
dedup on the target, and broker-outage semantics (local unaffected, remote
fails fast, gateway stays soft-state). §8 keeps what is genuinely still open,
with control *authorization* now separated from the resolved *authenticity*.

The broker-level half (U8 broadcast, exclusive consumers, inbox caps, TLS,
schema version, trace id) lands in wiki chapter 10 §10 — pointer bumped
(also picks up 1710a77, chapter 11 Features).

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_013ZGgxLQ2VpwZhEYoru8rkf
2026-08-10 21:54:30 +07:00
kevin 8a549d8610 Release 1.0.0 — one leader, one host, complete
CI / contract (push) Successful in 1m13s
CI / build (push) Successful in 1m36s
Bump bridged to 1.0.0 and add the release notes: the single-leader,
single-host scope is closed — gateway, lifecycle, two-way delivery,
pluggable peers, auth/authz/audit, supervision, CI. Cross-host
federation (CB-308) is the next major line.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_013ZGgxLQ2VpwZhEYoru8rkf
2026-08-10 20:58:06 +07:00
Dai Ha 84081b2bd8 CB-526: make a shipped capability undocumentable-by-accident
CI / contract (push) Successful in 49s
CI / build (push) Successful in 1m23s
CLAUDE.md already carries a mandatory before-done checklist ("the prompt is part
of the product"). It covered the instruction surface but not the operator-facing
one, and the result was measurable: CB-506 through CB-525 shipped without a
single wiki mention, while the Roadmap went on claiming Stage 5 was finished.

Discipline is what already failed, so this rides the existing gate rather than
adding a new habit to remember: one more row, firing when a change touches
anything an operator can use, configure, or observe. The row names where the
other two kinds of change go too (contracts to Implementation, coverage to the
Roadmap), so "nothing to document" is a decision the table makes rather than a
default you fall into.

Addendum-only — the canonical block is untouched and still byte-identical to the
wiki template (verified).
2026-08-09 19:38:34 +02:00
Dai Ha defe3365c4 CB-519: re-point pane assertions at the protocol-19 coordinate
CI / build (push) Successful in 1m22s
CI / contract (push) Successful in 2m32s
Rebase integration only, no behaviour change. CB-519's tests named the pane the
pre-protocol-19 fake produced (w9:pW_n); upstream's herdr 0.8.0 port creates the
pane through tab.create and starts the agent into it, so the fake now reports
w9:pRoot_n. Five assertions were therefore counting closes of a pane that never
existed and reading 0.

mvn clean install: Tests run: 399, Failures: 0, Errors: 0 — BUILD SUCCESS
2026-08-09 05:59:14 +02:00
Dai Ha 4e6201ecd1 CB-525: isolate a worker's tool surface to what its launcher mounts
A worker in a provisioned worktree was inheriting the primary's MCP servers by
two independent routes: the repo commits a .mcp.json declaring the IDE servers,
so a fresh checkout mounts them, and the default parity overlay then copied the
primary's own copy over the top.

Those servers are bound to the primary's IntelliJ project, so every path they
hand back points into the primary's checkout. A CB-523 worker made all 59 of its
edits there while running `mvn -f bridged/pom.xml` against its worktree — every
build it ran was of code that did not contain its changes, and it passed. The
worker's own `ls` of the file it had "edited" returned "No such file".

GitWorktrees now neutralizes .mcp.json at provisioning: an explicitly empty
server map, --skip-worktree'd when tracked so it never reads as pending work a
worker might commit. Unconditional, because the overlay was only half the leak.
The bridge itself is unaffected — it reaches a worker through the launcher's
--mcp-config flag, not the project file, so bridge_reply still works.

- BridgedConfig: .mcp.json out of the default parity overlay
- GitWorktrees: isolateToolSurface() on add(), with the rationale in javadoc
- GitWorktreesTest: 4 real-git acceptance tests (2 fail if the call is removed)
- implementer skill: work from $PWD, and quote a green unpiped `mvn clean
  install` from the worktree as the acceptance criterion

mvn clean install: Tests run: 392, Failures: 0, Errors: 0 — BUILD SUCCESS
2026-08-09 05:57:30 +02:00
Dai Ha ccf50f950e CB-524: make worker placement order deterministic across JVM runs
The weighted policy breaks an exact-weight tie on candidate list order
(WeightedRoundRobinPolicy picks the first candidate with a strictly greater
score), and that list comes from CompositePeerLauncher.candidates(), which
iterates profileConfigs. Both that map and BridgedConfig.workerProfiles() were
built with Map.copyOf, whose iteration order is salted per JVM run — so the
"in definition order" contract candidates() documents was not held.

Two consequences. In production, a config with equal weights (ollama 0.5 /
gx10 0.5) placed its first worker on a profile chosen at random on every daemon
restart. In the suite, CompositePeerLauncherTest.failoverRetriesNextCandidate-
WhenProfileIsUnreachable failed roughly one run in four, because whether
profile "a" was tried first depended on the salt.

Preserve definition order at every layer: unmodifiable LinkedHashMap for
workerProfiles(), profileConfigs, and byProfile (which also feeds the
user-visible bridge_profiles listing). The tests build profile maps with an
ordered helper rather than Map.of, which is salted for the same reason.

Guarded by a pair of tests declaring the same two profiles in opposite order
and asserting opposite first attempts, so any order-scrambling implementation
must fail one of them. Verified by mutation: reverting profileConfigs to
Map.copyOf fails 8/8 runs (6 caught by the original test, 2 only by the new
reversed-order one); with the fix, 10/10 fresh JVMs pass, 388 tests green.
2026-08-09 05:57:30 +02:00
Dai Ha a7f0211e2f CB-518: weighted placement policy for worker spawns 2026-08-09 05:57:30 +02:00
Dai Ha 049ce4828c CB-520: split ReplyInbox into explicit own/release and publish halves 2026-08-09 05:57:12 +02:00
Dai Ha c1173346ef CB-521: make the AMQP contract test runnable locally and in CI 2026-08-09 05:56:45 +02:00
Dai Ha 7ace184fe6 CB-519: make PeerHandle.id() a host-unique opaque UUID, decoupled from the herdr pane id 2026-08-09 05:56:45 +02:00
Dai Ha 54b314ace5 CB-522: let the primary run inside a herdr pane
Caller identity resolved any loopback PID that mapped to a herdr pane as a
WORKER, and PaneLocator scans every pane -- not just bridged-spawned ones. A
primary running inside a herdr pane therefore classified itself as a worker and
was refused SPAWN/SEND/STOP, i.e. every orchestration verb it exists to call.

The failure is self-locking: PrimaryRegistry only learns the primary's terminal
from bridge_send/bridge_spawn, the exact calls being refused, so the learned
value can never bootstrap. Only an operator-set pin breaks the cycle.

CallerResolver now consults primary.terminal from config *before* the pane
lookup. Deliberately the pinned value only, never the learned one -- the learned
terminal is populated by the callers this method is itself classifying, so
trusting it would be circular. Config is operator input, never network input, so
this widens no attack surface; bridge_whoami and the authz gate still share one
resolution.

Fixing that exposed a second, older bug. BridgeMcp's context extractor forwards
the caller's terminal into markPresent on every MCP call, documented as "no-op
for the primary (null terminal)". WorkerPresence.markPresent honours that, but
PresenceBridge overrides it and forwards the same null into SessionManager.
onReady -> transitionByTerminal -> findByTerminal, which called
terminalId.equals(...) unguarded. It only reached the scan once the registry was
non-empty, so the primary's first spawn succeeded and every later call NPE'd
with an HTTP 500 -- and it would have fired for ANY primary not living in a
herdr pane, pinned or not.

findByTerminal is now total. That covers onReady, onDelivered, onTurnComplete
and onTurnFailed at once; a null id could never match a registered session
anyway, so "no match" is the honest answer rather than taking down an unrelated
tool call.

Also drops two dead pass-throughs on CallerResolver (cwdForPid, tokenMode) that
IDE inspections flagged -- callers use ConnectionIdentity and BridgedConfig.Auth
directly.

The example config now states that primary.terminal is REQUIRED, not just a
push-loop optimisation, when the primary shares a herdr pane.

mvn clean install: 360 tests, 0 failures. Verified live: daemon restarted on
this jar, bridge_whoami reports primary, and four concurrent worktree spawns --
the exact shape that NPE'd -- now all succeed.
2026-08-09 05:55:12 +02:00
kevin 224b344445 CB-522: resolve the pinned primary.terminal pane as the primary, not a worker
CI / build (push) Successful in 2m54s
A primary running INSIDE a herdr pane was resolved as a worker by the
pane-match rule and refused every orchestration tool — the exact lockout
bridge_whoami surfaced on this deployment. The CB-307 primary.terminal pin
always claimed to replace connection-derived identity but only fed the push
loop; it now short-circuits CallerResolver ahead of the pane→worker rule
(the pane mapping is as unforgeable as a worker's, so no credential needed,
even in token mode). bridged.example.yaml documents the block.

Also guard the presence bridge against the primary's null terminal: the MCP
context extractor marks presence on every request, and the first genuine
primary contact NPEd into the SPAWNING→READY transition.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01SUTLvxtRPr2iT5u5g45BEs
2026-08-08 21:52:38 +07:00
kevin 0b28b4cb0f CB-521: port the herdr adapter to protocol 19 (herdr 0.8.0)
herdr 0.8.0 redesigned the agent API out from under the daemon: agent.start
now launches a supported kind INTO an existing pane, env/cwd move to pane
creation (tab.create / pane.split — the subscription-boundary seam now),
agent.send is replaced by agent.prompt (self-submitting) plus agent.send_keys
for the Enter nudge, and terminal ids are no longer valid agent.* targets.

- AgentControl: start(name, kind, args, paneId); prompt/send_keys delivery;
  cached terminal→pane target translation (invalidated on agent_not_found).
- WorkspaceControl: tab.create carries cwd+env; pane.split for legacy placement.
- HerdrPeerLauncher: the seed pane IS the worker pane (no drop step); retry
  agent.start while the seed shell boots (agent_pane_busy).
- FakeHerdr and the test suite model protocol 19 (unique seed panes, required
  kind/pane_id, prompt-based delivery); contract tests probe the seed shell
  instead of arbitrary-command agents, which protocol 19 removed.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01SUTLvxtRPr2iT5u5g45BEs
2026-08-08 21:52:25 +07:00
kevin 83129e165c Merge CB-518: state the primary's orchestration as an explicit, ordered flow
CI / build (push) Successful in 1m19s
Turns the primary's half of the bridge charter from a bullet list of
policies into a numbered 0-8 procedure, and splits delegated review out
of the merge step it used to sit beside. Wiki template kept byte-identical
by splicing; pointer bumped to 0c896eb in the same commit.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_017Kw1FosEt3Noix5GG9wJ2r
2026-08-04 22:37:17 +07:00
112 changed files with 13426 additions and 1865 deletions
+18
View File
@@ -0,0 +1,18 @@
{
"name": "claude-bridge",
"description": "Tooling for orchestrating a fleet of delegated coding agents through the bridged MCP gateway.",
"owner": {
"name": "LTMS"
},
"plugins": [
{
"name": "claude-bridge",
"source": "./plugin",
"description": "Make a project bridge-ready: mount the bridged MCP gateway and apply standard Claude Code settings so a session can orchestrate delegated workers. Ships no credentials.",
"version": "0.1.0",
"author": {
"name": "LTMS"
}
}
]
}
+37 -9
View File
@@ -9,11 +9,13 @@ The turn contract (one `bridge_reply`, `bridge_ask` for the lead's decisions, ho
never merge, never commit `.mcp.json` or `wiki/`) is in **`CLAUDE.md` → Bridge communication →
Worker** and already applies. This skill is only the *implement-and-hand-off procedure*.
You run in an **isolated git worktree on your own branch** — a full peer of the primary (same
repo, `CLAUDE.md`, skills, MCP), differing in the model behind you and the branch you sit on.
You run in an **isolated git worktree on your own branch** — a full peer of the primary (same repo,
`CLAUDE.md`, skills), differing in the model behind you and the branch you sit on. Your MCP surface
is **only what your launcher mounted** (the bridge): the primary's IDE and forge servers are not
yours, and the worktree's `.mcp.json` is deliberately emptied so you cannot inherit them.
The worktree model is documented in [`docs/Worker-Git-Workflow.md`](../../../docs/Worker-Git-Workflow.md).
## 1. Confirm where you are
## 1. Confirm where you are — then never leave
Before touching anything:
@@ -26,13 +28,36 @@ git status # should be clean at the start
Do **all** work here, on this branch. Never `git checkout main`, never rebase onto or push to
`main`. The branch is your isolation — respect it.
**Every path you read, edit, or build is relative to that root.** Work from `$PWD`; if a tool, a
brief, or your own memory hands you an absolute path, check it starts with your worktree root
before you touch it, and stop if it doesn't. An absolute path pointing anywhere else is the
primary's checkout — editing there while building here means **every build you run is of code that
does not contain your changes**, and it passes while your work goes nowhere. This has happened:
a worker made all 59 of its edits in the primary's tree and never noticed.
```bash
test "$(git rev-parse --show-toplevel)" = "$PWD" || cd "$(git rev-parse --show-toplevel)"
```
## 2. Implement
- Implement exactly the scope the lead named. Keep the diff focused; note anything out of scope
in your reply instead of widening it.
- Match the surrounding code's style, naming, and idioms.
- Run whatever build/test you can — `mvn clean install` from the module root. Read its **full**
output; a piped `mvn ... | tail` hides failures.
**Acceptance criterion — a green build, quoted.** Your work is not done until this passes *inside
your worktree*:
```bash
cd "$(git rev-parse --show-toplevel)/bridged" && mvn clean install
echo "exit=$?"
```
Read its **full** output — never pipe it through `tail`/`head`/`grep`, which hide a failure behind
a zero exit. Then quote the real `Tests run: … Failures: … Errors: …` line and the
`BUILD SUCCESS`/`FAILURE` verbatim in your reply. If it does not go green, say so with the actual
error; a failing build honestly reported is a usable result, a claimed-green one is not. You have
no IDE MCP tools, so `mvn` is your only verification — never claim a check you had no way to run.
## 3. Commit
@@ -41,7 +66,8 @@ git add <the files you changed> # explicitly — never `git add -A` / `git a
git commit -m "<ticket>: <clear one-line summary>"
```
`.mcp.json` will show as modified. Leave it — it is `--skip-worktree` and not yours to commit.
`.mcp.json` is neutralized and `--skip-worktree` in your worktree — never `git add` it, and never
"restore" it from the primary's copy. Same for `wiki/` (a submodule with its own remote).
## 4. Push
@@ -84,8 +110,9 @@ The reply is the entire handoff; the lead cannot see your terminal.
```
PR: <html_url from step 5, or "not created: <reason>" + branch name>
branch: <your branch>
files: <the files you changed>
tests: <what you ran and its REAL result — or "not run: <why>">
root: <git rev-parse --show-toplevel — proves you worked in your own worktree>
files: <worktree-relative paths you changed>
build: <the verbatim "Tests run: …" and BUILD SUCCESS/FAILURE lines — or "not run: <why>">
summary: <2-3 lines: what you implemented and any caveat the reviewer needs>
```
@@ -97,7 +124,8 @@ sequenceDiagram
participant G as git / gitea
L->>I: delegated task (you are in a worktree on your branch)
I->>I: implement + build/test here
I->>I: "implement here — every path under $PWD"
I->>I: "mvn clean install in this worktree, unpiped, until green"
I->>G: git commit (never .mcp.json / wiki)
I->>G: git push -u origin HEAD
I->>G: POST /pulls (GITEA_TOKEN) — open PR to main
+180
View File
@@ -0,0 +1,180 @@
---
name: port-to-opencode
description: Make an OpenCode session a first-class participant in a Claude Code workspace — instructions, MCP servers, and secrets — without duplicating config. Use when onboarding opencode to a project that already has CLAUDE.md and .mcp.json, or when an opencode peer needs the same tools and rules as the Claude session.
---
# Porting a Claude Code workspace to OpenCode
**The headline: there is almost nothing to port.** OpenCode reads `CLAUDE.md` natively. The only
artifact you create is one `opencode.json` mapping MCP servers. Do not translate instructions, do
not generate a second rules file, and do not install a sync tool — every one of those makes the
workspace worse.
Everything below was verified against `opencode 1.18.16` and the OpenCode docs.
## 1. Know what you get for free
OpenCode's instruction search order:
```
1. walking up from cwd: AGENTS.md , then CLAUDE.md
2. global: ~/.config/opencode/AGENTS.md
3. Claude Code global: ~/.claude/CLAUDE.md (unless disabled)
```
*"The first matching file wins in each category."*
Consequences that decide the whole procedure:
- **A project `CLAUDE.md` is already read.** No port needed.
- **Your user-level `~/.claude/CLAUDE.md` is already read too.** Global preferences carry over.
- **An `AGENTS.md` in the repo SHADOWS `CLAUDE.md`.** If one exists from a previous Codex port,
**delete it** — otherwise opencode reads the stale translated copy instead of the real rules.
This is the single most likely way to get this wrong.
## 2. Create `opencode.json` for MCP servers only
Project config lives at `opencode.json` in the repo root; the global one is
`~/.config/opencode/opencode.json`. **Configs merge, they do not replace** — so machine-local
servers belong in the global file and shared ones in the project file.
Map each entry from `.mcp.json`:
| `.mcp.json` | `opencode.json` |
|---|---|
| `"type": "http"` / `"sse"` | `"type": "remote"`, `"url"` |
| `"type": "stdio"` | `"type": "local"`, `"command": ["bin", "arg"]` |
| `"command"` + `"args"` | single `"command"` array |
| `"env"` | `"environment"` |
| `"headers"` | `"headers"` |
```json
{
"$schema": "https://opencode.ai/config.json",
"instructions": ["CLAUDE.md"],
"mcp": {
"bridged": { "type": "remote", "url": "http://127.0.0.1:8765/mcp", "enabled": true },
"context7": { "type": "remote", "url": "https://example.dev/mcp", "enabled": true,
"headers": { "Authorization": "Bearer {env:CONTEXT7_TOKEN}" } },
"gitea": { "type": "local", "command": ["gitea-mcp", "-t", "stdio"], "enabled": true,
"environment": { "GITEA_ACCESS_TOKEN": "{env:GITEA_ACCESS_TOKEN}" } }
}
}
```
Set `instructions` explicitly even though `CLAUDE.md` is found anyway — the fallback only applies
while no `AGENTS.md` exists, and being explicit survives someone adding one later.
## 3. Reference secrets, never embed them
OpenCode substitutes at load time, in both `headers` and `environment`:
```
{env:VARIABLE_NAME} value from the environment
{file:~/.secrets/token} value read from a file
```
**No credential ever belongs in `opencode.json`.** With `{env:…}` there is no reason to write one,
which is what makes this file safe to commit — and it must be committable, because a peer running
in a git worktree receives tracked files only.
**But `{env:…}` reads OpenCode's *process* environment — and OpenCode has no env store of its
own.** A Claude Code `env` block in `~/.claude/settings.json` or `.claude/settings.local.json` does
**not** reach it: those files are Claude Code's, and the variables exist only inside processes
Claude Code spawned. Verify with a clean login shell, not the shell your agent hands you:
```bash
env -u MY_TOKEN zsh -lc 'echo "${MY_TOKEN:-NOT IN PROFILE}"'
```
So `{env:…}` only works if something puts the variable in the environment first. Pick the supply
route by who launches opencode:
| Launcher | Route |
|---|---|
| a human, from a terminal | one central store, sourced by the login shell |
| a spawner (bridge, CI, IDE) | `{env:…}`, with the spawner injecting the variable |
**For the human case, keep every credential in one file the login shell sources.** Here that file
is `${SHARED_ENV}/tools/secrets.sh`, sourced from `${SHARED_ENV}/.ltms`, kept at mode 600 and never
committed. `opencode.json` then names variables and holds no values:
```json
"headers": { "Authorization": "Bearer {env:CONTEXT7_TOKEN}" }
```
Both routes end at the same syntax, and that is the point. The file does not change when a human
launches opencode instead of the bridge.
**`{file:…}` also works, and this project moved away from it.** Opencode resolves a relative
`{file:}` path against the project root, so a gitignored `.secrets/` beside `opencode.json` needs no
shell setup at all. It did not fail; the problem is that it makes a second copy of the token. The
same secret then lives in two places, and the copy you forget is the one that leaks or goes stale.
One store with many references is easier to rotate and to audit.
**One catch survives either choice, so state it out loud:** what a human's shell exports does not
reach a spawned peer, and neither does a gitignored `.secrets/` — a git worktree receives tracked
files only. Spawned peers must be fed through `{env:…}` by whatever launches them. Check the
variable *names* match: a spawner often injects under a different name than your shell uses, and
the config has no fallback. In this repo the bridge goes further: it neutralizes a worktree's
`opencode.json`, so a member cannot inherit the primary's credentials by accident.
## 4. Do not port machine-local MCP servers
IDE indexes, language servers, editor bridges — anything bound to *your* checkout — stay out of the
project file. Put them in `~/.config/opencode/opencode.json` if you want them personally.
A committed project config reaches every worktree. A worker that mounts servers whose paths point
into the primary's checkout will edit the primary's files while building in its own — every build
passes, every change lands in the wrong tree.
Port the servers the work needs. Leave the rest.
## 5. Verify against the running agent, not the file
A file on disk proves nothing about what the agent loaded.
```bash
opencode run "In one line: state a rule from this project's instructions."
```
The answer must reflect the actual `CLAUDE.md`. If it answers generically, the instructions did not
reach the model and everything after this is built on sand.
Then confirm the tools are mounted:
```bash
opencode mcp list
```
**"connected" does not mean "working".** A stdio server with a missing credential still completes
the MCP handshake and reports green; only a real tool call reveals it. Verified: `gitea` showed
`✓ connected` with no token, then failed the first call with `token is required`. A remote server
is more honest (`⚠ needs authentication`), but do not rely on that difference — **exercise one
authenticated tool per server**:
```bash
opencode run "Call <server>'s <tool>. Report the result or the exact error. One line."
```
Check too that no server you deliberately withheld is present.
## 6. Report
State what you changed, which servers crossed and which you withheld and why, and quote the
verification answer verbatim. If any server failed to connect, say so plainly — a partially mounted
peer is worse than a missing one, because it looks configured.
## Gotchas
- **A leftover `AGENTS.md` silently wins over `CLAUDE.md`.** Check for one before anything else.
- **`opencode.json` is merged, not overridden** — a global entry and a project entry with the same
server name both matter; keep names distinct unless you intend to layer them.
- **`enabled: false`** turns a server off without deleting its config — prefer it over removal when
you may want the server back.
- **OAuth-based servers** store tokens in `~/.local/share/opencode/mcp-auth.json` after
`opencode mcp auth <server>`; that is machine state, never config to commit.
- **Skills and subagents do not port.** OpenCode uses its own agent markdown under
`.opencode/agents/`; `.claude/skills/**` is not read. If a delegation brief tells a peer to load a
skill by name, that instruction has no effect on an opencode peer — spell the procedure out in the
brief, or author the equivalent agent file.
+52
View File
@@ -53,3 +53,55 @@ jobs:
grep -qE "Failures: [1-9]|Errors: [1-9]" "$f" && { echo "===== $f ====="; cat "$f"; }
done
exit 0
# CB-521 — actually run the AMQP contract test in CI, against a REAL broker. The broker is a
# RabbitMQ SERVICE CONTAINER, not Testcontainers-with-Docker: the runner image has no Docker, so
# AmqpReplyInboxContractTest reads AMQP_URI (set below to the service's network alias) and binds
# straight to it — no Docker, no skipped tests. This separation (build job hermetic and
# Docker-free; contract job broker-provided) is deliberate — see the default-excludes/contract
# profiles in bridged/pom.xml. `setup-java` provides the JDK only; Maven is installed separately,
# exactly as in the build job above.
contract:
runs-on: ubuntu-latest
services:
rabbitmq:
image: rabbitmq:3.13 # same AMQP 0-9-1 engine the local Testcontainers fixture uses
env:
RABBITMQ_DEFAULT_USER: guest
RABBITMQ_DEFAULT_PASS: guest
ports:
- 5672:5672
env:
# Service containers are reachable from the job by their network alias on their internal port.
AMQP_URI: amqp://guest:guest@rabbitmq:5672
steps:
- uses: actions/checkout@v4
- name: Set up JDK 25
uses: actions/setup-java@v4
with:
distribution: temurin
java-version: '25'
cache: maven
- name: Install Maven
run: |
apt-get update && apt-get install -y --no-install-recommends maven
mvn -version
# The `contract` profile clears the default-excludes group, so the @Tag("contract") AMQP test
# runs against the RabbitMQ service container (AMQP_URI). Pinned to the one contract test to
# avoid re-running the unit suite already covered by the `build` job.
- name: Contract tests
working-directory: bridged
run: mvn -B -Pcontract test -Dtest=AmqpReplyInboxContractTest
- name: Failing test output
if: failure()
working-directory: bridged
run: |
for f in target/surefire-reports/*.txt; do
[ -f "$f" ] || continue
grep -qE "Failures: [1-9]|Errors: [1-9]" "$f" && { echo "===== $f ====="; cat "$f"; }
done
exit 0
+13
View File
@@ -0,0 +1,13 @@
# No secret belongs in this repo any more: every credential lives in one shell-level store
# (${SHARED_ENV}/tools/secrets.sh), and opencode.json reads it as {env:...}. This line stays as a
# backstop, so a workspace-scoped copy that someone re-creates by habit still cannot be committed.
.secrets/
# Settings backups inherit the env block — and secrets with it.
.claude/settings.local.json.bak*
# Daemon runtime artefacts. bridged appends its log wherever it is launched from, so both the
# repo root and bridged/ collect one; neither belongs in git.
bridged.out
bridged/bridged.out
logs/
+44 -3
View File
@@ -41,8 +41,9 @@ nothing. Fail toward the recoverable error.
2. **The bridge is the only channel.** Text you print in your terminal reaches nobody — the other
side cannot see your screen. An answer that isn't in a `bridge_*` call is silently discarded.
3. **Identity comes from the connection, never an argument.** Workers never pass a target; you
cannot act as another session. Spawn/stop/send/drain are primary-only; reply/ask are
worker-only-and-only-as-itself. A call outside your role is refused, not queued.
cannot act as another session. Spawn/stop/send/drain are lead-only; reply/ask are
only-as-itself — any peer may answer for its own pane, and for no other. A call outside your
role is refused, not queued.
4. **Delivery is status-gated: one message per turn.** Don't busy-poll a peer's terminal and don't
re-send because a call looks slow — the bridge delivers when the peer is `idle`/`blocked`.
5. **Never drive the terminal multiplexer directly** (no `herdr` CLI, no socket). The bridge owns
@@ -101,13 +102,42 @@ the merge — and merging on a reviewer's word is delegating it by proxy.
| Confirm your own role | `bridge_whoami` |
| See backends available | `bridge_profiles` |
| Start a worker | `bridge_spawn{profile?, cwd?, worktree?, ticket?}` → `sessionId` + `paneId` |
| See the fleet | `bridge_list` · one worker's state: `bridge_status{sessionId}` |
| See the fleet | `bridge_list` → `leads` (your peers) + `workers` · one peer's state: `bridge_status{sessionId}` |
| Delegate (blocking) | `bridge_send{sessionId, content}` |
| Delegate (long task) | `bridge_send{sessionId, content, wait:false}` → ticket → `bridge_poll{ticket}` |
| Answer a worker's `bridge_ask` | `bridge_send{turnId, content}` — **not** `sessionId` |
| Message a **peer lead** | `bridge_send{sessionId: <their terminal>, content}` — `bridge_list` → `leads` reports it. Coordination only, **never** a task |
| Answer a peer lead that messaged you | `bridge_reply{content}` — the one case a lead replies |
| Collect a held reply | `bridge_poll{target}` · then `bridge_ack{target, msgId}` |
| Tear down | `bridge_stop{paneId}` |
### Lead ↔ lead — coordinate, never delegate
`bridge_list` returns `leads` alongside `workers`; your own row carries `self: true`. Every other row
is a peer — an orchestrator with its own context, its own workers, and its own judgment. An empty
`workers` array means no workers are spawned; it says nothing about peers.
**A lead never assigns a task to another lead.** Work goes to workers — only ever downward, never
sideways. Sending a peer a brief with acceptance criteria is a category error: a brief is a worker's
artefact, and a peer is not yours to task. If a unit needs doing and it falls in your area, spawn a
worker and delegate it yourself; if it falls in the peer's area, say so and let the peer assign it.
The traffic between leads is coordination and nothing else:
1. **Divide the map, not the work.** Agree who owns which area, then each of you assigns inside your
own. Split by **context ownership** — whoever already holds the context owns that area — and say
who takes what, in one message, before either of you starts. Two leads silently working the same
unit is the failure mode here, and neither notices until the merge.
2. **Share findings, hazards, and corrections.** What you have already discovered, what broke, what
the next person will trip on. This is the traffic that actually pays for the channel: it costs one
message and saves a peer a rediscovery.
3. **Verify a peer exactly as you verify yourself.** Peer status buys nothing: check the claim
against the code, and re-run the build. A peer's correction gets the same treatment — right or
wrong on the evidence, not on who said it. Neither of you merges the other's work unreviewed.
Being messaged by a peer does not make you its worker: answer with `bridge_reply`, and push back on
the substance if it is wrong. A peer that simply complies has thrown away the reason there are two of
you.
### Worker — the turn contract
1. **Load the playbook skill the lead named** before doing anything else.
@@ -143,6 +173,8 @@ charter, not here.
`mcp/ConnectionIdentity` (connection→role), and `worker/*Launcher` (`REPLY_CHARTER`).
- **Skills available to delegate:** `implementer` (worktree → commit → push → own PR) and
`reviewer` (scoped review → one structured finding). Name one in every delegation.
- **Primary-side skills** (not delegation playbooks — a worker cannot use them):
`port-to-opencode` (make an OpenCode session a participant in this workspace).
- **Never commit** `.mcp.json` (the primary's local copy, flagged `--skip-worktree`) or `wiki/`
(a submodule with its own remote).
- **Flows and the error model** — rendezvous, `bridge_ask`, detached delivery, turn-done fallback —
@@ -168,6 +200,15 @@ Before you call any work done, check the row that matches what you touched:
| worktree provisioning or the parity overlay | the "both roles read this file" premise — it rests on the worker's worktree being a checkout of this repo |
| `.claude/skills/**` | the addendum's skill list, and the "name the playbook" rule |
| a new peer kind (non-Claude adapter) | what that peer can read — anything it must obey belongs in its charter, not in the block |
| **anything an operator can use, configure, or observe** — an MCP tool, a `bridged.yaml` knob, an endpoint, a visible behaviour | **[Features](wiki/11-Features.md)** — one entry: what it does · the knob that turns it on · **why it exists** · the gotcha |
That last row is not bookkeeping. Chapters 1–10 answer *how is this built* and *why this way*;
none of them has a home for *what can it do and how do I turn it on*, so for twenty tickets a
shipped capability landed nowhere and the Roadmap went on claiming the stage was finished. The
*why* line is the one that matters — without it a decision gets re-litigated from scratch a month
later. Internal contract changes go to `wiki/9-Implementation.md` instead; test and coverage work
is a Roadmap line. A change that touches none of the three earns no entry, and that is a normal
outcome rather than an omission.
Then **propagate**: the block in this file and the template in the wiki
([Use Cases](https://git.ltms.dev/lms/claude-bridge/wiki/7-Use-Cases) → *The portable `CLAUDE.md`
+202 -14
View File
@@ -31,15 +31,71 @@ bind:
# mode: token
# tokenEnv: BRIDGED_API_TOKEN
# Optional pinned primary terminal (CB-307). Names the herdr pane the PRIMARY itself runs in:
# a caller whose connection maps to this pane resolves as the primary (no credential needed —
# the pane mapping is as unforgeable as a worker's), and reply nudges are pushed to it.
# REQUIRED when the primary runs inside a herdr pane — without it the pane match reads the
# primary as a worker and refuses spawn/send/stop. Get the id from bridge_whoami; re-pin if
# the primary moves panes.
# primary:
# terminal: term_0123456789abcd
# pushReminders: 5 # max nudges before giving up (default 5)
# pushBackoffMs: 15000 # delay between nudges (default 15000)
# CB-530: MORE THAN ONE LEAD. `primary:` above is singular by construction — every other pane
# resolves as a worker — which is right for one lead driving a fleet and wrong the moment two leads
# (say a Claude lead and an opencode lead) work as peers: the second is silently demoted and refused
# every orchestration call. List each lead's pane here and all of them resolve as leads.
#
# terminal → the ONLY field identity depends on; get it from that session's bridge_whoami
# kind/model → descriptive; they document what runs in the pane and are echoed by bridge_whoami
#
# A lead is never spawned — it pre-exists, which is exactly why it must be named rather than created.
# `bridge_whoami` reports `{"role":"primary","leader":"<name>"}`; role stays "primary" because a lead
# IS a primary for authorization, so nothing that keys on the role breaks.
#
# KEEP `primary:` when adding leads: it still addresses the CB-307 push loop, which needs a single
# destination for its nudges. If both name the same terminal, the `fleet.leaders:` entry wins.
#
# Leads are configured under `fleet.leaders:` — see THE FLEET further down.
#
# Two things stop the tab-name convention from becoming a way to claim leadership: the configured
# member spaces are excluded from the scan, so nothing bridged places can land in a matching tab;
# and startup REFUSES a `tabPrefix` that the fleet tabLabel template, or any per-profile `tabLabel`
# override, also matches — so the two namespaces cannot overlap by accident. The label is a NAME,
# never a capability: what a pane may do is decided by the role the daemon resolves for it.
# CB-551: IDLE-LEAD HEARTBEAT — nudge the single lead back to work when it has been continuously
# idle (no open bridge_send driving it) past the quiet period. The fleet is one lead + architects +
# workers, so a lead that stalls is a single point of failure; the ReplyPushLoop only nudges when a
# reply lands, and this timer catches the gap where nothing lands and the lead just sits idle.
#
# Opt-in on purpose — it SPENDS the operator's subscription on its own initiative (each nudge starts
# a lead turn nobody asked for), so upgrading the daemon must never switch it on for you. Absent
# block = feature off, exactly as before.
#
# Three knobs, each with a default that errs on the side of not burning context:
# idleAfterSeconds: 300 # how long the lead must stay idle before the FIRST nudge (default 300 —
# # absorbs normal post-turn pauses; re-prompting every pause burns context)
# backoffMs: 60000 # re-check cadence / spacing between nudges past the quiet period (default 60000)
# quietNudgeCap: 3 # cap on consecutive nudges that find NOTHING pending, then it stops
# # until real state appears (default 3 — never nag an empty fleet forever)
# leadHeartbeat:
# idleAfterSeconds: 300
# backoffMs: 60000
# quietNudgeCap: 3
# herdr Unix socket. Omit to use the client default
# (${HERDR_SOCKET_PATH:-~/.config/herdr/herdr.sock}).
herdrSocket: ~/.config/herdr/herdr.sock
# How worker sessions are spawned. Define one or more named profiles (backends) under
# `workers`; each key is the profile name (also the ccs profile). `defaultWorker` picks
# which one a no-argument spawn uses (bridge_spawn with no profile / POST /workers).
# How member sessions are spawned. Define one or more named profiles (backends) under
# `profiles`; each key is the profile name (also the ccs profile). A profile says only WHICH
# BACKEND — model, CLI adapter, credentials, cost. It says nothing about what a member spawned on
# it is for; that is the member's role, and roles live under `fleet:` below. Which profile an
# unqualified spawn lands on comes from that role's pool, not from a global default.
#
# Shared knobs (placement/workspace/tabLabel) can be repeated per profile; they usually match.
# Shared knobs (placement/workspace) can be repeated per profile; they usually match.
# placement: tab → each worker lands in its OWN tab in a dedicated worker space (default).
# Use `pane` for the legacy behaviour (split the focused tab).
# mcpUrl → bridged mounts the bridge MCP (--mcp-config, inline) + reply charter
@@ -53,7 +109,16 @@ herdrSocket: ~/.config/herdr/herdr.sock
# skills/MCP/hooks. Omit to leave the worker on the host default.
# parityOverlay → repo-relative paths copied primary→worktree so a worker in a provisioned
# worktree sees the same local config (CB-301-ext). Omit for the default set:
# [.mcp.json, .claude/settings.local.json, .env, .envrc].
# [.claude/settings.local.json, .env, .envrc].
#
# Do NOT add .mcp.json (CB-525). A worker's tools are whatever its launcher
# mounts — the bridge, and nothing else. Replicating the primary's MCP config
# handed a worker the primary's IDE servers, which are bound to the primary's
# checkout, so its navigation returned paths OUTSIDE its own worktree: one
# worker made all 59 of its edits in the primary tree while compiling its
# worktree, and every build it ran was of code that did not contain them.
# bridged neutralizes a provisioned worktree's .mcp.json for this reason;
# listing it here would copy the primary's back over that.
# gitTokenEnv → host env var holding the git-forge API token. When set, its value is injected
# as GITEA_TOKEN so the worker can open its OWN PR at checkpoint (CB-302).
# Opt-in by design — omit and the worker gets no PR-create grant (push over
@@ -78,29 +143,39 @@ herdrSocket: ~/.config/herdr/herdr.sock
# entry cannot repoint a worker past the SubscriptionGuard — which is checked
# against `baseUrl` alone.
# Put `defaultMode: "auto"` in each ccs profile so the worker runs autonomously.
workers:
profiles:
gx10: # ccs profile name (NOT a hostname)
kind: claude-code # which adapter spawns this profile (default; may omit)
baseUrl: http://gx01.gw:8000 # the vLLM host this profile targets (gx00.gw / gx01.gw)
model: coder
placement: tab
workspace: bridged-workers
tabLabel: "worker: {profile} #{n}" # {profile}/{model}/{n} substituted; {n} keeps sibling tabs distinct
# tabLabel: an optional per-profile override; the fleet template usually covers it
mcpUrl: http://127.0.0.1:8765/mcp
tokenEnv: BRIDGED_WORKER_TOKEN
argv: ["ccs", "gx10"]
weight: 0.5 # relative selection weight for placement: weighted
maxLoad: 2 # max live workers on this profile (omit for unlimited)
# gitTokenEnv: GITEA_TOKEN # opt-in: let this profile's workers open their own PR (CB-302)
# gitHostEnv: GITEA_HOST # defaults to GITEA_HOST; injected only with gitTokenEnv
# configDir: /Users/me/.ccs/instances/gx10 # CLAUDE_CONFIG_DIR — inherit that profile's skills/MCP
# cwd: /Users/me/src/myrepo # pin the working dir; omit to inherit the primary's
# parityOverlay: [".mcp.json", ".claude/settings.local.json", ".env", ".envrc"]
ollama:
baseUrl: http://ollama.ltms.dev # local/self-hosted; usually no token
# parityOverlay: [".claude/settings.local.json", ".env", ".envrc"] # never add .mcp.json — see above
gx11: # a second backend, so `placement: weighted` has a choice
baseUrl: http://gx01.gw:8000 # self-hosted; ccs handles the model + token
placement: tab
workspace: bridged-workers
tabLabel: "worker: {profile} #{n}"
# tabLabel: an optional per-profile override; the fleet template usually covers it
mcpUrl: http://127.0.0.1:8765/mcp
argv: ["ccs", "ollama"]
argv: ["ccs", "gx11"]
weight: 0.5
maxLoad: 2
# Pin an auto-compact window BELOW the served model's context ceiling. The global
# ~/.claude/settings.json value is shared by every ccs instance and the primary, so the
# per-profile override belongs here. Equal to the ceiling means auto-compact never fires
# before the server rejects the prompt, which kills a worker mid-turn (CB-523).
env:
CLAUDE_CODE_AUTO_COMPACT_WINDOW: "280000"
# CB-402: a second coding-agent kind, proving the PeerLauncher SPI is provider-neutral.
# opencode is provider-agnostic and uses NONE of Claude's private seams: no ANTHROPIC_BASE_URL /
# SubscriptionGuard (so it needs no `guard` host entry), no --mcp-config / --append-system-prompt.
@@ -144,7 +219,112 @@ workers:
# tabLabel: "opencode: {profile} #{n}"
# mcpUrl: http://127.0.0.1:8765/mcp
# argv: ["opencode"]
defaultWorker: gx10
# How an unqualified spawn chooses a profile: fixed (default, reproduces pre-CB-518 behaviour),
# round-robin, or weighted. Omitting this key is a strict no-op for existing configs.
placement: weighted
# Re-read this file without restarting the daemon (CB-559). Off unless you add this block, so an
# upgraded bridged keeps the old behaviour: the file is read once at boot and never again.
# enabled → turn the watch on. bridged checks the file's modified time on a timer and
# reloads when it moves.
# intervalSeconds → how often to check (default 10). One `stat` per tick, so this is cheap.
#
# Not every key can move under a running daemon, and the difference is about what already exists
# when the reload happens — not about how important the key is:
# HOT → takes effect on the next spawn: the whole `fleet:` block (every role pool and
# `tabLabel`), `placement:`, and an existing profile's weight / maxLoad. Those are
# hot because the placement policy reads them through a supplier — being config is
# not by itself enough to make a key hot.
# DEFERRED → accepted into the new config, but the wiring built at startup keeps the old value
# until you restart: `lifecycle:`, `leadHeartbeat:`, `guard:`, `worktreeRoot:`,
# `spawnReadyTimeoutMs` / `spawnReadyPollMs`, ADDING or REMOVING a profile (a new
# backend needs its own launcher, and launchers are built once), AND an existing
# profile's launch settings — model, baseUrl, argv, env, configDir, mcpUrl, tabLabel.
# The launcher takes a copy of `profiles:` at startup and resolves every spawn out of
# that copy, so those never reach a launch until you restart. The reload logs them by
# name rather than pretending they applied.
# COLD → cannot change at all: `bind:`, `herdrSocket:`, `broker:` and `auth:`. The socket is
# bound, the broker connection is open, and the auth mode decides who may reach the
# port that is already listening.
#
# A changed COLD key refuses the WHOLE reload — not the hot half applied and the cold half warned
# about. A half-applied reload would leave the daemon matching no file on disk, which is the worst
# thing a reload can do to an operator debugging one. A file that fails to parse or fails a startup
# validator is refused the same way, and the running config stays live.
# configReload:
# enabled: true
# intervalSeconds: 10
# THE FLEET (CB-557) — who the daemon may run, and under which role. This one block replaced four
# older keys: `leaders:`, `members:`, `leadScan:` and `defaultProfile:`.
#
# A member is anything a lead spawns, and every member has two INDEPENDENT attributes:
# role — which contract: architect, dev or reviewer. It picks the launch charter, the role
# file, the playbook skill and the authz row.
# profile — which backend: one of the `profiles:` keys above (model, CLI adapter, cost).
# They vary on their own. A reviewer may run on the same profile as the dev whose diff it reads,
# which is why the two cannot be one field.
#
# The ROLE IS THE CONTAINING KEY, not a `role:` field. That is not only tidier: a misspelled role
# used to parse into a member with no contract at all, while a misspelled pool name here simply
# declares nothing.
#
# Each pool lists the profiles that role MAY run on — these are pools, not identities. That is also
# what replaced `defaultProfile:`: an unqualified spawn names a role, and that role's pool supplies
# the candidates, in definition order. A dev and a reviewer staying anonymous is exactly compatible
# with being listed here; the entry key just names the entry.
fleet:
# Optional. Template for a member tab's label; {role}, {profile}, {model} and {n} are substituted.
# {n} counts per role+profile, so `dev: sonnet #2` really is the second sonnet dev. Because {role}
# comes from a closed enum, a generated label can never begin with a lead's tabPrefix.
# tabLabel: "{role}: {profile} #{n}"
# Panes that orchestrate rather than are orchestrated. A lead may now be CREATED as well as
# recognised: give it a `profile:` and the daemon launches the shortfall when fewer than
# `instances` are live. Give it only a `terminal:` and it is recognise-only, as before.
#
# `tabPrefix` is the naming convention that finds a lead without pasting a terminal id: label the
# tab `lead: <name>` when you open it and the pane is recognised on the next rescan. Reopen the
# tab later and the id changes; the label does not.
#
# A lead the daemon launches is labelled BY the daemon, using the same convention, so it is found
# by the same scan. A lead counts as live only when herdr also reports a running agent in that
# tab — a label left behind by a session that died does not block the relaunch.
#
# An auto-launched lead is NOT a member: it gets no worker reply charter, is never registered with
# the session lifecycle (the idle reaper would kill your orchestrator), and stays on the
# subscription — ANTHROPIC_BASE_URL/AUTH_TOKEN are stripped from its env whatever the profile says.
# leaders:
# opus-5.0:
# profile: opus # omit to never create this lead, only recognise it
# instances: 1 # desired live count; only the shortfall is launched. 0 = off
# terminal: term_0123456789abcd # optional hand-pin; usually found by tabPrefix instead.
# # A running agent on this terminal also counts as live, so a
# # lead you opened by hand is not relaunched under you.
# tabPrefix: "lead:" # `lead: opus-5.0` ⇒ a lead named opus-5.0 (case-insensitive)
# scanIntervalSeconds: 10 # rescan cadence, and the worst case before a new tab is seen
# workspace: leads # where a launched lead's tab is created (default "leads").
# # MUST NOT be a member workspace — those are excluded from the
# # scan, so a lead placed in one is never found again.
# cwd: /path/to/repo # the launched lead's working directory (default: bridged's own)
# kind: claude # descriptive; reported by bridge_whoami
# gpt-sol-5.6:
# terminal: term_fedcba9876543
# kind: opencode
# model: openai/gpt-5.6-terra
# architects:
# architect-1:
# profile: opus # a strong model, on the operator's subscription
# architect-2:
# profile: sol # a different vendor on purpose — two architects that share a
# # model share its blind spots
developers:
gx10:
profile: gx10
# reviewers:
# gx10:
# profile: gx10 # the same backend may serve two roles; that is the point
# Subscription boundary. A worker's base_url host MUST be one of these; the primary
# must carry none. Every profile above must have its host listed here.
@@ -152,7 +332,6 @@ guard:
offSubscriptionHosts:
- gx00.gw
- gx01.gw
- ollama.ltms.dev
# Spawn-readiness gate (CB-306). The launcher blocks until the worker's herdr status is
# injectable (IDLE/BLOCKED/DONE) or the timeout elapses. 0 disables the gate.
@@ -195,6 +374,15 @@ guard:
# the first orchestration-side MCP call (the normal case). An off-host or
# non-herdr primary leaves this unresolved → the loop is a no-op and delivery
# degrades to pull; the reply is still never lost.
#
# REQUIRED (CB-522) if the primary itself runs inside a herdr pane. Caller
# identity resolves a loopback PID to its herdr pane, and PaneLocator scans
# EVERY pane — not just bridged-spawned ones — so such a primary is otherwise
# classified as a WORKER and refused SPAWN/SEND/STOP. That failure is
# self-locking: the learned terminal is populated by the very orchestration
# calls being refused, so only this pinned value can break the cycle. Read the
# id off bridge_whoami (it reports the current terminal even while
# misclassified) and re-pin whenever the primary moves panes.
# pushReminders → max nudges before giving up (default 5)
# pushBackoffMs → delay between nudges in ms (default 15000)
# primary:
+21 -1
View File
@@ -6,7 +6,7 @@
<groupId>dev.ltms</groupId>
<artifactId>bridged</artifactId>
<version>0.1.0-SNAPSHOT</version>
<version>1.0.0</version>
<packaging>jar</packaging>
<name>bridged</name>
@@ -236,6 +236,26 @@
<profile>
<id>contract</id>
<properties><excludedGroups/></properties>
<build>
<plugins>
<plugin>
<artifactId>maven-surefire-plugin</artifactId>
<configuration>
<!-- Docker-engine compat (see "Running the contract tests" in
docs/CB-307-Reliable-Delivery.md): Testcontainers 1.20.4's docker-java
client defaults to Docker API 1.32 when no version is set, but modern
engines (OrbStack on this dev host, min 1.40) reject that as too old —
which surfaces as "Could not find a valid Docker environment". Pinning
api.version=1.43 works on OrbStack and Docker 24+, and is overridable
per-host via -Dapi.version. Only active under -Pcontract, so the
default hermetic build never sets it. -->
<systemPropertyVariables>
<api.version>1.43</api.version>
</systemPropertyVariables>
</configuration>
</plugin>
</plugins>
</build>
</profile>
</profiles>
</project>
@@ -1,10 +1,14 @@
package dev.ltms.bridged;
import dev.ltms.bridged.config.BridgedConfig;
import dev.ltms.bridged.config.ConfigRef;
import dev.ltms.bridged.config.ConfigWatcher;
import dev.ltms.bridged.guard.SubscriptionGuard;
import dev.ltms.bridged.herdr.AgentControl;
import dev.ltms.bridged.herdr.HerdrClient;
import dev.ltms.bridged.herdr.HerdrException;
import dev.ltms.bridged.herdr.LeadTabScanner;
import dev.ltms.bridged.lead.LeadLauncher;
import dev.ltms.bridged.herdr.PaneLocator;
import dev.ltms.bridged.herdr.UnixSocketHerdrClient;
import dev.ltms.bridged.herdr.WorkspaceControl;
@@ -12,7 +16,8 @@ import dev.ltms.bridged.inject.CompletionResolver;
import dev.ltms.bridged.inject.Injector;
import dev.ltms.bridged.inject.StatusPoller;
import dev.ltms.bridged.inject.TurnListener;
import dev.ltms.bridged.inject.WorkerPresence;
import dev.ltms.bridged.inject.MemberPresence;
import dev.ltms.bridged.auth.MemberRegistry;
import dev.ltms.bridged.auth.CallerResolver;
import dev.ltms.bridged.mcp.BridgeMcp;
import dev.ltms.bridged.mcp.ConnectionIdentity;
@@ -26,16 +31,17 @@ import dev.ltms.bridged.msg.InMemoryReplyInbox;
import dev.ltms.bridged.msg.MessageService;
import dev.ltms.bridged.msg.Rendezvous;
import dev.ltms.bridged.msg.ReplyInbox;
import dev.ltms.bridged.msg.LeadHeartbeatLoop;
import dev.ltms.bridged.msg.ReplyPushLoop;
import dev.ltms.bridged.rest.BridgedApp;
import dev.ltms.bridged.session.GitWorktrees;
import dev.ltms.bridged.session.SessionManager;
import dev.ltms.bridged.peer.PeerLauncher;
import dev.ltms.bridged.session.SessionReaper;
import dev.ltms.bridged.worker.ClaudeCodeLauncher;
import dev.ltms.bridged.worker.CompositePeerLauncher;
import dev.ltms.bridged.worker.HerdrPeerLauncher;
import dev.ltms.bridged.worker.OpenCodeLauncher;
import dev.ltms.bridged.member.ClaudeCodeLauncher;
import dev.ltms.bridged.member.CompositePeerLauncher;
import dev.ltms.bridged.member.HerdrPeerLauncher;
import dev.ltms.bridged.member.OpenCodeLauncher;
import io.javalin.Javalin;
import org.slf4j.Logger;
import org.slf4j.LoggerFactory;
@@ -45,7 +51,15 @@ import java.util.ArrayList;
import java.util.LinkedHashMap;
import java.util.List;
import java.util.Map;
import java.util.Objects;
import java.util.Set;
import java.util.concurrent.Executors;
import java.util.concurrent.TimeUnit;
import java.util.concurrent.atomic.AtomicReference;
import java.util.function.Function;
import java.util.function.Predicate;
import java.util.function.Supplier;
import java.util.stream.Collectors;
/**
* {@code bridged} entry point. Wires the real herdr socket client to the REST app and
@@ -56,9 +70,6 @@ public final class Bridged {
private static final Logger log = LoggerFactory.getLogger(Bridged.class);
/** How often the injector samples a busy worker's status while it has queued work. */
private static final long INJECT_POLL_MILLIS = 250;
/** CB-504: how long to wait at startup for herdr's socket before serving degraded. */
private static final long HERDR_WAIT_SECONDS = 30;
private static final long HERDR_WAIT_POLL_MILLIS = 500;
@@ -66,6 +77,11 @@ public final class Bridged {
static void main(String[] args) {
Path configPath = Path.of(args.length > 0 ? args[0] : "bridged.yaml");
BridgedConfig cfg = BridgedConfig.load(configPath);
// CB-559: `cfg` stays the startup snapshot — every validation and every piece of one-time
// wiring below reads it, and must, because those decisions cannot be unmade. `config` is the
// live reference the hot paths read per use. Which keys can actually move is ConfigRef's
// contract; adding a reader here does not make a key reloadable by itself.
ConfigRef config = new ConfigRef(configPath, cfg);
// The primary/host env that launched bridged must not be tainted.
SubscriptionGuard guard = new SubscriptionGuard(cfg.guard().hostSet());
@@ -76,6 +92,13 @@ public final class Bridged {
// refuses remote connections to a loopback socket. This throws rather than warns so the
// dangerous configuration cannot be reached by ignoring a log line.
cfg.validateAuthExposure();
cfg.validateLeadTabPrefixes();
// CB-542: a subscription:true profile whose env: reseats ANTHROPIC_BASE_URL/AUTH_TOKEN would
// reach an unguarded endpoint (the launcher skips SubscriptionGuard for it). Refuse at load.
cfg.validateSubscriptionProfiles();
// CB-548: every architect slot must name a configured workers: profile — the strong-model
// backend the future spawn lifecycle would read. A stale reference dies here, not later.
cfg.validateMembers();
Path socket = cfg.herdrSocket() != null && !cfg.herdrSocket().isBlank()
? Path.of(cfg.herdrSocket())
@@ -88,9 +111,9 @@ public final class Bridged {
// CB-402: one adapter per configured peer kind, fronted by a composite router. A profile's
// `kind:` selects its adapter — claude-code (the default) and opencode partition the profile
// set — and the composite dispatches each SPI call to the adapter that owns the profile/pane.
Map<String, BridgedConfig.Worker> claudeProfiles = new LinkedHashMap<>();
Map<String, BridgedConfig.Worker> opencodeProfiles = new LinkedHashMap<>();
cfg.workerProfiles().forEach((name, w) -> {
Map<String, BridgedConfig.Profile> claudeProfiles = new LinkedHashMap<>();
Map<String, BridgedConfig.Profile> opencodeProfiles = new LinkedHashMap<>();
cfg.profiles().forEach((name, w) -> {
if (w.isOpenCode()) {
opencodeProfiles.put(name, w);
} else {
@@ -103,21 +126,29 @@ public final class Bridged {
// unless opencode is the only kind configured.
if (!claudeProfiles.isEmpty() || opencodeProfiles.isEmpty()) {
adapters.add(new ClaudeCodeLauncher(agents, spaces, guard,
claudeProfiles, cfg.defaultProfile(), System::getenv,
cfg.spawnReadyTimeoutMs(), cfg.spawnReadyPollMs()));
claudeProfiles, cfg.effectiveDefaultProfile(), System::getenv,
cfg.spawnReadyTimeoutMs(), cfg.spawnReadyPollMs(),
() -> config.get().fleet().tabLabel()));
}
if (!opencodeProfiles.isEmpty()) {
adapters.add(new OpenCodeLauncher(agents, spaces,
opencodeProfiles, cfg.defaultProfile(), System::getenv,
cfg.spawnReadyTimeoutMs(), cfg.spawnReadyPollMs()));
opencodeProfiles, cfg.effectiveDefaultProfile(), System::getenv,
cfg.spawnReadyTimeoutMs(), cfg.spawnReadyPollMs(),
() -> config.get().fleet().tabLabel()));
}
PeerLauncher workers = new CompositePeerLauncher(adapters, cfg.defaultProfile());
AtomicReference<Function<String, Integer>> liveCountRef = new AtomicReference<>(_ -> 0);
PeerLauncher workers = new CompositePeerLauncher(
adapters,
cfg.effectiveDefaultProfile(),
config,
profileName -> liveCountRef.get().apply(profileName));
// CB-504: under supervision (launchd/systemd) bridged can start before herdr's socket
// exists. The client itself is lazy — it connects per call — but the orphan reap below is
// the first thing that actually talks to herdr, so without this wait a boot-order race
// would crash the daemon into a restart loop. Wait, then degrade rather than die: serving
// with /healthz reporting "degraded" is strictly more useful than exiting.
if (awaitHerdr(herdr)) {
boolean herdrUp = awaitHerdr(herdr);
if (herdrUp) {
// CB-117: herdr keeps worker panes alive across a daemon restart, and their ids died
// with the previous process — reap those leaked orphans now, before we start serving.
workers.reapOrphanWorkers();
@@ -135,7 +166,12 @@ public final class Bridged {
&& cfg.lifecycle().contextCap() > 0) {
contextCap = cfg.lifecycle().contextCap();
}
SessionManager sessions = new SessionManager(workers, new GitWorktrees(cfg.worktreeRoot()), contextCap);
boolean clearAfterTurn = cfg.lifecycle() != null && cfg.lifecycle().clearAfterTurn();
SessionManager sessions = new SessionManager(workers, new GitWorktrees(cfg.worktreeRoot()),
System::nanoTime, contextCap, clearAfterTurn);
liveCountRef.set(profileName -> (int) sessions.roster().stream()
.filter(s -> profileName.equals(s.profile()))
.count());
// CB-303 part 1: idle-ttl reaper — only when configured, defaults to disabled.
final SessionReaper reaper;
@@ -148,6 +184,69 @@ public final class Bridged {
reaper = null;
}
// CB-530: every pane the config names as a lead, merged from `leaders:` and the legacy
// singular pin. PrimaryRegistry below still tracks ONE terminal — it addresses the push
// loop's nudges, which need a single destination — so it keeps the legacy pin.
Map<String, String> leadTerminals = cfg.leaderTerminals();
if (leadTerminals.size() > 1) {
log.info("leads: {} panes recognised {}", leadTerminals.size(), leadTerminals.values());
}
// CB-531: on top of the static registry, discover leads by the tab labels the operator
// writes. CB-557 moved the settings onto the lead they describe, so scanning is on whenever
// a `fleet.leaders:` entry exists — with no leads configured the supplier is a constant and
// never touches herdr, exactly as a missing `leadScan:` block used to behave.
final Supplier<Map<String, String>> leads;
var leaders = cfg.fleet().leaders();
if (!leaders.isEmpty()) {
// One scanner, so one prefix and one interval. Distinct per-lead prefixes would need a
// scanner each; until a config actually wants that, take the first entry's settings and
// say so, rather than silently honouring one lead's prefix and dropping another's.
var scan = leaders.values().iterator().next();
Set<String> memberSpaces = cfg.profiles().values().stream()
.map(BridgedConfig.Profile::workspace)
.filter(Objects::nonNull)
.collect(Collectors.toSet());
leads = new LeadTabScanner(herdr, scan.tabPrefix(), memberSpaces, leadTerminals,
TimeUnit.SECONDS.toNanos(scan.scanIntervalSeconds()), System::nanoTime);
log.info("lead scan: tabs labelled '{}…' host a lead (rescan every {}s, member spaces {} "
+ "excluded)",
scan.tabPrefix(), scan.scanIntervalSeconds(), memberSpaces);
long distinctPrefixes = leaders.values().stream()
.map(BridgedConfig.Leader::tabPrefix).distinct().count();
if (distinctPrefixes > 1) {
log.warn("fleet.leaders declares {} different tabPrefix values; only '{}' is scanned "
+ "for. Give every lead the same tabPrefix, or leads under the others "
+ "will not be discovered.",
distinctPrefixes, scan.tabPrefix());
}
} else {
leads = () -> leadTerminals;
}
// CB-558: start any declared lead that is not already running. After the scanner is built,
// because both read the same tab labels and the ordering makes that dependency visible; and
// only when herdr answered, because the launcher's whole safety property is that it can
// count live leads first — it must never guess and risk a second orchestrator.
if (herdrUp && !leaders.isEmpty()) {
int launched = new LeadLauncher(agents, spaces, cfg).ensureLeads();
if (launched > 0) {
log.info("lead auto-launch: {} lead(s) started", launched);
}
}
// CB-548: config-declared architect slots. Config supplies only the stable name → profile
// map; the terminal → slot binding is owned by the registry and is empty at startup, so no
// pane resolves to an architect until the later spawn lifecycle binds one. The registry is
// what CallerResolver resolves against and what that lifecycle will read profiles from;
// nothing here spawns a slot.
MemberRegistry members = new MemberRegistry(cfg.fleet());
sessions.setMemberLifecycle(members);
if (!members.slots().isEmpty()) {
log.info("member slots: {} configured {} — none bound yet (a slot is idle until the "
+ "spawn lifecycle binds a live terminal to it)",
members.slots().size(), members.slots().keySet());
}
// Status-gated injector (CB-103): the single writer into workers, fed by a poller.
// The blocking message endpoint (CB-104) is the producer; the poller is inert until then.
// CB-106: a confirmed turn completion resolves a blocked send whose worker never replied.
@@ -155,7 +254,7 @@ public final class Bridged {
CompletionResolver completion = new CompletionResolver(agents, rendezvous);
// CB-113: deliver only to an available worker (its MCP is connected), never its boot window.
// CB-301: the manager's presence bridge records availability and drives SPAWNING → READY.
WorkerPresence presence = sessions.asPresence();
MemberPresence presence = sessions.asPresence();
TurnListener turnListener = new TurnListener() {
@Override
public void onTurnComplete(String target) {
@@ -163,6 +262,17 @@ public final class Bridged {
sessions.onTurnComplete(target);
}
@Override
public boolean hasPostTurnAction(String target) {
return sessions.hasPostTurnAction(target);
}
@Override
public boolean onTurnCompleteWithPostAction(String target) {
completion.resolveBeforePostAction(target);
return sessions.onTurnCompleteWithPostAction(target);
}
@Override
public void onDelivered(String target) {
completion.onDelivered(target);
@@ -175,8 +285,9 @@ public final class Bridged {
sessions.onTurnFailed(target);
}
};
Injector injector = new Injector(agents, turnListener, presence::isPresent, presence::forget);
StatusPoller poller = new StatusPoller(agents, injector, INJECT_POLL_MILLIS);
Injector injector = new Injector(agents, turnListener, deliverableTo(presence, leads),
presence::forget);
StatusPoller poller = new StatusPoller(agents, injector, Injector.POLL_INTERVAL_MILLIS);
poller.start();
// CB-307: reply inbox. A broker: block (with a uri) selects the AMQP-backed durable adapter;
@@ -191,9 +302,19 @@ public final class Bridged {
log.info("reply inbox: in-memory (soft-state)");
}
// CB-307: learn the primary's terminal from orchestration tool calls (or pin from config).
PrimaryRegistry primaryRegistry = new PrimaryRegistry(
cfg.primary() != null ? cfg.primary().terminal() : null);
// The pin also feeds CallerResolver below: a primary running inside a herdr pane would
// otherwise resolve as a worker and be refused every orchestration tool.
String pinnedPrimaryTerminal = cfg.primary() != null ? cfg.primary().terminal() : null;
PrimaryRegistry primaryRegistry = new PrimaryRegistry(pinnedPrimaryTerminal);
// CB-532: `primary.terminal` is superseded and no longer needed for either of its jobs —
// identity comes from `leaders:`/`leadScan:`, and reply nudges now follow the delegating
// lead. Say so once at startup rather than leaving a redundant pin to look load-bearing.
if (pinnedPrimaryTerminal != null && !pinnedPrimaryTerminal.isBlank()) {
log.warn("primary.terminal is DEPRECATED (CB-532) and can be deleted: identity now comes "
+ "from leaders:/leadScan:, and reply nudges follow the lead that delegated. "
+ "It still works, and is still the fallback nudge destination when a restart "
+ "has lost the delegation map. Its pushReminders/pushBackoffMs stay valid.");
}
// CB-307: active push-to-primary loop — nudge the primary when replies land without an
// open bridge_send. Uses its own lightweight scheduled executor, separate from the injector.
int maxReminders = cfg.primary() != null ? cfg.primary().remindersOrDefault() : 5;
@@ -206,14 +327,37 @@ public final class Bridged {
Metrics metrics = BridgedMetrics.create(sessions, replyInbox);
var pushLoop = new ReplyPushLoop(primaryRegistry, agents, replyInbox,
pushScheduler, maxReminders, backoffMs, metrics);
// CB-551: idle-lead heartbeat. Opt-in; absent `leadHeartbeat:` this is never constructed, so
// an upgraded daemon cannot silently start spending subscription on nudging an idle lead.
// It has its own single-thread scheduler and holds its own scheduler shutdown via close().
final LeadHeartbeatLoop heartbeat;
var heartbeatScheduler = Executors.newSingleThreadScheduledExecutor(r ->
Thread.ofVirtual().name("bridge-heartbeat-").unstarted(r));
if (cfg.leadHeartbeat() != null) {
var hb = cfg.leadHeartbeat();
heartbeat = new LeadHeartbeatLoop(primaryRegistry, agents, replyInbox, sessions::roster,
pushLoop, heartbeatScheduler, System::nanoTime,
TimeUnit.SECONDS.toNanos(hb.idleAfterSeconds()), hb.backoffMs(), hb.quietNudgeCap(),
metrics);
heartbeat.start();
} else {
heartbeat = null;
heartbeatScheduler.shutdownNow();
}
MessageService messages = new MessageService(agents, injector, rendezvous, replyInbox,
pushLoop, metrics);
// CB-520: the reply inbox only consumes for agents this gateway owns. own on acquire,
// release on teardown. Do this before CB-516 so the inbox is owned before any reply can land.
sessions.onAcquire(replyInbox::own);
// CB-516: releasing a worker must fail whatever send was waiting on it. Without this a
// torn-down delegation kept reporting PENDING until the 30-minute async timeout, and never
// reached /metrics — the delegation was unresolvable and nothing said so.
sessions.onRelease(terminal ->
messages.abandon(terminal, "the worker session was released before it replied"));
sessions.onRelease(terminal -> {
messages.abandon(terminal, "the worker session was released before it replied");
replyInbox.release(terminal);
primaryRegistry.forgetDelegation(terminal); // CB-532: don't leak the lead binding
});
// MCP server face (CB-105): bridge_send/bridge_reply/bridge_status, mounted at /mcp.
// Caller identity is resolved from the connection (peer PID → herdr pane), not arguments.
@@ -229,17 +373,27 @@ public final class Bridged {
throw new IllegalStateException("auth.mode=token but env var " + cfg.auth().tokenEnv()
+ " is unset or empty — export it before starting bridged");
}
callers = new CallerResolver(identity, true, token);
callers = CallerResolver.withLeadsAndMembers(identity, true, token, leads, members);
log.info("auth: token mode (bearer required for non-worker callers, env {})",
cfg.auth().tokenEnv());
} else {
callers = new CallerResolver(identity);
callers = CallerResolver.withLeadsAndMembers(identity, false, null, leads, members);
log.info("auth: loopback-trust (any loopback non-worker caller is the primary)");
}
BridgeMcp mcp = new BridgeMcp(messages, workers, sessions, identity, presence,
primaryRegistry, callers, metrics);
// CB-559: opt-in config reload. With no `configReload:` block nothing is constructed, so an
// upgraded daemon behaves exactly as before — the file is read once at boot and never again.
final ConfigWatcher configWatcher;
if (cfg.configReload() != null && cfg.configReload().isEnabled()) {
configWatcher = new ConfigWatcher(config, cfg.configReload().intervalSeconds());
configWatcher.start();
} else {
configWatcher = null;
}
// CB-303 part 3: single ordered shutdown hook. Drain sessions first while herdr is still
// open (so releases reach the daemon), then stop poller/message/mcp/reaper, and close herdr
// last. This replaces the earlier independent hooks that could race and close herdr early.
@@ -248,6 +402,8 @@ public final class Bridged {
poller.stop();
messages.close();
pushLoop.close();
if (heartbeat != null) heartbeat.close(); // CB-551: stop the idle-lead heartbeat scheduler
if (configWatcher != null) configWatcher.stop(); // CB-559: stop polling the config file
mcp.close();
if (reaper != null) reaper.stop();
// Release the broker connection last among message resources (no-op for the in-memory inbox).
@@ -268,6 +424,29 @@ public final class Bridged {
cfg.bind().host(), cfg.bind().port(), socket);
}
/**
* The {@link Injector}'s readiness gate (CB-534): a target is deliverable if it is a spawned
* member whose agent has connected the bridge MCP, <em>or</em> a lead.
*
* <p>The gate exists for one reason — to hold a delivery out of a <em>spawned</em> member's boot
* window, where herdr already reports {@code idle} but the TUI would drop an injected paste. That
* hazard is a property of spawning. A lead is never spawned: the operator started it and named it
* (or labelled its tab) only once it was up, so there is no boot window to guard.
*
* <p>A lead is also never enrolled in {@link MemberPresence} — {@code BridgeMcp} marks presence
* for every spawned member (worker and architect), deliberately, since that map doubles as the
* member roster's availability signal and a lead counted there would show up as an available
* member. So without the second disjunct a lead is permanently un-deliverable: every
* lead→lead send sat on the gate for {@code READINESS_GRACE_POLLS} (~60s) and then failed
* having never been typed into the pane.
*
* <p>The lead set is read through the supplier on each call rather than snapshotted, so a lead
* discovered by {@code leadScan} after startup becomes deliverable without a restart.
*/
static Predicate<String> deliverableTo(MemberPresence presence, Supplier<Map<String, String>> leads) {
return target -> presence.isPresent(target) || leads.get().containsKey(target);
}
/**
* Poll herdr's {@code ping} until it answers or {@link #HERDR_WAIT_SECONDS} elapses (CB-504).
*
@@ -46,18 +46,28 @@ public final class Authz {
return false; // authenticated as nothing ⇒ authorized for nothing
}
return switch (action) {
// Orchestration is the primary's alone. A worker driving spawn/stop/send would be a
// worker escalating into the orchestrator role.
case SPAWN, STOP, SEND, DRAIN -> caller.isPrimary();
// Fleet lifecycle is the primary's alone — spawn, stop, drain. An architect
// deliberately does NOT get these (CB-548), so it cannot tear down or stand up workers
// even though it coordinates them; and a worker driving any of these would be a worker
// escalating into the orchestrator role.
case SPAWN, STOP, DRAIN -> caller.isPrimary();
// The load-bearing rule: a worker acts only as itself. The primary is deliberately
// excluded — a reply/ask is a worker's own turn output, and letting the primary forge
// one would corrupt the rendezvous correlation it is itself waiting on.
// Delivering a turn is open to the primary and the architect: an architect delegates
// to workers (that is the role's point) but still has no lifecycle rights. A worker is
// excluded — sending would be it escalating.
case SEND -> caller.isPrimary() || caller.isArchitect();
// The load-bearing rule: a caller acts only as the pane it occupies. CB-532 widened who
// that can be — a lead answering another lead is replying for its OWN terminal, which
// this already permits — while the rule itself is unchanged, and is what stops anyone
// forging a reply for a rendezvous someone else is waiting on. An architect's own pane
// passes through the same check, so it can answer a funnel that delegated to it. An
// unnamed primary (token/loopback, no pane) owns nothing and is still excluded.
case REPLY, ASK -> caller.ownsSession(targetSession);
// Observation is open to both authenticated roles: a worker legitimately polls its own
// Observation is open to every authenticated role: a worker legitimately polls its own
// status, and the roster carries no secrets.
case READ, METRICS -> caller.isPrimary() || caller.isWorker();
case READ, METRICS -> caller.isPrimary() || caller.isWorker() || caller.isArchitect();
};
}
@@ -1,9 +1,13 @@
package dev.ltms.bridged.auth;
import dev.ltms.bridged.mcp.ConnectionIdentity;
import dev.ltms.bridged.peer.MemberRole;
import java.nio.charset.StandardCharsets;
import java.security.MessageDigest;
import java.util.Map;
import java.util.function.Function;
import java.util.function.Supplier;
/**
* Resolves every caller to a {@link Principal}, for both entry paths into the core (CB-501).
@@ -15,7 +19,18 @@ import java.security.MessageDigest;
*
* <p><strong>Resolution order</strong> — connection identity first, token second, nothing third:
* <ol>
* <li>A loopback peer PID that maps to a herdr worker pane ⇒ {@link Role#WORKER}. This is
* <li>A loopback peer PID that maps to a pane named by {@code leaders:}, by the legacy
* {@code primary.terminal} pin, or by an operator-labelled lead tab (CB-307, CB-530, CB-531)
* ⇒ {@link Role#PRIMARY}, carrying that lead's
* name. The pane mapping is as unforgeable as a worker's, and the config explicitly names
* that pane as a lead's own — without this rule a lead running <em>inside</em> a herdr pane
* is misread as a worker and locked out of orchestration. More than one pane may be named,
* so two leads can work as peers rather than one being demoted.</li>
* <li>A loopback peer PID that maps to a pane bound to a CB-548 architect slot ⇒
* {@link Role#ARCHITECT}, carrying the slot name. Just unforgeable as a worker's, and
* resolved from the <em>live</em> terminal→slot binding (never a request argument), before
* the generic worker fallback.</li>
* <li>A loopback peer PID that maps to any other herdr pane ⇒ {@link Role#WORKER}. This is
* unforgeable (the OS reports the PID, herdr owns the PID→pane map) and is honoured
* regardless of auth mode, so enabling auth never breaks the fleet.</li>
* <li>Otherwise, under {@code token} mode, a valid bearer token ⇒ {@link Role#PRIMARY}.</li>
@@ -29,18 +44,122 @@ public final class CallerResolver {
private final ConnectionIdentity identity;
private final boolean tokenMode;
private final byte[] expectedToken; // null unless tokenMode
/**
* terminal_id → lead name; empty when nothing is pinned. CB-530.
*
* <p>A supplier rather than a map because the registry is no longer fixed at startup: CB-531
* discovers leads by scanning herdr for operator-labelled tabs, so a lead that opens its tab
* after the daemon booted must still be recognised. Consulted per resolve; the scanner behind
* it is TTL-cached, so this is a map lookup in the common case.
*/
private final Supplier<Map<String, String>> leadTerminals;
/**
* terminal_id → architect slot name; empty when nothing is configured. CB-548.
*
* <p>Like {@link #leadTerminals}, a supplier rather than a fixed map, so a binding injected
* after startup — when the later spawn lifecycle establishes a live architect session, or an
* operator pins one — takes effect without a restart. Consulted per resolve; today's wiring
* in {@code Bridged} reads a constant from config, which is the degenerate live case.
*/
private final Supplier<Map<String, String>> architectTerminals;
private final Function<String, MemberRole> memberSlotRoles;
private final Function<String, String> memberSlotNames;
/** Loopback-trust resolver: no token required, historical behaviour. */
public CallerResolver(ConnectionIdentity identity) {
this(identity, false, null);
/** Loopback-trust resolver: no token required, historical behaviour. Test-only. */
CallerResolver(ConnectionIdentity identity) {
this(identity, false, null, Map.of());
}
/** As {@link #CallerResolver(ConnectionIdentity, boolean, String, Map)} with no leads pinned. Test-only. */
CallerResolver(ConnectionIdentity identity, boolean tokenMode, String token) {
this(identity, tokenMode, token, Map.of());
}
/**
* @param identity connection-based worker identification
* @param tokenMode when true, a non-worker caller must present a valid bearer token
* @param token the expected bearer token; required (non-blank) when {@code tokenMode}
* Single-pin form for the legacy {@code primary.terminal}-only configuration — one lead, named
* {@code primary}.
*
* <p>A static factory rather than a fourth constructor overload on purpose: {@code String} and
* {@code Map} overloads are ambiguous for a literal {@code null} argument, which is a compile
* error at the call site and exactly the shape "unpinned" is written in.
*
* @param pinnedPrimaryTerminal the primary's own herdr {@code terminal_id}
* ({@code null}/blank = unpinned)
*/
public CallerResolver(ConnectionIdentity identity, boolean tokenMode, String token) {
static CallerResolver pinnedTo(ConnectionIdentity identity, boolean tokenMode,
String token, String pinnedPrimaryTerminal) {
return new CallerResolver(identity, tokenMode, token,
pinnedPrimaryTerminal == null || pinnedPrimaryTerminal.isBlank()
? Map.of() : Map.of(pinnedPrimaryTerminal, "primary"));
}
/**
* @param identity connection-based worker identification
* @param tokenMode when true, a non-worker caller must present a valid bearer token
* @param token the expected bearer token; required (non-blank) when {@code tokenMode}
* @param leadTerminals herdr {@code terminal_id} → lead name for every configured lead
* (CB-530). A caller resolving to one of these panes is that lead — a
* {@link Role#PRIMARY} — rather than a worker. Empty = nothing pinned,
* so every pane resolves as a worker.
*/
CallerResolver(ConnectionIdentity identity, boolean tokenMode, String token,
Map<String, String> leadTerminals) {
this(identity, tokenMode, token, fixed(leadTerminals), null);
}
/**
* Live-registry form: {@code leadTerminals} is consulted on every resolve, so leads discovered
* after startup (CB-531's tab scan) take effect without a restart.
*
* <p>A static factory rather than a fourth constructor overload, for the same reason as
* {@link #pinnedTo}: {@code Map} and {@code Supplier} overloads are ambiguous for a literal
* {@code null}.
*/
static CallerResolver withLeads(ConnectionIdentity identity, boolean tokenMode,
String token,
Supplier<Map<String, String>> leadTerminals) {
return new CallerResolver(identity, tokenMode, token, leadTerminals, null);
}
/**
* Live registry form that can confirm a bound slot is an architect slot.
*
* <p>This is the only public construction path. It keeps terminal bindings and slot roles in
* the same {@link MemberRegistry}, so a configured architect can resolve as an architect.
*/
public static CallerResolver withLeadsAndMembers(ConnectionIdentity identity,
boolean tokenMode, String token,
Supplier<Map<String, String>> leadTerminals,
MemberRegistry members) {
return new CallerResolver(identity, tokenMode, token, leadTerminals,
members == null ? null : members::snapshot,
members == null ? null : members::roleForSlot,
members == null ? null : members::nameForSlot);
}
private static Supplier<Map<String, String>> fixed(Map<String, String> leadTerminals) {
Map<String, String> snapshot = leadTerminals == null ? Map.of() : Map.copyOf(leadTerminals);
return () -> snapshot;
}
private CallerResolver(ConnectionIdentity identity, boolean tokenMode, String token,
Supplier<Map<String, String>> leadTerminals,
Supplier<Map<String, String>> architectTerminals) {
this(identity, tokenMode, token, leadTerminals, architectTerminals, null);
}
private CallerResolver(ConnectionIdentity identity, boolean tokenMode, String token,
Supplier<Map<String, String>> leadTerminals,
Supplier<Map<String, String>> architectTerminals,
Function<String, MemberRole> memberSlotRoles) {
this(identity, tokenMode, token, leadTerminals, architectTerminals, memberSlotRoles, null);
}
private CallerResolver(ConnectionIdentity identity, boolean tokenMode, String token,
Supplier<Map<String, String>> leadTerminals,
Supplier<Map<String, String>> architectTerminals,
Function<String, MemberRole> memberSlotRoles,
Function<String, String> memberSlotNames) {
if (tokenMode && (token == null || token.isBlank())) {
throw new IllegalArgumentException(
"auth.mode=token requires a non-empty token; check that the env var named by "
@@ -49,6 +168,34 @@ public final class CallerResolver {
this.identity = identity;
this.tokenMode = tokenMode;
this.expectedToken = tokenMode ? token.getBytes(StandardCharsets.UTF_8) : null;
this.leadTerminals = leadTerminals == null ? Map::of : leadTerminals;
this.architectTerminals = architectTerminals == null ? Map::of : architectTerminals;
this.memberSlotRoles = memberSlotRoles == null ? _ -> null : memberSlotRoles;
this.memberSlotNames = memberSlotNames == null ? Function.identity() : memberSlotNames;
}
/**
* The currently-recognised leads, {@code terminal_id → name} (CB-535).
*
* <p>Deliberately read from the same supplier {@link #resolve} consults, rather than from a
* second copy handed to the roster: a lead that is <em>listed</em> but would not <em>resolve</em>
* (or the reverse) is an address a peer cannot actually reach, and the two answers drifting apart
* is precisely the confusion this exists to end. Live, so a lead discovered by the tab scan after
* startup appears without a restart.
*/
public Map<String, String> leads() {
return leadTerminals.get();
}
/**
* The currently-recognised architect slots, {@code terminal_id → slot name} (CB-548).
*
* <p>Read from the same supplier {@link #resolve} consults, so a slot that is <em>listed</em>
* here but would not <em>resolve</em> (or the reverse) cannot drift apart. Live for the same
* reason as {@link #leads()}.
*/
public Map<String, String> members() {
return architectTerminals.get();
}
/**
@@ -61,6 +208,22 @@ public final class CallerResolver {
public Principal resolve(String remoteAddr, int remotePort, String authorizationHeader) {
ConnectionIdentity.Caller c = identity.resolve(remoteAddr, remotePort);
if (c.terminal() != null) {
String lead = leadTerminals.get().get(c.terminal());
if (lead != null) {
// The config names this pane as a lead's own. The pane mapping is exactly as
// unforgeable as a worker's, so it outranks the token path — no credential needed.
// Checked before the architect registry so a pane named in BOTH is still the lead
// (CB-548 preserves every existing leader behaviour).
return Principal.leader(lead, c.terminal(), c.pid());
}
String slot = architectTerminals.get().get(c.terminal());
if (slot != null && memberSlotRoles.apply(slot) == MemberRole.ARCHITECT) {
// The config/live binding names this pane as an architect slot's own. Same
// unforgeable pane mapping; the live binding, never a request argument, decides.
// Check the slot role too: this defence in depth prevents a bad lifecycle bind from
// escalating a dev or reviewer into an architect. Checked before the worker fallback.
return Principal.architect(memberSlotNames.apply(slot), c.terminal(), c.pid());
}
return Principal.worker(c.terminal(), c.pid()); // unforgeable; never token-gated
}
@@ -76,16 +239,6 @@ public final class CallerResolver {
return isLoopback(remoteAddr) ? Principal.primary(c.pid()) : Principal.anonymous();
}
/** The working directory of the calling process (CB-112 spawn cwd inheritance), or {@code null}. */
public String cwdForPid(long pid) {
return identity.cwdForPid(pid);
}
/** True when auth requires a bearer token of non-worker callers. */
public boolean tokenMode() {
return tokenMode;
}
private boolean presentedTokenMatches(String authorizationHeader) {
String presented = bearerValue(authorizationHeader);
if (presented == null) {
@@ -0,0 +1,21 @@
package dev.ltms.bridged.auth;
import dev.ltms.bridged.peer.MemberRole;
/** Optional session lifecycle hook for live member-slot bindings. */
public interface MemberLifecycle {
MemberLifecycle NONE = new MemberLifecycle() {
@Override
public void acquired(MemberRole role, String profile, String terminal) {
}
@Override
public void released(String terminal) {
}
};
void acquired(MemberRole role, String profile, String terminal);
void released(String terminal);
}
@@ -0,0 +1,232 @@
package dev.ltms.bridged.auth;
import dev.ltms.bridged.config.BridgedConfig;
import dev.ltms.bridged.peer.MemberRole;
import org.slf4j.Logger;
import org.slf4j.LoggerFactory;
import java.util.Collections;
import java.util.HashMap;
import java.util.LinkedHashMap;
import java.util.Map;
import java.util.Objects;
/**
* The architect-slot registry (CB-548): every gateway-local architect name and the strong-model
* profile it points at, plus the <em>live</em> bindings from a live architect's herdr terminal to
* its slot.
*
* <p>Two halves, split by who owns each:
* <ul>
* <li><b>slots</b> — configured once, keyed by the gateway-local unique name; each carries the
* {@code profile} reference the spawn lifecycle reads when it stands the slot up. A read-only
* snapshot taken at construction.</li>
* <li><b>terminal bindings</b> — owned by this registry and initially <em>empty</em>. Config
* declares no architect terminal, so at startup every slot is idle and nothing resolves to an
* architect; a session only becomes one when the spawn lifecycle {@linkplain #bind(String,
* String) binds} its terminal to a slot. {@link CallerResolver} reads this through
* {@link #snapshot()} to turn a pane into an {@link Role#ARCHITECT}.</li>
* </ul>
*
* <p>Spawning/lifecycle is deliberately a separate unit: this class only owns the bindings and
* exposes the map the resolver resolves against plus the profile lookup lifecycle will call.
* Nothing here creates or manages an architect session.
*/
public final class MemberRegistry implements MemberLifecycle {
private static final Logger log = LoggerFactory.getLogger(MemberRegistry.class);
/**
* One flattened {@code fleet:} entry.
*
* <p>Flattened because a slot name is unique only <em>within</em> its pool — {@code sonnet} may
* legitimately be both a developer and a reviewer — while a terminal binds to exactly one thing.
* The qualified {@link #key()} is what that binding uses.
*
* @param name the slot's key inside its pool
* @param role the pool it came from
* @param profile the backend it runs on
*/
public record Entry(String name, MemberRole role, String profile) {
/** {@code "architect:opus"} — unique across pools, unlike {@link #name()}. */
public String key() {
return role.wireName() + ":" + name;
}
}
private final Map<String, Entry> slots;
/** Live {@code terminal_id → qualified slot key}; guarded by {@code this}. */
private final Map<String, String> terminalToSlot = new HashMap<>();
/** Flatten every role pool in {@code fleet} into one registry. Leaders are not members. */
public MemberRegistry(BridgedConfig.Fleet fleet) {
Map<String, Entry> flat = new LinkedHashMap<>();
if (fleet != null) {
for (MemberRole role : MemberRole.values()) {
fleet.pool(role).forEach((name, slot) -> {
if (slot != null) {
Entry e = new Entry(name, role, slot.profile());
flat.put(e.key(), e);
}
});
}
}
this.slots = Collections.unmodifiableMap(flat);
}
/** The configured slots, keyed by qualified {@link Entry#key()}. Unmodifiable snapshot. */
public Map<String, Entry> slots() {
return slots;
}
/** The slots belonging to {@code role}, in definition order. */
public Map<String, Entry> slotsFor(MemberRole role) {
Map<String, Entry> out = new LinkedHashMap<>();
slots.forEach((key, e) -> {
if (e.role() == role) {
out.put(key, e);
}
});
return Collections.unmodifiableMap(out);
}
/**
* An immutable copy of the live {@code terminal_id → slot name} bindings.
*
* <p>Passed to {@link CallerResolver} as the source of architect identity, and what
* {@code bridge_whoami}/the roster will read to say which slot a pane hosts. Empty until the
* spawn lifecycle binds a slot.
*/
public Map<String, String> snapshot() {
synchronized (terminalToSlot) {
return Map.copyOf(terminalToSlot);
}
}
/** The slot a live terminal is bound to, or {@code null} if it is not an architect slot. */
public String slotForTerminal(String terminal) {
if (terminal == null) {
return null;
}
synchronized (terminalToSlot) {
return terminalToSlot.get(terminal);
}
}
/**
* The strong-model profile a slot runs under — what the spawn lifecycle reads.
*
* @return the slot's configured {@code profile}, or {@code null} if the slot is unknown or
* declares none
*/
public String profileForSlot(String slotName) {
Entry e = slots.get(slotName);
return (e == null || e.profile() == null) ? null : e.profile();
}
/** The role a qualified slot key belongs to, or {@code null} when the key is unknown. */
public MemberRole roleForSlot(String slotName) {
Entry e = slots.get(slotName);
return e == null ? null : e.role();
}
/** The unqualified configured name for a slot, or {@code null} if it is unknown. */
public String nameForSlot(String slotName) {
Entry e = slots.get(slotName);
return e == null ? null : e.name();
}
/** True when {@code slotName} is a configured architect slot. */
public boolean isSlot(String slotName) {
return slots.containsKey(slotName);
}
/**
* Bind {@code terminal} to {@code slot} (CB-548).
*
* <p>The spawn lifecycle calls this when it stands a slot up. The bind is atomic and preserves
* the two cardinality invariants: a terminal may occupy at most one slot, and a slot may host at
* most one terminal. Binding the same terminal to the same slot again is a harmless no-op.
*
* @param slot a configured slot name, or the bind is refused
* @param terminal the pane that will act as this architect
* @return {@code true} if the binding is now {@code terminal → slot}; {@code false} if it was
* refused — an unknown slot, a terminal already bound to a different slot, or a slot
* already hosting a different terminal
*/
public boolean bind(String slot, String terminal) {
if (slot == null || terminal == null || terminal.isBlank()) {
return false;
}
synchronized (terminalToSlot) {
if (!isSlot(slot)) {
return false; // unknown slot — nothing to bind to
}
String existingSlot = terminalToSlot.get(terminal);
if (existingSlot != null) {
return slot.equals(existingSlot); // already this slot (idempotent) or a different one
}
if (terminalToSlot.containsValue(slot)) {
return false; // slot already hosts a terminal — no second one
}
terminalToSlot.put(terminal, slot);
return true;
}
}
/**
* Compare-safe unbind of {@code expectedTerminal} from {@code slot} (CB-548).
*
* <p>The spawn lifecycle calls this when it tears a slot down. Only the exact binding
* {@code expectedTerminal → slot} is removed; if that terminal was since rebound to a different
* slot (or the slot to a different terminal), the call is a no-op returning {@code false} — a
* stale unbind must never remove a replacement.
*
* @param slot the slot the caller believes the terminal is bound to
* @param expectedTerminal the terminal it expects to be bound there
* @return {@code true} if {@code expectedTerminal → slot} was removed; {@code false} if nothing
* was (no such binding, or the binding had already moved)
*/
public boolean unbind(String slot, String expectedTerminal) {
if (slot == null || expectedTerminal == null) {
return false;
}
synchronized (terminalToSlot) {
String current = terminalToSlot.get(expectedTerminal);
if (current == null || !slot.equals(current)) {
return false; // absent, or a replacement/moved binding — leave it in place
}
terminalToSlot.remove(expectedTerminal);
return true;
}
}
/**
* Bind only architect sessions to a free slot with the resolved profile.
*
* <p>The role check is lifecycle policy. {@link CallerResolver} repeats it when resolving a
* binding, so a later lifecycle regression cannot turn a worker into an architect.
*/
@Override
public void acquired(MemberRole role, String profile, String terminal) {
if (role != MemberRole.ARCHITECT || terminal == null || terminal.isBlank()) {
return;
}
// slotsFor preserves definition order, so duplicate-profile slots use the first free one.
for (Entry entry : slotsFor(MemberRole.ARCHITECT).values()) {
if (Objects.equals(profile, entry.profile()) && bind(entry.key(), terminal)) {
return;
}
}
log.info("member slot: no free architect slot for profile={}; session remains a worker", profile);
}
/** Unbind a released terminal using the compare-safe registry operation. */
@Override
public void released(String terminal) {
String slot = slotForTerminal(terminal);
if (slot != null) {
unbind(slot, terminal);
}
}
}
@@ -5,52 +5,122 @@ package dev.ltms.bridged.auth;
* identifies which worker it is (CB-501).
*
* @param role what this caller is authorized to act as
* @param terminal the worker's herdr terminal id; {@code null} for {@code PRIMARY}/{@code ANONYMOUS}
* @param terminal the herdr {@code terminal_id} of the pane this caller occupies — a worker's, or
* (since CB-532) a named lead's; {@code null} for an unnamed primary resolved off a
* token or loopback trust, and for {@code ANONYMOUS}
* @param pid the connecting process id, or {@code -1} when not resolvable (audit context)
* @param name for a lead resolved from the CB-530 {@code leaders:} registry, which lead it is;
* for an architect resolved from the CB-548 {@code architects:} registry, which
* slot it occupies; {@code null} for every other caller, including an unnamed primary
*/
public record Principal(Role role, String terminal, long pid) {
public record Principal(Role role, String terminal, long pid, String name) {
/**
* Three-arg form for the callers that have no name to carry (workers, anonymous, and the
* token/loopback primary paths). Kept so adding CB-530's {@code name} did not churn every
* construction site — and so a reconstructed principal without a stashed name still works.
*/
public Principal(Role role, String terminal, long pid) {
this(role, terminal, pid, null);
}
/** A caller authenticated as nothing — the default when no check establishes anything else. */
public static Principal anonymous() {
return new Principal(Role.ANONYMOUS, null, -1);
}
/** The orchestrating session. */
/** The orchestrating session, unnamed (token or loopback-trust path). */
public static Principal primary(long pid) {
return new Principal(Role.PRIMARY, null, pid);
}
/**
* A named lead from the {@code leaders:} registry (CB-530).
*
* <p>Carries {@link Role#PRIMARY}: a lead <em>is</em> a primary as far as authorization goes,
* so every existing {@code isPrimary()} gate keeps working unchanged and the role table needed
* no new entry. The name is reporting only — it lets {@code bridge_whoami} say <em>which</em>
* lead is asking once more than one is configured.
*
* <p><strong>CB-532: a lead now carries the terminal it was matched by.</strong> Under CB-530 it
* deliberately did not, because {@code terminal} meant "which worker pane" everywhere and a
* non-null one would have enrolled the lead in the worker presence map. That reading was what
* made a lead unaddressable: {@link #ownsSession} could never be true for it, so
* {@code bridge_reply} was refused and one lead could send to another but never be answered.
* The terminal now means "which pane is this caller", the presence map keys on
* {@link #isSpawnedMember()} instead, and a lead is a peer that can both send and receive.
*/
public static Principal leader(String name, String terminal, long pid) {
return new Principal(Role.PRIMARY, terminal, pid, name);
}
/** A worker peer, identified by its herdr pane. */
public static Principal worker(String terminal, long pid) {
return new Principal(Role.WORKER, terminal, pid);
}
/**
* An architect (CB-548), identified by the slot it occupies and the pane bound to it.
*
* <p>Carries {@link Role#ARCHITECT}. {@code slotName} is reporting only — it lets
* {@code bridge_whoami} say <em>which</em> architect slot is asking, and it is the key the
* (future) spawn lifecycle reads a profile back from. Identity is the {@code terminal}: like a
* worker's it comes from the connection and the live terminal→slot binding, so
* {@code ownsSession} works exactly as it does for a worker — an architect acts as its own
* pane and no other.
*/
public static Principal architect(String slotName, String terminal, long pid) {
return new Principal(Role.ARCHITECT, terminal, pid, slotName);
}
public boolean isPrimary() {
return role == Role.PRIMARY;
}
public boolean isArchitect() {
return role == Role.ARCHITECT;
}
public boolean isWorker() {
return role == Role.WORKER;
}
/**
* Whether this caller is a spawned member with its own pane.
*
* <p>Both workers and architects are spawned members. A lead is excluded because recording it
* as present would count it as an available member in the roster.
*/
public boolean isSpawnedMember() {
return role == Role.WORKER || role == Role.ARCHITECT;
}
public boolean isAnonymous() {
return role == Role.ANONYMOUS;
}
/**
* Whether this caller may act <em>as</em> {@code sessionId} — the "own session only" rule that
* keeps one worker from replying or asking on another's behalf. Only a worker can own a
* session, and only its own.
* keeps one peer from replying or asking on another's behalf.
*
* <p>The rule is about <em>identity</em>, not rank: a caller may act as the pane it demonstrably
* occupies, and as no other. CB-532 dropped the extra {@code isWorker()} conjunct that used to
* be here. It was not what enforced the rule — {@code terminal.equals(sessionId)} is, and that
* terminal comes from the connection, so it cannot be forged either way. All the conjunct did
* was make a lead permanently unable to answer anyone, since a lead's terminal was null and a
* lead is not a worker. A caller with no terminal at all (an off-host or token-authenticated
* primary) still owns nothing, which is the case the null check covers.
*/
public boolean ownsSession(String sessionId) {
return isWorker() && terminal != null && terminal.equals(sessionId);
return terminal != null && terminal.equals(sessionId);
}
/** Short, non-sensitive description for audit lines and error details. */
public String describe() {
return switch (role) {
case WORKER -> "worker:" + terminal;
case PRIMARY -> "primary";
case ARCHITECT -> "architect:" + name;
case PRIMARY -> name == null ? "primary" : "leader:" + name;
case ANONYMOUS -> "anonymous";
};
}
@@ -24,6 +24,16 @@ public enum Role {
*/
WORKER,
/**
* A config-declared architect slot (CB-548): a gateway-local named session on a strong-model
* profile that coordinates and delegates turns but does not own the fleet. Unforgeable like a
* worker's — derived from the connection's pane and the live terminal→slot binding, never from
* a request argument. May {@code SEND} a turn, {@code REPLY}/{@code ASK} only as its own pane,
* and {@code READ}/{@code METRICS}; may <em>not</em> {@code SPAWN}/{@code STOP}/{@code DRAIN}
* (those stay the primary's, to keep lifecycle in one pair of hands).
*/
ARCHITECT,
/** Authenticated as nothing. Authorized for nothing but {@code /healthz}. */
ANONYMOUS
}
File diff suppressed because it is too large Load Diff
@@ -0,0 +1,269 @@
package dev.ltms.bridged.config;
import org.slf4j.Logger;
import org.slf4j.LoggerFactory;
import java.nio.file.Path;
import java.util.ArrayList;
import java.util.LinkedHashSet;
import java.util.List;
import java.util.Map;
import java.util.Objects;
import java.util.Set;
import java.util.concurrent.atomic.AtomicReference;
import java.util.function.Supplier;
/**
* The daemon's live configuration, re-readable without a restart (CB-559).
*
* <p>Consumers hold this, not a {@link BridgedConfig}, and read through {@link #get()} at the point
* of use. A component that captures {@code ref.get()} into a field at construction has opted out of
* reload — which is sometimes right (see <em>deferred</em> below), but it must then be a deliberate
* choice rather than an accident of where the field was initialised.
*
* <h2>Not every key can change under a running daemon</h2>
* Keys fall into three classes, and the difference is about what already exists when the reload
* happens — not about how important the key is.
*
* <ul>
* <li><strong>Hot</strong> — re-read per use, so a reload takes effect on the next spawn:
* {@code fleet:} (every role pool and {@code tabLabel}), {@code placement:}, and an existing
* profile's {@code weight} / {@code maxLoad}. Those three are read through a supplier on
* {@code CompositePeerLauncher}, which is what makes them hot — not the fact that they are
* config.</li>
* <li><strong>Deferred</strong> — accepted into the new snapshot, but the wiring built at startup
* keeps the old value until a restart: {@code lifecycle:}, {@code leadHeartbeat:},
* {@code spawnReadyTimeoutMs} / {@code spawnReadyPollMs}, {@code guard:},
* {@code worktreeRoot:}, adding or removing a profile (a new backend needs its own launcher,
* which is constructed once), <em>and an existing profile's launch settings</em> —
* {@code model}, {@code baseUrl}, {@code argv}, {@code env}, {@code mcpUrl} and the rest.
* {@code HerdrPeerLauncher} takes {@code Map.copyOf(profiles)} at construction and resolves
* each spawn out of that copy, so those never reach a launch until the daemon restarts. A
* reload logs these rather than pretending they applied.</li>
* <li><strong>Cold</strong> — cannot change at all under a running daemon: {@code bind:},
* {@code herdrSocket:}, {@code broker:} and {@code auth:}. The socket is bound, the broker
* connection is open, and the auth mode decides who may reach the port that is already
* listening.</li>
* </ul>
*
* <p><strong>A cold change refuses the whole reload.</strong> Not the hot half applied and the cold
* half warned about: that would leave the running daemon in a state matching no file on disk, which
* is the worst thing a reload can do to an operator debugging one. Refusing keeps the invariant that
* the live config is always some version of the file, and the message names the keys that must
* change through a restart.
*
* <p>A reload that fails to parse or fails validation is also refused, and the previous config keeps
* running. A config file being edited is normally read once mid-save; degrading a working daemon
* because it caught a half-written file would be a bad trade.
*/
public final class ConfigRef implements Supplier<BridgedConfig> {
private static final Logger log = LoggerFactory.getLogger(ConfigRef.class);
/** Keys that cannot change under a running daemon — see the class doc. */
private static final Set<String> COLD_KEYS =
Set.of("bind", "herdrSocket", "broker", "auth");
private final Path path;
private final AtomicReference<BridgedConfig> current;
public ConfigRef(Path path, BridgedConfig initial) {
this.path = path;
this.current = new AtomicReference<>(Objects.requireNonNull(initial, "initial config"));
}
/** A fixed reference that never reloads — for tests and for wiring built from a config in code. */
public static ConfigRef fixed(BridgedConfig cfg) {
return new ConfigRef(null, cfg);
}
/** The live configuration. Read this per use; do not cache it in a field. */
@Override
public BridgedConfig get() {
return current.get();
}
/** The file this ref reloads from, or {@code null} for a {@link #fixed} ref. */
public Path path() {
return path;
}
/**
* What a reload attempt did.
*
* @param applied true when the new config is now live
* @param coldKeys cold keys whose value changed, which is why an unapplied reload was refused
* @param deferred keys that changed and were accepted, but whose effect waits for a restart
* @param error the parse or validation failure that refused the reload, else {@code null}
*/
public record Outcome(boolean applied, List<String> coldKeys, List<String> deferred,
String error) {
public Outcome {
coldKeys = List.copyOf(coldKeys);
deferred = List.copyOf(deferred);
}
static Outcome refusedCold(List<String> keys) {
return new Outcome(false, keys, List.of(), null);
}
static Outcome failed(String error) {
return new Outcome(false, List.of(), List.of(), error);
}
/** A one-line summary for the operator — the reason, not just the verdict. */
public String summary() {
if (error != null) {
return "config reload refused — " + error;
}
if (!applied) {
return "config reload refused — these keys cannot change under a running daemon: "
+ String.join(", ", coldKeys) + ". Restart bridged to apply them.";
}
if (!deferred.isEmpty()) {
return "config reloaded; these changes need a restart to take effect: "
+ String.join(", ", deferred);
}
return "config reloaded";
}
}
/**
* Re-read the file, validate it, and swap it in when nothing cold changed.
*
* <p>Never throws: a reload is a best-effort operation on a daemon that is already serving, and
* a bad edit must not take it down. Every failure path leaves the previous config live and is
* reported through the returned {@link Outcome}.
*/
public Outcome reload() {
if (path == null) {
return Outcome.failed("this config was built in code and has no file to reload from");
}
BridgedConfig old = current.get();
BridgedConfig fresh;
try {
fresh = BridgedConfig.load(path);
// The same gate startup runs. A config that would have refused to boot must not be able
// to slip in through a reload — that is how a daemon ends up in a state it could never
// have started in, which is the hardest kind to debug.
fresh.validateAuthExposure();
fresh.validateLeadTabPrefixes();
fresh.validateSubscriptionProfiles();
fresh.validateMembers();
} catch (RuntimeException e) {
String msg = e.getMessage() == null ? e.toString() : e.getMessage();
log.warn("config reload from {} refused, keeping the running config: {}", path, msg);
return Outcome.failed(msg);
}
List<String> cold = changedColdKeys(old, fresh);
if (!cold.isEmpty()) {
Outcome out = Outcome.refusedCold(cold);
log.warn(out.summary());
return out;
}
List<String> deferred = changedDeferredKeys(old, fresh);
current.set(fresh);
Outcome out = new Outcome(true, List.of(), deferred, null);
log.info(out.summary());
return out;
}
/** Cold keys whose value differs between the running config and the candidate. */
private static List<String> changedColdKeys(BridgedConfig old, BridgedConfig fresh) {
List<String> changed = new ArrayList<>();
if (!Objects.equals(old.bind(), fresh.bind())) {
changed.add("bind");
}
if (!Objects.equals(old.herdrSocket(), fresh.herdrSocket())) {
changed.add("herdrSocket");
}
if (!Objects.equals(old.broker(), fresh.broker())) {
changed.add("broker");
}
if (!Objects.equals(old.auth(), fresh.auth())) {
changed.add("auth");
}
// Kept in step with COLD_KEYS so the doc and the code cannot drift apart silently.
assert COLD_KEYS.containsAll(changed) : "a cold key was reported that COLD_KEYS omits";
return changed;
}
/** Changed keys that were accepted but whose effect waits for a restart. */
private static List<String> changedDeferredKeys(BridgedConfig old, BridgedConfig fresh) {
List<String> changed = new ArrayList<>();
if (!Objects.equals(old.lifecycle(), fresh.lifecycle())) {
changed.add("lifecycle");
}
if (!Objects.equals(old.leadHeartbeat(), fresh.leadHeartbeat())) {
changed.add("leadHeartbeat");
}
if (!Objects.equals(old.guard(), fresh.guard())) {
changed.add("guard");
}
if (!Objects.equals(old.worktreeRoot(), fresh.worktreeRoot())) {
changed.add("worktreeRoot");
}
if (!Objects.equals(old.spawnReadyTimeoutMs(), fresh.spawnReadyTimeoutMs())
|| !Objects.equals(old.spawnReadyPollMs(), fresh.spawnReadyPollMs())) {
changed.add("spawnReady*");
}
Map<String, BridgedConfig.Profile> before =
old.profiles() == null ? Map.of() : old.profiles();
Map<String, BridgedConfig.Profile> after =
fresh.profiles() == null ? Map.of() : fresh.profiles();
// Adding or removing a profile is deferred: a new backend needs its own launcher, and
// launchers are built once at startup.
if (!before.keySet().equals(after.keySet())) {
Set<String> diff = new LinkedHashSet<>(before.keySet());
diff.addAll(after.keySet());
diff.removeIf(p -> before.containsKey(p) && after.containsKey(p));
changed.add("profiles (added/removed: " + String.join(", ", diff) + ")");
}
// An EXISTING profile's launch settings are deferred too, and this is easy to get wrong:
// `HerdrPeerLauncher` takes `Map.copyOf(profiles)` at construction and `spawn` resolves the
// profile out of that snapshot, so a reloaded model/baseUrl/argv/env never reaches a launch.
// Only weight and maxLoad are genuinely hot, because placement reads them through the
// supplier on the composite rather than from the adapter's copy. Without this check a
// changed model would report "config reloaded" and silently do nothing — the worst outcome
// a reload can produce, because the operator has no reason to doubt it.
List<String> relaunch = new ArrayList<>();
before.forEach((name, was) -> {
BridgedConfig.Profile now = after.get(name);
if (now != null && !sameLaunchSettings(was, now)) {
relaunch.add(name);
}
});
if (!relaunch.isEmpty()) {
changed.add("profiles." + String.join("/", relaunch) + " launch settings "
+ "(model, baseUrl, argv, env, …) — the launcher holds a startup snapshot");
}
return changed;
}
/**
* Whether two versions of a profile would launch a peer identically. Compares every component
* the launcher reads at spawn; {@code weight} and {@code maxLoad} are excluded because those are
* read live by the placement policy and really do take effect on the next spawn.
*/
private static boolean sameLaunchSettings(BridgedConfig.Profile a, BridgedConfig.Profile b) {
return Objects.equals(a.baseUrl(), b.baseUrl())
&& Objects.equals(a.model(), b.model())
&& Objects.equals(a.configDir(), b.configDir())
&& Objects.equals(a.tokenEnv(), b.tokenEnv())
&& Objects.equals(a.argv(), b.argv())
&& Objects.equals(a.placement(), b.placement())
&& Objects.equals(a.workspace(), b.workspace())
&& Objects.equals(a.tabLabel(), b.tabLabel())
&& Objects.equals(a.mcpUrl(), b.mcpUrl())
&& Objects.equals(a.cwd(), b.cwd())
&& Objects.equals(a.parityOverlay(), b.parityOverlay())
&& Objects.equals(a.gitTokenEnv(), b.gitTokenEnv())
&& Objects.equals(a.gitHostEnv(), b.gitHostEnv())
&& Objects.equals(a.kind(), b.kind())
&& Objects.equals(a.env(), b.env())
&& Objects.equals(a.subscription(), b.subscription());
}
}
@@ -0,0 +1,96 @@
package dev.ltms.bridged.config;
import org.slf4j.Logger;
import org.slf4j.LoggerFactory;
import java.io.IOException;
import java.nio.file.Files;
import java.nio.file.Path;
import java.util.concurrent.Executors;
import java.util.concurrent.ScheduledExecutorService;
import java.util.concurrent.TimeUnit;
/**
* Polls {@code bridged.yaml}'s modified time and asks {@link ConfigRef} to reload when it moves
* (CB-559). Opt-in through {@code configReload.enabled}.
*
* <p><strong>Why polling and not a filesystem watch.</strong> {@code WatchService} on macOS has no
* native backend — it falls back to polling internally anyway, at an interval this code does not
* control — and editors save config files in ways that produce a different event mix per editor
* (write-in-place, write-and-rename, write-temp-and-swap). A modified-time check treats all of them
* the same and is a single {@code stat} per tick, which at a ten-second cadence costs nothing worth
* measuring.
*
* <p><strong>A missing or unreadable file is not a reason to act.</strong> Many editors briefly
* unlink the file during a save. Reloading on "it vanished" would mean reloading from a file that no
* longer exists; reporting an error every tick would bury the log. So an unreadable file is skipped
* silently and the next tick tries again — the running config stays live, which is the correct
* outcome either way.
*/
public final class ConfigWatcher {
private static final Logger log = LoggerFactory.getLogger(ConfigWatcher.class);
private final ConfigRef ref;
private final long intervalSeconds;
private final ScheduledExecutorService scheduler;
private volatile long lastSeenMillis;
public ConfigWatcher(ConfigRef ref, long intervalSeconds) {
this.ref = ref;
this.intervalSeconds = intervalSeconds;
this.lastSeenMillis = modifiedMillis(ref.path());
this.scheduler = Executors.newSingleThreadScheduledExecutor(r -> {
Thread t = new Thread(r, "config-watcher");
// A daemon thread: an operator's config watch must never be the reason the JVM refuses
// to exit after everything else has shut down.
t.setDaemon(true);
return t;
});
}
/** Begin watching. A ref with no file (a fixed one) is a no-op rather than an error. */
public void start() {
if (ref.path() == null) {
log.debug("config watch not started — this config has no file behind it");
return;
}
scheduler.scheduleWithFixedDelay(this::tick, intervalSeconds, intervalSeconds,
TimeUnit.SECONDS);
log.info("config watch: {} re-read when it changes (every {}s)", ref.path(), intervalSeconds);
}
/** One poll. Never throws — an exception here would silently cancel the schedule. */
void tick() {
try {
long now = modifiedMillis(ref.path());
if (now == 0 || now == lastSeenMillis) {
return;
}
// Stamp BEFORE reloading. A file whose reload is refused (a bad edit, or a cold key)
// must not be retried every tick — that would log the same refusal forever. The next
// save moves the timestamp again and earns a fresh attempt.
lastSeenMillis = now;
ref.reload();
} catch (RuntimeException e) {
log.warn("config watch tick failed, still watching: {}", e.getMessage());
}
}
private static long modifiedMillis(Path path) {
if (path == null) {
return 0;
}
try {
return Files.getLastModifiedTime(path).toMillis();
} catch (IOException e) {
return 0; // mid-save, or gone: say nothing and try again next tick
}
}
/** Stop polling. Called from the daemon's ordered shutdown hook, alongside the other loops. */
public void stop() {
scheduler.shutdownNow();
}
}
@@ -6,93 +6,120 @@ import java.util.ArrayList;
import java.util.LinkedHashMap;
import java.util.List;
import java.util.Map;
import java.util.concurrent.ConcurrentHashMap;
/**
* Domain layer over herdr's native {@code agent.*} namespace — the worker south side.
* Chosen in the CB-102 spike over the pane + {@code send_text} fallback because
* {@code agent.start} takes a first-class {@code env} map (clean, guard-checked
* subscription injection) and herdr tracks each worker's Claude session UUID itself.
* Chosen in the CB-102 spike over the pane + {@code send_text} fallback because herdr
* tracks each worker's Claude session UUID itself.
*
* <p>Ported to herdr protocol 19 (herdr 0.8.0, CB-521): {@code agent.start} now starts a
* <em>supported</em> agent ({@code kind}) into an <em>existing</em> pane, so the worker's
* {@code env}/{@code cwd} move to pane creation ({@code tab.create}/{@code pane.split} — see
* {@link WorkspaceControl}), and {@code agent.send} is replaced by {@code agent.prompt}
* (which submits in one call) plus {@code agent.send_keys} for the raw Enter nudge.
*
* <p>Every method is one herdr call through the injected {@link HerdrClient}, so this
* layer is unit-testable with a fake and contract-tested against a live daemon.
*/
public final class AgentControl {
/**
* The keystroke that submits a prompt in the Claude Code TUI: a carriage return (Enter).
* It must be delivered as its <em>own</em> {@code agent.send} call — herdr delivers a message's
* text as a bracketed paste, and a {@code "\r"} appended to that same text is swallowed as
* literal newline content, not a submit. Sent as a separate keystroke event it lands outside
* the paste and submits. (A bare {@code "\n"} inserts a newline either way.) Verified live
* against Claude Code v2.1.210: an injected task stayed unsubmitted with {@code "text\r"} in
* one call, and submitted the instant a standalone {@code "\r"} was sent.
*/
static final String SUBMIT_KEY = "\r";
private final HerdrClient herdr;
/**
* Protocol 19 dropped {@code terminal_id} as an {@code agent.*} target — herdr now resolves
* targets by pane id or agent name only, while the bridge keys every session on the terminal.
* This caches the terminal→pane mapping (stable for a worker's lifetime) so callers keep
* addressing agents by terminal; entries are invalidated on {@code agent_not_found}.
*/
private final Map<String, String> paneByTerminal = new ConcurrentHashMap<>();
public AgentControl(HerdrClient herdr) {
this.herdr = herdr;
}
/** One agent-targeted call, translating a terminal id to its pane id (retrying once fresh). */
private JsonNode agentCall(String method, String target, Map<String, Object> extra) {
String resolved = resolveTarget(target);
try {
return herdr.call(method, withTarget(resolved, extra));
} catch (HerdrException e) {
if (!"agent_not_found".equals(e.code()) || resolved.equals(target)) throw e;
paneByTerminal.remove(target); // the cached pane went away — re-resolve once
String fresh = resolveTarget(target);
if (fresh.equals(resolved)) throw e;
return herdr.call(method, withTarget(fresh, extra));
}
}
private static Map<String, Object> withTarget(String target, Map<String, Object> extra) {
Map<String, Object> m = new LinkedHashMap<>();
m.put("target", target);
m.putAll(extra);
return m;
}
/** The pane id behind a terminal-id target, or the target verbatim for pane ids / names. */
private String resolveTarget(String target) {
if (target == null || !target.startsWith("term_")) {
return target;
}
String cached = paneByTerminal.get(target);
if (cached != null) {
return cached;
}
for (JsonNode a : herdr.call("agent.list").path("agents")) {
if (target.equals(a.path("terminal_id").asText(null))) {
String pane = a.path("pane_id").asText(null);
if (pane != null) {
paneByTerminal.put(target, pane);
return pane;
}
}
}
return target; // unknown terminal — let herdr report it against the original target
}
/**
* Spawn an agent. {@code env} is applied to the process environment verbatim — this
* is where a worker's {@code ANTHROPIC_BASE_URL} lives, and the ONLY place it should.
* Start an agent into {@code paneId}, which must be sitting at its interactive shell prompt —
* the seed pane of a freshly-created worker tab, or a fresh split. The pane's shell already
* carries the worker's env ({@code ANTHROPIC_BASE_URL}, token, …) and cwd from pane creation;
* herdr resolves the executable from {@code kind} and waits (its default timeout) until the
* agent is detected and ready for input.
*
* @param name label/kind for herdr status detection (e.g. {@code "claude"})
* @param argv launch command, e.g. {@code ["claude"]}
* @param env process environment additions ({@code ANTHROPIC_BASE_URL}, token, …)
* @param name unique label for this agent ({@code <kind>-<profile>-<nonce>-<seq>})
* @param kind supported agent kind and canonical executable, e.g. {@code "claude"},
* {@code "opencode"}
* @param args extra arguments after the executable, e.g. {@code --mcp-config …}
* @param paneId the pane to start the agent in
*/
public Agent start(String name, List<String> argv, Map<String, String> env) {
return start(name, argv, env, null);
}
/** Spawn an agent into {@code tabId} at herdr's default cwd. */
public Agent start(String name, List<String> argv, Map<String, String> env, String tabId) {
return start(name, argv, env, tabId, null);
}
/**
* Spawn an agent. With a non-null {@code tabId} the worker lands in that tab (the placement
* policy's dedicated worker tab); with {@code null} herdr splits the currently-focused tab
* (legacy pane placement). A non-blank {@code cwd} sets the worker process's working directory —
* {@code agent.start} honours {@code cwd} directly (an agent pane does <em>not</em> inherit the
* tab's or workspace's cwd, so this is the only way to root a worker in the primary's directory;
* CB-112).
*/
public Agent start(String name, List<String> argv, Map<String, String> env, String tabId, String cwd) {
public Agent start(String name, String kind, List<String> args, String paneId) {
Map<String, Object> params = new LinkedHashMap<>();
params.put("name", name);
params.put("argv", argv);
params.put("env", env);
if (tabId != null) {
params.put("tab_id", tabId);
}
if (cwd != null && !cwd.isBlank()) {
params.put("cwd", cwd);
}
params.put("kind", kind);
params.put("pane_id", paneId);
params.put("args", args);
JsonNode result = herdr.call("agent.start", params);
return Agent.from(result.get("agent"));
}
/**
* Deliver {@code text} to an agent as its next prompt <em>and submit it</em> — two keystroke
* events: the message (a bracketed paste, so any embedded newlines are preserved verbatim),
* then a standalone {@link #SUBMIT_KEY} (Enter) that actually submits it. Without the second
* event the text just sits in the worker's input box, never processed (see {@link #SUBMIT_KEY}).
* Deliver {@code text} to an agent as its next prompt <em>and submit it</em> — herdr's
* {@code agent.prompt} pastes the text (embedded newlines preserved verbatim) and submits it
* in the same call, replacing the pre-protocol-19 two-event {@code agent.send} dance.
*/
public void send(String target, String text) {
herdr.call("agent.send", Map.of("target", target, "text", text));
herdr.call("agent.send", Map.of("target", target, "text", SUBMIT_KEY));
agentCall("agent.prompt", target, Map.of("text", text));
}
/**
* Re-send the submit keystroke (Enter) to {@code target}. The Enter that accompanies a delivery
* can race the paste — especially right as the worker's TUI becomes interactive — leaving the
* text unsubmitted; the injector nudges it with this until the worker actually picks up (CB-113).
* Re-send the submit keystroke (Enter) to {@code target}. The submit that accompanies a
* delivery can race the paste — especially right as the worker's TUI becomes interactive —
* leaving the text unsubmitted; the injector nudges it with this until the worker actually
* picks up (CB-113).
*/
public void submit(String target) {
herdr.call("agent.send", Map.of("target", target, "text", SUBMIT_KEY));
agentCall("agent.send_keys", target, Map.of("keys", List.of("enter")));
}
/**
@@ -101,13 +128,13 @@ public final class AgentControl {
* @param source one of {@code visible|recent|recent_unwrapped|detection}
*/
public String read(String target, String source) {
JsonNode result = herdr.call("agent.read", Map.of("target", target, "source", source));
JsonNode result = agentCall("agent.read", target, Map.of("source", source));
return result.path("read").path("text").asText("");
}
/** Current agent record (status, session UUID, pane). */
public Agent get(String target) {
return Agent.from(herdr.call("agent.get", Map.of("target", target)).get("agent"));
return Agent.from(agentCall("agent.get", target, Map.of()).get("agent"));
}
/** Just the lifecycle status — what the status-gated injector checks before send. */
@@ -0,0 +1,175 @@
package dev.ltms.bridged.herdr;
import com.fasterxml.jackson.databind.JsonNode;
import org.slf4j.Logger;
import org.slf4j.LoggerFactory;
import java.util.Collections;
import java.util.LinkedHashMap;
import java.util.Map;
import java.util.Set;
import java.util.function.LongSupplier;
import java.util.function.Supplier;
/**
* Discovers which panes host a lead by scanning herdr for tabs the operator labelled by convention
* (CB-531), and hands {@link dev.ltms.bridged.auth.CallerResolver} the resulting
* {@code terminal_id → lead name} map.
*
* <p><strong>Why scan at all.</strong> A lead is never spawned — a human opens a tab and starts an
* agent in it — so the daemon cannot learn a lead's {@code terminal_id} at creation time the way it
* does a worker's. CB-530 solved that by having the operator paste each id into {@code leaders:},
* which works but costs a config edit and a daemon restart per lead, and the id is only obtainable
* by first starting the session and asking it. Scanning closes that loop: label the tab, and the
* pane is recognised on the next resolve.
*
* <p><strong>Direction of trust.</strong> The label names the lead; it never <em>grants</em>
* anything a pane could take for itself. Three properties keep that honest:
* <ol>
* <li>Worker spaces are excluded wholesale ({@code excludedWorkspaceLabels}), so a worker cannot
* become a lead by being placed — as a split, say — inside a matching tab.</li>
* <li>A worker cannot rename a tab: {@code tab.rename} is reachable only through
* {@link WorkspaceControl}, which no {@code bridge_*} tool exposes. The label is writable by
* the human at the terminal and by nobody the bridge is defending against.</li>
* <li>The label is a <em>name</em>, not a capability. What a pane may do is decided by
* {@code Authz} against the role {@code CallerResolver} returns; a tab that calls itself a
* lead still cannot act as one unless the daemon's own registry agrees.</li>
* </ol>
*
* <p><strong>CB-558 — bridged now writes lead labels too.</strong> This class used to be able to say
* that bridged never renames a lead tab, so the label was always the human's own writing and there
* was no round-trip from the daemon's rename back into its next decision.
* {@code dev.ltms.bridged.lead.LeadLauncher} ends that: an auto-launched lead is labelled by the
* daemon and found again by this scan. The trust direction above is unaffected — bridged writing a
* name for a lead it just started is not a pane promoting itself — but <em>staleness</em> becomes
* real: a label left behind by a session that has since died would read as a live lead forever.
* This scanner does not solve that (its job is naming, and a stale name costs nothing here); the
* launcher does, by requiring a running agent in the tab before it counts the lead as live. If you
* ever make a decision that <em>removes</em> something based on this map, add the same check.
* The remaining hazard is an <em>operator</em> one — a worker {@code tabLabel} template that
* happens to start with the same prefix would promote the whole fleet — and that is refused at
* startup by {@code BridgedConfig.validateLeadScan} rather than documented here.
*
* <p><strong>Caching.</strong> {@link #get()} is on the request path (every resolve), so the scan
* is TTL-cached and a stale-but-valid map is preferred to a herdr round-trip. A failed scan keeps
* the previous answer instead of emptying it — a herdr hiccup must not silently demote a live lead
* mid-session.
*/
public final class LeadTabScanner implements Supplier<Map<String, String>> {
private static final Logger log = LoggerFactory.getLogger(LeadTabScanner.class);
private final HerdrClient herdr;
private final String tabPrefix;
private final Set<String> excludedWorkspaceLabels;
private final Map<String, String> configuredLeads;
private final long ttlNanos;
private final LongSupplier clock;
private Map<String, String> cached;
private long scannedAtNanos;
private boolean everScanned;
/**
* @param herdr the herdr client to query ({@code workspace.list},
* {@code tab.list}, {@code pane.list} — all read-only)
* @param tabPrefix a tab whose label starts with this (case-insensitively) hosts a
* lead; the rest of the label, trimmed, is the lead's name
* @param excludedWorkspaceLabels workspaces never scanned — the configured worker spaces
* @param configuredLeads the static {@code leaders:}/{@code primary:} registry, merged
* over every scan result. Explicit config outranks the
* convention, and survives a scan that cannot run at all
* @param ttlNanos how long a scan result is reused before the next one
* @param clock nanosecond time source ({@code System::nanoTime} in production)
*/
public LeadTabScanner(HerdrClient herdr, String tabPrefix, Set<String> excludedWorkspaceLabels,
Map<String, String> configuredLeads, long ttlNanos, LongSupplier clock) {
this.herdr = herdr;
this.tabPrefix = tabPrefix == null || tabPrefix.isBlank() ? "lead:" : tabPrefix.strip();
this.excludedWorkspaceLabels = excludedWorkspaceLabels == null
? Set.of() : Set.copyOf(excludedWorkspaceLabels);
this.configuredLeads = configuredLeads == null ? Map.of() : Map.copyOf(configuredLeads);
this.ttlNanos = ttlNanos;
this.clock = clock;
this.cached = this.configuredLeads;
}
/**
* The current {@code terminal_id → lead name} map, rescanning when the cache has expired.
*
* <p>Synchronized so a burst of concurrent calls produces one scan rather than one each; a scan
* is a handful of RPCs over a Unix socket and is rate-limited to one per TTL.
*/
@Override
public synchronized Map<String, String> get() {
long now = clock.getAsLong();
if (everScanned && now - scannedAtNanos < ttlNanos) {
return cached;
}
// Stamp before scanning, not after: a herdr that is down must cost one attempt per TTL, not
// one per request.
scannedAtNanos = now;
everScanned = true;
try {
Map<String, String> fresh = scan();
if (!fresh.equals(cached)) {
log.info("lead panes: {}", fresh);
}
cached = fresh;
} catch (HerdrException e) {
log.warn("lead-tab scan failed, keeping the {} lead(s) already known: {}",
cached.size(), e.getMessage());
}
return cached;
}
/** One full pass: labelled tabs → their panes → those panes' terminals. */
private Map<String, String> scan() {
Map<String, String> nameByTab = new LinkedHashMap<>();
for (JsonNode w : herdr.call("workspace.list").path("workspaces")) {
Workspace ws = Workspace.from(w);
if (ws.workspaceId() == null || excludedWorkspaceLabels.contains(ws.label())) {
continue;
}
for (JsonNode t : herdr.call("tab.list", Map.of("workspace_id", ws.workspaceId())).path("tabs")) {
Tab tab = Tab.from(t);
String name = leadNameOf(tab.label());
if (name != null && tab.tabId() != null) {
nameByTab.put(tab.tabId(), name);
}
}
}
Map<String, String> byTerminal = new LinkedHashMap<>();
if (!nameByTab.isEmpty()) {
// One pane.list for every tab: panes carry tab_id, so the join is local.
for (JsonNode p : herdr.call("pane.list", Map.of()).path("panes")) {
String name = nameByTab.get(p.path("tab_id").asText(null));
String terminal = p.path("terminal_id").asText(null);
if (name != null && terminal != null && !terminal.isBlank()) {
byTerminal.put(terminal, name);
}
}
}
byTerminal.putAll(configuredLeads); // an explicit pin outranks a label
return Collections.unmodifiableMap(byTerminal);
}
/**
* The lead name a tab label declares, or {@code null} if it declares none.
*
* <p>{@code "lead: opus-5.0"} → {@code "opus-5.0"}. A bare {@code "lead:"} names nobody and is
* rejected: an unnamed lead would resolve as {@code PRIMARY} with nothing to attribute it to.
*/
private String leadNameOf(String label) {
if (label == null) {
return null;
}
String l = label.strip();
if (!l.regionMatches(true, 0, tabPrefix, 0, tabPrefix.length())) {
return null;
}
String name = l.substring(tabPrefix.length()).strip();
return name.isEmpty() ? null : name;
}
}
@@ -23,9 +23,9 @@ public record Tab(String tabId, String workspaceId, String label, int paneCount)
}
/**
* A freshly-created tab together with the placeholder shell pane herdr seeds it with.
* The caller starts the worker into {@link #tab()} then closes {@link #rootPaneId()} so
* only the worker pane remains.
* A freshly-created tab together with the shell pane herdr seeds it with. Under protocol 19
* the caller starts the worker <em>into</em> {@link #rootPaneId()} — the seed pane's shell
* carries the worker's cwd and env from {@code tab.create}, and becomes the worker pane.
*/
public record Created(Tab tab, String rootPaneId) {
/** Project a {@code tab_created} result ({@code {tab, root_pane}}). */
@@ -5,6 +5,7 @@ import org.slf4j.Logger;
import org.slf4j.LoggerFactory;
import java.util.ArrayList;
import java.util.LinkedHashMap;
import java.util.List;
import java.util.Map;
import java.util.Optional;
@@ -40,6 +41,16 @@ public final class WorkspaceControl {
return out;
}
/** Every tab in {@code workspaceId}, in herdr's order. */
public List<Tab> listTabs(String workspaceId) {
JsonNode result = herdr.call("tab.list", Map.of("workspace_id", workspaceId));
List<Tab> out = new ArrayList<>();
for (JsonNode t : result.path("tabs")) {
out.add(Tab.from(t));
}
return out;
}
/** The first workspace with this exact label, if any. */
public Optional<Workspace> findByLabel(String label) {
return listWorkspaces().stream()
@@ -65,13 +76,37 @@ public final class WorkspaceControl {
}
/**
* A brand-new tab in {@code workspaceId} plus the placeholder shell pane herdr seeds it with.
* Start the worker into the tab, then {@code pane.close} the root pane so the tab holds only the
* worker. (The worker's own cwd is set on {@code agent.start}, not here — an {@code agent.start}
* pane does not inherit the tab's cwd; see {@code AgentControl.start}.)
* A brand-new tab in {@code workspaceId} plus the shell pane herdr seeds it with. Under
* protocol 19 that seed pane is where the worker <em>starts</em>: its shell carries
* {@code cwd} and {@code env} (the worker's {@code ANTHROPIC_BASE_URL} — this is the
* subscription-injection seam now), and {@code agent.start} launches the agent into it.
*/
public Tab.Created createTab(String workspaceId) {
return Tab.Created.from(herdr.call("tab.create", Map.of("workspace_id", workspaceId)));
public Tab.Created createTab(String workspaceId, String cwd, Map<String, String> env) {
Map<String, Object> params = new LinkedHashMap<>();
params.put("workspace_id", workspaceId);
if (cwd != null && !cwd.isBlank()) {
params.put("cwd", cwd);
}
if (env != null && !env.isEmpty()) {
params.put("env", env);
}
return Tab.Created.from(herdr.call("tab.create", params));
}
/**
* Split the currently-focused tab and return the new pane's id — the legacy pane placement's
* seed pane, carrying {@code cwd} and {@code env} exactly as {@link #createTab}'s does.
*/
public String splitPane(String cwd, Map<String, String> env) {
Map<String, Object> params = new LinkedHashMap<>();
params.put("direction", "right");
if (cwd != null && !cwd.isBlank()) {
params.put("cwd", cwd);
}
if (env != null && !env.isEmpty()) {
params.put("env", env);
}
return herdr.call("pane.split", params).path("pane").path("pane_id").asText(null);
}
/** Give a worker's tab a human label in the tab bar. */
@@ -55,6 +55,9 @@ public final class CompletionResolver implements TurnListener {
/** Cap the scraped tail so a long transcript can't return an unbounded blob. */
static final int MAX_SCRAPE_CHARS = 4000;
private static final String CLIPPED_PANE_TAIL_MARKER =
"[Pane tail clipped: member did not call bridge_reply.]";
private final AgentControl agents;
private final Rendezvous rendezvous;
@@ -122,6 +125,15 @@ public final class CompletionResolver implements TurnListener {
Thread.ofVirtual().name("completion-" + target).start(() -> resolve(target, turn));
}
/**
* Resolve the completed turn before adapter housekeeping can erase its rendered output. This is
* intentionally synchronous and used only when a post-turn context reset is enabled; the normal
* path remains off-loaded so polling is not blocked by a scrape.
*/
public void resolveBeforePostAction(String target) {
resolve(target, inFlight.get(target));
}
@Override
public void onTurnFailed(String target) {
InFlight turn = inFlight.get(target);
@@ -138,9 +150,14 @@ public final class CompletionResolver implements TurnListener {
return;
}
String tail;
int originalLength = 0;
boolean clipped = false;
boolean scrapeFailed = false;
try {
tail = clip(lastAssistantBlock(agents.read(target, SCRAPE_SOURCE)));
String assistantBlock = lastAssistantBlock(agents.read(target, SCRAPE_SOURCE));
originalLength = assistantBlock.strip().length();
clipped = originalLength > MAX_SCRAPE_CHARS;
tail = clip(assistantBlock);
} catch (RuntimeException e) {
// The worker finished but we couldn't read its screen — still resolve the send so the
// caller unblocks; an empty tail beats hanging until the caller's timeout.
@@ -160,8 +177,14 @@ public final class CompletionResolver implements TurnListener {
target);
return; // keep the in-flight record: a later genuine completion still needs it
}
if (rendezvous.resolveCompletion(waiter, tail)) {
String completion = clipped ? tail + "\n" + CLIPPED_PANE_TAIL_MARKER : tail;
if (rendezvous.resolveCompletion(waiter, completion)) {
inFlight.remove(target, turn);
if (clipped) {
log.warn("completion scrape for {} clipped from {} chars to the {} char cap; "
+ "member did not call bridge_reply, so the pane tail is partial",
target, originalLength, MAX_SCRAPE_CHARS);
}
log.debug("resolved send to {} via turn-completion fallback ({} chars scraped)",
target, tail.length());
}
@@ -78,6 +78,15 @@ public final class Injector {
*/
private static final int READINESS_GRACE_POLLS = 240;
/**
* The single source for the injector poll cadence — how often the {@link StatusPoller} drives
* {@link #onStatus} at. {@code Bridged} passes this to every {@link StatusPoller} it constructs,
* and this class reads it to state the readiness grace in seconds on the CB-562 expiry log
* instead of hardcoding "60s". One constant, so a cadence change cannot silently desync a log
* that claims a grace duration.
*/
public static final long POLL_INTERVAL_MILLIS = 250;
private final AgentControl agents;
private final TurnListener turnListener;
private final Predicate<String> ready; // CB-113: a target is deliverable only when available
@@ -132,6 +141,10 @@ public final class Injector {
boolean turnObserved; // saw a real `working` sample since that delivery (turn ran)
int unknownSinceTurn; // consecutive `unknown` samples while a delegation is outstanding (CB-109)
int notReadySincePoll; // consecutive injectable samples a queued message waited on the readiness gate (CB-114)
boolean postTurnPending; // completion observed; adapter housekeeping has not started yet
boolean awaitingPostTurnPickup;
boolean postTurnObserved;
int injectableSincePostTurnPickup;
synchronized void add(Pending p) {
queue.add(p);
@@ -173,9 +186,15 @@ public final class Injector {
boolean turnCompleted = false;
boolean turnFailed = false;
boolean resubmit = false;
boolean startPostTurn = false;
List<Pending> notReady = null; // queued messages failed because the worker never became ready
synchronized (t) {
if (status == AgentStatus.WORKING) {
if (t.awaitingPostTurnPickup) {
t.awaitingPostTurnPickup = false;
t.injectableSincePostTurnPickup = 0;
t.postTurnObserved = true;
}
// Definitive pickup: the worker is busy on our last message, and (if a delivery is
// outstanding) a real turn is now confirmed to be running.
t.awaitingPickup = false;
@@ -185,6 +204,16 @@ public final class Injector {
if (t.awaitingCompletion) t.turnObserved = true;
} else if (status.injectable()) { // IDLE or BLOCKED
t.unknownSinceTurn = 0;
if (t.awaitingPostTurnPickup) {
if (++t.injectableSincePostTurnPickup >= PICKUP_GRACE_POLLS) {
t.awaitingPostTurnPickup = false;
t.injectableSincePostTurnPickup = 0;
} else {
resubmit = true;
}
} else if (t.postTurnObserved) {
t.postTurnObserved = false;
}
if (t.awaitingPickup) {
if (++t.injectableSincePickup >= PICKUP_GRACE_POLLS) {
// Pickup edge was never sampled (turn faster than the poll, or status lag).
@@ -209,11 +238,16 @@ public final class Injector {
t.awaitingCompletion = false;
t.turnObserved = false;
turnCompleted = true;
if (turnListener.hasPostTurnAction(target)) {
t.postTurnPending = true;
startPostTurn = true;
}
}
// Deliver the next queued message only once the prior turn is fully settled, so a
// completion is never confused with the pickup of the following message — and only
// once the worker is available (CB-113), so we never paste into its boot window.
if (!t.awaitingCompletion) {
if (!t.awaitingCompletion && !t.postTurnPending
&& !t.awaitingPostTurnPickup && !t.postTurnObserved) {
Pending p = t.queue.peek();
if (p != null && ready.test(target)) {
t.notReadySincePoll = 0;
@@ -239,6 +273,11 @@ public final class Injector {
// fail every queued message and release the target (CB-114) instead of
// polling it indefinitely with the caller's future never completing.
notReady = new ArrayList<>(t.queue);
log.warn("readiness grace for {} expired after {} polls ({}s): target never "
+ "became deliverable, so failing {} queued message(s) that never "
+ "reached its pane",
target, READINESS_GRACE_POLLS,
READINESS_GRACE_POLLS * POLL_INTERVAL_MILLIS / 1000, notReady.size());
t.queue.clear();
t.notReadySincePoll = 0;
}
@@ -263,7 +302,8 @@ public final class Injector {
// Reclaim the entry once the worker is fully quiescent (nothing queued, no pickup or
// completion awaited), so the map cannot grow without bound across short-lived workers.
if (t.queue.isEmpty() && !t.awaitingPickup && !t.awaitingCompletion) {
if (t.queue.isEmpty() && !t.awaitingPickup && !t.awaitingCompletion
&& !t.postTurnPending && !t.awaitingPostTurnPickup && !t.postTurnObserved) {
targets.remove(target, t);
}
}
@@ -290,7 +330,21 @@ public final class Injector {
turnListener.onTurnFailed(target);
}
if (turnCompleted) {
turnListener.onTurnComplete(target);
if (startPostTurn) {
boolean started = turnListener.onTurnCompleteWithPostAction(target);
synchronized (t) {
t.postTurnPending = false;
if (started) {
t.awaitingPostTurnPickup = true;
t.injectableSincePostTurnPickup = 0;
}
if (t.queue.isEmpty() && !t.awaitingPostTurnPickup) {
targets.remove(target, t);
}
}
} else {
turnListener.onTurnComplete(target);
}
}
if (turnFailed) {
turnListener.onTurnFailed(target);
@@ -317,7 +371,8 @@ public final class Injector {
.filter(e -> {
synchronized (e.getValue()) {
Target t = e.getValue();
return !t.queue.isEmpty() || t.awaitingPickup || t.awaitingCompletion;
return !t.queue.isEmpty() || t.awaitingPickup || t.awaitingCompletion
|| t.postTurnPending || t.awaitingPostTurnPickup || t.postTurnObserved;
}
})
.map(java.util.Map.Entry::getKey)
@@ -343,6 +398,9 @@ public final class Injector {
hadDeliveredTurn = t.awaitingCompletion;
t.awaitingCompletion = false;
t.awaitingPickup = false;
t.postTurnPending = false;
t.awaitingPostTurnPickup = false;
t.postTurnObserved = false;
}
forget.accept(target); // the worker is gone — clear its readiness/presence too (CB-114)
for (Pending p : pending) {
@@ -15,7 +15,7 @@ import java.util.Set;
* mounts the bridge MCP is never marked present — its sends stay queued until they time out, which is
* correct (it could not have replied anyway).
*/
public class WorkerPresence {
public class MemberPresence {
private final Set<String> present = ConcurrentHashMap.newKeySet();
@@ -13,6 +13,25 @@ public interface TurnListener {
/** A worker's delegated turn finished (worker returned to idle after visibly working). */
void onTurnComplete(String target);
/**
* Whether completion must pause ordinary delivery while adapter-specific housekeeping starts.
* This is queried before the injector considers the next queued message, closing the same-tick
* ordering gap. Implementations must not mutate state here.
*/
default boolean hasPostTurnAction(String target) {
return false;
}
/**
* Complete the delegated turn and start its non-turn housekeeping operation.
*
* @return true when the injector must observe that operation settle before the next delivery
*/
default boolean onTurnCompleteWithPostAction(String target) {
onTurnComplete(target);
return false;
}
/**
* A worker that visibly ran a delegated turn then wedged in a non-idle, non-working state
* (CB-109) — e.g. an error screen herdr classifies as {@code unknown} — so no
@@ -0,0 +1,304 @@
package dev.ltms.bridged.lead;
import dev.ltms.bridged.config.BridgedConfig;
import dev.ltms.bridged.herdr.Agent;
import dev.ltms.bridged.herdr.AgentControl;
import dev.ltms.bridged.herdr.HerdrException;
import dev.ltms.bridged.herdr.Tab;
import dev.ltms.bridged.herdr.Workspace;
import dev.ltms.bridged.herdr.WorkspaceControl;
import org.slf4j.Logger;
import org.slf4j.LoggerFactory;
import java.util.ArrayList;
import java.util.LinkedHashMap;
import java.util.List;
import java.util.Map;
import java.util.Objects;
import java.util.Set;
import java.util.stream.Collectors;
/**
* Starts the leads {@code fleet.leaders:} declares, when none is already running (CB-558).
*
* <p><strong>A lead is not a member, and this class exists to keep it that way.</strong> Every other
* spawn path in the daemon goes through {@code HerdrPeerLauncher}, which does three things a lead
* must never receive:
* <ol>
* <li>it appends the <em>reply charter</em> — "you are an off-subscription worker … end every turn
* with {@code bridge_reply}". A lead is the orchestrator; telling it that it is a worker is
* exactly backwards.</li>
* <li>it registers the session with {@code SessionManager}, which subjects it to the idle reaper,
* the context cap and the shutdown drain. An idle lead is the normal state of a lead, so the
* reaper would kill the orchestrator for doing its job.</li>
* <li>it can move a peer off the subscription via {@code ANTHROPIC_BASE_URL}. A lead stays on the
* operator's subscription, always.</li>
* </ol>
* So this launcher talks to {@link AgentControl}/{@link WorkspaceControl} directly. The duplication
* with the member launchers (argv and env assembly) is deliberate and is the cheaper half of the
* trade: entangling the worker path with a not-a-worker case is how the three rules above get
* broken later, quietly.
*
* <p><strong>Liveness, and the label round-trip.</strong> {@code LeadTabScanner} used to be able to
* promise that bridged never writes a lead label. That is no longer true — an auto-launched lead is
* labelled by this class, and the scanner reads that label back. The risk this opens is not
* privilege escalation (the tab label never granted anything a pane could take for itself; see that
* class's javadoc), but <em>staleness</em>: a label left behind by a crashed session would otherwise
* read as a live lead forever, and the lead would never be relaunched. So a lead counts as live only
* when herdr also reports a <em>running agent</em> in that tab — see {@link #liveLeads}. A labelled
* tab with no agent in it is not a lead.
*/
public final class LeadLauncher {
private static final Logger log = LoggerFactory.getLogger(LeadLauncher.class);
private final AgentControl agents;
private final WorkspaceControl spaces;
private final BridgedConfig cfg;
/**
* @param agents herdr agent control (start, list)
* @param spaces workspace / tab control (ensure, create, label, list)
* @param cfg the loaded config — {@code fleet.leaders}, {@code profiles} and the lead pins
*/
public LeadLauncher(AgentControl agents, WorkspaceControl spaces, BridgedConfig cfg) {
this.agents = agents;
this.spaces = spaces;
this.cfg = cfg;
}
/**
* Bring every declared lead up to its {@code instances} count, and return how many were started.
*
* <p>Never throws: a daemon that cannot start a lead must still serve. herdr being unreachable,
* a profile that does not exist, a failed {@code agent.start} — each is logged and skipped, and
* the remaining leads are still attempted.
*/
public int ensureLeads() {
Map<String, BridgedConfig.Leader> leaders = cfg.fleet().leaders();
if (leaders.isEmpty()) {
return 0;
}
Map<String, Integer> live;
try {
live = liveLeads(leaders);
} catch (HerdrException e) {
// Counting is the whole safety mechanism against double-spawning. If we cannot count, we
// must not guess — spawning a second orchestrator is worse than starting none.
log.warn("lead auto-launch skipped — cannot tell which leads are live: {}", e.getMessage());
return 0;
}
int started = 0;
for (Map.Entry<String, BridgedConfig.Leader> e : leaders.entrySet()) {
String name = e.getKey();
BridgedConfig.Leader lead = e.getValue();
int running = live.getOrDefault(name, 0);
int wanted = lead.instances();
if (running >= wanted) {
log.info("lead '{}': {} live, {} wanted — nothing to start", name, running, wanted);
continue;
}
if (!lead.isCreatable()) {
// A lead with a `terminal:` pin and no `profile:` is recognise-only by design: the
// operator opens it by hand. Say so once rather than looking like a silent failure.
log.info("lead '{}' is not live, and names no profile — it can be recognised but not "
+ "launched. Add `profile:` under fleet.leaders.{} to have bridged start it.",
name, name);
continue;
}
BridgedConfig.Profile profile = cfg.profiles().get(lead.profile());
if (profile == null) {
log.warn("lead '{}' names profile '{}', which is not configured — not launching",
name, lead.profile());
continue;
}
for (int i = running; i < wanted; i++) {
if (launch(name, lead, profile)) {
started++;
}
}
}
return started;
}
/**
* How many live leads exist per configured name.
*
* <p>Two independent pieces of evidence, because either alone double-spawns:
* <ul>
* <li>a running agent in a tab labelled {@code "<tabPrefix> <name>"} — how an auto-launched
* lead, or an operator following the labelling convention, is found;</li>
* <li>a running agent on a terminal the config pins in {@code fleet.leaders.<name>.terminal} —
* how a lead the operator opened and pinned by hand is found. Without this, a pinned lead
* whose tab carries no matching label would be relaunched on every boot.</li>
* </ul>
* Member workspaces are excluded, exactly as the scanner excludes them: a member must not be
* counted as a lead because it happens to sit in a matching tab.
*/
private Map<String, Integer> liveLeads(Map<String, BridgedConfig.Leader> leaders) {
Set<String> memberSpaces = cfg.profiles().values().stream()
.map(BridgedConfig.Profile::workspace)
.filter(w -> w != null && !w.isBlank())
.collect(Collectors.toSet());
// tabId → the lead name its label declares.
Map<String, String> nameByTab = new LinkedHashMap<>();
for (Workspace ws : spaces.listWorkspaces()) {
if (ws.workspaceId() == null || memberSpaces.contains(ws.label())) {
continue;
}
for (Tab tab : spaces.listTabs(ws.workspaceId())) {
String declared = leadNameOf(tab.label(), leaders);
if (declared != null && tab.tabId() != null) {
nameByTab.put(tab.tabId(), declared);
}
}
}
// terminalId → the lead name the config pins it to.
Map<String, String> nameByPinnedTerminal = new LinkedHashMap<>();
leaders.forEach((name, lead) -> {
if (lead.terminal() != null && !lead.terminal().isBlank()) {
nameByPinnedTerminal.put(lead.terminal().strip(), name);
}
});
Map<String, Integer> counts = new LinkedHashMap<>();
for (Agent a : agents.list()) {
String name = nameByTab.get(a.tabId());
if (name == null) {
name = nameByPinnedTerminal.get(a.terminalId());
}
if (name != null) {
counts.merge(name, 1, Integer::sum);
}
}
return counts;
}
/**
* The configured lead a tab label names, or {@code null} for a label that names none.
*
* <p>Matched against the declared lead names rather than by splitting on the prefix, so an
* operator's {@code "lead: something-else"} tab is not mistaken for a configured lead.
*/
private String leadNameOf(String label, Map<String, BridgedConfig.Leader> leaders) {
if (label == null) {
return null;
}
String l = label.strip();
for (Map.Entry<String, BridgedConfig.Leader> e : leaders.entrySet()) {
if (l.equalsIgnoreCase(e.getValue().tabLabel(e.getKey()).strip())) {
return e.getKey();
}
}
return null;
}
/** Start one lead. Returns false (having logged) rather than throwing on any failure. */
private boolean launch(String name, BridgedConfig.Leader lead, BridgedConfig.Profile profile) {
String label = lead.tabLabel(name);
String cwd = (lead.cwd() == null || lead.cwd().isBlank())
? System.getProperty("user.dir") : lead.cwd();
Tab.Created tab = null;
try {
Workspace ws = spaces.ensureWorkspace(lead.workspace());
tab = spaces.createTab(ws.workspaceId(), cwd, leadEnv(profile));
if (tab.rootPaneId() == null) {
throw new IllegalStateException("tab " + tab.tab().tabId()
+ " came back with no seed pane — nowhere to start the lead");
}
// Same shape as the member launchers: herdr resolves the executable from `kind`, so
// argv[0] (the configured launcher, e.g. `ccs`) is dropped and only the rest is passed.
List<String> argv = leadArgv(profile);
Agent started = agents.start("lead-" + name, herdrKind(profile),
argv.isEmpty() ? argv : argv.subList(1, argv.size()), tab.rootPaneId());
// Label AFTER the start succeeds. A label written before would survive a failed start
// and then read back as a live lead on the next boot, which is the exact staleness the
// agent-liveness check exists to prevent — no need to create the case ourselves.
spaces.renameTab(tab.tab().tabId(), label);
log.info("lead '{}' launched: profile={} tab={} pane={} terminal={} label='{}' cwd={}",
name, profile.profile(), tab.tab().tabId(), started.paneId(),
started.terminalId(), label, cwd);
return true;
} catch (RuntimeException e) {
log.warn("lead '{}' failed to launch on profile '{}': {}",
name, profile.profile(), e.getMessage());
if (tab != null && tab.tab() != null && tab.tab().tabId() != null) {
try {
spaces.closeTab(tab.tab().tabId());
} catch (RuntimeException cleanup) {
log.warn("could not close the orphaned lead tab {}: {}",
tab.tab().tabId(), cleanup.getMessage());
}
}
return false;
}
}
/**
* The herdr agent kind for this profile — the same value the matching member adapter passes, so
* herdr resolves the same executable for a lead as it does for a member on that backend.
*/
private static String herdrKind(BridgedConfig.Profile profile) {
return profile.isOpenCode() ? "opencode" : "claude";
}
/**
* The lead's argv: the profile's own command, the model pin, and the bridge MCP mount.
*
* <p>No {@code --append-system-prompt}. That flag carries the worker reply charter, and a lead
* is not a worker — it reads its orchestration rules from the project's {@code CLAUDE.md} like
* any other primary. This is the single most important difference from the member launchers;
* do not "unify" it back.
*/
private List<String> leadArgv(BridgedConfig.Profile profile) {
List<String> argv = new ArrayList<>(profile.argv());
if (profile.hasMcp()) {
argv.add("--mcp-config");
argv.add("{\"mcpServers\":{\"bridge\":{\"type\":\"http\",\"url\":\""
+ profile.mcpUrl() + "\"}}}");
}
// Appended last, for the same reason the member launcher does it (CB-533): the argv is
// usually a wrapper such as `ccs <profile>`, which exports its own model family over
// whatever it inherited, and --model outranks the environment.
if (profile.model() != null && !profile.model().isBlank()) {
argv.add("--model");
argv.add(profile.model());
}
return argv;
}
/**
* The lead's environment: the profile's {@code env:} block, and nothing that could move it off
* the subscription.
*
* <p>{@code ANTHROPIC_BASE_URL} and {@code ANTHROPIC_AUTH_TOKEN} are stripped unconditionally —
* not defaulted, not guarded, stripped. A lead runs on the operator's subscription by
* definition, so there is no configuration under which pointing it elsewhere is correct, and a
* profile that carries them (a member profile reused as a lead's backend) must not leak them in.
*/
private Map<String, String> leadEnv(BridgedConfig.Profile profile) {
Map<String, String> out = new LinkedHashMap<>();
if (profile.env() != null) {
out.putAll(profile.env());
}
out.remove("ANTHROPIC_BASE_URL");
out.remove("ANTHROPIC_AUTH_TOKEN");
if (profile.configDir() != null && !profile.configDir().isBlank()) {
out.put("CLAUDE_CONFIG_DIR", profile.configDir());
}
// Deliberately no git token: a lead reviews and merges through the operator's own
// credentials, and never needs the scoped write:repository token a member is granted.
out.values().removeIf(Objects::isNull);
return out;
}
}
@@ -9,14 +9,15 @@ import dev.ltms.bridged.guard.GuardException;
import dev.ltms.bridged.herdr.Agent;
import dev.ltms.bridged.metrics.BridgedMetrics;
import dev.ltms.bridged.metrics.Metrics;
import dev.ltms.bridged.inject.WorkerPresence;
import dev.ltms.bridged.inject.MemberPresence;
import dev.ltms.bridged.herdr.HerdrException;
import dev.ltms.bridged.msg.MessageService;
import dev.ltms.bridged.msg.Rendezvous;
import dev.ltms.bridged.peer.PeerUnreachableException;
import dev.ltms.bridged.session.SessionManager;
import dev.ltms.bridged.session.WorkerSession;
import dev.ltms.bridged.session.MemberSession;
import dev.ltms.bridged.session.WorktreeRequest;
import dev.ltms.bridged.peer.MemberRole;
import dev.ltms.bridged.peer.PeerLauncher;
import io.modelcontextprotocol.common.McpTransportContext;
import io.modelcontextprotocol.json.McpJsonMapper;
@@ -65,6 +66,8 @@ public final class BridgeMcp {
static final String CALLER_PID = "callerPid";
/** Transport-context key under which the extractor stashes the resolved {@link Role} (CB-501). */
static final String CALLER_ROLE = "callerRole";
/** Transport-context key for the lead's configured name, when the caller is one (CB-530). */
static final String CALLER_NAME = "callerName";
private final HttpServletStreamableServerTransportProvider transport;
private final McpSyncServer server;
@@ -76,7 +79,7 @@ public final class BridgeMcp {
* without an auth fixture.
*/
public BridgeMcp(MessageService messages, PeerLauncher workers,
SessionManager sessions, ConnectionIdentity identity, WorkerPresence presence,
SessionManager sessions, ConnectionIdentity identity, MemberPresence presence,
PrimaryRegistry primaryRegistry) {
this(messages, workers, sessions, identity, presence, primaryRegistry, null, null);
}
@@ -89,7 +92,7 @@ public final class BridgeMcp {
* @param metrics registry for auth-failure counting; may be {@code null}
*/
public BridgeMcp(MessageService messages, PeerLauncher workers,
SessionManager sessions, ConnectionIdentity identity, WorkerPresence presence,
SessionManager sessions, ConnectionIdentity identity, MemberPresence presence,
PrimaryRegistry primaryRegistry, CallerResolver callers, Metrics metrics) {
McpJsonMapper json = new JacksonMcpJsonMapperSupplier().get();
this.transport = HttpServletStreamableServerTransportProvider.builder()
@@ -106,11 +109,15 @@ public final class BridgeMcp {
? callers.resolve(req.getRemoteAddr(), req.getRemotePort(),
req.getHeader("Authorization"))
: legacyPrincipal(identity, req.getRemoteAddr(), req.getRemotePort());
presence.markPresent(p.terminal()); // no-op for the primary (null terminal)
// CB-532: guard on the ROLE, not on the terminal being null. This excludes a
// lead, which carries its pane too, while including every spawned member role.
// Enrolling a lead would count it as an available member in the roster.
markSpawnedMemberPresent(p, presence);
return McpTransportContext.create(Map.of(
CALLER_TERMINAL, orEmpty(p.terminal()),
CALLER_PID, Long.toString(p.pid()),
CALLER_ROLE, p.role().name()));
CALLER_ROLE, p.role().name(),
CALLER_NAME, orEmpty(p.name())));
})
.build();
this.server = McpServer.sync(transport)
@@ -121,18 +128,31 @@ public final class BridgeMcp {
str(req.arguments(), "sessionId"));
if (denied != null) return denied;
String caller = callerTerminal(exchange);
if (caller != null) primaryRegistry.record(caller);
// CB-548: only a PRIMARY caller may claim the legacy singleton "primary" fallback.
// An architect delegates as its own pane but must never become the fallback that
// no-delegation inbox nudges target as if it were the primary (the per-target
// delegation map does not cure the singleton).
recordPrimarySingleton(primaryRegistry, caller, principal(exchange));
Map<String, Object> a = req.arguments();
String target = str(a, "sessionId");
String content = str(a, "content");
String turnId = str(a, "turnId");
if (turnId != null && !turnId.isBlank()) {
// Answering a worker's bridge_ask (CB-205): resolve its blocked question and
// block for the worker's reply as it resumes the same turn.
return answer(messages, turnId, str(a, "content"), timeoutMs(a));
// block for the worker's reply as it resumes the same turn. This is the same
// delegation, so ownership is left untouched (CB-548) — never re-recorded.
return answer(messages, turnId, content, timeoutMs(a));
}
// CB-548: delegator ownership (which lead's reply nudge this worker routes to,
// CB-532) is recorded only once the send is ACCEPTED — MessageService has won the
// session lock and queued delivery — via the accepted-delivery callback, never at
// request time. A concurrent sender that times out BUSY therefore cannot steal a
// live turn's reply routing without ever owning the turn.
Runnable onAccepted = () -> primaryRegistry.recordDelegation(target, caller);
// wait defaults to true (block for the reply); wait:false is fire-and-poll.
return Boolean.FALSE.equals(a.get("wait"))
? sendAsync(messages, str(a, "sessionId"), str(a, "content"))
: send(messages, str(a, "sessionId"), str(a, "content"), timeoutMs(a));
? sendAsync(messages, target, content, onAccepted)
: send(messages, target, content, timeoutMs(a), onAccepted);
})
// bridge_reply's identity is the CONNECTION, never an argument — so the authz check
// is "is this caller a worker at all", and it can only ever reply as itself.
@@ -173,19 +193,23 @@ public final class BridgeMcp {
McpSchema.CallToolResult denied = deny(exchange, Authz.Action.SPAWN, null);
if (denied != null) return denied;
String caller = callerTerminal(exchange);
if (caller != null) primaryRegistry.record(caller);
// SPAWN is already auth-gated to PRIMARY (architects can never call it), but
// enforce the same invariant here: only a PRIMARY may claim the legacy singleton.
recordPrimarySingleton(primaryRegistry, caller, principal(exchange));
Map<String, Object> a = req.arguments();
// CB-112: worker inherits the primary's cwd unless the call pins one.
// CB-301: carry the caller's identity as the session owner (null for the primary).
// CB-301-ext: optional isolated worktree for parallel implementers.
String callerCwd = identity.cwdForPid(callerPid(exchange));
return spawn(sessions, str(a, "profile"), str(a, "cwd"), callerCwd,
return spawn(sessions, str(a, "profile"), str(a, "role"), str(a, "cwd"), callerCwd,
callerTerminal(exchange), worktreeRequest(a));
})
.toolCall(listTool(), (exchange, _) -> {
McpSchema.CallToolResult denied = deny(exchange, Authz.Action.READ, null);
if (denied != null) return denied;
return listWorkers(workers, sessions);
return listFleet(workers, sessions,
callers == null ? Map.of() : callers.leads(),
callerTerminal(exchange));
})
.toolCall(stopTool(), (exchange, req) -> {
String paneId = str(req.arguments(), "paneId");
@@ -222,7 +246,7 @@ public final class BridgeMcp {
/** The caller reconstructed from the transport context. */
private static Principal principal(McpSyncServerExchange exchange) {
return principalFrom(exchange.transportContext().get(CALLER_ROLE),
callerTerminal(exchange), callerPid(exchange));
callerTerminal(exchange), callerPid(exchange), callerName(exchange));
}
/**
@@ -237,11 +261,16 @@ public final class BridgeMcp {
* @param pid the calling pid, or {@code -1}
*/
static Principal principalFrom(Object role, String terminal, long pid) {
return principalFrom(role, terminal, pid, null);
}
/** As {@link #principalFrom(Object, String, long)}, carrying a lead's name (CB-530). */
static Principal principalFrom(Object role, String terminal, long pid, String name) {
if (role == null) {
// No role stashed (legacy path): fall back to the historical interpretation.
return terminal != null ? Principal.worker(terminal, pid) : Principal.primary(pid);
}
return new Principal(Role.valueOf(role.toString()), terminal, pid);
return new Principal(Role.valueOf(role.toString()), terminal, pid, name);
}
/**
@@ -285,6 +314,24 @@ public final class BridgeMcp {
return error(reason + ": " + caller.describe() + " may not " + action);
}
/**
* Update the legacy singleton "primary" fallback used for no-delegation inbox nudges (CB-548).
*
* <p>Only {@link Role#PRIMARY} callers — the unnamed primary and named leads alike — may claim
* it. An architect delegates as its own pane but must never become the fallback: the per-target
* delegation map ({@code PrimaryRegistry#recordDelegation}) does not cure the singleton, so an
* architect left here would draw nudges that belong to a primary. The decision uses the resolved
* role, never name/kind sniffing. A null {@code caller} (legacy/no-auth path) records nothing.
*
* <p>Split out of the tool handlers so the guard is unit-testable without fabricating an SDK
* {@code McpSyncServerExchange} (same pattern as {@link #denyFor}/{@link #principalFrom}).
*/
static void recordPrimarySingleton(PrimaryRegistry registry, String callerTerminal, Principal caller) {
if (caller != null && caller.isPrimary()) {
registry.record(callerTerminal);
}
}
/** The worker identity resolved from this call's connection, or {@code null} if the primary. */
private static String callerTerminal(McpSyncServerExchange exchange) {
Object v = exchange.transportContext().get(CALLER_TERMINAL);
@@ -292,6 +339,13 @@ public final class BridgeMcp {
return (s == null || s.isBlank()) ? null : s;
}
/** The lead name resolved from this call's connection, or {@code null} (CB-530). */
private static String callerName(McpSyncServerExchange exchange) {
Object v = exchange.transportContext().get(CALLER_NAME);
String s = v == null ? null : v.toString();
return (s == null || s.isBlank()) ? null : s;
}
/** The caller's PID resolved from this call's connection, or {@code -1} if unknown. */
private static long callerPid(McpSyncServerExchange exchange) {
Object v = exchange.transportContext().get(CALLER_PID);
@@ -311,6 +365,13 @@ public final class BridgeMcp {
return transport;
}
/** Mark a connected spawned member available for the injector readiness gate. */
static void markSpawnedMemberPresent(Principal caller, MemberPresence presence) {
if (caller.isSpawnedMember()) {
presence.markPresent(caller.terminal());
}
}
/** Graceful shutdown of the MCP server. */
public void close() {
server.closeGracefully();
@@ -320,12 +381,22 @@ public final class BridgeMcp {
/** {@code bridge_send}: delegate {@code content} to a worker session and block for its reply. */
static McpSchema.CallToolResult send(MessageService messages, String sessionId, String content, Long timeoutMs) {
return send(messages, sessionId, content, timeoutMs, null);
}
/**
* As {@link #send(MessageService, String, String, Long)}, wiring an accepted-delivery hook
* (CB-548): {@code onAccepted} records delegator ownership the instant the send is accepted, so
* a BUSY interloper never claims a turn it did not win. {@code null} disables recording.
*/
static McpSchema.CallToolResult send(MessageService messages, String sessionId, String content,
Long timeoutMs, Runnable onAccepted) {
if (isBlank(sessionId) || isBlank(content)) {
return error("sessionId and content are required");
}
long timeout = clamp(timeoutMs == null ? DEFAULT_TIMEOUT_MS : timeoutMs);
try {
return formatReply(messages.send(sessionId, content, timeout), timeout);
return formatReply(messages.send(sessionId, content, timeout, onAccepted), timeout);
} catch (HerdrException e) {
return error("herdr error contacting session " + sessionId + ": " + e.getMessage());
}
@@ -394,10 +465,19 @@ public final class BridgeMcp {
* immediately (fire-and-poll), so a long task isn't cut off by the caller's MCP call timeout.
*/
static McpSchema.CallToolResult sendAsync(MessageService messages, String sessionId, String content) {
return sendAsync(messages, sessionId, content, null);
}
/**
* As {@link #sendAsync(MessageService, String, String)}, wiring the accepted-delivery hook
* (CB-548) so an async flooding send records delegator ownership exactly once it is accepted.
*/
static McpSchema.CallToolResult sendAsync(MessageService messages, String sessionId, String content,
Runnable onAccepted) {
if (isBlank(sessionId) || isBlank(content)) {
return error("sessionId and content are required");
}
String ticket = messages.sendAsync(sessionId, content);
String ticket = messages.sendAsync(sessionId, content, onAccepted);
return text("accepted — task delegated. Poll bridge_poll with ticket=" + ticket);
}
@@ -484,7 +564,29 @@ public final class BridgeMcp {
static McpSchema.CallToolResult whoami(Principal caller, SessionManager sessions) {
Map<String, Object> m = new LinkedHashMap<>();
m.put("role", caller.role().name().toLowerCase());
if (caller.isArchitect()) {
// CB-548: the role reads "architect"; the name is the gateway-local slot the pane is
// bound to, and the pane itself so a peer knows where to reach it.
if (caller.name() != null) {
m.put("architect", caller.name());
}
if (caller.terminal() != null) {
m.put("sessionId", caller.terminal());
}
return text(json(m));
}
if (!caller.isWorker()) {
// CB-530: which lead, once more than one pane is configured as one. `role` deliberately
// still reads "primary" — the fallback ladder in CLAUDE.md keys on it, and a lead IS a
// primary for authorization; the name is additive so no existing reader breaks.
if (caller.name() != null) {
m.put("leader", caller.name());
}
// CB-532: a lead's own pane, so it can tell a peer where to reach it — and so an
// operator can read off which tab hosts which lead without going to herdr.
if (caller.terminal() != null) {
m.put("sessionId", caller.terminal());
}
return text(json(m));
}
m.put("sessionId", caller.terminal());
@@ -512,23 +614,33 @@ public final class BridgeMcp {
/** {@code bridge_spawn} without cwd/caller context (default resolution). */
static McpSchema.CallToolResult spawn(SessionManager sessions, String profile) {
return spawn(sessions, profile, null, null, null, null);
return spawn(sessions, profile, null, null, null, null, null);
}
/**
* {@code bridge_spawn}: launch a guard-checked worker for {@code profile} (blank → the default
* profile) and return its session id + pane id. The worker's cwd is {@code requestedCwd} if given,
* else the profile's config, else {@code callerCwd} (the primary's directory), else the daemon's.
* {@code bridge_spawn}: launch a guard-checked member for {@code profile} (blank → the default
* profile) under {@code role} (blank → {@code dev}), and return its session id + pane id. The
* member's cwd is {@code requestedCwd} if given, else the profile's config, else
* {@code callerCwd} (the primary's directory), else the daemon's.
* CB-301: the session is registered with {@code ownerTerminal} as its owner.
* CB-301-ext: {@code worktreeRequest} non-null provisions an isolated git worktree.
*
* <p>{@code role} and {@code profile} are independent: the role picks the contract, the profile
* picks the backend. A reviewer on the same profile as the dev it reviews is a normal spawn.
*/
static McpSchema.CallToolResult spawn(SessionManager sessions, String profile,
static McpSchema.CallToolResult spawn(SessionManager sessions, String profile, String role,
String requestedCwd, String callerCwd,
String ownerTerminal, WorktreeRequest worktreeRequest) {
MemberRole memberRole;
try {
WorkerSession worker = sessions.acquire(isBlank(profile) ? null : profile,
memberRole = isBlank(role) ? MemberRole.DEV : MemberRole.parse(role);
} catch (IllegalArgumentException e) {
return error(e.getMessage());
}
try {
MemberSession member = sessions.acquire(isBlank(profile) ? null : profile, memberRole,
requestedCwd, callerCwd, ownerTerminal, worktreeRequest);
return text(json(workerView(worker)));
return text(json(memberView(member)));
} catch (GuardException e) {
return error("subscription boundary: " + e.getMessage());
} catch (IllegalArgumentException e) {
@@ -575,22 +687,66 @@ public final class BridgeMcp {
"default", workers.defaultProfile() == null ? "" : workers.defaultProfile())));
}
/** {@code bridge_list}: bridge-owned roster merged with live herdr status by paneId. */
static McpSchema.CallToolResult listWorkers(PeerLauncher workers, SessionManager sessions) {
/**
* {@code bridge_list}: the whole fleet — {@code leads} and {@code workers} — each merged with
* live herdr status. CB-519 decoupled the registry key (a host-unique id) from the herdr pane
* coordinate, so the join is on the terminal id, which both the session and the live agent carry.
*
* <p>CB-535 added the {@code leads} half. Until then this listed the worker roster alone, and a
* lead asking "who else is here?" got an empty array — which reads as <em>no peers</em> but
* actually means <em>no workers spawned</em>. There was no way at all for a lead to learn a
* peer's address; it had to be carried across by a human. Both halves are reported even when a
* half is empty, so an empty {@code workers} can no longer be mistaken for an empty fleet.
*
* <p>Leads are drawn from the resolver rather than from a second registry, so an address listed
* here is one that would actually resolve as a lead — see {@link CallerResolver#leads()}. The
* caller's own row is flagged {@code "self": true}: a peer needs to tell its own pane apart from
* a peer's, and the alternative is every lead calling {@code bridge_whoami} to subtract itself.
*
* @param leads terminal_id → lead name, live from the resolver
* @param selfTerm the calling pane's terminal id, or blank for a caller with no pane
*/
static McpSchema.CallToolResult listFleet(PeerLauncher workers, SessionManager sessions,
Map<String, String> leads, String selfTerm) {
try {
Map<String, Agent> live = workers.list().stream()
.map(Agent.class::cast)
.filter(a -> a.paneId() != null)
.collect(Collectors.toMap(Agent::paneId, Function.identity(), (_, b) -> b));
List<Map<String, Object>> out = sessions.roster().stream()
.map(s -> SessionManager.rosterView(s, live.get(s.paneId())))
.filter(a -> a.terminalId() != null)
.collect(Collectors.toMap(Agent::terminalId, Function.identity(), (_, b) -> b));
List<Map<String, Object>> leadRows = leads.entrySet().stream()
.sorted(Map.Entry.comparingByValue())
.map(e -> leadView(e.getKey(), e.getValue(), live.get(e.getKey()), selfTerm))
.toList();
return text(json(Map.of("workers", out)));
List<Map<String, Object>> out = sessions.roster().stream()
.map(s -> SessionManager.rosterView(s, live.get(s.terminalId())))
.toList();
return text(json(Map.of("leads", leadRows, "members", out)));
} catch (HerdrException e) {
return error("herdr error listing workers: " + e.getMessage());
return error("herdr error listing the fleet: " + e.getMessage());
}
}
/**
* One lead's row: its address, its name, and whether it can be reached right now.
*
* <p>{@code status} is herdr's live view, and {@code unknown} when herdr is not tracking that
* pane as an agent — the honest answer, and the one that matters: a lead whose pane herdr cannot
* see is a lead a {@code bridge_send} cannot be typed into. It is reported rather than hidden,
* because a peer that has gone unreachable is exactly what the sender needs to know.
*/
private static Map<String, Object> leadView(String terminal, String name, Agent live,
String selfTerm) {
Map<String, Object> m = new LinkedHashMap<>();
m.put("sessionId", terminal);
m.put("name", name);
m.put("status", live == null || live.status() == null
? "unknown" : live.status().name().toLowerCase());
if (terminal.equals(selfTerm)) {
m.put("self", true);
}
return m;
}
/** {@code bridge_stop}: tear a worker down by its pane id. */
static McpSchema.CallToolResult stop(SessionManager sessions, String paneId) {
if (isBlank(paneId)) {
@@ -605,10 +761,14 @@ public final class BridgeMcp {
}
/** CB-301 projection from the authoritative session registry. */
private static Map<String, Object> workerView(WorkerSession s) {
private static Map<String, Object> memberView(MemberSession s) {
Map<String, Object> m = new LinkedHashMap<>();
m.put("sessionId", s.terminalId());
m.put("paneId", s.paneId());
// Echo the role back so the caller can see what it actually got, not what it meant to ask
// for — a spawn that silently fell back to dev is otherwise invisible.
m.put("role", s.role() == null ? "dev" : s.role().wireName());
m.put("profile", s.profile());
m.put("status", s.state().name().toLowerCase());
if (s.worktree() != null) {
m.put("worktree", s.worktree());
@@ -685,14 +845,20 @@ public final class BridgeMcp {
private static McpSchema.Tool spawnTool() {
return tool("bridge_spawn",
"Spawn a new off-subscription worker session. Pass a profile (from bridge_profiles) to "
+ "pick the backend, or omit it for the default. The worker opens your current "
+ "directory by default; pass cwd to pin a different one. Pass worktree:true (with "
+ "ticket) or worktree:<ticket-slug> to provision an isolated git worktree. "
+ "Returns the worker's sessionId (use with bridge_send) and paneId (use with bridge_stop).",
"Spawn a new off-subscription member session. A member has two independent attributes: "
+ "role (what it is for) and profile (which backend it runs on). Pass role to pick "
+ "the contract — 'dev' implements a unit and opens its own PR, 'reviewer' reviews a "
+ "diff it did not write, 'architect' refines a ticket before anyone builds it; omit "
+ "it for 'dev'. Pass profile (from bridge_profiles) to pick the backend, or omit it "
+ "for the default. The two are independent: a reviewer may run on the same profile "
+ "as the dev it reviews. The member opens your current directory by default; pass "
+ "cwd to pin a different one. Pass worktree:true (with ticket) or "
+ "worktree:<ticket-slug> to provision an isolated git worktree. Returns the member's "
+ "sessionId (use with bridge_send) and paneId (use with bridge_stop).",
objectSchema(Map.of(
"profile", stringProp("Worker profile to spawn (omit for the default profile)"),
"cwd", stringProp("Working directory for the worker (omit to inherit yours)"),
"role", stringProp("What the member is for: architect, dev or reviewer (default dev)"),
"profile", stringProp("Which backend to run it on (omit for the default profile)"),
"cwd", stringProp("Working directory for the member (omit to inherit yours)"),
"worktree", Map.of("type", "string", "description", "'true' or a ticket slug — requests an isolated git worktree"),
"ticket", stringProp("Ticket slug when worktree:true")),
List.of()));
@@ -706,8 +872,14 @@ public final class BridgeMcp {
private static McpSchema.Tool listTool() {
return tool("bridge_list",
"List the worker sessions the bridge tracks — each with its sessionId, paneId, profile, "
+ "state, optional worktree/branch/owner, and live herdr status.",
"List the whole fleet the bridge tracks, in two parts. 'leads' are your PEERS — other "
+ "orchestrators, each with its sessionId (the address to bridge_send to), "
+ "name, live status, and 'self': true on your own row; this is how you "
+ "discover a peer lead without being told its address. 'members' are the "
+ "sessions delegated to — each with sessionId, paneId, role (architect/dev/"
+ "reviewer), profile (the backend it runs on), state, optional "
+ "worktree/branch/owner, and live herdr status. An empty 'members' "
+ "means no members are spawned; it says nothing about peers.",
objectSchema(Map.of(), List.of()));
}
@@ -720,10 +892,12 @@ public final class BridgeMcp {
}
private static McpSchema.Tool replyTool() {
// No session/target arg — the worker's identity is resolved from the connection.
// No session/target arg — the caller's identity is resolved from the connection.
return tool("bridge_reply",
"Return your structured answer for the task you were delegated, "
+ "resolving the caller's blocked bridge_send.",
"Return your structured answer for a message you were sent, resolving the sender's "
+ "blocked bridge_send. A worker MUST end every delegated turn with exactly "
+ "one of these. A lead uses it only to answer another lead that messaged "
+ "it — never to answer a worker, whose turn it is not.",
objectSchema(Map.of(
"content", stringProp("Your reply/answer")),
List.of("content")));
@@ -741,11 +915,14 @@ public final class BridgeMcp {
return tool("bridge_whoami",
"Report who YOU are on the bridge — your role is resolved from your connection "
+ "(unforgeable), never from anything you claim. Returns role 'primary' (you "
+ "orchestrate: spawn/send/stop, and you must never call bridge_reply) or "
+ "'worker' (you were delegated to: you must end every turn with exactly one "
+ "bridge_reply, and cannot spawn or send), plus your own sessionId, profile, "
+ "worktree and branch when you are a worker. Call this first when following "
+ "role-conditional instructions rather than guessing your role.",
+ "orchestrate: spawn/send/stop; reply ONLY to answer a peer lead that "
+ "messaged you, never to answer a worker), 'architect' (you delegate turns "
+ "and reply/ask as your own pane, but cannot spawn/stop/drain), or 'worker' "
+ "(you were delegated to: you must end every turn with exactly one "
+ "bridge_reply, and cannot spawn or send), plus 'leader'/'architect' naming "
+ "which one you are, your own sessionId, and profile/worktree/branch when "
+ "you are a worker. Call this first when following role-conditional "
+ "instructions rather than guessing.",
objectSchema(Map.of(), List.of()));
}
@@ -4,6 +4,7 @@ import org.slf4j.Logger;
import org.slf4j.LoggerFactory;
import java.util.Optional;
import java.util.concurrent.ConcurrentHashMap;
import java.util.concurrent.atomic.AtomicReference;
/**
@@ -25,6 +26,15 @@ public final class PrimaryRegistry {
private final AtomicReference<String> terminal = new AtomicReference<>();
private final boolean pinned;
/**
* CB-532: worker terminal → the lead that delegated to it. The single slot above answers "who is
* THE primary", a question with no correct answer once two leads orchestrate the same fleet:
* whichever called {@code bridge_send} first captured every nudge, including nudges for the
* other lead's delegations. This map answers the question that actually matters — "who is
* waiting on THIS worker" — and is what lets {@code primary.terminal} be retired.
*/
private final ConcurrentHashMap<String, String> leadByTarget = new ConcurrentHashMap<>();
/**
* @param pinnedTerminal an optional pinned terminal from config ({@code null}/blank = unpinned)
*/
@@ -56,6 +66,45 @@ public final class PrimaryRegistry {
}
}
/**
* Record that {@code leadTerminal} owns the accepted delegation of worker {@code target} (CB-532).
*
* <p>Called from the {@code MessageService} accepted-delivery hook — only after a send has won
* the session's send lock and queued delivery — where both halves are known (CB-548). It is
* deliberately <em>not</em> called at {@code bridge_send} request time: a concurrent sender that
* times out {@code BUSY} must not steal a live delegation's reply routing without ever owning
* the turn. Last writer wins — if a second lead's later send is accepted, replies follow the
* lead that most recently delegated to it, which is the one waiting.
*/
public void recordDelegation(String target, String leadTerminal) {
if (target == null || target.isBlank() || leadTerminal == null || leadTerminal.isBlank()) {
return;
}
leadByTarget.put(target, leadTerminal);
}
/** Forget a worker's delegating lead — call on release, so a torn-down session leaks nothing. */
public void forgetDelegation(String target) {
if (target != null) {
leadByTarget.remove(target);
}
}
/**
* Where a nudge about {@code target}'s reply should go: the lead that delegated to it, falling
* back to the single known primary.
*
* <p>The fallback matters after a daemon restart, which loses the map while the durable inbox
* keeps the reply. With one lead the fallback is unambiguous and correct. With several and no
* recorded delegation there is no right answer, so this returns empty rather than guessing —
* delivery degrades to pull, which is exactly what the durable inbox is for, instead of
* interrupting the wrong lead with someone else's result.
*/
public Optional<String> nudgeTargetFor(String target) {
String lead = target == null ? null : leadByTarget.get(target);
return lead != null ? Optional.of(lead) : Optional.ofNullable(terminal.get());
}
/** The known primary terminal, or empty if not yet learned (and not pinned). */
public Optional<String> primaryTerminal() {
return Optional.ofNullable(terminal.get());
@@ -0,0 +1,350 @@
package dev.ltms.bridged.member;
import dev.ltms.bridged.config.BridgedConfig;
import dev.ltms.bridged.guard.SubscriptionGuard;
import dev.ltms.bridged.herdr.Agent;
import dev.ltms.bridged.herdr.AgentControl;
import dev.ltms.bridged.herdr.WorkspaceControl;
import dev.ltms.bridged.peer.Capability;
import org.slf4j.Logger;
import org.slf4j.LoggerFactory;
import java.util.EnumSet;
import java.util.List;
import java.util.Map;
import java.util.Set;
import java.util.UUID;
import java.util.function.Function;
import java.util.function.LongSupplier;
import java.util.function.Supplier;
/**
* The {@link HerdrPeerLauncher} adapter for <strong>Claude Code</strong> — the safe path from a
* delegation request to a running off-subscription Claude.
*
* <p>Everything transport-related (tab/pane placement, the CB-306 spawn-readiness gate, unique
* naming, CB-117 orphan reap, teardown, listing, cwd resolution) lives in the base. This class
* supplies only the two Claude-specific seams:
* <ul>
* <li>the {@code claude} name prefix (so reap matches {@code claude-*} panes, never another
* adapter's), and</li>
* <li>{@link #buildLaunch}, which encodes the subscription boundary: build the worker env with
* {@code ANTHROPIC_BASE_URL}, assert that host is on the allowlist <em>before</em> touching
* herdr, and mount the bridge MCP + reply charter as inline launch flags. A worker's base_url
* lives in the env map handed to herdr and nowhere else; {@code bridged}'s own environment is
* never mutated, and nothing is written to the worker's profile.</li>
* </ul>
*/
public final class ClaudeCodeLauncher extends HerdrPeerLauncher {
/** Label prefix for this adapter's herdr agent names (drives naming + orphan reap). */
private static final String NAME_PREFIX = "claude";
private static final Logger log = LoggerFactory.getLogger(ClaudeCodeLauncher.class);
private final SubscriptionGuard guard;
/**
* Standing instruction appended to the worker's system prompt so it returns its result via
* {@code bridge_reply}. Injected as a launch flag, so nothing is written to the worker's
* profile — it is guidance, and a worker that never replies is caught by the send's timeout.
*/
static final String REPLY_CHARTER =
"You are an off-subscription worker in the claude-bridge fleet. Every message you "
+ "receive arrives through the bridge, and the ONLY channel back to the sender is the "
+ "bridge_reply MCP tool. Text you write in your terminal is NOT sent anywhere — the "
+ "sender cannot see your screen, so an in-terminal answer is silently discarded. "
+ "Therefore you MUST end EVERY turn by calling bridge_reply with `content` set to your "
+ "complete response. This holds for every message without exception — tasks, questions, "
+ "clarifications, acknowledgements, and ordinary back-and-forth conversation. Call "
+ "bridge_reply exactly once, as the final action of your turn, with your full answer in "
+ "`content`; never wait for confirmation first. If you end a turn without calling "
+ "bridge_reply, the sender receives nothing and the exchange stalls.";
/**
* Production constructor — disables the spawn-ready gate ({@code spawnReadyTimeoutMs == 0}) so
* existing deployments and tests keep the legacy non-blocking spawn semantics.
*/
public ClaudeCodeLauncher(AgentControl agents, WorkspaceControl spaces, SubscriptionGuard guard,
Map<String, BridgedConfig.Profile> profiles, String defaultProfile,
Function<String, String> env) {
this(agents, spaces, guard, profiles, defaultProfile, env, 0,
System::currentTimeMillis, () -> sleepUninterruptibly(300));
}
/**
* Production constructor with spawn-ready gate enabled. The gate polls {@code agents.status()}
* until the pane reports an injectable state or {@code spawnReadyTimeoutMs} elapses.
*/
public ClaudeCodeLauncher(AgentControl agents, WorkspaceControl spaces, SubscriptionGuard guard,
Map<String, BridgedConfig.Profile> profiles, String defaultProfile,
Function<String, String> env,
long spawnReadyTimeoutMs, long spawnReadyPollMs) {
this(agents, spaces, guard, profiles, defaultProfile, env,
spawnReadyTimeoutMs, spawnReadyPollMs, null);
}
/**
* Production constructor carrying the fleet-wide tab-label template (CB-557). The template comes
* from {@code fleet.tabLabel}, which a profile cannot know because it names the member's
* <em>role</em>; a profile may still override it with its own {@code tabLabel}.
*/
public ClaudeCodeLauncher(AgentControl agents, WorkspaceControl spaces, SubscriptionGuard guard,
Map<String, BridgedConfig.Profile> profiles, String defaultProfile,
Function<String, String> env,
long spawnReadyTimeoutMs, long spawnReadyPollMs,
Supplier<String> tabLabelTemplate) {
this(agents, spaces, guard, profiles, defaultProfile, env,
spawnReadyTimeoutMs,
System::currentTimeMillis, () -> sleepUninterruptibly(spawnReadyPollMs),
tabLabelTemplate);
}
/**
* Full testability constructor. Every injectable collaborator is explicit so unit tests supply
* fakes for the clock ({@code nowMillis}) and poll-loop wait ({@code sleeper}). The
* {@code sleeper} is never called when the gate is disabled ({@code spawnReadyTimeoutMs == 0}).
*
* @param agents herdr agent control (start, status, close)
* @param spaces workspace / tab control (ensure, create, close)
* @param guard subscription-boundary guard (checked before spawning)
* @param profiles configured worker profiles
* @param defaultProfile profile a no-argument spawn uses (nullable)
* @param env host env lookup (injectable for tests)
* @param spawnReadyTimeoutMs max ms to wait for injectable state (0 disables the gate)
* @param nowMillis monotonic clock source (e.g. {@code System::currentTimeMillis})
* @param sleeper sleep/wait hook (e.g. {@code () -> Thread.sleep(pollMs)}); it
* already encodes the poll interval, so the 8th positional argument
* (poll ms) is accepted for API symmetry but otherwise unused here
*/
public ClaudeCodeLauncher(AgentControl agents, WorkspaceControl spaces, SubscriptionGuard guard,
Map<String, BridgedConfig.Profile> profiles, String defaultProfile,
Function<String, String> env,
long spawnReadyTimeoutMs,
LongSupplier nowMillis, Runnable sleeper) {
this(agents, spaces, guard, profiles, defaultProfile, env,
spawnReadyTimeoutMs, nowMillis, sleeper, null);
}
/**
* Full testability constructor, plus the fleet-wide tab-label template (CB-557).
*
* @param tabLabelTemplate {@code fleet.tabLabel}; {@code null}/blank ⇒
* {@link BridgedConfig.Fleet#DEFAULT_TAB_LABEL}
*/
public ClaudeCodeLauncher(AgentControl agents, WorkspaceControl spaces, SubscriptionGuard guard,
Map<String, BridgedConfig.Profile> profiles, String defaultProfile,
Function<String, String> env,
long spawnReadyTimeoutMs,
LongSupplier nowMillis, Runnable sleeper,
Supplier<String> tabLabelTemplate) {
super(NAME_PREFIX, agents, spaces, profiles, defaultProfile, env,
spawnReadyTimeoutMs, nowMillis, sleeper, tabLabelTemplate);
this.guard = guard;
}
/**
* {@inheritDoc}
*
* <p>A legacy spawn with no session identity is a fresh, launcher-derived session — delegate to
* the session-aware form with no name and no resume id.
*/
@Override
protected Launch buildLaunch(BridgedConfig.Profile cfg) {
return buildLaunch(cfg, null, null);
}
/**
* {@inheritDoc}
*
* <p>The spawn sequence encodes the subscription boundary: assert the profile's base_url is on
* the allowlist <em>before</em> any herdr call, then build the worker env with
* {@code ANTHROPIC_*}, the parity-neutral git-forge grant, and the bridge MCP + reply charter
* mounted as inline launch flags. When the request carries session identity (CB-547a) it is
* applied here — see {@link #applySessionIdentity}.
*/
@Override
protected Launch buildLaunch(BridgedConfig.Profile cfg, String sessionName, String resumeSessionId) {
// CB-539: a profile may deliberately opt into the subscription (subscription: true) when no
// off-subscription endpoint exists for it — e.g. `sonnet` on `ccs`. That profile gets no
// ANTHROPIC_BASE_URL/AUTH_TOKEN (there is nothing to point them at) and the guard's base_url
// requirement is skipped FOR IT ONLY. Every other profile keeps the hard boundary below.
boolean onSubscription = cfg.isSubscription();
String baseUrl = cfg.baseUrl();
if (onSubscription) {
// NO SILENT CONTRADICTION: subscription:true + a baseUrl state opposite intents; refuse
// loudly rather than pick a winner.
if (baseUrl != null && !baseUrl.isBlank()) {
throw new IllegalStateException("profile '" + cfg.profile()
+ "' sets both subscription: true and a baseUrl ('" + baseUrl + "') — the two "
+ "are contradictory: a subscription profile must not point at an endpoint. "
+ "Drop baseUrl, or drop subscription: true.");
}
// Visible without anyone going looking for it: this worker bills the subscription.
log.warn("spawning profile '{}' on the Claude subscription (subscription: true) — this "
+ "worker WILL bill the operator's subscription", cfg.profile());
} else {
guard.assertWorker(baseUrl); // hard stop before we spawn anything
}
Map<String, String> workerEnv = baseEnv(cfg);
if (onSubscription) {
// CB-542 belt-and-braces: on the subscription path no guard vets these two keys, and the
// profile's env: is layered in by baseEnv — so strip any that rode in there. Config load
// already rejects this (loudly, naming the profile); this makes the boundary hold even
// for a profile built in code that never passed through that validation.
workerEnv.remove("ANTHROPIC_BASE_URL");
workerEnv.remove("ANTHROPIC_AUTH_TOKEN");
} else {
workerEnv.put("ANTHROPIC_BASE_URL", baseUrl);
putIfPresent(workerEnv, "ANTHROPIC_AUTH_TOKEN", env.apply(cfg.tokenEnv()));
}
putIfPresent(workerEnv, "ANTHROPIC_MODEL", cfg.model());
putIfPresent(workerEnv, "CLAUDE_CONFIG_DIR", cfg.configDir());
applyGitToken(workerEnv, cfg);
// CB-547a: Claude Code can MINT its own session id, so bridged chooses it — a fresh spawn
// gets a UUID we pass as --session-id and return from agentSessionId(), so the resume
// handle is known BEFORE the agent has written anything; a resume spawn adopts its prior
// id via -r and passes no --session-id (the two conflict). Both are injected before the
// model flag so --model keeps outranking the operator's own argv.
// mutableArgv: argvWithBridge may hand back the profile's own (immutable) List.of when it
// has no MCP — session flags must be added into a list we own.
List<String> argv = mutableArgv(argvWithBridge(cfg));
String agentSessionId = applySessionIdentity(argv, sessionName, resumeSessionId);
return new Launch(workerEnv, argvWithModel(argv, cfg), agentSessionId);
}
/**
* Add the Claude-specific session-identity flags to {@code argv} and return the peer's OWN
* session id — the resume handle. A resume request passes the prior id via {@code -r} and
* returns that id; a fresh named session mints a new UUID, passes it via {@code --session-id},
* and returns the mint. The bridge's logical name rides along as {@code -n} when present. When
* <em>no</em> identity is requested (sessionName and resumeSessionId both blank) this adds
* nothing and returns {@code null}, keeping the legacy no-identity launch byte-identical.
*/
private static String applySessionIdentity(List<String> argv, String sessionName, String resumeSessionId) {
boolean resuming = resumeSessionId != null && !resumeSessionId.isBlank();
boolean named = sessionName != null && !sessionName.isBlank();
if (!resuming && !named) {
return null; // no identity requested — keep the legacy launch byte-identical
}
if (named) {
argv.add("-n");
argv.add(sessionName);
}
if (resuming) {
argv.add("-r");
argv.add(resumeSessionId);
return resumeSessionId;
}
String minted = UUID.randomUUID().toString();
argv.add("--session-id");
argv.add(minted);
return minted;
}
/**
* The launch argv, plus — when {@code worker.mcpUrl} is set — inline {@code --mcp-config} for
* the bridge server and {@code --append-system-prompt} for the {@link #REPLY_CHARTER}. Neither
* touches the profile's config; both are pure command-line flags. This inline-flag mount is
* Claude Code specific — other adapters mount MCP and instructions their own way.
*/
private List<String> argvWithBridge(BridgedConfig.Profile cfg) {
if (!cfg.hasMcp()) {
return cfg.argv();
}
String mcpJson = "{\"mcpServers\":{\"bridge\":{\"type\":\"http\",\"url\":\""
+ cfg.mcpUrl() + "\"}}}";
List<String> argv = mutableArgv(cfg.argv());
argv.add("--mcp-config");
argv.add(mcpJson);
argv.add("--append-system-prompt");
argv.add(REPLY_CHARTER);
return argv;
}
/**
* Pin the model on the command line as well as in {@code ANTHROPIC_MODEL} (CB-533).
*
* <p>The env var alone is not a reliable pin for this adapter, because the argv is usually a
* launcher rather than {@code claude} itself — {@code ["ccs", "<profile>"]} — and {@code ccs}
* exports its profile's own model family ({@code ANTHROPIC_MODEL}, {@code DEFAULT_OPUS/SONNET/
* HAIKU}, {@code CLAUDE_CODE_SUBAGENT_MODEL}) over whatever it inherited. A worker profile that
* set {@code model:} therefore got silently overruled by its own launcher. Claude Code's
* {@code --model} flag outranks the environment, and {@code ccs <profile> [claude-args...]}
* passes trailing arguments through, so the flag survives the wrapper.
*
* <p>Appended last so it also outranks anything in the operator's own {@code argv}. Profiles
* that deliberately leave {@code model:} unset (letting {@code ccs} own model selection, as
* {@code gx10} does) are untouched — this adds nothing when there is nothing to add. This is
* the {@code kind: claude} counterpart of the opencode adapter's {@code -m provider/model}.
*/
private static List<String> argvWithModel(List<String> argv, BridgedConfig.Profile cfg) {
if (cfg.model() == null || cfg.model().isBlank()) {
return argv;
}
List<String> withModel = mutableArgv(argv);
withModel.add("--model");
withModel.add(cfg.model());
return withModel;
}
// --- Agent-returning convenience spawns (used by callers/tests that want the herdr Agent) ---
/** Spawn a worker for the default profile in the resolved default cwd. */
public Agent spawn() {
return spawnInternal(null, null, null);
}
/** Spawn a worker for a named profile (null → default) in the resolved default cwd. */
public Agent spawn(String profileName) {
return spawnInternal(profileName, null, null);
}
/** Spawn a worker for a named profile with an explicit requested/caller cwd (CB-112). */
public Agent spawn(String profileName, String requestedCwd, String callerCwd) {
return spawnInternal(profileName, requestedCwd, callerCwd);
}
// --- capabilities --------------------------------------------------------------------------
@Override
public Set<Capability> capabilities() {
Set<Capability> caps = EnumSet.of(Capability.MID_TURN_ASK, Capability.WORKTREE,
Capability.CONTEXT_RESET, Capability.ORPHAN_REAP,
Capability.SESSION_NAME, Capability.SESSION_RESUME);
if (hasGitTokenProfile()) {
caps.add(Capability.SELF_PR);
}
return Set.copyOf(caps);
}
@Override
public boolean clearContext(String id) {
String target = agentTarget(id);
if (target == null) {
return false;
}
// This deliberately bypasses Injector: /clear is housekeeping, not a delegated turn.
agents().send(target, "/clear");
return true;
}
/** Whether any configured profile opts into a git-forge token (required for {@link Capability#SELF_PR}). */
private boolean hasGitTokenProfile() {
return profileConfigs().stream().anyMatch(BridgedConfig.Profile::hasGitToken);
}
// --- CB-117 reap predicate (Claude prefix), kept for direct unit testing -------------------
/**
* Whether {@code name} is a Claude Code bridge worker started by a <em>different</em> process
* than {@code currentNonce}. A thin {@code claude}-prefix binding of
* {@link HerdrPeerLauncher#isForeignWorker(String, String, String)}.
*/
static boolean isForeignWorker(String name, String currentNonce) {
return HerdrPeerLauncher.isForeignWorker(NAME_PREFIX, name, currentNonce);
}
}
@@ -0,0 +1,422 @@
package dev.ltms.bridged.member;
import dev.ltms.bridged.config.BridgedConfig;
import dev.ltms.bridged.herdr.Agent;
import dev.ltms.bridged.peer.Capability;
import dev.ltms.bridged.peer.MemberRole;
import dev.ltms.bridged.peer.PeerHandle;
import dev.ltms.bridged.peer.PeerLauncher;
import dev.ltms.bridged.peer.PeerUnreachableException;
import dev.ltms.bridged.peer.SpawnRequest;
import dev.ltms.bridged.placement.PlacementCandidate;
import dev.ltms.bridged.placement.PlacementContext;
import dev.ltms.bridged.placement.PlacementException;
import dev.ltms.bridged.placement.PlacementPolicies;
import dev.ltms.bridged.placement.PlacementPolicy;
import org.slf4j.Logger;
import org.slf4j.LoggerFactory;
import java.util.ArrayList;
import java.util.Collections;
import java.util.EnumSet;
import java.util.HashSet;
import java.util.LinkedHashMap;
import java.util.List;
import java.util.Map;
import java.util.Set;
import java.util.concurrent.ConcurrentHashMap;
import java.util.function.Function;
import java.util.function.Supplier;
/**
* The {@link PeerLauncher} the core actually holds when more than one adapter is configured — a thin
* router in front of one {@link HerdrPeerLauncher} per peer {@code kind} (Claude Code, opencode, …).
* It owns no transport of its own; it dispatches each SPI call to the delegate that owns the profile
* involved, and fans the fleet-wide queries (list/reap/caps/profiles) across all delegates.
*
* <p>Routing rules:
* <ul>
* <li><strong>By profile</strong> — {@link #spawn}, {@link #effectiveCwd}, {@link #parityOverlay}
* resolve the profile (a null/blank name → the global {@link #defaultProfile}) and delegate to
* the single adapter that declares it. Profiles partition cleanly across adapters: the
* constructor rejects a name claimed by two.</li>
* <li><strong>By pane id</strong> — {@link #stop} routes to the adapter that spawned that pane
* (recorded at spawn time). A pane the composite never spawned (only real for a caller that
* hand-rolls an id) falls back to the first delegate; teardown is pane-id addressed and
* tab cleanup is single-occupant guarded, so it is safe either way.</li>
* <li><strong>Fleet-wide</strong> — {@link #reapOrphanWorkers} and {@link #capabilities} fan out
* and combine. {@link #list} is deduplicated by pane id because every herdr-backed delegate
* shares one herdr connection and so reports the same global agent set.</li>
* </ul>
*
* <p>CB-518: an unqualified spawn is routed through a {@link PlacementPolicy}. The default
* {@code fixed} policy reproduces the historical default-profile behaviour; {@code weighted} uses
* smooth weighted round-robin with {@code maxLoad} gating. If a chosen profile fails with
* {@link PeerUnreachableException}, the composite advances to the next available candidate and
* retries, bounded by the number of candidates.
*/
public final class CompositePeerLauncher implements PeerLauncher {
private static final Logger log = LoggerFactory.getLogger(CompositePeerLauncher.class);
private final List<HerdrPeerLauncher> delegates;
private final Map<String, HerdrPeerLauncher> byProfile;
private final String defaultProfile;
/** paneId → the delegate that spawned it, so {@link #stop} tears down through the right adapter. */
private final Map<String, HerdrPeerLauncher> spawnedBy = new ConcurrentHashMap<>();
private final Function<String, Integer> liveCount;
/**
* CB-559: the placement inputs are read <em>per spawn</em>, not captured at construction, so a
* config reload changes where the next member lands without a restart. These are the hot keys —
* role pools, an existing profile's weight/maxLoad, and the placement policy. What cannot change
* this way is the set of adapters ({@link #byProfile}), because a new backend needs a launcher
* and launchers are built once; {@code ConfigRef} classifies that as deferred and says so.
*/
private final Supplier<Map<String, BridgedConfig.Profile>> profileConfigs;
private final Supplier<PlacementPolicy> placementPolicy;
/**
* CB-557: the role pools an unqualified spawn draws its candidates from. A supplier that yields
* {@code null}, and an empty pool for a role, both fall back to every configured profile — the
* pre-CB-557 behaviour.
*/
private final Supplier<BridgedConfig.Fleet> fleet;
/**
* Backward-compatible constructor: fixed placement, no live-counting. Use this for tests and
* simple wiring; it preserves the pre-CB-518 behaviour exactly.
*
* @param delegates one adapter per configured peer kind; must be non-empty and declare
* disjoint profile-name sets
* @param defaultProfile the profile a no-argument spawn resolves to (may be null)
* @throws IllegalArgumentException if {@code delegates} is empty or two adapters claim one profile
*/
public CompositePeerLauncher(List<HerdrPeerLauncher> delegates, String defaultProfile) {
this(delegates, defaultProfile, Map.of(), PlacementPolicies.fixed(), _ -> 0);
}
/**
* Production constructor with a placement policy and live-worker counter.
*
* @param delegates one adapter per configured peer kind; must be non-empty and declare
* disjoint profile-name sets
* @param defaultProfile the profile a no-argument spawn resolves to under {@code fixed} policy
* @param profileConfigs all configured worker profiles (used for candidate weights/caps)
* @param placementPolicy which policy governs unqualified spawns
* @param liveCount live worker count per profile (must never return {@code null})
*/
public CompositePeerLauncher(List<HerdrPeerLauncher> delegates,
String defaultProfile,
Map<String, BridgedConfig.Profile> profileConfigs,
PlacementPolicy placementPolicy,
Function<String, Integer> liveCount) {
this(delegates, defaultProfile, profileConfigs, placementPolicy, liveCount, null);
}
/**
* Production constructor with role pools (CB-557). An unqualified spawn draws its candidates from
* {@code fleet.<role>} instead of from every configured profile, so a reviewer is placed on a
* reviewer backend and never on, say, the architect-only one.
*
* @param fleet the configured role pools; {@code null} ⇒ every profile is a candidate for every
* role, which is the pre-CB-557 behaviour
*/
public CompositePeerLauncher(List<HerdrPeerLauncher> delegates,
String defaultProfile,
Map<String, BridgedConfig.Profile> profileConfigs,
PlacementPolicy placementPolicy,
Function<String, Integer> liveCount,
BridgedConfig.Fleet fleet) {
// LinkedHashMap, not Map.copyOf: candidates() promises definition order and the weighted
// policy breaks exact-weight ties on it, so a salted iteration order would make placement
// differ from one JVM run to the next.
this(delegates, defaultProfile,
constant(Collections.unmodifiableMap(new LinkedHashMap<>(profileConfigs))),
constant(placementPolicy), liveCount, constant(fleet));
}
/**
* Production constructor that re-reads its placement inputs per spawn (CB-559), so a config
* reload retargets the next member without a restart.
*
* @param config the live configuration — read at every spawn, never captured
*/
public CompositePeerLauncher(List<HerdrPeerLauncher> delegates,
String defaultProfile,
Supplier<BridgedConfig> config,
Function<String, Integer> liveCount) {
this(delegates, defaultProfile,
() -> config.get().profiles(),
() -> PlacementPolicies.fromName(config.get().placement()),
liveCount,
() -> config.get().fleet());
}
/** The all-suppliers form every other constructor funnels into. */
private CompositePeerLauncher(List<HerdrPeerLauncher> delegates,
String defaultProfile,
Supplier<Map<String, BridgedConfig.Profile>> profileConfigs,
Supplier<PlacementPolicy> placementPolicy,
Function<String, Integer> liveCount,
Supplier<BridgedConfig.Fleet> fleet) {
this.fleet = fleet;
if (delegates.isEmpty()) {
throw new IllegalArgumentException("at least one peer adapter must be configured");
}
this.delegates = List.copyOf(delegates);
this.defaultProfile = defaultProfile;
this.profileConfigs = profileConfigs;
this.placementPolicy = placementPolicy;
this.liveCount = liveCount;
Map<String, HerdrPeerLauncher> index = new LinkedHashMap<>();
for (HerdrPeerLauncher d : this.delegates) {
for (String profile : d.profiles()) {
HerdrPeerLauncher prev = index.putIfAbsent(profile, d);
if (prev != null) {
throw new IllegalArgumentException(
"worker profile '" + profile + "' is claimed by two peer adapters");
}
}
}
// Order-preserving for the same reason, and because profiles() is user-visible (bridge_profiles).
this.byProfile = Collections.unmodifiableMap(index);
}
/** A supplier of a value fixed at construction — how the non-reloading constructors funnel in. */
private static <T> Supplier<T> constant(T value) {
return () -> value;
}
/**
* The currently-configured profiles, never null.
*
* <p>Read fresh on every call so a reload is visible; a caller that needs two consistent reads
* takes one local, as {@link #poolFor} does.
*/
private Map<String, BridgedConfig.Profile> profiles0() {
Map<String, BridgedConfig.Profile> m = profileConfigs.get();
return m == null ? Map.of() : m;
}
/** The adapter owning {@code profileName} (null/blank → the default). Throws on an unknown profile. */
private HerdrPeerLauncher route(String profileName) {
String resolved = (profileName == null || profileName.isBlank()) ? defaultProfile : profileName;
if (resolved == null) {
// No profile and no default configured — hand to the first delegate so it raises the
// same "no default" error it would on its own; keeps the SPI contract single-sourced.
return delegates.getFirst();
}
HerdrPeerLauncher d = byProfile.get(resolved);
if (d == null) {
throw new IllegalArgumentException("unknown worker profile: " + resolved);
}
return d;
}
@Override
public PeerHandle spawn(SpawnRequest req) {
String requestedProfile = req.profileName();
if (requestedProfile != null && !requestedProfile.isBlank()) {
// An explicit profile bypasses the placement policy, but not the capacity cap: maxLoad
// is documented as an unconditional limit on this profile (BridgedConfig.Profile), and
// the charter makes explicit-profile spawns the normal path — so skipping the check
// here would leave the cap dead config in real operation.
HerdrPeerLauncher d = route(requestedProfile);
enforceMaxLoad(requestedProfile);
PeerHandle handle = d.spawn(req);
spawnedBy.put(handle.id(), d);
return handle;
}
// CB-557: an unqualified spawn is placed inside the pool of the role it asked for, not across
// the whole profile list. An EXPLICIT profile (above) is left alone on purpose — it is the
// operator overriding, and refusing it would break `bridge_spawn{profile:"opus"}`, which
// carries no role and so would be judged against the dev pool it was never meant for.
List<PlacementCandidate> candidates = candidates(req.role());
String roleDefault = defaultProfileFor(req.role());
Set<String> unreachable = new HashSet<>();
PlacementContext ctx = new PlacementContext(roleDefault, candidates, liveCount, unreachable);
int maxAttempts = candidates.isEmpty() ? 1 : candidates.size();
for (int attempt = 0; attempt < maxAttempts; attempt++) {
// Deliberately uncaught: when no candidate is left (all at cap, or all unreachable) the
// policy already throws a clear message. Catching it to rethrow a generic
// PeerUnreachableException would replace a precise diagnosis with a vague one.
PlacementCandidate chosen = placementPolicy.get().select(ctx);
HerdrPeerLauncher d = byProfile.get(chosen.profile());
if (d == null) {
// A configured profile with no adapter is a wiring bug; fail fast.
unreachable.add(chosen.profile());
continue;
}
// CB-547a: route the chosen profile but keep the caller's session identity — dropping it
// here would silently sever the resume handle on every policy-routed spawn. CB-557: the
// role rides along for the same reason, or a routed spawn would be labelled as a dev.
SpawnRequest routedReq = new SpawnRequest(chosen.profile(), req.requestedCwd(), req.callerCwd(),
req.sessionName(), req.resumeSessionId(), req.role());
try {
PeerHandle handle = d.spawn(routedReq);
spawnedBy.put(handle.id(), d);
return handle;
} catch (PeerUnreachableException e) {
log.warn("spawn on profile {} unreachable, will retry next candidate if any: {}",
chosen.profile(), e.getMessage());
unreachable.add(chosen.profile());
// Update the context for the next selection so the policy excludes this profile.
ctx = new PlacementContext(roleDefault, candidates, liveCount, unreachable);
}
}
throw new PeerUnreachableException(
"no reachable worker profile available after trying " + unreachable.size()
+ " candidate(s): " + String.join(", ", unreachable));
}
/**
* Refuse an explicit-profile spawn when the profile is at its {@code maxLoad} cap.
*
* <p>maxLoad is a documented, unconditional capacity limit (see {@code BridgedConfig.Profile#maxLoad}),
* and the charter makes explicit-profile spawns the normal path — so enforcing it only in placement
* ({@code PlacementPolicyUtil}, package-private, hence not linked) would leave the cap dead config
* on every call that names a profile. Same rule as placement: {@code live >= cap} is at capacity.
*
* <p>Deliberately no fallback to another profile: the caller named {@code profile} for a cost/model
* reason, and silently re-routing a paid-tier (subscription) request elsewhere is worse than
* refusing it. A caller that wants placement should omit the profile and let the policy pick.
*
* <p>Known TOCTOU limitation — documented, not fixed. {@link #liveCount} is read outside any lock and
* {@code SessionManager} registers a session only after {@code launcher.spawn} returns, so two
* genuinely concurrent spawns can both pass this check. The race already exists on the placement
* path. Closing it needs slot reservation in the registry; serializing spawn here would block on
* the readiness gate and is a far worse trade.
*
* @param profile the profile the caller explicitly named
* @throws PlacementException when the profile is at capacity
*/
private void enforceMaxLoad(String profile) {
// Absent config, or a config whose maxLoad normalized to null (non-positive ⇒ unlimited at
// load), means no cap — never cap what wasn't configured.
BridgedConfig.Profile cfg = profiles0().get(profile);
Integer cap = (cfg == null) ? null : cfg.maxLoad();
if (cap == null) {
return;
}
int live = liveCount.apply(profile);
if (live >= cap) {
throw new PlacementException("worker profile '" + profile + "' is at maxLoad: " + live
+ " live >= " + cap + " cap; refusing spawn — no fallback to another profile");
}
}
/**
* The profile names {@code role} may be placed on, in definition order.
*
* <p>An empty or absent pool means "unconstrained", not "nothing allowed": a config that declares
* no pool for a role must keep spawning, so it falls back to every configured profile. Names in a
* pool that no adapter declares are dropped here rather than thrown — config load already rejects
* a pool entry with no profile, so a survivor is a profile this particular composite does not own.
*/
private List<String> poolFor(MemberRole role) {
Map<String, BridgedConfig.Profile> configured = profiles0();
BridgedConfig.Fleet f = fleet.get();
List<String> pool = (f == null) ? List.of() : f.profilesFor(role);
List<String> known = pool.stream().filter(configured::containsKey).toList();
return known.isEmpty() ? List.copyOf(configured.keySet()) : known;
}
/** The profile an unqualified spawn for {@code role} falls back to under {@code fixed} placement. */
private String defaultProfileFor(MemberRole role) {
List<String> pool = poolFor(role);
return pool.isEmpty() ? defaultProfile : pool.getFirst();
}
/** Build the candidate list from {@code role}'s pool, in definition order. */
private List<PlacementCandidate> candidates(MemberRole role) {
List<PlacementCandidate> out = new ArrayList<>();
for (String name : poolFor(role)) {
BridgedConfig.Profile w = profiles0().get(name);
if (w != null) {
out.add(new PlacementCandidate(name, null, w.weight(), w.maxLoad()));
}
}
return out;
}
@Override
public String effectiveCwd(SpawnRequest req) {
return route(req.profileName()).effectiveCwd(req);
}
@Override
public List<String> parityOverlay(String profileName) {
return route(profileName).parityOverlay(profileName);
}
@Override
public void stop(String id) {
HerdrPeerLauncher d = spawnedBy.remove(id);
if (d == null) {
log.debug("stop({}) — no recorded owner, routing to the first adapter (pane-addressed)", id);
d = delegates.getFirst();
}
d.stop(id);
}
@Override
public boolean clearContext(String id) {
HerdrPeerLauncher delegate = spawnedBy.get(id);
if (delegate == null) {
log.debug("clearContext({}) ignored — no recorded owning adapter", id);
return false;
}
return delegate.clearContext(id);
}
@Override
public Set<String> profiles() {
return byProfile.keySet();
}
@Override
public String defaultProfile() {
return defaultProfile;
}
/** Every herdr agent, deduplicated by pane id (all delegates share one herdr and list globally). */
@Override
public List<Agent> list() {
Map<String, Agent> byPane = new LinkedHashMap<>();
for (HerdrPeerLauncher d : delegates) {
for (Agent a : d.list()) {
if (a.paneId() != null) {
byPane.putIfAbsent(a.paneId(), a);
}
}
}
return List.copyOf(byPane.values());
}
@Override
public int reapOrphanWorkers() {
int reaped = 0;
for (HerdrPeerLauncher d : delegates) {
reaped += d.reapOrphanWorkers();
}
return reaped;
}
/** The union of every adapter's capabilities — a capability any adapter offers, the fleet offers. */
@Override
public Set<Capability> capabilities() {
EnumSet<Capability> caps = EnumSet.noneOf(Capability.class);
for (HerdrPeerLauncher d : delegates) {
caps.addAll(d.capabilities());
}
return Set.copyOf(caps);
}
}
@@ -1,4 +1,4 @@
package dev.ltms.bridged.worker;
package dev.ltms.bridged.member;
import dev.ltms.bridged.config.BridgedConfig;
import dev.ltms.bridged.herdr.Agent;
@@ -7,6 +7,7 @@ import dev.ltms.bridged.herdr.HerdrException;
import dev.ltms.bridged.herdr.Tab;
import dev.ltms.bridged.herdr.Workspace;
import dev.ltms.bridged.herdr.WorkspaceControl;
import dev.ltms.bridged.peer.MemberRole;
import dev.ltms.bridged.peer.PeerHandle;
import dev.ltms.bridged.peer.PeerLauncher;
import dev.ltms.bridged.peer.PeerUnreachableException;
@@ -21,9 +22,14 @@ import java.util.LinkedHashMap;
import java.util.List;
import java.util.Map;
import java.util.Set;
import java.util.UUID;
import java.util.concurrent.ConcurrentHashMap;
import java.util.concurrent.ConcurrentMap;
import java.util.concurrent.atomic.AtomicBoolean;
import java.util.concurrent.atomic.AtomicLong;
import java.util.function.Function;
import java.util.function.LongSupplier;
import java.util.function.Supplier;
import java.util.regex.Matcher;
import java.util.regex.Pattern;
@@ -38,7 +44,7 @@ import java.util.regex.Pattern;
* <li>{@code namePrefix} (constructor arg) — the label prefix ({@code claude}, {@code opencode})
* that drives both unique naming and the orphan-reap pattern, so each adapter reaps only its
* own kind of pane and never another's.</li>
* <li>{@link #buildLaunch(BridgedConfig.Worker)} — the peer-specific env map + argv, including any
* <li>{@link #buildLaunch(BridgedConfig.Profile)} — the peer-specific env map + argv, including any
* subscription/guard check, MCP mount, and instruction injection. The base never sees how the
* peer is configured; it only places and starts the returned {@link Launch}.</li>
* </ul>
@@ -55,16 +61,48 @@ public abstract class HerdrPeerLauncher implements PeerLauncher {
/** herdr rejects a duplicate agent {@code name}; we retry a bumped name this many times. */
private static final int NAME_RETRIES = 8;
/**
* Retries for {@code agent.start} against a seed pane whose shell has not reached its prompt
* yet — {@code tab.create}/{@code pane.split} return as soon as the pane exists, and herdr
* refuses to start an agent in a pane that is not "an available shell" ({@code agent_pane_busy}).
*/
private static final int SHELL_READY_RETRIES = 20;
private final String namePrefix; // label prefix: naming + reap scheme
private final AgentControl agents;
private final WorkspaceControl spaces;
private final Map<String, BridgedConfig.Worker> profiles; // profile name → spawn settings
private final Map<String, BridgedConfig.Profile> profiles; // profile name → spawn settings
private final String defaultProfile; // profile a no-arg spawn uses (nullable)
/** Host env lookup (injectable for tests); adapters read it in {@link #buildLaunch}. */
protected final Function<String, String> env;
private final AtomicLong nameSeq = new AtomicLong(); // per-peer counter (also the tab #)
private final AtomicLong nameSeq = new AtomicLong(); // per-peer counter (herdr agent names only)
/**
* The {@code fleet.tabLabel} template; a {@code null} supplier or a {@code null}/blank value ⇒
* {@link BridgedConfig.Fleet#DEFAULT_TAB_LABEL}. A profile's own {@code tabLabel} still
* overrides it.
*
* <p>CB-559: a supplier rather than a String, so a config reload renames the <em>next</em> tab
* without a restart. Existing tabs keep the label they were given — bridged does not rewrite a
* label it already wrote.
*/
private final Supplier<String> tabLabelTemplate;
/**
* Tab numbers, counted per {@code role/profile} pair (CB-557).
*
* <p>Deliberately not {@link #nameSeq}. That counter is shared by every profile this launcher
* serves, because its job is to make herdr <em>agent names</em> unique. Reusing it for the tab
* label made the numbers global, so sibling tabs read {@code #4}, {@code #9}, {@code #17} — gaps
* that look like a member died. Counting per role+profile makes {@code dev: sonnet #2} mean the
* second sonnet dev, which is what a reader assumes it means.
*
* <p>Resets when the daemon restarts, and that is fine: the label is a human-facing hint, not an
* identity. Identity is {@link PeerHandle#id()}.
*/
private final ConcurrentMap<String, AtomicLong> labelSeq = new ConcurrentHashMap<>();
private final long spawnReadyTimeoutMs; // 0 = disable gate (legacy non-blocking spawn)
private final LongSupplier nowMillis; // monotonic clock (injectable for tests)
@@ -74,6 +112,14 @@ public abstract class HerdrPeerLauncher implements PeerLauncher {
// collide with same-profile peers that outlived a restart. See startUniquelyNamed.
private final String nameNonce = String.format("%06x", new SecureRandom().nextInt(1 << 24));
// CB-519: PeerHandle.id() is a host-unique opaque UUID, decoupled from the herdr pane id. The
// routing/registry key is the UUID; the herdr pane id is a launcher-private placement/teardown
// coordinate. This map bridges the two so stop(id) can resolve a host-unique key back to the
// exact pane it must tear down. The pane id is launcher-private (never the routing key) — see
// PeerHandle.id().
private final ConcurrentMap<String, String> paneByAgentId = new ConcurrentHashMap<>();
private final AtomicBoolean resetUnsupportedLogged = new AtomicBoolean();
/**
* @param namePrefix label prefix for this peer kind (drives naming and reap)
* @param agents herdr agent control (start, status, close)
@@ -87,10 +133,30 @@ public abstract class HerdrPeerLauncher implements PeerLauncher {
* interval is baked into this hook, so the base needs no poll field
*/
protected HerdrPeerLauncher(String namePrefix, AgentControl agents, WorkspaceControl spaces,
Map<String, BridgedConfig.Worker> profiles, String defaultProfile,
Map<String, BridgedConfig.Profile> profiles, String defaultProfile,
Function<String, String> env,
long spawnReadyTimeoutMs,
LongSupplier nowMillis, Runnable sleeper) {
this(namePrefix, agents, spaces, profiles, defaultProfile, env, spawnReadyTimeoutMs,
nowMillis, sleeper, null);
}
/**
* As above, plus the {@code fleet.tabLabel} template (CB-557).
*
* @param tabLabelTemplate fleet-wide tab-label template, read per spawn (CB-559); {@code null},
* or a supplier yielding {@code null}/blank ⇒
* {@link BridgedConfig.Fleet#DEFAULT_TAB_LABEL}. A separate constructor
* rather than a new parameter on the one above, so every existing call
* site keeps the default without an edit.
*/
protected HerdrPeerLauncher(String namePrefix, AgentControl agents, WorkspaceControl spaces,
Map<String, BridgedConfig.Profile> profiles, String defaultProfile,
Function<String, String> env,
long spawnReadyTimeoutMs,
LongSupplier nowMillis, Runnable sleeper,
Supplier<String> tabLabelTemplate) {
this.tabLabelTemplate = tabLabelTemplate;
this.namePrefix = namePrefix;
this.agents = agents;
this.spaces = spaces;
@@ -109,10 +175,54 @@ public abstract class HerdrPeerLauncher implements PeerLauncher {
* Any subscription/guard check, MCP mount, and instruction injection happen here. The env map
* and argv are adapter-private; the base only places and starts what is returned.
*/
protected abstract Launch buildLaunch(BridgedConfig.Worker cfg);
protected abstract Launch buildLaunch(BridgedConfig.Profile cfg);
/** A peer-specific launch: the herdr {@code env} map and {@code argv}. */
protected record Launch(Map<String, String> env, List<String> argv) {
/**
* Session-aware variant of {@link #buildLaunch(BridgedConfig.Profile)} (CB-547a). Default
* discards the session identity and delegates to the profile-only form, so an adapter that
* carries no durable peer session (opencode, say) inherits byte-identical behaviour and needs
* no change. An adapter that does (Claude Code) overrides this to mint/resume the id and to
* surface it on the returned {@link Launch#agentSessionId()}.
*
* @param cfg the resolved profile to spawn
* @param sessionName the bridge's logical session name, or null/blank for launcher-derived
* @param resumeSessionId the peer's own prior session id to resume, or null/blank for fresh
*/
protected Launch buildLaunch(BridgedConfig.Profile cfg, String sessionName, String resumeSessionId) {
return buildLaunch(cfg);
}
/** Direct transport access for peer-specific, non-turn control operations. */
protected final AgentControl agents() {
return agents;
}
/** Resolve the public peer id to the launcher's private herdr target. */
protected final String agentTarget(String id) {
return paneByAgentId.get(id);
}
@Override
public boolean clearContext(String id) {
if (resetUnsupportedLogged.compareAndSet(false, true)) {
log.warn("context reset is unsupported for peer kind {}; clearAfterTurn is a no-op",
namePrefix);
}
return false;
}
/**
* A peer-specific launch: the herdr {@code env} map and {@code argv}, plus — for an adapter
* that carries durable session identity (CB-547a) — the peer's OWN session id
* ({@link PeerHandle#agentSessionId()}), known before the peer has written anything. Null for
* a launch that carries no identity.
*/
protected record Launch(Map<String, String> env, List<String> argv, String agentSessionId) {
/** A launch without a discoverable agent session id (an adapter that carries none). */
Launch(Map<String, String> env, List<String> argv) {
this(env, argv, null);
}
}
// --- profile surface -----------------------------------------------------------------------
@@ -130,7 +240,7 @@ public abstract class HerdrPeerLauncher implements PeerLauncher {
if (name == null || name.isBlank()) {
return List.of();
}
BridgedConfig.Worker cfg = profiles.get(name);
BridgedConfig.Profile cfg = profiles.get(name);
return cfg == null ? List.of() : cfg.parityOverlay();
}
@@ -141,18 +251,18 @@ public abstract class HerdrPeerLauncher implements PeerLauncher {
}
/** The configured profiles, for adapter capability decisions (e.g. any git-token grant). */
protected Collection<BridgedConfig.Worker> profileConfigs() {
protected Collection<BridgedConfig.Profile> profileConfigs() {
return profiles.values();
}
/** Resolve {@code profileName} (null/blank → default) to its config, or throw with the options. */
protected BridgedConfig.Worker requireProfile(String profileName) {
protected BridgedConfig.Profile requireProfile(String profileName) {
String name = (profileName == null || profileName.isBlank()) ? defaultProfile : profileName;
if (name == null || name.isBlank()) {
throw new IllegalArgumentException("no default worker profile is configured — "
+ "pass a profile; configured: " + profiles.keySet());
}
BridgedConfig.Worker cfg = profiles.get(name);
BridgedConfig.Profile cfg = profiles.get(name);
if (cfg == null) {
throw new IllegalArgumentException("unknown worker profile '" + name
+ "' — configured: " + profiles.keySet());
@@ -162,6 +272,10 @@ public abstract class HerdrPeerLauncher implements PeerLauncher {
// --- spawn ---------------------------------------------------------------------------------
/** A started peer plus the launch's agent-session id (the resume handle, or null). */
private record Spawned(Agent agent, String agentSessionId) {
}
/**
* Spawn a peer. {@code profileName} null/blank → the default profile. The working directory
* (CB-112) is resolved by {@link #resolveCwd}: an explicit {@code requestedCwd}, else the
@@ -170,31 +284,70 @@ public abstract class HerdrPeerLauncher implements PeerLauncher {
* The adapter's {@link #buildLaunch} runs before any herdr call.
*/
protected Agent spawnInternal(String profileName, String requestedCwd, String callerCwd) {
BridgedConfig.Worker cfg = requireProfile(profileName);
Launch launch = buildLaunch(cfg);
return spawnInternal(profileName, requestedCwd, callerCwd, null, null).agent();
}
/** Pre-CB-557 shape: no explicit role, so the tab is labelled as a {@code dev}. */
protected Spawned spawnInternal(String profileName, String requestedCwd, String callerCwd,
String sessionName, String resumeSessionId) {
return spawnInternal(profileName, requestedCwd, callerCwd, sessionName, resumeSessionId,
MemberRole.DEV);
}
/**
* Spawn a peer with session identity (CB-547a). {@code sessionName} and {@code resumeSessionId}
* are threaded from the {@link SpawnRequest} into {@link #buildLaunch(BridgedConfig.Profile,
* String, String)}, and the launch's resolved agent-session id is returned alongside the agent
* so the caller can put it on the {@link PeerHandle}.
*/
protected Spawned spawnInternal(String profileName, String requestedCwd, String callerCwd,
String sessionName, String resumeSessionId, MemberRole role) {
BridgedConfig.Profile cfg = requireProfile(profileName);
Launch launch = buildLaunch(cfg, sessionName, resumeSessionId);
String cwd = resolveCwd(requestedCwd, cfg, callerCwd);
return cfg.tabPlacement()
? spawnInTab(cfg, launch.env(), launch.argv(), cwd)
Agent agent = cfg.tabPlacement()
? spawnInTab(cfg, launch.env(), launch.argv(), cwd, role)
: spawnAsPane(cfg, launch.env(), launch.argv(), cwd);
return new Spawned(agent, launch.agentSessionId());
}
/**
* The next tab number for {@code role} on {@code profile}, starting at 1.
*
* <p>Starts at 1 rather than 0 because the number is read by a person: {@code "dev: sonnet #1"}
* is the first one, and {@code #0} invites the question of where {@code #1} went.
*/
private long nextLabelSeq(MemberRole role, String profile) {
String key = (role == null ? "" : role.wireName()) + "/" + profile;
return labelSeq.computeIfAbsent(key, _ -> new AtomicLong()).incrementAndGet();
}
/**
* {@inheritDoc}
*
* <p>Delegates to {@link #spawnInternal} and wraps the resulting herdr {@link Agent} in a
* {@link WorkerHandle} whose {@link PeerHandle#id()} equals the agent's paneId. When
* {@code spawnReadyTimeoutMs > 0}, blocks until the peer's herdr status is injectable or the
* timeout elapses; on timeout the pane is closed (no orphan) and a
* {@link PeerUnreachableException} is thrown.
* {@link WorkerHandle} whose {@link PeerHandle#id()} is a fresh <em>host-unique</em> opaque
* UUID (CB-519), deliberately decoupled from the herdr pane id: the id is the registry/routing
* key and must never collide across daemon processes on the same host, while the herdr pane id
* stays a launcher-private placement/teardown coordinate, remembered here so {@link #stop}
* can resolve the host-unique key back to its pane. When {@code spawnReadyTimeoutMs > 0},
* blocks until the peer's herdr status is injectable or the timeout elapses; on timeout the
* pane is closed (no orphan) and a {@link PeerUnreachableException} is thrown.
*/
@Override
public PeerHandle spawn(SpawnRequest req) {
Agent agent = spawnInternal(req.profileName(), req.requestedCwd(), req.callerCwd());
Spawned spawned = spawnInternal(req.profileName(), req.requestedCwd(), req.callerCwd(),
req.sessionName(), req.resumeSessionId(), req.role());
Agent agent = spawned.agent();
String paneId = agent.paneId();
if (spawnReadyTimeoutMs > 0) {
waitUntilInjectableOrThrow(paneId);
}
return new WorkerHandle(paneId, agent.terminalId());
// CB-519: the handle id is a host-unique UUID; the herdr pane it maps to stays internal.
String id = UUID.randomUUID().toString();
paneByAgentId.put(id, paneId);
return new WorkerHandle(id, agent.terminalId(), requireProfile(req.profileName()).profile(),
req.sessionName(), spawned.agentSessionId());
}
@Override
@@ -216,7 +369,7 @@ public abstract class HerdrPeerLauncher implements PeerLauncher {
* guaranteed last resort so a pathological environment with an unset {@code user.dir} still
* honours the "never assume {@code $HOME}" contract rather than letting herdr default the pane.
*/
private static String resolveCwd(String requestedCwd, BridgedConfig.Worker cfg, String callerCwd) {
private static String resolveCwd(String requestedCwd, BridgedConfig.Profile cfg, String callerCwd) {
return firstNonBlank(requestedCwd, cfg.cwd(), callerCwd, System.getProperty("user.dir"), ".");
}
@@ -227,17 +380,23 @@ public abstract class HerdrPeerLauncher implements PeerLauncher {
return null;
}
/** Dedicated worker space → own tab → start the peer (rooted at {@code cwd}) → drop the shell. */
private Agent spawnInTab(BridgedConfig.Worker cfg, Map<String, String> workerEnv,
List<String> argv, String cwd) {
/** Dedicated worker space → own tab (carrying cwd+env) → start the peer into the seed pane. */
private Agent spawnInTab(BridgedConfig.Profile cfg, Map<String, String> workerEnv,
List<String> argv, String cwd, MemberRole role) {
Workspace space = spaces.ensureWorkspace(cfg.workspace());
Tab.Created tab = spaces.createTab(space.workspaceId());
Tab.Created tab = spaces.createTab(space.workspaceId(), cwd, workerEnv);
log.info("spawning {} profile={} space={} tab={} cwd={}",
namePrefix, cfg.profile(), space.workspaceId(), tab.tab().tabId(), cwd);
Started started;
try {
started = startUniquelyNamed(cfg, workerEnv, argv, tab.tab().tabId(), cwd);
if (tab.rootPaneId() == null) {
// Protocol 19 starts the agent INTO the seed pane — without one there is nowhere
// to start, and a partial tab would be left behind.
throw new IllegalStateException("tab " + tab.tab().tabId()
+ " had no seed pane in the create response — cannot start a peer in it");
}
started = startUniquelyNamed(cfg, argv, tab.rootPaneId());
} catch (RuntimeException e) {
// The peer never started — don't leave the tab we just created orphaned.
// Best-effort cleanup; never let it mask the real spawn failure.
@@ -250,18 +409,14 @@ public abstract class HerdrPeerLauncher implements PeerLauncher {
throw e;
}
// The peer is LIVE now. The remaining steps are cosmetic (drop herdr's seed shell so the
// tab holds only the peer; label the tab). They must not fail the spawn or orphan the
// running peer — on error we log and still return it so the caller gets its paneId and can
// tear it down.
if (tab.rootPaneId() != null) {
tidy("close seed pane " + tab.rootPaneId(), () -> agents.close(tab.rootPaneId()));
} else {
log.warn("tab {} had no seed pane in the create response; peer tab may hold an extra pane",
tab.tab().tabId());
}
// The peer is LIVE now, in the seed pane itself (no shell pane to drop — protocol 19).
// Labelling is cosmetic: it must not fail the spawn or orphan the running peer — on error
// we log and still return it so the caller gets its paneId and can tear it down.
tidy("label tab " + tab.tab().tabId(),
() -> spaces.renameTab(tab.tab().tabId(), cfg.renderTabLabel(started.seq())));
() -> spaces.renameTab(tab.tab().tabId(),
cfg.renderTabLabel(
tabLabelTemplate == null ? null : tabLabelTemplate.get(),
role, nextLabelSeq(role, cfg.profile()))));
log.info("{} started pane={} tab={} terminal={}",
namePrefix, started.agent().paneId(), started.agent().tabId(), started.agent().terminalId());
return started.agent();
@@ -276,12 +431,16 @@ public abstract class HerdrPeerLauncher implements PeerLauncher {
}
}
/** Legacy placement: herdr splits the currently-focused tab; the peer still starts in {@code cwd}. */
private Agent spawnAsPane(BridgedConfig.Worker cfg, Map<String, String> workerEnv,
/** Legacy placement: split the currently-focused tab; the peer still starts in {@code cwd}. */
private Agent spawnAsPane(BridgedConfig.Profile cfg, Map<String, String> workerEnv,
List<String> argv, String cwd) {
log.info("spawning {} (pane placement) profile={} cwd={} argv={}",
namePrefix, cfg.profile(), cwd, argv);
Agent peer = startUniquelyNamed(cfg, workerEnv, argv, null, cwd).agent();
String paneId = spaces.splitPane(cwd, workerEnv);
if (paneId == null) {
throw new IllegalStateException("pane.split returned no pane — cannot start a peer");
}
Agent peer = startUniquelyNamed(cfg, argv, paneId).agent();
log.info("{} started pane={} terminal={}", namePrefix, peer.paneId(), peer.terminalId());
return peer;
}
@@ -300,14 +459,16 @@ public abstract class HerdrPeerLauncher implements PeerLauncher {
* backstop for the astronomically unlikely nonce+seq clash; the name is a label only — herdr
* detects kind and status from terminal output, not from it.
*/
private Started startUniquelyNamed(BridgedConfig.Worker cfg, Map<String, String> workerEnv,
List<String> argv, String tabId, String cwd) {
private Started startUniquelyNamed(BridgedConfig.Profile cfg, List<String> argv, String paneId) {
// Protocol 19 resolves the executable from the agent kind (== namePrefix here), so
// argv[0] — the configured executable — is dropped and only the extra args are passed.
List<String> args = argv.isEmpty() ? argv : argv.subList(1, argv.size());
HerdrException last = null;
for (int attempt = 0; attempt < NAME_RETRIES; attempt++) {
long seq = nameSeq.incrementAndGet();
String name = namePrefix + "-" + cfg.profile() + "-" + nameNonce + "-" + seq;
try {
return new Started(agents.start(name, argv, workerEnv, tabId, cwd), seq);
return new Started(startAwaitingShellPrompt(name, args, paneId), seq);
} catch (HerdrException e) {
if (!"agent_name_taken".equals(e.code())) throw e;
log.debug("peer name '{}' taken, retrying", name);
@@ -317,6 +478,22 @@ public abstract class HerdrPeerLauncher implements PeerLauncher {
throw last;
}
/** Start the agent into {@code paneId}, waiting out the seed shell's boot with the sleeper. */
private Agent startAwaitingShellPrompt(String name, List<String> args, String paneId) {
HerdrException busy = null;
for (int attempt = 0; attempt < SHELL_READY_RETRIES; attempt++) {
try {
return agents.start(name, namePrefix, args, paneId);
} catch (HerdrException e) {
if (!"agent_pane_busy".equals(e.code())) throw e;
log.debug("pane {} not at its shell prompt yet, retrying agent.start", paneId);
busy = e;
sleeper.run();
}
}
throw busy;
}
// --- discovery + reap ----------------------------------------------------------------------
/** All herdr-tracked agents — discovery for "what peers exist". */
@@ -398,21 +575,31 @@ public abstract class HerdrPeerLauncher implements PeerLauncher {
// --- teardown ------------------------------------------------------------------------------
/**
* Tear a peer down by pane id: close the pane, and close its tab <em>only</em> when the peer is
* that tab's sole occupant. The single-pane check is what makes this safe regardless of how the
* peer was placed (or a placement-config change across a restart): a pane-placement peer sitting
* in one of the user's shared tabs has siblings, so its tab is never closed — we only ever
* remove a tab we created to hold one peer.
* Tear a peer down: close the pane, and close its tab <em>only</em> when the peer is that tab's
* sole occupant. The single-pane check is what makes this safe regardless of how the peer was
* placed (or a placement-config change across a restart): a pane-placement peer sitting in one
* of the user's shared tabs has siblings, so its tab is never closed — we only ever remove a
* tab we created to hold one peer.
*
* <p>{@code idOrPane} is the {@link PeerHandle#id()} of a peer this launcher spawned (CB-519's
* host-unique opaque UUID), resolved through {@link #paneByAgentId} to the pane it must tear
* down. An argument that is not one of our ids is treated as a raw herdr pane id — the
* {@link #reapOrphanWorkers() orphan-reap} and spawn-gate-timeout paths, plus any caller that
* passes a pane directly, keep working without an owning id.
*
* <p>Resolves the tab from the pane <em>before</em> closing it. An already-gone pane/tab
* (repeated DELETE, crashed peer) is treated as success; any other failure propagates so a
* genuinely failed teardown is not reported as done.
*/
@Override
public void stop(String paneId) {
// Teardown knows only the paneId, not which profile spawned it. Attempt tab cleanup when any
public void stop(String idOrPane) {
// Teardown knows only the pane, not which profile spawned it. Attempt tab cleanup when any
// profile uses tab placement (so the bridge may have created a dedicated peer tab); the
// single-occupant check below is what actually protects the user's shared tabs.
String paneId = paneByAgentId.remove(idOrPane);
if (paneId == null) {
paneId = idOrPane; // raw-pane fallback (reap, gate timeout, pane-addressed callers)
}
WorkspaceControl.PaneLocation loc = usesTabPlacement() ? spaces.locatePane(paneId) : null;
try {
agents.close(paneId);
@@ -430,7 +617,7 @@ public abstract class HerdrPeerLauncher implements PeerLauncher {
/** Whether any configured profile places peers in their own tab (so tabs may need cleanup). */
private boolean usesTabPlacement() {
return profiles.values().stream().anyMatch(BridgedConfig.Worker::tabPlacement);
return profiles.values().stream().anyMatch(BridgedConfig.Profile::tabPlacement);
}
/** True when a herdr error means the target is already gone (safe to treat as done). */
@@ -460,8 +647,13 @@ public abstract class HerdrPeerLauncher implements PeerLauncher {
+ spawnReadyTimeoutMs + "ms");
}
/** A concrete {@link PeerHandle} wrapping herdr agent coordinates. */
private record WorkerHandle(String id, String terminalId) implements PeerHandle {
/**
* A concrete {@link PeerHandle} wrapping herdr agent coordinates, the profile that spawned it,
* and the session identity the launch resolved (CB-547a): the bridge's logical name and the
* peer's own session id, both null when the spawn carried no identity.
*/
private record WorkerHandle(String id, String terminalId, String profile,
String sessionName, String agentSessionId) implements PeerHandle {
}
// --- shared helpers ------------------------------------------------------------------------
@@ -484,7 +676,7 @@ public abstract class HerdrPeerLauncher implements PeerLauncher {
* {@code GITEA_HOST}. Push over SSH is unaffected; the only incremental grant is PR-create.
* Peer-neutral, so every herdr adapter reuses it unchanged.
*/
protected void applyGitToken(Map<String, String> workerEnv, BridgedConfig.Worker cfg) {
protected void applyGitToken(Map<String, String> workerEnv, BridgedConfig.Profile cfg) {
if (!cfg.hasGitToken()) {
return;
}
@@ -513,7 +705,7 @@ public abstract class HerdrPeerLauncher implements PeerLauncher {
* dev.ltms.bridged.guard.SubscriptionGuard}, which is checked against the profile's
* {@code baseUrl} and nothing else.
*/
protected Map<String, String> baseEnv(BridgedConfig.Worker cfg) {
protected Map<String, String> baseEnv(BridgedConfig.Profile cfg) {
Map<String, String> workerEnv = new LinkedHashMap<>();
String path = env.apply("PATH");
if (path != null && !path.isBlank()) {
@@ -1,4 +1,4 @@
package dev.ltms.bridged.worker;
package dev.ltms.bridged.member;
import com.fasterxml.jackson.databind.ObjectMapper;
import com.fasterxml.jackson.databind.node.ObjectNode;
@@ -7,6 +7,8 @@ import dev.ltms.bridged.herdr.Agent;
import dev.ltms.bridged.herdr.AgentControl;
import dev.ltms.bridged.herdr.WorkspaceControl;
import dev.ltms.bridged.peer.Capability;
import dev.ltms.bridged.peer.PeerHandle;
import dev.ltms.bridged.peer.SpawnRequest;
import java.io.IOException;
import java.io.UncheckedIOException;
@@ -18,6 +20,7 @@ import java.util.Map;
import java.util.Set;
import java.util.function.Function;
import java.util.function.LongSupplier;
import java.util.function.Supplier;
/**
* The {@link HerdrPeerLauncher} adapter for <strong>opencode</strong> — an open-source,
@@ -72,16 +75,34 @@ public final class OpenCodeLauncher extends HerdrPeerLauncher {
/** Root under which per-spawn opencode config dirs are created (injectable for tests). */
private final Path configRoot;
/**
* The current spawn's resume-target session id, threaded from {@link #spawn(SpawnRequest)} to
* {@link #buildLaunch} across the base's {@code spawn -> spawnInternal -> buildLaunch} chain,
* which carries no request. A plain field would race under concurrent spawns (the base supports
* them), so it is thread-local: each spawn captures its own request's id on its own thread, and
* {@code buildLaunch}, synchronous and same-thread, reads exactly that one. Set only around the
* {@code super.spawn} call and cleared in {@code finally}, so a paused/leftover value can never
* bleed into the next spawn.
*/
private final ThreadLocal<String> resumeSessionId = new ThreadLocal<>();
/**
* Session discovery against opencode's on-disk storage ({@link OpenCodeSessionDiscovery}) —
* the one seam that knows opencode's private session-file layout. Its root is injectable for
* tests so they never touch the operator's real {@code ~/.local/share/opencode}.
*/
private final OpenCodeSessionDiscovery discovery;
/**
* Production constructor — disables the spawn-ready gate ({@code spawnReadyTimeoutMs == 0}) so it
* matches the legacy non-blocking spawn semantics. Config dirs are created under the JVM temp dir.
*/
public OpenCodeLauncher(AgentControl agents, WorkspaceControl spaces,
Map<String, BridgedConfig.Worker> profiles, String defaultProfile,
Map<String, BridgedConfig.Profile> profiles, String defaultProfile,
Function<String, String> env) {
this(agents, spaces, profiles, defaultProfile, env, 0,
System::currentTimeMillis, () -> sleepUninterruptibly(300),
defaultConfigRoot());
defaultConfigRoot(), defaultDiscoveryRoot());
}
/**
@@ -89,12 +110,26 @@ public final class OpenCodeLauncher extends HerdrPeerLauncher {
* the pane reports an injectable state or {@code spawnReadyTimeoutMs} elapses.
*/
public OpenCodeLauncher(AgentControl agents, WorkspaceControl spaces,
Map<String, BridgedConfig.Worker> profiles, String defaultProfile,
Map<String, BridgedConfig.Profile> profiles, String defaultProfile,
Function<String, String> env,
long spawnReadyTimeoutMs, long spawnReadyPollMs) {
this(agents, spaces, profiles, defaultProfile, env,
spawnReadyTimeoutMs, spawnReadyPollMs, null);
}
/**
* Production constructor carrying the fleet-wide tab-label template (CB-557). The template comes
* from {@code fleet.tabLabel}, which a profile cannot know because it names the member's
* <em>role</em>; a profile may still override it with its own {@code tabLabel}.
*/
public OpenCodeLauncher(AgentControl agents, WorkspaceControl spaces,
Map<String, BridgedConfig.Profile> profiles, String defaultProfile,
Function<String, String> env,
long spawnReadyTimeoutMs, long spawnReadyPollMs,
Supplier<String> tabLabelTemplate) {
this(agents, spaces, profiles, defaultProfile, env, spawnReadyTimeoutMs,
System::currentTimeMillis, () -> sleepUninterruptibly(spawnReadyPollMs),
defaultConfigRoot());
defaultConfigRoot(), defaultDiscoveryRoot(), tabLabelTemplate);
}
/**
@@ -112,21 +147,48 @@ public final class OpenCodeLauncher extends HerdrPeerLauncher {
* @param sleeper sleep/wait hook (encodes the poll interval; never called when the
* gate is disabled)
* @param configRoot existing directory under which per-spawn config dirs are created
* @param discoveryRoot opencode's on-disk storage root to scan for session records
* (injectable for tests; opencode's layout is matched at
* {@link OpenCodeSessionDiscovery})
*/
public OpenCodeLauncher(AgentControl agents, WorkspaceControl spaces,
Map<String, BridgedConfig.Worker> profiles, String defaultProfile,
Map<String, BridgedConfig.Profile> profiles, String defaultProfile,
Function<String, String> env,
long spawnReadyTimeoutMs,
LongSupplier nowMillis, Runnable sleeper, Path configRoot) {
LongSupplier nowMillis, Runnable sleeper,
Path configRoot, Path discoveryRoot) {
this(agents, spaces, profiles, defaultProfile, env, spawnReadyTimeoutMs,
nowMillis, sleeper, configRoot, discoveryRoot, null);
}
/**
* Full testability constructor, plus the fleet-wide tab-label template (CB-557).
*
* @param tabLabelTemplate {@code fleet.tabLabel}; {@code null}/blank ⇒
* {@link BridgedConfig.Fleet#DEFAULT_TAB_LABEL}
*/
public OpenCodeLauncher(AgentControl agents, WorkspaceControl spaces,
Map<String, BridgedConfig.Profile> profiles, String defaultProfile,
Function<String, String> env,
long spawnReadyTimeoutMs,
LongSupplier nowMillis, Runnable sleeper,
Path configRoot, Path discoveryRoot,
Supplier<String> tabLabelTemplate) {
super(NAME_PREFIX, agents, spaces, profiles, defaultProfile, env,
spawnReadyTimeoutMs, nowMillis, sleeper);
spawnReadyTimeoutMs, nowMillis, sleeper, tabLabelTemplate);
this.configRoot = configRoot;
this.discovery = new OpenCodeSessionDiscovery(discoveryRoot);
}
private static Path defaultConfigRoot() {
return Path.of(System.getProperty("java.io.tmpdir"));
}
/** The default opencode storage root: {@code ~/.local/share/opencode} (the XDG data dir). */
private static Path defaultDiscoveryRoot() {
return Path.of(System.getProperty("user.home"), ".local", "share", "opencode");
}
/**
* {@inheritDoc}
*
@@ -137,14 +199,14 @@ public final class OpenCodeLauncher extends HerdrPeerLauncher {
* with {@code -m}.
*/
@Override
protected Launch buildLaunch(BridgedConfig.Worker cfg) {
protected Launch buildLaunch(BridgedConfig.Profile cfg) {
Map<String, String> workerEnv = baseEnv(cfg);
// A config file is needed for the bridge MCP mount, for a pinned endpoint (CB-508), or both.
if (cfg.hasMcp() || hasCustomProvider(cfg)) {
workerEnv.put("OPENCODE_CONFIG", writeConfig(cfg).toString());
}
applyGitToken(workerEnv, cfg);
return new Launch(workerEnv, argvWithModel(cfg));
return new Launch(workerEnv, argvWithResume(argvWithModel(argvWithAuto(cfg), cfg)));
}
/**
@@ -157,13 +219,46 @@ public final class OpenCodeLauncher extends HerdrPeerLauncher {
* primary's Anthropic subscription, and an opencode process has no Anthropic credential path
* at all. Pointing it at a local vLLM cannot leak the subscription.
*/
private static boolean hasCustomProvider(BridgedConfig.Worker cfg) {
private static boolean hasCustomProvider(BridgedConfig.Profile cfg) {
return cfg.baseUrl() != null && !cfg.baseUrl().isBlank();
}
/** The launch argv plus, when a model is configured, the opencode {@code -m provider/model} flag. */
private List<String> argvWithModel(BridgedConfig.Worker cfg) {
/**
* The launch argv plus the unconditional {@code --auto} flag, which auto-approves the
* permissions opencode does not explicitly deny. It is unconditional, not a preference: a
* spawned peer has no human at its pane — the bridge spawned it — so one that stops at an
* approval prompt is a wedged agent, indistinguishable from a legitimate mid-turn wait and
* unable to end its turn with {@code bridge_reply}. opencode's own help calls this
* "dangerous!", but the blast radius here is already bounded by design: a worker runs in its
* own git worktree on its own branch, is off-subscription, and cannot merge — the lead is the
* gate.
*/
private List<String> argvWithAuto(BridgedConfig.Profile cfg) {
List<String> argv = mutableArgv(cfg.argv());
argv.add("--auto");
return argv;
}
/**
* The launch argv plus, on a resumed spawn, opencode's {@code -s <id>} flag to continue a prior
* conversation by its session id. {@code -s, --session <id>} resumes an existing session; on a
* fresh spawn (no resume target) no flag is added, letting opencode start a brand-new session.
* The id comes from the current spawn request's {@code resumeSessionId}, threaded per-thread by
* {@link #spawn(SpawnRequest)}.
*/
private List<String> argvWithResume(List<String> argv) {
String id = resumeSessionId.get();
if (id == null || id.isBlank()) {
return argv;
}
List<String> withResume = mutableArgv(argv);
withResume.add("-s");
withResume.add(id);
return withResume;
}
/** The launch argv plus, when a model is configured, the opencode {@code -m provider/model} flag. */
private List<String> argvWithModel(List<String> argv, BridgedConfig.Profile cfg) {
if (cfg.model() != null && !cfg.model().isBlank()) {
argv.add("-m");
argv.add(cfg.model());
@@ -177,13 +272,28 @@ public final class OpenCodeLauncher extends HerdrPeerLauncher {
* {@code OPENCODE_CONFIG}. The dir is unique per spawn so concurrent workers never race on it;
* it is best-effort cleaned on JVM exit (worker config is disposable — regenerated every spawn).
*/
private Path writeConfig(BridgedConfig.Worker cfg) {
private Path writeConfig(BridgedConfig.Profile cfg) {
try {
Path dir = Files.createTempDirectory(configRoot, "bridged-opencode-");
dir.toFile().deleteOnExit();
ObjectNode root = JSON.createObjectNode();
root.put("$schema", "https://opencode.ai/config.json");
// CB-523, opencode side: a worker that runs out of context dies mid-turn, and its reply
// — the entire point of the turn — is lost with it. Auto-compaction is therefore not an
// operator preference for a bridged worker, it is a condition of the turn contract.
//
// Stated deliberately even though it is redundant today: OPENCODE_CONFIG is MERGED over
// ~/.config/opencode/config.json rather than replacing it, so a worker already inherits
// an `auto: true` set at home. We do not want that inheritance to be what the guarantee
// rests on — the home file is outside this repo, differs per machine, and is not ours.
//
// Know the cost before removing it: this key WINS over the home config (verified — an
// OPENCODE_CONFIG value overrides the home value, it does not defer to it), so an
// operator who sets `compaction.auto: false` at home cannot turn it off for bridged
// workers. That is the intended trade for peers we spawn and whose turns we must land;
// if per-profile control is ever wanted, add a profile knob rather than dropping this.
root.putObject("compaction").put("auto", true);
if (cfg.hasMcp()) {
Path charter = dir.resolve("reply-charter.md");
@@ -219,7 +329,7 @@ public final class OpenCodeLauncher extends HerdrPeerLauncher {
* <p>The provider id comes from the {@code provider/model} selector in {@code model:}, so one
* field drives both the declaration and the {@code -m} flag and they cannot drift apart.
*/
private void addCustomProvider(ObjectNode root, BridgedConfig.Worker cfg) {
private void addCustomProvider(ObjectNode root, BridgedConfig.Profile cfg) {
String[] parts = splitModelSelector(cfg);
String providerId = parts[0];
String modelId = parts[1];
@@ -242,7 +352,7 @@ public final class OpenCodeLauncher extends HerdrPeerLauncher {
* needs both, so a bare model name is rejected loudly rather than silently falling back to the
* default gateway — a worker quietly talking to the wrong endpoint is the failure this avoids.
*/
private static String[] splitModelSelector(BridgedConfig.Worker cfg) {
private static String[] splitModelSelector(BridgedConfig.Profile cfg) {
String model = cfg.model();
int slash = model == null ? -1 : model.indexOf('/');
if (model == null || model.isBlank() || slash <= 0 || slash == model.length() - 1) {
@@ -270,6 +380,82 @@ public final class OpenCodeLauncher extends HerdrPeerLauncher {
return afterScheme.contains("/") ? trimmed : trimmed + "/v1";
}
/**
* {@inheritDoc}
*
* <p>adds this adapter's session-identity work around the base's spawn — as opencode cannot be
* told its session id at spawn (see {@link Capability#SESSION_RESUME} vs
* {@link Capability#SESSION_NAME}), identity is only ever adopted after the fact:
* <ul>
* <li>the request's {@code resumeSessionId} is remembered for {@link #buildLaunch} to turn
* into {@code -s <id>}; and</li>
* <li>the returned handle is wrapped so its
* {@link dev.ltms.bridged.peer.PeerHandle#agentSessionId()} performs lazy session
* discovery against opencode's storage (see {@link OpenCodeSessionDiscovery}) — always
* non-blocking, {@code null} until opencode has persisted the session record.</li>
* </ul>
*/
@Override
public PeerHandle spawn(SpawnRequest req) {
resumeSessionId.set(req.resumeSessionId());
try {
PeerHandle inner = super.spawn(req);
return new SessionAwareHandle(inner, discovery, effectiveCwd(req));
} finally {
// Never let a paused/leftover resume id bleed into the next spawn on this thread.
resumeSessionId.remove();
}
}
/**
* A {@link PeerHandle} that delegates everything to the base's worker handle but resolves
* {@link #agentSessionId()} lazily through opencode session discovery. Delegate-only, so the
* base's id/terminalId/profile semantics (CB-519's host-unique routing key, herdr coordinates)
* are untouched — only the opencode-specific identity answer is added. {@code sessionName()}
* stays null: opencode has no display-name seam, so the logical name lives only in the bridge's
* roster (see the SESSION_NAME capability).
*/
private static final class SessionAwareHandle implements PeerHandle {
private final PeerHandle delegate;
private final OpenCodeSessionDiscovery discovery;
private final String cwd;
SessionAwareHandle(PeerHandle delegate, OpenCodeSessionDiscovery discovery, String cwd) {
this.delegate = delegate;
this.discovery = discovery;
this.cwd = cwd;
}
@Override
public String id() {
return delegate.id();
}
@Override
public String terminalId() {
return delegate.terminalId();
}
@Override
public String profile() {
return delegate.profile();
}
@Override
public String sessionName() {
return delegate.sessionName();
}
@Override
public String agentSessionId() {
// Lazy + retried, never a spawn-time blocker: opencode writes the session record only
// when the session is first persisted, so null here is the correct interim answer and
// the caller re-calls later (each call re-scans, picking up a record that has since
// appeared).
return discovery.sessionIdForDirectory(cwd);
}
}
// --- Agent-returning convenience spawns (used by callers/tests that want the herdr Agent) ---
/** Spawn a worker for the default profile in the resolved default cwd. */
@@ -291,16 +477,21 @@ public final class OpenCodeLauncher extends HerdrPeerLauncher {
@Override
public Set<Capability> capabilities() {
Set<Capability> caps = EnumSet.of(Capability.MID_TURN_ASK, Capability.WORKTREE, Capability.ORPHAN_REAP);
Set<Capability> caps = EnumSet.of(Capability.MID_TURN_ASK, Capability.WORKTREE,
Capability.ORPHAN_REAP, Capability.SESSION_RESUME);
if (hasGitTokenProfile()) {
caps.add(Capability.SELF_PR);
}
// Deliberately NOT SESSION_NAME: opencode has no display-name flag, so the bridge's logical
// name can't surface in the peer's own UI — declaring the capability would hide that
// asymmetry rather than make it honest. For opencode the name lives only in the bridge's
// roster (see PeerHandle.sessionName() returning null).
return Set.copyOf(caps);
}
/** Whether any configured profile opts into a git-forge token (required for {@link Capability#SELF_PR}). */
private boolean hasGitTokenProfile() {
return profileConfigs().stream().anyMatch(BridgedConfig.Worker::hasGitToken);
return profileConfigs().stream().anyMatch(BridgedConfig.Profile::hasGitToken);
}
// --- CB-117 reap predicate (opencode prefix), kept for direct unit testing -----------------
@@ -0,0 +1,122 @@
package dev.ltms.bridged.member;
import com.fasterxml.jackson.databind.JsonNode;
import com.fasterxml.jackson.databind.ObjectMapper;
import java.io.IOException;
import java.nio.file.Files;
import java.nio.file.Path;
import java.util.stream.Stream;
/**
* Resolves the opencode session id for a bridged worker from opencode's on-disk storage — the
* only place this adapter touches opencode's private layout, and deliberately the <em>only</em>
* class that does.
*
* <p><strong>Why this is isolated behind one seam.</strong> The layout is version-coupled and not a
* stable contract: opencode writes one JSON file per session under
* {@code <storageRoot>/session/<projectID>/<ses_*.json>}, and each record carries a
* {@code "version"} field (e.g. {@code "1.1.31"}), so the exact directory shape, file naming, and
* field names can move between opencode releases. opencode also ships a headless HTTP server that
* may supersede file scanning entirely. Everything this adapter knows about that private storage —
* its shape, naming, and field names — lives here, so a layout change, or a switch to the HTTP
* server, changes exactly one class and nothing in {@link OpenCodeLauncher}.
*
* <p>The determinism that makes this useful is structural, not a guess: every bridged worker runs
* in its own unique git worktree, so the record's {@code directory} (its project root) equals the
* worker's cwd identifies <em>its</em> session unambiguously. We match on {@code directory} rather
* than diffing {@code opencode session list} before/after — that races under concurrent spawns, and
* the CLI listing does not even show the directory.
*
* <p>All reads are best-effort and never throw: a missing or unreadable storage root, a record that
* fails to parse, or a directory with no record yet all yield {@code null}, and the caller (the
* session handle) treats that as "identity not resolved yet" and retries later.
*/
final class OpenCodeSessionDiscovery {
private final Path storageRoot; // e.g. ~/.local/share/opencode (injectable for tests)
private final ObjectMapper json;
OpenCodeSessionDiscovery(Path storageRoot) {
this.storageRoot = storageRoot;
this.json = new ObjectMapper();
}
/**
* The opencode session id whose record references {@code directory} (the worker's cwd), or
* {@code null} when no record matches yet. When several records share the directory — e.g.
* repeated spawns into the same worktree — the <em>most recently modified</em> one wins: it is
* the session the pane most likely corresponds to.
*
* <p>Never throws: a missing {@code storageRoot}, an unreadable/malformed record, or a
* directory that has not been persisted yet all resolve to {@code null} rather than failing a
* spawn. A bridged worker's session record is written lazily (when the session is first
* persisted), so {@code null} here is the normal answer right after the pane is ready, and the
* caller retries later.
*
* @param directory the worker's cwd, as resolved for this spawn
* @return the matching session id, or {@code null} if none is known yet
*/
String sessionIdForDirectory(String directory) {
if (directory == null || directory.isBlank()) {
return null;
}
Path sessionRoot = storageRoot.resolve("session");
if (!Files.isDirectory(sessionRoot)) {
return null;
}
String best = null;
long bestMtime = Long.MIN_VALUE;
try (Stream<Path> projectDirs = Files.list(sessionRoot)) {
for (Path projectDir : projectDirs.filter(Files::isDirectory).toList()) {
try (Stream<Path> records = Files.list(projectDir)) {
for (Path record : records.toList()) {
String id = matchId(record, directory);
if (id == null) {
continue;
}
long mtime = lastModifiedEpochMillis(record);
if (mtime > bestMtime) {
bestMtime = mtime;
best = id;
}
}
} catch (IOException ignored) {
// one project dir unreadable — skip it; another may still match
}
}
} catch (IOException ignored) {
// storage root vanished or became unreadable — "no session known yet"
return null;
}
return best;
}
/**
* The record's session id when it references {@code directory}, else {@code null}. A record
* that is not JSON, lacks {@code id}/{@code directory}, or points at a different directory is
* simply not our session; a malformed one is skipped, never fatal.
*/
private String matchId(Path record, String directory) {
try {
JsonNode node = json.readTree(record.toFile());
JsonNode id = node == null ? null : node.get("id");
JsonNode dir = node == null ? null : node.get("directory");
if (id == null || dir == null || !directory.equals(dir.asText())) {
return null;
}
return id.asText();
} catch (IOException e) {
return null;
}
}
/** The record's last-modified epoch ms, or {@code Long.MIN_VALUE} if unreadable (never wins). */
private static long lastModifiedEpochMillis(Path record) {
try {
return Files.getLastModifiedTime(record).toMillis();
} catch (IOException e) {
return Long.MIN_VALUE;
}
}
}
@@ -2,7 +2,7 @@ package dev.ltms.bridged.metrics;
import dev.ltms.bridged.msg.ReplyInbox;
import dev.ltms.bridged.session.SessionManager;
import dev.ltms.bridged.session.WorkerSession;
import dev.ltms.bridged.session.MemberSession;
import java.util.LinkedHashMap;
import java.util.Map;
@@ -26,6 +26,8 @@ public final class BridgedMetrics {
public static final String REPLIES = "bridged_replies_total";
/** Counter: push-loop nudges to the primary, by outcome. */
public static final String PUSH_NUDGES = "bridged_push_nudges_total";
/** Counter: idle-lead heartbeat nudges to the lead, by outcome (CB-551). */
public static final String HEARTBEAT_NUDGES = "bridged_lead_heartbeat_nudges_total";
/** Counter: spawn attempts by peer kind and outcome. */
public static final String SPAWNS = "bridged_spawns_total";
/** Counter: herdr socket calls by method and outcome. */
@@ -55,6 +57,9 @@ public final class BridgedMetrics {
"Worker replies by delivery path (rendezvous=resolved an open send, inbox=stranded and held).");
m.describe(PUSH_NUDGES, "counter",
"CB-307 push-loop nudges to the primary (delivered|exhausted).");
m.describe(HEARTBEAT_NUDGES, "counter",
"CB-551 idle-lead heartbeat nudges (delivered|failed|exhausted). Quiet-cap exhaustion "
+ "means the lead idled with nothing pending and was told to stand down.");
m.describe(SPAWNS, "counter",
"Worker spawn attempts by peer kind and outcome (ready|timeout|guard_rejected).");
m.describe(HERDR_CALLS, "counter",
@@ -69,7 +74,7 @@ public final class BridgedMetrics {
// One gauge per state so a scrape shows the whole census even when a state is empty —
// an absent series and a zero series read very differently on a dashboard.
for (WorkerSession.State state : WorkerSession.State.values()) {
for (MemberSession.State state : MemberSession.State.values()) {
String label = state.name().toLowerCase();
m.gauge(SESSIONS, () -> countIn(sessions, state), "state", label);
}
@@ -78,7 +83,7 @@ public final class BridgedMetrics {
// port's non-destructive read — scraping metrics must never ack a reply out of the inbox.
m.collector(INBOX_DEPTH, "target", () -> {
Map<String, Number> depths = new LinkedHashMap<>();
for (WorkerSession s : sessions.roster()) {
for (MemberSession s : sessions.roster()) {
String target = s.terminalId();
if (target == null) {
continue;
@@ -90,7 +95,7 @@ public final class BridgedMetrics {
return m;
}
private static long countIn(SessionManager sessions, WorkerSession.State state) {
private static long countIn(SessionManager sessions, MemberSession.State state) {
return sessions.roster().stream().filter(s -> s.state() == state).count();
}
}
@@ -14,25 +14,30 @@ import java.io.IOException;
import java.nio.charset.StandardCharsets;
import java.util.LinkedHashMap;
import java.util.List;
import java.util.Set;
import java.util.concurrent.ConcurrentHashMap;
/**
* AMQP-backed {@link ReplyInbox} (CB-307 Stage 2): genuine cross-restart durability behind the same
* port {@link InMemoryReplyInbox} implements as soft state.
*
* <p><strong>Mapping — consume-and-hold with deferred manual ack.</strong> Each target owns a durable
* queue {@code agent.<target>.inbox}. A manual-ack consumer pulls persistent messages off that queue
* into an in-memory <em>held</em> map (keyed by {@code msgId}) but does <em>not</em> ack them.
* {@link #peek} returns that snapshot; {@link #ack} acks the broker delivery-tag and drops the entry.
* Because messages stay unacked until the primary actually drains them, a crash (or a {@code java -jar}
* bounce) before caller-ack leaves them on the broker — it redelivers on reconnect. That is the
* durability the in-memory adapter cannot give, with the port contract preserved.
* <p><strong>Mapping — consume-and-hold with deferred manual ack.</strong> Each target has a durable
* queue {@code agent.<target>.inbox}. The gateway that owns the target starts a manual-ack consumer
* ({@link #own}) that pulls persistent messages off that queue into an in-memory <em>held</em> map
* (keyed by {@code msgId}) but does <em>not</em> ack them. {@link #peek} returns that snapshot;
* {@link #ack} acks the broker delivery-tag and drops the entry. Because messages stay unacked until
* the owning gateway actually drains them, a crash (or a {@code java -jar} bounce) before caller-ack
* leaves them on the broker — it redelivers on reconnect. That is the durability the in-memory
* adapter cannot give, with the port contract preserved.
*
* <p><strong>Ownership is explicit.</strong> {@link #own} declares the queue and starts the consumer;
* {@link #release} cancels it. {@link #publish} sends to the queue but does <em>not</em> imply ownership
* and does not attach a consumer. This split is required by CB-308 federation, where one gateway may
* publish to an agent owned by another gateway; in that case the publisher must not compete for
* deliveries.
*
* <p><strong>Dedup.</strong> The consumer keys the held map by {@code msgId}; a redelivered duplicate
* (at-least-once, or a producer double-publish) is acked-and-dropped on arrival, so it never
* double-queues. {@link #publish} additionally short-circuits an already-held {@code msgId} — a
* fast path; the consumer-side check is the real guarantee.
* double-queues.
*
* <p><strong>Visibility.</strong> Unlike the in-memory adapter, publish → broker → consumer is
* asynchronous, so a {@link #peek} immediately after {@link #publish} may not yet see the message
@@ -51,12 +56,12 @@ public final class AmqpReplyInbox implements ReplyInbox, AutoCloseable {
private final Connection connection;
private final Channel channel;
/** All channel operations (publish/declare/ack) serialize on this — a Channel is not thread-safe. */
/** All channel operations (publish/declare/ack/cancel) serialize on this — a Channel is not thread-safe. */
private final Object channelLock = new Object();
/** target → (msgId → held delivery). Per-target map is guarded by synchronizing on itself. */
private final ConcurrentHashMap<String, LinkedHashMap<String, Held>> held = new ConcurrentHashMap<>();
/** Targets whose queue is declared and consumer is running. */
private final Set<String> consuming = ConcurrentHashMap.newKeySet();
/** Targets whose queue is declared and consumer is running, mapped to their broker consumer tag. */
private final ConcurrentHashMap<String, String> consumerTags = new ConcurrentHashMap<>();
/** A message pulled off the broker but not yet acked: its delivery-tag plus the port payload. */
private record Held(long deliveryTag, InboxMessage message) {}
@@ -103,16 +108,41 @@ public final class AmqpReplyInbox implements ReplyInbox, AutoCloseable {
}
@Override
public void publish(String target, String msgId, String content) {
ensureConsuming(target);
var perTarget = held.get(target);
if (perTarget != null) {
synchronized (perTarget) {
if (perTarget.containsKey(msgId)) {
return; // already held — producer-side fast dedup
}
public void own(String target) {
synchronized (channelLock) {
if (consumerTags.containsKey(target)) {
return; // already owning this target
}
String queue = queueName(target);
try {
channel.queueDeclare(queue, true, false, false, null); // durable, non-exclusive, keep on idle
String tag = channel.basicConsume(queue, false, deliverCallback(target), _ -> { });
consumerTags.put(target, tag);
log.debug("AMQP inbox owns queue {} for target {}", queue, target);
} catch (IOException e) {
throw new IllegalStateException("cannot own queue " + queue, e);
}
}
}
@Override
public void release(String target) {
synchronized (channelLock) {
String tag = consumerTags.remove(target);
held.remove(target); // stale delivery tags must not survive release
if (tag == null) {
return;
}
try {
channel.basicCancel(tag);
} catch (IOException e) {
throw new IllegalStateException("cannot cancel consumer for " + target, e);
}
}
}
@Override
public void publish(String target, String msgId, String content) {
AMQP.BasicProperties props = new AMQP.BasicProperties.Builder()
.messageId(msgId)
.deliveryMode(2) // persistent — survives a broker restart
@@ -129,7 +159,6 @@ public final class AmqpReplyInbox implements ReplyInbox, AutoCloseable {
@Override
public List<InboxMessage> peek(String target) {
ensureConsuming(target);
var perTarget = held.get(target);
if (perTarget == null) {
return List.of();
@@ -166,26 +195,6 @@ public final class AmqpReplyInbox implements ReplyInbox, AutoCloseable {
}
}
/** Declare the durable per-target queue and start its manual-ack consumer, once per target. */
private void ensureConsuming(String target) {
if (consuming.contains(target)) {
return;
}
synchronized (channelLock) {
if (!consuming.add(target)) {
return; // another thread just set it up
}
String queue = queueName(target);
try {
channel.queueDeclare(queue, true, false, false, null); // durable, non-exclusive, keep on idle
channel.basicConsume(queue, false, deliverCallback(target), _ -> { });
} catch (IOException e) {
consuming.remove(target);
throw new IllegalStateException("cannot consume queue " + queue, e);
}
}
}
private DeliverCallback deliverCallback(String target) {
return (_, delivery) -> {
String msgId = delivery.getProperties().getMessageId();
@@ -2,6 +2,7 @@ package dev.ltms.bridged.msg;
import java.util.LinkedHashMap;
import java.util.List;
import java.util.Set;
import java.util.concurrent.ConcurrentHashMap;
/**
@@ -9,12 +10,29 @@ import java.util.concurrent.ConcurrentHashMap;
* Per-target FIFO ordering (insertion order via {@link LinkedHashMap}). Dedup by {@code msgId}
* within a target. Thread-safe for concurrent publish vs. drain.
*
* <p><strong>Ownership is explicit.</strong> {@link #own} marks a target as locally owned so that
* {@link #peek} and {@link #ack} operate on it; {@link #publish} works whether or not the target is
* owned. {@link #release} clears the local snapshot. This mirrors the AMQP adapter's contract so the
* non-broker path stays interchangeable.
*
* <p><strong>This is soft-state, NOT persistence.</strong> Lost on a {@code java -jar} bounce — that
* is correct and consistent with "bridged stays soft-state." The Stage-2 AMQP adapter replaces this.
*/
public final class InMemoryReplyInbox implements ReplyInbox {
private final ConcurrentHashMap<String, LinkedHashMap<String, InboxMessage>> store = new ConcurrentHashMap<>();
private final Set<String> owned = ConcurrentHashMap.newKeySet();
@Override
public void own(String target) {
owned.add(target);
}
@Override
public void release(String target) {
owned.remove(target);
store.remove(target);
}
@Override
public void publish(String target, String msgId, String content) {
@@ -27,6 +45,9 @@ public final class InMemoryReplyInbox implements ReplyInbox {
@Override
public List<InboxMessage> peek(String target) {
if (!owned.contains(target)) {
return List.of();
}
var perTarget = store.get(target);
if (perTarget == null) {
return List.of();
@@ -39,6 +60,9 @@ public final class InMemoryReplyInbox implements ReplyInbox {
@Override
public void ack(String target, String msgId) {
if (!owned.contains(target)) {
return;
}
var perTarget = store.get(target);
if (perTarget != null) {
//noinspection SynchronizationOnLocalVariableOrMethodParameter
@@ -0,0 +1,351 @@
package dev.ltms.bridged.msg;
import dev.ltms.bridged.herdr.AgentControl;
import dev.ltms.bridged.herdr.AgentStatus;
import dev.ltms.bridged.mcp.PrimaryRegistry;
import dev.ltms.bridged.metrics.BridgedMetrics;
import dev.ltms.bridged.metrics.Metrics;
import dev.ltms.bridged.session.MemberSession;
import org.slf4j.Logger;
import org.slf4j.LoggerFactory;
import java.util.ArrayList;
import java.util.List;
import java.util.concurrent.ScheduledExecutorService;
import java.util.concurrent.TimeUnit;
import java.util.function.LongSupplier;
import java.util.function.Supplier;
/**
* CB-551: an opt-in heartbeat that nudges the single idle lead back to work once it has been
* continuously idle past a quiet period with no open {@code bridge_send} driving it.
*
* <p>Why this exists: the fleet is ONE lead + architects + workers, so an idle, stalled lead is a
* single point of failure for the fleet's progress. {@link ReplyPushLoop} nudges the lead only when
* a worker reply lands; this loop is the timer that catches the gap where nothing lands and the
* lead simply sits idle with nobody to prompt it onward.
*
* <p>Mechanism mirrors {@link ReplyPushLoop}: a pure, unit-testable decision function
* ({@link #decide(long, Long, int, AgentStatus, boolean, boolean, FleetState)}) plus a thin
* scheduler around it. The loop is <em>entirely opt-in</em> ({@code leadHeartbeat:} in config); with
* no such block it is never constructed, so upgrading the daemon cannot silently acquire a behaviour
* that spends the operator's model subscription on its own initiative (constraint 1).
*
* <p>Four invariants keep it from becoming a runaway subscription burner:
* <ol>
* <li><b>Status-gated</b> — a {@code WORKING} lead is making progress and is never touched; only an
* injectable (idle/done/blocked) lead is even considered (constraint 2).</li>
* <li><b>Debounced</b> — a nudge only fires once the lead has been <em>continuously</em> idle past
* {@code idleAfterSeconds}, so a lead that just finished a turn is not re-prompted into every
* natural pause (constraint 3).</li>
* <li><b>Bounded when quiet</b> — consecutive nudges that find no pending fleet state are capped at
* {@code quietNudgeCap}, then the loop stops nudging until real state returns (constraint 4).</li>
* <li><b>Never races {@link ReplyPushLoop}</b> — while that loop is actively nudging any target this
* loop stands down, so two competing injections never start two turns in the same pane
* (constraint 6).</li>
* </ol>
*/
public final class LeadHeartbeatLoop {
private static final Logger log = LoggerFactory.getLogger(LeadHeartbeatLoop.class);
/** Sentinel for {@link #idleSinceNanos}: the lead is not currently in an idle stretch. */
private static final long NOT_IDLE = Long.MIN_VALUE;
private final PrimaryRegistry primaryRegistry;
private final AgentControl agents;
private final ReplyInbox inbox;
private final Supplier<List<MemberSession>> roster;
private final ReplyPushLoop pushLoop;
private final ScheduledExecutorService scheduler;
private final LongSupplier clock;
private final long idleAfterNanos;
private final long backoffMs;
private final int quietNudgeCap;
private final Metrics metrics; // CB-512 pattern: nullable — no registry in unit tests
/** When the current idle stretch began (nanos), or {@link #NOT_IDLE}. Single scheduler thread only. */
private long idleSinceNanos = NOT_IDLE;
/** Consecutive nudges that found no pending fleet state. Single scheduler thread only. */
private int quietCount = 0;
/** Constructor with an injectable clock and no metric registry (unit tests, or wiring that opts out). */
public LeadHeartbeatLoop(PrimaryRegistry primaryRegistry, AgentControl agents, ReplyInbox inbox,
Supplier<List<MemberSession>> roster, ReplyPushLoop pushLoop,
ScheduledExecutorService scheduler, LongSupplier clock,
long idleAfterNanos, long backoffMs, int quietNudgeCap) {
this(primaryRegistry, agents, inbox, roster, pushLoop, scheduler, clock,
idleAfterNanos, backoffMs, quietNudgeCap, null);
}
/** As above, with a metric registry (the CB-512 pattern) so nudge outcomes are counted. */
public LeadHeartbeatLoop(PrimaryRegistry primaryRegistry, AgentControl agents, ReplyInbox inbox,
Supplier<List<MemberSession>> roster, ReplyPushLoop pushLoop,
ScheduledExecutorService scheduler, LongSupplier clock,
long idleAfterNanos, long backoffMs, int quietNudgeCap, Metrics metrics) {
this.primaryRegistry = primaryRegistry;
this.agents = agents;
this.inbox = inbox;
this.roster = roster;
this.pushLoop = pushLoop;
this.scheduler = scheduler;
this.clock = clock;
this.idleAfterNanos = idleAfterNanos;
this.backoffMs = backoffMs;
this.quietNudgeCap = quietNudgeCap;
this.metrics = metrics;
}
/** Count one nudge outcome when a registry is wired; a no-op in unit tests. */
private void countNudge(String outcome) {
if (metrics != null) {
metrics.inc(BridgedMetrics.HEARTBEAT_NUDGES, "outcome", outcome);
}
}
// --- decision logic (package-private for unit-testing) --------------------------------------
/** The action the loop should take on one tick. */
enum Action {
/** Send a heartbeat nudge to the lead now. */
INJECT,
/** Lead is idle but not yet past the quiet period — keep waiting, do nothing. */
WAIT_IDLE,
/** Lead is making progress (or its status cannot be read) — reset the idle window and re-arm. */
LEAD_BUSY,
/** Idle past the quiet period with nothing pending and the cap exhausted — stop until new state appears. */
QUIET_DONE,
/** {@link ReplyPushLoop} is actively nudging — stand aside rather than start a competing turn. */
STAND_DOWN
}
/** The outcome of one decision: the action plus the state to persist for the next tick. */
record Decision(Action action, Long idleSinceNanos, int quietCount) {}
/**
* Pure decision function: given the current loop state and fleet/lead facts, return what to do
* next and the exact state to carry forward. Pure means no I/O and no mutation — the caller
* ({@link #tick}) applies {@link Decision} to its fields. This is what makes the loop unit-testable
* with a fake clock and no sleeping.
*
* @param nowNanos the current time (an injected clock in tests, {@link System#nanoTime()} live)
* @param idleSinceNanos when the current idle stretch began, or {@code null} if the lead is not idle
* @param quietCount consecutive nudges so far that found no pending fleet state
* @param status the lead's herdr status
* @param pushLoopActive whether {@link ReplyPushLoop} is currently nudging some target (constraint 6)
* @param leadKnown whether a lead terminal is known to nudge at all
* @param fleet a snapshot of the pending fleet state (constraint 5)
* @return the action to take and the state to persist
*/
Decision decide(long nowNanos, Long idleSinceNanos, int quietCount, AgentStatus status,
boolean pushLoopActive, boolean leadKnown, FleetState fleet) {
// Constraint 6: while ReplyPushLoop is actively nudging the lead, injecting a second,
// competing prompt into the same pane would start a second turn — racing loops multiply
// turns and context burn. Stand aside, and treat the active push as real state (re-arm the
// quiet counter), because the reply that drove it is exactly the kind of new state that
// should reset the cap.
if (pushLoopActive) {
return new Decision(Action.STAND_DOWN, idleSinceNanos, 0);
}
// Constraint 2: a WORKING lead is making progress and must NOT be touched; an unreadable
// status (read failure, or the agent is gone) is safest treated the same way — never inject
// into a state we cannot read. Either way, reset the idle window and the quiet counter: the
// lead was / may be active, so the next idle stretch must count its own quiet period fresh.
if (status == null || !status.injectable()) {
return new Decision(Action.LEAD_BUSY, null, 0);
}
if (idleSinceNanos == null) {
// The lead just became injectable — record the start of an idle stretch and wait out the
// debounce quiet period before ever nudging (constraint 3).
return new Decision(Action.WAIT_IDLE, nowNanos, quietCount);
}
if (nowNanos - idleSinceNanos < idleAfterNanos) {
// Still within the quiet period: the lead that just finished a turn sits momentarily idle
// and must not be re-prompted into every natural pause.
return new Decision(Action.WAIT_IDLE, idleSinceNanos, quietCount);
}
// Past the quiet period with an injectable lead: it is a genuine candidate for a nudge. Two
// gating facts decide whether and how:
if (!leadKnown) {
// No lead terminal is known yet (e.g. the registry has not learned one) — there is nobody
// to nudge. Keep waiting; the window stays open so discovery re-arms it without a fresh
// quiet period.
return new Decision(Action.WAIT_IDLE, idleSinceNanos, quietCount);
}
if (fleet.hasPending()) {
// Real fleet state is waiting — a worker reply or a DONE session. This is new state, so
// it resets the quiet counter (constraint 4) and the lead is nudged to go collect it.
return new Decision(Action.INJECT, idleSinceNanos, 0);
}
if (quietCount < quietNudgeCap) {
// Nothing is pending, but the cap is not exhausted: nudge anyway, telling the lead
// exactly that nothing is waiting so it can choose to stand down rather than hunt
// (constraint 5). Count it toward the consecutive-quiet cap.
return new Decision(Action.INJECT, idleSinceNanos, quietCount + 1);
}
// Nothing pending and the cap is exhausted: stop nudging until real state appears again
// (constraint 4). The loop still ticks on backoff so a genuinely new reply or session change
// re-arms it — QUIET_DONE stops injection, not observation.
return new Decision(Action.QUIET_DONE, idleSinceNanos, quietCount);
}
// --- loop ----------------------------------------------------------------------------------
/**
* Start the heartbeat. Schedules the first tick after one backoff so a freshly-booting daemon
* does not evaluate the lead's idle state before the fleet has settled.
*/
public void start() {
log.info("idle-lead heartbeat: on — nudge lead after {}s idle (recheck {}ms, quiet cap {})",
TimeUnit.NANOSECONDS.toSeconds(idleAfterNanos), backoffMs, quietNudgeCap);
scheduler.schedule(this::tick, backoffMs, TimeUnit.MILLISECONDS);
}
/** One loop tick, every {@link #backoffMs} — the thin scheduler around {@link #decide}. */
private void tick() {
boolean leadKnown = primaryRegistry.primaryTerminal().isPresent();
FleetState fleet = snapshot(inbox, roster);
AgentStatus status = AgentStatus.UNKNOWN;
if (leadKnown) {
try {
status = agents.status(primaryRegistry.primaryTerminal().orElseThrow());
} catch (RuntimeException e) {
// A failed status read degrades to "unknown" — decide() treats that like a busy lead
// and never injects into a state it cannot read. Retry on the next backoff.
log.debug("idle-heartbeat: status check failed for lead, will retry: {}", e.toString());
}
}
Decision d = decide(clock.getAsLong(),
idleSinceNanos == NOT_IDLE ? null : idleSinceNanos,
quietCount, status, pushLoop.isActive(), leadKnown, fleet);
applyDecision(d);
switch (d.action()) {
case INJECT -> injectNudge(fleet);
case QUIET_DONE -> countNudge("exhausted");
case WAIT_IDLE, LEAD_BUSY, STAND_DOWN -> { /* nothing to inject, nothing to count */ }
}
scheduleNext();
}
/** Persist the state a decision returned, so the next tick starts from it. */
private void applyDecision(Decision d) {
idleSinceNanos = d.idleSinceNanos() == null ? NOT_IDLE : d.idleSinceNanos();
quietCount = d.quietCount();
}
/** Send the nudge to the known lead. */
private void injectNudge(FleetState fleet) {
var lead = primaryRegistry.primaryTerminal();
if (lead.isEmpty()) {
return; // the lead disappeared between the decision and the injection
}
String leadTerminal = lead.get();
try {
agents.send(leadTerminal, fleet.nudgeText());
log.debug("idle-heartbeat: nudge sent to lead {} (quiet nudges so far in this stretch: {})",
leadTerminal, quietCount);
countNudge("delivered");
} catch (RuntimeException e) {
log.warn("idle-heartbeat: failed to nudge lead {}: {}", leadTerminal, e.toString());
countNudge("failed");
}
}
/** Schedule the next tick on the scheduler thread pool. */
private void scheduleNext() {
scheduler.schedule(this::tick, backoffMs, TimeUnit.MILLISECONDS);
}
// --- lifecycle -----------------------------------------------------------------------------
/** Shut down the scheduler; outstanding ticks are cancelled. */
public void stop() {
scheduler.shutdownNow();
}
/** @see #stop() */
public void close() {
stop();
}
/**
* A read-only snapshot of the fleet facts the heartbeat reports to an idle lead. Carries exact
* counts and names the targets that hold replies, so the nudge is actionable rather than a
* content-free "keep working" that would cost the lead a turn just to discover what there is to
* do (constraint 5).
*
* @param pendingReplies total undrained worker replies across owned targets
* @param doneSessions session count in the DONE state — finished turns awaiting teardown
* @param liveWorkers sessions still registered (acquired minus released)
* @param replyTargets the terminals whose inboxes currently hold at least one reply
*/
record FleetState(int pendingReplies, int doneSessions, int liveWorkers, List<String> replyTargets) {
/** True when the lead has something real to collect — a reply or a session awaiting teardown. */
boolean hasPending() {
return pendingReplies > 0 || doneSessions > 0;
}
/** The nudge body, phrased for the two cases the heartbeat distinguishes. */
String nudgeText() {
StringBuilder sb = new StringBuilder(
"Heartbeat: you are idle and no bridge_send is waiting on you.");
if (hasPending()) {
sb.append(" The fleet has state to collect: ").append(pendingDetail());
} else {
sb.append(" Fleet is quiet — nothing pending to collect (")
.append(pendingReplies).append(" worker replies, ")
.append(doneSessions).append(" DONE sessions); ")
.append(liveWorkers).append(" workers live. You may stand down until new state re-arms me.");
}
return sb.toString();
}
private String pendingDetail() {
StringBuilder sb = new StringBuilder();
sb.append(pendingReplies).append(" worker ")
.append(pendingReplies == 1 ? "reply" : "replies")
.append(" pending collection");
if (!replyTargets.isEmpty()) {
// Render each as the exact command so the lead can act without parsing: the nearest
// analogue to ReplyPushLoop's bridge_poll(target=...) nudge.
sb.append(" (")
.append(String.join(", ",
replyTargets.stream().map(t -> "bridge_poll(target=" + t + ")").toList()))
.append(")");
}
sb.append(", ").append(doneSessions).append(" DONE session").append(doneSessions == 1 ? "" : "s")
.append(" awaiting teardown")
.append(", ").append(liveWorkers).append(" worker").append(liveWorkers == 1 ? "" : "s")
.append(" live");
return sb.toString();
}
}
/**
* Build a {@link FleetState} snapshot over the live roster and the reply inbox — the same
* sources the metrics collector uses, so the heartbeat reports facts the operator can cross-check
* against {@code /metrics}. Package-private and pure (reads only) so it is unit-testable.
*/
static FleetState snapshot(ReplyInbox inbox, Supplier<List<MemberSession>> roster) {
List<MemberSession> sessions = roster.get();
int replies = 0;
int done = 0;
List<String> replyTargets = new ArrayList<>();
for (MemberSession s : sessions) {
if (s.terminalId() == null) {
continue;
}
int depth = inbox.peek(s.terminalId()).size();
if (depth > 0) {
replies += depth;
replyTargets.add(s.terminalId());
}
if (s.state() == MemberSession.State.DONE) {
done++;
}
}
return new FleetState(replies, done, sessions.size(), List.copyOf(replyTargets));
}
}
@@ -310,6 +310,22 @@ public final class MessageService {
* worker replies via {@link Rendezvous} or {@code timeoutMillis} elapses.
*/
public Reply send(String target, String content, long timeoutMillis) {
return send(target, content, timeoutMillis, null);
}
/**
* As {@link #send(String, String, long)}, but with an accepted-delivery hook.
*
* <p>{@code onAccepted} is invoked exactly once, once this send has won {@code target}'s send
* lock and so become the <em>accepted target turn</em> — it runs <em>before</em> delivery is
* queued, so a throwing hook fails the send cleanly (the waiter it already opened is closed and
* nothing is left queued). It is <em>not</em> invoked when the send is {@link Outcome#BUSY}
* (lock never taken). A caller uses this to record that <em>it</em> now owns the delegation's
* reply routing (CB-548: {@code PrimaryRegistry} delegator ownership) — recording only on
* acceptance means a concurrent sender that times out {@code BUSY} can never steal ownership it
* never earned. {@code null} disables the hook.
*/
public Reply send(String target, String content, long timeoutMillis, Runnable onAccepted) {
long deadlineNanos = System.nanoTime() + timeoutMillis * 1_000_000L;
ReentrantLock lock = sessionLocks.computeIfAbsent(target, _ -> new ReentrantLock());
@@ -317,22 +333,36 @@ public final class MessageService {
return new Reply(Outcome.BUSY, null); // another send held the session the whole window
}
try {
CompletableFuture<Void> delivered = injector.enqueue(target, content);
// Open the waiter BEFORE queueing delivery (CB-548). A fast reply — the worker already
// injectable the instant we enqueue — otherwise arrives before the waiter is registered
// and orphans into the inbox while this send blocks to the timeout (the enqueue-before-
// open race). Opening first also means a throwing onAccepted (fired before enqueue) or an
// enqueue failure is safely closed by the finally below: nothing is left queued, and the
// failed send leaves no stale waiter behind.
CompletableFuture<Rendezvous.Resolution> reply = rendezvous.open(target);
try {
Rendezvous.Resolution r = reply.get(remainingMillis(deadlineNanos), TimeUnit.MILLISECONDS);
return recorded(new Reply(outcomeOf(r.kind()), r.text(), r.turnId()));
} catch (TimeoutException e) {
boolean wasDelivered = delivered.isDone() && !delivered.isCompletedExceptionally();
log.debug("send to {} timed out (delivered={})", target, wasDelivered);
return recorded(new Reply(
wasDelivered ? Outcome.TIMED_OUT_WORKING : Outcome.TIMED_OUT_QUEUED, null));
} catch (ExecutionException e) {
Throwable cause = e.getCause();
throw cause instanceof RuntimeException re ? re : new IllegalStateException(cause);
} catch (InterruptedException e) {
Thread.currentThread().interrupt();
throw new IllegalStateException("interrupted awaiting reply from " + target, e);
// The send has won the lock; the accepted-delivery hook records delegator ownership
// here (CB-548). It runs BEFORE enqueue so a throwing hook — onAccepted is now a
// public callback — fails the send without queuing a message that would orphan.
if (onAccepted != null) {
onAccepted.run();
}
CompletableFuture<Void> delivered = injector.enqueue(target, content);
try {
Rendezvous.Resolution r = reply.get(remainingMillis(deadlineNanos), TimeUnit.MILLISECONDS);
return recorded(new Reply(outcomeOf(r.kind()), r.text(), r.turnId()));
} catch (TimeoutException e) {
boolean wasDelivered = delivered.isDone() && !delivered.isCompletedExceptionally();
log.debug("send to {} timed out (delivered={})", target, wasDelivered);
return recorded(new Reply(
wasDelivered ? Outcome.TIMED_OUT_WORKING : Outcome.TIMED_OUT_QUEUED, null));
} catch (ExecutionException e) {
Throwable cause = e.getCause();
throw cause instanceof RuntimeException re ? re : new IllegalStateException(cause);
} catch (InterruptedException e) {
Thread.currentThread().interrupt();
throw new IllegalStateException("interrupted awaiting reply from " + target, e);
}
} finally {
rendezvous.close(target, reply);
}
@@ -440,9 +470,21 @@ public final class MessageService {
* @return the ticket to poll for the eventual result
*/
public String sendAsync(String target, String content) {
return sendAsync(target, content, null);
}
/**
* As {@link #sendAsync(String, String)}, with the accepted-delivery hook of
* {@link #send(String, String, long, Runnable)} — the running {@code send} invokes {@code onAccepted}
* the moment it becomes the accepted target turn, so async flooding records delegator ownership
* exactly as the blocking path does (CB-548).
*
* @return the ticket to poll for the eventual result
*/
public String sendAsync(String target, String content, Runnable onAccepted) {
String ticket = "task-" + ticketSeq.incrementAndGet();
CompletableFuture<Reply> future =
CompletableFuture.supplyAsync(() -> send(target, content, ASYNC_TIMEOUT_MS), asyncExecutor);
CompletableFuture<Reply> future = CompletableFuture.supplyAsync(
() -> send(target, content, ASYNC_TIMEOUT_MS, onAccepted), asyncExecutor);
tasks.put(ticket, new Task(target, future, System.nanoTime()));
pruneTerminalTickets();
log.debug("async send {} -> {}", ticket, target);
@@ -74,15 +74,30 @@ public final class Rendezvous {
/**
* Register a waiter for {@code session} — the await side of the public {@code resolve*} methods.
* The caller must hold that session's send lock.
*
* <p>Atomic fail-if-present (CB-548): if a waiter is already registered for {@code session}, an
* {@link IllegalStateException} is thrown rather than replacing the first — so any future
* invariant violation fails loudly instead of silently swapping the waiter another send is
* blocked on. {@code MessageService} serializes sends per session (the send lock), so in correct
* code a double open is impossible; this is a tripwire for the day that no longer holds.
*/
public CompletableFuture<Resolution> open(String session) {
CompletableFuture<Resolution> waiter = new CompletableFuture<>();
waiters.put(session, waiter);
CompletableFuture<Resolution> existing = waiters.putIfAbsent(session, waiter);
if (existing != null) {
throw new IllegalStateException(
"rendezvous double-open for session " + session + " — a waiter is already registered");
}
return waiter;
}
/** Remove {@code waiter} for {@code session} (only if it is still the registered one). */
void close(String session, CompletableFuture<Resolution> waiter) {
/**
* Remove {@code waiter} for {@code session}, only if it is still the registered one. The
* symmetric complement of {@link #open}: a terminal send deregisters its waiter so the next
* send on the session may {@link #open} a fresh one (CB-548 makes double-open an error, so a
* successful {@code open} after a finished turn requires this close to have happened first).
*/
public void close(String session, CompletableFuture<Resolution> waiter) {
waiters.remove(session, waiter);
}
@@ -10,16 +10,36 @@ import java.util.List;
* <p><strong>This interface is the port.</strong> {@link InMemoryReplyInbox} is the Stage-1 adapter;
* an AMQP-backed adapter (Stage 2) must implement the same contract (idempotent publish, FIFO peek,
* at-least-once ack).
*
* <p><strong>Ownership is explicit.</strong> A gateway {@link #own owns} the inbox for each agent it
* spawned; only the owner consumes and drains it. {@link #publish} sends a reply to the target's
* inbox but does <em>not</em> imply ownership or start a consumer. This separation is required by
* CB-308 federation, where one gateway may publish to an agent owned by another gateway.
*/
public interface ReplyInbox {
/** A queued reply: an idempotency id, the worker session it came from, and the reply text. */
record InboxMessage(String msgId, String target, String content) {}
/**
* Start owning (consuming) the inbox for {@code target}. Idempotent: multiple calls for the same
* target are no-ops. The owner is the only gateway that may {@link #peek} and {@link #ack} replies
* for this target.
*/
void own(String target);
/**
* Stop owning (consuming) the inbox for {@code target}. Idempotent. Any replies held locally but
* not yet acked are dropped from the local snapshot; the underlying durable queue keeps
* unacked messages for redelivery when the target is re-owned.
*/
void release(String target);
/**
* Queue {@code content} from worker {@code target} under {@code msgId}. Idempotent: publishing an
* already-present {@code msgId} for {@code target} is a no-op (dedup), so an at-least-once Stage-2
* redelivery cannot double-queue.
* redelivery cannot double-queue. Publishing does <em>not</em> imply ownership and must not start a
* consumer.
*/
void publish(String target, String msgId, String content);
@@ -59,10 +59,10 @@ public final class ReplyPushLoop {
this.metrics = metrics;
}
/** Record a counter sample when a registry is wired; a no-op in unit tests. */
private void count(String name, String... labels) {
/** Count one nudge outcome when a registry is wired; a no-op in unit tests. */
private void countNudge(String outcome) {
if (metrics != null) {
metrics.inc(name, labels);
metrics.inc(BridgedMetrics.PUSH_NUDGES, "outcome", outcome);
}
}
@@ -79,8 +79,13 @@ public final class ReplyPushLoop {
* @return the action the caller should take
*/
Action decide(String target, int reminderCount) {
if (!primaryRegistry.isKnown()) {
log.debug("push: primary unknown, stopping reminder for {}", target);
// CB-532: the destination is per-delegation — the lead that sent this worker its work, not
// "the primary". With two leads orchestrating one fleet the singular question has no right
// answer, and answering it anyway interrupted whichever lead happened to call bridge_send
// first with results it never asked for.
var nudgeTarget = primaryRegistry.nudgeTargetFor(target);
if (nudgeTarget.isEmpty()) {
log.debug("push: no lead is known to be waiting on {}, stopping reminder", target);
return Action.STOP;
}
if (inbox.peek(target).isEmpty()) {
@@ -89,21 +94,21 @@ public final class ReplyPushLoop {
}
if (reminderCount >= maxReminders) {
log.debug("push: reminder cap ({}) reached for {}, stopping", maxReminders, target);
count(BridgedMetrics.PUSH_NUDGES, "outcome", "exhausted");
countNudge("exhausted");
return Action.STOP;
}
var primaryTerminal = primaryRegistry.primaryTerminal().orElseThrow();
String leadTerminal = nudgeTarget.get();
AgentStatus status;
try {
status = agents.status(primaryTerminal);
status = agents.status(leadTerminal);
} catch (RuntimeException e) {
log.debug("push: status check failed for primary {}, will retry", primaryTerminal, e);
log.debug("push: status check failed for lead {}, will retry", leadTerminal, e);
return Action.WAIT_BUSY;
}
if (status.injectable()) {
return Action.INJECT;
}
log.debug("push: primary {} is {} (not injectable), waiting", primaryTerminal, status);
log.debug("push: lead {} is {} (not injectable), waiting", leadTerminal, status);
return Action.WAIT_BUSY;
}
@@ -143,16 +148,23 @@ public final class ReplyPushLoop {
/** Send the nudge and log the event. */
private void injectNudge(String target, int reminderCount) {
var primaryTerminal = primaryRegistry.primaryTerminal().orElseThrow();
// Re-read rather than threading it down from decide(): the delegating lead can change
// between the decision and the injection, and the nudge should follow the current one.
var lead = primaryRegistry.nudgeTargetFor(target);
if (lead.isEmpty()) {
log.debug("push: lead for {} disappeared before the nudge could be sent", target);
return;
}
String leadTerminal = lead.get();
String nudge = NUDGE_FORMAT.formatted(target, target);
try {
agents.send(primaryTerminal, nudge);
log.debug("push: nudge {}/{} sent to primary {} for target {}",
reminderCount + 1, maxReminders, primaryTerminal, target);
count(BridgedMetrics.PUSH_NUDGES, "outcome", "delivered");
agents.send(leadTerminal, nudge);
log.debug("push: nudge {}/{} sent to lead {} for target {}",
reminderCount + 1, maxReminders, leadTerminal, target);
countNudge("delivered");
} catch (RuntimeException e) {
log.warn("push: failed to nudge primary {} for target {} (reminder {}/{}): {}",
primaryTerminal, target, reminderCount + 1, maxReminders, e.toString());
log.warn("push: failed to nudge lead {} for target {} (reminder {}/{}): {}",
leadTerminal, target, reminderCount + 1, maxReminders, e.toString());
}
}
@@ -163,6 +175,17 @@ public final class ReplyPushLoop {
// --- lifecycle -----------------------------------------------------------------------------
/**
* Whether any reminder loop is currently active for some target (CB-551). The idle-lead heartbeat
* uses this to stand aside: while the push loop is actively nudging the lead, a concurrent
* heartbeat injection would start a second competing turn in the same pane — racing loops multiply
* turns and context burn. "Active" means a schedule exists in {@link #activeTargets}; the set is
* bounded by what has been triggered, not by any persistent state.
*/
public boolean isActive() {
return !activeTargets.isEmpty();
}
/** Shut down the scheduler. Outstanding reminders are cancelled. */
public void stop() {
scheduler.shutdownNow();
@@ -26,10 +26,33 @@ public enum Capability {
*/
WORKTREE,
/**
* The peer can discard its current conversation context without starting a delegated bridge
* turn. The command or API used to do that is adapter-specific.
*/
CONTEXT_RESET,
/**
* The spawner can reconcile orphaned peers on boot — workers that outlived a prior daemon
* process and whose pane ids died with it (CB-117). Claude Code over herdr supports this
* via name-based matching against the herdr agent list.
*/
ORPHAN_REAP
ORPHAN_REAP,
/**
* The peer surfaces the bridge's logical name ({@link SpawnRequest#sessionName()}) in its own
* UI at spawn — e.g. Claude Code's {@code -n} display name, which shows in the prompt box, the
* {@code /resume} picker, and the terminal title. This is what lets an operator tell the
* bridge's worker apart from a user's own session in the same terminal after a restart.
*/
SESSION_NAME,
/**
* The peer can be relaunched onto a prior conversation via that conversation's own session id
* ({@link SpawnRequest#resumeSessionId()}) — e.g. Claude Code's {@code -r}, which adopts the id
* as the agent's own identity rather than starting a new conversation. A launcher that declares
* this mints/resolves that id at spawn and exposes it on the returned
* {@link PeerHandle#agentSessionId()}, so a later resume addresses the same conversation.
*/
SESSION_RESUME
}
@@ -0,0 +1,122 @@
package dev.ltms.bridged.peer;
import java.util.Locale;
/**
* What a member is <em>for</em> — the contract it runs under.
*
* <p>A member is anything a lead spawns. Every member carries two independent attributes:
*
* <ul>
* <li><b>role</b> (this enum) — <em>which contract</em>: the launch charter it is given, the role
* file it reads, the playbook skill it loads, and its authorization row.</li>
* <li><b>profile</b> (a {@code profiles:} key) — <em>which backend</em>: model, CLI adapter,
* credentials, cost.</li>
* </ul>
*
* <p>These are separate axes on purpose. A {@code REVIEWER} may run on the same profile as the
* {@code DEV} whose diff it reviews — same backend, different contract. That case is what proves
* role and profile cannot be collapsed into one field.
*
* <p>A lead is deliberately <em>not</em> a role here. A lead is not spawned: it is a pre-existing,
* human-facing session that config recognises. Only spawned peers have a member role.
*/
public enum MemberRole {
/**
* Refines a ticket before anyone implements it: scope, acceptance criteria, risks, unit split.
*
* <p>Reads the repo and writes analysis. Never commits code and never opens a pull request —
* an architect that starts implementing has stopped doing the job that makes it useful.
*
* <p>Architects are the one member kind declared in config, because a lead addresses the same
* slots across many tickets and needs a stable name for them.
*/
ARCHITECT,
/**
* Implements one unit of work: provisions a worktree, writes the code, commits, pushes, and
* opens its own pull request.
*
* <p>Never merges. The lead is the gate, and a member that merged its own work would remove the
* only independent check in the flow.
*/
DEV,
/**
* Reviews a diff it did not write and reports one structured finding.
*
* <p>Never commits, never merges, and is never the member that wrote the scope under review.
* Briefed from the diff rather than from the implementer's rationale, because that rationale
* carries the same blind spot that produced the bug.
*/
REVIEWER;
/** The lowercase spelling used in config and on the wire ({@code architect}, {@code dev}, …). */
public String wireName() {
return name().toLowerCase(Locale.ROOT);
}
/**
* The {@code fleet:} block that holds this role's pool — {@code architects},
* {@code developers}, {@code reviewers}.
*
* <p>Plural, and not always the wire name: the pool of things a {@code dev} may run on reads
* naturally as {@code developers:}. The wire name stays the singular {@code dev}, because that
* is what a tab label and a roster row say.
*/
public String configKey() {
return switch (this) {
case ARCHITECT -> "architects";
case DEV -> "developers";
case REVIEWER -> "reviewers";
};
}
/**
* The role owning the {@code fleet:} pool named {@code key}, or {@code null} when the key is not
* a role pool ({@code leaders}, {@code tabLabel}, …). Null rather than a throw: callers use this
* to sort a fleet block's children, where a non-pool key is normal rather than an error.
*/
public static MemberRole fromConfigKey(String key) {
if (key == null) {
return null;
}
String k = key.trim().toLowerCase(Locale.ROOT);
for (MemberRole r : values()) {
if (r.configKey().equals(k)) {
return r;
}
}
return null;
}
/**
* Parse a config/wire spelling, case-insensitively.
*
* @param s the spelling to parse
* @return the matching role
* @throws IllegalArgumentException when {@code s} is null, blank, or not a known role — the
* message lists the valid spellings, because a typo'd role in
* config should fail at startup with the fix in the error
*/
public static MemberRole parse(String s) {
if (s != null && !s.isBlank()) {
String t = s.trim().toLowerCase(Locale.ROOT);
for (MemberRole r : values()) {
if (r.wireName().equals(t)) {
return r;
}
}
}
StringBuilder valid = new StringBuilder();
for (MemberRole r : values()) {
if (!valid.isEmpty()) {
valid.append(", ");
}
valid.append(r.wireName());
}
throw new IllegalArgumentException(
"unknown member role '" + s + "'; valid roles are: " + valid);
}
}
@@ -12,9 +12,13 @@ package dev.ltms.bridged.peer;
public interface PeerHandle {
/**
* The registry/routing key — an opaque, launcher-assigned identifier. For the herdr-backed
* launcher this is the herdr pane id; for other launchers it is whatever their transport
* uses. Guaranteed to be non-null and unique among live peers within a single daemon process.
* The registry/routing key — an opaque, launcher-assigned identifier (CB-519). Multiple
* daemon processes may run on one host, so the contract is <em>host-unique</em>, not merely
* process-unique: the herdr-backed launcher mints a fresh UUID per spawn, and a non-herdr
* launcher is likewise expected to return an identifier that cannot collide across processes
* on the same host. This id is the routing key and is deliberately decoupled from any launcher
* transport coordinate (e.g. a herdr pane id), which stays launcher-private. Guaranteed to be
* non-null and unique among live peers on the host.
*/
String id();
@@ -26,4 +30,41 @@ public interface PeerHandle {
default String terminalId() {
return null;
}
/**
* The worker profile that spawned this peer, if the launcher resolved one. A launcher that
* performs dynamic profile selection (e.g. CB-518 weighted placement) sets this so the
* session registry records the actual profile rather than the requested/default one.
*
* @return the profile name, or {@code null} when the launcher leaves it unspecified
*/
default String profile() {
return null;
}
/**
* The bridge's logical name for this session, as assigned at spawn
* ({@link SpawnRequest#sessionName()}). Stable across restarts and meaningful to an operator,
* unlike the transport identifiers above; a launch with no name leaves the peer's display
* identity to the launcher to derive.
*
* @return the bridge-assigned logical session name, or {@code null} if none was assigned
*/
default String sessionName() {
return null;
}
/**
* The peer's OWN session id — the handle that resumes this conversation later (the id a
* later {@link Capability#SESSION_RESUME resume} spawn would pass back). Null when the adapter
* cannot determine it — the contract for an adapter that declines
* {@link Capability#SESSION_RESUME}; an adapter that declares that capability returns this
* non-null for a spawn that requested session identity, because it knows the id before the
* peer has written anything.
*
* @return the peer's own session id, or {@code null} when not determinable
*/
default String agentSessionId() {
return null;
}
}
@@ -82,4 +82,13 @@ public interface PeerLauncher {
* when safe to do so.
*/
void stop(String id);
/**
* Discard the context of the peer identified by {@code id}. Implementations must bypass normal
* bridge delivery/turn accounting. Unsupported peer kinds return {@code false} without sending
* a guessed command.
*
* @return {@code true} when a reset was sent and its status transition must settle before reuse
*/
boolean clearContext(String id);
}
@@ -1,13 +1,32 @@
package dev.ltms.bridged.peer;
/**
* Parameters for a {@link PeerLauncher#spawn(SpawnRequest)} call — the peer-neutral
* aggregation of what the core knows at delegation time: which profile to use, the caller's
* requested working directory, and the caller's own cwd (to inherit when no other cwd is set).
* What a caller asks for when spawning a peer.
*
* <p>A null or blank {@code profileName} means "use the launcher's default profile."
* A null or blank {@code requestedCwd} means "inherit from config or caller."
* A null {@code callerCwd} means "the request came from the daemon itself (not a primary)."
* @param profileName the {@code profiles:} entry to spawn on; {@code null}/blank ⇒ the caller
* did not choose, and the role's pool supplies the candidates
* @param requestedCwd working directory asked for by the caller ({@code null} ⇒ unset)
* @param callerCwd the caller's own working directory, used when nothing else pins one
* @param sessionName agent session name (CB-547a); {@code null} ⇒ the launcher mints one
* @param resumeSessionId prior agent session to resume; {@code null} ⇒ a fresh session
* @param role the contract this member runs under (CB-557). Picks the tab label and,
* with the profile, the counter its tab number comes from. {@code null} is
* read as {@link MemberRole#DEV} — an unqualified spawn is a unit of work
*/
public record SpawnRequest(String profileName, String requestedCwd, String callerCwd) {
public record SpawnRequest(String profileName, String requestedCwd, String callerCwd,
String sessionName, String resumeSessionId, MemberRole role) {
public SpawnRequest {
role = (role == null) ? MemberRole.DEV : role;
}
public SpawnRequest(String profileName, String requestedCwd, String callerCwd) {
this(profileName, requestedCwd, callerCwd, null, null, null);
}
/** Pre-CB-557 shape: session identity without an explicit role (defaults to {@code dev}). */
public SpawnRequest(String profileName, String requestedCwd, String callerCwd,
String sessionName, String resumeSessionId) {
this(profileName, requestedCwd, callerCwd, sessionName, resumeSessionId, null);
}
}
@@ -0,0 +1,22 @@
package dev.ltms.bridged.placement;
/**
* Backward-compatible placement: an unqualified spawn always resolves to the configured default
* profile, exactly as {@code CompositePeerLauncher} did before CB-518. This ignores caps and
* reachability so that a pre-existing config behaves identically after upgrade.
*/
final class FixedPlacementPolicy implements PlacementPolicy {
@Override
public PlacementCandidate select(PlacementContext ctx) {
String d = ctx.defaultProfile();
if (d != null && !d.isBlank()) {
return new PlacementCandidate(d, null, 1.0f, null);
}
if (!ctx.candidates().isEmpty()) {
PlacementCandidate first = ctx.candidates().getFirst();
return new PlacementCandidate(first.profile(), null, first.weight(), first.maxLoad());
}
throw new PlacementException("no worker profiles configured");
}
}
@@ -0,0 +1,20 @@
package dev.ltms.bridged.placement;
/**
* A profile (and, in CB-308, a host) that can be chosen by a {@link PlacementPolicy}.
*
* <p>Keeping this as a small descriptor rather than a bare profile name lets CB-308 widen
* selection to {@code (host, profile)} pairs without changing the policy interface.
*/
public record PlacementCandidate(String profile, String host, float weight, Integer maxLoad) {
/** A candidate with no explicit host (the single-host default) and the given weight/cap. */
public static PlacementCandidate profile(String profile, float weight, Integer maxLoad) {
return new PlacementCandidate(profile, null, weight, maxLoad);
}
/** A candidate with no explicit host, unit weight, and no cap. */
public static PlacementCandidate profile(String profile) {
return new PlacementCandidate(profile, null, 1.0f, null);
}
}
@@ -0,0 +1,19 @@
package dev.ltms.bridged.placement;
import java.util.List;
import java.util.Set;
import java.util.function.Function;
/**
* Everything a {@link PlacementPolicy} needs to make one selection.
*
* @param defaultProfile profile a {@code fixed} policy should return (may be {@code null})
* @param candidates every configured candidate; the policy filters out those at cap or unreachable
* @param liveCount current live worker count per profile (from the session registry)
* @param unreachable profiles already known to have failed in this spawn attempt
*/
public record PlacementContext(String defaultProfile,
List<PlacementCandidate> candidates,
Function<String, Integer> liveCount,
Set<String> unreachable) {
}
@@ -0,0 +1,12 @@
package dev.ltms.bridged.placement;
/**
* Thrown when a {@link PlacementPolicy} has no candidate available. Kept as a distinct type so
* callers can distinguish "no capacity" from a spawn-time transport failure.
*/
public final class PlacementException extends IllegalStateException {
public PlacementException(String message) {
super(message);
}
}
@@ -0,0 +1,41 @@
package dev.ltms.bridged.placement;
/**
* Factory for the built-in placement policies.
*/
public final class PlacementPolicies {
private PlacementPolicies() {
}
/**
* Resolve a policy name from config. Absent/blank values and {@code "fixed"} return the
* backward-compatible fixed policy; unknown names throw.
*/
public static PlacementPolicy fromName(String name) {
String n = (name == null) ? "" : name.toLowerCase();
if (n.isBlank() || "fixed".equals(n)) {
return fixed();
}
if ("weighted".equals(n)) {
return weighted();
}
if ("round-robin".equals(n)) {
return roundRobin();
}
throw new IllegalArgumentException("unknown placement policy '" + name
+ "' — must be one of: fixed, round-robin, weighted");
}
public static PlacementPolicy fixed() {
return new FixedPlacementPolicy();
}
public static PlacementPolicy weighted() {
return new WeightedRoundRobinPolicy();
}
public static PlacementPolicy roundRobin() {
return new RoundRobinPlacementPolicy();
}
}
@@ -0,0 +1,16 @@
package dev.ltms.bridged.placement;
/**
* How {@code bridged} chooses a worker profile when a spawn names none. Implementations are
* deterministic and unit-testable; the caller (the composite launcher) handles failover retries.
*/
public interface PlacementPolicy {
/**
* Pick one candidate from the configured set.
*
* @throws java.lang.IllegalStateException when no candidate is available, with a message naming
* whether every profile is at capacity or unreachable
*/
PlacementCandidate select(PlacementContext ctx);
}
@@ -0,0 +1,65 @@
package dev.ltms.bridged.placement;
import java.util.ArrayList;
import java.util.List;
/**
* Shared filtering and empty-set reporting used by the built-in placement policies.
*/
final class PlacementPolicyUtil {
private PlacementPolicyUtil() {
}
/**
* Candidates that are not known-unreachable and have not reached their maxLoad.
* A {@code null} maxLoad means unlimited.
*/
static List<PlacementCandidate> available(PlacementContext ctx) {
List<PlacementCandidate> out = new ArrayList<>();
for (PlacementCandidate c : ctx.candidates()) {
if (ctx.unreachable().contains(c.profile())) {
continue;
}
Integer cap = c.maxLoad();
if (cap != null) {
int live = ctx.liveCount().apply(c.profile());
if (live >= cap) {
continue;
}
}
out.add(c);
}
return out;
}
/**
* Build a clear exception describing why every candidate was dropped: all at capacity,
* all unreachable, or a mix.
*/
static PlacementException emptyException(PlacementContext ctx) {
int atCap = 0;
int unreachable = 0;
for (PlacementCandidate c : ctx.candidates()) {
Integer cap = c.maxLoad();
if (ctx.unreachable().contains(c.profile())) {
unreachable++;
} else if (cap != null && ctx.liveCount().apply(c.profile()) >= cap) {
atCap++;
}
}
int total = ctx.candidates().size();
if (total == 0) {
return new PlacementException("no worker profiles configured");
}
if (atCap == total) {
return new PlacementException("all worker profiles are at maxLoad");
}
if (unreachable == total) {
return new PlacementException("all worker profiles are unreachable");
}
return new PlacementException("no worker profile available: " + atCap + " at maxLoad, "
+ unreachable + " unreachable, " + (total - atCap - unreachable) + " remaining");
}
}
@@ -0,0 +1,23 @@
package dev.ltms.bridged.placement;
import java.util.List;
import java.util.concurrent.atomic.AtomicInteger;
/**
* Deterministic round-robin over the profiles that still have capacity and are not known to be
* unreachable. The index advances only on successful selections so the distribution stays even
* across spawns.
*/
final class RoundRobinPlacementPolicy implements PlacementPolicy {
private final AtomicInteger index = new AtomicInteger(0);
@Override
public synchronized PlacementCandidate select(PlacementContext ctx) {
List<PlacementCandidate> available = PlacementPolicyUtil.available(ctx);
if (available.isEmpty()) {
throw PlacementPolicyUtil.emptyException(ctx);
}
return available.get(index.getAndIncrement() % available.size());
}
}
@@ -0,0 +1,50 @@
package dev.ltms.bridged.placement;
import java.util.List;
import java.util.Map;
import java.util.concurrent.ConcurrentHashMap;
/**
* Smooth weighted round-robin (nginx-style): for each selection, add the candidate's weight to
* its current score, pick the highest score, then subtract the total weight of all available
* candidates from the winner. Weights are not required to sum to 1.0; only their ratios matter.
*
* <p>The state is per-policy instance and protected by {@code synchronized} so concurrent spawns
* see a consistent, deterministic sequence rather than interleaving updates.
*/
final class WeightedRoundRobinPolicy implements PlacementPolicy {
private final Map<String, Double> current = new ConcurrentHashMap<>();
@Override
public synchronized PlacementCandidate select(PlacementContext ctx) {
List<PlacementCandidate> available = PlacementPolicyUtil.available(ctx);
if (available.isEmpty()) {
throw PlacementPolicyUtil.emptyException(ctx);
}
double total = 0.0;
for (PlacementCandidate c : available) {
total += c.weight();
}
if (total <= 0.0) {
throw new PlacementException("all available profiles have non-positive weight");
}
PlacementCandidate best = null;
double bestScore = Double.NEGATIVE_INFINITY;
for (PlacementCandidate c : available) {
double score = current.merge(c.profile(), (double) c.weight(), (old, add) -> old + add);
if (score > bestScore) {
bestScore = score;
best = c;
}
}
if (best == null) {
throw new PlacementException("no placement candidate could be selected");
}
current.put(best.profile(), current.get(best.profile()) - total);
return best;
}
}
@@ -11,11 +11,12 @@ import dev.ltms.bridged.metrics.Metrics;
import dev.ltms.bridged.herdr.Agent;
import dev.ltms.bridged.herdr.HerdrClient;
import dev.ltms.bridged.herdr.HerdrException;
import dev.ltms.bridged.inject.WorkerPresence;
import dev.ltms.bridged.inject.MemberPresence;
import dev.ltms.bridged.peer.PeerUnreachableException;
import dev.ltms.bridged.msg.MessageService;
import dev.ltms.bridged.session.SessionManager;
import dev.ltms.bridged.session.WorkerSession;
import dev.ltms.bridged.peer.MemberRole;
import dev.ltms.bridged.session.MemberSession;
import dev.ltms.bridged.session.WorktreeRequest;
import dev.ltms.bridged.peer.PeerLauncher;
import io.javalin.Javalin;
@@ -55,7 +56,7 @@ public final class BridgedApp {
private final PeerLauncher workers;
private final SessionManager sessions; // CB-301: authoritative session registry
private final MessageService messages;
private final WorkerPresence presence; // CB-113: which workers are MCP-connected (available)
private final MemberPresence presence; // CB-113: which workers are MCP-connected (available)
private final HttpServlet mcpServlet; // MCP Streamable-HTTP endpoint, mounted at /mcp (nullable)
private final CallerResolver auth; // CB-501: null → authz not enforced (legacy behaviour)
private final Metrics metrics; // CB-502: null → /metrics not exposed
@@ -67,7 +68,7 @@ public final class BridgedApp {
* behaviour without each needing an auth fixture.
*/
public BridgedApp(HerdrClient herdr, PeerLauncher workers, SessionManager sessions,
MessageService messages, WorkerPresence presence,
MessageService messages, MemberPresence presence,
HttpServlet mcpServlet) {
this(herdr, workers, sessions, messages, presence, mcpServlet, null, null);
}
@@ -79,7 +80,7 @@ public final class BridgedApp {
* the endpoint
*/
public BridgedApp(HerdrClient herdr, PeerLauncher workers, SessionManager sessions,
MessageService messages, WorkerPresence presence,
MessageService messages, MemberPresence presence,
HttpServlet mcpServlet, CallerResolver auth, Metrics metrics) {
this.herdr = herdr;
this.workers = workers;
@@ -115,10 +116,10 @@ public final class BridgedApp {
}
app.get("/sessions", this::sessions);
app.get("/agents", this::agents);
app.get("/workers", this::listWorkers); // CB-304: registry roster + live herdr status
app.get("/profiles", this::profiles); // configured worker profiles
app.post("/workers", this::spawnWorker); // optional ?profile= or {"profile":…}
app.delete("/workers/{paneId}", this::stopWorker);
app.get("/members", this::listMembers); // CB-304: registry roster + live herdr status
app.get("/profiles", this::profiles); // configured backend profiles
app.post("/members", this::spawnMember); // optional ?role=&profile= or {"role":…,"profile":…}
app.delete("/members/{paneId}", this::stopMember);
app.post("/sessions/{id}/message", this::sendMessage); // bridge_send (primary; blocking, wait:false, or answer via turnId)
app.post("/sessions/{id}/reply", this::replyMessage); // bridge_reply (worker)
app.get("/sessions/{id}/replies", this::drainReplies); // drain reply inbox (CB-307)
@@ -221,16 +222,17 @@ public final class BridgedApp {
}
/** CB-304: bridge-owned roster merged with live herdr status by paneId. */
private void listWorkers(Context ctx) {
private void listMembers(Context ctx) {
if (!allow(ctx, Authz.Action.READ, null)) {
return;
}
// CB-519: the registry key is a host-unique id, not the pane coordinate — join on terminal.
Map<String, Agent> live = workers.list().stream()
.map(Agent.class::cast)
.filter(a -> a.paneId() != null)
.collect(Collectors.toMap(Agent::paneId, Function.identity(), (_, b) -> b));
.filter(a -> a.terminalId() != null)
.collect(Collectors.toMap(Agent::terminalId, Function.identity(), (_, b) -> b));
List<Map<String, Object>> out = sessions.roster().stream()
.map(s -> SessionManager.rosterView(s, live.get(s.paneId())))
.map(s -> SessionManager.rosterView(s, live.get(s.terminalId())))
.toList();
ctx.status(200).json(Map.of("workers", out));
}
@@ -250,10 +252,11 @@ public final class BridgedApp {
* body) picks which configured profile; omitted → the default. 403 if the base_url would breach
* the subscription boundary, 400 for an unknown profile.
*/
private void spawnWorker(Context ctx) {
private void spawnMember(Context ctx) {
if (!allow(ctx, Authz.Action.SPAWN, null)) {
return;
}
String role = ctx.queryParam("role");
String profile = ctx.queryParam("profile");
String cwd = ctx.queryParam("cwd");
String worktree = ctx.queryParam("worktree");
@@ -264,6 +267,7 @@ public final class BridgedApp {
String body = ctx.body();
if (!body.isBlank()) {
JsonNode b = mapper.readTree(body);
if (role == null || role.isBlank()) role = b.path("role").asText(null);
if (profile == null || profile.isBlank()) profile = b.path("profile").asText(null);
if (cwd == null || cwd.isBlank()) cwd = b.path("cwd").asText(null);
if (worktree == null || worktree.isBlank()) worktree = b.path("worktree").asText(null);
@@ -274,10 +278,18 @@ public final class BridgedApp {
}
}
WorktreeRequest wt = worktreeRequest(worktree, ticket);
MemberRole memberRole;
try {
memberRole = (role == null || role.isBlank()) ? MemberRole.DEV : MemberRole.parse(role);
} catch (IllegalArgumentException e) {
ctx.status(400).json(Map.of("error", "unknown_role", "detail", e.getMessage()));
return;
}
try {
// No MCP caller over REST, so callerCwd and ownerTerminal are null.
WorkerSession worker = sessions.acquire(blankToNull(profile), blankToNull(cwd), null, null, wt);
ctx.status(201).json(view(worker));
MemberSession member = sessions.acquire(blankToNull(profile), memberRole,
blankToNull(cwd), null, null, wt);
ctx.status(201).json(view(member));
} catch (GuardException e) {
ctx.status(403).json(Map.of("error", "subscription_boundary", "detail", e.getMessage()));
} catch (IllegalArgumentException e) {
@@ -305,7 +317,7 @@ public final class BridgedApp {
}
/** Tear a worker down by pane id. */
private void stopWorker(Context ctx) {
private void stopMember(Context ctx) {
String paneId = ctx.pathParam("paneId");
if (!allow(ctx, Authz.Action.STOP, paneId)) {
return;
@@ -543,7 +555,7 @@ public final class BridgedApp {
}
/** CB-301 projection of an authoritative bridge-owned session. */
private static Map<String, Object> view(WorkerSession s) {
private static Map<String, Object> view(MemberSession s) {
Map<String, Object> m = new LinkedHashMap<>();
m.put("terminalId", s.terminalId());
m.put("paneId", s.paneId());
@@ -27,6 +27,54 @@ public final class GitWorktrees implements Worktrees {
private static final Logger log = LoggerFactory.getLogger(GitWorktrees.class);
/** Project-level MCP config. Present in the repo, so every worktree would otherwise inherit the
* primary's IDE server mounts (a CB-523 worker edited the primary checkout; see the isolation
* javadoc). Neutralized unconditionally. */
private static final String MCP_CONFIG = ".mcp.json";
/** What {@link #isolateToolSurface} writes for {@code .mcp.json}: a valid, explicitly empty server map. */
private static final String NEUTRAL_MCP_CONFIG = "{\n \"mcpServers\": {}\n}\n";
/** OpenCode's repo-level config. Tracked here, so it lands in every worktree, and it mounts the
* primary's gitea and context7 servers with the primary's credentials. Neutralized so the worker
* gets only the config its launcher writes via {@code OPENCODE_CONFIG}.
*
* <p>The reason has changed shape and is now stronger. It used to be a crash: the file carried
* {@code {file:.secrets/...}} references to gitignored files that never reached a worktree, and
* opencode refuses to start on a dangling reference (CB-543). Those credentials now live in one
* shell-level store and the file reads them as {@code {env:...}}, so in a worktree the reference
* resolves instead of failing. That is worse, not better: a member would silently inherit the
* primary's admin-scoped {@code GITEA_ACCESS_TOKEN}. A loud crash became a quiet privilege leak,
* so this entry protects a boundary now rather than papering over a startup error. */
private static final String OPENCODE_CONFIG = "opencode.json";
/** What {@link #isolateToolSurface} writes for {@code opencode.json}: a valid, empty JSON object. */
private static final String NEUTRAL_OPENCODE_CONFIG = "{}\n";
/** Autoenv's repo-level config. Not tracked today, but re-landing it must stay safe: autoenv
* authorizes by path, so a fresh worktree path is always unauthorized and its interactive prompt
* would block every spawn — neutralize it so it can never be committed. */
private static final String AUTOENV_CONFIG = ".autoenv";
/** What {@link #isolateToolSurface} writes for {@code .autoenv}: a valid, empty env file. */
private static final String NEUTRAL_AUTOENV_CONFIG = "";
/**
* A tracked project config that is hostile in a provisioned worktree, and what to replace it
* with. {@link #file} is the repo-relative path; {@link #stub} is a neutral but VALID payload for
* that file's format — a malformed stub would only trade one crash for another;
* {@link #createIfAbsent} keeps {@code .mcp.json}'s long-standing behaviour of writing its stub
* even when the repo carries no such file, whereas the others are only touched when present.
*/
private record WorktreeHostileConfig(String file, String stub, boolean createIfAbsent) {}
/** The worktree-hostile configs neutralized in every provisioned worktree, in order. */
private static final List<WorktreeHostileConfig> WORKTREE_HOSTILE_CONFIGS = List.of(
new WorktreeHostileConfig(MCP_CONFIG, NEUTRAL_MCP_CONFIG, true),
new WorktreeHostileConfig(OPENCODE_CONFIG, NEUTRAL_OPENCODE_CONFIG, false),
new WorktreeHostileConfig(AUTOENV_CONFIG, NEUTRAL_AUTOENV_CONFIG, false)
);
private final String configuredRoot;
private final SecureRandom random = new SecureRandom();
private final AtomicLong seq = new AtomicLong();
@@ -55,9 +103,59 @@ public final class GitWorktrees implements Worktrees {
String wt = path.toAbsolutePath().toString();
log.info("adding worktree branch={} path={} base={}", branch, wt, base);
exec("git", "-C", repoRoot, "worktree", "add", wt, "-b", branch, base);
isolateToolSurface(wt);
return wt;
}
/**
* Neutralize the worktree's worktree-hostile project configs so a worker inherits only the tools
* and environment its launcher mounts (the bridge via {@code --mcp-config}, the opencode config
* via {@code OPENCODE_CONFIG}) — never the primary's.
*
* <p>This is unconditional, and it is not the same job as the parity overlay. The repo's own
* committed {@code .mcp.json} declares the primary's IDE servers, so a fresh checkout mounts them
* whether or not the overlay copies anything; a worker that inherits them navigates and edits
* through tools bound to the <em>primary's</em> IntelliJ project, which silently hands it absolute
* paths outside its own worktree. That is not hypothetical: a CB-523 worker made all 59 of its
* edits in the primary checkout while compiling its worktree, so every build it ran was of code
* that did not contain its changes. {@code opencode.json} is the same trap one tool over — tracked,
* so it lands in every worktree, and it mounts gitea and context7 with the primary's own
* credentials, which a member must never hold. {@code .autoenv} extends the principle to a
* config that is not tracked today: autoenv authorizes by path, so a fresh worktree path is always
* unauthorized and its interactive prompt would block every spawn, so re-landing one must be safe.
*
* <p>Where a config exists it is replaced by a valid neutral stub (an explicitly empty
* map/object, or an empty env file — never a deletion, which would still let a later
* {@code git checkout} restore the hostile copy). The {@code --skip-worktree} bit keeps the
* neutralized copy from ever showing up as a local modification the worker might commit. A config
* the repo does not carry is skipped silently — no stub is invented for a file the repo does not
* have, and one missing file must never fail provisioning.
*/
private void isolateToolSurface(String worktreePath) {
Path root = Path.of(worktreePath).toAbsolutePath().normalize();
for (WorktreeHostileConfig cfg : WORKTREE_HOSTILE_CONFIGS) {
neutralize(root, worktreePath, cfg);
}
}
private void neutralize(Path root, String worktreePath, WorktreeHostileConfig cfg) {
Path target = root.resolve(cfg.file());
if (!Files.exists(target) && !cfg.createIfAbsent()) {
log.debug("{} absent in the worktree — skipping (repo does not carry it)", cfg.file());
return;
}
try {
Files.writeString(target, cfg.stub());
} catch (IOException e) {
throw new WorktreeException("cannot neutralize " + cfg.file() + " in the worktree: "
+ e.getMessage(), e);
}
if (isTracked(root, cfg.file())) {
exec("git", "-C", worktreePath, "update-index", "--skip-worktree", cfg.file());
}
log.debug("neutralized {} — worker tool surface is launcher-mounted only", cfg.file());
}
@Override
public void remove(String repoRoot, String worktreePath) {
Path p = Path.of(worktreePath);
@@ -1,13 +1,22 @@
package dev.ltms.bridged.session;
import dev.ltms.bridged.peer.MemberRole;
/**
* A bridge-owned worker session — the authoritative in-daemon record of a worker this
* process spawned. Immutable; state transitions are performed by replacing the record in
* {@link SessionManager}'s registry.
*
* @param paneId herdr pane handle — the registry key and the argument to teardown
* @param paneId the host-unique opaque id (CB-519) — the registry key and the argument to
* teardown. Despite the historical name this is the {@link
* dev.ltms.bridged.peer.PeerHandle#id()}, a UUID, and is distinct from the
* launcher-private herdr pane coordinate.
* @param terminalId herdr terminal handle — the {@code target} for send/read/status
* @param profile the worker profile name that spawned this session
* @param profile the profile name that spawned this session — WHICH BACKEND
* @param role the contract this member runs under — WHAT IT IS FOR. Independent of
* {@code profile}: a reviewer may run on the same profile as the dev
* whose diff it reads. {@code null} only for a session recorded before
* the role was known
* @param cwd the resolved working directory the worker started in
* @param ownerTerminal the caller that requested this worker ({@code null} = daemon/anon)
* @param spawnedAtNanos {@link System#nanoTime()} when the session was registered
@@ -15,10 +24,11 @@ package dev.ltms.bridged.session;
* @param turnCount number of delegated turns that have been delivered to this session
* @param state current lifecycle state in the one-shot FSM
*/
public record WorkerSession(
public record MemberSession(
String paneId,
String terminalId,
String profile,
MemberRole role,
String cwd,
String ownerTerminal,
long spawnedAtNanos,
@@ -39,20 +49,20 @@ public record WorkerSession(
}
/** Return a copy of this session in {@code state}. */
public WorkerSession withState(State state) {
return new WorkerSession(paneId, terminalId, profile, cwd, ownerTerminal, spawnedAtNanos,
public MemberSession withState(State state) {
return new MemberSession(paneId, terminalId, profile, role, cwd, ownerTerminal, spawnedAtNanos,
lastActivityAtNanos, turnCount, state, worktree, branch);
}
/** Return a copy with {@code lastActivityAtNanos} updated to {@code nowNanos}. */
public WorkerSession withActivity(long nowNanos) {
return new WorkerSession(paneId, terminalId, profile, cwd, ownerTerminal, spawnedAtNanos,
public MemberSession withActivity(long nowNanos) {
return new MemberSession(paneId, terminalId, profile, role, cwd, ownerTerminal, spawnedAtNanos,
nowNanos, turnCount, state, worktree, branch);
}
/** Return a copy with the turn count incremented and activity timestamped at {@code nowNanos}. */
public WorkerSession bumpTurn(long nowNanos) {
return new WorkerSession(paneId, terminalId, profile, cwd, ownerTerminal, spawnedAtNanos,
public MemberSession bumpTurn(long nowNanos) {
return new MemberSession(paneId, terminalId, profile, role, cwd, ownerTerminal, spawnedAtNanos,
nowNanos, turnCount + 1, state, worktree, branch);
}
}
@@ -1,8 +1,10 @@
package dev.ltms.bridged.session;
import dev.ltms.bridged.auth.MemberLifecycle;
import dev.ltms.bridged.herdr.Agent;
import dev.ltms.bridged.inject.TurnListener;
import dev.ltms.bridged.inject.WorkerPresence;
import dev.ltms.bridged.inject.MemberPresence;
import dev.ltms.bridged.peer.MemberRole;
import dev.ltms.bridged.peer.PeerHandle;
import dev.ltms.bridged.peer.PeerLauncher;
import dev.ltms.bridged.peer.SpawnRequest;
@@ -31,7 +33,7 @@ import java.util.function.LongSupplier;
* for {@code release + acquire} with a new distinct pane id.
*
* <p>The manager implements {@link TurnListener} so the injector's turn boundaries drive
* {@code READY → BUSY → DONE} (or {@code FAILED}). It exposes a {@link WorkerPresence} view via
* {@code READY → BUSY → DONE} (or {@code FAILED}). It exposes a {@link MemberPresence} view via
* {@link #asPresence()}: any MCP contact from a worker marks it present and simultaneously
* transitions the session {@code SPAWNING → READY}.
*/
@@ -41,128 +43,241 @@ public final class SessionManager implements TurnListener {
private final PeerLauncher launcher;
private final Worktrees worktrees;
private final ConcurrentHashMap<String /*paneId*/, WorkerSession> registry = new ConcurrentHashMap<>();
private final WorkerPresence presence;
private final ConcurrentHashMap<String /*paneId*/, MemberSession> registry = new ConcurrentHashMap<>();
private final MemberPresence presence;
private final SecureRandom nonceRandom = new SecureRandom();
private final AtomicLong nonceSeq = new AtomicLong();
private final LongSupplier nowNanos;
private final int contextCap;
private final boolean clearAfterTurn;
private volatile MemberLifecycle memberLifecycle = MemberLifecycle.NONE;
/** CB-520: notified with a terminalId on every acquire; no-op until wired. */
private final List<Consumer<String>> acquireListeners = new java.util.concurrent.CopyOnWriteArrayList<>();
/** CB-516: notified with a terminalId on every release; no-op until wired. */
private volatile Consumer<String> releaseListener = _ -> { };
private final List<Consumer<String>> releaseListeners = new java.util.concurrent.CopyOnWriteArrayList<>();
/** Backward-compatible constructor: shared-tree sessions, production git seam. */
public SessionManager(PeerLauncher launcher) {
this(launcher, new GitWorktrees(), System::nanoTime, 0);
this(launcher, new GitWorktrees(), System::nanoTime, 0, false);
}
/** Backward-compatible constructor with an injectable worktree seam. */
public SessionManager(PeerLauncher launcher, Worktrees worktrees) {
this(launcher, worktrees, System::nanoTime, 0);
this(launcher, worktrees, System::nanoTime, 0, false);
}
/** Test constructor with an injectable clock. */
public SessionManager(PeerLauncher launcher, Worktrees worktrees, LongSupplier nowNanos) {
this(launcher, worktrees, nowNanos, 0);
this(launcher, worktrees, nowNanos, 0, false);
}
/** Production constructor with a configured context turn cap. */
public SessionManager(PeerLauncher launcher, Worktrees worktrees, int contextCap) {
this(launcher, worktrees, System::nanoTime, contextCap);
this(launcher, worktrees, System::nanoTime, contextCap, false);
}
public SessionManager(PeerLauncher launcher, Worktrees worktrees, LongSupplier nowNanos,
int contextCap) {
int contextCap) {
this(launcher, worktrees, nowNanos, contextCap, false);
}
public SessionManager(PeerLauncher launcher, Worktrees worktrees, LongSupplier nowNanos,
int contextCap, boolean clearAfterTurn) {
this.launcher = launcher;
this.worktrees = worktrees;
this.presence = new PresenceBridge(this);
this.nowNanos = nowNanos;
this.contextCap = contextCap;
this.clearAfterTurn = clearAfterTurn;
}
/**
* The single {@link WorkerPresence} view of this manager: it records availability and forwards
* The single {@link MemberPresence} view of this manager: it records availability and forwards
* the signal to the {@code SPAWNING → READY} transition. Pass this to the {@code Injector} and
* {@code BridgeMcp} where they previously accepted a plain {@link WorkerPresence}. The same
* {@code BridgeMcp} where they previously accepted a plain {@link MemberPresence}. The same
* instance is returned every call — presence is shared state, so a fresh bridge per call would
* fragment the {@code present} set and lose signals across callers.
*/
public WorkerPresence asPresence() {
public MemberPresence asPresence() {
return presence;
}
/**
* Spawn a worker and register it as {@link WorkerSession.State#SPAWNING}. The caller's
* Spawn a worker and register it as {@link MemberSession.State#SPAWNING}. The caller's
* identity is recorded as {@code ownerTerminal} ({@code null} for daemon/anon callers).
*/
public WorkerSession acquire(String profile, String requestedCwd, String callerCwd,
public MemberSession acquire(String profile, String requestedCwd, String callerCwd,
String ownerTerminal) {
return acquire(profile, requestedCwd, callerCwd, ownerTerminal, null);
}
/**
* Spawn a worker, optionally inside a fresh git worktree. When {@code wt} is non-null the
* worktree is provisioned, parity-overlaid, and its path becomes the worker's cwd. On any
* failure before registration the worktree is removed so no dangling checkout is left.
* Spawn a member, optionally inside a fresh git worktree, defaulting the role to
* {@link MemberRole#DEV}.
*
* <p>{@code DEV} is the right default because it is exactly what the old "worker" meant: an
* unqualified spawn is a unit of implementation work. An architect or a reviewer is always
* asked for on purpose, so neither is ever what a caller silently gets.
*/
public WorkerSession acquire(String profile, String requestedCwd, String callerCwd,
public MemberSession acquire(String profile, String requestedCwd, String callerCwd,
String ownerTerminal, WorktreeRequest wt) {
return acquire(profile, MemberRole.DEV, requestedCwd, callerCwd, ownerTerminal, wt);
}
/**
* Spawn a member, optionally inside a fresh git worktree. When {@code wt} is non-null the
* worktree is provisioned, parity-overlaid, and its path becomes the member's cwd. On any
* failure before registration the worktree is removed so no dangling checkout is left.
*
* @param profile which backend to run on — a {@code profiles:} key
* @param role which contract the member runs under; never {@code null}
*/
public MemberSession acquire(String profile, MemberRole role, String requestedCwd, String callerCwd,
String ownerTerminal, WorktreeRequest wt) {
MemberRole memberRole = (role == null) ? MemberRole.DEV : role;
if (wt == null) {
SpawnRequest req = new SpawnRequest(profile, requestedCwd, callerCwd);
// CB-557: the role must ride on the SpawnRequest, not stay a local. The launcher needs it
// to pick the profile out of that role's pool and to label the tab; a role kept only on
// the MemberSession is recorded after the spawn it was supposed to steer.
SpawnRequest req = new SpawnRequest(profile, requestedCwd, callerCwd, null, null, memberRole);
PeerHandle handle = launcher.spawn(req);
String resolvedProfile = (profile == null || profile.isBlank())
? launcher.defaultProfile() : profile;
String cwd = launcher.effectiveCwd(req);
String resolvedProfile = resolveProfile(handle, profile);
String cwd = launcher.effectiveCwd(new SpawnRequest(resolvedProfile, requestedCwd, callerCwd));
long now = nowNanos.getAsLong();
WorkerSession session = new WorkerSession(
MemberSession session = new MemberSession(
handle.id(),
handle.terminalId(),
resolvedProfile,
memberRole,
cwd,
ownerTerminal,
now,
now,
0,
WorkerSession.State.SPAWNING,
MemberSession.State.SPAWNING,
null,
null);
registry.put(handle.id(), session);
memberLifecycle.acquired(session.role(), session.profile(), session.terminalId());
log.debug("acquired session id={} terminal={} profile={} owner={}",
handle.id(), handle.terminalId(), session.profile(), session.ownerTerminal());
notifyAcquired(session.terminalId());
return session;
}
return acquireWithWorktree(profile, requestedCwd, callerCwd, ownerTerminal, wt);
return acquireWithWorktree(profile, memberRole, requestedCwd, callerCwd, ownerTerminal, wt);
}
/** Tear a worker down by pane id and remove it from the registry. Idempotent. */
public void release(String paneId) {
WorkerSession removed = registry.remove(paneId);
release(paneId, ReleaseCause.COMPLETED);
}
/**
* Core teardown: always stops the worker pane and deregisters the session; whether the worker's
* git worktree is also removed depends on {@code cause}.
*
* <p>CB-544: these are two concerns that used to be fused. Stopping the pane is correct on every
* teardown — the worker process must end. Removing the worktree is a destructive act that is only
* correct for a deliberately-finished teardown (an explicit stop of a completed session, the
* reaper releasing a genuinely idle one, a context-capped or recycled session). A shutdown drain
* must stop panes but preserve worktrees: a worker's uncommitted work exists in exactly one
* place — its worktree — so deleting it while the daemon simply goes down is silent data loss,
* with no copy and no error. Do NOT fuse these back together; the cost of an orphaned worktree
* is a logged path an operator can reclaim, the cost of a deleted one is unrecoverable work.
*/
private void release(String paneId, ReleaseCause cause) {
MemberSession removed = registry.remove(paneId);
boolean preserveWorktree = cause == ReleaseCause.SHUTDOWN;
if (removed != null) {
log.debug("releasing session pane={} terminal={} state={}",
removed.paneId(), removed.terminalId(), removed.state());
memberLifecycle.released(removed.terminalId());
log.debug("releasing session pane={} terminal={} state={} cause={}",
removed.paneId(), removed.terminalId(), removed.state(), cause);
if (preserveWorktree && removed.worktree() != null) {
logPreservedForShutdown(removed);
}
// CB-516: a send still waiting on this worker can never be answered now. Tell the
// listener BEFORE the pane is torn down, so a blocked caller fails fast with a real
// reason instead of sitting on a rendezvous nothing will ever resolve.
notifyReleased(removed.terminalId());
}
launcher.stop(paneId);
if (removed != null && removed.worktree() != null) {
if (removed != null && !preserveWorktree && removed.worktree() != null) {
worktrees.remove(worktrees.repoRoot(removed.cwd()), removed.worktree());
}
}
/**
* Why a session is being released — governs whether its worktree is preserved or removed.
* Worktree removal is reserved for the one case that is genuinely finished; everything else
* must keep the worker's only copy of its work.
*/
public enum ReleaseCause {
/** Deliberate teardown of a finished session. Stops the pane and removes the worktree. */
COMPLETED,
/** Daemon shutdown drain. Stops the pane but PRESERVES the worktree. */
SHUTDOWN
}
/**
* CB-544 shutdown drain log for a worktree we deliberately kept. A session still {@code BUSY}
* when the drain timeout expired was abandoned mid-turn — that work may be uncommitted and is
* the only copy — so the message is loud and points at the path an operator needs to reclaim.
*/
private void logPreservedForShutdown(MemberSession session) {
if (session.state() == MemberSession.State.BUSY) {
log.warn("shutdown drain abandoned BUSY session pane={} terminal={} mid-turn; "
+ "worktree preserved at {}", session.paneId(), session.terminalId(),
session.worktree());
} else {
log.info("shutdown drain preserved worktree at {} for pane={}",
session.worktree(), session.paneId());
}
}
/**
* Register a callback invoked with a session's {@code terminalId} whenever it is acquired
* (CB-520). This is the hook that lets the reply inbox {@code own} a target's queue.
*/
public void onAcquire(Consumer<String> listener) {
if (listener != null) {
acquireListeners.add(listener);
}
}
/**
* Register a callback invoked with a session's {@code terminalId} whenever it is released
* (CB-516). Every teardown path funnels through {@link #release}, so one hook covers the REST
* and MCP stop tools, the idle-TTL reaper, {@code recycle}, and shutdown drain alike.
*
* <p>Set rather than injected because {@code MessageService} — the intended listener — is
* <p>Added rather than injected because {@code MessageService} — one intended listener — is
* constructed after this manager (it needs the injector and rendezvous, which need the session
* presence view this manager exposes). Wiring it at construction would require breaking that
* cycle for one callback.
*/
public void onRelease(Consumer<String> listener) {
this.releaseListener = (listener == null) ? _ -> { } : listener;
if (listener != null) {
releaseListeners.add(listener);
}
}
/** Inject the optional member-slot lifecycle after construction without changing constructors. */
public void setMemberLifecycle(MemberLifecycle memberLifecycle) {
this.memberLifecycle = memberLifecycle == null ? MemberLifecycle.NONE : memberLifecycle;
}
/** A listener failure must never prevent the acquisition it is reacting to. */
private void notifyAcquired(String terminalId) {
if (terminalId == null) {
return;
}
for (Consumer<String> listener : acquireListeners) {
try {
listener.accept(terminalId);
} catch (RuntimeException e) {
log.warn("acquire listener failed for terminal {}: {}", terminalId, e.toString());
}
}
}
/** A listener failure must never prevent the teardown it is reacting to. */
@@ -170,16 +285,18 @@ public final class SessionManager implements TurnListener {
if (terminalId == null) {
return;
}
try {
releaseListener.accept(terminalId);
} catch (RuntimeException e) {
log.warn("release listener failed for terminal {}: {}", terminalId, e.toString());
for (Consumer<String> listener : releaseListeners) {
try {
listener.accept(terminalId);
} catch (RuntimeException e) {
log.warn("release listener failed for terminal {}: {}", terminalId, e.toString());
}
}
}
private WorkerSession acquireWithWorktree(String profile, String requestedCwd, String callerCwd,
private MemberSession acquireWithWorktree(String profile, MemberRole memberRole, String requestedCwd, String callerCwd,
String ownerTerminal, WorktreeRequest wt) {
String resolvedProfile = (profile == null || profile.isBlank())
String preResolvedProfile = (profile == null || profile.isBlank())
? launcher.defaultProfile() : profile;
// CB-507: resolve through the launcher's CB-112 chain (requested → profile cwd → caller →
// daemon cwd → "."), never the raw args. A plain REST spawn supplies neither a requested
@@ -187,14 +304,14 @@ public final class SessionManager implements TurnListener {
// `git -C null` on the command line — an NPE out of ProcessBuilder, surfacing as HTTP 500.
// The non-worktree path always used this chain; only this branch was missed.
String repoRoot = worktrees.repoRoot(
launcher.effectiveCwd(new SpawnRequest(resolvedProfile, requestedCwd, callerCwd)));
launcher.effectiveCwd(new SpawnRequest(preResolvedProfile, requestedCwd, callerCwd)));
String branch = "worker/" + slug(wt.ticketSlug()) + "-" + nonce();
String path = null;
PeerHandle handle;
try {
path = worktrees.add(repoRoot, branch, wt.baseRef());
worktrees.overlayParity(repoRoot, path, launcher.parityOverlay(resolvedProfile));
handle = launcher.spawn(new SpawnRequest(profile, path, callerCwd));
worktrees.overlayParity(repoRoot, path, launcher.parityOverlay(preResolvedProfile));
handle = launcher.spawn(new SpawnRequest(profile, path, callerCwd, null, null, memberRole));
} catch (RuntimeException e) {
if (path != null) {
try {
@@ -205,22 +322,27 @@ public final class SessionManager implements TurnListener {
}
throw e;
}
String resolvedProfile = resolveProfile(handle, profile);
String cwd = launcher.effectiveCwd(new SpawnRequest(resolvedProfile, path, callerCwd));
long now = nowNanos.getAsLong();
WorkerSession session = new WorkerSession(
MemberSession session = new MemberSession(
handle.id(),
handle.terminalId(),
resolvedProfile,
resolveCwd(path, profile, callerCwd),
memberRole,
cwd,
ownerTerminal,
now,
now,
0,
WorkerSession.State.SPAWNING,
MemberSession.State.SPAWNING,
path,
branch);
registry.put(handle.id(), session);
memberLifecycle.acquired(session.role(), session.profile(), session.terminalId());
log.debug("acquired worktree session id={} terminal={} profile={} branch={} path={}",
handle.id(), handle.terminalId(), session.profile(), session.branch(), session.worktree());
notifyAcquired(session.terminalId());
return session;
}
@@ -232,12 +354,28 @@ public final class SessionManager implements TurnListener {
return String.format("%06x", nonceRandom.nextInt(1 << 24)) + "-" + nonceSeq.incrementAndGet();
}
/**
* The profile to record for a session. A launcher that performed dynamic selection tells us
* the actual profile via {@link PeerHandle#profile()}; otherwise fall back to what the caller
* requested (or the launcher's default for a no-profile spawn).
*/
private String resolveProfile(PeerHandle handle, String requestedProfile) {
String fromHandle = handle.profile();
if (fromHandle != null && !fromHandle.isBlank()) {
return fromHandle;
}
if (requestedProfile != null && !requestedProfile.isBlank()) {
return requestedProfile;
}
return launcher.defaultProfile();
}
/**
* Release the old session and acquire a fresh one with the same profile and working directory.
* The new session is guaranteed to have a pane id distinct from the old one (no-reuse invariant).
*/
public WorkerSession recycle(String paneId) {
WorkerSession old = registry.get(paneId);
public MemberSession recycle(String paneId) {
MemberSession old = registry.get(paneId);
if (old == null) {
throw new IllegalArgumentException("no session for paneId " + paneId);
}
@@ -246,12 +384,12 @@ public final class SessionManager implements TurnListener {
}
/** The session for {@code paneId}, if it is still registered and not released. */
public Optional<WorkerSession> get(String paneId) {
public Optional<MemberSession> get(String paneId) {
return Optional.ofNullable(registry.get(paneId));
}
/** Bridge-owned roster: all registered sessions (acquired minus released). */
public List<WorkerSession> roster() {
public List<MemberSession> roster() {
return List.copyOf(registry.values());
}
@@ -259,11 +397,15 @@ public final class SessionManager implements TurnListener {
* CB-304 merged roster+live view. The registry is authoritative for worktree, branch,
* profile, owner, and state; the optional live agent supplies the herdr-reported status.
*/
public static Map<String, Object> rosterView(WorkerSession session, Agent live) {
public static Map<String, Object> rosterView(MemberSession session, Agent live) {
Map<String, Object> m = new LinkedHashMap<>();
m.put("sessionId", session.terminalId());
m.put("paneId", session.paneId());
// Both axes, always: profile says which backend this member runs on, role says what it is
// for. A lead reading the roster needs both — two rows may share a profile and still be
// allowed to do entirely different things.
m.put("profile", session.profile());
m.put("role", session.role() == null ? "dev" : session.role().wireName());
m.put("state", session.state().name().toLowerCase());
if (session.worktree() != null) {
m.put("worktree", session.worktree());
@@ -280,7 +422,7 @@ public final class SessionManager implements TurnListener {
/** Lifecycle hook: worker became available on the bridge MCP. */
void onReady(String terminalId) {
transitionByTerminal(terminalId, WorkerSession.State.SPAWNING, WorkerSession.State.READY);
transitionByTerminal(terminalId, MemberSession.State.SPAWNING, MemberSession.State.READY);
}
/**
@@ -290,13 +432,13 @@ public final class SessionManager implements TurnListener {
*/
@Override
public void onDelivered(String target) {
WorkerSession current = findByTerminal(target);
MemberSession current = findByTerminal(target);
if (current == null) return;
if (current.state() != WorkerSession.State.READY && current.state() != WorkerSession.State.DONE) {
if (current.state() != MemberSession.State.READY && current.state() != MemberSession.State.DONE) {
return;
}
long now = nowNanos.getAsLong();
WorkerSession updated = current.withState(WorkerSession.State.BUSY).bumpTurn(now);
MemberSession updated = current.withState(MemberSession.State.BUSY).bumpTurn(now);
if (replace(current, updated)) {
log.debug("session transitioned terminal={} pane={} {} -> BUSY turn={}",
target, current.paneId(), current.state(), updated.turnCount());
@@ -306,16 +448,42 @@ public final class SessionManager implements TurnListener {
/** Lifecycle hook: the worker's delegated turn completed successfully. */
@Override
public void onTurnComplete(String target) {
WorkerSession current = findByTerminal(target);
if (current == null || current.state() != WorkerSession.State.BUSY) return;
completeTurn(target, false);
}
@Override
public boolean hasPostTurnAction(String target) {
if (!clearAfterTurn) return false;
MemberSession current = findByTerminal(target);
return current != null && current.state() == MemberSession.State.BUSY
&& (contextCap <= 0 || current.turnCount() < contextCap);
}
@Override
public boolean onTurnCompleteWithPostAction(String target) {
return completeTurn(target, true);
}
private boolean completeTurn(String target, boolean startContextReset) {
MemberSession current = findByTerminal(target);
if (current == null || current.state() != MemberSession.State.BUSY) return false;
long now = nowNanos.getAsLong();
WorkerSession updated = current.withState(WorkerSession.State.DONE).withActivity(now);
MemberSession updated = current.withState(MemberSession.State.DONE).withActivity(now);
if (replace(current, updated)) {
log.debug("session transitioned terminal={} pane={} BUSY -> DONE turn={}",
target, current.paneId(), updated.turnCount());
}
if (contextCap > 0 && updated.turnCount() >= contextCap) {
release(current.paneId());
return false;
}
if (!startContextReset || !clearAfterTurn) return false;
try {
return launcher.clearContext(current.paneId());
} catch (RuntimeException e) {
log.warn("context reset failed for terminal={} pane={}; continuing without reset: {}",
target, current.paneId(), e.getMessage());
return false;
}
}
@@ -327,10 +495,10 @@ public final class SessionManager implements TurnListener {
/** Lifecycle hook: the worker vanished or was dropped mid-life. */
void onFailed(String target) {
WorkerSession current = findByTerminal(target);
MemberSession current = findByTerminal(target);
if (current == null) return;
if (current.state() == WorkerSession.State.RELEASED) return;
if (replace(current, current.withState(WorkerSession.State.FAILED))) {
if (current.state() == MemberSession.State.RELEASED) return;
if (replace(current, current.withState(MemberSession.State.FAILED))) {
log.debug("session marked failed terminal={} pane={}", target, current.paneId());
}
}
@@ -343,8 +511,8 @@ public final class SessionManager implements TurnListener {
int reapIdle(long idleTtlNanos) {
long now = nowNanos.getAsLong();
int reaped = 0;
for (WorkerSession s : roster()) {
if (s.state() != WorkerSession.State.READY && s.state() != WorkerSession.State.DONE) {
for (MemberSession s : roster()) {
if (s.state() != MemberSession.State.READY && s.state() != MemberSession.State.DONE) {
continue;
}
if (now - s.lastActivityAtNanos() > idleTtlNanos) {
@@ -356,19 +524,25 @@ public final class SessionManager implements TurnListener {
}
/**
* Gracefully drain all registered sessions. For each session that is {@code BUSY}, poll up to
* {@code timeoutNanos} for it to leave {@code BUSY}, then release it regardless. Non-busy
* sessions are released immediately. A failure releasing one session is logged and does not
* abort the rest.
* Gracefully drain all registered sessions on daemon shutdown. For each session that is
* {@code BUSY}, poll up to {@code timeoutNanos} for it to leave {@code BUSY}, then release it
* regardless. Non-busy sessions are released immediately. A failure releasing one session is
* logged and does not abort the rest.
*
* <p>CB-544: this is a {@link ReleaseCause#SHUTDOWN} release — the worker's pane is stopped
* (the process must end) but its worktree is preserved and its path logged. Shutdown is never
* a reason to delete a worker's only copy of its uncommitted work. A session still {@code BUSY}
* when the timeout expired is abandoned mid-turn and logged loudly so an operator can find its
* kept worktree.
*/
void drainAll(long timeoutNanos) {
long deadline = System.nanoTime() + timeoutNanos;
for (WorkerSession s : roster()) {
for (MemberSession s : roster()) {
try {
if (s.state() == WorkerSession.State.BUSY) {
if (s.state() == MemberSession.State.BUSY) {
while (System.nanoTime() < deadline) {
WorkerSession current = registry.get(s.paneId());
if (current == null || current.state() != WorkerSession.State.BUSY) {
MemberSession current = registry.get(s.paneId());
if (current == null || current.state() != MemberSession.State.BUSY) {
break;
}
try {
@@ -380,7 +554,7 @@ public final class SessionManager implements TurnListener {
}
}
}
release(s.paneId());
release(s.paneId(), ReleaseCause.SHUTDOWN);
} catch (RuntimeException e) {
log.warn("drain failed for pane={}; continuing with remaining sessions", s.paneId(), e);
}
@@ -401,16 +575,28 @@ public final class SessionManager implements TurnListener {
return registry.size();
}
private WorkerSession findByTerminal(String terminalId) {
for (WorkerSession s : registry.values()) {
/**
* The registered session owning {@code terminalId}, or {@code null} if none does.
*
* <p>A null {@code terminalId} is a normal input, not a caller bug: every lifecycle hook here is
* fed from the MCP transport, where the <em>primary</em> resolves to a {@link
* dev.ltms.bridged.auth.Principal} with no terminal. {@code BridgeMcp} documents that contact as
* a no-op, and {@link dev.ltms.bridged.inject.MemberPresence#markPresent} honours it — but
* {@code PresenceBridge} then forwards the same null here. Matching on a null id can never
* succeed anyway (a registered session always has a terminal), so answer "no match" rather than
* throwing: an NPE on this path takes down an unrelated tool call for the primary.
*/
private MemberSession findByTerminal(String terminalId) {
if (terminalId == null) return null;
for (MemberSession s : registry.values()) {
if (terminalId.equals(s.terminalId())) return s;
}
return null;
}
private void transitionByTerminal(String terminalId, WorkerSession.State from,
WorkerSession.State to) {
WorkerSession current = findByTerminal(terminalId);
private void transitionByTerminal(String terminalId, MemberSession.State from,
MemberSession.State to) {
MemberSession current = findByTerminal(terminalId);
if (current == null || current.state() != from) return;
long now = nowNanos.getAsLong();
if (replace(current, current.withState(to).withActivity(now))) {
@@ -419,16 +605,12 @@ public final class SessionManager implements TurnListener {
}
}
private boolean replace(WorkerSession expected, WorkerSession updated) {
private boolean replace(MemberSession expected, MemberSession updated) {
return registry.replace(expected.paneId(), expected, updated);
}
private String resolveCwd(String requestedCwd, String profileName, String callerCwd) {
return launcher.effectiveCwd(new SpawnRequest(profileName, requestedCwd, callerCwd));
}
/** WorkerPresence bridge that also drives the manager's READY transition. */
private static final class PresenceBridge extends WorkerPresence {
/** MemberPresence bridge that also drives the manager's READY transition. */
private static final class PresenceBridge extends MemberPresence {
private final SessionManager sessions;
PresenceBridge(SessionManager sessions) {
@@ -437,6 +619,9 @@ public final class SessionManager implements TurnListener {
@Override
public void markPresent(String terminal) {
if (terminal == null || terminal.isBlank()) {
return; // the primary's contact carries no worker terminal — not a readiness signal
}
super.markPresent(terminal);
sessions.onReady(terminal);
}
@@ -1,195 +0,0 @@
package dev.ltms.bridged.worker;
import dev.ltms.bridged.config.BridgedConfig;
import dev.ltms.bridged.guard.SubscriptionGuard;
import dev.ltms.bridged.herdr.Agent;
import dev.ltms.bridged.herdr.AgentControl;
import dev.ltms.bridged.herdr.WorkspaceControl;
import dev.ltms.bridged.peer.Capability;
import java.util.EnumSet;
import java.util.List;
import java.util.Map;
import java.util.Set;
import java.util.function.Function;
import java.util.function.LongSupplier;
/**
* The {@link HerdrPeerLauncher} adapter for <strong>Claude Code</strong> — the safe path from a
* delegation request to a running off-subscription Claude.
*
* <p>Everything transport-related (tab/pane placement, the CB-306 spawn-readiness gate, unique
* naming, CB-117 orphan reap, teardown, listing, cwd resolution) lives in the base. This class
* supplies only the two Claude-specific seams:
* <ul>
* <li>the {@code claude} name prefix (so reap matches {@code claude-*} panes, never another
* adapter's), and</li>
* <li>{@link #buildLaunch}, which encodes the subscription boundary: build the worker env with
* {@code ANTHROPIC_BASE_URL}, assert that host is on the allowlist <em>before</em> touching
* herdr, and mount the bridge MCP + reply charter as inline launch flags. A worker's base_url
* lives in the env map handed to herdr and nowhere else; {@code bridged}'s own environment is
* never mutated, and nothing is written to the worker's profile.</li>
* </ul>
*/
public final class ClaudeCodeLauncher extends HerdrPeerLauncher {
/** Label prefix for this adapter's herdr agent names (drives naming + orphan reap). */
private static final String NAME_PREFIX = "claude";
private final SubscriptionGuard guard;
/**
* Standing instruction appended to the worker's system prompt so it returns its result via
* {@code bridge_reply}. Injected as a launch flag, so nothing is written to the worker's
* profile — it is guidance, and a worker that never replies is caught by the send's timeout.
*/
static final String REPLY_CHARTER =
"You are an off-subscription worker in the claude-bridge fleet. Every message you "
+ "receive arrives through the bridge, and the ONLY channel back to the sender is the "
+ "bridge_reply MCP tool. Text you write in your terminal is NOT sent anywhere — the "
+ "sender cannot see your screen, so an in-terminal answer is silently discarded. "
+ "Therefore you MUST end EVERY turn by calling bridge_reply with `content` set to your "
+ "complete response. This holds for every message without exception — tasks, questions, "
+ "clarifications, acknowledgements, and ordinary back-and-forth conversation. Call "
+ "bridge_reply exactly once, as the final action of your turn, with your full answer in "
+ "`content`; never wait for confirmation first. If you end a turn without calling "
+ "bridge_reply, the sender receives nothing and the exchange stalls.";
/**
* Production constructor — disables the spawn-ready gate ({@code spawnReadyTimeoutMs == 0}) so
* existing deployments and tests keep the legacy non-blocking spawn semantics.
*/
public ClaudeCodeLauncher(AgentControl agents, WorkspaceControl spaces, SubscriptionGuard guard,
Map<String, BridgedConfig.Worker> profiles, String defaultProfile,
Function<String, String> env) {
this(agents, spaces, guard, profiles, defaultProfile, env, 0,
System::currentTimeMillis, () -> sleepUninterruptibly(300));
}
/**
* Production constructor with spawn-ready gate enabled. The gate polls {@code agents.status()}
* until the pane reports an injectable state or {@code spawnReadyTimeoutMs} elapses.
*/
public ClaudeCodeLauncher(AgentControl agents, WorkspaceControl spaces, SubscriptionGuard guard,
Map<String, BridgedConfig.Worker> profiles, String defaultProfile,
Function<String, String> env,
long spawnReadyTimeoutMs, long spawnReadyPollMs) {
this(agents, spaces, guard, profiles, defaultProfile, env,
spawnReadyTimeoutMs,
System::currentTimeMillis, () -> sleepUninterruptibly(spawnReadyPollMs));
}
/**
* Full testability constructor. Every injectable collaborator is explicit so unit tests supply
* fakes for the clock ({@code nowMillis}) and poll-loop wait ({@code sleeper}). The
* {@code sleeper} is never called when the gate is disabled ({@code spawnReadyTimeoutMs == 0}).
*
* @param agents herdr agent control (start, status, close)
* @param spaces workspace / tab control (ensure, create, close)
* @param guard subscription-boundary guard (checked before spawning)
* @param profiles configured worker profiles
* @param defaultProfile profile a no-argument spawn uses (nullable)
* @param env host env lookup (injectable for tests)
* @param spawnReadyTimeoutMs max ms to wait for injectable state (0 disables the gate)
* @param nowMillis monotonic clock source (e.g. {@code System::currentTimeMillis})
* @param sleeper sleep/wait hook (e.g. {@code () -> Thread.sleep(pollMs)}); it
* already encodes the poll interval, so the 8th positional argument
* (poll ms) is accepted for API symmetry but otherwise unused here
*/
public ClaudeCodeLauncher(AgentControl agents, WorkspaceControl spaces, SubscriptionGuard guard,
Map<String, BridgedConfig.Worker> profiles, String defaultProfile,
Function<String, String> env,
long spawnReadyTimeoutMs,
LongSupplier nowMillis, Runnable sleeper) {
super(NAME_PREFIX, agents, spaces, profiles, defaultProfile, env,
spawnReadyTimeoutMs, nowMillis, sleeper);
this.guard = guard;
}
/**
* {@inheritDoc}
*
* <p>The spawn sequence encodes the subscription boundary: assert the profile's base_url is on
* the allowlist <em>before</em> any herdr call, then build the worker env with
* {@code ANTHROPIC_*}, the parity-neutral git-forge grant, and the bridge MCP + reply charter
* mounted as inline launch flags.
*/
@Override
protected Launch buildLaunch(BridgedConfig.Worker cfg) {
String baseUrl = cfg.baseUrl();
guard.assertWorker(baseUrl); // hard stop before we spawn anything
Map<String, String> workerEnv = baseEnv(cfg);
workerEnv.put("ANTHROPIC_BASE_URL", baseUrl);
putIfPresent(workerEnv, "ANTHROPIC_MODEL", cfg.model());
putIfPresent(workerEnv, "CLAUDE_CONFIG_DIR", cfg.configDir());
putIfPresent(workerEnv, "ANTHROPIC_AUTH_TOKEN", env.apply(cfg.tokenEnv()));
applyGitToken(workerEnv, cfg);
return new Launch(workerEnv, argvWithBridge(cfg));
}
/**
* The launch argv, plus — when {@code worker.mcpUrl} is set — inline {@code --mcp-config} for
* the bridge server and {@code --append-system-prompt} for the {@link #REPLY_CHARTER}. Neither
* touches the profile's config; both are pure command-line flags. This inline-flag mount is
* Claude Code specific — other adapters mount MCP and instructions their own way.
*/
private List<String> argvWithBridge(BridgedConfig.Worker cfg) {
if (!cfg.hasMcp()) {
return cfg.argv();
}
String mcpJson = "{\"mcpServers\":{\"bridge\":{\"type\":\"http\",\"url\":\""
+ cfg.mcpUrl() + "\"}}}";
List<String> argv = mutableArgv(cfg.argv());
argv.add("--mcp-config");
argv.add(mcpJson);
argv.add("--append-system-prompt");
argv.add(REPLY_CHARTER);
return argv;
}
// --- Agent-returning convenience spawns (used by callers/tests that want the herdr Agent) ---
/** Spawn a worker for the default profile in the resolved default cwd. */
public Agent spawn() {
return spawnInternal(null, null, null);
}
/** Spawn a worker for a named profile (null → default) in the resolved default cwd. */
public Agent spawn(String profileName) {
return spawnInternal(profileName, null, null);
}
/** Spawn a worker for a named profile with an explicit requested/caller cwd (CB-112). */
public Agent spawn(String profileName, String requestedCwd, String callerCwd) {
return spawnInternal(profileName, requestedCwd, callerCwd);
}
// --- capabilities --------------------------------------------------------------------------
@Override
public Set<Capability> capabilities() {
Set<Capability> caps = EnumSet.of(Capability.MID_TURN_ASK, Capability.WORKTREE, Capability.ORPHAN_REAP);
if (hasGitTokenProfile()) {
caps.add(Capability.SELF_PR);
}
return Set.copyOf(caps);
}
/** Whether any configured profile opts into a git-forge token (required for {@link Capability#SELF_PR}). */
private boolean hasGitTokenProfile() {
return profileConfigs().stream().anyMatch(BridgedConfig.Worker::hasGitToken);
}
// --- CB-117 reap predicate (Claude prefix), kept for direct unit testing -------------------
/**
* Whether {@code name} is a Claude Code bridge worker started by a <em>different</em> process
* than {@code currentNonce}. A thin {@code claude}-prefix binding of
* {@link HerdrPeerLauncher#isForeignWorker(String, String, String)}.
*/
static boolean isForeignWorker(String name, String currentNonce) {
return HerdrPeerLauncher.isForeignWorker(NAME_PREFIX, name, currentNonce);
}
}
@@ -1,160 +0,0 @@
package dev.ltms.bridged.worker;
import dev.ltms.bridged.herdr.Agent;
import dev.ltms.bridged.peer.Capability;
import dev.ltms.bridged.peer.PeerHandle;
import dev.ltms.bridged.peer.PeerLauncher;
import dev.ltms.bridged.peer.SpawnRequest;
import org.slf4j.Logger;
import org.slf4j.LoggerFactory;
import java.util.EnumSet;
import java.util.LinkedHashMap;
import java.util.List;
import java.util.Map;
import java.util.Set;
import java.util.concurrent.ConcurrentHashMap;
/**
* The {@link PeerLauncher} the core actually holds when more than one adapter is configured — a thin
* router in front of one {@link HerdrPeerLauncher} per peer {@code kind} (Claude Code, opencode, …).
* It owns no transport of its own; it dispatches each SPI call to the delegate that owns the profile
* involved, and fans the fleet-wide queries (list/reap/caps/profiles) across all delegates.
*
* <p>Routing rules:
* <ul>
* <li><strong>By profile</strong> — {@link #spawn}, {@link #effectiveCwd}, {@link #parityOverlay}
* resolve the profile (a null/blank name → the global {@link #defaultProfile}) and delegate to
* the single adapter that declares it. Profiles partition cleanly across adapters: the
* constructor rejects a name claimed by two.</li>
* <li><strong>By pane id</strong> — {@link #stop} routes to the adapter that spawned that pane
* (recorded at spawn time). A pane the composite never spawned (only real for a caller that
* hand-rolls an id) falls back to the first delegate; teardown is pane-id addressed and
* tab cleanup is single-occupant guarded, so it is safe either way.</li>
* <li><strong>Fleet-wide</strong> — {@link #reapOrphanWorkers} and {@link #capabilities} fan out
* and combine. {@link #list} is deduplicated by pane id because every herdr-backed delegate
* shares one herdr connection and so reports the same global agent set.</li>
* </ul>
*/
public final class CompositePeerLauncher implements PeerLauncher {
private static final Logger log = LoggerFactory.getLogger(CompositePeerLauncher.class);
private final List<HerdrPeerLauncher> delegates;
private final Map<String, HerdrPeerLauncher> byProfile;
private final String defaultProfile;
/** paneId → the delegate that spawned it, so {@link #stop} tears down through the right adapter. */
private final Map<String, HerdrPeerLauncher> spawnedBy = new ConcurrentHashMap<>();
/**
* @param delegates one adapter per configured peer kind; must be non-empty and declare
* disjoint profile-name sets
* @param defaultProfile the profile a no-argument spawn resolves to (may be null)
* @throws IllegalArgumentException if {@code delegates} is empty or two adapters claim one profile
*/
public CompositePeerLauncher(List<HerdrPeerLauncher> delegates, String defaultProfile) {
if (delegates.isEmpty()) {
throw new IllegalArgumentException("at least one peer adapter must be configured");
}
this.delegates = List.copyOf(delegates);
this.defaultProfile = defaultProfile;
Map<String, HerdrPeerLauncher> index = new LinkedHashMap<>();
for (HerdrPeerLauncher d : this.delegates) {
for (String profile : d.profiles()) {
HerdrPeerLauncher prev = index.putIfAbsent(profile, d);
if (prev != null) {
throw new IllegalArgumentException(
"worker profile '" + profile + "' is claimed by two peer adapters");
}
}
}
this.byProfile = Map.copyOf(index);
}
/** The adapter owning {@code profileName} (null/blank → the default). Throws on an unknown profile. */
private HerdrPeerLauncher route(String profileName) {
String resolved = (profileName == null || profileName.isBlank()) ? defaultProfile : profileName;
if (resolved == null) {
// No profile and no default configured — hand to the first delegate so it raises the
// same "no default" error it would on its own; keeps the SPI contract single-sourced.
return delegates.getFirst();
}
HerdrPeerLauncher d = byProfile.get(resolved);
if (d == null) {
throw new IllegalArgumentException("unknown worker profile: " + resolved);
}
return d;
}
@Override
public PeerHandle spawn(SpawnRequest req) {
HerdrPeerLauncher d = route(req.profileName());
PeerHandle handle = d.spawn(req);
spawnedBy.put(handle.id(), d);
return handle;
}
@Override
public String effectiveCwd(SpawnRequest req) {
return route(req.profileName()).effectiveCwd(req);
}
@Override
public List<String> parityOverlay(String profileName) {
return route(profileName).parityOverlay(profileName);
}
@Override
public void stop(String id) {
HerdrPeerLauncher d = spawnedBy.remove(id);
if (d == null) {
log.debug("stop({}) — no recorded owner, routing to the first adapter (pane-addressed)", id);
d = delegates.getFirst();
}
d.stop(id);
}
@Override
public Set<String> profiles() {
return byProfile.keySet();
}
@Override
public String defaultProfile() {
return defaultProfile;
}
/** Every herdr agent, deduplicated by pane id (all delegates share one herdr and list globally). */
@Override
public List<Agent> list() {
Map<String, Agent> byPane = new LinkedHashMap<>();
for (HerdrPeerLauncher d : delegates) {
for (Agent a : d.list()) {
if (a.paneId() != null) {
byPane.putIfAbsent(a.paneId(), a);
}
}
}
return List.copyOf(byPane.values());
}
@Override
public int reapOrphanWorkers() {
int reaped = 0;
for (HerdrPeerLauncher d : delegates) {
reaped += d.reapOrphanWorkers();
}
return reaped;
}
/** The union of every adapter's capabilities — a capability any adapter offers, the fleet offers. */
@Override
public Set<Capability> capabilities() {
EnumSet<Capability> caps = EnumSet.noneOf(Capability.class);
for (HerdrPeerLauncher d : delegates) {
caps.addAll(d.capabilities());
}
return Set.copyOf(caps);
}
}
@@ -0,0 +1,86 @@
package dev.ltms.bridged;
import dev.ltms.bridged.inject.MemberPresence;
import org.junit.jupiter.api.DisplayName;
import org.junit.jupiter.api.Test;
import java.util.HashMap;
import java.util.Map;
import java.util.function.Predicate;
import java.util.function.Supplier;
import static org.junit.jupiter.api.Assertions.assertFalse;
import static org.junit.jupiter.api.Assertions.assertTrue;
/**
* CB-534: the injector's readiness gate must open for a lead as well as for a present worker.
*
* <p>The bug these cover was silent and slow: a lead was never marked present (only workers are), so
* every lead→lead delivery sat on the gate for the full readiness grace and failed ~60s later without
* a keystroke ever reaching the pane.
*/
class BridgedDeliverabilityTest {
private static Supplier<Map<String, String>> leads(Map<String, String> m) {
return () -> m;
}
@Test
@DisplayName("a worker that has connected its MCP is deliverable")
void presentWorkerIsDeliverable() {
MemberPresence presence = new MemberPresence();
presence.markPresent("term_worker");
assertTrue(Bridged.deliverableTo(presence, leads(Map.of())).test("term_worker"));
}
@Test
@DisplayName("a worker still in its boot window is held back")
void absentWorkerIsNotDeliverable() {
assertFalse(Bridged.deliverableTo(new MemberPresence(), leads(Map.of())).test("term_booting"));
}
@Test
@DisplayName("a lead is deliverable without ever being marked present")
void leadIsDeliverableWithoutPresence() {
MemberPresence presence = new MemberPresence();
Predicate<String> deliverable =
Bridged.deliverableTo(presence, leads(Map.of("term_lead", "opus-5.0")));
assertFalse(presence.isPresent("term_lead"), "a lead is never enrolled in worker presence");
assertTrue(deliverable.test("term_lead"), "…and must be deliverable anyway");
}
@Test
@DisplayName("an unknown terminal is deliverable to neither")
void strangerIsNotDeliverable() {
MemberPresence presence = new MemberPresence();
presence.markPresent("term_worker");
assertFalse(Bridged.deliverableTo(presence, leads(Map.of("term_lead", "opus-5.0")))
.test("term_stranger"));
}
@Test
@DisplayName("a lead discovered after startup becomes deliverable with no restart")
void leadSetIsReadThroughOnEveryCall() {
Map<String, String> discovered = new HashMap<>();
Predicate<String> deliverable = Bridged.deliverableTo(new MemberPresence(), leads(discovered));
assertFalse(deliverable.test("term_late"));
discovered.put("term_late", "gpt-sol-5.6"); // leadScan picks up a newly labelled tab
assertTrue(deliverable.test("term_late"), "the supplier must be re-read, not snapshotted");
}
@Test
@DisplayName("forgetting a torn-down worker does not strip a lead of its deliverability")
void forgetDoesNotDisarmALead() {
MemberPresence presence = new MemberPresence();
Predicate<String> deliverable =
Bridged.deliverableTo(presence, leads(Map.of("term_lead", "opus-5.0")));
presence.forget("term_lead"); // the injector's cleanup path runs against every target
assertTrue(deliverable.test("term_lead"));
}
}
@@ -12,6 +12,8 @@ class AuthzTest {
private static final Principal WORKER_A = Principal.worker("term_a", 200);
private static final Principal WORKER_B = Principal.worker("term_b", 300);
private static final Principal ANON = Principal.anonymous();
private static final Principal ARCH_DESIGN = Principal.architect("lead-designer", "term_design", 400);
private static final Principal ARCH_OTHER = Principal.architect("reviewer", "term_review", 500);
@Test
void anonymousIsAuthorizedForNothing() {
@@ -61,6 +63,48 @@ class AuthzTest {
"an absent session id must not satisfy the own-session rule");
}
// ── CB-548: the architect matrix ───────────────────────────────────────────────────────────
@Test
void anArchitectMaySendButNotSpawnStopOrDrain() {
assertTrue(Authz.permits(ARCH_DESIGN, SEND, "term_worker"),
"delegating a turn to a worker IS the architect's job");
assertTrue(Authz.permits(ARCH_DESIGN, SEND, null));
for (Authz.Action a : new Authz.Action[]{SPAWN, STOP, DRAIN}) {
assertFalse(Authz.permits(ARCH_DESIGN, a, null),
"an architect must not " + a + " — fleet lifecycle is the primary's alone, so "
+ "a coordinator cannot also stand up or tear down the fleet");
}
}
@Test
void anArchitectMayReplyAndAskOnlyAsItsOwnPane() {
assertTrue(Authz.permits(ARCH_DESIGN, REPLY, "term_design"), "its own pane is its own");
assertTrue(Authz.permits(ARCH_DESIGN, ASK, "term_design"));
assertFalse(Authz.permits(ARCH_DESIGN, REPLY, "term_review"),
"architect 'lead-designer' must not reply on reviewer's pane");
assertFalse(Authz.permits(ARCH_OTHER, ASK, "term_design"),
"reviewer must not ask as lead-designer — no terminal is another's");
assertFalse(Authz.permits(ARCH_DESIGN, REPLY, null),
"an absent target must not pass the own-session rule");
}
@Test
void anArchitectMayReadAndScrapeMetrics() {
assertTrue(Authz.permits(ARCH_DESIGN, READ, null));
assertTrue(Authz.permits(ARCH_DESIGN, METRICS, null));
}
@Test
void anArchitectIsNotCountedAsPrimaryOrWorker() {
assertFalse(Authz.permits(ARCH_DESIGN, SPAWN, null), "not a primary — no lifecycle");
assertFalse(ARCH_DESIGN.isPrimary());
assertFalse(ARCH_DESIGN.isWorker(), "an architect is its own role, not a widened worker");
assertTrue(ARCH_DESIGN.isArchitect());
}
@Test
void observationIsOpenToBothAuthenticatedRoles() {
assertTrue(Authz.permits(PRIMARY, READ, null));
@@ -3,8 +3,12 @@ package dev.ltms.bridged.auth;
import dev.ltms.bridged.herdr.FakeHerdr;
import dev.ltms.bridged.herdr.PaneLocator;
import dev.ltms.bridged.mcp.ConnectionIdentity;
import dev.ltms.bridged.config.BridgedConfig;
import dev.ltms.bridged.peer.MemberRole;
import org.junit.jupiter.api.Test;
import java.util.Map;
import static org.junit.jupiter.api.Assertions.*;
/**
@@ -30,6 +34,17 @@ class CallerResolverTest {
return identity(999_999);
}
private static MemberRegistry boundMembers(String slot, MemberRole role) {
Map<String, BridgedConfig.Slot> architects = role == MemberRole.ARCHITECT
? Map.of("lead-designer", new BridgedConfig.Slot("sonnet")) : Map.of();
Map<String, BridgedConfig.Slot> devs = role == MemberRole.DEV
? Map.of("builder", new BridgedConfig.Slot("sonnet")) : Map.of();
MemberRegistry members = new MemberRegistry(
new BridgedConfig.Fleet(Map.of(), architects, devs, Map.of(), null));
assertTrue(members.bind(slot, "term_a"));
return members;
}
@Test
void aLoopbackWorkerPaneResolvesToWorkerRegardlessOfAuthMode() {
Principal underTrust = new CallerResolver(workerIdentity()).resolve("127.0.0.1", 42, null);
@@ -44,6 +59,43 @@ class CallerResolverTest {
assertEquals("term_a", underToken.terminal());
}
@Test
void aPinnedPrimaryTerminalResolvesToPrimaryNotWorker() {
// The primary's own session lives in a herdr pane (term_a here). Without the pin the pane
// match wins and the primary is locked out of spawn/send/stop as a misread worker.
Principal p = CallerResolver.pinnedTo(workerIdentity(), false, null, "term_a")
.resolve("127.0.0.1", 42, null);
assertEquals(Role.PRIMARY, p.role());
}
@Test
void aPinnedPrimaryTerminalNeedsNoTokenEvenInTokenMode() {
Principal p = CallerResolver.pinnedTo(workerIdentity(), true, "s3cret", "term_a")
.resolve("127.0.0.1", 42, null);
assertEquals(Role.PRIMARY, p.role(),
"the pane mapping is as unforgeable as a worker's — the pin outranks the token path");
}
@Test
void otherPanesRemainWorkersWhenAPinIsSet() {
Principal p = CallerResolver.pinnedTo(workerIdentity(), false, null, "term_someone_else")
.resolve("127.0.0.1", 42, null);
assertEquals(Role.WORKER, p.role());
assertEquals("term_a", p.terminal());
}
/** The pin is optional config, so an absent or whitespace one must change nothing at all. */
@Test
void aBlankPinLeavesWorkerResolutionUntouched() {
assertEquals(Role.WORKER,
CallerResolver.pinnedTo(workerIdentity(), false, null, " ").resolve("127.0.0.1", 42, null).role());
assertEquals(Role.WORKER,
CallerResolver.pinnedTo(workerIdentity(), false, null, null).resolve("127.0.0.1", 42, null).role());
}
@Test
void loopbackTrustTreatsANonWorkerLoopbackCallerAsThePrimary() {
Principal p = new CallerResolver(nonWorkerIdentity()).resolve("127.0.0.1", 99, null);
@@ -96,6 +148,264 @@ class CallerResolverTest {
assertEquals(Role.ANONYMOUS, p.role());
}
// ── CB-530: the leaders registry ────────────────────────────────────────────────────────────
// The fake resolves exactly one pane (term_a) from a PID, so "two leads both resolve" is
// asserted at the config layer (BridgedConfigTest#leaderTerminals…). What matters here is that
// resolution is a REGISTRY LOOKUP rather than a single equality test against one pin.
@Test
void aRegisteredLeadPaneResolvesToPrimaryCarryingItsName() {
Principal p = new CallerResolver(workerIdentity(), false, null, Map.of("term_a", "opus-5.0"))
.resolve("127.0.0.1", 42, null);
assertEquals(Role.PRIMARY, p.role());
assertEquals("opus-5.0", p.name(), "whoami must be able to say WHICH lead is asking");
assertEquals("term_a", p.terminal(),
"CB-532: a lead carries the pane it was matched by. Without it ownsSession() can "
+ "never be true for a lead, so it can send to a peer but never answer one");
}
@Test
void aSecondLeadIsRecognisedRatherThanSilentlyDemoted() {
// The regression this feature exists for: with a singular pin, whichever lead was not the
// pin resolved as a worker and was refused every orchestration call.
Principal p = new CallerResolver(workerIdentity(), false, null,
Map.of("term_elsewhere", "gpt-sol-5.6", "term_a", "opus-5.0"))
.resolve("127.0.0.1", 42, null);
assertEquals(Role.PRIMARY, p.role());
assertEquals("opus-5.0", p.name());
}
@Test
void aPaneAbsentFromTheRegistryIsStillAWorker() {
Principal p = new CallerResolver(workerIdentity(), false, null,
Map.of("term_elsewhere", "gpt-sol-5.6"))
.resolve("127.0.0.1", 42, null);
assertEquals(Role.WORKER, p.role());
assertEquals("term_a", p.terminal());
assertNull(p.name());
}
@Test
void aRegisteredLeadNeedsNoTokenEvenInTokenMode() {
Principal p = new CallerResolver(workerIdentity(), true, "s3cret", Map.of("term_a", "opus"))
.resolve("127.0.0.1", 42, null);
assertEquals(Role.PRIMARY, p.role(),
"the pane mapping is as unforgeable as a worker's — it outranks the token path");
assertEquals("opus", p.name());
}
/** The pre-CB-530 spelling must keep working, exactly, including for configs that never migrate. */
@Test
void theLegacySinglePinBehavesAsALeadNamedPrimary() {
Principal p = CallerResolver.pinnedTo(workerIdentity(), false, null, "term_a")
.resolve("127.0.0.1", 42, null);
assertEquals(Role.PRIMARY, p.role());
assertEquals("primary", p.name());
}
@Test
void anEmptyRegistryLeavesEveryPaneAWorker() {
Map<String, String> noLeads = null;
assertEquals(Role.WORKER,
new CallerResolver(workerIdentity(), false, null, Map.of())
.resolve("127.0.0.1", 42, null).role());
assertEquals(Role.WORKER,
new CallerResolver(workerIdentity(), false, null, noLeads)
.resolve("127.0.0.1", 42, null).role());
// CB-531: and the same for the live-registry form, whose supplier may also be absent.
assertEquals(Role.WORKER,
CallerResolver.withLeads(workerIdentity(), false, null, null)
.resolve("127.0.0.1", 42, null).role());
}
/** The audit line must distinguish leads once several exist, or a log says nothing useful. */
@Test
void describeNamesTheLeadButStillReadsPrimaryWhenUnnamed() {
assertEquals("leader:opus-5.0", Principal.leader("opus-5.0", "term_a", 1).describe());
assertEquals("primary", Principal.primary(1).describe());
assertEquals("worker:term_a", Principal.worker("term_a", 1).describe());
}
// ── CB-532: a lead is an addressable peer, not only a sender ────────────────────────────────
/**
* The regression this ticket exists for: two leads could both be recognised (CB-530/531) and
* still not converse, because REPLY is gated on ownsSession() and a lead owned nothing.
*/
@Test
void aLeadOwnsItsOwnPaneSoItMayAnswerAPeer() {
Principal lead = new CallerResolver(workerIdentity(), false, null, Map.of("term_a", "opus-5.0"))
.resolve("127.0.0.1", 42, null);
assertTrue(lead.ownsSession("term_a"));
assertTrue(Authz.permits(lead, Authz.Action.REPLY, "term_a"),
"a lead answering a peer replies for its OWN terminal — the rendezvous the sender "
+ "opened is keyed on exactly that");
assertTrue(Authz.permits(lead, Authz.Action.ASK, "term_a"));
}
@Test
void aLeadStillCannotActAsAnyoneElse() {
Principal lead = new CallerResolver(workerIdentity(), false, null, Map.of("term_a", "opus-5.0"))
.resolve("127.0.0.1", 42, null);
assertFalse(lead.ownsSession("term_someone_else"));
assertFalse(Authz.permits(lead, Authz.Action.REPLY, "term_someone_else"),
"widening WHO may reply must not widen WHAT they may reply as");
}
/** A primary with no pane — token mode, or off-host — owns nothing and must stay a sender only. */
@Test
void anUnnamedPrimaryWithNoPaneOwnsNothing() {
Principal p = new CallerResolver(nonWorkerIdentity(), true, "s3cret")
.resolve("127.0.0.1", 99, "Bearer s3cret");
assertEquals(Role.PRIMARY, p.role());
assertNull(p.terminal());
assertFalse(p.ownsSession(null), "a null terminal must never match a null session id");
assertFalse(Authz.permits(p, Authz.Action.REPLY, null));
}
@Test
void aLeadKeepsEveryOrchestrationRightItAlreadyHad() {
Principal lead = new CallerResolver(workerIdentity(), false, null, Map.of("term_a", "opus-5.0"))
.resolve("127.0.0.1", 42, null);
assertTrue(Authz.permits(lead, Authz.Action.SPAWN, null));
assertTrue(Authz.permits(lead, Authz.Action.SEND, "term_worker"));
assertTrue(Authz.permits(lead, Authz.Action.STOP, null));
assertTrue(Authz.permits(lead, Authz.Action.DRAIN, null));
}
/**
* CB-531: the registry is read per resolve, not snapshotted at construction — a lead that
* labels its tab after the daemon booted is recognised without a restart.
*/
@Test
void aLeadRegisteredAfterConstructionIsHonouredWithoutRebuildingTheResolver() {
Map<String, String> live = new java.util.HashMap<>();
CallerResolver r = CallerResolver.withLeads(workerIdentity(), false, null, () -> live);
assertEquals(Role.WORKER, r.resolve("127.0.0.1", 42, null).role());
live.put("term_a", "gpt-sol-5.6"); // the scanner sees a newly-labelled tab
Principal p = r.resolve("127.0.0.1", 42, null);
assertEquals(Role.PRIMARY, p.role());
assertEquals("gpt-sol-5.6", p.name());
}
/** The map form must stay a snapshot: a caller handing over a map is not offering live state. */
@Test
void theMapFormIsCopiedSoLaterMutationCannotGrantLeadership() {
Map<String, String> mutable = new java.util.HashMap<>();
CallerResolver r = new CallerResolver(workerIdentity(), false, null, mutable);
mutable.put("term_a", "sneaky");
assertEquals(Role.WORKER, r.resolve("127.0.0.1", 42, null).role());
}
// ── CB-548: architect slots ─────────────────────────────────────────────────────────────────
@Test
void aBoundArchitectPaneResolvesToArchitectBeforeTheWorkerFallback() {
// This is the production construction path used by Bridged.
Principal p = CallerResolver.withLeadsAndMembers(workerIdentity(), false, null,
Map::of, boundMembers("architect:lead-designer", MemberRole.ARCHITECT))
.resolve("127.0.0.1", 42, null);
assertEquals(Role.ARCHITECT, p.role(),
"a terminal bound to an architect slot is an architect, NOT the generic worker it "
+ "would otherwise resolve to");
assertEquals("lead-designer", p.name(), "whoami must say WHICH slot is asking");
assertEquals("term_a", p.terminal(), "the pane identity is carried so ownsSession works");
}
@Test
void anArchitectNeedsNoTokenEvenInTokenMode() {
Principal p = CallerResolver.withLeadsAndMembers(workerIdentity(), true, "s3cret",
Map::of, boundMembers("architect:lead-designer", MemberRole.ARCHITECT))
.resolve("127.0.0.1", 42, null);
assertEquals(Role.ARCHITECT, p.role(),
"the pane mapping is as unforgeable as a worker's — it outranks the token path");
}
@Test
void anUnboundPaneStillResolvesAsAWorker() {
MemberRegistry members = new MemberRegistry(new BridgedConfig.Fleet(Map.of(),
Map.of("lead-designer", new BridgedConfig.Slot("sonnet")), Map.of(), Map.of(), null));
Principal p = CallerResolver.withLeadsAndMembers(workerIdentity(), false, null,
Map::of, members).resolve("127.0.0.1", 42, null);
assertEquals(Role.WORKER, p.role());
assertNull(p.name());
}
/** CB-548 precedence: lead > architect > worker, so a pane named in BOTH is still a lead. */
@Test
void aLeadWinsOverAnArchitectBindingForTheSamePane() {
Principal p = CallerResolver.withLeadsAndMembers(workerIdentity(), false, null,
() -> Map.of("term_a", "opus-5.0"),
boundMembers("architect:lead-designer", MemberRole.ARCHITECT))
.resolve("127.0.0.1", 42, null);
assertEquals(Role.PRIMARY, p.role(),
"a pane the config calls a lead must keep resolving as a lead — no behaviour change "
+ "when an architect binding is added to an existing fleet");
assertEquals("opus-5.0", p.name());
}
/** The registry is live, like leads: a binding injected after construction is honoured. */
@Test
void anArchitectBoundAfterConstructionIsHonouredWithoutRebuildingTheResolver() {
MemberRegistry members = new MemberRegistry(new BridgedConfig.Fleet(Map.of(),
Map.of("lead-designer", new BridgedConfig.Slot("sonnet")), Map.of(), Map.of(), null));
CallerResolver r = CallerResolver.withLeadsAndMembers(workerIdentity(), false, null,
Map::of, members);
assertEquals(Role.WORKER, r.resolve("127.0.0.1", 42, null).role());
assertTrue(members.bind("architect:lead-designer", "term_a")); // the later lifecycle binds the slot
assertEquals(Role.ARCHITECT, r.resolve("127.0.0.1", 42, null).role());
assertEquals("architect:lead-designer", r.members().get("term_a"));
}
@Test
void describeNamesTheArchitectSlot() {
assertEquals("architect:lead-designer",
Principal.architect("lead-designer", "term_a", 1).describe());
}
/** An architect acts only as its own pane — the same ownsSession rule as a worker or lead. */
@Test
void anArchitectOwnsItsOwnPaneAndNoOther() {
Principal arch = CallerResolver.withLeadsAndMembers(workerIdentity(), false, null,
Map::of, boundMembers("architect:lead-designer", MemberRole.ARCHITECT))
.resolve("127.0.0.1", 42, null);
assertTrue(arch.ownsSession("term_a"));
assertTrue(Authz.permits(arch, Authz.Action.REPLY, "term_a"));
assertFalse(arch.ownsSession("term_b"));
assertFalse(Authz.permits(arch, Authz.Action.REPLY, "term_b"));
}
@Test
void aBoundNonArchitectSlotStillResolvesAsAWorker() {
Principal p = CallerResolver.withLeadsAndMembers(workerIdentity(), false, null,
Map::of, boundMembers("dev:builder", MemberRole.DEV))
.resolve("127.0.0.1", 42, null);
assertEquals(Role.WORKER, p.role(), "a dev binding must never grant architect rights");
}
@Test
void tokenModeRequiresANonEmptyConfiguredToken() {
ConnectionIdentity id = nonWorkerIdentity();
@@ -0,0 +1,269 @@
package dev.ltms.bridged.auth;
import dev.ltms.bridged.config.BridgedConfig;
import dev.ltms.bridged.peer.MemberRole;
import org.junit.jupiter.api.Test;
import java.util.ArrayList;
import java.util.LinkedHashMap;
import java.util.List;
import java.util.Map;
import java.util.Set;
import java.util.concurrent.CountDownLatch;
import java.util.concurrent.ExecutorService;
import java.util.concurrent.Executors;
import java.util.concurrent.Future;
import static org.junit.jupiter.api.Assertions.*;
/**
* CB-548 — the architect-slot registry: the config snapshot of slot → profile, and the live
* terminal → slot bindings it owns. The role a binding produces is asserted in
* {@link CallerResolverTest}; this pins the registry object itself — its invariants and their
* thread-safety.
*/
class MemberRegistryTest {
/**
* Slot keys are qualified by role (CB-557): a bare name is unique only within its pool, so
* {@code sonnet} can be both a developer and a reviewer, while a terminal binds to exactly one.
*/
private static final String DESIGNER = "architect:lead-designer";
private static final String REVIEWER = "architect:code-reviewer";
private static BridgedConfig.Fleet fleetWith(Map<String, BridgedConfig.Slot> architects) {
return new BridgedConfig.Fleet(Map.of(), architects, Map.of(), Map.of(), null);
}
private static Map<String, BridgedConfig.Slot> architects() {
Map<String, BridgedConfig.Slot> pool = new LinkedHashMap<>();
pool.put("lead-designer", new BridgedConfig.Slot("sonnet"));
pool.put("code-reviewer", new BridgedConfig.Slot("gx10"));
return pool;
}
private final MemberRegistry registry = new MemberRegistry(fleetWith(architects()));
@Test
void exposesTheConfiguredSlotsQualifiedByRole() {
assertEquals(Set.of(DESIGNER, REVIEWER), registry.slots().keySet());
assertTrue(registry.isSlot(REVIEWER));
assertFalse(registry.isSlot("nope"));
assertFalse(registry.isSlot("lead-designer"),
"the bare name is not the key — it is unique only inside its pool");
}
/** The case the pool shape exists for: one profile serving two roles is not a duplicate. */
@Test
void oneProfileMayServeTwoRolesUnderTheSameSlotName() {
Map<String, BridgedConfig.Slot> devs = Map.of("sonnet", new BridgedConfig.Slot("sonnet"));
Map<String, BridgedConfig.Slot> revs = Map.of("sonnet", new BridgedConfig.Slot("sonnet"));
MemberRegistry r = new MemberRegistry(
new BridgedConfig.Fleet(Map.of(), Map.of(), devs, revs, null));
assertEquals(Set.of("dev:sonnet", "reviewer:sonnet"), r.slots().keySet());
assertEquals(MemberRole.DEV, r.roleForSlot("dev:sonnet"));
assertEquals(MemberRole.REVIEWER, r.roleForSlot("reviewer:sonnet"));
}
@Test
void theSpawnLifecycleReadsTheProfileBackFromASlot() {
assertEquals("sonnet", registry.profileForSlot(DESIGNER));
assertEquals("gx10", registry.profileForSlot(REVIEWER));
assertNull(registry.profileForSlot("unknown"), "an unknown slot has no profile");
}
@Test
void startsEmptySoNoTerminalResolvesToAnArchitect() {
assertTrue(registry.snapshot().isEmpty());
assertNull(registry.slotForTerminal("term_design"),
"config declares no architect terminal — nothing is recognised until a bind");
assertNull(registry.slotForTerminal(null), "no terminal ⇒ no slot");
}
// ── bind ──────────────────────────────────────────────────────────────────────────────────
@Test
void bindResolvesTheTerminalToTheSlot() {
assertTrue(registry.bind(DESIGNER, "term_design"));
assertEquals(DESIGNER, registry.slotForTerminal("term_design"));
assertEquals(Map.of("term_design", DESIGNER), registry.snapshot());
}
@Test
void bindRefusesAnUnknownSlot() {
assertFalse(registry.bind("nope", "term_x"),
"a slot that is not configured must be refused — bind is not a way to invent one");
assertNull(registry.slotForTerminal("term_x"));
}
@Test
void bindRefusesATerminalInTwoSlots() {
assertTrue(registry.bind(DESIGNER, "term_design"));
assertFalse(registry.bind(REVIEWER, "term_design"),
"a terminal may occupy at most one slot");
assertEquals(DESIGNER, registry.slotForTerminal("term_design"),
"the first binding survives the refused second");
}
@Test
void bindRefusesASlotWithTwoTerminals() {
assertTrue(registry.bind(DESIGNER, "term_design"));
assertFalse(registry.bind(DESIGNER, "term_other"),
"a slot may host at most one terminal");
assertEquals(DESIGNER, registry.slotForTerminal("term_design"),
"the first binding survives the refused second");
assertNull(registry.slotForTerminal("term_other"));
}
@Test
void rebindingTheSamePairIsAnIdempotentNoOp() {
assertTrue(registry.bind(DESIGNER, "term_design"));
assertTrue(registry.bind(DESIGNER, "term_design"),
"the same terminal → slot is harmless to repeat");
assertEquals(1, registry.snapshot().size());
}
// ── unbind ────────────────────────────────────────────────────────────────────────────────
@Test
void unbindRemovesTheExactBinding() {
assertTrue(registry.bind(DESIGNER, "term_design"));
assertTrue(registry.unbind(DESIGNER, "term_design"));
assertNull(registry.slotForTerminal("term_design"));
assertTrue(registry.snapshot().isEmpty());
}
@Test
void aStaleUnbindDoesNotRemoveAReplacement() {
// Bind, tear down, and stand the slot back up with a NEW terminal.
assertTrue(registry.bind(DESIGNER, "term_design"));
registry.unbind(DESIGNER, "term_design");
assertTrue(registry.bind(DESIGNER, "term_new"));
// A late unbind naming the OLD terminal must not remove the replacement binding.
assertFalse(registry.unbind(DESIGNER, "term_design"));
assertEquals(DESIGNER, registry.slotForTerminal("term_new"),
"the replacement terminal stays bound");
}
@Test
void aStaleUnbindForATerminalThatMovedSlotsDoesNothing() {
// term_design starts in lead-designer, is torn down, and stands back up in a FREE slot.
assertTrue(registry.bind(DESIGNER, "term_design"));
registry.unbind(DESIGNER, "term_design");
assertTrue(registry.bind(REVIEWER, "term_design"));
// Unbinding against the slot it no longer occupies is refused; the new binding is intact.
assertFalse(registry.unbind(DESIGNER, "term_design"),
"the old slot must not unbind a terminal that moved elsewhere");
assertEquals(REVIEWER, registry.slotForTerminal("term_design"));
}
@Test
void unbindOfNothingIsAFalseNoOp() {
assertFalse(registry.unbind(DESIGNER, "term_design"),
"nothing was bound, so nothing is removed");
}
// ── snapshot ─────────────────────────────────────────────────────────────────────────────
@Test
void theSnapshotIsAnImmutableCopyNotAliveState() {
assertTrue(registry.bind(DESIGNER, "term_design"));
Map<String, String> snap = registry.snapshot();
assertThrows(UnsupportedOperationException.class, () -> snap.put("x", "y"),
"a handed-out snapshot cannot be mutated in place");
// Later binds must not leak into an earlier snapshot.
assertTrue(registry.bind(REVIEWER, "term_review"));
assertFalse(snap.containsKey("term_review"),
"a snapshot is a point-in-time copy, not a live view");
}
// ── concurrency (CB-548 invariants hold under contention) ─────────────────────────────────
@Test
void concurrentBindsNeverGiveASlotTwoTerminals() throws Exception {
int n = 16;
ExecutorService pool = Executors.newFixedThreadPool(n);
try {
CountDownLatch go = new CountDownLatch(1);
List<Future<Boolean>> results = new ArrayList<>();
for (int i = 0; i < n; i++) {
final String term = "term_" + i; // every thread races for the SAME slot
results.add(pool.submit(() -> {
go.await();
return registry.bind(DESIGNER, term);
}));
}
go.countDown();
int won = 0;
for (Future<Boolean> r : results) {
if (r.get()) {
won++;
}
}
assertEquals(1, won, "exactly one terminal may win the sole slot, got " + won);
assertEquals(1, registry.snapshot().size(),
"the slot hosts at most one terminal after the race");
} finally {
pool.shutdownNow();
}
}
@Test
void concurrentBindsNeverPutOneTerminalInTwoSlots() throws Exception {
int n = 16;
ExecutorService pool = Executors.newFixedThreadPool(n);
try {
CountDownLatch go = new CountDownLatch(1);
List<Future<String>> results = new ArrayList<>();
for (int i = 0; i < n; i++) {
final String slot = (i % 2 == 0) ? DESIGNER : REVIEWER; // all race for ONE terminal
results.add(pool.submit(() -> {
go.await();
return registry.bind(slot, "shared_term")
? registry.slotForTerminal("shared_term") : null;
}));
}
go.countDown();
// Rebinding the same terminal to the same slot is a harmless idempotent true, so count
// winners is not the assertion — agreement is: every thread that reported success must
// have seen the terminal in the SAME slot, never in two at once.
String bound = null;
boolean conflict = false;
for (Future<String> r : results) {
String s = r.get();
if (s != null) {
if (bound == null) {
bound = s;
} else if (!bound.equals(s)) {
conflict = true;
}
}
}
assertFalse(conflict, "a terminal was observed in two slots at once");
assertNotNull(bound, "at least one thread bound the terminal");
assertEquals(1, registry.snapshot().size(),
"the terminal occupies exactly one slot in the final snapshot");
assertEquals(bound, registry.slotForTerminal("shared_term"));
} finally {
pool.shutdownNow();
}
}
/** A handed-over slot map is snapshotted at construction, not offered as live state. */
@Test
void theSlotSnapshotIsFixedByConstruction() {
Map<String, BridgedConfig.Slot> mutable = architects();
MemberRegistry r = new MemberRegistry(fleetWith(mutable));
mutable.put("hijack", new BridgedConfig.Slot("gx10"));
assertFalse(r.isSlot("architect:hijack"), "a handed-over map is not offered as live state");
}
}
File diff suppressed because it is too large Load Diff
@@ -0,0 +1,295 @@
package dev.ltms.bridged.config;
import org.junit.jupiter.api.Test;
import org.junit.jupiter.api.io.TempDir;
import java.nio.file.Files;
import java.nio.file.Path;
import static org.junit.jupiter.api.Assertions.*;
/**
* CB-559: re-reading {@code bridged.yaml} under a running daemon.
*
* <p>The tests that matter here are the refusals. A reload that applies a good file is the easy
* half; the half that protects an operator is the one that keeps the running config when the new
* file is bad, and the one that refuses a change the running daemon cannot honour.
*/
class ConfigRefTest {
/** A minimal file that loads and passes every startup validator. */
private static String yaml(String extra) {
return """
bind:
host: 127.0.0.1
port: 8765
herdrSocket: ~/.config/herdr/herdr.sock
profiles:
sonnet:
baseUrl: http://gx00.gw:8000
model: sonnet
guard:
offSubscriptionHosts:
- gx00.gw
""" + extra;
}
private static ConfigRef refFor(Path f) {
return new ConfigRef(f, BridgedConfig.load(f));
}
@Test
void aHotChangeIsAppliedAndReadThroughGet(@TempDir Path dir) throws Exception {
Path f = dir.resolve("bridged.yaml");
Files.writeString(f, yaml("""
fleet:
tabLabel: "{role}: {profile} #{n}"
developers:
a:
profile: sonnet
"""));
ConfigRef ref = refFor(f);
assertEquals("{role}: {profile} #{n}", ref.get().fleet().tabLabel());
Files.writeString(f, yaml("""
fleet:
tabLabel: "[{profile}] {role}"
developers:
a:
profile: sonnet
"""));
ConfigRef.Outcome out = ref.reload();
assertTrue(out.applied());
assertTrue(out.deferred().isEmpty());
assertEquals("config reloaded", out.summary());
assertEquals("[{profile}] {role}", ref.get().fleet().tabLabel());
}
/**
* The point of the whole class: a consumer holding the ref sees the new value without being
* rebuilt. A component that captured {@code get()} into a field would still show the old one.
*/
@Test
void aConsumerHoldingTheRefSeesTheNewValue(@TempDir Path dir) throws Exception {
Path f = dir.resolve("bridged.yaml");
Files.writeString(f, yaml("placement: weighted\n"));
ConfigRef ref = refFor(f);
java.util.function.Supplier<String> reader = () -> ref.get().placement();
assertEquals("weighted", reader.get());
Files.writeString(f, yaml("placement: fixed\n"));
assertTrue(ref.reload().applied());
assertEquals("fixed", reader.get());
}
@Test
void aChangedColdKeyRefusesTheWholeReload(@TempDir Path dir) throws Exception {
Path f = dir.resolve("bridged.yaml");
Files.writeString(f, yaml("placement: weighted\n"));
ConfigRef ref = refFor(f);
// Two changes in one file: a cold one (the port) and a hot one (placement).
Files.writeString(f, yaml("placement: fixed\n").replace("port: 8765", "port: 9999"));
ConfigRef.Outcome out = ref.reload();
assertFalse(out.applied());
assertEquals(java.util.List.of("bind"), out.coldKeys());
assertTrue(out.summary().contains("Restart bridged"), out.summary());
// The hot half must NOT have leaked in. A half-applied reload leaves the daemon matching no
// file on disk, which is worse for an operator than no reload at all.
assertEquals("weighted", ref.get().placement());
assertEquals(8765, ref.get().bind().port());
}
@Test
void aFileThatNoLongerParsesKeepsTheRunningConfig(@TempDir Path dir) throws Exception {
Path f = dir.resolve("bridged.yaml");
Files.writeString(f, yaml("placement: weighted\n"));
ConfigRef ref = refFor(f);
BridgedConfig before = ref.get();
Files.writeString(f, "profiles:\n sonnet:\n baseUrl: \"unclosed\n");
ConfigRef.Outcome out = ref.reload();
assertFalse(out.applied());
assertNotNull(out.error());
assertTrue(out.summary().startsWith("config reload refused"), out.summary());
assertSame(before, ref.get());
}
/** A file that would have refused to boot must not be able to slip in through a reload. */
@Test
void aFileThatFailsAValidatorKeepsTheRunningConfig(@TempDir Path dir) throws Exception {
Path f = dir.resolve("bridged.yaml");
Files.writeString(f, yaml("placement: weighted\n"));
ConfigRef ref = refFor(f);
BridgedConfig before = ref.get();
// A member slot naming a profile that does not exist — validateMembers refuses this at
// startup, so it must refuse it here too.
Files.writeString(f, yaml("""
fleet:
developers:
a:
profile: no-such-profile
"""));
ConfigRef.Outcome out = ref.reload();
assertFalse(out.applied());
assertNotNull(out.error());
assertSame(before, ref.get());
}
@Test
void aDeletedFileIsRefusedRatherThanCrashing(@TempDir Path dir) throws Exception {
Path f = dir.resolve("bridged.yaml");
Files.writeString(f, yaml(""));
ConfigRef ref = refFor(f);
BridgedConfig before = ref.get();
Files.delete(f);
ConfigRef.Outcome out = ref.reload();
assertFalse(out.applied());
assertNotNull(out.error());
assertSame(before, ref.get());
}
/** A deferred change applies to the snapshot but the operator is told it needs a restart. */
@Test
void aDeferredChangeIsAppliedAndReported(@TempDir Path dir) throws Exception {
Path f = dir.resolve("bridged.yaml");
Files.writeString(f, yaml("""
lifecycle:
drainTimeoutSeconds: 30
"""));
ConfigRef ref = refFor(f);
Files.writeString(f, yaml("""
lifecycle:
drainTimeoutSeconds: 60
"""));
ConfigRef.Outcome out = ref.reload();
assertTrue(out.applied());
assertEquals(java.util.List.of("lifecycle"), out.deferred());
assertTrue(out.summary().contains("needs a restart") || out.summary().contains("need a restart"),
out.summary());
assertEquals(60, ref.get().lifecycle().drainTimeoutSeconds());
}
/**
* Adding a profile is deferred, not hot: a new backend needs its own launcher, and launchers are
* built once at startup. The snapshot carries it so a restart picks it up.
*/
@Test
void addingAProfileIsReportedAsDeferred(@TempDir Path dir) throws Exception {
Path f = dir.resolve("bridged.yaml");
Files.writeString(f, yaml(""));
ConfigRef ref = refFor(f);
Files.writeString(f, """
bind:
host: 127.0.0.1
port: 8765
herdrSocket: ~/.config/herdr/herdr.sock
profiles:
sonnet:
baseUrl: http://gx00.gw:8000
model: sonnet
haiku:
baseUrl: http://gx00.gw:8000
model: haiku
guard:
offSubscriptionHosts:
- gx00.gw
""");
ConfigRef.Outcome out = ref.reload();
assertTrue(out.applied());
assertEquals(1, out.deferred().size());
assertTrue(out.deferred().getFirst().contains("haiku"), out.deferred().toString());
}
/**
* Changing an existing profile's weight or maxLoad IS hot: placement reads those live through
* the supplier on the composite, so the next spawn already sees them.
*/
@Test
void changingAProfilesWeightOrMaxLoadIsHot(@TempDir Path dir) throws Exception {
Path f = dir.resolve("bridged.yaml");
Files.writeString(f, yaml(""));
ConfigRef ref = refFor(f);
Files.writeString(f, """
bind:
host: 127.0.0.1
port: 8765
herdrSocket: ~/.config/herdr/herdr.sock
profiles:
sonnet:
baseUrl: http://gx00.gw:8000
model: sonnet
maxLoad: 7
weight: 3.0
guard:
offSubscriptionHosts:
- gx00.gw
""");
ConfigRef.Outcome out = ref.reload();
assertTrue(out.applied());
assertTrue(out.deferred().isEmpty(), out.deferred().toString());
assertEquals(7, ref.get().profiles().get("sonnet").maxLoad());
}
/**
* Changing an existing profile's MODEL is deferred, not hot — and saying so is the whole point.
* HerdrPeerLauncher takes Map.copyOf(profiles) at construction and resolves every spawn out of
* that copy, so a reloaded model never reaches a launch. Reporting it as applied would be the
* worst outcome a reload can produce: the operator has no reason to doubt a clean "reloaded".
*/
@Test
void changingAProfilesLaunchSettingsIsReportedAsDeferred(@TempDir Path dir) throws Exception {
Path f = dir.resolve("bridged.yaml");
Files.writeString(f, yaml(""));
ConfigRef ref = refFor(f);
Files.writeString(f, """
bind:
host: 127.0.0.1
port: 8765
herdrSocket: ~/.config/herdr/herdr.sock
profiles:
sonnet:
baseUrl: http://gx00.gw:8000
model: deepseek-v4-flash
guard:
offSubscriptionHosts:
- gx00.gw
""");
ConfigRef.Outcome out = ref.reload();
assertTrue(out.applied());
assertEquals(1, out.deferred().size(), out.deferred().toString());
assertTrue(out.deferred().getFirst().contains("sonnet"), out.deferred().toString());
assertTrue(out.deferred().getFirst().contains("launch settings"), out.deferred().toString());
// The snapshot still carries the new value — a restart is what makes it take effect.
assertEquals("deepseek-v4-flash", ref.get().profiles().get("sonnet").model());
}
@Test
void aFixedRefHasNoFileAndRefusesToReload() {
BridgedConfig cfg = new BridgedConfig(null, null, null, null, null, null,
null, null, null, null, null, null, null, null, null).withDefaults();
ConfigRef ref = ConfigRef.fixed(cfg);
assertNull(ref.path());
assertSame(cfg, ref.get());
ConfigRef.Outcome out = ref.reload();
assertFalse(out.applied());
assertNotNull(out.error());
}
}
@@ -0,0 +1,162 @@
package dev.ltms.bridged.config;
import org.junit.jupiter.api.Test;
import org.junit.jupiter.api.io.TempDir;
import java.nio.file.Files;
import java.nio.file.Path;
import java.nio.file.attribute.FileTime;
import static org.junit.jupiter.api.Assertions.*;
/**
* CB-559: the mtime poller behind {@code configReload:}. The tests drive {@link
* ConfigWatcher#tick()} directly rather than the scheduler, so nothing here sleeps.
*/
class ConfigWatcherTest {
private static String yaml(String placement) {
return """
bind:
host: 127.0.0.1
port: 8765
herdrSocket: ~/.config/herdr/herdr.sock
profiles:
sonnet:
baseUrl: http://gx00.gw:8000
model: sonnet
guard:
offSubscriptionHosts:
- gx00.gw
""" + "placement: " + placement + "\n";
}
/** Write and stamp an mtime, so a test never depends on the filesystem's clock resolution. */
private static void write(Path f, String content, long millis) throws Exception {
Files.writeString(f, content);
Files.setLastModifiedTime(f, FileTime.fromMillis(millis));
}
@Test
void aChangedMtimeTriggersAReload(@TempDir Path dir) throws Exception {
Path f = dir.resolve("bridged.yaml");
write(f, yaml("weighted"), 1_000L);
ConfigRef ref = new ConfigRef(f, BridgedConfig.load(f));
ConfigWatcher watcher = new ConfigWatcher(ref, 10);
try {
write(f, yaml("fixed"), 2_000L);
watcher.tick();
assertEquals("fixed", ref.get().placement());
} finally {
watcher.stop();
}
}
@Test
void anUnchangedFileIsNotReloaded(@TempDir Path dir) throws Exception {
Path f = dir.resolve("bridged.yaml");
write(f, yaml("weighted"), 1_000L);
ConfigRef ref = new ConfigRef(f, BridgedConfig.load(f));
BridgedConfig before = ref.get();
ConfigWatcher watcher = new ConfigWatcher(ref, 10);
try {
watcher.tick();
watcher.tick();
// Same instance, so no reload happened — a reload always swaps in a fresh object.
assertSame(before, ref.get());
} finally {
watcher.stop();
}
}
/**
* A refused reload must not be retried every tick. Without the stamp-before-reload order the
* same refusal would be logged forever, which buries the log an operator needs.
*/
@Test
void aRefusedReloadIsNotRetriedUntilTheFileChangesAgain(@TempDir Path dir) throws Exception {
Path f = dir.resolve("bridged.yaml");
write(f, yaml("weighted"), 1_000L);
ConfigRef ref = new ConfigRef(f, BridgedConfig.load(f));
BridgedConfig before = ref.get();
ConfigWatcher watcher = new ConfigWatcher(ref, 10);
try {
// A cold change: refused, and the running config stays.
write(f, yaml("fixed").replace("port: 8765", "port: 9999"), 2_000L);
watcher.tick();
assertSame(before, ref.get());
// The next tick sees the same mtime, so it does nothing at all.
watcher.tick();
assertSame(before, ref.get());
// A fresh save earns a fresh attempt — and this one is hot, so it applies.
write(f, yaml("fixed"), 3_000L);
watcher.tick();
assertEquals("fixed", ref.get().placement());
} finally {
watcher.stop();
}
}
/**
* Editors briefly unlink the file mid-save. A missing file is skipped, not an error and not a
* reload — the running config stays live, which is right either way.
*/
@Test
void aMissingFileIsSkippedAndTheNextTickTriesAgain(@TempDir Path dir) throws Exception {
Path f = dir.resolve("bridged.yaml");
write(f, yaml("weighted"), 1_000L);
ConfigRef ref = new ConfigRef(f, BridgedConfig.load(f));
BridgedConfig before = ref.get();
ConfigWatcher watcher = new ConfigWatcher(ref, 10);
try {
Files.delete(f);
assertDoesNotThrow(watcher::tick);
assertSame(before, ref.get());
write(f, yaml("fixed"), 2_000L);
watcher.tick();
assertEquals("fixed", ref.get().placement());
} finally {
watcher.stop();
}
}
/** A ref with no file behind it must not start a scheduler that could never do anything. */
@Test
void aFixedRefStartsNoWatch() {
BridgedConfig cfg = new BridgedConfig(null, null, null, null, null, null,
null, null, null, null, null, null, null, null, null).withDefaults();
ConfigWatcher watcher = new ConfigWatcher(ConfigRef.fixed(cfg), 10);
try {
assertDoesNotThrow(watcher::start);
assertDoesNotThrow(watcher::tick);
} finally {
watcher.stop();
}
}
@Test
void configReloadIsOffUnlessTheBlockSaysOtherwise(@TempDir Path dir) throws Exception {
Path f = dir.resolve("bridged.yaml");
write(f, yaml("weighted"), 1_000L);
assertNull(BridgedConfig.load(f).configReload());
write(f, yaml("weighted") + "configReload:\n enabled: true\n", 2_000L);
BridgedConfig.ConfigReload on = BridgedConfig.load(f).configReload();
assertTrue(on.isEnabled());
assertEquals(10, on.intervalSeconds(), "an absent interval defaults to 10s");
write(f, yaml("weighted") + "configReload:\n intervalSeconds: 30\n", 3_000L);
BridgedConfig.ConfigReload off = BridgedConfig.load(f).configReload();
assertFalse(off.isEnabled(), "a block that only sets the interval does not enable the watch");
assertEquals(30, off.intervalSeconds());
}
}
@@ -11,10 +11,11 @@ import static org.junit.jupiter.api.Assertions.*;
import static org.junit.jupiter.api.Assumptions.assumeTrue;
/**
* Contract test for the {@code agent.*} south side against a REAL herdr, locking in
* the CB-102 spike findings. It spawns a HARMLESS probe command (never {@code claude},
* so no subscription/token involvement), proves the {@code env} map reaches the process
* environment, exercises status/read, and always tears the pane down.
* Contract test for the worker env seam against a REAL herdr. Under protocol 19 (CB-521) the
* env map is injected at PANE CREATION ({@code tab.create}), not {@code agent.start} — and
* {@code agent.start} now only launches supported agent kinds, so this probes the seed pane's
* SHELL directly (never {@code claude}, so no subscription/token involvement) and always tears
* the throwaway space down.
*
* <p>Tagged {@code contract}; run with {@code mvn test -Pcontract}.
*/
@@ -26,38 +27,30 @@ class AgentControlContractTest {
}
@Test
void startInjectsEnvThenReadAndClose() throws Exception {
void tabCreateInjectsEnvIntoTheSeedShell() throws Exception {
assumeTrue(!noSocket(), "no herdr socket — skipping");
try (UnixSocketHerdrClient herdr = UnixSocketHerdrClient.connect()) {
AgentControl agents = new AgentControl(herdr);
Agent probe = agents.start(
"__contract__",
List.of("bash", "-c", "printf 'PROBE_BASE=[%s]\\n' \"$ANTHROPIC_BASE_URL\"; sleep 20"),
WorkspaceControl spaces = new WorkspaceControl(herdr);
Workspace space = spaces.ensureWorkspace("__bridged_env_contract__");
Tab.Created tab = spaces.createTab(space.workspaceId(), null,
Map.of("ANTHROPIC_BASE_URL", "http://gx00.gw:8000"));
assertNotNull(probe.terminalId());
assertNotNull(probe.paneId());
try {
// Give the shell a moment to print, then confirm env reached the process.
assertNotNull(tab.rootPaneId(), "tab.create must return the seed pane");
Thread.sleep(1000); // let the seed shell reach its prompt
herdr.call("pane.send_input", Map.of(
"pane_id", tab.rootPaneId(),
"text", "printf 'PROBE_BASE=[%s]\\n' \"$ANTHROPIC_BASE_URL\"",
"keys", List.of("enter")));
Thread.sleep(800);
String visible = agents.read(probe.terminalId(), "visible");
String visible = herdr.call("pane.read",
Map.of("pane_id", tab.rootPaneId(), "source", "visible"))
.path("read").path("text").asText("");
assertTrue(visible.contains("PROBE_BASE=[http://gx00.gw:8000]"),
"env map must reach the process; saw: " + visible);
// Status is queryable; the probe appears in the agent list.
assertNotNull(agents.status(probe.terminalId()));
assertTrue(agents.list().stream()
.anyMatch(a -> probe.terminalId().equals(a.terminalId())),
"spawned probe should appear in agent.list");
"env map must reach the seed shell; saw: " + visible);
} finally {
agents.close(probe.paneId());
spaces.closeTab(tab.tab().tabId());
herdr.call("workspace.close", Map.of("workspace_id", space.workspaceId()));
}
// After close the pane is gone.
assertFalse(agents.list().stream()
.anyMatch(a -> probe.terminalId().equals(a.terminalId())),
"closed probe should no longer be listed");
}
}
}
@@ -7,34 +7,68 @@ import java.util.Map;
import static org.junit.jupiter.api.Assertions.assertEquals;
/** Unit-level behaviour of {@link AgentControl} over a fake herdr. */
/** Unit-level behaviour of {@link AgentControl} over a fake herdr (protocol 19). */
class AgentControlTest {
/** The {@code text} of every agent.send, in call order. */
/** The {@code text} of every agent.prompt, in call order. */
@SuppressWarnings("unchecked")
private static List<String> sendTexts(FakeHerdr herdr) {
private static List<String> promptTexts(FakeHerdr herdr) {
return herdr.calls.stream()
.filter(c -> c.method().equals("agent.send"))
.filter(c -> c.method().equals("agent.prompt"))
.map(c -> ((Map<String, Object>) c.params()).get("text").toString())
.toList();
}
@Test
void sendDeliversThePayloadThenAStandaloneSubmitKey() {
void sendDeliversThePayloadAsOnePromptThatSubmitsItself() {
FakeHerdr herdr = new FakeHerdr();
new AgentControl(herdr).send("term_x", "do the thing");
// The Enter must be its own event — appended to the paste it would be swallowed as text.
assertEquals(List.of("do the thing", "\r"), sendTexts(herdr),
"payload paste first, then a separate carriage-return keystroke to submit it");
// agent.prompt pastes AND submits in one call — no separate Enter event to assert.
assertEquals(List.of("do the thing"), promptTexts(herdr),
"exactly one agent.prompt carrying the payload");
}
@Test
void sendPreservesEmbeddedNewlinesAndSubmitsOnlyOnce() {
void sendPreservesEmbeddedNewlinesVerbatim() {
FakeHerdr herdr = new FakeHerdr();
new AgentControl(herdr).send("term_x", "line1\nline2");
assertEquals(List.of("line1\nline2", "\r"), sendTexts(herdr),
"multiline content is delivered verbatim; a single trailing Enter submits it");
assertEquals(List.of("line1\nline2"), promptTexts(herdr),
"multiline content is delivered verbatim in the single prompt");
}
@Test
void submitNudgesWithAStandaloneEnterKeystroke() {
FakeHerdr herdr = new FakeHerdr();
new AgentControl(herdr).submit("term_x");
FakeHerdr.Call keys = herdr.lastCall("agent.send_keys");
assertEquals(Map.of("target", "term_x", "keys", List.of("enter")), keys.params(),
"the raced-Enter nudge is a raw send_keys, not a second prompt");
}
@Test
@SuppressWarnings("unchecked")
void aTerminalIdTargetIsTranslatedToItsPaneId() {
// Protocol 19 rejects terminal_id as an agent.* target; the fake's agent.list maps
// term_a to pane w2:p7, and the control layer must address herdr by that pane.
FakeHerdr herdr = new FakeHerdr();
new AgentControl(herdr).send("term_a", "hello");
Map<String, Object> prompt = (Map<String, Object>) herdr.lastCall("agent.prompt").params();
assertEquals("w2:p7", prompt.get("target"), "terminal target resolved to the agent's pane id");
}
@Test
@SuppressWarnings("unchecked")
void theTerminalToPaneMappingIsCachedAcrossCalls() {
FakeHerdr herdr = new FakeHerdr();
AgentControl agents = new AgentControl(herdr);
agents.send("term_a", "one");
agents.send("term_a", "two");
long lists = herdr.calls.stream().filter(c -> c.method().equals("agent.list")).count();
assertEquals(1, lists, "one agent.list resolution serves every later call to the same terminal");
}
}
@@ -4,12 +4,14 @@ import com.fasterxml.jackson.databind.JsonNode;
import com.fasterxml.jackson.databind.ObjectMapper;
import java.util.ArrayList;
import java.util.LinkedHashMap;
import java.util.List;
import java.util.Map;
/**
* Recording fake {@link HerdrClient} for unit/acceptance tests. Returns canned frames
* captured from the real herdr 0.7.0 daemon and records every call so tests can assert
* both behaviour and that guard-blocked paths never reached herdr.
* matching the real herdr 0.8.0 daemon (protocol 19) and records every call so tests can
* assert both behaviour and that guard-blocked paths never reached herdr.
*/
public final class FakeHerdr implements HerdrClient {
@@ -24,12 +26,18 @@ public final class FakeHerdr implements HerdrClient {
private boolean healthy = true;
private final List<String> extraWorkspaces = new ArrayList<>();
private final List<String> extraAgents = new ArrayList<>();
/** workspaceId → extra tabs that {@code tab.list} reports for it (CB-558 lead scans). */
private final Map<String, List<String>> extraTabs = new LinkedHashMap<>();
private int agentNameTakenFor = 0;
private int agentPaneBusyFor = 0;
private int workerTabPaneCount = 1;
private String paneCloseErrorCode = null;
private String agentSendErrorCode = null;
private volatile String agentStatus = "idle"; // steady-state agent.get status
private volatile String readText = "worker transcript tail"; // canned agent.read output
private int pinnedStarts = 0; // how many upcoming agent.start calls report a fixed pane
private String pinnedStartTerminal;
private String pinnedStartPane;
public FakeHerdr healthy(boolean h) {
this.healthy = h;
@@ -42,6 +50,12 @@ public final class FakeHerdr implements HerdrClient {
return this;
}
/** Reject the first {@code n} {@code agent.start} calls with {@code agent_pane_busy}. */
public FakeHerdr agentPaneBusyTimes(int n) {
this.agentPaneBusyFor = n;
return this;
}
/** Make the worker tab (w9:t2) report this many panes in {@code tab.list} (default 1). */
public FakeHerdr withWorkerTabPaneCount(int n) {
this.workerTabPaneCount = n;
@@ -66,12 +80,25 @@ public final class FakeHerdr implements HerdrClient {
return this;
}
/** Make {@code agent.send} fail with this herdr error code. */
/** Make delivery ({@code agent.prompt} / {@code agent.send_keys}) fail with this error code. */
public FakeHerdr agentSendFailsWith(String code) {
this.agentSendErrorCode = code;
return this;
}
/**
* Force the next {@code n} {@code agent.start} calls to report this terminal/pane coordinate,
* instead of the fake's usual incrementing {@code term_new_n}/{@code w9:pRoot_n}. Lets a test make
* two spawns report the <em>same</em> herdr pane, to prove the host-unique id (CB-519) never
* collides on that coordinate.
*/
public FakeHerdr pinNextStarts(int n, String terminalId, String paneId) {
this.pinnedStarts = n;
this.pinnedStartTerminal = terminalId;
this.pinnedStartPane = paneId;
return this;
}
/**
* Seed a named agent into {@code agent.list} (e.g. an orphaned worker for CB-117 reaper tests).
@@ -86,6 +113,18 @@ public final class FakeHerdr implements HerdrClient {
return this;
}
/**
* Seed a labelled tab into {@code tab.list} for one workspace (e.g. an existing {@code lead: x}
* tab). Pair it with {@link #withAgent} on the same {@code tabId} to make the lead <em>live</em>;
* seeding the tab alone models the stale-label case.
*/
public FakeHerdr withTab(String workspaceId, String tabId, String label) {
extraTabs.computeIfAbsent(workspaceId, _ -> new ArrayList<>())
.add(("{\"tab_id\":\"%s\",\"workspace_id\":\"%s\",\"label\":\"%s\",\"pane_count\":1}")
.formatted(tabId, workspaceId, label));
return this;
}
/** Seed an additional workspace into {@code workspace.list} (e.g. a pre-existing worker space). */
public FakeHerdr withWorkspace(String id, String label) {
extraWorkspaces.add(("{\"workspace_id\":\"%s\",\"label\":\"%s\",\"focused\":false,"
@@ -110,7 +149,7 @@ public final class FakeHerdr implements HerdrClient {
try {
return switch (method) {
case "ping" -> mapper.readTree(
"{\"type\":\"pong\",\"version\":\"0.7.0\",\"protocol\":14}");
"{\"type\":\"pong\",\"version\":\"0.8.0\",\"protocol\":19}");
case "workspace.list" -> mapper.readTree(("""
{"type":"workspace_list","workspaces":[
{"workspace_id":"w1","label":"dev-mgnl","focused":true,"pane_count":7,"agent_status":"unknown"},
@@ -122,9 +161,19 @@ public final class FakeHerdr implements HerdrClient {
"agent_session":{"kind":"id","value":"sess-1111"},
"workspace_id":"w2","tab_id":"w2:t7","pane_id":"w2:p7"}%s]}""")
.formatted(extraAgents.isEmpty() ? "" : "," + String.join(",", extraAgents)));
case "agent.send" -> {
case "agent.prompt" -> {
if (agentSendErrorCode != null) {
throw new HerdrException("herdr error [" + agentSendErrorCode + "]: agent.send failed",
throw new HerdrException("herdr error [" + agentSendErrorCode + "]: agent.prompt failed",
agentSendErrorCode, null);
}
yield mapper.readTree(("""
{"type":"agent_prompted","agent":{"terminal_id":"term_a","agent":"claude",
"agent_status":"%s","workspace_id":"w2","tab_id":"w2:t7","pane_id":"w2:p7"}}""")
.formatted(agentStatus));
}
case "agent.send_keys" -> {
if (agentSendErrorCode != null) {
throw new HerdrException("herdr error [" + agentSendErrorCode + "]: agent.send_keys failed",
agentSendErrorCode, null);
}
yield mapper.readTree("{\"type\":\"ok\"}");
@@ -136,37 +185,78 @@ public final class FakeHerdr implements HerdrClient {
case "agent.read" -> mapper.readTree(mapper.writeValueAsString(
java.util.Map.of("type", "agent_read", "read", java.util.Map.of("text", readText))));
case "agent.start" -> {
// Protocol 19: kind and pane_id are required — reject like the real daemon.
java.util.Map<?, ?> p = params instanceof java.util.Map<?, ?> m ? m : java.util.Map.of();
for (String required : new String[]{"kind", "pane_id"}) {
if (p.get(required) == null) {
throw new HerdrException(
"herdr error [invalid_request]: invalid request: missing field `"
+ required + "`", "invalid_request", null);
}
}
long starts = calls.stream().filter(c -> c.method().equals("agent.start")).count();
if (starts <= agentNameTakenFor) {
if (starts <= agentPaneBusyFor) {
throw new HerdrException(
"herdr error [agent_pane_busy]: agent target pane is not an available shell",
"agent_pane_busy", null);
}
long busyAdjusted = starts - agentPaneBusyFor;
if (busyAdjusted <= agentNameTakenFor) {
throw new HerdrException(
"herdr error [agent_name_taken]: agent name already used",
"agent_name_taken", null);
}
long n = starts - agentNameTakenFor;
long n = busyAdjusted - agentNameTakenFor;
boolean pinned = pinnedStarts > 0;
if (pinned) {
pinnedStarts--;
}
// Protocol 19: the agent starts INTO the requested pane, so its pane_id normally
// echoes the param. A pin overrides both coordinates, which is the only way to
// make two spawns report one pane — what CB-519's collision test needs.
String terminal = pinned ? pinnedStartTerminal : ("term_new_" + n);
Object pane = pinned ? pinnedStartPane : p.get("pane_id");
yield mapper.readTree(("""
{"type":"agent_started","agent":{
"terminal_id":"term_new_%d","name":"claude","agent_status":"unknown",
"workspace_id":"w9","tab_id":"w9:t2","pane_id":"w9:pW_%d"}}""")
.formatted(n, n));
"terminal_id":"%s","name":"claude","agent_status":"unknown",
"workspace_id":"w9","tab_id":"w9:t2","pane_id":"%s"}}""")
.formatted(terminal, pane));
}
case "pane.split" -> mapper.readTree("""
{"type":"pane_info","pane":{"pane_id":"w1:pSplit","workspace_id":"w1",
"tab_id":"w1:t1"}}""");
case "workspace.create" -> mapper.readTree("""
{"type":"workspace_created",
"workspace":{"workspace_id":"w9","label":"bridged-workers","focused":false,
"pane_count":1,"tab_count":1,"active_tab_id":"w9:t1","agent_status":"unknown"},
"tab":{"tab_id":"w9:t1","workspace_id":"w9","label":"1","pane_count":1},
"root_pane":{"pane_id":"w9:p1","workspace_id":"w9","tab_id":"w9:t1"}}""");
case "tab.create" -> mapper.readTree("""
case "tab.create" -> {
// Each tab gets its own seed pane — under protocol 19 that pane becomes the
// worker pane, so distinct spawns must yield distinct pane ids.
long tabs = calls.stream().filter(c -> c.method().equals("tab.create")).count();
yield mapper.readTree(("""
{"type":"tab_created",
"tab":{"tab_id":"w9:t2","workspace_id":"w9","label":"2","pane_count":1},
"root_pane":{"pane_id":"w9:pRoot","workspace_id":"w9","tab_id":"w9:t2"}}""");
"root_pane":{"pane_id":"w9:pRoot_%d","workspace_id":"w9","tab_id":"w9:t2"}}""")
.formatted(tabs));
}
case "tab.rename" -> mapper.readTree("""
{"type":"tab_info","tab":{"tab_id":"w9:t2","workspace_id":"w9",
"label":"worker: ltms-local","pane_count":1}}""");
case "tab.list" -> mapper.readTree(("""
case "tab.list" -> {
// The base pair is returned for every workspace, exactly as before. Seeded tabs
// are appended only for the workspace they were registered against, so a test
// that seeds none sees the historical response byte for byte.
Object wsId = params instanceof java.util.Map<?, ?> m ? m.get("workspace_id") : null;
List<String> seeded = extraTabs.getOrDefault(String.valueOf(wsId), List.of());
yield mapper.readTree(("""
{"type":"tab_list","tabs":[
{"tab_id":"w9:t1","workspace_id":"w9","label":"1","pane_count":1},
{"tab_id":"w9:t2","workspace_id":"w9","label":"worker: ltms-local","pane_count":%d}]}""")
.formatted(workerTabPaneCount));
{"tab_id":"w9:t2","workspace_id":"w9","label":"worker: ltms-local","pane_count":%d}%s]}""")
.formatted(workerTabPaneCount,
seeded.isEmpty() ? "" : "," + String.join(",", seeded)));
}
case "tab.close" -> mapper.readTree("{\"type\":\"ok\"}");
case "pane.get" -> mapper.readTree("""
{"type":"pane_info","pane":{"pane_id":"w9:pW","workspace_id":"w9",
@@ -0,0 +1,262 @@
package dev.ltms.bridged.herdr;
import com.fasterxml.jackson.databind.JsonNode;
import com.fasterxml.jackson.databind.ObjectMapper;
import org.junit.jupiter.api.Test;
import java.util.ArrayList;
import java.util.LinkedHashMap;
import java.util.List;
import java.util.Map;
import java.util.Set;
import java.util.concurrent.TimeUnit;
import java.util.concurrent.atomic.AtomicLong;
import static org.junit.jupiter.api.Assertions.*;
/**
* CB-531. A lead is never spawned, so the daemon has to <em>find</em> it: these assert that an
* operator-labelled tab is what makes a pane a lead, and — just as importantly — what does not.
*/
class LeadTabScannerTest {
private static final ObjectMapper MAPPER = new ObjectMapper();
private static final long TTL = TimeUnit.SECONDS.toNanos(10);
/**
* A herdr whose workspace/tab/pane topology is declared per test. Counts calls so the caching
* contract can be asserted, and can be made to fail on demand.
*/
private static final class TopologyHerdr implements HerdrClient {
/** workspace_id → label. */
final Map<String, String> workspaces = new LinkedHashMap<>();
/** tab_id → [workspace_id, label]. */
final Map<String, String[]> tabs = new LinkedHashMap<>();
/** pane_id → [tab_id, terminal_id]. */
final Map<String, String[]> panes = new LinkedHashMap<>();
int calls;
boolean failing;
TopologyHerdr workspace(String id, String label) {
workspaces.put(id, label);
return this;
}
TopologyHerdr tab(String tabId, String workspaceId, String label) {
tabs.put(tabId, new String[]{workspaceId, label});
return this;
}
TopologyHerdr pane(String paneId, String tabId, String terminalId) {
panes.put(paneId, new String[]{tabId, terminalId});
return this;
}
@Override
public JsonNode call(String method, Object params) {
calls++;
if (failing) {
throw new HerdrException("socket closed");
}
List<String> items = new ArrayList<>();
switch (method) {
case "workspace.list" -> {
workspaces.forEach((id, label) -> items.add(
"{\"workspace_id\":\"%s\",\"label\":\"%s\"}".formatted(id, label)));
return read("{\"workspaces\":[%s]}".formatted(String.join(",", items)));
}
case "tab.list" -> {
String ws = String.valueOf(((Map<?, ?>) params).get("workspace_id"));
tabs.forEach((id, t) -> {
if (ws.equals(t[0])) {
items.add(("{\"tab_id\":\"%s\",\"workspace_id\":\"%s\",\"label\":%s,"
+ "\"pane_count\":1}").formatted(id, t[0],
t[1] == null ? "null" : "\"" + t[1] + "\""));
}
});
return read("{\"tabs\":[%s]}".formatted(String.join(",", items)));
}
case "pane.list" -> {
panes.forEach((id, p) -> items.add(
"{\"pane_id\":\"%s\",\"tab_id\":\"%s\",\"terminal_id\":\"%s\"}"
.formatted(id, p[0], p[1])));
return read("{\"panes\":[%s]}".formatted(String.join(",", items)));
}
default -> throw new AssertionError("unexpected herdr call: " + method);
}
}
private static JsonNode read(String json) {
try {
return MAPPER.readTree(json);
} catch (Exception e) {
throw new AssertionError(e);
}
}
@Override
public void close() {
}
}
/**
* The usual shape: one user space with lead tabs, one worker space bridged owns.
*
* <p>Not closed: {@code close()} is a no-op on this fake, and every test needs the handle after
* the scanner is built (to mutate the topology or read {@code calls}).
*/
@SuppressWarnings("resource")
private TopologyHerdr twoLeads() {
return new TopologyHerdr()
.workspace("w1", "main")
.workspace("w9", "bridged-workers")
.tab("w1:t1", "w1", "lead: opus-5.0")
.tab("w1:t2", "w1", "lead: gpt-sol-5.6")
.tab("w1:t3", "w1", "notes")
.tab("w9:t1", "w9", "worker: gx10 #1")
.pane("w1:p1", "w1:t1", "term_opus")
.pane("w1:p2", "w1:t2", "term_gpt")
.pane("w1:p3", "w1:t3", "term_notes")
.pane("w9:p1", "w9:t1", "term_worker");
}
private LeadTabScanner scanner(TopologyHerdr herdr, Map<String, String> configured,
AtomicLong clock) {
return new LeadTabScanner(herdr, "lead:", Set.of("bridged-workers"), configured, TTL,
clock::get);
}
@Test
void everyLabelledTabBecomesALeadNamedByItsLabel() {
Map<String, String> leads = scanner(twoLeads(), Map.of(), new AtomicLong()).get();
assertEquals(Map.of("term_opus", "opus-5.0", "term_gpt", "gpt-sol-5.6"), leads,
"two leads discovered from labels alone — no terminal_id was ever configured");
}
@Test
void anUnlabelledTabContributesNothing() {
assertFalse(scanner(twoLeads(), Map.of(), new AtomicLong()).get().containsKey("term_notes"));
}
/**
* The guard that matters: bridged labels its own worker tabs, so if a worker space were scanned
* a naming accident would promote the fleet. The exclusion is by workspace, not by hoping the
* worker template never collides.
*/
@Test
void aTabInAWorkerSpaceIsNeverALeadEvenWhenItsLabelMatches() {
TopologyHerdr herdr = twoLeads().tab("w9:t2", "w9", "lead: impostor")
.pane("w9:p2", "w9:t2", "term_impostor");
assertFalse(scanner(herdr, Map.of(), new AtomicLong()).get().containsKey("term_impostor"));
}
@Test
void aBarePrefixNamesNobodyAndIsRejected() {
TopologyHerdr herdr = new TopologyHerdr().workspace("w1", "main")
.tab("w1:t1", "w1", "lead:").pane("w1:p1", "w1:t1", "term_a");
assertEquals(Map.of(), scanner(herdr, Map.of(), new AtomicLong()).get(),
"a lead with no name would resolve as PRIMARY with nothing to attribute it to");
}
@Test
void thePrefixMatchesCaseInsensitivelyAndTheNameIsTrimmed() {
TopologyHerdr herdr = new TopologyHerdr().workspace("w1", "main")
.tab("w1:t1", "w1", " LEAD: opus-5.0 ").pane("w1:p1", "w1:t1", "term_a");
assertEquals(Map.of("term_a", "opus-5.0"), scanner(herdr, Map.of(), new AtomicLong()).get());
}
@Test
void everyPaneInALeadTabResolvesAsThatLead() {
// A human may split their own lead tab. Both panes are theirs, so both are that lead —
// nothing bridged placed can land here (see the worker-space test above).
TopologyHerdr herdr = twoLeads().pane("w1:p1b", "w1:t1", "term_opus_split");
assertEquals("opus-5.0", scanner(herdr, Map.of(), new AtomicLong()).get().get("term_opus_split"));
}
@Test
void anExplicitlyConfiguredLeadIsMergedInAndOutranksALabel() {
Map<String, String> configured = Map.of("term_opus", "pinned-name", "term_extra", "from-config");
Map<String, String> leads = scanner(twoLeads(), configured, new AtomicLong()).get();
assertEquals("pinned-name", leads.get("term_opus"), "an explicit pin is the operator's last word");
assertEquals("from-config", leads.get("term_extra"), "a configured lead needs no tab at all");
assertEquals("gpt-sol-5.6", leads.get("term_gpt"));
}
// ── caching ─────────────────────────────────────────────────────────────────────────────────
@Test
void aSecondLookupWithinTheTtlDoesNotTouchHerdr() {
TopologyHerdr herdr = twoLeads();
AtomicLong clock = new AtomicLong();
LeadTabScanner s = scanner(herdr, Map.of(), clock);
s.get();
int afterFirst = herdr.calls;
clock.addAndGet(TTL - 1);
s.get();
assertEquals(afterFirst, herdr.calls,
"resolve() runs on every request — an un-cached scan would put herdr on that path");
}
@Test
void aTabLabelledAfterStartupIsPickedUpOnceTheTtlExpires() {
TopologyHerdr herdr = twoLeads();
AtomicLong clock = new AtomicLong();
LeadTabScanner s = scanner(herdr, Map.of(), clock);
assertFalse(s.get().containsKey("term_notes"));
herdr.tab("w1:t3", "w1", "lead: late-arrival"); // the operator renames their tab
clock.addAndGet(TTL);
assertEquals("late-arrival", s.get().get("term_notes"),
"the whole point over `leaders:`: no config edit, no restart");
}
@Test
void aFailedScanKeepsTheLeadsAlreadyKnownRatherThanDemotingThem() {
TopologyHerdr herdr = twoLeads();
AtomicLong clock = new AtomicLong();
LeadTabScanner s = scanner(herdr, Map.of(), clock);
Map<String, String> before = s.get();
herdr.failing = true;
clock.addAndGet(TTL);
assertEquals(before, s.get(),
"a herdr hiccup must not silently demote a live lead to a worker mid-session");
}
@Test
void aFailedFirstScanStillHonoursTheConfiguredLeads() {
TopologyHerdr herdr = twoLeads();
herdr.failing = true;
Map<String, String> leads = scanner(herdr, Map.of("term_x", "opus-5.0"), new AtomicLong()).get();
assertEquals(Map.of("term_x", "opus-5.0"), leads,
"config-named leads must not depend on herdr answering at all");
}
@Test
void aDownHerdrIsRetriedOncePerTtlNotOncePerRequest() {
TopologyHerdr herdr = twoLeads();
herdr.failing = true;
AtomicLong clock = new AtomicLong();
LeadTabScanner s = scanner(herdr, Map.of(), clock);
s.get();
int afterFirst = herdr.calls;
s.get();
s.get();
assertEquals(afterFirst, herdr.calls, "the failure path must be rate-limited too");
}
}
@@ -5,16 +5,15 @@ import org.junit.jupiter.api.Tag;
import org.junit.jupiter.api.Test;
import java.nio.file.Files;
import java.util.List;
import java.util.Map;
import static org.junit.jupiter.api.Assertions.*;
import static org.junit.jupiter.api.Assumptions.assumeTrue;
/**
* Contract test for the herdr half of connection-based identity against a REAL herdr: spawn a
* harmless probe, read its actual {@code shell_pid} from {@code pane.process_info}, and confirm
* {@link PaneLocator} resolves that PID back to the probe's own {@code terminal_id}.
* Contract test for {@link PaneLocator} against a REAL herdr: the PID→pane mapping that
* connection identity rests on. Uses a throwaway tab's seed shell as the probe process
* (protocol 19 removed arbitrary-command agents), and always tears the space down.
*
* <p>Tagged {@code contract}; run with {@code mvn test -Pcontract}.
*/
@@ -25,18 +24,24 @@ class PaneLocatorContractTest {
void resolvesTheTerminalOwningARealProcessPid() throws Exception {
assumeTrue(Files.exists(UnixSocketHerdrClient.defaultSocketPath()), "no herdr socket — skipping");
try (UnixSocketHerdrClient herdr = UnixSocketHerdrClient.connect()) {
AgentControl agents = new AgentControl(herdr);
Agent probe = agents.start("__pidprobe__", List.of("bash", "-c", "sleep 20"), Map.of());
WorkspaceControl spaces = new WorkspaceControl(herdr);
Workspace space = spaces.ensureWorkspace("__bridged_pid_contract__");
Tab.Created tab = spaces.createTab(space.workspaceId(), null, Map.of());
try {
JsonNode info = herdr.call("pane.process_info", Map.of("pane_id", probe.paneId()))
JsonNode info = herdr.call("pane.process_info", Map.of("pane_id", tab.rootPaneId()))
.path("process_info");
long shellPid = info.path("shell_pid").asLong(-1);
assertTrue(shellPid > 0, "probe pane should report a shell pid");
assertTrue(shellPid > 0, "seed pane should report a shell pid");
assertEquals(probe.terminalId(), new PaneLocator(herdr).terminalForPid(shellPid),
String terminalId = herdr.call("pane.get", Map.of("pane_id", tab.rootPaneId()))
.path("pane").path("terminal_id").asText(null);
assertNotNull(terminalId, "seed pane should carry a terminal_id");
assertEquals(terminalId, new PaneLocator(herdr).terminalForPid(shellPid),
"a real PID must resolve back to its own pane's terminal_id");
} finally {
agents.close(probe.paneId());
spaces.closeTab(tab.tab().tabId());
herdr.call("workspace.close", Map.of("workspace_id", space.workspaceId()));
}
}
}
@@ -5,7 +5,6 @@ import org.junit.jupiter.api.Tag;
import org.junit.jupiter.api.Test;
import java.nio.file.Files;
import java.util.List;
import java.util.Map;
import static org.junit.jupiter.api.Assertions.*;
@@ -13,9 +12,9 @@ import static org.junit.jupiter.api.Assumptions.assumeTrue;
/**
* Contract test for the placement layer ({@code workspace.*}/{@code tab.*}) against a
* REAL herdr, locking in the "one clean tab per worker" recipe: find-or-create a worker
* space, give the worker its own tab, drop herdr's seed shell so the tab holds only the
* worker, and tear it all down. Uses a HARMLESS probe (never {@code claude}) in a
* REAL herdr, locking in the "one clean tab per worker" recipe under protocol 19: find-or-create
* a worker space, give the worker its own tab, and the SEED pane is where the worker starts —
* the tab holds exactly that one pane from creation. Uses no agent (never {@code claude}) in a
* throwaway space that is fully removed at the end.
*
* <p>Tagged {@code contract}; run with {@code mvn test -Pcontract}.
@@ -41,7 +40,6 @@ class WorkspacePlacementContractTest {
void workerGetsOwnCleanTabAndTearsDownCompletely() throws Exception {
assumeTrue(!noSocket(), "no herdr socket — skipping");
try (UnixSocketHerdrClient herdr = UnixSocketHerdrClient.connect()) {
AgentControl agents = new AgentControl(herdr);
WorkspaceControl spaces = new WorkspaceControl(herdr);
Workspace space = spaces.ensureWorkspace(LABEL);
@@ -49,36 +47,27 @@ class WorkspacePlacementContractTest {
// Idempotent: a second ensure finds the same space, never creates a duplicate.
assertEquals(space.workspaceId(), spaces.ensureWorkspace(LABEL).workspaceId());
Tab.Created tab = spaces.createTab(space.workspaceId());
Agent worker = agents.start(
"__contract__",
List.of("bash", "-c", "sleep 20"),
Map.of(),
tab.tab().tabId());
Tab.Created tab = spaces.createTab(space.workspaceId(), null, Map.of());
try {
// The worker landed in its dedicated tab in the worker space.
assertEquals(tab.tab().tabId(), worker.tabId());
assertEquals(space.workspaceId(), worker.workspaceId());
assertNotNull(tab.rootPaneId(), "tab.create must return the seed pane");
// Drop the seed shell; the tab now holds exactly the worker pane.
agents.close(tab.rootPaneId());
// Protocol 19: the seed pane IS the worker pane — the tab holds exactly it.
spaces.renameTab(tab.tab().tabId(), "worker: contract");
assertEquals(1, paneCount(herdr, space.workspaceId(), tab.tab().tabId()),
"worker tab must hold only the worker pane after the seed shell is dropped");
"worker tab must hold exactly the seed/worker pane");
// Teardown resolves the tab from the pane, and sees it holds exactly one pane.
WorkspaceControl.PaneLocation loc = spaces.locatePane(worker.paneId());
WorkspaceControl.PaneLocation loc = spaces.locatePane(tab.rootPaneId());
assertNotNull(loc);
assertEquals(tab.tab().tabId(), loc.tabId());
assertEquals(1, loc.tabPaneCount(), "worker is the tab's sole occupant");
assertEquals(1, loc.tabPaneCount(), "worker pane is the tab's sole occupant");
} finally {
agents.close(worker.paneId());
spaces.closeTab(tab.tab().tabId());
}
// Tolerant teardown: closing an already-gone tab / reading a gone pane is a no-op.
spaces.closeTab(tab.tab().tabId());
assertNull(spaces.locatePane(worker.paneId()), "closed worker pane must be gone");
assertNull(spaces.locatePane(tab.rootPaneId()), "closed worker pane must be gone");
// Remove the throwaway space entirely so the test leaves no residue.
herdr.call("workspace.close", Map.of("workspace_id", space.workspaceId()));
@@ -156,13 +156,56 @@ class CompletionResolverTest {
assertEquals("No, 391 = 17 × 23.", waiter.getNow(null).text());
}
@Test
void marksAClippedCompletionPaneTail() {
String block = "⏺ " + "x".repeat(CompletionResolver.MAX_SCRAPE_CHARS + 1) + "\n❯ ";
FakeHerdr herdr = new FakeHerdr().readText(block);
Rendezvous rendezvous = new Rendezvous();
CompletionResolver resolver = new CompletionResolver(new AgentControl(herdr), rendezvous);
var waiter = rendezvous.open("term_a");
resolver.resolve("term_a", new CompletionResolver.InFlight(waiter, null));
assertEquals("x".repeat(CompletionResolver.MAX_SCRAPE_CHARS)
+ "\n[Pane tail clipped: member did not call bridge_reply.]",
waiter.getNow(null).text());
}
@Test
void leavesAnUnclippedCompletionPaneTailUnmarked() {
FakeHerdr herdr = new FakeHerdr().readText("⏺ complete report\n❯ ");
Rendezvous rendezvous = new Rendezvous();
CompletionResolver resolver = new CompletionResolver(new AgentControl(herdr), rendezvous);
var waiter = rendezvous.open("term_a");
resolver.resolve("term_a", new CompletionResolver.InFlight(waiter, null));
assertEquals("complete report", waiter.getNow(null).text());
}
@Test
void resolvesSynchronouslyBeforePostTurnContextClearing() {
FakeHerdr herdr = new FakeHerdr().readText("⏺ previous answer\n❯ ");
Rendezvous rendezvous = new Rendezvous();
CompletionResolver resolver = new CompletionResolver(new AgentControl(herdr), rendezvous);
var waiter = rendezvous.open("term_a");
resolver.captureBaseline("term_a");
herdr.readText("⏺ answer that /clear would erase\n❯ ");
resolver.resolveBeforePostAction("term_a");
assertTrue(waiter.isDone(), "the answer is captured before the adapter sends /clear");
assertEquals("answer that /clear would erase", waiter.getNow(null).text());
}
@Test
void suppressesAnUnchangedCompletionEvenWhenTheBlockExceedsTheScrapeCap() {
// The fan-out issue-hunt finding: captureBaseline once stored the RAW (unclipped) assistant
// block while resolve compares against a clip()'d tail. For a block longer than MAX_SCRAPE_CHARS
// the two capped representations differ even when the pane never changed, so the CB-115
// byte-identical guard failed to fire and a stale completion could resolve the send. Both sides
// must clip identically; here an unchanged >cap block on rapid back-to-back turns stays suppressed.
// must clip identically. The returned-text marker is added only after this comparison, so an
// unchanged >cap block on rapid back-to-back turns still stays suppressed.
String longBlock = "⏺ " + "x".repeat(CompletionResolver.MAX_SCRAPE_CHARS + 500) + "\n❯ ";
FakeHerdr herdr = new FakeHerdr().readText(longBlock);
Rendezvous rendezvous = new Rendezvous();
@@ -271,10 +314,12 @@ class CompletionResolverTest {
// The turn as the injector captured it at delivery (waiter + pre-turn baseline).
var turnN = new CompletionResolver.InFlight(waiterN, "an earlier answer");
// Turn N is resolved by the worker's explicit reply.
// Turn N is resolved by the worker's explicit reply, and its send deregisters the waiter.
assertTrue(rendezvous.resolve("term_a", "N replied"));
rendezvous.close("term_a", waiterN); // the sender's finally, before the next turn opens
// Turn N+1's send opens its own waiter on the same session (replacing the registered one).
// Turn N+1's send opens its own waiter on the same session (CB-548: open fails if the
// previous waiter is still registered, so a clean turn deregisters it first as above).
var waiterN1 = rendezvous.open("term_a");
resolver.resolve("term_a", turnN); // turn N's completion fallback finally fires
@@ -1,10 +1,15 @@
package dev.ltms.bridged.inject;
import ch.qos.logback.classic.Level;
import ch.qos.logback.classic.LoggerContext;
import ch.qos.logback.classic.spi.ILoggingEvent;
import ch.qos.logback.core.read.ListAppender;
import dev.ltms.bridged.herdr.AgentControl;
import dev.ltms.bridged.herdr.AgentStatus;
import dev.ltms.bridged.herdr.FakeHerdr;
import dev.ltms.bridged.herdr.HerdrException;
import org.junit.jupiter.api.Test;
import org.slf4j.LoggerFactory;
import java.util.ArrayList;
import java.util.List;
@@ -28,17 +33,15 @@ class InjectorTest {
private final Injector injector = new Injector(new AgentControl(herdr));
/**
* The logical messages delivered, in order. AgentControl.send emits each delivery as two
* agent.send calls — the payload, then a standalone Enter keystroke ({@code "\r"}) to submit
* it; these tests assert delivery ordering/gating, not the submit event, so drop the bare
* carriage returns.
* The logical messages delivered, in order. Under protocol 19 each delivery is one
* {@code agent.prompt} carrying the payload (it submits itself); the Enter nudge is a
* separate {@code agent.send_keys} and never appears here.
*/
@SuppressWarnings("unchecked")
private List<String> sent() {
return herdr.calls.stream()
.filter(c -> c.method().equals("agent.send"))
.filter(c -> c.method().equals("agent.prompt"))
.map(c -> ((Map<String, Object>) c.params()).get("text").toString())
.filter(t -> !t.equals("\r"))
.toList();
}
@@ -66,11 +69,9 @@ class InjectorTest {
assertEquals(List.of("task"), sent(), "delivers once the worker is available");
}
@SuppressWarnings("unchecked")
private long enterKeystrokes() {
return herdr.calls.stream()
.filter(c -> c.method().equals("agent.send"))
.filter(c -> "\r".equals(((Map<String, Object>) c.params()).get("text")))
.filter(c -> c.method().equals("agent.send_keys"))
.count();
}
@@ -200,6 +201,47 @@ class InjectorTest {
assertEquals(List.of(T), completed, "a confirmed working→idle fires exactly one completion");
}
@Test
void postTurnResetSettlesBeforeNextDelegationWithoutBecomingATurn() {
AgentControl agents = new AgentControl(herdr);
class ResetListener implements TurnListener {
int completed;
@Override
public void onTurnComplete(String target) {
completed++;
}
@Override
public boolean hasPostTurnAction(String target) {
return true;
}
@Override
public boolean onTurnCompleteWithPostAction(String target) {
completed++;
agents.send(target, "/clear"); // direct housekeeping, never Injector.enqueue
return true;
}
}
ResetListener listener = new ResetListener();
Injector inj = new Injector(agents, listener);
inj.enqueue(T, "first");
inj.enqueue(T, "second");
inj.onStatus(T, AgentStatus.IDLE); // first delegation
inj.onStatus(T, AgentStatus.WORKING);
inj.onStatus(T, AgentStatus.IDLE); // first complete; reset dispatched
assertEquals(List.of("first", "/clear"), sent(),
"same-tick completion must not let the queued delegation overtake reset");
inj.onStatus(T, AgentStatus.WORKING); // reset picked up, but this is not a bridge turn
inj.onStatus(T, AgentStatus.IDLE); // reset settled; second may now deliver
assertEquals(List.of("first", "/clear", "second"), sent());
assertEquals(1, listener.completed,
"reset settlement must not recursively emit another turn completion");
}
@Test
void doesNotSynthesizeCompletionFromAnUnconfirmedTurn() {
List<String> completed = new ArrayList<>();
@@ -332,6 +374,40 @@ class InjectorTest {
assertTrue(inj.activeTargets().isEmpty(), "the target is reclaimed, not polled forever");
}
@Test
void readinessGraceExpiryIsLogged() {
// CB-562: the grace-expiry path used to clear the queue silently, so a message that never
// reached the worker's pane surfaced elsewhere as an unrelated turn-stall failure. Assert the
// expiry now names the real cause. (ListAppender capture pattern mirrors AuditLogTest.)
LoggerContext ctx = (LoggerContext) LoggerFactory.getILoggerFactory();
ch.qos.logback.classic.Logger injectorLog =
(ch.qos.logback.classic.Logger) LoggerFactory.getLogger(Injector.class);
ListAppender<ILoggingEvent> appender = new ListAppender<>();
appender.setContext(ctx);
appender.start();
injectorLog.addAppender(appender);
injectorLog.setLevel(Level.WARN);
try {
Injector inj = new Injector(new AgentControl(herdr), TurnListener.NOOP, _ -> false, _ -> {
});
inj.enqueue(T, "task");
for (int i = 0; i < READINESS_SAMPLES; i++) inj.onStatus(T, AgentStatus.IDLE);
String warn = appender.list.stream()
.filter(e -> e.getLevel().equals(Level.WARN))
.map(ILoggingEvent::getFormattedMessage)
.findFirst()
.orElse("no grace-expiry WARN logged");
assertTrue(warn.contains(T), "the log names the target terminal: " + warn);
assertTrue(warn.contains("never reached"), "the log names the real cause: " + warn);
assertTrue(warn.contains("1 queued message"),
"the log carries the failed message count: " + warn);
} finally {
injectorLog.detachAppender(appender);
}
}
@Test
void aWorkerThatBecomesReadyWithinTheGraceIsDeliveredNormally() {
// The readiness grace must not fail a worker that is merely slow to boot: once it becomes
@@ -352,7 +428,7 @@ class InjectorTest {
@Test
void dropClearsWorkerPresence() {
// CB-114 (finding #1): a vanished worker's readiness must be forgotten so a stale entry cannot
// linger past the worker's life (WorkerPresence.forget had no caller before this).
// linger past the worker's life (MemberPresence.forget had no caller before this).
List<String> forgotten = new ArrayList<>();
Injector inj = new Injector(new AgentControl(herdr), TurnListener.NOOP, _ -> true, forgotten::add);
inj.enqueue(T, "orphan");
@@ -374,13 +450,12 @@ class InjectorTest {
poller.stop();
}
assertEquals(List.of("via-poller"), idle.calls.stream()
.filter(c -> c.method().equals("agent.send"))
.filter(c -> c.method().equals("agent.prompt"))
.map(c -> {
@SuppressWarnings("unchecked")
Map<String, Object> p = (Map<String, Object>) c.params();
return p.get("text").toString();
})
.filter(t -> !t.equals("\r")) // drop the standalone submit keystroke
.toList());
}
@@ -11,7 +11,7 @@ class WorkerPresenceTest {
@Test
void tracksPresenceAndForgets() {
WorkerPresence p = new WorkerPresence();
MemberPresence p = new MemberPresence();
assertFalse(p.isPresent("term_a"), "unseen worker is not available");
p.markPresent("term_a");
@@ -23,7 +23,7 @@ class WorkerPresenceTest {
@Test
void nullOrBlankMarkIsANoOp() {
WorkerPresence p = new WorkerPresence();
MemberPresence p = new MemberPresence();
assertDoesNotThrow(() -> {
p.markPresent(null);
p.markPresent(" ");
@@ -0,0 +1,255 @@
package dev.ltms.bridged.lead;
import dev.ltms.bridged.config.BridgedConfig;
import dev.ltms.bridged.herdr.AgentControl;
import dev.ltms.bridged.herdr.FakeHerdr;
import dev.ltms.bridged.herdr.WorkspaceControl;
import org.junit.jupiter.api.Test;
import java.util.LinkedHashMap;
import java.util.List;
import java.util.Map;
import static org.junit.jupiter.api.Assertions.*;
/**
* CB-558 — the daemon starts a declared lead when none is running.
*
* <p>Two properties carry the whole feature. It must not double-spawn (a second orchestrator is
* worse than none), and what it starts must be a <em>lead</em> and not a member: no reply charter,
* no off-subscription env, and never registered with the session lifecycle.
*/
class LeadLauncherTest {
/** The lead's backend: a subscription profile with the bridge mounted, as `opus` really is. */
private static BridgedConfig.Profile opusProfile() {
return new BridgedConfig.Profile(
"opus", null, "claude-opus-5", null, "BRIDGED_WORKER_TOKEN",
List.of("ccs", "ltms"), "tab", "bridged-workers", null,
"http://127.0.0.1:8765/mcp", null, null,
null, null, null,
Map.of("CLAUDE_CODE_AUTO_COMPACT_WINDOW", "300000"), null, null, true);
}
private static BridgedConfig configWith(BridgedConfig.Leader lead) {
Map<String, BridgedConfig.Leader> leaders = new LinkedHashMap<>();
leaders.put("opus", lead);
BridgedConfig.Fleet fleet =
new BridgedConfig.Fleet(leaders, Map.of(), Map.of(), Map.of(), null);
return new BridgedConfig(
null, null, Map.of("opus", opusProfile()), null, null, null, null, null,
null, null, fleet, null, "fixed", null).withDefaults();
}
private static BridgedConfig.Leader lead(String profile, String terminal, int instances) {
return new BridgedConfig.Leader(profile, terminal, instances, "lead:", 10, null, null,
"leads", "/repo");
}
private static LeadLauncher launcher(FakeHerdr herdr, BridgedConfig cfg) {
return new LeadLauncher(new AgentControl(herdr), new WorkspaceControl(herdr), cfg);
}
@SuppressWarnings("unchecked")
private static List<String> startedArgs(FakeHerdr herdr) {
return (List<String>) ((Map<String, Object>) herdr.lastCall("agent.start").params()).get("args");
}
private static String startedName(FakeHerdr herdr) {
return (String) ((Map<?, ?>) herdr.lastCall("agent.start").params()).get("name");
}
@SuppressWarnings("unchecked")
private static Map<String, String> tabEnv(FakeHerdr herdr) {
return (Map<String, String>) ((Map<String, Object>) herdr.lastCall("tab.create").params()).get("env");
}
// ── it starts a lead when none is live ────────────────────────────────────────────────────
@Test
void startsTheDeclaredLeadWhenNoneIsRunning() {
FakeHerdr herdr = new FakeHerdr();
assertEquals(1, launcher(herdr, configWith(lead("opus", null, 1))).ensureLeads());
assertTrue(herdr.called("agent.start"), "a lead must actually be started");
assertEquals("lead-opus", startedName(herdr));
}
/** The tab is labelled so the scanner finds the lead on the next resolve. */
@Test
void labelsTheTabWithThePrefixTheScannerReadsBack() {
FakeHerdr herdr = new FakeHerdr();
launcher(herdr, configWith(lead("opus", null, 1))).ensureLeads();
assertEquals("lead: opus",
((Map<?, ?>) herdr.lastCall("tab.rename").params()).get("label"));
}
/** `instances: 2` with none live means two starts, not one. */
@Test
void startsAsManyInstancesAsAreDeclared() {
FakeHerdr herdr = new FakeHerdr();
assertEquals(2, launcher(herdr, configWith(lead("opus", null, 2))).ensureLeads());
assertEquals(2, herdr.calls.stream().filter(c -> c.method().equals("agent.start")).count());
}
// ── it must not double-spawn ──────────────────────────────────────────────────────────────
/** A labelled tab WITH a running agent in it is a live lead — leave it alone. */
@Test
void doesNotStartASecondLeadWhenOneIsAlreadyRunning() {
FakeHerdr herdr = new FakeHerdr()
.withWorkspace("wL", "leads")
.withTab("wL", "wL:t1", "lead: opus")
.withAgent("lead-opus", "term_lead", "wL:p1", "wL:t1");
assertEquals(0, launcher(herdr, configWith(lead("opus", null, 1))).ensureLeads());
assertFalse(herdr.called("agent.start"), "the live lead must not be duplicated");
}
/**
* The reason liveness is not "does the label exist". A tab left labelled by a session that has
* since died must not block the relaunch, or one crash disables auto-launch permanently.
*/
@Test
void aLabelledTabWithNoRunningAgentIsNotALiveLead() {
FakeHerdr herdr = new FakeHerdr()
.withWorkspace("wL", "leads")
.withTab("wL", "wL:t1", "lead: opus"); // label only — nothing running in it
assertEquals(1, launcher(herdr, configWith(lead("opus", null, 1))).ensureLeads(),
"a stale label is not a lead; the lead must be relaunched");
}
/**
* A lead the operator opened by hand and pinned with `terminal:` is live even though its tab
* carries no matching label. Counting labels alone would relaunch it on every boot.
*/
@Test
void aPinnedTerminalWithARunningAgentCountsAsLive() {
FakeHerdr herdr = new FakeHerdr()
.withAgent("hand-opened", "term_pinned", "wX:p1", "wX:t1");
assertEquals(0, launcher(herdr, configWith(lead("opus", "term_pinned", 1))).ensureLeads());
assertFalse(herdr.called("agent.start"));
}
/** A member sitting in a matching tab must never be counted — or spawn a lead — as one. */
@Test
void aMemberWorkspaceIsNeverScannedForLeads() {
FakeHerdr herdr = new FakeHerdr()
.withWorkspace("wM", "bridged-workers") // a configured member space
.withTab("wM", "wM:t1", "lead: opus") // a member tab that looks like a lead
.withAgent("claude-opus-x", "term_m", "wM:p1", "wM:t1");
assertEquals(1, launcher(herdr, configWith(lead("opus", null, 1))).ensureLeads(),
"a member in a lead-labelled tab is not a lead, so the real lead is still missing");
}
/** If herdr cannot be counted, start nothing: guessing risks a second orchestrator. */
@Test
void anUncountableHerdrStartsNothing() {
FakeHerdr herdr = new FakeHerdr().healthy(false);
assertEquals(0, launcher(herdr, configWith(lead("opus", null, 1))).ensureLeads());
assertFalse(herdr.called("agent.start"));
}
// ── what it starts is a LEAD, not a member ────────────────────────────────────────────────
/**
* The single most important assertion here. The worker charter tells its reader it is an
* off-subscription worker that must end every turn with bridge_reply — the opposite of what an
* orchestrator is. A lead must never receive it.
*/
@Test
void theLeadNeverReceivesTheWorkerReplyCharter() {
FakeHerdr herdr = new FakeHerdr();
launcher(herdr, configWith(lead("opus", null, 1))).ensureLeads();
List<String> args = startedArgs(herdr);
assertFalse(args.contains("--append-system-prompt"),
"the reply charter is a worker contract and must not be injected into a lead");
assertTrue(args.stream().noneMatch(a -> a.contains("bridge_reply")), args.toString());
}
/** It still mounts the bridge — a lead that cannot orchestrate is pointless. */
@Test
void theLeadMountsTheBridgeMcpAndPinsItsModel() {
FakeHerdr herdr = new FakeHerdr();
launcher(herdr, configWith(lead("opus", null, 1))).ensureLeads();
List<String> args = startedArgs(herdr);
assertTrue(args.contains("--mcp-config"));
assertTrue(args.stream().anyMatch(a -> a.contains("http://127.0.0.1:8765/mcp")), args.toString());
assertEquals("claude-opus-5", args.get(args.indexOf("--model") + 1));
assertTrue(args.indexOf("--model") > args.indexOf("--mcp-config"),
"--model is appended last so it outranks the ccs wrapper (CB-533)");
}
/** A lead runs on the operator's subscription. Nothing may move it off. */
@Test
void theLeadEnvCarriesNoAnthropicBinding() {
FakeHerdr herdr = new FakeHerdr();
launcher(herdr, configWith(lead("opus", null, 1))).ensureLeads();
Map<String, String> env = tabEnv(herdr);
assertNull(env.get("ANTHROPIC_BASE_URL"));
assertNull(env.get("ANTHROPIC_AUTH_TOKEN"));
assertEquals("300000", env.get("CLAUDE_CODE_AUTO_COMPACT_WINDOW"),
"the profile's own env: still applies");
}
/** The lead's tab goes in its own workspace, never a member one — the scanner skips those. */
@Test
void theLeadTabIsCreatedOutsideEveryMemberWorkspace() {
FakeHerdr herdr = new FakeHerdr();
launcher(herdr, configWith(lead("opus", null, 1))).ensureLeads();
String label = (String) ((Map<?, ?>) herdr.lastCall("workspace.create").params()).get("label");
assertEquals("leads", label);
assertNotEquals("bridged-workers", label);
}
// ── recognise-only and misconfiguration ───────────────────────────────────────────────────
/** A lead with a pin but no profile is recognise-only by design — not an error, not a launch. */
@Test
void aLeadThatNamesNoProfileIsRecognisedButNeverLaunched() {
FakeHerdr herdr = new FakeHerdr();
assertEquals(0, launcher(herdr, configWith(lead(null, "term_dead", 1))).ensureLeads());
assertFalse(herdr.called("agent.start"));
}
/** `instances: 0` is a deliberate off switch. */
@Test
void zeroInstancesLaunchesNothing() {
FakeHerdr herdr = new FakeHerdr();
assertEquals(0, launcher(herdr, configWith(lead("opus", null, 0))).ensureLeads());
assertFalse(herdr.called("agent.start"));
}
/** A profile name with no matching profile is logged and skipped, never a daemon crash. */
@Test
void anUnknownProfileIsSkippedRatherThanThrown() {
FakeHerdr herdr = new FakeHerdr();
assertEquals(0, launcher(herdr, configWith(lead("nope", null, 1))).ensureLeads());
assertFalse(herdr.called("agent.start"));
}
/** No leads declared at all: not a herdr call in sight. */
@Test
void noLeadersConfiguredTouchesHerdrNotAtAll() {
FakeHerdr herdr = new FakeHerdr();
BridgedConfig cfg = new BridgedConfig(
null, null, Map.of("opus", opusProfile()), null, null, null, null, null,
null, null, null, null, "fixed", null).withDefaults();
assertEquals(0, launcher(herdr, cfg).ensureLeads());
assertTrue(herdr.calls.isEmpty(), "nothing declared ⇒ nothing scanned");
}
}
@@ -2,6 +2,7 @@ package dev.ltms.bridged.mcp;
import dev.ltms.bridged.auth.Authz;
import dev.ltms.bridged.auth.CallerResolver;
import dev.ltms.bridged.auth.MemberRegistry;
import dev.ltms.bridged.auth.Principal;
import dev.ltms.bridged.auth.Role;
import dev.ltms.bridged.config.BridgedConfig;
@@ -18,7 +19,7 @@ import dev.ltms.bridged.msg.MessageService;
import dev.ltms.bridged.msg.Rendezvous;
import dev.ltms.bridged.session.FakeWorktrees;
import dev.ltms.bridged.session.SessionManager;
import dev.ltms.bridged.worker.ClaudeCodeLauncher;
import dev.ltms.bridged.member.ClaudeCodeLauncher;
import io.modelcontextprotocol.spec.McpSchema;
import org.junit.jupiter.api.AfterEach;
import org.junit.jupiter.api.Test;
@@ -55,7 +56,7 @@ class BridgeMcpAuthzTest {
/** A fully wired BridgeMcp on fakes — constructing it is itself part of what is under test. */
private BridgeMcp mcp(boolean enforce) {
BridgedConfig.Worker cfg = new BridgedConfig.Worker(
BridgedConfig.Profile cfg = new BridgedConfig.Profile(
"ltms-local", "http://gx00.gw:8000", "coder", null, "BRIDGED_WORKER_TOKEN", null,
"tab", "bridged-workers", "worker: {profile} #{n}", null, null, null);
ClaudeCodeLauncher workers = new ClaudeCodeLauncher(agents, new WorkspaceControl(herdr),
@@ -69,7 +70,8 @@ class BridgeMcpAuthzTest {
mcp = new BridgeMcp(messages, workers, sessions, identity, sessions.asPresence(),
new PrimaryRegistry(null),
enforce ? new CallerResolver(identity) : null,
enforce ? CallerResolver.withLeadsAndMembers(identity, false, null,
Map::of, new MemberRegistry(null)) : null,
metrics);
return mcp;
}
@@ -77,6 +79,7 @@ class BridgeMcpAuthzTest {
private static final Principal PRIMARY = Principal.primary(100);
private static final Principal WORKER_A = Principal.worker("term_a", 200);
private static final Principal ANON = Principal.anonymous();
private static final Principal ARCH_DESIGN = Principal.architect("lead-designer", "term_design", 400);
// --- the table, enforced on THIS path too ---------------------------------------------------
@@ -119,6 +122,33 @@ class BridgeMcpAuthzTest {
assertNotNull(m.denyFor(PRIMARY, Authz.Action.ASK, "term_a"));
}
// --- CB-548: the architect on this path ------------------------------------------------
@Test
void anArchitectMaySendAndReadButNotOrchestrateOverMcp() {
BridgeMcp m = mcp(true);
assertNull(m.denyFor(ARCH_DESIGN, Authz.Action.SEND, "term_a"),
"delegating a turn is the architect's job");
assertNull(m.denyFor(ARCH_DESIGN, Authz.Action.READ, null));
for (Authz.Action a : new Authz.Action[]{Authz.Action.SPAWN, Authz.Action.STOP,
Authz.Action.DRAIN}) {
McpSchema.CallToolResult denied = m.denyFor(ARCH_DESIGN, a, null);
assertNotNull(denied, a + " must be refused to an architect");
assertTrue(denied.isError(), "a refusal is returned as an MCP tool error");
}
}
@Test
void anArchitectMayReplyAndAskOnlyAsItsOwnPaneOverMcp() {
BridgeMcp m = mcp(true);
assertNull(m.denyFor(ARCH_DESIGN, Authz.Action.REPLY, "term_design"));
assertNull(m.denyFor(ARCH_DESIGN, Authz.Action.ASK, "term_design"));
assertNotNull(m.denyFor(ARCH_DESIGN, Authz.Action.REPLY, "term_a"),
"architect 'lead-designer' must not reply on worker term_a's session");
}
@Test
void anonymousIsRefusedEverythingAndCountedAsUnauthenticated() {
BridgeMcp m = mcp(true);
@@ -156,6 +186,11 @@ class BridgeMcpAuthzTest {
assertEquals("term_a", BridgeMcp.principalFrom("WORKER", "term_a", 7).terminal());
assertEquals(Role.PRIMARY, BridgeMcp.principalFrom("PRIMARY", null, 7).role());
assertEquals(Role.ANONYMOUS, BridgeMcp.principalFrom("ANONYMOUS", null, -1).role());
// CB-548: an architect round-trips through the same stash, carrying its slot name.
Principal arch = BridgeMcp.principalFrom("ARCHITECT", "term_design", 7, "lead-designer");
assertEquals(Role.ARCHITECT, arch.role());
assertEquals("lead-designer", arch.name());
assertEquals("term_design", arch.terminal());
}
@Test
@@ -11,10 +11,14 @@ import dev.ltms.bridged.msg.MessageService;
import dev.ltms.bridged.msg.Rendezvous;
import dev.ltms.bridged.session.FakeWorktrees;
import dev.ltms.bridged.session.SessionManager;
import dev.ltms.bridged.session.WorkerSession;
import dev.ltms.bridged.peer.MemberRole;
import dev.ltms.bridged.inject.MemberPresence;
import dev.ltms.bridged.session.MemberSession;
import dev.ltms.bridged.session.WorktreeRequest;
import dev.ltms.bridged.worker.ClaudeCodeLauncher;
import dev.ltms.bridged.member.ClaudeCodeLauncher;
import io.modelcontextprotocol.spec.McpSchema;
import dev.ltms.bridged.msg.InMemoryReplyInbox;
import org.junit.jupiter.api.BeforeEach;
import org.junit.jupiter.api.Test;
import java.util.Map;
@@ -31,17 +35,26 @@ import static org.junit.jupiter.api.Assertions.*;
*/
class BridgeMcpTest {
private static final String T = "term_a";
private final FakeHerdr herdr = new FakeHerdr();
private final AgentControl agents = new AgentControl(herdr);
private final Rendezvous rendezvous = new Rendezvous();
private final MessageService messages = new MessageService(agents, new Injector(agents), rendezvous);
private final InMemoryReplyInbox inbox = new InMemoryReplyInbox();
private final MessageService messages = new MessageService(agents, new Injector(agents), rendezvous, inbox);
@BeforeEach
void setUp() {
// CB-520: the inbox only peeks/acks targets it owns.
inbox.own(T);
}
private static String textOf(McpSchema.CallToolResult r) {
return ((McpSchema.TextContent) r.content().getFirst()).text();
}
private static ClaudeCodeLauncher workerService(FakeHerdr h, String baseUrl, Set<String> allow) {
BridgedConfig.Worker cfg = new BridgedConfig.Worker(
BridgedConfig.Profile cfg = new BridgedConfig.Profile(
"ltms-local", baseUrl, "coder", null, "BRIDGED_WORKER_TOKEN", null,
"tab", "bridged-workers", "worker: {profile} #{n}", null, null, null);
return new ClaudeCodeLauncher(new AgentControl(h), new WorkspaceControl(h),
@@ -216,12 +229,15 @@ class BridgeMcpTest {
@Test
void spawnReturnsTheNewWorkersSessionAndPane() {
FakeHerdr h = new FakeHerdr();
McpSchema.CallToolResult res = BridgeMcp.spawn(
sessionManager(h, "http://gx00.gw:8000", Set.of("gx00.gw")), null);
SessionManager sm = sessionManager(h, "http://gx00.gw:8000", Set.of("gx00.gw"));
McpSchema.CallToolResult res = BridgeMcp.spawn(sm, null);
assertNotEquals(Boolean.TRUE, res.isError());
String out = textOf(res);
assertTrue(out.contains("\"sessionId\":\"term_new_1\""), out);
assertTrue(out.contains("\"paneId\":\"w9:pW_1\""), out);
// CB-519: the "paneId" wire field now carries the host-unique opaque id, not the herdr pane.
MemberSession s = sm.roster().getFirst();
assertTrue(out.contains("\"paneId\":\"" + s.paneId() + "\""), out);
assertNotEquals("w9:pRoot_1", s.paneId(), "the id is decoupled from the herdr pane coordinate");
assertTrue(out.contains("\"status\":\"spawning\""), out);
}
@@ -248,11 +264,12 @@ class BridgeMcpTest {
void spawnPassesTheRequestedCwdToTheWorker() {
FakeHerdr h = new FakeHerdr();
McpSchema.CallToolResult res = BridgeMcp.spawn(
sessionManager(h, "http://gx00.gw:8000", Set.of("gx00.gw")), null, "/req/dir", null, null, null);
sessionManager(h, "http://gx00.gw:8000", Set.of("gx00.gw")), null, null, "/req/dir", null, null, null);
assertNotEquals(Boolean.TRUE, res.isError());
// Protocol 19: the requested cwd roots the worker's pane at creation (tab.create).
@SuppressWarnings("unchecked")
Map<String, Object> start = (Map<String, Object>) h.lastCall("agent.start").params();
assertEquals("/req/dir", start.get("cwd"));
Map<String, Object> create = (Map<String, Object>) h.lastCall("tab.create").params();
assertEquals("/req/dir", create.get("cwd"));
}
@Test
@@ -270,10 +287,11 @@ class BridgeMcpTest {
FakeHerdr h = new FakeHerdr();
FakeWorktrees worktrees = new FakeWorktrees().withRepoRoot("/repo").withPrefix("/wt");
SessionManager sessions = new SessionManager(workerService(h, "http://gx00.gw:8000", Set.of("gx00.gw")), worktrees);
WorkerSession s = sessions.acquire("ltms-local", null, "/caller/proj", "term_primary",
MemberSession s = sessions.acquire("ltms-local", null, "/caller/proj", "term_primary",
new WorktreeRequest("cb-304", null));
McpSchema.CallToolResult res = BridgeMcp.listWorkers(workerService(h, "http://gx00.gw:8000", Set.of("gx00.gw")), sessions);
McpSchema.CallToolResult res = BridgeMcp.listFleet(
workerService(h, "http://gx00.gw:8000", Set.of("gx00.gw")), sessions, Map.of(), "");
assertNotEquals(Boolean.TRUE, res.isError());
String out = textOf(res);
@@ -287,6 +305,58 @@ class BridgeMcpTest {
assertTrue(out.contains("\"liveStatus\":\"unknown\""), out);
}
@Test
void listReportsLeadsAndFlagsTheCallersOwnRow() {
FakeHerdr h = new FakeHerdr();
SessionManager sessions = new SessionManager(
workerService(h, "http://gx00.gw:8000", Set.of("gx00.gw")));
McpSchema.CallToolResult res = BridgeMcp.listFleet(
workerService(h, "http://gx00.gw:8000", Set.of("gx00.gw")), sessions,
Map.of("term_me", "opus-5.0", "term_peer", "gpt-sol-5.6"), "term_me");
assertNotEquals(Boolean.TRUE, res.isError());
String out = textOf(res);
assertTrue(out.contains("\"name\":\"opus-5.0\""), out);
assertTrue(out.contains("\"name\":\"gpt-sol-5.6\""), out);
assertTrue(out.contains("\"sessionId\":\"term_peer\""), out);
// The caller's own row is flagged, and only the caller's — a peer must be distinguishable
// from self without a second bridge_whoami call.
assertEquals(1, out.split("\"self\":true", -1).length - 1, out);
assertTrue(out.indexOf("term_me") < out.indexOf("\"self\":true"), out);
}
@Test
void listReportsBothHalvesEvenWhenEmpty() {
FakeHerdr h = new FakeHerdr();
SessionManager sessions = new SessionManager(
workerService(h, "http://gx00.gw:8000", Set.of("gx00.gw")));
McpSchema.CallToolResult res = BridgeMcp.listFleet(
workerService(h, "http://gx00.gw:8000", Set.of("gx00.gw")), sessions, Map.of(), "");
// An absent "leads" key is what made an empty member roster read as "no peers" (CB-535).
String out = textOf(res);
assertTrue(out.contains("\"leads\":[]"), out);
assertTrue(out.contains("\"members\":[]"), out);
}
@Test
void listReportsALeadHerdrCannotSeeAsUnknown() {
FakeHerdr h = new FakeHerdr();
SessionManager sessions = new SessionManager(
workerService(h, "http://gx00.gw:8000", Set.of("gx00.gw")));
McpSchema.CallToolResult res = BridgeMcp.listFleet(
workerService(h, "http://gx00.gw:8000", Set.of("gx00.gw")), sessions,
Map.of("term_ghost", "gone-away"), "term_me");
// Reported, not hidden: an unreachable peer is exactly what a would-be sender needs to see.
String out = textOf(res);
assertTrue(out.contains("\"name\":\"gone-away\""), out);
assertTrue(out.contains("\"status\":\"unknown\""), out);
}
@Test
void stopTearsDownAWorkerByPane() {
FakeHerdr h = new FakeHerdr();
@@ -335,6 +405,30 @@ class BridgeMcpTest {
assertDoesNotThrow(() -> messages.ackReply("term_a", msgId));
}
@Test
void spawnedMembersAreMarkedPresent() {
Principal worker = Principal.worker("term_worker", 200);
Principal architect = Principal.architect("lead-designer", "term_design", 400);
MemberPresence presence = new MemberPresence();
BridgeMcp.markSpawnedMemberPresent(worker, presence);
BridgeMcp.markSpawnedMemberPresent(architect, presence);
assertTrue(presence.isPresent("term_worker"));
assertTrue(presence.isPresent("term_design"));
}
@Test
void nonMembersAreNotMarkedPresent() {
Principal lead = Principal.leader("opus", "term_lead", 100);
MemberPresence presence = new MemberPresence();
BridgeMcp.markSpawnedMemberPresent(lead, presence);
BridgeMcp.markSpawnedMemberPresent(Principal.anonymous(), presence);
assertFalse(presence.isPresent("term_lead"));
}
@Test
void statusReportsLiveAgentStatus() {
FakeHerdr blocked = new FakeHerdr().agentStatus("blocked");
@@ -366,7 +460,7 @@ class BridgeMcpTest {
FakeWorktrees worktrees = new FakeWorktrees().withRepoRoot("/repo").withPrefix("/wt");
SessionManager sessions = new SessionManager(
workerService(h, "http://gx00.gw:8000", Set.of("gx00.gw")), worktrees);
WorkerSession s = sessions.acquire("ltms-local", null, "/caller/proj", "term_primary",
MemberSession s = sessions.acquire("ltms-local", null, "/caller/proj", "term_primary",
new WorktreeRequest("cb-517", null));
McpSchema.CallToolResult res = BridgeMcp.whoami(Principal.worker(s.terminalId(), 200), sessions);
@@ -398,4 +492,124 @@ class BridgeMcpTest {
assertTrue(out.contains("\"sessionId\":\"term_orphan\""), out);
assertFalse(out.contains("profile"), out); // nothing invented for a session we don't track
}
/**
* CB-548: an architect reports its role and which gateway-local slot its pane is bound to —
* the same shape as a lead, under the architect key, so it can tell a peer where to reach it.
*/
@Test
void whoamiReportsAnArchitectWithItsSlotAndPane() {
FakeHerdr h = new FakeHerdr();
McpSchema.CallToolResult res = BridgeMcp.whoami(
Principal.architect("lead-designer", "term_design", 400),
sessionManager(h, "http://gx00.gw:8000", Set.of("gx00.gw")));
assertNotEquals(Boolean.TRUE, res.isError());
String out = textOf(res);
assertTrue(out.contains("\"role\":\"architect\""), out);
assertTrue(out.contains("\"architect\":\"lead-designer\""), out);
assertTrue(out.contains("\"sessionId\":\"term_design\""), out);
}
/**
* CB-548: an architect SEND delegates as its own pane (recording the per-target delegation) but
* must NEVER become the legacy singleton "primary" fallback — the per-target map does not cure
* the singleton, so an architect left there would draw no-delegation inbox nudges meant for a
* primary. Only PRIMARY callers (the unnamed primary and named leads alike) may claim it, and
* the decision keys on the resolved role, not name/kind sniffing.
*/
@Test
void architectSendDoesNotClaimThePrimarySingletonButALeadSendStillCan() {
// Architect SEND: does not change the legacy primary fallback.
PrimaryRegistry reg = new PrimaryRegistry(null);
BridgeMcp.recordPrimarySingleton(reg, "term_design", Principal.architect("design", "term_design", 400));
assertTrue(reg.primaryTerminal().isEmpty(),
"an architect must never become the legacy primary fallback");
// Lead SEND (a named PRIMARY) still claims it — preserved from CB-530/CB-532.
PrimaryRegistry leadReg = new PrimaryRegistry(null);
BridgeMcp.recordPrimarySingleton(leadReg, "term_lead_opus", Principal.leader("opus", "term_lead_opus", 100));
assertEquals("term_lead_opus", leadReg.primaryTerminal().orElseThrow(),
"a named lead is a primary and may claim the fallback");
// Unnamed primary likewise.
PrimaryRegistry primaryReg = new PrimaryRegistry(null);
BridgeMcp.recordPrimarySingleton(primaryReg, "term_p", Principal.primary(50));
assertEquals("term_p", primaryReg.primaryTerminal().orElseThrow(),
"an unnamed primary may claim the fallback");
// A null caller (legacy/no-auth path) records nothing.
PrimaryRegistry legacy = new PrimaryRegistry(null);
BridgeMcp.recordPrimarySingleton(legacy, "term_x", null);
assertTrue(legacy.primaryTerminal().isEmpty(), "no caller means nothing is recorded");
}
// --- member taxonomy (CB-557) ---------------------------------------------------------
@Test
void bridgeListReportsMembersNotWorkersAndCarriesEachRole() {
FakeHerdr h = new FakeHerdr();
SessionManager sessions = new SessionManager(workerService(h, "http://gx00.gw:8000", Set.of("gx00.gw")));
sessions.acquire("ltms-local", MemberRole.REVIEWER, null, "/caller/proj", "term_primary", null);
McpSchema.CallToolResult res = BridgeMcp.listFleet(
workerService(h, "http://gx00.gw:8000", Set.of("gx00.gw")), sessions, Map.of(), "");
String out = textOf(res);
assertTrue(out.contains("\"members\":"), "the roster half is named members: " + out);
assertFalse(out.contains("\"workers\":"), "the old key must be gone: " + out);
assertTrue(out.contains("\"role\":\"reviewer\""), out);
}
/**
* Role and profile are separate axes, so the roster has to report both. Two members on one
* backend may still be allowed to do entirely different things.
*/
@Test
void aRosterRowCarriesBothItsRoleAndItsProfile() {
FakeHerdr h = new FakeHerdr();
SessionManager sessions = new SessionManager(workerService(h, "http://gx00.gw:8000", Set.of("gx00.gw")));
sessions.acquire("ltms-local", MemberRole.DEV, null, "/caller/proj", "term_primary", null);
sessions.acquire("ltms-local", MemberRole.REVIEWER, null, "/caller/proj", "term_primary", null);
String out = textOf(BridgeMcp.listFleet(
workerService(h, "http://gx00.gw:8000", Set.of("gx00.gw")), sessions, Map.of(), ""));
assertTrue(out.contains("\"role\":\"dev\""), out);
assertTrue(out.contains("\"role\":\"reviewer\""), out);
assertEquals(2, out.split("\"profile\":\"ltms-local\"", -1).length - 1,
"both members share one profile — that is the point: " + out);
}
@Test
void spawnDefaultsTheRoleToDev() {
FakeHerdr h = new FakeHerdr();
McpSchema.CallToolResult res = BridgeMcp.spawn(
sessionManager(h, "http://gx00.gw:8000", Set.of("gx00.gw")), null, null, null, null, null, null);
assertNotEquals(Boolean.TRUE, res.isError());
assertTrue(textOf(res).contains("\"role\":\"dev\""), textOf(res));
}
@Test
void spawnAcceptsAnExplicitRole() {
FakeHerdr h = new FakeHerdr();
McpSchema.CallToolResult res = BridgeMcp.spawn(
sessionManager(h, "http://gx00.gw:8000", Set.of("gx00.gw")), null, "architect",
null, null, null, null);
assertNotEquals(Boolean.TRUE, res.isError());
assertTrue(textOf(res).contains("\"role\":\"architect\""), textOf(res));
}
@Test
void spawnRejectsAnUnknownRoleAndNamesTheValidOnes() {
FakeHerdr h = new FakeHerdr();
McpSchema.CallToolResult res = BridgeMcp.spawn(
sessionManager(h, "http://gx00.gw:8000", Set.of("gx00.gw")), null, "worker",
null, null, null, null);
assertEquals(Boolean.TRUE, res.isError());
assertTrue(textOf(res).contains("architect, dev, reviewer"), textOf(res));
}
}
@@ -105,4 +105,68 @@ class PrimaryRegistryTest {
reg.record("term_found");
assertEquals("term_found", reg.primaryTerminal().get());
}
// ── CB-532: nudges follow the delegating lead, not "the primary" ────────────────────────────
/**
* The bug that made `primary.terminal` unretirable: with one slot, whichever lead called
* bridge_send first captured every nudge — including nudges for the other lead's delegations.
*/
@Test
void aNudgeGoesToTheLeadThatDelegatedToThatWorker() {
var reg = new PrimaryRegistry(null);
reg.recordDelegation("term_worker_a", "term_lead_opus");
reg.recordDelegation("term_worker_b", "term_lead_sol");
assertEquals("term_lead_opus", reg.nudgeTargetFor("term_worker_a").orElseThrow());
assertEquals("term_lead_sol", reg.nudgeTargetFor("term_worker_b").orElseThrow());
}
@Test
void theMostRecentDelegatorWinsWhenAWorkerChangesHands() {
var reg = new PrimaryRegistry(null);
reg.recordDelegation("term_worker", "term_lead_opus");
reg.recordDelegation("term_worker", "term_lead_sol");
assertEquals("term_lead_sol", reg.nudgeTargetFor("term_worker").orElseThrow(),
"the lead waiting on the reply is the one that sent the work most recently");
}
/** After a restart the map is empty while the durable inbox still holds the reply. */
@Test
void anUnknownDelegationFallsBackToThePinnedPrimary() {
var reg = new PrimaryRegistry("term_pinned");
assertEquals("term_pinned", reg.nudgeTargetFor("term_never_seen").orElseThrow());
}
@Test
void withNoPinAndNoDelegationNobodyIsNudged() {
var reg = new PrimaryRegistry(null);
assertTrue(reg.nudgeTargetFor("term_worker").isEmpty(),
"guessing would interrupt the wrong lead with someone else's result; the durable "
+ "inbox makes pull the correct degradation");
}
@Test
void releasingASessionForgetsItsLead() {
var reg = new PrimaryRegistry(null);
reg.recordDelegation("term_worker", "term_lead");
reg.forgetDelegation("term_worker");
assertTrue(reg.nudgeTargetFor("term_worker").isEmpty());
}
@Test
void aBlankOrNullDelegationIsIgnoredRatherThanStored() {
var reg = new PrimaryRegistry(null);
reg.recordDelegation("term_worker", " ");
reg.recordDelegation(null, "term_lead");
reg.recordDelegation(" ", "term_lead");
assertTrue(reg.nudgeTargetFor("term_worker").isEmpty());
assertTrue(reg.nudgeTargetFor(null).isEmpty());
}
}
@@ -0,0 +1,881 @@
package dev.ltms.bridged.member;
import dev.ltms.bridged.config.BridgedConfig;
import dev.ltms.bridged.guard.GuardException;
import dev.ltms.bridged.guard.SubscriptionGuard;
import dev.ltms.bridged.herdr.AgentControl;
import dev.ltms.bridged.herdr.FakeHerdr;
import dev.ltms.bridged.herdr.WorkspaceControl;
import dev.ltms.bridged.peer.Capability;
import dev.ltms.bridged.peer.MemberRole;
import dev.ltms.bridged.peer.PeerHandle;
import dev.ltms.bridged.peer.PeerUnreachableException;
import dev.ltms.bridged.peer.SpawnRequest;
import org.junit.jupiter.api.Test;
import java.util.List;
import java.util.Map;
import java.util.Set;
import java.util.UUID;
import java.util.concurrent.atomic.AtomicReference;
import java.util.function.Function;
import java.util.function.Supplier;
import static org.junit.jupiter.api.Assertions.*;
/** The step-4 launch-flag injection: the bridge MCP + reply charter are appended to the argv. */
class ClaudeCodeLauncherTest {
private ClaudeCodeLauncher service(FakeHerdr herdr, List<String> argv, String mcpUrl) {
BridgedConfig.Profile cfg = new BridgedConfig.Profile(
"ltms-local", "http://gx00.gw:8000", "coder", null, "BRIDGED_WORKER_TOKEN",
argv, "tab", "bridged-workers", "worker: {profile} #{n}", mcpUrl, null, null);
return new ClaudeCodeLauncher(new AgentControl(herdr), new WorkspaceControl(herdr),
new SubscriptionGuard(Set.of("gx00.gw")), Map.of(cfg.profile(), cfg), cfg.profile(), _ -> null);
}
/** The {@code args} of the last agent.start — protocol 19: everything after the executable. */
@SuppressWarnings("unchecked")
private List<String> spawnedArgs(FakeHerdr herdr) {
return (List<String>) ((Map<String, Object>) herdr.lastCall("agent.start").params()).get("args");
}
@Test
void appendsBridgeMcpAndReplyCharterWhenMcpUrlSet() {
FakeHerdr herdr = new FakeHerdr();
service(herdr, List.of("claude"), "http://127.0.0.1:8765/mcp").spawn();
List<String> args = spawnedArgs(herdr);
assertTrue(args.contains("--mcp-config"));
assertTrue(args.stream().anyMatch(a -> a.contains("\"bridge\"") && a.contains("http://127.0.0.1:8765/mcp")),
"inline bridge MCP config present");
assertTrue(args.contains("--append-system-prompt"));
assertTrue(args.stream().anyMatch(a -> a.contains("bridge_reply")), "reply charter present");
}
@Test
void startRetriesWhileTheSeedShellBoots() {
// tab.create returns before the seed shell reaches its prompt; herdr refuses agent.start
// into a not-ready pane with agent_pane_busy. The launcher must wait it out, not fail.
FakeHerdr herdr = new FakeHerdr().agentPaneBusyTimes(2);
long[] clock = {0};
ClaudeCodeLauncher svc = new ClaudeCodeLauncher(
new AgentControl(herdr), new WorkspaceControl(herdr),
new SubscriptionGuard(Set.of("gx00.gw")),
Map.of("ltms-local", new BridgedConfig.Profile(
"ltms-local", "http://gx00.gw:8000", "coder", null, "BRIDGED_WORKER_TOKEN",
List.of("claude"), "tab", "bridged-workers", "w #{n}", null, null, null)),
"ltms-local", _ -> null,
0, () -> clock[0], () -> clock[0] += 50);
PeerHandle handle = svc.spawn(new SpawnRequest(null, null, null));
assertNotNull(handle, "spawn succeeds once the shell is ready");
assertEquals(3, herdr.calls.stream().filter(c -> c.method().equals("agent.start")).count(),
"two busy rejections, then the successful start");
}
@Test
void startResolvesTheExecutableFromKindAndDropsArgvZero() {
FakeHerdr herdr = new FakeHerdr();
service(herdr, List.of("claude"), null).spawn();
Map<?, ?> start = (Map<?, ?>) herdr.lastCall("agent.start").params();
assertEquals("claude", start.get("kind"), "herdr launches the canonical executable by kind");
// CB-533: the shared fixture pins model "coder", so the model flag is the whole args list.
// What this test guards is that argv[0] is NOT repeated — herdr supplies it from `kind`.
assertEquals(List.of("--model", "coder"), start.get("args"),
"the configured executable is not repeated in args");
}
@Test
void noBridgeFlagsWhenMcpUrlAbsent() {
FakeHerdr herdr = new FakeHerdr();
service(herdr, List.of("claude", "--verbose"), null).spawn();
List<String> args = spawnedArgs(herdr);
assertFalse(args.contains("--mcp-config"), "no bridge mount without mcpUrl");
assertFalse(args.contains("--append-system-prompt"), "no reply charter without mcpUrl");
// CB-533: the model flag is independent of the MCP mount — pinning the model is not part of
// "mount the bridge", so an unmounted worker still runs the model its profile names.
assertEquals(List.of("--verbose", "--model", "coder"), args,
"the operator's own args are preserved, in order, ahead of the model flag");
}
private ClaudeCodeLauncher multiProfile(FakeHerdr herdr) {
BridgedConfig.Profile gx10 = new BridgedConfig.Profile("gx10", "http://gx10.gw:8000", "coder",
null, "BRIDGED_WORKER_TOKEN", List.of("claude"), "tab", "bridged-workers", "w #{n}", null, null, null);
BridgedConfig.Profile ollama = new BridgedConfig.Profile("ollama", "http://ollama.ltms.dev", null,
null, "BRIDGED_WORKER_TOKEN", List.of("claude"), "tab", "bridged-workers", "w #{n}", null, null, null);
return new ClaudeCodeLauncher(new AgentControl(herdr), new WorkspaceControl(herdr),
new SubscriptionGuard(Set.of("gx10.gw", "ollama.ltms.dev")),
Map.of("gx10", gx10, "ollama", ollama), "gx10", _ -> "tok");
}
@Test
void spawnPicksTheNamedProfilesBaseUrl() {
FakeHerdr herdr = new FakeHerdr();
multiProfile(herdr).spawn("ollama");
assertEquals("http://ollama.ltms.dev", startEnv(herdr).get("ANTHROPIC_BASE_URL"),
"the named profile's base_url");
}
@Test
void spawnRejectsAnUnknownProfile() {
try (FakeHerdr herdr = new FakeHerdr()) {
assertThrows(IllegalArgumentException.class, () -> multiProfile(herdr).spawn("nope"));
}
}
@SuppressWarnings("unchecked")
private static String startCwd(FakeHerdr herdr) {
// Protocol 19: the worker's cwd is set at pane creation (tab.create), where the seed
// shell — which the agent starts into — is rooted.
Object v = ((Map<String, Object>) herdr.lastCall("tab.create").params()).get("cwd");
return v == null ? null : v.toString();
}
@Test
void requestedCwdRootsTheWorker() {
FakeHerdr herdr = new FakeHerdr();
service(herdr, List.of("ccs", "ltms-local"), null).spawn("ltms-local", "/work/proj", "/caller/home");
assertEquals("/work/proj", startCwd(herdr), "an explicit spawn cwd wins over everything");
}
@Test
void profileConfigCwdBeatsTheCallerCwd() {
FakeHerdr herdr = new FakeHerdr();
BridgedConfig.Profile cfg = new BridgedConfig.Profile("ltms-local", "http://gx00.gw:8000", "coder",
null, "BRIDGED_WORKER_TOKEN", List.of("ccs", "ltms-local"), "tab", "bridged-workers",
"w #{n}", null, "/pinned/dir", null);
ClaudeCodeLauncher svc = new ClaudeCodeLauncher(new AgentControl(herdr), new WorkspaceControl(herdr),
new SubscriptionGuard(Set.of("gx00.gw")), Map.of("ltms-local", cfg), "ltms-local", _ -> null);
svc.spawn("ltms-local", null, "/caller/home");
assertEquals("/pinned/dir", startCwd(herdr), "a profile-pinned cwd overrides the caller's");
}
@Test
void inheritsTheCallerCwdWhenNothingElseIsSet() {
FakeHerdr herdr = new FakeHerdr();
service(herdr, List.of("ccs", "ltms-local"), null).spawn("ltms-local", null, "/primary/project");
assertEquals("/primary/project", startCwd(herdr), "no explicit/config cwd → inherit the primary's");
}
// --- CB-302 git-forge token injection (worker checkpoint grant) ------------
/** Protocol 19: the worker's env is injected at pane creation (tab.create), not agent.start. */
@SuppressWarnings("unchecked")
private static Map<String, String> startEnv(FakeHerdr herdr) {
return (Map<String, String>) ((Map<String, Object>) herdr.lastCall("tab.create").params()).get("env");
}
@Test
void injectsForgeTokenAndHostWhenProfileGrantsIt() {
FakeHerdr herdr = new FakeHerdr();
BridgedConfig.Profile cfg = new BridgedConfig.Profile(
"impl", "http://gx00.gw:8000", "coder", null, "BRIDGED_WORKER_TOKEN",
List.of("ccs", "impl"), "tab", "bridged-workers", "w #{n}", null, null, null,
"GITEA_ACCESS_TOKEN", null); // parityOverlay null; gitHostEnv null → defaults to GITEA_HOST
Function<String, String> host = name -> switch (name) {
case "GITEA_ACCESS_TOKEN" -> "gt-secret";
case "GITEA_HOST" -> "git.ltms.dev";
default -> null;
};
new ClaudeCodeLauncher(new AgentControl(herdr), new WorkspaceControl(herdr),
new SubscriptionGuard(Set.of("gx00.gw")), Map.of("impl", cfg), "impl", host).spawn();
Map<String, String> env = startEnv(herdr);
assertEquals("gt-secret", env.get("GITEA_TOKEN"), "the forge token is injected for a granting profile");
assertEquals("git.ltms.dev", env.get("GITEA_HOST"), "the paired forge host rides along with the token");
}
@Test
void noForgeTokenWhenProfileDoesNotGrantIt() {
FakeHerdr herdr = new FakeHerdr();
// gitTokenEnv unset (12-arg ctor); the env would resolve a token if asked, proving the gate
// is the profile config, not a missing env var.
BridgedConfig.Profile cfg = new BridgedConfig.Profile("ltms-local", "http://gx00.gw:8000", "coder",
null, "BRIDGED_WORKER_TOKEN", List.of("ccs"), "tab", "bridged-workers", "w #{n}",
null, null, null);
new ClaudeCodeLauncher(new AgentControl(herdr), new WorkspaceControl(herdr),
new SubscriptionGuard(Set.of("gx00.gw")), Map.of("ltms-local", cfg), "ltms-local",
_ -> "would-be-secret").spawn();
Map<String, String> env = startEnv(herdr);
assertNull(env.get("GITEA_TOKEN"), "no forge token when the profile does not opt in");
assertNull(env.get("GITEA_HOST"), "no forge host without a granted token");
}
// --- CB-117 orphan reap: the pure predicate --------------------------------
@Test
void isForeignWorkerMatchesOurSchemeWithANonSelfNonce() {
assertTrue(ClaudeCodeLauncher.isForeignWorker("claude-ollama-be09c2-2", "aaaaaa"),
"a bridge worker name with a different nonce is a prior daemon's orphan");
assertTrue(ClaudeCodeLauncher.isForeignWorker("claude-gx10-4127af-11", "aaaaaa"),
"profile and multi-digit seq are still parsed; foreign nonce ⇒ reap");
}
@Test
void isForeignWorkerSparesOurOwnLiveWorkersAndNonWorkers() {
assertFalse(ClaudeCodeLauncher.isForeignWorker("claude-ollama-abcdef-3", "abcdef"),
"a worker with THIS process's nonce is ours and live — never reap it");
assertFalse(ClaudeCodeLauncher.isForeignWorker(null, "abcdef"), "an unnamed agent is not a worker");
assertFalse(ClaudeCodeLauncher.isForeignWorker("claude", "abcdef"), "a bare kind name is not a worker");
assertFalse(ClaudeCodeLauncher.isForeignWorker("my-repl", "abcdef"), "a user's own label is not a worker");
assertFalse(ClaudeCodeLauncher.isForeignWorker("claude-ollama-XYZ123-2", "abcdef"),
"a non-hex nonce does not match our scheme");
}
// --- CB-117 orphan reap: the wiring through stop() -------------------------
private static long paneCloseCount(FakeHerdr herdr, String paneId) {
return herdr.calls.stream()
.filter(c -> c.method().equals("pane.close"))
.filter(c -> paneId.equals(((Map<?, ?>) c.params()).get("pane_id")))
.count();
}
@Test
void reapsAForeignOrphanButSparesOurOwnWorkerAndUserSessions() {
FakeHerdr herdr = new FakeHerdr();
ClaudeCodeLauncher svc = multiProfile(herdr);
herdr.withAgent("claude-ollama-be09c2-2", "term_orphan", "wQ:pF", "wQ:t8") // prior daemon's leak
.withAgent("claude-gx10-" + svc.nameNonce() + "-1", "term_mine", "wQ:pMine", "wQ:tMine"); // ours, live
// (the fake's default unnamed term_a stands in for a user's own Claude session)
int reaped = svc.reapOrphanWorkers();
assertEquals(1, reaped, "exactly the one foreign-nonce orphan is reaped");
assertEquals(1, paneCloseCount(herdr, "wQ:pF"), "the orphan's pane is closed");
assertEquals(0, paneCloseCount(herdr, "wQ:pMine"), "our own live worker's pane is left running");
assertEquals(0, paneCloseCount(herdr, "w2:p7"), "a user's own session is never touched");
assertTrue(herdr.called("tab.close"), "the orphan's now-empty dedicated tab is closed too");
}
@Test
void reapCountsAnAlreadyGoneOrphanAsReaped() {
FakeHerdr herdr = new FakeHerdr().paneCloseFailsWith("pane_not_found");
ClaudeCodeLauncher svc = multiProfile(herdr);
herdr.withAgent("claude-ollama-0d856d-3", "term_gone", "wQ:pS", "wQ:tD");
assertEquals(1, svc.reapOrphanWorkers(),
"a pane that vanished between list and close is a successful reap, not a failure");
}
@Test
void reapIsSkippedWhenHerdrCannotBeListed() {
FakeHerdr herdr = new FakeHerdr().healthy(false); // agent.list throws
assertEquals(0, multiProfile(herdr).reapOrphanWorkers(), "a listing failure reaps nothing and does not throw");
}
// --- PeerHandle indirection ----------------------------------------------------------------
@Test
void spawnReturnsPeerHandleWithHostUniqueOpaqueId() {
FakeHerdr herdr = new FakeHerdr();
ClaudeCodeLauncher svc = service(herdr, List.of("ccs", "ltms-local"), null);
PeerHandle handle = svc.spawn(new SpawnRequest(null, null, null));
assertNotNull(handle, "spawn must return a non-null handle");
// CB-519: id() is a host-unique opaque UUID, decoupled from the herdr pane coordinate.
assertNotEquals("w9:pRoot_1", handle.id(),
"handle.id() must NOT be the herdr pane id");
assertDoesNotThrow(() -> UUID.fromString(handle.id()),
"handle.id() must be a UUID: " + handle.id());
}
@Test
void spawnReturnsPeerHandleWithCorrectTerminalId() {
FakeHerdr herdr = new FakeHerdr();
ClaudeCodeLauncher svc = service(herdr, List.of("ccs", "ltms-local"), null);
PeerHandle handle = svc.spawn(new SpawnRequest("ltms-local", null, "/caller"));
assertEquals("term_new_1", handle.terminalId(), "handle.terminalId() must equal the agent's terminalId");
}
@Test
void capabilitiesIncludeMidTurnAskWorktreeOrphanReap() {
FakeHerdr herdr = new FakeHerdr();
ClaudeCodeLauncher svc = service(herdr, List.of("ccs", "ltms-local"), null);
Set<Capability> caps = svc.capabilities();
assertTrue(caps.contains(Capability.MID_TURN_ASK), "every Claude Code peer supports mid-turn ask");
assertTrue(caps.contains(Capability.WORKTREE), "every CLI peer supports worktree cwd");
assertTrue(caps.contains(Capability.ORPHAN_REAP), "every herdr launcher supports orphan reap");
assertTrue(caps.contains(Capability.CONTEXT_RESET), "Claude Code supports /clear");
}
@Test
void clearContextUsesTheClaudeCommandThroughTheOwningHandle() {
FakeHerdr herdr = new FakeHerdr();
ClaudeCodeLauncher svc = service(herdr, List.of("claude"), null);
PeerHandle handle = svc.spawn(new SpawnRequest(null, null, null));
assertTrue(svc.clearContext(handle.id()));
Map<?, ?> prompt = (Map<?, ?>) herdr.lastCall("agent.prompt").params();
assertEquals("/clear", prompt.get("text"));
assertEquals("w9:pRoot_1", prompt.get("target"));
}
@Test
void capabilitiesIncludeSelfPrWhenProfileHasGitToken() {
FakeHerdr herdr = new FakeHerdr();
BridgedConfig.Profile cfg = new BridgedConfig.Profile(
"impl", "http://gx00.gw:8000", "coder", null, "BRIDGED_WORKER_TOKEN",
List.of("ccs", "impl"), "tab", "bridged-workers", "w #{n}", null, null, null,
"GITEA_ACCESS_TOKEN", null);
ClaudeCodeLauncher svc = new ClaudeCodeLauncher(new AgentControl(herdr), new WorkspaceControl(herdr),
new SubscriptionGuard(Set.of("gx00.gw")), Map.of("impl", cfg), "impl",
_ -> "tok");
assertTrue(svc.capabilities().contains(Capability.SELF_PR),
"a profile with a git token grants SELF_PR");
}
@Test
void capabilitiesExcludeSelfPrWhenNoGitToken() {
FakeHerdr herdr = new FakeHerdr();
ClaudeCodeLauncher svc = service(herdr, List.of("ccs", "ltms-local"), null);
assertFalse(svc.capabilities().contains(Capability.SELF_PR),
"no git token profile → no SELF_PR capability");
}
@Test
void effectiveCwdViaSpawnRequestMatchesExistingResolution() {
FakeHerdr herdr = new FakeHerdr();
ClaudeCodeLauncher svc = service(herdr, List.of("ccs", "ltms-local"), null);
String cwd = svc.effectiveCwd(new SpawnRequest("ltms-local", "/work/proj", "/caller/home"));
assertEquals("/work/proj", cwd, "effectiveCwd via SpawnRequest must match the three-arg resolution");
}
// --- CB-547a: durable session identity (mint / resume / no-identity legacy) -----------------
@Test
void freshSpawnMintsASessionIdAndPassesTheName() {
FakeHerdr herdr = new FakeHerdr();
ClaudeCodeLauncher svc = service(herdr, List.of("ccs", "ltms-local"), null);
PeerHandle handle = svc.spawn(new SpawnRequest("ltms-local", null, null, "my-session", null));
List<String> args = spawnedArgs(herdr);
int flag = args.indexOf("--session-id");
assertTrue(flag >= 0 && flag + 1 < args.size(), "--session-id present: " + args);
String minted = args.get(flag + 1);
assertDoesNotThrow(() -> UUID.fromString(minted), "--session-id is a valid UUID: " + minted);
assertEquals("my-session", args.get(args.indexOf("-n") + 1), "the logical name rides as -n");
assertEquals(minted, handle.agentSessionId(),
"the resume handle is the minted id, known before the agent has written anything");
assertEquals("my-session", handle.sessionName(), "the handle carries the logical name");
}
@Test
void resumeSpawnPassesDashRAndNeverASessionId() {
FakeHerdr herdr = new FakeHerdr();
ClaudeCodeLauncher svc = service(herdr, List.of("ccs", "ltms-local"), null);
PeerHandle handle = svc.spawn(new SpawnRequest("ltms-local", null, null, "my-session", "cb-resume-1"));
List<String> args = spawnedArgs(herdr);
assertFalse(args.contains("--session-id"), "--session-id must NOT be passed on a resume (conflicts with -r)");
assertEquals("cb-resume-1", args.get(args.indexOf("-r") + 1), "-r carries the prior session id");
assertEquals("cb-resume-1", handle.agentSessionId(), "a resume adopts the prior id as its own");
assertEquals("my-session", handle.sessionName(), "the logical name survives a resume");
}
@Test
void noIdentitySpawnKeepsTheLegacyArgvAndCarriesNoSessionHandle() {
FakeHerdr herdr = new FakeHerdr();
ClaudeCodeLauncher svc = service(herdr, List.of("ccs", "ltms-local"), null);
PeerHandle handle = svc.spawn(new SpawnRequest("ltms-local", null, null));
List<String> args = spawnedArgs(herdr);
assertFalse(args.contains("--session-id"), "no identity → no --session-id");
assertFalse(args.contains("-n"), "no identity → no -n");
assertFalse(args.contains("-r"), "no identity → no -r");
assertNull(handle.agentSessionId(), "no identity → no resume handle");
assertNull(handle.sessionName(), "no identity → no logical name");
}
@Test
void capabilitiesIncludeSessionNameAndSessionResume() {
FakeHerdr herdr = new FakeHerdr();
ClaudeCodeLauncher svc = service(herdr, List.of("ccs", "ltms-local"), null);
Set<Capability> caps = svc.capabilities();
assertTrue(caps.contains(Capability.SESSION_NAME), "Claude Code surfaces the bridge's logical name (-n)");
assertTrue(caps.contains(Capability.SESSION_RESUME), "Claude Code can relaunch onto a prior conversation (-r)");
}
@Test
void profilesViaPeerLauncherMatchesExistingApi() {
FakeHerdr herdr = new FakeHerdr();
ClaudeCodeLauncher svc = multiProfile(herdr);
assertEquals(Set.of("gx10", "ollama"), svc.profiles(), "profiles() via PeerLauncher must match");
}
@Test
void defaultProfileViaPeerLauncherMatches() {
FakeHerdr herdr = new FakeHerdr();
ClaudeCodeLauncher svc = multiProfile(herdr);
assertEquals("gx10", svc.defaultProfile(), "defaultProfile() via PeerLauncher must match");
}
@Test
void stopViaPeerLauncherTearsDownByHandleId() {
FakeHerdr herdr = new FakeHerdr();
ClaudeCodeLauncher svc = service(herdr, List.of("ccs", "ltms-local"), null);
PeerHandle handle = svc.spawn(new SpawnRequest(null, null, null));
svc.stop(handle.id());
assertTrue(herdr.called("pane.close"), "stop via handle.id() must close the pane");
}
// --- CB-519: host-unique id, decoupled from the pane coordinate ------------------------------
@Test
void twoSpawnsOnTheSamePaneNeverCollideOnHostUniqueId() {
// Two spawns may be placed on the same herdr pane coordinate (e.g. a pane that was reused
// or re-reported after a restart); the host-unique id must not collide even then.
FakeHerdr herdr = new FakeHerdr().pinNextStarts(2, "term_shared", "w9:pShared");
ClaudeCodeLauncher svc = service(herdr, List.of("ccs", "ltms-local"), null);
PeerHandle a = svc.spawn(new SpawnRequest(null, null, null));
PeerHandle b = svc.spawn(new SpawnRequest(null, null, null));
assertNotEquals(a.id(), b.id(),
"two spawns on the same pane coordinate get distinct host-unique ids");
assertNotEquals("w9:pShared", a.id(), "id is not the pane coordinate");
assertNotEquals("w9:pShared", b.id(), "id is not the pane coordinate");
}
@Test
void stopResolvesTheHostUniqueIdToThePaneThatSpawnedIt() {
// CB-519: id() != paneId, so stop(id) must tear down the exact pane the id names — and no
// other live peer's pane.
FakeHerdr herdr = new FakeHerdr(); // deterministic panes w9:pRoot_1, w9:pRoot_2 per spawn
ClaudeCodeLauncher svc = service(herdr, List.of("ccs", "ltms-local"), null);
PeerHandle a = svc.spawn(new SpawnRequest(null, null, null));
PeerHandle b = svc.spawn(new SpawnRequest(null, null, null));
svc.stop(b.id());
assertEquals(1, paneCloseCount(herdr, "w9:pRoot_2"), "stop(b.id()) closes only b's pane");
assertEquals(0, paneCloseCount(herdr, "w9:pRoot_1"), "a's pane is untouched");
}
// --- CB-306 spawn-readiness gate -----------------------------------------------------------
private static Map<String, BridgedConfig.Profile> workerConfigMap(String profile, String mcpUrl) {
BridgedConfig.Profile cfg = new BridgedConfig.Profile(
profile, "http://gx00.gw:8000", "coder", null, "BRIDGED_WORKER_TOKEN",
List.of("ccs", profile), "tab", "bridged-workers",
"worker: {profile} #{n}", mcpUrl, null, null);
return Map.of(cfg.profile(), cfg);
}
@Test
void spawnWaitsUntilInjectableThenReturnsHandle() {
FakeHerdr herdr = new FakeHerdr();
herdr.agentStatus("unknown"); // first status call sees UNKNOWN
long[] clock = {0};
boolean[] firstSleep = {true};
// The sleeper: advance the fake clock, and on the first call flip the
// agent status to IDLE so the next poll succeeds.
Runnable sleeper = () -> {
clock[0] += 300;
if (firstSleep[0]) {
herdr.agentStatus("idle");
firstSleep[0] = false;
}
};
ClaudeCodeLauncher svc = new ClaudeCodeLauncher(
new AgentControl(herdr), new WorkspaceControl(herdr),
new SubscriptionGuard(Set.of("gx00.gw")),
workerConfigMap("ltms-local", null), "ltms-local", _ -> null,
5000, () -> clock[0], sleeper);
PeerHandle handle = svc.spawn(new SpawnRequest(null, null, null));
assertNotNull(handle, "spawn returns a handle when worker becomes injectable");
assertNotEquals("w9:pRoot_1", handle.id(),
"handle id is a host-unique opaque id, not the started pane");
assertEquals(0, paneCloseCount(herdr, "w9:pRoot_1"),
"no pane.close when worker becomes injectable before timeout");
}
@Test
void spawnThrowsPeerUnreachableWhenNeverInjectableAndReapsPane() {
FakeHerdr herdr = new FakeHerdr();
herdr.agentStatus("unknown"); // always UNKNOWN
long[] clock = {0};
ClaudeCodeLauncher svc = new ClaudeCodeLauncher(
new AgentControl(herdr), new WorkspaceControl(herdr),
new SubscriptionGuard(Set.of("gx00.gw")),
workerConfigMap("ltms-local", null), "ltms-local", _ -> null,
1000, () -> clock[0], () -> clock[0] += 50);
PeerUnreachableException ex = assertThrows(
PeerUnreachableException.class,
() -> svc.spawn(new SpawnRequest(null, null, null)));
assertTrue(ex.getMessage().contains("w9:pRoot_1"),
"exception message references the paneId: " + ex.getMessage());
assertTrue(ex.getMessage().contains("1000"),
"exception message references the timeout: " + ex.getMessage());
assertTrue(clock[0] >= 1000, "fake clock advanced past the timeout: " + clock[0]);
assertEquals(1, paneCloseCount(herdr, "w9:pRoot_1"),
"pane was closed on timeout (no orphan left behind)");
}
@Test
void spawnReturnsImmediatelyWhenGateIsDisabled() {
FakeHerdr herdr = new FakeHerdr();
// The default 6-arg constructor has spawnReadyTimeoutMs=0 (gate disabled).
ClaudeCodeLauncher svc = service(herdr, List.of("ccs", "ltms-local"), null);
PeerHandle handle = svc.spawn(new SpawnRequest(null, null, null));
assertNotNull(handle, "spawn returns a handle when the gate is disabled");
assertFalse(herdr.called("agent.get"),
"agent.get is never called when the gate is disabled (no polling)");
}
@Test
void spawnGateRespectsZeroTimeoutEvenWithFullConstructor() {
FakeHerdr herdr = new FakeHerdr();
long[] clock = {0};
// Explicit zero timeout with the full testability constructor — should
// skip polling entirely, just like the legacy default path.
ClaudeCodeLauncher svc = new ClaudeCodeLauncher(
new AgentControl(herdr), new WorkspaceControl(herdr),
new SubscriptionGuard(Set.of("gx00.gw")),
workerConfigMap("ltms-local", null), "ltms-local", _ -> null,
0, () -> clock[0], () -> clock[0] += 1);
PeerHandle handle = svc.spawn(new SpawnRequest(null, null, null));
assertNotNull(handle, "spawn still succeeds with zero timeout");
assertEquals(0, paneCloseCount(herdr, "w9:pRoot_1"),
"no orphan pane close from the gate path");
assertDoesNotThrow(() -> UUID.fromString(handle.id()));
}
// --- CB-511: worker environment seeding -----------------------------------------------------
@Test
void workerInheritsTheDaemonPath() {
FakeHerdr herdr = new FakeHerdr();
BridgedConfig.Profile cfg = new BridgedConfig.Profile(
"ltms-local", "http://gx00.gw:8000", "coder", null, "BRIDGED_WORKER_TOKEN",
List.of("claude"), "tab", "bridged-workers", "w #{n}", null, null, null);
new ClaudeCodeLauncher(new AgentControl(herdr), new WorkspaceControl(herdr),
new SubscriptionGuard(Set.of("gx00.gw")), Map.of(cfg.profile(), cfg), cfg.profile(),
k -> "PATH".equals(k) ? "/opt/tools/bin:/usr/bin" : null).spawn();
assertEquals("/opt/tools/bin:/usr/bin", startEnv(herdr).get("PATH"),
"a worker with no PATH cannot run the build it is asked to run");
}
@Test
void profileEnvIsInjectedIntoTheWorker() {
FakeHerdr herdr = new FakeHerdr();
BridgedConfig.Profile cfg = new BridgedConfig.Profile(
"ltms-local", "http://gx00.gw:8000", "coder", null, "BRIDGED_WORKER_TOKEN",
List.of("claude"), "tab", "bridged-workers", "w #{n}", null, null, null, null, null,
null, Map.of("JAVA_HOME", "/opt/jdk", "PATH", "/profile/bin"), null, null);
new ClaudeCodeLauncher(new AgentControl(herdr), new WorkspaceControl(herdr),
new SubscriptionGuard(Set.of("gx00.gw")), Map.of(cfg.profile(), cfg), cfg.profile(),
k -> "PATH".equals(k) ? "/daemon/bin" : null).spawn();
Map<String, String> env = startEnv(herdr);
assertEquals("/opt/jdk", env.get("JAVA_HOME"), "profile env: is passed through");
assertEquals("/profile/bin", env.get("PATH"), "an explicit profile PATH overrides the daemon's");
}
/**
* The security-relevant ordering. {@code SubscriptionGuard} is checked against the profile's
* {@code baseUrl} only, so if a profile's {@code env:} could overwrite ANTHROPIC_BASE_URL a
* worker could be pointed at an unguarded host while the guard passed on a benign one.
*/
@Test
void profileEnvCannotOverrideGuardCheckedAnthropicVars() {
FakeHerdr herdr = new FakeHerdr();
BridgedConfig.Profile cfg = new BridgedConfig.Profile(
"ltms-local", "http://gx00.gw:8000", "coder", null, "BRIDGED_WORKER_TOKEN",
List.of("claude"), "tab", "bridged-workers", "w #{n}", null, null, null, null, null,
null, Map.of("ANTHROPIC_BASE_URL", "http://evil.example.com"), null, null);
new ClaudeCodeLauncher(new AgentControl(herdr), new WorkspaceControl(herdr),
new SubscriptionGuard(Set.of("gx00.gw")), Map.of(cfg.profile(), cfg), cfg.profile(),
_ -> null).spawn();
assertEquals("http://gx00.gw:8000", startEnv(herdr).get("ANTHROPIC_BASE_URL"),
"the guard-checked baseUrl must win over any env: entry, or the boundary is bypassable");
}
// ── CB-533: the model is pinned on the command line, not only in the environment ────────────
/** A launcher for a profile identical but for its {@code model:} — the only variable here. */
private ClaudeCodeLauncher serviceWithModel(FakeHerdr herdr, String model) {
BridgedConfig.Profile cfg = profileWithModel(model);
return new ClaudeCodeLauncher(new AgentControl(herdr), new WorkspaceControl(herdr),
new SubscriptionGuard(Set.of("gx00.gw")), Map.of(cfg.profile(), cfg), cfg.profile(),
_ -> null);
}
private static BridgedConfig.Profile profileWithModel(String model) {
return new BridgedConfig.Profile("sonnet", "http://gx00.gw:8000", model, null,
"BRIDGED_WORKER_TOKEN", List.of("ccs", "sonnet"), "tab", "bridged-workers",
"w #{n}", "http://127.0.0.1:8765/mcp", null, null);
}
@Test
void aConfiguredModelIsPassedAsAModelFlagAsWellAsTheEnvVar() {
// ANTHROPIC_MODEL alone loses to `ccs`, which exports its own model family over whatever it
// inherited — so a profile that set model: was silently overruled by its own launcher.
FakeHerdr herdr = new FakeHerdr();
serviceWithModel(herdr, "claude-sonnet-5").spawn("sonnet", null, null);
assertEquals("claude-sonnet-5", startEnv(herdr).get("ANTHROPIC_MODEL"));
List<String> args = spawnedArgs(herdr);
int flag = args.indexOf("--model");
assertTrue(flag >= 0, "the flag is what survives a wrapper argv like [ccs, sonnet]");
assertEquals("claude-sonnet-5", args.get(flag + 1));
}
@Test
void theModelFlagComesLastSoItOutranksTheOperatorsOwnArgv() {
FakeHerdr herdr = new FakeHerdr();
serviceWithModel(herdr, "claude-sonnet-5").spawn("sonnet", null, null);
List<String> args = spawnedArgs(herdr);
assertEquals(args.size() - 2, args.indexOf("--model"));
}
@Test
void aProfileWithNoModelGetsNoModelFlag() {
// gx10 deliberately leaves model: unset so ccs owns selection; adding a flag would make
// this file a second source of truth for exactly the thing it declines to decide.
FakeHerdr herdr = new FakeHerdr();
serviceWithModel(herdr, null).spawn("sonnet", null, null);
assertFalse(spawnedArgs(herdr).contains("--model"));
assertNull(startEnv(herdr).get("ANTHROPIC_MODEL"));
}
// --- CB-539: subscription-profile opt-in ----------------------------------------------------
/** A claude-code profile on the subscription: no baseUrl (by design), no off-sub endpoint. */
private static BridgedConfig.Profile subscriptionCfg(String profile, String baseUrl) {
return new BridgedConfig.Profile(
profile, baseUrl, "sonnet", null, "BRIDGED_WORKER_TOKEN",
List.of("ccs", profile), "tab", "bridged-workers", "w #{n}", null, null, null,
null, null, null, Map.of(), null, null, true);
}
@Test
void defaultRefusalIsPreservedForClaudeProfileWithNoBaseUrl() {
// Requirement 1: absent subscription:true ⇒ byte-identical refusal to today. A claude-code
// profile with no baseUrl and no subscription must still be refused (it would bill the sub).
FakeHerdr herdr = new FakeHerdr();
BridgedConfig.Profile cfg = new BridgedConfig.Profile(
"ltms-local", null, "coder", null, "BRIDGED_WORKER_TOKEN",
List.of("ccs", "ltms-local"), "tab", "bridged-workers", "w #{n}", null, null, null);
ClaudeCodeLauncher svc = new ClaudeCodeLauncher(new AgentControl(herdr), new WorkspaceControl(herdr),
new SubscriptionGuard(Set.of("gx00.gw")), Map.of(cfg.profile(), cfg), cfg.profile(), _ -> null);
GuardException ex = assertThrows(GuardException.class, () -> svc.spawn("ltms-local", null, null));
assertTrue(ex.getMessage().contains("no ANTHROPIC_BASE_URL"),
"the refusal names the missing baseUrl: " + ex.getMessage());
assertEquals(0, herdr.calls.stream().filter(c -> c.method().equals("agent.start")).count(),
"nothing was spawned before the refusal");
}
@Test
void subscriptionProfileSpawnsWithoutInjectedAnthropicVars() {
// Requirement on subscription:true: no baseUrl is required (or injected), and neither
// ANTHROPIC_BASE_URL nor ANTHROPIC_AUTH_TOKEN is injected even though the token env would
// resolve one if asked.
FakeHerdr herdr = new FakeHerdr();
BridgedConfig.Profile cfg = subscriptionCfg("sonnet", null);
new ClaudeCodeLauncher(new AgentControl(herdr), new WorkspaceControl(herdr),
new SubscriptionGuard(Set.of("gx00.gw")), Map.of(cfg.profile(), cfg), cfg.profile(),
_ -> "would-be-token").spawn();
Map<String, String> env = startEnv(herdr);
assertNull(env.get("ANTHROPIC_BASE_URL"), "no baseUrl injected for a subscription profile");
assertNull(env.get("ANTHROPIC_AUTH_TOKEN"), "no auth token injected for a subscription profile");
assertEquals("sonnet", env.get("ANTHROPIC_MODEL"),
"the model alias is still injected; only the subscription-boundary vars are dropped");
}
@Test
void subscriptionPlusBaseUrlIsRefused() {
// Requirement 2: subscription:true + a baseUrl state opposite intents — refuse at spawn,
// naming the profile, rather than silently picking a winner.
FakeHerdr herdr = new FakeHerdr();
BridgedConfig.Profile cfg = subscriptionCfg("sonnet", "http://gx00.gw:8000");
ClaudeCodeLauncher svc = new ClaudeCodeLauncher(new AgentControl(herdr), new WorkspaceControl(herdr),
new SubscriptionGuard(Set.of("gx00.gw")), Map.of(cfg.profile(), cfg), cfg.profile(), _ -> null);
IllegalStateException ex = assertThrows(IllegalStateException.class,
() -> svc.spawn("sonnet", null, null));
assertTrue(ex.getMessage().contains("sonnet"), "refusal names the profile: " + ex.getMessage());
assertTrue(ex.getMessage().contains("subscription"), "refusal explains the contradiction: " + ex.getMessage());
assertEquals(0, herdr.calls.stream().filter(c -> c.method().equals("agent.start")).count(),
"nothing was spawned before the contradiction was refused");
}
@Test
void nonSubscriptionProfilesAreStillAllowlistChecked() {
// Requirement 3: the guard keeps its teeth for every other profile — a base_url whose host is
// not on the allowlist is still refused, whether or not any subscription profile exists.
FakeHerdr herdr = new FakeHerdr();
BridgedConfig.Profile rogue = new BridgedConfig.Profile(
"rogue", "http://evil.example.com:8000", "coder", null, "BRIDGED_WORKER_TOKEN",
List.of("ccs", "rogue"), "tab", "bridged-workers", "w #{n}", null, null, null);
ClaudeCodeLauncher svc = new ClaudeCodeLauncher(new AgentControl(herdr), new WorkspaceControl(herdr),
new SubscriptionGuard(Set.of("gx00.gw")), Map.of(rogue.profile(), rogue), rogue.profile(), _ -> null);
GuardException ex = assertThrows(GuardException.class, () -> svc.spawn("rogue", null, null));
assertTrue(ex.getMessage().contains("not on the"), "refusal cites the allowlist: " + ex.getMessage());
assertEquals(0, herdr.calls.stream().filter(c -> c.method().equals("agent.start")).count(),
"nothing was spawned before the allowlist refusal");
}
@Test
void aSubscriptionProfileHasAnEnvSuppliedAnthropicBindingStripped() {
// CB-542: even a subscription profile whose env: carries ANTHROPIC_BASE_URL (or AUTH_TOKEN)
// must not hand them to the worker — on the subscription path no guard would vet them. Config
// load refuses this loudly; this launcher-side strip is the belt-and-braces that makes the
// invariant hold for a profile built in code that never passed through that validation.
FakeHerdr herdr = new FakeHerdr();
BridgedConfig.Profile cfg = new BridgedConfig.Profile(
"sonnet", null, "sonnet", null, "BRIDGED_WORKER_TOKEN",
List.of("ccs", "sonnet"), "tab", "bridged-workers", "w #{n}", null, null, null,
null, null, null,
Map.of("ANTHROPIC_BASE_URL", "http://evil.example.com",
"ANTHROPIC_AUTH_TOKEN", "sk-ant-bad", "JAVA_HOME", "/opt/jdk"),
null, null, true);
new ClaudeCodeLauncher(new AgentControl(herdr), new WorkspaceControl(herdr),
new SubscriptionGuard(Set.of("gx00.gw")), Map.of(cfg.profile(), cfg), cfg.profile(),
_ -> "would-be-token").spawn();
Map<String, String> env = startEnv(herdr);
assertNull(env.get("ANTHROPIC_BASE_URL"),
"the unguarded endpoint must not survive into the worker");
assertNull(env.get("ANTHROPIC_AUTH_TOKEN"),
"the unguarded token must not survive into the worker");
assertEquals("/opt/jdk", env.get("JAVA_HOME"),
"only the Anthropic binding keys are stripped; the rest of env: still applies");
}
// ── CB-557: role-aware tab labels ─────────────────────────────────────────────────────────
/** The {@code label} of every {@code tab.rename}, in call order. */
private List<String> tabLabels(FakeHerdr herdr) {
return herdr.calls.stream()
.filter(c -> c.method().equals("tab.rename"))
.map(c -> (String) ((Map<?, ?>) c.params()).get("label"))
.toList();
}
/** A profile with no {@code tabLabel:} of its own — the fleet template decides. */
private ClaudeCodeLauncher labelService(FakeHerdr herdr, Supplier<String> fleetTemplate) {
BridgedConfig.Profile cfg = new BridgedConfig.Profile(
"sonnet", "http://gx00.gw:8000", "sonnet", null, "BRIDGED_WORKER_TOKEN",
List.of("claude"), "tab", "bridged-workers", null, null, null, null);
return new ClaudeCodeLauncher(new AgentControl(herdr), new WorkspaceControl(herdr),
new SubscriptionGuard(Set.of("gx00.gw")), Map.of(cfg.profile(), cfg), cfg.profile(),
_ -> null, 0, 0L, fleetTemplate);
}
/**
* The knob must reach the rename call. It was inert once — {@code HerdrPeerLauncher} accepted a
* template while {@code Bridged} passed none, and the label stayed right only because the
* fallback happened to match. Pin the wiring, not the coincidence.
*/
@Test
void theFleetTemplateNamesTheRoleTheMemberWasSpawnedFor() {
FakeHerdr herdr = new FakeHerdr();
ClaudeCodeLauncher svc = labelService(herdr, () -> "{role}: {profile} #{n}");
svc.spawn(new SpawnRequest("sonnet", null, null, null, null, MemberRole.REVIEWER));
assertEquals(List.of("reviewer: sonnet #1"), tabLabels(herdr));
}
/** The counter is per role+profile, so a dev and a reviewer on one profile both start at #1. */
@Test
void theCounterRunsPerRoleAndProfileNotPerFleet() {
FakeHerdr herdr = new FakeHerdr();
ClaudeCodeLauncher svc = labelService(herdr, () -> "{role}: {profile} #{n}");
svc.spawn(new SpawnRequest("sonnet", null, null, null, null, MemberRole.DEV));
svc.spawn(new SpawnRequest("sonnet", null, null, null, null, MemberRole.REVIEWER));
svc.spawn(new SpawnRequest("sonnet", null, null, null, null, MemberRole.DEV));
assertEquals(List.of("dev: sonnet #1", "reviewer: sonnet #1", "dev: sonnet #2"),
tabLabels(herdr));
}
/** No fleet template configured ⇒ the built-in default, still role-first. */
@Test
void aBlankFleetTemplateFallsBackToTheRoleFirstDefault() {
FakeHerdr herdr = new FakeHerdr();
labelService(herdr, () -> null).spawn(
new SpawnRequest("sonnet", null, null, null, null, MemberRole.ARCHITECT));
assertEquals(List.of("architect: sonnet #1"), tabLabels(herdr));
assertEquals("{role}: {profile} #{n}", BridgedConfig.Fleet.DEFAULT_TAB_LABEL);
}
/** A profile that wants its own label still outranks the fleet template. */
@Test
void aProfileTabLabelOverridesTheFleetTemplate() {
FakeHerdr herdr = new FakeHerdr();
BridgedConfig.Profile cfg = new BridgedConfig.Profile(
"sonnet", "http://gx00.gw:8000", "sonnet", null, "BRIDGED_WORKER_TOKEN",
List.of("claude"), "tab", "bridged-workers", "pinned {profile}", null, null, null);
new ClaudeCodeLauncher(new AgentControl(herdr), new WorkspaceControl(herdr),
new SubscriptionGuard(Set.of("gx00.gw")), Map.of(cfg.profile(), cfg), cfg.profile(),
_ -> null, 0, 0L, () -> "{role}: {profile} #{n}")
.spawn(new SpawnRequest("sonnet", null, null, null, null, MemberRole.REVIEWER));
assertEquals(List.of("pinned sonnet"), tabLabels(herdr));
}
/**
* CB-559: the template is read per spawn, not captured at construction. This is what makes
* {@code fleet.tabLabel} a hot key — a launcher built at boot must see an edit made an hour later
* without being rebuilt.
*/
@Test
void theTemplateIsReadOnEverySpawnSoAnEditTakesEffect() {
FakeHerdr herdr = new FakeHerdr();
AtomicReference<String> template = new AtomicReference<>("{role}: {profile} #{n}");
ClaudeCodeLauncher svc = labelService(herdr, template::get);
svc.spawn(new SpawnRequest("sonnet", null, null, null, null, MemberRole.DEV));
template.set("[{profile}] {role} {n}");
svc.spawn(new SpawnRequest("sonnet", null, null, null, null, MemberRole.DEV));
assertEquals(List.of("dev: sonnet #1", "[sonnet] dev 2"), tabLabels(herdr));
}
}
@@ -0,0 +1,555 @@
package dev.ltms.bridged.member;
import ch.qos.logback.classic.Logger;
import ch.qos.logback.classic.spi.ILoggingEvent;
import ch.qos.logback.core.read.ListAppender;
import dev.ltms.bridged.config.BridgedConfig;
import dev.ltms.bridged.guard.SubscriptionGuard;
import dev.ltms.bridged.herdr.Agent;
import dev.ltms.bridged.herdr.AgentControl;
import dev.ltms.bridged.herdr.FakeHerdr;
import dev.ltms.bridged.herdr.WorkspaceControl;
import dev.ltms.bridged.peer.Capability;
import dev.ltms.bridged.peer.MemberRole;
import dev.ltms.bridged.peer.PeerHandle;
import dev.ltms.bridged.peer.PeerLauncher;
import dev.ltms.bridged.peer.PeerUnreachableException;
import dev.ltms.bridged.peer.SpawnRequest;
import dev.ltms.bridged.placement.PlacementException;
import dev.ltms.bridged.placement.PlacementPolicies;
import org.junit.jupiter.api.Test;
import org.slf4j.LoggerFactory;
import java.util.EnumSet;
import java.util.HashMap;
import java.util.LinkedHashMap;
import java.util.List;
import java.util.Map;
import java.util.Set;
import static org.junit.jupiter.api.Assertions.*;
/**
* The composite router: profile → owning adapter for spawn/cwd/parity, pane id → owner for stop,
* and fleet-wide union/dedup for list/reap/caps/profiles. Exercised through two real adapters —
* claude-code + opencode — over one FakeHerdr, so each call is observed reaching the right adapter
* (the started herdr agent name carries that adapter's {@code claude-}/{@code opencode-} prefix).
*/
class CompositePeerLauncherTest {
private ClaudeCodeLauncher claudeAdapter(FakeHerdr herdr) {
// 12-arg back-compat Worker ctor → kind defaults to claude-code.
BridgedConfig.Profile claude = new BridgedConfig.Profile("claude", "http://gx00.gw:8000", "coder",
null, "BRIDGED_WORKER_TOKEN", List.of("claude"), "tab", "bridged-workers", "w #{n}",
null, null, null);
return new ClaudeCodeLauncher(new AgentControl(herdr), new WorkspaceControl(herdr),
new SubscriptionGuard(Set.of("gx00.gw")), Map.of("claude", claude), "claude", _ -> null);
}
private OpenCodeLauncher opencodeAdapter(FakeHerdr herdr) {
BridgedConfig.Profile gemini = new BridgedConfig.Profile("gemini", null, "google/gemini-2.5-pro",
null, "BRIDGED_WORKER_TOKEN", List.of("opencode"), "tab", "bridged-workers", "w #{n}",
null, null, null, "GITEA_ACCESS_TOKEN", null, BridgedConfig.Profile.KIND_OPENCODE);
return new OpenCodeLauncher(new AgentControl(herdr), new WorkspaceControl(herdr),
Map.of("gemini", gemini), "gemini", _ -> "tok");
}
private CompositePeerLauncher composite(FakeHerdr herdr) {
return new CompositePeerLauncher(
List.of(claudeAdapter(herdr), opencodeAdapter(herdr)), "claude");
}
/**
* A minimal concrete HerdrPeerLauncher for policy tests. It either returns a fake handle for the
* requested profile or throws, depending on {@code failProfiles}. buildLaunch is a stub; only
* spawn/stop/list/caps/reap are exercised by the composite.
*/
private static final class StubLauncher extends HerdrPeerLauncher {
private final Set<String> failProfiles;
private final Map<String, Integer> spawnCounts = new HashMap<>();
StubLauncher(String prefix, FakeHerdr herdr,
Map<String, BridgedConfig.Profile> profiles, String defaultProfile,
Set<String> failProfiles) {
super(prefix, new AgentControl(herdr), new WorkspaceControl(herdr),
profiles, defaultProfile, _ -> null, 0L, System::currentTimeMillis, () -> { });
this.failProfiles = Set.copyOf(failProfiles);
}
@Override
protected Launch buildLaunch(BridgedConfig.Profile cfg) {
return new Launch(Map.of(), List.of());
}
@Override
public PeerHandle spawn(SpawnRequest req) {
String p = (req.profileName() == null || req.profileName().isBlank())
? defaultProfile() : req.profileName();
spawnCounts.merge(p, 1, Integer::sum);
if (failProfiles.contains(p)) {
throw new PeerUnreachableException(p + " is down");
}
return new PeerHandle() {
@Override public String id() { return "pane-" + p; }
@Override public String terminalId() { return "term-" + p; }
@Override public String profile() { return p; }
};
}
@Override
public void stop(String id) { }
@Override
public List<Agent> list() { return List.of(); }
@Override
public Set<Capability> capabilities() { return EnumSet.noneOf(Capability.class); }
@Override
public int reapOrphanWorkers() { return 0; }
int spawnCount(String profile) {
return spawnCounts.getOrDefault(profile, 0);
}
}
private static BridgedConfig.Profile stubWorker(String profile) {
return new BridgedConfig.Profile(profile, "http://gx00.gw:8000", "coder",
null, "BRIDGED_WORKER_TOKEN", List.of("claude"), "tab", "bridged-workers",
"w #{n}", null, null, null, null, null, null, null, null, null);
}
private static BridgedConfig.Profile stubWorker(String profile, float weight, Integer maxLoad) {
return new BridgedConfig.Profile(profile, "http://gx00.gw:8000", "coder",
null, "BRIDGED_WORKER_TOKEN", List.of("claude"), "tab", "bridged-workers",
"w #{n}", null, null, null, null, null, null, null,
weight, maxLoad);
}
/**
* An <em>order-preserving</em> profile map. Never {@code Map.of} here: its iteration order is
* salted per JVM run, and the weighted policy breaks an exact-weight tie on candidate order —
* so a {@code Map.of} would make "which profile is tried first" a coin flip per run and any
* assertion about the first attempt intermittently false.
*/
private static Map<String, BridgedConfig.Profile> ordered(String first, BridgedConfig.Profile a,
String second, BridgedConfig.Profile b) {
Map<String, BridgedConfig.Profile> m = new LinkedHashMap<>();
m.put(first, a);
m.put(second, b);
return m;
}
@SuppressWarnings("unchecked")
private static String startedName(FakeHerdr herdr) {
return (String) ((Map<String, Object>) herdr.lastCall("agent.start").params()).get("name");
}
@Test
void spawnRoutesEachProfileToItsOwningAdapter() {
FakeHerdr herdr = new FakeHerdr();
PeerLauncher composite = composite(herdr);
composite.spawn(new SpawnRequest("gemini", null, null));
assertTrue(startedName(herdr).startsWith("opencode-"),
"the gemini profile is spawned by the opencode adapter: " + startedName(herdr));
composite.spawn(new SpawnRequest("claude", null, null));
assertTrue(startedName(herdr).startsWith("claude-"),
"the claude profile is spawned by the claude-code adapter: " + startedName(herdr));
}
@Test
void nullProfileResolvesTheDefaultAndRoutesToItsOwner() {
FakeHerdr herdr = new FakeHerdr();
composite(herdr).spawn(new SpawnRequest(null, null, null));
assertTrue(startedName(herdr).startsWith("claude-"),
"a no-profile spawn resolves the default (claude) and routes to its adapter");
}
@Test
void unknownProfileIsRejected() {
FakeHerdr herdr = new FakeHerdr();
PeerLauncher composite = composite(herdr);
assertThrows(IllegalArgumentException.class,
() -> composite.spawn(new SpawnRequest("nope", null, null)),
"a profile no adapter declares is an error");
}
@Test
void profilesAndDefaultAreExposedAcrossAdapters() {
FakeHerdr herdr = new FakeHerdr();
PeerLauncher composite = composite(herdr);
assertEquals(Set.of("claude", "gemini"), composite.profiles(),
"profiles are the union of every adapter's profiles");
assertEquals("claude", composite.defaultProfile());
}
@Test
void capabilitiesAreTheUnionOfEveryAdapter() {
FakeHerdr herdr = new FakeHerdr();
ClaudeCodeLauncher claude = claudeAdapter(herdr);
OpenCodeLauncher opencode = opencodeAdapter(herdr);
PeerLauncher composite = new CompositePeerLauncher(List.of(claude, opencode), "claude");
assertTrue(composite.capabilities().containsAll(claude.capabilities()),
"the fleet offers every claude-code capability");
assertTrue(composite.capabilities().containsAll(opencode.capabilities()),
"the fleet offers every opencode capability (incl. SELF_PR from its git-token profile)");
}
@Test
void listIsDeduplicatedByPaneIdAcrossAdaptersSharingHerdr() {
FakeHerdr herdr = new FakeHerdr();
PeerLauncher composite = composite(herdr);
// Both adapters wrap the same herdr, so each list() returns the same global agent set;
// the composite must return each pane once, not once per adapter.
assertEquals(1, composite.list().size(),
"the single herdr-tracked pane appears once, not duplicated per adapter");
}
@Test
void reapSumsAcrossAdaptersAndEachAdapterReapsOnlyItsOwnPrefix() {
// One foreign opencode orphan + one foreign claude orphan, from a prior daemon (different nonce).
FakeHerdr herdr = new FakeHerdr()
.withAgent("opencode-gemini-ffffff-1", "term_o", "wQ:pO", "wQ:tO")
.withAgent("claude-claude-eeeeee-1", "term_c", "wQ:pC", "wQ:tC");
PeerLauncher composite = composite(herdr);
assertEquals(2, composite.reapOrphanWorkers(),
"both orphans are reaped — one by each adapter, summed by the composite");
}
@Test
void stopTearsDownAPaneSpawnedThroughTheComposite() {
// CB-519: handle.id() is a host-unique opaque UUID, not the herdr pane — stop(id) must
// resolve it through the owning adapter down to the actual pane coordinate it spawned.
FakeHerdr herdr = new FakeHerdr();
PeerLauncher composite = composite(herdr);
PeerHandle handle = composite.spawn(new SpawnRequest("gemini", null, null));
assertNotEquals("w9:pRoot_1", handle.id(), "the id is decoupled from the pane coordinate");
composite.stop(handle.id());
assertTrue(herdr.calls.stream()
.anyMatch(c -> c.method().equals("pane.close")
&& "w9:pRoot_1".equals(((Map<?, ?>) c.params()).get("pane_id"))),
"stop routes to the spawning adapter and closes exactly that worker's pane");
}
@Test
void opencodeContextResetIsANoOpAndWarnsOnlyOnce() {
FakeHerdr herdr = new FakeHerdr();
CompositePeerLauncher composite = composite(herdr);
PeerHandle handle = composite.spawn(new SpawnRequest("gemini", null, null));
Logger logger = (Logger) LoggerFactory.getLogger(HerdrPeerLauncher.class);
ListAppender<ILoggingEvent> appender = new ListAppender<>();
appender.start();
logger.addAppender(appender);
try {
assertFalse(composite.clearContext(handle.id()));
assertFalse(composite.clearContext(handle.id()));
} finally {
logger.detachAppender(appender);
}
assertFalse(opencodeAdapter(herdr).capabilities().contains(Capability.CONTEXT_RESET));
assertTrue(herdr.calls.stream().noneMatch(c -> "agent.prompt".equals(c.method())),
"never type Claude's /clear into an opencode prompt");
assertEquals(1, appender.list.stream()
.filter(e -> e.getFormattedMessage().contains("context reset is unsupported"))
.count(), "unsupported reset is logged once per adapter, not once per turn");
}
@Test
void constructorRejectsAProfileClaimedByTwoAdapters() {
FakeHerdr herdr = new FakeHerdr();
// Two opencode adapters both declaring "gemini" — a profile-name collision.
OpenCodeLauncher a = opencodeAdapter(herdr);
OpenCodeLauncher b = opencodeAdapter(herdr);
assertThrows(IllegalArgumentException.class,
() -> new CompositePeerLauncher(List.of(a, b), "gemini"),
"a profile two adapters both claim is a configuration error");
}
@Test
void constructorRejectsAnEmptyAdapterList() {
assertThrows(IllegalArgumentException.class,
() -> new CompositePeerLauncher(List.of(), "claude"),
"at least one adapter must be configured");
}
@Test
void fixedDefaultIsNoOpForUnqualifiedSpawns() {
FakeHerdr herdr = new FakeHerdr();
CompositePeerLauncher composite = composite(herdr);
PeerHandle h = composite.spawn(new SpawnRequest(null, null, null));
assertTrue(startedName(herdr).startsWith("claude-"),
"fixed placement still routes an unqualified spawn to the default profile");
assertEquals("claude", h.profile(), "the returned handle carries the resolved default profile");
}
@Test
void weightedPolicyGatesProfileAtMaxLoad() {
FakeHerdr herdr = new FakeHerdr();
Map<String, BridgedConfig.Profile> profiles = ordered(
"a", stubWorker("a", 1.0f, 1),
"b", stubWorker("b", 1.0f, null));
StubLauncher adapter = new StubLauncher("claude", herdr, profiles, "a", Set.of());
CompositePeerLauncher composite = new CompositePeerLauncher(
List.of(adapter), "a", profiles, PlacementPolicies.weighted(), name -> "a".equals(name) ? 1 : 0);
for (int i = 0; i < 5; i++) {
PeerHandle h = composite.spawn(new SpawnRequest(null, null, null));
assertEquals("b", h.profile(), "profile a is at maxLoad, so every spawn must land on b");
}
}
@Test
void weightedPolicyDistributesAccordingToWeightRatio() {
FakeHerdr herdr = new FakeHerdr();
Map<String, BridgedConfig.Profile> profiles = ordered(
"a", stubWorker("a", 0.75f, null),
"b", stubWorker("b", 0.25f, null));
StubLauncher adapter = new StubLauncher("claude", herdr, profiles, "a", Set.of());
CompositePeerLauncher composite = new CompositePeerLauncher(
List.of(adapter), "a", profiles, PlacementPolicies.weighted(), _ -> 0);
int a = 0, b = 0;
for (int i = 0; i < 40; i++) {
String p = composite.spawn(new SpawnRequest(null, null, null)).profile();
if ("a".equals(p)) a++;
else if ("b".equals(p)) b++;
}
assertEquals(30, a, "weighted distribution should hold the 3:1 ratio");
assertEquals(10, b);
}
@Test
void failoverRetriesNextCandidateWhenProfileIsUnreachable() {
FakeHerdr herdr = new FakeHerdr();
Map<String, BridgedConfig.Profile> profiles = ordered(
"a", stubWorker("a"),
"b", stubWorker("b"));
StubLauncher adapter = new StubLauncher("claude", herdr, profiles, "a", Set.of("a"));
CompositePeerLauncher composite = new CompositePeerLauncher(
List.of(adapter), "a", profiles, PlacementPolicies.weighted(), _ -> 0);
PeerHandle h = composite.spawn(new SpawnRequest(null, null, null));
assertEquals("b", h.profile(), "the spawn must fail over from unreachable a to b");
assertEquals(1, adapter.spawnCount("a"), "a was tried once and failed");
assertEquals(1, adapter.spawnCount("b"), "b was tried once and succeeded");
}
/**
* Definition order — not hash order — decides an exact-weight tie. Paired with the test above
* (same two profiles, opposite declaration order, opposite expected first attempt) this pins the
* ordering contract from both sides: under a salted map one of the two must fail on every run.
*/
@Test
void reversingDefinitionOrderReversesWhichProfileIsTriedFirst() {
FakeHerdr herdr = new FakeHerdr();
Map<String, BridgedConfig.Profile> profiles = ordered(
"b", stubWorker("b"),
"a", stubWorker("a"));
StubLauncher adapter = new StubLauncher("claude", herdr, profiles, "a", Set.of("b"));
CompositePeerLauncher composite = new CompositePeerLauncher(
List.of(adapter), "a", profiles, PlacementPolicies.weighted(), _ -> 0);
PeerHandle h = composite.spawn(new SpawnRequest(null, null, null));
assertEquals("a", h.profile(), "b is declared first and unreachable, so the spawn lands on a");
assertEquals(1, adapter.spawnCount("b"), "b, declared first, is the one tried first");
}
@Test
void failoverBoundedByCandidateCount() {
FakeHerdr herdr = new FakeHerdr();
Map<String, BridgedConfig.Profile> profiles = ordered(
"a", stubWorker("a"),
"b", stubWorker("b"));
StubLauncher adapter = new StubLauncher("claude", herdr, profiles, "a", Set.of("a", "b"));
CompositePeerLauncher composite = new CompositePeerLauncher(
List.of(adapter), "a", profiles, PlacementPolicies.weighted(), _ -> 0);
PeerUnreachableException e = assertThrows(PeerUnreachableException.class,
() -> composite.spawn(new SpawnRequest(null, null, null)));
assertTrue(e.getMessage().contains("no reachable worker profile"), e.getMessage());
assertEquals(1, adapter.spawnCount("a"));
assertEquals(1, adapter.spawnCount("b"));
}
@Test
void explicitSpawnAtMaxLoadThrowsPlacementExceptionNamingProfileLiveAndCap() {
FakeHerdr herdr = new FakeHerdr();
Map<String, BridgedConfig.Profile> profiles = ordered(
"a", stubWorker("a", 1.0f, 2),
"b", stubWorker("b"));
StubLauncher adapter = new StubLauncher("claude", herdr, profiles, "a", Set.of());
CompositePeerLauncher composite = new CompositePeerLauncher(
List.of(adapter), "a", profiles, PlacementPolicies.fixed(), name -> "a".equals(name) ? 2 : 0);
PlacementException e = assertThrows(PlacementException.class,
() -> composite.spawn(new SpawnRequest("a", null, null)));
assertTrue(e.getMessage().contains("'a'"), "message names the profile: " + e.getMessage());
assertTrue(e.getMessage().contains("2 live"), "message names the live count: " + e.getMessage());
assertTrue(e.getMessage().contains("2 cap"), "message names the cap: " + e.getMessage());
assertEquals(0, adapter.spawnCount("a"), "at cap, the spawn is refused before any delegation");
}
@Test
void explicitSpawnUnderMaxLoadStillSucceeds() {
FakeHerdr herdr = new FakeHerdr();
Map<String, BridgedConfig.Profile> profiles = ordered(
"a", stubWorker("a", 1.0f, 2),
"b", stubWorker("b"));
StubLauncher adapter = new StubLauncher("claude", herdr, profiles, "a", Set.of());
CompositePeerLauncher composite = new CompositePeerLauncher(
List.of(adapter), "a", profiles, PlacementPolicies.fixed(), name -> "a".equals(name) ? 1 : 0);
PeerHandle h = composite.spawn(new SpawnRequest("a", null, null));
assertEquals("a", h.profile(), "a profile under its cap accepts an explicit spawn");
assertEquals(1, adapter.spawnCount("a"), "the under-cap spawn is delegated");
}
@Test
void explicitSpawnWithNullMaxLoadIsNeverCapped() {
FakeHerdr herdr = new FakeHerdr();
Map<String, BridgedConfig.Profile> profiles = ordered(
"a", stubWorker("a", 1.0f, null),
"b", stubWorker("b"));
StubLauncher adapter = new StubLauncher("claude", herdr, profiles, "a", Set.of());
// A deliberately absurd live count: an unset maxLoad means unlimited, so it must never refuse.
CompositePeerLauncher composite = new CompositePeerLauncher(
List.of(adapter), "a", profiles, PlacementPolicies.fixed(), _ -> 1000);
PeerHandle h = composite.spawn(new SpawnRequest("a", null, null));
assertEquals("a", h.profile(), "a profile with no maxLoad is never capped, however many live workers");
}
@Test
void emptyCandidateSetThrowsClearException() {
FakeHerdr herdr = new FakeHerdr();
Map<String, BridgedConfig.Profile> profiles = ordered(
"a", stubWorker("a", 1.0f, 1),
"b", stubWorker("b", 1.0f, 1));
StubLauncher adapter = new StubLauncher("claude", herdr, profiles, "a", Set.of());
CompositePeerLauncher composite = new CompositePeerLauncher(
List.of(adapter), "a", profiles, PlacementPolicies.weighted(), _ -> 1);
PlacementException e = assertThrows(PlacementException.class,
() -> composite.spawn(new SpawnRequest(null, null, null)));
assertTrue(e.getMessage().contains("maxLoad"), e.getMessage());
}
// ── CB-557: an unqualified spawn is placed inside its role's pool ─────────────────────────
/** Three profiles in definition order — pools are carved out of this set. */
private static Map<String, BridgedConfig.Profile> threeProfiles() {
Map<String, BridgedConfig.Profile> m = new LinkedHashMap<>();
m.put("opus", stubWorker("opus", 1.0f, null));
m.put("sonnet", stubWorker("sonnet", 1.0f, null));
m.put("terra", stubWorker("terra", 1.0f, null));
return m;
}
private static Map<String, BridgedConfig.Slot> pool(String... names) {
Map<String, BridgedConfig.Slot> m = new LinkedHashMap<>();
for (String n : names) {
m.put(n, new BridgedConfig.Slot(n));
}
return m;
}
private static CompositePeerLauncher withPools(FakeHerdr herdr, BridgedConfig.Fleet fleet) {
Map<String, BridgedConfig.Profile> profiles = threeProfiles();
StubLauncher adapter = new StubLauncher("claude", herdr, profiles, "opus", Set.of());
return new CompositePeerLauncher(List.of(adapter), "opus", profiles,
PlacementPolicies.fixed(), _ -> 0, fleet);
}
/**
* The point of the pools: a role is placed only on a backend its pool names. Before CB-557 an
* unqualified spawn ranged over every configured profile, so a reviewer could land on the
* architect-only one.
*/
@Test
void anUnqualifiedSpawnIsPlacedInsideItsRolePool() {
FakeHerdr herdr = new FakeHerdr();
CompositePeerLauncher composite = withPools(herdr, new BridgedConfig.Fleet(
Map.of(), pool("opus"), pool("terra"), pool("sonnet"), null));
assertEquals("opus", composite.spawn(
new SpawnRequest(null, null, null, null, null, MemberRole.ARCHITECT)).profile());
assertEquals("terra", composite.spawn(
new SpawnRequest(null, null, null, null, null, MemberRole.DEV)).profile());
assertEquals("sonnet", composite.spawn(
new SpawnRequest(null, null, null, null, null, MemberRole.REVIEWER)).profile());
}
/** Under `fixed`, the pool's first entry wins — not the global defaultProfile. */
@Test
void theRolePoolOutranksTheGlobalDefaultProfile() {
FakeHerdr herdr = new FakeHerdr();
CompositePeerLauncher composite = withPools(herdr, new BridgedConfig.Fleet(
Map.of(), Map.of(), pool("sonnet", "terra"), Map.of(), null));
assertEquals("sonnet", composite.spawn(
new SpawnRequest(null, null, null, null, null, MemberRole.DEV)).profile(),
"the dev pool starts at sonnet, so the global default 'opus' must not win");
}
/**
* A role with no pool is unconstrained, not blocked. A config that declares pools for some roles
* and not others must keep spawning the rest.
*/
@Test
void aRoleWithNoPoolFallsBackToEveryProfile() {
FakeHerdr herdr = new FakeHerdr();
CompositePeerLauncher composite = withPools(herdr, new BridgedConfig.Fleet(
Map.of(), pool("sonnet"), Map.of(), Map.of(), null));
assertEquals("opus", composite.spawn(
new SpawnRequest(null, null, null, null, null, MemberRole.DEV)).profile(),
"no dev pool ⇒ all profiles are candidates, so `fixed` takes the first one");
}
/** No fleet at all is the pre-CB-557 wiring, and must behave exactly as it did. */
@Test
void noFleetConfiguredKeepsTheOldWholeProfileListBehaviour() {
FakeHerdr herdr = new FakeHerdr();
CompositePeerLauncher composite = withPools(herdr, null);
assertEquals("opus", composite.spawn(new SpawnRequest(null, null, null)).profile());
}
/**
* An explicit profile is the operator overriding and is NOT judged against the pool. It must
* stay that way: an unrolled `bridge_spawn{profile:"opus"}` carries no role, so it defaults to
* DEV, and enforcing the pool here would refuse a spawn the operator asked for by name.
*/
@Test
void anExplicitProfileIsNotConfinedToTheRolePool() {
FakeHerdr herdr = new FakeHerdr();
CompositePeerLauncher composite = withPools(herdr, new BridgedConfig.Fleet(
Map.of(), pool("opus"), pool("terra"), Map.of(), null));
assertEquals("opus", composite.spawn(new SpawnRequest("opus", null, null)).profile(),
"naming opus explicitly must work even though the dev pool holds only terra");
}
/** Placement still respects maxLoad, but only across the pool — never by escaping it. */
@Test
void aFullPoolIsRefusedRatherThanSpilledOntoAnotherRolesProfile() {
FakeHerdr herdr = new FakeHerdr();
Map<String, BridgedConfig.Profile> profiles = new LinkedHashMap<>();
profiles.put("opus", stubWorker("opus", 1.0f, null)); // architect-only, uncapped
profiles.put("terra", stubWorker("terra", 1.0f, 1)); // the sole dev, capped at 1
StubLauncher adapter = new StubLauncher("claude", herdr, profiles, "opus", Set.of());
CompositePeerLauncher composite = new CompositePeerLauncher(List.of(adapter), "opus", profiles,
PlacementPolicies.weighted(), name -> "terra".equals(name) ? 1 : 0,
new BridgedConfig.Fleet(Map.of(), pool("opus"), pool("terra"), Map.of(), null));
PlacementException e = assertThrows(PlacementException.class, () -> composite.spawn(
new SpawnRequest(null, null, null, null, null, MemberRole.DEV)));
assertTrue(e.getMessage().contains("maxLoad"), e.getMessage());
}
}
@@ -1,4 +1,4 @@
package dev.ltms.bridged.worker;
package dev.ltms.bridged.member;
import com.fasterxml.jackson.databind.JsonNode;
import com.fasterxml.jackson.databind.ObjectMapper;
@@ -27,17 +27,17 @@ import static org.junit.jupiter.api.Assertions.*;
*/
class OpenCodeLauncherTest {
private static BridgedConfig.Worker opencodeCfg(String model, String mcpUrl, String gitTokenEnv) {
return new BridgedConfig.Worker("gemini", null, model, null, "BRIDGED_WORKER_TOKEN",
private static BridgedConfig.Profile opencodeCfg(String model, String mcpUrl, String gitTokenEnv) {
return new BridgedConfig.Profile("gemini", null, model, null, "BRIDGED_WORKER_TOKEN",
List.of("opencode"), "tab", "bridged-workers", "opencode: {model} #{n}", mcpUrl,
null, null, gitTokenEnv, null, BridgedConfig.Worker.KIND_OPENCODE);
null, null, gitTokenEnv, null, BridgedConfig.Profile.KIND_OPENCODE);
}
/** Gate-disabled launcher whose per-spawn config dirs land under an inspectable temp root. */
private OpenCodeLauncher service(FakeHerdr herdr, Path configRoot, BridgedConfig.Worker cfg) {
private OpenCodeLauncher service(FakeHerdr herdr, Path configRoot, BridgedConfig.Profile cfg) {
return new OpenCodeLauncher(new AgentControl(herdr), new WorkspaceControl(herdr),
Map.of(cfg.profile(), cfg), cfg.profile(), k -> "GITEA_ACCESS_TOKEN".equals(k) ? "tok" : null,
0, System::currentTimeMillis, () -> { }, configRoot);
0, System::currentTimeMillis, () -> { }, configRoot, configRoot);
}
@SuppressWarnings("unchecked")
@@ -45,14 +45,18 @@ class OpenCodeLauncherTest {
return (Map<String, Object>) herdr.lastCall("agent.start").params();
}
/** Protocol 19: the worker's env is injected at pane creation (tab.create), not agent.start. */
@SuppressWarnings("unchecked")
private static Map<String, String> startEnv(FakeHerdr herdr) {
return (Map<String, String>) lastStart(herdr).get("env");
Map<String, String> env =
(Map<String, String>) ((Map<String, Object>) herdr.lastCall("tab.create").params()).get("env");
return env == null ? Map.of() : env;
}
/** Protocol 19: agent.start carries only the args after the kind-resolved executable. */
@SuppressWarnings("unchecked")
private static List<String> startArgv(FakeHerdr herdr) {
return (List<String>) lastStart(herdr).get("argv");
private static List<String> startArgs(FakeHerdr herdr) {
return (List<String>) lastStart(herdr).get("args");
}
@Test
@@ -70,6 +74,8 @@ class OpenCodeLauncherTest {
// Assert on parsed structure, not substrings: the generated config is real JSON and its
// whitespace is the formatter's business, not the contract's.
JsonNode json = new ObjectMapper().readTree(Path.of(cfgPath).toFile());
assertTrue(json.path("compaction").path("auto").asBoolean(),
"spawned opencode peers explicitly enable automatic compaction");
JsonNode bridge = json.path("mcp").path("bridge");
assertEquals("remote", bridge.path("type").asText(), "bridge is mounted as a remote MCP server");
assertEquals("http://127.0.0.1:8765/mcp", bridge.path("url").asText(),
@@ -95,22 +101,24 @@ class OpenCodeLauncherTest {
}
@Test
void passesTheModelAsDashMFlag(@TempDir Path root) {
void passesTheModelAsDashMFlagAlongsideAutoApprove(@TempDir Path root) {
FakeHerdr herdr = new FakeHerdr();
service(herdr, root, opencodeCfg("google/gemini-2.5-pro", null, null)).spawn();
List<String> argv = startArgv(herdr);
assertEquals("opencode", argv.getFirst(), "base opencode command preserved first");
int m = argv.indexOf("-m");
List<String> args = startArgs(herdr);
assertTrue(args.contains("--auto"),
"--auto is present alongside -m so a spawned peer never blocks on approval");
int m = args.indexOf("-m");
assertTrue(m >= 0, "model is selected with -m");
assertEquals("google/gemini-2.5-pro", argv.get(m + 1), "the provider/model selector follows -m");
assertEquals("google/gemini-2.5-pro", args.get(m + 1), "the provider/model selector follows -m");
}
@Test
void noModelFlagWhenModelBlank(@TempDir Path root) {
void autoApproveIsUnconditionalWhenModelBlank(@TempDir Path root) {
FakeHerdr herdr = new FakeHerdr();
service(herdr, root, opencodeCfg(null, null, null)).spawn();
assertEquals(List.of("opencode"), startArgv(herdr), "no model → argv is the bare opencode command");
assertEquals(List.of("--auto"), startArgs(herdr),
"--auto is unconditional: a model-less worker still must never block on approval");
}
@Test
@@ -124,14 +132,65 @@ class OpenCodeLauncherTest {
@Test
void capabilitiesDeclareOrphanReapAndMcpAskAndConditionalSelfPr(@TempDir Path root) {
FakeHerdr herdr = new FakeHerdr();
assertEquals(java.util.Set.of(Capability.MID_TURN_ASK, Capability.WORKTREE, Capability.ORPHAN_REAP),
assertEquals(java.util.Set.of(Capability.MID_TURN_ASK, Capability.WORKTREE, Capability.ORPHAN_REAP,
Capability.SESSION_RESUME),
service(herdr, root, opencodeCfg(null, null, null)).capabilities(),
"no git token → no SELF_PR");
"opencode can be resumed by its own session id, so SESSION_RESUME is always declared");
assertFalse(service(herdr, root, opencodeCfg(null, null, null))
.capabilities().contains(Capability.SESSION_NAME),
"opencode has no display-name flag, so SESSION_NAME must NOT be declared");
assertTrue(service(herdr, root, opencodeCfg(null, null, "GITEA_ACCESS_TOKEN"))
.capabilities().contains(Capability.SELF_PR),
"a git-token profile adds SELF_PR");
}
// --- CB-547: resume + post-hoc session discovery --------------------------------------------
@Test
void aResumeSpawnPassesTheSessionIdAsDashS(@TempDir Path root) {
FakeHerdr herdr = new FakeHerdr();
service(herdr, root, opencodeCfg("google/gemini-2.5-pro", null, null))
.spawn(new SpawnRequest(null, null, null, null, "ses_41b79fc90ffeI9E8uZv6VprUn2"));
List<String> args = startArgs(herdr);
int s = args.indexOf("-s");
assertTrue(s >= 0, "a resumed spawn carries opencode's -s flag");
assertEquals("ses_41b79fc90ffeI9E8uZv6VprUn2", args.get(s + 1),
"the resume target id follows -s");
}
@Test
void aFreshSpawnCarriesNoSessionFlag(@TempDir Path root) {
FakeHerdr herdr = new FakeHerdr();
service(herdr, root, opencodeCfg(null, null, null))
.spawn(new SpawnRequest(null, null, null, null, null));
assertFalse(startArgs(herdr).contains("-s"),
"no resume target → a fresh session with no -s flag");
}
@Test
void theHandleDiscoversTheSessionIdForTheWorkersCwdOnlyAfterItAppears(@TempDir Path root,
@TempDir Path discRoot)
throws Exception {
FakeHerdr herdr = new FakeHerdr();
OpenCodeLauncher launcher = new OpenCodeLauncher(new AgentControl(herdr),
new WorkspaceControl(herdr), Map.of("gemini", opencodeCfg(null, null, null)),
"gemini", _ -> null, 0, System::currentTimeMillis, () -> { }, root, discRoot);
PeerHandle handle = launcher.spawn(new SpawnRequest(null, "/work/dir", null));
// opencode writes the record only when the session is first persisted — the instant the
// pane is ready it does not exist, so agentSessionId() is null (never a spawn failure).
assertNull(handle.agentSessionId(), "no record yet → null, not a spawn-time block");
// Once the record appears (here: same cwd), lazy discovery resolves it — the handle's
// session id matches its own worktree, not another's.
OpenCodeSessionDiscoveryTest.writeRecord(discRoot, "p1", "ses_a.json",
"ses_resolved", "/work/dir", 1000L);
assertEquals("ses_resolved", handle.agentSessionId(),
"agentSessionId() re-scans and picks up a record that has since been written");
}
@Test
void foreignWorkerMatchesOpencodePrefixButNotClaude() {
String nonce = "abc123";
@@ -146,7 +205,7 @@ class OpenCodeLauncherTest {
@Test
void productionConstructorsWireThroughToTheBase() {
FakeHerdr herdr = new FakeHerdr();
BridgedConfig.Worker cfg = opencodeCfg(null, null, null);
BridgedConfig.Profile cfg = opencodeCfg(null, null, null);
// 5-arg (gate disabled) and 7-arg (gate enabled) production constructors both expose the profile.
OpenCodeLauncher disabled = new OpenCodeLauncher(new AgentControl(herdr), new WorkspaceControl(herdr),
Map.of(cfg.profile(), cfg), cfg.profile(), _ -> null);
@@ -163,14 +222,14 @@ class OpenCodeLauncherTest {
long[] clock = {0};
OpenCodeLauncher svc = new OpenCodeLauncher(new AgentControl(herdr), new WorkspaceControl(herdr),
Map.of("gemini", opencodeCfg(null, null, null)), "gemini", _ -> null,
1000, () -> clock[0], () -> clock[0] += 50, root);
1000, () -> clock[0], () -> clock[0] += 50, root, root);
PeerUnreachableException ex = assertThrows(PeerUnreachableException.class,
() -> svc.spawn(new SpawnRequest(null, null, null)));
assertTrue(clock[0] >= 1000, "the fake clock advanced past the timeout: " + clock[0]);
long closes = herdr.calls.stream()
.filter(c -> c.method().equals("pane.close"))
.filter(c -> "w9:pW_1".equals(((Map<?, ?>) c.params()).get("pane_id")))
.filter(c -> "w9:pRoot_1".equals(((Map<?, ?>) c.params()).get("pane_id")))
.count();
assertEquals(1, closes, "the worker pane was reaped on timeout (no orphan)");
assertNotNull(ex.getMessage());
@@ -188,10 +247,10 @@ class OpenCodeLauncherTest {
// --- CB-508: pinned OpenAI-compatible endpoint (e.g. a local vLLM) ---------------------------
/** A profile with a baseUrl but no model provider prefix cannot be resolved — fail loudly. */
private static BridgedConfig.Worker pinnedCfg(String model, String baseUrl, String mcpUrl) {
return new BridgedConfig.Worker("local", baseUrl, model, null, "BRIDGED_WORKER_TOKEN",
private static BridgedConfig.Profile pinnedCfg(String model, String baseUrl, String mcpUrl) {
return new BridgedConfig.Profile("local", baseUrl, model, null, "BRIDGED_WORKER_TOKEN",
List.of("opencode"), "tab", "bridged-workers", "opencode: {model} #{n}", mcpUrl,
null, null, null, null, BridgedConfig.Worker.KIND_OPENCODE);
null, null, null, null, BridgedConfig.Profile.KIND_OPENCODE);
}
@Test
@@ -0,0 +1,90 @@
package dev.ltms.bridged.member;
import org.junit.jupiter.api.Test;
import org.junit.jupiter.api.io.TempDir;
import java.nio.file.Files;
import java.nio.file.Path;
import java.nio.file.attribute.FileTime;
import static org.junit.jupiter.api.Assertions.*;
/**
* {@link OpenCodeSessionDiscovery} matches an opencode session record by the worker's cwd (its
* {@code directory}) against opencode's on-disk storage. These tests populate a TEMP storage root
* themselves — never the operator's real {@code ~/.local/share/opencode}.
*/
class OpenCodeSessionDiscoveryTest {
/**
* Write a session record {@code {"id":..., "directory":...}} under
* {@code <root>/session/<projectID>/<fileName>} and stamp it with a known last-modified time,
* so "most recently modified wins" is deterministic. Static so the launcher test can reuse it.
*/
static void writeRecord(Path root, String projectId, String fileName, String id,
String directory, long lastModifiedEpochMillis) throws Exception {
Path dir = root.resolve("session").resolve(projectId);
Files.createDirectories(dir);
Path file = dir.resolve(fileName);
Files.writeString(file, "{\"id\":\"" + id + "\",\"directory\":\"" + directory
+ "\",\"projectID\":\"" + projectId + "\",\"version\":\"1.1.31\"}");
Files.setLastModifiedTime(file, FileTime.fromMillis(lastModifiedEpochMillis));
}
@Test
void findsTheRecordWhoseDirectoryEqualsTheCwd(@TempDir Path root) throws Exception {
writeRecord(root, "p1", "ses_a.json", "ses_aaa", "/w/a", 1000L);
writeRecord(root, "p2", "ses_b.json", "ses_bbb", "/w/b", 2000L);
assertEquals("ses_bbb", new OpenCodeSessionDiscovery(root).sessionIdForDirectory("/w/b"),
"the record whose directory equals the cwd is the one found");
assertEquals("ses_aaa", new OpenCodeSessionDiscovery(root).sessionIdForDirectory("/w/a"));
}
@Test
void aNonMatchingDirectoryYieldsNullRatherThanAMismatch(@TempDir Path root) throws Exception {
writeRecord(root, "p1", "ses_a.json", "ses_aaa", "/w/a", 1000L);
assertNull(new OpenCodeSessionDiscovery(root).sessionIdForDirectory("/w/other"),
"no record for this cwd yet → null, not a wrong session");
}
@Test
void prefersTheMostRecentlyModifiedRecordWhenSeveralMatch(@TempDir Path root) throws Exception {
writeRecord(root, "p1", "old.json", "ses_old", "/w/a", 1000L);
writeRecord(root, "p2", "new.json", "ses_new", "/w/a", 5000L);
assertEquals("ses_new", new OpenCodeSessionDiscovery(root).sessionIdForDirectory("/w/a"),
"the freshest record for the cwd wins");
}
@Test
void aMissingOrEmptyStorageRootYieldsNullWithoutThrowing(@TempDir Path root) throws Exception {
// Missing: no session dir at all under the root.
assertNull(new OpenCodeSessionDiscovery(root).sessionIdForDirectory("/w/a"));
// Present but empty: a session dir with nothing in it produces no match, not a throw.
Path emptyRoot = root.resolve("empty");
Files.createDirectories(emptyRoot.resolve("session"));
assertNull(new OpenCodeSessionDiscovery(emptyRoot).sessionIdForDirectory("/w/a"));
}
@Test
void aBlankOrNullDirectoryYieldsNull(@TempDir Path root) {
OpenCodeSessionDiscovery discovery = new OpenCodeSessionDiscovery(root);
assertNull(discovery.sessionIdForDirectory(null));
assertNull(discovery.sessionIdForDirectory(" "));
}
@Test
void aMalformedRecordIsSkippedRatherThanFatal(@TempDir Path root) throws Exception {
// A record that fails to parse must not abort the scan of its siblings.
Path dir = root.resolve("session").resolve("p1");
Files.createDirectories(dir);
Files.writeString(dir.resolve("broken.json"), "{not valid json");
writeRecord(root, "p1", "good.json", "ses_good", "/w/a", 1000L);
assertEquals("ses_good", new OpenCodeSessionDiscovery(root).sessionIdForDirectory("/w/a"),
"an unreadable record is skipped; a later valid one still matches");
}
}
@@ -1,5 +1,10 @@
package dev.ltms.bridged.msg;
import com.rabbitmq.client.AMQP;
import com.rabbitmq.client.Channel;
import com.rabbitmq.client.Connection;
import com.rabbitmq.client.ConnectionFactory;
import org.junit.jupiter.api.BeforeAll;
import org.junit.jupiter.api.Tag;
import org.junit.jupiter.api.Test;
import org.testcontainers.containers.RabbitMQContainer;
@@ -19,19 +24,46 @@ import static org.junit.jupiter.api.Assertions.assertTrue;
* excluded from {@code mvn test}/{@code mvn clean install} (which stay hermetic and need no Docker);
* run it with Docker present via {@code mvn test -Pcontract}.
*
* <p>Two broker modes:
* <ul>
* <li><b>Locally</b> ({@code AMQP_URI} unset): Testcontainers spins a RabbitMQ container. Requires
* a working Docker engine; see the "Running the contract tests" note in
* {@code docs/CB-307-Reliable-Delivery.md} for the {@code api.version} engine-compat pin.</li>
* <li><b>In CI</b> ({@code AMQP_URI} set): a RabbitMQ service container provisions the broker and
* {@code AMQP_URI} points at it, so the contract job needs <em>no</em> Docker on the runner.</li>
* </ul>
*
* <p>It proves the port contract on genuine infrastructure: eventual visibility of a published reply,
* ack removal, msgId dedup, and — the reason Stage 2 exists — cross-restart durability: an unacked
* reply survives closing the inbox and is redelivered to a fresh connection.
* ack removal, msgId dedup, explicit ownership, and — the reason Stage 2 exists — cross-restart
* durability: an unacked reply survives closing the inbox and is redelivered to a fresh connection.
*/
@Tag("contract")
@Testcontainers
// disabledWithoutDocker=false on purpose: on the CI path (AMQP_URI set) no container is started and
// the class must still RUN against the external broker even though the runner has no Docker — a
// disabled-without-docker check would silently skip the whole contract suite there.
@Testcontainers(disabledWithoutDocker = false)
class AmqpReplyInboxContractTest {
@Container
static final RabbitMQContainer BROKER =
// When a broker is provisioned out-of-band (CI service container), AMQP_URI takes us straight to
// it and we never touch Testcontainers. Unset locally → Testcontainers starts the container below.
private static final String EXTERNAL_URI = System.getenv("AMQP_URI");
private static final RabbitMQContainer BROKER =
new RabbitMQContainer(DockerImageName.parse("rabbitmq:3.13-management"));
// No @Container on BROKER: the JUnit 5 extension would force-start it even when AMQP_URI is set.
// Start it manually only on the local (no-external-broker) path; Ryuk reaps it on JVM exit.
@BeforeAll
static void startBrokerUnlessExternal() {
if (EXTERNAL_URI == null) {
BROKER.start();
}
}
private static String uri() {
if (EXTERNAL_URI != null) {
return EXTERNAL_URI;
}
// guest/guest against the mapped AMQP port. No trailing slash: an empty path is vhost "",
// which does not exist — omitting it selects the default vhost "/".
return "amqp://guest:guest@" + BROKER.getHost() + ":" + BROKER.getAmqpPort();
@@ -41,6 +73,7 @@ class AmqpReplyInboxContractTest {
void publishThenPeekThenAck() throws Exception {
String target = "worker-pub-" + System.nanoTime();
try (AmqpReplyInbox inbox = AmqpReplyInbox.open(uri())) {
inbox.own(target);
inbox.publish(target, "m1", "hello primary");
List<ReplyInbox.InboxMessage> got = awaitPeek(inbox, target);
@@ -58,6 +91,7 @@ class AmqpReplyInboxContractTest {
void duplicateMsgIdIsNotDoubleQueued() throws Exception {
String target = "worker-dedup-" + System.nanoTime();
try (AmqpReplyInbox inbox = AmqpReplyInbox.open(uri())) {
inbox.own(target);
inbox.publish(target, "dup", "first");
awaitPeek(inbox, target);
inbox.publish(target, "dup", "second"); // same msgId — must be a no-op
@@ -74,8 +108,9 @@ class AmqpReplyInboxContractTest {
void unackedReplySurvivesRestartAndIsRedelivered() throws Exception {
String target = "worker-durable-" + System.nanoTime();
// First "process life": publish, see it held, but crash before acking.
// First "process life": own, publish, see it held, but crash before acking.
try (AmqpReplyInbox first = AmqpReplyInbox.open(uri())) {
first.own(target);
first.publish(target, "persist-1", "survive me");
assertEquals(1, awaitPeek(first, target).size());
// no ack — simulate a java -jar bounce with the reply still pending
@@ -83,6 +118,7 @@ class AmqpReplyInboxContractTest {
// Second "process life": a fresh connection to the same broker must be redelivered the reply.
try (AmqpReplyInbox second = AmqpReplyInbox.open(uri())) {
second.own(target);
List<ReplyInbox.InboxMessage> got = awaitPeek(second, target);
assertEquals(1, got.size(), "an unacked persistent reply is redelivered after restart");
assertEquals("persist-1", got.getFirst().msgId());
@@ -93,11 +129,43 @@ class AmqpReplyInboxContractTest {
// Third life: once acked, it is gone for good — durability is not endless replay.
try (AmqpReplyInbox third = AmqpReplyInbox.open(uri())) {
third.own(target);
Thread.sleep(500);
assertTrue(third.peek(target).isEmpty(), "an acked reply does not come back on the next restart");
}
}
@Test
void publishDoesNotAttachAConsumer() throws Exception {
String target = "worker-pub-no-consumer-" + System.nanoTime();
try (AmqpReplyInbox inbox = AmqpReplyInbox.open(uri());
Connection inspect = newConnection()) {
inbox.own(target);
inbox.publish(target, "m1", "published");
awaitPeek(inbox, target); // ensure the owner's consumer received it
try (Channel ch = inspect.createChannel()) {
AMQP.Queue.DeclareOk ok = ch.queueDeclare(queueName(target), true, false, false, null);
assertEquals(1, ok.getConsumerCount(),
"publish must not attach a consumer; only the owner's consumer should exist");
}
}
}
@Test
void releaseCancelsConsumerAndClearsHeld() throws Exception {
String target = "worker-release-" + System.nanoTime();
try (AmqpReplyInbox inbox = AmqpReplyInbox.open(uri())) {
inbox.own(target);
inbox.publish(target, "m1", "release me");
assertEquals(1, awaitPeek(inbox, target).size(), "owned target holds the reply");
inbox.release(target);
assertTrue(inbox.peek(target).isEmpty(),
"release clears the local held snapshot");
}
}
/** Poll peek (broker delivery is async) until a reply for {@code target} appears or ~10s elapse. */
@SuppressWarnings("BusyWait") // deliberate poll for async broker delivery, bounded by the deadline
private static List<ReplyInbox.InboxMessage> awaitPeek(AmqpReplyInbox inbox, String target)
@@ -110,4 +178,15 @@ class AmqpReplyInboxContractTest {
}
return msgs;
}
/** A separate broker connection for inspecting queue state without disturbing the inbox. */
private static Connection newConnection() throws Exception {
ConnectionFactory factory = new ConnectionFactory();
factory.setUri(uri());
return factory.newConnection("contract-inspector");
}
private static String queueName(String target) {
return "agent." + target + ".inbox";
}
}
@@ -1,5 +1,6 @@
package dev.ltms.bridged.msg;
import org.junit.jupiter.api.BeforeEach;
import org.junit.jupiter.api.Test;
import java.util.concurrent.CountDownLatch;
@@ -10,13 +11,18 @@ import java.util.concurrent.atomic.AtomicReference;
import static org.junit.jupiter.api.Assertions.*;
/**
* Unit tests for {@link InMemoryReplyInbox}: publish, peek, ack, dedup, FIFO ordering, and thread
* safety under concurrent publish vs. drain.
* Unit tests for {@link InMemoryReplyInbox}: publish, peek, ack, dedup, FIFO ordering, thread
* safety under concurrent publish vs. drain, and explicit ownership.
*/
class InMemoryReplyInboxTest {
private final ReplyInbox inbox = new InMemoryReplyInbox();
@BeforeEach
void setUp() {
inbox.own("term_a");
}
@Test
void publishThenPeekReturnsTheMessage() {
inbox.publish("term_a", "m1", "hello");
@@ -72,6 +78,7 @@ class InMemoryReplyInboxTest {
@Test
void perTargetIsolation() {
inbox.own("term_b");
inbox.publish("term_a", "m1", "for-a");
inbox.publish("term_b", "m2", "for-b");
assertEquals(1, inbox.peek("term_a").size());
@@ -142,4 +149,40 @@ class InMemoryReplyInboxTest {
exec.shutdown();
}
}
@Test
void publishWithoutOwnDoesNotClaimOwnership() {
String unowned = "term_unowned";
inbox.publish(unowned, "m1", "hello");
// Without an owner, peek returns nothing — publish did not imply consume.
assertTrue(inbox.peek(unowned).isEmpty(),
"publishing to an unowned target must not make it peekable");
// Owning afterwards makes the already-published message visible.
inbox.own(unowned);
var msgs = inbox.peek(unowned);
assertEquals(1, msgs.size());
assertEquals("m1", msgs.getFirst().msgId());
}
@Test
void peekAndAckAreNoOpsForUnownedTarget() {
assertTrue(inbox.peek("term_not_owned").isEmpty());
inbox.ack("term_not_owned", "m1"); // no-op, should not throw
}
@Test
void releaseStopsConsumingAndClearsHeld() {
inbox.publish("term_a", "m1", "hello");
assertEquals(1, inbox.peek("term_a").size());
inbox.release("term_a");
assertTrue(inbox.peek("term_a").isEmpty(),
"after release, the local snapshot is cleared");
}
@Test
void ownIsIdempotent() {
inbox.own("term_a");
inbox.publish("term_a", "m1", "hello");
assertEquals(1, inbox.peek("term_a").size());
}
}
@@ -0,0 +1,216 @@
package dev.ltms.bridged.msg;
import dev.ltms.bridged.herdr.AgentStatus;
import dev.ltms.bridged.mcp.PrimaryRegistry;
import dev.ltms.bridged.peer.MemberRole;
import dev.ltms.bridged.session.MemberSession;
import org.junit.jupiter.api.AfterEach;
import org.junit.jupiter.api.BeforeEach;
import org.junit.jupiter.api.Test;
import java.util.List;
import java.util.concurrent.Executors;
import java.util.concurrent.ScheduledExecutorService;
import java.util.concurrent.TimeUnit;
import static org.junit.jupiter.api.Assertions.*;
/**
* Unit tests for the CB-551 idle-lead heartbeat: the pure {@link LeadHeartbeatLoop#decide} decision
* function, the nudge text, and the {@link LeadHeartbeatLoop#snapshot} fleet snapshot.
*
* <p>All decision tests call {@code decide} directly with explicit nanoTime values from an injected
* clock — no sleeping, no scheduler races. This mirrors how {@code ReplyPushLoopTest} pins the pure
* decision before exercising the loop.
*/
class LeadHeartbeatLoopTest {
/** Arbitrary nanoTime origin for the fake clock. */
private static final long NOW = 1_000_000_000L;
/** 300s (the default quiet period) in nanos — under it the lead is "within the quiet period". */
private static final long IDLE_AFTER_NANOS = TimeUnit.SECONDS.toNanos(300);
/** 400s of idle — clearly past the quiet period. */
private static final long IDLE_PAST = NOW - TimeUnit.SECONDS.toNanos(400);
/** 10s of idle — clearly within the quiet period. */
private static final long IDLE_WITHIN = NOW - TimeUnit.SECONDS.toNanos(10);
private ScheduledExecutorService scheduler;
@BeforeEach
void setUp() {
scheduler = Executors.newSingleThreadScheduledExecutor();
}
@AfterEach
void tearDown() {
scheduler.shutdownNow();
}
/** A fleet with nothing pending. */
private static LeadHeartbeatLoop.FleetState quietFleet() {
return new LeadHeartbeatLoop.FleetState(0, 0, 2, List.of());
}
/** A fleet with a pending reply and a DONE session. */
private static LeadHeartbeatLoop.FleetState pendingFleet() {
return new LeadHeartbeatLoop.FleetState(2, 1, 3, List.of("term_a", "term_b"));
}
/** The loop under test; the scheduler is never invoked on the pure decide path. */
private static LeadHeartbeatLoop loop(int quietCap) {
return new LeadHeartbeatLoop(
new PrimaryRegistry("term_lead"), null /*agents — unused on the decide path*/,
null /*inbox*/, List::of, null /*pushLoop*/, null /*scheduler*/, () -> 0L,
IDLE_AFTER_NANOS, 1_000L, quietCap);
}
// ── (a) a working lead is never injected ───────────────────────────────────────────────────
@Test
void aWorkingLeadIsNeverInjected() {
LeadHeartbeatLoop.Decision d = loop(3).decide(
NOW, IDLE_PAST, 0, AgentStatus.WORKING, false, true, quietFleet());
assertEquals(LeadHeartbeatLoop.Action.LEAD_BUSY, d.action(),
"a WORKING lead is making progress and must not be touched");
assertNull(d.idleSinceNanos(), "a busy lead resets the idle window");
assertEquals(0, d.quietCount(), "a busy lead re-arms the quiet counter");
}
@Test
void anUnreadableStatusIsNeverInjectedEither() {
// A failed status read (or a gone agent) must degrade to "do not inject", never hammer the pane.
LeadHeartbeatLoop.Decision d = loop(3).decide(
NOW, IDLE_PAST, 0, AgentStatus.UNKNOWN, false, true, pendingFleet());
assertEquals(LeadHeartbeatLoop.Action.LEAD_BUSY, d.action(),
"never inject into a state the loop cannot read");
}
// ── (b) an idle lead within the quiet period is not yet injected ───────────────────────────
@Test
void justBecameIdleStartsTheDebounceWindow() {
LeadHeartbeatLoop.Decision d = loop(3).decide(
NOW, null, 0, AgentStatus.IDLE, false, true, quietFleet());
assertEquals(LeadHeartbeatLoop.Action.WAIT_IDLE, d.action(),
"the first injectable tick only records the start of the idle stretch");
assertEquals(NOW, d.idleSinceNanos(), "the idle window opens at the moment the lead became injectable");
}
@Test
void idleWithinQuietPeriodIsNotInjected() {
LeadHeartbeatLoop.Decision d = loop(3).decide(
NOW, IDLE_WITHIN, 0, AgentStatus.IDLE, false, true, quietFleet());
assertEquals(LeadHeartbeatLoop.Action.WAIT_IDLE, d.action(),
"a lead idle for 10s (< 300s) has just finished a turn — do not re-prompt it");
}
// ── (c) an idle lead past the quiet period is injected ─────────────────────────────────────
@Test
void idlePastQuietPeriodIsInjected() {
LeadHeartbeatLoop.Decision d = loop(3).decide(
NOW, IDLE_PAST, 0, AgentStatus.IDLE, false, true, quietFleet());
assertEquals(LeadHeartbeatLoop.Action.INJECT, d.action(),
"a lead continuously idle past the quiet period is the reason to nudge");
}
@Test
void blockedAndDoneAreInjectableViewsOfIdle() {
assertEquals(LeadHeartbeatLoop.Action.INJECT, loop(3).decide(
NOW, IDLE_PAST, 0, AgentStatus.BLOCKED, false, true, quietFleet()).action());
assertEquals(LeadHeartbeatLoop.Action.INJECT, loop(3).decide(
NOW, IDLE_PAST, 0, AgentStatus.DONE, false, true, quietFleet()).action());
}
@Test
void nothingIsInjectedWhenNoLeadIsKnown() {
var d = loop(3).decide(NOW, IDLE_PAST, 0, AgentStatus.IDLE, false, false, pendingFleet());
assertEquals(LeadHeartbeatLoop.Action.WAIT_IDLE, d.action(),
"with no known lead there is nobody to nudge — keep waiting until one is discovered");
assertEquals(IDLE_PAST, d.idleSinceNanos(), "the idle window stays open so discovery re-arms it");
}
// ── (d) the quiet-nudge cap stops the loop ──────────────────────────────────────────────────
@Test
void quietNudgeCapStopsTheLoopWhenNothingIsPending() {
LeadHeartbeatLoop.Decision d = loop(3).decide(
NOW, IDLE_PAST, 3, AgentStatus.IDLE, false, true, quietFleet());
assertEquals(LeadHeartbeatLoop.Action.QUIET_DONE, d.action(),
"3 consecutive nothing-pending nudges have already happened — stop nagging an empty fleet");
}
// ── (e) new pending state resets the cap and re-arms the loop ───────────────────────────────
@Test
void newPendingStateResetsTheQuietCap() {
LeadHeartbeatLoop.Decision d = loop(3).decide(
NOW, IDLE_PAST, 3, AgentStatus.IDLE, false, true, pendingFleet());
assertEquals(LeadHeartbeatLoop.Action.INJECT, d.action(),
"real state appearing re-arms the loop past an exhausted cap");
assertEquals(0, d.quietCount(), "the pending state resets the consecutive-quiet counter");
}
// ── constraint 6: never race ReplyPushLoop ─────────────────────────────────────────────────
@Test
void standsDownWhileReplyPushLoopIsActive() {
LeadHeartbeatLoop.Decision d = loop(3).decide(
NOW, IDLE_PAST, 2, AgentStatus.IDLE, true, true, quietFleet());
assertEquals(LeadHeartbeatLoop.Action.STAND_DOWN, d.action(),
"a second injection would start a competing turn — stand aside instead");
assertEquals(0, d.quietCount(), "the active push is real state, so it re-arms the cap");
}
// ── nudge text (constraint 5: the nudge must carry state) ──────────────────────────────────
@Test
void pendingNudgeCarriesTheFleetFacts() {
String t = pendingFleet().nudgeText();
assertTrue(t.startsWith("Heartbeat:"), "the nudge identifies itself as a heartbeat");
assertTrue(t.contains("2 worker replies pending collection"), t);
assertTrue(t.contains("bridge_poll(target=term_a)"), t);
assertTrue(t.contains("bridge_poll(target=term_b)"), t);
assertTrue(t.contains("1 DONE session awaiting teardown"), t);
assertTrue(t.contains("3 workers live"), t);
}
@Test
void nothingPendingNudgeSaysExactlyThatSoTheLeadCanStandDown() {
String t = quietFleet().nudgeText();
assertTrue(t.contains("nothing pending to collect"), t);
assertTrue(t.contains("0 worker replies"), t);
assertTrue(t.contains("0 DONE sessions"), t);
assertTrue(t.contains("2 workers live"), t);
assertTrue(t.contains("stand down"), "allow the lead to stand down rather than hunt");
}
// ── snapshot over roster + inbox ───────────────────────────────────────────────────────────
@Test
void snapshotCountsRepliesDoneSessionsAndLiveWorkers() {
InMemoryReplyInbox inbox = new InMemoryReplyInbox();
inbox.own("term_w1");
inbox.publish("term_w1", "m1", "hello");
MemberSession done = new MemberSession("p1", "term_w1", "prof", MemberRole.DEV, "/cwd", null,
0, 0, 0, MemberSession.State.DONE, null, null);
MemberSession ready = new MemberSession("p2", "term_w2", "prof", MemberRole.DEV, "/cwd", null,
0, 0, 0, MemberSession.State.READY, null, null);
LeadHeartbeatLoop.FleetState fs = LeadHeartbeatLoop.snapshot(inbox, () -> List.of(done, ready));
assertEquals(1, fs.pendingReplies(), "one undrained reply on term_w1");
assertEquals(1, fs.doneSessions(), "term_w1 is DONE, awaiting teardown");
assertEquals(2, fs.liveWorkers(), "both sessions are still registered");
assertEquals(List.of("term_w1"), fs.replyTargets(), "the target with a pending reply is named");
assertTrue(fs.hasPending(), "a pending reply counts as fleet state to collect");
}
@Test
void snapshotWithEmptyRosterHasNoPendingState() {
InMemoryReplyInbox inbox = new InMemoryReplyInbox();
LeadHeartbeatLoop.FleetState fs = LeadHeartbeatLoop.snapshot(inbox, List::of);
assertFalse(fs.hasPending());
assertEquals(0, fs.liveWorkers());
}
}
@@ -5,7 +5,9 @@ import dev.ltms.bridged.herdr.AgentStatus;
import dev.ltms.bridged.herdr.FakeHerdr;
import dev.ltms.bridged.herdr.HerdrException;
import dev.ltms.bridged.inject.CompletionResolver;
import dev.ltms.bridged.mcp.PrimaryRegistry;
import dev.ltms.bridged.inject.Injector;
import org.junit.jupiter.api.BeforeEach;
import org.junit.jupiter.api.Test;
import java.util.concurrent.CompletableFuture;
@@ -15,6 +17,7 @@ import static org.junit.jupiter.api.Assertions.assertEquals;
import static org.junit.jupiter.api.Assertions.assertFalse;
import static org.junit.jupiter.api.Assertions.assertNotNull;
import static org.junit.jupiter.api.Assertions.assertNull;
import static org.junit.jupiter.api.Assertions.assertThrows;
import static org.junit.jupiter.api.Assertions.assertTrue;
/**
@@ -30,7 +33,14 @@ class MessageServiceTest {
private final Rendezvous rendezvous = new Rendezvous();
private final CompletionResolver completion = new CompletionResolver(agents, rendezvous);
private final Injector injector = new Injector(agents, completion);
private final MessageService messages = new MessageService(agents, injector, rendezvous);
private final InMemoryReplyInbox inbox = new InMemoryReplyInbox();
private final MessageService messages = new MessageService(agents, injector, rendezvous, inbox);
@BeforeEach
void setUp() {
// CB-520: the inbox only peeks/acks targets it owns.
inbox.own(T);
}
/** Run {@code send} on a background thread; the current thread drives the worker's turn. */
private CompletableFuture<MessageService.Reply> sendAsync() {
@@ -313,6 +323,141 @@ class MessageServiceTest {
assertEquals("first done", firstReply.text());
}
// --- CB-548: delegator ownership is recorded only on an ACCEPTED send ----------------------
private static final String LEAD_L = "term_lead_l";
private static final String LEAD_A = "term_lead_a";
/**
* The bug CB-548 fixes: L holds worker W, then architect A attempts W and times out BUSY. With
* delegator ownership recorded at {@code bridge_send} <em>request</em> time, A's rejected call
* would overwrite L — and W's late no-waiter reply would be pushed to A, who never owned the
* turn. The accepted-delivery hook must not fire for a BUSY send, so L stays the delegator.
*/
@Test
void busySenderDoesNotBecomeTheDelegatingOwner() throws Exception {
PrimaryRegistry reg = new PrimaryRegistry(null);
// L accepts a delegation to W: the send wins the lock and queues delivery → L is recorded.
CompletableFuture<MessageService.Reply> first = CompletableFuture.supplyAsync(
() -> messages.send(T, "first", 5000, () -> reg.recordDelegation(T, LEAD_L)));
awaitWaiting();
assertEquals(LEAD_L, reg.nudgeTargetFor(T).orElseThrow(),
"an accepted send owns the delegation");
// A attempts W while L holds it → BUSY (lock never taken) → its hook never fires.
MessageService.Reply busy = messages.send(T, "second", 100, () -> reg.recordDelegation(T, LEAD_A));
assertEquals(MessageService.Outcome.BUSY, busy.outcome());
assertEquals(LEAD_L, reg.nudgeTargetFor(T).orElseThrow(),
"a BUSY send must not steal the delegator ownership it never earned");
// L completes so the test thread is not left pinned.
injector.onStatus(T, AgentStatus.IDLE);
injector.onStatus(T, AgentStatus.WORKING);
assertTrue(rendezvous.resolve(T, "first done"));
MessageService.Reply firstReply = first.get(5, TimeUnit.SECONDS);
assertEquals(MessageService.Outcome.REPLIED, firstReply.outcome());
}
/**
* Once L's accepted delegation is fully done, a later <em>accepted</em> send from A may
* legitimately become the new delegator — ownership follows the turn, not the first caller.
*/
@Test
void anAcceptedSendAfterThePriorOwnerFinishesBecomesTheNewOwner() throws Exception {
PrimaryRegistry reg = new PrimaryRegistry(null);
CompletableFuture<MessageService.Reply> first = CompletableFuture.supplyAsync(
() -> messages.send(T, "first", 5000, () -> reg.recordDelegation(T, LEAD_L)));
awaitWaiting();
injector.onStatus(T, AgentStatus.IDLE);
injector.onStatus(T, AgentStatus.WORKING);
assertTrue(rendezvous.resolve(T, "first done"));
MessageService.Reply firstReply = first.get(5, TimeUnit.SECONDS);
assertEquals(MessageService.Outcome.REPLIED, firstReply.outcome());
assertEquals(LEAD_L, reg.nudgeTargetFor(T).orElseThrow(), "L owned the first turn");
// L finished; A's later accepted send takes the delegation over.
CompletableFuture<MessageService.Reply> second = CompletableFuture.supplyAsync(
() -> messages.send(T, "second", 5000, () -> reg.recordDelegation(T, LEAD_A)));
awaitWaiting();
assertEquals(LEAD_A, reg.nudgeTargetFor(T).orElseThrow(),
"an accepted send after the owner finished becomes the new delegator");
injector.onStatus(T, AgentStatus.IDLE);
injector.onStatus(T, AgentStatus.WORKING);
assertTrue(rendezvous.resolve(T, "second done"));
MessageService.Reply secondReply = second.get(5, TimeUnit.SECONDS);
assertEquals(MessageService.Outcome.REPLIED, secondReply.outcome());
}
/**
* CB-548 requirement: answering an existing {@code bridge_ask} is the SAME delegation, so it must
* not rewrite ownership. L accepted the send (owned), the worker paused to ask, and L answers via
* turnId — ownership stays L throughout; the answer path never touches the registry.
*/
@Test
void answeringAnAskDoesNotRewriteDelegatorOwnership() throws Exception {
PrimaryRegistry reg = new PrimaryRegistry(null);
CompletableFuture<MessageService.Reply> send = CompletableFuture.supplyAsync(
() -> messages.send(T, "do X", 5000, () -> reg.recordDelegation(T, LEAD_L)));
awaitWaiting();
assertEquals(LEAD_L, reg.nudgeTargetFor(T).orElseThrow(), "L owns the delegation");
injector.onStatus(T, AgentStatus.IDLE); // deliver
injector.onStatus(T, AgentStatus.WORKING); // worker picks it up, then pauses to ask
CompletableFuture<MessageService.AskResult> ask =
CompletableFuture.supplyAsync(() -> messages.ask(T, "which config?", 5000));
MessageService.Reply q = send.get(5, TimeUnit.SECONDS);
assertEquals(MessageService.Outcome.QUESTION, q.outcome());
assertNotNull(q.turnId());
// L answers the ask on the same turn; the answer path must not touch ownership.
CompletableFuture<MessageService.Reply> answer =
CompletableFuture.supplyAsync(() -> messages.answer(q.turnId(), "config.yaml", 5000));
assertEquals("config.yaml", ask.get(5, TimeUnit.SECONDS).answer());
awaitWaiting(); // the answering send reopened its forward waiter
assertEquals(LEAD_L, reg.nudgeTargetFor(T).orElseThrow(),
"answering an ask keeps L as the delegator — ownership is not rewritten");
assertTrue(rendezvous.resolve(T, "done"));
assertEquals(MessageService.Outcome.REPLIED, answer.get(5, TimeUnit.SECONDS).outcome());
}
/**
* CB-548: {@code onAccepted} is a public callback, so a throwing one must not orphan the turn.
* The waiter is opened first, then the hook runs BEFORE delivery is queued — so a throw fails
* the send loudly, closes its waiter, and never enqueues a message the worker would pick up and
* reply into the void.
*/
@Test
void aThrowingAcceptedHookLeavesNoStaleWaiterOrQueuedOrphan() {
assertThrows(IllegalStateException.class,
() -> messages.send(T, "doomed", 500,
() -> { throw new IllegalStateException("ownership hook failed"); }),
"a throwing ownership hook fails the send loudly");
assertFalse(rendezvous.isWaiting(T), "the failed send must not leave a stale rendezvous waiter");
// Give the injector a delivery window: with nothing enqueued, nothing may reach the worker.
injector.onStatus(T, AgentStatus.IDLE);
boolean doomedQueued = herdr.calls.stream()
.anyMatch(c -> c.method().equals("agent.prompt")
&& String.valueOf(c.params()).contains("doomed"));
assertFalse(doomedQueued, "a throwing ownership hook must not leave a queued, orphanable message");
}
/**
* The async (fire-and-poll) path runs the same {@code send} on a background thread, so the
* accepted-delivery hook must thread through it — ownership is recorded exactly as blocking sends.
*/
@Test
void asyncSendRecordsOwnershipOnAcceptance() throws Exception {
PrimaryRegistry reg = new PrimaryRegistry(null);
messages.sendAsync(T, "async task", () -> reg.recordDelegation(T, LEAD_L));
awaitWaiting(); // the background send won the lock, queued, and opened its waiter
assertEquals(LEAD_L, reg.nudgeTargetFor(T).orElseThrow(),
"the async path records delegator ownership on acceptance, like the blocking path");
}
// --- CB-307 reply inbox ----------------------------------------------------------------
@Test
@@ -8,6 +8,8 @@ import static org.junit.jupiter.api.Assertions.assertEquals;
import static org.junit.jupiter.api.Assertions.assertFalse;
import static org.junit.jupiter.api.Assertions.assertNotEquals;
import static org.junit.jupiter.api.Assertions.assertNull;
import static org.junit.jupiter.api.Assertions.assertSame;
import static org.junit.jupiter.api.Assertions.assertThrows;
import static org.junit.jupiter.api.Assertions.assertTrue;
/**
@@ -20,6 +22,26 @@ class RendezvousTest {
private final Rendezvous rendezvous = new Rendezvous();
/**
* CB-548: an {@code open} is atomic fail-if-present so a double open can never replace the first
* waiter another send is blocked on. {@code MessageService} serializes sends per session, so in
* correct code this cannot happen — the rejection is a loud tripwire for an invariant violation,
* and the first waiter must survive it and stay resolvable.
*/
@Test
void openRejectsADoubleOpenAndKeepsTheFirstWaiterRegisteredAndResolvable() {
CompletableFuture<Rendezvous.Resolution> first = rendezvous.open(W);
assertThrows(IllegalStateException.class, () -> rendezvous.open(W),
"a second open while one is registered is rejected loudly, not a silent replace");
assertSame(first, rendezvous.currentWaiter(W), "the first waiter remains the registered one");
assertTrue(rendezvous.resolve(W, "first wins"), "the first waiter is still resolvable");
Rendezvous.Resolution r = first.getNow(null);
assertEquals(Rendezvous.Kind.REPLY, r.kind());
assertEquals("first wins", r.text(), "the resolution lands on the first waiter, not the rejected one");
}
@Test
void openAskMintsAUniqueTurnScopedToItsSessionAndCoalescesDuplicates() {
Rendezvous.AskTicket t1 = rendezvous.openAsk(W);
@@ -45,6 +45,7 @@ class ReplyPushLoopTest {
void setUp() {
registry = new PrimaryRegistry(PRIMARY);
inbox = new InMemoryReplyInbox();
inbox.own(WORKER); // CB-520: the inbox only peeks/acks targets it owns
scheduler = Executors.newSingleThreadScheduledExecutor();
}
@@ -133,10 +134,10 @@ class ReplyPushLoopTest {
loop(1, 50).onReplyQueued(WORKER);
assertTrue(rec.sendLatch.await(3, TimeUnit.SECONDS),
"one nudge (2 agent.send calls) should have been sent");
"one nudge (1 agent.prompt call) should have been sent");
// Exactly one nudge = exactly 2 agent.send calls (text + submit)
assertEquals(2, rec.sendCount());
// Exactly one nudge = exactly 1 agent.prompt call (it submits itself)
assertEquals(1, rec.sendCount());
assertTrue(rec.sentParams().stream()
.anyMatch(e -> e.getValue().toString().contains("bridge_poll")),
"nudge text should contain bridge_poll");
@@ -153,9 +154,9 @@ class ReplyPushLoopTest {
loop.onReplyQueued(WORKER); // second call — should be a no-op
assertTrue(rec.sendLatch.await(3, TimeUnit.SECONDS),
"expected exactly one nudge (2 sends)");
"expected exactly one nudge (1 prompt)");
Thread.sleep(200);
assertEquals(2, rec.sendCount(),
assertEquals(1, rec.sendCount(),
"second onReplyQueued must not trigger another nudge");
}
@@ -165,19 +166,33 @@ class ReplyPushLoopTest {
var rec = recordingClient();
agents = new AgentControl(rec);
inbox.publish(WORKER, "m1", "hello");
rec.sendLatch = new CountDownLatch(cap * 2);
rec.sendLatch = new CountDownLatch(cap);
loop(cap, 50).onReplyQueued(WORKER);
assertTrue(rec.sendLatch.await(5, TimeUnit.SECONDS),
cap + " nudges (" + (cap * 2) + " sends) should have fired");
cap + " nudges (" + cap + " prompts) should have fired");
Thread.sleep(300);
assertEquals(cap * 2, rec.sendCount(),
"exactly " + (cap * 2) + " agent.send calls (cap=" + cap + ")");
assertEquals(cap, rec.sendCount(),
"exactly " + cap + " agent.prompt calls (cap=" + cap + ")");
}
// --- nudge format --------------------------------------------------------------------------
@Test
void isActiveReflectsALiveScheduleForTheHeartbeatStandDown() {
var rec = recordingClient();
agents = new AgentControl(rec);
inbox.publish(WORKER, "m1", "hello");
ReplyPushLoop loop = loop(1, 100_000); // long backoff so the tick cannot fire mid-test
loop.onReplyQueued(WORKER);
assertTrue(loop.isActive(), "CB-551: the heartbeat must stand aside while a reminder is live");
loop.stop();
assertFalse(loop.isActive(), "stopping clears the active schedule");
}
@Test
void nudgeFormatIsCorrect() {
String nudge = ReplyPushLoop.NUDGE_FORMAT.formatted(WORKER, WORKER);
@@ -197,7 +212,7 @@ class ReplyPushLoopTest {
loop(1, 50, metrics).onReplyQueued(WORKER);
assertTrue(rec.sendLatch.await(3, TimeUnit.SECONDS),
"one nudge (2 agent.send calls) should have been sent");
"one nudge (1 agent.prompt call) should have been sent");
// The delivered count is bumped on the scheduler thread right after the send that releases
// the latch — settle briefly so the counter is published before we read it.
Thread.sleep(200);
@@ -261,13 +276,13 @@ class ReplyPushLoopTest {
}
/**
* Thread-safe recording fake that counts agent.send calls. Uses synchronized access
* so the scheduler thread and test thread never race.
* Thread-safe recording fake that counts agent.prompt calls (protocol 19: one nudge = one
* prompt). Uses synchronized access so the scheduler thread and test thread never race.
*/
private static final class RecordingHerdrClient implements HerdrClient {
private final List<Map.Entry<String, Object>> calls =
Collections.synchronizedList(new ArrayList<>());
volatile CountDownLatch sendLatch = new CountDownLatch(2);
volatile CountDownLatch sendLatch = new CountDownLatch(1);
@Override
public JsonNode call(String method, Object params) {
@@ -277,7 +292,7 @@ class ReplyPushLoopTest {
.put("terminal_id", PRIMARY)
.put("agent_status", "idle")); // recording double is always injectable
}
if ("agent.send".equals(method)) {
if ("agent.prompt".equals(method)) {
calls.add(Map.entry(method, params));
sendLatch.countDown();
}
@@ -0,0 +1,67 @@
package dev.ltms.bridged.peer;
import org.junit.jupiter.api.Test;
import static org.junit.jupiter.api.Assertions.assertEquals;
import static org.junit.jupiter.api.Assertions.assertSame;
import static org.junit.jupiter.api.Assertions.assertThrows;
import static org.junit.jupiter.api.Assertions.assertTrue;
class MemberRoleTest {
@Test
void theThreeRolesAreArchitectDevAndReviewer() {
assertEquals(3, MemberRole.values().length,
"a new role changes the charter, the role file, the skill and the authz row — "
+ "adding one is a deliberate act, so this count is meant to fail first");
assertEquals("architect", MemberRole.ARCHITECT.wireName());
assertEquals("dev", MemberRole.DEV.wireName());
assertEquals("reviewer", MemberRole.REVIEWER.wireName());
}
@Test
void parseAcceptsTheWireSpelling() {
for (MemberRole r : MemberRole.values()) {
assertSame(r, MemberRole.parse(r.wireName()));
}
}
@Test
void parseIsCaseInsensitiveAndTrimsSurroundingSpace() {
assertSame(MemberRole.ARCHITECT, MemberRole.parse("Architect"));
assertSame(MemberRole.DEV, MemberRole.parse(" DEV "));
assertSame(MemberRole.REVIEWER, MemberRole.parse("ReViEwEr"));
}
@Test
void parseRejectsAnUnknownRoleAndListsTheValidOnes() {
IllegalArgumentException e =
assertThrows(IllegalArgumentException.class, () -> MemberRole.parse("archtiect"));
assertTrue(e.getMessage().contains("archtiect"), e.getMessage());
assertTrue(e.getMessage().contains("architect, dev, reviewer"),
"a typo in config should be fixable from the message alone: " + e.getMessage());
}
@Test
void parseRejectsNullAndBlank() {
assertThrows(IllegalArgumentException.class, () -> MemberRole.parse(null));
assertThrows(IllegalArgumentException.class, () -> MemberRole.parse(""));
assertThrows(IllegalArgumentException.class, () -> MemberRole.parse(" "));
}
/**
* A lead is not a member role. A lead is not spawned — it is a pre-existing session that config
* recognises — so it has no launch charter to pick and never appears in a {@code members:} slot.
*/
@Test
void leadIsNotAMemberRole() {
assertThrows(IllegalArgumentException.class, () -> MemberRole.parse("lead"));
assertThrows(IllegalArgumentException.class, () -> MemberRole.parse("primary"));
}
/** "worker" is the name this taxonomy replaced; it must not quietly keep working. */
@Test
void workerIsNoLongerARole() {
assertThrows(IllegalArgumentException.class, () -> MemberRole.parse("worker"));
}
}
@@ -0,0 +1,183 @@
package dev.ltms.bridged.placement;
import org.junit.jupiter.api.Test;
import java.util.HashSet;
import java.util.List;
import java.util.Map;
import java.util.Set;
import java.util.function.Function;
import static org.junit.jupiter.api.Assertions.*;
/**
* Unit tests for the placement policies. They run with no herdr and no launcher — pure selection
* logic exercised through the descriptor type so CB-308 host expansion will not need to rewrite
* these assertions.
*/
class PlacementPolicyTest {
private static Function<String, Integer> noSessions() {
return name -> 0;
}
private static PlacementContext ctx(List<PlacementCandidate> candidates,
Function<String, Integer> liveCount,
Set<String> unreachable) {
return new PlacementContext("b", candidates, liveCount, unreachable);
}
private static PlacementContext ctx(List<PlacementCandidate> candidates,
Function<String, Integer> liveCount) {
return ctx(candidates, liveCount, Set.of());
}
@Test
void fixedReturnsDefaultEvenIfOtherProfilesExist() {
PlacementPolicy policy = PlacementPolicies.fixed();
PlacementContext ctx = ctx(List.of(
PlacementCandidate.profile("a"),
PlacementCandidate.profile("b")), noSessions());
assertEquals("b", policy.select(ctx).profile());
}
@Test
void fixedFallsBackToFirstCandidateWhenNoDefault() {
PlacementPolicy policy = PlacementPolicies.fixed();
PlacementContext ctx = new PlacementContext(null,
List.of(PlacementCandidate.profile("a"), PlacementCandidate.profile("b")),
noSessions(), Set.of());
assertEquals("a", policy.select(ctx).profile());
}
@Test
void fixedThrowsWhenNoProfilesAndNoDefault() {
PlacementPolicy policy = PlacementPolicies.fixed();
PlacementContext ctx = new PlacementContext(null, List.of(), noSessions(), Set.of());
assertThrows(PlacementException.class, () -> policy.select(ctx));
}
@Test
void roundRobinCyclesThroughAvailableProfiles() {
PlacementPolicy policy = PlacementPolicies.roundRobin();
List<PlacementCandidate> candidates = List.of(
PlacementCandidate.profile("a"),
PlacementCandidate.profile("b"),
PlacementCandidate.profile("c"));
assertEquals("a", policy.select(ctx(candidates, noSessions())).profile());
assertEquals("b", policy.select(ctx(candidates, noSessions())).profile());
assertEquals("c", policy.select(ctx(candidates, noSessions())).profile());
assertEquals("a", policy.select(ctx(candidates, noSessions())).profile());
}
@Test
void roundRobinSkipsProfilesAtMaxLoad() {
PlacementPolicy policy = PlacementPolicies.roundRobin();
List<PlacementCandidate> candidates = List.of(
PlacementCandidate.profile("a", 1.0f, 2),
PlacementCandidate.profile("b", 1.0f, null));
Function<String, Integer> liveCount = Map.of("a", 2)::get;
for (int i = 0; i < 5; i++) {
assertEquals("b", policy.select(ctx(candidates, liveCount)).profile());
}
}
@Test
void roundRobinThrowsWhenAllAtMaxLoad() {
PlacementPolicy policy = PlacementPolicies.roundRobin();
List<PlacementCandidate> candidates = List.of(
PlacementCandidate.profile("a", 1.0f, 1),
PlacementCandidate.profile("b", 1.0f, 1));
Function<String, Integer> liveCount = Map.of("a", 1, "b", 1)::get;
PlacementException e = assertThrows(PlacementException.class,
() -> policy.select(ctx(candidates, liveCount)));
assertTrue(e.getMessage().contains("maxLoad"), e.getMessage());
}
@Test
void weightedAlternatesEvenlyWithEqualWeights() {
PlacementPolicy policy = PlacementPolicies.weighted();
List<PlacementCandidate> candidates = List.of(
PlacementCandidate.profile("a", 0.5f, null),
PlacementCandidate.profile("b", 0.5f, null));
int a = 0, b = 0;
for (int i = 0; i < 100; i++) {
String p = policy.select(ctx(candidates, noSessions())).profile();
if ("a".equals(p)) a++;
else if ("b".equals(p)) b++;
}
assertEquals(50, a, "equal weights should split 50/50");
assertEquals(50, b);
}
@Test
void weightedHoldsThreeToOneRatio() {
PlacementPolicy policy = PlacementPolicies.weighted();
List<PlacementCandidate> candidates = List.of(
PlacementCandidate.profile("a", 0.75f, null),
PlacementCandidate.profile("b", 0.25f, null));
int a = 0, b = 0;
for (int i = 0; i < 40; i++) {
String p = policy.select(ctx(candidates, noSessions())).profile();
if ("a".equals(p)) a++;
else if ("b".equals(p)) b++;
}
assertEquals(30, a, "0.75/0.25 should yield a 3:1 ratio over a multiple of 4");
assertEquals(10, b);
}
@Test
void weightedSkipsProfileAtMaxLoad() {
PlacementPolicy policy = PlacementPolicies.weighted();
List<PlacementCandidate> candidates = List.of(
PlacementCandidate.profile("a", 1.0f, 1),
PlacementCandidate.profile("b", 1.0f, null));
Function<String, Integer> liveCount = name -> "a".equals(name) ? 1 : 0;
for (int i = 0; i < 5; i++) {
assertEquals("b", policy.select(ctx(candidates, liveCount)).profile());
}
}
@Test
void weightedThrowsWhenAllAtMaxLoad() {
PlacementPolicy policy = PlacementPolicies.weighted();
List<PlacementCandidate> candidates = List.of(
PlacementCandidate.profile("a", 1.0f, 1),
PlacementCandidate.profile("b", 1.0f, 1));
Function<String, Integer> liveCount = name -> 1;
PlacementException e = assertThrows(PlacementException.class,
() -> policy.select(ctx(candidates, liveCount)));
assertTrue(e.getMessage().contains("maxLoad"), e.getMessage());
}
@Test
void weightedThrowsWhenAllUnreachable() {
PlacementPolicy policy = PlacementPolicies.weighted();
List<PlacementCandidate> candidates = List.of(
PlacementCandidate.profile("a"),
PlacementCandidate.profile("b"));
PlacementException e = assertThrows(PlacementException.class,
() -> policy.select(ctx(candidates, noSessions(), Set.of("a", "b"))));
assertTrue(e.getMessage().contains("unreachable"), e.getMessage());
}
@Test
void mixedExclusionMessageNamesBothReasons() {
PlacementPolicy policy = PlacementPolicies.weighted();
List<PlacementCandidate> candidates = List.of(
PlacementCandidate.profile("a", 1.0f, 1),
PlacementCandidate.profile("b", 1.0f, null));
Function<String, Integer> liveCount = name -> "a".equals(name) ? 1 : 0;
Set<String> unreachable = new HashSet<>();
unreachable.add("b");
PlacementException e = assertThrows(PlacementException.class,
() -> policy.select(ctx(candidates, liveCount, unreachable)));
assertTrue(e.getMessage().contains("1 at maxLoad"), e.getMessage());
assertTrue(e.getMessage().contains("1 unreachable"), e.getMessage());
}
@Test
void unknownPolicyNameThrows() {
assertThrows(IllegalArgumentException.class, () -> PlacementPolicies.fromName("random"));
}
}
@@ -1,6 +1,7 @@
package dev.ltms.bridged.rest;
import dev.ltms.bridged.auth.CallerResolver;
import dev.ltms.bridged.auth.MemberRegistry;
import dev.ltms.bridged.config.BridgedConfig;
import dev.ltms.bridged.guard.SubscriptionGuard;
import dev.ltms.bridged.herdr.AgentControl;
@@ -15,7 +16,7 @@ import dev.ltms.bridged.msg.MessageService;
import dev.ltms.bridged.msg.Rendezvous;
import dev.ltms.bridged.session.FakeWorktrees;
import dev.ltms.bridged.session.SessionManager;
import dev.ltms.bridged.worker.ClaudeCodeLauncher;
import dev.ltms.bridged.member.ClaudeCodeLauncher;
import io.javalin.Javalin;
import org.junit.jupiter.api.AfterEach;
import org.junit.jupiter.api.Test;
@@ -52,7 +53,7 @@ class BridgedAppAuthTest {
*/
private int start(long pid, boolean tokenMode, String token) {
FakeHerdr herdr = new FakeHerdr();
BridgedConfig.Worker wcfg = new BridgedConfig.Worker(
BridgedConfig.Profile wcfg = new BridgedConfig.Profile(
"ltms-local", "http://gx00.gw:8000", "coder", null, "BRIDGED_WORKER_TOKEN", null,
"tab", "bridged-workers", "worker: {profile} #{n}", null, null, null);
AgentControl agents = new AgentControl(herdr);
@@ -65,9 +66,8 @@ class BridgedAppAuthTest {
MessageService messages = new MessageService(agents, injector, new Rendezvous());
ConnectionIdentity identity = new ConnectionIdentity(new PaneLocator(herdr), _ -> pid);
CallerResolver callers = tokenMode
? new CallerResolver(identity, true, token)
: new CallerResolver(identity);
CallerResolver callers = CallerResolver.withLeadsAndMembers(identity, tokenMode, token,
Map::of, new MemberRegistry(null));
metrics = BridgedMetrics.create(sessions, new dev.ltms.bridged.msg.InMemoryReplyInbox());
app = new BridgedApp(herdr, workers, sessions, messages, sessions.asPresence(), null,
@@ -128,9 +128,9 @@ class BridgedAppAuthTest {
void aWorkerMayNotOrchestrate() throws Exception {
int port = start(FakeHerdr.WORKER_PID, false, null);
assertEquals(403, send(port, "POST", "/workers", null, null).statusCode(),
assertEquals(403, send(port, "POST", "/members", null, null).statusCode(),
"a worker spawning workers would be escalating into the orchestrator role");
assertEquals(403, send(port, "DELETE", "/workers/w2:p7", null, null).statusCode());
assertEquals(403, send(port, "DELETE", "/members/w2:p7", null, null).statusCode());
assertEquals(403, send(port, "POST", "/sessions/term_b/message",
"{\"content\":\"hi\"}", null).statusCode());
assertEquals(403, send(port, "GET", "/sessions/term_a/replies", null, null).statusCode(),
@@ -206,7 +206,7 @@ class BridgedAppAuthTest {
// The 29 pre-existing acceptance tests rely on this: no auth fixture, no enforcement.
FakeHerdr herdr = new FakeHerdr();
AgentControl agents = new AgentControl(herdr);
BridgedConfig.Worker wcfg = new BridgedConfig.Worker(
BridgedConfig.Profile wcfg = new BridgedConfig.Profile(
"ltms-local", "http://gx00.gw:8000", "coder", null, "BRIDGED_WORKER_TOKEN", null,
"tab", "bridged-workers", "worker: {profile} #{n}", null, null, null);
ClaudeCodeLauncher workers = new ClaudeCodeLauncher(
@@ -9,14 +9,15 @@ import dev.ltms.bridged.herdr.FakeHerdr;
import dev.ltms.bridged.herdr.WorkspaceControl;
import dev.ltms.bridged.inject.Injector;
import dev.ltms.bridged.inject.StatusPoller;
import dev.ltms.bridged.inject.WorkerPresence;
import dev.ltms.bridged.inject.MemberPresence;
import dev.ltms.bridged.msg.InMemoryReplyInbox;
import dev.ltms.bridged.msg.MessageService;
import dev.ltms.bridged.msg.Rendezvous;
import dev.ltms.bridged.session.FakeWorktrees;
import dev.ltms.bridged.session.GitWorktrees;
import dev.ltms.bridged.session.SessionManager;
import dev.ltms.bridged.session.Worktrees;
import dev.ltms.bridged.worker.ClaudeCodeLauncher;
import dev.ltms.bridged.member.ClaudeCodeLauncher;
import io.javalin.Javalin;
import org.junit.jupiter.api.AfterEach;
import org.junit.jupiter.api.Test;
@@ -28,6 +29,7 @@ import java.net.http.HttpResponse;
import java.util.List;
import java.util.Map;
import java.util.Set;
import java.util.UUID;
import static org.junit.jupiter.api.Assertions.*;
@@ -41,7 +43,7 @@ class BridgedAppTest {
private final ObjectMapper mapper = new ObjectMapper();
private final HttpClient http = HttpClient.newHttpClient();
private WorkerPresence presence;
private MemberPresence presence;
private Javalin app;
private StatusPoller poller;
@@ -60,7 +62,7 @@ class BridgedAppTest {
}
private int start(FakeHerdr herdr, String workerBaseUrl, Set<String> allow, String placement, Worktrees worktrees) {
BridgedConfig.Worker wcfg = new BridgedConfig.Worker(
BridgedConfig.Profile wcfg = new BridgedConfig.Profile(
"ltms-local", workerBaseUrl, "coder", null, "BRIDGED_WORKER_TOKEN", null,
placement, "bridged-workers", "worker: {profile} #{n}", null, null, null);
AgentControl agents = new AgentControl(herdr);
@@ -74,7 +76,12 @@ class BridgedAppTest {
poller = new StatusPoller(agents, injector, 5); // delivers when the fake reports idle
poller.start();
Rendezvous rendezvous = new Rendezvous();
MessageService messages = new MessageService(agents, injector, rendezvous);
InMemoryReplyInbox inbox = new InMemoryReplyInbox();
sessions.onAcquire(inbox::own);
// The static REST tests address "term_a" without acquiring it through SessionManager, so own
// it directly so the inbox contract holds for those endpoints.
inbox.own("term_a");
MessageService messages = new MessageService(agents, injector, rendezvous, inbox);
app = new BridgedApp(herdr, workers, sessions, messages, this.presence, null)
.build().start("127.0.0.1", 0);
return app.port();
@@ -118,7 +125,7 @@ class BridgedAppTest {
assertEquals(200, res.statusCode());
JsonNode body = mapper.readTree(res.body());
assertEquals("ok", body.get("status").asText());
assertEquals(14, body.get("herdr").get("protocol").asInt());
assertEquals(19, body.get("herdr").get("protocol").asInt());
}
@Test
@@ -153,25 +160,30 @@ class BridgedAppTest {
FakeHerdr herdr = new FakeHerdr();
int port = start(herdr, "http://gx00.gw:8000", Set.of("gx00.gw"));
HttpResponse<String> res = req(port, "POST", "/workers");
HttpResponse<String> res = req(port, "POST", "/members");
assertEquals(201, res.statusCode());
JsonNode body = mapper.readTree(res.body());
assertEquals("w9:pW_1", body.get("paneId").asText());
// CB-519: the responded paneId is a host-unique opaque UUID, not the herdr pane coordinate.
String id = body.get("paneId").asText();
assertNotEquals("w9:pRoot_1", id, "paneId is the host-unique id, not the herdr pane");
assertDoesNotThrow(() -> UUID.fromString(id), "paneId must be a UUID: " + id);
assertEquals("spawning", body.get("state").asText());
// Subscription boundary: agent.start carried base_url + token in its env map.
Map<String, Object> start = params(herdr, "agent.start");
// Subscription boundary (protocol 19): tab.create carried base_url + token in its env
// map — the seed shell the agent starts into is what inherits them.
Map<String, Object> create = params(herdr, "tab.create");
@SuppressWarnings("unchecked")
Map<String, String> env = (Map<String, String>) start.get("env");
Map<String, String> env = (Map<String, String>) create.get("env");
assertEquals("http://gx00.gw:8000", env.get("ANTHROPIC_BASE_URL"));
assertEquals("tok-abc", env.get("ANTHROPIC_AUTH_TOKEN"));
assertEquals(List.of("claude"), start.get("argv"));
// Placement: worker space ensured, worker started INTO its own tab, seed shell
// dropped, and the tab given a friendly label.
// Placement: worker space ensured, worker started INTO its tab's seed pane (which
// becomes the worker pane — nothing is dropped), and the tab given a friendly label.
Map<String, Object> start = params(herdr, "agent.start");
assertTrue(herdr.called("workspace.create"), "worker space must be found-or-created");
assertEquals("w9:t2", start.get("tab_id"), "worker must start into its dedicated tab");
assertEquals("w9:pRoot", params(herdr, "pane.close").get("pane_id"), "seed shell pane dropped");
assertEquals("claude", start.get("kind"), "herdr resolves the executable from kind");
assertEquals("w9:pRoot_1", start.get("pane_id"), "worker must start into its tab's seed pane");
assertFalse(herdr.called("pane.close"), "the seed pane IS the worker pane — never dropped");
assertEquals("worker: ltms-local #1", params(herdr, "tab.rename").get("label"),
"tab label carries the worker number so siblings stay distinct");
}
@@ -190,12 +202,12 @@ class BridgedAppTest {
int port = start(herdr, "http://gx00.gw:8000", Set.of("gx00.gw"), "tab",
new FakeWorktrees().withRepoRoot("/repo").withPrefix("/wt"));
HttpResponse<String> spawn = req(port, "POST", "/workers?worktree=true&ticket=cb-304");
HttpResponse<String> spawn = req(port, "POST", "/members?worktree=true&ticket=cb-304");
assertEquals(201, spawn.statusCode());
JsonNode spawned = mapper.readTree(spawn.body());
String paneId = spawned.get("paneId").asText();
HttpResponse<String> res = req(port, "GET", "/workers");
HttpResponse<String> res = req(port, "GET", "/members");
assertEquals(200, res.statusCode());
JsonNode workers = mapper.readTree(res.body()).get("workers");
assertEquals(1, workers.size());
@@ -214,15 +226,15 @@ class BridgedAppTest {
void spawnWithACwdParamRootsTheWorkerThere() throws Exception {
FakeHerdr herdr = new FakeHerdr();
int port = start(herdr, "http://gx00.gw:8000", Set.of("gx00.gw"));
assertEquals(201, req(port, "POST", "/workers?cwd=/tmp/proj").statusCode());
assertEquals("/tmp/proj", params(herdr, "agent.start").get("cwd"), "the worker starts in cwd");
assertEquals(201, req(port, "POST", "/members?cwd=/tmp/proj").statusCode());
assertEquals("/tmp/proj", params(herdr, "tab.create").get("cwd"), "the worker starts in cwd");
}
@Test
void spawnWithAnUnknownProfileIs400() throws Exception {
FakeHerdr herdr = new FakeHerdr();
int port = start(herdr, "http://gx00.gw:8000", Set.of("gx00.gw"));
HttpResponse<String> res = req(port, "POST", "/workers?profile=nope");
HttpResponse<String> res = req(port, "POST", "/members?profile=nope");
assertEquals(400, res.statusCode());
assertEquals("unknown_profile", mapper.readTree(res.body()).get("error").asText());
assertFalse(herdr.called("agent.start"), "an unknown profile must not spawn anything");
@@ -234,7 +246,7 @@ class BridgedAppTest {
FakeHerdr herdr = new FakeHerdr().withWorkspace("w9", "bridged-workers");
int port = start(herdr, "http://gx00.gw:8000", Set.of("gx00.gw"));
assertEquals(201, req(port, "POST", "/workers").statusCode());
assertEquals(201, req(port, "POST", "/members").statusCode());
assertFalse(herdr.called("workspace.create"), "existing worker space must be reused, not recreated");
assertTrue(herdr.called("tab.create"), "a fresh tab is still created for the worker");
}
@@ -246,7 +258,7 @@ class BridgedAppTest {
FakeHerdr herdr = new FakeHerdr().agentNameTakenTimes(2);
int port = start(herdr, "http://gx00.gw:8000", Set.of("gx00.gw"));
assertEquals(201, req(port, "POST", "/workers").statusCode());
assertEquals(201, req(port, "POST", "/members").statusCode());
List<String> names = herdr.calls.stream()
.filter(c -> c.method().equals("agent.start"))
.map(c -> ((Map<String, Object>) c.params()).get("name").toString())
@@ -261,7 +273,7 @@ class BridgedAppTest {
FakeHerdr herdr = new FakeHerdr().agentNameTakenTimes(99);
int port = start(herdr, "http://gx00.gw:8000", Set.of("gx00.gw"));
assertEquals(500, req(port, "POST", "/workers").statusCode());
assertEquals(500, req(port, "POST", "/members").statusCode());
assertTrue(herdr.called("tab.create"), "a tab was created before the failed start");
assertEquals("w9:t2", params(herdr, "tab.close").get("tab_id"), "orphaned tab must be closed");
}
@@ -272,7 +284,7 @@ class BridgedAppTest {
// base_url points at the subscription — guard must block before any herdr call.
int port = start(herdr, "https://api.anthropic.com", Set.of("gx00.gw"));
HttpResponse<String> res = req(port, "POST", "/workers");
HttpResponse<String> res = req(port, "POST", "/members");
assertEquals(403, res.statusCode());
assertEquals("subscription_boundary", mapper.readTree(res.body()).get("error").asText());
assertFalse(herdr.called("agent.start"), "guard must stop the spawn before herdr");
@@ -284,7 +296,7 @@ class BridgedAppTest {
FakeHerdr herdr = new FakeHerdr();
int port = start(herdr, "http://gx00.gw:8000", Set.of("gx00.gw"), "pane");
assertEquals(201, req(port, "POST", "/workers").statusCode());
assertEquals(201, req(port, "POST", "/members").statusCode());
Map<String, Object> start = params(herdr, "agent.start");
assertFalse(start.containsKey("tab_id"), "pane placement must not target a tab");
assertFalse(herdr.called("workspace.create"), "pane placement uses no dedicated space");
@@ -296,7 +308,7 @@ class BridgedAppTest {
FakeHerdr herdr = new FakeHerdr();
int port = start(herdr, "http://gx00.gw:8000", Set.of("gx00.gw"));
assertEquals(204, req(port, "DELETE", "/workers/w9:pW").statusCode());
assertEquals(204, req(port, "DELETE", "/members/w9:pW").statusCode());
assertTrue(herdr.called("pane.close"));
// Tab resolved from the pane (pane.get), then closed.
assertEquals("w9:t2", params(herdr, "tab.close").get("tab_id"));
@@ -322,7 +334,7 @@ class BridgedAppTest {
HttpResponse<String> res = send.get(6, java.util.concurrent.TimeUnit.SECONDS);
assertEquals(200, res.statusCode());
assertEquals("LGTM ship it", mapper.readTree(res.body()).get("reply").asText());
// (injection via agent.send is covered deterministically by the timeout-working test)
// (injection via agent.prompt is covered deterministically by the timeout-working test)
}
@Test
@@ -394,7 +406,7 @@ class BridgedAppTest {
HttpResponse<String> res = postMessage(port, "{\"content\":\"hi\",\"timeoutMs\":150}");
assertEquals(202, res.statusCode());
assertEquals("queued", mapper.readTree(res.body()).get("status").asText());
assertFalse(herdr.called("agent.send"), "no injection while the worker is mid-turn");
assertFalse(herdr.called("agent.prompt"), "no injection while the worker is mid-turn");
}
@Test
@@ -405,7 +417,7 @@ class BridgedAppTest {
HttpResponse<String> res = postMessage(port, "{\"content\":\"hi\",\"timeoutMs\":250}");
assertEquals(202, res.statusCode());
assertEquals("working", mapper.readTree(res.body()).get("status").asText());
assertTrue(herdr.called("agent.send"), "message was injected");
assertTrue(herdr.called("agent.prompt"), "message was injected");
}
@Test
@@ -438,7 +450,7 @@ class BridgedAppTest {
FakeHerdr herdr = new FakeHerdr();
int port = start(herdr, "http://gx00.gw:8000", Set.of("gx00.gw"), "pane");
assertEquals(204, req(port, "DELETE", "/workers/w9:pW").statusCode());
assertEquals(204, req(port, "DELETE", "/members/w9:pW").statusCode());
assertTrue(herdr.called("pane.close"));
assertFalse(herdr.called("tab.close"), "pane placement owns no tab to close");
assertFalse(herdr.called("pane.get"), "no tab resolution in pane placement");
@@ -450,7 +462,7 @@ class BridgedAppTest {
FakeHerdr herdr = new FakeHerdr().withWorkerTabPaneCount(2);
int port = start(herdr, "http://gx00.gw:8000", Set.of("gx00.gw"));
assertEquals(204, req(port, "DELETE", "/workers/w9:pW").statusCode());
assertEquals(204, req(port, "DELETE", "/members/w9:pW").statusCode());
assertTrue(herdr.called("pane.close"), "the worker's own pane is still closed");
assertFalse(herdr.called("tab.close"), "must not close a tab that holds the user's other panes");
}
@@ -461,7 +473,7 @@ class BridgedAppTest {
int port = start(herdr, "http://gx00.gw:8000", Set.of("gx00.gw"));
// A genuine teardown failure must surface, not be reported as a successful 204.
assertEquals(500, req(port, "DELETE", "/workers/w9:pW").statusCode());
assertEquals(500, req(port, "DELETE", "/members/w9:pW").statusCode());
assertFalse(herdr.called("tab.close"), "tab is not removed when the pane close failed");
}
@@ -471,7 +483,7 @@ class BridgedAppTest {
int port = start(herdr, "http://gx00.gw:8000", Set.of("gx00.gw"));
// Already-gone is success; the (now-empty) tab is still cleaned up.
assertEquals(204, req(port, "DELETE", "/workers/w9:pW").statusCode());
assertEquals(204, req(port, "DELETE", "/members/w9:pW").statusCode());
assertTrue(herdr.called("tab.close"));
}
}
@@ -0,0 +1,208 @@
package dev.ltms.bridged.session;
import org.junit.jupiter.api.Test;
import org.junit.jupiter.api.io.TempDir;
import java.nio.file.Files;
import java.nio.file.Path;
import java.util.List;
import java.util.concurrent.TimeUnit;
import static org.junit.jupiter.api.Assertions.*;
/**
* CB-525 acceptance test for tool-surface isolation. This is one of the few tests that drives real
* {@code git} — the behaviour under test is precisely what {@link GitWorktrees} does to a checkout,
* so a fake would assert nothing. Everything happens inside a {@link TempDir} throwaway repo.
*/
class GitWorktreesTest {
/** A project MCP config with servers in it — what this repo actually commits. */
private static final String WITH_SERVERS = """
{
"mcpServers": {
"jetbrains": { "type": "sse", "url": "http://localhost:64342/sse" }
}
}
""";
/** An opencode config carrying a {@code {file:.secrets/...}} reference — CB-543's crash repro. */
private static final String OPENCODE_WITH_FILE_REF = """
{
"env": {
"CONTEXT7_TOKEN": "{file:.secrets/context7-token}"
}
}
""";
/** A non-empty autoenv file — the form that would prompt for authorization in a worktree. */
private static final String AUTOENV_WITH_DIRECTIVE = "export HELLO=world\n";
private static Path initRepo(Path dir) throws Exception {
Files.createDirectories(dir);
git(dir, "init", "-q", "-b", "main");
git(dir, "config", "user.email", "test@example.invalid");
git(dir, "config", "user.name", "Test");
Files.writeString(dir.resolve(".mcp.json"), WITH_SERVERS);
Files.writeString(dir.resolve("README.md"), "seed\n");
git(dir, "add", ".mcp.json", "README.md");
git(dir, "commit", "-q", "-m", "seed");
return dir;
}
private static void git(Path cwd, String... args) throws Exception {
List<String> cmd = new java.util.ArrayList<>(List.of("git"));
cmd.addAll(List.of(args));
Process p = new ProcessBuilder(cmd).directory(cwd.toFile()).redirectErrorStream(true).start();
String out = new String(p.getInputStream().readAllBytes());
assertTrue(p.waitFor(30, TimeUnit.SECONDS), "git timed out: " + String.join(" ", cmd));
assertEquals(0, p.exitValue(), "git " + String.join(" ", args) + " failed:\n" + out);
}
/** Pending changes to {@code file} in {@code cwd}, empty when git considers it unmodified. */
private static String status(Path cwd, String file) throws Exception {
Process p = new ProcessBuilder("git", "status", "--porcelain", "--", file)
.directory(cwd.toFile()).redirectErrorStream(true).start();
String out = new String(p.getInputStream().readAllBytes());
assertTrue(p.waitFor(30, TimeUnit.SECONDS), "git status timed out");
return out;
}
/**
* The heart of CB-525: a provisioned worktree must not inherit the primary's MCP servers. Without
* the isolation step the checked-out {@code .mcp.json} carries them in, and a worker navigating
* through the primary's IDE servers edits the primary's tree while building its own.
*/
@Test
void aProvisionedWorktreeInheritsNoMcpServers(@TempDir Path tmp) throws Exception {
Path repo = initRepo(tmp.resolve("repo"));
String wt = new GitWorktrees(tmp.resolve("wts").toString())
.add(repo.toString(), "cb-525-a", "HEAD");
Path mcp = Path.of(wt).resolve(".mcp.json");
assertTrue(Files.exists(mcp), ".mcp.json must still exist — present and explicitly empty");
String body = Files.readString(mcp);
assertFalse(body.contains("jetbrains"), "worktree inherited the primary's MCP servers:\n" + body);
assertTrue(body.replaceAll("\\s+", "").contains("\"mcpServers\":{}"),
"expected an explicitly empty server map, got:\n" + body);
}
/** Neutralizing must not look like work in progress, or a worker would commit it into its PR. */
@Test
void theNeutralizedConfigIsNotAPendingLocalModification(@TempDir Path tmp) throws Exception {
Path repo = initRepo(tmp.resolve("repo"));
String wt = new GitWorktrees(tmp.resolve("wts").toString())
.add(repo.toString(), "cb-525-b", "HEAD");
assertEquals("", status(Path.of(wt), ".mcp.json"),
"the neutralized .mcp.json shows as modified — --skip-worktree did not take");
}
/** Isolation is the worktree's business only; the primary's own checkout must be untouched. */
@Test
void thePrimaryCheckoutIsLeftAlone(@TempDir Path tmp) throws Exception {
Path repo = initRepo(tmp.resolve("repo"));
new GitWorktrees(tmp.resolve("wts").toString()).add(repo.toString(), "cb-525-c", "HEAD");
assertEquals(WITH_SERVERS, Files.readString(repo.resolve(".mcp.json")),
"the primary's .mcp.json was rewritten — isolation reached out of the worktree");
}
/** A repo that commits no {@code .mcp.json} still gets one, so nothing can be inherited later. */
@Test
void aRepoWithoutAnMcpConfigStillGetsANeutralOne(@TempDir Path tmp) throws Exception {
Path repo = tmp.resolve("repo");
Files.createDirectories(repo);
git(repo, "init", "-q", "-b", "main");
git(repo, "config", "user.email", "test@example.invalid");
git(repo, "config", "user.name", "Test");
Files.writeString(repo.resolve("README.md"), "seed\n");
git(repo, "add", "README.md");
git(repo, "commit", "-q", "-m", "seed");
String wt = new GitWorktrees(tmp.resolve("wts").toString())
.add(repo.toString(), "cb-525-d", "HEAD");
// Untracked is the normal case here, so the --skip-worktree branch must be skipped rather
// than run and fail: `update-index --skip-worktree` on an unknown path exits non-zero.
String body = Files.readString(Path.of(wt).resolve(".mcp.json"));
assertTrue(body.replaceAll("\\s+", "").contains("\"mcpServers\":{}"), body);
}
/**
* CB-543's repro: a tracked {@code opencode.json} carries a {@code {file:.secrets/...}} reference
* to a gitignored secret that never reaches a worktree, and opencode refuses to start on it. The
* worktree's copy must be neutralized and hidden like {@code .mcp.json}.
*/
@Test
void aTrackedOpencodeConfigIsNeutralizedAndHidden(@TempDir Path tmp) throws Exception {
Path repo = tmp.resolve("repo");
Files.createDirectories(repo);
git(repo, "init", "-q", "-b", "main");
git(repo, "config", "user.email", "test@example.invalid");
git(repo, "config", "user.name", "Test");
Files.writeString(repo.resolve(".mcp.json"), WITH_SERVERS);
Files.writeString(repo.resolve("opencode.json"), OPENCODE_WITH_FILE_REF);
Files.writeString(repo.resolve("README.md"), "seed\n");
git(repo, "add", ".mcp.json", "opencode.json", "README.md");
git(repo, "commit", "-q", "-m", "seed");
String wt = new GitWorktrees(tmp.resolve("wts").toString())
.add(repo.toString(), "cb-543-a", "HEAD");
String body = Files.readString(Path.of(wt).resolve("opencode.json"));
assertFalse(body.contains(".secrets"),
"worktree kept a dangling {file:...} secret reference:\n" + body);
assertEquals("{}", body.replaceAll("\\s+", ""),
"expected an empty JSON object stub, got:\n" + body);
assertEquals("", status(Path.of(wt), "opencode.json"),
"the neutralized opencode.json shows as modified — --skip-worktree did not take");
}
/** A config the repo does not carry must be skipped — no stub invented, provisioning still succeeds. */
@Test
void anAbsentConfigIsSkippedWithoutError(@TempDir Path tmp) throws Exception {
Path repo = initRepo(tmp.resolve("repo")); // only .mcp.json + README are committed
String wt = new GitWorktrees(tmp.resolve("wts").toString())
.add(repo.toString(), "cb-543-b", "HEAD");
assertFalse(Files.exists(Path.of(wt).resolve("opencode.json")),
"a stub was invented for a config the repo does not carry");
assertFalse(Files.exists(Path.of(wt).resolve(".autoenv")),
"a stub was invented for a config the repo does not carry");
// .mcp.json's long-standing create-always behaviour must be unchanged.
assertTrue(Files.exists(Path.of(wt).resolve(".mcp.json")), ".mcp.json stub was dropped");
}
/** All three protected configs are covered: each one present in a worktree is neutralized and hidden. */
@Test
void allThreeConfigsAreNeutralizedWhenPresent(@TempDir Path tmp) throws Exception {
Path repo = tmp.resolve("repo");
Files.createDirectories(repo);
git(repo, "init", "-q", "-b", "main");
git(repo, "config", "user.email", "test@example.invalid");
git(repo, "config", "user.name", "Test");
Files.writeString(repo.resolve(".mcp.json"), WITH_SERVERS);
Files.writeString(repo.resolve("opencode.json"), OPENCODE_WITH_FILE_REF);
Files.writeString(repo.resolve(".autoenv"), AUTOENV_WITH_DIRECTIVE);
Files.writeString(repo.resolve("README.md"), "seed\n");
git(repo, "add", ".mcp.json", "opencode.json", ".autoenv", "README.md");
git(repo, "commit", "-q", "-m", "seed");
String wt = new GitWorktrees(tmp.resolve("wts").toString())
.add(repo.toString(), "cb-543-c", "HEAD");
assertTrue(Files.readString(Path.of(wt).resolve(".mcp.json"))
.replaceAll("\\s+", "").contains("\"mcpServers\":{}"),
".mcp.json was not neutralized");
assertEquals("{}", Files.readString(Path.of(wt).resolve("opencode.json")).replaceAll("\\s+", ""),
"opencode.json was not neutralized");
assertEquals("", Files.readString(Path.of(wt).resolve(".autoenv")),
".autoenv was not neutralized");
assertEquals("", status(Path.of(wt), ".mcp.json"), ".mcp.json still shows as modified");
assertEquals("", status(Path.of(wt), "opencode.json"), "opencode.json still shows as modified");
assertEquals("", status(Path.of(wt), ".autoenv"), ".autoenv still shows as modified");
}
}
@@ -5,7 +5,7 @@ import dev.ltms.bridged.guard.SubscriptionGuard;
import dev.ltms.bridged.herdr.AgentControl;
import dev.ltms.bridged.herdr.FakeHerdr;
import dev.ltms.bridged.herdr.WorkspaceControl;
import dev.ltms.bridged.worker.ClaudeCodeLauncher;
import dev.ltms.bridged.member.ClaudeCodeLauncher;
import dev.ltms.bridged.peer.PeerUnreachableException;
import org.junit.jupiter.api.Test;
@@ -25,7 +25,7 @@ import static org.junit.jupiter.api.Assertions.*;
class SessionManagerTest {
private SessionManager sessionManager(FakeHerdr herdr) {
BridgedConfig.Worker cfg = new BridgedConfig.Worker(
BridgedConfig.Profile cfg = new BridgedConfig.Profile(
"ltms-local", "http://gx00.gw:8000", "coder", null, "BRIDGED_WORKER_TOKEN",
List.of("ccs", "ltms-local"), "tab", "bridged-workers",
"worker: {profile} #{n}", null, null, null);
@@ -39,13 +39,29 @@ class SessionManagerTest {
}
private SessionManager sessionManager(FakeHerdr herdr, LongSupplier clock, int contextCap) {
BridgedConfig.Worker cfg = new BridgedConfig.Worker(
return sessionManager(herdr, clock, contextCap, false);
}
private SessionManager sessionManager(FakeHerdr herdr, LongSupplier clock, int contextCap,
boolean clearAfterTurn) {
BridgedConfig.Profile cfg = new BridgedConfig.Profile(
"ltms-local", "http://gx00.gw:8000", "coder", null, "BRIDGED_WORKER_TOKEN",
List.of("ccs", "ltms-local"), "tab", "bridged-workers",
"worker: {profile} #{n}", null, null, null);
ClaudeCodeLauncher workers = new ClaudeCodeLauncher(new AgentControl(herdr), new WorkspaceControl(herdr),
new SubscriptionGuard(Set.of("gx00.gw")), Map.of(cfg.profile(), cfg), cfg.profile(), _ -> null);
return new SessionManager(workers, new GitWorktrees(), clock, contextCap);
return new SessionManager(workers, new GitWorktrees(), clock, contextCap, clearAfterTurn);
}
@Test
void primaryContactWithNoTerminalIsNotAReadinessSignal() {
// The MCP context extractor calls presence.markPresent(p.terminal()) on EVERY request,
// and the primary's terminal is null — the presence bridge must treat that as a no-op,
// not feed it into the READY transition (which NPEd on the first real primary contact).
SessionManager sessions = sessionManager(new FakeHerdr());
assertDoesNotThrow(() -> sessions.asPresence().markPresent(null));
assertDoesNotThrow(() -> sessions.asPresence().markPresent(" "));
}
@Test
@@ -53,10 +69,10 @@ class SessionManagerTest {
FakeHerdr herdr = new FakeHerdr();
SessionManager sessions = sessionManager(herdr);
WorkerSession a = sessions.acquire("ltms-local", "/work/a", "/caller/a", "term_primary");
WorkerSession b = sessions.acquire("ltms-local", "/work/b", "/caller/b", "term_primary");
MemberSession a = sessions.acquire("ltms-local", "/work/a", "/caller/a", "term_primary");
MemberSession b = sessions.acquire("ltms-local", "/work/b", "/caller/b", "term_primary");
assertEquals(WorkerSession.State.SPAWNING, a.state(), "fresh session starts spawning");
assertEquals(MemberSession.State.SPAWNING, a.state(), "fresh session starts spawning");
assertEquals("ltms-local", a.profile());
assertEquals("/work/a", a.cwd(), "explicit requested cwd is recorded");
assertEquals("term_primary", a.ownerTerminal());
@@ -69,24 +85,43 @@ class SessionManagerTest {
assertEquals(2, sessions.roster().size(), "both sessions are registered");
}
@Test
void aNullTerminalFromThePrimaryIsANoOpEvenWithSessionsRegistered() {
// The primary resolves to a Principal with no terminal, and BridgeMcp's context extractor
// forwards that null into markPresent on EVERY MCP call. It only reached the registry scan
// once a session existed, so this NPE'd the primary's second spawn while the first passed.
FakeHerdr herdr = new FakeHerdr();
SessionManager sessions = sessionManager(herdr);
MemberSession session = sessions.acquire("ltms-local", null, "/caller", "term_primary");
assertDoesNotThrow(() -> sessions.asPresence().markPresent(null),
"the primary's null terminal must not blow up an unrelated tool call");
assertDoesNotThrow(() -> sessions.onDelivered(null));
assertDoesNotThrow(() -> sessions.onTurnComplete(null));
assertDoesNotThrow(() -> sessions.onTurnFailed(null));
assertEquals(MemberSession.State.SPAWNING, sessions.get(session.paneId()).orElseThrow().state(),
"and must not transition any registered session");
}
@Test
void presenceMovesSpawningToReadyAndDeliveredTurnMovesToDone() {
FakeHerdr herdr = new FakeHerdr();
SessionManager sessions = sessionManager(herdr);
WorkerSession session = sessions.acquire("ltms-local", null, "/caller", "term_primary");
MemberSession session = sessions.acquire("ltms-local", null, "/caller", "term_primary");
String terminal = session.terminalId();
sessions.asPresence().markPresent(terminal);
assertEquals(WorkerSession.State.READY, sessions.get(session.paneId()).orElseThrow().state(),
assertEquals(MemberSession.State.READY, sessions.get(session.paneId()).orElseThrow().state(),
"MCP presence moves SPAWNING → READY");
assertTrue(sessions.asPresence().isPresent(terminal), "presence is also recorded");
sessions.onDelivered(terminal);
assertEquals(WorkerSession.State.BUSY, sessions.get(session.paneId()).orElseThrow().state(),
assertEquals(MemberSession.State.BUSY, sessions.get(session.paneId()).orElseThrow().state(),
"delivery moves READY → BUSY");
sessions.onTurnComplete(terminal);
assertEquals(WorkerSession.State.DONE, sessions.get(session.paneId()).orElseThrow().state(),
assertEquals(MemberSession.State.DONE, sessions.get(session.paneId()).orElseThrow().state(),
"turn completion moves BUSY → DONE");
}
@@ -94,7 +129,7 @@ class SessionManagerTest {
void releaseTearsDownWorkerAndRemovesFromRosterAndIsIdempotent() {
FakeHerdr herdr = new FakeHerdr();
SessionManager sessions = sessionManager(herdr);
WorkerSession session = sessions.acquire("ltms-local", null, "/caller", null);
MemberSession session = sessions.acquire("ltms-local", null, "/caller", null);
String paneId = session.paneId();
sessions.release(paneId);
@@ -110,15 +145,15 @@ class SessionManagerTest {
void onTurnFailedMovesSessionToFailed() {
FakeHerdr herdr = new FakeHerdr();
SessionManager sessions = sessionManager(herdr);
WorkerSession session = sessions.acquire("ltms-local", null, "/caller", "term_primary");
MemberSession session = sessions.acquire("ltms-local", null, "/caller", "term_primary");
String terminal = session.terminalId();
sessions.asPresence().markPresent(terminal);
sessions.onDelivered(terminal);
sessions.onTurnFailed(terminal);
WorkerSession updated = sessions.get(session.paneId()).orElseThrow();
assertEquals(WorkerSession.State.FAILED, updated.state(), "turn failure moves to FAILED");
MemberSession updated = sessions.get(session.paneId()).orElseThrow();
assertEquals(MemberSession.State.FAILED, updated.state(), "turn failure moves to FAILED");
assertTrue(sessions.roster().contains(updated), "FAILED is still in acquired-minus-released roster");
}
@@ -126,11 +161,11 @@ class SessionManagerTest {
void recycleProducesNewPaneIdAndOldOneIsGone() {
FakeHerdr herdr = new FakeHerdr();
SessionManager sessions = sessionManager(herdr);
WorkerSession oldSession = sessions.acquire("ltms-local", null, "/caller", "term_primary");
MemberSession oldSession = sessions.acquire("ltms-local", null, "/caller", "term_primary");
String oldPane = oldSession.paneId();
String oldTerminal = oldSession.terminalId();
WorkerSession fresh = sessions.recycle(oldPane);
MemberSession fresh = sessions.recycle(oldPane);
assertNotEquals(oldPane, fresh.paneId(), "recycle yields a new pane id");
assertNotEquals(oldTerminal, fresh.terminalId(), "recycle yields a new terminal id");
@@ -142,9 +177,11 @@ class SessionManagerTest {
assertEquals(1, sessions.roster().size(), "only the fresh session remains");
assertEquals(fresh.paneId(), sessions.roster().getFirst().paneId());
// The old session was the first spawn → pane w9:pRoot_1 (CB-519: the registry key is the
// uuid id, so teardown is asserted on the real pane coordinate).
long paneCloseCount = herdr.calls.stream()
.filter(c -> "pane.close".equals(c.method()))
.filter(c -> oldPane.equals(((Map<?, ?>) c.params()).get("pane_id")))
.filter(c -> "w9:pRoot_1".equals(((Map<?, ?>) c.params()).get("pane_id")))
.count();
assertEquals(1, paneCloseCount, "the old worker was torn down");
}
@@ -153,8 +190,8 @@ class SessionManagerTest {
void rosterReflectsAcquiredMinusReleased() {
FakeHerdr herdr = new FakeHerdr();
SessionManager sessions = sessionManager(herdr);
WorkerSession a = sessions.acquire("ltms-local", "/a", "/caller", "ownerA");
WorkerSession b = sessions.acquire("ltms-local", "/b", "/caller", "ownerB");
MemberSession a = sessions.acquire("ltms-local", "/a", "/caller", "ownerA");
MemberSession b = sessions.acquire("ltms-local", "/b", "/caller", "ownerB");
assertEquals(2, sessions.roster().size());
assertTrue(sessions.roster().stream().anyMatch(s -> s.paneId().equals(a.paneId())));
@@ -182,7 +219,7 @@ class SessionManagerTest {
long[] clock = {0};
FakeHerdr herdr = new FakeHerdr();
SessionManager sessions = sessionManager(herdr, () -> clock[0]);
WorkerSession session = sessions.acquire("ltms-local", null, "/caller", "term_primary");
MemberSession session = sessions.acquire("ltms-local", null, "/caller", "term_primary");
String terminal = session.terminalId();
sessions.asPresence().markPresent(terminal);
@@ -197,13 +234,13 @@ class SessionManagerTest {
long[] clock = {0};
FakeHerdr herdr = new FakeHerdr();
SessionManager sessions = sessionManager(herdr, () -> clock[0]);
WorkerSession session = sessions.acquire("ltms-local", null, "/caller", "term_primary");
MemberSession session = sessions.acquire("ltms-local", null, "/caller", "term_primary");
String terminal = session.terminalId();
sessions.asPresence().markPresent(terminal);
clock[0] = 5;
assertEquals(0, sessions.reapIdle(10), "READY session within TTL is not reaped");
assertEquals(WorkerSession.State.READY,
assertEquals(MemberSession.State.READY,
sessions.get(session.paneId()).orElseThrow().state(),
"READY session survives");
}
@@ -213,14 +250,14 @@ class SessionManagerTest {
long[] clock = {0};
FakeHerdr herdr = new FakeHerdr();
SessionManager sessions = sessionManager(herdr, () -> clock[0]);
WorkerSession session = sessions.acquire("ltms-local", null, "/caller", "term_primary");
MemberSession session = sessions.acquire("ltms-local", null, "/caller", "term_primary");
String terminal = session.terminalId();
sessions.asPresence().markPresent(terminal);
sessions.onDelivered(terminal);
clock[0] = 100;
assertEquals(0, sessions.reapIdle(10), "BUSY session past TTL is never reaped");
assertEquals(WorkerSession.State.BUSY,
assertEquals(MemberSession.State.BUSY,
sessions.get(session.paneId()).orElseThrow().state(),
"BUSY session remains");
}
@@ -230,7 +267,7 @@ class SessionManagerTest {
long[] clock = {0};
FakeHerdr herdr = new FakeHerdr();
SessionManager sessions = sessionManager(herdr, () -> clock[0]);
WorkerSession session = sessions.acquire("ltms-local", null, "/caller", "term_primary");
MemberSession session = sessions.acquire("ltms-local", null, "/caller", "term_primary");
String terminal = session.terminalId();
sessions.asPresence().markPresent(terminal);
sessions.onDelivered(terminal);
@@ -247,8 +284,8 @@ class SessionManagerTest {
FakeHerdr herdr = new FakeHerdr();
SessionManager sessions = sessionManager(herdr, () -> clock[0]);
WorkerSession ready = sessions.acquire("ltms-local", "/ready", "/caller", "owner1");
WorkerSession busy = sessions.acquire("ltms-local", "/busy", "/caller", "owner2");
MemberSession ready = sessions.acquire("ltms-local", "/ready", "/caller", "owner1");
MemberSession busy = sessions.acquire("ltms-local", "/busy", "/caller", "owner2");
sessions.asPresence().markPresent(ready.terminalId());
sessions.asPresence().markPresent(busy.terminalId());
sessions.onDelivered(busy.terminalId());
@@ -256,7 +293,7 @@ class SessionManagerTest {
clock[0] = 50;
assertEquals(1, sessions.reapIdle(30), "only READY past TTL is reaped");
assertTrue(sessions.get(ready.paneId()).isEmpty(), "READY session is gone");
assertEquals(WorkerSession.State.BUSY,
assertEquals(MemberSession.State.BUSY,
sessions.get(busy.paneId()).orElseThrow().state(),
"BUSY session is still registered");
}
@@ -265,7 +302,7 @@ class SessionManagerTest {
void contextCapDisabledSessionSurvivesMultipleTurns() {
FakeHerdr herdr = new FakeHerdr();
SessionManager sessions = sessionManager(herdr, () -> 0L, 0);
WorkerSession session = sessions.acquire("ltms-local", null, "/caller", "term_primary");
MemberSession session = sessions.acquire("ltms-local", null, "/caller", "term_primary");
String terminal = session.terminalId();
sessions.asPresence().markPresent(terminal);
@@ -274,10 +311,10 @@ class SessionManagerTest {
sessions.onDelivered(terminal);
sessions.onTurnComplete(terminal);
WorkerSession updated = sessions.get(session.paneId()).orElseThrow();
assertEquals(WorkerSession.State.DONE, updated.state(), "session finishes second turn");
MemberSession updated = sessions.get(session.paneId()).orElseThrow();
assertEquals(MemberSession.State.DONE, updated.state(), "session finishes second turn");
assertEquals(2, updated.turnCount(), "turn count tracks both deliveries");
long releaseCloseCount = paneCloseCallsFor(herdr, session.paneId());
long releaseCloseCount = paneCloseCallsFor(herdr, "w9:pRoot_1"); // the real pane coordinate
assertEquals(0, releaseCloseCount, "cap disabled — no forced release of the worker pane");
}
@@ -285,13 +322,13 @@ class SessionManagerTest {
void contextCapTwoReleasesAfterSecondComplete() {
FakeHerdr herdr = new FakeHerdr();
SessionManager sessions = sessionManager(herdr, () -> 0L, 2);
WorkerSession session = sessions.acquire("ltms-local", null, "/caller", "term_primary");
MemberSession session = sessions.acquire("ltms-local", null, "/caller", "term_primary");
String terminal = session.terminalId();
sessions.asPresence().markPresent(terminal);
sessions.onDelivered(terminal);
sessions.onTurnComplete(terminal);
assertEquals(WorkerSession.State.DONE,
assertEquals(MemberSession.State.DONE,
sessions.get(session.paneId()).orElseThrow().state(),
"first turn completes without release");
@@ -300,18 +337,62 @@ class SessionManagerTest {
assertTrue(sessions.get(session.paneId()).isEmpty(), "session released after cap reached");
assertTrue(sessions.roster().isEmpty(), "released session leaves roster");
assertEquals(1, paneCloseCallsFor(herdr, session.paneId()),
assertEquals(1, paneCloseCallsFor(herdr, "w9:pRoot_1"),
"forced release tears the worker pane down exactly once");
}
@Test
void clearAfterTurnResetsContextWithoutDoubleCountingTheTurn() {
FakeHerdr herdr = new FakeHerdr();
SessionManager sessions = sessionManager(herdr, () -> 0L, 0, true);
MemberSession session = sessions.acquire("ltms-local", null, "/caller", "term_primary");
sessions.asPresence().markPresent(session.terminalId());
sessions.onDelivered(session.terminalId());
assertTrue(sessions.onTurnCompleteWithPostAction(session.terminalId()));
MemberSession updated = sessions.get(session.paneId()).orElseThrow();
assertEquals(1, updated.turnCount(), "the reset is housekeeping, not a second delegation");
assertEquals(List.of("/clear"), promptTexts(herdr));
}
@Test
void contextCapReleaseWinsOverClearAfterTurn() {
FakeHerdr herdr = new FakeHerdr();
SessionManager sessions = sessionManager(herdr, () -> 0L, 1, true);
MemberSession session = sessions.acquire("ltms-local", null, "/caller", "term_primary");
sessions.asPresence().markPresent(session.terminalId());
sessions.onDelivered(session.terminalId());
assertFalse(sessions.hasPostTurnAction(session.terminalId()),
"a session at its cap will be released, not reset for reuse");
assertFalse(sessions.onTurnCompleteWithPostAction(session.terminalId()));
assertTrue(sessions.get(session.paneId()).isEmpty());
assertTrue(promptTexts(herdr).isEmpty(), "never send /clear into a worker being torn down");
}
@Test
void clearAfterTurnFalsePreservesCompletionWithoutAControlPrompt() {
FakeHerdr herdr = new FakeHerdr();
SessionManager sessions = sessionManager(herdr, () -> 0L, 0, false);
MemberSession session = sessions.acquire("ltms-local", null, "/caller", "term_primary");
sessions.asPresence().markPresent(session.terminalId());
sessions.onDelivered(session.terminalId());
sessions.onTurnComplete(session.terminalId());
assertEquals(MemberSession.State.DONE, sessions.get(session.paneId()).orElseThrow().state());
assertTrue(promptTexts(herdr).isEmpty());
}
@Test
void drainAllReleasesBusyAndReadySessionsAndWaitsForBusy() {
long[] clock = {0};
FakeHerdr herdr = new FakeHerdr();
SessionManager sessions = sessionManager(herdr, () -> clock[0]);
WorkerSession ready = sessions.acquire("ltms-local", "/ready", "/caller", "ownerR");
WorkerSession busy = sessions.acquire("ltms-local", "/busy", "/caller", "ownerB");
MemberSession ready = sessions.acquire("ltms-local", "/ready", "/caller", "ownerR");
MemberSession busy = sessions.acquire("ltms-local", "/busy", "/caller", "ownerB");
sessions.asPresence().markPresent(ready.terminalId());
sessions.asPresence().markPresent(busy.terminalId());
sessions.onDelivered(busy.terminalId());
@@ -321,9 +402,10 @@ class SessionManagerTest {
assertTrue(sessions.roster().isEmpty(), "drain clears the roster");
assertTrue(sessions.get(ready.paneId()).isEmpty(), "ready session is released");
assertTrue(sessions.get(busy.paneId()).isEmpty(), "busy session is released after timeout");
assertEquals(1, paneCloseCallsFor(herdr, ready.paneId()),
// ready is the first spawn → pane w9:pRoot_1, busy the second → w9:pRoot_2 (FakeHerdr order).
assertEquals(1, paneCloseCallsFor(herdr, "w9:pRoot_1"),
"ready worker pane is torn down");
assertEquals(1, paneCloseCallsFor(herdr, busy.paneId()),
assertEquals(1, paneCloseCallsFor(herdr, "w9:pRoot_2"),
"busy worker pane is torn down");
}
@@ -334,6 +416,13 @@ class SessionManagerTest {
.count();
}
private static List<String> promptTexts(FakeHerdr herdr) {
return herdr.calls.stream()
.filter(c -> "agent.prompt".equals(c.method()))
.map(c -> String.valueOf(((Map<?, ?>) c.params()).get("text")))
.toList();
}
// --- CB-306 spawn-readiness gate: no half-registered session on timeout ----------------
@Test
@@ -343,7 +432,7 @@ class SessionManagerTest {
long[] clock = {0};
// Gate-enabled launcher (1 ms timeout + no-op sleeper that advances clock past deadline)
BridgedConfig.Worker cfg = new BridgedConfig.Worker(
BridgedConfig.Profile cfg = new BridgedConfig.Profile(
"ltms-local", "http://gx00.gw:8000", "coder", null, "BRIDGED_WORKER_TOKEN",
List.of("ccs", "ltms-local"), "tab", "bridged-workers",
"worker: {profile} #{n}", null, null, null);
@@ -372,7 +461,7 @@ class SessionManagerTest {
java.util.List<String> released = new java.util.concurrent.CopyOnWriteArrayList<>();
sessions.onRelease(released::add);
WorkerSession s = sessions.acquire("ltms-local", null, "/caller", null);
MemberSession s = sessions.acquire("ltms-local", null, "/caller", null);
sessions.release(s.paneId());
assertEquals(java.util.List.of(s.terminalId()), released,
@@ -399,7 +488,7 @@ class SessionManagerTest {
throw new IllegalStateException("listener blew up");
});
WorkerSession s = sessions.acquire("ltms-local", null, "/caller", null);
MemberSession s = sessions.acquire("ltms-local", null, "/caller", null);
assertDoesNotThrow(() -> sessions.release(s.paneId()),
"a listener failure must never prevent the teardown it is reacting to");
assertTrue(sessions.get(s.paneId()).isEmpty(), "and the session is still deregistered");
@@ -5,7 +5,7 @@ import dev.ltms.bridged.guard.SubscriptionGuard;
import dev.ltms.bridged.herdr.AgentControl;
import dev.ltms.bridged.herdr.FakeHerdr;
import dev.ltms.bridged.herdr.WorkspaceControl;
import dev.ltms.bridged.worker.ClaudeCodeLauncher;
import dev.ltms.bridged.member.ClaudeCodeLauncher;
import org.junit.jupiter.api.Test;
import java.util.List;
@@ -29,7 +29,7 @@ class SessionReaperTest {
private static ClaudeCodeLauncher launcher() {
FakeHerdr herdr = new FakeHerdr();
BridgedConfig.Worker cfg = new BridgedConfig.Worker(
BridgedConfig.Profile cfg = new BridgedConfig.Profile(
"ltms-local", "http://gx00.gw:8000", "coder", null, "BRIDGED_WORKER_TOKEN",
List.of("ccs", "ltms-local"), "tab", "bridged-workers",
"worker: {profile} #{n}", null, null, null);
@@ -1,16 +1,19 @@
package dev.ltms.bridged.session;
import dev.ltms.bridged.auth.MemberRegistry;
import dev.ltms.bridged.config.BridgedConfig;
import dev.ltms.bridged.guard.SubscriptionGuard;
import dev.ltms.bridged.herdr.AgentControl;
import dev.ltms.bridged.herdr.FakeHerdr;
import dev.ltms.bridged.herdr.WorkspaceControl;
import dev.ltms.bridged.worker.ClaudeCodeLauncher;
import dev.ltms.bridged.member.ClaudeCodeLauncher;
import dev.ltms.bridged.peer.MemberRole;
import org.junit.jupiter.api.Test;
import java.util.List;
import java.util.Map;
import java.util.Set;
import java.util.concurrent.TimeUnit;
import static org.junit.jupiter.api.Assertions.*;
@@ -21,8 +24,15 @@ import static org.junit.jupiter.api.Assertions.*;
*/
class WorktreeSessionManagerTest {
private static MemberRegistry members() {
return new MemberRegistry(new BridgedConfig.Fleet(Map.of(),
Map.of("architect", new BridgedConfig.Slot("ltms-local")),
Map.of("dev", new BridgedConfig.Slot("ltms-local")),
Map.of("reviewer", new BridgedConfig.Slot("ltms-local")), null));
}
private static ClaudeCodeLauncher workerService(FakeHerdr herdr) {
BridgedConfig.Worker cfg = new BridgedConfig.Worker(
BridgedConfig.Profile cfg = new BridgedConfig.Profile(
"ltms-local", "http://gx00.gw:8000", "coder", null, "BRIDGED_WORKER_TOKEN",
List.of("ccs", "ltms-local"), "tab", "bridged-workers",
"worker: {profile} #{n}", null, null, null);
@@ -32,7 +42,8 @@ class WorktreeSessionManagerTest {
private static String startCwd(FakeHerdr herdr) {
@SuppressWarnings("unchecked")
Map<String, Object> start = (Map<String, Object>) herdr.lastCall("agent.start").params();
// Protocol 19: the worker's cwd rides on pane creation (tab.create), not agent.start.
Map<String, Object> start = (Map<String, Object>) herdr.lastCall("tab.create").params();
Object cwd = start.get("cwd");
return cwd == null ? null : cwd.toString();
}
@@ -43,7 +54,7 @@ class WorktreeSessionManagerTest {
FakeWorktrees worktrees = new FakeWorktrees();
SessionManager sessions = new SessionManager(workerService(herdr), worktrees);
WorkerSession s = sessions.acquire("ltms-local", null, "/caller/proj", "term_primary");
MemberSession s = sessions.acquire("ltms-local", null, "/caller/proj", "term_primary");
assertTrue(worktrees.addCalls().isEmpty(), "shared-tree acquire never adds a worktree");
assertTrue(worktrees.repoRootCalls().isEmpty(), "shared-tree acquire never resolves a repo root");
@@ -54,13 +65,43 @@ class WorktreeSessionManagerTest {
assertEquals("/caller/proj", startCwd(herdr), "spawn receives the caller's cwd");
}
@Test
void onlyArchitectsBindAndReleaseMakesTheirSlotReusable() {
FakeHerdr herdr = new FakeHerdr();
MemberRegistry members = members();
SessionManager sessions = new SessionManager(workerService(herdr), new FakeWorktrees());
sessions.setMemberLifecycle(members);
MemberSession architect = sessions.acquire("ltms-local", MemberRole.ARCHITECT,
null, "/caller/proj", null, null);
MemberSession dev = sessions.acquire("ltms-local", MemberRole.DEV,
null, "/caller/proj", null, null);
MemberSession reviewer = sessions.acquire("ltms-local", MemberRole.REVIEWER,
null, "/caller/proj", null, null);
assertEquals("architect:architect", members.slotForTerminal(architect.terminalId()));
assertNull(members.slotForTerminal(dev.terminalId()), "a dev must never receive architect rights");
assertNull(members.slotForTerminal(reviewer.terminalId()),
"a reviewer must never receive architect rights");
sessions.release(architect.paneId());
MemberSession replacement = sessions.acquire("ltms-local", MemberRole.ARCHITECT,
null, "/caller/proj", null, null);
assertEquals("architect:architect", members.slotForTerminal(replacement.terminalId()));
MemberSession overflow = sessions.acquire("ltms-local", MemberRole.ARCHITECT,
null, "/caller/proj", null, null);
assertNull(members.slotForTerminal(overflow.terminalId()),
"a full slot pool must not stop the architect spawn");
}
@Test
void worktreeAcquireProvisionsAndRecordsPathAndBranch() {
FakeHerdr herdr = new FakeHerdr();
FakeWorktrees worktrees = new FakeWorktrees().withRepoRoot("/repo").withPrefix("/wt");
SessionManager sessions = new SessionManager(workerService(herdr), worktrees);
WorkerSession s = sessions.acquire("ltms-local", null, "/caller/proj", "term_primary",
MemberSession s = sessions.acquire("ltms-local", null, "/caller/proj", "term_primary",
new WorktreeRequest("cb-999", null));
assertEquals(1, worktrees.addCalls().size(), "one worktree was added");
@@ -77,11 +118,24 @@ class WorktreeSessionManagerTest {
assertEquals(expectedPath, s.cwd(), "session cwd is the worktree path");
}
@Test
void worktreeArchitectAcquireAlsoBindsItsSlot() {
FakeHerdr herdr = new FakeHerdr();
MemberRegistry members = members();
SessionManager sessions = new SessionManager(workerService(herdr), new FakeWorktrees());
sessions.setMemberLifecycle(members);
MemberSession architect = sessions.acquire("ltms-local", MemberRole.ARCHITECT,
null, "/caller/proj", null, new WorktreeRequest("cb-548", null));
assertEquals("architect:architect", members.slotForTerminal(architect.terminalId()));
}
@Test
void worktreeAcquireRunsParityOverlayWithProfileDefaults() {
FakeHerdr herdr = new FakeHerdr();
FakeWorktrees worktrees = new FakeWorktrees().withRepoRoot("/repo").withPrefix("/wt")
.track(".mcp.json")
.track(".envrc")
.exists(".claude/settings.local.json");
SessionManager sessions = new SessionManager(workerService(herdr), worktrees);
@@ -92,11 +146,14 @@ class WorktreeSessionManagerTest {
FakeWorktrees.OverlayCall overlay = worktrees.lastOverlay();
assertNotNull(overlay);
assertEquals("/repo", overlay.repoRoot());
assertEquals(List.of(".mcp.json", ".claude/settings.local.json", ".env", ".envrc"),
assertEquals(List.of(".env", ".envrc"),
overlay.requested(), "default parity overlay is used when unset");
assertEquals(List.of(".mcp.json", ".claude/settings.local.json"), overlay.copied(),
assertFalse(overlay.requested().contains(".mcp.json"),
"CB-525: replicating the primary's MCP config gives a worker the primary's IDE "
+ "servers, which navigate its edits out of its own worktree");
assertEquals(List.of(".envrc"), overlay.copied(),
"existing paths are copied; missing paths are skipped");
assertEquals(List.of(".mcp.json"), overlay.skipWorktree(),
assertEquals(List.of(".envrc"), overlay.skipWorktree(),
"tracked copied paths are --skip-worktree'd");
}
@@ -105,7 +162,7 @@ class WorktreeSessionManagerTest {
FakeHerdr herdr = new FakeHerdr();
FakeWorktrees worktrees = new FakeWorktrees().withRepoRoot("/repo").withPrefix("/wt");
SessionManager sessions = new SessionManager(workerService(herdr), worktrees);
WorkerSession s = sessions.acquire("ltms-local", null, "/caller/proj", null,
MemberSession s = sessions.acquire("ltms-local", null, "/caller/proj", null,
new WorktreeRequest("cb-666", null));
String paneId = s.paneId();
@@ -121,12 +178,48 @@ class WorktreeSessionManagerTest {
assertTrue(sessions.get(paneId).isEmpty(), "released session is no longer retrievable");
}
@Test
void drainAllPreservesWorktreeOfIdleSession() {
FakeHerdr herdr = new FakeHerdr();
FakeWorktrees worktrees = new FakeWorktrees().withRepoRoot("/repo").withPrefix("/wt");
SessionManager sessions = new SessionManager(workerService(herdr), worktrees);
MemberSession s = sessions.acquire("ltms-local", null, "/caller/proj", null,
new WorktreeRequest("cb-544", null));
sessions.asPresence().markPresent(s.terminalId()); // READY (idle)
sessions.drainAll(TimeUnit.MILLISECONDS.toNanos(100));
assertTrue(herdr.called("pane.close"), "shutdown drain still stops the worker pane");
assertTrue(worktrees.removeCalls().isEmpty(),
"shutdown drain must NOT remove the worktree — it is the only copy of the work");
assertTrue(sessions.roster().isEmpty(), "shutdown drain still deregisters the session");
}
@Test
void drainAllPreservesWorktreeOfSessionStillBusyAtTimeout() {
FakeHerdr herdr = new FakeHerdr();
FakeWorktrees worktrees = new FakeWorktrees().withRepoRoot("/repo").withPrefix("/wt");
SessionManager sessions = new SessionManager(workerService(herdr), worktrees);
MemberSession s = sessions.acquire("ltms-local", null, "/caller/proj", null,
new WorktreeRequest("cb-544", null));
String terminal = s.terminalId();
sessions.asPresence().markPresent(terminal);
sessions.onDelivered(terminal); // BUSY, never completes → still BUSY when the timeout hits
sessions.drainAll(TimeUnit.MILLISECONDS.toNanos(100));
assertTrue(herdr.called("pane.close"), "shutdown drain stops the worker pane");
assertTrue(worktrees.removeCalls().isEmpty(),
"a session still BUSY at timeout must preserve its worktree unconditionally");
assertTrue(sessions.roster().isEmpty(), "shutdown drain still deregisters the session");
}
@Test
void sharedTreeReleaseMakesNoWorktreesCalls() {
FakeHerdr herdr = new FakeHerdr();
FakeWorktrees worktrees = new FakeWorktrees();
SessionManager sessions = new SessionManager(workerService(herdr), worktrees);
WorkerSession s = sessions.acquire("ltms-local", null, "/caller/proj", null);
MemberSession s = sessions.acquire("ltms-local", null, "/caller/proj", null);
sessions.release(s.paneId());
@@ -155,9 +248,9 @@ class WorktreeSessionManagerTest {
FakeWorktrees worktrees = new FakeWorktrees().withRepoRoot("/repo").withPrefix("/wt");
SessionManager sessions = new SessionManager(workerService(herdr), worktrees);
WorkerSession a = sessions.acquire("ltms-local", null, "/caller/proj", null,
MemberSession a = sessions.acquire("ltms-local", null, "/caller/proj", null,
new WorktreeRequest("cb-444", null));
WorkerSession b = sessions.acquire("ltms-local", null, "/caller/proj", null,
MemberSession b = sessions.acquire("ltms-local", null, "/caller/proj", null,
new WorktreeRequest("cb-444", null));
assertNotEquals(a.branch(), b.branch(), "branches are distinct");
@@ -204,7 +297,7 @@ class WorktreeSessionManagerTest {
FakeHerdr herdr = new FakeHerdr();
FakeWorktrees worktrees = new FakeWorktrees().withRepoRoot("/repo").withPrefix("/wt");
// Argument order matters: configDir is the 4th parameter, cwd the 11th (after mcpUrl).
BridgedConfig.Worker cfg = new BridgedConfig.Worker(
BridgedConfig.Profile cfg = new BridgedConfig.Profile(
"ltms-local", "http://gx00.gw:8000", "coder", null, "BRIDGED_WORKER_TOKEN",
List.of("ccs", "ltms-local"), "tab", "bridged-workers",
"worker: {profile} #{n}", null, "/pinned/dir", null);

Some files were not shown because too many files have changed in this diff Show More