Two changes ship together here.
1. One shared herdr workspace. The lead and every worker now live in one
workspace called "fleet", so the operator sees one "session" with many
windows, not two. Before, the lead sat in a "leads" workspace and workers
in "bridged-workers", which read as two sessions. The lead is still told
apart from workers by its exact tab label ("lead: <name>"), so putting them
in one space is safe. LeadTabScanner keeps the exclude-by-label mechanism
for split layouts; Fleetd now passes an empty exclude set.
2. Rename the daemon from "bridged" to "fleetd" (the binary, config, scripts,
launchd/systemd units, module dir, and MCP mount).
- Module dir bridged/ -> fleetd/; jar finalName -> fleetd.jar.
- Log line, comments, docs, and CLAUDE.md updated to say fleetd.
- Scripts renamed: redeploy-bridged.sh -> redeploy-fleetd.sh,
bridged-launchd-wrapper.sh -> fleetd-launchd-wrapper.sh.
- Deploy units renamed: dev.ltms.bridged.plist -> dev.ltms.fleetd.plist,
bridged.service -> fleetd.service; launchd Label -> dev.ltms.fleetd.
- Config default bridged.yaml -> fleetd.yaml; the legacy bridged.yaml is
still read as a fallback, and still gitignored.
- MCP: drop the deprecated bridge_* tool twins; only fleet_* remain. The
server name is "fleet". The mount name in the local .mcp.json becomes
"fleet" (gitignored, not in this commit).
- Env var defaults BRIDGED_API_TOKEN -> FLEETD_API_TOKEN, fixture
BRIDGED_WORKER_TOKEN -> FLEETD_WORKER_TOKEN.
Kept on purpose: the BRIDGED_MEMBER marker. Renaming it is a coupled change to
the credential-scrub security control (an operator secrets.sh may guard on it),
so it stays until that migration is done on its own.
Metrics were already fleet_* (CB-632); MetricNamesTest still guards that no
name says bridged_.
The canonical CLAUDE.md block and the wiki template stay byte-identical
(wiki working tree edited, committed to the wiki repo separately).
949 tests pass (mvn clean install). 4 fewer than before = the 4 removed
bridge_* alias tests.
Unit 3 renamed bridged.example.yaml to fleetd.example.yaml and taught Fleetd to
read fleetd.yaml first. Two things it left behind.
bridged/.gitignore still ignored only bridged.yaml. An operator who follows the
new comment and copies the example to fleetd.yaml gets an untracked live config
holding tokens, and git offers to commit it. Both names are ignored now, because
both names work until the cutover.
Three design docs still pointed readers at bridged.example.yaml, a file that no
longer exists under that name.
Part of #145 (CB-632). Doing this NOW, ahead of the rest of the path
renames, for one reason: the vms lead is about to wire a monitoring
dashboard to these series. Renaming a metric after a dashboard points at
it breaks continuity and silently leaves a dead panel. Renaming it before
costs nothing, so it goes first rather than at the cutover.
Nine series renamed, all declared in FleetMetrics.
Two real defects found while doing it:
- FleetApp had "bridged_auth_failures_total" written as a LITERAL
instead of using FleetMetrics.AUTH_FAILURES -- a second hand-written
copy of a name, which is how these drift. It now uses the constant,
so there is one source for that name again.
- Nothing guarded the prefix. One test does assert a wire name
(FleetAppAuthTest checks the real /metrics body for fleet_sessions),
which is good, but it covers one series out of nine. The literal
above was covered by nothing at all.
So this adds MetricNamesTest, which reads the constants reflectively
rather than listing them -- a test that lists the nine names is itself a
second hand-written copy, and would pass while a tenth went unchecked.
It asserts its own denominator too: "no name starts with bridged_" is
true of an empty set, so a sweep that found nothing would pass loudly.
Asserting the count of 9 makes a broken sweep fail instead.
Also renamed three herdr contract-test workspace labels, __bridged_* ->
__fleet_*. Those are throwaway workspaces created by ensureWorkspace, not
metrics, but they are the same word.
Verified: mvn clean install green, 52 classes, 881 tests, 0 failures.
Test count is up by 3 -- the new prefix, denominator and uniqueness
checks. No bridged_ string remains anywhere outside wiki/.
Part of #145 (CB-632). Documentation only, plus one internal literal.
Unit 1 renamed the package and classes, which left every doc describing
classes that no longer exist. This fixes the prose across README.md,
docs/ and bridged/docs/ -- 18 files.
Renamed: dev.ltms.bridged -> dev.ltms.fleet, the five class names, and
"bridged" where it names the daemon as a product rather than a path.
Also renamed two literals, because a doc that disagrees with the code is
worse than one that is out of date:
- bridged-local-noauth -> fleetd-local-noauth. A placeholder apiKey
OpenCodeLauncher sends when a profile resolves no token, to a local
endpoint that does not check it. No test asserts the old string.
- the vnd.ltms.bridged.* media type in the M4 design doc. It appears
in no Java file, so nothing implements it yet.
Deliberately NOT renamed, because each is still literally true today and
changes only at the cutover:
- paths: bridged/, bridged.yaml, bridged.example.yaml, bridged.jar,
.bridged-worktrees, deploy/dev.ltms.bridged.plist,
scripts/redeploy-bridged.sh, bridged-launchd-wrapper.sh
- bridged_* metric names -- renaming these after the monitoring is
wired would break dashboard continuity, so they move before it is
- bridge_* MCP tool names, which answer alongside fleet_* on purpose
- BRIDGED_* environment variables, read by a file outside this repo
Method note: perl, not sed. BSD sed has no \b and no lookaround, and a
word-boundary expression there fails silently. The prose replace uses
(?<![\w./-])bridged(?![\w./-]) so it cannot touch a path or an
identifier, then every remaining hit was read by hand.
Verified: mvn clean install green, 51 classes, 878 tests, 0 failures.
The repo moved lms/claude-bridge -> fleet/claude-bridge -> fleet/fleetd.
Gitea redirects hold, so most of this is not urgent, but one line was a
real break: the implementer skill posts a worker's PR to a hardcoded
repo path, so every worker PR would have gone to the old address.
.claude/skills/implementer/SKILL.md the worker PR endpoint (functional)
.gitmodules wiki submodule URL
CLAUDE.md + wiki/7-Use-Cases.md the canonical block, kept byte-identical
README.md clone command and wiki link
deploy/bridged.service Documentation=
plugin/.claude-plugin/plugin.json homepage + repository
docs/*.md issue and wiki links
The wiki is not a separate repo. /repos/lms/claude-bridge.wiki returns 404
and lms owned no .wiki entity, so the wiki moved with the repo; both the old
and the new wiki SSH URLs resolve to the same sha. Ticket step 4 assumed a
second transfer that does not exist.
1. opencode.json — mount key bridged -> fleetd. Unit C could not do this:
the file is neutralized by the worktree overlay, so every worker sees a
stub and correctly reported it held no mount key. Tracked as CB-628.
2. e2e/bridge_ask_transcript.md keeps its name. It is a dated record of a
run on 2026-07-16 that really did call bridge_ask, and its first line
says so. The harness now writes fleet_ask_transcript.md for new runs;
its header already says 'Live fleet_ask', so the two now agree.
3. docs/MCP-Contract.md line 20 — the historical banner describes section
6, and section 6 now says fleet_ask.
Rename the eleven MCP tool names (bridge_ack/ask/list/poll/profiles/reply/
send/spawn/status/stop/whoami) to their fleet_* names across the Markdown
documentation. fleet_* is written as the normal name; one deprecation note
in README.md says bridge_* still works for one release.
docs/MCP-Contract.md is renamed only inside section 6 (lines 210-293):
sections 1-5 and 7-11 are stale pre-build design text (CB-609) and are
deliberately left with old names so dead text does not look maintained.
Also renames e2e/bridge_ask_transcript.md to e2e/fleet_ask_transcript.md
to match its content. CLAUDE.md, wiki/, plugin/skills/setup/SKILL.md and
.claude/skills/port-to-opencode/SKILL.md are owned by other units and are
untouched.
A page audit found the dead host still named in four places outside the wiki.
ollama.ltms.dev no longer exists; the one front door is llm.ltms.dev, and the
path differs by backend - /anthropic for claude-code, /v1 for opencode.
Anyone following README.md today points a member at nothing.
README.md also now links the new wiki chapter 13, the operator guide, since
the README's own next step used to be the never-written Setup page.
Not touched: SubscriptionGuard.java:14 still names ollama.ltms.dev, but only
in a javadoc line describing the Stage-1 example allowlist. It is a comment
about history, not a live default, so it stays until that class is next
edited for its own reasons.
The instruction surface named a parameter the tool refuses. bridge_ack requires
target and msgId (BridgeMcp.java:1001) and errors with 'target and msgId are
required' otherwise, but step 5 of the primary procedure said bridge_ack{ticket,
msgId}. The table two sections below it was already correct, so the file
contradicted itself and the wrong half was in the numbered steps a lead follows.
The same line is fixed in the byte-identical wiki template.
docs/MCP-Contract.md still opened with 'Greenfield - no MCP code exists yet' and
CLAUDE.md pointed every session at it. Audited against BridgeMcp.java: it names
two tools that do not exist, omits four that do, gets nearly every parameter name
wrong, and uses /workers paths the daemon does not serve. Its status banner now
says so and names the live schema as the authority; CLAUDE.md's pointer is
narrowed to section 6, which is the part that did survive. Rewrite is CB-609.
systems/vms moved the LLM route timeout again, 1800s -> 86400s (24h), after
the silent-truncation risk was discussed. They tried request: 0s first: it
removes the total-duration timer, but on an AIGatewayRoute the idle timeout
is derived from the request timeout, so 0s also removed any bound on a
stalled connection.
At 86400s our own MessageService.ASYNC_TIMEOUT_MS (30 min) binds first, so a
runaway request now ends as a clean FAILED ticket we raised instead of a
silently truncated 200. While the gateway sat at 1800s the two numbers were
equal and did not nest.
`local` runs on /anthropic and `gx` on /v1, both weight 100; `local-direct`
stays weight 0 as the escape hatch.
systems/vms fixed both blockers, and each was re-checked from this side rather
than taken on trust:
listener buffer 32 KiB -> 32 Mi ours: 1.2 MB body -> 200 (was 413)
LLM route timeout 60s -> 1800s ours: 101s stream -> 200,
message_stop present, 4000/4000
Neither was deliberate: 32 KiB was Envoy Gateway's default
per_connection_buffer_limit_bytes, and 60s was Envoy AI Gateway's own default.
The 60s bounded GENERATION as well as prompt size — a tiny prompt with a long
answer returned 504 at 60.05s.
Verified with real workloads, not liveness probes. A `local` member read this
document and CLAUDE.md in full — 48,344 bytes of file content, comfortably past
the old 32,768 ceiling — and answered four questions correctly, including the
document's length (said ~456, actual 455). A `gx` member did the same. The
trivial 3-question probe is what hid the 32 KiB ceiling for an afternoon, so it
no longer counts as proof here.
§7.2 is new and is the part that matters later. One risk is ACCEPTED, not
solved: on a mid-response timeout over chunked HTTP/1.1, Envoy ends the chunked
encoding cleanly instead of resetting, so a truncated answer arrives as HTTP 200
with no error and no terminator (envoyproxy/envoy#17186, acknowledged 2021,
never fixed; the Dec 2025 fix#42269 is HTTP/2 only and SSE here is HTTP/1.1).
Measured at the old 60s: 200, 61.07s, 2473 of 4000 emitted, message_stop 0,
error events 0, ending on a well-formed frame.
The recommended defence — reject a stream with no terminator — does NOT
transfer to us: Claude Code and opencode are third-party clients and we do not
own their SSE parsing. So this is acceptable because a request would have to run
1800s to trip it, not because we could detect it. If a member ever returns a
confident but truncated answer, suspect this before anything in our own code.
Also recorded, from the upstream bisection: ClientTrafficPolicy is honoured in
standalone `aigw run` but BackendTrafficPolicy is silently ignored, and nothing
external distinguishes them (envoyproxy/gateway#9513). Same silent-default shape
this repo keeps hitting.
bridged.yaml carries the same notes inline (gitignored, so not in this commit).
Refs: gitea #76
I wrote "Caddy request_body max_size and/or Envoy's own" and marked it
unverified. The Caddy half was wrong, and an unverified guess still points the
next reader at the wrong component.
Confirmed by the systems/vms side: Envoy Gateway defaults a listener's
per_connection_buffer_limit_bytes to 32768, and aigw buffers the WHOLE request
body before it can route on the model name. So that default is not a network
tuning knob — it is a hard ceiling on prompt size. From the live config_dump:
listener default/llm/http per_connection_buffer_limit_bytes: 32768
Nobody chose 32 KiB; it was inherited.
Both TLS edges are innocent, and the technique that showed it is better than
mine: both 413s carry x-llm-consumer, a header their auth proxy sets only AFTER
authenticating, so the body cleared both edges and the auth. On llm.vm, aigw
413s at 39 KB while the vLLM backend answers 200 at the same size. I found the
boundary; they found the component, by reading the failure's response headers.
Consequences recorded in the doc:
* DO NOT plan around 32 KiB. The intended ceiling is far higher, so sizing our
profiles to it would be designing around a bug.
* Their fix (ClientTrafficPolicy, bufferLimit: 8Mi) is written but NOT
deployed, pending their operator's approval. We do not re-test until they
confirm — a half-changed system gives a number neither side can trust.
* In standalone `aigw run` a SecurityPolicy is accepted and then silently
ignored, so "the config was accepted" proves nothing there. They will verify
by re-reading the live config_dump and sending a large request. Same
silent-default shape this repo keeps hitting, one layer down.
bridged.yaml carries the same correction (gitignored, so not in this commit).
Refs: gitea #76
Deployed U1-U2c, restarted, spawned both new profiles for real, then reverted.
llm.ltms.dev answers HTTP 413 above 32 KiB (32768 bytes), on BOTH surfaces:
/v1 32695 bytes -> 200 /anthropic 32095 bytes -> 200
/v1 32795 bytes -> 413 /anthropic 32855 bytes -> 413
That is far below one agent turn. It is an edge limit (Caddy request_body
max_size, and/or Envoy), so the fix is in systems/vms, not here.
The part worth recording is how it nearly passed. Two members, same message,
same moment: `local` finished in 66s, `gx` never finished at all. `local`
passed only because the probe was three trivial questions in a fresh session,
so the request fit under 32 KiB — the profile looked healthy and was a
landmine set to fire on the first turn that reads a file. So §7's checklist
was not wrong, it was too easy; it now says to use a file-reading task.
opencode's failure mode is worse than a crash: it catches the 413, compacts
its context, retries, and loops. Observed 10+ minutes BUSY with no reply. From
the lead's side that is indistinguishable from a slow worker. Reproduced
outside the bridge with the launcher's own generated config, which is how it
became a one-line error instead of a hang; §7.1 records that procedure.
Everything else about the migration checked out and is recorded so it is not
re-tested: token accepted on both surfaces, unauthenticated 401 (the Caddy
proxy does gate, whatever the gateway's own fail-open policy does),
/v1/models exactly ["deepseek-v4-flash"], the guard allowlist accepted
llm.ltms.dev, and the generated opencode provider block is correct with a real
llmk- key.
Also answers §3b's open question: reasoning survives BOTH surfaces —
/anthropic returns a real "type":"thinking" block and /v1 returns a populated
reasoning_content. The feared /v1 translation loss did not happen.
Config state (bridged.yaml is gitignored, so it is described rather than
committed): `local` back on http://gx00.gw:8000, `gx` kept at weight 0,
`local-direct` kept, llm.ltms.dev left in the guard allowlist. The file
carries these numbers and the exact two-key edit to switch back.
Verified after the revert with a task that reads two large files: correct on
all three questions. Daemon pid 66745, jar f1fd659423e6.
Refs: gitea #76
The gateway (llm.ltms.dev) replaced Bifrost on 2026-08-15 and serves an Anthropic
surface and an OpenAI surface, so both member kinds can point at it. The plan is in
docs/CB-591-Gateway-Migration.md; gitea #76 tracks the work.
The opencode half needs no code: OpenCodeLauncher already pins an OpenAI-compatible
endpoint (CB-508), so baseUrl + tokenEnv + provider/model is a config change. That
matters more than it looks — every opencode member today is sol or terra, and both sit
on one OpenAI account via credentialId: openai-shared, so an exhaustion on either locks
out both. A gateway-backed opencode profile is free and off that credential, which
retires a single point of failure rather than only adding capacity.
Also extends the redeploy script's --check to AI_GATEWAY_TOKEN. A profile's tokenEnv is
resolved from the DAEMON's own environment by HerdrPeerLauncher.resolveEnv, so a token
added to secrets.sh after the daemon started is simply absent: the launcher injects an
empty token and the gateway answers 401, long after the restart and with nothing tying
the two together. That is the same trap as WORKER_GITEA_TOKEN, and it gets the same
login-shell check that never prints the value.
Unit 2 was written as one block, but CB-571/576/580 and unit 2a have since
built parts of it. A symbol survey of main source plus the merge history puts
it at 1 of 21 criteria done, with three partials.
Criterion 15 says a normal COMPLETED release removes the worktree. Since
CB-576 that is false on purpose — a dirty worktree is preserved, because
deleting it destroys work that cannot be recovered. The criterion is wrong,
not the code.
The criterion required the token to bind the session turn number. Three
independent refusals from the implementer showed why that is not
implementable at this layer: MessageService owns acceptance but never
learns of delivery, and CompletionResolver.onDelivered runs before
SessionManager.onDelivered, so the turn number does not exist yet at the
only point the token could capture it.
Records both rejected alternatives and why, so the next reader does not
re-derive them: a target-keyed registry restores the ambiguity the token
exists to remove, and injecting a turn counter couples layers to fill a
field nothing reads yet.
Four new resolved decisions (§7.7-7.10): turn state split by where the signals
are, with ABANDONED explicitly belt-and-braces over the waiter timeout; the
dual ack model with spawn idempotence by construction (gid stored IN the herdr
pane — the load-bearing detail of the no-ledger position — plus an in-flight
reservation for redelivery during a slow spawn); enforced publish semantics
(confirms + mandatory on a separate channel, return-before-confirm caveat);
queue lifecycle = session lifecycle with .v2 names for the redeclare hazard.
§8 reworked: global id scheme resolved and moved up; control authorization
sharpened into the hard gate on U4 (per-host allowlist beside the peer keys);
key distribution/rotation added. New §9: implementation order, each step
verifiable single-host, U4 gated, U8 last.
Wiki pointer bumped to 4320c1c (chapter 10 same-pass changes).
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_013ZGgxLQ2VpwZhEYoru8rkf
Six decisions recorded in §7, replacing the matching open questions: signed
messages (identity from the key, extending identity-from-connection across the
broker), target-host-owned profiles advertised via presence, repo provisioning
by pinned forge clone, live-only asks with TTL + TOO_LATE notice, spawn-id
dedup on the target, and broker-outage semantics (local unaffected, remote
fails fast, gateway stays soft-state). §8 keeps what is genuinely still open,
with control *authorization* now separated from the resolved *authenticity*.
The broker-level half (U8 broadcast, exclusive consumers, inbox caps, TLS,
schema version, trace id) lands in wiki chapter 10 §10 — pointer bumped
(also picks up 1710a77, chapter 11 Features).
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_013ZGgxLQ2VpwZhEYoru8rkf
Bump bridged to 1.0.0 and add the release notes: the single-leader,
single-host scope is closed — gateway, lifecycle, two-way delivery,
pluggable peers, auth/authz/audit, supervision, CI. Cross-host
federation (CB-308) is the next major line.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_013ZGgxLQ2VpwZhEYoru8rkf
Closes the one known-unverified item before cross-host. CB-402 merged in
ded226a with increment 5 (the §5 live checklist) deferred; it has now run.
Provider question (§7 Q1) resolved with no credentials needed: opencode's own
gateway serves free-tier models. `opencode auth list` reports 0 credentials,
yet `opencode run -m opencode/north-mini-code-free` answers. Distinct from the
primary's subscription by construction, and needs no guard entry — opencode
carries no ANTHROPIC_BASE_URL, so SubscriptionGuard never applies to it.
The schema-drift risk was the real one and it did not bite. The adapter was
designed against opencode 1.1.31; installed is 1.18.5. The generated config
still validates unchanged (type:"remote" + instructions:[path]), and
`OPENCODE_CONFIG=… opencode mcp list` reports the bridge connected. Pinned as
a verified fact for 1.18.5.
Full lifecycle through REST: spawn (201, kind-routed to OpenCodeLauncher) ->
CB-306 gate passed ~0.6s -> ready -> send -> {"replySource":"reply"} (a
STRUCTURED bridge_reply, not the CB-115 completion fallback) -> delete (204,
tolerant teardown).
Unplanned cross-validation with CB-501: the audit trail recorded the reply as
role=WORKER actor=worker:term_657c… — connection-based identity classified an
opencode process as a worker with no opencode-specific handling. The identity
model is peer-kind-agnostic, which is what CB-308 needs when the roster
stretches across hosts.
Stage 5 verified live on the same run: /workers (CB-304) answers where the
13-day-old daemon 404'd, /metrics counted the delegation
(sends_total{outcome=replied} 1, replies_total{path=rendezvous} 1,
inbox_depth 0), and the audit log captured SPAWN/SEND/REPLY with correct roles.
Adds the opencode-free dogfood profile to the local bridged.yaml (gitignored;
recorded here for reproducibility) and docs/CB-402 §8 as-built.
Closes out single-host before the cross-host work. Sequenced BEFORE CB-308
deliberately: federation's own gating concern is the trust model, and it
inherits whatever identity shape lands here.
The finding this stage is built around: bridged had exactly ONE security
control, the loopback bind. ConnectionIdentity resolves a worker from its
connection (unforgeable), but every caller that was not a recognised worker
pane fell through to being treated as the PRIMARY -- the most privileged role
on the bus. Latent today; load-bearing the moment a bind widens.
CB-501 auth:
- Role/Principal/CallerResolver: connection identity first, bearer token
second, ANONYMOUS third. Inverts the old default so absence of identity
means nothing, not everything.
- Worker identity is never token-gated, so enabling auth cannot lock the
fleet out of bridge_reply.
- Constant-time token compare (MessageDigest.isEqual).
- validateAuthExposure(): a non-loopback bind under loopback-trust now
REFUSES TO START. Makes the dangerous config unrepresentable rather than
merely documented.
- TLS terminates at a reverse proxy by design (D3), not in the JVM.
CB-505 authz + audit, enforced on BOTH entry paths:
- The docs describe MCP as "a thin adapter over the REST core"; at code level
it is not. BridgeMcp calls MessageService directly, and /mcp is a raw
servlet on Jetty's context handler that never traverses Javalin's before
filter. Enforcing only at REST would have left /mcp open.
- Load-bearing rule is own-session-only: a worker may reply/ask only as
itself. Structurally true over MCP already; over REST the session id in the
URL path had simply been trusted.
- Audit: JSON lines to a dedicated appender, additivity=false. Never records
message content -- this bus carries source and prompts.
CB-502 metrics: zero new dependencies. A ~150-line Prometheus text renderer
instead of the specced Micrometer, because this pom already hand-pins
jackson-annotations to reconcile Jackson 2/3, imports a Jetty BOM against
skew, and carries four accepted-CVE advisories -- and CLAUDE.md's mandated
dependency CVE gate could not be run (no JetBrains MCP server connected).
Instrumented at MessageService, the single funnel both surfaces share.
CB-503 CI: .gitea/workflows/ci.yml against the already-running Gitea runner.
Needs no contract-exclusion flag -- the pom's default-excludes profile
already sets excludedGroups=contract, so plain `mvn clean install` IS the
mock-socket surface. Provisions JDK 25 explicitly (runner default-jdk is older).
CB-504 supervision: launchd agent (the real target -- this host is macOS,
there is no systemd) plus a systemd unit for the Linux gateways CB-308 adds.
Ordering directives are advisory, so the actual fix is that startup now waits
up to 30s for the herdr socket and then serves degraded, instead of crashing
into a restart loop on a boot-order race.
Also fixes drift found while surveying:
- bridged.example.yaml documented spawn_ready_timeout_ms in snake_case; config
binds via plain Jackson with ignoreUnknown, so uncommenting it would have
been silently dropped and the default kept. Now camelCase, with a test that
loads the shipped example and one that pins every documented knob's
spelling -- no test had ever loaded that file.
- Added the 6 shipped-but-undocumented knobs (worktreeRoot, parityOverlay,
gitTokenEnv, gitHostEnv, configDir, primary:).
- README "Next" listed bridge_ask and session lifecycle as upcoming; both
shipped long ago.
- docs/CB-301-ext and docs/CB-402 status headers said "design"/"pre-
implementation" for work already merged.
307 unit/acceptance tests green (was 266), mvn clean install BUILD SUCCESS.
Note: CLAUDE.md's per-file ide_diagnostics gate and the pom Mend.io CVE check
could not be run -- no JetBrains/intellij-index MCP server is connected this
session. mvn clean install is the only gate that ran.
Clarifies the open fork from §4/§9: Development A (sandbox launcher) and CB-308
(per-host federation) COMPOSE — each host runs a bridged gateway whose launcher
spawns agents into that host's LOCAL sandboxes; the broker moves messages +
presence, never keystrokes.
The forcing fact: delivery is herdr keystroke-injection into a locally-owned PTY,
so "a sandboxed agent on another host" ≡ "a sandbox spawned by that host's
gateway" (a remote container with no local herdr can't be injected into). Rules
out a central daemon reaching remote PTYs.
Adds Figure 11 (composed topology) + Figure 12 (remote-delegation sequence: the
local?inject:publish fork with a sandboxed far side — both injection points stay
local, only the middle hop crosses the broker), the two forced reachability
changes (host-routable mcpUrl; PTY in the local gateway's herdr), and a
maps-to-existing-seams table (CB-308 gateway × Dev-A launcher, CB-117 reap,
CB-303 container lifecycle, CB-308 #5 trust). No new pillars. Both diagrams
mmdc-validated; §4/§9 updated to point at §11.
Scopes the lead's next direction on top of the peer-launcher arc, as a
proposal (ticket split deferred):
- A · sandboxed, role-specific workers — a sandbox is a *placement*, so it
slots into the CB-401 SPI as a new kind: sandbox adapter exactly the way
CB-402's opencode slotted in as a new provider (SPI proven placement-neutral,
not just provider-neutral). Ownership line held: bridge launches INTO a
peer-owned image, never provisions the IDE/dev-tools inside it.
- B · main-agent pairs (Opus + cloud) — both mains are MCP clients so neither
can be called into; each needs a pull inbox → depends on CB-308's per-agent
channels. PrimaryRegistry single-slot → multi-slot.
- C · orchestrator tier — SessionManager recursed one tier up
(orchestrator:mains :: main:workers) + context scoping; re-roots the human
from a live primary to the orchestrator.
10 mermaid diagrams (component + sequence per development, ownership guardrail,
staging graph), all mmdc-validated and theme-safe. §7 pins the bus-vs-env-manager
boundary as an acceptance criterion on A; §8 stages A → CB-308 substrate → B → C.
Design-only. Audits the as-built Claude/herdr coupling (concentrated in
WorkerService), defines a PeerLauncher SPI + opaque PeerHandle + capability
model so the bus delegates peer materialization to a config-selected adapter.
ClaudeCodeLauncher = adapted WorkerService. Stages A/B/C with the Stage-C
plugin-loading security gate called out. No production code touched.
Worktree is code-only isolation; SessionManager hydrates it to full config
parity (overlay untracked local settings/.mcp.json/.env) so a worker is a
full peer of the primary, differing only in the LLM provider. Worker commits,
pushes over SSH, and opens its own PR (gitea REST + repo-scoped token).
The as-built audit surfaced that WorkerService keeps no registry of what it
spawned (its own Javadoc: 'there is no registry; list() only asks herdr').
CB-301 adds a SessionManager wrapping WorkerService: an authoritative in-daemon
roster with a per-session lifecycle FSM (SPAWNING/READY/BUSY/DONE/RELEASED/
FAILED), deterministic release, and recycle (= release + fresh acquire, no
reuse). Leaves clean seams for CB-302 (checkpoint on release), CB-303 (idle_ttl/
context_cap/drain policy over roster), CB-304 (bridge_list reads roster).
A worker now opens the same directory the primary is in, unless told otherwise.
Resolution: explicit spawn cwd → per-profile config cwd → the primary's cwd
(auto-detected from the bridge_spawn caller's PID via lsof -d cwd) → the daemon's
cwd. Never $HOME.
Mechanism (found by live probe, corrects the earlier assumption): an agent.start
pane does NOT inherit its tab's or workspace's cwd — it starts in $HOME. herdr's
agent.start honours an (undocumented) cwd param, so the resolved cwd is threaded
onto agent.start {cwd} (both tab and pane placement), not tab.create.
Surfaces: bridge_spawn {cwd?} + auto-detect via ConnectionIdentity.resolve (peer
PID) + ProcessCwdLookup (lsof); REST POST /workers ?cwd= / body cwd; per-profile
'cwd:' config. Validated live: explicit cwd → worker rooted there; no cwd over
REST → daemon cwd, not $HOME. Also clears the ccs folder-trust prompt when the
project dir is already trusted (see docs/Worker-Startup-and-Trust.md).
Documents (a) the rule that a worker inherits the PRIMARY's working directory,
never $HOME — resolved from an explicit cwd, else the bridge_spawn caller's PID
(lsof -d cwd), else the daemon cwd; herdr's seam is workspace.create {cwd}, since
agent.start has no cwd; (b) ccs (Claude Code) as the assumed launcher and its
folder-trust model (hasTrustDialogAccepted per project in each instance's
.claude.json), so trusting the project dir once per profile clears the prompt;
(c) that other CLIs have their own startup gates, documented per launcher.
Marks the cwd-inheritance as target design (follow-up), not yet wired.