Compare commits

..

6 Commits

Author SHA1 Message Date
Dai Ha 6f5b3a1fd6 CB-528b: provision an isolated CODEX_HOME for a Codex peer 2026-08-10 21:25:12 +02:00
Dai Ha ef49835c4f CB-528: the CodexHome seam, ahead of the adapter that uses it
Codex reads everything from CODEX_HOME — config, credentials, sessions, skills,
plugins, state. Pointing a peer at the operator's own ~/.codex would hand it the
operator's tool surface and let it write into the operator's session history:
the same failure CB-525 exists to prevent on the Claude side, in a runtime where
there is no --mcp-config to neutralize.

Extracted as an interface rather than a launcher method because provisioning is
filesystem work with its own failure modes. The common one is a missing
credential, which Codex surfaces as an opaque 401 mid-turn instead of a spawn
error — so the contract says provision() must fail loudly there. Splitting it
also lets the launcher be tested without touching a real home directory.

Lands before the adapter so the launcher and the provisioner can be built
against a fixed seam instead of against each other.
2026-08-10 20:47:26 +02:00
Dai Ha ef1e014b41 CB-527: ship the bridge as an installable Claude Code plugin
CI / build (push) Successful in 1m0s
CI / contract (push) Successful in 1m21s
The orchestration contract had no distributable form. Every consuming project
hand-copied a block of CLAUDE.md and hand-wrote an .mcp.json, and we keep a
script whose only job is to notice those copies drifting apart. A plugin is
versioned, installed once, and updates in place.

Ships no credentials, deliberately: every secret is referenced by environment
variable NAME and the value never enters a file, which is what makes the
artifact safe to publish. The setup skill states the two rules that are easy to
get wrong for the right-sounding reasons — the PR token must not be able to
merge (a worker opens, the primary gates), and ANTHROPIC_BASE_URL must never be
set by setup, because mounting the bridge must not move a session off
subscription.

The plugin root is plugin/, not the repo root. An installed plugin's .mcp.json
is a committed file, while this repo's root .mcp.json is local-only and
--skip-worktree; rooting the plugin at the repo would commit the primary's IDE
servers and hand them to every worker — the exact failure CB-525 exists to
prevent.

Scope is client-side setup only. herdr and bridged stay separate services with
their own lifecycles, and the skill refuses to install them rather than guess.
It also refuses to accept /healthz as proof: health reports only that the daemon
can reach herdr, and CB-521 showed it staying green while every spawn failed, so
verification ends with a real spawn.

Both manifests pass `claude plugin validate --strict`.
2026-08-10 19:55:17 +02:00
Dai Ha e3c8393d1b CB-527: retire the ollama profile from the worker choice
The ollama backend is decommissioned, so the example config stops pointing
readers at a dead host and the guard allowlist stops carrying an entry with
no profile behind it — a stale entry there is dead permission, and that list
is the only thing keeping a worker off the primary's subscription.

The second illustrative profile survives as gx11: the example exists to show
`placement: weighted` having something to choose between, and a one-profile
example would quietly stop demonstrating that.

It also moves the CB-523 auto-compact override onto the surviving profile.
That guard had been attached to `ollama` alone, so retiring the profile would
have removed the fleet's only protection against the failure it was written
for — a worker whose prompt is rejected before auto-compact ever fires. The
window belongs on every profile, not on whichever one happened to hit it.
2026-08-10 19:55:17 +02:00
kevin 379e03f9d0 CB-308: fold in the adversarial review; bump wiki to 4320c1c
CI / build (push) Successful in 59s
CI / contract (push) Successful in 1m16s
Four new resolved decisions (§7.7-7.10): turn state split by where the signals
are, with ABANDONED explicitly belt-and-braces over the waiter timeout; the
dual ack model with spawn idempotence by construction (gid stored IN the herdr
pane — the load-bearing detail of the no-ledger position — plus an in-flight
reservation for redelivery during a slow spawn); enforced publish semantics
(confirms + mandatory on a separate channel, return-before-confirm caveat);
queue lifecycle = session lifecycle with .v2 names for the redeclare hazard.

§8 reworked: global id scheme resolved and moved up; control authorization
sharpened into the hard gate on U4 (per-host allowlist beside the peer keys);
key distribution/rotation added. New §9: implementation order, each step
verifiable single-host, U4 gated, U8 last.

Wiki pointer bumped to 4320c1c (chapter 10 same-pass changes).

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_013ZGgxLQ2VpwZhEYoru8rkf
2026-08-10 22:32:00 +07:00
kevin e13921aa8a CB-308: resolve the cross-host design review; bump wiki to 36bb865
CI / build (push) Successful in 54s
CI / contract (push) Successful in 1m5s
Six decisions recorded in §7, replacing the matching open questions: signed
messages (identity from the key, extending identity-from-connection across the
broker), target-host-owned profiles advertised via presence, repo provisioning
by pinned forge clone, live-only asks with TTL + TOO_LATE notice, spawn-id
dedup on the target, and broker-outage semantics (local unaffected, remote
fails fast, gateway stays soft-state). §8 keeps what is genuinely still open,
with control *authorization* now separated from the resolved *authenticity*.

The broker-level half (U8 broadcast, exclusive consumers, inbox caps, TLS,
schema version, trace id) lands in wiki chapter 10 §10 — pointer bumped
(also picks up 1710a77, chapter 11 Features).

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_013ZGgxLQ2VpwZhEYoru8rkf
2026-08-10 21:54:30 +07:00
12 changed files with 889 additions and 13 deletions
+18
View File
@@ -0,0 +1,18 @@
{
"name": "claude-bridge",
"description": "Tooling for orchestrating a fleet of delegated coding agents through the bridged MCP gateway.",
"owner": {
"name": "LTMS"
},
"plugins": [
{
"name": "claude-bridge",
"source": "./plugin",
"description": "Make a project bridge-ready: mount the bridged MCP gateway and apply standard Claude Code settings so a session can orchestrate delegated workers. Ships no credentials.",
"version": "0.1.0",
"author": {
"name": "LTMS"
}
}
]
}
+9 -4
View File
@@ -116,15 +116,21 @@ workers:
# configDir: /Users/me/.ccs/instances/gx10 # CLAUDE_CONFIG_DIR — inherit that profile's skills/MCP
# cwd: /Users/me/src/myrepo # pin the working dir; omit to inherit the primary's
# parityOverlay: [".claude/settings.local.json", ".env", ".envrc"] # never add .mcp.json — see above
ollama:
baseUrl: http://ollama.ltms.dev # local/self-hosted; usually no token
gx11: # a second backend, so `placement: weighted` has a choice
baseUrl: http://gx01.gw:8000 # self-hosted; ccs handles the model + token
placement: tab
workspace: bridged-workers
tabLabel: "worker: {profile} #{n}"
mcpUrl: http://127.0.0.1:8765/mcp
argv: ["ccs", "ollama"]
argv: ["ccs", "gx11"]
weight: 0.5
maxLoad: 2
# Pin an auto-compact window BELOW the served model's context ceiling. The global
# ~/.claude/settings.json value is shared by every ccs instance and the primary, so the
# per-profile override belongs here. Equal to the ceiling means auto-compact never fires
# before the server rejects the prompt, which kills a worker mid-turn (CB-523).
env:
CLAUDE_CODE_AUTO_COMPACT_WINDOW: "280000"
# CB-402: a second coding-agent kind, proving the PeerLauncher SPI is provider-neutral.
# opencode is provider-agnostic and uses NONE of Claude's private seams: no ANTHROPIC_BASE_URL /
# SubscriptionGuard (so it needs no `guard` host entry), no --mcp-config / --append-system-prompt.
@@ -179,7 +185,6 @@ guard:
offSubscriptionHosts:
- gx00.gw
- gx01.gw
- ollama.ltms.dev
# Spawn-readiness gate (CB-306). The launcher blocks until the worker's herdr status is
# injectable (IDLE/BLOCKED/DONE) or the timeout elapses. 0 disables the gate.
@@ -126,6 +126,12 @@ public record BridgedConfig(
public static final String KIND_CLAUDE_CODE = "claude-code";
/** Peer kind spawned by the opencode adapter (CB-402). */
public static final String KIND_OPENCODE = "opencode";
/**
* Peer kind spawned by the Codex adapter (CB-528). Like {@link #KIND_OPENCODE} it carries
* its own argv and never inherits the Claude binary, and it sits outside the
* {@code ANTHROPIC_BASE_URL} subscription guard because Codex has no such seam.
*/
public static final String KIND_CODEX = "codex";
public Worker {
// A claude-code worker defaults its launch command to `claude`; other kinds carry their own
@@ -0,0 +1,48 @@
package dev.ltms.bridged.worker;
import dev.ltms.bridged.config.BridgedConfig;
import java.nio.file.Path;
/**
* Provisions an isolated {@code CODEX_HOME} for one Codex peer (CB-528).
*
* <p>Codex reads <em>everything</em> from {@code CODEX_HOME} — its config, credentials, sessions,
* skills, plugins, and state. Pointing a peer at the operator's own {@code ~/.codex} would give it
* the operator's tool surface and let it write into the operator's session history, which is the
* same class of failure CB-525 exists to prevent on the Claude side. So every peer gets its own
* directory, and this is the seam that builds it.
*
* <p>It is an interface rather than a method on the launcher for two reasons: provisioning is
* filesystem work with its own failure modes (a missing credential is the most common, and it
* surfaces as an opaque {@code 401} from Codex rather than a spawn error), and keeping it separate
* lets the launcher be tested without touching a real home directory.
*
* <p>Three things the implementation must put in the home, because Codex has no launch flag for
* any of them:
* <ul>
* <li>the bridge MCP server, as {@code [mcp_servers.*]} in {@code config.toml};</li>
* <li>the reply charter, as {@code AGENTS.md} — Codex has no {@code --append-system-prompt},
* so the standing instruction has to reach it as a file;</li>
* <li>credentials, since a freshly created home has none and Codex fails closed.</li>
* </ul>
*/
public interface CodexHome {
/**
* Build a fresh, isolated home for a peer launching under {@code cfg} and return its path,
* suitable for the {@code CODEX_HOME} environment variable.
*
* @param cfg the profile being launched; supplies the MCP URL and any bearer-token variable
* @return the provisioned directory
* @throws RuntimeException if the home cannot be provisioned — including when no credential is
* available, which must fail loudly here rather than as a 401 later
*/
Path provision(BridgedConfig.Worker cfg);
/**
* Remove a home previously returned by {@link #provision}. Idempotent: releasing an already
* released or never provisioned path is not an error, because teardown races teardown.
*/
void release(Path home);
}
@@ -0,0 +1,205 @@
package dev.ltms.bridged.worker;
import dev.ltms.bridged.config.BridgedConfig;
import java.io.IOException;
import java.io.UncheckedIOException;
import java.nio.file.Files;
import java.nio.file.Path;
import java.nio.file.StandardCopyOption;
import java.util.Comparator;
import java.util.function.Function;
/**
* The default {@link CodexHome} provisioner (CB-528): builds a fresh, isolated {@code CODEX_HOME}
* for one Codex peer under a configurable root (default the JVM temp dir, mirroring how
* {@link OpenCodeLauncher} picks its config root).
*
* <p>Codex reads <em>everything</em> from {@code CODEX_HOME} — config, credentials, sessions,
* skills, plugins, state — so a peer must never be pointed at the operator's own {@code ~/.codex}.
* Provisioning therefore assembles, in one throwaway directory:
* <ul>
* <li>{@code config.toml} registering the bridge as a streamable-HTTP MCP server (with the
* profile's bearer-token env var, when one is configured);</li>
* <li>{@code AGENTS.md} carrying the reply charter — codex has no {@code --append-system-prompt},
* so the standing instruction has to reach the peer as a file;</li>
* <li>{@code auth.json} copied from the operator's real codex home, since a fresh home has no
* credential and codex then fails closed with an opaque {@code 401} mid-run.</li>
* </ul>
*
* <p>The credential is <em>copied</em>, never symlinked: a peer that could write through a symlink
* would be able to modify the operator's real {@code auth.json}. Each peer gets its own on-disk
* copy, so nothing outside the provisioned home is ever touched on write.
*/
public final class DefaultCodexHome implements CodexHome {
/** Writer for the generated {@code AGENTS.md} (kept as a field for unit-test inspection). */
static final String AGENTS_HEADER = "# Standing instruction\n\n";
/** Root under which per-peer Codex homes are created (injectable for tests). */
private final Path root;
/** Host env lookup used to resolve the operator's codex home (injectable for tests). */
private final Function<String, String> env;
/** Production constructor — homes land under the JVM temp dir and the operator's credentials
* are resolved from the real environment ({@code CODEX_HOME}, else {@code ~/.codex}). */
public DefaultCodexHome() {
this(Path.of(System.getProperty("java.io.tmpdir")), System::getenv);
}
/**
* Full testability constructor: an injectable {@code root} (so tests never touch the JVM temp
* dir's real real estate) and an injectable env lookup (so the credential source can be pointed
* at a temp dir instead of the operator's real {@code ~/.codex}).
*
* @param root directory under which per-peer Codex homes are created (must exist)
* @param env host environment lookup, used to resolve the operator's codex home
*/
public DefaultCodexHome(Path root, Function<String, String> env) {
this.root = root;
this.env = env;
}
/** Standing instruction written to {@code AGENTS.md}. Adapted from
* {@link OpenCodeLauncher#REPLY_CHARTER}: codex has no {@code --append-system-prompt}, so the
* only way a codex peer receives the reply contract is as a file in its home. The substance must
* survive — every message arrives through the bridge, terminal text reaches nobody, so the turn
* MUST end with exactly one {@code bridge_reply} carrying the complete answer. Kept as a single
* logical line so it reads the same way the opencode charter does (Markdown wraps it on render). */
private static final String REPLY_CHARTER =
"You are an off-subscription worker in the claude-bridge fleet, running under codex. "
+ "Every message you receive arrives through the bridge, and the ONLY channel back to the "
+ "sender is the bridge_reply MCP tool. Text you write in your terminal is NOT sent "
+ "anywhere — the sender cannot see your screen, so an in-terminal answer is silently "
+ "discarded. Therefore you MUST end EVERY turn by calling bridge_reply with `content` set "
+ "to your complete response. This holds for every message without exception — tasks, "
+ "questions, clarifications, acknowledgements, and ordinary back-and-forth conversation. "
+ "Call bridge_reply exactly once, as the final action of your turn, with your full answer "
+ "in `content`; never wait for confirmation first. If you end a turn without calling "
+ "bridge_reply, the sender receives nothing and the exchange stalls.";
@Override
public Path provision(BridgedConfig.Worker cfg) {
Path home;
try {
home = Files.createTempDirectory(root, "bridged-codex-");
} catch (IOException e) {
throw new UncheckedIOException("cannot create Codex home for profile " + cfg.profile(), e);
}
try {
if (cfg.hasMcp()) {
writeConfig(home, cfg);
}
writeCharter(home);
provisionCredential(home);
return home;
} catch (RuntimeException e) {
// Every failure below is already a RuntimeException (wrapped IOException, or the missing-
// credential IllegalStateException). Best-effort remove the partially-built home so a
// failed spawn does not leak a temp dir, then rethrow.
try {
release(home);
} catch (RuntimeException ignored) {
// A cleanup failure must not mask the provisioning failure that got us here.
}
throw e;
}
}
@Override
public void release(Path home) {
if (home == null || !Files.exists(home)) {
return; // already released, or never provisioned — teardown races teardown
}
try (var stream = Files.walk(home)) {
// Delete deepest-first so directories are empty when their turn comes.
stream.sorted(Comparator.reverseOrder()).forEach(p -> {
try {
Files.deleteIfExists(p);
} catch (IOException e) {
throw new UncheckedIOException("cannot remove " + p, e);
}
});
} catch (IOException e) {
throw new UncheckedIOException("cannot walk " + home + " for release", e);
}
}
/** Write {@code config.toml} registering the bridge as a streamable-HTTP MCP server. */
private void writeConfig(Path home, BridgedConfig.Worker cfg) {
StringBuilder sb = new StringBuilder("[mcp_servers.bridged]\n");
sb.append("url = ").append(tomlString(cfg.mcpUrl())).append('\n');
if (cfg.tokenEnv() != null && !cfg.tokenEnv().isBlank()) {
sb.append("bearer_token_env_var = ").append(tomlString(cfg.tokenEnv())).append('\n');
}
writeString(home.resolve("config.toml"), sb.toString());
}
/** Write {@code AGENTS.md} carrying the reply charter (always — the standing instruction applies
* to every codex peer, not only those that mount the bridge). */
private void writeCharter(Path home) {
writeString(home.resolve("AGENTS.md"), AGENTS_HEADER + REPLY_CHARTER + "\n");
}
/** Copy {@code auth.json} from the operator's real codex home into the fresh home. A freshly
* created home has no credential, and codex then fails closed with an opaque {@code 401} mid-run
* rather than a spawn error — so a missing source is thrown from here, loudly, naming the path. */
private void provisionCredential(Path home) {
Path source = credentialSource();
if (!Files.isRegularFile(source)) {
throw new IllegalStateException("no Codex credential for the peer: looked for "
+ source + " but it does not exist. A fresh CODEX_HOME has no credentials and the "
+ "peer would otherwise fail with an opaque 401 Unauthorized mid-run; provision one "
+ "at that path or mount the operator's codex home before spawning codex peers.");
}
try {
Files.copy(source, home.resolve("auth.json"), StandardCopyOption.COPY_ATTRIBUTES);
} catch (IOException e) {
throw new UncheckedIOException("cannot copy Codex credential from " + source, e);
}
}
/** The operator's codex credential source: {@code CODEX_HOME}/auth.json when the env var is set,
* else {@code ~/.codex}/auth.json. */
private Path credentialSource() {
String codexHome = env.apply("CODEX_HOME");
Path homeDir = (codexHome == null || codexHome.isBlank())
? Path.of(System.getProperty("user.home"), ".codex")
: Path.of(codexHome);
return homeDir.resolve("auth.json");
}
private static void writeString(Path file, String content) {
try {
Files.writeString(file, content);
} catch (IOException e) {
throw new UncheckedIOException("cannot write " + file, e);
}
}
/** Render {@code s} as a TOML basic string, escaping what TOML requires. Values here are
* operator- or config-supplied (a URL, an env-var name), so escaping must be real — a malformed
* config.toml makes codex fail in a way that looks like a network problem. */
private static String tomlString(String s) {
StringBuilder sb = new StringBuilder("\"");
for (int i = 0; i < s.length(); i++) {
char c = s.charAt(i);
switch (c) {
case '\\' -> sb.append("\\\\");
case '"' -> sb.append("\\\"");
case '\n' -> sb.append("\\n");
case '\r' -> sb.append("\\r");
case '\t' -> sb.append("\\t");
default -> {
if (c < 0x20) {
sb.append(String.format("\\u%04X", (int) c));
} else {
sb.append(c);
}
}
}
}
return sb.append('"').toString();
}
}
@@ -0,0 +1,189 @@
package dev.ltms.bridged.worker;
import dev.ltms.bridged.config.BridgedConfig;
import org.junit.jupiter.api.Test;
import org.junit.jupiter.api.io.TempDir;
import java.io.IOException;
import java.nio.file.Files;
import java.nio.file.Path;
import java.util.List;
import java.util.Map;
import static org.junit.jupiter.api.Assertions.assertEquals;
import static org.junit.jupiter.api.Assertions.assertFalse;
import static org.junit.jupiter.api.Assertions.assertThrows;
import static org.junit.jupiter.api.Assertions.assertTrue;
/**
* The {@link CodexHome} provisioner: a fresh, isolated {@code CODEX_HOME} per peer, holding the
* bridge MCP config, the reply charter ({@code AGENTS.md} — the only way a codex peer receives a
* standing instruction), and a <em>copy</em> of the operator's credential. The credential source is
* always injected (via {@code CODEX_HOME}) so these tests never read the real {@code ~/.codex}, and
* the operator's home is asserted byte-for-byte untouched after provisioning.
*/
class DefaultCodexHomeTest {
private static final String MCP_URL = "http://127.0.0.1:8765/mcp";
/** A codex-profile worker; {@code tokenEnv} may be {@code null} to exercise the record default. */
private static BridgedConfig.Worker codexCfg(String mcpUrl, String tokenEnv) {
return new BridgedConfig.Worker("codex-peer", null, null, null, tokenEnv, List.of("codex"),
"tab", "bridged-workers", "codex: {profile} #{n}", mcpUrl, null, null, null, null,
BridgedConfig.Worker.KIND_CODEX);
}
/** A provisioner whose homes land under a temp root and whose env maps {@code CODEX_HOME} to a
* temp operator home — so every assertion runs against temp dirs, never the real {@code ~/.codex}. */
private static DefaultCodexHome service(Path root, String operatorCodexHome) {
return new DefaultCodexHome(root, key -> "CODEX_HOME".equals(key) ? operatorCodexHome : null);
}
private static String read(Path home, String name) throws IOException {
return Files.readString(home.resolve(name));
}
/** Snapshot of a directory's contents (relative paths + hashes) — for the isolation assertion. */
private static Map<String, String> snapshot(Path dir) throws IOException {
Map<String, String> out = new java.util.HashMap<>();
try (var stream = Files.walk(dir)) {
for (Path p : stream.filter(Files::isRegularFile).toList()) {
out.put(dir.relativize(p).toString(), hex(Files.readAllBytes(p)));
}
}
return out;
}
private static String hex(byte[] b) {
StringBuilder sb = new StringBuilder();
for (byte x : b) {
sb.append(String.format("%02x", x));
}
return sb.toString();
}
@Test
void configTomlRegistersBridgeWithExactUrl(@TempDir Path root, @TempDir Path operator) throws IOException {
Files.writeString(operator.resolve("auth.json"), "{\"token\":\"x\"}");
DefaultCodexHome svc = service(root, operator.toString());
Path home = svc.provision(codexCfg(MCP_URL, "MY_BRIDGED_TOKEN"));
String config = read(home, "config.toml");
assertTrue(config.contains("[mcp_servers.bridged]"), config);
assertTrue(config.contains("url = \"" + MCP_URL + "\""), config);
// The schema is the one codex itself writes for `codex mcp add --url ... --bearer-token-env-var`.
assertTrue(config.contains("bearer_token_env_var = \"MY_BRIDGED_TOKEN\""), config);
}
@Test
void noConfigTomlWhenMcpUrlUnset(@TempDir Path root, @TempDir Path operator) throws IOException {
Files.writeString(operator.resolve("auth.json"), "{\"token\":\"x\"}");
DefaultCodexHome svc = service(root, operator.toString());
Path home = svc.provision(codexCfg(null, "MY_BRIDGED_TOKEN"));
assertFalse(Files.exists(home.resolve("config.toml")));
}
/**
* The bearer-token key is emitted only when {@code tokenEnv} is set. Note the {@code Worker}
* record normalises a blank {@code tokenEnv} to {@code "BRIDGED_WORKER_TOKEN"}
* ({@link BridgedConfig.Worker} compact constructor), so a truly blank token env is not
* constructible through the public API; the negative is therefore exercised by the no-MCP case
* above (no {@code config.toml}, hence no key), and here we pin the emitted value to the
* configured env-var name — including that a {@code null} tokenEnv resolves to the record default.
*/
@Test
void bearerKeyNamesConfiguredTokenEnvAndDefaultsWhenUnset(@TempDir Path root, @TempDir Path operator)
throws IOException {
Files.writeString(operator.resolve("auth.json"), "{\"token\":\"x\"}");
DefaultCodexHome svc = service(root, operator.toString());
Path home1 = svc.provision(codexCfg(MCP_URL, "CUSTOM_TOKEN"));
assertTrue(read(home1, "config.toml").contains("bearer_token_env_var = \"CUSTOM_TOKEN\""));
// null tokenEnv -> record default BRIDGED_WORKER_TOKEN, which is itself a meaningful name.
Path home2 = svc.provision(codexCfg(MCP_URL, null));
assertTrue(read(home2, "config.toml").contains("bearer_token_env_var = \"BRIDGED_WORKER_TOKEN\""));
}
@Test
void agentsMdCarriesTheReplyCharter(@TempDir Path root, @TempDir Path operator) throws IOException {
Files.writeString(operator.resolve("auth.json"), "{\"token\":\"x\"}");
DefaultCodexHome svc = service(root, operator.toString());
Path home = svc.provision(codexCfg(MCP_URL, "MY_BRIDGED_TOKEN"));
String agents = read(home, "AGENTS.md");
assertTrue(agents.contains("bridge_reply"), agents);
assertTrue(agents.contains("running under codex"), agents);
}
@Test
void credentialCopiedFromSourceWhenPresent(@TempDir Path root, @TempDir Path operator) throws IOException {
String auth = "{\"account\":\"dai.ha\",\"token\":\"secret-copy\"}";
Files.writeString(operator.resolve("auth.json"), auth);
DefaultCodexHome svc = service(root, operator.toString());
Path home = svc.provision(codexCfg(MCP_URL, "MY_BRIDGED_TOKEN"));
assertTrue(Files.exists(home.resolve("auth.json")));
assertEquals(auth, read(home, "auth.json"));
}
@Test
void provisionThrowsNamingPathWhenNoCredentialSource(@TempDir Path root, @TempDir Path operator) {
// operator home exists but holds no auth.json.
DefaultCodexHome svc = service(root, operator.toString());
Path missing = operator.resolve("auth.json");
IllegalStateException ex =
assertThrows(IllegalStateException.class, () -> svc.provision(codexCfg(MCP_URL, "TOK")));
assertTrue(ex.getMessage().contains(missing.toString()), ex.getMessage());
assertTrue(ex.getMessage().contains("401"), ex.getMessage());
}
@Test
void releaseRemovesHomeAndIsIdempotent(@TempDir Path root, @TempDir Path operator) throws IOException {
Files.writeString(operator.resolve("auth.json"), "{\"token\":\"x\"}");
DefaultCodexHome svc = service(root, operator.toString());
Path home = svc.provision(codexCfg(MCP_URL, "MY_BRIDGED_TOKEN"));
assertTrue(Files.exists(home));
svc.release(home);
assertFalse(Files.exists(home));
// second release of the same path must not throw (teardown races teardown).
svc.release(home);
// releasing a never-provisioned path must not throw either.
svc.release(root.resolve("bridged-codex-never-made"));
}
/**
* Isolation: provisioning writes only under the returned home. The operator's codex home (here a
* temp dir, the same code path that production resolves to {@code ~/.codex}) must be byte-for-byte
* untouched afterwards — no leaked {@code config.toml}/AGENTS.md, and its {@code auth.json} only
* ever read, never written. This is a real assertion: a future change that starts writing into
* the operator's home fails here.
*/
@Test
void nothingWrittenToTheOperatorsCodexHome(@TempDir Path root, @TempDir Path operator) throws IOException {
Files.writeString(operator.resolve("auth.json"), "{\"token\":\"x\"}");
Map<String, String> before = snapshot(operator);
DefaultCodexHome svc = service(root, operator.toString());
Path home = svc.provision(codexCfg(MCP_URL, "MY_BRIDGED_TOKEN"));
assertEquals(before, snapshot(operator), "operator codex home must be untouched");
// The peer's artifacts live only inside the returned home, never in the operator's home.
assertFalse(Files.exists(operator.resolve("config.toml")));
assertFalse(Files.exists(operator.resolve("AGENTS.md")));
// And the returned home carries all three artifacts.
assertTrue(Files.exists(home.resolve("auth.json")));
assertTrue(Files.exists(home.resolve("config.toml")));
assertTrue(Files.exists(home.resolve("AGENTS.md")));
}
}
+106 -8
View File
@@ -1,6 +1,6 @@
# CB-308 — Multi-Host Federation (Stage 5)
**Status:** design note (proposal)
**Status:** design note (proposal) — core design decisions resolved 2026-08-10 (§7)
**Depends on:** CB-307 (broker-based reliable delivery) — CB-308 is the multi-host layer built *on*
CB-307's broker fabric.
**Relates to:** CB-401 (`PeerHandle` opaque id), CB-304 (`rosterView`), CB-306 (spawn-readiness),
@@ -143,6 +143,8 @@ extending the same broker from "worker→primary reliability" to "gateway↔gate
5. **Trust** — the broker connection is now the security boundary. A gateway injects env/tokens at
daemon privilege (the CB-401 Stage-C concern), so a **remote-triggered spawn/send** needs
authn/authz: who may act on which host, and which control channels a gateway will honour.
*Authenticity* is resolved — signed messages, §7.1; *authorization* (who may do what) remains
open — §8.
## 5. The one thing the broker does NOT dissolve
@@ -191,15 +193,111 @@ useful), and build CB-308's items on top once the broker fabric exists. Choose C
channel naming **multi-host-ready** now (per-agent routing keys, a `roster.*` topic namespace) so
CB-308 doesn't have to repaint the topology.
## 7. Open questions
## 7. Resolved design decisions (2026-08-10)
Settled in a design review of this note + wiki chapter 10. The broker-level operational rules
(inbox caps, TLS + private broker, schema versioning, trace id, exclusive consumers, U8 broadcast)
are recorded in wiki 10 §10; the CB-308-side decisions are below. Entries 1–6 are the first-pass
decisions; 7–10 came out of the adversarial second-pass review (same day) and supersede 1–6 where
they overlap (notably: the envelope is no longer optional, and dedup is split by path).
1. **Sender authenticity — sign every message.** Each gateway holds its own signing key and signs
what it publishes (sender gid, `msgId`, timestamp). The receiving gateway verifies the
signature **and** checks against the roster that the claimed sender lives on the signing
gateway's host. This extends the single-host invariant — *identity comes from the connection,
never an argument* — across the broker: cross-host, identity comes from the key. Complements
(not replaces) per-gateway broker logins over TLS.
2. **Profiles are owned by the worker's host.** `bridge_spawn(profile, host)` resolves the name in
the *target* gateway's `bridged.yaml`. Gateways advertise their profile names in presence
heartbeats, so a leader sees what each host offers before spawning; an unknown name is a clear
error from the target. Secrets (base URLs, tokens) never leave the host that uses them.
3. **Repo provisioning — clone from the forge, pinned.** A cross-host spawn names the repo URL and
the exact commit. The target gateway clones from the forge into a local cache (first spawn
only), then cuts a per-worker worktree — the CB-301-ext flow with a clone step in front,
covered by the same repo-scoped forge token (CB-302). Git stays the only channel code moves
through.
4. **Asks are live-only, with expiry.** `ASK`/`ANSWER` (U2) traverse the broker as short-lived
(TTL'd) messages carrying the `turn_id`, and are never held durably — the single-host rule
kept. An answer arriving after its turn ended is **not** injected; it is dropped and the leader
gets a `TOO_LATE` notice, so the one failure case is loud rather than weird. Only terminal
replies are durable. Walkthrough: wiki 10 §7.4.
5. **Spawn dedup — a spawn id, remembered on the target.** The control queue redelivers like any
queue; a replayed `SpawnRequest` must not double-spawn. Requests carry a unique spawn id; the
target gateway keeps a short memory of handled ids and answers a redelivery with the existing
`PeerHandle`. CB-117's orphan reap stays as the backstop.
6. **Broker down — local unaffected, remote fails fast.** The routing fork (§3.2) means same-host
traffic never touches the broker; that is now a written promise. A send to a remote agent while
the broker is unreachable **fails immediately** with a clear error — the gateway never buffers
on the broker's behalf (it stays soft-state, so a crash cannot lose messages it claimed to
deliver). Gateways auto-reconnect; remote hosts read as unknown in the roster meanwhile. Broker
HA is a later ops choice, not a design requirement.
7. **Turn state — split by where the signals are.** The *worker's* gateway owns the turn record
(turnId minting, ask coalescing, STALE_TURN, the completion/failure fallbacks, CB-516 abandon):
every input to those decisions — pane status, injection, teardown — is local to it. The
*sender's* gateway owns only the waiter. The two are stitched by terminal-outcome envelope
kinds (`REPLY` / `FAILED` / `ABANDONED`) published to the sender's inbox: a worker dying on B
fails A's waiter fast because gateway B sees the death synchronously and says so.
**`ABANDONED` is belt-and-braces over the waiter's own timeout and roster expiry, never a
replacement** — the case where the waiter hangs longest is gateway B itself dying, which is
exactly when B can publish nothing.
8. **Dual ack model + spawn idempotence by construction.** Forward path (a brief into a worker):
ack **before** the inject — at-most-once, duplicates structurally impossible; the loss window
is closed by an `INJECTED` confirmation published after the inject lands (no `INJECTED` within
a bound = loud fast failure at the sender, not a silent send-timeout). Reply/pull path keeps
ack-after-drain — a duplicate reply is benign, deduped by `msgId`. Spawn: the requester mints
**spawn id = the new worker's gid**; the target checks it against the **live pane registry**,
and the gid is **stored in the herdr pane itself** (label/env, readable back), so a restarted
gateway rebuilds gid↔pane from herdr and the check survives restarts with *no persisted
ledger* — this storage point is the load-bearing detail of the no-ledger position. An
**in-flight reservation set**, entered before the launcher call, absorbs a redelivery arriving
while the first spawn is still inside CB-306's readiness gate; a crash mid-spawn leaves a
half-built pane, which is exactly what CB-117 reaps.
9. **Publish is enforced, not fire-and-forget.** Publisher confirms + the `mandatory` flag + a
return listener, on a **publish channel separate from the consume/ack channel** — synchronous
confirms on the single shared channel would hold its lock across a broker round trip and
serialize acks fleet-wide. Ordering caveat: a *return* (unroutable) arrives **before** the
confirm, so "confirmed" ≠ "routed"; the sender checks the returned-set at confirm time.
`mandatory` is false only for `BROADCAST`, where an empty group is legal silence.
10. **Queue lifecycle is session lifecycle.** `bridge_stop`/reap deletes the worker's inbox queue
(its `broadcast.*` bindings die with it — no broadcasts to the dead); `x-expires` collects
queues orphaned by a crashed gateway (long for main/orchestrator inboxes, short for workers).
Queue names carry a version suffix (`.v2`): AMQP refuses to redeclare an existing durable
queue with new arguments (`PRECONDITION_FAILED` — a crash loop on an in-place upgrade from
v1.0.0), and the suffix keeps old sender-keyed and new recipient-keyed queues apart during
the keying migration (wiki 10 §3 footnote).
## 8. Still open
- **Directory ground-truth:** pure soft-state presence (heartbeats) vs. also treating broker queue
existence as authoritative. Lean soft-state to preserve the persistence boundary; revisit if
split-brain roster views cause mis-routing.
- **Global id scheme:** `<host>/<paneId>` (human-legible, leaks host) vs. opaque UUID (clean, needs
the directory to resolve host). Probably UUID in the protocol, host as directory metadata.
- **Gateway discovery:** how gateways find the broker and each other (static config vs. discovery).
- **Trust model shape:** per-host shared secret vs. mTLS on the broker vs. a capability token per
control action — ties into CB-401 Stage-C.
- **Failure semantics:** a host/gateway dies mid-turn — how the federated roster reaps it (missed
heartbeat) and whether in-flight primary-bound messages survive (broker durability = yes).
- **Control authorization — THE GATE ON U4.** Signing (§7.1) settles *who sent it*; authorization
is *who may do what*. **Cross-host spawn must not land before the minimal version exists**: a
per-host allowlist in `bridged.yaml` — beside the peer public keys — of gateway ids permitted to
publish control to this host, checked against the verified signature. A few lines of config and
check; without them, any principal holding broker credentials can start processes on every host
in the fleet.
- **Key distribution & rotation:** static config (host → public key in each `bridged.yaml`) is
fine at the current 2–3 host scale; rotation is manual. A refinement, not a blocker.
- **Gateway death mid-turn:** the roster reaps it by missed heartbeat, and in-flight primary-bound
messages survive by broker durability; still open is reconciling *worker* state when the dead
gateway's host comes back (orphaned panes vs. still-valid sessions).
*(Resolved and moved up: the global id scheme — an opaque UUID minted by the spawn requester as
the spawn id, host carried as roster metadata; §7.8.)*
## 9. Implementation order (each step verifiable single-host)
1. **As-built fixes, independent of CB-308** (v1.0.x tickets): `basicQos` prefetch on the AMQP
consumer (today the queue drains into gateway heap, so any cap would guard an empty queue);
publisher confirms + `mandatory` (§7.9); the `drainReplies` javadoc that claims "the ack is
local" — false for the AMQP adapter.
2. Envelope + signing (wiki 10 §2.1) — testable against the single-host broker.
3. Recipient-keyed queue migration (`.v2` names, drain-by-`from`, `ReplyPushLoop` rekeyed).
4. Global id + queue lifecycle (§7.8, §7.10).
5. Roster: host-level heartbeat + signed presence; then the routing fork (§3.2).
6. U2 cross-host with the terminal-outcome kinds (§7.4, §7.7).
7. U4 cross-host spawn — **gated on the control allowlist (§8)**.
8. U8 broadcast **last** — it is the feature that punishes an unfinished queue lifecycle.
+18
View File
@@ -0,0 +1,18 @@
{
"name": "claude-bridge",
"description": "Make a project bridge-ready: mount the bridged MCP gateway and set up standard Claude Code settings so this session can orchestrate a fleet of delegated workers. Ships no credentials.",
"version": "0.1.0",
"author": {
"name": "LTMS"
},
"homepage": "https://git.ltms.dev/lms/claude-bridge",
"repository": "https://git.ltms.dev/lms/claude-bridge",
"license": "MIT",
"keywords": [
"mcp",
"orchestration",
"multi-agent",
"delegation",
"codex"
]
}
+8
View File
@@ -0,0 +1,8 @@
{
"mcpServers": {
"bridged": {
"type": "http",
"url": "http://127.0.0.1:8765/mcp"
}
}
}
+54
View File
@@ -0,0 +1,54 @@
# claude-bridge (Claude Code plugin)
Makes a project **bridge-ready**: mounts the `bridged` MCP gateway and applies standard Claude Code
settings, so the session can orchestrate a fleet of delegated workers.
**This plugin ships no credentials.** Every secret is referenced by environment-variable *name*;
the values stay with the user. Nothing the plugin writes is unsafe to commit.
## What it is not
The plugin is the **client-side setup**, not the bridge. `bridged` is a separate daemon and `herdr`
is a separate PTY multiplexer, each with its own lifecycle and install. The plugin mounts an
already-running daemon and tells you what is missing when one isn't there — it deliberately does
not try to install system services on your behalf.
## Install
```shell
/plugin marketplace add ltms/claude-bridge
/plugin install claude-bridge@claude-bridge
```
Then, in the project you want to onboard:
```shell
/claude-bridge:setup
```
## What you get
| Component | Effect |
|---|---|
| `.mcp.json` | mounts `bridged` at `http://127.0.0.1:8765/mcp` for any session with the plugin enabled |
| `skills/setup` | `/claude-bridge:setup` — preflight, project settings, credential guidance, and verification |
Because the plugin carries its own `.mcp.json`, an installed plugin needs no project-level MCP
file at all. The setup skill writes one only when you want the mount to work *without* the plugin —
for teammates who haven't installed it, or for CI.
## Verifying a setup
The setup skill ends by requiring a **real spawn**, not a health check. `/healthz` only reports
that the daemon can reach herdr; a protocol mismatch between the daemon's adapter and the herdr
binary leaves health green while every spawn fails. Only a spawn that reaches `ready` proves the
fleet.
## Local development
```shell
claude --plugin-dir ./plugin
claude plugin validate ./plugin
```
`/reload-plugins` picks up edits without restarting the session.
+227
View File
@@ -0,0 +1,227 @@
---
name: setup
description: Make the current project bridge-ready — check the prerequisites, mount the bridged MCP gateway into the project's .mcp.json, apply standard Claude Code settings, and verify this session resolves as the primary. Writes no credentials. Load this when asked to set up, install, configure, or onboard a project onto claude-bridge, or when bridge_* tools are expected but absent.
---
# Bridge setup — make this project bridge-ready
This skill configures **the project you are currently in** so that this Claude Code session can
orchestrate a fleet of delegated workers through `bridged`.
**It writes no credentials, ever.** Every secret is referenced by environment-variable *name*, and
the user exports the value themselves. Nothing this skill creates is unsafe to commit. If you are
ever about to write a token, key, or password into a file, you have misread this skill — stop.
Work through the steps in order. Each one has a check; **report what actually happened**, including
failures. A setup that half-worked and was reported as done is worse than one that failed loudly.
## 0. Establish where you are
```bash
pwd
git rev-parse --show-toplevel 2>/dev/null || echo "(not a git repo)"
ls -a | head -30
```
Everything below is written into **this** project root. If the user meant a different directory,
confirm before writing anything.
## 1. Preflight — what must already exist
The bridge is three moving parts, and the plugin is only one of them. Check all of it before
changing any file, so you can tell the user the whole story at once instead of failing one step at
a time.
```bash
command -v herdr && herdr --version 2>&1 | head -1 || echo "MISSING: herdr"
command -v ccs && ccs version 2>&1 | head -1 || echo "MISSING: ccs (needed for worker profiles)"
command -v codex && codex --version 2>&1 | head -1 || echo "absent: codex (optional)"
curl -s -m 5 http://127.0.0.1:8765/healthz || echo "MISSING: bridged daemon is not reachable"
```
A healthy daemon answers with its status **and the herdr protocol it negotiated**:
```json
{"status":"ok","herdr":{"version":"0.8.0","protocol":19}}
```
| Missing | What to tell the user |
|---|---|
| `herdr` | The PTY multiplexer that owns worker terminals. Install it first; nothing else works without it. |
| `bridged` | The daemon. It is a separate service, not part of this plugin — the plugin only *mounts* it. Point the user at the project's own install instructions. |
| `ccs` | Only needed to launch worker profiles. The bridge itself will still start. |
| `codex` | Optional. Only needed if this fleet will run Codex peers. |
**Do not attempt to install these yourself.** They are system services with their own lifecycles;
guessing at an install is how you end up with two daemons on one socket. Report what is missing and
let the user install it.
## 2. Mount the bridge MCP — merge, never overwrite
The project's `.mcp.json` may already declare servers. **Read it first and merge**; clobbering
someone's existing MCP config is not a recoverable mistake.
```bash
cat .mcp.json 2>/dev/null || echo "(no .mcp.json yet)"
```
The entry to add, exactly:
```json
{
"mcpServers": {
"bridged": {
"type": "http",
"url": "http://127.0.0.1:8765/mcp"
}
}
}
```
If `.mcp.json` already exists, add only the `bridged` key and leave every other server untouched.
If a `bridged` entry is already there with a different URL, **ask** rather than assuming yours is
right — a non-default port usually means a deliberate second daemon.
> **If this plugin is installed, you can skip this step entirely.** The plugin ships its own
> `.mcp.json`, so `bridged` is already mounted for any session with the plugin enabled. Write the
> project-level file only when the user wants the mount to work *without* the plugin — for
> teammates who have not installed it, or for CI.
**Before writing it, settle whether `.mcp.json` is committed here:**
```bash
git ls-files --error-unmatch .mcp.json 2>/dev/null && echo "TRACKED" || echo "untracked"
```
A tracked `.mcp.json` is inherited by every checkout of this repo — including git worktrees the
bridge provisions for workers. Servers bound to *your* machine (an IDE index, a local language
server) will then be mounted by workers too, and every path they return points into **your**
checkout rather than the worker's. That failure is silent and expensive: it has produced a worker
that made all of its edits in the wrong tree while its builds passed, because it was building the
tree it was not editing. Keep machine-local servers out of a tracked `.mcp.json`, or keep the file
untracked.
## 3. Standard project settings
Create or merge `.claude/settings.json`. These are defaults, not requirements — keep anything the
project already set.
```json
{
"$schema": "https://json.schemastore.org/claude-code-settings.json",
"permissions": {
"allow": [
"mcp__bridged__bridge_whoami",
"mcp__bridged__bridge_list",
"mcp__bridged__bridge_status",
"mcp__bridged__bridge_profiles",
"mcp__bridged__bridge_poll"
]
}
}
```
Only the **read-only** bridge verbs are pre-allowed. `bridge_spawn`, `bridge_send`, and
`bridge_stop` start processes, deliver work, and tear down terminals — those stay behind a prompt
on purpose. Do not "helpfully" add them.
Never write `settings.local.json` on the user's behalf; that file is personal and usually
gitignored.
## 4. Credentials — by reference only
The bridge takes every secret from the **environment**, and the daemon's config names the variable
rather than holding the value. Your job is to tell the user which variables to export, not to
collect or store them.
| Variable | Needed for | Notes |
|---|---|---|
| `BRIDGED_WORKER_TOKEN` | authenticating a worker to the daemon | only when the daemon is configured with `tokenEnv` |
| `GITEA_TOKEN` / equivalent | letting a worker open its own PR | **minimal `write:repository` scope** — see below |
| `GITEA_HOST` | the forge base URL | no secret; safe anywhere |
Two rules to state plainly to the user:
- **The PR token must not be able to merge.** A worker opens a PR; the primary is the gate. A token
that can merge makes the gate decorative. Mint a narrow, repo-scoped token — never reuse a
personal admin token.
- **Never set `ANTHROPIC_BASE_URL` or `ANTHROPIC_AUTH_TOKEN`** in this project, this shell, or any
settings file. The primary stays on subscription; only the daemon moves a *worker* off it, at
spawn. Mounting the bridge must never move a session across that boundary — if setup appears to
need this, something is wrong and you should stop and say so.
Write none of these into any file. Show the user the `export` lines to run themselves.
## 5. Verify — and do not trust a green health check
Reconnect MCP if needed (`/mcp`), then confirm the tools are live and this session is the primary:
```
bridge_whoami
```
- `{"role":"primary"}` — correct, you are done with this step.
- `{"role":"worker", …}` — **this is the trap.** If the primary runs inside a herdr pane, the
daemon resolves it to a terminal and classifies it as a worker, refusing `spawn`/`send`/`stop`:
every verb an orchestrator exists to call. It is **self-locking**, because the daemon can only
*learn* the primary's terminal from those same refused calls. The only way out is an
operator-set pin in the daemon's config:
```yaml
primary:
terminal: term_xxxxxxxxxxxx # the terminalId bridge_whoami just reported
```
The daemon reads this **at boot**, so it needs a restart. Re-pin whenever the primary moves
panes — a stale pin fails exactly as silently as no pin.
Then prove the fleet actually works, with a real spawn:
```
bridge_profiles → the configured backends
bridge_spawn{profile: "<one of them>"} → must reach state "ready"
bridge_stop{paneId: "<from spawn>"}
```
**`/healthz` reporting `ok` is not evidence that spawning works.** It reports that the daemon can
reach herdr — nothing more. A version mismatch between the daemon's adapter and the herdr binary
leaves health green while every single spawn fails. Only a real spawn proves the fleet. Do this
even when everything above looked fine.
## 6. Optional — Codex parity
Only if the user wants Codex and Claude Code to share instructions, skills, and MCP config:
```bash
npm install -g ai-config-sync-manager
ai-config-sync connect
ai-config-sync status # compare both hosts
ai-config-sync sync --dry-run # preview — always look before applying
ai-config-sync sync --apply
```
It maps `~/.claude/CLAUDE.md` ↔ `~/.codex/AGENTS.md`, `~/.claude/skills/` ↔ `~/.codex/skills/`, and
Claude's MCP servers ↔ `[mcp_servers.*]` in `~/.codex/config.toml`.
**Raise the boundary before running it.** That sync is *user-level* and bidirectional, while the
bridge deliberately isolates each worker's tool surface (§2). Syncing your MCP servers into
`~/.codex/config.toml` gives every Codex session your machine-local servers — the same
wrong-tree failure as §2, in a different runtime. Use the sync for the two CLIs *you* drive
interactively; leave anything the bridge spawns isolated. Always `--dry-run` first.
## 7. Report
State plainly:
```
prereqs: herdr <version> · bridged <protocol> · ccs <version> · codex <version|absent>
written: <files created or merged, or "none">
role: <bridge_whoami result — and the pin, if one was needed>
spawn: <real spawn result: profile, state reached, torn down>
env: <variables the USER still needs to export — names only, never values>
skipped: <anything not done, and why>
```
Never report a step as done that you did not verify. If the daemon was unreachable, say so and stop
— the remaining steps cannot be checked, and guessing at them is how a broken setup gets called
finished.
+1 -1
Submodule wiki updated: 0c896eb49b...4320c1ca52