Closes out single-host before the cross-host work. Sequenced BEFORE CB-308 deliberately: federation's own gating concern is the trust model, and it inherits whatever identity shape lands here. The finding this stage is built around: bridged had exactly ONE security control, the loopback bind. ConnectionIdentity resolves a worker from its connection (unforgeable), but every caller that was not a recognised worker pane fell through to being treated as the PRIMARY -- the most privileged role on the bus. Latent today; load-bearing the moment a bind widens. CB-501 auth: - Role/Principal/CallerResolver: connection identity first, bearer token second, ANONYMOUS third. Inverts the old default so absence of identity means nothing, not everything. - Worker identity is never token-gated, so enabling auth cannot lock the fleet out of bridge_reply. - Constant-time token compare (MessageDigest.isEqual). - validateAuthExposure(): a non-loopback bind under loopback-trust now REFUSES TO START. Makes the dangerous config unrepresentable rather than merely documented. - TLS terminates at a reverse proxy by design (D3), not in the JVM. CB-505 authz + audit, enforced on BOTH entry paths: - The docs describe MCP as "a thin adapter over the REST core"; at code level it is not. BridgeMcp calls MessageService directly, and /mcp is a raw servlet on Jetty's context handler that never traverses Javalin's before filter. Enforcing only at REST would have left /mcp open. - Load-bearing rule is own-session-only: a worker may reply/ask only as itself. Structurally true over MCP already; over REST the session id in the URL path had simply been trusted. - Audit: JSON lines to a dedicated appender, additivity=false. Never records message content -- this bus carries source and prompts. CB-502 metrics: zero new dependencies. A ~150-line Prometheus text renderer instead of the specced Micrometer, because this pom already hand-pins jackson-annotations to reconcile Jackson 2/3, imports a Jetty BOM against skew, and carries four accepted-CVE advisories -- and CLAUDE.md's mandated dependency CVE gate could not be run (no JetBrains MCP server connected). Instrumented at MessageService, the single funnel both surfaces share. CB-503 CI: .gitea/workflows/ci.yml against the already-running Gitea runner. Needs no contract-exclusion flag -- the pom's default-excludes profile already sets excludedGroups=contract, so plain `mvn clean install` IS the mock-socket surface. Provisions JDK 25 explicitly (runner default-jdk is older). CB-504 supervision: launchd agent (the real target -- this host is macOS, there is no systemd) plus a systemd unit for the Linux gateways CB-308 adds. Ordering directives are advisory, so the actual fix is that startup now waits up to 30s for the herdr socket and then serves degraded, instead of crashing into a restart loop on a boot-order race. Also fixes drift found while surveying: - bridged.example.yaml documented spawn_ready_timeout_ms in snake_case; config binds via plain Jackson with ignoreUnknown, so uncommenting it would have been silently dropped and the default kept. Now camelCase, with a test that loads the shipped example and one that pins every documented knob's spelling -- no test had ever loaded that file. - Added the 6 shipped-but-undocumented knobs (worktreeRoot, parityOverlay, gitTokenEnv, gitHostEnv, configDir, primary:). - README "Next" listed bridge_ask and session lifecycle as upcoming; both shipped long ago. - docs/CB-301-ext and docs/CB-402 status headers said "design"/"pre- implementation" for work already merged. 307 unit/acceptance tests green (was 266), mvn clean install BUILD SUCCESS. Note: CLAUDE.md's per-file ide_diagnostics gate and the pom Mend.io CVE check could not be run -- no JetBrains/intellij-index MCP server is connected this session. mvn clean install is the only gate that ran.
14 KiB
CB-5xx — Stage 5 Hardening (auth · metrics · CI · supervision · authz+audit)
Status: design note (pre-implementation) — the single-host close-out before cross-host work.
Covers: CB-501 (bearer auth + TLS) · CB-502 (/metrics) · CB-503 (mock-socket CI) ·
CB-504 (service supervision) · CB-505 (per-session authz + audit log).
Depends on: everything shipped through CB-402. Nothing here changes messaging semantics.
Blocks: CB-308. The cross-host trust model is CB-308's own gating concern, and it inherits
whatever identity/authz shape lands here — so this stage is deliberately before federation,
not after it.
1. Why this stage is not optional bookkeeping
bridged today has exactly one security control: the loopback bind. Every other guarantee
rests on it.
The identity model (mcp/ConnectionIdentity.java) resolves a caller from the connection alone —
the OS reports the connecting PID, herdr owns the PID→pane map, so a worker cannot forge another
worker. Its own javadoc is explicit: "Single-host only (the herd shares the bridged host); the
token path is the split-host fallback." The token path does not exist yet.
That leaves a seam that is latent today and load-bearing the moment the bind moves:
// ConnectionIdentity.resolve — non-loopback callers get a null terminal
if (!isLoopback(remoteAddr)) return new Caller(null, -1);
…and null terminal is interpreted downstream as "this caller is the primary". Combined:
Any caller that is not a recognised on-host worker pane is treated as the primary — including, if
bind.hostis ever widened, an arbitrary remote client.
Today bind defaults to 127.0.0.1 so this is unreachable. But CB-308 exists precisely to widen
the boundary, and the primary is the most privileged role on the bus (it spawns, stops, sends to
any session, and drains any inbox). Shipping federation on top of "unauthenticated ⇒ primary"
would be building the security boundary backwards.
So CB-501 is not "add a token header". It is: make identity explicit, and make the absence of identity mean nothing, not everything.
2. Decisions (locked)
D1 — Three caller roles, one resolution path
Introduce Role { PRIMARY, WORKER, ANONYMOUS } resolved by a single CallerResolver that both
REST and MCP go through. Resolution order:
- Connection identity wins where it applies. A loopback peer PID that maps to a herdr worker
pane ⇒
WORKERwith that terminal. Unforgeable, unchanged from today, zero config. - Token, if presented. A valid bearer token ⇒ the role that token is provisioned for.
- Otherwise
ANONYMOUS— notPRIMARY.
This inverts today's default. PRIMARY becomes something you must prove (by being a loopback
non-worker process when auth is disabled, or by presenting a primary-scoped token when it is
enabled), rather than something you get by failing every other check.
D2 — Auth is opt-in by config, but the default must stay zero-friction on loopback
The daemon is dogfooded constantly on one machine. If enabling hardening breaks the local setup, it will be disabled and the stage is wasted. So:
auth:
mode: loopback-trust # default — behaves exactly like today: loopback ⇒ PRIMARY, no token needed
# mode: token # every non-worker caller must present a valid bearer token
# tokenEnv: BRIDGED_API_TOKEN # host env var holding the token; never the literal value
mode: loopback-trust is the current behaviour, named honestly and now chosen rather than
implied. mode: token is what a non-loopback bind requires. A non-loopback bind.host with
mode: loopback-trust must fail fast at startup — that check is the single highest-value line
in this stage, because it makes the dangerous configuration unrepresentable rather than merely
discouraged.
D3 — TLS terminates outside the JVM
Do not add TLS config to Javalin/Jetty. The deployment story for a cross-host gateway is a
reverse proxy (or an SSH/WireGuard tunnel) in front of the daemon; the AMQP link has its own TLS
via the broker URI (amqps://). Adding keystore handling here would mean certificate lifecycle
code in a daemon whose whole value is being small, and would duplicate what the proxy does better.
CB-501 therefore ships bearer auth + the fail-fast bind check, and documents TLS as a deployment concern with a worked reverse-proxy example. This is a deliberate narrowing of the roadmap's "auth/TLS" wording — flagged in §6 for the lead.
D4 — Metrics without a new dependency
The roadmap's tech-stack table says Micrometer→Prometheus. Recommend not taking that dep:
- This pom already carries an unusually heavy dependency-reconciliation burden (a hand-pinned
jackson-annotations3.0-rc5 to reconcile the MCP SDK's Jackson 3 with our Jackson 2.19, a Jetty BOM import to stop version skew, plus four documented accepted-CVE advisories). Every new transitive tree is a real cost here, not a hypothetical one. - The CVE gate that CLAUDE.md mandates for dependency changes (
jetbrains get_file_problems→ Mend.io) cannot currently be run — no JetBrains MCP server is connected. Adding a dependency tree we cannot scan violates the project's own stated policy. - The metric set is small and fully known (§4). Prometheus text exposition is a trivial, stable, well-specified format.
So: a ~120-line metrics/Metrics.java holding LongAdder counters and gauge suppliers, rendered
to the Prometheus text format at GET /metrics. If Micrometer is wanted later for its
registry/push ecosystem, this stays a drop-in swap behind the same endpoint. Flagged in §6 —
this deviates from a documented tech-stack choice.
D5 — Supervision targets launchd first, systemd second
The roadmap says "systemd unit". This host is macOS — there is no systemd on it (systemctl
not found), and the daemon that has been dogfooded for weeks runs as a bare foreground
java -jar. Ship both:
deploy/dev.ltms.bridged.plist— launchd agent, the actual runtime here, withKeepAliveand ordered start after herdr.deploy/bridged.service— systemd unit for the Linux gateways CB-308 introduces.
Ordering after herdr is advisory in both: the herdr socket may not exist at boot, so the daemon must retry the socket rather than exit — supervision ordering is a nicety, socket-retry is the actual fix. That retry behaviour is part of CB-504, not a separate ticket.
D6 — Audit log is a separate append-only stream, not the app log
Privileged actions (spawn, stop, send, reply-drain, ack) emit a structured JSON line to a
dedicated audit SLF4J logger with its own appender, carrying {ts, role, terminal, pid, action, target, outcome}. Keeping it off the chatty app logger is what makes it greppable and, later,
shippable. No message content in the audit record — the bridge carries the user's source
code and prompts; an audit trail that quietly becomes a transcript archive is a liability, not a
control. Content stays out; correlation ids go in.
3. Authorization model (CB-505)
With D1's roles, the rules are small enough to state completely:
| Action | REST | PRIMARY | WORKER | ANONYMOUS |
|---|---|---|---|---|
| spawn worker | POST /workers |
✅ | ❌ | ❌ |
| stop worker | DELETE /workers/{paneId} |
✅ | ❌ | ❌ |
| send to a session | POST /sessions/{id}/message |
✅ | ❌ | ❌ |
| reply | POST /sessions/{id}/reply |
❌ | ✅ own session only | ❌ |
| ask | POST /sessions/{id}/ask |
❌ | ✅ own session only | ❌ |
| drain replies | GET /sessions/{id}/replies |
✅ | ❌ | ❌ |
| status / list / profiles | GET … |
✅ | ✅ | ❌ |
| health | GET /healthz |
✅ | ✅ | ✅ (unauthenticated by design) |
| metrics | GET /metrics |
✅ | ✅ | ❌ |
The load-bearing row is "own session only": a worker may only reply or ask as itself. That is
already true de facto — ConnectionIdentity derives the terminal rather than reading it from the
body — so CB-505 mostly asserts an existing invariant explicitly and adds the test that pins
it. The one real change is rejecting a worker that names a different session id in the path.
/healthz stays open: it must answer for a load balancer or supervisor before any credential is
configured. It already leaks nothing but herdr's version and up/down.
3.1 There are TWO entry paths, and only one of them has identity today
The wiki describes MCP as "a thin adapter over the REST core". At the code level that is not
literally true, and the difference is security-relevant. BridgeMcp calls MessageService /
SessionManager directly; it never issues an HTTP request against a Javalin route. And /mcp is
mounted as a raw servlet on Jetty's ServletContextHandler
(BridgedApp.build → cfg.jetty.modifyServletContextHandler), so it does not pass through
Javalin's before filters at all.
The current split is the mirror image of what you'd expect:
| Path | Caller identity today | Authz today |
|---|---|---|
MCP /mcp |
✅ resolved per call (ConnectionIdentity via the transport-context extractor) |
❌ none |
| REST routes | ❌ none at all — the session id is taken from the URL path and trusted | ❌ none |
So REST is the more exposed surface: POST /sessions/{id}/reply accepts any {id} from the
path, whereas the MCP bridge_reply derives the worker from the connection and refuses to read it
from an argument. Loopback-only bind is what makes this safe today.
Therefore CB-505 must enforce on both paths against one shared resolver — not at a single
choke point. Concretely: a Javalin before filter for REST, and the existing transport-context
extractor for MCP, both delegating to auth.CallerResolver. Any authz check that lives in only
one of the two is not a control.
4. Metric set (CB-502)
Deliberately small; every one maps to a failure mode we have actually hit.
| Metric | Type | Why it exists |
|---|---|---|
bridged_sends_total{outcome} |
counter | outcome ∈ replied|completion_fallback|timeout|failed — the completion-fallback rate is the health signal for turn detection (CB-115/116/118) |
bridged_send_duration_seconds |
histogram | delegated turn latency |
bridged_replies_total{path} |
counter | path ∈ rendezvous|inbox — how often a reply strands (CB-307's whole reason to exist) |
bridged_inbox_depth{target} |
gauge | undrained replies; steady-state should be 0 |
bridged_push_nudges_total{outcome} |
counter | outcome ∈ delivered|exhausted — a rising exhausted means the primary is not draining |
bridged_spawns_total{kind,outcome} |
counter | outcome ∈ ready|timeout|guard_rejected; per peer kind (CB-402) |
bridged_sessions{state} |
gauge | SPAWNING/READY/BUSY/DONE census |
bridged_herdr_calls_total{method,outcome} |
counter | socket health — the dependency everything rests on |
bridged_auth_failures_total{reason} |
counter | only meaningful once CB-501 lands; catches misconfigured workers |
5. Increment plan
Ordered so each step is independently mergeable and the risky one lands first.
- CB-501a —
CallerResolver+Role. Pure refactor: route today's connection identity through the new type,ANONYMOUSnot yet reachable (loopback-trust default preserves behaviour). Green build, no behaviour change. - CB-501b — token mode + fail-fast bind check. Config block, bearer parsing, the non-loopback-bind guard. This is the security-relevant commit; keep it small and reviewable.
- CB-505 — authz table + audit logger. Enforce §3 on both entry paths (see §3.1); add the audit appender.
- CB-502 —
Metrics+/metrics. Instrument the paths in §4. - CB-503 — CI. Runs
mvn -B clean installwith-Dgroups='!contract'so the live-herdr and RabbitMQ contract tests are excluded; the mock-UDS suite is the CI surface, exactly as the roadmap's testability section intends. - CB-504 — launchd plist + systemd unit + herdr-socket retry.
6. Open questions for the lead
All four resolved 2026-07-29 — the lead confirmed D3, D4, and the D6 sink; the CI runner question was answered from the forge itself. Kept here as the decision record.
- ✅ TLS scope (D3) — confirmed. Bearer auth + the fail-fast bind guard ship in the daemon;
TLS terminates at a reverse proxy, documented with a worked example. AMQP gets TLS via an
amqps://URI. No keystore handling inbridged. - ✅ Micrometer (D4) — confirmed dropped. Zero-dependency Prometheus text renderer, for the reasons in D4 (pom reconciliation burden + the mandated CVE gate being un-runnable this session). Revisit if a push-gateway or JVM-metrics requirement appears; the endpoint is the swap seam.
CI runner (CB-503).✅ Resolved during design — a Gitea Actions runner is registered and healthy (lms/alms-memoryhas 28 completed runs;lms/almsruns on push and pull_request). CB-503 targets.gitea/workflows/ci.ymlwithruns-on: ubuntu-latest, matching the sibling repo's convention. Note the runner's image ships an olderdefault-jdk, so the workflow must provision JDK 25 explicitly rather than apt-installing the default. Contract-test exclusion needs no CI flag: the pom'sdefault-excludesprofile already setsexcludedGroups=contract, so a plainmvn -B clean installis the mock-socket surface.- ✅ Audit sink — confirmed dedicated file. Its own logback appender writing JSON lines beside the daemon, separate from the app log, per D6. Content still never enters the record.