Files
fleetd/docs/CB-5xx-Hardening.md
kevin 9daf1ec5ba CB-5xx: Stage 5 hardening — auth, authz+audit, metrics, CI, supervision
Closes out single-host before the cross-host work. Sequenced BEFORE CB-308
deliberately: federation's own gating concern is the trust model, and it
inherits whatever identity shape lands here.

The finding this stage is built around: bridged had exactly ONE security
control, the loopback bind. ConnectionIdentity resolves a worker from its
connection (unforgeable), but every caller that was not a recognised worker
pane fell through to being treated as the PRIMARY -- the most privileged role
on the bus. Latent today; load-bearing the moment a bind widens.

CB-501 auth:
- Role/Principal/CallerResolver: connection identity first, bearer token
  second, ANONYMOUS third. Inverts the old default so absence of identity
  means nothing, not everything.
- Worker identity is never token-gated, so enabling auth cannot lock the
  fleet out of bridge_reply.
- Constant-time token compare (MessageDigest.isEqual).
- validateAuthExposure(): a non-loopback bind under loopback-trust now
  REFUSES TO START. Makes the dangerous config unrepresentable rather than
  merely documented.
- TLS terminates at a reverse proxy by design (D3), not in the JVM.

CB-505 authz + audit, enforced on BOTH entry paths:
- The docs describe MCP as "a thin adapter over the REST core"; at code level
  it is not. BridgeMcp calls MessageService directly, and /mcp is a raw
  servlet on Jetty's context handler that never traverses Javalin's before
  filter. Enforcing only at REST would have left /mcp open.
- Load-bearing rule is own-session-only: a worker may reply/ask only as
  itself. Structurally true over MCP already; over REST the session id in the
  URL path had simply been trusted.
- Audit: JSON lines to a dedicated appender, additivity=false. Never records
  message content -- this bus carries source and prompts.

CB-502 metrics: zero new dependencies. A ~150-line Prometheus text renderer
instead of the specced Micrometer, because this pom already hand-pins
jackson-annotations to reconcile Jackson 2/3, imports a Jetty BOM against
skew, and carries four accepted-CVE advisories -- and CLAUDE.md's mandated
dependency CVE gate could not be run (no JetBrains MCP server connected).
Instrumented at MessageService, the single funnel both surfaces share.

CB-503 CI: .gitea/workflows/ci.yml against the already-running Gitea runner.
Needs no contract-exclusion flag -- the pom's default-excludes profile
already sets excludedGroups=contract, so plain `mvn clean install` IS the
mock-socket surface. Provisions JDK 25 explicitly (runner default-jdk is older).

CB-504 supervision: launchd agent (the real target -- this host is macOS,
there is no systemd) plus a systemd unit for the Linux gateways CB-308 adds.
Ordering directives are advisory, so the actual fix is that startup now waits
up to 30s for the herdr socket and then serves degraded, instead of crashing
into a restart loop on a boot-order race.

Also fixes drift found while surveying:
- bridged.example.yaml documented spawn_ready_timeout_ms in snake_case; config
  binds via plain Jackson with ignoreUnknown, so uncommenting it would have
  been silently dropped and the default kept. Now camelCase, with a test that
  loads the shipped example and one that pins every documented knob's
  spelling -- no test had ever loaded that file.
- Added the 6 shipped-but-undocumented knobs (worktreeRoot, parityOverlay,
  gitTokenEnv, gitHostEnv, configDir, primary:).
- README "Next" listed bridge_ask and session lifecycle as upcoming; both
  shipped long ago.
- docs/CB-301-ext and docs/CB-402 status headers said "design"/"pre-
  implementation" for work already merged.

307 unit/acceptance tests green (was 266), mvn clean install BUILD SUCCESS.
Note: CLAUDE.md's per-file ide_diagnostics gate and the pom Mend.io CVE check
could not be run -- no JetBrains/intellij-index MCP server is connected this
session. mvn clean install is the only gate that ran.
2026-07-29 22:29:26 +07:00

14 KiB

CB-5xx — Stage 5 Hardening (auth · metrics · CI · supervision · authz+audit)

Status: design note (pre-implementation) — the single-host close-out before cross-host work. Covers: CB-501 (bearer auth + TLS) · CB-502 (/metrics) · CB-503 (mock-socket CI) · CB-504 (service supervision) · CB-505 (per-session authz + audit log). Depends on: everything shipped through CB-402. Nothing here changes messaging semantics. Blocks: CB-308. The cross-host trust model is CB-308's own gating concern, and it inherits whatever identity/authz shape lands here — so this stage is deliberately before federation, not after it.


1. Why this stage is not optional bookkeeping

bridged today has exactly one security control: the loopback bind. Every other guarantee rests on it.

The identity model (mcp/ConnectionIdentity.java) resolves a caller from the connection alone — the OS reports the connecting PID, herdr owns the PID→pane map, so a worker cannot forge another worker. Its own javadoc is explicit: "Single-host only (the herd shares the bridged host); the token path is the split-host fallback." The token path does not exist yet.

That leaves a seam that is latent today and load-bearing the moment the bind moves:

// ConnectionIdentity.resolve — non-loopback callers get a null terminal
if (!isLoopback(remoteAddr)) return new Caller(null, -1);

…and null terminal is interpreted downstream as "this caller is the primary". Combined:

Any caller that is not a recognised on-host worker pane is treated as the primary — including, if bind.host is ever widened, an arbitrary remote client.

Today bind defaults to 127.0.0.1 so this is unreachable. But CB-308 exists precisely to widen the boundary, and the primary is the most privileged role on the bus (it spawns, stops, sends to any session, and drains any inbox). Shipping federation on top of "unauthenticated ⇒ primary" would be building the security boundary backwards.

So CB-501 is not "add a token header". It is: make identity explicit, and make the absence of identity mean nothing, not everything.


2. Decisions (locked)

D1 — Three caller roles, one resolution path

Introduce Role { PRIMARY, WORKER, ANONYMOUS } resolved by a single CallerResolver that both REST and MCP go through. Resolution order:

  1. Connection identity wins where it applies. A loopback peer PID that maps to a herdr worker pane ⇒ WORKER with that terminal. Unforgeable, unchanged from today, zero config.
  2. Token, if presented. A valid bearer token ⇒ the role that token is provisioned for.
  3. Otherwise ANONYMOUS — not PRIMARY.

This inverts today's default. PRIMARY becomes something you must prove (by being a loopback non-worker process when auth is disabled, or by presenting a primary-scoped token when it is enabled), rather than something you get by failing every other check.

D2 — Auth is opt-in by config, but the default must stay zero-friction on loopback

The daemon is dogfooded constantly on one machine. If enabling hardening breaks the local setup, it will be disabled and the stage is wasted. So:

auth:
  mode: loopback-trust   # default — behaves exactly like today: loopback ⇒ PRIMARY, no token needed
  # mode: token          # every non-worker caller must present a valid bearer token
  # tokenEnv: BRIDGED_API_TOKEN   # host env var holding the token; never the literal value

mode: loopback-trust is the current behaviour, named honestly and now chosen rather than implied. mode: token is what a non-loopback bind requires. A non-loopback bind.host with mode: loopback-trust must fail fast at startup — that check is the single highest-value line in this stage, because it makes the dangerous configuration unrepresentable rather than merely discouraged.

D3 — TLS terminates outside the JVM

Do not add TLS config to Javalin/Jetty. The deployment story for a cross-host gateway is a reverse proxy (or an SSH/WireGuard tunnel) in front of the daemon; the AMQP link has its own TLS via the broker URI (amqps://). Adding keystore handling here would mean certificate lifecycle code in a daemon whose whole value is being small, and would duplicate what the proxy does better.

CB-501 therefore ships bearer auth + the fail-fast bind check, and documents TLS as a deployment concern with a worked reverse-proxy example. This is a deliberate narrowing of the roadmap's "auth/TLS" wording — flagged in §6 for the lead.

D4 — Metrics without a new dependency

The roadmap's tech-stack table says Micrometer→Prometheus. Recommend not taking that dep:

  • This pom already carries an unusually heavy dependency-reconciliation burden (a hand-pinned jackson-annotations 3.0-rc5 to reconcile the MCP SDK's Jackson 3 with our Jackson 2.19, a Jetty BOM import to stop version skew, plus four documented accepted-CVE advisories). Every new transitive tree is a real cost here, not a hypothetical one.
  • The CVE gate that CLAUDE.md mandates for dependency changes (jetbrains get_file_problems → Mend.io) cannot currently be run — no JetBrains MCP server is connected. Adding a dependency tree we cannot scan violates the project's own stated policy.
  • The metric set is small and fully known (§4). Prometheus text exposition is a trivial, stable, well-specified format.

So: a ~120-line metrics/Metrics.java holding LongAdder counters and gauge suppliers, rendered to the Prometheus text format at GET /metrics. If Micrometer is wanted later for its registry/push ecosystem, this stays a drop-in swap behind the same endpoint. Flagged in §6 — this deviates from a documented tech-stack choice.

D5 — Supervision targets launchd first, systemd second

The roadmap says "systemd unit". This host is macOS — there is no systemd on it (systemctl not found), and the daemon that has been dogfooded for weeks runs as a bare foreground java -jar. Ship both:

  • deploy/dev.ltms.bridged.plist — launchd agent, the actual runtime here, with KeepAlive and ordered start after herdr.
  • deploy/bridged.service — systemd unit for the Linux gateways CB-308 introduces.

Ordering after herdr is advisory in both: the herdr socket may not exist at boot, so the daemon must retry the socket rather than exit — supervision ordering is a nicety, socket-retry is the actual fix. That retry behaviour is part of CB-504, not a separate ticket.

D6 — Audit log is a separate append-only stream, not the app log

Privileged actions (spawn, stop, send, reply-drain, ack) emit a structured JSON line to a dedicated audit SLF4J logger with its own appender, carrying {ts, role, terminal, pid, action, target, outcome}. Keeping it off the chatty app logger is what makes it greppable and, later, shippable. No message content in the audit record — the bridge carries the user's source code and prompts; an audit trail that quietly becomes a transcript archive is a liability, not a control. Content stays out; correlation ids go in.


3. Authorization model (CB-505)

With D1's roles, the rules are small enough to state completely:

Action REST PRIMARY WORKER ANONYMOUS
spawn worker POST /workers ✅ ❌ ❌
stop worker DELETE /workers/{paneId} ✅ ❌ ❌
send to a session POST /sessions/{id}/message ✅ ❌ ❌
reply POST /sessions/{id}/reply ❌ ✅ own session only ❌
ask POST /sessions/{id}/ask ❌ ✅ own session only ❌
drain replies GET /sessions/{id}/replies ✅ ❌ ❌
status / list / profiles GET … ✅ ✅ ❌
health GET /healthz ✅ ✅ ✅ (unauthenticated by design)
metrics GET /metrics ✅ ✅ ❌

The load-bearing row is "own session only": a worker may only reply or ask as itself. That is already true de facto — ConnectionIdentity derives the terminal rather than reading it from the body — so CB-505 mostly asserts an existing invariant explicitly and adds the test that pins it. The one real change is rejecting a worker that names a different session id in the path.

/healthz stays open: it must answer for a load balancer or supervisor before any credential is configured. It already leaks nothing but herdr's version and up/down.

3.1 There are TWO entry paths, and only one of them has identity today

The wiki describes MCP as "a thin adapter over the REST core". At the code level that is not literally true, and the difference is security-relevant. BridgeMcp calls MessageService / SessionManager directly; it never issues an HTTP request against a Javalin route. And /mcp is mounted as a raw servlet on Jetty's ServletContextHandler (BridgedApp.build → cfg.jetty.modifyServletContextHandler), so it does not pass through Javalin's before filters at all.

The current split is the mirror image of what you'd expect:

Path Caller identity today Authz today
MCP /mcp ✅ resolved per call (ConnectionIdentity via the transport-context extractor) ❌ none
REST routes ❌ none at all — the session id is taken from the URL path and trusted ❌ none

So REST is the more exposed surface: POST /sessions/{id}/reply accepts any {id} from the path, whereas the MCP bridge_reply derives the worker from the connection and refuses to read it from an argument. Loopback-only bind is what makes this safe today.

Therefore CB-505 must enforce on both paths against one shared resolver — not at a single choke point. Concretely: a Javalin before filter for REST, and the existing transport-context extractor for MCP, both delegating to auth.CallerResolver. Any authz check that lives in only one of the two is not a control.


4. Metric set (CB-502)

Deliberately small; every one maps to a failure mode we have actually hit.

Metric Type Why it exists
bridged_sends_total{outcome} counter outcome ∈ replied|completion_fallback|timeout|failed — the completion-fallback rate is the health signal for turn detection (CB-115/116/118)
bridged_send_duration_seconds histogram delegated turn latency
bridged_replies_total{path} counter path ∈ rendezvous|inbox — how often a reply strands (CB-307's whole reason to exist)
bridged_inbox_depth{target} gauge undrained replies; steady-state should be 0
bridged_push_nudges_total{outcome} counter outcome ∈ delivered|exhausted — a rising exhausted means the primary is not draining
bridged_spawns_total{kind,outcome} counter outcome ∈ ready|timeout|guard_rejected; per peer kind (CB-402)
bridged_sessions{state} gauge SPAWNING/READY/BUSY/DONE census
bridged_herdr_calls_total{method,outcome} counter socket health — the dependency everything rests on
bridged_auth_failures_total{reason} counter only meaningful once CB-501 lands; catches misconfigured workers

5. Increment plan

Ordered so each step is independently mergeable and the risky one lands first.

  1. CB-501a — CallerResolver + Role. Pure refactor: route today's connection identity through the new type, ANONYMOUS not yet reachable (loopback-trust default preserves behaviour). Green build, no behaviour change.
  2. CB-501b — token mode + fail-fast bind check. Config block, bearer parsing, the non-loopback-bind guard. This is the security-relevant commit; keep it small and reviewable.
  3. CB-505 — authz table + audit logger. Enforce §3 on both entry paths (see §3.1); add the audit appender.
  4. CB-502 — Metrics + /metrics. Instrument the paths in §4.
  5. CB-503 — CI. Runs mvn -B clean install with -Dgroups='!contract' so the live-herdr and RabbitMQ contract tests are excluded; the mock-UDS suite is the CI surface, exactly as the roadmap's testability section intends.
  6. CB-504 — launchd plist + systemd unit + herdr-socket retry.

6. Open questions for the lead

All four resolved 2026-07-29 — the lead confirmed D3, D4, and the D6 sink; the CI runner question was answered from the forge itself. Kept here as the decision record.

  1. ✅ TLS scope (D3) — confirmed. Bearer auth + the fail-fast bind guard ship in the daemon; TLS terminates at a reverse proxy, documented with a worked example. AMQP gets TLS via an amqps:// URI. No keystore handling in bridged.
  2. ✅ Micrometer (D4) — confirmed dropped. Zero-dependency Prometheus text renderer, for the reasons in D4 (pom reconciliation burden + the mandated CVE gate being un-runnable this session). Revisit if a push-gateway or JVM-metrics requirement appears; the endpoint is the swap seam.
  3. CI runner (CB-503). ✅ Resolved during design — a Gitea Actions runner is registered and healthy (lms/alms-memory has 28 completed runs; lms/alms runs on push and pull_request). CB-503 targets .gitea/workflows/ci.yml with runs-on: ubuntu-latest, matching the sibling repo's convention. Note the runner's image ships an older default-jdk, so the workflow must provision JDK 25 explicitly rather than apt-installing the default. Contract-test exclusion needs no CI flag: the pom's default-excludes profile already sets excludedGroups=contract, so a plain mvn -B clean install is the mock-socket surface.
  4. ✅ Audit sink — confirmed dedicated file. Its own logback appender writing JSON lines beside the daemon, separate from the app log, per D6. Content still never enters the record.