# CB-5xx — Stage 5 Hardening (auth · metrics · CI · supervision · authz+audit) **Status:** design note (pre-implementation) — the single-host close-out before cross-host work. **Covers:** CB-501 (bearer auth + TLS) · CB-502 (`/metrics`) · CB-503 (mock-socket CI) · CB-504 (service supervision) · CB-505 (per-session authz + audit log). **Depends on:** everything shipped through CB-402. Nothing here changes messaging semantics. **Blocks:** CB-308. The cross-host trust model is CB-308's own gating concern, and it inherits whatever identity/authz shape lands here — so this stage is deliberately *before* federation, not after it. --- ## 1. Why this stage is not optional bookkeeping `fleetd` today has **exactly one security control: the loopback bind**. Every other guarantee rests on it. The identity model (`mcp/ConnectionIdentity.java`) resolves a caller from the connection alone — the OS reports the connecting PID, herdr owns the PID→pane map, so a worker cannot forge another worker. Its own javadoc is explicit: *"Single-host only (the herd shares the `fleetd` host); the token path is the split-host fallback."* The token path does not exist yet. That leaves a seam that is **latent today and load-bearing the moment the bind moves**: ```java // ConnectionIdentity.resolve — non-loopback callers get a null terminal if (!isLoopback(remoteAddr)) return new Caller(null, -1); ``` …and `null` terminal is interpreted downstream as **"this caller is the primary"**. Combined: > Any caller that is not a recognised on-host worker pane is treated as the primary — including, > if `bind.host` is ever widened, an arbitrary remote client. Today `bind` defaults to `127.0.0.1` so this is unreachable. But CB-308 exists precisely to widen the boundary, and the primary is the *most* privileged role on the bus (it spawns, stops, sends to any session, and drains any inbox). Shipping federation on top of "unauthenticated ⇒ primary" would be building the security boundary backwards. **So CB-501 is not "add a token header". It is: make identity explicit, and make the absence of identity mean *nothing*, not *everything*.** --- ## 2. Decisions (locked) ### D1 — Three caller roles, one resolution path Introduce `Role { PRIMARY, WORKER, ANONYMOUS }` resolved by a single `CallerResolver` that both REST and MCP go through. Resolution order: 1. **Connection identity wins where it applies.** A loopback peer PID that maps to a herdr worker pane ⇒ `WORKER` with that terminal. Unforgeable, unchanged from today, zero config. 2. **Token, if presented.** A valid bearer token ⇒ the role that token is provisioned for. 3. **Otherwise `ANONYMOUS`** — *not* `PRIMARY`. This inverts today's default. `PRIMARY` becomes something you must *prove* (by being a loopback non-worker process when auth is disabled, or by presenting a primary-scoped token when it is enabled), rather than something you get by failing every other check. ### D2 — Auth is opt-in by config, but the *default* must stay zero-friction on loopback The daemon is dogfooded constantly on one machine. If enabling hardening breaks the local setup, it will be disabled and the stage is wasted. So: ```yaml auth: mode: loopback-trust # default — behaves exactly like today: loopback ⇒ PRIMARY, no token needed # mode: token # every non-worker caller must present a valid bearer token # tokenEnv: FLEETD_API_TOKEN # host env var holding the token; never the literal value ``` `mode: loopback-trust` is the current behaviour, named honestly and now *chosen* rather than implied. `mode: token` is what a non-loopback bind requires. **A non-loopback `bind.host` with `mode: loopback-trust` must fail fast at startup** — that check is the single highest-value line in this stage, because it makes the dangerous configuration unrepresentable rather than merely discouraged. ### D3 — TLS terminates *outside* the JVM Do **not** add TLS config to Javalin/Jetty. The deployment story for a cross-host gateway is a reverse proxy (or an SSH/WireGuard tunnel) in front of the daemon; the AMQP link has its own TLS via the broker URI (`amqps://`). Adding keystore handling here would mean certificate lifecycle code in a daemon whose whole value is being small, and would duplicate what the proxy does better. **CB-501 therefore ships bearer auth + the fail-fast bind check, and documents TLS as a deployment concern with a worked reverse-proxy example.** This is a deliberate narrowing of the roadmap's "auth/TLS" wording — flagged in §6 for the lead. ### D4 — Metrics without a new dependency The roadmap's tech-stack table says Micrometer→Prometheus. Recommend **not** taking that dep: - This pom already carries an unusually heavy dependency-reconciliation burden (a hand-pinned `jackson-annotations` 3.0-rc5 to reconcile the MCP SDK's Jackson 3 with our Jackson 2.19, a Jetty BOM import to stop version skew, plus four documented accepted-CVE advisories). Every new transitive tree is a real cost here, not a hypothetical one. - The CVE gate that CLAUDE.md mandates for dependency changes (`jetbrains get_file_problems` → Mend.io) **cannot currently be run** — no JetBrains MCP server is connected. Adding a dependency tree we cannot scan violates the project's own stated policy. - The metric set is small and fully known (§4). Prometheus text exposition is a trivial, stable, well-specified format. So: a ~120-line `metrics/Metrics.java` holding `LongAdder` counters and gauge suppliers, rendered to the Prometheus text format at `GET /metrics`. If Micrometer is wanted later for its registry/push ecosystem, this stays a drop-in swap behind the same endpoint. **Flagged in §6 — this deviates from a documented tech-stack choice.** ### D5 — Supervision targets launchd first, systemd second The roadmap says "systemd unit". **This host is macOS — there is no systemd on it** (`systemctl` not found), and the daemon that has been dogfooded for weeks runs as a bare foreground `java -jar`. Ship **both**: - `deploy/dev.ltms.fleet.plist` — launchd agent, the *actual* runtime here, with `KeepAlive` and ordered start after herdr. - `deploy/fleetd.service` — systemd unit for the Linux gateways CB-308 introduces. Ordering after herdr is advisory in both: the herdr socket may not exist at boot, so the daemon must **retry the socket rather than exit** — supervision ordering is a nicety, socket-retry is the actual fix. That retry behaviour is part of CB-504, not a separate ticket. ### D6 — Audit log is a separate append-only stream, not the app log Privileged actions (spawn, stop, send, reply-drain, ack) emit a structured JSON line to a dedicated `audit` SLF4J logger with its own appender, carrying `{ts, role, terminal, pid, action, target, outcome}`. Keeping it off the chatty app logger is what makes it greppable and, later, shippable. **No message *content* in the audit record** — the bridge carries the user's source code and prompts; an audit trail that quietly becomes a transcript archive is a liability, not a control. Content stays out; correlation ids go in. --- ## 3. Authorization model (CB-505) With D1's roles, the rules are small enough to state completely: | Action | REST | PRIMARY | WORKER | ANONYMOUS | |---|---|---|---|---| | spawn worker | `POST /workers` | ✅ | ❌ | ❌ | | stop worker | `DELETE /workers/{paneId}` | ✅ | ❌ | ❌ | | send to a session | `POST /sessions/{id}/message` | ✅ | ❌ | ❌ | | reply | `POST /sessions/{id}/reply` | ❌ | ✅ **own session only** | ❌ | | ask | `POST /sessions/{id}/ask` | ❌ | ✅ **own session only** | ❌ | | drain replies | `GET /sessions/{id}/replies` | ✅ | ❌ | ❌ | | status / list / profiles | `GET …` | ✅ | ✅ | ❌ | | health | `GET /healthz` | ✅ | ✅ | ✅ (unauthenticated by design) | | metrics | `GET /metrics` | ✅ | ✅ | ❌ | The load-bearing row is **"own session only"**: a worker may only reply or ask *as itself*. That is already true de facto — `ConnectionIdentity` derives the terminal rather than reading it from the body — so CB-505 mostly **asserts an existing invariant explicitly** and adds the test that pins it. The one real change is rejecting a worker that names a *different* session id in the path. `/healthz` stays open: it must answer for a load balancer or supervisor before any credential is configured. It already leaks nothing but herdr's version and up/down. ### 3.1 There are TWO entry paths, and only one of them has identity today The wiki describes MCP as "a thin adapter over the REST core". **At the code level that is not literally true, and the difference is security-relevant.** `FleetMcp` calls `MessageService` / `SessionManager` *directly*; it never issues an HTTP request against a Javalin route. And `/mcp` is mounted as a raw servlet on Jetty's `ServletContextHandler` (`FleetApp.build → cfg.jetty.modifyServletContextHandler`), so it does **not** pass through Javalin's `before` filters at all. The current split is the mirror image of what you'd expect: | Path | Caller identity today | Authz today | |---|---|---| | MCP `/mcp` | ✅ resolved per call (`ConnectionIdentity` via the transport-context extractor) | ❌ none | | REST routes | ❌ **none at all** — the session id is taken from the URL path and trusted | ❌ none | So REST is the *more* exposed surface: `POST /sessions/{id}/reply` accepts any `{id}` from the path, whereas the MCP `fleet_reply` derives the worker from the connection and refuses to read it from an argument. Loopback-only bind is what makes this safe today. **Therefore CB-505 must enforce on both paths against one shared resolver** — not at a single choke point. Concretely: a Javalin `before` filter for REST, and the existing transport-context extractor for MCP, both delegating to `auth.CallerResolver`. Any authz check that lives in only one of the two is not a control. --- ## 4. Metric set (CB-502) Deliberately small; every one maps to a failure mode we have actually hit. | Metric | Type | Why it exists | |---|---|---| | `fleet_sends_total{outcome}` | counter | outcome ∈ replied\|completion_fallback\|timeout\|failed — the completion-fallback rate is the health signal for turn detection (CB-115/116/118) | | `fleet_send_duration_seconds` | histogram | delegated turn latency | | `fleet_replies_total{path}` | counter | path ∈ rendezvous\|inbox — how often a reply strands (CB-307's whole reason to exist) | | `fleet_inbox_depth{target}` | gauge | undrained replies; steady-state should be 0 | | `fleet_push_nudges_total{outcome}` | counter | outcome ∈ sent\|exhausted — a rising `exhausted` means the primary is not draining. `sent` was called `delivered` until fleetd #365; it counts the herdr paste-and-submit call returning, never a confirmation the pane read it | | `fleet_spawns_total{kind,outcome}` | counter | outcome ∈ ready\|timeout\|guard_rejected; per peer kind (CB-402) | | `fleet_sessions{state}` | gauge | SPAWNING/READY/BUSY/DONE census | | `fleet_herdr_calls_total{method,outcome}` | counter | socket health — the dependency everything rests on | | `fleet_auth_failures_total{reason}` | counter | only meaningful once CB-501 lands; catches misconfigured workers | --- ## 5. Increment plan Ordered so each step is independently mergeable and the risky one lands first. 1. **CB-501a — `CallerResolver` + `Role`.** Pure refactor: route today's connection identity through the new type, `ANONYMOUS` not yet reachable (loopback-trust default preserves behaviour). Green build, no behaviour change. 2. **CB-501b — token mode + fail-fast bind check.** Config block, bearer parsing, the non-loopback-bind guard. This is the security-relevant commit; keep it small and reviewable. 3. **CB-505 — authz table + audit logger.** Enforce §3 on **both** entry paths (see §3.1); add the audit appender. 4. **CB-502 — `Metrics` + `/metrics`.** Instrument the paths in §4. 5. **CB-503 — CI.** Runs `mvn -B clean install` with `-Dgroups='!contract'` so the live-herdr and RabbitMQ contract tests are excluded; the mock-UDS suite is the CI surface, exactly as the roadmap's testability section intends. 6. **CB-504 — launchd plist + systemd unit + herdr-socket retry.** --- ## 6. Open questions for the lead *All four resolved 2026-07-29 — the lead confirmed D3, D4, and the D6 sink; the CI runner question was answered from the forge itself. Kept here as the decision record.* 1. ✅ **TLS scope (D3) — confirmed.** Bearer auth + the fail-fast bind guard ship in the daemon; TLS terminates at a reverse proxy, documented with a worked example. AMQP gets TLS via an `amqps://` URI. No keystore handling in `fleetd`. 2. ✅ **Micrometer (D4) — confirmed dropped.** Zero-dependency Prometheus text renderer, for the reasons in D4 (pom reconciliation burden + the mandated CVE gate being un-runnable this session). Revisit if a push-gateway or JVM-metrics requirement appears; the endpoint is the swap seam. 3. ~~**CI runner (CB-503).**~~ ✅ **Resolved during design** — a Gitea Actions runner *is* registered and healthy (`lms/alms-memory` has 28 completed runs; `lms/alms` runs on push and pull_request). CB-503 targets `.gitea/workflows/ci.yml` with `runs-on: ubuntu-latest`, matching the sibling repo's convention. Note the runner's image ships an older `default-jdk`, so the workflow must provision **JDK 25** explicitly rather than apt-installing the default. Contract-test exclusion needs no CI flag: the pom's `default-excludes` profile already sets `excludedGroups=contract`, so a plain `mvn -B clean install` *is* the mock-socket surface. 4. ✅ **Audit sink — confirmed dedicated file.** Its own logback appender writing JSON lines beside the daemon, separate from the app log, per D6. Content still never enters the record.