Files
fleetd/docs/CB-5xx-Hardening.md
T
Dai Ha cbc732444f
CI / contract (push) Successful in 1m22s
CI / build (push) Successful in 1m31s
fleetd #365: the push-nudge metric outcome is 'sent', not 'delivered'
The label was renamed in the #365 merge. This table still named the old one.
'sent' counts the herdr paste-and-submit call returning, never a confirmation
that the pane read it.
2026-09-09 07:40:57 +07:00

14 KiB

CB-5xx — Stage 5 Hardening (auth · metrics · CI · supervision · authz+audit)

Status: design note (pre-implementation) — the single-host close-out before cross-host work. Covers: CB-501 (bearer auth + TLS) · CB-502 (/metrics) · CB-503 (mock-socket CI) · CB-504 (service supervision) · CB-505 (per-session authz + audit log). Depends on: everything shipped through CB-402. Nothing here changes messaging semantics. Blocks: CB-308. The cross-host trust model is CB-308's own gating concern, and it inherits whatever identity/authz shape lands here — so this stage is deliberately before federation, not after it.


1. Why this stage is not optional bookkeeping

fleetd today has exactly one security control: the loopback bind. Every other guarantee rests on it.

The identity model (mcp/ConnectionIdentity.java) resolves a caller from the connection alone — the OS reports the connecting PID, herdr owns the PID→pane map, so a worker cannot forge another worker. Its own javadoc is explicit: "Single-host only (the herd shares the fleetd host); the token path is the split-host fallback." The token path does not exist yet.

That leaves a seam that is latent today and load-bearing the moment the bind moves:

// ConnectionIdentity.resolve — non-loopback callers get a null terminal
if (!isLoopback(remoteAddr)) return new Caller(null, -1);

…and null terminal is interpreted downstream as "this caller is the primary". Combined:

Any caller that is not a recognised on-host worker pane is treated as the primary — including, if bind.host is ever widened, an arbitrary remote client.

Today bind defaults to 127.0.0.1 so this is unreachable. But CB-308 exists precisely to widen the boundary, and the primary is the most privileged role on the bus (it spawns, stops, sends to any session, and drains any inbox). Shipping federation on top of "unauthenticated ⇒ primary" would be building the security boundary backwards.

So CB-501 is not "add a token header". It is: make identity explicit, and make the absence of identity mean nothing, not everything.


2. Decisions (locked)

D1 — Three caller roles, one resolution path

Introduce Role { PRIMARY, WORKER, ANONYMOUS } resolved by a single CallerResolver that both REST and MCP go through. Resolution order:

  1. Connection identity wins where it applies. A loopback peer PID that maps to a herdr worker pane ⇒ WORKER with that terminal. Unforgeable, unchanged from today, zero config.
  2. Token, if presented. A valid bearer token ⇒ the role that token is provisioned for.
  3. Otherwise ANONYMOUS — not PRIMARY.

This inverts today's default. PRIMARY becomes something you must prove (by being a loopback non-worker process when auth is disabled, or by presenting a primary-scoped token when it is enabled), rather than something you get by failing every other check.

D2 — Auth is opt-in by config, but the default must stay zero-friction on loopback

The daemon is dogfooded constantly on one machine. If enabling hardening breaks the local setup, it will be disabled and the stage is wasted. So:

auth:
  mode: loopback-trust   # default — behaves exactly like today: loopback ⇒ PRIMARY, no token needed
  # mode: token          # every non-worker caller must present a valid bearer token
  # tokenEnv: FLEETD_API_TOKEN   # host env var holding the token; never the literal value

mode: loopback-trust is the current behaviour, named honestly and now chosen rather than implied. mode: token is what a non-loopback bind requires. A non-loopback bind.host with mode: loopback-trust must fail fast at startup — that check is the single highest-value line in this stage, because it makes the dangerous configuration unrepresentable rather than merely discouraged.

D3 — TLS terminates outside the JVM

Do not add TLS config to Javalin/Jetty. The deployment story for a cross-host gateway is a reverse proxy (or an SSH/WireGuard tunnel) in front of the daemon; the AMQP link has its own TLS via the broker URI (amqps://). Adding keystore handling here would mean certificate lifecycle code in a daemon whose whole value is being small, and would duplicate what the proxy does better.

CB-501 therefore ships bearer auth + the fail-fast bind check, and documents TLS as a deployment concern with a worked reverse-proxy example. This is a deliberate narrowing of the roadmap's "auth/TLS" wording — flagged in §6 for the lead.

D4 — Metrics without a new dependency

The roadmap's tech-stack table says Micrometer→Prometheus. Recommend not taking that dep:

  • This pom already carries an unusually heavy dependency-reconciliation burden (a hand-pinned jackson-annotations 3.0-rc5 to reconcile the MCP SDK's Jackson 3 with our Jackson 2.19, a Jetty BOM import to stop version skew, plus four documented accepted-CVE advisories). Every new transitive tree is a real cost here, not a hypothetical one.
  • The CVE gate that CLAUDE.md mandates for dependency changes (jetbrains get_file_problems → Mend.io) cannot currently be run — no JetBrains MCP server is connected. Adding a dependency tree we cannot scan violates the project's own stated policy.
  • The metric set is small and fully known (§4). Prometheus text exposition is a trivial, stable, well-specified format.

So: a ~120-line metrics/Metrics.java holding LongAdder counters and gauge suppliers, rendered to the Prometheus text format at GET /metrics. If Micrometer is wanted later for its registry/push ecosystem, this stays a drop-in swap behind the same endpoint. Flagged in §6 — this deviates from a documented tech-stack choice.

D5 — Supervision targets launchd first, systemd second

The roadmap says "systemd unit". This host is macOS — there is no systemd on it (systemctl not found), and the daemon that has been dogfooded for weeks runs as a bare foreground java -jar. Ship both:

  • deploy/dev.ltms.fleet.plist — launchd agent, the actual runtime here, with KeepAlive and ordered start after herdr.
  • deploy/fleetd.service — systemd unit for the Linux gateways CB-308 introduces.

Ordering after herdr is advisory in both: the herdr socket may not exist at boot, so the daemon must retry the socket rather than exit — supervision ordering is a nicety, socket-retry is the actual fix. That retry behaviour is part of CB-504, not a separate ticket.

D6 — Audit log is a separate append-only stream, not the app log

Privileged actions (spawn, stop, send, reply-drain, ack) emit a structured JSON line to a dedicated audit SLF4J logger with its own appender, carrying {ts, role, terminal, pid, action, target, outcome}. Keeping it off the chatty app logger is what makes it greppable and, later, shippable. No message content in the audit record — the bridge carries the user's source code and prompts; an audit trail that quietly becomes a transcript archive is a liability, not a control. Content stays out; correlation ids go in.


3. Authorization model (CB-505)

With D1's roles, the rules are small enough to state completely:

Action REST PRIMARY WORKER ANONYMOUS
spawn worker POST /workers ✅ ❌ ❌
stop worker DELETE /workers/{paneId} ✅ ❌ ❌
send to a session POST /sessions/{id}/message ✅ ❌ ❌
reply POST /sessions/{id}/reply ❌ ✅ own session only ❌
ask POST /sessions/{id}/ask ❌ ✅ own session only ❌
drain replies GET /sessions/{id}/replies ✅ ❌ ❌
status / list / profiles GET … ✅ ✅ ❌
health GET /healthz ✅ ✅ ✅ (unauthenticated by design)
metrics GET /metrics ✅ ✅ ❌

The load-bearing row is "own session only": a worker may only reply or ask as itself. That is already true de facto — ConnectionIdentity derives the terminal rather than reading it from the body — so CB-505 mostly asserts an existing invariant explicitly and adds the test that pins it. The one real change is rejecting a worker that names a different session id in the path.

/healthz stays open: it must answer for a load balancer or supervisor before any credential is configured. It already leaks nothing but herdr's version and up/down.

3.1 There are TWO entry paths, and only one of them has identity today

The wiki describes MCP as "a thin adapter over the REST core". At the code level that is not literally true, and the difference is security-relevant. FleetMcp calls MessageService / SessionManager directly; it never issues an HTTP request against a Javalin route. And /mcp is mounted as a raw servlet on Jetty's ServletContextHandler (FleetApp.build → cfg.jetty.modifyServletContextHandler), so it does not pass through Javalin's before filters at all.

The current split is the mirror image of what you'd expect:

Path Caller identity today Authz today
MCP /mcp ✅ resolved per call (ConnectionIdentity via the transport-context extractor) ❌ none
REST routes ❌ none at all — the session id is taken from the URL path and trusted ❌ none

So REST is the more exposed surface: POST /sessions/{id}/reply accepts any {id} from the path, whereas the MCP fleet_reply derives the worker from the connection and refuses to read it from an argument. Loopback-only bind is what makes this safe today.

Therefore CB-505 must enforce on both paths against one shared resolver — not at a single choke point. Concretely: a Javalin before filter for REST, and the existing transport-context extractor for MCP, both delegating to auth.CallerResolver. Any authz check that lives in only one of the two is not a control.


4. Metric set (CB-502)

Deliberately small; every one maps to a failure mode we have actually hit.

Metric Type Why it exists
fleet_sends_total{outcome} counter outcome ∈ replied|completion_fallback|timeout|failed — the completion-fallback rate is the health signal for turn detection (CB-115/116/118)
fleet_send_duration_seconds histogram delegated turn latency
fleet_replies_total{path} counter path ∈ rendezvous|inbox — how often a reply strands (CB-307's whole reason to exist)
fleet_inbox_depth{target} gauge undrained replies; steady-state should be 0
fleet_push_nudges_total{outcome} counter outcome ∈ sent|exhausted — a rising exhausted means the primary is not draining. sent was called delivered until fleetd #365; it counts the herdr paste-and-submit call returning, never a confirmation the pane read it
fleet_spawns_total{kind,outcome} counter outcome ∈ ready|timeout|guard_rejected; per peer kind (CB-402)
fleet_sessions{state} gauge SPAWNING/READY/BUSY/DONE census
fleet_herdr_calls_total{method,outcome} counter socket health — the dependency everything rests on
fleet_auth_failures_total{reason} counter only meaningful once CB-501 lands; catches misconfigured workers

5. Increment plan

Ordered so each step is independently mergeable and the risky one lands first.

  1. CB-501a — CallerResolver + Role. Pure refactor: route today's connection identity through the new type, ANONYMOUS not yet reachable (loopback-trust default preserves behaviour). Green build, no behaviour change.
  2. CB-501b — token mode + fail-fast bind check. Config block, bearer parsing, the non-loopback-bind guard. This is the security-relevant commit; keep it small and reviewable.
  3. CB-505 — authz table + audit logger. Enforce §3 on both entry paths (see §3.1); add the audit appender.
  4. CB-502 — Metrics + /metrics. Instrument the paths in §4.
  5. CB-503 — CI. Runs mvn -B clean install with -Dgroups='!contract' so the live-herdr and RabbitMQ contract tests are excluded; the mock-UDS suite is the CI surface, exactly as the roadmap's testability section intends.
  6. CB-504 — launchd plist + systemd unit + herdr-socket retry.

6. Open questions for the lead

All four resolved 2026-07-29 — the lead confirmed D3, D4, and the D6 sink; the CI runner question was answered from the forge itself. Kept here as the decision record.

  1. ✅ TLS scope (D3) — confirmed. Bearer auth + the fail-fast bind guard ship in the daemon; TLS terminates at a reverse proxy, documented with a worked example. AMQP gets TLS via an amqps:// URI. No keystore handling in fleetd.
  2. ✅ Micrometer (D4) — confirmed dropped. Zero-dependency Prometheus text renderer, for the reasons in D4 (pom reconciliation burden + the mandated CVE gate being un-runnable this session). Revisit if a push-gateway or JVM-metrics requirement appears; the endpoint is the swap seam.
  3. CI runner (CB-503). ✅ Resolved during design — a Gitea Actions runner is registered and healthy (lms/alms-memory has 28 completed runs; lms/alms runs on push and pull_request). CB-503 targets .gitea/workflows/ci.yml with runs-on: ubuntu-latest, matching the sibling repo's convention. Note the runner's image ships an older default-jdk, so the workflow must provision JDK 25 explicitly rather than apt-installing the default. Contract-test exclusion needs no CI flag: the pom's default-excludes profile already sets excludedGroups=contract, so a plain mvn -B clean install is the mock-socket surface.
  4. ✅ Audit sink — confirmed dedicated file. Its own logback appender writing JSON lines beside the daemon, separate from the app log, per D6. Content still never enters the record.