Files
fleetd/docs/CB-5xx-Hardening.md
T
Dai Ha cbc732444f
CI / contract (push) Successful in 1m22s
CI / build (push) Successful in 1m31s
fleetd #365: the push-nudge metric outcome is 'sent', not 'delivered'
The label was renamed in the #365 merge. This table still named the old one.
'sent' counts the herdr paste-and-submit call returning, never a confirmation
that the pane read it.
2026-09-09 07:40:57 +07:00

243 lines
14 KiB
Markdown

# CB-5xx — Stage 5 Hardening (auth · metrics · CI · supervision · authz+audit)
**Status:** design note (pre-implementation) — the single-host close-out before cross-host work.
**Covers:** CB-501 (bearer auth + TLS) · CB-502 (`/metrics`) · CB-503 (mock-socket CI) ·
CB-504 (service supervision) · CB-505 (per-session authz + audit log).
**Depends on:** everything shipped through CB-402. Nothing here changes messaging semantics.
**Blocks:** CB-308. The cross-host trust model is CB-308's own gating concern, and it inherits
whatever identity/authz shape lands here — so this stage is deliberately *before* federation,
not after it.
---
## 1. Why this stage is not optional bookkeeping
`fleetd` today has **exactly one security control: the loopback bind**. Every other guarantee
rests on it.
The identity model (`mcp/ConnectionIdentity.java`) resolves a caller from the connection alone —
the OS reports the connecting PID, herdr owns the PID→pane map, so a worker cannot forge another
worker. Its own javadoc is explicit: *"Single-host only (the herd shares the `fleetd` host); the
token path is the split-host fallback."* The token path does not exist yet.
That leaves a seam that is **latent today and load-bearing the moment the bind moves**:
```java
// ConnectionIdentity.resolve — non-loopback callers get a null terminal
if (!isLoopback(remoteAddr)) return new Caller(null, -1);
```
…and `null` terminal is interpreted downstream as **"this caller is the primary"**. Combined:
> Any caller that is not a recognised on-host worker pane is treated as the primary — including,
> if `bind.host` is ever widened, an arbitrary remote client.
Today `bind` defaults to `127.0.0.1` so this is unreachable. But CB-308 exists precisely to widen
the boundary, and the primary is the *most* privileged role on the bus (it spawns, stops, sends to
any session, and drains any inbox). Shipping federation on top of "unauthenticated ⇒ primary"
would be building the security boundary backwards.
**So CB-501 is not "add a token header". It is: make identity explicit, and make the absence of
identity mean *nothing*, not *everything*.**
---
## 2. Decisions (locked)
### D1 — Three caller roles, one resolution path
Introduce `Role { PRIMARY, WORKER, ANONYMOUS }` resolved by a single `CallerResolver` that both
REST and MCP go through. Resolution order:
1. **Connection identity wins where it applies.** A loopback peer PID that maps to a herdr worker
pane ⇒ `WORKER` with that terminal. Unforgeable, unchanged from today, zero config.
2. **Token, if presented.** A valid bearer token ⇒ the role that token is provisioned for.
3. **Otherwise `ANONYMOUS`** — *not* `PRIMARY`.
This inverts today's default. `PRIMARY` becomes something you must *prove* (by being a loopback
non-worker process when auth is disabled, or by presenting a primary-scoped token when it is
enabled), rather than something you get by failing every other check.
### D2 — Auth is opt-in by config, but the *default* must stay zero-friction on loopback
The daemon is dogfooded constantly on one machine. If enabling hardening breaks the local setup,
it will be disabled and the stage is wasted. So:
```yaml
auth:
mode: loopback-trust # default — behaves exactly like today: loopback ⇒ PRIMARY, no token needed
# mode: token # every non-worker caller must present a valid bearer token
# tokenEnv: FLEETD_API_TOKEN # host env var holding the token; never the literal value
```
`mode: loopback-trust` is the current behaviour, named honestly and now *chosen* rather than
implied. `mode: token` is what a non-loopback bind requires. **A non-loopback `bind.host` with
`mode: loopback-trust` must fail fast at startup** — that check is the single highest-value line
in this stage, because it makes the dangerous configuration unrepresentable rather than merely
discouraged.
### D3 — TLS terminates *outside* the JVM
Do **not** add TLS config to Javalin/Jetty. The deployment story for a cross-host gateway is a
reverse proxy (or an SSH/WireGuard tunnel) in front of the daemon; the AMQP link has its own TLS
via the broker URI (`amqps://`). Adding keystore handling here would mean certificate lifecycle
code in a daemon whose whole value is being small, and would duplicate what the proxy does better.
**CB-501 therefore ships bearer auth + the fail-fast bind check, and documents TLS as a
deployment concern with a worked reverse-proxy example.** This is a deliberate narrowing of the
roadmap's "auth/TLS" wording — flagged in §6 for the lead.
### D4 — Metrics without a new dependency
The roadmap's tech-stack table says Micrometer→Prometheus. Recommend **not** taking that dep:
- This pom already carries an unusually heavy dependency-reconciliation burden (a hand-pinned
`jackson-annotations` 3.0-rc5 to reconcile the MCP SDK's Jackson 3 with our Jackson 2.19, a
Jetty BOM import to stop version skew, plus four documented accepted-CVE advisories). Every new
transitive tree is a real cost here, not a hypothetical one.
- The CVE gate that CLAUDE.md mandates for dependency changes (`jetbrains get_file_problems` →
Mend.io) **cannot currently be run** — no JetBrains MCP server is connected. Adding a dependency
tree we cannot scan violates the project's own stated policy.
- The metric set is small and fully known (§4). Prometheus text exposition is a trivial,
stable, well-specified format.
So: a ~120-line `metrics/Metrics.java` holding `LongAdder` counters and gauge suppliers, rendered
to the Prometheus text format at `GET /metrics`. If Micrometer is wanted later for its
registry/push ecosystem, this stays a drop-in swap behind the same endpoint. **Flagged in §6 —
this deviates from a documented tech-stack choice.**
### D5 — Supervision targets launchd first, systemd second
The roadmap says "systemd unit". **This host is macOS — there is no systemd on it** (`systemctl`
not found), and the daemon that has been dogfooded for weeks runs as a bare foreground
`java -jar`. Ship **both**:
- `deploy/dev.ltms.fleet.plist` — launchd agent, the *actual* runtime here, with `KeepAlive` and
ordered start after herdr.
- `deploy/fleetd.service` — systemd unit for the Linux gateways CB-308 introduces.
Ordering after herdr is advisory in both: the herdr socket may not exist at boot, so the daemon
must **retry the socket rather than exit** — supervision ordering is a nicety, socket-retry is the
actual fix. That retry behaviour is part of CB-504, not a separate ticket.
### D6 — Audit log is a separate append-only stream, not the app log
Privileged actions (spawn, stop, send, reply-drain, ack) emit a structured JSON line to a
dedicated `audit` SLF4J logger with its own appender, carrying `{ts, role, terminal, pid, action,
target, outcome}`. Keeping it off the chatty app logger is what makes it greppable and, later,
shippable. **No message *content* in the audit record** — the bridge carries the user's source
code and prompts; an audit trail that quietly becomes a transcript archive is a liability, not a
control. Content stays out; correlation ids go in.
---
## 3. Authorization model (CB-505)
With D1's roles, the rules are small enough to state completely:
| Action | REST | PRIMARY | WORKER | ANONYMOUS |
|---|---|---|---|---|
| spawn worker | `POST /workers` | ✅ | ❌ | ❌ |
| stop worker | `DELETE /workers/{paneId}` | ✅ | ❌ | ❌ |
| send to a session | `POST /sessions/{id}/message` | ✅ | ❌ | ❌ |
| reply | `POST /sessions/{id}/reply` | ❌ | ✅ **own session only** | ❌ |
| ask | `POST /sessions/{id}/ask` | ❌ | ✅ **own session only** | ❌ |
| drain replies | `GET /sessions/{id}/replies` | ✅ | ❌ | ❌ |
| status / list / profiles | `GET …` | ✅ | ✅ | ❌ |
| health | `GET /healthz` | ✅ | ✅ | ✅ (unauthenticated by design) |
| metrics | `GET /metrics` | ✅ | ✅ | ❌ |
The load-bearing row is **"own session only"**: a worker may only reply or ask *as itself*. That is
already true de facto — `ConnectionIdentity` derives the terminal rather than reading it from the
body — so CB-505 mostly **asserts an existing invariant explicitly** and adds the test that pins
it. The one real change is rejecting a worker that names a *different* session id in the path.
`/healthz` stays open: it must answer for a load balancer or supervisor before any credential is
configured. It already leaks nothing but herdr's version and up/down.
### 3.1 There are TWO entry paths, and only one of them has identity today
The wiki describes MCP as "a thin adapter over the REST core". **At the code level that is not
literally true, and the difference is security-relevant.** `FleetMcp` calls `MessageService` /
`SessionManager` *directly*; it never issues an HTTP request against a Javalin route. And `/mcp` is
mounted as a raw servlet on Jetty's `ServletContextHandler`
(`FleetApp.build → cfg.jetty.modifyServletContextHandler`), so it does **not** pass through
Javalin's `before` filters at all.
The current split is the mirror image of what you'd expect:
| Path | Caller identity today | Authz today |
|---|---|---|
| MCP `/mcp` | ✅ resolved per call (`ConnectionIdentity` via the transport-context extractor) | ❌ none |
| REST routes | ❌ **none at all** — the session id is taken from the URL path and trusted | ❌ none |
So REST is the *more* exposed surface: `POST /sessions/{id}/reply` accepts any `{id}` from the
path, whereas the MCP `fleet_reply` derives the worker from the connection and refuses to read it
from an argument. Loopback-only bind is what makes this safe today.
**Therefore CB-505 must enforce on both paths against one shared resolver** — not at a single
choke point. Concretely: a Javalin `before` filter for REST, and the existing transport-context
extractor for MCP, both delegating to `auth.CallerResolver`. Any authz check that lives in only
one of the two is not a control.
---
## 4. Metric set (CB-502)
Deliberately small; every one maps to a failure mode we have actually hit.
| Metric | Type | Why it exists |
|---|---|---|
| `fleet_sends_total{outcome}` | counter | outcome ∈ replied\|completion_fallback\|timeout\|failed — the completion-fallback rate is the health signal for turn detection (CB-115/116/118) |
| `fleet_send_duration_seconds` | histogram | delegated turn latency |
| `fleet_replies_total{path}` | counter | path ∈ rendezvous\|inbox — how often a reply strands (CB-307's whole reason to exist) |
| `fleet_inbox_depth{target}` | gauge | undrained replies; steady-state should be 0 |
| `fleet_push_nudges_total{outcome}` | counter | outcome ∈ sent\|exhausted — a rising `exhausted` means the primary is not draining. `sent` was called `delivered` until fleetd #365; it counts the herdr paste-and-submit call returning, never a confirmation the pane read it |
| `fleet_spawns_total{kind,outcome}` | counter | outcome ∈ ready\|timeout\|guard_rejected; per peer kind (CB-402) |
| `fleet_sessions{state}` | gauge | SPAWNING/READY/BUSY/DONE census |
| `fleet_herdr_calls_total{method,outcome}` | counter | socket health — the dependency everything rests on |
| `fleet_auth_failures_total{reason}` | counter | only meaningful once CB-501 lands; catches misconfigured workers |
---
## 5. Increment plan
Ordered so each step is independently mergeable and the risky one lands first.
1. **CB-501a — `CallerResolver` + `Role`.** Pure refactor: route today's connection identity through
the new type, `ANONYMOUS` not yet reachable (loopback-trust default preserves behaviour).
Green build, no behaviour change.
2. **CB-501b — token mode + fail-fast bind check.** Config block, bearer parsing, the
non-loopback-bind guard. This is the security-relevant commit; keep it small and reviewable.
3. **CB-505 — authz table + audit logger.** Enforce §3 on **both** entry paths (see §3.1); add the
audit appender.
4. **CB-502 — `Metrics` + `/metrics`.** Instrument the paths in §4.
5. **CB-503 — CI.** Runs `mvn -B clean install` with `-Dgroups='!contract'` so the live-herdr and
RabbitMQ contract tests are excluded; the mock-UDS suite is the CI surface, exactly as the
roadmap's testability section intends.
6. **CB-504 — launchd plist + systemd unit + herdr-socket retry.**
---
## 6. Open questions for the lead
*All four resolved 2026-07-29 — the lead confirmed D3, D4, and the D6 sink; the CI runner question
was answered from the forge itself. Kept here as the decision record.*
1. ✅ **TLS scope (D3) — confirmed.** Bearer auth + the fail-fast bind guard ship in the daemon;
TLS terminates at a reverse proxy, documented with a worked example. AMQP gets TLS via an
`amqps://` URI. No keystore handling in `fleetd`.
2. ✅ **Micrometer (D4) — confirmed dropped.** Zero-dependency Prometheus text renderer, for the
reasons in D4 (pom reconciliation burden + the mandated CVE gate being un-runnable this
session). Revisit if a push-gateway or JVM-metrics requirement appears; the endpoint is the
swap seam.
3. ~~**CI runner (CB-503).**~~ ✅ **Resolved during design** — a Gitea Actions runner *is*
registered and healthy (`lms/alms-memory` has 28 completed runs; `lms/alms` runs on push and
pull_request). CB-503 targets `.gitea/workflows/ci.yml` with `runs-on: ubuntu-latest`, matching
the sibling repo's convention. Note the runner's image ships an older `default-jdk`, so the
workflow must provision **JDK 25** explicitly rather than apt-installing the default.
Contract-test exclusion needs no CI flag: the pom's `default-excludes` profile already sets
`excludedGroups=contract`, so a plain `mvn -B clean install` *is* the mock-socket surface.
4. ✅ **Audit sink — confirmed dedicated file.** Its own logback appender writing JSON lines beside
the daemon, separate from the app log, per D6. Content still never enters the record.