cbc732444f
The label was renamed in the #365 merge. This table still named the old one. 'sent' counts the herdr paste-and-submit call returning, never a confirmation that the pane read it.
243 lines
14 KiB
Markdown
243 lines
14 KiB
Markdown
# CB-5xx — Stage 5 Hardening (auth · metrics · CI · supervision · authz+audit)
|
|
|
|
**Status:** design note (pre-implementation) — the single-host close-out before cross-host work.
|
|
**Covers:** CB-501 (bearer auth + TLS) · CB-502 (`/metrics`) · CB-503 (mock-socket CI) ·
|
|
CB-504 (service supervision) · CB-505 (per-session authz + audit log).
|
|
**Depends on:** everything shipped through CB-402. Nothing here changes messaging semantics.
|
|
**Blocks:** CB-308. The cross-host trust model is CB-308's own gating concern, and it inherits
|
|
whatever identity/authz shape lands here — so this stage is deliberately *before* federation,
|
|
not after it.
|
|
|
|
---
|
|
|
|
## 1. Why this stage is not optional bookkeeping
|
|
|
|
`fleetd` today has **exactly one security control: the loopback bind**. Every other guarantee
|
|
rests on it.
|
|
|
|
The identity model (`mcp/ConnectionIdentity.java`) resolves a caller from the connection alone —
|
|
the OS reports the connecting PID, herdr owns the PID→pane map, so a worker cannot forge another
|
|
worker. Its own javadoc is explicit: *"Single-host only (the herd shares the `fleetd` host); the
|
|
token path is the split-host fallback."* The token path does not exist yet.
|
|
|
|
That leaves a seam that is **latent today and load-bearing the moment the bind moves**:
|
|
|
|
```java
|
|
// ConnectionIdentity.resolve — non-loopback callers get a null terminal
|
|
if (!isLoopback(remoteAddr)) return new Caller(null, -1);
|
|
```
|
|
|
|
…and `null` terminal is interpreted downstream as **"this caller is the primary"**. Combined:
|
|
|
|
> Any caller that is not a recognised on-host worker pane is treated as the primary — including,
|
|
> if `bind.host` is ever widened, an arbitrary remote client.
|
|
|
|
Today `bind` defaults to `127.0.0.1` so this is unreachable. But CB-308 exists precisely to widen
|
|
the boundary, and the primary is the *most* privileged role on the bus (it spawns, stops, sends to
|
|
any session, and drains any inbox). Shipping federation on top of "unauthenticated ⇒ primary"
|
|
would be building the security boundary backwards.
|
|
|
|
**So CB-501 is not "add a token header". It is: make identity explicit, and make the absence of
|
|
identity mean *nothing*, not *everything*.**
|
|
|
|
---
|
|
|
|
## 2. Decisions (locked)
|
|
|
|
### D1 — Three caller roles, one resolution path
|
|
|
|
Introduce `Role { PRIMARY, WORKER, ANONYMOUS }` resolved by a single `CallerResolver` that both
|
|
REST and MCP go through. Resolution order:
|
|
|
|
1. **Connection identity wins where it applies.** A loopback peer PID that maps to a herdr worker
|
|
pane ⇒ `WORKER` with that terminal. Unforgeable, unchanged from today, zero config.
|
|
2. **Token, if presented.** A valid bearer token ⇒ the role that token is provisioned for.
|
|
3. **Otherwise `ANONYMOUS`** — *not* `PRIMARY`.
|
|
|
|
This inverts today's default. `PRIMARY` becomes something you must *prove* (by being a loopback
|
|
non-worker process when auth is disabled, or by presenting a primary-scoped token when it is
|
|
enabled), rather than something you get by failing every other check.
|
|
|
|
### D2 — Auth is opt-in by config, but the *default* must stay zero-friction on loopback
|
|
|
|
The daemon is dogfooded constantly on one machine. If enabling hardening breaks the local setup,
|
|
it will be disabled and the stage is wasted. So:
|
|
|
|
```yaml
|
|
auth:
|
|
mode: loopback-trust # default — behaves exactly like today: loopback ⇒ PRIMARY, no token needed
|
|
# mode: token # every non-worker caller must present a valid bearer token
|
|
# tokenEnv: FLEETD_API_TOKEN # host env var holding the token; never the literal value
|
|
```
|
|
|
|
`mode: loopback-trust` is the current behaviour, named honestly and now *chosen* rather than
|
|
implied. `mode: token` is what a non-loopback bind requires. **A non-loopback `bind.host` with
|
|
`mode: loopback-trust` must fail fast at startup** — that check is the single highest-value line
|
|
in this stage, because it makes the dangerous configuration unrepresentable rather than merely
|
|
discouraged.
|
|
|
|
### D3 — TLS terminates *outside* the JVM
|
|
|
|
Do **not** add TLS config to Javalin/Jetty. The deployment story for a cross-host gateway is a
|
|
reverse proxy (or an SSH/WireGuard tunnel) in front of the daemon; the AMQP link has its own TLS
|
|
via the broker URI (`amqps://`). Adding keystore handling here would mean certificate lifecycle
|
|
code in a daemon whose whole value is being small, and would duplicate what the proxy does better.
|
|
|
|
**CB-501 therefore ships bearer auth + the fail-fast bind check, and documents TLS as a
|
|
deployment concern with a worked reverse-proxy example.** This is a deliberate narrowing of the
|
|
roadmap's "auth/TLS" wording — flagged in §6 for the lead.
|
|
|
|
### D4 — Metrics without a new dependency
|
|
|
|
The roadmap's tech-stack table says Micrometer→Prometheus. Recommend **not** taking that dep:
|
|
|
|
- This pom already carries an unusually heavy dependency-reconciliation burden (a hand-pinned
|
|
`jackson-annotations` 3.0-rc5 to reconcile the MCP SDK's Jackson 3 with our Jackson 2.19, a
|
|
Jetty BOM import to stop version skew, plus four documented accepted-CVE advisories). Every new
|
|
transitive tree is a real cost here, not a hypothetical one.
|
|
- The CVE gate that CLAUDE.md mandates for dependency changes (`jetbrains get_file_problems` →
|
|
Mend.io) **cannot currently be run** — no JetBrains MCP server is connected. Adding a dependency
|
|
tree we cannot scan violates the project's own stated policy.
|
|
- The metric set is small and fully known (§4). Prometheus text exposition is a trivial,
|
|
stable, well-specified format.
|
|
|
|
So: a ~120-line `metrics/Metrics.java` holding `LongAdder` counters and gauge suppliers, rendered
|
|
to the Prometheus text format at `GET /metrics`. If Micrometer is wanted later for its
|
|
registry/push ecosystem, this stays a drop-in swap behind the same endpoint. **Flagged in §6 —
|
|
this deviates from a documented tech-stack choice.**
|
|
|
|
### D5 — Supervision targets launchd first, systemd second
|
|
|
|
The roadmap says "systemd unit". **This host is macOS — there is no systemd on it** (`systemctl`
|
|
not found), and the daemon that has been dogfooded for weeks runs as a bare foreground
|
|
`java -jar`. Ship **both**:
|
|
|
|
- `deploy/dev.ltms.fleet.plist` — launchd agent, the *actual* runtime here, with `KeepAlive` and
|
|
ordered start after herdr.
|
|
- `deploy/fleetd.service` — systemd unit for the Linux gateways CB-308 introduces.
|
|
|
|
Ordering after herdr is advisory in both: the herdr socket may not exist at boot, so the daemon
|
|
must **retry the socket rather than exit** — supervision ordering is a nicety, socket-retry is the
|
|
actual fix. That retry behaviour is part of CB-504, not a separate ticket.
|
|
|
|
### D6 — Audit log is a separate append-only stream, not the app log
|
|
|
|
Privileged actions (spawn, stop, send, reply-drain, ack) emit a structured JSON line to a
|
|
dedicated `audit` SLF4J logger with its own appender, carrying `{ts, role, terminal, pid, action,
|
|
target, outcome}`. Keeping it off the chatty app logger is what makes it greppable and, later,
|
|
shippable. **No message *content* in the audit record** — the bridge carries the user's source
|
|
code and prompts; an audit trail that quietly becomes a transcript archive is a liability, not a
|
|
control. Content stays out; correlation ids go in.
|
|
|
|
---
|
|
|
|
## 3. Authorization model (CB-505)
|
|
|
|
With D1's roles, the rules are small enough to state completely:
|
|
|
|
| Action | REST | PRIMARY | WORKER | ANONYMOUS |
|
|
|---|---|---|---|---|
|
|
| spawn worker | `POST /workers` | ✅ | ❌ | ❌ |
|
|
| stop worker | `DELETE /workers/{paneId}` | ✅ | ❌ | ❌ |
|
|
| send to a session | `POST /sessions/{id}/message` | ✅ | ❌ | ❌ |
|
|
| reply | `POST /sessions/{id}/reply` | ❌ | ✅ **own session only** | ❌ |
|
|
| ask | `POST /sessions/{id}/ask` | ❌ | ✅ **own session only** | ❌ |
|
|
| drain replies | `GET /sessions/{id}/replies` | ✅ | ❌ | ❌ |
|
|
| status / list / profiles | `GET …` | ✅ | ✅ | ❌ |
|
|
| health | `GET /healthz` | ✅ | ✅ | ✅ (unauthenticated by design) |
|
|
| metrics | `GET /metrics` | ✅ | ✅ | ❌ |
|
|
|
|
The load-bearing row is **"own session only"**: a worker may only reply or ask *as itself*. That is
|
|
already true de facto — `ConnectionIdentity` derives the terminal rather than reading it from the
|
|
body — so CB-505 mostly **asserts an existing invariant explicitly** and adds the test that pins
|
|
it. The one real change is rejecting a worker that names a *different* session id in the path.
|
|
|
|
`/healthz` stays open: it must answer for a load balancer or supervisor before any credential is
|
|
configured. It already leaks nothing but herdr's version and up/down.
|
|
|
|
### 3.1 There are TWO entry paths, and only one of them has identity today
|
|
|
|
The wiki describes MCP as "a thin adapter over the REST core". **At the code level that is not
|
|
literally true, and the difference is security-relevant.** `FleetMcp` calls `MessageService` /
|
|
`SessionManager` *directly*; it never issues an HTTP request against a Javalin route. And `/mcp` is
|
|
mounted as a raw servlet on Jetty's `ServletContextHandler`
|
|
(`FleetApp.build → cfg.jetty.modifyServletContextHandler`), so it does **not** pass through
|
|
Javalin's `before` filters at all.
|
|
|
|
The current split is the mirror image of what you'd expect:
|
|
|
|
| Path | Caller identity today | Authz today |
|
|
|---|---|---|
|
|
| MCP `/mcp` | ✅ resolved per call (`ConnectionIdentity` via the transport-context extractor) | ❌ none |
|
|
| REST routes | ❌ **none at all** — the session id is taken from the URL path and trusted | ❌ none |
|
|
|
|
So REST is the *more* exposed surface: `POST /sessions/{id}/reply` accepts any `{id}` from the
|
|
path, whereas the MCP `fleet_reply` derives the worker from the connection and refuses to read it
|
|
from an argument. Loopback-only bind is what makes this safe today.
|
|
|
|
**Therefore CB-505 must enforce on both paths against one shared resolver** — not at a single
|
|
choke point. Concretely: a Javalin `before` filter for REST, and the existing transport-context
|
|
extractor for MCP, both delegating to `auth.CallerResolver`. Any authz check that lives in only
|
|
one of the two is not a control.
|
|
|
|
---
|
|
|
|
## 4. Metric set (CB-502)
|
|
|
|
Deliberately small; every one maps to a failure mode we have actually hit.
|
|
|
|
| Metric | Type | Why it exists |
|
|
|---|---|---|
|
|
| `fleet_sends_total{outcome}` | counter | outcome ∈ replied\|completion_fallback\|timeout\|failed — the completion-fallback rate is the health signal for turn detection (CB-115/116/118) |
|
|
| `fleet_send_duration_seconds` | histogram | delegated turn latency |
|
|
| `fleet_replies_total{path}` | counter | path ∈ rendezvous\|inbox — how often a reply strands (CB-307's whole reason to exist) |
|
|
| `fleet_inbox_depth{target}` | gauge | undrained replies; steady-state should be 0 |
|
|
| `fleet_push_nudges_total{outcome}` | counter | outcome ∈ sent\|exhausted — a rising `exhausted` means the primary is not draining. `sent` was called `delivered` until fleetd #365; it counts the herdr paste-and-submit call returning, never a confirmation the pane read it |
|
|
| `fleet_spawns_total{kind,outcome}` | counter | outcome ∈ ready\|timeout\|guard_rejected; per peer kind (CB-402) |
|
|
| `fleet_sessions{state}` | gauge | SPAWNING/READY/BUSY/DONE census |
|
|
| `fleet_herdr_calls_total{method,outcome}` | counter | socket health — the dependency everything rests on |
|
|
| `fleet_auth_failures_total{reason}` | counter | only meaningful once CB-501 lands; catches misconfigured workers |
|
|
|
|
---
|
|
|
|
## 5. Increment plan
|
|
|
|
Ordered so each step is independently mergeable and the risky one lands first.
|
|
|
|
1. **CB-501a — `CallerResolver` + `Role`.** Pure refactor: route today's connection identity through
|
|
the new type, `ANONYMOUS` not yet reachable (loopback-trust default preserves behaviour).
|
|
Green build, no behaviour change.
|
|
2. **CB-501b — token mode + fail-fast bind check.** Config block, bearer parsing, the
|
|
non-loopback-bind guard. This is the security-relevant commit; keep it small and reviewable.
|
|
3. **CB-505 — authz table + audit logger.** Enforce §3 on **both** entry paths (see §3.1); add the
|
|
audit appender.
|
|
4. **CB-502 — `Metrics` + `/metrics`.** Instrument the paths in §4.
|
|
5. **CB-503 — CI.** Runs `mvn -B clean install` with `-Dgroups='!contract'` so the live-herdr and
|
|
RabbitMQ contract tests are excluded; the mock-UDS suite is the CI surface, exactly as the
|
|
roadmap's testability section intends.
|
|
6. **CB-504 — launchd plist + systemd unit + herdr-socket retry.**
|
|
|
|
---
|
|
|
|
## 6. Open questions for the lead
|
|
|
|
*All four resolved 2026-07-29 — the lead confirmed D3, D4, and the D6 sink; the CI runner question
|
|
was answered from the forge itself. Kept here as the decision record.*
|
|
|
|
1. ✅ **TLS scope (D3) — confirmed.** Bearer auth + the fail-fast bind guard ship in the daemon;
|
|
TLS terminates at a reverse proxy, documented with a worked example. AMQP gets TLS via an
|
|
`amqps://` URI. No keystore handling in `fleetd`.
|
|
2. ✅ **Micrometer (D4) — confirmed dropped.** Zero-dependency Prometheus text renderer, for the
|
|
reasons in D4 (pom reconciliation burden + the mandated CVE gate being un-runnable this
|
|
session). Revisit if a push-gateway or JVM-metrics requirement appears; the endpoint is the
|
|
swap seam.
|
|
3. ~~**CI runner (CB-503).**~~ ✅ **Resolved during design** — a Gitea Actions runner *is*
|
|
registered and healthy (`lms/alms-memory` has 28 completed runs; `lms/alms` runs on push and
|
|
pull_request). CB-503 targets `.gitea/workflows/ci.yml` with `runs-on: ubuntu-latest`, matching
|
|
the sibling repo's convention. Note the runner's image ships an older `default-jdk`, so the
|
|
workflow must provision **JDK 25** explicitly rather than apt-installing the default.
|
|
Contract-test exclusion needs no CI flag: the pom's `default-excludes` profile already sets
|
|
`excludedGroups=contract`, so a plain `mvn -B clean install` *is* the mock-socket surface.
|
|
4. ✅ **Audit sink — confirmed dedicated file.** Its own logback appender writing JSON lines beside
|
|
the daemon, separate from the app log, per D6. Content still never enters the record.
|