2e138a199b
Two changes ship together here.
1. One shared herdr workspace. The lead and every worker now live in one
workspace called "fleet", so the operator sees one "session" with many
windows, not two. Before, the lead sat in a "leads" workspace and workers
in "bridged-workers", which read as two sessions. The lead is still told
apart from workers by its exact tab label ("lead: <name>"), so putting them
in one space is safe. LeadTabScanner keeps the exclude-by-label mechanism
for split layouts; Fleetd now passes an empty exclude set.
2. Rename the daemon from "bridged" to "fleetd" (the binary, config, scripts,
launchd/systemd units, module dir, and MCP mount).
- Module dir bridged/ -> fleetd/; jar finalName -> fleetd.jar.
- Log line, comments, docs, and CLAUDE.md updated to say fleetd.
- Scripts renamed: redeploy-bridged.sh -> redeploy-fleetd.sh,
bridged-launchd-wrapper.sh -> fleetd-launchd-wrapper.sh.
- Deploy units renamed: dev.ltms.bridged.plist -> dev.ltms.fleetd.plist,
bridged.service -> fleetd.service; launchd Label -> dev.ltms.fleetd.
- Config default bridged.yaml -> fleetd.yaml; the legacy bridged.yaml is
still read as a fallback, and still gitignored.
- MCP: drop the deprecated bridge_* tool twins; only fleet_* remain. The
server name is "fleet". The mount name in the local .mcp.json becomes
"fleet" (gitignored, not in this commit).
- Env var defaults BRIDGED_API_TOKEN -> FLEETD_API_TOKEN, fixture
BRIDGED_WORKER_TOKEN -> FLEETD_WORKER_TOKEN.
Kept on purpose: the BRIDGED_MEMBER marker. Renaming it is a coupled change to
the credential-scrub security control (an operator secrets.sh may guard on it),
so it stays until that migration is done on its own.
Metrics were already fleet_* (CB-632); MetricNamesTest still guards that no
name says bridged_.
The canonical CLAUDE.md block and the wiki template stay byte-identical
(wiki working tree edited, committed to the wiki repo separately).
949 tests pass (mvn clean install). 4 fewer than before = the 4 removed
bridge_* alias tests.
243 lines
14 KiB
Markdown
243 lines
14 KiB
Markdown
# CB-5xx — Stage 5 Hardening (auth · metrics · CI · supervision · authz+audit)
|
|
|
|
**Status:** design note (pre-implementation) — the single-host close-out before cross-host work.
|
|
**Covers:** CB-501 (bearer auth + TLS) · CB-502 (`/metrics`) · CB-503 (mock-socket CI) ·
|
|
CB-504 (service supervision) · CB-505 (per-session authz + audit log).
|
|
**Depends on:** everything shipped through CB-402. Nothing here changes messaging semantics.
|
|
**Blocks:** CB-308. The cross-host trust model is CB-308's own gating concern, and it inherits
|
|
whatever identity/authz shape lands here — so this stage is deliberately *before* federation,
|
|
not after it.
|
|
|
|
---
|
|
|
|
## 1. Why this stage is not optional bookkeeping
|
|
|
|
`fleetd` today has **exactly one security control: the loopback bind**. Every other guarantee
|
|
rests on it.
|
|
|
|
The identity model (`mcp/ConnectionIdentity.java`) resolves a caller from the connection alone —
|
|
the OS reports the connecting PID, herdr owns the PID→pane map, so a worker cannot forge another
|
|
worker. Its own javadoc is explicit: *"Single-host only (the herd shares the `fleetd` host); the
|
|
token path is the split-host fallback."* The token path does not exist yet.
|
|
|
|
That leaves a seam that is **latent today and load-bearing the moment the bind moves**:
|
|
|
|
```java
|
|
// ConnectionIdentity.resolve — non-loopback callers get a null terminal
|
|
if (!isLoopback(remoteAddr)) return new Caller(null, -1);
|
|
```
|
|
|
|
…and `null` terminal is interpreted downstream as **"this caller is the primary"**. Combined:
|
|
|
|
> Any caller that is not a recognised on-host worker pane is treated as the primary — including,
|
|
> if `bind.host` is ever widened, an arbitrary remote client.
|
|
|
|
Today `bind` defaults to `127.0.0.1` so this is unreachable. But CB-308 exists precisely to widen
|
|
the boundary, and the primary is the *most* privileged role on the bus (it spawns, stops, sends to
|
|
any session, and drains any inbox). Shipping federation on top of "unauthenticated ⇒ primary"
|
|
would be building the security boundary backwards.
|
|
|
|
**So CB-501 is not "add a token header". It is: make identity explicit, and make the absence of
|
|
identity mean *nothing*, not *everything*.**
|
|
|
|
---
|
|
|
|
## 2. Decisions (locked)
|
|
|
|
### D1 — Three caller roles, one resolution path
|
|
|
|
Introduce `Role { PRIMARY, WORKER, ANONYMOUS }` resolved by a single `CallerResolver` that both
|
|
REST and MCP go through. Resolution order:
|
|
|
|
1. **Connection identity wins where it applies.** A loopback peer PID that maps to a herdr worker
|
|
pane ⇒ `WORKER` with that terminal. Unforgeable, unchanged from today, zero config.
|
|
2. **Token, if presented.** A valid bearer token ⇒ the role that token is provisioned for.
|
|
3. **Otherwise `ANONYMOUS`** — *not* `PRIMARY`.
|
|
|
|
This inverts today's default. `PRIMARY` becomes something you must *prove* (by being a loopback
|
|
non-worker process when auth is disabled, or by presenting a primary-scoped token when it is
|
|
enabled), rather than something you get by failing every other check.
|
|
|
|
### D2 — Auth is opt-in by config, but the *default* must stay zero-friction on loopback
|
|
|
|
The daemon is dogfooded constantly on one machine. If enabling hardening breaks the local setup,
|
|
it will be disabled and the stage is wasted. So:
|
|
|
|
```yaml
|
|
auth:
|
|
mode: loopback-trust # default — behaves exactly like today: loopback ⇒ PRIMARY, no token needed
|
|
# mode: token # every non-worker caller must present a valid bearer token
|
|
# tokenEnv: FLEETD_API_TOKEN # host env var holding the token; never the literal value
|
|
```
|
|
|
|
`mode: loopback-trust` is the current behaviour, named honestly and now *chosen* rather than
|
|
implied. `mode: token` is what a non-loopback bind requires. **A non-loopback `bind.host` with
|
|
`mode: loopback-trust` must fail fast at startup** — that check is the single highest-value line
|
|
in this stage, because it makes the dangerous configuration unrepresentable rather than merely
|
|
discouraged.
|
|
|
|
### D3 — TLS terminates *outside* the JVM
|
|
|
|
Do **not** add TLS config to Javalin/Jetty. The deployment story for a cross-host gateway is a
|
|
reverse proxy (or an SSH/WireGuard tunnel) in front of the daemon; the AMQP link has its own TLS
|
|
via the broker URI (`amqps://`). Adding keystore handling here would mean certificate lifecycle
|
|
code in a daemon whose whole value is being small, and would duplicate what the proxy does better.
|
|
|
|
**CB-501 therefore ships bearer auth + the fail-fast bind check, and documents TLS as a
|
|
deployment concern with a worked reverse-proxy example.** This is a deliberate narrowing of the
|
|
roadmap's "auth/TLS" wording — flagged in §6 for the lead.
|
|
|
|
### D4 — Metrics without a new dependency
|
|
|
|
The roadmap's tech-stack table says Micrometer→Prometheus. Recommend **not** taking that dep:
|
|
|
|
- This pom already carries an unusually heavy dependency-reconciliation burden (a hand-pinned
|
|
`jackson-annotations` 3.0-rc5 to reconcile the MCP SDK's Jackson 3 with our Jackson 2.19, a
|
|
Jetty BOM import to stop version skew, plus four documented accepted-CVE advisories). Every new
|
|
transitive tree is a real cost here, not a hypothetical one.
|
|
- The CVE gate that CLAUDE.md mandates for dependency changes (`jetbrains get_file_problems` →
|
|
Mend.io) **cannot currently be run** — no JetBrains MCP server is connected. Adding a dependency
|
|
tree we cannot scan violates the project's own stated policy.
|
|
- The metric set is small and fully known (§4). Prometheus text exposition is a trivial,
|
|
stable, well-specified format.
|
|
|
|
So: a ~120-line `metrics/Metrics.java` holding `LongAdder` counters and gauge suppliers, rendered
|
|
to the Prometheus text format at `GET /metrics`. If Micrometer is wanted later for its
|
|
registry/push ecosystem, this stays a drop-in swap behind the same endpoint. **Flagged in §6 —
|
|
this deviates from a documented tech-stack choice.**
|
|
|
|
### D5 — Supervision targets launchd first, systemd second
|
|
|
|
The roadmap says "systemd unit". **This host is macOS — there is no systemd on it** (`systemctl`
|
|
not found), and the daemon that has been dogfooded for weeks runs as a bare foreground
|
|
`java -jar`. Ship **both**:
|
|
|
|
- `deploy/dev.ltms.fleet.plist` — launchd agent, the *actual* runtime here, with `KeepAlive` and
|
|
ordered start after herdr.
|
|
- `deploy/fleetd.service` — systemd unit for the Linux gateways CB-308 introduces.
|
|
|
|
Ordering after herdr is advisory in both: the herdr socket may not exist at boot, so the daemon
|
|
must **retry the socket rather than exit** — supervision ordering is a nicety, socket-retry is the
|
|
actual fix. That retry behaviour is part of CB-504, not a separate ticket.
|
|
|
|
### D6 — Audit log is a separate append-only stream, not the app log
|
|
|
|
Privileged actions (spawn, stop, send, reply-drain, ack) emit a structured JSON line to a
|
|
dedicated `audit` SLF4J logger with its own appender, carrying `{ts, role, terminal, pid, action,
|
|
target, outcome}`. Keeping it off the chatty app logger is what makes it greppable and, later,
|
|
shippable. **No message *content* in the audit record** — the bridge carries the user's source
|
|
code and prompts; an audit trail that quietly becomes a transcript archive is a liability, not a
|
|
control. Content stays out; correlation ids go in.
|
|
|
|
---
|
|
|
|
## 3. Authorization model (CB-505)
|
|
|
|
With D1's roles, the rules are small enough to state completely:
|
|
|
|
| Action | REST | PRIMARY | WORKER | ANONYMOUS |
|
|
|---|---|---|---|---|
|
|
| spawn worker | `POST /workers` | ✅ | ❌ | ❌ |
|
|
| stop worker | `DELETE /workers/{paneId}` | ✅ | ❌ | ❌ |
|
|
| send to a session | `POST /sessions/{id}/message` | ✅ | ❌ | ❌ |
|
|
| reply | `POST /sessions/{id}/reply` | ❌ | ✅ **own session only** | ❌ |
|
|
| ask | `POST /sessions/{id}/ask` | ❌ | ✅ **own session only** | ❌ |
|
|
| drain replies | `GET /sessions/{id}/replies` | ✅ | ❌ | ❌ |
|
|
| status / list / profiles | `GET …` | ✅ | ✅ | ❌ |
|
|
| health | `GET /healthz` | ✅ | ✅ | ✅ (unauthenticated by design) |
|
|
| metrics | `GET /metrics` | ✅ | ✅ | ❌ |
|
|
|
|
The load-bearing row is **"own session only"**: a worker may only reply or ask *as itself*. That is
|
|
already true de facto — `ConnectionIdentity` derives the terminal rather than reading it from the
|
|
body — so CB-505 mostly **asserts an existing invariant explicitly** and adds the test that pins
|
|
it. The one real change is rejecting a worker that names a *different* session id in the path.
|
|
|
|
`/healthz` stays open: it must answer for a load balancer or supervisor before any credential is
|
|
configured. It already leaks nothing but herdr's version and up/down.
|
|
|
|
### 3.1 There are TWO entry paths, and only one of them has identity today
|
|
|
|
The wiki describes MCP as "a thin adapter over the REST core". **At the code level that is not
|
|
literally true, and the difference is security-relevant.** `FleetMcp` calls `MessageService` /
|
|
`SessionManager` *directly*; it never issues an HTTP request against a Javalin route. And `/mcp` is
|
|
mounted as a raw servlet on Jetty's `ServletContextHandler`
|
|
(`FleetApp.build → cfg.jetty.modifyServletContextHandler`), so it does **not** pass through
|
|
Javalin's `before` filters at all.
|
|
|
|
The current split is the mirror image of what you'd expect:
|
|
|
|
| Path | Caller identity today | Authz today |
|
|
|---|---|---|
|
|
| MCP `/mcp` | ✅ resolved per call (`ConnectionIdentity` via the transport-context extractor) | ❌ none |
|
|
| REST routes | ❌ **none at all** — the session id is taken from the URL path and trusted | ❌ none |
|
|
|
|
So REST is the *more* exposed surface: `POST /sessions/{id}/reply` accepts any `{id}` from the
|
|
path, whereas the MCP `fleet_reply` derives the worker from the connection and refuses to read it
|
|
from an argument. Loopback-only bind is what makes this safe today.
|
|
|
|
**Therefore CB-505 must enforce on both paths against one shared resolver** — not at a single
|
|
choke point. Concretely: a Javalin `before` filter for REST, and the existing transport-context
|
|
extractor for MCP, both delegating to `auth.CallerResolver`. Any authz check that lives in only
|
|
one of the two is not a control.
|
|
|
|
---
|
|
|
|
## 4. Metric set (CB-502)
|
|
|
|
Deliberately small; every one maps to a failure mode we have actually hit.
|
|
|
|
| Metric | Type | Why it exists |
|
|
|---|---|---|
|
|
| `fleet_sends_total{outcome}` | counter | outcome ∈ replied\|completion_fallback\|timeout\|failed — the completion-fallback rate is the health signal for turn detection (CB-115/116/118) |
|
|
| `fleet_send_duration_seconds` | histogram | delegated turn latency |
|
|
| `fleet_replies_total{path}` | counter | path ∈ rendezvous\|inbox — how often a reply strands (CB-307's whole reason to exist) |
|
|
| `fleet_inbox_depth{target}` | gauge | undrained replies; steady-state should be 0 |
|
|
| `fleet_push_nudges_total{outcome}` | counter | outcome ∈ delivered\|exhausted — a rising `exhausted` means the primary is not draining |
|
|
| `fleet_spawns_total{kind,outcome}` | counter | outcome ∈ ready\|timeout\|guard_rejected; per peer kind (CB-402) |
|
|
| `fleet_sessions{state}` | gauge | SPAWNING/READY/BUSY/DONE census |
|
|
| `fleet_herdr_calls_total{method,outcome}` | counter | socket health — the dependency everything rests on |
|
|
| `fleet_auth_failures_total{reason}` | counter | only meaningful once CB-501 lands; catches misconfigured workers |
|
|
|
|
---
|
|
|
|
## 5. Increment plan
|
|
|
|
Ordered so each step is independently mergeable and the risky one lands first.
|
|
|
|
1. **CB-501a — `CallerResolver` + `Role`.** Pure refactor: route today's connection identity through
|
|
the new type, `ANONYMOUS` not yet reachable (loopback-trust default preserves behaviour).
|
|
Green build, no behaviour change.
|
|
2. **CB-501b — token mode + fail-fast bind check.** Config block, bearer parsing, the
|
|
non-loopback-bind guard. This is the security-relevant commit; keep it small and reviewable.
|
|
3. **CB-505 — authz table + audit logger.** Enforce §3 on **both** entry paths (see §3.1); add the
|
|
audit appender.
|
|
4. **CB-502 — `Metrics` + `/metrics`.** Instrument the paths in §4.
|
|
5. **CB-503 — CI.** Runs `mvn -B clean install` with `-Dgroups='!contract'` so the live-herdr and
|
|
RabbitMQ contract tests are excluded; the mock-UDS suite is the CI surface, exactly as the
|
|
roadmap's testability section intends.
|
|
6. **CB-504 — launchd plist + systemd unit + herdr-socket retry.**
|
|
|
|
---
|
|
|
|
## 6. Open questions for the lead
|
|
|
|
*All four resolved 2026-07-29 — the lead confirmed D3, D4, and the D6 sink; the CI runner question
|
|
was answered from the forge itself. Kept here as the decision record.*
|
|
|
|
1. ✅ **TLS scope (D3) — confirmed.** Bearer auth + the fail-fast bind guard ship in the daemon;
|
|
TLS terminates at a reverse proxy, documented with a worked example. AMQP gets TLS via an
|
|
`amqps://` URI. No keystore handling in `fleetd`.
|
|
2. ✅ **Micrometer (D4) — confirmed dropped.** Zero-dependency Prometheus text renderer, for the
|
|
reasons in D4 (pom reconciliation burden + the mandated CVE gate being un-runnable this
|
|
session). Revisit if a push-gateway or JVM-metrics requirement appears; the endpoint is the
|
|
swap seam.
|
|
3. ~~**CI runner (CB-503).**~~ ✅ **Resolved during design** — a Gitea Actions runner *is*
|
|
registered and healthy (`lms/alms-memory` has 28 completed runs; `lms/alms` runs on push and
|
|
pull_request). CB-503 targets `.gitea/workflows/ci.yml` with `runs-on: ubuntu-latest`, matching
|
|
the sibling repo's convention. Note the runner's image ships an older `default-jdk`, so the
|
|
workflow must provision **JDK 25** explicitly rather than apt-installing the default.
|
|
Contract-test exclusion needs no CI flag: the pom's `default-excludes` profile already sets
|
|
`excludedGroups=contract`, so a plain `mvn -B clean install` *is* the mock-socket surface.
|
|
4. ✅ **Audit sink — confirmed dedicated file.** Its own logback appender writing JSON lines beside
|
|
the daemon, separate from the app log, per D6. Content still never enters the record.
|