Two changes ship together here.
1. One shared herdr workspace. The lead and every worker now live in one
workspace called "fleet", so the operator sees one "session" with many
windows, not two. Before, the lead sat in a "leads" workspace and workers
in "bridged-workers", which read as two sessions. The lead is still told
apart from workers by its exact tab label ("lead: <name>"), so putting them
in one space is safe. LeadTabScanner keeps the exclude-by-label mechanism
for split layouts; Fleetd now passes an empty exclude set.
2. Rename the daemon from "bridged" to "fleetd" (the binary, config, scripts,
launchd/systemd units, module dir, and MCP mount).
- Module dir bridged/ -> fleetd/; jar finalName -> fleetd.jar.
- Log line, comments, docs, and CLAUDE.md updated to say fleetd.
- Scripts renamed: redeploy-bridged.sh -> redeploy-fleetd.sh,
bridged-launchd-wrapper.sh -> fleetd-launchd-wrapper.sh.
- Deploy units renamed: dev.ltms.bridged.plist -> dev.ltms.fleetd.plist,
bridged.service -> fleetd.service; launchd Label -> dev.ltms.fleetd.
- Config default bridged.yaml -> fleetd.yaml; the legacy bridged.yaml is
still read as a fallback, and still gitignored.
- MCP: drop the deprecated bridge_* tool twins; only fleet_* remain. The
server name is "fleet". The mount name in the local .mcp.json becomes
"fleet" (gitignored, not in this commit).
- Env var defaults BRIDGED_API_TOKEN -> FLEETD_API_TOKEN, fixture
BRIDGED_WORKER_TOKEN -> FLEETD_WORKER_TOKEN.
Kept on purpose: the BRIDGED_MEMBER marker. Renaming it is a coupled change to
the credential-scrub security control (an operator secrets.sh may guard on it),
so it stays until that migration is done on its own.
Metrics were already fleet_* (CB-632); MetricNamesTest still guards that no
name says bridged_.
The canonical CLAUDE.md block and the wiki template stay byte-identical
(wiki working tree edited, committed to the wiki repo separately).
949 tests pass (mvn clean install). 4 fewer than before = the 4 removed
bridge_* alias tests.
14 KiB
CB-5xx — Stage 5 Hardening (auth · metrics · CI · supervision · authz+audit)
Status: design note (pre-implementation) — the single-host close-out before cross-host work.
Covers: CB-501 (bearer auth + TLS) · CB-502 (/metrics) · CB-503 (mock-socket CI) ·
CB-504 (service supervision) · CB-505 (per-session authz + audit log).
Depends on: everything shipped through CB-402. Nothing here changes messaging semantics.
Blocks: CB-308. The cross-host trust model is CB-308's own gating concern, and it inherits
whatever identity/authz shape lands here — so this stage is deliberately before federation,
not after it.
1. Why this stage is not optional bookkeeping
fleetd today has exactly one security control: the loopback bind. Every other guarantee
rests on it.
The identity model (mcp/ConnectionIdentity.java) resolves a caller from the connection alone —
the OS reports the connecting PID, herdr owns the PID→pane map, so a worker cannot forge another
worker. Its own javadoc is explicit: "Single-host only (the herd shares the fleetd host); the
token path is the split-host fallback." The token path does not exist yet.
That leaves a seam that is latent today and load-bearing the moment the bind moves:
// ConnectionIdentity.resolve — non-loopback callers get a null terminal
if (!isLoopback(remoteAddr)) return new Caller(null, -1);
…and null terminal is interpreted downstream as "this caller is the primary". Combined:
Any caller that is not a recognised on-host worker pane is treated as the primary — including, if
bind.hostis ever widened, an arbitrary remote client.
Today bind defaults to 127.0.0.1 so this is unreachable. But CB-308 exists precisely to widen
the boundary, and the primary is the most privileged role on the bus (it spawns, stops, sends to
any session, and drains any inbox). Shipping federation on top of "unauthenticated ⇒ primary"
would be building the security boundary backwards.
So CB-501 is not "add a token header". It is: make identity explicit, and make the absence of identity mean nothing, not everything.
2. Decisions (locked)
D1 — Three caller roles, one resolution path
Introduce Role { PRIMARY, WORKER, ANONYMOUS } resolved by a single CallerResolver that both
REST and MCP go through. Resolution order:
- Connection identity wins where it applies. A loopback peer PID that maps to a herdr worker
pane ⇒
WORKERwith that terminal. Unforgeable, unchanged from today, zero config. - Token, if presented. A valid bearer token ⇒ the role that token is provisioned for.
- Otherwise
ANONYMOUS— notPRIMARY.
This inverts today's default. PRIMARY becomes something you must prove (by being a loopback
non-worker process when auth is disabled, or by presenting a primary-scoped token when it is
enabled), rather than something you get by failing every other check.
D2 — Auth is opt-in by config, but the default must stay zero-friction on loopback
The daemon is dogfooded constantly on one machine. If enabling hardening breaks the local setup, it will be disabled and the stage is wasted. So:
auth:
mode: loopback-trust # default — behaves exactly like today: loopback ⇒ PRIMARY, no token needed
# mode: token # every non-worker caller must present a valid bearer token
# tokenEnv: FLEETD_API_TOKEN # host env var holding the token; never the literal value
mode: loopback-trust is the current behaviour, named honestly and now chosen rather than
implied. mode: token is what a non-loopback bind requires. A non-loopback bind.host with
mode: loopback-trust must fail fast at startup — that check is the single highest-value line
in this stage, because it makes the dangerous configuration unrepresentable rather than merely
discouraged.
D3 — TLS terminates outside the JVM
Do not add TLS config to Javalin/Jetty. The deployment story for a cross-host gateway is a
reverse proxy (or an SSH/WireGuard tunnel) in front of the daemon; the AMQP link has its own TLS
via the broker URI (amqps://). Adding keystore handling here would mean certificate lifecycle
code in a daemon whose whole value is being small, and would duplicate what the proxy does better.
CB-501 therefore ships bearer auth + the fail-fast bind check, and documents TLS as a deployment concern with a worked reverse-proxy example. This is a deliberate narrowing of the roadmap's "auth/TLS" wording — flagged in §6 for the lead.
D4 — Metrics without a new dependency
The roadmap's tech-stack table says Micrometer→Prometheus. Recommend not taking that dep:
- This pom already carries an unusually heavy dependency-reconciliation burden (a hand-pinned
jackson-annotations3.0-rc5 to reconcile the MCP SDK's Jackson 3 with our Jackson 2.19, a Jetty BOM import to stop version skew, plus four documented accepted-CVE advisories). Every new transitive tree is a real cost here, not a hypothetical one. - The CVE gate that CLAUDE.md mandates for dependency changes (
jetbrains get_file_problems→ Mend.io) cannot currently be run — no JetBrains MCP server is connected. Adding a dependency tree we cannot scan violates the project's own stated policy. - The metric set is small and fully known (§4). Prometheus text exposition is a trivial, stable, well-specified format.
So: a ~120-line metrics/Metrics.java holding LongAdder counters and gauge suppliers, rendered
to the Prometheus text format at GET /metrics. If Micrometer is wanted later for its
registry/push ecosystem, this stays a drop-in swap behind the same endpoint. Flagged in §6 —
this deviates from a documented tech-stack choice.
D5 — Supervision targets launchd first, systemd second
The roadmap says "systemd unit". This host is macOS — there is no systemd on it (systemctl
not found), and the daemon that has been dogfooded for weeks runs as a bare foreground
java -jar. Ship both:
deploy/dev.ltms.fleet.plist— launchd agent, the actual runtime here, withKeepAliveand ordered start after herdr.deploy/fleetd.service— systemd unit for the Linux gateways CB-308 introduces.
Ordering after herdr is advisory in both: the herdr socket may not exist at boot, so the daemon must retry the socket rather than exit — supervision ordering is a nicety, socket-retry is the actual fix. That retry behaviour is part of CB-504, not a separate ticket.
D6 — Audit log is a separate append-only stream, not the app log
Privileged actions (spawn, stop, send, reply-drain, ack) emit a structured JSON line to a
dedicated audit SLF4J logger with its own appender, carrying {ts, role, terminal, pid, action, target, outcome}. Keeping it off the chatty app logger is what makes it greppable and, later,
shippable. No message content in the audit record — the bridge carries the user's source
code and prompts; an audit trail that quietly becomes a transcript archive is a liability, not a
control. Content stays out; correlation ids go in.
3. Authorization model (CB-505)
With D1's roles, the rules are small enough to state completely:
| Action | REST | PRIMARY | WORKER | ANONYMOUS |
|---|---|---|---|---|
| spawn worker | POST /workers |
✅ | ❌ | ❌ |
| stop worker | DELETE /workers/{paneId} |
✅ | ❌ | ❌ |
| send to a session | POST /sessions/{id}/message |
✅ | ❌ | ❌ |
| reply | POST /sessions/{id}/reply |
❌ | ✅ own session only | ❌ |
| ask | POST /sessions/{id}/ask |
❌ | ✅ own session only | ❌ |
| drain replies | GET /sessions/{id}/replies |
✅ | ❌ | ❌ |
| status / list / profiles | GET … |
✅ | ✅ | ❌ |
| health | GET /healthz |
✅ | ✅ | ✅ (unauthenticated by design) |
| metrics | GET /metrics |
✅ | ✅ | ❌ |
The load-bearing row is "own session only": a worker may only reply or ask as itself. That is
already true de facto — ConnectionIdentity derives the terminal rather than reading it from the
body — so CB-505 mostly asserts an existing invariant explicitly and adds the test that pins
it. The one real change is rejecting a worker that names a different session id in the path.
/healthz stays open: it must answer for a load balancer or supervisor before any credential is
configured. It already leaks nothing but herdr's version and up/down.
3.1 There are TWO entry paths, and only one of them has identity today
The wiki describes MCP as "a thin adapter over the REST core". At the code level that is not
literally true, and the difference is security-relevant. FleetMcp calls MessageService /
SessionManager directly; it never issues an HTTP request against a Javalin route. And /mcp is
mounted as a raw servlet on Jetty's ServletContextHandler
(FleetApp.build → cfg.jetty.modifyServletContextHandler), so it does not pass through
Javalin's before filters at all.
The current split is the mirror image of what you'd expect:
| Path | Caller identity today | Authz today |
|---|---|---|
MCP /mcp |
✅ resolved per call (ConnectionIdentity via the transport-context extractor) |
❌ none |
| REST routes | ❌ none at all — the session id is taken from the URL path and trusted | ❌ none |
So REST is the more exposed surface: POST /sessions/{id}/reply accepts any {id} from the
path, whereas the MCP fleet_reply derives the worker from the connection and refuses to read it
from an argument. Loopback-only bind is what makes this safe today.
Therefore CB-505 must enforce on both paths against one shared resolver — not at a single
choke point. Concretely: a Javalin before filter for REST, and the existing transport-context
extractor for MCP, both delegating to auth.CallerResolver. Any authz check that lives in only
one of the two is not a control.
4. Metric set (CB-502)
Deliberately small; every one maps to a failure mode we have actually hit.
| Metric | Type | Why it exists |
|---|---|---|
fleet_sends_total{outcome} |
counter | outcome ∈ replied|completion_fallback|timeout|failed — the completion-fallback rate is the health signal for turn detection (CB-115/116/118) |
fleet_send_duration_seconds |
histogram | delegated turn latency |
fleet_replies_total{path} |
counter | path ∈ rendezvous|inbox — how often a reply strands (CB-307's whole reason to exist) |
fleet_inbox_depth{target} |
gauge | undrained replies; steady-state should be 0 |
fleet_push_nudges_total{outcome} |
counter | outcome ∈ delivered|exhausted — a rising exhausted means the primary is not draining |
fleet_spawns_total{kind,outcome} |
counter | outcome ∈ ready|timeout|guard_rejected; per peer kind (CB-402) |
fleet_sessions{state} |
gauge | SPAWNING/READY/BUSY/DONE census |
fleet_herdr_calls_total{method,outcome} |
counter | socket health — the dependency everything rests on |
fleet_auth_failures_total{reason} |
counter | only meaningful once CB-501 lands; catches misconfigured workers |
5. Increment plan
Ordered so each step is independently mergeable and the risky one lands first.
- CB-501a —
CallerResolver+Role. Pure refactor: route today's connection identity through the new type,ANONYMOUSnot yet reachable (loopback-trust default preserves behaviour). Green build, no behaviour change. - CB-501b — token mode + fail-fast bind check. Config block, bearer parsing, the non-loopback-bind guard. This is the security-relevant commit; keep it small and reviewable.
- CB-505 — authz table + audit logger. Enforce §3 on both entry paths (see §3.1); add the audit appender.
- CB-502 —
Metrics+/metrics. Instrument the paths in §4. - CB-503 — CI. Runs
mvn -B clean installwith-Dgroups='!contract'so the live-herdr and RabbitMQ contract tests are excluded; the mock-UDS suite is the CI surface, exactly as the roadmap's testability section intends. - CB-504 — launchd plist + systemd unit + herdr-socket retry.
6. Open questions for the lead
All four resolved 2026-07-29 — the lead confirmed D3, D4, and the D6 sink; the CI runner question was answered from the forge itself. Kept here as the decision record.
- ✅ TLS scope (D3) — confirmed. Bearer auth + the fail-fast bind guard ship in the daemon;
TLS terminates at a reverse proxy, documented with a worked example. AMQP gets TLS via an
amqps://URI. No keystore handling infleetd. - ✅ Micrometer (D4) — confirmed dropped. Zero-dependency Prometheus text renderer, for the reasons in D4 (pom reconciliation burden + the mandated CVE gate being un-runnable this session). Revisit if a push-gateway or JVM-metrics requirement appears; the endpoint is the swap seam.
CI runner (CB-503).✅ Resolved during design — a Gitea Actions runner is registered and healthy (lms/alms-memoryhas 28 completed runs;lms/almsruns on push and pull_request). CB-503 targets.gitea/workflows/ci.ymlwithruns-on: ubuntu-latest, matching the sibling repo's convention. Note the runner's image ships an olderdefault-jdk, so the workflow must provision JDK 25 explicitly rather than apt-installing the default. Contract-test exclusion needs no CI flag: the pom'sdefault-excludesprofile already setsexcludedGroups=contract, so a plainmvn -B clean installis the mock-socket surface.- ✅ Audit sink — confirmed dedicated file. Its own logback appender writing JSON lines beside the daemon, separate from the app log, per D6. Content still never enters the record.