CB-595: the last five Features entries, from a read of the code

Adds /metrics and /healthz, bearer auth and the bind fail-fast, the
authz table and audit log, systemd supervision, and multi-profile kind:
routing. Every claim traced to a file:line on main.

Two defects turned up while writing them, both the same shape - config
accepted, does nothing useful, reports no error: kind: is never
validated, and the systemd unit has the login-shell secret defect that
launchd's wrapper already fixes.

Also records why a worker cannot read this page: the submodule pointer
is deliberately never updated, so a worker's checkout is months old.
Dai Ha
2026-08-16 18:32:08 +02:00
parent 9fe66176b0
commit 073f01088f
+205 -10
@@ -1553,6 +1553,177 @@ as *"everything else was checked and cleared"*.
---
## `/metrics` and `/healthz`
**What.** Two read-only diagnostic endpoints. `GET /healthz` calls herdr's `ping` and reports
only whether herdr answered: `200 {"status":"ok","herdr":{"version":…,"protocol":…}}`, or
`503 {"status":"degraded","herdr":"unreachable",…}` on any `HerdrException`. `GET /metrics`
renders every registered series as Prometheus text (CB-502): counters
`bridged_sends_total`, `bridged_replies_total`, `bridged_push_nudges_total`,
`bridged_lead_heartbeat_nudges_total`, `bridged_spawns_total`, `bridged_herdr_calls_total`,
`bridged_auth_failures_total`, plus gauges `bridged_sessions` (one series per lifecycle state)
and `bridged_inbox_depth` (one series per undrained target).
**On.** Both are always registered. `/healthz` carries no authorization check at all — it is
reachable by an unauthenticated caller by design. `/metrics` is gated on `Authz.Action.METRICS`:
open to `PRIMARY`/`WORKER`/`ARCHITECT`, refused to `ANONYMOUS`.
**Why.** `BridgedMetrics`'s own class doc states the design intent directly: "each series maps
to a failure mode this project has actually hit, not to whatever was easy to count," and names
the two worth watching — a rising `bridged_sends_total{outcome="completion_fallback"}` share
(turn detection degrading) and `bridged_push_nudges_total{outcome="exhausted"}` (the primary
stopped draining its inbox). `/healthz`'s narrow scope traces to CB-504: under supervision the
daemon must serve before herdr's socket even exists, so "degraded but alive" needed one cheap,
reliable signal.
**Gotcha.** Neither endpoint proves the fleet actually works. `/healthz` echoes back whatever
`protocol` number herdr reports, but nothing in the codebase compares that number against what
bridged's own herdr calls need — and the two have already drifted apart in the source itself:
`AgentControl`'s class doc says it was "ported to herdr protocol 19 (herdr 0.8.0, CB-521)," while
`HerdrClient`'s class doc still says "protocol 14, herdr 0.7.0." This is the exact CB-521
incident: herdr answers `ping` correctly and `/healthz` goes green, while `agent.start` and the
rest of the protocol-19 surface fail because the adapter and the herdr binary disagree on
protocol version. `/metrics` has no counter or gauge for that mismatch either — a resulting spawn
failure only shows up as a `bridged_spawns_total{outcome=…}` tick, and only once something
actually tries to spawn.
---
## Bearer-token auth and the non-loopback-bind fail-fast
**What.** `auth.mode` decides how a non-worker caller proves it is the primary:
`loopback-trust` (default) — any loopback caller that is not a known worker/lead/architect pane
is the primary, no credential needed; or `token` — such a caller must present
`Authorization: Bearer <token>` or resolves to `ANONYMOUS`. A worker/lead/architect pane is
always resolved from the unforgeable loopback-PID→herdr-pane mapping regardless of `auth.mode`.
Separately, the daemon refuses to start at all when `bind.host` is non-loopback and `auth.mode`
is still `loopback-trust`.
**On.** `auth.mode: loopback-trust | token` (default `loopback-trust`); `auth.tokenEnv` names the
host env var holding the token (default `BRIDGED_API_TOKEN`, read only under `token` mode). The
bind fail-fast has no separate switch — it always runs in `main()`.
**Why.** Stated directly in the code: loopback-trust's safety depends entirely on the OS refusing
non-local connections to a loopback socket. Widen the bind without switching to `token` mode and
"not a known worker" silently becomes "any client that can reach this port is the primary" — the
most privileged role on the bus (spawn/stop/send/drain on any session). Rather than document the
hazard, the config makes it unrepresentable: it throws instead of starting.
**Gotcha.** There are two separate fail-fast throws, both inline in `main()`, both before the
daemon binds its port — so a bad config never opens the socket at all. Under `launchd`, that
repeats forever: the plist's own comment warns launchd retries a fast-failing job every
`ThrottleInterval` (10s) with no give-up count, until a human unloads the agent or fixes the
cause. And `auth.tokenEnv` naming an unset/empty var throws a *different* message than the
bind-mismatch check — don't assume one error class covers both.
---
## The per-session authz table and the audit log
**What.** `Authz.permits(Principal caller, Action action, String targetSession)` is one static
table stating, for each of eight actions (`SPAWN`, `STOP`, `SEND`, `REPLY`, `ASK`, `DRAIN`,
`READ`, `METRICS`), which role may call it: `SPAWN`/`STOP`/`DRAIN` are the primary alone; `SEND`
is primary or architect; `REPLY`/`ASK` require the caller to own the target session (its own
pane, checked structurally, never by argument); `READ`/`METRICS` are open to any authenticated
role. It is the single gate behind both entry paths (REST and MCP) — `BridgedApp.allow()` and
`BridgeMcp`'s own check both call into it, so the rule can't drift between the two surfaces.
Refusals are recorded by `AuditLog`, an append-only JSON-lines trail written by a dedicated
`audit` logger to `logs/audit.log` (daily rolling, 30-day retention, 100MB cap), independent of
the daemon's normal app log.
**On.** Always on; not configurable. Every request through `BridgedApp` or `BridgeMcp` passes
through `Authz.permits()`.
**Why.** The class doc states this plainly: most of the rule was already true de facto — a
worker's identity comes from its connection, never an argument, so it could never reply as
another worker over MCP — but the REST surface used to trust the session id in the URL path
outright, and neither surface checked role at all. This makes the invariant explicit and
testable instead of emergent, and gives every privileged action one recorded outcome
(allowed/denied/failed) instead of none.
**Gotcha.** Message content is *never* written to the audit log by design — only
who/what/target/outcome/reason and a correlation id, because the bus carries user source code,
diffs, and prompts, and an audit trail that quietly accumulated those would be a transcript
archive wearing a security control's clothing. `READ` actions are deliberately excluded from the
*allowed* audit trail ("reads would drown the trail") — only denials of `READ`/`METRICS` are
recorded, not successes. And the 401-vs-403 split matters if you're debugging a refusal: 401
means "you presented no usable identity" (fixable by the caller), 403 means "you are
authenticated, but this isn't yours" — a worker reaching for another worker's session, or for
orchestration it was never granted.
---
## Multi-profile routing and `kind:` adapter selection
**What.** Each `workers:` profile carries a `kind:` field selecting which backend launcher spawns
it — `claude-code` (the default) or `opencode` (CB-402). At startup, `main()` partitions every
configured profile into two maps by `Profile.isOpenCode()`, builds one `ClaudeCodeLauncher` and/or
one `OpenCodeLauncher` accordingly, and wraps both in a `CompositePeerLauncher` that routes each
call to whichever adapter declares the profile the call names.
**On.** `kind: claude-code | opencode` on a profile; absent or blank defaults to `claude-code`.
The claude-code adapter is built even with zero claude-code profiles configured, unless opencode
is the *only* kind present — so a bridge with no `workers:` at all still has a well-defined base
adapter.
**Why.** Not stated as a single "why" comment beyond the CB-402 changelog note that opencode
"proves the `PeerLauncher` SPI is genuinely provider-neutral rather than Claude-shaped" (see the
existing *Pin an opencode endpoint* entry). The partition-by-kind design itself reads as the
natural consequence: profiles fully own their backend, so routing is a lookup, not a branch.
**Gotcha.** `kind:` is lower-cased but **never validated against the two known values.** A typo —
`kind: opencod`, say — is silently accepted, normalized, and (because it doesn't equal
`"opencode"`) routed into the **claude-code** adapter bucket. If `argv:` was also left unset, the
launch command defaults to `List.of(k)` — literally the misspelled string itself — rather than
`claude`, because the argv-defaulting logic only special-cases the exact string
`"claude-code"`. There is no config-load check anywhere that would catch this before spawn.
Separately, the composite constructor does refuse two adapters claiming the same profile name
("worker profile '…' is claimed by two peer adapters"), so that failure mode is caught loud.
---
## Supervise the daemon on Linux (systemd)
**What.** `deploy/bridged.service` is a systemd **user** unit (not system-level — "bridged drives
the user's herdr, not a system daemon") that runs `java -jar target/bridged.jar bridged.yaml`,
restarts on failure (`Restart=on-failure`, `RestartSec=10s`, capped at 5 restarts per 120s via
`StartLimitBurst`/`StartLimitIntervalSec`), waits on `herdr.service` only advisorially
(`Wants=`, not `Requires=`, so a herdr restart never takes bridged down with it), and applies a
sandboxing profile (`NoNewPrivileges`, `ProtectSystem=strict`, `ProtectHome=read-write`, etc.).
**On.** Manual install: copy to `~/.config/systemd/user/`, edit `ExecStart`/`WorkingDirectory`/
`Environment`, then `systemctl --user daemon-reload && systemctl --user enable --now bridged`. Not
currently the live supervision target — the unit file's own header comment says the dogfooded
daemon runs on macOS under launchd; this unit is for the Linux gateways CB-308 introduces.
**Why.** Not stated beyond the practical need: a per-host gateway topology (CB-308, noted
elsewhere in the wiki) needs Linux hosts, and those need systemd rather than launchd.
**Gotcha.** The unit hard-codes `Environment=PATH=…` with an explicit comment explaining why:
"systemd does not source a login shell, so without it the daemon — and every worker — gets a bare
default with no JDK/Maven" (same defect class CB-511/CB-594 already fixed for `PATH` on
launchd). For secrets, the unit's own comment says plainly "Secrets are NOT set here" and points
the operator at a `systemctl --user edit bridged` drop-in or an `EnvironmentFile=`. **It answers
the login-shell defect only for `PATH`, not for `WORKER_GITEA_TOKEN`/`AI_GATEWAY_TOKEN`.**
**Systemd answer — YES, it has the underlying defect, undocumented for those two variables
specifically.** Comparing to the launchd side: launchd had the identical problem
(`deploy/dev.ltms.bridged.plist`'s own comment: "launchd does NOT source .zprofile/.zshrc") and it
was fixed by `scripts/bridged-launchd-wrapper.sh`, which execs `zsh -l` so
`${SHARED_ENV}/tools/secrets.sh` gets sourced — its own header comment names exactly
`WORKER_GITEA_TOKEN` and `AI_GATEWAY_TOKEN` as the two secrets this closes the gap for. The
systemd unit has **no equivalent wrapper** and does not source that file at all. Its own comment
mentions only `BRIDGED_API_TOKEN` as the secret to add via drop-in — it never names
`WORKER_GITEA_TOKEN` or `AI_GATEWAY_TOKEN`. An operator following the unit file's own guidance
verbatim would set `BRIDGED_API_TOKEN` and stop there: the daemon boots fine, and the failure
surfaces only later as a worker that cannot open a PR or a gateway profile returning 401 — the
exact failure mode CB-594 fixed for launchd, left open here. `Bridged.reportRequiredSecrets`
(called at the top of `main()`) *does* log which secret env-var names resolved on either
platform — but that log line can only tell the truth about the names it names; it doesn't fix
the sourcing gap, and the unit file gives the operator no prompt to look for it.
---
## Backfill status
This page was started after the fact, so it is **not yet complete**. Entries above are written from
@@ -1563,19 +1734,43 @@ between `v1.0.0` and now that had landed nowhere. Those entries were written fro
on `main`, not from commit messages, and the three keys that changed *meaning* were pulled to the
front because an upgrading operator meets them first.
Still to catalogue — each needs its config surface and gotcha confirmed against the code before it
earns an entry:
The second CB-595 pass (also 2026-08-16) cleared the five areas that pass had left open: `/metrics`
and `/healthz`, bearer auth and the bind fail-fast, the authz table and audit log, systemd
supervision, and multi-profile `kind:` routing. Every claim in those five was traced to a `file:line`
on `main`, and two of them turned up defects the catalogue work was not looking for — see below.
Still to catalogue:
- `/metrics` and `/healthz`, and what each does *not* tell you (CB-502; `/healthz` reports herdr
reachability only — it went green while every spawn failed, see the CB-521 version-coupling note)
- bearer-token auth and the non-loopback-bind fail-fast (CB-501)
- the per-session authz table and the audit log (CB-505)
- systemd supervision (`deploy/bridged.service`) — the launchd half is catalogued above (CB-504,
CB-594); the Linux unit has not been checked for the same login-shell secret defect, and the
gateways it targets are CB-308 work
- multi-profile routing and `kind:` adapter selection (CB-305, CB-401/402)
- the reply push loop and its nudge budget (CB-307 step 2, CB-590, CB-598) — the budget is now
tracked **per pending item**, not per lead and not per source. A newly-arrived item keeps its
source eligible even when an older, still-undrained item has used up its own budget. The practical
consequence to write up: the cap bounds nudges *about one item*, so a lead with a steady arrival of
new work keeps being nudged — which is correct, but is not what the knob's name suggests
### Found while cataloguing, not by looking for bugs
Two of these are filed as their own tickets. They are recorded here because both are the same shape:
**a configuration that is accepted, does nothing useful, and reports no error.**
- **`kind:` is never validated.** A typo such as `kind: opencod` is lower-cased, accepted, and — not
matching `"opencode"` — routed to the **claude-code** adapter. If `argv:` is unset, the launch
command defaults to the misspelled string itself rather than `claude`. No config-load check catches
it. See *Multi-profile routing* above.
- **The systemd unit has the login-shell secret defect that launchd's had.** `deploy/bridged.service`
fixes `PATH` explicitly and names only `BRIDGED_API_TOKEN` as a secret to add. It never sources the
secret store, and never mentions `WORKER_GITEA_TOKEN` or `AI_GATEWAY_TOKEN`. An operator following
the unit's own guidance gets a daemon that boots cleanly and members that cannot open a PR. launchd
got `scripts/bridged-launchd-wrapper.sh` for exactly this; the Linux unit has no equivalent.
One more piece of drift worth knowing while reading these entries: `AgentControl`'s class doc says
herdr protocol 19 (herdr 0.8.0), while `HerdrClient`'s still says protocol 14 (0.7.0). Nothing
compares the protocol number herdr reports against what bridged actually needs — which is precisely
how `/healthz` once went green while every spawn failed.
### A note for anyone briefing a worker to read this page
**A worker cannot see the current version of this file.** `wiki/` is a submodule, and the parent
repo's recorded pointer is deliberately never updated (committing it is on the never-commit list). So
`git submodule update --init` in a worker's worktree checks out a **months-old** snapshot. During
this very backfill a worker reported that an entry "does not exist anywhere in the repo" when it had
been on the page for hours. Paste the relevant text into the brief instead of pointing at the file.