CB-595: the last five Features entries, from a read of the code
Adds /metrics and /healthz, bearer auth and the bind fail-fast, the authz table and audit log, systemd supervision, and multi-profile kind: routing. Every claim traced to a file:line on main. Two defects turned up while writing them, both the same shape - config accepted, does nothing useful, reports no error: kind: is never validated, and the systemd unit has the login-shell secret defect that launchd's wrapper already fixes. Also records why a worker cannot read this page: the submodule pointer is deliberately never updated, so a worker's checkout is months old.
+205
-10
@@ -1553,6 +1553,177 @@ as *"everything else was checked and cleared"*.
|
||||
|
||||
---
|
||||
|
||||
## `/metrics` and `/healthz`
|
||||
|
||||
**What.** Two read-only diagnostic endpoints. `GET /healthz` calls herdr's `ping` and reports
|
||||
only whether herdr answered: `200 {"status":"ok","herdr":{"version":…,"protocol":…}}`, or
|
||||
`503 {"status":"degraded","herdr":"unreachable",…}` on any `HerdrException`. `GET /metrics`
|
||||
renders every registered series as Prometheus text (CB-502): counters
|
||||
`bridged_sends_total`, `bridged_replies_total`, `bridged_push_nudges_total`,
|
||||
`bridged_lead_heartbeat_nudges_total`, `bridged_spawns_total`, `bridged_herdr_calls_total`,
|
||||
`bridged_auth_failures_total`, plus gauges `bridged_sessions` (one series per lifecycle state)
|
||||
and `bridged_inbox_depth` (one series per undrained target).
|
||||
|
||||
**On.** Both are always registered. `/healthz` carries no authorization check at all — it is
|
||||
reachable by an unauthenticated caller by design. `/metrics` is gated on `Authz.Action.METRICS`:
|
||||
open to `PRIMARY`/`WORKER`/`ARCHITECT`, refused to `ANONYMOUS`.
|
||||
|
||||
**Why.** `BridgedMetrics`'s own class doc states the design intent directly: "each series maps
|
||||
to a failure mode this project has actually hit, not to whatever was easy to count," and names
|
||||
the two worth watching — a rising `bridged_sends_total{outcome="completion_fallback"}` share
|
||||
(turn detection degrading) and `bridged_push_nudges_total{outcome="exhausted"}` (the primary
|
||||
stopped draining its inbox). `/healthz`'s narrow scope traces to CB-504: under supervision the
|
||||
daemon must serve before herdr's socket even exists, so "degraded but alive" needed one cheap,
|
||||
reliable signal.
|
||||
|
||||
**Gotcha.** Neither endpoint proves the fleet actually works. `/healthz` echoes back whatever
|
||||
`protocol` number herdr reports, but nothing in the codebase compares that number against what
|
||||
bridged's own herdr calls need — and the two have already drifted apart in the source itself:
|
||||
`AgentControl`'s class doc says it was "ported to herdr protocol 19 (herdr 0.8.0, CB-521)," while
|
||||
`HerdrClient`'s class doc still says "protocol 14, herdr 0.7.0." This is the exact CB-521
|
||||
incident: herdr answers `ping` correctly and `/healthz` goes green, while `agent.start` and the
|
||||
rest of the protocol-19 surface fail because the adapter and the herdr binary disagree on
|
||||
protocol version. `/metrics` has no counter or gauge for that mismatch either — a resulting spawn
|
||||
failure only shows up as a `bridged_spawns_total{outcome=…}` tick, and only once something
|
||||
actually tries to spawn.
|
||||
|
||||
---
|
||||
|
||||
## Bearer-token auth and the non-loopback-bind fail-fast
|
||||
|
||||
**What.** `auth.mode` decides how a non-worker caller proves it is the primary:
|
||||
`loopback-trust` (default) — any loopback caller that is not a known worker/lead/architect pane
|
||||
is the primary, no credential needed; or `token` — such a caller must present
|
||||
`Authorization: Bearer <token>` or resolves to `ANONYMOUS`. A worker/lead/architect pane is
|
||||
always resolved from the unforgeable loopback-PID→herdr-pane mapping regardless of `auth.mode`.
|
||||
Separately, the daemon refuses to start at all when `bind.host` is non-loopback and `auth.mode`
|
||||
is still `loopback-trust`.
|
||||
|
||||
**On.** `auth.mode: loopback-trust | token` (default `loopback-trust`); `auth.tokenEnv` names the
|
||||
host env var holding the token (default `BRIDGED_API_TOKEN`, read only under `token` mode). The
|
||||
bind fail-fast has no separate switch — it always runs in `main()`.
|
||||
|
||||
**Why.** Stated directly in the code: loopback-trust's safety depends entirely on the OS refusing
|
||||
non-local connections to a loopback socket. Widen the bind without switching to `token` mode and
|
||||
"not a known worker" silently becomes "any client that can reach this port is the primary" — the
|
||||
most privileged role on the bus (spawn/stop/send/drain on any session). Rather than document the
|
||||
hazard, the config makes it unrepresentable: it throws instead of starting.
|
||||
|
||||
**Gotcha.** There are two separate fail-fast throws, both inline in `main()`, both before the
|
||||
daemon binds its port — so a bad config never opens the socket at all. Under `launchd`, that
|
||||
repeats forever: the plist's own comment warns launchd retries a fast-failing job every
|
||||
`ThrottleInterval` (10s) with no give-up count, until a human unloads the agent or fixes the
|
||||
cause. And `auth.tokenEnv` naming an unset/empty var throws a *different* message than the
|
||||
bind-mismatch check — don't assume one error class covers both.
|
||||
|
||||
---
|
||||
|
||||
## The per-session authz table and the audit log
|
||||
|
||||
**What.** `Authz.permits(Principal caller, Action action, String targetSession)` is one static
|
||||
table stating, for each of eight actions (`SPAWN`, `STOP`, `SEND`, `REPLY`, `ASK`, `DRAIN`,
|
||||
`READ`, `METRICS`), which role may call it: `SPAWN`/`STOP`/`DRAIN` are the primary alone; `SEND`
|
||||
is primary or architect; `REPLY`/`ASK` require the caller to own the target session (its own
|
||||
pane, checked structurally, never by argument); `READ`/`METRICS` are open to any authenticated
|
||||
role. It is the single gate behind both entry paths (REST and MCP) — `BridgedApp.allow()` and
|
||||
`BridgeMcp`'s own check both call into it, so the rule can't drift between the two surfaces.
|
||||
Refusals are recorded by `AuditLog`, an append-only JSON-lines trail written by a dedicated
|
||||
`audit` logger to `logs/audit.log` (daily rolling, 30-day retention, 100MB cap), independent of
|
||||
the daemon's normal app log.
|
||||
|
||||
**On.** Always on; not configurable. Every request through `BridgedApp` or `BridgeMcp` passes
|
||||
through `Authz.permits()`.
|
||||
|
||||
**Why.** The class doc states this plainly: most of the rule was already true de facto — a
|
||||
worker's identity comes from its connection, never an argument, so it could never reply as
|
||||
another worker over MCP — but the REST surface used to trust the session id in the URL path
|
||||
outright, and neither surface checked role at all. This makes the invariant explicit and
|
||||
testable instead of emergent, and gives every privileged action one recorded outcome
|
||||
(allowed/denied/failed) instead of none.
|
||||
|
||||
**Gotcha.** Message content is *never* written to the audit log by design — only
|
||||
who/what/target/outcome/reason and a correlation id, because the bus carries user source code,
|
||||
diffs, and prompts, and an audit trail that quietly accumulated those would be a transcript
|
||||
archive wearing a security control's clothing. `READ` actions are deliberately excluded from the
|
||||
*allowed* audit trail ("reads would drown the trail") — only denials of `READ`/`METRICS` are
|
||||
recorded, not successes. And the 401-vs-403 split matters if you're debugging a refusal: 401
|
||||
means "you presented no usable identity" (fixable by the caller), 403 means "you are
|
||||
authenticated, but this isn't yours" — a worker reaching for another worker's session, or for
|
||||
orchestration it was never granted.
|
||||
|
||||
---
|
||||
|
||||
## Multi-profile routing and `kind:` adapter selection
|
||||
|
||||
**What.** Each `workers:` profile carries a `kind:` field selecting which backend launcher spawns
|
||||
it — `claude-code` (the default) or `opencode` (CB-402). At startup, `main()` partitions every
|
||||
configured profile into two maps by `Profile.isOpenCode()`, builds one `ClaudeCodeLauncher` and/or
|
||||
one `OpenCodeLauncher` accordingly, and wraps both in a `CompositePeerLauncher` that routes each
|
||||
call to whichever adapter declares the profile the call names.
|
||||
|
||||
**On.** `kind: claude-code | opencode` on a profile; absent or blank defaults to `claude-code`.
|
||||
The claude-code adapter is built even with zero claude-code profiles configured, unless opencode
|
||||
is the *only* kind present — so a bridge with no `workers:` at all still has a well-defined base
|
||||
adapter.
|
||||
|
||||
**Why.** Not stated as a single "why" comment beyond the CB-402 changelog note that opencode
|
||||
"proves the `PeerLauncher` SPI is genuinely provider-neutral rather than Claude-shaped" (see the
|
||||
existing *Pin an opencode endpoint* entry). The partition-by-kind design itself reads as the
|
||||
natural consequence: profiles fully own their backend, so routing is a lookup, not a branch.
|
||||
|
||||
**Gotcha.** `kind:` is lower-cased but **never validated against the two known values.** A typo —
|
||||
`kind: opencod`, say — is silently accepted, normalized, and (because it doesn't equal
|
||||
`"opencode"`) routed into the **claude-code** adapter bucket. If `argv:` was also left unset, the
|
||||
launch command defaults to `List.of(k)` — literally the misspelled string itself — rather than
|
||||
`claude`, because the argv-defaulting logic only special-cases the exact string
|
||||
`"claude-code"`. There is no config-load check anywhere that would catch this before spawn.
|
||||
Separately, the composite constructor does refuse two adapters claiming the same profile name
|
||||
("worker profile '…' is claimed by two peer adapters"), so that failure mode is caught loud.
|
||||
|
||||
---
|
||||
|
||||
## Supervise the daemon on Linux (systemd)
|
||||
|
||||
**What.** `deploy/bridged.service` is a systemd **user** unit (not system-level — "bridged drives
|
||||
the user's herdr, not a system daemon") that runs `java -jar target/bridged.jar bridged.yaml`,
|
||||
restarts on failure (`Restart=on-failure`, `RestartSec=10s`, capped at 5 restarts per 120s via
|
||||
`StartLimitBurst`/`StartLimitIntervalSec`), waits on `herdr.service` only advisorially
|
||||
(`Wants=`, not `Requires=`, so a herdr restart never takes bridged down with it), and applies a
|
||||
sandboxing profile (`NoNewPrivileges`, `ProtectSystem=strict`, `ProtectHome=read-write`, etc.).
|
||||
|
||||
**On.** Manual install: copy to `~/.config/systemd/user/`, edit `ExecStart`/`WorkingDirectory`/
|
||||
`Environment`, then `systemctl --user daemon-reload && systemctl --user enable --now bridged`. Not
|
||||
currently the live supervision target — the unit file's own header comment says the dogfooded
|
||||
daemon runs on macOS under launchd; this unit is for the Linux gateways CB-308 introduces.
|
||||
|
||||
**Why.** Not stated beyond the practical need: a per-host gateway topology (CB-308, noted
|
||||
elsewhere in the wiki) needs Linux hosts, and those need systemd rather than launchd.
|
||||
|
||||
**Gotcha.** The unit hard-codes `Environment=PATH=…` with an explicit comment explaining why:
|
||||
"systemd does not source a login shell, so without it the daemon — and every worker — gets a bare
|
||||
default with no JDK/Maven" (same defect class CB-511/CB-594 already fixed for `PATH` on
|
||||
launchd). For secrets, the unit's own comment says plainly "Secrets are NOT set here" and points
|
||||
the operator at a `systemctl --user edit bridged` drop-in or an `EnvironmentFile=`. **It answers
|
||||
the login-shell defect only for `PATH`, not for `WORKER_GITEA_TOKEN`/`AI_GATEWAY_TOKEN`.**
|
||||
|
||||
**Systemd answer — YES, it has the underlying defect, undocumented for those two variables
|
||||
specifically.** Comparing to the launchd side: launchd had the identical problem
|
||||
(`deploy/dev.ltms.bridged.plist`'s own comment: "launchd does NOT source .zprofile/.zshrc") and it
|
||||
was fixed by `scripts/bridged-launchd-wrapper.sh`, which execs `zsh -l` so
|
||||
`${SHARED_ENV}/tools/secrets.sh` gets sourced — its own header comment names exactly
|
||||
`WORKER_GITEA_TOKEN` and `AI_GATEWAY_TOKEN` as the two secrets this closes the gap for. The
|
||||
systemd unit has **no equivalent wrapper** and does not source that file at all. Its own comment
|
||||
mentions only `BRIDGED_API_TOKEN` as the secret to add via drop-in — it never names
|
||||
`WORKER_GITEA_TOKEN` or `AI_GATEWAY_TOKEN`. An operator following the unit file's own guidance
|
||||
verbatim would set `BRIDGED_API_TOKEN` and stop there: the daemon boots fine, and the failure
|
||||
surfaces only later as a worker that cannot open a PR or a gateway profile returning 401 — the
|
||||
exact failure mode CB-594 fixed for launchd, left open here. `Bridged.reportRequiredSecrets`
|
||||
(called at the top of `main()`) *does* log which secret env-var names resolved on either
|
||||
platform — but that log line can only tell the truth about the names it names; it doesn't fix
|
||||
the sourcing gap, and the unit file gives the operator no prompt to look for it.
|
||||
|
||||
---
|
||||
|
||||
## Backfill status
|
||||
|
||||
This page was started after the fact, so it is **not yet complete**. Entries above are written from
|
||||
@@ -1563,19 +1734,43 @@ between `v1.0.0` and now that had landed nowhere. Those entries were written fro
|
||||
on `main`, not from commit messages, and the three keys that changed *meaning* were pulled to the
|
||||
front because an upgrading operator meets them first.
|
||||
|
||||
Still to catalogue — each needs its config surface and gotcha confirmed against the code before it
|
||||
earns an entry:
|
||||
The second CB-595 pass (also 2026-08-16) cleared the five areas that pass had left open: `/metrics`
|
||||
and `/healthz`, bearer auth and the bind fail-fast, the authz table and audit log, systemd
|
||||
supervision, and multi-profile `kind:` routing. Every claim in those five was traced to a `file:line`
|
||||
on `main`, and two of them turned up defects the catalogue work was not looking for — see below.
|
||||
|
||||
Still to catalogue:
|
||||
|
||||
- `/metrics` and `/healthz`, and what each does *not* tell you (CB-502; `/healthz` reports herdr
|
||||
reachability only — it went green while every spawn failed, see the CB-521 version-coupling note)
|
||||
- bearer-token auth and the non-loopback-bind fail-fast (CB-501)
|
||||
- the per-session authz table and the audit log (CB-505)
|
||||
- systemd supervision (`deploy/bridged.service`) — the launchd half is catalogued above (CB-504,
|
||||
CB-594); the Linux unit has not been checked for the same login-shell secret defect, and the
|
||||
gateways it targets are CB-308 work
|
||||
- multi-profile routing and `kind:` adapter selection (CB-305, CB-401/402)
|
||||
- the reply push loop and its nudge budget (CB-307 step 2, CB-590, CB-598) — the budget is now
|
||||
tracked **per pending item**, not per lead and not per source. A newly-arrived item keeps its
|
||||
source eligible even when an older, still-undrained item has used up its own budget. The practical
|
||||
consequence to write up: the cap bounds nudges *about one item*, so a lead with a steady arrival of
|
||||
new work keeps being nudged — which is correct, but is not what the knob's name suggests
|
||||
|
||||
### Found while cataloguing, not by looking for bugs
|
||||
|
||||
Two of these are filed as their own tickets. They are recorded here because both are the same shape:
|
||||
**a configuration that is accepted, does nothing useful, and reports no error.**
|
||||
|
||||
- **`kind:` is never validated.** A typo such as `kind: opencod` is lower-cased, accepted, and — not
|
||||
matching `"opencode"` — routed to the **claude-code** adapter. If `argv:` is unset, the launch
|
||||
command defaults to the misspelled string itself rather than `claude`. No config-load check catches
|
||||
it. See *Multi-profile routing* above.
|
||||
- **The systemd unit has the login-shell secret defect that launchd's had.** `deploy/bridged.service`
|
||||
fixes `PATH` explicitly and names only `BRIDGED_API_TOKEN` as a secret to add. It never sources the
|
||||
secret store, and never mentions `WORKER_GITEA_TOKEN` or `AI_GATEWAY_TOKEN`. An operator following
|
||||
the unit's own guidance gets a daemon that boots cleanly and members that cannot open a PR. launchd
|
||||
got `scripts/bridged-launchd-wrapper.sh` for exactly this; the Linux unit has no equivalent.
|
||||
|
||||
One more piece of drift worth knowing while reading these entries: `AgentControl`'s class doc says
|
||||
herdr protocol 19 (herdr 0.8.0), while `HerdrClient`'s still says protocol 14 (0.7.0). Nothing
|
||||
compares the protocol number herdr reports against what bridged actually needs — which is precisely
|
||||
how `/healthz` once went green while every spawn failed.
|
||||
|
||||
### A note for anyone briefing a worker to read this page
|
||||
|
||||
**A worker cannot see the current version of this file.** `wiki/` is a submodule, and the parent
|
||||
repo's recorded pointer is deliberately never updated (committing it is on the never-commit list). So
|
||||
`git submodule update --init` in a worker's worktree checks out a **months-old** snapshot. During
|
||||
this very backfill a worker reported that an entry "does not exist anywhere in the repo" when it had
|
||||
been on the page for hours. Paste the relevant text into the brief instead of pointing at the file.
|
||||
|
||||
Reference in New Issue
Block a user