diff --git a/11-Features.md b/11-Features.md index aa4b564..5850a48 100644 --- a/11-Features.md +++ b/11-Features.md @@ -1553,6 +1553,177 @@ as *"everything else was checked and cleared"*. --- +## `/metrics` and `/healthz` + +**What.** Two read-only diagnostic endpoints. `GET /healthz` calls herdr's `ping` and reports +only whether herdr answered: `200 {"status":"ok","herdr":{"version":…,"protocol":…}}`, or +`503 {"status":"degraded","herdr":"unreachable",…}` on any `HerdrException`. `GET /metrics` +renders every registered series as Prometheus text (CB-502): counters +`bridged_sends_total`, `bridged_replies_total`, `bridged_push_nudges_total`, +`bridged_lead_heartbeat_nudges_total`, `bridged_spawns_total`, `bridged_herdr_calls_total`, +`bridged_auth_failures_total`, plus gauges `bridged_sessions` (one series per lifecycle state) +and `bridged_inbox_depth` (one series per undrained target). + +**On.** Both are always registered. `/healthz` carries no authorization check at all — it is +reachable by an unauthenticated caller by design. `/metrics` is gated on `Authz.Action.METRICS`: +open to `PRIMARY`/`WORKER`/`ARCHITECT`, refused to `ANONYMOUS`. + +**Why.** `BridgedMetrics`'s own class doc states the design intent directly: "each series maps +to a failure mode this project has actually hit, not to whatever was easy to count," and names +the two worth watching — a rising `bridged_sends_total{outcome="completion_fallback"}` share +(turn detection degrading) and `bridged_push_nudges_total{outcome="exhausted"}` (the primary +stopped draining its inbox). `/healthz`'s narrow scope traces to CB-504: under supervision the +daemon must serve before herdr's socket even exists, so "degraded but alive" needed one cheap, +reliable signal. + +**Gotcha.** Neither endpoint proves the fleet actually works. `/healthz` echoes back whatever +`protocol` number herdr reports, but nothing in the codebase compares that number against what +bridged's own herdr calls need — and the two have already drifted apart in the source itself: +`AgentControl`'s class doc says it was "ported to herdr protocol 19 (herdr 0.8.0, CB-521)," while +`HerdrClient`'s class doc still says "protocol 14, herdr 0.7.0." This is the exact CB-521 +incident: herdr answers `ping` correctly and `/healthz` goes green, while `agent.start` and the +rest of the protocol-19 surface fail because the adapter and the herdr binary disagree on +protocol version. `/metrics` has no counter or gauge for that mismatch either — a resulting spawn +failure only shows up as a `bridged_spawns_total{outcome=…}` tick, and only once something +actually tries to spawn. + +--- + +## Bearer-token auth and the non-loopback-bind fail-fast + +**What.** `auth.mode` decides how a non-worker caller proves it is the primary: +`loopback-trust` (default) — any loopback caller that is not a known worker/lead/architect pane +is the primary, no credential needed; or `token` — such a caller must present +`Authorization: Bearer ` or resolves to `ANONYMOUS`. A worker/lead/architect pane is +always resolved from the unforgeable loopback-PID→herdr-pane mapping regardless of `auth.mode`. +Separately, the daemon refuses to start at all when `bind.host` is non-loopback and `auth.mode` +is still `loopback-trust`. + +**On.** `auth.mode: loopback-trust | token` (default `loopback-trust`); `auth.tokenEnv` names the +host env var holding the token (default `BRIDGED_API_TOKEN`, read only under `token` mode). The +bind fail-fast has no separate switch — it always runs in `main()`. + +**Why.** Stated directly in the code: loopback-trust's safety depends entirely on the OS refusing +non-local connections to a loopback socket. Widen the bind without switching to `token` mode and +"not a known worker" silently becomes "any client that can reach this port is the primary" — the +most privileged role on the bus (spawn/stop/send/drain on any session). Rather than document the +hazard, the config makes it unrepresentable: it throws instead of starting. + +**Gotcha.** There are two separate fail-fast throws, both inline in `main()`, both before the +daemon binds its port — so a bad config never opens the socket at all. Under `launchd`, that +repeats forever: the plist's own comment warns launchd retries a fast-failing job every +`ThrottleInterval` (10s) with no give-up count, until a human unloads the agent or fixes the +cause. And `auth.tokenEnv` naming an unset/empty var throws a *different* message than the +bind-mismatch check — don't assume one error class covers both. + +--- + +## The per-session authz table and the audit log + +**What.** `Authz.permits(Principal caller, Action action, String targetSession)` is one static +table stating, for each of eight actions (`SPAWN`, `STOP`, `SEND`, `REPLY`, `ASK`, `DRAIN`, +`READ`, `METRICS`), which role may call it: `SPAWN`/`STOP`/`DRAIN` are the primary alone; `SEND` +is primary or architect; `REPLY`/`ASK` require the caller to own the target session (its own +pane, checked structurally, never by argument); `READ`/`METRICS` are open to any authenticated +role. It is the single gate behind both entry paths (REST and MCP) — `BridgedApp.allow()` and +`BridgeMcp`'s own check both call into it, so the rule can't drift between the two surfaces. +Refusals are recorded by `AuditLog`, an append-only JSON-lines trail written by a dedicated +`audit` logger to `logs/audit.log` (daily rolling, 30-day retention, 100MB cap), independent of +the daemon's normal app log. + +**On.** Always on; not configurable. Every request through `BridgedApp` or `BridgeMcp` passes +through `Authz.permits()`. + +**Why.** The class doc states this plainly: most of the rule was already true de facto — a +worker's identity comes from its connection, never an argument, so it could never reply as +another worker over MCP — but the REST surface used to trust the session id in the URL path +outright, and neither surface checked role at all. This makes the invariant explicit and +testable instead of emergent, and gives every privileged action one recorded outcome +(allowed/denied/failed) instead of none. + +**Gotcha.** Message content is *never* written to the audit log by design — only +who/what/target/outcome/reason and a correlation id, because the bus carries user source code, +diffs, and prompts, and an audit trail that quietly accumulated those would be a transcript +archive wearing a security control's clothing. `READ` actions are deliberately excluded from the +*allowed* audit trail ("reads would drown the trail") — only denials of `READ`/`METRICS` are +recorded, not successes. And the 401-vs-403 split matters if you're debugging a refusal: 401 +means "you presented no usable identity" (fixable by the caller), 403 means "you are +authenticated, but this isn't yours" — a worker reaching for another worker's session, or for +orchestration it was never granted. + +--- + +## Multi-profile routing and `kind:` adapter selection + +**What.** Each `workers:` profile carries a `kind:` field selecting which backend launcher spawns +it — `claude-code` (the default) or `opencode` (CB-402). At startup, `main()` partitions every +configured profile into two maps by `Profile.isOpenCode()`, builds one `ClaudeCodeLauncher` and/or +one `OpenCodeLauncher` accordingly, and wraps both in a `CompositePeerLauncher` that routes each +call to whichever adapter declares the profile the call names. + +**On.** `kind: claude-code | opencode` on a profile; absent or blank defaults to `claude-code`. +The claude-code adapter is built even with zero claude-code profiles configured, unless opencode +is the *only* kind present — so a bridge with no `workers:` at all still has a well-defined base +adapter. + +**Why.** Not stated as a single "why" comment beyond the CB-402 changelog note that opencode +"proves the `PeerLauncher` SPI is genuinely provider-neutral rather than Claude-shaped" (see the +existing *Pin an opencode endpoint* entry). The partition-by-kind design itself reads as the +natural consequence: profiles fully own their backend, so routing is a lookup, not a branch. + +**Gotcha.** `kind:` is lower-cased but **never validated against the two known values.** A typo — +`kind: opencod`, say — is silently accepted, normalized, and (because it doesn't equal +`"opencode"`) routed into the **claude-code** adapter bucket. If `argv:` was also left unset, the +launch command defaults to `List.of(k)` — literally the misspelled string itself — rather than +`claude`, because the argv-defaulting logic only special-cases the exact string +`"claude-code"`. There is no config-load check anywhere that would catch this before spawn. +Separately, the composite constructor does refuse two adapters claiming the same profile name +("worker profile '…' is claimed by two peer adapters"), so that failure mode is caught loud. + +--- + +## Supervise the daemon on Linux (systemd) + +**What.** `deploy/bridged.service` is a systemd **user** unit (not system-level — "bridged drives +the user's herdr, not a system daemon") that runs `java -jar target/bridged.jar bridged.yaml`, +restarts on failure (`Restart=on-failure`, `RestartSec=10s`, capped at 5 restarts per 120s via +`StartLimitBurst`/`StartLimitIntervalSec`), waits on `herdr.service` only advisorially +(`Wants=`, not `Requires=`, so a herdr restart never takes bridged down with it), and applies a +sandboxing profile (`NoNewPrivileges`, `ProtectSystem=strict`, `ProtectHome=read-write`, etc.). + +**On.** Manual install: copy to `~/.config/systemd/user/`, edit `ExecStart`/`WorkingDirectory`/ +`Environment`, then `systemctl --user daemon-reload && systemctl --user enable --now bridged`. Not +currently the live supervision target — the unit file's own header comment says the dogfooded +daemon runs on macOS under launchd; this unit is for the Linux gateways CB-308 introduces. + +**Why.** Not stated beyond the practical need: a per-host gateway topology (CB-308, noted +elsewhere in the wiki) needs Linux hosts, and those need systemd rather than launchd. + +**Gotcha.** The unit hard-codes `Environment=PATH=…` with an explicit comment explaining why: +"systemd does not source a login shell, so without it the daemon — and every worker — gets a bare +default with no JDK/Maven" (same defect class CB-511/CB-594 already fixed for `PATH` on +launchd). For secrets, the unit's own comment says plainly "Secrets are NOT set here" and points +the operator at a `systemctl --user edit bridged` drop-in or an `EnvironmentFile=`. **It answers +the login-shell defect only for `PATH`, not for `WORKER_GITEA_TOKEN`/`AI_GATEWAY_TOKEN`.** + +**Systemd answer — YES, it has the underlying defect, undocumented for those two variables +specifically.** Comparing to the launchd side: launchd had the identical problem +(`deploy/dev.ltms.bridged.plist`'s own comment: "launchd does NOT source .zprofile/.zshrc") and it +was fixed by `scripts/bridged-launchd-wrapper.sh`, which execs `zsh -l` so +`${SHARED_ENV}/tools/secrets.sh` gets sourced — its own header comment names exactly +`WORKER_GITEA_TOKEN` and `AI_GATEWAY_TOKEN` as the two secrets this closes the gap for. The +systemd unit has **no equivalent wrapper** and does not source that file at all. Its own comment +mentions only `BRIDGED_API_TOKEN` as the secret to add via drop-in — it never names +`WORKER_GITEA_TOKEN` or `AI_GATEWAY_TOKEN`. An operator following the unit file's own guidance +verbatim would set `BRIDGED_API_TOKEN` and stop there: the daemon boots fine, and the failure +surfaces only later as a worker that cannot open a PR or a gateway profile returning 401 — the +exact failure mode CB-594 fixed for launchd, left open here. `Bridged.reportRequiredSecrets` +(called at the top of `main()`) *does* log which secret env-var names resolved on either +platform — but that log line can only tell the truth about the names it names; it doesn't fix +the sourcing gap, and the unit file gives the operator no prompt to look for it. + +--- + ## Backfill status This page was started after the fact, so it is **not yet complete**. Entries above are written from @@ -1563,19 +1734,43 @@ between `v1.0.0` and now that had landed nowhere. Those entries were written fro on `main`, not from commit messages, and the three keys that changed *meaning* were pulled to the front because an upgrading operator meets them first. -Still to catalogue — each needs its config surface and gotcha confirmed against the code before it -earns an entry: +The second CB-595 pass (also 2026-08-16) cleared the five areas that pass had left open: `/metrics` +and `/healthz`, bearer auth and the bind fail-fast, the authz table and audit log, systemd +supervision, and multi-profile `kind:` routing. Every claim in those five was traced to a `file:line` +on `main`, and two of them turned up defects the catalogue work was not looking for — see below. + +Still to catalogue: -- `/metrics` and `/healthz`, and what each does *not* tell you (CB-502; `/healthz` reports herdr - reachability only — it went green while every spawn failed, see the CB-521 version-coupling note) -- bearer-token auth and the non-loopback-bind fail-fast (CB-501) -- the per-session authz table and the audit log (CB-505) -- systemd supervision (`deploy/bridged.service`) — the launchd half is catalogued above (CB-504, - CB-594); the Linux unit has not been checked for the same login-shell secret defect, and the - gateways it targets are CB-308 work -- multi-profile routing and `kind:` adapter selection (CB-305, CB-401/402) - the reply push loop and its nudge budget (CB-307 step 2, CB-590, CB-598) — the budget is now tracked **per pending item**, not per lead and not per source. A newly-arrived item keeps its source eligible even when an older, still-undrained item has used up its own budget. The practical consequence to write up: the cap bounds nudges *about one item*, so a lead with a steady arrival of new work keeps being nudged — which is correct, but is not what the knob's name suggests + +### Found while cataloguing, not by looking for bugs + +Two of these are filed as their own tickets. They are recorded here because both are the same shape: +**a configuration that is accepted, does nothing useful, and reports no error.** + +- **`kind:` is never validated.** A typo such as `kind: opencod` is lower-cased, accepted, and — not + matching `"opencode"` — routed to the **claude-code** adapter. If `argv:` is unset, the launch + command defaults to the misspelled string itself rather than `claude`. No config-load check catches + it. See *Multi-profile routing* above. +- **The systemd unit has the login-shell secret defect that launchd's had.** `deploy/bridged.service` + fixes `PATH` explicitly and names only `BRIDGED_API_TOKEN` as a secret to add. It never sources the + secret store, and never mentions `WORKER_GITEA_TOKEN` or `AI_GATEWAY_TOKEN`. An operator following + the unit's own guidance gets a daemon that boots cleanly and members that cannot open a PR. launchd + got `scripts/bridged-launchd-wrapper.sh` for exactly this; the Linux unit has no equivalent. + +One more piece of drift worth knowing while reading these entries: `AgentControl`'s class doc says +herdr protocol 19 (herdr 0.8.0), while `HerdrClient`'s still says protocol 14 (0.7.0). Nothing +compares the protocol number herdr reports against what bridged actually needs — which is precisely +how `/healthz` once went green while every spawn failed. + +### A note for anyone briefing a worker to read this page + +**A worker cannot see the current version of this file.** `wiki/` is a submodule, and the parent +repo's recorded pointer is deliberately never updated (committing it is on the never-commit list). So +`git submodule update --init` in a worker's worktree checks out a **months-old** snapshot. During +this very backfill a worker reported that an entry "does not exist anywhere in the repo" when it had +been on the page for hours. Paste the relevant text into the brief instead of pointing at the file.