CB-594/CB-527/CB-528/CB-590: catalogue supervision, the startup secret report, and the AMQP guarantees

Four entries' worth of shipped, operator-facing behaviour that had no home:

- launchd supervision that actually carries the fleet's secrets, plus the
  redeploy script's launchctl branch and why a bare kill was wrong (exit 143)
- the startup secret report, and why it warns rather than refuses to boot
- broker.prefetch and publisher confirms, with the Return-before-Confirm
  ordering trap and the fact that these tests only run in a separate CI job

Also corrected two entries the merges invalidated: the async-ticket nudge now
shares one schedule per lead with a budget per source, and the systemd half of
CB-504 is still unchecked for the same login-shell defect the launchd half had.
Dai Ha
2026-08-16 17:35:17 +02:00
parent ac5c10d0be
commit f7dd817b97
+98 -2
@@ -1144,6 +1144,12 @@ collected, so a ticket you already read is never nudged about again.
(default 5) and `push_backoff_ms` (default 15000). A lead that is not a herdr pane leaves the
registry empty, the loop becomes a no-op, and delivery degrades to pull — nothing is lost.
Since CB-590 there is **one nudge schedule per lead**, shared with the CB-307 reply nudge, so the two
can no longer inject into the same pane at once. `push_reminders` is a budget **per source**, not one
shared counter: reply work and ticket work each get their own, so a busy reply stream cannot spend the
budget a ticket needs. Worst case a lead sees up to twice `push_reminders` nudges, which is the
deliberate price of that isolation.
**Why.** The charter tells leads to prefer `wait:false` for anything non-trivial, because a blocking
`bridge_send` is capped by the caller's own MCP client timeout of about 60 seconds. But until this
landed, that preferred mode was the one mode with **no notification at all**: `MessageService.reply`
@@ -1186,6 +1192,93 @@ difference beforehand. **Present-and-useless looks identical to absent.**
---
## Supervise the daemon without breaking the fleet
**What.** A launchd unit that restarts `bridged` if it dies, and that still gets the fleet's secrets.
`deploy/dev.ltms.bridged.plist` runs `scripts/bridged-launchd-wrapper.sh`, which execs one login shell
in place (`exec /bin/zsh -lc 'exec "$@"' -- "$@"`) and then execs the real java command. One `exec`
chain, so launchd keeps tracking the right PID. `scripts/redeploy-bridged.sh` detects whether the
agent is loaded and switches stop/start to `launchctl unload -w` / `load -w`, falling back to its
original kill + `nohup` when it is not.
**On.** Not automatic, and deliberately so. Install it yourself:
```bash
cp deploy/dev.ltms.bridged.plist ~/Library/LaunchAgents/
launchctl load -w ~/Library/LaunchAgents/dev.ltms.bridged.plist
launchctl list | grep bridged
```
`scripts/redeploy-bridged.sh --check` reports whether the agent is installed and whether it is loaded.
It is read-only.
**Why.** The unit shipped with CB-504 was never installed, and could not have worked if it were.
launchd does not source a login shell, so a launchd-started daemon would have had no
`WORKER_GITEA_TOKEN` and no `AI_GATEWAY_TOKEN`. It would have started fine and looked healthy; the
failure would have appeared hours later as workers unable to open a PR. So the operator's real choice
was a supervised daemon with a broken fleet, or a working fleet with no supervision. The wrapper
removes that choice. The `launchctl` branch in the redeploy script matters just as much: a `SIGTERM`ed
daemon exits **143** even when its shutdown hook completes normally, so under
`KeepAlive{SuccessfulExit: false}` a bare `kill` makes launchd restart the **old** jar, racing the
script's own restart.
**Gotcha.** The script computes its log path from where the script file sits; the plist hard-codes an
absolute `StandardOutPath`. Nothing checks that the two agree. If they ever diverge — a worktree, a
renamed clone — the script's post-restart ERROR check reads the wrong file, finds nothing, and reports
"ok" while the daemon crash-loops. And the loop really is unbounded: `ThrottleInterval: 10` paces
restarts to one per ten seconds, it does not cap how many. Fix **CB-600** before installing the agent.
---
## See at startup which secrets the daemon actually got
**What.** `bridged` logs, at startup, every secret environment variable it needs, and whether each one
resolved or is `MISSING`. The required set is derived from the loaded config — each non-subscription
profile's `tokenEnv`, plus every profile's `gitTokenEnv` — not hard-coded, so a new profile is covered
the day it is added.
**On.** Automatic. Read the `startup secret …` lines at the top of `bridged/bridged.out`.
**Why.** An empty token used to be completely invisible. The daemon started, `/healthz` went green,
and the first sign of trouble came much later and somewhere else — a worker that could not open a PR,
or a gateway profile that could not authenticate. Neither symptom points back at the shell the daemon
was started from, which is the actual cause. This turns a silent, delayed, misattributed failure into
one line at startup.
**Gotcha.** It reports **names and set/MISSING only** — never a value, a prefix, or a length. That is
deliberate and must stay that way; the log is not a secret store. Also, `MISSING` is a warning, not a
refusal: the daemon starts anyway, because refusing to boot over a credential that half the fleet may
not need would be worse. So the line has to actually be read. One known false alarm: a non-subscription
profile that never sets `tokenEnv` inherits the default name `BRIDGED_WORKER_TOKEN` and is reported
missing — which is nearly always a real misconfiguration rather than a bug in the report.
---
## Bound the AMQP backlog, and know a reply was really published
**What.** Two guarantees on the durable reply inbox. The consumer calls `basicQos` before
`basicConsume`, so unacked messages beyond the window stay **on the queue** instead of being pushed
into the daemon's heap. And publishing runs on its own confirm-mode channel with `mandatory=true` and
a return listener, so an unroutable or unconfirmed publish raises an error instead of vanishing.
**On.** `broker.prefetch` sets the window, default 32. The confirm behaviour is automatic whenever
`broker:` is configured at all.
**Why.** Both existed to make the word "durable" true. Without prefetch the queue sat near-empty while
the real backlog lived in an in-memory map with nothing capping it — so queue-depth metrics read
healthy, and any queue-level limit would have guarded an empty queue. Without confirms, a publish to a
queue that was never declared was a silent black hole, and `deliveryMode(2)` bought nothing, because
"persisted" is only true after the broker says so.
**Gotcha.** A broker `Return` always arrives **before** its matching `Confirm`, so an ack alone does
not mean routed — the confirm path has to check a per-message returned flag, and that ordering is easy
to get wrong when editing this code. Note also that these tests are `@Tag("contract")`: they are
excluded from the default build and run in a **separate CI job** against a real `rabbitmq:3.13`
container. A green default build says nothing about them. And the broker here is **LavinMQ**; the
RabbitMQ client library is used only because LavinMQ speaks the same protocol.
---
## Backfill status
This page was started after the fact, so it is **not yet complete**. Entries above are written from
@@ -1196,6 +1289,9 @@ the code before it earns an entry:
reachability only — it went green while every spawn failed, see the CB-521 version-coupling note)
- bearer-token auth and the non-loopback-bind fail-fast (CB-501)
- the per-session authz table and the audit log (CB-505)
- launchd / systemd supervision (CB-504)
- systemd supervision (`deploy/bridged.service`) — the launchd half is catalogued above (CB-504,
CB-594); the Linux unit has not been checked for the same login-shell secret defect, and the
gateways it targets are CB-308 work
- multi-profile routing and `kind:` adapter selection (CB-305, CB-401/402)
- the reply push loop and its nudge budget (CB-307 step 2)
- the reply push loop and its nudge budget (CB-307 step 2, CB-590) — note the budget is now **per
source**, a reply budget and a ticket budget, not one shared counter