CB-594/CB-527/CB-528/CB-590: catalogue supervision, the startup secret report, and the AMQP guarantees
Four entries' worth of shipped, operator-facing behaviour that had no home: - launchd supervision that actually carries the fleet's secrets, plus the redeploy script's launchctl branch and why a bare kill was wrong (exit 143) - the startup secret report, and why it warns rather than refuses to boot - broker.prefetch and publisher confirms, with the Return-before-Confirm ordering trap and the fact that these tests only run in a separate CI job Also corrected two entries the merges invalidated: the async-ticket nudge now shares one schedule per lead with a budget per source, and the systemd half of CB-504 is still unchecked for the same login-shell defect the launchd half had.
+98
-2
@@ -1144,6 +1144,12 @@ collected, so a ticket you already read is never nudged about again.
|
||||
(default 5) and `push_backoff_ms` (default 15000). A lead that is not a herdr pane leaves the
|
||||
registry empty, the loop becomes a no-op, and delivery degrades to pull — nothing is lost.
|
||||
|
||||
Since CB-590 there is **one nudge schedule per lead**, shared with the CB-307 reply nudge, so the two
|
||||
can no longer inject into the same pane at once. `push_reminders` is a budget **per source**, not one
|
||||
shared counter: reply work and ticket work each get their own, so a busy reply stream cannot spend the
|
||||
budget a ticket needs. Worst case a lead sees up to twice `push_reminders` nudges, which is the
|
||||
deliberate price of that isolation.
|
||||
|
||||
**Why.** The charter tells leads to prefer `wait:false` for anything non-trivial, because a blocking
|
||||
`bridge_send` is capped by the caller's own MCP client timeout of about 60 seconds. But until this
|
||||
landed, that preferred mode was the one mode with **no notification at all**: `MessageService.reply`
|
||||
@@ -1186,6 +1192,93 @@ difference beforehand. **Present-and-useless looks identical to absent.**
|
||||
|
||||
---
|
||||
|
||||
## Supervise the daemon without breaking the fleet
|
||||
|
||||
**What.** A launchd unit that restarts `bridged` if it dies, and that still gets the fleet's secrets.
|
||||
`deploy/dev.ltms.bridged.plist` runs `scripts/bridged-launchd-wrapper.sh`, which execs one login shell
|
||||
in place (`exec /bin/zsh -lc 'exec "$@"' -- "$@"`) and then execs the real java command. One `exec`
|
||||
chain, so launchd keeps tracking the right PID. `scripts/redeploy-bridged.sh` detects whether the
|
||||
agent is loaded and switches stop/start to `launchctl unload -w` / `load -w`, falling back to its
|
||||
original kill + `nohup` when it is not.
|
||||
|
||||
**On.** Not automatic, and deliberately so. Install it yourself:
|
||||
|
||||
```bash
|
||||
cp deploy/dev.ltms.bridged.plist ~/Library/LaunchAgents/
|
||||
launchctl load -w ~/Library/LaunchAgents/dev.ltms.bridged.plist
|
||||
launchctl list | grep bridged
|
||||
```
|
||||
|
||||
`scripts/redeploy-bridged.sh --check` reports whether the agent is installed and whether it is loaded.
|
||||
It is read-only.
|
||||
|
||||
**Why.** The unit shipped with CB-504 was never installed, and could not have worked if it were.
|
||||
launchd does not source a login shell, so a launchd-started daemon would have had no
|
||||
`WORKER_GITEA_TOKEN` and no `AI_GATEWAY_TOKEN`. It would have started fine and looked healthy; the
|
||||
failure would have appeared hours later as workers unable to open a PR. So the operator's real choice
|
||||
was a supervised daemon with a broken fleet, or a working fleet with no supervision. The wrapper
|
||||
removes that choice. The `launchctl` branch in the redeploy script matters just as much: a `SIGTERM`ed
|
||||
daemon exits **143** even when its shutdown hook completes normally, so under
|
||||
`KeepAlive{SuccessfulExit: false}` a bare `kill` makes launchd restart the **old** jar, racing the
|
||||
script's own restart.
|
||||
|
||||
**Gotcha.** The script computes its log path from where the script file sits; the plist hard-codes an
|
||||
absolute `StandardOutPath`. Nothing checks that the two agree. If they ever diverge — a worktree, a
|
||||
renamed clone — the script's post-restart ERROR check reads the wrong file, finds nothing, and reports
|
||||
"ok" while the daemon crash-loops. And the loop really is unbounded: `ThrottleInterval: 10` paces
|
||||
restarts to one per ten seconds, it does not cap how many. Fix **CB-600** before installing the agent.
|
||||
|
||||
---
|
||||
|
||||
## See at startup which secrets the daemon actually got
|
||||
|
||||
**What.** `bridged` logs, at startup, every secret environment variable it needs, and whether each one
|
||||
resolved or is `MISSING`. The required set is derived from the loaded config — each non-subscription
|
||||
profile's `tokenEnv`, plus every profile's `gitTokenEnv` — not hard-coded, so a new profile is covered
|
||||
the day it is added.
|
||||
|
||||
**On.** Automatic. Read the `startup secret …` lines at the top of `bridged/bridged.out`.
|
||||
|
||||
**Why.** An empty token used to be completely invisible. The daemon started, `/healthz` went green,
|
||||
and the first sign of trouble came much later and somewhere else — a worker that could not open a PR,
|
||||
or a gateway profile that could not authenticate. Neither symptom points back at the shell the daemon
|
||||
was started from, which is the actual cause. This turns a silent, delayed, misattributed failure into
|
||||
one line at startup.
|
||||
|
||||
**Gotcha.** It reports **names and set/MISSING only** — never a value, a prefix, or a length. That is
|
||||
deliberate and must stay that way; the log is not a secret store. Also, `MISSING` is a warning, not a
|
||||
refusal: the daemon starts anyway, because refusing to boot over a credential that half the fleet may
|
||||
not need would be worse. So the line has to actually be read. One known false alarm: a non-subscription
|
||||
profile that never sets `tokenEnv` inherits the default name `BRIDGED_WORKER_TOKEN` and is reported
|
||||
missing — which is nearly always a real misconfiguration rather than a bug in the report.
|
||||
|
||||
---
|
||||
|
||||
## Bound the AMQP backlog, and know a reply was really published
|
||||
|
||||
**What.** Two guarantees on the durable reply inbox. The consumer calls `basicQos` before
|
||||
`basicConsume`, so unacked messages beyond the window stay **on the queue** instead of being pushed
|
||||
into the daemon's heap. And publishing runs on its own confirm-mode channel with `mandatory=true` and
|
||||
a return listener, so an unroutable or unconfirmed publish raises an error instead of vanishing.
|
||||
|
||||
**On.** `broker.prefetch` sets the window, default 32. The confirm behaviour is automatic whenever
|
||||
`broker:` is configured at all.
|
||||
|
||||
**Why.** Both existed to make the word "durable" true. Without prefetch the queue sat near-empty while
|
||||
the real backlog lived in an in-memory map with nothing capping it — so queue-depth metrics read
|
||||
healthy, and any queue-level limit would have guarded an empty queue. Without confirms, a publish to a
|
||||
queue that was never declared was a silent black hole, and `deliveryMode(2)` bought nothing, because
|
||||
"persisted" is only true after the broker says so.
|
||||
|
||||
**Gotcha.** A broker `Return` always arrives **before** its matching `Confirm`, so an ack alone does
|
||||
not mean routed — the confirm path has to check a per-message returned flag, and that ordering is easy
|
||||
to get wrong when editing this code. Note also that these tests are `@Tag("contract")`: they are
|
||||
excluded from the default build and run in a **separate CI job** against a real `rabbitmq:3.13`
|
||||
container. A green default build says nothing about them. And the broker here is **LavinMQ**; the
|
||||
RabbitMQ client library is used only because LavinMQ speaks the same protocol.
|
||||
|
||||
---
|
||||
|
||||
## Backfill status
|
||||
|
||||
This page was started after the fact, so it is **not yet complete**. Entries above are written from
|
||||
@@ -1196,6 +1289,9 @@ the code before it earns an entry:
|
||||
reachability only — it went green while every spawn failed, see the CB-521 version-coupling note)
|
||||
- bearer-token auth and the non-loopback-bind fail-fast (CB-501)
|
||||
- the per-session authz table and the audit log (CB-505)
|
||||
- launchd / systemd supervision (CB-504)
|
||||
- systemd supervision (`deploy/bridged.service`) — the launchd half is catalogued above (CB-504,
|
||||
CB-594); the Linux unit has not been checked for the same login-shell secret defect, and the
|
||||
gateways it targets are CB-308 work
|
||||
- multi-profile routing and `kind:` adapter selection (CB-305, CB-401/402)
|
||||
- the reply push loop and its nudge budget (CB-307 step 2)
|
||||
- the reply push loop and its nudge budget (CB-307 step 2, CB-590) — note the budget is now **per
|
||||
source**, a reply budget and a ticket budget, not one shared counter
|
||||
|
||||
Reference in New Issue
Block a user