diff --git a/11-Features.md b/11-Features.md index c1fb32b..fc3bac3 100644 --- a/11-Features.md +++ b/11-Features.md @@ -1144,6 +1144,12 @@ collected, so a ticket you already read is never nudged about again. (default 5) and `push_backoff_ms` (default 15000). A lead that is not a herdr pane leaves the registry empty, the loop becomes a no-op, and delivery degrades to pull — nothing is lost. +Since CB-590 there is **one nudge schedule per lead**, shared with the CB-307 reply nudge, so the two +can no longer inject into the same pane at once. `push_reminders` is a budget **per source**, not one +shared counter: reply work and ticket work each get their own, so a busy reply stream cannot spend the +budget a ticket needs. Worst case a lead sees up to twice `push_reminders` nudges, which is the +deliberate price of that isolation. + **Why.** The charter tells leads to prefer `wait:false` for anything non-trivial, because a blocking `bridge_send` is capped by the caller's own MCP client timeout of about 60 seconds. But until this landed, that preferred mode was the one mode with **no notification at all**: `MessageService.reply` @@ -1186,6 +1192,93 @@ difference beforehand. **Present-and-useless looks identical to absent.** --- +## Supervise the daemon without breaking the fleet + +**What.** A launchd unit that restarts `bridged` if it dies, and that still gets the fleet's secrets. +`deploy/dev.ltms.bridged.plist` runs `scripts/bridged-launchd-wrapper.sh`, which execs one login shell +in place (`exec /bin/zsh -lc 'exec "$@"' -- "$@"`) and then execs the real java command. One `exec` +chain, so launchd keeps tracking the right PID. `scripts/redeploy-bridged.sh` detects whether the +agent is loaded and switches stop/start to `launchctl unload -w` / `load -w`, falling back to its +original kill + `nohup` when it is not. + +**On.** Not automatic, and deliberately so. Install it yourself: + +```bash +cp deploy/dev.ltms.bridged.plist ~/Library/LaunchAgents/ +launchctl load -w ~/Library/LaunchAgents/dev.ltms.bridged.plist +launchctl list | grep bridged +``` + +`scripts/redeploy-bridged.sh --check` reports whether the agent is installed and whether it is loaded. +It is read-only. + +**Why.** The unit shipped with CB-504 was never installed, and could not have worked if it were. +launchd does not source a login shell, so a launchd-started daemon would have had no +`WORKER_GITEA_TOKEN` and no `AI_GATEWAY_TOKEN`. It would have started fine and looked healthy; the +failure would have appeared hours later as workers unable to open a PR. So the operator's real choice +was a supervised daemon with a broken fleet, or a working fleet with no supervision. The wrapper +removes that choice. The `launchctl` branch in the redeploy script matters just as much: a `SIGTERM`ed +daemon exits **143** even when its shutdown hook completes normally, so under +`KeepAlive{SuccessfulExit: false}` a bare `kill` makes launchd restart the **old** jar, racing the +script's own restart. + +**Gotcha.** The script computes its log path from where the script file sits; the plist hard-codes an +absolute `StandardOutPath`. Nothing checks that the two agree. If they ever diverge — a worktree, a +renamed clone — the script's post-restart ERROR check reads the wrong file, finds nothing, and reports +"ok" while the daemon crash-loops. And the loop really is unbounded: `ThrottleInterval: 10` paces +restarts to one per ten seconds, it does not cap how many. Fix **CB-600** before installing the agent. + +--- + +## See at startup which secrets the daemon actually got + +**What.** `bridged` logs, at startup, every secret environment variable it needs, and whether each one +resolved or is `MISSING`. The required set is derived from the loaded config — each non-subscription +profile's `tokenEnv`, plus every profile's `gitTokenEnv` — not hard-coded, so a new profile is covered +the day it is added. + +**On.** Automatic. Read the `startup secret …` lines at the top of `bridged/bridged.out`. + +**Why.** An empty token used to be completely invisible. The daemon started, `/healthz` went green, +and the first sign of trouble came much later and somewhere else — a worker that could not open a PR, +or a gateway profile that could not authenticate. Neither symptom points back at the shell the daemon +was started from, which is the actual cause. This turns a silent, delayed, misattributed failure into +one line at startup. + +**Gotcha.** It reports **names and set/MISSING only** — never a value, a prefix, or a length. That is +deliberate and must stay that way; the log is not a secret store. Also, `MISSING` is a warning, not a +refusal: the daemon starts anyway, because refusing to boot over a credential that half the fleet may +not need would be worse. So the line has to actually be read. One known false alarm: a non-subscription +profile that never sets `tokenEnv` inherits the default name `BRIDGED_WORKER_TOKEN` and is reported +missing — which is nearly always a real misconfiguration rather than a bug in the report. + +--- + +## Bound the AMQP backlog, and know a reply was really published + +**What.** Two guarantees on the durable reply inbox. The consumer calls `basicQos` before +`basicConsume`, so unacked messages beyond the window stay **on the queue** instead of being pushed +into the daemon's heap. And publishing runs on its own confirm-mode channel with `mandatory=true` and +a return listener, so an unroutable or unconfirmed publish raises an error instead of vanishing. + +**On.** `broker.prefetch` sets the window, default 32. The confirm behaviour is automatic whenever +`broker:` is configured at all. + +**Why.** Both existed to make the word "durable" true. Without prefetch the queue sat near-empty while +the real backlog lived in an in-memory map with nothing capping it — so queue-depth metrics read +healthy, and any queue-level limit would have guarded an empty queue. Without confirms, a publish to a +queue that was never declared was a silent black hole, and `deliveryMode(2)` bought nothing, because +"persisted" is only true after the broker says so. + +**Gotcha.** A broker `Return` always arrives **before** its matching `Confirm`, so an ack alone does +not mean routed — the confirm path has to check a per-message returned flag, and that ordering is easy +to get wrong when editing this code. Note also that these tests are `@Tag("contract")`: they are +excluded from the default build and run in a **separate CI job** against a real `rabbitmq:3.13` +container. A green default build says nothing about them. And the broker here is **LavinMQ**; the +RabbitMQ client library is used only because LavinMQ speaks the same protocol. + +--- + ## Backfill status This page was started after the fact, so it is **not yet complete**. Entries above are written from @@ -1196,6 +1289,9 @@ the code before it earns an entry: reachability only — it went green while every spawn failed, see the CB-521 version-coupling note) - bearer-token auth and the non-loopback-bind fail-fast (CB-501) - the per-session authz table and the audit log (CB-505) -- launchd / systemd supervision (CB-504) +- systemd supervision (`deploy/bridged.service`) — the launchd half is catalogued above (CB-504, + CB-594); the Linux unit has not been checked for the same login-shell secret defect, and the + gateways it targets are CB-308 work - multi-profile routing and `kind:` adapter selection (CB-305, CB-401/402) -- the reply push loop and its nudge budget (CB-307 step 2) +- the reply push loop and its nudge budget (CB-307 step 2, CB-590) — note the budget is now **per + source**, a reply budget and a ticket budget, not one shared counter