diff --git a/11-Features.md b/11-Features.md index 7c03310..2519607 100644 --- a/11-Features.md +++ b/11-Features.md @@ -48,6 +48,8 @@ six weeks, and the table alone will not carry it. | [Keep a worktree that still holds work](#keep-a-worktree-that-still-holds-work) | automatic | CB-576 | `session/GitWorktrees` | | [Fail a ticket when its member dies](#fail-a-ticket-when-its-member-dies) | `health.enabled: true` | CB-580 | `health/FleetHealthMonitor` | | [Durable reply inbox](#durable-reply-inbox) | `broker:` | CB-307 | `msg/AmqpReplyInbox` | +| [Broker password out of the config](#broker-password-out-of-the-config) | `broker.uriEnv:` | CB-635 | `config/FleetConfig` | +| [An unreachable broker does not stop the daemon](#an-unreachable-broker-does-not-stop-the-daemon) | (always on) | CB-635 | `Fleetd.selectReplyInbox` | | [Reject overlapping rendezvous](#reject-overlapping-rendezvous) | automatic | CB-548 | `msg/Rendezvous` | | [Pin an opencode endpoint](#pin-an-opencode-endpoint) | profile `baseUrl:` | CB-508 | `worker/OpenCodeLauncher` | | [Onboard a project with the plugin](#onboard-a-project-with-the-plugin) | `/plugin install claude-bridge` → `/claude-bridge:setup` | CB-527 | `plugin/` | @@ -356,6 +358,44 @@ real task's runtime. Without a durable inbox, a reply arriving after that window writing into someone else's broker. Also see the known hole: an async ticket that times out while the session is still BUSY currently discards the later completion rather than parking it. +## Broker password out of the config + +**What.** `broker.uriEnv:` names an environment variable that holds the AMQP URI, instead of writing +the URI into the config file. An AMQP URI carries `user:password@` inline, so the old `broker.uri:` +put a live password in clear text in `bridged.yaml`. + +**On.** `broker: { uriEnv: LAVINMQ_URI }`. The variable must be on the **daemon's own** environment, +so start the daemon from a login shell — `scripts/redeploy-bridged.sh --check` reports whether the +named variable resolves, without ever printing its value. + +**Why.** The same reason `auth.tokenEnv` and `Profile.tokenEnv` exist: a config file gets read, +copied and pasted into tickets far more often than a secret store does. This was the last credential +still living in clear text in the config. + +**Gotcha.** `uriEnv` **wins over `uri` whenever it is set**, and it does not fall back. If the named +variable is unset or blank the daemon treats the broker as unconfigured and uses the in-memory inbox +— it does not quietly drop back onto a stale `uri:` left in the file. That is deliberate: an operator +who moved to the secret store must never be silently returned to clear text. Both set ⇒ the log says +`broker.uri is ignored`. + +## An unreachable broker does not stop the daemon + +**What.** If the broker cannot be reached **at boot**, bridged logs a loud warning and starts on the +in-memory reply inbox for that process lifetime, instead of failing to start at all. + +**On.** Always on. There is no background retry: fix the broker and restart to get durability back. + +**Why.** `AmqpReplyInbox.open` throws `IllegalStateException`, and nothing caught it — so a broker +that was down took the whole daemon with it, and with the daemon went every lead, every member and +every pane. Losing durable replies is bad; losing the fleet because the *reply store* is down is far +worse. A broker that drops *after* startup already self-heals through the AMQP client's automatic +recovery, so only the boot path needed this. + +**Gotcha.** The fallback is **not** silent, but it is easy to miss in a busy startup log. What you +lose is real: replies become soft-state and a held report does not survive the next restart. Grep the +startup log for `reply inbox:` — it says which adapter won, every time. The warning names the failing +URI **with the credentials stripped**. + ## Reject overlapping rendezvous **What.** A second attempt to open a reply waiter for the same worker session fails atomically instead