diff --git a/11-Features.md b/11-Features.md index fc3bac3..820f2f9 100644 --- a/11-Features.md +++ b/11-Features.md @@ -1279,11 +1279,209 @@ RabbitMQ client library is used only because LavinMQ speaks the same protocol. --- +## Three config keys that changed meaning — read this before upgrading + +**What.** `weight`, `maxLoad` and `fleet.leaders.*` all changed what they *mean*, not just what they +do. A config file that worked before an upgrade can behave differently, or refuse to start, with no +edit to it. The three are grouped here because an upgrading operator meets them together. + +| Key | Used to mean | Now means | +|---|---|---| +| `profiles..weight: 0` | coerced to 1.0 — "pick me as often as anyone else" | excluded from automatic selection | +| `profiles..maxLoad: 0` | coerced to unlimited | capped at zero live members | +| `fleet.leaders..terminal` | a terminal-id pin | **rejected at startup** — use `tab` | + +**On.** Nothing to switch on. These are the meanings now. + +**Why.** In the first two cases the old behaviour was the exact opposite of what the key reads like. +`weight: 0` looked like "never pick this" and meant "pick it normally". `maxLoad: 0` looked like +"never run anything here" and meant "unlimited" — and that was the only throttle a +`subscription: true` profile had against the operator's own paid plan. A key whose plain reading +inverts its behaviour is a trap, so both were fixed toward the reading. A negative `maxLoad` is now a +startup error rather than being silently normalised away. + +`fleet.leaders.*.terminal` went further and is **refused**, naming the offending entries, instead of +being ignored. A silently-ignored lead pin means the lead is not recognised, gets demoted to worker, +and every orchestration call is refused — a failure that looks nothing like its cause. Failing at +startup is the kinder outcome. + +**Gotcha.** `weight: 0` excludes a profile from *automatic* selection only. An explicit +`bridge_spawn{profile:"..."}` bypasses placement entirely and still resolves it, so `weight: 0` is +**not** a way to disable a profile. `maxLoad: 0` is, because that cap now applies to explicit spawns +too. And a lead's `tab` must match a real herdr tab label exactly (case-insensitively) — the old +`tabPrefix` no longer finds a lead, it only guards against a worker's label colliding with the +convention. + +--- + +## Stop an explicit spawn from busting the cap + +**What.** `maxLoad` is now an unconditional cap. An explicit `bridge_spawn{profile:"X"}` used to skip +the check entirely — the cap applied only to automatic placement — so naming a profile was a way +around it. Now an explicit spawn is refused when `live >= cap`, and it is **not** re-routed to +another profile. + +**On.** `profiles..maxLoad`. + +**Why.** A cap that any caller can opt out of by naming the profile is not a cap. The no-fallback part +is deliberate: a caller who named a profile did so for a cost or model reason, so quietly moving the +work to a different backend would defeat the reason they named it. Better to refuse and let the +caller decide. + +**Gotcha.** There is a known TOCTOU race: `liveCount` is read outside a lock, so two genuinely +concurrent spawns can both pass the check and briefly exceed the cap. It is documented in the code +rather than fixed. Also note the refusal message only became visible to callers in CB-599 — before +that it was a blank 500 with the reason in the log. + +--- + +## Nudge an idle lead back to work + +**What.** An opt-in loop that nudges the one idle lead after it has been continuously injectable for +a quiet period with nothing driving it. It fills the gap the reply push loop leaves: that loop only +fires when a worker reply lands, so a lead that is simply sitting idle with nothing arriving is +invisible to it. + +**On.** The `leadHeartbeat:` block — `idleAfterSeconds` (default 300), `backoffMs` (default 60000), +`quietNudgeCap` (default 3). **Absent means off**, and the loop is not even constructed. Changing it +needs a restart. + +**Why.** Off by default, deliberately. This loop spends the operator's model subscription on the +daemon's own initiative, so it must never switch itself on during an upgrade. That constraint shaped +the whole design. + +**Gotcha.** It is status-gated (a WORKING lead is never touched), debounced, caps consecutive quiet +nudges, and stands aside while the reply push loop is already nudging. That last one depends on +`ReplyPushLoop.isActive()` being honest — see CB-598, where a schedule could die with work still +pending and report inactive. + +--- + +## Prove which charter a member got, without logging the prose + +**What.** Per-role launch charters are read from the **live** config and composed once per spawn: the +role charter plus the reply charter. Claude Code gets it as an appended system prompt; opencode gets +an ephemeral `member-charter.md` referenced from its generated config. A `CharterReceipt` records a +sha256 and byte count of the exact composed bytes, and rides on both the spawn log and the roster. + +**On.** `fleet.charters.`, where role is `dev`, `architect` or `reviewer`. Read live per spawn, +so a change applies to the **next spawn with no restart**. + +**Why.** Two problems at once. Charters had been hardcoded, which meant changing what a role is told +required a rebuild. And an operator asking "what was this member actually told?" had no answer that +did not involve printing the prose into a log, where it does not belong. The receipt answers the +question with a digest instead. + +**Gotcha.** The role charter is combined with the reply charter **only when the profile mounts the +bridge MCP**. A non-MCP profile gets the role charter alone — which is correct, since the reply +charter tells a member to call `bridge_reply`, and a member with no bridge cannot. + +--- + +## Tell "busy" from "refusing" in the capacity view + +**What.** In `bridge_list`'s `capacity` view, a quarantined profile's row is forced to `free: 0` +whatever its `maxLoad` and `live` counts say, and gains `credentialId` and `quarantinedForSeconds`. + +**On.** No new knob. The quarantine facts come from the existing `exhaustedPattern`, `credentialId` +and `quarantineCooldownSeconds` keys. + +**Why.** A quarantined profile used to show free slots it would refuse to fill. A lead reading that +row would keep trying and keep failing. The two states need different reactions — "busy, will free +up" means wait, "refusing for N seconds" means go elsewhere — and the row could not express the +difference. + +**Gotcha.** The two new fields appear **only** when a profile is actually quarantined, so an ordinary +fleet's rows are byte-identical to before. Do not write a client that expects them. The view shares +its `QuarantineSource` with `bridge_profiles` so the two surfaces cannot disagree. + +--- + +## Resume a member onto its previous conversation + +**What.** A member's own agent-session id is captured at spawn, survives onto the roster as +`agentSessionId` in `bridge_list` and `GET /members`, and `bridge_spawn` accepts `sessionName` and +`resumeSessionId` to relaunch onto that same conversation. + +**On.** `bridge_spawn{sessionName, resumeSessionId}`. + +**Why.** The pieces existed but were connected at neither end: nothing persisted the id, and nothing +exposed a way to pass it back. So a member that died took its context with it even though the backend +could have resumed it. + +**Gotcha.** A `resumeSessionId` **requires an explicit `profile`**, and that profile's adapter must +declare `Capability.SESSION_RESUME`. A placement-routed spawn carrying a `resumeSessionId` is +refused, because a resumed conversation is tied to the specific backend that started it. If the +adapter lacks the capability the spawn is refused naming it, rather than silently starting a cold +session that looks resumed. + +--- + +## opencode members run with `--auto` and forced auto-compaction + +**What.** opencode workers launch with `--auto`, which auto-approves the permissions opencode does +not explicitly deny, and the bridge-generated config pins `compaction.auto: true`. + +**On.** Neither is configurable. Both are unconditional for the `opencode` kind. + +**Why.** A worker that stops to ask for permission on a tool call has no one to ask — the lead is not +watching its pane, and an unanswered prompt burns the whole turn. Auto-compaction is the same +argument: a worker that runs out of context dies mid-turn and loses its `bridge_reply`, so its report +is gone even though the work was done. + +**Gotcha, and it is a real trade.** opencode's own help calls `--auto` "dangerous!". The blast radius +is bounded by everything else about a member — its own worktree, its own branch, off-subscription, +and it cannot merge — not by the flag. And `compaction.auto: true` in the generated config **wins +over the operator's home config**, so auto-compaction cannot be turned off for bridged workers from +there. That was chosen knowingly: a lost report was judged worse than an unwanted compaction. + +--- + +## Silent member failures now say why + +**What.** Logging only, no behaviour change. The paths that used to fail a member with no log, a bare +DEBUG, or a message naming only the symptom now log at WARN with the real cause and the numbers — +across the injector, the completion resolver, session acquire/reap/failure, and ticket abandonment. + +**On.** Automatic. + +**Why.** A member could fail and leave nothing to read. The lead saw a session that stopped +responding and had no way to tell a crashed backend from a wedged pane from a reaped session. + +**Gotcha.** You will see more WARN lines than before if you filter by level. That is the feature, not +noise — but it will change what a log-volume alert sees. + +--- + +## Sessions are one-shot — `recycle()` is gone + +**What.** `SessionManager.recycle()` — release plus immediate re-acquire under a new pane id — was +removed. A released or finished session is torn down, never pooled or reused. Callers must `release` +then `acquire` explicitly. + +**On.** Nothing to configure; this is a removal. + +**Why.** Recycling quietly reused a session across two unrelated pieces of work. The lifecycle is +meant to be strictly one-shot, and a method that bypassed it was a standing invitation to leak one +delegation's context into the next. + +**Gotcha.** None left in-tree — no references remain. It had no external surface, so only internal +callers were ever affected. + +--- + ## Backfill status This page was started after the fact, so it is **not yet complete**. Entries above are written from -verified behaviour. Still to catalogue — each needs its config surface and gotcha confirmed against -the code before it earns an entry: +verified behaviour. + +The CB-595 backfill (2026-08-16) cleared the largest gap: the ~14 operator-facing tickets shipped +between `v1.0.0` and now that had landed nowhere. Those entries were written from a read of the code +on `main`, not from commit messages, and the three keys that changed *meaning* were pulled to the +front because an upgrading operator meets them first. + +Still to catalogue — each needs its config surface and gotcha confirmed against the code before it +earns an entry: - `/metrics` and `/healthz`, and what each does *not* tell you (CB-502; `/healthz` reports herdr reachability only — it went green while every spawn failed, see the CB-521 version-coupling note)