From bec9fdffd8f8ee853b2d6d45803511bd4c04dffc Mon Sep 17 00:00:00 2001 From: Dai Ha Date: Sun, 16 Aug 2026 17:58:02 +0200 Subject: [PATCH] CB-595: backfill the ~14 operator-facing tickets shipped since v1.0.0 Nine new entries, written from a read of the code on main rather than from commit messages. The three keys that changed MEANING lead the batch, because that is what an upgrading operator meets first and none of it is visible in a diff of their own config: weight: 0 and maxLoad: 0 both used to do the opposite of what they read like, and fleet.leaders.*.terminal is now refused at startup rather than ignored. Then: the unconditional maxLoad cap on explicit spawns, the opt-in idle-lead heartbeat and why it must stay opt-in, charter receipts, the quarantine-aware capacity view, session resume and why it demands an explicit profile, opencode's --auto and forced auto-compaction with the trade stated plainly, the silent member-failure logging, and the removal of recycle(). Each entry carries the why, which is the half that stops a decision being re-litigated from scratch a month later. --- 11-Features.md | 202 ++++++++++++++++++++++++++++++++++++++++++++++++- 1 file changed, 200 insertions(+), 2 deletions(-) diff --git a/11-Features.md b/11-Features.md index fc3bac3..820f2f9 100644 --- a/11-Features.md +++ b/11-Features.md @@ -1279,11 +1279,209 @@ RabbitMQ client library is used only because LavinMQ speaks the same protocol. --- +## Three config keys that changed meaning — read this before upgrading + +**What.** `weight`, `maxLoad` and `fleet.leaders.*` all changed what they *mean*, not just what they +do. A config file that worked before an upgrade can behave differently, or refuse to start, with no +edit to it. The three are grouped here because an upgrading operator meets them together. + +| Key | Used to mean | Now means | +|---|---|---| +| `profiles..weight: 0` | coerced to 1.0 — "pick me as often as anyone else" | excluded from automatic selection | +| `profiles..maxLoad: 0` | coerced to unlimited | capped at zero live members | +| `fleet.leaders..terminal` | a terminal-id pin | **rejected at startup** — use `tab` | + +**On.** Nothing to switch on. These are the meanings now. + +**Why.** In the first two cases the old behaviour was the exact opposite of what the key reads like. +`weight: 0` looked like "never pick this" and meant "pick it normally". `maxLoad: 0` looked like +"never run anything here" and meant "unlimited" — and that was the only throttle a +`subscription: true` profile had against the operator's own paid plan. A key whose plain reading +inverts its behaviour is a trap, so both were fixed toward the reading. A negative `maxLoad` is now a +startup error rather than being silently normalised away. + +`fleet.leaders.*.terminal` went further and is **refused**, naming the offending entries, instead of +being ignored. A silently-ignored lead pin means the lead is not recognised, gets demoted to worker, +and every orchestration call is refused — a failure that looks nothing like its cause. Failing at +startup is the kinder outcome. + +**Gotcha.** `weight: 0` excludes a profile from *automatic* selection only. An explicit +`bridge_spawn{profile:"..."}` bypasses placement entirely and still resolves it, so `weight: 0` is +**not** a way to disable a profile. `maxLoad: 0` is, because that cap now applies to explicit spawns +too. And a lead's `tab` must match a real herdr tab label exactly (case-insensitively) — the old +`tabPrefix` no longer finds a lead, it only guards against a worker's label colliding with the +convention. + +--- + +## Stop an explicit spawn from busting the cap + +**What.** `maxLoad` is now an unconditional cap. An explicit `bridge_spawn{profile:"X"}` used to skip +the check entirely — the cap applied only to automatic placement — so naming a profile was a way +around it. Now an explicit spawn is refused when `live >= cap`, and it is **not** re-routed to +another profile. + +**On.** `profiles..maxLoad`. + +**Why.** A cap that any caller can opt out of by naming the profile is not a cap. The no-fallback part +is deliberate: a caller who named a profile did so for a cost or model reason, so quietly moving the +work to a different backend would defeat the reason they named it. Better to refuse and let the +caller decide. + +**Gotcha.** There is a known TOCTOU race: `liveCount` is read outside a lock, so two genuinely +concurrent spawns can both pass the check and briefly exceed the cap. It is documented in the code +rather than fixed. Also note the refusal message only became visible to callers in CB-599 — before +that it was a blank 500 with the reason in the log. + +--- + +## Nudge an idle lead back to work + +**What.** An opt-in loop that nudges the one idle lead after it has been continuously injectable for +a quiet period with nothing driving it. It fills the gap the reply push loop leaves: that loop only +fires when a worker reply lands, so a lead that is simply sitting idle with nothing arriving is +invisible to it. + +**On.** The `leadHeartbeat:` block — `idleAfterSeconds` (default 300), `backoffMs` (default 60000), +`quietNudgeCap` (default 3). **Absent means off**, and the loop is not even constructed. Changing it +needs a restart. + +**Why.** Off by default, deliberately. This loop spends the operator's model subscription on the +daemon's own initiative, so it must never switch itself on during an upgrade. That constraint shaped +the whole design. + +**Gotcha.** It is status-gated (a WORKING lead is never touched), debounced, caps consecutive quiet +nudges, and stands aside while the reply push loop is already nudging. That last one depends on +`ReplyPushLoop.isActive()` being honest — see CB-598, where a schedule could die with work still +pending and report inactive. + +--- + +## Prove which charter a member got, without logging the prose + +**What.** Per-role launch charters are read from the **live** config and composed once per spawn: the +role charter plus the reply charter. Claude Code gets it as an appended system prompt; opencode gets +an ephemeral `member-charter.md` referenced from its generated config. A `CharterReceipt` records a +sha256 and byte count of the exact composed bytes, and rides on both the spawn log and the roster. + +**On.** `fleet.charters.`, where role is `dev`, `architect` or `reviewer`. Read live per spawn, +so a change applies to the **next spawn with no restart**. + +**Why.** Two problems at once. Charters had been hardcoded, which meant changing what a role is told +required a rebuild. And an operator asking "what was this member actually told?" had no answer that +did not involve printing the prose into a log, where it does not belong. The receipt answers the +question with a digest instead. + +**Gotcha.** The role charter is combined with the reply charter **only when the profile mounts the +bridge MCP**. A non-MCP profile gets the role charter alone — which is correct, since the reply +charter tells a member to call `bridge_reply`, and a member with no bridge cannot. + +--- + +## Tell "busy" from "refusing" in the capacity view + +**What.** In `bridge_list`'s `capacity` view, a quarantined profile's row is forced to `free: 0` +whatever its `maxLoad` and `live` counts say, and gains `credentialId` and `quarantinedForSeconds`. + +**On.** No new knob. The quarantine facts come from the existing `exhaustedPattern`, `credentialId` +and `quarantineCooldownSeconds` keys. + +**Why.** A quarantined profile used to show free slots it would refuse to fill. A lead reading that +row would keep trying and keep failing. The two states need different reactions — "busy, will free +up" means wait, "refusing for N seconds" means go elsewhere — and the row could not express the +difference. + +**Gotcha.** The two new fields appear **only** when a profile is actually quarantined, so an ordinary +fleet's rows are byte-identical to before. Do not write a client that expects them. The view shares +its `QuarantineSource` with `bridge_profiles` so the two surfaces cannot disagree. + +--- + +## Resume a member onto its previous conversation + +**What.** A member's own agent-session id is captured at spawn, survives onto the roster as +`agentSessionId` in `bridge_list` and `GET /members`, and `bridge_spawn` accepts `sessionName` and +`resumeSessionId` to relaunch onto that same conversation. + +**On.** `bridge_spawn{sessionName, resumeSessionId}`. + +**Why.** The pieces existed but were connected at neither end: nothing persisted the id, and nothing +exposed a way to pass it back. So a member that died took its context with it even though the backend +could have resumed it. + +**Gotcha.** A `resumeSessionId` **requires an explicit `profile`**, and that profile's adapter must +declare `Capability.SESSION_RESUME`. A placement-routed spawn carrying a `resumeSessionId` is +refused, because a resumed conversation is tied to the specific backend that started it. If the +adapter lacks the capability the spawn is refused naming it, rather than silently starting a cold +session that looks resumed. + +--- + +## opencode members run with `--auto` and forced auto-compaction + +**What.** opencode workers launch with `--auto`, which auto-approves the permissions opencode does +not explicitly deny, and the bridge-generated config pins `compaction.auto: true`. + +**On.** Neither is configurable. Both are unconditional for the `opencode` kind. + +**Why.** A worker that stops to ask for permission on a tool call has no one to ask — the lead is not +watching its pane, and an unanswered prompt burns the whole turn. Auto-compaction is the same +argument: a worker that runs out of context dies mid-turn and loses its `bridge_reply`, so its report +is gone even though the work was done. + +**Gotcha, and it is a real trade.** opencode's own help calls `--auto` "dangerous!". The blast radius +is bounded by everything else about a member — its own worktree, its own branch, off-subscription, +and it cannot merge — not by the flag. And `compaction.auto: true` in the generated config **wins +over the operator's home config**, so auto-compaction cannot be turned off for bridged workers from +there. That was chosen knowingly: a lost report was judged worse than an unwanted compaction. + +--- + +## Silent member failures now say why + +**What.** Logging only, no behaviour change. The paths that used to fail a member with no log, a bare +DEBUG, or a message naming only the symptom now log at WARN with the real cause and the numbers — +across the injector, the completion resolver, session acquire/reap/failure, and ticket abandonment. + +**On.** Automatic. + +**Why.** A member could fail and leave nothing to read. The lead saw a session that stopped +responding and had no way to tell a crashed backend from a wedged pane from a reaped session. + +**Gotcha.** You will see more WARN lines than before if you filter by level. That is the feature, not +noise — but it will change what a log-volume alert sees. + +--- + +## Sessions are one-shot — `recycle()` is gone + +**What.** `SessionManager.recycle()` — release plus immediate re-acquire under a new pane id — was +removed. A released or finished session is torn down, never pooled or reused. Callers must `release` +then `acquire` explicitly. + +**On.** Nothing to configure; this is a removal. + +**Why.** Recycling quietly reused a session across two unrelated pieces of work. The lifecycle is +meant to be strictly one-shot, and a method that bypassed it was a standing invitation to leak one +delegation's context into the next. + +**Gotcha.** None left in-tree — no references remain. It had no external surface, so only internal +callers were ever affected. + +--- + ## Backfill status This page was started after the fact, so it is **not yet complete**. Entries above are written from -verified behaviour. Still to catalogue — each needs its config surface and gotcha confirmed against -the code before it earns an entry: +verified behaviour. + +The CB-595 backfill (2026-08-16) cleared the largest gap: the ~14 operator-facing tickets shipped +between `v1.0.0` and now that had landed nowhere. Those entries were written from a read of the code +on `main`, not from commit messages, and the three keys that changed *meaning* were pulled to the +front because an upgrading operator meets them first. + +Still to catalogue — each needs its config surface and gotcha confirmed against the code before it +earns an entry: - `/metrics` and `/healthz`, and what each does *not* tell you (CB-502; `/healthz` reports herdr reachability only — it went green while every spawn failed, see the CB-521 version-coupling note)