CB-595: backfill the ~14 operator-facing tickets shipped since v1.0.0

Nine new entries, written from a read of the code on main rather than from
commit messages.

The three keys that changed MEANING lead the batch, because that is what an
upgrading operator meets first and none of it is visible in a diff of their own
config: weight: 0 and maxLoad: 0 both used to do the opposite of what they read
like, and fleet.leaders.*.terminal is now refused at startup rather than
ignored.

Then: the unconditional maxLoad cap on explicit spawns, the opt-in idle-lead
heartbeat and why it must stay opt-in, charter receipts, the quarantine-aware
capacity view, session resume and why it demands an explicit profile, opencode's
--auto and forced auto-compaction with the trade stated plainly, the silent
member-failure logging, and the removal of recycle().

Each entry carries the why, which is the half that stops a decision being
re-litigated from scratch a month later.
Dai Ha
2026-08-16 17:58:02 +02:00
parent f7dd817b97
commit bec9fdffd8
+200 -2
@@ -1279,11 +1279,209 @@ RabbitMQ client library is used only because LavinMQ speaks the same protocol.
---
## Three config keys that changed meaning — read this before upgrading
**What.** `weight`, `maxLoad` and `fleet.leaders.*` all changed what they *mean*, not just what they
do. A config file that worked before an upgrade can behave differently, or refuse to start, with no
edit to it. The three are grouped here because an upgrading operator meets them together.
| Key | Used to mean | Now means |
|---|---|---|
| `profiles.<n>.weight: 0` | coerced to 1.0 — "pick me as often as anyone else" | excluded from automatic selection |
| `profiles.<n>.maxLoad: 0` | coerced to unlimited | capped at zero live members |
| `fleet.leaders.<n>.terminal` | a terminal-id pin | **rejected at startup** — use `tab` |
**On.** Nothing to switch on. These are the meanings now.
**Why.** In the first two cases the old behaviour was the exact opposite of what the key reads like.
`weight: 0` looked like "never pick this" and meant "pick it normally". `maxLoad: 0` looked like
"never run anything here" and meant "unlimited" — and that was the only throttle a
`subscription: true` profile had against the operator's own paid plan. A key whose plain reading
inverts its behaviour is a trap, so both were fixed toward the reading. A negative `maxLoad` is now a
startup error rather than being silently normalised away.
`fleet.leaders.*.terminal` went further and is **refused**, naming the offending entries, instead of
being ignored. A silently-ignored lead pin means the lead is not recognised, gets demoted to worker,
and every orchestration call is refused — a failure that looks nothing like its cause. Failing at
startup is the kinder outcome.
**Gotcha.** `weight: 0` excludes a profile from *automatic* selection only. An explicit
`bridge_spawn{profile:"..."}` bypasses placement entirely and still resolves it, so `weight: 0` is
**not** a way to disable a profile. `maxLoad: 0` is, because that cap now applies to explicit spawns
too. And a lead's `tab` must match a real herdr tab label exactly (case-insensitively) — the old
`tabPrefix` no longer finds a lead, it only guards against a worker's label colliding with the
convention.
---
## Stop an explicit spawn from busting the cap
**What.** `maxLoad` is now an unconditional cap. An explicit `bridge_spawn{profile:"X"}` used to skip
the check entirely — the cap applied only to automatic placement — so naming a profile was a way
around it. Now an explicit spawn is refused when `live >= cap`, and it is **not** re-routed to
another profile.
**On.** `profiles.<name>.maxLoad`.
**Why.** A cap that any caller can opt out of by naming the profile is not a cap. The no-fallback part
is deliberate: a caller who named a profile did so for a cost or model reason, so quietly moving the
work to a different backend would defeat the reason they named it. Better to refuse and let the
caller decide.
**Gotcha.** There is a known TOCTOU race: `liveCount` is read outside a lock, so two genuinely
concurrent spawns can both pass the check and briefly exceed the cap. It is documented in the code
rather than fixed. Also note the refusal message only became visible to callers in CB-599 — before
that it was a blank 500 with the reason in the log.
---
## Nudge an idle lead back to work
**What.** An opt-in loop that nudges the one idle lead after it has been continuously injectable for
a quiet period with nothing driving it. It fills the gap the reply push loop leaves: that loop only
fires when a worker reply lands, so a lead that is simply sitting idle with nothing arriving is
invisible to it.
**On.** The `leadHeartbeat:` block — `idleAfterSeconds` (default 300), `backoffMs` (default 60000),
`quietNudgeCap` (default 3). **Absent means off**, and the loop is not even constructed. Changing it
needs a restart.
**Why.** Off by default, deliberately. This loop spends the operator's model subscription on the
daemon's own initiative, so it must never switch itself on during an upgrade. That constraint shaped
the whole design.
**Gotcha.** It is status-gated (a WORKING lead is never touched), debounced, caps consecutive quiet
nudges, and stands aside while the reply push loop is already nudging. That last one depends on
`ReplyPushLoop.isActive()` being honest — see CB-598, where a schedule could die with work still
pending and report inactive.
---
## Prove which charter a member got, without logging the prose
**What.** Per-role launch charters are read from the **live** config and composed once per spawn: the
role charter plus the reply charter. Claude Code gets it as an appended system prompt; opencode gets
an ephemeral `member-charter.md` referenced from its generated config. A `CharterReceipt` records a
sha256 and byte count of the exact composed bytes, and rides on both the spawn log and the roster.
**On.** `fleet.charters.<role>`, where role is `dev`, `architect` or `reviewer`. Read live per spawn,
so a change applies to the **next spawn with no restart**.
**Why.** Two problems at once. Charters had been hardcoded, which meant changing what a role is told
required a rebuild. And an operator asking "what was this member actually told?" had no answer that
did not involve printing the prose into a log, where it does not belong. The receipt answers the
question with a digest instead.
**Gotcha.** The role charter is combined with the reply charter **only when the profile mounts the
bridge MCP**. A non-MCP profile gets the role charter alone — which is correct, since the reply
charter tells a member to call `bridge_reply`, and a member with no bridge cannot.
---
## Tell "busy" from "refusing" in the capacity view
**What.** In `bridge_list`'s `capacity` view, a quarantined profile's row is forced to `free: 0`
whatever its `maxLoad` and `live` counts say, and gains `credentialId` and `quarantinedForSeconds`.
**On.** No new knob. The quarantine facts come from the existing `exhaustedPattern`, `credentialId`
and `quarantineCooldownSeconds` keys.
**Why.** A quarantined profile used to show free slots it would refuse to fill. A lead reading that
row would keep trying and keep failing. The two states need different reactions — "busy, will free
up" means wait, "refusing for N seconds" means go elsewhere — and the row could not express the
difference.
**Gotcha.** The two new fields appear **only** when a profile is actually quarantined, so an ordinary
fleet's rows are byte-identical to before. Do not write a client that expects them. The view shares
its `QuarantineSource` with `bridge_profiles` so the two surfaces cannot disagree.
---
## Resume a member onto its previous conversation
**What.** A member's own agent-session id is captured at spawn, survives onto the roster as
`agentSessionId` in `bridge_list` and `GET /members`, and `bridge_spawn` accepts `sessionName` and
`resumeSessionId` to relaunch onto that same conversation.
**On.** `bridge_spawn{sessionName, resumeSessionId}`.
**Why.** The pieces existed but were connected at neither end: nothing persisted the id, and nothing
exposed a way to pass it back. So a member that died took its context with it even though the backend
could have resumed it.
**Gotcha.** A `resumeSessionId` **requires an explicit `profile`**, and that profile's adapter must
declare `Capability.SESSION_RESUME`. A placement-routed spawn carrying a `resumeSessionId` is
refused, because a resumed conversation is tied to the specific backend that started it. If the
adapter lacks the capability the spawn is refused naming it, rather than silently starting a cold
session that looks resumed.
---
## opencode members run with `--auto` and forced auto-compaction
**What.** opencode workers launch with `--auto`, which auto-approves the permissions opencode does
not explicitly deny, and the bridge-generated config pins `compaction.auto: true`.
**On.** Neither is configurable. Both are unconditional for the `opencode` kind.
**Why.** A worker that stops to ask for permission on a tool call has no one to ask — the lead is not
watching its pane, and an unanswered prompt burns the whole turn. Auto-compaction is the same
argument: a worker that runs out of context dies mid-turn and loses its `bridge_reply`, so its report
is gone even though the work was done.
**Gotcha, and it is a real trade.** opencode's own help calls `--auto` "dangerous!". The blast radius
is bounded by everything else about a member — its own worktree, its own branch, off-subscription,
and it cannot merge — not by the flag. And `compaction.auto: true` in the generated config **wins
over the operator's home config**, so auto-compaction cannot be turned off for bridged workers from
there. That was chosen knowingly: a lost report was judged worse than an unwanted compaction.
---
## Silent member failures now say why
**What.** Logging only, no behaviour change. The paths that used to fail a member with no log, a bare
DEBUG, or a message naming only the symptom now log at WARN with the real cause and the numbers —
across the injector, the completion resolver, session acquire/reap/failure, and ticket abandonment.
**On.** Automatic.
**Why.** A member could fail and leave nothing to read. The lead saw a session that stopped
responding and had no way to tell a crashed backend from a wedged pane from a reaped session.
**Gotcha.** You will see more WARN lines than before if you filter by level. That is the feature, not
noise — but it will change what a log-volume alert sees.
---
## Sessions are one-shot — `recycle()` is gone
**What.** `SessionManager.recycle()` — release plus immediate re-acquire under a new pane id — was
removed. A released or finished session is torn down, never pooled or reused. Callers must `release`
then `acquire` explicitly.
**On.** Nothing to configure; this is a removal.
**Why.** Recycling quietly reused a session across two unrelated pieces of work. The lifecycle is
meant to be strictly one-shot, and a method that bypassed it was a standing invitation to leak one
delegation's context into the next.
**Gotcha.** None left in-tree — no references remain. It had no external surface, so only internal
callers were ever affected.
---
## Backfill status
This page was started after the fact, so it is **not yet complete**. Entries above are written from
verified behaviour. Still to catalogue — each needs its config surface and gotcha confirmed against
the code before it earns an entry:
verified behaviour.
The CB-595 backfill (2026-08-16) cleared the largest gap: the ~14 operator-facing tickets shipped
between `v1.0.0` and now that had landed nowhere. Those entries were written from a read of the code
on `main`, not from commit messages, and the three keys that changed *meaning* were pulled to the
front because an upgrading operator meets them first.
Still to catalogue — each needs its config surface and gotcha confirmed against the code before it
earns an entry:
- `/metrics` and `/healthz`, and what each does *not* tell you (CB-502; `/healthz` reports herdr
reachability only — it went green while every spawn failed, see the CB-521 version-coupling note)