11, 7: the multi-lead arc — two leads as peers, and the fleet they can see

Chapter 11 gains seven entries covering CB-530..536, the arc that turns one
primary-plus-workers into two leads working as peers:

  CB-530  `leaders:` — a registry, not a singleton pin
  CB-530  unknown top-level config keys are named at load (the bug that let a
          hand-written `leaders:` block look configured while being inert)
  CB-531  `leadScan:` — discover a lead by the tab label a human typed
  CB-532  leads message each other and are answered; `primary:` retired, with
          reply nudges now following the lead that delegated
  CB-534  a lead is deliverable — the CB-113 readiness gate opens for one
  CB-535  `bridge_list` returns `leads` alongside `workers`, `self` on your row

Each carries the why, not just the knob. CB-534's and CB-535's exist because both
failed *silently*: the gate held every lead-to-lead send for ~60s and then failed
it as a stalled turn, and an empty `workers` array read as "no peers" to a lead
that had one. Signatures of both are recorded so a recurrence is recognisable.

Also corrected in place: the CB-530 gotcha still said to keep `primary:` beside
`leaders:`. CB-532 retired it, so the file contradicted itself two sections apart
— it now says delete it and points at the entry that explains why.

Chapter 11's lifecycle entry picks up `clearAfterTurn` (CB-537), which is in the
Lifecycle record as shipped. Note its design is already superseded: per-delivery
inherit|fresh|thread policy applied pre-delivery, because a post-turn reset races
by construction.

Chapter 7's copy of the portable CLAUDE.md block is re-synced byte-identically
(verified) with the lead-to-lead section: coordinate, never delegate sideways.
Dai Ha
2026-08-13 10:43:06 +02:00
parent eef96843a9
commit 7b5381bdb4
2 changed files with 230 additions and 7 deletions
+197 -4
@@ -21,6 +21,13 @@ six weeks, and the table alone will not carry it.
|---|---|---|---|
| [Ask the bridge who you are](#ask-the-bridge-who-you-are) | `bridge_whoami` | CB-517 | `mcp/BridgeMcp` |
| [Primary inside a herdr pane](#primary-inside-a-herdr-pane) | `primary.terminal:` | CB-522 | `auth/CallerResolver` |
| [More than one lead](#more-than-one-lead) | `leaders:` | CB-530 | `auth/CallerResolver` |
| [Unknown config keys are named](#unknown-config-keys-are-named) | (always on) | CB-530 | `config/BridgedConfig` |
| [Find leads by tab name](#find-leads-by-tab-name) | `leadScan:` | CB-531 | `herdr/LeadTabScanner` |
| [Leads talk to each other](#leads-talk-to-each-other) | (always on, two leads) | CB-532 | `auth/Principal` |
| [A lead can be delivered to](#a-lead-can-be-delivered-to) | automatic | CB-534 | `Bridged.deliverableTo` |
| [Leads are visible in bridge_list](#leads-are-visible-in-bridge_list) | automatic | CB-535 | `mcp/BridgeMcp.listFleet` |
| [Reply nudges follow the delegating lead](#reply-nudges-follow-the-delegating-lead) | automatic (retires `primary:`) | CB-532 | `mcp/PrimaryRegistry` |
| [Weighted worker placement](#weighted-worker-placement) | `placement: weighted` + `weight` / `maxLoad` | CB-518 | `placement/` |
| [Give workers a toolchain](#give-workers-a-toolchain) | per-profile `env:` | CB-511 | `worker/HerdrPeerLauncher` |
| [Isolated worktree per worker](#isolated-worktree-per-worker) | `bridge_spawn{worktree, ticket}` | CB-301-ext | `session/GitWorktrees` |
@@ -167,14 +174,20 @@ gated.
## Session lifecycle caps
**What.** Reaps idle sessions, caps turns per session, and drains cleanly on shutdown.
**What.** Reaps idle sessions, caps turns per session, optionally clears a reused Claude Code
worker's conversation after each delegation, and drains cleanly on shutdown.
**On.** `lifecycle: { idleTtlSeconds, contextCap, drainTimeoutSeconds, clearAfterTurn }`.
**Why.** Workers are disposable but not free; without caps an abandoned session holds a pane and a
context indefinitely.
context indefinitely. `clearAfterTurn: true` keeps the pane and process warm while preventing
delegation N+1 from inheriting delegation N's conversation.
**Gotcha.** Read at boot — changes need a daemon restart.
**Gotcha.** `clearAfterTurn` defaults to `false`. It uses Claude Code's `/clear` command without
counting that housekeeping as a delegated turn, and waits for it to settle before delivering the
next task. Peer kinds without a known context-reset operation (currently opencode) treat the knob as
a no-op and log that once; bridged never guesses a command. Lifecycle config is read at boot, so
changes need a daemon restart.
## Durable reply inbox
@@ -253,6 +266,187 @@ Delete it before anything else. And skills do not cross: `.claude/skills/**` is
so a brief telling a peer to "load the *implementer* skill" is a no-op there — spell the procedure
out in the brief instead.
## More than one lead
**What.** Recognises several panes as leads, so two orchestrators — say a Claude lead and an
opencode lead — work as peers instead of one being demoted.
**On.**
```yaml
leaders:
opus-5.0:
terminal: term_0123456789abcd
kind: claude
gpt-sol-5.6:
terminal: term_fedcba9876543
kind: opencode
model: openai/gpt-5.6-terra
```
`terminal` is the only field identity depends on; `kind`/`model` are descriptive and are echoed back
by `bridge_whoami` as `leader: <name>`. `role` still reads `primary` — a lead **is** a primary for
authorization, so nothing keying on the role breaks.
**Why.** `primary.terminal` is singular by construction: one pane is the lead and every other pane
resolving to a herdr terminal is a worker. That is right while one lead drives a fleet, and wrong the
moment two leads collaborate — the second is silently demoted and refused every orchestration call it
makes. Resolution is now a registry lookup rather than an equality test against one pin.
**Gotcha.** This entry originally said to keep `primary:` alongside `leaders:`, because the CB-307
push loop needed exactly one nudge destination while `leaders:` only widened who was *recognised*.
That is no longer true: [reply nudges now follow the delegating
lead](#reply-nudges-follow-the-delegating-lead), and `primary:` is retired. Delete it. If both are
present and name the same terminal the `leaders:` entry wins. And a lead is never spawned — it
pre-exists, which is why it
must be named here rather than created; `argv`/`placement` are worker-profile keys and mean nothing
in this block. Read at boot, so a change needs a restart.
## Find leads by tab name
**What.** Discovers leads by scanning herdr for tabs *you* labelled, instead of you pasting each
lead's `terminal_id` into [`leaders:`](#more-than-one-lead). Label a tab `lead: gpt-sol-5.6`, start
an agent in it, and that pane resolves as a lead named `gpt-sol-5.6` within one rescan — no config
edit, no daemon restart.
**On.**
```yaml
leadScan:
tabPrefix: "lead:" # matched case-insensitively; the rest of the label is the lead's name
intervalSeconds: 10 # rescan cadence, and the worst case before a new tab is recognised
```
Opt-in: no block means leads come only from `leaders:`/`primary:`, exactly as before. Both sources
merge, and an explicit `leaders:` entry outranks a label for the same terminal.
**Why.** A lead is never spawned — a human opens a tab and starts an agent in it — so unlike a
worker, the daemon cannot learn its `terminal_id` at creation. `leaders:` therefore costs a
four-step ritual per lead: start the session, ask it `bridge_whoami` for its id, edit config,
restart. Naming the tab is one step, taken at the moment the operator is already there. The label
also survives what the id does not: close and reopen the tab and the `terminal_id` changes, while
the label is retyped as-is.
**Gotcha.** The direction of trust is what makes this safe, and it is one-way: bridged **reads**
lead tab labels and never writes them, so what is in the tab bar is always what a human typed.
Two guards keep that from eroding — the configured worker spaces (where bridged *does* write
labels, via `tabLabel`) are excluded from the scan wholesale, so nothing the bridge places can land
in a matching tab; and startup **refuses** a `tabPrefix` that any worker `tabLabel` also matches,
because overlapping those two namespaces would have the daemon label its own workers as leads and
promote the entire fleet. Every pane in a labelled tab is that lead, so split a lead tab only with
panes you mean to be leads. A failed scan keeps the leads already known rather than emptying the
registry — a herdr hiccup must not demote a live lead mid-session.
## Leads talk to each other
**What.** A lead can message another lead and *be answered*. `bridge_send{sessionId: <peer's
terminal>}` reaches a peer, and the peer closes the exchange with `bridge_reply` — the same
rendezvous a worker uses. `bridge_whoami` now reports a lead's own `sessionId`, which is how a lead
learns the address to give a peer.
**On.** Nothing to configure; it applies as soon as two panes resolve as leads (via
[`leaders:`](#more-than-one-lead) or [`leadScan:`](#find-leads-by-tab-name)).
**Why.** CB-530 and CB-531 widened *recognition* — both leads are seen — but nothing had widened
*addressing*, so collaboration was one-way and silently so: the send was accepted, the peer's
`bridge_reply` was refused, and the sender waited out its timeout. The cause was that a lead's
`Principal` carried no terminal, so `ownsSession()` could never be true for it and the `REPLY`/`ASK`
rules excluded it by construction. A lead now carries the pane it was matched by, and the rule it
must satisfy is unchanged: **you may act as the pane you occupy, and as no other**. That was always
the real control — `terminal.equals(sessionId)`, against a terminal that comes from the connection —
and the extra "…and you must be a worker" conjunct beside it protected nothing.
**Gotcha.** A lead replies **only** to answer a peer that messaged it — never to answer a worker,
whose turn it is not, and never as a way to end its own turn. Widening who may reply did not widen
what they may reply *as*: a lead still cannot act for another pane, and an unnamed primary (token
mode, or off-host, with no pane at all) owns nothing and remains a sender only.
## Leads are visible in bridge_list
**What.** `bridge_list` returns `leads` alongside `workers`. Each lead row carries its `sessionId`
(the address to `bridge_send` to), its `name`, its live status, and `self: true` on the caller's own
row.
**On.** Automatic, wherever a pane resolves as a lead.
**Why.** A lead had no way to discover a peer. `bridge_list` enumerated the worker roster alone, so a
lead asking "who else is here?" got an empty array — which reads as *no peers* but only ever meant
*no workers spawned*. The peer lead on this bridge drew exactly that wrong conclusion and reported
itself alone in a two-lead fleet. Addresses had to be carried between panes by a human, which is not
a protocol. Reporting both halves — even when a half is empty — also removes the ambiguity that
caused the misreading.
**Gotcha.** The lead rows come from the same registry that *resolves* identity
([`CallerResolver.leads()`](#find-leads-by-tab-name)), not from a second copy, so a listed address is
one that would actually resolve as a lead. A lead herdr is not tracking as an agent is listed with
`status: unknown` rather than hidden — it cannot be delivered to, and a would-be sender needs to see
that rather than infer it from silence.
## A lead can be delivered to
**What.** The injector's readiness gate opens for a lead as well as for a booted worker, so a message
addressed to a lead is actually typed into its pane.
**On.** Automatic, wherever a pane resolves as a lead.
**Why.** The gate (CB-113) holds a delivery out of a *spawned* worker's boot window: herdr reports
`idle` while the agent is still starting, and a paste into that window is lost. Membership in it
comes from `WorkerPresence`, which `BridgeMcp` populates **only for a worker** — deliberately, since
that map doubles as the roster's availability signal and a lead counted there would appear as an
available worker. The two rules composed into a dead end: a lead is never marked present, so the
gate never opened for one, so [lead-to-lead messaging](#leads-talk-to-each-other) — shipped and
authorized in CB-532 — still could not deliver a single keystroke. The gate's premise simply does not
apply to a lead: a lead is never spawned, so it has no boot window to guard.
**Gotcha.** The failure this fixes was slow and mute, which is worth recognising if it recurs in
another form: the send was accepted, the pane stayed `idle`, nothing was ever typed, and ~60s later
(`READINESS_GRACE_POLLS`, 240 × 250ms) it failed through the same path as a stalled turn — so the log
said *turn-stall fallback* while the truth was that delivery had never been attempted. A repeating
62-second gap between send and failure is the signature of the gate, not of a peer that ignored you.
It was also intermittently masked: presence is a sticky set, so a pane that was seen as a worker
*before* being recognised as a lead stayed deliverable until the next restart cleared the set.
## Reply nudges follow the delegating lead
**What.** When a worker's reply lands with no `bridge_send` open, the CB-307 nudge goes to the lead
that delegated that worker — not to a globally-configured "the primary". This is what retires
`primary.terminal:`, which now logs a deprecation warning at startup.
**On.** Automatic. Delete `primary:` from `bridged.yaml`; keep it only if you still want
`pushReminders`/`pushBackoffMs`, or a fallback nudge destination across restarts.
**Why.** `PrimaryRegistry` held one slot answering "who is the primary" — a question with no correct
answer once two leads drive one fleet. Whichever lead called `bridge_send` first captured *every*
nudge thereafter, so the other lead's results were announced to the wrong pane. The binding that
actually matters is per-delegation and is known exactly where it is created: at `bridge_send`, where
the target is the argument and the lead is resolved from the connection.
**Gotcha.** A restart loses the delegation map while the durable inbox keeps the reply. With one
lead the old pin (or the first lead to send) is an unambiguous fallback; with several and no
recorded delegation the daemon nudges **nobody** rather than guessing, and delivery degrades to
`bridge_poll`. That is the correct degradation — interrupting the wrong lead with someone else's
result is worse than a quiet inbox — but it does mean a post-restart reply may need an explicit
poll.
## Unknown config keys are named
**What.** A top-level key in `bridged.yaml` that this build does not understand is logged as a WARN
naming it, at load.
**On.** Always on; nothing to configure.
**Why.** Every config record is `@JsonIgnoreProperties(ignoreUnknown = true)` — deliberate, so config
may run ahead of the code and a rolled-back daemon still starts. The cost is that a whole block can be
written, parsed, dropped, and never mentioned again. That is exactly how a hand-written `leaders:`
registry came to look configured while being inert: the daemon started, nothing complained, and the
only way to find out was reading the config class. A dropped block and a working one were
indistinguishable.
**Gotcha.** A warning, not a failure — deliberately. Failing closed would turn "the config names
something this build has not learned yet" into a daemon that will not boot, destroying the
forward-compatibility the annotation exists for. Nested unknown keys are still silent; only the top
level is checked.
---
## Backfill status
@@ -268,4 +462,3 @@ the code before it earns an entry:
- launchd / systemd supervision (CB-504)
- multi-profile routing and `kind:` adapter selection (CB-305, CB-401/402)
- the reply push loop and its nudge budget (CB-307 step 2)
+33 -3
@@ -329,8 +329,9 @@ nothing. Fail toward the recoverable error.
2. **The bridge is the only channel.** Text you print in your terminal reaches nobody — the other
side cannot see your screen. An answer that isn't in a `bridge_*` call is silently discarded.
3. **Identity comes from the connection, never an argument.** Workers never pass a target; you
cannot act as another session. Spawn/stop/send/drain are primary-only; reply/ask are
worker-only-and-only-as-itself. A call outside your role is refused, not queued.
cannot act as another session. Spawn/stop/send/drain are lead-only; reply/ask are
only-as-itself — any peer may answer for its own pane, and for no other. A call outside your
role is refused, not queued.
4. **Delivery is status-gated: one message per turn.** Don't busy-poll a peer's terminal and don't
re-send because a call looks slow — the bridge delivers when the peer is `idle`/`blocked`.
5. **Never drive the terminal multiplexer directly** (no `herdr` CLI, no socket). The bridge owns
@@ -389,13 +390,42 @@ the merge — and merging on a reviewer's word is delegating it by proxy.
| Confirm your own role | `bridge_whoami` |
| See backends available | `bridge_profiles` |
| Start a worker | `bridge_spawn{profile?, cwd?, worktree?, ticket?}` → `sessionId` + `paneId` |
| See the fleet | `bridge_list` · one worker's state: `bridge_status{sessionId}` |
| See the fleet | `bridge_list` → `leads` (your peers) + `workers` · one peer's state: `bridge_status{sessionId}` |
| Delegate (blocking) | `bridge_send{sessionId, content}` |
| Delegate (long task) | `bridge_send{sessionId, content, wait:false}` → ticket → `bridge_poll{ticket}` |
| Answer a worker's `bridge_ask` | `bridge_send{turnId, content}` — **not** `sessionId` |
| Message a **peer lead** | `bridge_send{sessionId: <their terminal>, content}` — `bridge_list` → `leads` reports it. Coordination only, **never** a task |
| Answer a peer lead that messaged you | `bridge_reply{content}` — the one case a lead replies |
| Collect a held reply | `bridge_poll{target}` · then `bridge_ack{target, msgId}` |
| Tear down | `bridge_stop{paneId}` |
### Lead ↔ lead — coordinate, never delegate
`bridge_list` returns `leads` alongside `workers`; your own row carries `self: true`. Every other row
is a peer — an orchestrator with its own context, its own workers, and its own judgment. An empty
`workers` array means no workers are spawned; it says nothing about peers.
**A lead never assigns a task to another lead.** Work goes to workers — only ever downward, never
sideways. Sending a peer a brief with acceptance criteria is a category error: a brief is a worker's
artefact, and a peer is not yours to task. If a unit needs doing and it falls in your area, spawn a
worker and delegate it yourself; if it falls in the peer's area, say so and let the peer assign it.
The traffic between leads is coordination and nothing else:
1. **Divide the map, not the work.** Agree who owns which area, then each of you assigns inside your
own. Split by **context ownership** — whoever already holds the context owns that area — and say
who takes what, in one message, before either of you starts. Two leads silently working the same
unit is the failure mode here, and neither notices until the merge.
2. **Share findings, hazards, and corrections.** What you have already discovered, what broke, what
the next person will trip on. This is the traffic that actually pays for the channel: it costs one
message and saves a peer a rediscovery.
3. **Verify a peer exactly as you verify yourself.** Peer status buys nothing: check the claim
against the code, and re-run the build. A peer's correction gets the same treatment — right or
wrong on the evidence, not on who said it. Neither of you merges the other's work unreviewed.
Being messaged by a peer does not make you its worker: answer with `bridge_reply`, and push back on
the substance if it is wrong. A peer that simply complies has thrown away the reason there are two of
you.
### Worker — the turn contract
1. **Load the playbook skill the lead named** before doing anything else.