Commit Graph

108 Commits

Author SHA1 Message Date
Dai Ha cea1183f75 CB-569: pass composed charter to Claude
CI / contract (pull_request) Successful in 43s
CI / build (pull_request) Successful in 50s
2026-08-15 05:09:30 +02:00
Dai Ha e81944cef6 CB-567: compose charters per spawn
CI / build (pull_request) Successful in 57s
CI / contract (pull_request) Successful in 1m3s
2026-08-15 05:00:48 +02:00
Dai Ha 799668e129 CB-566: add fleet charter config
CI / contract (pull_request) Successful in 43s
CI / build (pull_request) Successful in 55s
2026-08-15 04:50:23 +02:00
Dai Ha f8522edacd Merge CB-564: give every silent failure path a voice
CI / build (push) Successful in 1m0s
CI / contract (push) Successful in 1m15s
The bridge cannot classify what it never emits. Fleet-health monitoring —
detect a wedged member, decide, escalate — is blocked on that, so this is the
foundation rather than the feature.

A survey of Injector, CompletionResolver, StatusPoller, SessionManager,
SessionReaper and MessageService found six conditions that ended a member's
usefulness while saying nothing useful:

  * Injector.drop()             — worker gone, queue cleared: SILENT
  * CompletionResolver.fail()   — "via turn-stall fallback" at DEBUG, no reason
  * SessionManager.onFailed()   — "session marked failed" at DEBUG, no stage
  * SessionManager.reapIdle()   — indistinguishable from any other release
  * SessionManager.acquire()    — spawn failure rethrown with no log at all
  * MessageService.abandon()    — failed a caller's request at DEBUG

The first two are the exact phrases that misled the CB-560 diagnosis: both name
a symptom and neither names a cause. They are now WARN and carry the reason,
the stage, and the counts.

Conditions already loud were left alone, and so were two by-design timeouts in
MessageService — an async model exists precisely for those, and promoting them
would turn healthy operation into noise.

Behaviour is unchanged: every edit is a log statement.

Merge note: the recycle test deleted by CB-565 conflicted with a test added
here. Resolved by keeping the new onTurnFailed assertion and dropping the
recycle test, which tests a method that no longer exists.
2026-08-15 04:36:09 +02:00
Dai Ha 078bde2c02 CB-564: give a voice to silent member-failure paths
Injector.drop, CompletionResolver.fail, SessionManager.onFailed/reapIdle/
acquire spawn failures, and MessageService.abandon used to fail a member
or a caller's request with no log, a bare DEBUG, or a log that named only
the symptom ("session marked failed", "failed send via turn-stall
fallback"). Each now logs at WARN and names the real cause and the
numbers involved. Observability only — no behaviour changed.
2026-08-15 04:31:42 +02:00
Dai Ha 0b10ea987b CB-565: remove unsafe session recycle
CI / contract (pull_request) Successful in 44s
CI / build (pull_request) Successful in 55s
2026-08-15 04:29:50 +02:00
Dai Ha d04b075996 CB-563: mark clipped completion scrapes
CI / contract (pull_request) Successful in 42s
CI / build (pull_request) Successful in 56s
2026-08-15 04:25:54 +02:00
Dai Ha 644927636d Merge CB-562: the readiness gate says why it gave up (PR #33)
CI / build (push) Successful in 1m1s
CI / contract (push) Successful in 1m14s
When the injector's readiness grace expired it cleared the queue and logged
nothing. The failure then surfaced elsewhere as a turn-stall, which names the
wrong cause. Diagnosing CB-560 cost two live spawns and a wrong first
hypothesis for exactly this reason: the logs said "session marked failed" and
"failed send via turn-stall fallback", and neither says the message was never
typed into the pane at all.

The expiry now logs the target, the number of messages being failed, the grace
in polls and seconds, and the real cause in plain words.

The grace in seconds is derived, not written down twice: Bridged's own
INJECT_POLL_MILLIS is deleted and Injector.POLL_INTERVAL_MILLIS is the single
source, passed to every StatusPoller. A cadence change can no longer leave a
log line confidently stating the wrong duration.

Behaviour is unchanged. This is the first structured health event in the
daemon, and the foundation the fleet-health work will build on.
2026-08-15 04:21:29 +02:00
Dai Ha 95e45007aa CB-562: single source for the injector poll cadence; tighten count assertion
CI / build (pull_request) Successful in 59s
CI / contract (pull_request) Successful in 1m17s
2026-08-15 04:20:29 +02:00
Dai Ha 7a583c4045 CB-562: log why the readiness gate gave up on a target
CI / build (pull_request) Successful in 57s
CI / contract (pull_request) Successful in 1m4s
2026-08-14 21:52:25 +02:00
Dai Ha 65bce058bb Merge CB-561: one public way to build a CallerResolver (PR #32)
CI / build (push) Successful in 1m17s
CI / contract (push) Successful in 1m16s
CallerResolver had 5 public constructors and 4 public factories, and only one
of them could ever produce an architect. The rest defaulted memberSlotRoles to
`_ -> null`, so every architect quietly fell through to Principal.worker().
Nothing logged, nothing threw — the role was simply off.

That is the same failure shape as CB-560, so the fix is structural rather than
a warning: withLeadsAndMembers(.., MemberRegistry) is now the only public
construction path. Two overloads with no caller at all are deleted; the rest
are package-private and marked test-only. No path remains that accepts
architect bindings without a slot-role lookup, so no runtime WARN is needed.

Also checked and closed: the suspected slot leak on shutdown drain is not
real. SessionManager.release calls memberLifecycle.released() for every
removed session, and drainAll routes every session through release,
SPAWNING included. Verified in code.
2026-08-14 21:50:04 +02:00
Dai Ha fa0612859b CB-560: document spawned member presence
CI / build (pull_request) Successful in 1m0s
CI / contract (pull_request) Successful in 1m4s
2026-08-14 21:48:58 +02:00
Dai Ha 4ffbcd0b7d CB-561: require member registry for architects
CI / contract (pull_request) Successful in 43s
CI / build (pull_request) Successful in 50s
2026-08-14 21:47:53 +02:00
Dai Ha a0cd053fd9 CB-560: mark architect members present
CI / build (pull_request) Successful in 55s
CI / contract (pull_request) Successful in 1m4s
2026-08-14 21:45:28 +02:00
Dai Ha e9bc192160 CB-548: bind architect sessions to member slots
CI / contract (pull_request) Successful in 45s
CI / build (pull_request) Successful in 55s
2026-08-14 21:10:48 +02:00
Dai Ha b414a74c26 CB-559: drop .claude/settings.local.json from the default parity overlay
CI / build (pull_request) Successful in 54s
CI / contract (pull_request) Successful in 1m7s
2026-08-14 20:56:16 +02:00
Dai Ha e18e002d2f CB-559: a profile's launch settings are deferred, not hot
The shipped docs and javadoc said a profile's `model` and `tabLabel` take effect on
the next spawn. They do not, and ConfigRef did not detect the change either, so a
reload logged a clean "config reloaded" and silently did nothing. That is the worst
outcome a reload can produce: the operator has no reason to doubt it.

What makes a key hot is who reads it and when, not that it is config. Placement
reads weight and maxLoad through a supplier on CompositePeerLauncher, so those are
genuinely hot. HerdrPeerLauncher takes Map.copyOf(profiles) at construction and
resolves every spawn out of that copy, so model, baseUrl, argv, env and the rest
cannot move until the daemon restarts.

changedDeferredKeys now compares every launch component of an existing profile,
excluding weight and maxLoad, and names the profiles that need a restart. The
javadoc and bridged.example.yaml say the same thing. Two tests pin the pair:
weight/maxLoad reports nothing deferred, a changed model reports the profile by name.
2026-08-14 20:37:42 +02:00
Dai Ha a2108a8a14 CB-559: re-read bridged.yaml without restarting the daemon
Tuning a fleet meant restarting bridged, and a restart tears down every lead
and worker it owns. Changing one pool's weight cost the whole fleet's state,
so in practice nobody changed it.

ConfigRef holds the live BridgedConfig in an AtomicReference. Consumers read
it at the point of use, so a change reaches the next spawn with nothing
rebuilt. The launchers that used to capture config into fields now take
suppliers: the fleet tabLabel template, the profile map, the placement policy
and the fleet block.

Keys fall into three classes, and the difference is what already exists when
the reload happens:

  hot       fleet: (pools + tabLabel), placement:, and an existing profile's
            weight / maxLoad / model / tabLabel — live on the next spawn.
  deferred  lifecycle:, leadHeartbeat:, guard:, worktreeRoot:, spawnReady*,
            and adding/removing a profile — accepted, but the startup wiring
            keeps the old value. The reload logs these by name.
  cold      bind:, herdrSocket:, broker:, auth: — refuses the WHOLE reload.

A cold change refuses everything rather than applying the hot half. A
half-applied reload leaves the daemon matching no file on disk, which is the
worst thing a reload can do to an operator reading that file to work out what
the daemon is doing. Refusing keeps the invariant that the live config is
always some version of the file.

A parse failure or a failed startup validator is refused the same way, and
the running config stays live: a file being saved is sometimes read
mid-write, and degrading a working daemon over a half-written file is a bad
trade. The same four validators startup runs are re-run, so a config that
could not have booted cannot slip in through a reload.

ConfigWatcher polls the modified time on a daemon thread, opt-in through
configReload.enabled (default off, so an upgraded daemon is unchanged). It
stamps the timestamp BEFORE reloading, so a refused file is not retried every
tick — the next save earns a fresh attempt. A missing file is skipped
silently, because editors unlink briefly mid-save.

MicroProfile Config was the first idea and does not fit: @ConfigMapping needs
interfaces, resolves once at bootstrap, and reload would still mean rebuild
and swap. The port would also lose the raw-YAML duplicate-key detection,
since duplicates have already collapsed once the tree is flattened to
properties.

634 tests.
2026-08-14 18:23:32 +02:00
Dai Ha 57f8fa257a CB-558: launch a declared lead at startup when none is live
`fleet.leaders.<name>.instances` was descriptive. Now the daemon reads it: a
lead that names a `profile:` is started when fewer than `instances` are running.
A lead with only a `terminal:` stays recognise-only, as before.

A lead is not a member, and LeadLauncher exists to keep it that way. Every other
spawn path goes through HerdrPeerLauncher, which does three things a lead must
never get: it appends the worker reply charter ("you are an off-subscription
worker … end every turn with bridge_reply" — the opposite of an orchestrator);
it registers the session with SessionManager, whose idle reaper would kill a
lead for being idle, which is a lead's normal state; and it can move a peer off
the subscription. So this launcher talks to AgentControl/WorkspaceControl
directly. The duplicated argv/env assembly is the cheaper half of that trade.

Not double-spawning is the safety property, so liveness needs two pieces of
evidence. A running agent in a tab labelled `lead: <name>` finds an
auto-launched lead. A running agent on a pinned `terminal:` finds one the
operator opened by hand — without it, a pinned lead whose tab carries no
matching label would be relaunched on every boot. Member workspaces are
excluded, so a member in a matching tab is never counted. If herdr cannot be
reached, nothing is started: a second orchestrator is worse than none.

Liveness deliberately requires the AGENT, not just the label. LeadTabScanner
used to promise that bridged never writes a lead label, so there was no
round-trip from the daemon's own rename back into its next decision. That is no
longer true, and its javadoc now says so. The trust direction is unaffected — a
label is a name, not a capability — but staleness becomes real: a label left by
a crashed session would otherwise read as a live lead forever and disable
auto-launch permanently.

Two new knobs. `workspace:` (default "leads") is where a launched lead's tab
goes; it must not be a member workspace, because those are excluded from the
scan and a lead placed in one would never be found again. `cwd:` defaults to
bridged's own working directory.

Also: WorkspaceControl.listTabs, and a FakeHerdr tab seeder that leaves the
canned response byte-identical when no tab is seeded.

617 tests pass (16 new), IDE-clean.
2026-08-14 16:47:57 +02:00
Dai Ha 61944fc045 CB-557: place an unqualified spawn inside its role's pool
The pools were config-only until now: the launchers still received one global
effectiveDefaultProfile and placement still ranged over every configured
profile, so a reviewer could be placed on an architect-only backend.

Three parts:

SessionManager computed the role, stored it on the MemberSession, and never put
it on the SpawnRequest. So the role reached the record that describes the spawn
but not the call that performs it — every launcher saw DEV. Both spawn paths
(plain and worktree) now carry it.

CompositePeerLauncher takes the Fleet and draws its candidates from
fleet.<role> instead of from all profiles. An absent or empty pool means
unconstrained, not blocked: a config that declares pools for some roles must
keep spawning the rest, so it falls back to every profile. A null Fleet is the
pre-CB-557 wiring and behaves exactly as before.

Bridged passes cfg.fleet() to the composite and cfg.fleet().tabLabel() to both
launchers. The tab-label knob was accepted by HerdrPeerLauncher but passed by
nobody, so it was inert — the label only looked right because the fallback
happened to match the configured template. Four tests now pin the wiring
instead of the coincidence.

An EXPLICIT profile stays exempt from the pool. `bridge_spawn{profile:"opus"}`
carries no role, so it defaults to DEV; judging it against the dev pool would
refuse a spawn the operator asked for by name. maxLoad still applies to it.

Also cleared the IDE warnings in the touched files: an immediately-rethrown
catch (the comment stays, the redundant block goes), unused lambda params, a
javadoc link to a package-private class, two unused imports.

601 tests pass.
2026-08-14 16:37:29 +02:00
Dai Ha 4b48d2d921 CB-557: fleet role pools — role is the config key, and the tab label says it
Four top-level keys (leaders:, members:, leadScan:, defaultProfile:) become one
`fleet:` block, and a member's role becomes the map key that contains it rather
than a `role:` field inside it.

Why the key and not a field: a misspelled `role: architct` used to produce a
member with no contract, which nothing rejected. A misspelled pool name declares
nothing, which is a shape the loader can see.

`fleet.architects/developers/reviewers` are pools of profiles a role MAY run on.
That replaces the single global `defaultProfile:`, so an unqualified spawn now
resolves its profile from the pool of the role it asked for. Role and profile
stay orthogonal: a reviewer may run on the same profile as the dev it reviews,
and one profile may appear in several pools.

Tab labels are role-first — `dev: sonnet #4`. The template lives on `fleet:`
because a profile cannot know the role of the member launched on it; a profile
may still override it. The `{n}` counter is scoped per role+profile, so a dev
and a reviewer on one profile each start at #1. Making {role} the first field
also turns the lead/member namespace check into a structural guarantee: roles
are a closed enum, so only hand-written templates can still collide with a lead
tabPrefix.

Removed keys are hard errors that name their successor. `defaultProfile:` has no
single successor key, so its message explains the new model instead of pointing
at a key that does not exist.

Map order is kept with LinkedHashMap, deliberately not Map.copyOf — the latter
salts iteration order per JVM run, which would destroy the YAML definition order
that `placement: fixed` selects on.

Not yet wired: SessionManager still hands the launchers one effectiveDefault-
Profile, so pools are not enforced at spawn time yet, and placement still ranges
over all profiles.

595 tests pass.
2026-08-14 16:32:54 +02:00
Dai Ha 4875127daa CB-557: adopt the member taxonomy in the API surface
Every spawned peer is now a member with a role, and the role travels with it
from the spawn call to the roster.

MCP:
  bridge_spawn gains role: architect | dev | reviewer (default dev). An
    unknown role is refused with the valid spellings in the message.
  bridge_list returns "members" instead of "workers"; each row carries both
    role (what it is for) and profile (which backend it runs on).
  The spawn result echoes the role back, so a spawn that fell back to dev is
    visible rather than silent.

REST:
  GET/POST /members and DELETE /members/{paneId} replace /workers.
  POST accepts role= as a query param or a body field; an unknown role is 400.

Code:
  dev.ltms.bridged.worker package -> dev.ltms.bridged.member
  WorkerSession   -> MemberSession, plus a MemberRole role component
  WorkerPresence  -> MemberPresence
  SessionManager.acquire gains a role parameter; the existing overloads keep
    working and default to DEV, which is exactly what "worker" used to mean.

ClaudeCodeLauncher and OpenCodeLauncher keep their names on purpose — they
are named after the backend, not the role.

Not done here: the launch charter is still one string for every role, so a
member is told its role by nobody yet. That is the next ticket.

mvn clean install: 583 tests, 0 failures, 0 errors, BUILD SUCCESS.
2026-08-14 07:08:26 +02:00
Dai Ha 246f50b778 CB-557: adopt the member taxonomy in config
A member is anything a lead spawns. Every member carries two independent
attributes:

  role    — which contract: architect, dev or reviewer. It picks the launch
            charter, the role file, the playbook skill and the authz row.
  profile — which backend: model, CLI adapter, credentials, cost.

They vary on their own. A reviewer may run on the same profile as the dev
whose diff it reads, which is the case that proves the two cannot be one
field.

Config changes (breaking — we are in active development, so no aliases):

  workers:       -> profiles:        it was never a list of workers; it is a
                                     catalogue of backends
  defaultWorker: -> defaultProfile:
  architects:    -> members:         each slot now names its role

An old config is rejected at load with the new key named, rather than being
warned about once and then running with zero profiles — that failure would
surface much later, at the first spawn, pointing nowhere near the cause.

Also:
  - BridgedConfig.Worker    -> BridgedConfig.Profile
  - ArchitectRegistry       -> MemberRegistry
  - new peer.MemberRole enum, validated at startup
  - profiles map is normalized once in the compact constructor, so the raw
    map and the derived one can no longer disagree
  - the legacy singular worker: block is dropped
  - workerProfiles() -> profiles(); defaultProfile() -> effectiveDefaultProfile()
    (the record component now owns the plain name)

mvn clean install: 578 tests, 0 failures, 0 errors, BUILD SUCCESS.
2026-08-14 07:02:17 +02:00
Dai Ha 3a7ef0adbd CB-553: enforce maxLoad on explicit-profile spawns (no cap bypass) 2026-08-14 06:50:28 +02:00
Dai Ha 293a305748 Merge CB-551: idle-lead heartbeat
CI / build (push) Successful in 1m15s
CI / contract (push) Successful in 1m21s
2026-08-13 21:21:33 +02:00
Dai Ha e5038c6d13 CB-551: idle-lead heartbeat — nudge an idle lead back to work on a timer
CI / contract (pull_request) Successful in 43s
CI / build (pull_request) Successful in 56s
2026-08-13 21:10:21 +02:00
Dai Ha 13ea79f6fd CB-523 (opencode): generated worker config pins compaction.auto
CI / build (push) Successful in 56s
CI / contract (push) Successful in 1m10s
A worker that exhausts its context dies mid-turn and its bridge_reply — the
whole point of the turn — dies with it. Auto-compaction is a condition of the
turn contract for a spawned peer, not an operator preference.

OPENCODE_CONFIG is merged over ~/.config/opencode/config.json rather than
replacing it, so workers already inherited auto:true from the home file. That
inheritance is exactly what this removes as a dependency: the home file is
outside this repo, differs per machine, and is not ours to rely on.

The trade is recorded in the code: an OPENCODE_CONFIG value overrides the home
value, so an operator cannot disable compaction for bridged workers from home.
Deliberate for peers we spawn and whose turns we must land; per-profile control
would be a profile knob, not the removal of this line.

Origin: fleet-wide auto-compaction audit by peer lead gpt-sol-5.6, which found
autoCompactEnabled:false alongside a 300k window across the Claude instance
configs. Handed over uncommitted; taken deliberately, with rationale added.
2026-08-13 21:03:52 +02:00
Dai Ha 04e21c9243 CB-544: shutdown drain preserves worker worktrees (no data loss)
CI / contract (pull_request) Successful in 41s
CI / build (pull_request) Successful in 1m29s
2026-08-13 20:36:32 +02:00
Dai Ha dd906526c0 CB-548: guard primary singleton to PRIMARY callers; open waiter before enqueue
CI / contract (pull_request) Successful in 42s
CI / build (pull_request) Successful in 53s
Fix two handler-level bugs found in PR #21:
- Only PRIMARY callers may update PrimaryRegistry.record (the legacy singleton 'primary'
  fallback for no-delegation inbox nudges). An architect SEND previously recorded its terminal
  as the fallback; the per-target delegation map does not cure the singleton. New
  BridgeMcp.recordPrimarySingleton uses the resolved role (caller.isPrimary()) — named leads
  (PRIMARY) still record, architects never do.
- MessageService.send now opens the rendezvous waiter BEFORE queueing delivery, fixing both the
  enqueue-before-open fast-reply race (a fast reply no longer orphans into the inbox) and
  callback-failure ordering: a throwing onAccepted (public callback) fails the send cleanly with
  no stale waiter and no queued, orphanable message.

Tests: architect SEND vs lead SEND primary-singleton regression; throwing onAccepted leaves no
stale waiter or queued orphan.
2026-08-13 19:42:48 +02:00
Dai Ha ec3001796a CB-548: record delegator ownership only when a send is accepted
Record PrimaryRegistry delegator ownership via a MessageService accepted-delivery
hook (won the session lock + queued delivery), never at bridge_send request time, so
a concurrent sender that times out BUSY cannot steal a live turn's reply routing.
Make Rendezvous.open atomic fail-if-present so a double open trips loudly instead of
replacing the waiter another send is blocked on. Answering a bridge_ask keeps the same
ownership (no rewrite). Adds ownership/rendezvous regression tests.
2026-08-13 19:42:48 +02:00
Dai Ha cd18887b69 CB-547b: opencode session identity — resume (-s) + post-hoc session discovery
CI / contract (pull_request) Successful in 42s
CI / build (pull_request) Successful in 1m16s
2026-08-13 19:40:17 +02:00
Dai Ha 0114bd1fa7 CB-543: neutralize tracked opencode.json and .autoenv in provisioned worktrees
CI / contract (pull_request) Successful in 45s
CI / build (pull_request) Successful in 1m26s
2026-08-13 19:30:19 +02:00
Dai Ha f004a0c654 CB-548: make duplicate-architect-slot detection top-level-only and depth-safe
CI / contract (pull_request) Successful in 42s
CI / build (pull_request) Successful in 1m13s
2026-08-13 18:31:29 +02:00
Dai Ha 6123576c68 CB-548: correct architect premise — profile-only slots, registry-owned bindings, dup-key rejection
CI / build (pull_request) Successful in 54s
CI / contract (pull_request) Successful in 1m6s
2026-08-13 18:05:53 +02:00
Dai Ha 21cfc09f8e CB-548: config-declared architect slots + Role.ARCHITECT authz
CI / build (pull_request) Successful in 55s
CI / contract (pull_request) Successful in 1m6s
2026-08-13 17:28:14 +02:00
Dai Ha d146a01422 CB-547a: peer-neutral session-identity contract + Claude Code adapter
CI / contract (pull_request) Successful in 44s
CI / build (pull_request) Successful in 1m31s
2026-08-13 16:43:47 +02:00
Dai Ha 5afe8e14d9 CB-542: close the subscription env: bypass, and fix the env: env-carries-anthropic docs
CI / contract (pull_request) Successful in 44s
CI / build (pull_request) Successful in 51s
2026-08-13 15:26:20 +02:00
Dai Ha 73f6b12dd8 CB-539: let a worker profile run on the subscription, explicitly
Add a per-profile subscription: true opt-in that lets a claude-code worker run on
the operator's Claude subscription when there is no off-subscription endpoint for it
(e.g. sonnet on ccs). When set, the launcher injects neither ANTHROPIC_BASE_URL nor
ANTHROPIC_AUTH_TOKEN and skips SubscriptionGuard's base_url requirement for that
profile only, logging a WARN naming the profile. subscription: true alongside a
baseUrl is refused as contradictory. The default (absent/false) keeps today's hard
refusal unchanged; every other profile stays allowlist-checked and SubscriptionGuard
is untouched.
2026-08-13 15:15:14 +02:00
Dai Ha cf48983046 Merge CB-538: opencode peers launch with --auto
CI / contract (push) Successful in 50s
CI / build (push) Successful in 1m30s
A spawned peer has no human at its pane, so an approval prompt is not a pause —
it is a wedge. The agent stops, looks identical to a legitimate mid-turn wait,
and can never reach its bridge_reply, so the delegation dies silently and the
lead learns nothing until the timeout.

Unconditional rather than a per-profile knob, which is the right call: there is
no configuration under which a bridge-spawned opencode worker WANTS to block on
an approval it has no way to answer. opencode's help calls --auto 'dangerous!',
and that warning is written for a human at a terminal; the blast radius here is
already bounded by the layer above — a worker runs in its own git worktree, on
its own branch, off-subscription, and cannot merge. The lead is the gate.

Reviewed by me rather than fanned out: 24 lines across one method and its two
tests, below the threshold where a reviewer pass pays for itself.

argvWithModel is re-signatured to take composed argv instead of building it,
so the two flag-appenders compose rather than each owning construction.
2026-08-13 15:03:59 +02:00
Dai Ha bc13b8e92c CB-530..536, CB-538 groundwork: leads as peers, and a fleet that can find itself
Two leads now work as peers rather than one primary plus workers. The arc:

  CB-530/531  lead identity: `leaders:` names panes, `leadScan:` discovers them by
              tab label (LeadTabScanner, TTL-cached, worker spaces excluded).
  CB-532      leads can message each other AND be answered. Principal.leader now
              carries its terminal, so ownsSession() can be true for a lead; the
              "and you must be a worker" conjunct beside it protected nothing.
              Retires `primary:` — reply nudges follow the delegating lead, a
              binding recorded at bridge_send where both halves are known.
  CB-533      ClaudeCodeLauncher passes --model. argv is usually a wrapper
              (`ccs <profile>`) that re-exports its own model family, so
              ANTHROPIC_MODEL alone was silently overruled.
  CB-534      a lead is deliverable. The CB-113 readiness gate only opened for
              terminals in WorkerPresence, which only workers ever enter, so every
              lead->lead send waited out the ~60s grace and failed having never
              been typed. The gate guards a *spawned* peer's boot window; a lead
              is never spawned.
  CB-535      bridge_list returns `leads` alongside `workers`, with `self` on the
              caller's row. An empty worker roster no longer reads as "no peers".
  CB-536      CLAUDE.md: lead<->lead is coordinate-only, never sideways delegation.
              Propagated byte-identically to wiki/7-Use-Cases.md.

MIXED PROVENANCE — recorded deliberately rather than hidden. This tree also carries
in-progress CB-537 (context separation) authored by the peer lead gpt-sol-5.6 and
its worker: Capability.CONTEXT_RESET, SessionManager.clearAfterTurn, and the
Injector/TurnListener/CompletionResolver/launcher changes around it. That work was
done in this shared working tree rather than a worktree, and is entangled with the
above in BridgedConfig.java, Bridged.java and ClaudeCodeLauncher.java, so neither
lead could stage its own half without sweeping in the other's. Committing the whole
green state is the honest resolution; the peer branches from here.

Note for whoever picks CB-537 up: the design in this commit is SUPERSEDED. Both
leads agreed to replace the global `clearAfterTurn` boolean with per-delivery
policy (inherit|fresh|thread) applied PRE-delivery, because a post-turn reset races
by construction — Injector.onStatus clears awaitingCompletion and dequeues the next
message in the same tick. `fresh` is also a correctness guarantee, so an adapter
without a reset capability must refuse it rather than log a no-op.

mvn clean install: Tests run: 464, Failures: 0, Errors: 0, Skipped: 0. BUILD SUCCESS.
2026-08-13 09:38:24 +02:00
Dai Ha fbe79258bf CB-538: opencode peers launch with --auto 2026-08-13 07:19:19 +02:00
Dai Ha defe3365c4 CB-519: re-point pane assertions at the protocol-19 coordinate
CI / build (push) Successful in 1m22s
CI / contract (push) Successful in 2m32s
Rebase integration only, no behaviour change. CB-519's tests named the pane the
pre-protocol-19 fake produced (w9:pW_n); upstream's herdr 0.8.0 port creates the
pane through tab.create and starts the agent into it, so the fake now reports
w9:pRoot_n. Five assertions were therefore counting closes of a pane that never
existed and reading 0.

mvn clean install: Tests run: 399, Failures: 0, Errors: 0 — BUILD SUCCESS
2026-08-09 05:59:14 +02:00
Dai Ha 4e6201ecd1 CB-525: isolate a worker's tool surface to what its launcher mounts
A worker in a provisioned worktree was inheriting the primary's MCP servers by
two independent routes: the repo commits a .mcp.json declaring the IDE servers,
so a fresh checkout mounts them, and the default parity overlay then copied the
primary's own copy over the top.

Those servers are bound to the primary's IntelliJ project, so every path they
hand back points into the primary's checkout. A CB-523 worker made all 59 of its
edits there while running `mvn -f bridged/pom.xml` against its worktree — every
build it ran was of code that did not contain its changes, and it passed. The
worker's own `ls` of the file it had "edited" returned "No such file".

GitWorktrees now neutralizes .mcp.json at provisioning: an explicitly empty
server map, --skip-worktree'd when tracked so it never reads as pending work a
worker might commit. Unconditional, because the overlay was only half the leak.
The bridge itself is unaffected — it reaches a worker through the launcher's
--mcp-config flag, not the project file, so bridge_reply still works.

- BridgedConfig: .mcp.json out of the default parity overlay
- GitWorktrees: isolateToolSurface() on add(), with the rationale in javadoc
- GitWorktreesTest: 4 real-git acceptance tests (2 fail if the call is removed)
- implementer skill: work from $PWD, and quote a green unpiped `mvn clean
  install` from the worktree as the acceptance criterion

mvn clean install: Tests run: 392, Failures: 0, Errors: 0 — BUILD SUCCESS
2026-08-09 05:57:30 +02:00
Dai Ha ccf50f950e CB-524: make worker placement order deterministic across JVM runs
The weighted policy breaks an exact-weight tie on candidate list order
(WeightedRoundRobinPolicy picks the first candidate with a strictly greater
score), and that list comes from CompositePeerLauncher.candidates(), which
iterates profileConfigs. Both that map and BridgedConfig.workerProfiles() were
built with Map.copyOf, whose iteration order is salted per JVM run — so the
"in definition order" contract candidates() documents was not held.

Two consequences. In production, a config with equal weights (ollama 0.5 /
gx10 0.5) placed its first worker on a profile chosen at random on every daemon
restart. In the suite, CompositePeerLauncherTest.failoverRetriesNextCandidate-
WhenProfileIsUnreachable failed roughly one run in four, because whether
profile "a" was tried first depended on the salt.

Preserve definition order at every layer: unmodifiable LinkedHashMap for
workerProfiles(), profileConfigs, and byProfile (which also feeds the
user-visible bridge_profiles listing). The tests build profile maps with an
ordered helper rather than Map.of, which is salted for the same reason.

Guarded by a pair of tests declaring the same two profiles in opposite order
and asserting opposite first attempts, so any order-scrambling implementation
must fail one of them. Verified by mutation: reverting profileConfigs to
Map.copyOf fails 8/8 runs (6 caught by the original test, 2 only by the new
reversed-order one); with the fix, 10/10 fresh JVMs pass, 388 tests green.
2026-08-09 05:57:30 +02:00
Dai Ha a7f0211e2f CB-518: weighted placement policy for worker spawns 2026-08-09 05:57:30 +02:00
Dai Ha 049ce4828c CB-520: split ReplyInbox into explicit own/release and publish halves 2026-08-09 05:57:12 +02:00
Dai Ha c1173346ef CB-521: make the AMQP contract test runnable locally and in CI 2026-08-09 05:56:45 +02:00
Dai Ha 7ace184fe6 CB-519: make PeerHandle.id() a host-unique opaque UUID, decoupled from the herdr pane id 2026-08-09 05:56:45 +02:00
Dai Ha 54b314ace5 CB-522: let the primary run inside a herdr pane
Caller identity resolved any loopback PID that mapped to a herdr pane as a
WORKER, and PaneLocator scans every pane -- not just bridged-spawned ones. A
primary running inside a herdr pane therefore classified itself as a worker and
was refused SPAWN/SEND/STOP, i.e. every orchestration verb it exists to call.

The failure is self-locking: PrimaryRegistry only learns the primary's terminal
from bridge_send/bridge_spawn, the exact calls being refused, so the learned
value can never bootstrap. Only an operator-set pin breaks the cycle.

CallerResolver now consults primary.terminal from config *before* the pane
lookup. Deliberately the pinned value only, never the learned one -- the learned
terminal is populated by the callers this method is itself classifying, so
trusting it would be circular. Config is operator input, never network input, so
this widens no attack surface; bridge_whoami and the authz gate still share one
resolution.

Fixing that exposed a second, older bug. BridgeMcp's context extractor forwards
the caller's terminal into markPresent on every MCP call, documented as "no-op
for the primary (null terminal)". WorkerPresence.markPresent honours that, but
PresenceBridge overrides it and forwards the same null into SessionManager.
onReady -> transitionByTerminal -> findByTerminal, which called
terminalId.equals(...) unguarded. It only reached the scan once the registry was
non-empty, so the primary's first spawn succeeded and every later call NPE'd
with an HTTP 500 -- and it would have fired for ANY primary not living in a
herdr pane, pinned or not.

findByTerminal is now total. That covers onReady, onDelivered, onTurnComplete
and onTurnFailed at once; a null id could never match a registered session
anyway, so "no match" is the honest answer rather than taking down an unrelated
tool call.

Also drops two dead pass-throughs on CallerResolver (cwdForPid, tokenMode) that
IDE inspections flagged -- callers use ConnectionIdentity and BridgedConfig.Auth
directly.

The example config now states that primary.terminal is REQUIRED, not just a
push-loop optimisation, when the primary shares a herdr pane.

mvn clean install: 360 tests, 0 failures. Verified live: daemon restarted on
this jar, bridge_whoami reports primary, and four concurrent worktree spawns --
the exact shape that NPE'd -- now all succeed.
2026-08-09 05:55:12 +02:00
kevin 224b344445 CB-522: resolve the pinned primary.terminal pane as the primary, not a worker
CI / build (push) Successful in 2m54s
A primary running INSIDE a herdr pane was resolved as a worker by the
pane-match rule and refused every orchestration tool — the exact lockout
bridge_whoami surfaced on this deployment. The CB-307 primary.terminal pin
always claimed to replace connection-derived identity but only fed the push
loop; it now short-circuits CallerResolver ahead of the pane→worker rule
(the pane mapping is as unforgeable as a worker's, so no credential needed,
even in token mode). bridged.example.yaml documents the block.

Also guard the presence bridge against the primary's null terminal: the MCP
context extractor marks presence on every request, and the first genuine
primary contact NPEd into the SPAWNING→READY transition.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01SUTLvxtRPr2iT5u5g45BEs
2026-08-08 21:52:38 +07:00