Commit Graph

128 Commits

Author SHA1 Message Date
Dai Ha 33a6e77f0e Merge CB-573: dormant fleet health monitor (M4 unit 1)
CI / build (push) Successful in 49s
CI / contract (push) Successful in 1m14s
An opt-in whole-fleet observer, separate from the 250ms delivery poller.
One AgentControl.list and one roster snapshot per tick, joined and fed to
the FleetHealth classifier, because a fault is a disagreement between the
two views at the same instant. Absent a health: block nothing is built
and no herdr call is made.

Adds bridge_list healthCoverage: off, detection-only, or full. Detection
is deliberately separate from notification, so a single-lead setup with
no webhook still gets detection and is told its coverage is partial
rather than being refused.

Two review fixes worth naming. tick() rescheduled itself as its last
statement with no try/catch, and a ScheduledExecutorService does not
re-run a task that threw — so the first agents.list failure would have
stopped health permanently and silently, which is exactly when the
control link is down. It now catches Throwable and reschedules in a
finally. And the snapshot fields this unit cannot supply are the named
constant NOT_YET_OBSERVED rather than bare false literals, because false
means no fault to this classifier.
2026-08-15 06:19:15 +02:00
Dai Ha c76b2f149e CB-573: keep health monitoring after failures
CI / contract (pull_request) Successful in 42s
CI / build (pull_request) Successful in 1m15s
2026-08-15 06:17:28 +02:00
Dai Ha 826fffe05b CB-573: add dormant fleet health monitor 2026-08-15 06:15:24 +02:00
Dai Ha 75f57cdba7 CB-568c: fail every queued async ticket
CI / contract (pull_request) Successful in 44s
CI / build (pull_request) Successful in 1m34s
2026-08-15 06:15:21 +02:00
Dai Ha f556af5d4e CB-568: fail queued async tickets on teardown 2026-08-15 06:14:28 +02:00
Dai Ha d5f33f0c6e Merge CB-568: a dropped send reports the real cause
CI / build (push) Successful in 55s
CI / contract (push) Successful in 1m0s
Injector.drop knew the precise cause (herdr agent_not_found) but the
sender was told only 'worker unreachable or stuck', so a lead could not
tell a dead pane from a stalled model.

TurnListener.onTurnFailed gains a reason, defaulting to the old one-arg
form. CompletionResolver prefers that reason, then the pane scrape, then
the old fixed text.

drop now fires onTurnFailed unconditionally. That is the substantive
fix: the sender blocks on the rendezvous waiter, not on the delivered
future, so failing delivered() alone never woke it and a queued send sat
until its timeout.
2026-08-15 06:08:41 +02:00
Dai Ha 16e17b32ad CB-568: preserve dropped turn causes
CI / contract (pull_request) Successful in 43s
CI / build (pull_request) Successful in 1m26s
2026-08-15 06:06:07 +02:00
Dai Ha 46fa4f38d5 CB-575: filter MCP cancellation warnings
CI / build (pull_request) Successful in 49s
CI / contract (pull_request) Successful in 1m1s
2026-08-15 05:57:45 +02:00
Dai Ha 8a53d5bfc6 Merge CB-574: an async delegation can now receive a worker's question
CI / contract (push) Successful in 1m2s
CI / build (push) Failing after 1m40s
A worker on a wait:false delegation called bridge_ask and the lead never
saw the question. Outcome.QUESTION is deliberately non-terminal, but
taskView tested r.completed() and fell into the failure branch, so the
ticket was marked FAILED and both the question text and its turnId were
discarded. The worker blocked for 55s, gave up, and had to abandon its
task. CLAUDE.md tells leads to prefer wait:false and to answer an ask with
bridge_send{turnId, content}; those two could not both be followed.

bridge_poll now returns a non-terminal ASKING phase carrying the question
and its turnId, and the ticket stays live so the worker's real reply still
lands on it. An unanswered ask returns the ticket to PENDING, because only
the question wait ended - the delegated turn continues. The 55s/115s ask
caps are unchanged: they exist because the worker's own MCP call would time
out, so widening them would only move the failure.

Two defects found reviewing the first revision, both from replacing
supplyAsync with a manually completed future:

- an exception inside the send left the future uncompleted, so the ticket
  stayed PENDING for the life of the daemon. Now caught and completed
  exceptionally.
- correlation was keyed by target, one entry per worker, registered before
  the session lock. With two tickets outstanding on one target the second
  overwrote the first, so a late reply could resolve the wrong ticket.
  Correlation is now per turn, the target entry exists only while that send
  owns the lock, and a reply with no live waiter still goes to the durable
  inbox as before.
2026-08-15 05:50:45 +02:00
Dai Ha abd26c796b CB-574: retain async task correlation
CI / build (pull_request) Failing after 1m14s
CI / contract (pull_request) Successful in 1m15s
2026-08-15 05:49:27 +02:00
Dai Ha 24559d81ac CB-573: require explicit capacity source
CI / contract (pull_request) Successful in 44s
CI / build (pull_request) Successful in 55s
2026-08-15 05:49:04 +02:00
Dai Ha 6c1c2c3994 CB-573: report empty configured profile capacity
CI / build (pull_request) Successful in 59s
CI / contract (pull_request) Successful in 1m0s
2026-08-15 05:45:39 +02:00
Dai Ha 01fab15713 CB-573: add fleet capacity view
CI / contract (pull_request) Successful in 1m1s
CI / build (pull_request) Successful in 1m25s
2026-08-15 05:43:46 +02:00
Dai Ha 0cd00e71c3 CB-574: surface async worker questions
CI / contract (pull_request) Successful in 42s
CI / build (pull_request) Successful in 56s
2026-08-15 05:43:43 +02:00
Dai Ha bf0ff2adbf CB-573: keep active delegations out of idle
CI / contract (pull_request) Successful in 45s
CI / build (pull_request) Successful in 1m27s
2026-08-15 05:40:37 +02:00
Dai Ha ed4bbc1c56 CB-573: add pure health classification model
CI / build (pull_request) Successful in 59s
CI / contract (pull_request) Successful in 1m15s
2026-08-15 05:37:30 +02:00
Dai Ha 4f0bf667b1 CB-572: require profiles for send validation
CI / build (pull_request) Successful in 1m0s
CI / contract (pull_request) Successful in 1m2s
2026-08-15 05:33:15 +02:00
Dai Ha 61af9aa574 CB-572: reject profile names as send targets
CI / build (pull_request) Successful in 59s
CI / contract (pull_request) Successful in 1m20s
2026-08-15 05:29:37 +02:00
Dai Ha 65b38997f7 Merge CB-570: OpenCode receives the composed charter (PR #40)
One file, member-charter.md, not two. Two files would have made the U5
digest non-comparable between the Claude adapter and this one, which is
the whole point of the receipt; and nobody verified how OpenCode merges
multiple instruction files, so array order was an unverified dependency.

The OPENCODE_CONFIG condition widens to include a charter. It used to be
hasMcp() || hasCustomProvider(cfg), so a profile with a role charter but
no MCP and no custom provider would have got no config file and therefore
no charter — the feature silently doing nothing for that profile.

The file stays in the per-spawn temp dir, never the worktree: the
worktree is removed on release, the parity overlay already writes into
it, and CB-525's lesson was that config the bridge copied into a worktree
made a worker operate on the wrong tree. Being outside the repo is also
what stops it being committed, which a .gitignore line does not.
2026-08-15 05:12:26 +02:00
Dai Ha 619792a81c CB-570: deliver composed OpenCode charter
CI / build (pull_request) Successful in 58s
CI / contract (pull_request) Successful in 1m13s
2026-08-15 05:10:49 +02:00
Dai Ha cea1183f75 CB-569: pass composed charter to Claude
CI / contract (pull_request) Successful in 43s
CI / build (pull_request) Successful in 50s
2026-08-15 05:09:30 +02:00
Dai Ha e81944cef6 CB-567: compose charters per spawn
CI / build (pull_request) Successful in 57s
CI / contract (pull_request) Successful in 1m3s
2026-08-15 05:00:48 +02:00
Dai Ha 799668e129 CB-566: add fleet charter config
CI / contract (pull_request) Successful in 43s
CI / build (pull_request) Successful in 55s
2026-08-15 04:50:23 +02:00
Dai Ha f8522edacd Merge CB-564: give every silent failure path a voice
CI / build (push) Successful in 1m0s
CI / contract (push) Successful in 1m15s
The bridge cannot classify what it never emits. Fleet-health monitoring —
detect a wedged member, decide, escalate — is blocked on that, so this is the
foundation rather than the feature.

A survey of Injector, CompletionResolver, StatusPoller, SessionManager,
SessionReaper and MessageService found six conditions that ended a member's
usefulness while saying nothing useful:

  * Injector.drop()             — worker gone, queue cleared: SILENT
  * CompletionResolver.fail()   — "via turn-stall fallback" at DEBUG, no reason
  * SessionManager.onFailed()   — "session marked failed" at DEBUG, no stage
  * SessionManager.reapIdle()   — indistinguishable from any other release
  * SessionManager.acquire()    — spawn failure rethrown with no log at all
  * MessageService.abandon()    — failed a caller's request at DEBUG

The first two are the exact phrases that misled the CB-560 diagnosis: both name
a symptom and neither names a cause. They are now WARN and carry the reason,
the stage, and the counts.

Conditions already loud were left alone, and so were two by-design timeouts in
MessageService — an async model exists precisely for those, and promoting them
would turn healthy operation into noise.

Behaviour is unchanged: every edit is a log statement.

Merge note: the recycle test deleted by CB-565 conflicted with a test added
here. Resolved by keeping the new onTurnFailed assertion and dropping the
recycle test, which tests a method that no longer exists.
2026-08-15 04:36:09 +02:00
Dai Ha 078bde2c02 CB-564: give a voice to silent member-failure paths
Injector.drop, CompletionResolver.fail, SessionManager.onFailed/reapIdle/
acquire spawn failures, and MessageService.abandon used to fail a member
or a caller's request with no log, a bare DEBUG, or a log that named only
the symptom ("session marked failed", "failed send via turn-stall
fallback"). Each now logs at WARN and names the real cause and the
numbers involved. Observability only — no behaviour changed.
2026-08-15 04:31:42 +02:00
Dai Ha 0b10ea987b CB-565: remove unsafe session recycle
CI / contract (pull_request) Successful in 44s
CI / build (pull_request) Successful in 55s
2026-08-15 04:29:50 +02:00
Dai Ha d04b075996 CB-563: mark clipped completion scrapes
CI / contract (pull_request) Successful in 42s
CI / build (pull_request) Successful in 56s
2026-08-15 04:25:54 +02:00
Dai Ha 644927636d Merge CB-562: the readiness gate says why it gave up (PR #33)
CI / build (push) Successful in 1m1s
CI / contract (push) Successful in 1m14s
When the injector's readiness grace expired it cleared the queue and logged
nothing. The failure then surfaced elsewhere as a turn-stall, which names the
wrong cause. Diagnosing CB-560 cost two live spawns and a wrong first
hypothesis for exactly this reason: the logs said "session marked failed" and
"failed send via turn-stall fallback", and neither says the message was never
typed into the pane at all.

The expiry now logs the target, the number of messages being failed, the grace
in polls and seconds, and the real cause in plain words.

The grace in seconds is derived, not written down twice: Bridged's own
INJECT_POLL_MILLIS is deleted and Injector.POLL_INTERVAL_MILLIS is the single
source, passed to every StatusPoller. A cadence change can no longer leave a
log line confidently stating the wrong duration.

Behaviour is unchanged. This is the first structured health event in the
daemon, and the foundation the fleet-health work will build on.
2026-08-15 04:21:29 +02:00
Dai Ha 95e45007aa CB-562: single source for the injector poll cadence; tighten count assertion
CI / build (pull_request) Successful in 59s
CI / contract (pull_request) Successful in 1m17s
2026-08-15 04:20:29 +02:00
Dai Ha 7a583c4045 CB-562: log why the readiness gate gave up on a target
CI / build (pull_request) Successful in 57s
CI / contract (pull_request) Successful in 1m4s
2026-08-14 21:52:25 +02:00
Dai Ha 65bce058bb Merge CB-561: one public way to build a CallerResolver (PR #32)
CI / build (push) Successful in 1m17s
CI / contract (push) Successful in 1m16s
CallerResolver had 5 public constructors and 4 public factories, and only one
of them could ever produce an architect. The rest defaulted memberSlotRoles to
`_ -> null`, so every architect quietly fell through to Principal.worker().
Nothing logged, nothing threw — the role was simply off.

That is the same failure shape as CB-560, so the fix is structural rather than
a warning: withLeadsAndMembers(.., MemberRegistry) is now the only public
construction path. Two overloads with no caller at all are deleted; the rest
are package-private and marked test-only. No path remains that accepts
architect bindings without a slot-role lookup, so no runtime WARN is needed.

Also checked and closed: the suspected slot leak on shutdown drain is not
real. SessionManager.release calls memberLifecycle.released() for every
removed session, and drainAll routes every session through release,
SPAWNING included. Verified in code.
2026-08-14 21:50:04 +02:00
Dai Ha fa0612859b CB-560: document spawned member presence
CI / build (pull_request) Successful in 1m0s
CI / contract (pull_request) Successful in 1m4s
2026-08-14 21:48:58 +02:00
Dai Ha 4ffbcd0b7d CB-561: require member registry for architects
CI / contract (pull_request) Successful in 43s
CI / build (pull_request) Successful in 50s
2026-08-14 21:47:53 +02:00
Dai Ha a0cd053fd9 CB-560: mark architect members present
CI / build (pull_request) Successful in 55s
CI / contract (pull_request) Successful in 1m4s
2026-08-14 21:45:28 +02:00
Dai Ha e9bc192160 CB-548: bind architect sessions to member slots
CI / contract (pull_request) Successful in 45s
CI / build (pull_request) Successful in 55s
2026-08-14 21:10:48 +02:00
Dai Ha b414a74c26 CB-559: drop .claude/settings.local.json from the default parity overlay
CI / build (pull_request) Successful in 54s
CI / contract (pull_request) Successful in 1m7s
2026-08-14 20:56:16 +02:00
Dai Ha e18e002d2f CB-559: a profile's launch settings are deferred, not hot
The shipped docs and javadoc said a profile's `model` and `tabLabel` take effect on
the next spawn. They do not, and ConfigRef did not detect the change either, so a
reload logged a clean "config reloaded" and silently did nothing. That is the worst
outcome a reload can produce: the operator has no reason to doubt it.

What makes a key hot is who reads it and when, not that it is config. Placement
reads weight and maxLoad through a supplier on CompositePeerLauncher, so those are
genuinely hot. HerdrPeerLauncher takes Map.copyOf(profiles) at construction and
resolves every spawn out of that copy, so model, baseUrl, argv, env and the rest
cannot move until the daemon restarts.

changedDeferredKeys now compares every launch component of an existing profile,
excluding weight and maxLoad, and names the profiles that need a restart. The
javadoc and bridged.example.yaml say the same thing. Two tests pin the pair:
weight/maxLoad reports nothing deferred, a changed model reports the profile by name.
2026-08-14 20:37:42 +02:00
Dai Ha a2108a8a14 CB-559: re-read bridged.yaml without restarting the daemon
Tuning a fleet meant restarting bridged, and a restart tears down every lead
and worker it owns. Changing one pool's weight cost the whole fleet's state,
so in practice nobody changed it.

ConfigRef holds the live BridgedConfig in an AtomicReference. Consumers read
it at the point of use, so a change reaches the next spawn with nothing
rebuilt. The launchers that used to capture config into fields now take
suppliers: the fleet tabLabel template, the profile map, the placement policy
and the fleet block.

Keys fall into three classes, and the difference is what already exists when
the reload happens:

  hot       fleet: (pools + tabLabel), placement:, and an existing profile's
            weight / maxLoad / model / tabLabel — live on the next spawn.
  deferred  lifecycle:, leadHeartbeat:, guard:, worktreeRoot:, spawnReady*,
            and adding/removing a profile — accepted, but the startup wiring
            keeps the old value. The reload logs these by name.
  cold      bind:, herdrSocket:, broker:, auth: — refuses the WHOLE reload.

A cold change refuses everything rather than applying the hot half. A
half-applied reload leaves the daemon matching no file on disk, which is the
worst thing a reload can do to an operator reading that file to work out what
the daemon is doing. Refusing keeps the invariant that the live config is
always some version of the file.

A parse failure or a failed startup validator is refused the same way, and
the running config stays live: a file being saved is sometimes read
mid-write, and degrading a working daemon over a half-written file is a bad
trade. The same four validators startup runs are re-run, so a config that
could not have booted cannot slip in through a reload.

ConfigWatcher polls the modified time on a daemon thread, opt-in through
configReload.enabled (default off, so an upgraded daemon is unchanged). It
stamps the timestamp BEFORE reloading, so a refused file is not retried every
tick — the next save earns a fresh attempt. A missing file is skipped
silently, because editors unlink briefly mid-save.

MicroProfile Config was the first idea and does not fit: @ConfigMapping needs
interfaces, resolves once at bootstrap, and reload would still mean rebuild
and swap. The port would also lose the raw-YAML duplicate-key detection,
since duplicates have already collapsed once the tree is flattened to
properties.

634 tests.
2026-08-14 18:23:32 +02:00
Dai Ha 57f8fa257a CB-558: launch a declared lead at startup when none is live
`fleet.leaders.<name>.instances` was descriptive. Now the daemon reads it: a
lead that names a `profile:` is started when fewer than `instances` are running.
A lead with only a `terminal:` stays recognise-only, as before.

A lead is not a member, and LeadLauncher exists to keep it that way. Every other
spawn path goes through HerdrPeerLauncher, which does three things a lead must
never get: it appends the worker reply charter ("you are an off-subscription
worker … end every turn with bridge_reply" — the opposite of an orchestrator);
it registers the session with SessionManager, whose idle reaper would kill a
lead for being idle, which is a lead's normal state; and it can move a peer off
the subscription. So this launcher talks to AgentControl/WorkspaceControl
directly. The duplicated argv/env assembly is the cheaper half of that trade.

Not double-spawning is the safety property, so liveness needs two pieces of
evidence. A running agent in a tab labelled `lead: <name>` finds an
auto-launched lead. A running agent on a pinned `terminal:` finds one the
operator opened by hand — without it, a pinned lead whose tab carries no
matching label would be relaunched on every boot. Member workspaces are
excluded, so a member in a matching tab is never counted. If herdr cannot be
reached, nothing is started: a second orchestrator is worse than none.

Liveness deliberately requires the AGENT, not just the label. LeadTabScanner
used to promise that bridged never writes a lead label, so there was no
round-trip from the daemon's own rename back into its next decision. That is no
longer true, and its javadoc now says so. The trust direction is unaffected — a
label is a name, not a capability — but staleness becomes real: a label left by
a crashed session would otherwise read as a live lead forever and disable
auto-launch permanently.

Two new knobs. `workspace:` (default "leads") is where a launched lead's tab
goes; it must not be a member workspace, because those are excluded from the
scan and a lead placed in one would never be found again. `cwd:` defaults to
bridged's own working directory.

Also: WorkspaceControl.listTabs, and a FakeHerdr tab seeder that leaves the
canned response byte-identical when no tab is seeded.

617 tests pass (16 new), IDE-clean.
2026-08-14 16:47:57 +02:00
Dai Ha 61944fc045 CB-557: place an unqualified spawn inside its role's pool
The pools were config-only until now: the launchers still received one global
effectiveDefaultProfile and placement still ranged over every configured
profile, so a reviewer could be placed on an architect-only backend.

Three parts:

SessionManager computed the role, stored it on the MemberSession, and never put
it on the SpawnRequest. So the role reached the record that describes the spawn
but not the call that performs it — every launcher saw DEV. Both spawn paths
(plain and worktree) now carry it.

CompositePeerLauncher takes the Fleet and draws its candidates from
fleet.<role> instead of from all profiles. An absent or empty pool means
unconstrained, not blocked: a config that declares pools for some roles must
keep spawning the rest, so it falls back to every profile. A null Fleet is the
pre-CB-557 wiring and behaves exactly as before.

Bridged passes cfg.fleet() to the composite and cfg.fleet().tabLabel() to both
launchers. The tab-label knob was accepted by HerdrPeerLauncher but passed by
nobody, so it was inert — the label only looked right because the fallback
happened to match the configured template. Four tests now pin the wiring
instead of the coincidence.

An EXPLICIT profile stays exempt from the pool. `bridge_spawn{profile:"opus"}`
carries no role, so it defaults to DEV; judging it against the dev pool would
refuse a spawn the operator asked for by name. maxLoad still applies to it.

Also cleared the IDE warnings in the touched files: an immediately-rethrown
catch (the comment stays, the redundant block goes), unused lambda params, a
javadoc link to a package-private class, two unused imports.

601 tests pass.
2026-08-14 16:37:29 +02:00
Dai Ha 4b48d2d921 CB-557: fleet role pools — role is the config key, and the tab label says it
Four top-level keys (leaders:, members:, leadScan:, defaultProfile:) become one
`fleet:` block, and a member's role becomes the map key that contains it rather
than a `role:` field inside it.

Why the key and not a field: a misspelled `role: architct` used to produce a
member with no contract, which nothing rejected. A misspelled pool name declares
nothing, which is a shape the loader can see.

`fleet.architects/developers/reviewers` are pools of profiles a role MAY run on.
That replaces the single global `defaultProfile:`, so an unqualified spawn now
resolves its profile from the pool of the role it asked for. Role and profile
stay orthogonal: a reviewer may run on the same profile as the dev it reviews,
and one profile may appear in several pools.

Tab labels are role-first — `dev: sonnet #4`. The template lives on `fleet:`
because a profile cannot know the role of the member launched on it; a profile
may still override it. The `{n}` counter is scoped per role+profile, so a dev
and a reviewer on one profile each start at #1. Making {role} the first field
also turns the lead/member namespace check into a structural guarantee: roles
are a closed enum, so only hand-written templates can still collide with a lead
tabPrefix.

Removed keys are hard errors that name their successor. `defaultProfile:` has no
single successor key, so its message explains the new model instead of pointing
at a key that does not exist.

Map order is kept with LinkedHashMap, deliberately not Map.copyOf — the latter
salts iteration order per JVM run, which would destroy the YAML definition order
that `placement: fixed` selects on.

Not yet wired: SessionManager still hands the launchers one effectiveDefault-
Profile, so pools are not enforced at spawn time yet, and placement still ranges
over all profiles.

595 tests pass.
2026-08-14 16:32:54 +02:00
Dai Ha 4875127daa CB-557: adopt the member taxonomy in the API surface
Every spawned peer is now a member with a role, and the role travels with it
from the spawn call to the roster.

MCP:
  bridge_spawn gains role: architect | dev | reviewer (default dev). An
    unknown role is refused with the valid spellings in the message.
  bridge_list returns "members" instead of "workers"; each row carries both
    role (what it is for) and profile (which backend it runs on).
  The spawn result echoes the role back, so a spawn that fell back to dev is
    visible rather than silent.

REST:
  GET/POST /members and DELETE /members/{paneId} replace /workers.
  POST accepts role= as a query param or a body field; an unknown role is 400.

Code:
  dev.ltms.bridged.worker package -> dev.ltms.bridged.member
  WorkerSession   -> MemberSession, plus a MemberRole role component
  WorkerPresence  -> MemberPresence
  SessionManager.acquire gains a role parameter; the existing overloads keep
    working and default to DEV, which is exactly what "worker" used to mean.

ClaudeCodeLauncher and OpenCodeLauncher keep their names on purpose — they
are named after the backend, not the role.

Not done here: the launch charter is still one string for every role, so a
member is told its role by nobody yet. That is the next ticket.

mvn clean install: 583 tests, 0 failures, 0 errors, BUILD SUCCESS.
2026-08-14 07:08:26 +02:00
Dai Ha 246f50b778 CB-557: adopt the member taxonomy in config
A member is anything a lead spawns. Every member carries two independent
attributes:

  role    — which contract: architect, dev or reviewer. It picks the launch
            charter, the role file, the playbook skill and the authz row.
  profile — which backend: model, CLI adapter, credentials, cost.

They vary on their own. A reviewer may run on the same profile as the dev
whose diff it reads, which is the case that proves the two cannot be one
field.

Config changes (breaking — we are in active development, so no aliases):

  workers:       -> profiles:        it was never a list of workers; it is a
                                     catalogue of backends
  defaultWorker: -> defaultProfile:
  architects:    -> members:         each slot now names its role

An old config is rejected at load with the new key named, rather than being
warned about once and then running with zero profiles — that failure would
surface much later, at the first spawn, pointing nowhere near the cause.

Also:
  - BridgedConfig.Worker    -> BridgedConfig.Profile
  - ArchitectRegistry       -> MemberRegistry
  - new peer.MemberRole enum, validated at startup
  - profiles map is normalized once in the compact constructor, so the raw
    map and the derived one can no longer disagree
  - the legacy singular worker: block is dropped
  - workerProfiles() -> profiles(); defaultProfile() -> effectiveDefaultProfile()
    (the record component now owns the plain name)

mvn clean install: 578 tests, 0 failures, 0 errors, BUILD SUCCESS.
2026-08-14 07:02:17 +02:00
Dai Ha 3a7ef0adbd CB-553: enforce maxLoad on explicit-profile spawns (no cap bypass) 2026-08-14 06:50:28 +02:00
Dai Ha 293a305748 Merge CB-551: idle-lead heartbeat
CI / build (push) Successful in 1m15s
CI / contract (push) Successful in 1m21s
2026-08-13 21:21:33 +02:00
Dai Ha e5038c6d13 CB-551: idle-lead heartbeat — nudge an idle lead back to work on a timer
CI / contract (pull_request) Successful in 43s
CI / build (pull_request) Successful in 56s
2026-08-13 21:10:21 +02:00
Dai Ha 13ea79f6fd CB-523 (opencode): generated worker config pins compaction.auto
CI / build (push) Successful in 56s
CI / contract (push) Successful in 1m10s
A worker that exhausts its context dies mid-turn and its bridge_reply — the
whole point of the turn — dies with it. Auto-compaction is a condition of the
turn contract for a spawned peer, not an operator preference.

OPENCODE_CONFIG is merged over ~/.config/opencode/config.json rather than
replacing it, so workers already inherited auto:true from the home file. That
inheritance is exactly what this removes as a dependency: the home file is
outside this repo, differs per machine, and is not ours to rely on.

The trade is recorded in the code: an OPENCODE_CONFIG value overrides the home
value, so an operator cannot disable compaction for bridged workers from home.
Deliberate for peers we spawn and whose turns we must land; per-profile control
would be a profile knob, not the removal of this line.

Origin: fleet-wide auto-compaction audit by peer lead gpt-sol-5.6, which found
autoCompactEnabled:false alongside a 300k window across the Claude instance
configs. Handed over uncommitted; taken deliberately, with rationale added.
2026-08-13 21:03:52 +02:00
Dai Ha 04e21c9243 CB-544: shutdown drain preserves worker worktrees (no data loss)
CI / contract (pull_request) Successful in 41s
CI / build (pull_request) Successful in 1m29s
2026-08-13 20:36:32 +02:00
Dai Ha dd906526c0 CB-548: guard primary singleton to PRIMARY callers; open waiter before enqueue
CI / contract (pull_request) Successful in 42s
CI / build (pull_request) Successful in 53s
Fix two handler-level bugs found in PR #21:
- Only PRIMARY callers may update PrimaryRegistry.record (the legacy singleton 'primary'
  fallback for no-delegation inbox nudges). An architect SEND previously recorded its terminal
  as the fallback; the per-target delegation map does not cure the singleton. New
  BridgeMcp.recordPrimarySingleton uses the resolved role (caller.isPrimary()) — named leads
  (PRIMARY) still record, architects never do.
- MessageService.send now opens the rendezvous waiter BEFORE queueing delivery, fixing both the
  enqueue-before-open fast-reply race (a fast reply no longer orphans into the inbox) and
  callback-failure ordering: a throwing onAccepted (public callback) fails the send cleanly with
  no stale waiter and no queued, orphanable message.

Tests: architect SEND vs lead SEND primary-singleton regression; throwing onAccepted leaves no
stale waiter or queued orphan.
2026-08-13 19:42:48 +02:00
Dai Ha ec3001796a CB-548: record delegator ownership only when a send is accepted
Record PrimaryRegistry delegator ownership via a MessageService accepted-delivery
hook (won the session lock + queued delivery), never at bridge_send request time, so
a concurrent sender that times out BUSY cannot steal a live turn's reply routing.
Make Rendezvous.open atomic fail-if-present so a double open trips loudly instead of
replacing the waiter another send is blocked on. Answering a bridge_ask keeps the same
ownership (no rewrite). Adds ownership/rendezvous regression tests.
2026-08-13 19:42:48 +02:00