Compare commits

...

118 Commits

Author SHA1 Message Date
Dai Ha 7930a31b94 CB-596 round 2: the exec-time argv-prefix fix has no seam — stop and report
CI / contract (pull_request) Successful in 47s
CI / build (pull_request) Successful in 1m6s
herdr protocol 19's agent.start takes a fixed `kind` (herdr resolves the
executable) plus trailing CLI args for that binary; only tab.create/pane.split
accept an env map, and that IS the round-1 pane-creation overlay already
shipped. There is no argv/env control point that runs after the pane's login
shell and before the agent process starts, so the proposed `env NAME=value ...`
argv prefix cannot be implemented against this API. Documented the finding and
corrected bridged.example.yaml's round-1 comments, which had overclaimed that
the overlay survives the login shell.

Kept everything else: added a startup WARN (Bridged.reportMemberCredentialsGap)
when memberCredentials: is absent or its known: list is empty, so CB-592's
protection loss is never silent, mirroring CB-594's reportRequiredSecrets.
2026-08-17 09:13:57 +02:00
Dai Ha 850fb12807 CB-596: member credential blocking is config-driven deny-by-default, not one hardcoded name
CI / build (pull_request) Successful in 1m1s
CI / contract (pull_request) Successful in 1m26s
Replaces CB-592's single hardcoded GITEA_ACCESS_TOKEN shadow in HerdrPeerLauncher
with BridgedConfig.MemberCredentials (memberCredentials: policy/allow/known).
Every known name not also allowed is overlaid with a sentinel; an unrecognized
policy value refuses at load; a credential-shaped host env var on neither list
is logged as a gap (name only, never a value). bridged.yaml is gitignored, so
the 31 measured names + 4-name allowlist ship as a commented block in
bridged.example.yaml for the operator to apply live.
2026-08-17 09:02:21 +02:00
Dai Ha b14b66ab03 CB-586: the retention sweep never ran — Long.MIN_VALUE overflowed the gate
CI / contract (push) Successful in 1m20s
CI / build (push) Successful in 1m36s
SessionReaper.lastWipSweepNanos started at Long.MIN_VALUE as a "never swept
yet" sentinel. That sentinel cannot be compared by subtraction. nanoTime() is
positive on this platform, so `now - Long.MIN_VALUE` wraps to a large negative
number, the gate `delta < WIP_SWEEP_INTERVAL_NANOS` reads it as "swept moments
ago", and the method returns before the assignment that would have fixed the
field. The sweep never ran once, for the life of the process, and nothing in
the log said so.

Measured:
  System.nanoTime()      = 31305820625625   (positive)
  now - Long.MIN_VALUE   = -9223340731034150183
  interval (6h in nanos) = 21600000000000
  gate 'delta < interval' -> true  => returns early, every iteration, forever

Fix: a separate `sweptOnce` boolean holds "never yet", so the subtraction only
runs once both operands come from nanoTime. The first pass always sweeps — a
restart is a fine moment for it, the 24h age floor keeps it safe, and the
feature becomes observable right after a redeploy instead of six hours later.

The existing tests all passed because they call SessionManager.sweepWipRefs
directly, which walks around the gate. The new test asserts through the reaper
loop instead: it spawns a worktree session, starts the reaper, and requires a
real prune call at the seam with the 24h floor intact. Removing the fix makes
it fail with "it never reached the seam".

860 tests, mvn clean install, BUILD SUCCESS.
2026-08-16 20:13:37 +02:00
Dai Ha 15ff6bcde5 Merge CB-586: prune refs/wip snapshots whose content is already on main
A refs/wip snapshot ref is deleted only when both hold: its commit's tree is
already reachable from main, and it is older than 24h. Reachability is the safety
floor — a snapshot exists because the work was committed nowhere else, so an
unreachable one is the last copy and is never swept. Every deletion logs the ref
and the sha.

/members gains wipRefs{count,costBytes} so the growth is visible.

Verified against real git, not only the fakes: a recoverable+old ref is deleted,
a recoverable+young one survives the age floor, and the last copy survives. A repo
with no main deletes nothing, and a repo with no snapshots is a clean no-op.

Closes #67
2026-08-16 20:09:16 +02:00
Dai Ha 7822772905 CB-582: bridge_status now reports an open question — say so, and keep the warning
CI / contract (push) Successful in 51s
CI / build (push) Successful in 1m7s
The prompt is part of the product: CB-582 changed what bridge_status returns, so
the primary's step 5 was no longer the whole truth.

The second half matters more than the first. A nudge makes a lead more likely to
notice an ask; it does not widen the ~55s window, which is bounded by the
WORKER's own MCP client timeout, not by anything the daemon chooses. Without that
sentence a lead reads 'the ask now nudges me' as 'asking works now' and briefs a
worker to ask — which is the failure CB-582 was filed about.
2026-08-16 19:05:06 +02:00
Dai Ha 0efe1567c0 CB-586: prune refs/wip/* older than 24h whose tree is reachable from main
CI / build (pull_request) Successful in 1m13s
CI / contract (pull_request) Successful in 1m21s
Add the CB-586 retention rule to GitWorktrees and drive it from the reaper:
a snapshot is deleted only when its tree content is already reachable from
main AND the ref is older than 24h. Reachability keeps the last copy of a
worker's work; the age floor stops a fresh snapshot being swept while a
lead is still looking at it. Every deletion logs the ref name and commit
sha so it is recoverable from the reflog. The /members response gains a
wipRefs{count,costBytes} census the operator can read without shelling
into the repo.
2026-08-16 19:02:33 +02:00
Dai Ha 837fed7690 CB-596: the credential probe, as one auditable command
CI / contract (push) Successful in 48s
CI / build (push) Successful in 1m39s
Issue #82 step 1 is a measurement, and the classifier refuses an ad-hoc pipeline
that enumerates credential names inside a member — correctly. This is the seam:
one file the operator reads once and then runs, instead of approving a shell
pipeline they have to take on trust.

It never prints a credential value or any part of one. #82's criterion 1 asked
for a 6-character prefix; this prints a truncated SHA-256 instead. A prefix of a
short secret is most of the secret and would end up pasted into a ticket, while
the hash answers every question the prefix was for — is it set, is it the same
value as over there, is it the CB-592 sentinel.

Refuses to run unless BRIDGED_MEMBER=1, since the finding is what a MEMBER holds;
--allow-outside-member takes the comparison reading and labels it as such.

I have not run the reading path. That is the operator's call, which is the whole
point of the ticket.
2026-08-16 19:01:42 +02:00
Dai Ha aa4ee64a34 Merge CB-606: refuse an unrecognized auth.mode or placement at config load
CI / contract (push) Successful in 51s
CI / build (push) Successful in 1m40s
Three fields had CB-604's shape — lower-cased, compared against one string,
never checked against the valid set. The auth.mode one was the worst: a typo of
'token' silently behaved as loopback-trust, and validateAuthExposure() only fires
on a non-loopback bind, so a loopback bind hid it end to end. The daemon started
clean and authenticated nobody while the operator believed token mode was on.

All three now refuse at config load, naming the value and the accepted set. The
top-level placement policy was already validated but only lazily at first spawn;
it now calls PlacementPolicies.fromName eagerly at load, so a bad name cannot
start a daemon that merely looks healthy.

Verified by probing the real BridgedConfig.load with eleven values, including the
critical auth.mode typo on a loopback bind, and by loading the live gitignored
bridged.yaml — which the worker cannot see and so could not check.

Closes #106
2026-08-16 18:58:49 +02:00
Dai Ha bb750cdba3 CB-606: refuse an unrecognized auth.mode, per-profile placement, or top-level placement policy at config load
CI / build (pull_request) Successful in 1m7s
CI / contract (pull_request) Successful in 1m25s
An auth.mode typo (e.g. "toekn") used to silently fall back to loopback-trust with no signal
anywhere — validateAuthExposure() only checks the pairing on a non-loopback bind, so on the
common loopback bind the daemon started cleanly and authenticated nobody. Per-profile placement
had the same shape, falling back to legacy pane placement. The top-level placement policy name
was already validated by PlacementPolicies.fromName, but only lazily at first spawn through
CompositePeerLauncher's Supplier; it is now checked eagerly at load, calling fromName itself as
the single source of truth.
2026-08-16 18:54:35 +02:00
Dai Ha d56c77b368 CB-582: close the question when ask() leaves by throwing
CI / contract (push) Successful in 48s
CI / build (push) Successful in 1m38s
Found reviewing CB-582 before the merge, not by the implementer.

ask() clears its question on three paths — no-waiter, timed out, and (from
answer()) answered. It can also leave by throwing: an interrupt while blocked on
the answer, or an ExecutionException from the answer future. Those run only the
finally block, which tore down the rendezvous turn but not the push loop's copy.

The result was a question that stayed pending for good: named in every nudge
until it hit its own cap, then left in pendingQuestions with no remover at all.

Teardown now happens where the rendezvous teardown already happens, so the two
cannot drift apart again. Closing a turnId that was never pending is a no-op, so
the normal paths are unaffected.

The new test fails on the pre-fix code with expected: <STOP> but was: <INJECT>.
2026-08-16 18:47:18 +02:00
Dai Ha 83e2ff06cf Merge CB-582: nudge the lead when a worker pauses on bridge_ask
An async (wait:false) delegation opens a ~55s reverse-rendezvous window when its
worker calls bridge_ask. A lead polling on its normal minutes-long cadence never
sees that window, so the worker times out and proceeds without an answer.

The question is now a third source in the CB-588 per-lead push schedule, with its
own per-item nudge count (CB-598's shape), and it is surfaced by bridge_status and
by REST /sessions/{id}/status and /tasks/{ticket} (which previously dropped turnId
on an ASKING phase, so a REST caller could see the question but not answer it).

Closes #61
2026-08-16 18:45:14 +02:00
Dai Ha f5deaafd06 Merge CB-604: refuse an unknown profile kind at config load
CI / build (push) Successful in 1m2s
CI / contract (push) Successful in 1m4s
kind: was lower-cased and compared against one string, so a typo like
'opencod' was accepted and routed to the claude-code adapter. With argv:
unset the launch command became the misspelled string itself, and the
daemon tried to run a program named after the typo. Nothing said a word
until the spawn failed.

It now throws at load, naming the profile, the bad value and the
accepted set - matching rejectNegativeMaxLoad and the duplicate-adapter
check, which already treat routing mistakes as fatal.

Verified here by probing the real BridgedConfig.load with five values:
opencod refused with the full message; opencode, OpenCode, claude-code
and an absent kind all accepted with the right adapter. 832 tests,
BUILD SUCCESS, unpiped.

Closes #102
2026-08-16 18:39:50 +02:00
Dai Ha fe46311266 CB-604: reject an unrecognized profile kind at config load
CI / contract (pull_request) Successful in 53s
CI / build (pull_request) Successful in 1m10s
2026-08-16 18:37:21 +02:00
Dai Ha 08968bb1b7 CB-582: make a pending bridge_ask question visible on the lead's poll cadence
CI / build (pull_request) Successful in 1m23s
CI / contract (pull_request) Successful in 1m23s
bridge_ask blocks the worker's turn for ~55s by default (BridgedApp.java,
BridgeMcp.java) — a value deliberately kept just under the worker's own MCP
client's ~60s call cap so the daemon can return a clean timeout before the
client severs the call, not a value that can usefully be widened. A lead
following the charter's wait:false + poll cadence is minutes away, so the
window closes long before a poll would ever see the question — and until now
bridge_poll on such a ticket just read as ordinary "pending" progress.

bridge_poll(ticket) already surfaced Phase.ASKING with the question and
turnId (CB-205); this ships the two pieces that were still missing:

- The lead's own pane is now nudged the instant a question opens, reusing
  the CB-588 ReplyPushLoop push mechanism (a third source alongside queued
  replies and terminal tickets) rather than a new path. The nudge is capped
  by the loop's existing maxReminders budget, and stops the moment the
  question is answered or lapses.
- bridge_status(sessionId) and REST GET /sessions/{id}/status now also show
  an open question and how to answer it, via a new
  MessageService.pendingAsk() lookup — covering the case where a lead checks
  status directly rather than the ticket.
- The REST /tasks/{ticket} endpoint was silently missing turnId on an ASKING
  phase (only the MCP layer's formatted text carried it) — fixed as part of
  making the state genuinely visible over both surfaces.

An unanswered question still behaves as today: the worker proceeds and its
reply says the ask went unanswered — not a hard failure.
2026-08-16 18:35:45 +02:00
Dai Ha 27bbd11f06 Merge CB-584: carry the agent session id on a failed ticket's detail
CI / contract (push) Successful in 47s
CI / build (push) Successful in 1m23s
Most of issue #65 was already shipped in 5d5b3bd - MemberSession
records agentSessionId, PeerHandle.agentSessionId() has no default, the
roster exposes it, and bridge_spawn accepts sessionName and
resumeSessionId. That commit left one item for follow-up: carrying the
id on ReleaseDetail.

This does that item. CB-578 stage C already lets a lead re-dispatch onto
the same worktree after a failure; without the session id that is a cold
start. With it, the work and the thread both survive.

Also updates the bridge_spawn and bridge_list rows in the CLAUDE.md
intent table, which the earlier commit missed.

Verified here: trial merge onto main builds 830 tests BUILD SUCCESS,
unpiped. Confirmed against the code that items 1-4 really were already
on main, so the scope-down is correct rather than work skipped.

Closes #65
2026-08-16 18:29:24 +02:00
Dai Ha 48d7841fbf CB-589: document that weighted placement is not cheapest-first
CI / contract (push) Successful in 50s
CI / build (push) Successful in 1m43s
The weight ratio does not express a preference order. weighted spreads
spawns across every profile with a free slot, so paid spawns happen
while the free box is idle - and a profile at maxLoad freezes its score,
so it can lose the next pick after a slot frees.

The real fix is a cost-first policy (CB-589). This documents the
workaround and its trap next to the key, because bridged.yaml is
gitignored: a fresh host starts without the workaround and quietly pays,
with nothing to tell the operator why.

Comments only. BridgedConfigTest: 85 tests, BUILD SUCCESS.
2026-08-16 18:27:19 +02:00
Dai Ha cc1df11f69 CB-584: carry agentSessionId on a failed ticket's ReleaseDetail
CI / contract (pull_request) Successful in 1m6s
CI / build (pull_request) Successful in 1m42s
Closes the one piece issue #65 deliberately left out of 5d5b3bd: a
released member's agentSessionId now rides alongside worktree, branch
and snapshotRef on ReleaseDetail, and Bridged's onRelease handler
names it in the abandon reason, so a lead can resume the member's
conversation instead of only re-dispatching a fresh one onto the same
files.

Also updates the CLAUDE.md bridge_spawn/bridge_list table row, which
5d5b3bd shipped the sessionName/resumeSessionId/agentSessionId surface
for but never updated.
2026-08-16 18:25:24 +02:00
Dai Ha fdfd4ac491 Merge CB-600: make installing the launchd agent safe
CI / contract (push) Successful in 51s
CI / build (push) Successful in 1m42s
Three gaps that only bite once the agent is loaded, plus one wrong
comment.

The script computed its log path from its own location while the plist
hard-codes one. Run from a different checkout, every post-restart check
would read the wrong file and report a clean restart while the daemon
crash-looped. It now compares the two and fails, not warns.

A failed 'launchctl load' after a successful 'unload -w' left the agent
stopped AND persistently disabled - worse than before the redeploy. It
now retries once, then dies naming the exact recovery command.

The plist now says plainly that ThrottleInterval paces restarts but does
not bound them, and what actually stops the loop.

Verified here: ran the script with --check from the merged tree and it
behaves exactly as before, so the unsupervised path - my only restart
route - is intact. Exercised the log-path check against match, mismatch
and missing-plist fixtures using a truncated copy with no mutating code
in it: ok/1/1. 829 tests BUILD SUCCESS.

Closes #91
2026-08-16 18:17:11 +02:00
Dai Ha d5dd5639ae Merge CB-602: guard against a config key that never reaches the example
bridged.yaml is gitignored, so bridged.example.yaml is the only
committed description of the config schema. Two tests already covered
example -> code; nothing covered code -> example, so a brand-new key
could ship undocumented and no test would notice.

A new test compares BridgedConfig.KNOWN_TOP_LEVEL_KEYS against the
example scanned as TEXT, so a key documented only as a comment counts as
documented. That is what makes the guard correct rather than annoying:
most of the example is commented on purpose.

Verified here: added an undocumented key and watched the test fail with
an actionable message naming it; then documented that key as a comment
only and watched it pass. Probe reverted, tree clean.

Closes #96
2026-08-16 18:17:11 +02:00
Dai Ha 7a120b3256 Merge CB-598: per-item reminder counts so backoff-window work is never orphaned
CI / contract (push) Successful in 45s
CI / build (push) Successful in 1m39s
The reminder count was one counter per lead per source, carried forward
across ticks. A counter carried forward has no memory of which item it
counted, so work arriving during the backoff window inherited an
already-capped count and was never named in a nudge.

tick() now recomputes each source's count fresh from the minimum count
among the items actually pending, tracked per item. A fresh item keeps
its source eligible; an older capped item still rides along in the text
without spending more budget. decide() is unchanged.

Verified here: read the diff; the bumped set is exactly the set named in
the nudge, and the empty early-return skips the bump. Trial merge onto
main builds 826 tests BUILD SUCCESS, unpiped.

Closes #87
2026-08-16 18:13:29 +02:00
Dai Ha 863d477966 CB-603: make FakeHerdr.calls thread-safe
Background loops call the fake from their own scheduler threads while a
test polls called() from the test thread. The list was a plain
ArrayList, so a nudge landing mid-stream threw
ConcurrentModificationException out of called().

It surfaced while I was verifying CB-598, which nudges more often, but
the race is on main today and is unrelated to that change.

824 tests, BUILD SUCCESS.
2026-08-16 18:12:23 +02:00
Dai Ha cec48832be CB-600: make it safe to install the launchd agent
CI / build (pull_request) Failing after 1m21s
CI / contract (pull_request) Successful in 1m26s
- redeploy-bridged.sh now refuses (not warns) a supervised restart when
  its computed log path disagrees with the loaded plist's StandardOutPath
  — otherwise every post-restart check reads the wrong file and can
  report a clean restart while the daemon crash-loops. The check is a
  pure, testable function; the script gained a source-for-test guard so
  it can be exercised without installing the agent or touching launchd.
- a failed 'launchctl load' after a successful 'unload' now retries once
  and, on ultimate failure, tells the operator the agent is stopped AND
  disabled plus the exact recovery command, instead of leaving that
  silently worse than the pre-redeploy state.
- the plist documents honestly that the crash loop launchd retries is
  unbounded (ThrottleInterval only paces it), and what actually stops it.
- fixed the requiredSecretEnvVars javadoc: the auth.tokenEnv startup
  throw is ~370 lines below its call site, not a few lines above it, and
  only fires in auth.mode: token.
2026-08-16 18:08:58 +02:00
Dai Ha a36b7ccd7c Merge CB-601: make the recovery-race test deterministic
CI / build (push) Successful in 1m5s
CI / contract (push) Successful in 1m7s
The test asserted one of two interleavings that are both correct, and
steered toward it with a 5 ms Thread.sleep. Under load the other
interleaving happened and main went red on a correct implementation.

The head start is now a latch counted down from inside the sweep's
guarded loop, so the ordering is guaranteed, not likely. Test file only;
the production guard is unchanged.

Verified here: 822 tests BUILD SUCCESS; 10/10 passes while a full clean
install ran in parallel; 5/5 failures with the guard removed, so the test
still catches the bug it exists for.

Closes #95
2026-08-16 18:08:18 +02:00
Dai Ha 32bf324a1e CB-602: guard against a config key that never reaches the example
CI / contract (pull_request) Successful in 49s
CI / build (pull_request) Failing after 1m24s
BridgedConfig.KNOWN_TOP_LEVEL_KEYS is now package-private so a test can assert
every key the parser accepts appears in bridged.example.yaml — live or
commented-out, since the file is gitignored and the example is the only
committed description of the config schema. The existing tests only checked
the example->code direction; this adds code->example.
2026-08-16 18:06:36 +02:00
Dai Ha 16de9df000 CB-601: make the recovery-race test's head start deterministic, not a sleep
CI / contract (pull_request) Successful in 1m8s
CI / build (pull_request) Successful in 1m9s
2026-08-16 18:02:57 +02:00
ltms 65c9deb4d1 Merge CB-597: correct bridged.example.yaml, including two knobs that do nothing
CI / contract (push) Successful in 42s
CI / build (push) Successful in 2m2s
My premise for this ticket was wrong and the worker corrected it. I reported five whole sections missing from the example; nothing was missing. My comparison script only counted uncommented lines, so every section documented as a commented-out example looked absent. Earlier tickets had each updated the example alongside their feature.

What it found instead is more useful than what I asked for — real inaccuracies, found by tracing each field through the parser and its consumers:

- `health.workingSuspectAfterSeconds` and `paneProbeIntervalSeconds` documented enforced minimums that do not exist. I checked: both names appear **only** in the `Health` record declaration and are read by nothing. Only `intervalSeconds` is clamped, and it is silently raised to 15 rather than rejected.
- `notifications.mode: webhook` only flips what `bridge_list` reports as `healthCoverage`. It sends no webhook — "webhook" appears in one `configured()` boolean and there is no delivery code in the repo.
- `lifecycle.clearAfterTurn` was undocumented, and is a no-op for any peer kind other than claude-code.
- The reload doc claimed the whole `fleet:` block is hot; `fleet.leaders` is built once at startup and is not rebuilt, so a change is silently accepted and does nothing until a restart.
- The `fleet.leaders` demotion consequence is now stated next to the block itself: an unmatched pane is silently an ordinary worker and every orchestration call it makes is refused, with no startup error.

Documenting a knob as dead is worth more than documenting it as working. Someone tuning `workingSuspectAfterSeconds` would otherwise have concluded their monitor was broken.

Comments only — no parsing or production code touched. Verified by the lead: parses cleanly under the project's own snakeyaml 1.30, top-level live keys `[bind, herdrSocket, profiles, placement, fleet, guard]`, the rest correctly commented examples. Both dead-knob claims verified by grep against `src/main` rather than taken on the worker's word.
2026-08-16 17:59:41 +02:00
ltms 28ae27b8e1 Merge CB-599: a capacity refusal now tells the caller why
CI / build (push) Failing after 1m19s
CI / contract (push) Successful in 1m25s
`PlacementException extends IllegalStateException`, and neither spawn path caught that type, so it escaped to Javalin's default handler as a bare `500 Server Error` with a text/plain body — while every other failure on the same endpoint returned structured JSON. The reason existed and was good, but only in the daemon log.

I hit this live while orchestrating: asked for a member on a full profile, got a blank 500, guessed another profile, got a blank 500 again, and spent two round trips learning things the daemon already knew.

Both surfaces now catch it. REST returns 503 with `{"error":"no_capacity","detail":...}`; MCP returns the same reason in the `isError` shape it already uses for every other spawn failure. 503 is right because the request was valid and will likely succeed later — the caller did nothing wrong, so 400 would have been a lie.

The other throw sites all funnel through the same type, so quarantine cooldowns, weight-0 exclusion, and the all-at-cap / all-quarantined / all-unreachable messages now reach callers too. That last group matters most: those three distinguish "wait a moment" from "your backends are gone", and all three used to arrive as the identical blank 500.

Tests assert the caller can read the *reason*, not merely that the status changed — one per surface.

Verified by the lead: `mvn -f bridged/pom.xml clean install` unpiped, exit code captured — Tests run: 824, Failures: 0, Errors: 0, Skipped: 0 — BUILD SUCCESS.

No exception message was reworded. This change delivers messages that were already written.
2026-08-16 17:54:44 +02:00
Dai Ha 81a0cf4710 CB-598: track reminder counts per pending item, not per lead per source
CI / build (pull_request) Successful in 1m12s
CI / contract (pull_request) Successful in 1m11s
Work that arrived during the ~15s push_backoff_ms window between two
ticks landed in the pending map before the next tick's start-of-tick
snapshot, so a shared per-lead-per-source counter (carried forward via
scheduleNext(lead, count+1, ...)) already treated it as exhausted
backlog even though no nudge had ever named it. ReplyPushLoop.tick now
recomputes each source's reminder count fresh every tick as the
minimum nudge count among that source's currently pending items, so a
freshly-arrived item (count 0) keeps its source eligible regardless of
how depleted an older, still-undrained sibling's count is. decide()
itself is unchanged.
2026-08-16 17:49:21 +02:00
Dai Ha 8a837a2830 CB-599: surface a capacity refusal's reason instead of a bare 500
CI / contract (pull_request) Successful in 50s
CI / build (pull_request) Successful in 1m35s
PlacementException extends IllegalStateException, which neither BridgedApp
nor BridgeMcp's spawn catch blocks handled, so a maxLoad/quarantine/
all-exhausted refusal fell through to a blank 500 on REST and lost its
message on MCP. Catch it on both surfaces, ahead of the unrelated
IllegalArgumentException(unknown_profile) mapping, and return its message
structured: REST as {"error":"no_capacity","detail":...} with status 503,
MCP as an isError result prefixed "no capacity: ...".
2026-08-16 17:43:07 +02:00
Dai Ha 0a2b3a4a56 CB-597: fix inaccuracies the example config already had, none actually missing
CI / contract (pull_request) Successful in 1m7s
CI / build (pull_request) Failing after 1m21s
Audited bridged.example.yaml against BridgedConfig's KNOWN_TOP_LEVEL_KEYS and
found every top-level key already documented (broker, health, lifecycle,
configReload, quarantineCooldownSeconds, fleet.leaders/architects/reviewers
included) — CB-573/CB-566/CB-559/CB-579/CB-527/528 each updated the example
alongside their feature. What was actually wrong:

- health.workingSuspectAfterSeconds/paneProbeIntervalSeconds claimed enforced
  minimums (300/60) that don't exist in code — only intervalSeconds is
  clamped (floor 15); the other two are parsed but never read anywhere.
- notifications.mode: webhook was undocumented as only flipping the
  healthCoverage label bridge_list reports — no webhook is ever sent.
- lifecycle.clearAfterTurn was missing entirely.
- the HOT bullet under configReload claimed the whole fleet: block reloads
  live, but ConfigRef's own javadoc carves out fleet.leaders as needing a
  restart with no deferred-list warning — added that exception.
- fleet.leaders' demotion consequence (unmatched tab -> silent WORKER
  demotion, no startup error) is now stated inline next to the block, not
  just implied by the multi-lead rationale higher up.
2026-08-16 17:37:28 +02:00
ltms 613ece92dc Merge CB-594: make supervision and a working fleet possible at the same time
CI / contract (push) Successful in 49s
CI / build (push) Failing after 1m45s
The launchd unit was a CHANGEME template that had never been installed, and it could not have worked if it were: launchd does not source a login shell, so the daemon would have started with no forge or gateway token, and the failure would only appear much later as workers unable to open a PR.

Four parts: a wrapper that execs one login shell in place so the job inherits the secret store; a startup report naming which required secret env vars resolved and which are MISSING, by name only, never a value; the plist filled in with this host's real verified paths; and `redeploy-bridged.sh` detecting the agent and switching stop/start to `launchctl unload -w` / `load -w`, falling back to the existing kill + nohup when it is not installed.

Verified by the lead. My own unpiped build of the branch: 813 tests, BUILD SUCCESS. Trial-merged onto current main (which had moved twice) and rebuilt: Tests run: 822, Failures: 0, Errors: 0, Skipped: 0 — BUILD SUCCESS. All three host paths in the plist exist; no CHANGEME remains; the secret report leaks no values.

An independent reviewer, briefed only from the diff, verified the parts that matter today and said merge. It confirmed live that the unsupervised path is unchanged, that `launchctl list` exits 113 for an absent agent so the detection reads correctly, that the wrapper round-trips arguments containing spaces and quotes and stays a single exec chain, and that neither branch of the secret report can print a value.

Its two open findings only bite once the agent is actually loaded, which has not happened and is the operator's call. Filed as #91 — that must land before the agent is ever installed. The important one is that the script computes its log path from its own location while the plist hard-codes an absolute one; if they ever disagree, the post-restart ERROR check reads the wrong file and reports "ok" while the daemon crash-loops.

The agent is deliberately NOT installed by this merge. Nothing here changes how the daemon runs today.
2026-08-16 17:32:57 +02:00
ltms 8d4206c2b5 Merge CB-590: one nudge schedule per lead, with a reminder budget per source
CI / contract (push) Successful in 1m7s
CI / build (push) Failing after 1m15s
Carries PR #84 (its head `78ca24d` is an ancestor of this one), so this single merge delivers both rounds.

CB-590 collapses the CB-307 reply schedule and the CB-588 ticket schedule into one per lead, which closes the double-injection race. The review round found that this also collapsed the two reminder caps into one shared budget — so a busy reply stream could exhaust the cap and a ticket arriving afterwards would never be nudged at all. That was a regression, not a pre-existing wart: before this PR the two sources had independent counters.

The follow-up keeps the single schedule and gives each source its own budget. `decide()` returns INJECT while either source has pending work under its own cap, and STOP only when neither does. `stopOrRestart` and its snapshot-diff race logic are untouched.

Verified by the lead: `mvn -f bridged/pom.xml clean install` unpiped, exit code captured — Tests run: 814, Failures: 0, Errors: 0, Skipped: 0 — BUILD SUCCESS. `ReplyPushLoopTest`: 41 tests.

Known and deliberately out of scope: an item arriving during the 15s backoff is already inside the tick's "before" snapshot, so at-cap work can be abandoned rather than nudged once. That shape predates this PR in the ticket-only loop and is filed as #87 (CB-598).
2026-08-16 17:28:57 +02:00
ltms 03286a589b Merge CB-527/CB-528: bound AMQP prefetch, confirm publishes, and close the recovery race
CI / contract (push) Successful in 1m4s
CI / build (push) Failing after 1m18s
Carries PR #83 (its head is an ancestor of this one), so this single merge delivers both rounds.

CB-527 bounds prefetch. CB-528 makes a publish wait for its broker confirm, so a failed publish is never reported as success, and correlates a `mandatory` Return back to the right publish.

The review round on #83 found one real race: `failPendingPublishesOnRecovery` swept the pending maps without holding `publishChannelLock`, so a publish issued on the already-recovered channel could be failed by the sweep — the exact inversion of what CB-528 exists to prevent, and undetectable by dedup because `MessageService.reply` mints a fresh msgId per call. Fixed by taking the same lock. `close()` now fails in-flight publishes promptly instead of letting them time out after 10s.

Verified by the lead, not taken on the worker's word:
- `mvn -f bridged/pom.xml clean install` unpiped: Tests run: 809, Failures: 0, Errors: 0, Skipped: 0 — BUILD SUCCESS
- `mvn test -Pcontract -Dtest=AmqpReplyInboxContractTest` against a real broker: Tests run: 8 — BUILD SUCCESS

That second run mattered: #83's own new tests are contract-tagged, so the default suite and CI prove nothing about them.
2026-08-16 17:27:21 +02:00
Dai Ha 88b9503c3b CB-590 follow-up: give each nudge source its own reminder budget
CI / contract (pull_request) Successful in 43s
CI / build (pull_request) Successful in 1m46s
decide(lead, reminderCount) shared one counter across the reply and
ticket sources after PR #84 collapsed both onto a single per-lead
schedule. A reply stream that used up the whole budget could then
make decide() STOP even for a ticket that had never been nudged and
had coalesced onto the same still-active schedule — stranding it with
no live schedule left, since stopOrRestart's racedIn check does not
save work that was already present in the "before" snapshot.

decide() now tracks a per-source count (replyReminderCount,
ticketReminderCount) and returns INJECT while either source is still
under its own cap, STOP only when both are exhausted. Still exactly
one schedule per lead; stopOrRestart's snapshot-diff logic is
untouched.
2026-08-16 17:26:07 +02:00
Dai Ha c553d795d8 CB-528: close the recovery race in AmqpReplyInbox
CI / contract (pull_request) Successful in 44s
CI / build (pull_request) Failing after 1m20s
failPendingPublishesOnRecovery walked and cleared pendingBySeq/pendingByMsgId
without holding publishChannelLock, so a publish() that registered while the
sweep was still iterating could be failed even though it published on the
already-recovered channel — a successful publish reported as failed, and
since MessageService.reply mints a fresh msgId per retry, dedup can't catch
the resulting duplicate. Guard the sweep with publishChannelLock: publish()
only holds it for the seq/map-put/basicPublish, so the sweep can only ever
wait for an in-flight basicPublish to return, never a broker round trip.

Also make close() fail in-flight publishes immediately with a clear message
instead of leaving them to idle out the 10s confirm timeout, and record the
(currently unreachable) msgId-uniqueness assumption pendingByMsgId relies on.

AmqpReplyInboxRecoveryRaceTest drives the sweep and a real publish() against
each other directly (no live broker reconnect) using Proxy-backed fake AMQP
channels and a large in-flight backlog to make the race window observable;
confirmed it fails without the guard (reply "fresh" wrongly failed as
"connection recovered mid-publish") and passes with it.
2026-08-16 17:24:43 +02:00
Dai Ha 3ba6d6784c CB-594: make supervision and a working fleet possible at the same time
CI / contract (pull_request) Successful in 44s
CI / build (pull_request) Successful in 1m40s
Adds scripts/bridged-launchd-wrapper.sh so the launchd-run daemon still gets
WORKER_GITEA_TOKEN/AI_GATEWAY_TOKEN by execing through a login shell (launchd
never sources secrets.sh itself). bridged now logs at startup which required
token env vars (derived from each profile's tokenEnv/gitTokenEnv, not a
hand-written list) resolved or are MISSING, by name only. Fills in the real
paths in deploy/dev.ltms.bridged.plist for this host and points it at the
wrapper. scripts/redeploy-bridged.sh now detects a loaded launchd agent and
uses launchctl unload/load instead of a raw kill+nohup, because a bare
SIGTERM exits this JVM at 143 (measured) which KeepAlive.SuccessfulExit=false
reads as a crash and would race the script's own restart; --check reports
installed/loaded state and stays read-only.
2026-08-16 17:13:17 +02:00
Dai Ha 78ca24dc3f CB-590: collapse the CB-307 and CB-588 nudge schedules into one per lead
CI / build (pull_request) Successful in 1m0s
CI / contract (pull_request) Successful in 1m33s
Both reply-queued and ticket-terminal nudges could independently decide
to inject into the same lead pane in the same window, since they ran as
two separate schedules keyed differently (worker target vs. lead) that
never checked each other. Replace both with a single per-lead schedule
(activeLeads) that drains pending reply targets and pending tickets
together, sends at most one combined nudge per tick, and shares one
reminder cap across both sources — so two injections into the same pane
can no longer overlap, and work queued while the lead is busy is never
lost, only deferred.
2026-08-16 16:56:04 +02:00
Dai Ha 4fa6553db5 CB-527/CB-528: bound AMQP prefetch and confirm publishes before claiming durable
CI / build (pull_request) Successful in 1m3s
CI / contract (pull_request) Successful in 1m15s
CB-527: basicQos(prefetch) on the consume channel before basicConsume, configurable
via broker.prefetch (default 32), so an undrained inbox backlog stays on the broker
instead of growing the JVM heap without limit.

CB-528: publish moves to its own confirm-mode channel with mandatory=true and a
return listener, so an unroutable or unconfirmed reply now throws instead of
vanishing silently. The confirm callback checks the per-message returned flag
(set by the return listener, which the broker always fires before the matching
confirm) so an acked-but-returned publish is still reported as a failure. The ack
path stays on its own channel/lock and never waits on a publish confirm.
2026-08-16 16:55:40 +02:00
Dai Ha 2124e043ce CB-593: correct the member MCP claim — measured, not assumed
CI / contract (push) Successful in 43s
CI / build (push) Successful in 1m19s
CLAUDE.md told every member 'You mount only the bridge MCP' and told the lead
'a worker mounts only the bridge MCP and cannot run your other tooling'. Both
were false for Claude Code members.

Measured by spawning one member per backend and asking each what it actually has:

  opencode (gx)      11 bridge tools only                     claim TRUE
  claude-code (local) 11 bridge + 45 gitea + 2 context7       claim FALSE

Source is ~/.claude.json user-scope mcpServers; --mcp-config adds to that scope
rather than replacing it, so the worktree parity overlay (which correctly
neutralises .mcp.json and opencode.json) cannot see or stop it.

The forge tools are mounted but not usable: CB-592's blocked sentinel means
get_me and list_issues both fail with 'invalid username, password or token'.
That is defence in depth working in a path it was not designed for, so the text
now says a mounted tool is not a working tool rather than pretending the tools
are absent.

Block propagated byte-identically to wiki 7-Use-Cases.md (wiki 1d95e3f).
2026-08-16 09:46:42 +02:00
Dai Ha e4f3620acb CB-591: record the final 86400s route timeout and the request:0s trap
CI / contract (push) Successful in 44s
CI / build (push) Successful in 1m19s
systems/vms moved the LLM route timeout again, 1800s -> 86400s (24h), after
the silent-truncation risk was discussed. They tried request: 0s first: it
removes the total-duration timer, but on an AIGatewayRoute the idle timeout
is derived from the request timeout, so 0s also removed any bound on a
stalled connection.

At 86400s our own MessageService.ASYNC_TIMEOUT_MS (30 min) binds first, so a
runaway request now ends as a clean FAILED ticket we raised instead of a
silently truncated 200. While the gateway sat at 1800s the two numbers were
equal and did not nest.
2026-08-15 21:29:49 +02:00
Dai Ha 032a59a34d CB-591: fleet moved onto the gateway — both ceilings fixed and re-verified
CI / contract (push) Successful in 40s
CI / build (push) Successful in 1m19s
`local` runs on /anthropic and `gx` on /v1, both weight 100; `local-direct`
stays weight 0 as the escape hatch.

systems/vms fixed both blockers, and each was re-checked from this side rather
than taken on trust:

    listener buffer    32 KiB -> 32 Mi   ours: 1.2 MB body -> 200 (was 413)
    LLM route timeout  60s    -> 1800s   ours: 101s stream -> 200,
                                               message_stop present, 4000/4000

Neither was deliberate: 32 KiB was Envoy Gateway's default
per_connection_buffer_limit_bytes, and 60s was Envoy AI Gateway's own default.
The 60s bounded GENERATION as well as prompt size — a tiny prompt with a long
answer returned 504 at 60.05s.

Verified with real workloads, not liveness probes. A `local` member read this
document and CLAUDE.md in full — 48,344 bytes of file content, comfortably past
the old 32,768 ceiling — and answered four questions correctly, including the
document's length (said ~456, actual 455). A `gx` member did the same. The
trivial 3-question probe is what hid the 32 KiB ceiling for an afternoon, so it
no longer counts as proof here.

§7.2 is new and is the part that matters later. One risk is ACCEPTED, not
solved: on a mid-response timeout over chunked HTTP/1.1, Envoy ends the chunked
encoding cleanly instead of resetting, so a truncated answer arrives as HTTP 200
with no error and no terminator (envoyproxy/envoy#17186, acknowledged 2021,
never fixed; the Dec 2025 fix #42269 is HTTP/2 only and SSE here is HTTP/1.1).
Measured at the old 60s: 200, 61.07s, 2473 of 4000 emitted, message_stop 0,
error events 0, ending on a well-formed frame.

The recommended defence — reject a stream with no terminator — does NOT
transfer to us: Claude Code and opencode are third-party clients and we do not
own their SSE parsing. So this is acceptable because a request would have to run
1800s to trip it, not because we could detect it. If a member ever returns a
confident but truncated answer, suspect this before anything in our own code.

Also recorded, from the upstream bisection: ClientTrafficPolicy is honoured in
standalone `aigw run` but BackendTrafficPolicy is silently ignored, and nothing
external distinguishes them (envoyproxy/gateway#9513). Same silent-default shape
this repo keeps hitting.

bridged.yaml carries the same notes inline (gitignored, so not in this commit).

Refs: gitea #76
2026-08-15 20:47:41 +02:00
Dai Ha e689090024 CB-591: correct the root cause — Envoy's buffer limit, not Caddy
CI / contract (push) Successful in 1m7s
CI / build (push) Successful in 1m35s
I wrote "Caddy request_body max_size and/or Envoy's own" and marked it
unverified. The Caddy half was wrong, and an unverified guess still points the
next reader at the wrong component.

Confirmed by the systems/vms side: Envoy Gateway defaults a listener's
per_connection_buffer_limit_bytes to 32768, and aigw buffers the WHOLE request
body before it can route on the model name. So that default is not a network
tuning knob — it is a hard ceiling on prompt size. From the live config_dump:

    listener default/llm/http    per_connection_buffer_limit_bytes: 32768

Nobody chose 32 KiB; it was inherited.

Both TLS edges are innocent, and the technique that showed it is better than
mine: both 413s carry x-llm-consumer, a header their auth proxy sets only AFTER
authenticating, so the body cleared both edges and the auth. On llm.vm, aigw
413s at 39 KB while the vLLM backend answers 200 at the same size. I found the
boundary; they found the component, by reading the failure's response headers.

Consequences recorded in the doc:

  * DO NOT plan around 32 KiB. The intended ceiling is far higher, so sizing our
    profiles to it would be designing around a bug.
  * Their fix (ClientTrafficPolicy, bufferLimit: 8Mi) is written but NOT
    deployed, pending their operator's approval. We do not re-test until they
    confirm — a half-changed system gives a number neither side can trust.
  * In standalone `aigw run` a SecurityPolicy is accepted and then silently
    ignored, so "the config was accepted" proves nothing there. They will verify
    by re-reading the live config_dump and sending a large request. Same
    silent-default shape this repo keeps hitting, one layer down.

bridged.yaml carries the same correction (gitignored, so not in this commit).

Refs: gitea #76
2026-08-15 19:47:01 +02:00
Dai Ha 1cc34888fd CB-591: record the live result — blocked by a 32 KiB body limit at the gateway
CI / contract (push) Successful in 44s
CI / build (push) Successful in 55s
Deployed U1-U2c, restarted, spawned both new profiles for real, then reverted.

llm.ltms.dev answers HTTP 413 above 32 KiB (32768 bytes), on BOTH surfaces:

    /v1        32695 bytes -> 200        /anthropic  32095 bytes -> 200
    /v1        32795 bytes -> 413        /anthropic  32855 bytes -> 413

That is far below one agent turn. It is an edge limit (Caddy request_body
max_size, and/or Envoy), so the fix is in systems/vms, not here.

The part worth recording is how it nearly passed. Two members, same message,
same moment: `local` finished in 66s, `gx` never finished at all. `local`
passed only because the probe was three trivial questions in a fresh session,
so the request fit under 32 KiB — the profile looked healthy and was a
landmine set to fire on the first turn that reads a file. So §7's checklist
was not wrong, it was too easy; it now says to use a file-reading task.

opencode's failure mode is worse than a crash: it catches the 413, compacts
its context, retries, and loops. Observed 10+ minutes BUSY with no reply. From
the lead's side that is indistinguishable from a slow worker. Reproduced
outside the bridge with the launcher's own generated config, which is how it
became a one-line error instead of a hang; §7.1 records that procedure.

Everything else about the migration checked out and is recorded so it is not
re-tested: token accepted on both surfaces, unauthenticated 401 (the Caddy
proxy does gate, whatever the gateway's own fail-open policy does),
/v1/models exactly ["deepseek-v4-flash"], the guard allowlist accepted
llm.ltms.dev, and the generated opencode provider block is correct with a real
llmk- key.

Also answers §3b's open question: reasoning survives BOTH surfaces —
/anthropic returns a real "type":"thinking" block and /v1 returns a populated
reasoning_content. The feared /v1 translation loss did not happen.

Config state (bridged.yaml is gitignored, so it is described rather than
committed): `local` back on http://gx00.gw:8000, `gx` kept at weight 0,
`local-direct` kept, llm.ltms.dev left in the guard allowlist. The file
carries these numbers and the exact two-key edit to switch back.

Verified after the revert with a task that reads two large files: correct on
all three questions. Daemon pid 66745, jar f1fd659423e6.

Refs: gitea #76
2026-08-15 19:27:45 +02:00
Dai Ha 0331ecd5d3 CB-592: add the BRIDGED_MEMBER marker — the sentinel alone cannot hold
CI / contract (push) Successful in 1m5s
CI / build (push) Successful in 1m39s
Live check on a member pane showed the CB-592 shadow did NOT take effect:
GITEA_ACCESS_TOKEN inside the pane was still the real admin token.

Measured cause. The overlay itself works — GITEA_TOKEN is injected the same
way, is exported by no shell file, and does reach the pane. The sentinel loses
one step later. A herdr pane runs a LOGIN shell, ~/.zprofile line 41 sources
${SHARED_ENV}/tools/secrets.sh, and that file does a plain unconditional
`export GITEA_ACCESS_TOKEN=...`. A login shell overwrites a value already in
the environment, so the real token is put back before the member starts.
Confirmed directly:

    GITEA_ACCESS_TOKEN=cb592-sentinel zsh -lc ...
    -> RESULT: sentinel was OVERWRITTEN by the login shell

This defeats any launcher-side overlay for any name secrets.sh exports. No
change in this repo can win it alone.

So this adds the half that does survive: BRIDGED_MEMBER=1, a name secrets.sh
never exports. It is a no-op until the operator guards the export:

    [ -n "${BRIDGED_MEMBER:-}" ] || export GITEA_ACCESS_TOKEN=...

Setting it now costs nothing and makes that one line the whole remaining fix.
The sentinel stays: it is correct for any peer kind whose pane does not start
a login shell, and it keeps the intent explicit where every adapter passes.

Also corrects the javadoc and the test javadoc, which both claimed a
protection that was measured not to hold.

The other reported failure was my own bad test, not a regression. The probe
called /api/v1/user, which a minimal write:repository token cannot read. Same
token on the repo endpoint answers 200, so CB-302 is intact:

    GITEA_ACCESS_TOKEN: /user=200  /repos/lms/claude-bridge=200
    WORKER_GITEA_TOKEN: /user=403  /repos/lms/claude-bridge=200

Tests 805 -> 807. Both new tests proved to discriminate by reverting the
marker: everySpawnMarksThePaneAsAMember and
aProfileEnvEntryCannotClearTheMemberMarker both fail without it.

Refs: gitea #77
2026-08-15 18:37:52 +02:00
Dai Ha 831a918c30 Merge CB-592: shadow the admin GITEA_ACCESS_TOKEN in every member's environment
CI / contract (push) Successful in 41s
CI / build (push) Successful in 1m22s
A live probe showed every spawned member carried the admin GITEA_ACCESS_TOKEN: 108
environment variables in a member's pane against 99 in the primary's. The operator's
rule is that only the leader and architects may use it; everyone else uses
WORKER_GITEA_TOKEN. We were not enforcing that at all.

The cause is invisible from inside the launcher. baseEnv builds a fresh map holding only
PATH and the profile's env:, so a member looks like it gets a small explicit environment.
That map is an OVERLAY: WorkspaceControl.createTab/splitPane send only the keys it
contains, and herdr spawns the pane from its own login-shell environment, so every key we
never mention passes straight through — admin token included.

The fix puts a non-blank sentinel over the key in baseEnv, applied AFTER the profile's
env: so no profile, present or future, can restore the real token by naming it in config.
One place, every adapter, including peer kinds not yet written — deliberately not a
per-profile bridged.yaml entry, which is the silent-default shape this repo has shipped
nine times.

A non-blank sentinel rather than the empty string, on purpose: whether an empty overlay
value overrides an inherited variable or is skipped as blank cannot be settled from this
repo, because herdr's merge happens in an external process. baseEnv's own PATH seeding
(CB-511) already relies on a non-blank value replacing an inherited one, so this reuses
the shape that is demonstrated to work rather than the one that is merely plausible.

CB-302's repo-scoped GITEA_TOKEN grant is untouched — a worker can still open its own PR.
The subscription boundary was checked and is unaffected: the primary's pane carries no
ANTHROPIC_* at all, so nothing is inherited there.

Closes gitea #77. Live verification follows separately: the daemon must be redeployed
before this reaches any pane.
2026-08-15 18:27:05 +02:00
Dai Ha 3db5277ae8 CB-592: shadow the admin GITEA_ACCESS_TOKEN in every member's herdr overlay
CI / contract (pull_request) Successful in 1m12s
CI / build (pull_request) Successful in 1m37s
herdr spawns a pane from its own login-shell process env and layers our map on
top, so any key baseEnv never mentions passes straight through — including the
admin forge token. baseEnv now puts a non-blank sentinel for
GITEA_ACCESS_TOKEN, applied after the profile's own env: so no profile can
restore it. One place, every adapter, every profile including future ones.
CB-302's GITEA_TOKEN grant (applyGitToken) is untouched.
2026-08-15 18:25:16 +02:00
Dai Ha 6939e0cbbc CB-592: the tracked opencode.json must name the worker forge token, not the admin one
Operator's rule, 2026-08-15: only the leader and architects may use GITEA_ACCESS_TOKEN;
everyone else uses WORKER_GITEA_TOKEN.

opencode.json is TRACKED, so it ships in every worker worktree, and it mounted the gitea
MCP with {env:GITEA_ACCESS_TOKEN}. A live probe confirmed that variable actually resolves
inside a member: herdr spawns each pane from its own login-shell environment and layers
the launcher's map on top, so a member sees 108 variables rather than the small explicit
set baseEnv appears to build. That gave an opencode member admin forge TOOLS — enough to
merge its own PR, which both CLAUDE.md and the member contract forbid.

This is the narrow half of the fix: it removes the tooling. The admin token is still
present as a string in every member's environment, which is the real defect and is
tracked as CB-592 (gitea #77) — that fix belongs in the launcher, in one place, not
per-profile in bridged.yaml where a sixth profile would silently reopen it.

.mcp.json keeps GITEA_ACCESS_TOKEN and is correct to: it is skip-worktree, the primary's
own local copy, and the primary is the lead. That is the pattern this change follows —
the shared tracked file grants least privilege, and anything needing more overrides
locally.
2026-08-15 18:20:55 +02:00
Dai Ha f0095bf8b2 CB-591: plan the move onto the LLM/MCP gateway, and check AI_GATEWAY_TOKEN
CI / contract (push) Successful in 1m5s
CI / build (push) Successful in 1m38s
The gateway (llm.ltms.dev) replaced Bifrost on 2026-08-15 and serves an Anthropic
surface and an OpenAI surface, so both member kinds can point at it. The plan is in
docs/CB-591-Gateway-Migration.md; gitea #76 tracks the work.

The opencode half needs no code: OpenCodeLauncher already pins an OpenAI-compatible
endpoint (CB-508), so baseUrl + tokenEnv + provider/model is a config change. That
matters more than it looks — every opencode member today is sol or terra, and both sit
on one OpenAI account via credentialId: openai-shared, so an exhaustion on either locks
out both. A gateway-backed opencode profile is free and off that credential, which
retires a single point of failure rather than only adding capacity.

Also extends the redeploy script's --check to AI_GATEWAY_TOKEN. A profile's tokenEnv is
resolved from the DAEMON's own environment by HerdrPeerLauncher.resolveEnv, so a token
added to secrets.sh after the daemon started is simply absent: the launcher injects an
empty token and the gateway answers 401, long after the restart and with nothing tying
the two together. That is the same trap as WORKER_GITEA_TOKEN, and it gets the same
login-shell check that never prints the value.
2026-08-15 17:06:31 +02:00
Dai Ha 5206679efd Merge CB-588: nudge the lead when an async ticket goes terminal
An async delegation ticket (bridge_send wait:false) resolves on MessageService.reply's
rendezvous fast path, which returns before onReplyQueued. So CB-307's push loop only ever
heard about the durable-inbox case, and the mode CLAUDE.md tells leads to prefer never
nudged anyone. Closes gitea #72.

Adds a second, independent reminder schedule keyed by the lead terminal, so several
tickets finishing together coalesce into one nudge. The CB-307 path is untouched.

Three defects were found in review and fixed before merge:
 * a pendingTickets entry outlived the ticket it named. poll() returns null once
   pruneTerminalTickets drops a ticket, so ticketCollected was never reached and the
   entry leaked for the daemon's life, riding along on every later nudge and sending
   the lead after a ticket bridge_poll can no longer find.
 * a lost nudge: a ticket landing between decideTickets returning STOP and
   activeLeads.remove coalesced onto a schedule that was about to die. That is the
   exact failure this ticket exists to remove, reintroduced in a narrow window.
 * the success direction was unpinned in tests, and the comment listing the paths that
   complete the future was short by several.

The obvious fix for the second one was wrong: restarting on any pending ticket defeats
the reminder cap, because a never-collected ticket at cap is expected to still be there.
The fix diffs against a snapshot taken before the decision, so only a ticket that truly
arrived during the window restarts the schedule.

Verified on my own unpiped build: 802 tests, 0 failures, BUILD SUCCESS.
Two reviewers on the diff; the loop-gating finding they raised is split out as CB-590.
2026-08-15 17:06:12 +02:00
Dai Ha ac044e7573 CB-588 round 3: close the STOP-vs-onTicketTerminal race, pin nudge polarity, fix comment
CI / build (pull_request) Successful in 1m19s
CI / contract (pull_request) Successful in 1m31s
- ReplyPushLoop.stopOrRestartTicketLoop: after releasing a lead's active-schedule
  slot on STOP, restart only if a ticket landed that the pre-decision snapshot
  did not already account for. A naive "restart on any pending ticket" version
  was tried first and reverted: it defeated the reminder cap by restarting
  forever on a stale, never-collected ticket (broke
  successfulTicketNudgeIncrementsDelivered and ticketNudgesSendUpToCapThenStop).
  Diffing against a pendingBefore snapshot distinguishes a genuine race arrival
  from stale cap-exhausted backlog.
- Two new ReplyPushLoopTest cases exercise stopOrRestartTicketLoop directly
  (now package-private) rather than forcing the underlying thread race:
  aTicketStillPendingWhenTheLoopStopsIsNotStranded (the race must restart) and
  aStaleUncollectedTicketAtCapDoesNotRestartTheLoop (the cap must still hold).
  Both were verified to fail against deliberately-reverted versions of the fix
  before being restored to green.
- MessageServiceTest: pin the success-path nudge test's negative direction too
  (must not contain "FAILED"), not just the failure-path test.
- MessageService: complete the whenComplete comment's list of completion paths
  (TIMED_OUT/BUSY/BACKEND_EXHAUSTED via finishAsyncTask, completeExceptionally
  on throw, and answer() -> finishAsyncTask(turnId, result)).
2026-08-15 17:02:05 +02:00
Dai Ha 6ebad2a91f CB-588 follow-up: reclaim a pruned ticket's pendingTickets entry too
CI / contract (pull_request) Successful in 45s
CI / build (pull_request) Successful in 1m0s
tasks is the sole authority on whether a ticket exists, but
pruneTerminalTickets dropped entries from it without telling
ReplyPushLoop. ticketCollected(ticket) was only ever called from
MessageService.poll's terminal branch, which pruneTerminalTickets
short-circuits past once a ticket is gone (poll returns null at the
top). A ticket the lead never polled — or one the reminder cap already
gave up on — was pruned from tasks but never collected in
ReplyPushLoop.pendingTickets, so it rode along on every later nudge to
the same lead forever, naming a ticket bridge_poll could no longer
find, and the map itself never shrank.

pruneTerminalTickets now calls pushLoop.ticketCollected for every
ticket it actually removes (guarded on pushLoop != null), so a pending
nudge entry lives exactly as long as its ticket is pollable. Reused
tasks as the only removal trigger rather than adding a second live
query back into MessageService — no new source of truth.

Added an injectable clock (LongSupplier nowNanos, defaulting to
System::nanoTime) to MessageService, mirroring the SessionManager/
SessionReaper nowNanos seam, so a test can cross the 10-minute
TICKET_TTL_NANOS deterministically instead of sleeping for real.
Confirmed the new regression test fails against the prior
pruneTerminalTickets (a stale ticket rides along on a later coalesced
nudge) before restoring the fix.
2026-08-15 16:41:23 +02:00
Dai Ha 3d10ed385c CB-588: nudge the lead when an async ticket reaches a terminal phase
CI / contract (pull_request) Successful in 41s
CI / build (pull_request) Successful in 1m46s
An async bridge_send(wait:false) registers a rendezvous waiter, so its
reply always takes MessageService.reply's fast path and returns before
ReplyPushLoop.onReplyQueued is ever called — the exact mode the charter
tells leads to prefer never nudged.

Add ReplyPushLoop.onTicketTerminal(ticket, target, failed), a second
entry point reusing the loop's status gating, bounded/backoff reminders
and metrics, keyed by the nudge-receiving lead so several tickets
finishing together coalesce into one nudge naming the count. MessageService
wires it via task.future.whenComplete in sendAsync (covers reply,
completion fallback, wedge, and abandon() alike) and calls the new
ticketCollected(ticket) from poll() once a terminal view is handed back,
so an already-collected ticket is never nudged again. The CB-307 inbox
path (onReplyQueued/decide/NUDGE_FORMAT) is untouched.
2026-08-15 16:10:22 +02:00
Dai Ha 6d0c94dbdb Correct enforceMaxLoad's comment after CB-585
CI / contract (push) Successful in 42s
CI / build (push) Successful in 1m16s
The comment said non-positive means unlimited at load. That stopped being
true when CB-585 made an explicit maxLoad: 0 survive as a real cap of zero
and made a negative value refuse config load. The code below it was already
right — only the comment described the old normalisation. Flagged by the
CB-585 worker, which correctly stayed out of a file not on its list.
2026-08-15 16:02:35 +02:00
Dai Ha e01563a550 Merge cb584-session-resume-b62247-4
CI / build (push) Successful in 1m20s
CI / contract (push) Successful in 1m44s
2026-08-15 15:58:29 +02:00
Dai Ha 30e3225a3c Merge cb585-maxload-zero-4585c4-3 2026-08-15 15:58:29 +02:00
Dai Ha ef186a1516 Merge cb587-snapshot-index-flags-24e236-2 2026-08-15 15:58:29 +02:00
Dai Ha 5d5b3bdc76 CB-584: persist agentSessionId at acquire and expose it via bridge_spawn/bridge_list
CI / contract (pull_request) Successful in 1m5s
CI / build (pull_request) Successful in 2m18s
Wires up the two links that made session resume unreachable: MemberSession
now records agentSessionId from PeerHandle at acquire (both the plain and
worktree paths), and it survives onto the roster (rosterView, so both
bridge_list and GET /members show it). bridge_spawn accepts sessionName and
resumeSessionId, threading them into the SpawnRequest fields that already
existed but were never reachable from the MCP surface.

A resumeSessionId now requires an explicit profile (a resumed conversation
is tied to the specific backend that started it, so an unqualified spawn
routed by placement has no safe candidate to check) and is refused, naming
Capability.SESSION_RESUME, when that profile's adapter does not declare it
— PeerLauncher gains capabilitiesFor(profileName) so a mixed fleet is
checked per-adapter rather than against the fleet-wide capability union.

PeerHandle.agentSessionId() loses its default, the same fix CB-571 (c00a86b)
made for charterReceipt() one method above it — both existing adapters
already overrode it, so this only closes the landmine for a future one.

Out of scope, left for follow-up: carrying agentSessionId on a failed
ticket's ReleaseDetail (issue #65 criterion 5) — CB-578 stage C just
landed in that file and this ticket deliberately stayed out of it.
2026-08-15 15:48:06 +02:00
Dai Ha 2f8c98dac9 CB-585: maxLoad: 0 caps a profile at zero, negative refused at load
CI / contract (pull_request) Successful in 41s
CI / build (pull_request) Successful in 1m18s
2026-08-15 15:38:50 +02:00
Dai Ha 2db7189067 CB-587: seed the snapshot's temp index from the worktree's real index
CI / build (pull_request) Successful in 1m1s
CI / contract (pull_request) Successful in 1m24s
--skip-worktree is an index flag. GitWorktrees.snapshot staged into a fresh
empty temp index, which carried none of the real index's skip-worktree bits,
so add -A staged local on-disk content for files git status correctly hides
(e.g. .mcp.json). Copy the real index (resolved via git rev-parse --git-path
index, correct for linked worktrees) into the temp index before staging, so
add -A skips exactly what git status skips.
2026-08-15 15:36:06 +02:00
Dai Ha 94476ac109 CB-529: drainReplies javadoc said the ack is local — false for AMQP
CI / build (push) Successful in 55s
CI / contract (push) Successful in 1m5s
The doc claimed an in-flight failure re-surfaces the messages on a later
drain, because the ack is local. That was true when InMemoryReplyInbox was
the only inbox, and became false without anyone noticing when the AMQP
adapter landed: there the ack is a broker-side basicAck, so a crash while
writing the response loses the reply outright. Re-polling cannot recover it,
since the broker has already forgotten it.

Behaviour is unchanged and the window stays accepted — acking on the next
poll instead would double-deliver on every normal drain. The point is that
the comment sat exactly where the next implementer would read it and said
the opposite of what happens.

Closes #12.
2026-08-15 15:34:59 +02:00
Dai Ha 5fe02b7c98 Fix redeploy script reporting failure on a successful restart
CI / contract (push) Successful in 45s
CI / build (push) Successful in 59s
Found by running it. The script launched the daemon with a relative jar path
(cwd is bridged/) but detected it with an absolute one, so pgrep never matched.
The daemon restarted correctly and booted clean, and the script still failed
with 'no process appeared' — the worst shape of bug for a deploy tool, because
it invites a second restart on a daemon that is already healthy.

Detection now matches both path forms, and the launch uses the absolute path so
ps names which checkout is running.
2026-08-15 15:18:44 +02:00
Dai Ha a1052f4fd1 Add scripts/redeploy-bridged.sh so the lead can deploy in one command
CI / build (push) Successful in 57s
CI / contract (push) Successful in 1m23s
The lead already owned the redeploy, but the command classifier refuses a bare
kill on the daemon, so in practice every deploy still needed the operator to
approve a stop and a start by hand. A single script is the seam that fixes
that: the operator allow-lists one auditable command instead of two ad-hoc
ones.

It also stops the procedure from living only in a checklist people read after
things go wrong. It builds before it stops anything, so a failed build never
leaves the fleet down; waits for the old process to exit instead of assuming;
polls /healthz; and anchors its log checks to a line marker taken before the
restart, so old errors cannot be misread as new ones.

The check with no log line anywhere in the daemon is the reason --check exists:
bridged inherits WORKER_GITEA_TOKEN from the shell that starts it, and starting
from a non-login shell empties it. The daemon boots fine, healthz is green, and
the failure only appears later as workers that cannot open a PR. --check tests
whether the name resolves and never prints the value.
2026-08-15 15:14:21 +02:00
Dai Ha 8c9904a7c4 Record that the classifier still blocks the daemon kill
CI / contract (push) Successful in 45s
CI / build (push) Successful in 59s
A CLAUDE.md rule grants intent, not tool permission. Tried the redeploy right
after writing the section and the classifier refused `kill`, so the section
would have been misleading as written. States the honest split until a Bash
permission rule exists: the lead builds and verifies, the operator runs the
stop and start with `!`.
2026-08-15 13:57:18 +02:00
Dai Ha 23af5dfdff The lead may redeploy the daemon — write down the procedure and its traps
A merge is not a deployment: the running bridged holds the jar it started
with, so merged code does nothing until the daemon is rebuilt and restarted.
Calling that work shipped is a false report. This makes the redeploy the
lead's job rather than something handed back to the operator, and records the
five things that have gone wrong doing it here — chiefly that starting the
daemon from a non-login shell empties WORKER_GITEA_TOKEN, which nothing logs
and which only surfaces later as workers that cannot open a PR.

Goes in the project addendum, not the canonical block; the block is unchanged
and still byte-identical with the wiki template.
2026-08-15 13:54:38 +02:00
Dai Ha 3a10f6ad17 Note that stage C's snapshot makes the parityOverlay gitignore rule load-bearing
CI / build (push) Successful in 56s
CI / contract (push) Successful in 1m19s
CB-578 stage C commits a preserved dirty worktree to refs/wip/<branch> with
git add -A. The parity overlay copies the primary's environment files into
every worktree, so a non-gitignored overlay path now reaches a durable git
object instead of only sitting on disk. add -A respects .gitignore, which is
what stops it — so the rule documented for CB-581 is now what keeps a secret
out of a commit, not just out of a directory.
2026-08-15 13:09:01 +02:00
Dai Ha a3842c873d Merge cb578c-92885c-1 2026-08-15 13:07:50 +02:00
Dai Ha c29c3f063d Merge cb554-e8e224-3 2026-08-15 13:07:50 +02:00
Dai Ha 758d62a396 Merge cb583-b82d11-2 2026-08-15 13:07:50 +02:00
Dai Ha 6725642274 CB-578 stage C: snapshot a dirty worktree into refs/wip before it can be lost
CI / build (pull_request) Successful in 54s
CI / contract (pull_request) Successful in 1m21s
Adds Worktrees.snapshot(worktreePath, branch, message): stages into a
temporary GIT_INDEX_FILE (never the worker's real index/HEAD), writes
the tree, commit-trees it onto the worktree's current HEAD, and points
refs/wip/<branch> at the result. add -A (never -f) respects .gitignore.

SessionManager.release() now snapshots any dirty worktree before the
preserve-or-remove decision, regardless of release cause (COMPLETED or
SHUTDOWN) — preserving on disk alone is one `worktree remove --force`
away from gone. A failing snapshot never escalates: the worktree is
still preserved, the pane still stops, and the release listener is
still notified.

onRelease's listener now receives a ReleaseDetail (terminal, worktree
path, branch, snapshot ref) instead of a bare terminal id, so Bridged's
WORKER_FAILED wiring can put the same three facts into a failed
ticket's detail — a lead can re-dispatch onto the same tree instead of
starting from the base commit.
2026-08-15 12:56:12 +02:00
Dai Ha a1f4dc365a CB-554: weight <= 0 excludes a profile from automatic placement, not 1.0
CI / contract (pull_request) Successful in 44s
CI / build (pull_request) Successful in 52s
2026-08-15 12:54:02 +02:00
Dai Ha 165b62ee20 CB-583: make bridge_list's capacity view quarantine-aware
CI / contract (pull_request) Successful in 44s
CI / build (pull_request) Successful in 1m32s
A quarantined profile's capacity row now forces free:0 and names the
quarantine (credentialId, quarantinedForSeconds), reusing the same
QuarantineSource bridge_profiles already reads instead of a second
lookup. An ordinary fleet's capacity rows are unchanged (no new keys).
2026-08-15 12:51:36 +02:00
Dai Ha 72d3481de3 Merge CB-578 stage B: quarantine the exhausted credential, not the profile
CI / contract (push) Successful in 41s
CI / build (push) Successful in 1m32s
Closes issue #50. Verified by the lead: own unpiped build of the branch merged
onto main — 740 tests, BUILD SUCCESS, exit 0.

Stage A only classified a usage-limit refusal. Nothing acted on it, so the
fleet would spawn another member onto the same exhausted account and fail the
same way. Now a BACKEND_EXHAUSTED classification quarantines the CREDENTIAL
for a cooldown: an explicit spawn onto it is refused with the profile, the
credential and the seconds remaining, placement skips it under every policy,
and it lifts itself on an injected clock.

Keyed by credential, not profile name, because sol and terra are two models on
one OpenAI account. Quarantining only the profile that reported the refusal
would leave its sibling live, and the next spawn walks onto the same dead
account. A profile that sets no credentialId quarantines alone under its own
name, so an existing config behaves exactly as before.

The implementer also found a real bug outside the brief: ConfigRef's
sameLaunchSettings never compared exhaustedPattern, so a reload changing only
that key reported 'config reloaded' while the value is actually deferred —
the precise failure that file's own doc calls the worst outcome a reload can
produce. Fixed with a regression test.

Two things I changed at review. The example.yaml conflict with e09cac6 was
resolved toward the branch, which was a strict superset. And
BackendQuarantine.none() held a clock frozen at 0, so a quarantine() call on
it would have locked a credential out for the life of the daemon — two
production CompositePeerLauncher constructors default to none(). It is now a
real no-op with a test.

Left open deliberately: the implementer's bridge_ask about the credentialId
name timed out after 55s with no answer, because I was polling minutes apart.
It proceeded on its own judgment and chose well. The 55s ask window against an
async lead is a bridge problem, not a worker problem.
2026-08-15 10:33:59 +02:00
Dai Ha 29ccb747ad CB-578 stage B: make BackendQuarantine.none() actually inert
Added at merge review. none() held a clock frozen at 0 with a 1ns cooldown, so
a quarantine() call on it recorded a deadline that could never pass — the
credential would be locked out for the life of the daemon. Two production
CompositePeerLauncher constructors default to none(), so that failure would
have been silent and permanent.

The implementer documented the limitation honestly rather than hiding it, but
a stand-in named none() should not need the caveat. quarantine() is now a
no-op on that instance, with a test asserting it. An inert value must omit the
fact, never invent one.
2026-08-15 10:33:44 +02:00
Dai Ha 4749e27453 Merge remote-tracking branch 'origin/main' into HEAD
# Conflicts:
#	bridged/bridged.example.yaml
2026-08-15 10:31:41 +02:00
Dai Ha e501d39988 CB-578 stage B: quarantine the exhausted credential, not the profile
CI / build (pull_request) Successful in 55s
CI / contract (pull_request) Successful in 1m21s
A BACKEND_EXHAUSTED classification (stage A) now puts that profile's
credential into a BackendQuarantine for a configurable cooldown. A spawn
onto a quarantined profile is refused naming the credential and roughly
when it lifts; weighted/round-robin/fixed placement skip a quarantined
candidate; the quarantine lifts itself on the injected clock; and it is
visible on bridge_profiles.

Keyed by credential, not by profile name, via the new Profile.credentialId
(profiles sharing one credential quarantine together — e.g. two models on
one account) and effectiveCredentialId() (unset ⇒ quarantines alone,
today's behaviour unchanged). Fixed a related gap along the way: a reload
changing exhaustedPattern was silently reported "applied" even though it's
deferred — sameLaunchSettings() now catches it too.
2026-08-15 10:29:54 +02:00
Dai Ha 6b6cf25862 CB-581: neutralise the parityOverlay false-preserve risk, and record the rule
CI / contract (push) Successful in 1m5s
CI / build (push) Successful in 1m32s
Issue #57 criterion 5 asked for an explicit decision on this rather than
silence. It is closer to live than the issue assumed.

The DEFAULT parityOverlay is List.of(".env", ".envrc"), so it applies to every
profile, and neither path was gitignored. Since CB-576 a release preserves any
worktree that git status --porcelain calls dirty, and that deliberately counts
untracked files — the work lost in CB-576 was a file nobody had added. So one
.env at the repo root would make every COMPLETED release preserve its
worktree, and worktrees would accumulate with no error to notice.

Inert today only because neither file exists here. Both are now gitignored,
which they deserve on their own as environment files. The javadoc carries the
rule for the next overlay path: it must be gitignored, or tracked and
skip-worktree'd.
2026-08-15 10:17:26 +02:00
Dai Ha 180de840eb Merge CB-581: a throw in release() no longer orphans the pane or aborts reapIdle
CI / contract (push) Successful in 43s
CI / build (push) Successful in 52s
Verified by the lead: own build of the branch merged onto main — 717 tests,
BUILD SUCCESS, exit 0.

release() ran an unprotected sequence. hasUncommitted shells out to git status
and throws on a non-zero exit, which skipped both notifyReleased and
launcher.stop. The session was already out of the registry, so nothing retried
it: a live pane kept burning a fleet slot while absent from the roster, and any
send blocked on it was never resolved. reapIdle called release bare inside a
loop, so one such session aborted the whole pass and skipped every session
after it.

Now the dirty-check block is guarded and fails toward preserving the worktree —
'we could not tell' must not be treated as 'it is clean', because deleting on a
guess destroys work with no other copy. notifyReleased runs in a finally and
launcher.stop runs unconditionally, so the pane always stops. reapIdle catches
per session, matching the shape drainAll already used.

Checked while reviewing: notifyReleased cannot throw out of the finally — it
already guards each listener and only logs. The pane stop is genuinely
unconditional.
2026-08-15 10:14:41 +02:00
Dai Ha 8fd2d7e5e7 CB-581: fail-safe release() so a throw never orphans the pane or aborts reapIdle
CI / build (pull_request) Successful in 52s
CI / contract (pull_request) Successful in 1m1s
hasUncommitted shells out to git and can throw; release() now catches that
inside a try/finally so notifyReleased and launcher.stop always run, and
defaults to preserving the worktree on a throw (can't tell dirty vs clean,
so don't risk deleting unrecoverable work). reapIdle wraps each per-session
release in try/catch, matching drainAll, so one bad session no longer
skips the rest of the reaping pass.
2026-08-15 10:09:32 +02:00
Dai Ha 129dd4a838 M4: record what Unit 2 has actually landed, and correct criterion 15
CI / build (push) Successful in 58s
CI / contract (push) Successful in 1m16s
Unit 2 was written as one block, but CB-571/576/580 and unit 2a have since
built parts of it. A symbol survey of main source plus the merge history puts
it at 1 of 21 criteria done, with three partials.

Criterion 15 says a normal COMPLETED release removes the worktree. Since
CB-576 that is false on purpose — a dirty worktree is preserved, because
deleting it destroys work that cannot be recovered. The criterion is wrong,
not the code.
2026-08-15 10:08:38 +02:00
Dai Ha e09cac6f1f CB-578 stage A: document exhaustedPattern as a deferred reload key
CI / contract (push) Successful in 44s
CI / build (push) Successful in 1m31s
The patterns are compiled once in Bridged.main from the startup config
snapshot, so adding one to a profile does nothing until a restart. Neither
the per-key docs nor the reload-class table said so.
2026-08-15 10:06:41 +02:00
Dai Ha c00a86b32c Merge CB-571: charter receipt on the roster, with no silent null for OpenCode
CI / contract (push) Successful in 42s
CI / build (push) Successful in 1m19s
Verified by the lead: own build of this branch merged onto main — 710 tests,
BUILD SUCCESS, exit 0.

Supersedes PR #54. A reviewer found that OpenCodeLauncher.SessionAwareHandle
wrapped the base WorkerHandle but never overrode charterReceipt(), so it
inherited the interface default of null while the real receipt sat on its
delegate — sol and terra would have shown no charterSource/charterSha256 on
the roster while Claude Code members showed both.

The fix is the root one, not the one-line override: PeerHandle.charterReceipt()
is no longer a default, so the compiler forces every implementation to answer.
This repo had shipped that same class of defect — a defaulted dependency that
compiles, passes tests, and quietly turns a feature off — eight times before
this one.
2026-08-15 10:02:36 +02:00
Dai Ha d3ae0350a2 Merge remote-tracking branch 'origin/main' into fix-charter 2026-08-15 10:01:10 +02:00
Dai Ha 2c2196a1f1 Merge CB-578 stage A: classify a usage-limit refusal instead of a completed reply
CI / build (push) Successful in 55s
CI / contract (push) Successful in 1m24s
Verified by the lead: own build of the branch merged onto main — 704 tests,
BUILD SUCCESS, exit 0. No vendor wording in any Java source (grep clean); a
profile with no exhaustedPattern keeps today's completion-fallback path exactly.

Accepted the implementer's deviation from the brief. The brief asked for a
terminal HealthState BACKEND_EXHAUSTED. FleetHealth.decide() is a pure
classifier over HealthSnapshot, which carries only booleans, AgentStatus and
MemberSession.State — never pane text. FleetHealth's own javadoc already says
ERROR_ON_SCREEN 'is not decided yet because it needs a bounded pane detection
read'. A second declared-but-unproduced value would repeat a gap the class
already documents as a problem, so the signal was put where the evidence
actually lives: the completion scrape.
2026-08-15 10:00:34 +02:00
Dai Ha 35ade14630 CB-571: make PeerHandle.charterReceipt() abstract, fix OpenCode adapter's silent null
CI / build (pull_request) Successful in 50s
CI / contract (pull_request) Successful in 1m21s
SessionAwareHandle wrapped the base's WorkerHandle but never overrode
charterReceipt(), so it silently inherited the interface default (null)
while the real receipt sat on its delegate. sol/terra never got a
charterSource/charterSha256 roster row.

Deletes the default so every PeerHandle must answer explicitly; the
compiler now catches this class of gap instead of a roster field
quietly going missing.
2026-08-15 09:55:52 +02:00
Dai Ha 541df87272 CB-578 stage A: classify a usage-limit refusal instead of a completed reply
CI / build (pull_request) Failing after 59s
CI / contract (pull_request) Successful in 1m10s
A backend that refuses on a subscription usage limit leaves the pane healthy but
the turn ends with no bridge_reply; the completion fallback used to scrape and
hand that refusal back as if it were a real answer. CompletionResolver now
matches the scrape against a per-profile exhaustedPattern (config, never a
vendor string) and resolves the send as Rendezvous.Kind/Outcome.BACKEND_EXHAUSTED
with a reason carrying the matched line, kept distinct from GONE/WORKER_FAILED.
A profile with no pattern configured is unaffected. Coverage is logged at
startup via CompletionResolver.coverage(...), naming which profiles have a
pattern and which don't, following FleetHealthMonitor.coverage's pattern.
2026-08-15 09:55:06 +02:00
Dai Ha 337b6ccd6e Merge remote-tracking branch 'origin/main' into fix-charter
# Conflicts:
#	bridged/src/test/java/dev/ltms/bridged/session/SessionManagerTest.java
2026-08-15 09:52:39 +02:00
ltms 42f46dfe9a Merge CB-579: resolve a lead by its tab name, drop the terminal-id pin
CI / build (push) Successful in 58s
CI / contract (push) Successful in 1m4s
Verified by the lead: merged onto main (9088d2b) in a scratch worktree, mvn -f bridged/pom.xml
clean install unpiped — MVN_EXIT=0, Tests run: 696, Failures: 0, BUILD SUCCESS. main alone measures
692, so this adds 4 net tests. Merges cleanly; Bridged.java auto-merged against CB-580.

Reviewed by the lead reading the full production diff and the three test files. The member's own
report was lost to the idle reaper before collection, so there was no author write-up.

Closes the live bug: LeadTabScanner.scan() no longer merges the config pin over the scan result, and
the cache no longer seeds from it, so a lead disappears once its tab is gone. The ghost this fixes
had begun throwing agent_not_found from ReplyPushLoop.decide on every tick, against a dead terminal
that still owned two live members.

Beyond the brief, and correct: `tab` is required for every leader, not only non-creatable ones,
because LeadLauncher also uses it to label a tab it creates. terminal: is rejected by a raw-YAML
check rather than by record shape — the only way to beat @JsonIgnoreProperties(ignoreUnknown = true).
The silent-default trap is avoided: the back-compat constructor still takes `tab` positionally.

Behaviour change worth knowing: the scanner's initial cache is now empty instead of the config pins,
so a herdr failure on the very first scan yields no leads until a scan succeeds. That is unavoidable
once the pins are gone, it fails loudly rather than silently, and it is covered by
aFailedFirstScanReturnsEmptyRatherThanThrowing.

Operator action required: bridged.yaml must replace fleet.leaders.<name>.terminal with tab. Already
done for this deployment.
2026-08-15 09:29:00 +02:00
ltms 9088d2b2c5 Merge CB-580: fail a ticket when its member reaches a terminal health state
CI / contract (push) Successful in 41s
CI / build (push) Successful in 1m32s
Verified by the lead on 0af902e: mvn -f bridged/pom.xml clean install, unpiped, in a scratch
worktree — MVN_EXIT=0, Tests run: 689, Failures: 0, Errors: 0, BUILD SUCCESS.

Reviewed by the lead reading the diff. The member's own report was lost to the idle reaper before
it was collected, so there was no author write-up to review against.

All three defects that sank 3b2f395 are absent:
  * failTarget is required — one constructor, Objects.requireNonNull, no defaulting overload
    anywhere in the repo, and Bridged.java:365 updated to pass messages::abandon.
  * The failure fires inside reportTransition after its `if (previous == next) return;` guard, so an
    unchanged tick cannot reach it.
  * No AtomicReference; MessageService is passed directly, so there is no empty window.

abandon(String target, String reason) confirmed as CB-568's target-wide operation that resolves
waiters as a failure rather than letting them time out.

Known limitation, accepted and covered by its own test: states.put records the new state before the
bounded retries run, so if all three attempts throw, the tickets stay pending and no later tick
retries. It is logged at WARN, and the retries carry no backoff.
2026-08-15 08:57:03 +02:00
ltms 500bfa2c33 Merge CB-576: release preserves a dirty worktree instead of deleting it
CI / build (push) Successful in 51s
CI / contract (push) Successful in 1m1s
Verified by the lead on b525b0f: mvn -f bridged/pom.xml clean install, unpiped, in a scratch
worktree — MVN_EXIT=0, Tests run: 686, Failures: 0, Errors: 0, BUILD SUCCESS.

Review accepted the required-interface-method shape (no defaulting overload) and the decision to
count untracked files as dirty — the work lost in the incident was a file that was never added.

One blocking defect was found and fixed in b525b0f: hasUncommitted called git with no existence
check, so a missing worktree threw WorktreeException from inside release() after registry.remove()
but before notifyReleased() and launcher.stop(), orphaning the pane and stranding a blocked
bridge_send caller. It now mirrors remove()'s already-gone tolerance.

Waived on merge, tracked as follow-up: a SessionManager-level test that teardown completes when the
worktree is gone, and the stronger fix behind it — reapIdle calls release() with no try/catch while
drainAll wraps it, so any exception in that window aborts the whole reaping pass.
2026-08-15 08:54:33 +02:00
Dai Ha 9ca9c43dfa CB-571: retag charter-receipt references from the taken CB-575
CI / contract (pull_request) Successful in 43s
CI / build (pull_request) Successful in 54s
CB-575 already names the merged MCP-cancellation-filter change, so the
charter-receipt comments used the wrong number. Retag to CB-571, the number
this work was authored against.
2026-08-15 08:48:57 +02:00
Dai Ha b525b0f08f CB-576: hasUncommitted tolerates an already-gone worktree
CI / contract (pull_request) Successful in 1m8s
CI / build (pull_request) Successful in 1m37s
2026-08-15 08:47:12 +02:00
Dai Ha 1966c69994 CB-575: charter receipt on spawn, in the roster and in the logs
CI / contract (pull_request) Successful in 46s
CI / build (pull_request) Successful in 1m34s
Record a CharterReceipt (role, source, sha-256 digest, byte count) for every
launch, store it on the MemberSession, expose it in bridge_list and GET
/members, and log it at spawn as digest+role only. Redact the charter argv
argument in the legacy pane-placement spawn log so the charter text never
reaches the daemon log. The charter prose itself is never recorded.
2026-08-15 08:10:56 +02:00
Dai Ha 976eff8ad1 CB-579: resolve a lead by its tab name, drop the terminal-id pin
CI / build (pull_request) Successful in 52s
CI / contract (pull_request) Successful in 1m18s
Leader.terminal -> Leader.tab (exact tab label, case-insensitive match).
LeadTabScanner matches an exact tab->name map instead of stripping a
shared tabPrefix, and no longer merges configured leads into every
scan result -- a stale pin can no longer outlive its tab.
LeadLauncher.tabLabel() returns the configured tab directly; the
terminalId pinned-terminal fallback in liveLeads() is gone.
Config load now rejects a leftover fleet.leaders.*.terminal key
instead of silently ignoring it. primary.terminal is untouched.
2026-08-15 08:05:21 +02:00
Dai Ha 0af902ec43 CB-580: fail a ticket when its member reaches a terminal health state
CI / build (pull_request) Successful in 54s
CI / contract (pull_request) Successful in 1m5s
FleetHealthMonitor now requires a failTarget BiConsumer<String,String>
collaborator (no defaulting overload) and calls it exactly once when a
member transitions into GONE or NEVER_READY, via CB-568's idempotent
target-wide abandon() operation. The reason string names the real
terminal state. failTarget invocation retries up to
MAX_FAIL_TARGET_ATTEMPTS (3) within the same transition if it throws,
and never refires on a later tick where the state is unchanged.

Bridged.java wires messages::abandon as the production failTarget.
2026-08-15 07:54:28 +02:00
Dai Ha 9118ce2537 CB-576: release preserves a dirty worktree instead of deleting it
CI / contract (pull_request) Successful in 1m2s
CI / build (pull_request) Successful in 1m35s
2026-08-15 07:34:06 +02:00
Dai Ha 2f48e08f1f Merge CB-577 follow-up: drop the target-keyed async index
CI / contract (push) Successful in 43s
CI / build (push) Successful in 55s
asyncTasksByWaiter correlates an async question by the exact rendezvous
waiter, so the target-keyed set it replaced can no longer decide
anything. Keeping it meant two indexes of the same fact, one of them
ambiguous whenever a target has two accepted tickets.

The race the old test modelled by reflection is gone with it: identity
keys make 'some other task reached this target' unrepresentable, so
there is no longer a wrong task for the question to land on.

Also carries the criterion-1 doc correction, which is identical to
aac29d6 on a different parent.
2026-08-15 06:40:31 +02:00
Dai Ha fec284e7cb Merge M4 unit 2a: accepted-turn identity carried to delivery
CI / build (push) Successful in 59s
CI / contract (push) Successful in 1m17s
TurnToken, owned by MessageService, binds a target to the exact
rendezvous waiter for one accepted send. Injector.Pending carries it and
the delivery callback hands it to CompletionResolver, so the baseline is
bound to the send it belongs to by construction rather than by a lookup
that could pick a different one.

The callback signature is required, not a defaulted overload: a delivery
with no token is exactly the unbound baseline this unit forbids, so a
default would let a caller silently produce it.

The token deliberately omits the session turn number. MessageService
owns acceptance but never learns of delivery, and
CompletionResolver.onDelivered runs before SessionManager.onDelivered,
so the number does not exist yet at the only point the token could
capture it. docs/M4-Fleet-Health.md criterion 1 records this and the two
rejected alternatives.

Still open for the next slice: the missing/post-restart baseline test and
the no-replay test.
2026-08-15 06:36:46 +02:00
Dai Ha 8e2e4c5e73 M4 unit 2a: migrate test call sites to the required turn token
The delivery callback now requires a TurnToken, so 54 test call sites
had to pass one. They use an explicit TestTurnTokens.inert(target)
rather than a defaulted overload, because a delivery with no token is
the unbound baseline this unit forbids.

The first version of inert() returned a fresh CompletableFuture as the
waiter, which turned captureBaselineSkipsTheReadWhenNoSendIsWaiting red:
the resolver saw a non-null waiter, concluded a turn was in flight, and
scraped a pane no send was blocked on. An inert value must omit the
fact, not invent it, so the waiter is now null and the production skip
fires as designed.
2026-08-15 06:36:35 +02:00
Dai Ha 0edc6615fc M4: correct unit 2 criterion 1 — TurnToken cannot carry the session turn
The criterion required the token to bind the session turn number. Three
independent refusals from the implementer showed why that is not
implementable at this layer: MessageService owns acceptance but never
learns of delivery, and CompletionResolver.onDelivered runs before
SessionManager.onDelivered, so the turn number does not exist yet at the
only point the token could capture it.

Records both rejected alternatives and why, so the next reader does not
re-derive them: a target-keyed registry restores the ambiguity the token
exists to remove, and injecting a turn counter couples layers to fill a
field nothing reads yet.
2026-08-15 06:32:02 +02:00
Dai Ha b745e159de CB-573: carry accepted turn tokens on delivery 2026-08-15 06:30:41 +02:00
Dai Ha aac29d604c M4: correct unit 2 criterion 1 — TurnToken cannot carry the session turn
CI / contract (push) Successful in 1m0s
CI / build (push) Successful in 1m32s
The criterion required the token to bind the session turn number. Three
independent refusals from the implementer showed why that is not
implementable at this layer: MessageService owns acceptance but never
learns of delivery, and CompletionResolver.onDelivered runs before
SessionManager.onDelivered, so the turn number does not exist yet at the
only point the token could capture it.

Records both rejected alternatives and why, so the next reader does not
re-derive them: a target-keyed registry restores the ambiguity the token
exists to remove, and injecting a turn counter couples layers to fill a
field nothing reads yet.
2026-08-15 06:29:27 +02:00
Dai Ha c884802b13 CB-577: remove obsolete async target tracking
CI / contract (pull_request) Successful in 43s
CI / build (pull_request) Successful in 56s
2026-08-15 06:28:53 +02:00
Dai Ha 5f5573a24e Merge CB-577: correlate an async question by its exact waiter
CI / contract (push) Failing after 0s
CI / build (push) Successful in 1m13s
markAsyncQuestion picked the first not-done task out of an unordered
set, so between resolveQuestion waking the first async send and the
question being recorded, a queued second send could join the set and
take the question. A lead answering with bridge_send{turnId} would then
resume a turn it did not mean to.

Each async task is now indexed by its exact rendezvous waiter, which has
identity semantics, so no other task can hold the same key. The question
is recorded before resolveQuestion, with a rollback when no waiter is
there, which closes the window rather than narrowing it.

An unanswered async question stays PENDING — the worker resumes after
its ask times out, so the delegation is not failed — and its stale
target tracking is now cleared instead of leaking.
2026-08-15 06:27:16 +02:00
Dai Ha 74b0087ebb CB-577: model async question ownership race
CI / build (pull_request) Successful in 1m29s
CI / contract (pull_request) Failing after 0s
2026-08-15 06:24:05 +02:00
Dai Ha 927e0151d4 CB-577: test async question waiter ownership 2026-08-15 06:23:22 +02:00
Dai Ha 5275922d1d CB-577: handle questions without async waiters 2026-08-15 06:22:22 +02:00
Dai Ha e186c7945a CB-577: correlate async questions to turns 2026-08-15 06:21:53 +02:00
Dai Ha 33a6e77f0e Merge CB-573: dormant fleet health monitor (M4 unit 1)
CI / build (push) Successful in 49s
CI / contract (push) Successful in 1m14s
An opt-in whole-fleet observer, separate from the 250ms delivery poller.
One AgentControl.list and one roster snapshot per tick, joined and fed to
the FleetHealth classifier, because a fault is a disagreement between the
two views at the same instant. Absent a health: block nothing is built
and no herdr call is made.

Adds bridge_list healthCoverage: off, detection-only, or full. Detection
is deliberately separate from notification, so a single-lead setup with
no webhook still gets detection and is told its coverage is partial
rather than being refused.

Two review fixes worth naming. tick() rescheduled itself as its last
statement with no try/catch, and a ScheduledExecutorService does not
re-run a task that threw — so the first agents.list failure would have
stopped health permanently and silently, which is exactly when the
control link is down. It now catches Throwable and reschedules in a
finally. And the snapshot fields this unit cannot supply are the named
constant NOT_YET_OBSERVED rather than bare false literals, because false
means no fault to this classifier.
2026-08-15 06:19:15 +02:00
Dai Ha 4e47489d53 Merge CB-568c: fail every pending async ticket on teardown
A released target left its second async ticket pending for the full
30-minute async timeout. abandon resolved only the rendezvous waiter,
and async tickets live in a separate map that could not even represent
two tasks on one target.

abandon now sweeps every non-question async ticket for the target, and
asyncTasksByTarget holds a set. The sweep is a plain loop: the first
attempt used Stream.anyMatch, which short-circuits on the first true, so
it completed one ticket and left the rest pending — the exact bug it was
fixing. Its test passed only because the rendezvous path failed the
first ticket anyway; the test now uses three tickets so a single
completion cannot satisfy it.

An ASKING ticket is an active turn, not a pending send, so the sweep
skips it and CB-574 is unaffected.
2026-08-15 06:18:28 +02:00
Dai Ha c76b2f149e CB-573: keep health monitoring after failures
CI / contract (pull_request) Successful in 42s
CI / build (pull_request) Successful in 1m15s
2026-08-15 06:17:28 +02:00
Dai Ha 826fffe05b CB-573: add dormant fleet health monitor 2026-08-15 06:15:24 +02:00
Dai Ha 75f57cdba7 CB-568c: fail every queued async ticket
CI / contract (pull_request) Successful in 44s
CI / build (pull_request) Successful in 1m34s
2026-08-15 06:15:21 +02:00
Dai Ha f556af5d4e CB-568: fail queued async tickets on teardown 2026-08-15 06:14:28 +02:00
Dai Ha d5f33f0c6e Merge CB-568: a dropped send reports the real cause
CI / build (push) Successful in 55s
CI / contract (push) Successful in 1m0s
Injector.drop knew the precise cause (herdr agent_not_found) but the
sender was told only 'worker unreachable or stuck', so a lead could not
tell a dead pane from a stalled model.

TurnListener.onTurnFailed gains a reason, defaulting to the old one-arg
form. CompletionResolver prefers that reason, then the pane scrape, then
the old fixed text.

drop now fires onTurnFailed unconditionally. That is the substantive
fix: the sender blocks on the rendezvous waiter, not on the delivered
future, so failing delivered() alone never woke it and a queued send sat
until its timeout.
2026-08-15 06:08:41 +02:00
Dai Ha 16e17b32ad CB-568: preserve dropped turn causes
CI / contract (pull_request) Successful in 43s
CI / build (pull_request) Successful in 1m26s
2026-08-15 06:06:07 +02:00
Dai Ha 988494e18b Merge M4 fleet-health design (docs only)
The full 5-unit design behind M4: evidence model, classification
precedence, the automatic-vs-lead action boundary, worktree safety on
release, typed inbox and lead routing, capacity, and human escalation.

Two decisions worth keeping visible. Detection is split from
notification, so health works in a single-lead setup with no webhook and
reports partial coverage instead of refusing to run. And capacity stays
a view: the bridge reports free slots but never spawns, reassigns, or
stops a member to improve utilisation, because only the lead holds the
work list.

Section 13 records eleven things nobody checked, including live LavinMQ,
OpenCode pane fixtures, and multi-lead routing.
2026-08-15 06:05:26 +02:00
Dai Ha c4549a5e20 Merge CB-575: filter the routine MCP cancellation WARN
The MCP SDK 2.0.0 registers no handler for notifications/cancelled, so every
client abort logged a WARN. M4 fleet health treats WARN as action-needed, so
that noise had a cost. A Logback TurboFilter denies only that one event:
right logger, WARN level, the SDK's exact format string, and a
JSONRPCNotification whose method is notifications/cancelled. Everything else
is NEUTRAL. If a later SDK handles cancellation the filter stops matching.
2026-08-15 06:00:40 +02:00
Dai Ha 46fa4f38d5 CB-575: filter MCP cancellation warnings
CI / build (pull_request) Successful in 49s
CI / contract (pull_request) Successful in 1m1s
2026-08-15 05:57:45 +02:00
80 changed files with 10391 additions and 688 deletions
+8
View File
@@ -6,6 +6,14 @@
# Settings backups inherit the env block — and secrets with it.
.claude/settings.local.json.bak*
# The default profile parityOverlay copies these primary→worktree, so they appear in EVERY worker
# worktree. Two reasons they must be ignored. They hold environment values, which is reason enough.
# And since CB-576 a release preserves any worktree that `git status --porcelain` calls dirty —
# untracked files included, deliberately. An untracked overlay file would therefore make every
# COMPLETED release preserve its worktree, and worktrees would pile up with no error to notice.
.env
.envrc
# Daemon runtime artefacts. bridged appends its log wherever it is launched from, so both the
# repo root and bridged/ collect one; neither belongs in git.
bridged.out
+73 -9
View File
@@ -80,10 +80,12 @@ below are the procedure — run them in order, every task, not only the big ones
of your context, your plan, or your screen.
5. **Collect** — `bridge_poll{ticket}` → `bridge_ack{ticket, msgId}`. Answer a worker's `bridge_ask`
with `bridge_send{turnId, content}` — **not** `sessionId`. A worker gone quiet is diagnosed with
`bridge_status`, never by reading its terminal.
6. **Verify yourself.** Re-run the build and the checks. A worker mounts only the bridge MCP and
cannot run your other tooling, and a piped command (`… | tail`) hides failures behind a zero
exit — never promote a worker's "clean" to a fact.
`bridge_status`, never by reading its terminal; it also reports an open question and the `turnId`
that answers it. **A worker's ask waits ~55 seconds, and no nudge makes that longer** — so never
brief a worker to "ask me". Decide before you delegate, or give it an explicit default.
6. **Verify yourself.** Re-run the build and the checks. A worker cannot run your IDE tooling, any
forge tools it appears to have hold a blocked credential and fail, and a piped command
(`… | tail`) hides failures behind a zero exit — never promote a worker's "clean" to a fact.
7. **Review — fan out.** Spawn reviewers against the diff, one per dimension or per file, with
`wait:false`. Never the implementer of the scope it reviews, and brief them from the diff — not
from the implementer's rationale, which carries its own blind spot. Dispatch each PR's reviewers
@@ -105,8 +107,8 @@ the merge — and merging on a reviewer's word is delegating it by proxy.
|---|---|
| Confirm your own role | `bridge_whoami` |
| See backends available | `bridge_profiles` |
| Start a member | `bridge_spawn{role?, profile?, cwd?, worktree?, ticket?}` → `sessionId` + `paneId` |
| See the fleet | `bridge_list` → `leads` (your peers) + `members` · one peer's state: `bridge_status{sessionId}` |
| Start a member | `bridge_spawn{role?, profile?, cwd?, worktree?, ticket?, sessionName?, resumeSessionId?}` → `sessionId` + `paneId` |
| See the fleet | `bridge_list` → `leads` (your peers) + `members` (each carries `agentSessionId` when its backend knows one) · one peer's state: `bridge_status{sessionId}` |
| Delegate (blocking) | `bridge_send{sessionId, content}` |
| Delegate (long task) | `bridge_send{sessionId, content, wait:false}` → ticket → `bridge_poll{ticket}` |
| Answer a member's `bridge_ask` | `bridge_send{turnId, content}` — **not** `sessionId` |
@@ -155,9 +157,13 @@ you.
without replying, the bridge scrapes your pane, and it can return only the last 4000 characters.
A clipped scrape is marked as partial, but the missing text is gone — your report reaches the
lead with its end cut off.
5. **Report honestly.** State only what you actually ran and its real output, including failures.
You mount **only** the bridge MCP — the primary's other servers (IDE, forge, docs) are not yours,
so never claim the result of a check you had no way to run.
5. **Report honestly.** State only what you actually ran and its real output, including failures,
and never claim the result of a check you had no way to run. **Measure your own tools; do not
assume them.** What you mount depends on your backend: an opencode member gets the bridge and
nothing else, while a Claude Code member also inherits the operator's user-scope MCP servers,
which the bridge never chose for you. Two rules follow. The primary's IDE tooling is still not
yours, whatever you see. And **a mounted tool is not a working tool** — the forge server you may
find there holds a deliberately blocked credential and fails every call, by design.
6. **Never merge.** Stage files explicitly — never `git add -A` — and leave alone anything the
project marks as not-yours-to-commit.
@@ -189,6 +195,64 @@ charter, not here.
are diagrammed in `docs/MCP-Contract.md` §6, kept out of this file because it loads into every
session's context.
### Redeploying the daemon — the lead may do this (primary only)
**A merge is not a deployment.** The running `bridged` holds the jar it was started with, so a
feature merged to `main` does nothing until the daemon is rebuilt and restarted. Saying "shipped"
about code the live daemon has never loaded is a false report. The lead **may and should** redeploy
rather than hand the job back to the operator.
Workers must never do this. A worker has no business restarting the daemon it is talking through,
and stopping it kills the worker's own channel mid-turn.
**Use the script — do not hand-roll the steps.**
```bash
scripts/redeploy-bridged.sh --check # report state, change nothing
scripts/redeploy-bridged.sh # build, confirm drain, restart, verify
scripts/redeploy-bridged.sh --yes # skip the drain prompt (fleet already checked)
```
It builds before it stops anything, so a failed build never leaves the fleet down; it waits for the
old process to exit rather than assuming; it polls `/healthz`; and it anchors its log checks to a
line marker taken before the restart, so old errors cannot be misread as new ones. Run `--check`
first — it is read-only and reports whether the forge token resolves, which nothing else tells you.
The script encodes the five things below, each of which has gone wrong here before. Read them anyway:
if the script is unavailable or a step fails, this is what it was protecting you from.
1. **Login shell, or workers silently lose their forge token.** The daemon inherits
`WORKER_GITEA_TOKEN` from the shell that starts it, and that comes from
`${SHARED_ENV}/tools/secrets.sh`. Start it from a non-login shell and the variable is empty, the
daemon starts fine, and the failure appears much later as workers that cannot open a PR. Nothing
logs this at startup — the script's `--check` is the only thing that reports it, and it checks
whether the name resolves without ever printing the value.
2. **Drain live members first.** `bridge_list`, then `bridge_stop` each member, and collect anything
you still want with `bridge_poll` before you kill anything. A restart drops in-flight tickets and
rendezvous, and a member's report is not recoverable once its ticket is gone.
3. **A restart is the only way deferred config keys take effect.** That is usually the reason to do
it. The startup log names which keys it accepted and which it deferred — read those lines rather
than assuming.
4. **Re-check identity afterwards.** Call `bridge_whoami` and confirm it still answers `primary`. The
lead is found by its tab label (`fleet.leaders.*.tab`), and a lead whose tab no longer matches is
demoted to worker, which refuses every orchestration call.
5. **Prove the new jar is the one running.** Confirm a *fresh* `bridged listening` line at the end of
`bridged/bridged.out`, dated after the restart. An old daemon that never died looks identical from
the outside.
**Permission.** A `CLAUDE.md` rule grants intent, not tool permission — the command classifier
refuses a bare `kill` on the daemon whatever this file says. The script is the seam that fixes that:
it is one auditable command, so the operator allow-lists it once instead of approving a stop and a
start every time. The rule lives in the operator's Claude Code settings:
```json
{ "permissions": { "allow": ["Bash(scripts/redeploy-bridged.sh:*)"] } }
```
Granted by the operator on 2026-08-15. If a call is still refused, do **not** route around it by
running the stop and start as separate commands — that is exactly the approval the script replaced.
Say what you were going to run and why, and let the operator decide.
### The prompt is part of the product — update it with the code (mandatory)
This repo *is* the bridge, so the canonical block above is not documentation about someone else's
+219 -26
View File
@@ -47,15 +47,18 @@ bind:
# (say a Claude lead and an opencode lead) work as peers: the second is silently demoted and refused
# every orchestration call. List each lead's pane here and all of them resolve as leads.
#
# terminal → the ONLY field identity depends on; get it from that session's bridge_whoami
# tab → the ONLY field identity depends on (CB-579); the exact label of the tab hosting the lead.
# Label the tab yourself, or let bridged label one it launches — see `fleet.leaders:` below.
# kind/model → descriptive; they document what runs in the pane and are echoed by bridge_whoami
#
# A lead is never spawned — it pre-exists, which is exactly why it must be named rather than created.
# A lead's tab must already carry its label (or be launched by bridged, which labels it) — there is
# no terminal id to paste in and nothing to re-pin when the session restarts: the tab survives, so
# the same label resolves the same lead again on the next scan.
# `bridge_whoami` reports `{"role":"primary","leader":"<name>"}`; role stays "primary" because a lead
# IS a primary for authorization, so nothing that keys on the role breaks.
#
# KEEP `primary:` when adding leads: it still addresses the CB-307 push loop, which needs a single
# destination for its nudges. If both name the same terminal, the `fleet.leaders:` entry wins.
# destination for its nudges, and is a separate mechanism from lead identity — see `fleet.leaders:`.
#
# Leads are configured under `fleet.leaders:` — see THE FLEET further down.
#
@@ -85,6 +88,29 @@ bind:
# backoffMs: 60000
# quietNudgeCap: 3
# Fleet health detection is dormant unless enabled (CB-573). It reads one whole-fleet agent list
# per tick.
# intervalSeconds → how often a tick runs (default 30). ENFORCED floor of 15: the code computes
# Math.max(15, intervalSeconds), so a lower value is silently raised, not
# rejected.
# workingSuspectAfterSeconds, paneProbeIntervalSeconds → accepted and parsed, but NOT YET READ by
# anything — the dormant monitor only consumes intervalSeconds today (CB-573
# shipped ahead of the evidence publishers these two knobs are for). Setting
# them changes nothing right now, and no minimum is enforced on either, because
# nothing reads them to enforce one. They exist so a later build can start
# honouring them without another config-shape change.
# notifications.mode → "webhook" flips what bridge_list REPORTS (healthCoverage: "full" instead
# of "detection-only") — it does NOT make bridged send any webhook call; no
# delivery mechanism is implemented yet. Any other value, or omitting the
# block, reports "detection-only".
# health:
# enabled: true
# intervalSeconds: 30
# workingSuspectAfterSeconds: 600
# paneProbeIntervalSeconds: 60
# notifications:
# mode: disabled
# herdr Unix socket. Omit to use the client default
# (${HERDR_SOCKET_PATH:-~/.config/herdr/herdr.sock}).
herdrSocket: ~/.config/herdr/herdr.sock
@@ -125,6 +151,23 @@ herdrSocket: ~/.config/herdr/herdr.sock
# SSH is unaffected). The token value itself is never stored in this file.
# gitHostEnv → host env var holding the forge host (default GITEA_HOST). Injected as
# GITEA_HOST *only* alongside a resolved gitTokenEnv.
# exhaustedPattern → regex matched against a completion-fallback scrape (CB-578 stage A) to
# classify a turn that ended with no bridge_reply as the backend having
# refused on a subscription usage limit, rather than a real answer. Opt-in —
# omit and this profile's completion fallback behaves exactly as before.
# Every backend words its refusal differently, so this is config, never a
# vendor string baked into bridged itself.
# DEFERRED: compiled once into a startup pattern map — editing it needs a
# daemon restart, same as this profile's model/baseUrl/argv.
# credentialId → CB-578 stage B: the credential this profile quarantines WITH when a
# BACKEND_EXHAUSTED classification fires. Two profiles that set the SAME
# credentialId share one quarantine — the case this exists for is two models
# on one account (e.g. sol and terra both billing one OpenAI credential): an
# exhaustion on either one must lock out both, or the fleet just walks onto
# the same dead account under the sibling's name. Opt-in — omit and this
# profile quarantines alone, under its own name, exactly as if the field did
# not exist. Cooldown length is the top-level quarantineCooldownSeconds below.
# HOT: read live at every spawn/exhaustion check — no restart needed.
# env → extra environment for this profile's workers, as a literal key/value map
# (CB-511). Use it to give workers a toolchain.
#
@@ -154,10 +197,21 @@ profiles:
mcpUrl: http://127.0.0.1:8765/mcp
tokenEnv: BRIDGED_WORKER_TOKEN
argv: ["ccs", "gx10"]
weight: 0.5 # relative selection weight for placement: weighted
maxLoad: 2 # max live workers on this profile (omit for unlimited)
# weight: relative selection weight for automatic placement (weighted, round-robin, and
# fixed's fallback walk). Absent defaults to 1.0. An explicit 0 or negative value means
# "never auto-select this profile" (CB-554) — it stays reachable via an explicit
# `bridge_spawn{profile:"gx10"}`, which bypasses placement entirely; only automatic
# selection skips it.
weight: 0.5
# maxLoad: max live workers on this profile. Omit for unlimited. An explicit 0 (CB-585) caps
# the profile at zero live members — it is excluded from automatic placement and an explicit
# `bridge_spawn{profile:"gx10"}` against it is refused too; a cap holds even when the profile
# is named directly. Negative is refused at config load — there is no sane meaning for it.
maxLoad: 2
# gitTokenEnv: GITEA_TOKEN # opt-in: let this profile's workers open their own PR (CB-302)
# gitHostEnv: GITEA_HOST # defaults to GITEA_HOST; injected only with gitTokenEnv
# exhaustedPattern: "usage limit has been reached" # opt-in: classify a usage-limit refusal (CB-578)
# credentialId: shared-openai # opt-in: quarantine together with every other profile sharing this id (CB-578)
# configDir: /Users/me/.ccs/instances/gx10 # CLAUDE_CONFIG_DIR — inherit that profile's skills/MCP
# cwd: /Users/me/src/myrepo # pin the working dir; omit to inherit the primary's
# parityOverlay: [".claude/settings.local.json", ".env", ".envrc"] # never add .mcp.json — see above
@@ -221,8 +275,37 @@ profiles:
# argv: ["opencode"]
# How an unqualified spawn chooses a profile: fixed (default, reproduces pre-CB-518 behaviour),
# round-robin, or weighted. Omitting this key is a strict no-op for existing configs.
#
# `weighted` IS NOT "cheapest first" — read this before you set weights (CB-589).
# It is smooth weighted round-robin: it spreads spawns across EVERY profile that has a free slot,
# in weight ratio. It has no idea which profile costs money. So with local:10 / paid:2 you do not
# get "use local, overflow to paid" — you get roughly one spawn in six going to the paid profile
# while the local box still has a free slot.
#
# There is a sharper second effect. The policy's running score map lives for the daemon's whole
# life. While a profile is at maxLoad it is filtered out and its score FREEZES, so the paid
# profiles keep accumulating against it. When the local slot frees up it returns with a stale
# score and can LOSE the next pick — a paid spawn while the free box sits idle.
#
# Until a real cost-first policy exists, the workaround is to make the ratio decisive rather than
# proportional: give the free profile a weight so large that it wins every pick it is eligible
# for, and paid profiles only ever take genuine overflow. On this host that is local weight 100
# against paid weights of ~1.
#
# The gotcha with that workaround: it expresses a PREFERENCE ORDER through a RATIO knob. Add a
# future profile at weight 150 and it silently outranks the free box, with nothing to warn you.
# Re-check the weights whenever you add a profile.
placement: weighted
# How long a credential sits out after a BACKEND_EXHAUSTED classification (CB-578 stage B), in
# seconds, before a spawn may land on it again. Applies to every profile's effective credential
# (its own name, or its credentialId if set above) — there is no per-profile override. Default
# 1800 (30 minutes) when omitted or non-positive.
# DEFERRED: baked once into the BackendQuarantine built at startup — a running quarantine keeps
# its original cooldown regardless; a new value only applies to a quarantine that starts after a
# restart. Editing this needs a daemon restart to take effect.
# quarantineCooldownSeconds: 1800
# Re-read this file without restarting the daemon (CB-559). Off unless you add this block, so an
# upgraded bridged keeps the old behaviour: the file is read once at boot and never again.
# enabled → turn the watch on. bridged checks the file's modified time on a timer and
@@ -232,17 +315,25 @@ placement: weighted
# Not every key can move under a running daemon, and the difference is about what already exists
# when the reload happens — not about how important the key is:
# HOT → takes effect on the next spawn: the whole `fleet:` block (every role pool,
# `charters`, and `tabLabel`), `placement:`, and an existing profile's weight / maxLoad. Those are
# hot because the placement policy reads them through a supplier — being config is
# `charters`, and `tabLabel`), `placement:`, and an existing profile's weight / maxLoad
# / credentialId. Those are hot because the placement policy (and, for credentialId,
# the CB-578 stage B quarantine check) reads them through a supplier — being config is
# not by itself enough to make a key hot.
# EXCEPT `fleet.leaders`: Bridged.main reads it once at startup to build the lead tab
# scanner and launcher, and neither is rebuilt on reload. A changed/added/removed
# `fleet.leaders` entry is silently accepted — the reload reports "config reloaded"
# with nothing in the deferred list — but has NO effect until you restart. Treat it
# as deferred in practice, even though today's reload output does not say so.
# DEFERRED → accepted into the new config, but the wiring built at startup keeps the old value
# until you restart: `lifecycle:`, `leadHeartbeat:`, `guard:`, `worktreeRoot:`,
# `spawnReadyTimeoutMs` / `spawnReadyPollMs`, ADDING or REMOVING a profile (a new
# backend needs its own launcher, and launchers are built once), AND an existing
# profile's launch settings — model, baseUrl, argv, env, configDir, mcpUrl, tabLabel.
# The launcher takes a copy of `profiles:` at startup and resolves every spawn out of
# that copy, so those never reach a launch until you restart. The reload logs them by
# name rather than pretending they applied.
# `spawnReadyTimeoutMs` / `spawnReadyPollMs`, `quarantineCooldownSeconds` (CB-578
# stage B — baked once into the quarantine tracker built at startup), ADDING or
# REMOVING a profile (a new backend needs its own launcher, and launchers are built
# once), AND an existing profile's launch settings — model, baseUrl, argv, env,
# configDir, mcpUrl, tabLabel, exhaustedPattern. The launcher takes a copy of
# `profiles:` at startup and resolves every spawn out of that copy, so those never
# reach a launch until you restart. The reload logs them by name rather than
# pretending they applied.
# COLD → cannot change at all: `bind:`, `herdrSocket:`, `broker:` and `auth:`. The socket is
# bound, the broker connection is open, and the auth mode decides who may reach the
# port that is already listening.
@@ -299,27 +390,35 @@ fleet:
# Panes that orchestrate rather than are orchestrated. A lead may now be CREATED as well as
# recognised: give it a `profile:` and the daemon launches the shortfall when fewer than
# `instances` are live. Give it only a `terminal:` and it is recognise-only, as before.
# `instances` are live. Omit `profile:` and it is recognise-only, as before.
#
# `tabPrefix` is the naming convention that finds a lead without pasting a terminal id: label the
# tab `lead: <name>` when you open it and the pane is recognised on the next rescan. Reopen the
# tab later and the id changes; the label does not.
# `tab:` (CB-579) is REQUIRED and is the only field identity depends on — the exact label of the
# tab hosting the lead, matched case-insensitively. Label the tab yourself and put that same
# string here, and the pane is recognised on the next rescan. Reopen the tab later, or the session
# inside it restarts — the terminal id changes; the tab, and its label, do not, so no config edit
# follows a restart.
#
# A lead the daemon launches is labelled BY the daemon, using the same convention, so it is found
# A lead the daemon launches is labelled BY the daemon with this same `tab:` value, so it is found
# by the same scan. A lead counts as live only when herdr also reports a running agent in that
# tab — a label left behind by a session that died does not block the relaunch.
# tab — a label left behind by a session that died does not block the relaunch, and a tab that is
# gone entirely drops out of the next scan rather than being remembered forever.
#
# An auto-launched lead is NOT a member: it gets no worker reply charter, is never registered with
# the session lifecycle (the idle reaper would kill your orchestrator), and stays on the
# subscription — ANTHROPIC_BASE_URL/AUTH_TOKEN are stripped from its env whatever the profile says.
#
# GET THE `tab:` VALUE RIGHT. A pane that does not match any configured `tab:` (a typo, a renamed
# tab, a pane no entry names at all) is not recognised as a lead — it resolves as an ordinary
# WORKER instead, silently, and every orchestration call it makes (spawn/stop/send/drain) is
# refused. There is no error at startup for this: an unmatched pane is simply not a lead. If your
# primary suddenly can't spawn or send, check this section first.
# leaders:
# opus-5.0:
# profile: opus # omit to never create this lead, only recognise it
# instances: 1 # desired live count; only the shortfall is launched. 0 = off
# terminal: term_0123456789abcd # optional hand-pin; usually found by tabPrefix instead.
# # A running agent on this terminal also counts as live, so a
# # lead you opened by hand is not relaunched under you.
# tabPrefix: "lead:" # `lead: opus-5.0` ⇒ a lead named opus-5.0 (case-insensitive)
# tab: "lead: opus-5.0" # REQUIRED — the exact tab label this lead lives in
# tabPrefix: "lead:" # only used to guard against a worker tabLabel colliding with
# # this convention at startup; plays no part in matching a lead
# scanIntervalSeconds: 10 # rescan cadence, and the worst case before a new tab is seen
# workspace: leads # where a launched lead's tab is created (default "leads").
# # MUST NOT be a member workspace — those are excluded from the
@@ -327,7 +426,7 @@ fleet:
# cwd: /path/to/repo # the launched lead's working directory (default: bridged's own)
# kind: claude # descriptive; reported by bridge_whoami
# gpt-sol-5.6:
# terminal: term_fedcba9876543
# tab: "lead: gpt-sol-5.6"
# kind: opencode
# model: openai/gpt-5.6-terra
@@ -351,6 +450,91 @@ guard:
- gx00.gw
- gx01.gw
# Member credential policy (CB-596, gitea issue #82). A herdr pane runs a LOGIN shell, and that
# shell re-sources the operator's own secret store — so a spawned member inherits every credential
# the operator's shell holds, not just the ones bridged means to give it. Measured on this host:
# 31 credential names, all set, with only ONE (GITEA_ACCESS_TOKEN) blocked before this — and that
# block was a single name hardcoded in HerdrPeerLauncher.java, not driven by this file. This block
# replaces that hardcoded shadow with a config-driven list of names.
#
# ROUND-2 CORRECTION, measured live: the pane-creation env overlay below (applied at tab.create /
# pane.split, BEFORE the pane's login shell runs) does NOT survive that login shell for any name
# secrets.sh actually exports — the shell re-exports it afterwards and overwrites the sentinel.
# Proof: GITEA_ACCESS_TOKEN comes back blocked only because secrets.sh itself carries a guarded
# export (`[ -n "${BRIDGED_MEMBER:-}" ] || export GITEA_ACCESS_TOKEN=...`) — that guard, not this
# file, is what wins. No other name in `known` below has a matching guard in secrets.sh yet (1
# guard measured against 33 export lines there). So today this block's overlay is REAL protection
# only for a name secrets.sh does not export, or a peer kind whose pane never runs a login shell —
# for everything secrets.sh exports and guards, the guard in secrets.sh (out of scope for this
# ticket) is what actually blocks it, not this list. An exec-time fix (winning after the login
# shell finishes, before the agent process starts) was attempted and found to have no seam in the
# current herdr protocol — AgentControl.start takes a fixed `kind` (herdr resolves the executable)
# plus trailing CLI args for that binary, not an arbitrary argv or an env map; only tab.create /
# pane.split accept `env`, and that is this same pane-creation overlay. See gitea #82 for the open
# design question this leaves.
#
# DENY-BY-DEFAULT, NOT A DENY-LIST. A deny-list (name the bad ones, let everything else through) is
# silently wrong the moment the operator's store gains a new secret — nothing would ever report it.
# Deny-by-default inverts that: `known` bounds the blast radius to names actually enumerated below,
# and EVERY one of them is blocked UNLESS it is also in `allow`. Omitting this block entirely (the
# shipped default) blocks NOTHING — unlike most optional blocks in this file, absence here is a real
# gap, not a safe "feature off". A name that is neither `known` nor `allow`-ed is not silently let
# through either: the daemon logs a WARN naming any credential-shaped env var it finds on neither
# list (never its value), so a secret added to the store later does not go unnoticed forever.
#
# policy → only "deny-by-default" exists today (an operator-authored deny-list was deliberately
# rejected — see above). An unrecognized value refuses to start, naming it.
# allow → credential names a member legitimately needs. Left OUT of the pane's env overlay
# entirely, so the value the pane's own (login) shell exports passes through untouched.
# known → every credential name the operator's store is known to export. Every name here NOT
# also in `allow` is overlaid with a non-secret sentinel value before the pane's login
# shell runs — real protection only for names the login shell does not itself re-export
# (see the ROUND-2 CORRECTION note above for the ones it does).
#
# HOT-RELOADABLE the same way `fleet:` is (CB-559): read fresh on every spawn, so editing this list
# and reloading config (or restarting) changes what the NEXT spawn inherits; already-running members
# are unaffected either way.
# memberCredentials:
# policy: deny-by-default
# allow:
# - AI_GATEWAY_TOKEN # named in a profile's tokenEnv (local/gx) — a member reaching the
# # gateway is by design, not a leak
# - WORKER_GITEA_TOKEN # the repo-scoped forge token a member needs to open its own PR (CB-302)
# - CONTEXT7_TOKEN # already decided as allowed by CB-593
# - GITEA_HOST # not a credential — a hostname, paired with the forge token above
# known:
# - AI_GATEWAY_TOKEN
# - BESZEL_ADMIN_EMAIL
# - BESZEL_ADMIN_PASSWORD
# - BESZEL_HUB_URL
# - BESZEL_KEY
# - BESZEL_UNIVERSAL_TOKEN
# - BRAIN_MCP_TOKEN
# - CF_ACCOUNT_ID
# - CF_API_TOKEN
# - CF_USER_TOKEN
# - CONFLUENCE_API_TOKEN
# - CONFLUENCE_USERNAME
# - CONTEXT7_TOKEN
# - GITEA_ACCESS_TOKEN
# - GITEA_HOST
# - GITLAB_OAUTH_CLIENT_SECRET
# - GITLAB_PERSONAL_ACCESS_TOKEN
# - GRAFANA_ADMIN_PASSWORD
# - GRAFANA_ADMIN_USER
# - HASS_TOKEN
# - HW_PASSWORD
# - HW_USER
# - LTMS_API_KEY
# - MEMORY_MCP_TOKEN
# - METRICS_PUSH_TOKEN
# - OPENCODE_AUTOMODE_MODEL
# - TELEGRAM_BOT_TOKEN
# - TELEGRAM_CHAT_ID
# - TS_API_KEY
# - TS_AUTHKEY
# - WORKER_GITEA_TOKEN
# Spawn-readiness gate (CB-306). The launcher blocks until the worker's herdr status is
# injectable (IDLE/BLOCKED/DONE) or the timeout elapses. 0 disables the gate.
# NOTE: keys are camelCase — config is bound by plain Jackson with no naming strategy and
@@ -368,10 +552,15 @@ guard:
# idleTtlSeconds → reap READY/DONE sessions idle longer than this (never BUSY/SPAWNING)
# contextCap → force-release a session after this many delegated turns
# drainTimeoutSeconds → seconds to wait for BUSY sessions on shutdown before forced teardown
# clearAfterTurn → whether a reusable worker discards its conversation context after every
# completed delegated turn (default false). Works for claude-code workers
# only — any other peer kind (e.g. opencode) logs "context reset is
# unsupported for peer kind …" once and the reset is a no-op.
# lifecycle:
# idleTtlSeconds: 300
# contextCap: 10
# drainTimeoutSeconds: 5
# clearAfterTurn: false
# Durable reply delivery (CB-307 Stage 2). OMIT this block entirely to keep the default
# in-memory, soft-state reply inbox (late worker replies are held only until a daemon bounce).
@@ -379,10 +568,14 @@ guard:
# on a durable per-target queue (agent.<target>.inbox) and survive a restart — the broker
# redelivers anything the primary had not yet drained. Production default is LavinMQ; a stock
# RabbitMQ speaks the same AMQP 0-9-1, so it is a URI-only swap.
# uri → AMQP connection URI. No trailing slash ⇒ the default vhost "/"; an empty path ("/")
# is vhost "" and will NOT connect. Encode a named vhost as .../%2Fmyvhost.
# uri → AMQP connection URI. No trailing slash ⇒ the default vhost "/"; an empty path ("/")
# is vhost "" and will NOT connect. Encode a named vhost as .../%2Fmyvhost.
# prefetch → CB-527: consumer basicQos, capping how many unacked messages the inbox holds
# in-heap per owned target (the rest sits on the broker's durable queue instead of
# growing the JVM heap). Default 32 when omitted.
# broker:
# uri: amqp://guest:guest@127.0.0.1:5672
# prefetch: 32
# Active push-to-primary (CB-307 Stage 3). When a worker reply lands with no open bridge_send,
# the ReplyPushLoop injects a *drain nudge* (never the payload) into the primary's own herdr
@@ -13,6 +13,8 @@ import dev.ltms.bridged.herdr.PaneLocator;
import dev.ltms.bridged.herdr.UnixSocketHerdrClient;
import dev.ltms.bridged.herdr.WorkspaceControl;
import dev.ltms.bridged.inject.CompletionResolver;
import dev.ltms.bridged.inject.ExhaustedPatternLookup;
import dev.ltms.bridged.inject.ExhaustionSink;
import dev.ltms.bridged.inject.Injector;
import dev.ltms.bridged.inject.StatusPoller;
import dev.ltms.bridged.inject.TurnListener;
@@ -23,6 +25,7 @@ import dev.ltms.bridged.mcp.BridgeMcp;
import dev.ltms.bridged.mcp.ConnectionIdentity;
import dev.ltms.bridged.metrics.BridgedMetrics;
import dev.ltms.bridged.metrics.Metrics;
import dev.ltms.bridged.health.FleetHealthMonitor;
import dev.ltms.bridged.mcp.PrimaryRegistry;
import dev.ltms.bridged.mcp.LsofPeerPidLookup;
import dev.ltms.bridged.mcp.LsofProcessCwdLookup;
@@ -35,6 +38,7 @@ import dev.ltms.bridged.msg.LeadHeartbeatLoop;
import dev.ltms.bridged.msg.ReplyPushLoop;
import dev.ltms.bridged.rest.BridgedApp;
import dev.ltms.bridged.session.GitWorktrees;
import dev.ltms.bridged.session.MemberSession;
import dev.ltms.bridged.session.SessionManager;
import dev.ltms.bridged.peer.PeerLauncher;
import dev.ltms.bridged.session.SessionReaper;
@@ -42,6 +46,7 @@ import dev.ltms.bridged.member.ClaudeCodeLauncher;
import dev.ltms.bridged.member.CompositePeerLauncher;
import dev.ltms.bridged.member.HerdrPeerLauncher;
import dev.ltms.bridged.member.OpenCodeLauncher;
import dev.ltms.bridged.placement.BackendQuarantine;
import io.javalin.Javalin;
import org.slf4j.Logger;
import org.slf4j.LoggerFactory;
@@ -59,6 +64,7 @@ import java.util.concurrent.atomic.AtomicReference;
import java.util.function.Function;
import java.util.function.Predicate;
import java.util.function.Supplier;
import java.util.regex.Pattern;
import java.util.stream.Collectors;
/**
@@ -77,6 +83,15 @@ public final class Bridged {
static void main(String[] args) {
Path configPath = Path.of(args.length > 0 ? args[0] : "bridged.yaml");
BridgedConfig cfg = BridgedConfig.load(configPath);
// CB-594: report which secret env vars the config actually needs, by name, before anything
// else can fail on a silently-empty one. A daemon started without a login shell (launchd)
// boots fine either way — this is the only thing that says so out loud.
reportRequiredSecrets(cfg);
// CB-596: an absent (or empty) memberCredentials: block blocks NOTHING — no credential
// name is hardcoded any more to fall back on. Say so loudly, the same way a missing
// secret is reported above, so upgrading past this commit never silently drops CB-592's
// protection.
reportMemberCredentialsGap(cfg);
// CB-559: `cfg` stays the startup snapshot — every validation and every piece of one-time
// wiring below reads it, and must, because those decisions cannot be unmade. `config` is the
// live reference the hot paths read per use. Which keys can actually move is ConfigRef's
@@ -129,20 +144,29 @@ public final class Bridged {
adapters.add(new ClaudeCodeLauncher(agents, spaces, guard,
claudeProfiles, cfg.effectiveDefaultProfile(), System::getenv,
cfg.spawnReadyTimeoutMs(), cfg.spawnReadyPollMs(),
() -> config.get().fleet()));
() -> config.get().fleet(),
() -> config.get().memberCredentials()));
}
if (!opencodeProfiles.isEmpty()) {
adapters.add(new OpenCodeLauncher(agents, spaces,
opencodeProfiles, cfg.effectiveDefaultProfile(), System::getenv,
cfg.spawnReadyTimeoutMs(), cfg.spawnReadyPollMs(),
() -> config.get().fleet()));
() -> config.get().fleet(),
() -> config.get().memberCredentials()));
}
AtomicReference<Function<String, Integer>> liveCountRef = new AtomicReference<>(_ -> 0);
// CB-578 stage B: one quarantine tracker for the whole daemon, shared between the launcher
// (checked at spawn) and the exhaustion sink wired in below (written on BACKEND_EXHAUSTED).
// The cooldown is deferred (see BridgedConfig#quarantineCooldownSeconds): it is read once
// here, at startup, and a config reload only changes it for a daemon restart.
BackendQuarantine quarantine = new BackendQuarantine(System::nanoTime,
TimeUnit.SECONDS.toNanos(cfg.quarantineCooldownSeconds()));
PeerLauncher workers = new CompositePeerLauncher(
adapters,
cfg.effectiveDefaultProfile(),
config,
profileName -> liveCountRef.get().apply(profileName));
profileName -> liveCountRef.get().apply(profileName),
quarantine);
// CB-504: under supervision (launchd/systemd) bridged can start before herdr's socket
// exists. The client itself is lazy — it connects per call — but the orphan reap below is
// the first thing that actually talks to herdr, so without this wait a boot-order race
@@ -192,34 +216,34 @@ public final class Bridged {
if (leadTerminals.size() > 1) {
log.info("leads: {} panes recognised {}", leadTerminals.size(), leadTerminals.values());
}
// CB-531: on top of the static registry, discover leads by the tab labels the operator
// writes. CB-557 moved the settings onto the lead they describe, so scanning is on whenever
// a `fleet.leaders:` entry exists — with no leads configured the supplier is a constant and
// never touches herdr, exactly as a missing `leadScan:` block used to behave.
// CB-531: on top of the legacy primary.terminal pin, discover leads by the tab labels the
// operator writes. CB-557 moved the settings onto the lead they describe, so scanning is on
// whenever a `fleet.leaders:` entry exists — with no leads configured the supplier is a
// constant and never touches herdr, exactly as a missing `leadScan:` block used to behave.
// CB-579: each lead now names its own exact `tab:` label, so one scanner discovers every
// configured lead regardless of how differently their tabs are labelled — the old
// single-shared-tabPrefix limitation (and its warning) is gone.
final Supplier<Map<String, String>> leads;
var leaders = cfg.fleet().leaders();
if (!leaders.isEmpty()) {
// One scanner, so one prefix and one interval. Distinct per-lead prefixes would need a
// scanner each; until a config actually wants that, take the first entry's settings and
// say so, rather than silently honouring one lead's prefix and dropping another's.
var scan = leaders.values().iterator().next();
Set<String> memberSpaces = cfg.profiles().values().stream()
.map(BridgedConfig.Profile::workspace)
.filter(Objects::nonNull)
.collect(Collectors.toSet());
leads = new LeadTabScanner(herdr, scan.tabPrefix(), memberSpaces, leadTerminals,
TimeUnit.SECONDS.toNanos(scan.scanIntervalSeconds()), System::nanoTime);
log.info("lead scan: tabs labelled '{}…' host a lead (rescan every {}s, member spaces {} "
+ "excluded)",
scan.tabPrefix(), scan.scanIntervalSeconds(), memberSpaces);
long distinctPrefixes = leaders.values().stream()
.map(BridgedConfig.Leader::tabPrefix).distinct().count();
if (distinctPrefixes > 1) {
log.warn("fleet.leaders declares {} different tabPrefix values; only '{}' is scanned "
+ "for. Give every lead the same tabPrefix, or leads under the others "
+ "will not be discovered.",
distinctPrefixes, scan.tabPrefix());
}
Map<String, String> tabToName = new LinkedHashMap<>();
leaders.forEach((name, leader) -> {
if (leader != null && leader.tab() != null && !leader.tab().isBlank()) {
tabToName.put(leader.tab(), name);
}
});
// One shared rescan cadence: still taken from the first entry, as before — it is an
// operational cadence, not identity, so there is no correctness reason to give every
// lead its own scanner.
int scanIntervalSeconds = leaders.values().iterator().next().scanIntervalSeconds();
leads = new LeadTabScanner(herdr, tabToName, memberSpaces,
TimeUnit.SECONDS.toNanos(scanIntervalSeconds), System::nanoTime);
log.info("lead scan: tabs {} host a lead (rescan every {}s, member spaces {} excluded)",
tabToName.keySet(), scanIntervalSeconds, memberSpaces);
} else {
leads = () -> leadTerminals;
}
@@ -252,7 +276,40 @@ public final class Bridged {
// The blocking message endpoint (CB-104) is the producer; the poller is inert until then.
// CB-106: a confirmed turn completion resolves a blocked send whose worker never replied.
Rendezvous rendezvous = new Rendezvous();
CompletionResolver completion = new CompletionResolver(agents, rendezvous);
// CB-578 stage A: classify a completion-fallback scrape that matches a profile's configured
// usage-limit refusal as BACKEND_EXHAUSTED rather than handing it back as a real answer.
// Compiled once at startup, keyed by profile name; a profile with no exhaustedPattern is
// simply absent here, so its workers keep today's completion-fallback behaviour unchanged.
Map<String, Pattern> exhaustedPatternsByProfile = new LinkedHashMap<>();
cfg.profiles().forEach((name, profile) -> {
if (profile.hasExhaustedPattern()) {
exhaustedPatternsByProfile.put(name, Pattern.compile(profile.exhaustedPattern()));
}
});
ExhaustedPatternLookup exhaustedPatterns = target -> sessions.roster().stream()
.filter(session -> target.equals(session.terminalId()))
.findFirst()
.map(session -> exhaustedPatternsByProfile.get(session.profile()))
.orElse(null);
log.info("backend-exhausted classification (CB-578 stage A): {}",
CompletionResolver.coverage(cfg.profiles().keySet(), exhaustedPatternsByProfile.keySet()));
// CB-578 stage B: on a classification that actually wins, quarantine the exhausted profile's
// CREDENTIAL — not the profile name — so a profile sharing that credential (e.g. two models
// on one OpenAI account) is refused too, not just the one that happened to report it. Reads
// the profile config live off `config`, so a credentialId edit is hot: no restart needed.
ExhaustionSink exhaustionSink = (target, reason) -> sessions.roster().stream()
.filter(session -> target.equals(session.terminalId()))
.findFirst()
.map(MemberSession::profile)
.map(profileName -> config.get().profiles().get(profileName))
.ifPresent(profile -> {
String credentialId = profile.effectiveCredentialId();
quarantine.quarantine(credentialId);
log.warn("credential '{}' quarantined for {}s (profile '{}' classified "
+ "BACKEND_EXHAUSTED): {}", credentialId,
cfg.quarantineCooldownSeconds(), profile.profile(), reason);
});
CompletionResolver completion = new CompletionResolver(agents, rendezvous, exhaustedPatterns, exhaustionSink);
// CB-113: deliver only to an available worker (its MCP is connected), never its boot window.
// CB-301: the manager's presence bridge records availability and drives SPAWNING → READY.
MemberPresence presence = sessions.asPresence();
@@ -275,9 +332,9 @@ public final class Bridged {
}
@Override
public void onDelivered(String target) {
completion.onDelivered(target);
sessions.onDelivered(target);
public void onDelivered(String target, dev.ltms.bridged.msg.TurnToken token) {
completion.onDelivered(target, token);
sessions.onDelivered(target, token);
}
@Override
@@ -285,6 +342,12 @@ public final class Bridged {
completion.onTurnFailed(target);
sessions.onTurnFailed(target);
}
@Override
public void onTurnFailed(String target, String reason) {
completion.onTurnFailed(target, reason);
sessions.onTurnFailed(target);
}
};
Injector injector = new Injector(agents, turnListener, deliverableTo(presence, leads),
presence::forget);
@@ -296,8 +359,9 @@ public final class Bridged {
// connection, so keep the reference to close it in the ordered shutdown hook.
final ReplyInbox replyInbox;
if (cfg.broker() != null && cfg.broker().isConfigured()) {
replyInbox = AmqpReplyInbox.open(cfg.broker().uri());
log.info("reply inbox: AMQP broker (durable) at {}", cfg.broker().uri());
replyInbox = AmqpReplyInbox.open(cfg.broker().uri(), cfg.broker().prefetchOrDefault());
log.info("reply inbox: AMQP broker (durable) at {} (prefetch={})",
cfg.broker().uri(), cfg.broker().prefetchOrDefault());
} else {
replyInbox = new InMemoryReplyInbox();
log.info("reply inbox: in-memory (soft-state)");
@@ -348,16 +412,50 @@ public final class Bridged {
MessageService messages = new MessageService(agents, injector, rendezvous, replyInbox,
pushLoop, metrics);
// Health is a slow whole-fleet observer. Keep it separate from the 250ms delivery poller.
final FleetHealthMonitor healthMonitor;
var healthScheduler = Executors.newSingleThreadScheduledExecutor(r ->
Thread.ofVirtual().name("bridge-health-").unstarted(r));
if (cfg.health() != null && cfg.health().isEnabled()) {
// CB-580: a member found GONE/NEVER_READY must fail whatever ticket is waiting on it,
// through the same idempotent target-wide operation CB-516 already uses on release.
healthMonitor = new FleetHealthMonitor(agents, sessions::roster, messages, healthScheduler,
System::nanoTime, cfg.health().intervalOrDefault(), messages::abandon);
String coverage = FleetHealthMonitor.coverage(true,
cfg.health().notifications() != null && cfg.health().notifications().configured());
if ("detection-only".equals(coverage)) {
log.warn("fleet health: {} (no notification sink configured)", coverage);
} else {
log.info("fleet health: {}", coverage);
}
healthMonitor.start();
} else {
healthMonitor = null;
healthScheduler.shutdownNow();
}
// CB-520: the reply inbox only consumes for agents this gateway owns. own on acquire,
// release on teardown. Do this before CB-516 so the inbox is owned before any reply can land.
sessions.onAcquire(replyInbox::own);
// CB-516: releasing a worker must fail whatever send was waiting on it. Without this a
// torn-down delegation kept reporting PENDING until the 30-minute async timeout, and never
// reached /metrics — the delegation was unresolvable and nothing said so.
sessions.onRelease(terminal -> {
messages.abandon(terminal, "the worker session was released before it replied");
replyInbox.release(terminal);
primaryRegistry.forgetDelegation(terminal); // CB-532: don't leak the lead binding
sessions.onRelease(detail -> {
// CB-578 stage C, acceptance criterion 10: a failed ticket's detail should tell a lead
// where to re-dispatch onto the same tree, not just that the worker vanished.
String reason = "the worker session was released before it replied";
if (detail.worktreePath() != null) {
reason += "; worktree=" + detail.worktreePath() + " branch=" + detail.branch()
+ " snapshot=" + (detail.snapshotRef() != null ? detail.snapshotRef() : "none");
}
// CB-584 (issue #65 criterion 5): also name the agent session, so a lead can resume the
// member's conversation instead of only re-dispatching a fresh one onto the same files.
if (detail.agentSessionId() != null) {
reason += " agentSessionId=" + detail.agentSessionId();
}
messages.abandon(detail.terminalId(), reason);
replyInbox.release(detail.terminalId());
primaryRegistry.forgetDelegation(detail.terminalId()); // CB-532: don't leak the lead binding
});
// MCP server face (CB-105): bridge_send/bridge_reply/bridge_status, mounted at /mcp.
@@ -387,7 +485,16 @@ public final class Bridged {
profile -> {
var configured = config.get().profiles().get(profile);
return configured == null ? null : configured.maxLoad();
}, () -> config.get().profiles().keySet(), System::nanoTime));
}, () -> config.get().profiles().keySet(), System::nanoTime),
new BridgeMcp.HealthCoverageSource(() -> {
var health = config.get().health();
return FleetHealthMonitor.coverage(health != null && health.isEnabled(),
health != null && health.notifications() != null && health.notifications().configured());
}),
new BridgeMcp.QuarantineSource(profile -> {
var configured = config.get().profiles().get(profile);
return configured == null ? null : configured.effectiveCredentialId();
}, quarantine));
// CB-559: opt-in config reload. With no `configReload:` block nothing is constructed, so an
// upgraded daemon behaves exactly as before — the file is read once at boot and never again.
@@ -408,6 +515,7 @@ public final class Bridged {
messages.close();
pushLoop.close();
if (heartbeat != null) heartbeat.close(); // CB-551: stop the idle-lead heartbeat scheduler
if (healthMonitor != null) healthMonitor.stop();
if (configWatcher != null) configWatcher.stop(); // CB-559: stop polling the config file
mcp.close();
if (reaper != null) reaper.stop();
@@ -452,6 +560,91 @@ public final class Bridged {
return target -> presence.isPresent(target) || leads.get().containsKey(target);
}
/**
* CB-594: which env vars the loaded config actually needs, and why — every non-{@code
* subscription} profile's {@code tokenEnv} (a subscription profile never reads one, see
* {@link BridgedConfig.Profile#isSubscription()}), plus every profile's {@code gitTokenEnv}
* where set (opt-in). Derived from the config, not hard-coded, so a new profile is covered for
* free. A var required by more than one profile is one entry naming every profile that needs
* it. Deliberately excludes {@code auth.tokenEnv}: that one is already enforced loudly, by a
* startup throw in {@code main()} — about 370 lines <em>below</em> this method's call site
* ({@link #reportRequiredSecrets(BridgedConfig)}), not a few lines above it. That throw only
* fires when {@code auth.mode: token} is configured; under the default loopback-trust mode it
* never runs, and {@code auth.tokenEnv} is simply not required.
*
* <p>Package-private and pure (no I/O, no logging) so the derivation is unit-testable without
* capturing log output; {@link #reportRequiredSecrets(BridgedConfig)} is the logging caller.
*/
static Map<String, List<String>> requiredSecretEnvVars(BridgedConfig cfg) {
Map<String, List<String>> requiredBy = new LinkedHashMap<>();
cfg.profiles().forEach((name, profile) -> {
if (!profile.isSubscription()) {
requiredBy.computeIfAbsent(profile.tokenEnv(), _ -> new ArrayList<>())
.add("profile '" + name + "' tokenEnv");
}
if (profile.hasGitToken()) {
requiredBy.computeIfAbsent(profile.gitTokenEnv(), _ -> new ArrayList<>())
.add("profile '" + name + "' gitTokenEnv");
}
});
return requiredBy;
}
/**
* CB-594: log, by name only, which required env vars (see {@link #requiredSecretEnvVars}) are
* set in the daemon's own process environment — the environment every profile's {@code
* tokenEnv}/{@code gitTokenEnv} is read from at spawn time (see
* {@code HerdrPeerLauncher.resolveEnv}). Never logs a value, a prefix, or a length.
*
* <p>A missing entry only warns — it must never refuse to start. A daemon that boots and says
* what is wrong is strictly more useful than one that will not boot at all.
*/
private static void reportRequiredSecrets(BridgedConfig cfg) {
Map<String, List<String>> requiredBy = requiredSecretEnvVars(cfg);
if (requiredBy.isEmpty()) {
log.info("startup secrets: no profile references a token env var — nothing to check");
return;
}
Map<String, String> env = System.getenv();
requiredBy.forEach((varName, sources) -> {
String value = env.get(varName);
if (value != null && !value.isBlank()) {
log.info("startup secret {}: set ({})", varName, String.join(", ", sources));
} else {
log.warn("startup secret {}: MISSING ({}) — the daemon will start anyway, and this "
+ "failure stays invisible until a worker actually needs it. Fix "
+ "${SHARED_ENV}/tools/secrets.sh and restart bridged from a LOGIN "
+ "shell (see scripts/redeploy-bridged.sh).",
varName, String.join(", ", sources));
}
});
}
/**
* CB-596: {@code known:} empty (block absent entirely, or present but empty) means {@link
* BridgedConfig.MemberCredentials#blockedSet()} is empty too — every member pane inherits the
* operator's whole secret store, unblocked, exactly the defect this ticket fixes. Unlike a
* missing token ({@link #reportRequiredSecrets}), there is no name to point at: the point is
* that the block itself is missing. Warn once at startup and say what to add; never refuse to
* start over it — see {@link #reportRequiredSecrets} for why a daemon that boots and says
* what is wrong beats one that will not boot at all.
*
* <p>Package-private so the test can capture the log directly, the same way {@link
* #requiredSecretEnvVars} is exposed for {@link #reportRequiredSecrets}'s own test.
*/
static void reportMemberCredentialsGap(BridgedConfig cfg) {
BridgedConfig.MemberCredentials creds = cfg.memberCredentials();
if (creds != null && !creds.known().isEmpty()) {
log.info("memberCredentials: {} known name(s), {} allowed — blocking {} on every spawn",
creds.known().size(), creds.allow().size(), creds.blockedSet().size());
return;
}
log.warn("memberCredentials: absent or empty — the daemon will start anyway, and every "
+ "member pane inherits the operator's WHOLE secret store, unblocked (CB-592's "
+ "protection is lost). Add a memberCredentials: block (policy/allow/known) to "
+ "bridged.yaml — see bridged.example.yaml — and restart.");
}
/**
* Poll herdr's {@code ping} until it answers or {@link #HERDR_WAIT_SECONDS} elapses (CB-504).
*
@@ -5,7 +5,9 @@ import com.fasterxml.jackson.core.JsonParser;
import com.fasterxml.jackson.core.JsonToken;
import com.fasterxml.jackson.databind.ObjectMapper;
import com.fasterxml.jackson.dataformat.yaml.YAMLFactory;
import dev.ltms.bridged.msg.AmqpReplyInbox;
import dev.ltms.bridged.peer.MemberRole;
import dev.ltms.bridged.placement.PlacementPolicies;
import org.slf4j.Logger;
import org.slf4j.LoggerFactory;
@@ -58,6 +60,15 @@ import java.util.Set;
* {@code fixed} (default), {@code round-robin}, or {@code weighted}
* @param auth API authentication mode ({@code null} → {@code loopback-trust}, the
* historical behaviour), CB-501
* @param quarantineCooldownSeconds how long a credential stays quarantined after a
* {@code BACKEND_EXHAUSTED} classification (CB-578 stage B); {@code null}/{@code
* <=0} → {@link #DEFAULT_QUARANTINE_COOLDOWN_SECONDS}. Baked once into the
* {@code BackendQuarantine} built at startup, so it is DEFERRED: changing it
* needs a restart, and a quarantine already running keeps whatever cooldown was
* live when it started.
* @param memberCredentials deny-by-default policy (CB-596) for which of the operator's own host
* credentials a spawned member's pane inherits. {@code null} (the block
* omitted) blocks nothing — see {@link MemberCredentials}.
*/
@JsonIgnoreProperties(ignoreUnknown = true)
public record BridgedConfig(
@@ -73,9 +84,26 @@ public record BridgedConfig(
Primary primary,
Fleet fleet,
LeadHeartbeat leadHeartbeat,
Health health,
String placement,
Auth auth,
ConfigReload configReload) {
ConfigReload configReload,
Integer quarantineCooldownSeconds,
MemberCredentials memberCredentials) {
/** Back-compat form before the CB-596 {@code memberCredentials:} block was added. */
public BridgedConfig(Bind bind, String herdrSocket, Map<String, Profile> profiles, Guard guard,
String worktreeRoot, Lifecycle lifecycle, Integer spawnReadyTimeoutMs,
Integer spawnReadyPollMs, Broker broker, Primary primary, Fleet fleet,
LeadHeartbeat leadHeartbeat, Health health, String placement, Auth auth,
ConfigReload configReload, Integer quarantineCooldownSeconds) {
this(bind, herdrSocket, profiles, guard, worktreeRoot, lifecycle, spawnReadyTimeoutMs,
spawnReadyPollMs, broker, primary, fleet, leadHeartbeat, health, placement, auth,
configReload, quarantineCooldownSeconds, null);
}
/** Default cooldown (CB-578 stage B) when {@code quarantineCooldownSeconds} is absent/non-positive. */
public static final int DEFAULT_QUARANTINE_COOLDOWN_SECONDS = 1800;
/** Back-compat 14-arg form — no {@code configReload:} block, so file watching stays off. */
public BridgedConfig(Bind bind, String herdrSocket, Map<String, Profile> profiles, Guard guard,
@@ -83,7 +111,27 @@ public record BridgedConfig(
Integer spawnReadyPollMs, Broker broker, Primary primary, Fleet fleet,
LeadHeartbeat leadHeartbeat, String placement, Auth auth) {
this(bind, herdrSocket, profiles, guard, worktreeRoot, lifecycle, spawnReadyTimeoutMs,
spawnReadyPollMs, broker, primary, fleet, leadHeartbeat, placement, auth, null);
spawnReadyPollMs, broker, primary, fleet, leadHeartbeat, null, placement, auth, null, null);
}
/** Back-compat form before the optional {@code health:} block was added. */
public BridgedConfig(Bind bind, String herdrSocket, Map<String, Profile> profiles, Guard guard,
String worktreeRoot, Lifecycle lifecycle, Integer spawnReadyTimeoutMs,
Integer spawnReadyPollMs, Broker broker, Primary primary, Fleet fleet,
LeadHeartbeat leadHeartbeat, String placement, Auth auth, ConfigReload configReload) {
this(bind, herdrSocket, profiles, guard, worktreeRoot, lifecycle, spawnReadyTimeoutMs,
spawnReadyPollMs, broker, primary, fleet, leadHeartbeat, null, placement, auth, configReload, null);
}
/** Back-compat form before the CB-578 stage B {@code quarantineCooldownSeconds} field was added. */
public BridgedConfig(Bind bind, String herdrSocket, Map<String, Profile> profiles, Guard guard,
String worktreeRoot, Lifecycle lifecycle, Integer spawnReadyTimeoutMs,
Integer spawnReadyPollMs, Broker broker, Primary primary, Fleet fleet,
LeadHeartbeat leadHeartbeat, Health health, String placement, Auth auth,
ConfigReload configReload) {
this(bind, herdrSocket, profiles, guard, worktreeRoot, lifecycle, spawnReadyTimeoutMs,
spawnReadyPollMs, broker, primary, fleet, leadHeartbeat, health, placement, auth,
configReload, null);
}
/**
@@ -141,19 +189,58 @@ public record BridgedConfig(
* @param cwd fixed working directory for this profile's workers (CB-112 "told otherwise");
* {@code null}/blank → inherit the primary's cwd, else the daemon's
* @param parityOverlay repo-relative paths copied primary→worktree for config parity; null/empty
* defaults to a sensible set of local config files
* defaults to a sensible set of local config files.
* <p><strong>Every overlay path must be gitignored or tracked-and-skipped.</strong>
* CB-576 made {@code release()} preserve a worktree that {@code git status
* --porcelain} reports as dirty, and it deliberately counts untracked files —
* the work lost in CB-576 was a new file nobody had added. So an overlay path
* that is neither gitignored nor tracked lands in every worktree as an
* untracked file, makes every {@code COMPLETED} release preserve, and
* worktrees then accumulate with no error anywhere.
* <p>Checked on 2026-08-15 (CB-581): inert as configured. Tracked overlay
* files carry {@code --skip-worktree} so {@code --porcelain} cannot see them,
* {@code bridged.yaml} is gitignored, and the default pair {@code .env} /
* {@code .envrc} does not exist in this repo. Note the default applies to
* <em>every</em> profile, so creating either file at the repo root is enough
* to make it live. Add a new overlay path to {@code .gitignore} in the same
* change that adds it here.
* <p>CB-578 stage C raised the stakes: a preserved dirty worktree is now also
* committed to {@code refs/wip/<branch>} via {@code git add -A}. The overlay
* carries the primary's own environment files into the worktree, so a
* non-gitignored overlay path no longer merely sits there as an untracked
* file — it gets committed into a git object that survives the worktree's
* removal. {@code add -A} respects {@code .gitignore}, which is exactly what
* keeps that from happening, so the gitignore rule above is now what stops a
* secret from reaching a durable commit.
* @param gitTokenEnv name of the host env var holding the git-forge API token; when set, its
* value is injected as {@code GITEA_TOKEN} so the worker can open its own PR
* at checkpoint (CB-302). {@code null}/blank ⇒ no token is injected
* (minimal-grant default — push over SSH stays free, PR-create is opt-in)
* @param gitHostEnv name of the host env var holding the forge host (default {@code GITEA_HOST});
* injected as {@code GITEA_HOST} <em>only</em> when {@code gitTokenEnv} is set
* @param weight relative selection weight for {@code placement: weighted}. Absent or
* non-positive ⇒ 1.0. Weights are normalised by the policy, so they need
* not sum to 1.0.
* @param maxLoad max live workers allowed on this profile at one time; absent or
* non-positive ⇒ unlimited. Live means any session the registry still owns
* (acquired and not yet released), in any state.
* @param weight relative selection weight for automatic placement ({@code weighted},
* {@code round-robin}, and {@code fixed}'s fallback walk). Absent ⇒ 1.0.
* An explicit {@code 0} or negative value (CB-554) means "never
* auto-select this profile": it is excluded from every automatic policy's
* pool the same way a quarantined candidate is (see
* {@code PlacementPolicyUtil}). This does not make the profile
* unreachable — an explicit {@code bridge_spawn{profile:"..."}} bypasses
* placement entirely and still resolves it. Weights among the remaining
* (non-excluded) candidates need not sum to 1.0; only their ratios matter.
* @param maxLoad max live workers allowed on this profile at one time. Absent
* (unset/{@code null}) ⇒ unlimited — most profiles rely on this. An
* explicit {@code 0} (CB-585) means "cap this profile at zero live
* members": it is excluded from every automatic policy's candidate pool
* the same way a {@code weight <= 0} profile is (see
* {@code PlacementPolicyUtil.available()}, which already treats "at cap"
* and "excluded" alike), and an explicit
* {@code bridge_spawn{profile:"..."}} against it is refused too (see
* {@code CompositePeerLauncher.enforceMaxLoad}) — a cap is a capacity
* statement that does not stop being true just because the profile was
* named directly. A negative value has no sane meaning (there is no
* "excluded" to degrade to below zero) and is refused at config load
* instead, naming the profile and the key. Live means any session the
* registry still owns (acquired and not yet released), in any state.
* @param kind which peer launcher spawns this profile: {@code "claude-code"} (default —
* the {@link dev.ltms.bridged.member.ClaudeCodeLauncher}) or {@code "opencode"}.
* The {@code CompositePeerLauncher} routes {@code spawn}/reap by this value, so
@@ -171,6 +258,23 @@ public record BridgedConfig(
* opposite intents). For the same reason, an {@code env:} entry naming
* {@code ANTHROPIC_BASE_URL} or {@code ANTHROPIC_AUTH_TOKEN} is refused at
* config load (CB-542): on the subscription path no guard would vet it.
* @param exhaustedPattern regex matched against a completion-fallback scrape (CB-578 stage A) to
* classify a turn that ended with no {@code bridge_reply} as the backend
* having refused on a subscription usage limit, rather than a real answer.
* {@code null}/blank ⇒ the classification never fires for this profile and
* today's completion-fallback behaviour is unchanged. Every backend words
* its refusal differently, so this is config, never a vendor string in code.
* @param credentialId the shared account this profile authenticates as (CB-578 stage B). Two
* or more profiles setting the <em>same</em> non-blank value are quarantined
* together by one {@code exhaustedPattern} classification on any one of
* them — the case a single OpenAI (or any other) credential backing several
* profiles ({@code sol}, {@code terra}, ...) needs, so a spawn does not walk
* straight onto the other profile sharing the same exhausted account.
* {@code null}/blank ⇒ this profile's {@link #effectiveCredentialId()} is
* its own name, so it quarantines alone — today's behaviour for every
* profile that does not opt in. Read live off the current config, so it is
* HOT: a change takes effect on the next exhaustion classification / spawn,
* no restart needed.
*/
@JsonIgnoreProperties(ignoreUnknown = true)
public record Profile(String profile, String baseUrl, String model,
@@ -183,7 +287,9 @@ public record BridgedConfig(
Map<String, String> env,
Float weight,
Integer maxLoad,
Boolean subscription) {
Boolean subscription,
String exhaustedPattern,
String credentialId) {
/** Peer kind spawned by {@link dev.ltms.bridged.member.ClaudeCodeLauncher} (the default). */
public static final String KIND_CLAUDE_CODE = "claude-code";
@@ -223,9 +329,25 @@ public record BridgedConfig(
// checkpoints need only set gitTokenEnv; it is injected only alongside a resolved token.
gitHostEnv = (gitHostEnv == null || gitHostEnv.isBlank()) ? "GITEA_HOST" : gitHostEnv;
env = (env == null) ? Map.of() : Map.copyOf(env);
weight = (weight == null || weight <= 0.0f) ? 1.0f : weight;
maxLoad = (maxLoad == null || maxLoad <= 0) ? null : maxLoad;
// CB-554: absent still means 1.0, but an explicit non-positive value must survive as
// "excluded from automatic selection" (PlacementCandidate.excluded(), weight <= 0), not
// get coerced back up to 1.0 — that coercion was the bug (weight: 0 looked like "never
// pick me" and actually meant "pick me as often as anyone else").
weight = (weight == null) ? 1.0f : Math.max(weight, 0.0f);
// CB-585: absent still means unlimited (null), but an explicit maxLoad: 0 must survive
// as "capped at zero," not get coerced back up to null/unlimited — that coercion was
// the bug (maxLoad: 0 read as "never run anything here" and behaved as the opposite,
// the one throttle a subscription: true profile has against the operator's own paid
// plan). A negative value is refused earlier, at config load (rejectNegativeMaxLoad),
// so it never reaches this constructor and needs no clamping here.
subscription = (subscription != null && subscription) ? Boolean.TRUE : Boolean.FALSE;
// exhaustedPattern stays null when unset/blank (opt-in) — no defaulting, no vendor
// wording: an unconfigured profile keeps today's completion-fallback behaviour exactly.
exhaustedPattern = (exhaustedPattern == null || exhaustedPattern.isBlank()) ? null : exhaustedPattern;
// credentialId stays null when unset/blank — effectiveCredentialId() is where the
// "quarantines alone" fallback actually lives, so today's behaviour needs no defaulting
// here at all.
credentialId = (credentialId == null || credentialId.isBlank()) ? null : credentialId;
}
/**
@@ -238,7 +360,7 @@ public record BridgedConfig(
String placement, String workspace, String tabLabel, String mcpUrl,
String cwd, List<String> parityOverlay) {
this(profile, baseUrl, model, configDir, tokenEnv, argv, placement, workspace, tabLabel,
mcpUrl, cwd, parityOverlay, null, null, null, null, null, null, null);
mcpUrl, cwd, parityOverlay, null, null, null, null, null, null, null, null);
}
/**
@@ -250,7 +372,7 @@ public record BridgedConfig(
String placement, String workspace, String tabLabel, String mcpUrl,
String cwd, List<String> parityOverlay, String gitTokenEnv, String gitHostEnv) {
this(profile, baseUrl, model, configDir, tokenEnv, argv, placement, workspace, tabLabel,
mcpUrl, cwd, parityOverlay, gitTokenEnv, gitHostEnv, null, null, null, null, null);
mcpUrl, cwd, parityOverlay, gitTokenEnv, gitHostEnv, null, null, null, null, null, null);
}
/**
@@ -263,13 +385,14 @@ public record BridgedConfig(
String cwd, List<String> parityOverlay, String gitTokenEnv, String gitHostEnv,
String kind) {
this(profile, baseUrl, model, configDir, tokenEnv, argv, placement, workspace, tabLabel,
mcpUrl, cwd, parityOverlay, gitTokenEnv, gitHostEnv, kind, null, null, null, null);
mcpUrl, cwd, parityOverlay, gitTokenEnv, gitHostEnv, kind, null, null, null, null, null);
}
/** A copy with {@code profile} set — used to default a profile to its {@code workers} key. */
public Profile withProfile(String p) {
return new Profile(p, baseUrl, model, configDir, tokenEnv, argv, placement, workspace, tabLabel,
mcpUrl, cwd, parityOverlay, gitTokenEnv, gitHostEnv, kind, env, weight, maxLoad, subscription);
mcpUrl, cwd, parityOverlay, gitTokenEnv, gitHostEnv, kind, env, weight, maxLoad, subscription,
exhaustedPattern, credentialId);
}
/** True when this profile is served by the Claude Code adapter (the default kind). */
@@ -302,7 +425,39 @@ public record BridgedConfig(
String cwd, List<String> parityOverlay, String gitTokenEnv, String gitHostEnv,
String kind, Map<String, String> env, Float weight, Integer maxLoad) {
this(profile, baseUrl, model, configDir, tokenEnv, argv, placement, workspace, tabLabel,
mcpUrl, cwd, parityOverlay, gitTokenEnv, gitHostEnv, kind, env, weight, maxLoad, null);
mcpUrl, cwd, parityOverlay, gitTokenEnv, gitHostEnv, kind, env, weight, maxLoad, null, null);
}
/**
* Backward-compatible constructor without the CB-578 stage B {@code credentialId} — the
* profile quarantines alone under its own name (see {@link #effectiveCredentialId()}). Keeps
* pre-stage-B call sites (and any YAML that omits the key) compiling and behaving identically.
*/
public Profile(String profile, String baseUrl, String model,
String configDir, String tokenEnv, List<String> argv,
String placement, String workspace, String tabLabel, String mcpUrl,
String cwd, List<String> parityOverlay, String gitTokenEnv, String gitHostEnv,
String kind, Map<String, String> env, Float weight, Integer maxLoad,
Boolean subscription, String exhaustedPattern) {
this(profile, baseUrl, model, configDir, tokenEnv, argv, placement, workspace, tabLabel,
mcpUrl, cwd, parityOverlay, gitTokenEnv, gitHostEnv, kind, env, weight, maxLoad,
subscription, exhaustedPattern, null);
}
/** True when this profile's CB-578 stage A backend-exhausted classification is configured. */
public boolean hasExhaustedPattern() {
return exhaustedPattern != null;
}
/**
* The credential group this profile quarantines with (CB-578 stage B): the configured
* {@link #credentialId} when set, else this profile's own name — so an unconfigured profile
* quarantines alone, exactly as it did before this field existed. Two profiles that set the
* same non-blank {@code credentialId} share one quarantine: a {@code BACKEND_EXHAUSTED}
* classification on either one quarantines both.
*/
public String effectiveCredentialId() {
return (credentialId == null || credentialId.isBlank()) ? profile : credentialId;
}
/** True when this profile's workers are granted a forge token to open their own PR (CB-302). */
@@ -377,22 +532,41 @@ public record BridgedConfig(
boolean clearAfterTurn) {
}
/** Optional fleet detection. A missing block stays dormant. */
@JsonIgnoreProperties(ignoreUnknown = true)
public record Health(Boolean enabled, Integer intervalSeconds, Integer workingSuspectAfterSeconds,
Integer paneProbeIntervalSeconds, Notifications notifications) {
public boolean isEnabled() { return Boolean.TRUE.equals(enabled); }
public int intervalOrDefault() { return Math.max(15, intervalSeconds == null ? 30 : intervalSeconds); }
public record Notifications(String mode) {
public boolean configured() { return "webhook".equalsIgnoreCase(mode); }
}
}
/**
* External AMQP broker for durable, cross-restart reply delivery (CB-307 Stage 2). Its mere
* presence swaps the in-memory {@code ReplyInbox} for the AMQP-backed adapter; absent, bridged
* stays soft-state. Production default is LavinMQ; a stock RabbitMQ speaks the same AMQP 0-9-1
* and is a URI-only swap.
*
* @param uri AMQP connection URI, e.g. {@code amqp://guest:guest@127.0.0.1:5672/}. Blank/{@code null}
* ⇒ the broker block is treated as absent (in-memory adapter).
* @param uri AMQP connection URI, e.g. {@code amqp://guest:guest@127.0.0.1:5672/}. Blank/
* {@code null} ⇒ the broker block is treated as absent (in-memory adapter).
* @param prefetch CB-527: the consumer's {@code basicQos} prefetch count, bounding how many
* unacked messages the AMQP inbox holds in-heap per owned target. {@code null}/
* non-positive ⇒ {@link AmqpReplyInbox#DEFAULT_PREFETCH}.
*/
@JsonIgnoreProperties(ignoreUnknown = true)
public record Broker(String uri) {
public record Broker(String uri, Integer prefetch) {
/** True when a usable broker URI is configured (an empty block does not enable AMQP). */
public boolean isConfigured() {
return uri != null && !uri.isBlank();
}
/** The prefetch to use, defaulting to {@link AmqpReplyInbox#DEFAULT_PREFETCH} when unset. */
public int prefetchOrDefault() {
return (prefetch != null && prefetch > 0) ? prefetch : AmqpReplyInbox.DEFAULT_PREFETCH;
}
}
/**
@@ -437,24 +611,32 @@ public record BridgedConfig(
* pre-existed, which is why it had to be recognised by configuration rather than created. With
* {@code profile} and {@code instances} the daemon may stand one up when none is live, so the
* pane no longer has to exist before the daemon does. Recognition still comes first: a lead
* already running under {@code tabPrefix} is adopted, and only the shortfall is launched.
* already running in its configured {@code tab} is adopted, and only the shortfall is launched.
*
* <p><b>{@code tab} replaced {@code terminal} (CB-579).</b> A herdr {@code terminal_id} changes
* every time the lead's session restarts, so pinning one cost a config edit and a daemon restart
* per restart. A tab is stable: a human opens it once, it holds exactly one pane, and its label
* survives restarts of the agent inside it — so identity is now the tab label alone.
*
* @param profile the {@code profiles:} entry to launch this lead on when one must
* be created; {@code null} ⇒ recognise-only, never create
* @param terminal the lead's herdr {@code terminal_id} when pinned by hand; the only
* field identity depends on. {@code null} ⇒ found by {@code tabPrefix}
* @param tab the exact tab label hosting this lead, matched case-insensitively;
* the only field identity depends on. Required — a lead with no
* {@code tab} can never be discovered, launched or not
* @param instances how many of this lead should be live (default 1). The daemon
* launches only the shortfall, so a restart adopts rather than doubles
* @param tabPrefix label prefix marking this lead's tab, matched case-insensitively;
* the remainder is the lead's name ({@code "lead: opus"} →
* {@code opus}). Default {@code "lead:"}
* @param tabPrefix no longer used to find a lead's tab — {@code tab} is matched
* exactly. Its only remaining job is the startup collision guard
* ({@link #validateLeadTabPrefixes()}), which still uses it to refuse
* a worker {@code tabLabel} template that could be misread as a lead.
* Default {@code "lead:"}
* @param scanIntervalSeconds how long a tab scan is cached before herdr is asked again; also the
* worst case before a newly-labelled tab is recognised. Default 10
* @param kind which agent runs there ({@code claude}, {@code opencode}, …)
* @param model the model or selector it runs, for operators reading the roster
*/
@JsonIgnoreProperties(ignoreUnknown = true)
public record Leader(String profile, String terminal, Integer instances, String tabPrefix,
public record Leader(String profile, String tab, Integer instances, String tabPrefix,
Integer scanIntervalSeconds, String kind, String model,
String workspace, String cwd) {
@@ -472,12 +654,13 @@ public record BridgedConfig(
(scanIntervalSeconds == null || scanIntervalSeconds <= 0) ? 10 : scanIntervalSeconds;
workspace = (workspace == null || workspace.isBlank())
? DEFAULT_WORKSPACE : workspace.strip();
tab = (tab == null || tab.isBlank()) ? null : tab.strip();
}
/** Back-compat 7-arg form — no workspace or cwd, so both take their defaults. */
public Leader(String profile, String terminal, Integer instances, String tabPrefix,
public Leader(String profile, String tab, Integer instances, String tabPrefix,
Integer scanIntervalSeconds, String kind, String model) {
this(profile, terminal, instances, tabPrefix, scanIntervalSeconds, kind, model, null, null);
this(profile, tab, instances, tabPrefix, scanIntervalSeconds, kind, model, null, null);
}
/** True when this lead may be launched by the daemon rather than only recognised. */
@@ -485,9 +668,9 @@ public record BridgedConfig(
return profile != null && !profile.isBlank() && instances > 0;
}
/** The tab label an auto-launched instance of this lead gets — what the scanner reads back. */
public String tabLabel(String name) {
return tabPrefix + " " + name;
/** The tab label an auto-launched instance of this lead gets — its configured {@code tab}. */
public String tabLabel() {
return tab;
}
}
@@ -686,28 +869,20 @@ public record BridgedConfig(
}
/**
* The terminal → lead-name map that {@link dev.ltms.bridged.auth.CallerResolver} resolves
* against, merging the {@code leaders:} registry with the legacy singular {@code primary:} pin.
* The terminal → lead-name map seeded from the legacy singular {@code primary:} pin (CB-530).
*
* <p>Precedence: an explicit {@code leaders:} entry wins over the {@code primary:} pin for the
* same terminal. The pin is the older, less expressive spelling of the same fact, so when both
* name a pane the named entry is the one an operator meant. The pin is still honoured on its
* own — a config carrying only {@code primary:} behaves exactly as it did before CB-530.
* <p>{@code fleet.leaders} no longer carries a per-entry terminal pin (CB-579): a lead's identity
* comes from its {@code tab} alone, resolved live by {@code LeadTabScanner}. This method now
* exists only for the {@code primary.terminal} fallback — a config that never migrated off it
* still resolves that one pane as a lead named {@code "primary"}, exactly as before CB-530.
*
* @return an unmodifiable map, empty when neither block is configured (nothing is pinned, and
* every pane therefore resolves as a worker — the pre-CB-307 behaviour)
* @return an unmodifiable map, empty when {@code primary.terminal} is not configured (nothing is
* pinned, and every pane therefore resolves as a worker — the pre-CB-307 behaviour)
*/
public Map<String, String> leaderTerminals() {
Map<String, String> byTerminal = new LinkedHashMap<>();
if (fleet != null) {
fleet.leaders().forEach((name, leader) -> {
if (leader != null && leader.terminal() != null && !leader.terminal().isBlank()) {
byTerminal.put(leader.terminal(), name);
}
});
}
if (primary != null && primary.terminal() != null && !primary.terminal().isBlank()) {
byTerminal.putIfAbsent(primary.terminal(), "primary");
byTerminal.put(primary.terminal(), "primary");
}
return Collections.unmodifiableMap(byTerminal);
}
@@ -766,6 +941,58 @@ public record BridgedConfig(
}
}
/**
* CB-596: which of the operator's own host credentials a spawned member's pane may inherit.
*
* <p>A herdr pane runs a login shell that re-sources the operator's own secret store, so a
* member inherits every credential the operator's shell holds — measured at 31 names on this
* host, of which only one ({@code GITEA_ACCESS_TOKEN}) used to be blocked, and that block was a
* single name hardcoded in {@link HerdrPeerLauncher} rather than driven by config (gitea issue
* #82). This record replaces that hardcoded shadow with a config-driven one.
*
* <p><b>deny-by-default, not a deny-list.</b> A deny-list (block these specific names, let
* everything else through) is silently wrong the moment a new secret is added to the operator's
* store — nothing would ever report it. Deny-by-default inverts that: {@link #known} bounds the
* blast radius to names the operator has actually enumerated, and every one of them is blocked
* UNLESS it is also in {@link #allow}. A name that shows up in neither list is not silently
* allowed — see {@code HerdrPeerLauncher}'s gap detector, which logs it.
*
* @param policy how the block is computed. Only {@link #POLICY_DENY_BY_DEFAULT} is understood
* today; {@code null}/blank defaults to it. An operator's own deny-list is
* deliberately not supported — see above.
* @param allow credential names a member legitimately needs (e.g. the gateway token it reaches
* the LLM through, the repo-scoped forge token it opens its own PR with). Every
* name here is left unmentioned in the pane's env overlay, so the value the pane's
* own (login) shell exports passes through untouched.
* @param known every credential name the operator's store is known to export. Every name here
* that is NOT also in {@link #allow} is overlaid with a non-secret sentinel value,
* shadowing whatever the pane's login shell would otherwise export for it.
*/
@JsonIgnoreProperties(ignoreUnknown = true)
public record MemberCredentials(String policy, List<String> allow, List<String> known) {
/** The only policy this build understands: block every {@code known} name not in {@code allow}. */
public static final String POLICY_DENY_BY_DEFAULT = "deny-by-default";
public MemberCredentials {
policy = (policy == null || policy.isBlank()) ? POLICY_DENY_BY_DEFAULT : policy.toLowerCase();
allow = allow == null ? List.of() : List.copyOf(allow);
known = known == null ? List.of() : List.copyOf(known);
}
/** {@link #allow} as a set, for membership checks. */
public Set<String> allowSet() {
return Set.copyOf(allow);
}
/** {@link #known} minus {@link #allow} — the names a spawn must shadow. */
public Set<String> blockedSet() {
Set<String> blocked = new java.util.LinkedHashSet<>(known);
blocked.removeAll(allowSet());
return blocked;
}
}
/**
* The candidate profiles an unqualified spawn of {@code role} chooses between, in definition
* order (CB-557).
@@ -808,20 +1035,35 @@ public record BridgedConfig(
/**
* Top-level keys this version understands. Used only to warn about the rest — see
* {@link #warnUnknownTopLevelKeys}. Keep in step with the record components.
*
* <p>Package-private (not {@code private}) so a test can assert every key here is documented in
* {@code bridged.example.yaml} — the only committed description of the config schema, since
* {@code bridged.yaml} itself is gitignored.
*/
private static final Set<String> KNOWN_TOP_LEVEL_KEYS = Set.of(
static final Set<String> KNOWN_TOP_LEVEL_KEYS = Set.of(
"bind", "herdrSocket", "profiles", "guard", "worktreeRoot",
"lifecycle", "spawnReadyTimeoutMs", "spawnReadyPollMs", "broker", "primary", "fleet",
"leadHeartbeat", "placement", "auth", "configReload");
"leadHeartbeat", "health", "placement", "auth", "configReload", "quarantineCooldownSeconds",
"memberCredentials");
/** Load and validate config from {@code path}. */
public static BridgedConfig load(Path path) {
try {
String yaml = Files.readString(path);
rejectRenamedTopLevelKeys(yaml);
rejectLeaderTerminalKey(yaml);
warnUnknownTopLevelKeys(yaml, path);
rejectDuplicateMemberSlots(yaml);
rejectNegativeMaxLoad(yaml);
rejectUnknownKind(yaml);
rejectUnknownAuthMode(yaml);
rejectUnknownPlacement(yaml);
rejectUnknownMemberCredentialsPolicy(yaml);
BridgedConfig cfg = YAML.readValue(yaml, BridgedConfig.class);
// CB-606: validated here, eagerly, using PlacementPolicies.fromName as the single source
// of truth — not lazily at first spawn (see CompositePeerLauncher's placementPolicy
// Supplier), where a bad name would still start a daemon that looks healthy.
rejectUnknownPlacementPolicy(cfg.placement());
return cfg.withDefaults();
} catch (IOException e) {
throw new UncheckedIOException("cannot read bridged config at " + path, e);
@@ -1043,6 +1285,273 @@ public record BridgedConfig(
}
}
/**
* Reject a config whose {@code fleet.leaders.<name>} still carries the retired {@code terminal:}
* pin (CB-579), naming {@code tab:} as its replacement.
*
* <p>{@code Leader} is {@code @JsonIgnoreProperties(ignoreUnknown = true)}, so simply dropping
* the record component would make a leftover {@code terminal:} key silently no-op — the daemon
* would start, the pin would never take effect, and nothing would say why. Fatal and specific
* instead, exactly like {@link #rejectRenamedTopLevelKeys}, which this mirrors for a key one
* level deeper than the ones that method covers.
*
* @param yaml the raw config text
* @throws IllegalStateException when any {@code fleet.leaders.<name>.terminal} key is present
*/
static void rejectLeaderTerminalKey(String yaml) {
Map<?, ?> raw;
try {
raw = YAML.readValue(yaml, Map.class);
} catch (IOException | IllegalArgumentException e) {
return; // a malformed file is reported by the real parse, not here
}
if (raw == null || !(raw.get("fleet") instanceof Map<?, ?> fleet)
|| !(fleet.get("leaders") instanceof Map<?, ?> leaders)) {
return;
}
List<String> bad = leaders.entrySet().stream()
.filter(e -> e.getValue() instanceof Map<?, ?> leader && leader.containsKey("terminal"))
.map(e -> String.valueOf(e.getKey()))
.sorted()
.toList();
if (!bad.isEmpty()) {
throw new IllegalStateException("refusing to start: fleet.leaders entries ["
+ String.join(", ", bad) + "] still use the retired 'terminal:' key — replace it "
+ "with 'tab:', the exact tab label hosting the lead. A terminal_id changes on "
+ "every restart of the lead's session; a tab label does not.");
}
}
/**
* Reject a profile whose {@code maxLoad:} is negative, naming both the profile and the key.
*
* <p>Unlike {@code weight} (CB-554), where a negative value degrades to the same "excluded"
* meaning as an explicit {@code 0}, {@code maxLoad} has nowhere lower to degrade to — {@code 0}
* already means "capped at zero live members" (CB-585), the strictest cap there is. Silently
* normalising a negative value to something else is exactly the shape of bug this ticket fixes
* for {@code 0}, so it is refused instead, loud and specific, rather than guessed at.
*
* @param yaml the raw config text
* @throws IllegalStateException when any profile's {@code maxLoad} is negative
*/
static void rejectNegativeMaxLoad(String yaml) {
Map<?, ?> raw;
try {
raw = YAML.readValue(yaml, Map.class);
} catch (IOException | IllegalArgumentException e) {
return; // a malformed file is reported by the real parse, not here
}
if (raw == null || !(raw.get("profiles") instanceof Map<?, ?> profiles)) {
return;
}
List<String> bad = profiles.entrySet().stream()
.filter(e -> e.getValue() instanceof Map<?, ?> p
&& p.get("maxLoad") instanceof Number n && n.doubleValue() < 0)
.map(e -> String.valueOf(e.getKey()))
.sorted()
.toList();
if (!bad.isEmpty()) {
throw new IllegalStateException("refusing to start: profile(s) [" + String.join(", ", bad)
+ "] set a negative maxLoad — maxLoad caps live workers at a non-negative count;"
+ " use 0 to cap a profile at zero live members, or omit the key for unlimited.");
}
}
/** The peer kinds this build has an adapter for — {@link Profile#kind()}'s only valid values. */
private static final Set<String> KNOWN_KINDS = Set.of(Profile.KIND_CLAUDE_CODE, Profile.KIND_OPENCODE);
/**
* Reject a profile whose {@code kind:} is not one of {@link #KNOWN_KINDS} (CB-604), naming the
* profile, the value it set, and the accepted set.
*
* <p>{@link Profile}'s compact constructor only lower-cases {@code kind} and compares it against
* {@code KIND_OPENCODE} — anything else, including a typo like {@code opencod}, silently falls
* into the claude-code bucket ({@link dev.ltms.bridged.member.CompositePeerLauncher} routes by
* exact adapter claim, not by membership in a known set). With {@code argv:} also unset, the argv
* default special-cases only the exact string {@code "claude-code"}, so the launch command falls
* back to {@code List.of(kind)} — the daemon then tries to run a program literally named after the
* typo. {@code CompositePeerLauncher}'s constructor already treats a profile claimed by two
* adapters as fatal (CB-402); an unrecognized kind is the same class of adapter-routing mistake
* and gets the same treatment here, at config load, rather than surfacing later as a failed spawn.
*
* @param yaml the raw config text
* @throws IllegalStateException when any profile's {@code kind} is a non-blank value not in
* {@link #KNOWN_KINDS} (case-insensitive)
*/
static void rejectUnknownKind(String yaml) {
Map<?, ?> raw;
try {
raw = YAML.readValue(yaml, Map.class);
} catch (IOException | IllegalArgumentException e) {
return; // a malformed file is reported by the real parse, not here
}
if (raw == null || !(raw.get("profiles") instanceof Map<?, ?> profiles)) {
return;
}
List<String> bad = profiles.entrySet().stream()
.filter(e -> e.getValue() instanceof Map<?, ?> p
&& p.get("kind") instanceof String k && !k.isBlank()
&& !KNOWN_KINDS.contains(k.toLowerCase()))
.map(e -> String.valueOf(e.getKey()) + "=" + ((Map<?, ?>) e.getValue()).get("kind"))
.sorted()
.toList();
if (!bad.isEmpty()) {
throw new IllegalStateException("refusing to start: profile(s) [" + String.join(", ", bad)
+ "] set an unrecognized kind — accepted values are "
+ String.join(", ", KNOWN_KINDS.stream().sorted().toList())
+ " (case-insensitive); an unrecognized kind would otherwise fall back to the"
+ " claude-code adapter and try to launch a program named after the typo.");
}
}
/** The auth modes this build understands — {@link Auth#mode()}'s only valid values. */
private static final Set<String> KNOWN_AUTH_MODES = Set.of(Auth.MODE_LOOPBACK_TRUST, Auth.MODE_TOKEN);
/**
* Reject an {@code auth.mode} that is not one of {@link #KNOWN_AUTH_MODES} (CB-606), naming the
* value and the accepted set.
*
* <p>{@link Auth}'s compact constructor only lower-cases {@code mode}, and
* {@link Auth#tokenMode()} only compares the result against {@code MODE_TOKEN} — anything else,
* including a typo like {@code toekn}, silently behaves as {@code loopback-trust}. That fallback
* is otherwise checked only by {@link #validateAuthExposure()}, and only when the bind is
* non-loopback: on a loopback bind (the common case) the typo is invisible end to end — the
* daemon starts cleanly and authenticates nobody while the operator believes {@code token} mode
* is active. Refuse it here, unconditionally, at config load, rather than let it hide behind the
* bind check.
*
* @param yaml the raw config text
* @throws IllegalStateException when {@code auth.mode} is a non-blank value not in
* {@link #KNOWN_AUTH_MODES} (case-insensitive)
*/
static void rejectUnknownAuthMode(String yaml) {
Map<?, ?> raw;
try {
raw = YAML.readValue(yaml, Map.class);
} catch (IOException | IllegalArgumentException e) {
return; // a malformed file is reported by the real parse, not here
}
if (raw == null || !(raw.get("auth") instanceof Map<?, ?> auth)) {
return;
}
if (!(auth.get("mode") instanceof String mode) || mode.isBlank()
|| KNOWN_AUTH_MODES.contains(mode.toLowerCase())) {
return;
}
throw new IllegalStateException("refusing to start: auth.mode=" + mode
+ " is not recognized — accepted values are "
+ String.join(", ", KNOWN_AUTH_MODES.stream().sorted().toList())
+ " (case-insensitive); an unrecognized mode would otherwise silently fall back to"
+ " loopback-trust, which authenticates nobody.");
}
/** The per-profile placements this build understands — {@link Profile#placement()}'s only valid values. */
private static final Set<String> KNOWN_PLACEMENTS = Set.of("tab", "pane");
/**
* Reject a profile whose {@code placement:} is not one of {@link #KNOWN_PLACEMENTS} (CB-606),
* naming the profile, the value it set, and the accepted set.
*
* <p>{@link Profile}'s compact constructor only lower-cases {@code placement}, and
* {@link Profile#tabPlacement()} only compares the result against {@code "tab"} — anything else,
* including a typo like {@code tabb}, silently falls back to the legacy pane placement with no
* signal anywhere.
*
* @param yaml the raw config text
* @throws IllegalStateException when any profile's {@code placement} is a non-blank value not in
* {@link #KNOWN_PLACEMENTS} (case-insensitive)
*/
static void rejectUnknownPlacement(String yaml) {
Map<?, ?> raw;
try {
raw = YAML.readValue(yaml, Map.class);
} catch (IOException | IllegalArgumentException e) {
return; // a malformed file is reported by the real parse, not here
}
if (raw == null || !(raw.get("profiles") instanceof Map<?, ?> profiles)) {
return;
}
List<String> bad = profiles.entrySet().stream()
.filter(e -> e.getValue() instanceof Map<?, ?> p
&& p.get("placement") instanceof String pl && !pl.isBlank()
&& !KNOWN_PLACEMENTS.contains(pl.toLowerCase()))
.map(e -> String.valueOf(e.getKey()) + "=" + ((Map<?, ?>) e.getValue()).get("placement"))
.sorted()
.toList();
if (!bad.isEmpty()) {
throw new IllegalStateException("refusing to start: profile(s) [" + String.join(", ", bad)
+ "] set an unrecognized placement — accepted values are "
+ String.join(", ", KNOWN_PLACEMENTS.stream().sorted().toList())
+ " (case-insensitive); an unrecognized placement would otherwise fall back to"
+ " legacy pane placement with no signal anywhere.");
}
}
/** The member-credential policies this build understands — {@link MemberCredentials#policy()}'s only valid value. */
private static final Set<String> KNOWN_MEMBER_CREDENTIALS_POLICIES =
Set.of(MemberCredentials.POLICY_DENY_BY_DEFAULT);
/**
* Reject a {@code memberCredentials.policy} that is not {@link #KNOWN_MEMBER_CREDENTIALS_POLICIES}
* (CB-596), naming the value and the accepted set.
*
* <p>{@link MemberCredentials}'s compact constructor only lower-cases {@code policy} and defaults
* a blank one to {@link MemberCredentials#POLICY_DENY_BY_DEFAULT} — nothing rejects an actual
* typo like {@code deny-by-defualt}. There is only one policy today, so such a typo would
* currently behave identically to the real value by accident; the day a second policy exists
* that accident becomes a silent behavior change. Refuse it now, at config load, following the
* same pattern as {@link #rejectUnknownAuthMode} and {@link #rejectUnknownPlacement}.
*
* @param yaml the raw config text
* @throws IllegalStateException when {@code memberCredentials.policy} is a non-blank value not in
* {@link #KNOWN_MEMBER_CREDENTIALS_POLICIES} (case-insensitive)
*/
static void rejectUnknownMemberCredentialsPolicy(String yaml) {
Map<?, ?> raw;
try {
raw = YAML.readValue(yaml, Map.class);
} catch (IOException | IllegalArgumentException e) {
return; // a malformed file is reported by the real parse, not here
}
if (raw == null || !(raw.get("memberCredentials") instanceof Map<?, ?> mc)) {
return;
}
if (!(mc.get("policy") instanceof String policy) || policy.isBlank()
|| KNOWN_MEMBER_CREDENTIALS_POLICIES.contains(policy.toLowerCase())) {
return;
}
throw new IllegalStateException("refusing to start: memberCredentials.policy=" + policy
+ " is not recognized — accepted values are "
+ String.join(", ", KNOWN_MEMBER_CREDENTIALS_POLICIES.stream().sorted().toList())
+ " (case-insensitive).");
}
/**
* Reject a top-level {@code placement:} policy name {@link PlacementPolicies#fromName} does not
* recognize (CB-606), at config load rather than lazily at first spawn.
*
* <p>{@code CompositePeerLauncher} only calls {@link PlacementPolicies#fromName} per spawn,
* through a {@code Supplier} that re-reads live config (CB-559, so a hot-reloaded placement
* policy takes effect without a restart) — so a bad name still starts a daemon that looks
* healthy and fails only the first time something spawns without naming a profile. Every other
* field this class validates fails here, at load; this one gets the same treatment, calling
* {@link PlacementPolicies#fromName} itself as the single source of truth for what is valid
* rather than duplicating its accepted set.
*
* @param placement the raw, possibly null/blank {@code placement} value as parsed (before
* {@link #withDefaults()} runs); {@code fromName} itself treats null/blank as
* {@code fixed}, so this call changes no default
* @throws IllegalStateException when {@code placement} is a name {@link PlacementPolicies} does
* not recognize
*/
private static void rejectUnknownPlacementPolicy(String placement) {
try {
PlacementPolicies.fromName(placement);
} catch (IllegalArgumentException e) {
throw new IllegalStateException("refusing to start: " + e.getMessage(), e);
}
}
static List<String> unknownTopLevelKeys(String yaml) {
Map<?, ?> raw;
try {
@@ -1081,8 +1590,23 @@ public record BridgedConfig(
// configReload is left as-is: null is "off", and ConfigReload's own compact constructor
// defaults the fields of a block that IS present. Defaulting it here would start watching
// the file for every config that never asked to be watched.
// quarantineCooldownSeconds IS defaulted, unlike leadHeartbeat/configReload above: it has no
// separate on/off switch of its own (BackendQuarantine only ever quarantines a credential
// after a BACKEND_EXHAUSTED classification, which stays opt-in via exhaustedPattern), so a
// config that never mentions it should still get a sane cooldown rather than a null one.
Integer quarantineCooldown = (quarantineCooldownSeconds != null && quarantineCooldownSeconds > 0)
? quarantineCooldownSeconds : DEFAULT_QUARANTINE_COOLDOWN_SECONDS;
// memberCredentials IS defaulted, like guard/lifecycle/auth above, so no reader ever sees a
// null. CB-596: an empty MemberCredentials (empty known, empty allow) blocks NOTHING — unlike
// guard/lifecycle, an absent block is not a safe "feature off" default here, it is a gap. It
// is deliberately not pre-populated with a Java-side name list (that would just reintroduce
// the hardcoded-list defect this record replaces); the block must be configured in
// bridged.yaml to protect anything. See bridged.example.yaml's memberCredentials: comment.
MemberCredentials mc = memberCredentials != null ? memberCredentials
: new MemberCredentials(null, List.of(), List.of());
return new BridgedConfig(b, herdrSocket, profiles, g, worktreeRoot, l, timeout, pollMs,
broker, primary, f, leadHeartbeat, placementOrDefault, a, configReload);
broker, primary, f, leadHeartbeat, health, placementOrDefault, a, configReload,
quarantineCooldown, mc);
}
/**
@@ -1292,10 +1816,9 @@ public record BridgedConfig(
+ "', which is not a configured profiles: entry (have: " + profiles.keySet()
+ ").");
}
if (!leader.isCreatable() && (leader.terminal() == null || leader.terminal().isBlank())) {
bad.add("fleet.leaders." + name + " can neither be found nor created — it pins no "
+ "terminal: and names no profile: to launch one on. Give it one or the "
+ "other, or drop the entry.");
if (leader.tab() == null || leader.tab().isBlank()) {
bad.add("fleet.leaders." + name + " has no tab: — a lead is now found (and, if "
+ "auto-launched, labelled) purely by its tab, so every entry must name one.");
}
});
if (!bad.isEmpty()) {
@@ -30,13 +30,22 @@ import java.util.function.Supplier;
* {@code fleet:} (every role pool, {@code charters}, and {@code tabLabel}),
* {@code placement:}, and an existing profile's {@code weight} / {@code maxLoad}. Those
* three are read through a supplier on {@code CompositePeerLauncher}, which is what makes
* them hot — not the fact that they are config.</li>
* them hot — not the fact that they are config. <strong>This does NOT include
* {@code fleet.leaders}</strong>: {@code Bridged.main} reads {@code cfg.fleet().leaders()}
* once at startup to build the {@code LeadTabScanner} and the {@code LeadLauncher}, and
* neither is reconstructed on reload — so a lead added, removed, or re-{@code tab}'d under
* {@code fleet.leaders} needs a restart, the same as any deferred key below.</li>
* <li><strong>Deferred</strong> — accepted into the new snapshot, but the wiring built at startup
* keeps the old value until a restart: {@code lifecycle:}, {@code leadHeartbeat:},
* {@code spawnReadyTimeoutMs} / {@code spawnReadyPollMs}, {@code guard:},
* {@code worktreeRoot:}, adding or removing a profile (a new backend needs its own launcher,
* {@code spawnReadyTimeoutMs} / {@code spawnReadyPollMs}, {@code quarantineCooldownSeconds}
* (CB-578 stage B — baked once into the {@code BackendQuarantine} built at startup),
* {@code guard:}, {@code worktreeRoot:}, adding or removing a profile (a new backend needs its own launcher,
* which is constructed once), <em>and an existing profile's launch settings</em> —
* {@code model}, {@code baseUrl}, {@code argv}, {@code env}, {@code mcpUrl} and the rest.
* {@code model}, {@code baseUrl}, {@code argv}, {@code env}, {@code mcpUrl},
* {@code exhaustedPattern} (CB-578 stage A — compiled once into {@code Bridged.main}'s
* pattern map at startup), and the rest. {@code credentialId} (CB-578 stage B) is NOT on
* this list — it is read live off the config supplier at every quarantine check and
* exhaustion event, exactly like {@code weight} / {@code maxLoad}, so it is hot instead.
* {@code HerdrPeerLauncher} takes {@code Map.copyOf(profiles)} at construction and resolves
* each spawn out of that copy, so those never reach a launch until the daemon restarts. A
* reload logs these rather than pretending they applied.</li>
@@ -211,6 +220,12 @@ public final class ConfigRef implements Supplier<BridgedConfig> {
|| !Objects.equals(old.spawnReadyPollMs(), fresh.spawnReadyPollMs())) {
changed.add("spawnReady*");
}
// CB-578 stage B: baked once into the BackendQuarantine built at startup — a running
// quarantine keeps its original cooldown regardless, and a new cooldown only applies to a
// quarantine that starts after a restart.
if (!Objects.equals(old.quarantineCooldownSeconds(), fresh.quarantineCooldownSeconds())) {
changed.add("quarantineCooldownSeconds");
}
Map<String, BridgedConfig.Profile> before =
old.profiles() == null ? Map.of() : old.profiles();
Map<String, BridgedConfig.Profile> after =
@@ -246,8 +261,10 @@ public final class ConfigRef implements Supplier<BridgedConfig> {
/**
* Whether two versions of a profile would launch a peer identically. Compares every component
* the launcher reads at spawn; {@code weight} and {@code maxLoad} are excluded because those are
* read live by the placement policy and really do take effect on the next spawn.
* the launcher reads at spawn; {@code weight}, {@code maxLoad} and {@code credentialId} are
* excluded because those are read live (by the placement policy and, for credentialId, by
* {@code CompositePeerLauncher}/the CB-578 stage B exhaustion sink) and really do take effect on
* the next spawn.
*/
private static boolean sameLaunchSettings(BridgedConfig.Profile a, BridgedConfig.Profile b) {
return Objects.equals(a.baseUrl(), b.baseUrl())
@@ -265,6 +282,10 @@ public final class ConfigRef implements Supplier<BridgedConfig> {
&& Objects.equals(a.gitHostEnv(), b.gitHostEnv())
&& Objects.equals(a.kind(), b.kind())
&& Objects.equals(a.env(), b.env())
&& Objects.equals(a.subscription(), b.subscription());
&& Objects.equals(a.subscription(), b.subscription())
// CB-578 stage B: exhaustedPattern is compiled once into Bridged.main's pattern map
// at startup (see ExhaustedPatternLookup wiring) — a reload never re-reads it, so a
// changed pattern must be reported as deferred, exactly like model/baseUrl/argv.
&& Objects.equals(a.exhaustedPattern(), b.exhaustedPattern());
}
}
@@ -0,0 +1,151 @@
package dev.ltms.bridged.health;
import dev.ltms.bridged.herdr.Agent;
import dev.ltms.bridged.herdr.AgentControl;
import dev.ltms.bridged.herdr.AgentStatus;
import dev.ltms.bridged.msg.MessageService;
import dev.ltms.bridged.session.MemberSession;
import org.slf4j.Logger;
import org.slf4j.LoggerFactory;
import java.util.HashMap;
import java.util.HashSet;
import java.util.List;
import java.util.Map;
import java.util.Objects;
import java.util.concurrent.ScheduledExecutorService;
import java.util.concurrent.TimeUnit;
import java.util.function.BiConsumer;
import java.util.function.LongSupplier;
import java.util.function.Supplier;
/** Slow whole-fleet evidence collection. It is deliberately separate from the delivery poller. */
public final class FleetHealthMonitor {
private static final Logger log = LoggerFactory.getLogger(FleetHealthMonitor.class);
/** Bounded attempts to run {@link #failTarget} for one transition. Never retried tick-to-tick (CB-580). */
static final int MAX_FAIL_TARGET_ATTEMPTS = 3;
private final AgentControl agents;
private final Supplier<List<MemberSession>> roster;
private final MessageService messages;
private final ScheduledExecutorService scheduler;
private final LongSupplier clock;
private final long intervalSeconds;
private final BiConsumer<String, String> failTarget;
private final Map<String, HealthPrior> priors = new HashMap<>();
private final Map<String, HealthState> states = new HashMap<>();
// These facts need the evidence publishers introduced by later M4 units. They are not negatives.
private static final boolean NOT_YET_OBSERVED = false;
/**
* @param failTarget CB-568's idempotent target-wide failure operation (e.g. {@code messages::abandon}),
* invoked once when a member transitions into a terminal health state. Required —
* there is deliberately no defaulting overload; a caller that does not want the
* fail-tickets-on-terminal-health behavior must pass an explicit inert value (see
* {@code TestTurnTokens.inert} / {@code BridgeMcp.CapacitySource.none()} for the pattern).
*/
public FleetHealthMonitor(AgentControl agents, Supplier<List<MemberSession>> roster, MessageService messages,
ScheduledExecutorService scheduler, LongSupplier clock, long intervalSeconds,
BiConsumer<String, String> failTarget) {
this.agents = agents;
this.roster = roster;
this.messages = messages;
this.scheduler = scheduler;
this.clock = clock;
this.intervalSeconds = intervalSeconds;
this.failTarget = Objects.requireNonNull(failTarget, "failTarget");
}
/** Pure per-member decision seam. */
static HealthDecision decide(HealthSnapshot snapshot, HealthPrior prior, long nowNanos) {
return FleetHealth.decide(snapshot, prior, nowNanos);
}
public void start() { scheduler.schedule(this::tick, intervalSeconds, TimeUnit.SECONDS); }
public void stop() { scheduler.shutdownNow(); }
// Package-private so tests can run one tick without waiting.
void tick() {
try {
List<Agent> agentsNow = agents.list(); // Exactly one list call for this complete observation.
List<MemberSession> rosterNow = roster.get(); // One in-memory roster snapshot for this tick.
Map<String, Agent> live = new HashMap<>();
for (Agent agent : agentsNow) live.put(agent.terminalId(), agent);
HashSet<String> current = new HashSet<>();
for (MemberSession session : rosterNow) {
current.add(session.terminalId());
Agent agent = live.get(session.terminalId());
AgentStatus status = agent == null ? AgentStatus.UNKNOWN : agent.status();
boolean accepted = messages.hasAcceptedDelivery(session.terminalId());
HealthSnapshot snapshot = new HealthSnapshot(session.state(), status, accepted, NOT_YET_OBSERVED,
messages.hasInboxMessage(session.terminalId()), agent != null, NOT_YET_OBSERVED,
NOT_YET_OBSERVED, NOT_YET_OBSERVED, NOT_YET_OBSERVED, NOT_YET_OBSERVED, NOT_YET_OBSERVED);
HealthDecision decision = decide(snapshot, priors.getOrDefault(session.terminalId(), HealthPrior.NONE),
clock.getAsLong());
priors.put(session.terminalId(), decision.prior());
reportTransition(session.terminalId(), decision.state());
}
priors.keySet().retainAll(current);
states.keySet().retainAll(current);
} catch (Throwable error) {
// A list failure is health evidence, and must never kill the monitor's only scheduler task.
log.warn("fleet health collection failed; will retry next tick", error);
} finally {
if (!scheduler.isShutdown()) {
scheduler.schedule(this::tick, intervalSeconds, TimeUnit.SECONDS);
}
}
}
void reportTransition(String target, HealthState next) {
HealthState previous = states.put(target, next);
if (previous == next) return;
if (fault(next)) {
log.warn("fleet health member={} state={} previous={}", target, next, previous);
} else if (previous != null && fault(previous)) {
log.info("fleet health member={} recovered state={} previous={}", target, next, previous);
}
// CB-580: a member entering GONE/NEVER_READY must not leave its waiting tickets pending
// forever. Fire exactly once per transition — never on a tick where the state is unchanged,
// which is what made the rejected commit call abandon() once per tick for as long as a
// member stayed terminal.
if (terminal(next)) {
failTerminalTarget(target, next);
}
}
private void failTerminalTarget(String target, HealthState state) {
String reason = "fleet health: member reached terminal state " + state.name();
RuntimeException last = null;
for (int attempt = 1; attempt <= MAX_FAIL_TARGET_ATTEMPTS; attempt++) {
try {
failTarget.accept(target, reason);
return;
} catch (RuntimeException error) {
last = error;
log.warn("fleet health: failTarget attempt {}/{} failed for member={} state={}",
attempt, MAX_FAIL_TARGET_ATTEMPTS, target, state, error);
}
}
log.warn("fleet health: giving up on failTarget for member={} state={} after {} attempts",
target, state, MAX_FAIL_TARGET_ATTEMPTS, last);
}
private static boolean terminal(HealthState state) {
return state == HealthState.GONE || state == HealthState.NEVER_READY;
}
private static boolean fault(HealthState state) {
return switch (state) {
case NEVER_READY, GONE, TURN_BOUNDARY_LOST, ERROR_ON_SCREEN, STALL_SUSPECTED,
MUTE, REPLY_STRANDED, DELEGATION_ORPHANED, CONTROL_LINK_DOWN -> true;
default -> false;
};
}
public static String coverage(boolean enabled, boolean notificationConfigured) {
return !enabled ? "off" : notificationConfigured ? "full" : "detection-only";
}
}
@@ -6,6 +6,7 @@ import org.slf4j.LoggerFactory;
import java.util.Collections;
import java.util.LinkedHashMap;
import java.util.Locale;
import java.util.Map;
import java.util.Set;
import java.util.function.LongSupplier;
@@ -23,6 +24,15 @@ import java.util.function.Supplier;
* by first starting the session and asking it. Scanning closes that loop: label the tab, and the
* pane is recognised on the next resolve.
*
* <p><strong>CB-579 — matched by name, not prefix.</strong> This used to strip one shared
* {@code tabPrefix} off a label to derive the lead's name, and merged a config-supplied
* {@code terminal_id} pin over every scan result so the pin could never expire. Both are gone: each
* lead now configures its own exact {@code tab} label ({@code fleet.leaders.<name>.tab}), so this
* class is handed a {@code tab → name} map up front and matches labels against it exactly
* (case-insensitively). There is no merge step — a scan result is the whole answer. That is the
* fix for the bug this replaces: a {@code terminal_id} pin surviving in config after the pane it
* named was gone, so the daemon kept treating a dead session as a live lead forever.
*
* <p><strong>Direction of trust.</strong> The label names the lead; it never <em>grants</em>
* anything a pane could take for itself. Three properties keep that honest:
* <ol>
@@ -48,7 +58,7 @@ import java.util.function.Supplier;
* ever make a decision that <em>removes</em> something based on this map, add the same check.
* The remaining hazard is an <em>operator</em> one — a worker {@code tabLabel} template that
* happens to start with the same prefix would promote the whole fleet — and that is refused at
* startup by {@code BridgedConfig.validateLeadScan} rather than documented here.
* startup by {@code BridgedConfig.validateLeadTabPrefixes} rather than documented here.
*
* <p><strong>Caching.</strong> {@link #get()} is on the request path (every resolve), so the scan
* is TTL-cached and a stale-but-valid map is preferred to a herdr round-trip. A failed scan keeps
@@ -60,38 +70,46 @@ public final class LeadTabScanner implements Supplier<Map<String, String>> {
private static final Logger log = LoggerFactory.getLogger(LeadTabScanner.class);
private final HerdrClient herdr;
private final String tabPrefix;
private final Map<String, String> tabToName;
private final Set<String> excludedWorkspaceLabels;
private final Map<String, String> configuredLeads;
private final long ttlNanos;
private final LongSupplier clock;
private Map<String, String> cached;
private Map<String, String> cached = Map.of();
private long scannedAtNanos;
private boolean everScanned;
/**
* @param herdr the herdr client to query ({@code workspace.list},
* {@code tab.list}, {@code pane.list} — all read-only)
* @param tabPrefix a tab whose label starts with this (case-insensitively) hosts a
* lead; the rest of the label, trimmed, is the lead's name
* @param tabToName every configured lead's exact tab label → its name
* ({@code fleet.leaders.<name>.tab}), matched case-insensitively
* @param excludedWorkspaceLabels workspaces never scanned — the configured worker spaces
* @param configuredLeads the static {@code leaders:}/{@code primary:} registry, merged
* over every scan result. Explicit config outranks the
* convention, and survives a scan that cannot run at all
* @param ttlNanos how long a scan result is reused before the next one
* @param clock nanosecond time source ({@code System::nanoTime} in production)
*/
public LeadTabScanner(HerdrClient herdr, String tabPrefix, Set<String> excludedWorkspaceLabels,
Map<String, String> configuredLeads, long ttlNanos, LongSupplier clock) {
public LeadTabScanner(HerdrClient herdr, Map<String, String> tabToName,
Set<String> excludedWorkspaceLabels, long ttlNanos, LongSupplier clock) {
this.herdr = herdr;
this.tabPrefix = tabPrefix == null || tabPrefix.isBlank() ? "lead:" : tabPrefix.strip();
this.tabToName = normalize(tabToName);
this.excludedWorkspaceLabels = excludedWorkspaceLabels == null
? Set.of() : Set.copyOf(excludedWorkspaceLabels);
this.configuredLeads = configuredLeads == null ? Map.of() : Map.copyOf(configuredLeads);
this.ttlNanos = ttlNanos;
this.clock = clock;
this.cached = this.configuredLeads;
}
/** Keys stripped and lower-cased once, so every lookup is a plain map hit. */
private static Map<String, String> normalize(Map<String, String> tabToName) {
if (tabToName == null || tabToName.isEmpty()) {
return Map.of();
}
Map<String, String> out = new LinkedHashMap<>();
tabToName.forEach((tab, name) -> {
if (tab != null && !tab.isBlank() && name != null && !name.isBlank()) {
out.put(tab.strip().toLowerCase(Locale.ROOT), name);
}
});
return Collections.unmodifiableMap(out);
}
/**
@@ -151,25 +169,20 @@ public final class LeadTabScanner implements Supplier<Map<String, String>> {
}
}
}
byTerminal.putAll(configuredLeads); // an explicit pin outranks a label
return Collections.unmodifiableMap(byTerminal);
}
/**
* The lead name a tab label declares, or {@code null} if it declares none.
* The lead name a tab label declares, or {@code null} if it names none of the configured leads.
*
* <p>{@code "lead: opus-5.0"} → {@code "opus-5.0"}. A bare {@code "lead:"} names nobody and is
* rejected: an unnamed lead would resolve as {@code PRIMARY} with nothing to attribute it to.
* <p>Exact match (case-insensitive, ends stripped) against {@link #tabToName} — no prefix
* stripping, so an operator's {@code "lead: something-else"} tab is never mistaken for a
* configured lead just because it shares a prefix.
*/
private String leadNameOf(String label) {
if (label == null) {
return null;
}
String l = label.strip();
if (!l.regionMatches(true, 0, tabPrefix, 0, tabPrefix.length())) {
return null;
}
String name = l.substring(tabPrefix.length()).strip();
return name.isEmpty() ? null : name;
return tabToName.get(label.strip().toLowerCase(Locale.ROOT));
}
}
@@ -2,11 +2,17 @@ package dev.ltms.bridged.inject;
import dev.ltms.bridged.herdr.AgentControl;
import dev.ltms.bridged.msg.Rendezvous;
import dev.ltms.bridged.msg.TurnToken;
import org.slf4j.Logger;
import org.slf4j.LoggerFactory;
import java.util.List;
import java.util.Objects;
import java.util.Set;
import java.util.TreeSet;
import java.util.concurrent.CompletableFuture;
import java.util.concurrent.ConcurrentHashMap;
import java.util.regex.Pattern;
/**
* The CB-106 completion fallback: bridges the {@link Injector}'s turn-completion signal to the
@@ -60,6 +66,8 @@ public final class CompletionResolver implements TurnListener {
private final AgentControl agents;
private final Rendezvous rendezvous;
private final ExhaustedPatternLookup exhaustedPatterns;
private final ExhaustionSink exhaustionSink;
/**
* Per-target record of the turn currently in flight: the exact {@link Rendezvous} waiter its
@@ -77,23 +85,36 @@ public final class CompletionResolver implements TurnListener {
private final ConcurrentHashMap<String, InFlight> inFlight = new ConcurrentHashMap<>();
public CompletionResolver(AgentControl agents, Rendezvous rendezvous) {
/**
* @param exhaustedPatterns CB-578 stage A: per-target lookup for a profile's configured
* usage-limit refusal pattern. Required — there is deliberately no
* defaulting overload; a caller that does not want the classification
* must pass an explicit inert value ({@link ExhaustedPatternLookup#none()}).
* @param exhaustionSink CB-578 stage B: notified when a {@code BACKEND_EXHAUSTED}
* classification actually resolves a waiter. Required for the same
* reason as {@code exhaustedPatterns} — pass {@link ExhaustionSink#none()}
* to opt out.
*/
public CompletionResolver(AgentControl agents, Rendezvous rendezvous, ExhaustedPatternLookup exhaustedPatterns,
ExhaustionSink exhaustionSink) {
this.agents = agents;
this.rendezvous = rendezvous;
this.exhaustedPatterns = Objects.requireNonNull(exhaustedPatterns, "exhaustedPatterns");
this.exhaustionSink = Objects.requireNonNull(exhaustionSink, "exhaustionSink");
}
@Override
public void onDelivered(String target) {
public void onDelivered(String target, TurnToken token) {
// Capture the exact waiter this turn belongs to (CB-116) and snapshot the pane's pre-turn
// content — what it shows *before* the just-delivered turn produces output — as the staleness
// reference (CB-115). Done synchronously (like the delivering send itself) so both are in
// place before this turn's completion can fire.
captureBaseline(target);
captureBaseline(target, token);
}
/** Capture the in-flight turn: its waiter and pre-turn baseline (the testable core of {@link #onDelivered}). */
void captureBaseline(String target) {
CompletableFuture<Rendezvous.Resolution> waiter = rendezvous.currentWaiter(target);
void captureBaseline(String target, TurnToken token) {
CompletableFuture<Rendezvous.Resolution> waiter = token.waiter();
if (waiter == null) {
inFlight.remove(target); // no send is waiting on this delivery — nothing to resolve later
return;
@@ -137,7 +158,13 @@ public final class CompletionResolver implements TurnListener {
@Override
public void onTurnFailed(String target) {
InFlight turn = inFlight.get(target);
Thread.ofVirtual().name("turn-failed-" + target).start(() -> fail(target, turn));
Thread.ofVirtual().name("turn-failed-" + target).start(() -> fail(target, turn, null));
}
@Override
public void onTurnFailed(String target, String reason) {
InFlight turn = inFlight.get(target);
Thread.ofVirtual().name("turn-failed-" + target).start(() -> fail(target, turn, reason));
}
/** Synchronous resolve (the unit-testable core of {@link #onTurnComplete}). */
@@ -150,11 +177,12 @@ public final class CompletionResolver implements TurnListener {
return;
}
String tail;
String assistantBlock = null;
int originalLength = 0;
boolean clipped = false;
boolean scrapeFailed = false;
try {
String assistantBlock = lastAssistantBlock(agents.read(target, SCRAPE_SOURCE));
assistantBlock = lastAssistantBlock(agents.read(target, SCRAPE_SOURCE));
originalLength = assistantBlock.strip().length();
clipped = originalLength > MAX_SCRAPE_CHARS;
tail = clip(assistantBlock);
@@ -177,6 +205,25 @@ public final class CompletionResolver implements TurnListener {
target);
return; // keep the in-flight record: a later genuine completion still needs it
}
// CB-578 stage A: a turn that ended with no bridge_reply AND whose scrape matches the
// backend's configured usage-limit pattern is a refusal, not an answer. Classify it as
// BACKEND_EXHAUSTED rather than handing the caller a scrape that reads like a real reply.
if (!scrapeFailed) {
Pattern exhausted = exhaustedPatterns.patternFor(target);
String matchedLine = exhausted == null ? null : firstMatchingLine(assistantBlock, exhausted);
if (matchedLine != null) {
String reason = "backend exhausted (usage limit): " + matchedLine;
if (rendezvous.resolveExhausted(waiter, reason)) {
inFlight.remove(target, turn);
log.warn("completion for {} classified BACKEND_EXHAUSTED (no bridge_reply; scrape "
+ "matched the profile's exhausted pattern): {}", target, reason);
// CB-578 stage B: only on the resolution that actually won the race — a late
// duplicate must never quarantine a credential twice for one refusal.
exhaustionSink.onExhausted(target, reason);
}
return;
}
}
String completion = clipped ? tail + "\n" + CLIPPED_PANE_TAIL_MARKER : tail;
if (rendezvous.resolveCompletion(waiter, completion)) {
inFlight.remove(target, turn);
@@ -192,6 +239,11 @@ public final class CompletionResolver implements TurnListener {
/** Synchronous fail (the unit-testable core of {@link #onTurnFailed}). */
void fail(String target, InFlight turn) {
fail(target, turn, null);
}
/** Synchronous fail with an optional reason supplied by a dropped worker queue. */
void fail(String target, InFlight turn, String explicitReason) {
// A never-delivered readiness failure has no in-flight record but still has a blocked send;
// fall back to the currently-registered waiter (unambiguous — that send never completed, so
// no next turn exists to confuse it with).
@@ -201,16 +253,18 @@ public final class CompletionResolver implements TurnListener {
inFlight.remove(target, turn); // nobody blocked on this worker — nothing to fail
return;
}
String reason;
try {
reason = clip(agents.read(target, SCRAPE_SOURCE));
} catch (RuntimeException e) {
reason = "";
}
if (reason.isBlank()) {
// No screen to scrape — either the worker is stuck (CB-109) or gone (CB-110).
reason = "worker did not reply; its turn ended in an unrecoverable state "
+ "(worker unreachable or stuck)";
String reason = explicitReason;
if (reason == null || reason.isBlank()) {
try {
reason = clip(agents.read(target, SCRAPE_SOURCE));
} catch (RuntimeException e) {
reason = "";
}
if (reason.isBlank()) {
// No screen to scrape — either the worker is stuck (CB-109) or gone (CB-110).
reason = "worker did not reply; its turn ended in an unrecoverable state "
+ "(worker unreachable or stuck)";
}
}
if (rendezvous.resolveFailure(waiter, reason)) {
inFlight.remove(target, turn);
@@ -218,6 +272,45 @@ public final class CompletionResolver implements TurnListener {
}
}
/**
* The first line of {@code text} matching {@code pattern}, stripped — the CB-578 stage A
* evidence carried in a {@code BACKEND_EXHAUSTED} reason so the operator sees the real refusal
* text, never a generic label. {@code null} if no line matches.
*/
static String firstMatchingLine(String text, Pattern pattern) {
if (text == null || text.isEmpty()) return null;
for (String line : text.split("\n", -1)) {
if (pattern.matcher(line).find()) {
return line.strip();
}
}
return null;
}
/**
* Coverage summary for the CB-578 stage A exhausted-pattern classification, logged at startup
* the way {@link dev.ltms.bridged.health.FleetHealthMonitor#coverage} is — so an operator can
* see whether the classification is on, and for which profiles, without reading every
* profile's config by hand.
*
* @param allProfiles every configured profile name
* @param configuredProfiles the subset of {@code allProfiles} that carry an exhausted pattern
*/
public static String coverage(Set<String> allProfiles, Set<String> configuredProfiles) {
if (configuredProfiles.isEmpty()) {
return "off (no profile has an exhaustedPattern configured; profiles: " + sorted(allProfiles) + ")";
}
Set<String> unconfigured = new TreeSet<>(allProfiles);
unconfigured.removeAll(configuredProfiles);
return unconfigured.isEmpty()
? "full (all profiles configured: " + sorted(allProfiles) + ")"
: "partial (configured: " + sorted(configuredProfiles) + "; not configured: " + sorted(unconfigured) + ")";
}
private static List<String> sorted(Set<String> names) {
return names.stream().sorted().toList();
}
private static String clip(String s) {
if (s == null) return "";
String trimmed = s.strip();
@@ -0,0 +1,27 @@
package dev.ltms.bridged.inject;
import java.util.regex.Pattern;
/**
* Per-target lookup for a profile's configured usage-limit refusal pattern (CB-578 stage A): how
* {@link CompletionResolver} tells a backend that refused on a subscription usage limit — the
* worker's pane stays healthy, but the account is exhausted — apart from a genuine completion.
*
* <p>The pattern is always profile config, never a vendor string in Java source: every backend
* words its refusal differently, so a hardcoded sentence would only ever match one of them.
*/
@FunctionalInterface
public interface ExhaustedPatternLookup {
/** The compiled pattern configured for {@code target}'s profile, or {@code null} if none. */
Pattern patternFor(String target);
/**
* Inert lookup — no profile has a pattern configured, so the classification never fires and
* the completion fallback behaves exactly as before CB-578 stage A. The explicit stand-in a
* caller (or a test not exercising this feature) passes instead of a defaulting overload.
*/
static ExhaustedPatternLookup none() {
return target -> null;
}
}
@@ -0,0 +1,31 @@
package dev.ltms.bridged.inject;
/**
* Notified when {@link CompletionResolver} actually delivers a {@code BACKEND_EXHAUSTED}
* classification to a waiting send (CB-578 stage B) — never on a race that lost (see
* {@link CompletionResolver#resolve}, which only calls this after
* {@code Rendezvous.resolveExhausted} returns {@code true}).
*
* <p>{@link CompletionResolver} knows only {@code target} (a herdr terminal id); it has no notion of
* profiles or credentials, so mapping {@code target} to whatever should be quarantined is entirely
* the sink's job — see {@code Bridged.main}'s wiring, which resolves target → session → profile →
* {@code effectiveCredentialId()} and calls {@code BackendQuarantine.quarantine} on it.
*/
@FunctionalInterface
public interface ExhaustionSink {
/**
* @param target the herdr terminal id whose turn was classified {@code BACKEND_EXHAUSTED}
* @param reason the matched-line reason carried by the classification
*/
void onExhausted(String target, String reason);
/**
* Inert sink — nothing happens on exhaustion. The explicit stand-in a caller (or a test not
* exercising this feature) passes instead of a defaulting overload, exactly like
* {@link ExhaustedPatternLookup#none()}.
*/
static ExhaustionSink none() {
return (target, reason) -> { };
}
}
@@ -2,6 +2,7 @@ package dev.ltms.bridged.inject;
import dev.ltms.bridged.herdr.AgentControl;
import dev.ltms.bridged.herdr.AgentStatus;
import dev.ltms.bridged.msg.TurnToken;
import org.slf4j.Logger;
import org.slf4j.LoggerFactory;
@@ -129,7 +130,7 @@ public final class Injector {
}
/** A pending message and the future that completes when it has been delivered. */
private record Pending(String text, CompletableFuture<Void> delivered) {
private record Pending(String text, TurnToken token, CompletableFuture<Void> delivered) {
}
/** Per-worker delivery state, guarded by its own monitor (single writer per worker). */
@@ -159,9 +160,9 @@ public final class Injector {
* <p>Uses an atomic map update so a concurrent {@link #drop} cannot slip between "find the
* target" and "queue the message" and orphan it in a target it just removed.
*/
public CompletableFuture<Void> enqueue(String target, String text) {
public CompletableFuture<Void> enqueue(String target, String text, TurnToken token) {
CompletableFuture<Void> delivered = new CompletableFuture<>();
Pending p = new Pending(text, delivered);
Pending p = new Pending(text, token, delivered);
targets.compute(target, (_, existing) -> {
Target t = (existing != null) ? existing : new Target();
t.add(p); // synchronized on the Target monitor — atomic with a concurrent drop
@@ -356,7 +357,7 @@ public final class Injector {
} else {
// Baseline the pane's pre-turn content so a misattributed completion (no new output)
// can't resolve this send with the previous turn's stale answer (CB-115).
turnListener.onDelivered(target);
turnListener.onDelivered(target, sent.token());
sent.delivered().complete(null);
}
}
@@ -410,8 +411,8 @@ public final class Injector {
for (Pending p : pending) {
p.delivered().completeExceptionally(cause);
}
if (hadDeliveredTurn) {
turnListener.onTurnFailed(target);
}
// A queued send has no in-flight record, while a delivered turn does. CompletionResolver
// handles both forms and resolves its waiter at most once.
turnListener.onTurnFailed(target, cause.getMessage());
}
}
@@ -1,5 +1,7 @@
package dev.ltms.bridged.inject;
import dev.ltms.bridged.msg.TurnToken;
/**
* Notified when a worker's delegated turn is observed to complete — a confirmed
* {@code WORKING → IDLE} transition after a delivery. This is the CB-106 completion signal the
@@ -41,6 +43,14 @@ public interface TurnListener {
default void onTurnFailed(String target) {
}
/**
* As {@link #onTurnFailed(String)}, carrying the reason a worker became unreachable. The default
* keeps existing listeners working while allowing the completion resolver to report a useful cause.
*/
default void onTurnFailed(String target, String reason) {
onTurnFailed(target);
}
/**
* A message was just delivered into {@code target}'s pane (CB-115). Fired so the completion
* resolver can snapshot the pane's pre-turn content: a later {@link #onTurnComplete} whose
@@ -49,7 +59,7 @@ public interface TurnListener {
* resolve the send with the previous turn's stale answer. A default no-op keeps the interface
* functional for callers that don't scrape.
*/
default void onDelivered(String target) {
default void onDelivered(String target, TurnToken token) {
}
/** No-op default for callers that only need delivery, not completion signalling. */
@@ -59,7 +59,7 @@ public final class LeadLauncher {
/**
* @param agents herdr agent control (start, list)
* @param spaces workspace / tab control (ensure, create, label, list)
* @param cfg the loaded config — {@code fleet.leaders}, {@code profiles} and the lead pins
* @param cfg the loaded config — {@code fleet.leaders}, {@code profiles} and each lead's tab
*/
public LeadLauncher(AgentControl agents, WorkspaceControl spaces, BridgedConfig cfg) {
this.agents = agents;
@@ -102,8 +102,8 @@ public final class LeadLauncher {
continue;
}
if (!lead.isCreatable()) {
// A lead with a `terminal:` pin and no `profile:` is recognise-only by design: the
// operator opens it by hand. Say so once rather than looking like a silent failure.
// A lead with a `tab:` but no `profile:` is recognise-only by design: the operator
// opens it by hand. Say so once rather than looking like a silent failure.
log.info("lead '{}' is not live, and names no profile — it can be recognised but not "
+ "launched. Add `profile:` under fleet.leaders.{} to have bridged start it.",
name, name);
@@ -127,18 +127,16 @@ public final class LeadLauncher {
}
/**
* How many live leads exist per configured name.
* How many live leads exist per configured name: a running agent in a tab labelled with that
* lead's exact {@code tab} (CB-579). Member workspaces are excluded, exactly as the scanner
* excludes them: a member must not be counted as a lead because it happens to sit in a matching
* tab.
*
* <p>Two independent pieces of evidence, because either alone double-spawns:
* <ul>
* <li>a running agent in a tab labelled {@code "<tabPrefix> <name>"} — how an auto-launched
* lead, or an operator following the labelling convention, is found;</li>
* <li>a running agent on a terminal the config pins in {@code fleet.leaders.<name>.terminal} —
* how a lead the operator opened and pinned by hand is found. Without this, a pinned lead
* whose tab carries no matching label would be relaunched on every boot.</li>
* </ul>
* Member workspaces are excluded, exactly as the scanner excludes them: a member must not be
* counted as a lead because it happens to sit in a matching tab.
* <p>There used to be a second path here — a running agent on the terminal a
* {@code fleet.leaders.<name>.terminal} pin named, for a lead opened and pinned by hand. That
* pin is retired: {@code tab} is now the only field identity depends on, and {@link Agent}
* already carries {@link Agent#tabId()} directly, so a hand-opened lead is found the same way an
* auto-launched one is — by labelling its tab to match.
*/
private Map<String, Integer> liveLeads(Map<String, BridgedConfig.Leader> leaders) {
Set<String> memberSpaces = cfg.profiles().values().stream()
@@ -160,20 +158,9 @@ public final class LeadLauncher {
}
}
// terminalId → the lead name the config pins it to.
Map<String, String> nameByPinnedTerminal = new LinkedHashMap<>();
leaders.forEach((name, lead) -> {
if (lead.terminal() != null && !lead.terminal().isBlank()) {
nameByPinnedTerminal.put(lead.terminal().strip(), name);
}
});
Map<String, Integer> counts = new LinkedHashMap<>();
for (Agent a : agents.list()) {
String name = nameByTab.get(a.tabId());
if (name == null) {
name = nameByPinnedTerminal.get(a.terminalId());
}
if (name != null) {
counts.merge(name, 1, Integer::sum);
}
@@ -184,7 +171,7 @@ public final class LeadLauncher {
/**
* The configured lead a tab label names, or {@code null} for a label that names none.
*
* <p>Matched against the declared lead names rather than by splitting on the prefix, so an
* <p>Matched exactly (case-insensitively) against each lead's configured {@code tab}, so an
* operator's {@code "lead: something-else"} tab is not mistaken for a configured lead.
*/
private String leadNameOf(String label, Map<String, BridgedConfig.Leader> leaders) {
@@ -193,7 +180,8 @@ public final class LeadLauncher {
}
String l = label.strip();
for (Map.Entry<String, BridgedConfig.Leader> e : leaders.entrySet()) {
if (l.equalsIgnoreCase(e.getValue().tabLabel(e.getKey()).strip())) {
String tab = e.getValue().tabLabel();
if (tab != null && l.equalsIgnoreCase(tab.strip())) {
return e.getKey();
}
}
@@ -202,7 +190,7 @@ public final class LeadLauncher {
/** Start one lead. Returns false (having logged) rather than throwing on any failure. */
private boolean launch(String name, BridgedConfig.Leader lead, BridgedConfig.Profile profile) {
String label = lead.tabLabel(name);
String label = lead.tabLabel();
String cwd = (lead.cwd() == null || lead.cwd().isBlank())
? System.getProperty("user.dir") : lead.cwd();
@@ -0,0 +1,30 @@
package dev.ltms.bridged.logging;
import ch.qos.logback.classic.Level;
import ch.qos.logback.classic.Logger;
import ch.qos.logback.classic.turbo.TurboFilter;
import ch.qos.logback.core.spi.FilterReply;
import io.modelcontextprotocol.spec.McpSchema;
import org.slf4j.Marker;
/** Suppresses only the SDK warning for the normal MCP cancellation notification. */
public final class McpCancelledNotificationFilter extends TurboFilter {
static final String LOGGER = "io.modelcontextprotocol.spec.McpStreamableServerSession";
static final String UNHANDLED_NOTIFICATION = "No handler registered for notification method: {}";
@Override
public FilterReply decide(Marker marker, Logger logger, Level level, String format, Object[] params,
Throwable throwable) {
if (level == Level.WARN
&& LOGGER.equals(logger.getName())
&& UNHANDLED_NOTIFICATION.equals(format)
&& params != null
&& params.length == 1
&& params[0] instanceof McpSchema.JSONRPCNotification notification
&& "notifications/cancelled".equals(notification.method())) {
return FilterReply.DENY;
}
return FilterReply.NEUTRAL;
}
}
@@ -14,6 +14,8 @@ import dev.ltms.bridged.herdr.HerdrException;
import dev.ltms.bridged.msg.MessageService;
import dev.ltms.bridged.msg.Rendezvous;
import dev.ltms.bridged.peer.PeerUnreachableException;
import dev.ltms.bridged.placement.BackendQuarantine;
import dev.ltms.bridged.placement.PlacementException;
import dev.ltms.bridged.session.SessionManager;
import dev.ltms.bridged.session.MemberSession;
import dev.ltms.bridged.session.WorktreeRequest;
@@ -33,6 +35,7 @@ import jakarta.servlet.http.HttpServlet;
import java.util.LinkedHashMap;
import java.util.List;
import java.util.Map;
import java.util.Objects;
import java.util.Set;
import java.util.function.Function;
import java.util.function.LongSupplier;
@@ -77,6 +80,8 @@ public final class BridgeMcp {
private final CallerResolver authz; // CB-501: null → authorization not enforced (legacy)
private final Metrics metrics; // CB-502: null → auth failures not counted
private final CapacitySource capacity;
private final HealthCoverageSource healthCoverage;
private final QuarantineSource quarantine;
/** Capacity facts used by {@code bridge_list}; production must supply the placement live count. */
public record CapacitySource(Function<String, Integer> liveCount, Function<String, Integer> maxLoad,
@@ -86,17 +91,34 @@ public final class BridgeMcp {
boolean available() { return !configuredProfiles.get().isEmpty(); }
}
/** Coverage is supplied by the health wiring, not inferred from a missing dependency. */
public record HealthCoverageSource(Supplier<String> value) { }
/**
* CB-578 stage B quarantine facts used by {@code bridge_profiles}: a profile → credential id
* lookup, plus the shared {@link BackendQuarantine} to read remaining cooldowns off.
*/
public record QuarantineSource(Function<String, String> credentialIdFor, BackendQuarantine quarantine) {
/** Inert source — no profile is ever reported quarantined. Explicit stand-in, not a default. */
public static QuarantineSource none() { return new QuarantineSource(_ -> null, BackendQuarantine.none()); }
}
/**
* @param callers resolves each call's {@link Principal}; {@code null} disables authorization.
* This surface needs its own enforcement: {@code /mcp} is a raw servlet on
* Jetty's context handler and never passes through Javalin's {@code before}
* filter, so the REST guard does not cover it.
* @param metrics registry for auth-failure counting; may be {@code null}
* @param quarantine CB-578 stage B facts for {@code bridge_profiles}; required — pass
* {@link QuarantineSource#none()} for a caller that does not want the feature
*/
public BridgeMcp(MessageService messages, PeerLauncher workers, SessionManager sessions,
ConnectionIdentity identity, MemberPresence presence, PrimaryRegistry primaryRegistry,
CallerResolver callers, Metrics metrics, CapacitySource capacity) {
CallerResolver callers, Metrics metrics, CapacitySource capacity, HealthCoverageSource healthCoverage,
QuarantineSource quarantine) {
this.capacity = capacity;
this.quarantine = Objects.requireNonNull(quarantine, "quarantine");
this.healthCoverage = healthCoverage;
McpJsonMapper json = new JacksonMcpJsonMapperSupplier().get();
this.transport = HttpServletStreamableServerTransportProvider.builder()
.jsonMapper(json)
@@ -205,12 +227,13 @@ public final class BridgeMcp {
// CB-301-ext: optional isolated worktree for parallel implementers.
String callerCwd = identity.cwdForPid(callerPid(exchange));
return spawn(sessions, str(a, "profile"), str(a, "role"), str(a, "cwd"), callerCwd,
callerTerminal(exchange), worktreeRequest(a));
callerTerminal(exchange), worktreeRequest(a),
str(a, "sessionName"), str(a, "resumeSessionId"));
})
.toolCall(listTool(), (exchange, _) -> {
McpSchema.CallToolResult denied = deny(exchange, Authz.Action.READ, null);
if (denied != null) return denied;
return listFleet(workers, sessions, messages, capacity,
return listFleet(workers, sessions, messages, capacity, healthCoverage, quarantine,
callers == null ? Map.of() : callers.leads(),
callerTerminal(exchange));
})
@@ -223,7 +246,7 @@ public final class BridgeMcp {
.toolCall(profilesTool(), (exchange, _) -> {
McpSchema.CallToolResult denied = deny(exchange, Authz.Action.READ, null);
if (denied != null) return denied;
return profiles(workers);
return profiles(workers, quarantine);
})
.toolCall(whoamiTool(), (exchange, _) -> {
McpSchema.CallToolResult denied = deny(exchange, Authz.Action.READ, null);
@@ -453,6 +476,11 @@ public final class BridgeMcp {
"[worker finished without a structured bridge_reply — transcript tail follows]\n" + r.text());
// The worker ran the turn then wedged (CB-109) — surface the error context.
case WORKER_FAILED -> text("[worker failed — turn ended in an unrecoverable state]\n" + r.text());
// The backend refused on a subscription usage limit (CB-578 stage A) — the worker's
// pane stayed healthy, but its account is exhausted. Distinct from WORKER_FAILED so the
// primary gets the real cause, not a generic wedge.
case BACKEND_EXHAUSTED -> text("[backend exhausted — the worker's account refused on a "
+ "usage limit]\n" + r.text());
// The worker paused mid-turn to ask (CB-205) — tell the primary how to answer in-turn.
case QUESTION -> text("[question] the worker paused to ask before it can finish:\n" + r.text()
+ "\n\nAnswer it by calling bridge_send again with turnId=\"" + r.turnId()
@@ -549,13 +577,25 @@ public final class BridgeMcp {
return text("acknowledged " + msgId);
}
/** {@code bridge_status}: the live lifecycle status of a worker session. */
/**
* {@code bridge_status}: the live lifecycle status of a worker session, plus — when the worker
* is paused mid-turn in an async {@code bridge_ask} (CB-582) — the open question and how to
* answer it, so a lead on its normal poll cadence does not need the ticket to notice.
*/
static McpSchema.CallToolResult status(MessageService messages, String sessionId) {
if (isBlank(sessionId)) {
return error("sessionId is required");
}
try {
return text(messages.status(sessionId).name().toLowerCase());
String base = messages.status(sessionId).name().toLowerCase();
MessageService.PendingAsk ask = messages.pendingAsk(sessionId);
if (ask == null) {
return text(base);
}
return text(base + "\n\n[question — worker is waiting for your answer]\n" + ask.question()
+ "\n\nAnswer it by calling bridge_send again with turnId=\"" + ask.turnId()
+ "\" and content set to your answer; the worker resumes the same turn."
+ " (ticket " + ask.ticket() + ")");
} catch (HerdrException e) {
return error("herdr error for session " + sessionId + ": " + e.getMessage());
}
@@ -630,7 +670,7 @@ public final class BridgeMcp {
/** {@code bridge_spawn} without cwd/caller context (default resolution). */
static McpSchema.CallToolResult spawn(SessionManager sessions, String profile) {
return spawn(sessions, profile, null, null, null, null, null);
return spawn(sessions, profile, null, null, null, null, null, null, null);
}
/**
@@ -640,13 +680,18 @@ public final class BridgeMcp {
* {@code callerCwd} (the primary's directory), else the daemon's.
* CB-301: the session is registered with {@code ownerTerminal} as its owner.
* CB-301-ext: {@code worktreeRequest} non-null provisions an isolated git worktree.
* CB-584: {@code sessionName}/{@code resumeSessionId} request agent session identity — a resumed
* conversation requires an explicit {@code profile} whose adapter declares
* {@link dev.ltms.bridged.peer.Capability#SESSION_RESUME}, or the spawn is refused rather than
* silently starting a cold session.
*
* <p>{@code role} and {@code profile} are independent: the role picks the contract, the profile
* picks the backend. A reviewer on the same profile as the dev it reviews is a normal spawn.
*/
static McpSchema.CallToolResult spawn(SessionManager sessions, String profile, String role,
String requestedCwd, String callerCwd,
String ownerTerminal, WorktreeRequest worktreeRequest) {
String ownerTerminal, WorktreeRequest worktreeRequest,
String sessionName, String resumeSessionId) {
MemberRole memberRole;
try {
memberRole = isBlank(role) ? MemberRole.DEV : MemberRole.parse(role);
@@ -655,12 +700,17 @@ public final class BridgeMcp {
}
try {
MemberSession member = sessions.acquire(isBlank(profile) ? null : profile, memberRole,
requestedCwd, callerCwd, ownerTerminal, worktreeRequest);
requestedCwd, callerCwd, ownerTerminal, worktreeRequest,
isBlank(sessionName) ? null : sessionName, isBlank(resumeSessionId) ? null : resumeSessionId);
return text(json(memberView(member)));
} catch (GuardException e) {
return error("subscription boundary: " + e.getMessage());
} catch (PlacementException e) {
// CB-599: no candidate had capacity (maxLoad, quarantine, or all-exhausted) — distinct
// from "profile does not exist" below.
return error("no capacity: " + e.getMessage());
} catch (IllegalArgumentException e) {
return error(e.getMessage()); // unknown / no-default profile
return error(e.getMessage()); // unknown / no-default profile, or a refused resumeSessionId
} catch (PeerUnreachableException e) {
return error("spawn timed out — worker pane never reached injectable state: " + e.getMessage());
} catch (HerdrException e) {
@@ -696,11 +746,33 @@ public final class BridgeMcp {
return null;
}
/** {@code bridge_profiles}: the configured worker profiles and the default. */
static McpSchema.CallToolResult profiles(PeerLauncher workers) {
return text(json(Map.of(
"profiles", workers.profiles(),
"default", workers.defaultProfile() == null ? "" : workers.defaultProfile())));
/**
* {@code bridge_profiles}: the configured worker profiles, the default, and — CB-578 stage B —
* which of them are currently quarantined (backend exhausted) and for how much longer. The
* {@code quarantined} key is present only when at least one profile is, so a fleet where nothing
* has ever been quarantined gets exactly the pre-stage-B shape.
*/
static McpSchema.CallToolResult profiles(PeerLauncher workers, QuarantineSource quarantine) {
Map<String, Object> result = new LinkedHashMap<>();
result.put("profiles", workers.profiles());
result.put("default", workers.defaultProfile() == null ? "" : workers.defaultProfile());
Map<String, Object> quarantined = new LinkedHashMap<>();
for (String profile : workers.profiles()) {
String credentialId = quarantine.credentialIdFor().apply(profile);
if (credentialId == null) {
continue;
}
quarantine.quarantine().remainingSeconds(credentialId).ifPresent(remaining -> {
Map<String, Object> row = new LinkedHashMap<>();
row.put("credentialId", credentialId);
row.put("quarantinedForSeconds", remaining);
quarantined.put(profile, row);
});
}
if (!quarantined.isEmpty()) {
result.put("quarantined", quarantined);
}
return text(json(result));
}
/**
@@ -719,17 +791,22 @@ public final class BridgeMcp {
* caller's own row is flagged {@code "self": true}: a peer needs to tell its own pane apart from
* a peer's, and the alternative is every lead calling {@code bridge_whoami} to subtract itself.
*
* <p>CB-583: the {@code capacity} rows reuse {@code quarantine} (the same {@link QuarantineSource}
* {@code bridge_profiles} reads) so the two surfaces cannot disagree about which profile is
* quarantined — see {@link #capacityView}.
*
* @param leads terminal_id → lead name, live from the resolver
* @param selfTerm the calling pane's terminal id, or blank for a caller with no pane
*/
static McpSchema.CallToolResult listFleet(PeerLauncher workers, SessionManager sessions,
Map<String, String> leads, String selfTerm) {
return listFleet(workers, sessions, null, CapacitySource.none(), leads, selfTerm);
return listFleet(workers, sessions, null, CapacitySource.none(), new HealthCoverageSource(() -> "off"),
QuarantineSource.none(), leads, selfTerm);
}
static McpSchema.CallToolResult listFleet(PeerLauncher workers, SessionManager sessions, MessageService messages,
CapacitySource capacity,
Map<String, String> leads, String selfTerm) {
CapacitySource capacity, HealthCoverageSource healthCoverage,
QuarantineSource quarantine, Map<String, String> leads, String selfTerm) {
try {
Map<String, Agent> live = workers.list().stream()
.map(Agent.class::cast)
@@ -747,9 +824,10 @@ public final class BridgeMcp {
roster.stream().map(MemberSession::profile).forEach(profiles::add);
Map<String, Object> result = new LinkedHashMap<>();
result.put("leads", leadRows); result.put("members", out);
result.put("healthCoverage", healthCoverage.value().get());
if (capacity.available()) result.put("capacity", profiles.stream()
.map(profile -> capacityView(profile, capacity.liveCount(), capacity.maxLoad(), roster, messages,
capacity.clock().getAsLong())).toList());
capacity.clock().getAsLong(), quarantine)).toList());
return text(json(result));
} catch (HerdrException e) {
return error("herdr error listing the fleet: " + e.getMessage());
@@ -774,9 +852,18 @@ public final class BridgeMcp {
return row;
}
/**
* CB-583: {@code free} alone cannot tell a lead "busy, will free up" from "refusing, and
* nothing changes for N seconds" — those need different decisions. So a quarantined profile
* forces {@code free} to 0, whatever its {@code maxLoad}/{@code live} say, and the row carries
* the same {@code credentialId}/{@code quarantinedForSeconds} facts {@code bridge_profiles}
* reports, reusing {@link QuarantineSource} rather than a second lookup. Both new keys are
* added only when the profile is actually quarantined, so an ordinary fleet's rows are
* byte-identical to before this change.
*/
private static Map<String, Object> capacityView(String profile, Function<String, Integer> liveCount,
Function<String, Integer> maxLoad, List<MemberSession> roster,
MessageService messages, long nowNanos) {
MessageService messages, long nowNanos, QuarantineSource quarantine) {
Integer cap = maxLoad.apply(profile);
int live = liveCount.apply(profile);
int reclaimable = (int) roster.stream().filter(s -> profile.equals(s.profile()))
@@ -786,6 +873,14 @@ public final class BridgeMcp {
Map<String, Object> row = new LinkedHashMap<>();
row.put("profile", profile); row.put("maxLoad", cap); row.put("live", live);
row.put("free", cap == null ? null : Math.max(0, cap - live)); row.put("reclaimable", reclaimable);
String credentialId = quarantine.credentialIdFor().apply(profile);
if (credentialId != null) {
quarantine.quarantine().remainingSeconds(credentialId).ifPresent(remaining -> {
row.put("free", 0);
row.put("credentialId", credentialId);
row.put("quarantinedForSeconds", remaining);
});
}
return row;
}
@@ -916,20 +1011,30 @@ public final class BridgeMcp {
+ "for the default. The two are independent: a reviewer may run on the same profile "
+ "as the dev it reviews. The member opens your current directory by default; pass "
+ "cwd to pin a different one. Pass worktree:true (with ticket) or "
+ "worktree:<ticket-slug> to provision an isolated git worktree. Returns the member's "
+ "sessionId (use with bridge_send) and paneId (use with bridge_stop).",
+ "worktree:<ticket-slug> to provision an isolated git worktree. Pass resumeSessionId "
+ "to relaunch onto a prior conversation instead of starting cold — this requires an "
+ "explicit profile whose backend supports it (bridge_list shows agentSessionId for "
+ "resumable members), and is refused otherwise rather than silently starting fresh. "
+ "sessionName gives the member a display name in its own UI when the backend supports "
+ "one. Returns the member's sessionId (use with bridge_send) and paneId (use with "
+ "bridge_stop).",
objectSchema(Map.of(
"role", stringProp("What the member is for: architect, dev or reviewer (default dev)"),
"profile", stringProp("Which backend to run it on (omit for the default profile)"),
"cwd", stringProp("Working directory for the member (omit to inherit yours)"),
"worktree", Map.of("type", "string", "description", "'true' or a ticket slug — requests an isolated git worktree"),
"ticket", stringProp("Ticket slug when worktree:true")),
"ticket", stringProp("Ticket slug when worktree:true"),
"sessionName", stringProp("Logical display name for the member's own session, when its backend supports one"),
"resumeSessionId", stringProp("A prior member's agentSessionId (from bridge_list) to resume — requires an explicit profile that supports it")),
List.of()));
}
private static McpSchema.Tool profilesTool() {
return tool("bridge_profiles",
"List the configured worker profiles (backends) and which one bridge_spawn uses by default.",
"List the configured worker profiles (backends) and which one bridge_spawn uses by "
+ "default. A 'quarantined' map is present when a backend-exhausted refusal put "
+ "a profile's credential on cooldown — bridge_spawn onto it is refused until "
+ "quarantinedForSeconds elapses; a profile sharing that credential is listed too.",
objectSchema(Map.of(), List.of()));
}
@@ -941,8 +1046,15 @@ public final class BridgeMcp {
+ "discover a peer lead without being told its address. 'members' are the "
+ "sessions delegated to — each with sessionId, paneId, role (architect/dev/"
+ "reviewer), profile (the backend it runs on), state, optional "
+ "worktree/branch/owner, and live herdr status. An empty 'members' "
+ "means no members are spawned; it says nothing about peers.",
+ "worktree/branch/owner/agentSessionId (the id to pass as bridge_spawn's "
+ "resumeSessionId to relaunch onto that same conversation, when the backend "
+ "supports it), and live herdr status. An empty 'members' "
+ "means no members are spawned; it says nothing about peers. When capacity "
+ "facts are configured, a 'capacity' row per profile also reports free: 0 for "
+ "a quarantined profile's credential (see bridge_profiles), whatever its "
+ "maxLoad/live — with credentialId and quarantinedForSeconds naming the "
+ "quarantine, so 'free: 0, busy' can be told apart from 'free: 0, refusing "
+ "for N seconds'.",
objectSchema(Map.of(), List.of()));
}
@@ -83,6 +83,21 @@ public final class ClaudeCodeLauncher extends HerdrPeerLauncher {
fleet);
}
/**
* Production constructor, plus the CB-596 {@code memberCredentials} policy supplier.
*/
public ClaudeCodeLauncher(AgentControl agents, WorkspaceControl spaces, SubscriptionGuard guard,
Map<String, BridgedConfig.Profile> profiles, String defaultProfile,
Function<String, String> env,
long spawnReadyTimeoutMs, long spawnReadyPollMs,
Supplier<BridgedConfig.Fleet> fleet,
Supplier<BridgedConfig.MemberCredentials> memberCredentials) {
this(agents, spaces, guard, profiles, defaultProfile, env,
spawnReadyTimeoutMs,
System::currentTimeMillis, () -> sleepUninterruptibly(spawnReadyPollMs),
fleet, memberCredentials);
}
/**
* Full testability constructor. Every injectable collaborator is explicit so unit tests supply
* fakes for the clock ({@code nowMillis}) and poll-loop wait ({@code sleeper}). The
@@ -125,6 +140,39 @@ public final class ClaudeCodeLauncher extends HerdrPeerLauncher {
this.guard = guard;
}
/**
* Full testability constructor, plus the CB-596 {@code memberCredentials} policy supplier.
*/
public ClaudeCodeLauncher(AgentControl agents, WorkspaceControl spaces, SubscriptionGuard guard,
Map<String, BridgedConfig.Profile> profiles, String defaultProfile,
Function<String, String> env,
long spawnReadyTimeoutMs,
LongSupplier nowMillis, Runnable sleeper,
Supplier<BridgedConfig.Fleet> fleet,
Supplier<BridgedConfig.MemberCredentials> memberCredentials) {
super(NAME_PREFIX, agents, spaces, profiles, defaultProfile, env,
spawnReadyTimeoutMs, nowMillis, sleeper, fleet, memberCredentials);
this.guard = guard;
}
/**
* Full testability constructor, plus an injectable host-env-names source for the CB-596
* criterion-4 gap detector. Test seam only — every production call site leaves this at the
* default (the real {@code System.getenv()} key set) via the constructor above.
*/
public ClaudeCodeLauncher(AgentControl agents, WorkspaceControl spaces, SubscriptionGuard guard,
Map<String, BridgedConfig.Profile> profiles, String defaultProfile,
Function<String, String> env,
long spawnReadyTimeoutMs,
LongSupplier nowMillis, Runnable sleeper,
Supplier<BridgedConfig.Fleet> fleet,
Supplier<BridgedConfig.MemberCredentials> memberCredentials,
Supplier<Set<String>> hostEnvNames) {
super(NAME_PREFIX, agents, spaces, profiles, defaultProfile, env,
spawnReadyTimeoutMs, nowMillis, sleeper, fleet, memberCredentials, hostEnvNames);
this.guard = guard;
}
/**
* {@inheritDoc}
*
@@ -8,6 +8,7 @@ import dev.ltms.bridged.peer.PeerHandle;
import dev.ltms.bridged.peer.PeerLauncher;
import dev.ltms.bridged.peer.PeerUnreachableException;
import dev.ltms.bridged.peer.SpawnRequest;
import dev.ltms.bridged.placement.BackendQuarantine;
import dev.ltms.bridged.placement.PlacementCandidate;
import dev.ltms.bridged.placement.PlacementContext;
import dev.ltms.bridged.placement.PlacementException;
@@ -23,10 +24,12 @@ import java.util.HashSet;
import java.util.LinkedHashMap;
import java.util.List;
import java.util.Map;
import java.util.Objects;
import java.util.Set;
import java.util.concurrent.ConcurrentHashMap;
import java.util.function.Function;
import java.util.function.Supplier;
import java.util.stream.Collectors;
/**
* The {@link PeerLauncher} the core actually holds when more than one adapter is configured — a thin
@@ -78,6 +81,9 @@ public final class CompositePeerLauncher implements PeerLauncher {
private final Supplier<Map<String, BridgedConfig.Profile>> profileConfigs;
private final Supplier<PlacementPolicy> placementPolicy;
/** CB-578 stage B: credential cooldown, checked before an explicit spawn and filtered into placement. */
private final BackendQuarantine quarantine;
/**
* CB-557: the role pools an unqualified spawn draws its candidates from. A supplier that yields
* {@code null}, and an empty pool for a role, both fall back to every configured profile — the
@@ -99,7 +105,10 @@ public final class CompositePeerLauncher implements PeerLauncher {
}
/**
* Production constructor with a placement policy and live-worker counter.
* Production constructor with a placement policy and live-worker counter. Quarantine (CB-578
* stage B) is off for this constructor — {@link BackendQuarantine#none()} — since it predates
* the feature and existing callers of this exact overload never exercised it; use the 7-arg
* overload below to wire a real {@link BackendQuarantine}.
*
* @param delegates one adapter per configured peer kind; must be non-empty and declare
* disjoint profile-name sets
@@ -113,13 +122,14 @@ public final class CompositePeerLauncher implements PeerLauncher {
Map<String, BridgedConfig.Profile> profileConfigs,
PlacementPolicy placementPolicy,
Function<String, Integer> liveCount) {
this(delegates, defaultProfile, profileConfigs, placementPolicy, liveCount, null);
this(delegates, defaultProfile, profileConfigs, placementPolicy, liveCount, null, BackendQuarantine.none());
}
/**
* Production constructor with role pools (CB-557). An unqualified spawn draws its candidates from
* {@code fleet.<role>} instead of from every configured profile, so a reviewer is placed on a
* reviewer backend and never on, say, the architect-only one.
* reviewer backend and never on, say, the architect-only one. Quarantine is off for this
* constructor too, for the same reason as the 5-arg overload above.
*
* @param fleet the configured role pools; {@code null} ⇒ every profile is a candidate for every
* role, which is the pre-CB-557 behaviour
@@ -130,29 +140,50 @@ public final class CompositePeerLauncher implements PeerLauncher {
PlacementPolicy placementPolicy,
Function<String, Integer> liveCount,
BridgedConfig.Fleet fleet) {
this(delegates, defaultProfile, profileConfigs, placementPolicy, liveCount, fleet, BackendQuarantine.none());
}
/**
* Production constructor with role pools and quarantine (CB-578 stage B). The full-featured
* non-reloading form; {@link #CompositePeerLauncher(List, String, Supplier, Function, BackendQuarantine)}
* is what {@code Bridged.main} actually wires up.
*
* @param quarantine required — pass {@link BackendQuarantine#none()} for a caller that does not
* want the feature, never a defaulting overload (CB-578 stage B's own rule).
*/
public CompositePeerLauncher(List<HerdrPeerLauncher> delegates,
String defaultProfile,
Map<String, BridgedConfig.Profile> profileConfigs,
PlacementPolicy placementPolicy,
Function<String, Integer> liveCount,
BridgedConfig.Fleet fleet,
BackendQuarantine quarantine) {
// LinkedHashMap, not Map.copyOf: candidates() promises definition order and the weighted
// policy breaks exact-weight ties on it, so a salted iteration order would make placement
// differ from one JVM run to the next.
this(delegates, defaultProfile,
constant(Collections.unmodifiableMap(new LinkedHashMap<>(profileConfigs))),
constant(placementPolicy), liveCount, constant(fleet));
constant(placementPolicy), liveCount, constant(fleet), quarantine);
}
/**
* Production constructor that re-reads its placement inputs per spawn (CB-559), so a config
* reload retargets the next member without a restart.
*
* @param config the live configuration — read at every spawn, never captured
* @param config the live configuration — read at every spawn, never captured
* @param quarantine required — CB-578 stage B; pass {@link BackendQuarantine#none()} to opt out
*/
public CompositePeerLauncher(List<HerdrPeerLauncher> delegates,
String defaultProfile,
Supplier<BridgedConfig> config,
Function<String, Integer> liveCount) {
Function<String, Integer> liveCount,
BackendQuarantine quarantine) {
this(delegates, defaultProfile,
() -> config.get().profiles(),
() -> PlacementPolicies.fromName(config.get().placement()),
liveCount,
() -> config.get().fleet());
() -> config.get().fleet(),
quarantine);
}
/** The all-suppliers form every other constructor funnels into. */
@@ -161,8 +192,10 @@ public final class CompositePeerLauncher implements PeerLauncher {
Supplier<Map<String, BridgedConfig.Profile>> profileConfigs,
Supplier<PlacementPolicy> placementPolicy,
Function<String, Integer> liveCount,
Supplier<BridgedConfig.Fleet> fleet) {
Supplier<BridgedConfig.Fleet> fleet,
BackendQuarantine quarantine) {
this.fleet = fleet;
this.quarantine = Objects.requireNonNull(quarantine, "quarantine");
if (delegates.isEmpty()) {
throw new IllegalArgumentException("at least one peer adapter must be configured");
}
@@ -225,6 +258,7 @@ public final class CompositePeerLauncher implements PeerLauncher {
// the charter makes explicit-profile spawns the normal path — so skipping the check
// here would leave the cap dead config in real operation.
HerdrPeerLauncher d = route(requestedProfile);
enforceNotQuarantined(requestedProfile);
enforceMaxLoad(requestedProfile);
PeerHandle handle = d.spawn(req);
spawnedBy.put(handle.id(), d);
@@ -238,7 +272,10 @@ public final class CompositePeerLauncher implements PeerLauncher {
List<PlacementCandidate> candidates = candidates(req.role());
String roleDefault = defaultProfileFor(req.role());
Set<String> unreachable = new HashSet<>();
PlacementContext ctx = new PlacementContext(roleDefault, candidates, liveCount, unreachable);
// CB-578 stage B: computed once up front — a quarantine's expiry cannot pass within one spawn
// call, so re-deriving it per retry would only cost work, never change the answer.
Set<String> quarantined = quarantinedProfiles(candidates);
PlacementContext ctx = new PlacementContext(roleDefault, candidates, liveCount, unreachable, quarantined);
int maxAttempts = candidates.isEmpty() ? 1 : candidates.size();
for (int attempt = 0; attempt < maxAttempts; attempt++) {
@@ -268,7 +305,7 @@ public final class CompositePeerLauncher implements PeerLauncher {
chosen.profile(), e.getMessage());
unreachable.add(chosen.profile());
// Update the context for the next selection so the policy excludes this profile.
ctx = new PlacementContext(roleDefault, candidates, liveCount, unreachable);
ctx = new PlacementContext(roleDefault, candidates, liveCount, unreachable, quarantined);
}
}
@@ -298,9 +335,48 @@ public final class CompositePeerLauncher implements PeerLauncher {
* @param profile the profile the caller explicitly named
* @throws PlacementException when the profile is at capacity
*/
/**
* Refuse an explicit-profile spawn whose credential is quarantined (CB-578 stage B): a prior
* {@code BACKEND_EXHAUSTED} classification on this profile, or on another profile sharing its
* {@code credentialId}, is still on cooldown.
*
* <p>Checked before {@link #enforceMaxLoad}, and for the same reason that check exists: an
* explicit profile bypasses the placement policy's filtering entirely, so without this the cap
* (there, quarantine here) would be dead config on the very path the charter calls normal.
* Deliberately no fallback to another profile, matching {@link #enforceMaxLoad}'s own reasoning —
* the caller named this profile for a cost/model reason.
*
* @throws PlacementException naming the profile, its credential, and the remaining cooldown
*/
private void enforceNotQuarantined(String profile) {
String credentialId = credentialIdFor(profile);
quarantine.remainingSeconds(credentialId).ifPresent(remaining -> {
throw new PlacementException("worker profile '" + profile + "' is quarantined "
+ "(credential '" + credentialId + "' exhausted; ~" + remaining
+ "s remaining) — refusing spawn");
});
}
/** {@code profile}'s credential group (CB-578 stage B), or the profile's own name if unconfigured. */
private String credentialIdFor(String profile) {
BridgedConfig.Profile cfg = profiles0().get(profile);
return cfg == null ? profile : cfg.effectiveCredentialId();
}
/** The subset of {@code candidates} whose credential is currently quarantined (CB-578 stage B). */
private Set<String> quarantinedProfiles(List<PlacementCandidate> candidates) {
return candidates.stream()
.map(PlacementCandidate::profile)
.filter(p -> quarantine.isQuarantined(credentialIdFor(p)))
.collect(Collectors.toSet());
}
private void enforceMaxLoad(String profile) {
// Absent config, or a config whose maxLoad normalized to null (non-positive ⇒ unlimited at
// load), means no cap — never cap what wasn't configured.
// Absent config, or a config whose maxLoad normalized to null (ABSENT ⇒ unlimited at load),
// means no cap — never cap what wasn't configured. Note "non-positive ⇒ unlimited" was true
// until CB-585: an explicit `maxLoad: 0` now survives as 0 and is a real cap of zero, so the
// check below refuses every spawn on that profile, and a negative value is refused at config
// load rather than normalized away.
BridgedConfig.Profile cfg = profiles0().get(profile);
Integer cap = (cfg == null) ? null : cfg.maxLoad();
if (cap == null) {
@@ -387,6 +463,18 @@ public final class CompositePeerLauncher implements PeerLauncher {
return defaultProfile;
}
/**
* {@inheritDoc}
*
* <p>Routes to the specific delegate {@code profileName} resolves to, not the fleet-wide
* union {@link #capabilities()} returns — the whole reason this method exists (CB-584): in a
* mixed fleet, one adapter's capability must never be read as every profile's.
*/
@Override
public Set<Capability> capabilitiesFor(String profileName) {
return route(profileName).capabilities();
}
/** Every herdr agent, deduplicated by pane id (all delegates share one herdr and list globally). */
@Override
public List<Agent> list() {
@@ -7,6 +7,8 @@ import dev.ltms.bridged.herdr.HerdrException;
import dev.ltms.bridged.herdr.Tab;
import dev.ltms.bridged.herdr.Workspace;
import dev.ltms.bridged.herdr.WorkspaceControl;
import dev.ltms.bridged.peer.Capability;
import dev.ltms.bridged.peer.CharterReceipt;
import dev.ltms.bridged.peer.MemberRole;
import dev.ltms.bridged.peer.PeerHandle;
import dev.ltms.bridged.peer.PeerLauncher;
@@ -18,6 +20,7 @@ import org.slf4j.LoggerFactory;
import java.security.SecureRandom;
import java.util.ArrayList;
import java.util.Collection;
import java.util.HashSet;
import java.util.LinkedHashMap;
import java.util.List;
import java.util.Map;
@@ -88,6 +91,27 @@ public abstract class HerdrPeerLauncher implements PeerLauncher {
*/
private final Supplier<BridgedConfig.Fleet> fleet;
/**
* CB-596: the live {@code memberCredentials:} policy, read once per spawn (same hot-reload shape
* as {@link #fleet}). {@code null} — either the supplier itself, or what it returns — means no
* policy is configured and {@link #applyMemberCredentialPolicy} shadows nothing.
*/
private final Supplier<BridgedConfig.MemberCredentials> memberCredentials;
/**
* Enumerates the daemon's own process environment variable NAMES ONLY, never values — the CB-596
* criterion-4 gap detector's data source (see {@link #logCredentialGap}). Injectable for tests;
* production always resolves to the real {@code System.getenv()} key set.
*
* <p>Deliberately the daemon's own environment, not the spawned pane's: nothing in the herdr
* client surface lets the daemon read back an arbitrary command's output from a pane before the
* peer starts in it, so there is no channel to inspect the pane's environment directly. The
* daemon's own process is started the same way (a login shell sourcing the same secret store —
* see CB-592's investigation of {@code secrets.sh}), so on a single-host deployment its env
* mirrors what the pane's login shell is about to export.
*/
private final Supplier<Set<String>> hostEnvNames;
/** The final instruction always requires a bridge reply when the bridge MCP is mounted. */
protected static final String REPLY_CHARTER =
"You are a spawned member in the claude-bridge fleet. Every message you receive arrives "
@@ -164,6 +188,41 @@ public abstract class HerdrPeerLauncher implements PeerLauncher {
long spawnReadyTimeoutMs,
LongSupplier nowMillis, Runnable sleeper,
Supplier<BridgedConfig.Fleet> fleet) {
this(namePrefix, agents, spaces, profiles, defaultProfile, env, spawnReadyTimeoutMs,
nowMillis, sleeper, fleet, null);
}
/**
* As above, plus the live {@code memberCredentials} policy (CB-596).
*
* @param memberCredentials live member-credential policy, read once per spawn; {@code null} ⇒
* no policy configured, so a spawn shadows nothing. A separate
* constructor rather than a new parameter on the one above, so every
* existing call site keeps the pre-CB-596 default without an edit.
*/
protected HerdrPeerLauncher(String namePrefix, AgentControl agents, WorkspaceControl spaces,
Map<String, BridgedConfig.Profile> profiles, String defaultProfile,
Function<String, String> env,
long spawnReadyTimeoutMs,
LongSupplier nowMillis, Runnable sleeper,
Supplier<BridgedConfig.Fleet> fleet,
Supplier<BridgedConfig.MemberCredentials> memberCredentials) {
this(namePrefix, agents, spaces, profiles, defaultProfile, env, spawnReadyTimeoutMs,
nowMillis, sleeper, fleet, memberCredentials, null);
}
/**
* As above, plus an injectable {@link #hostEnvNames} source for the CB-596 gap detector. Test
* seam only — every production call site leaves this {@code null} and gets the real host env.
*/
protected HerdrPeerLauncher(String namePrefix, AgentControl agents, WorkspaceControl spaces,
Map<String, BridgedConfig.Profile> profiles, String defaultProfile,
Function<String, String> env,
long spawnReadyTimeoutMs,
LongSupplier nowMillis, Runnable sleeper,
Supplier<BridgedConfig.Fleet> fleet,
Supplier<BridgedConfig.MemberCredentials> memberCredentials,
Supplier<Set<String>> hostEnvNames) {
this.fleet = fleet;
this.namePrefix = namePrefix;
this.agents = agents;
@@ -174,6 +233,8 @@ public abstract class HerdrPeerLauncher implements PeerLauncher {
this.spawnReadyTimeoutMs = spawnReadyTimeoutMs;
this.nowMillis = nowMillis;
this.sleeper = sleeper;
this.memberCredentials = memberCredentials;
this.hostEnvNames = hostEnvNames != null ? hostEnvNames : () -> System.getenv().keySet();
}
// --- adapter seams -------------------------------------------------------------------------
@@ -247,6 +308,19 @@ public abstract class HerdrPeerLauncher implements PeerLauncher {
return defaultProfile;
}
/**
* {@inheritDoc}
*
* <p>One {@link HerdrPeerLauncher} instance always serves exactly one adapter kind, so every
* profile it owns shares that adapter's {@link #capabilities()} — {@code profileName} only
* needs validating (throwing on an unknown profile, same as {@link #spawn}), not routing.
*/
@Override
public Set<Capability> capabilitiesFor(String profileName) {
requireProfile(profileName);
return capabilities();
}
/** The configured profiles, for adapter capability decisions (e.g. any git-token grant). */
protected Collection<BridgedConfig.Profile> profileConfigs() {
return profiles.values();
@@ -269,8 +343,11 @@ public abstract class HerdrPeerLauncher implements PeerLauncher {
// --- spawn ---------------------------------------------------------------------------------
/** A started peer plus the launch's agent-session id (the resume handle, or null). */
private record Spawned(Agent agent, String agentSessionId) {
/**
* A started peer plus the launch's agent-session id (the resume handle, or null) and the
* charter receipt (CB-571) the base composed for it.
*/
private record Spawned(Agent agent, String agentSessionId, CharterReceipt receipt) {
}
/**
@@ -305,12 +382,41 @@ public abstract class HerdrPeerLauncher implements PeerLauncher {
String replyCharter = cfg.hasMcp() ? REPLY_CHARTER : null;
String charter = roleCharter == null ? replyCharter
: replyCharter == null ? roleCharter : roleCharter + "\n\n" + replyCharter;
Launch launch = buildLaunch(cfg, new LaunchSpec(sessionName, resumeSessionId, role, charter));
String cwd = resolveCwd(requestedCwd, cfg, callerCwd);
Agent agent = cfg.tabPlacement()
? spawnInTab(cfg, launch.env(), launch.argv(), cwd, role, liveFleet)
: spawnAsPane(cfg, launch.env(), launch.argv(), cwd);
return new Spawned(agent, launch.agentSessionId());
// CB-571: fingerprint the exact composed charter bytes once, here in the base, before the
// string leaves for an adapter — so Claude and OpenCode derive the same digest. A failed
// start has no bridge_spawn result and no roster row, so the failure log below is the only
// surface the byte count can appear on. The charter text itself is never logged.
CharterReceipt receipt = CharterReceipt.compose(role, cfg.profile(), roleCharter, charter);
try {
Launch launch = buildLaunch(cfg, new LaunchSpec(sessionName, resumeSessionId, role, charter));
String cwd = resolveCwd(requestedCwd, cfg, callerCwd);
Agent agent = cfg.tabPlacement()
? spawnInTab(cfg, launch.env(), launch.argv(), cwd, role, liveFleet)
: spawnAsPane(cfg, launch.env(), launch.argv(), cwd, charter);
logCharterReceipt(receipt, true);
return new Spawned(agent, launch.agentSessionId(), receipt);
} catch (RuntimeException e) {
logCharterReceipt(receipt, false);
throw e;
}
}
/**
* The one place the charter's size and digest appear in the logs. {@code success} true after a
* start, false from the failure path of {@link #spawnInternal} where no handle or roster row
* exists to carry the receipt. Always metadata only — never the charter text.
*/
private static void logCharterReceipt(CharterReceipt receipt, boolean success) {
String role = receipt.role() == null ? "" : receipt.role().wireName();
if (success) {
log.info("spawned role={} profile={} charterSource={} charterSha256={} charterBytes={}",
role, receipt.profile(), receipt.charterSource(),
receipt.charterSha256(), receipt.charterBytes());
} else {
log.warn("spawn failed; charter role={} profile={} charterSource={} charterSha256={} charterBytes={}",
role, receipt.profile(), receipt.charterSource(),
receipt.charterSha256(), receipt.charterBytes());
}
}
/**
@@ -349,7 +455,7 @@ public abstract class HerdrPeerLauncher implements PeerLauncher {
String id = UUID.randomUUID().toString();
paneByAgentId.put(id, paneId);
return new WorkerHandle(id, agent.terminalId(), requireProfile(req.profileName()).profile(),
req.sessionName(), spawned.agentSessionId());
req.sessionName(), spawned.agentSessionId(), spawned.receipt());
}
@Override
@@ -433,11 +539,19 @@ public abstract class HerdrPeerLauncher implements PeerLauncher {
}
}
/** Legacy placement: split the currently-focused tab; the peer still starts in {@code cwd}. */
/**
* Legacy placement: split the currently-focused tab; the peer still starts in {@code cwd}.
*
* <p>CB-571: this is the one legacy log that printed the full argv, and the charter travels
* inside argv — so the charter text went to the daemon log on every pane-placement spawn. The
* {@code spawnInTab} path never logs argv, so only this site is fixed. {@code charter} is the
* composed charter, if any; its argv element is replaced by its digest so the log still shows
* which args were passed without exposing the charter prose.
*/
private Agent spawnAsPane(BridgedConfig.Profile cfg, Map<String, String> workerEnv,
List<String> argv, String cwd) {
List<String> argv, String cwd, String charter) {
log.info("spawning {} (pane placement) profile={} cwd={} argv={}",
namePrefix, cfg.profile(), cwd, argv);
namePrefix, cfg.profile(), cwd, redactCharter(argv, charter));
String paneId = spaces.splitPane(cwd, workerEnv);
if (paneId == null) {
throw new IllegalStateException("pane.split returned no pane — cannot start a peer");
@@ -447,6 +561,21 @@ public abstract class HerdrPeerLauncher implements PeerLauncher {
return peer;
}
/**
* A copy of {@code argv} with an element equal to {@code charter} replaced by its digest, so
* the pane log never prints the charter prose. The charter is handed to an adapter as one argv
* element, so exact-equality is the right match; every other argument passes through unchanged.
*/
private static List<String> redactCharter(List<String> argv, String charter) {
if (charter == null || charter.isBlank() || argv == null || argv.isEmpty()) {
return argv;
}
String digest = CharterReceipt.digestOf(charter);
return argv.stream()
.map(a -> a.equals(charter) ? "<charter sha256=" + digest + ">" : a)
.toList();
}
/** A started peer together with the sequence its unique name/label used. */
private record Started(Agent agent, long seq) {
}
@@ -651,11 +780,18 @@ public abstract class HerdrPeerLauncher implements PeerLauncher {
/**
* A concrete {@link PeerHandle} wrapping herdr agent coordinates, the profile that spawned it,
* and the session identity the launch resolved (CB-547a): the bridge's logical name and the
* peer's own session id, both null when the spawn carried no identity.
* the session identity the launch resolved (CB-547a): the bridge's logical name and the peer's
* own session id, both null when the spawn carried no identity — and the charter receipt
* (CB-571) the base computed for this launch.
*/
private record WorkerHandle(String id, String terminalId, String profile,
String sessionName, String agentSessionId) implements PeerHandle {
String sessionName, String agentSessionId,
CharterReceipt receipt) implements PeerHandle {
@Override
public CharterReceipt charterReceipt() {
return receipt;
}
}
// --- shared helpers ------------------------------------------------------------------------
@@ -689,10 +825,62 @@ public abstract class HerdrPeerLauncher implements PeerLauncher {
}
}
/** A fresh mutable env map — the conventional starting point for {@link #buildLaunch}. */
/**
* CB-596: overlay value that shadows any host credential a herdr pane otherwise inherits from
* herdr's own login-shell process environment (gitea issue #82, superseding CB-592's single
* hardcoded {@code GITEA_ACCESS_TOKEN} name — see {@link #applyMemberCredentialPolicy}). herdr
* spawns a pane from its <em>own</em> process environment and layers our map on top —
* {@link dev.ltms.bridged.herdr.WorkspaceControl#createTab} and {@code #splitPane} send only
* the keys we put in that map, so any key we never mention passes straight through from
* herdr's own shell, admin credentials included.
*
* <p>Deliberately a non-blank sentinel, not {@code ""}. Whether an empty-string overlay value
* overrides an inherited variable or is skipped as blank could not be settled by reading this
* codebase — herdr's server-side merge is an external process, not something in this repo.
* A non-blank replacement sidesteps that ambiguity entirely: {@link #baseEnv}'s own {@code
* PATH} seeding already depends on the overlay reliably replacing an inherited value (see its
* javadoc), and that is only demonstrated for a non-blank value, so this reuses the same,
* proven-reliable shape rather than the unverified one.
*
* <p><b>MEASURED ON A LIVE PANE, 2026-08-15 (CB-592): this sentinel alone does NOT hold.</b> The
* overlay itself works — {@code GITEA_TOKEN} is injected here, is exported by no shell file, and
* does reach the pane. The sentinel loses one step later. A herdr pane runs a <em>login</em>
* shell, {@code ~/.zprofile} sources {@code ${SHARED_ENV}/tools/secrets.sh}, and that file does a
* plain unconditional {@code export GITEA_ACCESS_TOKEN=...}. A login shell overwrites a value
* already in the environment, so the real admin token is put back over this sentinel before the
* member process ever starts. That defeat applies to <em>every</em> name {@code secrets.sh}
* exports, and no launcher-side overlay can win against it.
*
* <p>So this constant is not the control on its own — {@link #MEMBER_MARKER} is the other half.
* Keeping the sentinel is still worth it: it is correct for any peer kind whose pane does not
* start a login shell, and it makes the intent explicit at the one place every adapter passes.
*/
private static final String BLOCKED_CREDENTIAL_SENTINEL =
"blocked-by-bridged-cb596-see-gitea-issue-82";
/**
* CB-592: marks a pane as a bridged member so a shell startup file can decline to export
* operator-only credentials into it (gitea issue #77).
*
* <p>This name is deliberately one that {@code secrets.sh} never exports, which is exactly why
* it survives the login shell that wipes {@link #BLOCKED_GITEA_ACCESS_TOKEN}. The mechanism is
* measured, not assumed: {@code GITEA_TOKEN} is injected the same way, is absent from a login
* shell of its own, and was observed set inside a live member pane.
*
* <p>It is a no-op until the operator guards the export, which is a one-line change in a file
* this repo does not own and must not edit unasked:
*
* <pre>{@code
* [ -n "${BRIDGED_MEMBER:-}" ] || export GITEA_ACCESS_TOKEN=...
* }</pre>
*
* <p>Setting the marker now costs nothing and means that edit is the whole remaining fix.
*/
static final String MEMBER_MARKER = "BRIDGED_MEMBER";
/**
* Seed a worker's environment (CB-511): the daemon's own {@code PATH}, then the profile's
* {@code env:} entries.
* {@code env:} entries, then the CB-592 admin-token shadow.
*
* <p>Why this exists: bridged passes herdr an explicit env map, and herdr merges it into
* <em>its own</em> process environment. So before this, a worker inherited whatever PATH the
@@ -706,6 +894,12 @@ public abstract class HerdrPeerLauncher implements PeerLauncher {
* overriding {@code ANTHROPIC_BASE_URL} and slipping past {@link
* dev.ltms.bridged.guard.SubscriptionGuard}, which is checked against the profile's
* {@code baseUrl} and nothing else.
*
* <p>The CB-596 credential shadow and the CB-592 marker are put in <em>last</em>, after the
* profile's own {@code env:}, so no profile — present or future — can restore a blocked
* credential, or hide that the pane is a member, by naming either in config. This is the one
* place both are applied: every {@code buildLaunch} in every adapter calls this first, so a new
* profile, and a peer kind not yet written, gets them for free.
*/
protected Map<String, String> baseEnv(BridgedConfig.Profile cfg) {
Map<String, String> workerEnv = new LinkedHashMap<>();
@@ -716,9 +910,70 @@ public abstract class HerdrPeerLauncher implements PeerLauncher {
if (cfg != null && cfg.env() != null) {
workerEnv.putAll(cfg.env());
}
applyMemberCredentialPolicy(workerEnv);
workerEnv.put(MEMBER_MARKER, "1");
return workerEnv;
}
/**
* CB-596: shadow every configured {@code memberCredentials.known} name that is not also
* {@code allow}-ed, replacing CB-592's single hardcoded {@code GITEA_ACCESS_TOKEN} name (gitea
* issue #82). An allow-listed name is deliberately left unmentioned here — see {@link
* #BLOCKED_CREDENTIAL_SENTINEL}'s javadoc for why an overlay entry is the only way to shadow an
* inherited value, which is exactly why an allowed name must get NO entry: any entry at all,
* blank or not, risks overriding the real value the pane needs.
*
* <p>No {@code memberCredentials} configured — the supplier is {@code null}, or it resolves to
* one whose {@code known} list is empty — shadows nothing. This is a real, config-driven gap
* (see {@link BridgedConfig.MemberCredentials}'s javadoc), not a safe default: deny-by-default
* only defends names the operator has actually enumerated in {@code known}.
*/
private void applyMemberCredentialPolicy(Map<String, String> workerEnv) {
BridgedConfig.MemberCredentials creds = memberCredentials == null ? null : memberCredentials.get();
if (creds == null) {
return;
}
for (String name : creds.blockedSet()) {
workerEnv.put(name, BLOCKED_CREDENTIAL_SENTINEL);
}
logCredentialGap(creds);
}
/** Credential-shaped env var name heuristic for {@link #logCredentialGap} — case-insensitive. */
private static final Pattern CREDENTIAL_SHAPED_NAME =
Pattern.compile("(?i).*(TOKEN|SECRET|_KEY|APIKEY|PASSWORD|CREDENTIAL|AUTH).*");
/** Guards {@link #logCredentialGap} to one WARN per launcher instance, not one per spawn. */
private final AtomicBoolean credentialGapLogged = new AtomicBoolean();
/**
* CB-596 criterion 4: a credential-shaped host env var name on neither {@code known} nor
* {@code allow} is not silently allowed — it is reported. {@link #hostEnvNames} enumerates the
* daemon's own environment (see that field's javadoc for why the daemon's env is read rather
* than the spawned pane's, which the daemon has no channel to inspect at spawn time); this logs
* every such NAME, at WARN, at most once per launcher instance — never a value, a prefix of a
* value, or a hash of a value, so the log itself cannot leak anything.
*/
private void logCredentialGap(BridgedConfig.MemberCredentials creds) {
Set<String> covered = new HashSet<>(creds.known());
covered.addAll(creds.allow());
List<String> gap = hostEnvNames.get().stream()
.filter(name -> CREDENTIAL_SHAPED_NAME.matcher(name).matches())
.filter(name -> !covered.contains(name))
.sorted()
.toList();
if (gap.isEmpty()) {
return;
}
if (credentialGapLogged.compareAndSet(false, true)) {
log.warn("memberCredentials gap: {} credential-shaped env var name(s) are on neither "
+ "known: nor allow: — every member pane inherits them UNBLOCKED — {}. "
+ "Add each to memberCredentials.known (blocked by default) or .allow "
+ "(if a member legitimately needs it).",
gap.size(), gap);
}
}
/** Defensive copy of {@code argv} plus room to append launch flags. */
protected static List<String> mutableArgv(List<String> argv) {
return new ArrayList<>(argv);
@@ -7,6 +7,7 @@ import dev.ltms.bridged.herdr.Agent;
import dev.ltms.bridged.herdr.AgentControl;
import dev.ltms.bridged.herdr.WorkspaceControl;
import dev.ltms.bridged.peer.Capability;
import dev.ltms.bridged.peer.CharterReceipt;
import dev.ltms.bridged.peer.PeerHandle;
import dev.ltms.bridged.peer.SpawnRequest;
@@ -103,6 +104,20 @@ public final class OpenCodeLauncher extends HerdrPeerLauncher {
defaultConfigRoot(), defaultDiscoveryRoot(), fleet);
}
/**
* Production constructor, plus the CB-596 {@code memberCredentials} policy supplier.
*/
public OpenCodeLauncher(AgentControl agents, WorkspaceControl spaces,
Map<String, BridgedConfig.Profile> profiles, String defaultProfile,
Function<String, String> env,
long spawnReadyTimeoutMs, long spawnReadyPollMs,
Supplier<BridgedConfig.Fleet> fleet,
Supplier<BridgedConfig.MemberCredentials> memberCredentials) {
this(agents, spaces, profiles, defaultProfile, env, spawnReadyTimeoutMs,
System::currentTimeMillis, () -> sleepUninterruptibly(spawnReadyPollMs),
defaultConfigRoot(), defaultDiscoveryRoot(), fleet, memberCredentials);
}
/**
* Full testability constructor. Every injectable collaborator is explicit so unit tests supply a
* fake clock ({@code nowMillis}), poll-loop wait ({@code sleeper}), and a temp {@code configRoot}
@@ -150,6 +165,23 @@ public final class OpenCodeLauncher extends HerdrPeerLauncher {
this.discovery = new OpenCodeSessionDiscovery(discoveryRoot);
}
/**
* Full testability constructor, plus the CB-596 {@code memberCredentials} policy supplier.
*/
public OpenCodeLauncher(AgentControl agents, WorkspaceControl spaces,
Map<String, BridgedConfig.Profile> profiles, String defaultProfile,
Function<String, String> env,
long spawnReadyTimeoutMs,
LongSupplier nowMillis, Runnable sleeper,
Path configRoot, Path discoveryRoot,
Supplier<BridgedConfig.Fleet> fleet,
Supplier<BridgedConfig.MemberCredentials> memberCredentials) {
super(NAME_PREFIX, agents, spaces, profiles, defaultProfile, env,
spawnReadyTimeoutMs, nowMillis, sleeper, fleet, memberCredentials);
this.configRoot = configRoot;
this.discovery = new OpenCodeSessionDiscovery(discoveryRoot);
}
private static Path defaultConfigRoot() {
return Path.of(System.getProperty("java.io.tmpdir"));
}
@@ -406,6 +438,11 @@ public final class OpenCodeLauncher extends HerdrPeerLauncher {
// appeared).
return discovery.sessionIdForDirectory(cwd);
}
@Override
public CharterReceipt charterReceipt() {
return delegate.charterReceipt();
}
}
// --- Agent-returning convenience spawns (used by callers/tests that want the herdr Agent) ---
@@ -7,6 +7,7 @@ import com.rabbitmq.client.ConnectionFactory;
import com.rabbitmq.client.DeliverCallback;
import com.rabbitmq.client.Recoverable;
import com.rabbitmq.client.RecoveryListener;
import com.rabbitmq.client.Return;
import org.slf4j.Logger;
import org.slf4j.LoggerFactory;
@@ -14,7 +15,13 @@ import java.io.IOException;
import java.nio.charset.StandardCharsets;
import java.util.LinkedHashMap;
import java.util.List;
import java.util.NavigableMap;
import java.util.concurrent.CompletableFuture;
import java.util.concurrent.ConcurrentHashMap;
import java.util.concurrent.ConcurrentSkipListMap;
import java.util.concurrent.ExecutionException;
import java.util.concurrent.TimeUnit;
import java.util.concurrent.TimeoutException;
/**
* AMQP-backed {@link ReplyInbox} (CB-307 Stage 2): genuine cross-restart durability behind the same
@@ -29,6 +36,27 @@ import java.util.concurrent.ConcurrentHashMap;
* leaves them on the broker — it redelivers on reconnect. That is the durability the in-memory
* adapter cannot give, with the port contract preserved.
*
* <p><strong>Prefetch bounds the held backlog (CB-527).</strong> The consumer channel calls
* {@code basicQos} with a configurable prefetch count ({@link #DEFAULT_PREFETCH} unless the caller
* passes another value to {@link #open(String, int)}) before starting any consumer. Without a bound,
* the broker pushes its entire queue into {@link #held} the instant a target is {@link #own owned},
* so an undrained primary grows the JVM heap without limit and any queue-level control
* ({@code x-max-length}, per-message TTL) never fires because the queue never actually holds a
* backlog. Prefetch keeps the backlog where it is visible — on the broker — until the owner drains it.
*
* <p><strong>Publishes require a confirmed, routable delivery (CB-528).</strong> {@link #publish}
* runs on a channel separate from the consume/ack channel ({@link #channel}), so a slow or blocked
* publish confirm can never hold {@link #channelLock} and stall an ack — the ack path never waits on
* a publish confirm. That publish channel is in publisher-confirm mode and every publish sets the
* {@code mandatory} flag, so an unroutable publish (queue not declared, e.g. the owner never called
* {@link #own}) is returned by the broker instead of silently dropped. The broker sends the
* <em>return</em> for an unroutable message before the <em>confirm</em> that covers it — the ack/nack
* callback checks the returned-set at confirm time rather than assuming an ack means routed — so
* "confirmed" here means "durably queued", not merely "accepted by the broker". A returned or nacked
* (or un-confirmed within the timeout) publish surfaces as an {@link IllegalStateException} on the
* caller's thread; the caller — {@link MessageService#reply} — must not report success for a
* black-holed reply.
*
* <p><strong>Ownership is explicit.</strong> {@link #own} declares the queue and starts the consumer;
* {@link #release} cancels it. {@link #publish} sends to the queue but does <em>not</em> imply ownership
* and does not attach a consumer. This split is required by CB-308 federation, where one gateway may
@@ -54,6 +82,12 @@ public final class AmqpReplyInbox implements ReplyInbox, AutoCloseable {
private static final String QUEUE_PREFIX = "agent.";
private static final String QUEUE_SUFFIX = ".inbox";
/** CB-527: the prefetch used when a caller does not pass an explicit value to {@link #open(String, int)}. */
public static final int DEFAULT_PREFETCH = 32;
/** How long {@link #publish} waits for its publisher confirm before failing the call (CB-528). */
private static final long CONFIRM_TIMEOUT_MS = 10_000L;
private final Connection connection;
private final Channel channel;
/** All channel operations (publish/declare/ack/cancel) serialize on this — a Channel is not thread-safe. */
@@ -63,39 +97,87 @@ public final class AmqpReplyInbox implements ReplyInbox, AutoCloseable {
/** Targets whose queue is declared and consumer is running, mapped to their broker consumer tag. */
private final ConcurrentHashMap<String, String> consumerTags = new ConcurrentHashMap<>();
/**
* CB-528: a dedicated channel for {@link #publish}, kept separate from {@link #channel} (consume
* + ack) so a publish confirm round trip never blocks under {@link #channelLock} and stalls an ack.
*/
private final Channel publishChannel;
private final Object publishChannelLock = new Object();
/** In-flight publishes awaiting their confirm, keyed by the publish channel's sequence number. */
private final ConcurrentSkipListMap<Long, Pending> pendingBySeq = new ConcurrentSkipListMap<>();
/**
* The same in-flight publishes, keyed by {@code msgId} — a broker {@code Return} carries no delivery
* tag. Assumes {@code msgId} is unique per in-flight publish: a second {@link #publish} for a
* {@code msgId} still awaiting its confirm would overwrite this entry and misdirect
* {@link #onReturn}'s lookup. Not reachable today — {@code MessageService.reply} generates a fresh
* {@code UUID} per call — so no guard is added for it.
*/
private final ConcurrentHashMap<String, Pending> pendingByMsgId = new ConcurrentHashMap<>();
/** A message pulled off the broker but not yet acked: its delivery-tag plus the port payload. */
private record Held(long deliveryTag, InboxMessage message) {}
/** Connect to {@code uri} (e.g. {@code amqp://guest:guest@127.0.0.1:5672/}) and open the inbox. */
/** A publish awaiting its confirm; {@link #returned} records whether the broker already returned it. */
private static final class Pending {
final String msgId;
final CompletableFuture<Void> confirmed = new CompletableFuture<>();
volatile boolean returned;
Pending(String msgId) {
this.msgId = msgId;
}
}
/** Connect to {@code uri} (e.g. {@code amqp://guest:guest@127.0.0.1:5672/}) with {@link #DEFAULT_PREFETCH}. */
public static AmqpReplyInbox open(String uri) {
return open(uri, DEFAULT_PREFETCH);
}
/** As {@link #open(String)}, with an explicit consumer prefetch (CB-527: caps the held backlog per target). */
public static AmqpReplyInbox open(String uri, int prefetch) {
try {
ConnectionFactory factory = new ConnectionFactory();
factory.setUri(uri);
// Self-heal transient blips; topology recovery re-declares queues and re-attaches consumers.
factory.setAutomaticRecoveryEnabled(true);
factory.setTopologyRecoveryEnabled(true);
return new AmqpReplyInbox(factory.newConnection("bridged-reply-inbox"));
return new AmqpReplyInbox(factory.newConnection("bridged-reply-inbox"), prefetch);
} catch (Exception e) {
throw new IllegalStateException("cannot connect to AMQP broker at " + uri, e);
}
}
/** Wrap an already-open connection (injection seam for the contract test). */
/** Wrap an already-open connection with {@link #DEFAULT_PREFETCH} (injection seam for the contract test). */
AmqpReplyInbox(Connection connection) {
this(connection, DEFAULT_PREFETCH);
}
/** As above, with an explicit prefetch (injection seam for the contract test). */
AmqpReplyInbox(Connection connection, int prefetch) {
this.connection = connection;
try {
this.channel = connection.createChannel();
// CB-527: bound the held backlog per owned target — must be set before any own()/basicConsume.
this.channel.basicQos(prefetch);
this.publishChannel = connection.createChannel();
this.publishChannel.confirmSelect();
this.publishChannel.addReturnListener(this::onReturn);
this.publishChannel.addConfirmListener(this::onAck, this::onNack);
} catch (IOException e) {
throw new IllegalStateException("cannot open AMQP channel", e);
}
// On automatic recovery the broker redelivers unacked messages with FRESH delivery-tags; the
// tags we were holding are now stale. Drop the held snapshot so the re-attached consumer
// repopulates it with valid tags (dedup by msgId still prevents any double-queue).
// repopulates it with valid tags (dedup by msgId still prevents any double-queue). Any publish
// confirm still in flight when the connection dropped is equally stale — its sequence number
// meant nothing on the old channel and means nothing on the recovered one, so fail it now
// rather than let it silently ride out CONFIRM_TIMEOUT_MS.
if (connection instanceof Recoverable recoverable) {
recoverable.addRecoveryListener(new RecoveryListener() {
@Override
public void handleRecovery(Recoverable recoverable) {
held.clear();
failPendingPublishesOnRecovery();
log.info("AMQP connection recovered; cleared held replies for fresh redelivery");
}
@@ -141,6 +223,12 @@ public final class AmqpReplyInbox implements ReplyInbox, AutoCloseable {
}
}
/**
* Publish {@code content} and block until the broker's publisher confirm for it lands (CB-528).
* Throws {@link IllegalStateException} if the message is returned as unroutable, nacked, or not
* confirmed within {@link #CONFIRM_TIMEOUT_MS} — the caller must treat that as a failed publish,
* not a lost-and-forgotten one.
*/
@Override
public void publish(String target, String msgId, String content) {
AMQP.BasicProperties props = new AMQP.BasicProperties.Builder()
@@ -148,12 +236,35 @@ public final class AmqpReplyInbox implements ReplyInbox, AutoCloseable {
.deliveryMode(2) // persistent — survives a broker restart
.contentType("text/plain")
.build();
try {
synchronized (channelLock) {
channel.basicPublish("", queueName(target), props, content.getBytes(StandardCharsets.UTF_8));
Pending pending = new Pending(msgId);
long seq;
synchronized (publishChannelLock) {
seq = publishChannel.getNextPublishSeqNo();
pendingBySeq.put(seq, pending);
pendingByMsgId.put(msgId, pending);
try {
publishChannel.basicPublish("", queueName(target), true, props,
content.getBytes(StandardCharsets.UTF_8));
} catch (IOException e) {
pendingBySeq.remove(seq, pending);
pendingByMsgId.remove(msgId, pending);
throw new IllegalStateException("cannot publish reply to " + queueName(target), e);
}
} catch (IOException e) {
throw new IllegalStateException("cannot publish reply to " + queueName(target), e);
}
try {
pending.confirmed.get(CONFIRM_TIMEOUT_MS, TimeUnit.MILLISECONDS);
} catch (ExecutionException e) {
Throwable cause = e.getCause();
throw cause instanceof RuntimeException re ? re : new IllegalStateException(cause);
} catch (TimeoutException e) {
throw new IllegalStateException("publish confirm for reply " + msgId + " to " + queueName(target)
+ " timed out after " + CONFIRM_TIMEOUT_MS + "ms — broker may be unreachable or overloaded", e);
} catch (InterruptedException e) {
Thread.currentThread().interrupt();
throw new IllegalStateException("interrupted awaiting publish confirm for " + msgId, e);
} finally {
pendingBySeq.remove(seq, pending);
pendingByMsgId.remove(msgId, pending);
}
}
@@ -222,17 +333,115 @@ public final class AmqpReplyInbox implements ReplyInbox, AutoCloseable {
};
}
/** Broker return for an unroutable {@code mandatory} publish — arrives BEFORE its confirm (CB-528). */
private void onReturn(Return r) {
String msgId = r.getProperties() == null ? null : r.getProperties().getMessageId();
Pending pending = msgId == null ? null : pendingByMsgId.get(msgId);
if (pending != null) {
pending.returned = true;
} else {
log.warn("AMQP return for reply {} (routingKey={}, {} {}) with no matching in-flight publish"
+ " — already resolved by a prior confirm", msgId, r.getRoutingKey(), r.getReplyCode(),
r.getReplyText());
}
}
private void onAck(long seq, boolean multiple) {
resolveConfirm(seq, multiple, true);
}
private void onNack(long seq, boolean multiple) {
resolveConfirm(seq, multiple, false);
}
/**
* Resolve every pending publish covered by this confirm (a single seq, or — {@code multiple} —
* every seq up to and including it). Checks {@link Pending#returned} at confirm time: since the
* broker's return for an unroutable message always precedes its confirm, an ack that arrives after
* a return means "confirmed but never routed", not "durably queued".
*/
private void resolveConfirm(long seq, boolean multiple, boolean ack) {
NavigableMap<Long, Pending> covered = multiple
? pendingBySeq.headMap(seq, true)
: pendingBySeq.subMap(seq, true, seq, true);
for (var it = covered.entrySet().iterator(); it.hasNext(); ) {
Pending pending = it.next().getValue();
it.remove();
pendingByMsgId.remove(pending.msgId, pending);
if (ack && !pending.returned) {
pending.confirmed.complete(null);
} else if (ack) {
pending.confirmed.completeExceptionally(new IllegalStateException(
"reply " + pending.msgId + " was returned as unroutable (queue not declared/owned)"));
} else {
pending.confirmed.completeExceptionally(new IllegalStateException(
"broker nacked publish of reply " + pending.msgId));
}
}
}
/**
* Fail every publish still awaiting its confirm — their sequence numbers are stale after recovery.
* Guarded by {@link #publishChannelLock}, the same lock {@link #publish} holds while it takes its
* sequence number and registers its {@link Pending}: without it, a {@link #publish} that starts
* after the connection has already recovered (so it publishes — and will be confirmed — on the
* <em>new</em> channel) can register between this sweep's iteration and its clear, and this sweep
* then fails a publish that actually succeeded. {@link #publish} only holds the lock for the
* seq/map-put/{@code basicPublish} — it awaits the confirm outside it — so this sweep can only ever
* wait for an in-flight {@code basicPublish} call to return, never for a broker round trip. No
* deadlock.
*
* <p>Package-private (rather than {@code private}) only so the unit test can drive it directly
* against a concurrent {@link #publish} without a live broker reconnect.
*/
void failPendingPublishesOnRecovery() {
synchronized (publishChannelLock) {
for (var it = pendingBySeq.entrySet().iterator(); it.hasNext(); ) {
Pending pending = it.next().getValue();
it.remove();
pendingByMsgId.remove(pending.msgId, pending);
pending.confirmed.completeExceptionally(new IllegalStateException(
"AMQP connection recovered mid-publish; confirm status of reply " + pending.msgId
+ " is unknown"));
}
}
}
/**
* Fail every publish still awaiting its confirm with a clear, immediate error instead of leaving it
* to time out after {@link #CONFIRM_TIMEOUT_MS} once the channels are closed underneath it. Guarded
* by {@link #publishChannelLock} for the same reason as {@link #failPendingPublishesOnRecovery}.
*/
private void failPendingPublishesOnClose() {
synchronized (publishChannelLock) {
for (var it = pendingBySeq.entrySet().iterator(); it.hasNext(); ) {
Pending pending = it.next().getValue();
it.remove();
pendingByMsgId.remove(pending.msgId, pending);
pending.confirmed.completeExceptionally(new IllegalStateException(
"AMQP reply inbox closed while publish of reply " + pending.msgId
+ " was still awaiting its confirm"));
}
}
}
private static String queueName(String target) {
return QUEUE_PREFIX + target + QUEUE_SUFFIX;
}
@Override
public void close() {
failPendingPublishesOnClose();
try {
channel.close();
} catch (Exception e) {
log.debug("AMQP channel close: {}", e.toString());
}
try {
publishChannel.close();
} catch (Exception e) {
log.debug("AMQP publish channel close: {}", e.toString());
}
try {
connection.close();
} catch (Exception e) {
@@ -20,6 +20,7 @@ import java.util.concurrent.TimeUnit;
import java.util.concurrent.TimeoutException;
import java.util.concurrent.atomic.AtomicLong;
import java.util.concurrent.locks.ReentrantLock;
import java.util.function.LongSupplier;
/**
* The blocking delegation feature (CB-104): deliver {@code content} into a worker and block until
@@ -52,8 +53,12 @@ public final class MessageService {
*/
private static final long ASYNC_TIMEOUT_MS = 30 * 60 * 1_000L;
/** How long a finished (terminal) ticket is retained for polling before it is pruned. */
private static final long TICKET_TTL_NANOS = 10 * 60 * 1_000_000_000L;
/**
* How long a finished (terminal) ticket is retained for polling before it is pruned. Package-
* private (not {@code private}) so a test can advance an injected clock past it deterministically
* instead of duplicating the magic number or sleeping for real.
*/
static final long TICKET_TTL_NANOS = 10 * 60 * 1_000_000_000L;
/** Outcome of a blocking send. */
public enum Outcome {
@@ -69,6 +74,14 @@ public final class MessageService {
* failure context (e.g. the error screen). Terminal, but not a successful completion.
*/
WORKER_FAILED,
/**
* The turn finished without a {@code bridge_reply} and the scrape matched the backend's
* configured usage-limit refusal pattern (CB-578 stage A); {@code text} is the reason,
* carrying the matched line. The worker's pane is healthy — only its account is refusing —
* so this is never reported as a completed reply, and is kept distinct from
* {@link #WORKER_FAILED} (a wedged worker) and a session simply going {@code GONE}.
*/
BACKEND_EXHAUSTED,
/**
* The worker paused mid-turn to ask the primary a question (CB-205); {@code text} is the
* question and {@code turnId} correlates the answer. Not terminal — the primary answers with
@@ -151,27 +164,44 @@ public final class MessageService {
/** An in-flight or finished async delegation, keyed by its ticket. */
private static final class Task {
private final String ticket;
private final String target;
private final CompletableFuture<Reply> future = new CompletableFuture<>();
private final long createdNanos = System.nanoTime();
private final long createdNanos;
private volatile Reply question;
private volatile String turnId;
private Task(String target) {
private Task(String ticket, String target, long createdNanos) {
this.ticket = ticket;
this.target = target;
this.createdNanos = createdNanos;
}
}
/**
* A worker session's currently-open {@code bridge_ask} question, surfaced so {@code bridge_status}
* can show it without the caller needing the ticket first (CB-582). Only covers async
* (fire-and-poll) delegations, which track the question on their {@link Task}; a blocking
* ({@code wait:true}) send already hands the question straight back to its own caller, so there is
* nothing hidden left for {@code bridge_status} to surface in that case.
*/
public record PendingAsk(String ticket, String question, String turnId) {
}
private final AgentControl agents;
private final Injector injector;
private final Rendezvous rendezvous;
private final ReplyInbox inbox;
private final ReplyPushLoop pushLoop;
private final Metrics metrics; // CB-502: nullable — no registry in unit tests
// CB-588: injectable so pruneTerminalTickets' 10-minute TICKET_TTL_NANOS can be exercised in a
// test without a real wait — same seam SessionManager already uses for its idle reaper (nowNanos).
private final LongSupplier nowNanos;
private final ConcurrentHashMap<String, ReentrantLock> sessionLocks = new ConcurrentHashMap<>();
private final ConcurrentHashMap<String, Task> tasks = new ConcurrentHashMap<>();
/** The async task that currently owns a target's send lock. */
private final ConcurrentHashMap<String, Task> asyncTasksByTarget = new ConcurrentHashMap<>();
/** Async task that owns each exact forward rendezvous waiter. */
private final ConcurrentHashMap<CompletableFuture<Rendezvous.Resolution>, Task> asyncTasksByWaiter =
new ConcurrentHashMap<>();
/** Async tickets paused on a specific {@code bridge_ask} turn. */
private final ConcurrentHashMap<String, Task> asyncTasksByTurn = new ConcurrentHashMap<>();
private final AtomicLong ticketSeq = new AtomicLong();
@@ -182,7 +212,11 @@ public final class MessageService {
* Create with an explicit {@link ReplyInbox} and optional {@link ReplyPushLoop}.
*
* @param pushLoop nullable — when non-null, the push loop is notified on the no-waiter reply
* branch ({@link #reply}) so it can nudge the primary to drain the inbox
* branch ({@link #reply}) so it can nudge the primary to drain the inbox,
* (CB-588) whenever an async ticket started by {@link #sendAsync} reaches a
* terminal phase, whenever {@link #poll} hands a terminal ticket to its caller,
* and (CB-582) whenever an async ticket's worker pauses mid-turn in
* {@code bridge_ask} or that pause ends (answered or lapsed)
*/
public MessageService(AgentControl agents, Injector injector, Rendezvous rendezvous,
ReplyInbox inbox, ReplyPushLoop pushLoop) {
@@ -197,12 +231,19 @@ public final class MessageService {
*/
public MessageService(AgentControl agents, Injector injector, Rendezvous rendezvous,
ReplyInbox inbox, ReplyPushLoop pushLoop, Metrics metrics) {
this(agents, injector, rendezvous, inbox, pushLoop, metrics, System::nanoTime);
}
/** Test constructor with an injectable clock (CB-588: exercise the ticket-prune TTL without a real wait). */
MessageService(AgentControl agents, Injector injector, Rendezvous rendezvous, ReplyInbox inbox,
ReplyPushLoop pushLoop, Metrics metrics, LongSupplier nowNanos) {
this.agents = agents;
this.injector = injector;
this.rendezvous = rendezvous;
this.inbox = inbox;
this.pushLoop = pushLoop;
this.metrics = metrics;
this.nowNanos = nowNanos;
}
/** Create with an explicit {@link ReplyInbox} and no push loop. */
@@ -279,6 +320,7 @@ public final class MessageService {
case COMPLETED_UNREPLIED -> "completion_fallback";
case TIMED_OUT_WORKING, TIMED_OUT_QUEUED, BUSY -> "timeout";
case WORKER_FAILED -> "failed";
case BACKEND_EXHAUSTED -> "backend_exhausted";
case STALE_TURN, QUESTION -> null; // not a completed delegation
};
}
@@ -300,14 +342,18 @@ public final class MessageService {
*/
public boolean abandon(String target, String reason) {
CompletableFuture<Rendezvous.Resolution> waiter = rendezvous.currentWaiter(target);
if (waiter == null || waiter.isDone()) {
return false; // nobody is blocked on this worker — nothing to abandon
boolean failed = waiter != null && !waiter.isDone() && rendezvous.resolveFailure(waiter, reason);
boolean asyncFailed = false;
for (Task task : tasks.values()) {
if (target.equals(task.target) && task.question == null
&& task.future.complete(new Reply(Outcome.WORKER_FAILED, reason))) {
asyncFailed = true;
}
}
boolean failed = rendezvous.resolveFailure(waiter, reason);
if (failed) {
log.warn("abandoning the blocked send to {}: {}", target, reason);
}
return failed;
return failed || asyncFailed;
}
/**
@@ -319,9 +365,30 @@ public final class MessageService {
}
/**
* Drain (peek + ack) all pending inbox replies for {@code target}. At-least-once: returns the
* messages and acknowledges them; an in-flight failure between returning and the caller
* processing them re-surfaces them on a subsequent drain (the ack is local).
* Drain (peek + ack) all pending inbox replies for {@code target}.
*
* <p><strong>The ack happens here, before the caller has the messages</strong> — before the MCP
* or REST response carrying them has been written, and long before the client has processed
* them. That ordering is what the two adapters disagree about, so do not read this method as
* "at-least-once" without qualifying which inbox is behind it (CB-529):
*
* <ul>
* <li>{@code InMemoryReplyInbox} — the ack only drops an entry from a local map. The messages
* are already in the returned list, so nothing can be lost after this point.
* <li>{@code AmqpReplyInbox} — the ack is a broker-side {@code basicAck}. Once it lands the
* broker has forgotten the message. If the daemon dies while writing the response, the
* reply is gone from the broker <em>and</em> the client never received it. Re-polling
* cannot recover it, because there is nothing left to re-deliver.
* </ul>
*
* <p>So the loss window is the response write, and it is a genuine loss rather than a
* redelivery. This is accepted, not overlooked: the alternative — ack on the next poll — turns
* every normal drain into a double delivery, which costs more than the window it closes. A
* caller that needs certainty re-polls; that is idempotent for every case except this one.
*
* <p>Any change here must be checked against <em>both</em> adapters. The previous version of
* this javadoc claimed "the ack is local", which was true when only the in-memory inbox existed
* and silently became false when the AMQP adapter landed.
*
* @return the drained messages, newest last (FIFO); empty list if none
*/
@@ -354,6 +421,11 @@ public final class MessageService {
* never earned. {@code null} disables the hook.
*/
public Reply send(String target, String content, long timeoutMillis, Runnable onAccepted) {
return send(target, content, timeoutMillis, onAccepted, null);
}
/** Run a send, optionally stopping an async task that teardown already failed before acceptance. */
private Reply send(String target, String content, long timeoutMillis, Runnable onAccepted, Task task) {
long deadlineNanos = System.nanoTime() + timeoutMillis * 1_000_000L;
ReentrantLock lock = sessionLocks.computeIfAbsent(target, _ -> new ReentrantLock());
@@ -361,6 +433,9 @@ public final class MessageService {
return new Reply(Outcome.BUSY, null); // another send held the session the whole window
}
try {
if (task != null && task.future.isDone()) {
return task.future.getNow(null);
}
if (hasAsyncQuestion(target)) {
return new Reply(Outcome.BUSY, null); // the worker's current turn is paused for its lead
}
@@ -372,13 +447,17 @@ public final class MessageService {
// failed send leaves no stale waiter behind.
CompletableFuture<Rendezvous.Resolution> reply = rendezvous.open(target);
try {
if (task != null) {
asyncTasksByWaiter.put(reply, task);
}
TurnToken token = new TurnToken(target, reply);
// The send has won the lock; the accepted-delivery hook records delegator ownership
// here (CB-548). It runs BEFORE enqueue so a throwing hook — onAccepted is now a
// public callback — fails the send without queuing a message that would orphan.
if (onAccepted != null) {
onAccepted.run();
}
CompletableFuture<Void> delivered = injector.enqueue(target, content);
CompletableFuture<Void> delivered = injector.enqueue(target, content, token);
try {
Rendezvous.Resolution r = reply.get(remainingMillis(deadlineNanos), TimeUnit.MILLISECONDS);
return recorded(new Reply(outcomeOf(r.kind()), r.text(), r.turnId()));
@@ -395,6 +474,7 @@ public final class MessageService {
throw new IllegalStateException("interrupted awaiting reply from " + target, e);
}
} finally {
asyncTasksByWaiter.remove(reply);
rendezvous.close(target, reply);
}
} finally {
@@ -420,11 +500,24 @@ public final class MessageService {
if (ticket.fresh()) {
// Register the reverse waiter first, then surface the question — so the answer, which can
// arrive the instant the primary reacts, always finds an open waiter to resolve.
CompletableFuture<Rendezvous.Resolution> waiter = rendezvous.currentWaiter(workerSession);
Task task = markAsyncQuestion(waiter, question, ticket.turnId());
if (!rendezvous.resolveQuestion(workerSession, question, ticket.turnId())) {
if (task != null) {
clearAsyncQuestion(ticket.turnId(), true);
}
rendezvous.closeAsk(ticket.turnId());
return new AskResult(AskOutcome.NO_WAITER, null); // no primary is blocked on this worker
}
markAsyncQuestion(workerSession, question, ticket.turnId());
// CB-582: the question just became visible via bridge_poll (Phase.ASKING) for an async
// (wait:false) delegation — nudge the lead's own pane the same way a terminal ticket does
// (CB-588), since the lead's normal poll cadence is minutes away and the reverse-rendezvous
// window (~55s, see BridgeMcp/BridgedApp) is far shorter. A blocking (wait:true) send has
// no Task and gets the question directly in its own reply, so task == null there — nothing
// to nudge.
if (task != null && pushLoop != null) {
pushLoop.onQuestionOpened(task.ticket, workerSession, ticket.turnId(), question);
}
}
try {
String answer = ticket.answer().get(timeoutMillis, TimeUnit.MILLISECONDS);
@@ -443,6 +536,14 @@ public final class MessageService {
// Only the fresh owner tears down the shared turn; a duplicate must leave it open.
if (ticket.fresh()) {
rendezvous.closeAsk(ticket.turnId());
// CB-582: tear the push loop's copy down at the same point, not only on the three
// paths that call clearAsyncQuestion. The answer future can complete exceptionally
// (ExecutionException) or the thread be interrupted, and both leave this method by
// throwing — the question would stay pending forever, keep being named in nudges
// until its own cap, and never be removed from the map. Already-closed is a no-op.
if (pushLoop != null) {
pushLoop.questionClosed(ticket.turnId());
}
}
}
}
@@ -519,25 +620,35 @@ public final class MessageService {
*/
public String sendAsync(String target, String content, Runnable onAccepted) {
String ticket = "task-" + ticketSeq.incrementAndGet();
Task task = new Task(target);
Task task = new Task(ticket, target, nowNanos.getAsLong());
tasks.put(ticket, task);
if (pushLoop != null) {
// CB-588: task.future only ever completes on a terminal phase (DONE or a failure) — a
// worker paused in bridge_ask leaves it running, per finishAsyncTask's own contract — so
// this fires exactly once, from whichever path completes it: finishAsyncTask(task, result)
// below on any non-QUESTION outcome of send() — a worker's bridge_reply, the CB-106
// completion fallback, a CB-109 wedge, TIMED_OUT, BUSY, or BACKEND_EXHAUSTED — the same
// finishAsyncTask reached via answer()'s finishAsyncTask(turnId, result) once a QUESTION
// is resolved, completeExceptionally(t) just below when send() itself throws, or a CB-516
// abandon() on teardown. Without this, MessageService.reply's rendezvous fast path (the
// one an async ticket always takes) never told the push loop anything happened — see the
// class javadoc on sendAsync/CB-107.
task.future.whenComplete((reply, ex) -> {
boolean failed = ex != null || reply == null || !reply.completed();
pushLoop.onTicketTerminal(ticket, target, failed);
});
}
asyncExecutor.submit(() -> {
try {
Runnable trackingAccepted = () -> {
if (onAccepted != null) {
onAccepted.run();
}
asyncTasksByTarget.put(target, task);
};
Reply result = send(target, content, ASYNC_TIMEOUT_MS, trackingAccepted);
Reply result = send(target, content, ASYNC_TIMEOUT_MS, onAccepted, task);
if (result.outcome() == Outcome.QUESTION) {
asyncTasksByTarget.remove(target, task);
// Keep the accepted owner until answer() finishes it. markAsyncQuestion may run
// just after resolveQuestion wakes this thread.
} else {
finishAsyncTask(task, result);
}
} catch (Throwable t) {
task.future.completeExceptionally(t);
asyncTasksByTarget.remove(target, task);
}
});
pruneTerminalTickets();
@@ -564,6 +675,11 @@ public final class MessageService {
}
return new TaskView(ticket, Phase.PENDING, null, null, "worker " + liveStatus(task.target), null);
}
// CB-588: the ticket is terminal and being handed to the caller right here — tell the push
// loop it is collected so a later tick's nudge never names a ticket the lead already has.
if (pushLoop != null) {
pushLoop.ticketCollected(ticket);
}
Reply r;
try {
r = f.getNow(null);
@@ -575,9 +691,11 @@ public final class MessageService {
String source = r.outcome() == Outcome.REPLIED ? "reply" : "transcript";
return new TaskView(ticket, Phase.DONE, r.text(), source, null, null);
}
// A wedged worker (CB-109) carries the error context as its reason; the timeout/busy
// outcomes carry none, so fall back to the outcome name.
String detail = r.outcome() == Outcome.WORKER_FAILED && r.text() != null
// A wedged worker (CB-109) or a backend-exhausted classification (CB-578 stage A) carries
// the real cause as its reason; the timeout/busy outcomes carry none, so fall back to the
// outcome name.
boolean carriesReason = r.outcome() == Outcome.WORKER_FAILED || r.outcome() == Outcome.BACKEND_EXHAUSTED;
String detail = carriesReason && r.text() != null
? r.text()
: "no reply — " + r.outcome().name().toLowerCase();
return new TaskView(ticket, Phase.FAILED, null, null, detail, null);
@@ -592,24 +710,48 @@ public final class MessageService {
}
}
/** Drop finished tickets older than the TTL so the registry cannot grow without bound. */
/**
* Drop finished tickets older than the TTL so {@link #tasks} cannot grow without bound.
*
* <p>{@code tasks} is the sole authority on whether a ticket still exists — {@link #poll} returns
* {@code null} the instant a ticket is gone from here, before it ever reaches the terminal branch
* that calls {@link ReplyPushLoop#ticketCollected}. Without telling the push loop about a prune
* too, its own {@code pendingTickets} entry would outlive the ticket it names: an unpolled ticket
* (or one the reminder cap already gave up on) is pruned here but never collected there, so it
* lingers in {@code pendingTickets} forever and rides along on every later nudge to the same lead
* — naming a ticket {@code bridge_poll} can no longer find (CB-588 follow-up).
*/
private void pruneTerminalTickets() {
long cutoff = System.nanoTime() - TICKET_TTL_NANOS;
tasks.values().removeIf(t -> t.future.isDone() && t.createdNanos < cutoff);
long cutoff = nowNanos.getAsLong() - TICKET_TTL_NANOS;
tasks.entrySet().removeIf(e -> {
Task t = e.getValue();
boolean expired = t.future.isDone() && t.createdNanos < cutoff;
if (expired && pushLoop != null) {
pushLoop.ticketCollected(e.getKey());
}
return expired;
});
}
/** Record the active question for an async ticket; blocking sends have no entry and stay unchanged. */
private void markAsyncQuestion(String target, String text, String turnId) {
Task task = asyncTasksByTarget.get(target);
private Task markAsyncQuestion(CompletableFuture<Rendezvous.Resolution> waiter, String text, String turnId) {
Task task = waiter == null ? null : asyncTasksByWaiter.get(waiter);
if (task != null) {
task.question = new Reply(Outcome.QUESTION, text, turnId);
task.turnId = turnId;
asyncTasksByTurn.put(turnId, task);
}
return task;
}
/** Clear an answered or lapsed question, but only when it matches the ticket's current turn. */
private void clearAsyncQuestion(String turnId, boolean forgetTurn) {
// CB-582: tell the push loop first — like ticketCollected, a removal for a turnId it never
// nudged about (or already dropped) is a harmless no-op, so this is safe to call unconditionally
// rather than threading the guard below through it.
if (pushLoop != null) {
pushLoop.questionClosed(turnId);
}
Task task = asyncTasksByTurn.get(turnId);
if (task != null && turnId.equals(task.turnId)) {
task.question = null;
@@ -623,7 +765,6 @@ public final class MessageService {
/** Complete and detach an async ticket after its worker's actual terminal reply. */
private void finishAsyncTask(Task task, Reply result) {
task.future.complete(result);
asyncTasksByTarget.remove(task.target, task);
if (task.turnId != null) {
asyncTasksByTurn.remove(task.turnId, task);
}
@@ -642,6 +783,23 @@ public final class MessageService {
return asyncTasksByTurn.values().stream().anyMatch(task -> target.equals(task.target));
}
/**
* The question {@code workerSession} is currently paused on via {@code bridge_ask}, if any
* (CB-582) — {@code bridge_status} uses this to show a pending question without the caller
* needing the ticket. {@code null} when the session has no open async question (including a
* session mid a <em>blocking</em> {@code bridge_ask}, which has no {@link Task} to look up — see
* {@link PendingAsk}).
*/
public PendingAsk pendingAsk(String workerSession) {
for (Task task : tasks.values()) {
Reply q = task.question;
if (q != null && workerSession.equals(task.target)) {
return new PendingAsk(task.ticket, q.text(), q.turnId());
}
}
return null;
}
/** Release the async executor. */
public void close() {
asyncExecutor.shutdown();
@@ -653,6 +811,7 @@ public final class MessageService {
case REPLY -> Outcome.REPLIED;
case COMPLETION -> Outcome.COMPLETED_UNREPLIED;
case FAILED -> Outcome.WORKER_FAILED;
case BACKEND_EXHAUSTED -> Outcome.BACKEND_EXHAUSTED;
case QUESTION -> Outcome.QUESTION;
};
}
@@ -33,6 +33,13 @@ public final class Rendezvous {
COMPLETION,
/** The worker ran the turn then wedged (CB-109); {@code text} is the failure context. */
FAILED,
/**
* The turn finished without a {@code bridge_reply}, and the scrape matched the backend's
* configured usage-limit refusal pattern (CB-578 stage A); {@code text} is the reason,
* carrying the matched line. The pane is healthy — only the account is refusing — so this
* is kept separate from a session simply going {@code GONE}.
*/
BACKEND_EXHAUSTED,
/**
* The worker paused mid-turn to ask the primary a question (CB-205 reverse rendezvous);
* {@code text} is the question and {@code turnId} correlates the primary's answer back to
@@ -224,6 +231,19 @@ public final class Rendezvous {
return waiter != null && waiter.complete(new Resolution(Kind.FAILED, reason));
}
/**
* Resolve a specific captured {@code waiter} as {@link Kind#BACKEND_EXHAUSTED} (CB-578 stage A):
* the turn finished with no {@code bridge_reply} and the scrape matched the backend's configured
* usage-limit pattern; {@code reason} carries the matched line. Like
* {@link #resolveCompletion(CompletableFuture, String)} it targets the exact captured send
* (CB-116). A no-op if that waiter was already resolved — first resolution wins.
*
* @return {@code true} if this call resolved the waiter, {@code false} if it was null or already resolved
*/
public boolean resolveExhausted(CompletableFuture<Resolution> waiter, String reason) {
return waiter != null && waiter.complete(new Resolution(Kind.BACKEND_EXHAUSTED, reason));
}
private boolean complete(String session, Resolution resolution) {
CompletableFuture<Resolution> waiter = waiters.get(session);
return waiter != null && waiter.complete(resolution);
@@ -8,26 +8,66 @@ import dev.ltms.bridged.metrics.Metrics;
import org.slf4j.Logger;
import org.slf4j.LoggerFactory;
import java.util.ArrayList;
import java.util.HashSet;
import java.util.List;
import java.util.Set;
import java.util.concurrent.ConcurrentHashMap;
import java.util.concurrent.ScheduledExecutorService;
import java.util.concurrent.TimeUnit;
import java.util.stream.Collectors;
/**
* Mechanism (b) of CB-307: a dedicated, status-gated push loop that nudges the primary's own
* herdr pane when a worker reply lands with no live {@code bridge_send} to resolve it.
* A status-gated push loop that nudges a lead's own herdr pane when it has uncollected work
* waiting: a worker reply queued with no live {@code bridge_send} to resolve it (CB-307), an
* async delegation ticket ({@code bridge_send(wait:false)}) that reached a terminal phase
* (CB-588), or an async ticket's worker pausing mid-turn in {@code bridge_ask} to await an answer
* (CB-582).
*
* <p>The loop is triggered by {@link #onReplyQueued(String)} (called from
* {@link MessageService#reply} after the durable inbox publish). It checks four conditions
* at each tick via {@link #decide(String, int)}, then either injects a drain nudge,
* waits for the primary to become injectable, or stops reminding.
* <p><strong>CB-590: one schedule per lead.</strong> All three kinds of work are triggered
* through their own entry point — {@link #onReplyQueued(String)},
* {@link #onTicketTerminal(String, String, boolean)}, and
* {@link #onQuestionOpened(String, String, String, String)} — but each resolves the lead that
* should be nudged and coalesces onto a single per-lead reminder schedule, tracked in
* {@link #activeLeads}. Earlier this was two independent schedules (one keyed by worker target
* for replies, one keyed by lead for tickets) that could both decide to inject into the same pane
* in the same window — a race, not routine behaviour, but the expensive kind: it interrupts the
* lead's live turn twice. Collapsing to one schedule per lead makes that structurally impossible:
* at most one scheduled tick chain is ever live for a given lead (guarded by {@link #activeLeads}'
* compare-and-set), so at most one {@code agents.send} to that lead's pane is ever in flight.
*
* <p>Bounded: at most {@link #maxReminders} nudges per target, with a configurable backoff
* between them. The reply is never lost — the durable inbox is the backstop.
* <p>Each tick examines <em>everything</em> pending for that lead — reply targets whose inbox
* still holds an unacked message ({@link #pendingReplies}), tickets not yet collected
* ({@link #pendingTickets}), and open questions not yet answered or lapsed
* ({@link #pendingQuestions}) — and sends at most one combined nudge per tick
* ({@link #injectNudge(String, int, int, int)}). Work that arrives while the lead is busy is
* never lost: it is re-read fresh on every tick until the lead is injectable or its own reminder
* cap ({@link #maxReminders}) is reached — each source spends from its own budget, so one source
* exhausting its cap does not stop nudges about the others (post-CB-590 regression fix; see
* {@link #decide}) — whichever the durable inbox / pending set doesn't already answer via
* {@code STOP}.
*/
public final class ReplyPushLoop {
private static final Logger log = LoggerFactory.getLogger(ReplyPushLoop.class);
static final String NUDGE_FORMAT = "Worker %s returned a reply — run bridge_poll(target=%s) to collect it";
/** Coalesced form, several uncollected replies for the same lead. */
static final String REPLIES_NUDGE_FORMAT =
"%d workers returned replies — run bridge_poll(target=...) for each to collect them: %s";
/** Singular form, one uncollected ticket. */
static final String TICKET_NUDGE_FORMAT =
"Ticket %s finished%s — run bridge_poll(ticket=%s) to collect it";
/** Coalesced form, several uncollected tickets for the same lead. */
static final String TICKETS_NUDGE_FORMAT =
"%d tickets finished%s — run bridge_poll(ticket=...) for each to collect them: %s";
/** Singular form, one worker paused mid-turn in bridge_ask (CB-582) — names the answer call directly. */
static final String QUESTION_NUDGE_FORMAT =
"Worker %s asked a question (ticket %s) — answer it with bridge_send(turnId=\"%s\", "
+ "content=...) to resume its turn:\n%s";
/** Coalesced form, several open questions for the same lead. */
static final String QUESTIONS_NUDGE_FORMAT =
"%d workers are paused on a question — run bridge_poll(ticket=...) for each, then answer "
+ "with bridge_send(turnId=..., content=...): %s";
private final PrimaryRegistry primaryRegistry;
private final AgentControl agents;
@@ -37,8 +77,24 @@ public final class ReplyPushLoop {
private final long backoffMs;
private final Metrics metrics; // CB-512: nullable — no registry in unit tests
/** Track targets that have an active schedule. */
private final ConcurrentHashMap<String, Boolean> activeTargets = new ConcurrentHashMap<>();
/**
* Worker targets with a reply queued, keyed by target. Each entry carries its own nudge
* count (CB-598) rather than sharing one counter per lead per source: a target's count only
* ever reflects nudges that actually named that target, so a target that joins while the
* schedule is already deep into another target's reminders still reads as fresh.
*/
private final ConcurrentHashMap<String, ReplyEntry> pendingReplies = new ConcurrentHashMap<>();
/** Tickets that have gone terminal but not yet been polled, keyed by ticket. */
private final ConcurrentHashMap<String, PendingTicket> pendingTickets = new ConcurrentHashMap<>();
/**
* Open {@code bridge_ask} questions not yet answered or lapsed, keyed by {@code turnId}
* (CB-582). A question's own nudge count is tracked the same per-item way as
* {@link #pendingTickets} (CB-598): a fresh question keeps its source eligible regardless of
* how depleted an older, still-open question's count is.
*/
private final ConcurrentHashMap<String, PendingQuestion> pendingQuestions = new ConcurrentHashMap<>();
/** CB-590: leads with an active combined reminder schedule (replies and/or tickets and/or questions). */
private final ConcurrentHashMap<String, Boolean> activeLeads = new ConcurrentHashMap<>();
public ReplyPushLoop(PrimaryRegistry primaryRegistry, AgentControl agents, ReplyInbox inbox,
ScheduledExecutorService scheduler,
@@ -66,130 +122,511 @@ public final class ReplyPushLoop {
}
}
// --- decision logic (package-private for unit-testing) -------------------------------------
// --- pending-work lookups (package-private for unit-testing) -------------------------------
/** The action the loop should take for a target at the given reminder count. */
/** The action the loop should take for a lead at the given reminder count. */
enum Action { INJECT, WAIT_BUSY, STOP }
/**
* Pure decision function: examine the current state and return what the loop should do.
* Reply targets still pending for {@code lead} — registered via {@link #onReplyQueued} and
* whose inbox still holds an unacked message. A target whose inbox has since drained (acked,
* or collected via a live {@code bridge_send} rendezvous instead) is dropped from
* {@link #pendingReplies} here rather than lingering forever; there is no explicit "reply
* collected" callback the way {@link #ticketCollected} exists for tickets, so the inbox itself
* is the only signal.
*/
private Set<String> pendingReplyTargetsFor(String lead) {
Set<String> result = new HashSet<>();
for (var entry : pendingReplies.entrySet()) {
String target = entry.getKey();
ReplyEntry owning = entry.getValue();
if (!lead.equals(owning.lead())) continue;
if (inbox.peek(target).isEmpty()) {
pendingReplies.remove(target, owning);
continue;
}
result.add(target);
}
return result;
}
/** A pending reply target: which lead to nudge, and how many nudges have named it so far. */
private record ReplyEntry(String lead, int nudgeCount) {
}
/**
* A ticket awaiting collection: which lead to nudge, whether it ended in failure, and how
* many nudges have named it so far (CB-598 — tracked per ticket, not per lead per source).
*/
private record PendingTicket(String ticket, String lead, boolean failed, int nudgeCount) {
}
/** Tickets still pending for {@code lead}, snapshotted fresh for one tick. */
private List<PendingTicket> pendingTicketsFor(String lead) {
return pendingTickets.values().stream().filter(t -> lead.equals(t.lead())).toList();
}
/** Ticket ids still pending for {@code lead} — a plain snapshot for race comparison. */
private Set<String> pendingTicketIdsFor(String lead) {
return pendingTicketsFor(lead).stream().map(PendingTicket::ticket)
.collect(Collectors.toUnmodifiableSet());
}
/**
* An open question awaiting the lead's answer: which ticket it belongs to, which worker asked,
* which lead to nudge, the question text, and how many nudges have named it so far (CB-598 —
* tracked per question, not per lead per source).
*/
private record PendingQuestion(String turnId, String ticket, String target, String lead,
String question, int nudgeCount) {
}
/** Questions still open for {@code lead}, snapshotted fresh for one tick. */
private List<PendingQuestion> pendingQuestionsFor(String lead) {
return pendingQuestions.values().stream().filter(q -> lead.equals(q.lead())).toList();
}
/** Question turnIds still open for {@code lead} — a plain snapshot for race comparison. */
private Set<String> pendingQuestionTurnIdsFor(String lead) {
return pendingQuestionsFor(lead).stream().map(PendingQuestion::turnId)
.collect(Collectors.toUnmodifiableSet());
}
/**
* The reply-source reminder count {@link #decide} should see for {@code lead} on this tick:
* the <em>minimum</em> nudge count among the reply targets currently pending for it (CB-598).
*
* @param target the worker session (target terminal id)
* @param reminderCount how many nudges have been sent so far for this target
* <p>Before this, the count passed to {@code decide} was a single counter carried forward
* across scheduled ticks ({@code scheduleNext(lead, count + 1, ...)}), incremented whenever
* the source had <em>any</em> pending work — not tied to which target that work was. A target
* that joined while an older target's count was already near the cap inherited that count on
* its very next tick, even though no nudge had ever named it. Taking the minimum over what is
* actually pending now means a fresh target (count 0) keeps the source eligible regardless of
* how many times an older, still-undrained target has already been nudged; that older target
* keeps riding along in the combined nudge text without spending any more of its own budget
* (see {@link #bumpNudgeCounts}). Returns 0 when nothing is pending — {@link #decide} never
* consults the count in that case, since {@code hasReplyWork} is false.
*/
private int minReplyNudgeCountFor(String lead) {
int min = Integer.MAX_VALUE;
for (String target : pendingReplyTargetsFor(lead)) {
ReplyEntry entry = pendingReplies.get(target);
if (entry != null) {
min = Math.min(min, entry.nudgeCount());
}
}
return min == Integer.MAX_VALUE ? 0 : min;
}
/** As {@link #minReplyNudgeCountFor}, for the ticket source. */
private int minTicketNudgeCountFor(String lead) {
int min = Integer.MAX_VALUE;
for (PendingTicket ticket : pendingTicketsFor(lead)) {
min = Math.min(min, ticket.nudgeCount());
}
return min == Integer.MAX_VALUE ? 0 : min;
}
/** As {@link #minReplyNudgeCountFor}, for the question source (CB-582). */
private int minQuestionNudgeCountFor(String lead) {
int min = Integer.MAX_VALUE;
for (PendingQuestion q : pendingQuestionsFor(lead)) {
min = Math.min(min, q.nudgeCount());
}
return min == Integer.MAX_VALUE ? 0 : min;
}
/**
* Pure decision function: examine everything pending for {@code lead} — reply targets and
* tickets alike — and return what the loop should do.
*
* <p><strong>CB-590-fix: one schedule, two budgets.</strong> The single per-lead schedule
* (CB-590) still ticks once for both sources, but each source is capped independently —
* {@code replyReminderCount} against a reply target still pending, {@code ticketReminderCount}
* against a ticket still pending. A busy reply stream that exhausts its own cap must not stop
* the loop from nudging about a ticket that still has budget left, and vice versa: either
* source being eligible (has pending work AND is under its own cap) is enough for
* {@link Action#INJECT}. Only when neither source has eligible work does the loop
* {@link Action#STOP}.
*
* <p><strong>CB-598: the counts are per-item, not per-tick.</strong> {@link #tick} no longer
* carries these counts forward across scheduled calls — it recomputes them fresh every tick via
* {@link #minReplyNudgeCountFor} / {@link #minTicketNudgeCountFor}, so this function itself did
* not need to change; only what its caller feeds it did.
*
* @param lead the lead terminal to nudge
* @param replyReminderCount the lowest nudge count among reply targets pending for this lead
* @param ticketReminderCount the lowest nudge count among tickets pending for this lead
* @return the action the caller should take
*/
Action decide(String target, int reminderCount) {
// CB-532: the destination is per-delegation — the lead that sent this worker its work, not
// "the primary". With two leads orchestrating one fleet the singular question has no right
// answer, and answering it anyway interrupted whichever lead happened to call bridge_send
// first with results it never asked for.
var nudgeTarget = primaryRegistry.nudgeTargetFor(target);
if (nudgeTarget.isEmpty()) {
log.debug("push: no lead is known to be waiting on {}, stopping reminder", target);
Action decide(String lead, int replyReminderCount, int ticketReminderCount) {
return decide(lead, replyReminderCount, ticketReminderCount, minQuestionNudgeCountFor(lead));
}
/**
* As {@link #decide(String, int, int)}, with the question source (CB-582) folded in on the
* same footing as replies and tickets: its own eligibility (has open questions AND under its
* own {@link #maxReminders} budget) is enough on its own to {@link Action#INJECT}, exactly like
* the other two.
*
* @param questionReminderCount the lowest nudge count among questions open for this lead
*/
Action decide(String lead, int replyReminderCount, int ticketReminderCount, int questionReminderCount) {
boolean hasReplyWork = !pendingReplyTargetsFor(lead).isEmpty();
boolean hasTicketWork = !pendingTicketIdsFor(lead).isEmpty();
boolean hasQuestionWork = !pendingQuestionTurnIdsFor(lead).isEmpty();
if (!hasReplyWork && !hasTicketWork && !hasQuestionWork) {
log.debug("push: nothing pending for lead {}, stopping reminder", lead);
return Action.STOP;
}
if (inbox.peek(target).isEmpty()) {
log.debug("push: inbox empty for {}, stopping reminder", target);
return Action.STOP;
}
if (reminderCount >= maxReminders) {
log.debug("push: reminder cap ({}) reached for {}, stopping", maxReminders, target);
boolean replyEligible = hasReplyWork && replyReminderCount < maxReminders;
boolean ticketEligible = hasTicketWork && ticketReminderCount < maxReminders;
boolean questionEligible = hasQuestionWork && questionReminderCount < maxReminders;
if (!replyEligible && !ticketEligible && !questionEligible) {
log.debug("push: reminder cap ({}) reached for lead {} on every source with pending work, stopping",
maxReminders, lead);
countNudge("exhausted");
return Action.STOP;
}
String leadTerminal = nudgeTarget.get();
AgentStatus status;
try {
status = agents.status(leadTerminal);
status = agents.status(lead);
} catch (RuntimeException e) {
log.debug("push: status check failed for lead {}, will retry", leadTerminal, e);
log.debug("push: status check failed for lead {}, will retry", lead, e);
return Action.WAIT_BUSY;
}
if (status.injectable()) {
return Action.INJECT;
}
log.debug("push: lead {} is {} (not injectable), waiting", leadTerminal, status);
log.debug("push: lead {} is {} (not injectable), waiting", lead, status);
return Action.WAIT_BUSY;
}
// --- public entrypoint ---------------------------------------------------------------------
// --- public entrypoints ----------------------------------------------------------------------
/**
* Called when a reply is queued for {@code target}. Idempotent per target: a second call while
* a schedule is active is a no-op. The schedule nudges the primary, then schedules a follow-up
* check (reminder on backoff, or re-check on WAIT_BUSY), until the inbox is empty or the cap
* is reached.
* Called when a reply is queued for {@code target}. Resolves the lead delegating to
* {@code target} (CB-532) and coalesces onto that lead's single reminder schedule — starting
* one if none is active, joining an already-active one otherwise. A no-op if no lead is known
* to be waiting on {@code target}: there is nobody to nudge yet, and the durable inbox is the
* backstop until a lead is recorded.
*/
public void onReplyQueued(String target) {
if (activeTargets.putIfAbsent(target, Boolean.TRUE) != null) {
log.debug("push: already active for {}, ignoring duplicate trigger", target);
return; // already scheduled
}
log.debug("push: starting reminder loop for {}", target);
scheduleNext(target, 0);
}
/** Execute one loop tick — called on the scheduler thread. */
private void tick(String target, int reminderCount) {
var action = decide(target, reminderCount);
switch (action) {
case INJECT -> {
injectNudge(target, reminderCount);
scheduleNext(target, reminderCount + 1);
}
// Re-check after the configured backoff; the primary may become injectable soon.
case WAIT_BUSY -> scheduleNext(target, reminderCount);
case STOP -> {
activeTargets.remove(target);
log.debug("push: reminder loop ended for {}", target);
}
}
}
/** Send the nudge and log the event. */
private void injectNudge(String target, int reminderCount) {
// Re-read rather than threading it down from decide(): the delegating lead can change
// between the decision and the injection, and the nudge should follow the current one.
var lead = primaryRegistry.nudgeTargetFor(target);
if (lead.isEmpty()) {
log.debug("push: lead for {} disappeared before the nudge could be sent", target);
log.debug("push: no lead is known to be waiting on {}, skipping reminder", target);
return;
}
String leadTerminal = lead.get();
String nudge = NUDGE_FORMAT.formatted(target, target);
pendingReplies.compute(target, (t, existing) ->
new ReplyEntry(lead.get(), existing == null ? 0 : existing.nudgeCount()));
startOrCoalesce(lead.get());
}
/**
* Called when an async delegation ticket ({@code bridge_send(wait:false)}, CB-107) reaches a
* terminal phase — DONE or a failure. Unlike {@link #onReplyQueued}, which nudges about the
* durable-inbox no-waiter path, this covers the path {@code MessageService.reply} takes when a
* fire-and-poll send's own rendezvous waiter resolves the reply directly: that path returns
* before {@link #onReplyQueued} is ever called, so without this entry point a ticket finishing
* that way never nudged anyone (CB-588 / gitea #72).
*
* <p>Resolves the delegating lead the same way {@link #onReplyQueued} does and coalesces onto
* the same per-lead schedule (CB-590) — several tickets, or a ticket and a reply, finishing
* for the same lead while its schedule is already active all ride the existing schedule's next
* tick rather than firing a nudge each.
*
* @param ticket the ticket to nudge about
* @param target the worker session the ticket was sent to — resolves which lead delegated it
* @param failed whether the ticket ended in a failure phase rather than {@code DONE}
*/
public void onTicketTerminal(String ticket, String target, boolean failed) {
var lead = primaryRegistry.nudgeTargetFor(target);
if (lead.isEmpty()) {
log.debug("push: no lead is known to be waiting on ticket {} (target {}), skipping nudge",
ticket, target);
return;
}
pendingTickets.compute(ticket, (id, existing) ->
new PendingTicket(ticket, lead.get(), failed, existing == null ? 0 : existing.nudgeCount()));
startOrCoalesce(lead.get());
}
/**
* Called when a ticket's terminal state has been collected via {@code bridge_poll}. Removes it
* from the pending set so a scheduled tick — and any nudge it sends — never names a ticket the
* lead already has. A ticket that was never pending (unknown ticket, or one nudged with no push
* loop configured) is a no-op.
*/
public void ticketCollected(String ticket) {
pendingTickets.remove(ticket);
}
/**
* Called when an async ticket's worker pauses mid-turn in {@code bridge_ask} (CB-582): the
* question is now visible via {@code bridge_poll} (Phase.ASKING), but the reverse-rendezvous
* window it opened with (~55s default, see {@code BridgeMcp}/{@code BridgedApp}) is far shorter
* than a lead's normal minutes-long poll cadence — exactly the gap this closes. Resolves the
* delegating lead the same way {@link #onTicketTerminal} does and coalesces onto the same
* per-lead schedule (CB-590).
*
* @param ticket the async ticket the question belongs to (for {@code bridge_poll})
* @param target the worker session that asked
* @param turnId correlation id the lead answers with ({@code bridge_send turnId=...})
* @param question the question text
*/
public void onQuestionOpened(String ticket, String target, String turnId, String question) {
var lead = primaryRegistry.nudgeTargetFor(target);
if (lead.isEmpty()) {
log.debug("push: no lead is known to be waiting on {}'s question (turnId {}), skipping nudge",
target, turnId);
return;
}
pendingQuestions.put(turnId, new PendingQuestion(turnId, ticket, target, lead.get(), question, 0));
startOrCoalesce(lead.get());
}
/**
* Called when a worker's {@code bridge_ask} resolves — answered or lapsed unanswered — so a
* scheduled tick never nudges about a question the lead already handled. A {@code turnId} that
* was never pending (never nudged, or already closed) is a no-op.
*/
public void questionClosed(String turnId) {
pendingQuestions.remove(turnId);
}
// --- the schedule ----------------------------------------------------------------------------
/** Start a reminder schedule for {@code lead}, or join the one already running. */
private void startOrCoalesce(String lead) {
if (activeLeads.putIfAbsent(lead, Boolean.TRUE) != null) {
log.debug("push: reminder loop already active for lead {}, work coalesced in", lead);
return;
}
log.debug("push: starting reminder loop for lead {}", lead);
scheduleNext(lead);
}
/**
* Execute one loop tick — called on the scheduler thread (or directly by a test; package-private
* for the same reason as {@link #stopOrRestart}).
*
* <p><strong>CB-598.</strong> The reminder counts fed into {@link #decide} are recomputed fresh
* every tick from what is actually pending right now ({@link #minReplyNudgeCountFor} /
* {@link #minTicketNudgeCountFor}), rather than carried forward as running counters across
* scheduled calls. A counter carried forward has no memory of which item it was counting for:
* a target or ticket that joined mid-backoff — after the previous tick fired but before this one
* did — is already sitting in {@code repliesBefore} / {@code ticketsBefore} below by the time this
* tick takes its snapshot, indistinguishable at that point from backlog the cap is meant to
* silence. Recomputing from the per-item counts fixes that: a newly-joined item's own count is
* still 0, so it keeps its source eligible regardless of how depleted an older, still-undrained
* item's count is.
*/
void tick(String lead) {
Set<String> repliesBefore = pendingReplyTargetsFor(lead);
Set<String> ticketsBefore = pendingTicketIdsFor(lead);
Set<String> questionsBefore = pendingQuestionTurnIdsFor(lead);
int replyReminderCount = minReplyNudgeCountFor(lead);
int ticketReminderCount = minTicketNudgeCountFor(lead);
int questionReminderCount = minQuestionNudgeCountFor(lead);
var action = decide(lead, replyReminderCount, ticketReminderCount, questionReminderCount);
switch (action) {
case INJECT -> {
injectNudge(lead, replyReminderCount, ticketReminderCount, questionReminderCount);
scheduleNext(lead);
}
// Re-check after the configured backoff; the lead may become injectable soon.
case WAIT_BUSY -> scheduleNext(lead);
case STOP -> stopOrRestart(lead, repliesBefore, ticketsBefore, questionsBefore);
}
}
/**
* Release {@code lead}'s active-schedule slot, then restart it only if work landed that
* {@code repliesBefore} / {@code ticketsBefore} — the snapshots taken just before this tick's
* decision — did not already account for. {@link #onReplyQueued} / {@link #onTicketTerminal}
* read {@link #activeLeads} to decide whether to coalesce onto an existing schedule or start
* one, so work that lands between {@link #decide} returning {@link Action#STOP} and this
* removal running sees the (soon-to-be-stale) slot as occupied, coalesces onto a schedule that
* is about to die, and gets no nudge scheduled at all — a lost nudge, exactly what CB-588 (and
* now CB-590) exist to remove (originally found in review, gitea PR #73, for the ticket-only
* loop; carried forward here for the unified one).
*
* <p>Restarting on ANY non-empty pending set would be wrong: when STOP is reached because the
* reminder cap was hit rather than the backlog draining, the same never-collected work is
* expected to still be sitting there — that is the cap doing its job — and restarting would
* nudge about it forever, defeating the bound. Diffing the current pending sets against the
* "before" snapshots tells the two cases apart: an item present before this tick's decision is
* stale backlog, not a race; only an item absent from the "before" snapshot can only have
* arrived during the decision-to-release window, which is exactly the race this method closes.
*
* <p>Package-private so a test can drive the interleaving directly rather than trying to force a
* genuine thread race.
*
* <p>Terminates rather than spinning: this method restarts the schedule at most once per call,
* and a fresh {@link #onReplyQueued} / {@link #onTicketTerminal} racing the recheck below still
* terminates in one of two ways — either it observes the slot already vacated (by the
* {@code activeLeads.remove} above, which happens-before this recheck in program order) and
* claims it itself, or it lands first and this recheck then observes its work in
* {@link #pendingReplies} / {@link #pendingTickets} and reclaims the slot instead. Exactly one
* side always wins; neither can miss the other, so this never loops on its own account.
*/
void stopOrRestart(String lead, Set<String> repliesBefore, Set<String> ticketsBefore) {
stopOrRestart(lead, repliesBefore, ticketsBefore, pendingQuestionTurnIdsFor(lead));
}
/**
* As {@link #stopOrRestart(String, Set, Set)}, with the question source's (CB-582) own "before"
* snapshot folded into the same race check: a question that raced in during the
* decision-to-release window reclaims the schedule slot exactly like a raced-in reply or ticket.
*/
void stopOrRestart(String lead, Set<String> repliesBefore, Set<String> ticketsBefore,
Set<String> questionsBefore) {
activeLeads.remove(lead);
boolean racedIn = pendingReplyTargetsFor(lead).stream().anyMatch(t -> !repliesBefore.contains(t))
|| pendingTicketIdsFor(lead).stream().anyMatch(t -> !ticketsBefore.contains(t))
|| pendingQuestionTurnIdsFor(lead).stream().anyMatch(t -> !questionsBefore.contains(t));
if (racedIn && activeLeads.putIfAbsent(lead, Boolean.TRUE) == null) {
log.debug("push: new work for lead {} raced the reminder loop's stop — restarting", lead);
scheduleNext(lead);
return;
}
log.debug("push: reminder loop ended for lead {}", lead);
}
/** Send one combined nudge covering everything currently pending for {@code lead}. */
private void injectNudge(String lead, int replyReminderCount, int ticketReminderCount,
int questionReminderCount) {
// Re-read rather than threading it down from decide(): a reply can drain, a ticket be
// collected, or a question be answered (or another arrive), between the decision and the
// injection.
Set<String> replyTargets = pendingReplyTargetsFor(lead);
List<PendingTicket> tickets = pendingTicketsFor(lead);
List<PendingQuestion> questions = pendingQuestionsFor(lead);
if (replyTargets.isEmpty() && tickets.isEmpty() && questions.isEmpty()) {
log.debug("push: pending work for lead {} drained before the nudge could be sent", lead);
return;
}
String nudge = formatNudge(replyTargets, tickets, questions);
try {
agents.send(leadTerminal, nudge);
log.debug("push: nudge {}/{} sent to lead {} for target {}",
reminderCount + 1, maxReminders, leadTerminal, target);
agents.send(lead, nudge);
log.debug("push: nudge sent to lead {} (reply {}/{}, ticket {}/{}, question {}/{}; "
+ "{} reply target(s), {} ticket(s), {} question(s))",
lead, replyReminderCount + 1, maxReminders, ticketReminderCount + 1, maxReminders,
questionReminderCount + 1, maxReminders,
replyTargets.size(), tickets.size(), questions.size());
countNudge("delivered");
} catch (RuntimeException e) {
log.warn("push: failed to nudge lead {} for target {} (reminder {}/{}): {}",
leadTerminal, target, reminderCount + 1, maxReminders, e.toString());
log.warn("push: failed to nudge lead {} (reply {}/{}, ticket {}/{}, question {}/{}): {}",
lead, replyReminderCount + 1, maxReminders, ticketReminderCount + 1, maxReminders,
questionReminderCount + 1, maxReminders, e.toString());
}
// Bump every item actually named in this nudge, not just whatever the shared source-level
// eligibility used to gate (CB-598) — each item's own count is what the next tick's
// minReplyNudgeCountFor / minTicketNudgeCountFor / minQuestionNudgeCountFor will read. An
// item already at or over the cap keeps riding along in the text (still pending, still
// named) but its extra bumps here are inert: decide() already treats it as ineligible once
// its count reaches maxReminders.
bumpNudgeCounts(replyTargets, tickets, questions);
}
/** Record that every one of these items was just named in a sent (or attempted) nudge. */
private void bumpNudgeCounts(Set<String> replyTargets, List<PendingTicket> tickets,
List<PendingQuestion> questions) {
for (String target : replyTargets) {
pendingReplies.computeIfPresent(target, (t, e) -> new ReplyEntry(e.lead(), e.nudgeCount() + 1));
}
for (PendingTicket ticket : tickets) {
pendingTickets.computeIfPresent(ticket.ticket(),
(id, e) -> new PendingTicket(e.ticket(), e.lead(), e.failed(), e.nudgeCount() + 1));
}
for (PendingQuestion question : questions) {
pendingQuestions.computeIfPresent(question.turnId(), (id, e) ->
new PendingQuestion(e.turnId(), e.ticket(), e.target(), e.lead(), e.question(),
e.nudgeCount() + 1));
}
}
/** Schedule the next tick on the scheduler thread pool. */
private void scheduleNext(String target, int nextReminderCount) {
scheduler.schedule(() -> tick(target, nextReminderCount), backoffMs, TimeUnit.MILLISECONDS);
private void scheduleNext(String lead) {
scheduler.schedule(() -> tick(lead),
backoffMs, TimeUnit.MILLISECONDS);
}
// --- nudge formatting ------------------------------------------------------------------------
/** Render everything pending for one lead as a single nudge line. */
private static String formatNudge(Set<String> replyTargets, List<PendingTicket> tickets,
List<PendingQuestion> questions) {
List<String> parts = new ArrayList<>();
if (!replyTargets.isEmpty()) {
parts.add(formatRepliesNudge(replyTargets));
}
if (!tickets.isEmpty()) {
parts.add(formatTicketsNudge(tickets));
}
if (!questions.isEmpty()) {
parts.add(formatQuestionsNudge(questions));
}
return String.join(" | ", parts);
}
/** Render one or several pending reply targets. */
private static String formatRepliesNudge(Set<String> targets) {
if (targets.size() == 1) {
String target = targets.iterator().next();
return NUDGE_FORMAT.formatted(target, target);
}
String ids = String.join(", ", targets);
return REPLIES_NUDGE_FORMAT.formatted(targets.size(), ids);
}
/** Render one or several pending tickets. */
private static String formatTicketsNudge(List<PendingTicket> pending) {
if (pending.size() == 1) {
PendingTicket t = pending.get(0);
return TICKET_NUDGE_FORMAT.formatted(t.ticket(), t.failed() ? " (FAILED)" : "", t.ticket());
}
long failedCount = pending.stream().filter(PendingTicket::failed).count();
String ids = pending.stream()
.map(t -> t.failed() ? t.ticket() + " (FAILED)" : t.ticket())
.collect(Collectors.joining(", "));
String failedNote = failedCount > 0 ? " (%d failed)".formatted(failedCount) : "";
return TICKETS_NUDGE_FORMAT.formatted(pending.size(), failedNote, ids);
}
/** Render one or several open questions (CB-582). */
private static String formatQuestionsNudge(List<PendingQuestion> pending) {
if (pending.size() == 1) {
PendingQuestion q = pending.get(0);
return QUESTION_NUDGE_FORMAT.formatted(q.target(), q.ticket(), q.turnId(), q.question());
}
String ids = pending.stream()
.map(q -> q.ticket() + " (turnId=" + q.turnId() + ")")
.collect(Collectors.joining(", "));
return QUESTIONS_NUDGE_FORMAT.formatted(pending.size(), ids);
}
// --- lifecycle -----------------------------------------------------------------------------
/**
* Whether any reminder loop is currently active for some target (CB-551). The idle-lead heartbeat
* uses this to stand aside: while the push loop is actively nudging the lead, a concurrent
* heartbeat injection would start a second competing turn in the same pane — racing loops multiply
* turns and context burn. "Active" means a schedule exists in {@link #activeTargets}; the set is
* bounded by what has been triggered, not by any persistent state.
* Whether any reminder loop is currently active for some lead (CB-551). The idle-lead heartbeat
* uses this to stand aside: while the push loop is actively nudging a lead, a concurrent
* heartbeat injection would start a second competing turn in the same pane — racing loops
* multiply turns and context burn. "Active" means a schedule exists in {@link #activeLeads},
* which now covers reply-queued (CB-307), ticket-terminal (CB-588), and question-open (CB-582)
* work (CB-590) — bounded by what has been triggered, not by any persistent state.
*/
public boolean isActive() {
return !activeTargets.isEmpty();
return !activeLeads.isEmpty();
}
/** Shut down the scheduler. Outstanding reminders are cancelled. */
public void stop() {
scheduler.shutdownNow();
activeTargets.clear();
activeLeads.clear();
pendingReplies.clear();
pendingTickets.clear();
pendingQuestions.clear();
}
/** @see #stop() */
@@ -0,0 +1,20 @@
package dev.ltms.bridged.msg;
import java.util.concurrent.CompletableFuture;
/**
* Identity for one accepted send. The session turn is deliberately absent: CompletionResolver's
* delivery callback runs before SessionManager.onDelivered, so binding it needs a later ordering design.
*/
public final class TurnToken {
private final String target;
private final CompletableFuture<Rendezvous.Resolution> waiter;
public TurnToken(String target, CompletableFuture<Rendezvous.Resolution> waiter) {
this.target = target;
this.waiter = waiter;
}
public String target() { return target; }
public CompletableFuture<Rendezvous.Resolution> waiter() { return waiter; }
}
@@ -0,0 +1,84 @@
package dev.ltms.bridged.peer;
import java.nio.charset.StandardCharsets;
import java.security.MessageDigest;
import java.security.NoSuchAlgorithmException;
import java.util.HexFormat;
/**
* CB-571: a fingerprint of the exact charter bytes handed to a spawned member.
*
* <p>Lets an operator prove <em>which</em> charter a member actually got, without ever logging the
* charter text. The digest covers the exact composed UTF-8 string {@code HerdrPeerLauncher} passes
* to its adapter as {@code LaunchSpec.charter()}, so every adapter that receives the same string —
* Claude inlining it, OpenCode writing it to a file — produces the same digest for the same config.
* Two spawns of the same role from the same config agree; editing the charter changes the digest.
*
* <p>Deliberately places no charter prose. A charter is operator-authored text that may name
* internal projects or unreleased plans, and logs get tailed, shipped, and pasted into tickets.
* The {@code charterSource} key is what the operator wants to confirm, and it carries no content.
*/
public record CharterReceipt(
MemberRole role,
String profile,
String charterSource,
String charterSha256,
int charterBytes) {
/** Source reported when the role has no configured charter, so the field is never omitted. */
public static final String NO_SOURCE = "none";
/**
* The config key that supplied the role's charter text, e.g. {@code fleet.charters.architect}.
*/
public static String sourceKey(MemberRole role) {
return "fleet.charters." + (role == null ? "?" : role.wireName());
}
/**
* Fingerprint the composed charter for {@code role} on {@code profile}. {@code configured} is
* the role's charter text as read from config ({@code null} when none is configured);
* {@code composed} is the exact string the launcher will pass to the adapter — the reply
* charter may be appended to {@code configured}, or stand alone when no role charter exists.
*
* <p>No composed charter at all is reported as an explicit absence — a {@code null} digest and
* a zero byte count — never a digest of the empty string, which would hide the fact that no
* text was supplied. {@code configured} being {@code null} while {@code composed} is the reply
* charter alone is a normal case, and the source says so.
*/
public static CharterReceipt compose(MemberRole role, String profile,
String configured, String composed) {
String source = (configured == null || configured.isBlank())
? NO_SOURCE : sourceKey(role);
if (composed == null) {
return new CharterReceipt(role, profile, source, null, 0);
}
byte[] bytes = composed.getBytes(StandardCharsets.UTF_8);
return new CharterReceipt(role, profile, source, digestOf(composed), bytes.length);
}
/** Whether the composed charter was absent (no text was given to the member). */
public boolean absent() {
return charterSha256 == null;
}
/**
* The stable SHA-256 hex digest of {@code text}, or {@code null} for null/blank text. Used both
* for the receipt's fingerprint and to redact a charter argument in a spawn log.
*/
public static String digestOf(String text) {
if (text == null || text.isBlank()) {
return null;
}
return sha256Hex(text.getBytes(StandardCharsets.UTF_8));
}
private static String sha256Hex(byte[] bytes) {
try {
MessageDigest md = MessageDigest.getInstance("SHA-256");
return HexFormat.of().formatHex(md.digest(bytes));
} catch (NoSuchAlgorithmException e) {
throw new IllegalStateException("SHA-256 is unavailable", e);
}
}
}
@@ -62,9 +62,27 @@ public interface PeerHandle {
* non-null for a spawn that requested session identity, because it knows the id before the
* peer has written anything.
*
* <p>Deliberately not a {@code default} (CB-584, the same fix CB-571 made for
* {@link #charterReceipt()} one method below): a decorator that forgets to override this
* silently answers {@code null} for a question it has no basis to answer, and the gap surfaces
* only as a resume that quietly starts a cold session, not a compile error. Every
* implementation must answer explicitly.
*
* @return the peer's own session id, or {@code null} when not determinable
*/
default String agentSessionId() {
return null;
}
String agentSessionId();
/**
* The charter receipt (CB-571) for this peer's launch — the fingerprint of the exact charter
* bytes it was started with. {@code null} when the launcher records none (a non-instrumented
* adapter, or a launcher before this field); the session registry stores it so the spawn result
* and the roster row can show an operator which charter a member actually got.
*
* <p>Deliberately not a {@code default}: a decorator that forgets to override this silently
* answers {@code null} for a question it has no basis to answer, and the gap surfaces only as
* a missing roster field, not a compile error. Every implementation must answer explicitly.
*
* @return the fingerprint, or {@code null} when the launcher carries none
*/
CharterReceipt charterReceipt();
}
@@ -25,6 +25,19 @@ public interface PeerLauncher {
*/
Set<Capability> capabilities();
/**
* The capabilities of the adapter that {@code profileName} resolves to (null/blank → the
* default profile, the same resolution {@link #spawn} uses). Distinct from {@link
* #capabilities()}, which unions every configured adapter: a caller that must know whether
* <em>this</em> profile's backend supports a capability — e.g. {@link Capability#SESSION_RESUME}
* before honoring {@link SpawnRequest#resumeSessionId()} — needs the per-profile answer, not
* the fleet-wide union, or a mixed fleet could OK a resume that lands on a non-supporting
* adapter (CB-584).
*
* @throws IllegalArgumentException if the profile is unknown and no default is configured
*/
Set<Capability> capabilitiesFor(String profileName);
/**
* {@code profileName}/requestedCwd null/blank → default resolution. Returns after the peer
* process is live (env + argv + placement complete). Never returns {@code null}.
@@ -0,0 +1,115 @@
package dev.ltms.bridged.placement;
import java.util.LinkedHashMap;
import java.util.Map;
import java.util.Objects;
import java.util.OptionalLong;
import java.util.concurrent.ConcurrentHashMap;
import java.util.function.LongSupplier;
/**
* Where a credential (not a profile — see {@code BridgedConfig.Profile#effectiveCredentialId()})
* sits out a cooldown after a {@code BACKEND_EXHAUSTED} classification (CB-578 stage B), so a fresh
* spawn does not walk straight back onto the account that just refused on a usage limit.
*
* <p>Keyed by credential id, never by profile name: two profiles sharing one credential (e.g. two
* models on the same OpenAI account) share one quarantine — {@link #quarantine} one credential id
* and every profile whose {@code effectiveCredentialId()} equals it is quarantined too, without this
* class knowing anything about profiles at all. That mapping is the caller's job (see
* {@code CompositePeerLauncher} and {@code dev.ltms.bridged.inject.ExhaustionSink}).
*
* <p>The clock is injected ({@link LongSupplier}, conventionally {@code System::nanoTime} like
* {@code FleetHealthMonitor}), never read inline, so a quarantine's expiry is testable without a
* real sleep.
*/
public final class BackendQuarantine {
private final ConcurrentHashMap<String, Long> quarantinedUntilNanos = new ConcurrentHashMap<>();
private final LongSupplier nowNanos;
private final long cooldownNanos;
/** True only for {@link #none()}. See {@link #quarantine} for why this exists. */
private final boolean inert;
/**
* @param nowNanos monotonic clock, injected for testability
* @param cooldownNanos how long a fresh {@link #quarantine} call blocks the credential for;
* must be positive
*/
public BackendQuarantine(LongSupplier nowNanos, long cooldownNanos) {
this(nowNanos, cooldownNanos, false);
}
private BackendQuarantine(LongSupplier nowNanos, long cooldownNanos, boolean inert) {
this.nowNanos = Objects.requireNonNull(nowNanos, "nowNanos");
if (cooldownNanos <= 0) {
throw new IllegalArgumentException("cooldownNanos must be positive: " + cooldownNanos);
}
this.cooldownNanos = cooldownNanos;
this.inert = inert;
}
/**
* Inert quarantine — {@link #quarantine} does nothing on this instance, so nothing is ever
* quarantined. The explicit stand-in a caller (or a test not exercising this feature) passes
* instead of a defaulting overload, exactly like {@code ExhaustedPatternLookup.none()}.
*/
public static BackendQuarantine none() {
return new BackendQuarantine(() -> 0L, 1, true);
}
/**
* Quarantine {@code credentialId} for the configured cooldown, starting now. A repeat call while
* already quarantined restarts the cooldown at full length — a fresh refusal is fresh evidence the
* account is still exhausted, not a reason to let an earlier, shorter wait stand.
*
* <p>On {@link #none()} this is a no-op. It has to be: that instance holds a clock frozen at 0,
* so recording a deadline would produce a quarantine that never expires — a credential locked out
* for the life of the daemon. Two production {@code CompositePeerLauncher} constructors default to
* {@code none()}, so the failure would be silent and permanent. An inert stand-in must omit the
* fact, never invent one.
*/
public void quarantine(String credentialId) {
Objects.requireNonNull(credentialId, "credentialId");
if (inert) {
return;
}
quarantinedUntilNanos.put(credentialId, nowNanos.getAsLong() + cooldownNanos);
}
/** Whether {@code credentialId} is quarantined right now. */
public boolean isQuarantined(String credentialId) {
return remainingNanos(credentialId) > 0;
}
/** Seconds left on {@code credentialId}'s quarantine, or empty when it is not quarantined. */
public OptionalLong remainingSeconds(String credentialId) {
long remaining = remainingNanos(credentialId);
return remaining > 0 ? OptionalLong.of(toSecondsRoundedUp(remaining)) : OptionalLong.empty();
}
/**
* Every currently-quarantined credential id and its remaining seconds (CB-578 stage B fleet
* reporting) — expired entries are never included. Not pruned from the backing map here: it stays
* small (bounded by the number of distinct credentials ever exhausted) and a lazily-stale entry is
* harmless, since every read already checks the deadline.
*/
public Map<String, Long> activeRemainingSeconds() {
Map<String, Long> out = new LinkedHashMap<>();
quarantinedUntilNanos.forEach((credentialId, deadline) -> {
long remaining = deadline - nowNanos.getAsLong();
if (remaining > 0) {
out.put(credentialId, toSecondsRoundedUp(remaining));
}
});
return out;
}
private long remainingNanos(String credentialId) {
Long deadline = quarantinedUntilNanos.get(credentialId);
return deadline == null ? 0L : deadline - nowNanos.getAsLong();
}
private static long toSecondsRoundedUp(long nanos) {
return (nanos + 999_999_999L) / 1_000_000_000L;
}
}
@@ -4,19 +4,63 @@ package dev.ltms.bridged.placement;
* Backward-compatible placement: an unqualified spawn always resolves to the configured default
* profile, exactly as {@code CompositePeerLauncher} did before CB-518. This ignores caps and
* reachability so that a pre-existing config behaves identically after upgrade.
*
* <p>Two exceptions walk past the default instead of returning it unconditionally:
* <ul>
* <li>Quarantine (CB-578 stage B): a quarantined default is a credential that just refused on
* a usage limit, not a transient capacity or reachability concern.
* <li>Weight 0 (CB-554): {@code fixed} is still automatic selection, so a profile the operator
* marked "never auto-select me" ({@code weight <= 0}) must be skipped here exactly as
* {@code weighted}/{@code round-robin} skip it — an explicit {@code bridge_spawn} naming
* the profile is unaffected, only this automatic fallback walk.
* </ul>
* A fleet where nothing is ever quarantined or weight-0 never exercises either path, so today's
* behaviour is unchanged.
*/
final class FixedPlacementPolicy implements PlacementPolicy {
@Override
public PlacementCandidate select(PlacementContext ctx) {
String d = ctx.defaultProfile();
if (d != null && !d.isBlank()) {
if (d != null && !d.isBlank() && !ctx.quarantined().contains(d) && !weightExcluded(ctx, d)) {
return new PlacementCandidate(d, null, 1.0f, null);
}
for (PlacementCandidate c : ctx.candidates()) {
if (!ctx.quarantined().contains(c.profile()) && !c.excluded()) {
return new PlacementCandidate(c.profile(), null, c.weight(), c.maxLoad());
}
}
if (d != null && !d.isBlank()) {
boolean dQuarantined = ctx.quarantined().contains(d);
boolean dWeightExcluded = weightExcluded(ctx, d);
if (dQuarantined && dWeightExcluded) {
throw new PlacementException("worker profile '" + d + "' is quarantined (backend "
+ "exhausted) and has weight 0 (excluded from automatic selection), and no "
+ "available candidate remains");
}
if (dWeightExcluded) {
throw new PlacementException("worker profile '" + d + "' has weight 0 (excluded "
+ "from automatic selection) and no available candidate remains");
}
if (dQuarantined) {
throw new PlacementException("worker profile '" + d + "' is quarantined (backend "
+ "exhausted) and no un-quarantined candidate is available");
}
}
if (!ctx.candidates().isEmpty()) {
PlacementCandidate first = ctx.candidates().getFirst();
return new PlacementCandidate(first.profile(), null, first.weight(), first.maxLoad());
throw new PlacementException(
"all worker profiles are excluded from automatic selection (quarantined or weight-0)");
}
throw new PlacementException("no worker profiles configured");
}
/** Whether {@code profile} carries {@code weight <= 0} (CB-554) among {@code ctx}'s candidates. */
private static boolean weightExcluded(PlacementContext ctx, String profile) {
for (PlacementCandidate c : ctx.candidates()) {
if (c.profile().equals(profile)) {
return c.excluded();
}
}
return false;
}
}
@@ -17,4 +17,18 @@ public record PlacementCandidate(String profile, String host, float weight, Inte
public static PlacementCandidate profile(String profile) {
return new PlacementCandidate(profile, null, 1.0f, null);
}
/**
* True when this candidate carries an explicit {@code weight <= 0} (CB-554) and must be
* skipped by every automatic policy — the same way a quarantined or unreachable candidate is
* skipped. {@code BridgedConfig.Profile}'s compact constructor already normalises "absent" to
* {@code 1.0} and "negative" to {@code 0.0}, so this is a plain threshold check here; it does
* not need to distinguish "explicit 0" from "absent" itself.
*
* <p>Exclusion is about <em>automatic</em> selection only — an explicit
* {@code bridge_spawn{profile:"..."}} bypasses placement entirely and is unaffected.
*/
public boolean excluded() {
return weight <= 0.0f;
}
}
@@ -11,9 +11,13 @@ import java.util.function.Function;
* @param candidates every configured candidate; the policy filters out those at cap or unreachable
* @param liveCount current live worker count per profile (from the session registry)
* @param unreachable profiles already known to have failed in this spawn attempt
* @param quarantined profiles whose credential is currently quarantined (CB-578 stage B) — a
* {@code BACKEND_EXHAUSTED} classification put it, or a profile it shares a
* credential with, on cooldown. Filtered the same way as {@code unreachable}.
*/
public record PlacementContext(String defaultProfile,
List<PlacementCandidate> candidates,
Function<String, Integer> liveCount,
Set<String> unreachable) {
Set<String> unreachable,
Set<String> quarantined) {
}
@@ -12,13 +12,16 @@ final class PlacementPolicyUtil {
}
/**
* Candidates that are not known-unreachable and have not reached their maxLoad.
* Candidates that are not weight-excluded (CB-554: explicit {@code weight <= 0}, checked
* first because it is a static config choice rather than transient state), not
* known-unreachable, not quarantined (CB-578 stage B), and have not reached their maxLoad.
* A {@code null} maxLoad means unlimited.
*/
static List<PlacementCandidate> available(PlacementContext ctx) {
List<PlacementCandidate> out = new ArrayList<>();
for (PlacementCandidate c : ctx.candidates()) {
if (ctx.unreachable().contains(c.profile())) {
if (c.excluded() || ctx.unreachable().contains(c.profile())
|| ctx.quarantined().contains(c.profile())) {
continue;
}
Integer cap = c.maxLoad();
@@ -34,15 +37,23 @@ final class PlacementPolicyUtil {
}
/**
* Build a clear exception describing why every candidate was dropped: all at capacity,
* all unreachable, or a mix.
* Build a clear exception describing why every candidate was dropped: all weight-0, all
* quarantined, all at capacity, all unreachable, or a mix. Each candidate is counted into
* exactly one bucket (weight-excluded takes priority) so a candidate excluded for more than
* one reason is never double-counted.
*/
static PlacementException emptyException(PlacementContext ctx) {
int weightExcluded = 0;
int atCap = 0;
int unreachable = 0;
int quarantined = 0;
for (PlacementCandidate c : ctx.candidates()) {
Integer cap = c.maxLoad();
if (ctx.unreachable().contains(c.profile())) {
if (c.excluded()) {
weightExcluded++;
} else if (ctx.quarantined().contains(c.profile())) {
quarantined++;
} else if (ctx.unreachable().contains(c.profile())) {
unreachable++;
} else if (cap != null && ctx.liveCount().apply(c.profile()) >= cap) {
atCap++;
@@ -53,6 +64,13 @@ final class PlacementPolicyUtil {
if (total == 0) {
return new PlacementException("no worker profiles configured");
}
if (weightExcluded == total) {
return new PlacementException(
"all worker profiles have weight 0 (excluded from automatic selection)");
}
if (quarantined == total) {
return new PlacementException("all worker profiles are quarantined (backend exhausted)");
}
if (atCap == total) {
return new PlacementException("all worker profiles are at maxLoad");
}
@@ -60,6 +78,8 @@ final class PlacementPolicyUtil {
return new PlacementException("all worker profiles are unreachable");
}
return new PlacementException("no worker profile available: " + atCap + " at maxLoad, "
+ unreachable + " unreachable, " + (total - atCap - unreachable) + " remaining");
+ unreachable + " unreachable, " + quarantined + " quarantined, "
+ weightExcluded + " weight-0, "
+ (total - atCap - unreachable - quarantined - weightExcluded) + " remaining");
}
}
@@ -13,6 +13,7 @@ import dev.ltms.bridged.herdr.HerdrClient;
import dev.ltms.bridged.herdr.HerdrException;
import dev.ltms.bridged.inject.MemberPresence;
import dev.ltms.bridged.peer.PeerUnreachableException;
import dev.ltms.bridged.placement.PlacementException;
import dev.ltms.bridged.msg.MessageService;
import dev.ltms.bridged.session.SessionManager;
import dev.ltms.bridged.peer.MemberRole;
@@ -234,7 +235,14 @@ public final class BridgedApp {
List<Map<String, Object>> out = sessions.roster().stream()
.map(s -> SessionManager.rosterView(s, live.get(s.terminalId())))
.toList();
ctx.status(200).json(Map.of("workers", out));
Map<String, Object> body = new LinkedHashMap<>();
body.put("workers", out);
// CB-586: operator visibility for the refs/wip snapshot store without shelling into the
// repo — how many snapshot refs exist and roughly what they cost. Present only once a
// worktree session has established the repo, so a never-snapshotted fleet reports nothing.
sessions.wipRefs().ifPresent(st -> body.put("wipRefs",
Map.of("count", st.count(), "costBytes", st.costBytes())));
ctx.status(200).json(body);
}
/** The configured worker profiles and which one a no-argument spawn uses. */
@@ -292,6 +300,11 @@ public final class BridgedApp {
ctx.status(201).json(view(member));
} catch (GuardException e) {
ctx.status(403).json(Map.of("error", "subscription_boundary", "detail", e.getMessage()));
} catch (PlacementException e) {
// CB-599: no candidate had capacity (maxLoad, quarantine, or all-exhausted) — a benign,
// likely-transient refusal, distinct from "profile does not exist" below. 503: the
// request was valid and will likely succeed later.
ctx.status(503).json(Map.of("error", "no_capacity", "detail", e.getMessage()));
} catch (IllegalArgumentException e) {
ctx.status(400).json(Map.of("error", "unknown_profile", "detail", e.getMessage()));
} catch (PeerUnreachableException e) {
@@ -403,9 +416,12 @@ public final class BridgedApp {
case TIMED_OUT_QUEUED -> "queued";
case BUSY -> "busy";
case WORKER_FAILED -> "failed";
case BACKEND_EXHAUSTED -> "backend_exhausted";
default -> "done"; // unreachable (terminal outcomes handled above)
},
"detail", reply.outcome() == MessageService.Outcome.WORKER_FAILED && reply.text() != null
"detail", (reply.outcome() == MessageService.Outcome.WORKER_FAILED
|| reply.outcome() == MessageService.Outcome.BACKEND_EXHAUSTED)
&& reply.text() != null
? reply.text()
: "no reply within " + timeout + "ms; poll status or retry"));
}
@@ -500,10 +516,20 @@ public final class BridgedApp {
return;
}
try {
ctx.status(200).json(Map.of(
"sessionId", id,
"status", messages.status(id).name().toLowerCase(),
"ready", presence.isPresent(id)));
Map<String, Object> body = new LinkedHashMap<>();
body.put("sessionId", id);
body.put("status", messages.status(id).name().toLowerCase());
body.put("ready", presence.isPresent(id));
// CB-582: a worker paused mid-turn in an async bridge_ask is otherwise invisible to a
// status poll — surface the open question and how to answer it, same as bridge_poll's
// Phase.ASKING view.
MessageService.PendingAsk ask = messages.pendingAsk(id);
if (ask != null) {
body.put("question", ask.question());
body.put("turnId", ask.turnId());
body.put("ticket", ask.ticket());
}
ctx.status(200).json(body);
} catch (HerdrException e) {
herdrError(ctx, e);
}
@@ -529,6 +555,11 @@ public final class BridgedApp {
if (v.detail() != null) {
body.put("detail", v.detail());
}
// CB-582: Phase.ASKING carries the question in v.reply() (handled above) and its answer-
// correlation id here — a REST caller polling this ticket otherwise has no way to answer it.
if (v.turnId() != null) {
body.put("turnId", v.turnId());
}
ctx.status(200).json(body);
}
@@ -12,7 +12,12 @@ import java.nio.file.Files;
import java.nio.file.Path;
import java.nio.file.StandardCopyOption;
import java.security.SecureRandom;
import java.util.ArrayList;
import java.util.HashSet;
import java.util.List;
import java.util.Map;
import java.util.Optional;
import java.util.Set;
import java.util.concurrent.TimeUnit;
import java.util.concurrent.atomic.AtomicLong;
import java.util.stream.Collectors;
@@ -167,6 +172,22 @@ public final class GitWorktrees implements Worktrees {
exec("git", "-C", repoRoot, "worktree", "remove", "--force", worktreePath);
}
@Override
public boolean hasUncommitted(String worktreePath) {
// A worktree that is already gone holds no work to lose, and it must not break teardown:
// git -C <missing-dir> status exits non-zero and would throw where release() is mid-way
// through stopping a pane. Mirror remove()'s already-gone tolerance by treating it as clean.
Path p = Path.of(worktreePath);
if (!Files.exists(p)) {
log.debug("worktree {} already gone — nothing can be uncommitted", worktreePath);
return false;
}
// No --untracked-files=no: the exact shape of the work lost in CB-576 was a new file
// that was never added, so an untracked-only worktree is still dirty.
String out = exec("git", "-C", worktreePath, "status", "--porcelain");
return !out.isBlank();
}
@Override
public void overlayParity(String repoRoot, String worktreePath, List<String> overlay) {
if (overlay == null || overlay.isEmpty()) {
@@ -201,6 +222,209 @@ public final class GitWorktrees implements Worktrees {
return Path.of(out.trim()).toAbsolutePath().normalize().toString();
}
/**
* CB-578 stage C, fixed by CB-587. Stages into a <em>temporary</em> index (never the worktree's
* real one, which the worker may still be writing to) — but that temp index is first <em>seeded</em>
* from the worktree's real one, rather than starting empty:
*
* <pre>
* cp $(git -C worktree rev-parse --git-path index) &lt;temp&gt;
* GIT_INDEX_FILE=&lt;temp&gt; git -C worktree add -A
* tree=$(GIT_INDEX_FILE=&lt;temp&gt; git -C worktree write-tree)
* commit=$(git -C worktree commit-tree $tree -p HEAD -m message)
* git -C worktree update-ref refs/wip/branch $commit
* </pre>
*
* A fresh empty index carries none of the real index's {@code --skip-worktree} /
* {@code --assume-unchanged} bits, so {@code add -A} into it stages a skip-worktree file's local
* on-disk content even though {@code git status --porcelain} correctly hides that file (CB-587).
* Seeding from the real index preserves those bits, so {@code add -A} then skips exactly what
* {@code git status} skips. {@code add -A} (never {@code -f}) also still respects
* {@code .gitignore} exactly as it would in the real index — a gitignored file staying ignored is
* what keeps secrets and local config out of the snapshot's tree. The temporary index file is
* removed afterwards regardless of outcome; the worker's real index is never opened for writing.
*/
@Override
public Optional<String> snapshot(String worktreePath, String branch, String message) {
if (!Files.exists(Path.of(worktreePath))) {
log.debug("worktree {} already gone — nothing to snapshot", worktreePath);
return Optional.empty();
}
Path tempIndex;
try {
tempIndex = Files.createTempFile("bridged-wip-index-", ".tmp");
} catch (IOException e) {
throw new WorktreeException("cannot create a temporary index for snapshot: " + e.getMessage(), e);
}
Map<String, String> indexEnv = Map.of("GIT_INDEX_FILE", tempIndex.toAbsolutePath().toString());
try {
Path realIndex = resolveRealIndex(worktreePath);
try {
Files.copy(realIndex, tempIndex, StandardCopyOption.REPLACE_EXISTING);
} catch (IOException e) {
throw new WorktreeException("cannot copy the worktree's real index (" + realIndex
+ ") into the temporary snapshot index: " + e.getMessage(), e);
}
exec(indexEnv, "git", "-C", worktreePath, "add", "-A");
String tree = exec(indexEnv, "git", "-C", worktreePath, "write-tree").trim();
String commit = exec("git", "-C", worktreePath, "commit-tree", tree, "-p", "HEAD", "-m", message).trim();
exec("git", "-C", worktreePath, "update-ref", "refs/wip/" + branch, commit);
log.info("snapshotted worktree {} to refs/wip/{} commit={}", worktreePath, branch, commit);
return Optional.of(commit);
} finally {
try {
Files.deleteIfExists(tempIndex);
} catch (IOException e) {
log.debug("could not delete temporary snapshot index {}: {}", tempIndex, e.getMessage());
}
}
}
/**
* Resolve the path of {@code worktreePath}'s real index. Never assume {@code <worktree>/.git/index}:
* in a linked worktree {@code .git} is a <em>file</em> pointing at the main repo's
* {@code worktrees/<name>/} directory, and that is where the real per-worktree index lives.
* {@code git rev-parse --git-path index} resolves this correctly for both a linked worktree and
* the main checkout. Throws {@link WorktreeException} — same as every other failure in this
* class — if the command fails or the resolved path does not exist, rather than silently
* snapshotting from an empty index.
*/
private Path resolveRealIndex(String worktreePath) {
String out = exec("git", "-C", worktreePath, "rev-parse", "--git-path", "index").trim();
Path index = Path.of(out);
if (!index.isAbsolute()) {
index = Path.of(worktreePath).resolve(index).normalize();
}
if (!Files.exists(index)) {
throw new WorktreeException("worktree's real index not found at resolved path " + index
+ " (git rev-parse --git-path index reported '" + out + "')");
}
return index;
}
/**
* One {@code refs/wip/<branch>} snapshot ref as read by {@link #listWipRefs}: its full ref name,
* the snapshot commit's sha, and that commit's committer time in unix millis (the age of the
* snapshot — a snapshot is written once and never rewritten, so the commit date is the ref's).
*/
private record WipRef(String refName, String sha, long committerMillis) {
String branch() {
return refName.substring("refs/wip/".length());
}
}
@Override
public WipRefStats wipRefs(String repoRoot) {
List<WipRef> refs = listWipRefs(repoRoot);
long costBytes = 0;
for (WipRef ref : refs) {
costBytes += treeSize(repoRoot, ref.sha());
}
return new WipRefStats(refs.size(), costBytes);
}
@Override
public int pruneWipRefs(String repoRoot, long minAgeMillis) {
// The rule is documented on Worktrees#pruneWipRefs: delete only a snapshot whose tree
// content is already reachable from main AND that is older than minAgeMillis. Reachability
// is the floor that keeps a worker's last copy; the age floor keeps a just-written snapshot
// from being swept while a lead may still be looking at it.
List<WipRef> refs = listWipRefs(repoRoot);
if (refs.isEmpty()) {
return 0;
}
long nowMillis = System.currentTimeMillis();
// Resolve what main carries once per sweep, not once per ref.
Set<String> mainObjects = reachableObjectsFromMain(repoRoot);
int deleted = 0;
for (WipRef ref : refs) {
long ageMillis = nowMillis - ref.committerMillis();
if (ageMillis <= minAgeMillis) {
continue; // too recent — never swept, even if it looks recoverable (CB-586)
}
String tree = exec("git", "-C", repoRoot, "rev-parse", ref.sha() + "^{tree}").trim();
if (!mainObjects.contains(tree)) {
// Last copy of the snapshot's content — the worker's work exists nowhere else.
// Never delete automatically (CB-586 criterion 2).
continue;
}
exec("git", "-C", repoRoot, "update-ref", "-d", ref.refName());
deleted++;
log.info("pruned snapshot ref refs/wip/{} commit={} (age {}h): its tree is already "
+ "reachable from main, so the work is preserved; recover from reflog via "
+ "git update-ref refs/wip/{} {}",
ref.branch(), ref.sha(), TimeUnit.MILLISECONDS.toHours(ageMillis),
ref.branch(), ref.sha());
}
return deleted;
}
/**
* Every {@code refs/wip/*} ref (see {@link WipRef}). The committer date is read as a unix
* count of seconds and converted to millis. {@code %00} (NUL) separates the fields because a
* branch name may contain spaces.
*/
private List<WipRef> listWipRefs(String repoRoot) {
String out = exec("git", "-C", repoRoot, "for-each-ref",
"--format=%(refname)%00%(objectname)%00%(committerdate:unix)", "refs/wip/");
List<WipRef> refs = new ArrayList<>();
for (String line : out.split("\\R")) {
if (line.isBlank()) {
continue;
}
String[] parts = line.split("\u0000", -1);
if (parts.length == 3 && !parts[1].isBlank()) {
refs.add(new WipRef(parts[0], parts[1], Long.parseLong(parts[2]) * 1000L));
}
}
return refs;
}
/**
* The set of object shas reachable from {@code main}, or an empty set when {@code main} cannot
* be resolved. An empty set is the safe direction: the retention sweep then concludes nothing
* is recoverable, so it deletes nothing — a repo with no {@code main} must never cause a
* worker's last copy of a snapshot to be dropped on a reachability misreading.
*/
private Set<String> reachableObjectsFromMain(String repoRoot) {
if (exitCode("git", "-C", repoRoot, "rev-parse", "--verify", "main") != 0) {
log.debug("refs/wip retention: no 'main' ref in {} — treating nothing as reachable", repoRoot);
return Set.of();
}
String out = exec("git", "-C", repoRoot, "rev-list", "--objects", "main");
Set<String> objects = new HashSet<>();
for (String line : out.split("\\R")) {
if (line.isBlank()) {
continue;
}
int sp = line.indexOf(' ');
objects.add(sp < 0 ? line : line.substring(0, sp));
}
return objects;
}
/** Approximate cost of a snapshot: the sum of every blob's size in its committed tree. */
private long treeSize(String repoRoot, String sha) {
String out = exec("git", "-C", repoRoot, "ls-tree", "-r", "-l", sha);
long total = 0;
for (String line : out.split("\\R")) {
if (line.isBlank()) {
continue;
}
// ls-tree -l row: "<mode> <type> <object> <size>\t<path>"; the size is only numeric for
// blobs (trees read "-"), so gate on the type token and take the 4th whitespace field.
String[] parts = line.split("\\s+");
if (parts.length >= 4 && "blob".equals(parts[1])) {
try {
total += Long.parseLong(parts[3]);
} catch (NumberFormatException ignored) {
// a '-' size (or any anomaly) contributes nothing to the rough figure
}
}
}
return total;
}
/** Resolve the directory that will hold per-session worktree checkouts. */
private Path resolveRoot(String repoRoot) {
if (configuredRoot != null && !configuredRoot.isBlank()) {
@@ -223,11 +447,20 @@ public final class GitWorktrees implements Worktrees {
* stdout and stderr (merged by redirectErrorStream).
*/
private String exec(String... command) {
return exec(Map.of(), command);
}
/** Same as {@link #exec(String...)}, with extra environment variables set on the child process. */
private String exec(Map<String, String> extraEnv, String... command) {
String out;
int code;
Process p;
try {
p = new ProcessBuilder(command).redirectErrorStream(true).start();
ProcessBuilder pb = new ProcessBuilder(command).redirectErrorStream(true);
if (extraEnv != null && !extraEnv.isEmpty()) {
pb.environment().putAll(extraEnv);
}
p = pb.start();
} catch (IOException e) {
throw new WorktreeException("failed to start " + command[0] + ": " + e.getMessage(), e);
}
@@ -1,5 +1,6 @@
package dev.ltms.bridged.session;
import dev.ltms.bridged.peer.CharterReceipt;
import dev.ltms.bridged.peer.MemberRole;
/**
@@ -23,6 +24,13 @@ import dev.ltms.bridged.peer.MemberRole;
* @param lastActivityAtNanos {@link System#nanoTime()} of the most recent lifecycle event
* @param turnCount number of delegated turns that have been delivered to this session
* @param state current lifecycle state in the one-shot FSM
* @param charterReceipt the fingerprint (CB-571) of the charter bytes this member was started
* with; {@code null} for a session whose launcher recorded none
* @param agentSessionId the peer's OWN session id (CB-584) — the handle a later
* {@code resumeSessionId} spawn would pass back to resume this exact
* conversation; {@code null} when the launcher could not determine one
* (an adapter that declines {@code Capability.SESSION_RESUME}, or one
* that resolves it lazily and has not yet)
*/
public record MemberSession(
String paneId,
@@ -36,7 +44,9 @@ public record MemberSession(
int turnCount,
State state,
String worktree,
String branch) {
String branch,
CharterReceipt charterReceipt,
String agentSessionId) {
/** One-shot worker lifecycle states. */
public enum State {
@@ -48,21 +58,35 @@ public record MemberSession(
RELEASED
}
/**
* Backward-compatible shape: a session with no charter receipt and no agent session id (a test
* or a launcher before CB-571 / CB-584). A separate constructor rather than a new parameter on
* the canonical one, so existing call sites that have nothing to record keep compiling
* unchanged.
*/
public MemberSession(String paneId, String terminalId, String profile, MemberRole role,
String cwd, String ownerTerminal, long spawnedAtNanos,
long lastActivityAtNanos, int turnCount, State state,
String worktree, String branch) {
this(paneId, terminalId, profile, role, cwd, ownerTerminal, spawnedAtNanos,
lastActivityAtNanos, turnCount, state, worktree, branch, null, null);
}
/** Return a copy of this session in {@code state}. */
public MemberSession withState(State state) {
return new MemberSession(paneId, terminalId, profile, role, cwd, ownerTerminal, spawnedAtNanos,
lastActivityAtNanos, turnCount, state, worktree, branch);
lastActivityAtNanos, turnCount, state, worktree, branch, charterReceipt, agentSessionId);
}
/** Return a copy with {@code lastActivityAtNanos} updated to {@code nowNanos}. */
public MemberSession withActivity(long nowNanos) {
return new MemberSession(paneId, terminalId, profile, role, cwd, ownerTerminal, spawnedAtNanos,
nowNanos, turnCount, state, worktree, branch);
nowNanos, turnCount, state, worktree, branch, charterReceipt, agentSessionId);
}
/** Return a copy with the turn count incremented and activity timestamped at {@code nowNanos}. */
public MemberSession bumpTurn(long nowNanos) {
return new MemberSession(paneId, terminalId, profile, role, cwd, ownerTerminal, spawnedAtNanos,
nowNanos, turnCount + 1, state, worktree, branch);
nowNanos, turnCount + 1, state, worktree, branch, charterReceipt, agentSessionId);
}
}
@@ -4,6 +4,8 @@ import dev.ltms.bridged.auth.MemberLifecycle;
import dev.ltms.bridged.herdr.Agent;
import dev.ltms.bridged.inject.TurnListener;
import dev.ltms.bridged.inject.MemberPresence;
import dev.ltms.bridged.msg.TurnToken;
import dev.ltms.bridged.peer.Capability;
import dev.ltms.bridged.peer.MemberRole;
import dev.ltms.bridged.peer.PeerHandle;
import dev.ltms.bridged.peer.PeerLauncher;
@@ -16,6 +18,7 @@ import java.util.LinkedHashMap;
import java.util.List;
import java.util.Map;
import java.util.Optional;
import java.util.Set;
import java.util.concurrent.ConcurrentHashMap;
import java.util.concurrent.TimeUnit;
import java.util.concurrent.atomic.AtomicLong;
@@ -50,11 +53,19 @@ public final class SessionManager implements TurnListener {
private final int contextCap;
private final boolean clearAfterTurn;
private volatile MemberLifecycle memberLifecycle = MemberLifecycle.NONE;
/**
* CB-586: the repo root the fleet actually works in, remembered the first time a worktree
* session is spawned (worktrees are checkouts of it). {@code refs/wip/*} live there, and this
* single cached value is what the snapshot retention sweep and the operator-visible census run
* against. The daemon is bridged into one project at a time, so "the first worktree's repo" is
* the repo; {@code null} until any worktree is spawned, meaning nothing to sweep or measure.
*/
private volatile String fleetRepoRoot;
/** CB-520: notified with a terminalId on every acquire; no-op until wired. */
private final List<Consumer<String>> acquireListeners = new java.util.concurrent.CopyOnWriteArrayList<>();
/** CB-516: notified with a terminalId on every release; no-op until wired. */
private final List<Consumer<String>> releaseListeners = new java.util.concurrent.CopyOnWriteArrayList<>();
/** CB-516: notified with a {@link ReleaseDetail} on every release; no-op until wired. */
private final List<Consumer<ReleaseDetail>> releaseListeners = new java.util.concurrent.CopyOnWriteArrayList<>();
/** Backward-compatible constructor: shared-tree sessions, production git seam. */
public SessionManager(PeerLauncher launcher) {
@@ -125,21 +136,49 @@ public final class SessionManager implements TurnListener {
}
/**
* Spawn a member, optionally inside a fresh git worktree. When {@code wt} is non-null the
* worktree is provisioned, parity-overlaid, and its path becomes the member's cwd. On any
* failure before registration the worktree is removed so no dangling checkout is left.
* Spawn a member, optionally inside a fresh git worktree, with no session identity requested.
* Equivalent to {@link #acquire(String, MemberRole, String, String, String, WorktreeRequest,
* String, String)} with both trailing args {@code null}.
*
* @param profile which backend to run on — a {@code profiles:} key
* @param role which contract the member runs under; never {@code null}
*/
public MemberSession acquire(String profile, MemberRole role, String requestedCwd, String callerCwd,
String ownerTerminal, WorktreeRequest wt) {
return acquire(profile, role, requestedCwd, callerCwd, ownerTerminal, wt, null, null);
}
/**
* Spawn a member, optionally inside a fresh git worktree, optionally onto a chosen or resumed
* agent session (CB-547a / CB-584). When {@code wt} is non-null the worktree is provisioned,
* parity-overlaid, and its path becomes the member's cwd. On any failure before registration the
* worktree is removed so no dangling checkout is left.
*
* <p>A non-blank {@code resumeSessionId} requires an explicit {@code profile}: a resumed
* conversation is tied to the specific backend that started it, so an unqualified spawn (whose
* backend a placement policy picks at spawn time) has no safe candidate to check the capability
* against. It also requires that profile's adapter to declare {@link Capability#SESSION_RESUME};
* refusing rather than silently delivering a cold session on a non-supporting adapter is the
* whole point of checking before spawn, not after (CB-584).
*
* @param profile which backend to run on — a {@code profiles:} key
* @param role which contract the member runs under; never {@code null}
* @param sessionName the bridge's logical name for the session, or {@code null}
* @param resumeSessionId the prior agent session to resume, or {@code null} for a fresh one
* @throws IllegalArgumentException if {@code resumeSessionId} is set with no explicit profile,
* or the resolved profile's adapter lacks
* {@link Capability#SESSION_RESUME}
*/
public MemberSession acquire(String profile, MemberRole role, String requestedCwd, String callerCwd,
String ownerTerminal, WorktreeRequest wt,
String sessionName, String resumeSessionId) {
MemberRole memberRole = (role == null) ? MemberRole.DEV : role;
requireResumeCapability(profile, resumeSessionId);
if (wt == null) {
// CB-557: the role must ride on the SpawnRequest, not stay a local. The launcher needs it
// to pick the profile out of that role's pool and to label the tab; a role kept only on
// the MemberSession is recorded after the spawn it was supposed to steer.
SpawnRequest req = new SpawnRequest(profile, requestedCwd, callerCwd, null, null, memberRole);
SpawnRequest req = new SpawnRequest(profile, requestedCwd, callerCwd, sessionName, resumeSessionId, memberRole);
PeerHandle handle;
try {
handle = launcher.spawn(req);
@@ -162,7 +201,9 @@ public final class SessionManager implements TurnListener {
0,
MemberSession.State.SPAWNING,
null,
null);
null,
handle.charterReceipt(),
handle.agentSessionId());
registry.put(handle.id(), session);
memberLifecycle.acquired(session.role(), session.profile(), session.terminalId());
log.debug("acquired session id={} terminal={} profile={} owner={}",
@@ -170,7 +211,30 @@ public final class SessionManager implements TurnListener {
notifyAcquired(session.terminalId());
return session;
}
return acquireWithWorktree(profile, memberRole, requestedCwd, callerCwd, ownerTerminal, wt);
return acquireWithWorktree(profile, memberRole, requestedCwd, callerCwd, ownerTerminal, wt,
sessionName, resumeSessionId);
}
/**
* Refuse a {@code resumeSessionId} that either names no explicit profile or names one whose
* adapter does not declare {@link Capability#SESSION_RESUME}. A no-op when
* {@code resumeSessionId} is blank — the ordinary, no-identity spawn path.
*/
private void requireResumeCapability(String profile, String resumeSessionId) {
if (resumeSessionId == null || resumeSessionId.isBlank()) {
return;
}
if (profile == null || profile.isBlank()) {
throw new IllegalArgumentException("resumeSessionId requires an explicit profile — "
+ "a resumed conversation is tied to the backend that started it, so it cannot "
+ "be left to placement to pick");
}
Set<Capability> caps = launcher.capabilitiesFor(profile);
if (!caps.contains(Capability.SESSION_RESUME)) {
throw new IllegalArgumentException("worker profile '" + profile + "' does not declare "
+ "Capability.SESSION_RESUME — refusing resumeSessionId rather than silently "
+ "starting a cold session");
}
}
/** Tear a worker down by pane id and remove it from the registry. Idempotent. */
@@ -194,24 +258,110 @@ public final class SessionManager implements TurnListener {
private void release(String paneId, ReleaseCause cause) {
MemberSession removed = registry.remove(paneId);
boolean preserveWorktree = cause == ReleaseCause.SHUTDOWN;
String snapshotRef = null;
if (removed != null) {
memberLifecycle.released(removed.terminalId());
log.debug("releasing session pane={} terminal={} state={} cause={}",
removed.paneId(), removed.terminalId(), removed.state(), cause);
if (preserveWorktree && removed.worktree() != null) {
logPreservedForShutdown(removed);
try {
memberLifecycle.released(removed.terminalId());
log.debug("releasing session pane={} terminal={} state={} cause={}",
removed.paneId(), removed.terminalId(), removed.state(), cause);
boolean dirty = removed.worktree() != null && worktrees.hasUncommitted(removed.worktree());
if (preserveWorktree && removed.worktree() != null) {
logPreservedForShutdown(removed);
} else if (dirty) {
// CB-576: a release that would otherwise remove the worktree finds it holding
// uncommitted work the bridge cannot see. A worker that ends a turn without
// committing (normally because it stopped to ask a question or refused the turn)
// has its only copy of that work in the worktree. Remove would --force-delete it,
// so preserve the directory and tell an operator where to find it.
preserveWorktree = true;
log.warn("release {} preserves dirty worktree {} for pane={} terminal={}: "
+ "the worktree holds uncommitted changes that --force remove would destroy",
cause, removed.worktree(), removed.paneId(), removed.terminalId());
}
if (dirty) {
// CB-578 stage C: preserving on disk is not saving — the directory is one
// `worktree remove --force`, or an operator tidying up, away from gone. Commit
// its full state to a ref before the preserve-or-remove decision above can be
// undone by anything else, regardless of why this release fired.
snapshotRef = trySnapshot(removed, cause);
}
} catch (RuntimeException e) {
// CB-581: hasUncommitted shells out to `git status` and can throw on a non-zero
// exit. We can no longer tell whether the worktree holds uncommitted work, so fail
// toward the safe answer and preserve it — deleting on a guess can destroy work
// that has no other copy (CB-576), while keeping it on a false alarm only costs
// disk. The exception must not propagate: the pane still has to stop below.
preserveWorktree = true;
log.warn("release {} could not tell whether worktree {} for pane={} terminal={} has "
+ "uncommitted changes; preserving it rather than risk destroying unsaved work: {}",
cause, removed.worktree(), removed.paneId(), removed.terminalId(), e.toString());
} finally {
// CB-516/CB-581: a send still waiting on this worker can never be answered now, no
// matter what happened above. Tell the listener BEFORE the pane is torn down, so a
// blocked caller fails fast with a real reason instead of sitting on a rendezvous
// nothing will ever resolve. CB-578 stage C: carry the worktree/branch/snapshot ref
// too, so a failed ticket's detail can point a lead at the same tree to re-dispatch.
// CB-584 (issue #65 criterion 5): carry agentSessionId alongside them, so a lead can
// also resume the member's conversation, not just re-dispatch onto its files.
notifyReleased(new ReleaseDetail(removed.terminalId(), removed.worktree(),
removed.branch(), snapshotRef, removed.agentSessionId()));
}
// CB-516: a send still waiting on this worker can never be answered now. Tell the
// listener BEFORE the pane is torn down, so a blocked caller fails fast with a real
// reason instead of sitting on a rendezvous nothing will ever resolve.
notifyReleased(removed.terminalId());
}
// CB-581: the pane must always stop, even if the dirty check above threw. A session removed
// from the registry with no pane stop is an orphaned pane — a live terminal burning a fleet
// slot that no longer appears in the roster and can never be reclaimed.
launcher.stop(paneId);
if (removed != null && !preserveWorktree && removed.worktree() != null) {
worktrees.remove(worktrees.repoRoot(removed.cwd()), removed.worktree());
}
}
/**
* Best-effort snapshot of a dirty worktree into {@code refs/wip/<branch>} (CB-578 stage C). A
* failure here must never escalate: the caller has already decided to preserve the worktree
* regardless of whether this succeeds, so the only cost of a failed snapshot is a WARN and a
* missing ref — never a lost pane stop or a lost release notification.
*/
private String trySnapshot(MemberSession session, ReleaseCause cause) {
if (session.worktree() == null || session.branch() == null) {
return null;
}
try {
Optional<String> ref = worktrees.snapshot(session.worktree(), session.branch(),
snapshotMessage(session, cause));
ref.ifPresent(sha -> log.info(
"snapshotted dirty worktree {} to refs/wip/{} commit={} for pane={} terminal={}",
session.worktree(), session.branch(), sha, session.paneId(), session.terminalId()));
return ref.orElse(null);
} catch (RuntimeException e) {
log.warn("snapshot of dirty worktree {} failed for pane={} terminal={} branch={}: the "
+ "worktree is still preserved on disk, just not committed to refs/wip/{}: {}",
session.worktree(), session.paneId(), session.terminalId(), session.branch(),
session.branch(), e.toString());
return null;
}
}
/** Commit message for a CB-578 stage C snapshot — names the member so an operator can tell runs apart. */
private String snapshotMessage(MemberSession session, ReleaseCause cause) {
return "CB-578 stage C: snapshot of a released worker\n\n"
+ "terminal: " + session.terminalId() + "\n"
+ "profile: " + session.profile() + "\n"
+ "branch: " + session.branch() + "\n"
+ "cause: " + cause;
}
/**
* Facts about a released session that a listener needs beyond the bare terminal id — enough
* for a caller to point a lead at where to re-dispatch onto the same tree after a failed
* release (CB-578 stage C, acceptance criterion 10). {@code worktreePath} and {@code branch}
* are {@code null} for a shared-tree session; {@code snapshotRef} is {@code null} unless this
* release snapshotted a dirty worktree into {@code refs/wip/<branch>}.
*/
public record ReleaseDetail(String terminalId, String worktreePath, String branch, String snapshotRef,
String agentSessionId) {
}
/**
* Why a session is being released — governs whether its worktree is preserved or removed.
* Worktree removal is reserved for the one case that is genuinely finished; everything else
@@ -260,7 +410,7 @@ public final class SessionManager implements TurnListener {
* presence view this manager exposes). Wiring it at construction would require breaking that
* cycle for one callback.
*/
public void onRelease(Consumer<String> listener) {
public void onRelease(Consumer<ReleaseDetail> listener) {
if (listener != null) {
releaseListeners.add(listener);
}
@@ -286,21 +436,22 @@ public final class SessionManager implements TurnListener {
}
/** A listener failure must never prevent the teardown it is reacting to. */
private void notifyReleased(String terminalId) {
if (terminalId == null) {
private void notifyReleased(ReleaseDetail detail) {
if (detail.terminalId() == null) {
return;
}
for (Consumer<String> listener : releaseListeners) {
for (Consumer<ReleaseDetail> listener : releaseListeners) {
try {
listener.accept(terminalId);
listener.accept(detail);
} catch (RuntimeException e) {
log.warn("release listener failed for terminal {}: {}", terminalId, e.toString());
log.warn("release listener failed for terminal {}: {}", detail.terminalId(), e.toString());
}
}
}
private MemberSession acquireWithWorktree(String profile, MemberRole memberRole, String requestedCwd, String callerCwd,
String ownerTerminal, WorktreeRequest wt) {
String ownerTerminal, WorktreeRequest wt,
String sessionName, String resumeSessionId) {
String preResolvedProfile = (profile == null || profile.isBlank())
? launcher.defaultProfile() : profile;
// CB-507: resolve through the launcher's CB-112 chain (requested → profile cwd → caller →
@@ -310,13 +461,18 @@ public final class SessionManager implements TurnListener {
// The non-worktree path always used this chain; only this branch was missed.
String repoRoot = worktrees.repoRoot(
launcher.effectiveCwd(new SpawnRequest(preResolvedProfile, requestedCwd, callerCwd)));
if (fleetRepoRoot == null) {
// CB-586: remember the repo whose worktrees the fleet spawns — its refs/wip/* are the
// snapshot store the retention sweep and the operator census operate on.
fleetRepoRoot = repoRoot;
}
String branch = "worker/" + slug(wt.ticketSlug()) + "-" + nonce();
String path = null;
PeerHandle handle;
try {
path = worktrees.add(repoRoot, branch, wt.baseRef());
worktrees.overlayParity(repoRoot, path, launcher.parityOverlay(preResolvedProfile));
handle = launcher.spawn(new SpawnRequest(profile, path, callerCwd, null, null, memberRole));
handle = launcher.spawn(new SpawnRequest(profile, path, callerCwd, sessionName, resumeSessionId, memberRole));
} catch (RuntimeException e) {
log.warn("spawn failed for profile={} role={} branch={} path={}: {}",
preResolvedProfile, memberRole, branch, path, e.getMessage());
@@ -344,7 +500,9 @@ public final class SessionManager implements TurnListener {
0,
MemberSession.State.SPAWNING,
path,
branch);
branch,
handle.charterReceipt(),
handle.agentSessionId());
registry.put(handle.id(), session);
memberLifecycle.acquired(session.role(), session.profile(), session.terminalId());
log.debug("acquired worktree session id={} terminal={} profile={} branch={} path={}",
@@ -410,6 +568,22 @@ public final class SessionManager implements TurnListener {
if (session.ownerTerminal() != null) {
m.put("owner", session.ownerTerminal());
}
// CB-584: which conversation this member holds — the id a later resumeSessionId spawn would
// pass back. Absent for an adapter that declines Capability.SESSION_RESUME, or one that
// resolves it lazily and has not yet.
if (session.agentSessionId() != null) {
m.put("agentSessionId", session.agentSessionId());
}
// CB-571: which charter this member was started with — never the charter text itself. The
// digest lets a lead tell at a glance whether all members got the same charter; the source
// records whether a role charter was configured ("fleet.charters.<role>") or only the reply
// charter was composed ("none").
if (session.charterReceipt() != null) {
m.put("charterSource", session.charterReceipt().charterSource());
if (session.charterReceipt().charterSha256() != null) {
m.put("charterSha256", session.charterReceipt().charterSha256());
}
}
m.put("liveStatus", live == null ? "unknown" : live.status().name().toLowerCase());
return m;
}
@@ -425,7 +599,7 @@ public final class SessionManager implements TurnListener {
* can be re-delivered for multi-turn reuse until it is released.
*/
@Override
public void onDelivered(String target) {
public void onDelivered(String target, TurnToken token) {
MemberSession current = findByTerminal(target);
if (current == null) return;
if (current.state() != MemberSession.State.READY && current.state() != MemberSession.State.DONE) {
@@ -516,8 +690,15 @@ public final class SessionManager implements TurnListener {
log.debug("reaping idle session terminal={} pane={}: idle {}s exceeds the {}s ttl",
s.terminalId(), s.paneId(), TimeUnit.NANOSECONDS.toSeconds(idleNanos),
TimeUnit.NANOSECONDS.toSeconds(idleTtlNanos));
release(s.paneId());
reaped++;
// CB-581: one session that fails to release must not abort the whole reaping pass —
// match drainAll's per-session try/catch so the rest of the roster still gets reaped.
try {
release(s.paneId());
reaped++;
} catch (RuntimeException e) {
log.warn("reap failed for pane={} terminal={} worktree={}; continuing with "
+ "remaining sessions", s.paneId(), s.terminalId(), s.worktree(), e);
}
}
}
return reaped;
@@ -575,6 +756,26 @@ public final class SessionManager implements TurnListener {
return registry.size();
}
/**
* CB-586: the operator-visible census of {@code refs/wip/*} in the repo the fleet works in —
* how many snapshot refs exist and roughly what they cost. Empty (no repo known) until at
* least one worktree session has been spawned, exactly so a fleet that has never snapshotted
* anything surfaces nothing new, as it did before CB-586.
*/
public Optional<Worktrees.WipRefStats> wipRefs() {
String repo = fleetRepoRoot;
return repo == null ? Optional.empty() : Optional.of(worktrees.wipRefs(repo));
}
/**
* CB-586: run the snapshot retention sweep in the fleet's repo (a no-op until a worktree has
* been spawned, which establishes the repo). Returns how many {@code refs/wip/*} it deleted.
*/
public int sweepWipRefs(long minAgeMillis) {
String repo = fleetRepoRoot;
return repo == null ? 0 : worktrees.pruneWipRefs(repo, minAgeMillis);
}
/**
* The registered session owning {@code terminalId}, or {@code null} if none does.
*
@@ -14,12 +14,29 @@ public final class SessionReaper {
private static final Logger log = LoggerFactory.getLogger(SessionReaper.class);
private static final long DEFAULT_INTERVAL_MILLIS = 5000;
/** CB-586: the refs/wip age floor — never sweep a snapshot younger than 24h (the CB-586 rule). */
private static final long WIP_MIN_AGE_MILLIS = TimeUnit.HOURS.toMillis(24);
/**
* CB-586: how often the retention sweep runs. Given the 24h age floor, running it every few
* hours means a ref is dropped within hours of becoming eligible, never within minutes.
*/
private static final long WIP_SWEEP_INTERVAL_NANOS = TimeUnit.HOURS.toNanos(6);
private final SessionManager sessions;
private final long idleTtlNanos;
private final long intervalMillis;
private volatile boolean running;
private Thread thread;
/**
* When the retention sweep last ran, and whether it ever has. The flag is not a convenience:
* a "never yet" sentinel value cannot be compared by subtraction. {@code Long.MIN_VALUE} was
* the obvious choice and it silently overflows — {@code System.nanoTime()} is positive on this
* platform, so {@code now - Long.MIN_VALUE} wraps to a large negative number, the interval gate
* reads it as "swept moments ago", and it returns before ever assigning the field. The sweep
* then never runs at all, for the life of the process, with nothing in the log to say so.
*/
private volatile boolean sweptOnce;
private volatile long lastWipSweepNanos;
/** Construct a reaper with the default 5-second polling interval. */
public SessionReaper(SessionManager sessions, long idleTtlSeconds) {
@@ -49,10 +66,39 @@ public final class SessionReaper {
} catch (RuntimeException e) {
log.warn("session reaper iteration failed; continuing", e);
}
maybeSweepWipRefs();
sleep();
}
}
/**
* CB-586: run the refs/wip retention sweep on a slow cadence (hours, not the per-iteration
* millisecond loop). Best-effort — a failure must never take the idle-reap loop down with it.
*/
private void maybeSweepWipRefs() {
long now = System.nanoTime();
// The first pass always sweeps: a restart is a fine moment to sweep, the 24h age floor
// makes it safe, and it means the feature is observable right after a redeploy instead of
// six hours later. Only after that does the interval gate apply, and by then both operands
// come from nanoTime, so the subtraction is well-defined.
if (sweptOnce && now - lastWipSweepNanos < WIP_SWEEP_INTERVAL_NANOS) {
return;
}
try {
int deleted = sessions.sweepWipRefs(WIP_MIN_AGE_MILLIS);
if (deleted > 0) {
log.info("refs/wip retention sweep deleted {} snapshot ref(s) older than 24h whose "
+ "content was already reachable from main", deleted);
}
} catch (RuntimeException e) {
log.warn("refs/wip retention sweep failed; continuing", e);
}
// Set even when the sweep threw, so a broken repo is retried on the slow cadence rather
// than hammering git on every 5-second iteration.
lastWipSweepNanos = now;
sweptOnce = true;
}
private void sleep() {
try {
Thread.sleep(intervalMillis);
@@ -1,6 +1,7 @@
package dev.ltms.bridged.session;
import java.util.List;
import java.util.Optional;
/** Seam between {@link SessionManager} and git worktree operations. Tests use a recording fake. */
public interface Worktrees {
@@ -10,9 +11,91 @@ public interface Worktrees {
/** git -C <repoRoot> worktree remove --force <path>. Idempotent (already-gone tolerated). */
void remove(String repoRoot, String worktreePath);
/**
* True when the worktree holds uncommitted changes the bridge cannot see: tracked
* modifications, staged files, or untracked files. {@code git status --porcelain} is the
* test; an empty result means clean. Callers use this to decide whether removing the
* worktree would silently destroy a worker's only copy of its work.
*
* <p>An already-gone worktree is reported as clean (no throw), matching {@link #remove}'s
* idempotent contract: a path that does not exist holds no work to lose, and must not break
* a teardown that is mid-way through stopping the pane.
*/
boolean hasUncommitted(String worktreePath);
/** Copy each existing overlay path repoRoot→worktree; mark tracked ones --skip-worktree. */
void overlayParity(String repoRoot, String worktreePath, List<String> overlay);
/** git -C <cwd> rev-parse --show-toplevel — the repo root that owns cwd. */
String repoRoot(String cwd);
/**
* Commit the worktree's full on-disk state — tracked and untracked, respecting
* {@code .gitignore} — to {@code refs/wip/<branch>}, so a release that would otherwise leave
* the work as loose, unprotected files has a durable git object to fall back on (CB-578 stage
* C). Built on a <em>temporary</em> index: the worker's own index, working tree, and HEAD are
* never touched, since the worker may still be mid-write. The commit is parented on the
* worktree's current HEAD.
*
* <p>Never writes under {@code refs/heads/} — the ref must not appear in {@code git branch},
* must not be pushed by default, and must not be swept by a later {@code git branch -d}.
*
* <p>Callers are expected to have already confirmed {@link #hasUncommitted} before reaching
* for this; it always stages and commits whatever {@code git add -A} finds, so calling it on
* a clean worktree still produces a (harmless, tree-identical-to-HEAD) commit rather than
* detecting cleanliness itself.
*
* @param worktreePath absolute path of the worktree to snapshot
* @param branch the worktree's own branch — keys {@code refs/wip/<branch>}
* @param message the commit message; should name the member, its branch, and the release
* cause so an operator can tell which run produced it
* @return the created commit's sha, or {@link Optional#empty()} if {@code worktreePath} does
* not exist (mirrors {@link #remove} and {@link #hasUncommitted}'s already-gone
* tolerance — a worktree that is gone holds nothing to snapshot)
*/
Optional<String> snapshot(String worktreePath, String branch, String message);
/**
* CB-586: how many {@code refs/wip/*} snapshot refs exist in {@code repoRoot} and roughly what
* they cost. This is the operator-visible surface for the snapshot growth CB-578 stage C left
* behind — counts of refs alone hide that each one pins a whole tree for {@code git gc}.
*
* @param repoRoot the repository to scan
* @return count of snapshot refs, and {@code costBytes} = the approximate total working-tree
* size of every snapshot's committed content (summed per ref, so shared objects are
* counted once per ref that carries them)
*/
WipRefStats wipRefs(String repoRoot);
/**
* CB-586: run the {@code refs/wip/*} retention sweep and return how many refs it deleted.
*
* <p>The retention rule is <em>reachability plus an age floor</em>. A snapshot ref is deleted
* only when <strong>both</strong> hold:
* <ol>
* <li>its commit's <em>tree content</em> is already reachable from {@code main} — the work
* the snapshot preserved has been recovered, so dropping the ref loses nothing; and</li>
* <li>the ref is older than {@code minAgeMillis} — a very recent snapshot is never swept
* while a lead may still be looking at it.</li>
* </ol>
*
* <p>Reachability is the safety property. A snapshot exists precisely because the work was not
* committed anywhere else, so a snapshot whose content is <em>not</em> reachable from
* {@code main} is the <strong>last copy</strong> of a worker's work and must never be deleted
* automatically — that is the failure CB-576 and CB-578 stage C were built to stop. Age alone
* must never drive a deletion, because age-based sweeping is exactly how the last copy gets
* destroyed. (Both numbers and the rule are CB-586's decision; this method only implements it.)
*
* <p>Every deletion logs the ref name and the commit sha, so an operator who finds they lost
* the wrong thing can still recover it from git's reflog.
*
* @param repoRoot the repository whose {@code refs/wip/*} to sweep
* @param minAgeMillis the age floor; a ref younger than this is never touched
* @return the number of snapshot refs deleted
*/
int pruneWipRefs(String repoRoot, long minAgeMillis);
/** CB-586: the operator-visible census of {@code refs/wip/*} in one repository. */
record WipRefStats(int count, long costBytes) {
}
}
+9
View File
@@ -1,4 +1,13 @@
<configuration>
<!--
CB-575: MCP SDK 2.0.0 has no public notification registration API. It registers only
notifications/initialized and notifications/roots/list_changed, so other client notifications
still warn when unhandled. Clients may legitimately send notifications/cancelled; suppress only
that SDK WARN because it would devalue the action-needed WARN level used by M4 fleet health.
If a later SDK handles cancellation, this filter simply stops matching and can be removed.
-->
<turboFilter class="dev.ltms.bridged.logging.McpCancelledNotificationFilter"/>
<appender name="STDOUT" class="ch.qos.logback.core.ConsoleAppender">
<encoder>
<pattern>%d{HH:mm:ss.SSS} %-5level [%thread] %logger{28} - %msg%n</pattern>
@@ -0,0 +1,108 @@
package dev.ltms.bridged;
import ch.qos.logback.classic.Logger;
import ch.qos.logback.classic.spi.ILoggingEvent;
import ch.qos.logback.core.read.ListAppender;
import dev.ltms.bridged.config.BridgedConfig;
import org.junit.jupiter.api.Test;
import org.junit.jupiter.api.io.TempDir;
import org.slf4j.LoggerFactory;
import java.nio.file.Files;
import java.nio.file.Path;
import static org.junit.jupiter.api.Assertions.assertFalse;
import static org.junit.jupiter.api.Assertions.assertTrue;
/**
* CB-596: an absent (or empty) {@code memberCredentials:} block blocks nothing — no name is
* hardcoded any more to fall back on, so the daemon must say so out loud at startup rather than
* silently dropping CB-592's protection. Mirrors {@link RequiredSecretEnvVarsTest}'s pattern for
* the CB-594 startup-secrets report, capturing the real log via a {@link ListAppender}.
*/
class MemberCredentialsGapReportTest {
private static BridgedConfig load(Path dir, String yaml) throws Exception {
Path f = dir.resolve("bridged.yaml");
Files.writeString(f, yaml);
return BridgedConfig.load(f);
}
private static ListAppender<ILoggingEvent> attach() {
Logger logger = (Logger) LoggerFactory.getLogger(Bridged.class);
ListAppender<ILoggingEvent> appender = new ListAppender<>();
appender.start();
logger.addAppender(appender);
return appender;
}
private static void detach(ListAppender<ILoggingEvent> appender) {
((Logger) LoggerFactory.getLogger(Bridged.class)).detachAppender(appender);
}
@Test
void anAbsentBlockWarnsThatEveryMemberInheritsTheWholeStore(@TempDir Path dir) throws Exception {
BridgedConfig cfg = load(dir, "bind:\n host: 127.0.0.1\n port: 8765\n");
ListAppender<ILoggingEvent> appender = attach();
try {
Bridged.reportMemberCredentialsGap(cfg);
} finally {
detach(appender);
}
assertTrue(appender.list.stream().anyMatch(e ->
e.getLevel() == ch.qos.logback.classic.Level.WARN
&& e.getFormattedMessage().contains("memberCredentials")
&& e.getFormattedMessage().contains("WHOLE secret store")),
"an absent block must WARN that protection is lost, not stay silent");
}
@Test
void anEmptyKnownListWarnsTheSameAsAbsent(@TempDir Path dir) throws Exception {
BridgedConfig cfg = load(dir, """
memberCredentials:
policy: deny-by-default
""");
ListAppender<ILoggingEvent> appender = attach();
try {
Bridged.reportMemberCredentialsGap(cfg);
} finally {
detach(appender);
}
assertTrue(appender.list.stream().anyMatch(e ->
e.getLevel() == ch.qos.logback.classic.Level.WARN
&& e.getFormattedMessage().contains("memberCredentials")),
"policy: with no known: names still blocks nothing and must warn the same way");
}
@Test
void aPopulatedKnownListLogsInfoNotWarn(@TempDir Path dir) throws Exception {
BridgedConfig cfg = load(dir, """
memberCredentials:
policy: deny-by-default
allow: [AI_GATEWAY_TOKEN]
known: [AI_GATEWAY_TOKEN, GITEA_ACCESS_TOKEN]
""");
// logback-test.xml pins dev.ltms.bridged to WARN (see its own comment); raise it here so
// the INFO line this test asserts on actually reaches the appender, and restore after.
Logger logger = (Logger) LoggerFactory.getLogger(Bridged.class);
ch.qos.logback.classic.Level original = logger.getLevel();
logger.setLevel(ch.qos.logback.classic.Level.INFO);
ListAppender<ILoggingEvent> appender = attach();
try {
Bridged.reportMemberCredentialsGap(cfg);
} finally {
detach(appender);
logger.setLevel(original);
}
assertFalse(appender.list.stream().anyMatch(e -> e.getLevel() == ch.qos.logback.classic.Level.WARN),
"a configured, non-empty known: list must not warn — the block is doing its job");
assertTrue(appender.list.stream().anyMatch(e -> e.getFormattedMessage().contains("blocking 1")),
"the INFO line should say how many names are actually blocked (known minus allow)");
}
}
@@ -0,0 +1,114 @@
package dev.ltms.bridged;
import dev.ltms.bridged.config.BridgedConfig;
import org.junit.jupiter.api.Test;
import org.junit.jupiter.api.io.TempDir;
import java.nio.file.Files;
import java.nio.file.Path;
import java.util.List;
import java.util.Map;
import static org.junit.jupiter.api.Assertions.assertEquals;
import static org.junit.jupiter.api.Assertions.assertFalse;
import static org.junit.jupiter.api.Assertions.assertTrue;
/**
* CB-594: {@link Bridged#requiredSecretEnvVars(BridgedConfig)} is what decides what the startup
* secret report checks — it must derive that set from the config, not a hand-written list, or a
* new profile's token silently stops being reported.
*/
class RequiredSecretEnvVarsTest {
private static BridgedConfig load(Path dir, String yaml) throws Exception {
Path f = dir.resolve("bridged.yaml");
Files.writeString(f, yaml);
return BridgedConfig.load(f);
}
@Test
void collectsATokenEnvPerNonSubscriptionProfile(@TempDir Path dir) throws Exception {
BridgedConfig cfg = load(dir, """
profiles:
local:
baseUrl: http://gx00.gw:8000
tokenEnv: AI_GATEWAY_TOKEN
""");
Map<String, List<String>> required = Bridged.requiredSecretEnvVars(cfg);
assertTrue(required.containsKey("AI_GATEWAY_TOKEN"));
assertEquals(List.of("profile 'local' tokenEnv"), required.get("AI_GATEWAY_TOKEN"));
}
@Test
void aSubscriptionProfileNeedsNoTokenEnv(@TempDir Path dir) throws Exception {
BridgedConfig cfg = load(dir, """
profiles:
opus:
subscription: true
model: claude-opus-5
""");
assertTrue(Bridged.requiredSecretEnvVars(cfg).isEmpty(),
"subscription: true never reads ANTHROPIC_AUTH_TOKEN — see Profile#isSubscription");
}
@Test
void gitTokenEnvIsOptInAndCollectedWhenSet(@TempDir Path dir) throws Exception {
BridgedConfig cfg = load(dir, """
profiles:
local:
baseUrl: http://gx00.gw:8000
tokenEnv: AI_GATEWAY_TOKEN
gitTokenEnv: WORKER_GITEA_TOKEN
""");
Map<String, List<String>> required = Bridged.requiredSecretEnvVars(cfg);
assertTrue(required.containsKey("WORKER_GITEA_TOKEN"));
assertEquals(List.of("profile 'local' gitTokenEnv"), required.get("WORKER_GITEA_TOKEN"));
}
@Test
void noGitTokenEnvMeansNothingIsRequiredForIt(@TempDir Path dir) throws Exception {
BridgedConfig cfg = load(dir, """
profiles:
local:
baseUrl: http://gx00.gw:8000
tokenEnv: AI_GATEWAY_TOKEN
""");
assertFalse(Bridged.requiredSecretEnvVars(cfg).containsKey("WORKER_GITEA_TOKEN"));
}
@Test
void aVarSharedByTwoProfilesIsReportedOnceNamingBoth(@TempDir Path dir) throws Exception {
BridgedConfig cfg = load(dir, """
profiles:
local:
baseUrl: http://gx00.gw:8000
tokenEnv: AI_GATEWAY_TOKEN
gitTokenEnv: WORKER_GITEA_TOKEN
gx:
kind: opencode
baseUrl: https://llm.ltms.dev/v1
tokenEnv: AI_GATEWAY_TOKEN
gitTokenEnv: WORKER_GITEA_TOKEN
""");
Map<String, List<String>> required = Bridged.requiredSecretEnvVars(cfg);
assertEquals(List.of("profile 'local' tokenEnv", "profile 'gx' tokenEnv"),
required.get("AI_GATEWAY_TOKEN"));
assertEquals(List.of("profile 'local' gitTokenEnv", "profile 'gx' gitTokenEnv"),
required.get("WORKER_GITEA_TOKEN"));
}
@Test
void noProfilesMeansNothingIsRequired(@TempDir Path dir) throws Exception {
BridgedConfig cfg = load(dir, "bind:\n host: 127.0.0.1\n port: 8765\n");
assertTrue(Bridged.requiredSecretEnvVars(cfg).isEmpty());
}
}
@@ -10,6 +10,7 @@ import java.nio.file.Path;
import java.util.List;
import java.util.Map;
import java.util.Set;
import java.util.regex.Pattern;
import static org.junit.jupiter.api.Assertions.*;
@@ -194,7 +195,7 @@ class BridgedConfigTest {
fleet:
leaders:
opus:
terminal: term_opus
tab: "lead: opus"
""");
BridgedConfig.Leader lead = BridgedConfig.load(f).fleet().leaders().get("opus");
@@ -212,7 +213,7 @@ class BridgedConfigTest {
fleet:
leaders:
opus:
terminal: term_opus
tab: "drive: opus"
tabPrefix: "drive:"
scanIntervalSeconds: 30
""");
@@ -224,7 +225,9 @@ class BridgedConfigTest {
/**
* The pane no longer has to exist before the daemon does (CB-557): a lead naming a profile may
* be launched, while one that names only a terminal is recognised and never created.
* be launched, while one that names no profile is recognised and never created. Either way it
* still needs its own {@code tab:} (CB-579) — that part is unconditional, see
* {@link #aLeadWithNoTabRefusesToStart}.
*/
@Test
void aLeadIsCreatableOnlyWhenItNamesAProfile(@TempDir Path dir) throws Exception {
@@ -239,8 +242,9 @@ class BridgedConfigTest {
leaders:
launched:
profile: opus
tab: "lead: launched"
pinned:
terminal: term_opus
tab: "lead: pinned"
""");
var leaders = BridgedConfig.load(f).fleet().leaders();
@@ -249,8 +253,12 @@ class BridgedConfigTest {
"no profile to launch on ⇒ recognise-only, the pre-CB-557 behaviour");
}
/**
* CB-579: {@code tab} is the only field a lead's identity depends on now, so it is required
* whether the entry is creatable or recognise-only — without it the entry can never be found.
*/
@Test
void aLeadThatCanBeNeitherFoundNorCreatedRefusesToStart(@TempDir Path dir) throws Exception {
void aLeadWithNoTabRefusesToStart(@TempDir Path dir) throws Exception {
Path f = dir.resolve("useless-lead.yaml");
Files.writeString(f, """
bind:
@@ -264,6 +272,80 @@ class BridgedConfigTest {
IllegalStateException e = assertThrows(IllegalStateException.class, cfg::validateMembers);
assertTrue(e.getMessage().contains("ghost"), "the message must name the useless entry");
assertTrue(e.getMessage().contains("tab:"), "the message must say what is missing");
}
/**
* CB-579 acceptance (2): a config still spelling {@code fleet.leaders.<name>.terminal} must fail
* loudly at load, not be silently dropped by {@code Leader}'s {@code @JsonIgnoreProperties}.
*/
@Test
void aLeaderTerminalKeyFailsLoadAndNamesTabAsTheReplacement(@TempDir Path dir) throws Exception {
Path f = dir.resolve("stale-terminal.yaml");
Files.writeString(f, """
bind:
port: 8080
fleet:
leaders:
opus:
terminal: term_opus
""");
IllegalStateException e =
assertThrows(IllegalStateException.class, () -> BridgedConfig.load(f));
assertTrue(e.getMessage().contains("opus"), "the message must name the offending entry");
assertTrue(e.getMessage().contains("tab:"), "the message must name the replacement key");
assertTrue(e.getMessage().contains("terminal"), "the message must name the retired key");
}
/** The same refusal, and it must name every offending entry, not just the first. */
@Test
void everyLeaderStillUsingTerminalIsReportedAtOnce(@TempDir Path dir) throws Exception {
Path f = dir.resolve("stale-terminals.yaml");
Files.writeString(f, """
bind:
port: 8080
fleet:
leaders:
opus:
terminal: term_opus
sol:
terminal: term_sol
""");
IllegalStateException e =
assertThrows(IllegalStateException.class, () -> BridgedConfig.load(f));
assertTrue(e.getMessage().contains("opus"));
assertTrue(e.getMessage().contains("sol"));
}
/** A {@code terminal:} anywhere else in the document (not under a leader entry) is unaffected. */
@Test
void aTerminalKeyOutsideFleetLeadersIsNotRejected(@TempDir Path dir) throws Exception {
Path f = dir.resolve("primary-terminal-ok.yaml");
Files.writeString(f, "bind:\n port: 8080\nprimary:\n terminal: term_fixed\n");
assertDoesNotThrow(() -> BridgedConfig.load(f));
}
/** CB-579 acceptance (3): distinct `tab:` labels need no shared prefix — one scanner finds both. */
@Test
void twoLeadersWithDifferentTabsAreBothConfigured(@TempDir Path dir) throws Exception {
Path f = dir.resolve("two-tabs.yaml");
Files.writeString(f, """
bind:
port: 8080
fleet:
leaders:
opus:
tab: "lead: opus"
sol:
tab: "captain: sol"
""");
var leaders = BridgedConfig.load(f).fleet().leaders();
assertEquals("lead: opus", leaders.get("opus").tab());
assertEquals("captain: sol", leaders.get("sol").tab());
}
// ── CB-551: the idle-lead heartbeat ─────────────────────────────────────────────────────────
@@ -328,7 +410,7 @@ class BridgedConfigTest {
fleet:
leaders:
opus:
terminal: term_opus
tab: "lead: opus"
tabPrefix: "lead:"
""");
BridgedConfig cfg = BridgedConfig.load(f);
@@ -349,7 +431,7 @@ class BridgedConfigTest {
tabLabel: "lead: {role} {profile}"
leaders:
opus:
terminal: term_opus
tab: "lead: opus"
""");
BridgedConfig cfg = BridgedConfig.load(f);
@@ -374,7 +456,7 @@ class BridgedConfigTest {
fleet:
leaders:
opus:
terminal: term_opus
tab: "lead: opus"
""");
assertDoesNotThrow(() -> BridgedConfig.load(f).validateLeadTabPrefixes());
@@ -400,7 +482,7 @@ class BridgedConfigTest {
"a label that collides with a convention nobody reads is not a problem");
}
// ── CB-530: the leaders registry ────────────────────────────────────────────────────────────
// ── CB-530/CB-579: the leaders registry ─────────────────────────────────────────────────────
@Test
void leadersBlockRegistersEveryPaneByName(@TempDir Path dir) throws Exception {
@@ -411,10 +493,10 @@ class BridgedConfigTest {
fleet:
leaders:
opus-5.0:
terminal: term_opus
tab: "lead: opus-5.0"
kind: claude
gpt-sol-5.6:
terminal: term_sol
tab: "lead: gpt-sol-5.6"
kind: opencode
model: openai/gpt-5.6-terra
""");
@@ -425,9 +507,9 @@ class BridgedConfigTest {
assertEquals(Set.of("opus-5.0", "gpt-sol-5.6"), leaders.keySet());
assertEquals("opencode", leaders.get("gpt-sol-5.6").kind());
assertEquals("openai/gpt-5.6-terra", leaders.get("gpt-sol-5.6").model());
// The whole point: BOTH panes resolve as leads, so neither is demoted to worker.
assertEquals(Map.of("term_opus", "opus-5.0", "term_sol", "gpt-sol-5.6"),
cfg.leaderTerminals());
// Identity is the tab now (CB-579) — both entries carry their own, distinct label.
assertEquals("lead: opus-5.0", leaders.get("opus-5.0").tab());
assertEquals("lead: gpt-sol-5.6", leaders.get("gpt-sol-5.6").tab());
}
@Test
@@ -439,42 +521,6 @@ class BridgedConfigTest {
"configs that never migrate must behave exactly as they did before CB-530");
}
@Test
void anExplicitLeadersEntryWinsOverThePinForTheSameTerminal(@TempDir Path dir) throws Exception {
Path f = dir.resolve("both.yaml");
Files.writeString(f, """
bind:
port: 8080
primary:
terminal: term_shared
fleet:
leaders:
opus-5.0:
terminal: term_shared
""");
assertEquals(Map.of("term_shared", "opus-5.0"), BridgedConfig.load(f).leaderTerminals(),
"the pin is the older spelling of the same fact; the named entry is what was meant");
}
@Test
void bothBlocksTogetherRegisterTheUnionOfTheirTerminals(@TempDir Path dir) throws Exception {
Path f = dir.resolve("union.yaml");
Files.writeString(f, """
bind:
port: 8080
primary:
terminal: term_pinned
fleet:
leaders:
gpt-sol-5.6:
terminal: term_sol
""");
assertEquals(Map.of("term_pinned", "primary", "term_sol", "gpt-sol-5.6"),
BridgedConfig.load(f).leaderTerminals());
}
@Test
void neitherBlockLeavesNothingRegistered(@TempDir Path dir) throws Exception {
Path f = dir.resolve("none.yaml");
@@ -483,22 +529,25 @@ class BridgedConfigTest {
assertTrue(BridgedConfig.load(f).leaderTerminals().isEmpty());
}
/** A lead entry with no terminal identifies nothing — it must not register a null key. */
/**
* CB-579: {@code fleet.leaders} no longer feeds {@code leaderTerminals()} at all — a lead's
* identity comes from the live tab scan, not a config-held terminal map. This method now exists
* only for the {@code primary.terminal} fallback.
*/
@Test
void aLeadWithoutATerminalIsNotRegistered(@TempDir Path dir) throws Exception {
Path f = dir.resolve("no-terminal.yaml");
void fleetLeadersNeverContributesToLeaderTerminals(@TempDir Path dir) throws Exception {
Path f = dir.resolve("leaders-only.yaml");
Files.writeString(f, """
bind:
port: 8080
fleet:
leaders:
sketch:
kind: opencode
real:
terminal: term_real
opus-5.0:
tab: "lead: opus-5.0"
""");
assertEquals(Map.of("term_real", "real"), BridgedConfig.load(f).leaderTerminals());
assertTrue(BridgedConfig.load(f).leaderTerminals().isEmpty(),
"no primary.terminal pin ⇒ nothing registered, even with fleet.leaders configured");
}
// ── CB-548: the architects registry ────────────────────────────────────────────────────────
@@ -991,6 +1040,44 @@ class BridgedConfigTest {
"an opencode worker with no argv defaults to the opencode binary, never claude");
}
/**
* CB-604: an unrecognized {@code kind:} used to silently fall into the claude-code bucket — not
* matching {@code "opencode"} was the only check. With {@code argv:} also unset that meant the
* daemon tried to launch a program literally named after the typo.
*/
@Test
void unknownKindIsRefusedAtLoadNamingTheValueAndTheAcceptedSet(@TempDir Path dir) throws Exception {
Path f = dir.resolve("kind-typo.yaml");
Files.writeString(f, """
profiles:
gemini:
kind: opencod
model: google/gemini-2.5-pro
""");
IllegalStateException e = assertThrows(IllegalStateException.class, () -> BridgedConfig.load(f));
assertTrue(e.getMessage().contains("gemini"), "error names the profile: " + e.getMessage());
assertTrue(e.getMessage().contains("opencod"), "error names the bad value: " + e.getMessage());
assertTrue(e.getMessage().contains("claude-code") && e.getMessage().contains("opencode"),
"error names the accepted set: " + e.getMessage());
}
@Test
void blankKindStillDefaultsToClaudeCode(@TempDir Path dir) throws Exception {
Path f = dir.resolve("kind-blank.yaml");
Files.writeString(f, """
profiles:
gx10:
baseUrl: http://gx10.gw:8000
kind: ""
argv: ["ccs", "gx10"]
""");
BridgedConfig cfg = BridgedConfig.load(f);
assertEquals(BridgedConfig.Profile.KIND_CLAUDE_CODE, cfg.profiles().get("gx10").kind(),
"a blank kind: is documented to behave exactly like an absent one");
}
@Test
void authDefaultsToLoopbackTrustSoExistingConfigsBehaveAsBefore(@TempDir Path dir) throws Exception {
Path f = dir.resolve("no-auth-block.yaml");
@@ -1092,6 +1179,7 @@ class BridgedConfigTest {
gitHostEnv: GITEA_HOST
weight: 0.5
maxLoad: 2
exhaustedPattern: "usage limit has been reached"
placement: weighted
lifecycle:
idleTtlSeconds: 300
@@ -1108,6 +1196,10 @@ class BridgedConfigTest {
idleAfterSeconds: 600
backoffMs: 45000
quietNudgeCap: 5
memberCredentials:
policy: deny-by-default
allow: [AI_GATEWAY_TOKEN]
known: [AI_GATEWAY_TOKEN, GITEA_ACCESS_TOKEN]
""");
BridgedConfig cfg = BridgedConfig.load(f);
@@ -1123,6 +1215,8 @@ class BridgedConfigTest {
assertEquals("GITEA_HOST", w.gitHostEnv());
assertEquals(0.5f, w.weight(), 0.0001f, "weight binds as a float");
assertEquals(2, w.maxLoad(), "maxLoad binds as an integer");
assertTrue(w.hasExhaustedPattern(), "exhaustedPattern binds and enables the CB-578 stage A classification");
assertEquals("usage limit has been reached", w.exhaustedPattern());
assertEquals("weighted", cfg.placement(), "placement binds at the top level");
assertEquals(300, cfg.lifecycle().idleTtlSeconds());
@@ -1136,6 +1230,84 @@ class BridgedConfigTest {
assertEquals(600, cfg.leadHeartbeat().idleAfterSeconds(), "leadHeartbeat binds at the top level");
assertEquals(45_000L, cfg.leadHeartbeat().backoffMs());
assertEquals(5, cfg.leadHeartbeat().quietNudgeCap());
assertEquals(BridgedConfig.MemberCredentials.POLICY_DENY_BY_DEFAULT, cfg.memberCredentials().policy(),
"memberCredentials binds at the top level");
assertEquals(Set.of("AI_GATEWAY_TOKEN"), cfg.memberCredentials().allowSet());
assertEquals(Set.of("GITEA_ACCESS_TOKEN"), cfg.memberCredentials().blockedSet());
}
/**
* A top-level key {@code BridgedConfig} reads but that appears nowhere in
* {@code bridged.example.yaml} — live or commented — is invisible drift: {@code bridged.yaml}
* is gitignored, so the example is the ONLY committed description of the config schema, and
* neither {@link #shippedExampleConfigParses} (example → code: does the example still parse)
* nor {@link #everyOptionalKnobDocumentedInTheExampleBinds} (a hand-maintained list of keys
* that must bind) can catch a brand-new key nobody added to either.
*
* <p>This test compares the OTHER direction: every key in {@link BridgedConfig#KNOWN_TOP_LEVEL_KEYS}
* (the parser's own accepted set, which backs the unknown-key WARN) must appear as a top-level
* key in the example text, live or commented-out — see {@link #topLevelKeyDocumented}.
*/
@Test
void everyKnownTopLevelKeyIsDocumentedInTheExample() throws Exception {
Path example = Path.of("bridged.example.yaml");
assertTrue(Files.exists(example), "bridged.example.yaml must ship next to the pom");
String text = Files.readString(example);
List<String> undocumented = BridgedConfig.KNOWN_TOP_LEVEL_KEYS.stream()
.filter(key -> !topLevelKeyDocumented(text, key))
.sorted()
.toList();
assertTrue(undocumented.isEmpty(), () -> "key(s) " + undocumented
+ " are read by BridgedConfig but appear nowhere in bridged.example.yaml — "
+ "document each one there, commented out if optional. bridged.yaml is "
+ "gitignored, so this file is the only committed description of the config "
+ "schema an operator or a worker can see.");
}
/**
* Most of {@code bridged.example.yaml} is deliberately commented out — optional sections are
* documented as commented blocks so the shipped file stays a working minimal config. A key
* documented ONLY as a comment must still count as documented; parsing the file as YAML and
* reading its live key set (as an earlier attempt at this guard did) gets this wrong, because
* every commented section then looks entirely absent.
*/
@Test
void commentedOnlyTopLevelKeyCountsAsDocumented() {
String yaml = """
bind:
port: 8765
# broker:
# uri: amqp://guest:guest@127.0.0.1:5672
""";
assertTrue(topLevelKeyDocumented(yaml, "broker"),
"a key documented only inside a commented-out block must still count as documented");
}
/** A key that appears in neither a live nor a commented top-level line must NOT count. */
@Test
void absentTopLevelKeyIsNotDocumented() {
String yaml = """
bind:
port: 8765
""";
assertFalse(topLevelKeyDocumented(yaml, "broker"),
"a key mentioned nowhere in the example must not be reported as documented");
}
/**
* True when {@code key} appears as a top-level YAML key in {@code yaml} — either live
* ({@code key:} at column 0) or commented out ({@code # key:}, also at column 0, with only
* whitespace between the {@code #} and the key). Anchoring on column 0 is what keeps this a
* top-level check: an indented occurrence (a nested field, or prose inside a comment that
* happens to end in a colon) never matches, because {@code ^} requires the key's own first
* character — or the sole leading {@code #} — to sit at the very start of the line.
*/
private static boolean topLevelKeyDocumented(String yaml, String key) {
Pattern p = Pattern.compile("(?m)^(?:#\\s*)?" + Pattern.quote(key) + ":");
return p.matcher(yaml).find();
}
@Test
@@ -1170,6 +1342,232 @@ class BridgedConfigTest {
assertNull(w.maxLoad(), "absent maxLoad defaults to unlimited (null)");
}
@Test
void explicitWeightZeroStaysZeroInsteadOfCoercingToOne(@TempDir Path dir) throws Exception {
// CB-554: weight: 0 used to be normalised to 1.0 by this same compact constructor, which
// made it read as "never auto-select" while behaving as "select like anyone else."
Path f = dir.resolve("weight-zero.yaml");
Files.writeString(f, """
bind:
port: 8080
profiles:
opus:
baseUrl: http://gx10.gw:8000
weight: 0
""");
BridgedConfig cfg = BridgedConfig.load(f);
BridgedConfig.Profile w = cfg.profiles().get("opus");
assertEquals(0.0f, w.weight(), 0.0001f, "an explicit weight: 0 must stay 0, not coerce to 1.0");
}
@Test
void negativeWeightNormalisesToZeroNotOne(@TempDir Path dir) throws Exception {
Path f = dir.resolve("weight-negative.yaml");
Files.writeString(f, """
bind:
port: 8080
profiles:
gx10:
baseUrl: http://gx10.gw:8000
weight: -3.0
""");
BridgedConfig cfg = BridgedConfig.load(f);
BridgedConfig.Profile w = cfg.profiles().get("gx10");
assertEquals(0.0f, w.weight(), 0.0001f, "a negative weight behaves as excluded (0), not as an error and not as 1.0");
}
@Test
void explicitMaxLoadZeroStaysZeroInsteadOfCoercingToUnlimited(@TempDir Path dir) throws Exception {
// CB-585: maxLoad: 0 used to be normalised to null (unlimited) by this same compact
// constructor, so an operator writing it to mean "never run anything here" got the exact
// opposite. On a subscription: true profile, maxLoad is the only throttle against the
// operator's own paid plan.
Path f = dir.resolve("maxload-zero.yaml");
Files.writeString(f, """
bind:
port: 8080
profiles:
opus:
baseUrl: http://gx10.gw:8000
maxLoad: 0
""");
BridgedConfig cfg = BridgedConfig.load(f);
BridgedConfig.Profile w = cfg.profiles().get("opus");
assertEquals(0, w.maxLoad(), "an explicit maxLoad: 0 must stay 0, not coerce to unlimited (null)");
}
@Test
void negativeMaxLoadIsRefusedAtLoadNamingTheProfileAndKey(@TempDir Path dir) throws Exception {
Path f = dir.resolve("maxload-negative.yaml");
Files.writeString(f, """
bind:
port: 8080
profiles:
opus:
baseUrl: http://gx10.gw:8000
maxLoad: -2
""");
IllegalStateException e = assertThrows(IllegalStateException.class, () -> BridgedConfig.load(f));
assertTrue(e.getMessage().contains("opus"), "error names the profile: " + e.getMessage());
assertTrue(e.getMessage().contains("maxLoad"), "error names the key: " + e.getMessage());
}
/**
* CB-606: an unrecognized {@code auth.mode} used to silently fall back to
* {@code loopback-trust} — {@link BridgedConfig.Auth#tokenMode()} only checked equality
* against {@code "token"}. On a loopback bind {@link BridgedConfig#validateAuthExposure()}
* never runs (it only fires for a non-loopback bind), so the typo was completely invisible:
* the daemon started cleanly and authenticated nobody while the operator believed token mode
* was active.
*/
@Test
void unknownAuthModeIsRefusedAtLoadNamingTheValueAndTheAcceptedSet(@TempDir Path dir) throws Exception {
Path f = dir.resolve("auth-mode-typo.yaml");
Files.writeString(f, """
bind:
host: 127.0.0.1
port: 8765
auth:
mode: toekn
""");
IllegalStateException e = assertThrows(IllegalStateException.class, () -> BridgedConfig.load(f));
assertTrue(e.getMessage().contains("toekn"), "error names the bad value: " + e.getMessage());
assertTrue(e.getMessage().contains("loopback-trust") && e.getMessage().contains("token"),
"error names the accepted set: " + e.getMessage());
}
/**
* CB-606: an unrecognized per-profile {@code placement:} used to silently fall back to legacy
* pane placement — {@link BridgedConfig.Profile#tabPlacement()} only checked equality against
* {@code "tab"}.
*/
@Test
void unknownProfilePlacementIsRefusedAtLoadNamingTheProfileAndTheAcceptedSet(@TempDir Path dir) throws Exception {
Path f = dir.resolve("placement-typo.yaml");
Files.writeString(f, """
profiles:
gx10:
baseUrl: http://gx10.gw:8000
placement: tabb
""");
IllegalStateException e = assertThrows(IllegalStateException.class, () -> BridgedConfig.load(f));
assertTrue(e.getMessage().contains("gx10"), "error names the profile: " + e.getMessage());
assertTrue(e.getMessage().contains("tabb"), "error names the bad value: " + e.getMessage());
assertTrue(e.getMessage().contains("tab") && e.getMessage().contains("pane"),
"error names the accepted set: " + e.getMessage());
}
/**
* CB-596: {@code memberCredentials.policy} is validated the same way {@code auth.mode} and
* per-profile {@code placement} are (CB-606's pattern) — a typo must not silently behave as the
* one real policy, because the day a second policy exists that silent fallback becomes a real
* behavior change instead of a happy accident.
*/
@Test
void unknownMemberCredentialsPolicyIsRefusedAtLoadNamingTheValueAndTheAcceptedSet(@TempDir Path dir) throws Exception {
Path f = dir.resolve("member-credentials-policy-typo.yaml");
Files.writeString(f, """
bind:
host: 127.0.0.1
port: 8765
memberCredentials:
policy: deny-by-defualt
known:
- GITEA_ACCESS_TOKEN
""");
IllegalStateException e = assertThrows(IllegalStateException.class, () -> BridgedConfig.load(f));
assertTrue(e.getMessage().contains("deny-by-defualt"), "error names the bad value: " + e.getMessage());
assertTrue(e.getMessage().contains("deny-by-default"), "error names the accepted set: " + e.getMessage());
}
/** {@code allow}/{@code known} bind and {@link BridgedConfig.MemberCredentials#blockedSet()} is known minus allow. */
@Test
void memberCredentialsBindsAllowAndKnownAndComputesBlockedSet(@TempDir Path dir) throws Exception {
Path f = dir.resolve("member-credentials.yaml");
Files.writeString(f, """
bind:
port: 8080
memberCredentials:
policy: deny-by-default
allow:
- AI_GATEWAY_TOKEN
- WORKER_GITEA_TOKEN
known:
- AI_GATEWAY_TOKEN
- WORKER_GITEA_TOKEN
- GITEA_ACCESS_TOKEN
- GITLAB_PERSONAL_ACCESS_TOKEN
""");
BridgedConfig.MemberCredentials mc = BridgedConfig.load(f).memberCredentials();
assertEquals(BridgedConfig.MemberCredentials.POLICY_DENY_BY_DEFAULT, mc.policy());
assertEquals(Set.of("AI_GATEWAY_TOKEN", "WORKER_GITEA_TOKEN"), mc.allowSet());
assertEquals(Set.of("GITEA_ACCESS_TOKEN", "GITLAB_PERSONAL_ACCESS_TOKEN"), mc.blockedSet(),
"blockedSet is known minus allow");
}
/**
* CB-596: omitting {@code memberCredentials:} entirely must NOT crash a reader that assumes a
* non-null block (the same "fill in nested defaults" contract every other structural field
* gets — see {@link BridgedConfig#withDefaults()}), but it also must not pretend anything is
* blocked: an empty {@code known} list blocks nothing, and that is a real gap the operator must
* close by configuring this block, not a safe default.
*/
@Test
void absentMemberCredentialsDefaultsToAnEmptyNonNullBlock(@TempDir Path dir) throws Exception {
Path f = dir.resolve("member-credentials-absent.yaml");
Files.writeString(f, "bind:\n port: 8080\n");
BridgedConfig.MemberCredentials mc = BridgedConfig.load(f).memberCredentials();
assertNotNull(mc, "withDefaults() must never leave this null");
assertTrue(mc.known().isEmpty(), "no known list configured — nothing is blocked");
assertTrue(mc.allowSet().isEmpty());
assertTrue(mc.blockedSet().isEmpty());
}
@Test
void absentProfilePlacementDefaultsToTab(@TempDir Path dir) throws Exception {
Path f = dir.resolve("placement-absent.yaml");
Files.writeString(f, """
profiles:
gx10:
baseUrl: http://gx10.gw:8000
""");
BridgedConfig cfg = BridgedConfig.load(f);
BridgedConfig.Profile w = cfg.profiles().get("gx10");
assertTrue(w.tabPlacement(), "an absent placement must keep defaulting to tab");
}
/**
* CB-606: the top-level {@code placement:} policy name WAS validated, but only lazily, by
* {@code PlacementPolicies.fromName} through {@code CompositePeerLauncher}'s per-spawn
* {@code Supplier} — so a bad name still started a daemon that looked healthy and failed only
* the first time something spawned without naming a profile. This must now fail at load.
*/
@Test
void unknownTopLevelPlacementPolicyIsRefusedAtLoadNotLazilyAtFirstSpawn(@TempDir Path dir) throws Exception {
Path f = dir.resolve("placement-policy-typo.yaml");
Files.writeString(f, """
bind:
port: 8080
placement: weightd
""");
IllegalStateException e = assertThrows(IllegalStateException.class, () -> BridgedConfig.load(f));
assertTrue(e.getMessage().contains("weightd"), "error names the bad value: " + e.getMessage());
assertTrue(e.getMessage().contains("fixed") && e.getMessage().contains("round-robin")
&& e.getMessage().contains("weighted"),
"error names the accepted set: " + e.getMessage());
}
@Test
void subscriptionFlagBindsAndDefaultsFalse(@TempDir Path dir) throws Exception {
Path f = dir.resolve("subscription.yaml");
@@ -338,6 +338,54 @@ class ConfigRefTest {
assertEquals("deepseek-v4-flash", ref.get().profiles().get("sonnet").model());
}
/**
* CB-578 stage B: exhaustedPattern is compiled once into Bridged.main's pattern map at startup
* (see ExhaustedPatternLookup), so a reload never re-reads it — a changed pattern must be
* reported deferred exactly like model/baseUrl, not silently claimed as applied.
*/
@Test
void changingAProfilesExhaustedPatternIsReportedAsDeferred(@TempDir Path dir) throws Exception {
Path f = dir.resolve("bridged.yaml");
Files.writeString(f, """
bind:
host: 127.0.0.1
port: 8765
herdrSocket: ~/.config/herdr/herdr.sock
profiles:
sonnet:
baseUrl: http://gx00.gw:8000
model: sonnet
exhaustedPattern: "usage limit has been reached"
guard:
offSubscriptionHosts:
- gx00.gw
""");
ConfigRef ref = refFor(f);
Files.writeString(f, """
bind:
host: 127.0.0.1
port: 8765
herdrSocket: ~/.config/herdr/herdr.sock
profiles:
sonnet:
baseUrl: http://gx00.gw:8000
model: sonnet
exhaustedPattern: "rate limit exceeded"
guard:
offSubscriptionHosts:
- gx00.gw
""");
ConfigRef.Outcome out = ref.reload();
assertTrue(out.applied());
assertEquals(1, out.deferred().size(), out.deferred().toString());
assertTrue(out.deferred().getFirst().contains("sonnet"), out.deferred().toString());
assertTrue(out.deferred().getFirst().contains("launch settings"), out.deferred().toString());
// The snapshot still carries the new value — a restart is what makes it take effect.
assertEquals("rate limit exceeded", ref.get().profiles().get("sonnet").exhaustedPattern());
}
@Test
void aFixedRefHasNoFileAndRefusesToReload() {
BridgedConfig cfg = new BridgedConfig(null, null, null, null, null, null,
@@ -0,0 +1,170 @@
package dev.ltms.bridged.health;
import ch.qos.logback.classic.Logger;
import ch.qos.logback.classic.spi.ILoggingEvent;
import ch.qos.logback.core.read.ListAppender;
import dev.ltms.bridged.herdr.AgentControl;
import dev.ltms.bridged.herdr.FakeHerdr;
import dev.ltms.bridged.inject.Injector;
import dev.ltms.bridged.msg.InMemoryReplyInbox;
import dev.ltms.bridged.msg.MessageService;
import dev.ltms.bridged.msg.Rendezvous;
import dev.ltms.bridged.session.SessionManager;
import dev.ltms.bridged.member.ClaudeCodeLauncher;
import dev.ltms.bridged.herdr.WorkspaceControl;
import dev.ltms.bridged.guard.SubscriptionGuard;
import dev.ltms.bridged.config.BridgedConfig;
import org.junit.jupiter.api.Test;
import org.slf4j.LoggerFactory;
import java.util.Map;
import java.util.Set;
import java.util.concurrent.Executors;
import java.util.function.BiConsumer;
import static org.junit.jupiter.api.Assertions.assertEquals;
import static org.junit.jupiter.api.Assertions.assertTrue;
class FleetHealthMonitorTest {
@Test void oneTickUsesOneFleetListForAnyRosterSize() {
FakeHerdr herdr = new FakeHerdr();
AgentControl agents = new AgentControl(herdr);
BridgedConfig.Profile profile = new BridgedConfig.Profile("test", "http://test:1", null,
null, null, null, null, null, null, null, null, null);
ClaudeCodeLauncher launcher = new ClaudeCodeLauncher(agents, new WorkspaceControl(herdr),
new SubscriptionGuard(Set.of("test")), Map.of("test", profile), "test", _ -> "token");
SessionManager sessions = new SessionManager(launcher);
sessions.acquire("test", null, null, null);
sessions.acquire("test", null, null, null);
herdr.calls.clear();
MessageService messages = new MessageService(agents, new Injector(agents), new Rendezvous(), new InMemoryReplyInbox());
var scheduler = Executors.newSingleThreadScheduledExecutor();
FleetHealthMonitor monitor = new FleetHealthMonitor(agents, sessions::roster, messages, scheduler, () -> 1, 60,
(_, _) -> { });
monitor.tick();
monitor.stop();
assertEquals(1, herdr.calls.stream().filter(call -> call.method().equals("agent.list")).count());
}
@Test void failedTickDoesNotStopTheNextTick() {
FakeHerdr herdr = new FakeHerdr().healthy(false);
AgentControl agents = new AgentControl(herdr);
var scheduler = Executors.newSingleThreadScheduledExecutor();
FleetHealthMonitor monitor = new FleetHealthMonitor(agents, java.util.List::of,
new MessageService(agents, new Injector(agents), new Rendezvous(), new InMemoryReplyInbox()),
scheduler, () -> 1, 60, (_, _) -> { });
monitor.tick();
herdr.healthy(true);
monitor.tick();
monitor.stop();
assertEquals(2, herdr.calls.stream().filter(call -> call.method().equals("agent.list")).count());
}
@Test void faultTransitionLogsOnlyOnceUntilItChanges() {
Logger logger = (Logger) LoggerFactory.getLogger(FleetHealthMonitor.class);
ListAppender<ILoggingEvent> appender = new ListAppender<>();
appender.start();
logger.addAppender(appender);
try {
FakeHerdr herdr = new FakeHerdr();
AgentControl agents = new AgentControl(herdr);
var scheduler = Executors.newSingleThreadScheduledExecutor();
FleetHealthMonitor monitor = new FleetHealthMonitor(agents, java.util.List::of,
new MessageService(agents, new Injector(agents), new Rendezvous(), new InMemoryReplyInbox()),
scheduler, () -> 1, 60, (_, _) -> { });
monitor.reportTransition("term_a", HealthState.TURN_BOUNDARY_LOST);
monitor.reportTransition("term_a", HealthState.TURN_BOUNDARY_LOST);
monitor.stop();
assertEquals(1, appender.list.stream().filter(event -> event.getFormattedMessage()
.contains("member=term_a state=TURN_BOUNDARY_LOST")).count());
} finally {
logger.detachAppender(appender);
}
}
// --- CB-580: a member that reaches GONE/NEVER_READY must fail its waiting tickets
private static FleetHealthMonitor monitorWith(BiConsumer<String, String> failTarget) {
FakeHerdr herdr = new FakeHerdr();
AgentControl agents = new AgentControl(herdr);
var scheduler = Executors.newSingleThreadScheduledExecutor();
return new FleetHealthMonitor(agents, java.util.List::of,
new MessageService(agents, new Injector(agents), new Rendezvous(), new InMemoryReplyInbox()),
scheduler, () -> 1, 60, failTarget);
}
@Test void terminalTransitionFailsTheTargetOnce() {
RecordingFailTarget failTarget = new RecordingFailTarget();
FleetHealthMonitor monitor = monitorWith(failTarget);
monitor.reportTransition("term_a", HealthState.GONE);
monitor.stop();
assertEquals(1, failTarget.calls.size());
assertEquals("term_a", failTarget.calls.get(0).target());
assertTrue(failTarget.calls.get(0).reason().contains("GONE"));
}
@Test void neverReadyNamesItselfAsTheReason() {
RecordingFailTarget failTarget = new RecordingFailTarget();
FleetHealthMonitor monitor = monitorWith(failTarget);
monitor.reportTransition("term_a", HealthState.NEVER_READY);
monitor.stop();
assertEquals(1, failTarget.calls.size());
assertTrue(failTarget.calls.get(0).reason().contains("NEVER_READY"));
}
@Test void stayingInATerminalStateProducesOneFailureNotN() {
RecordingFailTarget failTarget = new RecordingFailTarget();
FleetHealthMonitor monitor = monitorWith(failTarget);
monitor.reportTransition("term_a", HealthState.GONE);
monitor.reportTransition("term_a", HealthState.GONE);
monitor.reportTransition("term_a", HealthState.GONE);
monitor.reportTransition("term_a", HealthState.GONE);
monitor.stop();
assertEquals(1, failTarget.calls.size());
}
@Test void aNonTerminalFaultStateDoesNotFailTheTarget() {
RecordingFailTarget failTarget = new RecordingFailTarget();
FleetHealthMonitor monitor = monitorWith(failTarget);
monitor.reportTransition("term_a", HealthState.TURN_BOUNDARY_LOST);
monitor.stop();
assertEquals(0, failTarget.calls.size());
}
@Test void failTargetRetryIsBounded() {
AlwaysThrowingFailTarget failTarget = new AlwaysThrowingFailTarget();
FleetHealthMonitor monitor = monitorWith(failTarget);
monitor.reportTransition("term_a", HealthState.GONE);
monitor.stop();
assertEquals(FleetHealthMonitor.MAX_FAIL_TARGET_ATTEMPTS, failTarget.calls);
}
@Test void exhaustedRetryStillDoesNotRefireOnAnUnchangedTick() {
AlwaysThrowingFailTarget failTarget = new AlwaysThrowingFailTarget();
FleetHealthMonitor monitor = monitorWith(failTarget);
monitor.reportTransition("term_a", HealthState.GONE);
int afterFirstTransition = failTarget.calls;
monitor.reportTransition("term_a", HealthState.GONE);
monitor.stop();
assertEquals(afterFirstTransition, failTarget.calls);
}
private record RecordedCall(String target, String reason) { }
private static final class RecordingFailTarget implements BiConsumer<String, String> {
final java.util.List<RecordedCall> calls = new java.util.ArrayList<>();
@Override public void accept(String target, String reason) {
calls.add(new RecordedCall(target, reason));
}
}
private static final class AlwaysThrowingFailTarget implements BiConsumer<String, String> {
int calls = 0;
@Override public void accept(String target, String reason) {
calls++;
throw new RuntimeException("boom");
}
}
}
@@ -7,6 +7,7 @@ import java.util.ArrayList;
import java.util.LinkedHashMap;
import java.util.List;
import java.util.Map;
import java.util.concurrent.CopyOnWriteArrayList;
/**
* Recording fake {@link HerdrClient} for unit/acceptance tests. Returns canned frames
@@ -22,7 +23,13 @@ public final class FakeHerdr implements HerdrClient {
public static final long WORKER_PID = 4242;
private final ObjectMapper mapper = new ObjectMapper();
public final List<Call> calls = new ArrayList<>();
/**
* Thread-safe on purpose. Background loops — {@link dev.ltms.bridged.msg.ReplyPushLoop} and the
* lead heartbeat — call this fake from their own scheduler threads while a test polls
* {@link #called} from the test thread. A plain {@code ArrayList} threw
* {@code ConcurrentModificationException} out of {@code called()} when a nudge landed mid-stream.
*/
public final List<Call> calls = new CopyOnWriteArrayList<>();
private boolean healthy = true;
private final List<String> extraWorkspaces = new ArrayList<>();
private final List<String> extraAgents = new ArrayList<>();
@@ -15,8 +15,9 @@ import java.util.concurrent.atomic.AtomicLong;
import static org.junit.jupiter.api.Assertions.*;
/**
* CB-531. A lead is never spawned, so the daemon has to <em>find</em> it: these assert that an
* operator-labelled tab is what makes a pane a lead, and — just as importantly — what does not.
* CB-531/CB-579. A lead is never spawned, so the daemon has to <em>find</em> it: these assert that
* an operator-labelled tab matching a configured {@code tab:} is what makes a pane a lead, and —
* just as importantly — what does not, and that a stale entry does not linger forever.
*/
class LeadTabScannerTest {
@@ -120,23 +121,48 @@ class LeadTabScannerTest {
.pane("w9:p1", "w9:t1", "term_worker");
}
private LeadTabScanner scanner(TopologyHerdr herdr, Map<String, String> configured,
/** The {@code tab:} → name map {@code twoLeads()}'s two lead tabs are configured under. */
private static Map<String, String> twoLeadsConfigured() {
return Map.of("lead: opus-5.0", "opus-5.0", "lead: gpt-sol-5.6", "gpt-sol-5.6");
}
private LeadTabScanner scanner(TopologyHerdr herdr, Map<String, String> tabToName,
AtomicLong clock) {
return new LeadTabScanner(herdr, "lead:", Set.of("bridged-workers"), configured, TTL,
clock::get);
return new LeadTabScanner(herdr, tabToName, Set.of("bridged-workers"), TTL, clock::get);
}
@Test
void everyLabelledTabBecomesALeadNamedByItsLabel() {
Map<String, String> leads = scanner(twoLeads(), Map.of(), new AtomicLong()).get();
void everyConfiguredTabBecomesALeadNamedByItsEntry() {
Map<String, String> leads = scanner(twoLeads(), twoLeadsConfigured(), new AtomicLong()).get();
assertEquals(Map.of("term_opus", "opus-5.0", "term_gpt", "gpt-sol-5.6"), leads,
"two leads discovered from labels alone — no terminal_id was ever configured");
"two leads discovered by their configured tab — no terminal_id was ever configured");
}
@Test
void anUnlabelledTabContributesNothing() {
assertFalse(scanner(twoLeads(), Map.of(), new AtomicLong()).get().containsKey("term_notes"));
void anUnconfiguredTabContributesNothing() {
assertFalse(scanner(twoLeads(), twoLeadsConfigured(), new AtomicLong())
.get().containsKey("term_notes"));
}
/**
* CB-579: matching is exact against the configured map now, not a shared prefix — two leads with
* completely different labels are both discovered by one scanner, no convention required.
*/
@Test
void twoLeadsWithCompletelyDifferentLabelsAreBothDiscovered() {
TopologyHerdr herdr = new TopologyHerdr()
.workspace("w1", "main")
.tab("w1:t1", "w1", "orchestrator: opus")
.tab("w1:t2", "w1", "captain: sol")
.pane("w1:p1", "w1:t1", "term_opus")
.pane("w1:p2", "w1:t2", "term_sol");
Map<String, String> tabToName = Map.of("orchestrator: opus", "opus", "captain: sol", "sol");
Map<String, String> leads = scanner(herdr, tabToName, new AtomicLong()).get();
assertEquals(Map.of("term_opus", "opus", "term_sol", "sol"), leads,
"no shared prefix needed — each lead is matched by its own configured tab");
}
/**
@@ -148,25 +174,28 @@ class LeadTabScannerTest {
void aTabInAWorkerSpaceIsNeverALeadEvenWhenItsLabelMatches() {
TopologyHerdr herdr = twoLeads().tab("w9:t2", "w9", "lead: impostor")
.pane("w9:p2", "w9:t2", "term_impostor");
Map<String, String> tabToName = new LinkedHashMap<>(twoLeadsConfigured());
tabToName.put("lead: impostor", "impostor");
assertFalse(scanner(herdr, Map.of(), new AtomicLong()).get().containsKey("term_impostor"));
assertFalse(scanner(herdr, tabToName, new AtomicLong()).get().containsKey("term_impostor"));
}
@Test
void aBarePrefixNamesNobodyAndIsRejected() {
void aLabelWithNoConfiguredEntryIsIgnored() {
TopologyHerdr herdr = new TopologyHerdr().workspace("w1", "main")
.tab("w1:t1", "w1", "lead:").pane("w1:p1", "w1:t1", "term_a");
.tab("w1:t1", "w1", "lead: nobody-configured").pane("w1:p1", "w1:t1", "term_a");
assertEquals(Map.of(), scanner(herdr, Map.of(), new AtomicLong()).get(),
"a lead with no name would resolve as PRIMARY with nothing to attribute it to");
assertEquals(Map.of(), scanner(herdr, twoLeadsConfigured(), new AtomicLong()).get(),
"a label that names no configured lead resolves nobody");
}
@Test
void thePrefixMatchesCaseInsensitivelyAndTheNameIsTrimmed() {
void matchingIsCaseInsensitiveAndToleratesSurroundingWhitespace() {
TopologyHerdr herdr = new TopologyHerdr().workspace("w1", "main")
.tab("w1:t1", "w1", " LEAD: opus-5.0 ").pane("w1:p1", "w1:t1", "term_a");
.tab("w1:t1", "w1", " LEAD: Opus-5.0 ").pane("w1:p1", "w1:t1", "term_a");
assertEquals(Map.of("term_a", "opus-5.0"), scanner(herdr, Map.of(), new AtomicLong()).get());
assertEquals(Map.of("term_a", "opus-5.0"),
scanner(herdr, Map.of("lead: Opus-5.0", "opus-5.0"), new AtomicLong()).get());
}
@Test
@@ -175,18 +204,51 @@ class LeadTabScannerTest {
// nothing bridged placed can land here (see the worker-space test above).
TopologyHerdr herdr = twoLeads().pane("w1:p1b", "w1:t1", "term_opus_split");
assertEquals("opus-5.0", scanner(herdr, Map.of(), new AtomicLong()).get().get("term_opus_split"));
assertEquals("opus-5.0",
scanner(herdr, twoLeadsConfigured(), new AtomicLong()).get().get("term_opus_split"));
}
/**
* CB-579 acceptance (6): this is the bug the ticket closes. A stale pin used to be merged back
* over every scan and never expire; now a scan is the whole answer, so a lead whose tab is gone
* drops out on the very next scan.
*/
@Test
void anExplicitlyConfiguredLeadIsMergedInAndOutranksALabel() {
Map<String, String> configured = Map.of("term_opus", "pinned-name", "term_extra", "from-config");
void aTabNoLongerPresentDropsTheLeadOnTheNextScan() {
TopologyHerdr herdr = twoLeads();
AtomicLong clock = new AtomicLong();
LeadTabScanner s = scanner(herdr, twoLeadsConfigured(), clock);
assertTrue(s.get().containsKey("term_opus"));
Map<String, String> leads = scanner(twoLeads(), configured, new AtomicLong()).get();
// The session behind term_opus restarted — herdr no longer reports that tab or pane at all.
herdr.tabs.remove("w1:t1");
herdr.panes.remove("w1:p1");
clock.addAndGet(TTL);
assertEquals("pinned-name", leads.get("term_opus"), "an explicit pin is the operator's last word");
assertEquals("from-config", leads.get("term_extra"), "a configured lead needs no tab at all");
assertEquals("gpt-sol-5.6", leads.get("term_gpt"));
assertFalse(s.get().containsKey("term_opus"),
"a stale entry must expire once the tab it named is gone, not be merged back forever");
}
/**
* CB-579 acceptance (5): the whole point of matching by tab instead of {@code terminal_id} — a
* restart changes the terminal, not the tab, so the lead resolves under the same name with no
* config edit.
*/
@Test
void aLeadRestartingInTheSameTabResolvesUnderTheSameName() {
TopologyHerdr herdr = twoLeads();
AtomicLong clock = new AtomicLong();
LeadTabScanner s = scanner(herdr, twoLeadsConfigured(), clock);
assertEquals("opus-5.0", s.get().get("term_opus"));
// The session restarts: herdr assigns the pane a new terminal_id, same tab (w1:t1).
herdr.panes.remove("w1:p1");
herdr.pane("w1:p1", "w1:t1", "term_opus_v2");
clock.addAndGet(TTL);
Map<String, String> leads = s.get();
assertEquals("opus-5.0", leads.get("term_opus_v2"), "the new terminal resolves immediately");
assertFalse(leads.containsKey("term_opus"), "the old terminal_id is simply gone, not carried");
}
// ── caching ─────────────────────────────────────────────────────────────────────────────────
@@ -195,7 +257,7 @@ class LeadTabScannerTest {
void aSecondLookupWithinTheTtlDoesNotTouchHerdr() {
TopologyHerdr herdr = twoLeads();
AtomicLong clock = new AtomicLong();
LeadTabScanner s = scanner(herdr, Map.of(), clock);
LeadTabScanner s = scanner(herdr, twoLeadsConfigured(), clock);
s.get();
int afterFirst = herdr.calls;
@@ -210,21 +272,23 @@ class LeadTabScannerTest {
void aTabLabelledAfterStartupIsPickedUpOnceTheTtlExpires() {
TopologyHerdr herdr = twoLeads();
AtomicLong clock = new AtomicLong();
LeadTabScanner s = scanner(herdr, Map.of(), clock);
Map<String, String> tabToName = new LinkedHashMap<>(twoLeadsConfigured());
tabToName.put("lead: late-arrival", "late-arrival");
LeadTabScanner s = scanner(herdr, tabToName, clock);
assertFalse(s.get().containsKey("term_notes"));
herdr.tab("w1:t3", "w1", "lead: late-arrival"); // the operator renames their tab
clock.addAndGet(TTL);
assertEquals("late-arrival", s.get().get("term_notes"),
"the whole point over `leaders:`: no config edit, no restart");
"the whole point over a config-held terminal_id: no config edit, no restart");
}
@Test
void aFailedScanKeepsTheLeadsAlreadyKnownRatherThanDemotingThem() {
TopologyHerdr herdr = twoLeads();
AtomicLong clock = new AtomicLong();
LeadTabScanner s = scanner(herdr, Map.of(), clock);
LeadTabScanner s = scanner(herdr, twoLeadsConfigured(), clock);
Map<String, String> before = s.get();
herdr.failing = true;
@@ -235,14 +299,14 @@ class LeadTabScannerTest {
}
@Test
void aFailedFirstScanStillHonoursTheConfiguredLeads() {
void aFailedFirstScanReturnsEmptyRatherThanThrowing() {
TopologyHerdr herdr = twoLeads();
herdr.failing = true;
Map<String, String> leads = scanner(herdr, Map.of("term_x", "opus-5.0"), new AtomicLong()).get();
Map<String, String> leads = scanner(herdr, twoLeadsConfigured(), new AtomicLong()).get();
assertEquals(Map.of("term_x", "opus-5.0"), leads,
"config-named leads must not depend on herdr answering at all");
assertEquals(Map.of(), leads,
"with nothing scanned yet and no override to fall back on, the map is simply empty");
}
@Test
@@ -250,7 +314,7 @@ class LeadTabScannerTest {
TopologyHerdr herdr = twoLeads();
herdr.failing = true;
AtomicLong clock = new AtomicLong();
LeadTabScanner s = scanner(herdr, Map.of(), clock);
LeadTabScanner s = scanner(herdr, twoLeadsConfigured(), clock);
s.get();
int afterFirst = herdr.calls;
@@ -7,9 +7,14 @@ import ch.qos.logback.core.read.ListAppender;
import dev.ltms.bridged.herdr.AgentControl;
import dev.ltms.bridged.herdr.FakeHerdr;
import dev.ltms.bridged.msg.Rendezvous;
import dev.ltms.bridged.msg.TestTurnTokens;
import dev.ltms.bridged.msg.TurnToken;
import org.junit.jupiter.api.Test;
import org.slf4j.LoggerFactory;
import java.util.Set;
import java.util.regex.Pattern;
import static org.junit.jupiter.api.Assertions.assertEquals;
import static org.junit.jupiter.api.Assertions.assertFalse;
import static org.junit.jupiter.api.Assertions.assertTrue;
@@ -21,7 +26,7 @@ class CompletionResolverTest {
void skipsTheScrapeWhenNoSendIsWaiting() {
FakeHerdr herdr = new FakeHerdr();
Rendezvous rendezvous = new Rendezvous(); // no waiter opened
CompletionResolver resolver = new CompletionResolver(new AgentControl(herdr), rendezvous);
CompletionResolver resolver = new CompletionResolver(new AgentControl(herdr), rendezvous, ExhaustedPatternLookup.none(), ExhaustionSink.none());
resolver.resolve("term_a", null); // no in-flight turn captured for this target
@@ -33,7 +38,7 @@ class CompletionResolverTest {
void failSkipsTheScrapeWhenNoSendIsWaiting() {
FakeHerdr herdr = new FakeHerdr();
Rendezvous rendezvous = new Rendezvous(); // no waiter opened
CompletionResolver resolver = new CompletionResolver(new AgentControl(herdr), rendezvous);
CompletionResolver resolver = new CompletionResolver(new AgentControl(herdr), rendezvous, ExhaustedPatternLookup.none(), ExhaustionSink.none());
resolver.fail("term_a", null); // no in-flight turn, and no registered waiter to fall back to
@@ -45,9 +50,9 @@ class CompletionResolverTest {
void captureBaselineSkipsTheReadWhenNoSendIsWaiting() {
FakeHerdr herdr = new FakeHerdr().readText("⏺ X\n❯ ");
Rendezvous rendezvous = new Rendezvous(); // no waiter opened
CompletionResolver resolver = new CompletionResolver(new AgentControl(herdr), rendezvous);
CompletionResolver resolver = new CompletionResolver(new AgentControl(herdr), rendezvous, ExhaustedPatternLookup.none(), ExhaustionSink.none());
resolver.captureBaseline("term_a"); // no send to attribute a later completion to
resolver.captureBaseline("term_a", TestTurnTokens.inert("term_a")); // no send to attribute a later completion to
assertFalse(herdr.called("agent.read"),
"with no waiting send there is no turn to baseline — skip the scrape");
@@ -134,7 +139,7 @@ class CompletionResolverTest {
// send must NOT be resolved with the stale answer.
FakeHerdr herdr = new FakeHerdr().readText("⏺ 391\n❯ ");
Rendezvous rendezvous = new Rendezvous();
CompletionResolver resolver = new CompletionResolver(new AgentControl(herdr), rendezvous);
CompletionResolver resolver = new CompletionResolver(new AgentControl(herdr), rendezvous, ExhaustedPatternLookup.none(), ExhaustionSink.none());
var waiter = rendezvous.open("term_a"); // a send is blocked on this turn
// The turn as captured at delivery: its waiter, and the previous turn's answer still on screen.
@@ -149,7 +154,7 @@ class CompletionResolverTest {
void resolvesACompletionWhoseScrapeChangedSinceDelivery() {
FakeHerdr herdr = new FakeHerdr().readText("⏺ No, 391 = 17 × 23.\n❯ "); // the worker's real answer
Rendezvous rendezvous = new Rendezvous();
CompletionResolver resolver = new CompletionResolver(new AgentControl(herdr), rendezvous);
CompletionResolver resolver = new CompletionResolver(new AgentControl(herdr), rendezvous, ExhaustedPatternLookup.none(), ExhaustionSink.none());
var waiter = rendezvous.open("term_a");
// Delivery baseline was the previous turn's "391"; the scrape now differs → resolve.
@@ -166,7 +171,7 @@ class CompletionResolverTest {
String block = "⏺ " + "x".repeat(CompletionResolver.MAX_SCRAPE_CHARS + 1) + "\n❯ ";
FakeHerdr herdr = new FakeHerdr().readText(block);
Rendezvous rendezvous = new Rendezvous();
CompletionResolver resolver = new CompletionResolver(new AgentControl(herdr), rendezvous);
CompletionResolver resolver = new CompletionResolver(new AgentControl(herdr), rendezvous, ExhaustedPatternLookup.none(), ExhaustionSink.none());
var waiter = rendezvous.open("term_a");
resolver.resolve("term_a", new CompletionResolver.InFlight(waiter, null));
@@ -180,7 +185,7 @@ class CompletionResolverTest {
void leavesAnUnclippedCompletionPaneTailUnmarked() {
FakeHerdr herdr = new FakeHerdr().readText("⏺ complete report\n❯ ");
Rendezvous rendezvous = new Rendezvous();
CompletionResolver resolver = new CompletionResolver(new AgentControl(herdr), rendezvous);
CompletionResolver resolver = new CompletionResolver(new AgentControl(herdr), rendezvous, ExhaustedPatternLookup.none(), ExhaustionSink.none());
var waiter = rendezvous.open("term_a");
resolver.resolve("term_a", new CompletionResolver.InFlight(waiter, null));
@@ -192,9 +197,9 @@ class CompletionResolverTest {
void resolvesSynchronouslyBeforePostTurnContextClearing() {
FakeHerdr herdr = new FakeHerdr().readText("⏺ previous answer\n❯ ");
Rendezvous rendezvous = new Rendezvous();
CompletionResolver resolver = new CompletionResolver(new AgentControl(herdr), rendezvous);
CompletionResolver resolver = new CompletionResolver(new AgentControl(herdr), rendezvous, ExhaustedPatternLookup.none(), ExhaustionSink.none());
var waiter = rendezvous.open("term_a");
resolver.captureBaseline("term_a");
resolver.captureBaseline("term_a", new TurnToken("term_a", waiter));
herdr.readText("⏺ answer that /clear would erase\n❯ ");
resolver.resolveBeforePostAction("term_a");
@@ -214,10 +219,10 @@ class CompletionResolverTest {
String longBlock = "⏺ " + "x".repeat(CompletionResolver.MAX_SCRAPE_CHARS + 500) + "\n❯ ";
FakeHerdr herdr = new FakeHerdr().readText(longBlock);
Rendezvous rendezvous = new Rendezvous();
CompletionResolver resolver = new CompletionResolver(new AgentControl(herdr), rendezvous);
CompletionResolver resolver = new CompletionResolver(new AgentControl(herdr), rendezvous, ExhaustedPatternLookup.none(), ExhaustionSink.none());
var waiter = rendezvous.open("term_a"); // a send is blocked on this turn
resolver.captureBaseline("term_a"); // baseline is the clipped >cap block
resolver.captureBaseline("term_a", new TurnToken("term_a", waiter)); // baseline is the clipped >cap block
var turn = resolver.inFlight("term_a");
assertEquals(CompletionResolver.MAX_SCRAPE_CHARS, turn.baseline().length(),
"the delivery baseline is clipped to the same cap resolve() applies to the tail");
@@ -234,7 +239,7 @@ class CompletionResolverTest {
// No delivery baseline (e.g. the pre-turn read failed) ⇒ never suppress; the completion resolves.
FakeHerdr herdr = new FakeHerdr().readText("⏺ hello\n❯ ");
Rendezvous rendezvous = new Rendezvous();
CompletionResolver resolver = new CompletionResolver(new AgentControl(herdr), rendezvous);
CompletionResolver resolver = new CompletionResolver(new AgentControl(herdr), rendezvous, ExhaustedPatternLookup.none(), ExhaustionSink.none());
var waiter = rendezvous.open("term_a");
resolver.resolve("term_a", new CompletionResolver.InFlight(waiter, null));
@@ -252,7 +257,7 @@ class CompletionResolverTest {
// byte-identical guard would wrongly match the empty tail and suppress.
FakeHerdr herdr = new FakeHerdr().healthy(false); // agent.read throws HerdrException
Rendezvous rendezvous = new Rendezvous();
CompletionResolver resolver = new CompletionResolver(new AgentControl(herdr), rendezvous);
CompletionResolver resolver = new CompletionResolver(new AgentControl(herdr), rendezvous, ExhaustedPatternLookup.none(), ExhaustionSink.none());
var waiter = rendezvous.open("term_a");
var turn = new CompletionResolver.InFlight(waiter, ""); // empty pane baselined at delivery
@@ -272,7 +277,7 @@ class CompletionResolverTest {
// fail must not overwrite that value, and must not even scrape the worker — nobody needs it.
FakeHerdr herdr = new FakeHerdr().readText("an error screen");
Rendezvous rendezvous = new Rendezvous();
CompletionResolver resolver = new CompletionResolver(new AgentControl(herdr), rendezvous);
CompletionResolver resolver = new CompletionResolver(new AgentControl(herdr), rendezvous, ExhaustedPatternLookup.none(), ExhaustionSink.none());
var waiter = rendezvous.open("term_a");
var turn = new CompletionResolver.InFlight(waiter, null);
@@ -293,7 +298,7 @@ class CompletionResolverTest {
// fail falls back to the waiter currently registered on the Rendezvous and fails it.
FakeHerdr herdr = new FakeHerdr().readText("stuck on an error screen");
Rendezvous rendezvous = new Rendezvous();
CompletionResolver resolver = new CompletionResolver(new AgentControl(herdr), rendezvous);
CompletionResolver resolver = new CompletionResolver(new AgentControl(herdr), rendezvous, ExhaustedPatternLookup.none(), ExhaustionSink.none());
var waiter = rendezvous.open("term_a"); // send registered, but no captureBaseline ever ran
resolver.fail("term_a", null); // no in-flight turn → fall back to the registered waiter
@@ -319,7 +324,7 @@ class CompletionResolverTest {
try {
FakeHerdr herdr = new FakeHerdr().readText("stuck on an error screen");
Rendezvous rendezvous = new Rendezvous();
CompletionResolver resolver = new CompletionResolver(new AgentControl(herdr), rendezvous);
CompletionResolver resolver = new CompletionResolver(new AgentControl(herdr), rendezvous, ExhaustedPatternLookup.none(), ExhaustionSink.none());
var waiter = rendezvous.open("term_a");
resolver.fail("term_a", null);
@@ -347,7 +352,7 @@ class CompletionResolverTest {
// scrape to turn N+1; targeting turn N's captured waiter makes the late completion a no-op.
FakeHerdr herdr = new FakeHerdr().readText("⏺ turn N answer\n❯ ");
Rendezvous rendezvous = new Rendezvous();
CompletionResolver resolver = new CompletionResolver(new AgentControl(herdr), rendezvous);
CompletionResolver resolver = new CompletionResolver(new AgentControl(herdr), rendezvous, ExhaustedPatternLookup.none(), ExhaustionSink.none());
var waiterN = rendezvous.open("term_a"); // turn N's send
// The turn as the injector captured it at delivery (waiter + pre-turn baseline).
@@ -368,4 +373,129 @@ class CompletionResolverTest {
"turn N stays resolved by its own reply");
assertTrue(rendezvous.isWaiting("term_a"), "turn N+1 is still awaiting its own resolution");
}
// --- CB-578 stage A: backend-exhausted classification ---------------------------------
@Test
void classifiesAMatchingScrapeAsBackendExhaustedInsteadOfACompletedReply() {
String block = "⏺ Working on it...\nThe usage limit has been reached. Try again later.\n❯ ";
FakeHerdr herdr = new FakeHerdr().readText(block);
Rendezvous rendezvous = new Rendezvous();
ExhaustedPatternLookup patterns = target -> Pattern.compile("usage limit has been reached");
CompletionResolver resolver = new CompletionResolver(new AgentControl(herdr), rendezvous, patterns, ExhaustionSink.none());
var waiter = rendezvous.open("term_a");
resolver.resolve("term_a", new CompletionResolver.InFlight(waiter, null));
assertTrue(waiter.isDone(), "a matching scrape still resolves the blocked send");
assertEquals(Rendezvous.Kind.BACKEND_EXHAUSTED, waiter.getNow(null).kind(),
"not reported as a completed reply — the classification is distinct");
}
@Test
void theExhaustedReasonCarriesTheMatchedLine() {
String block = "⏺ Working on it...\nThe usage limit has been reached. Try again later.\n❯ ";
FakeHerdr herdr = new FakeHerdr().readText(block);
Rendezvous rendezvous = new Rendezvous();
ExhaustedPatternLookup patterns = target -> Pattern.compile("usage limit has been reached");
CompletionResolver resolver = new CompletionResolver(new AgentControl(herdr), rendezvous, patterns, ExhaustionSink.none());
var waiter = rendezvous.open("term_a");
resolver.resolve("term_a", new CompletionResolver.InFlight(waiter, null));
assertEquals("backend exhausted (usage limit): The usage limit has been reached. Try again later.",
waiter.getNow(null).text(), "the reason names the real cause and carries the matched line");
}
@Test
void aWinningBackendExhaustedClassificationNotifiesTheExhaustionSink() {
String block = "⏺ Working on it...\nThe usage limit has been reached. Try again later.\n❯ ";
FakeHerdr herdr = new FakeHerdr().readText(block);
Rendezvous rendezvous = new Rendezvous();
ExhaustedPatternLookup patterns = target -> Pattern.compile("usage limit has been reached");
java.util.List<String> notified = new java.util.ArrayList<>();
ExhaustionSink sink = (target, reason) -> notified.add(target + ": " + reason);
CompletionResolver resolver = new CompletionResolver(new AgentControl(herdr), rendezvous, patterns, sink);
var waiter = rendezvous.open("term_a");
resolver.resolve("term_a", new CompletionResolver.InFlight(waiter, null));
assertEquals(1, notified.size(), "the sink is notified exactly once for the winning classification");
assertTrue(notified.get(0).startsWith("term_a: "), "the sink is told which target exhausted");
assertTrue(notified.get(0).contains("The usage limit has been reached"),
"the sink is told the matched reason: " + notified.get(0));
}
@Test
void aLosingBackendExhaustedClassificationNeverNotifiesTheExhaustionSink() {
// The waiter was already resolved (e.g. by the worker's own reply) before this scrape landed —
// resolveExhausted loses the race and must return false, so the sink must not fire either.
String block = "⏺ Working on it...\nThe usage limit has been reached. Try again later.\n❯ ";
FakeHerdr herdr = new FakeHerdr().readText(block);
Rendezvous rendezvous = new Rendezvous();
ExhaustedPatternLookup patterns = target -> Pattern.compile("usage limit has been reached");
java.util.List<String> notified = new java.util.ArrayList<>();
ExhaustionSink sink = (target, reason) -> notified.add(target + ": " + reason);
CompletionResolver resolver = new CompletionResolver(new AgentControl(herdr), rendezvous, patterns, sink);
var waiter = rendezvous.open("term_a");
var turn = new CompletionResolver.InFlight(waiter, null);
assertTrue(rendezvous.resolveCompletion(waiter, "already replied"));
resolver.resolve("term_a", turn);
assertTrue(notified.isEmpty(), "a classification that loses the race must not quarantine anything");
assertEquals("already replied", waiter.getNow(null).text(), "the earlier resolution stands untouched");
}
@Test
void aNonMatchingScrapeResolvesAsAnOrdinaryCompletion() {
FakeHerdr herdr = new FakeHerdr().readText("⏺ complete report\n❯ ");
Rendezvous rendezvous = new Rendezvous();
ExhaustedPatternLookup patterns = target -> Pattern.compile("usage limit has been reached");
CompletionResolver resolver = new CompletionResolver(new AgentControl(herdr), rendezvous, patterns, ExhaustionSink.none());
var waiter = rendezvous.open("term_a");
resolver.resolve("term_a", new CompletionResolver.InFlight(waiter, null));
assertEquals(Rendezvous.Kind.COMPLETION, waiter.getNow(null).kind(),
"a scrape that does not match the pattern is an ordinary completion");
assertEquals("complete report", waiter.getNow(null).text());
}
@Test
void aProfileWithNoConfiguredPatternKeepsTodaysCompletionFallbackUnchanged() {
// Even a scrape that WOULD have matched some other profile's pattern must resolve as a
// plain completion when this target's own profile has none configured (CB-578 criterion 4).
String block = "⏺ The usage limit has been reached.\n❯ ";
FakeHerdr herdr = new FakeHerdr().readText(block);
Rendezvous rendezvous = new Rendezvous();
CompletionResolver resolver =
new CompletionResolver(new AgentControl(herdr), rendezvous, ExhaustedPatternLookup.none(), ExhaustionSink.none());
var waiter = rendezvous.open("term_a");
resolver.resolve("term_a", new CompletionResolver.InFlight(waiter, null));
assertEquals(Rendezvous.Kind.COMPLETION, waiter.getNow(null).kind(),
"no pattern configured for this target's profile ⇒ unchanged completion-fallback behaviour");
assertEquals("The usage limit has been reached.", waiter.getNow(null).text());
}
@Test
void coverageIsOffWhenNoProfileHasAPatternConfigured() {
assertEquals("off (no profile has an exhaustedPattern configured; profiles: [terra])",
CompletionResolver.coverage(Set.of("terra"), Set.of()));
}
@Test
void coverageIsFullWhenEveryProfileHasAPatternConfigured() {
assertEquals("full (all profiles configured: [gx10, terra])",
CompletionResolver.coverage(Set.of("terra", "gx10"), Set.of("terra", "gx10")));
}
@Test
void coverageIsPartialAndNamesWhichProfilesAreConfigured() {
assertEquals("partial (configured: [terra]; not configured: [gx10])",
CompletionResolver.coverage(Set.of("terra", "gx10"), Set.of("terra")));
}
}
@@ -8,6 +8,7 @@ import dev.ltms.bridged.herdr.AgentControl;
import dev.ltms.bridged.herdr.AgentStatus;
import dev.ltms.bridged.herdr.FakeHerdr;
import dev.ltms.bridged.herdr.HerdrException;
import dev.ltms.bridged.msg.TestTurnTokens;
import org.junit.jupiter.api.Test;
import org.slf4j.LoggerFactory;
@@ -47,7 +48,7 @@ class InjectorTest {
@Test
void deliversWhenIdle() {
CompletableFuture<Void> f = injector.enqueue(T, "hello");
CompletableFuture<Void> f = injector.enqueue(T, "hello", TestTurnTokens.inert(T));
assertFalse(f.isDone(), "not delivered until an injectable status arrives");
injector.onStatus(T, AgentStatus.IDLE);
assertTrue(f.isDone());
@@ -59,7 +60,7 @@ class InjectorTest {
// CB-113: idle alone is not enough — hold until the worker's MCP is connected (ready).
java.util.Set<String> ready = new java.util.HashSet<>();
Injector inj = new Injector(new AgentControl(herdr), TurnListener.NOOP, ready::contains);
inj.enqueue(T, "task");
inj.enqueue(T, "task", TestTurnTokens.inert(T));
inj.onStatus(T, AgentStatus.IDLE); // idle but not yet available → held out of the boot window
assertEquals(List.of(), sent(), "must not deliver into a not-yet-available worker");
@@ -79,7 +80,7 @@ class InjectorTest {
void resubmitsEnterWhenADeliveredMessageIsNotPickedUp() {
// CB-113: the Enter at delivery can race the paste; while the worker stays idle (not picked
// up), the injector re-nudges Enter so the pending paste submits.
injector.enqueue(T, "task");
injector.enqueue(T, "task", TestTurnTokens.inert(T));
injector.onStatus(T, AgentStatus.IDLE); // deliver: paste + one Enter
long afterDeliver = enterKeystrokes();
@@ -95,7 +96,7 @@ class InjectorTest {
@Test
void holdsWhileWorkingThenDeliversOnIdle() {
injector.enqueue(T, "later");
injector.enqueue(T, "later", TestTurnTokens.inert(T));
injector.onStatus(T, AgentStatus.WORKING);
assertEquals(List.of(), sent(), "must not inject mid-turn");
injector.onStatus(T, AgentStatus.IDLE);
@@ -104,7 +105,7 @@ class InjectorTest {
@Test
void blockedIsInjectableButUnknownIsNot() {
injector.enqueue(T, "answer");
injector.enqueue(T, "answer", TestTurnTokens.inert(T));
injector.onStatus(T, AgentStatus.UNKNOWN);
assertEquals(List.of(), sent(), "unknown status is not safe to inject");
injector.onStatus(T, AgentStatus.BLOCKED);
@@ -113,8 +114,8 @@ class InjectorTest {
@Test
void twoRapidDeliveriesNeverInterleave() {
injector.enqueue(T, "m1");
injector.enqueue(T, "m2");
injector.enqueue(T, "m1", TestTurnTokens.inert(T));
injector.enqueue(T, "m2", TestTurnTokens.inert(T));
// First idle window delivers only m1, even if idle is observed twice before pickup.
injector.onStatus(T, AgentStatus.IDLE);
@@ -129,8 +130,8 @@ class InjectorTest {
@Test
void transientUnknownDoesNotReleaseThePickupLatch() {
injector.enqueue(T, "m1");
injector.enqueue(T, "m2");
injector.enqueue(T, "m1", TestTurnTokens.inert(T));
injector.enqueue(T, "m2", TestTurnTokens.inert(T));
injector.onStatus(T, AgentStatus.IDLE); // m1 sent, awaiting pickup
assertEquals(List.of("m1"), sent());
@@ -145,8 +146,8 @@ class InjectorTest {
@Test
void missedPickupEdgeIsReleasedByGraceSoTheQueueNeverWedges() {
injector.enqueue(T, "m1");
injector.enqueue(T, "m2");
injector.enqueue(T, "m1", TestTurnTokens.inert(T));
injector.enqueue(T, "m2", TestTurnTokens.inert(T));
injector.onStatus(T, AgentStatus.IDLE); // m1 sent
assertEquals(List.of("m1"), sent());
@@ -158,9 +159,9 @@ class InjectorTest {
@Test
void fifoOrderAcrossManyTurns() {
injector.enqueue(T, "a");
injector.enqueue(T, "b");
injector.enqueue(T, "c");
injector.enqueue(T, "a", TestTurnTokens.inert(T));
injector.enqueue(T, "b", TestTurnTokens.inert(T));
injector.enqueue(T, "c", TestTurnTokens.inert(T));
for (int i = 0; i < 3; i++) {
injector.onStatus(T, AgentStatus.IDLE); // deliver one
injector.onStatus(T, AgentStatus.WORKING); // pickup
@@ -172,7 +173,7 @@ class InjectorTest {
@Test
void activeWhileQueuedOrInFlightThenQuietAfterTurnCompletes() {
assertTrue(injector.activeTargets().isEmpty());
injector.enqueue(T, "x");
injector.enqueue(T, "x", TestTurnTokens.inert(T));
assertEquals(Set.of(T), injector.activeTargets(), "active while a message is queued");
injector.onStatus(T, AgentStatus.IDLE); // delivers; awaiting pickup
@@ -191,7 +192,7 @@ class InjectorTest {
void firesTurnCompleteOnAConfirmedWorkingThenIdle() {
List<String> completed = new ArrayList<>();
Injector inj = new Injector(new AgentControl(herdr), completed::add);
inj.enqueue(T, "task");
inj.enqueue(T, "task", TestTurnTokens.inert(T));
inj.onStatus(T, AgentStatus.IDLE); // deliver
inj.onStatus(T, AgentStatus.WORKING); // pickup + turn running
@@ -226,8 +227,8 @@ class InjectorTest {
}
ResetListener listener = new ResetListener();
Injector inj = new Injector(agents, listener);
inj.enqueue(T, "first");
inj.enqueue(T, "second");
inj.enqueue(T, "first", TestTurnTokens.inert(T));
inj.enqueue(T, "second", TestTurnTokens.inert(T));
inj.onStatus(T, AgentStatus.IDLE); // first delegation
inj.onStatus(T, AgentStatus.WORKING);
@@ -246,7 +247,7 @@ class InjectorTest {
void doesNotSynthesizeCompletionFromAnUnconfirmedTurn() {
List<String> completed = new ArrayList<>();
Injector inj = new Injector(new AgentControl(herdr), completed::add);
inj.enqueue(T, "task");
inj.enqueue(T, "task", TestTurnTokens.inert(T));
// Deliver, then only ever idle — a `working` sample is never seen. The pickup grace unwedges
// the queue but must NOT invent a completion: without a sampled turn there is no trustworthy
@@ -259,6 +260,7 @@ class InjectorTest {
private static final class Captor implements TurnListener {
final List<String> completed = new ArrayList<>();
final List<String> failed = new ArrayList<>();
final List<String> failureReasons = new ArrayList<>();
@Override
public void onTurnComplete(String target) {
@@ -269,6 +271,12 @@ class InjectorTest {
public void onTurnFailed(String target) {
failed.add(target);
}
@Override
public void onTurnFailed(String target, String reason) {
failed.add(target);
failureReasons.add(reason);
}
}
// ~30s of unknown at the 250ms prod poll interval; enough onStatus samples to trip the stall.
@@ -278,7 +286,7 @@ class InjectorTest {
void failsAnOutstandingDelegationWhoseWorkerWedgesInUnknown() {
Captor cap = new Captor();
Injector inj = new Injector(new AgentControl(herdr), cap);
inj.enqueue(T, "task");
inj.enqueue(T, "task", TestTurnTokens.inert(T));
inj.onStatus(T, AgentStatus.IDLE); // deliver
inj.onStatus(T, AgentStatus.WORKING); // worker starts the turn
@@ -293,7 +301,7 @@ class InjectorTest {
void aTransientUnknownGlitchNeitherFailsNorBlocksCompletion() {
Captor cap = new Captor();
Injector inj = new Injector(new AgentControl(herdr), cap);
inj.enqueue(T, "task");
inj.enqueue(T, "task", TestTurnTokens.inert(T));
inj.onStatus(T, AgentStatus.IDLE); // deliver
inj.onStatus(T, AgentStatus.WORKING); // confirmed turn
@@ -308,7 +316,7 @@ class InjectorTest {
void sendFailureDropsMessageAndFailsItsFuture() {
FakeHerdr failing = new FakeHerdr().agentSendFailsWith("send_failed");
Injector inj = new Injector(new AgentControl(failing));
CompletableFuture<Void> f = inj.enqueue(T, "boom");
CompletableFuture<Void> f = inj.enqueue(T, "boom", TestTurnTokens.inert(T));
inj.onStatus(T, AgentStatus.IDLE);
assertTrue(f.isCompletedExceptionally());
@@ -317,18 +325,35 @@ class InjectorTest {
@Test
void dropFailsPendingWaiters() {
CompletableFuture<Void> f = injector.enqueue(T, "orphan");
CompletableFuture<Void> f = injector.enqueue(T, "orphan", TestTurnTokens.inert(T));
injector.drop(T, new HerdrException("worker gone", "pane_not_found", null));
assertTrue(f.isCompletedExceptionally(), "queued waiters unblock when the worker vanishes");
}
@Test
void dropPassesTheRealCauseForQueuedAndDeliveredWork() {
Captor cap = new Captor();
Injector inj = new Injector(new AgentControl(herdr), cap);
CompletableFuture<Void> delivered = inj.enqueue(T, "delivered", TestTurnTokens.inert(T));
CompletableFuture<Void> queued = inj.enqueue(T, "queued", TestTurnTokens.inert(T));
inj.onStatus(T, AgentStatus.IDLE); // deliver the first message
inj.onStatus(T, AgentStatus.WORKING); // its turn is now in flight; one remains queued
inj.drop(T, new HerdrException("agent target sol not found", "agent_not_found", null));
assertEquals(List.of(T), cap.failed, "drop signals one turn failure for both affected states");
assertEquals(List.of("agent target sol not found"), cap.failureReasons);
assertTrue(delivered.isDone(), "the delivered future has already completed");
assertTrue(queued.isCompletedExceptionally(), "the queued future fails with the drop cause");
}
@Test
void dropFailsTheTurnOfADeliveredMessageWhenTheWorkerVanishes() {
// CB-110: the message was delivered (no longer queued), so failing queued waiters alone would
// leave its send hanging. A vanished worker must fail that in-flight turn too.
Captor cap = new Captor();
Injector inj = new Injector(new AgentControl(herdr), cap);
inj.enqueue(T, "task");
inj.enqueue(T, "task", TestTurnTokens.inert(T));
inj.onStatus(T, AgentStatus.IDLE); // deliver
inj.onStatus(T, AgentStatus.WORKING); // turn running
@@ -343,7 +368,7 @@ class InjectorTest {
// by an off-sub worker's review of CB-110, delegated through the bridge.)
Captor cap = new Captor();
Injector inj = new Injector(new AgentControl(herdr), cap);
inj.enqueue(T, "task");
inj.enqueue(T, "task", TestTurnTokens.inert(T));
inj.onStatus(T, AgentStatus.IDLE); // deliver; pickup never confirmed
inj.drop(T, new HerdrException("worker gone", "pane_not_found", null));
@@ -362,7 +387,7 @@ class InjectorTest {
Captor cap = new Captor();
List<String> forgotten = new ArrayList<>();
Injector inj = new Injector(new AgentControl(herdr), cap, _ -> false, forgotten::add);
CompletableFuture<Void> f = inj.enqueue(T, "task");
CompletableFuture<Void> f = inj.enqueue(T, "task", TestTurnTokens.inert(T));
for (int i = 0; i < READINESS_SAMPLES; i++) inj.onStatus(T, AgentStatus.IDLE);
@@ -390,7 +415,7 @@ class InjectorTest {
try {
Injector inj = new Injector(new AgentControl(herdr), TurnListener.NOOP, _ -> false, _ -> {
});
inj.enqueue(T, "task");
inj.enqueue(T, "task", TestTurnTokens.inert(T));
for (int i = 0; i < READINESS_SAMPLES; i++) inj.onStatus(T, AgentStatus.IDLE);
@@ -415,7 +440,7 @@ class InjectorTest {
Set<String> ready = new java.util.HashSet<>();
Injector inj = new Injector(new AgentControl(herdr), TurnListener.NOOP, ready::contains, _ -> {
});
inj.enqueue(T, "task");
inj.enqueue(T, "task", TestTurnTokens.inert(T));
for (int i = 0; i < 100; i++) inj.onStatus(T, AgentStatus.IDLE); // still booting, well under grace
assertEquals(List.of(), sent());
@@ -431,7 +456,7 @@ class InjectorTest {
// linger past the worker's life (MemberPresence.forget had no caller before this).
List<String> forgotten = new ArrayList<>();
Injector inj = new Injector(new AgentControl(herdr), TurnListener.NOOP, _ -> true, forgotten::add);
inj.enqueue(T, "orphan");
inj.enqueue(T, "orphan", TestTurnTokens.inert(T));
inj.drop(T, new HerdrException("worker gone", "pane_not_found", null));
assertEquals(List.of(T), forgotten, "drop clears the gone worker's presence");
}
@@ -452,7 +477,7 @@ class InjectorTest {
try {
Injector inj = new Injector(new AgentControl(herdr), TurnListener.NOOP, _ -> true, _ -> {
});
inj.enqueue(T, "orphan");
inj.enqueue(T, "orphan", TestTurnTokens.inert(T));
inj.drop(T, new HerdrException("worker gone", "pane_not_found", null));
String warn = appender.list.stream()
@@ -476,7 +501,7 @@ class InjectorTest {
StatusPoller poller = new StatusPoller(new AgentControl(idle), inj, 10);
poller.start();
try {
CompletableFuture<Void> delivered = inj.enqueue(T, "via-poller");
CompletableFuture<Void> delivered = inj.enqueue(T, "via-poller", TestTurnTokens.inert(T));
delivered.get(2, TimeUnit.SECONDS); // completes when the poller drives the send
} finally {
poller.stop();
@@ -495,7 +520,7 @@ class InjectorTest {
void deliveredFutureCarriesSendFailure() {
FakeHerdr failing = new FakeHerdr().agentSendFailsWith("send_failed");
Injector inj = new Injector(new AgentControl(failing));
CompletableFuture<Void> f = inj.enqueue(T, "boom");
CompletableFuture<Void> f = inj.enqueue(T, "boom", TestTurnTokens.inert(T));
inj.onStatus(T, AgentStatus.IDLE);
ExecutionException ex = assertThrows(ExecutionException.class, f::get);
assertInstanceOf(HerdrException.class, ex.getCause());
@@ -28,7 +28,7 @@ class LeadLauncherTest {
List.of("ccs", "ltms"), "tab", "bridged-workers", null,
"http://127.0.0.1:8765/mcp", null, null,
null, null, null,
Map.of("CLAUDE_CODE_AUTO_COMPACT_WINDOW", "300000"), null, null, true);
Map.of("CLAUDE_CODE_AUTO_COMPACT_WINDOW", "300000"), null, null, true, null);
}
private static BridgedConfig configWith(BridgedConfig.Leader lead) {
@@ -41,8 +41,8 @@ class LeadLauncherTest {
null, null, fleet, null, "fixed", null).withDefaults();
}
private static BridgedConfig.Leader lead(String profile, String terminal, int instances) {
return new BridgedConfig.Leader(profile, terminal, instances, "lead:", 10, null, null,
private static BridgedConfig.Leader lead(String profile, String tab, int instances) {
return new BridgedConfig.Leader(profile, tab, instances, "lead:", 10, null, null,
"leads", "/repo");
}
@@ -70,16 +70,16 @@ class LeadLauncherTest {
void startsTheDeclaredLeadWhenNoneIsRunning() {
FakeHerdr herdr = new FakeHerdr();
assertEquals(1, launcher(herdr, configWith(lead("opus", null, 1))).ensureLeads());
assertEquals(1, launcher(herdr, configWith(lead("opus", "lead: opus", 1))).ensureLeads());
assertTrue(herdr.called("agent.start"), "a lead must actually be started");
assertEquals("lead-opus", startedName(herdr));
}
/** The tab is labelled so the scanner finds the lead on the next resolve. */
/** The tab is labelled with the configured `tab:` so the scanner finds the lead on the next resolve. */
@Test
void labelsTheTabWithThePrefixTheScannerReadsBack() {
void labelsTheTabWithTheConfiguredTabValue() {
FakeHerdr herdr = new FakeHerdr();
launcher(herdr, configWith(lead("opus", null, 1))).ensureLeads();
launcher(herdr, configWith(lead("opus", "lead: opus", 1))).ensureLeads();
assertEquals("lead: opus",
((Map<?, ?>) herdr.lastCall("tab.rename").params()).get("label"));
@@ -90,7 +90,7 @@ class LeadLauncherTest {
void startsAsManyInstancesAsAreDeclared() {
FakeHerdr herdr = new FakeHerdr();
assertEquals(2, launcher(herdr, configWith(lead("opus", null, 2))).ensureLeads());
assertEquals(2, launcher(herdr, configWith(lead("opus", "lead: opus", 2))).ensureLeads());
assertEquals(2, herdr.calls.stream().filter(c -> c.method().equals("agent.start")).count());
}
@@ -104,7 +104,7 @@ class LeadLauncherTest {
.withTab("wL", "wL:t1", "lead: opus")
.withAgent("lead-opus", "term_lead", "wL:p1", "wL:t1");
assertEquals(0, launcher(herdr, configWith(lead("opus", null, 1))).ensureLeads());
assertEquals(0, launcher(herdr, configWith(lead("opus", "lead: opus", 1))).ensureLeads());
assertFalse(herdr.called("agent.start"), "the live lead must not be duplicated");
}
@@ -118,20 +118,23 @@ class LeadLauncherTest {
.withWorkspace("wL", "leads")
.withTab("wL", "wL:t1", "lead: opus"); // label only — nothing running in it
assertEquals(1, launcher(herdr, configWith(lead("opus", null, 1))).ensureLeads(),
assertEquals(1, launcher(herdr, configWith(lead("opus", "lead: opus", 1))).ensureLeads(),
"a stale label is not a lead; the lead must be relaunched");
}
/**
* A lead the operator opened by hand and pinned with `terminal:` is live even though its tab
* carries no matching label. Counting labels alone would relaunch it on every boot.
* A lead the operator opened by hand is live once its tab carries the configured `tab:` label —
* CB-579 retired the `terminal:` pin, so a hand-opened lead is found the same way an
* auto-launched one is, by its tab, not by a terminal id nobody wrote down in advance.
*/
@Test
void aPinnedTerminalWithARunningAgentCountsAsLive() {
void aHandOpenedLeadWithTheConfiguredTabLabelCountsAsLive() {
FakeHerdr herdr = new FakeHerdr()
.withAgent("hand-opened", "term_pinned", "wX:p1", "wX:t1");
.withWorkspace("wX", "main")
.withTab("wX", "wX:t1", "lead: opus")
.withAgent("hand-opened", "term_hand", "wX:p1", "wX:t1");
assertEquals(0, launcher(herdr, configWith(lead("opus", "term_pinned", 1))).ensureLeads());
assertEquals(0, launcher(herdr, configWith(lead("opus", "lead: opus", 1))).ensureLeads());
assertFalse(herdr.called("agent.start"));
}
@@ -143,7 +146,7 @@ class LeadLauncherTest {
.withTab("wM", "wM:t1", "lead: opus") // a member tab that looks like a lead
.withAgent("claude-opus-x", "term_m", "wM:p1", "wM:t1");
assertEquals(1, launcher(herdr, configWith(lead("opus", null, 1))).ensureLeads(),
assertEquals(1, launcher(herdr, configWith(lead("opus", "lead: opus", 1))).ensureLeads(),
"a member in a lead-labelled tab is not a lead, so the real lead is still missing");
}
@@ -152,7 +155,7 @@ class LeadLauncherTest {
void anUncountableHerdrStartsNothing() {
FakeHerdr herdr = new FakeHerdr().healthy(false);
assertEquals(0, launcher(herdr, configWith(lead("opus", null, 1))).ensureLeads());
assertEquals(0, launcher(herdr, configWith(lead("opus", "lead: opus", 1))).ensureLeads());
assertFalse(herdr.called("agent.start"));
}
@@ -166,7 +169,7 @@ class LeadLauncherTest {
@Test
void theLeadNeverReceivesTheWorkerReplyCharter() {
FakeHerdr herdr = new FakeHerdr();
launcher(herdr, configWith(lead("opus", null, 1))).ensureLeads();
launcher(herdr, configWith(lead("opus", "lead: opus", 1))).ensureLeads();
List<String> args = startedArgs(herdr);
assertFalse(args.contains("--append-system-prompt"),
@@ -178,7 +181,7 @@ class LeadLauncherTest {
@Test
void theLeadMountsTheBridgeMcpAndPinsItsModel() {
FakeHerdr herdr = new FakeHerdr();
launcher(herdr, configWith(lead("opus", null, 1))).ensureLeads();
launcher(herdr, configWith(lead("opus", "lead: opus", 1))).ensureLeads();
List<String> args = startedArgs(herdr);
assertTrue(args.contains("--mcp-config"));
@@ -192,7 +195,7 @@ class LeadLauncherTest {
@Test
void theLeadEnvCarriesNoAnthropicBinding() {
FakeHerdr herdr = new FakeHerdr();
launcher(herdr, configWith(lead("opus", null, 1))).ensureLeads();
launcher(herdr, configWith(lead("opus", "lead: opus", 1))).ensureLeads();
Map<String, String> env = tabEnv(herdr);
assertNull(env.get("ANTHROPIC_BASE_URL"));
@@ -205,7 +208,7 @@ class LeadLauncherTest {
@Test
void theLeadTabIsCreatedOutsideEveryMemberWorkspace() {
FakeHerdr herdr = new FakeHerdr();
launcher(herdr, configWith(lead("opus", null, 1))).ensureLeads();
launcher(herdr, configWith(lead("opus", "lead: opus", 1))).ensureLeads();
String label = (String) ((Map<?, ?>) herdr.lastCall("workspace.create").params()).get("label");
assertEquals("leads", label);
@@ -214,12 +217,12 @@ class LeadLauncherTest {
// ── recognise-only and misconfiguration ───────────────────────────────────────────────────
/** A lead with a pin but no profile is recognise-only by design — not an error, not a launch. */
/** A lead with a tab but no profile is recognise-only by design — not an error, not a launch. */
@Test
void aLeadThatNamesNoProfileIsRecognisedButNeverLaunched() {
FakeHerdr herdr = new FakeHerdr();
assertEquals(0, launcher(herdr, configWith(lead(null, "term_dead", 1))).ensureLeads());
assertEquals(0, launcher(herdr, configWith(lead(null, "lead: dead", 1))).ensureLeads());
assertFalse(herdr.called("agent.start"));
}
@@ -228,7 +231,7 @@ class LeadLauncherTest {
void zeroInstancesLaunchesNothing() {
FakeHerdr herdr = new FakeHerdr();
assertEquals(0, launcher(herdr, configWith(lead("opus", null, 0))).ensureLeads());
assertEquals(0, launcher(herdr, configWith(lead("opus", "lead: opus", 0))).ensureLeads());
assertFalse(herdr.called("agent.start"));
}
@@ -237,7 +240,7 @@ class LeadLauncherTest {
void anUnknownProfileIsSkippedRatherThanThrown() {
FakeHerdr herdr = new FakeHerdr();
assertEquals(0, launcher(herdr, configWith(lead("nope", null, 1))).ensureLeads());
assertEquals(0, launcher(herdr, configWith(lead("nope", "lead: opus", 1))).ensureLeads());
assertFalse(herdr.called("agent.start"));
}
@@ -0,0 +1,37 @@
package dev.ltms.bridged.logging;
import ch.qos.logback.classic.Level;
import ch.qos.logback.classic.Logger;
import ch.qos.logback.classic.spi.ILoggingEvent;
import ch.qos.logback.core.read.ListAppender;
import io.modelcontextprotocol.spec.McpSchema;
import org.junit.jupiter.api.Test;
import org.slf4j.LoggerFactory;
import java.util.Map;
import static org.junit.jupiter.api.Assertions.assertEquals;
import static org.junit.jupiter.api.Assertions.assertTrue;
class McpCancelledNotificationFilterTest {
@Test
void suppressesOnlyTheCancelledNotificationWarning() {
Logger logger = (Logger) LoggerFactory.getLogger(McpCancelledNotificationFilter.LOGGER);
ListAppender<ILoggingEvent> appender = new ListAppender<>();
appender.start();
logger.addAppender(appender);
try {
logger.warn(McpCancelledNotificationFilter.UNHANDLED_NOTIFICATION,
new McpSchema.JSONRPCNotification("notifications/cancelled", Map.of("requestId", 7)));
logger.warn(McpCancelledNotificationFilter.UNHANDLED_NOTIFICATION,
new McpSchema.JSONRPCNotification("notifications/progress", Map.of("progress", 1)));
assertEquals(1, appender.list.size());
assertEquals(Level.WARN, appender.list.getFirst().getLevel());
assertTrue(appender.list.getFirst().getFormattedMessage().contains("notifications/progress"));
} finally {
logger.detachAppender(appender);
}
}
}
@@ -72,7 +72,8 @@ class BridgeMcpAuthzTest {
new PrimaryRegistry(null),
enforce ? CallerResolver.withLeadsAndMembers(identity, false, null,
Map::of, new MemberRegistry(null)) : null,
metrics, BridgeMcp.CapacitySource.none());
metrics, BridgeMcp.CapacitySource.none(), new BridgeMcp.HealthCoverageSource(() -> "off"),
BridgeMcp.QuarantineSource.none());
return mcp;
}
@@ -16,11 +16,15 @@ import dev.ltms.bridged.inject.MemberPresence;
import dev.ltms.bridged.session.MemberSession;
import dev.ltms.bridged.session.WorktreeRequest;
import dev.ltms.bridged.member.ClaudeCodeLauncher;
import dev.ltms.bridged.member.CompositePeerLauncher;
import dev.ltms.bridged.placement.BackendQuarantine;
import dev.ltms.bridged.placement.PlacementPolicies;
import io.modelcontextprotocol.spec.McpSchema;
import dev.ltms.bridged.msg.InMemoryReplyInbox;
import org.junit.jupiter.api.BeforeEach;
import org.junit.jupiter.api.Test;
import java.util.List;
import java.util.Map;
import java.util.Set;
import java.util.concurrent.CompletableFuture;
@@ -422,11 +426,41 @@ class BridgeMcpTest {
assertTrue(textOf(res).contains("unknown worker profile"), textOf(res));
}
/**
* CB-599: a profile at its {@code maxLoad} cap must surface a readable reason on the MCP
* surface too, not merely flip {@code isError} with an opaque or absent message.
*/
@Test
void spawnAtMaxLoadSurfacesTheCapacityReason() {
FakeHerdr h = new FakeHerdr();
BridgedConfig.Profile wcfg = new BridgedConfig.Profile(
"ltms-local", "http://gx00.gw:8000", "coder", null, "BRIDGED_WORKER_TOKEN", null,
"tab", "bridged-workers", "worker: {profile} #{n}", null,
null, null, null, null, null, null, null, 0, null, null, null);
Map<String, BridgedConfig.Profile> profiles = Map.of(wcfg.profile(), wcfg);
ClaudeCodeLauncher delegate = new ClaudeCodeLauncher(
new AgentControl(h), new WorkspaceControl(h), new SubscriptionGuard(Set.of("gx00.gw")),
profiles, wcfg.profile(), k -> "BRIDGED_WORKER_TOKEN".equals(k) ? "tok" : null);
CompositePeerLauncher composite = new CompositePeerLauncher(
List.of(delegate), wcfg.profile(), profiles, PlacementPolicies.fixed(), _ -> 0);
SessionManager sm = new SessionManager(composite);
McpSchema.CallToolResult res = BridgeMcp.spawn(sm, "ltms-local");
assertTrue(res.isError());
String text = textOf(res);
assertTrue(text.contains("no capacity"), "surfaces a capacity reason, not a bare error: " + text);
assertTrue(text.contains("ltms-local"), "names the profile: " + text);
assertTrue(text.contains("maxLoad"), "explains the refusal: " + text);
assertFalse(h.called("agent.start"), "at cap, the spawn is refused before any herdr call");
}
@Test
void spawnPassesTheRequestedCwdToTheWorker() {
FakeHerdr h = new FakeHerdr();
McpSchema.CallToolResult res = BridgeMcp.spawn(
sessionManager(h, "http://gx00.gw:8000", Set.of("gx00.gw")), null, null, "/req/dir", null, null, null);
sessionManager(h, "http://gx00.gw:8000", Set.of("gx00.gw")), null, null, "/req/dir", null, null, null,
null, null);
assertNotEquals(Boolean.TRUE, res.isError());
// Protocol 19: the requested cwd roots the worker's pane at creation (tab.create).
@SuppressWarnings("unchecked")
@@ -437,11 +471,29 @@ class BridgeMcpTest {
@Test
void profilesListsConfiguredProfilesAndDefault() {
FakeHerdr h = new FakeHerdr();
McpSchema.CallToolResult res = BridgeMcp.profiles(workerService(h, "http://gx00.gw:8000", Set.of("gx00.gw")));
McpSchema.CallToolResult res = BridgeMcp.profiles(
workerService(h, "http://gx00.gw:8000", Set.of("gx00.gw")), BridgeMcp.QuarantineSource.none());
assertNotEquals(Boolean.TRUE, res.isError());
String out = textOf(res);
assertTrue(out.contains("ltms-local"), out);
assertTrue(out.contains("\"default\":\"ltms-local\""), out);
assertFalse(out.contains("quarantined"), "no profile is quarantined, so the key is omitted: " + out);
}
@Test
void profilesReportsAQuarantinedCredential() {
FakeHerdr h = new FakeHerdr();
BackendQuarantine quarantine = new BackendQuarantine(() -> 0L, TimeUnit.MINUTES.toNanos(30));
quarantine.quarantine("shared-openai");
BridgeMcp.QuarantineSource source = new BridgeMcp.QuarantineSource(
profile -> "ltms-local".equals(profile) ? "shared-openai" : null, quarantine);
McpSchema.CallToolResult res = BridgeMcp.profiles(
workerService(h, "http://gx00.gw:8000", Set.of("gx00.gw")), source);
assertNotEquals(Boolean.TRUE, res.isError());
String out = textOf(res);
assertTrue(out.contains("\"quarantined\""), out);
assertTrue(out.contains("shared-openai"), out);
assertTrue(out.contains("\"quarantinedForSeconds\":1800"), out);
}
@Test
@@ -474,11 +526,14 @@ class BridgeMcpTest {
sessions.acquire("ltms-local", null, null, null);
McpSchema.CallToolResult res = BridgeMcp.listFleet(workerService(h, "http://gx00.gw:8000", Set.of("gx00.gw")),
sessions, null, new BridgeMcp.CapacitySource(profile -> 2, profile -> 2,
() -> Set.of("ltms-local"), () -> 0), Map.of(), "");
() -> Set.of("ltms-local"), () -> 0), new BridgeMcp.HealthCoverageSource(() -> "off"),
BridgeMcp.QuarantineSource.none(), Map.of(), "");
String out = textOf(res);
assertTrue(out.contains("\"maxLoad\":2"), out);
assertTrue(out.contains("\"live\":2"), out);
assertTrue(out.contains("\"free\":0"), out);
assertFalse(out.contains("credentialId"), "nothing is quarantined, so no new key appears: " + out);
assertFalse(out.contains("quarantinedForSeconds"), out);
}
@Test
@@ -487,7 +542,8 @@ class BridgeMcpTest {
SessionManager sessions = new SessionManager(workerService(h, "http://gx00.gw:8000", Set.of("gx00.gw")));
String out = textOf(BridgeMcp.listFleet(workerService(h, "http://gx00.gw:8000", Set.of("gx00.gw")),
sessions, null, new BridgeMcp.CapacitySource(profile -> 0, profile -> 2,
() -> Set.of("terra"), () -> 0), Map.of(), ""));
() -> Set.of("terra"), () -> 0), new BridgeMcp.HealthCoverageSource(() -> "off"),
BridgeMcp.QuarantineSource.none(), Map.of(), ""));
assertTrue(out.contains("\"profile\":\"terra\""), out);
assertTrue(out.contains("\"live\":0"), out);
assertTrue(out.contains("\"free\":2"), out);
@@ -499,10 +555,90 @@ class BridgeMcpTest {
FakeHerdr h = new FakeHerdr();
String out = textOf(BridgeMcp.listFleet(workerService(h, "http://gx00.gw:8000", Set.of("gx00.gw")),
new SessionManager(workerService(h, "http://gx00.gw:8000", Set.of("gx00.gw"))), null,
BridgeMcp.CapacitySource.none(), Map.of(), ""));
BridgeMcp.CapacitySource.none(), new BridgeMcp.HealthCoverageSource(() -> "off"),
BridgeMcp.QuarantineSource.none(), Map.of(), ""));
assertFalse(out.contains("\"capacity\":"), out);
}
@Test
void quarantinedProfileReportsZeroFreeRegardlessOfMaxLoadAndLive() {
FakeHerdr h = new FakeHerdr();
SessionManager sessions = new SessionManager(workerService(h, "http://gx00.gw:8000", Set.of("gx00.gw")));
BackendQuarantine quarantine = new BackendQuarantine(() -> 0L, TimeUnit.MINUTES.toNanos(20));
quarantine.quarantine("shared-openai");
BridgeMcp.QuarantineSource source = new BridgeMcp.QuarantineSource(
profile -> "terra".equals(profile) ? "shared-openai" : null, quarantine);
String out = textOf(BridgeMcp.listFleet(workerService(h, "http://gx00.gw:8000", Set.of("gx00.gw")),
sessions, null, new BridgeMcp.CapacitySource(profile -> 0, profile -> 2,
() -> Set.of("terra"), () -> 0), new BridgeMcp.HealthCoverageSource(() -> "off"),
source, Map.of(), ""));
assertTrue(out.contains("\"profile\":\"terra\""), out);
assertTrue(out.contains("\"maxLoad\":2"), out);
assertTrue(out.contains("\"live\":0"), out);
assertTrue(out.contains("\"free\":0"), out);
assertTrue(out.contains("\"credentialId\":\"shared-openai\""), out);
assertTrue(out.contains("\"quarantinedForSeconds\":1200"), out);
}
@Test
void everyProfileSharingTheQuarantinedCredentialReportsZeroFree() {
FakeHerdr h = new FakeHerdr();
SessionManager sessions = new SessionManager(workerService(h, "http://gx00.gw:8000", Set.of("gx00.gw")));
BackendQuarantine quarantine = new BackendQuarantine(() -> 0L, TimeUnit.MINUTES.toNanos(20));
quarantine.quarantine("shared-openai");
BridgeMcp.QuarantineSource source = new BridgeMcp.QuarantineSource(_ -> "shared-openai", quarantine);
String out = textOf(BridgeMcp.listFleet(workerService(h, "http://gx00.gw:8000", Set.of("gx00.gw")),
sessions, null, new BridgeMcp.CapacitySource(profile -> 0, profile -> 2,
() -> Set.of("terra", "sol"), () -> 0), new BridgeMcp.HealthCoverageSource(() -> "off"),
source, Map.of(), ""));
assertEquals(2, out.split("\"free\":0", -1).length - 1, out);
}
@Test
void listAndProfilesAgreeOnWhatIsQuarantined() {
FakeHerdr h = new FakeHerdr();
SessionManager sessions = new SessionManager(workerService(h, "http://gx00.gw:8000", Set.of("gx00.gw")));
BackendQuarantine quarantine = new BackendQuarantine(() -> 0L, TimeUnit.MINUTES.toNanos(20));
quarantine.quarantine("shared-openai");
BridgeMcp.QuarantineSource source = new BridgeMcp.QuarantineSource(
profile -> "ltms-local".equals(profile) ? "shared-openai" : null, quarantine);
var workers = workerService(h, "http://gx00.gw:8000", Set.of("gx00.gw"));
String listOut = textOf(BridgeMcp.listFleet(workers, sessions, null,
new BridgeMcp.CapacitySource(profile -> 0, profile -> 2, () -> Set.of("ltms-local"), () -> 0),
new BridgeMcp.HealthCoverageSource(() -> "off"), source, Map.of(), ""));
String profilesOut = textOf(BridgeMcp.profiles(workers, source));
assertTrue(listOut.contains("\"free\":0"), listOut);
assertTrue(listOut.contains("\"credentialId\":\"shared-openai\""), listOut);
assertTrue(profilesOut.contains("\"quarantined\""), profilesOut);
assertTrue(profilesOut.contains("shared-openai"), profilesOut);
}
@Test
void expiredQuarantineLeavesTheCapacityRowOrdinary() {
FakeHerdr h = new FakeHerdr();
SessionManager sessions = new SessionManager(workerService(h, "http://gx00.gw:8000", Set.of("gx00.gw")));
long[] clockNanos = {0L};
BackendQuarantine quarantine = new BackendQuarantine(() -> clockNanos[0], TimeUnit.MINUTES.toNanos(20));
quarantine.quarantine("shared-openai");
clockNanos[0] = TimeUnit.MINUTES.toNanos(21); // clock advances past the cooldown
BridgeMcp.QuarantineSource source = new BridgeMcp.QuarantineSource(
profile -> "terra".equals(profile) ? "shared-openai" : null, quarantine);
String out = textOf(BridgeMcp.listFleet(workerService(h, "http://gx00.gw:8000", Set.of("gx00.gw")),
sessions, null, new BridgeMcp.CapacitySource(profile -> 0, profile -> 2,
() -> Set.of("terra"), () -> 0), new BridgeMcp.HealthCoverageSource(() -> "off"),
source, Map.of(), ""));
assertTrue(out.contains("\"free\":2"), out);
assertFalse(out.contains("quarantinedForSeconds"), out);
}
@Test
void listReportsLeadsAndFlagsTheCallersOwnRow() {
FakeHerdr h = new FakeHerdr();
@@ -637,6 +773,52 @@ class BridgeMcpTest {
assertEquals("blocked", textOf(res));
}
/**
* CB-582: a lead polling {@code bridge_status} on its normal cadence — not {@code bridge_poll}
* — must also see a worker's open async {@code bridge_ask} question, since the reverse-rendezvous
* window it opened with is far shorter than that cadence.
*/
@Test
void statusReportsAnOpenQuestionWhenTheWorkerIsMidAsk() throws Exception {
String ticket = messages.sendAsync(T, "task that asks");
long deadline = System.currentTimeMillis() + 3000;
while (!rendezvous.isWaiting(T) && System.currentTimeMillis() < deadline) {
Thread.sleep(5);
}
assertTrue(rendezvous.isWaiting(T), "sendAsync should have opened its rendezvous waiter");
CompletableFuture<MessageService.AskResult> ask =
CompletableFuture.supplyAsync(() -> messages.ask(T, "which config file?", 5000));
MessageService.TaskView asking;
deadline = System.currentTimeMillis() + 3000;
do {
asking = messages.poll(ticket);
Thread.sleep(5);
} while (asking.phase() != MessageService.Phase.ASKING && System.currentTimeMillis() < deadline);
assertEquals(MessageService.Phase.ASKING, asking.phase());
McpSchema.CallToolResult res = BridgeMcp.status(messages, T);
assertNotEquals(Boolean.TRUE, res.isError());
String out = textOf(res);
assertTrue(out.startsWith("idle"), "the live status must still lead the text: " + out);
assertTrue(out.contains("which config file?"), "the question text must be shown: " + out);
assertTrue(out.contains("turnId=\"" + asking.turnId() + "\""), "the turnId must be shown: " + out);
assertTrue(out.contains("(ticket " + ticket + ")"), "the ticket must be shown: " + out);
// Clean up the still-open ask so the background thread does not linger past the test.
String turnId = asking.turnId();
CompletableFuture<MessageService.Reply> answer = CompletableFuture.supplyAsync(
() -> messages.answer(turnId, "config.yaml", 5000));
assertEquals("config.yaml", ask.get(5, TimeUnit.SECONDS).answer());
deadline = System.currentTimeMillis() + 3000;
while (!rendezvous.isWaiting(T) && System.currentTimeMillis() < deadline) {
Thread.sleep(5);
}
assertTrue(rendezvous.resolve(T, "done"));
answer.get(5, TimeUnit.SECONDS);
}
// --- bridge_whoami: the caller's own identity, so an agent never has to guess its role -------
@Test
@@ -783,7 +965,8 @@ class BridgeMcpTest {
void spawnDefaultsTheRoleToDev() {
FakeHerdr h = new FakeHerdr();
McpSchema.CallToolResult res = BridgeMcp.spawn(
sessionManager(h, "http://gx00.gw:8000", Set.of("gx00.gw")), null, null, null, null, null, null);
sessionManager(h, "http://gx00.gw:8000", Set.of("gx00.gw")), null, null, null, null, null, null,
null, null);
assertNotEquals(Boolean.TRUE, res.isError());
assertTrue(textOf(res).contains("\"role\":\"dev\""), textOf(res));
@@ -794,7 +977,7 @@ class BridgeMcpTest {
FakeHerdr h = new FakeHerdr();
McpSchema.CallToolResult res = BridgeMcp.spawn(
sessionManager(h, "http://gx00.gw:8000", Set.of("gx00.gw")), null, "architect",
null, null, null, null);
null, null, null, null, null, null);
assertNotEquals(Boolean.TRUE, res.isError());
assertTrue(textOf(res).contains("\"role\":\"architect\""), textOf(res));
@@ -805,9 +988,36 @@ class BridgeMcpTest {
FakeHerdr h = new FakeHerdr();
McpSchema.CallToolResult res = BridgeMcp.spawn(
sessionManager(h, "http://gx00.gw:8000", Set.of("gx00.gw")), null, "worker",
null, null, null, null);
null, null, null, null, null, null);
assertEquals(Boolean.TRUE, res.isError());
assertTrue(textOf(res).contains("architect, dev, reviewer"), textOf(res));
}
// ── CB-584: bridge_spawn accepts sessionName/resumeSessionId; roster shows agentSessionId ──
@Test
void spawnWithResumeSessionIdPutsTheIdOnTheRoster() {
FakeHerdr h = new FakeHerdr();
SessionManager sessions = sessionManager(h, "http://gx00.gw:8000", Set.of("gx00.gw"));
McpSchema.CallToolResult res = BridgeMcp.spawn(sessions, "ltms-local", null, null, null, null, null,
null, "cb-resume-99");
assertNotEquals(Boolean.TRUE, res.isError(), textOf(res));
String listOut = textOf(BridgeMcp.listFleet(
workerService(h, "http://gx00.gw:8000", Set.of("gx00.gw")), sessions, Map.of(), ""));
assertTrue(listOut.contains("\"agentSessionId\":\"cb-resume-99\""), listOut);
}
@Test
void spawnRefusesResumeSessionIdWithoutAnExplicitProfile() {
FakeHerdr h = new FakeHerdr();
McpSchema.CallToolResult res = BridgeMcp.spawn(
sessionManager(h, "http://gx00.gw:8000", Set.of("gx00.gw")), null, null,
null, null, null, null, null, "cb-resume-99");
assertEquals(Boolean.TRUE, res.isError());
assertTrue(textOf(res).contains("explicit profile"), textOf(res));
}
}
@@ -1,5 +1,8 @@
package dev.ltms.bridged.member;
import ch.qos.logback.classic.Logger;
import ch.qos.logback.classic.spi.ILoggingEvent;
import ch.qos.logback.core.read.ListAppender;
import dev.ltms.bridged.config.BridgedConfig;
import dev.ltms.bridged.guard.GuardException;
import dev.ltms.bridged.guard.SubscriptionGuard;
@@ -12,6 +15,7 @@ import dev.ltms.bridged.peer.PeerHandle;
import dev.ltms.bridged.peer.PeerUnreachableException;
import dev.ltms.bridged.peer.SpawnRequest;
import org.junit.jupiter.api.Test;
import org.slf4j.LoggerFactory;
import java.util.List;
import java.util.Map;
@@ -676,6 +680,174 @@ class ClaudeCodeLauncherTest {
"the guard-checked baseUrl must win over any env: entry, or the boundary is bypassable");
}
// --- CB-596: config-driven member-credential policy (replaces CB-592's hardcoded single name) --
/** A representative {@code memberCredentials} — 4 allowed, 4 blocked, matching the real ticket shape. */
private static final BridgedConfig.MemberCredentials TEST_MEMBER_CREDENTIALS = new BridgedConfig.MemberCredentials(
null,
List.of("AI_GATEWAY_TOKEN", "WORKER_GITEA_TOKEN", "CONTEXT7_TOKEN", "GITEA_HOST"),
List.of("AI_GATEWAY_TOKEN", "WORKER_GITEA_TOKEN", "CONTEXT7_TOKEN", "GITEA_HOST",
"GITEA_ACCESS_TOKEN", "GITLAB_PERSONAL_ACCESS_TOKEN", "TS_AUTHKEY", "HASS_TOKEN"));
private ClaudeCodeLauncher serviceWithCredentials(FakeHerdr herdr, BridgedConfig.MemberCredentials creds) {
BridgedConfig.Profile cfg = new BridgedConfig.Profile(
"ltms-local", "http://gx00.gw:8000", "coder", null, "BRIDGED_WORKER_TOKEN",
List.of("claude"), "tab", "bridged-workers", "worker: {profile} #{n}", null, null, null);
return new ClaudeCodeLauncher(new AgentControl(herdr), new WorkspaceControl(herdr),
new SubscriptionGuard(Set.of("gx00.gw")), Map.of(cfg.profile(), cfg), cfg.profile(),
_ -> null, 0, System::currentTimeMillis, () -> {}, null, () -> creds);
}
/**
* herdr's env map is an overlay onto its own (login-shell) process environment, so a worker
* inherits whatever the daemon's shell carries — including the operator's own credentials — for
* every key {@code baseEnv} does not explicitly shadow. This pins that every {@code known} name
* NOT also {@code allow}-ed gets an explicit (non-blank) sentinel overlay, whatever the profile
* is. Asserted against what tab.create's params actually carry, not an internal map built in the
* test (gitea #82).
*
* <p>Scope, measured on a live pane 2026-08-15 (CB-592): this pins what the launcher SENDS, and
* that is all a unit test can pin. It does not prove the value survives, and for names the
* operator's secrets.sh also exports it does not: the pane runs a login shell that puts the real
* value back over this sentinel unless the export is guarded on BRIDGED_MEMBER — see
* everySpawnMarksThePaneAsAMember below.
*/
@Test
void everyKnownNameNotAllowedIsShadowedWithTheSentinel() {
FakeHerdr herdr = new FakeHerdr();
serviceWithCredentials(herdr, TEST_MEMBER_CREDENTIALS).spawn();
Map<String, String> env = startEnv(herdr);
for (String blocked : List.of("GITEA_ACCESS_TOKEN", "GITLAB_PERSONAL_ACCESS_TOKEN", "TS_AUTHKEY", "HASS_TOKEN")) {
String shadowed = env.get(blocked);
assertNotNull(shadowed, blocked + " must be explicitly overlaid, not left unmentioned");
assertFalse(shadowed.isBlank(), blocked + "'s overlay value must be non-blank");
}
}
/**
* An {@code allow}-ed name must get NO overlay entry at all — any entry, blank or not, risks
* overriding the real value the pane needs, and the whole point of {@code allow} is that the
* pane's own inherited value passes through untouched.
*/
@Test
void everyAllowedNameGetsNoOverlayEntrySoTheRealValuePassesThrough() {
FakeHerdr herdr = new FakeHerdr();
serviceWithCredentials(herdr, TEST_MEMBER_CREDENTIALS).spawn();
Map<String, String> env = startEnv(herdr);
for (String allowed : List.of("AI_GATEWAY_TOKEN", "WORKER_GITEA_TOKEN", "CONTEXT7_TOKEN", "GITEA_HOST")) {
assertFalse(env.containsKey(allowed),
allowed + " is allow-listed — the launcher must not mention it at all");
}
}
/**
* No {@code memberCredentials} configured (the pre-CB-596 constructor overloads still used
* throughout this file, and the shape a fresh {@code bridged.yaml} with no memberCredentials:
* block resolves to) blocks NOTHING. This documents the transitional gap rather than hiding it —
* see {@code HerdrPeerLauncher#applyMemberCredentialPolicy}'s javadoc.
*/
@Test
void noMemberCredentialsConfiguredBlocksNothing() {
FakeHerdr herdr = new FakeHerdr();
service(herdr, List.of("claude"), null).spawn();
assertNull(startEnv(herdr).get("GITEA_ACCESS_TOKEN"),
"with no memberCredentials configured, nothing is shadowed — config must supply the policy");
}
/**
* No profile — present or future — may restore a blocked name by naming it in {@code env:}.
* The shadow is applied after the profile's own env in {@link HerdrPeerLauncher#baseEnv}
* precisely so this can never happen; this test pins that ordering.
*/
@Test
void aProfileEnvEntryCannotRestoreABlockedName() {
FakeHerdr herdr = new FakeHerdr();
BridgedConfig.Profile cfg = new BridgedConfig.Profile(
"ltms-local", "http://gx00.gw:8000", "coder", null, "BRIDGED_WORKER_TOKEN",
List.of("claude"), "tab", "bridged-workers", "w #{n}", null, null, null, null, null,
null, Map.of("GITEA_ACCESS_TOKEN", "admin-secret-from-profile-config"), null, null);
new ClaudeCodeLauncher(new AgentControl(herdr), new WorkspaceControl(herdr),
new SubscriptionGuard(Set.of("gx00.gw")), Map.of(cfg.profile(), cfg), cfg.profile(),
_ -> null, 0, System::currentTimeMillis, () -> {}, null, () -> TEST_MEMBER_CREDENTIALS)
.spawn(cfg.profile(), null, null);
assertNotEquals("admin-secret-from-profile-config", startEnv(herdr).get("GITEA_ACCESS_TOKEN"),
"a profile's own env: must not be able to smuggle a blocked name back in");
}
/**
* CB-596 criterion 4: a credential-shaped host env var name on neither {@code known} nor
* {@code allow} is not silently allowed — it must be reported (never its value). This pins the
* WARN naming the gap, using an injected host-env-names source rather than the real
* {@code System.getenv()} so the test is deterministic.
*/
@Test
void aCredentialShapedNameOnNeitherListIsLoggedAsAGap() {
FakeHerdr herdr = new FakeHerdr();
BridgedConfig.Profile cfg = new BridgedConfig.Profile(
"ltms-local", "http://gx00.gw:8000", "coder", null, "BRIDGED_WORKER_TOKEN",
List.of("claude"), "tab", "bridged-workers", "worker: {profile} #{n}", null, null, null);
ClaudeCodeLauncher svc = new ClaudeCodeLauncher(new AgentControl(herdr), new WorkspaceControl(herdr),
new SubscriptionGuard(Set.of("gx00.gw")), Map.of(cfg.profile(), cfg), cfg.profile(),
_ -> null, 0, System::currentTimeMillis, () -> {}, null, () -> TEST_MEMBER_CREDENTIALS,
() -> Set.of("PATH", "HOME", "AI_GATEWAY_TOKEN", "A_BRAND_NEW_SECRET_TOKEN"));
Logger logger = (Logger) LoggerFactory.getLogger(HerdrPeerLauncher.class);
ListAppender<ILoggingEvent> appender = new ListAppender<>();
appender.start();
logger.addAppender(appender);
try {
svc.spawn();
} finally {
logger.detachAppender(appender);
}
assertTrue(appender.list.stream().anyMatch(e ->
e.getFormattedMessage().contains("memberCredentials gap")
&& e.getFormattedMessage().contains("A_BRAND_NEW_SECRET_TOKEN")),
"the gap must name the unrecognized credential-shaped var, never a value");
assertFalse(appender.list.stream().anyMatch(e -> e.getFormattedMessage().contains("PATH")),
"PATH/HOME are not credential-shaped and must not be reported as a gap");
assertFalse(appender.list.stream().anyMatch(e -> e.getFormattedMessage().contains("AI_GATEWAY_TOKEN")),
"a name already on allow: is covered, not a gap");
}
/**
* The half of CB-592 that can actually survive the pane's login shell. BRIDGED_MEMBER is a name
* secrets.sh never exports, so nothing overwrites it — measured: GITEA_TOKEN is injected the
* same way, is absent from a login shell of its own, and was observed set inside a live member
* pane. It lets the operator guard the admin export with
* `[ -n "${BRIDGED_MEMBER:-}" ] || export GITEA_ACCESS_TOKEN=...`, which is the whole fix.
* Pinned here so a refactor cannot drop the marker and quietly un-guard every member (#77).
*/
@Test
void everySpawnMarksThePaneAsAMember() {
FakeHerdr herdr = new FakeHerdr();
service(herdr, List.of("claude"), null).spawn();
assertEquals("1", startEnv(herdr).get("BRIDGED_MEMBER"),
"every member pane must be marked, or a shell file cannot tell it apart from the operator's");
}
/** A profile must not be able to hide that its pane is a member, for the same reason as above. */
@Test
void aProfileEnvEntryCannotClearTheMemberMarker() {
FakeHerdr herdr = new FakeHerdr();
BridgedConfig.Profile cfg = new BridgedConfig.Profile(
"ltms-local", "http://gx00.gw:8000", "coder", null, "BRIDGED_WORKER_TOKEN",
List.of("claude"), "tab", "bridged-workers", "w #{n}", null, null, null, null, null,
null, Map.of("BRIDGED_MEMBER", ""), null, null);
new ClaudeCodeLauncher(new AgentControl(herdr), new WorkspaceControl(herdr),
new SubscriptionGuard(Set.of("gx00.gw")), Map.of(cfg.profile(), cfg), cfg.profile(),
_ -> null).spawn();
assertEquals("1", startEnv(herdr).get("BRIDGED_MEMBER"),
"a profile's own env: must not be able to unmark its pane");
}
// ── CB-533: the model is pinned on the command line, not only in the environment ────────────
/** A launcher for a profile identical but for its {@code model:} — the only variable here. */
@@ -733,7 +905,7 @@ class ClaudeCodeLauncherTest {
return new BridgedConfig.Profile(
profile, baseUrl, "sonnet", null, "BRIDGED_WORKER_TOKEN",
List.of("ccs", profile), "tab", "bridged-workers", "w #{n}", null, null, null,
null, null, null, Map.of(), null, null, true);
null, null, null, Map.of(), null, null, true, null);
}
@Test
@@ -819,7 +991,7 @@ class ClaudeCodeLauncherTest {
null, null, null,
Map.of("ANTHROPIC_BASE_URL", "http://evil.example.com",
"ANTHROPIC_AUTH_TOKEN", "sk-ant-bad", "JAVA_HOME", "/opt/jdk"),
null, null, true);
null, null, true, null);
new ClaudeCodeLauncher(new AgentControl(herdr), new WorkspaceControl(herdr),
new SubscriptionGuard(Set.of("gx00.gw")), Map.of(cfg.profile(), cfg), cfg.profile(),
_ -> "would-be-token").spawn();
@@ -10,11 +10,13 @@ import dev.ltms.bridged.herdr.AgentControl;
import dev.ltms.bridged.herdr.FakeHerdr;
import dev.ltms.bridged.herdr.WorkspaceControl;
import dev.ltms.bridged.peer.Capability;
import dev.ltms.bridged.peer.CharterReceipt;
import dev.ltms.bridged.peer.MemberRole;
import dev.ltms.bridged.peer.PeerHandle;
import dev.ltms.bridged.peer.PeerLauncher;
import dev.ltms.bridged.peer.PeerUnreachableException;
import dev.ltms.bridged.peer.SpawnRequest;
import dev.ltms.bridged.placement.BackendQuarantine;
import dev.ltms.bridged.placement.PlacementException;
import dev.ltms.bridged.placement.PlacementPolicies;
import org.junit.jupiter.api.Test;
@@ -26,6 +28,8 @@ import java.util.LinkedHashMap;
import java.util.List;
import java.util.Map;
import java.util.Set;
import java.util.concurrent.TimeUnit;
import java.util.concurrent.atomic.AtomicLong;
import static org.junit.jupiter.api.Assertions.*;
@@ -93,6 +97,8 @@ class CompositePeerLauncherTest {
@Override public String id() { return "pane-" + p; }
@Override public String terminalId() { return "term-" + p; }
@Override public String profile() { return p; }
@Override public String agentSessionId() { return null; }
@Override public CharterReceipt charterReceipt() { return null; }
};
}
@@ -126,6 +132,13 @@ class CompositePeerLauncherTest {
weight, maxLoad);
}
private static BridgedConfig.Profile stubWorker(String profile, String credentialId) {
return new BridgedConfig.Profile(profile, "http://gx00.gw:8000", "coder",
null, "BRIDGED_WORKER_TOKEN", List.of("claude"), "tab", "bridged-workers",
"w #{n}", null, null, null, null, null, null, null, null, null,
null, null, credentialId);
}
/**
* An <em>order-preserving</em> profile map. Never {@code Map.of} here: its iteration order is
* salted per JVM run, and the weighted policy breaks an exact-weight tie on candidate order —
@@ -324,6 +337,42 @@ class CompositePeerLauncherTest {
assertEquals(10, b);
}
@Test
void weightedPolicySkipsWeightZeroProfileOnUnqualifiedSpawn() {
// CB-554: weight: 0 must exclude a profile from automatic placement, not coerce to 1.0.
FakeHerdr herdr = new FakeHerdr();
Map<String, BridgedConfig.Profile> profiles = ordered(
"a", stubWorker("a", 0.0f, null),
"b", stubWorker("b", 1.0f, null));
StubLauncher adapter = new StubLauncher("claude", herdr, profiles, "a", Set.of());
CompositePeerLauncher composite = new CompositePeerLauncher(
List.of(adapter), "a", profiles, PlacementPolicies.weighted(), _ -> 0);
for (int i = 0; i < 5; i++) {
PeerHandle h = composite.spawn(new SpawnRequest(null, null, null));
assertEquals("b", h.profile(), "profile a has weight 0, so every unqualified spawn must land on b");
}
assertEquals(0, adapter.spawnCount("a"), "a is never chosen automatically");
}
@Test
void explicitSpawnStillSucceedsOnWeightZeroProfile() {
// CB-554: weight: 0 excludes a profile from AUTOMATIC selection only — an explicit
// bridge_spawn{profile:"a"} must still work exactly as today (e.g. `opus` on the
// operator's own subscription, kept weight-0 so it is never picked automatically).
FakeHerdr herdr = new FakeHerdr();
Map<String, BridgedConfig.Profile> profiles = ordered(
"a", stubWorker("a", 0.0f, null),
"b", stubWorker("b", 1.0f, null));
StubLauncher adapter = new StubLauncher("claude", herdr, profiles, "b", Set.of());
CompositePeerLauncher composite = new CompositePeerLauncher(
List.of(adapter), "b", profiles, PlacementPolicies.weighted(), _ -> 0);
PeerHandle h = composite.spawn(new SpawnRequest("a", null, null));
assertEquals("a", h.profile(), "naming a weight-0 profile explicitly bypasses placement and still spawns it");
assertEquals(1, adapter.spawnCount("a"));
}
@Test
void failoverRetriesNextCandidateWhenProfileIsUnreachable() {
FakeHerdr herdr = new FakeHerdr();
@@ -425,6 +474,27 @@ class CompositePeerLauncherTest {
assertEquals("a", h.profile(), "a profile with no maxLoad is never capped, however many live workers");
}
@Test
void explicitSpawnOnMaxLoadZeroProfileIsRefusedEvenWithZeroLiveWorkers() {
// CB-585: before the fix, maxLoad: 0 normalised to null (unlimited) in the compact
// constructor, so this exact case — naming a zero-cap profile explicitly, with nothing
// live on it yet — would have spawned instead of refusing. "at most zero members" must
// hold even when the profile is named directly, not only against automatic placement.
FakeHerdr herdr = new FakeHerdr();
Map<String, BridgedConfig.Profile> profiles = ordered(
"a", stubWorker("a", 1.0f, 0),
"b", stubWorker("b"));
StubLauncher adapter = new StubLauncher("claude", herdr, profiles, "a", Set.of());
CompositePeerLauncher composite = new CompositePeerLauncher(
List.of(adapter), "a", profiles, PlacementPolicies.fixed(), _ -> 0);
PlacementException e = assertThrows(PlacementException.class,
() -> composite.spawn(new SpawnRequest("a", null, null)));
assertTrue(e.getMessage().contains("'a'"), "message names the profile: " + e.getMessage());
assertTrue(e.getMessage().contains("0 cap"), "message names the zero cap: " + e.getMessage());
assertEquals(0, adapter.spawnCount("a"), "a maxLoad: 0 profile accepts no explicit spawn");
}
@Test
void emptyCandidateSetThrowsClearException() {
FakeHerdr herdr = new FakeHerdr();
@@ -552,4 +622,105 @@ class CompositePeerLauncherTest {
new SpawnRequest(null, null, null, null, null, MemberRole.DEV)));
assertTrue(e.getMessage().contains("maxLoad"), e.getMessage());
}
// ── CB-578 stage B: a BACKEND_EXHAUSTED classification quarantines the credential ──────────
@Test
void explicitSpawnOntoAQuarantinedProfileIsRefusedNamingTheCredentialAndRemainingTime() {
FakeHerdr herdr = new FakeHerdr();
Map<String, BridgedConfig.Profile> profiles = ordered(
"sol", stubWorker("sol", "shared-openai"),
"terra", stubWorker("terra", "shared-openai"));
StubLauncher adapter = new StubLauncher("claude", herdr, profiles, "sol", Set.of());
BackendQuarantine quarantine = new BackendQuarantine(() -> 0L, TimeUnit.MINUTES.toNanos(30));
quarantine.quarantine("shared-openai");
CompositePeerLauncher composite = new CompositePeerLauncher(List.of(adapter), "sol", profiles,
PlacementPolicies.fixed(), _ -> 0, null, quarantine);
PlacementException e = assertThrows(PlacementException.class,
() -> composite.spawn(new SpawnRequest("sol", null, null)));
assertTrue(e.getMessage().contains("sol"), "message names the profile: " + e.getMessage());
assertTrue(e.getMessage().contains("shared-openai"), "message names the credential: " + e.getMessage());
assertTrue(e.getMessage().contains("1800"), "message names roughly when it lifts: " + e.getMessage());
assertEquals(0, adapter.spawnCount("sol"), "the quarantined profile is never delegated to");
}
/**
* The part CB-578 stage B calls out as easy to get wrong: sol and terra are two different
* profiles sharing one OpenAI credential. Quarantining because of an exhaustion classified on
* ONE of them must lock out the other too, or the fleet just walks onto the same dead account
* under the sibling's name.
*/
@Test
void twoProfilesSharingACredentialAreBothQuarantinedByOneExhaustionEvent() {
FakeHerdr herdr = new FakeHerdr();
Map<String, BridgedConfig.Profile> profiles = ordered(
"sol", stubWorker("sol", "shared-openai"),
"terra", stubWorker("terra", "shared-openai"));
StubLauncher adapter = new StubLauncher("claude", herdr, profiles, "sol", Set.of());
BackendQuarantine quarantine = new BackendQuarantine(() -> 0L, TimeUnit.MINUTES.toNanos(30));
// Only "sol" was classified BACKEND_EXHAUSTED — but the two profiles share one credential.
quarantine.quarantine("shared-openai");
CompositePeerLauncher composite = new CompositePeerLauncher(List.of(adapter), "sol", profiles,
PlacementPolicies.fixed(), _ -> 0, null, quarantine);
assertThrows(PlacementException.class, () -> composite.spawn(new SpawnRequest("sol", null, null)),
"sol was the one classified exhausted");
assertThrows(PlacementException.class, () -> composite.spawn(new SpawnRequest("terra", null, null)),
"terra shares sol's credential, so it must be locked out too");
assertEquals(0, adapter.spawnCount("sol"));
assertEquals(0, adapter.spawnCount("terra"));
}
@Test
void placementSkipsAQuarantinedProfileAndRoutesToAnUnquarantinedOne() {
FakeHerdr herdr = new FakeHerdr();
Map<String, BridgedConfig.Profile> profiles = ordered(
"sol", stubWorker("sol", "shared-openai"),
"b", stubWorker("b"));
StubLauncher adapter = new StubLauncher("claude", herdr, profiles, "sol", Set.of());
BackendQuarantine quarantine = new BackendQuarantine(() -> 0L, TimeUnit.MINUTES.toNanos(30));
quarantine.quarantine("shared-openai");
CompositePeerLauncher composite = new CompositePeerLauncher(List.of(adapter), "sol", profiles,
PlacementPolicies.weighted(), _ -> 0, null, quarantine);
PeerHandle h = composite.spawn(new SpawnRequest(null, null, null));
assertEquals("b", h.profile(), "sol is quarantined, so an unqualified spawn must land on b");
assertEquals(0, adapter.spawnCount("sol"));
}
@Test
void aQuarantineLiftsOnTheInjectedClockAndTheProfileBecomesSpawnableAgain() {
FakeHerdr herdr = new FakeHerdr();
Map<String, BridgedConfig.Profile> profiles = ordered(
"sol", stubWorker("sol", "shared-openai"),
"b", stubWorker("b"));
StubLauncher adapter = new StubLauncher("claude", herdr, profiles, "sol", Set.of());
AtomicLong nowNanos = new AtomicLong(0L);
BackendQuarantine quarantine = new BackendQuarantine(nowNanos::get, TimeUnit.MINUTES.toNanos(30));
quarantine.quarantine("shared-openai");
CompositePeerLauncher composite = new CompositePeerLauncher(List.of(adapter), "sol", profiles,
PlacementPolicies.fixed(), _ -> 0, null, quarantine);
assertThrows(PlacementException.class, () -> composite.spawn(new SpawnRequest("sol", null, null)),
"still inside the cooldown");
nowNanos.set(TimeUnit.MINUTES.toNanos(31));
PeerHandle h = composite.spawn(new SpawnRequest("sol", null, null));
assertEquals("sol", h.profile(), "the cooldown expired on the injected clock — sol is spawnable again");
assertEquals(1, adapter.spawnCount("sol"));
}
@Test
void aFleetWithNoExhaustedPatternAnywhereBehavesExactlyAsBeforeQuarantineExisted() {
// BackendQuarantine.none() is the inert stand-in every constructor already defaults to when
// no quarantine is wired — the 2/5/6-arg constructors used throughout this file all exercise
// it. This test pins that an explicit .none() also never refuses a spawn, for any profile.
FakeHerdr herdr = new FakeHerdr();
PeerLauncher composite = composite(herdr);
assertDoesNotThrow(() -> composite.spawn(new SpawnRequest("claude", null, null)));
assertDoesNotThrow(() -> composite.spawn(new SpawnRequest("gemini", null, null)));
}
}
@@ -1,13 +1,19 @@
package dev.ltms.bridged.member;
import ch.qos.logback.classic.Level;
import ch.qos.logback.classic.Logger;
import ch.qos.logback.classic.spi.ILoggingEvent;
import ch.qos.logback.core.read.ListAppender;
import dev.ltms.bridged.config.BridgedConfig;
import dev.ltms.bridged.herdr.AgentControl;
import dev.ltms.bridged.herdr.FakeHerdr;
import dev.ltms.bridged.herdr.WorkspaceControl;
import dev.ltms.bridged.peer.Capability;
import dev.ltms.bridged.peer.CharterReceipt;
import dev.ltms.bridged.peer.MemberRole;
import dev.ltms.bridged.peer.SpawnRequest;
import org.junit.jupiter.api.Test;
import org.slf4j.LoggerFactory;
import java.util.ArrayList;
import java.util.List;
@@ -17,6 +23,8 @@ import java.util.concurrent.atomic.AtomicReference;
import java.util.function.Supplier;
import static org.junit.jupiter.api.Assertions.assertEquals;
import static org.junit.jupiter.api.Assertions.assertFalse;
import static org.junit.jupiter.api.Assertions.assertTrue;
class HerdrPeerLauncherCharterTest {
@@ -38,10 +46,72 @@ class HerdrPeerLauncherCharterTest {
"a role charter does not depend on an MCP mount");
}
@Test
void panePlacementSpawnLogNeverContainsTheCharterText() {
// A pane-placement spawn used to log the whole argv (CB-571), and the charter travels
// inside argv — so the charter text leaked to the daemon log. Prove the legacy pane path
// now redacts it to its digest.
String secret = "TOP SECRET charter marker 99x"; // distinctive, so a leak is unambiguous
AtomicReference<BridgedConfig.Fleet> fleet = new AtomicReference<>(fleet(Map.of("dev", secret)));
CharterArgLauncher launcher = new CharterArgLauncher(fleet::get);
Logger logger = (Logger) LoggerFactory.getLogger(HerdrPeerLauncher.class);
Level previous = logger.getLevel();
ListAppender<ILoggingEvent> appender = new ListAppender<>();
appender.start();
logger.addAppender(appender);
logger.setLevel(Level.INFO); // the test logback sets dev.ltms.bridged to WARN; a leak lives at INFO
try {
launcher.spawn(new SpawnRequest("mcp", null, null, null, null, MemberRole.DEV));
String all = String.join("\n", appender.list.stream().map(ILoggingEvent::getFormattedMessage).toList());
assertFalse(all.contains(secret),
"the pane-placement spawn log must not contain the charter text; got:\n" + all);
// The "mcp" profile composes role + reply charter; the digest must match that composed
// string (the exact bytes the adapter receives), proving the redaction hashes and
// removes the real, full charter — not some placeholder.
String composed = secret + "\n\n" + HerdrPeerLauncher.REPLY_CHARTER;
assertTrue(all.contains("<charter sha256=" + CharterReceipt.digestOf(composed) + ">"),
"the charter argv argument should be replaced by its digest; got:\n" + all);
} finally {
logger.setLevel(previous);
logger.detachAppender(appender);
}
}
private static BridgedConfig.Fleet fleet(Map<String, String> charters) {
return new BridgedConfig.Fleet(Map.of(), Map.of(), Map.of(), Map.of(), charters, null);
}
private static BridgedConfig.Profile profile(String name, String mcpUrl) {
return new BridgedConfig.Profile(name, "http://gx00.gw:8000", null, null,
"BRIDGED_WORKER_TOKEN", List.of("test"), "pane", null, null, mcpUrl, null, null);
}
/**
* A launcher whose {@code buildLaunch} hands the composed charter to herdr as one argv element
* (what the claude-cod adapter does), so a pane-placement spawn log would print it unless the
* base redacts it.
*/
private static final class CharterArgLauncher extends HerdrPeerLauncher {
CharterArgLauncher(Supplier<BridgedConfig.Fleet> fleet) {
super("test", new AgentControl(new FakeHerdr()), new WorkspaceControl(new FakeHerdr()),
Map.of("mcp", profile("mcp", "http://bridge")),
"mcp", _ -> null, 0, () -> 0L, () -> { }, fleet);
}
@Override
protected Launch buildLaunch(BridgedConfig.Profile cfg, LaunchSpec spec) {
return new Launch(Map.of(), List.of("test", spec.charter() == null ? "none" : spec.charter()));
}
@Override
public Set<Capability> capabilities() {
return Set.of();
}
}
private static final class CapturingLauncher extends HerdrPeerLauncher {
private final List<LaunchSpec> specs = new ArrayList<>();
@@ -62,10 +132,5 @@ class HerdrPeerLauncherCharterTest {
public Set<Capability> capabilities() {
return Set.of();
}
private static BridgedConfig.Profile profile(String name, String mcpUrl) {
return new BridgedConfig.Profile(name, "http://gx00.gw:8000", null, null,
"BRIDGED_WORKER_TOKEN", List.of("test"), "pane", null, null, mcpUrl, null, null);
}
}
}
@@ -7,6 +7,7 @@ import dev.ltms.bridged.herdr.AgentControl;
import dev.ltms.bridged.herdr.FakeHerdr;
import dev.ltms.bridged.herdr.WorkspaceControl;
import dev.ltms.bridged.peer.Capability;
import dev.ltms.bridged.peer.CharterReceipt;
import dev.ltms.bridged.peer.PeerHandle;
import dev.ltms.bridged.peer.PeerUnreachableException;
import dev.ltms.bridged.peer.SpawnRequest;
@@ -51,6 +52,14 @@ class OpenCodeLauncherTest {
0, System::currentTimeMillis, () -> { }, configRoot, configRoot, fleet);
}
private static OpenCodeLauncher serviceWithCredentials(FakeHerdr herdr, Path configRoot,
BridgedConfig.Profile cfg,
BridgedConfig.MemberCredentials creds) {
return new OpenCodeLauncher(new AgentControl(herdr), new WorkspaceControl(herdr),
Map.of(cfg.profile(), cfg), cfg.profile(), k -> "GITEA_ACCESS_TOKEN".equals(k) ? "tok" : null,
0, System::currentTimeMillis, () -> { }, configRoot, configRoot, null, () -> creds);
}
@SuppressWarnings("unchecked")
private static Map<String, Object> lastStart(FakeHerdr herdr) {
return (Map<String, Object>) herdr.lastCall("agent.start").params();
@@ -207,6 +216,36 @@ class OpenCodeLauncherTest {
"a git-token profile gets the peer-neutral GITEA_TOKEN grant, same as Claude");
}
/**
* CB-596: the config-driven shadow lives in {@link HerdrPeerLauncher#baseEnv}, shared by every
* adapter — this pins that the opencode path gets it too, not just Claude's. See the matching
* tests in {@code ClaudeCodeLauncherTest} for the full rationale (gitea #82, superseding CB-592's
* single hardcoded name).
*/
@Test
void aKnownNameNotAllowedIsShadowedWithTheSentinel(@TempDir Path root) {
FakeHerdr herdr = new FakeHerdr();
BridgedConfig.MemberCredentials creds = new BridgedConfig.MemberCredentials(
null, List.of("AI_GATEWAY_TOKEN"), List.of("AI_GATEWAY_TOKEN", "GITEA_ACCESS_TOKEN"));
serviceWithCredentials(herdr, root, opencodeCfg(null, null, null), creds).spawn();
String shadowed = startEnv(herdr).get("GITEA_ACCESS_TOKEN");
assertNotNull(shadowed, "GITEA_ACCESS_TOKEN must be explicitly overlaid, not left unmentioned");
assertFalse(shadowed.isBlank(), "a blank overlay value's override behaviour is unverified — must be non-blank");
assertFalse(startEnv(herdr).containsKey("AI_GATEWAY_TOKEN"),
"an allow-listed name must get no overlay entry at all");
}
/** No {@code memberCredentials} configured (the pre-CB-596 constructor overloads) blocks nothing. */
@Test
void noMemberCredentialsConfiguredBlocksNothing(@TempDir Path root) {
FakeHerdr herdr = new FakeHerdr();
service(herdr, root, opencodeCfg(null, null, null)).spawn();
assertNull(startEnv(herdr).get("GITEA_ACCESS_TOKEN"),
"with no memberCredentials configured, nothing is shadowed — config must supply the policy");
}
@Test
void capabilitiesDeclareOrphanReapAndMcpAskAndConditionalSelfPr(@TempDir Path root) {
FakeHerdr herdr = new FakeHerdr();
@@ -322,6 +361,26 @@ class OpenCodeLauncherTest {
assertFalse(herdr.called("agent.get"), "no polling when the gate is disabled");
}
@Test
void handleCarriesTheRealCharterReceiptNotTheInterfaceDefault(@TempDir Path root) {
// The base's WorkerHandle computes a real CharterReceipt (CB-571), but the opencode adapter
// wraps it in SessionAwareHandle for lazy session discovery. Before this fix that decorator
// did not override charterReceipt(), so it silently inherited PeerHandle's `null` default
// and the real receipt sitting on its delegate was lost.
FakeHerdr herdr = new FakeHerdr();
BridgedConfig.Fleet fleet = new BridgedConfig.Fleet(Map.of(), Map.of(), Map.of(), Map.of(),
Map.of("dev", "role rule"), null);
PeerHandle handle = service(herdr, root,
opencodeCfg("google/gemini-2.5-pro", "http://127.0.0.1:8765/mcp", null), () -> fleet)
.spawn(new SpawnRequest(null, null, null));
assertNotNull(handle.charterReceipt(),
"an opencode spawn's charterReceipt() must not silently be null");
String composed = "role rule\n\n" + HerdrPeerLauncher.REPLY_CHARTER;
assertEquals(CharterReceipt.digestOf(composed), handle.charterReceipt().charterSha256(),
"the receipt on the wrapped handle must match the exact composed charter bytes");
}
// --- CB-508: pinned OpenAI-compatible endpoint (e.g. a local vLLM) ---------------------------
/** A profile with a baseUrl but no model provider prefix cannot be resolved — fail loudly. */
@@ -16,6 +16,7 @@ import java.util.List;
import java.util.concurrent.TimeUnit;
import static org.junit.jupiter.api.Assertions.assertEquals;
import static org.junit.jupiter.api.Assertions.assertThrows;
import static org.junit.jupiter.api.Assertions.assertTrue;
/**
@@ -166,6 +167,88 @@ class AmqpReplyInboxContractTest {
}
}
@Test
void prefetchBoundsTheHeldBacklog() throws Exception {
String target = "worker-prefetch-" + System.nanoTime();
int prefetch = 4;
int published = 10;
try (AmqpReplyInbox inbox = new AmqpReplyInbox(newConnection(), prefetch);
Connection inspect = newConnection()) {
inbox.own(target);
for (int i = 0; i < published; i++) {
inbox.publish(target, "m" + i, "payload " + i);
}
awaitHeldAtLeast(inbox, target, prefetch);
long depth;
try (Channel ch = inspect.createChannel()) {
depth = ch.queueDeclarePassive(queueName(target)).getMessageCount();
}
assertTrue(depth >= published - prefetch,
"broker should still hold at least " + (published - prefetch)
+ " undelivered messages behind a prefetch of " + prefetch + ", saw " + depth);
drainUntilEmpty(inbox, target, inspect);
}
}
@Test
void unroutablePublishReportsFailureNotSilentSuccess() throws Exception {
String target = "worker-unroutable-" + System.nanoTime();
try (AmqpReplyInbox inbox = AmqpReplyInbox.open(uri())) {
// Deliberately never own(target): the queue is never declared, so the default-exchange
// route to agent.<target>.inbox does not exist and the broker must return the publish.
IllegalStateException ex = assertThrows(IllegalStateException.class,
() -> inbox.publish(target, "m1", "nobody home"));
assertTrue(ex.getMessage() != null && ex.getMessage().toLowerCase().contains("unroutable"),
"expected an unroutable-publish failure, got: " + ex.getMessage());
}
}
@Test
void confirmedPublishDeliversNormally() throws Exception {
String target = "worker-confirm-" + System.nanoTime();
try (AmqpReplyInbox inbox = AmqpReplyInbox.open(uri())) {
inbox.own(target);
inbox.publish(target, "m1", "confirmed delivery"); // must return normally: routed and confirmed
List<ReplyInbox.InboxMessage> got = awaitPeek(inbox, target);
assertEquals(1, got.size());
assertEquals("confirmed delivery", got.getFirst().content());
inbox.ack(target, "m1");
}
}
/** Poll peek until at least {@code n} replies for {@code target} are held, or ~10s elapse. */
@SuppressWarnings("BusyWait")
private static void awaitHeldAtLeast(AmqpReplyInbox inbox, String target, int n) throws InterruptedException {
long deadline = System.nanoTime() + TimeUnit.SECONDS.toNanos(10);
while (inbox.peek(target).size() < n && System.nanoTime() < deadline) {
Thread.sleep(50);
}
}
/**
* Repeatedly ack whatever is currently held (freeing prefetch slots for the next delivery) until
* both the local snapshot and the broker's own queue depth are empty, or ~10s elapse.
*/
@SuppressWarnings("BusyWait")
private static void drainUntilEmpty(AmqpReplyInbox inbox, String target, Connection inspect) throws Exception {
long deadline = System.nanoTime() + TimeUnit.SECONDS.toNanos(10);
while (System.nanoTime() < deadline) {
for (ReplyInbox.InboxMessage msg : inbox.peek(target)) {
inbox.ack(target, msg.msgId());
}
try (Channel ch = inspect.createChannel()) {
if (ch.queueDeclarePassive(queueName(target)).getMessageCount() == 0 && inbox.peek(target).isEmpty()) {
return;
}
}
Thread.sleep(50);
}
throw new AssertionError("did not drain " + target + " to empty within the deadline");
}
/** Poll peek (broker delivery is async) until a reply for {@code target} appears or ~10s elapse. */
@SuppressWarnings("BusyWait") // deliberate poll for async broker delivery, bounded by the deadline
private static List<ReplyInbox.InboxMessage> awaitPeek(AmqpReplyInbox inbox, String target)
@@ -0,0 +1,288 @@
package dev.ltms.bridged.msg;
import com.rabbitmq.client.AMQP;
import com.rabbitmq.client.Channel;
import com.rabbitmq.client.ConfirmCallback;
import com.rabbitmq.client.Connection;
import org.junit.jupiter.api.Test;
import org.junit.jupiter.api.Timeout;
import java.lang.reflect.InvocationHandler;
import java.lang.reflect.Proxy;
import java.util.List;
import java.util.concurrent.CopyOnWriteArrayList;
import java.util.concurrent.CountDownLatch;
import java.util.concurrent.TimeUnit;
import java.util.concurrent.atomic.AtomicInteger;
import java.util.concurrent.atomic.AtomicLong;
import java.util.concurrent.atomic.AtomicReference;
import static org.junit.jupiter.api.Assertions.assertNotNull;
import static org.junit.jupiter.api.Assertions.assertNull;
import static org.junit.jupiter.api.Assertions.assertTrue;
/**
* CB-528 follow-up: {@link AmqpReplyInbox#failPendingPublishesOnRecovery()} must not fail a publish
* that registers concurrently with (and was not yet visible when) the recovery sweep began — that
* would report a publish that actually succeeded as failed, and {@code MessageService.reply} retries
* with a fresh {@code msgId}, so the reply is delivered twice. A live broker reconnect cannot be
* forced reliably, so this drives {@link AmqpReplyInbox#failPendingPublishesOnRecovery} and a real
* {@link AmqpReplyInbox#publish} against each other directly, against fake AMQP channels built with
* {@link Proxy} (no mocking library is on the classpath).
*
* <p>Also covers Finding 2 (CB-528 follow-up): {@link AmqpReplyInbox#close()} must fail an in-flight
* publish promptly instead of leaving it to idle out the 10s confirm timeout.
*/
class AmqpReplyInboxRecoveryRaceTest {
/** Large enough that thousands of entries are still unprocessed by the time the very first one
* is observed as failed (see {@code sweepIsHoldingTheLock} below) — that gap is what makes the
* head start deterministic instead of a coin flip. 100,000 gave the same guarantee but made the
* test far more expensive than the guarantee needs; the ordering no longer depends on a timing
* window sized to the full backlog; just to the tail of it. */
private static final int STALE_PUBLISHES = 2_000;
@Test
@Timeout(30)
void recoverySweepDoesNotFailAPublishThatRegistersWhileItIsRunning() throws Exception {
AtomicLong seqCounter = new AtomicLong();
List<Long> seqOrder = new CopyOnWriteArrayList<>();
List<String> msgIdOrder = new CopyOnWriteArrayList<>();
AtomicReference<ConfirmCallback> ackCallback = new AtomicReference<>();
AtomicReference<ConfirmCallback> nackCallback = new AtomicReference<>();
Channel publishChannel = fakeChannel(seqCounter, seqOrder, msgIdOrder, ackCallback, nackCallback);
Channel consumeChannel = fakeChannel(seqCounter, seqOrder, msgIdOrder, ackCallback, nackCallback);
Connection connection = fakeConnection(consumeChannel, publishChannel);
AmqpReplyInbox inbox = new AmqpReplyInbox(connection, AmqpReplyInbox.DEFAULT_PREFETCH);
// STALE_PUBLISHES in-flight publishes that never get confirmed — they sit in pendingBySeq /
// pendingByMsgId exactly like publishes whose confirm never arrived before a connection drop.
// Virtual threads make this many concurrent blocking publish() calls cheap.
CountDownLatch staleStarted = new CountDownLatch(STALE_PUBLISHES);
// Counted down by the FIRST stale publish thread to observe its own failure. That can only
// happen from inside failPendingPublishesOnRecovery() — nothing else in this test ever
// completes a stale Pending exceptionally (no nack/return is simulated for any "stale-*"
// msgId) — so seeing it fire is direct, observable proof the sweep is inside its loop, not a
// timing guess. It replaces the old fixed Thread.sleep(5) head start.
CountDownLatch sweepIsHoldingTheLock = new CountDownLatch(1);
for (int i = 0; i < STALE_PUBLISHES; i++) {
String msgId = "stale-" + i;
Thread.ofVirtual().start(() -> {
staleStarted.countDown();
try {
inbox.publish("worker-stale", msgId, "x");
} catch (IllegalStateException expected) {
sweepIsHoldingTheLock.countDown();
}
});
}
staleStarted.await();
// Let the registrations (the synchronized put into pendingBySeq/pendingByMsgId) actually land
// for all of them before the sweep starts, so the sweep begins with a large, real backlog.
Thread.sleep(300);
AtomicReference<Throwable> sweepError = new AtomicReference<>();
Thread sweepThread = new Thread(() -> {
try {
inbox.failPendingPublishesOnRecovery();
} catch (Throwable t) {
sweepError.set(t);
}
}, "recovery-sweep");
sweepThread.start();
// Deterministic head start: block until the sweep has actually failed one of the stale
// publishes. failPendingPublishesOnRecovery() (once guarded, as it is on main) holds
// publishChannelLock for its ENTIRE loop, not just per entry — so this failure proves the
// sweep is, at this instant, still holding that lock. With STALE_PUBLISHES this large, the
// remaining ~1,999 entries give an enormous margin between "first failure observed" and "sweep
// releases the lock": there is no window left for "fresh" to slip in before the sweep starts,
// or to win the lock ahead of it — see the case-2 note below. This also means Case 1 (the sweep
// is already inside its loop, holding the lock, when "fresh" tries to register) is now
// guaranteed by construction rather than merely likely under a fixed sleep.
assertTrue(sweepIsHoldingTheLock.await(20, TimeUnit.SECONDS),
"the sweep never failed a single stale publish — it may not have started");
// Case 2 ("fresh" wins publishChannelLock before the sweep even starts, so it genuinely
// published on the stale channel and the sweep correctly fails it) is impossible by
// construction in this test: freshThread.start() below is reached only after
// sweepIsHoldingTheLock has counted down, which can only happen once
// failPendingPublishesOnRecovery() is already running and has already failed a stale entry.
// There is no code path that lets "fresh" start before the sweep starts. That case is real
// and correct production behaviour (see AmqpReplyInbox#failPendingPublishesOnRecovery's
// javadoc), it is just not reachable from this deterministic ordering, so it does not need a
// separate assertion here.
// This is the exact interleaving CB-528's follow-up describes: "the still-running recovery
// sweep" racing a publish that registers while it is mid-flight.
AtomicReference<Throwable> publishError = new AtomicReference<>();
Thread freshThread = new Thread(() -> {
try {
inbox.publish("worker-fresh", "fresh", "hello");
} catch (Throwable t) {
publishError.set(t);
}
}, "fresh-publish");
freshThread.start();
sweepThread.join(20_000);
assertNull(sweepError.get(), "sweep threw: " + sweepError.get());
// Simulate the broker's real confirm for "fresh" now that the sweep is done, so a correct
// implementation's publish() returns normally instead of idling out CONFIRM_TIMEOUT_MS. Poll
// for the registration rather than checking once: freshThread may still be contending for
// publishChannelLock (behind the sweep's own hold on it, and possibly other stale threads
// still unwinding) even though the sweep itself has already finished.
long deadline = System.nanoTime() + TimeUnit.SECONDS.toNanos(9);
int idx = -1;
while (idx < 0 && System.nanoTime() < deadline) {
idx = msgIdOrder.indexOf("fresh");
if (idx < 0) {
Thread.sleep(20);
}
}
if (idx >= 0 && ackCallback.get() != null) {
ackCallback.get().handle(seqOrder.get(idx), false);
}
freshThread.join(15_000);
assertNull(publishError.get(),
"a publish that registered while the recovery sweep was running must not be failed by "
+ "it, but got: " + publishError.get());
}
@Test
@Timeout(15)
void closeFailsInFlightPublishPromptlyInsteadOfWaitingOutTheConfirmTimeout() throws Exception {
AtomicLong seqCounter = new AtomicLong();
List<Long> seqOrder = new CopyOnWriteArrayList<>();
List<String> msgIdOrder = new CopyOnWriteArrayList<>();
AtomicReference<ConfirmCallback> ackCallback = new AtomicReference<>();
AtomicReference<ConfirmCallback> nackCallback = new AtomicReference<>();
Channel publishChannel = fakeChannel(seqCounter, seqOrder, msgIdOrder, ackCallback, nackCallback);
Channel consumeChannel = fakeChannel(seqCounter, seqOrder, msgIdOrder, ackCallback, nackCallback);
Connection connection = fakeConnection(consumeChannel, publishChannel);
AmqpReplyInbox inbox = new AmqpReplyInbox(connection, AmqpReplyInbox.DEFAULT_PREFETCH);
CountDownLatch publishReturned = new CountDownLatch(1);
AtomicReference<Throwable> publishError = new AtomicReference<>();
AtomicLong elapsedMillis = new AtomicLong();
Thread publishThread = new Thread(() -> {
long start = System.nanoTime();
try {
inbox.publish("worker-close", "never-confirmed", "x");
} catch (Throwable t) {
publishError.set(t);
} finally {
elapsedMillis.set((System.nanoTime() - start) / 1_000_000);
publishReturned.countDown();
}
}, "publish-during-close");
publishThread.start();
Thread.sleep(200); // let publish() register before close() runs
inbox.close();
assertTrue(publishReturned.await(5, TimeUnit.SECONDS), "publish() did not return after close()");
assertNotNull(publishError.get(), "a publish in flight when close() runs must fail, not hang");
assertTrue(publishError.get().getMessage() != null
&& publishError.get().getMessage().toLowerCase().contains("closed"),
"expected a clear closed-inbox message, got: " + publishError.get());
assertTrue(elapsedMillis.get() < 5_000,
"close() should fail the in-flight publish promptly, not wait out the confirm timeout — took "
+ elapsedMillis.get() + "ms");
}
/** A {@link Proxy}-backed {@link Channel}: only the calls {@link AmqpReplyInbox} actually makes
* are meaningfully implemented; everything else returns a harmless default. */
private static Channel fakeChannel(AtomicLong seqCounter, List<Long> seqOrder, List<String> msgIdOrder,
AtomicReference<ConfirmCallback> ackCallback,
AtomicReference<ConfirmCallback> nackCallback) {
InvocationHandler handler = (proxy, method, args) -> {
String name = method.getName();
if (name.equals("getNextPublishSeqNo")) {
long value = seqCounter.incrementAndGet();
seqOrder.add(value);
return value;
}
if (name.equals("basicPublish")) {
AMQP.BasicProperties props = (AMQP.BasicProperties) args[3];
msgIdOrder.add(props.getMessageId());
return null;
}
if (name.equals("addConfirmListener")) {
ackCallback.set((ConfirmCallback) args[0]);
nackCallback.set((ConfirmCallback) args[1]);
return null;
}
if (name.equals("equals")) {
return proxy == args[0];
}
if (name.equals("hashCode")) {
return System.identityHashCode(proxy);
}
if (name.equals("toString")) {
return "FakeChannel";
}
return defaultValue(method.getReturnType());
};
return (Channel) Proxy.newProxyInstance(AmqpReplyInboxRecoveryRaceTest.class.getClassLoader(),
new Class<?>[] {Channel.class}, handler);
}
/** A {@link Proxy}-backed {@link Connection} handing out {@code first} then {@code second} from
* successive {@code createChannel()} calls, matching {@link AmqpReplyInbox}'s constructor. */
private static Connection fakeConnection(Channel first, Channel second) {
AtomicInteger calls = new AtomicInteger();
InvocationHandler handler = (proxy, method, args) -> {
String name = method.getName();
if (name.equals("createChannel") && (args == null || args.length == 0)) {
return calls.getAndIncrement() == 0 ? first : second;
}
if (name.equals("equals")) {
return proxy == args[0];
}
if (name.equals("hashCode")) {
return System.identityHashCode(proxy);
}
if (name.equals("toString")) {
return "FakeConnection";
}
return defaultValue(method.getReturnType());
};
return (Connection) Proxy.newProxyInstance(AmqpReplyInboxRecoveryRaceTest.class.getClassLoader(),
new Class<?>[] {Connection.class}, handler);
}
private static Object defaultValue(Class<?> type) {
if (!type.isPrimitive() || type == void.class) {
return null;
}
if (type == boolean.class) {
return Boolean.FALSE;
}
if (type == long.class) {
return 0L;
}
if (type == short.class) {
return (short) 0;
}
if (type == byte.class) {
return (byte) 0;
}
if (type == char.class) {
return (char) 0;
}
if (type == double.class) {
return 0.0d;
}
if (type == float.class) {
return 0.0f;
}
return 0;
}
}
@@ -5,6 +5,8 @@ import dev.ltms.bridged.herdr.AgentStatus;
import dev.ltms.bridged.herdr.FakeHerdr;
import dev.ltms.bridged.herdr.HerdrException;
import dev.ltms.bridged.inject.CompletionResolver;
import dev.ltms.bridged.inject.ExhaustedPatternLookup;
import dev.ltms.bridged.inject.ExhaustionSink;
import dev.ltms.bridged.mcp.PrimaryRegistry;
import dev.ltms.bridged.inject.Injector;
import org.junit.jupiter.api.BeforeEach;
@@ -31,7 +33,8 @@ class MessageServiceTest {
private final FakeHerdr herdr = new FakeHerdr().readText("BUILD GREEN: 391 files");
private final AgentControl agents = new AgentControl(herdr);
private final Rendezvous rendezvous = new Rendezvous();
private final CompletionResolver completion = new CompletionResolver(agents, rendezvous);
private final CompletionResolver completion =
new CompletionResolver(agents, rendezvous, ExhaustedPatternLookup.none(), ExhaustionSink.none());
private final Injector injector = new Injector(agents, completion);
private final InMemoryReplyInbox inbox = new InMemoryReplyInbox();
private final MessageService messages = new MessageService(agents, injector, rendezvous, inbox);
@@ -120,6 +123,36 @@ class MessageServiceTest {
assertFalse(reply.completed());
}
@Test
void droppedQueuedAndDeliveredTurnsExposeTheRealCauseExactlyOnce() throws Exception {
CompletableFuture<MessageService.Reply> first = sendAsync();
awaitWaiting();
injector.onStatus(T, AgentStatus.IDLE); // first delivery
injector.onStatus(T, AgentStatus.WORKING); // first turn in flight
CompletableFuture<Void> queued = injector.enqueue(T, "second task", TestTurnTokens.inert(T));
CompletableFuture<Rendezvous.Resolution> waiter = rendezvous.currentWaiter(T);
injector.drop(T, new HerdrException("agent target sol not found", "agent_not_found", null));
MessageService.Reply reply = first.get(5, TimeUnit.SECONDS);
assertEquals(MessageService.Outcome.WORKER_FAILED, reply.outcome());
assertEquals("agent target sol not found", reply.text());
assertTrue(queued.isCompletedExceptionally(), "the queued delivery future also fails");
assertFalse(rendezvous.resolveFailure(waiter, "second failure"), "the waiter fails exactly once");
}
@Test
void noDropReasonKeepsTheExistingFallbackText() throws Exception {
herdr.readText("");
CompletableFuture<Rendezvous.Resolution> waiter = rendezvous.open(T);
completion.onTurnFailed(T);
Rendezvous.Resolution resolution = waiter.get(2, TimeUnit.SECONDS);
assertEquals("worker did not reply; its turn ended in an unrecoverable state "
+ "(worker unreachable or stuck)", resolution.text());
}
// --- bridge_ask reverse rendezvous (CB-205) ------------------------------------------------
@Test
@@ -609,4 +642,443 @@ class MessageServiceTest {
assertTrue(view.detail() != null && view.detail().contains("released"),
"and the detail says why, rather than 'worker unknown'");
}
@Test
void abandonFailsEveryPendingAsyncTicketForTheReleasedTarget() throws Exception {
String first = messages.sendAsync(T, "first task");
awaitWaiting(); // first task owns the target lock and rendezvous waiter
String second = messages.sendAsync(T, "second task"); // parked on the same lock, not yet queued
String third = messages.sendAsync(T, "third task"); // a second queued ticket proves the full sweep
assertTrue(messages.abandon(T, "agent target term_a not found"));
assertFailedTicket(first, "agent target term_a not found");
assertFailedTicket(second, "agent target term_a not found");
assertFailedTicket(third, "agent target term_a not found");
}
@Test
void abandonDoesNotFailAnAsyncTicketWaitingForAnAnswer() throws Exception {
String ticket = messages.sendAsync(T, "task that asks");
awaitWaiting();
injectDelivery();
CompletableFuture<MessageService.AskResult> ask =
CompletableFuture.supplyAsync(() -> messages.ask(T, "which config?", 5000));
MessageService.TaskView asking = awaitTicketPhase(ticket, MessageService.Phase.ASKING);
assertFalse(messages.abandon(T, "agent target term_a not found"),
"an asking ticket is an active turn, not a pending send to sweep");
assertEquals(MessageService.Phase.ASKING, messages.poll(ticket).phase());
CompletableFuture<MessageService.Reply> answer = CompletableFuture.supplyAsync(
() -> messages.answer(asking.turnId(), "config.yaml", 5000));
assertEquals("config.yaml", ask.get(5, TimeUnit.SECONDS).answer());
awaitWaiting();
assertTrue(rendezvous.resolve(T, "done"));
assertEquals(MessageService.Outcome.REPLIED, answer.get(5, TimeUnit.SECONDS).outcome());
}
@Test
void unansweredAsyncQuestionReturnsTheTicketToPendingAndReleasesItsTarget() throws Exception {
String ticket = messages.sendAsync(T, "task that asks");
awaitWaiting();
injectDelivery();
assertEquals(MessageService.AskOutcome.TIMED_OUT,
messages.ask(T, "which config?", 200).outcome());
assertEquals(MessageService.Phase.PENDING, messages.poll(ticket).phase(),
"only the question wait ended; the delegated turn may still finish");
String next = messages.sendAsync(T, "next task");
awaitWaiting();
assertTrue(rendezvous.resolve(T, "done"));
assertEquals(MessageService.Phase.DONE, awaitTicketPhase(next, MessageService.Phase.DONE).phase());
}
@Test
void asyncQuestionBelongsToTheTaskThatOwnsItsForwardWaiter() throws Exception {
String first = messages.sendAsync(T, "first task");
awaitWaiting();
CompletableFuture<MessageService.AskResult> ask =
CompletableFuture.supplyAsync(() -> messages.ask(T, "which config?", 5000));
assertEquals(MessageService.Phase.ASKING, awaitTicketPhase(first, MessageService.Phase.ASKING).phase());
MessageService.TaskView asking = messages.poll(first);
CompletableFuture<MessageService.Reply> answer = CompletableFuture.supplyAsync(
() -> messages.answer(asking.turnId(), "config.yaml", 5000));
assertEquals("config.yaml", ask.get(5, TimeUnit.SECONDS).answer());
awaitWaiting();
assertTrue(rendezvous.resolve(T, "done"));
assertEquals(MessageService.Outcome.REPLIED, answer.get(5, TimeUnit.SECONDS).outcome());
}
// --- CB-582: bridge_status pendingAsk() ------------------------------------------------------
@Test
void pendingAskReturnsNullWhenNoQuestionIsOpen() throws Exception {
assertNull(messages.pendingAsk(T), "no async ticket at all -> no pending ask");
String ticket = messages.sendAsync(T, "long task");
awaitWaiting();
assertNull(messages.pendingAsk(T), "a plain pending delegation is not a question");
injectDelivery();
assertTrue(rendezvous.resolve(T, "done"));
awaitTicketPhase(ticket, MessageService.Phase.DONE);
assertNull(messages.pendingAsk(T), "a finished ticket carries no open question either");
}
@Test
void pendingAskReturnsTheOpenQuestionForAnAsyncTicket() throws Exception {
String ticket = messages.sendAsync(T, "task that asks");
awaitWaiting();
injectDelivery();
CompletableFuture<MessageService.AskResult> ask =
CompletableFuture.supplyAsync(() -> messages.ask(T, "which config file?", 5000));
MessageService.TaskView asking = awaitTicketPhase(ticket, MessageService.Phase.ASKING);
MessageService.PendingAsk pending = messages.pendingAsk(T);
assertNotNull(pending, "bridge_status should see the open question");
assertEquals(ticket, pending.ticket());
assertEquals("which config file?", pending.question());
assertEquals(asking.turnId(), pending.turnId());
CompletableFuture<MessageService.Reply> answer = CompletableFuture.supplyAsync(
() -> messages.answer(asking.turnId(), "config.yaml", 5000));
assertEquals("config.yaml", ask.get(5, TimeUnit.SECONDS).answer());
assertNull(messages.pendingAsk(T), "an answered question is no longer pending");
awaitWaiting();
assertTrue(rendezvous.resolve(T, "done"));
assertEquals(MessageService.Outcome.REPLIED, answer.get(5, TimeUnit.SECONDS).outcome());
}
// --- CB-588: async ticket terminal nudges ---------------------------------------------------
//
// MessageService.reply's rendezvous fast path is exactly what an async ticket always takes
// (sendAsync registers a rendezvous waiter — see asyncTasksByWaiter), so it never reached
// ReplyPushLoop.onReplyQueued. These prove the ticket reaches ReplyPushLoop through the new
// onTicketTerminal entry point instead, with no bridge_poll from the lead first.
private static final String LEAD = "term_lead";
/** A MessageService wired to a real ReplyPushLoop pointed at a private herdr fake for the lead. */
private record PushWiring(MessageService service, FakeHerdr leadHerdr,
java.util.concurrent.ScheduledExecutorService scheduler) implements AutoCloseable {
@Override
public void close() {
scheduler.shutdownNow();
}
}
private PushWiring wireWithPushLoop(int maxReminders, long backoffMs) {
return wireWithPushLoop(maxReminders, backoffMs, System::nanoTime);
}
/** As above, with an injectable clock (CB-588 follow-up: exercise pruneTerminalTickets' TTL). */
private PushWiring wireWithPushLoop(int maxReminders, long backoffMs, java.util.function.LongSupplier nowNanos) {
PrimaryRegistry registry = new PrimaryRegistry(null);
registry.recordDelegation(T, LEAD);
FakeHerdr leadHerdr = new FakeHerdr();
AgentControl leadAgents = new AgentControl(leadHerdr);
var scheduler = java.util.concurrent.Executors.newSingleThreadScheduledExecutor();
ReplyPushLoop pushLoop = new ReplyPushLoop(registry, leadAgents, inbox, scheduler, maxReminders, backoffMs);
MessageService service = new MessageService(agents, injector, rendezvous, inbox, pushLoop, null, nowNanos);
return new PushWiring(service, leadHerdr, scheduler);
}
private void awaitNudge(FakeHerdr leadHerdr) throws InterruptedException {
long deadline = System.currentTimeMillis() + 3000;
while (!leadHerdr.called("agent.prompt") && System.currentTimeMillis() < deadline) {
Thread.sleep(10);
}
assertTrue(leadHerdr.called("agent.prompt"), "expected a nudge in the lead's pane");
}
@Test
void anAsyncTicketThatFinishesNudgesTheLeadWithNoPriorPollCall() throws Exception {
try (var wiring = wireWithPushLoop(1, 50)) {
String ticket = wiring.service().sendAsync(T, "long task");
awaitWaiting();
injectDelivery();
assertTrue(rendezvous.resolve(T, "async result"));
awaitNudge(wiring.leadHerdr());
String nudge = wiring.leadHerdr().lastCall("agent.prompt").params().toString();
assertTrue(nudge.contains(ticket), "the nudge should name the ticket: " + nudge);
assertTrue(nudge.contains("bridge_poll(ticket="),
"the nudge should name the exact ticket-collecting call: " + nudge);
assertFalse(nudge.toUpperCase().contains("FAILED"),
"a successfully-replied ticket's nudge must not say it failed: " + nudge);
}
}
@Test
void aFailedAsyncTicketAlsoNudgesAndSaysSo() throws Exception {
try (var wiring = wireWithPushLoop(1, 50)) {
wiring.service().sendAsync(T, "long task");
awaitWaiting();
wiring.service().abandon(T, "session released"); // a terminal failure phase
awaitNudge(wiring.leadHerdr());
String nudge = wiring.leadHerdr().lastCall("agent.prompt").params().toString();
assertTrue(nudge.toUpperCase().contains("FAILED"), "a failed ticket's nudge must say so: " + nudge);
}
}
@Test
void anAlreadyCollectedTicketProducesNoNudge() throws Exception {
try (var wiring = wireWithPushLoop(1, 300)) { // wide backoff: poll before the first tick fires
String ticket = wiring.service().sendAsync(T, "long task");
awaitWaiting();
injectDelivery();
assertTrue(rendezvous.resolve(T, "async result"));
MessageService.TaskView view = awaitTicketPhaseOn(wiring.service(), ticket, MessageService.Phase.DONE);
assertEquals(MessageService.Phase.DONE, view.phase());
Thread.sleep(400); // let the scheduled tick run — it must find nothing pending
assertFalse(wiring.leadHerdr().called("agent.prompt"),
"a ticket the lead already polled must never be nudged");
}
}
@Test
void severalAsyncTicketsFinishingTogetherProduceOneCoalescedNudge() throws Exception {
try (var wiring = wireWithPushLoop(1, 300)) { // wide backoff: both tickets land before the tick fires
String first = wiring.service().sendAsync(T, "first task");
awaitWaiting();
injectDelivery();
assertTrue(rendezvous.resolve(T, "first done"));
// Settle without polling: poll() itself marks a ticket collected (that's the point of
// anAlreadyCollectedTicketProducesNoNudge above) — using it here to detect completion
// would collect the ticket before the coalescing this test checks ever gets a chance.
Thread.sleep(100);
String second = wiring.service().sendAsync(T, "second task");
awaitWaiting();
injectDelivery();
assertTrue(rendezvous.resolve(T, "second done"));
Thread.sleep(100);
awaitNudge(wiring.leadHerdr());
Thread.sleep(200); // settle — nothing more should arrive beyond the one coalesced nudge
long nudgeCount = wiring.leadHerdr().calls.stream()
.filter(c -> c.method().equals("agent.prompt")).count();
assertEquals(1, nudgeCount, "two tickets finishing together must produce ONE nudge, not two");
String nudge = wiring.leadHerdr().lastCall("agent.prompt").params().toString();
assertTrue(nudge.contains(first) && nudge.contains(second),
"the coalesced nudge should name both tickets: " + nudge);
}
}
// --- CB-582: bridge_ask question-open nudges --------------------------------------------------
@Test
void anAsyncTicketThatPausesOnAQuestionNudgesTheLeadWithNoPriorPollCall() throws Exception {
try (var wiring = wireWithPushLoop(5, 50)) {
String ticket = wiring.service().sendAsync(T, "task that asks");
awaitWaiting();
injectDelivery();
CompletableFuture<MessageService.AskResult> ask = CompletableFuture.supplyAsync(
() -> wiring.service().ask(T, "which config file?", 5000));
MessageService.TaskView asking = awaitTicketPhaseOn(wiring.service(), ticket, MessageService.Phase.ASKING);
awaitNudge(wiring.leadHerdr());
String nudge = wiring.leadHerdr().lastCall("agent.prompt").params().toString();
assertTrue(nudge.contains(ticket), "the nudge should name the ticket: " + nudge);
assertTrue(nudge.contains(asking.turnId()), "the nudge should name the turnId: " + nudge);
assertTrue(nudge.contains("bridge_send(turnId="),
"the nudge should name the exact answer call: " + nudge);
assertTrue(nudge.contains("which config file?"), "the nudge should include the question: " + nudge);
// Clean up the still-open ask so the background thread does not linger past the test.
CompletableFuture<MessageService.Reply> answer = CompletableFuture.supplyAsync(
() -> wiring.service().answer(asking.turnId(), "config.yaml", 5000));
assertEquals("config.yaml", ask.get(5, TimeUnit.SECONDS).answer());
awaitWaiting();
assertTrue(rendezvous.resolve(T, "done"));
answer.get(5, TimeUnit.SECONDS);
}
}
@Test
void answeringAQuestionStopsFurtherNudgesAboutIt() throws Exception {
try (var wiring = wireWithPushLoop(5, 50)) {
String ticket = wiring.service().sendAsync(T, "task that asks");
awaitWaiting();
injectDelivery();
CompletableFuture<MessageService.AskResult> ask = CompletableFuture.supplyAsync(
() -> wiring.service().ask(T, "which config file?", 5000));
MessageService.TaskView asking = awaitTicketPhaseOn(wiring.service(), ticket, MessageService.Phase.ASKING);
awaitNudge(wiring.leadHerdr());
long callsBeforeAnswer = wiring.leadHerdr().calls.stream()
.filter(c -> c.method().equals("agent.prompt")).count();
CompletableFuture<MessageService.Reply> answer = CompletableFuture.supplyAsync(
() -> wiring.service().answer(asking.turnId(), "config.yaml", 5000));
assertEquals("config.yaml", ask.get(5, TimeUnit.SECONDS).answer());
awaitWaiting();
assertTrue(rendezvous.resolve(T, "done"));
answer.get(5, TimeUnit.SECONDS);
// Let several more ticks (and the ticket's own now-legitimate terminal nudge) fire —
// none of them may still name the question's turnId, which is closed.
Thread.sleep(300);
boolean anyNamesClosedQuestion = wiring.leadHerdr().calls.stream()
.filter(c -> c.method().equals("agent.prompt"))
.skip(callsBeforeAnswer)
.anyMatch(c -> c.params().toString().contains(asking.turnId()));
assertFalse(anyNamesClosedQuestion,
"no nudge sent after the answer may still name the now-closed turnId " + asking.turnId());
}
}
@Test
void anAskThatLeavesByThrowingStillClosesItsQuestion() throws Exception {
// CB-582 follow-up. ask() calls clearAsyncQuestion on three paths — no-waiter, timed out,
// and (from answer()) answered — but it can also leave by *throwing*: an interrupt while
// blocked on the answer, or an ExecutionException from the answer future. Those paths run
// only the finally block, so before the fix the push loop kept the question pending for
// good: named in every nudge until its own cap, then never removed from the map at all.
PrimaryRegistry registry = new PrimaryRegistry(null);
registry.recordDelegation(T, LEAD);
var scheduler = java.util.concurrent.Executors.newSingleThreadScheduledExecutor();
// A backoff far longer than the test: the schedule is started but no tick ever fires, so
// decide() is read directly and nothing here depends on timing.
ReplyPushLoop pushLoop = new ReplyPushLoop(registry, new AgentControl(new FakeHerdr()), inbox,
scheduler, 5, 60_000);
MessageService service = new MessageService(agents, injector, rendezvous, inbox, pushLoop);
try {
String ticket = service.sendAsync(T, "task that asks");
awaitWaiting();
injectDelivery();
Thread asker = new Thread(() -> assertThrows(IllegalStateException.class,
() -> service.ask(T, "which config file?", 30_000)));
asker.start();
awaitTicketPhaseOn(service, ticket, MessageService.Phase.ASKING);
assertEquals(ReplyPushLoop.Action.INJECT, pushLoop.decide(LEAD, 0, 0, 0),
"the open question should be the one thing keeping this lead's schedule alive");
asker.interrupt();
asker.join(5000);
assertFalse(asker.isAlive(), "the interrupted ask should have left ask() by throwing");
assertEquals(ReplyPushLoop.Action.STOP, pushLoop.decide(LEAD, 0, 0, 0),
"an ask that threw must still close its question, or the loop nudges about it for good");
} finally {
service.close();
scheduler.shutdownNow();
}
}
@Test
void aFleetWithNoPushLoopConfiguredBehavesExactlyAsToday() throws Exception {
// `messages` (the shared field) uses the no-pushLoop constructor — poll() must not throw,
// and no nudge mechanism exists to fire regardless of how the ticket resolves.
String ticket = messages.sendAsync(T, "long task");
awaitWaiting();
injectDelivery();
assertTrue(rendezvous.resolve(T, "async result"));
MessageService.TaskView view = awaitTicketPhase(ticket, MessageService.Phase.DONE);
assertEquals(MessageService.Phase.DONE, view.phase());
assertEquals("async result", view.reply());
}
/**
* CB-588 follow-up: {@code tasks} is the sole authority on whether a ticket exists, and
* {@code pruneTerminalTickets} drops entries from it once {@link MessageService#TICKET_TTL_NANOS}
* elapses. Before this test, that prune never told {@code ReplyPushLoop} — its own
* {@code pendingTickets} entry for a pruned, never-collected ticket had no remover at all, so it
* rode along on every later nudge to the same lead, naming a ticket {@code bridge_poll} could no
* longer find. Uses the injectable clock (mirroring {@code SessionManager}'s {@code nowNanos} seam
* for its idle reaper) to cross the 10-minute TTL without a real wait.
*/
@Test
void aPrunedTicketIsReclaimedFromThePushLoopNotLeakedForever() throws Exception {
java.util.concurrent.atomic.AtomicLong clock = new java.util.concurrent.atomic.AtomicLong(1_000_000_000L);
// maxReminders=1 + a short backoff: the stale ticket gets its one legitimate reminder, then
// decideTickets hits the cap and STOPs — activeLeads drops the lead, but (before the fix)
// pendingTickets never drops the ticket. That is the exact "cap already STOPped" branch of
// the bug report, reached deterministically rather than by timing it against a live tick.
try (var wiring = wireWithPushLoop(1, 50, clock::get)) {
String stale = wiring.service().sendAsync(T, "first task");
awaitWaiting();
injectDelivery();
assertTrue(rendezvous.resolve(T, "stale result"));
// Let the reminder loop fire its one nudge and hit the cap (STOP removes it from
// activeLeads; pendingTickets is untouched either way — that asymmetry is the bug).
awaitNudge(wiring.leadHerdr());
Thread.sleep(300);
assertTrue(wiring.leadHerdr().lastCall("agent.prompt").params().toString().contains(stale),
"sanity: the stale ticket's own reminder must have fired first");
// Cross the TTL — a real clock would need 10 minutes; the injected one does it instantly.
clock.addAndGet(MessageService.TICKET_TTL_NANOS + TimeUnit.SECONDS.toNanos(1));
// A second, unrelated ticket to the same target/lead reaches sendAsync, which prunes.
String fresh = wiring.service().sendAsync(T, "second task");
awaitWaiting();
injectDelivery();
assertTrue(rendezvous.resolve(T, "fresh result"));
// The fresh ticket restarts the (now-dormant) reminder loop with its own nudge.
long before = wiring.leadHerdr().calls.stream().filter(c -> c.method().equals("agent.prompt")).count();
long deadline = System.currentTimeMillis() + 3000;
while (wiring.leadHerdr().calls.stream().filter(c -> c.method().equals("agent.prompt")).count() <= before
&& System.currentTimeMillis() < deadline) {
Thread.sleep(10);
}
String latestNudge = wiring.leadHerdr().lastCall("agent.prompt").params().toString();
assertTrue(latestNudge.contains(fresh), "the fresh ticket's nudge must still arrive: " + latestNudge);
assertFalse(latestNudge.contains(stale),
"a pruned ticket must never be named in a later nudge — it is gone and bridge_poll "
+ "on it would return nothing: " + latestNudge);
// And bridge_poll(ticket=stale) really does return nothing now — the nudge would have lied.
assertNull(wiring.service().poll(stale), "the pruned ticket must actually be gone, not just unmentioned");
}
}
private MessageService.TaskView awaitTicketPhaseOn(MessageService svc, String ticket,
MessageService.Phase phase) throws Exception {
long deadline = System.currentTimeMillis() + 3000;
MessageService.TaskView view;
do {
view = svc.poll(ticket);
if (view.phase() == phase) {
return view;
}
Thread.sleep(5);
} while (System.currentTimeMillis() < deadline);
assertEquals(phase, view.phase());
return view;
}
private void assertFailedTicket(String ticket, String reason) throws Exception {
MessageService.TaskView view = awaitTicketPhase(ticket, MessageService.Phase.FAILED);
assertEquals(reason, view.detail());
}
private MessageService.TaskView awaitTicketPhase(String ticket, MessageService.Phase phase) throws Exception {
long deadline = System.currentTimeMillis() + 3000;
MessageService.TaskView view;
do {
view = messages.poll(ticket);
if (view.phase() == phase) {
return view;
}
Thread.sleep(5);
} while (System.currentTimeMillis() < deadline);
assertEquals(phase, view.phase());
return view;
}
}
@@ -114,6 +114,17 @@ class RendezvousTest {
"the first resolution wins; the stored value is unchanged");
}
@Test
void resolveExhaustedTwiceIsANoOpTheSecondTime() {
CompletableFuture<Rendezvous.Resolution> waiter = rendezvous.open(W);
assertTrue(rendezvous.resolveExhausted(waiter, "first reason"), "the first classification resolves");
assertFalse(rendezvous.resolveExhausted(waiter, "second reason"),
"a second exhausted resolution on an already-resolved waiter returns false");
assertEquals(Rendezvous.Kind.BACKEND_EXHAUSTED, waiter.getNow(null).kind());
assertEquals("first reason", waiter.getNow(null).text(),
"the first resolution wins; the stored value is unchanged");
}
@Test
void closeAskRemovesTheTurn() {
Rendezvous.AskTicket t = rendezvous.openAsk(W);
@@ -15,10 +15,12 @@ import java.util.ArrayList;
import java.util.Collections;
import java.util.List;
import java.util.Map;
import java.util.Set;
import java.util.concurrent.CountDownLatch;
import java.util.concurrent.Executors;
import java.util.concurrent.ScheduledExecutorService;
import java.util.concurrent.TimeUnit;
import java.util.concurrent.atomic.AtomicInteger;
import static org.junit.jupiter.api.Assertions.*;
@@ -26,6 +28,12 @@ import static org.junit.jupiter.api.Assertions.*;
* Unit tests for {@link ReplyPushLoop}: decision logic, nudge injection, idempotency,
* bounded reminders, and stop conditions.
*
* <p>CB-590 collapsed the CB-307 reply-nudge schedule and the CB-588 ticket-nudge schedule into
* one schedule per lead ({@link ReplyPushLoop#decide}), so most tests below register pending work
* through the public entry points ({@code onReplyQueued} / {@code onTicketTerminal}) before
* exercising {@code decide} directly, mirroring how the two entry points now share one decision
* function keyed by the lead terminal rather than by worker target.
*
* <p>Uses a {@link RecordingHerdrClient} that synchronizes access to its call list so the
* scheduler thread and test thread never have memory ordering issues. The {@code decide()}
* tests use a simple client with no concurrency concern.
@@ -34,6 +42,7 @@ class ReplyPushLoopTest {
private static final String PRIMARY = "term_primary";
private static final String WORKER = "term_worker";
private static final String WORKER2 = "term_worker2";
private static final ObjectMapper MAPPER = new ObjectMapper();
private PrimaryRegistry registry;
@@ -57,38 +66,50 @@ class ReplyPushLoopTest {
// --- decide() logic ------------------------------------------------------------------------
@Test
void decideWithoutPrimaryIsStop() {
agents = agentWithStatus("idle");
var loop = new ReplyPushLoop(
new PrimaryRegistry(null), agents, inbox, scheduler, 5, 100);
assertEquals(ReplyPushLoop.Action.STOP, loop.decide(WORKER, 0));
void onReplyQueuedWithNoKnownLeadNeverStartsASchedule() throws Exception {
var rec = recordingClient();
agents = new AgentControl(rec);
inbox.publish(WORKER, "m1", "hello");
var loop = new ReplyPushLoop(new PrimaryRegistry(null), agents, inbox, scheduler, 5, 50);
loop.onReplyQueued(WORKER); // no lead known -> never registered, never scheduled
assertFalse(loop.isActive(), "no lead known means nothing to nudge yet");
Thread.sleep(150);
assertEquals(0, rec.sendCount(), "must not nudge when no lead is known to be waiting");
}
@Test
void decideWithEmptyInboxIsStop() {
void decideWithNothingPendingIsStop() {
agents = agentWithStatus("idle");
assertEquals(ReplyPushLoop.Action.STOP, loop().decide(WORKER, 0));
assertEquals(ReplyPushLoop.Action.STOP, loop().decide(PRIMARY, 0, 0));
}
@Test
void decideAtCapIsStop() {
agents = agentWithStatus("idle");
inbox.publish(WORKER, "m1", "hello");
assertEquals(ReplyPushLoop.Action.STOP, loop(2, 100).decide(WORKER, 2));
var loop = loop(2, 100);
loop.onReplyQueued(WORKER);
assertEquals(ReplyPushLoop.Action.STOP, loop.decide(PRIMARY, 2, 0));
}
@Test
void decideUnderCapWithInjectablePrimaryIsInject() {
agents = agentWithStatus("idle");
inbox.publish(WORKER, "m1", "hello");
assertEquals(ReplyPushLoop.Action.INJECT, loop().decide(WORKER, 0));
var loop = loop();
loop.onReplyQueued(WORKER);
assertEquals(ReplyPushLoop.Action.INJECT, loop.decide(PRIMARY, 0, 0));
}
@Test
void decideUnderCapWithBlockedPrimaryIsInject() {
agents = agentWithStatus("blocked");
inbox.publish(WORKER, "m1", "hello");
assertEquals(ReplyPushLoop.Action.INJECT, loop().decide(WORKER, 0),
var loop = loop();
loop.onReplyQueued(WORKER);
assertEquals(ReplyPushLoop.Action.INJECT, loop.decide(PRIMARY, 0, 0),
"BLOCKED is injectable");
}
@@ -96,7 +117,9 @@ class ReplyPushLoopTest {
void decideUnderCapWithDonePrimaryIsInject() {
agents = agentWithStatus("done");
inbox.publish(WORKER, "m1", "hello");
assertEquals(ReplyPushLoop.Action.INJECT, loop().decide(WORKER, 0),
var loop = loop();
loop.onReplyQueued(WORKER);
assertEquals(ReplyPushLoop.Action.INJECT, loop.decide(PRIMARY, 0, 0),
"DONE is injectable");
}
@@ -104,23 +127,29 @@ class ReplyPushLoopTest {
void decideUnderCapWithBusyPrimaryIsWaitBusy() {
agents = agentWithStatus("working");
inbox.publish(WORKER, "m1", "hello");
assertEquals(ReplyPushLoop.Action.WAIT_BUSY, loop().decide(WORKER, 0));
var loop = loop();
loop.onReplyQueued(WORKER);
assertEquals(ReplyPushLoop.Action.WAIT_BUSY, loop.decide(PRIMARY, 0, 0));
}
@Test
void decideUnderCapWithUnknownPrimaryIsWaitBusy() {
agents = agentWithStatus("unknown");
inbox.publish(WORKER, "m1", "hello");
assertEquals(ReplyPushLoop.Action.WAIT_BUSY, loop().decide(WORKER, 0));
var loop = loop();
loop.onReplyQueued(WORKER);
assertEquals(ReplyPushLoop.Action.WAIT_BUSY, loop.decide(PRIMARY, 0, 0));
}
@Test
void decideStopsAfterInboxIsEmptied() {
agents = agentWithStatus("idle");
inbox.publish(WORKER, "m1", "hello");
assertEquals(ReplyPushLoop.Action.INJECT, loop().decide(WORKER, 0));
var loop = loop();
loop.onReplyQueued(WORKER);
assertEquals(ReplyPushLoop.Action.INJECT, loop.decide(PRIMARY, 0, 0));
inbox.ack(WORKER, "m1");
assertEquals(ReplyPushLoop.Action.STOP, loop().decide(WORKER, 0));
assertEquals(ReplyPushLoop.Action.STOP, loop.decide(PRIMARY, 0, 0));
}
// --- onReplyQueued integration -------------------------------------------------------------
@@ -200,6 +229,560 @@ class ReplyPushLoopTest {
assertTrue(nudge.contains("bridge_poll(target=term_worker)"));
}
@Test
void repliesNudgeFormatIsCorrect() {
String multi = ReplyPushLoop.REPLIES_NUDGE_FORMAT.formatted(2, "term_worker1, term_worker2");
assertTrue(multi.contains("2 workers"));
assertTrue(multi.contains("bridge_poll(target=...)"));
}
// --- CB-588: async ticket terminal nudges — decide() logic on tickets -----------------------
@Test
void decideTicketsWithNothingPendingIsStop() {
agents = agentWithStatus("idle");
assertEquals(ReplyPushLoop.Action.STOP, loop().decide(PRIMARY, 0, 0));
}
@Test
void decideTicketsAtCapIsStop() {
agents = agentWithStatus("idle");
var loop = loop(2, 100_000);
loop.onTicketTerminal("task-1", WORKER, false);
assertEquals(ReplyPushLoop.Action.STOP, loop.decide(PRIMARY, 0, 2));
}
@Test
void decideTicketsUnderCapWithInjectableLeadIsInject() {
agents = agentWithStatus("idle");
var loop = loop(5, 100_000);
loop.onTicketTerminal("task-1", WORKER, false);
assertEquals(ReplyPushLoop.Action.INJECT, loop.decide(PRIMARY, 0, 0));
}
@Test
void decideTicketsUnderCapWithBusyLeadIsWaitBusy() {
agents = agentWithStatus("working");
var loop = loop(5, 100_000);
loop.onTicketTerminal("task-1", WORKER, false);
assertEquals(ReplyPushLoop.Action.WAIT_BUSY, loop.decide(PRIMARY, 0, 0));
}
@Test
void decideTicketsWithoutAKnownLeadIsStop() {
agents = agentWithStatus("idle");
var loop = new ReplyPushLoop(new PrimaryRegistry(null), agents, inbox, scheduler, 5, 100_000);
loop.onTicketTerminal("task-1", WORKER, false); // no lead known -> never registered as pending
assertEquals(ReplyPushLoop.Action.STOP, loop.decide(PRIMARY, 0, 0));
}
// --- CB-588: onTicketTerminal integration ---------------------------------------------------
@Test
void onTicketTerminalCausesExactlyOneNudgeNamingTheTicketAndThePollCall() throws Exception {
var rec = recordingClient();
agents = new AgentControl(rec);
loop(1, 50).onTicketTerminal("task-1", WORKER, false);
assertTrue(rec.sendLatch.await(3, TimeUnit.SECONDS), "one ticket nudge should have been sent");
assertEquals(1, rec.sendCount());
String nudge = rec.sentParams().getFirst().getValue().toString();
assertTrue(nudge.contains("task-1"), "nudge should name the ticket");
assertTrue(nudge.contains("bridge_poll(ticket="), "nudge should name the exact ticket-poll call");
assertFalse(nudge.contains("bridge_poll(target="), "a ticket-only nudge must not tell the lead to run the target-poll call");
}
@Test
void aFailedTicketNudgeSaysItWasAFailure() throws Exception {
var rec = recordingClient();
agents = new AgentControl(rec);
loop(1, 50).onTicketTerminal("task-1", WORKER, true);
assertTrue(rec.sendLatch.await(3, TimeUnit.SECONDS));
String nudge = rec.sentParams().getFirst().getValue().toString();
assertTrue(nudge.toUpperCase().contains("FAILED"), "a failed ticket's nudge must say so: " + nudge);
}
@Test
void severalTicketsFinishingTogetherProduceOneCoalescedNudge() throws Exception {
var rec = recordingClient();
agents = new AgentControl(rec);
var loop = loop(1, 300); // backoff wide enough that both onTicketTerminal calls land first
loop.onTicketTerminal("task-1", WORKER, false);
loop.onTicketTerminal("task-2", WORKER, true); // arrives while the schedule is already active
assertTrue(rec.sendLatch.await(3, TimeUnit.SECONDS), "one coalesced nudge should have been sent");
assertEquals(1, rec.sendCount(), "two tickets finishing together must produce ONE nudge, not two");
String nudge = rec.sentParams().getFirst().getValue().toString();
assertTrue(nudge.contains("task-1") && nudge.contains("task-2"),
"the coalesced nudge should name both tickets: " + nudge);
assertTrue(nudge.contains("2 tickets"), "the coalesced nudge should name the count: " + nudge);
}
@Test
void onTicketTerminalIsIdempotentPerLeadWhileActive() throws Exception {
var rec = recordingClient();
agents = new AgentControl(rec);
rec.sendLatch = new CountDownLatch(1);
var loop = loop(1, 100);
loop.onTicketTerminal("task-1", WORKER, false);
loop.onTicketTerminal("task-1", WORKER, false); // duplicate — should not start a second schedule
assertTrue(rec.sendLatch.await(3, TimeUnit.SECONDS));
Thread.sleep(200);
assertEquals(1, rec.sendCount(), "a duplicate onTicketTerminal for the same lead must not double-nudge");
}
@Test
void aTicketAlreadyCollectedProducesNoNudge() throws Exception {
var rec = recordingClient();
agents = new AgentControl(rec);
var loop = loop(1, 100);
loop.onTicketTerminal("task-1", WORKER, false);
loop.ticketCollected("task-1"); // the lead polled before the first tick fired
Thread.sleep(300); // let the scheduled tick run
assertEquals(0, rec.sendCount(), "an already-collected ticket must never be nudged");
}
@Test
void ticketNudgesSendUpToCapThenStop() throws Exception {
int cap = 2;
var rec = recordingClient();
agents = new AgentControl(rec);
rec.sendLatch = new CountDownLatch(cap);
loop(cap, 50).onTicketTerminal("task-1", WORKER, false);
assertTrue(rec.sendLatch.await(5, TimeUnit.SECONDS), cap + " ticket nudges should have fired");
Thread.sleep(300);
assertEquals(cap, rec.sendCount(), "exactly " + cap + " ticket nudges (cap=" + cap + ")");
}
@Test
void isActiveReflectsALiveTicketScheduleForTheHeartbeatStandDown() {
var rec = recordingClient();
agents = new AgentControl(rec);
ReplyPushLoop loop = loop(1, 100_000); // long backoff so the tick cannot fire mid-test
loop.onTicketTerminal("task-1", WORKER, false);
assertTrue(loop.isActive(), "an active ticket-reminder schedule must also stand off the heartbeat");
loop.stop();
assertFalse(loop.isActive(), "stopping clears the active ticket schedule too");
}
@Test
void aTicketStillPendingWhenTheLoopStopsIsNotStranded() {
// Regression for the race a reviewer found in gitea PR #73: onTicketTerminal's
// activeLeads.putIfAbsent can see the lead's slot as still occupied a moment before
// decide's STOP releases it, so the ticket coalesces onto a schedule that is about to
// die and nothing ever nudges about it. Forcing that exact thread interleaving is not
// reliable, so this drives stopOrRestart — the STOP path's own release-and-recheck —
// directly, arranging the state it must not lose a ticket in: a ticket pending for the lead
// that was NOT part of the pre-decision snapshot (ticketsBefore=empty), standing in for one
// that races in during the decision-to-release window.
var rec = recordingClient();
agents = new AgentControl(rec);
ReplyPushLoop loop = loop(1, 100_000); // long backoff — no natural tick fires during this test
loop.onTicketTerminal("task-1", WORKER, false); // pendingTickets={task-1}; activeLeads={PRIMARY}
// Stand in for the scheduler thread reaching decide==STOP for this lead — with nothing
// pending at decide time — while task-1 races in before the release below runs.
loop.stopOrRestart(PRIMARY, Set.of(), Set.of());
assertTrue(loop.isActive(), "a ticket that raced the loop's stop must reclaim the schedule "
+ "slot, not be stranded with no schedule left to ever nudge about it");
}
@Test
void aStaleUncollectedTicketAtCapDoesNotRestartTheLoop() {
// The other direction of the same fix: restarting on ANY non-empty pending set would be
// wrong. When STOP is reached because the reminder cap was hit, the same never-collected
// ticket is expected to still be there — that is the cap doing its job (acceptance criterion
// #5: no spin / nudges stay bounded). task-1 here was already accounted for at decide time
// (it is in ticketsBefore), so it must not restart the loop just because it is still sitting
// there.
var rec = recordingClient();
agents = new AgentControl(rec);
ReplyPushLoop loop = loop(1, 100_000);
loop.onTicketTerminal("task-1", WORKER, false); // pendingTickets={task-1}; activeLeads={PRIMARY}
loop.stopOrRestart(PRIMARY, Set.of(), Set.of("task-1"));
assertFalse(loop.isActive(), "a stale ticket already accounted for at decide time must not "
+ "restart the loop — that would defeat the reminder cap");
}
// --- mirror of the two stopOrRestart tests above, for the reply arm ------------------------
//
// Both tests above only ever passed Set.of() for repliesBefore, so racedIn's reply branch
// (`pendingReplyTargetsFor(lead).stream().anyMatch(t -> !repliesBefore.contains(t))`) was
// never exercised by anything other than an always-empty snapshot. The reviewer flagged this:
// racedIn is symmetric in the code, and only half of it was pinned by a test.
@Test
void aReplyStillPendingWhenTheLoopStopsIsNotStranded() {
// Mirrors aTicketStillPendingWhenTheLoopStopsIsNotStranded: a reply target that raced in
// during the decision-to-release window (absent from the "before" snapshot) must reclaim
// the schedule slot rather than being stranded with no schedule left to nudge about it.
var rec = recordingClient();
agents = new AgentControl(rec);
inbox.publish(WORKER, "m1", "hello");
ReplyPushLoop loop = loop(1, 100_000); // long backoff — no natural tick fires during this test
loop.onReplyQueued(WORKER); // pendingReplies={term_worker}; activeLeads={PRIMARY}
loop.stopOrRestart(PRIMARY, Set.of(), Set.of());
assertTrue(loop.isActive(), "a reply that raced the loop's stop must reclaim the schedule "
+ "slot, not be stranded with no schedule left to ever nudge about it");
}
@Test
void aStaleUncollectedReplyAtCapDoesNotRestartTheLoop() {
// Mirrors aStaleUncollectedTicketAtCapDoesNotRestartTheLoop: a reply target already
// accounted for at decide time (present in repliesBefore) must not restart the loop —
// that is the reminder cap doing its job, not a race.
var rec = recordingClient();
agents = new AgentControl(rec);
inbox.publish(WORKER, "m1", "hello");
ReplyPushLoop loop = loop(1, 100_000);
loop.onReplyQueued(WORKER); // pendingReplies={term_worker}; activeLeads={PRIMARY}
loop.stopOrRestart(PRIMARY, Set.of(WORKER), Set.of());
assertFalse(loop.isActive(), "a stale reply target already accounted for at decide time "
+ "must not restart the loop — that would defeat the reminder cap");
}
@Test
void ticketNudgeFormatIsCorrect() {
String single = ReplyPushLoop.TICKET_NUDGE_FORMAT.formatted("task-1", "", "task-1");
assertTrue(single.contains("Ticket task-1"));
assertTrue(single.contains("bridge_poll(ticket=task-1)"));
String multi = ReplyPushLoop.TICKETS_NUDGE_FORMAT.formatted(2, "", "task-1, task-2");
assertTrue(multi.contains("2 tickets"));
assertTrue(multi.contains("bridge_poll(ticket=...)"));
}
// --- CB-307 nudge path is unchanged (regression) --------------------------------------------
@Test
void inboxNudgeStillUsesTheOriginalTargetPollCall() {
String nudge = ReplyPushLoop.NUDGE_FORMAT.formatted(WORKER, WORKER);
assertTrue(nudge.contains("bridge_poll(target=" + WORKER + ")"),
"CB-588/CB-590 must not change the CB-307 inbox nudge's call shape");
}
// --- CB-590: one schedule per lead — no overlap, no lost nudges -----------------------------
@Test
void replyAndTicketForTheSameLeadCoalesceIntoOneSendNeverOverlapping() throws Exception {
var rec = recordingClient();
agents = new AgentControl(rec);
inbox.publish(WORKER, "m1", "hello");
var loop = loop(1, 300); // backoff wide enough that both entry points land before the first tick
loop.onReplyQueued(WORKER);
loop.onTicketTerminal("task-1", WORKER, false);
assertTrue(rec.sendLatch.await(3, TimeUnit.SECONDS), "one combined nudge should have been sent");
Thread.sleep(300);
assertEquals(1, rec.sendCount(),
"a reply and a ticket for the same lead must coalesce onto ONE schedule — "
+ "two nudge injections into the same lead pane must never overlap");
String nudge = rec.sentParams().getFirst().getValue().toString();
assertTrue(nudge.contains("bridge_poll(target=" + WORKER + ")"),
"the combined nudge must still mention the reply: " + nudge);
assertTrue(nudge.contains("task-1"), "the combined nudge must still mention the ticket: " + nudge);
}
@Test
void aReplyQueuedWhileTheLeadIsBusyIsNotLostWhenATicketArrivesToo() throws Exception {
// The lead is busy for its first two status checks, then becomes injectable. A reply is
// queued while busy; a ticket for the same lead arrives before the lead frees up. Neither
// may be dropped — deferred is fine, lost is not (acceptance criterion #2).
var rec = new BusyThenIdleHerdrClient(2);
agents = new AgentControl(rec);
inbox.publish(WORKER, "m1", "hello");
var loop = loop(1, 50); // cap=1: WAIT_BUSY doesn't count against it, so exactly one send once injectable
loop.onReplyQueued(WORKER); // schedule starts, first tick(s) WAIT_BUSY
loop.onTicketTerminal("task-1", WORKER, false); // coalesces onto the same waiting schedule
assertTrue(rec.sendLatch.await(3, TimeUnit.SECONDS),
"once the lead becomes injectable, the deferred work must still be nudged");
Thread.sleep(200);
assertEquals(1, rec.sendCount(), "exactly one nudge once injectable — reply and ticket coalesced");
String nudge = rec.sentParams().getFirst().getValue().toString();
assertTrue(nudge.contains("bridge_poll(target=" + WORKER + ")"), "the reply must not be dropped: " + nudge);
assertTrue(nudge.contains("task-1"), "the ticket must not be dropped: " + nudge);
}
// --- CB-590 follow-up: per-source reminder budgets — the regression this round exists for ---
@Test
void oneExhaustedSourceDoesNotBlockANudgeForTheOtherSource() {
// The live trace this ticket was filed from: an undrained reply target got nudged up to
// its cap (5 reminders), then a ticket for the SAME lead went terminal shortly before the
// next scheduled tick — so it coalesced onto the still-active schedule (arriving BEFORE
// that tick's "before" snapshot, not during the decision-to-release race stopOrRestart
// guards). With CB-590's single shared reminder counter, that tick's decide() saw
// reminderCount already at the cap and returned STOP regardless of the ticket, and because
// the ticket was already present in that tick's "before" snapshot, stopOrRestart's
// racedIn check (proven correct on its own above) did not save it either — it is not a
// race, it looks like ordinary stale backlog. The ticket was then stranded: pending
// forever with no live schedule, never named in any nudge.
//
// Fixed by giving each source its own counter. Here the reply source is AT its cap (2/2)
// and the ticket source has NEVER been nudged (0/2) — decide() must still return INJECT,
// because the ticket is still eligible on its own budget.
agents = agentWithStatus("idle");
inbox.publish(WORKER, "m1", "hello");
var loop = loop(2, 100_000); // huge backoff — this test drives decide()/isActive() directly
loop.onReplyQueued(WORKER);
loop.onTicketTerminal("task-1", WORKER, false);
assertEquals(ReplyPushLoop.Action.INJECT, loop.decide(PRIMARY, 2, 0),
"the reply source is exhausted (2/2), but the ticket source has never been "
+ "nudged (0/2) — the lead must still be injected so the ticket is not "
+ "lost, exactly the CB-590 follow-up regression");
// isActive() (criterion #5): a real tick that takes the INJECT branch above never calls
// stopOrRestart, so the schedule started by onTicketTerminal above stays live — the
// ticket is not left stranded with isActive()==false while it is still pending.
assertTrue(loop.isActive(), "the schedule must stay active while the ticket source still "
+ "has budget left, even though the reply source sharing it is exhausted");
}
@Test
void bothSourcesExhaustedIsStillStop() {
// The flip side: per-source budgets must not turn into unbounded nudging. When BOTH
// sources are at their cap, decide() must still STOP — a per-source budget is still a
// budget.
agents = agentWithStatus("idle");
inbox.publish(WORKER, "m1", "hello");
var loop = loop(2, 100_000);
loop.onReplyQueued(WORKER);
loop.onTicketTerminal("task-1", WORKER, false);
assertEquals(ReplyPushLoop.Action.STOP, loop.decide(PRIMARY, 2, 2),
"both the reply and the ticket source are at their own cap — must still stop");
}
// --- CB-598: work arriving during a backoff must not read as stale backlog -------------------
@Test
void aTargetArrivingDuringTheBackoffGetsNudgedDespiteAnAlreadyCappedSibling() {
// The bug: reminder counts used to be a single counter per lead per source, carried
// forward across scheduled ticks (scheduleNext(lead, count + 1, ...)) rather than tracked
// per pending item. WORKER gets nudged once here, which — with cap=1 — exhausts the
// shared reply-source counter for this lead. WORKER2 then queues a reply for the SAME
// lead "during the backoff": while the schedule from WORKER's tick is still active, before
// the next tick's own start-of-tick snapshot runs. At that next tick, the OLD code passed
// the already-exhausted shared counter into decide() regardless of WORKER2 never having
// been named in any nudge, and — because WORKER2 was already present in that tick's
// "before" snapshot — stopOrRestart's race check (proven correct on its own elsewhere in
// this file) does not save it either: it looks like ordinary stale backlog, not a race.
// WORKER2 was then stranded forever with no live schedule and no nudge ever naming it.
//
// tick() is driven directly (package-private, same reasoning as stopOrRestart being
// directly testable) so the exact interleaving is deterministic instead of racing the
// scheduler thread over a real ~15s backoff.
//
// Before the fix, this test fails on the second assertEquals: rec.sendCount() stays at 1
// (decide() returns STOP on the second tick(), so injectNudge is never called a second
// time) and the "must still get one" assertion never even runs.
int cap = 1;
var rec = recordingClient();
agents = new AgentControl(rec);
inbox.own(WORKER2);
inbox.publish(WORKER, "m1", "hello");
var loop = loop(cap, 100_000); // huge backoff — nothing fires on its own; we drive tick()
loop.onReplyQueued(WORKER);
loop.tick(PRIMARY); // first tick: nudges WORKER alone; WORKER's own count reaches the cap
assertEquals(1, rec.sendCount(), "the first tick should nudge about WORKER");
// WORKER2 "arrives during the backoff": queued for the same lead while the schedule from
// the tick above is still active (activeLeads still holds PRIMARY), before the next tick
// (simulated below) takes its own start-of-tick snapshot.
inbox.publish(WORKER2, "m2", "hello2");
loop.onReplyQueued(WORKER2);
loop.tick(PRIMARY); // the tick that would fire once that backoff elapsed
assertEquals(2, rec.sendCount(),
"WORKER2 was never named in any nudge yet and must still get one, even though "
+ "WORKER's own reminder count is already at the cap");
String secondNudge = rec.sentParams().get(1).getValue().toString();
assertTrue(secondNudge.contains(WORKER2), "the never-named target must be named: " + secondNudge);
// Criterion #3: isActive() must reflect that this lead still had a live nudge to give —
// the second tick took the INJECT branch, so the schedule stayed live rather than being
// torn down under WORKER2.
assertTrue(loop.isActive(), "the schedule must stay active after nudging the fresh target");
}
@Test
void aTicketArrivingDuringTheBackoffGetsNudgedDespiteAnAlreadyCappedSibling() {
// Mirrors the reply-side test above for the ticket source.
int cap = 1;
var rec = recordingClient();
agents = new AgentControl(rec);
var loop = loop(cap, 100_000);
loop.onTicketTerminal("task-1", WORKER, false);
loop.tick(PRIMARY); // first tick: nudges task-1 alone; its count reaches the cap
assertEquals(1, rec.sendCount(), "the first tick should nudge about task-1");
loop.onTicketTerminal("task-2", WORKER, false); // arrives during the backoff, same lead
loop.tick(PRIMARY);
assertEquals(2, rec.sendCount(),
"task-2 was never named in any nudge yet and must still get one, even though "
+ "task-1's reminder count is already at the cap");
String secondNudge = rec.sentParams().get(1).getValue().toString();
assertTrue(secondNudge.contains("task-2"), "the never-named ticket must be named: " + secondNudge);
assertTrue(loop.isActive(), "the schedule must stay active after nudging the fresh ticket");
}
// --- CB-582: bridge_ask question-open nudges -------------------------------------------------
@Test
void onQuestionOpenedWithNoKnownLeadNeverStartsASchedule() throws Exception {
var rec = recordingClient();
agents = new AgentControl(rec);
var loop = new ReplyPushLoop(new PrimaryRegistry(null), agents, inbox, scheduler, 5, 50);
loop.onQuestionOpened("task-1", WORKER, "term_worker#1", "which config?");
assertFalse(loop.isActive(), "no lead known means nothing to nudge yet");
Thread.sleep(150);
assertEquals(0, rec.sendCount(), "must not nudge when no lead is known to be waiting");
}
@Test
void decideQuestionsWithNothingPendingIsStop() {
agents = agentWithStatus("idle");
assertEquals(ReplyPushLoop.Action.STOP, loop().decide(PRIMARY, 0, 0, 0));
}
@Test
void decideQuestionsAtCapIsStop() {
agents = agentWithStatus("idle");
var loop = loop(2, 100_000);
loop.onQuestionOpened("task-1", WORKER, "term_worker#1", "which config?");
assertEquals(ReplyPushLoop.Action.STOP, loop.decide(PRIMARY, 0, 0, 2));
}
@Test
void decideQuestionsUnderCapWithInjectableLeadIsInject() {
agents = agentWithStatus("idle");
var loop = loop(5, 100_000);
loop.onQuestionOpened("task-1", WORKER, "term_worker#1", "which config?");
assertEquals(ReplyPushLoop.Action.INJECT, loop.decide(PRIMARY, 0, 0, 0));
}
@Test
void onQuestionOpenedCausesExactlyOneNudgeNamingTheTicketAndTurnId() throws Exception {
var rec = recordingClient();
agents = new AgentControl(rec);
loop(1, 50).onQuestionOpened("task-1", WORKER, "term_worker#1", "which config file?");
assertTrue(rec.sendLatch.await(3, TimeUnit.SECONDS), "one question nudge should have been sent");
assertEquals(1, rec.sendCount());
String nudge = rec.sentParams().getFirst().getValue().toString();
assertTrue(nudge.contains("task-1"), "nudge should name the ticket: " + nudge);
assertTrue(nudge.contains("term_worker#1"), "nudge should name the turnId: " + nudge);
assertTrue(nudge.contains("bridge_send(turnId="), "nudge should name the exact answer call: " + nudge);
assertTrue(nudge.contains("which config file?"), "nudge should include the question text: " + nudge);
}
@Test
void questionClosedPreventsFurtherNudging() throws Exception {
var rec = recordingClient();
agents = new AgentControl(rec);
var loop = loop(1, 100);
loop.onQuestionOpened("task-1", WORKER, "term_worker#1", "which config?");
loop.questionClosed("term_worker#1"); // answered/lapsed before the first tick fired
Thread.sleep(300); // let the scheduled tick run
assertEquals(0, rec.sendCount(), "an already-closed question must never be nudged");
}
@Test
void questionNudgesSendUpToCapThenStop() throws Exception {
int cap = 2;
var rec = recordingClient();
agents = new AgentControl(rec);
rec.sendLatch = new CountDownLatch(cap);
loop(cap, 50).onQuestionOpened("task-1", WORKER, "term_worker#1", "which config?");
assertTrue(rec.sendLatch.await(5, TimeUnit.SECONDS), cap + " question nudges should have fired");
Thread.sleep(300);
assertEquals(cap, rec.sendCount(), "exactly " + cap + " question nudges (cap=" + cap + ")");
}
@Test
void questionAndTicketForTheSameLeadCoalesceIntoOneSend() throws Exception {
var rec = recordingClient();
agents = new AgentControl(rec);
var loop = loop(1, 300); // backoff wide enough that both entry points land before the first tick
loop.onTicketTerminal("task-1", WORKER, false);
loop.onQuestionOpened("task-2", WORKER, "term_worker#1", "which config?");
assertTrue(rec.sendLatch.await(3, TimeUnit.SECONDS), "one combined nudge should have been sent");
Thread.sleep(300);
assertEquals(1, rec.sendCount(),
"a ticket and a question for the same lead must coalesce onto ONE schedule");
String nudge = rec.sentParams().getFirst().getValue().toString();
assertTrue(nudge.contains("task-1"), "the combined nudge must still mention the ticket: " + nudge);
assertTrue(nudge.contains("term_worker#1"), "the combined nudge must still mention the question: " + nudge);
}
@Test
void oneExhaustedQuestionSourceDoesNotBlockANudgeForTheOtherSources() {
// Mirrors oneExhaustedSourceDoesNotBlockANudgeForTheOtherSource for the question source:
// the question source is at its cap (2/2), but the ticket source has never been nudged
// (0/2) — decide() must still INJECT so the ticket is not stranded.
agents = agentWithStatus("idle");
var loop = loop(2, 100_000);
loop.onQuestionOpened("task-1", WORKER, "term_worker#1", "which config?");
loop.onTicketTerminal("task-2", WORKER, false);
assertEquals(ReplyPushLoop.Action.INJECT, loop.decide(PRIMARY, 0, 0, 2),
"the question source is exhausted (2/2), but the ticket source has never been "
+ "nudged (0/2) — the lead must still be injected so the ticket is not lost");
}
@Test
void questionNudgeFormatIsCorrect() {
String single = ReplyPushLoop.QUESTION_NUDGE_FORMAT.formatted(
WORKER, "task-1", "term_worker#1", "which config?");
assertTrue(single.contains("Worker term_worker"));
assertTrue(single.contains("bridge_send(turnId=\"term_worker#1\""));
assertTrue(single.contains("which config?"));
String multi = ReplyPushLoop.QUESTIONS_NUDGE_FORMAT.formatted(2, "task-1 (turnId=t1), task-2 (turnId=t2)");
assertTrue(multi.contains("2 workers"));
assertTrue(multi.contains("bridge_poll(ticket=...)"));
}
// --- metrics (CB-512) ----------------------------------------------------------------------
@Test
@@ -225,14 +808,43 @@ class ReplyPushLoopTest {
agents = agentWithStatus("idle");
inbox.publish(WORKER, "m1", "hello");
Metrics metrics = new Metrics();
var loop = loop(2, 100, metrics);
loop.onReplyQueued(WORKER);
assertEquals(ReplyPushLoop.Action.STOP, loop(2, 100, metrics).decide(WORKER, 2));
assertEquals(ReplyPushLoop.Action.STOP, loop.decide(PRIMARY, 2, 0));
assertEquals(1, metrics.count(BridgedMetrics.PUSH_NUDGES, "outcome", "exhausted"),
"hitting the reminder cap must count as exhausted");
assertEquals(0, metrics.count(BridgedMetrics.PUSH_NUDGES, "outcome", "delivered"));
}
@Test
void successfulTicketNudgeIncrementsDelivered() throws Exception {
var rec = recordingClient();
agents = new AgentControl(rec);
Metrics metrics = new Metrics();
loop(1, 50, metrics).onTicketTerminal("task-1", WORKER, false);
assertTrue(rec.sendLatch.await(3, TimeUnit.SECONDS), "one ticket nudge should have been sent");
Thread.sleep(200);
assertEquals(1, metrics.count(BridgedMetrics.PUSH_NUDGES, "outcome", "delivered"),
"a successfully sent ticket nudge must count as delivered, same metric as CB-307");
}
@Test
void ticketReminderCapIncrementsExhausted() {
agents = agentWithStatus("idle");
Metrics metrics = new Metrics();
var loop = loop(2, 100_000, metrics);
loop.onTicketTerminal("task-1", WORKER, false);
assertEquals(ReplyPushLoop.Action.STOP, loop.decide(PRIMARY, 0, 2));
assertEquals(1, metrics.count(BridgedMetrics.PUSH_NUDGES, "outcome", "exhausted"),
"hitting the ticket reminder cap must count as exhausted");
}
// --- helpers -------------------------------------------------------------------------------
private ReplyPushLoop loop() {
@@ -315,4 +927,49 @@ class ReplyPushLoopTest {
private static RecordingHerdrClient recordingClient() {
return new RecordingHerdrClient();
}
/**
* Thread-safe fake that reports {@code working} (not injectable) for its first
* {@code busyChecks} status calls, then {@code idle} forever after — used to prove work queued
* while the lead is busy is deferred, not dropped, once it becomes injectable.
*/
private static final class BusyThenIdleHerdrClient implements HerdrClient {
private final List<Map.Entry<String, Object>> calls =
Collections.synchronizedList(new ArrayList<>());
private final AtomicInteger statusChecks = new AtomicInteger();
private final int busyChecks;
volatile CountDownLatch sendLatch = new CountDownLatch(1);
BusyThenIdleHerdrClient(int busyChecks) {
this.busyChecks = busyChecks;
}
@Override
public JsonNode call(String method, Object params) {
if ("agent.get".equals(method)) {
String status = statusChecks.getAndIncrement() < busyChecks ? "working" : "idle";
return MAPPER.createObjectNode()
.set("agent", MAPPER.createObjectNode()
.put("terminal_id", PRIMARY)
.put("agent_status", status));
}
if ("agent.prompt".equals(method)) {
calls.add(Map.entry(method, params));
sendLatch.countDown();
}
return MAPPER.createObjectNode();
}
long sendCount() {
return calls.size();
}
List<Map.Entry<String, Object>> sentParams() {
return List.copyOf(calls);
}
@Override
public void close() {
}
}
}
@@ -0,0 +1,20 @@
package dev.ltms.bridged.msg;
/**
* Explicit unbound tokens for tests that exercise delivery without an accepted send.
*
* <p>The waiter is {@code null} on purpose. "No accepted send" is an <em>absence</em>, and a helper
* that handed back a fresh {@code CompletableFuture} would invent one — which is how the first
* version of this class turned {@code captureBaselineSkipsTheReadWhenNoSendIsWaiting} red: the
* resolver saw a non-null waiter, decided a turn was in flight, and scraped a pane that no send was
* blocked on. An inert value must omit the fact, never fabricate it.
*/
public final class TestTurnTokens {
private TestTurnTokens() {
}
/** A token for a delivery that no send is waiting on: it authorises nothing. */
public static TurnToken inert(String target) {
return new TurnToken(target, null);
}
}
@@ -0,0 +1,57 @@
package dev.ltms.bridged.peer;
import org.junit.jupiter.api.Test;
import static org.junit.jupiter.api.Assertions.assertEquals;
import static org.junit.jupiter.api.Assertions.assertFalse;
import static org.junit.jupiter.api.Assertions.assertNotEquals;
import static org.junit.jupiter.api.Assertions.assertNull;
import static org.junit.jupiter.api.Assertions.assertTrue;
class CharterReceiptTest {
@Test
void recordsNoCharterConfiguredDistinctFromCharterDelivered() {
// No role charter configured — only the reply charter is composed. Source is "none", but
// text was still delivered, so absent() is false and the digest is present.
CharterReceipt viaReply = CharterReceipt.compose(MemberRole.DEV, "s", null, "reply charter");
// A role charter was configured AND delivered.
CharterReceipt delivered = CharterReceipt.compose(MemberRole.DEV, "s", "role charter",
"role charter\n\nreply charter");
// The two cases must not collapse: the no-role-charter case reports "none", the delivered
// case reports the config key, and their digests differ.
assertEquals(CharterReceipt.NO_SOURCE, viaReply.charterSource());
assertEquals("fleet.charters.dev", delivered.charterSource());
assertNotEquals(viaReply.charterSource(), delivered.charterSource());
assertNotEquals(viaReply.charterSha256(), delivered.charterSha256());
// Both actually delivered text — the distinction is the source and digest, not absence.
assertFalse(viaReply.absent());
assertFalse(delivered.absent());
}
@Test
void recordsExplicitAbsenceWhenNoCharterIsComposed() {
CharterReceipt none = CharterReceipt.compose(MemberRole.REVIEWER, "s", null, null);
assertTrue(none.absent());
assertNull(none.charterSha256());
assertEquals(0, none.charterBytes());
assertEquals(CharterReceipt.NO_SOURCE, none.charterSource(),
"no configured charter and nothing composed still reports a source, never a gap");
}
@Test
void digestIsStableForSameTextAndDiffersForDifferentText() {
assertEquals(CharterReceipt.digestOf("charter-aaa"), CharterReceipt.digestOf("charter-aaa"),
"the same text must always produce the same digest");
assertNotEquals(CharterReceipt.digestOf("charter-aaa"), CharterReceipt.digestOf("charter-bbb"),
"different text must produce a different digest");
assertNull(CharterReceipt.digestOf(""), "blank text carries no digest");
// The record's fingerprint matches the standalone digest for the same composed string.
CharterReceipt r = CharterReceipt.compose(MemberRole.DEV, "s", "role", "the composed text");
assertEquals(CharterReceipt.digestOf("the composed text"), r.charterSha256());
assertFalse(r.absent());
}
}
@@ -0,0 +1,119 @@
package dev.ltms.bridged.placement;
import org.junit.jupiter.api.Test;
import java.util.Map;
import java.util.OptionalLong;
import java.util.concurrent.TimeUnit;
import java.util.concurrent.atomic.AtomicLong;
import static org.junit.jupiter.api.Assertions.assertEquals;
import static org.junit.jupiter.api.Assertions.assertFalse;
import static org.junit.jupiter.api.Assertions.assertThrows;
import static org.junit.jupiter.api.Assertions.assertTrue;
/**
* CB-578 stage B: the credential-keyed quarantine tracker itself, isolated from placement/spawn
* wiring (that's {@code CompositePeerLauncherTest}). The clock is a plain {@link AtomicLong} of
* nanos so expiry is exercised without a real sleep.
*/
class BackendQuarantineTest {
@Test
void aFreshCredentialIsNotQuarantined() {
BackendQuarantine q = new BackendQuarantine(() -> 0L, TimeUnit.MINUTES.toNanos(30));
assertFalse(q.isQuarantined("shared-openai"));
assertEquals(OptionalLong.empty(), q.remainingSeconds("shared-openai"));
}
@Test
void quarantineBlocksTheCredentialForTheFullCooldown() {
BackendQuarantine q = new BackendQuarantine(() -> 0L, TimeUnit.MINUTES.toNanos(30));
q.quarantine("shared-openai");
assertTrue(q.isQuarantined("shared-openai"));
assertEquals(OptionalLong.of(1800L), q.remainingSeconds("shared-openai"));
}
@Test
void onlyTheQuarantinedCredentialIsAffected() {
BackendQuarantine q = new BackendQuarantine(() -> 0L, TimeUnit.MINUTES.toNanos(30));
q.quarantine("shared-openai");
assertFalse(q.isQuarantined("some-other-credential"),
"an unrelated credential must not be swept into the quarantine");
}
@Test
void expiresOnTheInjectedClock() {
AtomicLong now = new AtomicLong(0L);
BackendQuarantine q = new BackendQuarantine(now::get, TimeUnit.MINUTES.toNanos(30));
q.quarantine("shared-openai");
assertTrue(q.isQuarantined("shared-openai"));
now.set(TimeUnit.MINUTES.toNanos(29));
assertTrue(q.isQuarantined("shared-openai"), "still inside the cooldown");
now.set(TimeUnit.MINUTES.toNanos(31));
assertFalse(q.isQuarantined("shared-openai"), "the cooldown has elapsed on the injected clock");
assertEquals(OptionalLong.empty(), q.remainingSeconds("shared-openai"));
}
@Test
void aRepeatQuarantineCallRestartsTheCooldownAtFullLength() {
AtomicLong now = new AtomicLong(0L);
BackendQuarantine q = new BackendQuarantine(now::get, TimeUnit.MINUTES.toNanos(30));
q.quarantine("shared-openai");
now.set(TimeUnit.MINUTES.toNanos(20));
q.quarantine("shared-openai");
now.set(TimeUnit.MINUTES.toNanos(45)); // 25 min after the second call, 45 after the first
assertTrue(q.isQuarantined("shared-openai"),
"a fresh exhaustion resets the cooldown to full length, not the earlier shorter wait");
}
@Test
void activeRemainingSecondsListsOnlyStillQuarantinedCredentials() {
AtomicLong now = new AtomicLong(0L);
BackendQuarantine q = new BackendQuarantine(now::get, TimeUnit.MINUTES.toNanos(30));
q.quarantine("shared-openai");
q.quarantine("another-credential");
now.set(TimeUnit.MINUTES.toNanos(31));
q.quarantine("shared-openai"); // re-quarantined after the first one expired
Map<String, Long> active = q.activeRemainingSeconds();
assertEquals(Map.of("shared-openai", 1800L), active,
"the expired credential is dropped; the re-quarantined one is reported");
}
@Test
void noneReportsNothingQuarantinedWhenNeverToldTo() {
BackendQuarantine q = BackendQuarantine.none();
assertFalse(q.isQuarantined("anything"));
assertTrue(q.activeRemainingSeconds().isEmpty());
}
@Test
void noneIgnoresAQuarantineCallInsteadOfLockingTheCredentialForever() {
// .none() holds a clock frozen at 0, so if #quarantine recorded a deadline the credential
// would never expire — locked out for the life of the daemon. Two production
// CompositePeerLauncher constructors default to none(), so that failure would be silent and
// permanent. An inert stand-in must omit the fact, never invent one.
BackendQuarantine q = BackendQuarantine.none();
q.quarantine("shared-openai");
assertFalse(q.isQuarantined("shared-openai"), "none() must not quarantine anything");
assertTrue(q.remainingSeconds("shared-openai").isEmpty());
assertTrue(q.activeRemainingSeconds().isEmpty());
}
@Test
void aNonPositiveCooldownIsRejected() {
assertThrows(IllegalArgumentException.class, () -> new BackendQuarantine(() -> 0L, 0L));
assertThrows(IllegalArgumentException.class, () -> new BackendQuarantine(() -> 0L, -1L));
}
}
@@ -24,7 +24,13 @@ class PlacementPolicyTest {
private static PlacementContext ctx(List<PlacementCandidate> candidates,
Function<String, Integer> liveCount,
Set<String> unreachable) {
return new PlacementContext("b", candidates, liveCount, unreachable);
return ctx(candidates, liveCount, unreachable, Set.of());
}
private static PlacementContext ctx(List<PlacementCandidate> candidates,
Function<String, Integer> liveCount,
Set<String> unreachable, Set<String> quarantined) {
return new PlacementContext("b", candidates, liveCount, unreachable, quarantined);
}
private static PlacementContext ctx(List<PlacementCandidate> candidates,
@@ -46,17 +52,37 @@ class PlacementPolicyTest {
PlacementPolicy policy = PlacementPolicies.fixed();
PlacementContext ctx = new PlacementContext(null,
List.of(PlacementCandidate.profile("a"), PlacementCandidate.profile("b")),
noSessions(), Set.of());
noSessions(), Set.of(), Set.of());
assertEquals("a", policy.select(ctx).profile());
}
@Test
void fixedThrowsWhenNoProfilesAndNoDefault() {
PlacementPolicy policy = PlacementPolicies.fixed();
PlacementContext ctx = new PlacementContext(null, List.of(), noSessions(), Set.of());
PlacementContext ctx = new PlacementContext(null, List.of(), noSessions(), Set.of(), Set.of());
assertThrows(PlacementException.class, () -> policy.select(ctx));
}
@Test
void fixedSkipsQuarantinedDefault() {
PlacementPolicy policy = PlacementPolicies.fixed();
PlacementContext ctx = new PlacementContext("b",
List.of(PlacementCandidate.profile("a"), PlacementCandidate.profile("b")),
noSessions(), Set.of(), Set.of("b"));
assertEquals("a", policy.select(ctx).profile(),
"the default 'b' is quarantined, so fixed falls through to the first un-quarantined candidate");
}
@Test
void fixedThrowsWhenDefaultAndEveryCandidateQuarantined() {
PlacementPolicy policy = PlacementPolicies.fixed();
PlacementContext ctx = new PlacementContext("b",
List.of(PlacementCandidate.profile("a"), PlacementCandidate.profile("b")),
noSessions(), Set.of(), Set.of("a", "b"));
PlacementException e = assertThrows(PlacementException.class, () -> policy.select(ctx));
assertTrue(e.getMessage().contains("quarantined"), e.getMessage());
}
@Test
void roundRobinCyclesThroughAvailableProfiles() {
PlacementPolicy policy = PlacementPolicies.roundRobin();
@@ -82,6 +108,18 @@ class PlacementPolicyTest {
}
}
@Test
void roundRobinSkipsQuarantinedProfiles() {
PlacementPolicy policy = PlacementPolicies.roundRobin();
List<PlacementCandidate> candidates = List.of(
PlacementCandidate.profile("a"),
PlacementCandidate.profile("b"));
PlacementContext ctx = ctx(candidates, noSessions(), Set.of(), Set.of("a"));
for (int i = 0; i < 5; i++) {
assertEquals("b", policy.select(ctx).profile(), "a is quarantined, so every pick lands on b");
}
}
@Test
void roundRobinThrowsWhenAllAtMaxLoad() {
PlacementPolicy policy = PlacementPolicies.roundRobin();
@@ -150,6 +188,29 @@ class PlacementPolicyTest {
assertTrue(e.getMessage().contains("maxLoad"), e.getMessage());
}
@Test
void weightedSkipsQuarantinedProfile() {
PlacementPolicy policy = PlacementPolicies.weighted();
List<PlacementCandidate> candidates = List.of(
PlacementCandidate.profile("a", 1.0f, null),
PlacementCandidate.profile("b", 1.0f, null));
PlacementContext ctx = ctx(candidates, noSessions(), Set.of(), Set.of("a"));
for (int i = 0; i < 5; i++) {
assertEquals("b", policy.select(ctx).profile(), "a is quarantined, so every pick lands on b");
}
}
@Test
void weightedThrowsWhenAllQuarantined() {
PlacementPolicy policy = PlacementPolicies.weighted();
List<PlacementCandidate> candidates = List.of(
PlacementCandidate.profile("a"),
PlacementCandidate.profile("b"));
PlacementException e = assertThrows(PlacementException.class,
() -> policy.select(ctx(candidates, noSessions(), Set.of(), Set.of("a", "b"))));
assertTrue(e.getMessage().contains("quarantined"), e.getMessage());
}
@Test
void weightedThrowsWhenAllUnreachable() {
PlacementPolicy policy = PlacementPolicies.weighted();
@@ -176,8 +237,134 @@ class PlacementPolicyTest {
assertTrue(e.getMessage().contains("1 unreachable"), e.getMessage());
}
// --- CB-554: weight <= 0 excludes a candidate from automatic selection ------------------
@Test
void weightedSkipsWeightZeroProfile() {
PlacementPolicy policy = PlacementPolicies.weighted();
List<PlacementCandidate> candidates = List.of(
PlacementCandidate.profile("a", 0.0f, null),
PlacementCandidate.profile("b", 1.0f, null));
for (int i = 0; i < 5; i++) {
assertEquals("b", policy.select(ctx(candidates, noSessions())).profile(),
"a has weight 0, so every pick lands on b");
}
}
@Test
void weightedThrowsWhenAllWeightZero() {
PlacementPolicy policy = PlacementPolicies.weighted();
List<PlacementCandidate> candidates = List.of(
PlacementCandidate.profile("a", 0.0f, null),
PlacementCandidate.profile("b", 0.0f, null));
PlacementException e = assertThrows(PlacementException.class,
() -> policy.select(ctx(candidates, noSessions())));
assertTrue(e.getMessage().contains("weight 0"), e.getMessage());
}
@Test
void weightedTreatsNegativeWeightAsExcludedNotError() {
PlacementPolicy policy = PlacementPolicies.weighted();
List<PlacementCandidate> candidates = List.of(
PlacementCandidate.profile("a", -1.0f, null),
PlacementCandidate.profile("b", 1.0f, null));
for (int i = 0; i < 5; i++) {
assertEquals("b", policy.select(ctx(candidates, noSessions())).profile(),
"a has a negative weight, so it is excluded like weight 0 — never picked, never an error");
}
}
@Test
void weightZeroAndQuarantinedIsNotDoubleCounted() {
PlacementPolicy policy = PlacementPolicies.weighted();
List<PlacementCandidate> candidates = List.of(
PlacementCandidate.profile("a", 0.0f, null),
PlacementCandidate.profile("b", 1.0f, null),
PlacementCandidate.profile("c", 1.0f, null));
// a is BOTH weight-0 and quarantined; b and c are quarantined only. If a were counted
// into both buckets, "quarantined" would read 3 (or the arithmetic would go negative)
// instead of the true count of 2 candidates actually quarantined.
PlacementContext mixedCtx = ctx(candidates, noSessions(), Set.of(), Set.of("a", "b", "c"));
PlacementException e = assertThrows(PlacementException.class, () -> policy.select(mixedCtx));
assertTrue(e.getMessage().contains("1 weight-0"), e.getMessage());
assertTrue(e.getMessage().contains("2 quarantined"), e.getMessage());
assertTrue(e.getMessage().contains("0 remaining"), e.getMessage());
}
@Test
void roundRobinSkipsWeightZeroProfiles() {
PlacementPolicy policy = PlacementPolicies.roundRobin();
List<PlacementCandidate> candidates = List.of(
PlacementCandidate.profile("a", 0.0f, null),
PlacementCandidate.profile("b", 1.0f, null));
for (int i = 0; i < 5; i++) {
assertEquals("b", policy.select(ctx(candidates, noSessions())).profile(),
"a has weight 0, so every pick lands on b");
}
}
@Test
void roundRobinThrowsWhenAllWeightZero() {
PlacementPolicy policy = PlacementPolicies.roundRobin();
List<PlacementCandidate> candidates = List.of(
PlacementCandidate.profile("a", 0.0f, null),
PlacementCandidate.profile("b", 0.0f, null));
PlacementException e = assertThrows(PlacementException.class,
() -> policy.select(ctx(candidates, noSessions())));
assertTrue(e.getMessage().contains("weight 0"), e.getMessage());
}
@Test
void fixedSkipsWeightZeroDefault() {
PlacementPolicy policy = PlacementPolicies.fixed();
PlacementContext ctx = new PlacementContext("b",
List.of(PlacementCandidate.profile("a", 1.0f, null),
PlacementCandidate.profile("b", 0.0f, null)),
noSessions(), Set.of(), Set.of());
assertEquals("a", policy.select(ctx).profile(),
"the default 'b' has weight 0, so fixed falls through to the first available candidate");
}
@Test
void fixedThrowsWhenDefaultAndEveryCandidateWeightZero() {
PlacementPolicy policy = PlacementPolicies.fixed();
PlacementContext ctx = new PlacementContext("b",
List.of(PlacementCandidate.profile("a", 0.0f, null),
PlacementCandidate.profile("b", 0.0f, null)),
noSessions(), Set.of(), Set.of());
PlacementException e = assertThrows(PlacementException.class, () -> policy.select(ctx));
assertTrue(e.getMessage().contains("weight 0"), e.getMessage());
}
@Test
void unknownPolicyNameThrows() {
assertThrows(IllegalArgumentException.class, () -> PlacementPolicies.fromName("random"));
}
// --- CB-585: maxLoad 0 caps a profile at zero live members, even with nothing live yet -----
@Test
void weightedSkipsMaxLoadZeroProfileEvenWithZeroLiveWorkers() {
PlacementPolicy policy = PlacementPolicies.weighted();
List<PlacementCandidate> candidates = List.of(
PlacementCandidate.profile("a", 1.0f, 0),
PlacementCandidate.profile("b", 1.0f, null));
Function<String, Integer> liveCount = _ -> 0;
for (int i = 0; i < 5; i++) {
assertEquals("b", policy.select(ctx(candidates, liveCount)).profile(),
"a has maxLoad 0, so it is already at its cap with nobody live — every pick lands on b");
}
}
@Test
void weightedThrowsWhenAllMaxLoadZero() {
PlacementPolicy policy = PlacementPolicies.weighted();
List<PlacementCandidate> candidates = List.of(
PlacementCandidate.profile("a", 1.0f, 0),
PlacementCandidate.profile("b", 1.0f, 0));
Function<String, Integer> liveCount = _ -> 0;
PlacementException e = assertThrows(PlacementException.class,
() -> policy.select(ctx(candidates, liveCount)));
assertTrue(e.getMessage().contains("maxLoad"), e.getMessage());
}
}
@@ -18,6 +18,8 @@ import dev.ltms.bridged.session.GitWorktrees;
import dev.ltms.bridged.session.SessionManager;
import dev.ltms.bridged.session.Worktrees;
import dev.ltms.bridged.member.ClaudeCodeLauncher;
import dev.ltms.bridged.member.CompositePeerLauncher;
import dev.ltms.bridged.placement.PlacementPolicies;
import io.javalin.Javalin;
import org.junit.jupiter.api.AfterEach;
import org.junit.jupiter.api.Test;
@@ -240,6 +242,43 @@ class BridgedAppTest {
assertFalse(herdr.called("agent.start"), "an unknown profile must not spawn anything");
}
/**
* CB-599: a profile at its {@code maxLoad} cap must not surface as a bare 500 — the caller
* needs a structured, readable reason, distinct from "unknown_profile".
*/
@Test
void spawnAtMaxLoadIs503WithTheCapacityReasonNotABare500() throws Exception {
FakeHerdr herdr = new FakeHerdr();
BridgedConfig.Profile wcfg = new BridgedConfig.Profile(
"ltms-local", "http://gx00.gw:8000", "coder", null, "BRIDGED_WORKER_TOKEN", null,
"tab", "bridged-workers", "worker: {profile} #{n}", null,
null, null, null, null, null, null, null, 0, null, null, null);
Map<String, BridgedConfig.Profile> profiles = Map.of(wcfg.profile(), wcfg);
ClaudeCodeLauncher delegate = new ClaudeCodeLauncher(
new AgentControl(herdr), new WorkspaceControl(herdr), new SubscriptionGuard(Set.of("gx00.gw")),
profiles, wcfg.profile(), k -> "BRIDGED_WORKER_TOKEN".equals(k) ? "tok-abc" : null);
CompositePeerLauncher workers = new CompositePeerLauncher(
List.of(delegate), wcfg.profile(), profiles, PlacementPolicies.fixed(), _ -> 0);
SessionManager sessions = new SessionManager(workers, new GitWorktrees());
this.presence = sessions.asPresence();
Injector injector = new Injector(new AgentControl(herdr));
InMemoryReplyInbox inbox = new InMemoryReplyInbox();
sessions.onAcquire(inbox::own);
MessageService messages = new MessageService(new AgentControl(herdr), injector, new Rendezvous(), inbox);
app = new BridgedApp(herdr, workers, sessions, messages, this.presence, null).build().start("127.0.0.1", 0);
int port = app.port();
HttpResponse<String> res = req(port, "POST", "/members?profile=ltms-local");
assertEquals(503, res.statusCode(), res.body());
JsonNode body = mapper.readTree(res.body());
assertEquals("no_capacity", body.get("error").asText());
String detail = body.get("detail").asText();
assertTrue(detail.contains("ltms-local"), "detail names the profile: " + detail);
assertTrue(detail.contains("maxLoad"), "detail explains the refusal: " + detail);
assertFalse(herdr.called("agent.start"), "at cap, the spawn is refused before any herdr call");
}
@Test
void spawnWorkerReusesExistingWorkerSpace() throws Exception {
// A space labelled "bridged-workers" already exists → no second workspace.create.
@@ -445,6 +484,67 @@ class BridgedAppTest {
assertTrue(mapper.readTree(req(port, "GET", "/sessions/term_a/status").body()).get("ready").asBoolean());
}
/**
* CB-582: a lead polling {@code GET /sessions/{id}/status} on its normal cadence — not the
* ticket-scoped {@code /tasks/{ticket}} — must also see a worker's open async {@code bridge_ask}
* question, since the reverse-rendezvous window it opened with is far shorter than that cadence.
*/
@Test
void sessionStatusReportsAnOpenQuestionWhenTheWorkerIsMidAsk() throws Exception {
FakeHerdr herdr = new FakeHerdr().agentStatus("idle"); // poller delivers the injection
int port = start(herdr, "http://gx00.gw:8000", Set.of("gx00.gw"));
HttpResponse<String> accepted = postMessage(port, "{\"content\":\"do it\",\"wait\":false}");
assertEquals(202, accepted.statusCode());
String ticket = mapper.readTree(accepted.body()).get("ticket").asText();
Thread.sleep(200); // let the background async send open its rendezvous waiter
var ask = java.util.concurrent.CompletableFuture.supplyAsync(() -> {
try {
return postJson(port, "/sessions/term_a/ask",
"{\"question\":\"which config file?\",\"timeoutMs\":5000}");
} catch (Exception e) {
throw new RuntimeException(e);
}
});
JsonNode task;
long deadline = System.currentTimeMillis() + 3000;
do {
task = mapper.readTree(req(port, "GET", "/tasks/" + ticket).body());
if ("asking".equals(task.path("phase").asText())) break;
//noinspection BusyWait
Thread.sleep(10);
} while (System.currentTimeMillis() < deadline);
assertEquals("asking", task.get("phase").asText());
String turnId = task.get("turnId").asText();
HttpResponse<String> status = req(port, "GET", "/sessions/term_a/status");
assertEquals(200, status.statusCode());
JsonNode body = mapper.readTree(status.body());
assertEquals("idle", body.get("status").asText(), "the live status must still be reported");
assertEquals("which config file?", body.get("question").asText());
assertEquals(turnId, body.get("turnId").asText());
assertEquals(ticket, body.get("ticket").asText());
// Answer it — via the same /message route bridge_send uses, keyed by turnId — so the
// background ask thread does not linger past the test.
var answer = java.util.concurrent.CompletableFuture.supplyAsync(() -> {
try {
return postJson(port, "/sessions/term_a/message",
"{\"content\":\"config.yaml\",\"turnId\":\"" + turnId + "\",\"timeoutMs\":4000}");
} catch (Exception e) {
throw new RuntimeException(e);
}
});
HttpResponse<String> askResponse = ask.get(6, java.util.concurrent.TimeUnit.SECONDS);
assertEquals(200, askResponse.statusCode());
assertEquals("config.yaml", mapper.readTree(askResponse.body()).get("answer").asText());
postJson(port, "/sessions/term_a/reply", "{\"content\":\"done\"}");
answer.get(6, java.util.concurrent.TimeUnit.SECONDS);
}
@Test
void stopWorkerInPanePlacementClosesOnlyThePane() throws Exception {
FakeHerdr herdr = new FakeHerdr();
@@ -2,9 +2,11 @@ package dev.ltms.bridged.session;
import java.util.Collections;
import java.util.List;
import java.util.Optional;
import java.util.Set;
import java.util.concurrent.ConcurrentHashMap;
import java.util.concurrent.CopyOnWriteArrayList;
import java.util.concurrent.atomic.AtomicLong;
/** Recording fake {@link Worktrees} for CB-301-ext acceptance tests (no live git). */
public final class FakeWorktrees implements Worktrees {
@@ -22,15 +24,28 @@ public final class FakeWorktrees implements Worktrees {
public record RepoRootCall(String cwd) {
}
public record SnapshotCall(String worktreePath, String branch, String message) {
}
public record PruneCall(String repoRoot, long minAgeMillis) {
}
private final List<AddCall> addCalls = new CopyOnWriteArrayList<>();
private final List<RemoveCall> removeCalls = new CopyOnWriteArrayList<>();
private final List<OverlayCall> overlayCalls = new CopyOnWriteArrayList<>();
private final List<RepoRootCall> repoRootCalls = new CopyOnWriteArrayList<>();
private final List<SnapshotCall> snapshotCalls = new CopyOnWriteArrayList<>();
private final List<PruneCall> pruneCalls = new CopyOnWriteArrayList<>();
private final Set<String> existingPaths = ConcurrentHashMap.newKeySet();
private final Set<String> trackedPaths = ConcurrentHashMap.newKeySet();
private final AtomicLong snapshotSeq = new AtomicLong();
private volatile RuntimeException addFailure;
private volatile RuntimeException snapshotFailure;
private volatile boolean dirty = false;
private volatile String repoRoot = "/repo";
private volatile String prefix = "/worktrees";
private volatile WipRefStats wipRefs = new WipRefStats(0, 0L);
private volatile int pruneResult = 0;
public FakeWorktrees withRepoRoot(String root) {
this.repoRoot = root;
@@ -61,6 +76,30 @@ public final class FakeWorktrees implements Worktrees {
return this;
}
/** Mark the worktree dirty so {@link #hasUncommitted} reports true (simulates uncommitted work). */
public FakeWorktrees withDirty(boolean dirty) {
this.dirty = dirty;
return this;
}
/** Make subsequent {@link #snapshot} calls throw (simulates a failing git snapshot). */
public FakeWorktrees failSnapshot(String message) {
this.snapshotFailure = new WorktreeException(message);
return this;
}
/** Configure the value returned by {@link #wipRefs}. */
public FakeWorktrees withWipRefs(WipRefStats stats) {
this.wipRefs = stats;
return this;
}
/** Configure the value returned by {@link #pruneWipRefs}. */
public FakeWorktrees withPruneResult(int deleted) {
this.pruneResult = deleted;
return this;
}
@Override
public String add(String repoRoot, String branch, String baseRef) {
addCalls.add(new AddCall(repoRoot, branch, baseRef));
@@ -77,6 +116,11 @@ public final class FakeWorktrees implements Worktrees {
removeCalls.add(new RemoveCall(repoRoot, worktreePath));
}
@Override
public boolean hasUncommitted(String worktreePath) {
return dirty;
}
@Override
public void overlayParity(String repoRoot, String worktreePath, List<String> overlay) {
List<String> copied = new java.util.ArrayList<>();
@@ -100,6 +144,26 @@ public final class FakeWorktrees implements Worktrees {
return repoRoot;
}
@Override
public Optional<String> snapshot(String worktreePath, String branch, String message) {
snapshotCalls.add(new SnapshotCall(worktreePath, branch, message));
if (snapshotFailure != null) {
throw snapshotFailure;
}
return Optional.of("wip" + snapshotSeq.incrementAndGet());
}
@Override
public WipRefStats wipRefs(String repoRoot) {
return wipRefs;
}
@Override
public int pruneWipRefs(String repoRoot, long minAgeMillis) {
pruneCalls.add(new PruneCall(repoRoot, minAgeMillis));
return pruneResult;
}
public List<AddCall> addCalls() {
return List.copyOf(addCalls);
}
@@ -127,4 +191,16 @@ public final class FakeWorktrees implements Worktrees {
public OverlayCall lastOverlay() {
return overlayCalls.isEmpty() ? null : overlayCalls.getLast();
}
public List<SnapshotCall> snapshotCalls() {
return List.copyOf(snapshotCalls);
}
public List<PruneCall> pruneCalls() {
return List.copyOf(pruneCalls);
}
public SnapshotCall lastSnapshot() {
return snapshotCalls.isEmpty() ? null : snapshotCalls.getLast();
}
}
@@ -3,9 +3,13 @@ package dev.ltms.bridged.session;
import org.junit.jupiter.api.Test;
import org.junit.jupiter.api.io.TempDir;
import java.nio.charset.StandardCharsets;
import java.nio.file.Files;
import java.nio.file.Path;
import java.util.HashSet;
import java.util.List;
import java.util.Optional;
import java.util.Set;
import java.util.concurrent.TimeUnit;
import static org.junit.jupiter.api.Assertions.*;
@@ -68,6 +72,120 @@ class GitWorktreesTest {
return out;
}
/** Every pending change in {@code cwd} — the whole-tree porcelain status, unlike {@link #status}. */
private static String fullStatus(Path cwd) throws Exception {
Process p = new ProcessBuilder("git", "status", "--porcelain")
.directory(cwd.toFile()).redirectErrorStream(true).start();
String out = new String(p.getInputStream().readAllBytes());
assertTrue(p.waitFor(30, TimeUnit.SECONDS), "git status timed out");
return out;
}
private static String revParse(Path cwd, String ref) throws Exception {
Process p = new ProcessBuilder("git", "-C", cwd.toString(), "rev-parse", ref)
.redirectErrorStream(true).start();
String out = new String(p.getInputStream().readAllBytes()).trim();
assertTrue(p.waitFor(30, TimeUnit.SECONDS), "git rev-parse timed out");
assertEquals(0, p.exitValue(), "git rev-parse " + ref + " failed:\n" + out);
return out;
}
/** The recursive file list of a commit's tree — used to check what a snapshot actually committed. */
private static String lsTree(Path cwd, String ref) throws Exception {
Process p = new ProcessBuilder("git", "-C", cwd.toString(), "ls-tree", "-r", "--name-only", ref)
.redirectErrorStream(true).start();
String out = new String(p.getInputStream().readAllBytes());
assertTrue(p.waitFor(30, TimeUnit.SECONDS), "git ls-tree timed out");
assertEquals(0, p.exitValue(), "git ls-tree " + ref + " failed:\n" + out);
return out;
}
/** The set of paths in {@code git diff --name-only from..to} — used to check exactly what a
* snapshot's tree changed relative to its parent, the same shape {@code git status --porcelain}
* reports for the worktree it was taken from. */
private static Set<String> diffNameOnly(Path cwd, String from, String to) throws Exception {
Process p = new ProcessBuilder("git", "-C", cwd.toString(), "diff", "--name-only", from, to)
.redirectErrorStream(true).start();
String out = new String(p.getInputStream().readAllBytes());
assertTrue(p.waitFor(30, TimeUnit.SECONDS), "git diff timed out");
assertEquals(0, p.exitValue(), "git diff " + from + ".." + to + " failed:\n" + out);
Set<String> paths = new HashSet<>();
for (String line : out.split("\\R")) {
if (!line.isBlank()) {
paths.add(line.trim());
}
}
return paths;
}
/** The set of paths a {@code git status --porcelain} listing names, stripping the two-char status
* code prefix each line carries. */
private static Set<String> porcelainPaths(String porcelain) {
Set<String> paths = new HashSet<>();
for (String line : porcelain.split("\\R")) {
if (!line.isBlank()) {
paths.add(line.substring(3).trim());
}
}
return paths;
}
private static String forEachRef(Path cwd, String pattern) throws Exception {
Process p = new ProcessBuilder("git", "-C", cwd.toString(), "for-each-ref", pattern)
.redirectErrorStream(true).start();
String out = new String(p.getInputStream().readAllBytes());
assertTrue(p.waitFor(30, TimeUnit.SECONDS), "git for-each-ref timed out");
assertEquals(0, p.exitValue(), "git for-each-ref " + pattern + " failed:\n" + out);
return out;
}
/** Write {@code content} as a blob into the object database; returns its sha. */
private static String blobOf(Path cwd, String content) throws Exception {
Process p = new ProcessBuilder("git", "-C", cwd.toString(), "hash-object", "-w", "--stdin")
.redirectErrorStream(true).start();
p.getOutputStream().write(content.getBytes(StandardCharsets.UTF_8));
p.getOutputStream().close();
String out = new String(p.getInputStream().readAllBytes()).trim();
assertTrue(p.waitFor(30, TimeUnit.SECONDS), "git hash-object timed out");
assertEquals(0, p.exitValue(), "git hash-object failed:\n" + out);
return out;
}
/** Build a single-file tree object from {@code blob}; returns the tree's sha. */
private static String treeOf(Path cwd, String path, String blob) throws Exception {
Process p = new ProcessBuilder("git", "-C", cwd.toString(), "mktree")
.redirectErrorStream(true).start();
p.getOutputStream().write(("100644 blob " + blob + "\t" + path + "\n").getBytes(StandardCharsets.UTF_8));
p.getOutputStream().close();
String out = new String(p.getInputStream().readAllBytes()).trim();
assertTrue(p.waitFor(30, TimeUnit.SECONDS), "git mktree timed out");
assertEquals(0, p.exitValue(), "git mktree failed:\n" + out);
return out;
}
/** {@code git commit-tree} rooted at {@code tree} with a chosen committer date; returns the sha. */
private static String commitTree(Path cwd, String tree, String parent, String committerDate,
String message) throws Exception {
ProcessBuilder pb = new ProcessBuilder("git", "-C", cwd.toString(), "commit-tree",
tree, "-p", parent, "-m", message);
pb.environment().put("GIT_COMMITTER_DATE", committerDate);
Process p = pb.redirectErrorStream(true).start();
String out = new String(p.getInputStream().readAllBytes()).trim();
assertTrue(p.waitFor(30, TimeUnit.SECONDS), "git commit-tree timed out");
assertEquals(0, p.exitValue(), "git commit-tree failed:\n" + out);
return out;
}
/** {@code git update-ref <ref> <sha>} — create the snapshot ref directly. */
private static void updateRef(Path cwd, String ref, String sha) throws Exception {
git(cwd, "update-ref", ref, sha);
}
/** True when {@code ref} exists in the repo (for-each-ref on a missing ref is empty, not an error). */
private static boolean refExists(Path cwd, String ref) throws Exception {
return !forEachRef(cwd, ref).trim().isEmpty();
}
/**
* The heart of CB-525: a provisioned worktree must not inherit the primary's MCP servers. Without
* the isolation step the checked-out {@code .mcp.json} carries them in, and a worker navigating
@@ -175,6 +293,46 @@ class GitWorktreesTest {
assertTrue(Files.exists(Path.of(wt).resolve(".mcp.json")), ".mcp.json stub was dropped");
}
/**
* CB-576. {@code hasUncommitted} must treat a freshly-provisioned worktree as clean, but a
* worktree holding a brand-new, never-added file as dirty. The untracked-file-only shape is
* exactly the work lost in the incident — a worker's draft that compiled but was never
* committed because it stopped to ask its lead a question.
*/
@Test
void anUntrackedOnlyWorktreeCountsAsDirty(@TempDir Path tmp) throws Exception {
Path repo = initRepo(tmp.resolve("repo"));
GitWorktrees gitWorktrees = new GitWorktrees(tmp.resolve("wts").toString());
String wt = gitWorktrees.add(repo.toString(), "cb-576-u", "HEAD");
assertFalse(gitWorktrees.hasUncommitted(wt),
"a freshly provisioned worktree must read as clean");
Files.writeString(Path.of(wt).resolve("brand-new.txt"), "draft that was never added\n");
assertTrue(gitWorktrees.hasUncommitted(wt),
"an untracked-only file must count as dirty");
Files.writeString(Path.of(wt).resolve("README.md"), "edited tracked file\n");
assertTrue(gitWorktrees.hasUncommitted(wt),
"a tracked modification must also count as dirty");
}
/**
* CB-576 review. {@code hasUncommitted} must tolerate a missing worktree exactly like
* {@code remove}: an already-gone directory holds no work to lose, and throwing here would
* break teardown — SessionManager.release() calls it before stopping the pane, so an
* exception would orphan a live pane and skip the release notification (CB-516).
*/
@Test
void hasUncommittedOnAMissingWorktreeReturnsFalseWithoutThrowing(@TempDir Path tmp) {
GitWorktrees gitWorktrees = new GitWorktrees(tmp.resolve("wts").toString());
String gone = tmp.resolve("wts").resolve("does-not-exist").toString();
assertFalse(gitWorktrees.hasUncommitted(gone),
"a missing worktree is reported clean, not an error");
}
/** All three protected configs are covered: each one present in a worktree is neutralized and hidden. */
@Test
void allThreeConfigsAreNeutralizedWhenPresent(@TempDir Path tmp) throws Exception {
@@ -205,4 +363,263 @@ class GitWorktreesTest {
assertEquals("", status(Path.of(wt), "opencode.json"), "opencode.json still shows as modified");
assertEquals("", status(Path.of(wt), ".autoenv"), ".autoenv still shows as modified");
}
/**
* CB-578 stage C, acceptance criterion 1. A dirty worktree — a tracked edit plus a brand-new
* untracked file, exactly the shape lost in CB-576 — must land in {@code refs/wip/<branch>}'s
* tree, and that ref must live outside {@code refs/heads} so it never shows up in
* {@code git branch} or gets swept by a branch cleanup.
*/
@Test
void dirtySnapshotCreatesARefWhoseTreeContainsUntrackedAndTrackedChanges(@TempDir Path tmp) throws Exception {
Path repo = initRepo(tmp.resolve("repo"));
GitWorktrees gitWorktrees = new GitWorktrees(tmp.resolve("wts").toString());
String branch = "cb-578-c-a";
String wt = gitWorktrees.add(repo.toString(), branch, "HEAD");
Files.writeString(Path.of(wt).resolve("untracked.txt"), "draft that was never added\n");
Files.writeString(Path.of(wt).resolve("README.md"), "edited tracked file\n");
Optional<String> ref = gitWorktrees.snapshot(wt, branch, "test snapshot");
assertTrue(ref.isPresent(), "a dirty worktree snapshot returns a commit sha");
String tree = lsTree(repo, "refs/wip/" + branch);
assertTrue(tree.contains("untracked.txt"), "snapshot tree must include the untracked file:\n" + tree);
assertTrue(tree.contains("README.md"), "snapshot tree must include the tracked edit:\n" + tree);
String heads = forEachRef(repo, "refs/heads");
assertFalse(heads.contains("refs/wip/"), "the snapshot ref must not live under refs/heads:\n" + heads);
}
/**
* CB-578 stage C, acceptance criterion 2. Building the commit through a temporary
* {@code GIT_INDEX_FILE} must leave the worker's own index, working tree, and HEAD exactly as
* they were — the worker may still be mid-write, and staging into the real index would corrupt
* that.
*/
@Test
void snapshotDoesNotTouchTheWorkersOwnIndexWorkingTreeOrHead(@TempDir Path tmp) throws Exception {
Path repo = initRepo(tmp.resolve("repo"));
GitWorktrees gitWorktrees = new GitWorktrees(tmp.resolve("wts").toString());
String branch = "cb-578-c-b";
String wt = gitWorktrees.add(repo.toString(), branch, "HEAD");
Files.writeString(Path.of(wt).resolve("untracked.txt"), "draft that was never added\n");
Files.writeString(Path.of(wt).resolve("README.md"), "edited tracked file\n");
String headBefore = revParse(Path.of(wt), "HEAD");
String statusBefore = fullStatus(Path.of(wt));
gitWorktrees.snapshot(wt, branch, "test snapshot");
assertEquals(headBefore, revParse(Path.of(wt), "HEAD"), "snapshot must not move HEAD");
assertEquals(statusBefore, fullStatus(Path.of(wt)),
"snapshot must not change the worker's own index or working tree status");
}
/**
* CB-578 stage C, acceptance criterion 3. {@code add -A} (never {@code -f}) respects
* {@code .gitignore}, and that is the only thing keeping a gitignored file (secrets, local
* config) out of a snapshot commit — an untracked file that is NOT ignored must still be
* included, so this isn't just "untracked files are dropped".
*/
@Test
void gitignoredFileIsExcludedButOtherUntrackedFilesAreNot(@TempDir Path tmp) throws Exception {
Path repo = tmp.resolve("repo");
Files.createDirectories(repo);
git(repo, "init", "-q", "-b", "main");
git(repo, "config", "user.email", "test@example.invalid");
git(repo, "config", "user.name", "Test");
Files.writeString(repo.resolve(".gitignore"), ".env\n");
Files.writeString(repo.resolve("README.md"), "seed\n");
git(repo, "add", ".gitignore", "README.md");
git(repo, "commit", "-q", "-m", "seed");
GitWorktrees gitWorktrees = new GitWorktrees(tmp.resolve("wts").toString());
String branch = "cb-578-c-c";
String wt = gitWorktrees.add(repo.toString(), branch, "HEAD");
Files.writeString(Path.of(wt).resolve(".env"), "SECRET=shh\n");
Files.writeString(Path.of(wt).resolve("untracked.txt"), "draft that was never added\n");
Optional<String> ref = gitWorktrees.snapshot(wt, branch, "test snapshot");
assertTrue(ref.isPresent());
String tree = lsTree(repo, "refs/wip/" + branch);
assertFalse(tree.contains(".env"), "a gitignored file must never enter the snapshot:\n" + tree);
assertTrue(tree.contains("untracked.txt"),
"a non-ignored untracked file must still be included:\n" + tree);
}
/** CB-578 stage C. A clean worktree still produces a valid, if tree-identical, commit — the caller
* (SessionManager) is the one that decides not to call this on a clean worktree. */
@Test
void snapshotOfAMissingWorktreeReturnsEmptyWithoutThrowing(@TempDir Path tmp) {
GitWorktrees gitWorktrees = new GitWorktrees(tmp.resolve("wts").toString());
String gone = tmp.resolve("wts").resolve("does-not-exist").toString();
assertTrue(gitWorktrees.snapshot(gone, "some-branch", "msg").isEmpty(),
"a missing worktree has nothing to snapshot, and must not throw");
}
/**
* CB-587, acceptance criteria 1 and 2. A file with a REAL {@code --skip-worktree} bit set in the
* worktree's own index must never enter the snapshot, even though its on-disk content has locally
* diverged from what is committed — that is exactly the divergence {@code --skip-worktree} exists
* to hide from {@code git status}, and a snapshot built from a fresh empty temp index (the bug)
* stages that local content anyway because the fresh index carries none of the real index's flags.
* The snapshot's diff against its parent must list exactly what {@code git status --porcelain}
* reports for the worktree — no more, no less.
*/
@Test
void dirtySnapshotHonoursARealSkipWorktreeBit(@TempDir Path tmp) throws Exception {
Path repo = tmp.resolve("repo");
Files.createDirectories(repo);
git(repo, "init", "-q", "-b", "main");
git(repo, "config", "user.email", "test@example.invalid");
git(repo, "config", "user.name", "Test");
Files.writeString(repo.resolve("protected.cfg"), "committed-value\n");
Files.writeString(repo.resolve("README.md"), "seed\n");
git(repo, "add", "protected.cfg", "README.md");
git(repo, "commit", "-q", "-m", "seed");
GitWorktrees gitWorktrees = new GitWorktrees(tmp.resolve("wts").toString());
String branch = "cb-587-a";
String wt = gitWorktrees.add(repo.toString(), branch, "HEAD");
git(Path.of(wt), "update-index", "--skip-worktree", "protected.cfg");
Files.writeString(Path.of(wt).resolve("protected.cfg"), "locally-diverged-never-commit\n");
Files.writeString(Path.of(wt).resolve("untracked.txt"), "draft that was never added\n");
Files.writeString(Path.of(wt).resolve("README.md"), "edited tracked file\n");
String porcelain = fullStatus(Path.of(wt));
assertFalse(porcelain.contains("protected.cfg"),
"test setup invalid — protected.cfg must not show in git status once skip-worktree is set:\n"
+ porcelain);
Optional<String> ref = gitWorktrees.snapshot(wt, branch, "test snapshot");
assertTrue(ref.isPresent(), "a dirty worktree snapshot returns a commit sha");
Set<String> diffPaths = diffNameOnly(repo, "HEAD", "refs/wip/" + branch);
assertFalse(diffPaths.contains("protected.cfg"),
"a --skip-worktree file's local drift leaked into the snapshot:\n" + diffPaths);
assertEquals(porcelainPaths(porcelain), diffPaths,
"snapshot diff must list exactly what git status --porcelain reports, no more, no less");
}
/**
* CB-587, acceptance criterion 6. If the worktree's real index cannot be resolved/read, snapshot
* must fail loudly (throw) rather than silently falling back to an empty temp index and producing
* a wrong snapshot. SessionManager's caller already catches and WARNs on any exception here — this
* only needs to confirm the failure is not swallowed inside snapshot() itself.
*/
@Test
void snapshotThrowsWhenTheWorktreesRealIndexCannotBeResolved(@TempDir Path tmp) throws Exception {
Path repo = initRepo(tmp.resolve("repo"));
GitWorktrees gitWorktrees = new GitWorktrees(tmp.resolve("wts").toString());
String branch = "cb-587-b";
String wt = gitWorktrees.add(repo.toString(), branch, "HEAD");
// Break git's ability to resolve the worktree's real index by removing the linked worktree's
// `.git` file (which normally points at the main repo's worktrees/<name>/ directory).
Files.delete(Path.of(wt).resolve(".git"));
assertThrows(WorktreeException.class, () -> gitWorktrees.snapshot(wt, branch, "test snapshot"),
"an unresolvable real index must fail loudly, not silently snapshot from an empty index");
}
/**
* CB-586, criterion 2. A snapshot whose content is NOT reachable from {@code main} is the last
* copy of a worker's work, and must never be deleted automatically — even when it is old and
* even when the caller passes a zero age floor. Uses the real snapshot path on a dirty worktree,
* so the unreachable tree is exactly the shape CB-576/CB-578 stage C exist to protect.
*/
@Test
void anUnreachableSnapshotIsNeverPruned(@TempDir Path tmp) throws Exception {
Path repo = initRepo(tmp.resolve("repo"));
GitWorktrees gitWorktrees = new GitWorktrees(tmp.resolve("wts").toString());
String branch = "cb-586-unreachable";
String wt = gitWorktrees.add(repo.toString(), branch, "HEAD");
Files.writeString(Path.of(wt).resolve("worker-draft.txt"), "work that exists nowhere else\n");
Optional<String> ref = gitWorktrees.snapshot(wt, branch, "snapshot with unreachable content");
assertTrue(ref.isPresent());
// Age floor 0 makes age a non-issue: only reachability can save it — and it must.
assertEquals(0, gitWorktrees.pruneWipRefs(repo.toString(), 0),
"the unreachable snapshot is the last copy and must not be pruned");
assertTrue(refExists(repo, "refs/wip/" + branch),
"an unreachable snapshot must survive the sweep");
}
/**
* CB-586, criterion 1 (the reachable half). A snapshot whose tree content IS already reachable
* from {@code main} and which is older than the age floor is pure duplication — the work is
* recovered — so it must be pruned.
*/
@Test
void aReachableSnapshotOlderThanTheFloorIsPruned(@TempDir Path tmp) throws Exception {
Path repo = initRepo(tmp.resolve("repo"));
GitWorktrees gitWorktrees = new GitWorktrees(tmp.resolve("wts").toString());
// A snapshot whose tree is exactly main's current tree: fully reachable from main.
String mainTree = revParse(repo, "main^{tree}");
String old = commitTree(repo, mainTree, revParse(repo, "HEAD"), "2020-01-01T00:00:00", "snapshot");
updateRef(repo, "refs/wip/recovered", old);
assertEquals(1, gitWorktrees.pruneWipRefs(repo.toString(), TimeUnit.HOURS.toMillis(24)),
"an old, main-reachable snapshot must be pruned");
assertFalse(refExists(repo, "refs/wip/recovered"),
"the reachable snapshot's ref must be gone after the sweep");
}
/**
* CB-586, the age floor. A snapshot whose content IS reachable from {@code main} but which is
* younger than the age floor must not be swept — a lead may still be looking at it.
*/
@Test
void aReachableButRecentSnapshotIsNotPruned(@TempDir Path tmp) throws Exception {
Path repo = initRepo(tmp.resolve("repo"));
GitWorktrees gitWorktrees = new GitWorktrees(tmp.resolve("wts").toString());
// Reachable from main, but committed "now" — a fresh snapshot. The 24h floor must protect it.
String mainTree = revParse(repo, "main^{tree}");
String fresh = commitTree(repo, mainTree, revParse(repo, "HEAD"),
"2038-01-01T00:00:00", "snapshot just taken");
updateRef(repo, "refs/wip/fresh", fresh);
assertEquals(0, gitWorktrees.pruneWipRefs(repo.toString(), TimeUnit.HOURS.toMillis(24)),
"a recent snapshot must be kept even when reachable");
assertTrue(refExists(repo, "refs/wip/fresh"),
"the recent reachable snapshot must survive the sweep");
}
/**
* CB-586, criterion 5. A fleet that has never snapshotted anything has no {@code refs/wip/*},
* so a sweep is a no-op and the census reports none — identical to before CB-586 existed.
*/
@Test
void aFleetWithNoSnapshotsPrunesNothingAndReportsNothing(@TempDir Path tmp) throws Exception {
Path repo = initRepo(tmp.resolve("repo"));
GitWorktrees gitWorktrees = new GitWorktrees(tmp.resolve("wts").toString());
assertEquals(0, gitWorktrees.pruneWipRefs(repo.toString(), 0),
"no snapshot refs means nothing to prune");
Worktrees.WipRefStats stats = gitWorktrees.wipRefs(repo.toString());
assertEquals(0, stats.count(), "a never-snapshotted fleet has zero refs/wip refs");
assertEquals(0L, stats.costBytes(), "a never-snapshotted fleet costs zero bytes");
}
/** CB-586, criterion 4: the census reports how many refs exist and roughly what they cost. */
@Test
void wipRefsReportsCountAndCost(@TempDir Path tmp) throws Exception {
Path repo = initRepo(tmp.resolve("repo"));
GitWorktrees gitWorktrees = new GitWorktrees(tmp.resolve("wts").toString());
String blob = blobOf(repo, "a recoverable snapshot's worth of content");
String tree = treeOf(repo, "snapshot.txt", blob);
updateRef(repo, "refs/wip/one", commitTree(repo, tree, revParse(repo, "HEAD"),
"2020-01-01T00:00:00", "snapshot"));
updateRef(repo, "refs/wip/two", commitTree(repo, tree, revParse(repo, "HEAD"),
"2020-01-02T00:00:00", "snapshot"));
Worktrees.WipRefStats stats = gitWorktrees.wipRefs(repo.toString());
assertEquals(2, stats.count(), "two snapshot refs are reported");
assertTrue(stats.costBytes() > 0, "the cost of the snapshots is a positive byte count");
}
}
@@ -10,7 +10,14 @@ import dev.ltms.bridged.herdr.AgentControl;
import dev.ltms.bridged.herdr.FakeHerdr;
import dev.ltms.bridged.herdr.WorkspaceControl;
import dev.ltms.bridged.member.ClaudeCodeLauncher;
import dev.ltms.bridged.msg.TestTurnTokens;
import dev.ltms.bridged.peer.Capability;
import dev.ltms.bridged.peer.CharterReceipt;
import dev.ltms.bridged.peer.MemberRole;
import dev.ltms.bridged.peer.PeerHandle;
import dev.ltms.bridged.peer.PeerLauncher;
import dev.ltms.bridged.peer.PeerUnreachableException;
import dev.ltms.bridged.peer.SpawnRequest;
import org.junit.jupiter.api.Test;
import org.slf4j.LoggerFactory;
@@ -43,6 +50,113 @@ class SessionManagerTest {
return sessionManager(herdr, clock, 0);
}
private SessionManager sessionManager(FakeHerdr herdr, Worktrees worktrees) {
return sessionManager(herdr, worktrees, System::nanoTime);
}
private SessionManager sessionManager(FakeHerdr herdr, Worktrees worktrees, LongSupplier clock) {
BridgedConfig.Profile cfg = new BridgedConfig.Profile(
"ltms-local", "http://gx00.gw:8000", "coder", null, "BRIDGED_WORKER_TOKEN",
List.of("ccs", "ltms-local"), "tab", "bridged-workers",
"worker: {profile} #{n}", null, null, null);
ClaudeCodeLauncher workers = new ClaudeCodeLauncher(new AgentControl(herdr), new WorkspaceControl(herdr),
new SubscriptionGuard(Set.of("gx00.gw")), Map.of(cfg.profile(), cfg), cfg.profile(), _ -> null);
return new SessionManager(workers, worktrees, clock);
}
/**
* CB-581: a {@link Worktrees} test double whose {@code hasUncommitted} and {@code remove} can
* be told to throw, so {@link SessionManager#release} can be exercised against exactly the
* failure {@code GitWorktrees} produces when {@code git status}/{@code git worktree remove}
* exits non-zero.
*/
private static final class RecordingWorktrees implements Worktrees {
private final List<String> removeCalls = new java.util.ArrayList<>();
private final List<String> snapshotCalls = new java.util.ArrayList<>();
private final java.util.Set<String> failRemoveFor = new java.util.HashSet<>();
private volatile boolean dirty = false;
private volatile RuntimeException hasUncommittedFailure;
private volatile RuntimeException snapshotFailure;
private final java.util.concurrent.atomic.AtomicLong snapshotSeq = new java.util.concurrent.atomic.AtomicLong();
RecordingWorktrees dirty(boolean dirty) {
this.dirty = dirty;
return this;
}
RecordingWorktrees failHasUncommittedWith(RuntimeException e) {
this.hasUncommittedFailure = e;
return this;
}
RecordingWorktrees failRemoveFor(String worktreePath) {
failRemoveFor.add(worktreePath);
return this;
}
RecordingWorktrees failSnapshotWith(RuntimeException e) {
this.snapshotFailure = e;
return this;
}
@Override
public String add(String repoRoot, String branch, String baseRef) {
return "/wt/" + branch.replace('/', '_');
}
@Override
public void remove(String repoRoot, String worktreePath) {
if (failRemoveFor.contains(worktreePath)) {
throw new WorktreeException("simulated remove failure for " + worktreePath);
}
removeCalls.add(worktreePath);
}
@Override
public boolean hasUncommitted(String worktreePath) {
if (hasUncommittedFailure != null) {
throw hasUncommittedFailure;
}
return dirty;
}
@Override
public void overlayParity(String repoRoot, String worktreePath, List<String> overlay) {
}
@Override
public String repoRoot(String cwd) {
return "/repo";
}
@Override
public java.util.Optional<String> snapshot(String worktreePath, String branch, String message) {
snapshotCalls.add(worktreePath);
if (snapshotFailure != null) {
throw snapshotFailure;
}
return java.util.Optional.of("wip" + snapshotSeq.incrementAndGet());
}
@Override
public WipRefStats wipRefs(String repoRoot) {
return new WipRefStats(0, 0L);
}
@Override
public int pruneWipRefs(String repoRoot, long minAgeMillis) {
return 0;
}
List<String> removeCalls() {
return List.copyOf(removeCalls);
}
List<String> snapshotCalls() {
return List.copyOf(snapshotCalls);
}
}
private SessionManager sessionManager(FakeHerdr herdr, LongSupplier clock, int contextCap) {
return sessionManager(herdr, clock, contextCap, false);
}
@@ -90,6 +204,24 @@ class SessionManagerTest {
assertEquals(2, sessions.roster().size(), "both sessions are registered");
}
@Test
void rosterViewExposesTheCharterReceiptButNeverTheCharterText() {
// The roster (bridge_list and GET /members both render through rosterView) must let a lead
// see which charter a member got, without ever carrying the charter prose itself (CB-571).
MemberSession s = new MemberSession("p1", "term1", "prof", MemberRole.DEV, "/cwd", null,
0, 0, 0, MemberSession.State.READY, null, null,
CharterReceipt.compose(MemberRole.DEV, "prof", "role charter", "role charter\n\nreply"), null);
Map<String, Object> view = SessionManager.rosterView(s, null);
assertEquals("fleet.charters.dev", view.get("charterSource"),
"the config key that supplied the role charter is reported");
assertEquals(CharterReceipt.digestOf("role charter\n\nreply"), view.get("charterSha256"),
"the digest of the exact composed charter bytes is reported");
assertFalse(view.values().toString().contains("role charter"),
"the roster row must not embed the charter text itself");
}
@Test
void aNullTerminalFromThePrimaryIsANoOpEvenWithSessionsRegistered() {
// The primary resolves to a Principal with no terminal, and BridgeMcp's context extractor
@@ -101,7 +233,7 @@ class SessionManagerTest {
assertDoesNotThrow(() -> sessions.asPresence().markPresent(null),
"the primary's null terminal must not blow up an unrelated tool call");
assertDoesNotThrow(() -> sessions.onDelivered(null));
assertDoesNotThrow(() -> sessions.onDelivered(null, TestTurnTokens.inert(null)));
assertDoesNotThrow(() -> sessions.onTurnComplete(null));
assertDoesNotThrow(() -> sessions.onTurnFailed(null));
@@ -121,7 +253,7 @@ class SessionManagerTest {
"MCP presence moves SPAWNING → READY");
assertTrue(sessions.asPresence().isPresent(terminal), "presence is also recorded");
sessions.onDelivered(terminal);
sessions.onDelivered(terminal, TestTurnTokens.inert(terminal));
assertEquals(MemberSession.State.BUSY, sessions.get(session.paneId()).orElseThrow().state(),
"delivery moves READY → BUSY");
@@ -153,7 +285,7 @@ class SessionManagerTest {
MemberSession session = sessions.acquire("ltms-local", null, "/caller", "term_primary");
String terminal = session.terminalId();
sessions.asPresence().markPresent(terminal);
sessions.onDelivered(terminal);
sessions.onDelivered(terminal, TestTurnTokens.inert(terminal));
sessions.onTurnFailed(terminal);
@@ -181,7 +313,7 @@ class SessionManagerTest {
MemberSession session = sessions.acquire("ltms-local", null, "/caller", "term_primary");
String terminal = session.terminalId();
sessions.asPresence().markPresent(terminal);
sessions.onDelivered(terminal);
sessions.onDelivered(terminal, TestTurnTokens.inert(terminal));
sessions.onTurnFailed(terminal);
@@ -264,7 +396,7 @@ class SessionManagerTest {
MemberSession session = sessions.acquire("ltms-local", null, "/caller", "term_primary");
String terminal = session.terminalId();
sessions.asPresence().markPresent(terminal);
sessions.onDelivered(terminal);
sessions.onDelivered(terminal, TestTurnTokens.inert(terminal));
clock[0] = 100;
assertEquals(0, sessions.reapIdle(10), "BUSY session past TTL is never reaped");
@@ -281,7 +413,7 @@ class SessionManagerTest {
MemberSession session = sessions.acquire("ltms-local", null, "/caller", "term_primary");
String terminal = session.terminalId();
sessions.asPresence().markPresent(terminal);
sessions.onDelivered(terminal);
sessions.onDelivered(terminal, TestTurnTokens.inert(terminal));
sessions.onTurnComplete(terminal);
clock[0] = 21;
@@ -299,7 +431,7 @@ class SessionManagerTest {
MemberSession busy = sessions.acquire("ltms-local", "/busy", "/caller", "owner2");
sessions.asPresence().markPresent(ready.terminalId());
sessions.asPresence().markPresent(busy.terminalId());
sessions.onDelivered(busy.terminalId());
sessions.onDelivered(busy.terminalId(), TestTurnTokens.inert(busy.terminalId()));
clock[0] = 50;
assertEquals(1, sessions.reapIdle(30), "only READY past TTL is reaped");
@@ -317,9 +449,9 @@ class SessionManagerTest {
String terminal = session.terminalId();
sessions.asPresence().markPresent(terminal);
sessions.onDelivered(terminal);
sessions.onDelivered(terminal, TestTurnTokens.inert(terminal));
sessions.onTurnComplete(terminal);
sessions.onDelivered(terminal);
sessions.onDelivered(terminal, TestTurnTokens.inert(terminal));
sessions.onTurnComplete(terminal);
MemberSession updated = sessions.get(session.paneId()).orElseThrow();
@@ -337,13 +469,13 @@ class SessionManagerTest {
String terminal = session.terminalId();
sessions.asPresence().markPresent(terminal);
sessions.onDelivered(terminal);
sessions.onDelivered(terminal, TestTurnTokens.inert(terminal));
sessions.onTurnComplete(terminal);
assertEquals(MemberSession.State.DONE,
sessions.get(session.paneId()).orElseThrow().state(),
"first turn completes without release");
sessions.onDelivered(terminal);
sessions.onDelivered(terminal, TestTurnTokens.inert(terminal));
sessions.onTurnComplete(terminal);
assertTrue(sessions.get(session.paneId()).isEmpty(), "session released after cap reached");
@@ -359,7 +491,7 @@ class SessionManagerTest {
MemberSession session = sessions.acquire("ltms-local", null, "/caller", "term_primary");
sessions.asPresence().markPresent(session.terminalId());
sessions.onDelivered(session.terminalId());
sessions.onDelivered(session.terminalId(), TestTurnTokens.inert(session.terminalId()));
assertTrue(sessions.onTurnCompleteWithPostAction(session.terminalId()));
MemberSession updated = sessions.get(session.paneId()).orElseThrow();
@@ -373,7 +505,7 @@ class SessionManagerTest {
SessionManager sessions = sessionManager(herdr, () -> 0L, 1, true);
MemberSession session = sessions.acquire("ltms-local", null, "/caller", "term_primary");
sessions.asPresence().markPresent(session.terminalId());
sessions.onDelivered(session.terminalId());
sessions.onDelivered(session.terminalId(), TestTurnTokens.inert(session.terminalId()));
assertFalse(sessions.hasPostTurnAction(session.terminalId()),
"a session at its cap will be released, not reset for reuse");
@@ -388,7 +520,7 @@ class SessionManagerTest {
SessionManager sessions = sessionManager(herdr, () -> 0L, 0, false);
MemberSession session = sessions.acquire("ltms-local", null, "/caller", "term_primary");
sessions.asPresence().markPresent(session.terminalId());
sessions.onDelivered(session.terminalId());
sessions.onDelivered(session.terminalId(), TestTurnTokens.inert(session.terminalId()));
sessions.onTurnComplete(session.terminalId());
@@ -406,7 +538,7 @@ class SessionManagerTest {
MemberSession busy = sessions.acquire("ltms-local", "/busy", "/caller", "ownerB");
sessions.asPresence().markPresent(ready.terminalId());
sessions.asPresence().markPresent(busy.terminalId());
sessions.onDelivered(busy.terminalId());
sessions.onDelivered(busy.terminalId(), TestTurnTokens.inert(busy.terminalId()));
sessions.drainAll(TimeUnit.MILLISECONDS.toNanos(100));
@@ -470,7 +602,7 @@ class SessionManagerTest {
FakeHerdr herdr = new FakeHerdr();
SessionManager sessions = sessionManager(herdr);
java.util.List<String> released = new java.util.concurrent.CopyOnWriteArrayList<>();
sessions.onRelease(released::add);
sessions.onRelease(detail -> released.add(detail.terminalId()));
MemberSession s = sessions.acquire("ltms-local", null, "/caller", null);
sessions.release(s.paneId());
@@ -484,7 +616,7 @@ class SessionManagerTest {
FakeHerdr herdr = new FakeHerdr();
SessionManager sessions = sessionManager(herdr);
java.util.List<String> released = new java.util.concurrent.CopyOnWriteArrayList<>();
sessions.onRelease(released::add);
sessions.onRelease(detail -> released.add(detail.terminalId()));
sessions.release("w9:p404"); // idempotent teardown of something already gone
@@ -504,4 +636,295 @@ class SessionManagerTest {
"a listener failure must never prevent the teardown it is reacting to");
assertTrue(sessions.get(s.paneId()).isEmpty(), "and the session is still deregistered");
}
// --- CB-581: a throw inside release() must not orphan the pane or abort reapIdle -----------
@Test
void releasePreservesWorktreeWhenDirtyCheckThrows() {
FakeHerdr herdr = new FakeHerdr();
RecordingWorktrees worktrees = new RecordingWorktrees();
SessionManager sessions = sessionManager(herdr, worktrees);
MemberSession s = sessions.acquire("ltms-local", null, "/caller/proj", null,
new WorktreeRequest("cb-581a", null));
worktrees.failHasUncommittedWith(new WorktreeException("git status exited 128"));
LoggerContext ctx = (LoggerContext) LoggerFactory.getILoggerFactory();
ch.qos.logback.classic.Logger sessionLog =
(ch.qos.logback.classic.Logger) LoggerFactory.getLogger(SessionManager.class);
ListAppender<ILoggingEvent> appender = new ListAppender<>();
appender.setContext(ctx);
appender.start();
sessionLog.addAppender(appender);
sessionLog.setLevel(Level.WARN);
try {
assertDoesNotThrow(() -> sessions.release(s.paneId()),
"a throwing dirty check must not abort the release");
assertTrue(worktrees.removeCalls().isEmpty(),
"the worktree is preserved when its dirty state cannot be determined");
String warn = appender.list.stream()
.filter(e -> e.getLevel().equals(Level.WARN))
.map(ILoggingEvent::getFormattedMessage)
.filter(m -> m.contains(s.worktree()))
.findFirst()
.orElse("no warn logged naming the worktree");
assertTrue(warn.contains(s.paneId()), "the WARN names the pane: " + warn);
assertTrue(warn.contains(s.terminalId()), "the WARN names the terminal: " + warn);
} finally {
sessionLog.detachAppender(appender);
}
}
@Test
void releaseStillStopsThePaneWhenDirtyCheckThrows() {
FakeHerdr herdr = new FakeHerdr();
RecordingWorktrees worktrees = new RecordingWorktrees();
SessionManager sessions = sessionManager(herdr, worktrees);
MemberSession s = sessions.acquire("ltms-local", null, "/caller/proj", null,
new WorktreeRequest("cb-581b", null));
worktrees.failHasUncommittedWith(new WorktreeException("git status exited 128"));
sessions.release(s.paneId());
assertEquals(1, paneCloseCallsFor(herdr, "w9:pRoot_1"),
"the pane is stopped exactly once even though the dirty check threw");
}
@Test
void releaseStillNotifiesTheListenerWhenDirtyCheckThrows() {
FakeHerdr herdr = new FakeHerdr();
RecordingWorktrees worktrees = new RecordingWorktrees();
SessionManager sessions = sessionManager(herdr, worktrees);
java.util.List<String> released = new java.util.concurrent.CopyOnWriteArrayList<>();
sessions.onRelease(detail -> released.add(detail.terminalId()));
MemberSession s = sessions.acquire("ltms-local", null, "/caller/proj", null,
new WorktreeRequest("cb-581c", null));
worktrees.failHasUncommittedWith(new WorktreeException("git status exited 128"));
sessions.release(s.paneId());
assertEquals(java.util.List.of(s.terminalId()), released,
"a blocked caller must still be told the terminal was released, even though the "
+ "dirty check threw");
}
@Test
void reapIdleSurvivesOneSessionThatFailsToRelease() {
long[] clock = {0};
FakeHerdr herdr = new FakeHerdr();
RecordingWorktrees worktrees = new RecordingWorktrees();
SessionManager sessions = sessionManager(herdr, worktrees, () -> clock[0]);
MemberSession a = sessions.acquire("ltms-local", null, "/caller/proj", null,
new WorktreeRequest("cb-581d", null));
MemberSession b = sessions.acquire("ltms-local", null, "/caller/proj", null,
new WorktreeRequest("cb-581e", null));
MemberSession c = sessions.acquire("ltms-local", null, "/caller/proj", null,
new WorktreeRequest("cb-581f", null));
sessions.asPresence().markPresent(a.terminalId());
sessions.asPresence().markPresent(b.terminalId());
sessions.asPresence().markPresent(c.terminalId());
// The middle session's worktree removal fails — release() propagates that, so this is the
// one call reapIdle's per-session guard must survive without skipping the rest of the pass.
worktrees.failRemoveFor(b.worktree());
LoggerContext ctx = (LoggerContext) LoggerFactory.getILoggerFactory();
ch.qos.logback.classic.Logger sessionLog =
(ch.qos.logback.classic.Logger) LoggerFactory.getLogger(SessionManager.class);
ListAppender<ILoggingEvent> appender = new ListAppender<>();
appender.setContext(ctx);
appender.start();
sessionLog.addAppender(appender);
sessionLog.setLevel(Level.WARN);
int reaped;
try {
clock[0] = 100;
reaped = sessions.reapIdle(10);
String warn = appender.list.stream()
.filter(e -> e.getLevel().equals(Level.WARN))
.map(ILoggingEvent::getFormattedMessage)
.filter(m -> m.contains(b.paneId()))
.findFirst()
.orElse("no reap-failure WARN logged");
assertTrue(warn.contains(b.terminalId()), "the WARN names the failed session's terminal: " + warn);
assertTrue(warn.contains(b.worktree()), "the WARN names the failed session's worktree: " + warn);
} finally {
sessionLog.detachAppender(appender);
}
assertEquals(2, reaped, "the middle session's failure is logged, not counted as reaped");
assertTrue(sessions.get(a.paneId()).isEmpty(), "the first session is still released");
assertTrue(sessions.get(c.paneId()).isEmpty(), "the third session is still released");
assertTrue(sessions.get(b.paneId()).isEmpty(),
"the middle session is still deregistered even though its worktree removal threw");
assertEquals(1, paneCloseCallsFor(herdr, "w9:pRoot_1"), "the first pane is stopped");
assertEquals(1, paneCloseCallsFor(herdr, "w9:pRoot_2"),
"the middle pane is still stopped even though its worktree removal failed");
assertEquals(1, paneCloseCallsFor(herdr, "w9:pRoot_3"), "the third pane is stopped");
}
@Test
void unchangedRegressionCleanCompletedReleaseStillRemovesTheWorktree() {
FakeHerdr herdr = new FakeHerdr();
RecordingWorktrees worktrees = new RecordingWorktrees();
SessionManager sessions = sessionManager(herdr, worktrees);
MemberSession s = sessions.acquire("ltms-local", null, "/caller/proj", null,
new WorktreeRequest("cb-581g", null));
sessions.release(s.paneId());
assertEquals(List.of(s.worktree()), worktrees.removeCalls(),
"COMPLETED release of a clean worktree still removes it");
}
@Test
void unchangedRegressionDirtyCompletedReleaseStillPreservesTheWorktree() {
FakeHerdr herdr = new FakeHerdr();
RecordingWorktrees worktrees = new RecordingWorktrees().dirty(true);
SessionManager sessions = sessionManager(herdr, worktrees);
MemberSession s = sessions.acquire("ltms-local", null, "/caller/proj", null,
new WorktreeRequest("cb-581h", null));
sessions.release(s.paneId());
assertTrue(worktrees.removeCalls().isEmpty(),
"COMPLETED release of a dirty worktree still preserves it");
}
@Test
void unchangedRegressionShutdownDrainStillPreservesTheWorktree() {
FakeHerdr herdr = new FakeHerdr();
RecordingWorktrees worktrees = new RecordingWorktrees();
SessionManager sessions = sessionManager(herdr, worktrees);
MemberSession s = sessions.acquire("ltms-local", null, "/caller/proj", null,
new WorktreeRequest("cb-581i", null));
sessions.asPresence().markPresent(s.terminalId());
sessions.drainAll(TimeUnit.MILLISECONDS.toNanos(100));
assertTrue(worktrees.removeCalls().isEmpty(), "SHUTDOWN drain still preserves the worktree");
}
// ── CB-584: agentSessionId recorded at acquire, and gated by Capability.SESSION_RESUME ─────
/**
* A minimal {@link PeerLauncher} whose capabilities exclude SESSION_RESUME — the "does not
* support it" case. {@link #requireResumeCapability} throws before ever reaching {@link #spawn},
* so every method beyond {@link #capabilitiesFor} is unreachable in these tests and left unimplemented.
*/
private static final class NoResumeLauncher implements PeerLauncher {
@Override
public Set<Capability> capabilities() {
return Set.of(Capability.WORKTREE); // deliberately no SESSION_RESUME
}
@Override
public Set<Capability> capabilitiesFor(String profileName) {
return capabilities();
}
@Override
public PeerHandle spawn(SpawnRequest req) {
throw new UnsupportedOperationException("not reachable — the capability check refuses first");
}
@Override
public Set<String> profiles() {
return Set.of("stub-profile");
}
@Override
public String defaultProfile() {
return "stub-profile";
}
@Override
public String effectiveCwd(SpawnRequest req) {
throw new UnsupportedOperationException("not reachable — the capability check refuses first");
}
@Override
public List<String> parityOverlay(String profileName) {
return List.of();
}
@Override
public List<?> list() {
return List.of();
}
@Override
public int reapOrphanWorkers() {
return 0;
}
@Override
public void stop(String id) {
}
@Override
public boolean clearContext(String id) {
return false;
}
}
@Test
void acquireRecordsTheAgentSessionIdFromTheHandle() {
FakeHerdr herdr = new FakeHerdr();
SessionManager sessions = sessionManager(herdr);
MemberSession s = sessions.acquire("ltms-local", MemberRole.DEV, null, null, null, null,
"my-session-name", null);
assertNotNull(s.agentSessionId(), "a spawn that asked for session identity gets one back");
Map<String, Object> view = SessionManager.rosterView(s, null);
assertEquals(s.agentSessionId(), view.get("agentSessionId"),
"the roster exposes the same id the session recorded");
}
@Test
void acquireWithNeitherSessionFieldLeavesAgentSessionIdNull() {
FakeHerdr herdr = new FakeHerdr();
SessionManager sessions = sessionManager(herdr);
MemberSession s = sessions.acquire("ltms-local", null, null, null);
assertNull(s.agentSessionId(), "no identity requested — unchanged from before CB-584");
assertFalse(SessionManager.rosterView(s, null).containsKey("agentSessionId"),
"a null id is omitted from the roster, like charterSha256 for a receipt-less session");
}
@Test
void resumeSessionIdResumesTheSameConversationOnASupportingAdapter() {
FakeHerdr herdr = new FakeHerdr();
SessionManager sessions = sessionManager(herdr);
MemberSession s = sessions.acquire("ltms-local", MemberRole.DEV, null, null, null, null,
null, "cb-resume-77");
assertEquals("cb-resume-77", s.agentSessionId(),
"a resume adopts the prior id as its own agentSessionId");
}
@Test
void resumeSessionIdWithoutAnExplicitProfileIsRefused() {
FakeHerdr herdr = new FakeHerdr();
SessionManager sessions = sessionManager(herdr);
IllegalArgumentException e = assertThrows(IllegalArgumentException.class, () ->
sessions.acquire(null, MemberRole.DEV, null, null, null, null, null, "cb-resume-1"));
assertTrue(e.getMessage().contains("explicit profile"), e.getMessage());
}
@Test
void resumeSessionIdOnAnAdapterWithoutTheCapabilityIsRefusedNamingIt() {
SessionManager sessions = new SessionManager(new NoResumeLauncher());
IllegalArgumentException e = assertThrows(IllegalArgumentException.class, () ->
sessions.acquire("stub-profile", MemberRole.DEV, null, null, null, null, null, "cb-resume-1"));
assertTrue(e.getMessage().contains("SESSION_RESUME"), e.getMessage());
assertTrue(e.getMessage().contains("stub-profile"), e.getMessage());
}
}
@@ -11,10 +11,12 @@ import org.junit.jupiter.api.Test;
import java.util.List;
import java.util.Map;
import java.util.Set;
import java.util.concurrent.TimeUnit;
import java.util.concurrent.atomic.AtomicLong;
import static org.junit.jupiter.api.Assertions.assertDoesNotThrow;
import static org.junit.jupiter.api.Assertions.assertEquals;
import static org.junit.jupiter.api.Assertions.assertFalse;
import static org.junit.jupiter.api.Assertions.assertTrue;
/**
@@ -90,6 +92,59 @@ class SessionReaperTest {
return ticks.get() >= n;
}
/**
* CB-586: the retention sweep must actually reach the seam when the loop runs.
*
* <p>The sweep's own tests call {@code SessionManager.sweepWipRefs} directly, which walks
* around the reaper's interval gate — and the gate is where it broke. {@code lastWipSweepNanos}
* started at {@code Long.MIN_VALUE}, so {@code now - lastWipSweepNanos} overflowed to a large
* negative number, the gate read that as "swept moments ago", and it returned <em>before</em>
* the assignment that would have fixed the field. The sweep never ran once, for the life of
* the process, and every direct-call test still passed.
*
* <p>So this asserts through the loop: spawn a worktree session (which is what tells the
* manager which repo holds {@code refs/wip/*}), start the reaper, and require a real prune call
* to arrive at the fake.
*/
@Test
void theLoopActuallyRunsTheWipRetentionSweep() throws InterruptedException {
FakeHerdr herdr = new FakeHerdr();
FakeWorktrees worktrees = new FakeWorktrees().withRepoRoot("/repo").withPrefix("/wt");
SessionManager sessions = new SessionManager(launcher(herdr), worktrees);
// Without a worktree session the repo is unknown and the sweep is a legitimate no-op, so
// this step is what makes the assertion below meaningful rather than vacuous.
sessions.acquire("ltms-local", null, "/caller/proj", null, new WorktreeRequest("cb-586", null));
assertTrue(worktrees.pruneCalls().isEmpty(), "nothing has swept before the reaper starts");
SessionReaper reaper = new SessionReaper(sessions, IDLE_TTL_SECONDS, SHORT_INTERVAL_MILLIS);
reaper.start();
try {
long deadline = System.currentTimeMillis() + 3000;
while (worktrees.pruneCalls().isEmpty() && System.currentTimeMillis() < deadline) {
Thread.sleep(10);
}
} finally {
reaper.stop();
}
assertFalse(worktrees.pruneCalls().isEmpty(),
"the reaper loop must run the refs/wip retention sweep; it never reached the seam");
assertEquals("/repo", worktrees.pruneCalls().getFirst().repoRoot(),
"the sweep must target the repo the fleet's worktrees came from");
assertEquals(TimeUnit.HOURS.toMillis(24), worktrees.pruneCalls().getFirst().minAgeMillis(),
"the 24h age floor is the safety rule and must reach the seam intact");
}
/** As {@link #launcher()}, on a caller-supplied herdr so the test can inspect it. */
private static ClaudeCodeLauncher launcher(FakeHerdr herdr) {
BridgedConfig.Profile cfg = new BridgedConfig.Profile(
"ltms-local", "http://gx00.gw:8000", "coder", null, "BRIDGED_WORKER_TOKEN",
List.of("ccs", "ltms-local"), "tab", "bridged-workers",
"worker: {profile} #{n}", null, null, null);
return new ClaudeCodeLauncher(new AgentControl(herdr), new WorkspaceControl(herdr),
new SubscriptionGuard(Set.of("gx00.gw")), Map.of(cfg.profile(), cfg), cfg.profile(), _ -> null);
}
@Test
void stopIsIdempotentAndSafeBeforeStart() {
SessionReaper reaper = reaper();
@@ -6,9 +6,15 @@ import dev.ltms.bridged.guard.SubscriptionGuard;
import dev.ltms.bridged.herdr.AgentControl;
import dev.ltms.bridged.herdr.FakeHerdr;
import dev.ltms.bridged.herdr.WorkspaceControl;
import ch.qos.logback.classic.Level;
import ch.qos.logback.classic.LoggerContext;
import ch.qos.logback.classic.spi.ILoggingEvent;
import ch.qos.logback.core.read.ListAppender;
import dev.ltms.bridged.member.ClaudeCodeLauncher;
import dev.ltms.bridged.msg.TestTurnTokens;
import dev.ltms.bridged.peer.MemberRole;
import org.junit.jupiter.api.Test;
import org.slf4j.LoggerFactory;
import java.util.List;
import java.util.Map;
@@ -178,6 +184,49 @@ class WorktreeSessionManagerTest {
assertTrue(sessions.get(paneId).isEmpty(), "released session is no longer retrievable");
}
/**
* CB-576. A normal {@code COMPLETED} release whose worktree holds uncommitted work must NOT
* remove it — {@code --force} would destroy the worker's only copy. The bridge cannot see
* uncommitted files, so the worktree is preserved and the release logged at WARN naming the
* path, the session, and the cause an operator needs to find the work.
*/
@Test
void releasePreservesDirtyWorktreeAndLogsWarn() {
FakeHerdr herdr = new FakeHerdr();
FakeWorktrees worktrees = new FakeWorktrees().withRepoRoot("/repo").withPrefix("/wt")
.withDirty(true);
SessionManager sessions = new SessionManager(workerService(herdr), worktrees);
MemberSession s = sessions.acquire("ltms-local", null, "/caller/proj", null,
new WorktreeRequest("cb-576", null));
LoggerContext ctx = (LoggerContext) LoggerFactory.getILoggerFactory();
ch.qos.logback.classic.Logger sessionLog =
(ch.qos.logback.classic.Logger) LoggerFactory.getLogger(SessionManager.class);
ListAppender<ILoggingEvent> appender = new ListAppender<>();
appender.setContext(ctx);
appender.start();
sessionLog.addAppender(appender);
sessionLog.setLevel(Level.WARN);
try {
sessions.release(s.paneId());
assertTrue(herdr.called("pane.close"), "release still tears the worker pane down");
assertTrue(worktrees.removeCalls().isEmpty(),
"a dirty worktree is never removed — it holds the only copy of the work");
String warn = appender.list.stream()
.filter(e -> e.getLevel().equals(Level.WARN))
.map(ILoggingEvent::getFormattedMessage)
.filter(m -> m.contains("dirty worktree"))
.findFirst()
.orElse("no dirty-release WARN logged");
assertTrue(warn.contains(s.worktree()), "the WARN names the worktree path: " + warn);
assertTrue(warn.contains(s.terminalId()), "the WARN names the session: " + warn);
assertTrue(warn.contains("COMPLETED"), "the WARN names the release cause: " + warn);
} finally {
sessionLog.detachAppender(appender);
}
}
@Test
void drainAllPreservesWorktreeOfIdleSession() {
FakeHerdr herdr = new FakeHerdr();
@@ -204,7 +253,7 @@ class WorktreeSessionManagerTest {
new WorktreeRequest("cb-544", null));
String terminal = s.terminalId();
sessions.asPresence().markPresent(terminal);
sessions.onDelivered(terminal); // BUSY, never completes → still BUSY when the timeout hits
sessions.onDelivered(terminal, TestTurnTokens.inert(terminal)); // BUSY, never completes → still BUSY when the timeout hits
sessions.drainAll(TimeUnit.MILLISECONDS.toNanos(100));
@@ -313,4 +362,141 @@ class WorktreeSessionManagerTest {
"the profile's configured cwd must reach repoRoot, not be ignored");
}
// --- CB-578 stage C: snapshot a dirty worktree into git before the preserve-or-remove decision ---
@Test
void releaseOfACleanWorktreeNeverSnapshots() {
FakeHerdr herdr = new FakeHerdr();
FakeWorktrees worktrees = new FakeWorktrees().withRepoRoot("/repo").withPrefix("/wt");
SessionManager sessions = new SessionManager(workerService(herdr), worktrees);
MemberSession s = sessions.acquire("ltms-local", null, "/caller/proj", null,
new WorktreeRequest("cb-578c-a", null));
sessions.release(s.paneId());
assertTrue(worktrees.snapshotCalls().isEmpty(),
"acceptance criterion 4: a clean release must create no snapshot ref or commit");
assertEquals(1, worktrees.removeCalls().size(), "a clean worktree is still removed as before");
}
@Test
void releaseOfADirtyWorktreeSnapshotsBeforePreserving() {
FakeHerdr herdr = new FakeHerdr();
FakeWorktrees worktrees = new FakeWorktrees().withRepoRoot("/repo").withPrefix("/wt")
.withDirty(true);
SessionManager sessions = new SessionManager(workerService(herdr), worktrees);
MemberSession s = sessions.acquire("ltms-local", null, "/caller/proj", null,
new WorktreeRequest("cb-578c-b", null));
sessions.release(s.paneId());
assertEquals(1, worktrees.snapshotCalls().size(),
"acceptance criterion 1: a dirty release snapshots the worktree exactly once");
FakeWorktrees.SnapshotCall call = worktrees.lastSnapshot();
assertEquals(s.worktree(), call.worktreePath());
assertEquals(s.branch(), call.branch());
assertTrue(worktrees.removeCalls().isEmpty(), "the dirty worktree is still preserved, not removed");
}
@Test
void releaseNotifiesTheListenerWithWorktreeBranchAndSnapshotRef() {
FakeHerdr herdr = new FakeHerdr();
FakeWorktrees worktrees = new FakeWorktrees().withRepoRoot("/repo").withPrefix("/wt")
.withDirty(true);
SessionManager sessions = new SessionManager(workerService(herdr), worktrees);
java.util.List<SessionManager.ReleaseDetail> released = new java.util.concurrent.CopyOnWriteArrayList<>();
sessions.onRelease(released::add);
MemberSession s = sessions.acquire("ltms-local", null, "/caller/proj", null,
new WorktreeRequest("cb-578c-c", null));
sessions.release(s.paneId());
assertEquals(1, released.size());
SessionManager.ReleaseDetail detail = released.getFirst();
assertEquals(s.terminalId(), detail.terminalId());
assertEquals(s.worktree(), detail.worktreePath(),
"acceptance criterion 6: a failed ticket's detail must carry the worktree path");
assertEquals(s.branch(), detail.branch(),
"acceptance criterion 6: a failed ticket's detail must carry the branch");
assertEquals("wip1", detail.snapshotRef(),
"acceptance criterion 6: a failed ticket's detail must carry the snapshot ref");
}
@Test
void releaseNotifiesTheListenerWithTheAgentSessionId() {
// CB-584 (issue #65 criterion 5): a failed ticket's detail must also name the agent
// session, alongside worktree/branch/snapshot, so a lead can resume the conversation
// rather than only re-dispatch a fresh member onto the same files.
FakeHerdr herdr = new FakeHerdr();
FakeWorktrees worktrees = new FakeWorktrees().withRepoRoot("/repo").withPrefix("/wt")
.withDirty(true);
SessionManager sessions = new SessionManager(workerService(herdr), worktrees);
java.util.List<SessionManager.ReleaseDetail> released = new java.util.concurrent.CopyOnWriteArrayList<>();
sessions.onRelease(released::add);
MemberSession s = sessions.acquire("ltms-local", MemberRole.DEV, null, "/caller/proj", null,
new WorktreeRequest("cb-584-e", null), "cb-584-session", null);
assertNotNull(s.agentSessionId(), "a named session must mint an agent session id to assert on");
sessions.release(s.paneId());
assertEquals(1, released.size());
SessionManager.ReleaseDetail detail = released.getFirst();
assertEquals(s.agentSessionId(), detail.agentSessionId());
}
@Test
void aFailingSnapshotStillPreservesTheWorktreeStopsThePaneAndNotifies() {
FakeHerdr herdr = new FakeHerdr();
FakeWorktrees worktrees = new FakeWorktrees().withRepoRoot("/repo").withPrefix("/wt")
.withDirty(true).failSnapshot("git commit-tree exited 128");
SessionManager sessions = new SessionManager(workerService(herdr), worktrees);
java.util.List<SessionManager.ReleaseDetail> released = new java.util.concurrent.CopyOnWriteArrayList<>();
sessions.onRelease(released::add);
MemberSession s = sessions.acquire("ltms-local", null, "/caller/proj", null,
new WorktreeRequest("cb-578c-d", null));
assertDoesNotThrow(() -> sessions.release(s.paneId()),
"acceptance criterion 5: a failing snapshot must not propagate out of release()");
assertTrue(herdr.called("pane.close"), "acceptance criterion 5: the pane must still stop");
assertTrue(worktrees.removeCalls().isEmpty(),
"acceptance criterion 5: the worktree must still be preserved when the snapshot fails");
assertEquals(1, released.size(), "acceptance criterion 5: the listener must still be notified");
assertNull(released.getFirst().snapshotRef(),
"a failed snapshot leaves no ref to report");
}
/**
* CB-586, criteria 4 and 1. Once a worktree session is spawned the repo is known, so the
* operator census and the retention sweep delegate to that repo's {@code refs/wip/*}.
*/
@Test
void wipRefsAndSweepDelegateToTheFleetRepoOnceKnown() {
FakeHerdr herdr = new FakeHerdr();
FakeWorktrees worktrees = new FakeWorktrees().withRepoRoot("/repo").withPrefix("/wt")
.withWipRefs(new Worktrees.WipRefStats(3, 42L)).withPruneResult(2);
SessionManager sessions = new SessionManager(workerService(herdr), worktrees);
assertTrue(sessions.wipRefs().isEmpty(),
"no worktree spawned yet means no repo is known and nothing to report");
assertEquals(0, sessions.sweepWipRefs(TimeUnit.HOURS.toMillis(24)),
"no worktree spawned yet means the sweep is a no-op");
sessions.acquire("ltms-local", null, "/caller/proj", null,
new WorktreeRequest("cb-586", null));
Worktrees.WipRefStats stats = sessions.wipRefs().orElseThrow();
assertEquals(3, stats.count(), "the census comes from the fleet repo");
assertEquals(42L, stats.costBytes(), "the cost comes from the fleet repo");
assertEquals(2, sessions.sweepWipRefs(TimeUnit.HOURS.toMillis(24)),
"the sweep runs against the fleet repo");
List<FakeWorktrees.PruneCall> prunes = worktrees.pruneCalls();
assertEquals(1, prunes.size(), "the no-op short-circuits before reaching the seam, so only "
+ "the repo-known sweep issues a call");
assertEquals("/repo", prunes.getFirst().repoRoot(), "the sweep targets the fleet repo");
assertEquals(TimeUnit.HOURS.toMillis(24), prunes.getFirst().minAgeMillis(),
"the caller's age floor is passed through");
}
}
+4 -1
View File
@@ -16,6 +16,9 @@
AuditLogTest, which attaches its own ListAppender and asserts on emitted records.
-->
<!-- Keep the test logger behaviour aligned with the production cancellation filter. -->
<turboFilter class="dev.ltms.bridged.logging.McpCancelledNotificationFilter"/>
<appender name="STDOUT" class="ch.qos.logback.core.ConsoleAppender">
<encoder>
<pattern>%d{HH:mm:ss.SSS} [%thread] %-5level %logger{36} - %msg%n</pattern>
@@ -34,4 +37,4 @@
<appender-ref ref="STDOUT"/>
</root>
</configuration>
</configuration>
+62 -15
View File
@@ -1,7 +1,7 @@
<?xml version="1.0" encoding="UTF-8"?>
<!DOCTYPE plist PUBLIC "-//Apple//DTD PLIST 1.0//EN" "http://www.apple.com/DTDs/PropertyList-1.0.dtd">
<!--
CB-504 — launchd agent for bridged (macOS).
CB-504 / CB-594 — launchd agent for bridged (macOS).
This is the real supervision target today: the dogfooded daemon runs on macOS, where there is
no systemd. A systemd unit ships alongside (deploy/bridged.service) for the Linux gateways
@@ -9,14 +9,33 @@
Install:
cp deploy/dev.ltms.bridged.plist ~/Library/LaunchAgents/
# edit the paths + JAVA_HOME below to match this host, then:
launchctl load -w ~/Library/LaunchAgents/dev.ltms.bridged.plist
launchctl list | grep bridged
The paths below are already filled in for this host (resolved 2026-08-16 from
`/usr/libexec/java_home`... except that reported the system Applet-plugin JVM, not the jenv-
managed JDK 25 actually used to build/run bridged, so JAVA_HOME here is the real one:
`JENV_VERSION=25.0.3 java -XshowSettings:properties -version 2>&1 | grep java.home`; `which mvn`;
`echo $HOME`). If this file is copied to a different host, re-resolve all three paths and check
no placeholder path is left behind; scripts/redeploy-bridged.sh's check mode does not (and
cannot) check this file for you.
CB-594 — launchd cannot run a login shell (see the PATH comment on EnvironmentVariables below,
and scripts/bridged-launchd-wrapper.sh for the fix): ProgramArguments below execs THAT wrapper,
not java directly, so WORKER_GITEA_TOKEN and AI_GATEWAY_TOKEN still get sourced from
${SHARED_ENV}/tools/secrets.sh even though launchd itself never sources anything.
Note on ordering: launchd has no "start after herdr" primitive for user agents, and neither
does systemd in a way that survives a socket appearing late. bridged retries the herdr socket
on startup instead, so an agent that comes up before herdr converges rather than dying — that
retry is the actual fix; KeepAlive below is the backstop.
CB-594 — KeepAlive vs. scripts/redeploy-bridged.sh: a bare SIGTERM makes this JVM exit 143 even
with its shutdown hook running to completion (measured, see the CB-594 report), which
SuccessfulExit:false below reads as a crash and races to restart the OLD jar. The redeploy
script now detects a loaded agent and uses `launchctl unload`/`load` instead of a raw kill, so
only one supervisor ever touches the process at a time — read that script's own output on a
redeploy for the confirmation.
-->
<plist version="1.0">
<dict>
@@ -25,22 +44,23 @@
<key>ProgramArguments</key>
<array>
<string>/Users/CHANGEME/Tool/jdk-25.0.2.jdk/Contents/Home/bin/java</string>
<string>/Users/dai.ha/LTMS/claude-bridge/scripts/bridged-launchd-wrapper.sh</string>
<string>/Users/dai.ha/Softwares/jdks/jdk-25.0.3.jdk/Contents/Home/bin/java</string>
<string>-jar</string>
<string>/Users/CHANGEME/src/claude-bridge/bridged/target/bridged.jar</string>
<string>/Users/dai.ha/LTMS/claude-bridge/bridged/target/bridged.jar</string>
<string>bridged.yaml</string>
</array>
<!-- Config path in ProgramArguments is relative, so the working directory must be the module. -->
<key>WorkingDirectory</key>
<string>/Users/CHANGEME/src/claude-bridge/bridged</string>
<string>/Users/dai.ha/LTMS/claude-bridge/bridged</string>
<key>EnvironmentVariables</key>
<dict>
<key>JAVA_HOME</key>
<string>/Users/CHANGEME/Tool/jdk-25.0.2.jdk/Contents/Home</string>
<string>/Users/dai.ha/Softwares/jdks/jdk-25.0.3.jdk/Contents/Home</string>
<key>HERDR_SOCKET_PATH</key>
<string>/Users/CHANGEME/.config/herdr/herdr.sock</string>
<string>/Users/dai.ha/.config/herdr/herdr.sock</string>
<!--
PATH matters more than it looks (CB-511): bridged propagates its own PATH to every worker
it spawns, so this line decides whether the fleet can run a build at all. launchd does NOT
@@ -48,19 +68,39 @@
bare /usr/bin:/bin and no JDK or Maven. Keep the toolchain entries first.
-->
<key>PATH</key>
<string>/Users/CHANGEME/Tool/jdk-25.0.2.jdk/Contents/Home/bin:/Users/CHANGEME/Tool/apache-maven-3.9.16/bin:/opt/homebrew/bin:/usr/local/bin:/usr/bin:/bin:/usr/sbin:/sbin</string>
<string>/Users/dai.ha/Softwares/jdks/jdk-25.0.3.jdk/Contents/Home/bin:/Users/dai.ha/Softwares/apache-maven/bin:/opt/homebrew/bin:/usr/local/bin:/usr/bin:/bin:/usr/sbin:/sbin</string>
<!--
Worker/API tokens are NOT set here: this file is committed. Export them from a private
launchd override or a wrapper script. bridged reads the API token from the env var named
by auth.tokenEnv (default BRIDGED_API_TOKEN) and only in auth.mode: token.
Worker/API tokens are NOT set here: this file is committed. CB-594 —
scripts/bridged-launchd-wrapper.sh (named in ProgramArguments above) is what supplies
them, by execing a login shell that sources ${SHARED_ENV}/tools/secrets.sh before the
daemon itself starts. bridged also reads the API token from the env var named by
auth.tokenEnv (default BRIDGED_API_TOKEN) and only in auth.mode: token — the wrapper
covers that one too, since it is the same login shell.
-->
</dict>
<key>RunAtLoad</key>
<true/>
<!-- Restart on crash, but not in a tight loop if the config is bad (bridged fails fast on a
non-loopback bind without token auth — that is a config error, not a transient one). -->
<!--
CB-600 — read this before assuming ThrottleInterval bounds anything. It paces restarts to at
most one per 10s; it does NOT cap how many times launchd retries. If bridged fails fast on
every start — a bad bridged.yaml, for example auth.mode: token with the token env var unset,
which throws in main() before the daemon ever binds a port — launchd restarts it forever,
once every 10s, until a human intervenes. LaunchAgents have no "give up after N attempts"
primitive, so this is not something a config change here can fix.
That loop stops only two ways: (1) `launchctl unload -w ~/Library/LaunchAgents/dev.ltms.bridged.plist`,
or (2) the underlying cause gets fixed, so the process starts successfully and stays up (no
more exits to restart). scripts/redeploy-bridged.sh does not add a third way — it does not
make bridged self-disable on a config error, on purpose: a fail-fast exit path that
sometimes decides "this is unrecoverable, stop trying" is one more thing that can misfire,
and a wrongly self-disabled daemon needs the exact same manual `launchctl load -w` recovery
this comment already names — so it buys nothing an operator watching for the crash loop
doesn't already have, at the cost of a new way to be silently down. Watch for it with
`launchctl list dev.ltms.bridged` (a high restart count) or by tailing bridged.out for the
same startup error repeating every ~10s.
-->
<key>KeepAlive</key>
<dict>
<key>SuccessfulExit</key>
@@ -69,10 +109,17 @@
<key>ThrottleInterval</key>
<integer>10</integer>
<!--
CB-594 — same file scripts/redeploy-bridged.sh already tails ($BRIDGED/bridged.out), and both
streams point at it, not two separate log files: the script's fresh-line / ERROR-count checks
after a restart read this one path regardless of whether launchd or the script started the
process, and a stdout/stderr split would make half of what happens during a launchd-driven
restart invisible to it.
-->
<key>StandardOutPath</key>
<string>/Users/CHANGEME/src/claude-bridge/bridged/logs/bridged.out.log</string>
<string>/Users/dai.ha/LTMS/claude-bridge/bridged/bridged.out</string>
<key>StandardErrorPath</key>
<string>/Users/CHANGEME/src/claude-bridge/bridged/logs/bridged.err.log</string>
<string>/Users/dai.ha/LTMS/claude-bridge/bridged/bridged.out</string>
<key>ProcessType</key>
<string>Background</string>
+467
View File
@@ -0,0 +1,467 @@
# CB-591 — move the fleet onto the LLM and MCP gateway
**Status: DONE — the fleet is on the gateway as of 2026-08-15.** `local` runs on `/anthropic` and
`gx` on `/v1`, both at `weight: 100`; `local-direct` stays at `weight: 0` as the escape hatch. Getting
here took a revert and two upstream fixes — see §7.1, which is the useful part of this document. One
risk is **accepted rather than solved**: a stream cut by any mid-response timer arrives as HTTP 200
with no terminator, and our third-party members cannot detect it (§7.2).
· **Upstream:** [systems/vms wiki → LLM and MCP Gateway](https://git.ltms.dev/systems/vms/wiki/LLM-and-MCP-Gateway)
· **Upstream issue:** [systems/vms#31](https://git.ltms.dev/systems/vms/issues/31)
The gateway went live on 2026-08-15 and replaced Bifrost. This plan says what that means for a
**member definition** in `bridged.yaml`, because that is the part of this repo the change actually
touches.
---
## 1. What changed upstream
One front door for every LLM and MCP client: `https://llm.ltms.dev`, one token per consumer.
| Surface | URL |
|---|---|
| OpenAI chat | `https://llm.ltms.dev/v1/chat/completions` |
| OpenAI models | `https://llm.ltms.dev/v1/models` |
| **Anthropic messages** | `https://llm.ltms.dev/anthropic/v1/messages` |
| MCP, all servers multiplexed | `https://llm.ltms.dev/mcp` |
Anything outside that list returns **404 before any token is checked**, on purpose — the gateway must
never become a blanket proxy.
The model backend is unchanged: GX10 vLLM at `10.10.10.26:8000` (`gx00.gw`), model name exactly
`deepseek-v4-flash`. The direct LAN path stays open on purpose as an escape hatch.
---
## 2. Where claude-bridge sits today
We do **not** use the gateway. The `local` profile talks straight to the vLLM:
```yaml
local:
kind: claude-code
baseUrl: http://gx00.gw:8000 # direct vLLM — no auth, LAN only
model: deepseek-v4-flash
configDir: /Users/dai.ha/.ccs/instances/gx10
```
Three facts about our side that decide the shape of this work:
1. **`baseUrl` becomes `ANTHROPIC_BASE_URL`** in the member's environment, and `tokenEnv` becomes
`ANTHROPIC_AUTH_TOKEN` (the value is read from a host env var and never stored in config).
`local` sets no `tokenEnv` today, because a direct vLLM needs no token.
2. **`SubscriptionGuard` refuses any host not on an allowlist**, and that allowlist is
`guard.offSubscriptionHosts: [gx00.gw]`. It is built once in `Bridged.java:93` and handed to the
launcher, so **it is a restart-required key**, not a hot one. Changing `baseUrl` without changing
this makes every `local` spawn throw.
3. **The wiki names us as a blocker.** Under *Not done yet*: retiring the shared `legacy` token is
blocked because "kb, brain, **claude-bridge** and the workstation still share it. Each needs its
own consumer first."
Context7 is mounted twice today, both times straight at `https://ct7.ltms.dev/mcp` — once in
`.mcp.json` (the primary) and once in `opencode.json` (the `sol` and `terra` members).
```mermaid
flowchart LR
subgraph now["Today"]
M1["local member<br/>claude-code"] -->|"ANTHROPIC_BASE_URL"| V1["vLLM gx00.gw:8000<br/>no auth, LAN only"]
M2["sol / terra<br/>opencode"] --> CT1["ct7.ltms.dev/mcp"]
P1["primary"] --> CT1
end
subgraph after["Proposed"]
M3["local member"] -->|"ANTHROPIC_BASE_URL<br/>+ ANTHROPIC_AUTH_TOKEN"| G["llm.ltms.dev/anthropic<br/>consumer: claude-bridge"]
G --> V2["vLLM gx00.gw:8000"]
M4["local-direct<br/>weight 0, escape hatch"] --> V2
end
```
*The member definition is the only thing that moves. The model behind it does not.*
---
## 3. The member definition change
The gateway serves an Anthropic surface *and* an OpenAI surface, so **both member kinds can point at
it**. That is the main opportunity here, and it is bigger than the `local` profile alone.
### 3a. `local` — claude-code, on `/anthropic`
| Key | Today | After | Note |
|---|---|---|---|
| `baseUrl` | `http://gx00.gw:8000` | `https://llm.ltms.dev/anthropic` | see the schema warning below |
| `tokenEnv` | *(unset)* | `AI_GATEWAY_TOKEN` | new consumer token, `llmk-claude-bridge-<32 hex>` |
| `model` | `deepseek-v4-flash` | unchanged | must stay **exact**; a regex match returns an empty `/v1/models` while completions keep working |
| `guard.offSubscriptionHosts` | `[gx00.gw]` | `[gx00.gw, llm.ltms.dev]` | **restart required** |
### 3b. A new opencode profile on `/v1` — no code needed
`OpenCodeLauncher` already supports a pinned OpenAI-compatible endpoint (CB-508). Given `baseUrl` it
writes a custom provider block into the worker's opencode config:
- `baseUrl` → `options.baseURL`. `openAiBaseUrl` appends `/v1` to a bare host, and takes a URL that
already has a path **as-is** — so `https://llm.ltms.dev/v1` works unchanged.
- `tokenEnv` → `options.apiKey` (falls back to a placeholder when unset, since a local vLLM ignores it).
- `model:` **must** be `<provider>/<model>` when `baseUrl` is set — a bare name is rejected loudly
rather than silently falling back to opencode's default gateway.
So the profile is pure config:
```yaml
gx:
kind: opencode
baseUrl: https://llm.ltms.dev/v1
tokenEnv: AI_GATEWAY_TOKEN
model: gx/deepseek-v4-flash # provider id is ours to choose; the half after / is the model
argv: ["opencode"]
mcpUrl: http://127.0.0.1:8765/mcp
gitTokenEnv: WORKER_GITEA_TOKEN
weight: 100 # same tier as `local` — free
maxLoad: 2
# deliberately NO credentialId — this is our own box, not the shared OpenAI account
```
**Why this matters more than it looks.** Today every opencode member is `sol` or `terra`, and those
are two models on **one** OpenAI account sharing `credentialId: openai-shared` — so an exhaustion on
either locks out both, and half the fleet's opencode capacity dies at once. A gateway-backed opencode
profile is free, is not on that credential, and therefore is not in that quarantine pair. It removes
a single point of failure rather than just adding capacity.
**Note the asymmetry, it is deliberate:** `SubscriptionGuard` does not apply to opencode at all — the
guard exists to stop a *Claude* worker borrowing the operator's subscription, and opencode reads its
own provider credentials. So 3b needs **no allowlist change**; only 3a does.
**Both still need a restart, for a different reason.** `tokenEnv` is resolved by
`HerdrPeerLauncher.resolveEnv` → `env.apply(name)`, which reads the **daemon's own process
environment**. The running `bridged` inherited its environment when it started, so a variable added to
`secrets.sh` afterwards is simply not there — the launcher would inject an empty token and the
gateway would answer 401. This is the same failure as trap 1 in `scripts/redeploy-bridged.sh`
(`WORKER_GITEA_TOKEN`), and it has the same fix: **restart from a login shell**, and use
`scripts/redeploy-bridged.sh --check` to confirm the name resolves before restarting anything.
### 3c. What this does to `ccs`
Once a profile carries `baseUrl`, `tokenEnv` and `model` itself, the ccs instance stops being what
routes a member. Be precise about what is left, though: `configDir` still supplies **folder trust**
and `settings.json`, and dropping it is what produced the trust dialog and the wrong-model error
recorded in `bridged.yaml`. So ccs goes from *deciding where the tokens go* to *holding client-side
state*. Less load-bearing, not removable.
### Why `/anthropic` and never `/v1/chat/completions`
The gateway declares its Anthropic backend as `schema.name: Anthropic`, which means **no
translation** — streaming, tool use and thinking blocks pass through exactly as they do against vLLM
directly.
Declared as `OpenAI`, Envoy's translator looks for a `thinking_blocks` field that our vLLM does not
send (it sends `reasoning_content`), and **every thinking delta disappears silently**. Claude Code
speaks the Anthropic protocol, so `/anthropic` is both correct and the only safe choice.
This is the exact failure shape this repo keeps hitting: it compiles, it answers, it looks healthy,
and a capability is quietly off. Treat it as a `silent-default` risk, not a config preference.
**Open question for 3b — ANSWERED, 2026-08-15.** The worry was that the OpenAI surface might drop
reasoning the way the wiki documents for a mis-declared Anthropic backend. It does not. Checked at
the API before any profile was switched:
| surface | request | result |
|---|---|---|
| `/anthropic/v1/messages` | `deepseek-v4-flash`, 64 tokens | 200, response carries a real `"type":"thinking"` block |
| `/v1/chat/completions` | same | 200, message carries a populated `reasoning_content` (and a `reasoning` field) |
| `/v1/models` | — | 200, exactly `["deepseek-v4-flash"]` — the exact-name trap is clear |
| `/v1/models`, **no token** | — | **401** — Caddy is gating, as designed |
So reasoning survives on **both** surfaces, and the `/anthropic` choice for `local` is about protocol
correctness rather than a repair for a known loss. The last row matters on its own: the wiki warns
the gateway's own `SecurityPolicy` fails open, so it is worth knowing the proxy in front really does
refuse an unauthenticated request here.
---
## 4. Decisions
### D1 — switch, but keep the direct path as an explicit profile · **recommended**
Switching buys four things we do not have:
- **Free opencode capacity, off the shared credential.** The largest single win. See §3b — it retires
a real single point of failure, not just a cost line.
- **Per-consumer usage figures.** The cockpit counts requests per consumer. That is the first real
measurement of what the fleet consumes, and it feeds [CB-589](https://git.ltms.dev/lms/claude-bridge/issues/74) Gap 2 directly.
- **Our own revocable token.** One consumer to revoke if a worker ever leaks it, instead of a shared
`legacy` token used by four systems.
- **It works off-LAN.** `gx00.gw` resolves on the LAN only.
The cost is honest and worth stating: we add a TLS edge, an auth proxy and a gateway to the path of
every member spawn. The wiki keeps the direct route open precisely because "if the gateway breaks,
nothing that matters is blocked."
So keep it. Add a second profile `local-direct` pointing at `http://gx00.gw:8000` with **`weight: 0`**
— never auto-selected, still spawnable with an explicit `bridge_spawn{profile: "local-direct"}`.
That is exactly what CB-554 made `weight: 0` mean, and it turns the escape hatch into something the
lead can actually reach during an incident.
### D2 — do members also mount the gateway's `/mcp`? · **OPEN, operator's call**
Not a detail. `CLAUDE.md` states in two places that a member mounts **only** the bridge MCP, and a
worker's honesty rule leans on it ("never claim the result of a check you had no way to run").
- **Keep bridge-only.** The invariant stays true and simple. Workers stay cheap and narrow.
- **Add the gateway MCP.** Implementers get context7 documentation lookups, which is genuinely useful
for library work. But `mcpUrl` in `BridgedConfig.Profile` is a **single `String`**, so a
claude-code member can mount exactly one MCP — this needs a code change, not a config edit.
Note the invariant is **already inaccurate**: `opencode.json` gives `sol` and `terra` both context7
and gitea. So the choice is really "make the rule true" or "make the rule match reality". Either is
defensible; picking one is not mine to do.
### D3 — token scope
One consumer, `claude-bridge`, its token in `${SHARED_ENV}/tools/secrets.sh` as `AI_GATEWAY_TOKEN`,
referenced by name only. Never the literal value in `bridged.yaml` — `tokenEnv` exists for this.
---
## 5. Units of work
```mermaid
flowchart TB
U1["U1 · consumer token<br/>issue via cockpit, add to secrets.sh"]
U2["U2 · profile + guard<br/>bridged.yaml, restart"]
U3["U3 · verify live<br/>spawn, prove thinking survives"]
U4["U4 · context7 via gateway<br/>.mcp.json + opencode.json"]
U5["U5 · docs<br/>CLAUDE.md, wiki 11-Features"]
U1 --> U2 --> U3
U4 --> U5
U3 --> U5
```
| # | Scope | Who | Why |
|---|---|---|---|
| U1 | Issue the `claude-bridge` consumer at `auth.ltms.dev`; store as `AI_GATEWAY_TOKEN` | **operator** | touches secrets and a host we do not own |
| U2a | New `gx` opencode profile on `/v1` — pure config, no guard change | **lead** | `bridged.yaml` is gitignored, so a worker cannot see or edit it |
| U2b | `local` → `/anthropic`; add `local-direct` weight 0; add `llm.ltms.dev` to the guard allowlist | **lead** | same |
| U2c | One restart from a **login shell**, after U2a and U2b | **lead** | picks up `AI_GATEWAY_TOKEN` into the daemon env *and* the guard allowlist, in one stop |
| U3 | Live spawn on both new profiles; confirm reasoning survives on each surface | **lead** | needs real spawns and the running daemon |
| U4 | Point `.mcp.json` and `opencode.json` context7 at the gateway `/mcp`; rename pinned tools | delegatable | tracked files, self-contained |
| U5 | Fix the "members mount only the bridge" claim; add a `wiki/11-Features.md` entry | delegatable | writing, clear criteria |
U1 blocks U2a, U2b and U3. U4 and U5 do not depend on it.
**Write U2a and U2b, then restart once (U2c), then verify `gx` before `local`.** Since both profiles
need the same restart there is no reason to do two, but there is still a reason to *verify* in order:
`gx` exercises the token and the gateway with no guard involved, so if it fails the cause is upstream.
`local` adds the guard allowlist on top, so a failure there points at our config instead. Testing them
in that order separates the two causes instead of confusing them.
> **U1 status, 2026-08-15:** the operator issued the consumer and exported it as `AI_GATEWAY_TOKEN`
> (one key for every agent and MCP client behind `llm.ltms.dev`). Confirmed: it resolves in a login
> shell, is 48 characters and carries the documented `llmk-` prefix. The value was never printed.
---
## 6. Traps carried over from the wiki
Each of these cost someone real debugging time upstream. They apply to us.
1. **Rotating a token restarts the auth proxy, which drops in-flight streaming responses.** For us
that means rotating `AI_GATEWAY_TOKEN` kills every live member mid-turn, and an async ticket's
report goes with it. This is the same rule as a daemon redeploy: **drain the fleet first**
(`bridge_list` → `bridge_poll` anything wanted → `bridge_stop`), then rotate.
2. **The gateway's own `SecurityPolicy` fails open.** Standalone `aigw run` accepts it and silently
ignores it — an unauthenticated request returned **200**. Auth is the Caddy proxy in front, and
nothing else. Never reason as if the gateway authenticates.
3. **Exact model name.** A regex match routes fine but returns an **empty** `/v1/models` list while
completions keep working. A wrong name returns a bare 404 that reads exactly like a dead gateway.
4. **MCP tool names changed prefix separator.** Bifrost used one dash (`ct7-resolve-library-id`); the
gateway uses **two underscores** (`ct7__resolve-library-id`). Relevant only if U4 is done.
5. **`/v1/models` 404 vs empty list are different faults.** 404 means no route loaded at all; empty
means the model match is a regex. Do not conflate them when diagnosing.
---
## 7. Verification — what would prove this works
Merging config is not proving it. The checks, in order:
1. `bridge_spawn{profile: "gx"}` succeeds and the member completes a real turn ending in
`bridge_reply`. This is the first proof of the token, the URL and the model name, and it risks
nothing the fleet depends on.
2. `bridge_spawn{profile: "local"}` succeeds. If the guard allowlist was missed, this **throws** — a
loud, self-correcting failure, which is the good kind. If the restart was missed, it also throws,
for the same reason.
3. A `local` member completes a turn. That exercises streaming through two TLS edges, the auth proxy
and the gateway.
4. **Reasoning survives, checked separately on each surface.** For `local` on `/anthropic` this is
the check that catches the `/v1` versus `/anthropic` mistake, and it is the only one that does —
nothing else distinguishes a working passthrough from a translator quietly dropping thinking
deltas. For `gx` on `/v1`, this answers the open question in §3 rather than assuming it.
5. The cockpit at `auth.ltms.dev` shows requests counted against the `claude-bridge` consumer, not
`legacy`. That is the whole point of taking our own token.
6. `bridge_spawn{profile: "local-direct"}` still works, so the escape hatch is real rather than
theoretical.
7. `bridge_list` shows `gx` carrying no `credentialId`, so a `sol`/`terra` exhaustion cannot
quarantine it. This is the single-point-of-failure claim in §3b, checked rather than asserted.
---
## 7.1 What the live run actually found — 2026-08-15
U1–U2c were done, the daemon restarted onto them, and both new profiles were spawned for real. The
migration was then **reverted**. This section is the result, so none of it has to be re-derived.
### The blocker
`llm.ltms.dev` answers **HTTP 413 Request Entity Too Large** above **32 KiB (32768 bytes)**, on both
surfaces:
```
/v1 32695 bytes -> 200 /anthropic 32095 bytes -> 200
/v1 32795 bytes -> 413 /anthropic 32855 bytes -> 413
```
32 KiB is far below one real agent turn.
**Root cause — confirmed by the systems/vms side, 2026-08-15.** My guess that it was a Caddy
`request_body max_size` was **wrong**. It is Envoy, inside `aigw` on `llm.vm`. Envoy Gateway defaults
a listener's `per_connection_buffer_limit_bytes` to **32768**, and the AI Gateway buffers the *whole*
request body before it can route on the model name — so that default is not a network tuning knob
here, it is a hard ceiling on prompt size. Read out of the live Envoy `config_dump`:
```
listener default/llm/http per_connection_buffer_limit_bytes: 32768
```
Nobody chose 32 KiB; it was inherited from the default. Both TLS edges are innocent: the same
boundary reproduces on the LAN path and the internet path, and both 413s carry an `x-llm-consumer`
header their auth proxy sets only *after* authenticating — so the body cleared both edges and the
auth. Directly on `llm.vm`, `aigw` 413s at 39 KB while the vLLM backend accepts the same 39 KB and
answers 200.
**Do not plan around 32 KiB.** The intended ceiling is far higher. Their fix — a `ClientTrafficPolicy`
setting `bufferLimit: 8Mi` — is written but **not deployed** as of this note, pending their operator's
approval. I have not re-tested and will not until they confirm, so as not to measure a half-changed
system. Fixed in **systems/vms**, not here.
### The part worth remembering
Two members were spawned at the same moment with the same message:
| | `local` (claude-code, `/anthropic`) | `gx` (opencode, `/v1`) |
|---|---|---|
| READY → BUSY | 19:07:26 | 19:07:45 |
| BUSY → DONE | **19:08:51 (66s)** | **never — 10+ min, ticket FAILED** |
**`local` passed.** It passed only because the probe was three trivial questions in a fresh session,
so the request fit under 32 KiB. The profile looked healthy and was a landmine set to fire on the
first turn that reads a file.
So §7's checklist was not wrong, it was **too easy**. Any future run of it must use a task that reads
a real file. A liveness probe proves the token and the URL; it does not prove the path.
`gx` did not fail loudly either. Reproduced outside the bridge by running `opencode` by hand with the
launcher's own generated config:
```
Error: Request Entity Too Large
...compacts context, retries...
Error: Request Entity Too Large
```
opencode **catches the 413, compacts, and retries — indefinitely**. A member that fails loudly costs
one turn; this one costs the whole task and is indistinguishable from a slow worker.
> **Diagnosing a stuck opencode member.** Do not read its pane. The launcher writes its config to a
> temp dir and passes it as `OPENCODE_CONFIG` — find it with
> `ls -dt /var/folders/*/*/T/bridged-opencode-* | head -1`, check the provider block and the key's
> length and prefix (never its value), then reproduce with `opencode run --auto -m <provider>/<model>`
> using the same `OPENCODE_CONFIG`. That is what turned "it hangs" into a one-line error.
### What checked out, and needs no re-testing
- Token accepted on both surfaces. **Unauthenticated → 401**, so the Caddy proxy really does gate —
the wiki's "SecurityPolicy fails open" warning is about the gateway itself, not the edge.
- `/v1/models` returns exactly `["deepseek-v4-flash"]`, so trap 3 is clear.
- **Reasoning survives both surfaces** — see §3b above.
- The launcher's generated opencode provider block is correct, carrying a real 48-character `llmk-`
key rather than the `bridged-local-noauth` placeholder.
- `SubscriptionGuard` accepted `llm.ltms.dev` after the allowlist edit and the restart: `local`
spawned without throwing, which is the check that catches a missed restart.
### Resolution — both ceilings fixed, migration completed
systems/vms fixed both, and each was re-checked from this side rather than taken on trust:
| ceiling | was | now | our own check |
|---|---|---|---|
| listener buffer | 32 KiB | 32 Mi | 1.2 MB body → **200** (was 413) |
| LLM route timeout | 60s | 86400s | the request that truncated: **101s, `message_stop` present, 4000/4000** |
The timeout moved in two steps on 2026-08-15: 60s → 1800s, then 1800s → **86400s (24 hours)** after
the truncation risk below was discussed. They tried `request: 0s` first, which removes the
total-duration timer completely. It works, but on an `AIGatewayRoute` the **idle timeout is derived
from the request timeout**, so `0s` also removed any bound on a stalled connection. 86400s keeps a
reaper for dead connections while putting the truncation timer out of practical reach.
Neither was deliberate. The 32 KiB was Envoy Gateway's default `per_connection_buffer_limit_bytes`;
the 60s was Envoy AI Gateway's own documented default. The 60s bounded **generation** as well as
prompt size — a tiny prompt with a long answer returned 504 at 60.05s.
Two configuration facts worth keeping, from their bisection:
- **`ClientTrafficPolicy` is honoured in standalone `aigw run`; `BackendTrafficPolicy` is NOT.** A
`BackendTrafficPolicy` setting `requestTimeout` is accepted, logs nothing, and leaves the routes
unchanged (upstream `envoyproxy/gateway#9513`). What works is `timeouts: {request: …}` on each
`AIGatewayRoute` rule. Nothing from the outside distinguishes the two — the same silent-default
shape as their `SecurityPolicy` caveat.
- In that stack, "the config was accepted" proves nothing. Read the live `config_dump`.
## 7.2 The risk we accepted, and why we could not remove it
Raising the timeout made the failure **rare, not impossible**, and the residual failure is silent.
On a mid-response timeout over chunked HTTP/1.1, Envoy ends the chunked encoding *cleanly* instead of
resetting the connection, so the client receives what looks like a complete transfer
(`envoyproxy/envoy#17186` — acknowledged as a bug in 2021, closed by a stale bot, never fixed). The
December 2025 fix `envoyproxy/envoy#42269` changes locally-originated resets from `NO_ERROR` to
`INTERNAL_ERROR`, but it is **HTTP/2 only** and SSE clients here speak HTTP/1.1.
Measured on our side while the timeout was still 60s:
```
HTTP 200 61.07s 141992 bytes
message_stop 0 message_delta 0 error events 0
emitted 2473 of 4000, ending on a WELL-FORMED SSE frame
```
A syntactically valid stream that simply stops. Any timer firing mid-stream — route timeout, idle
timeout, `max_stream_duration` — fails this same way.
**The recommended defence does not transfer to us.** The right fix is to treat a stream with no
`message_stop` / `[DONE]` / `finish_reason` as failed. We cannot: our members are Claude Code and
opencode, third-party clients whose SSE parsing we do not own, and there is no seam to insert the
check. Whether either detects a missing terminator is unverified — and opencode's handling of the 413
(swallow, compact, retry forever, never surface an error) does not suggest it is strict.
So the honest statement of our position:
> Gateway traffic is acceptable at 86400s because a single request would have to run for 24 hours to
> trip the bug — **not** because we could detect it if it did.
At 86400s our **own** limit binds first, which is the ordering we want. `MessageService.ASYNC_TIMEOUT_MS`
caps a turn at 30 minutes, so a runaway request ends as a clean `FAILED` ticket that we raised, rather
than as a silently truncated `200` that we cannot see. While the gateway sat at 1800s the two numbers
were equal and did not nest, so a gateway-side stall could have been misread as a bug in our own ticket
handling. That ambiguity is now gone.
**If a member ever returns a confident but truncated answer, suspect this before anything in our own
code.** That is the whole reason this section exists.
---
## 8. Related
- [CB-589 / #74](https://git.ltms.dev/lms/claude-bridge/issues/74) — cost-first placement and a
gateway that reports live capacity. The per-consumer figures this migration unlocks are the first
input that ticket actually needs.
- `docs/CB-500-Multi-Tier-Coordination.md` §11 — the distributed-sandbox topology this gateway is
part of.
+59 -4
View File
@@ -1,7 +1,9 @@
# M4 - Fleet health, recovery, routing, and capacity
**Status:** Design accepted on 2026-08-15. CB-573 part 1 has shipped the classification model and
the `bridge_list` capacity view; the remaining M4 units are not yet shipped.
the `bridge_list` capacity view; the remaining M4 units are not yet shipped. See
[Unit 2 — what has landed so far](#unit-2---what-has-landed-so-far) before planning Unit 2 work:
some of its criteria were met by separate CB tickets, and one of them contradicts the unit text.
**Scope:** Fleet evidence, safe mechanical repair, lead routing, capacity reporting, and optional
human notification.
**Grounded in:** `health/FleetHealth`, `health/PaneBudget`, `inject/StatusPoller`,
@@ -237,7 +239,7 @@ state never presents stop as the only action.
| Release cause | Process action | Provisioned worktree |
|---|---|---|
| `SPAWN_ROLLBACK` before registration or delivery | Stop and clean up | Remove |
| `COMPLETED` for `READY` or `DONE` without pending work, idle TTL, or successful context-cap completion | Stop | Remove under completed policy |
| `COMPLETED` for `READY` or `DONE` without pending work, idle TTL, or successful context-cap completion | Stop | Remove only if clean; preserve a dirty worktree (CB-576) |
| `NEVER_READY` | Stop | Preserve |
| `GONE` | Best-effort stop | Preserve |
| `TURN_FAILED` or lead abort while `BUSY` or `FAILED` | Stop | Preserve |
@@ -744,8 +746,28 @@ worktree discovery.
Acceptance criteria:
1. Every accepted send receives a stable `TurnToken` tied to target, exact waiter, session turn,
delivery baseline, and task outcome.
1. Every accepted send receives a stable `TurnToken` tied to target, exact waiter, and delivery
baseline.
**Corrected during implementation (2026-08-15).** This criterion first also required the session
turn number and the task outcome. That is not implementable at this layer, and the implementer
refused it three times rather than fabricate a value — correctly. The reason is an ordering fact
that is invisible from any single class: `MessageService` owns acceptance and holds the waiter and
the async `Task`, but it learns nothing about delivery, because the delivery event goes to
`CompletionResolver` through `TurnListener.onDelivered`. And `CompletionResolver.onDelivered` runs
*before* `SessionManager.onDelivered`, so the session turn number does not exist yet at the only
point where the token could capture it.
Two ways out were rejected. A shared registry keyed by target reintroduces exactly the "whichever
send happens to be waiting" ambiguity the token exists to remove — the same weak claim
`Rendezvous.currentWaiter` warns about. Injecting a turn counter into `MessageService` adds a
required cross-layer dependency to populate a field that nothing in this slice reads, which is
speculative coupling across a boundary already shown to be fragile.
So the token identifies the **accepted send**, and `SessionManager` keeps verifying its own
delivery separately. Repair (criterion 2) does need the session turn; binding it means resolving
that acceptance-versus-delivery ordering first, and that work belongs to the repair unit, not
here. The token record carries a comment saying the field is deliberately absent.
2. Repair requires the same `BUSY` token, two raw `IDLE` or `DONE` snapshots, no conflicting
observation, exact open waiter, successful baseline, and new recognised assistant output.
3. Missing, failed, late, or post-restart baseline never authorises repair.
@@ -779,6 +801,39 @@ Acceptance criteria:
21. Tests cover both real traces, all repair refusals, clipping, explicit-reply and next-turn races,
restart without capture, concurrent send and release, and preserved discovery after restart.
#### Unit 2 - what has landed so far
Checked against `main` at `e09cac6` on 2026-08-15. Unit 2 was written as one block, but parts of it
have since been built by separate CB tickets. Read this before planning the rest, or that work gets
done twice.
The check was a symbol survey of `bridged/src/main/java` plus the merge history. It tells you whether
the machinery exists at all. It is **not** a line-by-line audit of whether each criterion is fully
met, and I did not run one.
| Criterion | Marker searched for | Found in main source | Reading |
|---|---|---|---|
| 1 | `TurnToken` | 8 files | **Done** — unit 2a, merged as `fec284e`. Criterion 1 was corrected first; see the note under it. |
| 2-5, 9 | `REPAIRED` | 0 files | Not started. The whole guarded-repair path is absent. |
| 6, 7, 10 | `reconcileLostBoundary` | 0 files | Not started. |
| 12 | CB-568 failure operation | via CB-580 | **Partial.** CB-580 (`0af902e`) routes `GONE` and `NEVER_READY` into the one idempotent target-wide failure. I did not check that release and abnormal stop go through the same call. |
| 14 | `DELEGATION_ORPHANED` | 3 files | **Partial.** The health state exists. The teardown-invariant check that creates it, and the retry rule, do not. |
| 15 | `SPAWN_ROLLBACK` | 0 files | **Contradicted — see below.** |
| 16 | — | — | Partial at best. CB-576 made release preserve a dirty worktree; whether explicit stop is state-aware is not checked. |
| 17, 18 | `preservedWorktrees` | 0 files | Not started. No manifest, and no lead-only `bridge_list` field. |
| 19 | `WORK_PRODUCT_AT_RISK` | 0 files | Not started. |
**Criterion 15 no longer matches the code, and the code is right.** It says "normal `COMPLETED`
remove worktrees". Since CB-576 (`500bfa2`) that is false on purpose: a `COMPLETED` release now
preserves the worktree when it still holds uncommitted work, because deleting it destroys work
nobody can get back. CB-576 was filed after exactly that loss. CB-581 goes further — if the
dirty-check itself fails, the worktree is preserved rather than removed, since "we could not tell"
must not be treated as "it is clean".
So criterion 15 should be rewritten as: `SPAWN_ROLLBACK` and a `COMPLETED` release with a **clean**
worktree remove it; abnormal causes, shutdown, a dirty worktree, and a failed dirty-check all
preserve it. `SPAWN_ROLLBACK` itself does not exist yet.
### Unit 3 - Typed inbox and member routing
Scope: semantic record, AMQP migration, both adapters, member routing, polling, and member health in
+1 -1
View File
@@ -26,7 +26,7 @@
],
"enabled": true,
"environment": {
"GITEA_ACCESS_TOKEN": "{env:GITEA_ACCESS_TOKEN}",
"GITEA_ACCESS_TOKEN": "{env:WORKER_GITEA_TOKEN}",
"GITEA_HOST": "{env:GITEA_HOST}"
}
}
+32
View File
@@ -0,0 +1,32 @@
#!/usr/bin/env bash
#
# CB-594 — the only reason this file exists: launchd does not run a login shell.
#
# WORKER_GITEA_TOKEN and AI_GATEWAY_TOKEN live in ${SHARED_ENV}/tools/secrets.sh, sourced only by a
# LOGIN shell (.zprofile/.zshrc etc). launchd execs a job's ProgramArguments directly — no shell, no
# profile, nothing sourced (the plist's own PATH comment documents the same gap one variable over).
# A daemon started that way boots fine and looks healthy; the failure is invisible until a worker
# tries to open a PR (WORKER_GITEA_TOKEN empty) or a gateway profile gets a 401 (AI_GATEWAY_TOKEN
# empty) — hours later, with nothing tying the two together (CB-591, CLAUDE.md "Redeploying the
# daemon"). Bridged now also logs which required secret names resolved at startup (see
# Bridged.reportRequiredSecrets), but that log line can only tell the truth if the tokens had a
# chance to be sourced in the first place — which is this script's entire job.
#
# So: launchd execs THIS script instead of java directly. This script execs a login shell
# ('zsh -l'), which sources secrets.sh, and that shell execs the real command in its place — one
# process throughout (exec, not a subshell fork), so launchd's PID tracking, KeepAlive, and
# StandardOut/ErrorPath all still see the one process they expect.
#
# The plist passes the full command as THIS script's own arguments, e.g.:
# ProgramArguments = [ .../bridged-launchd-wrapper.sh, /path/to/java, -jar, /path/to/bridged.jar,
# bridged.yaml ]
# so the wrapper stays generic and the actual command lives in exactly one place (the plist), not
# duplicated here.
set -euo pipefail
if [ "$#" -eq 0 ]; then
echo "bridged-launchd-wrapper.sh: no command given — check the plist's ProgramArguments" >&2
exit 2
fi
exec /bin/zsh -lc 'exec "$@"' -- "$@"
+135
View File
@@ -0,0 +1,135 @@
#!/usr/bin/env bash
#
# CB-596 step 1: measure which credentials a live member actually holds.
#
# The claim under test is that a member's herdr pane starts a LOGIN shell, that shell sources
# ${SHARED_ENV}/tools/secrets.sh, and so a member inherits every name that file exports — while
# CB-592 blocks exactly one of them (GITEA_ACCESS_TOKEN). That is an inference from the code, not a
# measurement, and issue #82 says plainly: do not build a fix on the inference. This is the
# measurement.
#
# WHY THIS IS A SCRIPT AND NOT A COMMAND SOMEONE TYPES
#
# Enumerating credential names inside a member is exactly the action that should need the operator's
# explicit approval, and the command classifier refuses it. That refusal is correct. This script is
# the seam: it is one auditable file the operator can read once, top to bottom, and then run — rather
# than approving an ad-hoc shell pipeline whose behaviour they have to take on trust.
#
# WHAT IT WILL NOT DO
#
# * It never prints a credential value, and never any prefix or suffix of one. Not one character.
# Issue #82's criterion 1 asked for a 6-character prefix; this prints a truncated SHA-256 instead.
# A prefix of a short secret is most of the secret, and it would end up pasted into a ticket. The
# hash answers every question the prefix was for — is it set, is it the same value as over there,
# is it the CB-592 sentinel — and answers none of the ones it should not.
# * It never writes anywhere, never contacts the network, and never touches secrets.sh, which is
# the operator's file.
#
# HOW TO RUN IT
#
# 1. As the operator, in a member's pane (a spawned worker's terminal):
# bash scripts/probe-member-credentials.sh
# 2. For the comparison row, in your OWN shell — a lead, not a member:
# bash scripts/probe-member-credentials.sh --allow-outside-member
#
# The two outputs side by side are the finding: any name whose hash matches between them is a
# credential the member holds in full.
#
set -uo pipefail
# The names ${SHARED_ENV}/tools/secrets.sh exports, recorded on 2026-08-16 (issue #82). Names only —
# this list contains no values and never should. If secrets.sh gains a name, this list goes stale and
# the probe silently stops asking about it; that staleness is itself part of what #82's criterion 4
# has to solve, so it is called out in the summary rather than hidden.
NAMES=(
AI_GATEWAY_TOKEN BESZEL_ADMIN_EMAIL BESZEL_ADMIN_PASSWORD
BESZEL_HUB_URL BESZEL_KEY BESZEL_UNIVERSAL_TOKEN
BRAIN_MCP_TOKEN CF_ACCOUNT_ID CF_API_TOKEN
CF_USER_TOKEN CONFLUENCE_API_TOKEN CONFLUENCE_USERNAME
CONTEXT7_TOKEN GITEA_HOST GITLAB_OAUTH_CLIENT_SECRET
GITLAB_PERSONAL_ACCESS_TOKEN GRAFANA_ADMIN_PASSWORD GRAFANA_ADMIN_USER
HASS_TOKEN HW_PASSWORD HW_USER
LTMS_API_KEY MEMORY_MCP_TOKEN METRICS_PUSH_TOKEN
OPENCODE_AUTOMODE_MODEL TELEGRAM_BOT_TOKEN TELEGRAM_CHAT_ID
TS_API_KEY TS_AUTHKEY WORKER_GITEA_TOKEN
GITEA_ACCESS_TOKEN
)
allow_outside=0
for arg in "$@"; do
case "$arg" in
--allow-outside-member) allow_outside=1 ;;
-h|--help) sed -n '2,40p' "$0"; exit 0 ;;
*) echo "unknown argument: $arg" >&2; exit 2 ;;
esac
done
if [ "${BRIDGED_MEMBER:-}" != "1" ] && [ "$allow_outside" -eq 0 ]; then
cat >&2 <<'EOF'
refusing to run: BRIDGED_MEMBER is not 1, so this is not a member's shell.
The finding this probe exists for is what a MEMBER holds. Run it in a spawned worker's pane. If you
meant to take the comparison reading from your own shell, pass --allow-outside-member and the output
will be labelled as such.
EOF
exit 1
fi
# Prefer sha256sum (Linux), fall back to shasum (macOS). If neither exists, report presence and
# length only — degraded, but never a value.
hasher=""
if command -v sha256sum >/dev/null 2>&1; then
hasher="sha256sum"
elif command -v shasum >/dev/null 2>&1; then
hasher="shasum -a 256"
fi
digest() { # value -> first 12 hex chars of its sha256, or "-" when no hasher is available
[ -z "$hasher" ] && { printf '%s' "-"; return; }
printf '%s' "$1" | $hasher | cut -c1-12
}
if [ "${BRIDGED_MEMBER:-}" = "1" ]; then
where="MEMBER (BRIDGED_MEMBER=1)"
else
where="NOT a member — comparison reading only"
fi
echo "CB-596 credential probe"
echo "reading from : $where"
echo "shell : ${SHELL:-unknown}"
echo "hash : ${hasher:-none available — lengths only}"
# Only printed so the two readings can be told apart when they are pasted side by side.
echo "host : $(hostname 2>/dev/null || echo unknown)"
echo
printf '%-30s %-7s %6s %s\n' "NAME" "STATE" "LEN" "SHA256-12"
printf '%-30s %-7s %6s %s\n' "------------------------------" "-------" "------" "------------"
set_count=0
for name in "${NAMES[@]}"; do
value="${!name:-}"
if [ -z "$value" ]; then
printf '%-30s %-7s %6s %s\n' "$name" "unset" "-" "-"
else
set_count=$((set_count + 1))
printf '%-30s %-7s %6s %s\n' "$name" "SET" "${#value}" "$(digest "$value")"
fi
done
echo
echo "$set_count of ${#NAMES[@]} names are set in this shell."
echo
cat <<'EOF'
How to read this:
* Take the MEMBER reading and the comparison reading, and line them up. A name whose SHA256-12
matches on both sides is a credential the member holds in full. That is the finding.
* GITEA_ACCESS_TOKEN is the control. CB-592 replaces it with a blocked sentinel, so its hash
should DIFFER between the two readings. If it matches, CB-592 is not working and that is the
most urgent thing on this page.
* AI_GATEWAY_TOKEN matching is expected and correct, not a leak: bridged.yaml names it in
`tokenEnv:` for the local and gx profiles, so a member reaching the gateway is by design.
* A name that is set here but is NOT in the list above will not appear at all. The list was
recorded on 2026-08-16 and does not update itself. Anything added to secrets.sh since then is
invisible to this probe — which is the same gap issue #82 criterion 4 asks to close properly.
EOF
+361
View File
@@ -0,0 +1,361 @@
#!/usr/bin/env bash
#
# Rebuild and restart the bridged daemon.
#
# A merge is not a deployment: the running daemon holds the jar it was started with, so code merged
# to main does nothing until this runs. See CLAUDE.md -> "Redeploying the daemon".
#
# This script exists to turn six remembered traps into one auditable command:
#
# 1. A piped `mvn` hides BUILD FAILURE behind a zero exit, so the build here is never piped.
# 2. The daemon must start from a LOGIN shell, or the tokens it hands to members are empty:
# WORKER_GITEA_TOKEN (workers cannot open a PR) and AI_GATEWAY_TOKEN (401 at llm.ltms.dev).
# Both are read from the DAEMON's own environment at spawn time, so a value added to
# secrets.sh after startup is absent. Nothing logs this here, so the script checks and says
# so — and since CB-594, bridged's own startup log says so too, by env var name.
# 3. An old daemon that never actually died looks identical from the outside, so the script waits
# for the process to exit and for the port to free before it starts a new one.
# 4. "It started" is not "it works": the script polls /healthz until it answers, and reports the
# herdr protocol number, because healthz can be green while every spawn fails on a protocol
# mismatch.
# 5. Restarting under live members drops their tickets, so the script refuses unless you confirm
# the fleet is drained.
# 6. CB-594 — the launchd agent (deploy/dev.ltms.bridged.plist), if installed and loaded, is a
# SECOND supervisor: its KeepAlive.SuccessfulExit=false restarts the daemon on any nonzero
# exit, and a bare SIGTERM makes this JVM exit 143 even with its shutdown hook running to
# completion (measured — see the CB-594 report). A plain `kill` here would race launchd's own
# restart of the OLD jar. So this script detects whether the agent is loaded and, only then,
# swaps `kill` + manual `nohup` for `launchctl unload`/`load` — the one supervisor in control
# at any moment is whichever one you asked to act, never both.
#
# Usage:
# scripts/redeploy-bridged.sh # build, confirm, restart, verify
# scripts/redeploy-bridged.sh --yes # skip the drain confirmation (fleet already checked)
# scripts/redeploy-bridged.sh --no-build # restart the jar already on disk
# scripts/redeploy-bridged.sh --check # report state and exit; changes nothing
#
# Exits non-zero on any failure. A failed build never stops the running daemon.
set -euo pipefail
REPO="$(cd "$(dirname "${BASH_SOURCE[0]}")/.." && pwd)"
BRIDGED="$REPO/bridged"
JAR="$BRIDGED/target/bridged.jar"
OUT="$BRIDGED/bridged.out"
# Matches BOTH the absolute form and the relative `java -jar target/bridged.jar` a hand-start
# produces from inside bridged/. Anchoring on the absolute path alone was a real bug: the daemon
# restarted correctly and the script still reported "no process appeared", because it launched with
# a relative path and then looked for an absolute one.
PATTERN='target/bridged.jar'
HEALTH='http://127.0.0.1:8765/healthz'
STOP_WAIT=30 # seconds to wait for a clean exit before reporting failure
HEALTH_WAIT=60 # seconds to wait for /healthz to answer after start
# CB-594: the launchd agent this script must not fight with (see trap 6 above).
LAUNCHD_LABEL='dev.ltms.bridged'
LAUNCHD_PLIST="$HOME/Library/LaunchAgents/$LAUNCHD_LABEL.plist"
DO_BUILD=1; ASSUME_YES=0; CHECK_ONLY=0
for arg in "$@"; do
case "$arg" in
--yes|-y) ASSUME_YES=1 ;;
--no-build) DO_BUILD=0 ;;
--check) CHECK_ONLY=1 ;;
-h|--help) sed -n '3,37p' "${BASH_SOURCE[0]}"; exit 0 ;;
*) echo "unknown option: $arg (try --help)" >&2; exit 2 ;;
esac
done
say() { printf '\n\033[1m== %s\033[0m\n' "$*"; }
ok() { printf ' ok %s\n' "$*"; }
warn() { printf ' WARN %s\n' "$*"; }
die() { printf '\n FAIL %s\n\n' "$*" >&2; exit 1; }
jar_id() { [ -f "$JAR" ] && shasum -a 256 "$JAR" | cut -c1-12 || echo "absent"; }
running_pid() { pgrep -f "$PATTERN" || true; }
# `launchctl list <label>` exits 0 iff the label is loaded (registered with launchd) — true whether
# or not it is currently running, which is exactly "supervision is active" for our purposes. Read-
# only: neither helper below changes anything, so both are also safe under --check.
launchd_installed() { [ -f "$LAUNCHD_PLIST" ]; }
launchd_loaded() { launchctl list "$LAUNCHD_LABEL" >/dev/null 2>&1; }
# CB-600: the script computes its own log path from where it sits on disk (REPO, above); the
# plist hard-codes an absolute StandardOutPath. Nothing forced the two to agree — if this script
# were ever run from a checkout other than the one the loaded plist names, launchd would start and
# log the daemon correctly, while every check below (the fresh "bridged listening" line, the
# ERROR-count scan) would read a different, empty or stale file and the script would report a
# clean restart while the daemon crash-loops. Pure and side-effect-free besides `die`/`ok` — reads
# the two paths, resolves them, compares — so it never touches launchd or the daemon and can be
# exercised by sourcing this script (see the SOURCED guard below) without installing the agent.
check_log_path_matches_plist() {
local script_out="$1" plist_path="$2"
local plist_out resolved_out resolved_plist_out
# Checked by exit status, not by emptiness: on a missing file/key PlistBuddy exits nonzero but
# still writes a message ("File Doesn't Exist, Will Create: ...") that command substitution
# would happily capture as if it were the real value — testing only `-z` missed that case.
if ! plist_out="$(/usr/libexec/PlistBuddy -c 'Print :StandardOutPath' "$plist_path" 2>/dev/null)" \
|| [ -z "$plist_out" ]; then
die "launchd agent is loaded but PlistBuddy could not read StandardOutPath from
$plist_path
— cannot verify the daemon logs where this script is about to look. Fix the plist before
redeploying supervised."
fi
resolved_out="$(cd "$(dirname "$script_out")" 2>/dev/null && pwd -P)/$(basename "$script_out")" || true
resolved_plist_out="$(cd "$(dirname "$plist_out")" 2>/dev/null && pwd -P)/$(basename "$plist_out")" || true
if [ -z "$resolved_out" ] || [ -z "$resolved_plist_out" ] || [ "$resolved_out" != "$resolved_plist_out" ]; then
die "log path mismatch — this script reads
$script_out (resolved: ${resolved_out:-<directory does not exist>})
but the loaded plist's StandardOutPath is
$plist_out (resolved: ${resolved_plist_out:-<directory does not exist>})
Under supervision the daemon writes to the PLIST's path, not necessarily this script's — every
post-restart check below (the fresh 'bridged listening' line, the ERROR-count scan) would read
the wrong file and could report a clean restart while the daemon crash-loops. Fix the mismatch
(move this checkout to match the plist, or edit the plist's StandardOutPath/StandardErrorPath)
before redeploying supervised."
fi
ok "log path check: script and plist agree ($resolved_out)"
}
# CB-600: sourceable for testing. When this file is SOURCED (not executed) it stops here — nothing
# below runs — so a test harness can `source` it to call check_log_path_matches_plist (or the
# other pure helpers above) against a throwaway plist fixture without ever reaching the mutating
# flow (build/stop/start) or touching the real daemon or launchd. On a normal `./redeploy-bridged.sh`
# invocation `(return 0 2>/dev/null)` fails (return is illegal at top level of an executed script),
# so this whole block is a no-op and every line below still runs exactly as before.
if (return 0 2>/dev/null); then
return 0
fi
# ---------------------------------------------------------------- report state
say "current state"
OLD_PID="$(running_pid)"
if [ -n "$OLD_PID" ]; then
ok "daemon running, pid $OLD_PID"
else
warn "no daemon running — this will be a cold start"
fi
ok "jar on disk: $(jar_id) ($([ -f "$JAR" ] && date -r "$JAR" '+%Y-%m-%d %H:%M:%S' || echo 'none'))"
ok "HEAD: $(git -C "$REPO" log --oneline -1)"
# CB-594: supervision state. Installed and loaded are different facts — a copied-but-never-loaded
# plist supervises nothing, and a loaded label with no file backing it (rare, but possible after an
# edited/moved plist) is still what launchd will act on.
if launchd_installed; then
ok "launchd agent installed: $LAUNCHD_PLIST"
else
warn "launchd agent NOT installed (no supervision — a crash will not restart the daemon)."
fi
SUPERVISED=0
if launchd_loaded; then
SUPERVISED=1
ok "launchd agent loaded ($LAUNCHD_LABEL) — launchd supervises this daemon"
# CB-600: fail loudly here, before ANY other check runs, if this script and the loaded plist
# would read different log files — every check after this point is worthless otherwise.
check_log_path_matches_plist "$OUT" "$LAUNCHD_PLIST"
else
warn "launchd agent not loaded — this script is the only thing that will restart the daemon."
fi
# The trap with no log line. Checked in a LOGIN shell, because that is how the daemon is started
# below. Never prints the value — only whether it resolved.
if zsh -lc '[ -n "${WORKER_GITEA_TOKEN:-}" ]' 2>/dev/null; then
ok "WORKER_GITEA_TOKEN resolves in a login shell"
else
warn "WORKER_GITEA_TOKEN is EMPTY in a login shell."
warn "The daemon will start fine and workers will silently fail to open PRs."
warn "Fix \${SHARED_ENV}/tools/secrets.sh before relying on worker checkpoints."
fi
# Same trap, second variable (CB-591). A profile's `tokenEnv:` is resolved from the DAEMON's own
# process environment by HerdrPeerLauncher.resolveEnv, so a token added to secrets.sh after the
# daemon started is simply absent. The launcher then injects an empty token and llm.ltms.dev answers
# 401 — long after the restart, and with nothing tying the two together.
if zsh -lc '[ -n "${AI_GATEWAY_TOKEN:-}" ]' 2>/dev/null; then
ok "AI_GATEWAY_TOKEN resolves in a login shell"
else
warn "AI_GATEWAY_TOKEN is EMPTY in a login shell."
warn "Any profile whose tokenEnv is AI_GATEWAY_TOKEN will get an empty token and 401 at the gateway."
warn "This only matters once a profile points at llm.ltms.dev — harmless before that."
fi
if [ "$CHECK_ONLY" = 1 ]; then
say "--check: nothing changed"
exit 0
fi
# ---------------------------------------------------------------------- build
# Deliberately before the stop: a failed build must never leave the fleet down.
if [ "$DO_BUILD" = 1 ]; then
say "build"
BUILD_LOG="$(mktemp -t bridged-build)"
echo " log: $BUILD_LOG"
if ! mvn -f "$BRIDGED/pom.xml" clean install > "$BUILD_LOG" 2>&1; then
grep -E 'ERROR|BUILD FAILURE|Tests run:.*Failures: [1-9]|Tests run:.*Errors: [1-9]' "$BUILD_LOG" \
| head -20 || true
die "build failed — the running daemon was NOT touched. Full log: $BUILD_LOG"
fi
grep -E '^\[INFO\] Tests run:.*Failures' "$BUILD_LOG" | tail -1 | sed 's/^\[INFO\] / /' || true
ok "BUILD SUCCESS"
ok "jar now: $(jar_id)"
else
say "build skipped (--no-build)"
fi
[ -f "$JAR" ] || die "no jar at $JAR — run without --no-build"
# ----------------------------------------------------------------- drain gate
if [ -n "$OLD_PID" ] && [ "$ASSUME_YES" = 0 ]; then
say "drain check"
echo " A restart drops every in-flight ticket and rendezvous. A member's report"
echo " is NOT recoverable once its ticket is gone."
echo
echo " Confirm with bridge_list that no members are live, and bridge_poll anything"
echo " you still want, BEFORE continuing."
echo
read -r -p " Fleet drained? type yes to restart: " reply
[ "$reply" = "yes" ] || die "aborted — nothing changed"
fi
# ------------------------------------------------------------------ stop
#
# CB-594: when SUPERVISED, launchd owns the stop — never a raw `kill` here. A bare SIGTERM makes
# this JVM exit 143 even with its shutdown hook running to completion (verified separately: a
# throwaway Java process with an equivalent shutdown hook, sent SIGTERM from a login shell that
# could `wait` on it directly, reported exit code 143 every time — never 0). launchd's
# KeepAlive.SuccessfulExit=false treats any nonzero exit as a crash and restarts the OLD jar,
# which would race this script's own restart of the NEW one. `launchctl unload` avoids that race
# by deregistering the job first, so no KeepAlive is left armed when the process actually stops.
if [ -n "$OLD_PID" ]; then
say "stop"
RESTART_MARK="$(wc -l < "$OUT" 2>/dev/null || echo 0)" # verify a FRESH line appears later
if [ "$SUPERVISED" = 1 ]; then
echo " supervision is ON: using 'launchctl unload' (not kill) so launchd's own KeepAlive"
echo " cannot restart the OLD jar out from under this script — see the CB-594 comment above."
launchctl unload -w "$LAUNCHD_PLIST" \
|| die "launchctl unload failed — the daemon may still be under supervision; investigate before retrying"
else
kill "$OLD_PID"
fi
for _ in $(seq "$STOP_WAIT"); do
[ -z "$(running_pid)" ] && break
sleep 1
done
if [ -n "$(running_pid)" ]; then
die "pid $OLD_PID still alive after ${STOP_WAIT}s. Not escalating to kill -9 automatically:
the shutdown hook releases sessions and worktrees in order, and killing it hard can
leave worktrees and panes behind. Investigate, then kill -9 by hand if you accept that."
fi
ok "pid $OLD_PID exited"
elif [ "$SUPERVISED" = 1 ]; then
# Loaded but not currently running (e.g. throttled after a crash loop). Unload it anyway so the
# start step below does a clean load, never a load stacked on an already-loaded label.
say "stop"
RESTART_MARK="$(wc -l < "$OUT" 2>/dev/null || echo 0)"
launchctl unload -w "$LAUNCHD_PLIST" 2>/dev/null || true
ok "launchd agent unloaded (was already not running)"
else
RESTART_MARK="$(wc -l < "$OUT" 2>/dev/null || echo 0)"
fi
# ------------------------------------------------------------------ start
# Unsupervised: login shell (zsh -l) is what puts the secrets on the daemon's environment, and cwd
# must be bridged/ because the daemon resolves bridged.yaml, logs/ and target/ relative to it.
# Supervised: launchd does both — deploy/dev.ltms.bridged.plist points ProgramArguments at
# scripts/bridged-launchd-wrapper.sh (CB-594), which is what execs the login shell in launchd's
# place, and WorkingDirectory in the plist already pins bridged/.
say "start"
if [ "$SUPERVISED" = 1 ]; then
echo " supervision is ON: using 'launchctl load' so launchd starts and keeps supervising this"
echo " process, instead of a manual nohup that launchd would know nothing about."
# CB-600: 'launchctl unload -w' above already persisted Disabled=true for this label. A load -w
# that succeeds clears it; a load -w that FAILS leaves the agent both stopped and disabled — worse
# than before this script ran, because a later reboot or login will not bring it back either. One
# retry covers a transient race (e.g. launchd not yet fully done deregistering); if it still fails,
# die with the exact recovery command rather than a bare "failed".
if ! launchctl load -w "$LAUNCHD_PLIST" 2>/dev/null; then
warn "launchctl load failed on the first attempt — retrying once after a short pause"
sleep 2
launchctl load -w "$LAUNCHD_PLIST" || die "launchctl load failed twice.
The agent is now STOPPED and DISABLED — it will NOT come back on its own, not even after a
reboot or login, because 'launchctl unload -w' above persisted Disabled=true and load -w
never got the chance to clear it. Recover with:
launchctl load -w \"$LAUNCHD_PLIST\"
If that still fails, check 'launchctl list $LAUNCHD_LABEL', validate the plist with
'plutil -lint \"$LAUNCHD_PLIST\"', and check $OUT before assuming a retry will succeed."
fi
else
# Absolute jar path so `ps` names which checkout is running.
( cd "$BRIDGED" && zsh -lc "nohup java -jar '$JAR' >> bridged.out 2>&1 &" )
fi
for _ in $(seq 10); do
NEW_PID="$(running_pid)"
[ -n "$NEW_PID" ] && break
sleep 1
done
[ -n "${NEW_PID:-}" ] || die "no process appeared. Last lines of $OUT:
$(tail -20 "$OUT" 2>/dev/null)"
[ "$NEW_PID" != "${OLD_PID:-}" ] || die "pid unchanged ($NEW_PID) — the old daemon never died"
ok "started, pid $NEW_PID"
# ------------------------------------------------------------------ verify
say "verify"
HEALTH_BODY=""
for _ in $(seq "$HEALTH_WAIT"); do
if HEALTH_BODY="$(curl -fsS --max-time 2 "$HEALTH" 2>/dev/null)"; then break; fi
HEALTH_BODY=""
sleep 1
done
if [ -z "$HEALTH_BODY" ]; then
# 503 still means the daemon is up — it means herdr is unreachable. Say which.
CODE="$(curl -s -o /dev/null -w '%{http_code}' --max-time 2 "$HEALTH" 2>/dev/null || echo 000)"
if [ "$CODE" = "503" ]; then
warn "/healthz answers 503 degraded — the daemon is up but herdr is unreachable."
warn "Spawns will fail. Check herdr before delegating anything."
curl -s --max-time 2 "$HEALTH" 2>/dev/null | head -3 || true
else
die "/healthz never answered within ${HEALTH_WAIT}s (last code: $CODE). Last lines of $OUT:
$(tail -30 "$OUT" 2>/dev/null)"
fi
else
ok "/healthz 200 — $HEALTH_BODY"
warn "healthz green only proves herdr ANSWERS. If its protocol number changed, spawns can still"
warn "fail — prove a real spawn before trusting the fleet."
fi
# A fresh listening line, strictly after the restart mark. An old daemon that never died would
# otherwise let an old line pass for a new one.
if tail -n "+$((RESTART_MARK + 1))" "$OUT" 2>/dev/null | grep -q 'bridged listening'; then
ok "$(tail -n "+$((RESTART_MARK + 1))" "$OUT" | grep 'bridged listening' | tail -1)"
else
warn "no fresh 'bridged listening' line after the restart — check $OUT yourself"
fi
# Config keys the daemon accepted or deferred at boot. This is usually WHY you restarted.
say "config at boot"
tail -n "+$((RESTART_MARK + 1))" "$OUT" 2>/dev/null \
| grep -iE 'deferred|classification:|fleet health:|coverage' | tail -8 | sed 's/^/ /' \
|| echo " (nothing reported)"
# Errors since the restart, anchored to the marker so old noise cannot leak in.
ERRS="$(tail -n "+$((RESTART_MARK + 1))" "$OUT" 2>/dev/null | grep -cE ' (ERROR|SEVERE) ' || true)"
say "result"
ok "pid $NEW_PID, jar $(jar_id)"
if [ "${ERRS:-0}" -gt 0 ]; then
warn "$ERRS ERROR lines since restart:"
tail -n "+$((RESTART_MARK + 1))" "$OUT" | grep -E ' (ERROR|SEVERE) ' | tail -5 | sed 's/^/ /'
else
ok "no ERROR lines since restart"
fi
echo
echo " Next: call bridge_whoami and confirm it still answers 'primary'. A lead whose tab label"
echo " no longer matches fleet.leaders.*.tab is demoted to worker and refuses orchestration."
echo