Advance the wiki submodule pointer to include the CB-307 Stage 2/3 as-built
docs plus the intervening CB-401/delegation-directive commits. Standalone
pointer bump — not bundled into a feature commit.
Primary verification pass over the delegated push-loop delivery:
- Remove the unused LongSupplier clock threaded into ReplyPushLoop
(timing is the scheduler's; the field was never read) from the
component, Bridged wiring, and both test call sites.
- Collapse the single-statement WAIT_BUSY switch arm (redundant block).
- Drop now-dead test scaffolding: the always-"idle" recordingClient
param and unused AgentStatus/AtomicReference imports.
IDE diagnostics 0/0 on all changed files; mvn clean install green
(242 tests, 0 failures).
Chose a small dedicated scheduled loop over AgentControl.send guarded by an
injectable status check, instead of reusing the worker Injector (which couples
to WorkerPresence/StatusPoller). Isolated + unit-testable via injected clock.
Records live ground truth: this primary resolves to term_656c8cc03e1f0b1 (w2:pY)
— confirms the primary runs in a herdr pane so the push path is exercisable.
The reliability layer over the durable inbox (Stage 2, 2bc5f3a): push a nudge
into the primary's own herdr pane the moment a no-waiter reply lands, remind on
a bounded backoff until the primary drains (ack=drain), degrade to pull when the
primary pane isn't resolvable. Grounds the seam: the primary terminal_id is
already derivable via ConnectionIdentity/PaneLocator, just discarded today.
Refs gitea #5.
Behind the existing ReplyInbox port, add an AMQP-backed adapter selected by a
`broker:` block in config (absent → the in-memory soft-state inbox; present →
AMQP). Mapping is consume-and-hold with deferred manual ack: each target owns a
durable queue `agent.<target>.inbox`; a manual-ack consumer pulls persistent
messages into an in-memory held map (dedup by msgId) but does not ack; peek
returns the snapshot; ack acks the broker delivery-tag and drops it. A crash
before caller-ack leaves messages unacked, so the broker redelivers on
reconnect — genuine durability with the port contract preserved. bridged still
owns no persistence; the broker does.
- msg/AmqpReplyInbox: the adapter (single synchronized channel; recovery
listener clears held on reconnect so fresh delivery-tags repopulate).
- config/BridgedConfig: nullable Broker(uri) record; isConfigured() gates it.
- Bridged.main: select adapter; close the AMQP connection in the ordered
shutdown hook (no-op for the in-memory inbox).
- deps: com.rabbitmq:amqp-client (main); testcontainers rabbitmq/junit-jupiter
(test). Pinned commons-compress 1.27.1 + commons-lang3 3.18.0 to clear the
test-scope CVEs those pull. Production default LavinMQ; RabbitMQ URI-swap.
- tests: BridgedConfigTest broker-selection cases (hermetic); AmqpReplyInbox
contract test (@Tag("contract"), Testcontainers RabbitMQ) proving
publish/peek/ack, msgId dedup, and cross-restart redelivery. Excluded from
the default build so `mvn clean install` stays hermetic (210 green).
Problem: the reverse (worker->primary) path was Rendezvous, a map of LIVE blocking
waiters only. A bridge_reply arriving with no open send hit Rendezvous.complete()
-> no waiter -> returned false -> the reply was silently DISCARDED (worker saw an
error / REST 409). No message-id/dedup/ack existed anywhere.
Stage 1 (no broker, soft-state) behind one port:
- ReplyInbox port + InboxMessage record; InMemoryReplyInbox adapter (per-target
FIFO via LinkedHashMap, dedup by msgId, thread-safe). Soft-state, not persistence.
- MessageService.reply(session, content): resolve an open send, else publish to the
inbox with a minted UUID (was a silent drop). drainReplies(target) = peek + ack.
- BridgeMcp.reply / BridgedApp.replyMessage repointed off bare Rendezvous.resolve
onto messages.reply -> no-waiter is now SUCCESS (queued), not error / 409.
- Drain surface: bridge_poll gains optional target; REST GET /sessions/{id}/replies.
- Rendezvous left untouched. QUESTION path (bridge_ask) NOT queued (interactive,
keeps NO_WAITER); completion/failure fallbacks NOT queued (captured-waiter).
- Bridged.main wires new InMemoryReplyInbox(); no broker: config yet (Stage 2 = AMQP).
Tests: +19 (188 -> 207), 0 failures/0 errors. New InMemoryReplyInboxTest (12) +
MessageService/BridgeMcp/BridgedApp coverage incl. guards proving a QUESTION and a
completion fallback are never queued.
Implemented via delegation to an off-sub gx10 worker in a pre-trusted worktree;
primary-verified (mvn clean install green, 207 tests) and committed by the primary
because the worker's completion replies were lost to the very bug this fixes.
Refs CB-307 (gitea #5), Stage 1 of 2.
The launcher now polls AgentControl.status(paneId) after starting the pane.
It returns the handle only once the worker reports an injectable state
(IDLE/BLOCKED/DONE). If the timeout elapses while still UNKNOWN, the
pane is self-reaped and a PeerUnreachableException is thrown — no orphan
left behind. The gate is disabled when spawnReadyTimeoutMs == 0 (legacy
non-blocking spawn, the default for the 6-arg constructor).
Key changes:
- PeerUnreachableException (new) in dev.ltms.bridged.peer
- BridgedConfig: spawnReadyTimeoutMs (default 20000), spawnReadyPollMs (default 300)
- ClaudeCodeLauncher: 3 constructor overloads:
(a) 6-arg backward-compat: gate disabled (timeout=0)
(b) 8-arg production: gate with config knobs + real clock/sleep
(c) 10-arg testability: full seam (LongSupplier clock + Runnable sleeper)
- waitUntilInjectableOrThrow() loop in spawn(SpawnRequest)
- sleepUninterruptibly() helper for the production sleeper
- BridgeMcp.spawn + BridgedApp.spawnWorker catch PeerUnreachableException
→ clean tool error / 502 response (not an uncaught 500)
- SessionManager.acquire inherently registers nothing on throw (both
worktree and non-worktree paths) — confirmed by new test
Tests:
- ClaudeCodeLauncherTest: 4 new tests
- unknown→idle: returns handle, no pane.close
- always-unknown: throws PeerUnreachableException, pane closed,
clock advanced past timeout
- timeout=0 (6-arg ctor): no agent.get calls, returns handle
- timeout=0 (10-arg ctor): no orphan pane close
- SessionManagerTest: 1 new test
- acquire → PeerUnreachableException: roster remains empty
Total: 188 tests, all pass (no existing test changed semantics)
Name the first-class Claude Code adapter explicitly, per the Peer Launcher SPI:
WorkerService was the de-facto Claude-Code launcher; as an in-tree PeerLauncher impl
it should say so. Pure IDE rename (class + file + WorkerServiceTest) plus stale
Javadoc/comment mentions swept to the new name. No behaviour change.
Gate: IDE diagnostics 0/0 on touched files; mvn clean install BUILD SUCCESS,
MVN_EXIT=0, 183 tests pass. Deferral #1 from issue #3 cleared.
Design-only. Audits the as-built Claude/herdr coupling (concentrated in
WorkerService), defines a PeerLauncher SPI + opaque PeerHandle + capability
model so the bus delegates peer materialization to a config-selected adapter.
ClaudeCodeLauncher = adapted WorkerService. Stages A/B/C with the Stage-C
plugin-loading security gate called out. No production code touched.
The worker "checkpoint" is commit → push → open its own PR. Push is free over SSH
(same user, same keys); the only incremental grant is PR-create, so the daemon injects
a repo-scoped gitea token into the worker env — opt-in per profile, never mutating
bridged's own environment.
- BridgedConfig.Worker: gitTokenEnv/gitHostEnv fields (opt-in; gitHostEnv defaults to
GITEA_HOST). Backward-compat 12-arg constructor keeps pre-CB-302 call sites + YAML
working. hasGitToken() gates injection.
- WorkerService.spawn: inject GITEA_TOKEN (and paired GITEA_HOST) only when the profile
grants a token AND the host env resolves one. resolveEnv() tolerates unset var names.
- WorkerServiceTest: injection present for a granting profile; absent when not (proving
the gate is config, not a missing env var).
- .claude/skills/implementer/SKILL.md: worktree-aware playbook — confirm the worktree/
branch, implement, commit (never .mcp.json/wiki), push, open PR via gitea REST with
GITEA_TOKEN, hand off the PR URL via bridge_reply. Never merge; workers can't run IDE
diagnostics so never claim IDE-clean.
Whole-project gate: mvn clean install green, 164 tests, 0 failures.
Opt-in isolated worktree so parallel implementers don't stomp the shared
tree, hydrated to config parity so a worker differs from the primary only
in LLM provider.
- Worktrees seam (interface) behind SessionManager; GitWorktrees shells git
via ProcessBuilder (non-zero exit -> WorktreeException), FakeWorktrees for
tests. No live git in unit tests.
- acquire() 5-arg overload provisions add -> overlayParity -> spawn(cwd=wt)
-> register, unwinding the worktree on any failure before registration.
4-arg overload and shared-tree behavior unchanged (backward compatible).
- release() removes the checkout but never deletes the branch (it holds the
worker's commits + PR, CB-302).
- overlayParity copies local config (.mcp.json, settings.local.json, .env/
.envrc) into the worktree; tracked ones get --skip-worktree so a worker
can never stage the parity overlay.
- WorkerSession gains nullable worktree/branch; BridgedConfig.Worker gains
parityOverlay (default list) + top-level worktreeRoot.
- bridge_spawn / POST /workers gain an optional worktree(+ticket) arg; the
worker view includes worktree/branch only when non-null.
Verify fixes on the delegated impl: strip trailing dashes in slug()
(^-+|-+$, was ^-+|^-+$); make FakeWorktrees.add a pure fn of the branch
(nonce already unique); MCP worktreeRequest treats blank/"false" string as
no-worktree, matching the REST builder.
162 tests, 0 failures.
Adds dev.ltms.bridged.session with WorkerSession (immutable record) and
SessionManager wrapping WorkerService: a ConcurrentHashMap registry keyed by
paneId, the one-shot lifecycle FSM (SPAWNING->READY->BUSY->DONE, ->FAILED on
drop/turn-failure, ->RELEASED on teardown), ownership (ownerTerminal), and
recycle = release + fresh acquire (no-reuse invariant). Driven by TurnListener
(BUSY/DONE/FAILED) and a WorkerPresence bridge (READY).
Wiring: Bridged.main constructs it and composes it into the TurnListener
alongside CompletionResolver; bridge_spawn / POST /workers route through
acquire (carrying caller identity as owner); bridge_stop / DELETE /workers
route through release. WorkerService gains effectiveCwd(); WorkerPresence
de-finalized so the manager can present a READY-driving view.
asPresence() returns a single cached bridge (a fresh one per call would
fragment the shared present set). roster() is the registry snapshot; the live
herdr join is left for CB-304. 6 fake-based acceptance tests; full suite green
(155/155).
Delegated to an off-subscription worker against docs/CB-301-Session-Manager.md;
primary verified (ide diagnostics clean, mvn clean install green) + fixed the
asPresence caching bug.
Worktree is code-only isolation; SessionManager hydrates it to full config
parity (overlay untracked local settings/.mcp.json/.env) so a worker is a
full peer of the primary, differing only in the LLM provider. Worker commits,
pushes over SSH, and opens its own PR (gitea REST + repo-scoped token).
The as-built audit surfaced that WorkerService keeps no registry of what it
spawned (its own Javadoc: 'there is no registry; list() only asks herdr').
CB-301 adds a SessionManager wrapping WorkerService: an authoritative in-daemon
roster with a per-session lifecycle FSM (SPAWNING/READY/BUSY/DONE/RELEASED/
FAILED), deterministic release, and recycle (= release + fresh acquire, no
reuse). Leaves clean seams for CB-302 (checkpoint on release), CB-303 (idle_ttl/
context_cap/drain policy over roster), CB-304 (bridge_list reads roster).
A worker's tool execution is single-threaded, so two bridge_ask calls from the
same session can only be a transport retry — yet openAsk minted a fresh turnId
each time and both raced to resolve the primary's single forward waiter, the
loser returning NO_WAITER and leaving a dangling turn (observed as turnId #2 in
the live e2e). Now openAsk is idempotent per session: a second open ask coalesces
onto the existing turnId + answer future (fresh=false), and only the fresh owner
surfaces the question and tears the turn down. closeAsk clears the per-session
index (conditional by value) so a later ask reopens fresh.
Implemented by an off-subscription worker delegated over the bridge; reviewed and
validated on the primary (IDE-clean, mvn clean install 149 green, +2 tests:
concurrent double-ask coalescing + post-close reopen).
The playbook a reviewer worker loads when the lead delegates a scoped review:
read the whole scope before judging, stay in the assigned lane, ask the lead
via bridge_ask when the call is genuinely theirs (resuming the same turn with
the answer), and report exactly one structured finding via bridge_reply. Pairs
the already-shipped bridge_reply/bridge_ask tools with the role guidance that
tells a worker how to use them. Mermaid validated with mmdc.
Drives the reverse path end to end over REST loopback: a worker is delegated
a task it cannot finish without asking, calls bridge_ask mid-turn, and the
primary answers on the surfaced turnId so the worker resumes the SAME turn.
Two blocking sends, no polling. Verified live: worker asked in ~9s, resumed
and replied CHOSEN=BLUE via clean bridge_reply after the primary answered.
Subscription-safe by construction (REST face only; never sets ANTHROPIC_BASE_URL).
A worker can now pause its delegated turn to ask the primary a question and
resume the same turn with the answer — the reverse of bridge_send.
- Rendezvous: QUESTION kind carrying a turnId, plus a reverse-ask registry
(openAsk/resolveQuestion/askSession/answerAsk/closeAsk).
- MessageService.ask(): surface a worker's question to the primary's open send,
block for the answer. answer(): resolve the worker's ask by turnId, then block
for its eventual bridge_reply (session derived from turnId, not an argument).
- BridgeMcp: bridge_ask tool (worker-only, identity from the connection);
bridge_send routes a turnId to the answer path. Shared formatReply().
- BridgedApp REST parity: POST /sessions/{id}/ask, turnId on /message.
- Lightweight CB-201: the kind vocabulary is the QUESTION outcome + turn_id
correlation, not a rigid from/to/corr envelope (the connection-identity
mechanism already covers addressing more robustly).
Tests: MessageServiceTest ask/answer round-trip + NO_WAITER/timeout/stale;
new RendezvousTest for the reverse registry; BridgeMcp ask/answer parity.
mvn: 147 green.
captureBaseline stored the raw, unclipped last-assistant block while resolve()
compares against clip(...) capped at MAX_SCRAPE_CHARS. For a block longer than
4000 chars the two capped representations never match even when the pane is
unchanged, defeating the CB-115 misattribution guard and letting a stale
completion resolve a rapid back-to-back send. Clip the baseline identically.
Regression test: an unchanged >cap block stays suppressed.
Surfaced by the fan-out issue-hunt E2E (1 primary -> 3 concurrent workers,
e2e/issue_hunt_test.py, added here). The same hunt's WorkerService.stop() and
Rendezvous.complete() findings were verified as false positives (locatePane is
already guarded; the sender's finally-close already removes the waiter).
Closes#2
herdr keeps worker panes alive across a daemon restart by design, and a
worker's paneId is held only by its spawner — so a worker whose owning
process exited before its DELETE leaks with nothing tracking it (there is
no registry; list() only asks herdr). Observed as three idle claude-ollama
panes left in the worker space from earlier runs.
On boot, WorkerService.reapOrphanWorkers() scans herdr for agents whose
name matches our claude-<profile>-<nonce>-<seq> scheme with a nonce other
than this process's nameNonce, and tears each down (pane + its now-empty
dedicated tab). A current-nonce worker is ours and live (spared); a user's
own claude session carries no such name (untouched). Keyed on the nonce so
it survives kill -9 and reaps a *previous* daemon's leaks — the actual case
shutdown-hook reaping and an in-memory registry both miss.
- Agent now projects herdr's 'name' (was dropped) so the reaper can key on it.
- isForeignWorker/workerNonce are pure + package-private for unit testing.
- FakeHerdr.withAgent seeds named agents into agent.list.
- Wired best-effort into Bridged startup before serving.
Closeslms/claude-bridge#1
A repeatable multi-turn primary↔worker conversation driven entirely through the
bridge's loopback REST face (async fire-and-poll) — never sets ANTHROPIC_BASE_URL
and never touches herdr, so it is subscription-safe by construction. Records every
turn to a transcript, grades each (OK / DEGRADED / EMPTY / FAILED / WEDGE), and
exits non-zero if any turn fails to deliver-and-reply, so it is CI-usable.
This is the repeatable form of the ad-hoc channel test that surfaced the CB-115 and
CB-116 gaps.
A 5-turn primary↔worker conversation test (see the e2e harness) surfaced three
delegation-channel gaps; this closes them.
CB-115 — status + scrape correctness:
- AgentStatus gains DONE (herdr's explicit turn-complete marker) so a finished
turn is no longer misread as UNKNOWN and left to wedge or false-fail.
- StatusRefiner reclassifies content-bearing UNKNOWN samples (StatusPoller wired
to it), and the Injector baselines pane content on delivery (TurnListener gains
onDelivered) to guard completion against previous-turn misattribution.
- CompletionResolver.lastAssistantBlock stops at the first hard TUI boundary, so a
scrape returns only the assistant answer — no input box, prompt echo, spinner,
tips or warnings.
CB-116 — waiter identity (the cross-turn stale reply):
- The completion/failure fallback ran on a virtual thread and resolved whichever
waiter was currently registered for the session. Since the rendezvous holds one
waiter per session and sends serialize, turn N's late completion could land on
turn N+1's waiter and deliver turn N's stale scrape as turn N+1's answer. The
baseline guard missed it because turn N was resolved by bridge_reply, which never
updates the completion baseline.
- Fix: capture the exact waiter (and pre-turn baseline) when a turn is delivered,
on the poller thread before any next-turn delivery can overwrite it, and resolve
THAT waiter — a no-op if it was already resolved. Rendezvous.resolveCompletion/
resolveFailure now take the captured CompletableFuture; currentWaiter exposes the
registered one for capture. A late completion for turn N can no longer touch turn
N+1's send.
Verified: 128 unit tests green (incl. a CB-116 regression asserting a late
completion never resolves the next turn's waiter); a re-run of the conversation
test passes with turn 5 resolving to its own reply rather than turn 4's text.
The Status section still read 'Design selected' — but bridged is built and
dogfooded (CB-101..114, 105 tests, live MCP tools, multi-profile, cwd inheritance,
readiness gate). Replace it with an accurate shipped/next breakdown, and drop the
bridge_ask overclaim from the gateway bullet (bridge_ask is roadmap, not built).
Final review-sweep finding: firstNonBlank(requestedCwd, cfg.cwd(), callerCwd,
user.dir) returns null if all are blank (pathological env with user.dir unset),
after which AgentControl drops the cwd and herdr defaults the pane to $HOME —
violating CB-112's documented contract. Append "." (the daemon's own cwd) as a
guaranteed non-blank last resort. Near-impossible trigger; makes the code honor its
own javadoc. 105 tests green.
An off-sub worker's review of CB-113 (delegated through the bridge) surfaced two
real gaps in the readiness gate:
- A worker that herdr reports idle but whose Claude never connects the bridge MCP
(crashed during boot, or wedged on a startup prompt) left its message queued
forever: ready.test() never passed, the target was polled indefinitely, and the
caller's future never completed (async waiter hung for the full 30-min window).
Injector now counts injectable-but-not-ready samples and, after a ~60s grace
(READINESS_GRACE_POLLS, deliberately longer than the UNKNOWN stall grace since a
first boot is slower than an in-turn blip), fails the queued messages, fires
onTurnFailed so blocking/async waiters resolve WORKER_FAILED, and reclaims the
target. Mirrors the CB-109 UNKNOWN-stall path.
- WorkerPresence.forget had no caller, so a worker's readiness lingered past its
life. Injector now clears presence via a forget callback on drop() (pane crash)
and on the readiness timeout.
+3 InjectorTest cases (never-ready failure, ready-within-grace delivery, drop clears
presence). 105 tests green.
Delegating right after spawn failed: herdr reports 'idle' during the worker's
boot, so the injector delivered into a not-ready TUI (paste lost) and wedged the
worker. Two fixes:
- Readiness gate: a worker is 'available' only once its Claude connects the bridge
MCP (the daemon observes it via peer-PID->terminal). WorkerPresence tracks it; the
injector holds the first delivery until present, so it never pastes into the boot
window. Exposed as 'ready' on GET /sessions/{id}/status.
- Submission nudge: the Enter accompanying a delivery can race the paste (esp. right
as the TUI becomes ready), leaving text unsubmitted. While a delivered message
stays idle (not picked up), the injector re-sends Enter each poll until the worker
starts (WORKING) or the grace expires.
Validated live: spawn + immediate delegate now holds during boot (ready=false),
delivers on MCP-connect, re-nudges Enter, worker replies. 102 tests green.
A worker now opens the same directory the primary is in, unless told otherwise.
Resolution: explicit spawn cwd → per-profile config cwd → the primary's cwd
(auto-detected from the bridge_spawn caller's PID via lsof -d cwd) → the daemon's
cwd. Never $HOME.
Mechanism (found by live probe, corrects the earlier assumption): an agent.start
pane does NOT inherit its tab's or workspace's cwd — it starts in $HOME. herdr's
agent.start honours an (undocumented) cwd param, so the resolved cwd is threaded
onto agent.start {cwd} (both tab and pane placement), not tab.create.
Surfaces: bridge_spawn {cwd?} + auto-detect via ConnectionIdentity.resolve (peer
PID) + ProcessCwdLookup (lsof); REST POST /workers ?cwd= / body cwd; per-profile
'cwd:' config. Validated live: explicit cwd → worker rooted there; no cwd over
REST → daemon cwd, not $HOME. Also clears the ccs folder-trust prompt when the
project dir is already trusted (see docs/Worker-Startup-and-Trust.md).
Documents (a) the rule that a worker inherits the PRIMARY's working directory,
never $HOME — resolved from an explicit cwd, else the bridge_spawn caller's PID
(lsof -d cwd), else the daemon cwd; herdr's seam is workspace.create {cwd}, since
agent.start has no cwd; (b) ccs (Claude Code) as the assumed launcher and its
folder-trust model (hasTrustDialogAccepted per project in each instance's
.claude.json), so trusting the project dir once per profile clears the prompt;
(c) that other CLIs have their own startup gates, documented per launcher.
Marks the cwd-inheritance as target design (follow-up), not yet wired.
The bridge was hard-wired to one worker profile. Config now takes a 'workers'
map keyed by profile name plus 'defaultWorker'; WorkerService holds the map and
gains spawn(profile) (spawn() uses the default). Selection threads through the
surfaces: REST POST /workers ?profile= / {"profile":…} + GET /profiles; MCP
bridge_spawn {profile?} + new bridge_profiles. Each profile's base_url is
guard-checked independently, so gx10 and ollama can run side by side and you
address each worker by its returned sessionId. Backward-compatible: the legacy
singular 'worker:' block still loads as a one-entry profile map.
Adds the drop test for the delivered/awaitingPickup/!turnObserved state (worker
vanished after delivery but before a WORKING sample) — a gap the other two drop
tests missed. Surfaced by an off-sub worker's code review of CB-110, delegated
through the bridge itself.
Companion to CB-109. When a worker disappears mid-turn (pane crash → herdr
*_not_found), the status poller drops the target, which failed only *queued*
messages — a message already DELIVERED is out of the queue, so its send's
rendezvous was left hanging until the 30-min async timeout. Injector.drop now
fires onTurnFailed for a delivered-but-unresolved turn (awaitingCompletion), so
the send resolves as WORKER_FAILED. Reuses the CB-109 resolver path; the failure
reason is neutral to cover both wedge (stuck) and vanish (gone).
Dogfood found the gap: a turn that dies into an error screen herdr reports as
'unknown' (e.g. the worker hitting API ENOTFOUND) never produces a working->idle
boundary, so CB-106 never fires and the async send rides its full 30-min timeout.
The injector now counts consecutive 'unknown' samples while a delegation is
outstanding; any working/idle sample resets the streak, so only a genuine wedge
(~30s continuous unknown) trips it. It then fires TurnListener.onTurnFailed;
CompletionResolver scrapes the error screen and resolves the send via
Rendezvous.resolveFailure (Kind.FAILED -> Outcome.WORKER_FAILED), surfaced as
async phase=failed / REST status=failed / a [worker failed] MCP note, with the
error context as the reason. This also frees a delivery that wedged before pickup,
which the injectable-only pickup grace could never release.