Commit Graph

49 Commits

Author SHA1 Message Date
kevin d67d30c58a CB-508: let an opencode profile pin its own OpenAI-compatible endpoint
CI / build (push) Successful in 1m57s
Points an opencode worker at a local vLLM (or llama.cpp / LM Studio / TGI)
instead of opencode's own gateway. opencode has no ANTHROPIC_BASE_URL seam, so
this could not be a config-only change: setting baseUrl on a kind: opencode
profile now makes the launcher emit a custom `provider` block into the
generated opencode.json, using @ai-sdk/openai-compatible.

The provider id comes from the provider half of the model: selector, so one
field drives both the generated declaration and the -m flag and the two cannot
drift apart. A bare model name with a baseUrl set is rejected at spawn with a
message saying how to fix it — silently falling back to the default gateway
would leave a worker talking to the wrong LLM while looking perfectly healthy.

A bare host:port gets /v1 appended (where these servers mount the API); a URL
that already carries a path is used verbatim. tokenEnv, when set, becomes the
provider apiKey; local servers generally ignore it but the AI SDK requires a
non-empty value, so a placeholder is used otherwise.

Two supporting changes:
- writeConfig previously ran only when a bridge MCP url was set. A pinned
  endpoint needs the config file too, so it now runs when either applies, and
  the mcp/instructions half is emitted conditionally.
- The config is now built with Jackson instead of string concatenation. The
  provider block is nested and interpolates operator-supplied values (URL,
  model id, api key), so escaping has to be real rather than a hand-rolled
  two-character replace.

No guard entry is required even with baseUrl set. SubscriptionGuard exists to
stop a worker borrowing the primary's Anthropic subscription, and an opencode
process has no Anthropic credential path at all — the asymmetry with the Claude
adapter reusing the same field is deliberate and documented at the call site.

Also fixes a brittle assertion in the existing MCP-mount test, which matched
the substring "\"type\": \"remote\"" and broke on Jackson's spacing. It now
parses the generated JSON and asserts on structure; whitespace is the
formatter's business, not the contract's.

318 tests (was 311): 5 new covering provider generation, /v1 normalisation,
path-preserving URLs, the missing-prefix rejection, MCP+provider coexistence,
and that no baseUrl still means no provider block.

Verified live end to end: daemon restarted on this build, worker spawned on the
opencode-local profile, generated config carries baseURL
http://127.0.0.1:8000/v1, and a blocking bridge_send returned
{"reply":"LOCAL-OK","replySource":"reply"} — a structured reply, not the
completion fallback. The worker pane reports
"Build · deepseek-v4-flash local-vllm (bridged)", confirming traffic reached
the local server rather than silently falling back.
2026-08-01 20:03:00 +07:00
kevin 6804676a96 CB-507: fix NPE on a worktree spawn with no cwd (HTTP 500 over REST)
POST /workers?worktree=true returned HTTP 500 with a NullPointerException out
of ProcessBuilder.start(): acquireWithWorktree resolved the repo root from
firstNonBlank(requestedCwd, callerCwd), and a plain REST spawn supplies
neither (BridgedApp hardcodes callerCwd=null, "no MCP caller over REST"). Both
null yielded null, putting `git -C null rev-parse --show-toplevel` on the
command line.

Now resolved through launcher.effectiveCwd, the CB-112 chain used everywhere
else (requested -> profile cwd -> caller -> daemon cwd -> "."), which is
documented never to return null. The non-worktree path in this same class
already went through it; only the worktree branch was missed.

Also fixes a second latent bug in the same line: firstNonBlank never consulted
the profile's configured cwd:, so a worktree spawn silently ignored a pinned
per-profile working directory. effectiveCwd honours it.

Removes firstNonBlank, now dead (this was its only call site) — javac ignores
an unused private method but IDE inspections flag it, and CLAUDE.md requires a
clean bill.

Why 311 tests missed it: the null/null case only arises over REST, and
WorktreeSessionManagerTest always passes an explicit cwd. Over MCP callerCwd is
populated from the caller PID, so the feature worked there. This is the third
REST-vs-MCP divergence found this month, after CB-505's path-trusted session id.

The one-line change was implemented by an opencode-free worker over the bridge
in an isolated worktree (branch worker/cb-507-worktree-cwd-npe-3e9c3b-3); the
dead-helper cleanup and the explanatory comment were added on integration.
A regression test is still outstanding and is being delegated separately.
2026-08-01 19:43:08 +07:00
kevin 19cdf8dc9f CB-505 fix: audit lines were not valid JSON
The first cut spliced the timestamp on via a logback pattern:

    {"ts":"%d{...}",%replace(%msg){'^\{',''}%n

Logback's variable substitution chokes on the literal braces
("All tokens consumed but was expecting }"), so the encoder failed to
configure. Caught by running the real jar and noticing logback had dumped its
internal status — which it only does when something failed to parse. The build
was green throughout: nothing asserted the audit trail was machine-readable.

AuditLog now emits the complete object including its own ISO-8601 "ts", and the
appender pattern is a bare %msg. Adds AuditLogTest, which parses each emitted
line with Jackson (so a malformed record fails the build) and pins that hostile
ids cannot escape their field to forge a second record.

311 tests green; logback now configures with zero internal errors.
2026-07-29 22:32:18 +07:00
kevin 9daf1ec5ba CB-5xx: Stage 5 hardening — auth, authz+audit, metrics, CI, supervision
Closes out single-host before the cross-host work. Sequenced BEFORE CB-308
deliberately: federation's own gating concern is the trust model, and it
inherits whatever identity shape lands here.

The finding this stage is built around: bridged had exactly ONE security
control, the loopback bind. ConnectionIdentity resolves a worker from its
connection (unforgeable), but every caller that was not a recognised worker
pane fell through to being treated as the PRIMARY -- the most privileged role
on the bus. Latent today; load-bearing the moment a bind widens.

CB-501 auth:
- Role/Principal/CallerResolver: connection identity first, bearer token
  second, ANONYMOUS third. Inverts the old default so absence of identity
  means nothing, not everything.
- Worker identity is never token-gated, so enabling auth cannot lock the
  fleet out of bridge_reply.
- Constant-time token compare (MessageDigest.isEqual).
- validateAuthExposure(): a non-loopback bind under loopback-trust now
  REFUSES TO START. Makes the dangerous config unrepresentable rather than
  merely documented.
- TLS terminates at a reverse proxy by design (D3), not in the JVM.

CB-505 authz + audit, enforced on BOTH entry paths:
- The docs describe MCP as "a thin adapter over the REST core"; at code level
  it is not. BridgeMcp calls MessageService directly, and /mcp is a raw
  servlet on Jetty's context handler that never traverses Javalin's before
  filter. Enforcing only at REST would have left /mcp open.
- Load-bearing rule is own-session-only: a worker may reply/ask only as
  itself. Structurally true over MCP already; over REST the session id in the
  URL path had simply been trusted.
- Audit: JSON lines to a dedicated appender, additivity=false. Never records
  message content -- this bus carries source and prompts.

CB-502 metrics: zero new dependencies. A ~150-line Prometheus text renderer
instead of the specced Micrometer, because this pom already hand-pins
jackson-annotations to reconcile Jackson 2/3, imports a Jetty BOM against
skew, and carries four accepted-CVE advisories -- and CLAUDE.md's mandated
dependency CVE gate could not be run (no JetBrains MCP server connected).
Instrumented at MessageService, the single funnel both surfaces share.

CB-503 CI: .gitea/workflows/ci.yml against the already-running Gitea runner.
Needs no contract-exclusion flag -- the pom's default-excludes profile
already sets excludedGroups=contract, so plain `mvn clean install` IS the
mock-socket surface. Provisions JDK 25 explicitly (runner default-jdk is older).

CB-504 supervision: launchd agent (the real target -- this host is macOS,
there is no systemd) plus a systemd unit for the Linux gateways CB-308 adds.
Ordering directives are advisory, so the actual fix is that startup now waits
up to 30s for the herdr socket and then serves degraded, instead of crashing
into a restart loop on a boot-order race.

Also fixes drift found while surveying:
- bridged.example.yaml documented spawn_ready_timeout_ms in snake_case; config
  binds via plain Jackson with ignoreUnknown, so uncommenting it would have
  been silently dropped and the default kept. Now camelCase, with a test that
  loads the shipped example and one that pins every documented knob's
  spelling -- no test had ever loaded that file.
- Added the 6 shipped-but-undocumented knobs (worktreeRoot, parityOverlay,
  gitTokenEnv, gitHostEnv, configDir, primary:).
- README "Next" listed bridge_ask and session lifecycle as upcoming; both
  shipped long ago.
- docs/CB-301-ext and docs/CB-402 status headers said "design"/"pre-
  implementation" for work already merged.

307 unit/acceptance tests green (was 266), mvn clean install BUILD SUCCESS.
Note: CLAUDE.md's per-file ide_diagnostics gate and the pom Mend.io CVE check
could not be run -- no JetBrains/intellij-index MCP server is connected this
session. mvn clean install is the only gate that ran.
2026-07-29 22:29:26 +07:00
Dai Ha 11f8709286 CB-402 Increment 4: CompositePeerLauncher — route the fleet by kind
Introduce the router the core holds when more than one adapter is
configured: one HerdrPeerLauncher per peer kind, dispatched by profile
(spawn/effectiveCwd/parityOverlay), by pane id (stop, via a spawn-time
owner map), and fanned out + combined for the fleet-wide queries
(list dedup by pane id, reap/caps union, profiles union). The ctor
rejects an empty adapter list and a profile two adapters both claim.

Wire it in Bridged.main: partition workerProfiles() by kind (claude-code
is the always-present default adapter; opencode is added when any profile
opts in) and front both with the composite. This lets BridgeMcp and
BridgedApp finally take the PeerLauncher SPI instead of a concrete
ClaudeCodeLauncher — the two (ClaudeCodeLauncher) casts in Bridged are
gone. list() elements are cast to herdr Agent at the point of the
herdr-specific roster view, where that assumption actually lives.

Add BridgedConfig.Worker.isClaudeCode()/isOpenCode() kind predicates
(the wiring uses isOpenCode; both are unit-tested). Drop the long-dead
'rendezvous' constructor param threaded into BridgeMcp and BridgedApp.

10 CompositePeerLauncherTest cases over two real adapters on one
FakeHerdr: profile routing (observed via the started agent's
claude-/opencode- name prefix), default resolution, unknown-profile
and duplicate-profile rejection, caps union, list dedup, reap sum,
and stop teardown. 266 tests green.
2026-07-22 05:43:28 +02:00
Dai Ha b034f105c0 CB-402 Increment 3: OpenCodeLauncher — the SPI-proving second adapter
A HerdrPeerLauncher subclass for opencode, a provider-agnostic terminal
coding agent. It reuses every line of shared base transport (tab/pane
placement, CB-306 readiness gate, unique naming + CB-117 reap, teardown,
listing, cwd) and diverges only in buildLaunch:

  - No subscription boundary: no ANTHROPIC_BASE_URL, no SubscriptionGuard
    (the guard is a Claude-private concern, not part of the SPI).
  - File-based MCP mount + instructions: writes an ephemeral opencode.json
    declaring the bridge as a remote MCP server + a reply-charter file under
    instructions, pointed at via OPENCODE_CONFIG (opencode has no inline
    --mcp-config / --append-system-prompt).
  - Model selected with -m provider/model, not an env var.
  - 'opencode' name prefix so reap matches opencode-* panes only.

configRoot is injectable so tests inspect the generated config/charter under
a @TempDir. 10 tests cover config content, model flag, git-token grant,
capabilities, reap predicate, both production ctors, and the readiness gate
(throws PeerUnreachable on timeout, reaps only the worker pane).
2026-07-22 05:30:46 +02:00
Dai Ha 6e37722383 CB-402 Increment 2: kind: discriminator on Worker profiles
Add a `kind` field to BridgedConfig.Worker — "claude-code" (default) or
"opencode" — the discriminator the CompositePeerLauncher will route spawn/reap
by so each adapter drives only its own peer kind. Normalised to lower-case;
blank/absent ⇒ claude-code, so every existing config and call site is
unchanged. argv now defaults to the kind's own binary (claude vs opencode)
rather than always `claude`, so an opencode profile never inherits the Claude
command.

kind is appended at the record tail; a new 14-arg back-compat constructor
(git fields, no kind) keeps the CB-302 call sites working, and the existing
12-arg constructor is untouched. Also drop the never-used Primary(String)
legacy constructor to keep the file warning-clean.

example.yaml documents the key and carries a commented opencode-gemini
profile. Tests cover default/normalisation/argv-defaulting. 245 tests green,
config files 0 IDE problems.
2026-07-22 05:20:04 +02:00
Dai Ha ffce30afa2 CB-402 Increment 1: extract HerdrPeerLauncher abstract base
Behaviour-preserving refactor ahead of the second-adapter work. All herdr
transport shared by any peer kind — tab/pane placement, the CB-306
spawn-readiness gate, unique naming + CB-117 orphan reap, teardown, listing,
cwd resolution, and the peer-neutral git-forge grant — moves into a new
abstract HerdrPeerLauncher (Template-Method base). ClaudeCodeLauncher becomes
a final subclass supplying only the two Claude-specific seams: the `claude`
name prefix and buildLaunch(), which encodes the subscription boundary
(ANTHROPIC_BASE_URL + SubscriptionGuard assert, inline --mcp-config and
--append-system-prompt reply charter).

The base owns the injectable clock + sleeper for the readiness gate; the poll
interval is baked into the sleeper, so the vestigial spawnReadyPollMs field is
dropped from the base and from the full testability constructor (the explicit
sleeper already encodes it). The 6-arg and 8-arg production constructors keep
their signatures; three full-ctor test sites drop the now-unused poll argument.

No behaviour change: 242 tests green, both refactored files 0 IDE problems.
2026-07-22 05:15:21 +02:00
Dai Ha d4c9704007 CB-307: primary-gate cleanup of push-loop (drop dead clock, IDE 0/0)
Primary verification pass over the delegated push-loop delivery:
- Remove the unused LongSupplier clock threaded into ReplyPushLoop
  (timing is the scheduler's; the field was never read) from the
  component, Bridged wiring, and both test call sites.
- Collapse the single-statement WAIT_BUSY switch arm (redundant block).
- Drop now-dead test scaffolding: the always-"idle" recordingClient
  param and unused AgentStatus/AtomicReference imports.

IDE diagnostics 0/0 on all changed files; mvn clean install green
(242 tests, 0 failures).
2026-07-19 11:03:34 +02:00
Dai Ha 7c252b5f5f CB-307 Increment 3: bridge_ack tool — per-msgId ack refinement
- MessageService.ackReply(target, msgId) delegates to inbox.ack
- BridgeMcp registers bridge_ack tool with target/msgId args
- Tests: valid/invalid args, ack surface via BridgeMcp
2026-07-19 10:21:06 +02:00
Dai Ha f756933879 CB-307 Increment 2: ReplyPushLoop — status-gated push loop (mechanism b)
- ReplyPushLoop: dedicated scheduled loop for nudging the primary
- Package-private decide() method for pure decision logic (unit-testable)
- Status-gated injection via AgentControl.status().injectable()
- Bounded reminders (cap + backoff), idempotent per target
- Wire into MessageService.reply after inbox.publish on no-waiter branch
- Wire into Bridged.main (constructor + shutdown hook)
- Primary config record updated with pushReminders/pushBackoffMs knobs
- Tests: decide() matrix, nudge injection, idempotency, cap enforcement
2026-07-19 10:16:24 +02:00
Dai Ha a1aecbf4fc CB-307 Increment 1: PrimaryRegistry + config + wiring
- PrimaryRegistry: thread-safe single-slot registry with pin support
- Primary config record (last positional, like Broker)
- Wire capture in BridgeMcp (bridge_send and bridge_spawn handlers)
- Construct PrimaryRegistry in Bridged.main
- Tests: PrimaryRegistryTest + BridgedConfigTest primary config cases
2026-07-19 09:59:02 +02:00
Dai Ha 2bc5f3a057 CB-307 Stage 2: AmqpReplyInbox — durable, cross-restart reply delivery
Behind the existing ReplyInbox port, add an AMQP-backed adapter selected by a
`broker:` block in config (absent → the in-memory soft-state inbox; present →
AMQP). Mapping is consume-and-hold with deferred manual ack: each target owns a
durable queue `agent.<target>.inbox`; a manual-ack consumer pulls persistent
messages into an in-memory held map (dedup by msgId) but does not ack; peek
returns the snapshot; ack acks the broker delivery-tag and drops it. A crash
before caller-ack leaves messages unacked, so the broker redelivers on
reconnect — genuine durability with the port contract preserved. bridged still
owns no persistence; the broker does.

- msg/AmqpReplyInbox: the adapter (single synchronized channel; recovery
  listener clears held on reconnect so fresh delivery-tags repopulate).
- config/BridgedConfig: nullable Broker(uri) record; isConfigured() gates it.
- Bridged.main: select adapter; close the AMQP connection in the ordered
  shutdown hook (no-op for the in-memory inbox).
- deps: com.rabbitmq:amqp-client (main); testcontainers rabbitmq/junit-jupiter
  (test). Pinned commons-compress 1.27.1 + commons-lang3 3.18.0 to clear the
  test-scope CVEs those pull. Production default LavinMQ; RabbitMQ URI-swap.
- tests: BridgedConfigTest broker-selection cases (hermetic); AmqpReplyInbox
  contract test (@Tag("contract"), Testcontainers RabbitMQ) proving
  publish/peek/ack, msgId dedup, and cross-restart redelivery. Excluded from
  the default build so `mvn clean install` stays hermetic (210 green).
2026-07-19 07:30:29 +02:00
Dai Ha ba6b4a5da9 CB-307 Stage 1: reply-inbox port + in-memory adapter — hold stranded worker replies instead of dropping them
Problem: the reverse (worker->primary) path was Rendezvous, a map of LIVE blocking
waiters only. A bridge_reply arriving with no open send hit Rendezvous.complete()
-> no waiter -> returned false -> the reply was silently DISCARDED (worker saw an
error / REST 409). No message-id/dedup/ack existed anywhere.

Stage 1 (no broker, soft-state) behind one port:
- ReplyInbox port + InboxMessage record; InMemoryReplyInbox adapter (per-target
  FIFO via LinkedHashMap, dedup by msgId, thread-safe). Soft-state, not persistence.
- MessageService.reply(session, content): resolve an open send, else publish to the
  inbox with a minted UUID (was a silent drop). drainReplies(target) = peek + ack.
- BridgeMcp.reply / BridgedApp.replyMessage repointed off bare Rendezvous.resolve
  onto messages.reply -> no-waiter is now SUCCESS (queued), not error / 409.
- Drain surface: bridge_poll gains optional target; REST GET /sessions/{id}/replies.
- Rendezvous left untouched. QUESTION path (bridge_ask) NOT queued (interactive,
  keeps NO_WAITER); completion/failure fallbacks NOT queued (captured-waiter).
- Bridged.main wires new InMemoryReplyInbox(); no broker: config yet (Stage 2 = AMQP).

Tests: +19 (188 -> 207), 0 failures/0 errors. New InMemoryReplyInboxTest (12) +
MessageService/BridgeMcp/BridgedApp coverage incl. guards proving a QUESTION and a
completion fallback are never queued.

Implemented via delegation to an off-sub gx10 worker in a pre-trusted worktree;
primary-verified (mvn clean install green, 207 tests) and committed by the primary
because the worker's completion replies were lost to the very bug this fixes.

Refs CB-307 (gitea #5), Stage 1 of 2.
2026-07-18 21:12:51 +02:00
Dai Ha 7dd6c46156 CB-306: spawn-readiness gate — ClaudeCodeLauncher blocks until the worker is injectable or throws PeerUnreachableException
The launcher now polls AgentControl.status(paneId) after starting the pane.
It returns the handle only once the worker reports an injectable state
(IDLE/BLOCKED/DONE). If the timeout elapses while still UNKNOWN, the
pane is self-reaped and a PeerUnreachableException is thrown — no orphan
left behind. The gate is disabled when spawnReadyTimeoutMs == 0 (legacy
non-blocking spawn, the default for the 6-arg constructor).

Key changes:
- PeerUnreachableException (new) in dev.ltms.bridged.peer
- BridgedConfig: spawnReadyTimeoutMs (default 20000), spawnReadyPollMs (default 300)
- ClaudeCodeLauncher: 3 constructor overloads:
  (a) 6-arg backward-compat: gate disabled (timeout=0)
  (b) 8-arg production: gate with config knobs + real clock/sleep
  (c) 10-arg testability: full seam (LongSupplier clock + Runnable sleeper)
- waitUntilInjectableOrThrow() loop in spawn(SpawnRequest)
- sleepUninterruptibly() helper for the production sleeper
- BridgeMcp.spawn + BridgedApp.spawnWorker catch PeerUnreachableException
  → clean tool error / 502 response (not an uncaught 500)
- SessionManager.acquire inherently registers nothing on throw (both
  worktree and non-worktree paths) — confirmed by new test

Tests:
- ClaudeCodeLauncherTest: 4 new tests
  - unknown→idle: returns handle, no pane.close
  - always-unknown: throws PeerUnreachableException, pane closed,
    clock advanced past timeout
  - timeout=0 (6-arg ctor): no agent.get calls, returns handle
  - timeout=0 (10-arg ctor): no orphan pane close
- SessionManagerTest: 1 new test
  - acquire → PeerUnreachableException: roster remains empty
Total: 188 tests, all pass (no existing test changed semantics)
2026-07-18 16:17:43 +02:00
Dai Ha 3aa69a9e32 CB-401 Stage A follow-up: rename WorkerService -> ClaudeCodeLauncher
Name the first-class Claude Code adapter explicitly, per the Peer Launcher SPI:
WorkerService was the de-facto Claude-Code launcher; as an in-tree PeerLauncher impl
it should say so. Pure IDE rename (class + file + WorkerServiceTest) plus stale
Javadoc/comment mentions swept to the new name. No behaviour change.

Gate: IDE diagnostics 0/0 on touched files; mvn clean install BUILD SUCCESS,
MVN_EXIT=0, 183 tests pass. Deferral #1 from issue #3 cleared.
2026-07-18 14:39:56 +02:00
Dai Ha e056c7e1fa CB-401: Stage A - extract PeerLauncher SPI in-tree 2026-07-18 07:48:55 +02:00
Dai Ha 0efb65ca0a CB-303: session lifecycle limits — idle_ttl reaper, context_cap, graceful drain
Verified on primary: ide_diagnostics clean (incl. weak warnings), mvn clean install
BUILD SUCCESS, 173 tests. Delegated impl (worker/cb-303-80ec1a-3, 3 parts), primary-gated.
2026-07-17 10:08:55 +02:00
Dai Ha 09d3948acf CB-303 part 3: graceful drain on shutdown 2026-07-17 09:59:57 +02:00
Dai Ha 954351a80b CB-303 part 2: context_cap turn budget 2026-07-17 09:56:46 +02:00
Dai Ha 8d51066ddd CB-303 part 1: idle_ttl session reaper (injectable clock + SessionReaper) 2026-07-17 09:52:40 +02:00
Dai Ha 84102baab4 CB-304: bridge_list roster + live herdr join (worktree/branch); add GET /workers 2026-07-17 09:50:12 +02:00
Dai Ha 64e70efdf1 CB-302: worker checkpoint — repo-scoped forge token injection + implementer skill
The worker "checkpoint" is commit → push → open its own PR. Push is free over SSH
(same user, same keys); the only incremental grant is PR-create, so the daemon injects
a repo-scoped gitea token into the worker env — opt-in per profile, never mutating
bridged's own environment.

- BridgedConfig.Worker: gitTokenEnv/gitHostEnv fields (opt-in; gitHostEnv defaults to
  GITEA_HOST). Backward-compat 12-arg constructor keeps pre-CB-302 call sites + YAML
  working. hasGitToken() gates injection.
- WorkerService.spawn: inject GITEA_TOKEN (and paired GITEA_HOST) only when the profile
  grants a token AND the host env resolves one. resolveEnv() tolerates unset var names.
- WorkerServiceTest: injection present for a granting profile; absent when not (proving
  the gate is config, not a missing env var).
- .claude/skills/implementer/SKILL.md: worktree-aware playbook — confirm the worktree/
  branch, implement, commit (never .mcp.json/wiki), push, open PR via gitea REST with
  GITEA_TOKEN, hand off the PR URL via bridge_reply. Never merge; workers can't run IDE
  diagnostics so never claim IDE-clean.

Whole-project gate: mvn clean install green, 164 tests, 0 failures.
2026-07-17 08:51:27 +02:00
Dai Ha 97ecc7136e CB-301-ext: per-worker git worktree + config-parity overlay
Opt-in isolated worktree so parallel implementers don't stomp the shared
tree, hydrated to config parity so a worker differs from the primary only
in LLM provider.

- Worktrees seam (interface) behind SessionManager; GitWorktrees shells git
  via ProcessBuilder (non-zero exit -> WorktreeException), FakeWorktrees for
  tests. No live git in unit tests.
- acquire() 5-arg overload provisions add -> overlayParity -> spawn(cwd=wt)
  -> register, unwinding the worktree on any failure before registration.
  4-arg overload and shared-tree behavior unchanged (backward compatible).
- release() removes the checkout but never deletes the branch (it holds the
  worker's commits + PR, CB-302).
- overlayParity copies local config (.mcp.json, settings.local.json, .env/
  .envrc) into the worktree; tracked ones get --skip-worktree so a worker
  can never stage the parity overlay.
- WorkerSession gains nullable worktree/branch; BridgedConfig.Worker gains
  parityOverlay (default list) + top-level worktreeRoot.
- bridge_spawn / POST /workers gain an optional worktree(+ticket) arg; the
  worker view includes worktree/branch only when non-null.

Verify fixes on the delegated impl: strip trailing dashes in slug()
(^-+|-+$, was ^-+|^-+$); make FakeWorktrees.add a pure fn of the branch
(nonce already unique); MCP worktreeRequest treats blank/"false" string as
no-worktree, matching the REST builder.

162 tests, 0 failures.
2026-07-17 06:46:20 +02:00
Dai Ha 54d907c314 CB-301: SessionManager — authoritative worker session registry + one-shot FSM
Adds dev.ltms.bridged.session with WorkerSession (immutable record) and
SessionManager wrapping WorkerService: a ConcurrentHashMap registry keyed by
paneId, the one-shot lifecycle FSM (SPAWNING->READY->BUSY->DONE, ->FAILED on
drop/turn-failure, ->RELEASED on teardown), ownership (ownerTerminal), and
recycle = release + fresh acquire (no-reuse invariant). Driven by TurnListener
(BUSY/DONE/FAILED) and a WorkerPresence bridge (READY).

Wiring: Bridged.main constructs it and composes it into the TurnListener
alongside CompletionResolver; bridge_spawn / POST /workers route through
acquire (carrying caller identity as owner); bridge_stop / DELETE /workers
route through release. WorkerService gains effectiveCwd(); WorkerPresence
de-finalized so the manager can present a READY-driving view.

asPresence() returns a single cached bridge (a fresh one per call would
fragment the shared present set). roster() is the registry snapshot; the live
herdr join is left for CB-304. 6 fake-based acceptance tests; full suite green
(155/155).

Delegated to an off-subscription worker against docs/CB-301-Session-Manager.md;
primary verified (ide diagnostics clean, mvn clean install green) + fixed the
asPresence caching bug.
2026-07-16 19:30:56 +02:00
Dai Ha 4aa9d03da2 CB-205: coalesce duplicate bridge_ask calls onto one turn
A worker's tool execution is single-threaded, so two bridge_ask calls from the
same session can only be a transport retry — yet openAsk minted a fresh turnId
each time and both raced to resolve the primary's single forward waiter, the
loser returning NO_WAITER and leaving a dangling turn (observed as turnId #2 in
the live e2e). Now openAsk is idempotent per session: a second open ask coalesces
onto the existing turnId + answer future (fresh=false), and only the fresh owner
surfaces the question and tears the turn down. closeAsk clears the per-session
index (conditional by value) so a later ask reopens fresh.

Implemented by an off-subscription worker delegated over the bridge; reviewed and
validated on the primary (IDE-clean, mvn clean install 149 green, +2 tests:
concurrent double-ask coalescing + post-close reopen).
2026-07-16 16:27:10 +02:00
Dai Ha 358c6970b5 CB-205/CB-201: bridge_ask reverse rendezvous + lightweight question/turn kind
A worker can now pause its delegated turn to ask the primary a question and
resume the same turn with the answer — the reverse of bridge_send.

- Rendezvous: QUESTION kind carrying a turnId, plus a reverse-ask registry
  (openAsk/resolveQuestion/askSession/answerAsk/closeAsk).
- MessageService.ask(): surface a worker's question to the primary's open send,
  block for the answer. answer(): resolve the worker's ask by turnId, then block
  for its eventual bridge_reply (session derived from turnId, not an argument).
- BridgeMcp: bridge_ask tool (worker-only, identity from the connection);
  bridge_send routes a turnId to the answer path. Shared formatReply().
- BridgedApp REST parity: POST /sessions/{id}/ask, turnId on /message.
- Lightweight CB-201: the kind vocabulary is the QUESTION outcome + turn_id
  correlation, not a rigid from/to/corr envelope (the connection-identity
  mechanism already covers addressing more robustly).

Tests: MessageServiceTest ask/answer round-trip + NO_WAITER/timeout/stale;
new RendezvousTest for the reverse registry; BridgeMcp ask/answer parity.
mvn: 147 green.
2026-07-16 15:32:56 +02:00
Dai Ha 2a61fe69f1 CB-118: clip the completion baseline so the CB-115 guard survives >cap blocks
captureBaseline stored the raw, unclipped last-assistant block while resolve()
compares against clip(...) capped at MAX_SCRAPE_CHARS. For a block longer than
4000 chars the two capped representations never match even when the pane is
unchanged, defeating the CB-115 misattribution guard and letting a stale
completion resolve a rapid back-to-back send. Clip the baseline identically.

Regression test: an unchanged >cap block stays suppressed.

Surfaced by the fan-out issue-hunt E2E (1 primary -> 3 concurrent workers,
e2e/issue_hunt_test.py, added here). The same hunt's WorkerService.stop() and
Rendezvous.complete() findings were verified as false positives (locatePane is
already guarded; the sender's finally-close already removes the waiter).

Closes #2
2026-07-16 09:11:20 +02:00
Dai Ha ff6aacdc78 CB-117: reap orphaned worker panes on startup
herdr keeps worker panes alive across a daemon restart by design, and a
worker's paneId is held only by its spawner — so a worker whose owning
process exited before its DELETE leaks with nothing tracking it (there is
no registry; list() only asks herdr). Observed as three idle claude-ollama
panes left in the worker space from earlier runs.

On boot, WorkerService.reapOrphanWorkers() scans herdr for agents whose
name matches our claude-<profile>-<nonce>-<seq> scheme with a nonce other
than this process's nameNonce, and tears each down (pane + its now-empty
dedicated tab). A current-nonce worker is ours and live (spared); a user's
own claude session carries no such name (untouched). Keyed on the nonce so
it survives kill -9 and reaps a *previous* daemon's leaks — the actual case
shutdown-hook reaping and an in-memory registry both miss.

- Agent now projects herdr's 'name' (was dropped) so the reaper can key on it.
- isForeignWorker/workerNonce are pure + package-private for unit testing.
- FakeHerdr.withAgent seeds named agents into agent.list.
- Wired best-effort into Bridged startup before serving.

Closes lms/claude-bridge#1
2026-07-16 08:39:37 +02:00
Dai Ha 31e34d177b CB-115/CB-116: reliable turn completion — status refinement, clean scrape, waiter identity
A 5-turn primary↔worker conversation test (see the e2e harness) surfaced three
delegation-channel gaps; this closes them.

CB-115 — status + scrape correctness:
- AgentStatus gains DONE (herdr's explicit turn-complete marker) so a finished
  turn is no longer misread as UNKNOWN and left to wedge or false-fail.
- StatusRefiner reclassifies content-bearing UNKNOWN samples (StatusPoller wired
  to it), and the Injector baselines pane content on delivery (TurnListener gains
  onDelivered) to guard completion against previous-turn misattribution.
- CompletionResolver.lastAssistantBlock stops at the first hard TUI boundary, so a
  scrape returns only the assistant answer — no input box, prompt echo, spinner,
  tips or warnings.

CB-116 — waiter identity (the cross-turn stale reply):
- The completion/failure fallback ran on a virtual thread and resolved whichever
  waiter was currently registered for the session. Since the rendezvous holds one
  waiter per session and sends serialize, turn N's late completion could land on
  turn N+1's waiter and deliver turn N's stale scrape as turn N+1's answer. The
  baseline guard missed it because turn N was resolved by bridge_reply, which never
  updates the completion baseline.
- Fix: capture the exact waiter (and pre-turn baseline) when a turn is delivered,
  on the poller thread before any next-turn delivery can overwrite it, and resolve
  THAT waiter — a no-op if it was already resolved. Rendezvous.resolveCompletion/
  resolveFailure now take the captured CompletableFuture; currentWaiter exposes the
  registered one for capture. A late completion for turn N can no longer touch turn
  N+1's send.

Verified: 128 unit tests green (incl. a CB-116 regression asserting a late
completion never resolves the next turn's waiter); a re-run of the conversation
test passes with turn 5 resolving to its own reply rather than turn 4's text.
2026-07-16 08:27:38 +02:00
Dai Ha 3c05823f19 CB-114: resolveCwd never returns null (honor the 'daemon cwd, never $HOME' contract)
Final review-sweep finding: firstNonBlank(requestedCwd, cfg.cwd(), callerCwd,
user.dir) returns null if all are blank (pathological env with user.dir unset),
after which AgentControl drops the cwd and herdr defaults the pane to $HOME —
violating CB-112's documented contract. Append "." (the daemon's own cwd) as a
guaranteed non-blank last resort. Near-impossible trigger; makes the code honor its
own javadoc. 105 tests green.
2026-07-16 06:43:05 +02:00
Dai Ha 2fb46f670f CB-114: readiness-gate timeout + presence cleanup (delegated review findings)
An off-sub worker's review of CB-113 (delegated through the bridge) surfaced two
real gaps in the readiness gate:

- A worker that herdr reports idle but whose Claude never connects the bridge MCP
  (crashed during boot, or wedged on a startup prompt) left its message queued
  forever: ready.test() never passed, the target was polled indefinitely, and the
  caller's future never completed (async waiter hung for the full 30-min window).
  Injector now counts injectable-but-not-ready samples and, after a ~60s grace
  (READINESS_GRACE_POLLS, deliberately longer than the UNKNOWN stall grace since a
  first boot is slower than an in-turn blip), fails the queued messages, fires
  onTurnFailed so blocking/async waiters resolve WORKER_FAILED, and reclaims the
  target. Mirrors the CB-109 UNKNOWN-stall path.

- WorkerPresence.forget had no caller, so a worker's readiness lingered past its
  life. Injector now clears presence via a forget callback on drop() (pane crash)
  and on the readiness timeout.

+3 InjectorTest cases (never-ready failure, ready-within-grace delivery, drop clears
presence). 105 tests green.
2026-07-16 05:01:58 +02:00
Dai Ha 37f7ad4185 CB-113: reliable worker readiness gate + submission nudge
Delegating right after spawn failed: herdr reports 'idle' during the worker's
boot, so the injector delivered into a not-ready TUI (paste lost) and wedged the
worker. Two fixes:

- Readiness gate: a worker is 'available' only once its Claude connects the bridge
  MCP (the daemon observes it via peer-PID->terminal). WorkerPresence tracks it; the
  injector holds the first delivery until present, so it never pastes into the boot
  window. Exposed as 'ready' on GET /sessions/{id}/status.
- Submission nudge: the Enter accompanying a delivery can race the paste (esp. right
  as the TUI becomes ready), leaving text unsubmitted. While a delivered message
  stays idle (not picked up), the injector re-sends Enter each poll until the worker
  starts (WORKING) or the grace expires.

Validated live: spawn + immediate delegate now holds during boot (ready=false),
delivers on MCP-connect, re-nudges Enter, worker replies. 102 tests green.
2026-07-15 19:14:57 +02:00
Dai Ha 926724a279 CB-112: workers inherit the primary's working directory (not $HOME)
A worker now opens the same directory the primary is in, unless told otherwise.
Resolution: explicit spawn cwd → per-profile config cwd → the primary's cwd
(auto-detected from the bridge_spawn caller's PID via lsof -d cwd) → the daemon's
cwd. Never $HOME.

Mechanism (found by live probe, corrects the earlier assumption): an agent.start
pane does NOT inherit its tab's or workspace's cwd — it starts in $HOME. herdr's
agent.start honours an (undocumented) cwd param, so the resolved cwd is threaded
onto agent.start {cwd} (both tab and pane placement), not tab.create.

Surfaces: bridge_spawn {cwd?} + auto-detect via ConnectionIdentity.resolve (peer
PID) + ProcessCwdLookup (lsof); REST POST /workers ?cwd= / body cwd; per-profile
'cwd:' config. Validated live: explicit cwd → worker rooted there; no cwd over
REST → daemon cwd, not $HOME. Also clears the ccs folder-trust prompt when the
project dir is already trusted (see docs/Worker-Startup-and-Trust.md).
2026-07-15 16:33:46 +02:00
Dai Ha 7f1b6b3a0d CB-111: multi-profile workers — named backends selectable at spawn
The bridge was hard-wired to one worker profile. Config now takes a 'workers'
map keyed by profile name plus 'defaultWorker'; WorkerService holds the map and
gains spawn(profile) (spawn() uses the default). Selection threads through the
surfaces: REST POST /workers ?profile= / {"profile":…} + GET /profiles; MCP
bridge_spawn {profile?} + new bridge_profiles. Each profile's base_url is
guard-checked independently, so gx10 and ollama can run side by side and you
address each worker by its returned sessionId. Backward-compatible: the legacy
singular 'worker:' block still loads as a one-entry profile map.
2026-07-15 15:54:31 +02:00
Dai Ha a629a7ee73 CB-110: fail an in-flight delegation when its worker vanishes
Companion to CB-109. When a worker disappears mid-turn (pane crash → herdr
*_not_found), the status poller drops the target, which failed only *queued*
messages — a message already DELIVERED is out of the queue, so its send's
rendezvous was left hanging until the 30-min async timeout. Injector.drop now
fires onTurnFailed for a delivered-but-unresolved turn (awaitingCompletion), so
the send resolves as WORKER_FAILED. Reuses the CB-109 resolver path; the failure
reason is neutral to cover both wedge (stuck) and vanish (gone).
2026-07-15 15:23:17 +02:00
Dai Ha 2052929768 CB-109: fail a delegation whose worker wedges in an unknown state
Dogfood found the gap: a turn that dies into an error screen herdr reports as
'unknown' (e.g. the worker hitting API ENOTFOUND) never produces a working->idle
boundary, so CB-106 never fires and the async send rides its full 30-min timeout.

The injector now counts consecutive 'unknown' samples while a delegation is
outstanding; any working/idle sample resets the streak, so only a genuine wedge
(~30s continuous unknown) trips it. It then fires TurnListener.onTurnFailed;
CompletionResolver scrapes the error screen and resolves the send via
Rendezvous.resolveFailure (Kind.FAILED -> Outcome.WORKER_FAILED), surfaced as
async phase=failed / REST status=failed / a [worker failed] MCP note, with the
error context as the reason. This also frees a delivery that wedged before pickup,
which the injectable-only pickup grace could never release.
2026-07-15 15:16:55 +02:00
Dai Ha a97c287aee CB-108: fleet-management MCP tools (bridge_spawn / bridge_list / bridge_stop)
The primary could delegate to a worker but not create or reap one over MCP —
spawning was a raw REST POST /workers. BridgeMcp now adapts WorkerService so a
worker's whole lifecycle runs through MCP: bridge_spawn returns the new worker's
sessionId (for bridge_send) and paneId (for bridge_stop); bridge_list projects the
tracked workers; bridge_stop tears one down. The subscription boundary stays
enforced inside WorkerService (bridge_spawn surfaces a guard breach as a tool error
without touching herdr). Tools are thin static adapters, unit-tested by parity.
2026-07-15 14:22:50 +02:00
Dai Ha 8ed2370fbe CB-107: async fire-and-poll delegation (wait:false + ticket poll)
A caller's MCP client caps a blocking bridge_send at ~60s, but a real delegated
task runs for minutes. sendAsync runs the same blocking send on a background
virtual thread and returns a ticket; poll(ticket) reports pending/done/failed.
Async reuses the blocking path (and its per-target serialization), so it inherits
reply + completion resolution for free. Surfaces: REST POST message wait:false ->
202 {ticket} + GET /tasks/{ticket}; MCP bridge_send wait flag + new bridge_poll.
Terminal tickets are pruned after a TTL so the registry stays bounded.
2026-07-15 14:19:14 +02:00
Dai Ha 5b26caca0c CB-106: completion fallback — resolve a send when the worker's turn ends without bridge_reply
The blocking send previously resolved only on an explicit bridge_reply; a real
delegated task (edit files, run a build) finishes and goes idle without ever
calling it, so the send always timed out. The injector now reports a confirmed
working -> idle turn boundary via a TurnListener; CompletionResolver scrapes the
worker's transcript tail and resolves the awaiting send (Rendezvous.resolveCompletion,
Kind.COMPLETION -> Outcome.COMPLETED_UNREPLIED), surfaced as replySource=transcript
at the REST/MCP edges. Completion is synthesized only from a confirmed turn (a
sampled 'working'), never from the pickup-grace path, so it can't race the explicit
reply or fire on a turn that never ran.
2026-07-15 14:14:04 +02:00
Dai Ha 6988bfe88f CB-103: injector submits the prompt (Enter as a separate keystroke)
A delegated message was delivered into the worker's input box but never
submitted, so no task was ever processed: bridge_send blocked until timeout
while the worker sat idle with the prompt typed but not entered.

herdr's agent.send delivers text as a bracketed paste; a trailing carriage
return in that same call is swallowed as literal newline content, not Enter.
AgentControl.send now emits two keystroke events — the payload, then a
standalone "\r" — so the worker actually submits and runs the task.

Verified live (Claude Code v2.1.210, off-sub worker): full hands-off
bridge_send -> auto-submit -> worker computes -> bridge_reply round trip
returns the reply in ~41s. 62 tests green.
2026-07-15 09:40:18 +02:00
Dai Ha 07722a6007 CB-1xx: worker launch via ccs + inline bridge MCP mount (step 4) + port 8765
Spawn a real worker with 'ccs ltms-local' (profile sets CLAUDE_CONFIG_DIR + off-sub base_url; auto mode preconfigured as defaultMode:auto). bridged appends the bridge MCP mount (--mcp-config, inline JSON) and the reply charter (--append-system-prompt) as launch FLAGS — non-invasive, nothing written to the worker's profile (safer than provisioning its config dir, which would clobber it). Identity is connection-based so the mount is shared. config: worker.mcpUrl. Default bind port 8080 -> 8765. The reply charter is guidance; the send timeout catches a non-cooperative worker. 60 tests green, IDE-clean.
2026-07-15 09:19:31 +02:00
Dai Ha fa570ab32a CB-105: connection-based MCP caller identity (peer PID → herdr pane)
Resolve who is calling an MCP tool from the connection, not a spoofable argument (per docs/MCP-Contract.md). ConnectionIdentity ties the loopback peer PID (LsofPeerPidLookup) to a herdr pane (PaneLocator via pane.list/pane.process_info) → the caller's terminal_id; a caller owning no pane is the primary. bridge_reply now takes only content and resolves the worker from the connection (no sessionId). Live-verified: real PID→pane (contract test), and a non-pane MCP caller correctly gets a workers-only error. Single-host; the token path stays the split-host fallback. 58 tests green, IDE-clean.
2026-07-15 09:19:31 +02:00
Dai Ha 8ffbfd1f06 CB-105: MCP server (bridge_send/bridge_reply/bridge_status) over the REST core
Streamable-HTTP MCP server (io.modelcontextprotocol.sdk:mcp 2.0.0) mounted on the daemon's Jetty at /mcp, exposing three tools as thin adapters over MessageService/Rendezvous — the primary calls bridge_send/bridge_status, the worker calls bridge_reply. Tool logic in unit-testable static methods (parity tests); SDK owns the wire protocol. Resolves the Jackson 2/3 split by pinning jackson-annotations 3.0-rc5 (works for both our Jackson 2.19 and the SDK's Jackson 3). Live-verified: initialize handshake + tools/list return all three tools. Documents the residual Jackson-3 CVE (loopback, trusted clients). 52 tests green, IDE-clean.
2026-07-14 14:21:11 +02:00
Dai Ha 597ac2e562 CB-104: blocking bridge_send with rendezvous reply (POST /sessions/{id}/message + /reply)
Per-session-serialized blocking send that enqueues via the CB-103 injector (poller delivers) and blocks on a rendezvous resolved by the worker's structured bridge_reply, or a typed 200/202 outcome. No status-polling completion, no terminal scrape. GET /sessions/{id}/status. Reworked from an initial poll+scrape draft after a high-effort review found the polling completion unreliable; all findings fixed.
2026-07-14 14:21:11 +02:00
Dai Ha 081893fb32 CB-103: status-gated injector — per-worker FIFO, safe-window delivery, single writer 2026-07-13 15:52:53 +02:00
Dai Ha 1ce0aba7fc CB-108: worker placement — one tab per worker in a dedicated worker space
Workers now land in their own herdr tab inside a dedicated, shared "worker
space" (workspace.create → tab.create → agent.start{tab_id} → close seed shell
→ rename), instead of splitting the user's currently-focused tab. Placement is
configurable (worker.placement tab|pane, worker.workspace, worker.tabLabel);
a future per-session space is just a distinct label.

New: Workspace/Tab records, WorkspaceControl over workspace.*/tab.*,
AgentControl.start tab_id overload, Agent.tabId.

Lifecycle carefulness:
- ensureWorkspace is a synchronized find-or-create (never double-creates)
- teardown closes the tab ONLY when the worker is its sole pane (never a
  shared user tab), resolving the tab via pane.get before closing
- tolerance is precise: only *_not_found is swallowed; real failures surface
- spawn failure closes the orphan tab; post-start cosmetic steps can't orphan
  a live worker or fail the spawn
- unique per-worker names (claude-<profile>-<nonce>-<seq>) with retry: herdr
  rejects duplicate agent names, and the per-process nonce survives a restart
  with lingering workers

herdr facts pinned: agent.start honors tab_id; agent name must be unique;
kind/status are detected from terminal output, not the name.

Reviewed at high effort (multi-agent); all findings addressed. 34 tests green
(29 unit/acceptance + 5 contract vs live herdr 0.7.0). Also bumps wiki.
2026-07-13 08:26:53 +02:00
Dai Ha 2a815ee64e CB-102: native agent.* worker spawn (env-injected, guard-checked)
Spike decided the worker south side in favour of herdr's native agent.*
namespace over pane+send_text. Proven against live herdr 0.7.0:
agent.start takes a first-class env map that reaches the process
environment (ANTHROPIC_BASE_URL confirmed via agent.read), and herdr
tracks each worker's Claude session UUID itself.

- AgentControl: start/send/read/get/status/list + pane.close over agent.*.
- Agent/AgentStatus: projection of herdr agent records (session UUID,
  injectable status gate).
- WorkerService: build worker env (base_url/token/model/config dir),
  assertWorker BEFORE any herdr call, then agent.start.
- REST: GET /agents (discovery by session UUID), POST /workers (201, or
  403 subscription_boundary), DELETE /workers/{paneId}.
- Full stack smoke-tested live: POST->guard->agent.start->new pane,
  GET /agents lists it, DELETE closes it.
- Tests: 21 unit/acceptance + 4 contract (incl. a live end-to-end probe
  spawn that proves env injection and cleans up its pane).

No agent.stop in herdr (use pane.close); herdr id must be a string;
one request per connection — all pinned by contract tests.
2026-07-12 20:18:18 +02:00
Dai Ha 93cfdb640f Stage-1 walking skeleton: bridged Java daemon (herdr client + REST + guard)
Maven/Java 25 module under bridged/. End-to-end verified against live
herdr 0.7.0 (protocol 14): GET /healthz and GET /sessions serve real
workspace data through the socket client.

- herdr client (CB-101): Unix-socket JSON-RPC via UnixDomainSocketAddress.
  Two contract facts pinned by tests against the real daemon:
  ids MUST be strings, and herdr is one-shot per connection
  (connection-per-call, which also makes the client lock-free).
- subscription guard: worker base_url must be on the off-subscription
  allowlist (gx00.gw, ollama.ltms.dev); primary env must carry no base_url.
- config (CB-106): Jackson YAML + Logback; example grounded in ltms-local.
- REST app (CB-104 start): injectable HerdrClient so acceptance tests run
  on an ephemeral port with a fake herdr, no daemon/Claude in the loop.
- tests: 17 unit/acceptance (mvn test) + 3 contract (mvn test -Pcontract).
2026-07-12 20:00:02 +02:00