Commit Graph

18 Commits

Author SHA1 Message Date
Dai Ha 2fb46f670f CB-114: readiness-gate timeout + presence cleanup (delegated review findings)
An off-sub worker's review of CB-113 (delegated through the bridge) surfaced two
real gaps in the readiness gate:

- A worker that herdr reports idle but whose Claude never connects the bridge MCP
  (crashed during boot, or wedged on a startup prompt) left its message queued
  forever: ready.test() never passed, the target was polled indefinitely, and the
  caller's future never completed (async waiter hung for the full 30-min window).
  Injector now counts injectable-but-not-ready samples and, after a ~60s grace
  (READINESS_GRACE_POLLS, deliberately longer than the UNKNOWN stall grace since a
  first boot is slower than an in-turn blip), fails the queued messages, fires
  onTurnFailed so blocking/async waiters resolve WORKER_FAILED, and reclaims the
  target. Mirrors the CB-109 UNKNOWN-stall path.

- WorkerPresence.forget had no caller, so a worker's readiness lingered past its
  life. Injector now clears presence via a forget callback on drop() (pane crash)
  and on the readiness timeout.

+3 InjectorTest cases (never-ready failure, ready-within-grace delivery, drop clears
presence). 105 tests green.
2026-07-16 05:01:58 +02:00
Dai Ha 37f7ad4185 CB-113: reliable worker readiness gate + submission nudge
Delegating right after spawn failed: herdr reports 'idle' during the worker's
boot, so the injector delivered into a not-ready TUI (paste lost) and wedged the
worker. Two fixes:

- Readiness gate: a worker is 'available' only once its Claude connects the bridge
  MCP (the daemon observes it via peer-PID->terminal). WorkerPresence tracks it; the
  injector holds the first delivery until present, so it never pastes into the boot
  window. Exposed as 'ready' on GET /sessions/{id}/status.
- Submission nudge: the Enter accompanying a delivery can race the paste (esp. right
  as the TUI becomes ready), leaving text unsubmitted. While a delivered message
  stays idle (not picked up), the injector re-sends Enter each poll until the worker
  starts (WORKING) or the grace expires.

Validated live: spawn + immediate delegate now holds during boot (ready=false),
delivers on MCP-connect, re-nudges Enter, worker replies. 102 tests green.
2026-07-15 19:14:57 +02:00
Dai Ha 926724a279 CB-112: workers inherit the primary's working directory (not $HOME)
A worker now opens the same directory the primary is in, unless told otherwise.
Resolution: explicit spawn cwd → per-profile config cwd → the primary's cwd
(auto-detected from the bridge_spawn caller's PID via lsof -d cwd) → the daemon's
cwd. Never $HOME.

Mechanism (found by live probe, corrects the earlier assumption): an agent.start
pane does NOT inherit its tab's or workspace's cwd — it starts in $HOME. herdr's
agent.start honours an (undocumented) cwd param, so the resolved cwd is threaded
onto agent.start {cwd} (both tab and pane placement), not tab.create.

Surfaces: bridge_spawn {cwd?} + auto-detect via ConnectionIdentity.resolve (peer
PID) + ProcessCwdLookup (lsof); REST POST /workers ?cwd= / body cwd; per-profile
'cwd:' config. Validated live: explicit cwd → worker rooted there; no cwd over
REST → daemon cwd, not $HOME. Also clears the ccs folder-trust prompt when the
project dir is already trusted (see docs/Worker-Startup-and-Trust.md).
2026-07-15 16:33:46 +02:00
Dai Ha 7f1b6b3a0d CB-111: multi-profile workers — named backends selectable at spawn
The bridge was hard-wired to one worker profile. Config now takes a 'workers'
map keyed by profile name plus 'defaultWorker'; WorkerService holds the map and
gains spawn(profile) (spawn() uses the default). Selection threads through the
surfaces: REST POST /workers ?profile= / {"profile":…} + GET /profiles; MCP
bridge_spawn {profile?} + new bridge_profiles. Each profile's base_url is
guard-checked independently, so gx10 and ollama can run side by side and you
address each worker by its returned sessionId. Backward-compatible: the legacy
singular 'worker:' block still loads as a one-entry profile map.
2026-07-15 15:54:31 +02:00
Dai Ha a629a7ee73 CB-110: fail an in-flight delegation when its worker vanishes
Companion to CB-109. When a worker disappears mid-turn (pane crash → herdr
*_not_found), the status poller drops the target, which failed only *queued*
messages — a message already DELIVERED is out of the queue, so its send's
rendezvous was left hanging until the 30-min async timeout. Injector.drop now
fires onTurnFailed for a delivered-but-unresolved turn (awaitingCompletion), so
the send resolves as WORKER_FAILED. Reuses the CB-109 resolver path; the failure
reason is neutral to cover both wedge (stuck) and vanish (gone).
2026-07-15 15:23:17 +02:00
Dai Ha 2052929768 CB-109: fail a delegation whose worker wedges in an unknown state
Dogfood found the gap: a turn that dies into an error screen herdr reports as
'unknown' (e.g. the worker hitting API ENOTFOUND) never produces a working->idle
boundary, so CB-106 never fires and the async send rides its full 30-min timeout.

The injector now counts consecutive 'unknown' samples while a delegation is
outstanding; any working/idle sample resets the streak, so only a genuine wedge
(~30s continuous unknown) trips it. It then fires TurnListener.onTurnFailed;
CompletionResolver scrapes the error screen and resolves the send via
Rendezvous.resolveFailure (Kind.FAILED -> Outcome.WORKER_FAILED), surfaced as
async phase=failed / REST status=failed / a [worker failed] MCP note, with the
error context as the reason. This also frees a delivery that wedged before pickup,
which the injectable-only pickup grace could never release.
2026-07-15 15:16:55 +02:00
Dai Ha a97c287aee CB-108: fleet-management MCP tools (bridge_spawn / bridge_list / bridge_stop)
The primary could delegate to a worker but not create or reap one over MCP —
spawning was a raw REST POST /workers. BridgeMcp now adapts WorkerService so a
worker's whole lifecycle runs through MCP: bridge_spawn returns the new worker's
sessionId (for bridge_send) and paneId (for bridge_stop); bridge_list projects the
tracked workers; bridge_stop tears one down. The subscription boundary stays
enforced inside WorkerService (bridge_spawn surfaces a guard breach as a tool error
without touching herdr). Tools are thin static adapters, unit-tested by parity.
2026-07-15 14:22:50 +02:00
Dai Ha 8ed2370fbe CB-107: async fire-and-poll delegation (wait:false + ticket poll)
A caller's MCP client caps a blocking bridge_send at ~60s, but a real delegated
task runs for minutes. sendAsync runs the same blocking send on a background
virtual thread and returns a ticket; poll(ticket) reports pending/done/failed.
Async reuses the blocking path (and its per-target serialization), so it inherits
reply + completion resolution for free. Surfaces: REST POST message wait:false ->
202 {ticket} + GET /tasks/{ticket}; MCP bridge_send wait flag + new bridge_poll.
Terminal tickets are pruned after a TTL so the registry stays bounded.
2026-07-15 14:19:14 +02:00
Dai Ha 5b26caca0c CB-106: completion fallback — resolve a send when the worker's turn ends without bridge_reply
The blocking send previously resolved only on an explicit bridge_reply; a real
delegated task (edit files, run a build) finishes and goes idle without ever
calling it, so the send always timed out. The injector now reports a confirmed
working -> idle turn boundary via a TurnListener; CompletionResolver scrapes the
worker's transcript tail and resolves the awaiting send (Rendezvous.resolveCompletion,
Kind.COMPLETION -> Outcome.COMPLETED_UNREPLIED), surfaced as replySource=transcript
at the REST/MCP edges. Completion is synthesized only from a confirmed turn (a
sampled 'working'), never from the pickup-grace path, so it can't race the explicit
reply or fire on a turn that never ran.
2026-07-15 14:14:04 +02:00
Dai Ha 6988bfe88f CB-103: injector submits the prompt (Enter as a separate keystroke)
A delegated message was delivered into the worker's input box but never
submitted, so no task was ever processed: bridge_send blocked until timeout
while the worker sat idle with the prompt typed but not entered.

herdr's agent.send delivers text as a bracketed paste; a trailing carriage
return in that same call is swallowed as literal newline content, not Enter.
AgentControl.send now emits two keystroke events — the payload, then a
standalone "\r" — so the worker actually submits and runs the task.

Verified live (Claude Code v2.1.210, off-sub worker): full hands-off
bridge_send -> auto-submit -> worker computes -> bridge_reply round trip
returns the reply in ~41s. 62 tests green.
2026-07-15 09:40:18 +02:00
Dai Ha 07722a6007 CB-1xx: worker launch via ccs + inline bridge MCP mount (step 4) + port 8765
Spawn a real worker with 'ccs ltms-local' (profile sets CLAUDE_CONFIG_DIR + off-sub base_url; auto mode preconfigured as defaultMode:auto). bridged appends the bridge MCP mount (--mcp-config, inline JSON) and the reply charter (--append-system-prompt) as launch FLAGS — non-invasive, nothing written to the worker's profile (safer than provisioning its config dir, which would clobber it). Identity is connection-based so the mount is shared. config: worker.mcpUrl. Default bind port 8080 -> 8765. The reply charter is guidance; the send timeout catches a non-cooperative worker. 60 tests green, IDE-clean.
2026-07-15 09:19:31 +02:00
Dai Ha fa570ab32a CB-105: connection-based MCP caller identity (peer PID → herdr pane)
Resolve who is calling an MCP tool from the connection, not a spoofable argument (per docs/MCP-Contract.md). ConnectionIdentity ties the loopback peer PID (LsofPeerPidLookup) to a herdr pane (PaneLocator via pane.list/pane.process_info) → the caller's terminal_id; a caller owning no pane is the primary. bridge_reply now takes only content and resolves the worker from the connection (no sessionId). Live-verified: real PID→pane (contract test), and a non-pane MCP caller correctly gets a workers-only error. Single-host; the token path stays the split-host fallback. 58 tests green, IDE-clean.
2026-07-15 09:19:31 +02:00
Dai Ha 8ffbfd1f06 CB-105: MCP server (bridge_send/bridge_reply/bridge_status) over the REST core
Streamable-HTTP MCP server (io.modelcontextprotocol.sdk:mcp 2.0.0) mounted on the daemon's Jetty at /mcp, exposing three tools as thin adapters over MessageService/Rendezvous — the primary calls bridge_send/bridge_status, the worker calls bridge_reply. Tool logic in unit-testable static methods (parity tests); SDK owns the wire protocol. Resolves the Jackson 2/3 split by pinning jackson-annotations 3.0-rc5 (works for both our Jackson 2.19 and the SDK's Jackson 3). Live-verified: initialize handshake + tools/list return all three tools. Documents the residual Jackson-3 CVE (loopback, trusted clients). 52 tests green, IDE-clean.
2026-07-14 14:21:11 +02:00
Dai Ha 597ac2e562 CB-104: blocking bridge_send with rendezvous reply (POST /sessions/{id}/message + /reply)
Per-session-serialized blocking send that enqueues via the CB-103 injector (poller delivers) and blocks on a rendezvous resolved by the worker's structured bridge_reply, or a typed 200/202 outcome. No status-polling completion, no terminal scrape. GET /sessions/{id}/status. Reworked from an initial poll+scrape draft after a high-effort review found the polling completion unreliable; all findings fixed.
2026-07-14 14:21:11 +02:00
Dai Ha 081893fb32 CB-103: status-gated injector — per-worker FIFO, safe-window delivery, single writer 2026-07-13 15:52:53 +02:00
Dai Ha 1ce0aba7fc CB-108: worker placement — one tab per worker in a dedicated worker space
Workers now land in their own herdr tab inside a dedicated, shared "worker
space" (workspace.create → tab.create → agent.start{tab_id} → close seed shell
→ rename), instead of splitting the user's currently-focused tab. Placement is
configurable (worker.placement tab|pane, worker.workspace, worker.tabLabel);
a future per-session space is just a distinct label.

New: Workspace/Tab records, WorkspaceControl over workspace.*/tab.*,
AgentControl.start tab_id overload, Agent.tabId.

Lifecycle carefulness:
- ensureWorkspace is a synchronized find-or-create (never double-creates)
- teardown closes the tab ONLY when the worker is its sole pane (never a
  shared user tab), resolving the tab via pane.get before closing
- tolerance is precise: only *_not_found is swallowed; real failures surface
- spawn failure closes the orphan tab; post-start cosmetic steps can't orphan
  a live worker or fail the spawn
- unique per-worker names (claude-<profile>-<nonce>-<seq>) with retry: herdr
  rejects duplicate agent names, and the per-process nonce survives a restart
  with lingering workers

herdr facts pinned: agent.start honors tab_id; agent name must be unique;
kind/status are detected from terminal output, not the name.

Reviewed at high effort (multi-agent); all findings addressed. 34 tests green
(29 unit/acceptance + 5 contract vs live herdr 0.7.0). Also bumps wiki.
2026-07-13 08:26:53 +02:00
Dai Ha 2a815ee64e CB-102: native agent.* worker spawn (env-injected, guard-checked)
Spike decided the worker south side in favour of herdr's native agent.*
namespace over pane+send_text. Proven against live herdr 0.7.0:
agent.start takes a first-class env map that reaches the process
environment (ANTHROPIC_BASE_URL confirmed via agent.read), and herdr
tracks each worker's Claude session UUID itself.

- AgentControl: start/send/read/get/status/list + pane.close over agent.*.
- Agent/AgentStatus: projection of herdr agent records (session UUID,
  injectable status gate).
- WorkerService: build worker env (base_url/token/model/config dir),
  assertWorker BEFORE any herdr call, then agent.start.
- REST: GET /agents (discovery by session UUID), POST /workers (201, or
  403 subscription_boundary), DELETE /workers/{paneId}.
- Full stack smoke-tested live: POST->guard->agent.start->new pane,
  GET /agents lists it, DELETE closes it.
- Tests: 21 unit/acceptance + 4 contract (incl. a live end-to-end probe
  spawn that proves env injection and cleans up its pane).

No agent.stop in herdr (use pane.close); herdr id must be a string;
one request per connection — all pinned by contract tests.
2026-07-12 20:18:18 +02:00
Dai Ha 93cfdb640f Stage-1 walking skeleton: bridged Java daemon (herdr client + REST + guard)
Maven/Java 25 module under bridged/. End-to-end verified against live
herdr 0.7.0 (protocol 14): GET /healthz and GET /sessions serve real
workspace data through the socket client.

- herdr client (CB-101): Unix-socket JSON-RPC via UnixDomainSocketAddress.
  Two contract facts pinned by tests against the real daemon:
  ids MUST be strings, and herdr is one-shot per connection
  (connection-per-call, which also makes the client lock-free).
- subscription guard: worker base_url must be on the off-subscription
  allowlist (gx00.gw, ollama.ltms.dev); primary env must carry no base_url.
- config (CB-106): Jackson YAML + Logback; example grounded in ltms-local.
- REST app (CB-104 start): injectable HerdrClient so acceptance tests run
  on an ephemeral port with a fake herdr, no daemon/Claude in the loop.
- tests: 17 unit/acceptance (mvn test) + 3 contract (mvn test -Pcontract).
2026-07-12 20:00:02 +02:00