Compare commits
88 Commits
| Author | SHA1 | Date | |
|---|---|---|---|
| 3b2f395d3d | |||
| 0edc6615fc | |||
| c884802b13 | |||
| 5f5573a24e | |||
| 74b0087ebb | |||
| 927e0151d4 | |||
| 5275922d1d | |||
| e186c7945a | |||
| 33a6e77f0e | |||
| 4e47489d53 | |||
| c76b2f149e | |||
| 826fffe05b | |||
| 75f57cdba7 | |||
| f556af5d4e | |||
| d5f33f0c6e | |||
| 16e17b32ad | |||
| 988494e18b | |||
| c4549a5e20 | |||
| 46fa4f38d5 | |||
| 20e0e68ad7 | |||
| 17468a234a | |||
| 8a53d5bfc6 | |||
| 695da7418e | |||
| abd26c796b | |||
| 24559d81ac | |||
| 6c1c2c3994 | |||
| 01fab15713 | |||
| 0cd00e71c3 | |||
| bf0ff2adbf | |||
| ed4bbc1c56 | |||
| c456402cc5 | |||
| 4f0bf667b1 | |||
| 61af9aa574 | |||
| 3b59b34e76 | |||
| 65b38997f7 | |||
| 7510f7649c | |||
| 619792a81c | |||
| cea1183f75 | |||
| e2af4c5ae4 | |||
| dfd5f82894 | |||
| e81944cef6 | |||
| 7bdd39ab9a | |||
| f4b38f040e | |||
| 799668e129 | |||
| 890190263e | |||
| f8522edacd | |||
| b958747855 | |||
| 078bde2c02 | |||
| 0b10ea987b | |||
| 1553d38182 | |||
| d04b075996 | |||
| 644927636d | |||
| 95e45007aa | |||
| 7a583c4045 | |||
| 65bce058bb | |||
| 48b437083b | |||
| fa0612859b | |||
| 4ffbcd0b7d | |||
| a0cd053fd9 | |||
| f9a5e066b5 | |||
| e9bc192160 | |||
| e0a57988ad | |||
| b414a74c26 | |||
| 2c3796d598 | |||
| 7e0ff9ab06 | |||
| 6d1565d8b6 | |||
| e18e002d2f | |||
| c18572ea9d | |||
| e9d02bc5e8 | |||
| a2108a8a14 | |||
| 57f8fa257a | |||
| 61944fc045 | |||
| 4b48d2d921 | |||
| 3aa145cbef | |||
| 4875127daa | |||
| d14a624421 | |||
| 246f50b778 | |||
| 15b53c6cfa | |||
| 3a7ef0adbd | |||
| 293a305748 | |||
| e5038c6d13 | |||
| 13ea79f6fd | |||
| f55d3c203b | |||
| 37c4d47d3a | |||
| 04e21c9243 | |||
| 83cac07f6e | |||
| f472e0f782 | |||
| 91c9f981c5 |
@@ -92,24 +92,32 @@ route by who launches opencode:
|
||||
|
||||
| Launcher | Route |
|
||||
|---|---|
|
||||
| a human, from a terminal | `{file:…}` — see below |
|
||||
| a human, from a terminal | one central store, sourced by the login shell |
|
||||
| a spawner (bridge, CI, IDE) | `{env:…}`, with the spawner injecting the variable |
|
||||
|
||||
For the human case prefer **`{file:…}` with a workspace-relative path**, kept in a gitignored
|
||||
`.secrets/` directory beside `opencode.json`:
|
||||
**For the human case, keep every credential in one file the login shell sources.** Here that file
|
||||
is `${SHARED_ENV}/tools/secrets.sh`, sourced from `${SHARED_ENV}/.ltms`, kept at mode 600 and never
|
||||
committed. `opencode.json` then names variables and holds no values:
|
||||
|
||||
```json
|
||||
"headers": { "Authorization": "Bearer {file:.secrets/api-token}" }
|
||||
"headers": { "Authorization": "Bearer {env:CONTEXT7_TOKEN}" }
|
||||
```
|
||||
|
||||
Verified: opencode resolves relative `{file:}` paths against the project root, so this needs no
|
||||
shell setup at all — no rc export leaking the secret to every process, no direnv dependency.
|
||||
Both routes end at the same syntax, and that is the point. The file does not change when a human
|
||||
launches opencode instead of the bridge.
|
||||
|
||||
**The catch, and state it out loud:** `.secrets/` is gitignored, so a peer running in a git worktree
|
||||
does **not** get it — worktrees receive tracked files only, the same rule that makes `opencode.json`
|
||||
itself worth committing. Spawned peers must therefore be fed through `{env:…}` by whatever launches
|
||||
them. Check the variable *names* match: a spawner often injects under a different name than your
|
||||
shell uses, and the config has no fallback.
|
||||
**`{file:…}` also works, and this project moved away from it.** Opencode resolves a relative
|
||||
`{file:}` path against the project root, so a gitignored `.secrets/` beside `opencode.json` needs no
|
||||
shell setup at all. It did not fail; the problem is that it makes a second copy of the token. The
|
||||
same secret then lives in two places, and the copy you forget is the one that leaks or goes stale.
|
||||
One store with many references is easier to rotate and to audit.
|
||||
|
||||
**One catch survives either choice, so state it out loud:** what a human's shell exports does not
|
||||
reach a spawned peer, and neither does a gitignored `.secrets/` — a git worktree receives tracked
|
||||
files only. Spawned peers must be fed through `{env:…}` by whatever launches them. Check the
|
||||
variable *names* match: a spawner often injects under a different name than your shell uses, and
|
||||
the config has no fallback. In this repo the bridge goes further: it neutralizes a worktree's
|
||||
`opencode.json`, so a member cannot inherit the primary's credentials by accident.
|
||||
|
||||
## 4. Do not port machine-local MCP servers
|
||||
|
||||
|
||||
+3
-2
@@ -1,5 +1,6 @@
|
||||
# Workspace-scoped secrets for tools that read no settings cascade of their own
|
||||
# (opencode resolves these via {file:.secrets/...} in opencode.json).
|
||||
# No secret belongs in this repo any more: every credential lives in one shell-level store
|
||||
# (${SHARED_ENV}/tools/secrets.sh), and opencode.json reads it as {env:...}. This line stays as a
|
||||
# backstop, so a workspace-scoped copy that someone re-creates by habit still cannot be committed.
|
||||
.secrets/
|
||||
|
||||
# Settings backups inherit the env block — and secrets with it.
|
||||
|
||||
@@ -11,41 +11,45 @@
|
||||
If no `bridge_*` MCP tools are mounted in this session, this section does not apply — skip it.
|
||||
|
||||
`bridged` is the **sole communication gateway** between agents here. The orchestrating session (the
|
||||
**primary**) and every delegated peer (a **worker**) mount the *same* MCP server and talk only
|
||||
**primary**) and every delegated peer (a **member**) mount the *same* MCP server and talk only
|
||||
through its `bridge_*` tools. No session addresses a peer, a broker, or the network directly.
|
||||
|
||||
### Which role am I? — settle this before acting
|
||||
|
||||
**Both roles read this file.** A worker runs in a git worktree of this same repo, so it inherits
|
||||
**Every role reads this file.** A member runs in a git worktree of this same repo, so it inherits
|
||||
this `CLAUDE.md` verbatim, and every rule below is role-conditional.
|
||||
|
||||
**Call `bridge_whoami`.** It returns `{"role":"primary"}` or `{"role":"worker","sessionId":…,
|
||||
"profile":…,"worktree":…,"branch":…}`, resolved by the daemon from your connection — unforgeable,
|
||||
and the same resolution its authorization gate uses. Don't infer what you can ask.
|
||||
**Call `bridge_whoami`.** It returns `primary`, `worker`, or `architect`, resolved by the daemon from
|
||||
your connection — unforgeable, and the same resolution its authorization gate uses. A worker also
|
||||
carries its `sessionId`, `profile`, `worktree` and `branch`; an architect carries the slot name it
|
||||
was bound to. Don't infer what you can ask.
|
||||
|
||||
Only if that call is unavailable, fall back to these — each is one-way, so keep reading until one
|
||||
fires: the reply charter in your system prompt (*"You are an off-subscription worker in the
|
||||
claude-bridge fleet"*) ⇒ **worker**; bridge tools prefixed `mcp__bridge__*` ⇒ **worker** (the
|
||||
launcher fixes that mount name; a primary's mount is named by whoever wrote its `.mcp.json`, so it
|
||||
varies); `ANTHROPIC_BASE_URL` set ⇒ **worker** (Claude-model workers run on a clean env, so its
|
||||
*absence* proves nothing). **Still unsure ⇒ act as a worker.** The two mistakes are not symmetric: a
|
||||
primary acting as a worker is refused by the authorization gate — loud and self-correcting — while a
|
||||
worker acting as the primary ends its turn with no `bridge_reply`, and the sender silently receives
|
||||
nothing. Fail toward the recoverable error.
|
||||
claude-bridge fleet"*) ⇒ **spawned member**; bridge tools prefixed `mcp__bridge__*` ⇒ **spawned
|
||||
member** (the launcher fixes that mount name; a primary's mount is named by whoever wrote its
|
||||
`.mcp.json`, so it varies); `ANTHROPIC_BASE_URL` set ⇒ **spawned member** (Claude-model members run
|
||||
on a clean env, so its *absence* proves nothing). None of these separate a worker from an architect —
|
||||
only `bridge_whoami` does. **Still unsure ⇒ act as a worker**, the most restricted member role. The
|
||||
two mistakes are not symmetric: a primary acting as a worker is refused by the authorization gate —
|
||||
loud and self-correcting — while a member acting as the primary ends its turn with no `bridge_reply`,
|
||||
and the sender silently receives nothing. Fail toward the recoverable error.
|
||||
|
||||
### Invariants — both roles, no exceptions
|
||||
|
||||
1. **Never set, export, or forward `ANTHROPIC_BASE_URL`** (or `ANTHROPIC_AUTH_TOKEN`). The primary
|
||||
stays on subscription; only the bridge puts a worker off it, at spawn. Mounting the bridge must
|
||||
stays on subscription; only the bridge puts a member off it, at spawn. Mounting the bridge must
|
||||
never move a session across that boundary.
|
||||
2. **The bridge is the only channel.** Text you print in your terminal reaches nobody — the other
|
||||
side cannot see your screen. An answer that isn't in a `bridge_*` call is silently discarded.
|
||||
3. **Identity comes from the connection, never an argument.** Workers never pass a target; you
|
||||
cannot act as another session. Spawn/stop/send/drain are lead-only; reply/ask are
|
||||
only-as-itself — any peer may answer for its own pane, and for no other. A call outside your
|
||||
role is refused, not queued.
|
||||
cannot act as another session. Spawn/stop/drain are lead-only; **send is lead or architect**;
|
||||
reply/ask are only-as-itself — any peer may answer for its own pane, and for no other. A call
|
||||
outside your role is refused, not queued.
|
||||
4. **Delivery is status-gated: one message per turn.** Don't busy-poll a peer's terminal and don't
|
||||
re-send because a call looks slow — the bridge delivers when the peer is `idle`/`blocked`.
|
||||
re-send because a call looks slow — the bridge delivers when the peer is `idle`, `blocked` or
|
||||
`done`. A spawned member must **also** have mounted the bridge MCP: until it has, it is not
|
||||
deliverable, and a send waits on that gate for ~60s and then fails without ever reaching its pane.
|
||||
5. **Never drive the terminal multiplexer directly** (no `herdr` CLI, no socket). The bridge owns
|
||||
policy; the multiplexer owns PTYs. Going around the bridge bypasses every rule above.
|
||||
|
||||
@@ -101,26 +105,26 @@ the merge — and merging on a reviewer's word is delegating it by proxy.
|
||||
|---|---|
|
||||
| Confirm your own role | `bridge_whoami` |
|
||||
| See backends available | `bridge_profiles` |
|
||||
| Start a worker | `bridge_spawn{profile?, cwd?, worktree?, ticket?}` → `sessionId` + `paneId` |
|
||||
| See the fleet | `bridge_list` → `leads` (your peers) + `workers` · one peer's state: `bridge_status{sessionId}` |
|
||||
| Start a member | `bridge_spawn{role?, profile?, cwd?, worktree?, ticket?}` → `sessionId` + `paneId` |
|
||||
| See the fleet | `bridge_list` → `leads` (your peers) + `members` · one peer's state: `bridge_status{sessionId}` |
|
||||
| Delegate (blocking) | `bridge_send{sessionId, content}` |
|
||||
| Delegate (long task) | `bridge_send{sessionId, content, wait:false}` → ticket → `bridge_poll{ticket}` |
|
||||
| Answer a worker's `bridge_ask` | `bridge_send{turnId, content}` — **not** `sessionId` |
|
||||
| Answer a member's `bridge_ask` | `bridge_send{turnId, content}` — **not** `sessionId` |
|
||||
| Message a **peer lead** | `bridge_send{sessionId: <their terminal>, content}` — `bridge_list` → `leads` reports it. Coordination only, **never** a task |
|
||||
| Answer a peer lead that messaged you | `bridge_reply{content}` — the one case a lead replies |
|
||||
| Collect a held reply | `bridge_poll{target}` · then `bridge_ack{target, msgId}` |
|
||||
| Tear down | `bridge_stop{paneId}` |
|
||||
| Tear down a member | `bridge_stop{paneId}` |
|
||||
|
||||
### Lead ↔ lead — coordinate, never delegate
|
||||
|
||||
`bridge_list` returns `leads` alongside `workers`; your own row carries `self: true`. Every other row
|
||||
is a peer — an orchestrator with its own context, its own workers, and its own judgment. An empty
|
||||
`workers` array means no workers are spawned; it says nothing about peers.
|
||||
`bridge_list` returns `leads` alongside `members`; your own row carries `self: true`. Every other row
|
||||
is a peer — an orchestrator with its own context, its own members, and its own judgment. An empty
|
||||
`members` array means no members are spawned; it says nothing about peers.
|
||||
|
||||
**A lead never assigns a task to another lead.** Work goes to workers — only ever downward, never
|
||||
sideways. Sending a peer a brief with acceptance criteria is a category error: a brief is a worker's
|
||||
**A lead never assigns a task to another lead.** Work goes to members — only ever downward, never
|
||||
sideways. Sending a peer a brief with acceptance criteria is a category error: a brief is a member's
|
||||
artefact, and a peer is not yours to task. If a unit needs doing and it falls in your area, spawn a
|
||||
worker and delegate it yourself; if it falls in the peer's area, say so and let the peer assign it.
|
||||
member and delegate it yourself; if it falls in the peer's area, say so and let the peer assign it.
|
||||
The traffic between leads is coordination and nothing else:
|
||||
|
||||
1. **Divide the map, not the work.** Agree who owns which area, then each of you assigns inside your
|
||||
@@ -138,7 +142,7 @@ Being messaged by a peer does not make you its worker: answer with `bridge_reply
|
||||
the substance if it is wrong. A peer that simply complies has thrown away the reason there are two of
|
||||
you.
|
||||
|
||||
### Worker — the turn contract
|
||||
### Member (worker or architect) — the turn contract
|
||||
|
||||
1. **Load the playbook skill the lead named** before doing anything else.
|
||||
2. **Do the assigned scope only.** Note anything you spot outside it in one line; don't go hunt it.
|
||||
@@ -147,6 +151,10 @@ you.
|
||||
Don't ask what you could decide yourself.
|
||||
4. **End the turn with exactly one `bridge_reply{content}`**, carrying your complete answer. This is
|
||||
the whole handoff. No `bridge_reply` ⇒ the sender gets nothing and the exchange stalls.
|
||||
Do **not** lean on the completion fallback to carry your answer for you: when you end a turn
|
||||
without replying, the bridge scrapes your pane, and it can return only the last 4000 characters.
|
||||
A clipped scrape is marked as partial, but the missing text is gone — your report reaches the
|
||||
lead with its end cut off.
|
||||
5. **Report honestly.** State only what you actually ran and its real output, including failures.
|
||||
You mount **only** the bridge MCP — the primary's other servers (IDE, forge, docs) are not yours,
|
||||
so never claim the result of a check you had no way to run.
|
||||
@@ -157,9 +165,9 @@ you.
|
||||
|
||||
| Layer | Scope | Reaches |
|
||||
|---|---|---|
|
||||
| the launcher's reply charter | the one rule that must survive with no repo: *end every turn with `bridge_reply`* | every worker, at launch, every peer kind |
|
||||
| **this section** | protocol + orchestration policy | primary **and** every Claude worker — tracked in git, so worktrees inherit it |
|
||||
| role playbook skills | per-job procedure (commit/PR recipe, finding format) | a worker told to load one |
|
||||
| the launcher's reply charter | the one rule that must survive with no repo: *end every turn with `bridge_reply`* | every spawned member, at launch, every peer kind — never a lead |
|
||||
| **this section** | protocol + orchestration policy | primary **and** every member that reads the repo — tracked in git, so worktrees inherit it |
|
||||
| role playbook skills | per-job procedure (commit/PR recipe, finding format) | a member told to load one |
|
||||
| the bridge's own docs | design detail, flows, error model | on demand |
|
||||
|
||||
A rule belongs in **exactly one** layer — the outermost one that must obey it. Peers that don't read
|
||||
|
||||
+166
-33
@@ -55,44 +55,57 @@ bind:
|
||||
# IS a primary for authorization, so nothing that keys on the role breaks.
|
||||
#
|
||||
# KEEP `primary:` when adding leads: it still addresses the CB-307 push loop, which needs a single
|
||||
# destination for its nudges. If both name the same terminal, the `leaders:` entry wins.
|
||||
# leaders:
|
||||
# opus-5.0:
|
||||
# terminal: term_0123456789abcd
|
||||
# kind: claude
|
||||
# gpt-sol-5.6:
|
||||
# terminal: term_fedcba9876543
|
||||
# kind: opencode
|
||||
# model: openai/gpt-5.6-terra
|
||||
# destination for its nudges. If both name the same terminal, the `fleet.leaders:` entry wins.
|
||||
#
|
||||
# Leads are configured under `fleet.leaders:` — see THE FLEET further down.
|
||||
#
|
||||
# Two things stop the tab-name convention from becoming a way to claim leadership: the configured
|
||||
# member spaces are excluded from the scan, so nothing bridged places can land in a matching tab;
|
||||
# and startup REFUSES a `tabPrefix` that the fleet tabLabel template, or any per-profile `tabLabel`
|
||||
# override, also matches — so the two namespaces cannot overlap by accident. The label is a NAME,
|
||||
# never a capability: what a pane may do is decided by the role the daemon resolves for it.
|
||||
|
||||
# CB-531: FIND LEADS BY TAB NAME instead of pasting terminal ids. `leaders:` above needs an id that
|
||||
# only exists once the session is running, so adding a lead is: open a tab, start the agent, ask it
|
||||
# bridge_whoami, edit this file, restart the daemon. This block replaces all of that with a naming
|
||||
# convention — label the tab `lead: <name>` when you open it and the pane is recognised on the next
|
||||
# rescan, with no config edit and no restart. Reopen the tab later and the id changes; the label
|
||||
# does not.
|
||||
# CB-551: IDLE-LEAD HEARTBEAT — nudge the single lead back to work when it has been continuously
|
||||
# idle (no open bridge_send driving it) past the quiet period. The fleet is one lead + architects +
|
||||
# workers, so a lead that stalls is a single point of failure; the ReplyPushLoop only nudges when a
|
||||
# reply lands, and this timer catches the gap where nothing lands and the lead just sits idle.
|
||||
#
|
||||
# bridged NEVER writes these labels. It renames worker tabs (see `tabLabel` below) but reads lead
|
||||
# tabs read-only, so what is in the tab bar is always what you typed. Two things keep the convention
|
||||
# from being a way to claim leadership: the configured worker spaces are excluded from the scan, so
|
||||
# nothing bridged places can land in a matching tab; and startup REFUSES a `tabPrefix` that any
|
||||
# worker `tabLabel` also matches, so the two namespaces cannot overlap by accident.
|
||||
# Opt-in on purpose — it SPENDS the operator's subscription on its own initiative (each nudge starts
|
||||
# a lead turn nobody asked for), so upgrading the daemon must never switch it on for you. Absent
|
||||
# block = feature off, exactly as before.
|
||||
#
|
||||
# Opt-in on purpose — this widens who resolves as a lead, so upgrading the daemon must never switch
|
||||
# it on for you. Absent block = leads come only from `leaders:`/`primary:`, exactly as before.
|
||||
# leadScan:
|
||||
# tabPrefix: "lead:" # `lead: opus-5.0` ⇒ a lead named opus-5.0 (case-insensitive; default "lead:")
|
||||
# intervalSeconds: 10 # rescan cadence, and the worst case before a new tab is recognised
|
||||
# Three knobs, each with a default that errs on the side of not burning context:
|
||||
# idleAfterSeconds: 300 # how long the lead must stay idle before the FIRST nudge (default 300 —
|
||||
# # absorbs normal post-turn pauses; re-prompting every pause burns context)
|
||||
# backoffMs: 60000 # re-check cadence / spacing between nudges past the quiet period (default 60000)
|
||||
# quietNudgeCap: 3 # cap on consecutive nudges that find NOTHING pending, then it stops
|
||||
# # until real state appears (default 3 — never nag an empty fleet forever)
|
||||
# leadHeartbeat:
|
||||
# idleAfterSeconds: 300
|
||||
# backoffMs: 60000
|
||||
# quietNudgeCap: 3
|
||||
|
||||
# Fleet health detection is dormant unless enabled. It reads one whole-fleet agent list per tick.
|
||||
# It can run without a webhook; bridge_list then reports healthCoverage: detection-only.
|
||||
# health:
|
||||
# enabled: true
|
||||
# intervalSeconds: 30 # minimum 15
|
||||
# workingSuspectAfterSeconds: 600 # minimum 300
|
||||
# paneProbeIntervalSeconds: 60 # minimum 60
|
||||
# notifications:
|
||||
# mode: disabled # disabled (default) or webhook
|
||||
|
||||
# herdr Unix socket. Omit to use the client default
|
||||
# (${HERDR_SOCKET_PATH:-~/.config/herdr/herdr.sock}).
|
||||
herdrSocket: ~/.config/herdr/herdr.sock
|
||||
|
||||
# How worker sessions are spawned. Define one or more named profiles (backends) under
|
||||
# `workers`; each key is the profile name (also the ccs profile). `defaultWorker` picks
|
||||
# which one a no-argument spawn uses (bridge_spawn with no profile / POST /workers).
|
||||
# How member sessions are spawned. Define one or more named profiles (backends) under
|
||||
# `profiles`; each key is the profile name (also the ccs profile). A profile says only WHICH
|
||||
# BACKEND — model, CLI adapter, credentials, cost. It says nothing about what a member spawned on
|
||||
# it is for; that is the member's role, and roles live under `fleet:` below. Which profile an
|
||||
# unqualified spawn lands on comes from that role's pool, not from a global default.
|
||||
#
|
||||
# Shared knobs (placement/workspace/tabLabel) can be repeated per profile; they usually match.
|
||||
# Shared knobs (placement/workspace) can be repeated per profile; they usually match.
|
||||
# placement: tab → each worker lands in its OWN tab in a dedicated worker space (default).
|
||||
# Use `pane` for the legacy behaviour (split the focused tab).
|
||||
# mcpUrl → bridged mounts the bridge MCP (--mcp-config, inline) + reply charter
|
||||
@@ -140,14 +153,14 @@ herdrSocket: ~/.config/herdr/herdr.sock
|
||||
# entry cannot repoint a worker past the SubscriptionGuard — which is checked
|
||||
# against `baseUrl` alone.
|
||||
# Put `defaultMode: "auto"` in each ccs profile so the worker runs autonomously.
|
||||
workers:
|
||||
profiles:
|
||||
gx10: # ccs profile name (NOT a hostname)
|
||||
kind: claude-code # which adapter spawns this profile (default; may omit)
|
||||
baseUrl: http://gx01.gw:8000 # the vLLM host this profile targets (gx00.gw / gx01.gw)
|
||||
model: coder
|
||||
placement: tab
|
||||
workspace: bridged-workers
|
||||
tabLabel: "worker: {profile} #{n}" # {profile}/{model}/{n} substituted; {n} keeps sibling tabs distinct
|
||||
# tabLabel: an optional per-profile override; the fleet template usually covers it
|
||||
mcpUrl: http://127.0.0.1:8765/mcp
|
||||
tokenEnv: BRIDGED_WORKER_TOKEN
|
||||
argv: ["ccs", "gx10"]
|
||||
@@ -162,7 +175,7 @@ workers:
|
||||
baseUrl: http://gx01.gw:8000 # self-hosted; ccs handles the model + token
|
||||
placement: tab
|
||||
workspace: bridged-workers
|
||||
tabLabel: "worker: {profile} #{n}"
|
||||
# tabLabel: an optional per-profile override; the fleet template usually covers it
|
||||
mcpUrl: http://127.0.0.1:8765/mcp
|
||||
argv: ["ccs", "gx11"]
|
||||
weight: 0.5
|
||||
@@ -219,7 +232,127 @@ workers:
|
||||
# How an unqualified spawn chooses a profile: fixed (default, reproduces pre-CB-518 behaviour),
|
||||
# round-robin, or weighted. Omitting this key is a strict no-op for existing configs.
|
||||
placement: weighted
|
||||
defaultWorker: gx10
|
||||
|
||||
# Re-read this file without restarting the daemon (CB-559). Off unless you add this block, so an
|
||||
# upgraded bridged keeps the old behaviour: the file is read once at boot and never again.
|
||||
# enabled → turn the watch on. bridged checks the file's modified time on a timer and
|
||||
# reloads when it moves.
|
||||
# intervalSeconds → how often to check (default 10). One `stat` per tick, so this is cheap.
|
||||
#
|
||||
# Not every key can move under a running daemon, and the difference is about what already exists
|
||||
# when the reload happens — not about how important the key is:
|
||||
# HOT → takes effect on the next spawn: the whole `fleet:` block (every role pool,
|
||||
# `charters`, and `tabLabel`), `placement:`, and an existing profile's weight / maxLoad. Those are
|
||||
# hot because the placement policy reads them through a supplier — being config is
|
||||
# not by itself enough to make a key hot.
|
||||
# DEFERRED → accepted into the new config, but the wiring built at startup keeps the old value
|
||||
# until you restart: `lifecycle:`, `leadHeartbeat:`, `guard:`, `worktreeRoot:`,
|
||||
# `spawnReadyTimeoutMs` / `spawnReadyPollMs`, ADDING or REMOVING a profile (a new
|
||||
# backend needs its own launcher, and launchers are built once), AND an existing
|
||||
# profile's launch settings — model, baseUrl, argv, env, configDir, mcpUrl, tabLabel.
|
||||
# The launcher takes a copy of `profiles:` at startup and resolves every spawn out of
|
||||
# that copy, so those never reach a launch until you restart. The reload logs them by
|
||||
# name rather than pretending they applied.
|
||||
# COLD → cannot change at all: `bind:`, `herdrSocket:`, `broker:` and `auth:`. The socket is
|
||||
# bound, the broker connection is open, and the auth mode decides who may reach the
|
||||
# port that is already listening.
|
||||
#
|
||||
# A changed COLD key refuses the WHOLE reload — not the hot half applied and the cold half warned
|
||||
# about. A half-applied reload would leave the daemon matching no file on disk, which is the worst
|
||||
# thing a reload can do to an operator debugging one. A file that fails to parse or fails a startup
|
||||
# validator is refused the same way, and the running config stays live.
|
||||
# configReload:
|
||||
# enabled: true
|
||||
# intervalSeconds: 10
|
||||
|
||||
# THE FLEET (CB-557) — who the daemon may run, and under which role. This one block replaced four
|
||||
# older keys: `leaders:`, `members:`, `leadScan:` and `defaultProfile:`.
|
||||
#
|
||||
# A member is anything a lead spawns, and every member has two INDEPENDENT attributes:
|
||||
# role — which contract: architect, dev or reviewer. It picks the launch charter, the role
|
||||
# file, the playbook skill and the authz row.
|
||||
# profile — which backend: one of the `profiles:` keys above (model, CLI adapter, cost).
|
||||
# They vary on their own. A reviewer may run on the same profile as the dev whose diff it reads,
|
||||
# which is why the two cannot be one field.
|
||||
#
|
||||
# The ROLE IS THE CONTAINING KEY, not a `role:` field. That is not only tidier: a misspelled role
|
||||
# used to parse into a member with no contract at all, while a misspelled pool name here simply
|
||||
# declares nothing.
|
||||
#
|
||||
# Each pool lists the profiles that role MAY run on — these are pools, not identities. That is also
|
||||
# what replaced `defaultProfile:`: an unqualified spawn names a role, and that role's pool supplies
|
||||
# the candidates, in definition order. A dev and a reviewer staying anonymous is exactly compatible
|
||||
# with being listed here; the entry key just names the entry.
|
||||
fleet:
|
||||
# Optional launch-charter text, keyed only by the singular role wire names: architect, dev,
|
||||
# reviewer. Changes are HOT and reach the next spawn without a daemon restart. Do not put secrets
|
||||
# here: a later launch step writes this text to a world-readable temp file, and ${ENV} interpolation
|
||||
# is deliberately not supported.
|
||||
charters:
|
||||
architect: |-
|
||||
You are an architect in this fleet. You refine work before anyone builds it:
|
||||
scope, acceptance criteria, risks, and a unit split. You read the repo and
|
||||
write analysis. You never commit production code and never open a PR.
|
||||
A design task is worked by two architects. Design alone first, then exchange
|
||||
and say plainly where you disagree. Do not concede just to agree.
|
||||
dev: |-
|
||||
You implement the one unit you were given, and nothing else. You test it,
|
||||
commit it, and open your own pull request. You never merge.
|
||||
reviewer: |-
|
||||
You review the diff you were given. You report bugs, risks and missing tests.
|
||||
You do not change code.
|
||||
|
||||
# Optional. Template for a member tab's label; {role}, {profile}, {model} and {n} are substituted.
|
||||
# {n} counts per role+profile, so `dev: sonnet #2` really is the second sonnet dev. Because {role}
|
||||
# comes from a closed enum, a generated label can never begin with a lead's tabPrefix.
|
||||
# tabLabel: "{role}: {profile} #{n}"
|
||||
|
||||
# Panes that orchestrate rather than are orchestrated. A lead may now be CREATED as well as
|
||||
# recognised: give it a `profile:` and the daemon launches the shortfall when fewer than
|
||||
# `instances` are live. Give it only a `terminal:` and it is recognise-only, as before.
|
||||
#
|
||||
# `tabPrefix` is the naming convention that finds a lead without pasting a terminal id: label the
|
||||
# tab `lead: <name>` when you open it and the pane is recognised on the next rescan. Reopen the
|
||||
# tab later and the id changes; the label does not.
|
||||
#
|
||||
# A lead the daemon launches is labelled BY the daemon, using the same convention, so it is found
|
||||
# by the same scan. A lead counts as live only when herdr also reports a running agent in that
|
||||
# tab — a label left behind by a session that died does not block the relaunch.
|
||||
#
|
||||
# An auto-launched lead is NOT a member: it gets no worker reply charter, is never registered with
|
||||
# the session lifecycle (the idle reaper would kill your orchestrator), and stays on the
|
||||
# subscription — ANTHROPIC_BASE_URL/AUTH_TOKEN are stripped from its env whatever the profile says.
|
||||
# leaders:
|
||||
# opus-5.0:
|
||||
# profile: opus # omit to never create this lead, only recognise it
|
||||
# instances: 1 # desired live count; only the shortfall is launched. 0 = off
|
||||
# terminal: term_0123456789abcd # optional hand-pin; usually found by tabPrefix instead.
|
||||
# # A running agent on this terminal also counts as live, so a
|
||||
# # lead you opened by hand is not relaunched under you.
|
||||
# tabPrefix: "lead:" # `lead: opus-5.0` ⇒ a lead named opus-5.0 (case-insensitive)
|
||||
# scanIntervalSeconds: 10 # rescan cadence, and the worst case before a new tab is seen
|
||||
# workspace: leads # where a launched lead's tab is created (default "leads").
|
||||
# # MUST NOT be a member workspace — those are excluded from the
|
||||
# # scan, so a lead placed in one is never found again.
|
||||
# cwd: /path/to/repo # the launched lead's working directory (default: bridged's own)
|
||||
# kind: claude # descriptive; reported by bridge_whoami
|
||||
# gpt-sol-5.6:
|
||||
# terminal: term_fedcba9876543
|
||||
# kind: opencode
|
||||
# model: openai/gpt-5.6-terra
|
||||
|
||||
# architects:
|
||||
# architect-1:
|
||||
# profile: opus # a strong model, on the operator's subscription
|
||||
# architect-2:
|
||||
# profile: sol # a different vendor on purpose — two architects that share a
|
||||
# # model share its blind spots
|
||||
developers:
|
||||
gx10:
|
||||
profile: gx10
|
||||
# reviewers:
|
||||
# gx10:
|
||||
# profile: gx10 # the same backend may serve two roles; that is the point
|
||||
|
||||
# Subscription boundary. A worker's base_url host MUST be one of these; the primary
|
||||
# must carry none. Every profile above must have its host listed here.
|
||||
|
||||
@@ -1,11 +1,14 @@
|
||||
package dev.ltms.bridged;
|
||||
|
||||
import dev.ltms.bridged.config.BridgedConfig;
|
||||
import dev.ltms.bridged.config.ConfigRef;
|
||||
import dev.ltms.bridged.config.ConfigWatcher;
|
||||
import dev.ltms.bridged.guard.SubscriptionGuard;
|
||||
import dev.ltms.bridged.herdr.AgentControl;
|
||||
import dev.ltms.bridged.herdr.HerdrClient;
|
||||
import dev.ltms.bridged.herdr.HerdrException;
|
||||
import dev.ltms.bridged.herdr.LeadTabScanner;
|
||||
import dev.ltms.bridged.lead.LeadLauncher;
|
||||
import dev.ltms.bridged.herdr.PaneLocator;
|
||||
import dev.ltms.bridged.herdr.UnixSocketHerdrClient;
|
||||
import dev.ltms.bridged.herdr.WorkspaceControl;
|
||||
@@ -13,13 +16,14 @@ import dev.ltms.bridged.inject.CompletionResolver;
|
||||
import dev.ltms.bridged.inject.Injector;
|
||||
import dev.ltms.bridged.inject.StatusPoller;
|
||||
import dev.ltms.bridged.inject.TurnListener;
|
||||
import dev.ltms.bridged.inject.WorkerPresence;
|
||||
import dev.ltms.bridged.auth.ArchitectRegistry;
|
||||
import dev.ltms.bridged.inject.MemberPresence;
|
||||
import dev.ltms.bridged.auth.MemberRegistry;
|
||||
import dev.ltms.bridged.auth.CallerResolver;
|
||||
import dev.ltms.bridged.mcp.BridgeMcp;
|
||||
import dev.ltms.bridged.mcp.ConnectionIdentity;
|
||||
import dev.ltms.bridged.metrics.BridgedMetrics;
|
||||
import dev.ltms.bridged.metrics.Metrics;
|
||||
import dev.ltms.bridged.health.FleetHealthMonitor;
|
||||
import dev.ltms.bridged.mcp.PrimaryRegistry;
|
||||
import dev.ltms.bridged.mcp.LsofPeerPidLookup;
|
||||
import dev.ltms.bridged.mcp.LsofProcessCwdLookup;
|
||||
@@ -28,17 +32,17 @@ import dev.ltms.bridged.msg.InMemoryReplyInbox;
|
||||
import dev.ltms.bridged.msg.MessageService;
|
||||
import dev.ltms.bridged.msg.Rendezvous;
|
||||
import dev.ltms.bridged.msg.ReplyInbox;
|
||||
import dev.ltms.bridged.msg.LeadHeartbeatLoop;
|
||||
import dev.ltms.bridged.msg.ReplyPushLoop;
|
||||
import dev.ltms.bridged.rest.BridgedApp;
|
||||
import dev.ltms.bridged.session.GitWorktrees;
|
||||
import dev.ltms.bridged.session.SessionManager;
|
||||
import dev.ltms.bridged.peer.PeerLauncher;
|
||||
import dev.ltms.bridged.session.SessionReaper;
|
||||
import dev.ltms.bridged.worker.ClaudeCodeLauncher;
|
||||
import dev.ltms.bridged.placement.PlacementPolicies;
|
||||
import dev.ltms.bridged.worker.CompositePeerLauncher;
|
||||
import dev.ltms.bridged.worker.HerdrPeerLauncher;
|
||||
import dev.ltms.bridged.worker.OpenCodeLauncher;
|
||||
import dev.ltms.bridged.member.ClaudeCodeLauncher;
|
||||
import dev.ltms.bridged.member.CompositePeerLauncher;
|
||||
import dev.ltms.bridged.member.HerdrPeerLauncher;
|
||||
import dev.ltms.bridged.member.OpenCodeLauncher;
|
||||
import io.javalin.Javalin;
|
||||
import org.slf4j.Logger;
|
||||
import org.slf4j.LoggerFactory;
|
||||
@@ -51,6 +55,7 @@ import java.util.Map;
|
||||
import java.util.Objects;
|
||||
import java.util.Set;
|
||||
import java.util.concurrent.Executors;
|
||||
import java.util.concurrent.atomic.AtomicReference;
|
||||
import java.util.concurrent.TimeUnit;
|
||||
import java.util.concurrent.atomic.AtomicReference;
|
||||
import java.util.function.Function;
|
||||
@@ -67,9 +72,6 @@ public final class Bridged {
|
||||
|
||||
private static final Logger log = LoggerFactory.getLogger(Bridged.class);
|
||||
|
||||
/** How often the injector samples a busy worker's status while it has queued work. */
|
||||
private static final long INJECT_POLL_MILLIS = 250;
|
||||
|
||||
/** CB-504: how long to wait at startup for herdr's socket before serving degraded. */
|
||||
private static final long HERDR_WAIT_SECONDS = 30;
|
||||
private static final long HERDR_WAIT_POLL_MILLIS = 500;
|
||||
@@ -77,6 +79,11 @@ public final class Bridged {
|
||||
static void main(String[] args) {
|
||||
Path configPath = Path.of(args.length > 0 ? args[0] : "bridged.yaml");
|
||||
BridgedConfig cfg = BridgedConfig.load(configPath);
|
||||
// CB-559: `cfg` stays the startup snapshot — every validation and every piece of one-time
|
||||
// wiring below reads it, and must, because those decisions cannot be unmade. `config` is the
|
||||
// live reference the hot paths read per use. Which keys can actually move is ConfigRef's
|
||||
// contract; adding a reader here does not make a key reloadable by itself.
|
||||
ConfigRef config = new ConfigRef(configPath, cfg);
|
||||
|
||||
// The primary/host env that launched bridged must not be tainted.
|
||||
SubscriptionGuard guard = new SubscriptionGuard(cfg.guard().hostSet());
|
||||
@@ -87,13 +94,14 @@ public final class Bridged {
|
||||
// refuses remote connections to a loopback socket. This throws rather than warns so the
|
||||
// dangerous configuration cannot be reached by ignoring a log line.
|
||||
cfg.validateAuthExposure();
|
||||
cfg.validateLeadScan();
|
||||
cfg.validateLeadTabPrefixes();
|
||||
// CB-542: a subscription:true profile whose env: reseats ANTHROPIC_BASE_URL/AUTH_TOKEN would
|
||||
// reach an unguarded endpoint (the launcher skips SubscriptionGuard for it). Refuse at load.
|
||||
cfg.validateSubscriptionProfiles();
|
||||
cfg.validateCharters();
|
||||
// CB-548: every architect slot must name a configured workers: profile — the strong-model
|
||||
// backend the future spawn lifecycle would read. A stale reference dies here, not later.
|
||||
cfg.validateArchitects();
|
||||
cfg.validateMembers();
|
||||
|
||||
Path socket = cfg.herdrSocket() != null && !cfg.herdrSocket().isBlank()
|
||||
? Path.of(cfg.herdrSocket())
|
||||
@@ -106,9 +114,9 @@ public final class Bridged {
|
||||
// CB-402: one adapter per configured peer kind, fronted by a composite router. A profile's
|
||||
// `kind:` selects its adapter — claude-code (the default) and opencode partition the profile
|
||||
// set — and the composite dispatches each SPI call to the adapter that owns the profile/pane.
|
||||
Map<String, BridgedConfig.Worker> claudeProfiles = new LinkedHashMap<>();
|
||||
Map<String, BridgedConfig.Worker> opencodeProfiles = new LinkedHashMap<>();
|
||||
cfg.workerProfiles().forEach((name, w) -> {
|
||||
Map<String, BridgedConfig.Profile> claudeProfiles = new LinkedHashMap<>();
|
||||
Map<String, BridgedConfig.Profile> opencodeProfiles = new LinkedHashMap<>();
|
||||
cfg.profiles().forEach((name, w) -> {
|
||||
if (w.isOpenCode()) {
|
||||
opencodeProfiles.put(name, w);
|
||||
} else {
|
||||
@@ -121,27 +129,29 @@ public final class Bridged {
|
||||
// unless opencode is the only kind configured.
|
||||
if (!claudeProfiles.isEmpty() || opencodeProfiles.isEmpty()) {
|
||||
adapters.add(new ClaudeCodeLauncher(agents, spaces, guard,
|
||||
claudeProfiles, cfg.defaultProfile(), System::getenv,
|
||||
cfg.spawnReadyTimeoutMs(), cfg.spawnReadyPollMs()));
|
||||
claudeProfiles, cfg.effectiveDefaultProfile(), System::getenv,
|
||||
cfg.spawnReadyTimeoutMs(), cfg.spawnReadyPollMs(),
|
||||
() -> config.get().fleet()));
|
||||
}
|
||||
if (!opencodeProfiles.isEmpty()) {
|
||||
adapters.add(new OpenCodeLauncher(agents, spaces,
|
||||
opencodeProfiles, cfg.defaultProfile(), System::getenv,
|
||||
cfg.spawnReadyTimeoutMs(), cfg.spawnReadyPollMs()));
|
||||
opencodeProfiles, cfg.effectiveDefaultProfile(), System::getenv,
|
||||
cfg.spawnReadyTimeoutMs(), cfg.spawnReadyPollMs(),
|
||||
() -> config.get().fleet()));
|
||||
}
|
||||
AtomicReference<Function<String, Integer>> liveCountRef = new AtomicReference<>(_ -> 0);
|
||||
PeerLauncher workers = new CompositePeerLauncher(
|
||||
adapters,
|
||||
cfg.defaultProfile(),
|
||||
cfg.workerProfiles(),
|
||||
PlacementPolicies.fromName(cfg.placement()),
|
||||
cfg.effectiveDefaultProfile(),
|
||||
config,
|
||||
profileName -> liveCountRef.get().apply(profileName));
|
||||
// CB-504: under supervision (launchd/systemd) bridged can start before herdr's socket
|
||||
// exists. The client itself is lazy — it connects per call — but the orphan reap below is
|
||||
// the first thing that actually talks to herdr, so without this wait a boot-order race
|
||||
// would crash the daemon into a restart loop. Wait, then degrade rather than die: serving
|
||||
// with /healthz reporting "degraded" is strictly more useful than exiting.
|
||||
if (awaitHerdr(herdr)) {
|
||||
boolean herdrUp = awaitHerdr(herdr);
|
||||
if (herdrUp) {
|
||||
// CB-117: herdr keeps worker panes alive across a daemon restart, and their ids died
|
||||
// with the previous process — reap those leaked orphans now, before we start serving.
|
||||
workers.reapOrphanWorkers();
|
||||
@@ -185,35 +195,59 @@ public final class Bridged {
|
||||
log.info("leads: {} panes recognised {}", leadTerminals.size(), leadTerminals.values());
|
||||
}
|
||||
// CB-531: on top of the static registry, discover leads by the tab labels the operator
|
||||
// writes. Opt-in, so a config with no `leadScan:` block resolves exactly as it did under
|
||||
// CB-530 — the supplier is then a constant and never touches herdr.
|
||||
// writes. CB-557 moved the settings onto the lead they describe, so scanning is on whenever
|
||||
// a `fleet.leaders:` entry exists — with no leads configured the supplier is a constant and
|
||||
// never touches herdr, exactly as a missing `leadScan:` block used to behave.
|
||||
final Supplier<Map<String, String>> leads;
|
||||
if (cfg.leadScan() != null) {
|
||||
var scan = cfg.leadScan();
|
||||
Set<String> workerSpaces = cfg.workerProfiles().values().stream()
|
||||
.map(BridgedConfig.Worker::workspace)
|
||||
var leaders = cfg.fleet().leaders();
|
||||
if (!leaders.isEmpty()) {
|
||||
// One scanner, so one prefix and one interval. Distinct per-lead prefixes would need a
|
||||
// scanner each; until a config actually wants that, take the first entry's settings and
|
||||
// say so, rather than silently honouring one lead's prefix and dropping another's.
|
||||
var scan = leaders.values().iterator().next();
|
||||
Set<String> memberSpaces = cfg.profiles().values().stream()
|
||||
.map(BridgedConfig.Profile::workspace)
|
||||
.filter(Objects::nonNull)
|
||||
.collect(Collectors.toSet());
|
||||
leads = new LeadTabScanner(herdr, scan.tabPrefix(), workerSpaces, leadTerminals,
|
||||
TimeUnit.SECONDS.toNanos(scan.intervalSeconds()), System::nanoTime);
|
||||
log.info("lead scan: tabs labelled '{}…' host a lead (rescan every {}s, worker spaces {} "
|
||||
leads = new LeadTabScanner(herdr, scan.tabPrefix(), memberSpaces, leadTerminals,
|
||||
TimeUnit.SECONDS.toNanos(scan.scanIntervalSeconds()), System::nanoTime);
|
||||
log.info("lead scan: tabs labelled '{}…' host a lead (rescan every {}s, member spaces {} "
|
||||
+ "excluded)",
|
||||
scan.tabPrefix(), scan.intervalSeconds(), workerSpaces);
|
||||
scan.tabPrefix(), scan.scanIntervalSeconds(), memberSpaces);
|
||||
long distinctPrefixes = leaders.values().stream()
|
||||
.map(BridgedConfig.Leader::tabPrefix).distinct().count();
|
||||
if (distinctPrefixes > 1) {
|
||||
log.warn("fleet.leaders declares {} different tabPrefix values; only '{}' is scanned "
|
||||
+ "for. Give every lead the same tabPrefix, or leads under the others "
|
||||
+ "will not be discovered.",
|
||||
distinctPrefixes, scan.tabPrefix());
|
||||
}
|
||||
} else {
|
||||
leads = () -> leadTerminals;
|
||||
}
|
||||
|
||||
// CB-558: start any declared lead that is not already running. After the scanner is built,
|
||||
// because both read the same tab labels and the ordering makes that dependency visible; and
|
||||
// only when herdr answered, because the launcher's whole safety property is that it can
|
||||
// count live leads first — it must never guess and risk a second orchestrator.
|
||||
if (herdrUp && !leaders.isEmpty()) {
|
||||
int launched = new LeadLauncher(agents, spaces, cfg).ensureLeads();
|
||||
if (launched > 0) {
|
||||
log.info("lead auto-launch: {} lead(s) started", launched);
|
||||
}
|
||||
}
|
||||
|
||||
// CB-548: config-declared architect slots. Config supplies only the stable name → profile
|
||||
// map; the terminal → slot binding is owned by the registry and is empty at startup, so no
|
||||
// pane resolves to an architect until the later spawn lifecycle binds one. The registry is
|
||||
// what CallerResolver resolves against and what that lifecycle will read profiles from;
|
||||
// nothing here spawns a slot.
|
||||
ArchitectRegistry architects = new ArchitectRegistry(
|
||||
cfg.architects() == null ? Map.of() : cfg.architects());
|
||||
if (!architects.slots().isEmpty()) {
|
||||
log.info("architect slots: {} configured {} — none bound yet (a slot is idle until the "
|
||||
MemberRegistry members = new MemberRegistry(cfg.fleet());
|
||||
sessions.setMemberLifecycle(members);
|
||||
if (!members.slots().isEmpty()) {
|
||||
log.info("member slots: {} configured {} — none bound yet (a slot is idle until the "
|
||||
+ "spawn lifecycle binds a live terminal to it)",
|
||||
architects.slots().size(), architects.slots().keySet());
|
||||
members.slots().size(), members.slots().keySet());
|
||||
}
|
||||
|
||||
// Status-gated injector (CB-103): the single writer into workers, fed by a poller.
|
||||
@@ -221,9 +255,10 @@ public final class Bridged {
|
||||
// CB-106: a confirmed turn completion resolves a blocked send whose worker never replied.
|
||||
Rendezvous rendezvous = new Rendezvous();
|
||||
CompletionResolver completion = new CompletionResolver(agents, rendezvous);
|
||||
AtomicReference<MessageService> messagesRef = new AtomicReference<>();
|
||||
// CB-113: deliver only to an available worker (its MCP is connected), never its boot window.
|
||||
// CB-301: the manager's presence bridge records availability and drives SPAWNING → READY.
|
||||
WorkerPresence presence = sessions.asPresence();
|
||||
MemberPresence presence = sessions.asPresence();
|
||||
TurnListener turnListener = new TurnListener() {
|
||||
@Override
|
||||
public void onTurnComplete(String target) {
|
||||
@@ -252,11 +287,19 @@ public final class Bridged {
|
||||
public void onTurnFailed(String target) {
|
||||
completion.onTurnFailed(target);
|
||||
sessions.onTurnFailed(target);
|
||||
failTarget(messagesRef, target, "worker turn failed");
|
||||
}
|
||||
|
||||
@Override
|
||||
public void onTurnFailed(String target, String reason) {
|
||||
completion.onTurnFailed(target, reason);
|
||||
sessions.onTurnFailed(target);
|
||||
failTarget(messagesRef, target, reason);
|
||||
}
|
||||
};
|
||||
Injector injector = new Injector(agents, turnListener, deliverableTo(presence, leads),
|
||||
presence::forget);
|
||||
StatusPoller poller = new StatusPoller(agents, injector, INJECT_POLL_MILLIS);
|
||||
StatusPoller poller = new StatusPoller(agents, injector, Injector.POLL_INTERVAL_MILLIS);
|
||||
poller.start();
|
||||
|
||||
// CB-307: reply inbox. A broker: block (with a uri) selects the AMQP-backed durable adapter;
|
||||
@@ -296,8 +339,46 @@ public final class Bridged {
|
||||
Metrics metrics = BridgedMetrics.create(sessions, replyInbox);
|
||||
var pushLoop = new ReplyPushLoop(primaryRegistry, agents, replyInbox,
|
||||
pushScheduler, maxReminders, backoffMs, metrics);
|
||||
// CB-551: idle-lead heartbeat. Opt-in; absent `leadHeartbeat:` this is never constructed, so
|
||||
// an upgraded daemon cannot silently start spending subscription on nudging an idle lead.
|
||||
// It has its own single-thread scheduler and holds its own scheduler shutdown via close().
|
||||
final LeadHeartbeatLoop heartbeat;
|
||||
var heartbeatScheduler = Executors.newSingleThreadScheduledExecutor(r ->
|
||||
Thread.ofVirtual().name("bridge-heartbeat-").unstarted(r));
|
||||
if (cfg.leadHeartbeat() != null) {
|
||||
var hb = cfg.leadHeartbeat();
|
||||
heartbeat = new LeadHeartbeatLoop(primaryRegistry, agents, replyInbox, sessions::roster,
|
||||
pushLoop, heartbeatScheduler, System::nanoTime,
|
||||
TimeUnit.SECONDS.toNanos(hb.idleAfterSeconds()), hb.backoffMs(), hb.quietNudgeCap(),
|
||||
metrics);
|
||||
heartbeat.start();
|
||||
} else {
|
||||
heartbeat = null;
|
||||
heartbeatScheduler.shutdownNow();
|
||||
}
|
||||
MessageService messages = new MessageService(agents, injector, rendezvous, replyInbox,
|
||||
pushLoop, metrics);
|
||||
messagesRef.set(messages);
|
||||
|
||||
// Health is a slow whole-fleet observer. Keep it separate from the 250ms delivery poller.
|
||||
final FleetHealthMonitor healthMonitor;
|
||||
var healthScheduler = Executors.newSingleThreadScheduledExecutor(r ->
|
||||
Thread.ofVirtual().name("bridge-health-").unstarted(r));
|
||||
if (cfg.health() != null && cfg.health().isEnabled()) {
|
||||
healthMonitor = new FleetHealthMonitor(agents, sessions::roster, messages, healthScheduler,
|
||||
System::nanoTime, cfg.health().intervalOrDefault(), messages::abandon);
|
||||
String coverage = FleetHealthMonitor.coverage(true,
|
||||
cfg.health().notifications() != null && cfg.health().notifications().configured());
|
||||
if ("detection-only".equals(coverage)) {
|
||||
log.warn("fleet health: {} (no notification sink configured)", coverage);
|
||||
} else {
|
||||
log.info("fleet health: {}", coverage);
|
||||
}
|
||||
healthMonitor.start();
|
||||
} else {
|
||||
healthMonitor = null;
|
||||
healthScheduler.shutdownNow();
|
||||
}
|
||||
|
||||
// CB-520: the reply inbox only consumes for agents this gateway owns. own on acquire,
|
||||
// release on teardown. Do this before CB-516 so the inbox is owned before any reply can land.
|
||||
@@ -325,18 +406,35 @@ public final class Bridged {
|
||||
throw new IllegalStateException("auth.mode=token but env var " + cfg.auth().tokenEnv()
|
||||
+ " is unset or empty — export it before starting bridged");
|
||||
}
|
||||
callers = CallerResolver.withLeadsAndArchitects(identity, true, token, leads,
|
||||
architects::snapshot);
|
||||
callers = CallerResolver.withLeadsAndMembers(identity, true, token, leads, members);
|
||||
log.info("auth: token mode (bearer required for non-worker callers, env {})",
|
||||
cfg.auth().tokenEnv());
|
||||
} else {
|
||||
callers = CallerResolver.withLeadsAndArchitects(identity, false, null, leads,
|
||||
architects::snapshot);
|
||||
callers = CallerResolver.withLeadsAndMembers(identity, false, null, leads, members);
|
||||
log.info("auth: loopback-trust (any loopback non-worker caller is the primary)");
|
||||
}
|
||||
|
||||
BridgeMcp mcp = new BridgeMcp(messages, workers, sessions, identity, presence,
|
||||
primaryRegistry, callers, metrics);
|
||||
primaryRegistry, callers, metrics, new BridgeMcp.CapacitySource(profile -> liveCountRef.get().apply(profile),
|
||||
profile -> {
|
||||
var configured = config.get().profiles().get(profile);
|
||||
return configured == null ? null : configured.maxLoad();
|
||||
}, () -> config.get().profiles().keySet(), System::nanoTime),
|
||||
new BridgeMcp.HealthCoverageSource(() -> {
|
||||
var health = config.get().health();
|
||||
return FleetHealthMonitor.coverage(health != null && health.isEnabled(),
|
||||
health != null && health.notifications() != null && health.notifications().configured());
|
||||
}));
|
||||
|
||||
// CB-559: opt-in config reload. With no `configReload:` block nothing is constructed, so an
|
||||
// upgraded daemon behaves exactly as before — the file is read once at boot and never again.
|
||||
final ConfigWatcher configWatcher;
|
||||
if (cfg.configReload() != null && cfg.configReload().isEnabled()) {
|
||||
configWatcher = new ConfigWatcher(config, cfg.configReload().intervalSeconds());
|
||||
configWatcher.start();
|
||||
} else {
|
||||
configWatcher = null;
|
||||
}
|
||||
|
||||
// CB-303 part 3: single ordered shutdown hook. Drain sessions first while herdr is still
|
||||
// open (so releases reach the daemon), then stop poller/message/mcp/reaper, and close herdr
|
||||
@@ -346,6 +444,9 @@ public final class Bridged {
|
||||
poller.stop();
|
||||
messages.close();
|
||||
pushLoop.close();
|
||||
if (heartbeat != null) heartbeat.close(); // CB-551: stop the idle-lead heartbeat scheduler
|
||||
if (healthMonitor != null) healthMonitor.stop();
|
||||
if (configWatcher != null) configWatcher.stop(); // CB-559: stop polling the config file
|
||||
mcp.close();
|
||||
if (reaper != null) reaper.stop();
|
||||
// Release the broker connection last among message resources (no-op for the in-memory inbox).
|
||||
@@ -367,27 +468,36 @@ public final class Bridged {
|
||||
}
|
||||
|
||||
/**
|
||||
* The {@link Injector}'s readiness gate (CB-534): a target is deliverable if it is a worker whose
|
||||
* agent has connected the bridge MCP, <em>or</em> a lead.
|
||||
* The {@link Injector}'s readiness gate (CB-534): a target is deliverable if it is a spawned
|
||||
* member whose agent has connected the bridge MCP, <em>or</em> a lead.
|
||||
*
|
||||
* <p>The gate exists for one reason — to hold a delivery out of a <em>spawned</em> worker's boot
|
||||
* <p>The gate exists for one reason — to hold a delivery out of a <em>spawned</em> member's boot
|
||||
* window, where herdr already reports {@code idle} but the TUI would drop an injected paste. That
|
||||
* hazard is a property of spawning. A lead is never spawned: the operator started it and named it
|
||||
* (or labelled its tab) only once it was up, so there is no boot window to guard.
|
||||
*
|
||||
* <p>A lead is also never enrolled in {@link WorkerPresence} — {@code BridgeMcp} marks presence
|
||||
* only for a worker, deliberately, since that map doubles as the worker roster's availability
|
||||
* signal and a lead counted there would show up as an available worker. So without the second
|
||||
* disjunct a lead is permanently un-deliverable: every lead→lead send sat on the gate for
|
||||
* {@code READINESS_GRACE_POLLS} (~60s) and then failed having never been typed into the pane.
|
||||
* <p>A lead is also never enrolled in {@link MemberPresence} — {@code BridgeMcp} marks presence
|
||||
* for every spawned member (worker and architect), deliberately, since that map doubles as the
|
||||
* member roster's availability signal and a lead counted there would show up as an available
|
||||
* member. So without the second disjunct a lead is permanently un-deliverable: every
|
||||
* lead→lead send sat on the gate for {@code READINESS_GRACE_POLLS} (~60s) and then failed
|
||||
* having never been typed into the pane.
|
||||
*
|
||||
* <p>The lead set is read through the supplier on each call rather than snapshotted, so a lead
|
||||
* discovered by {@code leadScan} after startup becomes deliverable without a restart.
|
||||
*/
|
||||
static Predicate<String> deliverableTo(WorkerPresence presence, Supplier<Map<String, String>> leads) {
|
||||
static Predicate<String> deliverableTo(MemberPresence presence, Supplier<Map<String, String>> leads) {
|
||||
return target -> presence.isPresent(target) || leads.get().containsKey(target);
|
||||
}
|
||||
|
||||
/** Fail outstanding tickets without changing the member lifecycle or worktree state. */
|
||||
private static void failTarget(AtomicReference<MessageService> messagesRef, String target, String reason) {
|
||||
MessageService messages = messagesRef.get();
|
||||
if (messages != null) {
|
||||
messages.abandon(target, reason);
|
||||
}
|
||||
}
|
||||
|
||||
/**
|
||||
* Poll herdr's {@code ping} until it answers or {@link #HERDR_WAIT_SECONDS} elapses (CB-504).
|
||||
*
|
||||
|
||||
@@ -1,10 +1,12 @@
|
||||
package dev.ltms.bridged.auth;
|
||||
|
||||
import dev.ltms.bridged.mcp.ConnectionIdentity;
|
||||
import dev.ltms.bridged.peer.MemberRole;
|
||||
|
||||
import java.nio.charset.StandardCharsets;
|
||||
import java.security.MessageDigest;
|
||||
import java.util.Map;
|
||||
import java.util.function.Function;
|
||||
import java.util.function.Supplier;
|
||||
|
||||
/**
|
||||
@@ -60,14 +62,16 @@ public final class CallerResolver {
|
||||
* in {@code Bridged} reads a constant from config, which is the degenerate live case.
|
||||
*/
|
||||
private final Supplier<Map<String, String>> architectTerminals;
|
||||
private final Function<String, MemberRole> memberSlotRoles;
|
||||
private final Function<String, String> memberSlotNames;
|
||||
|
||||
/** Loopback-trust resolver: no token required, historical behaviour. */
|
||||
public CallerResolver(ConnectionIdentity identity) {
|
||||
/** Loopback-trust resolver: no token required, historical behaviour. Test-only. */
|
||||
CallerResolver(ConnectionIdentity identity) {
|
||||
this(identity, false, null, Map.of());
|
||||
}
|
||||
|
||||
/** As {@link #CallerResolver(ConnectionIdentity, boolean, String, Map)} with no leads pinned. */
|
||||
public CallerResolver(ConnectionIdentity identity, boolean tokenMode, String token) {
|
||||
/** As {@link #CallerResolver(ConnectionIdentity, boolean, String, Map)} with no leads pinned. Test-only. */
|
||||
CallerResolver(ConnectionIdentity identity, boolean tokenMode, String token) {
|
||||
this(identity, tokenMode, token, Map.of());
|
||||
}
|
||||
|
||||
@@ -82,8 +86,8 @@ public final class CallerResolver {
|
||||
* @param pinnedPrimaryTerminal the primary's own herdr {@code terminal_id}
|
||||
* ({@code null}/blank = unpinned)
|
||||
*/
|
||||
public static CallerResolver pinnedTo(ConnectionIdentity identity, boolean tokenMode,
|
||||
String token, String pinnedPrimaryTerminal) {
|
||||
static CallerResolver pinnedTo(ConnectionIdentity identity, boolean tokenMode,
|
||||
String token, String pinnedPrimaryTerminal) {
|
||||
return new CallerResolver(identity, tokenMode, token,
|
||||
pinnedPrimaryTerminal == null || pinnedPrimaryTerminal.isBlank()
|
||||
? Map.of() : Map.of(pinnedPrimaryTerminal, "primary"));
|
||||
@@ -98,21 +102,11 @@ public final class CallerResolver {
|
||||
* {@link Role#PRIMARY} — rather than a worker. Empty = nothing pinned,
|
||||
* so every pane resolves as a worker.
|
||||
*/
|
||||
public CallerResolver(ConnectionIdentity identity, boolean tokenMode, String token,
|
||||
Map<String, String> leadTerminals) {
|
||||
CallerResolver(ConnectionIdentity identity, boolean tokenMode, String token,
|
||||
Map<String, String> leadTerminals) {
|
||||
this(identity, tokenMode, token, fixed(leadTerminals), null);
|
||||
}
|
||||
|
||||
/**
|
||||
* Map-form of both registries (CB-548): lead terminals and the initial architect terminal
|
||||
* bindings, each snapshotted at construction (a handed-over map is not offered as live state).
|
||||
*/
|
||||
public CallerResolver(ConnectionIdentity identity, boolean tokenMode, String token,
|
||||
Map<String, String> leadTerminals,
|
||||
Map<String, String> architectTerminals) {
|
||||
this(identity, tokenMode, token, fixed(leadTerminals), fixed(architectTerminals));
|
||||
}
|
||||
|
||||
/**
|
||||
* Live-registry form: {@code leadTerminals} is consulted on every resolve, so leads discovered
|
||||
* after startup (CB-531's tab scan) take effect without a restart.
|
||||
@@ -121,26 +115,26 @@ public final class CallerResolver {
|
||||
* {@link #pinnedTo}: {@code Map} and {@code Supplier} overloads are ambiguous for a literal
|
||||
* {@code null}.
|
||||
*/
|
||||
public static CallerResolver withLeads(ConnectionIdentity identity, boolean tokenMode,
|
||||
String token,
|
||||
Supplier<Map<String, String>> leadTerminals) {
|
||||
static CallerResolver withLeads(ConnectionIdentity identity, boolean tokenMode,
|
||||
String token,
|
||||
Supplier<Map<String, String>> leadTerminals) {
|
||||
return new CallerResolver(identity, tokenMode, token, leadTerminals, null);
|
||||
}
|
||||
|
||||
/**
|
||||
* Live-registry form for both {@code leadTerminals} and the CB-548 architect registry: both
|
||||
* are consulted on every resolve, so a slot binding injected after startup takes effect
|
||||
* without a restart.
|
||||
* Live registry form that can confirm a bound slot is an architect slot.
|
||||
*
|
||||
* <p>A static factory rather than a constructor overload, for the same reason as
|
||||
* {@link #pinnedTo}: too many {@code Map}/{@code Supplier} combinations to make {@code null}
|
||||
* unambiguous.
|
||||
* <p>This is the only public construction path. It keeps terminal bindings and slot roles in
|
||||
* the same {@link MemberRegistry}, so a configured architect can resolve as an architect.
|
||||
*/
|
||||
public static CallerResolver withLeadsAndArchitects(ConnectionIdentity identity,
|
||||
boolean tokenMode, String token,
|
||||
Supplier<Map<String, String>> leadTerminals,
|
||||
Supplier<Map<String, String>> architectTerminals) {
|
||||
return new CallerResolver(identity, tokenMode, token, leadTerminals, architectTerminals);
|
||||
public static CallerResolver withLeadsAndMembers(ConnectionIdentity identity,
|
||||
boolean tokenMode, String token,
|
||||
Supplier<Map<String, String>> leadTerminals,
|
||||
MemberRegistry members) {
|
||||
return new CallerResolver(identity, tokenMode, token, leadTerminals,
|
||||
members == null ? null : members::snapshot,
|
||||
members == null ? null : members::roleForSlot,
|
||||
members == null ? null : members::nameForSlot);
|
||||
}
|
||||
|
||||
private static Supplier<Map<String, String>> fixed(Map<String, String> leadTerminals) {
|
||||
@@ -151,6 +145,21 @@ public final class CallerResolver {
|
||||
private CallerResolver(ConnectionIdentity identity, boolean tokenMode, String token,
|
||||
Supplier<Map<String, String>> leadTerminals,
|
||||
Supplier<Map<String, String>> architectTerminals) {
|
||||
this(identity, tokenMode, token, leadTerminals, architectTerminals, null);
|
||||
}
|
||||
|
||||
private CallerResolver(ConnectionIdentity identity, boolean tokenMode, String token,
|
||||
Supplier<Map<String, String>> leadTerminals,
|
||||
Supplier<Map<String, String>> architectTerminals,
|
||||
Function<String, MemberRole> memberSlotRoles) {
|
||||
this(identity, tokenMode, token, leadTerminals, architectTerminals, memberSlotRoles, null);
|
||||
}
|
||||
|
||||
private CallerResolver(ConnectionIdentity identity, boolean tokenMode, String token,
|
||||
Supplier<Map<String, String>> leadTerminals,
|
||||
Supplier<Map<String, String>> architectTerminals,
|
||||
Function<String, MemberRole> memberSlotRoles,
|
||||
Function<String, String> memberSlotNames) {
|
||||
if (tokenMode && (token == null || token.isBlank())) {
|
||||
throw new IllegalArgumentException(
|
||||
"auth.mode=token requires a non-empty token; check that the env var named by "
|
||||
@@ -161,6 +170,8 @@ public final class CallerResolver {
|
||||
this.expectedToken = tokenMode ? token.getBytes(StandardCharsets.UTF_8) : null;
|
||||
this.leadTerminals = leadTerminals == null ? Map::of : leadTerminals;
|
||||
this.architectTerminals = architectTerminals == null ? Map::of : architectTerminals;
|
||||
this.memberSlotRoles = memberSlotRoles == null ? _ -> null : memberSlotRoles;
|
||||
this.memberSlotNames = memberSlotNames == null ? Function.identity() : memberSlotNames;
|
||||
}
|
||||
|
||||
/**
|
||||
@@ -183,7 +194,7 @@ public final class CallerResolver {
|
||||
* here but would not <em>resolve</em> (or the reverse) cannot drift apart. Live for the same
|
||||
* reason as {@link #leads()}.
|
||||
*/
|
||||
public Map<String, String> architects() {
|
||||
public Map<String, String> members() {
|
||||
return architectTerminals.get();
|
||||
}
|
||||
|
||||
@@ -206,11 +217,12 @@ public final class CallerResolver {
|
||||
return Principal.leader(lead, c.terminal(), c.pid());
|
||||
}
|
||||
String slot = architectTerminals.get().get(c.terminal());
|
||||
if (slot != null) {
|
||||
if (slot != null && memberSlotRoles.apply(slot) == MemberRole.ARCHITECT) {
|
||||
// The config/live binding names this pane as an architect slot's own. Same
|
||||
// unforgeable pane mapping; the live binding, never a request argument, decides.
|
||||
// Checked before the generic worker fallback, per the CB-548 precedence order.
|
||||
return Principal.architect(slot, c.terminal(), c.pid());
|
||||
// Check the slot role too: this defence in depth prevents a bad lifecycle bind from
|
||||
// escalating a dev or reviewer into an architect. Checked before the worker fallback.
|
||||
return Principal.architect(memberSlotNames.apply(slot), c.terminal(), c.pid());
|
||||
}
|
||||
return Principal.worker(c.terminal(), c.pid()); // unforgeable; never token-gated
|
||||
}
|
||||
|
||||
@@ -0,0 +1,21 @@
|
||||
package dev.ltms.bridged.auth;
|
||||
|
||||
import dev.ltms.bridged.peer.MemberRole;
|
||||
|
||||
/** Optional session lifecycle hook for live member-slot bindings. */
|
||||
public interface MemberLifecycle {
|
||||
|
||||
MemberLifecycle NONE = new MemberLifecycle() {
|
||||
@Override
|
||||
public void acquired(MemberRole role, String profile, String terminal) {
|
||||
}
|
||||
|
||||
@Override
|
||||
public void released(String terminal) {
|
||||
}
|
||||
};
|
||||
|
||||
void acquired(MemberRole role, String profile, String terminal);
|
||||
|
||||
void released(String terminal);
|
||||
}
|
||||
+100
-10
@@ -1,9 +1,15 @@
|
||||
package dev.ltms.bridged.auth;
|
||||
|
||||
import dev.ltms.bridged.config.BridgedConfig;
|
||||
import dev.ltms.bridged.peer.MemberRole;
|
||||
import org.slf4j.Logger;
|
||||
import org.slf4j.LoggerFactory;
|
||||
|
||||
import java.util.Collections;
|
||||
import java.util.HashMap;
|
||||
import java.util.LinkedHashMap;
|
||||
import java.util.Map;
|
||||
import java.util.Objects;
|
||||
|
||||
/**
|
||||
* The architect-slot registry (CB-548): every gateway-local architect name and the strong-model
|
||||
@@ -26,21 +32,64 @@ import java.util.Map;
|
||||
* exposes the map the resolver resolves against plus the profile lookup lifecycle will call.
|
||||
* Nothing here creates or manages an architect session.
|
||||
*/
|
||||
public final class ArchitectRegistry {
|
||||
public final class MemberRegistry implements MemberLifecycle {
|
||||
|
||||
private final Map<String, BridgedConfig.Architect> slots;
|
||||
/** Live {@code terminal_id → slot name}; guarded by {@code this}. */
|
||||
private final Map<String, String> terminalToSlot = new HashMap<>();
|
||||
private static final Logger log = LoggerFactory.getLogger(MemberRegistry.class);
|
||||
|
||||
public ArchitectRegistry(Map<String, BridgedConfig.Architect> slots) {
|
||||
this.slots = slots == null ? Map.of() : Map.copyOf(slots);
|
||||
/**
|
||||
* One flattened {@code fleet:} entry.
|
||||
*
|
||||
* <p>Flattened because a slot name is unique only <em>within</em> its pool — {@code sonnet} may
|
||||
* legitimately be both a developer and a reviewer — while a terminal binds to exactly one thing.
|
||||
* The qualified {@link #key()} is what that binding uses.
|
||||
*
|
||||
* @param name the slot's key inside its pool
|
||||
* @param role the pool it came from
|
||||
* @param profile the backend it runs on
|
||||
*/
|
||||
public record Entry(String name, MemberRole role, String profile) {
|
||||
/** {@code "architect:opus"} — unique across pools, unlike {@link #name()}. */
|
||||
public String key() {
|
||||
return role.wireName() + ":" + name;
|
||||
}
|
||||
}
|
||||
|
||||
/** The configured slots, keyed by gateway-local unique name. Unmodifiable snapshot. */
|
||||
public Map<String, BridgedConfig.Architect> slots() {
|
||||
private final Map<String, Entry> slots;
|
||||
/** Live {@code terminal_id → qualified slot key}; guarded by {@code this}. */
|
||||
private final Map<String, String> terminalToSlot = new HashMap<>();
|
||||
|
||||
/** Flatten every role pool in {@code fleet} into one registry. Leaders are not members. */
|
||||
public MemberRegistry(BridgedConfig.Fleet fleet) {
|
||||
Map<String, Entry> flat = new LinkedHashMap<>();
|
||||
if (fleet != null) {
|
||||
for (MemberRole role : MemberRole.values()) {
|
||||
fleet.pool(role).forEach((name, slot) -> {
|
||||
if (slot != null) {
|
||||
Entry e = new Entry(name, role, slot.profile());
|
||||
flat.put(e.key(), e);
|
||||
}
|
||||
});
|
||||
}
|
||||
}
|
||||
this.slots = Collections.unmodifiableMap(flat);
|
||||
}
|
||||
|
||||
/** The configured slots, keyed by qualified {@link Entry#key()}. Unmodifiable snapshot. */
|
||||
public Map<String, Entry> slots() {
|
||||
return slots;
|
||||
}
|
||||
|
||||
/** The slots belonging to {@code role}, in definition order. */
|
||||
public Map<String, Entry> slotsFor(MemberRole role) {
|
||||
Map<String, Entry> out = new LinkedHashMap<>();
|
||||
slots.forEach((key, e) -> {
|
||||
if (e.role() == role) {
|
||||
out.put(key, e);
|
||||
}
|
||||
});
|
||||
return Collections.unmodifiableMap(out);
|
||||
}
|
||||
|
||||
/**
|
||||
* An immutable copy of the live {@code terminal_id → slot name} bindings.
|
||||
*
|
||||
@@ -71,8 +120,20 @@ public final class ArchitectRegistry {
|
||||
* declares none
|
||||
*/
|
||||
public String profileForSlot(String slotName) {
|
||||
BridgedConfig.Architect a = slots.get(slotName);
|
||||
return (a == null || a.profile() == null) ? null : a.profile();
|
||||
Entry e = slots.get(slotName);
|
||||
return (e == null || e.profile() == null) ? null : e.profile();
|
||||
}
|
||||
|
||||
/** The role a qualified slot key belongs to, or {@code null} when the key is unknown. */
|
||||
public MemberRole roleForSlot(String slotName) {
|
||||
Entry e = slots.get(slotName);
|
||||
return e == null ? null : e.role();
|
||||
}
|
||||
|
||||
/** The unqualified configured name for a slot, or {@code null} if it is unknown. */
|
||||
public String nameForSlot(String slotName) {
|
||||
Entry e = slots.get(slotName);
|
||||
return e == null ? null : e.name();
|
||||
}
|
||||
|
||||
/** True when {@code slotName} is a configured architect slot. */
|
||||
@@ -139,4 +200,33 @@ public final class ArchitectRegistry {
|
||||
return true;
|
||||
}
|
||||
}
|
||||
|
||||
/**
|
||||
* Bind only architect sessions to a free slot with the resolved profile.
|
||||
*
|
||||
* <p>The role check is lifecycle policy. {@link CallerResolver} repeats it when resolving a
|
||||
* binding, so a later lifecycle regression cannot turn a worker into an architect.
|
||||
*/
|
||||
@Override
|
||||
public void acquired(MemberRole role, String profile, String terminal) {
|
||||
if (role != MemberRole.ARCHITECT || terminal == null || terminal.isBlank()) {
|
||||
return;
|
||||
}
|
||||
// slotsFor preserves definition order, so duplicate-profile slots use the first free one.
|
||||
for (Entry entry : slotsFor(MemberRole.ARCHITECT).values()) {
|
||||
if (Objects.equals(profile, entry.profile()) && bind(entry.key(), terminal)) {
|
||||
return;
|
||||
}
|
||||
}
|
||||
log.info("member slot: no free architect slot for profile={}; session remains a worker", profile);
|
||||
}
|
||||
|
||||
/** Unbind a released terminal using the compare-safe registry operation. */
|
||||
@Override
|
||||
public void released(String terminal) {
|
||||
String slot = slotForTerminal(terminal);
|
||||
if (slot != null) {
|
||||
unbind(slot, terminal);
|
||||
}
|
||||
}
|
||||
}
|
||||
@@ -48,7 +48,7 @@ public record Principal(Role role, String terminal, long pid, String name) {
|
||||
* made a lead unaddressable: {@link #ownsSession} could never be true for it, so
|
||||
* {@code bridge_reply} was refused and one lead could send to another but never be answered.
|
||||
* The terminal now means "which pane is this caller", the presence map keys on
|
||||
* {@link #isWorker()} instead, and a lead is a peer that can both send and receive.
|
||||
* {@link #isSpawnedMember()} instead, and a lead is a peer that can both send and receive.
|
||||
*/
|
||||
public static Principal leader(String name, String terminal, long pid) {
|
||||
return new Principal(Role.PRIMARY, terminal, pid, name);
|
||||
@@ -85,6 +85,16 @@ public record Principal(Role role, String terminal, long pid, String name) {
|
||||
return role == Role.WORKER;
|
||||
}
|
||||
|
||||
/**
|
||||
* Whether this caller is a spawned member with its own pane.
|
||||
*
|
||||
* <p>Both workers and architects are spawned members. A lead is excluded because recording it
|
||||
* as present would count it as an available member in the roster.
|
||||
*/
|
||||
public boolean isSpawnedMember() {
|
||||
return role == Role.WORKER || role == Role.ARCHITECT;
|
||||
}
|
||||
|
||||
public boolean isAnonymous() {
|
||||
return role == Role.ANONYMOUS;
|
||||
}
|
||||
|
||||
File diff suppressed because it is too large
Load Diff
@@ -0,0 +1,270 @@
|
||||
package dev.ltms.bridged.config;
|
||||
|
||||
import org.slf4j.Logger;
|
||||
import org.slf4j.LoggerFactory;
|
||||
|
||||
import java.nio.file.Path;
|
||||
import java.util.ArrayList;
|
||||
import java.util.LinkedHashSet;
|
||||
import java.util.List;
|
||||
import java.util.Map;
|
||||
import java.util.Objects;
|
||||
import java.util.Set;
|
||||
import java.util.concurrent.atomic.AtomicReference;
|
||||
import java.util.function.Supplier;
|
||||
|
||||
/**
|
||||
* The daemon's live configuration, re-readable without a restart (CB-559).
|
||||
*
|
||||
* <p>Consumers hold this, not a {@link BridgedConfig}, and read through {@link #get()} at the point
|
||||
* of use. A component that captures {@code ref.get()} into a field at construction has opted out of
|
||||
* reload — which is sometimes right (see <em>deferred</em> below), but it must then be a deliberate
|
||||
* choice rather than an accident of where the field was initialised.
|
||||
*
|
||||
* <h2>Not every key can change under a running daemon</h2>
|
||||
* Keys fall into three classes, and the difference is about what already exists when the reload
|
||||
* happens — not about how important the key is.
|
||||
*
|
||||
* <ul>
|
||||
* <li><strong>Hot</strong> — re-read per use, so a reload takes effect on the next spawn:
|
||||
* {@code fleet:} (every role pool, {@code charters}, and {@code tabLabel}),
|
||||
* {@code placement:}, and an existing profile's {@code weight} / {@code maxLoad}. Those
|
||||
* three are read through a supplier on {@code CompositePeerLauncher}, which is what makes
|
||||
* them hot — not the fact that they are config.</li>
|
||||
* <li><strong>Deferred</strong> — accepted into the new snapshot, but the wiring built at startup
|
||||
* keeps the old value until a restart: {@code lifecycle:}, {@code leadHeartbeat:},
|
||||
* {@code spawnReadyTimeoutMs} / {@code spawnReadyPollMs}, {@code guard:},
|
||||
* {@code worktreeRoot:}, adding or removing a profile (a new backend needs its own launcher,
|
||||
* which is constructed once), <em>and an existing profile's launch settings</em> —
|
||||
* {@code model}, {@code baseUrl}, {@code argv}, {@code env}, {@code mcpUrl} and the rest.
|
||||
* {@code HerdrPeerLauncher} takes {@code Map.copyOf(profiles)} at construction and resolves
|
||||
* each spawn out of that copy, so those never reach a launch until the daemon restarts. A
|
||||
* reload logs these rather than pretending they applied.</li>
|
||||
* <li><strong>Cold</strong> — cannot change at all under a running daemon: {@code bind:},
|
||||
* {@code herdrSocket:}, {@code broker:} and {@code auth:}. The socket is bound, the broker
|
||||
* connection is open, and the auth mode decides who may reach the port that is already
|
||||
* listening.</li>
|
||||
* </ul>
|
||||
*
|
||||
* <p><strong>A cold change refuses the whole reload.</strong> Not the hot half applied and the cold
|
||||
* half warned about: that would leave the running daemon in a state matching no file on disk, which
|
||||
* is the worst thing a reload can do to an operator debugging one. Refusing keeps the invariant that
|
||||
* the live config is always some version of the file, and the message names the keys that must
|
||||
* change through a restart.
|
||||
*
|
||||
* <p>A reload that fails to parse or fails validation is also refused, and the previous config keeps
|
||||
* running. A config file being edited is normally read once mid-save; degrading a working daemon
|
||||
* because it caught a half-written file would be a bad trade.
|
||||
*/
|
||||
public final class ConfigRef implements Supplier<BridgedConfig> {
|
||||
|
||||
private static final Logger log = LoggerFactory.getLogger(ConfigRef.class);
|
||||
|
||||
/** Keys that cannot change under a running daemon — see the class doc. */
|
||||
private static final Set<String> COLD_KEYS =
|
||||
Set.of("bind", "herdrSocket", "broker", "auth");
|
||||
|
||||
private final Path path;
|
||||
private final AtomicReference<BridgedConfig> current;
|
||||
|
||||
public ConfigRef(Path path, BridgedConfig initial) {
|
||||
this.path = path;
|
||||
this.current = new AtomicReference<>(Objects.requireNonNull(initial, "initial config"));
|
||||
}
|
||||
|
||||
/** A fixed reference that never reloads — for tests and for wiring built from a config in code. */
|
||||
public static ConfigRef fixed(BridgedConfig cfg) {
|
||||
return new ConfigRef(null, cfg);
|
||||
}
|
||||
|
||||
/** The live configuration. Read this per use; do not cache it in a field. */
|
||||
@Override
|
||||
public BridgedConfig get() {
|
||||
return current.get();
|
||||
}
|
||||
|
||||
/** The file this ref reloads from, or {@code null} for a {@link #fixed} ref. */
|
||||
public Path path() {
|
||||
return path;
|
||||
}
|
||||
|
||||
/**
|
||||
* What a reload attempt did.
|
||||
*
|
||||
* @param applied true when the new config is now live
|
||||
* @param coldKeys cold keys whose value changed, which is why an unapplied reload was refused
|
||||
* @param deferred keys that changed and were accepted, but whose effect waits for a restart
|
||||
* @param error the parse or validation failure that refused the reload, else {@code null}
|
||||
*/
|
||||
public record Outcome(boolean applied, List<String> coldKeys, List<String> deferred,
|
||||
String error) {
|
||||
|
||||
public Outcome {
|
||||
coldKeys = List.copyOf(coldKeys);
|
||||
deferred = List.copyOf(deferred);
|
||||
}
|
||||
|
||||
static Outcome refusedCold(List<String> keys) {
|
||||
return new Outcome(false, keys, List.of(), null);
|
||||
}
|
||||
|
||||
static Outcome failed(String error) {
|
||||
return new Outcome(false, List.of(), List.of(), error);
|
||||
}
|
||||
|
||||
/** A one-line summary for the operator — the reason, not just the verdict. */
|
||||
public String summary() {
|
||||
if (error != null) {
|
||||
return "config reload refused — " + error;
|
||||
}
|
||||
if (!applied) {
|
||||
return "config reload refused — these keys cannot change under a running daemon: "
|
||||
+ String.join(", ", coldKeys) + ". Restart bridged to apply them.";
|
||||
}
|
||||
if (!deferred.isEmpty()) {
|
||||
return "config reloaded; these changes need a restart to take effect: "
|
||||
+ String.join(", ", deferred);
|
||||
}
|
||||
return "config reloaded";
|
||||
}
|
||||
}
|
||||
|
||||
/**
|
||||
* Re-read the file, validate it, and swap it in when nothing cold changed.
|
||||
*
|
||||
* <p>Never throws: a reload is a best-effort operation on a daemon that is already serving, and
|
||||
* a bad edit must not take it down. Every failure path leaves the previous config live and is
|
||||
* reported through the returned {@link Outcome}.
|
||||
*/
|
||||
public Outcome reload() {
|
||||
if (path == null) {
|
||||
return Outcome.failed("this config was built in code and has no file to reload from");
|
||||
}
|
||||
BridgedConfig old = current.get();
|
||||
BridgedConfig fresh;
|
||||
try {
|
||||
fresh = BridgedConfig.load(path);
|
||||
// The same gate startup runs. A config that would have refused to boot must not be able
|
||||
// to slip in through a reload — that is how a daemon ends up in a state it could never
|
||||
// have started in, which is the hardest kind to debug.
|
||||
fresh.validateAuthExposure();
|
||||
fresh.validateLeadTabPrefixes();
|
||||
fresh.validateSubscriptionProfiles();
|
||||
fresh.validateCharters();
|
||||
fresh.validateMembers();
|
||||
} catch (RuntimeException e) {
|
||||
String msg = e.getMessage() == null ? e.toString() : e.getMessage();
|
||||
log.warn("config reload from {} refused, keeping the running config: {}", path, msg);
|
||||
return Outcome.failed(msg);
|
||||
}
|
||||
|
||||
List<String> cold = changedColdKeys(old, fresh);
|
||||
if (!cold.isEmpty()) {
|
||||
Outcome out = Outcome.refusedCold(cold);
|
||||
log.warn(out.summary());
|
||||
return out;
|
||||
}
|
||||
|
||||
List<String> deferred = changedDeferredKeys(old, fresh);
|
||||
current.set(fresh);
|
||||
Outcome out = new Outcome(true, List.of(), deferred, null);
|
||||
log.info(out.summary());
|
||||
return out;
|
||||
}
|
||||
|
||||
/** Cold keys whose value differs between the running config and the candidate. */
|
||||
private static List<String> changedColdKeys(BridgedConfig old, BridgedConfig fresh) {
|
||||
List<String> changed = new ArrayList<>();
|
||||
if (!Objects.equals(old.bind(), fresh.bind())) {
|
||||
changed.add("bind");
|
||||
}
|
||||
if (!Objects.equals(old.herdrSocket(), fresh.herdrSocket())) {
|
||||
changed.add("herdrSocket");
|
||||
}
|
||||
if (!Objects.equals(old.broker(), fresh.broker())) {
|
||||
changed.add("broker");
|
||||
}
|
||||
if (!Objects.equals(old.auth(), fresh.auth())) {
|
||||
changed.add("auth");
|
||||
}
|
||||
// Kept in step with COLD_KEYS so the doc and the code cannot drift apart silently.
|
||||
assert COLD_KEYS.containsAll(changed) : "a cold key was reported that COLD_KEYS omits";
|
||||
return changed;
|
||||
}
|
||||
|
||||
/** Changed keys that were accepted but whose effect waits for a restart. */
|
||||
private static List<String> changedDeferredKeys(BridgedConfig old, BridgedConfig fresh) {
|
||||
List<String> changed = new ArrayList<>();
|
||||
if (!Objects.equals(old.lifecycle(), fresh.lifecycle())) {
|
||||
changed.add("lifecycle");
|
||||
}
|
||||
if (!Objects.equals(old.leadHeartbeat(), fresh.leadHeartbeat())) {
|
||||
changed.add("leadHeartbeat");
|
||||
}
|
||||
if (!Objects.equals(old.guard(), fresh.guard())) {
|
||||
changed.add("guard");
|
||||
}
|
||||
if (!Objects.equals(old.worktreeRoot(), fresh.worktreeRoot())) {
|
||||
changed.add("worktreeRoot");
|
||||
}
|
||||
if (!Objects.equals(old.spawnReadyTimeoutMs(), fresh.spawnReadyTimeoutMs())
|
||||
|| !Objects.equals(old.spawnReadyPollMs(), fresh.spawnReadyPollMs())) {
|
||||
changed.add("spawnReady*");
|
||||
}
|
||||
Map<String, BridgedConfig.Profile> before =
|
||||
old.profiles() == null ? Map.of() : old.profiles();
|
||||
Map<String, BridgedConfig.Profile> after =
|
||||
fresh.profiles() == null ? Map.of() : fresh.profiles();
|
||||
// Adding or removing a profile is deferred: a new backend needs its own launcher, and
|
||||
// launchers are built once at startup.
|
||||
if (!before.keySet().equals(after.keySet())) {
|
||||
Set<String> diff = new LinkedHashSet<>(before.keySet());
|
||||
diff.addAll(after.keySet());
|
||||
diff.removeIf(p -> before.containsKey(p) && after.containsKey(p));
|
||||
changed.add("profiles (added/removed: " + String.join(", ", diff) + ")");
|
||||
}
|
||||
// An EXISTING profile's launch settings are deferred too, and this is easy to get wrong:
|
||||
// `HerdrPeerLauncher` takes `Map.copyOf(profiles)` at construction and `spawn` resolves the
|
||||
// profile out of that snapshot, so a reloaded model/baseUrl/argv/env never reaches a launch.
|
||||
// Only weight and maxLoad are genuinely hot, because placement reads them through the
|
||||
// supplier on the composite rather than from the adapter's copy. Without this check a
|
||||
// changed model would report "config reloaded" and silently do nothing — the worst outcome
|
||||
// a reload can produce, because the operator has no reason to doubt it.
|
||||
List<String> relaunch = new ArrayList<>();
|
||||
before.forEach((name, was) -> {
|
||||
BridgedConfig.Profile now = after.get(name);
|
||||
if (now != null && !sameLaunchSettings(was, now)) {
|
||||
relaunch.add(name);
|
||||
}
|
||||
});
|
||||
if (!relaunch.isEmpty()) {
|
||||
changed.add("profiles." + String.join("/", relaunch) + " launch settings "
|
||||
+ "(model, baseUrl, argv, env, …) — the launcher holds a startup snapshot");
|
||||
}
|
||||
return changed;
|
||||
}
|
||||
|
||||
/**
|
||||
* Whether two versions of a profile would launch a peer identically. Compares every component
|
||||
* the launcher reads at spawn; {@code weight} and {@code maxLoad} are excluded because those are
|
||||
* read live by the placement policy and really do take effect on the next spawn.
|
||||
*/
|
||||
private static boolean sameLaunchSettings(BridgedConfig.Profile a, BridgedConfig.Profile b) {
|
||||
return Objects.equals(a.baseUrl(), b.baseUrl())
|
||||
&& Objects.equals(a.model(), b.model())
|
||||
&& Objects.equals(a.configDir(), b.configDir())
|
||||
&& Objects.equals(a.tokenEnv(), b.tokenEnv())
|
||||
&& Objects.equals(a.argv(), b.argv())
|
||||
&& Objects.equals(a.placement(), b.placement())
|
||||
&& Objects.equals(a.workspace(), b.workspace())
|
||||
&& Objects.equals(a.tabLabel(), b.tabLabel())
|
||||
&& Objects.equals(a.mcpUrl(), b.mcpUrl())
|
||||
&& Objects.equals(a.cwd(), b.cwd())
|
||||
&& Objects.equals(a.parityOverlay(), b.parityOverlay())
|
||||
&& Objects.equals(a.gitTokenEnv(), b.gitTokenEnv())
|
||||
&& Objects.equals(a.gitHostEnv(), b.gitHostEnv())
|
||||
&& Objects.equals(a.kind(), b.kind())
|
||||
&& Objects.equals(a.env(), b.env())
|
||||
&& Objects.equals(a.subscription(), b.subscription());
|
||||
}
|
||||
}
|
||||
@@ -0,0 +1,96 @@
|
||||
package dev.ltms.bridged.config;
|
||||
|
||||
import org.slf4j.Logger;
|
||||
import org.slf4j.LoggerFactory;
|
||||
|
||||
import java.io.IOException;
|
||||
import java.nio.file.Files;
|
||||
import java.nio.file.Path;
|
||||
import java.util.concurrent.Executors;
|
||||
import java.util.concurrent.ScheduledExecutorService;
|
||||
import java.util.concurrent.TimeUnit;
|
||||
|
||||
/**
|
||||
* Polls {@code bridged.yaml}'s modified time and asks {@link ConfigRef} to reload when it moves
|
||||
* (CB-559). Opt-in through {@code configReload.enabled}.
|
||||
*
|
||||
* <p><strong>Why polling and not a filesystem watch.</strong> {@code WatchService} on macOS has no
|
||||
* native backend — it falls back to polling internally anyway, at an interval this code does not
|
||||
* control — and editors save config files in ways that produce a different event mix per editor
|
||||
* (write-in-place, write-and-rename, write-temp-and-swap). A modified-time check treats all of them
|
||||
* the same and is a single {@code stat} per tick, which at a ten-second cadence costs nothing worth
|
||||
* measuring.
|
||||
*
|
||||
* <p><strong>A missing or unreadable file is not a reason to act.</strong> Many editors briefly
|
||||
* unlink the file during a save. Reloading on "it vanished" would mean reloading from a file that no
|
||||
* longer exists; reporting an error every tick would bury the log. So an unreadable file is skipped
|
||||
* silently and the next tick tries again — the running config stays live, which is the correct
|
||||
* outcome either way.
|
||||
*/
|
||||
public final class ConfigWatcher {
|
||||
|
||||
private static final Logger log = LoggerFactory.getLogger(ConfigWatcher.class);
|
||||
|
||||
private final ConfigRef ref;
|
||||
private final long intervalSeconds;
|
||||
private final ScheduledExecutorService scheduler;
|
||||
|
||||
private volatile long lastSeenMillis;
|
||||
|
||||
public ConfigWatcher(ConfigRef ref, long intervalSeconds) {
|
||||
this.ref = ref;
|
||||
this.intervalSeconds = intervalSeconds;
|
||||
this.lastSeenMillis = modifiedMillis(ref.path());
|
||||
this.scheduler = Executors.newSingleThreadScheduledExecutor(r -> {
|
||||
Thread t = new Thread(r, "config-watcher");
|
||||
// A daemon thread: an operator's config watch must never be the reason the JVM refuses
|
||||
// to exit after everything else has shut down.
|
||||
t.setDaemon(true);
|
||||
return t;
|
||||
});
|
||||
}
|
||||
|
||||
/** Begin watching. A ref with no file (a fixed one) is a no-op rather than an error. */
|
||||
public void start() {
|
||||
if (ref.path() == null) {
|
||||
log.debug("config watch not started — this config has no file behind it");
|
||||
return;
|
||||
}
|
||||
scheduler.scheduleWithFixedDelay(this::tick, intervalSeconds, intervalSeconds,
|
||||
TimeUnit.SECONDS);
|
||||
log.info("config watch: {} re-read when it changes (every {}s)", ref.path(), intervalSeconds);
|
||||
}
|
||||
|
||||
/** One poll. Never throws — an exception here would silently cancel the schedule. */
|
||||
void tick() {
|
||||
try {
|
||||
long now = modifiedMillis(ref.path());
|
||||
if (now == 0 || now == lastSeenMillis) {
|
||||
return;
|
||||
}
|
||||
// Stamp BEFORE reloading. A file whose reload is refused (a bad edit, or a cold key)
|
||||
// must not be retried every tick — that would log the same refusal forever. The next
|
||||
// save moves the timestamp again and earns a fresh attempt.
|
||||
lastSeenMillis = now;
|
||||
ref.reload();
|
||||
} catch (RuntimeException e) {
|
||||
log.warn("config watch tick failed, still watching: {}", e.getMessage());
|
||||
}
|
||||
}
|
||||
|
||||
private static long modifiedMillis(Path path) {
|
||||
if (path == null) {
|
||||
return 0;
|
||||
}
|
||||
try {
|
||||
return Files.getLastModifiedTime(path).toMillis();
|
||||
} catch (IOException e) {
|
||||
return 0; // mid-save, or gone: say nothing and try again next tick
|
||||
}
|
||||
}
|
||||
|
||||
/** Stop polling. Called from the daemon's ordered shutdown hook, alongside the other loops. */
|
||||
public void stop() {
|
||||
scheduler.shutdownNow();
|
||||
}
|
||||
}
|
||||
@@ -0,0 +1,40 @@
|
||||
package dev.ltms.bridged.health;
|
||||
|
||||
import dev.ltms.bridged.herdr.AgentStatus;
|
||||
import dev.ltms.bridged.session.MemberSession;
|
||||
|
||||
/**
|
||||
* Pure classifier. Collection and repair are deliberately outside this package.
|
||||
* {@link HealthState#ERROR_ON_SCREEN} is not decided yet because it needs a bounded pane detection
|
||||
* read and an adapter-specific fatal signature; status facts alone must not guess it.
|
||||
*/
|
||||
public final class FleetHealth {
|
||||
private FleetHealth() { }
|
||||
|
||||
public static HealthDecision decide(HealthSnapshot s, HealthPrior prior, long nowNanos) {
|
||||
if (s.controlLinkDown()) return result(HealthState.CONTROL_LINK_DOWN, false);
|
||||
if (s.targetNotFound()) return result(HealthState.GONE, false);
|
||||
if (s.sessionState() == MemberSession.State.SPAWNING && !s.present() && s.readinessGraceElapsed()) {
|
||||
return result(HealthState.NEVER_READY, false);
|
||||
}
|
||||
if (s.orphanedDelegation()) return result(HealthState.DELEGATION_ORPHANED, false);
|
||||
boolean disagreement = s.sessionState() == MemberSession.State.BUSY && s.acceptedDelivery()
|
||||
&& (s.liveStatus() == AgentStatus.IDLE || s.liveStatus() == AgentStatus.DONE);
|
||||
if (disagreement && prior.busyButDone()) return result(HealthState.TURN_BOUNDARY_LOST, true);
|
||||
if (s.stalled()) return result(HealthState.STALL_SUSPECTED, disagreement);
|
||||
if (s.replyStranded()) return result(HealthState.REPLY_STRANDED, disagreement);
|
||||
if (s.queuedDelivery() || s.inboxMessage()) return result(HealthState.WORK_PENDING, disagreement);
|
||||
if (s.sessionState() == MemberSession.State.SPAWNING) return result(HealthState.STARTING, disagreement);
|
||||
if (s.acceptedDelivery() && s.liveStatus() == AgentStatus.BLOCKED) {
|
||||
return result(HealthState.BLOCKED_AMBIGUOUS, disagreement);
|
||||
}
|
||||
// An accepted delivery remains bridge work even when herdr is late, unknown, or has already
|
||||
// reported DONE once. It cannot be IDLE until the delegation has resolved.
|
||||
if (s.acceptedDelivery()) return result(HealthState.WORKING, disagreement);
|
||||
return result(HealthState.IDLE, disagreement);
|
||||
}
|
||||
|
||||
private static HealthDecision result(HealthState state, boolean disagreement) {
|
||||
return new HealthDecision(state, new HealthPrior(disagreement));
|
||||
}
|
||||
}
|
||||
@@ -0,0 +1,119 @@
|
||||
package dev.ltms.bridged.health;
|
||||
|
||||
import dev.ltms.bridged.herdr.Agent;
|
||||
import dev.ltms.bridged.herdr.AgentControl;
|
||||
import dev.ltms.bridged.herdr.AgentStatus;
|
||||
import dev.ltms.bridged.msg.MessageService;
|
||||
import dev.ltms.bridged.session.MemberSession;
|
||||
import org.slf4j.Logger;
|
||||
import org.slf4j.LoggerFactory;
|
||||
|
||||
import java.util.HashMap;
|
||||
import java.util.HashSet;
|
||||
import java.util.List;
|
||||
import java.util.Map;
|
||||
import java.util.concurrent.ScheduledExecutorService;
|
||||
import java.util.concurrent.TimeUnit;
|
||||
import java.util.function.LongSupplier;
|
||||
import java.util.function.Supplier;
|
||||
import java.util.function.BiConsumer;
|
||||
|
||||
/** Slow whole-fleet evidence collection. It is deliberately separate from the delivery poller. */
|
||||
public final class FleetHealthMonitor {
|
||||
private static final Logger log = LoggerFactory.getLogger(FleetHealthMonitor.class);
|
||||
private final AgentControl agents;
|
||||
private final Supplier<List<MemberSession>> roster;
|
||||
private final MessageService messages;
|
||||
private final ScheduledExecutorService scheduler;
|
||||
private final LongSupplier clock;
|
||||
private final long intervalSeconds;
|
||||
private final BiConsumer<String, String> failTarget;
|
||||
private final Map<String, HealthPrior> priors = new HashMap<>();
|
||||
private final Map<String, HealthState> states = new HashMap<>();
|
||||
|
||||
// These facts need the evidence publishers introduced by later M4 units. They are not negatives.
|
||||
private static final boolean NOT_YET_OBSERVED = false;
|
||||
|
||||
public FleetHealthMonitor(AgentControl agents, Supplier<List<MemberSession>> roster, MessageService messages,
|
||||
ScheduledExecutorService scheduler, LongSupplier clock, long intervalSeconds) {
|
||||
this(agents, roster, messages, scheduler, clock, intervalSeconds, (_, _) -> { });
|
||||
}
|
||||
|
||||
public FleetHealthMonitor(AgentControl agents, Supplier<List<MemberSession>> roster, MessageService messages,
|
||||
ScheduledExecutorService scheduler, LongSupplier clock, long intervalSeconds,
|
||||
BiConsumer<String, String> failTarget) {
|
||||
this.agents = agents;
|
||||
this.roster = roster;
|
||||
this.messages = messages;
|
||||
this.scheduler = scheduler;
|
||||
this.clock = clock;
|
||||
this.intervalSeconds = intervalSeconds;
|
||||
this.failTarget = failTarget;
|
||||
}
|
||||
|
||||
/** Pure per-member decision seam. */
|
||||
static HealthDecision decide(HealthSnapshot snapshot, HealthPrior prior, long nowNanos) {
|
||||
return FleetHealth.decide(snapshot, prior, nowNanos);
|
||||
}
|
||||
|
||||
public void start() { scheduler.schedule(this::tick, intervalSeconds, TimeUnit.SECONDS); }
|
||||
public void stop() { scheduler.shutdownNow(); }
|
||||
|
||||
// Package-private so tests can run one tick without waiting.
|
||||
void tick() {
|
||||
try {
|
||||
List<Agent> agentsNow = agents.list(); // Exactly one list call for this complete observation.
|
||||
List<MemberSession> rosterNow = roster.get(); // One in-memory roster snapshot for this tick.
|
||||
Map<String, Agent> live = new HashMap<>();
|
||||
for (Agent agent : agentsNow) live.put(agent.terminalId(), agent);
|
||||
HashSet<String> current = new HashSet<>();
|
||||
for (MemberSession session : rosterNow) {
|
||||
current.add(session.terminalId());
|
||||
Agent agent = live.get(session.terminalId());
|
||||
AgentStatus status = agent == null ? AgentStatus.UNKNOWN : agent.status();
|
||||
boolean accepted = messages.hasAcceptedDelivery(session.terminalId());
|
||||
HealthSnapshot snapshot = new HealthSnapshot(session.state(), status, accepted, NOT_YET_OBSERVED,
|
||||
messages.hasInboxMessage(session.terminalId()), agent != null, NOT_YET_OBSERVED,
|
||||
NOT_YET_OBSERVED, NOT_YET_OBSERVED, NOT_YET_OBSERVED, NOT_YET_OBSERVED, NOT_YET_OBSERVED);
|
||||
HealthDecision decision = decide(snapshot, priors.getOrDefault(session.terminalId(), HealthPrior.NONE),
|
||||
clock.getAsLong());
|
||||
priors.put(session.terminalId(), decision.prior());
|
||||
if (decision.state() == HealthState.GONE || decision.state() == HealthState.NEVER_READY) {
|
||||
failTarget.accept(session.terminalId(), "fleet health detected " + decision.state());
|
||||
}
|
||||
reportTransition(session.terminalId(), decision.state());
|
||||
}
|
||||
priors.keySet().retainAll(current);
|
||||
states.keySet().retainAll(current);
|
||||
} catch (Throwable error) {
|
||||
// A list failure is health evidence, and must never kill the monitor's only scheduler task.
|
||||
log.warn("fleet health collection failed; will retry next tick", error);
|
||||
} finally {
|
||||
if (!scheduler.isShutdown()) {
|
||||
scheduler.schedule(this::tick, intervalSeconds, TimeUnit.SECONDS);
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
void reportTransition(String target, HealthState next) {
|
||||
HealthState previous = states.put(target, next);
|
||||
if (previous == next) return;
|
||||
if (fault(next)) {
|
||||
log.warn("fleet health member={} state={} previous={}", target, next, previous);
|
||||
} else if (previous != null && fault(previous)) {
|
||||
log.info("fleet health member={} recovered state={} previous={}", target, next, previous);
|
||||
}
|
||||
}
|
||||
|
||||
private static boolean fault(HealthState state) {
|
||||
return switch (state) {
|
||||
case NEVER_READY, GONE, TURN_BOUNDARY_LOST, ERROR_ON_SCREEN, STALL_SUSPECTED,
|
||||
MUTE, REPLY_STRANDED, DELEGATION_ORPHANED, CONTROL_LINK_DOWN -> true;
|
||||
default -> false;
|
||||
};
|
||||
}
|
||||
|
||||
public static String coverage(boolean enabled, boolean notificationConfigured) {
|
||||
return !enabled ? "off" : notificationConfigured ? "full" : "detection-only";
|
||||
}
|
||||
}
|
||||
@@ -0,0 +1,4 @@
|
||||
package dev.ltms.bridged.health;
|
||||
|
||||
/** Classification plus the private fact that the next pure decision needs. */
|
||||
public record HealthDecision(HealthState state, HealthPrior prior) { }
|
||||
@@ -0,0 +1,6 @@
|
||||
package dev.ltms.bridged.health;
|
||||
|
||||
/** Private cross-tick observation. It is deliberately not a reported health value. */
|
||||
public record HealthPrior(boolean busyButDone) {
|
||||
public static final HealthPrior NONE = new HealthPrior(false);
|
||||
}
|
||||
@@ -0,0 +1,11 @@
|
||||
package dev.ltms.bridged.health;
|
||||
|
||||
import dev.ltms.bridged.herdr.AgentStatus;
|
||||
import dev.ltms.bridged.session.MemberSession;
|
||||
|
||||
/** Read-only facts from one fleet collection tick. */
|
||||
public record HealthSnapshot(MemberSession.State sessionState, AgentStatus liveStatus,
|
||||
boolean acceptedDelivery, boolean queuedDelivery, boolean inboxMessage,
|
||||
boolean present, boolean targetNotFound, boolean controlLinkDown,
|
||||
boolean readinessGraceElapsed, boolean orphanedDelegation,
|
||||
boolean replyStranded, boolean stalled) { }
|
||||
@@ -0,0 +1,8 @@
|
||||
package dev.ltms.bridged.health;
|
||||
|
||||
/** Health classifications reported for a member. */
|
||||
public enum HealthState {
|
||||
STARTING, IDLE, WORKING, WORK_PENDING, BLOCKED_AMBIGUOUS,
|
||||
NEVER_READY, GONE, TURN_BOUNDARY_LOST, ERROR_ON_SCREEN, STALL_SUSPECTED,
|
||||
MUTE, REPLY_STRANDED, DELEGATION_ORPHANED, CONTROL_LINK_DOWN
|
||||
}
|
||||
@@ -0,0 +1,25 @@
|
||||
package dev.ltms.bridged.health;
|
||||
|
||||
import dev.ltms.bridged.msg.Rendezvous;
|
||||
|
||||
import java.util.Map;
|
||||
import java.util.concurrent.ConcurrentHashMap;
|
||||
|
||||
/**
|
||||
* Counts turns that ended via the completion fallback instead of {@code bridge_reply}.
|
||||
* MUTE is an observation by target and profile, not a classifier state and never suppresses faults.
|
||||
*/
|
||||
public final class MuteCounter {
|
||||
private final Map<String, Integer> byTarget = new ConcurrentHashMap<>();
|
||||
private final Map<String, Integer> byProfile = new ConcurrentHashMap<>();
|
||||
|
||||
/** Record only fallback completion; a structured reply does not make a member mute. */
|
||||
public void observe(String target, String profile, Rendezvous.Kind kind) {
|
||||
if (kind != Rendezvous.Kind.COMPLETION) return;
|
||||
byTarget.merge(target, 1, Integer::sum);
|
||||
byProfile.merge(profile, 1, Integer::sum);
|
||||
}
|
||||
|
||||
public int forTarget(String target) { return byTarget.getOrDefault(target, 0); }
|
||||
public int forProfile(String profile) { return byProfile.getOrDefault(profile, 0); }
|
||||
}
|
||||
@@ -0,0 +1,26 @@
|
||||
package dev.ltms.bridged.health;
|
||||
|
||||
import java.util.ArrayList;
|
||||
import java.util.HashMap;
|
||||
import java.util.List;
|
||||
import java.util.Map;
|
||||
|
||||
/** Fixed pane-probe limits. Pane content is never retained here. */
|
||||
public final class PaneBudget {
|
||||
public static final long COOLDOWN_NANOS = 60_000_000_000L;
|
||||
public static final int MAX_PER_TICK = 2;
|
||||
private final Map<String, Long> lastProbe = new HashMap<>();
|
||||
private int cursor;
|
||||
|
||||
public List<String> choose(List<String> candidates, long nowNanos, long configuredCooldownNanos) {
|
||||
long cooldown = Math.max(COOLDOWN_NANOS, configuredCooldownNanos);
|
||||
List<String> out = new ArrayList<>();
|
||||
for (int n = 0; n < candidates.size() && out.size() < MAX_PER_TICK; n++) {
|
||||
String target = candidates.get((cursor + n) % candidates.size());
|
||||
Long last = lastProbe.get(target);
|
||||
if (last == null || nowNanos - last >= cooldown) { out.add(target); lastProbe.put(target, nowNanos); }
|
||||
}
|
||||
if (!candidates.isEmpty()) cursor = (cursor + 1) % candidates.size();
|
||||
return List.copyOf(out);
|
||||
}
|
||||
}
|
||||
@@ -26,15 +26,26 @@ import java.util.function.Supplier;
|
||||
* <p><strong>Direction of trust.</strong> The label names the lead; it never <em>grants</em>
|
||||
* anything a pane could take for itself. Three properties keep that honest:
|
||||
* <ol>
|
||||
* <li>bridged never renames a lead tab. The operator's label is read-only input, so what is in
|
||||
* the tab bar is always what the human wrote — no round-trip where the daemon's own rename
|
||||
* becomes the evidence for its next decision.</li>
|
||||
* <li>Worker spaces are excluded wholesale ({@code excludedWorkspaceLabels}), so a worker cannot
|
||||
* become a lead by being placed — as a split, say — inside a matching tab.</li>
|
||||
* <li>A worker cannot rename a tab: {@code tab.rename} is reachable only through
|
||||
* {@link WorkspaceControl}, which no {@code bridge_*} tool exposes. The label is writable by
|
||||
* the human at the terminal and by nobody the bridge is defending against.</li>
|
||||
* <li>The label is a <em>name</em>, not a capability. What a pane may do is decided by
|
||||
* {@code Authz} against the role {@code CallerResolver} returns; a tab that calls itself a
|
||||
* lead still cannot act as one unless the daemon's own registry agrees.</li>
|
||||
* </ol>
|
||||
*
|
||||
* <p><strong>CB-558 — bridged now writes lead labels too.</strong> This class used to be able to say
|
||||
* that bridged never renames a lead tab, so the label was always the human's own writing and there
|
||||
* was no round-trip from the daemon's rename back into its next decision.
|
||||
* {@code dev.ltms.bridged.lead.LeadLauncher} ends that: an auto-launched lead is labelled by the
|
||||
* daemon and found again by this scan. The trust direction above is unaffected — bridged writing a
|
||||
* name for a lead it just started is not a pane promoting itself — but <em>staleness</em> becomes
|
||||
* real: a label left behind by a session that has since died would read as a live lead forever.
|
||||
* This scanner does not solve that (its job is naming, and a stale name costs nothing here); the
|
||||
* launcher does, by requiring a running agent in the tab before it counts the lead as live. If you
|
||||
* ever make a decision that <em>removes</em> something based on this map, add the same check.
|
||||
* The remaining hazard is an <em>operator</em> one — a worker {@code tabLabel} template that
|
||||
* happens to start with the same prefix would promote the whole fleet — and that is refused at
|
||||
* startup by {@code BridgedConfig.validateLeadScan} rather than documented here.
|
||||
|
||||
@@ -41,6 +41,16 @@ public final class WorkspaceControl {
|
||||
return out;
|
||||
}
|
||||
|
||||
/** Every tab in {@code workspaceId}, in herdr's order. */
|
||||
public List<Tab> listTabs(String workspaceId) {
|
||||
JsonNode result = herdr.call("tab.list", Map.of("workspace_id", workspaceId));
|
||||
List<Tab> out = new ArrayList<>();
|
||||
for (JsonNode t : result.path("tabs")) {
|
||||
out.add(Tab.from(t));
|
||||
}
|
||||
return out;
|
||||
}
|
||||
|
||||
/** The first workspace with this exact label, if any. */
|
||||
public Optional<Workspace> findByLabel(String label) {
|
||||
return listWorkspaces().stream()
|
||||
|
||||
@@ -55,6 +55,9 @@ public final class CompletionResolver implements TurnListener {
|
||||
/** Cap the scraped tail so a long transcript can't return an unbounded blob. */
|
||||
static final int MAX_SCRAPE_CHARS = 4000;
|
||||
|
||||
private static final String CLIPPED_PANE_TAIL_MARKER =
|
||||
"[Pane tail clipped: member did not call bridge_reply.]";
|
||||
|
||||
private final AgentControl agents;
|
||||
private final Rendezvous rendezvous;
|
||||
|
||||
@@ -134,7 +137,13 @@ public final class CompletionResolver implements TurnListener {
|
||||
@Override
|
||||
public void onTurnFailed(String target) {
|
||||
InFlight turn = inFlight.get(target);
|
||||
Thread.ofVirtual().name("turn-failed-" + target).start(() -> fail(target, turn));
|
||||
Thread.ofVirtual().name("turn-failed-" + target).start(() -> fail(target, turn, null));
|
||||
}
|
||||
|
||||
@Override
|
||||
public void onTurnFailed(String target, String reason) {
|
||||
InFlight turn = inFlight.get(target);
|
||||
Thread.ofVirtual().name("turn-failed-" + target).start(() -> fail(target, turn, reason));
|
||||
}
|
||||
|
||||
/** Synchronous resolve (the unit-testable core of {@link #onTurnComplete}). */
|
||||
@@ -147,9 +156,14 @@ public final class CompletionResolver implements TurnListener {
|
||||
return;
|
||||
}
|
||||
String tail;
|
||||
int originalLength = 0;
|
||||
boolean clipped = false;
|
||||
boolean scrapeFailed = false;
|
||||
try {
|
||||
tail = clip(lastAssistantBlock(agents.read(target, SCRAPE_SOURCE)));
|
||||
String assistantBlock = lastAssistantBlock(agents.read(target, SCRAPE_SOURCE));
|
||||
originalLength = assistantBlock.strip().length();
|
||||
clipped = originalLength > MAX_SCRAPE_CHARS;
|
||||
tail = clip(assistantBlock);
|
||||
} catch (RuntimeException e) {
|
||||
// The worker finished but we couldn't read its screen — still resolve the send so the
|
||||
// caller unblocks; an empty tail beats hanging until the caller's timeout.
|
||||
@@ -169,8 +183,14 @@ public final class CompletionResolver implements TurnListener {
|
||||
target);
|
||||
return; // keep the in-flight record: a later genuine completion still needs it
|
||||
}
|
||||
if (rendezvous.resolveCompletion(waiter, tail)) {
|
||||
String completion = clipped ? tail + "\n" + CLIPPED_PANE_TAIL_MARKER : tail;
|
||||
if (rendezvous.resolveCompletion(waiter, completion)) {
|
||||
inFlight.remove(target, turn);
|
||||
if (clipped) {
|
||||
log.warn("completion scrape for {} clipped from {} chars to the {} char cap; "
|
||||
+ "member did not call bridge_reply, so the pane tail is partial",
|
||||
target, originalLength, MAX_SCRAPE_CHARS);
|
||||
}
|
||||
log.debug("resolved send to {} via turn-completion fallback ({} chars scraped)",
|
||||
target, tail.length());
|
||||
}
|
||||
@@ -178,6 +198,11 @@ public final class CompletionResolver implements TurnListener {
|
||||
|
||||
/** Synchronous fail (the unit-testable core of {@link #onTurnFailed}). */
|
||||
void fail(String target, InFlight turn) {
|
||||
fail(target, turn, null);
|
||||
}
|
||||
|
||||
/** Synchronous fail with an optional reason supplied by a dropped worker queue. */
|
||||
void fail(String target, InFlight turn, String explicitReason) {
|
||||
// A never-delivered readiness failure has no in-flight record but still has a blocked send;
|
||||
// fall back to the currently-registered waiter (unambiguous — that send never completed, so
|
||||
// no next turn exists to confuse it with).
|
||||
@@ -187,20 +212,22 @@ public final class CompletionResolver implements TurnListener {
|
||||
inFlight.remove(target, turn); // nobody blocked on this worker — nothing to fail
|
||||
return;
|
||||
}
|
||||
String reason;
|
||||
try {
|
||||
reason = clip(agents.read(target, SCRAPE_SOURCE));
|
||||
} catch (RuntimeException e) {
|
||||
reason = "";
|
||||
}
|
||||
if (reason.isBlank()) {
|
||||
// No screen to scrape — either the worker is stuck (CB-109) or gone (CB-110).
|
||||
reason = "worker did not reply; its turn ended in an unrecoverable state "
|
||||
+ "(worker unreachable or stuck)";
|
||||
String reason = explicitReason;
|
||||
if (reason == null || reason.isBlank()) {
|
||||
try {
|
||||
reason = clip(agents.read(target, SCRAPE_SOURCE));
|
||||
} catch (RuntimeException e) {
|
||||
reason = "";
|
||||
}
|
||||
if (reason.isBlank()) {
|
||||
// No screen to scrape — either the worker is stuck (CB-109) or gone (CB-110).
|
||||
reason = "worker did not reply; its turn ended in an unrecoverable state "
|
||||
+ "(worker unreachable or stuck)";
|
||||
}
|
||||
}
|
||||
if (rendezvous.resolveFailure(waiter, reason)) {
|
||||
inFlight.remove(target, turn);
|
||||
log.debug("failed send to {} via turn-stall fallback", target);
|
||||
log.warn("failing send to {} via turn-stall fallback: {}", target, reason);
|
||||
}
|
||||
}
|
||||
|
||||
|
||||
@@ -78,6 +78,15 @@ public final class Injector {
|
||||
*/
|
||||
private static final int READINESS_GRACE_POLLS = 240;
|
||||
|
||||
/**
|
||||
* The single source for the injector poll cadence — how often the {@link StatusPoller} drives
|
||||
* {@link #onStatus} at. {@code Bridged} passes this to every {@link StatusPoller} it constructs,
|
||||
* and this class reads it to state the readiness grace in seconds on the CB-562 expiry log
|
||||
* instead of hardcoding "60s". One constant, so a cadence change cannot silently desync a log
|
||||
* that claims a grace duration.
|
||||
*/
|
||||
public static final long POLL_INTERVAL_MILLIS = 250;
|
||||
|
||||
private final AgentControl agents;
|
||||
private final TurnListener turnListener;
|
||||
private final Predicate<String> ready; // CB-113: a target is deliverable only when available
|
||||
@@ -264,6 +273,11 @@ public final class Injector {
|
||||
// fail every queued message and release the target (CB-114) instead of
|
||||
// polling it indefinitely with the caller's future never completing.
|
||||
notReady = new ArrayList<>(t.queue);
|
||||
log.warn("readiness grace for {} expired after {} polls ({}s): target never "
|
||||
+ "became deliverable, so failing {} queued message(s) that never "
|
||||
+ "reached its pane",
|
||||
target, READINESS_GRACE_POLLS,
|
||||
READINESS_GRACE_POLLS * POLL_INTERVAL_MILLIS / 1000, notReady.size());
|
||||
t.queue.clear();
|
||||
t.notReadySincePoll = 0;
|
||||
}
|
||||
@@ -388,12 +402,16 @@ public final class Injector {
|
||||
t.awaitingPostTurnPickup = false;
|
||||
t.postTurnObserved = false;
|
||||
}
|
||||
log.warn("{} is gone, dropping its queue: {} message(s) failed{}; cause: {}", target,
|
||||
pending.size(),
|
||||
hadDeliveredTurn ? " (including one turn already in flight whose completion was never confirmed)" : "",
|
||||
cause.getMessage());
|
||||
forget.accept(target); // the worker is gone — clear its readiness/presence too (CB-114)
|
||||
for (Pending p : pending) {
|
||||
p.delivered().completeExceptionally(cause);
|
||||
}
|
||||
if (hadDeliveredTurn) {
|
||||
turnListener.onTurnFailed(target);
|
||||
}
|
||||
// A queued send has no in-flight record, while a delivered turn does. CompletionResolver
|
||||
// handles both forms and resolves its waiter at most once.
|
||||
turnListener.onTurnFailed(target, cause.getMessage());
|
||||
}
|
||||
}
|
||||
|
||||
+1
-1
@@ -15,7 +15,7 @@ import java.util.Set;
|
||||
* mounts the bridge MCP is never marked present — its sends stay queued until they time out, which is
|
||||
* correct (it could not have replied anyway).
|
||||
*/
|
||||
public class WorkerPresence {
|
||||
public class MemberPresence {
|
||||
|
||||
private final Set<String> present = ConcurrentHashMap.newKeySet();
|
||||
|
||||
@@ -41,6 +41,14 @@ public interface TurnListener {
|
||||
default void onTurnFailed(String target) {
|
||||
}
|
||||
|
||||
/**
|
||||
* As {@link #onTurnFailed(String)}, carrying the reason a worker became unreachable. The default
|
||||
* keeps existing listeners working while allowing the completion resolver to report a useful cause.
|
||||
*/
|
||||
default void onTurnFailed(String target, String reason) {
|
||||
onTurnFailed(target);
|
||||
}
|
||||
|
||||
/**
|
||||
* A message was just delivered into {@code target}'s pane (CB-115). Fired so the completion
|
||||
* resolver can snapshot the pane's pre-turn content: a later {@link #onTurnComplete} whose
|
||||
|
||||
@@ -0,0 +1,304 @@
|
||||
package dev.ltms.bridged.lead;
|
||||
|
||||
import dev.ltms.bridged.config.BridgedConfig;
|
||||
import dev.ltms.bridged.herdr.Agent;
|
||||
import dev.ltms.bridged.herdr.AgentControl;
|
||||
import dev.ltms.bridged.herdr.HerdrException;
|
||||
import dev.ltms.bridged.herdr.Tab;
|
||||
import dev.ltms.bridged.herdr.Workspace;
|
||||
import dev.ltms.bridged.herdr.WorkspaceControl;
|
||||
import org.slf4j.Logger;
|
||||
import org.slf4j.LoggerFactory;
|
||||
|
||||
import java.util.ArrayList;
|
||||
import java.util.LinkedHashMap;
|
||||
import java.util.List;
|
||||
import java.util.Map;
|
||||
import java.util.Objects;
|
||||
import java.util.Set;
|
||||
import java.util.stream.Collectors;
|
||||
|
||||
/**
|
||||
* Starts the leads {@code fleet.leaders:} declares, when none is already running (CB-558).
|
||||
*
|
||||
* <p><strong>A lead is not a member, and this class exists to keep it that way.</strong> Every other
|
||||
* spawn path in the daemon goes through {@code HerdrPeerLauncher}, which does three things a lead
|
||||
* must never receive:
|
||||
* <ol>
|
||||
* <li>it appends the <em>reply charter</em> — "you are an off-subscription worker … end every turn
|
||||
* with {@code bridge_reply}". A lead is the orchestrator; telling it that it is a worker is
|
||||
* exactly backwards.</li>
|
||||
* <li>it registers the session with {@code SessionManager}, which subjects it to the idle reaper,
|
||||
* the context cap and the shutdown drain. An idle lead is the normal state of a lead, so the
|
||||
* reaper would kill the orchestrator for doing its job.</li>
|
||||
* <li>it can move a peer off the subscription via {@code ANTHROPIC_BASE_URL}. A lead stays on the
|
||||
* operator's subscription, always.</li>
|
||||
* </ol>
|
||||
* So this launcher talks to {@link AgentControl}/{@link WorkspaceControl} directly. The duplication
|
||||
* with the member launchers (argv and env assembly) is deliberate and is the cheaper half of the
|
||||
* trade: entangling the worker path with a not-a-worker case is how the three rules above get
|
||||
* broken later, quietly.
|
||||
*
|
||||
* <p><strong>Liveness, and the label round-trip.</strong> {@code LeadTabScanner} used to be able to
|
||||
* promise that bridged never writes a lead label. That is no longer true — an auto-launched lead is
|
||||
* labelled by this class, and the scanner reads that label back. The risk this opens is not
|
||||
* privilege escalation (the tab label never granted anything a pane could take for itself; see that
|
||||
* class's javadoc), but <em>staleness</em>: a label left behind by a crashed session would otherwise
|
||||
* read as a live lead forever, and the lead would never be relaunched. So a lead counts as live only
|
||||
* when herdr also reports a <em>running agent</em> in that tab — see {@link #liveLeads}. A labelled
|
||||
* tab with no agent in it is not a lead.
|
||||
*/
|
||||
public final class LeadLauncher {
|
||||
|
||||
private static final Logger log = LoggerFactory.getLogger(LeadLauncher.class);
|
||||
|
||||
private final AgentControl agents;
|
||||
private final WorkspaceControl spaces;
|
||||
private final BridgedConfig cfg;
|
||||
|
||||
/**
|
||||
* @param agents herdr agent control (start, list)
|
||||
* @param spaces workspace / tab control (ensure, create, label, list)
|
||||
* @param cfg the loaded config — {@code fleet.leaders}, {@code profiles} and the lead pins
|
||||
*/
|
||||
public LeadLauncher(AgentControl agents, WorkspaceControl spaces, BridgedConfig cfg) {
|
||||
this.agents = agents;
|
||||
this.spaces = spaces;
|
||||
this.cfg = cfg;
|
||||
}
|
||||
|
||||
/**
|
||||
* Bring every declared lead up to its {@code instances} count, and return how many were started.
|
||||
*
|
||||
* <p>Never throws: a daemon that cannot start a lead must still serve. herdr being unreachable,
|
||||
* a profile that does not exist, a failed {@code agent.start} — each is logged and skipped, and
|
||||
* the remaining leads are still attempted.
|
||||
*/
|
||||
public int ensureLeads() {
|
||||
Map<String, BridgedConfig.Leader> leaders = cfg.fleet().leaders();
|
||||
if (leaders.isEmpty()) {
|
||||
return 0;
|
||||
}
|
||||
|
||||
Map<String, Integer> live;
|
||||
try {
|
||||
live = liveLeads(leaders);
|
||||
} catch (HerdrException e) {
|
||||
// Counting is the whole safety mechanism against double-spawning. If we cannot count, we
|
||||
// must not guess — spawning a second orchestrator is worse than starting none.
|
||||
log.warn("lead auto-launch skipped — cannot tell which leads are live: {}", e.getMessage());
|
||||
return 0;
|
||||
}
|
||||
|
||||
int started = 0;
|
||||
for (Map.Entry<String, BridgedConfig.Leader> e : leaders.entrySet()) {
|
||||
String name = e.getKey();
|
||||
BridgedConfig.Leader lead = e.getValue();
|
||||
int running = live.getOrDefault(name, 0);
|
||||
int wanted = lead.instances();
|
||||
|
||||
if (running >= wanted) {
|
||||
log.info("lead '{}': {} live, {} wanted — nothing to start", name, running, wanted);
|
||||
continue;
|
||||
}
|
||||
if (!lead.isCreatable()) {
|
||||
// A lead with a `terminal:` pin and no `profile:` is recognise-only by design: the
|
||||
// operator opens it by hand. Say so once rather than looking like a silent failure.
|
||||
log.info("lead '{}' is not live, and names no profile — it can be recognised but not "
|
||||
+ "launched. Add `profile:` under fleet.leaders.{} to have bridged start it.",
|
||||
name, name);
|
||||
continue;
|
||||
}
|
||||
|
||||
BridgedConfig.Profile profile = cfg.profiles().get(lead.profile());
|
||||
if (profile == null) {
|
||||
log.warn("lead '{}' names profile '{}', which is not configured — not launching",
|
||||
name, lead.profile());
|
||||
continue;
|
||||
}
|
||||
|
||||
for (int i = running; i < wanted; i++) {
|
||||
if (launch(name, lead, profile)) {
|
||||
started++;
|
||||
}
|
||||
}
|
||||
}
|
||||
return started;
|
||||
}
|
||||
|
||||
/**
|
||||
* How many live leads exist per configured name.
|
||||
*
|
||||
* <p>Two independent pieces of evidence, because either alone double-spawns:
|
||||
* <ul>
|
||||
* <li>a running agent in a tab labelled {@code "<tabPrefix> <name>"} — how an auto-launched
|
||||
* lead, or an operator following the labelling convention, is found;</li>
|
||||
* <li>a running agent on a terminal the config pins in {@code fleet.leaders.<name>.terminal} —
|
||||
* how a lead the operator opened and pinned by hand is found. Without this, a pinned lead
|
||||
* whose tab carries no matching label would be relaunched on every boot.</li>
|
||||
* </ul>
|
||||
* Member workspaces are excluded, exactly as the scanner excludes them: a member must not be
|
||||
* counted as a lead because it happens to sit in a matching tab.
|
||||
*/
|
||||
private Map<String, Integer> liveLeads(Map<String, BridgedConfig.Leader> leaders) {
|
||||
Set<String> memberSpaces = cfg.profiles().values().stream()
|
||||
.map(BridgedConfig.Profile::workspace)
|
||||
.filter(w -> w != null && !w.isBlank())
|
||||
.collect(Collectors.toSet());
|
||||
|
||||
// tabId → the lead name its label declares.
|
||||
Map<String, String> nameByTab = new LinkedHashMap<>();
|
||||
for (Workspace ws : spaces.listWorkspaces()) {
|
||||
if (ws.workspaceId() == null || memberSpaces.contains(ws.label())) {
|
||||
continue;
|
||||
}
|
||||
for (Tab tab : spaces.listTabs(ws.workspaceId())) {
|
||||
String declared = leadNameOf(tab.label(), leaders);
|
||||
if (declared != null && tab.tabId() != null) {
|
||||
nameByTab.put(tab.tabId(), declared);
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
// terminalId → the lead name the config pins it to.
|
||||
Map<String, String> nameByPinnedTerminal = new LinkedHashMap<>();
|
||||
leaders.forEach((name, lead) -> {
|
||||
if (lead.terminal() != null && !lead.terminal().isBlank()) {
|
||||
nameByPinnedTerminal.put(lead.terminal().strip(), name);
|
||||
}
|
||||
});
|
||||
|
||||
Map<String, Integer> counts = new LinkedHashMap<>();
|
||||
for (Agent a : agents.list()) {
|
||||
String name = nameByTab.get(a.tabId());
|
||||
if (name == null) {
|
||||
name = nameByPinnedTerminal.get(a.terminalId());
|
||||
}
|
||||
if (name != null) {
|
||||
counts.merge(name, 1, Integer::sum);
|
||||
}
|
||||
}
|
||||
return counts;
|
||||
}
|
||||
|
||||
/**
|
||||
* The configured lead a tab label names, or {@code null} for a label that names none.
|
||||
*
|
||||
* <p>Matched against the declared lead names rather than by splitting on the prefix, so an
|
||||
* operator's {@code "lead: something-else"} tab is not mistaken for a configured lead.
|
||||
*/
|
||||
private String leadNameOf(String label, Map<String, BridgedConfig.Leader> leaders) {
|
||||
if (label == null) {
|
||||
return null;
|
||||
}
|
||||
String l = label.strip();
|
||||
for (Map.Entry<String, BridgedConfig.Leader> e : leaders.entrySet()) {
|
||||
if (l.equalsIgnoreCase(e.getValue().tabLabel(e.getKey()).strip())) {
|
||||
return e.getKey();
|
||||
}
|
||||
}
|
||||
return null;
|
||||
}
|
||||
|
||||
/** Start one lead. Returns false (having logged) rather than throwing on any failure. */
|
||||
private boolean launch(String name, BridgedConfig.Leader lead, BridgedConfig.Profile profile) {
|
||||
String label = lead.tabLabel(name);
|
||||
String cwd = (lead.cwd() == null || lead.cwd().isBlank())
|
||||
? System.getProperty("user.dir") : lead.cwd();
|
||||
|
||||
Tab.Created tab = null;
|
||||
try {
|
||||
Workspace ws = spaces.ensureWorkspace(lead.workspace());
|
||||
tab = spaces.createTab(ws.workspaceId(), cwd, leadEnv(profile));
|
||||
if (tab.rootPaneId() == null) {
|
||||
throw new IllegalStateException("tab " + tab.tab().tabId()
|
||||
+ " came back with no seed pane — nowhere to start the lead");
|
||||
}
|
||||
// Same shape as the member launchers: herdr resolves the executable from `kind`, so
|
||||
// argv[0] (the configured launcher, e.g. `ccs`) is dropped and only the rest is passed.
|
||||
List<String> argv = leadArgv(profile);
|
||||
Agent started = agents.start("lead-" + name, herdrKind(profile),
|
||||
argv.isEmpty() ? argv : argv.subList(1, argv.size()), tab.rootPaneId());
|
||||
|
||||
// Label AFTER the start succeeds. A label written before would survive a failed start
|
||||
// and then read back as a live lead on the next boot, which is the exact staleness the
|
||||
// agent-liveness check exists to prevent — no need to create the case ourselves.
|
||||
spaces.renameTab(tab.tab().tabId(), label);
|
||||
|
||||
log.info("lead '{}' launched: profile={} tab={} pane={} terminal={} label='{}' cwd={}",
|
||||
name, profile.profile(), tab.tab().tabId(), started.paneId(),
|
||||
started.terminalId(), label, cwd);
|
||||
return true;
|
||||
} catch (RuntimeException e) {
|
||||
log.warn("lead '{}' failed to launch on profile '{}': {}",
|
||||
name, profile.profile(), e.getMessage());
|
||||
if (tab != null && tab.tab() != null && tab.tab().tabId() != null) {
|
||||
try {
|
||||
spaces.closeTab(tab.tab().tabId());
|
||||
} catch (RuntimeException cleanup) {
|
||||
log.warn("could not close the orphaned lead tab {}: {}",
|
||||
tab.tab().tabId(), cleanup.getMessage());
|
||||
}
|
||||
}
|
||||
return false;
|
||||
}
|
||||
}
|
||||
|
||||
/**
|
||||
* The herdr agent kind for this profile — the same value the matching member adapter passes, so
|
||||
* herdr resolves the same executable for a lead as it does for a member on that backend.
|
||||
*/
|
||||
private static String herdrKind(BridgedConfig.Profile profile) {
|
||||
return profile.isOpenCode() ? "opencode" : "claude";
|
||||
}
|
||||
|
||||
/**
|
||||
* The lead's argv: the profile's own command, the model pin, and the bridge MCP mount.
|
||||
*
|
||||
* <p>No {@code --append-system-prompt}. That flag carries the worker reply charter, and a lead
|
||||
* is not a worker — it reads its orchestration rules from the project's {@code CLAUDE.md} like
|
||||
* any other primary. This is the single most important difference from the member launchers;
|
||||
* do not "unify" it back.
|
||||
*/
|
||||
private List<String> leadArgv(BridgedConfig.Profile profile) {
|
||||
List<String> argv = new ArrayList<>(profile.argv());
|
||||
if (profile.hasMcp()) {
|
||||
argv.add("--mcp-config");
|
||||
argv.add("{\"mcpServers\":{\"bridge\":{\"type\":\"http\",\"url\":\""
|
||||
+ profile.mcpUrl() + "\"}}}");
|
||||
}
|
||||
// Appended last, for the same reason the member launcher does it (CB-533): the argv is
|
||||
// usually a wrapper such as `ccs <profile>`, which exports its own model family over
|
||||
// whatever it inherited, and --model outranks the environment.
|
||||
if (profile.model() != null && !profile.model().isBlank()) {
|
||||
argv.add("--model");
|
||||
argv.add(profile.model());
|
||||
}
|
||||
return argv;
|
||||
}
|
||||
|
||||
/**
|
||||
* The lead's environment: the profile's {@code env:} block, and nothing that could move it off
|
||||
* the subscription.
|
||||
*
|
||||
* <p>{@code ANTHROPIC_BASE_URL} and {@code ANTHROPIC_AUTH_TOKEN} are stripped unconditionally —
|
||||
* not defaulted, not guarded, stripped. A lead runs on the operator's subscription by
|
||||
* definition, so there is no configuration under which pointing it elsewhere is correct, and a
|
||||
* profile that carries them (a member profile reused as a lead's backend) must not leak them in.
|
||||
*/
|
||||
private Map<String, String> leadEnv(BridgedConfig.Profile profile) {
|
||||
Map<String, String> out = new LinkedHashMap<>();
|
||||
if (profile.env() != null) {
|
||||
out.putAll(profile.env());
|
||||
}
|
||||
out.remove("ANTHROPIC_BASE_URL");
|
||||
out.remove("ANTHROPIC_AUTH_TOKEN");
|
||||
if (profile.configDir() != null && !profile.configDir().isBlank()) {
|
||||
out.put("CLAUDE_CONFIG_DIR", profile.configDir());
|
||||
}
|
||||
// Deliberately no git token: a lead reviews and merges through the operator's own
|
||||
// credentials, and never needs the scoped write:repository token a member is granted.
|
||||
out.values().removeIf(Objects::isNull);
|
||||
return out;
|
||||
}
|
||||
}
|
||||
@@ -0,0 +1,30 @@
|
||||
package dev.ltms.bridged.logging;
|
||||
|
||||
import ch.qos.logback.classic.Level;
|
||||
import ch.qos.logback.classic.Logger;
|
||||
import ch.qos.logback.classic.turbo.TurboFilter;
|
||||
import ch.qos.logback.core.spi.FilterReply;
|
||||
import io.modelcontextprotocol.spec.McpSchema;
|
||||
import org.slf4j.Marker;
|
||||
|
||||
/** Suppresses only the SDK warning for the normal MCP cancellation notification. */
|
||||
public final class McpCancelledNotificationFilter extends TurboFilter {
|
||||
|
||||
static final String LOGGER = "io.modelcontextprotocol.spec.McpStreamableServerSession";
|
||||
static final String UNHANDLED_NOTIFICATION = "No handler registered for notification method: {}";
|
||||
|
||||
@Override
|
||||
public FilterReply decide(Marker marker, Logger logger, Level level, String format, Object[] params,
|
||||
Throwable throwable) {
|
||||
if (level == Level.WARN
|
||||
&& LOGGER.equals(logger.getName())
|
||||
&& UNHANDLED_NOTIFICATION.equals(format)
|
||||
&& params != null
|
||||
&& params.length == 1
|
||||
&& params[0] instanceof McpSchema.JSONRPCNotification notification
|
||||
&& "notifications/cancelled".equals(notification.method())) {
|
||||
return FilterReply.DENY;
|
||||
}
|
||||
return FilterReply.NEUTRAL;
|
||||
}
|
||||
}
|
||||
@@ -9,14 +9,15 @@ import dev.ltms.bridged.guard.GuardException;
|
||||
import dev.ltms.bridged.herdr.Agent;
|
||||
import dev.ltms.bridged.metrics.BridgedMetrics;
|
||||
import dev.ltms.bridged.metrics.Metrics;
|
||||
import dev.ltms.bridged.inject.WorkerPresence;
|
||||
import dev.ltms.bridged.inject.MemberPresence;
|
||||
import dev.ltms.bridged.herdr.HerdrException;
|
||||
import dev.ltms.bridged.msg.MessageService;
|
||||
import dev.ltms.bridged.msg.Rendezvous;
|
||||
import dev.ltms.bridged.peer.PeerUnreachableException;
|
||||
import dev.ltms.bridged.session.SessionManager;
|
||||
import dev.ltms.bridged.session.WorkerSession;
|
||||
import dev.ltms.bridged.session.MemberSession;
|
||||
import dev.ltms.bridged.session.WorktreeRequest;
|
||||
import dev.ltms.bridged.peer.MemberRole;
|
||||
import dev.ltms.bridged.peer.PeerLauncher;
|
||||
import io.modelcontextprotocol.common.McpTransportContext;
|
||||
import io.modelcontextprotocol.json.McpJsonMapper;
|
||||
@@ -32,7 +33,10 @@ import jakarta.servlet.http.HttpServlet;
|
||||
import java.util.LinkedHashMap;
|
||||
import java.util.List;
|
||||
import java.util.Map;
|
||||
import java.util.Set;
|
||||
import java.util.function.Function;
|
||||
import java.util.function.LongSupplier;
|
||||
import java.util.function.Supplier;
|
||||
import java.util.stream.Collectors;
|
||||
|
||||
/**
|
||||
@@ -72,17 +76,20 @@ public final class BridgeMcp {
|
||||
private final McpSyncServer server;
|
||||
private final CallerResolver authz; // CB-501: null → authorization not enforced (legacy)
|
||||
private final Metrics metrics; // CB-502: null → auth failures not counted
|
||||
private final CapacitySource capacity;
|
||||
private final HealthCoverageSource healthCoverage;
|
||||
|
||||
/**
|
||||
* Legacy constructor — no authorization. Retained so existing tests exercise tool behaviour
|
||||
* without an auth fixture.
|
||||
*/
|
||||
public BridgeMcp(MessageService messages, PeerLauncher workers,
|
||||
SessionManager sessions, ConnectionIdentity identity, WorkerPresence presence,
|
||||
PrimaryRegistry primaryRegistry) {
|
||||
this(messages, workers, sessions, identity, presence, primaryRegistry, null, null);
|
||||
/** Capacity facts used by {@code bridge_list}; production must supply the placement live count. */
|
||||
public record CapacitySource(Function<String, Integer> liveCount, Function<String, Integer> maxLoad,
|
||||
Supplier<Set<String>> configuredProfiles, LongSupplier clock) {
|
||||
/** Inert test-only source. It omits capacity rather than inventing zero live counts. */
|
||||
public static CapacitySource none() { return new CapacitySource(_ -> 0, _ -> null, Set::of, System::nanoTime); }
|
||||
boolean available() { return !configuredProfiles.get().isEmpty(); }
|
||||
}
|
||||
|
||||
/** Coverage is supplied by the health wiring, not inferred from a missing dependency. */
|
||||
public record HealthCoverageSource(Supplier<String> value) { }
|
||||
|
||||
/**
|
||||
* @param callers resolves each call's {@link Principal}; {@code null} disables authorization.
|
||||
* This surface needs its own enforcement: {@code /mcp} is a raw servlet on
|
||||
@@ -90,9 +97,11 @@ public final class BridgeMcp {
|
||||
* filter, so the REST guard does not cover it.
|
||||
* @param metrics registry for auth-failure counting; may be {@code null}
|
||||
*/
|
||||
public BridgeMcp(MessageService messages, PeerLauncher workers,
|
||||
SessionManager sessions, ConnectionIdentity identity, WorkerPresence presence,
|
||||
PrimaryRegistry primaryRegistry, CallerResolver callers, Metrics metrics) {
|
||||
public BridgeMcp(MessageService messages, PeerLauncher workers, SessionManager sessions,
|
||||
ConnectionIdentity identity, MemberPresence presence, PrimaryRegistry primaryRegistry,
|
||||
CallerResolver callers, Metrics metrics, CapacitySource capacity, HealthCoverageSource healthCoverage) {
|
||||
this.capacity = capacity;
|
||||
this.healthCoverage = healthCoverage;
|
||||
McpJsonMapper json = new JacksonMcpJsonMapperSupplier().get();
|
||||
this.transport = HttpServletStreamableServerTransportProvider.builder()
|
||||
.jsonMapper(json)
|
||||
@@ -108,10 +117,10 @@ public final class BridgeMcp {
|
||||
? callers.resolve(req.getRemoteAddr(), req.getRemotePort(),
|
||||
req.getHeader("Authorization"))
|
||||
: legacyPrincipal(identity, req.getRemoteAddr(), req.getRemotePort());
|
||||
// CB-532: guard on the ROLE, not on the terminal being null. A named lead now
|
||||
// carries its pane too, and enrolling a lead in the worker presence map would
|
||||
// have it counted as an available worker.
|
||||
if (p.isWorker()) presence.markPresent(p.terminal());
|
||||
// CB-532: guard on the ROLE, not on the terminal being null. This excludes a
|
||||
// lead, which carries its pane too, while including every spawned member role.
|
||||
// Enrolling a lead would count it as an available member in the roster.
|
||||
markSpawnedMemberPresent(p, presence);
|
||||
return McpTransportContext.create(Map.of(
|
||||
CALLER_TERMINAL, orEmpty(p.terminal()),
|
||||
CALLER_PID, Long.toString(p.pid()),
|
||||
@@ -150,8 +159,8 @@ public final class BridgeMcp {
|
||||
Runnable onAccepted = () -> primaryRegistry.recordDelegation(target, caller);
|
||||
// wait defaults to true (block for the reply); wait:false is fire-and-poll.
|
||||
return Boolean.FALSE.equals(a.get("wait"))
|
||||
? sendAsync(messages, target, content, onAccepted)
|
||||
: send(messages, target, content, timeoutMs(a), onAccepted);
|
||||
? sendAsync(messages, target, content, onAccepted, workers.profiles())
|
||||
: send(messages, target, content, timeoutMs(a), onAccepted, workers.profiles());
|
||||
})
|
||||
// bridge_reply's identity is the CONNECTION, never an argument — so the authz check
|
||||
// is "is this caller a worker at all", and it can only ever reply as itself.
|
||||
@@ -200,13 +209,13 @@ public final class BridgeMcp {
|
||||
// CB-301: carry the caller's identity as the session owner (null for the primary).
|
||||
// CB-301-ext: optional isolated worktree for parallel implementers.
|
||||
String callerCwd = identity.cwdForPid(callerPid(exchange));
|
||||
return spawn(sessions, str(a, "profile"), str(a, "cwd"), callerCwd,
|
||||
return spawn(sessions, str(a, "profile"), str(a, "role"), str(a, "cwd"), callerCwd,
|
||||
callerTerminal(exchange), worktreeRequest(a));
|
||||
})
|
||||
.toolCall(listTool(), (exchange, _) -> {
|
||||
McpSchema.CallToolResult denied = deny(exchange, Authz.Action.READ, null);
|
||||
if (denied != null) return denied;
|
||||
return listFleet(workers, sessions,
|
||||
return listFleet(workers, sessions, messages, capacity, healthCoverage,
|
||||
callers == null ? Map.of() : callers.leads(),
|
||||
callerTerminal(exchange));
|
||||
})
|
||||
@@ -364,6 +373,13 @@ public final class BridgeMcp {
|
||||
return transport;
|
||||
}
|
||||
|
||||
/** Mark a connected spawned member available for the injector readiness gate. */
|
||||
static void markSpawnedMemberPresent(Principal caller, MemberPresence presence) {
|
||||
if (caller.isSpawnedMember()) {
|
||||
presence.markPresent(caller.terminal());
|
||||
}
|
||||
}
|
||||
|
||||
/** Graceful shutdown of the MCP server. */
|
||||
public void close() {
|
||||
server.closeGracefully();
|
||||
@@ -371,21 +387,22 @@ public final class BridgeMcp {
|
||||
|
||||
// --- tool logic (thin adapters over the services; unit-testable) ---------------------------
|
||||
|
||||
/** {@code bridge_send}: delegate {@code content} to a worker session and block for its reply. */
|
||||
static McpSchema.CallToolResult send(MessageService messages, String sessionId, String content, Long timeoutMs) {
|
||||
return send(messages, sessionId, content, timeoutMs, null);
|
||||
}
|
||||
|
||||
/**
|
||||
* As {@link #send(MessageService, String, String, Long)}, wiring an accepted-delivery hook
|
||||
* {@code bridge_send}: delegate {@code content} to a worker session and block for its reply.
|
||||
* The configured profiles are required so a profile name can never bypass target validation.
|
||||
*
|
||||
* (CB-548): {@code onAccepted} records delegator ownership the instant the send is accepted, so
|
||||
* a BUSY interloper never claims a turn it did not win. {@code null} disables recording.
|
||||
*/
|
||||
static McpSchema.CallToolResult send(MessageService messages, String sessionId, String content,
|
||||
Long timeoutMs, Runnable onAccepted) {
|
||||
Long timeoutMs, Runnable onAccepted, Set<String> profiles) {
|
||||
if (isBlank(sessionId) || isBlank(content)) {
|
||||
return error("sessionId and content are required");
|
||||
}
|
||||
McpSchema.CallToolResult targetError = profileTargetError(sessionId, profiles);
|
||||
if (targetError != null) {
|
||||
return targetError;
|
||||
}
|
||||
long timeout = clamp(timeoutMs == null ? DEFAULT_TIMEOUT_MS : timeoutMs);
|
||||
try {
|
||||
return formatReply(messages.send(sessionId, content, timeout, onAccepted), timeout);
|
||||
@@ -455,24 +472,33 @@ public final class BridgeMcp {
|
||||
/**
|
||||
* {@code bridge_send} with {@code wait:false}: delegate {@code content} and return a ticket
|
||||
* immediately (fire-and-poll), so a long task isn't cut off by the caller's MCP call timeout.
|
||||
*/
|
||||
static McpSchema.CallToolResult sendAsync(MessageService messages, String sessionId, String content) {
|
||||
return sendAsync(messages, sessionId, content, null);
|
||||
}
|
||||
|
||||
/**
|
||||
* As {@link #sendAsync(MessageService, String, String)}, wiring the accepted-delivery hook
|
||||
* The configured profiles are required so a profile name can never bypass target validation.
|
||||
*
|
||||
* This wires the accepted-delivery hook
|
||||
* (CB-548) so an async flooding send records delegator ownership exactly once it is accepted.
|
||||
*/
|
||||
static McpSchema.CallToolResult sendAsync(MessageService messages, String sessionId, String content,
|
||||
Runnable onAccepted) {
|
||||
Runnable onAccepted, Set<String> profiles) {
|
||||
if (isBlank(sessionId) || isBlank(content)) {
|
||||
return error("sessionId and content are required");
|
||||
}
|
||||
McpSchema.CallToolResult targetError = profileTargetError(sessionId, profiles);
|
||||
if (targetError != null) {
|
||||
return targetError;
|
||||
}
|
||||
String ticket = messages.sendAsync(sessionId, content, onAccepted);
|
||||
return text("accepted — task delegated. Poll bridge_poll with ticket=" + ticket);
|
||||
}
|
||||
|
||||
/** A configured profile is never a send target; other unknown values may be herdr-owned panes. */
|
||||
private static McpSchema.CallToolResult profileTargetError(String sessionId, Set<String> profiles) {
|
||||
if (profiles.contains(sessionId)) {
|
||||
return error("unknown send target \"" + sessionId + "\": it is a configured profile name, not a "
|
||||
+ "session id. Call bridge_list to find a member or lead sessionId.");
|
||||
}
|
||||
return null;
|
||||
}
|
||||
|
||||
/** {@code bridge_poll}: check an async delegation by ticket, or drain a worker's inbox by target. */
|
||||
static McpSchema.CallToolResult poll(MessageService messages, String ticket, String target) {
|
||||
if (!isBlank(target)) {
|
||||
@@ -494,6 +520,9 @@ public final class BridgeMcp {
|
||||
? "[done — worker finished without a structured bridge_reply; transcript tail follows]\n" + v.reply()
|
||||
: v.reply());
|
||||
case PENDING -> text("[pending — " + v.detail() + "]");
|
||||
case ASKING -> text("[question — worker is waiting for your answer]\n" + v.reply()
|
||||
+ "\n\nAnswer it by calling bridge_send again with turnId=\"" + v.turnId()
|
||||
+ "\" and content set to your answer; the worker resumes the same turn.");
|
||||
case FAILED -> text("[failed — " + v.detail() + "]");
|
||||
};
|
||||
}
|
||||
@@ -606,23 +635,33 @@ public final class BridgeMcp {
|
||||
|
||||
/** {@code bridge_spawn} without cwd/caller context (default resolution). */
|
||||
static McpSchema.CallToolResult spawn(SessionManager sessions, String profile) {
|
||||
return spawn(sessions, profile, null, null, null, null);
|
||||
return spawn(sessions, profile, null, null, null, null, null);
|
||||
}
|
||||
|
||||
/**
|
||||
* {@code bridge_spawn}: launch a guard-checked worker for {@code profile} (blank → the default
|
||||
* profile) and return its session id + pane id. The worker's cwd is {@code requestedCwd} if given,
|
||||
* else the profile's config, else {@code callerCwd} (the primary's directory), else the daemon's.
|
||||
* {@code bridge_spawn}: launch a guard-checked member for {@code profile} (blank → the default
|
||||
* profile) under {@code role} (blank → {@code dev}), and return its session id + pane id. The
|
||||
* member's cwd is {@code requestedCwd} if given, else the profile's config, else
|
||||
* {@code callerCwd} (the primary's directory), else the daemon's.
|
||||
* CB-301: the session is registered with {@code ownerTerminal} as its owner.
|
||||
* CB-301-ext: {@code worktreeRequest} non-null provisions an isolated git worktree.
|
||||
*
|
||||
* <p>{@code role} and {@code profile} are independent: the role picks the contract, the profile
|
||||
* picks the backend. A reviewer on the same profile as the dev it reviews is a normal spawn.
|
||||
*/
|
||||
static McpSchema.CallToolResult spawn(SessionManager sessions, String profile,
|
||||
static McpSchema.CallToolResult spawn(SessionManager sessions, String profile, String role,
|
||||
String requestedCwd, String callerCwd,
|
||||
String ownerTerminal, WorktreeRequest worktreeRequest) {
|
||||
MemberRole memberRole;
|
||||
try {
|
||||
WorkerSession worker = sessions.acquire(isBlank(profile) ? null : profile,
|
||||
memberRole = isBlank(role) ? MemberRole.DEV : MemberRole.parse(role);
|
||||
} catch (IllegalArgumentException e) {
|
||||
return error(e.getMessage());
|
||||
}
|
||||
try {
|
||||
MemberSession member = sessions.acquire(isBlank(profile) ? null : profile, memberRole,
|
||||
requestedCwd, callerCwd, ownerTerminal, worktreeRequest);
|
||||
return text(json(workerView(worker)));
|
||||
return text(json(memberView(member)));
|
||||
} catch (GuardException e) {
|
||||
return error("subscription boundary: " + e.getMessage());
|
||||
} catch (IllegalArgumentException e) {
|
||||
@@ -690,6 +729,12 @@ public final class BridgeMcp {
|
||||
*/
|
||||
static McpSchema.CallToolResult listFleet(PeerLauncher workers, SessionManager sessions,
|
||||
Map<String, String> leads, String selfTerm) {
|
||||
return listFleet(workers, sessions, null, CapacitySource.none(), new HealthCoverageSource(() -> "off"), leads, selfTerm);
|
||||
}
|
||||
|
||||
static McpSchema.CallToolResult listFleet(PeerLauncher workers, SessionManager sessions, MessageService messages,
|
||||
CapacitySource capacity, HealthCoverageSource healthCoverage,
|
||||
Map<String, String> leads, String selfTerm) {
|
||||
try {
|
||||
Map<String, Agent> live = workers.list().stream()
|
||||
.map(Agent.class::cast)
|
||||
@@ -699,15 +744,57 @@ public final class BridgeMcp {
|
||||
.sorted(Map.Entry.comparingByValue())
|
||||
.map(e -> leadView(e.getKey(), e.getValue(), live.get(e.getKey()), selfTerm))
|
||||
.toList();
|
||||
List<Map<String, Object>> out = sessions.roster().stream()
|
||||
.map(s -> SessionManager.rosterView(s, live.get(s.terminalId())))
|
||||
List<MemberSession> roster = sessions.roster();
|
||||
List<Map<String, Object>> out = roster.stream()
|
||||
.map(s -> memberCapacityView(s, live.get(s.terminalId()), messages, capacity.clock().getAsLong()))
|
||||
.toList();
|
||||
return text(json(Map.of("leads", leadRows, "workers", out)));
|
||||
Set<String> profiles = new java.util.TreeSet<>(capacity.configuredProfiles().get());
|
||||
roster.stream().map(MemberSession::profile).forEach(profiles::add);
|
||||
Map<String, Object> result = new LinkedHashMap<>();
|
||||
result.put("leads", leadRows); result.put("members", out);
|
||||
result.put("healthCoverage", healthCoverage.value().get());
|
||||
if (capacity.available()) result.put("capacity", profiles.stream()
|
||||
.map(profile -> capacityView(profile, capacity.liveCount(), capacity.maxLoad(), roster, messages,
|
||||
capacity.clock().getAsLong())).toList());
|
||||
return text(json(result));
|
||||
} catch (HerdrException e) {
|
||||
return error("herdr error listing the fleet: " + e.getMessage());
|
||||
}
|
||||
}
|
||||
|
||||
/**
|
||||
* Capacity is advisory only. {@code reclaimable} says there is no bridge work, not that bridged
|
||||
* may stop the member: the bridge has capacity facts but no work list, and choosing work needs
|
||||
* authority it does not have. {@code idleForSeconds} is derived from monotonic nanoTime and has
|
||||
* no wall-clock meaning across a daemon restart.
|
||||
*/
|
||||
private static Map<String, Object> memberCapacityView(MemberSession session, Agent live,
|
||||
MessageService messages, long nowNanos) {
|
||||
Map<String, Object> row = SessionManager.rosterView(session, live);
|
||||
boolean open = messages != null && messages.hasAcceptedDelivery(session.terminalId());
|
||||
boolean inbox = messages != null && messages.hasInboxMessage(session.terminalId());
|
||||
boolean reclaimable = (session.state() == MemberSession.State.READY || session.state() == MemberSession.State.DONE)
|
||||
&& !open && !inbox;
|
||||
row.put("reclaimable", reclaimable);
|
||||
row.put("idleForSeconds", reclaimable ? Math.max(0, (nowNanos - session.lastActivityAtNanos()) / 1_000_000_000L) : null);
|
||||
return row;
|
||||
}
|
||||
|
||||
private static Map<String, Object> capacityView(String profile, Function<String, Integer> liveCount,
|
||||
Function<String, Integer> maxLoad, List<MemberSession> roster,
|
||||
MessageService messages, long nowNanos) {
|
||||
Integer cap = maxLoad.apply(profile);
|
||||
int live = liveCount.apply(profile);
|
||||
int reclaimable = (int) roster.stream().filter(s -> profile.equals(s.profile()))
|
||||
.filter(s -> (s.state() == MemberSession.State.READY || s.state() == MemberSession.State.DONE))
|
||||
.filter(s -> messages == null || (!messages.hasAcceptedDelivery(s.terminalId()) && !messages.hasInboxMessage(s.terminalId())))
|
||||
.count();
|
||||
Map<String, Object> row = new LinkedHashMap<>();
|
||||
row.put("profile", profile); row.put("maxLoad", cap); row.put("live", live);
|
||||
row.put("free", cap == null ? null : Math.max(0, cap - live)); row.put("reclaimable", reclaimable);
|
||||
return row;
|
||||
}
|
||||
|
||||
/**
|
||||
* One lead's row: its address, its name, and whether it can be reached right now.
|
||||
*
|
||||
@@ -743,10 +830,14 @@ public final class BridgeMcp {
|
||||
}
|
||||
|
||||
/** CB-301 projection from the authoritative session registry. */
|
||||
private static Map<String, Object> workerView(WorkerSession s) {
|
||||
private static Map<String, Object> memberView(MemberSession s) {
|
||||
Map<String, Object> m = new LinkedHashMap<>();
|
||||
m.put("sessionId", s.terminalId());
|
||||
m.put("paneId", s.paneId());
|
||||
// Echo the role back so the caller can see what it actually got, not what it meant to ask
|
||||
// for — a spawn that silently fell back to dev is otherwise invisible.
|
||||
m.put("role", s.role() == null ? "dev" : s.role().wireName());
|
||||
m.put("profile", s.profile());
|
||||
m.put("status", s.state().name().toLowerCase());
|
||||
if (s.worktree() != null) {
|
||||
m.put("worktree", s.worktree());
|
||||
@@ -823,14 +914,20 @@ public final class BridgeMcp {
|
||||
|
||||
private static McpSchema.Tool spawnTool() {
|
||||
return tool("bridge_spawn",
|
||||
"Spawn a new off-subscription worker session. Pass a profile (from bridge_profiles) to "
|
||||
+ "pick the backend, or omit it for the default. The worker opens your current "
|
||||
+ "directory by default; pass cwd to pin a different one. Pass worktree:true (with "
|
||||
+ "ticket) or worktree:<ticket-slug> to provision an isolated git worktree. "
|
||||
+ "Returns the worker's sessionId (use with bridge_send) and paneId (use with bridge_stop).",
|
||||
"Spawn a new off-subscription member session. A member has two independent attributes: "
|
||||
+ "role (what it is for) and profile (which backend it runs on). Pass role to pick "
|
||||
+ "the contract — 'dev' implements a unit and opens its own PR, 'reviewer' reviews a "
|
||||
+ "diff it did not write, 'architect' refines a ticket before anyone builds it; omit "
|
||||
+ "it for 'dev'. Pass profile (from bridge_profiles) to pick the backend, or omit it "
|
||||
+ "for the default. The two are independent: a reviewer may run on the same profile "
|
||||
+ "as the dev it reviews. The member opens your current directory by default; pass "
|
||||
+ "cwd to pin a different one. Pass worktree:true (with ticket) or "
|
||||
+ "worktree:<ticket-slug> to provision an isolated git worktree. Returns the member's "
|
||||
+ "sessionId (use with bridge_send) and paneId (use with bridge_stop).",
|
||||
objectSchema(Map.of(
|
||||
"profile", stringProp("Worker profile to spawn (omit for the default profile)"),
|
||||
"cwd", stringProp("Working directory for the worker (omit to inherit yours)"),
|
||||
"role", stringProp("What the member is for: architect, dev or reviewer (default dev)"),
|
||||
"profile", stringProp("Which backend to run it on (omit for the default profile)"),
|
||||
"cwd", stringProp("Working directory for the member (omit to inherit yours)"),
|
||||
"worktree", Map.of("type", "string", "description", "'true' or a ticket slug — requests an isolated git worktree"),
|
||||
"ticket", stringProp("Ticket slug when worktree:true")),
|
||||
List.of()));
|
||||
@@ -847,10 +944,11 @@ public final class BridgeMcp {
|
||||
"List the whole fleet the bridge tracks, in two parts. 'leads' are your PEERS — other "
|
||||
+ "orchestrators, each with its sessionId (the address to bridge_send to), "
|
||||
+ "name, live status, and 'self': true on your own row; this is how you "
|
||||
+ "discover a peer lead without being told its address. 'workers' are the "
|
||||
+ "sessions delegated to — each with sessionId, paneId, profile, state, "
|
||||
+ "optional worktree/branch/owner, and live herdr status. An empty 'workers' "
|
||||
+ "means no workers are spawned; it says nothing about peers.",
|
||||
+ "discover a peer lead without being told its address. 'members' are the "
|
||||
+ "sessions delegated to — each with sessionId, paneId, role (architect/dev/"
|
||||
+ "reviewer), profile (the backend it runs on), state, optional "
|
||||
+ "worktree/branch/owner, and live herdr status. An empty 'members' "
|
||||
+ "means no members are spawned; it says nothing about peers.",
|
||||
objectSchema(Map.of(), List.of()));
|
||||
}
|
||||
|
||||
|
||||
+56
-49
@@ -1,4 +1,4 @@
|
||||
package dev.ltms.bridged.worker;
|
||||
package dev.ltms.bridged.member;
|
||||
|
||||
import dev.ltms.bridged.config.BridgedConfig;
|
||||
import dev.ltms.bridged.guard.SubscriptionGuard;
|
||||
@@ -16,6 +16,7 @@ import java.util.Set;
|
||||
import java.util.UUID;
|
||||
import java.util.function.Function;
|
||||
import java.util.function.LongSupplier;
|
||||
import java.util.function.Supplier;
|
||||
|
||||
/**
|
||||
* The {@link HerdrPeerLauncher} adapter for <strong>Claude Code</strong> — the safe path from a
|
||||
@@ -43,29 +44,12 @@ public final class ClaudeCodeLauncher extends HerdrPeerLauncher {
|
||||
|
||||
private final SubscriptionGuard guard;
|
||||
|
||||
/**
|
||||
* Standing instruction appended to the worker's system prompt so it returns its result via
|
||||
* {@code bridge_reply}. Injected as a launch flag, so nothing is written to the worker's
|
||||
* profile — it is guidance, and a worker that never replies is caught by the send's timeout.
|
||||
*/
|
||||
static final String REPLY_CHARTER =
|
||||
"You are an off-subscription worker in the claude-bridge fleet. Every message you "
|
||||
+ "receive arrives through the bridge, and the ONLY channel back to the sender is the "
|
||||
+ "bridge_reply MCP tool. Text you write in your terminal is NOT sent anywhere — the "
|
||||
+ "sender cannot see your screen, so an in-terminal answer is silently discarded. "
|
||||
+ "Therefore you MUST end EVERY turn by calling bridge_reply with `content` set to your "
|
||||
+ "complete response. This holds for every message without exception — tasks, questions, "
|
||||
+ "clarifications, acknowledgements, and ordinary back-and-forth conversation. Call "
|
||||
+ "bridge_reply exactly once, as the final action of your turn, with your full answer in "
|
||||
+ "`content`; never wait for confirmation first. If you end a turn without calling "
|
||||
+ "bridge_reply, the sender receives nothing and the exchange stalls.";
|
||||
|
||||
/**
|
||||
* Production constructor — disables the spawn-ready gate ({@code spawnReadyTimeoutMs == 0}) so
|
||||
* existing deployments and tests keep the legacy non-blocking spawn semantics.
|
||||
*/
|
||||
public ClaudeCodeLauncher(AgentControl agents, WorkspaceControl spaces, SubscriptionGuard guard,
|
||||
Map<String, BridgedConfig.Worker> profiles, String defaultProfile,
|
||||
Map<String, BridgedConfig.Profile> profiles, String defaultProfile,
|
||||
Function<String, String> env) {
|
||||
this(agents, spaces, guard, profiles, defaultProfile, env, 0,
|
||||
System::currentTimeMillis, () -> sleepUninterruptibly(300));
|
||||
@@ -76,12 +60,27 @@ public final class ClaudeCodeLauncher extends HerdrPeerLauncher {
|
||||
* until the pane reports an injectable state or {@code spawnReadyTimeoutMs} elapses.
|
||||
*/
|
||||
public ClaudeCodeLauncher(AgentControl agents, WorkspaceControl spaces, SubscriptionGuard guard,
|
||||
Map<String, BridgedConfig.Worker> profiles, String defaultProfile,
|
||||
Map<String, BridgedConfig.Profile> profiles, String defaultProfile,
|
||||
Function<String, String> env,
|
||||
long spawnReadyTimeoutMs, long spawnReadyPollMs) {
|
||||
this(agents, spaces, guard, profiles, defaultProfile, env,
|
||||
spawnReadyTimeoutMs, spawnReadyPollMs, null);
|
||||
}
|
||||
|
||||
/**
|
||||
* Production constructor carrying the fleet-wide tab-label template (CB-557). The template comes
|
||||
* from {@code fleet.tabLabel}, which a profile cannot know because it names the member's
|
||||
* <em>role</em>; a profile may still override it with its own {@code tabLabel}.
|
||||
*/
|
||||
public ClaudeCodeLauncher(AgentControl agents, WorkspaceControl spaces, SubscriptionGuard guard,
|
||||
Map<String, BridgedConfig.Profile> profiles, String defaultProfile,
|
||||
Function<String, String> env,
|
||||
long spawnReadyTimeoutMs, long spawnReadyPollMs,
|
||||
Supplier<BridgedConfig.Fleet> fleet) {
|
||||
this(agents, spaces, guard, profiles, defaultProfile, env,
|
||||
spawnReadyTimeoutMs,
|
||||
System::currentTimeMillis, () -> sleepUninterruptibly(spawnReadyPollMs));
|
||||
System::currentTimeMillis, () -> sleepUninterruptibly(spawnReadyPollMs),
|
||||
fleet);
|
||||
}
|
||||
|
||||
/**
|
||||
@@ -102,24 +101,28 @@ public final class ClaudeCodeLauncher extends HerdrPeerLauncher {
|
||||
* (poll ms) is accepted for API symmetry but otherwise unused here
|
||||
*/
|
||||
public ClaudeCodeLauncher(AgentControl agents, WorkspaceControl spaces, SubscriptionGuard guard,
|
||||
Map<String, BridgedConfig.Worker> profiles, String defaultProfile,
|
||||
Map<String, BridgedConfig.Profile> profiles, String defaultProfile,
|
||||
Function<String, String> env,
|
||||
long spawnReadyTimeoutMs,
|
||||
LongSupplier nowMillis, Runnable sleeper) {
|
||||
super(NAME_PREFIX, agents, spaces, profiles, defaultProfile, env,
|
||||
spawnReadyTimeoutMs, nowMillis, sleeper);
|
||||
this.guard = guard;
|
||||
this(agents, spaces, guard, profiles, defaultProfile, env,
|
||||
spawnReadyTimeoutMs, nowMillis, sleeper, null);
|
||||
}
|
||||
|
||||
/**
|
||||
* {@inheritDoc}
|
||||
* Full testability constructor, plus the fleet-wide tab-label template (CB-557).
|
||||
*
|
||||
* <p>A legacy spawn with no session identity is a fresh, launcher-derived session — delegate to
|
||||
* the session-aware form with no name and no resume id.
|
||||
* @param fleet live fleet config, read once for each spawn
|
||||
*/
|
||||
@Override
|
||||
protected Launch buildLaunch(BridgedConfig.Worker cfg) {
|
||||
return buildLaunch(cfg, null, null);
|
||||
public ClaudeCodeLauncher(AgentControl agents, WorkspaceControl spaces, SubscriptionGuard guard,
|
||||
Map<String, BridgedConfig.Profile> profiles, String defaultProfile,
|
||||
Function<String, String> env,
|
||||
long spawnReadyTimeoutMs,
|
||||
LongSupplier nowMillis, Runnable sleeper,
|
||||
Supplier<BridgedConfig.Fleet> fleet) {
|
||||
super(NAME_PREFIX, agents, spaces, profiles, defaultProfile, env,
|
||||
spawnReadyTimeoutMs, nowMillis, sleeper, fleet);
|
||||
this.guard = guard;
|
||||
}
|
||||
|
||||
/**
|
||||
@@ -132,7 +135,7 @@ public final class ClaudeCodeLauncher extends HerdrPeerLauncher {
|
||||
* applied here — see {@link #applySessionIdentity}.
|
||||
*/
|
||||
@Override
|
||||
protected Launch buildLaunch(BridgedConfig.Worker cfg, String sessionName, String resumeSessionId) {
|
||||
protected Launch buildLaunch(BridgedConfig.Profile cfg, LaunchSpec spec) {
|
||||
// CB-539: a profile may deliberately opt into the subscription (subscription: true) when no
|
||||
// off-subscription endpoint exists for it — e.g. `sonnet` on `ccs`. That profile gets no
|
||||
// ANTHROPIC_BASE_URL/AUTH_TOKEN (there is nothing to point them at) and the guard's base_url
|
||||
@@ -178,9 +181,9 @@ public final class ClaudeCodeLauncher extends HerdrPeerLauncher {
|
||||
// id via -r and passes no --session-id (the two conflict). Both are injected before the
|
||||
// model flag so --model keeps outranking the operator's own argv.
|
||||
// mutableArgv: argvWithBridge may hand back the profile's own (immutable) List.of when it
|
||||
// has no MCP — session flags must be added into a list we own.
|
||||
List<String> argv = mutableArgv(argvWithBridge(cfg));
|
||||
String agentSessionId = applySessionIdentity(argv, sessionName, resumeSessionId);
|
||||
// has neither MCP nor a charter — session flags must be added into a list we own.
|
||||
List<String> argv = mutableArgv(argvWithBridge(cfg, spec.charter()));
|
||||
String agentSessionId = applySessionIdentity(argv, spec.sessionName(), spec.resumeSessionId());
|
||||
return new Launch(workerEnv, argvWithModel(argv, cfg), agentSessionId);
|
||||
}
|
||||
|
||||
@@ -214,22 +217,26 @@ public final class ClaudeCodeLauncher extends HerdrPeerLauncher {
|
||||
}
|
||||
|
||||
/**
|
||||
* The launch argv, plus — when {@code worker.mcpUrl} is set — inline {@code --mcp-config} for
|
||||
* the bridge server and {@code --append-system-prompt} for the {@link #REPLY_CHARTER}. Neither
|
||||
* touches the profile's config; both are pure command-line flags. This inline-flag mount is
|
||||
* Claude Code specific — other adapters mount MCP and instructions their own way.
|
||||
* The launch argv, plus an inline {@code --mcp-config} when {@code worker.mcpUrl} is set and
|
||||
* {@code --append-system-prompt} when the base composed a charter. Neither touches the profile's
|
||||
* config; both are pure command-line flags. This inline-flag mount is Claude Code specific —
|
||||
* other adapters mount MCP and instructions their own way.
|
||||
*/
|
||||
private List<String> argvWithBridge(BridgedConfig.Worker cfg) {
|
||||
if (!cfg.hasMcp()) {
|
||||
private List<String> argvWithBridge(BridgedConfig.Profile cfg, String charter) {
|
||||
if (!cfg.hasMcp() && charter == null) {
|
||||
return cfg.argv();
|
||||
}
|
||||
String mcpJson = "{\"mcpServers\":{\"bridge\":{\"type\":\"http\",\"url\":\""
|
||||
+ cfg.mcpUrl() + "\"}}}";
|
||||
List<String> argv = mutableArgv(cfg.argv());
|
||||
argv.add("--mcp-config");
|
||||
argv.add(mcpJson);
|
||||
argv.add("--append-system-prompt");
|
||||
argv.add(REPLY_CHARTER);
|
||||
if (cfg.hasMcp()) {
|
||||
String mcpJson = "{\"mcpServers\":{\"bridge\":{\"type\":\"http\",\"url\":\""
|
||||
+ cfg.mcpUrl() + "\"}}}";
|
||||
argv.add("--mcp-config");
|
||||
argv.add(mcpJson);
|
||||
}
|
||||
if (charter != null) {
|
||||
argv.add("--append-system-prompt");
|
||||
argv.add(charter);
|
||||
}
|
||||
return argv;
|
||||
}
|
||||
|
||||
@@ -249,7 +256,7 @@ public final class ClaudeCodeLauncher extends HerdrPeerLauncher {
|
||||
* {@code gx10} does) are untouched — this adds nothing when there is nothing to add. This is
|
||||
* the {@code kind: claude} counterpart of the opencode adapter's {@code -m provider/model}.
|
||||
*/
|
||||
private static List<String> argvWithModel(List<String> argv, BridgedConfig.Worker cfg) {
|
||||
private static List<String> argvWithModel(List<String> argv, BridgedConfig.Profile cfg) {
|
||||
if (cfg.model() == null || cfg.model().isBlank()) {
|
||||
return argv;
|
||||
}
|
||||
@@ -302,7 +309,7 @@ public final class ClaudeCodeLauncher extends HerdrPeerLauncher {
|
||||
|
||||
/** Whether any configured profile opts into a git-forge token (required for {@link Capability#SELF_PR}). */
|
||||
private boolean hasGitTokenProfile() {
|
||||
return profileConfigs().stream().anyMatch(BridgedConfig.Worker::hasGitToken);
|
||||
return profileConfigs().stream().anyMatch(BridgedConfig.Profile::hasGitToken);
|
||||
}
|
||||
|
||||
// --- CB-117 reap predicate (Claude prefix), kept for direct unit testing -------------------
|
||||
+175
-28
@@ -1,14 +1,16 @@
|
||||
package dev.ltms.bridged.worker;
|
||||
package dev.ltms.bridged.member;
|
||||
|
||||
import dev.ltms.bridged.config.BridgedConfig;
|
||||
import dev.ltms.bridged.herdr.Agent;
|
||||
import dev.ltms.bridged.peer.Capability;
|
||||
import dev.ltms.bridged.peer.MemberRole;
|
||||
import dev.ltms.bridged.peer.PeerHandle;
|
||||
import dev.ltms.bridged.peer.PeerLauncher;
|
||||
import dev.ltms.bridged.peer.PeerUnreachableException;
|
||||
import dev.ltms.bridged.peer.SpawnRequest;
|
||||
import dev.ltms.bridged.placement.PlacementCandidate;
|
||||
import dev.ltms.bridged.placement.PlacementContext;
|
||||
import dev.ltms.bridged.placement.PlacementException;
|
||||
import dev.ltms.bridged.placement.PlacementPolicies;
|
||||
import dev.ltms.bridged.placement.PlacementPolicy;
|
||||
import org.slf4j.Logger;
|
||||
@@ -24,6 +26,7 @@ import java.util.Map;
|
||||
import java.util.Set;
|
||||
import java.util.concurrent.ConcurrentHashMap;
|
||||
import java.util.function.Function;
|
||||
import java.util.function.Supplier;
|
||||
|
||||
/**
|
||||
* The {@link PeerLauncher} the core actually holds when more than one adapter is configured — a thin
|
||||
@@ -63,10 +66,25 @@ public final class CompositePeerLauncher implements PeerLauncher {
|
||||
/** paneId → the delegate that spawned it, so {@link #stop} tears down through the right adapter. */
|
||||
private final Map<String, HerdrPeerLauncher> spawnedBy = new ConcurrentHashMap<>();
|
||||
|
||||
private final Map<String, BridgedConfig.Worker> profileConfigs;
|
||||
private final PlacementPolicy placementPolicy;
|
||||
private final Function<String, Integer> liveCount;
|
||||
|
||||
/**
|
||||
* CB-559: the placement inputs are read <em>per spawn</em>, not captured at construction, so a
|
||||
* config reload changes where the next member lands without a restart. These are the hot keys —
|
||||
* role pools, an existing profile's weight/maxLoad, and the placement policy. What cannot change
|
||||
* this way is the set of adapters ({@link #byProfile}), because a new backend needs a launcher
|
||||
* and launchers are built once; {@code ConfigRef} classifies that as deferred and says so.
|
||||
*/
|
||||
private final Supplier<Map<String, BridgedConfig.Profile>> profileConfigs;
|
||||
private final Supplier<PlacementPolicy> placementPolicy;
|
||||
|
||||
/**
|
||||
* CB-557: the role pools an unqualified spawn draws its candidates from. A supplier that yields
|
||||
* {@code null}, and an empty pool for a role, both fall back to every configured profile — the
|
||||
* pre-CB-557 behaviour.
|
||||
*/
|
||||
private final Supplier<BridgedConfig.Fleet> fleet;
|
||||
|
||||
/**
|
||||
* Backward-compatible constructor: fixed placement, no live-counting. Use this for tests and
|
||||
* simple wiring; it preserves the pre-CB-518 behaviour exactly.
|
||||
@@ -77,7 +95,7 @@ public final class CompositePeerLauncher implements PeerLauncher {
|
||||
* @throws IllegalArgumentException if {@code delegates} is empty or two adapters claim one profile
|
||||
*/
|
||||
public CompositePeerLauncher(List<HerdrPeerLauncher> delegates, String defaultProfile) {
|
||||
this(delegates, defaultProfile, Map.of(), PlacementPolicies.fixed(), name -> 0);
|
||||
this(delegates, defaultProfile, Map.of(), PlacementPolicies.fixed(), _ -> 0);
|
||||
}
|
||||
|
||||
/**
|
||||
@@ -92,18 +110,65 @@ public final class CompositePeerLauncher implements PeerLauncher {
|
||||
*/
|
||||
public CompositePeerLauncher(List<HerdrPeerLauncher> delegates,
|
||||
String defaultProfile,
|
||||
Map<String, BridgedConfig.Worker> profileConfigs,
|
||||
Map<String, BridgedConfig.Profile> profileConfigs,
|
||||
PlacementPolicy placementPolicy,
|
||||
Function<String, Integer> liveCount) {
|
||||
this(delegates, defaultProfile, profileConfigs, placementPolicy, liveCount, null);
|
||||
}
|
||||
|
||||
/**
|
||||
* Production constructor with role pools (CB-557). An unqualified spawn draws its candidates from
|
||||
* {@code fleet.<role>} instead of from every configured profile, so a reviewer is placed on a
|
||||
* reviewer backend and never on, say, the architect-only one.
|
||||
*
|
||||
* @param fleet the configured role pools; {@code null} ⇒ every profile is a candidate for every
|
||||
* role, which is the pre-CB-557 behaviour
|
||||
*/
|
||||
public CompositePeerLauncher(List<HerdrPeerLauncher> delegates,
|
||||
String defaultProfile,
|
||||
Map<String, BridgedConfig.Profile> profileConfigs,
|
||||
PlacementPolicy placementPolicy,
|
||||
Function<String, Integer> liveCount,
|
||||
BridgedConfig.Fleet fleet) {
|
||||
// LinkedHashMap, not Map.copyOf: candidates() promises definition order and the weighted
|
||||
// policy breaks exact-weight ties on it, so a salted iteration order would make placement
|
||||
// differ from one JVM run to the next.
|
||||
this(delegates, defaultProfile,
|
||||
constant(Collections.unmodifiableMap(new LinkedHashMap<>(profileConfigs))),
|
||||
constant(placementPolicy), liveCount, constant(fleet));
|
||||
}
|
||||
|
||||
/**
|
||||
* Production constructor that re-reads its placement inputs per spawn (CB-559), so a config
|
||||
* reload retargets the next member without a restart.
|
||||
*
|
||||
* @param config the live configuration — read at every spawn, never captured
|
||||
*/
|
||||
public CompositePeerLauncher(List<HerdrPeerLauncher> delegates,
|
||||
String defaultProfile,
|
||||
Supplier<BridgedConfig> config,
|
||||
Function<String, Integer> liveCount) {
|
||||
this(delegates, defaultProfile,
|
||||
() -> config.get().profiles(),
|
||||
() -> PlacementPolicies.fromName(config.get().placement()),
|
||||
liveCount,
|
||||
() -> config.get().fleet());
|
||||
}
|
||||
|
||||
/** The all-suppliers form every other constructor funnels into. */
|
||||
private CompositePeerLauncher(List<HerdrPeerLauncher> delegates,
|
||||
String defaultProfile,
|
||||
Supplier<Map<String, BridgedConfig.Profile>> profileConfigs,
|
||||
Supplier<PlacementPolicy> placementPolicy,
|
||||
Function<String, Integer> liveCount,
|
||||
Supplier<BridgedConfig.Fleet> fleet) {
|
||||
this.fleet = fleet;
|
||||
if (delegates.isEmpty()) {
|
||||
throw new IllegalArgumentException("at least one peer adapter must be configured");
|
||||
}
|
||||
this.delegates = List.copyOf(delegates);
|
||||
this.defaultProfile = defaultProfile;
|
||||
// LinkedHashMap, not Map.copyOf: candidates() promises definition order and the weighted
|
||||
// policy breaks exact-weight ties on it, so a salted iteration order would make placement
|
||||
// differ from one JVM run to the next.
|
||||
this.profileConfigs = Collections.unmodifiableMap(new LinkedHashMap<>(profileConfigs));
|
||||
this.profileConfigs = profileConfigs;
|
||||
this.placementPolicy = placementPolicy;
|
||||
this.liveCount = liveCount;
|
||||
Map<String, HerdrPeerLauncher> index = new LinkedHashMap<>();
|
||||
@@ -120,6 +185,22 @@ public final class CompositePeerLauncher implements PeerLauncher {
|
||||
this.byProfile = Collections.unmodifiableMap(index);
|
||||
}
|
||||
|
||||
/** A supplier of a value fixed at construction — how the non-reloading constructors funnel in. */
|
||||
private static <T> Supplier<T> constant(T value) {
|
||||
return () -> value;
|
||||
}
|
||||
|
||||
/**
|
||||
* The currently-configured profiles, never null.
|
||||
*
|
||||
* <p>Read fresh on every call so a reload is visible; a caller that needs two consistent reads
|
||||
* takes one local, as {@link #poolFor} does.
|
||||
*/
|
||||
private Map<String, BridgedConfig.Profile> profiles0() {
|
||||
Map<String, BridgedConfig.Profile> m = profileConfigs.get();
|
||||
return m == null ? Map.of() : m;
|
||||
}
|
||||
|
||||
/** The adapter owning {@code profileName} (null/blank → the default). Throws on an unknown profile. */
|
||||
private HerdrPeerLauncher route(String profileName) {
|
||||
String resolved = (profileName == null || profileName.isBlank()) ? defaultProfile : profileName;
|
||||
@@ -139,27 +220,32 @@ public final class CompositePeerLauncher implements PeerLauncher {
|
||||
public PeerHandle spawn(SpawnRequest req) {
|
||||
String requestedProfile = req.profileName();
|
||||
if (requestedProfile != null && !requestedProfile.isBlank()) {
|
||||
// An explicit profile bypasses the policy entirely.
|
||||
// An explicit profile bypasses the placement policy, but not the capacity cap: maxLoad
|
||||
// is documented as an unconditional limit on this profile (BridgedConfig.Profile), and
|
||||
// the charter makes explicit-profile spawns the normal path — so skipping the check
|
||||
// here would leave the cap dead config in real operation.
|
||||
HerdrPeerLauncher d = route(requestedProfile);
|
||||
enforceMaxLoad(requestedProfile);
|
||||
PeerHandle handle = d.spawn(req);
|
||||
spawnedBy.put(handle.id(), d);
|
||||
return handle;
|
||||
}
|
||||
|
||||
List<PlacementCandidate> candidates = candidates();
|
||||
// CB-557: an unqualified spawn is placed inside the pool of the role it asked for, not across
|
||||
// the whole profile list. An EXPLICIT profile (above) is left alone on purpose — it is the
|
||||
// operator overriding, and refusing it would break `bridge_spawn{profile:"opus"}`, which
|
||||
// carries no role and so would be judged against the dev pool it was never meant for.
|
||||
List<PlacementCandidate> candidates = candidates(req.role());
|
||||
String roleDefault = defaultProfileFor(req.role());
|
||||
Set<String> unreachable = new HashSet<>();
|
||||
PlacementContext ctx = new PlacementContext(defaultProfile, candidates, liveCount, unreachable);
|
||||
PlacementContext ctx = new PlacementContext(roleDefault, candidates, liveCount, unreachable);
|
||||
|
||||
int maxAttempts = candidates.isEmpty() ? 1 : candidates.size();
|
||||
for (int attempt = 0; attempt < maxAttempts; attempt++) {
|
||||
PlacementCandidate chosen;
|
||||
try {
|
||||
chosen = placementPolicy.select(ctx);
|
||||
} catch (RuntimeException e) {
|
||||
// No candidate left (all at cap or all unreachable). The policy already threw a clear
|
||||
// message; do not wrap it in a generic PeerUnreachableException.
|
||||
throw e;
|
||||
}
|
||||
// Deliberately uncaught: when no candidate is left (all at cap, or all unreachable) the
|
||||
// policy already throws a clear message. Catching it to rethrow a generic
|
||||
// PeerUnreachableException would replace a precise diagnosis with a vague one.
|
||||
PlacementCandidate chosen = placementPolicy.get().select(ctx);
|
||||
|
||||
HerdrPeerLauncher d = byProfile.get(chosen.profile());
|
||||
if (d == null) {
|
||||
@@ -169,9 +255,10 @@ public final class CompositePeerLauncher implements PeerLauncher {
|
||||
}
|
||||
|
||||
// CB-547a: route the chosen profile but keep the caller's session identity — dropping it
|
||||
// here would silently sever the resume handle on every policy-routed spawn.
|
||||
// here would silently sever the resume handle on every policy-routed spawn. CB-557: the
|
||||
// role rides along for the same reason, or a routed spawn would be labelled as a dev.
|
||||
SpawnRequest routedReq = new SpawnRequest(chosen.profile(), req.requestedCwd(), req.callerCwd(),
|
||||
req.sessionName(), req.resumeSessionId());
|
||||
req.sessionName(), req.resumeSessionId(), req.role());
|
||||
try {
|
||||
PeerHandle handle = d.spawn(routedReq);
|
||||
spawnedBy.put(handle.id(), d);
|
||||
@@ -181,7 +268,7 @@ public final class CompositePeerLauncher implements PeerLauncher {
|
||||
chosen.profile(), e.getMessage());
|
||||
unreachable.add(chosen.profile());
|
||||
// Update the context for the next selection so the policy excludes this profile.
|
||||
ctx = new PlacementContext(defaultProfile, candidates, liveCount, unreachable);
|
||||
ctx = new PlacementContext(roleDefault, candidates, liveCount, unreachable);
|
||||
}
|
||||
}
|
||||
|
||||
@@ -190,12 +277,72 @@ public final class CompositePeerLauncher implements PeerLauncher {
|
||||
+ " candidate(s): " + String.join(", ", unreachable));
|
||||
}
|
||||
|
||||
/** Build the candidate list from the configured profiles, in definition order. */
|
||||
private List<PlacementCandidate> candidates() {
|
||||
/**
|
||||
* Refuse an explicit-profile spawn when the profile is at its {@code maxLoad} cap.
|
||||
*
|
||||
* <p>maxLoad is a documented, unconditional capacity limit (see {@code BridgedConfig.Profile#maxLoad}),
|
||||
* and the charter makes explicit-profile spawns the normal path — so enforcing it only in placement
|
||||
* ({@code PlacementPolicyUtil}, package-private, hence not linked) would leave the cap dead config
|
||||
* on every call that names a profile. Same rule as placement: {@code live >= cap} is at capacity.
|
||||
*
|
||||
* <p>Deliberately no fallback to another profile: the caller named {@code profile} for a cost/model
|
||||
* reason, and silently re-routing a paid-tier (subscription) request elsewhere is worse than
|
||||
* refusing it. A caller that wants placement should omit the profile and let the policy pick.
|
||||
*
|
||||
* <p>Known TOCTOU limitation — documented, not fixed. {@link #liveCount} is read outside any lock and
|
||||
* {@code SessionManager} registers a session only after {@code launcher.spawn} returns, so two
|
||||
* genuinely concurrent spawns can both pass this check. The race already exists on the placement
|
||||
* path. Closing it needs slot reservation in the registry; serializing spawn here would block on
|
||||
* the readiness gate and is a far worse trade.
|
||||
*
|
||||
* @param profile the profile the caller explicitly named
|
||||
* @throws PlacementException when the profile is at capacity
|
||||
*/
|
||||
private void enforceMaxLoad(String profile) {
|
||||
// Absent config, or a config whose maxLoad normalized to null (non-positive ⇒ unlimited at
|
||||
// load), means no cap — never cap what wasn't configured.
|
||||
BridgedConfig.Profile cfg = profiles0().get(profile);
|
||||
Integer cap = (cfg == null) ? null : cfg.maxLoad();
|
||||
if (cap == null) {
|
||||
return;
|
||||
}
|
||||
int live = liveCount.apply(profile);
|
||||
if (live >= cap) {
|
||||
throw new PlacementException("worker profile '" + profile + "' is at maxLoad: " + live
|
||||
+ " live >= " + cap + " cap; refusing spawn — no fallback to another profile");
|
||||
}
|
||||
}
|
||||
|
||||
/**
|
||||
* The profile names {@code role} may be placed on, in definition order.
|
||||
*
|
||||
* <p>An empty or absent pool means "unconstrained", not "nothing allowed": a config that declares
|
||||
* no pool for a role must keep spawning, so it falls back to every configured profile. Names in a
|
||||
* pool that no adapter declares are dropped here rather than thrown — config load already rejects
|
||||
* a pool entry with no profile, so a survivor is a profile this particular composite does not own.
|
||||
*/
|
||||
private List<String> poolFor(MemberRole role) {
|
||||
Map<String, BridgedConfig.Profile> configured = profiles0();
|
||||
BridgedConfig.Fleet f = fleet.get();
|
||||
List<String> pool = (f == null) ? List.of() : f.profilesFor(role);
|
||||
List<String> known = pool.stream().filter(configured::containsKey).toList();
|
||||
return known.isEmpty() ? List.copyOf(configured.keySet()) : known;
|
||||
}
|
||||
|
||||
/** The profile an unqualified spawn for {@code role} falls back to under {@code fixed} placement. */
|
||||
private String defaultProfileFor(MemberRole role) {
|
||||
List<String> pool = poolFor(role);
|
||||
return pool.isEmpty() ? defaultProfile : pool.getFirst();
|
||||
}
|
||||
|
||||
/** Build the candidate list from {@code role}'s pool, in definition order. */
|
||||
private List<PlacementCandidate> candidates(MemberRole role) {
|
||||
List<PlacementCandidate> out = new ArrayList<>();
|
||||
for (Map.Entry<String, BridgedConfig.Worker> e : profileConfigs.entrySet()) {
|
||||
BridgedConfig.Worker w = e.getValue();
|
||||
out.add(new PlacementCandidate(e.getKey(), null, w.weight(), w.maxLoad()));
|
||||
for (String name : poolFor(role)) {
|
||||
BridgedConfig.Profile w = profiles0().get(name);
|
||||
if (w != null) {
|
||||
out.add(new PlacementCandidate(name, null, w.weight(), w.maxLoad()));
|
||||
}
|
||||
}
|
||||
return out;
|
||||
}
|
||||
+114
-44
@@ -1,4 +1,4 @@
|
||||
package dev.ltms.bridged.worker;
|
||||
package dev.ltms.bridged.member;
|
||||
|
||||
import dev.ltms.bridged.config.BridgedConfig;
|
||||
import dev.ltms.bridged.herdr.Agent;
|
||||
@@ -7,6 +7,7 @@ import dev.ltms.bridged.herdr.HerdrException;
|
||||
import dev.ltms.bridged.herdr.Tab;
|
||||
import dev.ltms.bridged.herdr.Workspace;
|
||||
import dev.ltms.bridged.herdr.WorkspaceControl;
|
||||
import dev.ltms.bridged.peer.MemberRole;
|
||||
import dev.ltms.bridged.peer.PeerHandle;
|
||||
import dev.ltms.bridged.peer.PeerLauncher;
|
||||
import dev.ltms.bridged.peer.PeerUnreachableException;
|
||||
@@ -28,6 +29,7 @@ import java.util.concurrent.atomic.AtomicBoolean;
|
||||
import java.util.concurrent.atomic.AtomicLong;
|
||||
import java.util.function.Function;
|
||||
import java.util.function.LongSupplier;
|
||||
import java.util.function.Supplier;
|
||||
import java.util.regex.Matcher;
|
||||
import java.util.regex.Pattern;
|
||||
|
||||
@@ -42,7 +44,7 @@ import java.util.regex.Pattern;
|
||||
* <li>{@code namePrefix} (constructor arg) — the label prefix ({@code claude}, {@code opencode})
|
||||
* that drives both unique naming and the orphan-reap pattern, so each adapter reaps only its
|
||||
* own kind of pane and never another's.</li>
|
||||
* <li>{@link #buildLaunch(BridgedConfig.Worker)} — the peer-specific env map + argv, including any
|
||||
* <li>{@link #buildLaunch(BridgedConfig.Profile, LaunchSpec)} — the peer-specific env map + argv, including any
|
||||
* subscription/guard check, MCP mount, and instruction injection. The base never sees how the
|
||||
* peer is configured; it only places and starts the returned {@link Launch}.</li>
|
||||
* </ul>
|
||||
@@ -69,13 +71,48 @@ public abstract class HerdrPeerLauncher implements PeerLauncher {
|
||||
private final String namePrefix; // label prefix: naming + reap scheme
|
||||
private final AgentControl agents;
|
||||
private final WorkspaceControl spaces;
|
||||
private final Map<String, BridgedConfig.Worker> profiles; // profile name → spawn settings
|
||||
private final Map<String, BridgedConfig.Profile> profiles; // profile name → spawn settings
|
||||
private final String defaultProfile; // profile a no-arg spawn uses (nullable)
|
||||
|
||||
/** Host env lookup (injectable for tests); adapters read it in {@link #buildLaunch}. */
|
||||
protected final Function<String, String> env;
|
||||
|
||||
private final AtomicLong nameSeq = new AtomicLong(); // per-peer counter (also the tab #)
|
||||
private final AtomicLong nameSeq = new AtomicLong(); // per-peer counter (herdr agent names only)
|
||||
|
||||
/**
|
||||
* Live fleet config, read once per spawn. A null supplier or value leaves tab labels at their
|
||||
* default and supplies no role charter. A profile's own {@code tabLabel} still overrides it.
|
||||
*
|
||||
* <p>CB-559: a supplier rather than a snapshot, so a config reload affects the next launch
|
||||
* without a restart. Existing tabs keep the label they were given.
|
||||
*/
|
||||
private final Supplier<BridgedConfig.Fleet> fleet;
|
||||
|
||||
/** The final instruction always requires a bridge reply when the bridge MCP is mounted. */
|
||||
protected static final String REPLY_CHARTER =
|
||||
"You are a spawned member in the claude-bridge fleet. Every message you receive arrives "
|
||||
+ "through the bridge, and the ONLY channel back to the sender is the bridge_reply MCP tool. "
|
||||
+ "Text you write in your terminal is NOT sent anywhere — the sender cannot see your screen, "
|
||||
+ "so an in-terminal answer is silently discarded. Therefore you MUST end EVERY turn by calling "
|
||||
+ "bridge_reply with `content` set to your complete response. This holds for every message without "
|
||||
+ "exception — tasks, questions, clarifications, acknowledgements, and ordinary back-and-forth "
|
||||
+ "conversation. Call bridge_reply exactly once, as the final action of your turn, with your full "
|
||||
+ "answer in `content`; never wait for confirmation first. If you end a turn without calling "
|
||||
+ "bridge_reply, the sender receives nothing and the exchange stalls.";
|
||||
|
||||
/**
|
||||
* Tab numbers, counted per {@code role/profile} pair (CB-557).
|
||||
*
|
||||
* <p>Deliberately not {@link #nameSeq}. That counter is shared by every profile this launcher
|
||||
* serves, because its job is to make herdr <em>agent names</em> unique. Reusing it for the tab
|
||||
* label made the numbers global, so sibling tabs read {@code #4}, {@code #9}, {@code #17} — gaps
|
||||
* that look like a member died. Counting per role+profile makes {@code dev: sonnet #2} mean the
|
||||
* second sonnet dev, which is what a reader assumes it means.
|
||||
*
|
||||
* <p>Resets when the daemon restarts, and that is fine: the label is a human-facing hint, not an
|
||||
* identity. Identity is {@link PeerHandle#id()}.
|
||||
*/
|
||||
private final ConcurrentMap<String, AtomicLong> labelSeq = new ConcurrentHashMap<>();
|
||||
|
||||
private final long spawnReadyTimeoutMs; // 0 = disable gate (legacy non-blocking spawn)
|
||||
private final LongSupplier nowMillis; // monotonic clock (injectable for tests)
|
||||
@@ -106,10 +143,28 @@ public abstract class HerdrPeerLauncher implements PeerLauncher {
|
||||
* interval is baked into this hook, so the base needs no poll field
|
||||
*/
|
||||
protected HerdrPeerLauncher(String namePrefix, AgentControl agents, WorkspaceControl spaces,
|
||||
Map<String, BridgedConfig.Worker> profiles, String defaultProfile,
|
||||
Map<String, BridgedConfig.Profile> profiles, String defaultProfile,
|
||||
Function<String, String> env,
|
||||
long spawnReadyTimeoutMs,
|
||||
LongSupplier nowMillis, Runnable sleeper) {
|
||||
this(namePrefix, agents, spaces, profiles, defaultProfile, env, spawnReadyTimeoutMs,
|
||||
nowMillis, sleeper, null);
|
||||
}
|
||||
|
||||
/**
|
||||
* As above, plus the live {@code fleet} config (CB-557).
|
||||
*
|
||||
* @param fleet live fleet config, read once per spawn; {@code null} ⇒ default tab label and no
|
||||
* role charter. A separate constructor rather than a new parameter on the one
|
||||
* above, so every existing call site keeps the default without an edit.
|
||||
*/
|
||||
protected HerdrPeerLauncher(String namePrefix, AgentControl agents, WorkspaceControl spaces,
|
||||
Map<String, BridgedConfig.Profile> profiles, String defaultProfile,
|
||||
Function<String, String> env,
|
||||
long spawnReadyTimeoutMs,
|
||||
LongSupplier nowMillis, Runnable sleeper,
|
||||
Supplier<BridgedConfig.Fleet> fleet) {
|
||||
this.fleet = fleet;
|
||||
this.namePrefix = namePrefix;
|
||||
this.agents = agents;
|
||||
this.spaces = spaces;
|
||||
@@ -128,22 +183,7 @@ public abstract class HerdrPeerLauncher implements PeerLauncher {
|
||||
* Any subscription/guard check, MCP mount, and instruction injection happen here. The env map
|
||||
* and argv are adapter-private; the base only places and starts what is returned.
|
||||
*/
|
||||
protected abstract Launch buildLaunch(BridgedConfig.Worker cfg);
|
||||
|
||||
/**
|
||||
* Session-aware variant of {@link #buildLaunch(BridgedConfig.Worker)} (CB-547a). Default
|
||||
* discards the session identity and delegates to the profile-only form, so an adapter that
|
||||
* carries no durable peer session (opencode, say) inherits byte-identical behaviour and needs
|
||||
* no change. An adapter that does (Claude Code) overrides this to mint/resume the id and to
|
||||
* surface it on the returned {@link Launch#agentSessionId()}.
|
||||
*
|
||||
* @param cfg the resolved profile to spawn
|
||||
* @param sessionName the bridge's logical session name, or null/blank for launcher-derived
|
||||
* @param resumeSessionId the peer's own prior session id to resume, or null/blank for fresh
|
||||
*/
|
||||
protected Launch buildLaunch(BridgedConfig.Worker cfg, String sessionName, String resumeSessionId) {
|
||||
return buildLaunch(cfg);
|
||||
}
|
||||
protected abstract Launch buildLaunch(BridgedConfig.Profile cfg, LaunchSpec spec);
|
||||
|
||||
/** Direct transport access for peer-specific, non-turn control operations. */
|
||||
protected final AgentControl agents() {
|
||||
@@ -178,6 +218,10 @@ public abstract class HerdrPeerLauncher implements PeerLauncher {
|
||||
}
|
||||
}
|
||||
|
||||
/** All per-spawn values adapters may need, including the base-composed effective charter. */
|
||||
protected record LaunchSpec(String sessionName, String resumeSessionId, MemberRole role, String charter) {
|
||||
}
|
||||
|
||||
// --- profile surface -----------------------------------------------------------------------
|
||||
|
||||
/** The configured peer profile names (what {@code spawn(profile)} accepts). */
|
||||
@@ -193,7 +237,7 @@ public abstract class HerdrPeerLauncher implements PeerLauncher {
|
||||
if (name == null || name.isBlank()) {
|
||||
return List.of();
|
||||
}
|
||||
BridgedConfig.Worker cfg = profiles.get(name);
|
||||
BridgedConfig.Profile cfg = profiles.get(name);
|
||||
return cfg == null ? List.of() : cfg.parityOverlay();
|
||||
}
|
||||
|
||||
@@ -204,18 +248,18 @@ public abstract class HerdrPeerLauncher implements PeerLauncher {
|
||||
}
|
||||
|
||||
/** The configured profiles, for adapter capability decisions (e.g. any git-token grant). */
|
||||
protected Collection<BridgedConfig.Worker> profileConfigs() {
|
||||
protected Collection<BridgedConfig.Profile> profileConfigs() {
|
||||
return profiles.values();
|
||||
}
|
||||
|
||||
/** Resolve {@code profileName} (null/blank → default) to its config, or throw with the options. */
|
||||
protected BridgedConfig.Worker requireProfile(String profileName) {
|
||||
protected BridgedConfig.Profile requireProfile(String profileName) {
|
||||
String name = (profileName == null || profileName.isBlank()) ? defaultProfile : profileName;
|
||||
if (name == null || name.isBlank()) {
|
||||
throw new IllegalArgumentException("no default worker profile is configured — "
|
||||
+ "pass a profile; configured: " + profiles.keySet());
|
||||
}
|
||||
BridgedConfig.Worker cfg = profiles.get(name);
|
||||
BridgedConfig.Profile cfg = profiles.get(name);
|
||||
if (cfg == null) {
|
||||
throw new IllegalArgumentException("unknown worker profile '" + name
|
||||
+ "' — configured: " + profiles.keySet());
|
||||
@@ -240,23 +284,46 @@ public abstract class HerdrPeerLauncher implements PeerLauncher {
|
||||
return spawnInternal(profileName, requestedCwd, callerCwd, null, null).agent();
|
||||
}
|
||||
|
||||
/**
|
||||
* Spawn a peer with session identity (CB-547a). {@code sessionName} and {@code resumeSessionId}
|
||||
* are threaded from the {@link SpawnRequest} into {@link #buildLaunch(BridgedConfig.Worker,
|
||||
* String, String)}, and the launch's resolved agent-session id is returned alongside the agent
|
||||
* so the caller can put it on the {@link PeerHandle}.
|
||||
*/
|
||||
/** Pre-CB-557 shape: no explicit role, so the tab is labelled as a {@code dev}. */
|
||||
protected Spawned spawnInternal(String profileName, String requestedCwd, String callerCwd,
|
||||
String sessionName, String resumeSessionId) {
|
||||
BridgedConfig.Worker cfg = requireProfile(profileName);
|
||||
Launch launch = buildLaunch(cfg, sessionName, resumeSessionId);
|
||||
return spawnInternal(profileName, requestedCwd, callerCwd, sessionName, resumeSessionId,
|
||||
MemberRole.DEV);
|
||||
}
|
||||
|
||||
/**
|
||||
* Spawn a peer with session identity (CB-547a). The session values, role, and charter are
|
||||
* threaded from the {@link SpawnRequest} into {@link #buildLaunch(BridgedConfig.Profile,
|
||||
* LaunchSpec)}, and the launch's resolved agent-session id is returned alongside the agent so
|
||||
* the caller can put it on the {@link PeerHandle}.
|
||||
*/
|
||||
protected Spawned spawnInternal(String profileName, String requestedCwd, String callerCwd,
|
||||
String sessionName, String resumeSessionId, MemberRole role) {
|
||||
BridgedConfig.Profile cfg = requireProfile(profileName);
|
||||
BridgedConfig.Fleet liveFleet = fleet == null ? null : fleet.get();
|
||||
String roleCharter = liveFleet == null ? null : liveFleet.charterFor(role);
|
||||
String replyCharter = cfg.hasMcp() ? REPLY_CHARTER : null;
|
||||
String charter = roleCharter == null ? replyCharter
|
||||
: replyCharter == null ? roleCharter : roleCharter + "\n\n" + replyCharter;
|
||||
Launch launch = buildLaunch(cfg, new LaunchSpec(sessionName, resumeSessionId, role, charter));
|
||||
String cwd = resolveCwd(requestedCwd, cfg, callerCwd);
|
||||
Agent agent = cfg.tabPlacement()
|
||||
? spawnInTab(cfg, launch.env(), launch.argv(), cwd)
|
||||
? spawnInTab(cfg, launch.env(), launch.argv(), cwd, role, liveFleet)
|
||||
: spawnAsPane(cfg, launch.env(), launch.argv(), cwd);
|
||||
return new Spawned(agent, launch.agentSessionId());
|
||||
}
|
||||
|
||||
/**
|
||||
* The next tab number for {@code role} on {@code profile}, starting at 1.
|
||||
*
|
||||
* <p>Starts at 1 rather than 0 because the number is read by a person: {@code "dev: sonnet #1"}
|
||||
* is the first one, and {@code #0} invites the question of where {@code #1} went.
|
||||
*/
|
||||
private long nextLabelSeq(MemberRole role, String profile) {
|
||||
String key = (role == null ? "" : role.wireName()) + "/" + profile;
|
||||
return labelSeq.computeIfAbsent(key, _ -> new AtomicLong()).incrementAndGet();
|
||||
}
|
||||
|
||||
/**
|
||||
* {@inheritDoc}
|
||||
*
|
||||
@@ -272,7 +339,7 @@ public abstract class HerdrPeerLauncher implements PeerLauncher {
|
||||
@Override
|
||||
public PeerHandle spawn(SpawnRequest req) {
|
||||
Spawned spawned = spawnInternal(req.profileName(), req.requestedCwd(), req.callerCwd(),
|
||||
req.sessionName(), req.resumeSessionId());
|
||||
req.sessionName(), req.resumeSessionId(), req.role());
|
||||
Agent agent = spawned.agent();
|
||||
String paneId = agent.paneId();
|
||||
if (spawnReadyTimeoutMs > 0) {
|
||||
@@ -304,7 +371,7 @@ public abstract class HerdrPeerLauncher implements PeerLauncher {
|
||||
* guaranteed last resort so a pathological environment with an unset {@code user.dir} still
|
||||
* honours the "never assume {@code $HOME}" contract rather than letting herdr default the pane.
|
||||
*/
|
||||
private static String resolveCwd(String requestedCwd, BridgedConfig.Worker cfg, String callerCwd) {
|
||||
private static String resolveCwd(String requestedCwd, BridgedConfig.Profile cfg, String callerCwd) {
|
||||
return firstNonBlank(requestedCwd, cfg.cwd(), callerCwd, System.getProperty("user.dir"), ".");
|
||||
}
|
||||
|
||||
@@ -316,8 +383,8 @@ public abstract class HerdrPeerLauncher implements PeerLauncher {
|
||||
}
|
||||
|
||||
/** Dedicated worker space → own tab (carrying cwd+env) → start the peer into the seed pane. */
|
||||
private Agent spawnInTab(BridgedConfig.Worker cfg, Map<String, String> workerEnv,
|
||||
List<String> argv, String cwd) {
|
||||
private Agent spawnInTab(BridgedConfig.Profile cfg, Map<String, String> workerEnv,
|
||||
List<String> argv, String cwd, MemberRole role, BridgedConfig.Fleet liveFleet) {
|
||||
Workspace space = spaces.ensureWorkspace(cfg.workspace());
|
||||
Tab.Created tab = spaces.createTab(space.workspaceId(), cwd, workerEnv);
|
||||
log.info("spawning {} profile={} space={} tab={} cwd={}",
|
||||
@@ -348,7 +415,10 @@ public abstract class HerdrPeerLauncher implements PeerLauncher {
|
||||
// Labelling is cosmetic: it must not fail the spawn or orphan the running peer — on error
|
||||
// we log and still return it so the caller gets its paneId and can tear it down.
|
||||
tidy("label tab " + tab.tab().tabId(),
|
||||
() -> spaces.renameTab(tab.tab().tabId(), cfg.renderTabLabel(started.seq())));
|
||||
() -> spaces.renameTab(tab.tab().tabId(),
|
||||
cfg.renderTabLabel(
|
||||
liveFleet == null ? null : liveFleet.tabLabel(),
|
||||
role, nextLabelSeq(role, cfg.profile()))));
|
||||
log.info("{} started pane={} tab={} terminal={}",
|
||||
namePrefix, started.agent().paneId(), started.agent().tabId(), started.agent().terminalId());
|
||||
return started.agent();
|
||||
@@ -364,7 +434,7 @@ public abstract class HerdrPeerLauncher implements PeerLauncher {
|
||||
}
|
||||
|
||||
/** Legacy placement: split the currently-focused tab; the peer still starts in {@code cwd}. */
|
||||
private Agent spawnAsPane(BridgedConfig.Worker cfg, Map<String, String> workerEnv,
|
||||
private Agent spawnAsPane(BridgedConfig.Profile cfg, Map<String, String> workerEnv,
|
||||
List<String> argv, String cwd) {
|
||||
log.info("spawning {} (pane placement) profile={} cwd={} argv={}",
|
||||
namePrefix, cfg.profile(), cwd, argv);
|
||||
@@ -391,7 +461,7 @@ public abstract class HerdrPeerLauncher implements PeerLauncher {
|
||||
* backstop for the astronomically unlikely nonce+seq clash; the name is a label only — herdr
|
||||
* detects kind and status from terminal output, not from it.
|
||||
*/
|
||||
private Started startUniquelyNamed(BridgedConfig.Worker cfg, List<String> argv, String paneId) {
|
||||
private Started startUniquelyNamed(BridgedConfig.Profile cfg, List<String> argv, String paneId) {
|
||||
// Protocol 19 resolves the executable from the agent kind (== namePrefix here), so
|
||||
// argv[0] — the configured executable — is dropped and only the extra args are passed.
|
||||
List<String> args = argv.isEmpty() ? argv : argv.subList(1, argv.size());
|
||||
@@ -549,7 +619,7 @@ public abstract class HerdrPeerLauncher implements PeerLauncher {
|
||||
|
||||
/** Whether any configured profile places peers in their own tab (so tabs may need cleanup). */
|
||||
private boolean usesTabPlacement() {
|
||||
return profiles.values().stream().anyMatch(BridgedConfig.Worker::tabPlacement);
|
||||
return profiles.values().stream().anyMatch(BridgedConfig.Profile::tabPlacement);
|
||||
}
|
||||
|
||||
/** True when a herdr error means the target is already gone (safe to treat as done). */
|
||||
@@ -608,7 +678,7 @@ public abstract class HerdrPeerLauncher implements PeerLauncher {
|
||||
* {@code GITEA_HOST}. Push over SSH is unaffected; the only incremental grant is PR-create.
|
||||
* Peer-neutral, so every herdr adapter reuses it unchanged.
|
||||
*/
|
||||
protected void applyGitToken(Map<String, String> workerEnv, BridgedConfig.Worker cfg) {
|
||||
protected void applyGitToken(Map<String, String> workerEnv, BridgedConfig.Profile cfg) {
|
||||
if (!cfg.hasGitToken()) {
|
||||
return;
|
||||
}
|
||||
@@ -637,7 +707,7 @@ public abstract class HerdrPeerLauncher implements PeerLauncher {
|
||||
* dev.ltms.bridged.guard.SubscriptionGuard}, which is checked against the profile's
|
||||
* {@code baseUrl} and nothing else.
|
||||
*/
|
||||
protected Map<String, String> baseEnv(BridgedConfig.Worker cfg) {
|
||||
protected Map<String, String> baseEnv(BridgedConfig.Profile cfg) {
|
||||
Map<String, String> workerEnv = new LinkedHashMap<>();
|
||||
String path = env.apply("PATH");
|
||||
if (path != null && !path.isBlank()) {
|
||||
+83
-84
@@ -1,4 +1,4 @@
|
||||
package dev.ltms.bridged.worker;
|
||||
package dev.ltms.bridged.member;
|
||||
|
||||
import com.fasterxml.jackson.databind.ObjectMapper;
|
||||
import com.fasterxml.jackson.databind.node.ObjectNode;
|
||||
@@ -20,6 +20,7 @@ import java.util.Map;
|
||||
import java.util.Set;
|
||||
import java.util.function.Function;
|
||||
import java.util.function.LongSupplier;
|
||||
import java.util.function.Supplier;
|
||||
|
||||
/**
|
||||
* The {@link HerdrPeerLauncher} adapter for <strong>opencode</strong> — an open-source,
|
||||
@@ -37,7 +38,7 @@ import java.util.function.LongSupplier;
|
||||
* <li><strong>File-based MCP mount + instructions.</strong> opencode has no inline
|
||||
* {@code --mcp-config}/{@code --append-system-prompt}. Instead the bridge writes an ephemeral
|
||||
* {@code opencode.json} that declares the bridge as a {@code remote} MCP server and lists a
|
||||
* reply-charter file under {@code instructions}, then points the worker at it with
|
||||
* member-charter file under {@code instructions}, then points the worker at it with
|
||||
* {@code OPENCODE_CONFIG}. This is the one place the launcher touches disk — Claude never did.</li>
|
||||
* <li><strong>Model as a flag.</strong> the {@code provider/model} selector is passed as
|
||||
* {@code -m}, not an env var.</li>
|
||||
@@ -53,38 +54,9 @@ public final class OpenCodeLauncher extends HerdrPeerLauncher {
|
||||
/** Writer for the generated {@code opencode.json}. */
|
||||
private static final ObjectMapper JSON = new ObjectMapper();
|
||||
|
||||
/**
|
||||
* Standing instruction written to the charter file and mounted via the config's
|
||||
* {@code instructions} so the worker returns its result through {@code bridge_reply}. Kept on
|
||||
* disk (not a launch flag) because opencode's {@code instructions} takes file paths, not inline
|
||||
* text — the file is regenerated per spawn and never touches the worker's own profile.
|
||||
*/
|
||||
static final String REPLY_CHARTER =
|
||||
"You are an off-subscription worker in the claude-bridge fleet, running under opencode. "
|
||||
+ "Every message you receive arrives through the bridge, and the ONLY channel back to the "
|
||||
+ "sender is the bridge_reply MCP tool. Text you write in your terminal is NOT sent "
|
||||
+ "anywhere — the sender cannot see your screen, so an in-terminal answer is silently "
|
||||
+ "discarded. Therefore you MUST end EVERY turn by calling bridge_reply with `content` set "
|
||||
+ "to your complete response. This holds for every message without exception — tasks, "
|
||||
+ "questions, clarifications, acknowledgements, and ordinary back-and-forth conversation. "
|
||||
+ "Call bridge_reply exactly once, as the final action of your turn, with your full answer "
|
||||
+ "in `content`; never wait for confirmation first. If you end a turn without calling "
|
||||
+ "bridge_reply, the sender receives nothing and the exchange stalls.";
|
||||
|
||||
/** Root under which per-spawn opencode config dirs are created (injectable for tests). */
|
||||
private final Path configRoot;
|
||||
|
||||
/**
|
||||
* The current spawn's resume-target session id, threaded from {@link #spawn(SpawnRequest)} to
|
||||
* {@link #buildLaunch} across the base's {@code spawn -> spawnInternal -> buildLaunch} chain,
|
||||
* which carries no request. A plain field would race under concurrent spawns (the base supports
|
||||
* them), so it is thread-local: each spawn captures its own request's id on its own thread, and
|
||||
* {@code buildLaunch}, synchronous and same-thread, reads exactly that one. Set only around the
|
||||
* {@code super.spawn} call and cleared in {@code finally}, so a paused/leftover value can never
|
||||
* bleed into the next spawn.
|
||||
*/
|
||||
private final ThreadLocal<String> resumeSessionId = new ThreadLocal<>();
|
||||
|
||||
/**
|
||||
* Session discovery against opencode's on-disk storage ({@link OpenCodeSessionDiscovery}) —
|
||||
* the one seam that knows opencode's private session-file layout. Its root is injectable for
|
||||
@@ -97,7 +69,7 @@ public final class OpenCodeLauncher extends HerdrPeerLauncher {
|
||||
* matches the legacy non-blocking spawn semantics. Config dirs are created under the JVM temp dir.
|
||||
*/
|
||||
public OpenCodeLauncher(AgentControl agents, WorkspaceControl spaces,
|
||||
Map<String, BridgedConfig.Worker> profiles, String defaultProfile,
|
||||
Map<String, BridgedConfig.Profile> profiles, String defaultProfile,
|
||||
Function<String, String> env) {
|
||||
this(agents, spaces, profiles, defaultProfile, env, 0,
|
||||
System::currentTimeMillis, () -> sleepUninterruptibly(300),
|
||||
@@ -109,12 +81,26 @@ public final class OpenCodeLauncher extends HerdrPeerLauncher {
|
||||
* the pane reports an injectable state or {@code spawnReadyTimeoutMs} elapses.
|
||||
*/
|
||||
public OpenCodeLauncher(AgentControl agents, WorkspaceControl spaces,
|
||||
Map<String, BridgedConfig.Worker> profiles, String defaultProfile,
|
||||
Map<String, BridgedConfig.Profile> profiles, String defaultProfile,
|
||||
Function<String, String> env,
|
||||
long spawnReadyTimeoutMs, long spawnReadyPollMs) {
|
||||
this(agents, spaces, profiles, defaultProfile, env,
|
||||
spawnReadyTimeoutMs, spawnReadyPollMs, null);
|
||||
}
|
||||
|
||||
/**
|
||||
* Production constructor carrying the fleet-wide tab-label template (CB-557). The template comes
|
||||
* from {@code fleet.tabLabel}, which a profile cannot know because it names the member's
|
||||
* <em>role</em>; a profile may still override it with its own {@code tabLabel}.
|
||||
*/
|
||||
public OpenCodeLauncher(AgentControl agents, WorkspaceControl spaces,
|
||||
Map<String, BridgedConfig.Profile> profiles, String defaultProfile,
|
||||
Function<String, String> env,
|
||||
long spawnReadyTimeoutMs, long spawnReadyPollMs,
|
||||
Supplier<BridgedConfig.Fleet> fleet) {
|
||||
this(agents, spaces, profiles, defaultProfile, env, spawnReadyTimeoutMs,
|
||||
System::currentTimeMillis, () -> sleepUninterruptibly(spawnReadyPollMs),
|
||||
defaultConfigRoot(), defaultDiscoveryRoot());
|
||||
defaultConfigRoot(), defaultDiscoveryRoot(), fleet);
|
||||
}
|
||||
|
||||
/**
|
||||
@@ -137,13 +123,29 @@ public final class OpenCodeLauncher extends HerdrPeerLauncher {
|
||||
* {@link OpenCodeSessionDiscovery})
|
||||
*/
|
||||
public OpenCodeLauncher(AgentControl agents, WorkspaceControl spaces,
|
||||
Map<String, BridgedConfig.Worker> profiles, String defaultProfile,
|
||||
Map<String, BridgedConfig.Profile> profiles, String defaultProfile,
|
||||
Function<String, String> env,
|
||||
long spawnReadyTimeoutMs,
|
||||
LongSupplier nowMillis, Runnable sleeper,
|
||||
Path configRoot, Path discoveryRoot) {
|
||||
this(agents, spaces, profiles, defaultProfile, env, spawnReadyTimeoutMs,
|
||||
nowMillis, sleeper, configRoot, discoveryRoot, null);
|
||||
}
|
||||
|
||||
/**
|
||||
* Full testability constructor, plus the fleet-wide tab-label template (CB-557).
|
||||
*
|
||||
* @param fleet live fleet config, read once for each spawn
|
||||
*/
|
||||
public OpenCodeLauncher(AgentControl agents, WorkspaceControl spaces,
|
||||
Map<String, BridgedConfig.Profile> profiles, String defaultProfile,
|
||||
Function<String, String> env,
|
||||
long spawnReadyTimeoutMs,
|
||||
LongSupplier nowMillis, Runnable sleeper,
|
||||
Path configRoot, Path discoveryRoot,
|
||||
Supplier<BridgedConfig.Fleet> fleet) {
|
||||
super(NAME_PREFIX, agents, spaces, profiles, defaultProfile, env,
|
||||
spawnReadyTimeoutMs, nowMillis, sleeper);
|
||||
spawnReadyTimeoutMs, nowMillis, sleeper, fleet);
|
||||
this.configRoot = configRoot;
|
||||
this.discovery = new OpenCodeSessionDiscovery(discoveryRoot);
|
||||
}
|
||||
@@ -161,20 +163,21 @@ public final class OpenCodeLauncher extends HerdrPeerLauncher {
|
||||
* {@inheritDoc}
|
||||
*
|
||||
* <p>Builds the opencode launch: no {@code ANTHROPIC_*} and no guard (opencode reads its own
|
||||
* provider credentials); when the profile mounts the bridge MCP, generate an ephemeral
|
||||
* {@code opencode.json} (remote MCP server + reply-charter instructions) and point the worker at
|
||||
* it via {@code OPENCODE_CONFIG}; carry the parity-neutral git-forge grant; and select the model
|
||||
* with {@code -m}.
|
||||
* provider credentials); when the profile mounts the bridge MCP or has a member charter,
|
||||
* generate an ephemeral {@code opencode.json} (remote MCP server + member-charter instructions)
|
||||
* and point the worker at it via {@code OPENCODE_CONFIG}; carry the parity-neutral git-forge
|
||||
* grant; and select the model with {@code -m}.
|
||||
*/
|
||||
@Override
|
||||
protected Launch buildLaunch(BridgedConfig.Worker cfg) {
|
||||
protected Launch buildLaunch(BridgedConfig.Profile cfg, LaunchSpec spec) {
|
||||
Map<String, String> workerEnv = baseEnv(cfg);
|
||||
// A config file is needed for the bridge MCP mount, for a pinned endpoint (CB-508), or both.
|
||||
if (cfg.hasMcp() || hasCustomProvider(cfg)) {
|
||||
workerEnv.put("OPENCODE_CONFIG", writeConfig(cfg).toString());
|
||||
// A config file is needed for the bridge MCP mount, a member charter, or a pinned endpoint (CB-508).
|
||||
if (cfg.hasMcp() || spec.charter() != null || hasCustomProvider(cfg)) {
|
||||
workerEnv.put("OPENCODE_CONFIG", writeConfig(cfg, spec.charter()).toString());
|
||||
}
|
||||
applyGitToken(workerEnv, cfg);
|
||||
return new Launch(workerEnv, argvWithResume(argvWithModel(argvWithAuto(cfg), cfg)));
|
||||
return new Launch(workerEnv,
|
||||
argvWithResume(argvWithModel(argvWithAuto(cfg), cfg), spec.resumeSessionId()));
|
||||
}
|
||||
|
||||
/**
|
||||
@@ -187,7 +190,7 @@ public final class OpenCodeLauncher extends HerdrPeerLauncher {
|
||||
* primary's Anthropic subscription, and an opencode process has no Anthropic credential path
|
||||
* at all. Pointing it at a local vLLM cannot leak the subscription.
|
||||
*/
|
||||
private static boolean hasCustomProvider(BridgedConfig.Worker cfg) {
|
||||
private static boolean hasCustomProvider(BridgedConfig.Profile cfg) {
|
||||
return cfg.baseUrl() != null && !cfg.baseUrl().isBlank();
|
||||
}
|
||||
|
||||
@@ -201,7 +204,7 @@ public final class OpenCodeLauncher extends HerdrPeerLauncher {
|
||||
* own git worktree on its own branch, is off-subscription, and cannot merge — the lead is the
|
||||
* gate.
|
||||
*/
|
||||
private List<String> argvWithAuto(BridgedConfig.Worker cfg) {
|
||||
private List<String> argvWithAuto(BridgedConfig.Profile cfg) {
|
||||
List<String> argv = mutableArgv(cfg.argv());
|
||||
argv.add("--auto");
|
||||
return argv;
|
||||
@@ -211,11 +214,9 @@ public final class OpenCodeLauncher extends HerdrPeerLauncher {
|
||||
* The launch argv plus, on a resumed spawn, opencode's {@code -s <id>} flag to continue a prior
|
||||
* conversation by its session id. {@code -s, --session <id>} resumes an existing session; on a
|
||||
* fresh spawn (no resume target) no flag is added, letting opencode start a brand-new session.
|
||||
* The id comes from the current spawn request's {@code resumeSessionId}, threaded per-thread by
|
||||
* {@link #spawn(SpawnRequest)}.
|
||||
* The id comes from the base launch spec.
|
||||
*/
|
||||
private List<String> argvWithResume(List<String> argv) {
|
||||
String id = resumeSessionId.get();
|
||||
private List<String> argvWithResume(List<String> argv, String id) {
|
||||
if (id == null || id.isBlank()) {
|
||||
return argv;
|
||||
}
|
||||
@@ -226,7 +227,7 @@ public final class OpenCodeLauncher extends HerdrPeerLauncher {
|
||||
}
|
||||
|
||||
/** The launch argv plus, when a model is configured, the opencode {@code -m provider/model} flag. */
|
||||
private List<String> argvWithModel(List<String> argv, BridgedConfig.Worker cfg) {
|
||||
private List<String> argvWithModel(List<String> argv, BridgedConfig.Profile cfg) {
|
||||
if (cfg.model() != null && !cfg.model().isBlank()) {
|
||||
argv.add("-m");
|
||||
argv.add(cfg.model());
|
||||
@@ -235,29 +236,47 @@ public final class OpenCodeLauncher extends HerdrPeerLauncher {
|
||||
}
|
||||
|
||||
/**
|
||||
* Write an ephemeral {@code opencode.json} (and the reply-charter file it references) into a
|
||||
* Write an ephemeral {@code opencode.json} (and the member-charter file it references) into a
|
||||
* fresh per-spawn directory under {@link #configRoot}, and return the config file's path for
|
||||
* {@code OPENCODE_CONFIG}. The dir is unique per spawn so concurrent workers never race on it;
|
||||
* it is best-effort cleaned on JVM exit (worker config is disposable — regenerated every spawn).
|
||||
*/
|
||||
private Path writeConfig(BridgedConfig.Worker cfg) {
|
||||
private Path writeConfig(BridgedConfig.Profile cfg, String charterText) {
|
||||
try {
|
||||
Path dir = Files.createTempDirectory(configRoot, "bridged-opencode-");
|
||||
dir.toFile().deleteOnExit();
|
||||
|
||||
ObjectNode root = JSON.createObjectNode();
|
||||
root.put("$schema", "https://opencode.ai/config.json");
|
||||
// CB-523, opencode side: a worker that runs out of context dies mid-turn, and its reply
|
||||
// — the entire point of the turn — is lost with it. Auto-compaction is therefore not an
|
||||
// operator preference for a bridged worker, it is a condition of the turn contract.
|
||||
//
|
||||
// Stated deliberately even though it is redundant today: OPENCODE_CONFIG is MERGED over
|
||||
// ~/.config/opencode/config.json rather than replacing it, so a worker already inherits
|
||||
// an `auto: true` set at home. We do not want that inheritance to be what the guarantee
|
||||
// rests on — the home file is outside this repo, differs per machine, and is not ours.
|
||||
//
|
||||
// Know the cost before removing it: this key WINS over the home config (verified — an
|
||||
// OPENCODE_CONFIG value overrides the home value, it does not defer to it), so an
|
||||
// operator who sets `compaction.auto: false` at home cannot turn it off for bridged
|
||||
// workers. That is the intended trade for peers we spawn and whose turns we must land;
|
||||
// if per-profile control is ever wanted, add a profile knob rather than dropping this.
|
||||
root.putObject("compaction").put("auto", true);
|
||||
|
||||
if (cfg.hasMcp()) {
|
||||
Path charter = dir.resolve("reply-charter.md");
|
||||
Files.writeString(charter, REPLY_CHARTER);
|
||||
if (charterText != null) {
|
||||
Path charter = dir.resolve("member-charter.md");
|
||||
Files.writeString(charter, charterText);
|
||||
charter.toFile().deleteOnExit();
|
||||
|
||||
root.putArray("instructions").add(charter.toAbsolutePath().toString());
|
||||
}
|
||||
|
||||
if (cfg.hasMcp()) {
|
||||
ObjectNode bridge = root.putObject("mcp").putObject("bridge");
|
||||
bridge.put("type", "remote");
|
||||
bridge.put("url", cfg.mcpUrl());
|
||||
bridge.put("enabled", true);
|
||||
root.putArray("instructions").add(charter.toAbsolutePath().toString());
|
||||
}
|
||||
if (hasCustomProvider(cfg)) {
|
||||
addCustomProvider(root, cfg);
|
||||
@@ -282,7 +301,7 @@ public final class OpenCodeLauncher extends HerdrPeerLauncher {
|
||||
* <p>The provider id comes from the {@code provider/model} selector in {@code model:}, so one
|
||||
* field drives both the declaration and the {@code -m} flag and they cannot drift apart.
|
||||
*/
|
||||
private void addCustomProvider(ObjectNode root, BridgedConfig.Worker cfg) {
|
||||
private void addCustomProvider(ObjectNode root, BridgedConfig.Profile cfg) {
|
||||
String[] parts = splitModelSelector(cfg);
|
||||
String providerId = parts[0];
|
||||
String modelId = parts[1];
|
||||
@@ -305,7 +324,7 @@ public final class OpenCodeLauncher extends HerdrPeerLauncher {
|
||||
* needs both, so a bare model name is rejected loudly rather than silently falling back to the
|
||||
* default gateway — a worker quietly talking to the wrong endpoint is the failure this avoids.
|
||||
*/
|
||||
private static String[] splitModelSelector(BridgedConfig.Worker cfg) {
|
||||
private static String[] splitModelSelector(BridgedConfig.Profile cfg) {
|
||||
String model = cfg.model();
|
||||
int slash = model == null ? -1 : model.indexOf('/');
|
||||
if (model == null || model.isBlank() || slash <= 0 || slash == model.length() - 1) {
|
||||
@@ -333,31 +352,11 @@ public final class OpenCodeLauncher extends HerdrPeerLauncher {
|
||||
return afterScheme.contains("/") ? trimmed : trimmed + "/v1";
|
||||
}
|
||||
|
||||
/**
|
||||
* {@inheritDoc}
|
||||
*
|
||||
* <p>adds this adapter's session-identity work around the base's spawn — as opencode cannot be
|
||||
* told its session id at spawn (see {@link Capability#SESSION_RESUME} vs
|
||||
* {@link Capability#SESSION_NAME}), identity is only ever adopted after the fact:
|
||||
* <ul>
|
||||
* <li>the request's {@code resumeSessionId} is remembered for {@link #buildLaunch} to turn
|
||||
* into {@code -s <id>}; and</li>
|
||||
* <li>the returned handle is wrapped so its
|
||||
* {@link dev.ltms.bridged.peer.PeerHandle#agentSessionId()} performs lazy session
|
||||
* discovery against opencode's storage (see {@link OpenCodeSessionDiscovery}) — always
|
||||
* non-blocking, {@code null} until opencode has persisted the session record.</li>
|
||||
* </ul>
|
||||
*/
|
||||
/** Add lazy on-disk session discovery to the base handle. */
|
||||
@Override
|
||||
public PeerHandle spawn(SpawnRequest req) {
|
||||
resumeSessionId.set(req.resumeSessionId());
|
||||
try {
|
||||
PeerHandle inner = super.spawn(req);
|
||||
return new SessionAwareHandle(inner, discovery, effectiveCwd(req));
|
||||
} finally {
|
||||
// Never let a paused/leftover resume id bleed into the next spawn on this thread.
|
||||
resumeSessionId.remove();
|
||||
}
|
||||
PeerHandle inner = super.spawn(req);
|
||||
return new SessionAwareHandle(inner, discovery, effectiveCwd(req));
|
||||
}
|
||||
|
||||
/**
|
||||
@@ -444,7 +443,7 @@ public final class OpenCodeLauncher extends HerdrPeerLauncher {
|
||||
|
||||
/** Whether any configured profile opts into a git-forge token (required for {@link Capability#SELF_PR}). */
|
||||
private boolean hasGitTokenProfile() {
|
||||
return profileConfigs().stream().anyMatch(BridgedConfig.Worker::hasGitToken);
|
||||
return profileConfigs().stream().anyMatch(BridgedConfig.Profile::hasGitToken);
|
||||
}
|
||||
|
||||
// --- CB-117 reap predicate (opencode prefix), kept for direct unit testing -----------------
|
||||
+1
-1
@@ -1,4 +1,4 @@
|
||||
package dev.ltms.bridged.worker;
|
||||
package dev.ltms.bridged.member;
|
||||
|
||||
import com.fasterxml.jackson.databind.JsonNode;
|
||||
import com.fasterxml.jackson.databind.ObjectMapper;
|
||||
@@ -2,7 +2,7 @@ package dev.ltms.bridged.metrics;
|
||||
|
||||
import dev.ltms.bridged.msg.ReplyInbox;
|
||||
import dev.ltms.bridged.session.SessionManager;
|
||||
import dev.ltms.bridged.session.WorkerSession;
|
||||
import dev.ltms.bridged.session.MemberSession;
|
||||
|
||||
import java.util.LinkedHashMap;
|
||||
import java.util.Map;
|
||||
@@ -26,6 +26,8 @@ public final class BridgedMetrics {
|
||||
public static final String REPLIES = "bridged_replies_total";
|
||||
/** Counter: push-loop nudges to the primary, by outcome. */
|
||||
public static final String PUSH_NUDGES = "bridged_push_nudges_total";
|
||||
/** Counter: idle-lead heartbeat nudges to the lead, by outcome (CB-551). */
|
||||
public static final String HEARTBEAT_NUDGES = "bridged_lead_heartbeat_nudges_total";
|
||||
/** Counter: spawn attempts by peer kind and outcome. */
|
||||
public static final String SPAWNS = "bridged_spawns_total";
|
||||
/** Counter: herdr socket calls by method and outcome. */
|
||||
@@ -55,6 +57,9 @@ public final class BridgedMetrics {
|
||||
"Worker replies by delivery path (rendezvous=resolved an open send, inbox=stranded and held).");
|
||||
m.describe(PUSH_NUDGES, "counter",
|
||||
"CB-307 push-loop nudges to the primary (delivered|exhausted).");
|
||||
m.describe(HEARTBEAT_NUDGES, "counter",
|
||||
"CB-551 idle-lead heartbeat nudges (delivered|failed|exhausted). Quiet-cap exhaustion "
|
||||
+ "means the lead idled with nothing pending and was told to stand down.");
|
||||
m.describe(SPAWNS, "counter",
|
||||
"Worker spawn attempts by peer kind and outcome (ready|timeout|guard_rejected).");
|
||||
m.describe(HERDR_CALLS, "counter",
|
||||
@@ -69,7 +74,7 @@ public final class BridgedMetrics {
|
||||
|
||||
// One gauge per state so a scrape shows the whole census even when a state is empty —
|
||||
// an absent series and a zero series read very differently on a dashboard.
|
||||
for (WorkerSession.State state : WorkerSession.State.values()) {
|
||||
for (MemberSession.State state : MemberSession.State.values()) {
|
||||
String label = state.name().toLowerCase();
|
||||
m.gauge(SESSIONS, () -> countIn(sessions, state), "state", label);
|
||||
}
|
||||
@@ -78,7 +83,7 @@ public final class BridgedMetrics {
|
||||
// port's non-destructive read — scraping metrics must never ack a reply out of the inbox.
|
||||
m.collector(INBOX_DEPTH, "target", () -> {
|
||||
Map<String, Number> depths = new LinkedHashMap<>();
|
||||
for (WorkerSession s : sessions.roster()) {
|
||||
for (MemberSession s : sessions.roster()) {
|
||||
String target = s.terminalId();
|
||||
if (target == null) {
|
||||
continue;
|
||||
@@ -90,7 +95,7 @@ public final class BridgedMetrics {
|
||||
return m;
|
||||
}
|
||||
|
||||
private static long countIn(SessionManager sessions, WorkerSession.State state) {
|
||||
private static long countIn(SessionManager sessions, MemberSession.State state) {
|
||||
return sessions.roster().stream().filter(s -> s.state() == state).count();
|
||||
}
|
||||
}
|
||||
|
||||
@@ -0,0 +1,351 @@
|
||||
package dev.ltms.bridged.msg;
|
||||
|
||||
import dev.ltms.bridged.herdr.AgentControl;
|
||||
import dev.ltms.bridged.herdr.AgentStatus;
|
||||
import dev.ltms.bridged.mcp.PrimaryRegistry;
|
||||
import dev.ltms.bridged.metrics.BridgedMetrics;
|
||||
import dev.ltms.bridged.metrics.Metrics;
|
||||
import dev.ltms.bridged.session.MemberSession;
|
||||
import org.slf4j.Logger;
|
||||
import org.slf4j.LoggerFactory;
|
||||
|
||||
import java.util.ArrayList;
|
||||
import java.util.List;
|
||||
import java.util.concurrent.ScheduledExecutorService;
|
||||
import java.util.concurrent.TimeUnit;
|
||||
import java.util.function.LongSupplier;
|
||||
import java.util.function.Supplier;
|
||||
|
||||
/**
|
||||
* CB-551: an opt-in heartbeat that nudges the single idle lead back to work once it has been
|
||||
* continuously idle past a quiet period with no open {@code bridge_send} driving it.
|
||||
*
|
||||
* <p>Why this exists: the fleet is ONE lead + architects + workers, so an idle, stalled lead is a
|
||||
* single point of failure for the fleet's progress. {@link ReplyPushLoop} nudges the lead only when
|
||||
* a worker reply lands; this loop is the timer that catches the gap where nothing lands and the
|
||||
* lead simply sits idle with nobody to prompt it onward.
|
||||
*
|
||||
* <p>Mechanism mirrors {@link ReplyPushLoop}: a pure, unit-testable decision function
|
||||
* ({@link #decide(long, Long, int, AgentStatus, boolean, boolean, FleetState)}) plus a thin
|
||||
* scheduler around it. The loop is <em>entirely opt-in</em> ({@code leadHeartbeat:} in config); with
|
||||
* no such block it is never constructed, so upgrading the daemon cannot silently acquire a behaviour
|
||||
* that spends the operator's model subscription on its own initiative (constraint 1).
|
||||
*
|
||||
* <p>Four invariants keep it from becoming a runaway subscription burner:
|
||||
* <ol>
|
||||
* <li><b>Status-gated</b> — a {@code WORKING} lead is making progress and is never touched; only an
|
||||
* injectable (idle/done/blocked) lead is even considered (constraint 2).</li>
|
||||
* <li><b>Debounced</b> — a nudge only fires once the lead has been <em>continuously</em> idle past
|
||||
* {@code idleAfterSeconds}, so a lead that just finished a turn is not re-prompted into every
|
||||
* natural pause (constraint 3).</li>
|
||||
* <li><b>Bounded when quiet</b> — consecutive nudges that find no pending fleet state are capped at
|
||||
* {@code quietNudgeCap}, then the loop stops nudging until real state returns (constraint 4).</li>
|
||||
* <li><b>Never races {@link ReplyPushLoop}</b> — while that loop is actively nudging any target this
|
||||
* loop stands down, so two competing injections never start two turns in the same pane
|
||||
* (constraint 6).</li>
|
||||
* </ol>
|
||||
*/
|
||||
public final class LeadHeartbeatLoop {
|
||||
|
||||
private static final Logger log = LoggerFactory.getLogger(LeadHeartbeatLoop.class);
|
||||
|
||||
/** Sentinel for {@link #idleSinceNanos}: the lead is not currently in an idle stretch. */
|
||||
private static final long NOT_IDLE = Long.MIN_VALUE;
|
||||
|
||||
private final PrimaryRegistry primaryRegistry;
|
||||
private final AgentControl agents;
|
||||
private final ReplyInbox inbox;
|
||||
private final Supplier<List<MemberSession>> roster;
|
||||
private final ReplyPushLoop pushLoop;
|
||||
private final ScheduledExecutorService scheduler;
|
||||
private final LongSupplier clock;
|
||||
private final long idleAfterNanos;
|
||||
private final long backoffMs;
|
||||
private final int quietNudgeCap;
|
||||
private final Metrics metrics; // CB-512 pattern: nullable — no registry in unit tests
|
||||
|
||||
/** When the current idle stretch began (nanos), or {@link #NOT_IDLE}. Single scheduler thread only. */
|
||||
private long idleSinceNanos = NOT_IDLE;
|
||||
/** Consecutive nudges that found no pending fleet state. Single scheduler thread only. */
|
||||
private int quietCount = 0;
|
||||
|
||||
/** Constructor with an injectable clock and no metric registry (unit tests, or wiring that opts out). */
|
||||
public LeadHeartbeatLoop(PrimaryRegistry primaryRegistry, AgentControl agents, ReplyInbox inbox,
|
||||
Supplier<List<MemberSession>> roster, ReplyPushLoop pushLoop,
|
||||
ScheduledExecutorService scheduler, LongSupplier clock,
|
||||
long idleAfterNanos, long backoffMs, int quietNudgeCap) {
|
||||
this(primaryRegistry, agents, inbox, roster, pushLoop, scheduler, clock,
|
||||
idleAfterNanos, backoffMs, quietNudgeCap, null);
|
||||
}
|
||||
|
||||
/** As above, with a metric registry (the CB-512 pattern) so nudge outcomes are counted. */
|
||||
public LeadHeartbeatLoop(PrimaryRegistry primaryRegistry, AgentControl agents, ReplyInbox inbox,
|
||||
Supplier<List<MemberSession>> roster, ReplyPushLoop pushLoop,
|
||||
ScheduledExecutorService scheduler, LongSupplier clock,
|
||||
long idleAfterNanos, long backoffMs, int quietNudgeCap, Metrics metrics) {
|
||||
this.primaryRegistry = primaryRegistry;
|
||||
this.agents = agents;
|
||||
this.inbox = inbox;
|
||||
this.roster = roster;
|
||||
this.pushLoop = pushLoop;
|
||||
this.scheduler = scheduler;
|
||||
this.clock = clock;
|
||||
this.idleAfterNanos = idleAfterNanos;
|
||||
this.backoffMs = backoffMs;
|
||||
this.quietNudgeCap = quietNudgeCap;
|
||||
this.metrics = metrics;
|
||||
}
|
||||
|
||||
/** Count one nudge outcome when a registry is wired; a no-op in unit tests. */
|
||||
private void countNudge(String outcome) {
|
||||
if (metrics != null) {
|
||||
metrics.inc(BridgedMetrics.HEARTBEAT_NUDGES, "outcome", outcome);
|
||||
}
|
||||
}
|
||||
|
||||
// --- decision logic (package-private for unit-testing) --------------------------------------
|
||||
|
||||
/** The action the loop should take on one tick. */
|
||||
enum Action {
|
||||
/** Send a heartbeat nudge to the lead now. */
|
||||
INJECT,
|
||||
/** Lead is idle but not yet past the quiet period — keep waiting, do nothing. */
|
||||
WAIT_IDLE,
|
||||
/** Lead is making progress (or its status cannot be read) — reset the idle window and re-arm. */
|
||||
LEAD_BUSY,
|
||||
/** Idle past the quiet period with nothing pending and the cap exhausted — stop until new state appears. */
|
||||
QUIET_DONE,
|
||||
/** {@link ReplyPushLoop} is actively nudging — stand aside rather than start a competing turn. */
|
||||
STAND_DOWN
|
||||
}
|
||||
|
||||
/** The outcome of one decision: the action plus the state to persist for the next tick. */
|
||||
record Decision(Action action, Long idleSinceNanos, int quietCount) {}
|
||||
|
||||
/**
|
||||
* Pure decision function: given the current loop state and fleet/lead facts, return what to do
|
||||
* next and the exact state to carry forward. Pure means no I/O and no mutation — the caller
|
||||
* ({@link #tick}) applies {@link Decision} to its fields. This is what makes the loop unit-testable
|
||||
* with a fake clock and no sleeping.
|
||||
*
|
||||
* @param nowNanos the current time (an injected clock in tests, {@link System#nanoTime()} live)
|
||||
* @param idleSinceNanos when the current idle stretch began, or {@code null} if the lead is not idle
|
||||
* @param quietCount consecutive nudges so far that found no pending fleet state
|
||||
* @param status the lead's herdr status
|
||||
* @param pushLoopActive whether {@link ReplyPushLoop} is currently nudging some target (constraint 6)
|
||||
* @param leadKnown whether a lead terminal is known to nudge at all
|
||||
* @param fleet a snapshot of the pending fleet state (constraint 5)
|
||||
* @return the action to take and the state to persist
|
||||
*/
|
||||
Decision decide(long nowNanos, Long idleSinceNanos, int quietCount, AgentStatus status,
|
||||
boolean pushLoopActive, boolean leadKnown, FleetState fleet) {
|
||||
// Constraint 6: while ReplyPushLoop is actively nudging the lead, injecting a second,
|
||||
// competing prompt into the same pane would start a second turn — racing loops multiply
|
||||
// turns and context burn. Stand aside, and treat the active push as real state (re-arm the
|
||||
// quiet counter), because the reply that drove it is exactly the kind of new state that
|
||||
// should reset the cap.
|
||||
if (pushLoopActive) {
|
||||
return new Decision(Action.STAND_DOWN, idleSinceNanos, 0);
|
||||
}
|
||||
// Constraint 2: a WORKING lead is making progress and must NOT be touched; an unreadable
|
||||
// status (read failure, or the agent is gone) is safest treated the same way — never inject
|
||||
// into a state we cannot read. Either way, reset the idle window and the quiet counter: the
|
||||
// lead was / may be active, so the next idle stretch must count its own quiet period fresh.
|
||||
if (status == null || !status.injectable()) {
|
||||
return new Decision(Action.LEAD_BUSY, null, 0);
|
||||
}
|
||||
if (idleSinceNanos == null) {
|
||||
// The lead just became injectable — record the start of an idle stretch and wait out the
|
||||
// debounce quiet period before ever nudging (constraint 3).
|
||||
return new Decision(Action.WAIT_IDLE, nowNanos, quietCount);
|
||||
}
|
||||
if (nowNanos - idleSinceNanos < idleAfterNanos) {
|
||||
// Still within the quiet period: the lead that just finished a turn sits momentarily idle
|
||||
// and must not be re-prompted into every natural pause.
|
||||
return new Decision(Action.WAIT_IDLE, idleSinceNanos, quietCount);
|
||||
}
|
||||
// Past the quiet period with an injectable lead: it is a genuine candidate for a nudge. Two
|
||||
// gating facts decide whether and how:
|
||||
if (!leadKnown) {
|
||||
// No lead terminal is known yet (e.g. the registry has not learned one) — there is nobody
|
||||
// to nudge. Keep waiting; the window stays open so discovery re-arms it without a fresh
|
||||
// quiet period.
|
||||
return new Decision(Action.WAIT_IDLE, idleSinceNanos, quietCount);
|
||||
}
|
||||
if (fleet.hasPending()) {
|
||||
// Real fleet state is waiting — a worker reply or a DONE session. This is new state, so
|
||||
// it resets the quiet counter (constraint 4) and the lead is nudged to go collect it.
|
||||
return new Decision(Action.INJECT, idleSinceNanos, 0);
|
||||
}
|
||||
if (quietCount < quietNudgeCap) {
|
||||
// Nothing is pending, but the cap is not exhausted: nudge anyway, telling the lead
|
||||
// exactly that nothing is waiting so it can choose to stand down rather than hunt
|
||||
// (constraint 5). Count it toward the consecutive-quiet cap.
|
||||
return new Decision(Action.INJECT, idleSinceNanos, quietCount + 1);
|
||||
}
|
||||
// Nothing pending and the cap is exhausted: stop nudging until real state appears again
|
||||
// (constraint 4). The loop still ticks on backoff so a genuinely new reply or session change
|
||||
// re-arms it — QUIET_DONE stops injection, not observation.
|
||||
return new Decision(Action.QUIET_DONE, idleSinceNanos, quietCount);
|
||||
}
|
||||
|
||||
// --- loop ----------------------------------------------------------------------------------
|
||||
|
||||
/**
|
||||
* Start the heartbeat. Schedules the first tick after one backoff so a freshly-booting daemon
|
||||
* does not evaluate the lead's idle state before the fleet has settled.
|
||||
*/
|
||||
public void start() {
|
||||
log.info("idle-lead heartbeat: on — nudge lead after {}s idle (recheck {}ms, quiet cap {})",
|
||||
TimeUnit.NANOSECONDS.toSeconds(idleAfterNanos), backoffMs, quietNudgeCap);
|
||||
scheduler.schedule(this::tick, backoffMs, TimeUnit.MILLISECONDS);
|
||||
}
|
||||
|
||||
/** One loop tick, every {@link #backoffMs} — the thin scheduler around {@link #decide}. */
|
||||
private void tick() {
|
||||
boolean leadKnown = primaryRegistry.primaryTerminal().isPresent();
|
||||
FleetState fleet = snapshot(inbox, roster);
|
||||
AgentStatus status = AgentStatus.UNKNOWN;
|
||||
if (leadKnown) {
|
||||
try {
|
||||
status = agents.status(primaryRegistry.primaryTerminal().orElseThrow());
|
||||
} catch (RuntimeException e) {
|
||||
// A failed status read degrades to "unknown" — decide() treats that like a busy lead
|
||||
// and never injects into a state it cannot read. Retry on the next backoff.
|
||||
log.debug("idle-heartbeat: status check failed for lead, will retry: {}", e.toString());
|
||||
}
|
||||
}
|
||||
|
||||
Decision d = decide(clock.getAsLong(),
|
||||
idleSinceNanos == NOT_IDLE ? null : idleSinceNanos,
|
||||
quietCount, status, pushLoop.isActive(), leadKnown, fleet);
|
||||
applyDecision(d);
|
||||
switch (d.action()) {
|
||||
case INJECT -> injectNudge(fleet);
|
||||
case QUIET_DONE -> countNudge("exhausted");
|
||||
case WAIT_IDLE, LEAD_BUSY, STAND_DOWN -> { /* nothing to inject, nothing to count */ }
|
||||
}
|
||||
scheduleNext();
|
||||
}
|
||||
|
||||
/** Persist the state a decision returned, so the next tick starts from it. */
|
||||
private void applyDecision(Decision d) {
|
||||
idleSinceNanos = d.idleSinceNanos() == null ? NOT_IDLE : d.idleSinceNanos();
|
||||
quietCount = d.quietCount();
|
||||
}
|
||||
|
||||
/** Send the nudge to the known lead. */
|
||||
private void injectNudge(FleetState fleet) {
|
||||
var lead = primaryRegistry.primaryTerminal();
|
||||
if (lead.isEmpty()) {
|
||||
return; // the lead disappeared between the decision and the injection
|
||||
}
|
||||
String leadTerminal = lead.get();
|
||||
try {
|
||||
agents.send(leadTerminal, fleet.nudgeText());
|
||||
log.debug("idle-heartbeat: nudge sent to lead {} (quiet nudges so far in this stretch: {})",
|
||||
leadTerminal, quietCount);
|
||||
countNudge("delivered");
|
||||
} catch (RuntimeException e) {
|
||||
log.warn("idle-heartbeat: failed to nudge lead {}: {}", leadTerminal, e.toString());
|
||||
countNudge("failed");
|
||||
}
|
||||
}
|
||||
|
||||
/** Schedule the next tick on the scheduler thread pool. */
|
||||
private void scheduleNext() {
|
||||
scheduler.schedule(this::tick, backoffMs, TimeUnit.MILLISECONDS);
|
||||
}
|
||||
|
||||
// --- lifecycle -----------------------------------------------------------------------------
|
||||
|
||||
/** Shut down the scheduler; outstanding ticks are cancelled. */
|
||||
public void stop() {
|
||||
scheduler.shutdownNow();
|
||||
}
|
||||
|
||||
/** @see #stop() */
|
||||
public void close() {
|
||||
stop();
|
||||
}
|
||||
|
||||
/**
|
||||
* A read-only snapshot of the fleet facts the heartbeat reports to an idle lead. Carries exact
|
||||
* counts and names the targets that hold replies, so the nudge is actionable rather than a
|
||||
* content-free "keep working" that would cost the lead a turn just to discover what there is to
|
||||
* do (constraint 5).
|
||||
*
|
||||
* @param pendingReplies total undrained worker replies across owned targets
|
||||
* @param doneSessions session count in the DONE state — finished turns awaiting teardown
|
||||
* @param liveWorkers sessions still registered (acquired minus released)
|
||||
* @param replyTargets the terminals whose inboxes currently hold at least one reply
|
||||
*/
|
||||
record FleetState(int pendingReplies, int doneSessions, int liveWorkers, List<String> replyTargets) {
|
||||
|
||||
/** True when the lead has something real to collect — a reply or a session awaiting teardown. */
|
||||
boolean hasPending() {
|
||||
return pendingReplies > 0 || doneSessions > 0;
|
||||
}
|
||||
|
||||
/** The nudge body, phrased for the two cases the heartbeat distinguishes. */
|
||||
String nudgeText() {
|
||||
StringBuilder sb = new StringBuilder(
|
||||
"Heartbeat: you are idle and no bridge_send is waiting on you.");
|
||||
if (hasPending()) {
|
||||
sb.append(" The fleet has state to collect: ").append(pendingDetail());
|
||||
} else {
|
||||
sb.append(" Fleet is quiet — nothing pending to collect (")
|
||||
.append(pendingReplies).append(" worker replies, ")
|
||||
.append(doneSessions).append(" DONE sessions); ")
|
||||
.append(liveWorkers).append(" workers live. You may stand down until new state re-arms me.");
|
||||
}
|
||||
return sb.toString();
|
||||
}
|
||||
|
||||
private String pendingDetail() {
|
||||
StringBuilder sb = new StringBuilder();
|
||||
sb.append(pendingReplies).append(" worker ")
|
||||
.append(pendingReplies == 1 ? "reply" : "replies")
|
||||
.append(" pending collection");
|
||||
if (!replyTargets.isEmpty()) {
|
||||
// Render each as the exact command so the lead can act without parsing: the nearest
|
||||
// analogue to ReplyPushLoop's bridge_poll(target=...) nudge.
|
||||
sb.append(" (")
|
||||
.append(String.join(", ",
|
||||
replyTargets.stream().map(t -> "bridge_poll(target=" + t + ")").toList()))
|
||||
.append(")");
|
||||
}
|
||||
sb.append(", ").append(doneSessions).append(" DONE session").append(doneSessions == 1 ? "" : "s")
|
||||
.append(" awaiting teardown")
|
||||
.append(", ").append(liveWorkers).append(" worker").append(liveWorkers == 1 ? "" : "s")
|
||||
.append(" live");
|
||||
return sb.toString();
|
||||
}
|
||||
}
|
||||
|
||||
/**
|
||||
* Build a {@link FleetState} snapshot over the live roster and the reply inbox — the same
|
||||
* sources the metrics collector uses, so the heartbeat reports facts the operator can cross-check
|
||||
* against {@code /metrics}. Package-private and pure (reads only) so it is unit-testable.
|
||||
*/
|
||||
static FleetState snapshot(ReplyInbox inbox, Supplier<List<MemberSession>> roster) {
|
||||
List<MemberSession> sessions = roster.get();
|
||||
int replies = 0;
|
||||
int done = 0;
|
||||
List<String> replyTargets = new ArrayList<>();
|
||||
for (MemberSession s : sessions) {
|
||||
if (s.terminalId() == null) {
|
||||
continue;
|
||||
}
|
||||
int depth = inbox.peek(s.terminalId()).size();
|
||||
if (depth > 0) {
|
||||
replies += depth;
|
||||
replyTargets.add(s.terminalId());
|
||||
}
|
||||
if (s.state() == MemberSession.State.DONE) {
|
||||
done++;
|
||||
}
|
||||
}
|
||||
return new FleetState(replies, done, sessions.size(), List.copyOf(replyTargets));
|
||||
}
|
||||
}
|
||||
@@ -127,6 +127,8 @@ public final class MessageService {
|
||||
public enum Phase {
|
||||
/** Delegated and in flight — queued for the worker or being worked. */
|
||||
PENDING,
|
||||
/** The worker is paused in {@code bridge_ask}; {@link TaskView#reply} and {@link TaskView#turnId} identify it. */
|
||||
ASKING,
|
||||
/** The worker's turn finished; {@link TaskView#reply} holds the answer. */
|
||||
DONE,
|
||||
/** The delegation could not complete (timed out, worker gone, or busy). */
|
||||
@@ -136,16 +138,28 @@ public final class MessageService {
|
||||
/**
|
||||
* A poll snapshot of an async delegation.
|
||||
*
|
||||
* @param reply the answer when {@link #phase} is {@link Phase#DONE}, else {@code null}
|
||||
* @param reply the answer when {@link #phase} is {@link Phase#DONE}, or the question when
|
||||
* {@link #phase} is {@link Phase#ASKING}; otherwise {@code null}
|
||||
* @param replySource {@code "reply"} (structured {@code bridge_reply}) or {@code "transcript"}
|
||||
* (completion scrape) when {@link Phase#DONE}, else {@code null}
|
||||
* @param detail a human note (live worker status while pending, or the failure reason)
|
||||
* @param detail a human note (live worker status while pending, ask state, or failure reason)
|
||||
* @param turnId correlation id for an {@link Phase#ASKING} ticket, else {@code null}
|
||||
*/
|
||||
public record TaskView(String ticket, Phase phase, String reply, String replySource, String detail) {
|
||||
public record TaskView(String ticket, Phase phase, String reply, String replySource, String detail,
|
||||
String turnId) {
|
||||
}
|
||||
|
||||
/** An in-flight or finished async delegation, keyed by its ticket. */
|
||||
private record Task(String target, CompletableFuture<Reply> future, long createdNanos) {
|
||||
private static final class Task {
|
||||
private final String target;
|
||||
private final CompletableFuture<Reply> future = new CompletableFuture<>();
|
||||
private final long createdNanos = System.nanoTime();
|
||||
private volatile Reply question;
|
||||
private volatile String turnId;
|
||||
|
||||
private Task(String target) {
|
||||
this.target = target;
|
||||
}
|
||||
}
|
||||
|
||||
private final AgentControl agents;
|
||||
@@ -156,6 +170,11 @@ public final class MessageService {
|
||||
private final Metrics metrics; // CB-502: nullable — no registry in unit tests
|
||||
private final ConcurrentHashMap<String, ReentrantLock> sessionLocks = new ConcurrentHashMap<>();
|
||||
private final ConcurrentHashMap<String, Task> tasks = new ConcurrentHashMap<>();
|
||||
/** Async task that owns each exact forward rendezvous waiter. */
|
||||
private final ConcurrentHashMap<CompletableFuture<Rendezvous.Resolution>, Task> asyncTasksByWaiter =
|
||||
new ConcurrentHashMap<>();
|
||||
/** Async tickets paused on a specific {@code bridge_ask} turn. */
|
||||
private final ConcurrentHashMap<String, Task> asyncTasksByTurn = new ConcurrentHashMap<>();
|
||||
private final AtomicLong ticketSeq = new AtomicLong();
|
||||
private final ExecutorService asyncExecutor = Executors.newThreadPerTaskExecutor(
|
||||
Thread.ofVirtual().name("bridge-async-", 0).factory());
|
||||
@@ -202,6 +221,16 @@ public final class MessageService {
|
||||
return agents.status(target);
|
||||
}
|
||||
|
||||
/** Read-only delegation fact for fleet views. */
|
||||
public boolean hasAcceptedDelivery(String target) {
|
||||
return rendezvous.isWaiting(target);
|
||||
}
|
||||
|
||||
/** Read-only inbox fact for fleet views. */
|
||||
public boolean hasInboxMessage(String target) {
|
||||
return !inbox.peek(target).isEmpty();
|
||||
}
|
||||
|
||||
/**
|
||||
* Route a worker's explicit {@code bridge_reply}: resolve an open send, or queue it in the
|
||||
* inbox if no send is currently open. Unlike the bare {@link Rendezvous#resolve}, a no-waiter
|
||||
@@ -272,14 +301,18 @@ public final class MessageService {
|
||||
*/
|
||||
public boolean abandon(String target, String reason) {
|
||||
CompletableFuture<Rendezvous.Resolution> waiter = rendezvous.currentWaiter(target);
|
||||
if (waiter == null || waiter.isDone()) {
|
||||
return false; // nobody is blocked on this worker — nothing to abandon
|
||||
boolean failed = waiter != null && !waiter.isDone() && rendezvous.resolveFailure(waiter, reason);
|
||||
boolean asyncFailed = false;
|
||||
for (Task task : tasks.values()) {
|
||||
if (target.equals(task.target) && task.question == null
|
||||
&& task.future.complete(new Reply(Outcome.WORKER_FAILED, reason))) {
|
||||
asyncFailed = true;
|
||||
}
|
||||
}
|
||||
boolean failed = rendezvous.resolveFailure(waiter, reason);
|
||||
if (failed) {
|
||||
log.debug("abandoned send to {}: {}", target, reason);
|
||||
log.warn("abandoning the blocked send to {}: {}", target, reason);
|
||||
}
|
||||
return failed;
|
||||
return failed || asyncFailed;
|
||||
}
|
||||
|
||||
/**
|
||||
@@ -326,6 +359,11 @@ public final class MessageService {
|
||||
* never earned. {@code null} disables the hook.
|
||||
*/
|
||||
public Reply send(String target, String content, long timeoutMillis, Runnable onAccepted) {
|
||||
return send(target, content, timeoutMillis, onAccepted, null);
|
||||
}
|
||||
|
||||
/** Run a send, optionally stopping an async task that teardown already failed before acceptance. */
|
||||
private Reply send(String target, String content, long timeoutMillis, Runnable onAccepted, Task task) {
|
||||
long deadlineNanos = System.nanoTime() + timeoutMillis * 1_000_000L;
|
||||
ReentrantLock lock = sessionLocks.computeIfAbsent(target, _ -> new ReentrantLock());
|
||||
|
||||
@@ -333,6 +371,12 @@ public final class MessageService {
|
||||
return new Reply(Outcome.BUSY, null); // another send held the session the whole window
|
||||
}
|
||||
try {
|
||||
if (task != null && task.future.isDone()) {
|
||||
return task.future.getNow(null);
|
||||
}
|
||||
if (hasAsyncQuestion(target)) {
|
||||
return new Reply(Outcome.BUSY, null); // the worker's current turn is paused for its lead
|
||||
}
|
||||
// Open the waiter BEFORE queueing delivery (CB-548). A fast reply — the worker already
|
||||
// injectable the instant we enqueue — otherwise arrives before the waiter is registered
|
||||
// and orphans into the inbox while this send blocks to the timeout (the enqueue-before-
|
||||
@@ -341,6 +385,9 @@ public final class MessageService {
|
||||
// failed send leaves no stale waiter behind.
|
||||
CompletableFuture<Rendezvous.Resolution> reply = rendezvous.open(target);
|
||||
try {
|
||||
if (task != null) {
|
||||
asyncTasksByWaiter.put(reply, task);
|
||||
}
|
||||
// The send has won the lock; the accepted-delivery hook records delegator ownership
|
||||
// here (CB-548). It runs BEFORE enqueue so a throwing hook — onAccepted is now a
|
||||
// public callback — fails the send without queuing a message that would orphan.
|
||||
@@ -364,6 +411,7 @@ public final class MessageService {
|
||||
throw new IllegalStateException("interrupted awaiting reply from " + target, e);
|
||||
}
|
||||
} finally {
|
||||
asyncTasksByWaiter.remove(reply);
|
||||
rendezvous.close(target, reply);
|
||||
}
|
||||
} finally {
|
||||
@@ -389,7 +437,12 @@ public final class MessageService {
|
||||
if (ticket.fresh()) {
|
||||
// Register the reverse waiter first, then surface the question — so the answer, which can
|
||||
// arrive the instant the primary reacts, always finds an open waiter to resolve.
|
||||
CompletableFuture<Rendezvous.Resolution> waiter = rendezvous.currentWaiter(workerSession);
|
||||
Task task = markAsyncQuestion(waiter, question, ticket.turnId());
|
||||
if (!rendezvous.resolveQuestion(workerSession, question, ticket.turnId())) {
|
||||
if (task != null) {
|
||||
clearAsyncQuestion(ticket.turnId(), true);
|
||||
}
|
||||
rendezvous.closeAsk(ticket.turnId());
|
||||
return new AskResult(AskOutcome.NO_WAITER, null); // no primary is blocked on this worker
|
||||
}
|
||||
@@ -399,6 +452,7 @@ public final class MessageService {
|
||||
return new AskResult(AskOutcome.ANSWERED, answer);
|
||||
} catch (TimeoutException e) {
|
||||
log.debug("bridge_ask from {} went unanswered in {}ms", workerSession, timeoutMillis);
|
||||
clearAsyncQuestion(ticket.turnId(), true);
|
||||
return new AskResult(AskOutcome.TIMED_OUT, null);
|
||||
} catch (ExecutionException e) {
|
||||
Throwable cause = e.getCause();
|
||||
@@ -441,9 +495,12 @@ public final class MessageService {
|
||||
rendezvous.close(workerSession, reply);
|
||||
return new Reply(Outcome.STALE_TURN, null); // lapsed between the lookup and the unblock
|
||||
}
|
||||
clearAsyncQuestion(turnId, false);
|
||||
try {
|
||||
Rendezvous.Resolution r = reply.get(remainingMillis(deadlineNanos), TimeUnit.MILLISECONDS);
|
||||
return new Reply(outcomeOf(r.kind()), r.text(), r.turnId());
|
||||
Reply result = new Reply(outcomeOf(r.kind()), r.text(), r.turnId());
|
||||
finishAsyncTask(turnId, result);
|
||||
return result;
|
||||
} catch (TimeoutException e) {
|
||||
// The worker resumed but hasn't replied yet — no completion fallback arms an answered
|
||||
// turn (it never re-entered the injector), so a silent worker rides out the window.
|
||||
@@ -483,9 +540,21 @@ public final class MessageService {
|
||||
*/
|
||||
public String sendAsync(String target, String content, Runnable onAccepted) {
|
||||
String ticket = "task-" + ticketSeq.incrementAndGet();
|
||||
CompletableFuture<Reply> future = CompletableFuture.supplyAsync(
|
||||
() -> send(target, content, ASYNC_TIMEOUT_MS, onAccepted), asyncExecutor);
|
||||
tasks.put(ticket, new Task(target, future, System.nanoTime()));
|
||||
Task task = new Task(target);
|
||||
tasks.put(ticket, task);
|
||||
asyncExecutor.submit(() -> {
|
||||
try {
|
||||
Reply result = send(target, content, ASYNC_TIMEOUT_MS, onAccepted, task);
|
||||
if (result.outcome() == Outcome.QUESTION) {
|
||||
// Keep the accepted owner until answer() finishes it. markAsyncQuestion may run
|
||||
// just after resolveQuestion wakes this thread.
|
||||
} else {
|
||||
finishAsyncTask(task, result);
|
||||
}
|
||||
} catch (Throwable t) {
|
||||
task.future.completeExceptionally(t);
|
||||
}
|
||||
});
|
||||
pruneTerminalTickets();
|
||||
log.debug("async send {} -> {}", ticket, target);
|
||||
return ticket;
|
||||
@@ -501,27 +570,32 @@ public final class MessageService {
|
||||
if (task == null) {
|
||||
return null;
|
||||
}
|
||||
CompletableFuture<Reply> f = task.future();
|
||||
CompletableFuture<Reply> f = task.future;
|
||||
if (!f.isDone()) {
|
||||
return new TaskView(ticket, Phase.PENDING, null, null, "worker " + liveStatus(task.target()));
|
||||
Reply question = task.question;
|
||||
if (question != null) {
|
||||
return new TaskView(ticket, Phase.ASKING, question.text(), null,
|
||||
"worker is waiting for your answer", question.turnId());
|
||||
}
|
||||
return new TaskView(ticket, Phase.PENDING, null, null, "worker " + liveStatus(task.target), null);
|
||||
}
|
||||
Reply r;
|
||||
try {
|
||||
r = f.getNow(null);
|
||||
} catch (CompletionException | java.util.concurrent.CancellationException e) {
|
||||
Throwable cause = (e instanceof CompletionException ce && ce.getCause() != null) ? ce.getCause() : e;
|
||||
return new TaskView(ticket, Phase.FAILED, null, null, cause.getMessage());
|
||||
return new TaskView(ticket, Phase.FAILED, null, null, cause.getMessage(), null);
|
||||
}
|
||||
if (r.completed()) {
|
||||
String source = r.outcome() == Outcome.REPLIED ? "reply" : "transcript";
|
||||
return new TaskView(ticket, Phase.DONE, r.text(), source, null);
|
||||
return new TaskView(ticket, Phase.DONE, r.text(), source, null, null);
|
||||
}
|
||||
// A wedged worker (CB-109) carries the error context as its reason; the timeout/busy
|
||||
// outcomes carry none, so fall back to the outcome name.
|
||||
String detail = r.outcome() == Outcome.WORKER_FAILED && r.text() != null
|
||||
? r.text()
|
||||
: "no reply — " + r.outcome().name().toLowerCase();
|
||||
return new TaskView(ticket, Phase.FAILED, null, null, detail);
|
||||
return new TaskView(ticket, Phase.FAILED, null, null, detail, null);
|
||||
}
|
||||
|
||||
/** Best-effort live worker status for a pending poll; never throws (a lookup error is just noise). */
|
||||
@@ -536,7 +610,51 @@ public final class MessageService {
|
||||
/** Drop finished tickets older than the TTL so the registry cannot grow without bound. */
|
||||
private void pruneTerminalTickets() {
|
||||
long cutoff = System.nanoTime() - TICKET_TTL_NANOS;
|
||||
tasks.values().removeIf(t -> t.future().isDone() && t.createdNanos() < cutoff);
|
||||
tasks.values().removeIf(t -> t.future.isDone() && t.createdNanos < cutoff);
|
||||
}
|
||||
|
||||
/** Record the active question for an async ticket; blocking sends have no entry and stay unchanged. */
|
||||
private Task markAsyncQuestion(CompletableFuture<Rendezvous.Resolution> waiter, String text, String turnId) {
|
||||
Task task = waiter == null ? null : asyncTasksByWaiter.get(waiter);
|
||||
if (task != null) {
|
||||
task.question = new Reply(Outcome.QUESTION, text, turnId);
|
||||
task.turnId = turnId;
|
||||
asyncTasksByTurn.put(turnId, task);
|
||||
}
|
||||
return task;
|
||||
}
|
||||
|
||||
/** Clear an answered or lapsed question, but only when it matches the ticket's current turn. */
|
||||
private void clearAsyncQuestion(String turnId, boolean forgetTurn) {
|
||||
Task task = asyncTasksByTurn.get(turnId);
|
||||
if (task != null && turnId.equals(task.turnId)) {
|
||||
task.question = null;
|
||||
if (forgetTurn) {
|
||||
asyncTasksByTurn.remove(turnId, task);
|
||||
task.turnId = null;
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
/** Complete and detach an async ticket after its worker's actual terminal reply. */
|
||||
private void finishAsyncTask(Task task, Reply result) {
|
||||
task.future.complete(result);
|
||||
if (task.turnId != null) {
|
||||
asyncTasksByTurn.remove(task.turnId, task);
|
||||
}
|
||||
}
|
||||
|
||||
/** Complete the async ticket correlated to a specific answered turn. */
|
||||
private void finishAsyncTask(String turnId, Reply result) {
|
||||
Task task = asyncTasksByTurn.get(turnId);
|
||||
if (task != null) {
|
||||
finishAsyncTask(task, result);
|
||||
}
|
||||
}
|
||||
|
||||
/** A new send must not open a waiter while an async ticket owns this worker's paused turn. */
|
||||
private boolean hasAsyncQuestion(String target) {
|
||||
return asyncTasksByTurn.values().stream().anyMatch(task -> target.equals(task.target));
|
||||
}
|
||||
|
||||
/** Release the async executor. */
|
||||
|
||||
@@ -175,6 +175,17 @@ public final class ReplyPushLoop {
|
||||
|
||||
// --- lifecycle -----------------------------------------------------------------------------
|
||||
|
||||
/**
|
||||
* Whether any reminder loop is currently active for some target (CB-551). The idle-lead heartbeat
|
||||
* uses this to stand aside: while the push loop is actively nudging the lead, a concurrent
|
||||
* heartbeat injection would start a second competing turn in the same pane — racing loops multiply
|
||||
* turns and context burn. "Active" means a schedule exists in {@link #activeTargets}; the set is
|
||||
* bounded by what has been triggered, not by any persistent state.
|
||||
*/
|
||||
public boolean isActive() {
|
||||
return !activeTargets.isEmpty();
|
||||
}
|
||||
|
||||
/** Shut down the scheduler. Outstanding reminders are cancelled. */
|
||||
public void stop() {
|
||||
scheduler.shutdownNow();
|
||||
|
||||
@@ -0,0 +1,122 @@
|
||||
package dev.ltms.bridged.peer;
|
||||
|
||||
import java.util.Locale;
|
||||
|
||||
/**
|
||||
* What a member is <em>for</em> — the contract it runs under.
|
||||
*
|
||||
* <p>A member is anything a lead spawns. Every member carries two independent attributes:
|
||||
*
|
||||
* <ul>
|
||||
* <li><b>role</b> (this enum) — <em>which contract</em>: the launch charter it is given, the role
|
||||
* file it reads, the playbook skill it loads, and its authorization row.</li>
|
||||
* <li><b>profile</b> (a {@code profiles:} key) — <em>which backend</em>: model, CLI adapter,
|
||||
* credentials, cost.</li>
|
||||
* </ul>
|
||||
*
|
||||
* <p>These are separate axes on purpose. A {@code REVIEWER} may run on the same profile as the
|
||||
* {@code DEV} whose diff it reviews — same backend, different contract. That case is what proves
|
||||
* role and profile cannot be collapsed into one field.
|
||||
*
|
||||
* <p>A lead is deliberately <em>not</em> a role here. A lead is not spawned: it is a pre-existing,
|
||||
* human-facing session that config recognises. Only spawned peers have a member role.
|
||||
*/
|
||||
public enum MemberRole {
|
||||
|
||||
/**
|
||||
* Refines a ticket before anyone implements it: scope, acceptance criteria, risks, unit split.
|
||||
*
|
||||
* <p>Reads the repo and writes analysis. Never commits code and never opens a pull request —
|
||||
* an architect that starts implementing has stopped doing the job that makes it useful.
|
||||
*
|
||||
* <p>Architects are the one member kind declared in config, because a lead addresses the same
|
||||
* slots across many tickets and needs a stable name for them.
|
||||
*/
|
||||
ARCHITECT,
|
||||
|
||||
/**
|
||||
* Implements one unit of work: provisions a worktree, writes the code, commits, pushes, and
|
||||
* opens its own pull request.
|
||||
*
|
||||
* <p>Never merges. The lead is the gate, and a member that merged its own work would remove the
|
||||
* only independent check in the flow.
|
||||
*/
|
||||
DEV,
|
||||
|
||||
/**
|
||||
* Reviews a diff it did not write and reports one structured finding.
|
||||
*
|
||||
* <p>Never commits, never merges, and is never the member that wrote the scope under review.
|
||||
* Briefed from the diff rather than from the implementer's rationale, because that rationale
|
||||
* carries the same blind spot that produced the bug.
|
||||
*/
|
||||
REVIEWER;
|
||||
|
||||
/** The lowercase spelling used in config and on the wire ({@code architect}, {@code dev}, …). */
|
||||
public String wireName() {
|
||||
return name().toLowerCase(Locale.ROOT);
|
||||
}
|
||||
|
||||
/**
|
||||
* The {@code fleet:} block that holds this role's pool — {@code architects},
|
||||
* {@code developers}, {@code reviewers}.
|
||||
*
|
||||
* <p>Plural, and not always the wire name: the pool of things a {@code dev} may run on reads
|
||||
* naturally as {@code developers:}. The wire name stays the singular {@code dev}, because that
|
||||
* is what a tab label and a roster row say.
|
||||
*/
|
||||
public String configKey() {
|
||||
return switch (this) {
|
||||
case ARCHITECT -> "architects";
|
||||
case DEV -> "developers";
|
||||
case REVIEWER -> "reviewers";
|
||||
};
|
||||
}
|
||||
|
||||
/**
|
||||
* The role owning the {@code fleet:} pool named {@code key}, or {@code null} when the key is not
|
||||
* a role pool ({@code leaders}, {@code tabLabel}, …). Null rather than a throw: callers use this
|
||||
* to sort a fleet block's children, where a non-pool key is normal rather than an error.
|
||||
*/
|
||||
public static MemberRole fromConfigKey(String key) {
|
||||
if (key == null) {
|
||||
return null;
|
||||
}
|
||||
String k = key.trim().toLowerCase(Locale.ROOT);
|
||||
for (MemberRole r : values()) {
|
||||
if (r.configKey().equals(k)) {
|
||||
return r;
|
||||
}
|
||||
}
|
||||
return null;
|
||||
}
|
||||
|
||||
/**
|
||||
* Parse a config/wire spelling, case-insensitively.
|
||||
*
|
||||
* @param s the spelling to parse
|
||||
* @return the matching role
|
||||
* @throws IllegalArgumentException when {@code s} is null, blank, or not a known role — the
|
||||
* message lists the valid spellings, because a typo'd role in
|
||||
* config should fail at startup with the fix in the error
|
||||
*/
|
||||
public static MemberRole parse(String s) {
|
||||
if (s != null && !s.isBlank()) {
|
||||
String t = s.trim().toLowerCase(Locale.ROOT);
|
||||
for (MemberRole r : values()) {
|
||||
if (r.wireName().equals(t)) {
|
||||
return r;
|
||||
}
|
||||
}
|
||||
}
|
||||
StringBuilder valid = new StringBuilder();
|
||||
for (MemberRole r : values()) {
|
||||
if (!valid.isEmpty()) {
|
||||
valid.append(", ");
|
||||
}
|
||||
valid.append(r.wireName());
|
||||
}
|
||||
throw new IllegalArgumentException(
|
||||
"unknown member role '" + s + "'; valid roles are: " + valid);
|
||||
}
|
||||
}
|
||||
@@ -1,24 +1,32 @@
|
||||
package dev.ltms.bridged.peer;
|
||||
|
||||
/**
|
||||
* Parameters for a {@link PeerLauncher#spawn(SpawnRequest)} call — the peer-neutral
|
||||
* aggregation of what the core knows at delegation time: which profile to use, the caller's
|
||||
* requested working directory, and the caller's own cwd (to inherit when no other cwd is set).
|
||||
* What a caller asks for when spawning a peer.
|
||||
*
|
||||
* <p>A null or blank {@code profileName} means "use the launcher's default profile."
|
||||
* A null or blank {@code requestedCwd} means "inherit from config or caller."
|
||||
* A null {@code callerCwd} means "the request came from the daemon itself (not a primary)."
|
||||
*
|
||||
* <p>{@code sessionName} and {@code resumeSessionId} carry the session's durable identity (CB-547a):
|
||||
* the bridge's LOGICAL name for the session (stable across restarts, meaningful to an operator)
|
||||
* and the peer's OWN prior session id to resume, respectively. Both are <em>opted in</em> — either
|
||||
* may be null/blank, in which case the launcher derives a display name and mints a fresh session.
|
||||
* @param profileName the {@code profiles:} entry to spawn on; {@code null}/blank ⇒ the caller
|
||||
* did not choose, and the role's pool supplies the candidates
|
||||
* @param requestedCwd working directory asked for by the caller ({@code null} ⇒ unset)
|
||||
* @param callerCwd the caller's own working directory, used when nothing else pins one
|
||||
* @param sessionName agent session name (CB-547a); {@code null} ⇒ the launcher mints one
|
||||
* @param resumeSessionId prior agent session to resume; {@code null} ⇒ a fresh session
|
||||
* @param role the contract this member runs under (CB-557). Picks the tab label and,
|
||||
* with the profile, the counter its tab number comes from. {@code null} is
|
||||
* read as {@link MemberRole#DEV} — an unqualified spawn is a unit of work
|
||||
*/
|
||||
public record SpawnRequest(String profileName, String requestedCwd, String callerCwd,
|
||||
String sessionName, String resumeSessionId) {
|
||||
String sessionName, String resumeSessionId, MemberRole role) {
|
||||
|
||||
public SpawnRequest {
|
||||
role = (role == null) ? MemberRole.DEV : role;
|
||||
}
|
||||
|
||||
/** Back-compat: a spawn with no session identity (fresh session, launcher-derived name). */
|
||||
public SpawnRequest(String profileName, String requestedCwd, String callerCwd) {
|
||||
this(profileName, requestedCwd, callerCwd, null, null);
|
||||
this(profileName, requestedCwd, callerCwd, null, null, null);
|
||||
}
|
||||
|
||||
/** Pre-CB-557 shape: session identity without an explicit role (defaults to {@code dev}). */
|
||||
public SpawnRequest(String profileName, String requestedCwd, String callerCwd,
|
||||
String sessionName, String resumeSessionId) {
|
||||
this(profileName, requestedCwd, callerCwd, sessionName, resumeSessionId, null);
|
||||
}
|
||||
}
|
||||
|
||||
@@ -11,11 +11,12 @@ import dev.ltms.bridged.metrics.Metrics;
|
||||
import dev.ltms.bridged.herdr.Agent;
|
||||
import dev.ltms.bridged.herdr.HerdrClient;
|
||||
import dev.ltms.bridged.herdr.HerdrException;
|
||||
import dev.ltms.bridged.inject.WorkerPresence;
|
||||
import dev.ltms.bridged.inject.MemberPresence;
|
||||
import dev.ltms.bridged.peer.PeerUnreachableException;
|
||||
import dev.ltms.bridged.msg.MessageService;
|
||||
import dev.ltms.bridged.session.SessionManager;
|
||||
import dev.ltms.bridged.session.WorkerSession;
|
||||
import dev.ltms.bridged.peer.MemberRole;
|
||||
import dev.ltms.bridged.session.MemberSession;
|
||||
import dev.ltms.bridged.session.WorktreeRequest;
|
||||
import dev.ltms.bridged.peer.PeerLauncher;
|
||||
import io.javalin.Javalin;
|
||||
@@ -55,7 +56,7 @@ public final class BridgedApp {
|
||||
private final PeerLauncher workers;
|
||||
private final SessionManager sessions; // CB-301: authoritative session registry
|
||||
private final MessageService messages;
|
||||
private final WorkerPresence presence; // CB-113: which workers are MCP-connected (available)
|
||||
private final MemberPresence presence; // CB-113: which workers are MCP-connected (available)
|
||||
private final HttpServlet mcpServlet; // MCP Streamable-HTTP endpoint, mounted at /mcp (nullable)
|
||||
private final CallerResolver auth; // CB-501: null → authz not enforced (legacy behaviour)
|
||||
private final Metrics metrics; // CB-502: null → /metrics not exposed
|
||||
@@ -67,7 +68,7 @@ public final class BridgedApp {
|
||||
* behaviour without each needing an auth fixture.
|
||||
*/
|
||||
public BridgedApp(HerdrClient herdr, PeerLauncher workers, SessionManager sessions,
|
||||
MessageService messages, WorkerPresence presence,
|
||||
MessageService messages, MemberPresence presence,
|
||||
HttpServlet mcpServlet) {
|
||||
this(herdr, workers, sessions, messages, presence, mcpServlet, null, null);
|
||||
}
|
||||
@@ -79,7 +80,7 @@ public final class BridgedApp {
|
||||
* the endpoint
|
||||
*/
|
||||
public BridgedApp(HerdrClient herdr, PeerLauncher workers, SessionManager sessions,
|
||||
MessageService messages, WorkerPresence presence,
|
||||
MessageService messages, MemberPresence presence,
|
||||
HttpServlet mcpServlet, CallerResolver auth, Metrics metrics) {
|
||||
this.herdr = herdr;
|
||||
this.workers = workers;
|
||||
@@ -115,10 +116,10 @@ public final class BridgedApp {
|
||||
}
|
||||
app.get("/sessions", this::sessions);
|
||||
app.get("/agents", this::agents);
|
||||
app.get("/workers", this::listWorkers); // CB-304: registry roster + live herdr status
|
||||
app.get("/profiles", this::profiles); // configured worker profiles
|
||||
app.post("/workers", this::spawnWorker); // optional ?profile= or {"profile":…}
|
||||
app.delete("/workers/{paneId}", this::stopWorker);
|
||||
app.get("/members", this::listMembers); // CB-304: registry roster + live herdr status
|
||||
app.get("/profiles", this::profiles); // configured backend profiles
|
||||
app.post("/members", this::spawnMember); // optional ?role=&profile= or {"role":…,"profile":…}
|
||||
app.delete("/members/{paneId}", this::stopMember);
|
||||
app.post("/sessions/{id}/message", this::sendMessage); // bridge_send (primary; blocking, wait:false, or answer via turnId)
|
||||
app.post("/sessions/{id}/reply", this::replyMessage); // bridge_reply (worker)
|
||||
app.get("/sessions/{id}/replies", this::drainReplies); // drain reply inbox (CB-307)
|
||||
@@ -221,7 +222,7 @@ public final class BridgedApp {
|
||||
}
|
||||
|
||||
/** CB-304: bridge-owned roster merged with live herdr status by paneId. */
|
||||
private void listWorkers(Context ctx) {
|
||||
private void listMembers(Context ctx) {
|
||||
if (!allow(ctx, Authz.Action.READ, null)) {
|
||||
return;
|
||||
}
|
||||
@@ -251,10 +252,11 @@ public final class BridgedApp {
|
||||
* body) picks which configured profile; omitted → the default. 403 if the base_url would breach
|
||||
* the subscription boundary, 400 for an unknown profile.
|
||||
*/
|
||||
private void spawnWorker(Context ctx) {
|
||||
private void spawnMember(Context ctx) {
|
||||
if (!allow(ctx, Authz.Action.SPAWN, null)) {
|
||||
return;
|
||||
}
|
||||
String role = ctx.queryParam("role");
|
||||
String profile = ctx.queryParam("profile");
|
||||
String cwd = ctx.queryParam("cwd");
|
||||
String worktree = ctx.queryParam("worktree");
|
||||
@@ -265,6 +267,7 @@ public final class BridgedApp {
|
||||
String body = ctx.body();
|
||||
if (!body.isBlank()) {
|
||||
JsonNode b = mapper.readTree(body);
|
||||
if (role == null || role.isBlank()) role = b.path("role").asText(null);
|
||||
if (profile == null || profile.isBlank()) profile = b.path("profile").asText(null);
|
||||
if (cwd == null || cwd.isBlank()) cwd = b.path("cwd").asText(null);
|
||||
if (worktree == null || worktree.isBlank()) worktree = b.path("worktree").asText(null);
|
||||
@@ -275,10 +278,18 @@ public final class BridgedApp {
|
||||
}
|
||||
}
|
||||
WorktreeRequest wt = worktreeRequest(worktree, ticket);
|
||||
MemberRole memberRole;
|
||||
try {
|
||||
memberRole = (role == null || role.isBlank()) ? MemberRole.DEV : MemberRole.parse(role);
|
||||
} catch (IllegalArgumentException e) {
|
||||
ctx.status(400).json(Map.of("error", "unknown_role", "detail", e.getMessage()));
|
||||
return;
|
||||
}
|
||||
try {
|
||||
// No MCP caller over REST, so callerCwd and ownerTerminal are null.
|
||||
WorkerSession worker = sessions.acquire(blankToNull(profile), blankToNull(cwd), null, null, wt);
|
||||
ctx.status(201).json(view(worker));
|
||||
MemberSession member = sessions.acquire(blankToNull(profile), memberRole,
|
||||
blankToNull(cwd), null, null, wt);
|
||||
ctx.status(201).json(view(member));
|
||||
} catch (GuardException e) {
|
||||
ctx.status(403).json(Map.of("error", "subscription_boundary", "detail", e.getMessage()));
|
||||
} catch (IllegalArgumentException e) {
|
||||
@@ -306,7 +317,7 @@ public final class BridgedApp {
|
||||
}
|
||||
|
||||
/** Tear a worker down by pane id. */
|
||||
private void stopWorker(Context ctx) {
|
||||
private void stopMember(Context ctx) {
|
||||
String paneId = ctx.pathParam("paneId");
|
||||
if (!allow(ctx, Authz.Action.STOP, paneId)) {
|
||||
return;
|
||||
@@ -544,7 +555,7 @@ public final class BridgedApp {
|
||||
}
|
||||
|
||||
/** CB-301 projection of an authoritative bridge-owned session. */
|
||||
private static Map<String, Object> view(WorkerSession s) {
|
||||
private static Map<String, Object> view(MemberSession s) {
|
||||
Map<String, Object> m = new LinkedHashMap<>();
|
||||
m.put("terminalId", s.terminalId());
|
||||
m.put("paneId", s.paneId());
|
||||
|
||||
@@ -35,10 +35,17 @@ public final class GitWorktrees implements Worktrees {
|
||||
/** What {@link #isolateToolSurface} writes for {@code .mcp.json}: a valid, explicitly empty server map. */
|
||||
private static final String NEUTRAL_MCP_CONFIG = "{\n \"mcpServers\": {}\n}\n";
|
||||
|
||||
/** OpenCode's repo-level config. Tracked here, so it lands in every worktree; it carries
|
||||
* {@code {file:.secrets/...}} references to gitignored secrets that never reach a worktree, and
|
||||
* opencode refuses to start on a dangling reference — so it is neutralized and the worker gets
|
||||
* only the config its launcher writes via {@code OPENCODE_CONFIG}. */
|
||||
/** OpenCode's repo-level config. Tracked here, so it lands in every worktree, and it mounts the
|
||||
* primary's gitea and context7 servers with the primary's credentials. Neutralized so the worker
|
||||
* gets only the config its launcher writes via {@code OPENCODE_CONFIG}.
|
||||
*
|
||||
* <p>The reason has changed shape and is now stronger. It used to be a crash: the file carried
|
||||
* {@code {file:.secrets/...}} references to gitignored files that never reached a worktree, and
|
||||
* opencode refuses to start on a dangling reference (CB-543). Those credentials now live in one
|
||||
* shell-level store and the file reads them as {@code {env:...}}, so in a worktree the reference
|
||||
* resolves instead of failing. That is worse, not better: a member would silently inherit the
|
||||
* primary's admin-scoped {@code GITEA_ACCESS_TOKEN}. A loud crash became a quiet privilege leak,
|
||||
* so this entry protects a boundary now rather than papering over a startup error. */
|
||||
private static final String OPENCODE_CONFIG = "opencode.json";
|
||||
|
||||
/** What {@link #isolateToolSurface} writes for {@code opencode.json}: a valid, empty JSON object. */
|
||||
@@ -112,8 +119,8 @@ public final class GitWorktrees implements Worktrees {
|
||||
* paths outside its own worktree. That is not hypothetical: a CB-523 worker made all 59 of its
|
||||
* edits in the primary checkout while compiling its worktree, so every build it ran was of code
|
||||
* that did not contain its changes. {@code opencode.json} is the same trap one tool over — tracked,
|
||||
* so it lands in every worktree, referencing gitignored {@code .secrets/} files that never do, and
|
||||
* opencode refuses to start on the dangling reference. {@code .autoenv} extends the principle to a
|
||||
* so it lands in every worktree, and it mounts gitea and context7 with the primary's own
|
||||
* credentials, which a member must never hold. {@code .autoenv} extends the principle to a
|
||||
* config that is not tracked today: autoenv authorizes by path, so a fresh worktree path is always
|
||||
* unauthorized and its interactive prompt would block every spawn, so re-landing one must be safe.
|
||||
*
|
||||
|
||||
+15
-8
@@ -1,5 +1,7 @@
|
||||
package dev.ltms.bridged.session;
|
||||
|
||||
import dev.ltms.bridged.peer.MemberRole;
|
||||
|
||||
/**
|
||||
* A bridge-owned worker session — the authoritative in-daemon record of a worker this
|
||||
* process spawned. Immutable; state transitions are performed by replacing the record in
|
||||
@@ -10,7 +12,11 @@ package dev.ltms.bridged.session;
|
||||
* dev.ltms.bridged.peer.PeerHandle#id()}, a UUID, and is distinct from the
|
||||
* launcher-private herdr pane coordinate.
|
||||
* @param terminalId herdr terminal handle — the {@code target} for send/read/status
|
||||
* @param profile the worker profile name that spawned this session
|
||||
* @param profile the profile name that spawned this session — WHICH BACKEND
|
||||
* @param role the contract this member runs under — WHAT IT IS FOR. Independent of
|
||||
* {@code profile}: a reviewer may run on the same profile as the dev
|
||||
* whose diff it reads. {@code null} only for a session recorded before
|
||||
* the role was known
|
||||
* @param cwd the resolved working directory the worker started in
|
||||
* @param ownerTerminal the caller that requested this worker ({@code null} = daemon/anon)
|
||||
* @param spawnedAtNanos {@link System#nanoTime()} when the session was registered
|
||||
@@ -18,10 +24,11 @@ package dev.ltms.bridged.session;
|
||||
* @param turnCount number of delegated turns that have been delivered to this session
|
||||
* @param state current lifecycle state in the one-shot FSM
|
||||
*/
|
||||
public record WorkerSession(
|
||||
public record MemberSession(
|
||||
String paneId,
|
||||
String terminalId,
|
||||
String profile,
|
||||
MemberRole role,
|
||||
String cwd,
|
||||
String ownerTerminal,
|
||||
long spawnedAtNanos,
|
||||
@@ -42,20 +49,20 @@ public record WorkerSession(
|
||||
}
|
||||
|
||||
/** Return a copy of this session in {@code state}. */
|
||||
public WorkerSession withState(State state) {
|
||||
return new WorkerSession(paneId, terminalId, profile, cwd, ownerTerminal, spawnedAtNanos,
|
||||
public MemberSession withState(State state) {
|
||||
return new MemberSession(paneId, terminalId, profile, role, cwd, ownerTerminal, spawnedAtNanos,
|
||||
lastActivityAtNanos, turnCount, state, worktree, branch);
|
||||
}
|
||||
|
||||
/** Return a copy with {@code lastActivityAtNanos} updated to {@code nowNanos}. */
|
||||
public WorkerSession withActivity(long nowNanos) {
|
||||
return new WorkerSession(paneId, terminalId, profile, cwd, ownerTerminal, spawnedAtNanos,
|
||||
public MemberSession withActivity(long nowNanos) {
|
||||
return new MemberSession(paneId, terminalId, profile, role, cwd, ownerTerminal, spawnedAtNanos,
|
||||
nowNanos, turnCount, state, worktree, branch);
|
||||
}
|
||||
|
||||
/** Return a copy with the turn count incremented and activity timestamped at {@code nowNanos}. */
|
||||
public WorkerSession bumpTurn(long nowNanos) {
|
||||
return new WorkerSession(paneId, terminalId, profile, cwd, ownerTerminal, spawnedAtNanos,
|
||||
public MemberSession bumpTurn(long nowNanos) {
|
||||
return new MemberSession(paneId, terminalId, profile, role, cwd, ownerTerminal, spawnedAtNanos,
|
||||
nowNanos, turnCount + 1, state, worktree, branch);
|
||||
}
|
||||
}
|
||||
@@ -1,8 +1,10 @@
|
||||
package dev.ltms.bridged.session;
|
||||
|
||||
import dev.ltms.bridged.auth.MemberLifecycle;
|
||||
import dev.ltms.bridged.herdr.Agent;
|
||||
import dev.ltms.bridged.inject.TurnListener;
|
||||
import dev.ltms.bridged.inject.WorkerPresence;
|
||||
import dev.ltms.bridged.inject.MemberPresence;
|
||||
import dev.ltms.bridged.peer.MemberRole;
|
||||
import dev.ltms.bridged.peer.PeerHandle;
|
||||
import dev.ltms.bridged.peer.PeerLauncher;
|
||||
import dev.ltms.bridged.peer.SpawnRequest;
|
||||
@@ -27,11 +29,10 @@ import java.util.function.LongSupplier;
|
||||
* teardown on top.
|
||||
*
|
||||
* <p>The state machine is intentionally one-shot / no-reuse: every acquired worker is fresh,
|
||||
* and a finished or released worker is torn down, never pooled. {@link #recycle} is a convenience
|
||||
* for {@code release + acquire} with a new distinct pane id.
|
||||
* and a finished or released worker is torn down, never pooled or reused.
|
||||
*
|
||||
* <p>The manager implements {@link TurnListener} so the injector's turn boundaries drive
|
||||
* {@code READY → BUSY → DONE} (or {@code FAILED}). It exposes a {@link WorkerPresence} view via
|
||||
* {@code READY → BUSY → DONE} (or {@code FAILED}). It exposes a {@link MemberPresence} view via
|
||||
* {@link #asPresence()}: any MCP contact from a worker marks it present and simultaneously
|
||||
* transitions the session {@code SPAWNING → READY}.
|
||||
*/
|
||||
@@ -41,13 +42,14 @@ public final class SessionManager implements TurnListener {
|
||||
|
||||
private final PeerLauncher launcher;
|
||||
private final Worktrees worktrees;
|
||||
private final ConcurrentHashMap<String /*paneId*/, WorkerSession> registry = new ConcurrentHashMap<>();
|
||||
private final WorkerPresence presence;
|
||||
private final ConcurrentHashMap<String /*paneId*/, MemberSession> registry = new ConcurrentHashMap<>();
|
||||
private final MemberPresence presence;
|
||||
private final SecureRandom nonceRandom = new SecureRandom();
|
||||
private final AtomicLong nonceSeq = new AtomicLong();
|
||||
private final LongSupplier nowNanos;
|
||||
private final int contextCap;
|
||||
private final boolean clearAfterTurn;
|
||||
private volatile MemberLifecycle memberLifecycle = MemberLifecycle.NONE;
|
||||
|
||||
/** CB-520: notified with a terminalId on every acquire; no-op until wired. */
|
||||
private final List<Consumer<String>> acquireListeners = new java.util.concurrent.CopyOnWriteArrayList<>();
|
||||
@@ -90,76 +92,154 @@ public final class SessionManager implements TurnListener {
|
||||
}
|
||||
|
||||
/**
|
||||
* The single {@link WorkerPresence} view of this manager: it records availability and forwards
|
||||
* The single {@link MemberPresence} view of this manager: it records availability and forwards
|
||||
* the signal to the {@code SPAWNING → READY} transition. Pass this to the {@code Injector} and
|
||||
* {@code BridgeMcp} where they previously accepted a plain {@link WorkerPresence}. The same
|
||||
* {@code BridgeMcp} where they previously accepted a plain {@link MemberPresence}. The same
|
||||
* instance is returned every call — presence is shared state, so a fresh bridge per call would
|
||||
* fragment the {@code present} set and lose signals across callers.
|
||||
*/
|
||||
public WorkerPresence asPresence() {
|
||||
public MemberPresence asPresence() {
|
||||
return presence;
|
||||
}
|
||||
|
||||
/**
|
||||
* Spawn a worker and register it as {@link WorkerSession.State#SPAWNING}. The caller's
|
||||
* Spawn a worker and register it as {@link MemberSession.State#SPAWNING}. The caller's
|
||||
* identity is recorded as {@code ownerTerminal} ({@code null} for daemon/anon callers).
|
||||
*/
|
||||
public WorkerSession acquire(String profile, String requestedCwd, String callerCwd,
|
||||
public MemberSession acquire(String profile, String requestedCwd, String callerCwd,
|
||||
String ownerTerminal) {
|
||||
return acquire(profile, requestedCwd, callerCwd, ownerTerminal, null);
|
||||
}
|
||||
|
||||
/**
|
||||
* Spawn a worker, optionally inside a fresh git worktree. When {@code wt} is non-null the
|
||||
* worktree is provisioned, parity-overlaid, and its path becomes the worker's cwd. On any
|
||||
* failure before registration the worktree is removed so no dangling checkout is left.
|
||||
* Spawn a member, optionally inside a fresh git worktree, defaulting the role to
|
||||
* {@link MemberRole#DEV}.
|
||||
*
|
||||
* <p>{@code DEV} is the right default because it is exactly what the old "worker" meant: an
|
||||
* unqualified spawn is a unit of implementation work. An architect or a reviewer is always
|
||||
* asked for on purpose, so neither is ever what a caller silently gets.
|
||||
*/
|
||||
public WorkerSession acquire(String profile, String requestedCwd, String callerCwd,
|
||||
public MemberSession acquire(String profile, String requestedCwd, String callerCwd,
|
||||
String ownerTerminal, WorktreeRequest wt) {
|
||||
return acquire(profile, MemberRole.DEV, requestedCwd, callerCwd, ownerTerminal, wt);
|
||||
}
|
||||
|
||||
/**
|
||||
* Spawn a member, optionally inside a fresh git worktree. When {@code wt} is non-null the
|
||||
* worktree is provisioned, parity-overlaid, and its path becomes the member's cwd. On any
|
||||
* failure before registration the worktree is removed so no dangling checkout is left.
|
||||
*
|
||||
* @param profile which backend to run on — a {@code profiles:} key
|
||||
* @param role which contract the member runs under; never {@code null}
|
||||
*/
|
||||
public MemberSession acquire(String profile, MemberRole role, String requestedCwd, String callerCwd,
|
||||
String ownerTerminal, WorktreeRequest wt) {
|
||||
MemberRole memberRole = (role == null) ? MemberRole.DEV : role;
|
||||
if (wt == null) {
|
||||
SpawnRequest req = new SpawnRequest(profile, requestedCwd, callerCwd);
|
||||
PeerHandle handle = launcher.spawn(req);
|
||||
// CB-557: the role must ride on the SpawnRequest, not stay a local. The launcher needs it
|
||||
// to pick the profile out of that role's pool and to label the tab; a role kept only on
|
||||
// the MemberSession is recorded after the spawn it was supposed to steer.
|
||||
SpawnRequest req = new SpawnRequest(profile, requestedCwd, callerCwd, null, null, memberRole);
|
||||
PeerHandle handle;
|
||||
try {
|
||||
handle = launcher.spawn(req);
|
||||
} catch (RuntimeException e) {
|
||||
log.warn("spawn failed for profile={} role={}: {}", profile, memberRole, e.getMessage());
|
||||
throw e;
|
||||
}
|
||||
String resolvedProfile = resolveProfile(handle, profile);
|
||||
String cwd = launcher.effectiveCwd(new SpawnRequest(resolvedProfile, requestedCwd, callerCwd));
|
||||
long now = nowNanos.getAsLong();
|
||||
WorkerSession session = new WorkerSession(
|
||||
MemberSession session = new MemberSession(
|
||||
handle.id(),
|
||||
handle.terminalId(),
|
||||
resolvedProfile,
|
||||
memberRole,
|
||||
cwd,
|
||||
ownerTerminal,
|
||||
now,
|
||||
now,
|
||||
0,
|
||||
WorkerSession.State.SPAWNING,
|
||||
MemberSession.State.SPAWNING,
|
||||
null,
|
||||
null);
|
||||
registry.put(handle.id(), session);
|
||||
memberLifecycle.acquired(session.role(), session.profile(), session.terminalId());
|
||||
log.debug("acquired session id={} terminal={} profile={} owner={}",
|
||||
handle.id(), handle.terminalId(), session.profile(), session.ownerTerminal());
|
||||
notifyAcquired(session.terminalId());
|
||||
return session;
|
||||
}
|
||||
return acquireWithWorktree(profile, requestedCwd, callerCwd, ownerTerminal, wt);
|
||||
return acquireWithWorktree(profile, memberRole, requestedCwd, callerCwd, ownerTerminal, wt);
|
||||
}
|
||||
|
||||
/** Tear a worker down by pane id and remove it from the registry. Idempotent. */
|
||||
public void release(String paneId) {
|
||||
WorkerSession removed = registry.remove(paneId);
|
||||
release(paneId, ReleaseCause.COMPLETED);
|
||||
}
|
||||
|
||||
/**
|
||||
* Core teardown: always stops the worker pane and deregisters the session; whether the worker's
|
||||
* git worktree is also removed depends on {@code cause}.
|
||||
*
|
||||
* <p>CB-544: these are two concerns that used to be fused. Stopping the pane is correct on every
|
||||
* teardown — the worker process must end. Removing the worktree is a destructive act that is only
|
||||
* correct for a deliberately-finished teardown (an explicit stop of a completed session, the
|
||||
* reaper releasing a genuinely idle one, or a context-capped session). A shutdown drain
|
||||
* must stop panes but preserve worktrees: a worker's uncommitted work exists in exactly one
|
||||
* place — its worktree — so deleting it while the daemon simply goes down is silent data loss,
|
||||
* with no copy and no error. Do NOT fuse these back together; the cost of an orphaned worktree
|
||||
* is a logged path an operator can reclaim, the cost of a deleted one is unrecoverable work.
|
||||
*/
|
||||
private void release(String paneId, ReleaseCause cause) {
|
||||
MemberSession removed = registry.remove(paneId);
|
||||
boolean preserveWorktree = cause == ReleaseCause.SHUTDOWN;
|
||||
if (removed != null) {
|
||||
log.debug("releasing session pane={} terminal={} state={}",
|
||||
removed.paneId(), removed.terminalId(), removed.state());
|
||||
memberLifecycle.released(removed.terminalId());
|
||||
log.debug("releasing session pane={} terminal={} state={} cause={}",
|
||||
removed.paneId(), removed.terminalId(), removed.state(), cause);
|
||||
if (preserveWorktree && removed.worktree() != null) {
|
||||
logPreservedForShutdown(removed);
|
||||
}
|
||||
// CB-516: a send still waiting on this worker can never be answered now. Tell the
|
||||
// listener BEFORE the pane is torn down, so a blocked caller fails fast with a real
|
||||
// reason instead of sitting on a rendezvous nothing will ever resolve.
|
||||
notifyReleased(removed.terminalId());
|
||||
}
|
||||
launcher.stop(paneId);
|
||||
if (removed != null && removed.worktree() != null) {
|
||||
if (removed != null && !preserveWorktree && removed.worktree() != null) {
|
||||
worktrees.remove(worktrees.repoRoot(removed.cwd()), removed.worktree());
|
||||
}
|
||||
}
|
||||
|
||||
/**
|
||||
* Why a session is being released — governs whether its worktree is preserved or removed.
|
||||
* Worktree removal is reserved for the one case that is genuinely finished; everything else
|
||||
* must keep the worker's only copy of its work.
|
||||
*/
|
||||
public enum ReleaseCause {
|
||||
/** Deliberate teardown of a finished session. Stops the pane and removes the worktree. */
|
||||
COMPLETED,
|
||||
/** Daemon shutdown drain. Stops the pane but PRESERVES the worktree. */
|
||||
SHUTDOWN
|
||||
}
|
||||
|
||||
/**
|
||||
* CB-544 shutdown drain log for a worktree we deliberately kept. A session still {@code BUSY}
|
||||
* when the drain timeout expired was abandoned mid-turn — that work may be uncommitted and is
|
||||
* the only copy — so the message is loud and points at the path an operator needs to reclaim.
|
||||
*/
|
||||
private void logPreservedForShutdown(MemberSession session) {
|
||||
if (session.state() == MemberSession.State.BUSY) {
|
||||
log.warn("shutdown drain abandoned BUSY session pane={} terminal={} mid-turn; "
|
||||
+ "worktree preserved at {}", session.paneId(), session.terminalId(),
|
||||
session.worktree());
|
||||
} else {
|
||||
log.info("shutdown drain preserved worktree at {} for pane={}",
|
||||
session.worktree(), session.paneId());
|
||||
}
|
||||
}
|
||||
|
||||
/**
|
||||
* Register a callback invoked with a session's {@code terminalId} whenever it is acquired
|
||||
* (CB-520). This is the hook that lets the reply inbox {@code own} a target's queue.
|
||||
@@ -173,7 +253,7 @@ public final class SessionManager implements TurnListener {
|
||||
/**
|
||||
* Register a callback invoked with a session's {@code terminalId} whenever it is released
|
||||
* (CB-516). Every teardown path funnels through {@link #release}, so one hook covers the REST
|
||||
* and MCP stop tools, the idle-TTL reaper, {@code recycle}, and shutdown drain alike.
|
||||
* and MCP stop tools, the idle-TTL reaper, and shutdown drain alike.
|
||||
*
|
||||
* <p>Added rather than injected because {@code MessageService} — one intended listener — is
|
||||
* constructed after this manager (it needs the injector and rendezvous, which need the session
|
||||
@@ -186,6 +266,11 @@ public final class SessionManager implements TurnListener {
|
||||
}
|
||||
}
|
||||
|
||||
/** Inject the optional member-slot lifecycle after construction without changing constructors. */
|
||||
public void setMemberLifecycle(MemberLifecycle memberLifecycle) {
|
||||
this.memberLifecycle = memberLifecycle == null ? MemberLifecycle.NONE : memberLifecycle;
|
||||
}
|
||||
|
||||
/** A listener failure must never prevent the acquisition it is reacting to. */
|
||||
private void notifyAcquired(String terminalId) {
|
||||
if (terminalId == null) {
|
||||
@@ -214,7 +299,7 @@ public final class SessionManager implements TurnListener {
|
||||
}
|
||||
}
|
||||
|
||||
private WorkerSession acquireWithWorktree(String profile, String requestedCwd, String callerCwd,
|
||||
private MemberSession acquireWithWorktree(String profile, MemberRole memberRole, String requestedCwd, String callerCwd,
|
||||
String ownerTerminal, WorktreeRequest wt) {
|
||||
String preResolvedProfile = (profile == null || profile.isBlank())
|
||||
? launcher.defaultProfile() : profile;
|
||||
@@ -231,8 +316,10 @@ public final class SessionManager implements TurnListener {
|
||||
try {
|
||||
path = worktrees.add(repoRoot, branch, wt.baseRef());
|
||||
worktrees.overlayParity(repoRoot, path, launcher.parityOverlay(preResolvedProfile));
|
||||
handle = launcher.spawn(new SpawnRequest(profile, path, callerCwd));
|
||||
handle = launcher.spawn(new SpawnRequest(profile, path, callerCwd, null, null, memberRole));
|
||||
} catch (RuntimeException e) {
|
||||
log.warn("spawn failed for profile={} role={} branch={} path={}: {}",
|
||||
preResolvedProfile, memberRole, branch, path, e.getMessage());
|
||||
if (path != null) {
|
||||
try {
|
||||
worktrees.remove(repoRoot, path);
|
||||
@@ -245,19 +332,21 @@ public final class SessionManager implements TurnListener {
|
||||
String resolvedProfile = resolveProfile(handle, profile);
|
||||
String cwd = launcher.effectiveCwd(new SpawnRequest(resolvedProfile, path, callerCwd));
|
||||
long now = nowNanos.getAsLong();
|
||||
WorkerSession session = new WorkerSession(
|
||||
MemberSession session = new MemberSession(
|
||||
handle.id(),
|
||||
handle.terminalId(),
|
||||
resolvedProfile,
|
||||
memberRole,
|
||||
cwd,
|
||||
ownerTerminal,
|
||||
now,
|
||||
now,
|
||||
0,
|
||||
WorkerSession.State.SPAWNING,
|
||||
MemberSession.State.SPAWNING,
|
||||
path,
|
||||
branch);
|
||||
registry.put(handle.id(), session);
|
||||
memberLifecycle.acquired(session.role(), session.profile(), session.terminalId());
|
||||
log.debug("acquired worktree session id={} terminal={} profile={} branch={} path={}",
|
||||
handle.id(), handle.terminalId(), session.profile(), session.branch(), session.worktree());
|
||||
notifyAcquired(session.terminalId());
|
||||
@@ -288,26 +377,13 @@ public final class SessionManager implements TurnListener {
|
||||
return launcher.defaultProfile();
|
||||
}
|
||||
|
||||
/**
|
||||
* Release the old session and acquire a fresh one with the same profile and working directory.
|
||||
* The new session is guaranteed to have a pane id distinct from the old one (no-reuse invariant).
|
||||
*/
|
||||
public WorkerSession recycle(String paneId) {
|
||||
WorkerSession old = registry.get(paneId);
|
||||
if (old == null) {
|
||||
throw new IllegalArgumentException("no session for paneId " + paneId);
|
||||
}
|
||||
release(paneId);
|
||||
return acquire(old.profile(), old.cwd(), old.cwd(), old.ownerTerminal());
|
||||
}
|
||||
|
||||
/** The session for {@code paneId}, if it is still registered and not released. */
|
||||
public Optional<WorkerSession> get(String paneId) {
|
||||
public Optional<MemberSession> get(String paneId) {
|
||||
return Optional.ofNullable(registry.get(paneId));
|
||||
}
|
||||
|
||||
/** Bridge-owned roster: all registered sessions (acquired minus released). */
|
||||
public List<WorkerSession> roster() {
|
||||
public List<MemberSession> roster() {
|
||||
return List.copyOf(registry.values());
|
||||
}
|
||||
|
||||
@@ -315,11 +391,15 @@ public final class SessionManager implements TurnListener {
|
||||
* CB-304 merged roster+live view. The registry is authoritative for worktree, branch,
|
||||
* profile, owner, and state; the optional live agent supplies the herdr-reported status.
|
||||
*/
|
||||
public static Map<String, Object> rosterView(WorkerSession session, Agent live) {
|
||||
public static Map<String, Object> rosterView(MemberSession session, Agent live) {
|
||||
Map<String, Object> m = new LinkedHashMap<>();
|
||||
m.put("sessionId", session.terminalId());
|
||||
m.put("paneId", session.paneId());
|
||||
// Both axes, always: profile says which backend this member runs on, role says what it is
|
||||
// for. A lead reading the roster needs both — two rows may share a profile and still be
|
||||
// allowed to do entirely different things.
|
||||
m.put("profile", session.profile());
|
||||
m.put("role", session.role() == null ? "dev" : session.role().wireName());
|
||||
m.put("state", session.state().name().toLowerCase());
|
||||
if (session.worktree() != null) {
|
||||
m.put("worktree", session.worktree());
|
||||
@@ -336,7 +416,7 @@ public final class SessionManager implements TurnListener {
|
||||
|
||||
/** Lifecycle hook: worker became available on the bridge MCP. */
|
||||
void onReady(String terminalId) {
|
||||
transitionByTerminal(terminalId, WorkerSession.State.SPAWNING, WorkerSession.State.READY);
|
||||
transitionByTerminal(terminalId, MemberSession.State.SPAWNING, MemberSession.State.READY);
|
||||
}
|
||||
|
||||
/**
|
||||
@@ -346,13 +426,13 @@ public final class SessionManager implements TurnListener {
|
||||
*/
|
||||
@Override
|
||||
public void onDelivered(String target) {
|
||||
WorkerSession current = findByTerminal(target);
|
||||
MemberSession current = findByTerminal(target);
|
||||
if (current == null) return;
|
||||
if (current.state() != WorkerSession.State.READY && current.state() != WorkerSession.State.DONE) {
|
||||
if (current.state() != MemberSession.State.READY && current.state() != MemberSession.State.DONE) {
|
||||
return;
|
||||
}
|
||||
long now = nowNanos.getAsLong();
|
||||
WorkerSession updated = current.withState(WorkerSession.State.BUSY).bumpTurn(now);
|
||||
MemberSession updated = current.withState(MemberSession.State.BUSY).bumpTurn(now);
|
||||
if (replace(current, updated)) {
|
||||
log.debug("session transitioned terminal={} pane={} {} -> BUSY turn={}",
|
||||
target, current.paneId(), current.state(), updated.turnCount());
|
||||
@@ -368,8 +448,8 @@ public final class SessionManager implements TurnListener {
|
||||
@Override
|
||||
public boolean hasPostTurnAction(String target) {
|
||||
if (!clearAfterTurn) return false;
|
||||
WorkerSession current = findByTerminal(target);
|
||||
return current != null && current.state() == WorkerSession.State.BUSY
|
||||
MemberSession current = findByTerminal(target);
|
||||
return current != null && current.state() == MemberSession.State.BUSY
|
||||
&& (contextCap <= 0 || current.turnCount() < contextCap);
|
||||
}
|
||||
|
||||
@@ -379,10 +459,10 @@ public final class SessionManager implements TurnListener {
|
||||
}
|
||||
|
||||
private boolean completeTurn(String target, boolean startContextReset) {
|
||||
WorkerSession current = findByTerminal(target);
|
||||
if (current == null || current.state() != WorkerSession.State.BUSY) return false;
|
||||
MemberSession current = findByTerminal(target);
|
||||
if (current == null || current.state() != MemberSession.State.BUSY) return false;
|
||||
long now = nowNanos.getAsLong();
|
||||
WorkerSession updated = current.withState(WorkerSession.State.DONE).withActivity(now);
|
||||
MemberSession updated = current.withState(MemberSession.State.DONE).withActivity(now);
|
||||
if (replace(current, updated)) {
|
||||
log.debug("session transitioned terminal={} pane={} BUSY -> DONE turn={}",
|
||||
target, current.paneId(), updated.turnCount());
|
||||
@@ -409,11 +489,13 @@ public final class SessionManager implements TurnListener {
|
||||
|
||||
/** Lifecycle hook: the worker vanished or was dropped mid-life. */
|
||||
void onFailed(String target) {
|
||||
WorkerSession current = findByTerminal(target);
|
||||
MemberSession current = findByTerminal(target);
|
||||
if (current == null) return;
|
||||
if (current.state() == WorkerSession.State.RELEASED) return;
|
||||
if (replace(current, current.withState(WorkerSession.State.FAILED))) {
|
||||
log.debug("session marked failed terminal={} pane={}", target, current.paneId());
|
||||
if (current.state() == MemberSession.State.RELEASED) return;
|
||||
MemberSession.State priorState = current.state();
|
||||
if (replace(current, current.withState(MemberSession.State.FAILED))) {
|
||||
log.warn("member terminal={} pane={} can no longer be delegated to: its turn never resolved "
|
||||
+ "(was {} when it failed)", target, current.paneId(), priorState);
|
||||
}
|
||||
}
|
||||
|
||||
@@ -425,11 +507,15 @@ public final class SessionManager implements TurnListener {
|
||||
int reapIdle(long idleTtlNanos) {
|
||||
long now = nowNanos.getAsLong();
|
||||
int reaped = 0;
|
||||
for (WorkerSession s : roster()) {
|
||||
if (s.state() != WorkerSession.State.READY && s.state() != WorkerSession.State.DONE) {
|
||||
for (MemberSession s : roster()) {
|
||||
if (s.state() != MemberSession.State.READY && s.state() != MemberSession.State.DONE) {
|
||||
continue;
|
||||
}
|
||||
if (now - s.lastActivityAtNanos() > idleTtlNanos) {
|
||||
long idleNanos = now - s.lastActivityAtNanos();
|
||||
if (idleNanos > idleTtlNanos) {
|
||||
log.debug("reaping idle session terminal={} pane={}: idle {}s exceeds the {}s ttl",
|
||||
s.terminalId(), s.paneId(), TimeUnit.NANOSECONDS.toSeconds(idleNanos),
|
||||
TimeUnit.NANOSECONDS.toSeconds(idleTtlNanos));
|
||||
release(s.paneId());
|
||||
reaped++;
|
||||
}
|
||||
@@ -438,19 +524,25 @@ public final class SessionManager implements TurnListener {
|
||||
}
|
||||
|
||||
/**
|
||||
* Gracefully drain all registered sessions. For each session that is {@code BUSY}, poll up to
|
||||
* {@code timeoutNanos} for it to leave {@code BUSY}, then release it regardless. Non-busy
|
||||
* sessions are released immediately. A failure releasing one session is logged and does not
|
||||
* abort the rest.
|
||||
* Gracefully drain all registered sessions on daemon shutdown. For each session that is
|
||||
* {@code BUSY}, poll up to {@code timeoutNanos} for it to leave {@code BUSY}, then release it
|
||||
* regardless. Non-busy sessions are released immediately. A failure releasing one session is
|
||||
* logged and does not abort the rest.
|
||||
*
|
||||
* <p>CB-544: this is a {@link ReleaseCause#SHUTDOWN} release — the worker's pane is stopped
|
||||
* (the process must end) but its worktree is preserved and its path logged. Shutdown is never
|
||||
* a reason to delete a worker's only copy of its uncommitted work. A session still {@code BUSY}
|
||||
* when the timeout expired is abandoned mid-turn and logged loudly so an operator can find its
|
||||
* kept worktree.
|
||||
*/
|
||||
void drainAll(long timeoutNanos) {
|
||||
long deadline = System.nanoTime() + timeoutNanos;
|
||||
for (WorkerSession s : roster()) {
|
||||
for (MemberSession s : roster()) {
|
||||
try {
|
||||
if (s.state() == WorkerSession.State.BUSY) {
|
||||
if (s.state() == MemberSession.State.BUSY) {
|
||||
while (System.nanoTime() < deadline) {
|
||||
WorkerSession current = registry.get(s.paneId());
|
||||
if (current == null || current.state() != WorkerSession.State.BUSY) {
|
||||
MemberSession current = registry.get(s.paneId());
|
||||
if (current == null || current.state() != MemberSession.State.BUSY) {
|
||||
break;
|
||||
}
|
||||
try {
|
||||
@@ -462,7 +554,7 @@ public final class SessionManager implements TurnListener {
|
||||
}
|
||||
}
|
||||
}
|
||||
release(s.paneId());
|
||||
release(s.paneId(), ReleaseCause.SHUTDOWN);
|
||||
} catch (RuntimeException e) {
|
||||
log.warn("drain failed for pane={}; continuing with remaining sessions", s.paneId(), e);
|
||||
}
|
||||
@@ -489,22 +581,22 @@ public final class SessionManager implements TurnListener {
|
||||
* <p>A null {@code terminalId} is a normal input, not a caller bug: every lifecycle hook here is
|
||||
* fed from the MCP transport, where the <em>primary</em> resolves to a {@link
|
||||
* dev.ltms.bridged.auth.Principal} with no terminal. {@code BridgeMcp} documents that contact as
|
||||
* a no-op, and {@link dev.ltms.bridged.inject.WorkerPresence#markPresent} honours it — but
|
||||
* a no-op, and {@link dev.ltms.bridged.inject.MemberPresence#markPresent} honours it — but
|
||||
* {@code PresenceBridge} then forwards the same null here. Matching on a null id can never
|
||||
* succeed anyway (a registered session always has a terminal), so answer "no match" rather than
|
||||
* throwing: an NPE on this path takes down an unrelated tool call for the primary.
|
||||
*/
|
||||
private WorkerSession findByTerminal(String terminalId) {
|
||||
private MemberSession findByTerminal(String terminalId) {
|
||||
if (terminalId == null) return null;
|
||||
for (WorkerSession s : registry.values()) {
|
||||
for (MemberSession s : registry.values()) {
|
||||
if (terminalId.equals(s.terminalId())) return s;
|
||||
}
|
||||
return null;
|
||||
}
|
||||
|
||||
private void transitionByTerminal(String terminalId, WorkerSession.State from,
|
||||
WorkerSession.State to) {
|
||||
WorkerSession current = findByTerminal(terminalId);
|
||||
private void transitionByTerminal(String terminalId, MemberSession.State from,
|
||||
MemberSession.State to) {
|
||||
MemberSession current = findByTerminal(terminalId);
|
||||
if (current == null || current.state() != from) return;
|
||||
long now = nowNanos.getAsLong();
|
||||
if (replace(current, current.withState(to).withActivity(now))) {
|
||||
@@ -513,12 +605,12 @@ public final class SessionManager implements TurnListener {
|
||||
}
|
||||
}
|
||||
|
||||
private boolean replace(WorkerSession expected, WorkerSession updated) {
|
||||
private boolean replace(MemberSession expected, MemberSession updated) {
|
||||
return registry.replace(expected.paneId(), expected, updated);
|
||||
}
|
||||
|
||||
/** WorkerPresence bridge that also drives the manager's READY transition. */
|
||||
private static final class PresenceBridge extends WorkerPresence {
|
||||
/** MemberPresence bridge that also drives the manager's READY transition. */
|
||||
private static final class PresenceBridge extends MemberPresence {
|
||||
private final SessionManager sessions;
|
||||
|
||||
PresenceBridge(SessionManager sessions) {
|
||||
|
||||
@@ -1,4 +1,13 @@
|
||||
<configuration>
|
||||
<!--
|
||||
CB-575: MCP SDK 2.0.0 has no public notification registration API. It registers only
|
||||
notifications/initialized and notifications/roots/list_changed, so other client notifications
|
||||
still warn when unhandled. Clients may legitimately send notifications/cancelled; suppress only
|
||||
that SDK WARN because it would devalue the action-needed WARN level used by M4 fleet health.
|
||||
If a later SDK handles cancellation, this filter simply stops matching and can be removed.
|
||||
-->
|
||||
<turboFilter class="dev.ltms.bridged.logging.McpCancelledNotificationFilter"/>
|
||||
|
||||
<appender name="STDOUT" class="ch.qos.logback.core.ConsoleAppender">
|
||||
<encoder>
|
||||
<pattern>%d{HH:mm:ss.SSS} %-5level [%thread] %logger{28} - %msg%n</pattern>
|
||||
|
||||
@@ -1,6 +1,6 @@
|
||||
package dev.ltms.bridged;
|
||||
|
||||
import dev.ltms.bridged.inject.WorkerPresence;
|
||||
import dev.ltms.bridged.inject.MemberPresence;
|
||||
import org.junit.jupiter.api.DisplayName;
|
||||
import org.junit.jupiter.api.Test;
|
||||
|
||||
@@ -28,7 +28,7 @@ class BridgedDeliverabilityTest {
|
||||
@Test
|
||||
@DisplayName("a worker that has connected its MCP is deliverable")
|
||||
void presentWorkerIsDeliverable() {
|
||||
WorkerPresence presence = new WorkerPresence();
|
||||
MemberPresence presence = new MemberPresence();
|
||||
presence.markPresent("term_worker");
|
||||
|
||||
assertTrue(Bridged.deliverableTo(presence, leads(Map.of())).test("term_worker"));
|
||||
@@ -37,13 +37,13 @@ class BridgedDeliverabilityTest {
|
||||
@Test
|
||||
@DisplayName("a worker still in its boot window is held back")
|
||||
void absentWorkerIsNotDeliverable() {
|
||||
assertFalse(Bridged.deliverableTo(new WorkerPresence(), leads(Map.of())).test("term_booting"));
|
||||
assertFalse(Bridged.deliverableTo(new MemberPresence(), leads(Map.of())).test("term_booting"));
|
||||
}
|
||||
|
||||
@Test
|
||||
@DisplayName("a lead is deliverable without ever being marked present")
|
||||
void leadIsDeliverableWithoutPresence() {
|
||||
WorkerPresence presence = new WorkerPresence();
|
||||
MemberPresence presence = new MemberPresence();
|
||||
Predicate<String> deliverable =
|
||||
Bridged.deliverableTo(presence, leads(Map.of("term_lead", "opus-5.0")));
|
||||
|
||||
@@ -54,7 +54,7 @@ class BridgedDeliverabilityTest {
|
||||
@Test
|
||||
@DisplayName("an unknown terminal is deliverable to neither")
|
||||
void strangerIsNotDeliverable() {
|
||||
WorkerPresence presence = new WorkerPresence();
|
||||
MemberPresence presence = new MemberPresence();
|
||||
presence.markPresent("term_worker");
|
||||
|
||||
assertFalse(Bridged.deliverableTo(presence, leads(Map.of("term_lead", "opus-5.0")))
|
||||
@@ -65,7 +65,7 @@ class BridgedDeliverabilityTest {
|
||||
@DisplayName("a lead discovered after startup becomes deliverable with no restart")
|
||||
void leadSetIsReadThroughOnEveryCall() {
|
||||
Map<String, String> discovered = new HashMap<>();
|
||||
Predicate<String> deliverable = Bridged.deliverableTo(new WorkerPresence(), leads(discovered));
|
||||
Predicate<String> deliverable = Bridged.deliverableTo(new MemberPresence(), leads(discovered));
|
||||
|
||||
assertFalse(deliverable.test("term_late"));
|
||||
discovered.put("term_late", "gpt-sol-5.6"); // leadScan picks up a newly labelled tab
|
||||
@@ -75,7 +75,7 @@ class BridgedDeliverabilityTest {
|
||||
@Test
|
||||
@DisplayName("forgetting a torn-down worker does not strip a lead of its deliverability")
|
||||
void forgetDoesNotDisarmALead() {
|
||||
WorkerPresence presence = new WorkerPresence();
|
||||
MemberPresence presence = new MemberPresence();
|
||||
Predicate<String> deliverable =
|
||||
Bridged.deliverableTo(presence, leads(Map.of("term_lead", "opus-5.0")));
|
||||
|
||||
|
||||
@@ -3,6 +3,8 @@ package dev.ltms.bridged.auth;
|
||||
import dev.ltms.bridged.herdr.FakeHerdr;
|
||||
import dev.ltms.bridged.herdr.PaneLocator;
|
||||
import dev.ltms.bridged.mcp.ConnectionIdentity;
|
||||
import dev.ltms.bridged.config.BridgedConfig;
|
||||
import dev.ltms.bridged.peer.MemberRole;
|
||||
import org.junit.jupiter.api.Test;
|
||||
|
||||
import java.util.Map;
|
||||
@@ -32,6 +34,17 @@ class CallerResolverTest {
|
||||
return identity(999_999);
|
||||
}
|
||||
|
||||
private static MemberRegistry boundMembers(String slot, MemberRole role) {
|
||||
Map<String, BridgedConfig.Slot> architects = role == MemberRole.ARCHITECT
|
||||
? Map.of("lead-designer", new BridgedConfig.Slot("sonnet")) : Map.of();
|
||||
Map<String, BridgedConfig.Slot> devs = role == MemberRole.DEV
|
||||
? Map.of("builder", new BridgedConfig.Slot("sonnet")) : Map.of();
|
||||
MemberRegistry members = new MemberRegistry(
|
||||
new BridgedConfig.Fleet(Map.of(), architects, devs, Map.of(), null));
|
||||
assertTrue(members.bind(slot, "term_a"));
|
||||
return members;
|
||||
}
|
||||
|
||||
@Test
|
||||
void aLoopbackWorkerPaneResolvesToWorkerRegardlessOfAuthMode() {
|
||||
Principal underTrust = new CallerResolver(workerIdentity()).resolve("127.0.0.1", 42, null);
|
||||
@@ -302,8 +315,9 @@ class CallerResolverTest {
|
||||
|
||||
@Test
|
||||
void aBoundArchitectPaneResolvesToArchitectBeforeTheWorkerFallback() {
|
||||
Principal p = CallerResolver.withLeadsAndArchitects(workerIdentity(), false, null,
|
||||
Map::of, () -> Map.of("term_a", "lead-designer"))
|
||||
// This is the production construction path used by Bridged.
|
||||
Principal p = CallerResolver.withLeadsAndMembers(workerIdentity(), false, null,
|
||||
Map::of, boundMembers("architect:lead-designer", MemberRole.ARCHITECT))
|
||||
.resolve("127.0.0.1", 42, null);
|
||||
|
||||
assertEquals(Role.ARCHITECT, p.role(),
|
||||
@@ -315,8 +329,8 @@ class CallerResolverTest {
|
||||
|
||||
@Test
|
||||
void anArchitectNeedsNoTokenEvenInTokenMode() {
|
||||
Principal p = CallerResolver.withLeadsAndArchitects(workerIdentity(), true, "s3cret",
|
||||
Map::of, () -> Map.of("term_a", "lead-designer"))
|
||||
Principal p = CallerResolver.withLeadsAndMembers(workerIdentity(), true, "s3cret",
|
||||
Map::of, boundMembers("architect:lead-designer", MemberRole.ARCHITECT))
|
||||
.resolve("127.0.0.1", 42, null);
|
||||
|
||||
assertEquals(Role.ARCHITECT, p.role(),
|
||||
@@ -325,9 +339,10 @@ class CallerResolverTest {
|
||||
|
||||
@Test
|
||||
void anUnboundPaneStillResolvesAsAWorker() {
|
||||
Map<String, String> arch = Map.of("term_elsewhere", "reviewer");
|
||||
Principal p = CallerResolver.withLeadsAndArchitects(workerIdentity(), false, null,
|
||||
Map::of, () -> arch).resolve("127.0.0.1", 42, null);
|
||||
MemberRegistry members = new MemberRegistry(new BridgedConfig.Fleet(Map.of(),
|
||||
Map.of("lead-designer", new BridgedConfig.Slot("sonnet")), Map.of(), Map.of(), null));
|
||||
Principal p = CallerResolver.withLeadsAndMembers(workerIdentity(), false, null,
|
||||
Map::of, members).resolve("127.0.0.1", 42, null);
|
||||
|
||||
assertEquals(Role.WORKER, p.role());
|
||||
assertNull(p.name());
|
||||
@@ -336,8 +351,9 @@ class CallerResolverTest {
|
||||
/** CB-548 precedence: lead > architect > worker, so a pane named in BOTH is still a lead. */
|
||||
@Test
|
||||
void aLeadWinsOverAnArchitectBindingForTheSamePane() {
|
||||
Principal p = CallerResolver.withLeadsAndArchitects(workerIdentity(), false, null,
|
||||
() -> Map.of("term_a", "opus-5.0"), () -> Map.of("term_a", "lead-designer"))
|
||||
Principal p = CallerResolver.withLeadsAndMembers(workerIdentity(), false, null,
|
||||
() -> Map.of("term_a", "opus-5.0"),
|
||||
boundMembers("architect:lead-designer", MemberRole.ARCHITECT))
|
||||
.resolve("127.0.0.1", 42, null);
|
||||
|
||||
assertEquals(Role.PRIMARY, p.role(),
|
||||
@@ -349,26 +365,17 @@ class CallerResolverTest {
|
||||
/** The registry is live, like leads: a binding injected after construction is honoured. */
|
||||
@Test
|
||||
void anArchitectBoundAfterConstructionIsHonouredWithoutRebuildingTheResolver() {
|
||||
Map<String, String> live = new java.util.HashMap<>();
|
||||
CallerResolver r = CallerResolver.withLeadsAndArchitects(workerIdentity(), false, null,
|
||||
Map::of, () -> live);
|
||||
MemberRegistry members = new MemberRegistry(new BridgedConfig.Fleet(Map.of(),
|
||||
Map.of("lead-designer", new BridgedConfig.Slot("sonnet")), Map.of(), Map.of(), null));
|
||||
CallerResolver r = CallerResolver.withLeadsAndMembers(workerIdentity(), false, null,
|
||||
Map::of, members);
|
||||
|
||||
assertEquals(Role.WORKER, r.resolve("127.0.0.1", 42, null).role());
|
||||
|
||||
live.put("term_a", "lead-designer"); // the later lifecycle binds the slot
|
||||
assertTrue(members.bind("architect:lead-designer", "term_a")); // the later lifecycle binds the slot
|
||||
|
||||
assertEquals(Role.ARCHITECT, r.resolve("127.0.0.1", 42, null).role());
|
||||
assertEquals("lead-designer", r.architects().get("term_a"));
|
||||
}
|
||||
|
||||
@Test
|
||||
void theArchitectMapFormIsCopiedSoLaterMutationCannotGrantArchitect() {
|
||||
Map<String, String> mutable = new java.util.LinkedHashMap<>();
|
||||
CallerResolver r = new CallerResolver(workerIdentity(), false, null, Map.of(), mutable);
|
||||
|
||||
mutable.put("term_a", "sneaky");
|
||||
|
||||
assertEquals(Role.WORKER, r.resolve("127.0.0.1", 42, null).role());
|
||||
assertEquals("architect:lead-designer", r.members().get("term_a"));
|
||||
}
|
||||
|
||||
@Test
|
||||
@@ -380,8 +387,9 @@ class CallerResolverTest {
|
||||
/** An architect acts only as its own pane — the same ownsSession rule as a worker or lead. */
|
||||
@Test
|
||||
void anArchitectOwnsItsOwnPaneAndNoOther() {
|
||||
Principal arch = CallerResolver.withLeadsAndArchitects(workerIdentity(), false, null,
|
||||
Map::of, () -> Map.of("term_a", "lead-designer")).resolve("127.0.0.1", 42, null);
|
||||
Principal arch = CallerResolver.withLeadsAndMembers(workerIdentity(), false, null,
|
||||
Map::of, boundMembers("architect:lead-designer", MemberRole.ARCHITECT))
|
||||
.resolve("127.0.0.1", 42, null);
|
||||
|
||||
assertTrue(arch.ownsSession("term_a"));
|
||||
assertTrue(Authz.permits(arch, Authz.Action.REPLY, "term_a"));
|
||||
@@ -389,6 +397,15 @@ class CallerResolverTest {
|
||||
assertFalse(Authz.permits(arch, Authz.Action.REPLY, "term_b"));
|
||||
}
|
||||
|
||||
@Test
|
||||
void aBoundNonArchitectSlotStillResolvesAsAWorker() {
|
||||
Principal p = CallerResolver.withLeadsAndMembers(workerIdentity(), false, null,
|
||||
Map::of, boundMembers("dev:builder", MemberRole.DEV))
|
||||
.resolve("127.0.0.1", 42, null);
|
||||
|
||||
assertEquals(Role.WORKER, p.role(), "a dev binding must never grant architect rights");
|
||||
}
|
||||
|
||||
@Test
|
||||
void tokenModeRequiresANonEmptyConfiguredToken() {
|
||||
ConnectionIdentity id = nonWorkerIdentity();
|
||||
|
||||
+74
-43
@@ -1,12 +1,14 @@
|
||||
package dev.ltms.bridged.auth;
|
||||
|
||||
import dev.ltms.bridged.config.BridgedConfig;
|
||||
import dev.ltms.bridged.peer.MemberRole;
|
||||
import org.junit.jupiter.api.Test;
|
||||
|
||||
import java.util.ArrayList;
|
||||
import java.util.HashMap;
|
||||
import java.util.LinkedHashMap;
|
||||
import java.util.List;
|
||||
import java.util.Map;
|
||||
import java.util.Set;
|
||||
import java.util.concurrent.CountDownLatch;
|
||||
import java.util.concurrent.ExecutorService;
|
||||
import java.util.concurrent.Executors;
|
||||
@@ -20,25 +22,54 @@ import static org.junit.jupiter.api.Assertions.*;
|
||||
* {@link CallerResolverTest}; this pins the registry object itself — its invariants and their
|
||||
* thread-safety.
|
||||
*/
|
||||
class ArchitectRegistryTest {
|
||||
class MemberRegistryTest {
|
||||
|
||||
private static final Map<String, BridgedConfig.Architect> SLOTS = Map.of(
|
||||
"lead-designer", new BridgedConfig.Architect("sonnet"),
|
||||
"reviewer", new BridgedConfig.Architect("gx10"));
|
||||
/**
|
||||
* Slot keys are qualified by role (CB-557): a bare name is unique only within its pool, so
|
||||
* {@code sonnet} can be both a developer and a reviewer, while a terminal binds to exactly one.
|
||||
*/
|
||||
private static final String DESIGNER = "architect:lead-designer";
|
||||
private static final String REVIEWER = "architect:code-reviewer";
|
||||
|
||||
private final ArchitectRegistry registry = new ArchitectRegistry(SLOTS);
|
||||
private static BridgedConfig.Fleet fleetWith(Map<String, BridgedConfig.Slot> architects) {
|
||||
return new BridgedConfig.Fleet(Map.of(), architects, Map.of(), Map.of(), null);
|
||||
}
|
||||
|
||||
private static Map<String, BridgedConfig.Slot> architects() {
|
||||
Map<String, BridgedConfig.Slot> pool = new LinkedHashMap<>();
|
||||
pool.put("lead-designer", new BridgedConfig.Slot("sonnet"));
|
||||
pool.put("code-reviewer", new BridgedConfig.Slot("gx10"));
|
||||
return pool;
|
||||
}
|
||||
|
||||
private final MemberRegistry registry = new MemberRegistry(fleetWith(architects()));
|
||||
|
||||
@Test
|
||||
void exposesTheConfiguredSlots() {
|
||||
assertEquals(SLOTS.keySet(), registry.slots().keySet());
|
||||
assertTrue(registry.isSlot("reviewer"));
|
||||
void exposesTheConfiguredSlotsQualifiedByRole() {
|
||||
assertEquals(Set.of(DESIGNER, REVIEWER), registry.slots().keySet());
|
||||
assertTrue(registry.isSlot(REVIEWER));
|
||||
assertFalse(registry.isSlot("nope"));
|
||||
assertFalse(registry.isSlot("lead-designer"),
|
||||
"the bare name is not the key — it is unique only inside its pool");
|
||||
}
|
||||
|
||||
/** The case the pool shape exists for: one profile serving two roles is not a duplicate. */
|
||||
@Test
|
||||
void oneProfileMayServeTwoRolesUnderTheSameSlotName() {
|
||||
Map<String, BridgedConfig.Slot> devs = Map.of("sonnet", new BridgedConfig.Slot("sonnet"));
|
||||
Map<String, BridgedConfig.Slot> revs = Map.of("sonnet", new BridgedConfig.Slot("sonnet"));
|
||||
MemberRegistry r = new MemberRegistry(
|
||||
new BridgedConfig.Fleet(Map.of(), Map.of(), devs, revs, null));
|
||||
|
||||
assertEquals(Set.of("dev:sonnet", "reviewer:sonnet"), r.slots().keySet());
|
||||
assertEquals(MemberRole.DEV, r.roleForSlot("dev:sonnet"));
|
||||
assertEquals(MemberRole.REVIEWER, r.roleForSlot("reviewer:sonnet"));
|
||||
}
|
||||
|
||||
@Test
|
||||
void theSpawnLifecycleReadsTheProfileBackFromASlot() {
|
||||
assertEquals("sonnet", registry.profileForSlot("lead-designer"));
|
||||
assertEquals("gx10", registry.profileForSlot("reviewer"));
|
||||
assertEquals("sonnet", registry.profileForSlot(DESIGNER));
|
||||
assertEquals("gx10", registry.profileForSlot(REVIEWER));
|
||||
assertNull(registry.profileForSlot("unknown"), "an unknown slot has no profile");
|
||||
}
|
||||
|
||||
@@ -54,9 +85,9 @@ class ArchitectRegistryTest {
|
||||
|
||||
@Test
|
||||
void bindResolvesTheTerminalToTheSlot() {
|
||||
assertTrue(registry.bind("lead-designer", "term_design"));
|
||||
assertEquals("lead-designer", registry.slotForTerminal("term_design"));
|
||||
assertEquals(Map.of("term_design", "lead-designer"), registry.snapshot());
|
||||
assertTrue(registry.bind(DESIGNER, "term_design"));
|
||||
assertEquals(DESIGNER, registry.slotForTerminal("term_design"));
|
||||
assertEquals(Map.of("term_design", DESIGNER), registry.snapshot());
|
||||
}
|
||||
|
||||
@Test
|
||||
@@ -68,27 +99,27 @@ class ArchitectRegistryTest {
|
||||
|
||||
@Test
|
||||
void bindRefusesATerminalInTwoSlots() {
|
||||
assertTrue(registry.bind("lead-designer", "term_design"));
|
||||
assertFalse(registry.bind("reviewer", "term_design"),
|
||||
assertTrue(registry.bind(DESIGNER, "term_design"));
|
||||
assertFalse(registry.bind(REVIEWER, "term_design"),
|
||||
"a terminal may occupy at most one slot");
|
||||
assertEquals("lead-designer", registry.slotForTerminal("term_design"),
|
||||
assertEquals(DESIGNER, registry.slotForTerminal("term_design"),
|
||||
"the first binding survives the refused second");
|
||||
}
|
||||
|
||||
@Test
|
||||
void bindRefusesASlotWithTwoTerminals() {
|
||||
assertTrue(registry.bind("lead-designer", "term_design"));
|
||||
assertFalse(registry.bind("lead-designer", "term_other"),
|
||||
assertTrue(registry.bind(DESIGNER, "term_design"));
|
||||
assertFalse(registry.bind(DESIGNER, "term_other"),
|
||||
"a slot may host at most one terminal");
|
||||
assertEquals("lead-designer", registry.slotForTerminal("term_design"),
|
||||
assertEquals(DESIGNER, registry.slotForTerminal("term_design"),
|
||||
"the first binding survives the refused second");
|
||||
assertNull(registry.slotForTerminal("term_other"));
|
||||
}
|
||||
|
||||
@Test
|
||||
void rebindingTheSamePairIsAnIdempotentNoOp() {
|
||||
assertTrue(registry.bind("lead-designer", "term_design"));
|
||||
assertTrue(registry.bind("lead-designer", "term_design"),
|
||||
assertTrue(registry.bind(DESIGNER, "term_design"));
|
||||
assertTrue(registry.bind(DESIGNER, "term_design"),
|
||||
"the same terminal → slot is harmless to repeat");
|
||||
assertEquals(1, registry.snapshot().size());
|
||||
}
|
||||
@@ -97,8 +128,8 @@ class ArchitectRegistryTest {
|
||||
|
||||
@Test
|
||||
void unbindRemovesTheExactBinding() {
|
||||
assertTrue(registry.bind("lead-designer", "term_design"));
|
||||
assertTrue(registry.unbind("lead-designer", "term_design"));
|
||||
assertTrue(registry.bind(DESIGNER, "term_design"));
|
||||
assertTrue(registry.unbind(DESIGNER, "term_design"));
|
||||
assertNull(registry.slotForTerminal("term_design"));
|
||||
assertTrue(registry.snapshot().isEmpty());
|
||||
}
|
||||
@@ -106,32 +137,32 @@ class ArchitectRegistryTest {
|
||||
@Test
|
||||
void aStaleUnbindDoesNotRemoveAReplacement() {
|
||||
// Bind, tear down, and stand the slot back up with a NEW terminal.
|
||||
assertTrue(registry.bind("lead-designer", "term_design"));
|
||||
registry.unbind("lead-designer", "term_design");
|
||||
assertTrue(registry.bind("lead-designer", "term_new"));
|
||||
assertTrue(registry.bind(DESIGNER, "term_design"));
|
||||
registry.unbind(DESIGNER, "term_design");
|
||||
assertTrue(registry.bind(DESIGNER, "term_new"));
|
||||
|
||||
// A late unbind naming the OLD terminal must not remove the replacement binding.
|
||||
assertFalse(registry.unbind("lead-designer", "term_design"));
|
||||
assertEquals("lead-designer", registry.slotForTerminal("term_new"),
|
||||
assertFalse(registry.unbind(DESIGNER, "term_design"));
|
||||
assertEquals(DESIGNER, registry.slotForTerminal("term_new"),
|
||||
"the replacement terminal stays bound");
|
||||
}
|
||||
|
||||
@Test
|
||||
void aStaleUnbindForATerminalThatMovedSlotsDoesNothing() {
|
||||
// term_design starts in lead-designer, is torn down, and stands back up in a FREE slot.
|
||||
assertTrue(registry.bind("lead-designer", "term_design"));
|
||||
registry.unbind("lead-designer", "term_design");
|
||||
assertTrue(registry.bind("reviewer", "term_design"));
|
||||
assertTrue(registry.bind(DESIGNER, "term_design"));
|
||||
registry.unbind(DESIGNER, "term_design");
|
||||
assertTrue(registry.bind(REVIEWER, "term_design"));
|
||||
|
||||
// Unbinding against the slot it no longer occupies is refused; the new binding is intact.
|
||||
assertFalse(registry.unbind("lead-designer", "term_design"),
|
||||
assertFalse(registry.unbind(DESIGNER, "term_design"),
|
||||
"the old slot must not unbind a terminal that moved elsewhere");
|
||||
assertEquals("reviewer", registry.slotForTerminal("term_design"));
|
||||
assertEquals(REVIEWER, registry.slotForTerminal("term_design"));
|
||||
}
|
||||
|
||||
@Test
|
||||
void unbindOfNothingIsAFalseNoOp() {
|
||||
assertFalse(registry.unbind("lead-designer", "term_design"),
|
||||
assertFalse(registry.unbind(DESIGNER, "term_design"),
|
||||
"nothing was bound, so nothing is removed");
|
||||
}
|
||||
|
||||
@@ -139,14 +170,14 @@ class ArchitectRegistryTest {
|
||||
|
||||
@Test
|
||||
void theSnapshotIsAnImmutableCopyNotAliveState() {
|
||||
assertTrue(registry.bind("lead-designer", "term_design"));
|
||||
assertTrue(registry.bind(DESIGNER, "term_design"));
|
||||
Map<String, String> snap = registry.snapshot();
|
||||
|
||||
assertThrows(UnsupportedOperationException.class, () -> snap.put("x", "y"),
|
||||
"a handed-out snapshot cannot be mutated in place");
|
||||
|
||||
// Later binds must not leak into an earlier snapshot.
|
||||
assertTrue(registry.bind("reviewer", "term_review"));
|
||||
assertTrue(registry.bind(REVIEWER, "term_review"));
|
||||
assertFalse(snap.containsKey("term_review"),
|
||||
"a snapshot is a point-in-time copy, not a live view");
|
||||
}
|
||||
@@ -164,7 +195,7 @@ class ArchitectRegistryTest {
|
||||
final String term = "term_" + i; // every thread races for the SAME slot
|
||||
results.add(pool.submit(() -> {
|
||||
go.await();
|
||||
return registry.bind("lead-designer", term);
|
||||
return registry.bind(DESIGNER, term);
|
||||
}));
|
||||
}
|
||||
go.countDown();
|
||||
@@ -191,7 +222,7 @@ class ArchitectRegistryTest {
|
||||
CountDownLatch go = new CountDownLatch(1);
|
||||
List<Future<String>> results = new ArrayList<>();
|
||||
for (int i = 0; i < n; i++) {
|
||||
final String slot = (i % 2 == 0) ? "lead-designer" : "reviewer"; // all race for ONE terminal
|
||||
final String slot = (i % 2 == 0) ? DESIGNER : REVIEWER; // all race for ONE terminal
|
||||
results.add(pool.submit(() -> {
|
||||
go.await();
|
||||
return registry.bind(slot, "shared_term")
|
||||
@@ -228,11 +259,11 @@ class ArchitectRegistryTest {
|
||||
/** A handed-over slot map is snapshotted at construction, not offered as live state. */
|
||||
@Test
|
||||
void theSlotSnapshotIsFixedByConstruction() {
|
||||
Map<String, BridgedConfig.Architect> mutable = new HashMap<>(SLOTS);
|
||||
ArchitectRegistry r = new ArchitectRegistry(mutable);
|
||||
Map<String, BridgedConfig.Slot> mutable = architects();
|
||||
MemberRegistry r = new MemberRegistry(fleetWith(mutable));
|
||||
|
||||
mutable.put("hijack", new BridgedConfig.Architect("gx10"));
|
||||
mutable.put("hijack", new BridgedConfig.Slot("gx10"));
|
||||
|
||||
assertFalse(r.isSlot("hijack"), "a handed-over map is not offered as live state");
|
||||
assertFalse(r.isSlot("architect:hijack"), "a handed-over map is not offered as live state");
|
||||
}
|
||||
}
|
||||
File diff suppressed because it is too large
Load Diff
@@ -0,0 +1,353 @@
|
||||
package dev.ltms.bridged.config;
|
||||
|
||||
import org.junit.jupiter.api.Test;
|
||||
import org.junit.jupiter.api.io.TempDir;
|
||||
|
||||
import java.nio.file.Files;
|
||||
import java.nio.file.Path;
|
||||
|
||||
import static org.junit.jupiter.api.Assertions.*;
|
||||
|
||||
/**
|
||||
* CB-559: re-reading {@code bridged.yaml} under a running daemon.
|
||||
*
|
||||
* <p>The tests that matter here are the refusals. A reload that applies a good file is the easy
|
||||
* half; the half that protects an operator is the one that keeps the running config when the new
|
||||
* file is bad, and the one that refuses a change the running daemon cannot honour.
|
||||
*/
|
||||
class ConfigRefTest {
|
||||
|
||||
/** A minimal file that loads and passes every startup validator. */
|
||||
private static String yaml(String extra) {
|
||||
return """
|
||||
bind:
|
||||
host: 127.0.0.1
|
||||
port: 8765
|
||||
herdrSocket: ~/.config/herdr/herdr.sock
|
||||
profiles:
|
||||
sonnet:
|
||||
baseUrl: http://gx00.gw:8000
|
||||
model: sonnet
|
||||
guard:
|
||||
offSubscriptionHosts:
|
||||
- gx00.gw
|
||||
""" + extra;
|
||||
}
|
||||
|
||||
private static ConfigRef refFor(Path f) {
|
||||
return new ConfigRef(f, BridgedConfig.load(f));
|
||||
}
|
||||
|
||||
@Test
|
||||
void aHotChangeIsAppliedAndReadThroughGet(@TempDir Path dir) throws Exception {
|
||||
Path f = dir.resolve("bridged.yaml");
|
||||
Files.writeString(f, yaml("""
|
||||
fleet:
|
||||
tabLabel: "{role}: {profile} #{n}"
|
||||
developers:
|
||||
a:
|
||||
profile: sonnet
|
||||
"""));
|
||||
ConfigRef ref = refFor(f);
|
||||
assertEquals("{role}: {profile} #{n}", ref.get().fleet().tabLabel());
|
||||
|
||||
Files.writeString(f, yaml("""
|
||||
fleet:
|
||||
tabLabel: "[{profile}] {role}"
|
||||
developers:
|
||||
a:
|
||||
profile: sonnet
|
||||
"""));
|
||||
ConfigRef.Outcome out = ref.reload();
|
||||
|
||||
assertTrue(out.applied());
|
||||
assertTrue(out.deferred().isEmpty());
|
||||
assertEquals("config reloaded", out.summary());
|
||||
assertEquals("[{profile}] {role}", ref.get().fleet().tabLabel());
|
||||
}
|
||||
|
||||
@Test
|
||||
void aCharterChangeIsHotAndReachesTheLiveConfig(@TempDir Path dir) throws Exception {
|
||||
Path f = dir.resolve("bridged.yaml");
|
||||
Files.writeString(f, yaml("""
|
||||
fleet:
|
||||
charters:
|
||||
architect: old charter
|
||||
"""));
|
||||
ConfigRef ref = refFor(f);
|
||||
assertEquals("old charter", ref.get().fleet().charterFor(
|
||||
dev.ltms.bridged.peer.MemberRole.ARCHITECT));
|
||||
|
||||
Files.writeString(f, yaml("""
|
||||
fleet:
|
||||
charters:
|
||||
architect: new charter
|
||||
"""));
|
||||
ConfigRef.Outcome out = ref.reload();
|
||||
|
||||
assertTrue(out.applied());
|
||||
assertTrue(out.deferred().isEmpty());
|
||||
assertEquals("new charter", ref.get().fleet().charterFor(
|
||||
dev.ltms.bridged.peer.MemberRole.ARCHITECT));
|
||||
}
|
||||
|
||||
@Test
|
||||
void invalidChartersRefuseReloadAndKeepTheRunningConfig(@TempDir Path dir) throws Exception {
|
||||
Path f = dir.resolve("bridged.yaml");
|
||||
Files.writeString(f, yaml("""
|
||||
fleet:
|
||||
charters:
|
||||
architect: valid charter
|
||||
"""));
|
||||
ConfigRef ref = refFor(f);
|
||||
BridgedConfig before = ref.get();
|
||||
|
||||
Files.writeString(f, yaml("""
|
||||
fleet:
|
||||
charters:
|
||||
architect: " "
|
||||
"""));
|
||||
ConfigRef.Outcome blank = ref.reload();
|
||||
assertFalse(blank.applied());
|
||||
assertTrue(blank.error().contains("fleet.charters.architect is blank"));
|
||||
assertSame(before, ref.get());
|
||||
|
||||
Files.writeString(f, yaml("""
|
||||
fleet:
|
||||
charters:
|
||||
architetc: valid charter
|
||||
"""));
|
||||
ConfigRef.Outcome unknown = ref.reload();
|
||||
assertFalse(unknown.applied());
|
||||
assertTrue(unknown.error().contains("architetc"));
|
||||
assertTrue(unknown.error().contains("architect"));
|
||||
assertSame(before, ref.get());
|
||||
}
|
||||
|
||||
/**
|
||||
* The point of the whole class: a consumer holding the ref sees the new value without being
|
||||
* rebuilt. A component that captured {@code get()} into a field would still show the old one.
|
||||
*/
|
||||
@Test
|
||||
void aConsumerHoldingTheRefSeesTheNewValue(@TempDir Path dir) throws Exception {
|
||||
Path f = dir.resolve("bridged.yaml");
|
||||
Files.writeString(f, yaml("placement: weighted\n"));
|
||||
ConfigRef ref = refFor(f);
|
||||
java.util.function.Supplier<String> reader = () -> ref.get().placement();
|
||||
assertEquals("weighted", reader.get());
|
||||
|
||||
Files.writeString(f, yaml("placement: fixed\n"));
|
||||
assertTrue(ref.reload().applied());
|
||||
|
||||
assertEquals("fixed", reader.get());
|
||||
}
|
||||
|
||||
@Test
|
||||
void aChangedColdKeyRefusesTheWholeReload(@TempDir Path dir) throws Exception {
|
||||
Path f = dir.resolve("bridged.yaml");
|
||||
Files.writeString(f, yaml("placement: weighted\n"));
|
||||
ConfigRef ref = refFor(f);
|
||||
|
||||
// Two changes in one file: a cold one (the port) and a hot one (placement).
|
||||
Files.writeString(f, yaml("placement: fixed\n").replace("port: 8765", "port: 9999"));
|
||||
ConfigRef.Outcome out = ref.reload();
|
||||
|
||||
assertFalse(out.applied());
|
||||
assertEquals(java.util.List.of("bind"), out.coldKeys());
|
||||
assertTrue(out.summary().contains("Restart bridged"), out.summary());
|
||||
// The hot half must NOT have leaked in. A half-applied reload leaves the daemon matching no
|
||||
// file on disk, which is worse for an operator than no reload at all.
|
||||
assertEquals("weighted", ref.get().placement());
|
||||
assertEquals(8765, ref.get().bind().port());
|
||||
}
|
||||
|
||||
@Test
|
||||
void aFileThatNoLongerParsesKeepsTheRunningConfig(@TempDir Path dir) throws Exception {
|
||||
Path f = dir.resolve("bridged.yaml");
|
||||
Files.writeString(f, yaml("placement: weighted\n"));
|
||||
ConfigRef ref = refFor(f);
|
||||
BridgedConfig before = ref.get();
|
||||
|
||||
Files.writeString(f, "profiles:\n sonnet:\n baseUrl: \"unclosed\n");
|
||||
ConfigRef.Outcome out = ref.reload();
|
||||
|
||||
assertFalse(out.applied());
|
||||
assertNotNull(out.error());
|
||||
assertTrue(out.summary().startsWith("config reload refused"), out.summary());
|
||||
assertSame(before, ref.get());
|
||||
}
|
||||
|
||||
/** A file that would have refused to boot must not be able to slip in through a reload. */
|
||||
@Test
|
||||
void aFileThatFailsAValidatorKeepsTheRunningConfig(@TempDir Path dir) throws Exception {
|
||||
Path f = dir.resolve("bridged.yaml");
|
||||
Files.writeString(f, yaml("placement: weighted\n"));
|
||||
ConfigRef ref = refFor(f);
|
||||
BridgedConfig before = ref.get();
|
||||
|
||||
// A member slot naming a profile that does not exist — validateMembers refuses this at
|
||||
// startup, so it must refuse it here too.
|
||||
Files.writeString(f, yaml("""
|
||||
fleet:
|
||||
developers:
|
||||
a:
|
||||
profile: no-such-profile
|
||||
"""));
|
||||
ConfigRef.Outcome out = ref.reload();
|
||||
|
||||
assertFalse(out.applied());
|
||||
assertNotNull(out.error());
|
||||
assertSame(before, ref.get());
|
||||
}
|
||||
|
||||
@Test
|
||||
void aDeletedFileIsRefusedRatherThanCrashing(@TempDir Path dir) throws Exception {
|
||||
Path f = dir.resolve("bridged.yaml");
|
||||
Files.writeString(f, yaml(""));
|
||||
ConfigRef ref = refFor(f);
|
||||
BridgedConfig before = ref.get();
|
||||
|
||||
Files.delete(f);
|
||||
ConfigRef.Outcome out = ref.reload();
|
||||
|
||||
assertFalse(out.applied());
|
||||
assertNotNull(out.error());
|
||||
assertSame(before, ref.get());
|
||||
}
|
||||
|
||||
/** A deferred change applies to the snapshot but the operator is told it needs a restart. */
|
||||
@Test
|
||||
void aDeferredChangeIsAppliedAndReported(@TempDir Path dir) throws Exception {
|
||||
Path f = dir.resolve("bridged.yaml");
|
||||
Files.writeString(f, yaml("""
|
||||
lifecycle:
|
||||
drainTimeoutSeconds: 30
|
||||
"""));
|
||||
ConfigRef ref = refFor(f);
|
||||
|
||||
Files.writeString(f, yaml("""
|
||||
lifecycle:
|
||||
drainTimeoutSeconds: 60
|
||||
"""));
|
||||
ConfigRef.Outcome out = ref.reload();
|
||||
|
||||
assertTrue(out.applied());
|
||||
assertEquals(java.util.List.of("lifecycle"), out.deferred());
|
||||
assertTrue(out.summary().contains("needs a restart") || out.summary().contains("need a restart"),
|
||||
out.summary());
|
||||
assertEquals(60, ref.get().lifecycle().drainTimeoutSeconds());
|
||||
}
|
||||
|
||||
/**
|
||||
* Adding a profile is deferred, not hot: a new backend needs its own launcher, and launchers are
|
||||
* built once at startup. The snapshot carries it so a restart picks it up.
|
||||
*/
|
||||
@Test
|
||||
void addingAProfileIsReportedAsDeferred(@TempDir Path dir) throws Exception {
|
||||
Path f = dir.resolve("bridged.yaml");
|
||||
Files.writeString(f, yaml(""));
|
||||
ConfigRef ref = refFor(f);
|
||||
|
||||
Files.writeString(f, """
|
||||
bind:
|
||||
host: 127.0.0.1
|
||||
port: 8765
|
||||
herdrSocket: ~/.config/herdr/herdr.sock
|
||||
profiles:
|
||||
sonnet:
|
||||
baseUrl: http://gx00.gw:8000
|
||||
model: sonnet
|
||||
haiku:
|
||||
baseUrl: http://gx00.gw:8000
|
||||
model: haiku
|
||||
guard:
|
||||
offSubscriptionHosts:
|
||||
- gx00.gw
|
||||
""");
|
||||
ConfigRef.Outcome out = ref.reload();
|
||||
|
||||
assertTrue(out.applied());
|
||||
assertEquals(1, out.deferred().size());
|
||||
assertTrue(out.deferred().getFirst().contains("haiku"), out.deferred().toString());
|
||||
}
|
||||
|
||||
/**
|
||||
* Changing an existing profile's weight or maxLoad IS hot: placement reads those live through
|
||||
* the supplier on the composite, so the next spawn already sees them.
|
||||
*/
|
||||
@Test
|
||||
void changingAProfilesWeightOrMaxLoadIsHot(@TempDir Path dir) throws Exception {
|
||||
Path f = dir.resolve("bridged.yaml");
|
||||
Files.writeString(f, yaml(""));
|
||||
ConfigRef ref = refFor(f);
|
||||
|
||||
Files.writeString(f, """
|
||||
bind:
|
||||
host: 127.0.0.1
|
||||
port: 8765
|
||||
herdrSocket: ~/.config/herdr/herdr.sock
|
||||
profiles:
|
||||
sonnet:
|
||||
baseUrl: http://gx00.gw:8000
|
||||
model: sonnet
|
||||
maxLoad: 7
|
||||
weight: 3.0
|
||||
guard:
|
||||
offSubscriptionHosts:
|
||||
- gx00.gw
|
||||
""");
|
||||
ConfigRef.Outcome out = ref.reload();
|
||||
|
||||
assertTrue(out.applied());
|
||||
assertTrue(out.deferred().isEmpty(), out.deferred().toString());
|
||||
assertEquals(7, ref.get().profiles().get("sonnet").maxLoad());
|
||||
}
|
||||
|
||||
/**
|
||||
* Changing an existing profile's MODEL is deferred, not hot — and saying so is the whole point.
|
||||
* HerdrPeerLauncher takes Map.copyOf(profiles) at construction and resolves every spawn out of
|
||||
* that copy, so a reloaded model never reaches a launch. Reporting it as applied would be the
|
||||
* worst outcome a reload can produce: the operator has no reason to doubt a clean "reloaded".
|
||||
*/
|
||||
@Test
|
||||
void changingAProfilesLaunchSettingsIsReportedAsDeferred(@TempDir Path dir) throws Exception {
|
||||
Path f = dir.resolve("bridged.yaml");
|
||||
Files.writeString(f, yaml(""));
|
||||
ConfigRef ref = refFor(f);
|
||||
|
||||
Files.writeString(f, """
|
||||
bind:
|
||||
host: 127.0.0.1
|
||||
port: 8765
|
||||
herdrSocket: ~/.config/herdr/herdr.sock
|
||||
profiles:
|
||||
sonnet:
|
||||
baseUrl: http://gx00.gw:8000
|
||||
model: deepseek-v4-flash
|
||||
guard:
|
||||
offSubscriptionHosts:
|
||||
- gx00.gw
|
||||
""");
|
||||
ConfigRef.Outcome out = ref.reload();
|
||||
|
||||
assertTrue(out.applied());
|
||||
assertEquals(1, out.deferred().size(), out.deferred().toString());
|
||||
assertTrue(out.deferred().getFirst().contains("sonnet"), out.deferred().toString());
|
||||
assertTrue(out.deferred().getFirst().contains("launch settings"), out.deferred().toString());
|
||||
// The snapshot still carries the new value — a restart is what makes it take effect.
|
||||
assertEquals("deepseek-v4-flash", ref.get().profiles().get("sonnet").model());
|
||||
}
|
||||
|
||||
@Test
|
||||
void aFixedRefHasNoFileAndRefusesToReload() {
|
||||
BridgedConfig cfg = new BridgedConfig(null, null, null, null, null, null,
|
||||
null, null, null, null, null, null, null, null, null).withDefaults();
|
||||
ConfigRef ref = ConfigRef.fixed(cfg);
|
||||
|
||||
assertNull(ref.path());
|
||||
assertSame(cfg, ref.get());
|
||||
ConfigRef.Outcome out = ref.reload();
|
||||
assertFalse(out.applied());
|
||||
assertNotNull(out.error());
|
||||
}
|
||||
}
|
||||
@@ -0,0 +1,162 @@
|
||||
package dev.ltms.bridged.config;
|
||||
|
||||
import org.junit.jupiter.api.Test;
|
||||
import org.junit.jupiter.api.io.TempDir;
|
||||
|
||||
import java.nio.file.Files;
|
||||
import java.nio.file.Path;
|
||||
import java.nio.file.attribute.FileTime;
|
||||
|
||||
import static org.junit.jupiter.api.Assertions.*;
|
||||
|
||||
/**
|
||||
* CB-559: the mtime poller behind {@code configReload:}. The tests drive {@link
|
||||
* ConfigWatcher#tick()} directly rather than the scheduler, so nothing here sleeps.
|
||||
*/
|
||||
class ConfigWatcherTest {
|
||||
|
||||
private static String yaml(String placement) {
|
||||
return """
|
||||
bind:
|
||||
host: 127.0.0.1
|
||||
port: 8765
|
||||
herdrSocket: ~/.config/herdr/herdr.sock
|
||||
profiles:
|
||||
sonnet:
|
||||
baseUrl: http://gx00.gw:8000
|
||||
model: sonnet
|
||||
guard:
|
||||
offSubscriptionHosts:
|
||||
- gx00.gw
|
||||
""" + "placement: " + placement + "\n";
|
||||
}
|
||||
|
||||
/** Write and stamp an mtime, so a test never depends on the filesystem's clock resolution. */
|
||||
private static void write(Path f, String content, long millis) throws Exception {
|
||||
Files.writeString(f, content);
|
||||
Files.setLastModifiedTime(f, FileTime.fromMillis(millis));
|
||||
}
|
||||
|
||||
@Test
|
||||
void aChangedMtimeTriggersAReload(@TempDir Path dir) throws Exception {
|
||||
Path f = dir.resolve("bridged.yaml");
|
||||
write(f, yaml("weighted"), 1_000L);
|
||||
ConfigRef ref = new ConfigRef(f, BridgedConfig.load(f));
|
||||
|
||||
ConfigWatcher watcher = new ConfigWatcher(ref, 10);
|
||||
try {
|
||||
write(f, yaml("fixed"), 2_000L);
|
||||
watcher.tick();
|
||||
|
||||
assertEquals("fixed", ref.get().placement());
|
||||
} finally {
|
||||
watcher.stop();
|
||||
}
|
||||
}
|
||||
|
||||
@Test
|
||||
void anUnchangedFileIsNotReloaded(@TempDir Path dir) throws Exception {
|
||||
Path f = dir.resolve("bridged.yaml");
|
||||
write(f, yaml("weighted"), 1_000L);
|
||||
ConfigRef ref = new ConfigRef(f, BridgedConfig.load(f));
|
||||
BridgedConfig before = ref.get();
|
||||
|
||||
ConfigWatcher watcher = new ConfigWatcher(ref, 10);
|
||||
try {
|
||||
watcher.tick();
|
||||
watcher.tick();
|
||||
|
||||
// Same instance, so no reload happened — a reload always swaps in a fresh object.
|
||||
assertSame(before, ref.get());
|
||||
} finally {
|
||||
watcher.stop();
|
||||
}
|
||||
}
|
||||
|
||||
/**
|
||||
* A refused reload must not be retried every tick. Without the stamp-before-reload order the
|
||||
* same refusal would be logged forever, which buries the log an operator needs.
|
||||
*/
|
||||
@Test
|
||||
void aRefusedReloadIsNotRetriedUntilTheFileChangesAgain(@TempDir Path dir) throws Exception {
|
||||
Path f = dir.resolve("bridged.yaml");
|
||||
write(f, yaml("weighted"), 1_000L);
|
||||
ConfigRef ref = new ConfigRef(f, BridgedConfig.load(f));
|
||||
BridgedConfig before = ref.get();
|
||||
|
||||
ConfigWatcher watcher = new ConfigWatcher(ref, 10);
|
||||
try {
|
||||
// A cold change: refused, and the running config stays.
|
||||
write(f, yaml("fixed").replace("port: 8765", "port: 9999"), 2_000L);
|
||||
watcher.tick();
|
||||
assertSame(before, ref.get());
|
||||
|
||||
// The next tick sees the same mtime, so it does nothing at all.
|
||||
watcher.tick();
|
||||
assertSame(before, ref.get());
|
||||
|
||||
// A fresh save earns a fresh attempt — and this one is hot, so it applies.
|
||||
write(f, yaml("fixed"), 3_000L);
|
||||
watcher.tick();
|
||||
assertEquals("fixed", ref.get().placement());
|
||||
} finally {
|
||||
watcher.stop();
|
||||
}
|
||||
}
|
||||
|
||||
/**
|
||||
* Editors briefly unlink the file mid-save. A missing file is skipped, not an error and not a
|
||||
* reload — the running config stays live, which is right either way.
|
||||
*/
|
||||
@Test
|
||||
void aMissingFileIsSkippedAndTheNextTickTriesAgain(@TempDir Path dir) throws Exception {
|
||||
Path f = dir.resolve("bridged.yaml");
|
||||
write(f, yaml("weighted"), 1_000L);
|
||||
ConfigRef ref = new ConfigRef(f, BridgedConfig.load(f));
|
||||
BridgedConfig before = ref.get();
|
||||
|
||||
ConfigWatcher watcher = new ConfigWatcher(ref, 10);
|
||||
try {
|
||||
Files.delete(f);
|
||||
assertDoesNotThrow(watcher::tick);
|
||||
assertSame(before, ref.get());
|
||||
|
||||
write(f, yaml("fixed"), 2_000L);
|
||||
watcher.tick();
|
||||
assertEquals("fixed", ref.get().placement());
|
||||
} finally {
|
||||
watcher.stop();
|
||||
}
|
||||
}
|
||||
|
||||
/** A ref with no file behind it must not start a scheduler that could never do anything. */
|
||||
@Test
|
||||
void aFixedRefStartsNoWatch() {
|
||||
BridgedConfig cfg = new BridgedConfig(null, null, null, null, null, null,
|
||||
null, null, null, null, null, null, null, null, null).withDefaults();
|
||||
ConfigWatcher watcher = new ConfigWatcher(ConfigRef.fixed(cfg), 10);
|
||||
try {
|
||||
assertDoesNotThrow(watcher::start);
|
||||
assertDoesNotThrow(watcher::tick);
|
||||
} finally {
|
||||
watcher.stop();
|
||||
}
|
||||
}
|
||||
|
||||
@Test
|
||||
void configReloadIsOffUnlessTheBlockSaysOtherwise(@TempDir Path dir) throws Exception {
|
||||
Path f = dir.resolve("bridged.yaml");
|
||||
write(f, yaml("weighted"), 1_000L);
|
||||
assertNull(BridgedConfig.load(f).configReload());
|
||||
|
||||
write(f, yaml("weighted") + "configReload:\n enabled: true\n", 2_000L);
|
||||
BridgedConfig.ConfigReload on = BridgedConfig.load(f).configReload();
|
||||
assertTrue(on.isEnabled());
|
||||
assertEquals(10, on.intervalSeconds(), "an absent interval defaults to 10s");
|
||||
|
||||
write(f, yaml("weighted") + "configReload:\n intervalSeconds: 30\n", 3_000L);
|
||||
BridgedConfig.ConfigReload off = BridgedConfig.load(f).configReload();
|
||||
assertFalse(off.isEnabled(), "a block that only sets the interval does not enable the watch");
|
||||
assertEquals(30, off.intervalSeconds());
|
||||
}
|
||||
}
|
||||
@@ -0,0 +1,81 @@
|
||||
package dev.ltms.bridged.health;
|
||||
|
||||
import ch.qos.logback.classic.Logger;
|
||||
import ch.qos.logback.classic.spi.ILoggingEvent;
|
||||
import ch.qos.logback.core.read.ListAppender;
|
||||
import dev.ltms.bridged.herdr.AgentControl;
|
||||
import dev.ltms.bridged.herdr.FakeHerdr;
|
||||
import dev.ltms.bridged.inject.Injector;
|
||||
import dev.ltms.bridged.msg.InMemoryReplyInbox;
|
||||
import dev.ltms.bridged.msg.MessageService;
|
||||
import dev.ltms.bridged.msg.Rendezvous;
|
||||
import dev.ltms.bridged.session.SessionManager;
|
||||
import dev.ltms.bridged.member.ClaudeCodeLauncher;
|
||||
import dev.ltms.bridged.herdr.WorkspaceControl;
|
||||
import dev.ltms.bridged.guard.SubscriptionGuard;
|
||||
import dev.ltms.bridged.config.BridgedConfig;
|
||||
import org.junit.jupiter.api.Test;
|
||||
import org.slf4j.LoggerFactory;
|
||||
|
||||
import java.util.Map;
|
||||
import java.util.Set;
|
||||
import java.util.concurrent.Executors;
|
||||
|
||||
import static org.junit.jupiter.api.Assertions.assertEquals;
|
||||
|
||||
class FleetHealthMonitorTest {
|
||||
@Test void oneTickUsesOneFleetListForAnyRosterSize() {
|
||||
FakeHerdr herdr = new FakeHerdr();
|
||||
AgentControl agents = new AgentControl(herdr);
|
||||
BridgedConfig.Profile profile = new BridgedConfig.Profile("test", "http://test:1", null,
|
||||
null, null, null, null, null, null, null, null, null);
|
||||
ClaudeCodeLauncher launcher = new ClaudeCodeLauncher(agents, new WorkspaceControl(herdr),
|
||||
new SubscriptionGuard(Set.of("test")), Map.of("test", profile), "test", _ -> "token");
|
||||
SessionManager sessions = new SessionManager(launcher);
|
||||
sessions.acquire("test", null, null, null);
|
||||
sessions.acquire("test", null, null, null);
|
||||
herdr.calls.clear();
|
||||
MessageService messages = new MessageService(agents, new Injector(agents), new Rendezvous(), new InMemoryReplyInbox());
|
||||
var scheduler = Executors.newSingleThreadScheduledExecutor();
|
||||
FleetHealthMonitor monitor = new FleetHealthMonitor(agents, sessions::roster, messages, scheduler, () -> 1, 60);
|
||||
monitor.tick();
|
||||
monitor.stop();
|
||||
assertEquals(1, herdr.calls.stream().filter(call -> call.method().equals("agent.list")).count());
|
||||
}
|
||||
|
||||
@Test void failedTickDoesNotStopTheNextTick() {
|
||||
FakeHerdr herdr = new FakeHerdr().healthy(false);
|
||||
AgentControl agents = new AgentControl(herdr);
|
||||
var scheduler = Executors.newSingleThreadScheduledExecutor();
|
||||
FleetHealthMonitor monitor = new FleetHealthMonitor(agents, java.util.List::of,
|
||||
new MessageService(agents, new Injector(agents), new Rendezvous(), new InMemoryReplyInbox()),
|
||||
scheduler, () -> 1, 60);
|
||||
monitor.tick();
|
||||
herdr.healthy(true);
|
||||
monitor.tick();
|
||||
monitor.stop();
|
||||
assertEquals(2, herdr.calls.stream().filter(call -> call.method().equals("agent.list")).count());
|
||||
}
|
||||
|
||||
@Test void faultTransitionLogsOnlyOnceUntilItChanges() {
|
||||
Logger logger = (Logger) LoggerFactory.getLogger(FleetHealthMonitor.class);
|
||||
ListAppender<ILoggingEvent> appender = new ListAppender<>();
|
||||
appender.start();
|
||||
logger.addAppender(appender);
|
||||
try {
|
||||
FakeHerdr herdr = new FakeHerdr();
|
||||
AgentControl agents = new AgentControl(herdr);
|
||||
var scheduler = Executors.newSingleThreadScheduledExecutor();
|
||||
FleetHealthMonitor monitor = new FleetHealthMonitor(agents, java.util.List::of,
|
||||
new MessageService(agents, new Injector(agents), new Rendezvous(), new InMemoryReplyInbox()),
|
||||
scheduler, () -> 1, 60);
|
||||
monitor.reportTransition("term_a", HealthState.TURN_BOUNDARY_LOST);
|
||||
monitor.reportTransition("term_a", HealthState.TURN_BOUNDARY_LOST);
|
||||
monitor.stop();
|
||||
assertEquals(1, appender.list.stream().filter(event -> event.getFormattedMessage()
|
||||
.contains("member=term_a state=TURN_BOUNDARY_LOST")).count());
|
||||
} finally {
|
||||
logger.detachAppender(appender);
|
||||
}
|
||||
}
|
||||
}
|
||||
@@ -0,0 +1,56 @@
|
||||
package dev.ltms.bridged.health;
|
||||
|
||||
import dev.ltms.bridged.herdr.AgentStatus;
|
||||
import dev.ltms.bridged.session.MemberSession;
|
||||
import dev.ltms.bridged.msg.Rendezvous;
|
||||
import org.junit.jupiter.api.Test;
|
||||
|
||||
import static org.junit.jupiter.api.Assertions.assertEquals;
|
||||
|
||||
class FleetHealthTest {
|
||||
@Test void muteCountsOnlyCompletionFallbacks() {
|
||||
MuteCounter mute = new MuteCounter();
|
||||
mute.observe("target", "terra", Rendezvous.Kind.REPLY);
|
||||
mute.observe("target", "terra", Rendezvous.Kind.COMPLETION);
|
||||
assertEquals(1, mute.forTarget("target"));
|
||||
assertEquals(1, mute.forProfile("terra"));
|
||||
}
|
||||
@Test void turnBoundaryNeedsTwoSnapshots() {
|
||||
HealthSnapshot s = snapshot(MemberSession.State.BUSY, AgentStatus.DONE, true);
|
||||
HealthDecision first = FleetHealth.decide(s, HealthPrior.NONE, 1);
|
||||
assertEquals(HealthState.WORKING, first.state());
|
||||
assertEquals(new HealthPrior(true), first.prior());
|
||||
assertEquals(HealthState.TURN_BOUNDARY_LOST, FleetHealth.decide(s, first.prior(), 2).state());
|
||||
}
|
||||
|
||||
@Test void unknownLiveStatusWithAcceptedDeliveryIsNotIdle() {
|
||||
assertEquals(HealthState.WORKING, FleetHealth.decide(
|
||||
snapshot(MemberSession.State.BUSY, AgentStatus.UNKNOWN, true), HealthPrior.NONE, 1).state());
|
||||
}
|
||||
|
||||
@Test void acceptedDeliveryNeverReportsIdle() {
|
||||
for (MemberSession.State session : MemberSession.State.values()) {
|
||||
for (AgentStatus live : AgentStatus.values()) {
|
||||
HealthSnapshot s = snapshot(session, live, true);
|
||||
assertEquals(false, FleetHealth.decide(s, HealthPrior.NONE, 1).state() == HealthState.IDLE,
|
||||
() -> "accepted delivery returned IDLE for " + session + "/" + live);
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
@Test void blockedDoesNotGuessPromptKind() {
|
||||
assertEquals(HealthState.BLOCKED_AMBIGUOUS, FleetHealth.decide(
|
||||
snapshot(MemberSession.State.BUSY, AgentStatus.BLOCKED, true), HealthPrior.NONE, 1).state());
|
||||
}
|
||||
|
||||
@Test void controlLinkOutranksMemberFault() {
|
||||
HealthSnapshot s = new HealthSnapshot(MemberSession.State.BUSY, AgentStatus.DONE, true, false,
|
||||
false, true, true, true, false, true, true, true);
|
||||
assertEquals(HealthState.CONTROL_LINK_DOWN, FleetHealth.decide(s, HealthPrior.NONE, 1).state());
|
||||
}
|
||||
|
||||
private static HealthSnapshot snapshot(MemberSession.State state, AgentStatus live, boolean accepted) {
|
||||
return new HealthSnapshot(state, live, accepted, false, false, true, false, false,
|
||||
false, false, false, false);
|
||||
}
|
||||
}
|
||||
@@ -0,0 +1,14 @@
|
||||
package dev.ltms.bridged.health;
|
||||
|
||||
import org.junit.jupiter.api.Test;
|
||||
import java.util.List;
|
||||
import static org.junit.jupiter.api.Assertions.assertEquals;
|
||||
|
||||
class PaneBudgetTest {
|
||||
@Test void fixedLimitsIgnoreWeakerConfig() {
|
||||
PaneBudget budget = new PaneBudget();
|
||||
List<String> targets = List.of("a", "b", "c");
|
||||
assertEquals(List.of("a", "b"), budget.choose(targets, 0, 0));
|
||||
assertEquals(List.of("c"), budget.choose(targets, 1, 0));
|
||||
}
|
||||
}
|
||||
@@ -4,7 +4,9 @@ import com.fasterxml.jackson.databind.JsonNode;
|
||||
import com.fasterxml.jackson.databind.ObjectMapper;
|
||||
|
||||
import java.util.ArrayList;
|
||||
import java.util.LinkedHashMap;
|
||||
import java.util.List;
|
||||
import java.util.Map;
|
||||
|
||||
/**
|
||||
* Recording fake {@link HerdrClient} for unit/acceptance tests. Returns canned frames
|
||||
@@ -24,6 +26,8 @@ public final class FakeHerdr implements HerdrClient {
|
||||
private boolean healthy = true;
|
||||
private final List<String> extraWorkspaces = new ArrayList<>();
|
||||
private final List<String> extraAgents = new ArrayList<>();
|
||||
/** workspaceId → extra tabs that {@code tab.list} reports for it (CB-558 lead scans). */
|
||||
private final Map<String, List<String>> extraTabs = new LinkedHashMap<>();
|
||||
private int agentNameTakenFor = 0;
|
||||
private int agentPaneBusyFor = 0;
|
||||
private int workerTabPaneCount = 1;
|
||||
@@ -109,6 +113,18 @@ public final class FakeHerdr implements HerdrClient {
|
||||
return this;
|
||||
}
|
||||
|
||||
/**
|
||||
* Seed a labelled tab into {@code tab.list} for one workspace (e.g. an existing {@code lead: x}
|
||||
* tab). Pair it with {@link #withAgent} on the same {@code tabId} to make the lead <em>live</em>;
|
||||
* seeding the tab alone models the stale-label case.
|
||||
*/
|
||||
public FakeHerdr withTab(String workspaceId, String tabId, String label) {
|
||||
extraTabs.computeIfAbsent(workspaceId, _ -> new ArrayList<>())
|
||||
.add(("{\"tab_id\":\"%s\",\"workspace_id\":\"%s\",\"label\":\"%s\",\"pane_count\":1}")
|
||||
.formatted(tabId, workspaceId, label));
|
||||
return this;
|
||||
}
|
||||
|
||||
/** Seed an additional workspace into {@code workspace.list} (e.g. a pre-existing worker space). */
|
||||
public FakeHerdr withWorkspace(String id, String label) {
|
||||
extraWorkspaces.add(("{\"workspace_id\":\"%s\",\"label\":\"%s\",\"focused\":false,"
|
||||
@@ -228,11 +244,19 @@ public final class FakeHerdr implements HerdrClient {
|
||||
case "tab.rename" -> mapper.readTree("""
|
||||
{"type":"tab_info","tab":{"tab_id":"w9:t2","workspace_id":"w9",
|
||||
"label":"worker: ltms-local","pane_count":1}}""");
|
||||
case "tab.list" -> mapper.readTree(("""
|
||||
case "tab.list" -> {
|
||||
// The base pair is returned for every workspace, exactly as before. Seeded tabs
|
||||
// are appended only for the workspace they were registered against, so a test
|
||||
// that seeds none sees the historical response byte for byte.
|
||||
Object wsId = params instanceof java.util.Map<?, ?> m ? m.get("workspace_id") : null;
|
||||
List<String> seeded = extraTabs.getOrDefault(String.valueOf(wsId), List.of());
|
||||
yield mapper.readTree(("""
|
||||
{"type":"tab_list","tabs":[
|
||||
{"tab_id":"w9:t1","workspace_id":"w9","label":"1","pane_count":1},
|
||||
{"tab_id":"w9:t2","workspace_id":"w9","label":"worker: ltms-local","pane_count":%d}]}""")
|
||||
.formatted(workerTabPaneCount));
|
||||
{"tab_id":"w9:t2","workspace_id":"w9","label":"worker: ltms-local","pane_count":%d}%s]}""")
|
||||
.formatted(workerTabPaneCount,
|
||||
seeded.isEmpty() ? "" : "," + String.join(",", seeded)));
|
||||
}
|
||||
case "tab.close" -> mapper.readTree("{\"type\":\"ok\"}");
|
||||
case "pane.get" -> mapper.readTree("""
|
||||
{"type":"pane_info","pane":{"pane_id":"w9:pW","workspace_id":"w9",
|
||||
|
||||
@@ -1,9 +1,14 @@
|
||||
package dev.ltms.bridged.inject;
|
||||
|
||||
import ch.qos.logback.classic.Level;
|
||||
import ch.qos.logback.classic.LoggerContext;
|
||||
import ch.qos.logback.classic.spi.ILoggingEvent;
|
||||
import ch.qos.logback.core.read.ListAppender;
|
||||
import dev.ltms.bridged.herdr.AgentControl;
|
||||
import dev.ltms.bridged.herdr.FakeHerdr;
|
||||
import dev.ltms.bridged.msg.Rendezvous;
|
||||
import org.junit.jupiter.api.Test;
|
||||
import org.slf4j.LoggerFactory;
|
||||
|
||||
import static org.junit.jupiter.api.Assertions.assertEquals;
|
||||
import static org.junit.jupiter.api.Assertions.assertFalse;
|
||||
@@ -156,6 +161,33 @@ class CompletionResolverTest {
|
||||
assertEquals("No, 391 = 17 × 23.", waiter.getNow(null).text());
|
||||
}
|
||||
|
||||
@Test
|
||||
void marksAClippedCompletionPaneTail() {
|
||||
String block = "⏺ " + "x".repeat(CompletionResolver.MAX_SCRAPE_CHARS + 1) + "\n❯ ";
|
||||
FakeHerdr herdr = new FakeHerdr().readText(block);
|
||||
Rendezvous rendezvous = new Rendezvous();
|
||||
CompletionResolver resolver = new CompletionResolver(new AgentControl(herdr), rendezvous);
|
||||
|
||||
var waiter = rendezvous.open("term_a");
|
||||
resolver.resolve("term_a", new CompletionResolver.InFlight(waiter, null));
|
||||
|
||||
assertEquals("x".repeat(CompletionResolver.MAX_SCRAPE_CHARS)
|
||||
+ "\n[Pane tail clipped: member did not call bridge_reply.]",
|
||||
waiter.getNow(null).text());
|
||||
}
|
||||
|
||||
@Test
|
||||
void leavesAnUnclippedCompletionPaneTailUnmarked() {
|
||||
FakeHerdr herdr = new FakeHerdr().readText("⏺ complete report\n❯ ");
|
||||
Rendezvous rendezvous = new Rendezvous();
|
||||
CompletionResolver resolver = new CompletionResolver(new AgentControl(herdr), rendezvous);
|
||||
|
||||
var waiter = rendezvous.open("term_a");
|
||||
resolver.resolve("term_a", new CompletionResolver.InFlight(waiter, null));
|
||||
|
||||
assertEquals("complete report", waiter.getNow(null).text());
|
||||
}
|
||||
|
||||
@Test
|
||||
void resolvesSynchronouslyBeforePostTurnContextClearing() {
|
||||
FakeHerdr herdr = new FakeHerdr().readText("⏺ previous answer\n❯ ");
|
||||
@@ -177,7 +209,8 @@ class CompletionResolverTest {
|
||||
// block while resolve compares against a clip()'d tail. For a block longer than MAX_SCRAPE_CHARS
|
||||
// the two capped representations differ even when the pane never changed, so the CB-115
|
||||
// byte-identical guard failed to fire and a stale completion could resolve the send. Both sides
|
||||
// must clip identically; here an unchanged >cap block on rapid back-to-back turns stays suppressed.
|
||||
// must clip identically. The returned-text marker is added only after this comparison, so an
|
||||
// unchanged >cap block on rapid back-to-back turns still stays suppressed.
|
||||
String longBlock = "⏺ " + "x".repeat(CompletionResolver.MAX_SCRAPE_CHARS + 500) + "\n❯ ";
|
||||
FakeHerdr herdr = new FakeHerdr().readText(longBlock);
|
||||
Rendezvous rendezvous = new Rendezvous();
|
||||
@@ -270,6 +303,40 @@ class CompletionResolverTest {
|
||||
assertEquals("stuck on an error screen", waiter.getNow(null).text());
|
||||
}
|
||||
|
||||
@Test
|
||||
void failIsLoggedAtWarnWithTheReason() {
|
||||
// CB-564: this used to be a bare DEBUG "failed send to X via turn-stall fallback" — a symptom
|
||||
// with no cause, and below the level anyone watching for member health would see. A fail that
|
||||
// resolves a caller's blocked send is at least WARN and must carry the reason.
|
||||
LoggerContext ctx = (LoggerContext) LoggerFactory.getILoggerFactory();
|
||||
ch.qos.logback.classic.Logger resolverLog =
|
||||
(ch.qos.logback.classic.Logger) LoggerFactory.getLogger(CompletionResolver.class);
|
||||
ListAppender<ILoggingEvent> appender = new ListAppender<>();
|
||||
appender.setContext(ctx);
|
||||
appender.start();
|
||||
resolverLog.addAppender(appender);
|
||||
resolverLog.setLevel(Level.WARN);
|
||||
try {
|
||||
FakeHerdr herdr = new FakeHerdr().readText("stuck on an error screen");
|
||||
Rendezvous rendezvous = new Rendezvous();
|
||||
CompletionResolver resolver = new CompletionResolver(new AgentControl(herdr), rendezvous);
|
||||
var waiter = rendezvous.open("term_a");
|
||||
|
||||
resolver.fail("term_a", null);
|
||||
|
||||
String warn = appender.list.stream()
|
||||
.filter(e -> e.getLevel().equals(Level.WARN))
|
||||
.map(ILoggingEvent::getFormattedMessage)
|
||||
.findFirst()
|
||||
.orElse("no turn-stall WARN logged");
|
||||
assertTrue(warn.contains("term_a"), "the log names the target: " + warn);
|
||||
assertTrue(warn.contains("stuck on an error screen"), "the log carries the reason: " + warn);
|
||||
assertTrue(waiter.isDone());
|
||||
} finally {
|
||||
resolverLog.detachAppender(appender);
|
||||
}
|
||||
}
|
||||
|
||||
// --- CB-116 waiter identity: a late completion never crosses into the next turn ---------
|
||||
|
||||
@Test
|
||||
|
||||
@@ -1,10 +1,15 @@
|
||||
package dev.ltms.bridged.inject;
|
||||
|
||||
import ch.qos.logback.classic.Level;
|
||||
import ch.qos.logback.classic.LoggerContext;
|
||||
import ch.qos.logback.classic.spi.ILoggingEvent;
|
||||
import ch.qos.logback.core.read.ListAppender;
|
||||
import dev.ltms.bridged.herdr.AgentControl;
|
||||
import dev.ltms.bridged.herdr.AgentStatus;
|
||||
import dev.ltms.bridged.herdr.FakeHerdr;
|
||||
import dev.ltms.bridged.herdr.HerdrException;
|
||||
import org.junit.jupiter.api.Test;
|
||||
import org.slf4j.LoggerFactory;
|
||||
|
||||
import java.util.ArrayList;
|
||||
import java.util.List;
|
||||
@@ -254,6 +259,7 @@ class InjectorTest {
|
||||
private static final class Captor implements TurnListener {
|
||||
final List<String> completed = new ArrayList<>();
|
||||
final List<String> failed = new ArrayList<>();
|
||||
final List<String> failureReasons = new ArrayList<>();
|
||||
|
||||
@Override
|
||||
public void onTurnComplete(String target) {
|
||||
@@ -264,6 +270,12 @@ class InjectorTest {
|
||||
public void onTurnFailed(String target) {
|
||||
failed.add(target);
|
||||
}
|
||||
|
||||
@Override
|
||||
public void onTurnFailed(String target, String reason) {
|
||||
failed.add(target);
|
||||
failureReasons.add(reason);
|
||||
}
|
||||
}
|
||||
|
||||
// ~30s of unknown at the 250ms prod poll interval; enough onStatus samples to trip the stall.
|
||||
@@ -317,6 +329,23 @@ class InjectorTest {
|
||||
assertTrue(f.isCompletedExceptionally(), "queued waiters unblock when the worker vanishes");
|
||||
}
|
||||
|
||||
@Test
|
||||
void dropPassesTheRealCauseForQueuedAndDeliveredWork() {
|
||||
Captor cap = new Captor();
|
||||
Injector inj = new Injector(new AgentControl(herdr), cap);
|
||||
CompletableFuture<Void> delivered = inj.enqueue(T, "delivered");
|
||||
CompletableFuture<Void> queued = inj.enqueue(T, "queued");
|
||||
|
||||
inj.onStatus(T, AgentStatus.IDLE); // deliver the first message
|
||||
inj.onStatus(T, AgentStatus.WORKING); // its turn is now in flight; one remains queued
|
||||
inj.drop(T, new HerdrException("agent target sol not found", "agent_not_found", null));
|
||||
|
||||
assertEquals(List.of(T), cap.failed, "drop signals one turn failure for both affected states");
|
||||
assertEquals(List.of("agent target sol not found"), cap.failureReasons);
|
||||
assertTrue(delivered.isDone(), "the delivered future has already completed");
|
||||
assertTrue(queued.isCompletedExceptionally(), "the queued future fails with the drop cause");
|
||||
}
|
||||
|
||||
@Test
|
||||
void dropFailsTheTurnOfADeliveredMessageWhenTheWorkerVanishes() {
|
||||
// CB-110: the message was delivered (no longer queued), so failing queued waiters alone would
|
||||
@@ -369,6 +398,40 @@ class InjectorTest {
|
||||
assertTrue(inj.activeTargets().isEmpty(), "the target is reclaimed, not polled forever");
|
||||
}
|
||||
|
||||
@Test
|
||||
void readinessGraceExpiryIsLogged() {
|
||||
// CB-562: the grace-expiry path used to clear the queue silently, so a message that never
|
||||
// reached the worker's pane surfaced elsewhere as an unrelated turn-stall failure. Assert the
|
||||
// expiry now names the real cause. (ListAppender capture pattern mirrors AuditLogTest.)
|
||||
LoggerContext ctx = (LoggerContext) LoggerFactory.getILoggerFactory();
|
||||
ch.qos.logback.classic.Logger injectorLog =
|
||||
(ch.qos.logback.classic.Logger) LoggerFactory.getLogger(Injector.class);
|
||||
ListAppender<ILoggingEvent> appender = new ListAppender<>();
|
||||
appender.setContext(ctx);
|
||||
appender.start();
|
||||
injectorLog.addAppender(appender);
|
||||
injectorLog.setLevel(Level.WARN);
|
||||
try {
|
||||
Injector inj = new Injector(new AgentControl(herdr), TurnListener.NOOP, _ -> false, _ -> {
|
||||
});
|
||||
inj.enqueue(T, "task");
|
||||
|
||||
for (int i = 0; i < READINESS_SAMPLES; i++) inj.onStatus(T, AgentStatus.IDLE);
|
||||
|
||||
String warn = appender.list.stream()
|
||||
.filter(e -> e.getLevel().equals(Level.WARN))
|
||||
.map(ILoggingEvent::getFormattedMessage)
|
||||
.findFirst()
|
||||
.orElse("no grace-expiry WARN logged");
|
||||
assertTrue(warn.contains(T), "the log names the target terminal: " + warn);
|
||||
assertTrue(warn.contains("never reached"), "the log names the real cause: " + warn);
|
||||
assertTrue(warn.contains("1 queued message"),
|
||||
"the log carries the failed message count: " + warn);
|
||||
} finally {
|
||||
injectorLog.detachAppender(appender);
|
||||
}
|
||||
}
|
||||
|
||||
@Test
|
||||
void aWorkerThatBecomesReadyWithinTheGraceIsDeliveredNormally() {
|
||||
// The readiness grace must not fail a worker that is merely slow to boot: once it becomes
|
||||
@@ -389,7 +452,7 @@ class InjectorTest {
|
||||
@Test
|
||||
void dropClearsWorkerPresence() {
|
||||
// CB-114 (finding #1): a vanished worker's readiness must be forgotten so a stale entry cannot
|
||||
// linger past the worker's life (WorkerPresence.forget had no caller before this).
|
||||
// linger past the worker's life (MemberPresence.forget had no caller before this).
|
||||
List<String> forgotten = new ArrayList<>();
|
||||
Injector inj = new Injector(new AgentControl(herdr), TurnListener.NOOP, _ -> true, forgotten::add);
|
||||
inj.enqueue(T, "orphan");
|
||||
@@ -397,6 +460,38 @@ class InjectorTest {
|
||||
assertEquals(List.of(T), forgotten, "drop clears the gone worker's presence");
|
||||
}
|
||||
|
||||
@Test
|
||||
void dropIsLogged() {
|
||||
// CB-564: a vanished worker used to drop its queue with no log at all — the only trace was
|
||||
// whatever failed downstream (e.g. a caller's send timing out with no clue why). Assert the
|
||||
// drop itself now names the cause and the number of messages it failed.
|
||||
LoggerContext ctx = (LoggerContext) LoggerFactory.getILoggerFactory();
|
||||
ch.qos.logback.classic.Logger injectorLog =
|
||||
(ch.qos.logback.classic.Logger) LoggerFactory.getLogger(Injector.class);
|
||||
ListAppender<ILoggingEvent> appender = new ListAppender<>();
|
||||
appender.setContext(ctx);
|
||||
appender.start();
|
||||
injectorLog.addAppender(appender);
|
||||
injectorLog.setLevel(Level.WARN);
|
||||
try {
|
||||
Injector inj = new Injector(new AgentControl(herdr), TurnListener.NOOP, _ -> true, _ -> {
|
||||
});
|
||||
inj.enqueue(T, "orphan");
|
||||
inj.drop(T, new HerdrException("worker gone", "pane_not_found", null));
|
||||
|
||||
String warn = appender.list.stream()
|
||||
.filter(e -> e.getLevel().equals(Level.WARN))
|
||||
.map(ILoggingEvent::getFormattedMessage)
|
||||
.findFirst()
|
||||
.orElse("no drop WARN logged");
|
||||
assertTrue(warn.contains(T), "the log names the target terminal: " + warn);
|
||||
assertTrue(warn.contains("1 message"), "the log carries the failed message count: " + warn);
|
||||
assertTrue(warn.contains("worker gone"), "the log carries the real cause: " + warn);
|
||||
} finally {
|
||||
injectorLog.detachAppender(appender);
|
||||
}
|
||||
}
|
||||
|
||||
@Test
|
||||
void pollerDeliversToAnIdleWorker() throws Exception {
|
||||
// End-to-end through the poller: idle worker → message delivered without manual onStatus.
|
||||
|
||||
+2
-2
@@ -11,7 +11,7 @@ class WorkerPresenceTest {
|
||||
|
||||
@Test
|
||||
void tracksPresenceAndForgets() {
|
||||
WorkerPresence p = new WorkerPresence();
|
||||
MemberPresence p = new MemberPresence();
|
||||
assertFalse(p.isPresent("term_a"), "unseen worker is not available");
|
||||
|
||||
p.markPresent("term_a");
|
||||
@@ -23,7 +23,7 @@ class WorkerPresenceTest {
|
||||
|
||||
@Test
|
||||
void nullOrBlankMarkIsANoOp() {
|
||||
WorkerPresence p = new WorkerPresence();
|
||||
MemberPresence p = new MemberPresence();
|
||||
assertDoesNotThrow(() -> {
|
||||
p.markPresent(null);
|
||||
p.markPresent(" ");
|
||||
@@ -0,0 +1,255 @@
|
||||
package dev.ltms.bridged.lead;
|
||||
|
||||
import dev.ltms.bridged.config.BridgedConfig;
|
||||
import dev.ltms.bridged.herdr.AgentControl;
|
||||
import dev.ltms.bridged.herdr.FakeHerdr;
|
||||
import dev.ltms.bridged.herdr.WorkspaceControl;
|
||||
import org.junit.jupiter.api.Test;
|
||||
|
||||
import java.util.LinkedHashMap;
|
||||
import java.util.List;
|
||||
import java.util.Map;
|
||||
|
||||
import static org.junit.jupiter.api.Assertions.*;
|
||||
|
||||
/**
|
||||
* CB-558 — the daemon starts a declared lead when none is running.
|
||||
*
|
||||
* <p>Two properties carry the whole feature. It must not double-spawn (a second orchestrator is
|
||||
* worse than none), and what it starts must be a <em>lead</em> and not a member: no reply charter,
|
||||
* no off-subscription env, and never registered with the session lifecycle.
|
||||
*/
|
||||
class LeadLauncherTest {
|
||||
|
||||
/** The lead's backend: a subscription profile with the bridge mounted, as `opus` really is. */
|
||||
private static BridgedConfig.Profile opusProfile() {
|
||||
return new BridgedConfig.Profile(
|
||||
"opus", null, "claude-opus-5", null, "BRIDGED_WORKER_TOKEN",
|
||||
List.of("ccs", "ltms"), "tab", "bridged-workers", null,
|
||||
"http://127.0.0.1:8765/mcp", null, null,
|
||||
null, null, null,
|
||||
Map.of("CLAUDE_CODE_AUTO_COMPACT_WINDOW", "300000"), null, null, true);
|
||||
}
|
||||
|
||||
private static BridgedConfig configWith(BridgedConfig.Leader lead) {
|
||||
Map<String, BridgedConfig.Leader> leaders = new LinkedHashMap<>();
|
||||
leaders.put("opus", lead);
|
||||
BridgedConfig.Fleet fleet =
|
||||
new BridgedConfig.Fleet(leaders, Map.of(), Map.of(), Map.of(), null);
|
||||
return new BridgedConfig(
|
||||
null, null, Map.of("opus", opusProfile()), null, null, null, null, null,
|
||||
null, null, fleet, null, "fixed", null).withDefaults();
|
||||
}
|
||||
|
||||
private static BridgedConfig.Leader lead(String profile, String terminal, int instances) {
|
||||
return new BridgedConfig.Leader(profile, terminal, instances, "lead:", 10, null, null,
|
||||
"leads", "/repo");
|
||||
}
|
||||
|
||||
private static LeadLauncher launcher(FakeHerdr herdr, BridgedConfig cfg) {
|
||||
return new LeadLauncher(new AgentControl(herdr), new WorkspaceControl(herdr), cfg);
|
||||
}
|
||||
|
||||
@SuppressWarnings("unchecked")
|
||||
private static List<String> startedArgs(FakeHerdr herdr) {
|
||||
return (List<String>) ((Map<String, Object>) herdr.lastCall("agent.start").params()).get("args");
|
||||
}
|
||||
|
||||
private static String startedName(FakeHerdr herdr) {
|
||||
return (String) ((Map<?, ?>) herdr.lastCall("agent.start").params()).get("name");
|
||||
}
|
||||
|
||||
@SuppressWarnings("unchecked")
|
||||
private static Map<String, String> tabEnv(FakeHerdr herdr) {
|
||||
return (Map<String, String>) ((Map<String, Object>) herdr.lastCall("tab.create").params()).get("env");
|
||||
}
|
||||
|
||||
// ── it starts a lead when none is live ────────────────────────────────────────────────────
|
||||
|
||||
@Test
|
||||
void startsTheDeclaredLeadWhenNoneIsRunning() {
|
||||
FakeHerdr herdr = new FakeHerdr();
|
||||
|
||||
assertEquals(1, launcher(herdr, configWith(lead("opus", null, 1))).ensureLeads());
|
||||
assertTrue(herdr.called("agent.start"), "a lead must actually be started");
|
||||
assertEquals("lead-opus", startedName(herdr));
|
||||
}
|
||||
|
||||
/** The tab is labelled so the scanner finds the lead on the next resolve. */
|
||||
@Test
|
||||
void labelsTheTabWithThePrefixTheScannerReadsBack() {
|
||||
FakeHerdr herdr = new FakeHerdr();
|
||||
launcher(herdr, configWith(lead("opus", null, 1))).ensureLeads();
|
||||
|
||||
assertEquals("lead: opus",
|
||||
((Map<?, ?>) herdr.lastCall("tab.rename").params()).get("label"));
|
||||
}
|
||||
|
||||
/** `instances: 2` with none live means two starts, not one. */
|
||||
@Test
|
||||
void startsAsManyInstancesAsAreDeclared() {
|
||||
FakeHerdr herdr = new FakeHerdr();
|
||||
|
||||
assertEquals(2, launcher(herdr, configWith(lead("opus", null, 2))).ensureLeads());
|
||||
assertEquals(2, herdr.calls.stream().filter(c -> c.method().equals("agent.start")).count());
|
||||
}
|
||||
|
||||
// ── it must not double-spawn ──────────────────────────────────────────────────────────────
|
||||
|
||||
/** A labelled tab WITH a running agent in it is a live lead — leave it alone. */
|
||||
@Test
|
||||
void doesNotStartASecondLeadWhenOneIsAlreadyRunning() {
|
||||
FakeHerdr herdr = new FakeHerdr()
|
||||
.withWorkspace("wL", "leads")
|
||||
.withTab("wL", "wL:t1", "lead: opus")
|
||||
.withAgent("lead-opus", "term_lead", "wL:p1", "wL:t1");
|
||||
|
||||
assertEquals(0, launcher(herdr, configWith(lead("opus", null, 1))).ensureLeads());
|
||||
assertFalse(herdr.called("agent.start"), "the live lead must not be duplicated");
|
||||
}
|
||||
|
||||
/**
|
||||
* The reason liveness is not "does the label exist". A tab left labelled by a session that has
|
||||
* since died must not block the relaunch, or one crash disables auto-launch permanently.
|
||||
*/
|
||||
@Test
|
||||
void aLabelledTabWithNoRunningAgentIsNotALiveLead() {
|
||||
FakeHerdr herdr = new FakeHerdr()
|
||||
.withWorkspace("wL", "leads")
|
||||
.withTab("wL", "wL:t1", "lead: opus"); // label only — nothing running in it
|
||||
|
||||
assertEquals(1, launcher(herdr, configWith(lead("opus", null, 1))).ensureLeads(),
|
||||
"a stale label is not a lead; the lead must be relaunched");
|
||||
}
|
||||
|
||||
/**
|
||||
* A lead the operator opened by hand and pinned with `terminal:` is live even though its tab
|
||||
* carries no matching label. Counting labels alone would relaunch it on every boot.
|
||||
*/
|
||||
@Test
|
||||
void aPinnedTerminalWithARunningAgentCountsAsLive() {
|
||||
FakeHerdr herdr = new FakeHerdr()
|
||||
.withAgent("hand-opened", "term_pinned", "wX:p1", "wX:t1");
|
||||
|
||||
assertEquals(0, launcher(herdr, configWith(lead("opus", "term_pinned", 1))).ensureLeads());
|
||||
assertFalse(herdr.called("agent.start"));
|
||||
}
|
||||
|
||||
/** A member sitting in a matching tab must never be counted — or spawn a lead — as one. */
|
||||
@Test
|
||||
void aMemberWorkspaceIsNeverScannedForLeads() {
|
||||
FakeHerdr herdr = new FakeHerdr()
|
||||
.withWorkspace("wM", "bridged-workers") // a configured member space
|
||||
.withTab("wM", "wM:t1", "lead: opus") // a member tab that looks like a lead
|
||||
.withAgent("claude-opus-x", "term_m", "wM:p1", "wM:t1");
|
||||
|
||||
assertEquals(1, launcher(herdr, configWith(lead("opus", null, 1))).ensureLeads(),
|
||||
"a member in a lead-labelled tab is not a lead, so the real lead is still missing");
|
||||
}
|
||||
|
||||
/** If herdr cannot be counted, start nothing: guessing risks a second orchestrator. */
|
||||
@Test
|
||||
void anUncountableHerdrStartsNothing() {
|
||||
FakeHerdr herdr = new FakeHerdr().healthy(false);
|
||||
|
||||
assertEquals(0, launcher(herdr, configWith(lead("opus", null, 1))).ensureLeads());
|
||||
assertFalse(herdr.called("agent.start"));
|
||||
}
|
||||
|
||||
// ── what it starts is a LEAD, not a member ────────────────────────────────────────────────
|
||||
|
||||
/**
|
||||
* The single most important assertion here. The worker charter tells its reader it is an
|
||||
* off-subscription worker that must end every turn with bridge_reply — the opposite of what an
|
||||
* orchestrator is. A lead must never receive it.
|
||||
*/
|
||||
@Test
|
||||
void theLeadNeverReceivesTheWorkerReplyCharter() {
|
||||
FakeHerdr herdr = new FakeHerdr();
|
||||
launcher(herdr, configWith(lead("opus", null, 1))).ensureLeads();
|
||||
|
||||
List<String> args = startedArgs(herdr);
|
||||
assertFalse(args.contains("--append-system-prompt"),
|
||||
"the reply charter is a worker contract and must not be injected into a lead");
|
||||
assertTrue(args.stream().noneMatch(a -> a.contains("bridge_reply")), args.toString());
|
||||
}
|
||||
|
||||
/** It still mounts the bridge — a lead that cannot orchestrate is pointless. */
|
||||
@Test
|
||||
void theLeadMountsTheBridgeMcpAndPinsItsModel() {
|
||||
FakeHerdr herdr = new FakeHerdr();
|
||||
launcher(herdr, configWith(lead("opus", null, 1))).ensureLeads();
|
||||
|
||||
List<String> args = startedArgs(herdr);
|
||||
assertTrue(args.contains("--mcp-config"));
|
||||
assertTrue(args.stream().anyMatch(a -> a.contains("http://127.0.0.1:8765/mcp")), args.toString());
|
||||
assertEquals("claude-opus-5", args.get(args.indexOf("--model") + 1));
|
||||
assertTrue(args.indexOf("--model") > args.indexOf("--mcp-config"),
|
||||
"--model is appended last so it outranks the ccs wrapper (CB-533)");
|
||||
}
|
||||
|
||||
/** A lead runs on the operator's subscription. Nothing may move it off. */
|
||||
@Test
|
||||
void theLeadEnvCarriesNoAnthropicBinding() {
|
||||
FakeHerdr herdr = new FakeHerdr();
|
||||
launcher(herdr, configWith(lead("opus", null, 1))).ensureLeads();
|
||||
|
||||
Map<String, String> env = tabEnv(herdr);
|
||||
assertNull(env.get("ANTHROPIC_BASE_URL"));
|
||||
assertNull(env.get("ANTHROPIC_AUTH_TOKEN"));
|
||||
assertEquals("300000", env.get("CLAUDE_CODE_AUTO_COMPACT_WINDOW"),
|
||||
"the profile's own env: still applies");
|
||||
}
|
||||
|
||||
/** The lead's tab goes in its own workspace, never a member one — the scanner skips those. */
|
||||
@Test
|
||||
void theLeadTabIsCreatedOutsideEveryMemberWorkspace() {
|
||||
FakeHerdr herdr = new FakeHerdr();
|
||||
launcher(herdr, configWith(lead("opus", null, 1))).ensureLeads();
|
||||
|
||||
String label = (String) ((Map<?, ?>) herdr.lastCall("workspace.create").params()).get("label");
|
||||
assertEquals("leads", label);
|
||||
assertNotEquals("bridged-workers", label);
|
||||
}
|
||||
|
||||
// ── recognise-only and misconfiguration ───────────────────────────────────────────────────
|
||||
|
||||
/** A lead with a pin but no profile is recognise-only by design — not an error, not a launch. */
|
||||
@Test
|
||||
void aLeadThatNamesNoProfileIsRecognisedButNeverLaunched() {
|
||||
FakeHerdr herdr = new FakeHerdr();
|
||||
|
||||
assertEquals(0, launcher(herdr, configWith(lead(null, "term_dead", 1))).ensureLeads());
|
||||
assertFalse(herdr.called("agent.start"));
|
||||
}
|
||||
|
||||
/** `instances: 0` is a deliberate off switch. */
|
||||
@Test
|
||||
void zeroInstancesLaunchesNothing() {
|
||||
FakeHerdr herdr = new FakeHerdr();
|
||||
|
||||
assertEquals(0, launcher(herdr, configWith(lead("opus", null, 0))).ensureLeads());
|
||||
assertFalse(herdr.called("agent.start"));
|
||||
}
|
||||
|
||||
/** A profile name with no matching profile is logged and skipped, never a daemon crash. */
|
||||
@Test
|
||||
void anUnknownProfileIsSkippedRatherThanThrown() {
|
||||
FakeHerdr herdr = new FakeHerdr();
|
||||
|
||||
assertEquals(0, launcher(herdr, configWith(lead("nope", null, 1))).ensureLeads());
|
||||
assertFalse(herdr.called("agent.start"));
|
||||
}
|
||||
|
||||
/** No leads declared at all: not a herdr call in sight. */
|
||||
@Test
|
||||
void noLeadersConfiguredTouchesHerdrNotAtAll() {
|
||||
FakeHerdr herdr = new FakeHerdr();
|
||||
BridgedConfig cfg = new BridgedConfig(
|
||||
null, null, Map.of("opus", opusProfile()), null, null, null, null, null,
|
||||
null, null, null, null, "fixed", null).withDefaults();
|
||||
|
||||
assertEquals(0, launcher(herdr, cfg).ensureLeads());
|
||||
assertTrue(herdr.calls.isEmpty(), "nothing declared ⇒ nothing scanned");
|
||||
}
|
||||
}
|
||||
+37
@@ -0,0 +1,37 @@
|
||||
package dev.ltms.bridged.logging;
|
||||
|
||||
import ch.qos.logback.classic.Level;
|
||||
import ch.qos.logback.classic.Logger;
|
||||
import ch.qos.logback.classic.spi.ILoggingEvent;
|
||||
import ch.qos.logback.core.read.ListAppender;
|
||||
import io.modelcontextprotocol.spec.McpSchema;
|
||||
import org.junit.jupiter.api.Test;
|
||||
import org.slf4j.LoggerFactory;
|
||||
|
||||
import java.util.Map;
|
||||
|
||||
import static org.junit.jupiter.api.Assertions.assertEquals;
|
||||
import static org.junit.jupiter.api.Assertions.assertTrue;
|
||||
|
||||
class McpCancelledNotificationFilterTest {
|
||||
|
||||
@Test
|
||||
void suppressesOnlyTheCancelledNotificationWarning() {
|
||||
Logger logger = (Logger) LoggerFactory.getLogger(McpCancelledNotificationFilter.LOGGER);
|
||||
ListAppender<ILoggingEvent> appender = new ListAppender<>();
|
||||
appender.start();
|
||||
logger.addAppender(appender);
|
||||
try {
|
||||
logger.warn(McpCancelledNotificationFilter.UNHANDLED_NOTIFICATION,
|
||||
new McpSchema.JSONRPCNotification("notifications/cancelled", Map.of("requestId", 7)));
|
||||
logger.warn(McpCancelledNotificationFilter.UNHANDLED_NOTIFICATION,
|
||||
new McpSchema.JSONRPCNotification("notifications/progress", Map.of("progress", 1)));
|
||||
|
||||
assertEquals(1, appender.list.size());
|
||||
assertEquals(Level.WARN, appender.list.getFirst().getLevel());
|
||||
assertTrue(appender.list.getFirst().getFormattedMessage().contains("notifications/progress"));
|
||||
} finally {
|
||||
logger.detachAppender(appender);
|
||||
}
|
||||
}
|
||||
}
|
||||
@@ -2,6 +2,7 @@ package dev.ltms.bridged.mcp;
|
||||
|
||||
import dev.ltms.bridged.auth.Authz;
|
||||
import dev.ltms.bridged.auth.CallerResolver;
|
||||
import dev.ltms.bridged.auth.MemberRegistry;
|
||||
import dev.ltms.bridged.auth.Principal;
|
||||
import dev.ltms.bridged.auth.Role;
|
||||
import dev.ltms.bridged.config.BridgedConfig;
|
||||
@@ -18,7 +19,7 @@ import dev.ltms.bridged.msg.MessageService;
|
||||
import dev.ltms.bridged.msg.Rendezvous;
|
||||
import dev.ltms.bridged.session.FakeWorktrees;
|
||||
import dev.ltms.bridged.session.SessionManager;
|
||||
import dev.ltms.bridged.worker.ClaudeCodeLauncher;
|
||||
import dev.ltms.bridged.member.ClaudeCodeLauncher;
|
||||
import io.modelcontextprotocol.spec.McpSchema;
|
||||
import org.junit.jupiter.api.AfterEach;
|
||||
import org.junit.jupiter.api.Test;
|
||||
@@ -55,7 +56,7 @@ class BridgeMcpAuthzTest {
|
||||
|
||||
/** A fully wired BridgeMcp on fakes — constructing it is itself part of what is under test. */
|
||||
private BridgeMcp mcp(boolean enforce) {
|
||||
BridgedConfig.Worker cfg = new BridgedConfig.Worker(
|
||||
BridgedConfig.Profile cfg = new BridgedConfig.Profile(
|
||||
"ltms-local", "http://gx00.gw:8000", "coder", null, "BRIDGED_WORKER_TOKEN", null,
|
||||
"tab", "bridged-workers", "worker: {profile} #{n}", null, null, null);
|
||||
ClaudeCodeLauncher workers = new ClaudeCodeLauncher(agents, new WorkspaceControl(herdr),
|
||||
@@ -69,8 +70,9 @@ class BridgeMcpAuthzTest {
|
||||
|
||||
mcp = new BridgeMcp(messages, workers, sessions, identity, sessions.asPresence(),
|
||||
new PrimaryRegistry(null),
|
||||
enforce ? new CallerResolver(identity) : null,
|
||||
metrics);
|
||||
enforce ? CallerResolver.withLeadsAndMembers(identity, false, null,
|
||||
Map::of, new MemberRegistry(null)) : null,
|
||||
metrics, BridgeMcp.CapacitySource.none(), new BridgeMcp.HealthCoverageSource(() -> "off"));
|
||||
return mcp;
|
||||
}
|
||||
|
||||
|
||||
@@ -11,9 +11,11 @@ import dev.ltms.bridged.msg.MessageService;
|
||||
import dev.ltms.bridged.msg.Rendezvous;
|
||||
import dev.ltms.bridged.session.FakeWorktrees;
|
||||
import dev.ltms.bridged.session.SessionManager;
|
||||
import dev.ltms.bridged.session.WorkerSession;
|
||||
import dev.ltms.bridged.peer.MemberRole;
|
||||
import dev.ltms.bridged.inject.MemberPresence;
|
||||
import dev.ltms.bridged.session.MemberSession;
|
||||
import dev.ltms.bridged.session.WorktreeRequest;
|
||||
import dev.ltms.bridged.worker.ClaudeCodeLauncher;
|
||||
import dev.ltms.bridged.member.ClaudeCodeLauncher;
|
||||
import io.modelcontextprotocol.spec.McpSchema;
|
||||
import dev.ltms.bridged.msg.InMemoryReplyInbox;
|
||||
import org.junit.jupiter.api.BeforeEach;
|
||||
@@ -51,8 +53,20 @@ class BridgeMcpTest {
|
||||
return ((McpSchema.TextContent) r.content().getFirst()).text();
|
||||
}
|
||||
|
||||
private void assertSendRoundTrips(String target, Set<String> profiles) throws Exception {
|
||||
CompletableFuture<McpSchema.CallToolResult> send = CompletableFuture.supplyAsync(
|
||||
() -> BridgeMcp.send(messages, target, "hi", 4000L, null, profiles));
|
||||
long deadline = System.currentTimeMillis() + 3000;
|
||||
while (!rendezvous.isWaiting(target) && System.currentTimeMillis() < deadline) {
|
||||
Thread.sleep(5);
|
||||
}
|
||||
assertTrue(rendezvous.isWaiting(target), "send should be accepted for " + target);
|
||||
BridgeMcp.reply(messages, target, "received");
|
||||
assertEquals("received", textOf(send.get(6, TimeUnit.SECONDS)));
|
||||
}
|
||||
|
||||
private static ClaudeCodeLauncher workerService(FakeHerdr h, String baseUrl, Set<String> allow) {
|
||||
BridgedConfig.Worker cfg = new BridgedConfig.Worker(
|
||||
BridgedConfig.Profile cfg = new BridgedConfig.Profile(
|
||||
"ltms-local", baseUrl, "coder", null, "BRIDGED_WORKER_TOKEN", null,
|
||||
"tab", "bridged-workers", "worker: {profile} #{n}", null, null, null);
|
||||
return new ClaudeCodeLauncher(new AgentControl(h), new WorkspaceControl(h),
|
||||
@@ -67,7 +81,7 @@ class BridgeMcpTest {
|
||||
void sendThenReplyRoundTrips() throws Exception {
|
||||
// bridge_send blocks; bridge_reply resolves it with the worker's structured answer.
|
||||
CompletableFuture<McpSchema.CallToolResult> send = CompletableFuture.supplyAsync(
|
||||
() -> BridgeMcp.send(messages, "term_a", "review this", 4000L));
|
||||
() -> BridgeMcp.send(messages, "term_a", "review this", 4000L, null, Set.of()));
|
||||
|
||||
// Wait until the send has opened its waiter so the reply resolves it (CB-307: reply now
|
||||
// queues in the inbox if no waiter is open, which would break the round-trip).
|
||||
@@ -89,7 +103,7 @@ class BridgeMcpTest {
|
||||
@Test
|
||||
void asyncSendReturnsATicketThenPollReportsTheReply() throws Exception {
|
||||
// wait:false parity — a ticket is issued, resolved by a reply, and surfaced by bridge_poll.
|
||||
McpSchema.CallToolResult accepted = BridgeMcp.sendAsync(messages, "term_a", "do it");
|
||||
McpSchema.CallToolResult accepted = BridgeMcp.sendAsync(messages, "term_a", "do it", null, Set.of());
|
||||
assertNotEquals(Boolean.TRUE, accepted.isError());
|
||||
String out = textOf(accepted);
|
||||
assertTrue(out.contains("ticket="), out);
|
||||
@@ -118,6 +132,131 @@ class BridgeMcpTest {
|
||||
assertEquals("async LGTM", textOf(polled));
|
||||
}
|
||||
|
||||
@Test
|
||||
void asyncSendSurfacesAnAskThenKeepsTheTicketForTheFinalReply() throws Exception {
|
||||
McpSchema.CallToolResult accepted = BridgeMcp.sendAsync(messages, "term_a", "do it", null, Set.of());
|
||||
String ticket = textOf(accepted).substring(textOf(accepted).indexOf("ticket=") + "ticket=".length()).trim();
|
||||
|
||||
long deadline = System.currentTimeMillis() + 3000;
|
||||
while (!rendezvous.isWaiting("term_a") && System.currentTimeMillis() < deadline) {
|
||||
Thread.sleep(5);
|
||||
}
|
||||
assertTrue(rendezvous.isWaiting("term_a"));
|
||||
|
||||
CompletableFuture<McpSchema.CallToolResult> ask = CompletableFuture.supplyAsync(
|
||||
() -> BridgeMcp.ask(messages, "term_a", "which config?", 5000L));
|
||||
McpSchema.CallToolResult question = BridgeMcp.poll(messages, ticket, null);
|
||||
deadline = System.currentTimeMillis() + 3000;
|
||||
while (!textOf(question).contains("[question") && System.currentTimeMillis() < deadline) {
|
||||
Thread.sleep(5);
|
||||
question = BridgeMcp.poll(messages, ticket, null);
|
||||
}
|
||||
assertTrue(textOf(question).contains("which config?"), textOf(question));
|
||||
String questionText = textOf(question);
|
||||
String afterTurnId = questionText.substring(questionText.indexOf("turnId=\"") + "turnId=\"".length());
|
||||
String turnId = afterTurnId.substring(0, afterTurnId.indexOf('"'));
|
||||
|
||||
CompletableFuture<McpSchema.CallToolResult> answer = CompletableFuture.supplyAsync(
|
||||
() -> BridgeMcp.answer(messages, turnId, "config.yaml", 5000L));
|
||||
assertEquals("config.yaml", textOf(ask.get(6, TimeUnit.SECONDS)));
|
||||
|
||||
deadline = System.currentTimeMillis() + 3000;
|
||||
while (!rendezvous.isWaiting("term_a") && System.currentTimeMillis() < deadline) {
|
||||
Thread.sleep(5);
|
||||
}
|
||||
assertTrue(rendezvous.isWaiting("term_a"));
|
||||
BridgeMcp.reply(messages, "term_a", "done");
|
||||
assertEquals("done", textOf(answer.get(6, TimeUnit.SECONDS)));
|
||||
|
||||
McpSchema.CallToolResult done = BridgeMcp.poll(messages, ticket, null);
|
||||
deadline = System.currentTimeMillis() + 3000;
|
||||
while (!"done".equals(textOf(done)) && System.currentTimeMillis() < deadline) {
|
||||
Thread.sleep(5);
|
||||
done = BridgeMcp.poll(messages, ticket, null);
|
||||
}
|
||||
assertEquals("done", textOf(done));
|
||||
}
|
||||
|
||||
@Test
|
||||
void unansweredAsyncAskReturnsTheTicketToPending() throws Exception {
|
||||
McpSchema.CallToolResult accepted = BridgeMcp.sendAsync(messages, "term_a", "do it", null, Set.of());
|
||||
String ticket = textOf(accepted).substring(textOf(accepted).indexOf("ticket=") + "ticket=".length()).trim();
|
||||
|
||||
long deadline = System.currentTimeMillis() + 3000;
|
||||
while (!rendezvous.isWaiting("term_a") && System.currentTimeMillis() < deadline) {
|
||||
Thread.sleep(5);
|
||||
}
|
||||
assertTrue(rendezvous.isWaiting("term_a"));
|
||||
|
||||
McpSchema.CallToolResult ask = BridgeMcp.ask(messages, "term_a", "still there?", 50L);
|
||||
assertTrue(textOf(ask).contains("no answer"), textOf(ask));
|
||||
assertTrue(textOf(BridgeMcp.poll(messages, ticket, null)).startsWith("[pending"));
|
||||
|
||||
BridgeMcp.reply(messages, "term_a", "finished after timeout");
|
||||
assertEquals("finished after timeout", messages.drainReplies("term_a").getFirst().content());
|
||||
}
|
||||
|
||||
@Test
|
||||
void asyncSendFailureDoesNotLeaveItsTicketPending() throws Exception {
|
||||
String ticket = messages.sendAsync("term_a", "do it", () -> {
|
||||
throw new IllegalStateException("accept failed");
|
||||
});
|
||||
|
||||
long deadline = System.currentTimeMillis() + 3000;
|
||||
MessageService.TaskView view = messages.poll(ticket);
|
||||
while (view.phase() == MessageService.Phase.PENDING && System.currentTimeMillis() < deadline) {
|
||||
Thread.sleep(5);
|
||||
view = messages.poll(ticket);
|
||||
}
|
||||
assertEquals(MessageService.Phase.FAILED, view.phase());
|
||||
assertEquals("accept failed", view.detail());
|
||||
}
|
||||
|
||||
@Test
|
||||
void anotherAsyncTicketCannotCaptureAReplyWhileTheFirstTicketIsAsking() throws Exception {
|
||||
McpSchema.CallToolResult firstAccepted = BridgeMcp.sendAsync(messages, "term_a", "first", null, Set.of());
|
||||
String firstTicket = textOf(firstAccepted).substring(textOf(firstAccepted).indexOf("ticket=") + "ticket=".length()).trim();
|
||||
|
||||
long deadline = System.currentTimeMillis() + 3000;
|
||||
while (!rendezvous.isWaiting("term_a") && System.currentTimeMillis() < deadline) {
|
||||
Thread.sleep(5);
|
||||
}
|
||||
assertTrue(rendezvous.isWaiting("term_a"));
|
||||
|
||||
CompletableFuture<McpSchema.CallToolResult> ask = CompletableFuture.supplyAsync(
|
||||
() -> BridgeMcp.ask(messages, "term_a", "which config?", 5000L));
|
||||
MessageService.TaskView first = messages.poll(firstTicket);
|
||||
deadline = System.currentTimeMillis() + 3000;
|
||||
while (first.phase() != MessageService.Phase.ASKING && System.currentTimeMillis() < deadline) {
|
||||
Thread.sleep(5);
|
||||
first = messages.poll(firstTicket);
|
||||
}
|
||||
assertEquals(MessageService.Phase.ASKING, first.phase());
|
||||
String firstTurnId = first.turnId();
|
||||
|
||||
McpSchema.CallToolResult secondAccepted = BridgeMcp.sendAsync(messages, "term_a", "second", null, Set.of());
|
||||
String secondTicket = textOf(secondAccepted).substring(textOf(secondAccepted).indexOf("ticket=") + "ticket=".length()).trim();
|
||||
MessageService.TaskView second = messages.poll(secondTicket);
|
||||
deadline = System.currentTimeMillis() + 3000;
|
||||
while (second.phase() == MessageService.Phase.PENDING && System.currentTimeMillis() < deadline) {
|
||||
Thread.sleep(5);
|
||||
second = messages.poll(secondTicket);
|
||||
}
|
||||
assertEquals(MessageService.Phase.FAILED, second.phase());
|
||||
|
||||
BridgeMcp.reply(messages, "term_a", "late reply");
|
||||
assertEquals("late reply", messages.drainReplies("term_a").getFirst().content());
|
||||
|
||||
CompletableFuture<McpSchema.CallToolResult> answer = CompletableFuture.supplyAsync(
|
||||
() -> BridgeMcp.answer(messages, firstTurnId, "config.yaml", 5000L));
|
||||
assertEquals("config.yaml", textOf(ask.get(6, TimeUnit.SECONDS)));
|
||||
while (!rendezvous.isWaiting("term_a") && System.currentTimeMillis() < deadline) {
|
||||
Thread.sleep(5);
|
||||
}
|
||||
BridgeMcp.reply(messages, "term_a", "done");
|
||||
assertEquals("done", textOf(answer.get(6, TimeUnit.SECONDS)));
|
||||
}
|
||||
|
||||
@Test
|
||||
void pollUnknownTicketIsAnError() {
|
||||
McpSchema.CallToolResult res = BridgeMcp.poll(messages, "task-999", null);
|
||||
@@ -127,15 +266,40 @@ class BridgeMcpTest {
|
||||
|
||||
@Test
|
||||
void sendTimesOutWithAWorkingNote() {
|
||||
McpSchema.CallToolResult res = BridgeMcp.send(messages, "term_a", "hi", 120L);
|
||||
McpSchema.CallToolResult res = BridgeMcp.send(messages, "term_a", "hi", 120L, null, Set.of());
|
||||
assertNotEquals(Boolean.TRUE, res.isError(), "a timeout is informational, not a tool error");
|
||||
assertTrue(textOf(res).contains("no reply"), "got: " + textOf(res));
|
||||
}
|
||||
|
||||
@Test
|
||||
void sendRejectsMissingArgs() {
|
||||
assertTrue(BridgeMcp.send(messages, null, "hi", null).isError());
|
||||
assertTrue(BridgeMcp.send(messages, "term_a", " ", null).isError());
|
||||
assertTrue(BridgeMcp.send(messages, null, "hi", null, null, Set.of()).isError());
|
||||
assertTrue(BridgeMcp.send(messages, "term_a", " ", null, null, Set.of()).isError());
|
||||
}
|
||||
|
||||
@Test
|
||||
void sendRejectsAConfiguredProfileNameBeforeAcceptingIt() {
|
||||
McpSchema.CallToolResult blocking = BridgeMcp.send(messages, "sol", "hi", 100L, null, Set.of("sol"));
|
||||
McpSchema.CallToolResult async = BridgeMcp.sendAsync(messages, "sol", "hi", null, Set.of("sol"));
|
||||
|
||||
assertTrue(blocking.isError());
|
||||
assertTrue(async.isError());
|
||||
assertTrue(textOf(blocking).contains("sol"));
|
||||
assertTrue(textOf(blocking).contains("configured profile name"));
|
||||
assertTrue(textOf(blocking).contains("bridge_list"));
|
||||
assertFalse(textOf(async).contains("ticket="));
|
||||
}
|
||||
|
||||
@Test
|
||||
void sendAllowsPeerLeadMemberAndUnclassifiedTargets() throws Exception {
|
||||
Set<String> profiles = Set.of("sol");
|
||||
|
||||
assertSendRoundTrips("term_peer_lead", profiles);
|
||||
assertSendRoundTrips("term_live_member", profiles);
|
||||
|
||||
// A herdr-owned pane outside the bridge roster cannot be classified at accept time.
|
||||
McpSchema.CallToolResult result = BridgeMcp.send(messages, "external-pane", "hi", 10L, null, profiles);
|
||||
assertFalse(result.isError(), "an unclassified target must not be rejected at acceptance time");
|
||||
}
|
||||
|
||||
@Test
|
||||
@@ -171,7 +335,7 @@ class BridgeMcpTest {
|
||||
void askThenAnswerRoundTrips() throws Exception {
|
||||
// The primary delegates and blocks; wait until its waiter is open before the worker asks.
|
||||
CompletableFuture<McpSchema.CallToolResult> send = CompletableFuture.supplyAsync(
|
||||
() -> BridgeMcp.send(messages, "term_a", "do X", 5000L));
|
||||
() -> BridgeMcp.send(messages, "term_a", "do X", 5000L, null, Set.of()));
|
||||
long deadline = System.currentTimeMillis() + 3000;
|
||||
while (!rendezvous.isWaiting("term_a") && System.currentTimeMillis() < deadline) {
|
||||
//noinspection BusyWait
|
||||
@@ -233,7 +397,7 @@ class BridgeMcpTest {
|
||||
String out = textOf(res);
|
||||
assertTrue(out.contains("\"sessionId\":\"term_new_1\""), out);
|
||||
// CB-519: the "paneId" wire field now carries the host-unique opaque id, not the herdr pane.
|
||||
WorkerSession s = sm.roster().getFirst();
|
||||
MemberSession s = sm.roster().getFirst();
|
||||
assertTrue(out.contains("\"paneId\":\"" + s.paneId() + "\""), out);
|
||||
assertNotEquals("w9:pRoot_1", s.paneId(), "the id is decoupled from the herdr pane coordinate");
|
||||
assertTrue(out.contains("\"status\":\"spawning\""), out);
|
||||
@@ -262,7 +426,7 @@ class BridgeMcpTest {
|
||||
void spawnPassesTheRequestedCwdToTheWorker() {
|
||||
FakeHerdr h = new FakeHerdr();
|
||||
McpSchema.CallToolResult res = BridgeMcp.spawn(
|
||||
sessionManager(h, "http://gx00.gw:8000", Set.of("gx00.gw")), null, "/req/dir", null, null, null);
|
||||
sessionManager(h, "http://gx00.gw:8000", Set.of("gx00.gw")), null, null, "/req/dir", null, null, null);
|
||||
assertNotEquals(Boolean.TRUE, res.isError());
|
||||
// Protocol 19: the requested cwd roots the worker's pane at creation (tab.create).
|
||||
@SuppressWarnings("unchecked")
|
||||
@@ -285,7 +449,7 @@ class BridgeMcpTest {
|
||||
FakeHerdr h = new FakeHerdr();
|
||||
FakeWorktrees worktrees = new FakeWorktrees().withRepoRoot("/repo").withPrefix("/wt");
|
||||
SessionManager sessions = new SessionManager(workerService(h, "http://gx00.gw:8000", Set.of("gx00.gw")), worktrees);
|
||||
WorkerSession s = sessions.acquire("ltms-local", null, "/caller/proj", "term_primary",
|
||||
MemberSession s = sessions.acquire("ltms-local", null, "/caller/proj", "term_primary",
|
||||
new WorktreeRequest("cb-304", null));
|
||||
|
||||
McpSchema.CallToolResult res = BridgeMcp.listFleet(
|
||||
@@ -303,6 +467,42 @@ class BridgeMcpTest {
|
||||
assertTrue(out.contains("\"liveStatus\":\"unknown\""), out);
|
||||
}
|
||||
|
||||
@Test
|
||||
void capacityUsesThePlacementLiveCount() {
|
||||
FakeHerdr h = new FakeHerdr();
|
||||
SessionManager sessions = new SessionManager(workerService(h, "http://gx00.gw:8000", Set.of("gx00.gw")));
|
||||
sessions.acquire("ltms-local", null, null, null);
|
||||
McpSchema.CallToolResult res = BridgeMcp.listFleet(workerService(h, "http://gx00.gw:8000", Set.of("gx00.gw")),
|
||||
sessions, null, new BridgeMcp.CapacitySource(profile -> 2, profile -> 2,
|
||||
() -> Set.of("ltms-local"), () -> 0), new BridgeMcp.HealthCoverageSource(() -> "off"), Map.of(), "");
|
||||
String out = textOf(res);
|
||||
assertTrue(out.contains("\"maxLoad\":2"), out);
|
||||
assertTrue(out.contains("\"live\":2"), out);
|
||||
assertTrue(out.contains("\"free\":0"), out);
|
||||
}
|
||||
|
||||
@Test
|
||||
void capacityIncludesConfiguredProfileWithoutMembers() {
|
||||
FakeHerdr h = new FakeHerdr();
|
||||
SessionManager sessions = new SessionManager(workerService(h, "http://gx00.gw:8000", Set.of("gx00.gw")));
|
||||
String out = textOf(BridgeMcp.listFleet(workerService(h, "http://gx00.gw:8000", Set.of("gx00.gw")),
|
||||
sessions, null, new BridgeMcp.CapacitySource(profile -> 0, profile -> 2,
|
||||
() -> Set.of("terra"), () -> 0), new BridgeMcp.HealthCoverageSource(() -> "off"), Map.of(), ""));
|
||||
assertTrue(out.contains("\"profile\":\"terra\""), out);
|
||||
assertTrue(out.contains("\"live\":0"), out);
|
||||
assertTrue(out.contains("\"free\":2"), out);
|
||||
assertTrue(out.contains("\"reclaimable\":0"), out);
|
||||
}
|
||||
|
||||
@Test
|
||||
void inertCapacitySourceOmitsCapacityBlock() {
|
||||
FakeHerdr h = new FakeHerdr();
|
||||
String out = textOf(BridgeMcp.listFleet(workerService(h, "http://gx00.gw:8000", Set.of("gx00.gw")),
|
||||
new SessionManager(workerService(h, "http://gx00.gw:8000", Set.of("gx00.gw"))), null,
|
||||
BridgeMcp.CapacitySource.none(), new BridgeMcp.HealthCoverageSource(() -> "off"), Map.of(), ""));
|
||||
assertFalse(out.contains("\"capacity\":"), out);
|
||||
}
|
||||
|
||||
@Test
|
||||
void listReportsLeadsAndFlagsTheCallersOwnRow() {
|
||||
FakeHerdr h = new FakeHerdr();
|
||||
@@ -333,10 +533,10 @@ class BridgeMcpTest {
|
||||
McpSchema.CallToolResult res = BridgeMcp.listFleet(
|
||||
workerService(h, "http://gx00.gw:8000", Set.of("gx00.gw")), sessions, Map.of(), "");
|
||||
|
||||
// An absent "leads" key is what made an empty worker roster read as "no peers" (CB-535).
|
||||
// An absent "leads" key is what made an empty member roster read as "no peers" (CB-535).
|
||||
String out = textOf(res);
|
||||
assertTrue(out.contains("\"leads\":[]"), out);
|
||||
assertTrue(out.contains("\"workers\":[]"), out);
|
||||
assertTrue(out.contains("\"members\":[]"), out);
|
||||
}
|
||||
|
||||
@Test
|
||||
@@ -403,6 +603,30 @@ class BridgeMcpTest {
|
||||
assertDoesNotThrow(() -> messages.ackReply("term_a", msgId));
|
||||
}
|
||||
|
||||
@Test
|
||||
void spawnedMembersAreMarkedPresent() {
|
||||
Principal worker = Principal.worker("term_worker", 200);
|
||||
Principal architect = Principal.architect("lead-designer", "term_design", 400);
|
||||
MemberPresence presence = new MemberPresence();
|
||||
|
||||
BridgeMcp.markSpawnedMemberPresent(worker, presence);
|
||||
BridgeMcp.markSpawnedMemberPresent(architect, presence);
|
||||
|
||||
assertTrue(presence.isPresent("term_worker"));
|
||||
assertTrue(presence.isPresent("term_design"));
|
||||
}
|
||||
|
||||
@Test
|
||||
void nonMembersAreNotMarkedPresent() {
|
||||
Principal lead = Principal.leader("opus", "term_lead", 100);
|
||||
MemberPresence presence = new MemberPresence();
|
||||
|
||||
BridgeMcp.markSpawnedMemberPresent(lead, presence);
|
||||
BridgeMcp.markSpawnedMemberPresent(Principal.anonymous(), presence);
|
||||
|
||||
assertFalse(presence.isPresent("term_lead"));
|
||||
}
|
||||
|
||||
@Test
|
||||
void statusReportsLiveAgentStatus() {
|
||||
FakeHerdr blocked = new FakeHerdr().agentStatus("blocked");
|
||||
@@ -434,7 +658,7 @@ class BridgeMcpTest {
|
||||
FakeWorktrees worktrees = new FakeWorktrees().withRepoRoot("/repo").withPrefix("/wt");
|
||||
SessionManager sessions = new SessionManager(
|
||||
workerService(h, "http://gx00.gw:8000", Set.of("gx00.gw")), worktrees);
|
||||
WorkerSession s = sessions.acquire("ltms-local", null, "/caller/proj", "term_primary",
|
||||
MemberSession s = sessions.acquire("ltms-local", null, "/caller/proj", "term_primary",
|
||||
new WorktreeRequest("cb-517", null));
|
||||
|
||||
McpSchema.CallToolResult res = BridgeMcp.whoami(Principal.worker(s.terminalId(), 200), sessions);
|
||||
@@ -517,4 +741,73 @@ class BridgeMcpTest {
|
||||
BridgeMcp.recordPrimarySingleton(legacy, "term_x", null);
|
||||
assertTrue(legacy.primaryTerminal().isEmpty(), "no caller means nothing is recorded");
|
||||
}
|
||||
|
||||
// --- member taxonomy (CB-557) ---------------------------------------------------------
|
||||
|
||||
@Test
|
||||
void bridgeListReportsMembersNotWorkersAndCarriesEachRole() {
|
||||
FakeHerdr h = new FakeHerdr();
|
||||
SessionManager sessions = new SessionManager(workerService(h, "http://gx00.gw:8000", Set.of("gx00.gw")));
|
||||
sessions.acquire("ltms-local", MemberRole.REVIEWER, null, "/caller/proj", "term_primary", null);
|
||||
|
||||
McpSchema.CallToolResult res = BridgeMcp.listFleet(
|
||||
workerService(h, "http://gx00.gw:8000", Set.of("gx00.gw")), sessions, Map.of(), "");
|
||||
|
||||
String out = textOf(res);
|
||||
assertTrue(out.contains("\"members\":"), "the roster half is named members: " + out);
|
||||
assertFalse(out.contains("\"workers\":"), "the old key must be gone: " + out);
|
||||
assertTrue(out.contains("\"role\":\"reviewer\""), out);
|
||||
}
|
||||
|
||||
/**
|
||||
* Role and profile are separate axes, so the roster has to report both. Two members on one
|
||||
* backend may still be allowed to do entirely different things.
|
||||
*/
|
||||
@Test
|
||||
void aRosterRowCarriesBothItsRoleAndItsProfile() {
|
||||
FakeHerdr h = new FakeHerdr();
|
||||
SessionManager sessions = new SessionManager(workerService(h, "http://gx00.gw:8000", Set.of("gx00.gw")));
|
||||
sessions.acquire("ltms-local", MemberRole.DEV, null, "/caller/proj", "term_primary", null);
|
||||
sessions.acquire("ltms-local", MemberRole.REVIEWER, null, "/caller/proj", "term_primary", null);
|
||||
|
||||
String out = textOf(BridgeMcp.listFleet(
|
||||
workerService(h, "http://gx00.gw:8000", Set.of("gx00.gw")), sessions, Map.of(), ""));
|
||||
|
||||
assertTrue(out.contains("\"role\":\"dev\""), out);
|
||||
assertTrue(out.contains("\"role\":\"reviewer\""), out);
|
||||
assertEquals(2, out.split("\"profile\":\"ltms-local\"", -1).length - 1,
|
||||
"inert capacity is omitted, leaving the two member rows: " + out);
|
||||
}
|
||||
|
||||
@Test
|
||||
void spawnDefaultsTheRoleToDev() {
|
||||
FakeHerdr h = new FakeHerdr();
|
||||
McpSchema.CallToolResult res = BridgeMcp.spawn(
|
||||
sessionManager(h, "http://gx00.gw:8000", Set.of("gx00.gw")), null, null, null, null, null, null);
|
||||
|
||||
assertNotEquals(Boolean.TRUE, res.isError());
|
||||
assertTrue(textOf(res).contains("\"role\":\"dev\""), textOf(res));
|
||||
}
|
||||
|
||||
@Test
|
||||
void spawnAcceptsAnExplicitRole() {
|
||||
FakeHerdr h = new FakeHerdr();
|
||||
McpSchema.CallToolResult res = BridgeMcp.spawn(
|
||||
sessionManager(h, "http://gx00.gw:8000", Set.of("gx00.gw")), null, "architect",
|
||||
null, null, null, null);
|
||||
|
||||
assertNotEquals(Boolean.TRUE, res.isError());
|
||||
assertTrue(textOf(res).contains("\"role\":\"architect\""), textOf(res));
|
||||
}
|
||||
|
||||
@Test
|
||||
void spawnRejectsAnUnknownRoleAndNamesTheValidOnes() {
|
||||
FakeHerdr h = new FakeHerdr();
|
||||
McpSchema.CallToolResult res = BridgeMcp.spawn(
|
||||
sessionManager(h, "http://gx00.gw:8000", Set.of("gx00.gw")), null, "worker",
|
||||
null, null, null, null);
|
||||
|
||||
assertEquals(Boolean.TRUE, res.isError());
|
||||
assertTrue(textOf(res).contains("architect, dev, reviewer"), textOf(res));
|
||||
}
|
||||
}
|
||||
|
||||
+171
-24
@@ -1,4 +1,4 @@
|
||||
package dev.ltms.bridged.worker;
|
||||
package dev.ltms.bridged.member;
|
||||
|
||||
import dev.ltms.bridged.config.BridgedConfig;
|
||||
import dev.ltms.bridged.guard.GuardException;
|
||||
@@ -7,6 +7,7 @@ import dev.ltms.bridged.herdr.AgentControl;
|
||||
import dev.ltms.bridged.herdr.FakeHerdr;
|
||||
import dev.ltms.bridged.herdr.WorkspaceControl;
|
||||
import dev.ltms.bridged.peer.Capability;
|
||||
import dev.ltms.bridged.peer.MemberRole;
|
||||
import dev.ltms.bridged.peer.PeerHandle;
|
||||
import dev.ltms.bridged.peer.PeerUnreachableException;
|
||||
import dev.ltms.bridged.peer.SpawnRequest;
|
||||
@@ -16,7 +17,9 @@ import java.util.List;
|
||||
import java.util.Map;
|
||||
import java.util.Set;
|
||||
import java.util.UUID;
|
||||
import java.util.concurrent.atomic.AtomicReference;
|
||||
import java.util.function.Function;
|
||||
import java.util.function.Supplier;
|
||||
|
||||
import static org.junit.jupiter.api.Assertions.*;
|
||||
|
||||
@@ -24,7 +27,7 @@ import static org.junit.jupiter.api.Assertions.*;
|
||||
class ClaudeCodeLauncherTest {
|
||||
|
||||
private ClaudeCodeLauncher service(FakeHerdr herdr, List<String> argv, String mcpUrl) {
|
||||
BridgedConfig.Worker cfg = new BridgedConfig.Worker(
|
||||
BridgedConfig.Profile cfg = new BridgedConfig.Profile(
|
||||
"ltms-local", "http://gx00.gw:8000", "coder", null, "BRIDGED_WORKER_TOKEN",
|
||||
argv, "tab", "bridged-workers", "worker: {profile} #{n}", mcpUrl, null, null);
|
||||
return new ClaudeCodeLauncher(new AgentControl(herdr), new WorkspaceControl(herdr),
|
||||
@@ -59,7 +62,7 @@ class ClaudeCodeLauncherTest {
|
||||
ClaudeCodeLauncher svc = new ClaudeCodeLauncher(
|
||||
new AgentControl(herdr), new WorkspaceControl(herdr),
|
||||
new SubscriptionGuard(Set.of("gx00.gw")),
|
||||
Map.of("ltms-local", new BridgedConfig.Worker(
|
||||
Map.of("ltms-local", new BridgedConfig.Profile(
|
||||
"ltms-local", "http://gx00.gw:8000", "coder", null, "BRIDGED_WORKER_TOKEN",
|
||||
List.of("claude"), "tab", "bridged-workers", "w #{n}", null, null, null)),
|
||||
"ltms-local", _ -> null,
|
||||
@@ -99,10 +102,57 @@ class ClaudeCodeLauncherTest {
|
||||
"the operator's own args are preserved, in order, ahead of the model flag");
|
||||
}
|
||||
|
||||
@Test
|
||||
void appendsTheBaseComposedRoleAndReplyCharter() {
|
||||
FakeHerdr herdr = new FakeHerdr();
|
||||
String roleCharter = "You review changes.";
|
||||
BridgedConfig.Profile cfg = new BridgedConfig.Profile(
|
||||
"sonnet", "http://gx00.gw:8000", null, null, "BRIDGED_WORKER_TOKEN",
|
||||
List.of("claude"), "tab", "bridged-workers", "w #{n}",
|
||||
"http://127.0.0.1:8765/mcp", null, null);
|
||||
ClaudeCodeLauncher svc = new ClaudeCodeLauncher(new AgentControl(herdr), new WorkspaceControl(herdr),
|
||||
new SubscriptionGuard(Set.of("gx00.gw")), Map.of(cfg.profile(), cfg), cfg.profile(), _ -> null,
|
||||
0, 0L, () -> fleet(Map.of("reviewer", roleCharter), null));
|
||||
|
||||
svc.spawn(new SpawnRequest("sonnet", null, null, null, null, MemberRole.REVIEWER));
|
||||
|
||||
List<String> args = spawnedArgs(herdr);
|
||||
int flag = args.indexOf("--append-system-prompt");
|
||||
assertEquals(1, args.stream().filter("--append-system-prompt"::equals).count(),
|
||||
"the composed charter is passed once");
|
||||
assertEquals(roleCharter + "\n\n" + HerdrPeerLauncher.REPLY_CHARTER, args.get(flag + 1),
|
||||
"the role charter comes first and the reply rule comes last");
|
||||
}
|
||||
|
||||
@Test
|
||||
void profileWithoutMcpOrRoleCharterGetsNoSystemPrompt() {
|
||||
FakeHerdr herdr = new FakeHerdr();
|
||||
ClaudeCodeLauncher svc = labelService(herdr, () -> fleet(Map.of(), null));
|
||||
|
||||
svc.spawn(new SpawnRequest("sonnet", null, null, null, null, MemberRole.DEV));
|
||||
|
||||
assertFalse(spawnedArgs(herdr).contains("--append-system-prompt"));
|
||||
}
|
||||
|
||||
@Test
|
||||
void profileWithoutMcpStillGetsItsRoleCharter() {
|
||||
FakeHerdr herdr = new FakeHerdr();
|
||||
String roleCharter = "You design changes.";
|
||||
ClaudeCodeLauncher svc = labelService(herdr, () -> fleet(Map.of("architect", roleCharter), null));
|
||||
|
||||
svc.spawn(new SpawnRequest("sonnet", null, null, null, null, MemberRole.ARCHITECT));
|
||||
|
||||
List<String> args = spawnedArgs(herdr);
|
||||
int flag = args.indexOf("--append-system-prompt");
|
||||
assertTrue(flag >= 0, "a role charter does not need an MCP mount");
|
||||
assertEquals(roleCharter, args.get(flag + 1));
|
||||
assertFalse(args.contains("--mcp-config"));
|
||||
}
|
||||
|
||||
private ClaudeCodeLauncher multiProfile(FakeHerdr herdr) {
|
||||
BridgedConfig.Worker gx10 = new BridgedConfig.Worker("gx10", "http://gx10.gw:8000", "coder",
|
||||
BridgedConfig.Profile gx10 = new BridgedConfig.Profile("gx10", "http://gx10.gw:8000", "coder",
|
||||
null, "BRIDGED_WORKER_TOKEN", List.of("claude"), "tab", "bridged-workers", "w #{n}", null, null, null);
|
||||
BridgedConfig.Worker ollama = new BridgedConfig.Worker("ollama", "http://ollama.ltms.dev", null,
|
||||
BridgedConfig.Profile ollama = new BridgedConfig.Profile("ollama", "http://ollama.ltms.dev", null,
|
||||
null, "BRIDGED_WORKER_TOKEN", List.of("claude"), "tab", "bridged-workers", "w #{n}", null, null, null);
|
||||
return new ClaudeCodeLauncher(new AgentControl(herdr), new WorkspaceControl(herdr),
|
||||
new SubscriptionGuard(Set.of("gx10.gw", "ollama.ltms.dev")),
|
||||
@@ -143,7 +193,7 @@ class ClaudeCodeLauncherTest {
|
||||
@Test
|
||||
void profileConfigCwdBeatsTheCallerCwd() {
|
||||
FakeHerdr herdr = new FakeHerdr();
|
||||
BridgedConfig.Worker cfg = new BridgedConfig.Worker("ltms-local", "http://gx00.gw:8000", "coder",
|
||||
BridgedConfig.Profile cfg = new BridgedConfig.Profile("ltms-local", "http://gx00.gw:8000", "coder",
|
||||
null, "BRIDGED_WORKER_TOKEN", List.of("ccs", "ltms-local"), "tab", "bridged-workers",
|
||||
"w #{n}", null, "/pinned/dir", null);
|
||||
ClaudeCodeLauncher svc = new ClaudeCodeLauncher(new AgentControl(herdr), new WorkspaceControl(herdr),
|
||||
@@ -170,7 +220,7 @@ class ClaudeCodeLauncherTest {
|
||||
@Test
|
||||
void injectsForgeTokenAndHostWhenProfileGrantsIt() {
|
||||
FakeHerdr herdr = new FakeHerdr();
|
||||
BridgedConfig.Worker cfg = new BridgedConfig.Worker(
|
||||
BridgedConfig.Profile cfg = new BridgedConfig.Profile(
|
||||
"impl", "http://gx00.gw:8000", "coder", null, "BRIDGED_WORKER_TOKEN",
|
||||
List.of("ccs", "impl"), "tab", "bridged-workers", "w #{n}", null, null, null,
|
||||
"GITEA_ACCESS_TOKEN", null); // parityOverlay null; gitHostEnv null → defaults to GITEA_HOST
|
||||
@@ -192,7 +242,7 @@ class ClaudeCodeLauncherTest {
|
||||
FakeHerdr herdr = new FakeHerdr();
|
||||
// gitTokenEnv unset (12-arg ctor); the env would resolve a token if asked, proving the gate
|
||||
// is the profile config, not a missing env var.
|
||||
BridgedConfig.Worker cfg = new BridgedConfig.Worker("ltms-local", "http://gx00.gw:8000", "coder",
|
||||
BridgedConfig.Profile cfg = new BridgedConfig.Profile("ltms-local", "http://gx00.gw:8000", "coder",
|
||||
null, "BRIDGED_WORKER_TOKEN", List.of("ccs"), "tab", "bridged-workers", "w #{n}",
|
||||
null, null, null);
|
||||
new ClaudeCodeLauncher(new AgentControl(herdr), new WorkspaceControl(herdr),
|
||||
@@ -323,7 +373,7 @@ class ClaudeCodeLauncherTest {
|
||||
@Test
|
||||
void capabilitiesIncludeSelfPrWhenProfileHasGitToken() {
|
||||
FakeHerdr herdr = new FakeHerdr();
|
||||
BridgedConfig.Worker cfg = new BridgedConfig.Worker(
|
||||
BridgedConfig.Profile cfg = new BridgedConfig.Profile(
|
||||
"impl", "http://gx00.gw:8000", "coder", null, "BRIDGED_WORKER_TOKEN",
|
||||
List.of("ccs", "impl"), "tab", "bridged-workers", "w #{n}", null, null, null,
|
||||
"GITEA_ACCESS_TOKEN", null);
|
||||
@@ -476,8 +526,8 @@ class ClaudeCodeLauncherTest {
|
||||
|
||||
// --- CB-306 spawn-readiness gate -----------------------------------------------------------
|
||||
|
||||
private static Map<String, BridgedConfig.Worker> workerConfigMap(String profile, String mcpUrl) {
|
||||
BridgedConfig.Worker cfg = new BridgedConfig.Worker(
|
||||
private static Map<String, BridgedConfig.Profile> workerConfigMap(String profile, String mcpUrl) {
|
||||
BridgedConfig.Profile cfg = new BridgedConfig.Profile(
|
||||
profile, "http://gx00.gw:8000", "coder", null, "BRIDGED_WORKER_TOKEN",
|
||||
List.of("ccs", profile), "tab", "bridged-workers",
|
||||
"worker: {profile} #{n}", mcpUrl, null, null);
|
||||
@@ -579,7 +629,7 @@ class ClaudeCodeLauncherTest {
|
||||
@Test
|
||||
void workerInheritsTheDaemonPath() {
|
||||
FakeHerdr herdr = new FakeHerdr();
|
||||
BridgedConfig.Worker cfg = new BridgedConfig.Worker(
|
||||
BridgedConfig.Profile cfg = new BridgedConfig.Profile(
|
||||
"ltms-local", "http://gx00.gw:8000", "coder", null, "BRIDGED_WORKER_TOKEN",
|
||||
List.of("claude"), "tab", "bridged-workers", "w #{n}", null, null, null);
|
||||
new ClaudeCodeLauncher(new AgentControl(herdr), new WorkspaceControl(herdr),
|
||||
@@ -593,7 +643,7 @@ class ClaudeCodeLauncherTest {
|
||||
@Test
|
||||
void profileEnvIsInjectedIntoTheWorker() {
|
||||
FakeHerdr herdr = new FakeHerdr();
|
||||
BridgedConfig.Worker cfg = new BridgedConfig.Worker(
|
||||
BridgedConfig.Profile cfg = new BridgedConfig.Profile(
|
||||
"ltms-local", "http://gx00.gw:8000", "coder", null, "BRIDGED_WORKER_TOKEN",
|
||||
List.of("claude"), "tab", "bridged-workers", "w #{n}", null, null, null, null, null,
|
||||
null, Map.of("JAVA_HOME", "/opt/jdk", "PATH", "/profile/bin"), null, null);
|
||||
@@ -614,7 +664,7 @@ class ClaudeCodeLauncherTest {
|
||||
@Test
|
||||
void profileEnvCannotOverrideGuardCheckedAnthropicVars() {
|
||||
FakeHerdr herdr = new FakeHerdr();
|
||||
BridgedConfig.Worker cfg = new BridgedConfig.Worker(
|
||||
BridgedConfig.Profile cfg = new BridgedConfig.Profile(
|
||||
"ltms-local", "http://gx00.gw:8000", "coder", null, "BRIDGED_WORKER_TOKEN",
|
||||
List.of("claude"), "tab", "bridged-workers", "w #{n}", null, null, null, null, null,
|
||||
null, Map.of("ANTHROPIC_BASE_URL", "http://evil.example.com"), null, null);
|
||||
@@ -630,14 +680,14 @@ class ClaudeCodeLauncherTest {
|
||||
|
||||
/** A launcher for a profile identical but for its {@code model:} — the only variable here. */
|
||||
private ClaudeCodeLauncher serviceWithModel(FakeHerdr herdr, String model) {
|
||||
BridgedConfig.Worker cfg = profileWithModel(model);
|
||||
BridgedConfig.Profile cfg = profileWithModel(model);
|
||||
return new ClaudeCodeLauncher(new AgentControl(herdr), new WorkspaceControl(herdr),
|
||||
new SubscriptionGuard(Set.of("gx00.gw")), Map.of(cfg.profile(), cfg), cfg.profile(),
|
||||
_ -> null);
|
||||
}
|
||||
|
||||
private static BridgedConfig.Worker profileWithModel(String model) {
|
||||
return new BridgedConfig.Worker("sonnet", "http://gx00.gw:8000", model, null,
|
||||
private static BridgedConfig.Profile profileWithModel(String model) {
|
||||
return new BridgedConfig.Profile("sonnet", "http://gx00.gw:8000", model, null,
|
||||
"BRIDGED_WORKER_TOKEN", List.of("ccs", "sonnet"), "tab", "bridged-workers",
|
||||
"w #{n}", "http://127.0.0.1:8765/mcp", null, null);
|
||||
}
|
||||
@@ -679,8 +729,8 @@ class ClaudeCodeLauncherTest {
|
||||
// --- CB-539: subscription-profile opt-in ----------------------------------------------------
|
||||
|
||||
/** A claude-code profile on the subscription: no baseUrl (by design), no off-sub endpoint. */
|
||||
private static BridgedConfig.Worker subscriptionCfg(String profile, String baseUrl) {
|
||||
return new BridgedConfig.Worker(
|
||||
private static BridgedConfig.Profile subscriptionCfg(String profile, String baseUrl) {
|
||||
return new BridgedConfig.Profile(
|
||||
profile, baseUrl, "sonnet", null, "BRIDGED_WORKER_TOKEN",
|
||||
List.of("ccs", profile), "tab", "bridged-workers", "w #{n}", null, null, null,
|
||||
null, null, null, Map.of(), null, null, true);
|
||||
@@ -691,7 +741,7 @@ class ClaudeCodeLauncherTest {
|
||||
// Requirement 1: absent subscription:true ⇒ byte-identical refusal to today. A claude-code
|
||||
// profile with no baseUrl and no subscription must still be refused (it would bill the sub).
|
||||
FakeHerdr herdr = new FakeHerdr();
|
||||
BridgedConfig.Worker cfg = new BridgedConfig.Worker(
|
||||
BridgedConfig.Profile cfg = new BridgedConfig.Profile(
|
||||
"ltms-local", null, "coder", null, "BRIDGED_WORKER_TOKEN",
|
||||
List.of("ccs", "ltms-local"), "tab", "bridged-workers", "w #{n}", null, null, null);
|
||||
ClaudeCodeLauncher svc = new ClaudeCodeLauncher(new AgentControl(herdr), new WorkspaceControl(herdr),
|
||||
@@ -710,7 +760,7 @@ class ClaudeCodeLauncherTest {
|
||||
// ANTHROPIC_BASE_URL nor ANTHROPIC_AUTH_TOKEN is injected even though the token env would
|
||||
// resolve one if asked.
|
||||
FakeHerdr herdr = new FakeHerdr();
|
||||
BridgedConfig.Worker cfg = subscriptionCfg("sonnet", null);
|
||||
BridgedConfig.Profile cfg = subscriptionCfg("sonnet", null);
|
||||
new ClaudeCodeLauncher(new AgentControl(herdr), new WorkspaceControl(herdr),
|
||||
new SubscriptionGuard(Set.of("gx00.gw")), Map.of(cfg.profile(), cfg), cfg.profile(),
|
||||
_ -> "would-be-token").spawn();
|
||||
@@ -727,7 +777,7 @@ class ClaudeCodeLauncherTest {
|
||||
// Requirement 2: subscription:true + a baseUrl state opposite intents — refuse at spawn,
|
||||
// naming the profile, rather than silently picking a winner.
|
||||
FakeHerdr herdr = new FakeHerdr();
|
||||
BridgedConfig.Worker cfg = subscriptionCfg("sonnet", "http://gx00.gw:8000");
|
||||
BridgedConfig.Profile cfg = subscriptionCfg("sonnet", "http://gx00.gw:8000");
|
||||
ClaudeCodeLauncher svc = new ClaudeCodeLauncher(new AgentControl(herdr), new WorkspaceControl(herdr),
|
||||
new SubscriptionGuard(Set.of("gx00.gw")), Map.of(cfg.profile(), cfg), cfg.profile(), _ -> null);
|
||||
|
||||
@@ -744,7 +794,7 @@ class ClaudeCodeLauncherTest {
|
||||
// Requirement 3: the guard keeps its teeth for every other profile — a base_url whose host is
|
||||
// not on the allowlist is still refused, whether or not any subscription profile exists.
|
||||
FakeHerdr herdr = new FakeHerdr();
|
||||
BridgedConfig.Worker rogue = new BridgedConfig.Worker(
|
||||
BridgedConfig.Profile rogue = new BridgedConfig.Profile(
|
||||
"rogue", "http://evil.example.com:8000", "coder", null, "BRIDGED_WORKER_TOKEN",
|
||||
List.of("ccs", "rogue"), "tab", "bridged-workers", "w #{n}", null, null, null);
|
||||
ClaudeCodeLauncher svc = new ClaudeCodeLauncher(new AgentControl(herdr), new WorkspaceControl(herdr),
|
||||
@@ -763,7 +813,7 @@ class ClaudeCodeLauncherTest {
|
||||
// load refuses this loudly; this launcher-side strip is the belt-and-braces that makes the
|
||||
// invariant hold for a profile built in code that never passed through that validation.
|
||||
FakeHerdr herdr = new FakeHerdr();
|
||||
BridgedConfig.Worker cfg = new BridgedConfig.Worker(
|
||||
BridgedConfig.Profile cfg = new BridgedConfig.Profile(
|
||||
"sonnet", null, "sonnet", null, "BRIDGED_WORKER_TOKEN",
|
||||
List.of("ccs", "sonnet"), "tab", "bridged-workers", "w #{n}", null, null, null,
|
||||
null, null, null,
|
||||
@@ -782,4 +832,101 @@ class ClaudeCodeLauncherTest {
|
||||
assertEquals("/opt/jdk", env.get("JAVA_HOME"),
|
||||
"only the Anthropic binding keys are stripped; the rest of env: still applies");
|
||||
}
|
||||
|
||||
// ── CB-557: role-aware tab labels ─────────────────────────────────────────────────────────
|
||||
|
||||
/** The {@code label} of every {@code tab.rename}, in call order. */
|
||||
private List<String> tabLabels(FakeHerdr herdr) {
|
||||
return herdr.calls.stream()
|
||||
.filter(c -> c.method().equals("tab.rename"))
|
||||
.map(c -> (String) ((Map<?, ?>) c.params()).get("label"))
|
||||
.toList();
|
||||
}
|
||||
|
||||
private static BridgedConfig.Fleet fleet(Map<String, String> charters, String tabLabel) {
|
||||
return new BridgedConfig.Fleet(Map.of(), Map.of(), Map.of(), Map.of(), charters, tabLabel);
|
||||
}
|
||||
|
||||
/** A profile with no {@code tabLabel:} of its own — the fleet template decides. */
|
||||
private ClaudeCodeLauncher labelService(FakeHerdr herdr, Supplier<BridgedConfig.Fleet> fleet) {
|
||||
BridgedConfig.Profile cfg = new BridgedConfig.Profile(
|
||||
"sonnet", "http://gx00.gw:8000", "sonnet", null, "BRIDGED_WORKER_TOKEN",
|
||||
List.of("claude"), "tab", "bridged-workers", null, null, null, null);
|
||||
return new ClaudeCodeLauncher(new AgentControl(herdr), new WorkspaceControl(herdr),
|
||||
new SubscriptionGuard(Set.of("gx00.gw")), Map.of(cfg.profile(), cfg), cfg.profile(),
|
||||
_ -> null, 0, 0L, fleet);
|
||||
}
|
||||
|
||||
/**
|
||||
* The knob must reach the rename call. It was inert once — {@code HerdrPeerLauncher} accepted a
|
||||
* template while {@code Bridged} passed none, and the label stayed right only because the
|
||||
* fallback happened to match. Pin the wiring, not the coincidence.
|
||||
*/
|
||||
@Test
|
||||
void theFleetTemplateNamesTheRoleTheMemberWasSpawnedFor() {
|
||||
FakeHerdr herdr = new FakeHerdr();
|
||||
ClaudeCodeLauncher svc = labelService(herdr, () -> fleet(null, "{role}: {profile} #{n}"));
|
||||
|
||||
svc.spawn(new SpawnRequest("sonnet", null, null, null, null, MemberRole.REVIEWER));
|
||||
|
||||
assertEquals(List.of("reviewer: sonnet #1"), tabLabels(herdr));
|
||||
}
|
||||
|
||||
/** The counter is per role+profile, so a dev and a reviewer on one profile both start at #1. */
|
||||
@Test
|
||||
void theCounterRunsPerRoleAndProfileNotPerFleet() {
|
||||
FakeHerdr herdr = new FakeHerdr();
|
||||
ClaudeCodeLauncher svc = labelService(herdr, () -> fleet(null, "{role}: {profile} #{n}"));
|
||||
|
||||
svc.spawn(new SpawnRequest("sonnet", null, null, null, null, MemberRole.DEV));
|
||||
svc.spawn(new SpawnRequest("sonnet", null, null, null, null, MemberRole.REVIEWER));
|
||||
svc.spawn(new SpawnRequest("sonnet", null, null, null, null, MemberRole.DEV));
|
||||
|
||||
assertEquals(List.of("dev: sonnet #1", "reviewer: sonnet #1", "dev: sonnet #2"),
|
||||
tabLabels(herdr));
|
||||
}
|
||||
|
||||
/** No fleet template configured ⇒ the built-in default, still role-first. */
|
||||
@Test
|
||||
void aBlankFleetTemplateFallsBackToTheRoleFirstDefault() {
|
||||
FakeHerdr herdr = new FakeHerdr();
|
||||
labelService(herdr, () -> fleet(null, null)).spawn(
|
||||
new SpawnRequest("sonnet", null, null, null, null, MemberRole.ARCHITECT));
|
||||
|
||||
assertEquals(List.of("architect: sonnet #1"), tabLabels(herdr));
|
||||
assertEquals("{role}: {profile} #{n}", BridgedConfig.Fleet.DEFAULT_TAB_LABEL);
|
||||
}
|
||||
|
||||
/** A profile that wants its own label still outranks the fleet template. */
|
||||
@Test
|
||||
void aProfileTabLabelOverridesTheFleetTemplate() {
|
||||
FakeHerdr herdr = new FakeHerdr();
|
||||
BridgedConfig.Profile cfg = new BridgedConfig.Profile(
|
||||
"sonnet", "http://gx00.gw:8000", "sonnet", null, "BRIDGED_WORKER_TOKEN",
|
||||
List.of("claude"), "tab", "bridged-workers", "pinned {profile}", null, null, null);
|
||||
new ClaudeCodeLauncher(new AgentControl(herdr), new WorkspaceControl(herdr),
|
||||
new SubscriptionGuard(Set.of("gx00.gw")), Map.of(cfg.profile(), cfg), cfg.profile(),
|
||||
_ -> null, 0, 0L, () -> fleet(null, "{role}: {profile} #{n}"))
|
||||
.spawn(new SpawnRequest("sonnet", null, null, null, null, MemberRole.REVIEWER));
|
||||
|
||||
assertEquals(List.of("pinned sonnet"), tabLabels(herdr));
|
||||
}
|
||||
|
||||
/**
|
||||
* CB-559: the template is read per spawn, not captured at construction. This is what makes
|
||||
* {@code fleet.tabLabel} a hot key — a launcher built at boot must see an edit made an hour later
|
||||
* without being rebuilt.
|
||||
*/
|
||||
@Test
|
||||
void theTemplateIsReadOnEverySpawnSoAnEditTakesEffect() {
|
||||
FakeHerdr herdr = new FakeHerdr();
|
||||
AtomicReference<String> template = new AtomicReference<>("{role}: {profile} #{n}");
|
||||
ClaudeCodeLauncher svc = labelService(herdr, () -> fleet(null, template.get()));
|
||||
|
||||
svc.spawn(new SpawnRequest("sonnet", null, null, null, null, MemberRole.DEV));
|
||||
template.set("[{profile}] {role} {n}");
|
||||
svc.spawn(new SpawnRequest("sonnet", null, null, null, null, MemberRole.DEV));
|
||||
|
||||
assertEquals(List.of("dev: sonnet #1", "[sonnet] dev 2"), tabLabels(herdr));
|
||||
}
|
||||
}
|
||||
+185
-25
@@ -1,4 +1,4 @@
|
||||
package dev.ltms.bridged.worker;
|
||||
package dev.ltms.bridged.member;
|
||||
|
||||
import ch.qos.logback.classic.Logger;
|
||||
import ch.qos.logback.classic.spi.ILoggingEvent;
|
||||
@@ -10,6 +10,7 @@ import dev.ltms.bridged.herdr.AgentControl;
|
||||
import dev.ltms.bridged.herdr.FakeHerdr;
|
||||
import dev.ltms.bridged.herdr.WorkspaceControl;
|
||||
import dev.ltms.bridged.peer.Capability;
|
||||
import dev.ltms.bridged.peer.MemberRole;
|
||||
import dev.ltms.bridged.peer.PeerHandle;
|
||||
import dev.ltms.bridged.peer.PeerLauncher;
|
||||
import dev.ltms.bridged.peer.PeerUnreachableException;
|
||||
@@ -21,12 +22,10 @@ import org.slf4j.LoggerFactory;
|
||||
|
||||
import java.util.EnumSet;
|
||||
import java.util.HashMap;
|
||||
import java.util.HashSet;
|
||||
import java.util.LinkedHashMap;
|
||||
import java.util.List;
|
||||
import java.util.Map;
|
||||
import java.util.Set;
|
||||
import java.util.function.Function;
|
||||
|
||||
import static org.junit.jupiter.api.Assertions.*;
|
||||
|
||||
@@ -40,7 +39,7 @@ class CompositePeerLauncherTest {
|
||||
|
||||
private ClaudeCodeLauncher claudeAdapter(FakeHerdr herdr) {
|
||||
// 12-arg back-compat Worker ctor → kind defaults to claude-code.
|
||||
BridgedConfig.Worker claude = new BridgedConfig.Worker("claude", "http://gx00.gw:8000", "coder",
|
||||
BridgedConfig.Profile claude = new BridgedConfig.Profile("claude", "http://gx00.gw:8000", "coder",
|
||||
null, "BRIDGED_WORKER_TOKEN", List.of("claude"), "tab", "bridged-workers", "w #{n}",
|
||||
null, null, null);
|
||||
return new ClaudeCodeLauncher(new AgentControl(herdr), new WorkspaceControl(herdr),
|
||||
@@ -48,9 +47,9 @@ class CompositePeerLauncherTest {
|
||||
}
|
||||
|
||||
private OpenCodeLauncher opencodeAdapter(FakeHerdr herdr) {
|
||||
BridgedConfig.Worker gemini = new BridgedConfig.Worker("gemini", null, "google/gemini-2.5-pro",
|
||||
BridgedConfig.Profile gemini = new BridgedConfig.Profile("gemini", null, "google/gemini-2.5-pro",
|
||||
null, "BRIDGED_WORKER_TOKEN", List.of("opencode"), "tab", "bridged-workers", "w #{n}",
|
||||
null, null, null, "GITEA_ACCESS_TOKEN", null, BridgedConfig.Worker.KIND_OPENCODE);
|
||||
null, null, null, "GITEA_ACCESS_TOKEN", null, BridgedConfig.Profile.KIND_OPENCODE);
|
||||
return new OpenCodeLauncher(new AgentControl(herdr), new WorkspaceControl(herdr),
|
||||
Map.of("gemini", gemini), "gemini", _ -> "tok");
|
||||
}
|
||||
@@ -70,7 +69,7 @@ class CompositePeerLauncherTest {
|
||||
private final Map<String, Integer> spawnCounts = new HashMap<>();
|
||||
|
||||
StubLauncher(String prefix, FakeHerdr herdr,
|
||||
Map<String, BridgedConfig.Worker> profiles, String defaultProfile,
|
||||
Map<String, BridgedConfig.Profile> profiles, String defaultProfile,
|
||||
Set<String> failProfiles) {
|
||||
super(prefix, new AgentControl(herdr), new WorkspaceControl(herdr),
|
||||
profiles, defaultProfile, _ -> null, 0L, System::currentTimeMillis, () -> { });
|
||||
@@ -78,7 +77,7 @@ class CompositePeerLauncherTest {
|
||||
}
|
||||
|
||||
@Override
|
||||
protected Launch buildLaunch(BridgedConfig.Worker cfg) {
|
||||
protected Launch buildLaunch(BridgedConfig.Profile cfg, LaunchSpec spec) {
|
||||
return new Launch(Map.of(), List.of());
|
||||
}
|
||||
|
||||
@@ -114,14 +113,14 @@ class CompositePeerLauncherTest {
|
||||
}
|
||||
}
|
||||
|
||||
private static BridgedConfig.Worker stubWorker(String profile) {
|
||||
return new BridgedConfig.Worker(profile, "http://gx00.gw:8000", "coder",
|
||||
private static BridgedConfig.Profile stubWorker(String profile) {
|
||||
return new BridgedConfig.Profile(profile, "http://gx00.gw:8000", "coder",
|
||||
null, "BRIDGED_WORKER_TOKEN", List.of("claude"), "tab", "bridged-workers",
|
||||
"w #{n}", null, null, null, null, null, null, null, null, null);
|
||||
}
|
||||
|
||||
private static BridgedConfig.Worker stubWorker(String profile, float weight, Integer maxLoad) {
|
||||
return new BridgedConfig.Worker(profile, "http://gx00.gw:8000", "coder",
|
||||
private static BridgedConfig.Profile stubWorker(String profile, float weight, Integer maxLoad) {
|
||||
return new BridgedConfig.Profile(profile, "http://gx00.gw:8000", "coder",
|
||||
null, "BRIDGED_WORKER_TOKEN", List.of("claude"), "tab", "bridged-workers",
|
||||
"w #{n}", null, null, null, null, null, null, null,
|
||||
weight, maxLoad);
|
||||
@@ -133,9 +132,9 @@ class CompositePeerLauncherTest {
|
||||
* so a {@code Map.of} would make "which profile is tried first" a coin flip per run and any
|
||||
* assertion about the first attempt intermittently false.
|
||||
*/
|
||||
private static Map<String, BridgedConfig.Worker> ordered(String first, BridgedConfig.Worker a,
|
||||
String second, BridgedConfig.Worker b) {
|
||||
Map<String, BridgedConfig.Worker> m = new LinkedHashMap<>();
|
||||
private static Map<String, BridgedConfig.Profile> ordered(String first, BridgedConfig.Profile a,
|
||||
String second, BridgedConfig.Profile b) {
|
||||
Map<String, BridgedConfig.Profile> m = new LinkedHashMap<>();
|
||||
m.put(first, a);
|
||||
m.put(second, b);
|
||||
return m;
|
||||
@@ -292,7 +291,7 @@ class CompositePeerLauncherTest {
|
||||
@Test
|
||||
void weightedPolicyGatesProfileAtMaxLoad() {
|
||||
FakeHerdr herdr = new FakeHerdr();
|
||||
Map<String, BridgedConfig.Worker> profiles = ordered(
|
||||
Map<String, BridgedConfig.Profile> profiles = ordered(
|
||||
"a", stubWorker("a", 1.0f, 1),
|
||||
"b", stubWorker("b", 1.0f, null));
|
||||
StubLauncher adapter = new StubLauncher("claude", herdr, profiles, "a", Set.of());
|
||||
@@ -308,12 +307,12 @@ class CompositePeerLauncherTest {
|
||||
@Test
|
||||
void weightedPolicyDistributesAccordingToWeightRatio() {
|
||||
FakeHerdr herdr = new FakeHerdr();
|
||||
Map<String, BridgedConfig.Worker> profiles = ordered(
|
||||
Map<String, BridgedConfig.Profile> profiles = ordered(
|
||||
"a", stubWorker("a", 0.75f, null),
|
||||
"b", stubWorker("b", 0.25f, null));
|
||||
StubLauncher adapter = new StubLauncher("claude", herdr, profiles, "a", Set.of());
|
||||
CompositePeerLauncher composite = new CompositePeerLauncher(
|
||||
List.of(adapter), "a", profiles, PlacementPolicies.weighted(), name -> 0);
|
||||
List.of(adapter), "a", profiles, PlacementPolicies.weighted(), _ -> 0);
|
||||
|
||||
int a = 0, b = 0;
|
||||
for (int i = 0; i < 40; i++) {
|
||||
@@ -328,12 +327,12 @@ class CompositePeerLauncherTest {
|
||||
@Test
|
||||
void failoverRetriesNextCandidateWhenProfileIsUnreachable() {
|
||||
FakeHerdr herdr = new FakeHerdr();
|
||||
Map<String, BridgedConfig.Worker> profiles = ordered(
|
||||
Map<String, BridgedConfig.Profile> profiles = ordered(
|
||||
"a", stubWorker("a"),
|
||||
"b", stubWorker("b"));
|
||||
StubLauncher adapter = new StubLauncher("claude", herdr, profiles, "a", Set.of("a"));
|
||||
CompositePeerLauncher composite = new CompositePeerLauncher(
|
||||
List.of(adapter), "a", profiles, PlacementPolicies.weighted(), name -> 0);
|
||||
List.of(adapter), "a", profiles, PlacementPolicies.weighted(), _ -> 0);
|
||||
|
||||
PeerHandle h = composite.spawn(new SpawnRequest(null, null, null));
|
||||
assertEquals("b", h.profile(), "the spawn must fail over from unreachable a to b");
|
||||
@@ -349,7 +348,7 @@ class CompositePeerLauncherTest {
|
||||
@Test
|
||||
void reversingDefinitionOrderReversesWhichProfileIsTriedFirst() {
|
||||
FakeHerdr herdr = new FakeHerdr();
|
||||
Map<String, BridgedConfig.Worker> profiles = ordered(
|
||||
Map<String, BridgedConfig.Profile> profiles = ordered(
|
||||
"b", stubWorker("b"),
|
||||
"a", stubWorker("a"));
|
||||
StubLauncher adapter = new StubLauncher("claude", herdr, profiles, "a", Set.of("b"));
|
||||
@@ -364,12 +363,12 @@ class CompositePeerLauncherTest {
|
||||
@Test
|
||||
void failoverBoundedByCandidateCount() {
|
||||
FakeHerdr herdr = new FakeHerdr();
|
||||
Map<String, BridgedConfig.Worker> profiles = ordered(
|
||||
Map<String, BridgedConfig.Profile> profiles = ordered(
|
||||
"a", stubWorker("a"),
|
||||
"b", stubWorker("b"));
|
||||
StubLauncher adapter = new StubLauncher("claude", herdr, profiles, "a", Set.of("a", "b"));
|
||||
CompositePeerLauncher composite = new CompositePeerLauncher(
|
||||
List.of(adapter), "a", profiles, PlacementPolicies.weighted(), name -> 0);
|
||||
List.of(adapter), "a", profiles, PlacementPolicies.weighted(), _ -> 0);
|
||||
|
||||
PeerUnreachableException e = assertThrows(PeerUnreachableException.class,
|
||||
() -> composite.spawn(new SpawnRequest(null, null, null)));
|
||||
@@ -378,18 +377,179 @@ class CompositePeerLauncherTest {
|
||||
assertEquals(1, adapter.spawnCount("b"));
|
||||
}
|
||||
|
||||
@Test
|
||||
void explicitSpawnAtMaxLoadThrowsPlacementExceptionNamingProfileLiveAndCap() {
|
||||
FakeHerdr herdr = new FakeHerdr();
|
||||
Map<String, BridgedConfig.Profile> profiles = ordered(
|
||||
"a", stubWorker("a", 1.0f, 2),
|
||||
"b", stubWorker("b"));
|
||||
StubLauncher adapter = new StubLauncher("claude", herdr, profiles, "a", Set.of());
|
||||
CompositePeerLauncher composite = new CompositePeerLauncher(
|
||||
List.of(adapter), "a", profiles, PlacementPolicies.fixed(), name -> "a".equals(name) ? 2 : 0);
|
||||
|
||||
PlacementException e = assertThrows(PlacementException.class,
|
||||
() -> composite.spawn(new SpawnRequest("a", null, null)));
|
||||
assertTrue(e.getMessage().contains("'a'"), "message names the profile: " + e.getMessage());
|
||||
assertTrue(e.getMessage().contains("2 live"), "message names the live count: " + e.getMessage());
|
||||
assertTrue(e.getMessage().contains("2 cap"), "message names the cap: " + e.getMessage());
|
||||
assertEquals(0, adapter.spawnCount("a"), "at cap, the spawn is refused before any delegation");
|
||||
}
|
||||
|
||||
@Test
|
||||
void explicitSpawnUnderMaxLoadStillSucceeds() {
|
||||
FakeHerdr herdr = new FakeHerdr();
|
||||
Map<String, BridgedConfig.Profile> profiles = ordered(
|
||||
"a", stubWorker("a", 1.0f, 2),
|
||||
"b", stubWorker("b"));
|
||||
StubLauncher adapter = new StubLauncher("claude", herdr, profiles, "a", Set.of());
|
||||
CompositePeerLauncher composite = new CompositePeerLauncher(
|
||||
List.of(adapter), "a", profiles, PlacementPolicies.fixed(), name -> "a".equals(name) ? 1 : 0);
|
||||
|
||||
PeerHandle h = composite.spawn(new SpawnRequest("a", null, null));
|
||||
assertEquals("a", h.profile(), "a profile under its cap accepts an explicit spawn");
|
||||
assertEquals(1, adapter.spawnCount("a"), "the under-cap spawn is delegated");
|
||||
}
|
||||
|
||||
@Test
|
||||
void explicitSpawnWithNullMaxLoadIsNeverCapped() {
|
||||
FakeHerdr herdr = new FakeHerdr();
|
||||
Map<String, BridgedConfig.Profile> profiles = ordered(
|
||||
"a", stubWorker("a", 1.0f, null),
|
||||
"b", stubWorker("b"));
|
||||
StubLauncher adapter = new StubLauncher("claude", herdr, profiles, "a", Set.of());
|
||||
// A deliberately absurd live count: an unset maxLoad means unlimited, so it must never refuse.
|
||||
CompositePeerLauncher composite = new CompositePeerLauncher(
|
||||
List.of(adapter), "a", profiles, PlacementPolicies.fixed(), _ -> 1000);
|
||||
|
||||
PeerHandle h = composite.spawn(new SpawnRequest("a", null, null));
|
||||
assertEquals("a", h.profile(), "a profile with no maxLoad is never capped, however many live workers");
|
||||
}
|
||||
|
||||
@Test
|
||||
void emptyCandidateSetThrowsClearException() {
|
||||
FakeHerdr herdr = new FakeHerdr();
|
||||
Map<String, BridgedConfig.Worker> profiles = ordered(
|
||||
Map<String, BridgedConfig.Profile> profiles = ordered(
|
||||
"a", stubWorker("a", 1.0f, 1),
|
||||
"b", stubWorker("b", 1.0f, 1));
|
||||
StubLauncher adapter = new StubLauncher("claude", herdr, profiles, "a", Set.of());
|
||||
CompositePeerLauncher composite = new CompositePeerLauncher(
|
||||
List.of(adapter), "a", profiles, PlacementPolicies.weighted(), name -> 1);
|
||||
List.of(adapter), "a", profiles, PlacementPolicies.weighted(), _ -> 1);
|
||||
|
||||
PlacementException e = assertThrows(PlacementException.class,
|
||||
() -> composite.spawn(new SpawnRequest(null, null, null)));
|
||||
assertTrue(e.getMessage().contains("maxLoad"), e.getMessage());
|
||||
}
|
||||
|
||||
// ── CB-557: an unqualified spawn is placed inside its role's pool ─────────────────────────
|
||||
|
||||
/** Three profiles in definition order — pools are carved out of this set. */
|
||||
private static Map<String, BridgedConfig.Profile> threeProfiles() {
|
||||
Map<String, BridgedConfig.Profile> m = new LinkedHashMap<>();
|
||||
m.put("opus", stubWorker("opus", 1.0f, null));
|
||||
m.put("sonnet", stubWorker("sonnet", 1.0f, null));
|
||||
m.put("terra", stubWorker("terra", 1.0f, null));
|
||||
return m;
|
||||
}
|
||||
|
||||
private static Map<String, BridgedConfig.Slot> pool(String... names) {
|
||||
Map<String, BridgedConfig.Slot> m = new LinkedHashMap<>();
|
||||
for (String n : names) {
|
||||
m.put(n, new BridgedConfig.Slot(n));
|
||||
}
|
||||
return m;
|
||||
}
|
||||
|
||||
private static CompositePeerLauncher withPools(FakeHerdr herdr, BridgedConfig.Fleet fleet) {
|
||||
Map<String, BridgedConfig.Profile> profiles = threeProfiles();
|
||||
StubLauncher adapter = new StubLauncher("claude", herdr, profiles, "opus", Set.of());
|
||||
return new CompositePeerLauncher(List.of(adapter), "opus", profiles,
|
||||
PlacementPolicies.fixed(), _ -> 0, fleet);
|
||||
}
|
||||
|
||||
/**
|
||||
* The point of the pools: a role is placed only on a backend its pool names. Before CB-557 an
|
||||
* unqualified spawn ranged over every configured profile, so a reviewer could land on the
|
||||
* architect-only one.
|
||||
*/
|
||||
@Test
|
||||
void anUnqualifiedSpawnIsPlacedInsideItsRolePool() {
|
||||
FakeHerdr herdr = new FakeHerdr();
|
||||
CompositePeerLauncher composite = withPools(herdr, new BridgedConfig.Fleet(
|
||||
Map.of(), pool("opus"), pool("terra"), pool("sonnet"), null));
|
||||
|
||||
assertEquals("opus", composite.spawn(
|
||||
new SpawnRequest(null, null, null, null, null, MemberRole.ARCHITECT)).profile());
|
||||
assertEquals("terra", composite.spawn(
|
||||
new SpawnRequest(null, null, null, null, null, MemberRole.DEV)).profile());
|
||||
assertEquals("sonnet", composite.spawn(
|
||||
new SpawnRequest(null, null, null, null, null, MemberRole.REVIEWER)).profile());
|
||||
}
|
||||
|
||||
/** Under `fixed`, the pool's first entry wins — not the global defaultProfile. */
|
||||
@Test
|
||||
void theRolePoolOutranksTheGlobalDefaultProfile() {
|
||||
FakeHerdr herdr = new FakeHerdr();
|
||||
CompositePeerLauncher composite = withPools(herdr, new BridgedConfig.Fleet(
|
||||
Map.of(), Map.of(), pool("sonnet", "terra"), Map.of(), null));
|
||||
|
||||
assertEquals("sonnet", composite.spawn(
|
||||
new SpawnRequest(null, null, null, null, null, MemberRole.DEV)).profile(),
|
||||
"the dev pool starts at sonnet, so the global default 'opus' must not win");
|
||||
}
|
||||
|
||||
/**
|
||||
* A role with no pool is unconstrained, not blocked. A config that declares pools for some roles
|
||||
* and not others must keep spawning the rest.
|
||||
*/
|
||||
@Test
|
||||
void aRoleWithNoPoolFallsBackToEveryProfile() {
|
||||
FakeHerdr herdr = new FakeHerdr();
|
||||
CompositePeerLauncher composite = withPools(herdr, new BridgedConfig.Fleet(
|
||||
Map.of(), pool("sonnet"), Map.of(), Map.of(), null));
|
||||
|
||||
assertEquals("opus", composite.spawn(
|
||||
new SpawnRequest(null, null, null, null, null, MemberRole.DEV)).profile(),
|
||||
"no dev pool ⇒ all profiles are candidates, so `fixed` takes the first one");
|
||||
}
|
||||
|
||||
/** No fleet at all is the pre-CB-557 wiring, and must behave exactly as it did. */
|
||||
@Test
|
||||
void noFleetConfiguredKeepsTheOldWholeProfileListBehaviour() {
|
||||
FakeHerdr herdr = new FakeHerdr();
|
||||
CompositePeerLauncher composite = withPools(herdr, null);
|
||||
|
||||
assertEquals("opus", composite.spawn(new SpawnRequest(null, null, null)).profile());
|
||||
}
|
||||
|
||||
/**
|
||||
* An explicit profile is the operator overriding and is NOT judged against the pool. It must
|
||||
* stay that way: an unrolled `bridge_spawn{profile:"opus"}` carries no role, so it defaults to
|
||||
* DEV, and enforcing the pool here would refuse a spawn the operator asked for by name.
|
||||
*/
|
||||
@Test
|
||||
void anExplicitProfileIsNotConfinedToTheRolePool() {
|
||||
FakeHerdr herdr = new FakeHerdr();
|
||||
CompositePeerLauncher composite = withPools(herdr, new BridgedConfig.Fleet(
|
||||
Map.of(), pool("opus"), pool("terra"), Map.of(), null));
|
||||
|
||||
assertEquals("opus", composite.spawn(new SpawnRequest("opus", null, null)).profile(),
|
||||
"naming opus explicitly must work even though the dev pool holds only terra");
|
||||
}
|
||||
|
||||
/** Placement still respects maxLoad, but only across the pool — never by escaping it. */
|
||||
@Test
|
||||
void aFullPoolIsRefusedRatherThanSpilledOntoAnotherRolesProfile() {
|
||||
FakeHerdr herdr = new FakeHerdr();
|
||||
Map<String, BridgedConfig.Profile> profiles = new LinkedHashMap<>();
|
||||
profiles.put("opus", stubWorker("opus", 1.0f, null)); // architect-only, uncapped
|
||||
profiles.put("terra", stubWorker("terra", 1.0f, 1)); // the sole dev, capped at 1
|
||||
StubLauncher adapter = new StubLauncher("claude", herdr, profiles, "opus", Set.of());
|
||||
CompositePeerLauncher composite = new CompositePeerLauncher(List.of(adapter), "opus", profiles,
|
||||
PlacementPolicies.weighted(), name -> "terra".equals(name) ? 1 : 0,
|
||||
new BridgedConfig.Fleet(Map.of(), pool("opus"), pool("terra"), Map.of(), null));
|
||||
|
||||
PlacementException e = assertThrows(PlacementException.class, () -> composite.spawn(
|
||||
new SpawnRequest(null, null, null, null, null, MemberRole.DEV)));
|
||||
assertTrue(e.getMessage().contains("maxLoad"), e.getMessage());
|
||||
}
|
||||
}
|
||||
@@ -0,0 +1,71 @@
|
||||
package dev.ltms.bridged.member;
|
||||
|
||||
import dev.ltms.bridged.config.BridgedConfig;
|
||||
import dev.ltms.bridged.herdr.AgentControl;
|
||||
import dev.ltms.bridged.herdr.FakeHerdr;
|
||||
import dev.ltms.bridged.herdr.WorkspaceControl;
|
||||
import dev.ltms.bridged.peer.Capability;
|
||||
import dev.ltms.bridged.peer.MemberRole;
|
||||
import dev.ltms.bridged.peer.SpawnRequest;
|
||||
import org.junit.jupiter.api.Test;
|
||||
|
||||
import java.util.ArrayList;
|
||||
import java.util.List;
|
||||
import java.util.Map;
|
||||
import java.util.Set;
|
||||
import java.util.concurrent.atomic.AtomicReference;
|
||||
import java.util.function.Supplier;
|
||||
|
||||
import static org.junit.jupiter.api.Assertions.assertEquals;
|
||||
|
||||
class HerdrPeerLauncherCharterTest {
|
||||
|
||||
@Test
|
||||
void readsAndComposesTheFleetCharterForEachSpawn() {
|
||||
AtomicReference<BridgedConfig.Fleet> fleet = new AtomicReference<>(fleet(Map.of()));
|
||||
CapturingLauncher launcher = new CapturingLauncher(fleet::get);
|
||||
|
||||
launcher.spawn(new SpawnRequest("mcp", null, null, null, null, MemberRole.DEV));
|
||||
fleet.set(fleet(Map.of("dev", "role charter")));
|
||||
launcher.spawn(new SpawnRequest("mcp", null, null, null, null, MemberRole.DEV));
|
||||
launcher.spawn(new SpawnRequest("no-mcp", null, null, null, null, MemberRole.DEV));
|
||||
|
||||
assertEquals(HerdrPeerLauncher.REPLY_CHARTER, launcher.specs.get(0).charter(),
|
||||
"without a role charter, MCP profiles receive only the reply charter");
|
||||
assertEquals("role charter\n\n" + HerdrPeerLauncher.REPLY_CHARTER, launcher.specs.get(1).charter(),
|
||||
"the changed supplier value is read for the next spawn and the reply rule is last");
|
||||
assertEquals("role charter", launcher.specs.get(2).charter(),
|
||||
"a role charter does not depend on an MCP mount");
|
||||
}
|
||||
|
||||
private static BridgedConfig.Fleet fleet(Map<String, String> charters) {
|
||||
return new BridgedConfig.Fleet(Map.of(), Map.of(), Map.of(), Map.of(), charters, null);
|
||||
}
|
||||
|
||||
private static final class CapturingLauncher extends HerdrPeerLauncher {
|
||||
private final List<LaunchSpec> specs = new ArrayList<>();
|
||||
|
||||
CapturingLauncher(Supplier<BridgedConfig.Fleet> fleet) {
|
||||
super("test", new AgentControl(new FakeHerdr()), new WorkspaceControl(new FakeHerdr()),
|
||||
Map.of("mcp", profile("mcp", "http://bridge"),
|
||||
"no-mcp", profile("no-mcp", null)),
|
||||
"mcp", _ -> null, 0, () -> 0L, () -> { }, fleet);
|
||||
}
|
||||
|
||||
@Override
|
||||
protected Launch buildLaunch(BridgedConfig.Profile cfg, LaunchSpec spec) {
|
||||
specs.add(spec);
|
||||
return new Launch(Map.of(), List.of("test"));
|
||||
}
|
||||
|
||||
@Override
|
||||
public Set<Capability> capabilities() {
|
||||
return Set.of();
|
||||
}
|
||||
|
||||
private static BridgedConfig.Profile profile(String name, String mcpUrl) {
|
||||
return new BridgedConfig.Profile(name, "http://gx00.gw:8000", null, null,
|
||||
"BRIDGED_WORKER_TOKEN", List.of("test"), "pane", null, null, mcpUrl, null, null);
|
||||
}
|
||||
}
|
||||
}
|
||||
+96
-16
@@ -1,4 +1,4 @@
|
||||
package dev.ltms.bridged.worker;
|
||||
package dev.ltms.bridged.member;
|
||||
|
||||
import com.fasterxml.jackson.databind.JsonNode;
|
||||
import com.fasterxml.jackson.databind.ObjectMapper;
|
||||
@@ -17,6 +17,10 @@ import java.nio.file.Files;
|
||||
import java.nio.file.Path;
|
||||
import java.util.List;
|
||||
import java.util.Map;
|
||||
import java.util.concurrent.ExecutorService;
|
||||
import java.util.concurrent.Executors;
|
||||
import java.util.concurrent.Future;
|
||||
import java.util.function.Supplier;
|
||||
|
||||
import static org.junit.jupiter.api.Assertions.*;
|
||||
|
||||
@@ -27,19 +31,26 @@ import static org.junit.jupiter.api.Assertions.*;
|
||||
*/
|
||||
class OpenCodeLauncherTest {
|
||||
|
||||
private static BridgedConfig.Worker opencodeCfg(String model, String mcpUrl, String gitTokenEnv) {
|
||||
return new BridgedConfig.Worker("gemini", null, model, null, "BRIDGED_WORKER_TOKEN",
|
||||
private static BridgedConfig.Profile opencodeCfg(String model, String mcpUrl, String gitTokenEnv) {
|
||||
return new BridgedConfig.Profile("gemini", null, model, null, "BRIDGED_WORKER_TOKEN",
|
||||
List.of("opencode"), "tab", "bridged-workers", "opencode: {model} #{n}", mcpUrl,
|
||||
null, null, gitTokenEnv, null, BridgedConfig.Worker.KIND_OPENCODE);
|
||||
null, null, gitTokenEnv, null, BridgedConfig.Profile.KIND_OPENCODE);
|
||||
}
|
||||
|
||||
/** Gate-disabled launcher whose per-spawn config dirs land under an inspectable temp root. */
|
||||
private OpenCodeLauncher service(FakeHerdr herdr, Path configRoot, BridgedConfig.Worker cfg) {
|
||||
private static OpenCodeLauncher service(FakeHerdr herdr, Path configRoot, BridgedConfig.Profile cfg) {
|
||||
return new OpenCodeLauncher(new AgentControl(herdr), new WorkspaceControl(herdr),
|
||||
Map.of(cfg.profile(), cfg), cfg.profile(), k -> "GITEA_ACCESS_TOKEN".equals(k) ? "tok" : null,
|
||||
0, System::currentTimeMillis, () -> { }, configRoot, configRoot);
|
||||
}
|
||||
|
||||
private static OpenCodeLauncher service(FakeHerdr herdr, Path configRoot, BridgedConfig.Profile cfg,
|
||||
Supplier<BridgedConfig.Fleet> fleet) {
|
||||
return new OpenCodeLauncher(new AgentControl(herdr), new WorkspaceControl(herdr),
|
||||
Map.of(cfg.profile(), cfg), cfg.profile(), _ -> null,
|
||||
0, System::currentTimeMillis, () -> { }, configRoot, configRoot, fleet);
|
||||
}
|
||||
|
||||
@SuppressWarnings("unchecked")
|
||||
private static Map<String, Object> lastStart(FakeHerdr herdr) {
|
||||
return (Map<String, Object>) herdr.lastCall("agent.start").params();
|
||||
@@ -62,8 +73,10 @@ class OpenCodeLauncherTest {
|
||||
@Test
|
||||
void writesRemoteMcpConfigAndCharterInstructionsWhenMcpUrlSet(@TempDir Path root) throws Exception {
|
||||
FakeHerdr herdr = new FakeHerdr();
|
||||
service(herdr, root, opencodeCfg("google/gemini-2.5-pro", "http://127.0.0.1:8765/mcp", null))
|
||||
.spawn();
|
||||
BridgedConfig.Fleet fleet = new BridgedConfig.Fleet(Map.of(), Map.of(), Map.of(), Map.of(),
|
||||
Map.of("dev", "role rule"), null);
|
||||
service(herdr, root, opencodeCfg("google/gemini-2.5-pro", "http://127.0.0.1:8765/mcp", null),
|
||||
() -> fleet).spawn();
|
||||
|
||||
Map<String, String> env = startEnv(herdr);
|
||||
assertNull(env.get("ANTHROPIC_BASE_URL"), "opencode carries no ANTHROPIC_* / subscription boundary");
|
||||
@@ -74,19 +87,23 @@ class OpenCodeLauncherTest {
|
||||
// Assert on parsed structure, not substrings: the generated config is real JSON and its
|
||||
// whitespace is the formatter's business, not the contract's.
|
||||
JsonNode json = new ObjectMapper().readTree(Path.of(cfgPath).toFile());
|
||||
assertTrue(json.path("compaction").path("auto").asBoolean(),
|
||||
"spawned opencode peers explicitly enable automatic compaction");
|
||||
JsonNode bridge = json.path("mcp").path("bridge");
|
||||
assertEquals("remote", bridge.path("type").asText(), "bridge is mounted as a remote MCP server");
|
||||
assertEquals("http://127.0.0.1:8765/mcp", bridge.path("url").asText(),
|
||||
"the profile's bridge MCP url is present");
|
||||
assertTrue(bridge.path("enabled").asBoolean(), "the bridge server is enabled");
|
||||
assertTrue(json.path("instructions").isArray() && !json.path("instructions").isEmpty(),
|
||||
"the reply charter is mounted via instructions");
|
||||
"the member charter is mounted via instructions");
|
||||
|
||||
// The instructions entry is a real file path holding the reply charter.
|
||||
Path charter = Path.of(cfgPath).resolveSibling("reply-charter.md");
|
||||
// The instructions entry is a real file path holding the composed member charter.
|
||||
Path charter = Path.of(cfgPath).resolveSibling("member-charter.md");
|
||||
assertTrue(Files.exists(charter), "the charter file the config references was written");
|
||||
assertTrue(Files.readString(charter).contains("bridge_reply"),
|
||||
"the charter instructs the worker to answer via bridge_reply");
|
||||
assertEquals("role rule\n\n" + HerdrPeerLauncher.REPLY_CHARTER, Files.readString(charter),
|
||||
"the composed charter keeps the role rule first and the reply rule last");
|
||||
assertEquals(charter.toAbsolutePath().toString(), json.path("instructions").get(0).asText(),
|
||||
"instructions names the charter file by its absolute path");
|
||||
}
|
||||
|
||||
@Test
|
||||
@@ -98,6 +115,69 @@ class OpenCodeLauncherTest {
|
||||
"no bridge MCP url → no config file and no OPENCODE_CONFIG");
|
||||
}
|
||||
|
||||
@Test
|
||||
void roleCharterWithoutMcpOrCustomProviderStillWritesAConfig(@TempDir Path root) throws Exception {
|
||||
FakeHerdr herdr = new FakeHerdr();
|
||||
BridgedConfig.Fleet fleet = new BridgedConfig.Fleet(Map.of(), Map.of(), Map.of(), Map.of(),
|
||||
Map.of("dev", "role rule"), null);
|
||||
Path configRoot = Files.createDirectory(root.resolve("configs"));
|
||||
Path checkout = Files.createDirectory(root.resolve("checkout"));
|
||||
service(herdr, configRoot, opencodeCfg("google/gemini-2.5-pro", null, null), () -> fleet).spawn();
|
||||
|
||||
String cfgPath = startEnv(herdr).get("OPENCODE_CONFIG");
|
||||
assertNotNull(cfgPath, "a role charter needs a config even without MCP or custom provider");
|
||||
JsonNode json = new ObjectMapper().readTree(Path.of(cfgPath).toFile());
|
||||
Path charter = Path.of(json.path("instructions").get(0).asText());
|
||||
assertEquals("role rule", Files.readString(charter), "the base-composed role charter is unchanged");
|
||||
assertTrue(json.path("mcp").isMissingNode(), "a charter does not add an MCP mount");
|
||||
assertTrue(charter.startsWith(configRoot), "the charter is written under the temp config root");
|
||||
try (var files = Files.walk(checkout)) {
|
||||
assertFalse(files.anyMatch(path -> path.getFileName().toString().equals("member-charter.md")),
|
||||
"the worker checkout receives no charter file");
|
||||
}
|
||||
}
|
||||
|
||||
@Test
|
||||
void nullCharterWritesNoCharterFileOrInstructions(@TempDir Path root) throws Exception {
|
||||
FakeHerdr herdr = new FakeHerdr();
|
||||
service(herdr, root, pinnedCfg("local-vllm/model", "http://127.0.0.1:8000", null),
|
||||
() -> new BridgedConfig.Fleet(Map.of(), Map.of(), Map.of(), Map.of(), Map.of(), null)).spawn();
|
||||
|
||||
String config = startEnv(herdr).get("OPENCODE_CONFIG");
|
||||
assertNotNull(config, "the custom provider still needs a config");
|
||||
JsonNode json = new ObjectMapper().readTree(Path.of(config).toFile());
|
||||
assertTrue(json.path("instructions").isMissingNode(), "a null charter adds no instructions entry");
|
||||
try (var files = Files.walk(root)) {
|
||||
assertFalse(files.anyMatch(path -> path.getFileName().toString().equals("member-charter.md")),
|
||||
"a null charter creates no charter file");
|
||||
}
|
||||
}
|
||||
|
||||
@Test
|
||||
void concurrentSpawnsWriteSeparateCharterDirectories(@TempDir Path root) throws Exception {
|
||||
BridgedConfig.Profile cfg = opencodeCfg("google/gemini-2.5-pro", "http://127.0.0.1:8765/mcp", null);
|
||||
ExecutorService executor = Executors.newFixedThreadPool(2);
|
||||
try {
|
||||
Future<String> first = executor.submit(() -> spawnConfigPath(root, cfg));
|
||||
Future<String> second = executor.submit(() -> spawnConfigPath(root, cfg));
|
||||
|
||||
Path firstCharter = Path.of(first.get()).resolveSibling("member-charter.md");
|
||||
Path secondCharter = Path.of(second.get()).resolveSibling("member-charter.md");
|
||||
assertNotEquals(firstCharter.getParent(), secondCharter.getParent(),
|
||||
"each concurrent spawn owns a separate config directory");
|
||||
assertTrue(Files.exists(firstCharter));
|
||||
assertTrue(Files.exists(secondCharter));
|
||||
} finally {
|
||||
executor.shutdownNow();
|
||||
}
|
||||
}
|
||||
|
||||
private static String spawnConfigPath(Path root, BridgedConfig.Profile cfg) {
|
||||
FakeHerdr herdr = new FakeHerdr();
|
||||
service(herdr, root, cfg).spawn();
|
||||
return startEnv(herdr).get("OPENCODE_CONFIG");
|
||||
}
|
||||
|
||||
@Test
|
||||
void passesTheModelAsDashMFlagAlongsideAutoApprove(@TempDir Path root) {
|
||||
FakeHerdr herdr = new FakeHerdr();
|
||||
@@ -203,7 +283,7 @@ class OpenCodeLauncherTest {
|
||||
@Test
|
||||
void productionConstructorsWireThroughToTheBase() {
|
||||
FakeHerdr herdr = new FakeHerdr();
|
||||
BridgedConfig.Worker cfg = opencodeCfg(null, null, null);
|
||||
BridgedConfig.Profile cfg = opencodeCfg(null, null, null);
|
||||
// 5-arg (gate disabled) and 7-arg (gate enabled) production constructors both expose the profile.
|
||||
OpenCodeLauncher disabled = new OpenCodeLauncher(new AgentControl(herdr), new WorkspaceControl(herdr),
|
||||
Map.of(cfg.profile(), cfg), cfg.profile(), _ -> null);
|
||||
@@ -245,10 +325,10 @@ class OpenCodeLauncherTest {
|
||||
// --- CB-508: pinned OpenAI-compatible endpoint (e.g. a local vLLM) ---------------------------
|
||||
|
||||
/** A profile with a baseUrl but no model provider prefix cannot be resolved — fail loudly. */
|
||||
private static BridgedConfig.Worker pinnedCfg(String model, String baseUrl, String mcpUrl) {
|
||||
return new BridgedConfig.Worker("local", baseUrl, model, null, "BRIDGED_WORKER_TOKEN",
|
||||
private static BridgedConfig.Profile pinnedCfg(String model, String baseUrl, String mcpUrl) {
|
||||
return new BridgedConfig.Profile("local", baseUrl, model, null, "BRIDGED_WORKER_TOKEN",
|
||||
List.of("opencode"), "tab", "bridged-workers", "opencode: {model} #{n}", mcpUrl,
|
||||
null, null, null, null, BridgedConfig.Worker.KIND_OPENCODE);
|
||||
null, null, null, null, BridgedConfig.Profile.KIND_OPENCODE);
|
||||
}
|
||||
|
||||
@Test
|
||||
+1
-1
@@ -1,4 +1,4 @@
|
||||
package dev.ltms.bridged.worker;
|
||||
package dev.ltms.bridged.member;
|
||||
|
||||
import org.junit.jupiter.api.Test;
|
||||
import org.junit.jupiter.api.io.TempDir;
|
||||
@@ -0,0 +1,216 @@
|
||||
package dev.ltms.bridged.msg;
|
||||
|
||||
import dev.ltms.bridged.herdr.AgentStatus;
|
||||
import dev.ltms.bridged.mcp.PrimaryRegistry;
|
||||
import dev.ltms.bridged.peer.MemberRole;
|
||||
import dev.ltms.bridged.session.MemberSession;
|
||||
import org.junit.jupiter.api.AfterEach;
|
||||
import org.junit.jupiter.api.BeforeEach;
|
||||
import org.junit.jupiter.api.Test;
|
||||
|
||||
import java.util.List;
|
||||
import java.util.concurrent.Executors;
|
||||
import java.util.concurrent.ScheduledExecutorService;
|
||||
import java.util.concurrent.TimeUnit;
|
||||
|
||||
import static org.junit.jupiter.api.Assertions.*;
|
||||
|
||||
/**
|
||||
* Unit tests for the CB-551 idle-lead heartbeat: the pure {@link LeadHeartbeatLoop#decide} decision
|
||||
* function, the nudge text, and the {@link LeadHeartbeatLoop#snapshot} fleet snapshot.
|
||||
*
|
||||
* <p>All decision tests call {@code decide} directly with explicit nanoTime values from an injected
|
||||
* clock — no sleeping, no scheduler races. This mirrors how {@code ReplyPushLoopTest} pins the pure
|
||||
* decision before exercising the loop.
|
||||
*/
|
||||
class LeadHeartbeatLoopTest {
|
||||
|
||||
/** Arbitrary nanoTime origin for the fake clock. */
|
||||
private static final long NOW = 1_000_000_000L;
|
||||
/** 300s (the default quiet period) in nanos — under it the lead is "within the quiet period". */
|
||||
private static final long IDLE_AFTER_NANOS = TimeUnit.SECONDS.toNanos(300);
|
||||
/** 400s of idle — clearly past the quiet period. */
|
||||
private static final long IDLE_PAST = NOW - TimeUnit.SECONDS.toNanos(400);
|
||||
/** 10s of idle — clearly within the quiet period. */
|
||||
private static final long IDLE_WITHIN = NOW - TimeUnit.SECONDS.toNanos(10);
|
||||
|
||||
private ScheduledExecutorService scheduler;
|
||||
|
||||
@BeforeEach
|
||||
void setUp() {
|
||||
scheduler = Executors.newSingleThreadScheduledExecutor();
|
||||
}
|
||||
|
||||
@AfterEach
|
||||
void tearDown() {
|
||||
scheduler.shutdownNow();
|
||||
}
|
||||
|
||||
/** A fleet with nothing pending. */
|
||||
private static LeadHeartbeatLoop.FleetState quietFleet() {
|
||||
return new LeadHeartbeatLoop.FleetState(0, 0, 2, List.of());
|
||||
}
|
||||
|
||||
/** A fleet with a pending reply and a DONE session. */
|
||||
private static LeadHeartbeatLoop.FleetState pendingFleet() {
|
||||
return new LeadHeartbeatLoop.FleetState(2, 1, 3, List.of("term_a", "term_b"));
|
||||
}
|
||||
|
||||
/** The loop under test; the scheduler is never invoked on the pure decide path. */
|
||||
private static LeadHeartbeatLoop loop(int quietCap) {
|
||||
return new LeadHeartbeatLoop(
|
||||
new PrimaryRegistry("term_lead"), null /*agents — unused on the decide path*/,
|
||||
null /*inbox*/, List::of, null /*pushLoop*/, null /*scheduler*/, () -> 0L,
|
||||
IDLE_AFTER_NANOS, 1_000L, quietCap);
|
||||
}
|
||||
|
||||
// ── (a) a working lead is never injected ───────────────────────────────────────────────────
|
||||
|
||||
@Test
|
||||
void aWorkingLeadIsNeverInjected() {
|
||||
LeadHeartbeatLoop.Decision d = loop(3).decide(
|
||||
NOW, IDLE_PAST, 0, AgentStatus.WORKING, false, true, quietFleet());
|
||||
assertEquals(LeadHeartbeatLoop.Action.LEAD_BUSY, d.action(),
|
||||
"a WORKING lead is making progress and must not be touched");
|
||||
assertNull(d.idleSinceNanos(), "a busy lead resets the idle window");
|
||||
assertEquals(0, d.quietCount(), "a busy lead re-arms the quiet counter");
|
||||
}
|
||||
|
||||
@Test
|
||||
void anUnreadableStatusIsNeverInjectedEither() {
|
||||
// A failed status read (or a gone agent) must degrade to "do not inject", never hammer the pane.
|
||||
LeadHeartbeatLoop.Decision d = loop(3).decide(
|
||||
NOW, IDLE_PAST, 0, AgentStatus.UNKNOWN, false, true, pendingFleet());
|
||||
assertEquals(LeadHeartbeatLoop.Action.LEAD_BUSY, d.action(),
|
||||
"never inject into a state the loop cannot read");
|
||||
}
|
||||
|
||||
// ── (b) an idle lead within the quiet period is not yet injected ───────────────────────────
|
||||
|
||||
@Test
|
||||
void justBecameIdleStartsTheDebounceWindow() {
|
||||
LeadHeartbeatLoop.Decision d = loop(3).decide(
|
||||
NOW, null, 0, AgentStatus.IDLE, false, true, quietFleet());
|
||||
assertEquals(LeadHeartbeatLoop.Action.WAIT_IDLE, d.action(),
|
||||
"the first injectable tick only records the start of the idle stretch");
|
||||
assertEquals(NOW, d.idleSinceNanos(), "the idle window opens at the moment the lead became injectable");
|
||||
}
|
||||
|
||||
@Test
|
||||
void idleWithinQuietPeriodIsNotInjected() {
|
||||
LeadHeartbeatLoop.Decision d = loop(3).decide(
|
||||
NOW, IDLE_WITHIN, 0, AgentStatus.IDLE, false, true, quietFleet());
|
||||
assertEquals(LeadHeartbeatLoop.Action.WAIT_IDLE, d.action(),
|
||||
"a lead idle for 10s (< 300s) has just finished a turn — do not re-prompt it");
|
||||
}
|
||||
|
||||
// ── (c) an idle lead past the quiet period is injected ─────────────────────────────────────
|
||||
|
||||
@Test
|
||||
void idlePastQuietPeriodIsInjected() {
|
||||
LeadHeartbeatLoop.Decision d = loop(3).decide(
|
||||
NOW, IDLE_PAST, 0, AgentStatus.IDLE, false, true, quietFleet());
|
||||
assertEquals(LeadHeartbeatLoop.Action.INJECT, d.action(),
|
||||
"a lead continuously idle past the quiet period is the reason to nudge");
|
||||
}
|
||||
|
||||
@Test
|
||||
void blockedAndDoneAreInjectableViewsOfIdle() {
|
||||
assertEquals(LeadHeartbeatLoop.Action.INJECT, loop(3).decide(
|
||||
NOW, IDLE_PAST, 0, AgentStatus.BLOCKED, false, true, quietFleet()).action());
|
||||
assertEquals(LeadHeartbeatLoop.Action.INJECT, loop(3).decide(
|
||||
NOW, IDLE_PAST, 0, AgentStatus.DONE, false, true, quietFleet()).action());
|
||||
}
|
||||
|
||||
@Test
|
||||
void nothingIsInjectedWhenNoLeadIsKnown() {
|
||||
var d = loop(3).decide(NOW, IDLE_PAST, 0, AgentStatus.IDLE, false, false, pendingFleet());
|
||||
assertEquals(LeadHeartbeatLoop.Action.WAIT_IDLE, d.action(),
|
||||
"with no known lead there is nobody to nudge — keep waiting until one is discovered");
|
||||
assertEquals(IDLE_PAST, d.idleSinceNanos(), "the idle window stays open so discovery re-arms it");
|
||||
}
|
||||
|
||||
// ── (d) the quiet-nudge cap stops the loop ──────────────────────────────────────────────────
|
||||
|
||||
@Test
|
||||
void quietNudgeCapStopsTheLoopWhenNothingIsPending() {
|
||||
LeadHeartbeatLoop.Decision d = loop(3).decide(
|
||||
NOW, IDLE_PAST, 3, AgentStatus.IDLE, false, true, quietFleet());
|
||||
assertEquals(LeadHeartbeatLoop.Action.QUIET_DONE, d.action(),
|
||||
"3 consecutive nothing-pending nudges have already happened — stop nagging an empty fleet");
|
||||
}
|
||||
|
||||
// ── (e) new pending state resets the cap and re-arms the loop ───────────────────────────────
|
||||
|
||||
@Test
|
||||
void newPendingStateResetsTheQuietCap() {
|
||||
LeadHeartbeatLoop.Decision d = loop(3).decide(
|
||||
NOW, IDLE_PAST, 3, AgentStatus.IDLE, false, true, pendingFleet());
|
||||
assertEquals(LeadHeartbeatLoop.Action.INJECT, d.action(),
|
||||
"real state appearing re-arms the loop past an exhausted cap");
|
||||
assertEquals(0, d.quietCount(), "the pending state resets the consecutive-quiet counter");
|
||||
}
|
||||
|
||||
// ── constraint 6: never race ReplyPushLoop ─────────────────────────────────────────────────
|
||||
|
||||
@Test
|
||||
void standsDownWhileReplyPushLoopIsActive() {
|
||||
LeadHeartbeatLoop.Decision d = loop(3).decide(
|
||||
NOW, IDLE_PAST, 2, AgentStatus.IDLE, true, true, quietFleet());
|
||||
assertEquals(LeadHeartbeatLoop.Action.STAND_DOWN, d.action(),
|
||||
"a second injection would start a competing turn — stand aside instead");
|
||||
assertEquals(0, d.quietCount(), "the active push is real state, so it re-arms the cap");
|
||||
}
|
||||
|
||||
// ── nudge text (constraint 5: the nudge must carry state) ──────────────────────────────────
|
||||
|
||||
@Test
|
||||
void pendingNudgeCarriesTheFleetFacts() {
|
||||
String t = pendingFleet().nudgeText();
|
||||
assertTrue(t.startsWith("Heartbeat:"), "the nudge identifies itself as a heartbeat");
|
||||
assertTrue(t.contains("2 worker replies pending collection"), t);
|
||||
assertTrue(t.contains("bridge_poll(target=term_a)"), t);
|
||||
assertTrue(t.contains("bridge_poll(target=term_b)"), t);
|
||||
assertTrue(t.contains("1 DONE session awaiting teardown"), t);
|
||||
assertTrue(t.contains("3 workers live"), t);
|
||||
}
|
||||
|
||||
@Test
|
||||
void nothingPendingNudgeSaysExactlyThatSoTheLeadCanStandDown() {
|
||||
String t = quietFleet().nudgeText();
|
||||
assertTrue(t.contains("nothing pending to collect"), t);
|
||||
assertTrue(t.contains("0 worker replies"), t);
|
||||
assertTrue(t.contains("0 DONE sessions"), t);
|
||||
assertTrue(t.contains("2 workers live"), t);
|
||||
assertTrue(t.contains("stand down"), "allow the lead to stand down rather than hunt");
|
||||
}
|
||||
|
||||
// ── snapshot over roster + inbox ───────────────────────────────────────────────────────────
|
||||
|
||||
@Test
|
||||
void snapshotCountsRepliesDoneSessionsAndLiveWorkers() {
|
||||
InMemoryReplyInbox inbox = new InMemoryReplyInbox();
|
||||
inbox.own("term_w1");
|
||||
inbox.publish("term_w1", "m1", "hello");
|
||||
MemberSession done = new MemberSession("p1", "term_w1", "prof", MemberRole.DEV, "/cwd", null,
|
||||
0, 0, 0, MemberSession.State.DONE, null, null);
|
||||
MemberSession ready = new MemberSession("p2", "term_w2", "prof", MemberRole.DEV, "/cwd", null,
|
||||
0, 0, 0, MemberSession.State.READY, null, null);
|
||||
|
||||
LeadHeartbeatLoop.FleetState fs = LeadHeartbeatLoop.snapshot(inbox, () -> List.of(done, ready));
|
||||
|
||||
assertEquals(1, fs.pendingReplies(), "one undrained reply on term_w1");
|
||||
assertEquals(1, fs.doneSessions(), "term_w1 is DONE, awaiting teardown");
|
||||
assertEquals(2, fs.liveWorkers(), "both sessions are still registered");
|
||||
assertEquals(List.of("term_w1"), fs.replyTargets(), "the target with a pending reply is named");
|
||||
assertTrue(fs.hasPending(), "a pending reply counts as fleet state to collect");
|
||||
}
|
||||
|
||||
@Test
|
||||
void snapshotWithEmptyRosterHasNoPendingState() {
|
||||
InMemoryReplyInbox inbox = new InMemoryReplyInbox();
|
||||
LeadHeartbeatLoop.FleetState fs = LeadHeartbeatLoop.snapshot(inbox, List::of);
|
||||
assertFalse(fs.hasPending());
|
||||
assertEquals(0, fs.liveWorkers());
|
||||
}
|
||||
}
|
||||
@@ -120,6 +120,36 @@ class MessageServiceTest {
|
||||
assertFalse(reply.completed());
|
||||
}
|
||||
|
||||
@Test
|
||||
void droppedQueuedAndDeliveredTurnsExposeTheRealCauseExactlyOnce() throws Exception {
|
||||
CompletableFuture<MessageService.Reply> first = sendAsync();
|
||||
awaitWaiting();
|
||||
injector.onStatus(T, AgentStatus.IDLE); // first delivery
|
||||
injector.onStatus(T, AgentStatus.WORKING); // first turn in flight
|
||||
|
||||
CompletableFuture<Void> queued = injector.enqueue(T, "second task");
|
||||
CompletableFuture<Rendezvous.Resolution> waiter = rendezvous.currentWaiter(T);
|
||||
injector.drop(T, new HerdrException("agent target sol not found", "agent_not_found", null));
|
||||
|
||||
MessageService.Reply reply = first.get(5, TimeUnit.SECONDS);
|
||||
assertEquals(MessageService.Outcome.WORKER_FAILED, reply.outcome());
|
||||
assertEquals("agent target sol not found", reply.text());
|
||||
assertTrue(queued.isCompletedExceptionally(), "the queued delivery future also fails");
|
||||
assertFalse(rendezvous.resolveFailure(waiter, "second failure"), "the waiter fails exactly once");
|
||||
}
|
||||
|
||||
@Test
|
||||
void noDropReasonKeepsTheExistingFallbackText() throws Exception {
|
||||
herdr.readText("");
|
||||
CompletableFuture<Rendezvous.Resolution> waiter = rendezvous.open(T);
|
||||
|
||||
completion.onTurnFailed(T);
|
||||
|
||||
Rendezvous.Resolution resolution = waiter.get(2, TimeUnit.SECONDS);
|
||||
assertEquals("worker did not reply; its turn ended in an unrecoverable state "
|
||||
+ "(worker unreachable or stuck)", resolution.text());
|
||||
}
|
||||
|
||||
// --- bridge_ask reverse rendezvous (CB-205) ------------------------------------------------
|
||||
|
||||
@Test
|
||||
@@ -609,4 +639,94 @@ class MessageServiceTest {
|
||||
assertTrue(view.detail() != null && view.detail().contains("released"),
|
||||
"and the detail says why, rather than 'worker unknown'");
|
||||
}
|
||||
|
||||
@Test
|
||||
void abandonFailsEveryPendingAsyncTicketForTheReleasedTarget() throws Exception {
|
||||
String first = messages.sendAsync(T, "first task");
|
||||
awaitWaiting(); // first task owns the target lock and rendezvous waiter
|
||||
String second = messages.sendAsync(T, "second task"); // parked on the same lock, not yet queued
|
||||
String third = messages.sendAsync(T, "third task"); // a second queued ticket proves the full sweep
|
||||
|
||||
assertTrue(messages.abandon(T, "agent target term_a not found"));
|
||||
|
||||
assertFailedTicket(first, "agent target term_a not found");
|
||||
assertFailedTicket(second, "agent target term_a not found");
|
||||
assertFailedTicket(third, "agent target term_a not found");
|
||||
}
|
||||
|
||||
@Test
|
||||
void abandonDoesNotFailAnAsyncTicketWaitingForAnAnswer() throws Exception {
|
||||
String ticket = messages.sendAsync(T, "task that asks");
|
||||
awaitWaiting();
|
||||
injectDelivery();
|
||||
|
||||
CompletableFuture<MessageService.AskResult> ask =
|
||||
CompletableFuture.supplyAsync(() -> messages.ask(T, "which config?", 5000));
|
||||
MessageService.TaskView asking = awaitTicketPhase(ticket, MessageService.Phase.ASKING);
|
||||
|
||||
assertFalse(messages.abandon(T, "agent target term_a not found"),
|
||||
"an asking ticket is an active turn, not a pending send to sweep");
|
||||
assertEquals(MessageService.Phase.ASKING, messages.poll(ticket).phase());
|
||||
|
||||
CompletableFuture<MessageService.Reply> answer = CompletableFuture.supplyAsync(
|
||||
() -> messages.answer(asking.turnId(), "config.yaml", 5000));
|
||||
assertEquals("config.yaml", ask.get(5, TimeUnit.SECONDS).answer());
|
||||
awaitWaiting();
|
||||
assertTrue(rendezvous.resolve(T, "done"));
|
||||
assertEquals(MessageService.Outcome.REPLIED, answer.get(5, TimeUnit.SECONDS).outcome());
|
||||
}
|
||||
|
||||
@Test
|
||||
void unansweredAsyncQuestionReturnsTheTicketToPendingAndReleasesItsTarget() throws Exception {
|
||||
String ticket = messages.sendAsync(T, "task that asks");
|
||||
awaitWaiting();
|
||||
injectDelivery();
|
||||
|
||||
assertEquals(MessageService.AskOutcome.TIMED_OUT,
|
||||
messages.ask(T, "which config?", 200).outcome());
|
||||
assertEquals(MessageService.Phase.PENDING, messages.poll(ticket).phase(),
|
||||
"only the question wait ended; the delegated turn may still finish");
|
||||
|
||||
String next = messages.sendAsync(T, "next task");
|
||||
awaitWaiting();
|
||||
assertTrue(rendezvous.resolve(T, "done"));
|
||||
assertEquals(MessageService.Phase.DONE, awaitTicketPhase(next, MessageService.Phase.DONE).phase());
|
||||
}
|
||||
|
||||
@Test
|
||||
void asyncQuestionBelongsToTheTaskThatOwnsItsForwardWaiter() throws Exception {
|
||||
String first = messages.sendAsync(T, "first task");
|
||||
awaitWaiting();
|
||||
|
||||
CompletableFuture<MessageService.AskResult> ask =
|
||||
CompletableFuture.supplyAsync(() -> messages.ask(T, "which config?", 5000));
|
||||
assertEquals(MessageService.Phase.ASKING, awaitTicketPhase(first, MessageService.Phase.ASKING).phase());
|
||||
|
||||
MessageService.TaskView asking = messages.poll(first);
|
||||
CompletableFuture<MessageService.Reply> answer = CompletableFuture.supplyAsync(
|
||||
() -> messages.answer(asking.turnId(), "config.yaml", 5000));
|
||||
assertEquals("config.yaml", ask.get(5, TimeUnit.SECONDS).answer());
|
||||
awaitWaiting();
|
||||
assertTrue(rendezvous.resolve(T, "done"));
|
||||
assertEquals(MessageService.Outcome.REPLIED, answer.get(5, TimeUnit.SECONDS).outcome());
|
||||
}
|
||||
|
||||
private void assertFailedTicket(String ticket, String reason) throws Exception {
|
||||
MessageService.TaskView view = awaitTicketPhase(ticket, MessageService.Phase.FAILED);
|
||||
assertEquals(reason, view.detail());
|
||||
}
|
||||
|
||||
private MessageService.TaskView awaitTicketPhase(String ticket, MessageService.Phase phase) throws Exception {
|
||||
long deadline = System.currentTimeMillis() + 3000;
|
||||
MessageService.TaskView view;
|
||||
do {
|
||||
view = messages.poll(ticket);
|
||||
if (view.phase() == phase) {
|
||||
return view;
|
||||
}
|
||||
Thread.sleep(5);
|
||||
} while (System.currentTimeMillis() < deadline);
|
||||
assertEquals(phase, view.phase());
|
||||
return view;
|
||||
}
|
||||
}
|
||||
|
||||
@@ -179,6 +179,20 @@ class ReplyPushLoopTest {
|
||||
|
||||
// --- nudge format --------------------------------------------------------------------------
|
||||
|
||||
@Test
|
||||
void isActiveReflectsALiveScheduleForTheHeartbeatStandDown() {
|
||||
var rec = recordingClient();
|
||||
agents = new AgentControl(rec);
|
||||
inbox.publish(WORKER, "m1", "hello");
|
||||
ReplyPushLoop loop = loop(1, 100_000); // long backoff so the tick cannot fire mid-test
|
||||
|
||||
loop.onReplyQueued(WORKER);
|
||||
assertTrue(loop.isActive(), "CB-551: the heartbeat must stand aside while a reminder is live");
|
||||
|
||||
loop.stop();
|
||||
assertFalse(loop.isActive(), "stopping clears the active schedule");
|
||||
}
|
||||
|
||||
@Test
|
||||
void nudgeFormatIsCorrect() {
|
||||
String nudge = ReplyPushLoop.NUDGE_FORMAT.formatted(WORKER, WORKER);
|
||||
|
||||
@@ -0,0 +1,67 @@
|
||||
package dev.ltms.bridged.peer;
|
||||
|
||||
import org.junit.jupiter.api.Test;
|
||||
|
||||
import static org.junit.jupiter.api.Assertions.assertEquals;
|
||||
import static org.junit.jupiter.api.Assertions.assertSame;
|
||||
import static org.junit.jupiter.api.Assertions.assertThrows;
|
||||
import static org.junit.jupiter.api.Assertions.assertTrue;
|
||||
|
||||
class MemberRoleTest {
|
||||
|
||||
@Test
|
||||
void theThreeRolesAreArchitectDevAndReviewer() {
|
||||
assertEquals(3, MemberRole.values().length,
|
||||
"a new role changes the charter, the role file, the skill and the authz row — "
|
||||
+ "adding one is a deliberate act, so this count is meant to fail first");
|
||||
assertEquals("architect", MemberRole.ARCHITECT.wireName());
|
||||
assertEquals("dev", MemberRole.DEV.wireName());
|
||||
assertEquals("reviewer", MemberRole.REVIEWER.wireName());
|
||||
}
|
||||
|
||||
@Test
|
||||
void parseAcceptsTheWireSpelling() {
|
||||
for (MemberRole r : MemberRole.values()) {
|
||||
assertSame(r, MemberRole.parse(r.wireName()));
|
||||
}
|
||||
}
|
||||
|
||||
@Test
|
||||
void parseIsCaseInsensitiveAndTrimsSurroundingSpace() {
|
||||
assertSame(MemberRole.ARCHITECT, MemberRole.parse("Architect"));
|
||||
assertSame(MemberRole.DEV, MemberRole.parse(" DEV "));
|
||||
assertSame(MemberRole.REVIEWER, MemberRole.parse("ReViEwEr"));
|
||||
}
|
||||
|
||||
@Test
|
||||
void parseRejectsAnUnknownRoleAndListsTheValidOnes() {
|
||||
IllegalArgumentException e =
|
||||
assertThrows(IllegalArgumentException.class, () -> MemberRole.parse("archtiect"));
|
||||
assertTrue(e.getMessage().contains("archtiect"), e.getMessage());
|
||||
assertTrue(e.getMessage().contains("architect, dev, reviewer"),
|
||||
"a typo in config should be fixable from the message alone: " + e.getMessage());
|
||||
}
|
||||
|
||||
@Test
|
||||
void parseRejectsNullAndBlank() {
|
||||
assertThrows(IllegalArgumentException.class, () -> MemberRole.parse(null));
|
||||
assertThrows(IllegalArgumentException.class, () -> MemberRole.parse(""));
|
||||
assertThrows(IllegalArgumentException.class, () -> MemberRole.parse(" "));
|
||||
}
|
||||
|
||||
/**
|
||||
* A lead is not a member role. A lead is not spawned — it is a pre-existing session that config
|
||||
* recognises — so it has no launch charter to pick and never appears in a {@code members:} slot.
|
||||
*/
|
||||
@Test
|
||||
void leadIsNotAMemberRole() {
|
||||
assertThrows(IllegalArgumentException.class, () -> MemberRole.parse("lead"));
|
||||
assertThrows(IllegalArgumentException.class, () -> MemberRole.parse("primary"));
|
||||
}
|
||||
|
||||
/** "worker" is the name this taxonomy replaced; it must not quietly keep working. */
|
||||
@Test
|
||||
void workerIsNoLongerARole() {
|
||||
assertThrows(IllegalArgumentException.class, () -> MemberRole.parse("worker"));
|
||||
}
|
||||
}
|
||||
@@ -1,6 +1,7 @@
|
||||
package dev.ltms.bridged.rest;
|
||||
|
||||
import dev.ltms.bridged.auth.CallerResolver;
|
||||
import dev.ltms.bridged.auth.MemberRegistry;
|
||||
import dev.ltms.bridged.config.BridgedConfig;
|
||||
import dev.ltms.bridged.guard.SubscriptionGuard;
|
||||
import dev.ltms.bridged.herdr.AgentControl;
|
||||
@@ -15,7 +16,7 @@ import dev.ltms.bridged.msg.MessageService;
|
||||
import dev.ltms.bridged.msg.Rendezvous;
|
||||
import dev.ltms.bridged.session.FakeWorktrees;
|
||||
import dev.ltms.bridged.session.SessionManager;
|
||||
import dev.ltms.bridged.worker.ClaudeCodeLauncher;
|
||||
import dev.ltms.bridged.member.ClaudeCodeLauncher;
|
||||
import io.javalin.Javalin;
|
||||
import org.junit.jupiter.api.AfterEach;
|
||||
import org.junit.jupiter.api.Test;
|
||||
@@ -52,7 +53,7 @@ class BridgedAppAuthTest {
|
||||
*/
|
||||
private int start(long pid, boolean tokenMode, String token) {
|
||||
FakeHerdr herdr = new FakeHerdr();
|
||||
BridgedConfig.Worker wcfg = new BridgedConfig.Worker(
|
||||
BridgedConfig.Profile wcfg = new BridgedConfig.Profile(
|
||||
"ltms-local", "http://gx00.gw:8000", "coder", null, "BRIDGED_WORKER_TOKEN", null,
|
||||
"tab", "bridged-workers", "worker: {profile} #{n}", null, null, null);
|
||||
AgentControl agents = new AgentControl(herdr);
|
||||
@@ -65,9 +66,8 @@ class BridgedAppAuthTest {
|
||||
MessageService messages = new MessageService(agents, injector, new Rendezvous());
|
||||
|
||||
ConnectionIdentity identity = new ConnectionIdentity(new PaneLocator(herdr), _ -> pid);
|
||||
CallerResolver callers = tokenMode
|
||||
? new CallerResolver(identity, true, token)
|
||||
: new CallerResolver(identity);
|
||||
CallerResolver callers = CallerResolver.withLeadsAndMembers(identity, tokenMode, token,
|
||||
Map::of, new MemberRegistry(null));
|
||||
metrics = BridgedMetrics.create(sessions, new dev.ltms.bridged.msg.InMemoryReplyInbox());
|
||||
|
||||
app = new BridgedApp(herdr, workers, sessions, messages, sessions.asPresence(), null,
|
||||
@@ -128,9 +128,9 @@ class BridgedAppAuthTest {
|
||||
void aWorkerMayNotOrchestrate() throws Exception {
|
||||
int port = start(FakeHerdr.WORKER_PID, false, null);
|
||||
|
||||
assertEquals(403, send(port, "POST", "/workers", null, null).statusCode(),
|
||||
assertEquals(403, send(port, "POST", "/members", null, null).statusCode(),
|
||||
"a worker spawning workers would be escalating into the orchestrator role");
|
||||
assertEquals(403, send(port, "DELETE", "/workers/w2:p7", null, null).statusCode());
|
||||
assertEquals(403, send(port, "DELETE", "/members/w2:p7", null, null).statusCode());
|
||||
assertEquals(403, send(port, "POST", "/sessions/term_b/message",
|
||||
"{\"content\":\"hi\"}", null).statusCode());
|
||||
assertEquals(403, send(port, "GET", "/sessions/term_a/replies", null, null).statusCode(),
|
||||
@@ -206,7 +206,7 @@ class BridgedAppAuthTest {
|
||||
// The 29 pre-existing acceptance tests rely on this: no auth fixture, no enforcement.
|
||||
FakeHerdr herdr = new FakeHerdr();
|
||||
AgentControl agents = new AgentControl(herdr);
|
||||
BridgedConfig.Worker wcfg = new BridgedConfig.Worker(
|
||||
BridgedConfig.Profile wcfg = new BridgedConfig.Profile(
|
||||
"ltms-local", "http://gx00.gw:8000", "coder", null, "BRIDGED_WORKER_TOKEN", null,
|
||||
"tab", "bridged-workers", "worker: {profile} #{n}", null, null, null);
|
||||
ClaudeCodeLauncher workers = new ClaudeCodeLauncher(
|
||||
|
||||
@@ -9,7 +9,7 @@ import dev.ltms.bridged.herdr.FakeHerdr;
|
||||
import dev.ltms.bridged.herdr.WorkspaceControl;
|
||||
import dev.ltms.bridged.inject.Injector;
|
||||
import dev.ltms.bridged.inject.StatusPoller;
|
||||
import dev.ltms.bridged.inject.WorkerPresence;
|
||||
import dev.ltms.bridged.inject.MemberPresence;
|
||||
import dev.ltms.bridged.msg.InMemoryReplyInbox;
|
||||
import dev.ltms.bridged.msg.MessageService;
|
||||
import dev.ltms.bridged.msg.Rendezvous;
|
||||
@@ -17,7 +17,7 @@ import dev.ltms.bridged.session.FakeWorktrees;
|
||||
import dev.ltms.bridged.session.GitWorktrees;
|
||||
import dev.ltms.bridged.session.SessionManager;
|
||||
import dev.ltms.bridged.session.Worktrees;
|
||||
import dev.ltms.bridged.worker.ClaudeCodeLauncher;
|
||||
import dev.ltms.bridged.member.ClaudeCodeLauncher;
|
||||
import io.javalin.Javalin;
|
||||
import org.junit.jupiter.api.AfterEach;
|
||||
import org.junit.jupiter.api.Test;
|
||||
@@ -43,7 +43,7 @@ class BridgedAppTest {
|
||||
|
||||
private final ObjectMapper mapper = new ObjectMapper();
|
||||
private final HttpClient http = HttpClient.newHttpClient();
|
||||
private WorkerPresence presence;
|
||||
private MemberPresence presence;
|
||||
private Javalin app;
|
||||
private StatusPoller poller;
|
||||
|
||||
@@ -62,7 +62,7 @@ class BridgedAppTest {
|
||||
}
|
||||
|
||||
private int start(FakeHerdr herdr, String workerBaseUrl, Set<String> allow, String placement, Worktrees worktrees) {
|
||||
BridgedConfig.Worker wcfg = new BridgedConfig.Worker(
|
||||
BridgedConfig.Profile wcfg = new BridgedConfig.Profile(
|
||||
"ltms-local", workerBaseUrl, "coder", null, "BRIDGED_WORKER_TOKEN", null,
|
||||
placement, "bridged-workers", "worker: {profile} #{n}", null, null, null);
|
||||
AgentControl agents = new AgentControl(herdr);
|
||||
@@ -160,7 +160,7 @@ class BridgedAppTest {
|
||||
FakeHerdr herdr = new FakeHerdr();
|
||||
int port = start(herdr, "http://gx00.gw:8000", Set.of("gx00.gw"));
|
||||
|
||||
HttpResponse<String> res = req(port, "POST", "/workers");
|
||||
HttpResponse<String> res = req(port, "POST", "/members");
|
||||
assertEquals(201, res.statusCode());
|
||||
JsonNode body = mapper.readTree(res.body());
|
||||
// CB-519: the responded paneId is a host-unique opaque UUID, not the herdr pane coordinate.
|
||||
@@ -202,12 +202,12 @@ class BridgedAppTest {
|
||||
int port = start(herdr, "http://gx00.gw:8000", Set.of("gx00.gw"), "tab",
|
||||
new FakeWorktrees().withRepoRoot("/repo").withPrefix("/wt"));
|
||||
|
||||
HttpResponse<String> spawn = req(port, "POST", "/workers?worktree=true&ticket=cb-304");
|
||||
HttpResponse<String> spawn = req(port, "POST", "/members?worktree=true&ticket=cb-304");
|
||||
assertEquals(201, spawn.statusCode());
|
||||
JsonNode spawned = mapper.readTree(spawn.body());
|
||||
String paneId = spawned.get("paneId").asText();
|
||||
|
||||
HttpResponse<String> res = req(port, "GET", "/workers");
|
||||
HttpResponse<String> res = req(port, "GET", "/members");
|
||||
assertEquals(200, res.statusCode());
|
||||
JsonNode workers = mapper.readTree(res.body()).get("workers");
|
||||
assertEquals(1, workers.size());
|
||||
@@ -226,7 +226,7 @@ class BridgedAppTest {
|
||||
void spawnWithACwdParamRootsTheWorkerThere() throws Exception {
|
||||
FakeHerdr herdr = new FakeHerdr();
|
||||
int port = start(herdr, "http://gx00.gw:8000", Set.of("gx00.gw"));
|
||||
assertEquals(201, req(port, "POST", "/workers?cwd=/tmp/proj").statusCode());
|
||||
assertEquals(201, req(port, "POST", "/members?cwd=/tmp/proj").statusCode());
|
||||
assertEquals("/tmp/proj", params(herdr, "tab.create").get("cwd"), "the worker starts in cwd");
|
||||
}
|
||||
|
||||
@@ -234,7 +234,7 @@ class BridgedAppTest {
|
||||
void spawnWithAnUnknownProfileIs400() throws Exception {
|
||||
FakeHerdr herdr = new FakeHerdr();
|
||||
int port = start(herdr, "http://gx00.gw:8000", Set.of("gx00.gw"));
|
||||
HttpResponse<String> res = req(port, "POST", "/workers?profile=nope");
|
||||
HttpResponse<String> res = req(port, "POST", "/members?profile=nope");
|
||||
assertEquals(400, res.statusCode());
|
||||
assertEquals("unknown_profile", mapper.readTree(res.body()).get("error").asText());
|
||||
assertFalse(herdr.called("agent.start"), "an unknown profile must not spawn anything");
|
||||
@@ -246,7 +246,7 @@ class BridgedAppTest {
|
||||
FakeHerdr herdr = new FakeHerdr().withWorkspace("w9", "bridged-workers");
|
||||
int port = start(herdr, "http://gx00.gw:8000", Set.of("gx00.gw"));
|
||||
|
||||
assertEquals(201, req(port, "POST", "/workers").statusCode());
|
||||
assertEquals(201, req(port, "POST", "/members").statusCode());
|
||||
assertFalse(herdr.called("workspace.create"), "existing worker space must be reused, not recreated");
|
||||
assertTrue(herdr.called("tab.create"), "a fresh tab is still created for the worker");
|
||||
}
|
||||
@@ -258,7 +258,7 @@ class BridgedAppTest {
|
||||
FakeHerdr herdr = new FakeHerdr().agentNameTakenTimes(2);
|
||||
int port = start(herdr, "http://gx00.gw:8000", Set.of("gx00.gw"));
|
||||
|
||||
assertEquals(201, req(port, "POST", "/workers").statusCode());
|
||||
assertEquals(201, req(port, "POST", "/members").statusCode());
|
||||
List<String> names = herdr.calls.stream()
|
||||
.filter(c -> c.method().equals("agent.start"))
|
||||
.map(c -> ((Map<String, Object>) c.params()).get("name").toString())
|
||||
@@ -273,7 +273,7 @@ class BridgedAppTest {
|
||||
FakeHerdr herdr = new FakeHerdr().agentNameTakenTimes(99);
|
||||
int port = start(herdr, "http://gx00.gw:8000", Set.of("gx00.gw"));
|
||||
|
||||
assertEquals(500, req(port, "POST", "/workers").statusCode());
|
||||
assertEquals(500, req(port, "POST", "/members").statusCode());
|
||||
assertTrue(herdr.called("tab.create"), "a tab was created before the failed start");
|
||||
assertEquals("w9:t2", params(herdr, "tab.close").get("tab_id"), "orphaned tab must be closed");
|
||||
}
|
||||
@@ -284,7 +284,7 @@ class BridgedAppTest {
|
||||
// base_url points at the subscription — guard must block before any herdr call.
|
||||
int port = start(herdr, "https://api.anthropic.com", Set.of("gx00.gw"));
|
||||
|
||||
HttpResponse<String> res = req(port, "POST", "/workers");
|
||||
HttpResponse<String> res = req(port, "POST", "/members");
|
||||
assertEquals(403, res.statusCode());
|
||||
assertEquals("subscription_boundary", mapper.readTree(res.body()).get("error").asText());
|
||||
assertFalse(herdr.called("agent.start"), "guard must stop the spawn before herdr");
|
||||
@@ -296,7 +296,7 @@ class BridgedAppTest {
|
||||
FakeHerdr herdr = new FakeHerdr();
|
||||
int port = start(herdr, "http://gx00.gw:8000", Set.of("gx00.gw"), "pane");
|
||||
|
||||
assertEquals(201, req(port, "POST", "/workers").statusCode());
|
||||
assertEquals(201, req(port, "POST", "/members").statusCode());
|
||||
Map<String, Object> start = params(herdr, "agent.start");
|
||||
assertFalse(start.containsKey("tab_id"), "pane placement must not target a tab");
|
||||
assertFalse(herdr.called("workspace.create"), "pane placement uses no dedicated space");
|
||||
@@ -308,7 +308,7 @@ class BridgedAppTest {
|
||||
FakeHerdr herdr = new FakeHerdr();
|
||||
int port = start(herdr, "http://gx00.gw:8000", Set.of("gx00.gw"));
|
||||
|
||||
assertEquals(204, req(port, "DELETE", "/workers/w9:pW").statusCode());
|
||||
assertEquals(204, req(port, "DELETE", "/members/w9:pW").statusCode());
|
||||
assertTrue(herdr.called("pane.close"));
|
||||
// Tab resolved from the pane (pane.get), then closed.
|
||||
assertEquals("w9:t2", params(herdr, "tab.close").get("tab_id"));
|
||||
@@ -450,7 +450,7 @@ class BridgedAppTest {
|
||||
FakeHerdr herdr = new FakeHerdr();
|
||||
int port = start(herdr, "http://gx00.gw:8000", Set.of("gx00.gw"), "pane");
|
||||
|
||||
assertEquals(204, req(port, "DELETE", "/workers/w9:pW").statusCode());
|
||||
assertEquals(204, req(port, "DELETE", "/members/w9:pW").statusCode());
|
||||
assertTrue(herdr.called("pane.close"));
|
||||
assertFalse(herdr.called("tab.close"), "pane placement owns no tab to close");
|
||||
assertFalse(herdr.called("pane.get"), "no tab resolution in pane placement");
|
||||
@@ -462,7 +462,7 @@ class BridgedAppTest {
|
||||
FakeHerdr herdr = new FakeHerdr().withWorkerTabPaneCount(2);
|
||||
int port = start(herdr, "http://gx00.gw:8000", Set.of("gx00.gw"));
|
||||
|
||||
assertEquals(204, req(port, "DELETE", "/workers/w9:pW").statusCode());
|
||||
assertEquals(204, req(port, "DELETE", "/members/w9:pW").statusCode());
|
||||
assertTrue(herdr.called("pane.close"), "the worker's own pane is still closed");
|
||||
assertFalse(herdr.called("tab.close"), "must not close a tab that holds the user's other panes");
|
||||
}
|
||||
@@ -473,7 +473,7 @@ class BridgedAppTest {
|
||||
int port = start(herdr, "http://gx00.gw:8000", Set.of("gx00.gw"));
|
||||
|
||||
// A genuine teardown failure must surface, not be reported as a successful 204.
|
||||
assertEquals(500, req(port, "DELETE", "/workers/w9:pW").statusCode());
|
||||
assertEquals(500, req(port, "DELETE", "/members/w9:pW").statusCode());
|
||||
assertFalse(herdr.called("tab.close"), "tab is not removed when the pane close failed");
|
||||
}
|
||||
|
||||
@@ -483,7 +483,7 @@ class BridgedAppTest {
|
||||
int port = start(herdr, "http://gx00.gw:8000", Set.of("gx00.gw"));
|
||||
|
||||
// Already-gone is success; the (now-empty) tab is still cleaned up.
|
||||
assertEquals(204, req(port, "DELETE", "/workers/w9:pW").statusCode());
|
||||
assertEquals(204, req(port, "DELETE", "/members/w9:pW").statusCode());
|
||||
assertTrue(herdr.called("tab.close"));
|
||||
}
|
||||
}
|
||||
|
||||
@@ -1,13 +1,18 @@
|
||||
package dev.ltms.bridged.session;
|
||||
|
||||
import ch.qos.logback.classic.Level;
|
||||
import ch.qos.logback.classic.LoggerContext;
|
||||
import ch.qos.logback.classic.spi.ILoggingEvent;
|
||||
import ch.qos.logback.core.read.ListAppender;
|
||||
import dev.ltms.bridged.config.BridgedConfig;
|
||||
import dev.ltms.bridged.guard.SubscriptionGuard;
|
||||
import dev.ltms.bridged.herdr.AgentControl;
|
||||
import dev.ltms.bridged.herdr.FakeHerdr;
|
||||
import dev.ltms.bridged.herdr.WorkspaceControl;
|
||||
import dev.ltms.bridged.worker.ClaudeCodeLauncher;
|
||||
import dev.ltms.bridged.member.ClaudeCodeLauncher;
|
||||
import dev.ltms.bridged.peer.PeerUnreachableException;
|
||||
import org.junit.jupiter.api.Test;
|
||||
import org.slf4j.LoggerFactory;
|
||||
|
||||
import java.util.List;
|
||||
import java.util.Map;
|
||||
@@ -25,7 +30,7 @@ import static org.junit.jupiter.api.Assertions.*;
|
||||
class SessionManagerTest {
|
||||
|
||||
private SessionManager sessionManager(FakeHerdr herdr) {
|
||||
BridgedConfig.Worker cfg = new BridgedConfig.Worker(
|
||||
BridgedConfig.Profile cfg = new BridgedConfig.Profile(
|
||||
"ltms-local", "http://gx00.gw:8000", "coder", null, "BRIDGED_WORKER_TOKEN",
|
||||
List.of("ccs", "ltms-local"), "tab", "bridged-workers",
|
||||
"worker: {profile} #{n}", null, null, null);
|
||||
@@ -44,7 +49,7 @@ class SessionManagerTest {
|
||||
|
||||
private SessionManager sessionManager(FakeHerdr herdr, LongSupplier clock, int contextCap,
|
||||
boolean clearAfterTurn) {
|
||||
BridgedConfig.Worker cfg = new BridgedConfig.Worker(
|
||||
BridgedConfig.Profile cfg = new BridgedConfig.Profile(
|
||||
"ltms-local", "http://gx00.gw:8000", "coder", null, "BRIDGED_WORKER_TOKEN",
|
||||
List.of("ccs", "ltms-local"), "tab", "bridged-workers",
|
||||
"worker: {profile} #{n}", null, null, null);
|
||||
@@ -69,10 +74,10 @@ class SessionManagerTest {
|
||||
FakeHerdr herdr = new FakeHerdr();
|
||||
SessionManager sessions = sessionManager(herdr);
|
||||
|
||||
WorkerSession a = sessions.acquire("ltms-local", "/work/a", "/caller/a", "term_primary");
|
||||
WorkerSession b = sessions.acquire("ltms-local", "/work/b", "/caller/b", "term_primary");
|
||||
MemberSession a = sessions.acquire("ltms-local", "/work/a", "/caller/a", "term_primary");
|
||||
MemberSession b = sessions.acquire("ltms-local", "/work/b", "/caller/b", "term_primary");
|
||||
|
||||
assertEquals(WorkerSession.State.SPAWNING, a.state(), "fresh session starts spawning");
|
||||
assertEquals(MemberSession.State.SPAWNING, a.state(), "fresh session starts spawning");
|
||||
assertEquals("ltms-local", a.profile());
|
||||
assertEquals("/work/a", a.cwd(), "explicit requested cwd is recorded");
|
||||
assertEquals("term_primary", a.ownerTerminal());
|
||||
@@ -92,7 +97,7 @@ class SessionManagerTest {
|
||||
// once a session existed, so this NPE'd the primary's second spawn while the first passed.
|
||||
FakeHerdr herdr = new FakeHerdr();
|
||||
SessionManager sessions = sessionManager(herdr);
|
||||
WorkerSession session = sessions.acquire("ltms-local", null, "/caller", "term_primary");
|
||||
MemberSession session = sessions.acquire("ltms-local", null, "/caller", "term_primary");
|
||||
|
||||
assertDoesNotThrow(() -> sessions.asPresence().markPresent(null),
|
||||
"the primary's null terminal must not blow up an unrelated tool call");
|
||||
@@ -100,7 +105,7 @@ class SessionManagerTest {
|
||||
assertDoesNotThrow(() -> sessions.onTurnComplete(null));
|
||||
assertDoesNotThrow(() -> sessions.onTurnFailed(null));
|
||||
|
||||
assertEquals(WorkerSession.State.SPAWNING, sessions.get(session.paneId()).orElseThrow().state(),
|
||||
assertEquals(MemberSession.State.SPAWNING, sessions.get(session.paneId()).orElseThrow().state(),
|
||||
"and must not transition any registered session");
|
||||
}
|
||||
|
||||
@@ -108,20 +113,20 @@ class SessionManagerTest {
|
||||
void presenceMovesSpawningToReadyAndDeliveredTurnMovesToDone() {
|
||||
FakeHerdr herdr = new FakeHerdr();
|
||||
SessionManager sessions = sessionManager(herdr);
|
||||
WorkerSession session = sessions.acquire("ltms-local", null, "/caller", "term_primary");
|
||||
MemberSession session = sessions.acquire("ltms-local", null, "/caller", "term_primary");
|
||||
String terminal = session.terminalId();
|
||||
|
||||
sessions.asPresence().markPresent(terminal);
|
||||
assertEquals(WorkerSession.State.READY, sessions.get(session.paneId()).orElseThrow().state(),
|
||||
assertEquals(MemberSession.State.READY, sessions.get(session.paneId()).orElseThrow().state(),
|
||||
"MCP presence moves SPAWNING → READY");
|
||||
assertTrue(sessions.asPresence().isPresent(terminal), "presence is also recorded");
|
||||
|
||||
sessions.onDelivered(terminal);
|
||||
assertEquals(WorkerSession.State.BUSY, sessions.get(session.paneId()).orElseThrow().state(),
|
||||
assertEquals(MemberSession.State.BUSY, sessions.get(session.paneId()).orElseThrow().state(),
|
||||
"delivery moves READY → BUSY");
|
||||
|
||||
sessions.onTurnComplete(terminal);
|
||||
assertEquals(WorkerSession.State.DONE, sessions.get(session.paneId()).orElseThrow().state(),
|
||||
assertEquals(MemberSession.State.DONE, sessions.get(session.paneId()).orElseThrow().state(),
|
||||
"turn completion moves BUSY → DONE");
|
||||
}
|
||||
|
||||
@@ -129,7 +134,7 @@ class SessionManagerTest {
|
||||
void releaseTearsDownWorkerAndRemovesFromRosterAndIsIdempotent() {
|
||||
FakeHerdr herdr = new FakeHerdr();
|
||||
SessionManager sessions = sessionManager(herdr);
|
||||
WorkerSession session = sessions.acquire("ltms-local", null, "/caller", null);
|
||||
MemberSession session = sessions.acquire("ltms-local", null, "/caller", null);
|
||||
String paneId = session.paneId();
|
||||
|
||||
sessions.release(paneId);
|
||||
@@ -145,53 +150,59 @@ class SessionManagerTest {
|
||||
void onTurnFailedMovesSessionToFailed() {
|
||||
FakeHerdr herdr = new FakeHerdr();
|
||||
SessionManager sessions = sessionManager(herdr);
|
||||
WorkerSession session = sessions.acquire("ltms-local", null, "/caller", "term_primary");
|
||||
MemberSession session = sessions.acquire("ltms-local", null, "/caller", "term_primary");
|
||||
String terminal = session.terminalId();
|
||||
sessions.asPresence().markPresent(terminal);
|
||||
sessions.onDelivered(terminal);
|
||||
|
||||
sessions.onTurnFailed(terminal);
|
||||
|
||||
WorkerSession updated = sessions.get(session.paneId()).orElseThrow();
|
||||
assertEquals(WorkerSession.State.FAILED, updated.state(), "turn failure moves to FAILED");
|
||||
MemberSession updated = sessions.get(session.paneId()).orElseThrow();
|
||||
assertEquals(MemberSession.State.FAILED, updated.state(), "turn failure moves to FAILED");
|
||||
assertTrue(sessions.roster().contains(updated), "FAILED is still in acquired-minus-released roster");
|
||||
}
|
||||
|
||||
@Test
|
||||
void recycleProducesNewPaneIdAndOldOneIsGone() {
|
||||
FakeHerdr herdr = new FakeHerdr();
|
||||
SessionManager sessions = sessionManager(herdr);
|
||||
WorkerSession oldSession = sessions.acquire("ltms-local", null, "/caller", "term_primary");
|
||||
String oldPane = oldSession.paneId();
|
||||
String oldTerminal = oldSession.terminalId();
|
||||
void onTurnFailedIsLoggedAtWarnWithThePriorState() {
|
||||
// CB-564: this transition used to be a bare DEBUG "session marked failed" — a symptom with no
|
||||
// cause. A member that can no longer be delegated to must be at least WARN, and should name
|
||||
// what stage it failed at (here: BUSY, i.e. a turn was in flight and never resolved).
|
||||
LoggerContext ctx = (LoggerContext) LoggerFactory.getILoggerFactory();
|
||||
ch.qos.logback.classic.Logger sessionLog =
|
||||
(ch.qos.logback.classic.Logger) LoggerFactory.getLogger(SessionManager.class);
|
||||
ListAppender<ILoggingEvent> appender = new ListAppender<>();
|
||||
appender.setContext(ctx);
|
||||
appender.start();
|
||||
sessionLog.addAppender(appender);
|
||||
sessionLog.setLevel(Level.WARN);
|
||||
try {
|
||||
FakeHerdr herdr = new FakeHerdr();
|
||||
SessionManager sessions = sessionManager(herdr);
|
||||
MemberSession session = sessions.acquire("ltms-local", null, "/caller", "term_primary");
|
||||
String terminal = session.terminalId();
|
||||
sessions.asPresence().markPresent(terminal);
|
||||
sessions.onDelivered(terminal);
|
||||
|
||||
WorkerSession fresh = sessions.recycle(oldPane);
|
||||
sessions.onTurnFailed(terminal);
|
||||
|
||||
assertNotEquals(oldPane, fresh.paneId(), "recycle yields a new pane id");
|
||||
assertNotEquals(oldTerminal, fresh.terminalId(), "recycle yields a new terminal id");
|
||||
assertEquals(oldSession.profile(), fresh.profile(), "profile is preserved");
|
||||
assertEquals(oldSession.cwd(), fresh.cwd(), "cwd is preserved");
|
||||
assertEquals(oldSession.ownerTerminal(), fresh.ownerTerminal(), "owner is preserved");
|
||||
|
||||
assertTrue(sessions.get(oldPane).isEmpty(), "old pane is deregistered");
|
||||
assertEquals(1, sessions.roster().size(), "only the fresh session remains");
|
||||
assertEquals(fresh.paneId(), sessions.roster().getFirst().paneId());
|
||||
|
||||
// The old session was the first spawn → pane w9:pRoot_1 (CB-519: the registry key is the
|
||||
// uuid id, so teardown is asserted on the real pane coordinate).
|
||||
long paneCloseCount = herdr.calls.stream()
|
||||
.filter(c -> "pane.close".equals(c.method()))
|
||||
.filter(c -> "w9:pRoot_1".equals(((Map<?, ?>) c.params()).get("pane_id")))
|
||||
.count();
|
||||
assertEquals(1, paneCloseCount, "the old worker was torn down");
|
||||
String warn = appender.list.stream()
|
||||
.filter(e -> e.getLevel().equals(Level.WARN))
|
||||
.map(ILoggingEvent::getFormattedMessage)
|
||||
.findFirst()
|
||||
.orElse("no turn-failed WARN logged");
|
||||
assertTrue(warn.contains(terminal), "the log names the member: " + warn);
|
||||
assertTrue(warn.contains("BUSY"), "the log names the stage it failed at: " + warn);
|
||||
} finally {
|
||||
sessionLog.detachAppender(appender);
|
||||
}
|
||||
}
|
||||
|
||||
@Test
|
||||
void rosterReflectsAcquiredMinusReleased() {
|
||||
FakeHerdr herdr = new FakeHerdr();
|
||||
SessionManager sessions = sessionManager(herdr);
|
||||
WorkerSession a = sessions.acquire("ltms-local", "/a", "/caller", "ownerA");
|
||||
WorkerSession b = sessions.acquire("ltms-local", "/b", "/caller", "ownerB");
|
||||
MemberSession a = sessions.acquire("ltms-local", "/a", "/caller", "ownerA");
|
||||
MemberSession b = sessions.acquire("ltms-local", "/b", "/caller", "ownerB");
|
||||
|
||||
assertEquals(2, sessions.roster().size());
|
||||
assertTrue(sessions.roster().stream().anyMatch(s -> s.paneId().equals(a.paneId())));
|
||||
@@ -219,7 +230,7 @@ class SessionManagerTest {
|
||||
long[] clock = {0};
|
||||
FakeHerdr herdr = new FakeHerdr();
|
||||
SessionManager sessions = sessionManager(herdr, () -> clock[0]);
|
||||
WorkerSession session = sessions.acquire("ltms-local", null, "/caller", "term_primary");
|
||||
MemberSession session = sessions.acquire("ltms-local", null, "/caller", "term_primary");
|
||||
String terminal = session.terminalId();
|
||||
sessions.asPresence().markPresent(terminal);
|
||||
|
||||
@@ -234,13 +245,13 @@ class SessionManagerTest {
|
||||
long[] clock = {0};
|
||||
FakeHerdr herdr = new FakeHerdr();
|
||||
SessionManager sessions = sessionManager(herdr, () -> clock[0]);
|
||||
WorkerSession session = sessions.acquire("ltms-local", null, "/caller", "term_primary");
|
||||
MemberSession session = sessions.acquire("ltms-local", null, "/caller", "term_primary");
|
||||
String terminal = session.terminalId();
|
||||
sessions.asPresence().markPresent(terminal);
|
||||
|
||||
clock[0] = 5;
|
||||
assertEquals(0, sessions.reapIdle(10), "READY session within TTL is not reaped");
|
||||
assertEquals(WorkerSession.State.READY,
|
||||
assertEquals(MemberSession.State.READY,
|
||||
sessions.get(session.paneId()).orElseThrow().state(),
|
||||
"READY session survives");
|
||||
}
|
||||
@@ -250,14 +261,14 @@ class SessionManagerTest {
|
||||
long[] clock = {0};
|
||||
FakeHerdr herdr = new FakeHerdr();
|
||||
SessionManager sessions = sessionManager(herdr, () -> clock[0]);
|
||||
WorkerSession session = sessions.acquire("ltms-local", null, "/caller", "term_primary");
|
||||
MemberSession session = sessions.acquire("ltms-local", null, "/caller", "term_primary");
|
||||
String terminal = session.terminalId();
|
||||
sessions.asPresence().markPresent(terminal);
|
||||
sessions.onDelivered(terminal);
|
||||
|
||||
clock[0] = 100;
|
||||
assertEquals(0, sessions.reapIdle(10), "BUSY session past TTL is never reaped");
|
||||
assertEquals(WorkerSession.State.BUSY,
|
||||
assertEquals(MemberSession.State.BUSY,
|
||||
sessions.get(session.paneId()).orElseThrow().state(),
|
||||
"BUSY session remains");
|
||||
}
|
||||
@@ -267,7 +278,7 @@ class SessionManagerTest {
|
||||
long[] clock = {0};
|
||||
FakeHerdr herdr = new FakeHerdr();
|
||||
SessionManager sessions = sessionManager(herdr, () -> clock[0]);
|
||||
WorkerSession session = sessions.acquire("ltms-local", null, "/caller", "term_primary");
|
||||
MemberSession session = sessions.acquire("ltms-local", null, "/caller", "term_primary");
|
||||
String terminal = session.terminalId();
|
||||
sessions.asPresence().markPresent(terminal);
|
||||
sessions.onDelivered(terminal);
|
||||
@@ -284,8 +295,8 @@ class SessionManagerTest {
|
||||
FakeHerdr herdr = new FakeHerdr();
|
||||
SessionManager sessions = sessionManager(herdr, () -> clock[0]);
|
||||
|
||||
WorkerSession ready = sessions.acquire("ltms-local", "/ready", "/caller", "owner1");
|
||||
WorkerSession busy = sessions.acquire("ltms-local", "/busy", "/caller", "owner2");
|
||||
MemberSession ready = sessions.acquire("ltms-local", "/ready", "/caller", "owner1");
|
||||
MemberSession busy = sessions.acquire("ltms-local", "/busy", "/caller", "owner2");
|
||||
sessions.asPresence().markPresent(ready.terminalId());
|
||||
sessions.asPresence().markPresent(busy.terminalId());
|
||||
sessions.onDelivered(busy.terminalId());
|
||||
@@ -293,7 +304,7 @@ class SessionManagerTest {
|
||||
clock[0] = 50;
|
||||
assertEquals(1, sessions.reapIdle(30), "only READY past TTL is reaped");
|
||||
assertTrue(sessions.get(ready.paneId()).isEmpty(), "READY session is gone");
|
||||
assertEquals(WorkerSession.State.BUSY,
|
||||
assertEquals(MemberSession.State.BUSY,
|
||||
sessions.get(busy.paneId()).orElseThrow().state(),
|
||||
"BUSY session is still registered");
|
||||
}
|
||||
@@ -302,7 +313,7 @@ class SessionManagerTest {
|
||||
void contextCapDisabledSessionSurvivesMultipleTurns() {
|
||||
FakeHerdr herdr = new FakeHerdr();
|
||||
SessionManager sessions = sessionManager(herdr, () -> 0L, 0);
|
||||
WorkerSession session = sessions.acquire("ltms-local", null, "/caller", "term_primary");
|
||||
MemberSession session = sessions.acquire("ltms-local", null, "/caller", "term_primary");
|
||||
String terminal = session.terminalId();
|
||||
sessions.asPresence().markPresent(terminal);
|
||||
|
||||
@@ -311,8 +322,8 @@ class SessionManagerTest {
|
||||
sessions.onDelivered(terminal);
|
||||
sessions.onTurnComplete(terminal);
|
||||
|
||||
WorkerSession updated = sessions.get(session.paneId()).orElseThrow();
|
||||
assertEquals(WorkerSession.State.DONE, updated.state(), "session finishes second turn");
|
||||
MemberSession updated = sessions.get(session.paneId()).orElseThrow();
|
||||
assertEquals(MemberSession.State.DONE, updated.state(), "session finishes second turn");
|
||||
assertEquals(2, updated.turnCount(), "turn count tracks both deliveries");
|
||||
long releaseCloseCount = paneCloseCallsFor(herdr, "w9:pRoot_1"); // the real pane coordinate
|
||||
assertEquals(0, releaseCloseCount, "cap disabled — no forced release of the worker pane");
|
||||
@@ -322,13 +333,13 @@ class SessionManagerTest {
|
||||
void contextCapTwoReleasesAfterSecondComplete() {
|
||||
FakeHerdr herdr = new FakeHerdr();
|
||||
SessionManager sessions = sessionManager(herdr, () -> 0L, 2);
|
||||
WorkerSession session = sessions.acquire("ltms-local", null, "/caller", "term_primary");
|
||||
MemberSession session = sessions.acquire("ltms-local", null, "/caller", "term_primary");
|
||||
String terminal = session.terminalId();
|
||||
sessions.asPresence().markPresent(terminal);
|
||||
|
||||
sessions.onDelivered(terminal);
|
||||
sessions.onTurnComplete(terminal);
|
||||
assertEquals(WorkerSession.State.DONE,
|
||||
assertEquals(MemberSession.State.DONE,
|
||||
sessions.get(session.paneId()).orElseThrow().state(),
|
||||
"first turn completes without release");
|
||||
|
||||
@@ -345,13 +356,13 @@ class SessionManagerTest {
|
||||
void clearAfterTurnResetsContextWithoutDoubleCountingTheTurn() {
|
||||
FakeHerdr herdr = new FakeHerdr();
|
||||
SessionManager sessions = sessionManager(herdr, () -> 0L, 0, true);
|
||||
WorkerSession session = sessions.acquire("ltms-local", null, "/caller", "term_primary");
|
||||
MemberSession session = sessions.acquire("ltms-local", null, "/caller", "term_primary");
|
||||
sessions.asPresence().markPresent(session.terminalId());
|
||||
|
||||
sessions.onDelivered(session.terminalId());
|
||||
assertTrue(sessions.onTurnCompleteWithPostAction(session.terminalId()));
|
||||
|
||||
WorkerSession updated = sessions.get(session.paneId()).orElseThrow();
|
||||
MemberSession updated = sessions.get(session.paneId()).orElseThrow();
|
||||
assertEquals(1, updated.turnCount(), "the reset is housekeeping, not a second delegation");
|
||||
assertEquals(List.of("/clear"), promptTexts(herdr));
|
||||
}
|
||||
@@ -360,7 +371,7 @@ class SessionManagerTest {
|
||||
void contextCapReleaseWinsOverClearAfterTurn() {
|
||||
FakeHerdr herdr = new FakeHerdr();
|
||||
SessionManager sessions = sessionManager(herdr, () -> 0L, 1, true);
|
||||
WorkerSession session = sessions.acquire("ltms-local", null, "/caller", "term_primary");
|
||||
MemberSession session = sessions.acquire("ltms-local", null, "/caller", "term_primary");
|
||||
sessions.asPresence().markPresent(session.terminalId());
|
||||
sessions.onDelivered(session.terminalId());
|
||||
|
||||
@@ -375,13 +386,13 @@ class SessionManagerTest {
|
||||
void clearAfterTurnFalsePreservesCompletionWithoutAControlPrompt() {
|
||||
FakeHerdr herdr = new FakeHerdr();
|
||||
SessionManager sessions = sessionManager(herdr, () -> 0L, 0, false);
|
||||
WorkerSession session = sessions.acquire("ltms-local", null, "/caller", "term_primary");
|
||||
MemberSession session = sessions.acquire("ltms-local", null, "/caller", "term_primary");
|
||||
sessions.asPresence().markPresent(session.terminalId());
|
||||
sessions.onDelivered(session.terminalId());
|
||||
|
||||
sessions.onTurnComplete(session.terminalId());
|
||||
|
||||
assertEquals(WorkerSession.State.DONE, sessions.get(session.paneId()).orElseThrow().state());
|
||||
assertEquals(MemberSession.State.DONE, sessions.get(session.paneId()).orElseThrow().state());
|
||||
assertTrue(promptTexts(herdr).isEmpty());
|
||||
}
|
||||
|
||||
@@ -391,8 +402,8 @@ class SessionManagerTest {
|
||||
FakeHerdr herdr = new FakeHerdr();
|
||||
SessionManager sessions = sessionManager(herdr, () -> clock[0]);
|
||||
|
||||
WorkerSession ready = sessions.acquire("ltms-local", "/ready", "/caller", "ownerR");
|
||||
WorkerSession busy = sessions.acquire("ltms-local", "/busy", "/caller", "ownerB");
|
||||
MemberSession ready = sessions.acquire("ltms-local", "/ready", "/caller", "ownerR");
|
||||
MemberSession busy = sessions.acquire("ltms-local", "/busy", "/caller", "ownerB");
|
||||
sessions.asPresence().markPresent(ready.terminalId());
|
||||
sessions.asPresence().markPresent(busy.terminalId());
|
||||
sessions.onDelivered(busy.terminalId());
|
||||
@@ -432,7 +443,7 @@ class SessionManagerTest {
|
||||
long[] clock = {0};
|
||||
|
||||
// Gate-enabled launcher (1 ms timeout + no-op sleeper that advances clock past deadline)
|
||||
BridgedConfig.Worker cfg = new BridgedConfig.Worker(
|
||||
BridgedConfig.Profile cfg = new BridgedConfig.Profile(
|
||||
"ltms-local", "http://gx00.gw:8000", "coder", null, "BRIDGED_WORKER_TOKEN",
|
||||
List.of("ccs", "ltms-local"), "tab", "bridged-workers",
|
||||
"worker: {profile} #{n}", null, null, null);
|
||||
@@ -461,7 +472,7 @@ class SessionManagerTest {
|
||||
java.util.List<String> released = new java.util.concurrent.CopyOnWriteArrayList<>();
|
||||
sessions.onRelease(released::add);
|
||||
|
||||
WorkerSession s = sessions.acquire("ltms-local", null, "/caller", null);
|
||||
MemberSession s = sessions.acquire("ltms-local", null, "/caller", null);
|
||||
sessions.release(s.paneId());
|
||||
|
||||
assertEquals(java.util.List.of(s.terminalId()), released,
|
||||
@@ -488,7 +499,7 @@ class SessionManagerTest {
|
||||
throw new IllegalStateException("listener blew up");
|
||||
});
|
||||
|
||||
WorkerSession s = sessions.acquire("ltms-local", null, "/caller", null);
|
||||
MemberSession s = sessions.acquire("ltms-local", null, "/caller", null);
|
||||
assertDoesNotThrow(() -> sessions.release(s.paneId()),
|
||||
"a listener failure must never prevent the teardown it is reacting to");
|
||||
assertTrue(sessions.get(s.paneId()).isEmpty(), "and the session is still deregistered");
|
||||
|
||||
@@ -5,7 +5,7 @@ import dev.ltms.bridged.guard.SubscriptionGuard;
|
||||
import dev.ltms.bridged.herdr.AgentControl;
|
||||
import dev.ltms.bridged.herdr.FakeHerdr;
|
||||
import dev.ltms.bridged.herdr.WorkspaceControl;
|
||||
import dev.ltms.bridged.worker.ClaudeCodeLauncher;
|
||||
import dev.ltms.bridged.member.ClaudeCodeLauncher;
|
||||
import org.junit.jupiter.api.Test;
|
||||
|
||||
import java.util.List;
|
||||
@@ -29,7 +29,7 @@ class SessionReaperTest {
|
||||
|
||||
private static ClaudeCodeLauncher launcher() {
|
||||
FakeHerdr herdr = new FakeHerdr();
|
||||
BridgedConfig.Worker cfg = new BridgedConfig.Worker(
|
||||
BridgedConfig.Profile cfg = new BridgedConfig.Profile(
|
||||
"ltms-local", "http://gx00.gw:8000", "coder", null, "BRIDGED_WORKER_TOKEN",
|
||||
List.of("ccs", "ltms-local"), "tab", "bridged-workers",
|
||||
"worker: {profile} #{n}", null, null, null);
|
||||
|
||||
@@ -1,16 +1,19 @@
|
||||
package dev.ltms.bridged.session;
|
||||
|
||||
import dev.ltms.bridged.auth.MemberRegistry;
|
||||
import dev.ltms.bridged.config.BridgedConfig;
|
||||
import dev.ltms.bridged.guard.SubscriptionGuard;
|
||||
import dev.ltms.bridged.herdr.AgentControl;
|
||||
import dev.ltms.bridged.herdr.FakeHerdr;
|
||||
import dev.ltms.bridged.herdr.WorkspaceControl;
|
||||
import dev.ltms.bridged.worker.ClaudeCodeLauncher;
|
||||
import dev.ltms.bridged.member.ClaudeCodeLauncher;
|
||||
import dev.ltms.bridged.peer.MemberRole;
|
||||
import org.junit.jupiter.api.Test;
|
||||
|
||||
import java.util.List;
|
||||
import java.util.Map;
|
||||
import java.util.Set;
|
||||
import java.util.concurrent.TimeUnit;
|
||||
|
||||
import static org.junit.jupiter.api.Assertions.*;
|
||||
|
||||
@@ -21,8 +24,15 @@ import static org.junit.jupiter.api.Assertions.*;
|
||||
*/
|
||||
class WorktreeSessionManagerTest {
|
||||
|
||||
private static MemberRegistry members() {
|
||||
return new MemberRegistry(new BridgedConfig.Fleet(Map.of(),
|
||||
Map.of("architect", new BridgedConfig.Slot("ltms-local")),
|
||||
Map.of("dev", new BridgedConfig.Slot("ltms-local")),
|
||||
Map.of("reviewer", new BridgedConfig.Slot("ltms-local")), null));
|
||||
}
|
||||
|
||||
private static ClaudeCodeLauncher workerService(FakeHerdr herdr) {
|
||||
BridgedConfig.Worker cfg = new BridgedConfig.Worker(
|
||||
BridgedConfig.Profile cfg = new BridgedConfig.Profile(
|
||||
"ltms-local", "http://gx00.gw:8000", "coder", null, "BRIDGED_WORKER_TOKEN",
|
||||
List.of("ccs", "ltms-local"), "tab", "bridged-workers",
|
||||
"worker: {profile} #{n}", null, null, null);
|
||||
@@ -44,7 +54,7 @@ class WorktreeSessionManagerTest {
|
||||
FakeWorktrees worktrees = new FakeWorktrees();
|
||||
SessionManager sessions = new SessionManager(workerService(herdr), worktrees);
|
||||
|
||||
WorkerSession s = sessions.acquire("ltms-local", null, "/caller/proj", "term_primary");
|
||||
MemberSession s = sessions.acquire("ltms-local", null, "/caller/proj", "term_primary");
|
||||
|
||||
assertTrue(worktrees.addCalls().isEmpty(), "shared-tree acquire never adds a worktree");
|
||||
assertTrue(worktrees.repoRootCalls().isEmpty(), "shared-tree acquire never resolves a repo root");
|
||||
@@ -55,13 +65,43 @@ class WorktreeSessionManagerTest {
|
||||
assertEquals("/caller/proj", startCwd(herdr), "spawn receives the caller's cwd");
|
||||
}
|
||||
|
||||
@Test
|
||||
void onlyArchitectsBindAndReleaseMakesTheirSlotReusable() {
|
||||
FakeHerdr herdr = new FakeHerdr();
|
||||
MemberRegistry members = members();
|
||||
SessionManager sessions = new SessionManager(workerService(herdr), new FakeWorktrees());
|
||||
sessions.setMemberLifecycle(members);
|
||||
|
||||
MemberSession architect = sessions.acquire("ltms-local", MemberRole.ARCHITECT,
|
||||
null, "/caller/proj", null, null);
|
||||
MemberSession dev = sessions.acquire("ltms-local", MemberRole.DEV,
|
||||
null, "/caller/proj", null, null);
|
||||
MemberSession reviewer = sessions.acquire("ltms-local", MemberRole.REVIEWER,
|
||||
null, "/caller/proj", null, null);
|
||||
|
||||
assertEquals("architect:architect", members.slotForTerminal(architect.terminalId()));
|
||||
assertNull(members.slotForTerminal(dev.terminalId()), "a dev must never receive architect rights");
|
||||
assertNull(members.slotForTerminal(reviewer.terminalId()),
|
||||
"a reviewer must never receive architect rights");
|
||||
|
||||
sessions.release(architect.paneId());
|
||||
MemberSession replacement = sessions.acquire("ltms-local", MemberRole.ARCHITECT,
|
||||
null, "/caller/proj", null, null);
|
||||
assertEquals("architect:architect", members.slotForTerminal(replacement.terminalId()));
|
||||
|
||||
MemberSession overflow = sessions.acquire("ltms-local", MemberRole.ARCHITECT,
|
||||
null, "/caller/proj", null, null);
|
||||
assertNull(members.slotForTerminal(overflow.terminalId()),
|
||||
"a full slot pool must not stop the architect spawn");
|
||||
}
|
||||
|
||||
@Test
|
||||
void worktreeAcquireProvisionsAndRecordsPathAndBranch() {
|
||||
FakeHerdr herdr = new FakeHerdr();
|
||||
FakeWorktrees worktrees = new FakeWorktrees().withRepoRoot("/repo").withPrefix("/wt");
|
||||
SessionManager sessions = new SessionManager(workerService(herdr), worktrees);
|
||||
|
||||
WorkerSession s = sessions.acquire("ltms-local", null, "/caller/proj", "term_primary",
|
||||
MemberSession s = sessions.acquire("ltms-local", null, "/caller/proj", "term_primary",
|
||||
new WorktreeRequest("cb-999", null));
|
||||
|
||||
assertEquals(1, worktrees.addCalls().size(), "one worktree was added");
|
||||
@@ -78,6 +118,19 @@ class WorktreeSessionManagerTest {
|
||||
assertEquals(expectedPath, s.cwd(), "session cwd is the worktree path");
|
||||
}
|
||||
|
||||
@Test
|
||||
void worktreeArchitectAcquireAlsoBindsItsSlot() {
|
||||
FakeHerdr herdr = new FakeHerdr();
|
||||
MemberRegistry members = members();
|
||||
SessionManager sessions = new SessionManager(workerService(herdr), new FakeWorktrees());
|
||||
sessions.setMemberLifecycle(members);
|
||||
|
||||
MemberSession architect = sessions.acquire("ltms-local", MemberRole.ARCHITECT,
|
||||
null, "/caller/proj", null, new WorktreeRequest("cb-548", null));
|
||||
|
||||
assertEquals("architect:architect", members.slotForTerminal(architect.terminalId()));
|
||||
}
|
||||
|
||||
@Test
|
||||
void worktreeAcquireRunsParityOverlayWithProfileDefaults() {
|
||||
FakeHerdr herdr = new FakeHerdr();
|
||||
@@ -93,12 +146,12 @@ class WorktreeSessionManagerTest {
|
||||
FakeWorktrees.OverlayCall overlay = worktrees.lastOverlay();
|
||||
assertNotNull(overlay);
|
||||
assertEquals("/repo", overlay.repoRoot());
|
||||
assertEquals(List.of(".claude/settings.local.json", ".env", ".envrc"),
|
||||
assertEquals(List.of(".env", ".envrc"),
|
||||
overlay.requested(), "default parity overlay is used when unset");
|
||||
assertFalse(overlay.requested().contains(".mcp.json"),
|
||||
"CB-525: replicating the primary's MCP config gives a worker the primary's IDE "
|
||||
+ "servers, which navigate its edits out of its own worktree");
|
||||
assertEquals(List.of(".claude/settings.local.json", ".envrc"), overlay.copied(),
|
||||
assertEquals(List.of(".envrc"), overlay.copied(),
|
||||
"existing paths are copied; missing paths are skipped");
|
||||
assertEquals(List.of(".envrc"), overlay.skipWorktree(),
|
||||
"tracked copied paths are --skip-worktree'd");
|
||||
@@ -109,7 +162,7 @@ class WorktreeSessionManagerTest {
|
||||
FakeHerdr herdr = new FakeHerdr();
|
||||
FakeWorktrees worktrees = new FakeWorktrees().withRepoRoot("/repo").withPrefix("/wt");
|
||||
SessionManager sessions = new SessionManager(workerService(herdr), worktrees);
|
||||
WorkerSession s = sessions.acquire("ltms-local", null, "/caller/proj", null,
|
||||
MemberSession s = sessions.acquire("ltms-local", null, "/caller/proj", null,
|
||||
new WorktreeRequest("cb-666", null));
|
||||
String paneId = s.paneId();
|
||||
|
||||
@@ -125,12 +178,48 @@ class WorktreeSessionManagerTest {
|
||||
assertTrue(sessions.get(paneId).isEmpty(), "released session is no longer retrievable");
|
||||
}
|
||||
|
||||
@Test
|
||||
void drainAllPreservesWorktreeOfIdleSession() {
|
||||
FakeHerdr herdr = new FakeHerdr();
|
||||
FakeWorktrees worktrees = new FakeWorktrees().withRepoRoot("/repo").withPrefix("/wt");
|
||||
SessionManager sessions = new SessionManager(workerService(herdr), worktrees);
|
||||
MemberSession s = sessions.acquire("ltms-local", null, "/caller/proj", null,
|
||||
new WorktreeRequest("cb-544", null));
|
||||
sessions.asPresence().markPresent(s.terminalId()); // READY (idle)
|
||||
|
||||
sessions.drainAll(TimeUnit.MILLISECONDS.toNanos(100));
|
||||
|
||||
assertTrue(herdr.called("pane.close"), "shutdown drain still stops the worker pane");
|
||||
assertTrue(worktrees.removeCalls().isEmpty(),
|
||||
"shutdown drain must NOT remove the worktree — it is the only copy of the work");
|
||||
assertTrue(sessions.roster().isEmpty(), "shutdown drain still deregisters the session");
|
||||
}
|
||||
|
||||
@Test
|
||||
void drainAllPreservesWorktreeOfSessionStillBusyAtTimeout() {
|
||||
FakeHerdr herdr = new FakeHerdr();
|
||||
FakeWorktrees worktrees = new FakeWorktrees().withRepoRoot("/repo").withPrefix("/wt");
|
||||
SessionManager sessions = new SessionManager(workerService(herdr), worktrees);
|
||||
MemberSession s = sessions.acquire("ltms-local", null, "/caller/proj", null,
|
||||
new WorktreeRequest("cb-544", null));
|
||||
String terminal = s.terminalId();
|
||||
sessions.asPresence().markPresent(terminal);
|
||||
sessions.onDelivered(terminal); // BUSY, never completes → still BUSY when the timeout hits
|
||||
|
||||
sessions.drainAll(TimeUnit.MILLISECONDS.toNanos(100));
|
||||
|
||||
assertTrue(herdr.called("pane.close"), "shutdown drain stops the worker pane");
|
||||
assertTrue(worktrees.removeCalls().isEmpty(),
|
||||
"a session still BUSY at timeout must preserve its worktree unconditionally");
|
||||
assertTrue(sessions.roster().isEmpty(), "shutdown drain still deregisters the session");
|
||||
}
|
||||
|
||||
@Test
|
||||
void sharedTreeReleaseMakesNoWorktreesCalls() {
|
||||
FakeHerdr herdr = new FakeHerdr();
|
||||
FakeWorktrees worktrees = new FakeWorktrees();
|
||||
SessionManager sessions = new SessionManager(workerService(herdr), worktrees);
|
||||
WorkerSession s = sessions.acquire("ltms-local", null, "/caller/proj", null);
|
||||
MemberSession s = sessions.acquire("ltms-local", null, "/caller/proj", null);
|
||||
|
||||
sessions.release(s.paneId());
|
||||
|
||||
@@ -159,9 +248,9 @@ class WorktreeSessionManagerTest {
|
||||
FakeWorktrees worktrees = new FakeWorktrees().withRepoRoot("/repo").withPrefix("/wt");
|
||||
SessionManager sessions = new SessionManager(workerService(herdr), worktrees);
|
||||
|
||||
WorkerSession a = sessions.acquire("ltms-local", null, "/caller/proj", null,
|
||||
MemberSession a = sessions.acquire("ltms-local", null, "/caller/proj", null,
|
||||
new WorktreeRequest("cb-444", null));
|
||||
WorkerSession b = sessions.acquire("ltms-local", null, "/caller/proj", null,
|
||||
MemberSession b = sessions.acquire("ltms-local", null, "/caller/proj", null,
|
||||
new WorktreeRequest("cb-444", null));
|
||||
|
||||
assertNotEquals(a.branch(), b.branch(), "branches are distinct");
|
||||
@@ -208,7 +297,7 @@ class WorktreeSessionManagerTest {
|
||||
FakeHerdr herdr = new FakeHerdr();
|
||||
FakeWorktrees worktrees = new FakeWorktrees().withRepoRoot("/repo").withPrefix("/wt");
|
||||
// Argument order matters: configDir is the 4th parameter, cwd the 11th (after mcpUrl).
|
||||
BridgedConfig.Worker cfg = new BridgedConfig.Worker(
|
||||
BridgedConfig.Profile cfg = new BridgedConfig.Profile(
|
||||
"ltms-local", "http://gx00.gw:8000", "coder", null, "BRIDGED_WORKER_TOKEN",
|
||||
List.of("ccs", "ltms-local"), "tab", "bridged-workers",
|
||||
"worker: {profile} #{n}", null, "/pinned/dir", null);
|
||||
|
||||
@@ -16,6 +16,9 @@
|
||||
AuditLogTest, which attaches its own ListAppender and asserts on emitted records.
|
||||
-->
|
||||
|
||||
<!-- Keep the test logger behaviour aligned with the production cancellation filter. -->
|
||||
<turboFilter class="dev.ltms.bridged.logging.McpCancelledNotificationFilter"/>
|
||||
|
||||
<appender name="STDOUT" class="ch.qos.logback.core.ConsoleAppender">
|
||||
<encoder>
|
||||
<pattern>%d{HH:mm:ss.SSS} [%thread] %-5level %logger{36} - %msg%n</pattern>
|
||||
@@ -34,4 +37,4 @@
|
||||
<appender-ref ref="STDOUT"/>
|
||||
</root>
|
||||
|
||||
</configuration>
|
||||
</configuration>
|
||||
|
||||
@@ -35,7 +35,7 @@ build on.
|
||||
- **No checkpoint content.** Writing `STATE.md` + commit on teardown is CB-302; CB-301 only exposes
|
||||
the release hook it will attach to.
|
||||
|
||||
"Recycle" under no-reuse is simply **release + fresh acquire** — a helper, not a pool operation.
|
||||
Under no-reuse, a released session is terminal. A new `acquire` always creates a fresh session.
|
||||
|
||||
## Design
|
||||
|
||||
@@ -46,11 +46,6 @@ ownership on top.
|
||||
**Package:** new `dev.ltms.bridged.session` — keeps the registry/lifecycle concern separate from
|
||||
the `worker` spawn mechanics. Holds `SessionManager` + `WorkerSession`.
|
||||
|
||||
**`recycle` is IN SCOPE for CB-301** (decided): implement `recycle(paneId, …)` = `release` the old
|
||||
session then `acquire` a fresh one, asserting a new distinct paneId (the no-reuse invariant). It is
|
||||
a thin convenience over the two primitives, shipped now so the no-reuse teardown+respawn path is
|
||||
covered by a test from day one.
|
||||
|
||||
### `WorkerSession` (record or small mutable holder)
|
||||
|
||||
| Field | Source | Notes |
|
||||
@@ -88,7 +83,6 @@ SPAWNING|READY|BUSY|DONE --vanished/drop--> FAILED
|
||||
final class SessionManager {
|
||||
WorkerSession acquire(String profile, String requestedCwd, String callerCwd, String ownerTerminal);
|
||||
void release(String paneId); // deterministic teardown + deregister
|
||||
WorkerSession recycle(String paneId, ...); // release + acquire (no-reuse convenience)
|
||||
Optional<WorkerSession> get(String paneId);
|
||||
List<WorkerSession> roster(); // bridge-owned view (CB-304 consumes this)
|
||||
// lifecycle hooks (package-private): onReady/onDelivered/onComplete/onFailed(target)
|
||||
@@ -120,8 +114,7 @@ final class SessionManager {
|
||||
3. `release` tears the worker down via `WorkerService.stop` and removes it from `roster()`;
|
||||
a second `release` on the same paneId is a harmless no-op.
|
||||
4. `onTurnFailed` / drop moves the session to `FAILED` and it is absent from the live roster.
|
||||
5. `recycle` produces a new paneId and the old one is gone (no-reuse invariant).
|
||||
6. `roster()` reflects exactly the sessions acquired-minus-released, joined with live status.
|
||||
5. `roster()` reflects exactly the sessions acquired-minus-released, joined with live status.
|
||||
|
||||
## Seams left open (deliberately)
|
||||
|
||||
|
||||
@@ -1,6 +1,8 @@
|
||||
# CB-500 — Multi-Tier Coordination (Stage 6)
|
||||
|
||||
**Status:** design note (proposal — ticket split deferred)
|
||||
**Status:** design note. Developments A/B remain proposals; Development C (§6 and Figures 7–8) is
|
||||
**SUPERSEDED** by the advisory-architect design in Gitea issue #16 and the `architects:` configuration
|
||||
block (CB-548).
|
||||
**Depends on:** CB-401/402 (Peer Launcher SPI + composite router — placement-neutral spawn),
|
||||
CB-308 (per-agent broker channels + global id + federated roster — the addressing substrate),
|
||||
CB-307 (durable inbox + push loop), CB-301/303 (session FSM + context-cap/idle-ttl), CB-304
|
||||
@@ -16,8 +18,9 @@ workers) into a **multi-tier** one, along three axes the lead has asked for:
|
||||
1. **Sandboxed workers** — each worker runs in a **separated, peer-owned sandbox** carrying its own
|
||||
toolchain (Claude routed via `ANTHROPIC_BASE_URL`, a headless IDE, git, MCP, dev-tools), with
|
||||
**per-role** sandboxes (a backend-agent image, a frontend-agent image).
|
||||
2. **Main-agent pairs** — the "main" tier becomes a **pair** (on-subscription Opus + one cloud
|
||||
module) collaborating, instead of a lone primary.
|
||||
2. **Main-agent pairs** — **SUPERSEDED.** The considered model made the "main" tier a pair
|
||||
(on-subscription Opus + one cloud module). The actual fleet is one human-driven lead plus two
|
||||
short-lived advisory architects on different model families.
|
||||
3. **An orchestrator tier** — a supervisor **above** the mains that owns their **session identity**
|
||||
(naming, resume) and **curates context**, so every main→worker delegation carries the *exact*
|
||||
slice of context it needs and nothing else.
|
||||
@@ -63,6 +66,10 @@ Four concrete bake-ins assume a single tier:
|
||||
|
||||
## 3. Target multi-tier architecture
|
||||
|
||||
> **SUPERSEDED fleet sketch.** Figure 2 records the former two-main model. The actual fleet is one
|
||||
> lead, two independent advisory architects, and N workers; architects are sideways peers, not leads
|
||||
> and not a tier above the lead. See Gitea issue #16 and the `architects:` block.
|
||||
|
||||
```mermaid
|
||||
flowchart TB
|
||||
human["human"]
|
||||
@@ -93,10 +100,9 @@ flowchart TB
|
||||
chan --- roster
|
||||
```
|
||||
|
||||
*Figure 2 — three tiers. Tier 0 owns the mains' session identity + context scope; Tier 1 is a
|
||||
collaborating pair, each an MCP client with its own pull inbox; Tier 2 is peer-owned sandboxes the
|
||||
bus launches into. The middle is CB-308's per-agent-channel + federated-roster substrate, now
|
||||
carrying tier-to-tier traffic, not just host-to-host.*
|
||||
*Figure 2 — **SUPERSEDED historical fleet sketch.** It proposed a collaborating pair of managed mains.
|
||||
The actual fleet keeps one human-driven lead and uses two independent, short-lived advisory architects
|
||||
on different model families, so agreement is evidence rather than correlated echo.*
|
||||
|
||||
The recursion is the key idea: **`orchestrator : mains :: main : workers`** — the same
|
||||
spawn/name/resume/scope verbs at two levels.
|
||||
@@ -166,6 +172,11 @@ gateway, because herdr keystroke-injection needs a locally-owned PTY.**
|
||||
|
||||
## 5. Development B — Main-agent pairs
|
||||
|
||||
> **SUPERSEDED — do not implement this model.** The two-main fleet was replaced by one human-driven
|
||||
> lead and two independent advisory architects. They are deliberately different model families (Claude
|
||||
> Sonnet 5 and GPT-5.6 through opencode), receive the same brief, and work independently so agreement
|
||||
> is evidence rather than correlated echo. See Gitea issue #16 and `architects:`.
|
||||
|
||||
Both mains are MCP **clients**, so **neither can be called into** — each needs a **pull-based
|
||||
per-agent inbox**, which is precisely CB-308 item #1 (per-agent AMQP channels). The primary machinery
|
||||
that is singular today (single-slot `PrimaryRegistry`, a push-loop aimed at one terminal, "these
|
||||
@@ -221,6 +232,19 @@ push-loop fan-out; relax "orchestration tools only the primary calls" to "any re
|
||||
|
||||
## 6. Development C — Orchestrator tier
|
||||
|
||||
> **SUPERSEDED — do not implement this model.** The operator rejected a supervisor above the lead.
|
||||
> The human continues to drive the pre-existing lead directly; bridged neither spawns nor resumes that
|
||||
> lead. What replaced this proposal is **one lead, two short-lived advisory architects, and N workers**:
|
||||
> the lead engages architects sideways for a strong-model assessment, then discards them. Architect
|
||||
> slots are declared in `architects:` (see Gitea issue #16), rather than making leads managed sessions.
|
||||
> The two architects deliberately use different model families — Claude Sonnet 5 and GPT-5.6 through
|
||||
> opencode — and receive the same brief independently. Agreement is evidence, not correlated echo
|
||||
> from one provider or one conversation.
|
||||
|
||||
> **Historical alternative retained.** The text and figures below record the considered model and why it
|
||||
> was rejected: it re-rooted the human-facing session above the lead, violating the still-true premise
|
||||
> that configured leaders pre-exist, are recognised, and cannot be resumed by bridged.
|
||||
|
||||
The orchestrator is **`SessionManager` recursed one tier up**: today it spawns/names/reaps *worker*
|
||||
sessions; the orchestrator does the same for *main* sessions, and adds **context scoping**.
|
||||
|
||||
@@ -249,11 +273,10 @@ flowchart TB
|
||||
m2 -->|"scoped delegation"| w
|
||||
```
|
||||
|
||||
*Figure 7 — the recursion. Tiers 1 and 2 run the identical spawn/name/resume machinery; the
|
||||
orchestrator merely operates it one level higher. **Re-rooting caveat:** today the primary IS the
|
||||
human's live session; here the human drives the orchestrator, and the mains become managed,
|
||||
resumable sessions. That moves the human-facing top up a tier — an intentional re-root, not an
|
||||
add-on.*
|
||||
*Figure 7 — **SUPERSEDED historical alternative.** The recursion re-rooted the human-facing session:
|
||||
the human drove an orchestrator and the mains became managed, resumable sessions. The operator rejected
|
||||
that re-root. The replacement keeps the human-driven, pre-existing lead and engages architects sideways
|
||||
as short-lived advisory peers; see Gitea issue #16 and `architects:`.*
|
||||
|
||||
```mermaid
|
||||
sequenceDiagram
|
||||
@@ -270,10 +293,9 @@ sequenceDiagram
|
||||
O->>O: fold into orchestrator context, pick next main/turn
|
||||
```
|
||||
|
||||
*Figure 8 — context focus. The orchestrator holds the global context and hands each main only the
|
||||
slice a given delegation needs, so the main→worker conversation stays on-point. Context *scoping* is
|
||||
coordination (the bus already owns session/turn lifecycle) — it stays inside the identity boundary
|
||||
(§7), unlike toolchain ownership which does not.*
|
||||
*Figure 8 — **SUPERSEDED historical alternative.** This proposed an orchestrator holding global context
|
||||
and slicing it for managed mains. The replacement has the human-driven lead send the same advisory brief
|
||||
issue #16 and `architects:`.*
|
||||
|
||||
**Deltas:** a second, higher `SessionManager` instance whose "peers" are mains; the orchestrator
|
||||
becomes the top MCP client; context-slice selection (new) layered on CB-303's `context_cap` +
|
||||
@@ -315,7 +337,7 @@ flowchart LR
|
||||
cb402["CB-401/402<br/>Peer Launcher SPI + composite<br/>(DONE / in-flight)"]
|
||||
A["A · SandboxLauncher<br/>(placement-neutral, independent)"]
|
||||
cb308["CB-308 substrate<br/>per-agent channels + global id<br/>+ federated roster"]
|
||||
B["B · main-agent pair<br/>(multi-slot PrimaryRegistry)"]
|
||||
B["B · main-agent pair (SUPERSEDED)<br/>(multi-slot PrimaryRegistry)"]
|
||||
C["C · orchestrator tier<br/>(SessionManager recursed up)"]
|
||||
cb402 --> A
|
||||
cb402 --> cb308
|
||||
@@ -333,7 +355,7 @@ flowchart LR
|
||||
2. **A · SandboxLauncher** — independent; a second proof of the SPI (placement-neutral). Ships anytime.
|
||||
3. **CB-308 substrate** — per-agent channels + global id + federated roster (the multi-host work,
|
||||
promoted from host-to-host to tier-to-tier).
|
||||
4. **B · main-agent pair** — multi-slot `PrimaryRegistry` + per-main inbox routing, on the substrate.
|
||||
4. **B · main-agent pair** — **SUPERSEDED** by lead + two advisory architects.
|
||||
5. **C · orchestrator tier** — the capstone; the recursive session manager + context scoping.
|
||||
|
||||
## 9. Open questions (to resolve at ticket-split)
|
||||
@@ -341,8 +363,8 @@ flowchart LR
|
||||
- **Sandbox mechanism:** container (`docker exec`) vs devcontainer — how the role→image mapping is
|
||||
expressed on the profile. *(Topology **resolved** in §11: distributed = gateway-per-host × local
|
||||
sandboxes; the remaining choice is only the local launch mechanism, not the shape.)*
|
||||
- **Pair semantics:** are the two mains fully symmetric peers, or is one a co-primary that may also
|
||||
delegate? Affects how `PrimaryRegistry` and the "orchestration tools" identity relax.
|
||||
- **Pair semantics:** **SUPERSEDED.** The two-main question is replaced by the architect role's
|
||||
least-privilege boundary: advisory architects can send/reply/ask/read but cannot spawn/stop/drain.
|
||||
- **Orchestrator drivenness:** the mains become programmatically spawned/resumed — does the human
|
||||
still ever type directly into a main, or only into the orchestrator? (The re-root caveat, Fig 7.)
|
||||
- **Context-slice selection:** who decides the slice — orchestrator heuristics, explicit tool args,
|
||||
|
||||
@@ -0,0 +1,937 @@
|
||||
# M4 - Fleet health, recovery, routing, and capacity
|
||||
|
||||
**Status:** Design accepted on 2026-08-15. CB-573 part 1 has shipped the classification model and
|
||||
the `bridge_list` capacity view; the remaining M4 units are not yet shipped.
|
||||
**Scope:** Fleet evidence, safe mechanical repair, lead routing, capacity reporting, and optional
|
||||
human notification.
|
||||
**Grounded in:** `health/FleetHealth`, `health/PaneBudget`, `inject/StatusPoller`,
|
||||
`inject/StatusRefiner`, `inject/CompletionResolver`, `inject/Injector`, `session/SessionManager`,
|
||||
`msg/MessageService`, `msg/ReplyInbox`, `msg/ReplyPushLoop`, `msg/LeadHeartbeatLoop`,
|
||||
`mcp/PrimaryRegistry`, and `herdr/AgentControl`.
|
||||
|
||||
## 1. Problem and decision boundary
|
||||
|
||||
The operator asked the bridge to detect idle agents, exceptions, stopped work, and broken
|
||||
communication. The bridge may read an agent pane from time to time. It must notify a person when
|
||||
the fleet cannot move forward.
|
||||
|
||||
The four operator terms are not four equal health states. `IDLE` is a normal mode. An exception is
|
||||
sometimes visible only as pane text. Stopped work may look the same as slow work. Broken
|
||||
communication can occur on several links.
|
||||
|
||||
M4 uses this boundary:
|
||||
|
||||
- The bridge detects facts and joins evidence.
|
||||
- The bridge repairs only mechanical failures with no judgement.
|
||||
- The lead decides whether to stop, retry, replace, or reassign a member.
|
||||
- A human is notified only when no healthy lead can act.
|
||||
- n8n may route an outbound incident. It never classifies state or chooses recovery.
|
||||
|
||||
An inbound n8n decider would need bridge authority. No narrow machine-decider role exists. Giving a
|
||||
workflow engine lead authority is unsafe, while adding a new role is a separate authorization
|
||||
design. An outbound sink needs no bridge role.
|
||||
|
||||
The bridge must never replay a delivered task. That task may already have changed files, pushed a
|
||||
branch, opened a pull request, or changed external state. A replay can run those side effects twice.
|
||||
This rule must remain true even if later code stores delivered prompt text.
|
||||
|
||||
## 2. Evidence model
|
||||
|
||||
A health state is mainly a comparison between two views:
|
||||
|
||||
- **herdr view:** current agents and raw live status from one `AgentControl.list()` call.
|
||||
- **bridge view:** session FSM, MCP presence, accepted turns, tasks, inbox state, and lead ownership.
|
||||
|
||||
A strong fault often appears as a disagreement between those views. For example, `BUSY` in the
|
||||
session FSM and `DONE` in herdr means the bridge missed a turn boundary. Pane reads support this
|
||||
model, but they are not the main monitor.
|
||||
|
||||
`SessionManager.rosterView` already joins session state and live status. `AgentControl.list()`
|
||||
already gets the whole live fleet in one call. M4 makes that join persistent and adds timers,
|
||||
accepted-turn state, and incident state.
|
||||
|
||||
### 2.1 Real traces behind the design
|
||||
|
||||
The first trace was an architect that stopped making progress:
|
||||
|
||||
```text
|
||||
profile=opus role=architect state=busy liveStatus=done
|
||||
```
|
||||
|
||||
The session moved from `DONE` to `BUSY` for turn 2. Eighteen minutes later, the session still said
|
||||
`BUSY`, herdr still said `DONE`, the async task still said `PENDING`, and no completion fallback had
|
||||
run. This is `TURN_BOUNDARY_LOST`, not a general slow-turn guess.
|
||||
|
||||
The second trace had two async sends to the same pane, one second apart. The pane was then stopped.
|
||||
One ticket became failed. The other stayed `pending - worker unknown`. Current
|
||||
`MessageService.abandon` resolves only `Rendezvous.currentWaiter(target)`, while async tasks live in
|
||||
a separate ticket map. CB-568 is intended to fix that bug. M4 still keeps an independent
|
||||
post-teardown invariant so a later regression becomes `DELEGATION_ORPHANED`.
|
||||
|
||||
### 2.2 Corrections made during design
|
||||
|
||||
The first state table missed `BUSY` in bridged plus `IDLE` or `DONE` in herdr. It would have found
|
||||
the real trace only through a late, weak stall timer. The final model adds
|
||||
`TURN_BOUNDARY_LOST` as a strong disagreement state.
|
||||
|
||||
The first notification design also required a webhook before `health.enabled` could turn on. That
|
||||
removed useful local detection to avoid a narrower human-notification gap. The final design splits
|
||||
detection from notification. Missing human escalation is shown as partial coverage instead of
|
||||
disabling health.
|
||||
|
||||
## 3. Classification precedence
|
||||
|
||||
Evidence is applied in this order. A lower rule cannot hide a higher one.
|
||||
|
||||
1. **Control link:** failed fleet list plus failed ping becomes `CONTROL_LINK_DOWN`.
|
||||
2. **Definitive target loss:** `_not_found` becomes `GONE` or `LEAD_UNREACHABLE` when the control
|
||||
link is healthy.
|
||||
3. **Startup and teardown invariants:** readiness expiry becomes `NEVER_READY`; surviving tasks
|
||||
after teardown become `DELEGATION_ORPHANED`.
|
||||
4. **Bridge/live disagreement:** `BUSY` plus stable raw `IDLE` or `DONE` becomes
|
||||
`TURN_BOUNDARY_LOST`.
|
||||
5. **Known screen evidence:** a tested fatal signature becomes `ERROR_ON_SCREEN`.
|
||||
6. **Timed suspicion:** unchanged sparse pane probes may become `STALL_SUSPECTED`.
|
||||
7. **Communication quality:** completion fallback becomes `MUTE`; an old inbox entry becomes
|
||||
`REPLY_STRANDED`.
|
||||
8. **Normal mode:** `STARTING`, `IDLE`, `WORKING`, `WORK_PENDING`, or `BLOCKED_AMBIGUOUS`.
|
||||
|
||||
The member flow in Figure 1 shows lifecycle states and the main fault exits. Fault states are
|
||||
reported beside the session FSM; most are not new FSM values.
|
||||
|
||||
```mermaid
|
||||
flowchart TD
|
||||
Registered["Member registered"] --> Starting["STARTING"]
|
||||
Starting -->|"MCP presence"| Idle["IDLE"]
|
||||
Starting -->|"Readiness grace expires"| NeverReady["NEVER_READY"]
|
||||
Idle -->|"Accepted delivery"| Working["WORKING"]
|
||||
Working -->|"Trusted turn boundary"| Idle
|
||||
Working -->|"Bridge BUSY and herdr IDLE or DONE"| Lost["TURN_BOUNDARY_LOST"]
|
||||
Working -->|"Known fatal screen"| Error["ERROR_ON_SCREEN"]
|
||||
Working -->|"Long age and unchanged sparse probes"| Stall["STALL_SUSPECTED"]
|
||||
Working -->|"Target not found"| Gone["GONE"]
|
||||
Idle -->|"Inbox or queued delivery exists"| Pending["WORK_PENDING"]
|
||||
Pending -->|"Delivery or collection finishes"| Idle
|
||||
Idle -->|"Raw BLOCKED with an open turn"| Blocked["BLOCKED_AMBIGUOUS"]
|
||||
Lost -->|"Strict guarded repair"| Repaired["DONE with reconciled completion"]
|
||||
Lost -->|"Repair refused"| LeadDecision["Lead decision required"]
|
||||
```
|
||||
|
||||
*Figure 1. The member lifecycle and the main health exits. Pane-based states never authorise an
|
||||
automatic retry of the task.*
|
||||
|
||||
## 4. State model
|
||||
|
||||
### 4.1 Normal and transitional member states
|
||||
|
||||
| State | Exact evidence | Meaning and certainty |
|
||||
|---|---|---|
|
||||
| `STARTING` | Session is `SPAWNING`; MCP presence is absent | Normal inside the startup grace. MCP contact is the readiness signal. |
|
||||
| `IDLE` | Session is `READY` or `DONE`; live status is `IDLE` or `DONE`; no open turn or inbox item exists | Normal. Idle is not a fault. |
|
||||
| `WORKING` | Session is `BUSY`; raw live status is `WORKING`; the accepted turn is open | Certain that herdr sees work. It does not prove useful progress. |
|
||||
| `WORK_PENDING` | Queued delivery or inbox content exists while the target is injectable | Transitional. Existing injector or push logic should move it. |
|
||||
| `BLOCKED_AMBIGUOUS` | An open turn exists and raw live status is `BLOCKED` | The bridge cannot tell whether this is permission, input, or a settled screen. |
|
||||
|
||||
Idle may drive configured resource cleanup. It never opens an incident and never pages a person.
|
||||
|
||||
### 4.2 Member fault and quality states
|
||||
|
||||
| State | Exact evidence | Certainty and action |
|
||||
|---|---|---|
|
||||
| `NEVER_READY` | `SPAWNING`, no MCP presence, and an accepted delivery waits through the existing readiness grace | Delivery never became possible. The exact cause is unknown. Fail the send, stop the process, and preserve a provisioned worktree. |
|
||||
| `GONE` | Per-target herdr call returns `_not_found` while fleet list or ping works | Certain target loss. Fail all target work. Do not replay it. |
|
||||
| `TURN_BOUNDARY_LOST` | Same session turn stays `BUSY`; same accepted task stays open; two raw snapshots show `IDLE` or `DONE` | Strong disagreement. Strict reconciliation may repair it. |
|
||||
| `ERROR_ON_SCREEN` | Suspicious non-working state survives grace; `detection` matches a tested adapter-specific fatal signature | Certain only for the matched signature. A bare word such as `Exception` is not enough. |
|
||||
| `STALL_SUSPECTED` | Open turn is older than the configured threshold; two normalised `recent_unwrapped` digests are unchanged; no boundary or reply occurs | Not certain. A long valid API call can look the same. Lead decides. |
|
||||
| `MUTE` | Turn resolves through completion fallback instead of `bridge_reply` | Certain that no structured reply won. It does not prove an MCP failure. A single event is a metric, not an incident. |
|
||||
| `REPLY_STRANDED` | Typed reply or health message remains after owning-lead push reaches its cap | Collection failed. This does not explain whether the lead is busy, dead, or ignoring the nudge. |
|
||||
| `DELEGATION_ORPHANED` | Target is gone, failed, or released, but one or more tasks remain `PENDING` after reconciliation grace | Certain bridge invariant failure. This is not an inbox-drain fault. |
|
||||
| `WORK_PRODUCT_AT_RISK` | Provisioned branch has commits after its recorded base; member is `DONE`, `FAILED`, or preserved after release; no turn or inbox item remains; long-idle threshold passed | A warning, not proof of loss. Work may already have an open pull request or a squash merge. |
|
||||
|
||||
`MUTE` opens an incident only after a small fixed rate threshold for one target or profile, or when
|
||||
it appears with another fault.
|
||||
|
||||
`WORK_PRODUCT_AT_RISK` must not become `WORK_PRODUCT_UNCOLLECTED`. The bridge does not know pull
|
||||
request or merge state. If committed work appears with `REPLY_STRANDED` or
|
||||
`DELEGATION_ORPHANED`, the existing incident gains `committedWorkAtRisk: true`.
|
||||
|
||||
### 4.3 Control-link state
|
||||
|
||||
| State | Exact evidence | Certainty and action |
|
||||
|---|---|---|
|
||||
| `CONTROL_LINK_DOWN` | Two full-fleet `agent.list` calls fail across the grace, and herdr `ping` also fails | Certain for the bridged-to-herdr link. Retry calls, record the incident, and use human escalation if no lead can be reached. |
|
||||
|
||||
A failed fleet list alone is not a dead-member claim. A single `_not_found` with a healthy global
|
||||
link is a target fault, not a control-link fault.
|
||||
|
||||
### 4.4 Lead states
|
||||
|
||||
| State | Exact evidence | Meaning and action |
|
||||
|---|---|---|
|
||||
| `LEAD_IDLE` | Expected lead is present with raw injectable status; no actionable state waits | Normal. Existing heartbeat may run under its own policy. |
|
||||
| `LEAD_WORKING` | Expected lead is present with raw `WORKING`; stall threshold is not met | Reachable and busy. Never inject into the live turn. |
|
||||
| `LEAD_STATUS_UNKNOWN` | Expected lead is present with raw `UNKNOWN` | Neither dead nor a healthy routing target. Retain evidence and retry. |
|
||||
| `LEAD_UNREACHABLE` | Expected lead is absent from two successful live-agent snapshots while ping works, or targeted lookup returns `_not_found` with a healthy control link | Route to a healthy peer. If none exists, use human escalation. |
|
||||
| `LEAD_UNRESPONSIVE` | Actionable state waits; lead stays injectable; bounded nudges exhaust; inbox remains uncollected | Route to a healthy peer or a person. |
|
||||
| `LEAD_STALL_SUSPECTED` | Lead stays `WORKING` past threshold; two sparse pane probes show no progress | Not certain. Never kill or restart automatically. Route to peer or person. |
|
||||
|
||||
The monitor retains the lead name and terminal, last successful sighting, raw status and age,
|
||||
consecutive list absences, targeted errors, pane-probe facts, pending incident age, and nudge
|
||||
outcomes. Current heartbeat and push loops discard much of this history.
|
||||
|
||||
Expected lead identity comes from the same supplier used by `CallerResolver`. It is not liveness
|
||||
evidence. `LeadTabScanner` keeps cached identity after a failed scan, so the health monitor compares
|
||||
that identity with a fresh successful agent list. A dynamic identity also survives a two-successful-
|
||||
snapshot retirement grace. This stops a dead lead from escaping health by disappearing from one map.
|
||||
|
||||
### 4.5 Evidence limits
|
||||
|
||||
M4 cannot tell these cases apart with current evidence:
|
||||
|
||||
- A valid long call and a hung call may have the same status and pane digest.
|
||||
- `BLOCKED` does not explain which input is needed.
|
||||
- An idle prompt after failure may look like an idle prompt after success.
|
||||
- A missing structured reply does not prove a broken MCP connection.
|
||||
- An undrained inbox does not explain why the lead did not collect it.
|
||||
- Arbitrary pane text cannot safely classify arbitrary exceptions.
|
||||
- A branch ahead of its base does not prove that work was not collected.
|
||||
|
||||
Logs are outputs, not classifier inputs. The monitor never parses its own logs.
|
||||
|
||||
## 5. Automatic action and lead action
|
||||
|
||||
### 5.1 Actions the bridge may take
|
||||
|
||||
The bridge may:
|
||||
|
||||
- retry transient herdr status, list, ping, and pane-read failures with bounded backoff;
|
||||
- re-submit Enter after the existing paste/submit race;
|
||||
- fail queued delivery after `NEVER_READY`;
|
||||
- stop a never-ready process while preserving its provisioned worktree;
|
||||
- fail all queued, accepted, and async tasks for a gone or released target;
|
||||
- reconcile one lost boundary when every strict gate in Section 8 passes;
|
||||
- hold typed messages, nudge the owning lead, and stop at the configured cap;
|
||||
- use the existing bounded idle-lead heartbeat;
|
||||
- deduplicate, route, update, and resolve incidents.
|
||||
|
||||
These actions do not choose new work and do not replay old work.
|
||||
|
||||
### 5.2 Decisions reserved for the lead
|
||||
|
||||
Only the lead may:
|
||||
|
||||
- stop or continue `BLOCKED_AMBIGUOUS`;
|
||||
- stop, inspect, or wait on `ERROR_ON_SCREEN`;
|
||||
- kill or continue `STALL_SUSPECTED`;
|
||||
- spawn a replacement or reassign work;
|
||||
- retry a delivered task;
|
||||
- choose how to use partial work in a worktree;
|
||||
- restart herdr or change network, model, credentials, backend, or configuration.
|
||||
|
||||
Reports include literal safe tool calls such as `bridge_status(sessionId="...")`,
|
||||
`bridge_poll(ticket="...")`, `bridge_list()`, and optional `bridge_stop(paneId="...")`. A judgement
|
||||
state never presents stop as the only action.
|
||||
|
||||
### 5.3 Release causes and worktree safety
|
||||
|
||||
| Release cause | Process action | Provisioned worktree |
|
||||
|---|---|---|
|
||||
| `SPAWN_ROLLBACK` before registration or delivery | Stop and clean up | Remove |
|
||||
| `COMPLETED` for `READY` or `DONE` without pending work, idle TTL, or successful context-cap completion | Stop | Remove under completed policy |
|
||||
| `NEVER_READY` | Stop | Preserve |
|
||||
| `GONE` | Best-effort stop | Preserve |
|
||||
| `TURN_FAILED` or lead abort while `BUSY` or `FAILED` | Stop | Preserve |
|
||||
| `RELEASE_WITH_PENDING_TASKS` | Stop | Preserve |
|
||||
| `SHUTDOWN` | Stop | Preserve |
|
||||
|
||||
Explicit stop is state-aware. `SPAWNING`, `BUSY`, `FAILED`, or any target with pending tasks uses a
|
||||
preserving cause.
|
||||
|
||||
Before abnormal release removes the live session, M4 writes an atomic manifest under the worktree
|
||||
root. It records session identity, owner, role, profile, repository, path, branch, base commit,
|
||||
release cause, release time, state, and pending task ids. `bridge_list.preservedWorktrees` loads these
|
||||
manifests after restart. Stop output and WARN logs also name the path and cause. M4 never
|
||||
auto-deletes a preserved worktree.
|
||||
|
||||
## 6. Fleet health monitor
|
||||
|
||||
Add `FleetHealthMonitor`. Do not widen `StatusPoller` into a policy loop.
|
||||
|
||||
`StatusPoller` has a 250 ms delivery cadence and samples only injector targets with outstanding
|
||||
work. Health needs all sessions, all leads, task state, inbox age, and global control evidence. One
|
||||
loop cannot serve both cadences safely.
|
||||
|
||||
Build the monitor like `LeadHeartbeatLoop`:
|
||||
|
||||
- pure `decide(snapshot, priorState, now)` logic;
|
||||
- a thin scheduler;
|
||||
- an injected clock;
|
||||
- edge-triggered state changes;
|
||||
- no network work in the pure function;
|
||||
- no sleeping in tests.
|
||||
|
||||
Each enabled fleet tick reads:
|
||||
|
||||
- one `AgentControl.list()` result for the whole fleet;
|
||||
- one in-memory `SessionManager.roster()` snapshot;
|
||||
- accepted turns and async task state;
|
||||
- typed inbox depth, kind, and age;
|
||||
- push and heartbeat outcomes;
|
||||
- configured and discovered leads.
|
||||
|
||||
Existing failure paths publish structured evidence to the monitor. The monitor does not infer events
|
||||
from log text.
|
||||
|
||||
### 6.1 Pane budget
|
||||
|
||||
Healthy idle members, recent working members, and quiet leads cause no pane reads.
|
||||
|
||||
A pane is eligible only for a stable lost boundary, sustained `BLOCKED` or `UNKNOWN`, work older
|
||||
than the suspect threshold, or one final evidence read for a confirmed fault when the pane exists.
|
||||
|
||||
Compiled brakes apply even if config asks for more:
|
||||
|
||||
- per-target pane cooldown is at least 60 seconds;
|
||||
- working age before the first progress probe is at least 300 seconds;
|
||||
- at most two pane reads occur in one fleet tick;
|
||||
- targets rotate fairly;
|
||||
- only a normalised digest and optional clipped local excerpt are stored;
|
||||
- no pane excerpt leaves bridged in a human webhook.
|
||||
|
||||
Use `detection` for tested screen signatures. Use normalised `recent_unwrapped` only for progress
|
||||
comparison.
|
||||
|
||||
## 7. Typed inbox and routing
|
||||
|
||||
### 7.1 Semantic record
|
||||
|
||||
The typed inbox record carries:
|
||||
|
||||
```text
|
||||
schemaVersion
|
||||
kind: reply | health
|
||||
msgId, target, subjectTerminal, recipientLead
|
||||
severity, state, evidence
|
||||
createdAtEpochMillis, firstSeenEpochMillis, lastSeenEpochMillis
|
||||
recoveryTried, suggestedToolCalls, content
|
||||
```
|
||||
|
||||
A health message never calls `Rendezvous.resolve`. It cannot look like the member's task result.
|
||||
|
||||
Both inbox adapters share field preservation, first-id-wins dedup, FIFO among decoded messages,
|
||||
explicit ownership, ack, and release rules. The in-memory adapter stores typed records directly. It
|
||||
does not copy AMQP migration logic.
|
||||
|
||||
### 7.2 AMQP migration
|
||||
|
||||
The reader uses AMQP `content_type`, never body sniffing:
|
||||
|
||||
```text
|
||||
Legacy v0: text/plain
|
||||
Typed family: application/vnd.ltms.bridged.inbox-message+json
|
||||
```
|
||||
|
||||
A legacy reply may begin with `{`. It remains plain text because its media type is `text/plain`.
|
||||
Legacy text becomes `kind=reply` with exact UTF-8 content and absent typed metadata.
|
||||
|
||||
Typed JSON has required integer `schemaVersion: 1`. Version 1 ignores unknown optional fields.
|
||||
Missing required fields, invalid enums, malformed UTF-8 or JSON, and property/body identity mismatch
|
||||
are invalid data.
|
||||
|
||||
An unknown schema version is not partly decoded. It remains unacknowledged on the original queue and
|
||||
creates one operator-visible `unsupported_version` failure. A newer daemon may read it later.
|
||||
|
||||
Invalid known-format data is copied byte-for-byte to durable queue
|
||||
`agent.<target>.inbox.quarantine`. A dedicated confirm-mode publisher confirms the persistent copy
|
||||
before the original is acknowledged. A failed quarantine handoff leaves the original unacknowledged.
|
||||
The raw body never enters logs.
|
||||
|
||||
Decode failure creates a redacted WARN, metric, `bridge_list` summary, and routed health incident.
|
||||
One bad entry never escapes the consumer callback and never stops later valid messages.
|
||||
|
||||
Safe downgrade is not supported. The previous build ignores `content_type` and would show typed JSON
|
||||
as ordinary reply text. If drained, it would acknowledge the message and lose typed meaning. Typed
|
||||
queues must be drained or preserved before an old jar runs.
|
||||
|
||||
The existing contract suite uses RabbitMQ. Production uses LavinMQ. The migration and lead-key
|
||||
ownership cases must run once against production LavinMQ before release, or the release must state
|
||||
that LavinMQ was not checked.
|
||||
|
||||
### 7.3 Member routing
|
||||
|
||||
A member incident first goes to the exact lead that owns its accepted delegation.
|
||||
`PrimaryRegistry` needs a no-fallback `delegatingLeadFor(memberTarget)` query. Health routing must not
|
||||
use the old singular-primary fallback when several leads exist.
|
||||
|
||||
Publish the incident under the affected member target. Trigger the existing bounded push route. The
|
||||
push waits until the owning lead is injectable, so it does not interrupt a live lead turn.
|
||||
|
||||
### 7.4 Peer lead routing
|
||||
|
||||
Figure 2 shows the route from incident to lead, peer, or person.
|
||||
|
||||
```mermaid
|
||||
flowchart TD
|
||||
Incident["Open incident"] --> Member{"Member incident?"}
|
||||
Member -->|"yes"| Known{"Exact delegation owner known?"}
|
||||
Known -->|"no"| Sink{"Human webhook enabled and healthy?"}
|
||||
Known -->|"yes"| Owner{"Owner lead healthy?"}
|
||||
Owner -->|"yes"| OwnerInbox["Publish to owner lead path"]
|
||||
Owner -->|"no"| PeerSet["Build healthy peer candidate set"]
|
||||
Member -->|"no, lead incident"| PeerSet
|
||||
PeerSet --> Peer{"Healthy peer exists?"}
|
||||
Peer -->|"yes"| Select["Choose fewest assigned incidents<br/>then stable name and terminal id"]
|
||||
Select --> PeerInbox["Publish to peer lead inbox<br/>and status-gated push"]
|
||||
Peer -->|"no"| Sink
|
||||
Sink -->|"yes"| Webhook["Send classified outbound incident"]
|
||||
Sink -->|"no"| Passive["Keep incident open<br/>show partial coverage on local surfaces"]
|
||||
```
|
||||
|
||||
*Figure 2. Routing keeps delegation ownership separate from temporary peer fallback.*
|
||||
|
||||
Peer candidates exclude the incident subject, failed owner, absent leads, raw-unknown leads, and
|
||||
leads with an open unhealthy state. A reachable `WORKING` peer may be selected; its push waits for an
|
||||
injectable window.
|
||||
|
||||
Choose the candidate with the fewest assigned foreign incidents. Break ties by stable lead name,
|
||||
then terminal id. Pin the recipient. Reassign only if that peer becomes unhealthy or retires. A
|
||||
routing generation marks a reassignment, and old pending assignments become superseded.
|
||||
|
||||
`bridge_list` lead rows show health, health age, assigned foreign incident count, and a bounded list
|
||||
of incident id, subject, state, severity, age, and routing generation. The top-level view also shows
|
||||
owner, recipient, and routing reason.
|
||||
|
||||
A peer incident is published under the recipient lead's inbox key, not the failed subject's key. Its
|
||||
status-gated nudge names the failed lead and gives the exact
|
||||
`bridge_poll(target="<recipient-terminal>")` call.
|
||||
|
||||
### 7.5 Lead inbox ownership
|
||||
|
||||
Add `LeadInboxRegistry`, driven by the same expected-lead supplier as `CallerResolver`.
|
||||
|
||||
It calls `replyInbox.own(leadTerminal)` at startup for configured leads, after successful discovery,
|
||||
after config adds a lead, and before publication. Ownership is not an authorization side effect.
|
||||
|
||||
A missing lead keeps its key owned. Release happens only after confirmed retirement, all incidents
|
||||
are reassigned or resolved, typed health messages move or ack, and the queue is empty. Own a
|
||||
replacement terminal before moving messages from the old key. Never release a non-empty in-memory
|
||||
lead key, because in-memory release clears local data.
|
||||
|
||||
### 7.6 Single-lead deployment
|
||||
|
||||
One lead and no peer is a normal mode, not an edge case.
|
||||
|
||||
An idle, reachable lead may receive the existing bounded nudge. An unreachable or stalled sole lead
|
||||
has no safe in-loop recovery. The bridge must not restart or replace it. A new lead would not have the
|
||||
failed lead's plan or context, and an uncertain relaunch could create two orchestrators.
|
||||
|
||||
With no webhook, only `bridge_list`, `/healthz`, metrics, WARN logs, and the incident journal remain.
|
||||
These are passive surfaces. They are not a human notification.
|
||||
|
||||
## 8. Lost-boundary reconciliation
|
||||
|
||||
This is the only M4 path that reconstructs a result. It must prefer a visible stall over a fabricated
|
||||
reply.
|
||||
|
||||
### 8.1 Why normal completion rules are not enough
|
||||
|
||||
Current `CompletionResolver.resolve` has two fail-open rules. It resolves when the delivery baseline
|
||||
is missing. It also resolves an empty completion when the pane read fails. Those choices are valid
|
||||
after a trusted `WORKING -> IDLE` boundary because the bridge knows the turn ran. They are unsafe
|
||||
when health only guesses that a boundary was lost.
|
||||
|
||||
M4 gives each accepted send an internal `TurnToken`. It ties target, exact waiter, session turn,
|
||||
delivery baseline, and task outcome together.
|
||||
|
||||
### 8.2 Delivery baseline
|
||||
|
||||
Capture the baseline immediately after prompt send and before the delivery future completes. Store:
|
||||
|
||||
```text
|
||||
TurnToken
|
||||
exact waiter identity
|
||||
capture time and pane source
|
||||
normalised assistant block clipped to MAX_SCRAPE_CHARS
|
||||
whether a supported assistant marker was recognised
|
||||
capture result: PRESENT | READ_FAILED | UNRECOGNISED
|
||||
```
|
||||
|
||||
A failed or missing baseline never authorises repair. A late baseline is not valid evidence. After a
|
||||
daemon restart, the old waiter, task, token, and baseline are gone, so the old turn cannot be
|
||||
repaired.
|
||||
|
||||
Automatic repair is enabled only for agent kinds with tested assistant-block fixtures. Current
|
||||
extraction is Claude Code-specific and falls back to arbitrary raw text without `⏺`. That raw fallback
|
||||
cannot authorise repair. OpenCode repair stays disabled until live pane fixtures exist.
|
||||
|
||||
### 8.3 Strict gates and resolver result
|
||||
|
||||
Figure 3 shows the repair gates. Any failed gate keeps the waiter unchanged.
|
||||
|
||||
The two raw snapshots must describe the same `TurnToken` and session turn. No `WORKING`,
|
||||
`BLOCKED`, `UNKNOWN`, missing-agent, reply, failure, or new-delivery observation may occur between
|
||||
them.
|
||||
|
||||
```mermaid
|
||||
flowchart TD
|
||||
Candidate["TURN_BOUNDARY_LOST candidate"] --> Stable{"Same TurnToken and BUSY turn<br/>across two raw IDLE or DONE snapshots?"}
|
||||
Stable -->|"no"| Resnapshot["Take a fresh snapshot"]
|
||||
Stable -->|"yes"| Waiter{"Exact captured waiter<br/>still open by identity?"}
|
||||
Waiter -->|"no"| Stale["STALE_TURN or ALREADY_RESOLVED"]
|
||||
Waiter -->|"yes"| Baseline{"Successful recognised<br/>delivery baseline exists?"}
|
||||
Baseline -->|"no"| Refuse["Refuse repair<br/>leave ticket pending"]
|
||||
Baseline -->|"yes"| Read{"Fresh pane read succeeds?"}
|
||||
Read -->|"no"| Refuse
|
||||
Read -->|"yes"| Output{"Recognised non-blank assistant block<br/>differs from clipped baseline?"}
|
||||
Output -->|"no"| Refuse
|
||||
Output -->|"yes"| Resolve["Shared CompletionResolver guard core<br/>resolves exact waiter"]
|
||||
Resolve -->|"won race"| Repaired["RECONCILED_COMPLETION<br/>same turn becomes DONE"]
|
||||
Resolve -->|"lost race"| Resnapshot
|
||||
```
|
||||
|
||||
*Figure 3. Repair needs stronger evidence than a normal observed turn boundary.*
|
||||
|
||||
Refactor the current resolver into one guard core with two policies:
|
||||
|
||||
```text
|
||||
resolveCaptured(target, inFlight, OBSERVED_BOUNDARY)
|
||||
resolveCaptured(target, inFlight, LOST_BOUNDARY_REPAIR)
|
||||
```
|
||||
|
||||
The health monitor calls only:
|
||||
|
||||
```text
|
||||
CompletionResolver.reconcileLostBoundary(target, expectedTurnToken)
|
||||
```
|
||||
|
||||
It returns `REPAIRED`, `ALREADY_RESOLVED`, `REFUSED_NO_CAPTURE`, `REFUSED_NO_BASELINE`,
|
||||
`REFUSED_UNREADABLE`, `REFUSED_UNCHANGED`, `REFUSED_AMBIGUOUS_OUTPUT`, `STALE_TURN`, or
|
||||
`RACE_LOST`.
|
||||
|
||||
Only `REPAIRED` and same-turn `ALREADY_RESOLVED` may move that turn from `BUSY` to `DONE`. A
|
||||
per-target reconciliation gate stops a queued second send from being accepted between waiter
|
||||
resolution and the FSM transition.
|
||||
|
||||
### 8.4 Lead-visible marker and refusal
|
||||
|
||||
A repaired result uses distinct `RECONCILED_COMPLETION` values in `Rendezvous`, `MessageService`,
|
||||
task poll source, and metrics. The lead sees:
|
||||
|
||||
```text
|
||||
[repaired completion - bridged detected a lost turn boundary. The member did not call
|
||||
bridge_reply; pane-derived text follows and may be partial]
|
||||
```
|
||||
|
||||
Clipped text also keeps the existing clipped-tail marker.
|
||||
|
||||
A refused repair leaves `TURN_BOUNDARY_LOST` open and the ticket pending. The report states that no
|
||||
reply was reconstructed and no task was replayed. `UNCHANGED`, `UNREADABLE`, and
|
||||
`AMBIGUOUS_OUTPUT` get at most one delayed retry for the same token. Missing capture or baseline gets
|
||||
no retry. After two refused scrapes, automatic repair stops for that token.
|
||||
|
||||
### 8.5 Target-wide teardown invariant
|
||||
|
||||
CB-568 owns the multi-ticket cancellation mechanism. M4 routes every terminal cause through that one
|
||||
idempotent operation and checks this independent invariant after teardown:
|
||||
|
||||
- no injector entry exists for the target;
|
||||
- no accepted turn or completion record exists;
|
||||
- no rendezvous waiter or ask exists;
|
||||
- every async task is terminal or was already terminal;
|
||||
- no thread waiting for the target send lock can later accept it;
|
||||
- new sends fail immediately;
|
||||
- each old task has one terminal outcome and one metric count.
|
||||
|
||||
A violation becomes `DELEGATION_ORPHANED`. The monitor may call the same idempotent target-wide
|
||||
failure operation once. It never recreates the task.
|
||||
|
||||
## 9. Capacity and utilisation
|
||||
|
||||
Capacity is a view, not a health state.
|
||||
|
||||
`bridge_list` adds one block per profile:
|
||||
|
||||
```text
|
||||
profile, maxLoad, live, free, reclaimable
|
||||
```
|
||||
|
||||
For an unlimited profile, `maxLoad` and `free` are null. `free` is
|
||||
`max(0, maxLoad - live)` for a capped profile.
|
||||
|
||||
The view must use the exact live-count function used by placement. A second calculation could show a
|
||||
free slot that placement then refuses. Member rows add `idleForSeconds` only when state is `READY` or
|
||||
`DONE`, no accepted turn exists, and the inbox is empty. `reclaimable` means only that the member
|
||||
holds capacity without open bridge work.
|
||||
|
||||
The existing idle-lead nudge gains a bounded capacity summary. It lists per-profile live, cap, free,
|
||||
and reclaimable counts, plus at most three long-idle members. Capacity does not make
|
||||
`FleetState.hasPending()` true. A changed capacity fingerprint may re-arm one capped heartbeat
|
||||
sequence. The fingerprint excludes changing idle durations, so a static idle fleet cannot reset the
|
||||
cap forever. Reply-push stand-down remains first.
|
||||
|
||||
The bridge must never:
|
||||
|
||||
- spawn a member because a slot is free;
|
||||
- generate a task or acceptance criteria;
|
||||
- move queued work to another member or profile;
|
||||
- treat a free slot or idle member as an incident;
|
||||
- stop an idle member only to improve utilisation.
|
||||
|
||||
The bridge knows capacity facts but has no work list. Only the lead has the plan, task context,
|
||||
side-effect history, and acceptance criteria.
|
||||
|
||||
Capacity calculation is in memory and adds no pane reads. Work-product checks run on a terminal
|
||||
session edge, not every fleet tick.
|
||||
|
||||
This capacity design adds no automatic stop. The accepted `NEVER_READY` cleanup can still stop a
|
||||
very slow startup after the existing grace, which is a known risk. Free capacity and long idle time
|
||||
never trigger that path.
|
||||
|
||||
## 10. Human escalation and notification
|
||||
|
||||
### 10.1 Escalation rule
|
||||
|
||||
Notify a person only when no healthy lead can act:
|
||||
|
||||
- `CONTROL_LINK_DOWN` survives grace;
|
||||
- a lead is unhealthy and no healthy peer can receive the incident;
|
||||
- a member incident has no known owning lead;
|
||||
- the only owning lead becomes unreachable, unresponsive, or stalled;
|
||||
- incident publication or routing itself fails.
|
||||
|
||||
Do not page a person for a member fault while a healthy owning lead exists. An uncollected member
|
||||
incident feeds lead-health evidence. If the lead then becomes unhealthy, peer or human routing starts.
|
||||
|
||||
### 10.2 Detection and notification switches
|
||||
|
||||
`health.enabled` controls detection and bridge-local reporting. It does not require a webhook.
|
||||
|
||||
`health.notifications.mode` is `disabled` or `webhook`. Disabled is valid and is the default.
|
||||
Webhook mode requires a resolved environment variable. Turning notification off stops outbound
|
||||
attempts but keeps incidents. Turning it back on resumes still-open human incidents.
|
||||
|
||||
Without a sink, `bridge_list.healthCoverage` states that human escalation is unavailable. `/healthz`
|
||||
keeps its existing HTTP liveness result and adds a nested `fleetHealth.status=partial` component.
|
||||
Metrics and one startup or reload WARN expose the same limit.
|
||||
|
||||
### 10.3 Incident and delivery deduplication
|
||||
|
||||
One open incident uses this key:
|
||||
|
||||
```text
|
||||
(scope, subjectStableId, state, causeFingerprint)
|
||||
```
|
||||
|
||||
The cause fingerprint includes stable error codes, dependency names, signature ids, or invariant
|
||||
names. It excludes times, ages, retry counts, pane text, and changing digests. A later recurrence
|
||||
after resolution gets a new generation and incident id.
|
||||
|
||||
Each outbound event uses:
|
||||
|
||||
```text
|
||||
Idempotency-Key = hash(incidentId, eventType, eventRevision)
|
||||
```
|
||||
|
||||
Event types are `open`, `severity_changed`, `reminder`, and `resolved`. Transport retries keep the
|
||||
same key.
|
||||
|
||||
An atomic owner-only journal beside the active config stores open incidents, routing, delivered
|
||||
revisions, retry state, and resolution state. It stores no pane or task content. Journal failure does
|
||||
not stop detection, but notification coverage becomes degraded.
|
||||
|
||||
### 10.4 Retry, reminder, and resolve
|
||||
|
||||
Send the first event immediately. Retry network errors, timeouts, HTTP 408, HTTP 429, and HTTP 5xx
|
||||
with full-jitter exponential backoff:
|
||||
|
||||
```text
|
||||
base: 5 seconds
|
||||
factor: 3
|
||||
maximum delay: 15 minutes
|
||||
one outstanding attempt per event
|
||||
```
|
||||
|
||||
Respect `Retry-After` up to 15 minutes. Other HTTP 4xx responses are permanent for that event until
|
||||
config changes or a person requests replay.
|
||||
|
||||
Transport retry is not an incident reminder. `humanRepeatSeconds` creates a new reminder revision
|
||||
for an unresolved critical incident after the last successful human event. Disabled mode does not
|
||||
build an unbounded reminder queue.
|
||||
|
||||
Send `resolved` only if at least one human event for that incident was delivered. If an incident
|
||||
resolves before its first successful delivery, cancel the pending open event and record local
|
||||
resolution.
|
||||
|
||||
### 10.5 Outbound payload boundary
|
||||
|
||||
An outbound payload may contain incident id and event type, severity, state, scope, stable bridge
|
||||
ids, role or profile, times, duration, structured evidence type and counts, recovery attempted,
|
||||
routing reason, coverage, and safe tool calls.
|
||||
|
||||
It must never contain:
|
||||
|
||||
- raw pane text, pane excerpts, or pane digests;
|
||||
- task briefs, prompts, or member reply content;
|
||||
- source files, diffs, or worktree file content;
|
||||
- worktree paths;
|
||||
- environment values, tokens, credentials, headers, or webhook URL;
|
||||
- raw exception messages or stack traces;
|
||||
- arbitrary model output.
|
||||
|
||||
The sink response body is ignored. A webhook cannot direct recovery. n8n remains outbound-only.
|
||||
|
||||
### 10.6 Metrics
|
||||
|
||||
M4 adds bounded-label series:
|
||||
|
||||
```text
|
||||
bridged_health_incidents{scope,state,severity}
|
||||
bridged_health_incidents_total{event}
|
||||
bridged_health_notifications_total{event,outcome}
|
||||
bridged_health_notification_queue_depth
|
||||
bridged_health_notification_last_success_seconds
|
||||
bridged_health_notification_capability{mode,status}
|
||||
bridged_lead_health{lead,state}
|
||||
bridged_lead_assigned_incidents{lead}
|
||||
```
|
||||
|
||||
Metric labels never include terminal ids, incident ids, URLs, or error text.
|
||||
|
||||
## 11. Configuration
|
||||
|
||||
The optional `health:` block is absent or disabled by default. The dormant monitor scheduler does no
|
||||
herdr or pane work while disabled. Every listed key is hot because the monitor reads `ConfigRef` on
|
||||
each tick or notification.
|
||||
|
||||
| Key | Class | Default and hard bound | Purpose |
|
||||
|---|---|---|---|
|
||||
| `health.enabled` | Hot | `false` | Enable detection and bridge-local reporting. |
|
||||
| `health.snapshotIntervalSeconds` | Hot | default 30, minimum 15 | Whole-fleet comparison cadence. |
|
||||
| `health.workingSuspectAfterSeconds` | Hot | default 600, minimum 300 | Age before working-pane probes. |
|
||||
| `health.paneProbeIntervalSeconds` | Hot | default 60, minimum 60 | Per-target pane cooldown. |
|
||||
| `health.leadUnresponsiveAfterSeconds` | Hot | default 300, minimum 120 | Delay after exhausted actionable nudges before lead fault. |
|
||||
| `health.humanRepeatSeconds` | Hot | default 3600, minimum 900 | Minimum repeat period for one open human incident. |
|
||||
| `health.capacityLongIdleAfterSeconds` | Hot | default 900, minimum 300 | Long-idle threshold for capacity summaries. |
|
||||
| `health.includePaneExcerpt` | Hot | `false` | Allow a clipped excerpt in local lead reports only. Human payloads still exclude it. |
|
||||
| `health.notifications.mode` | Hot | `disabled` | Select `disabled` or `webhook`. |
|
||||
| `health.notifications.webhookUrlEnv` | Hot | required in webhook mode | Name of the environment variable that holds the sink URL. |
|
||||
| `health.notifications.requestTimeoutMs` | Hot | default 10000, range 1000-30000 | Whole webhook request limit. |
|
||||
|
||||
Two consecutive snapshots are compiled floors for lost boundary, lead disappearance, and control
|
||||
link failure. The two-pane-reads-per-tick limit is also compiled and cannot be weakened by config.
|
||||
|
||||
## 12. Delivery units and acceptance
|
||||
|
||||
### Unit 1 - Evidence model and fleet snapshot
|
||||
|
||||
Scope: health state model, fleet join, clocks, evidence retention, and pane budget.
|
||||
|
||||
Acceptance criteria:
|
||||
|
||||
1. One `agent.list` call covers one enabled fleet tick.
|
||||
2. Pure decision tests cover every state and every evidence limit in Section 4.
|
||||
3. `BUSY` plus stable raw `DONE` opens `TURN_BOUNDARY_LOST` after two snapshots.
|
||||
4. Healthy fleet snapshots perform zero pane reads.
|
||||
5. Pane cooldown, two-read fleet budget, and fair rotation cannot be disabled by config.
|
||||
6. Logs are outputs only; no log parsing exists.
|
||||
7. Fleet snapshots expose the same profile live-count calculation that placement uses.
|
||||
8. Capacity rows report cap, live, free, and reclaimable values without opening incidents.
|
||||
|
||||
### Unit 2 - Lost boundary and task reconciliation
|
||||
|
||||
Scope: accepted-turn identity, guarded repair, target-wide teardown, release causes, and preserved
|
||||
worktree discovery.
|
||||
|
||||
Acceptance criteria:
|
||||
|
||||
1. Every accepted send receives a stable `TurnToken` tied to target, exact waiter, and delivery
|
||||
baseline.
|
||||
|
||||
**Corrected during implementation (2026-08-15).** This criterion first also required the session
|
||||
turn number and the task outcome. That is not implementable at this layer, and the implementer
|
||||
refused it three times rather than fabricate a value — correctly. The reason is an ordering fact
|
||||
that is invisible from any single class: `MessageService` owns acceptance and holds the waiter and
|
||||
the async `Task`, but it learns nothing about delivery, because the delivery event goes to
|
||||
`CompletionResolver` through `TurnListener.onDelivered`. And `CompletionResolver.onDelivered` runs
|
||||
*before* `SessionManager.onDelivered`, so the session turn number does not exist yet at the only
|
||||
point where the token could capture it.
|
||||
|
||||
Two ways out were rejected. A shared registry keyed by target reintroduces exactly the "whichever
|
||||
send happens to be waiting" ambiguity the token exists to remove — the same weak claim
|
||||
`Rendezvous.currentWaiter` warns about. Injecting a turn counter into `MessageService` adds a
|
||||
required cross-layer dependency to populate a field that nothing in this slice reads, which is
|
||||
speculative coupling across a boundary already shown to be fragile.
|
||||
|
||||
So the token identifies the **accepted send**, and `SessionManager` keeps verifying its own
|
||||
delivery separately. Repair (criterion 2) does need the session turn; binding it means resolving
|
||||
that acceptance-versus-delivery ordering first, and that work belongs to the repair unit, not
|
||||
here. The token record carries a comment saying the field is deliberately absent.
|
||||
2. Repair requires the same `BUSY` token, two raw `IDLE` or `DONE` snapshots, no conflicting
|
||||
observation, exact open waiter, successful baseline, and new recognised assistant output.
|
||||
3. Missing, failed, late, or post-restart baseline never authorises repair.
|
||||
4. Repair is enabled only for agent kinds with tested assistant-block extraction. Raw-text fallback
|
||||
without a recognised marker refuses repair.
|
||||
5. Normal completion and repair use one resolver guard core. Waiter, scrape, clipping, unchanged, and
|
||||
exact-turn guards are not duplicated.
|
||||
6. `reconcileLostBoundary` returns every typed result named in Section 8.3.
|
||||
7. Only `REPAIRED` and same-turn `ALREADY_RESOLVED` may move the same turn to `DONE`.
|
||||
8. A per-target reconciliation gate blocks a queued second send during repair and FSM update.
|
||||
9. Repaired completion has distinct rendezvous kind, message outcome, poll source, lead marker, and
|
||||
metric. Clipping keeps its extra marker.
|
||||
10. Unchanged, unreadable, or ambiguous evidence gets at most one delayed retry. Missing capture or
|
||||
baseline gets none.
|
||||
11. Refusal leaves the ticket pending and tells the lead that no result was rebuilt or replayed.
|
||||
12. Release, gone, never-ready, and abnormal stop use CB-568's one idempotent target-wide failure
|
||||
operation.
|
||||
13. The post-teardown invariant in Section 8.5 is tested independently of CB-568 internals.
|
||||
14. A violated teardown invariant creates `DELEGATION_ORPHANED` and retries only the idempotent
|
||||
failure operation.
|
||||
15. `SPAWN_ROLLBACK` and normal `COMPLETED` remove worktrees. Abnormal and shutdown causes preserve
|
||||
them.
|
||||
16. Explicit stop is state-aware. Any pending task or non-terminal state preserves the worktree.
|
||||
17. Atomic preserved-worktree manifests reload after restart and appear in lead-only
|
||||
`bridge_list.preservedWorktrees`.
|
||||
18. Manifest failure preserves the worktree and opens an operator-visible health failure.
|
||||
19. Provision records the base commit. Terminal, long-idle worktrees report
|
||||
`WORK_PRODUCT_AT_RISK` only under the evidence in Section 4.2 and never auto-delete work.
|
||||
20. No path replays a delivered task, rebuilds its brief, or retargets it, even when prompt text is
|
||||
available.
|
||||
21. Tests cover both real traces, all repair refusals, clipping, explicit-reply and next-turn races,
|
||||
restart without capture, concurrent send and release, and preserved discovery after restart.
|
||||
|
||||
### Unit 3 - Typed inbox and member routing
|
||||
|
||||
Scope: semantic record, AMQP migration, both adapters, member routing, polling, and member health in
|
||||
`bridge_list`.
|
||||
|
||||
Acceptance criteria:
|
||||
|
||||
1. AMQP selects legacy or typed decoding only from `content_type`; it never sniffs the body.
|
||||
2. Persistent `text/plain` from the old build becomes `kind=reply` with exact UTF-8 content,
|
||||
including content beginning with `{`.
|
||||
3. New entries use the vendor media type, `schemaVersion: 1`, UTF-8, persistent delivery, and AMQP
|
||||
message ids.
|
||||
4. Version 1 ignores unknown optional fields but rejects missing fields and identity mismatch.
|
||||
5. Unknown versions are not decoded or acked. They remain on the original queue and create one
|
||||
deduplicated failure.
|
||||
6. Invalid known data never escapes the callback, appears as a reply, or blocks later valid messages.
|
||||
7. Invalid data reaches durable per-target quarantine before original ack. Failed handoff leaves the
|
||||
original unacked.
|
||||
8. Decode failures create redacted WARN, metric, `bridge_list` summary, and routed incident without
|
||||
raw content.
|
||||
9. Both adapters pass one semantic contract for fields, FIFO, dedup, ownership, ack, and release.
|
||||
10. Lead keys require explicit ownership. Publication never claims a queue.
|
||||
11. Unit codec tests cover legacy `{`, Unicode, malformed UTF-8, typed round trip, additive fields,
|
||||
malformed JSON, missing fields, identity mismatch, media type, version, and dedup.
|
||||
12. A live broker contract writes old wire data and reads it with the new adapter after reconnect.
|
||||
13. Live contract tests cover mixed entries, quarantine confirm-before-ack, unsupported redelivery,
|
||||
later progress past poison, property persistence, lead ownership, and ack removal.
|
||||
14. Safe downgrade is documented as unsupported.
|
||||
15. RabbitMQ contract tests pass with `mvn test -Pcontract`. The same cases run once on production
|
||||
LavinMQ, or the release states that LavinMQ was not checked.
|
||||
16. Member incidents route to the exact delegating lead and never resolve a task rendezvous.
|
||||
17. `bridge_list` shows compact member health and capacity without pane content. Member
|
||||
`idleForSeconds` is present only when no accepted turn or inbox item exists.
|
||||
|
||||
### Unit 4 - Lead health and peer routing
|
||||
|
||||
Scope: lead evidence, exact ownership, peer selection, explicit-recipient push, and lead inbox
|
||||
lifecycle.
|
||||
|
||||
Acceptance criteria:
|
||||
|
||||
1. Lead identity uses the `CallerResolver` supplier. Liveness uses successful current agent data.
|
||||
2. Two successful-list absences with healthy ping become `LEAD_UNREACHABLE`; global link failure does
|
||||
not.
|
||||
3. Raw `WORKING`, raw `UNKNOWN`, first-seen time, failures, last success, and error class persist
|
||||
across ticks.
|
||||
4. Heartbeat and push publish status and nudge outcomes before safe no-injection decisions.
|
||||
5. Dynamic lead identity survives a two-successful-snapshot retirement grace.
|
||||
6. Member incidents first use exact delegation ownership with no singular-primary fallback.
|
||||
7. Peer selection follows the exclusions, load rule, and stable tie break in Section 7.4.
|
||||
8. A selected working peer is not interrupted. Its push waits for an injectable window.
|
||||
9. Recipient assignment stays pinned. Reassignment increments generation and supersedes old pending
|
||||
assignment.
|
||||
10. `bridge_list` shows bounded foreign assignments, recipient, reason, and generation without pane
|
||||
content.
|
||||
11. `LeadInboxRegistry` owns configured and discovered lead keys before publication.
|
||||
12. Missing leads keep ownership. Retirement needs an empty queue and handled incidents.
|
||||
13. Replacement owns the new key before messages move. Non-empty in-memory keys are not released.
|
||||
14. Tests cover dead versus busy, unknown, global failure, stale scan cache, disappearance, one peer,
|
||||
several peers, reassignment, and no peer.
|
||||
15. Adapter tests cover lead ownership, restart re-ownership, retirement, and terminal replacement.
|
||||
LavinMQ is checked or named as unchecked.
|
||||
16. A sole unreachable or stalled lead is never restarted or replaced. Without a sink, only passive
|
||||
evidence remains and every coverage surface says so.
|
||||
|
||||
### Unit 5 - Human sink, hot config, metrics, and operator coverage
|
||||
|
||||
Scope: generic webhook, config split, incident journal, retry, resolve, metrics, example config, and
|
||||
operator documentation.
|
||||
|
||||
Acceptance criteria:
|
||||
|
||||
1. `health.enabled` works without a human sink.
|
||||
2. Notification mode is hot, defaults to disabled, and supports disabled or webhook.
|
||||
3. Webhook mode requires a resolved environment value. Bad notification config does not disable an
|
||||
already valid detector.
|
||||
4. Mode changes keep open incidents. Re-enable resumes eligible incidents.
|
||||
5. `bridge_list`, `/healthz`, metrics, and one WARN show partial coverage without a sink. HTTP
|
||||
liveness behavior stays unchanged.
|
||||
6. One-lead, no-sink coverage states that lead failure has no active notification or recovery.
|
||||
7. Incident and outbound dedupe use the stable keys in Section 10.3.
|
||||
8. The owner-only local journal survives restart and contains no pane or task content.
|
||||
9. Journal failure keeps detection running but marks notification coverage degraded.
|
||||
10. Retry tests cover network failure, timeout, 408, 429, `Retry-After`, 5xx, permanent 4xx, jitter,
|
||||
delay cap, config re-arm, and one outstanding attempt.
|
||||
11. Reminders and transport retries remain separate. Disabled mode does not build an unbounded queue.
|
||||
12. Resolve sends only after an earlier human event succeeded. Resolve-before-delivery cancels stale
|
||||
open delivery.
|
||||
13. Metrics use bounded labels and exclude ids, URLs, and error text.
|
||||
14. Payload tests reject every content type forbidden in Section 10.5.
|
||||
15. Webhook response bodies are ignored and cannot direct recovery.
|
||||
16. Tests cover disabled mode, one lead without sink, open/update/reminder/resolve, restart, dedup,
|
||||
reassignment, disable/re-enable, and sink failure while local health continues.
|
||||
17. `bridged.example.yaml` documents all hot keys and compiled floors.
|
||||
18. The operator Features wiki is updated separately. The portable `CLAUDE.md` block is checked and
|
||||
changed only if shipped tool or inbox semantics make it untrue.
|
||||
19. `mvn clean install` passes.
|
||||
|
||||
## 13. Not checked and release gates
|
||||
|
||||
These limits are part of the design, not optional follow-up notes.
|
||||
|
||||
- **OpenCode pane status and assistant markers were not checked.** OpenCode lost-boundary repair is
|
||||
disabled until live fixtures exist.
|
||||
- **Permission-prompt status was not checked** for Claude Code or OpenCode. `BLOCKED` remains
|
||||
ambiguous and has no automatic action.
|
||||
- **`recent_unwrapped` stability was not checked** across all supported agent kinds. If normalisation
|
||||
is not stable, `STALL_SUSPECTED` must say its evidence is weaker.
|
||||
- **The real `BUSY + DONE` trace was not replayed against live herdr.** The design uses the observed
|
||||
production trace and current poller behavior.
|
||||
- **CB-568 was not present when Unit 2 was designed.** Unit 2 must inspect the landed API and keep its
|
||||
independent teardown invariant.
|
||||
- **Production LavinMQ was not checked.** Existing durable-inbox contracts use RabbitMQ. Migration,
|
||||
quarantine, redelivery, lead ownership, and reassignment must run on LavinMQ before release or be
|
||||
recorded as unchecked.
|
||||
- **Live multi-lead routing was not checked.** Peer choice and reassignment are design rules backed by
|
||||
fake-clock and adapter tests until a live exercise runs.
|
||||
- **A live sole-lead failure with a webhook was not checked.** The no-peer path is a design result,
|
||||
not a tested recovery.
|
||||
- **No n8n, Slack, PagerDuty, or other receiver was checked.** The webhook remains generic and
|
||||
outbound-only.
|
||||
- **Deployment supervisor behavior for nested `/healthz` fields was not checked.** HTTP liveness
|
||||
status stays unchanged to reduce this risk.
|
||||
- **Incident-journal crash behavior was not checked** because the journal does not exist yet. Unit 5
|
||||
must test atomic replacement and restart recovery.
|
||||
- **Worktree merge state cannot be checked reliably** without forge or explicit collection evidence.
|
||||
`WORK_PRODUCT_AT_RISK` stays a warning.
|
||||
|
||||
## 14. Locked exclusions
|
||||
|
||||
M4 does not expose `agent.read` as a bridge tool. It does not add a workflow engine, inbound n8n
|
||||
authority, automatic task assignment, task replay, automatic lead replacement, or automatic member
|
||||
spawn for free capacity.
|
||||
|
||||
The bridge remains a message bus with evidence and bounded mechanical repair. The lead remains the
|
||||
place where judgement and work planning happen.
|
||||
+3
-3
@@ -14,7 +14,7 @@
|
||||
"url": "https://ct7.ltms.dev/mcp",
|
||||
"enabled": true,
|
||||
"headers": {
|
||||
"Authorization": "Bearer {file:.secrets/context7-token}"
|
||||
"Authorization": "Bearer {env:CONTEXT7_TOKEN}"
|
||||
}
|
||||
},
|
||||
"gitea": {
|
||||
@@ -26,8 +26,8 @@
|
||||
],
|
||||
"enabled": true,
|
||||
"environment": {
|
||||
"GITEA_ACCESS_TOKEN": "{file:.secrets/gitea-token}",
|
||||
"GITEA_HOST": "{file:.secrets/gitea-host}"
|
||||
"GITEA_ACCESS_TOKEN": "{env:GITEA_ACCESS_TOKEN}",
|
||||
"GITEA_HOST": "{env:GITEA_HOST}"
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
+1
-1
Submodule wiki updated: 8c63db5da6...7c50cce52e
Reference in New Issue
Block a user